Execution resources

Models & routing

Which models Primitive uses, why, where OpenRouter fits, and the privacy rules that sit above model preference.

TL;DR

Primitive routes by task fit, privacy, live capacity, cost and reviewer independence. Subscription capacity is used before API for OpenAI-family work. OpenRouter is mainly the cheap multi-family pool and non-OpenAI review/build lane.

ELI5

We do not ask the most expensive AI to do every job. Cheap workers build the well-described pieces, a fast inspector checks them, and stronger models step in only when the problem is hard, risky or disputed. Private material only goes through approved private routes.

Synthesis

The model list is evidence-driven rather than a permanent leaderboard. ROUTING.md is the policy; models.yaml records observed strengths, failures and traps. The same model can be a good builder and a bad reviewer, and a cheap model can be the correct choice when the work packet is precise. Review independence is based on model family, not account name.

Default ladder

Cheap first, frontier on failure

Current routing ladder
ArchitectOpus 5.5 for architecture/authority
Build pairGLM 5.3 Flash + DeepSeek V4 Flash 0731
AuditLuna 6
EscalateSol 6 or Opus when failure/disagreement survives
AstraHard novel maths only, owner approval
Subscription before API: for OpenAI-family work, Control currently routes to the Codex subscription with the most headroom before spending OpenRouter/API credit.
OpenRouter shortlist

Models we actually have a reason to call

This is a documentation snapshot of Control's current routing evidence, not a universal ranking. The canonical source is ROUTING.md + models.yaml and should be updated when ledger evidence changes.

OpenRouter slugFamilyBest current useImportant note
z-ai/glm-5.3-flashZhipuDefault cheap-pair builder/reviewerStrong bounded writes and one-shot reviews. Prefer one-shot review when the long-running harness stalls.
deepseek/deepseek-v4-flash-0731DeepSeekDefault cheap-pair builder/readerCheap, useful intake/review pair member. Current notes prefer pinned low-cost endpoints where appropriate.
deepseek/deepseek-v4.1-flashDeepSeekCheap stronger reviewerGood repo-aware reviewer. For large diffs use enough answer budget; prior small answer budgets produced empty output.
openai/gpt-6-lunaOpenAIAudit / tidy / synthesisFast, reliable cheap reader. Subscription capacity is preferred before API for OpenAI-family work.
openai/gpt-oss-20bOpenAILow-risk triple reviewerVery cheap but ACCEPTs can be thin and HIGHs can be false. Verify HIGHs by execution/code trace.
openai/gpt-oss-120bOpenAIWhole-diff checklistUseful as a final checklist over small/medium diffs; Control notes say not to use it for triage/intake synthesis.
deepseek/deepseek-v4-proDeepSeekStrong independent reviewUseful for restricted/high-risk review. Cap reasoning so it cannot consume the whole response budget.
moonshotai/kimi-k3MoonshotStrong independent review / alternativeIncluded in restricted-review allowlist; useful as a non-OpenAI/non-Anthropic family.
inclusionai/ling-3.0-flashInclusionAICheap floor / readerVery low-cost candidate for bounded reading/review tasks.
xiaomi/mimo-v2.6-flashXiaomiAlternate cheap reviewerHas found real medium issues, but a Codex-harness build also hit retry/patch problems. Not default pair until retested.
google/gemini-3.8-flashGoogleFourth-family fast review/writeUseful independent family; considerably more expensive than the cheap pair in the recorded Control snapshot.
qwen/qwen3.8-flashQwenBenchmark / alternateCandidate in the benchmark pool; do not promote above observed routes without ledger evidence.
Privacy

Three data classes sit above model preference

ClassWhat may enterOpenRouter rule
C3 — RestrictedAuth, permits, gateway, kernel, migrations, production configZDR enforced; narrow reviewer allowlist; low spend cap
C2 — InternalPrimitive code, packets, reviews, theoryZDR enforced; paid ZDR-capable model allowlist
C1 — OpenPublic/synthetic material onlyTraining/logging may be allowed; never use C1 for private material
Secrets are a hard deny. A model being capable or cheap never overrides the data-class boundary.
Review independence

Family matters more than account

Two Claude accounts are still Anthropic. Two Codex accounts are still OpenAI. When independent review is required, the reviewer's family must differ from every author's family.

Example: Anthropic-authored change
Anthropic authorOpus / Sonnet / Fable
OpenAI reviewerLuna / Sol
DeepSeek / Google / xAIIndependent second family when required
Non-chat decision model

Jev is not a reviewer

Control tracks typesafe/jev-1.13 as a cheap typed decision model for prepared closed-choice questions. It does not produce reasoning text and cannot replace an explanatory reviewer. Any error or low-confidence result should escalate rather than silently decide.

Reference / evidence

Canonical sources