Models & routing
Which models Primitive uses, why, where OpenRouter fits, and the privacy rules that sit above model preference.
Primitive routes by task fit, privacy, live capacity, cost and reviewer independence. Subscription capacity is used before API for OpenAI-family work. OpenRouter is mainly the cheap multi-family pool and non-OpenAI review/build lane.
We do not ask the most expensive AI to do every job. Cheap workers build the well-described pieces, a fast inspector checks them, and stronger models step in only when the problem is hard, risky or disputed. Private material only goes through approved private routes.
The model list is evidence-driven rather than a permanent leaderboard. ROUTING.md is the policy; models.yaml records observed strengths, failures and traps. The same model can be a good builder and a bad reviewer, and a cheap model can be the correct choice when the work packet is precise. Review independence is based on model family, not account name.
Cheap first, frontier on failure
Models we actually have a reason to call
This is a documentation snapshot of Control's current routing evidence, not a universal ranking. The canonical source is ROUTING.md + models.yaml and should be updated when ledger evidence changes.
| OpenRouter slug | Family | Best current use | Important note |
|---|---|---|---|
| z-ai/glm-5.3-flash | Zhipu | Default cheap-pair builder/reviewer | Strong bounded writes and one-shot reviews. Prefer one-shot review when the long-running harness stalls. |
| deepseek/deepseek-v4-flash-0731 | DeepSeek | Default cheap-pair builder/reader | Cheap, useful intake/review pair member. Current notes prefer pinned low-cost endpoints where appropriate. |
| deepseek/deepseek-v4.1-flash | DeepSeek | Cheap stronger reviewer | Good repo-aware reviewer. For large diffs use enough answer budget; prior small answer budgets produced empty output. |
| openai/gpt-6-luna | OpenAI | Audit / tidy / synthesis | Fast, reliable cheap reader. Subscription capacity is preferred before API for OpenAI-family work. |
| openai/gpt-oss-20b | OpenAI | Low-risk triple reviewer | Very cheap but ACCEPTs can be thin and HIGHs can be false. Verify HIGHs by execution/code trace. |
| openai/gpt-oss-120b | OpenAI | Whole-diff checklist | Useful as a final checklist over small/medium diffs; Control notes say not to use it for triage/intake synthesis. |
| deepseek/deepseek-v4-pro | DeepSeek | Strong independent review | Useful for restricted/high-risk review. Cap reasoning so it cannot consume the whole response budget. |
| moonshotai/kimi-k3 | Moonshot | Strong independent review / alternative | Included in restricted-review allowlist; useful as a non-OpenAI/non-Anthropic family. |
| inclusionai/ling-3.0-flash | InclusionAI | Cheap floor / reader | Very low-cost candidate for bounded reading/review tasks. |
| xiaomi/mimo-v2.6-flash | Xiaomi | Alternate cheap reviewer | Has found real medium issues, but a Codex-harness build also hit retry/patch problems. Not default pair until retested. |
| google/gemini-3.8-flash | Fourth-family fast review/write | Useful independent family; considerably more expensive than the cheap pair in the recorded Control snapshot. | |
| qwen/qwen3.8-flash | Qwen | Benchmark / alternate | Candidate in the benchmark pool; do not promote above observed routes without ledger evidence. |
Three data classes sit above model preference
| Class | What may enter | OpenRouter rule |
|---|---|---|
| C3 — Restricted | Auth, permits, gateway, kernel, migrations, production config | ZDR enforced; narrow reviewer allowlist; low spend cap |
| C2 — Internal | Primitive code, packets, reviews, theory | ZDR enforced; paid ZDR-capable model allowlist |
| C1 — Open | Public/synthetic material only | Training/logging may be allowed; never use C1 for private material |
Family matters more than account
Two Claude accounts are still Anthropic. Two Codex accounts are still OpenAI. When independent review is required, the reviewer's family must differ from every author's family.
Jev is not a reviewer
Control tracks typesafe/jev-1.13 as a cheap typed decision model for prepared closed-choice questions. It does not produce reasoning text and cannot replace an explanatory reviewer. Any error or low-confidence result should escalate rather than silently decide.