# Choosing a Claude Code model and effort level Claude Code exposes three choices this repo has an opinion on: which model a session itself runs as, what model a spawned subagent gets, and which `/code-review` effort level to pick. None of them are enforced anywhere - the gates in [instructions/gates.md](../instructions/gates.md) are code precisely because a model cannot be talked out of them ([why-gates-are-code.md](why-gates-are-code.md) makes that argument for gates; this page applies the same axis to who is holding the keyboard). What follows is a reference for making that choice well, not a rule anything checks. The axis worth tracking is not how important a task feels, but **what would catch a mistake in it**. Work behind `pytest`, `docs verify`, `instructions verify` or CI surfaces a bad call within one more round. Work behind nothing but a session reading prose does not surface at all - it ships, and stays until someone happens to notice. That asymmetry, not task size, is what the phase guide below is built on. ## Contents - [Phase guide](#phase-guide) - [Model per session, not per phase](#model-per-session-not-per-phase) - [Subagent models](#subagent-models) - [`/code-review` effort](#code-review-effort) - [When it's unclear](#when-its-unclear) - [Scope](#scope) ## Phase guide This repo's own stack-development work splits the axis into three phases, one skill each - `stack-dev` (design), `stack-build` (build) and `stack-close` (closing) - handed over through states in the issue tracker rather than inside one session: | Phase / task | What would catch a mistake | Suggested model | Effort | |---|---|---|---| | `wiki-status`, simple `wiki-query` lookups | the answer is re-checkable against the corpus | Sonnet | default | | `wiki-lint` | `lint` itself is the check | Sonnet | default | | `wiki-ingest`, `wiki-manage`, judgment-heavy `wiki-query` | `lint` and `docs verify`, partly | Sonnet | high | | Stack dev: design, the version part, a boundary-crossing judgment | nothing mechanical | Opus | high | | Stack dev: code, tests, mechanical doc sync, waiting for CI | `pytest`, `docs verify`, `instructions verify`, CI | Opus (open - see below) | high; medium for a small change | | Stack dev: closing an issue, `docs/` staleness, changelog prose | nothing, by construction | Opus | high | The two unchecked rows are short - minutes, not hours - so keeping them on the strongest model costs little and protects the only work that fails silently. The checked middle row is where the tokens are, which makes it the tempting one to run cheaper. It is also the row with the least settled answer: a build phase of this stack reads a lot of the tree, and a smaller context window does not hold it - it runs into compaction, which costs more than the cheaper model saves. Sonnet at high effort remains possible, but only as a session of its own (below). Which default the middle row should have is left to the record: every closed work package in this repo names the model, effort and session shape of each phase, and the table follows that evidence rather than the other way round. **Effort is the cheaper lever than the model.** A reduced effort level is what gives up multi-file consistency first, so `high` is a reasonable floor for anything touching more than one file or a contract; `default` or `medium` suits a small mechanical change with a test behind it. ## Model per session, not per phase A session cannot switch its own model - that is the user's `/model` - and it should not be asked to mid-flow either. Two reasons: - **A switch throws away the prompt cache.** A cache entry belongs to the model that wrote it, so a new model starts the session's whole history from cold. The same holds for effort: changing it always invalidates the cached message history - by far the largest part of a long session - and, on some models, the tool and system prefix too (Anthropic's prompt-caching documentation lists effort and the thinking configuration among what invalidates the cache). A switch from one effort to another costs the same re-read of the whole session as a switch of model. - **An offered switch is rarely taken.** A sentence in the output at the moment a phase changes is easy to read past - for the agent writing it and for the user reading it - and the session just carries on in whatever it started as. So the choice is made once, when a session starts, and phases that want different models or effort levels are separated by a session boundary instead: `/clear`, then the next phase in a session started the right way. That is only cheap if the next phase does not depend on the previous session's context - which is why the handover has to live somewhere outside the session (an issue body, a page) and be kept current at fixed points, not reconstructed at the end. In this repo the handover points are a ready issue body (design → build) and a green CI run with the body updated (build → closing); the two later skills can only be started by the user's slash command, so each phase change is a real stop at which that choice is made. ## Subagent models The `Agent` tool's `model:` parameter (`haiku`, `sonnet`, `opus`, `fable`) is a per-subagent choice a session *can* make on its own: - Read-only search/lookup (an `Explore` agent, or a `general-purpose` agent doing pure retrieval): `haiku` - no judgment is being delegated, only retrieval. - A subagent that writes pages, reviews code, or decides something: leave `model:` off so it inherits the parent session's model. - A fork (`subagent_type: "fork"`) always inherits the parent's model; a `model:` override on a fork is ignored. ## `/code-review` effort - A routine diff: `low` or `medium` - fewer, high-confidence findings are enough. - Gate code, the compiler, or a change about to ship in a version bump: `high` and up - broader coverage is worth it when the blast radius of a missed bug is a safety gate. - `ultra` is user-triggered and billed separately - worth recommending, not assuming. ## When it's unclear - A task spans both a mechanical step and a judgment call: weigh it by the judgment call, not the mechanical one - the tooling carries the mechanical part regardless of which model supervises. - No row fits cleanly: Sonnet at high effort is a safer default than the most capable model at the highest effort. Under-provisioning where a check exists costs one worse answer once; reflexively over-provisioning is a standing cost every session pays. - Not sure whether a phase is checked: treat it as unchecked - a needless Opus phase costs money once, an unchecked Sonnet phase can ship something nobody looks at again. - The phase changed and the session runs on a model or effort the table would not pick: keep working, and cut the session at the next handover rather than switching mid-flow - never block a publish or an issue close on a choice the session cannot make itself. Naming which model and effort ran which phase in the handover keeps the gap visible instead of silent. ## Scope Specific to Claude Code: the model names, the `/code-review` dial and the `Agent` tool's `model:` override have no equivalent in this repo's other supported harnesses (Codex CLI, GitHub Copilot CLI, Mistral Vibe). Does not set the classifier model behind Claude Code's own `auto` permission mode - that is a harness internal, not a per-task choice this repo controls.