stack-build and stack-close carry disable-model-invocation, so each phase change is the operator's slash command; no skill offers a mid-session /model or /effort switch. Mode rules and the phase table move to instructions/dev/stack-mode.md, publish and CI waiting to instructions/dev/publish-and-ci.md, the ready definition to issue-tracking.md. Files changed: - AGENTS.md - CHANGES.md - DEVELOPMENT.md - README.md - VERSION - docs/model-and-effort-selection.md - instructions/CONTRACT.md - instructions/dev/commonplace-kb.md - instructions/dev/dev-setup.md - instructions/dev/doc-pull-through.md - instructions/dev/issue-tracking.md - instructions/dev/publish-and-ci.md - instructions/dev/stack-build/SKILL.md - instructions/dev/stack-close/SKILL.md - instructions/dev/stack-dev/SKILL.md - instructions/dev/stack-mode.md - instructions/dev/testing-conventions.md - instructions/dev/version-parts.md - tools/chemenu/commands/git_publish.py - tools/chemenu/tests/test_git_publish.py - tools/chemenu/tests/test_instructions_cmd.py Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SnAJ7Z3CpVD3PRbN73QtU2
124 lines
7.4 KiB
Markdown
124 lines
7.4 KiB
Markdown
# Choosing a Claude Code model and effort level
|
|
|
|
Claude Code exposes three choices this repo has an opinion on: which model a session itself
|
|
runs as, what model a spawned subagent gets, and which `/code-review` effort level to pick.
|
|
None of them are enforced anywhere - the gates in [instructions/gates.md](../instructions/gates.md)
|
|
are code precisely because a model cannot be talked out of them
|
|
([why-gates-are-code.md](why-gates-are-code.md) makes that argument for gates; this page applies
|
|
the same axis to who is holding the keyboard). What follows is a reference for making that choice
|
|
well, not a rule anything checks.
|
|
|
|
The axis worth tracking is not how important a task feels, but **what would catch a mistake in
|
|
it**. Work behind `pytest`, `docs verify`, `instructions verify` or CI surfaces a bad call within
|
|
one more round. Work behind nothing but a session reading prose does not surface at all - it
|
|
ships, and stays until someone happens to notice. That asymmetry, not task size, is what the
|
|
phase guide below is built on.
|
|
|
|
<!-- wikitool:toc -->
|
|
## Contents
|
|
|
|
- [Phase guide](#phase-guide)
|
|
- [Model per session, not per phase](#model-per-session-not-per-phase)
|
|
- [Subagent models](#subagent-models)
|
|
- [`/code-review` effort](#code-review-effort)
|
|
- [When it's unclear](#when-its-unclear)
|
|
- [Scope](#scope)
|
|
<!-- /wikitool:toc -->
|
|
|
|
## Phase guide
|
|
|
|
<!-- dist:strip-start -->
|
|
This repo's own stack-development work splits the axis into three phases, one skill each -
|
|
`stack-dev` (design), `stack-build` (build) and `stack-close` (closing) - handed over through
|
|
states in the issue tracker rather than inside one session:
|
|
|
|
<!-- dist:strip-end -->
|
|
| Phase / task | What would catch a mistake | Suggested model | Effort |
|
|
|---|---|---|---|
|
|
| `wiki-status`, simple `wiki-query` lookups | the answer is re-checkable against the corpus | Sonnet | default |
|
|
| `wiki-lint` | `lint` itself is the check | Sonnet | default |
|
|
| `wiki-ingest`, `wiki-manage`, judgment-heavy `wiki-query` | `lint` and `docs verify`, partly | Sonnet | high |
|
|
| Stack dev: design, the version part, a boundary-crossing judgment | nothing mechanical | Opus | high |
|
|
| Stack dev: code, tests, mechanical doc sync, waiting for CI | `pytest`, `docs verify`, `instructions verify`, CI | Opus (open - see below) | high; medium for a small change |
|
|
| Stack dev: closing an issue, `docs/` staleness, changelog prose | nothing, by construction | Opus | high |
|
|
|
|
The two unchecked rows are short - minutes, not hours - so keeping them on the strongest model
|
|
costs little and protects the only work that fails silently. The checked middle row is where the
|
|
tokens are, which makes it the tempting one to run cheaper. It is also the row with the least
|
|
settled answer: a build phase of this stack reads a lot of the tree, and a smaller context window
|
|
does not hold it - it runs into compaction, which costs more than the cheaper model saves. Sonnet
|
|
at high effort remains possible, but only as a session of its own (below). Which default the
|
|
middle row should have is left to the record: every closed work package names the model, effort
|
|
and session shape of each phase, and the table follows that evidence rather than the other way
|
|
round.
|
|
|
|
**Effort is the cheaper lever than the model.** A reduced effort level is what gives up
|
|
multi-file consistency first, so `high` is a reasonable floor for anything touching more than one
|
|
file or a contract; `default` or `medium` suits a small mechanical change with a test behind it.
|
|
|
|
## Model per session, not per phase
|
|
|
|
A session cannot switch its own model - that is the user's `/model` - and it should not be asked
|
|
to mid-flow either. Two reasons:
|
|
|
|
- **A switch throws away the prompt cache.** A cache entry belongs to the model that wrote it,
|
|
so a new model starts the session's whole history from cold. The same holds for effort: changing it always invalidates
|
|
the cached message history - by far the largest part of a long session - and, on some models,
|
|
the tool and system prefix too (Anthropic's prompt-caching documentation lists effort and the
|
|
thinking configuration among what invalidates the cache). A switch from one effort to another
|
|
costs the same re-read of the whole session as a switch of model.
|
|
- **An offered switch is rarely taken.** A sentence in the output at the moment a phase changes
|
|
is easy to read past - for the agent writing it and for the user reading it - and the session
|
|
just carries on in whatever it started as.
|
|
|
|
So the choice is made once, when a session starts, and phases that want different models or
|
|
effort levels are separated by a session boundary instead: `/clear`, then the next phase in a
|
|
session started the right way. That is only cheap if the next phase does not depend on the
|
|
previous session's context - which is why the handover has to live somewhere outside the session
|
|
(an issue body, a page) and be kept current at fixed points, not reconstructed at the end.
|
|
<!-- dist:strip-start -->
|
|
In this repo the handover points are a ready issue body (design → build) and a green CI run with
|
|
the body updated (build → closing); the two later skills can only be started by the user's slash
|
|
command, so each phase change is a real stop at which that choice is made.
|
|
<!-- dist:strip-end -->
|
|
|
|
## Subagent models
|
|
|
|
The `Agent` tool's `model:` parameter (`haiku`, `sonnet`, `opus`, `fable`) is a per-subagent
|
|
choice a session *can* make on its own:
|
|
|
|
- Read-only search/lookup (an `Explore` agent, or a `general-purpose` agent doing pure
|
|
retrieval): `haiku` - no judgment is being delegated, only retrieval.
|
|
- A subagent that writes pages, reviews code, or decides something: leave `model:` off so it
|
|
inherits the parent session's model.
|
|
- A fork (`subagent_type: "fork"`) always inherits the parent's model; a `model:` override on a
|
|
fork is ignored.
|
|
|
|
## `/code-review` effort
|
|
|
|
- A routine diff: `low` or `medium` - fewer, high-confidence findings are enough.
|
|
- Gate code, the compiler, or a change about to ship in a version bump: `high` and up - broader
|
|
coverage is worth it when the blast radius of a missed bug is a safety gate.
|
|
- `ultra` is user-triggered and billed separately - worth recommending, not assuming.
|
|
|
|
## When it's unclear
|
|
|
|
- A task spans both a mechanical step and a judgment call: weigh it by the judgment call, not the
|
|
mechanical one - the tooling carries the mechanical part regardless of which model supervises.
|
|
- No row fits cleanly: Sonnet at high effort is a safer default than the most capable model at
|
|
the highest effort. Under-provisioning where a check exists costs one worse answer once;
|
|
reflexively over-provisioning is a standing cost every session pays.
|
|
- Not sure whether a phase is checked: treat it as unchecked - a needless Opus phase costs money
|
|
once, an unchecked Sonnet phase can ship something nobody looks at again.
|
|
- The phase changed and the session runs on a model or effort the table would not pick: keep
|
|
working, and cut the session at the next handover rather than switching mid-flow - never block a
|
|
publish or an issue close on a choice the session cannot make itself. Naming which model and
|
|
effort ran which phase in the handover keeps the gap visible instead of silent.
|
|
|
|
## Scope
|
|
|
|
Specific to Claude Code: the model names, the `/code-review` dial and the `Agent` tool's `model:`
|
|
override have no equivalent in this repo's other supported harnesses (Codex CLI, GitHub Copilot
|
|
CLI, Mistral Vibe). Does not set the classifier model behind Claude Code's own `auto` permission
|
|
mode - that is a harness internal, not a per-task choice this repo controls.
|