docs: Modellwahl nach Pruefbarkeit; stack-dev bricht an den Phasenwechseln fuer den Model-Switch (4.3.3)
Files changed: - CHANGES.md - VERSION - instructions/claude-code-model-selection.md - instructions/dev/stack-dev/SKILL.md
This commit is contained in:
@@ -6,10 +6,24 @@ description: Which Claude model and effort level to run a Claude Code session, a
|
||||
|
||||
# Pick the Claude model and effort level for the task at hand
|
||||
|
||||
Scale the model and effort to how much judgment the task actually needs. Running everything at
|
||||
the most capable model and highest effort is safe but wasteful: the gates in [gates.md](gates.md)
|
||||
are enforced in code, not by model judgment, so a weaker model cannot bypass them - it can only
|
||||
do a worse job of the calls the gates don't cover.
|
||||
Scale the model and effort to **what catches a mistake in this part of the work** - not to how
|
||||
important the task feels, and not to its name. Running everything at the most capable model and
|
||||
highest effort is safe but wasteful: the gates in [gates.md](gates.md) are enforced in code, not
|
||||
by model judgment, so a weaker model cannot bypass them - it can only do a worse job of the calls
|
||||
the gates don't cover.
|
||||
|
||||
That last clause is the whole rule, turned into a test. Where a check lives in code - `pytest`,
|
||||
`docs verify`, `instructions verify`, CI, the gates - a weaker model's mistake surfaces and costs
|
||||
one more round. Where the only enforcement is a session reading prose, the same mistake does not
|
||||
surface at all: it ships, and it stays until someone happens to notice. The two are not the same
|
||||
risk, and they should not get the same model. This is the argument
|
||||
[docs/why-gates-are-code.md](../docs/why-gates-are-code.md) makes about gates, applied to who is
|
||||
holding the keyboard.
|
||||
|
||||
Both directions cost something, which is why the axis matters rather than a blanket answer:
|
||||
over-provisioning is a standing cost paid every session, while under-provisioning in an unchecked
|
||||
phase is a silent error with a long tail. A corrective session, its bump, its CI runs and its
|
||||
release together cost more compute than the model difference they were saving.
|
||||
|
||||
Claude-Code-only, and imported by CLAUDE.md rather than linked from AGENTS.md: the model names,
|
||||
the `/code-review` effort dial and the `Agent` tool's `model:` override have no equivalent in the
|
||||
@@ -28,17 +42,41 @@ to *make*, not a setting to apply.
|
||||
|
||||
## Steps
|
||||
|
||||
1. **Recommend the session's model and effort by the skill in use**, when asked or when the
|
||||
mismatch is worth one sentence. Say it once and continue working either way - a session that
|
||||
argues about its own model instead of doing the task has already cost more than the model
|
||||
difference:
|
||||
1. **Recommend the session's model and effort by what catches a mistake in the phase it is in**,
|
||||
when asked or when the mismatch is worth one sentence. Say it once and continue working either
|
||||
way - a session that argues about its own model instead of doing the task has already cost
|
||||
more than the model difference:
|
||||
|
||||
| Skill / task | Model | Effort |
|
||||
|---|---|---|
|
||||
| `wiki-status`, simple `wiki-query` lookups | Sonnet | default |
|
||||
| `wiki-lint` | Sonnet | default |
|
||||
| `wiki-ingest`, `wiki-manage`, judgment-heavy `wiki-query` | Sonnet | high |
|
||||
| Stack development: `tools/`, `types/`, `instructions/` as code | Opus | high |
|
||||
| Phase / task | What catches a mistake here | Model | Effort |
|
||||
|---|---|---|---|
|
||||
| `wiki-status`, simple `wiki-query` lookups | the answer is re-checkable against the corpus | Sonnet | default |
|
||||
| `wiki-lint` | `lint` itself is the check | Sonnet | default |
|
||||
| `wiki-ingest`, `wiki-manage`, judgment-heavy `wiki-query` | `lint` and `docs verify`, partly - the judgment about a claim is not covered | Sonnet | high |
|
||||
| Stack dev: design, the version part, a boundary-crossing judgment | nothing - `docs verify` checks that a crossing documents itself, never that the part was right | Opus | high |
|
||||
| Stack dev: code, tests, mechanical doc sync (command tables, contract rows) | `pytest`, `docs verify`, `instructions verify`, CI | Sonnet | high |
|
||||
| Stack dev: closing an issue, `docs/` staleness, changelog prose | nothing, by construction - see below | Opus | high |
|
||||
|
||||
**Stack development is not one row**, which is the point of splitting it. The middle phase is
|
||||
where the tokens are and where the checks are, so it is the phase worth running cheaper. The
|
||||
two around it have no mechanical guard at all - a `docs/` page carries no normative sentence,
|
||||
so there is nothing for `docs verify` to check ([AGENTS.md](../AGENTS.md) § File naming), and
|
||||
the same holds for whatever tracker an instance keeps its open work in, which `wikitool`
|
||||
deliberately knows nothing about. Those two phases are short - minutes, not hours - so keeping
|
||||
them on the stronger model is cheap, and it protects the only work in the session that fails
|
||||
silently.
|
||||
|
||||
**Effort is the cheaper lever than the model.** Reach for it first: `medium` deliberately does
|
||||
not appear in this table for stack work, because multi-file consistency is exactly what a
|
||||
reduced effort level gives up. Sonnet at `high` is the floor for anything touching more than
|
||||
one file or a contract; `default` is for a single-file mechanical edit with a test behind it.
|
||||
|
||||
**A session cannot switch its own model**, so these rows only become real if someone offers the
|
||||
switch at the moment the phase changes - once, without arguing about it, and never as a reason
|
||||
to stop work that is already underway.
|
||||
<!-- dist:strip-start -->
|
||||
In this repo those moments are named: the `stack-dev` skill breaks for them at its steps 3
|
||||
(design settled, work turns mechanical) and 6 (publish done, the unchecked tail begins).
|
||||
<!-- dist:strip-end -->
|
||||
|
||||
2. **Pick a spawned subagent's model by what it does**, via the `Agent` tool's `model:`
|
||||
parameter - the values are `haiku`, `sonnet`, `opus`, `fable`:
|
||||
@@ -66,8 +104,14 @@ to *make*, not a setting to apply.
|
||||
mechanical one - `wikitool` carries the mechanical part regardless of which model is
|
||||
supervising it.
|
||||
- **Unsure which row applies?** Default to Sonnet at high effort, not the most capable model at
|
||||
the highest effort. Under-provisioning costs one worse answer in one session; reflexively
|
||||
over-provisioning is a standing cost paid every session.
|
||||
the highest effort. Under-provisioning *where a check exists* costs one worse answer in one
|
||||
session; reflexively over-provisioning is a standing cost paid every session.
|
||||
- **Unsure whether the phase is checked?** Treat it as unchecked. The asymmetry is not symmetric:
|
||||
a needless Opus phase costs money once, an unchecked Sonnet phase can ship something nobody
|
||||
looks at again.
|
||||
- **Mid-session and the phase changed, but nobody switched?** Do the work anyway - never block a
|
||||
publish or an issue close on a model the session cannot change itself. Say which phase ran on
|
||||
which model in the handover, so the gap is visible rather than silent.
|
||||
|
||||
## Scope
|
||||
|
||||
|
||||
Reference in New Issue
Block a user