diff --git a/docs/model-and-effort-selection.md b/docs/model-and-effort-selection.md index b7f55a7..b52f33b 100644 --- a/docs/model-and-effort-selection.md +++ b/docs/model-and-effort-selection.md @@ -47,10 +47,12 @@ costs little and protects the only work that fails silently. The checked middle tokens are, which makes it the tempting one to run cheaper. It is also the row with the least settled answer: a build phase of this stack reads a lot of the tree, and a smaller context window does not hold it - it runs into compaction, which costs more than the cheaper model saves. Sonnet -at high effort remains possible, but only as a session of its own (below). Which default the -middle row should have is left to the record: every closed work package names the model, effort -and session shape of each phase, and the table follows that evidence rather than the other way -round. +at high effort remains possible, but only as a session of its own (below). + +Which default the middle row should have is left to the record: every closed work package in this +repo names the model, effort and session shape of each phase, and the table follows that evidence +rather than the other way round. + **Effort is the cheaper lever than the model.** A reduced effort level is what gives up multi-file consistency first, so `high` is a reasonable floor for anything touching more than one @@ -62,11 +64,12 @@ A session cannot switch its own model - that is the user's `/model` - and it sho to mid-flow either. Two reasons: - **A switch throws away the prompt cache.** A cache entry belongs to the model that wrote it, - so a new model starts the session's whole history from cold. The same holds for effort: changing it always invalidates - the cached message history - by far the largest part of a long session - and, on some models, - the tool and system prefix too (Anthropic's prompt-caching documentation lists effort and the - thinking configuration among what invalidates the cache). A switch from one effort to another - costs the same re-read of the whole session as a switch of model. + so a new model starts the session's whole history from cold. The same holds for effort: + changing it always invalidates the cached message history - by far the largest part of a long + session - and, on some models, the tool and system prefix too (Anthropic's prompt-caching + documentation lists effort and the thinking configuration among what invalidates the cache). A + switch from one effort to another costs the same re-read of the whole session as a switch of + model. - **An offered switch is rarely taken.** A sentence in the output at the moment a phase changes is easy to read past - for the agent writing it and for the user reading it - and the session just carries on in whatever it started as.