feat: stack development in three phases - stack-dev (design), stack-build, stack-close, handed over through tracker states (#168)
CI / verify (push) Successful in 5m15s
CI / pwsh (push) Successful in 2m2s
Release / release (push) Successful in 34s

stack-build and stack-close carry disable-model-invocation, so each phase change is
the operator's slash command; no skill offers a mid-session /model or /effort switch.
Mode rules and the phase table move to instructions/dev/stack-mode.md, publish and CI
waiting to instructions/dev/publish-and-ci.md, the ready definition to issue-tracking.md.

Files changed:
- AGENTS.md
- CHANGES.md
- DEVELOPMENT.md
- README.md
- VERSION
- docs/model-and-effort-selection.md
- instructions/CONTRACT.md
- instructions/dev/commonplace-kb.md
- instructions/dev/dev-setup.md
- instructions/dev/doc-pull-through.md
- instructions/dev/issue-tracking.md
- instructions/dev/publish-and-ci.md
- instructions/dev/stack-build/SKILL.md
- instructions/dev/stack-close/SKILL.md
- instructions/dev/stack-dev/SKILL.md
- instructions/dev/stack-mode.md
- instructions/dev/testing-conventions.md
- instructions/dev/version-parts.md
- tools/chemenu/commands/git_publish.py
- tools/chemenu/tests/test_git_publish.py
- tools/chemenu/tests/test_instructions_cmd.py

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SnAJ7Z3CpVD3PRbN73QtU2
This commit is contained in:
torbenandClaude Opus 5.5 committed 2026-10-02 20:33:07 +02:00
1 parent 80488bee38
commit d8224ee2ab
21 files changed
+631 -331

No files matched your search

+54 -13
View File
@@ -14,11 +14,23 @@ one more round. Work behind nothing but a session reading prose does not surface
ships, and stays until someone happens to notice. That asymmetry, not task size, is what the
phase guide below is built on.
<!-- wikitool:toc -->
## Contents
- [Phase guide](#phase-guide)
- [Model per session, not per phase](#model-per-session-not-per-phase)
- [Subagent models](#subagent-models)
- [`/code-review` effort](#code-review-effort)
- [When it's unclear](#when-its-unclear)
- [Scope](#scope)
<!-- /wikitool:toc -->
## Phase guide
<!-- dist:strip-start -->
This repo's own stack-development work splits the axis into three phases, one per switch point
in its `stack-dev`/`stack-close` skills:
This repo's own stack-development work splits the axis into three phases, one skill each -
`stack-dev` (design), `stack-build` (build) and `stack-close` (closing) - handed over through
states in the issue tracker rather than inside one session:
<!-- dist:strip-end -->
| Phase / task | What would catch a mistake | Suggested model | Effort |
@@ -27,20 +39,48 @@ in its `stack-dev`/`stack-close` skills:
| `wiki-lint` | `lint` itself is the check | Sonnet | default |
| `wiki-ingest`, `wiki-manage`, judgment-heavy `wiki-query` | `lint` and `docs verify`, partly | Sonnet | high |
| Stack dev: design, the version part, a boundary-crossing judgment | nothing mechanical | Opus | high |
| Stack dev: code, tests, mechanical doc sync | `pytest`, `docs verify`, `instructions verify`, CI | Sonnet | high |
| Stack dev: code, tests, mechanical doc sync, waiting for CI | `pytest`, `docs verify`, `instructions verify`, CI | Opus (open - see below) | high; medium for a small change |
| Stack dev: closing an issue, `docs/` staleness, changelog prose | nothing, by construction | Opus | high |
The middle stack-dev row is where the tokens are and where the checks are, so it is the one worth
running cheaper. The two rows around it are short - minutes, not hours - so keeping them on the
stronger model costs little and protects the only work in the session that fails silently.
The two unchecked rows are short - minutes, not hours - so keeping them on the strongest model
costs little and protects the only work that fails silently. The checked middle row is where the
tokens are, which makes it the tempting one to run cheaper. It is also the row with the least
settled answer: a build phase of this stack reads a lot of the tree, and a smaller context window
does not hold it - it runs into compaction, which costs more than the cheaper model saves. Sonnet
at high effort remains possible, but only as a session of its own (below). Which default the
middle row should have is left to the record: every closed work package names the model, effort
and session shape of each phase, and the table follows that evidence rather than the other way
round.
**Effort is the cheaper lever than the model.** A reduced effort level is what gives up
multi-file consistency first, so `high` is a reasonable floor for anything touching more than one
file or a contract; `default` suits a single-file mechanical edit with a test behind it.
file or a contract; `default` or `medium` suits a small mechanical change with a test behind it.
A session cannot switch its own model - that is the user's `/model` - so this table only pays off
if someone offers the switch at the moment a phase changes, once, without turning it into a
debate.
## Model per session, not per phase
A session cannot switch its own model - that is the user's `/model` - and it should not be asked
to mid-flow either. Two reasons:
- **A switch throws away the prompt cache.** A cache entry belongs to the model that wrote it,
so a new model starts the session's whole history from cold. The same holds for effort: changing it always invalidates
the cached message history - by far the largest part of a long session - and, on some models,
the tool and system prefix too (Anthropic's prompt-caching documentation lists effort and the
thinking configuration among what invalidates the cache). A switch from one effort to another
costs the same re-read of the whole session as a switch of model.
- **An offered switch is rarely taken.** A sentence in the output at the moment a phase changes
is easy to read past - for the agent writing it and for the user reading it - and the session
just carries on in whatever it started as.
So the choice is made once, when a session starts, and phases that want different models or
effort levels are separated by a session boundary instead: `/clear`, then the next phase in a
session started the right way. That is only cheap if the next phase does not depend on the
previous session's context - which is why the handover has to live somewhere outside the session
(an issue body, a page) and be kept current at fixed points, not reconstructed at the end.
<!-- dist:strip-start -->
In this repo the handover points are a ready issue body (design → build) and a green CI run with
the body updated (build → closing); the two later skills can only be started by the user's slash
command, so each phase change is a real stop at which that choice is made.
<!-- dist:strip-end -->
## Subagent models
@@ -70,9 +110,10 @@ choice a session *can* make on its own:
reflexively over-provisioning is a standing cost every session pays.
- Not sure whether a phase is checked: treat it as unchecked - a needless Opus phase costs money
once, an unchecked Sonnet phase can ship something nobody looks at again.
- Mid-session and the phase changed but nobody switched: keep working - never block a publish or
an issue close on a model the session cannot change itself. Naming which model ran which phase
in the handover keeps the gap visible instead of silent.
- The phase changed and the session runs on a model or effort the table would not pick: keep
working, and cut the session at the next handover rather than switching mid-flow - never block a
publish or an issue close on a choice the session cannot make itself. Naming which model and
effort ran which phase in the handover keeps the gap visible instead of silent.
## Scope