Files changed: - CHANGES.md - VERSION - docs/ownership-and-templates.md - docs/why-gates-are-code.md
8.6 KiB
Why gates are code
Chemenu has four hard limits - the Mass-Update Gate, the Publish-Remote Gate, the Upload Review
Gate, and the Iteration Budget Gate - and all four live inside tools/wikitool, not in a
paragraph of instructions an agent reads and follows. The rules themselves, and what to do when
one trips, are in AGENTS.md § Gates and
instructions/gates.md. This page is only about the design choice
underneath them: why code, and why these four mechanisms in particular.
Contents
- A suggestion an agent can talk itself past
- Why four different mechanisms, not one
- Exit 42 is a posture, and it outgrew the gates
- A gate in code still has to be reachable
- Numbers that come from measurement, not intuition
A suggestion an agent can talk itself past
An instruction like "don't publish too much at once" or "don't loop forever" lives in the same place as every other piece of guidance a session is holding - alongside the task, the user's last message, and whatever context made the moment feel urgent. Under pressure, or with a plausible-sounding reason ("this batch is different, it's mechanical"), that guidance can be reasoned around without anyone deciding to break a rule. Nothing enforces it; it just competes for attention with everything else in the context window, and sometimes loses.
A check compiled into the tool doesn't have that problem, because it isn't part of the conversation at all. It runs before the command dispatches, regardless of how convincing the case for skipping it seemed a moment earlier. The difference isn't that code is smarter than a well-written instruction - it's that code doesn't get talked into anything.
Why four different mechanisms, not one
The four gates ask four different questions, and each one's shape follows from what kind of question it is.
The Mass-Update Gate asks is this change too large to publish unreviewed - a judgment that
varies changeset by changeset, so it clears with a --confirm token tied to the specific
output the user just read. Approval is scoped to that one publish.
The Publish-Remote Gate asks something underneath that: is this even the right repository.
That's not a per-push judgment, it's a standing property of the checkout - true or false for
every publish that checkout will ever attempt, not just this one. A confirm token would let an
agent clear it once and then treat the answer as settled, which is exactly backwards for a
question whose answer shouldn't move at all mid-session. The only way past it is the user
editing .wikitool-remotes.json directly, outside the gate's own flow.
The Upload Review Gate asks the Mass-Update Gate's own question - is this change right? - at
the opposite end of its size range: one file from a stranger instead of a changeset from the
session's own work. That similarity is exactly why it reuses the same shape (a --confirm token
digesting the thing being approved) rather than inventing a fourth one: the two gates differ in
who produced the change and how much of it there is, not in what kind of question either one
is answering, so nothing about the mechanism needed to change - only the boundary it sits behind
did, since the submission lives in a quarantine the ordinary pipeline never reads at all rather
than in the working tree publish is about to commit.
The Iteration Budget Gate asks a fourth kind of question - not "is this instance correct" but "has this session stopped making progress." That's read from the shape of the call history itself (call count, repeated identical calls), not from anything about the content of any one call.
Exit 42 is a posture, and it outgrew the gates
Those four are the named gates, and they are not the only thing that exits 42 any more. When the
task-tracker provider layer arrived, it brought a case that looks like a gate from the outside and
is not one: a provider whose API cannot create a project (Super Productivity's local REST API
reads projects but does not write them) raises HumanInterventionRequired, and the command prints
what a human has to do and exits 42.
Reusing the code was deliberate, and so was not calling it a fifth gate. What the four gates share is a refusal: the operation was possible and the tool declined to perform it unreviewed. This is the opposite situation - the operation is not possible at all, and no token could make it possible. What the two have in common is only what the exit code actually communicates: stop, show this to a human, do not improvise a way around it. That sentence is the whole meaning of 42 here, and it is worth more as a shared convention than as a number reserved for one mechanism.
The alternative was worse in a specific way. A provider that cannot do something could have been described in the instruction layer instead - "if you are on this tracker, create the project by hand first" - which is exactly the prose-shaped rule this page argues against, with the added cost that every instruction would then have to know which provider an instance runs. The capability gap belongs where the capability is, and reaches the session as an exit code rather than as a paragraph it has to remember to apply.
A gate in code still has to be reachable
Code beats prose for the reason above, but on its own it buys less than it looks like: a check that runs on every call is only as good as the thing it counts under. The Iteration Budget Gate scopes its counter to a session, and "session" was approximated by the parent process id whenever nothing set an explicit one. On a harness that runs every tool call in a freshly initialised shell, that approximation hands out a new session per call - so a traced run of thirty-three calls arrived as twenty-one sessions of one to three calls each, the ceiling of sixty was never approached, and the loop-breaker's window never held three calls at once to compare. The gate ran on every one of those calls, exactly as written, and refused nothing.
That failure has no symptom of its own. A gate that fires announces that it exists; a gate that
cannot fire looks identical to a gate nobody happened to need - the same clean runs, the same
silence - and what finally told the two apart was reading a trace for an unrelated reason. So
there is a third property to keep alongside living in code and carrying measured numbers: each
gate has to leave evidence that it can still fire. The three that clear by token or by a
deliberate edit have it by construction, because clearing one is a visible event in somebody's
terminal. The budget gate, whose ordinary outcome is silence, is the one that had to be given
it - which is why its session id now carries where it came from, into both the trace and
budget status, so a session's own record answers the question instead of an investigation
having to.
Numbers that come from measurement, not intuition
The iteration ceiling didn't start where it sits now. It used to run 15-25, borrowed from a
general rule of thumb, until four real ingest runs measured 24, 26, 29 and 30 calls apiece -
every one of them an ordinary workflow doing nothing wrong, and every one of them at or past
where the old ceiling would have refused it. A limit that the normal case keeps tripping stops
functioning as a limit; it becomes background noise a session learns to route --override-budget
around as a matter of course, and the whole point of a hard-coded check is that it isn't supposed
to feel routine.
That's the deeper reason these numbers live in a tool rather than in prose: prose is read once and remembered loosely, but a threshold enforced every call is tested by every call, and a threshold that fails its own test gets noticed and re-measured rather than quietly ignored.
The suite's coverage floor is the same argument run forwards instead of backwards. The ceiling above was wrong first and measured afterwards; the floor was withheld on purpose until the number existed - measured, then watched across 38 runs while the code grew by a quarter, and only then written down as 85 against an observed 87.0%. The two points of daylight are the same consideration as the ceiling's headroom: a limit the ordinary case keeps tripping stops being a limit. A coverage floor set at the measured number goes red on the next thin command wrapper, which is not a regression, and a threshold that goes red for a non-reason gets lowered rather than earned - the failure mode above, reached from the other direction.