Files
chemenu/instructions/dev/issue-tracking.md
T
torben a51d7a322f
CI / verify (push) Successful in 59s
Release / release (push) Successful in 37s
docs: ausgelieferte Doku zitiert keine Issue-Nummern mehr, docs verify prueft es (schliesst #77)
Files changed:
- .gitignore
- CHANGES.md
- EVALS.md
- INSTALL.md
- README.md
- VERSION
- docs/pipeline-rationale.md
- instructions/CONTRACT.md
- instructions/bootstrap.md
- instructions/dev/issue-tracking.md
- instructions/evolve-subtypes.md
- instructions/kb-profiles.md
- instructions/mcp-read-server.md
- instructions/wiki-ingest/SKILL.md
- kb/CONTRACT.md
- kb/concepts/COLLECTION.md
- kb/sources/COLLECTION.md
- raw/CONTRACT.md
- tools/.coveragerc
- tools/CONTRACT.md
- tools/README.md
- tools/chemenu/commands/docs_verify.py
- tools/chemenu/tests/test_docs_verify.py
- types/source.schema.yaml
- types/type-spec.md
2026-09-09 18:52:34 +02:00

402 lines
23 KiB
Markdown

---
type: types/instruction.md
name: issue-tracking
description: Where open work on this stack is tracked, what the four mandatory area/kind/prio/size labels and the three status flags on a Gitea issue mean, why a status/incoming stub is never implemented as it stands, and how to keep an issue body current across sessions.
---
# Track open work as Gitea issues, not as prose in the repo
Open work on this stack lives at
<https://gitea.nehmer.net/torben/chemenu/issues>, one issue per work
package, and nowhere else. There is no `TODO.md`; there was, and every item in
it either became an issue or was already one, described twice.
That is the whole reason for this file: a second list is a second thing to
maintain, and the one that drifts is always the one nobody reads first. The
issue tracker wins that comparison outright - it has state, comments, labels,
and a link that survives the change it describes. A markdown file in the repo
has none of it, and it costs a publish to touch.
This instruction exists only in the dev repo. A distributed instance has no
issues at that URL, which is exactly why `dist export` excludes
`instructions/dev/` wholesale (see [tools/CONTRACT.md](../../tools/CONTRACT.md)).
## When to run
- Something is worth doing but not now. Open an issue; do not write it down in
the repo.
- A session's findings outgrow the change it was making - a gap in the tooling,
an assumption nobody has checked, a decision that needs the user.
- Picking an issue up: before doing anything else, read the body as the current
spec, and re-label it if the ground has moved since.
- **The issue carries `status/incoming`:** it is a human's stub, not a spec, and
it is worked out and triaged before anything is built from it
(§ Incoming stubs).
- **While working on one:** the body is updated as the state moves, not at the
end (step 2). A session that is interrupted leaves the body as its handover.
- Prioritising: deciding what to pick up next, or re-labelling after the ground
moved.
- Closing one: the body is rewritten to its final state first, and only then
closed (step 7).
- A rename or move ships: sweep the open issues for text that assumed the old
name or path (§ Renames and other decay in the tracker).
## Steps
1. **Write the issue so it survives without you.** What is broken or missing,
why it matters, what "done" looks like as acceptance criteria, and the
specific files or commands involved. An issue that only makes sense to
whoever wrote it is a note, and notes were the problem.
**Destructive steps carry the invariant they must not violate.** A body
that prescribes a mechanism gets built as prescribed - including its
bugs. Where a step deletes, overwrites, resets or moves, name the
property that must still hold afterwards, not only the command that gets
there. "Remove the working directory, then `git checkout HEAD --
<stage>`" is a mechanism; "the content stages must afterwards match
`HEAD` exactly, without any untracked or ignored file being touched" is
the same instruction plus its test - a build instruction and an
acceptance criterion at once, so the defect surfaces while the test is
written rather than in review afterwards. #30's `upstream merge` body
wrote the mechanism and got exactly that bug: a working-directory removal
that took a stage's gitignored, unrecoverable data with it.
**An acceptance criterion states a checkable property, not an activity.**
"Implement X" is done when someone says so; "after `upstream merge`,
`reports/` still holds every file it held before" is done when it is
true. This is not a ban on imperative steps - a numbered procedure can
still produce a correct control flow, and that is its merit - it binds
the destructive steps, and every box in the criteria list.
2. **The body is the working state, not a historical first post - keep it
current as you go.** It is this stack's plan file: the same thing a harness's
own plan document is, and it is maintained the same way. Not written once,
not brought up to date at the end, but **updated whenever something in it
stops being true** - a decision made, a criterion met, an approach ruled out,
a new constraint found.
The test is an abort, not a milestone. A session can end at any moment - an
interrupt, a context limit, a crash, a human walking away - and whatever the
body says at that instant is the entire handover. So the standard is: **at
every point, a fresh session must be able to open the body and pick the work
up from there**, without a human re-explaining it and without reading back
through the comments. If the body would mislead someone who read it right
now, it is already out of date, whether or not the work is finished.
That means updating *during* the work, not only at its end:
- a decision gets made → the decision and its reasoning replace the question
- an acceptance criterion is done → tick it, in the same session that did it
- something turns out differently than the issue assumed → the assumption is
corrected where it stands, not contradicted three paragraphs later
- work is deferred or dropped → say so, with the reason, where the criterion is
**Rewrite, never append.** Do not add to a text that has become wrong: an
additively grown log forces every later reader to reconstruct the current
state by filtering the whole history, which is the exact cost the body exists
to remove. Comments carry the history (step 3); the body carries the state.
Body rewrites and comments are an LLM session's job. A human normally
touches only labels and metadata directly.
**Reading an issue, the body is the state and comments are history.** A
session picking an issue up reads the body as the spec; comments are read
for provenance - why something was decided, what was tried - never as
the current instruction. A recommendation in a comment can be older than
the body's decision and read just as convincingly: on #30 an earlier
comment recommended a smaller, `verify`-only command, while the body had
since settled on building the full `merge` command. A session trusting
the comment would have built the wrong thing, with a plausible
justification out of this repo's own tracker.
**A body that is demonstrably wrong is corrected first, not worked
around.** "Body beats comment" is a rule of precedence, not a licence to
execute a stale spec. Where a comment or the tree proves a claim in the
body false, the body is rewritten before the work starts - the rewrite
above is the fix; leaning on the comments as the "real" state is not.
#10 is the case: its body claimed coverage had never been measured while
three comments carried a percentage, a statement count and a CI run
number.
**Where two comments contradict each other, evidence decides, not
recency.** On #10, one comment showed a retrieved artifact with zero
items on a finished run - the report was not actually retrievable - and
a later comment declared the same criterion met without re-checking. The
later comment is not the newer truth, only the unchecked one. Resolve it
into the body with the evidence named, or mark the point open.
3. **Comment a changelog, never a copy.** A body rewrite gets one short comment
naming only what changed against the previous state - what is new, what is
gone, what was corrected. Do not snapshot the old body into a comment: a full
copy per revision forces a human to diff two prose texts, which is not a
readable history, only another copy.
One comment per *session's worth* of change, not per edit. Step 2 asks the
body to be kept current continuously, and a comment for every tick would bury
the board in noise; the changelog line summarises what that session moved.
Trivial upkeep - a typo, a tightened sentence - needs no comment at all.
```
**Changelog:** Decision 2 tightened - `kind/` may now change over an
issue's life. Old acceptance criterion 3 dropped (covered by #42).
```
4. **Give it all four mandatory labels: one `area/`, one `kind/`, one `prio/`,
one `size/`.** All four, always. Machine maintenance by an LLM session is
what makes four axes affordable - the original objection to a third and
fourth axis was the upkeep cost for a single human maintainer, and that
objection no longer holds.
| `area/` | Means |
|---|---|
| `area/kb` | The `kb/` schema, contract, confidence machinery, lint - the knowledge base as a system. |
| `area/distribution` | Shipping, upgrading and versioning an instance. |
| `area/corpus` | The content and scope of `kb/` in this instance, and the demo/testbed question. |
| `area/workflow` | Git, merging, branching, publish, PRs. |
| `area/process` | The development process itself, rather than the stack as an artefact. |
There is deliberately no `area/tools`: tooling is filed under the domain it
serves, not under where its code sits. The axis follows the stage split in
[AGENTS.md](../../AGENTS.md).
| `kind/` | Means |
|---|---|
| `kind/decision` | Waiting on an operator decision. |
| `kind/build` | Specified; waiting only on implementation time. |
| `kind/defect` | A finding: documentation and reality, or two documents, contradict each other. |
`kind/` is expected to change over an issue's life - `decision` becomes
`build` once the decision is made. That is session memory working, not a
labelling failure.
| `prio/` | Means |
|---|---|
| `prio/blocking` | Blocks or damages work in progress. Next. |
| `prio/planned` | Accrues interest. Planned. |
| `prio/waiting` | Worth doing, waiting on a trigger. |
`prio/waiting` is not a graveyard. It means the issue's value is real but
gated on something outside it - a decision, another issue, a second instance
existing. Name that trigger in the issue, or the label is a polite no.
| `size/` | Means |
|---|---|
| `size/S` | One session, one publish, a clear cut. |
| `size/M` | Several files; a contract or instruction change; its own test effort. |
| `size/L` | Several sessions, or open design questions before the first commit. |
Size is effort, not importance. A `prio/blocking size/S` is the best thing
on the board; a `prio/waiting size/L` is a thing to talk about before anyone
starts.
5. **Add a `status/` flag only when it applies.** All three are optional, because
each describes a temporary condition rather than a property every issue has.
| `status/` | Means |
|---|---|
| `status/blocked` | Waiting on another, still-open issue - not workable on its own, whatever its `prio/` says. Name the blocking issue in the body. |
| `status/unconfirmed` | A reported suspicion, not yet checked against actual behaviour. Applies to any `kind/`, not just `kind/defect`. |
| `status/incoming` | A human's stub: a request or a thought, filed at whatever length it arrived, deliberately short of everything step 1 asks for. **Never implemented as it stands** - § Incoming stubs. |
While `status/unconfirmed` is set, `size/` and `prio/` are provisional. Triage
ends it one of two ways: the flag comes off and `size`/`prio` are set for
real, or the issue is closed with the reason. An unverified suspicion does not
stay open indefinitely - the process-level analogue of AGENTS.md invariant 3.
`status/incoming` is the one flag that **suspends step 4** rather than
qualifying it. The four mandatory labels are not missing from such an issue,
they are not yet due: `area/` may be obvious, but `kind/`, `prio/` and `size/`
are answers to questions the stub has not been read against the tree to
settle. Labelling it all four on sight is the failure, not the omission - it
makes an unexamined stub look triaged. It is also the one flag a session never
*adds*: an issue a session files meets step 1 or it does not get filed.
6. **Re-label when the ground moves, and say why in a comment.** A trigger that
fired turns `prio/waiting` into `prio/planned`. A design question that got
answered can drop a size and move `kind/decision` to `kind/build`. Silent
re-labelling is how a board stops meaning anything.
7. **Closing is the last body update, not a comment.** If step 2 was followed
the body is already nearly there, and closing only settles what the final
run established. If it was not, closing is where the whole debt comes due -
and it comes due at the worst moment, because a closed body is the version
everyone reads afterwards and nobody revisits.
Either way the body reaches its final state *before* the issue closes:
proposals that were decided read as decided, a "to decide" section has become
the decision with its reasoning, acceptance criteria are ticked or struck with
a reason, and what was verified is named. Then close, with the one-line
changelog comment step 3 asks for.
Record what actually happened, not a commit hash alone: which proposals were
implemented, which were deliberately left out and why, and what was verified.
The issue is the only place that record survives - a changelog entry says
what changed, not what was decided against.
**A closing report in a comment does not satisfy this.** It reads as
complete to whoever writes it and leaves a body still phrased as open work:
unticked boxes, an undecided decision section, present tense about a defect
that no longer exists. #44 closed exactly that way, with a thorough comment
above a body that still asked for a decision that had already been made and
shipped. Nothing mechanical catches it (see below), which is why it is a step
rather than a habit.
## Incoming stubs
**A `status/incoming` issue is never implemented as it stands.** It is worked
out and triaged first, in a session, and only the result of that is built.
The flag exists because the tracker is also the human's inbox, and the two have
different entry costs. Step 1 asks for a body that survives without its author -
acceptance criteria, files, commands - and a thought worth keeping is not worth
that much work at the moment it occurs. So a stub is admitted at whatever
quality it arrives, and `status/incoming` is the receipt: this text was not held
to step 1, and nobody should read it as if it had been.
That is the whole danger. A stub *looks* like a body, and a body is what a
session trusts (step 2). What it actually holds is a symptom or a wish - #60
says the confidence defaults "feel too high", #61 says a mechanism from one
instruction "would be interesting" elsewhere. Neither states what done means,
and the parts they leave out are exactly the parts the human left to be worked
out. Building straight from one produces something that matches the sentence,
misses the intent, and closes the issue - so the question the stub was standing
in for is never asked again. It is step 1's mechanism-versus-invariant lesson
one stage earlier: there, a body prescribed a mechanism and got its bugs built;
here, a body prescribes nothing at all and gets the gap filled by whoever read
it fastest.
Working one out:
1. **Read the stub as a statement of intent, not a specification.** Its wording
is the only evidence of what was actually asked for. Reinterpret it and the
record of the request is gone - what remains is the session's reading of it,
indistinguishable from the human's.
2. **Check it against the tree before rewriting anything.** A stub may be a
suspicion (`status/unconfirmed` applies on top where it is), a duplicate of
something already built, or a premise that no longer holds. This is the step
that decides which of the two exits below the issue takes.
3. **Quote the stub verbatim in the elaboration comment, then rewrite the body.**
Step 2's "rewrite, never append" holds here as everywhere - but the rewrite
overwrites the only record of the request, and comments are where history
lives (step 3). Here the history *is* the request.
4. **Name the open questions; do not answer them.** Where the stub leaves
something a session cannot settle from the tree, it stays a question in the
body and the issue becomes `kind/decision`. Guessing turns the human's open
question into a spec that reads as decided, which is worse than the stub was:
the stub at least announced that it was incomplete.
5. **Then step 4 comes due** - all four mandatory labels, set against a body that
has been read against the tree. That is the moment the stub becomes a work
package.
6. **Remove `status/incoming`** and leave the one-line changelog comment step 3
asks for.
Triage ends a stub one of two ways, the same two `status/unconfirmed` has: it is
worked out, labelled and the flag comes off, or it is closed with the reason. A
stub does not sit in the inbox indefinitely.
Elaboration touches no file in the working tree, so it needs no version bump and
no publish - it is tracker work, and several stubs can be worked out in one pass.
What comes *after* it is an ordinary work package, picked up on its merits like
any other.
## Renames and other decay in the tracker
A rename is not finished when the tree is green. Renaming a package, a path,
a command, a flag or the repository itself moves text that lives outside the
working tree, and the open issues are the largest such text. Nothing catches
them - `wikitool` does not know this tracker exists and must not learn (see
"What no tool checks" below) - so a pass over the open issues is part of the
rename, in the session that did it, not a follow-up someone remembers.
Distinguish a wayfinder from a piece of evidence: a path meant to point at
where something *is* gets pulled through; a path quoted for what was true at
a time is left standing and dated. Note per corrected body what was pulled
through and when, so the next pass can tell a checked body from one that
merely looks right. Closed issues are out of scope - they guide nobody.
Renames are not the only thing that ages an issue text. A page a body cites
can vanish from `kb/` (`wikitool search` against the cited titles is the
second pass), and an old body can carry private infrastructure detail into
what is now a public tracker - both found in the same issue, both worth the
same look.
## Citing an issue in the repo
**No file `dist export` ships may cite an issue number.** The board is reachable only from the
origin repo, and this very file - the only one that says where it lives - is pruned along with
the rest of `instructions/dev/`. A "#66" that survives into a distributed instance is therefore
worse than a dead link: the reader cannot resolve it *and* cannot tell that it is unresolvable,
so a rule appears to rest on evidence nobody can produce. `instructions/CONTRACT.md` § "Writing
an instruction" asks the opposite ("self-contained enough for an agent with no prior context"),
and an issue number is the exact counter-example to it.
Which is the same wayfinder/evidence split as in the section below, applied one layer out - but
both halves land in the same place here:
- **A wayfinder** ("see #66 for the reasoning") is resolved: the reasoning goes into the text,
and the number goes.
- **A piece of evidence** ("removed in #66") is dated in words instead - "removed when the
schema default was dropped". The sentence carries itself, and the number stays reachable
through `git blame` -> the commit message, which names the issue anyway.
Where a pointer is genuinely worth having *here* and would leave nothing behind in words, keep
it in a `<!-- dist:strip-start/end -->` block ([instructions/CONTRACT.md](../CONTRACT.md)
§ `instructions/dev/`): visible in this repo, removed on export. Two passages use it today.
`tools/**/*.py` is deliberately outside all of this. A code comment addresses whoever edits that
line, and that only ever happens in the origin repo, because `dist export` prunes the
`stack-dev` skill together with this directory; a distributed `tools/` tree is runtime
machinery, not reading material. The same holds for `.gitignore` and `tools/.coveragerc` -
config, not documentation.
`docs verify` enforces the rule (below).
## What no tool checks
`wikitool` does not know this tracker exists, and should not learn. It ships to
instances that have no issues at that URL, while this file and the workflow it
describes are pruned by `dist export` - a Gitea client inside the shipped tool
would be a dev-only dependency carried by every instance, to check a board none
of them have. The tracker is reachable only through the `gitea-mcp` server, in a
session, by an agent.
So there is no `docs verify` for the board. Nothing reports a closed issue whose
body still reads as open, a body that contradicts its own comments, an issue
missing one of the four mandatory labels, or a `status/incoming` stub that got
built as it stood. Every one of those is caught by a
session following this file, or not at all - which is the argument for the
sequence in step 7 being explicit about the order (body first, then close),
rather than leaving it to be inferred from step 2.
The one rule here that *is* checked is § Citing an issue in the repo, and it is
worth being clear about why that is not a contradiction. `docs verify`'s
`check_no_issue_references` compiles `#\d+` and reads the text
`dist_cmd.build_plan()` would write. It has no client, no URL and no notion of
an issue's state - it cannot tell an open issue from a closed one, or a real
number from an invented one. What it knows is that a shipped document is making
a reference its reader cannot follow, which is a property of the *document*, not
of the board. That is the line: a check may look at what this repo writes about
the tracker; none may look at the tracker.
## Decision points
- **Issue or changelog?** An issue is work that is *not done*. `CHANGES.md` is
what shipped. A finished change needs both: the entry, and the issue closed
with the reasoning.
- **Issue or `kb/` page?** An issue is about *this stack* and is ephemeral - it
closes, and it records a wish. A `kb/` page is verified knowledge that stays
true. Never put wiki content findings in an issue, and never file a work item
as a page.
- **Rewrite the body, or add a comment?** Rewrite whenever a reader of the body
alone would otherwise be misled - a changed decision, a dropped criterion, a
new constraint. A comment carries the changelog line for that rewrite, and
nothing else that a future session needs in order to act. Closing an issue is
always a rewrite - see step 7.
- **A `status/incoming` stub looks trivially implementable?** Work it out anyway.
"Trivial" is a judgement about the sentence, and the sentence is the part the
human wrote down cheaply; what it omits is not visible from it. The elaboration
of an obvious stub is short - that is the argument for doing it, not for
skipping it.
- **An old issue carries only `prio/` and `size/`?** Complete it to all four
when you touch it, rather than in a sweep. The board reaches the new scheme
issue by issue, as each is picked up.