Compare commits

...

2 Commits

Author SHA1 Message Date
torben 9b461421e8 feat: Korpus-Kuratierungsrichtlinie - Floors und Leitplanke fuer reaktive Fixes (4.2.0, #28)
CI / verify (push) Successful in 52s
Release / release (push) Successful in 37s
Files changed:
- CHANGES.md
- VERSION
- instructions/dev/corpus-policy.md
- instructions/dev/stack-dev/SKILL.md
2026-09-03 08:11:16 +02:00
torben 41f5dfe1cd docs: Issue-Body ist das Plan-File - fortlaufend aktuell, Abschluss ist die letzte Aktualisierung (4.1.2, #44)
CI / verify (push) Successful in 52s
Release / release (push) Successful in 38s
Files changed:
- CHANGES.md
- VERSION
- instructions/dev/issue-tracking.md
- instructions/dev/stack-dev/SKILL.md
2026-09-03 06:39:16 +02:00
5 changed files with 274 additions and 22 deletions
+87
View File
@@ -20,6 +20,93 @@ their date-only headings.
---
## 4.2.0 - 2026-09-03 - Korpus-Kuratierungsrichtlinie: Untergrenzen und Leitplanke für reaktive Fixes
**Author:** Torben Nehmer
Ein Demo-Korpus will klein und stabil sein, ein Testbett groß, unordentlich und in Bewegung -
dieses Repo verlangt seit der Veröffentlichung beides vom selben `kb/` (Gitea #28). Die Sitzung
vom 2026-09-02 hatte Fixture, `--with-demo` und ein zweites Repo bereits verworfen; offen blieb
nur, wie kuratiert "kuratiert genug" heißt und welche Leitplanke reaktive Fixes bekommen.
**Neu:** `instructions/dev/corpus-policy.md`. Fünf Untergrenzen, jede mit einer bestehenden
`wikitool`-Prüfung messbar, keine davon durch neuen Tool-Code: jeder Seitentyp und jeder
deklarierte Subtyp mit mindestens einer Seite, mindestens fünf Seiten mit mindestens drei
Quellen, ein bis zehn Orphan-Seiten, im Schnitt mindestens vier ausgehende Wikilinks pro Seite.
Gemessen am 2026-09-03: 181 Seiten, alle Typ-/Subtyp-Floors erfüllt, 12 Seiten mit ≥3 Quellen, 3
Orphans, Ø 6,2 ausgehende Links - der Korpus war bereits groß genug, ohne dass eine einzige
Seite eigens dafür angelegt werden musste. Eine Untergrenze wird nie durch eine erfundene Seite
gefüllt, sondern durch eine echte Quelle beim nächsten passenden Ingest - Invariante 3 gilt
unverändert.
**Die Leitplanke für reaktive Fixes** unterscheidet drei Stufen: punktuelle Änderungen (immer
erlaubt, gewöhnliche Arbeit), korpusweite Änderungen (nur geplant, mit eigenem Issue und
`work/`-Run - trifft eine Session das Mass-Update-Gate während sie etwas anderes tat, holt sie
sich nicht den `--confirm`-Token, sondern stoppt und legt ein Issue an) und reaktive Eingriffe
in Korpusinhalt, um einen Test grün zu machen oder einen Tool-Bug zu umgehen (nie erlaubt,
Invariante 7). Das Verhältnis zu `kb_dir`/`raw_dir` und `test_pipeline_l0.py` bleibt wie im
ursprünglichen Befund: kleiner, isolierter Fall in der Fixture, großer, vernetzter Fall in
`kb/` - keine Fixture-Extraktion aus dem Korpus.
Dev-only und rein additiv - kein Feld, kein Kommando, keine Datei außerhalb von
`instructions/dev/` ändert sich, daher `--minor` ohne `--breaking`.
**Migration:** none required.
Berührt: `instructions/dev/corpus-policy.md` (neu),
`instructions/dev/stack-dev/SKILL.md` (Schritt 2, Routing-Zeile).
---
## 4.1.2 - 2026-09-03 - Issue-Abschluss ist ein Body-Rewrite, nicht nur ein Kommentar
**Author:** Torben Nehmer
Aufgefallen beim Schließen von #44: der Abschlussbericht stand als Kommentar da, der Body
darunter weiterhin als offene Arbeit — Abschnitt „Zu entscheiden" über eine längst getroffene
Entscheidung, ungehakte Checkliste, Präsens über einen Defekt, den es nicht mehr gab.
Die Regel gab es dafür schon: Schritt 2 von `instructions/dev/issue-tracking.md` sagt, der Body
ist die aktuelle Wahrheit und wird umgeschrieben, wenn sich der Stand ändert. Nur ließ die
Formulierung offen, *wann* — und Schritt 7 („Close with what actually happened") war vollständig
erfüllbar, ohne den Body anzufassen. Ein Abschlussbericht im Kommentar fühlt sich beim Schreiben
vollständig an; dass der Body dabei zurückbleibt, merkt erst der nächste Leser.
**Schritt 2 ist deshalb schärfer geworden: der Body ist das Plan-File dieses Stacks.** Dasselbe,
was das Plan-Dokument eines Harness ist, und genauso gepflegt — fortlaufend, sobald etwas darin
nicht mehr stimmt, nicht am Ende. Der Maßstab ist der Abbruch, nicht der Meilenstein: eine
Session kann jederzeit enden, und was der Body in diesem Moment sagt, ist die vollständige
Übergabe. Eine frische Session muss zu **jedem** Zeitpunkt allein aus dem Body weiterarbeiten
können, ohne Kommentare rückwärts zu lesen und ohne einen Menschen, der es neu erklärt. Entschieden
ersetzt die Frage, erledigt hakt das Kriterium ab, verworfen steht mit Begründung dort, wo das
Kriterium stand.
Schritt 7 ist damit kein Sonderakt mehr, sondern die letzte dieser Aktualisierungen: erst Body
auf den Endstand, dann schließen, dann die Changelog-Zeile aus Schritt 3. Wer Schritt 2 befolgt
hat, ist fast fertig; wer nicht, zahlt die ganze Schuld im schlechtesten Moment — der
geschlossene Body ist die Fassung, die danach alle lesen und niemand mehr aufsucht. #44 steht
als Beispiel drin.
Schritt 3 zieht die Konsequenz: **ein Kommentar pro Session-Umfang, nicht pro Edit.** Ein
fortlaufend gepflegter Body mit einem Changelog-Kommentar je Änderung wäre Lärm; triviale Pflege
braucht gar keinen. Der `stack-dev`-Skill sagt es beim Aufgreifen mit, weil dort die Entscheidung
fällt, ob eine Session den Body überhaupt anfasst.
**Und die ehrliche Antwort auf die Frage nach dem Tooling: es gibt keins, und es soll keins
geben.** `wikitool` kennt diesen Tracker nicht. Es wird an Instanzen ausgeliefert, die unter
dieser URL keine Issues haben, während `instructions/dev/` von `dist export` gepruned wird —
ein Gitea-Client im ausgelieferten Tool wäre eine Dev-Abhängigkeit, die jede Instanz mitträgt,
um ein Board zu prüfen, das keine von ihnen hat. Der Tracker ist ausschließlich über
`gitea-mcp` erreichbar, also in einer Session, durch einen Agenten.
Kein `docs verify` fängt also einen geschlossenen Issue, dessen Body offen klingt, einen Body,
der seinen eigenen Kommentaren widerspricht, oder ein fehlendes Pflichtlabel. Das steht jetzt
als eigener Abschnitt „What no tool checks" in der Instruktion — nicht als Bedauern, sondern als
Begründung dafür, warum die Reihenfolge in Schritt 7 ausgeschrieben ist statt aus Schritt 2
erschlossen zu werden.
---
## 4.1.1 - 2026-09-03 - Testisolation: kb_dir repointet config.ROOT, lint löst Kollektionen gegen den übergebenen Baum auf
**Author:** Torben Nehmer
+1 -1
View File
@@ -1 +1 @@
4.1.1
4.2.0
+94
View File
@@ -0,0 +1,94 @@
---
type: types/instruction.md
name: corpus-policy
description: What "curated enough" means for kb/ when it is demo and testbed at once, the measurable floors that define it, and what a reactive fix to the corpus may and may not do.
---
# Keep kb/ curated enough to develop against, without a second corpus
This instance runs one `kb/` for two purposes at once: a public demo and the testbed this stack
is developed against. There is deliberately no fixture corpus, no `--with-demo` export, and no
second repository - see Gitea #28. The corpus's size and shape are set by what targeted
development needs, not by a synthetic fixture size or a demo aesthetic.
## When to run
- Before judging whether the corpus can exercise a change under development - ranking, index
scaling, orphan detection, a new label, a new type-spec.
- Before a reactive fix touches `kb/` content rather than the failing code - the floors below
are what decides whether the fix may proceed as-is.
- Picking up Gitea #28 or #30, or any issue that references this file.
## The floors
Each is mechanically checkable with an existing `wikitool` command; none needs new tool code.
A floor exists to keep some class of bug observable, not to describe an aesthetic target - so
when a session is about to make one of these numbers *worse*, that is the signal to stop and
think, not a number to defend for its own sake.
| Floor | Check | Why this number |
|---|---|---|
| Every page type has ≥1 page | `wikitool search --field type=types/<t>.md` | A type with zero pages means its schema, its collection contract and its lint rules are unexercised |
| Every declared subtype has ≥1 page | `wikitool search --field <x>_type=<v>` | Same reasoning, one level down - `entity_type`, `concept_type`, `source_type` |
| ≥5 pages corpus-wide with ≥3 `sources:` entries | one-off script, see below | Provenance fan-in - multiple sources backing one claim - is a real case only a handful of pages exercise; fewer than 5 and a provenance-index bug can hide |
| Orphan pages (no inbound link) between 1 and 10 | `wikitool lint` | Zero orphans makes orphan detection itself unobservable; more than 10 means the corpus stopped being curated |
| Average outbound wikilinks per page ≥4 | one-off script, see below | Below this, ranking and graph-traversal work has too little structure to exercise |
A floor is a lower bound only. There is no upper bound on page count or on any of these numbers
except the orphan ceiling above - a corpus that outgrows these floors through real ingests is
not a problem this file cares about.
**Measured 2026-09-03** (see Gitea #28): 181 pages, 14/14 types and subtypes covered, 12 pages
with ≥3 sources, 3 orphans, 6.2 average outbound links. All floors held without any manufactured
content - the corpus was already big enough when the question was asked.
A type or subtype sitting at exactly the floor - one page - shows no set-level bugs, only that
the type is *reachable*. That is a soft target for the next `wiki-ingest` that happens to
produce a matching page, never a reason to write one: filing an unsourced page to clear a floor
is exactly what AGENTS.md invariant 3 forbids, floor or no floor. The same holds for an
authorised link label with zero live uses (`wikitool xref` reports these) - fill it when a real
edge calls for it, never manufacture one to exercise the label.
To check the two floors without a dedicated command, walk `kb/**/*.md` (excluding
`INDEX.md`/`COLLECTION.md`/`CONTRACT.md`/`CONVENTIONS.md`), parse frontmatter, and: count pages
whose `related:` array (resolved against page titles) has ≥3 entries for outbound density; count
`sources:` array length ≥3 for the provenance floor. `wikitool search` and `wikitool lint`
cover everything else in the table.
## What a reactive fix may do to kb/ content
Three tiers, by how much of the corpus a change touches:
1. **Pointwise - always allowed.** Creating, updating, renaming or deleting a single page
through the normal tools (`new`, `touch`, the page-lifecycle procedure), below the
Mass-Update Gate's threshold. This is ordinary work and needs no special permission.
2. **Corpus-wide - planned only, never reactive.** A migration, a vocabulary sweep, a bulk
`touch` across many pages. This needs its own issue and, per `work/CONTRACT.md`, a `work/`
run - never a same-session reaction to whatever the session was originally doing. If a
session hits the Mass-Update Gate (exit 42, see `instructions/gates.md`) while working on
something else, it does not fetch the `--confirm` token to push through: it stops, opens an
issue for the corpus-wide change, and finishes the original task without it.
3. **Reactive - never allowed.** Deleting or reshaping a page to make a failing test pass;
restructuring corpus content to route around a tool bug (AGENTS.md invariant 7); using
`kb/` as a scratch surface for a tool experiment. If a stack change under development needs a
corpus shape that does not exist, build it as a pytest fixture (see the next section) -
never manufacture it in `kb/`.
## Relationship to the test fixtures
`tools/chemenu/tests/conftest.py`'s `kb_dir`/`raw_dir` fixtures and `test_pipeline_l0.py` cover
the **small, isolated** case: a handful of pages, built fresh per test, hermetic. `kb/` covers
the **large, connected** case: 181+ pages, grown link density, real provenance history that no
per-test fixture reconstructs economically. The cut: if a `tmp_path` tree can reproduce what the
test needs, it belongs in a fixture; if the test needs density or scale that only a grown corpus
has, it belongs against `kb/`. Neither absorbs the other's job - see
[testing-conventions.md](testing-conventions.md).
## Decision points
- **A floor would be violated by an in-progress change - is that a blocker?** Only for the
orphan ceiling and the type/subtype floors, since those two can go to zero. The density and
provenance floors move gradually with ordinary ingests and are not gating on any single
session.
- **Corpus is "too small" for a feature under development?** That is not this file's problem to
solve by adding pages - see tier 3 above. Either the feature waits for a real ingest to supply
the shape, or it gets a pytest fixture.
+83 -18
View File
@@ -28,8 +28,12 @@ issues at that URL, which is exactly why `dist export` excludes
an assumption nobody has checked, a decision that needs the user.
- Picking an issue up: before doing anything else, read the body as the current
spec, and re-label it if the ground has moved since.
- **While working on one:** the body is updated as the state moves, not at the
end (step 2). A session that is interrupted leaves the body as its handover.
- Prioritising: deciding what to pick up next, or re-labelling after the ground
moved.
- Closing one: the body is rewritten to its final state first, and only then
closed (step 7).
## Steps
@@ -38,23 +42,47 @@ issues at that URL, which is exactly why `dist export` excludes
specific files or commands involved. An issue that only makes sense to
whoever wrote it is a note, and notes were the problem.
2. **Treat the body as the current truth, not as a historical first post.**
Work on one issue spans several sessions, often weeks apart, and the body is
the only thing that connects them: a session opening the issue must be able
to reconstruct what is decided and what is still open from the body alone,
without a human re-explaining it. So when the state changes, **rewrite the
body** - do not append to a text that has become wrong. An additively grown
log forces every later reader to reconstruct the current state by filtering
the whole history.
2. **The body is the working state, not a historical first post - keep it
current as you go.** It is this stack's plan file: the same thing a harness's
own plan document is, and it is maintained the same way. Not written once,
not brought up to date at the end, but **updated whenever something in it
stops being true** - a decision made, a criterion met, an approach ruled out,
a new constraint found.
The test is an abort, not a milestone. A session can end at any moment - an
interrupt, a context limit, a crash, a human walking away - and whatever the
body says at that instant is the entire handover. So the standard is: **at
every point, a fresh session must be able to open the body and pick the work
up from there**, without a human re-explaining it and without reading back
through the comments. If the body would mislead someone who read it right
now, it is already out of date, whether or not the work is finished.
That means updating *during* the work, not only at its end:
- a decision gets made → the decision and its reasoning replace the question
- an acceptance criterion is done → tick it, in the same session that did it
- something turns out differently than the issue assumed → the assumption is
corrected where it stands, not contradicted three paragraphs later
- work is deferred or dropped → say so, with the reason, where the criterion is
**Rewrite, never append.** Do not add to a text that has become wrong: an
additively grown log forces every later reader to reconstruct the current
state by filtering the whole history, which is the exact cost the body exists
to remove. Comments carry the history (step 3); the body carries the state.
Body rewrites and comments are an LLM session's job. A human normally
touches only labels and metadata directly.
3. **Comment a changelog, never a copy.** Every body rewrite gets one short
comment naming only what changed against the previous state - what is new,
what is gone, what was corrected. Do not snapshot the old body into a
comment: a full copy per revision forces a human to diff two prose texts,
which is not a readable history, only another copy.
3. **Comment a changelog, never a copy.** A body rewrite gets one short comment
naming only what changed against the previous state - what is new, what is
gone, what was corrected. Do not snapshot the old body into a comment: a full
copy per revision forces a human to diff two prose texts, which is not a
readable history, only another copy.
One comment per *session's worth* of change, not per edit. Step 2 asks the
body to be kept current continuously, and a comment for every tick would bury
the board in noise; the changelog line summarises what that session moved.
Trivial upkeep - a typo, a tightened sentence - needs no comment at all.
```
**Changelog:** Decision 2 tightened - `kind/` may now change over an
@@ -127,10 +155,46 @@ issues at that URL, which is exactly why `dist export` excludes
answered can drop a size and move `kind/decision` to `kind/build`. Silent
re-labelling is how a board stops meaning anything.
7. **Close with what actually happened**, not with a commit hash alone: which
proposals were implemented, which were deliberately left out and why, and
what was verified. The issue is the only place that record survives - a
changelog entry says what changed, not what was decided against.
7. **Closing is the last body update, not a comment.** If step 2 was followed
the body is already nearly there, and closing only settles what the final
run established. If it was not, closing is where the whole debt comes due -
and it comes due at the worst moment, because a closed body is the version
everyone reads afterwards and nobody revisits.
Either way the body reaches its final state *before* the issue closes:
proposals that were decided read as decided, a "to decide" section has become
the decision with its reasoning, acceptance criteria are ticked or struck with
a reason, and what was verified is named. Then close, with the one-line
changelog comment step 3 asks for.
Record what actually happened, not a commit hash alone: which proposals were
implemented, which were deliberately left out and why, and what was verified.
The issue is the only place that record survives - a changelog entry says
what changed, not what was decided against.
**A closing report in a comment does not satisfy this.** It reads as
complete to whoever writes it and leaves a body still phrased as open work:
unticked boxes, an undecided decision section, present tense about a defect
that no longer exists. #44 closed exactly that way, with a thorough comment
above a body that still asked for a decision that had already been made and
shipped. Nothing mechanical catches it (see below), which is why it is a step
rather than a habit.
## What no tool checks
`wikitool` does not know this tracker exists, and should not learn. It ships to
instances that have no issues at that URL, while this file and the workflow it
describes are pruned by `dist export` - a Gitea client inside the shipped tool
would be a dev-only dependency carried by every instance, to check a board none
of them have. The tracker is reachable only through the `gitea-mcp` server, in a
session, by an agent.
So there is no `docs verify` for the board. Nothing reports a closed issue whose
body still reads as open, a body that contradicts its own comments, or an issue
missing one of the four mandatory labels. Every one of those is caught by a
session following this file, or not at all - which is the argument for the
sequence in step 7 being explicit about the order (body first, then close),
rather than leaving it to be inferred from step 2.
## Decision points
@@ -144,7 +208,8 @@ issues at that URL, which is exactly why `dist export` excludes
- **Rewrite the body, or add a comment?** Rewrite whenever a reader of the body
alone would otherwise be misled - a changed decision, a dropped criterion, a
new constraint. A comment carries the changelog line for that rewrite, and
nothing else that a future session needs in order to act.
nothing else that a future session needs in order to act. Closing an issue is
always a rewrite - see step 7.
- **An old issue carries only `prio/` and `size/`?** Complete it to all four
when you touch it, rather than in a sweep. The board reaches the new scheme
issue by issue, as each is picked up.
+9 -3
View File
@@ -43,15 +43,21 @@ stack development happens in the origin repo instead (see AGENTS.md's routing li
engineering, memory and deploy-time learning; consult before a design decision in those
areas.
[issue-tracking.md](../issue-tracking.md) - open work lives in Gitea issues, one per work
package, labelled `area/`, `kind/`, `prio/` and `size/`, with the body kept as the current
truth rather than as a first post. There is no `TODO.md`. Read it before filing something
for later, before editing an issue, or before deciding what to pick up next.
package, labelled `area/`, `kind/`, `prio/` and `size/`. There is no `TODO.md`. **The body
of the issue you are working on is this session's plan file:** keep it current as the state
moves, not at the end, so an interrupted session leaves a body the next one can resume from.
Read it before filing something for later, before editing or closing an issue, or before
deciding what to pick up next.
[testing-conventions.md](../testing-conventions.md) - the suite runs against a deliberately
empty machine; what the autouse fixture already neutralizes, and what a test still has to
establish itself. Read it before adding or changing a test.
[version-parts.md](../version-parts.md) - which part a change bumps: the drop-in test, the
catalogue of breaks that cross the compatibility boundary with `kb/` untouched, and what to
put in front of the user before a breaking bump. Read it before step 3.
[corpus-policy.md](../corpus-policy.md) - what "curated enough" means for the shared
demo/testbed `kb/`, the measurable floors that define it, and what a reactive fix may and may
not do to corpus content. Read it before judging whether the corpus can exercise a change, or
before any fix that would touch `kb/` content.
More instructions are added here incrementally as stack-development needs come up - this
list grows without needing this skill file to change shape.
3. **Raise the version, if the change ships.** A change under `tools/`, `types/`,