Compare commits

...

1 Commits

Author SHA1 Message Date
torben 177c7e9ce8 feat: Prosa ist kein Identifier - Link-Taxonomie als Enum, generierte Regionen mit Markern (4.0.0)
CI / verify (push) Successful in 55s
Release / release (push) Successful in 38s
Files changed:
- .gitea/workflows/ci.yml
- AGENTS.md
- CHANGES.md
- VERSION
- instructions/CONTRACT.md
- instructions/link-taxonomy.md
- instructions/migrations/4.0.0-link-taxonomy.md
- instructions/setup-instance.md
- kb/CONTRACT.md
- kb/CONVENTIONS.md
- kb/CONVENTIONS.md.template
- kb/comparisons/COLLECTION.md
- kb/concepts/COLLECTION.md
- kb/entities/COLLECTION.md
- kb/sources/COLLECTION.md
- tools/CONTRACT.md
- tools/README.md
- tools/chemenu/blocks.py
- tools/chemenu/cli.py
- tools/chemenu/commands/cite_cmd.py
- tools/chemenu/commands/dist_cmd.py
- tools/chemenu/commands/docs_verify.py
- tools/chemenu/commands/doctor.py
- tools/chemenu/commands/links_cmd.py
- tools/chemenu/commands/migrate_cmd.py
- tools/chemenu/commands/new_page.py
- tools/chemenu/commands/page_ops.py
- tools/chemenu/commands/run_budget.py
- tools/chemenu/commands/xref.py
- tools/chemenu/conventions.py
- tools/chemenu/corpus_diff.py
- tools/chemenu/frontmatter_io.py
- tools/chemenu/kb_collections.py
- tools/chemenu/kb_state.py
- tools/chemenu/links.py
- tools/chemenu/lint_core.py
- tools/chemenu/provenance.py
- tools/chemenu/sections.py
- tools/chemenu/tests/conftest.py
- tools/chemenu/tests/test_blocks.py
- tools/chemenu/tests/test_cite_cmd.py
- tools/chemenu/tests/test_conventions.py
- tools/chemenu/tests/test_dist_cmd.py
- tools/chemenu/tests/test_doctor.py
- tools/chemenu/tests/test_migrate_cmd.py
- tools/chemenu/tests/test_new_page.py
- tools/chemenu/tests/test_pipeline_l0.py
- tools/chemenu/tests/test_types_cmd.py
- tools/chemenu/tests/test_xref.py
- types/concept.schema.yaml
- types/entity.md
- types/entity.schema.yaml
- types/instruction.schema.yaml
- types/type-spec.md
- work/link-taxonomy-migration/README.md
- work/link-taxonomy-migration/plan.md
2026-09-02 18:39:22 +02:00
56 changed files with 2692 additions and 750 deletions
+1 -1
View File
@@ -228,7 +228,7 @@ jobs:
# contracts are adopted verbatim - the shipped text is a working # contracts are adopted verbatim - the shipped text is a working
# default, unlike a personalization file. # default, unlike a personalization file.
grep -v 'wikitool:template-unfilled' kb/CONVENTIONS.md.template > kb/CONVENTIONS.md grep -v 'wikitool:template-unfilled' kb/CONVENTIONS.md.template > kb/CONVENTIONS.md
for template in kb/*/COLLECTION.md.template; do for template in kb/*/COLLECTION.md.template types/*.template; do
cp "$template" "${template%.template}" cp "$template" "${template%.template}"
done done
python3 -m venv tools/.venv python3 -m venv tools/.venv
+1 -1
View File
@@ -81,7 +81,7 @@ What a file is called says who it is for and how it is loaded. This is a rule, n
| `kb/<collection>/COLLECTION.md` | Agents | When writing in that collection. Instance-owned in the same way, and declares in frontmatter which profile it adopted | | `kb/<collection>/COLLECTION.md` | Agents | When writing in that collection. Instance-owned in the same way, and declares in frontmatter which profile it adopted |
| `instructions/<name>.md` | Agents | By link, or on explicit request | | `instructions/<name>.md` | Agents | By link, or on explicit request |
| `instructions/<name>/SKILL.md` | Agents | By the harness, once published | | `instructions/<name>/SKILL.md` | Agents | By the harness, once published |
| `types/<name>.md` | Agents + validator | Via `tools/wikitool types describe` | | `types/<name>.md` | Agents + validator | Via `tools/wikitool types describe`. Split by `root:`: a page type-spec (`root: kb`) belongs to the instance and ships as `.template`; one describing a stack artifact ships verbatim |
| `INDEX.md` | Both | Generated - never hand-edited | | `INDEX.md` | Both | Generated - never hand-edited |
A stage may carry both a `README.md` and a `CONTRACT.md`: different readers, different A stage may carry both a `README.md` and a `CONTRACT.md`: different readers, different
+73
View File
@@ -18,6 +18,79 @@ heading, and `wikitool docs verify` refuses a tree whose `VERSION` and newest
versioned entry disagree. Entries below `0.1.0` predate versioning and keep versioned entry disagree. Entries below `0.1.0` predate versioning and keep
their date-only headings. their date-only headings.
---
## 4.0.0 - 2026-09-02 - Prosa ist kein Identifier: Link-Taxonomie als Enum, generierte Regionen mit Markern
**Author:** Torben Nehmer
**Breaking Change:** Beziehungslabel sind Enum-Werte in related: statt Freitext im Body-Bullet, toolgefuehrte Abschnitte liegen zwischen Marker-Paaren statt hinter ihrer Ueberschrift, und xref add schreibt nur noch eine Kante statt beider Richtungen. tools/chemenu/sections.py ist geloescht. Eine bestehende Instanz muss sections: in kb/CONVENTIONS.md auf links/footnotes umstellen, outbound: in jede COLLECTION.md eintragen, die {section.*}-Variablen aus ihren Page-Type-Templates entfernen und den Korpus umstellen - sonst scaffoldet new die Variablen woertlich in neue Seiten. Ablauf: instructions/migrations/4.0.0-link-taxonomy.md
Der Stack benutzte an drei Stellen **Prosa als Identifier**, und jede hat messbar etwas
gekostet. Die Überschrift eines Abschnitts war seine Adresse (`^## Beziehungen$`), was die
KB-Sprache zu einer Compiler-Konstante machte *und* das Ende der Region zur Schätzung - sie lief
bis zur nächsten Überschrift, davor bis zum Dateiende, und hat auf acht Seiten still Inhalt
gelöscht. Das Beziehungslabel stand nur im Body-Bullet, also konnte nichts das Vokabular prüfen:
gemessen am Korpus **152 distinkte Label in 337 Bullets** gegen dreizehn dokumentierte, 102 davon
genau einmal vorkommend. Und `xref add` spiegelte jede Kante, was `## Siehe auch` mit 555
Bullets ohne Label füllte - 353 davon beweisbar redundant.
**Was jetzt Identifier ist.** Eine Region liegt zwischen `<!-- wikitool:links -->` bzw.
`<!-- wikitool:footnotes -->` und wird vollständig aus dem Frontmatter gerendert, Überschrift
eingeschlossen. Ein Label ist ein Maschinenwert in `related:` (`- depends-on: Hermes`), gezogen
aus `instructions/link-taxonomy.md` und **pro Ziel autorisiert von der Quell-Collection**
(`outbound:` im `COLLECTION.md`, Commonplaces ADR-019). Der Body-Bullet ist eine Darstellung
dieser Daten, nicht ihr zweiter Aufbewahrungsort.
**Gelöscht, ersatzlos:** `tools/chemenu/sections.py` komplett, `heading_re`, der
Alias-Mechanismus, `PRE_CONVENTIONS_NAMES`, `cite_block_heading`, `provenance.__getattr__`, die
`{section.*}`-Template-Variablen, `xref`s Abschnittssuche. Kein Überschriftentext liegt mehr in
Python - bis auf zwei kosmetische Fallbacks, und die sind harmlos geworden: der Marker trägt die
Identität, also rendert ein falscher Default falsche Wörter statt Struktur zu zerlegen, und der
nächste Write repariert es.
**Kanten sind direktional, und das war keine Geschmacksfrage.** Die per-Collection-Autorisierung
ist mit einer automatisch gespiegelten Gegenkante logisch unverträglich: die Spiegelhälfte
entsteht in einer Collection, deren Regeln der Autor nie gelesen hat. Entweder schriebe das
Werkzeug unautorisierte Kanten, oder die Regel "die Quellcollection entscheidet" löst sich auf.
Der Navigationseinwand wird dabei *besser* beantwortet als vorher: `wikitool links show --page`
berechnet die Eingangssicht über den Korpus, vollständig und ohne Pflege, und das gerenderte
Bullet ist ein gewöhnlicher `[[wikilink]]` - ein Backlink-Panel zeigt es ohnehin. Die erzwungene
Gegenkante garantierte nie Vollständigkeit, nur dass jemand daran gedacht hat.
**Der Orphan-Check meldet dadurch mehr,** und das ist die Prüfung bei der Arbeit: sie misst jetzt
Erreichbarkeit statt "ist `xref` gelaufen".
**`obligation:` trennt zwei Achsen, die vorher eine waren.** `migration_kind:` sagt *wie*
gearbeitet wird, neu `obligation: required|offered` *ob* überhaupt. Eine `offered`-Migration ist
ein Angebot für eine Datei, die der Instanz gehört - sie blockiert nie, steht nicht in der Kette,
und `migrate done` verbucht sie im Ledger, **ohne** `kb_version` zu bewegen. Genau daran hing ein
Entwurfsfehler, den erst der Test gezeigt hat: Offers gegen `kb_version` zu filtern hätte jede
Offer verschwinden lassen, sobald irgendein unbeteiligter Pflichtschritt lief. Dazu ist die
Erkennungshälfte aktiviert, die seit ihrer Einführung ungelesen dalag - die sha256 pro Datei in
`.wikitool-release.json` beantwortet jetzt "editiert oder nur empfangen", also ob eine Offer
kopiert werden darf oder von Hand abgeglichen werden muss.
**`types/` teilt sich entlang `root:`.** `root: kb` heißt Wissensseite heißt Instanz: die vier
Page-Type-Specs samt Schemas gehen als `.template`, `instruction`/`lint-report`/`type-spec`
verbatim. Damit ist die deutsche Prosa in jenen vier Dateien **korrekt statt Migrationsschuld** -
es war die richtige Sprache an einem Ort mit falsch deklariertem Eigentümer. Was der Stack von
der Type-Schicht noch verlangt, ist eine Zeile: ein Type-Spec `name: source`, dessen Schema
`raw_files` fordert. `STACK_REQUIRED_COLLECTIONS` entfällt als separate Liste - die pflichtige
Collection wird aus dem `base_dir` dieses Typs abgeleitet.
**Warum das MAJOR ist.** Vorwärts: `sections:` hat eine andere Form, `outbound:` fehlt, und die
in 3.0.0 übernommenen Page-Type-Templates enthalten `{section.*}`-Variablen, die es nicht mehr
gibt - `new` schriebe sie wörtlich in neue Seiten. Rückwärts: 4.0.0 schreibt gelabelte Kanten,
die 3.0.0s Schema als `type: string` ablehnt. Beide Hälften des Drop-in-Tests fallen.
**Der Korpus dieser Instanz ist noch nicht umgestellt.** Diese Version liefert die Maschinerie;
`lint` meldet die 480 noch ungelabelten Kanten als Findings, nicht als Fehler, weil das genau das
Fenster ist, für das `.wikitool-kb.json` existiert. `malformed_edges` und `unbalanced_markers`
sind dagegen sofort hart - keines beschreibt eine unkonvertierte Seite, nur eine kaputte. Die
Beförderung der beiden anderen kommt, wenn der Korpus sie bestehen kann.
--- ---
## 3.0.0 - 2026-09-02 - Autorenkonventionen nach Eigentum geschnitten: kb/CONVENTIONS.md, deklarierte Collections ## 3.0.0 - 2026-09-02 - Autorenkonventionen nach Eigentum geschnitten: kb/CONVENTIONS.md, deklarierte Collections
+1 -1
View File
@@ -1 +1 @@
3.0.0 4.0.0
+22 -2
View File
@@ -65,11 +65,31 @@ catalogue of authoring profiles an instance may adopt into its own `kb/CONVENTIO
## `instructions/migrations/` ## `instructions/migrations/`
A content migration is a Manual instruction with two extra frontmatter fields A content migration is a Manual instruction with three extra frontmatter fields
(`types/instruction.schema.yaml`): `migrates_to:`, the stack version whose content shape it (`types/instruction.schema.yaml`): `migrates_to:`, the stack version whose content shape it
produces, and `migration_kind:` (`mechanical` | `assisted`). It lives at produces; `migration_kind:` (`mechanical` | `assisted`); and `obligation:`
(`required` | `offered`, default `required`). It lives at
`instructions/migrations/<version>-<slug>.md`. `instructions/migrations/<version>-<slug>.md`.
`migration_kind:` and `obligation:` are **two axes, not one**. The first says how the work is
carried out, the second whether it has to happen at all:
| `obligation:` | Means | `migrate status` |
|---|---|---|
| `required` | The content must reach the new shape or it no longer fits the machinery | Counted as outstanding; `migrate done` advances `kb_version` through it, in chain order |
| `offered` | A file the instance owns still works as it is, and the stack proposes a better default | Listed separately, never blocks, no ordering rule. `migrate done` records it in the applied ledger and leaves `kb_version` where it is |
Keeping them apart is what stops `migrate status` crying wolf: an instance nagged about an
improvement it declined stops reading the nag that means its content no longer fits its
machinery. And because taking an offer deliberately does not move the version, the **applied
ledger** - not `kb_version` - is what makes an offer stop being offered; without that record
there is no way to tell a taken offer from an ignored one.
An `offered` migration is what makes an instance-owned file upgradeable at all. `dist export`
records a sha256 per shipped file in `.wikitool-release.json`, so `migrate status` can say which
of those files the instance edited and which it merely received - the first have to be
reconciled by a person, the second can simply be copied over.
The tier fits exactly: a migration must never be picked up implicitly - it rewrites the corpus - The tier fits exactly: a migration must never be picked up implicitly - it rewrites the corpus -
and it is referenced by nothing, because `tools/wikitool migrate status` finds it by reading the and it is referenced by nothing, because `tools/wikitool migrate status` finds it by reading the
directory and comparing `migrates_to:` against this instance's `kb_version`. That is also why directory and comparing `migrates_to:` against this instance's `kb_version`. That is also why
+199
View File
@@ -0,0 +1,199 @@
---
type: types/instruction.md
name: link-taxonomy
description: The link-label catalogue - every relationship label a page may declare in related:, grouped by register, with the reader need each one names. A palette to authorise from in a COLLECTION.md, never binding on its own.
manual: true
---
# Pick a link label
**This page is a palette, not an enum.** It lists every label this stack ships with and what
each one asserts. What a page may actually *use* is decided by its own collection: each
`kb/<name>/COLLECTION.md` authorises a subset per destination, and `wikitool lint` checks
`related:` against that authorisation rather than against this file. A collection that
authorises six labels has six, however long this list gets.
A label is an **identifier, not prose**. It is written into `related:` as a machine value and
rendered verbatim into the page body, so it is never translated - not in a German wiki, not in
any other. Which words a page is *written* in stays [kb/CONVENTIONS.md](../kb/CONVENTIONS.md)'s;
this is not one of them.
## The invariant every label obeys
Every label completes, with the page carrying the link as the grammatical subject:
> `[source] <label> [target]`
The page containing the link asserts something **about** the target. `Hermes depends-on
PostgreSQL` reads correctly on Hermes' page; the same fact written on PostgreSQL's page is a
different label (`required-by`), not the same one pointing back. Omitted helper verbs ("is",
"a") are fine where they do not reverse the endpoints.
This is Commonplace's ADR-058, adopted wholesale, and it is what makes a label checkable rather
than a matter of taste: read the sentence out loud, and if it says the opposite of what you
meant, the label is wrong.
## Direction is authored, never mirrored
Each direction is a separate decision. A link back from the target is welcome when it
independently helps a reader *there* - and unnecessary when it does not. **Do not add a reverse
edge merely to mirror the first one.** The inbound view is rendered from the graph by
`index rebuild` and `search`, so a reader landing on the target sees what points at it whether
or not anyone wrote a second edge.
That is why most labels below have no inverse. Only two pairs do, because in each the reverse
direction is a genuine primary statement someone would write on its own: `depends-on` /
`required-by` and `runs-on` / `hosts`.
## When to run
Adding or changing a `related:` entry, authorising labels in a `COLLECTION.md`, or judging
whether a relationship is worth naming as a formal edge at all.
## Steps
1. **Decide whether this is an edge.** Not every mention is one. An edge is a reader aid: it
says *follow this if you need X*. A subject mentioned once in passing is prose with a
`[[wikilink]]`, not a declared relationship. Over-declaring is how a graph becomes a list of
everything adjacent to everything.
2. **Say the sentence.** `[this page] <label> [that page]`. If it reads backwards, you want the
other page to carry the edge, or a different label.
3. **Pick from the register that fits the pair**, below. Prefer the most specific label that is
true; fall back outward only when nothing fits.
4. **Check the collection authorises it** for that destination -
`kb/<name>/COLLECTION.md`'s `outbound:` block. If the label you want is not authorised and
should be, that is a collection-contract change, made deliberately, not a lint error to
route around.
5. **Write it with the tool**, never by hand:
```bash
tools/wikitool xref add --a "<This Page>" --b "<That Page>" --rel <label>
```
## The catalogue
### Operational
Concrete things and how they stand to one another - the register this instance runs on. Mostly
entity to entity.
| label | inverse | asserts |
|---|---|---|
| `depends-on` | `required-by` | cannot function without the target |
| `required-by` | `depends-on` | the target cannot function without this |
| `runs-on` | `hosts` | executes on the target as its substrate |
| `hosts` | `runs-on` | provides the substrate the target executes on |
| `uses` | — | employs the target at runtime, but survives without it |
| `produces` | — | emits the target as an artifact or data |
| `consumes` | — | reads the target as an artifact or data |
| `maintains` | — | carries the upkeep of the target |
| `owns` | — | is accountable for the target's existence and decisions |
`uses` versus `depends-on` is the distinction worth keeping sharp: if removing the target breaks
this thing, it is `depends-on`. `owns` versus `maintains`: accountability versus labour, and
they are often different people.
### Realization
How an idea becomes a running thing. Usually concept to entity or the reverse.
| label | asserts |
|---|---|
| `implements` | is a concrete realization of the target |
| `operationalized-from` | is the prescriptive form of the target's theory |
| `mechanism` | is the mechanism by which the target works |
| `procedure` | is the procedure for carrying out the target |
| `applies-when` | applies under the condition the target describes |
| `operates-on` | acts upon the target as its subject matter |
| `invokes` | calls the target as a step within itself |
### Conceptual
Inference and comparison between ideas.
| label | asserts |
|---|---|
| `extends` | develops the target's argument further |
| `grounds` | provides the basis the target rests on |
| `rests-on` | takes the target as its premise |
| `enables` | is the operational prerequisite that makes the target possible |
| `precondition` | must hold before the target applies |
| `exemplifies` | is an instance of the general claim the target makes |
| `abstracted-from` | generalizes from the target |
| `contrasts` | differs from the target in a way worth reading both for |
| `compares-with` | is weighed against the target on shared dimensions |
| `contradicts` | asserts something the target denies |
| `composition` | is composed of the target |
| `part-of` | is a component of the target |
`grounds` / `rests-on` is a genuine pair and both directions are primary statements; they are
listed separately rather than as inverses because either page may legitimately carry only its
own side.
### Lineage
Where something came from, and what replaced it.
| label | asserts |
|---|---|
| `supersedes` | replaces the target, which is now historical |
| `derived-from` | was produced from the target |
| `adapted-from` | was reworked from the target for a different purpose |
| `defined-in` | takes its definition from the target |
A superseded page is never deleted or rewritten - see the collection contract for
`kb/concepts/`.
### Evidence
The provenance register. Distinct from `sources:` and `[^cite-id]`, which are the *mechanical*
provenance path: these two are authored claims about how strongly something is backed.
| label | asserts |
|---|---|
| `evidenced-by` | is supported by the target as evidence |
| `is-evidence-for` | serves as evidence for the target's claim |
### Universal
| label | asserts |
|---|---|
| `see-also` | nothing more specific applies, and a reader here would still want the target |
**`see-also` is the last resort and should stay rare.** A collection where it is the commonest
label has a vocabulary problem, not a lot of loosely related pages. The previous vocabulary's
`verwandt mit` was exactly that, and it is the reason this catalogue exists.
## Extending it
Adding a label is a line of data, never a code change:
1. Add a row here, in the register it belongs to, with the sentence it completes.
2. Authorise it in the `COLLECTION.md` of every collection that may use it.
The registers are advisory groupings for readers, not a schema - nothing checks that a label is
used only within its register. Invent an intra-collection label the work needs and propose it
here afterwards; the architecture is deliberately loose, because the link theory is still
developing.
## Decision points
- **Two labels both fit?** Take the more specific one. If they are equally specific and mean
different things, the relationship is probably two edges.
- **The relationship reads better from the other page?** Write it there. Nothing is lost - the
inbound view renders it here.
- **You want a reverse edge for navigation?** You do not need one. That is what the rendered
inbound view is for, and it is complete in a way an authored mirror never was.
- **Nothing fits at all?** Use `see-also` and say so in the commit, or propose a label. Do not
stretch a label whose sentence reads false - a wrong edge is worse than a weak one, because
it is machine-readable and will be believed.
## Scope
Covers labels on `related:` edges between pages. Says nothing about `sources:` (the provenance
field, unlabelled by construction), `[^cite-id]` footnotes
([kb/CONTRACT.md](../kb/CONTRACT.md#provenance-and-citation)), or `tags:` (search keys, not
relationships).
@@ -0,0 +1,146 @@
---
type: types/instruction.md
name: 4.0.0-link-taxonomy
description: Move every relationship from free-text prose in a body bullet to a labelled edge in related:, and every tool-owned body region from heading-matching to a marker pair.
manual: true
migrates_to: 4.0.0
migration_kind: assisted
obligation: required
---
# Move relationships into the data, and generated regions behind markers (4.0.0)
Until 4.0.0 the stack used **prose as an identifier** in three places, and each one cost
something measurable:
| Was the identifier | Cost |
|---|---|
| A section's heading text (`## Beziehungen`) | The KB language was a compiler constant, and the region's *end* was a guess. Content sitting after it was silently deleted on eight pages |
| A relationship label in a body bullet (`- **hängt ab von:**`) | Nothing could check the vocabulary, so it drifted to **152 distinct labels** across 337 bullets against thirteen that were documented |
| The reciprocal half of every edge | `xref add` mirrored every link, which made per-collection label authorisation impossible and filled `## Siehe auch` with 555 unlabelled bullets, 353 of them provably redundant |
4.0.0 replaces all three. A region is delimited by a marker pair and rendered from frontmatter;
a label is a machine value in `related:`, drawn from a catalogue and authorised per destination
by the source collection; an edge is authored in one direction and the inbound view is computed.
**This one touches pages.** Unlike 3.0.0 it is not a contract reshuffle: every `related:` entry
and every tool-owned body region changes. It is `assisted` because there is no mapping table -
mapping free-text German onto a 35-label catalogue is a judgment call per edge, and a large
minority of the old labels are reverse directions that under the new model are not stored at all.
## When to run
After installing 4.0.0 over an instance on 3.x. `tools/wikitool migrate status` names it, and
`lint` reports `unlabelled_edges` for every unconverted edge - that count reaching zero is how
you know the run is finished.
**Nothing breaks while it is outstanding.** Unlabelled edges and undelimited regions are read,
not rejected: `links.py` treats a bare title as an edge whose label is not declared yet, and
`provenance.split_cite_block` falls back to the pre-marker layout. That is deliberate - a corpus
has to stay readable while it is being converted - and it is why the two lint findings are
advisory until step 6 promotes them.
## Steps
1. **Rewrite `kb/CONVENTIONS.md`'s `sections:` block.** Three slots become two, because the
See Also region is gone:
```yaml
sections:
links: <your heading for declared relationships>
footnotes: <your heading for citation definitions>
```
Delete `section_aliases:` if you have one - nothing matches on heading text any more, so
there is nothing to alias. The heading is now a *rendering* value: changing it re-renders
the words above each region on the next write and can no longer split a page.
2. **Add an `outbound:` block to every `kb/<name>/COLLECTION.md`.** Which labels a page may use,
per destination collection, with `any` as a wildcard:
```yaml
outbound:
entities: [depends-on, runs-on, uses, see-also]
concepts: [implements, see-also]
```
The catalogue to draw from is [link-taxonomy.md](../link-taxonomy.md); the four contracts in
the origin repo are worked examples. **The source collection decides** - that is what makes a
35-label palette usable, and it is why the reverse edge can no longer be written
automatically. A destination you do not list authorises nothing, which is a real answer.
3. **Fix your page type-spec templates.** If you adopted the 3.0.0 templates, they contain
`## {section.relationships}` and `## {section.see_also}`. Those variables no longer exist and
would be written into new pages literally. **Delete both sections from the `## Template`
block** - a template must not scaffold a tool-owned region at all: it is generated between
markers on the first `xref add` / `cite add` and re-rendered on every write.
4. **Convert the corpus**, following [migrate-corpus.md](../migrate-corpus.md). Cut it into
units sized against the iteration budget; the origin repo used four, ~45 pages each. Per page:
- For each labelled bullet under the old relationships heading: say the sentence
`[this page] <label> [target]` and pick the catalogue label that makes it true. If it only
reads true **backwards**, the edge belongs on the other page - move it there rather than
inventing an inverse label the catalogue does not have.
- For each bare `- [[X]]` bullet under the old See Also heading: drop it if a labelled edge
already connects the pair. Otherwise decide - a real label, or dropped with the reason
recorded. **Do not convert them to `see-also` in bulk.** That is the one shortcut this
migration explicitly refuses: it would start the new taxonomy with most of its edges on its
weakest label, which is the sediment the change exists to remove.
- Write edges with `tools/wikitool xref add --a "<A>" --b "<B>" --rel <label>`, never by
hand. The body region is rendered from `related:`; editing inside a marker pair is
overwritten without warning.
- `cite sync` converts a page's old footnote block into a marked region in passing.
5. **Check each unit mechanically before anything else:**
```bash
tools/wikitool migrate verify --from <pre-migration rev> --path kb/<area> --fail-on-error
```
It compares wikilink and citation **counts**, footnote definitions, H1, structural
frontmatter, and - new in 4.0.0 - the **count of marker pairs per region**. A dropped marker
is otherwise silent: the region becomes ordinary prose and the next write appends a second
one beside it.
6. **Record it, then tighten the checks:**
```bash
tools/wikitool lint # unlabelled_edges and unauthorised_labels must be 0
tools/wikitool migrate done 4.0.0 --pages <N>
```
Only once `lint` reports zero of both is the run finished. The two findings are advisory
during the window and become hard errors afterwards - the same path
`legacy_citation_markers` took after the citation migration.
## How to tell a migrated page from an unmigrated one
Its `related:` entries are `- <label>: <title>` rather than bare titles, and its relationship
and footnote sections sit between `<!-- wikitool:links -->` / `<!-- wikitool:footnotes -->`
marker pairs. `tools/wikitool links show --page "<Title>"` prints `unlabelled` for every edge
still waiting, and `lint`'s `unlabelled_edges` count is the corpus-wide version of the same
question.
## Decision points
- **A label you want is not in the catalogue?** Add it - a row in `link-taxonomy.md` and an
entry in the authorising `COLLECTION.md`. No code change is involved, and the registers are
advisory groupings rather than a schema. Do not stretch a label whose sentence reads false: a
wrong edge is worse than a weak one, because it is machine-readable and will be believed.
- **`related:` holds an entry with no body bullet to derive a label from?** Expected - the
origin repo found 480 edges against 337 bullets, because frontmatter and body had already
drifted apart while the label lived only in prose. Read the page and decide; that drift is
itself part of what this migration repairs.
- **A page loses its last inbound edge?** The orphan check will now report it, and that is the
check working: directional edges mean a page nothing points at is genuinely unreachable, where
the old mirrored model always manufactured a back-link. Either something should point at it,
or it is reached through the catalog and that is fine.
- **Tempted to keep writing reverse edges for navigation?** Do not. `links show` computes the
inbound view, and the rendered bullet on the asserting page is an ordinary `[[wikilink]]`, so
a backlink panel in an editor already shows it.
## Scope
The corpus under `kb/`, plus the three instance-owned declarations in steps 1-3. It does not
touch `raw/`, and it learns nothing new: the same knowledge is restated in a form that can be
checked. Installing the 4.0.0 machinery itself is `INSTALL.md`'s and must have happened first.
+19 -6
View File
@@ -68,17 +68,23 @@ bereit für den ersten `Ingest`.
Ablauf: Ablauf:
1. Die Collection-Contracts übernehmen - vier Kopien, keine Frage an den Nutzer, denn was 1. Die Collection-Contracts **und die Page-Type-Specs** übernehmen - Kopien, keine Frage an
dort steht ist unabhängig von der Sprache brauchbar: den Nutzer, denn was dort steht ist als Ausgangspunkt unabhängig von der Sprache brauchbar:
```bash ```bash
for template in kb/*/COLLECTION.md.template; do for template in kb/*/COLLECTION.md.template types/*.template; do
cp "$template" "${template%.template}" cp "$template" "${template%.template}"
done done
``` ```
Die `.template`-Dateien bleiben liegen; sie sind die Vorlage für den nächsten Export. Die `.template`-Dateien bleiben liegen; sie sind die Vorlage für den nächsten Export.
Unter `types/` betrifft das genau die Type-Specs mit `root: kb` - `entity`, `concept`,
`source`, `comparison` - samt ihrer `.schema.yaml`. Sie beschreiben Seiten, die *diese*
Instanz schreibt, also gehören sie ihr: Prosa, Template und Sprache dürfen umgeschrieben
werden. `instruction`, `lint-report` und `type-spec` beschreiben Stack-Artefakte und
kommen unverändert.
2. Den Nutzer nach der KB-Sprache fragen. `kb/CONVENTIONS.md.template` ist auf **Englisch** 2. Den Nutzer nach der KB-Sprache fragen. `kb/CONVENTIONS.md.template` ist auf **Englisch**
voreingestellt; [kb-profiles.md](kb-profiles.md) hält daneben ein vollständiges voreingestellt; [kb-profiles.md](kb-profiles.md) hält daneben ein vollständiges
deutsches Profil bereit, und dessen Volltext ist die `kb/CONVENTIONS.md` des Quell-Repos. deutsches Profil bereit, und dessen Volltext ist die `kb/CONVENTIONS.md` des Quell-Repos.
@@ -99,9 +105,16 @@ bereit für den ersten `Ingest`.
Migration jeder vorhandenen Seite (`section_aliases:` trägt die alten Namen, siehe Migration jeder vorhandenen Seite (`section_aliases:` trägt die alten Namen, siehe
[migrate-corpus.md](migrate-corpus.md)). [migrate-corpus.md](migrate-corpus.md)).
**Nichts davon liegt unter `tools/` oder `types/`.** Der Compiler liest die Abschnittsnamen **Nichts davon liegt in einer Stack-Datei.** Der Compiler liest die Abschnittsnamen aus
aus `kb/CONVENTIONS.md`, und die vier Page-Type-Templates setzen sie über `kb/CONVENTIONS.md`; die vier Page-Type-Specs gehören ab Schritt 1 dieser Instanz. Eine
`{section.…}`-Variablen ein - eine anderssprachige Instanz ändert dort keine Datei. anderssprachige Instanz übersetzt sie einfach - das ist kein lokaler Patch an etwas
Ausgeliefertem mehr, sondern Arbeit an den eigenen Dateien, und ein Upgrade nimmt sie ihr
nicht wieder weg.
Was der Stack von `types/` überhaupt noch verlangt, ist eine Zeile: es muss einen Type-Spec
mit `name: source` geben, dessen Schema `raw_files` fordert. Daran hängt der gesamte
`raw/``kb/`-Provenance-Pfad (`sources coverage`, `[^cite-id]`-Auflösung, `kb/provenance.md`),
und `docs verify` prüft genau das - nicht mehr.
Unverändert bleibt in jedem Fall die Regel, die dem Stack gehört: **jede Zeile einer Seite Unverändert bleibt in jedem Fall die Regel, die dem Stack gehört: **jede Zeile einer Seite
ist Prosa oder Identifier, und nur Prosa wird übersetzt** ([kb/CONTRACT.md § Language and ist Prosa oder Identifier, und nur Prosa wird übersetzt** ([kb/CONTRACT.md § Language and
+53 -27
View File
@@ -12,9 +12,8 @@ second half is the cut: what is written here is enforced by `tools/wikitool` or
how it works, so it is identical everywhere and `dist export` ships it verbatim. how it works, so it is identical everywhere and `dist export` ships it verbatim.
**What an instance decides for itself is next door, in **What an instance decides for itself is next door, in
[kb/CONVENTIONS.md](CONVENTIONS.md)** - the language pages are written in and its three [kb/CONVENTIONS.md](CONVENTIONS.md)** - the language pages are written in, the headings its two
tool-owned section headings, the naming forms, the tone, the relationship-label vocabulary, the generated regions render under, the naming forms, the tone, the confidence rubric. That file binds exactly as this one does; it is simply owned by the instance
confidence rubric. That file binds exactly as this one does; it is simply owned by the instance
rather than by the stack, so the distribution ships only its `.template` and the instance writes rather than by the stack, so the distribution ships only its `.template` and the instance writes
the real one. Read both, plus the target collection's `kb/<name>/COLLECTION.md` (also the real one. Read both, plus the target collection's `kb/<name>/COLLECTION.md` (also
instance-owned), before writing or editing a page. instance-owned), before writing or editing a page.
@@ -99,7 +98,8 @@ decision record - is the instance's, in
- [ ] Carry a clear, descriptive title and a summary near the top - [ ] Carry a clear, descriptive title and a summary near the top
- [ ] Use consistent terminology with the rest of the wiki - [ ] Use consistent terminology with the rest of the wiki
- [ ] Link to every entity and concept it mentions, and be linked to in return - [ ] Link to the entities and concepts it mentions, and declare an edge where the relationship
is worth naming - in the direction this page asserts it, not in both
- [ ] Cite its hard facts (see [Provenance and citation](#provenance-and-citation)) - [ ] Cite its hard facts (see [Provenance and citation](#provenance-and-citation))
- [ ] Duplicate no existing page - [ ] Duplicate no existing page
- [ ] Appear in the catalog (guaranteed by `wikitool index rebuild`) - [ ] Appear in the catalog (guaranteed by `wikitool index rebuild`)
@@ -141,36 +141,62 @@ instance records - see [kb/CONVENTIONS.md § Language](CONVENTIONS.md#language).
evidence *about* a source, not a substitute for it. Quote verbatim in the original language and evidence *about* a source, not a substitute for it. Quote verbatim in the original language and
record the raw file's language in `source_language:`. record the raw file's language in `source_language:`.
### Section headings ### Generated regions
Three headings are a vocabulary the tool owns rather than prose an author picks: `xref add` Two regions of a page body are **generated**, not authored: the links region `xref` owns and the
writes into Relationships and See Also, and `cite add` owns the trailing Footnotes block. They footnotes region `cite` owns. Each sits between a marker pair:
follow the KB language like everything else, so **the instance names them**, in
`kb/CONVENTIONS.md`'s `sections:` frontmatter. `tools/chemenu/conventions.py` reads that
declaration and `tools/chemenu/sections.py` is what the rest of the compiler asks - there is no
heading text in the compiler itself.
Each has aliases the tool still *recognizes* but no longer writes, which is what lets the corpus ```markdown
be translated page by page: a page still carrying `## Relationships` is found and appended to <!-- wikitool:links -->
correctly, and `cite sync` leaves an untranslated `## Footnotes` heading alone rather than ## Beziehungen
retitling it. The recognized set is the canonical name, any `section_aliases:` the instance
declared, and the names this stack wrote before the declaration existed. Renaming a heading is - **depends-on:** [[Hermes]]
the translation pass's job, never a side effect of another command. Any *other* heading an <!-- /wikitool:links -->
author adds is ordinary prose and is translated with the rest. ```
The marker is what the tool locates the region by, and everything between the markers -
**heading included** - is replaced wholesale on the next write. An author never edits inside
them; anything left there is overwritten without warning, exactly as in `kb/index.md`. A region
with nothing to show is absent rather than empty.
The heading is therefore a *rendering* value, taken from `kb/CONVENTIONS.md`'s `sections:`. No
heading text exists in the compiler, and nothing matches on it: changing the declaration
re-renders the words on the next write and cannot split a page.
That is not how it used to work. The tool located these regions by matching their heading text,
which made a translated heading a structural fact - and made the region's *end* a guess. It ran
to the next heading, and before that to the end of the file, which silently deleted whatever sat
after it on eight pages. Any *other* heading a page carries is ordinary prose.
## Linking ## Linking
Every page links to what it mentions, in both directions. Cross-references are created with **An edge is authored in one direction**, on the page that asserts it, and carries a label that
`tools/wikitool xref add --a "<A>" --b "<B>" --rel-a "<label>" --rel-b "<label>"`, never by is a machine value rather than prose:
hand-editing the `related:` array or the Relationships/See Also bullets.
Use a typed relationship label rather than a generic one. The label is free text as far as the ```yaml
tool is concerned - it is written into a `- **label:** [[Title]]` bullet and no code matches on related:
it - so which vocabulary this instance uses is - depends-on: Hermes
[kb/CONVENTIONS.md § Relationship labels](CONVENTIONS.md#relationship-labels)'s to list. ```
A page is expected to have at least one inbound link; `wikitool lint` reports orphans. Created with `tools/wikitool xref add --a "<A>" --b "<B>" --rel <label>`, never by hand-editing
Comparison pages are exempt - they are reached through the catalog. `related:` or the rendered bullet. Say the sentence before choosing the label - `[A] <label>
[B]` - and if it only reads true backwards, the edge belongs on the other page.
**A reverse edge is a separate decision, not a mirror.** Write one when it independently helps a
reader at the other end; do not write one to make the graph symmetric. Navigation does not
depend on it either way: `index rebuild` renders the inbound view from the graph, completely and
without maintenance.
Which labels exist is [instructions/link-taxonomy.md](../instructions/link-taxonomy.md), a
palette that binds nothing. Which of them a page may *use* is its own collection's `outbound:`
block, per destination - the **source** collection decides, because the rules that govern an
edge are the rules of the collection asserting it. `xref add` refuses an unauthorised label and
`lint` reports one.
A page is expected to have at least one inbound edge; `wikitool lint` reports orphans.
Comparison pages are exempt - they are reached through the catalog. Directional edges mean more
pages qualify than under the old mirrored model, and that is the check measuring reachability
rather than measuring whether `xref` ran.
Renaming a page, deleting one, or dropping a single reference are tool operations with their Renaming a page, deleting one, or dropping a single reference are tool operations with their
own procedure: see [instructions/page-lifecycle.md](../instructions/page-lifecycle.md). own procedure: see [instructions/page-lifecycle.md](../instructions/page-lifecycle.md).
+22 -29
View File
@@ -2,8 +2,7 @@
language: de language: de
profile: german profile: german
sections: sections:
relationships: Beziehungen links: Beziehungen
see_also: Siehe auch
footnotes: Fußnoten footnotes: Fußnoten
--- ---
@@ -21,13 +20,11 @@ Adopted from the `german` profile in
[instructions/kb-profiles.md](../instructions/kb-profiles.md). That catalogue is a palette, not [instructions/kb-profiles.md](../instructions/kb-profiles.md). That catalogue is a palette, not
an enum - what is written here is what holds, whether or not a profile says the same thing. an enum - what is written here is what holds, whether or not a profile says the same thing.
The frontmatter above is the one machine-read part. `sections:` names the three headings The frontmatter above is the one machine-read part. `sections:` names the headings the two
`wikitool xref` and `wikitool cite` write into; `tools/chemenu/conventions.py` reads them and **generated regions** render under - the links region `wikitool xref` owns and the footnotes
`tools/chemenu/sections.py` is what the rest of the compiler asks. Renaming one here changes region `wikitool cite` owns. Each sits between a marker pair, and the marker is what the tool
what the tool *writes*; what it still *recognizes* is the union of that name, any locates it by, so the heading here is a display value: changing it re-renders the words above
`section_aliases:` declared beside it, and the names this stack wrote before this file existed. those regions and nothing else. Nothing matches on this text.
That asymmetry is the translation path: a page keeps working under its old heading until it is
itself translated.
## Language ## Language
@@ -52,12 +49,13 @@ material, not a second rule - every entry in it is a decision that was made wron
### Section headings ### Section headings
The canonical names are the frontmatter's: `## Beziehungen`, `## Siehe auch`, `## Fußnoten`. The two generated regions render under `## Beziehungen` and `## Fußnoten`. An author never
The English forms this stack wrote before the corpus was translated are still recognized, so a writes inside them - they are rebuilt from frontmatter on every write, exactly like
page carrying `## Relationships` is found and appended to correctly and `cite sync` leaves an `kb/index.md` - and never has to write the heading either. Any *other* heading on a page is
untranslated `## Footnotes` alone. Renaming such a heading is the translation pass's job, never ordinary prose and is translated with the rest.
a side effect of another command. Any *other* heading an author adds is ordinary prose and is
translated with the rest. There is no `## Siehe auch` region any more. It was the reciprocal half of a bidirectional
`xref add`; under authored directional edges, `see-also` is a *label* inside the links region.
## Naming ## Naming
@@ -92,15 +90,11 @@ The blockquote cap is not here: `wikitool lint` reports it, so it is the contrac
## Relationship labels ## Relationship labels
`tools/wikitool xref add --rel-a/--rel-b` takes a free-text label. This instance uses a typed **Not this file's to list, and not localized.** A label is a machine value in `related:`, drawn
one rather than a generic one: from [instructions/link-taxonomy.md](../instructions/link-taxonomy.md) and authorised per
destination in each `kb/<name>/COLLECTION.md`'s `outbound:` block. `- **depends-on:** [[Hermes]]`
`hängt ab von` · `verwendet` · `implementiert` · `erweitert` · `ersetzt` · `steht in Konflikt mit` is what a German page carries, and that is deliberate: the label is an identifier, so translating
· `benötigt` · `erzeugt` · `konsumiert` · `besitzt` · `pflegt` · `läuft auf` · `verwandt mit` it would make the graph's semantics depend on the prose again.
(last resort)
The labels are prose written into a `- **label:** [[Title]]` bullet; no code matches on them, so
an untranslated page's English label is stale wording, not a broken reference.
## Confidence rubric ## Confidence rubric
@@ -119,8 +113,7 @@ write "unsicher"/"unbestätigt".
## Keeping this file honest ## Keeping this file honest
Change it when a convention actually changes, and treat a change to `sections:` as a corpus Change it when a convention actually changes. `sections:` is safe to change at any time - the
migration rather than an edit: existing pages keep their old headings until something translates regions are located by their markers and re-rendered under the new words on the next write.
them, and the alias list is what carries them in the meantime. `wikitool doctor` FAILs on a `wikitool doctor` FAILs on a missing or unfilled file, and `wikitool docs verify` refuses a
missing or unfilled file, and `wikitool docs verify` refuses a `sections:` block that does not `sections:` block that does not name both regions.
name all three slots.
+11 -26
View File
@@ -3,16 +3,8 @@
language: en language: en
profile: none profile: none
sections: sections:
relationships: Relationships links: Relationships
see_also: See Also
footnotes: Footnotes footnotes: Footnotes
# Headings this instance no longer writes but still recognizes, so a corpus can
# be translated page by page instead of all at once. Optional; the names this
# stack wrote before this file existed are always recognized anyway.
# section_aliases:
# relationships: [Beziehungen]
# see_also: [Siehe auch]
# footnotes: [Fußnoten]
--- ---
# kb/ - Authoring Conventions of This Instance # kb/ - Authoring Conventions of This Instance
@@ -29,9 +21,9 @@ Ready-made answers to every section below - including a complete German profile
[instructions/kb-profiles.md](../instructions/kb-profiles.md). That catalogue is a palette, not [instructions/kb-profiles.md](../instructions/kb-profiles.md). That catalogue is a palette, not
an enum: adopt an entry, adapt it, or write your own. What is written *here* is what holds. an enum: adopt an entry, adapt it, or write your own. What is written *here* is what holds.
The frontmatter above is the one machine-read part. `sections:` names the three headings The frontmatter above is the one machine-read part. `sections:` names the headings the two
`wikitool xref` and `wikitool cite` write into. Set them before the first page is written: generated regions render under. Safe to change at any time - each region is located by its
afterwards, changing one is a corpus migration rather than an edit. marker pair, so a rename re-renders words and nothing else.
## Language ## Language
@@ -52,10 +44,8 @@ is the one those terms are already in.}
### Section headings ### Section headings
The canonical names are the frontmatter's. Any name this instance previously wrote stays The two generated regions render under the frontmatter's headings. An author never writes inside
recognized through `section_aliases:`, which is what lets a corpus be translated page by page. them - they are rebuilt from frontmatter on every write. Any *other* heading is ordinary prose.
Renaming such a heading is the translation pass's job, never a side effect of another command.
Any *other* heading an author adds is ordinary prose.
## Naming ## Naming
@@ -80,13 +70,9 @@ Bad: {the same sentence written the way it must not be.}
## Relationship labels ## Relationship labels
`tools/wikitool xref add --rel-a/--rel-b` takes a free-text label. Listing the ones this **Not this file's to list, and not localized.** A label is a machine value in `related:`, drawn
instance uses is what keeps a graph typed rather than a wiki full of "related to": from [instructions/link-taxonomy.md](../instructions/link-taxonomy.md) and authorised per
destination in each `kb/<name>/COLLECTION.md`'s `outbound:` block.
{the label vocabulary, in the KB language}
No code matches on these, so an old label on an untranslated page is stale wording, not a
broken reference.
## Confidence rubric ## Confidence rubric
@@ -99,6 +85,5 @@ contract's. What the number *means* is this instance's:
## Keeping this file honest ## Keeping this file honest
Change it when a convention actually changes, and treat a change to `sections:` as a corpus Change it when a convention actually changes. `wikitool doctor` FAILs on a missing or unfilled
migration rather than an edit. `wikitool doctor` FAILs on a missing or unfilled file, and file, and `wikitool docs verify` refuses a `sections:` block that does not name both regions.
`wikitool docs verify` refuses a `sections:` block that does not name all three slots.
+16 -3
View File
@@ -1,5 +1,7 @@
--- ---
profile: comparisons profile: comparisons
outbound:
any: [compares-with, contrasts, see-also]
required_by_stack: false required_by_stack: false
--- ---
@@ -37,11 +39,22 @@ alphabetically.
- State the trade-off, not a winner. Where a recommendation is genuinely warranted, scope it: - State the trade-off, not a winner. Where a recommendation is genuinely warranted, scope it:
"for X workload", not "better". "for X workload", not "better".
## Authorised labels
The `outbound:` block above is what `wikitool lint` and `xref add` check: which labels a page in
this collection may use, per destination. The catalogue they are drawn from - and what each one
asserts - is [instructions/link-taxonomy.md](../../instructions/link-taxonomy.md), which binds
nothing on its own.
Narrow for the opposite reason: a comparison's substance is its table, and its links to the compared subjects are the one relationship it asserts.
Adding a label here is a deliberate contract change, not a way around a refusal.
## Outbound linking ## Outbound linking
A comparison links to every subject with `related to`, and each subject links back. Comparison A comparison links to every subject with `compares-with`. The subjects do not have to link back:
pages are **exempt from the orphan check** - they are reached through `index.md` rather than a comparison is reached through the catalog, and each subject's inbound view renders the edge
through inbound prose links. anyway. Comparison pages are **exempt from the orphan check** for the same reason.
## What does not belong here ## What does not belong here
+18 -2
View File
@@ -1,5 +1,10 @@
--- ---
profile: concepts profile: concepts
outbound:
concepts: [extends, grounds, rests-on, enables, precondition, exemplifies, abstracted-from, contrasts, compares-with, contradicts, composition, part-of, supersedes, derived-from, adapted-from, see-also]
entities: [operationalized-from, mechanism, procedure, applies-when, operates-on, invokes, exemplifies, see-also]
sources: [evidenced-by, derived-from, adapted-from, defined-in, see-also]
comparisons: [compares-with, see-also]
required_by_stack: false required_by_stack: false
--- ---
@@ -33,8 +38,19 @@ An architectural decision is a concept page, prefixed as
- **Status** - proposed / accepted / deprecated / superseded. - **Status** - proposed / accepted / deprecated / superseded.
- Links to every entity the decision affects. - Links to every entity the decision affects.
A superseded ADR is never deleted or rewritten; a new one supersedes it and both link to the A superseded ADR is never deleted or rewritten. The new one declares `supersedes` pointing at
other with `replaces` / `replaced by`. it; the old one needs no edge back, because its inbound view renders the replacement.
## Authorised labels
The `outbound:` block above is what `wikitool lint` and `xref add` check: which labels a page in
this collection may use, per destination. The catalogue they are drawn from - and what each one
asserts - is [instructions/link-taxonomy.md](../../instructions/link-taxonomy.md), which binds
nothing on its own.
The widest authorisation in this instance, because argumentation is what concept pages do. Note that the operational labels are absent: a concept does not `depend-on` anything - the entity implementing it does.
Adding a label here is a deliberate contract change, not a way around a refusal.
## Outbound linking ## Outbound linking
+16
View File
@@ -1,5 +1,10 @@
--- ---
profile: entities profile: entities
outbound:
entities: [depends-on, required-by, runs-on, hosts, uses, produces, consumes, maintains, owns, part-of, composition, supersedes, see-also]
concepts: [implements, exemplifies, rests-on, applies-when, operates-on, invokes, see-also]
sources: [evidenced-by, defined-in, see-also]
comparisons: [compares-with, see-also]
required_by_stack: false required_by_stack: false
--- ---
@@ -43,6 +48,17 @@ These are areas, not collections: they inherit this contract and carry no `COLLE
- **People** - role, affiliation, and the projects or decisions they are connected to. Nothing - **People** - role, affiliation, and the projects or decisions they are connected to. Nothing
personal beyond what the source states. personal beyond what the source states.
## Authorised labels
The `outbound:` block above is what `wikitool lint` and `xref add` check: which labels a page in
this collection may use, per destination. The catalogue they are drawn from - and what each one
asserts - is [instructions/link-taxonomy.md](../../instructions/link-taxonomy.md), which binds
nothing on its own.
Operational labels dominate here because an entity's relationships are mostly to other concrete things. `implements` points *out* to a concept; the concept does not point back unless that direction is a statement of its own.
Adding a label here is a deliberate contract change, not a way around a refusal.
## Outbound linking ## Outbound linking
An entity links to the technologies it uses, the systems it runs on, the projects that depend An entity links to the technologies it uses, the systems it runs on, the projects that depend
+13
View File
@@ -1,5 +1,7 @@
--- ---
profile: sources profile: sources
outbound:
any: [is-evidence-for, defined-in, see-also]
required_by_stack: true required_by_stack: true
--- ---
@@ -38,6 +40,17 @@ The `raw_files:`/`source_url:`/citation rules are shared and live in
- `tools/wikitool sources trace --raw <path>` answers "what did we learn from this?"; - `tools/wikitool sources trace --raw <path>` answers "what did we learn from this?";
`tools/wikitool sources coverage` lists raw files no source page claims yet. `tools/wikitool sources coverage` lists raw files no source page claims yet.
## Authorised labels
The `outbound:` block above is what `wikitool lint` and `xref add` check: which labels a page in
this collection may use, per destination. The catalogue they are drawn from - and what each one
asserts - is [instructions/link-taxonomy.md](../../instructions/link-taxonomy.md), which binds
nothing on its own.
Deliberately narrow. A source page is evidence *about* a source; almost everything it would want to say is already carried by `raw_files:`, `sources:` and `[^cite-id]`, which are the mechanical provenance path rather than authored edges.
Adding a label here is a deliberate contract change, not a way around a refusal.
## Outbound linking ## Outbound linking
A source page links to every entity and concept it produced or updated. A source page links to every entity and concept it produced or updated.
+13 -11
View File
@@ -40,12 +40,13 @@ tools/wikitool <command> --help
| `touch --page "<Title>" [--summary "..."] [--provenance <v>] [--confidence-base <n>] [--date YYYY-MM-DD] [--set field=value ...] [--add field=value ...] [--remove field=value ...] [--no-date] [--dry-run]` | Update a page's own frontmatter: bump `modified:` and optionally rewrite any field its type declares. `--summary`/`--provenance`/`--confidence-base` are shorthands; `--set` reaches every other field and **replaces** its value, while `--add`/`--remove` change single elements of an array field (removing an absent element succeeds and says so). Repeating `--set` for one array field appends *within the call*, and `\,` is a literal comma - same rules as `new --set`. Refused with the command that owns them instead: `type:` (page-lifecycle), `confidence:` (derived - set `--confidence-base`), and the page-ref arrays `related:`/`sources:`/`entities:`/`concepts:` (`xref`). Everything else the schema declares is settable, and an unknown field lists what the page actually has. Schema-validates the fields it writes, and `raw_files:` entries must exist on disk. A source declares `date:` instead of `modified:`, and that is the *publication* date of the raw material - it is never bumped to today, and changes only when `--date` names a value explicitly. | | `touch --page "<Title>" [--summary "..."] [--provenance <v>] [--confidence-base <n>] [--date YYYY-MM-DD] [--set field=value ...] [--add field=value ...] [--remove field=value ...] [--no-date] [--dry-run]` | Update a page's own frontmatter: bump `modified:` and optionally rewrite any field its type declares. `--summary`/`--provenance`/`--confidence-base` are shorthands; `--set` reaches every other field and **replaces** its value, while `--add`/`--remove` change single elements of an array field (removing an absent element succeeds and says so). Repeating `--set` for one array field appends *within the call*, and `\,` is a literal comma - same rules as `new --set`. Refused with the command that owns them instead: `type:` (page-lifecycle), `confidence:` (derived - set `--confidence-base`), and the page-ref arrays `related:`/`sources:`/`entities:`/`concepts:` (`xref`). Everything else the schema declares is settable, and an unknown field lists what the page actually has. Schema-validates the fields it writes, and `raw_files:` entries must exist on disk. A source declares `date:` instead of `modified:`, and that is the *publication* date of the raw material - it is never bumped to today, and changes only when `--date` names a value explicitly. |
| `rename --from "<Old>" --to "<New>" [--dry-run]` | Rename a page and repoint every reference to it: body `[[wikilinks]]` (aliases and anchors preserved), a `[^cite-id]` whose id was derived from the old title (refreshed to match the new one, both in its Footnotes definition and every reference to it), the page's own H1, and every page-ref frontmatter array declared by the type's `page_ref_fields:`. If `--from` is *not* a page but is referenced, it instead repoints those references onto the existing `--to` page and moves nothing - the fix for a reference spelled `act_runner` when the page is `Act Runner` | | `rename --from "<Old>" --to "<New>" [--dry-run]` | Rename a page and repoint every reference to it: body `[[wikilinks]]` (aliases and anchors preserved), a `[^cite-id]` whose id was derived from the old title (refreshed to match the new one, both in its Footnotes definition and every reference to it), the page's own H1, and every page-ref frontmatter array declared by the type's `page_ref_fields:`. If `--from` is *not* a page but is referenced, it instead repoints those references onto the existing `--to` page and moves nothing - the fix for a reference spelled `act_runner` when the page is `Act Runner` |
| `rm --page "<Title>" [--yes] [--dry-run]` | Delete a page and mechanically de-link it. Refuses without `--yes` while other pages still reference it. Strips ref-array entries and bare `- [[Title]]` / `- **label:** [[Title]]` bullets; leaves prose and inline citations in place and reports them | | `rm --page "<Title>" [--yes] [--dry-run]` | Delete a page and mechanically de-link it. Refuses without `--yes` while other pages still reference it. Strips ref-array entries and bare `- [[Title]]` / `- **label:** [[Title]]` bullets; leaves prose and inline citations in place and reports them |
| `xref add --a "<A>" --b "<B>" --rel-a "<label>" --rel-b "<label>"` | Bidirectionally link two pages: frontmatter `related:` + body Relationships/See Also bullets. Idempotent. Refuses, before writing either side, when a page's type does not declare `related:` - a source page declares `entities:`/`concepts:` instead, and writing `related:` there produced frontmatter the schema rejects; the refusal names the fields the type does declare and points at `link-source`. | | `xref add --a "<A>" --b "<B>" --rel <label>` | Declare **one** edge: `A <label> B`, written into A's `related:` as `- <label>: B` and rendered into A's generated links region. B is not touched and does not point back - its inbound view is rendered from the graph. Idempotent, and re-running with a different label *relabels* rather than appending, since one page asserts one thing about another. Refuses before writing when the type does not declare `related:` (a source page declares `entities:`/`concepts:` - the refusal names them and points at `link-source`), and when `<label>` is not authorised by the source collection's `outbound:` block for the target's collection; that refusal lists the authorised set and points at `instructions/link-taxonomy.md` |
| `xref remove --a "<A>" --b "<B>" [--dry-run]` | Inverse of `xref add` *and* `xref link-source`: clears `<B>` from every page-ref frontmatter field `<A>`'s type declares (`related:`, `sources:`, `entities:`, `concepts:`) plus the matching bullets. It also sweeps a field the type does *not* declare but some other type does, and drops that key outright once empty - a leftover written before the check above existed has to stay repairable, or the page is a dead end. `--b` need not still exist as a page, so this is how a reference left by a hand-deleted or hand-renamed page gets cleared without hand-editing frontmatter. Idempotent. | | `xref remove --a "<A>" --b "<B>" [--dry-run]` | Clears the reference in **both** directions - it is the cleanup command for a deleted or hand-renamed page rather than the strict inverse of a one-directional `add`. Clears `<B>` from every page-ref frontmatter field `<A>`'s type declares (`related:`, `sources:`, `entities:`, `concepts:`) plus the matching bullets. It also sweeps a field the type does *not* declare but some other type does, and drops that key outright once empty - a leftover written before the check above existed has to stay repairable, or the page is a dead end. `--b` need not still exist as a page, so this is how a reference left by a hand-deleted or hand-renamed page gets cleared without hand-editing frontmatter. Idempotent. |
| `xref link-source --source "Source - X" --entities A,B,C` | Batch-link a source page to every entity/concept it mentions, **in both directions**: each target gets `sources:` + a See Also bullet, and the source page records each target in its own `entities:`/`concepts:`. Which of the two is chosen follows the target's collection (`kb/entities/` -> `entities:`), so a new collection needs no code change here. A target whose collection matches no reference field the source type declares is linked one-way and named in the output. Idempotent in both directions | | `xref link-source --source "Source - X" --entities A,B,C` | Batch-link a source page to every entity/concept it mentions: each target gets `sources:`, and the source page records each target in its own `entities:`/`concepts:`. No body bullet is written on either side - `sources:` *is* the record, and the See Also bullet this used to add was the reciprocal half of a model that no longer exists. Which of the two is chosen follows the target's collection (`kb/entities/` -> `entities:`), so a new collection needs no code change here. A target whose collection matches no reference field the source type declares is linked one-way and named in the output. Idempotent in both directions |
| `links show --page "<Title>" [--json]` | The declared graph around one page in both directions: the edges it asserts (from its own `related:`, with labels) and the edges other pages assert about it (computed across the corpus). The inbound half is derived rather than stored - that is what makes it complete, and it is the answer authored directional edges would otherwise have nowhere to come from. Read-only, exempt from the Iteration Budget Gate |
| `cite id --title "Source - X" [--file <qualifier>]` | Print the deterministic footnote id `cite add` would use for this (title, file) pair. Read-only, exempt from the Iteration Budget Gate | | `cite id --title "Source - X" [--file <qualifier>]` | Print the deterministic footnote id `cite add` would use for this (title, file) pair. Read-only, exempt from the Iteration Budget Gate |
| `cite add --page "<Title>" --source "Source - X" [--file <qualifier>] [--dry-run]` | Upsert a `[^cite-id]: [[Source - X]]` definition in the page's Footnotes block (reusing the id if the page already cites this exact source/file pair) and add `Source - X` to frontmatter `sources:`. Prints the `[^cite-id]` marker - pasting it into the prose is still a manual, editorial step | | `cite add --page "<Title>" --source "Source - X" [--file <qualifier>] [--dry-run]` | Upsert a `[^cite-id]: [[Source - X]]` definition in the page's generated footnotes region, creating it between `<!-- wikitool:footnotes -->` markers if absent (reusing the id if the page already cites this exact source/file pair) and add `Source - X` to frontmatter `sources:`. Prints the `[^cite-id]` marker - pasting it into the prose is still a manual, editorial step |
| `cite sync [--page "<Title>" \| --all] [--dry-run]` | Reconcile each page's Footnotes block against its actual `[^id]` references: prune definitions nothing references any more, re-render the block in first-reference order, and report any `[^id]` reference left with no definition | | `cite sync [--page "<Title>" \| --all] [--dry-run]` | Reconcile each page's footnotes region against its actual `[^id]` references: prune definitions nothing references any more, re-render the region in first-reference order, and report any `[^id]` reference left with no definition. A page still carrying the pre-4.0.0 undelimited block is converted to a marked region in the same pass - the marker carries the region's identity now, so re-rendering it under this instance's heading is a repair rather than a rename |
| `index rebuild [--dry-run]` | Regenerate the catalog from every page's frontmatter: `kb/index.md` becomes a map (statistics, one row per collection and per area, links to the shards) and the page tables are written to a generated `INDEX.md` in each collection. An area past 50 rows gets its own shard. Stale shards from removed collections/areas are deleted in the same pass | | `index rebuild [--dry-run]` | Regenerate the catalog from every page's frontmatter: `kb/index.md` becomes a map (statistics, one row per collection and per area, links to the shards) and the page tables are written to a generated `INDEX.md` in each collection. An area past 50 rows gets its own shard. Stale shards from removed collections/areas are deleted in the same pass |
| `log append --op ingest\|query\|lint\|create\|update\|delete\|rename --title "..." [--body "..."\|--body-file path]` | Append a formatted entry to `kb/log.md` | | `log append --op ingest\|query\|lint\|create\|update\|delete\|rename --title "..." [--body "..."\|--body-file path]` | Append a formatted entry to `kb/log.md` |
| `log status` | Read-only: count `ingest` entries logged since the last `lint` entry - the deterministic trigger behind the Maintenance Schedule's "every 10 sources" full-lint cadence | | `log status` | Read-only: count `ingest` entries logged since the last `lint` entry - the deterministic trigger behind the Maintenance Schedule's "every 10 sources" full-lint cadence |
@@ -70,15 +71,15 @@ tools/wikitool <command> --help
| `docs verify` | Check the docs that mirror the code: every CLI command documented here (and vice versa), every directory under `kb/` has a `COLLECTION.md` and no directory outside it does, every collection declaring `profile:` and a `required_by_stack:` that agrees with the stack's own list, `kb/CONVENTIONS.md` naming all three tool-owned section headings if it exists at all, every stage contract present, no pre-migration `type: entity` blocks left in the contracts, and the `.gitignore` canaries clear in both directions (nothing ignored under `raw/`/`kb/`, everything ignored under `reports/` and the published skill directories) | | `docs verify` | Check the docs that mirror the code: every CLI command documented here (and vice versa), every directory under `kb/` has a `COLLECTION.md` and no directory outside it does, every collection declaring `profile:` and a `required_by_stack:` that agrees with the stack's own list, `kb/CONVENTIONS.md` naming all three tool-owned section headings if it exists at all, every stage contract present, no pre-migration `type: entity` blocks left in the contracts, and the `.gitignore` canaries clear in both directions (nothing ignored under `raw/`/`kb/`, everything ignored under `reports/` and the published skill directories) |
| `eval sessions [--json]` | List the sessions that have a trace under `reports/telemetry/`, most recent first. Read-only and exempt from the Iteration Budget Gate | | `eval sessions [--json]` | List the sessions that have a trace under `reports/telemetry/`, most recent first. Read-only and exempt from the Iteration Budget Gate |
| `eval score [--session <id>] [--json] [--markdown out.md] [--save] [--fail-on-error]` | Score one traced session: structural state from `lint`'s own checks (L1) plus trajectory rules over the trace (L2) - was a refused call repeated unchanged, was a gate flag passed without that gate having refused anything, did a publish of `kb/` pages go unlogged. Defaults to the current session. `--save` writes `reports/evals/<date>/<session>.{json,md}`. Read-only over `kb/` and exempt from the budget; see [../EVALS.md](../EVALS.md) | | `eval score [--session <id>] [--json] [--markdown out.md] [--save] [--fail-on-error]` | Score one traced session: structural state from `lint`'s own checks (L1) plus trajectory rules over the trace (L2) - was a refused call repeated unchanged, was a gate flag passed without that gate having refused anything, did a publish of `kb/` pages go unlogged. Defaults to the current session. `--save` writes `reports/evals/<date>/<session>.{json,md}`. Read-only over `kb/` and exempt from the budget; see [../EVALS.md](../EVALS.md) |
| `dist export <target> [--dry-run] [--source-repo U] [--source-commit SHA] [--release-url U] [--update-url U]` | Write a contentless, distributable copy of this repo's machinery into an empty `<target>` directory: `AGENTS.md`/`README.md`/`EVALS.md` with any `<!-- dist:strip-start -->...<!-- dist:strip-end -->` region removed, `instructions/` (minus `instructions/dev/`), `types/`, `tools/` (no venv/caches), the `.github/hooks/`+`.vibe/` session-tracing config plus `.claude/settings.json`, `kb/CONTRACT.md` (no pages, no areas), empty `raw/{articles,documents,notes,assets}/`, `VERSION`, `USER.md.template`/`SOUL.md.template` plus `kb/CONVENTIONS.md.template` and each collection's contract re-keyed as `kb/<name>/COLLECTION.md.template` (the templates ship; the filled `USER.md`/`SOUL.md`/`kb/CONVENTIONS.md`/`kb/<name>/COLLECTION.md` never do - all four bind their instance and none of them are the stack's to decide, and `find_leaks` refuses a plan carrying one), and a generated `.wikitool-release.json` stamp (version, export date, origin, and a sha256 per exported file - the base a later upgrade would compare against). The four origin options only fill stamp fields: `export` never calls git and cannot discover them. Refuses a non-empty target, and a tree with no `VERSION`. See `instructions/setup-instance.md`. One-way: there is no command that reconstructs a distributed instance into a dev instance - work on the stack in the origin repo (or a new dev instance exported from it) instead | | `dist export <target> [--dry-run] [--source-repo U] [--source-commit SHA] [--release-url U] [--update-url U]` | Write a contentless, distributable copy of this repo's machinery into an empty `<target>` directory: `AGENTS.md`/`README.md`/`EVALS.md` with any `<!-- dist:strip-start -->...<!-- dist:strip-end -->` region removed, `instructions/` (minus `instructions/dev/`), `types/` (the `root: kb` page type-specs and their schemas re-keyed as `.template`, the stack's own verbatim), `tools/` (no venv/caches), the `.github/hooks/`+`.vibe/` session-tracing config plus `.claude/settings.json`, `kb/CONTRACT.md` (no pages, no areas), empty `raw/{articles,documents,notes,assets}/`, `VERSION`, `USER.md.template`/`SOUL.md.template` plus `kb/CONVENTIONS.md.template` and each collection's contract re-keyed as `kb/<name>/COLLECTION.md.template` (the templates ship; the filled `USER.md`/`SOUL.md`/`kb/CONVENTIONS.md`/`kb/<name>/COLLECTION.md`/`types/<page-type>.md` never do - all of them bind their instance and none are the stack's to decide, and `find_leaks` refuses a plan carrying one), and a generated `.wikitool-release.json` stamp (version, export date, origin, and a sha256 per exported file - the base a later upgrade would compare against). The four origin options only fill stamp fields: `export` never calls git and cannot discover them. Refuses a non-empty target, and a tree with no `VERSION`. See `instructions/setup-instance.md`. One-way: there is no command that reconstructs a distributed instance into a dev instance - work on the stack in the origin repo (or a new dev instance exported from it) instead |
| `version show [--json]` | Print this instance's stack version and where it came from (development tree, or a distribution with its export date and origin). Bare `wikitool version` is an alias for this. Read-only, offline, and **exempt from the Iteration Budget Gate** | | `version show [--json]` | Print this instance's stack version and where it came from (development tree, or a distribution with its export date and origin). Bare `wikitool version` is an alias for this. Read-only, offline, and **exempt from the Iteration Budget Gate** |
| `version check [--url U] [--timeout S] [--json]` | Ask the origin's release feed whether a newer stack exists, and whether the step crosses a compatibility boundary (`state: current\|update\|migration\|ahead`). **The only command in `wikitool` that makes a network call** - never reached implicitly from another command, needs no key, times out, and reports an unreachable feed as an error rather than as "up to date". The feed is `$WIKITOOL_UPDATE_URL`, else the release stamp's, else the built-in origin; `$WIKITOOL_UPDATE_TOKEN` is only needed if that feed is not readable anonymously. Read-only and exempt from the budget gate | | `version check [--url U] [--timeout S] [--json]` | Ask the origin's release feed whether a newer stack exists, and whether the step crosses a compatibility boundary (`state: current\|update\|migration\|ahead`). **The only command in `wikitool` that makes a network call** - never reached implicitly from another command, needs no key, times out, and reports an unreachable feed as an error rather than as "up to date". The feed is `$WIKITOOL_UPDATE_URL`, else the release stamp's, else the built-in origin; `$WIKITOOL_UPDATE_TOKEN` is only needed if that feed is not readable anonymously. Read-only and exempt from the budget gate |
| `version notes [--version X.Y.Z]` | Print one version's `CHANGES.md` entry, for use as release notes (default: this tree's `VERSION`). Read-only and exempt from the budget gate | | `version notes [--version X.Y.Z]` | Print one version's `CHANGES.md` entry, for use as release notes (default: this tree's `VERSION`). Read-only and exempt from the budget gate |
| `version bump --major\|--minor\|--patch --title "<...>" [--breaking "<what breaks>"] [--no-migration "<reason>"] [--dry-run]` | Raise `VERSION` and open the matching `CHANGES.md` entry - heading, date and author only; the body stays the author's to write, the way `new` writes frontmatter and leaves the prose. Refuses more or fewer than one part, an empty title, and a changelog already documenting a version that is not older than the new one. Compatibility follows the **leftmost non-zero component**, which for this stack (at `1.0.0` and up, no pre-release suffixes anywhere) means MAJOR: PATCH is a fix, MINOR a compatible capability, MAJOR a version that is **not a drop-in replacement** - any hand-work on update, or a downgrade that no longer works. Whether content must be migrated is a second, independent question. A MAJOR bump therefore requires `--breaking "<what stops working>"`, which is refused on any other part, and on top of it a migration document targeting the new version or `--no-migration "<reason>"`; both are recorded in the entry. Which part a change earns stays a judgment call: the command enforces that a crossing documents itself, never that the part was chosen correctly | | `version bump --major\|--minor\|--patch --title "<...>" [--breaking "<what breaks>"] [--no-migration "<reason>"] [--dry-run]` | Raise `VERSION` and open the matching `CHANGES.md` entry - heading, date and author only; the body stays the author's to write, the way `new` writes frontmatter and leaves the prose. Refuses more or fewer than one part, an empty title, and a changelog already documenting a version that is not older than the new one. Compatibility follows the **leftmost non-zero component**, which for this stack (at `1.0.0` and up, no pre-release suffixes anywhere) means MAJOR: PATCH is a fix, MINOR a compatible capability, MAJOR a version that is **not a drop-in replacement** - any hand-work on update, or a downgrade that no longer works. Whether content must be migrated is a second, independent question. A MAJOR bump therefore requires `--breaking "<what stops working>"`, which is refused on any other part, and on top of it a migration document targeting the new version or `--no-migration "<reason>"`; both are recorded in the entry. Which part a change earns stays a judgment call: the command enforces that a crossing documents itself, never that the part was chosen correctly |
| `migrate list [--json]` | List every migration document under `instructions/migrations/`, oldest target first, with its kind. Read-only and **exempt from the Iteration Budget Gate** | | `migrate list [--json]` | List every migration document under `instructions/migrations/`, oldest target first, with its kind and obligation. Read-only and **exempt from the Iteration Budget Gate** |
| `migrate status [--json]` | Show the migrations this instance still owes, in the order they must run: every document whose `migrates_to` lies in `(kb_version, VERSION]`. Exits 1 only when `.wikitool-kb.json` is missing - the content's shape is a question the tool refuses to answer by guessing. Read-only and exempt from the budget gate | | `migrate status [--json]` | Show the migrations this instance still owes, in the order they must run: every **required** document whose `migrates_to` lies in `(kb_version, VERSION]`. `offered` documents are listed separately above the chain and never block, never count as owed, and are bounded by the applied ledger rather than by `kb_version` - taking one deliberately does not move the version, so the version cannot say whether it was taken. When a release stamp is present, also reports which shipped files this instance has since edited (from the per-file sha256 in `.wikitool-release.json`), which is what says whether an offer may be copied over or has to be reconciled by hand; without a stamp that question is reported as unanswerable rather than answered. Exits 1 only when `.wikitool-kb.json` is missing - the content's shape is a question the tool refuses to answer by guessing. Read-only and exempt from the budget gate |
| `migrate verify --from <rev> [--path P ...] [--expect-body-change] [--json] [--fail-on-error]` | Compare `kb/` against a git revision on the invariants a content migration must not change: wikilink and citation **counts** (not sets), footnote definitions, H1, and structural frontmatter. Reports added/removed pages without failing on them. `--expect-body-change` additionally flags a page whose body did not change at all. Not migration-specific - worth running after any bulk rewrite, and the one question `lint` cannot answer, since it reads a single revision and so cannot see that something went missing. Read-only and exempt from the budget gate | | `migrate verify --from <rev> [--path P ...] [--expect-body-change] [--json] [--fail-on-error]` | Compare `kb/` against a git revision on the invariants a content migration must not change: wikilink and citation **counts** (not sets), footnote definitions, H1, structural frontmatter, and the **count of generated-region marker pairs** - a page that went from one links region to two has the same set of region names and a different count, and a lost marker turns a generated region into prose the next write appends a second one beside. Reports added/removed pages without failing on them. `--expect-body-change` additionally flags a page whose body did not change at all. Not migration-specific - worth running after any bulk rewrite, and the one question `lint` cannot answer, since it reads a single revision and so cannot see that something went missing. Read-only and exempt from the budget gate |
| `migrate done <version> [--pages N] [--dry-run]` | Record one migration as applied, advancing `kb_version` in `.wikitool-kb.json` to its target. **Refuses any version that is not the next link in the chain** - skipping one leaves the corpus in a shape no version describes, and an interrupted multi-step upgrade has to be resumable rather than guessable | | `migrate done <version> [--pages N] [--dry-run]` | Record one migration as applied, advancing `kb_version` in `.wikitool-kb.json` to its target. **Refuses any version that is not the next link in the chain** - skipping one leaves the corpus in a shape no version describes, and an interrupted multi-step upgrade has to be resumable rather than guessable. An `offered` migration is recorded in the applied ledger *without* moving `kb_version` and with no ordering rule applied: it is not a link in the chain, so there is nothing to skip, and requiring the chain first would make an unrelated file upgrade wait on it. Re-recording one already in the ledger is a no-op, not an error |
| `migrate baseline <version> [--force]` | Declare `kb_version` once, for an instance predating `.wikitool-kb.json`. Refuses to overwrite an existing declaration without `--force`: advancing after a migration is `done`, which checks the chain, and this command must not become the quiet way around it | | `migrate baseline <version> [--force]` | Declare `kb_version` once, for an instance predating `.wikitool-kb.json`. Refuses to overwrite an existing declaration without `--force`: advancing after a migration is `done`, which checks the chain, and this command must not become the quiet way around it |
| `doctor [--json]` | Check that this instance is correctly configured: dependencies (Python, ripgrep), author resolution, stack version, git identity/branch/remote, published skills, kb/raw/reports/work/instructions structure, personalization (`USER.md`/`SOUL.md` present **and** filled - a file still carrying the template's sentinel is a `FAIL`, since a renamed template is not a filled one), the KB conventions (`kb/CONVENTIONS.md` present, unsentinelled, and naming all three tool-owned section headings - a `FAIL` on any of the three, because `xref`/`cite` write out of it), the environment note (`ENVIRONMENT.md` - optional, so absent is `OK`; a still-templated one is a `WARN`), generated files, and `WIKITOOL_SESSION_ID`. Read-only, exit 1 only on a `FAIL` (a missing remote, session id, or `VERSION` is a `WARN`, not a fault). Exempt from the Iteration Budget Gate | | `doctor [--json]` | Check that this instance is correctly configured: dependencies (Python, ripgrep), author resolution, stack version, git identity/branch/remote, published skills, kb/raw/reports/work/instructions structure, personalization (`USER.md`/`SOUL.md` present **and** filled - a file still carrying the template's sentinel is a `FAIL`, since a renamed template is not a filled one), the KB conventions (`kb/CONVENTIONS.md` present, unsentinelled, and naming all three tool-owned section headings - a `FAIL` on any of the three, because `xref`/`cite` write out of it), the environment note (`ENVIRONMENT.md` - optional, so absent is `OK`; a still-templated one is a `WARN`), generated files, and `WIKITOOL_SESSION_ID`. Read-only, exit 1 only on a `FAIL` (a missing remote, session id, or `VERSION` is a `WARN`, not a fault). Exempt from the Iteration Budget Gate |
@@ -191,9 +192,10 @@ is atomic, and whether a retry is safe.
| `version show` / `version notes` | `VERSION` is missing or unparseable; for `notes`, no `CHANGES.md` entry names the version asked for | Read-only | Fix `VERSION`, or write the changelog entry (`version bump` writes its heading). Safe to retry | | `version show` / `version notes` | `VERSION` is missing or unparseable; for `notes`, no `CHANGES.md` entry names the version asked for | Read-only | Fix `VERSION`, or write the changelog entry (`version bump` writes its heading). Safe to retry |
| `version check` | The feed could not be reached, answered non-JSON, or carried no `tag_name`. **Never** answers "up to date" for a question it could not ask | Read-only, no local writes | A network failure is transient - retry once, then report it. HTTP 401/403 names `$WIKITOOL_UPDATE_TOKEN`; 404 means no release exists yet or the URL points at the wrong repo | | `version check` | The feed could not be reached, answered non-JSON, or carried no `tag_name`. **Never** answers "up to date" for a question it could not ask | Read-only, no local writes | A network failure is transient - retry once, then report it. HTTP 401/403 names `$WIKITOOL_UPDATE_TOKEN`; 404 means no release exists yet or the URL points at the wrong repo |
| `version bump` | More or fewer than one of `--major/--minor/--patch`, an empty `--title`, a missing `VERSION`/`CHANGES.md`, a changelog already documenting a version not older than the new one, a boundary-crossing bump without `--breaking` or with neither a migration document nor `--no-migration`, or `--breaking`/`--no-migration` on a bump that crosses nothing | No - `VERSION` then `CHANGES.md` | **Not idempotent**: a second run bumps again. If the outcome is uncertain, read `VERSION` and the top of `CHANGES.md` before retrying | | `version bump` | More or fewer than one of `--major/--minor/--patch`, an empty `--title`, a missing `VERSION`/`CHANGES.md`, a changelog already documenting a version not older than the new one, a boundary-crossing bump without `--breaking` or with neither a migration document nor `--no-migration`, or `--breaking`/`--no-migration` on a bump that crosses nothing | No - `VERSION` then `CHANGES.md` | **Not idempotent**: a second run bumps again. If the outcome is uncertain, read `VERSION` and the top of `CHANGES.md` before retrying |
| `links show` | Page not found | Read-only | Check the exact title with `search`; a wikilink target is not always the page's stem |
| `migrate list` / `migrate status` | `list` never fails; `status` exits 1 when `.wikitool-kb.json` is missing or unreadable, or `VERSION` is | Read-only | For a missing declaration: run `migrate baseline <version>` once, then retry. Safe to retry freely otherwise | | `migrate list` / `migrate status` | `list` never fails; `status` exits 1 when `.wikitool-kb.json` is missing or unreadable, or `VERSION` is | Read-only | For a missing declaration: run `migrate baseline <version>` once, then retry. Safe to retry freely otherwise |
| `migrate verify` | Only with `--fail-on-error`: an invariant changed. Also exits 1 if `--from` is not a revision in this repository | Read-only | Exit 1 from `--fail-on-error` means "act on the findings", not "the tool is broken". A finding is never fixed by re-running - it names a page and what changed on it | | `migrate verify` | Only with `--fail-on-error`: an invariant changed. Also exits 1 if `--from` is not a revision in this repository | Read-only | Exit 1 from `--fail-on-error` means "act on the findings", not "the tool is broken". A finding is never fixed by re-running - it names a page and what changed on it |
| `migrate done` | Unknown version, no `.wikitool-kb.json`, nothing outstanding, or a version that is not the next link in the chain | Yes - single file write | **Not idempotent**: it advances the chain. For "not the next link", run `migrate status` and apply them in the order it prints - never force the order | | `migrate done` | Unknown version, no `.wikitool-kb.json`, nothing outstanding, or a *required* version that is not the next link in the chain | Yes - single file write | **Not idempotent** for a required migration: it advances the chain. For "not the next link", run `migrate status` and apply them in the order it prints - never force the order. Recording an `offered` migration *is* idempotent and safe to repeat |
| `migrate baseline` | Unparseable version, or a declaration already exists and `--force` was not passed | Yes - single file write | Safe to re-run with the same version. If a declaration exists, it is almost always `migrate done` that was wanted | | `migrate baseline` | Unparseable version, or a declaration already exists and `--force` was not passed | Yes - single file write | Safe to re-run with the same version. If a declaration exists, it is almost always `migrate done` that was wanted |
| `doctor` | At least one check reported `FAIL` (a `WARN`, e.g. no remote or no `WIKITOOL_SESSION_ID`, does not exit 1) | Read-only | Each finding names its own fix command; re-run after applying it | | `doctor` | At least one check reported `FAIL` (a `WARN`, e.g. no remote or no `WIKITOOL_SESSION_ID`, does not exit 1) | Read-only | Each finding names its own fix command; re-run after applying it |
| `budget status` / `budget reset` | `reset` without `--yes`; `status` never fails | Read/rewrite of one JSON file | `status` is safe to retry. For `reset`: get the user's approval, then re-run with `--yes` | | `budget status` / `budget reset` | `reset` without `--yes`; `status` never fails | Read/rewrite of one JSON file | `status` is safe to retry. For `reset`: get the user's approval, then re-run with `--yes` |
+23 -16
View File
@@ -39,12 +39,13 @@ tools/
errors.py ChemenuError / ValidationError / BackendError errors.py ChemenuError / ValidationError / BackendError
corpus_cache.py one parsed corpus per commit, never cached while the tree is dirty corpus_cache.py one parsed corpus per commit, never cached while the tree is dirty
kb_scan.py page iteration/loading over kb/ kb_scan.py page iteration/loading over kb/
blocks.py generated regions in a page body, found by marker rather than by heading
links.py labelled edges in `related:` - the graph's semantics as data, not prose
kb_collections.py collection discovery (a directory with COLLECTION.md), and what one declares about itself kb_collections.py collection discovery (a directory with COLLECTION.md), and what one declares about itself
conventions.py kb/CONVENTIONS.md: what this instance decided about authoring, as opposed to what the stack enforces conventions.py kb/CONVENTIONS.md: what this instance decided about authoring, as opposed to what the stack enforces
type_resolver.py type-spec loading and schema resolution type_resolver.py type-spec loading and schema resolution
lint_core.py the lint checks and the report, with no CLI attached lint_core.py the lint checks and the report, with no CLI attached
types_core.py type-spec listing/description, with no CLI attached types_core.py type-spec listing/description, with no CLI attached
sections.py the section headings the tool reads and writes in a page body
markdown_code.py masks code spans/fences so a page may show wiki notation, not only use it markdown_code.py masks code spans/fences so a page may show wiki notation, not only use it
version.py the stack version: VERSION, the release stamp, the compatibility rule version.py the stack version: VERSION, the release stamp, the compatibility rule
kb_state.py the KB version (.wikitool-kb.json) and the migration chain kb_state.py the KB version (.wikitool-kb.json) and the migration chain
@@ -111,23 +112,29 @@ procedure written down in advance is one an agent can complete alone. Whether
a human *actually* saw it is not enforced here - that question is answered in a human *actually* saw it is not enforced here - that question is answered in
the eval layer (`evals/trajectory.py`, `clearance-ended-the-turn`). the eval layer (`evals/trajectory.py`, `clearance-ended-the-turn`).
**Section names are a vocabulary, not literals - and not the stack's.** `xref add` writes into **Nothing locates a region by its prose.** `xref` owns the links region and `cite` the footnotes
Relationships and See Also, and `cite add` owns the trailing Footnotes block, so those three region, and each is delimited by a `<!-- wikitool:<name> -->` marker pair (`blocks.py`). The
headings are structure the tool matches on. *Which words they are* is the corpus's own answer: heading inside is rendered from `kb/CONVENTIONS.md` and is replaced along with the rest of the
`conventions.py` reads them from `kb/CONVENTIONS.md`, `sections.py` resolves them on access region on every write - so no heading text exists in Python, and changing the declaration cannot
(PEP 562, the way `config` resolves its paths), and no heading text is written down in Python split a page.
except the pre-conventions fallback for an instance that has not declared one yet.
Each slot has one canonical spelling - what the tool writes - plus aliases it still recognizes. Both halves of that mattered. Matching on the heading made the KB language a compiler constant;
That asymmetry is what let the wiki be translated page by page instead of atomically: an *guessing* where the region ended - at the next heading, and before that at the end of the file -
untranslated `## Relationships` is still found and appended to. Dropping an alias is therefore a silently deleted content sitting after it on eight pages. `migrate verify` compares marker-pair
breaking change for any page not yet converted, not a cleanup. Renaming a heading is a counts for the same reason it compares wikilink counts: a dropped marker is invisible otherwise.
migration's job; no other command may do it as a side effect (see `cite_block_heading` in
`provenance.py`, which exists solely so `cite sync` stays a no-op on an untranslated page).
Because the value is resolved rather than bound, nothing may capture it at import time - not a **A relationship label is data, not prose.** `related:` carries `- <label>: <target>`
module constant, not an evaluated default argument. That is why `provenance.CITE_BLOCK_HEADING` (`links.py`), the label drawn from `instructions/link-taxonomy.md` and authorised per
is a module `__getattr__` and `render_cite_block(heading=None)` resolves inside the call. destination by the *source* collection's `outbound:` block. The body bullet is a rendering of
that, which is what removed the need to parse a German phrase back into a relationship - and why
the vocabulary can be checked at all, after drifting to 152 distinct labels while it could not
be. Both readers accept a bare title as an unlabelled edge: that is the shape a page is in
between the machinery landing and the migration reaching it, and `lint` is what reports it.
**An edge is authored in one direction.** `xref add` writes one, on the asserting page. The
inbound view is rendered from the graph rather than stored, so navigation does not depend on
anyone writing a mirror - and per-collection authorisation stays coherent, which it cannot be if
the tool writes edges into a collection whose rules the author never read.
**Generated output is never committed.** `reports/`, `.agents/skills/` and **Generated output is never committed.** `reports/`, `.agents/skills/` and
`.claude/skills/` are build output; `docs verify` carries canaries in both `.claude/skills/` are build output; `docs verify` carries canaries in both
+136
View File
@@ -0,0 +1,136 @@
"""Generated regions inside a page body, found by delimiter rather than by prose.
`xref` owns the links block and `cite` owns the footnotes block. Both used to be
located by matching their **heading text** - `^## Beziehungen$` - which made a
translated heading a structural fact and put the KB language into the compiler.
It also made the region's *end* a guess: the footnotes block ran to the next
heading, and before that to the end of the file, which silently deleted whatever
sat after it on eight pages.
A marker pair answers both questions exactly:
<!-- wikitool:links -->
## Beziehungen
- **depends-on:** [[Hermes]]
<!-- /wikitool:links -->
Everything between the markers is generated and is replaced wholesale on the
next write - heading included, which is why the heading text is a *rendering*
value from `kb/CONVENTIONS.md` rather than something the tool searches for. An
author never edits inside the markers; anything they put there is overwritten
without warning, exactly like `kb/index.md`.
The markers are HTML comments: invisible in every renderer this corpus is read
through, and the same convention `dist:strip-start`/`-end` already uses in
`AGENTS.md`. They cost a reader nothing and cost an LLM about twenty tokens a
page - the price of not having to guess where a generated region ends.
"""
from __future__ import annotations
import re
from typing import Optional
# The two regions the tool owns. `see-also` is deliberately absent: it was the
# reciprocal half of the old bidirectional `xref add`, and under authored
# directional edges it is a *label* (`see-also`) inside the links block, not a
# section of its own.
LINKS = "links"
FOOTNOTES = "footnotes"
BLOCKS = (LINKS, FOOTNOTES)
_NAME = r"[a-z][a-z0-9-]*"
def open_marker(name: str) -> str:
return f"<!-- wikitool:{name} -->"
def close_marker(name: str) -> str:
return f"<!-- /wikitool:{name} -->"
def _region_re(name: str) -> re.Pattern[str]:
"""The whole region including both markers and the blank line around it."""
return re.compile(
r"\n*"
+ re.escape(open_marker(name))
+ r".*?"
+ re.escape(close_marker(name))
+ r"\n*",
re.DOTALL,
)
_ANY_OPEN_RE = re.compile(rf"<!-- wikitool:({_NAME}) -->")
_ANY_CLOSE_RE = re.compile(rf"<!-- /wikitool:({_NAME}) -->")
def find(body: str, name: str) -> Optional[str]:
"""The generated content of `name`'s region, markers excluded, or None."""
match = _region_re(name).search(body)
if not match:
return None
text = match.group(0)
start = text.index(open_marker(name)) + len(open_marker(name))
end = text.index(close_marker(name))
return text[start:end].strip("\n")
def render(name: str, heading: str, lines: list[str]) -> str:
"""A whole region, ready to place into a body. Empty `lines` renders "".
An empty region is no region at all rather than a heading with nothing under
it: a page that cites nothing should not carry an empty Footnotes section,
and the same holds for a page with no declared edges.
"""
if not lines:
return ""
parts = [open_marker(name), f"## {heading}", "", *lines, close_marker(name)]
return "\n".join(parts)
def replace(body: str, name: str, region: str) -> str:
"""Put `region` where `name`'s region is, or append it if there is none.
Appending at the end is right for both blocks: they are the page's trailing
machine-owned material, and an author's prose never follows them. A region
that is `""` removes what was there.
"""
existing = _region_re(name).search(body)
if existing:
replacement = f"\n\n{region}\n" if region else "\n"
return (body[: existing.start()] + replacement + body[existing.end():]).rstrip("\n") + "\n"
if not region:
return body
return body.rstrip("\n") + "\n\n" + region + "\n"
def strip(body: str, name: str) -> str:
"""The body with `name`'s region removed entirely."""
return replace(body, name, "")
def marker_pairs(body: str) -> dict[str, int]:
"""How many complete open/close pairs each region name has in `body`.
The invariant `migrate verify` checks. An agent rewriting prose next to a
boundary can drop or duplicate a marker, and the failure is otherwise silent:
a lost opening marker turns a generated region into ordinary prose that the
next write appends a second copy beside.
"""
opens = [m.group(1) for m in _ANY_OPEN_RE.finditer(body)]
closes = [m.group(1) for m in _ANY_CLOSE_RE.finditer(body)]
names = set(opens) | set(closes)
return {name: min(opens.count(name), closes.count(name)) for name in sorted(names)}
def unbalanced_markers(body: str) -> list[str]:
"""Region names whose open and close markers do not pair up."""
opens = [m.group(1) for m in _ANY_OPEN_RE.finditer(body)]
closes = [m.group(1) for m in _ANY_CLOSE_RE.finditer(body)]
return sorted(
name
for name in set(opens) | set(closes)
if opens.count(name) != closes.count(name)
)
+2
View File
@@ -21,6 +21,7 @@ try:
index_build, index_build,
instructions_cmd, instructions_cmd,
lint as lint_module, lint as lint_module,
links_cmd,
log_append, log_append,
migrate_cmd, migrate_cmd,
new_page, new_page,
@@ -54,6 +55,7 @@ app = typer.Typer(
app.add_typer(xref.app, name="xref") app.add_typer(xref.app, name="xref")
app.add_typer(cite_cmd.app, name="cite") app.add_typer(cite_cmd.app, name="cite")
app.add_typer(links_cmd.app, name="links")
app.add_typer(index_build.app, name="index") app.add_typer(index_build.app, name="index")
app.add_typer(log_append.app, name="log") app.add_typer(log_append.app, name="log")
app.add_typer(confidence_decay.app, name="confidence") app.add_typer(confidence_decay.app, name="confidence")
+2 -3
View File
@@ -28,7 +28,6 @@ from chemenu.kb_scan import load_kb_pages
from chemenu.provenance import ( from chemenu.provenance import (
CITE_REF_RE, CITE_REF_RE,
cite_id, cite_id,
cite_block_heading,
render_page_body, render_page_body,
split_cite_block, split_cite_block,
unique_cite_id, unique_cite_id,
@@ -81,7 +80,7 @@ def upsert_citation(page: Page, source_title: str, qualifier: Optional[str]) ->
if sources_changed: if sources_changed:
sources.append(source_title) sources.append(source_title)
new_body = render_page_body(head, definitions, cite_block_heading(page.body)) new_body = render_page_body(head, definitions)
changed = block_changed or sources_changed or new_body != page.body changed = block_changed or sources_changed or new_body != page.body
return marker_id, new_body, changed return marker_id, new_body, changed
@@ -152,7 +151,7 @@ def sync_page(page: Page) -> tuple[str, bool, list[str], list[str]]:
ordered[cid] = definitions[cid] ordered[cid] = definitions[cid]
seen.add(cid) seen.add(cid)
new_body = render_page_body(head, ordered, cite_block_heading(page.body)) new_body = render_page_body(head, ordered)
changed = new_body != page.body changed = new_body != page.body
return new_body, changed, pruned, undefined return new_body, changed, pruned, undefined
+62 -1
View File
@@ -264,6 +264,60 @@ class Origin(NamedTuple):
update_url: Optional[str] = None update_url: Optional[str] = None
def instance_owned_type_stems() -> set[str]:
"""Type-spec stems whose instances are knowledge pages, and which therefore
belong to the instance rather than to the stack.
The line is `root:`, and it was already in the frontmatter before anyone
drew it: `root: kb` means the type describes a page the instance writes, so
its prose, its template and its language are the instance's business.
Anything else - `instruction` (`root: repo`), `lint-report` (no `base_dir`
at all), `type-spec` itself - describes a stack artifact and ships verbatim.
Read from `types/` rather than listed, so an instance adding its own page
type gets the same treatment without a code change.
"""
from chemenu.type_resolver import resolver
stems: set[str] = set()
for type_path, frontmatter in resolver.list_type_specs():
stem = Path(type_path).stem
if stem == "type-spec":
continue
if not frontmatter.get("base_dir"):
continue
if (frontmatter.get("root") or "kb") != "kb":
continue
stems.add(stem)
return stems
def _plan_types() -> dict[str, PlannedFile]:
"""`types/`, with the page type-specs re-keyed as templates.
Same split as the collection contracts, for the same reason and by the same
mechanism: the shipped content is a working default rather than something
wrong for the receiver, so the file itself crosses - under a name that has
to be adopted before it counts. A type-spec's `.schema.yaml` travels with
it, because the two are one type (see types/type-spec.md § Anatomy) and
adopting half of it would leave a spec validated by a file it does not own.
"""
plan = _copy_tree(config.TYPES_DIR, "types", frozenset())
stems = instance_owned_type_stems()
if not stems:
return plan
rekeyed: dict[str, PlannedFile] = {}
for relative, planned in plan.items():
name = relative.rsplit("/", 1)[-1]
stem = name.split(".", 1)[0]
if stem in stems:
rekeyed[f"{relative}.template"] = planned
else:
rekeyed[relative] = planned
return rekeyed
def build_plan(origin: Optional[Origin] = None) -> dict[str, PlannedFile]: def build_plan(origin: Optional[Origin] = None) -> dict[str, PlannedFile]:
"""Every (destination-relative path -> planned file) the export writes.""" """Every (destination-relative path -> planned file) the export writes."""
plan: dict[str, PlannedFile] = {} plan: dict[str, PlannedFile] = {}
@@ -286,7 +340,7 @@ def build_plan(origin: Optional[Origin] = None) -> dict[str, PlannedFile]:
plan[name] = _read_planned_file(source, name) plan[name] = _read_planned_file(source, name)
plan.update(_copy_tree(config.INSTRUCTIONS_DIR, "instructions", frozenset(INSTRUCTIONS_EXCLUDE_DIRS))) plan.update(_copy_tree(config.INSTRUCTIONS_DIR, "instructions", frozenset(INSTRUCTIONS_EXCLUDE_DIRS)))
plan.update(_copy_tree(config.TYPES_DIR, "types", frozenset())) plan.update(_plan_types())
plan.update(_copy_tree( plan.update(_copy_tree(
config.ROOT / "tools", "tools", frozenset(TOOLS_EXCLUDE_DIRS), _is_coverage_output config.ROOT / "tools", "tools", frozenset(TOOLS_EXCLUDE_DIRS), _is_coverage_output
)) ))
@@ -382,6 +436,7 @@ _INSTANCE_OWNED_KB_FILES = (kb_collections.CONTRACT_NAME, conventions.CONVENTION
def find_leaks(plan: dict[str, PlannedFile]) -> list[str]: def find_leaks(plan: dict[str, PlannedFile]) -> list[str]:
"""Planned paths that carry one instance's own data instead of machinery.""" """Planned paths that carry one instance's own data instead of machinery."""
owned_types = instance_owned_type_stems()
leaks: list[str] = [] leaks: list[str] = []
for relative in sorted(plan): for relative in sorted(plan):
name = relative.rsplit("/", 1)[-1] name = relative.rsplit("/", 1)[-1]
@@ -389,6 +444,12 @@ def find_leaks(plan: dict[str, PlannedFile]) -> list[str]:
leaks.append(f"{relative} (one instance's own personalization)") leaks.append(f"{relative} (one instance's own personalization)")
elif relative.startswith("kb/") and name in _INSTANCE_OWNED_KB_FILES: elif relative.startswith("kb/") and name in _INSTANCE_OWNED_KB_FILES:
leaks.append(f"{relative} (this instance's authoring conventions; ship the .template)") leaks.append(f"{relative} (this instance's authoring conventions; ship the .template)")
elif (
relative.startswith("types/")
and not relative.endswith(".template")
and name.split(".", 1)[0] in owned_types
):
leaks.append(f"{relative} (this instance's page type-spec; ship the .template)")
elif relative.startswith("instructions/dev/"): elif relative.startswith("instructions/dev/"):
leaks.append(f"{relative} (stack-development only)") leaks.append(f"{relative} (stack-development only)")
elif relative.startswith(_CONTENT_PREFIXES) and name not in _CONTENT_ALLOWED_NAMES: elif relative.startswith(_CONTENT_PREFIXES) and name not in _CONTENT_ALLOWED_NAMES:
+43
View File
@@ -260,10 +260,53 @@ def check_collection_contracts() -> list[str]:
issues += kb_collections.declaration_issues() issues += kb_collections.declaration_issues()
issues += conventions.declaration_issues() issues += conventions.declaration_issues()
issues += check_stack_required_types()
return issues return issues
def check_stack_required_types() -> list[str]:
"""The minimum the stack asks of the type layer, and nothing beyond it.
The four page type-specs belong to the instance: it may translate them,
rewrite their templates, add sections. What it may not do is remove the one
type the provenance path is built on, or drop the field that path reads.
Everything else about `types/source.md` - its prose, its template, its title
prefix, its directory - is the instance's, and is deliberately not checked
here.
"""
from chemenu.type_resolver import resolver
issues: list[str] = []
for type_name in kb_collections.STACK_REQUIRED_TYPES:
try:
type_path = resolver.find_type_by_name(type_name)
except (ValueError, OSError) as exc:
issues.append(f"types/ could not be read to find the `{type_name}` type: {exc}")
continue
if not type_path:
issues.append(
f"no type-spec declares `name: {type_name}` - `sources coverage`, `[^cite-id]` "
f"resolution and `kb/provenance.md` all ask `page.kind == \"{type_name}\"`, so "
f"without it the whole raw/ -> kb/ provenance path resolves against nothing"
)
continue
try:
schema = resolver.get_schema(type_path) or {}
except (ValueError, OSError) as exc:
issues.append(f"{type_path}: its schema could not be read: {exc}")
continue
declared = set(schema.get("required") or [])
for field in kb_collections.STACK_REQUIRED_TYPE_FIELDS.get(type_name, ()):
if field not in declared:
issues.append(
f"{type_path}: its schema must require `{field}` - it is what the "
f"provenance path reads, and a `{type_name}` page without it claims no "
f"raw material at all"
)
return issues
def check_legacy_type_blocks() -> list[str]: def check_legacy_type_blocks() -> list[str]:
issues = [] issues = []
guarded = [ guarded = [
+14 -11
View File
@@ -217,17 +217,18 @@ def check_conventions() -> Check:
"""Whether this instance has said how its own pages are written. """Whether this instance has said how its own pages are written.
`kb/CONVENTIONS.md` carries the decisions `kb/CONTRACT.md` deliberately no `kb/CONVENTIONS.md` carries the decisions `kb/CONTRACT.md` deliberately no
longer makes: the KB language and its three tool-owned section headings, the longer makes: the KB language and the headings its two generated regions
relationship-label vocabulary, the tone examples, the confidence rubric, the render under, the tone examples, the confidence rubric, the naming forms.
ADR prefix. The compiler reads the section names out of it, so an instance
without one is not merely undocumented - `xref add` and `cite add` fall back
to the names this stack hardcoded before the file existed, which is right
only for a corpus that was written under them.
Hence `FAIL` rather than `WARN`, and hence the same two failure modes the `FAIL` rather than `WARN` because those decisions bind every page, and
personalization pair has: the distribution can ship the template but never because it has the same two failure modes the personalization pair has: the
the filled file, so a template renamed and left unanswered looks present and distribution can ship the template but never the filled file, so a template
decides nothing. renamed and left unanswered looks present and decides nothing.
The headings themselves are only cosmetic now - the marker pair carries each
region's identity, so a default renders wrong words rather than corrupting
structure. That is why this check is about the *file*, not about rescuing a
lookup the compiler can no longer get wrong.
""" """
path = conventions.conventions_file() path = conventions.conventions_file()
fix = ( fix = (
@@ -246,7 +247,9 @@ def check_conventions() -> Check:
if issues: if issues:
return Check("conventions", "FAIL", "; ".join(issues), fix) return Check("conventions", "FAIL", "; ".join(issues), fix)
declared = conventions.language() or "unspecified" declared = conventions.language() or "unspecified"
headings = ", ".join(conventions.canonical(slot) for slot in conventions.SLOTS) from chemenu import blocks
headings = ", ".join(conventions.heading(block) for block in blocks.BLOCKS)
return Check( return Check(
"conventions", "OK", "conventions", "OK",
f"kb/{conventions.CONVENTIONS_FILENAME} present, language {declared}, " f"kb/{conventions.CONVENTIONS_FILENAME} present, language {declared}, "
+95
View File
@@ -0,0 +1,95 @@
"""`wikitool links` - the declared graph around one page, both directions.
The half that makes authored directional edges liveable. An edge is written once,
on the page that asserts it, so the question "what points at *this* page" has no
answer stored anywhere - it is computed from the graph, which is the only way it
is ever complete. A mirrored edge only ever recorded what someone remembered to
mirror.
Read-only, and exempt from the iteration budget for the same reason `search` is:
it answers a question rather than changing anything, and an agent that has to
ration looking things up starts guessing instead.
"""
from __future__ import annotations
import json as _json
from typing import Optional
import typer
from chemenu import config, links
from chemenu.commands._util import console, fail
from chemenu.kb_scan import load_kb_pages
from chemenu.page import Page
app = typer.Typer(help="Show the declared edges into and out of a page.")
EDGE_FIELD = "related"
def _collection_of(page: Page) -> Optional[str]:
try:
return page.path.relative_to(config.KB_DIR).parts[0]
except (ValueError, IndexError):
return None
def outbound(pages: dict[str, Page], title: str) -> list[dict]:
"""Edges this page asserts, in file order."""
page = pages[title]
return [
{"target": edge.target, "label": edge.label, "resolves": edge.target in pages}
for edge in links.edges(page.frontmatter, EDGE_FIELD)
]
def inbound(pages: dict[str, Page], title: str) -> list[dict]:
"""Edges other pages assert *about* this one.
A full scan of the corpus rather than a stored list, deliberately: the whole
argument for dropping mirrored edges is that this answer is derived and
therefore cannot go stale or be half-written.
"""
found = [
{"source": other, "label": edge.label, "collection": _collection_of(page)}
for other, page in pages.items()
for edge in links.edges(page.frontmatter, EDGE_FIELD)
if edge.target == title
]
return sorted(found, key=lambda item: (item["label"] or "", item["source"]))
@app.command("show")
def links_show(
page: str = typer.Option(..., "--page", help="Exact page title"),
json_out: bool = typer.Option(False, "--json", help="Print the edges as JSON"),
):
"""Show the edges out of and into a page.
Outbound is what the page declares in `related:`. Inbound is computed across
the corpus - nothing stores it, which is exactly why it is complete."""
pages = load_kb_pages(config.KB_DIR)
if page not in pages:
fail(f"No page titled '{page}' found under kb/.")
out, back = outbound(pages, page), inbound(pages, page)
if json_out:
typer.echo(_json.dumps({"page": page, "outbound": out, "inbound": back}, indent=2))
return
console.print(f"[bold]{page}[/bold]")
console.print(f"\n[cyan]asserts ({len(out)})[/cyan]")
if not out:
console.print(" (none)")
for edge in out:
label = edge["label"] or "[dim]unlabelled[/dim]"
missing = "" if edge["resolves"] else " [red](no such page)[/red]"
console.print(f" {label} -> [[{edge['target']}]]{missing}")
console.print(f"\n[cyan]asserted about it ({len(back)})[/cyan]")
if not back:
console.print(" (none - nothing in the corpus declares an edge to this page)")
for edge in back:
label = edge["label"] or "[dim]unlabelled[/dim]"
console.print(f" [[{edge['source']}]] {label} ->")
+91 -3
View File
@@ -64,6 +64,7 @@ def list_command(
"name": m.name, "name": m.name,
"migrates_to": str(m.target), "migrates_to": str(m.target),
"migration_kind": m.kind, "migration_kind": m.kind,
"obligation": m.obligation,
"description": m.description, "description": m.description,
"path": m.relative_path, "path": m.relative_path,
} }
@@ -78,7 +79,10 @@ def list_command(
success(f"No migration documents under {rel_path(kb_state.migrations_dir())}.") success(f"No migration documents under {rel_path(kb_state.migrations_dir())}.")
return return
for migration in migrations: for migration in migrations:
console.print(f"[bold]{migration.target}[/bold] {migration.name} ({migration.kind})") console.print(
f"[bold]{migration.target}[/bold] {migration.name} "
f"({migration.kind}, {migration.obligation})"
)
if migration.description: if migration.description:
console.print(f" {migration.description}") console.print(f" {migration.description}")
@@ -86,6 +90,50 @@ def list_command(
# --- migrate status -------------------------------------------------------- # --- migrate status --------------------------------------------------------
def _report_offers(
offered: list["kb_state.Migration"], divergent: Optional[list[str]]
) -> None:
"""Print the optional half of `status`, above the outstanding chain.
Deliberately never affects the exit code and never says "outstanding". An
offer is the stack proposing a better default for a file the instance owns;
an instance that keeps its own version is in a correct state, not a late
one. Mixing the two is how the message that actually matters - your content
no longer fits your machinery - stops being read.
"""
if not offered:
return
console.print(
f"[cyan]{len(offered)} optional upgrade(s) available[/cyan] - none of them block:"
)
for migration in offered:
console.print(f" {migration.target} {migration.name} ({migration.kind})")
if migration.description:
console.print(f" {migration.description}")
console.print(f" {migration.relative_path}")
if divergent is None:
console.print(
" [dim]This tree carries no release stamp, so which of your files still match "
"what you were given cannot be answered here.[/dim]"
)
return
if divergent:
console.print(
f" [dim]{len(divergent)} file(s) differ from the release you installed - those are "
"yours to reconcile by hand rather than overwrite:[/dim]"
)
for relative in divergent[:10]:
console.print(f" [dim]{relative}[/dim]")
if len(divergent) > 10:
console.print(f" [dim]... and {len(divergent) - 10} more[/dim]")
else:
console.print(
" [dim]No file differs from the release you installed, so an offer can be taken "
"by copying.[/dim]"
)
@app.command("status") @app.command("status")
def status_command( def status_command(
json_out: bool = typer.Option(False, "--json", help="Print the chain as JSON"), json_out: bool = typer.Option(False, "--json", help="Print the chain as JSON"),
@@ -113,6 +161,8 @@ def status_command(
return return
pending = kb_state.chain(migrations, kb_version, stack) pending = kb_state.chain(migrations, kb_version, stack)
offered = kb_state.offers(migrations, kb_state.applied_names(kb_state.read_kb_state()))
divergent = kb_state.divergent_files()
if json_out: if json_out:
typer.echo( typer.echo(
@@ -124,6 +174,11 @@ def status_command(
{"name": m.name, "migrates_to": str(m.target), "migration_kind": m.kind} {"name": m.name, "migrates_to": str(m.target), "migration_kind": m.kind}
for m in pending for m in pending
], ],
"offered": [
{"name": m.name, "migrates_to": str(m.target), "migration_kind": m.kind}
for m in offered
],
"divergent_files": divergent,
}, },
indent=2, indent=2,
) )
@@ -131,6 +186,7 @@ def status_command(
return return
console.print(f"stack {stack}, content {kb_version}") console.print(f"stack {stack}, content {kb_version}")
_report_offers(offered, divergent)
if not pending: if not pending:
if kb_version < stack: if kb_version < stack:
console.print( console.print(
@@ -166,7 +222,12 @@ def done_command(
Refuses any version that is not the *next* link in the chain: skipping a Refuses any version that is not the *next* link in the chain: skipping a
migration is how a corpus ends up in a shape no version describes, and an migration is how a corpus ends up in a shape no version describes, and an
interrupted multi-step upgrade has to be resumable rather than guessable.""" interrupted multi-step upgrade has to be resumable rather than guessable.
An `offered` migration is recorded but does not move the version, and no
ordering rule applies to it - it is not a link in the chain. The record is
the only thing that distinguishes an offer someone took from one they
ignored, precisely because the version stays put."""
stack, kb_version = _versions() stack, kb_version = _versions()
if kb_version is None: if kb_version is None:
fail( fail(
@@ -182,6 +243,34 @@ def done_command(
return return
migrations = kb_state.load_migrations() migrations = kb_state.load_migrations()
state = kb_state.read_kb_state() or {}
# An offer is recorded but does not advance the version: it is not a link in
# the chain, so there is no ordering rule to check and nothing to skip. The
# ledger is what makes it stop being offered - without that record there
# would be no way to tell a taken offer from an ignored one, because
# `kb_version` deliberately does not move.
offered = {m.name: m for m in migrations if not m.is_required}
taken = next((m for m in offered.values() if str(m.target) == version), None)
if taken is not None:
if taken.name in kb_state.applied_names(state):
success(f"{taken.name} is already recorded as taken. Nothing to do.")
return
if dry_run:
success(f"Dry run: would record the optional {taken.name}. Nothing written.")
return
applied = list(state.get("applied") or [])
entry = {"migration": taken.name, "at": today_iso(), "obligation": kb_state.OFFERED}
if pages is not None:
entry["pages"] = pages
applied.append(entry)
kb_state.write_kb_state(kb_version, applied)
success(
f"Recorded the optional {taken.name}. Content stays at {kb_version} - an offer "
"changes a file you own, not the shape of your content."
)
return
expected = kb_state.next_link(migrations, kb_version, stack) expected = kb_state.next_link(migrations, kb_version, stack)
if expected is None: if expected is None:
fail( fail(
@@ -197,7 +286,6 @@ def done_command(
) )
return return
state = kb_state.read_kb_state() or {}
applied = list(state.get("applied") or []) applied = list(state.get("applied") or [])
entry = {"migration": expected.name, "at": today_iso()} entry = {"migration": expected.name, "at": today_iso()}
if pages is not None: if pages is not None:
+2 -7
View File
@@ -24,7 +24,7 @@ import re
import typer import typer
from chemenu import config, conventions from chemenu import config
from chemenu.commands._util import ( from chemenu.commands._util import (
check_collision, check_collision,
check_raw_files_exist, check_raw_files_exist,
@@ -166,11 +166,7 @@ def _apply_template_variables(template: str, variables: Dict[str, Any]) -> str:
"""Apply variable substitutions to a template string. """Apply variable substitutions to a template string.
Supports: Supports:
- `{field}` - plain substitution from `variables[field]`, including the - `{field}` - plain substitution from `variables[field]`
`{section.<slot>}` names this instance gave the three tool-owned
headings (see chemenu.conventions). Those are what took the KB language
out of `types/*.md`: a template writes `## {section.relationships}`, so
scaffolding a page in another language needs no edit under `types/`
- `{field|filter}` - apply a named filter (bullets, join, capitalize) - `{field|filter}` - apply a named filter (bullets, join, capitalize)
to `variables[field]`'s value, so templates can render list/enum to `variables[field]`'s value, so templates can render list/enum
frontmatter fields directly instead of the caller precomputing a frontmatter fields directly instead of the caller precomputing a
@@ -333,7 +329,6 @@ def new_page_command(
**frontmatter, **frontmatter,
"name": name, "name": name,
"today": today.isoformat(), "today": today.isoformat(),
**conventions.section_variables(),
}, },
) )
+11 -8
View File
@@ -25,7 +25,7 @@ from typing import Optional
import typer import typer
from chemenu import config from chemenu import config, links
from chemenu.commands._util import check_collision, fail, rel_path, success from chemenu.commands._util import check_collision, fail, rel_path, success
from chemenu.frontmatter_io import write_page from chemenu.frontmatter_io import write_page
from chemenu.page import Page from chemenu.page import Page
@@ -33,7 +33,6 @@ from chemenu.kb_scan import load_kb_pages
from chemenu.provenance import ( from chemenu.provenance import (
CITE_REF_RE, CITE_REF_RE,
cite_id, cite_id,
cite_block_heading,
render_page_body, render_page_body,
split_cite_block, split_cite_block,
unique_cite_id, unique_cite_id,
@@ -103,7 +102,7 @@ def retarget_cite_ids(body: str, old: str, new: str) -> str:
return body return body
new_head = CITE_REF_RE.sub(lambda m: f"[^{renames.get(m.group(1), m.group(1))}]", head) new_head = CITE_REF_RE.sub(lambda m: f"[^{renames.get(m.group(1), m.group(1))}]", head)
return render_page_body(new_head, new_definitions, cite_block_heading(body)) return render_page_body(new_head, new_definitions)
def retarget_frontmatter(page: Page, old: str, new: str) -> bool: def retarget_frontmatter(page: Page, old: str, new: str) -> bool:
@@ -114,9 +113,11 @@ def retarget_frontmatter(page: Page, old: str, new: str) -> bool:
values = page.frontmatter.get(field) values = page.frontmatter.get(field)
if not values: if not values:
continue continue
updated = [new if value == old else value for value in values] # Through `links` so a labelled edge keeps its label across a rename:
if updated != values: # the entry is `{label: target}`, and a plain equality swap would have
page.frontmatter[field] = updated # compared the mapping against a title and silently left it pointing at
# the old page.
if links.retarget(page.frontmatter, field, old, new):
changed = True changed = True
return changed return changed
@@ -157,8 +158,10 @@ def strip_frontmatter_ref(page: Page, title: str) -> bool:
values = page.frontmatter.get(field) values = page.frontmatter.get(field)
if not values: if not values:
continue continue
updated = [value for value in values if value != title] before = list(values)
if updated == values: links.remove(page.frontmatter, field, title)
updated = page.frontmatter.get(field) or []
if updated == before:
continue continue
if not updated and field not in declared: if not updated and field not in declared:
del page.frontmatter[field] del page.frontmatter[field]
+4
View File
@@ -82,6 +82,10 @@ SKIP_COMMAND_PATHS = {
("eval", "score"), ("eval", "score"),
("eval", "sessions"), ("eval", "sessions"),
("cite", "id"), ("cite", "id"),
# Retrieval, like `search`: an agent that has to ration looking up what
# points at a page starts guessing instead - and under authored directional
# edges this is the *only* way to ask that question.
("links", "show"),
("version", "show"), ("version", "show"),
("version", "check"), ("version", "check"),
("version", "notes"), ("version", "notes"),
+116 -96
View File
@@ -1,9 +1,22 @@
"""Bidirectional cross-reference management between wiki pages. """Cross-reference management between wiki pages.
`xref add` keeps two pages' frontmatter `related:` lists AND their body `xref add` writes **one** edge: a label plus a target, into the asserting page's
"## Relationships" sections in sync in one operation, instead of the 3-5 `related:` frontmatter, and re-renders that page's generated links region from
separate manual edits this used to take per pair of pages. It is idempotent: it. It is idempotent, and re-running with a different label relabels rather than
re-running it never duplicates a link. duplicating.
It used to write four things at once - `related:` and a Relationships bullet on
both pages, plus reciprocal See Also bullets. That made every edge symmetric by
construction, which is not what a link means: an edge is an authored reader aid,
and "follow this to verify the premise" rarely reads the same from the other
end. Worse, it is incompatible with per-collection label authorisation, because
the mirrored half is written into a collection whose rules the author never
read.
The reverse direction is therefore authored separately, when it is a primary
statement of its own - and navigation does not depend on anyone bothering:
`index rebuild` renders the inbound view from the graph, completely and without
maintenance. See instructions/link-taxonomy.md.
""" """
from __future__ import annotations from __future__ import annotations
@@ -12,7 +25,7 @@ from pathlib import Path
import typer import typer
from chemenu import config, sections from chemenu import blocks, config, conventions, kb_collections, links
from chemenu.commands._util import fail, parse_list, success from chemenu.commands._util import fail, parse_list, success
from chemenu.commands.page_ops import strip_frontmatter_ref from chemenu.commands.page_ops import strip_frontmatter_ref
from chemenu.frontmatter_io import write_page from chemenu.frontmatter_io import write_page
@@ -71,121 +84,119 @@ def _back_reference_field(source: Page, target: Page) -> str | None:
return collection if collection in _declared_ref_fields(source) else None return collection if collection in _declared_ref_fields(source) else None
def add_related(frontmatter: dict, other_title: str) -> bool: def _collection_of(page: Page) -> str | None:
"""Add other_title to frontmatter['related'] if not already present. """The collection a page lives in, or None if it is outside `kb/`."""
Returns True if a change was made.""" try:
related = frontmatter.setdefault("related", []) return page.path.relative_to(config.KB_DIR).parts[0]
if other_title in related: except (ValueError, IndexError):
return False
related.append(other_title)
return True
def _section_bounds(body: str, heading: str) -> tuple[int, int] | None:
match = sections.heading_re(heading).search(body)
if not match:
return None return None
start = match.end()
next_heading = re.search(r"^## ", body[start:], re.MULTILINE)
end = start + next_heading.start() if next_heading else len(body)
return start, end
def add_bullet_to_section(body: str, heading: str, bullet: str, dedup_link: str) -> str: def render_links_block(page: Page) -> str:
"""Insert `bullet` into the `## {heading}` section of body, unless a """The page's generated links region, built from its `related:` edges.
wikilink to dedup_link already appears there. Creates the section
(before the See Also section if present, else at the end) if missing.
`heading` is a canonical name from `sections`; an existing section is found The body is a *rendering* of the frontmatter, not a second place the graph
under its aliases too, so a page that has not been translated yet is still is stored. That is what removed the need to parse a German bullet back into
appended to rather than given a duplicate section. A section this creates a relationship: the label lives in the data, and this writes it out.
always carries the canonical name.""" """
bounds = _section_bounds(body, heading) lines = []
if bounds is None: for edge in links.edges(page.frontmatter, "related"):
section = f"## {heading}\n\n{bullet}\n\n" if edge.is_labelled:
see_also = sections.heading_re(sections.SEE_ALSO).search(body) lines.append(f"- **{edge.label}:** [[{edge.target}]]")
if heading != sections.SEE_ALSO and see_also: else:
return body[: see_also.start()] + section + body[see_also.start() :] lines.append(f"- [[{edge.target}]]")
return body.rstrip("\n") + "\n\n" + section.rstrip("\n") + "\n" return blocks.render(blocks.LINKS, conventions.heading(blocks.LINKS), lines)
start, end = bounds
section_text = body[start:end]
if f"[[{dedup_link}]]" in section_text:
return body
trimmed = section_text.rstrip("\n")
new_section = trimmed + "\n" + bullet + "\n\n"
return body[:start] + new_section + body[end:]
def add_relationship_bullet(body: str, label: str, other_title: str) -> str: def apply_links_block(page: Page, body: str | None = None) -> str:
bullet = f"- **{label}:** [[{other_title}]]" """`body` with the links region re-rendered from `page.frontmatter`."""
return add_bullet_to_section(body, sections.RELATIONSHIPS, bullet, other_title) return blocks.replace(
page.body if body is None else body, blocks.LINKS, render_links_block(page)
)
def add_see_also_bullet(body: str, other_title: str) -> str: def _check_authorised(source: Page, target: Page, label: str) -> None:
return add_bullet_to_section(body, sections.SEE_ALSO, f"- [[{other_title}]]", other_title) """Refuse a label the source collection has not authorised for that
destination.
Checked here rather than only in `lint` because this is the moment the
author is present: a refusal names the authorised set and can be answered by
picking a better label, while a lint finding a day later is answered by
whoever is holding the report.
"""
source_collection = _collection_of(source)
destination = _collection_of(target)
if source_collection is None or destination is None:
return
allowed = kb_collections.authorised_labels(source_collection, destination)
if not allowed:
fail(
f"kb/{source_collection}/COLLECTION.md authorises no labels for edges into "
f"kb/{destination}/. Add an `outbound:` entry for it, or do not link there "
f"from this collection."
)
if label not in allowed:
fail(
f"'{label}' is not authorised for kb/{source_collection}/ -> kb/{destination}/.\n"
f" Authorised: {', '.join(sorted(allowed))}\n"
f" The catalogue and what each label asserts: instructions/link-taxonomy.md\n"
f" Authorising a further label is a deliberate edit to "
f"kb/{source_collection}/COLLECTION.md, not a way around this refusal."
)
@app.command("add") @app.command("add")
def xref_add( def xref_add(
a: str = typer.Option(..., "--a", help="Exact title of page A"), a: str = typer.Option(..., "--a", help="Exact title of the page that asserts the edge"),
b: str = typer.Option(..., "--b", help="Exact title of page B"), b: str = typer.Option(..., "--b", help="Exact title of the page it points at"),
rel_a: str = typer.Option("related to", "--rel-a", help="Relationship label on A pointing to B"), rel: str = typer.Option(
rel_b: str = typer.Option("related to", "--rel-b", help="Relationship label on B pointing to A"), ..., "--rel", help="Label from instructions/link-taxonomy.md, e.g. depends-on"
see_also: bool = typer.Option(True, "--see-also/--no-see-also", help="Also add reciprocal 'See Also' bullets"), ),
dry_run: bool = typer.Option(False, "--dry-run", help="Preview changes to both pages instead of writing"), dry_run: bool = typer.Option(False, "--dry-run", help="Preview the change instead of writing"),
): ):
"""Declare that A <rel> B. One edge, on A only.
Say the sentence before choosing the label: `[A] <rel> [B]`. If it only
reads true backwards, the edge belongs on B - run this the other way round
rather than reaching for an inverse label.
B is not modified and does not need to point back. Its inbound view is
rendered from the graph.
"""
pages = load_kb_pages(config.KB_DIR) pages = load_kb_pages(config.KB_DIR)
page_a = _find_page(pages, a) page_a = _find_page(pages, a)
page_b = _find_page(pages, b) page_b = _find_page(pages, b)
# Both refusals before either write, so a rejected pair leaves no half-link. # Every refusal before the single write, so a rejected edge leaves nothing.
_require_related_field(page_a, a) _require_related_field(page_a, a)
_require_related_field(page_b, b) _check_authorised(page_a, page_b, rel)
related_changed_a = add_related(page_a.frontmatter, b) changed = links.upsert(page_a.frontmatter, "related", links.Edge(b, rel))
related_changed_b = add_related(page_b.frontmatter, a) body = apply_links_block(page_a)
changed = changed or body != page_a.body
body_a = add_relationship_bullet(page_a.body, rel_a, b)
body_b = add_relationship_bullet(page_b.body, rel_b, a)
if see_also:
body_a = add_see_also_bullet(body_a, b)
body_b = add_see_also_bullet(body_b, a)
changed_a = related_changed_a or body_a != page_a.body
changed_b = related_changed_b or body_b != page_b.body
if dry_run: if dry_run:
state_a = "would update" if changed_a else "already up to date" typer.echo(
state_b = "would update" if changed_b else "already up to date" f"[dry-run] '{a}': {'would declare' if changed else 'already declares'} "
typer.echo(f"[dry-run] '{a}': {state_a} (related / Relationships / See Also)") f"{rel} -> '{b}'"
typer.echo(f"[dry-run] '{b}': {state_b} (related / Relationships / See Also)") )
typer.echo("No files written (--dry-run).") typer.echo("No files written (--dry-run).")
return return
try: if not changed:
write_page(page_a.path, page_a.frontmatter, body_a) success(f"'{a}' already declares {rel} -> '{b}'; nothing changed.")
except OSError as exc: return
fail(f"Failed to write '{a}': {exc}. '{b}' was not touched - fix the write failure and retry once.")
try: try:
write_page(page_b.path, page_b.frontmatter, body_b) write_page(page_a.path, page_a.frontmatter, body)
except OSError as exc: except OSError as exc:
fail( fail(f"Failed to write '{a}': {exc}")
f"'{a}' was updated but writing '{b}' failed: {exc}. The link is now one-directional - " success(f"'{a}' {rel} '{b}'")
f"fix the write failure, then re-run `xref add --a \"{a}\" --b \"{b}\"` (idempotent, safe to retry)."
)
success(f"Linked '{a}' <-> '{b}' ({rel_a} / {rel_b})")
def remove_related(frontmatter: dict, other_title: str) -> bool: def remove_related(frontmatter: dict, other_title: str) -> bool:
"""Drop other_title from frontmatter['related'] if present. Returns True if """Drop every edge pointing at other_title. True if a change was made."""
a change was made.""" return links.remove(frontmatter, "related", other_title)
related = frontmatter.get("related")
if not related or other_title not in related:
return False
frontmatter["related"] = [title for title in related if title != other_title]
return True
def remove_link_bullets(body: str, other_title: str) -> str: def remove_link_bullets(body: str, other_title: str) -> str:
"""Remove the whole-line Relationships/See Also bullets `xref add` writes - """Remove the whole-line Relationships/See Also bullets `xref add` writes -
@@ -223,14 +234,19 @@ def xref_remove(
page_a = _find_page(pages, a) page_a = _find_page(pages, a)
page_b = pages.get(b) page_b = pages.get(b)
body_a = remove_link_bullets(page_a.body, b) # The frontmatter first, then the region re-rendered from it - the body is a
changed_a = strip_frontmatter_ref(page_a, b) or body_a != page_a.body # rendering, so editing the bullet out directly would leave an empty region
# behind and, worse, put the two out of step.
changed_a = strip_frontmatter_ref(page_a, b)
body_a = apply_links_block(page_a, remove_link_bullets(page_a.body, b))
changed_a = changed_a or body_a != page_a.body
changed_b = False changed_b = False
body_b = "" body_b = ""
if page_b is not None: if page_b is not None:
body_b = remove_link_bullets(page_b.body, a) changed_b = strip_frontmatter_ref(page_b, a)
changed_b = strip_frontmatter_ref(page_b, a) or body_b != page_b.body body_b = apply_links_block(page_b, remove_link_bullets(page_b.body, a))
changed_b = changed_b or body_b != page_b.body
if dry_run: if dry_run:
typer.echo(f"[dry-run] '{a}': {'would update' if changed_a else 'no reference to remove'}") typer.echo(f"[dry-run] '{a}': {'would update' if changed_a else 'no reference to remove'}")
@@ -278,7 +294,11 @@ def xref_link_source(
sources = page.frontmatter.setdefault("sources", []) sources = page.frontmatter.setdefault("sources", [])
if source not in sources: if source not in sources:
sources.append(source) sources.append(source)
body = add_see_also_bullet(page.body, source) # No body bullet. `sources:` *is* the record, and the See Also bullet
# this used to add was the reciprocal half of a bidirectional model
# that no longer exists - 353 of the corpus's 555 such bullets were
# provably redundant with an edge that already said the same thing.
body = page.body
# The way back. Until this existed the command wrote only the targets, # The way back. Until this existed the command wrote only the targets,
# so a source page's own `entities:`/`concepts:` stayed as `new` left # so a source page's own `entities:`/`concepts:` stayed as `new` left
+37 -84
View File
@@ -35,30 +35,30 @@ from chemenu.frontmatter_io import read_page
CONVENTIONS_FILENAME = "CONVENTIONS.md" CONVENTIONS_FILENAME = "CONVENTIONS.md"
CONVENTIONS_TEMPLATE = f"{CONVENTIONS_FILENAME}.template" CONVENTIONS_TEMPLATE = f"{CONVENTIONS_FILENAME}.template"
# The three tool-owned headings, by slot name. The slot is the stable # The two tool-owned regions, keyed by the block name in `chemenu.blocks`. The
# identifier - it is what code, the type-spec templates and the conventions # block name is the identifier - it is what the marker pair carries and what the
# file all key on - while the heading text itself is the instance's to choose. # tool locates the region by - while the heading text below it is prose the
RELATIONSHIPS = "relationships" # instance chooses.
SEE_ALSO = "see_also" #
FOOTNOTES = "footnotes" # `see_also` is gone as a section: it was the reciprocal half of the old
SLOTS = (RELATIONSHIPS, SEE_ALSO, FOOTNOTES) # bidirectional `xref add`, and under authored directional edges it is a *label*
# inside the links block rather than a region of its own.
# Frontmatter keys read out of kb/CONVENTIONS.md.
SECTIONS_KEY = "sections" SECTIONS_KEY = "sections"
SECTION_ALIASES_KEY = "section_aliases"
LANGUAGE_KEY = "language" LANGUAGE_KEY = "language"
# Every heading name this stack has ever written as canonical, newest first. # What a heading renders as when the instance has not said. Purely cosmetic, and
# Two jobs, and they are separate: the first entry is the fallback for an # that is a genuine change from before: while the tool located a region by
# instance that has no conventions file yet, and the whole tuple is an implicit # matching this text, a wrong default silently split a page into two sections and
# alias set that every instance recognizes regardless of what it declares. The # `xref add` appended to the wrong one. Now the marker pair carries the identity,
# second is what makes a corpus translatable page by page - a page still # so a region rendered under the wrong words is a *display* fault that the next
# carrying `## Footnotes` is untranslated, not broken, and `cite sync` has to # write repairs by itself once `kb/CONVENTIONS.md` says otherwise.
# stay a no-op on it. #
PRE_CONVENTIONS_NAMES: dict[str, tuple[str, ...]] = { # So this is a fallback for the window between installing the machinery and
RELATIONSHIPS: ("Beziehungen", "Relationships"), # writing the conventions file - `doctor` is what makes that window loud - and
SEE_ALSO: ("Siehe auch", "See Also"), # not a language the compiler has an opinion about.
FOOTNOTES: ("Fußnoten", "Footnotes"), DEFAULT_HEADINGS: dict[str, str] = {
"links": "Relationships",
"footnotes": "Footnotes",
} }
@@ -119,44 +119,12 @@ def language() -> Optional[str]:
return str(value).strip() or None return str(value).strip() or None
def canonical(slot: str) -> str: def heading(block: str) -> str:
"""The heading name this instance writes for `slot`.""" """The heading this instance renders above `block`'s generated region."""
declared = _mapping(SECTIONS_KEY).get(slot) declared = _mapping(SECTIONS_KEY).get(block)
if isinstance(declared, str) and declared.strip(): if isinstance(declared, str) and declared.strip():
return declared.strip() return declared.strip()
return PRE_CONVENTIONS_NAMES[slot][0] return DEFAULT_HEADINGS.get(block, block.title())
def names(slot: str) -> tuple[str, ...]:
"""Every heading name `slot` is recognized under, canonical first.
The canonical name, then any `section_aliases:` the instance declared, then
the names this stack wrote before the conventions file existed. Deduplicated
while preserving that order, so an instance declaring English does not end
up with `Relationships` listed twice.
"""
declared_aliases = _mapping(SECTION_ALIASES_KEY).get(slot)
extra = declared_aliases if isinstance(declared_aliases, list) else []
ordered = [
canonical(slot),
*(str(name).strip() for name in extra if str(name).strip()),
*PRE_CONVENTIONS_NAMES[slot],
]
seen: dict[str, None] = {}
for name in ordered:
seen.setdefault(name, None)
return tuple(seen)
def section_variables() -> dict[str, str]:
"""The `{section.<slot>}` substitutions a type-spec template can use.
This is what took the three German headings out of `types/*.md`: a template
writes `## {section.relationships}` and the instance's own conventions fill
it in, so scaffolding a page in another language needs no edit under
`types/`.
"""
return {f"section.{slot}": canonical(slot) for slot in SLOTS}
def declaration_issues() -> list[str]: def declaration_issues() -> list[str]:
@@ -167,6 +135,8 @@ def declaration_issues() -> list[str]:
looks like. An absent file is *not* reported here - that is a separate looks like. An absent file is *not* reported here - that is a separate
finding with a separate fix, and only `doctor` makes it one. finding with a separate fix, and only `doctor` makes it one.
""" """
from chemenu import blocks
path = conventions_file() path = conventions_file()
if not path.is_file(): if not path.is_file():
return [] return []
@@ -176,44 +146,27 @@ def declaration_issues() -> list[str]:
if not frontmatter: if not frontmatter:
return [ return [
f"kb/{CONVENTIONS_FILENAME} has no readable frontmatter - it must declare " f"kb/{CONVENTIONS_FILENAME} has no readable frontmatter - it must declare "
f"`{SECTIONS_KEY}:` with the heading names this instance writes" f"`{SECTIONS_KEY}:` with the headings this instance renders"
] ]
declared = frontmatter.get(SECTIONS_KEY) declared = frontmatter.get(SECTIONS_KEY)
if not isinstance(declared, dict): if not isinstance(declared, dict):
return [ return [
f"kb/{CONVENTIONS_FILENAME}: `{SECTIONS_KEY}:` must be a mapping of " f"kb/{CONVENTIONS_FILENAME}: `{SECTIONS_KEY}:` must be a mapping of "
f"{'/'.join(SLOTS)} to the heading text this instance writes" f"{'/'.join(blocks.BLOCKS)} to the heading this instance renders above it"
] ]
for slot in SLOTS: for block in blocks.BLOCKS:
value = declared.get(slot) value = declared.get(block)
if not isinstance(value, str) or not value.strip(): if not isinstance(value, str) or not value.strip():
issues.append( issues.append(
f"kb/{CONVENTIONS_FILENAME}: `{SECTIONS_KEY}.{slot}` is missing or empty - " f"kb/{CONVENTIONS_FILENAME}: `{SECTIONS_KEY}.{block}` is missing or empty - "
"`xref add` and `cite add` write into a heading this instance has not named" f"the generated `{block}` region would render under a default heading rather "
"than this instance's own"
) )
for slot in sorted(set(declared) - set(SLOTS)): for block in sorted(set(declared) - set(blocks.BLOCKS)):
issues.append( issues.append(
f"kb/{CONVENTIONS_FILENAME}: `{SECTIONS_KEY}.{slot}` is not a section the tool " f"kb/{CONVENTIONS_FILENAME}: `{SECTIONS_KEY}.{block}` is not a region the tool "
f"owns; the slots are {', '.join(SLOTS)}" f"generates; the regions are {', '.join(blocks.BLOCKS)}"
)
aliases = frontmatter.get(SECTION_ALIASES_KEY, {})
if not isinstance(aliases, dict):
issues.append(
f"kb/{CONVENTIONS_FILENAME}: `{SECTION_ALIASES_KEY}:` must be a mapping of a "
"slot to the list of headings still recognized under it"
)
else:
for slot, value in sorted(aliases.items()):
if slot not in SLOTS:
issues.append(
f"kb/{CONVENTIONS_FILENAME}: `{SECTION_ALIASES_KEY}.{slot}` is not a "
f"section the tool owns; the slots are {', '.join(SLOTS)}"
)
elif not isinstance(value, list):
issues.append(
f"kb/{CONVENTIONS_FILENAME}: `{SECTION_ALIASES_KEY}.{slot}` must be a list"
) )
if config.TEMPLATE_SENTINEL in path.read_text(encoding="utf-8"): if config.TEMPLATE_SENTINEL in path.read_text(encoding="utf-8"):
+20 -1
View File
@@ -26,7 +26,7 @@ from collections import Counter
from dataclasses import dataclass, field from dataclasses import dataclass, field
from typing import Any, Optional from typing import Any, Optional
from chemenu import kb_scan, provenance from chemenu import blocks, kb_scan, provenance
from chemenu.page import Page from chemenu.page import Page
from chemenu.type_resolver import resolver from chemenu.type_resolver import resolver
@@ -59,6 +59,7 @@ class PageShape:
cite_refs: Counter cite_refs: Counter
cite_defs: dict[str, str] cite_defs: dict[str, str]
fields: dict[str, Any] fields: dict[str, Any]
markers: dict[str, int]
body: str body: str
@classmethod @classmethod
@@ -81,6 +82,7 @@ class PageShape:
cite_refs=Counter(m.group(1) for m in provenance.CITE_REF_RE.finditer(head)), cite_refs=Counter(m.group(1) for m in provenance.CITE_REF_RE.finditer(head)),
cite_defs={cite_id: source for cite_id, (source, _) in definitions.items()}, cite_defs={cite_id: source for cite_id, (source, _) in definitions.items()},
fields=fields, fields=fields,
markers=blocks.marker_pairs(page.body),
body=page.body, body=page.body,
) )
@@ -163,6 +165,23 @@ def compare_page(path: str, before: PageShape, after: PageShape) -> list[PageFin
changed.append(f"[^{cite_id}] {was!r} -> {now!r}") changed.append(f"[^{cite_id}] {was!r} -> {now!r}")
findings.append(PageFinding(path, "cite-defs", ", ".join(changed))) findings.append(PageFinding(path, "cite-defs", ", ".join(changed)))
# A generated region that lost or gained a marker is the failure mode the
# delimiters were introduced against, and it is silent: a lost opening
# marker turns the region into ordinary prose, and the next write appends a
# second region beside it. An agent rewriting prose at the boundary is
# exactly how that happens, which is what makes it a migration invariant
# rather than a lint nicety.
#
# Counts, not presence - the same reasoning as the wikilink counter. A page
# that goes from one links region to two has the same *set* of region names.
if before.markers != after.markers:
changed_regions = []
for name in sorted(set(before.markers) | set(after.markers)):
was, now = before.markers.get(name, 0), after.markers.get(name, 0)
if was != now:
changed_regions.append(f"{name}: {was} -> {now}")
findings.append(PageFinding(path, "markers", ", ".join(changed_regions)))
changed_fields = [] changed_fields = []
for name in sorted(set(before.fields) | set(after.fields)): for name in sorted(set(before.fields) | set(after.fields)):
was, now = before.fields.get(name), after.fields.get(name) was, now = before.fields.get(name), after.fields.get(name)
+32 -1
View File
@@ -267,10 +267,41 @@ def _format_list(items: list[Any]) -> str:
return "[" + ", ".join(_format_scalar(v, flow=True) for v in items) + "]" return "[" + ", ".join(_format_scalar(v, flow=True) for v in items) + "]"
def _is_single_key_mapping(value: Any) -> bool:
return isinstance(value, dict) and len(value) == 1
def _format_mapping_list(key: str, items: list[Any]) -> str:
"""A list holding `label: target` pairs, rendered block-style.
The inline `[...]` form this file uses everywhere else cannot carry a
mapping without quoting rules nobody reading the file would guess, so a
labelled edge list is the one place block style earns its keep:
related:
- depends-on: Hermes
- Borealis
Bare strings mixed in stay bare - that is an edge whose label has not been
declared yet, and promoting it to some default here would erase exactly what
`lint` is looking for.
"""
lines = [f"{key}:"]
for item in items:
if _is_single_key_mapping(item):
(label, target), = item.items()
lines.append(f" - {_format_scalar(label)}: {_format_scalar(target)}")
else:
lines.append(f" - {_format_scalar(item)}")
return "\n".join(lines)
def dump_frontmatter(frontmatter: dict[str, Any]) -> str: def dump_frontmatter(frontmatter: dict[str, Any]) -> str:
lines = [] lines = []
for key, value in frontmatter.items(): for key, value in frontmatter.items():
if isinstance(value, list): if isinstance(value, list) and any(_is_single_key_mapping(v) for v in value):
lines.append(_format_mapping_list(key, value))
elif isinstance(value, list):
lines.append(f"{key}: {_format_list(value)}") lines.append(f"{key}: {_format_list(value)}")
else: else:
lines.append(f"{key}: {_format_scalar(value)}") lines.append(f"{key}: {_format_scalar(value)}")
+87 -20
View File
@@ -37,14 +37,59 @@ CONTRACT_NAME = "COLLECTION.md"
PROFILE_FIELD = "profile" PROFILE_FIELD = "profile"
REQUIRED_BY_STACK_FIELD = "required_by_stack" REQUIRED_BY_STACK_FIELD = "required_by_stack"
# Collections `wikitool` itself depends on by name, as opposed to ones that # Which link labels a page in this collection may use, per destination
# merely hold pages. `sources` is here because three parts of the stack resolve # collection. The **source** collection decides, which is the whole point: an
# against it rather than against a page's type: `sources coverage` asks which # edge is an authored reader-aid written on the page that asserts it, so the
# raw files no source page claims, every `[^cite-id]` footnote resolves to a # rules that govern it are the rules of the collection that page lives in. A
# page in it, and `sources rebuild-index` writes `kb/provenance.md` from it. An # destination is another collection's name, or `any`.
# instance may add, rename or drop any collection that is not on this list; #
# renaming one that is leaves those three with nothing to resolve against. # This is Commonplace's ADR-019 adopted directly, and it is what makes a
STACK_REQUIRED_COLLECTIONS = ("sources",) # 35-label catalogue usable: a collection authorises the six that make sense
# from it, and the rest of the palette is simply not on its menu.
OUTBOUND_FIELD = "outbound"
ANY_DESTINATION = "any"
# The types `wikitool` itself depends on existing, as opposed to ones an
# instance keeps because they are useful. `source` is here because the whole
# `raw/ -> kb/` provenance path is built on it: `sources coverage` asks which
# raw files no source page claims, every `[^cite-id]` resolves to a source page,
# and `sources rebuild-index` writes `kb/provenance.md` from them. All three ask
# `page.kind == "source"`, so what is load-bearing is the type-spec's `name:`
# and its schema requiring `raw_files:` - not the directory, not the title
# prefix, and not a word of its prose or its template.
#
# That is the whole anchor, and it is deliberately this small: the four page
# type-specs belong to the instance (see types/type-spec.md), so anything more
# would be the stack reaching into a file it does not own.
STACK_REQUIRED_TYPES = ("source",)
STACK_REQUIRED_TYPE_FIELDS = {"source": ("raw_files",)}
def stack_required_collections() -> tuple[str, ...]:
"""Collection names an instance may not rename or drop.
**Derived, not listed.** The required collection is whichever one the
required type writes into - so an instance that legitimately renames
`kb/sources/` to something else, and says so in the type-spec's `base_dir:`,
stays consistent instead of tripping a constant that hardcoded the old name.
A second literal list would only be a copy that drifts.
"""
from chemenu.type_resolver import resolver
names: list[str] = []
for type_name in STACK_REQUIRED_TYPES:
try:
type_path = resolver.find_type_by_name(type_name)
if not type_path:
continue
if resolver.get_root(type_path) != "kb":
continue
base_dir = resolver.get_base_dir(type_path)
except (ValueError, OSError):
continue
if base_dir:
names.append(str(base_dir).strip("/"))
return tuple(dict.fromkeys(names))
def iter_kb_collections(kb_dir: Path | None = None) -> list[Path]: def iter_kb_collections(kb_dir: Path | None = None) -> list[Path]:
@@ -119,6 +164,25 @@ def collection_declaration(collection: Path) -> dict[str, Any]:
return frontmatter return frontmatter
def authorised_labels(source: str, destination: str, kb_dir: Path | None = None) -> set[str]:
"""Labels a page in `source` may use on an edge into `destination`.
The union of the destination's own entry and `any`. An empty result means
the collection authorises nothing for that destination - which is a real
answer ("do not link there from here"), not a missing declaration.
"""
root = kb_dir if kb_dir is not None else config.KB_DIR
declared = collection_declaration(root / source).get(OUTBOUND_FIELD)
if not isinstance(declared, dict):
return set()
labels: set[str] = set()
for key in (destination, ANY_DESTINATION):
entry = declared.get(key)
if isinstance(entry, list):
labels.update(str(label).strip() for label in entry if str(label).strip())
return labels
def declaration_issues(kb_dir: Path | None = None) -> list[str]: def declaration_issues(kb_dir: Path | None = None) -> list[str]:
"""What each `COLLECTION.md` fails to declare about itself. """What each `COLLECTION.md` fails to declare about itself.
@@ -127,19 +191,22 @@ def declaration_issues(kb_dir: Path | None = None) -> list[str]:
free text, because the profile catalogue is a palette rather than an enum, free text, because the profile catalogue is a palette rather than an enum,
and a collection an instance invented has no entry there to name. and a collection an instance invented has no entry there to name.
`required_by_stack:` is not the instance's to choose at all: it must agree `required_by_stack:` is not the instance's to choose at all: it must agree
with `STACK_REQUIRED_COLLECTIONS`, so a collection whose contract claims the with what `stack_required_collections()` derives from the required types, so
stack depends on it - or one the stack does depend on and that says it does a collection whose contract claims the stack depends on it - or one the
not - is a finding rather than a preference. stack does depend on and that says it does not - is a finding rather than a
preference.
""" """
root = kb_dir if kb_dir is not None else config.KB_DIR root = kb_dir if kb_dir is not None else config.KB_DIR
issues: list[str] = [] issues: list[str] = []
required = stack_required_collections()
present = {path.name for path in iter_kb_collections(root)} present = {path.name for path in iter_kb_collections(root)}
for name in STACK_REQUIRED_COLLECTIONS: for name in required:
if name not in present: if name not in present:
issues.append( issues.append(
f"kb/{name}/ is missing - `sources coverage`, `[^cite-id]` resolution and " f"kb/{name}/ is missing - it is where the stack-required `source` type writes, "
f"`kb/provenance.md` all resolve against it by name" f"and `sources coverage`, `[^cite-id]` resolution and `kb/provenance.md` all "
f"depend on those pages existing"
) )
for collection in iter_kb_collections(root): for collection in iter_kb_collections(root):
@@ -159,17 +226,17 @@ def declaration_issues(kb_dir: Path | None = None) -> list[str]:
f"instructions/kb-profiles.md this collection adopted, or `none`" f"instructions/kb-profiles.md this collection adopted, or `none`"
) )
required = declared.get(REQUIRED_BY_STACK_FIELD) required_flag = declared.get(REQUIRED_BY_STACK_FIELD)
expected = collection.name in STACK_REQUIRED_COLLECTIONS expected = collection.name in required
if not isinstance(required, bool): if not isinstance(required_flag, bool):
issues.append( issues.append(
f"{relative}: `{REQUIRED_BY_STACK_FIELD}:` is missing or not a boolean - " f"{relative}: `{REQUIRED_BY_STACK_FIELD}:` is missing or not a boolean - "
f"it must be {str(expected).lower()} for this collection" f"it must be {str(expected).lower()} for this collection"
) )
elif required != expected: elif required_flag != expected:
issues.append( issues.append(
f"{relative}: `{REQUIRED_BY_STACK_FIELD}: {str(required).lower()}` contradicts " f"{relative}: `{REQUIRED_BY_STACK_FIELD}: {str(required_flag).lower()}` "
f"the stack, which " f"contradicts the stack, which "
+ ( + (
"does depend on this collection by name" "does depend on this collection by name"
if expected if expected
+101 -6
View File
@@ -39,6 +39,16 @@ def kb_state_file() -> Path:
return config.ROOT / KB_STATE_FILENAME return config.ROOT / KB_STATE_FILENAME
# Whether a migration has to run, as opposed to how it is carried out. The two
# are independent: a `mechanical` migration can be optional and an `assisted`
# one mandatory. Keeping them on one axis is what would make `migrate status`
# cry wolf - an instance nagged about an improvement it declined stops reading
# the nag that means its content no longer fits the machinery.
REQUIRED = "required"
OFFERED = "offered"
OBLIGATIONS = (REQUIRED, OFFERED)
@dataclass(frozen=True) @dataclass(frozen=True)
class Migration: class Migration:
"""One migration document under `instructions/migrations/`.""" """One migration document under `instructions/migrations/`."""
@@ -48,6 +58,11 @@ class Migration:
kind: str # "mechanical" | "assisted" kind: str # "mechanical" | "assisted"
description: str description: str
path: Path path: Path
obligation: str = REQUIRED
@property
def is_required(self) -> bool:
return self.obligation != OFFERED
@property @property
def relative_path(self) -> str: def relative_path(self) -> str:
@@ -130,6 +145,7 @@ def load_migrations() -> list[Migration]:
target = Version.parse(str(raw_target)) target = Version.parse(str(raw_target))
except VersionError: except VersionError:
continue continue
obligation = str(frontmatter.get("obligation") or REQUIRED)
migrations.append( migrations.append(
Migration( Migration(
name=str(frontmatter.get("name") or path.stem), name=str(frontmatter.get("name") or path.stem),
@@ -137,6 +153,7 @@ def load_migrations() -> list[Migration]:
kind=str(frontmatter.get("migration_kind") or "assisted"), kind=str(frontmatter.get("migration_kind") or "assisted"),
description=str(frontmatter.get("description") or ""), description=str(frontmatter.get("description") or ""),
path=path, path=path,
obligation=obligation if obligation in OBLIGATIONS else REQUIRED,
) )
) )
return sorted(migrations, key=lambda m: m.target) return sorted(migrations, key=lambda m: m.target)
@@ -147,13 +164,46 @@ def chain(
) -> list[Migration]: ) -> list[Migration]:
"""The migrations still owed, in the order they must run. """The migrations still owed, in the order they must run.
Every migration whose target lies in `(kb_version, stack_version]`, oldest Every **required** migration whose target lies in
first. An instance at 1.3.1 upgrading to 2.0.0 gets 1.4.0, 1.7.0, 2.0.0 - `(kb_version, stack_version]`, oldest first. An instance at 1.3.1 upgrading
and the absence of any migration targeting 1.3.x is not a special case, it to 2.0.0 gets 1.4.0, 1.7.0, 2.0.0 - and the absence of any migration
simply is not in the interval. Targets above the installed machinery are targeting 1.3.x is not a special case, it simply is not in the interval.
excluded: the instance has no code for them yet. Targets above the installed machinery are excluded: the instance has no code
for them yet.
`offered` migrations are deliberately absent. They are not links in the
version chain: declining one leaves the content in a shape the machinery
still accepts, so counting it as owed would make `kb_version` unreachable
for an instance that simply kept its own file.
""" """
return [m for m in migrations if kb_version < m.target <= stack_version] return [
m for m in migrations if m.is_required and kb_version < m.target <= stack_version
]
def applied_names(state: Optional[dict]) -> set[str]:
"""Every migration this instance has recorded as carried out."""
entries = (state or {}).get("applied") or []
return {
str(entry.get("migration"))
for entry in entries
if isinstance(entry, dict) and entry.get("migration")
}
def offers(migrations: list[Migration], applied: set[str]) -> list[Migration]:
"""Optional upgrades this instance has not taken, oldest target first.
Bounded by the **applied ledger**, not by `kb_version`, and that is not a
detail: taking an offer deliberately does not move `kb_version`, so the
version says nothing about whether an offer was taken. Filtering by it
would hide every offer the moment some unrelated required migration ran.
Not bounded above by the stack version either. An offer is about a file the
instance owns rather than about the shape of its content, so it stays on the
table until it is recorded - or until the operator deletes the document.
"""
return [m for m in migrations if not m.is_required and m.name not in applied]
def next_link( def next_link(
@@ -161,3 +211,48 @@ def next_link(
) -> Optional[Migration]: ) -> Optional[Migration]:
pending = chain(migrations, kb_version, stack_version) pending = chain(migrations, kb_version, stack_version)
return pending[0] if pending else None return pending[0] if pending else None
# --- what this instance changed about what it was given --------------------
def divergent_files() -> Optional[list[str]]:
"""Files whose content no longer matches the release this instance installed.
Reads the per-file sha256 in `.wikitool-release.json`, which `dist export`
has been writing since the stamp existed and which nothing has read until
now. Its own docstring says why it is there: it is the only way a later
upgrade can tell a file the instance *edited* from one it merely *received*.
That distinction is what makes an `offered` migration actionable. The stack
proposing a better `entity` template needs to know whether it may be copied
over or whether the instance has its own version that a person has to
reconcile - and only the recorded hash can answer that.
Returns None when the question is unanswerable (a development tree, which
carries no stamp), which is different from `[]` (nothing diverged).
"""
import hashlib
from chemenu import version as version_mod
try:
stamp = version_mod.read_stamp()
except VersionError:
return None
if not stamp:
return None
recorded = stamp.get("files")
if not isinstance(recorded, dict):
return None
divergent: list[str] = []
for relative, digest in sorted(recorded.items()):
path = config.ROOT / relative
if not path.is_file():
divergent.append(relative)
continue
current = "sha256:" + hashlib.sha256(path.read_bytes()).hexdigest()
if current != digest:
divergent.append(relative)
return divergent
+143
View File
@@ -0,0 +1,143 @@
"""Labelled edges in a page's `related:` frontmatter.
An edge is a **label plus a target**, and the label is an identifier rather than
prose:
related:
- depends-on: Hermes
- implements: Hybrid Search
It used to be a bare list of titles with the label written only into a body
bullet - which meant the graph's semantics lived in German prose the tool had to
parse back, and the vocabulary drifted to 152 distinct labels in 337 bullets
because nothing could check it. The label moves into the data; the body bullet
becomes a rendering of the data.
**Both shapes read.** A bare string is an edge whose label is not yet declared,
which is exactly the state a page is in between the machinery landing and the
corpus migration reaching that page. Readers therefore never crash on the old
shape, and `lint` is what reports it - the migration is finished when no
unlabelled edge is left.
Direction is authored, never mirrored: an edge lives on the page that asserts
it, and the inbound view is rendered from the graph rather than stored. See
instructions/link-taxonomy.md.
"""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any, Iterable, Optional
# The label a not-yet-migrated bare-string edge reports as. Deliberately not a
# real catalogue label: it must be impossible for an instance to authorise it,
# so `lint` cannot be satisfied by declaring the placeholder legal.
UNLABELLED = None
@dataclass(frozen=True)
class Edge:
"""One declared relationship: what this page asserts about `target`."""
target: str
label: Optional[str] = UNLABELLED
@property
def is_labelled(self) -> bool:
return bool(self.label)
def parse_entry(entry: Any) -> Optional[Edge]:
"""One `related:` element as an Edge, or None if it is not one at all.
A single-key mapping is a labelled edge; a bare string is an unlabelled one.
Anything else - a multi-key mapping, a list, a number - is malformed, and
returning None rather than guessing is what lets `lint` report it as a
finding instead of a reader silently inventing an edge.
"""
if isinstance(entry, str):
title = entry.strip()
return Edge(title) if title else None
if isinstance(entry, dict) and len(entry) == 1:
(label, target), = entry.items()
label, target = str(label).strip(), str(target).strip()
return Edge(target, label) if label and target else None
return None
def edges(frontmatter: dict[str, Any], field: str) -> list[Edge]:
"""Every well-formed edge in `field`, in file order."""
parsed = (parse_entry(entry) for entry in (frontmatter.get(field) or []))
return [edge for edge in parsed if edge is not None]
def malformed(frontmatter: dict[str, Any], field: str) -> list[Any]:
"""Elements of `field` that are neither a title nor a `label: target` pair."""
return [
entry for entry in (frontmatter.get(field) or []) if parse_entry(entry) is None
]
def targets(frontmatter: dict[str, Any], field: str) -> list[str]:
"""Just the page titles in `field`, labelled or not.
The compatibility seam. Every caller that only ever wanted "which pages does
this reference" - dangling-reference checks, `rename`, `rm`, the link graph -
goes through here and is untouched by the label carried alongside.
"""
return [edge.target for edge in edges(frontmatter, field)]
def render(edge_list: Iterable[Edge]) -> list[Any]:
"""Edges back into frontmatter form, ready for `dump_frontmatter`.
An unlabelled edge round-trips as a bare string rather than being promoted
to some default label: inventing one here would erase the very thing `lint`
is looking for.
"""
rendered: list[Any] = []
for edge in edge_list:
rendered.append({edge.label: edge.target} if edge.is_labelled else edge.target)
return rendered
def upsert(frontmatter: dict[str, Any], field: str, edge: Edge) -> bool:
"""Add or relabel `edge` in `field`. True if anything changed.
Idempotent by target: one page asserts one thing about another, so a second
call with a different label *replaces* rather than appends. Two edges to the
same target would render two bullets and leave no way to say which is meant.
"""
current = edges(frontmatter, field)
for position, existing in enumerate(current):
if existing.target == edge.target:
if existing.label == edge.label:
return False
current[position] = edge
frontmatter[field] = render(current)
return True
current.append(edge)
frontmatter[field] = render(current)
return True
def remove(frontmatter: dict[str, Any], field: str, target: str) -> bool:
"""Drop every edge pointing at `target`. True if anything changed."""
current = edges(frontmatter, field)
kept = [edge for edge in current if edge.target != target]
if len(kept) == len(current):
return False
frontmatter[field] = render(kept)
return True
def retarget(frontmatter: dict[str, Any], field: str, old: str, new: str) -> bool:
"""Repoint every edge from `old` to `new`, keeping its label."""
current = edges(frontmatter, field)
changed = False
for position, edge in enumerate(current):
if edge.target == old:
current[position] = Edge(new, edge.label)
changed = True
if changed:
frontmatter[field] = render(current)
return changed
+86 -2
View File
@@ -16,7 +16,7 @@ from __future__ import annotations
from datetime import date from datetime import date
from pathlib import Path from pathlib import Path
from chemenu import config from chemenu import blocks, config, kb_collections, links
from chemenu.frontmatter_io import frontmatter_error from chemenu.frontmatter_io import frontmatter_error
from chemenu.markdown_code import strip_code_spans from chemenu.markdown_code import strip_code_spans
from chemenu.provenance import broken_raw_refs as find_broken_raw_refs from chemenu.provenance import broken_raw_refs as find_broken_raw_refs
@@ -164,7 +164,16 @@ def run_lint(kb_dir: Path) -> dict:
# propagated, a deleted page, or a URL pasted where a title belongs - used # propagated, a deleted page, or a URL pasted where a title belongs - used
# to pass every check. Which fields hold page titles is declared by each # to pass every check. Which fields hold page titles is declared by each
# type-spec's `page_ref_fields:`, not hardcoded here. # type-spec's `page_ref_fields:`, not hardcoded here.
def _collection_of(page):
try:
return page.path.relative_to(config.KB_DIR).parts[0]
except (ValueError, IndexError):
return None
dangling_frontmatter_refs = [] dangling_frontmatter_refs = []
malformed_edges: list[dict] = []
unlabelled_edges: list[dict] = []
unauthorised_labels: list[dict] = []
for title, page in sorted(pages.items()): for title, page in sorted(pages.items()):
type_path = page.frontmatter.get("type") type_path = page.frontmatter.get("type")
if not type_path: if not type_path:
@@ -174,11 +183,53 @@ def run_lint(kb_dir: Path) -> dict:
except ValueError: except ValueError:
continue # unresolvable type is already reported as type_resolution_errors continue # unresolvable type is already reported as type_resolution_errors
for field in ref_fields: for field in ref_fields:
for target in page.frontmatter.get(field) or []: # Through `links` so a labelled edge (`- depends-on: Hermes`) is read
# as its target rather than as a mapping - the entry carries the
# label alongside the title now, and comparing the whole entry would
# report every declared edge as dangling.
for target in links.targets(page.frontmatter, field):
if target not in pages: if target not in pages:
dangling_frontmatter_refs.append( dangling_frontmatter_refs.append(
{"page": title, "field": field, "target": target} {"page": title, "field": field, "target": target}
) )
for entry in links.malformed(page.frontmatter, field):
malformed_edges.append(
{"page": title, "field": field, "entry": str(entry)}
)
# Labels are checked on `related:` only. `sources:`/`entities:`/
# `concepts:` are the provenance path, unlabelled by construction.
if "related" in ref_fields:
source_collection = _collection_of(page)
for edge in links.edges(page.frontmatter, "related"):
if not edge.is_labelled:
unlabelled_edges.append({"page": title, "target": edge.target})
continue
target_page = pages.get(edge.target)
if source_collection is None or target_page is None:
continue
destination = _collection_of(target_page)
if destination is None:
continue
allowed = kb_collections.authorised_labels(source_collection, destination)
if edge.label not in allowed:
unauthorised_labels.append(
{
"page": title,
"target": edge.target,
"label": edge.label,
"destination": destination,
}
)
# A generated region whose markers do not pair up is not a tidiness problem:
# the next write appends a second region beside it instead of replacing it,
# and the page then carries two. An agent rewriting prose at the boundary is
# how a marker goes missing, which is why this is a hard error.
unbalanced_marker_findings = [
{"page": title, "region": name}
for title, page in sorted(pages.items())
for name in blocks.unbalanced_markers(page.body)
]
quote_limit_violations = [] quote_limit_violations = []
for title, page in sorted(pages.items()): for title, page in sorted(pages.items()):
@@ -238,6 +289,10 @@ def run_lint(kb_dir: Path) -> dict:
"undefined_footnote_refs": undefined_footnote_refs, "undefined_footnote_refs": undefined_footnote_refs,
"orphan_footnote_defs": orphan_footnote_defs, "orphan_footnote_defs": orphan_footnote_defs,
"dangling_frontmatter_refs": dangling_frontmatter_refs, "dangling_frontmatter_refs": dangling_frontmatter_refs,
"malformed_edges": malformed_edges,
"unlabelled_edges": unlabelled_edges,
"unauthorised_labels": unauthorised_labels,
"unbalanced_markers": unbalanced_marker_findings,
"quote_limit_violations": quote_limit_violations, "quote_limit_violations": quote_limit_violations,
"invalid_type_paths": invalid_type_paths, "invalid_type_paths": invalid_type_paths,
"type_resolution_errors": type_resolution_errors, "type_resolution_errors": type_resolution_errors,
@@ -322,6 +377,23 @@ def render_markdown(report: dict) -> str:
lines, "Orphan Footnote Definitions", report["orphan_footnote_defs"], lines, "Orphan Footnote Definitions", report["orphan_footnote_defs"],
lambda i: f"[[{i['page']}]] defines `[^{i['id']}]` (-> [[{i['source']}]]) but nothing references it - run `wikitool cite sync`", lambda i: f"[[{i['page']}]] defines `[^{i['id']}]` (-> [[{i['source']}]]) but nothing references it - run `wikitool cite sync`",
) )
_section(
lines, "Malformed Edges", report.get("malformed_edges", []),
lambda i: f"[[{i['page']}]] `{i['field']}`: {i['entry']}",
)
_section(
lines, "Unbalanced Generated-Region Markers", report.get("unbalanced_markers", []),
lambda i: f"[[{i['page']}]]: `{i['region']}`",
)
_section(
lines, "Unlabelled Edges", report.get("unlabelled_edges", []),
lambda i: f"[[{i['page']}]] -> [[{i['target']}]]",
)
_section(
lines, "Labels Not Authorised by the Source Collection",
report.get("unauthorised_labels", []),
lambda i: f"[[{i['page']}]] `{i['label']}` -> kb/{i['destination']}/ ([[{i['target']}]])",
)
_section( _section(
lines, "Dangling Frontmatter References", report["dangling_frontmatter_refs"], lines, "Dangling Frontmatter References", report["dangling_frontmatter_refs"],
lambda i: f"[[{i['page']}]] `{i['field']}:` names `{i['target']}`, which is not a page", lambda i: f"[[{i['page']}]] `{i['field']}:` names `{i['target']}`, which is not a page",
@@ -406,6 +478,16 @@ def default_report_path(report: dict) -> Path:
# through the index or navigation only. `quote_limit_violations` is advisory # through the index or navigation only. `quote_limit_violations` is advisory
# too - it flags a habit, not a broken tree. # too - it flags a habit, not a broken tree.
# #
# `unlabelled_edges` and `unauthorised_labels` are advisory **for now**, and
# that is a dated decision rather than a judgment about severity: they describe
# exactly the state a corpus is in between the 4.0.0 machinery landing and the
# migration reaching each page, which is the window `.wikitool-kb.json` exists
# to represent. They become hard errors once the migration is recorded - the
# same path `legacy_citation_markers` took.
#
# `malformed_edges` and `unbalanced_markers` are hard from the start: neither
# describes an unconverted page, only a broken one.
#
# One definition, used by `lint --fail-on-error` and by the eval scorecard: if # One definition, used by `lint --fail-on-error` and by the eval scorecard: if
# the two disagreed, a run could pass its score while lint refused it. # the two disagreed, a run could pass its score while lint refused it.
HARD_ERROR_KEYS = ( HARD_ERROR_KEYS = (
@@ -421,6 +503,8 @@ HARD_ERROR_KEYS = (
"undefined_footnote_refs", "undefined_footnote_refs",
"orphan_footnote_defs", "orphan_footnote_defs",
"dangling_frontmatter_refs", "dangling_frontmatter_refs",
"malformed_edges",
"unbalanced_markers",
"invalid_type_paths", "invalid_type_paths",
"type_resolution_errors", "type_resolution_errors",
"schema_validation_errors", "schema_validation_errors",
+104 -95
View File
@@ -20,7 +20,7 @@ import unicodedata
from pathlib import Path from pathlib import Path
from typing import Optional from typing import Optional
from chemenu import config, sections from chemenu import blocks, config, conventions
from chemenu.markdown_code import strip_code_spans from chemenu.markdown_code import strip_code_spans
from chemenu.page import Page from chemenu.page import Page
@@ -45,10 +45,6 @@ CITE_DEF_RE = re.compile(
) )
CITE_REF_RE = re.compile(rf"\[\^({_CITE_ID_PATTERN})\]") CITE_REF_RE = re.compile(rf"\[\^({_CITE_ID_PATTERN})\]")
# Where the Footnotes block stops: the next ATX heading of any level. Without
# this the block ran to the end of the file and took any following section with
# it - see split_cite_block().
_NEXT_HEADING_RE = re.compile(r"^#{1,6} ", re.MULTILINE)
# The pre-migration marker: `^[[Source - X]]` or `^[[Source - X|file.md]]`, # The pre-migration marker: `^[[Source - X]]` or `^[[Source - X|file.md]]`,
# read by a Pandoc-style parser as an inline footnote wrapping a broken # read by a Pandoc-style parser as an inline footnote wrapping a broken
@@ -61,27 +57,29 @@ LEGACY_CITE_RE = re.compile(r"\^\[\[([^\]|#]+)(?:\|([^\]]+))?\]\]")
# footnote definitions regardless of the heading text; this heading is purely # footnote definitions regardless of the heading text; this heading is purely
# for human readability when the raw markdown is read directly. # for human readability when the raw markdown is read directly.
# #
# Written under the canonical name, but split_cite_block() matches the aliases # The prefix a source page's title carries, stripped when minting a cite id so
# too - a page whose block still says "## Footnotes" keeps working until it is # the id is not "s-source-x". It is the `source` type-spec's own
# translated. See chemenu/sections.py. # `title_prefix:`, asked for at call time rather than written down here: the
# type-spec belongs to the instance, so hardcoding the string made a documented
# instance decision into a compiler constant - the same leak `sections.py` had.
# #
# Resolved on access rather than bound at import (PEP 562), because the # The literal survives as the fallback for a tree with no resolvable `source`
# canonical name is now this instance's own - `kb/CONVENTIONS.md`, via # type (a fixture, a half-built instance). It is what this stack shipped, so a
# chemenu.conventions - and a module constant would freeze whichever corpus the # corpus that can reach the fallback was minted under it, and ids stay stable.
# process started in. The functions below take it as a default the same way, via _FALLBACK_SOURCE_TITLE_PREFIX = "Source - "
# None rather than an evaluated default argument.
def __getattr__(name: str) -> str:
if name == "CITE_BLOCK_HEADING":
return f"## {sections.FOOTNOTES}"
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
def cite_block_heading_default() -> str: def source_title_prefix() -> str:
"""The Footnotes heading this instance writes, `## ` included.""" """This instance's source-page title prefix, from the type-spec."""
return f"## {sections.FOOTNOTES}" from chemenu.type_resolver import resolver
try:
_SOURCE_TITLE_PREFIX = "Source - " type_path = resolver.find_type_by_name("source")
if type_path:
return resolver.get_title_prefix(type_path)
except (ValueError, OSError):
pass
return _FALLBACK_SOURCE_TITLE_PREFIX
@@ -117,7 +115,8 @@ def cite_id(title: str, qualifier: Optional[str] = None) -> str:
NFKD transliteration is lossy), so callers resolving a real page use NFKD transliteration is lossy), so callers resolving a real page use
unique_cite_id() to add a `-2`/`-3` suffix on collision. unique_cite_id() to add a `-2`/`-3` suffix on collision.
""" """
base_title = title[len(_SOURCE_TITLE_PREFIX):] if title.startswith(_SOURCE_TITLE_PREFIX) else title prefix = source_title_prefix()
base_title = title[len(prefix):] if prefix and title.startswith(prefix) else title
slug = "s-" + _slugify(base_title) slug = "s-" + _slugify(base_title)
if qualifier: if qualifier:
slug += "--" + _slugify(qualifier) slug += "--" + _slugify(qualifier)
@@ -138,62 +137,64 @@ def unique_cite_id(existing_ids: set[str], title: str, qualifier: Optional[str]
return f"{base}-{suffix}" return f"{base}-{suffix}"
def split_cite_block(body: str) -> tuple[str, dict[str, tuple[str, Optional[str]]]]: # Headings a pre-4.0.0 page carries above its citation definitions, for the
"""Split the Footnotes block off `body`. # migration window only. Before the block was delimited it was *located* by this
# text, which is why there are four of them - two languages times two eras. The
# list is read, never written, and `instructions/migrations/` removes the need
# for it once every page carries markers.
_LEGACY_FOOTNOTE_HEADINGS = ("Fußnoten", "Footnotes", "Fussnoten", "Notes")
Returns (body_without_block, definitions), where definitions maps _LEGACY_HEADING_RE = re.compile(
cite_id -> (source_title, qualifier_or_None) in file order. If there is r"^## (?:" + "|".join(re.escape(name) for name in _LEGACY_FOOTNOTE_HEADINGS) + r")[ \t]*$",
no Footnotes block, definitions is {} and body is returned with trailing re.MULTILINE,
blank lines trimmed (so re-rendering after emptying the block is stable). )
_NEXT_HEADING_RE = re.compile(r"^#{1,6} ", re.MULTILINE)
**The block is not "everything to the end of the file".** It used to be,
and every caller here reassembles a page as `head + rendered block` - so a
section that happened to sit after the block was silently deleted on the
next `cite add`, `cite sync` or `rename`. That is not hypothetical: `xref
add` appends its Relationships and See Also sections at the end of the
file, so whether a page kept its cross-references came down to which of the
two commands ran last. Eight pages were carrying content in that position
when this was found.
So the block ends where the next heading begins, and everything after it - def _definitions_in(block: str) -> dict[str, tuple[str, Optional[str]]]:
plus anything inside it that is not a citation definition - is folded back """Every `[^id]: [[Target]]` definition in one region, code masked out.
on to `head`. Nothing is discarded, and because the rendered block is
always emitted last, a page that had drifted into the broken layout is A fenced example of a definition line is an illustration, not a definition.
normalised the first time any of these commands touches it. `strip_code_spans` preserves offsets and line structure, so the masked text
reads line-for-line against the real one.
""" """
# Where the block *starts* is decided on the unmasked body, deliberately. masked = strip_code_spans(block)
# Masking first would mean one unclosed fence anywhere in the prose blanks return {
# the real `## Footnotes` heading too, and the page then reads as having no m.group(1): (m.group(2).strip(), m.group(3).strip() if m.group(3) else None)
# definitions at all - every citation on it undefined, from a single typo. for m in CITE_DEF_RE.finditer(masked)
# A fenced example of the heading itself is the rarer accident and the }
# cheaper one: it costs one page its block, not every citation on it.
match = sections.heading_re(sections.FOOTNOTES).search(body)
def _split_legacy_block(body: str) -> tuple[str, dict[str, tuple[str, Optional[str]]]]:
"""The pre-marker layout: a heading, then definitions, ending at the next
heading.
Kept only so the corpus stays readable between this machinery landing and
the migration reaching each page. Every weakness of the old approach lives
here - it guesses the region's end, and it can be fooled by a fenced example
of the heading - which is the argument the marker pair settles.
"""
match = _LEGACY_HEADING_RE.search(body)
if not match: if not match:
return body.rstrip("\n"), {} return body.rstrip("\n"), {}
head, rest = body[: match.start()], body[match.end():] head, rest = body[: match.start()], body[match.end():]
next_section = _NEXT_HEADING_RE.search(rest) following = _NEXT_HEADING_RE.search(rest)
block, trailing = (rest[: next_section.start()], rest[next_section.start():]) if next_section else (rest, "") block, trailing = (
(rest[: following.start()], rest[following.start():]) if following else (rest, "")
)
# Inside the block, code is masked: a fenced example of a definition line is definitions = _definitions_in(block)
# an illustration, not a definition. strip_code_spans() preserves offsets
# and line structure, so the masked block can be read line-for-line against
# the real one.
masked_block = strip_code_spans(block) masked_block = strip_code_spans(block)
definitions = {
m.group(1): (m.group(2).strip(), m.group(3).strip() if m.group(3) else None)
for m in CITE_DEF_RE.finditer(masked_block)
}
# Lines inside the block that are not definitions are content too - prose # Lines inside the block that are not definitions are content too - prose
# someone left there, a stray bullet. Rescued rather than rejected: this # someone left there, a stray bullet. Rescued rather than rejected: this runs
# runs under `lint` and `corpus_diff` as well, where raising would refuse # under `lint` and `corpus_diff` as well, where raising would refuse to read
# to read a page instead of reporting it. # a page instead of reporting it.
stray = "\n".join( stray = "\n".join(
line line
for line, masked in zip(block.splitlines(), masked_block.splitlines()) for line, masked in zip(block.splitlines(), masked_block.splitlines())
if line.strip() and not CITE_DEF_RE.match(masked) if line.strip() and not CITE_DEF_RE.match(masked)
) )
rescued = "\n\n".join(part.strip("\n") for part in (stray, trailing) if part.strip()) rescued = "\n\n".join(part.strip("\n") for part in (stray, trailing) if part.strip())
head = head.rstrip("\n") head = head.rstrip("\n")
if rescued: if rescued:
@@ -201,49 +202,57 @@ def split_cite_block(body: str) -> tuple[str, dict[str, tuple[str, Optional[str]
return head, definitions return head, definitions
def cite_block_heading(body: str) -> str: def split_cite_block(body: str) -> tuple[str, dict[str, tuple[str, Optional[str]]]]:
"""The Footnotes heading `body` actually carries, canonical if it has none. """Split the citation region off `body`.
Rewriting a page must not silently retitle its block: a page still using an Returns (body_without_region, definitions), where definitions maps
alias is untranslated, not broken, and `cite sync` has to stay a no-op on cite_id -> (source_title, qualifier_or_None) in file order.
it. Translating the heading is the migration's job, not the tool's."""
match = sections.heading_re(sections.FOOTNOTES).search(body) **The region is delimited, not guessed.** It used to end "at the next
return match.group(0).strip() if match else cite_block_heading_default() heading", and before that "at the end of the file" - and every caller here
reassembles a page as `head + rendered region`, so a section that happened to
sit after it was silently deleted on the next `cite add`, `cite sync` or
`rename`. Eight pages were carrying content in that position when it was
found. A marker pair answers where the region stops exactly, which is the
whole reason for it.
A page with no markers is read through the legacy path instead, so the
corpus stays readable until the migration reaches it.
"""
region = blocks.find(body, blocks.FOOTNOTES)
if region is None:
return _split_legacy_block(body)
return blocks.strip(body, blocks.FOOTNOTES).rstrip("\n"), _definitions_in(region)
def render_cite_block( def render_cite_block(definitions: dict[str, tuple[str, Optional[str]]]) -> str:
definitions: dict[str, tuple[str, Optional[str]]], heading: Optional[str] = None """The citation region for `definitions`, markers included, in dict order.
) -> str:
"""Render the Footnotes block for `definitions` (cite_id -> (title,
qualifier)), preserving dict order. Empty dict renders "" - a page with
no citations carries no block at all.
`heading=None` means this instance's canonical Footnotes heading, resolved An empty dict renders "" - a page with no citations carries no region at
at call time. It cannot be an evaluated default: the name comes from all, rather than a heading with nothing under it.
`kb/CONVENTIONS.md`, so a default bound at import would answer for whichever """
corpus the process started in.""" lines = []
if not definitions:
return ""
lines = [heading or cite_block_heading_default(), ""]
for cid, (title, qualifier) in definitions.items(): for cid, (title, qualifier) in definitions.items():
target = f"{title}|{qualifier}" if qualifier else title target = f"{title}|{qualifier}" if qualifier else title
lines.append(f"[^{cid}]: [[{target}]]") lines.append(f"[^{cid}]: [[{target}]]")
return "\n".join(lines) + "\n" return blocks.render(
blocks.FOOTNOTES, conventions.heading(blocks.FOOTNOTES), lines
)
def render_page_body( def render_page_body(
head: str, head: str, definitions: dict[str, tuple[str, Optional[str]]]
definitions: dict[str, tuple[str, Optional[str]]],
heading: Optional[str] = None,
) -> str: ) -> str:
"""Reassemble a page body from its non-Footnotes content and citation """Reassemble a page body from its non-citation content and its definitions -
definitions - the inverse of split_cite_block(). Pass the original body's the inverse of `split_cite_block`.
`cite_block_heading()` to preserve an alias the page still uses."""
head = head.rstrip("\n") The heading is no longer threaded through from the caller. It used to be, so
block = render_cite_block(definitions, heading) that rewriting a page would not silently retitle a block whose text the tool
if not block: was *matching on*; now the marker carries the identity and the heading is a
return head + "\n" rendering value, so re-rendering it under this instance's own words is a
return head + "\n\n" + block repair rather than a rename.
"""
return blocks.replace(head.rstrip("\n") + "\n", blocks.FOOTNOTES, render_cite_block(definitions))
def extract_inline_cites(body: str) -> set[tuple[str, Optional[str]]]: def extract_inline_cites(body: str) -> set[tuple[str, Optional[str]]]:
-84
View File
@@ -1,84 +0,0 @@
"""The section headings wikitool reads and writes inside a page body.
These headings are structural, not prose: `xref add` locates Relationships and
See Also by name, and `cite add` owns the trailing Footnotes block. An author
may add any other heading they like - only the ones named here are matched by
the tool, and only these have to stay predictable.
**Which words they are is the instance's decision, not the stack's.** They
follow the KB language, and the KB language is declared in `kb/CONVENTIONS.md`
(see `chemenu.conventions`). This module used to hold `RELATIONSHIPS =
"Beziehungen"` as a Python constant, which made an instance writing its pages
in any other language edit the compiler to say so - the one place a documented
instance convention had leaked into code.
Each heading has one **canonical** name - what the tool writes - and any number
of **aliases** it still recognizes. That asymmetry is what lets a corpus migrate
page by page instead of all at once: a page still carrying `## Relationships` is
found and appended to correctly, and only takes the canonical name when the page
itself is translated. Removing an alias is therefore a breaking change for every
page not yet converted, not a cleanup.
The three module attributes below resolve on access (PEP 562), the same way
`config`'s paths do and for the same reason: a caller that repoints `KB_DIR`
must not be answered out of a value bound at import time by whichever tree the
process started in.
"""
from __future__ import annotations
import re
from chemenu import conventions
# The slots, re-exported so a caller keeps using `sections.RELATIONSHIPS` as an
# opaque handle. The value it resolves to is the heading text; the name it is
# looked up under is stable.
_SLOT_ATTRS = {
"RELATIONSHIPS": conventions.RELATIONSHIPS,
"SEE_ALSO": conventions.SEE_ALSO,
"FOOTNOTES": conventions.FOOTNOTES,
}
def __getattr__(name: str) -> str:
slot = _SLOT_ATTRS.get(name)
if slot is None:
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
return conventions.canonical(slot)
def __dir__() -> list[str]:
return sorted([*globals(), *_SLOT_ATTRS])
def _slot_of(canonical: str) -> str:
"""The slot whose current canonical name is `canonical`.
Callers hold on to the resolved heading text (`sections.FOOTNOTES`), not to
the slot, so the lookup has to go back the other way. Falls back to matching
against every name a slot is recognized under, so a caller that resolved the
attribute before the conventions file changed still lands on the right slot.
"""
for slot in conventions.SLOTS:
if canonical == conventions.canonical(slot):
return slot
for slot in conventions.SLOTS:
if canonical in conventions.names(slot):
return slot
raise ValueError(f"{canonical!r} is not a tool-owned section heading")
def names(canonical: str) -> tuple[str, ...]:
"""Every name `canonical` is recognized under, canonical first."""
return conventions.names(_slot_of(canonical))
def heading_re(canonical: str) -> re.Pattern[str]:
"""Match a `## <heading>` line for `canonical` or any of its aliases."""
alternation = "|".join(re.escape(name) for name in names(canonical))
return re.compile(rf"^## (?:{alternation})[ \t]*$", re.MULTILINE)
def is_known(heading: str) -> bool:
"""True if `heading` is a canonical name or an alias of one."""
return any(heading in conventions.names(slot) for slot in conventions.SLOTS)
+13 -1
View File
@@ -171,9 +171,21 @@ def kb_dir(tmp_path: Path) -> Path:
"entities/technologies", "entities/people", "entities/technologies", "entities/people",
"concepts", "sources", "comparisons"): "concepts", "sources", "comparisons"):
(kb / sub).mkdir(parents=True) (kb / sub).mkdir(parents=True)
# The contracts carry a real declaration, because three things now read one:
# `docs verify` checks `profile:`/`required_by_stack:`, and `xref add` asks
# `outbound:` whether a label is authorised from this collection. A fixture
# contract without it would make every `xref add` in the suite fail for a
# reason that has nothing to do with what the test is about.
for collection in ("entities", "concepts", "sources", "comparisons"): for collection in ("entities", "concepts", "sources", "comparisons"):
(kb / collection / "COLLECTION.md").write_text( (kb / collection / "COLLECTION.md").write_text(
f"# kb/{collection}/ - Collection Contract\n", encoding="utf-8" "---\n"
f"profile: {collection}\n"
f"required_by_stack: {'true' if collection == 'sources' else 'false'}\n"
"outbound:\n"
" any: [depends-on, required-by, runs-on, hosts, uses, implements, see-also]\n"
"---\n\n"
f"# kb/{collection}/ - Collection Contract\n",
encoding="utf-8",
) )
write_page( write_page(
+90
View File
@@ -0,0 +1,90 @@
"""Tests for generated regions - the delimiters that replaced heading matching.
The whole point is that a region's *identity* stops depending on its heading
text. Everything here is about the two questions the old approach answered by
guessing: where does the region start, and where does it stop.
"""
from __future__ import annotations
from chemenu import blocks
PROSE = "# Page\n\n## Beschreibung\n\nProse.\n"
def _links(heading="Beziehungen", lines=("- **uses:** [[X]]",)):
return blocks.render(blocks.LINKS, heading, list(lines))
def test_a_region_round_trips():
body = blocks.replace(PROSE, blocks.LINKS, _links())
assert blocks.find(body, blocks.LINKS) == "## Beziehungen\n\n- **uses:** [[X]]"
assert blocks.strip(body, blocks.LINKS) == PROSE
def test_replacing_does_not_append_a_second_region():
body = blocks.replace(PROSE, blocks.LINKS, _links())
again = blocks.replace(body, blocks.LINKS, _links(lines=["- **uses:** [[Y]]"]))
assert again.count(blocks.open_marker(blocks.LINKS)) == 1
assert "[[X]]" not in again and "[[Y]]" in again
def test_the_heading_inside_a_region_is_not_how_it_is_found():
"""A page whose region carries a heading the instance never declared - an
unconverted page, a hand-edit, another language - is still located exactly.
Under heading matching this was the case that silently created a second
section."""
body = blocks.replace(PROSE, blocks.LINKS, _links(heading="Something Else Entirely"))
assert "[[X]]" in blocks.find(body, blocks.LINKS)
assert blocks.strip(body, blocks.LINKS) == PROSE
def test_content_after_a_region_survives_a_rewrite():
"""The eight-page bug, as a test. The old block ran to the next heading -
and before that to the end of the file - so anything sitting after it was
deleted on the next write."""
body = blocks.replace(PROSE, blocks.LINKS, _links()) + "\n## Afterwards\n\nKeep me.\n"
rewritten = blocks.replace(body, blocks.LINKS, _links(lines=["- **uses:** [[Z]]"]))
assert "Keep me." in rewritten
assert rewritten.count("## Afterwards") == 1
def test_two_regions_coexist_without_reading_each_other():
body = blocks.replace(PROSE, blocks.LINKS, _links())
body = blocks.replace(
body, blocks.FOOTNOTES,
blocks.render(blocks.FOOTNOTES, "Fußnoten", ["[^s-x]: [[Source - X]]"]),
)
assert "[[X]]" in blocks.find(body, blocks.LINKS)
assert "[^s-x]" in blocks.find(body, blocks.FOOTNOTES)
assert blocks.marker_pairs(body) == {"links": 1, "footnotes": 1}
def test_an_empty_region_is_no_region_at_all():
"""A page that cites nothing must not carry an empty Footnotes heading."""
assert blocks.render(blocks.LINKS, "Beziehungen", []) == ""
body = blocks.replace(PROSE, blocks.LINKS, _links())
assert blocks.replace(body, blocks.LINKS, "") == PROSE
def test_an_absent_region_reads_as_none_not_as_empty():
"""None and "" have to stay distinguishable: one means the page has no
region, the other that it has one holding nothing."""
assert blocks.find(PROSE, blocks.LINKS) is None
def test_a_dropped_marker_is_detectable():
"""An agent rewriting prose at the boundary can lose one. Silent otherwise:
the region becomes ordinary prose and the next write appends a second one
beside it."""
body = blocks.replace(PROSE, blocks.LINKS, _links())
assert blocks.unbalanced_markers(body) == []
assert blocks.unbalanced_markers(body.replace(blocks.close_marker(blocks.LINKS), "")) == ["links"]
assert blocks.unbalanced_markers(body.replace(blocks.open_marker(blocks.LINKS), "")) == ["links"]
def test_marker_pairs_counts_rather_than_sets():
"""A page that went from one region to two has the same set of names and a
different count - which is why `migrate verify` compares counts."""
body = blocks.replace(PROSE, blocks.LINKS, _links())
doubled = body + "\n" + _links() + "\n"
assert blocks.marker_pairs(doubled) == {"links": 2}
+31 -8
View File
@@ -133,27 +133,50 @@ def test_sync_page_is_idempotent_once_clean():
from chemenu.page import Page from chemenu.page import Page
body = "\n# X\n\n## Definition\n\nCites [^s-a].\n\n## Fußnoten\n\n[^s-a]: [[Source - A]]\n" from chemenu import blocks
body = blocks.replace(
"\n# X\n\n## Definition\n\nCites [^s-a].\n",
blocks.FOOTNOTES,
blocks.render(blocks.FOOTNOTES, "Fußnoten", ["[^s-a]: [[Source - A]]"]),
)
page = Page(path=Path("/tmp/X.md"), frontmatter={}, body=body) page = Page(path=Path("/tmp/X.md"), frontmatter={}, body=body)
new_body, changed, pruned, undefined = sync_page(page) new_body, changed, pruned, undefined = sync_page(page)
assert changed is False assert changed is False
assert new_body == body
assert pruned == [] assert pruned == []
assert undefined == [] assert undefined == []
def test_sync_page_leaves_an_untranslated_footnotes_heading_alone(): def test_sync_page_upgrades_a_pre_marker_block_to_a_delimited_region():
"""`cite sync` must not retitle a block just because the page has not been """The mechanical half of the marker migration, done by the command that
translated yet - that would make it rewrite the whole corpus on one run.""" already owns the block.
This inverts an older rule. While the tool *located* the block by matching
its heading, re-rendering one under a different name was a rewrite of the
whole corpus on a single run, so `cite sync` had to leave an untranslated
heading alone. Now the marker carries the identity: converting the region is
a repair, and the heading follows `kb/CONVENTIONS.md` from then on."""
from pathlib import Path from pathlib import Path
from chemenu import blocks
from chemenu.page import Page from chemenu.page import Page
body = "\n# X\n\n## Definition\n\nCites [^s-a].\n\n## Footnotes\n\n[^s-a]: [[Source - A]]\n" body = "\n# X\n\n## Definition\n\nCites [^s-a].\n\n## Footnotes\n\n[^s-a]: [[Source - A]]\n"
page = Page(path=Path("/tmp/X.md"), frontmatter={}, body=body) page = Page(path=Path("/tmp/X.md"), frontmatter={}, body=body)
new_body, changed, pruned, undefined = sync_page(page) new_body, changed, _pruned, undefined = sync_page(page)
assert changed is False
assert "## Footnotes" in new_body assert changed is True
assert "## Fußnoten" not in new_body assert undefined == []
assert blocks.unbalanced_markers(new_body) == []
assert "[^s-a]: [[Source - A]]" in blocks.find(new_body, blocks.FOOTNOTES)
# The prose above it is untouched, and the legacy heading is not left behind
# as a second, now-empty section.
assert "## Definition" in new_body
# The legacy heading is not left behind as a second, now-empty section: the
# region carries its own heading, rendered from `kb/CONVENTIONS.md`.
assert "## Footnotes" not in new_body
assert new_body.count("<!-- wikitool:footnotes -->") == 1
def test_cite_sync_command_over_kb(kb_dir, raw_dir, monkeypatch): def test_cite_sync_command_over_kb(kb_dir, raw_dir, monkeypatch):
+84 -57
View File
@@ -1,9 +1,9 @@
"""Tests for `kb/CONVENTIONS.md` - the instance-owned half of the authoring rules. """Tests for `kb/CONVENTIONS.md` - the instance-owned half of the authoring rules.
Two things are under test here, and they are the two the split exists for: the Two things are under test here, and they are the two the split exists for: the
compiler reads its section headings from the corpus rather than from Python, and headings the compiler *renders* come from the corpus rather than from Python,
a collection declares who owns its rules rather than having it inferred from the and a collection declares who owns its rules rather than having it inferred from
directory name. the directory name.
""" """
from __future__ import annotations from __future__ import annotations
@@ -11,15 +11,15 @@ from pathlib import Path
import pytest import pytest
from chemenu import config, conventions, kb_collections, sections from chemenu import blocks, config, conventions, kb_collections
from chemenu.tests.conftest import use_shipped_type_specs
GERMAN = ( GERMAN = (
"---\n" "---\n"
"language: de\n" "language: de\n"
"profile: german\n" "profile: german\n"
"sections:\n" "sections:\n"
" relationships: Beziehungen\n" " links: Beziehungen\n"
" see_also: Siehe auch\n"
" footnotes: Fußnoten\n" " footnotes: Fußnoten\n"
"---\n\n# conventions\n" "---\n\n# conventions\n"
) )
@@ -29,11 +29,8 @@ FRENCH = (
"language: fr\n" "language: fr\n"
"profile: none\n" "profile: none\n"
"sections:\n" "sections:\n"
" relationships: Relations\n" " links: Relations\n"
" see_also: Voir aussi\n"
" footnotes: Notes\n" " footnotes: Notes\n"
"section_aliases:\n"
" relationships: [Beziehungen]\n"
"---\n\n# conventions\n" "---\n\n# conventions\n"
) )
@@ -44,6 +41,11 @@ def kb_root(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path:
kb.mkdir() kb.mkdir()
monkeypatch.setattr(config, "ROOT", tmp_path) monkeypatch.setattr(config, "ROOT", tmp_path)
monkeypatch.setattr(config, "KB_DIR", kb) monkeypatch.setattr(config, "KB_DIR", kb)
# Which collection the stack requires is *derived* from where the required
# `source` type writes, so these tests need the shipped `types/` reachable -
# a fixture tree without one derives an empty requirement and would assert
# against a rule that is not running. See conftest.use_shipped_type_specs.
use_shipped_type_specs(monkeypatch)
conventions.reset_cache() conventions.reset_cache()
yield kb yield kb
conventions.reset_cache() conventions.reset_cache()
@@ -64,65 +66,53 @@ def _collection(kb: Path, name: str, profile: str = "none", required: bool = Fal
return directory return directory
def test_missing_file_falls_back_to_what_the_stack_used_to_hardcode(kb_root): def test_a_missing_file_renders_under_a_cosmetic_default(kb_root):
"""The state between installing this machinery and running the migration """The window between installing the machinery and writing the conventions
that writes the file. Every command has to keep working through it, and the file. It has to render *something*, and a wrong heading is now merely wrong
only corpus that can be in it was written under these names.""" words: the marker pair carries the region's identity, so the next write
assert conventions.canonical(conventions.FOOTNOTES) == "Fußnoten" repairs it once the instance declares one. Before markers, the same mistake
assert sections.FOOTNOTES == "Fußnoten" split a page into two sections."""
assert conventions.heading(blocks.FOOTNOTES) == "Footnotes"
assert conventions.heading(blocks.LINKS) == "Relationships"
def test_the_compiler_writes_the_headings_the_instance_declared(kb_root): def test_the_compiler_renders_the_headings_the_instance_declared(kb_root):
_write(kb_root, FRENCH) _write(kb_root, FRENCH)
assert sections.RELATIONSHIPS == "Relations" assert conventions.heading(blocks.LINKS) == "Relations"
assert sections.SEE_ALSO == "Voir aussi" assert conventions.heading(blocks.FOOTNOTES) == "Notes"
assert sections.FOOTNOTES == "Notes"
def test_declared_aliases_and_the_pre_conventions_names_are_both_recognized(kb_root):
"""The translation path. A page still carrying the old heading has to be
found and appended to, or a language change would silently split every page
into two Relationships sections."""
_write(kb_root, FRENCH)
pattern = sections.heading_re(sections.RELATIONSHIPS)
for heading in ("## Relations", "## Beziehungen", "## Relationships"):
assert pattern.search(f"# Page\n\n{heading}\n\n- x\n"), heading
def test_the_canonical_name_is_not_duplicated_among_its_aliases(kb_root):
"""An instance declaring the pre-conventions name gets it once, not twice -
otherwise `heading_re`'s alternation carries a redundant branch and
`names()` misreports what a page could be carrying."""
_write(
kb_root,
"---\nsections:\n relationships: Relationships\n"
" see_also: See Also\n footnotes: Footnotes\n---\n",
)
names = conventions.names(conventions.RELATIONSHIPS)
assert names[0] == "Relationships"
assert len(names) == len(set(names))
def test_section_variables_are_what_a_type_spec_template_substitutes(kb_root):
_write(kb_root, GERMAN)
assert conventions.section_variables() == {
"section.relationships": "Beziehungen",
"section.see_also": "Siehe auch",
"section.footnotes": "Fußnoten",
}
def test_a_rewritten_file_is_not_answered_out_of_the_cache(kb_root): def test_a_rewritten_file_is_not_answered_out_of_the_cache(kb_root):
_write(kb_root, GERMAN) _write(kb_root, GERMAN)
assert sections.FOOTNOTES == "Fußnoten" assert conventions.heading(blocks.FOOTNOTES) == "Fußnoten"
_write(kb_root, FRENCH) _write(kb_root, FRENCH)
assert sections.FOOTNOTES == "Notes" assert conventions.heading(blocks.FOOTNOTES) == "Notes"
def test_a_region_is_found_by_its_marker_not_by_its_heading(kb_root):
"""The point of the whole change. A page whose heading says something the
instance never declared - an untranslated page, a hand-edit, another
language entirely - is still located exactly."""
_write(kb_root, FRENCH)
body = blocks.replace(
"# Page\n\nProse.\n",
blocks.LINKS,
blocks.render(blocks.LINKS, "Ganz andere Wörter", ["- **uses:** [[X]]"]),
)
assert "- **uses:** [[X]]" in blocks.find(body, blocks.LINKS)
def test_an_unknown_section_key_is_reported(kb_root):
_write(
kb_root,
"---\nsections:\n links: L\n footnotes: F\n see_also: S\n---\n",
)
assert any("see_also" in issue for issue in conventions.declaration_issues())
def test_an_incomplete_sections_block_is_reported(kb_root): def test_an_incomplete_sections_block_is_reported(kb_root):
_write(kb_root, "---\nlanguage: de\nsections:\n relationships: Beziehungen\n---\n") _write(kb_root, "---\nlanguage: de\nsections:\n links: Beziehungen\n---\n")
issues = conventions.declaration_issues() issues = conventions.declaration_issues()
assert any("sections.see_also" in issue for issue in issues)
assert any("sections.footnotes" in issue for issue in issues) assert any("sections.footnotes" in issue for issue in issues)
@@ -165,3 +155,40 @@ def test_a_correct_declaration_reports_nothing(kb_root):
_collection(kb_root, "sources", profile="sources", required=True) _collection(kb_root, "sources", profile="sources", required=True)
_collection(kb_root, "entities", profile="entities") _collection(kb_root, "entities", profile="entities")
assert kb_collections.declaration_issues(kb_root) == [] assert kb_collections.declaration_issues(kb_root) == []
# --- outbound authorisation ------------------------------------------------
def _authorising(kb: Path, name: str, outbound: str, required: bool = False) -> Path:
directory = kb / name
directory.mkdir(parents=True, exist_ok=True)
(directory / kb_collections.CONTRACT_NAME).write_text(
f"---\nprofile: {name}\nrequired_by_stack: {str(required).lower()}\n"
f"outbound:\n{outbound}\n---\n\n# {name}\n",
encoding="utf-8",
)
return directory
def test_the_source_collection_decides_which_labels_may_be_used(kb_root):
"""Commonplace ADR-019, adopted: the rules that govern an edge are the rules
of the collection the *asserting* page lives in. That is also why the reverse
edge cannot be written automatically - it would be governed by a contract the
author never read."""
_authorising(kb_root, "entities", " concepts: [implements]\n entities: [uses]")
assert kb_collections.authorised_labels("entities", "concepts") == {"implements"}
assert kb_collections.authorised_labels("entities", "entities") == {"uses"}
def test_any_widens_every_destination(kb_root):
_authorising(kb_root, "entities", " any: [see-also]\n concepts: [implements]")
assert kb_collections.authorised_labels("entities", "concepts") == {"implements", "see-also"}
assert kb_collections.authorised_labels("entities", "sources") == {"see-also"}
def test_an_undeclared_destination_authorises_nothing(kb_root):
"""An empty result is a real answer - "do not link there from here" - not a
missing declaration to be filled in with a permissive default."""
_authorising(kb_root, "entities", " concepts: [implements]")
assert kb_collections.authorised_labels("entities", "sources") == set()
+37
View File
@@ -76,6 +76,19 @@ def repo(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path:
types_dir = root / "types" types_dir = root / "types"
types_dir.mkdir() types_dir.mkdir()
(types_dir / "entity.schema.yaml").write_text("type: object\n", encoding="utf-8") (types_dir / "entity.schema.yaml").write_text("type: object\n", encoding="utf-8")
# Two real type-specs, one on each side of the `root:` line, so the export's
# split has something to split. `entity` writes into kb/ and is therefore
# the instance's; `instruction` writes into the repo and is the stack's.
(types_dir / "entity.md").write_text(
"---\ntype: types/type-spec.md\nname: entity\ndescription: d\n"
"schema: types/entity.schema.yaml\nbase_dir: entities\n---\n\n# Entity\n",
encoding="utf-8",
)
(types_dir / "instruction.md").write_text(
"---\ntype: types/type-spec.md\nname: instruction\ndescription: d\n"
"schema: null\nbase_dir: instructions\nroot: repo\n---\n\n# Instruction\n",
encoding="utf-8",
)
tools_dir = root / "tools" tools_dir = root / "tools"
(tools_dir / "chemenu").mkdir(parents=True) (tools_dir / "chemenu").mkdir(parents=True)
@@ -274,6 +287,10 @@ def test_find_leaks_is_silent_on_a_clean_plan(repo):
# filled name must never cross - only the `.template` beside it does. # filled name must never cross - only the `.template` beside it does.
"kb/CONVENTIONS.md", "kb/CONVENTIONS.md",
"kb/entities/COLLECTION.md", "kb/entities/COLLECTION.md",
# A page type-spec under its filled name: the instance's, not the
# stack's, so shipping it would hand a new instance this one's
# authoring language as though the stack had decided it.
"types/entity.md",
], ],
) )
def test_find_leaks_catches_one_instance_own_data(repo, relative): def test_find_leaks_catches_one_instance_own_data(repo, relative):
@@ -318,6 +335,26 @@ def test_plan_ships_the_conventions_template_and_not_the_filled_file(repo):
assert "kb/CONVENTIONS.md" not in plan assert "kb/CONVENTIONS.md" not in plan
def test_page_type_specs_ship_as_templates_and_stack_types_do_not(repo, monkeypatch):
"""The `root:` line, applied. A type-spec whose instances are pages under
`kb/` describes what this instance writes, so its prose, template and
language are the instance's; one whose instances are stack artifacts ships
verbatim. The `.schema.yaml` travels with its spec - the two are one type,
and adopting half would leave a spec validated by a file it does not own."""
from chemenu.type_resolver import resolver
monkeypatch.setattr(resolver, "_repo_root", config.ROOT)
plan = dist_cmd.build_plan()
assert "types/entity.md.template" in plan
assert "types/entity.schema.yaml.template" in plan
assert "types/entity.md" not in plan
assert "types/entity.schema.yaml" not in plan
assert "types/instruction.md" in plan
assert "types/instruction.md.template" not in plan
def test_plan_creates_empty_raw_subdirs_not_real_content(repo): def test_plan_creates_empty_raw_subdirs_not_real_content(repo):
plan = dist_cmd.build_plan() plan = dist_cmd.build_plan()
for sub in ("articles", "documents", "notes", "assets"): for sub in ("articles", "documents", "notes", "assets"):
+7 -3
View File
@@ -25,14 +25,18 @@ def instance(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path:
kb = root / "kb" kb = root / "kb"
for sub in ("entities", "concepts", "sources", "comparisons"): for sub in ("entities", "concepts", "sources", "comparisons"):
(kb / sub).mkdir(parents=True) (kb / sub).mkdir(parents=True)
(kb / sub / "COLLECTION.md").write_text(f"# {sub}\n", encoding="utf-8") (kb / sub / "COLLECTION.md").write_text(
f"---\nprofile: {sub}\nrequired_by_stack: "
f"{'true' if sub == 'sources' else 'false'}\n---\n\n# {sub}\n",
encoding="utf-8",
)
(kb / "index.md").write_text("# Index\n", encoding="utf-8") (kb / "index.md").write_text("# Index\n", encoding="utf-8")
(kb / "log.md").write_text("# Log\n", encoding="utf-8") (kb / "log.md").write_text("# Log\n", encoding="utf-8")
(kb / "provenance.md").write_text("# Provenance\n", encoding="utf-8") (kb / "provenance.md").write_text("# Provenance\n", encoding="utf-8")
(kb / "CONTRACT.md").write_text("# kb contract\n", encoding="utf-8") (kb / "CONTRACT.md").write_text("# kb contract\n", encoding="utf-8")
(kb / "CONVENTIONS.md").write_text( (kb / "CONVENTIONS.md").write_text(
"---\nlanguage: en\nprofile: none\nsections:\n" "---\nlanguage: en\nprofile: none\nsections:\n"
" relationships: Relationships\n see_also: See Also\n footnotes: Footnotes\n" " links: Relationships\n footnotes: Footnotes\n"
"---\n\n# conventions\n", "---\n\n# conventions\n",
encoding="utf-8", encoding="utf-8",
) )
@@ -220,7 +224,7 @@ def test_conventions_with_an_incomplete_sections_block_fail(instance):
"""Present and deciding nothing - the same failure mode the personalization """Present and deciding nothing - the same failure mode the personalization
sentinel check exists for, one directory down.""" sentinel check exists for, one directory down."""
(config.KB_DIR / conventions.CONVENTIONS_FILENAME).write_text( (config.KB_DIR / conventions.CONVENTIONS_FILENAME).write_text(
"---\nlanguage: en\nsections:\n relationships: Relationships\n---\n", encoding="utf-8" "---\nlanguage: en\nsections:\n links: Relationships\n---\n", encoding="utf-8"
) )
conventions.reset_cache() conventions.reset_cache()
assert _status(doctor.run_doctor(), "conventions") == "FAIL" assert _status(doctor.run_doctor(), "conventions") == "FAIL"
+144 -1
View File
@@ -16,7 +16,13 @@ from chemenu.version import Version
CHANGES = "# Changelog\n\n---\n\n## 1.0.0 - 2026-08-30 - First\n\nBody.\n" CHANGES = "# Changelog\n\n---\n\n## 1.0.0 - 2026-08-30 - First\n\nBody.\n"
def write_migration(directory: Path, target: str, slug: str, kind: str = "assisted") -> Path: def write_migration(
directory: Path,
target: str,
slug: str,
kind: str = "assisted",
obligation: str = "required",
) -> Path:
directory.mkdir(parents=True, exist_ok=True) directory.mkdir(parents=True, exist_ok=True)
path = directory / f"{target}-{slug}.md" path = directory / f"{target}-{slug}.md"
path.write_text( path.write_text(
@@ -27,6 +33,7 @@ def write_migration(directory: Path, target: str, slug: str, kind: str = "assist
"manual: true\n" "manual: true\n"
f"migrates_to: {target}\n" f"migrates_to: {target}\n"
f"migration_kind: {kind}\n" f"migration_kind: {kind}\n"
f"obligation: {obligation}\n"
"---\n\n# Migration\n\nSteps.\n", "---\n\n# Migration\n\nSteps.\n",
encoding="utf-8", encoding="utf-8",
) )
@@ -259,3 +266,139 @@ def test_verify_reports_an_unknown_revision(git_instance):
json_out=False, json_out=False,
fail_on_error=False, fail_on_error=False,
) )
# --- obligation: required vs offered ---------------------------------------
def test_an_offered_migration_is_not_in_the_outstanding_chain(instance):
"""An offer is the stack proposing a better default for a file the instance
owns. Declining it leaves the content in a shape the machinery accepts, so
counting it as owed would make `kb_version` unreachable for an instance that
simply kept its own file."""
write_migration(
instance / "instructions" / "migrations", "1.9.0", "nicer-template",
kind="mechanical", obligation="offered",
)
migrations = kb_state.load_migrations()
pending = kb_state.chain(migrations, Version.parse("1.3.0"), Version.parse("2.0.0"))
assert "1.9.0-nicer-template" not in [m.name for m in pending]
assert [str(m.target) for m in pending] == ["1.4.0", "1.7.0", "2.0.0"]
def test_an_offered_migration_is_listed_separately(instance):
write_migration(
instance / "instructions" / "migrations", "1.9.0", "nicer-template",
kind="mechanical", obligation="offered",
)
offered = kb_state.offers(kb_state.load_migrations(), applied=set())
assert [m.name for m in offered] == ["1.9.0-nicer-template"]
def test_an_offer_stays_on_the_table_regardless_of_the_version(instance):
"""Offers are bounded by the applied ledger, not by kb_version - taking one
deliberately does not move the version, so the version can say nothing about
whether it was taken. Nor are they bounded above by the stack: an offer is
about a file the instance owns, not about the content shape."""
directory = instance / "instructions" / "migrations"
write_migration(directory, "1.1.0", "old-default", obligation="offered")
write_migration(directory, "2.1.0", "later-default", obligation="offered")
offered = kb_state.offers(kb_state.load_migrations(), applied=set())
assert [m.name for m in offered] == ["1.1.0-old-default", "2.1.0-later-default"]
def test_taking_an_offer_records_it_without_moving_the_version(instance):
write_migration(
instance / "instructions" / "migrations", "1.9.0", "nicer-template",
obligation="offered",
)
set_kb_version(instance, "1.3.1")
migrate_cmd.done_command(version="1.9.0", pages=None, dry_run=False)
state = kb_state.read_kb_state()
assert state["kb_version"] == "1.3.1"
assert state["applied"][-1]["migration"] == "1.9.0-nicer-template"
assert kb_state.offers(kb_state.load_migrations(), kb_state.applied_names(state)) == []
def test_an_offer_out_of_order_is_not_refused(instance):
"""The chain's ordering rule exists because skipping a link leaves the
corpus in an undescribed shape. An offer is not a link, so there is nothing
to skip - and refusing it would make the required chain a prerequisite for
an unrelated file upgrade."""
write_migration(
instance / "instructions" / "migrations", "1.9.0", "nicer-template",
obligation="offered",
)
set_kb_version(instance, "1.3.1") # 1.4.0 is the next *required* link
migrate_cmd.done_command(version="1.9.0", pages=None, dry_run=False)
assert kb_state.read_kb_version() == Version(1, 3, 1)
def test_obligation_defaults_to_required_when_undeclared(instance):
"""Every migration written before this axis existed is mandatory, and an
unreadable value must not silently downgrade one."""
directory = instance / "instructions" / "migrations"
path = write_migration(directory, "1.5.0", "legacy")
path.write_text(
path.read_text(encoding="utf-8").replace("obligation: required\n", ""), encoding="utf-8"
)
bogus = write_migration(directory, "1.6.0", "bogus", obligation="whatever")
assert bogus.is_file()
by_name = {m.name: m for m in kb_state.load_migrations()}
assert by_name["1.5.0-legacy"].obligation == kb_state.REQUIRED
assert by_name["1.6.0-bogus"].obligation == kb_state.REQUIRED
def test_status_never_blocks_on_an_offer(instance, capsys):
write_migration(
instance / "instructions" / "migrations", "1.9.0", "nicer-template",
obligation="offered",
)
set_kb_version(instance, "2.0.0")
migrate_cmd.status_command(json_out=True)
result = json.loads(capsys.readouterr().out)
assert result["pending"] == []
assert [m["name"] for m in result["offered"]] == ["1.9.0-nicer-template"]
# The chain is empty and the offer is listed: `status` reports both without
# the offer ever counting as owed.
# --- divergence against the release stamp ----------------------------------
def _write_stamp(root: Path, files: dict[str, str]) -> None:
import hashlib
digests = {}
for relative, content in files.items():
path = root / relative
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(content, encoding="utf-8")
digests[relative] = "sha256:" + hashlib.sha256(content.encode("utf-8")).hexdigest()
(root / ".wikitool-release.json").write_text(
json.dumps({"schema": 1, "version": "2.0.0", "files": digests}), encoding="utf-8"
)
def test_divergent_files_tells_an_edited_file_from_a_received_one(instance):
"""The half of the release stamp that has existed since it was written and
that nothing read until offers needed it: may this file be overwritten, or
does a person have to reconcile it?"""
_write_stamp(instance, {"types/entity.md": "shipped\n", "types/concept.md": "shipped\n"})
(instance / "types" / "entity.md").write_text("locally changed\n", encoding="utf-8")
assert kb_state.divergent_files() == ["types/entity.md"]
def test_a_deleted_file_counts_as_divergent(instance):
_write_stamp(instance, {"types/entity.md": "shipped\n"})
(instance / "types" / "entity.md").unlink()
assert kb_state.divergent_files() == ["types/entity.md"]
def test_divergence_is_unanswerable_without_a_stamp(instance):
"""None, not []. A development tree carries no stamp, and reporting
"nothing diverged" there would be a fabricated answer."""
assert kb_state.divergent_files() is None
+8 -17
View File
@@ -75,29 +75,20 @@ def test_new_entity_creates_page_with_expected_frontmatter(monkeypatch, kb_dir):
assert "# gateway.example.net" in body assert "# gateway.example.net" in body
def test_scaffolded_body_carries_the_headings_this_instance_declared(monkeypatch, kb_dir): def test_a_scaffolded_body_carries_no_tool_owned_region(monkeypatch, kb_dir):
"""The type-spec writes `## {section.relationships}`, not a heading text, so """A template must not scaffold the links or footnotes regions. They are
an instance in another language scaffolds its own headings without editing generated between markers from frontmatter and re-rendered on every write,
anything under `types/`. This is that path end to end.""" so a scaffolded copy would be a section the author may not edit and the tool
from chemenu import conventions would replace anyway - and, before the markers existed, a second one it
appended beside."""
(kb_dir / conventions.CONVENTIONS_FILENAME).write_text(
"---\nlanguage: fr\nprofile: none\nsections:\n relationships: Relations\n"
" see_also: Voir aussi\n footnotes: Notes\n---\n\n# conventions\n",
encoding="utf-8",
)
conventions.reset_cache()
try:
result = _invoke_new(monkeypatch, kb_dir, [ result = _invoke_new(monkeypatch, kb_dir, [
"new", "entity", "--name", "passerelle", "--set", "entity_type=system", "new", "entity", "--name", "passerelle", "--set", "entity_type=system",
]) ])
assert result.exit_code == 0, result.output assert result.exit_code == 0, result.output
_fm, body = read_page(kb_dir / "entities/systems/passerelle.md") _fm, body = read_page(kb_dir / "entities/systems/passerelle.md")
assert "## Relations" in body assert "wikitool:links" not in body
assert "## Voir aussi" in body assert "wikitool:footnotes" not in body
assert "{section." not in body assert "{section." not in body
finally:
conventions.reset_cache()
def test_new_entity_applies_schema_declared_defaults(monkeypatch, kb_dir): def test_new_entity_applies_schema_declared_defaults(monkeypatch, kb_dir):
+22 -6
View File
@@ -41,7 +41,11 @@ def empty_kb(tmp_path, monkeypatch):
for collection in COLLECTIONS: for collection in COLLECTIONS:
(kb / collection).mkdir(parents=True) (kb / collection).mkdir(parents=True)
(kb / collection / "COLLECTION.md").write_text( (kb / collection / "COLLECTION.md").write_text(
f"# kb/{collection}/ - Collection Contract\n", encoding="utf-8" f"---\nprofile: {collection}\n"
f"required_by_stack: {'true' if collection == 'sources' else 'false'}\n"
"outbound:\n any: [implements, uses, see-also]\n---\n\n"
f"# kb/{collection}/ - Collection Contract\n",
encoding="utf-8",
) )
raw = tmp_path / "raw" raw = tmp_path / "raw"
raw.mkdir() raw.mkdir()
@@ -76,7 +80,8 @@ def build_wiki(kb):
"--set", "summary=A concept created by the pipeline test"]) "--set", "summary=A concept created by the pipeline test"])
for page in kb.rglob("Pipeline *.md"): for page in kb.rglob("Pipeline *.md"):
finish_page(page) finish_page(page)
invoke(["xref", "add", "--a", "Pipeline Host", "--b", "Pipeline Concept"]) invoke(["xref", "add", "--a", "Pipeline Host", "--b", "Pipeline Concept",
"--rel", "implements"])
invoke(["index", "rebuild"]) invoke(["index", "rebuild"])
@@ -101,12 +106,23 @@ def test_a_wiki_built_by_the_tools_lints_clean(empty_kb):
assert not has_hard_errors(report), f"hard errors after a clean build: {found}" assert not has_hard_errors(report), f"hard errors after a clean build: {found}"
def test_the_pages_reach_each_other(empty_kb): def test_an_edge_points_one_way_and_the_far_end_stops_being_an_orphan(empty_kb):
"""`xref add` is what makes two pages findable from one another; if it and """`xref add` declares one direction, and that is what the orphan check now
the link checker disagreed, the lint above would report a broken link.""" measures: reachability.
It used to assert that *neither* page was an orphan, which only held because
`xref add` wrote a mirror edge on the target. With authored directional
edges the source of the only edge in a two-page wiki genuinely has nothing
pointing at it - so the check reporting it is the check working, not a
regression. A real corpus answers this by having entry points that other
pages point at."""
build_wiki(empty_kb) build_wiki(empty_kb)
assert run_lint(empty_kb)["orphan_pages"] == [] report = run_lint(empty_kb)
assert "Pipeline Concept" not in report["orphan_pages"]
assert report["orphan_pages"] == ["Pipeline Host"]
assert report["broken_links"] == []
assert report["dangling_frontmatter_refs"] == []
def test_the_catalog_covers_what_was_created(empty_kb): def test_the_catalog_covers_what_was_created(empty_kb):
+7 -6
View File
@@ -1,6 +1,5 @@
from typer.testing import CliRunner from typer.testing import CliRunner
from chemenu import conventions
from chemenu.cli import app from chemenu.cli import app
runner = CliRunner() runner = CliRunner()
@@ -44,11 +43,13 @@ def test_types_describe_entity_reports_schema_and_body():
"project", "system", "tool", "technology", "person", "project", "system", "tool", "technology", "person",
] ]
assert fields_by_name["tags"]["required"] is False assert fields_by_name["tags"]["required"] is False
# The body must carry the page skeleton an authoring LLM works from. Anchored on the # The body must carry the page skeleton an authoring LLM works from...
# template *variable* rather than on any heading text: the spec no longer names the assert "## Kerndaten" in data["body"]
# tool-owned sections at all - `kb/CONVENTIONS.md` does, and `new` substitutes it - so a # ...and must *not* carry a tool-owned region. Those are generated between
# literal here would assert the very coupling that was removed. # markers from frontmatter, so scaffolding one would create a section the
assert f"## {{section.{conventions.RELATIONSHIPS}}}" in data["body"] # author may not edit and the next write would replace anyway.
assert "wikitool:links" not in data["body"]
assert "wikitool:footnotes" not in data["body"]
def test_types_describe_unknown_name_fails_cleanly(): def test_types_describe_unknown_name_fails_cleanly():
+194 -73
View File
@@ -1,25 +1,56 @@
from chemenu import blocks, links
from chemenu.frontmatter_io import read_page from chemenu.frontmatter_io import read_page
from chemenu.commands.xref import ( from chemenu.commands.xref import (
add_related, apply_links_block,
add_relationship_bullet,
add_see_also_bullet,
remove_link_bullets, remove_link_bullets,
remove_related, remove_related,
render_links_block,
) )
from chemenu.kb_scan import load_kb_pages from chemenu.kb_scan import load_kb_pages
from chemenu.page import Page
from pathlib import Path
def test_add_related_is_deduplicated(): def _page(related, body="\n# X\n\nProse.\n"):
fm = {"related": ["A"]} return Page(Path("kb/entities/X.md"), {"related": related}, body)
assert add_related(fm, "B") is True
assert add_related(fm, "B") is False
assert fm["related"] == ["A", "B"]
def test_remove_related_is_the_inverse_of_add(): def test_an_edge_carries_its_label_in_the_data():
fm = {"related": ["A", "B"]} fm = {}
assert links.upsert(fm, "related", links.Edge("B", "depends-on")) is True
assert links.upsert(fm, "related", links.Edge("B", "depends-on")) is False
assert fm["related"] == [{"depends-on": "B"}]
def test_relabelling_replaces_rather_than_appends():
"""One page asserts one thing about another. Two edges to the same target
would render two bullets with no way to say which is meant."""
fm = {"related": [{"uses": "B"}]}
assert links.upsert(fm, "related", links.Edge("B", "depends-on")) is True
assert fm["related"] == [{"depends-on": "B"}]
def test_a_bare_title_reads_as_an_unlabelled_edge():
"""The shape every page is in between this machinery landing and the
migration reaching it. Readers must not crash on it, and it must stay
visibly unlabelled so `lint` can report it."""
fm = {"related": ["B", {"uses": "C"}]}
edges = links.edges(fm, "related")
assert [(e.target, e.label) for e in edges] == [("B", None), ("C", "uses")]
assert links.targets(fm, "related") == ["B", "C"]
assert [e.is_labelled for e in edges] == [False, True]
def test_a_malformed_entry_is_reported_rather_than_guessed_at():
fm = {"related": [{"a": "X", "b": "Y"}, 42, "Fine"]}
assert links.targets(fm, "related") == ["Fine"]
assert len(links.malformed(fm, "related")) == 2
def test_remove_related_is_the_inverse_of_upsert():
fm = {"related": [{"uses": "A"}, {"uses": "B"}]}
assert remove_related(fm, "B") is True assert remove_related(fm, "B") is True
assert fm["related"] == ["A"] assert fm["related"] == [{"uses": "A"}]
assert remove_related(fm, "B") is False assert remove_related(fm, "B") is False
@@ -27,44 +58,58 @@ def test_remove_related_tolerates_a_missing_field():
assert remove_related({}, "B") is False assert remove_related({}, "B") is False
def test_remove_link_bullets_removes_what_add_wrote(): def test_retarget_keeps_the_label():
body = "\n# X\n\n## Relationships\n\n- **uses:** [[B]]\n\n## See Also\n\n- [[B]]\n" fm = {"related": [{"depends-on": "Old"}]}
assert links.retarget(fm, "related", "Old", "New") is True
assert fm["related"] == [{"depends-on": "New"}]
def test_the_body_block_is_rendered_from_the_frontmatter():
"""The body is a rendering of the graph, not a second place it is stored.
That is what removed the need to parse a German bullet back into a
relationship."""
page = _page([{"depends-on": "Hermes"}, "Unlabelled"])
body = apply_links_block(page)
region = blocks.find(body, blocks.LINKS)
assert "- **depends-on:** [[Hermes]]" in region
assert "- [[Unlabelled]]" in region
def test_rendering_is_idempotent_and_replaces_rather_than_appends():
page = _page([{"uses": "A"}])
once = apply_links_block(page)
twice = apply_links_block(page, once)
assert once == twice
page.frontmatter["related"] = [{"uses": "B"}]
thrice = apply_links_block(page, once)
assert thrice.count("<!-- wikitool:links -->") == 1
assert "[[A]]" not in thrice and "[[B]]" in thrice
def test_a_page_with_no_edges_carries_no_region():
page = _page([])
body = apply_links_block(page)
assert "wikitool:links" not in body
assert render_links_block(page) == ""
def test_prose_after_the_region_survives_a_rewrite():
"""The failure the marker pair exists to make impossible. The old block ran
to the next heading - and before that to the end of the file - so a section
sitting after it was deleted on the next write. Eight pages were carrying
content in that position when it was found."""
page = _page([{"uses": "A"}])
body = apply_links_block(page) + "\n## Afterwards\n\nKeep me.\n"
page.frontmatter["related"] = [{"uses": "B"}]
rewritten = apply_links_block(page, body)
assert "Keep me." in rewritten
assert rewritten.count("## Afterwards") == 1
def test_remove_link_bullets_removes_what_the_renderer_wrote():
body = "\n# X\n\n## Beziehungen\n\n- **uses:** [[B]]\n- [[B]]\n"
result = remove_link_bullets(body, "B") result = remove_link_bullets(body, "B")
assert "[[B]]" not in result assert "[[B]]" not in result
assert "## Relationships" in result and "## See Also" in result
def test_relationship_bullet_idempotent():
body = "\n# X\n\n## Relationships\n\n- **Related to:** [[A]]\n\n## See Also\n\n- [[A]]\n"
once = add_relationship_bullet(body, "hosts", "B")
twice = add_relationship_bullet(once, "hosts", "B")
assert once == twice
assert "[[B]]" in once
def test_see_also_bullet_creates_section_if_missing():
body = "\n# X\n\n## Description\n\nSomething.\n"
updated = add_see_also_bullet(body, "Y")
assert "## Siehe auch" in updated
assert "[[Y]]" in updated
def test_see_also_bullet_appends_to_an_untranslated_section():
"""A page still carrying the English heading is appended to, not given a
second section - that is what lets the corpus migrate page by page."""
body = "\n# X\n\n## Description\n\nSomething.\n\n## See Also\n\n- [[A]]\n"
updated = add_see_also_bullet(body, "Y")
assert updated.count("## See Also") == 1
assert "## Siehe auch" not in updated
assert "[[Y]]" in updated
def test_relationship_bullet_appends_to_an_untranslated_section():
body = "\n# X\n\n## Relationships\n\n- **Related to:** [[A]]\n"
updated = add_relationship_bullet(body, "hosts", "B")
assert updated.count("## Relationships") == 1
assert "## Beziehungen" not in updated
assert "- **hosts:** [[B]]" in updated
def test_xref_add_updates_both_pages_on_disk(kb_dir): def test_xref_add_updates_both_pages_on_disk(kb_dir):
@@ -76,23 +121,25 @@ def test_xref_add_updates_both_pages_on_disk(kb_dir):
config.INDEX_FILE = kb_dir / "index.md" config.INDEX_FILE = kb_dir / "index.md"
runner = CliRunner() runner = CliRunner()
result = runner.invoke(app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel-a", "uses", "--rel-b", "used by"]) modbus_before = (kb_dir / "concepts/Modbus.md").read_text(encoding="utf-8")
result = runner.invoke(app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel", "uses"])
assert result.exit_code == 0, result.output assert result.exit_code == 0, result.output
pages = load_kb_pages(kb_dir) pages = load_kb_pages(kb_dir)
assert "Modbus" in pages["gdeploy"].frontmatter["related"] assert links.targets(pages["gdeploy"].frontmatter, "related") == ["Modbus"]
assert "gdeploy" in pages["Modbus"].frontmatter["related"] assert "- **uses:** [[Modbus]]" in pages["gdeploy"].body
assert "[[Modbus]]" in pages["gdeploy"].body
assert "[[gdeploy]]" in pages["Modbus"].body
fm_before, body_before = read_page(kb_dir / "entities/tools/gdeploy.md") # B is not touched at all. Its inbound view is rendered from the graph, so
link_count_before = body_before.count("[[Modbus]]") # one in Relationships, one in See Also # nothing has to be written there for a reader to find its way back.
assert (kb_dir / "concepts/Modbus.md").read_text(encoding="utf-8") == modbus_before
# Re-running must not duplicate the relationship or See Also bullets. _fm_before, body_before = read_page(kb_dir / "entities/tools/gdeploy.md")
result2 = runner.invoke(app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel-a", "uses", "--rel-b", "used by"]) link_count_before = body_before.count("[[Modbus]]")
result2 = runner.invoke(app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel", "uses"])
assert result2.exit_code == 0 assert result2.exit_code == 0
fm_after, body_after = read_page(kb_dir / "entities/tools/gdeploy.md") fm_after, body_after = read_page(kb_dir / "entities/tools/gdeploy.md")
assert fm_after["related"].count("Modbus") == 1 assert links.targets(fm_after, "related").count("Modbus") == 1
assert body_after.count("[[Modbus]]") == link_count_before assert body_after.count("[[Modbus]]") == link_count_before
@@ -107,14 +154,14 @@ def test_xref_remove_undoes_xref_add(kb_dir):
runner = CliRunner() runner = CliRunner()
before = (kb_dir / "entities/tools/gdeploy.md").read_text(encoding="utf-8") before = (kb_dir / "entities/tools/gdeploy.md").read_text(encoding="utf-8")
added = runner.invoke(app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus"]) added = runner.invoke(app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel", "uses"])
assert added.exit_code == 0, added.output assert added.exit_code == 0, added.output
removed = runner.invoke(app, ["xref", "remove", "--a", "gdeploy", "--b", "Modbus"]) removed = runner.invoke(app, ["xref", "remove", "--a", "gdeploy", "--b", "Modbus"])
assert removed.exit_code == 0, removed.output assert removed.exit_code == 0, removed.output
pages = load_kb_pages(kb_dir) pages = load_kb_pages(kb_dir)
assert "Modbus" not in pages["gdeploy"].frontmatter["related"] assert "Modbus" not in links.targets(pages["gdeploy"].frontmatter, "related")
assert "gdeploy" not in pages["Modbus"].frontmatter["related"] assert "gdeploy" not in links.targets(pages["Modbus"].frontmatter, "related")
assert "[[Modbus]]" not in pages["gdeploy"].body assert "[[Modbus]]" not in pages["gdeploy"].body
assert before # sanity: fixture page was non-empty assert before # sanity: fixture page was non-empty
@@ -198,10 +245,10 @@ def test_xref_add_dry_run_writes_nothing(kb_dir):
runner = CliRunner() runner = CliRunner()
result = runner.invoke( result = runner.invoke(
app, app,
["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel-a", "uses", "--rel-b", "used by", "--dry-run"], ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel", "uses", "--dry-run"],
) )
assert result.exit_code == 0, result.output assert result.exit_code == 0, result.output
assert "would update" in result.output assert "would declare" in result.output
assert "No files written" in result.output assert "No files written" in result.output
assert gdeploy_path.read_text(encoding="utf-8") == gdeploy_before assert gdeploy_path.read_text(encoding="utf-8") == gdeploy_before
@@ -230,9 +277,11 @@ def test_xref_link_source_dry_run_writes_nothing(kb_dir):
assert gdeploy_path.read_text(encoding="utf-8") == gdeploy_before assert gdeploy_path.read_text(encoding="utf-8") == gdeploy_before
def test_xref_add_reports_a_write_failure_without_silently_leaving_a_one_way_link(kb_dir, monkeypatch): def test_xref_add_reports_a_write_failure(kb_dir, monkeypatch):
"""If writing B fails after A already succeeded, the command must fail """One edge, one write - so there is no half-written pair to report any
loudly (not silently succeed with a one-directional link) and say so.""" more. The old two-sided `xref add` could update A and fail on B, leaving a
link the user had to be told was one-directional; a directional edge has
nothing to be half of."""
from typer.testing import CliRunner from typer.testing import CliRunner
from chemenu.cli import app from chemenu.cli import app
from chemenu.commands import xref as xref_module from chemenu.commands import xref as xref_module
@@ -241,25 +290,40 @@ def test_xref_add_reports_a_write_failure_without_silently_leaving_a_one_way_lin
config.KB_DIR = kb_dir config.KB_DIR = kb_dir
config.INDEX_FILE = kb_dir / "index.md" config.INDEX_FILE = kb_dir / "index.md"
real_write_page = xref_module.write_page
def flaky_write_page(path, frontmatter, body): def flaky_write_page(path, frontmatter, body):
if path.name == "Modbus.md":
raise OSError("disk full") raise OSError("disk full")
return real_write_page(path, frontmatter, body)
monkeypatch.setattr(xref_module, "write_page", flaky_write_page) monkeypatch.setattr(xref_module, "write_page", flaky_write_page)
runner = CliRunner() runner = CliRunner()
result = runner.invoke( result = runner.invoke(
app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel-a", "uses", "--rel-b", "used by"] app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel", "uses"]
) )
assert result.exit_code == 1 assert result.exit_code == 1
assert "disk full" in result.output assert "disk full" in result.output
assert "one-directional" in result.output
pages = load_kb_pages(kb_dir) pages = load_kb_pages(kb_dir)
assert "Modbus" in pages["gdeploy"].frontmatter["related"] # A's write already happened assert "Modbus" not in links.targets(pages["gdeploy"].frontmatter, "related")
def test_xref_add_refuses_a_label_the_collection_does_not_authorise(kb_dir):
"""The source collection decides which labels may be used from it. Refused
here rather than only in `lint`, because this is the moment the author is
present and can pick a better one."""
from typer.testing import CliRunner
from chemenu.cli import app
import chemenu.config as config
config.KB_DIR = kb_dir
config.INDEX_FILE = kb_dir / "index.md"
runner = CliRunner()
result = runner.invoke(
app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel", "hängt ab von"]
)
assert result.exit_code == 1
assert "not authorised" in result.output
assert "link-taxonomy" in result.output
def test_xref_link_source_distinguishes_write_failures_from_missing_pages(kb_dir, monkeypatch): def test_xref_link_source_distinguishes_write_failures_from_missing_pages(kb_dir, monkeypatch):
@@ -334,7 +398,9 @@ def test_xref_add_refuses_a_type_without_a_related_field(kb_dir):
`related:` there produced frontmatter the schema rejects, and `xref remove` `related:` there produced frontmatter the schema rejects, and `xref remove`
could not clear it - one command creating a state another could not undo.""" could not clear it - one command creating a state another could not undo."""
runner, app = _runner_env(kb_dir) runner, app = _runner_env(kb_dir)
result = runner.invoke(app, ["xref", "add", "--a", "Source - Aurora", "--b", "aurora"]) result = runner.invoke(
app, ["xref", "add", "--a", "Source - Aurora", "--b", "aurora", "--rel", "uses"]
)
assert result.exit_code == 1 assert result.exit_code == 1
# Rich wraps the message to the terminal width, so compare on collapsed # Rich wraps the message to the terminal width, so compare on collapsed
# whitespace rather than pinning the line breaks. # whitespace rather than pinning the line breaks.
@@ -414,3 +480,58 @@ def test_xref_link_source_dry_run_leaves_the_source_page_alone(kb_dir):
"--dry-run"], "--dry-run"],
) )
assert path.read_text(encoding="utf-8") == before assert path.read_text(encoding="utf-8") == before
# --- the inbound view ------------------------------------------------------
def test_the_inbound_view_is_derived_not_stored(kb_dir):
"""The load-bearing half of dropping mirrored edges. Nothing writes an edge
onto the target, so the only way "what points at this page" can be answered
completely is by computing it - which is also why it cannot go stale or be
half-written the way a mirror could."""
from typer.testing import CliRunner
from chemenu.cli import app
import chemenu.config as config
config.KB_DIR = kb_dir
config.INDEX_FILE = kb_dir / "index.md"
runner = CliRunner()
assert runner.invoke(
app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel", "uses"]
).exit_code == 0
import json
result = runner.invoke(app, ["links", "show", "--page", "Modbus", "--json"])
assert result.exit_code == 0, result.output
data = json.loads(result.output)
assert data["inbound"] == [{"source": "gdeploy", "label": "uses", "collection": "entities"}]
# Nothing was written onto Modbus itself to make that answer possible.
assert "gdeploy" not in links.targets(
load_kb_pages(kb_dir)["Modbus"].frontmatter, "related"
)
def test_an_unresolvable_outbound_edge_is_marked(kb_dir):
"""`lint` reports dangling references corpus-wide; this reports it for the
one page someone is looking at, which is where it gets fixed."""
from typer.testing import CliRunner
from chemenu.cli import app
import chemenu.config as config
import json
config.KB_DIR = kb_dir
config.INDEX_FILE = kb_dir / "index.md"
page = load_kb_pages(kb_dir)["gdeploy"]
links.upsert(page.frontmatter, "related", links.Edge("Gone", "uses"))
from chemenu.frontmatter_io import write_page
write_page(page.path, page.frontmatter, page.body)
result = CliRunner().invoke(app, ["links", "show", "--page", "gdeploy", "--json"])
assert json.loads(result.output)["outbound"] == [
{"target": "Gone", "label": "uses", "resolves": False}
]
+14 -1
View File
@@ -27,8 +27,21 @@ properties:
related: related:
type: array type: array
items: items:
oneOf:
- type: string
- type: object
minProperties: 1
maxProperties: 1
additionalProperties:
type: string type: string
description: Related concept and entity titles description: >-
Declared outbound edges. Each entry is either `<label>: <page title>` - the
label drawn from instructions/link-taxonomy.md and authorised per
destination by the source collection's `outbound:` block - or a bare page
title for an edge whose label has not been declared yet. The bare form is
the pre-4.0.0 shape and is what `lint` reports until the migration reaches
the page; it is accepted rather than rejected so that a corpus stays
readable while it is being converted.
sources: sources:
type: array type: array
items: items:
+4 -15
View File
@@ -77,12 +77,6 @@ TODO: 1-2 Absätze dazu, was diese Entity ist und wozu sie dient.
- **Verantwortlich:** TODO (falls zutreffend) - **Verantwortlich:** TODO (falls zutreffend)
- **Repository:** TODO (falls zutreffend) - **Repository:** TODO (falls zutreffend)
## {section.relationships}
- **Hängt ab von:** TODO
- **Verwendet von:** TODO
- **Verwandt mit:** TODO
## Details ## Details
TODO: Ausführliche Informationen, nach sinnvollen Abschnitten gegliedert TODO: Ausführliche Informationen, nach sinnvollen Abschnitten gegliedert
@@ -90,17 +84,12 @@ TODO: Ausführliche Informationen, nach sinnvollen Abschnitten gegliedert
## Historie ## Historie
- [{today}] - Page created via wikitool - [{today}] - Page created via wikitool
## {section.see_also}
- TODO: Verwandte Seiten
``` ```
Die beiden `{section.…}`-Platzhalter sind toolgeführte Abschnitte: `wikitool xref` schreibt in Der Beziehungsabschnitt steht bewusst **nicht** im Template: er ist eine generierte Region, die
genau sie hinein, und wie sie heißen, entscheidet die Instanz in `kb/CONVENTIONS.md` `wikitool xref` beim ersten Kanteneintrag zwischen Markern anlegt und aus `related:` neu
(`sections:`) - nicht dieser Type-Spec und nicht der Compiler. `wikitool new` setzt den rendert. Ein Autor schreibt dort nie hinein. Der Wert hinter `**Typ:**` bleibt der englische
aktuellen Namen ein. Der Wert hinter `**Typ:**` bleibt der englische Enum-Wert - danach filtert Enum-Wert - danach filtert `search --field`.
`search --field`.
--- ---
+14 -1
View File
@@ -27,8 +27,21 @@ properties:
related: related:
type: array type: array
items: items:
oneOf:
- type: string
- type: object
minProperties: 1
maxProperties: 1
additionalProperties:
type: string type: string
description: Related page titles description: >-
Declared outbound edges. Each entry is either `<label>: <page title>` - the
label drawn from instructions/link-taxonomy.md and authorised per
destination by the source collection's `outbound:` block - or a bare page
title for an edge whose label has not been declared yet. The bare form is
the pre-4.0.0 shape and is what `lint` reports until the migration reaches
the page; it is accepted rather than rejected so that a corpus stays
readable while it is being converted.
sources: sources:
type: array type: array
items: items:
+15
View File
@@ -43,6 +43,21 @@ properties:
scriptable; `assisted` needs a judgment call per page and is therefore an scriptable; `assisted` needs a judgment call per page and is therefore an
agent procedure. Today this is a description rather than an execution agent procedure. Today this is a description rather than an execution
promise - there is no `migrate run`. promise - there is no `migrate run`.
obligation:
type: string
enum: [required, offered]
default: required
description: >-
Whether the migration must run at all - a separate axis from
`migration_kind`, which says only how the work is done. `required`
(the default) is the original meaning: the content must reach the new
shape or it no longer fits the machinery, so `migrate status` counts it
as outstanding and `migrate done` advances `kb_version` through it.
`offered` is an upgrade the instance may decline: a file it owns still
works as it is, and the stack is proposing a better default. An offered
migration never blocks, never appears in the outstanding chain, and is
not a link in the version chain - it is listed separately so an operator
can take it when they want it.
required: required:
- type - type
- name - name
+28 -5
View File
@@ -40,6 +40,30 @@ under `kb/`. It is deliberately **not a collection** and carries no `COLLECTION.
no per-collection type surface, and `tools/wikitool docs verify` fails if a contract appears no per-collection type surface, and `tools/wikitool docs verify` fails if a contract appears
here. here.
### Who owns a type-spec
`types/` holds two kinds of file, and the line between them is `root:` — already in the
frontmatter before anyone drew it:
| Type-spec | Describes | Owned by | Ships as |
|---|---|---|---|
| `root: kb` (`entity`, `concept`, `source`, `comparison`) | A page **this instance** writes | The instance | `types/<name>.md.template` plus its `.schema.yaml.template`, adopted by a rename |
| `root: repo` (`instruction`), no `base_dir` (`lint-report`), and `type-spec` itself | A stack artifact | The stack | Verbatim |
A page type-spec's prose, its `## Template` body and its language are therefore the instance's
to rewrite — an instance writing its pages in another language simply translates the file, and
an upgrade does not take that back. Improvements to a shipped default reach it as an *offered*
migration ([instructions/CONTRACT.md](../instructions/CONTRACT.md#instructionsmigrations)),
never by overwriting.
**What the stack still requires of the type layer is one line.** There must be a type-spec
declaring `name: source` whose schema requires `raw_files:` — the whole `raw/``kb/`
provenance path (`sources coverage`, `[^cite-id]` resolution, `kb/provenance.md`) asks
`page.kind == "source"`, so without it nothing resolves. `docs verify` checks exactly that and
nothing beyond it: not the directory, not the title prefix, not a word of the prose. Which
collection is stack-required is *derived* from where that type writes rather than listed
separately, so renaming it stays consistent instead of tripping a hardcoded name.
**Quality goal:** a type-spec is the single source of truth for its type. No structural fact **Quality goal:** a type-spec is the single source of truth for its type. No structural fact
about a page type may be restated anywhere else — not in `AGENTS.md`, not in a skill, not in about a page type may be restated anywhere else — not in `AGENTS.md`, not in a skill, not in
Python. Adding a type must require no code change. Python. Adding a type must require no code change.
@@ -98,11 +122,10 @@ The `## Template` block is filled from the page's own frontmatter, plus `{name}`
`{entities|table_cells}`. `{field|literal text}` falls back to the literal when the field is `{entities|table_cells}`. `{field|literal text}` falls back to the literal when the field is
absent. absent.
Three further variables come from the instance rather than from the page: **A template never contains a tool-owned region.** The links and footnotes regions are generated
`{section.relationships}`, `{section.see_also}` and `{section.footnotes}`, filled from between markers by `xref` and `cite`, rendered from frontmatter, and re-rendered on every write -
`kb/CONVENTIONS.md`'s `sections:` declaration. A template writes a tool-owned heading through so scaffolding them would create a section an author is forbidden to edit and the tool would
one of those and never as literal text - that is what lets an instance change the KB language replace anyway. See `tools/chemenu/blocks.py`.
without editing anything under `types/`.
### Ownership boundary ### Ownership boundary
+88
View File
@@ -0,0 +1,88 @@
# Workshop: link-taxonomy-migration
- **Run key:** `link-taxonomy-migration` (this directory's name - there is no other identifier)
- **Input:** none - this run is not an ingest
- **Started:** 2026-09-02
- **Session id form:** `WIKITOOL_SESSION_ID="link-taxonomy-migration/u<N>"`, one per unit
- **Issue:** Gitea #40, sections *Label werden Enum* and *Toolgeführte Blöcke*
## Goal
Move every relationship in `kb/` from free-text German prose in a body bullet to a machine
value in `related:`, and every tool-owned body region from heading-matching to a marker pair.
Afterwards the compiler contains no heading text and no relationship label, and `lint` can
enforce the vocabulary because there is one.
## Why this is `assisted` and not `mechanical`
Measured on 2026-09-02, against the corpus rather than against the documentation:
| | |
|---|---|
| Pages | 180 |
| Distinct relationship labels in body bullets | **152** |
| Labelled bullets | 337 |
| Labels occurring exactly once | 102 |
| Top 20 labels cover | 167 of 337 |
| Bare `- [[X]]` bullets under `## Siehe auch` | 555 |
| ...of those, provably redundant (a labelled edge already exists) | 353 |
| ...of those, the only connection between the two pages | **202**, across 63 target pages |
`kb/CONVENTIONS.md` documents thirteen labels. Nothing ever checked that, and the corpus does
not follow it - so there is no mapping table to apply, and roughly 539 edges need a judgment
call each. A large minority are reverse directions (`Verwendet von` 20x, `implementiert durch`,
`Ersetzt durch`), which under authored directional edges are exactly the edges that stop being
stored and start being rendered.
Rejected alternative, recorded so it is not re-proposed: map the top 20 mechanically and set
everything else to `see-also`. That would start the new taxonomy with ~370 of ~539 edges on its
weakest label - the `verwandt mit` sediment this whole change exists to end, re-created as the
documented initial state.
## Closes when
Every unit in `plan.md` is published, and:
- `wikitool lint` reports zero unlabelled edges and zero labels outside the authorising
collection's `outbound:` block
- `wikitool migrate verify --from <pre-migration rev>` reports no wikilink or citation count
change, and no unbalanced marker
- `wikitool migrate done 4.0.0 --pages <N>` has run
## Checklist
- [x] u0 mechanism - taxonomy, `links.py`, `blocks.py`, `xref` rewrite, lint checks,
`links show` for the inbound view, deletion of the matching layer. No page touched.
- [ ] u1 `kb/entities/` (systems, tools, technologies)
- [ ] u2 `kb/entities/` (projects, people) + `kb/comparisons/`
- [ ] u3 `kb/concepts/`
- [ ] u4 `kb/sources/`
- [ ] u5 close-out - `migrate done`, version bump, `CHANGES.md`, workshop close
## Measured after u0
`lint` against the corpus, with the machinery in place and no page touched:
| Finding | Count |
|---|---|
| `unlabelled_edges` | **480** - every `related:` entry, since none carries a label yet |
| `unauthorised_labels` | 0 - nothing declares a label at all, so nothing can be unauthorised |
| `malformed_edges` | 0 |
| `unbalanced_markers` | 0 |
| `orphan_pages` | 1 (`GRUB`, pre-existing) |
| `broken_links`, `dangling_frontmatter_refs`, `schema_validation_errors` | 0 |
480 is the number u1-u4 have to bring to zero. It is larger than the 337 labelled body bullets
because `related:` also holds entries whose bullet was lost or never written - which is itself a
finding: the frontmatter and the body had already drifted apart under the old model, and nothing
could see it while the label lived only in the prose.
## Open decisions
- **Settled 2026-09-02:** edges are directional; the reverse edge is authored only when it is a
primary statement on its own page. The inbound view is rendered, not stored.
- **Settled 2026-09-02:** the 353 provably-redundant `## Siehe auch` edges are dropped
mechanically. The 202 that are the only connection get a real label each, or are dropped with
a reason - never converted to `see-also` in bulk.
- **Settled 2026-09-02:** labels are not localized. `- **depends-on:** [[Hermes]]` is what a
German page carries.
+48
View File
@@ -0,0 +1,48 @@
# Plan: link-taxonomy-migration
One unit is one session id and one `publish`, sized against the 60-call iteration budget.
Per-page cost here is roughly `1 touch` + the edges on it; the corpus-wide commands
(`index rebuild`, `sources rebuild-index`, `log append`, `publish` twice for the gate) are
per unit, not per page. That puts the ceiling near 45 pages and the target at 40.
| # | Unit | Pages | Job | Done when |
|---|------|------:|-----|-----------|
| u0 | mechanism | 0 | Taxonomy catalogue, `links.py`, `blocks.py`, `outbound:` in every `COLLECTION.md`, `xref` rewritten to one directional edge, lint checks, `links show` for the inbound view, `migrate verify` marker invariant, deletion of `sections.py` and the matching layer | Suite green; `lint` reports the corpus's unlabelled edges as findings rather than crashing |
| u1 | `kb/entities/systems`, `tools`, `technologies` | ~45 | Label every edge, drop redundant see-also, wrap markers | `migrate verify --path kb/entities --fail-on-error` clean |
| u2 | `kb/entities/projects`, `people`, `kb/comparisons` | ~40 | as u1 | as u1 |
| u3 | `kb/concepts` | ~45 | as u1 | `migrate verify --path kb/concepts --fail-on-error` clean |
| u4 | `kb/sources` | ~50 | as u1, plus `entities:`/`concepts:` on source pages | `migrate verify --path kb/sources --fail-on-error` clean |
| u5 | close-out | 0 | `migrate done 4.0.0`, `version bump --major`, `CHANGES.md` body, promote nothing, close workshop | `docs verify` + `lint --fail-on-error` green, workshop deleted |
Unit boundaries are written down here *before* the run so that publishing several units
together stays a planned batch rather than a way around a Mass-Update Gate refusal - see
`instructions/gates.md`.
## Per-page procedure
1. Read the page's `## Beziehungen` and `## Siehe auch` blocks.
2. For each labelled bullet: say the sentence `[this page] <label> [target]`. Pick the catalogue
label that makes it true. If it only reads true backwards, the edge belongs on the other
page - move it, do not invert the label into something the catalogue does not have.
3. For each bare `## Siehe auch` bullet: drop it if a labelled edge already connects the pair
(the tooling lists these). Otherwise decide - a real label, or dropped with the reason
recorded in the unit's notes.
4. Write the edges with `xref add --rel`, never by hand.
5. The body blocks are then *generated*: no hand-editing inside a marker pair.
## Vocabulary carried between units
`glossary.md` in this directory. A mapping decided in u1 and re-decided in u3 is the failure the
file exists to prevent - add to it **before** dispatching the next unit.
## Deliberately excluded from this run
- **Commonplace's articulation test and the `connect` report workflow.** They change how ingest
proposes links, not how links are stored. Separate question, separate issue.
- **Promoting the lint checks to hard errors.** During this run an unlabelled edge is a finding,
because that is precisely the migration window `.wikitool-kb.json` exists to represent. The
promotion is a later version's change, once the corpus can pass it.
- **`sources:` and `[^cite-id]`.** The provenance path is unlabelled by construction and is not
part of the link taxonomy.
- **Any change to page prose.** This run restates relationships in a new form; it learns
nothing new, and a body edit outside a marker pair is out of scope.