# kb/ - Knowledge Layer Contract The compiled knowledge layer, and the third stage of the pipeline `raw/` -> `kb/` -> `reports/`. Everything here is written and maintained by the LLM from material in `raw/`, and is expected to stay correct without being re-derived. **Quality goal:** a page should answer a future question *without* re-reading the source it came from. If answering still requires the raw file, the page is incomplete. This file holds the rules that apply in **every** collection **and in every instance**. That second half is the cut: what is written here is enforced by `tools/wikitool` or follows from how it works, so it is identical everywhere and `dist export` ships it verbatim. **What an instance decides for itself is next door, in [kb/CONVENTIONS.md](CONVENTIONS.md)** - the language pages are written in, the headings its two generated regions render under, the naming forms, the tone, the confidence rubric. That file binds exactly as this one does; it is simply owned by the instance rather than by the stack, so the distribution ships only its `.template` and the instance writes the real one. Read both, plus the target collection's `kb//COLLECTION.md` (also instance-owned), before writing or editing a page. The split is by **who may change the sentence**, not by what it is about. Language, tone and naming used to sit here, which meant every instance that answered "not German" to `setup-instance.md` was locally editing a file the stack also ships - and a merge from upstream would quietly hand it back. Structural facts (which frontmatter fields exist, which are required, what the body skeleton looks like) are in neither - they belong to the type-specs and are printed by `tools/wikitool types describe `. Never hand-write frontmatter; scaffold with `tools/wikitool new --name "" --set field=value ...`. ## Contents - [Collections](#collections) - [Generated files](#generated-files) - [Titles are identifiers](#titles-are-identifiers) - [Every page should](#every-page-should) - [Quotation cap](#quotation-cap) - [Language and identifiers](#language-and-identifiers) - [Generated regions](#generated-regions) - [Linking](#linking) - [Provenance and citation](#provenance-and-citation) - [Confidence](#confidence) - [Confidence against source standing](#confidence-against-source-standing) - [What does not belong here](#what-does-not-belong-here) ## Collections `kb/` is a **namespace, not a collection**. It carries no `COLLECTION.md` of its own. A directory under `kb/` is a **collection** exactly when it contains a `COLLECTION.md`. That file is the local authoring contract for every page in the subtree, and it belongs to the instance: it declares in its frontmatter which profile from [instructions/kb-profiles.md](../instructions/kb-profiles.md) it adopted, and whether the stack resolves against it by name. | Field | Means | |---|---| | `profile:` | Which catalogue entry this contract started from, or `none`. Free text - the catalogue is a palette, not an enum, and a collection an instance invented has no entry to name | | `required_by_stack:` | Whether `wikitool` itself depends on this collection *by name*. Not the instance's to choose: `docs verify` checks it against the stack's own list. `kb/sources/` is `true` - `sources coverage`, `[^cite-id]` resolution and `kb/provenance.md` all resolve against that name - and everything else is `false` | - A subdirectory *inside* a collection is an **area**. It inherits the enclosing contract and must not carry a `COLLECTION.md` of its own - `kb/entities/systems/` is an area of `kb/entities/`. - **An area is as deep as a page goes.** `kb//.md` and `kb///.md` are the two depths a page may sit at; nothing goes a level deeper. A further subdirectory is not a second-level area - it is invisible to the generated catalog, which reads exactly two path segments below `kb/` and folds anything past them into the area's own table silently, with no location of its own. That is why `wikitool lint`'s `nested_pages` finding is a hard error rather than an advisory one like `misplaced_pages`: a misplaced page still catalogs correctly from the wrong place, a nested one makes the catalog itself wrong. A grouping axis that does not come from a type-spec's `layout:` - project owner was the case that surfaced this - does not earn a second directory level; it goes into frontmatter instead. - A `COLLECTION.md` nested inside another collection is invalid. - `COLLECTION.md` appears **nowhere outside `kb/`**. `raw/`, `types/`, `tools/`, `reports/` and `instructions/` are not collections and carry a `CONTRACT.md` or a root type-spec instead. `tools/wikitool docs verify` enforces all three. | Collection | Holds | Contract | |------------|-------|----------| | `kb/entities/` | Concrete things: projects, deployed systems, tools, technologies, people | [entities/COLLECTION.md](entities/COLLECTION.md) | | `kb/concepts/` | Architectures, patterns, protocols, workflows, decisions, recurring problems | [concepts/COLLECTION.md](concepts/COLLECTION.md) | | `kb/sources/` | One summary page per ingested source, carrying its `raw_files:` provenance | [sources/COLLECTION.md](sources/COLLECTION.md) | | `kb/comparisons/` | Structured comparisons of two or more existing pages | [comparisons/COLLECTION.md](comparisons/COLLECTION.md) | The four rows above are this instance's collections, not a fixed set. **Adding one:** `mkdir kb/` and write a `kb//COLLECTION.md` with the two fields above. Collections are discovered by contract presence, so no code change is needed. A collection only becomes *writable* once some type-spec declares a matching `base_dir:`. Renaming or dropping one is the instance's call too - except where `required_by_stack: true` says otherwise. **Where a page goes** is decided by its type-spec, never by hand - see [types/type-spec.md](../types/type-spec.md). ## Generated files Never hand-edit these; they are produced by `tools/wikitool`: | File | Produced by | |------|-------------| | `kb/index.md` | `wikitool index rebuild` - the catalog **map**: statistics, counts, links | | `kb//INDEX.md` and `kb///INDEX.md` | `wikitool index rebuild` - the page tables | | `kb/log.md` | `wikitool log append` | | `kb/provenance.md` | `wikitool sources rebuild-index` | To *find* a page, search rather than read the catalog: `tools/wikitool search ""`, or `tools/wikitool search --field ` for a structured query over frontmatter. ## Titles are identifiers **The filename stem *is* the page title, and `[[wikilinks]]` must match it exactly.** That is not a naming preference; it is the wiki's only way to address a page. `wikitool lint` reports an H1 that stops matching its title, `rename`/`rm` rewrite every reference to a stem, and a `[^cite-id]` resolves through one. Which *form* those titles take - spaces or kebab-case, singular or plural, what prefixes a decision record - is the instance's, in [kb/CONVENTIONS.md § Naming](CONVENTIONS.md#naming). ## Every page should - [ ] Carry a clear, descriptive title and a summary near the top - [ ] Use consistent terminology with the rest of the wiki - [ ] Link to the entities and concepts it mentions, and declare an edge where the relationship is worth naming - in the direction this page asserts it, not in both - [ ] Cite its hard facts (see [Provenance and citation](#provenance-and-citation)) - [ ] Duplicate no existing page - [ ] Appear in the catalog (guaranteed by `wikitool index rebuild`) ## Quotation cap At most 2 blockquoted lines per page. `wikitool lint` reports overages as advisory, since exceeding the cap can be a legitimate judgment call - but the page should carry the knowledge itself, not delegate it to quotations. The cap is about how much of the page you let quotes carry; it does not apply to text you are citing verbatim from a source. The register those lines are written in - what counts as a buzzword, what filler is refused - is the instance's, in [kb/CONVENTIONS.md § Tone](CONVENTIONS.md#tone). ## Language and identifiers *Which* language pages are written in is [kb/CONVENTIONS.md](CONVENTIONS.md)'s to say. What follows here is the part that is not a choice, because the tool resolves against it. Every line of a page is either **prose** or an **identifier**. Only prose is translated. **Prose:** descriptions and definitions, `## Key Information` values, `## Details` body text, a source page's Summary / Key Takeaways / Action Items / Not Extracted, and `summary:`. **Identifiers - never translated, in any language:** | Identifier | Why | |---|---| | Page titles, and the H1 that repeats one | A title is the wiki's only identifier for a page and follows the subject's own established name - see [Titles are identifiers](#titles-are-identifiers). `wikitool lint` reports an H1 that stops matching its title | | The subtype value on the generated `**Typ:**` line | It renders a schema enum value (`technology`, `workflow`), which `search --field` filters on. The label is prose; the value is not | | `tags:` | Search keys, not prose | | Commands, paths, config keys, hostnames, code | They are what they are | | Quotations | Quoted verbatim in the source's own language | Which foreign technical terms stay untranslated inside that prose is a judgment call the instance records - see [kb/CONVENTIONS.md § Language](CONVENTIONS.md#language). **A source in another language** is still summarized in the KB language: a source page is evidence *about* a source, not a substitute for it. Quote verbatim in the original language and record the raw file's language in `source_language:`. ### Generated regions Two regions of a page body are **generated**, not authored: the links region `xref` owns and the footnotes region `cite` owns. Each sits between a marker pair: ```markdown ## Beziehungen - **depends-on:** [[Hermes]] ``` The marker is what the tool locates the region by, and everything between the markers - **heading included** - is replaced wholesale on the next write. An author never edits inside them; anything left there is overwritten without warning, exactly as in `kb/index.md`. A region with nothing to show is absent rather than empty. The heading is therefore a *rendering* value, taken from `kb/CONVENTIONS.md`'s `sections:`. No heading text exists in the compiler, and nothing matches on it: changing the declaration re-renders the words on the next write and cannot split a page. That is not how it used to work. The tool located these regions by matching their heading text, which made a translated heading a structural fact - and made the region's *end* a guess. It ran to the next heading, and before that to the end of the file, which silently deleted whatever sat after it on eight pages. Any *other* heading a page carries is ordinary prose. ## Linking **An edge is authored in one direction**, on the page that asserts it, and carries a label that is a machine value rather than prose: ```yaml related: - depends-on: Hermes ``` Created with `tools/wikitool xref add --a "" --b "" --rel