kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59)
CI / verify (push) Successful in 52s
Release / release (push) Successful in 36s

Files changed:
- CHANGES.md
- README.md
- VERSION
- kb/concepts/Ambient Environment Dependency.md
- kb/concepts/Anti-Cramming Heuristic.md
- kb/concepts/Audit Trail.md
- kb/concepts/BM25.md
- kb/concepts/Bulk Operations.md
- kb/concepts/CI Integration.md
- kb/concepts/COLLECTION.md
- kb/concepts/CPPC.md
- kb/concepts/Checkpoint Audit.md
- kb/concepts/Claude Code Auto Mode.md
- kb/concepts/Command Round-Trip Integrity.md
- kb/concepts/Confidence Scoring.md
- kb/concepts/Consolidation Tiers.md
- kb/concepts/Content Quality Control.md
- kb/concepts/Context Isolation.md
- kb/concepts/Contradiction Resolution.md
- kb/concepts/Cross-platform Agent Skills.md
- kb/concepts/Crystallization.md
- kb/concepts/Delete Rather Than Anonymize.md
- kb/concepts/Denylist over Allowlist.md
- kb/concepts/Detect-Repair Asymmetry.md
- kb/concepts/Diff-Reviewable Agent Edits.md
- kb/concepts/Dual Licensing by File Plan.md
- kb/concepts/Entity Extraction.md
- kb/concepts/Episodic Memory.md
- kb/concepts/Event-Driven Automation.md
- kb/concepts/Filter on Ingest.md
- kb/concepts/Forgetting.md
- kb/concepts/Graph Traversal.md
- kb/concepts/Green Suite Blind Spot.md
- kb/concepts/Hooks.md
- kb/concepts/Hybrid Search.md
- kb/concepts/INDEX.md
- kb/concepts/Implementation Spectrum.md
- kb/concepts/Index Scaling.md
- kb/concepts/Issue Label Scheme.md
- kb/concepts/Iteration and Cost Limits.md
- kb/concepts/KB Migration.md
- kb/concepts/KB Stack Versioning.md
- kb/concepts/Knowledge Compounding.md
- kb/concepts/Knowledge Graph.md
- kb/concepts/LLM Wiki Pattern.md
- kb/concepts/Lint Workflow.md
- kb/concepts/MCP-Leseserver.md
- kb/concepts/Mass-Update Gate.md
- kb/concepts/Memory Lifecycle.md
- kb/concepts/Mesh Sync.md
- kb/concepts/Modbus.md
- kb/concepts/Multi-Agent Collaboration.md
- kb/concepts/Naming Convention Conflict.md
- kb/concepts/OKF Compatibility.md
- kb/concepts/Optional Instance Context File.md
- kb/concepts/Personalization Plane.md
- kb/concepts/Privacy and Governance.md
- kb/concepts/Procedural Memory.md
- kb/concepts/Publish-Remote Gate.md
- kb/concepts/Quality Scoring.md
- kb/concepts/Quality and Self-Correction.md
- kb/concepts/RAG.md
- kb/concepts/Reciprocal Rank Fusion.md
- kb/concepts/SSD TRIM.md
- kb/concepts/Scale Ceiling.md
- kb/concepts/Self-Healing.md
- kb/concepts/Semantic Lint Automation.md
- kb/concepts/Semantic Memory.md
- kb/concepts/Session Orientation.md
- kb/concepts/Shared vs Private.md
- kb/concepts/Split Merge Reclassify.md
- kb/concepts/Split Threshold.md
- kb/concepts/Structural Enforcement over Documented Rule.md
- kb/concepts/Stub Threshold.md
- kb/concepts/Supersession.md
- kb/concepts/Three-Layer Architecture.md
- kb/concepts/Token Economics.md
- kb/concepts/Typed Relationships.md
- kb/concepts/User Management.md
- kb/concepts/Vector Search.md
- kb/concepts/Work Coordination.md
- kb/concepts/Workflow Extraction.md
- kb/concepts/Workflow Orchestration.md
- kb/concepts/Working Memory.md
- kb/concepts/Write-Once Frontmatter Fields.md
- kb/concepts/architectures/Consolidation Tiers.md
- kb/concepts/architectures/Context Isolation.md
- kb/concepts/architectures/Cross-platform Agent Skills.md
- kb/concepts/architectures/Episodic Memory.md
- kb/concepts/architectures/Hybrid Search.md
- kb/concepts/architectures/Implementation Spectrum.md
- kb/concepts/architectures/Knowledge Graph.md
- kb/concepts/architectures/LLM Wiki Pattern.md
- kb/concepts/architectures/MCP-Leseserver.md
- kb/concepts/architectures/Memory Lifecycle.md
- kb/concepts/architectures/OKF Compatibility.md
- kb/concepts/architectures/Optional Instance Context File.md
- kb/concepts/architectures/Personalization Plane.md
- kb/concepts/architectures/Procedural Memory.md
- kb/concepts/architectures/RAG.md
- kb/concepts/architectures/Scale Ceiling.md
- kb/concepts/architectures/Semantic Memory.md
- kb/concepts/architectures/Three-Layer Architecture.md
- kb/concepts/architectures/Token Economics.md
- kb/concepts/architectures/Working Memory.md
- kb/concepts/decisions/Delete Rather Than Anonymize.md
- kb/concepts/decisions/Denylist over Allowlist.md
- kb/concepts/decisions/Diff-Reviewable Agent Edits.md
- kb/concepts/decisions/Dual Licensing by File Plan.md
- kb/concepts/decisions/Issue Label Scheme.md
- kb/concepts/decisions/KB Stack Versioning.md
- kb/concepts/decisions/Structural Enforcement over Documented Rule.md
- kb/concepts/patterns/Audit Trail.md
- kb/concepts/patterns/BM25.md
- kb/concepts/patterns/Command Round-Trip Integrity.md
- kb/concepts/patterns/Confidence Scoring.md
- kb/concepts/patterns/Contradiction Resolution.md
- kb/concepts/patterns/Entity Extraction.md
- kb/concepts/patterns/Filter on Ingest.md
- kb/concepts/patterns/Forgetting.md
- kb/concepts/patterns/Graph Traversal.md
- kb/concepts/patterns/Mesh Sync.md
- kb/concepts/patterns/Quality Scoring.md
- kb/concepts/patterns/Reciprocal Rank Fusion.md
- kb/concepts/patterns/Self-Healing.md
- kb/concepts/patterns/Shared vs Private.md
- kb/concepts/patterns/Typed Relationships.md
- kb/concepts/patterns/Vector Search.md
- kb/concepts/patterns/Work Coordination.md
- kb/concepts/problems/Ambient Environment Dependency.md
- kb/concepts/problems/Detect-Repair Asymmetry.md
- kb/concepts/problems/Green Suite Blind Spot.md
- kb/concepts/problems/Naming Convention Conflict.md
- kb/concepts/problems/Write-Once Frontmatter Fields.md
- kb/concepts/protocols/CPPC.md
- kb/concepts/protocols/Modbus.md
- kb/concepts/protocols/SSD TRIM.md
- kb/concepts/workflows/Anti-Cramming Heuristic.md
- kb/concepts/workflows/Bulk Operations.md
- kb/concepts/workflows/CI Integration.md
- kb/concepts/workflows/Checkpoint Audit.md
- kb/concepts/workflows/Claude Code Auto Mode.md
- kb/concepts/workflows/Content Quality Control.md
- kb/concepts/workflows/Crystallization.md
- kb/concepts/workflows/Event-Driven Automation.md
- kb/concepts/workflows/Hooks.md
- kb/concepts/workflows/Index Scaling.md
- kb/concepts/workflows/Iteration and Cost Limits.md
- kb/concepts/workflows/KB Migration.md
- kb/concepts/workflows/Knowledge Compounding.md
- kb/concepts/workflows/Lint Workflow.md
- kb/concepts/workflows/Mass-Update Gate.md
- kb/concepts/workflows/Multi-Agent Collaboration.md
- kb/concepts/workflows/Privacy and Governance.md
- kb/concepts/workflows/Publish-Remote Gate.md
- kb/concepts/workflows/Quality and Self-Correction.md
- kb/concepts/workflows/Semantic Lint Automation.md
- kb/concepts/workflows/Session Orientation.md
- kb/concepts/workflows/Split Merge Reclassify.md
- kb/concepts/workflows/Split Threshold.md
- kb/concepts/workflows/Stub Threshold.md
- kb/concepts/workflows/Supersession.md
- kb/concepts/workflows/User Management.md
- kb/concepts/workflows/Workflow Extraction.md
- kb/concepts/workflows/Workflow Orchestration.md
- kb/index.md
- kb/log.md
- tools/CONTRACT.md
- tools/README.md
- tools/chemenu/catalog.py
- tools/chemenu/commands/index_build.py
- tools/chemenu/lint_core.py
- tools/chemenu/tests/conftest.py
- tools/chemenu/tests/test_cite_cmd.py
- tools/chemenu/tests/test_git_publish.py
- tools/chemenu/tests/test_index_build.py
- tools/chemenu/tests/test_lint.py
- tools/chemenu/tests/test_new_page.py
- tools/chemenu/tests/test_provenance.py
- tools/chemenu/tests/test_type_resolver.py
- tools/chemenu/tests/test_xref.py
- types/concept.md
- types/type-spec.md
This commit is contained in:
2026-09-08 10:07:46 +02:00
parent 63b4bb82d9
commit 7f74303a00
103 changed files with 703 additions and 180 deletions
+77 -1
View File
@@ -35,7 +35,7 @@ dev-checkout concern - readable here, never shipped as something to parse.
---
## 4.8.0-beta.5 - 2026-09-05 - update entity naming conventions to use singular form for consistency
## 4.8.0-beta.6 - 2026-09-08 - kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59)
**Author:** Torben Nehmer
@@ -46,6 +46,7 @@ dev-checkout concern - readable here, never shipped as something to parse.
- raw accept: incoming/ als abgeleiteter Rohablage-Eingang (schliesst #58)
- raw accept: Stem-Eindeutigkeit im Typverzeichnis erzwingen, --replaces als einziger Weg daran vorbei (schliesst #64)
- update entity naming conventions to use singular form for consistency
- kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59)
<!-- /wikitool:bumps -->
Das Label `status/incoming` gibt es seit heute in Gitea: der Mensch legt einen
@@ -301,6 +302,81 @@ keine Kompatibilitätsfrage, solange sie vor `4.8.0` landet.
Schließt #64.
**`kb/concepts/` bekommt Areas** (#59): Sharden ist längst automatisch —
`index_build.SHARD_THRESHOLD = 50`, hergeleitet aus der wikieigenen Seite
`Index Scaling` — aber es passiert **pro Area**, und eine Area legt niemand an.
`kb/concepts/` hatte keine, also war die Schwelle dort ein toter Wert: 80 Seiten
in einer einzigen Tabelle, weit über der eigenen Grenze, ohne dass je etwas
gefeuert hätte. Die Ursache war eine Asymmetrie in den Type-Specs — `entity`
deklarierte ein `layout:`, `concept` nicht, obwohl das Subtype-Feld fertig dalag.
`types/concept.md` deklariert es jetzt für alle sechs `concept_type`-Werte
(`architectures/`, `patterns/`, `protocols/`, `workflows/`, `decisions/`,
`problems/`). Für `source` bewusst **nicht**: 25 von 29 Seiten sind `notes`, die
Aufteilung ergäbe eine Area und vier Splitter, und `kb/sources/` liegt mit 29
Seiten ohnehin unter der Schwelle. Ein Subtype-Feld zu haben ist kein Grund, es
als Achse zu benutzen.
Zwei Dinge im Code, beide Folgen desselben Befunds. `_area_titles()` in
`index_build.py` löste `entity` fest über `find_type_by_name("entity")` auf und
las nur dessen `layout:` — jeder zweite Typ mit einem `layout:` hätte
`.title()`-Namen auf dem Verzeichnisnamen bekommen statt der deklarierten Titel.
Es liest jetzt jedes Type-Spec, und zwar **pro Collection** geschlüsselt, damit
zwei Typen denselben Area-Namen für Verschiedenes benutzen dürfen. Und der neue
`lint`-Befund meldet eine Collection über der Schwelle **ohne** Areas, mit der
Verteilung ihres Subtype-Felds — als Empfehlung, nicht als Failure, und nur
dann, wenn die Aufteilung jede entstehende Area unter die Schwelle drückt. Das
begrenzt sich selbst in beide Richtungen: `kb/comparisons/` mit einer Seite
feuert nie, und die schlechte Aufteilung nach `source_type` unterbleibt von
allein, ohne dass der Check etwas über Sources wüsste.
Zwei Dinge fielen unterwegs an, die das Issue nicht vorhergesehen hatte.
`lint_core.py` durfte `SHARD_THRESHOLD`/`group_pages` nicht aus
`commands/index_build.py` importieren — `test_api.py` prüft strukturell, dass
`chemenu.api` kein Modul unter `chemenu.commands` lädt, und der Import hätte den
ganzen CLI-Kopf mitgezogen. Die Gruppierung liegt deshalb neu in
`tools/chemenu/catalog.py`, entlang derselben Linie wie `lint_core.py`:
Korpusform hier, Darstellung dort. Und `_anchor()` strich mit `[^a-z0-9\s-]`
jeden Nicht-ASCII-Buchstaben ersatzlos — die Karte verlinkte auf `#ablufe`,
während die Überschrift im Shard `#abläufe` heißt. Vorher fiel das keinem auf,
weil alle Entity-Area-Titel zufällig ASCII sind; `Abläufe` ist der erste, der es
nicht ist.
Auf dieser Instanz angewendet: `wikitool move --reconcile` hat alle 80
Concept-Seiten in ihre Area gezogen, `migrate verify --from HEAD` bestätigt
`182 compared, 0 added, 0 removed, 80 moved, 0 findings` — kein Titel, kein
Body, kein Frontmatter-Feld angefasst. `index rebuild` erzeugt sechs Areas
(Abläufe 28, Architekturen 20, Muster 17, Entscheidungen 7, Problemstellungen 5,
Protokolle 3); keine über der Schwelle, also kein eigener Shard, und die
Schwelle wirkt wieder als Schwelle.
Geändert: `types/concept.md` (`layout:`), `types/type-spec.md` (wann ein
`layout:` sich lohnt), `tools/chemenu/catalog.py` (neu),
`tools/chemenu/commands/index_build.py` (`area_titles`, `_anchor`),
`tools/chemenu/lint_core.py` (`unsharded_collections`),
`tools/chemenu/tests/` (Fixture-Concept liegt jetzt in seiner Area, plus neun
neue Tests), `tools/CONTRACT.md`, `tools/README.md`, `README.md`,
`kb/concepts/COLLECTION.md`, sowie die 80 bewegten Seiten unter `kb/concepts/`.
**MINOR**, nicht MAJOR: der Umzug ist ein **Angebot**, kein Zwang. Eine
bestehende Instanz, die `move --reconcile` nicht laufen lässt, bleibt
funktionsfähig — `group_pages` liest das Dateisystem, nicht das `layout:`, also
landen flache Bestandsseiten in der Area „All" und neu angelegte in ihrer
eigenen; beides rendert. Der gemischte Zustand meldet sich als `lint`-Befund
*Misplaced Pages*, der seit jeher advisory ist. Und ein Downgrade auf einen
Stack ohne dieses `layout:` funktioniert weiter: die Verzeichnisse bleiben
Verzeichnisse, nur die Anzeigetitel fallen auf `.title()` zurück. Kosmetik, kein
Bruch der Austauschbarkeit in beiden Richtungen.
Schließt #59.
**Nachzug an `63b4bb8`:** die Umstellung der Namenskonvention auf `HA Integration`
hatte in `README.md` das Gegenbeispiel verloren — die Zeile las
``Use singular for entities: `HA Integration.md` (not `HA Integration.md`)``, beide
Seiten des „not" identisch, also eine Regel ohne Fall, an dem sie greift.
`kb/CONVENTIONS.md` und `kb/entities/COLLECTION.md` hatten im selben Commit das
korrekte Paar bekommen; `README.md` zieht jetzt mit `HA Integrations.md` nach.
---
## 4.7.4 - 2026-09-04 - bootstrap.md nennt den session-id-WARN nach frischem Bootstrap explizit als erwartet
+19 -4
View File
@@ -105,8 +105,14 @@ chemenu/
│ │ ├── tools/ # own INDEX.md once past 50 pages
│ │ ├── technologies/
│ │ └── people/
│ ├── concepts/ # COLLECTION.md - architectures, patterns, protocols
├── sources/ # COLLECTION.md - source summaries
│ ├── concepts/ # COLLECTION.md + INDEX.md + areas below
│ ├── architectures/
│ │ ├── patterns/
│ │ ├── protocols/
│ │ ├── workflows/
│ │ ├── decisions/
│ │ └── problems/
│ ├── sources/ # COLLECTION.md - source summaries, no areas by choice
│ └── comparisons/ # COLLECTION.md - comparison pages
├── work/ # WORKSHOP: one directory per multi-session run, tracked
│ └── CONTRACT.md # Run keys, required files, how a run closes
@@ -130,7 +136,16 @@ A directory under `kb/` is a **collection** exactly when it holds a `COLLECTION.
subdirectory inside one is an **area** that inherits it, and that is as deep as a page goes -
nothing nests below an area, because the generated catalog reads exactly two path segments
under `kb/` and would fold a deeper page into the area silently (`kb/CONTRACT.md` § Collections
has the rule; `wikitool lint` reports a violation as a hard error). `COLLECTION.md` never
has the rule; `wikitool lint` reports a violation as a hard error).
Which areas a collection has is not chosen per page: a type-spec's `layout:` maps its subtype
field onto directories, and `wikitool new` writes the page straight into the one its subtype
names. That is also what makes the catalog's shard threshold do anything - `index rebuild`
splits **per area**, so a collection with no areas keeps one table however large it grows.
`wikitool lint` reports such a collection once it is past the threshold, as a recommendation
rather than an error, together with the split its subtype field would produce; it stays quiet
when the split would not actually help. `kb/sources/` is the worked example of the second case
and deliberately has no areas. `COLLECTION.md` never
appears outside `kb/` - the other layers carry a `CONTRACT.md` or a root type-spec instead. A stage may carry
both a `README.md` and a `CONTRACT.md`: they have different readers. The README is for humans
working *on* that layer, the contract is what binds an agent working *with* it.
@@ -247,7 +262,7 @@ Ingest incoming/notes/my-notes.md
### Naming
- Use human-readable titles with spaces for files: `Hybrid Search.md`, not kebab-case
- Use singular for entities: `HA Integration.md` (not `HA Integration.md`)
- Use singular for entities: `HA Integration.md` (not `HA Integrations.md`)
- Use wikilinks matching the file name exactly: `[[Entity Name]]`
- **Titles follow the subject's own established name, not the wiki's language.** `Act Runner` and
`GitOps Ownership Model` keep theirs. A title is the only identifier a page has - it also lives
+1 -1
View File
@@ -1 +1 @@
4.8.0-beta.5
4.8.0-beta.6
+26 -1
View File
@@ -25,7 +25,32 @@ tone, relationship labels, the confidence rubric. Neither is restated here.
## Types offered
`concept` (`tools/wikitool types describe concept`).
`concept` (`tools/wikitool types describe concept`). Das Feld `concept_type:`
wählt die Area:
| Area | Hält |
|------|------|
| `architectures/` | Aufbau und Struktur: wie ein System geschnitten ist und warum die Schnitte dort liegen |
| `patterns/` | Wiederverwendbare Lösungsformen, die über mehr als einen Gegenstand hinweg gelten |
| `protocols/` | Kommunikationsprotokolle und Standards, in ihrer üblichen Schreibweise benannt |
| `workflows/` | Abläufe und Prozesse, die projektübergreifend wiederkehren |
| `decisions/` | Architektur- und Entwurfsentscheidungen (siehe unten) |
| `problems/` | Wiederkehrende Problemstellungen und ihre Lösungsansätze |
Das sind Areas, keine Collections: sie erben diesen Contract und tragen keine
eigene `COLLECTION.md`.
Die Zuordnung trifft niemand von Hand — sie steht als `layout:` in
`types/concept.md`, und `wikitool new` legt eine neue Seite direkt dort ab.
Eine Seite, die anderswo liegt, meldet `wikitool lint` als *misplaced*;
`wikitool move --page "<Titel>"` bringt sie an ihren berechneten Ort.
Die Aufteilung ist keine Geschmacksfrage, sondern das, was die Shard-Schwelle
des Katalogs überhaupt wirksam macht: `index rebuild` teilt **pro Area**, und
eine Collection ohne Areas teilt sich nie — mit 80 Seiten in einer einzigen
Tabelle war die Schwelle hier ein toter Wert (Gitea #59). Keine der sechs
Areas liegt derzeit über der Schwelle, also bekommt auch keine einen eigenen
Shard; wächst eine hinein, passiert das ohne Zutun.
## Decisions
+76 -51
View File
@@ -4,88 +4,113 @@
80 page(s). Regenerated by `wikitool index rebuild`.
## All
## Abläufe
| Page | Type | Summary | Last Modified |
|------|------|---------|----------------|
| [[Ambient Environment Dependency]] | problem | Fehlerklasse, in der ein Test gruen ist, weil die Maschine zufaellig passt statt weil der Code stimmt - abgegrenzt gegen den Green Suite Blind Spot, belegt an vier Faellen unter Gitea-Issue #8 | 2026-08-31 |
| [[Anti-Cramming Heuristic]] | workflow | Regel gegen überladene Seiten: ab dem dritten Absatz zu einem Unterthema eine eigene Seite anlegen | 2026-08-29 |
| [[Audit Trail]] | pattern | Unveränderliches chronologisches Log aller Wiki-Operationen (Ingest, Bearbeitung, Löschung, Abfrage) mit Zeitstempel, Akteur, Ziel und Änderungsbeschreibung. | 2026-08-29 |
| [[BM25]] | pattern | Schlüsselwortbasiertes Retrieval-Verfahren, das über Termfrequenz, inverse Dokumentfrequenz und Stemming exakte oder teilweise Übereinstimmungen findet. | 2026-08-29 |
| [[Bulk Operations]] | workflow | Umkehrbare, protokollierte Operationen zum Massenlöschen, Exportieren, Zusammenführen oder Archivieren von Wiki-Inhalten, mit Freigabepflicht und Undo. | 2026-08-29 |
| [[Checkpoint Audit]] | workflow | Regelmäßiger Qualitätsrhythmus: Index und Backlinks alle 15 Einträge neu aufbauen, auf 0 neue Artikel prüfen, die 3 meistgeänderten erneut lesen | 2026-08-29 |
| [[CI Integration]] | workflow | CI/CD-Hooks vor dem Publish: ci.yml (Push/PR, Stack-Pfade, seit 1.8.1 mit Coverage-Messung ohne Schwelle) und nightly.yml (Zeitplan, schliesst die paths-ignore-Luecke fuer Content-Drift; schedule-Ausloesung seit 2026-09-01 bestaetigt) setzen Quality Gates durch | 2026-09-01 |
| [[Claude Code Auto Mode]] | workflow | auto-Berechtigungsmodus von Claude Code: ein Klassifikator genehmigt Aktionen vor der Ausfuehrung statt nachzufragen; die Beschreibung stammt weit ueberwiegend aus zweiter Hand ueber einen Doku-Subagenten | 2026-08-31 |
| [[Command Round-Trip Integrity]] | pattern | Anforderung, dass zwei Befehle auf derselben Datei in jeder Reihenfolge zusammenpassen und jeder erzeugte Zustand einen Gegenbefehl hat - 2026-08-31 in wikitool zweimal verletzt | 2026-08-31 |
| [[Confidence Scoring]] | pattern | Mechanismus, der faktischen Aussagen quantitative Werte nach Quellenzahl, Aktualität, Qualität und Bestätigung zuweist, um gut gestütztes Wissen zu erkennen. | 2026-08-29 |
| [[Consolidation Tiers]] | architecture | Hierarchische Speicherarchitektur, die Informationen durch zunehmend verdichtete Schichten vom Working Memory bis zum Semantic und Procedural Memory befördert. | 2026-08-29 |
| [[Content Quality Control]] | workflow | Regeln und Schwellenwerte für die Seitenqualität: Mindestumfang für Stubs, Aufteilungsschwellen und Zielwerte für die Zeilenzahl | 2026-08-29 |
| [[Context Isolation]] | architecture | Grundsatz, für jede Aufgabe nur den jeweils benötigten Kontext zu laden | 2026-08-29 |
| [[Contradiction Resolution]] | pattern | Automatisches Erkennen und Auflösen widersprüchlicher Aussagen anhand von Konfidenz, Aktualität und Autorität der Quelle. | 2026-08-29 |
| [[CPPC]] | protocol | Hardwareschnittstelle Collaborative Processor Performance Control für feingranulares CPU-Power-Management zwischen Betriebssystem und AMD-Prozessor. | 2026-08-29 |
| [[Cross-platform Agent Skills]] | architecture | Architektur fuer Agent-Skills, die ueber mehrere LLM-Werkzeuge hinweg funktionieren; in Chemenu selbst am 2026-08-04 umgesetzt und ueberprueft | 2026-09-01 |
| [[Crystallization]] | workflow | Verdichten abgeschlossener Erkundungen, Debugging-Sitzungen und Recherchen zu strukturierten Wiki-Auszügen als eigenständige Wissensquellen. | 2026-08-29 |
| [[Delete Rather Than Anonymize]] | decision | Private Korpusinhalte per Loeschung entfernen statt zu anonymisieren: ein Seitentitel ist der einzige Identifier eines Wikis, Umbenennen ist die volle page-lifecycle-Prozedur je Seite, Loeschen ist ein unterstuetztes Kommando. | 2026-09-01 |
| [[Denylist over Allowlist]] | decision | Entscheidung, schreibbare Felder als Schema minus kurzer Sperrliste zu bestimmen statt als gepflegte Positivliste, weil die Positivliste eine zweite Kopie des Schemas waere | 2026-08-31 |
| [[Detect-Repair Asymmetry]] | problem | Werkzeugluecke, in der ein Check einen Defekt zuverlaessig meldet, aber kein Befehl ihn behebt - womit die Handeditierung der einzige verbleibende Ausweg ist | 2026-08-31 |
| [[Diff-Reviewable Agent Edits]] | decision | Entscheidung, Dateiaenderungen ueber Edit/Write statt ueber Shell-Heredocs zu fahren, weil nur das erste eine pruefbare Diff hinterlaesst | 2026-08-31 |
| [[Dual Licensing by File Plan]] | decision | Ein Repo mit Code- und Inhaltsanteil erhaelt zwei Lizenzen; die Grenze zwischen ihnen ist kein zweiter, gepflegter Pfadkatalog, sondern der ohnehin vorhandene Dateiplan des Distributionswerkzeugs. | 2026-09-01 |
| [[Entity Extraction]] | pattern | Erkennen und Strukturieren von Entities (Personen, Projekte, Bibliotheken, Concepts, Dateien, Entscheidungen, Systeme, Werkzeuge) samt typspezifischer Attribute aus Rohquellen. | 2026-08-29 |
| [[Episodic Memory]] | architecture | Speicherschicht für verdichtete Sitzungszusammenfassungen und Befunde; Brücke zwischen rohem Working Memory und langlebigem Semantic Memory. | 2026-08-29 |
| [[Event-Driven Automation]] | workflow | Muster, das automatische Auslöser an Wiki-Lebenszyklusereignisse hängt, um manuellen Pflegeaufwand und das Risiko der Verwahrlosung zu senken. | 2026-08-29 |
| [[Filter on Ingest]] | pattern | Automatisches Erkennen und Entfernen sensibler Daten (API-Schlüssel, Token, Credentials, personenbezogene Daten) vor der Aufnahme ins Wiki, per Regex und ML-Erkennung. | 2026-08-29 |
| [[Forgetting]] | pattern | Muster zur Wissensbindung, das selten abgerufene Fakten schrittweise zurückstuft, modelliert nach der Ebbinghausschen Vergessenskurve. | 2026-08-29 |
| [[Graph Traversal]] | pattern | Verfahren, verbundene Entities im Wissensgraphen über typisierte Beziehungen (uses, depends-on, contradicts, caused) zu finden und strukturelle Fragen zu beantworten. | 2026-08-29 |
| [[Green Suite Blind Spot]] | problem | Defekt, der eine vollstaendig gruene Testsuite ueberlebt, weil nie ein Test das richtige Verhalten behauptet hat - belegt an drei prio/1-2-Defekten (Round-Trip, Zitat-Notation-als-Code, Zitat-Limit) | 2026-08-31 |
| [[Hooks]] | workflow | Mechanismus von Event-Listenern, der bei Wiki-Lebenszyklusereignissen wie Quellen-Ingest, Seitenänderung und Sitzungsende automatisch Aktionen auslöst. | 2026-08-29 |
| [[Hybrid Search]] | architecture | Multimodale Suche, die BM25-Schlüsselwortabgleich, Vektor-Embeddings und Graph Traversal verbindet, um Wissensabruf im Wiki skalierbar zu machen. | 2026-08-29 |
| [[Implementation Spectrum]] | architecture | Modularer Einführungspfad für die Funktionen von LLM Wiki v2, vom minimal tragfähigen Wiki bis zur vollen Umsetzung mit Automatisierung und Governance. | 2026-08-29 |
| [[Index Scaling]] | workflow | Skalierungsregeln für Indexseiten: Tabellenabschnitte ab 50 Einträgen teilen, ab 200 Seiten _meta/topic-map.md anlegen | 2026-08-29 |
| [[Issue Label Scheme]] | decision | Pflicht-Labelschema fuer das Gitea-Board: vier Achsen (area/kind/prio/size) plus seit 2026-09-04 drei optionale status/-Flags, darunter status/incoming fuer unausgearbeitete Stubs, die die Vier-Achsen-Pflicht aussetzen statt sie zu ergaenzen; die Regel liegt in instructions/dev/, weil sie keine ausgelieferte Instanz erreichen darf | 2026-09-04 |
| [[Iteration and Cost Limits]] | workflow | Im Code durchgesetzte Obergrenze von 60 wikitool-Aufrufen je Session, Loop-Breaker bei 3 identischen Wiederholungen, Slot-Erstattung, ein gemessenes Kalibrierungsband, und Retrieval sowie der MCP-Leseserver bleiben ausgenommen | 2026-09-02 |
| [[KB Migration]] | workflow | Migration des KB-Inhalts entlang einer geordneten Versionskette; abgegrenzt gegen offene Instanz-Aktionen, die in den doctor-Check gehoeren statt in die Kette | 2026-08-31 |
| [[KB Stack Versioning]] | decision | Semantische Versionierung des Wiki-Stacks: VERSION beschreibt die Maschinerie, Kompatibilitaet (Drop-in-Ersatz) und Inhaltsmigration sind seit 2.5.0 getrennte, unabhaengig geprueft Fragen | 2026-09-02 |
| [[Knowledge Compounding]] | workflow | Effekt, bei dem Wissen im Wiki an Wert gewinnt, weil jede neue Quelle an bestehende, untereinander verwiesene Seiten anknüpft und sie ergänzt. | 2026-08-29 |
| [[Knowledge Graph]] | architecture | Typisierte Schicht aus Entities und Beziehungen über den Wiki-Seiten, die eine reichere Wissensdarstellung und graphbasierte Abfragen ermöglicht. | 2026-08-29 |
| [[Lint Workflow]] | workflow | Deterministischer Health-Check rund um wikitool lint; seit 1.7.2 maskiert es Code vor dem Notation-Match und zaehlt Zitat-Bloecke statt Zeilen | 2026-09-01 |
| [[LLM Wiki Pattern]] | architecture | Methodik für persönliches Wissensmanagement, bei der ein LLM aus Rohquellen ein dauerhaftes Wiki aufbaut - Wissen wird kompiliert statt per RAG neu hergeleitet. | 2026-08-29 |
| [[Mass-Update Gate]] | workflow | Mass-Update Gate: publish endet mit 42 (Freigabe durch den Menschen noetig) ab 10 gezaehlten Dateien; generierte Dateien und work/ werden committet\, aber seit 1.5.0 nicht gezaehlt; freigegeben per --confirm <token> | 2026-09-01 |
| [[Multi-Agent Collaboration]] | workflow | Wissensmanagement mit mehreren Agenten; erweitert das LLM-Wiki-Muster um Mesh Sync, die Trennung von geteiltem und privatem Wissen und leichtgewichtige Arbeitskoordination. | 2026-08-29 |
| [[Privacy and Governance]] | workflow | Rahmenwerk zur Absicherung von Wiki-Inhalten über Datenfilterung beim Ingest, Audit-Trail-Protokollierung und umkehrbare Massenoperationen. | 2026-08-29 |
| [[Publish-Remote Gate]] | workflow | Drittes, im Code durchgesetztes Gate: publish bricht mit Exit 42 ab, wenn die aufgeloeste Push-URL nicht in einer optionalen, gitignoreten Allowlist steht; doctor benennt seit 2026-09-02 den Gate-Zustand statt nur die Dateiexistenz | 2026-09-02 |
| [[Quality and Self-Correction]] | workflow | Automatische Qualitätssicherung für Wikis mit Inhaltsbewertung, Selbstheilung und Widerspruchserkennung. | 2026-08-29 |
| [[Semantic Lint Automation]] | workflow | Maschinelle Heuristiken zur Priorisierung der semantischen Prüfung: veraltete Aussagen, hohe Änderungsdichte und schwache Verlinkung | 2026-08-29 |
| [[Session Orientation]] | workflow | Verbindliche Vorabprüfung, die vor Query- und Update-Operationen einen Kontextbericht erzeugt (Index, jüngste Logs, Umfang) | 2026-08-29 |
| [[Split Merge Reclassify]] | workflow | Eigene Befehle zum Teilen, Zusammenführen und Umklassifizieren von Seiten, mit automatischer Korrektur von Links und Frontmatter | 2026-08-29 |
| [[Split Threshold]] | workflow | Maximale Seitengröße, ab der eine Aufteilung empfohlen wird (Farza: >120-150 Zeilen, Pascalandy: 200 Zeilen) | 2026-08-29 |
| [[Stub Threshold]] | workflow | Mindestumfang, ab dem eine Wiki-Seite nicht mehr als Stub gilt (Farza: ≥3 Sätze oder 15 Zeilen) | 2026-08-29 |
| [[Supersession]] | workflow | Ablösen alten Wissens durch neue, widersprechende Information; gibt dem Wiki eine Versionierung mit ausdrücklicher Verknüpfung und Erhalt der Historie. | 2026-08-29 |
| [[User Management]] | workflow | Linux-Ablauf zum Anlegen, Ändern, Überwachen und Löschen von Benutzerkonten mit useradd, usermod und userdel, samt Gruppenverwaltung und sudoers-Konfiguration. | 2026-08-29 |
| [[Workflow Extraction]] | workflow | Herauslösen von Workflow-Abschnitten aus monolithischer Dokumentation | 2026-09-01 |
| [[Workflow Orchestration]] | workflow | Orchestrierte Einzelbefehle für vollständige Operationen (ingest run, lint run, update run) mit Dry-Run-Vorschau vor dem Schreiben | 2026-08-29 |
## Architekturen
| Page | Type | Summary | Last Modified |
|------|------|---------|----------------|
| [[Consolidation Tiers]] | architecture | Hierarchische Speicherarchitektur, die Informationen durch zunehmend verdichtete Schichten vom Working Memory bis zum Semantic und Procedural Memory befördert. | 2026-08-29 |
| [[Context Isolation]] | architecture | Grundsatz, für jede Aufgabe nur den jeweils benötigten Kontext zu laden | 2026-08-29 |
| [[Cross-platform Agent Skills]] | architecture | Architektur fuer Agent-Skills, die ueber mehrere LLM-Werkzeuge hinweg funktionieren; in Chemenu selbst am 2026-08-04 umgesetzt und ueberprueft | 2026-09-01 |
| [[Episodic Memory]] | architecture | Speicherschicht für verdichtete Sitzungszusammenfassungen und Befunde; Brücke zwischen rohem Working Memory und langlebigem Semantic Memory. | 2026-08-29 |
| [[Hybrid Search]] | architecture | Multimodale Suche, die BM25-Schlüsselwortabgleich, Vektor-Embeddings und Graph Traversal verbindet, um Wissensabruf im Wiki skalierbar zu machen. | 2026-08-29 |
| [[Implementation Spectrum]] | architecture | Modularer Einführungspfad für die Funktionen von LLM Wiki v2, vom minimal tragfähigen Wiki bis zur vollen Umsetzung mit Automatisierung und Governance. | 2026-08-29 |
| [[Knowledge Graph]] | architecture | Typisierte Schicht aus Entities und Beziehungen über den Wiki-Seiten, die eine reichere Wissensdarstellung und graphbasierte Abfragen ermöglicht. | 2026-08-29 |
| [[LLM Wiki Pattern]] | architecture | Methodik für persönliches Wissensmanagement, bei der ein LLM aus Rohquellen ein dauerhaftes Wiki aufbaut - Wissen wird kompiliert statt per RAG neu hergeleitet. | 2026-08-29 |
| [[MCP-Leseserver]] | architecture | Zweiter Konsument von chemenu ueber MCP: search/types/describe_type/lint/status auf chemenu.api.Corpus, strukturell ohne Schreibpfad, jede Antwort trage einen Commit-Stempel. | 2026-09-02 |
| [[Memory Lifecycle]] | architecture | Architektur des Wissenslebenszyklus mit Confidence Scoring, Supersession, Forgetting und Consolidation Tiers zur Pflege von Fakten über die Zeit. | 2026-08-29 |
| [[Mesh Sync]] | pattern | Abgleichsmechanismus, der Beobachtungen paralleler Agenten in ein gemeinsames Wiki überführt; Last-Write-Wins mit Konflikterkennung und manuellem Eingriff. | 2026-08-29 |
| [[Modbus]] | protocol | Industrielles Kommunikationsprotokoll von 1979 zur Anbindung speicherprogrammierbarer Steuerungen und Geräte über serielle oder TCP-Netze. | 2026-08-29 |
| [[Multi-Agent Collaboration]] | workflow | Wissensmanagement mit mehreren Agenten; erweitert das LLM-Wiki-Muster um Mesh Sync, die Trennung von geteiltem und privatem Wissen und leichtgewichtige Arbeitskoordination. | 2026-08-29 |
| [[Naming Convention Conflict]] | problem | Widerspruch zwischen README.md (kebab-case) und AGENTS.md (lesbar mit Leerzeichen), der zu Drift bei der Validierung führt | 2026-08-29 |
| [[OKF Compatibility]] | architecture | Optionale Kompatibilität zum Open Knowledge Framework als Export- und Prüfmodus, ohne das interne Modell zu ersetzen | 2026-08-29 |
| [[Optional Instance Context File]] | architecture | Muster fuer eine Datei, die eine Instanz ueber ihre Umgebung informiert, ohne Betriebsvoraussetzung zu sein: Health-Check meldet ohne zu scheitern, pro Checkout statt pro Repo | 2026-08-31 |
| [[Personalization Plane]] | architecture | Schicht fuer Instanz-Identitaet: USER.md/SOUL.md werden als Template ausgeliefert, im Setup-Interview woertlich befuellt und vom doctor-Check auf fehlend wie unbefuellt geprueft | 2026-08-31 |
| [[Privacy and Governance]] | workflow | Rahmenwerk zur Absicherung von Wiki-Inhalten über Datenfilterung beim Ingest, Audit-Trail-Protokollierung und umkehrbare Massenoperationen. | 2026-08-29 |
| [[Procedural Memory]] | architecture | Langlebigste Speicherschicht für Abläufe, Muster, bewährte Vorgehensweisen und Rezepte, gewonnen aus wiederholten semantischen Beobachtungen. | 2026-08-29 |
| [[Publish-Remote Gate]] | workflow | Drittes, im Code durchgesetztes Gate: publish bricht mit Exit 42 ab, wenn die aufgeloeste Push-URL nicht in einer optionalen, gitignoreten Allowlist steht; doctor benennt seit 2026-09-02 den Gate-Zustand statt nur die Dateiexistenz | 2026-09-02 |
| [[Quality and Self-Correction]] | workflow | Automatische Qualitätssicherung für Wikis mit Inhaltsbewertung, Selbstheilung und Widerspruchserkennung. | 2026-08-29 |
| [[Quality Scoring]] | pattern | Quantitative Bewertung aller vom LLM geschriebenen Inhalte nach struktureller Qualität, Vollständigkeit der Quellenangaben, Konsistenz mit dem Wiki und Themenabdeckung. | 2026-08-29 |
| [[RAG]] | architecture | Architekturmuster, bei dem LLMs die Generierung um Dokumente aus einer Wissensbasis anreichern. | 2026-08-29 |
| [[Reciprocal Rank Fusion]] | pattern | Verfahren, das Ergebnislisten mehrerer Suchmodalitäten zu einem gemeinsamen Ranking verbindet, ohne Gewichte zwischen den Modalitäten justieren zu müssen. | 2026-08-29 |
| [[Scale Ceiling]] | architecture | Punkt, ab dem Wiki-Ansätze mit einem einzigen Kontext qualitativ abfallen | 2026-09-01 |
| [[Self-Healing]] | pattern | Automatisches Beheben von Mängeln, die beim Lint auffallen: verwaiste Seiten, veraltete Aussagen, kaputte Links und Formatverstöße. | 2026-08-29 |
| [[Semantic Lint Automation]] | workflow | Maschinelle Heuristiken zur Priorisierung der semantischen Prüfung: veraltete Aussagen, hohe Änderungsdichte und schwache Verlinkung | 2026-08-29 |
| [[Semantic Memory]] | architecture | Langlebige Schicht für sitzungsübergreifend verdichtete Fakten aus mehreren Episoden, mit höherer Konfidenz und stärkerer Verdichtung als episodische Erinnerungen. | 2026-08-29 |
| [[Session Orientation]] | workflow | Verbindliche Vorabprüfung, die vor Query- und Update-Operationen einen Kontextbericht erzeugt (Index, jüngste Logs, Umfang) | 2026-08-29 |
| [[Shared vs Private]] | pattern | Abgrenzung persönlicher Beobachtungen (privat) von Team- und Projektwissen (geteilt), mit Regeln zum Hochstufen geprüften Wissens. | 2026-08-29 |
| [[Split Merge Reclassify]] | workflow | Eigene Befehle zum Teilen, Zusammenführen und Umklassifizieren von Seiten, mit automatischer Korrektur von Links und Frontmatter | 2026-08-29 |
| [[Split Threshold]] | workflow | Maximale Seitengröße, ab der eine Aufteilung empfohlen wird (Farza: >120-150 Zeilen, Pascalandy: 200 Zeilen) | 2026-08-29 |
| [[SSD TRIM]] | protocol | Datenträgerbefehl, mit dem SSDs ungenutzte Blöcke zurückgewinnen - für gleichbleibende Leistung und längere Lebensdauer. | 2026-08-29 |
| [[Structural Enforcement over Documented Rule]] | decision | Entscheidung, eine wiederkehrende Fehlerregel in die Ausfuehrung einzubauen statt sie aufzuschreiben - Rangfolge erzwingen vor melden vor erinnern, belegt an einer Regel, die gelesen wurde und nicht wirkte | 2026-08-31 |
| [[Stub Threshold]] | workflow | Mindestumfang, ab dem eine Wiki-Seite nicht mehr als Stub gilt (Farza: ≥3 Sätze oder 15 Zeilen) | 2026-08-29 |
| [[Supersession]] | workflow | Ablösen alten Wissens durch neue, widersprechende Information; gibt dem Wiki eine Versionierung mit ausdrücklicher Verknüpfung und Erhalt der Historie. | 2026-08-29 |
| [[Three-Layer Architecture]] | architecture | Strukturmodell des LLM-Wiki-Musters mit drei Schichten: unveränderliche Rohquellen, vom LLM gepflegtes Wiki und Schemakonfiguration, die Knowledge Compounding trägt. | 2026-08-29 |
| [[Token Economics]] | architecture | Kosten- und Effizienzüberlegungen zum Tokenverbrauch von LLMs - die genannten Werte 5-8x/61%/100% sind unbestätigt, keine gesicherten Fakten | 2026-09-01 |
| [[Working Memory]] | architecture | Kurzlebige Speicherschicht für jüngste Beobachtungen und vorläufige Befunde vor der Verdichtung; niedrigste Konfidenz, keine Verdichtung, wird zum Sitzungsende erneuert. | 2026-08-29 |
## Entscheidungen
| Page | Type | Summary | Last Modified |
|------|------|---------|----------------|
| [[Delete Rather Than Anonymize]] | decision | Private Korpusinhalte per Loeschung entfernen statt zu anonymisieren: ein Seitentitel ist der einzige Identifier eines Wikis, Umbenennen ist die volle page-lifecycle-Prozedur je Seite, Loeschen ist ein unterstuetztes Kommando. | 2026-09-01 |
| [[Denylist over Allowlist]] | decision | Entscheidung, schreibbare Felder als Schema minus kurzer Sperrliste zu bestimmen statt als gepflegte Positivliste, weil die Positivliste eine zweite Kopie des Schemas waere | 2026-08-31 |
| [[Diff-Reviewable Agent Edits]] | decision | Entscheidung, Dateiaenderungen ueber Edit/Write statt ueber Shell-Heredocs zu fahren, weil nur das erste eine pruefbare Diff hinterlaesst | 2026-08-31 |
| [[Dual Licensing by File Plan]] | decision | Ein Repo mit Code- und Inhaltsanteil erhaelt zwei Lizenzen; die Grenze zwischen ihnen ist kein zweiter, gepflegter Pfadkatalog, sondern der ohnehin vorhandene Dateiplan des Distributionswerkzeugs. | 2026-09-01 |
| [[Issue Label Scheme]] | decision | Pflicht-Labelschema fuer das Gitea-Board: vier Achsen (area/kind/prio/size) plus seit 2026-09-04 drei optionale status/-Flags, darunter status/incoming fuer unausgearbeitete Stubs, die die Vier-Achsen-Pflicht aussetzen statt sie zu ergaenzen; die Regel liegt in instructions/dev/, weil sie keine ausgelieferte Instanz erreichen darf | 2026-09-04 |
| [[KB Stack Versioning]] | decision | Semantische Versionierung des Wiki-Stacks: VERSION beschreibt die Maschinerie, Kompatibilitaet (Drop-in-Ersatz) und Inhaltsmigration sind seit 2.5.0 getrennte, unabhaengig geprueft Fragen | 2026-09-02 |
| [[Structural Enforcement over Documented Rule]] | decision | Entscheidung, eine wiederkehrende Fehlerregel in die Ausfuehrung einzubauen statt sie aufzuschreiben - Rangfolge erzwingen vor melden vor erinnern, belegt an einer Regel, die gelesen wurde und nicht wirkte | 2026-08-31 |
## Muster
| Page | Type | Summary | Last Modified |
|------|------|---------|----------------|
| [[Audit Trail]] | pattern | Unveränderliches chronologisches Log aller Wiki-Operationen (Ingest, Bearbeitung, Löschung, Abfrage) mit Zeitstempel, Akteur, Ziel und Änderungsbeschreibung. | 2026-08-29 |
| [[BM25]] | pattern | Schlüsselwortbasiertes Retrieval-Verfahren, das über Termfrequenz, inverse Dokumentfrequenz und Stemming exakte oder teilweise Übereinstimmungen findet. | 2026-08-29 |
| [[Command Round-Trip Integrity]] | pattern | Anforderung, dass zwei Befehle auf derselben Datei in jeder Reihenfolge zusammenpassen und jeder erzeugte Zustand einen Gegenbefehl hat - 2026-08-31 in wikitool zweimal verletzt | 2026-08-31 |
| [[Confidence Scoring]] | pattern | Mechanismus, der faktischen Aussagen quantitative Werte nach Quellenzahl, Aktualität, Qualität und Bestätigung zuweist, um gut gestütztes Wissen zu erkennen. | 2026-08-29 |
| [[Contradiction Resolution]] | pattern | Automatisches Erkennen und Auflösen widersprüchlicher Aussagen anhand von Konfidenz, Aktualität und Autorität der Quelle. | 2026-08-29 |
| [[Entity Extraction]] | pattern | Erkennen und Strukturieren von Entities (Personen, Projekte, Bibliotheken, Concepts, Dateien, Entscheidungen, Systeme, Werkzeuge) samt typspezifischer Attribute aus Rohquellen. | 2026-08-29 |
| [[Filter on Ingest]] | pattern | Automatisches Erkennen und Entfernen sensibler Daten (API-Schlüssel, Token, Credentials, personenbezogene Daten) vor der Aufnahme ins Wiki, per Regex und ML-Erkennung. | 2026-08-29 |
| [[Forgetting]] | pattern | Muster zur Wissensbindung, das selten abgerufene Fakten schrittweise zurückstuft, modelliert nach der Ebbinghausschen Vergessenskurve. | 2026-08-29 |
| [[Graph Traversal]] | pattern | Verfahren, verbundene Entities im Wissensgraphen über typisierte Beziehungen (uses, depends-on, contradicts, caused) zu finden und strukturelle Fragen zu beantworten. | 2026-08-29 |
| [[Mesh Sync]] | pattern | Abgleichsmechanismus, der Beobachtungen paralleler Agenten in ein gemeinsames Wiki überführt; Last-Write-Wins mit Konflikterkennung und manuellem Eingriff. | 2026-08-29 |
| [[Quality Scoring]] | pattern | Quantitative Bewertung aller vom LLM geschriebenen Inhalte nach struktureller Qualität, Vollständigkeit der Quellenangaben, Konsistenz mit dem Wiki und Themenabdeckung. | 2026-08-29 |
| [[Reciprocal Rank Fusion]] | pattern | Verfahren, das Ergebnislisten mehrerer Suchmodalitäten zu einem gemeinsamen Ranking verbindet, ohne Gewichte zwischen den Modalitäten justieren zu müssen. | 2026-08-29 |
| [[Self-Healing]] | pattern | Automatisches Beheben von Mängeln, die beim Lint auffallen: verwaiste Seiten, veraltete Aussagen, kaputte Links und Formatverstöße. | 2026-08-29 |
| [[Shared vs Private]] | pattern | Abgrenzung persönlicher Beobachtungen (privat) von Team- und Projektwissen (geteilt), mit Regeln zum Hochstufen geprüften Wissens. | 2026-08-29 |
| [[Typed Relationships]] | pattern | Verwendung semantisch aussagekräftiger Beziehungstypen (uses, depends-on, contradicts, caused, fixed, supersedes, replaces) statt undifferenzierter Wikilinks. | 2026-08-29 |
| [[User Management]] | workflow | Linux-Ablauf zum Anlegen, Ändern, Überwachen und Löschen von Benutzerkonten mit useradd, usermod und userdel, samt Gruppenverwaltung und sudoers-Konfiguration. | 2026-08-29 |
| [[Vector Search]] | pattern | Semantische Ähnlichkeitssuche über Embedding-Vektoren, die inhaltlich verwandte Seiten auch ohne exakte Schlüsselwortübereinstimmung findet. | 2026-08-29 |
| [[Work Coordination]] | pattern | Leichtgewichtige Erfassung von Aufgabenstatus (in Arbeit, blockiert, erledigt, prüfbedürftig) und Zuweisung, um Doppelarbeit bei mehreren Agenten zu vermeiden. | 2026-08-29 |
| [[Workflow Extraction]] | workflow | Herauslösen von Workflow-Abschnitten aus monolithischer Dokumentation | 2026-09-01 |
| [[Workflow Orchestration]] | workflow | Orchestrierte Einzelbefehle für vollständige Operationen (ingest run, lint run, update run) mit Dry-Run-Vorschau vor dem Schreiben | 2026-08-29 |
| [[Working Memory]] | architecture | Kurzlebige Speicherschicht für jüngste Beobachtungen und vorläufige Befunde vor der Verdichtung; niedrigste Konfidenz, keine Verdichtung, wird zum Sitzungsende erneuert. | 2026-08-29 |
## Problemstellungen
| Page | Type | Summary | Last Modified |
|------|------|---------|----------------|
| [[Ambient Environment Dependency]] | problem | Fehlerklasse, in der ein Test gruen ist, weil die Maschine zufaellig passt statt weil der Code stimmt - abgegrenzt gegen den Green Suite Blind Spot, belegt an vier Faellen unter Gitea-Issue #8 | 2026-08-31 |
| [[Detect-Repair Asymmetry]] | problem | Werkzeugluecke, in der ein Check einen Defekt zuverlaessig meldet, aber kein Befehl ihn behebt - womit die Handeditierung der einzige verbleibende Ausweg ist | 2026-08-31 |
| [[Green Suite Blind Spot]] | problem | Defekt, der eine vollstaendig gruene Testsuite ueberlebt, weil nie ein Test das richtige Verhalten behauptet hat - belegt an drei prio/1-2-Defekten (Round-Trip, Zitat-Notation-als-Code, Zitat-Limit) | 2026-08-31 |
| [[Naming Convention Conflict]] | problem | Widerspruch zwischen README.md (kebab-case) und AGENTS.md (lesbar mit Leerzeichen), der zu Drift bei der Validierung führt | 2026-08-29 |
| [[Write-Once Frontmatter Fields]] | problem | Defektklasse, in der ein Feld nur beim Anlegen der Seite schreibbar ist und danach unerreichbar bleibt, weil kein Mutationsbefehl es kennt und new nicht idempotent ist | 2026-08-31 |
## Protokolle
| Page | Type | Summary | Last Modified |
|------|------|---------|----------------|
| [[CPPC]] | protocol | Hardwareschnittstelle Collaborative Processor Performance Control für feingranulares CPU-Power-Management zwischen Betriebssystem und AMD-Prozessor. | 2026-08-29 |
| [[Modbus]] | protocol | Industrielles Kommunikationsprotokoll von 1979 zur Anbindung speicherprogrammierbarer Steuerungen und Geräte über serielle oder TCP-Netze. | 2026-08-29 |
| [[SSD TRIM]] | protocol | Datenträgerbefehl, mit dem SSDs ungenutzte Blöcke zurückgewinnen - für gleichbleibende Leistung und längere Lebensdauer. | 2026-08-29 |
+12 -1
View File
@@ -18,7 +18,7 @@ The page tables live in a generated `INDEX.md` inside each collection, linked be
- **Concepts:** 80
- **Entities:** 72
- **Sources:** 29
- **Last Updated:** 2026-09-04
- **Last Updated:** 2026-09-08
---
@@ -31,6 +31,17 @@ The page tables live in a generated `INDEX.md` inside each collection, linked be
| `entities/` | 72 | [entities/INDEX.md](entities/INDEX.md) |
| `sources/` | 29 | [sources/INDEX.md](sources/INDEX.md) |
### concepts/
| Area | Pages | Index |
|------|------:|-------|
| Abläufe | 28 | [concepts/INDEX.md#abläufe](concepts/INDEX.md#abläufe) |
| Architekturen | 20 | [concepts/INDEX.md#architekturen](concepts/INDEX.md#architekturen) |
| Entscheidungen | 7 | [concepts/INDEX.md#entscheidungen](concepts/INDEX.md#entscheidungen) |
| Muster | 17 | [concepts/INDEX.md#muster](concepts/INDEX.md#muster) |
| Problemstellungen | 5 | [concepts/INDEX.md#problemstellungen](concepts/INDEX.md#problemstellungen) |
| Protokolle | 3 | [concepts/INDEX.md#protokolle](concepts/INDEX.md#protokolle) |
### entities/
| Area | Pages | Index |
+6
View File
@@ -155,3 +155,9 @@ Drittes status/-Flag status/incoming ergaenzt (Gitea #63): status/-Tabelle auf d
wikitool move --reconcile hat die drei nach kb/entities/projects/{kfchou,vanillaflava,yugasun}/*.md verschachtelten Seiten (wiki-skills, wiki-skills-vanillaflava, llm-wiki-skills) nach kb/entities/projects/ hochgezogen und die drei geleerten Owner-Verzeichnisse entfernt. index rebuild und sources rebuild-index liefen danach; lint meldet weder misplaced_pages noch nested_pages noch duplicate_titles; migrate verify --from HEAD bestaetigt compared=182, added=0, removed=0, alle drei als moved.
---
## [2026-09-08] move | kb/concepts/ bekommt Areas: 80 Seiten in ihre concept_type-Verzeichnisse (#59)
types/concept.md deklariert ein layout: fuer alle sechs concept_type-Werte; wikitool move --reconcile hat daraufhin alle 80 Seiten aus kb/concepts/ in ihre Area gezogen (architectures 20, patterns 17, protocols 3, workflows 28, decisions 7, problems 5). Kein Titel, kein Body, kein Frontmatter-Feld angefasst: migrate verify --from HEAD bestaetigt compared=182, added=0, removed=0, 80 moved, 0 findings. index rebuild erzeugt sechs Areas, keine ueber der Shard-Schwelle von 50, also kein eigener Shard - die Schwelle wirkt wieder als Schwelle statt als toter Wert. lint meldet danach weder misplaced_pages noch den neuen unsharded_collections-Befund.
---
+1 -1
View File
@@ -51,7 +51,7 @@ tools/wikitool <command> --help
| `index rebuild [--dry-run]` | Regenerate the catalog from every page's frontmatter: `kb/index.md` becomes a map (statistics, one row per collection and per area, links to the shards) and the page tables are written to a generated `INDEX.md` in each collection. An area past 50 rows gets its own shard. Stale shards from removed collections/areas are deleted in the same pass |
| `log append --op ingest\|query\|lint\|create\|update\|delete\|rename\|move --title "..." [--body "..."\|--body-file path]` | Append a formatted entry to `kb/log.md` |
| `log status` | Read-only: count `ingest` entries logged since the last `lint` entry - the deterministic trigger behind the Maintenance Schedule's "every 10 sources" full-lint cadence |
| `lint [--json] [--markdown out.md] [--full] [--fail-on-error]` | Structural + provenance checks: broken wikilinks, dangling frontmatter references, orphan pages, index drift, schema gaps, duplicate titles, title mismatches, pages nested more than one directory below their collection (hard - the generated catalog folds these into their area silently rather than merely reading it, see #57), uncovered raw files, broken `raw_files:` refs, raw files claimed by more than one source page, unmarked provenance, citation/frontmatter drift, unbalanced generated-region markers, edges whose label is missing or not authorised by the source collection's `outbound:` (both hard once `kb_version` has reached the release that introduced labelled edges - advisory below it, so a corpus mid-migration is not refused by the check measuring it), `see-also` edges whose reverse direction already carries a specific label (advisory only - redundant rather than wrong, and never migration-gated, since no version turns the redundancy into an error), quote-limit overages (>2 blockquoted lines/page, advisory only). Prints only the sections that found something and always writes the full report to `reports/Lint Report <date>.md` (or `--markdown`), naming the path - `--full` prints everything, `--json` prints the findings and writes nothing |
| `lint [--json] [--markdown out.md] [--full] [--fail-on-error]` | Structural + provenance checks: broken wikilinks, dangling frontmatter references, orphan pages, index drift, schema gaps, duplicate titles, title mismatches, pages nested more than one directory below their collection (hard - the generated catalog folds these into their area silently rather than merely reading it, see #57), uncovered raw files, broken `raw_files:` refs, raw files claimed by more than one source page, unmarked provenance, citation/frontmatter drift, unbalanced generated-region markers, edges whose label is missing or not authorised by the source collection's `outbound:` (both hard once `kb_version` has reached the release that introduced labelled edges - advisory below it, so a corpus mid-migration is not refused by the check measuring it), `see-also` edges whose reverse direction already carries a specific label (advisory only - redundant rather than wrong, and never migration-gated, since no version turns the redundancy into an error), a collection past the catalog's per-area shard threshold that has no areas to shard (advisory only - sharding is automatic but per *area*, so a collection nobody gave areas keeps one table however large it grows, #59; reported with the split its subtype field would produce, and only when that split puts every resulting area at or under the threshold, so a lopsided or small collection stays silent), quote-limit overages (>2 blockquoted lines/page, advisory only). Prints only the sections that found something and always writes the full report to `reports/Lint Report <date>.md` (or `--markdown`), naming the path - `--full` prints everything, `--json` prints the findings and writes nothing |
| `search ["<text>"] [--field <predicate> ...] [--kind/--subtype/--collection/--tag <v>] [--regex] [--limit N] [--sort [-]<field>] [--backend <name>] [--matches] [--json]` | Find pages in `kb/` without reading the index. Text search runs through a pluggable backend (`rg` today); `--field` predicates are evaluated on frontmatter - `f=v`, `f~substring`, `'f>=v'`, `'f:*'` (present), `'!f'` (absent), repeatable and ANDed. With no text this is a pure structured query. Results carry kind/summary/confidence so a hit can be judged without opening the page. A page whose frontmatter does not parse can match no positive predicate, so it is **named** rather than dropped: `--json` always carries an `unreadable` list of `{path, reason}` (usually empty), and the table form writes the same lines to stderr. `--regex` is applied by `rg` alone, whose engine is linear; the ranking boosts for title and summary are literal-containment only, so a non-literal pattern is ranked by match count. `rg` is killed after 30 s and reported as a failure. Read-only, and **exempt from the Iteration Budget Gate** |
| `confidence decay [--apply]` | Recompute every page's derived `confidence` as `confidence_base * (1 - 0.01/month)`, floored at 0.2; dry-run by default |
| `confidence init-base [--apply]` | One-time backfill: set `confidence_base` from the current `confidence` on pages that predate the derived-confidence model |
+3 -2
View File
@@ -45,6 +45,7 @@ tools/
conventions.py kb/CONVENTIONS.md: what this instance decided about authoring, as opposed to what the stack enforces
ownership.py the stack-vs-instance boundary under a content stage - one predicate, read by `dist_cmd.py` and `commands/upstream_cmd.py` so the two cannot answer it differently
type_resolver.py type-spec loading and schema resolution
catalog.py how the corpus groups into collections and areas, and the shard threshold - with no CLI attached
lint_core.py the lint checks and the report, with no CLI attached
types_core.py type-spec listing/description, with no CLI attached
markdown_code.py masks code spans/fences so a page may show wiki notation, not only use it
@@ -57,8 +58,8 @@ tools/
```
**Two consumers, one core.** The CLI is not the only caller any more. The cores
(`search/service.py`, `lint_core.py`, `types_core.py`) hold what decides an
answer and import no `typer` and no `rich`; the modules under `commands/` turn
(`search/service.py`, `lint_core.py`, `types_core.py`, `catalog.py`) hold what
decides an answer and import no `typer` and no `rich`; the modules under `commands/` turn
those values into terminal output and those exceptions into exit codes.
`api.Corpus` is the in-process entry point over the same functions - it takes a
corpus root, returns exactly the structures the `--json` forms print, and
+140
View File
@@ -0,0 +1,140 @@
"""How the corpus groups into collections and areas, with no CLI attached.
Split out of `commands/index_build.py` for the reason `lint_core.py` gives at
the top of itself: this is a pure function over a corpus directory, and it was
sitting in a module that imports `typer` and `rich`. `lint` needs the same
grouping - it is what answers "does this collection have areas, and is it over
the threshold?" (Gitea #59) - and `chemenu.api`, the read surface, may not
reach a command module at all. Importing it from there would have pulled the
whole CLI head in behind it.
So the split runs along the same line as lint's: everything that decides *how
the corpus is shaped* lives here; everything that decides *what the catalog
looks like* - the tables, the map, the shard files - stays in
`commands/index_build.py`, which imports from here.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from pathlib import Path
from chemenu.kb_collections import iter_kb_collections
from chemenu.page import Page
from chemenu.type_resolver import resolver
# Rows per area before it is split into its own shard. From the wiki's own
# `Index Scaling` page ("split table sections at >50 entries"), kept as a plain
# number so growth is handled by arithmetic rather than by a judgment call.
SHARD_THRESHOLD = 50
# Display title for pages sitting directly in a collection root rather than in
# an area subdirectory.
UNGROUPED_TITLE = "All"
@dataclass
class Area:
"""One grouping inside a collection: a subdirectory, or the collection root
for pages that sit directly in it."""
name: str
title: str
pages: list[Page] = field(default_factory=list)
own_shard: bool = False
@property
def count(self) -> int:
return len(self.pages)
@dataclass
class Collection:
name: str
areas: list[Area] = field(default_factory=list)
@property
def count(self) -> int:
return sum(area.count for area in self.areas)
def area_titles() -> dict[str, dict[str, str]]:
"""Display titles per collection: `{collection: {area_dir: title}}`, taken
from each type-spec's own `layout:` rather than a hardcoded map - so a new
subtype names its own section by adding a type-spec, with no code change.
Every type-spec is read, not just `entity`'s. That hardcoding was the
asymmetry behind Gitea #59: the axis a collection splits along is declared
in `layout:`, and a second type declaring one would have had its areas
titled by `.title()` on the directory name while entity's got their real
names.
Keyed by collection rather than by directory name alone, because two types
writing into two collections may legitimately use the same area name for
different things (`kb/entities/tools/` and a hypothetical
`kb/concepts/tools/`); a flat map would hand the second one the first's
title. The collection key is the type's `base_dir:`, which is what put the
page in that directory to begin with.
"""
titles: dict[str, dict[str, str]] = {}
for type_path, _frontmatter in resolver.list_type_specs():
try:
if resolver.get_root(type_path) != "kb":
continue
base_dir = resolver.get_base_dir(type_path)
layout = resolver.get_layout(type_path)
except (ValueError, OSError):
continue
if not base_dir or not layout:
continue
per_collection = titles.setdefault(str(base_dir).strip("/"), {})
for key, spec in layout.items():
per_collection.setdefault(spec.get("dir", key), spec.get("title", str(key).title()))
return titles
def group_pages(kb_dir: Path, pages: dict[str, Page]) -> list[Collection]:
"""Group pages by their physical location: collection directory, then area
subdirectory.
Location rather than `kind` because a shard lives in the directory it
describes, and the two agree by construction: a type-spec's `base_dir:` is
what put the page there.
"""
titles = area_titles()
grouped: dict[str, dict[str, Area]] = {}
# Seed from the collections that exist on disk, not only from the ones that
# happen to hold pages: an empty collection is a real (if unfilled) part of
# the wiki, and dropping it from the map would hide it from every reader.
for collection_dir in iter_kb_collections(kb_dir):
grouped.setdefault(collection_dir.name, {})
for page in sorted(pages.values(), key=lambda p: p.title.lower()):
try:
parts = page.path.relative_to(kb_dir).parts
except ValueError: # pragma: no cover - pages always live under kb_dir
continue
if len(parts) < 2:
collection_name, area_name = "(kb root)", ""
else:
collection_name = parts[0]
area_name = parts[1] if len(parts) > 2 else ""
areas = grouped.setdefault(collection_name, {})
area = areas.get(area_name)
if area is None:
title = (
titles.get(collection_name, {}).get(area_name, area_name.title())
if area_name
else UNGROUPED_TITLE
)
area = Area(name=area_name, title=title)
areas[area_name] = area
area.pages.append(page)
collections = []
for name in sorted(grouped):
ordered = sorted(grouped[name].values(), key=lambda a: (a.name == "", a.title.lower()))
for area in ordered:
area.own_shard = bool(area.name) and area.count > SHARD_THRESHOLD
collections.append(Collection(name=name, areas=ordered))
return collections
+10 -90
View File
@@ -18,18 +18,16 @@ The map stays small enough to browse; `wikitool search` answers everything else.
from __future__ import annotations
import re
from dataclasses import dataclass, field
from datetime import date
from pathlib import Path
import typer
from chemenu import config
from chemenu.catalog import SHARD_THRESHOLD, Area, Collection, group_pages
from chemenu.commands._util import rel_path, success
from chemenu.kb_collections import iter_kb_collections
from chemenu.page import Page
from chemenu.kb_scan import GENERATED_INDEX, find_nested_pages, load_kb_pages
from chemenu.type_resolver import resolver
app = typer.Typer(help="Manage the generated wiki catalog (kb/index.md + per-collection INDEX.md).")
@@ -38,15 +36,6 @@ TABLE_SEP = "|------|------|---------|----------------|"
SUMMARY_HEADINGS = ("Description", "Definition", "Summary")
# Rows per area before it is split into its own shard. From the wiki's own
# `Index Scaling` page ("split table sections at >50 entries"), kept as a plain
# number so growth is handled by arithmetic rather than by a judgment call.
SHARD_THRESHOLD = 50
# Display title for pages sitting directly in a collection root rather than in
# an area subdirectory.
UNGROUPED_TITLE = "All"
DO_NOT_EDIT = "<!-- Generated by `wikitool index rebuild`. Do not hand-edit. -->"
@@ -85,86 +74,17 @@ def _table(pages: list[Page]) -> list[str]:
def _anchor(title: str) -> str:
"""GitHub-style heading anchor, so the map can deep-link into a shard."""
slug = re.sub(r"[^a-z0-9\s-]", "", title.lower())
return re.sub(r"\s+", "-", slug.strip())
"""GitHub-style heading anchor, so the map can deep-link into a shard.
@dataclass
class Area:
"""One grouping inside a collection: a subdirectory, or the collection root
for pages that sit directly in it."""
name: str
title: str
pages: list[Page] = field(default_factory=list)
own_shard: bool = False
@property
def count(self) -> int:
return len(self.pages)
@dataclass
class Collection:
name: str
areas: list[Area] = field(default_factory=list)
@property
def count(self) -> int:
return sum(area.count for area in self.areas)
def _area_titles() -> dict[str, str]:
"""Display titles for entity areas, taken from the entity type-spec's own
`layout:` rather than a hardcoded map - so a new subtype names its own
section by adding a type-spec, with no code change."""
layout = resolver.get_layout(resolver.find_type_by_name("entity")) or {}
return {spec.get("dir", key): spec.get("title", key.title()) for key, spec in layout.items()}
def group_pages(kb_dir: Path, pages: dict[str, Page]) -> list[Collection]:
"""Group pages by their physical location: collection directory, then area
subdirectory.
Location rather than `kind` because a shard lives in the directory it
describes, and the two agree by construction: a type-spec's `base_dir:` is
what put the page there.
`\\w` rather than `a-z0-9`, which is not cosmetic: an area title follows the
KB language, and the first non-English one (`Abläufe`) had its umlaut
*deleted* rather than kept, so the map linked at `#ablufe` and the anchor it
was aiming at was `#abläufe`. Every deep link into a shard whose title
carries a non-ASCII letter was silently dead. Nothing surfaced it while the
only areas were entity ones, whose titles happen to be ASCII throughout.
"""
titles = _area_titles()
grouped: dict[str, dict[str, Area]] = {}
# Seed from the collections that exist on disk, not only from the ones that
# happen to hold pages: an empty collection is a real (if unfilled) part of
# the wiki, and dropping it from the map would hide it from every reader.
for collection_dir in iter_kb_collections(kb_dir):
grouped.setdefault(collection_dir.name, {})
for page in sorted(pages.values(), key=lambda p: p.title.lower()):
try:
parts = page.path.relative_to(kb_dir).parts
except ValueError: # pragma: no cover - pages always live under kb_dir
continue
if len(parts) < 2:
collection_name, area_name = "(kb root)", ""
else:
collection_name = parts[0]
area_name = parts[1] if len(parts) > 2 else ""
areas = grouped.setdefault(collection_name, {})
area = areas.get(area_name)
if area is None:
title = titles.get(area_name, area_name.title()) if area_name else UNGROUPED_TITLE
area = Area(name=area_name, title=title)
areas[area_name] = area
area.pages.append(page)
collections = []
for name in sorted(grouped):
ordered = sorted(grouped[name].values(), key=lambda a: (a.name == "", a.title.lower()))
for area in ordered:
area.own_shard = bool(area.name) and area.count > SHARD_THRESHOLD
collections.append(Collection(name=name, areas=ordered))
return collections
slug = re.sub(r"[^\w\s-]", "", title.lower(), flags=re.UNICODE)
return re.sub(r"\s+", "-", slug.strip())
def build_area_shard(area: Area) -> str:
+111
View File
@@ -16,7 +16,10 @@ from __future__ import annotations
from datetime import date
from pathlib import Path
from collections import Counter
from chemenu import blocks, config, kb_collections, links
from chemenu.catalog import SHARD_THRESHOLD, group_pages
from chemenu.frontmatter_io import frontmatter_error
from chemenu.markdown_code import strip_code_spans
from chemenu.provenance import broken_raw_refs as find_broken_raw_refs
@@ -141,6 +144,88 @@ def nested_pages(kb_dir: Path, pages: dict[str, Page]) -> list[dict]:
]
def unsharded_collections(kb_dir: Path, pages: dict[str, Page]) -> list[dict]:
"""Collections past the catalog's shard threshold that have no areas to
shard, together with the subtype split that would give them some.
Sharding is already automatic, and it is per *area*: `index rebuild` hands
an area over `SHARD_THRESHOLD` rows its own `INDEX.md`. Creating an area is
not automatic and nothing ever asked for one - so a collection that never
grew any keeps its whole catalog in a single table, past the threshold,
forever. The threshold is then not a threshold but a dead value (Gitea
#59), and this is the only check that can notice: an ingest sees one
source and cannot see a collection's size, while `lint` sees the corpus and
runs every 10 sources anyway.
**A recommendation, not a failure** (it is deliberately absent from
`HARD_ERROR_KEYS`), and narrow enough to stay one: it fires only where the
split actually helps - every area it would create, the ungrouped remainder
included, lands at or under the threshold. That self-limits in both
directions. A collection under the threshold never fires, so a small
`kb/comparisons/` is not permanently in the report; and a collection whose
subtype values are lopsided (25 of 29 `source_type: notes`) does not fire
either, because splitting it would produce one area over the threshold and
a handful of splinters. What is left is a finding that appears when a
collection grows into it and is silent when it does not.
"""
findings: list[dict] = []
for collection in group_pages(kb_dir, pages):
if collection.count <= SHARD_THRESHOLD:
continue
# An area already exists, so the collection has been split once and
# `index rebuild` shards whatever outgrows the threshold from here.
# A page still sitting in the root is `misplaced_pages`' finding, not
# this one.
if any(area.name for area in collection.areas):
continue
counts: Counter[str] = Counter()
fields: set[str] = set()
# Whether every type writing here already declares the `layout:` that
# turns the subtype into a directory. It decides which half of the fix
# is still owed: without it there is nothing for `move` to compute a
# destination from, with it the move is all that is left.
layouts: set[bool] = set()
for area in collection.areas:
for page in area.pages:
type_path = page.frontmatter.get("type")
if not type_path:
continue
try:
field = resolver.get_subtype_field(type_path, page.path)
layout = resolver.get_layout(type_path, page.path)
except ValueError:
continue
value = page.frontmatter.get(field) if field else None
if not value:
continue
counts[str(value)] += 1
fields.add(str(field))
layouts.add(bool(layout))
if not counts:
continue
# The pages the subtype cannot place stay in the collection root, so
# they are an area of their own for the purpose of this test.
unplaced = collection.count - sum(counts.values())
if max([*counts.values(), unplaced]) > SHARD_THRESHOLD:
continue
findings.append(
{
"collection": collection.name,
"count": collection.count,
"field": ", ".join(sorted(fields)),
"layout_declared": layouts == {True},
"distribution": [
{"value": value, "count": count}
for value, count in sorted(counts.items(), key=lambda kv: (-kv[1], kv[0]))
],
}
)
return findings
def run_lint(kb_dir: Path) -> dict:
pages = load_kb_pages(kb_dir)
duplicate_titles = find_duplicate_title_paths(kb_dir, config.ROOT)
@@ -388,6 +473,7 @@ def run_lint(kb_dir: Path) -> dict:
"duplicate_titles": duplicate_titles,
"misplaced_pages": misplaced,
"nested_pages": nested,
"unsharded_collections": unsharded_collections(kb_dir, pages),
"uncovered_raw_files": find_uncovered_raw_files(config.RAW_DIR, pages),
"broken_raw_refs": find_broken_raw_refs(pages),
"duplicate_raw_file_owners": find_duplicate_raw_file_owners(pages),
@@ -464,6 +550,23 @@ def render_markdown(report: dict) -> str:
"the catalog folds this into its area silently; `wikitool move --reconcile` fixes it "
"when the page's type resolves to a shallower directory, otherwise move it up by hand",
)
_section(
lines, f"Collections Past the Shard Threshold (>{SHARD_THRESHOLD}) With No Areas "
"- recommendation, not an error",
report.get("unsharded_collections", []),
lambda i: f"`kb/{i['collection']}/` holds {i['count']} pages in a single table and has no "
f"areas, so the per-area shard threshold never fires. Splitting on `{i['field']}` would "
"give: "
+ ", ".join(f"{d['value']} {d['count']}" for d in i["distribution"])
+ f" - all at or under {SHARD_THRESHOLD}. "
+ (
"The type-spec already declares the `layout:` for those values, so "
"`wikitool move --reconcile` and `wikitool index rebuild` are the whole fix"
if i.get("layout_declared")
else "Declare a `layout:` for those values in the type-spec, then "
"`wikitool move --reconcile` and `wikitool index rebuild`"
),
)
_section(
lines, "Uncovered Raw Files (no source page)", report["uncovered_raw_files"],
lambda i: f"`{i}`",
@@ -617,6 +720,14 @@ def default_report_path(report: dict) -> Path:
# and there is no version at which "not under the computed directory" becomes
# wrong - only `wikitool move` someone does or does not get to run.
#
# `unsharded_collections` is advisory by construction rather than by tolerance:
# it does not describe anything that is wrong, only a collection that has grown
# past the size at which areas start paying for themselves. Whether to split it
# is an authoring decision about how the corpus is organised - the tool can see
# that the split would work and say so, and that is the whole of its authority.
# Failing on it would also make `lint` red on a corpus that is entirely
# self-consistent, which is the state the recommendation is asking to improve.
#
# `malformed_edges` and `unbalanced_markers` are hard from the start: neither
# describes an unconverted page, only a broken one.
#
+2 -2
View File
@@ -254,7 +254,7 @@ def kb_dir(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path:
kb = tmp_path / "kb"
for sub in ("entities/projects", "entities/systems", "entities/tools",
"entities/technologies", "entities/people",
"concepts", "sources", "comparisons"):
"concepts/protocols", "sources", "comparisons"):
(kb / sub).mkdir(parents=True)
# The contracts carry a real declaration, because three things now read one:
# `docs verify` checks `profile:`/`required_by_stack:`, and `xref add` asks
@@ -302,7 +302,7 @@ def kb_dir(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path:
"\n# gdeploy\n\n## Description\n\nDeploy tool.\n",
)
write_page(
kb / "concepts/Modbus.md",
kb / "concepts/protocols/Modbus.md",
{
"type": "types/concept.md", "concept_type": "protocol",
"tags": [], "created": "2026-07-25", "modified": "2026-07-25",
+3 -3
View File
@@ -54,7 +54,7 @@ def test_cite_add_command_writes_definition_and_prints_marker(kb_dir, raw_dir, m
marker_id = cite_id("Source - Aurora")
assert f"[^{marker_id}]" in result.output
fm, body = read_page(kb_dir / "concepts/Modbus.md")
fm, body = read_page(kb_dir / "concepts/protocols/Modbus.md")
assert "Source - Aurora" in fm["sources"]
assert f"[^{marker_id}]: [[Source - Aurora]]" in body
@@ -194,11 +194,11 @@ def test_cite_sync_command_over_kb(kb_dir, raw_dir, monkeypatch):
assert add_result.exit_code == 0, add_result.output
marker_id = cite_id("Source - Aurora")
fm, body = read_page(kb_dir / "concepts/Modbus.md")
fm, body = read_page(kb_dir / "concepts/protocols/Modbus.md")
body = body.replace("Industrial protocol.", f"Industrial protocol [^{marker_id}].")
from chemenu.frontmatter_io import write_page
write_page(kb_dir / "concepts/Modbus.md", fm, body)
write_page(kb_dir / "concepts/protocols/Modbus.md", fm, body)
result = runner.invoke(app, ["cite", "sync", "--all", "--dry-run"])
assert result.exit_code == 0, result.output
+1 -1
View File
@@ -494,7 +494,7 @@ def test_generated_files_are_recognised_wherever_they_sit():
assert is_generated("kb/provenance.md")
assert is_generated("kb/concepts/INDEX.md")
assert is_generated("kb/entities/tools/INDEX.md")
assert not is_generated("kb/concepts/Modbus.md")
assert not is_generated("kb/concepts/protocols/Modbus.md")
def test_paths_land_in_the_group_a_reviewer_expects():
+25 -3
View File
@@ -15,14 +15,14 @@ from chemenu.kb_scan import GENERATED_INDEX, iter_kb_pages
from chemenu.type_resolver import resolver
def _area_title(subtype: str) -> str:
"""The display title `index rebuild` will use for an entity subtype.
def _area_title(subtype: str, type_path: str = "types/entity.md") -> str:
"""The display title `index rebuild` will use for a subtype's area.
Read from the type-spec rather than written out, because these titles follow
the KB language: hard-coding them made translating the wiki fail tests that
are not about wording at all.
"""
return resolver.get_layout("types/entity.md")[subtype]["title"]
return resolver.get_layout(type_path)[subtype]["title"]
@pytest.fixture
@@ -116,6 +116,28 @@ def test_area_titles_come_from_the_entity_type_spec_layout(plan, kb_dir):
assert f"## {_area_title('tool')}" in entities
def test_anchor_keeps_non_ascii_letters():
"""An area title follows the KB language, so it may carry a letter outside
`a-z`. Deleting it - which is what the old `[^a-z0-9\\s-]` did - produced a
map link (`#ablufe`) that pointed at no heading in the shard it named."""
assert _anchor("Abläufe") == "abläufe"
assert _anchor("Größere Muster") == "größere-muster"
# Punctuation is still dropped and any run of whitespace still collapses to
# a single hyphen.
assert _anchor("Tools & Utilities (v2)") == "tools-utilities-v2"
def test_area_titles_are_read_from_every_type_spec_not_only_entity(plan, kb_dir):
"""Gitea #59: the title lookup used to resolve `entity` by name and read
only its `layout:`, so a second type declaring one got `.title()` on its
directory name (`Protocols`) instead of the title it declared. The concept
areas are the first case; nothing about them is special."""
concepts = _shard(plan, kb_dir, "concepts")
declared = _area_title("protocol", "types/concept.md")
assert f"## {declared}" in concepts
assert "## Protocols" not in concepts
def test_summary_prefers_frontmatter_then_falls_back_to_body(plan, kb_dir):
entities = _shard(plan, kb_dir, "entities")
assert "Server hosting DocStore with ZFS storage" in entities # frontmatter
+128 -4
View File
@@ -236,9 +236,133 @@ def test_lint_is_silent_about_pages_directly_in_an_area(kb_dir):
assert run_lint(kb_dir)["nested_pages"] == []
# --- collections that outgrew the shard threshold without areas (Gitea #59) ---
def _fill_collection(kb_dir, collection: str, type_path: str, field: str, distribution: dict):
"""Write pages flat into `kb/<collection>/`, `distribution` many per subtype
value - the shape a collection is in when nobody ever created an area.
The precondition is established rather than assumed: the fixture corpus
places its concept page in an area (as the real one does now), and a single
area is enough to make this finding stand down.
"""
for existing in sorted((kb_dir / collection).rglob("*.md")):
if existing.parent != kb_dir / collection:
existing.rename(kb_dir / collection / existing.name)
for area in sorted(p for p in (kb_dir / collection).iterdir() if p.is_dir()):
area.rmdir()
n = 0
for value, count in distribution.items():
for _ in range(count):
n += 1
write_page(
kb_dir / collection / f"page-{n:03d}.md",
{
"type": type_path, field: value,
"created": "2026-08-01", "modified": "2026-08-01",
"provenance": "general", "summary": f"Page {n}",
},
f"\n# page-{n:03d}\n",
)
def test_lint_recommends_areas_for_a_collection_past_the_threshold(kb_dir):
"""The concepts shape: 80 pages in one table, no areas, and a subtype axis
whose largest value (28) lands well under the threshold. Sharding is
per-area and automatic, so a collection with no areas never splits however
large it grows - the threshold is a dead value until someone makes areas."""
_fill_collection(
kb_dir, "concepts", "types/concept.md", "concept_type",
{"workflow": 28, "architecture": 20, "pattern": 17,
"decision": 7, "problem": 5, "protocol": 3},
)
finding = next(
i for i in run_lint(kb_dir)["unsharded_collections"] if i["collection"] == "concepts"
)
assert finding["field"] == "concept_type"
# Largest first, so the reader sees the area that decides whether it helps.
assert finding["distribution"][0] == {"value": "workflow", "count": 28}
assert {d["value"] for d in finding["distribution"]} == {
"workflow", "architecture", "pattern", "decision", "problem", "protocol"
}
assert finding["count"] == sum(d["count"] for d in finding["distribution"])
# `types/concept.md` carries the layout, so only the move is still owed -
# the report line says which half of the fix that is.
assert finding["layout_declared"] is True
def test_the_area_recommendation_says_the_layout_is_missing_when_it_is(kb_dir):
"""`types/source.md` deliberately declares no `layout:`, so a sources
collection that grew past the threshold has nothing for `move` to compute a
destination from - the fix starts one step earlier, and the report says so."""
_fill_collection(
kb_dir, "sources", "types/source.md", "source_type",
{"notes": 30, "article": 21},
)
report = run_lint(kb_dir)
finding = next(i for i in report["unsharded_collections"] if i["collection"] == "sources")
assert finding["layout_declared"] is False
assert "Declare a `layout:`" in render_markdown(report)
def test_the_area_recommendation_is_not_a_failure(kb_dir):
"""It reports a collection that has outgrown a layout, not a broken one.
A corpus whose only finding is this must stay green, or every instance
goes red on the release that shipped the check."""
_fill_collection(
kb_dir, "concepts", "types/concept.md", "concept_type",
{"workflow": 28, "architecture": 24},
)
report = run_lint(kb_dir)
assert report["unsharded_collections"]
assert "unsharded_collections" not in HARD_ERROR_KEYS
assert not has_hard_errors({**{key: [] for key in HARD_ERROR_KEYS},
"unsharded_collections": report["unsharded_collections"]})
def test_lint_is_silent_about_a_collection_under_the_threshold(kb_dir):
"""The sources shape: 29 pages, lopsided across `source_type` - and under
the threshold anyway, so it never fires. That is what keeps the bad split
(one area of 25 plus four splinters) from ever being recommended, without
the check needing to know anything about sources."""
_fill_collection(
kb_dir, "sources", "types/source.md", "source_type",
{"notes": 25, "article": 3, "document": 1},
)
assert run_lint(kb_dir)["unsharded_collections"] == []
def test_lint_is_silent_when_the_split_would_not_help(kb_dir):
"""Past the threshold, but 70 of 80 share one subtype value: splitting
produces one area still over the threshold plus splinters, which is not an
improvement. The second half of the criterion, and the one a
threshold-only check would have got wrong."""
_fill_collection(
kb_dir, "concepts", "types/concept.md", "concept_type",
{"workflow": 70, "architecture": 6, "pattern": 5},
)
assert run_lint(kb_dir)["unsharded_collections"] == []
def test_lint_is_silent_once_the_collection_has_areas(kb_dir):
"""After the fix - the pages sit in their areas - the finding goes away,
and `index rebuild` shards whatever outgrows the threshold from here."""
_fill_collection(
kb_dir, "concepts", "types/concept.md", "concept_type",
{"workflow": 28, "architecture": 24},
)
for page in sorted((kb_dir / "concepts").glob("page-*.md")):
area = "workflows" if "workflow" in page.read_text(encoding="utf-8") else "architectures"
(kb_dir / "concepts" / area).mkdir(exist_ok=True)
page.rename(kb_dir / "concepts" / area / page.name)
assert run_lint(kb_dir)["unsharded_collections"] == []
def test_lint_flags_legacy_citation_marker_as_hard_error(kb_dir):
write_page(
kb_dir / "concepts/Modbus.md",
kb_dir / "concepts/protocols/Modbus.md",
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7},
"\n# Modbus\n\n## Definition\n\nUses port 502 ^[[Source - Aurora]].\n",
@@ -250,7 +374,7 @@ def test_lint_flags_legacy_citation_marker_as_hard_error(kb_dir):
def test_lint_flags_undefined_footnote_ref_as_hard_error(kb_dir):
write_page(
kb_dir / "concepts/Modbus.md",
kb_dir / "concepts/protocols/Modbus.md",
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
"modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7},
"\n# Modbus\n\n## Definition\n\nUses port 502 [^s-ghost].\n",
@@ -264,7 +388,7 @@ def test_lint_flags_orphan_footnote_def_as_hard_error(kb_dir):
cid = cite_id("Source - Aurora")
block = render_cite_block({cid: ("Source - Aurora", None)})
write_page(
kb_dir / "concepts/Modbus.md",
kb_dir / "concepts/protocols/Modbus.md",
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7},
f"\n# Modbus\n\n## Definition\n\nIndustrial protocol, no citation here.\n\n{block}",
@@ -278,7 +402,7 @@ def test_lint_clean_footnote_citation_has_no_hard_errors(kb_dir):
cid = cite_id("Source - Aurora")
block = render_cite_block({cid: ("Source - Aurora", None)})
write_page(
kb_dir / "concepts/Modbus.md",
kb_dir / "concepts/protocols/Modbus.md",
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7},
f"\n# Modbus\n\n## Definition\n\nUses port 502 [^{cid}].\n\n{block}",
+9 -4
View File
@@ -43,8 +43,8 @@ def test_page_subdir_falls_back_for_unmapped_subtype():
def test_page_subdir_is_none_for_types_without_layout():
assert _page_subdir(None, "types/concept.md") is None
assert _page_subdir("anything", "types/concept.md") is None
assert _page_subdir(None, "types/source.md") is None
assert _page_subdir("anything", "types/source.md") is None
def test_coerce_set_value_uses_declared_schema_type():
@@ -258,12 +258,17 @@ def test_new_source_rejects_invalid_source_type(monkeypatch, kb_dir):
assert result.exit_code != 0
def test_new_concept_creates_page(monkeypatch, kb_dir):
def test_new_concept_creates_page_in_its_subtype_area(monkeypatch, kb_dir):
"""Gitea #59: `types/concept.md` declares a `layout:` now, so a new concept
reaches its area with nothing else asked of the author - the same rule that
has always placed an entity. Nothing about `new` changed to make this true;
the type-spec did."""
result = _invoke_new(monkeypatch, kb_dir, [
"new", "concept", "--name", "Event Sourcing", "--set", "concept_type=pattern",
])
assert result.exit_code == 0, result.output
assert (kb_dir / "concepts/Event Sourcing.md").exists()
assert (kb_dir / "concepts/patterns/Event Sourcing.md").exists()
assert not (kb_dir / "concepts/Event Sourcing.md").exists()
def test_new_concept_rejects_invalid_concept_type(monkeypatch, kb_dir):
+4 -4
View File
@@ -310,7 +310,7 @@ def test_citing_pages_via_frontmatter_and_inline(kb_dir, raw_dir):
)
refs, block = _footnote_block(("Source - Aurora", None))
write_page(
kb_dir / "concepts/Modbus.md",
kb_dir / "concepts/protocols/Modbus.md",
{
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
"modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7,
@@ -325,7 +325,7 @@ def test_citing_pages_via_frontmatter_and_inline(kb_dir, raw_dir):
def test_page_raw_files_resolves_through_sources_and_inline(kb_dir, raw_dir):
write_page(
kb_dir / "concepts/Modbus.md",
kb_dir / "concepts/protocols/Modbus.md",
{
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7,
@@ -429,7 +429,7 @@ def test_lint_flags_citation_not_in_frontmatter_sources(kb_dir, raw_dir, monkeyp
refs, block = _footnote_block(("Source - Aurora", None))
write_page(
kb_dir / "concepts/Modbus.md",
kb_dir / "concepts/protocols/Modbus.md",
{
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
"modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7,
@@ -451,7 +451,7 @@ def test_lint_no_drift_when_source_declared(kb_dir, raw_dir, monkeypatch):
refs, block = _footnote_block(("Source - Aurora", None))
write_page(
kb_dir / "concepts/Modbus.md",
kb_dir / "concepts/protocols/Modbus.md",
{
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7,
+29 -4
View File
@@ -86,9 +86,34 @@ def test_get_layout_reads_entity_type_specs_own_layout_field():
assert list(layout) == ["project", "system", "tool", "technology", "person"]
def test_concept_layout_covers_every_declared_concept_type():
"""Gitea #59: `kb/concepts/` had no areas, so the catalog's per-area shard
threshold could never fire however large it grew. A value missing from the
layout would still be *placed* (`subtype_dir` pluralizes the fallback), but
into an area with no declared title - so the schema's enum and the layout
have to agree, and this is what checks that they do."""
layout = resolver.get_layout("types/concept.md")
assert layout is not None
assert set(layout) == set(resolver.get_enum("types/concept.md", "concept_type"))
# `dir` is structural; `title` is display text following the KB language, so
# it is checked for presence rather than wording (see the entity test above).
assert {key: spec["dir"] for key, spec in layout.items()} == {
"architecture": "architectures",
"pattern": "patterns",
"protocol": "protocols",
"workflow": "workflows",
"decision": "decisions",
"problem": "problems",
}
assert all(spec.get("title") for spec in layout.values())
def test_get_layout_is_none_for_types_without_one():
"""`comparison` has no subtype field at all; `source` has one and
deliberately declares no `layout:` anyway - 25 of its 29 pages carry the
same `source_type`, so splitting on it would make one area and four
splinters (Gitea #59). Having a subtype axis is not a reason to use it."""
assert resolver.get_layout("types/comparison.md") is None
assert resolver.get_layout("types/concept.md") is None
assert resolver.get_layout("types/source.md") is None
@@ -237,7 +262,7 @@ def test_subtype_dir_falls_back_for_unmapped_subtype():
def test_subtype_dir_is_none_without_layout_or_subtype():
assert resolver.subtype_dir("types/concept.md", "workflow") is None
assert resolver.subtype_dir("types/source.md", "notes") is None
assert resolver.subtype_dir("types/entity.md", None) is None
@@ -247,8 +272,8 @@ def test_compute_target_dir_applies_layout_subdirectory():
def test_compute_target_dir_is_flat_for_a_type_without_layout():
target = resolver.compute_target_dir("types/concept.md", {"concept_type": "workflow"})
assert target == config.KB_DIR / "concepts"
target = resolver.compute_target_dir("types/source.md", {"source_type": "notes"})
assert target == config.KB_DIR / "sources"
def test_compute_target_dir_resolves_against_repo_root_for_root_repo_types():

Some files were not shown because too many files have changed in this diff Show More