kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59)
Files changed: - CHANGES.md - README.md - VERSION - kb/concepts/Ambient Environment Dependency.md - kb/concepts/Anti-Cramming Heuristic.md - kb/concepts/Audit Trail.md - kb/concepts/BM25.md - kb/concepts/Bulk Operations.md - kb/concepts/CI Integration.md - kb/concepts/COLLECTION.md - kb/concepts/CPPC.md - kb/concepts/Checkpoint Audit.md - kb/concepts/Claude Code Auto Mode.md - kb/concepts/Command Round-Trip Integrity.md - kb/concepts/Confidence Scoring.md - kb/concepts/Consolidation Tiers.md - kb/concepts/Content Quality Control.md - kb/concepts/Context Isolation.md - kb/concepts/Contradiction Resolution.md - kb/concepts/Cross-platform Agent Skills.md - kb/concepts/Crystallization.md - kb/concepts/Delete Rather Than Anonymize.md - kb/concepts/Denylist over Allowlist.md - kb/concepts/Detect-Repair Asymmetry.md - kb/concepts/Diff-Reviewable Agent Edits.md - kb/concepts/Dual Licensing by File Plan.md - kb/concepts/Entity Extraction.md - kb/concepts/Episodic Memory.md - kb/concepts/Event-Driven Automation.md - kb/concepts/Filter on Ingest.md - kb/concepts/Forgetting.md - kb/concepts/Graph Traversal.md - kb/concepts/Green Suite Blind Spot.md - kb/concepts/Hooks.md - kb/concepts/Hybrid Search.md - kb/concepts/INDEX.md - kb/concepts/Implementation Spectrum.md - kb/concepts/Index Scaling.md - kb/concepts/Issue Label Scheme.md - kb/concepts/Iteration and Cost Limits.md - kb/concepts/KB Migration.md - kb/concepts/KB Stack Versioning.md - kb/concepts/Knowledge Compounding.md - kb/concepts/Knowledge Graph.md - kb/concepts/LLM Wiki Pattern.md - kb/concepts/Lint Workflow.md - kb/concepts/MCP-Leseserver.md - kb/concepts/Mass-Update Gate.md - kb/concepts/Memory Lifecycle.md - kb/concepts/Mesh Sync.md - kb/concepts/Modbus.md - kb/concepts/Multi-Agent Collaboration.md - kb/concepts/Naming Convention Conflict.md - kb/concepts/OKF Compatibility.md - kb/concepts/Optional Instance Context File.md - kb/concepts/Personalization Plane.md - kb/concepts/Privacy and Governance.md - kb/concepts/Procedural Memory.md - kb/concepts/Publish-Remote Gate.md - kb/concepts/Quality Scoring.md - kb/concepts/Quality and Self-Correction.md - kb/concepts/RAG.md - kb/concepts/Reciprocal Rank Fusion.md - kb/concepts/SSD TRIM.md - kb/concepts/Scale Ceiling.md - kb/concepts/Self-Healing.md - kb/concepts/Semantic Lint Automation.md - kb/concepts/Semantic Memory.md - kb/concepts/Session Orientation.md - kb/concepts/Shared vs Private.md - kb/concepts/Split Merge Reclassify.md - kb/concepts/Split Threshold.md - kb/concepts/Structural Enforcement over Documented Rule.md - kb/concepts/Stub Threshold.md - kb/concepts/Supersession.md - kb/concepts/Three-Layer Architecture.md - kb/concepts/Token Economics.md - kb/concepts/Typed Relationships.md - kb/concepts/User Management.md - kb/concepts/Vector Search.md - kb/concepts/Work Coordination.md - kb/concepts/Workflow Extraction.md - kb/concepts/Workflow Orchestration.md - kb/concepts/Working Memory.md - kb/concepts/Write-Once Frontmatter Fields.md - kb/concepts/architectures/Consolidation Tiers.md - kb/concepts/architectures/Context Isolation.md - kb/concepts/architectures/Cross-platform Agent Skills.md - kb/concepts/architectures/Episodic Memory.md - kb/concepts/architectures/Hybrid Search.md - kb/concepts/architectures/Implementation Spectrum.md - kb/concepts/architectures/Knowledge Graph.md - kb/concepts/architectures/LLM Wiki Pattern.md - kb/concepts/architectures/MCP-Leseserver.md - kb/concepts/architectures/Memory Lifecycle.md - kb/concepts/architectures/OKF Compatibility.md - kb/concepts/architectures/Optional Instance Context File.md - kb/concepts/architectures/Personalization Plane.md - kb/concepts/architectures/Procedural Memory.md - kb/concepts/architectures/RAG.md - kb/concepts/architectures/Scale Ceiling.md - kb/concepts/architectures/Semantic Memory.md - kb/concepts/architectures/Three-Layer Architecture.md - kb/concepts/architectures/Token Economics.md - kb/concepts/architectures/Working Memory.md - kb/concepts/decisions/Delete Rather Than Anonymize.md - kb/concepts/decisions/Denylist over Allowlist.md - kb/concepts/decisions/Diff-Reviewable Agent Edits.md - kb/concepts/decisions/Dual Licensing by File Plan.md - kb/concepts/decisions/Issue Label Scheme.md - kb/concepts/decisions/KB Stack Versioning.md - kb/concepts/decisions/Structural Enforcement over Documented Rule.md - kb/concepts/patterns/Audit Trail.md - kb/concepts/patterns/BM25.md - kb/concepts/patterns/Command Round-Trip Integrity.md - kb/concepts/patterns/Confidence Scoring.md - kb/concepts/patterns/Contradiction Resolution.md - kb/concepts/patterns/Entity Extraction.md - kb/concepts/patterns/Filter on Ingest.md - kb/concepts/patterns/Forgetting.md - kb/concepts/patterns/Graph Traversal.md - kb/concepts/patterns/Mesh Sync.md - kb/concepts/patterns/Quality Scoring.md - kb/concepts/patterns/Reciprocal Rank Fusion.md - kb/concepts/patterns/Self-Healing.md - kb/concepts/patterns/Shared vs Private.md - kb/concepts/patterns/Typed Relationships.md - kb/concepts/patterns/Vector Search.md - kb/concepts/patterns/Work Coordination.md - kb/concepts/problems/Ambient Environment Dependency.md - kb/concepts/problems/Detect-Repair Asymmetry.md - kb/concepts/problems/Green Suite Blind Spot.md - kb/concepts/problems/Naming Convention Conflict.md - kb/concepts/problems/Write-Once Frontmatter Fields.md - kb/concepts/protocols/CPPC.md - kb/concepts/protocols/Modbus.md - kb/concepts/protocols/SSD TRIM.md - kb/concepts/workflows/Anti-Cramming Heuristic.md - kb/concepts/workflows/Bulk Operations.md - kb/concepts/workflows/CI Integration.md - kb/concepts/workflows/Checkpoint Audit.md - kb/concepts/workflows/Claude Code Auto Mode.md - kb/concepts/workflows/Content Quality Control.md - kb/concepts/workflows/Crystallization.md - kb/concepts/workflows/Event-Driven Automation.md - kb/concepts/workflows/Hooks.md - kb/concepts/workflows/Index Scaling.md - kb/concepts/workflows/Iteration and Cost Limits.md - kb/concepts/workflows/KB Migration.md - kb/concepts/workflows/Knowledge Compounding.md - kb/concepts/workflows/Lint Workflow.md - kb/concepts/workflows/Mass-Update Gate.md - kb/concepts/workflows/Multi-Agent Collaboration.md - kb/concepts/workflows/Privacy and Governance.md - kb/concepts/workflows/Publish-Remote Gate.md - kb/concepts/workflows/Quality and Self-Correction.md - kb/concepts/workflows/Semantic Lint Automation.md - kb/concepts/workflows/Session Orientation.md - kb/concepts/workflows/Split Merge Reclassify.md - kb/concepts/workflows/Split Threshold.md - kb/concepts/workflows/Stub Threshold.md - kb/concepts/workflows/Supersession.md - kb/concepts/workflows/User Management.md - kb/concepts/workflows/Workflow Extraction.md - kb/concepts/workflows/Workflow Orchestration.md - kb/index.md - kb/log.md - tools/CONTRACT.md - tools/README.md - tools/chemenu/catalog.py - tools/chemenu/commands/index_build.py - tools/chemenu/lint_core.py - tools/chemenu/tests/conftest.py - tools/chemenu/tests/test_cite_cmd.py - tools/chemenu/tests/test_git_publish.py - tools/chemenu/tests/test_index_build.py - tools/chemenu/tests/test_lint.py - tools/chemenu/tests/test_new_page.py - tools/chemenu/tests/test_provenance.py - tools/chemenu/tests/test_type_resolver.py - tools/chemenu/tests/test_xref.py - types/concept.md - types/type-spec.md
This commit is contained in:
+77
-1
@@ -35,7 +35,7 @@ dev-checkout concern - readable here, never shipped as something to parse.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 4.8.0-beta.5 - 2026-09-05 - update entity naming conventions to use singular form for consistency
|
## 4.8.0-beta.6 - 2026-09-08 - kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59)
|
||||||
|
|
||||||
**Author:** Torben Nehmer
|
**Author:** Torben Nehmer
|
||||||
|
|
||||||
@@ -46,6 +46,7 @@ dev-checkout concern - readable here, never shipped as something to parse.
|
|||||||
- raw accept: incoming/ als abgeleiteter Rohablage-Eingang (schliesst #58)
|
- raw accept: incoming/ als abgeleiteter Rohablage-Eingang (schliesst #58)
|
||||||
- raw accept: Stem-Eindeutigkeit im Typverzeichnis erzwingen, --replaces als einziger Weg daran vorbei (schliesst #64)
|
- raw accept: Stem-Eindeutigkeit im Typverzeichnis erzwingen, --replaces als einziger Weg daran vorbei (schliesst #64)
|
||||||
- update entity naming conventions to use singular form for consistency
|
- update entity naming conventions to use singular form for consistency
|
||||||
|
- kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59)
|
||||||
<!-- /wikitool:bumps -->
|
<!-- /wikitool:bumps -->
|
||||||
|
|
||||||
Das Label `status/incoming` gibt es seit heute in Gitea: der Mensch legt einen
|
Das Label `status/incoming` gibt es seit heute in Gitea: der Mensch legt einen
|
||||||
@@ -301,6 +302,81 @@ keine Kompatibilitätsfrage, solange sie vor `4.8.0` landet.
|
|||||||
|
|
||||||
Schließt #64.
|
Schließt #64.
|
||||||
|
|
||||||
|
**`kb/concepts/` bekommt Areas** (#59): Sharden ist längst automatisch —
|
||||||
|
`index_build.SHARD_THRESHOLD = 50`, hergeleitet aus der wikieigenen Seite
|
||||||
|
`Index Scaling` — aber es passiert **pro Area**, und eine Area legt niemand an.
|
||||||
|
`kb/concepts/` hatte keine, also war die Schwelle dort ein toter Wert: 80 Seiten
|
||||||
|
in einer einzigen Tabelle, weit über der eigenen Grenze, ohne dass je etwas
|
||||||
|
gefeuert hätte. Die Ursache war eine Asymmetrie in den Type-Specs — `entity`
|
||||||
|
deklarierte ein `layout:`, `concept` nicht, obwohl das Subtype-Feld fertig dalag.
|
||||||
|
|
||||||
|
`types/concept.md` deklariert es jetzt für alle sechs `concept_type`-Werte
|
||||||
|
(`architectures/`, `patterns/`, `protocols/`, `workflows/`, `decisions/`,
|
||||||
|
`problems/`). Für `source` bewusst **nicht**: 25 von 29 Seiten sind `notes`, die
|
||||||
|
Aufteilung ergäbe eine Area und vier Splitter, und `kb/sources/` liegt mit 29
|
||||||
|
Seiten ohnehin unter der Schwelle. Ein Subtype-Feld zu haben ist kein Grund, es
|
||||||
|
als Achse zu benutzen.
|
||||||
|
|
||||||
|
Zwei Dinge im Code, beide Folgen desselben Befunds. `_area_titles()` in
|
||||||
|
`index_build.py` löste `entity` fest über `find_type_by_name("entity")` auf und
|
||||||
|
las nur dessen `layout:` — jeder zweite Typ mit einem `layout:` hätte
|
||||||
|
`.title()`-Namen auf dem Verzeichnisnamen bekommen statt der deklarierten Titel.
|
||||||
|
Es liest jetzt jedes Type-Spec, und zwar **pro Collection** geschlüsselt, damit
|
||||||
|
zwei Typen denselben Area-Namen für Verschiedenes benutzen dürfen. Und der neue
|
||||||
|
`lint`-Befund meldet eine Collection über der Schwelle **ohne** Areas, mit der
|
||||||
|
Verteilung ihres Subtype-Felds — als Empfehlung, nicht als Failure, und nur
|
||||||
|
dann, wenn die Aufteilung jede entstehende Area unter die Schwelle drückt. Das
|
||||||
|
begrenzt sich selbst in beide Richtungen: `kb/comparisons/` mit einer Seite
|
||||||
|
feuert nie, und die schlechte Aufteilung nach `source_type` unterbleibt von
|
||||||
|
allein, ohne dass der Check etwas über Sources wüsste.
|
||||||
|
|
||||||
|
Zwei Dinge fielen unterwegs an, die das Issue nicht vorhergesehen hatte.
|
||||||
|
`lint_core.py` durfte `SHARD_THRESHOLD`/`group_pages` nicht aus
|
||||||
|
`commands/index_build.py` importieren — `test_api.py` prüft strukturell, dass
|
||||||
|
`chemenu.api` kein Modul unter `chemenu.commands` lädt, und der Import hätte den
|
||||||
|
ganzen CLI-Kopf mitgezogen. Die Gruppierung liegt deshalb neu in
|
||||||
|
`tools/chemenu/catalog.py`, entlang derselben Linie wie `lint_core.py`:
|
||||||
|
Korpusform hier, Darstellung dort. Und `_anchor()` strich mit `[^a-z0-9\s-]`
|
||||||
|
jeden Nicht-ASCII-Buchstaben ersatzlos — die Karte verlinkte auf `#ablufe`,
|
||||||
|
während die Überschrift im Shard `#abläufe` heißt. Vorher fiel das keinem auf,
|
||||||
|
weil alle Entity-Area-Titel zufällig ASCII sind; `Abläufe` ist der erste, der es
|
||||||
|
nicht ist.
|
||||||
|
|
||||||
|
Auf dieser Instanz angewendet: `wikitool move --reconcile` hat alle 80
|
||||||
|
Concept-Seiten in ihre Area gezogen, `migrate verify --from HEAD` bestätigt
|
||||||
|
`182 compared, 0 added, 0 removed, 80 moved, 0 findings` — kein Titel, kein
|
||||||
|
Body, kein Frontmatter-Feld angefasst. `index rebuild` erzeugt sechs Areas
|
||||||
|
(Abläufe 28, Architekturen 20, Muster 17, Entscheidungen 7, Problemstellungen 5,
|
||||||
|
Protokolle 3); keine über der Schwelle, also kein eigener Shard, und die
|
||||||
|
Schwelle wirkt wieder als Schwelle.
|
||||||
|
|
||||||
|
Geändert: `types/concept.md` (`layout:`), `types/type-spec.md` (wann ein
|
||||||
|
`layout:` sich lohnt), `tools/chemenu/catalog.py` (neu),
|
||||||
|
`tools/chemenu/commands/index_build.py` (`area_titles`, `_anchor`),
|
||||||
|
`tools/chemenu/lint_core.py` (`unsharded_collections`),
|
||||||
|
`tools/chemenu/tests/` (Fixture-Concept liegt jetzt in seiner Area, plus neun
|
||||||
|
neue Tests), `tools/CONTRACT.md`, `tools/README.md`, `README.md`,
|
||||||
|
`kb/concepts/COLLECTION.md`, sowie die 80 bewegten Seiten unter `kb/concepts/`.
|
||||||
|
|
||||||
|
**MINOR**, nicht MAJOR: der Umzug ist ein **Angebot**, kein Zwang. Eine
|
||||||
|
bestehende Instanz, die `move --reconcile` nicht laufen lässt, bleibt
|
||||||
|
funktionsfähig — `group_pages` liest das Dateisystem, nicht das `layout:`, also
|
||||||
|
landen flache Bestandsseiten in der Area „All" und neu angelegte in ihrer
|
||||||
|
eigenen; beides rendert. Der gemischte Zustand meldet sich als `lint`-Befund
|
||||||
|
*Misplaced Pages*, der seit jeher advisory ist. Und ein Downgrade auf einen
|
||||||
|
Stack ohne dieses `layout:` funktioniert weiter: die Verzeichnisse bleiben
|
||||||
|
Verzeichnisse, nur die Anzeigetitel fallen auf `.title()` zurück. Kosmetik, kein
|
||||||
|
Bruch der Austauschbarkeit in beiden Richtungen.
|
||||||
|
|
||||||
|
Schließt #59.
|
||||||
|
|
||||||
|
**Nachzug an `63b4bb8`:** die Umstellung der Namenskonvention auf `HA Integration`
|
||||||
|
hatte in `README.md` das Gegenbeispiel verloren — die Zeile las
|
||||||
|
``Use singular for entities: `HA Integration.md` (not `HA Integration.md`)``, beide
|
||||||
|
Seiten des „not" identisch, also eine Regel ohne Fall, an dem sie greift.
|
||||||
|
`kb/CONVENTIONS.md` und `kb/entities/COLLECTION.md` hatten im selben Commit das
|
||||||
|
korrekte Paar bekommen; `README.md` zieht jetzt mit `HA Integrations.md` nach.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 4.7.4 - 2026-09-04 - bootstrap.md nennt den session-id-WARN nach frischem Bootstrap explizit als erwartet
|
## 4.7.4 - 2026-09-04 - bootstrap.md nennt den session-id-WARN nach frischem Bootstrap explizit als erwartet
|
||||||
|
|||||||
@@ -105,8 +105,14 @@ chemenu/
|
|||||||
│ │ ├── tools/ # own INDEX.md once past 50 pages
|
│ │ ├── tools/ # own INDEX.md once past 50 pages
|
||||||
│ │ ├── technologies/
|
│ │ ├── technologies/
|
||||||
│ │ └── people/
|
│ │ └── people/
|
||||||
│ ├── concepts/ # COLLECTION.md - architectures, patterns, protocols
|
│ ├── concepts/ # COLLECTION.md + INDEX.md + areas below
|
||||||
│ ├── sources/ # COLLECTION.md - source summaries
|
│ │ ├── architectures/
|
||||||
|
│ │ ├── patterns/
|
||||||
|
│ │ ├── protocols/
|
||||||
|
│ │ ├── workflows/
|
||||||
|
│ │ ├── decisions/
|
||||||
|
│ │ └── problems/
|
||||||
|
│ ├── sources/ # COLLECTION.md - source summaries, no areas by choice
|
||||||
│ └── comparisons/ # COLLECTION.md - comparison pages
|
│ └── comparisons/ # COLLECTION.md - comparison pages
|
||||||
├── work/ # WORKSHOP: one directory per multi-session run, tracked
|
├── work/ # WORKSHOP: one directory per multi-session run, tracked
|
||||||
│ └── CONTRACT.md # Run keys, required files, how a run closes
|
│ └── CONTRACT.md # Run keys, required files, how a run closes
|
||||||
@@ -130,7 +136,16 @@ A directory under `kb/` is a **collection** exactly when it holds a `COLLECTION.
|
|||||||
subdirectory inside one is an **area** that inherits it, and that is as deep as a page goes -
|
subdirectory inside one is an **area** that inherits it, and that is as deep as a page goes -
|
||||||
nothing nests below an area, because the generated catalog reads exactly two path segments
|
nothing nests below an area, because the generated catalog reads exactly two path segments
|
||||||
under `kb/` and would fold a deeper page into the area silently (`kb/CONTRACT.md` § Collections
|
under `kb/` and would fold a deeper page into the area silently (`kb/CONTRACT.md` § Collections
|
||||||
has the rule; `wikitool lint` reports a violation as a hard error). `COLLECTION.md` never
|
has the rule; `wikitool lint` reports a violation as a hard error).
|
||||||
|
|
||||||
|
Which areas a collection has is not chosen per page: a type-spec's `layout:` maps its subtype
|
||||||
|
field onto directories, and `wikitool new` writes the page straight into the one its subtype
|
||||||
|
names. That is also what makes the catalog's shard threshold do anything - `index rebuild`
|
||||||
|
splits **per area**, so a collection with no areas keeps one table however large it grows.
|
||||||
|
`wikitool lint` reports such a collection once it is past the threshold, as a recommendation
|
||||||
|
rather than an error, together with the split its subtype field would produce; it stays quiet
|
||||||
|
when the split would not actually help. `kb/sources/` is the worked example of the second case
|
||||||
|
and deliberately has no areas. `COLLECTION.md` never
|
||||||
appears outside `kb/` - the other layers carry a `CONTRACT.md` or a root type-spec instead. A stage may carry
|
appears outside `kb/` - the other layers carry a `CONTRACT.md` or a root type-spec instead. A stage may carry
|
||||||
both a `README.md` and a `CONTRACT.md`: they have different readers. The README is for humans
|
both a `README.md` and a `CONTRACT.md`: they have different readers. The README is for humans
|
||||||
working *on* that layer, the contract is what binds an agent working *with* it.
|
working *on* that layer, the contract is what binds an agent working *with* it.
|
||||||
@@ -247,7 +262,7 @@ Ingest incoming/notes/my-notes.md
|
|||||||
### Naming
|
### Naming
|
||||||
|
|
||||||
- Use human-readable titles with spaces for files: `Hybrid Search.md`, not kebab-case
|
- Use human-readable titles with spaces for files: `Hybrid Search.md`, not kebab-case
|
||||||
- Use singular for entities: `HA Integration.md` (not `HA Integration.md`)
|
- Use singular for entities: `HA Integration.md` (not `HA Integrations.md`)
|
||||||
- Use wikilinks matching the file name exactly: `[[Entity Name]]`
|
- Use wikilinks matching the file name exactly: `[[Entity Name]]`
|
||||||
- **Titles follow the subject's own established name, not the wiki's language.** `Act Runner` and
|
- **Titles follow the subject's own established name, not the wiki's language.** `Act Runner` and
|
||||||
`GitOps Ownership Model` keep theirs. A title is the only identifier a page has - it also lives
|
`GitOps Ownership Model` keep theirs. A title is the only identifier a page has - it also lives
|
||||||
|
|||||||
@@ -25,7 +25,32 @@ tone, relationship labels, the confidence rubric. Neither is restated here.
|
|||||||
|
|
||||||
## Types offered
|
## Types offered
|
||||||
|
|
||||||
`concept` (`tools/wikitool types describe concept`).
|
`concept` (`tools/wikitool types describe concept`). Das Feld `concept_type:`
|
||||||
|
wählt die Area:
|
||||||
|
|
||||||
|
| Area | Hält |
|
||||||
|
|------|------|
|
||||||
|
| `architectures/` | Aufbau und Struktur: wie ein System geschnitten ist und warum die Schnitte dort liegen |
|
||||||
|
| `patterns/` | Wiederverwendbare Lösungsformen, die über mehr als einen Gegenstand hinweg gelten |
|
||||||
|
| `protocols/` | Kommunikationsprotokolle und Standards, in ihrer üblichen Schreibweise benannt |
|
||||||
|
| `workflows/` | Abläufe und Prozesse, die projektübergreifend wiederkehren |
|
||||||
|
| `decisions/` | Architektur- und Entwurfsentscheidungen (siehe unten) |
|
||||||
|
| `problems/` | Wiederkehrende Problemstellungen und ihre Lösungsansätze |
|
||||||
|
|
||||||
|
Das sind Areas, keine Collections: sie erben diesen Contract und tragen keine
|
||||||
|
eigene `COLLECTION.md`.
|
||||||
|
|
||||||
|
Die Zuordnung trifft niemand von Hand — sie steht als `layout:` in
|
||||||
|
`types/concept.md`, und `wikitool new` legt eine neue Seite direkt dort ab.
|
||||||
|
Eine Seite, die anderswo liegt, meldet `wikitool lint` als *misplaced*;
|
||||||
|
`wikitool move --page "<Titel>"` bringt sie an ihren berechneten Ort.
|
||||||
|
|
||||||
|
Die Aufteilung ist keine Geschmacksfrage, sondern das, was die Shard-Schwelle
|
||||||
|
des Katalogs überhaupt wirksam macht: `index rebuild` teilt **pro Area**, und
|
||||||
|
eine Collection ohne Areas teilt sich nie — mit 80 Seiten in einer einzigen
|
||||||
|
Tabelle war die Schwelle hier ein toter Wert (Gitea #59). Keine der sechs
|
||||||
|
Areas liegt derzeit über der Schwelle, also bekommt auch keine einen eigenen
|
||||||
|
Shard; wächst eine hinein, passiert das ohne Zutun.
|
||||||
|
|
||||||
## Decisions
|
## Decisions
|
||||||
|
|
||||||
|
|||||||
+76
-51
@@ -4,88 +4,113 @@
|
|||||||
|
|
||||||
80 page(s). Regenerated by `wikitool index rebuild`.
|
80 page(s). Regenerated by `wikitool index rebuild`.
|
||||||
|
|
||||||
## All
|
## Abläufe
|
||||||
|
|
||||||
| Page | Type | Summary | Last Modified |
|
| Page | Type | Summary | Last Modified |
|
||||||
|------|------|---------|----------------|
|
|------|------|---------|----------------|
|
||||||
| [[Ambient Environment Dependency]] | problem | Fehlerklasse, in der ein Test gruen ist, weil die Maschine zufaellig passt statt weil der Code stimmt - abgegrenzt gegen den Green Suite Blind Spot, belegt an vier Faellen unter Gitea-Issue #8 | 2026-08-31 |
|
|
||||||
| [[Anti-Cramming Heuristic]] | workflow | Regel gegen überladene Seiten: ab dem dritten Absatz zu einem Unterthema eine eigene Seite anlegen | 2026-08-29 |
|
| [[Anti-Cramming Heuristic]] | workflow | Regel gegen überladene Seiten: ab dem dritten Absatz zu einem Unterthema eine eigene Seite anlegen | 2026-08-29 |
|
||||||
| [[Audit Trail]] | pattern | Unveränderliches chronologisches Log aller Wiki-Operationen (Ingest, Bearbeitung, Löschung, Abfrage) mit Zeitstempel, Akteur, Ziel und Änderungsbeschreibung. | 2026-08-29 |
|
|
||||||
| [[BM25]] | pattern | Schlüsselwortbasiertes Retrieval-Verfahren, das über Termfrequenz, inverse Dokumentfrequenz und Stemming exakte oder teilweise Übereinstimmungen findet. | 2026-08-29 |
|
|
||||||
| [[Bulk Operations]] | workflow | Umkehrbare, protokollierte Operationen zum Massenlöschen, Exportieren, Zusammenführen oder Archivieren von Wiki-Inhalten, mit Freigabepflicht und Undo. | 2026-08-29 |
|
| [[Bulk Operations]] | workflow | Umkehrbare, protokollierte Operationen zum Massenlöschen, Exportieren, Zusammenführen oder Archivieren von Wiki-Inhalten, mit Freigabepflicht und Undo. | 2026-08-29 |
|
||||||
| [[Checkpoint Audit]] | workflow | Regelmäßiger Qualitätsrhythmus: Index und Backlinks alle 15 Einträge neu aufbauen, auf 0 neue Artikel prüfen, die 3 meistgeänderten erneut lesen | 2026-08-29 |
|
| [[Checkpoint Audit]] | workflow | Regelmäßiger Qualitätsrhythmus: Index und Backlinks alle 15 Einträge neu aufbauen, auf 0 neue Artikel prüfen, die 3 meistgeänderten erneut lesen | 2026-08-29 |
|
||||||
| [[CI Integration]] | workflow | CI/CD-Hooks vor dem Publish: ci.yml (Push/PR, Stack-Pfade, seit 1.8.1 mit Coverage-Messung ohne Schwelle) und nightly.yml (Zeitplan, schliesst die paths-ignore-Luecke fuer Content-Drift; schedule-Ausloesung seit 2026-09-01 bestaetigt) setzen Quality Gates durch | 2026-09-01 |
|
| [[CI Integration]] | workflow | CI/CD-Hooks vor dem Publish: ci.yml (Push/PR, Stack-Pfade, seit 1.8.1 mit Coverage-Messung ohne Schwelle) und nightly.yml (Zeitplan, schliesst die paths-ignore-Luecke fuer Content-Drift; schedule-Ausloesung seit 2026-09-01 bestaetigt) setzen Quality Gates durch | 2026-09-01 |
|
||||||
| [[Claude Code Auto Mode]] | workflow | auto-Berechtigungsmodus von Claude Code: ein Klassifikator genehmigt Aktionen vor der Ausfuehrung statt nachzufragen; die Beschreibung stammt weit ueberwiegend aus zweiter Hand ueber einen Doku-Subagenten | 2026-08-31 |
|
| [[Claude Code Auto Mode]] | workflow | auto-Berechtigungsmodus von Claude Code: ein Klassifikator genehmigt Aktionen vor der Ausfuehrung statt nachzufragen; die Beschreibung stammt weit ueberwiegend aus zweiter Hand ueber einen Doku-Subagenten | 2026-08-31 |
|
||||||
| [[Command Round-Trip Integrity]] | pattern | Anforderung, dass zwei Befehle auf derselben Datei in jeder Reihenfolge zusammenpassen und jeder erzeugte Zustand einen Gegenbefehl hat - 2026-08-31 in wikitool zweimal verletzt | 2026-08-31 |
|
|
||||||
| [[Confidence Scoring]] | pattern | Mechanismus, der faktischen Aussagen quantitative Werte nach Quellenzahl, Aktualität, Qualität und Bestätigung zuweist, um gut gestütztes Wissen zu erkennen. | 2026-08-29 |
|
|
||||||
| [[Consolidation Tiers]] | architecture | Hierarchische Speicherarchitektur, die Informationen durch zunehmend verdichtete Schichten vom Working Memory bis zum Semantic und Procedural Memory befördert. | 2026-08-29 |
|
|
||||||
| [[Content Quality Control]] | workflow | Regeln und Schwellenwerte für die Seitenqualität: Mindestumfang für Stubs, Aufteilungsschwellen und Zielwerte für die Zeilenzahl | 2026-08-29 |
|
| [[Content Quality Control]] | workflow | Regeln und Schwellenwerte für die Seitenqualität: Mindestumfang für Stubs, Aufteilungsschwellen und Zielwerte für die Zeilenzahl | 2026-08-29 |
|
||||||
| [[Context Isolation]] | architecture | Grundsatz, für jede Aufgabe nur den jeweils benötigten Kontext zu laden | 2026-08-29 |
|
|
||||||
| [[Contradiction Resolution]] | pattern | Automatisches Erkennen und Auflösen widersprüchlicher Aussagen anhand von Konfidenz, Aktualität und Autorität der Quelle. | 2026-08-29 |
|
|
||||||
| [[CPPC]] | protocol | Hardwareschnittstelle Collaborative Processor Performance Control für feingranulares CPU-Power-Management zwischen Betriebssystem und AMD-Prozessor. | 2026-08-29 |
|
|
||||||
| [[Cross-platform Agent Skills]] | architecture | Architektur fuer Agent-Skills, die ueber mehrere LLM-Werkzeuge hinweg funktionieren; in Chemenu selbst am 2026-08-04 umgesetzt und ueberprueft | 2026-09-01 |
|
|
||||||
| [[Crystallization]] | workflow | Verdichten abgeschlossener Erkundungen, Debugging-Sitzungen und Recherchen zu strukturierten Wiki-Auszügen als eigenständige Wissensquellen. | 2026-08-29 |
|
| [[Crystallization]] | workflow | Verdichten abgeschlossener Erkundungen, Debugging-Sitzungen und Recherchen zu strukturierten Wiki-Auszügen als eigenständige Wissensquellen. | 2026-08-29 |
|
||||||
| [[Delete Rather Than Anonymize]] | decision | Private Korpusinhalte per Loeschung entfernen statt zu anonymisieren: ein Seitentitel ist der einzige Identifier eines Wikis, Umbenennen ist die volle page-lifecycle-Prozedur je Seite, Loeschen ist ein unterstuetztes Kommando. | 2026-09-01 |
|
|
||||||
| [[Denylist over Allowlist]] | decision | Entscheidung, schreibbare Felder als Schema minus kurzer Sperrliste zu bestimmen statt als gepflegte Positivliste, weil die Positivliste eine zweite Kopie des Schemas waere | 2026-08-31 |
|
|
||||||
| [[Detect-Repair Asymmetry]] | problem | Werkzeugluecke, in der ein Check einen Defekt zuverlaessig meldet, aber kein Befehl ihn behebt - womit die Handeditierung der einzige verbleibende Ausweg ist | 2026-08-31 |
|
|
||||||
| [[Diff-Reviewable Agent Edits]] | decision | Entscheidung, Dateiaenderungen ueber Edit/Write statt ueber Shell-Heredocs zu fahren, weil nur das erste eine pruefbare Diff hinterlaesst | 2026-08-31 |
|
|
||||||
| [[Dual Licensing by File Plan]] | decision | Ein Repo mit Code- und Inhaltsanteil erhaelt zwei Lizenzen; die Grenze zwischen ihnen ist kein zweiter, gepflegter Pfadkatalog, sondern der ohnehin vorhandene Dateiplan des Distributionswerkzeugs. | 2026-09-01 |
|
|
||||||
| [[Entity Extraction]] | pattern | Erkennen und Strukturieren von Entities (Personen, Projekte, Bibliotheken, Concepts, Dateien, Entscheidungen, Systeme, Werkzeuge) samt typspezifischer Attribute aus Rohquellen. | 2026-08-29 |
|
|
||||||
| [[Episodic Memory]] | architecture | Speicherschicht für verdichtete Sitzungszusammenfassungen und Befunde; Brücke zwischen rohem Working Memory und langlebigem Semantic Memory. | 2026-08-29 |
|
|
||||||
| [[Event-Driven Automation]] | workflow | Muster, das automatische Auslöser an Wiki-Lebenszyklusereignisse hängt, um manuellen Pflegeaufwand und das Risiko der Verwahrlosung zu senken. | 2026-08-29 |
|
| [[Event-Driven Automation]] | workflow | Muster, das automatische Auslöser an Wiki-Lebenszyklusereignisse hängt, um manuellen Pflegeaufwand und das Risiko der Verwahrlosung zu senken. | 2026-08-29 |
|
||||||
| [[Filter on Ingest]] | pattern | Automatisches Erkennen und Entfernen sensibler Daten (API-Schlüssel, Token, Credentials, personenbezogene Daten) vor der Aufnahme ins Wiki, per Regex und ML-Erkennung. | 2026-08-29 |
|
|
||||||
| [[Forgetting]] | pattern | Muster zur Wissensbindung, das selten abgerufene Fakten schrittweise zurückstuft, modelliert nach der Ebbinghausschen Vergessenskurve. | 2026-08-29 |
|
|
||||||
| [[Graph Traversal]] | pattern | Verfahren, verbundene Entities im Wissensgraphen über typisierte Beziehungen (uses, depends-on, contradicts, caused) zu finden und strukturelle Fragen zu beantworten. | 2026-08-29 |
|
|
||||||
| [[Green Suite Blind Spot]] | problem | Defekt, der eine vollstaendig gruene Testsuite ueberlebt, weil nie ein Test das richtige Verhalten behauptet hat - belegt an drei prio/1-2-Defekten (Round-Trip, Zitat-Notation-als-Code, Zitat-Limit) | 2026-08-31 |
|
|
||||||
| [[Hooks]] | workflow | Mechanismus von Event-Listenern, der bei Wiki-Lebenszyklusereignissen wie Quellen-Ingest, Seitenänderung und Sitzungsende automatisch Aktionen auslöst. | 2026-08-29 |
|
| [[Hooks]] | workflow | Mechanismus von Event-Listenern, der bei Wiki-Lebenszyklusereignissen wie Quellen-Ingest, Seitenänderung und Sitzungsende automatisch Aktionen auslöst. | 2026-08-29 |
|
||||||
| [[Hybrid Search]] | architecture | Multimodale Suche, die BM25-Schlüsselwortabgleich, Vektor-Embeddings und Graph Traversal verbindet, um Wissensabruf im Wiki skalierbar zu machen. | 2026-08-29 |
|
|
||||||
| [[Implementation Spectrum]] | architecture | Modularer Einführungspfad für die Funktionen von LLM Wiki v2, vom minimal tragfähigen Wiki bis zur vollen Umsetzung mit Automatisierung und Governance. | 2026-08-29 |
|
|
||||||
| [[Index Scaling]] | workflow | Skalierungsregeln für Indexseiten: Tabellenabschnitte ab 50 Einträgen teilen, ab 200 Seiten _meta/topic-map.md anlegen | 2026-08-29 |
|
| [[Index Scaling]] | workflow | Skalierungsregeln für Indexseiten: Tabellenabschnitte ab 50 Einträgen teilen, ab 200 Seiten _meta/topic-map.md anlegen | 2026-08-29 |
|
||||||
| [[Issue Label Scheme]] | decision | Pflicht-Labelschema fuer das Gitea-Board: vier Achsen (area/kind/prio/size) plus seit 2026-09-04 drei optionale status/-Flags, darunter status/incoming fuer unausgearbeitete Stubs, die die Vier-Achsen-Pflicht aussetzen statt sie zu ergaenzen; die Regel liegt in instructions/dev/, weil sie keine ausgelieferte Instanz erreichen darf | 2026-09-04 |
|
|
||||||
| [[Iteration and Cost Limits]] | workflow | Im Code durchgesetzte Obergrenze von 60 wikitool-Aufrufen je Session, Loop-Breaker bei 3 identischen Wiederholungen, Slot-Erstattung, ein gemessenes Kalibrierungsband, und Retrieval sowie der MCP-Leseserver bleiben ausgenommen | 2026-09-02 |
|
| [[Iteration and Cost Limits]] | workflow | Im Code durchgesetzte Obergrenze von 60 wikitool-Aufrufen je Session, Loop-Breaker bei 3 identischen Wiederholungen, Slot-Erstattung, ein gemessenes Kalibrierungsband, und Retrieval sowie der MCP-Leseserver bleiben ausgenommen | 2026-09-02 |
|
||||||
| [[KB Migration]] | workflow | Migration des KB-Inhalts entlang einer geordneten Versionskette; abgegrenzt gegen offene Instanz-Aktionen, die in den doctor-Check gehoeren statt in die Kette | 2026-08-31 |
|
| [[KB Migration]] | workflow | Migration des KB-Inhalts entlang einer geordneten Versionskette; abgegrenzt gegen offene Instanz-Aktionen, die in den doctor-Check gehoeren statt in die Kette | 2026-08-31 |
|
||||||
| [[KB Stack Versioning]] | decision | Semantische Versionierung des Wiki-Stacks: VERSION beschreibt die Maschinerie, Kompatibilitaet (Drop-in-Ersatz) und Inhaltsmigration sind seit 2.5.0 getrennte, unabhaengig geprueft Fragen | 2026-09-02 |
|
|
||||||
| [[Knowledge Compounding]] | workflow | Effekt, bei dem Wissen im Wiki an Wert gewinnt, weil jede neue Quelle an bestehende, untereinander verwiesene Seiten anknüpft und sie ergänzt. | 2026-08-29 |
|
| [[Knowledge Compounding]] | workflow | Effekt, bei dem Wissen im Wiki an Wert gewinnt, weil jede neue Quelle an bestehende, untereinander verwiesene Seiten anknüpft und sie ergänzt. | 2026-08-29 |
|
||||||
| [[Knowledge Graph]] | architecture | Typisierte Schicht aus Entities und Beziehungen über den Wiki-Seiten, die eine reichere Wissensdarstellung und graphbasierte Abfragen ermöglicht. | 2026-08-29 |
|
|
||||||
| [[Lint Workflow]] | workflow | Deterministischer Health-Check rund um wikitool lint; seit 1.7.2 maskiert es Code vor dem Notation-Match und zaehlt Zitat-Bloecke statt Zeilen | 2026-09-01 |
|
| [[Lint Workflow]] | workflow | Deterministischer Health-Check rund um wikitool lint; seit 1.7.2 maskiert es Code vor dem Notation-Match und zaehlt Zitat-Bloecke statt Zeilen | 2026-09-01 |
|
||||||
| [[LLM Wiki Pattern]] | architecture | Methodik für persönliches Wissensmanagement, bei der ein LLM aus Rohquellen ein dauerhaftes Wiki aufbaut - Wissen wird kompiliert statt per RAG neu hergeleitet. | 2026-08-29 |
|
|
||||||
| [[Mass-Update Gate]] | workflow | Mass-Update Gate: publish endet mit 42 (Freigabe durch den Menschen noetig) ab 10 gezaehlten Dateien; generierte Dateien und work/ werden committet\, aber seit 1.5.0 nicht gezaehlt; freigegeben per --confirm <token> | 2026-09-01 |
|
| [[Mass-Update Gate]] | workflow | Mass-Update Gate: publish endet mit 42 (Freigabe durch den Menschen noetig) ab 10 gezaehlten Dateien; generierte Dateien und work/ werden committet\, aber seit 1.5.0 nicht gezaehlt; freigegeben per --confirm <token> | 2026-09-01 |
|
||||||
|
| [[Multi-Agent Collaboration]] | workflow | Wissensmanagement mit mehreren Agenten; erweitert das LLM-Wiki-Muster um Mesh Sync, die Trennung von geteiltem und privatem Wissen und leichtgewichtige Arbeitskoordination. | 2026-08-29 |
|
||||||
|
| [[Privacy and Governance]] | workflow | Rahmenwerk zur Absicherung von Wiki-Inhalten über Datenfilterung beim Ingest, Audit-Trail-Protokollierung und umkehrbare Massenoperationen. | 2026-08-29 |
|
||||||
|
| [[Publish-Remote Gate]] | workflow | Drittes, im Code durchgesetztes Gate: publish bricht mit Exit 42 ab, wenn die aufgeloeste Push-URL nicht in einer optionalen, gitignoreten Allowlist steht; doctor benennt seit 2026-09-02 den Gate-Zustand statt nur die Dateiexistenz | 2026-09-02 |
|
||||||
|
| [[Quality and Self-Correction]] | workflow | Automatische Qualitätssicherung für Wikis mit Inhaltsbewertung, Selbstheilung und Widerspruchserkennung. | 2026-08-29 |
|
||||||
|
| [[Semantic Lint Automation]] | workflow | Maschinelle Heuristiken zur Priorisierung der semantischen Prüfung: veraltete Aussagen, hohe Änderungsdichte und schwache Verlinkung | 2026-08-29 |
|
||||||
|
| [[Session Orientation]] | workflow | Verbindliche Vorabprüfung, die vor Query- und Update-Operationen einen Kontextbericht erzeugt (Index, jüngste Logs, Umfang) | 2026-08-29 |
|
||||||
|
| [[Split Merge Reclassify]] | workflow | Eigene Befehle zum Teilen, Zusammenführen und Umklassifizieren von Seiten, mit automatischer Korrektur von Links und Frontmatter | 2026-08-29 |
|
||||||
|
| [[Split Threshold]] | workflow | Maximale Seitengröße, ab der eine Aufteilung empfohlen wird (Farza: >120-150 Zeilen, Pascalandy: 200 Zeilen) | 2026-08-29 |
|
||||||
|
| [[Stub Threshold]] | workflow | Mindestumfang, ab dem eine Wiki-Seite nicht mehr als Stub gilt (Farza: ≥3 Sätze oder 15 Zeilen) | 2026-08-29 |
|
||||||
|
| [[Supersession]] | workflow | Ablösen alten Wissens durch neue, widersprechende Information; gibt dem Wiki eine Versionierung mit ausdrücklicher Verknüpfung und Erhalt der Historie. | 2026-08-29 |
|
||||||
|
| [[User Management]] | workflow | Linux-Ablauf zum Anlegen, Ändern, Überwachen und Löschen von Benutzerkonten mit useradd, usermod und userdel, samt Gruppenverwaltung und sudoers-Konfiguration. | 2026-08-29 |
|
||||||
|
| [[Workflow Extraction]] | workflow | Herauslösen von Workflow-Abschnitten aus monolithischer Dokumentation | 2026-09-01 |
|
||||||
|
| [[Workflow Orchestration]] | workflow | Orchestrierte Einzelbefehle für vollständige Operationen (ingest run, lint run, update run) mit Dry-Run-Vorschau vor dem Schreiben | 2026-08-29 |
|
||||||
|
|
||||||
|
## Architekturen
|
||||||
|
|
||||||
|
| Page | Type | Summary | Last Modified |
|
||||||
|
|------|------|---------|----------------|
|
||||||
|
| [[Consolidation Tiers]] | architecture | Hierarchische Speicherarchitektur, die Informationen durch zunehmend verdichtete Schichten vom Working Memory bis zum Semantic und Procedural Memory befördert. | 2026-08-29 |
|
||||||
|
| [[Context Isolation]] | architecture | Grundsatz, für jede Aufgabe nur den jeweils benötigten Kontext zu laden | 2026-08-29 |
|
||||||
|
| [[Cross-platform Agent Skills]] | architecture | Architektur fuer Agent-Skills, die ueber mehrere LLM-Werkzeuge hinweg funktionieren; in Chemenu selbst am 2026-08-04 umgesetzt und ueberprueft | 2026-09-01 |
|
||||||
|
| [[Episodic Memory]] | architecture | Speicherschicht für verdichtete Sitzungszusammenfassungen und Befunde; Brücke zwischen rohem Working Memory und langlebigem Semantic Memory. | 2026-08-29 |
|
||||||
|
| [[Hybrid Search]] | architecture | Multimodale Suche, die BM25-Schlüsselwortabgleich, Vektor-Embeddings und Graph Traversal verbindet, um Wissensabruf im Wiki skalierbar zu machen. | 2026-08-29 |
|
||||||
|
| [[Implementation Spectrum]] | architecture | Modularer Einführungspfad für die Funktionen von LLM Wiki v2, vom minimal tragfähigen Wiki bis zur vollen Umsetzung mit Automatisierung und Governance. | 2026-08-29 |
|
||||||
|
| [[Knowledge Graph]] | architecture | Typisierte Schicht aus Entities und Beziehungen über den Wiki-Seiten, die eine reichere Wissensdarstellung und graphbasierte Abfragen ermöglicht. | 2026-08-29 |
|
||||||
|
| [[LLM Wiki Pattern]] | architecture | Methodik für persönliches Wissensmanagement, bei der ein LLM aus Rohquellen ein dauerhaftes Wiki aufbaut - Wissen wird kompiliert statt per RAG neu hergeleitet. | 2026-08-29 |
|
||||||
| [[MCP-Leseserver]] | architecture | Zweiter Konsument von chemenu ueber MCP: search/types/describe_type/lint/status auf chemenu.api.Corpus, strukturell ohne Schreibpfad, jede Antwort trage einen Commit-Stempel. | 2026-09-02 |
|
| [[MCP-Leseserver]] | architecture | Zweiter Konsument von chemenu ueber MCP: search/types/describe_type/lint/status auf chemenu.api.Corpus, strukturell ohne Schreibpfad, jede Antwort trage einen Commit-Stempel. | 2026-09-02 |
|
||||||
| [[Memory Lifecycle]] | architecture | Architektur des Wissenslebenszyklus mit Confidence Scoring, Supersession, Forgetting und Consolidation Tiers zur Pflege von Fakten über die Zeit. | 2026-08-29 |
|
| [[Memory Lifecycle]] | architecture | Architektur des Wissenslebenszyklus mit Confidence Scoring, Supersession, Forgetting und Consolidation Tiers zur Pflege von Fakten über die Zeit. | 2026-08-29 |
|
||||||
| [[Mesh Sync]] | pattern | Abgleichsmechanismus, der Beobachtungen paralleler Agenten in ein gemeinsames Wiki überführt; Last-Write-Wins mit Konflikterkennung und manuellem Eingriff. | 2026-08-29 |
|
|
||||||
| [[Modbus]] | protocol | Industrielles Kommunikationsprotokoll von 1979 zur Anbindung speicherprogrammierbarer Steuerungen und Geräte über serielle oder TCP-Netze. | 2026-08-29 |
|
|
||||||
| [[Multi-Agent Collaboration]] | workflow | Wissensmanagement mit mehreren Agenten; erweitert das LLM-Wiki-Muster um Mesh Sync, die Trennung von geteiltem und privatem Wissen und leichtgewichtige Arbeitskoordination. | 2026-08-29 |
|
|
||||||
| [[Naming Convention Conflict]] | problem | Widerspruch zwischen README.md (kebab-case) und AGENTS.md (lesbar mit Leerzeichen), der zu Drift bei der Validierung führt | 2026-08-29 |
|
|
||||||
| [[OKF Compatibility]] | architecture | Optionale Kompatibilität zum Open Knowledge Framework als Export- und Prüfmodus, ohne das interne Modell zu ersetzen | 2026-08-29 |
|
| [[OKF Compatibility]] | architecture | Optionale Kompatibilität zum Open Knowledge Framework als Export- und Prüfmodus, ohne das interne Modell zu ersetzen | 2026-08-29 |
|
||||||
| [[Optional Instance Context File]] | architecture | Muster fuer eine Datei, die eine Instanz ueber ihre Umgebung informiert, ohne Betriebsvoraussetzung zu sein: Health-Check meldet ohne zu scheitern, pro Checkout statt pro Repo | 2026-08-31 |
|
| [[Optional Instance Context File]] | architecture | Muster fuer eine Datei, die eine Instanz ueber ihre Umgebung informiert, ohne Betriebsvoraussetzung zu sein: Health-Check meldet ohne zu scheitern, pro Checkout statt pro Repo | 2026-08-31 |
|
||||||
| [[Personalization Plane]] | architecture | Schicht fuer Instanz-Identitaet: USER.md/SOUL.md werden als Template ausgeliefert, im Setup-Interview woertlich befuellt und vom doctor-Check auf fehlend wie unbefuellt geprueft | 2026-08-31 |
|
| [[Personalization Plane]] | architecture | Schicht fuer Instanz-Identitaet: USER.md/SOUL.md werden als Template ausgeliefert, im Setup-Interview woertlich befuellt und vom doctor-Check auf fehlend wie unbefuellt geprueft | 2026-08-31 |
|
||||||
| [[Privacy and Governance]] | workflow | Rahmenwerk zur Absicherung von Wiki-Inhalten über Datenfilterung beim Ingest, Audit-Trail-Protokollierung und umkehrbare Massenoperationen. | 2026-08-29 |
|
|
||||||
| [[Procedural Memory]] | architecture | Langlebigste Speicherschicht für Abläufe, Muster, bewährte Vorgehensweisen und Rezepte, gewonnen aus wiederholten semantischen Beobachtungen. | 2026-08-29 |
|
| [[Procedural Memory]] | architecture | Langlebigste Speicherschicht für Abläufe, Muster, bewährte Vorgehensweisen und Rezepte, gewonnen aus wiederholten semantischen Beobachtungen. | 2026-08-29 |
|
||||||
| [[Publish-Remote Gate]] | workflow | Drittes, im Code durchgesetztes Gate: publish bricht mit Exit 42 ab, wenn die aufgeloeste Push-URL nicht in einer optionalen, gitignoreten Allowlist steht; doctor benennt seit 2026-09-02 den Gate-Zustand statt nur die Dateiexistenz | 2026-09-02 |
|
|
||||||
| [[Quality and Self-Correction]] | workflow | Automatische Qualitätssicherung für Wikis mit Inhaltsbewertung, Selbstheilung und Widerspruchserkennung. | 2026-08-29 |
|
|
||||||
| [[Quality Scoring]] | pattern | Quantitative Bewertung aller vom LLM geschriebenen Inhalte nach struktureller Qualität, Vollständigkeit der Quellenangaben, Konsistenz mit dem Wiki und Themenabdeckung. | 2026-08-29 |
|
|
||||||
| [[RAG]] | architecture | Architekturmuster, bei dem LLMs die Generierung um Dokumente aus einer Wissensbasis anreichern. | 2026-08-29 |
|
| [[RAG]] | architecture | Architekturmuster, bei dem LLMs die Generierung um Dokumente aus einer Wissensbasis anreichern. | 2026-08-29 |
|
||||||
| [[Reciprocal Rank Fusion]] | pattern | Verfahren, das Ergebnislisten mehrerer Suchmodalitäten zu einem gemeinsamen Ranking verbindet, ohne Gewichte zwischen den Modalitäten justieren zu müssen. | 2026-08-29 |
|
|
||||||
| [[Scale Ceiling]] | architecture | Punkt, ab dem Wiki-Ansätze mit einem einzigen Kontext qualitativ abfallen | 2026-09-01 |
|
| [[Scale Ceiling]] | architecture | Punkt, ab dem Wiki-Ansätze mit einem einzigen Kontext qualitativ abfallen | 2026-09-01 |
|
||||||
| [[Self-Healing]] | pattern | Automatisches Beheben von Mängeln, die beim Lint auffallen: verwaiste Seiten, veraltete Aussagen, kaputte Links und Formatverstöße. | 2026-08-29 |
|
|
||||||
| [[Semantic Lint Automation]] | workflow | Maschinelle Heuristiken zur Priorisierung der semantischen Prüfung: veraltete Aussagen, hohe Änderungsdichte und schwache Verlinkung | 2026-08-29 |
|
|
||||||
| [[Semantic Memory]] | architecture | Langlebige Schicht für sitzungsübergreifend verdichtete Fakten aus mehreren Episoden, mit höherer Konfidenz und stärkerer Verdichtung als episodische Erinnerungen. | 2026-08-29 |
|
| [[Semantic Memory]] | architecture | Langlebige Schicht für sitzungsübergreifend verdichtete Fakten aus mehreren Episoden, mit höherer Konfidenz und stärkerer Verdichtung als episodische Erinnerungen. | 2026-08-29 |
|
||||||
| [[Session Orientation]] | workflow | Verbindliche Vorabprüfung, die vor Query- und Update-Operationen einen Kontextbericht erzeugt (Index, jüngste Logs, Umfang) | 2026-08-29 |
|
|
||||||
| [[Shared vs Private]] | pattern | Abgrenzung persönlicher Beobachtungen (privat) von Team- und Projektwissen (geteilt), mit Regeln zum Hochstufen geprüften Wissens. | 2026-08-29 |
|
|
||||||
| [[Split Merge Reclassify]] | workflow | Eigene Befehle zum Teilen, Zusammenführen und Umklassifizieren von Seiten, mit automatischer Korrektur von Links und Frontmatter | 2026-08-29 |
|
|
||||||
| [[Split Threshold]] | workflow | Maximale Seitengröße, ab der eine Aufteilung empfohlen wird (Farza: >120-150 Zeilen, Pascalandy: 200 Zeilen) | 2026-08-29 |
|
|
||||||
| [[SSD TRIM]] | protocol | Datenträgerbefehl, mit dem SSDs ungenutzte Blöcke zurückgewinnen - für gleichbleibende Leistung und längere Lebensdauer. | 2026-08-29 |
|
|
||||||
| [[Structural Enforcement over Documented Rule]] | decision | Entscheidung, eine wiederkehrende Fehlerregel in die Ausfuehrung einzubauen statt sie aufzuschreiben - Rangfolge erzwingen vor melden vor erinnern, belegt an einer Regel, die gelesen wurde und nicht wirkte | 2026-08-31 |
|
|
||||||
| [[Stub Threshold]] | workflow | Mindestumfang, ab dem eine Wiki-Seite nicht mehr als Stub gilt (Farza: ≥3 Sätze oder 15 Zeilen) | 2026-08-29 |
|
|
||||||
| [[Supersession]] | workflow | Ablösen alten Wissens durch neue, widersprechende Information; gibt dem Wiki eine Versionierung mit ausdrücklicher Verknüpfung und Erhalt der Historie. | 2026-08-29 |
|
|
||||||
| [[Three-Layer Architecture]] | architecture | Strukturmodell des LLM-Wiki-Musters mit drei Schichten: unveränderliche Rohquellen, vom LLM gepflegtes Wiki und Schemakonfiguration, die Knowledge Compounding trägt. | 2026-08-29 |
|
| [[Three-Layer Architecture]] | architecture | Strukturmodell des LLM-Wiki-Musters mit drei Schichten: unveränderliche Rohquellen, vom LLM gepflegtes Wiki und Schemakonfiguration, die Knowledge Compounding trägt. | 2026-08-29 |
|
||||||
| [[Token Economics]] | architecture | Kosten- und Effizienzüberlegungen zum Tokenverbrauch von LLMs - die genannten Werte 5-8x/61%/100% sind unbestätigt, keine gesicherten Fakten | 2026-09-01 |
|
| [[Token Economics]] | architecture | Kosten- und Effizienzüberlegungen zum Tokenverbrauch von LLMs - die genannten Werte 5-8x/61%/100% sind unbestätigt, keine gesicherten Fakten | 2026-09-01 |
|
||||||
|
| [[Working Memory]] | architecture | Kurzlebige Speicherschicht für jüngste Beobachtungen und vorläufige Befunde vor der Verdichtung; niedrigste Konfidenz, keine Verdichtung, wird zum Sitzungsende erneuert. | 2026-08-29 |
|
||||||
|
|
||||||
|
## Entscheidungen
|
||||||
|
|
||||||
|
| Page | Type | Summary | Last Modified |
|
||||||
|
|------|------|---------|----------------|
|
||||||
|
| [[Delete Rather Than Anonymize]] | decision | Private Korpusinhalte per Loeschung entfernen statt zu anonymisieren: ein Seitentitel ist der einzige Identifier eines Wikis, Umbenennen ist die volle page-lifecycle-Prozedur je Seite, Loeschen ist ein unterstuetztes Kommando. | 2026-09-01 |
|
||||||
|
| [[Denylist over Allowlist]] | decision | Entscheidung, schreibbare Felder als Schema minus kurzer Sperrliste zu bestimmen statt als gepflegte Positivliste, weil die Positivliste eine zweite Kopie des Schemas waere | 2026-08-31 |
|
||||||
|
| [[Diff-Reviewable Agent Edits]] | decision | Entscheidung, Dateiaenderungen ueber Edit/Write statt ueber Shell-Heredocs zu fahren, weil nur das erste eine pruefbare Diff hinterlaesst | 2026-08-31 |
|
||||||
|
| [[Dual Licensing by File Plan]] | decision | Ein Repo mit Code- und Inhaltsanteil erhaelt zwei Lizenzen; die Grenze zwischen ihnen ist kein zweiter, gepflegter Pfadkatalog, sondern der ohnehin vorhandene Dateiplan des Distributionswerkzeugs. | 2026-09-01 |
|
||||||
|
| [[Issue Label Scheme]] | decision | Pflicht-Labelschema fuer das Gitea-Board: vier Achsen (area/kind/prio/size) plus seit 2026-09-04 drei optionale status/-Flags, darunter status/incoming fuer unausgearbeitete Stubs, die die Vier-Achsen-Pflicht aussetzen statt sie zu ergaenzen; die Regel liegt in instructions/dev/, weil sie keine ausgelieferte Instanz erreichen darf | 2026-09-04 |
|
||||||
|
| [[KB Stack Versioning]] | decision | Semantische Versionierung des Wiki-Stacks: VERSION beschreibt die Maschinerie, Kompatibilitaet (Drop-in-Ersatz) und Inhaltsmigration sind seit 2.5.0 getrennte, unabhaengig geprueft Fragen | 2026-09-02 |
|
||||||
|
| [[Structural Enforcement over Documented Rule]] | decision | Entscheidung, eine wiederkehrende Fehlerregel in die Ausfuehrung einzubauen statt sie aufzuschreiben - Rangfolge erzwingen vor melden vor erinnern, belegt an einer Regel, die gelesen wurde und nicht wirkte | 2026-08-31 |
|
||||||
|
|
||||||
|
## Muster
|
||||||
|
|
||||||
|
| Page | Type | Summary | Last Modified |
|
||||||
|
|------|------|---------|----------------|
|
||||||
|
| [[Audit Trail]] | pattern | Unveränderliches chronologisches Log aller Wiki-Operationen (Ingest, Bearbeitung, Löschung, Abfrage) mit Zeitstempel, Akteur, Ziel und Änderungsbeschreibung. | 2026-08-29 |
|
||||||
|
| [[BM25]] | pattern | Schlüsselwortbasiertes Retrieval-Verfahren, das über Termfrequenz, inverse Dokumentfrequenz und Stemming exakte oder teilweise Übereinstimmungen findet. | 2026-08-29 |
|
||||||
|
| [[Command Round-Trip Integrity]] | pattern | Anforderung, dass zwei Befehle auf derselben Datei in jeder Reihenfolge zusammenpassen und jeder erzeugte Zustand einen Gegenbefehl hat - 2026-08-31 in wikitool zweimal verletzt | 2026-08-31 |
|
||||||
|
| [[Confidence Scoring]] | pattern | Mechanismus, der faktischen Aussagen quantitative Werte nach Quellenzahl, Aktualität, Qualität und Bestätigung zuweist, um gut gestütztes Wissen zu erkennen. | 2026-08-29 |
|
||||||
|
| [[Contradiction Resolution]] | pattern | Automatisches Erkennen und Auflösen widersprüchlicher Aussagen anhand von Konfidenz, Aktualität und Autorität der Quelle. | 2026-08-29 |
|
||||||
|
| [[Entity Extraction]] | pattern | Erkennen und Strukturieren von Entities (Personen, Projekte, Bibliotheken, Concepts, Dateien, Entscheidungen, Systeme, Werkzeuge) samt typspezifischer Attribute aus Rohquellen. | 2026-08-29 |
|
||||||
|
| [[Filter on Ingest]] | pattern | Automatisches Erkennen und Entfernen sensibler Daten (API-Schlüssel, Token, Credentials, personenbezogene Daten) vor der Aufnahme ins Wiki, per Regex und ML-Erkennung. | 2026-08-29 |
|
||||||
|
| [[Forgetting]] | pattern | Muster zur Wissensbindung, das selten abgerufene Fakten schrittweise zurückstuft, modelliert nach der Ebbinghausschen Vergessenskurve. | 2026-08-29 |
|
||||||
|
| [[Graph Traversal]] | pattern | Verfahren, verbundene Entities im Wissensgraphen über typisierte Beziehungen (uses, depends-on, contradicts, caused) zu finden und strukturelle Fragen zu beantworten. | 2026-08-29 |
|
||||||
|
| [[Mesh Sync]] | pattern | Abgleichsmechanismus, der Beobachtungen paralleler Agenten in ein gemeinsames Wiki überführt; Last-Write-Wins mit Konflikterkennung und manuellem Eingriff. | 2026-08-29 |
|
||||||
|
| [[Quality Scoring]] | pattern | Quantitative Bewertung aller vom LLM geschriebenen Inhalte nach struktureller Qualität, Vollständigkeit der Quellenangaben, Konsistenz mit dem Wiki und Themenabdeckung. | 2026-08-29 |
|
||||||
|
| [[Reciprocal Rank Fusion]] | pattern | Verfahren, das Ergebnislisten mehrerer Suchmodalitäten zu einem gemeinsamen Ranking verbindet, ohne Gewichte zwischen den Modalitäten justieren zu müssen. | 2026-08-29 |
|
||||||
|
| [[Self-Healing]] | pattern | Automatisches Beheben von Mängeln, die beim Lint auffallen: verwaiste Seiten, veraltete Aussagen, kaputte Links und Formatverstöße. | 2026-08-29 |
|
||||||
|
| [[Shared vs Private]] | pattern | Abgrenzung persönlicher Beobachtungen (privat) von Team- und Projektwissen (geteilt), mit Regeln zum Hochstufen geprüften Wissens. | 2026-08-29 |
|
||||||
| [[Typed Relationships]] | pattern | Verwendung semantisch aussagekräftiger Beziehungstypen (uses, depends-on, contradicts, caused, fixed, supersedes, replaces) statt undifferenzierter Wikilinks. | 2026-08-29 |
|
| [[Typed Relationships]] | pattern | Verwendung semantisch aussagekräftiger Beziehungstypen (uses, depends-on, contradicts, caused, fixed, supersedes, replaces) statt undifferenzierter Wikilinks. | 2026-08-29 |
|
||||||
| [[User Management]] | workflow | Linux-Ablauf zum Anlegen, Ändern, Überwachen und Löschen von Benutzerkonten mit useradd, usermod und userdel, samt Gruppenverwaltung und sudoers-Konfiguration. | 2026-08-29 |
|
|
||||||
| [[Vector Search]] | pattern | Semantische Ähnlichkeitssuche über Embedding-Vektoren, die inhaltlich verwandte Seiten auch ohne exakte Schlüsselwortübereinstimmung findet. | 2026-08-29 |
|
| [[Vector Search]] | pattern | Semantische Ähnlichkeitssuche über Embedding-Vektoren, die inhaltlich verwandte Seiten auch ohne exakte Schlüsselwortübereinstimmung findet. | 2026-08-29 |
|
||||||
| [[Work Coordination]] | pattern | Leichtgewichtige Erfassung von Aufgabenstatus (in Arbeit, blockiert, erledigt, prüfbedürftig) und Zuweisung, um Doppelarbeit bei mehreren Agenten zu vermeiden. | 2026-08-29 |
|
| [[Work Coordination]] | pattern | Leichtgewichtige Erfassung von Aufgabenstatus (in Arbeit, blockiert, erledigt, prüfbedürftig) und Zuweisung, um Doppelarbeit bei mehreren Agenten zu vermeiden. | 2026-08-29 |
|
||||||
| [[Workflow Extraction]] | workflow | Herauslösen von Workflow-Abschnitten aus monolithischer Dokumentation | 2026-09-01 |
|
|
||||||
| [[Workflow Orchestration]] | workflow | Orchestrierte Einzelbefehle für vollständige Operationen (ingest run, lint run, update run) mit Dry-Run-Vorschau vor dem Schreiben | 2026-08-29 |
|
## Problemstellungen
|
||||||
| [[Working Memory]] | architecture | Kurzlebige Speicherschicht für jüngste Beobachtungen und vorläufige Befunde vor der Verdichtung; niedrigste Konfidenz, keine Verdichtung, wird zum Sitzungsende erneuert. | 2026-08-29 |
|
|
||||||
|
| Page | Type | Summary | Last Modified |
|
||||||
|
|------|------|---------|----------------|
|
||||||
|
| [[Ambient Environment Dependency]] | problem | Fehlerklasse, in der ein Test gruen ist, weil die Maschine zufaellig passt statt weil der Code stimmt - abgegrenzt gegen den Green Suite Blind Spot, belegt an vier Faellen unter Gitea-Issue #8 | 2026-08-31 |
|
||||||
|
| [[Detect-Repair Asymmetry]] | problem | Werkzeugluecke, in der ein Check einen Defekt zuverlaessig meldet, aber kein Befehl ihn behebt - womit die Handeditierung der einzige verbleibende Ausweg ist | 2026-08-31 |
|
||||||
|
| [[Green Suite Blind Spot]] | problem | Defekt, der eine vollstaendig gruene Testsuite ueberlebt, weil nie ein Test das richtige Verhalten behauptet hat - belegt an drei prio/1-2-Defekten (Round-Trip, Zitat-Notation-als-Code, Zitat-Limit) | 2026-08-31 |
|
||||||
|
| [[Naming Convention Conflict]] | problem | Widerspruch zwischen README.md (kebab-case) und AGENTS.md (lesbar mit Leerzeichen), der zu Drift bei der Validierung führt | 2026-08-29 |
|
||||||
| [[Write-Once Frontmatter Fields]] | problem | Defektklasse, in der ein Feld nur beim Anlegen der Seite schreibbar ist und danach unerreichbar bleibt, weil kein Mutationsbefehl es kennt und new nicht idempotent ist | 2026-08-31 |
|
| [[Write-Once Frontmatter Fields]] | problem | Defektklasse, in der ein Feld nur beim Anlegen der Seite schreibbar ist und danach unerreichbar bleibt, weil kein Mutationsbefehl es kennt und new nicht idempotent ist | 2026-08-31 |
|
||||||
|
|
||||||
|
## Protokolle
|
||||||
|
|
||||||
|
| Page | Type | Summary | Last Modified |
|
||||||
|
|------|------|---------|----------------|
|
||||||
|
| [[CPPC]] | protocol | Hardwareschnittstelle Collaborative Processor Performance Control für feingranulares CPU-Power-Management zwischen Betriebssystem und AMD-Prozessor. | 2026-08-29 |
|
||||||
|
| [[Modbus]] | protocol | Industrielles Kommunikationsprotokoll von 1979 zur Anbindung speicherprogrammierbarer Steuerungen und Geräte über serielle oder TCP-Netze. | 2026-08-29 |
|
||||||
|
| [[SSD TRIM]] | protocol | Datenträgerbefehl, mit dem SSDs ungenutzte Blöcke zurückgewinnen - für gleichbleibende Leistung und längere Lebensdauer. | 2026-08-29 |
|
||||||
|
|
||||||
|
|||||||
+12
-1
@@ -18,7 +18,7 @@ The page tables live in a generated `INDEX.md` inside each collection, linked be
|
|||||||
- **Concepts:** 80
|
- **Concepts:** 80
|
||||||
- **Entities:** 72
|
- **Entities:** 72
|
||||||
- **Sources:** 29
|
- **Sources:** 29
|
||||||
- **Last Updated:** 2026-09-04
|
- **Last Updated:** 2026-09-08
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -31,6 +31,17 @@ The page tables live in a generated `INDEX.md` inside each collection, linked be
|
|||||||
| `entities/` | 72 | [entities/INDEX.md](entities/INDEX.md) |
|
| `entities/` | 72 | [entities/INDEX.md](entities/INDEX.md) |
|
||||||
| `sources/` | 29 | [sources/INDEX.md](sources/INDEX.md) |
|
| `sources/` | 29 | [sources/INDEX.md](sources/INDEX.md) |
|
||||||
|
|
||||||
|
### concepts/
|
||||||
|
|
||||||
|
| Area | Pages | Index |
|
||||||
|
|------|------:|-------|
|
||||||
|
| Abläufe | 28 | [concepts/INDEX.md#abläufe](concepts/INDEX.md#abläufe) |
|
||||||
|
| Architekturen | 20 | [concepts/INDEX.md#architekturen](concepts/INDEX.md#architekturen) |
|
||||||
|
| Entscheidungen | 7 | [concepts/INDEX.md#entscheidungen](concepts/INDEX.md#entscheidungen) |
|
||||||
|
| Muster | 17 | [concepts/INDEX.md#muster](concepts/INDEX.md#muster) |
|
||||||
|
| Problemstellungen | 5 | [concepts/INDEX.md#problemstellungen](concepts/INDEX.md#problemstellungen) |
|
||||||
|
| Protokolle | 3 | [concepts/INDEX.md#protokolle](concepts/INDEX.md#protokolle) |
|
||||||
|
|
||||||
### entities/
|
### entities/
|
||||||
|
|
||||||
| Area | Pages | Index |
|
| Area | Pages | Index |
|
||||||
|
|||||||
@@ -155,3 +155,9 @@ Drittes status/-Flag status/incoming ergaenzt (Gitea #63): status/-Tabelle auf d
|
|||||||
wikitool move --reconcile hat die drei nach kb/entities/projects/{kfchou,vanillaflava,yugasun}/*.md verschachtelten Seiten (wiki-skills, wiki-skills-vanillaflava, llm-wiki-skills) nach kb/entities/projects/ hochgezogen und die drei geleerten Owner-Verzeichnisse entfernt. index rebuild und sources rebuild-index liefen danach; lint meldet weder misplaced_pages noch nested_pages noch duplicate_titles; migrate verify --from HEAD bestaetigt compared=182, added=0, removed=0, alle drei als moved.
|
wikitool move --reconcile hat die drei nach kb/entities/projects/{kfchou,vanillaflava,yugasun}/*.md verschachtelten Seiten (wiki-skills, wiki-skills-vanillaflava, llm-wiki-skills) nach kb/entities/projects/ hochgezogen und die drei geleerten Owner-Verzeichnisse entfernt. index rebuild und sources rebuild-index liefen danach; lint meldet weder misplaced_pages noch nested_pages noch duplicate_titles; migrate verify --from HEAD bestaetigt compared=182, added=0, removed=0, alle drei als moved.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## [2026-09-08] move | kb/concepts/ bekommt Areas: 80 Seiten in ihre concept_type-Verzeichnisse (#59)
|
||||||
|
|
||||||
|
types/concept.md deklariert ein layout: fuer alle sechs concept_type-Werte; wikitool move --reconcile hat daraufhin alle 80 Seiten aus kb/concepts/ in ihre Area gezogen (architectures 20, patterns 17, protocols 3, workflows 28, decisions 7, problems 5). Kein Titel, kein Body, kein Frontmatter-Feld angefasst: migrate verify --from HEAD bestaetigt compared=182, added=0, removed=0, 80 moved, 0 findings. index rebuild erzeugt sechs Areas, keine ueber der Shard-Schwelle von 50, also kein eigener Shard - die Schwelle wirkt wieder als Schwelle statt als toter Wert. lint meldet danach weder misplaced_pages noch den neuen unsharded_collections-Befund.
|
||||||
|
|
||||||
|
---
|
||||||
|
|||||||
+1
-1
@@ -51,7 +51,7 @@ tools/wikitool <command> --help
|
|||||||
| `index rebuild [--dry-run]` | Regenerate the catalog from every page's frontmatter: `kb/index.md` becomes a map (statistics, one row per collection and per area, links to the shards) and the page tables are written to a generated `INDEX.md` in each collection. An area past 50 rows gets its own shard. Stale shards from removed collections/areas are deleted in the same pass |
|
| `index rebuild [--dry-run]` | Regenerate the catalog from every page's frontmatter: `kb/index.md` becomes a map (statistics, one row per collection and per area, links to the shards) and the page tables are written to a generated `INDEX.md` in each collection. An area past 50 rows gets its own shard. Stale shards from removed collections/areas are deleted in the same pass |
|
||||||
| `log append --op ingest\|query\|lint\|create\|update\|delete\|rename\|move --title "..." [--body "..."\|--body-file path]` | Append a formatted entry to `kb/log.md` |
|
| `log append --op ingest\|query\|lint\|create\|update\|delete\|rename\|move --title "..." [--body "..."\|--body-file path]` | Append a formatted entry to `kb/log.md` |
|
||||||
| `log status` | Read-only: count `ingest` entries logged since the last `lint` entry - the deterministic trigger behind the Maintenance Schedule's "every 10 sources" full-lint cadence |
|
| `log status` | Read-only: count `ingest` entries logged since the last `lint` entry - the deterministic trigger behind the Maintenance Schedule's "every 10 sources" full-lint cadence |
|
||||||
| `lint [--json] [--markdown out.md] [--full] [--fail-on-error]` | Structural + provenance checks: broken wikilinks, dangling frontmatter references, orphan pages, index drift, schema gaps, duplicate titles, title mismatches, pages nested more than one directory below their collection (hard - the generated catalog folds these into their area silently rather than merely reading it, see #57), uncovered raw files, broken `raw_files:` refs, raw files claimed by more than one source page, unmarked provenance, citation/frontmatter drift, unbalanced generated-region markers, edges whose label is missing or not authorised by the source collection's `outbound:` (both hard once `kb_version` has reached the release that introduced labelled edges - advisory below it, so a corpus mid-migration is not refused by the check measuring it), `see-also` edges whose reverse direction already carries a specific label (advisory only - redundant rather than wrong, and never migration-gated, since no version turns the redundancy into an error), quote-limit overages (>2 blockquoted lines/page, advisory only). Prints only the sections that found something and always writes the full report to `reports/Lint Report <date>.md` (or `--markdown`), naming the path - `--full` prints everything, `--json` prints the findings and writes nothing |
|
| `lint [--json] [--markdown out.md] [--full] [--fail-on-error]` | Structural + provenance checks: broken wikilinks, dangling frontmatter references, orphan pages, index drift, schema gaps, duplicate titles, title mismatches, pages nested more than one directory below their collection (hard - the generated catalog folds these into their area silently rather than merely reading it, see #57), uncovered raw files, broken `raw_files:` refs, raw files claimed by more than one source page, unmarked provenance, citation/frontmatter drift, unbalanced generated-region markers, edges whose label is missing or not authorised by the source collection's `outbound:` (both hard once `kb_version` has reached the release that introduced labelled edges - advisory below it, so a corpus mid-migration is not refused by the check measuring it), `see-also` edges whose reverse direction already carries a specific label (advisory only - redundant rather than wrong, and never migration-gated, since no version turns the redundancy into an error), a collection past the catalog's per-area shard threshold that has no areas to shard (advisory only - sharding is automatic but per *area*, so a collection nobody gave areas keeps one table however large it grows, #59; reported with the split its subtype field would produce, and only when that split puts every resulting area at or under the threshold, so a lopsided or small collection stays silent), quote-limit overages (>2 blockquoted lines/page, advisory only). Prints only the sections that found something and always writes the full report to `reports/Lint Report <date>.md` (or `--markdown`), naming the path - `--full` prints everything, `--json` prints the findings and writes nothing |
|
||||||
| `search ["<text>"] [--field <predicate> ...] [--kind/--subtype/--collection/--tag <v>] [--regex] [--limit N] [--sort [-]<field>] [--backend <name>] [--matches] [--json]` | Find pages in `kb/` without reading the index. Text search runs through a pluggable backend (`rg` today); `--field` predicates are evaluated on frontmatter - `f=v`, `f~substring`, `'f>=v'`, `'f:*'` (present), `'!f'` (absent), repeatable and ANDed. With no text this is a pure structured query. Results carry kind/summary/confidence so a hit can be judged without opening the page. A page whose frontmatter does not parse can match no positive predicate, so it is **named** rather than dropped: `--json` always carries an `unreadable` list of `{path, reason}` (usually empty), and the table form writes the same lines to stderr. `--regex` is applied by `rg` alone, whose engine is linear; the ranking boosts for title and summary are literal-containment only, so a non-literal pattern is ranked by match count. `rg` is killed after 30 s and reported as a failure. Read-only, and **exempt from the Iteration Budget Gate** |
|
| `search ["<text>"] [--field <predicate> ...] [--kind/--subtype/--collection/--tag <v>] [--regex] [--limit N] [--sort [-]<field>] [--backend <name>] [--matches] [--json]` | Find pages in `kb/` without reading the index. Text search runs through a pluggable backend (`rg` today); `--field` predicates are evaluated on frontmatter - `f=v`, `f~substring`, `'f>=v'`, `'f:*'` (present), `'!f'` (absent), repeatable and ANDed. With no text this is a pure structured query. Results carry kind/summary/confidence so a hit can be judged without opening the page. A page whose frontmatter does not parse can match no positive predicate, so it is **named** rather than dropped: `--json` always carries an `unreadable` list of `{path, reason}` (usually empty), and the table form writes the same lines to stderr. `--regex` is applied by `rg` alone, whose engine is linear; the ranking boosts for title and summary are literal-containment only, so a non-literal pattern is ranked by match count. `rg` is killed after 30 s and reported as a failure. Read-only, and **exempt from the Iteration Budget Gate** |
|
||||||
| `confidence decay [--apply]` | Recompute every page's derived `confidence` as `confidence_base * (1 - 0.01/month)`, floored at 0.2; dry-run by default |
|
| `confidence decay [--apply]` | Recompute every page's derived `confidence` as `confidence_base * (1 - 0.01/month)`, floored at 0.2; dry-run by default |
|
||||||
| `confidence init-base [--apply]` | One-time backfill: set `confidence_base` from the current `confidence` on pages that predate the derived-confidence model |
|
| `confidence init-base [--apply]` | One-time backfill: set `confidence_base` from the current `confidence` on pages that predate the derived-confidence model |
|
||||||
|
|||||||
+3
-2
@@ -45,6 +45,7 @@ tools/
|
|||||||
conventions.py kb/CONVENTIONS.md: what this instance decided about authoring, as opposed to what the stack enforces
|
conventions.py kb/CONVENTIONS.md: what this instance decided about authoring, as opposed to what the stack enforces
|
||||||
ownership.py the stack-vs-instance boundary under a content stage - one predicate, read by `dist_cmd.py` and `commands/upstream_cmd.py` so the two cannot answer it differently
|
ownership.py the stack-vs-instance boundary under a content stage - one predicate, read by `dist_cmd.py` and `commands/upstream_cmd.py` so the two cannot answer it differently
|
||||||
type_resolver.py type-spec loading and schema resolution
|
type_resolver.py type-spec loading and schema resolution
|
||||||
|
catalog.py how the corpus groups into collections and areas, and the shard threshold - with no CLI attached
|
||||||
lint_core.py the lint checks and the report, with no CLI attached
|
lint_core.py the lint checks and the report, with no CLI attached
|
||||||
types_core.py type-spec listing/description, with no CLI attached
|
types_core.py type-spec listing/description, with no CLI attached
|
||||||
markdown_code.py masks code spans/fences so a page may show wiki notation, not only use it
|
markdown_code.py masks code spans/fences so a page may show wiki notation, not only use it
|
||||||
@@ -57,8 +58,8 @@ tools/
|
|||||||
```
|
```
|
||||||
|
|
||||||
**Two consumers, one core.** The CLI is not the only caller any more. The cores
|
**Two consumers, one core.** The CLI is not the only caller any more. The cores
|
||||||
(`search/service.py`, `lint_core.py`, `types_core.py`) hold what decides an
|
(`search/service.py`, `lint_core.py`, `types_core.py`, `catalog.py`) hold what
|
||||||
answer and import no `typer` and no `rich`; the modules under `commands/` turn
|
decides an answer and import no `typer` and no `rich`; the modules under `commands/` turn
|
||||||
those values into terminal output and those exceptions into exit codes.
|
those values into terminal output and those exceptions into exit codes.
|
||||||
`api.Corpus` is the in-process entry point over the same functions - it takes a
|
`api.Corpus` is the in-process entry point over the same functions - it takes a
|
||||||
corpus root, returns exactly the structures the `--json` forms print, and
|
corpus root, returns exactly the structures the `--json` forms print, and
|
||||||
|
|||||||
@@ -0,0 +1,140 @@
|
|||||||
|
"""How the corpus groups into collections and areas, with no CLI attached.
|
||||||
|
|
||||||
|
Split out of `commands/index_build.py` for the reason `lint_core.py` gives at
|
||||||
|
the top of itself: this is a pure function over a corpus directory, and it was
|
||||||
|
sitting in a module that imports `typer` and `rich`. `lint` needs the same
|
||||||
|
grouping - it is what answers "does this collection have areas, and is it over
|
||||||
|
the threshold?" (Gitea #59) - and `chemenu.api`, the read surface, may not
|
||||||
|
reach a command module at all. Importing it from there would have pulled the
|
||||||
|
whole CLI head in behind it.
|
||||||
|
|
||||||
|
So the split runs along the same line as lint's: everything that decides *how
|
||||||
|
the corpus is shaped* lives here; everything that decides *what the catalog
|
||||||
|
looks like* - the tables, the map, the shard files - stays in
|
||||||
|
`commands/index_build.py`, which imports from here.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass, field
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from chemenu.kb_collections import iter_kb_collections
|
||||||
|
from chemenu.page import Page
|
||||||
|
from chemenu.type_resolver import resolver
|
||||||
|
|
||||||
|
# Rows per area before it is split into its own shard. From the wiki's own
|
||||||
|
# `Index Scaling` page ("split table sections at >50 entries"), kept as a plain
|
||||||
|
# number so growth is handled by arithmetic rather than by a judgment call.
|
||||||
|
SHARD_THRESHOLD = 50
|
||||||
|
|
||||||
|
# Display title for pages sitting directly in a collection root rather than in
|
||||||
|
# an area subdirectory.
|
||||||
|
UNGROUPED_TITLE = "All"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Area:
|
||||||
|
"""One grouping inside a collection: a subdirectory, or the collection root
|
||||||
|
for pages that sit directly in it."""
|
||||||
|
|
||||||
|
name: str
|
||||||
|
title: str
|
||||||
|
pages: list[Page] = field(default_factory=list)
|
||||||
|
own_shard: bool = False
|
||||||
|
|
||||||
|
@property
|
||||||
|
def count(self) -> int:
|
||||||
|
return len(self.pages)
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Collection:
|
||||||
|
name: str
|
||||||
|
areas: list[Area] = field(default_factory=list)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def count(self) -> int:
|
||||||
|
return sum(area.count for area in self.areas)
|
||||||
|
|
||||||
|
|
||||||
|
def area_titles() -> dict[str, dict[str, str]]:
|
||||||
|
"""Display titles per collection: `{collection: {area_dir: title}}`, taken
|
||||||
|
from each type-spec's own `layout:` rather than a hardcoded map - so a new
|
||||||
|
subtype names its own section by adding a type-spec, with no code change.
|
||||||
|
|
||||||
|
Every type-spec is read, not just `entity`'s. That hardcoding was the
|
||||||
|
asymmetry behind Gitea #59: the axis a collection splits along is declared
|
||||||
|
in `layout:`, and a second type declaring one would have had its areas
|
||||||
|
titled by `.title()` on the directory name while entity's got their real
|
||||||
|
names.
|
||||||
|
|
||||||
|
Keyed by collection rather than by directory name alone, because two types
|
||||||
|
writing into two collections may legitimately use the same area name for
|
||||||
|
different things (`kb/entities/tools/` and a hypothetical
|
||||||
|
`kb/concepts/tools/`); a flat map would hand the second one the first's
|
||||||
|
title. The collection key is the type's `base_dir:`, which is what put the
|
||||||
|
page in that directory to begin with.
|
||||||
|
"""
|
||||||
|
titles: dict[str, dict[str, str]] = {}
|
||||||
|
for type_path, _frontmatter in resolver.list_type_specs():
|
||||||
|
try:
|
||||||
|
if resolver.get_root(type_path) != "kb":
|
||||||
|
continue
|
||||||
|
base_dir = resolver.get_base_dir(type_path)
|
||||||
|
layout = resolver.get_layout(type_path)
|
||||||
|
except (ValueError, OSError):
|
||||||
|
continue
|
||||||
|
if not base_dir or not layout:
|
||||||
|
continue
|
||||||
|
per_collection = titles.setdefault(str(base_dir).strip("/"), {})
|
||||||
|
for key, spec in layout.items():
|
||||||
|
per_collection.setdefault(spec.get("dir", key), spec.get("title", str(key).title()))
|
||||||
|
return titles
|
||||||
|
|
||||||
|
|
||||||
|
def group_pages(kb_dir: Path, pages: dict[str, Page]) -> list[Collection]:
|
||||||
|
"""Group pages by their physical location: collection directory, then area
|
||||||
|
subdirectory.
|
||||||
|
|
||||||
|
Location rather than `kind` because a shard lives in the directory it
|
||||||
|
describes, and the two agree by construction: a type-spec's `base_dir:` is
|
||||||
|
what put the page there.
|
||||||
|
"""
|
||||||
|
titles = area_titles()
|
||||||
|
grouped: dict[str, dict[str, Area]] = {}
|
||||||
|
|
||||||
|
# Seed from the collections that exist on disk, not only from the ones that
|
||||||
|
# happen to hold pages: an empty collection is a real (if unfilled) part of
|
||||||
|
# the wiki, and dropping it from the map would hide it from every reader.
|
||||||
|
for collection_dir in iter_kb_collections(kb_dir):
|
||||||
|
grouped.setdefault(collection_dir.name, {})
|
||||||
|
|
||||||
|
for page in sorted(pages.values(), key=lambda p: p.title.lower()):
|
||||||
|
try:
|
||||||
|
parts = page.path.relative_to(kb_dir).parts
|
||||||
|
except ValueError: # pragma: no cover - pages always live under kb_dir
|
||||||
|
continue
|
||||||
|
if len(parts) < 2:
|
||||||
|
collection_name, area_name = "(kb root)", ""
|
||||||
|
else:
|
||||||
|
collection_name = parts[0]
|
||||||
|
area_name = parts[1] if len(parts) > 2 else ""
|
||||||
|
areas = grouped.setdefault(collection_name, {})
|
||||||
|
area = areas.get(area_name)
|
||||||
|
if area is None:
|
||||||
|
title = (
|
||||||
|
titles.get(collection_name, {}).get(area_name, area_name.title())
|
||||||
|
if area_name
|
||||||
|
else UNGROUPED_TITLE
|
||||||
|
)
|
||||||
|
area = Area(name=area_name, title=title)
|
||||||
|
areas[area_name] = area
|
||||||
|
area.pages.append(page)
|
||||||
|
|
||||||
|
collections = []
|
||||||
|
for name in sorted(grouped):
|
||||||
|
ordered = sorted(grouped[name].values(), key=lambda a: (a.name == "", a.title.lower()))
|
||||||
|
for area in ordered:
|
||||||
|
area.own_shard = bool(area.name) and area.count > SHARD_THRESHOLD
|
||||||
|
collections.append(Collection(name=name, areas=ordered))
|
||||||
|
return collections
|
||||||
@@ -18,18 +18,16 @@ The map stays small enough to browse; `wikitool search` answers everything else.
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import re
|
import re
|
||||||
from dataclasses import dataclass, field
|
|
||||||
from datetime import date
|
from datetime import date
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
import typer
|
import typer
|
||||||
|
|
||||||
from chemenu import config
|
from chemenu import config
|
||||||
|
from chemenu.catalog import SHARD_THRESHOLD, Area, Collection, group_pages
|
||||||
from chemenu.commands._util import rel_path, success
|
from chemenu.commands._util import rel_path, success
|
||||||
from chemenu.kb_collections import iter_kb_collections
|
|
||||||
from chemenu.page import Page
|
from chemenu.page import Page
|
||||||
from chemenu.kb_scan import GENERATED_INDEX, find_nested_pages, load_kb_pages
|
from chemenu.kb_scan import GENERATED_INDEX, find_nested_pages, load_kb_pages
|
||||||
from chemenu.type_resolver import resolver
|
|
||||||
|
|
||||||
app = typer.Typer(help="Manage the generated wiki catalog (kb/index.md + per-collection INDEX.md).")
|
app = typer.Typer(help="Manage the generated wiki catalog (kb/index.md + per-collection INDEX.md).")
|
||||||
|
|
||||||
@@ -38,15 +36,6 @@ TABLE_SEP = "|------|------|---------|----------------|"
|
|||||||
|
|
||||||
SUMMARY_HEADINGS = ("Description", "Definition", "Summary")
|
SUMMARY_HEADINGS = ("Description", "Definition", "Summary")
|
||||||
|
|
||||||
# Rows per area before it is split into its own shard. From the wiki's own
|
|
||||||
# `Index Scaling` page ("split table sections at >50 entries"), kept as a plain
|
|
||||||
# number so growth is handled by arithmetic rather than by a judgment call.
|
|
||||||
SHARD_THRESHOLD = 50
|
|
||||||
|
|
||||||
# Display title for pages sitting directly in a collection root rather than in
|
|
||||||
# an area subdirectory.
|
|
||||||
UNGROUPED_TITLE = "All"
|
|
||||||
|
|
||||||
DO_NOT_EDIT = "<!-- Generated by `wikitool index rebuild`. Do not hand-edit. -->"
|
DO_NOT_EDIT = "<!-- Generated by `wikitool index rebuild`. Do not hand-edit. -->"
|
||||||
|
|
||||||
|
|
||||||
@@ -85,86 +74,17 @@ def _table(pages: list[Page]) -> list[str]:
|
|||||||
|
|
||||||
|
|
||||||
def _anchor(title: str) -> str:
|
def _anchor(title: str) -> str:
|
||||||
"""GitHub-style heading anchor, so the map can deep-link into a shard."""
|
"""GitHub-style heading anchor, so the map can deep-link into a shard.
|
||||||
slug = re.sub(r"[^a-z0-9\s-]", "", title.lower())
|
|
||||||
return re.sub(r"\s+", "-", slug.strip())
|
|
||||||
|
|
||||||
|
`\\w` rather than `a-z0-9`, which is not cosmetic: an area title follows the
|
||||||
@dataclass
|
KB language, and the first non-English one (`Abläufe`) had its umlaut
|
||||||
class Area:
|
*deleted* rather than kept, so the map linked at `#ablufe` and the anchor it
|
||||||
"""One grouping inside a collection: a subdirectory, or the collection root
|
was aiming at was `#abläufe`. Every deep link into a shard whose title
|
||||||
for pages that sit directly in it."""
|
carries a non-ASCII letter was silently dead. Nothing surfaced it while the
|
||||||
|
only areas were entity ones, whose titles happen to be ASCII throughout.
|
||||||
name: str
|
|
||||||
title: str
|
|
||||||
pages: list[Page] = field(default_factory=list)
|
|
||||||
own_shard: bool = False
|
|
||||||
|
|
||||||
@property
|
|
||||||
def count(self) -> int:
|
|
||||||
return len(self.pages)
|
|
||||||
|
|
||||||
|
|
||||||
@dataclass
|
|
||||||
class Collection:
|
|
||||||
name: str
|
|
||||||
areas: list[Area] = field(default_factory=list)
|
|
||||||
|
|
||||||
@property
|
|
||||||
def count(self) -> int:
|
|
||||||
return sum(area.count for area in self.areas)
|
|
||||||
|
|
||||||
|
|
||||||
def _area_titles() -> dict[str, str]:
|
|
||||||
"""Display titles for entity areas, taken from the entity type-spec's own
|
|
||||||
`layout:` rather than a hardcoded map - so a new subtype names its own
|
|
||||||
section by adding a type-spec, with no code change."""
|
|
||||||
layout = resolver.get_layout(resolver.find_type_by_name("entity")) or {}
|
|
||||||
return {spec.get("dir", key): spec.get("title", key.title()) for key, spec in layout.items()}
|
|
||||||
|
|
||||||
|
|
||||||
def group_pages(kb_dir: Path, pages: dict[str, Page]) -> list[Collection]:
|
|
||||||
"""Group pages by their physical location: collection directory, then area
|
|
||||||
subdirectory.
|
|
||||||
|
|
||||||
Location rather than `kind` because a shard lives in the directory it
|
|
||||||
describes, and the two agree by construction: a type-spec's `base_dir:` is
|
|
||||||
what put the page there.
|
|
||||||
"""
|
"""
|
||||||
titles = _area_titles()
|
slug = re.sub(r"[^\w\s-]", "", title.lower(), flags=re.UNICODE)
|
||||||
grouped: dict[str, dict[str, Area]] = {}
|
return re.sub(r"\s+", "-", slug.strip())
|
||||||
|
|
||||||
# Seed from the collections that exist on disk, not only from the ones that
|
|
||||||
# happen to hold pages: an empty collection is a real (if unfilled) part of
|
|
||||||
# the wiki, and dropping it from the map would hide it from every reader.
|
|
||||||
for collection_dir in iter_kb_collections(kb_dir):
|
|
||||||
grouped.setdefault(collection_dir.name, {})
|
|
||||||
|
|
||||||
for page in sorted(pages.values(), key=lambda p: p.title.lower()):
|
|
||||||
try:
|
|
||||||
parts = page.path.relative_to(kb_dir).parts
|
|
||||||
except ValueError: # pragma: no cover - pages always live under kb_dir
|
|
||||||
continue
|
|
||||||
if len(parts) < 2:
|
|
||||||
collection_name, area_name = "(kb root)", ""
|
|
||||||
else:
|
|
||||||
collection_name = parts[0]
|
|
||||||
area_name = parts[1] if len(parts) > 2 else ""
|
|
||||||
areas = grouped.setdefault(collection_name, {})
|
|
||||||
area = areas.get(area_name)
|
|
||||||
if area is None:
|
|
||||||
title = titles.get(area_name, area_name.title()) if area_name else UNGROUPED_TITLE
|
|
||||||
area = Area(name=area_name, title=title)
|
|
||||||
areas[area_name] = area
|
|
||||||
area.pages.append(page)
|
|
||||||
|
|
||||||
collections = []
|
|
||||||
for name in sorted(grouped):
|
|
||||||
ordered = sorted(grouped[name].values(), key=lambda a: (a.name == "", a.title.lower()))
|
|
||||||
for area in ordered:
|
|
||||||
area.own_shard = bool(area.name) and area.count > SHARD_THRESHOLD
|
|
||||||
collections.append(Collection(name=name, areas=ordered))
|
|
||||||
return collections
|
|
||||||
|
|
||||||
|
|
||||||
def build_area_shard(area: Area) -> str:
|
def build_area_shard(area: Area) -> str:
|
||||||
|
|||||||
@@ -16,7 +16,10 @@ from __future__ import annotations
|
|||||||
from datetime import date
|
from datetime import date
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
|
from collections import Counter
|
||||||
|
|
||||||
from chemenu import blocks, config, kb_collections, links
|
from chemenu import blocks, config, kb_collections, links
|
||||||
|
from chemenu.catalog import SHARD_THRESHOLD, group_pages
|
||||||
from chemenu.frontmatter_io import frontmatter_error
|
from chemenu.frontmatter_io import frontmatter_error
|
||||||
from chemenu.markdown_code import strip_code_spans
|
from chemenu.markdown_code import strip_code_spans
|
||||||
from chemenu.provenance import broken_raw_refs as find_broken_raw_refs
|
from chemenu.provenance import broken_raw_refs as find_broken_raw_refs
|
||||||
@@ -141,6 +144,88 @@ def nested_pages(kb_dir: Path, pages: dict[str, Page]) -> list[dict]:
|
|||||||
]
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def unsharded_collections(kb_dir: Path, pages: dict[str, Page]) -> list[dict]:
|
||||||
|
"""Collections past the catalog's shard threshold that have no areas to
|
||||||
|
shard, together with the subtype split that would give them some.
|
||||||
|
|
||||||
|
Sharding is already automatic, and it is per *area*: `index rebuild` hands
|
||||||
|
an area over `SHARD_THRESHOLD` rows its own `INDEX.md`. Creating an area is
|
||||||
|
not automatic and nothing ever asked for one - so a collection that never
|
||||||
|
grew any keeps its whole catalog in a single table, past the threshold,
|
||||||
|
forever. The threshold is then not a threshold but a dead value (Gitea
|
||||||
|
#59), and this is the only check that can notice: an ingest sees one
|
||||||
|
source and cannot see a collection's size, while `lint` sees the corpus and
|
||||||
|
runs every 10 sources anyway.
|
||||||
|
|
||||||
|
**A recommendation, not a failure** (it is deliberately absent from
|
||||||
|
`HARD_ERROR_KEYS`), and narrow enough to stay one: it fires only where the
|
||||||
|
split actually helps - every area it would create, the ungrouped remainder
|
||||||
|
included, lands at or under the threshold. That self-limits in both
|
||||||
|
directions. A collection under the threshold never fires, so a small
|
||||||
|
`kb/comparisons/` is not permanently in the report; and a collection whose
|
||||||
|
subtype values are lopsided (25 of 29 `source_type: notes`) does not fire
|
||||||
|
either, because splitting it would produce one area over the threshold and
|
||||||
|
a handful of splinters. What is left is a finding that appears when a
|
||||||
|
collection grows into it and is silent when it does not.
|
||||||
|
"""
|
||||||
|
findings: list[dict] = []
|
||||||
|
for collection in group_pages(kb_dir, pages):
|
||||||
|
if collection.count <= SHARD_THRESHOLD:
|
||||||
|
continue
|
||||||
|
# An area already exists, so the collection has been split once and
|
||||||
|
# `index rebuild` shards whatever outgrows the threshold from here.
|
||||||
|
# A page still sitting in the root is `misplaced_pages`' finding, not
|
||||||
|
# this one.
|
||||||
|
if any(area.name for area in collection.areas):
|
||||||
|
continue
|
||||||
|
|
||||||
|
counts: Counter[str] = Counter()
|
||||||
|
fields: set[str] = set()
|
||||||
|
# Whether every type writing here already declares the `layout:` that
|
||||||
|
# turns the subtype into a directory. It decides which half of the fix
|
||||||
|
# is still owed: without it there is nothing for `move` to compute a
|
||||||
|
# destination from, with it the move is all that is left.
|
||||||
|
layouts: set[bool] = set()
|
||||||
|
for area in collection.areas:
|
||||||
|
for page in area.pages:
|
||||||
|
type_path = page.frontmatter.get("type")
|
||||||
|
if not type_path:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
field = resolver.get_subtype_field(type_path, page.path)
|
||||||
|
layout = resolver.get_layout(type_path, page.path)
|
||||||
|
except ValueError:
|
||||||
|
continue
|
||||||
|
value = page.frontmatter.get(field) if field else None
|
||||||
|
if not value:
|
||||||
|
continue
|
||||||
|
counts[str(value)] += 1
|
||||||
|
fields.add(str(field))
|
||||||
|
layouts.add(bool(layout))
|
||||||
|
|
||||||
|
if not counts:
|
||||||
|
continue
|
||||||
|
# The pages the subtype cannot place stay in the collection root, so
|
||||||
|
# they are an area of their own for the purpose of this test.
|
||||||
|
unplaced = collection.count - sum(counts.values())
|
||||||
|
if max([*counts.values(), unplaced]) > SHARD_THRESHOLD:
|
||||||
|
continue
|
||||||
|
|
||||||
|
findings.append(
|
||||||
|
{
|
||||||
|
"collection": collection.name,
|
||||||
|
"count": collection.count,
|
||||||
|
"field": ", ".join(sorted(fields)),
|
||||||
|
"layout_declared": layouts == {True},
|
||||||
|
"distribution": [
|
||||||
|
{"value": value, "count": count}
|
||||||
|
for value, count in sorted(counts.items(), key=lambda kv: (-kv[1], kv[0]))
|
||||||
|
],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return findings
|
||||||
|
|
||||||
|
|
||||||
def run_lint(kb_dir: Path) -> dict:
|
def run_lint(kb_dir: Path) -> dict:
|
||||||
pages = load_kb_pages(kb_dir)
|
pages = load_kb_pages(kb_dir)
|
||||||
duplicate_titles = find_duplicate_title_paths(kb_dir, config.ROOT)
|
duplicate_titles = find_duplicate_title_paths(kb_dir, config.ROOT)
|
||||||
@@ -388,6 +473,7 @@ def run_lint(kb_dir: Path) -> dict:
|
|||||||
"duplicate_titles": duplicate_titles,
|
"duplicate_titles": duplicate_titles,
|
||||||
"misplaced_pages": misplaced,
|
"misplaced_pages": misplaced,
|
||||||
"nested_pages": nested,
|
"nested_pages": nested,
|
||||||
|
"unsharded_collections": unsharded_collections(kb_dir, pages),
|
||||||
"uncovered_raw_files": find_uncovered_raw_files(config.RAW_DIR, pages),
|
"uncovered_raw_files": find_uncovered_raw_files(config.RAW_DIR, pages),
|
||||||
"broken_raw_refs": find_broken_raw_refs(pages),
|
"broken_raw_refs": find_broken_raw_refs(pages),
|
||||||
"duplicate_raw_file_owners": find_duplicate_raw_file_owners(pages),
|
"duplicate_raw_file_owners": find_duplicate_raw_file_owners(pages),
|
||||||
@@ -464,6 +550,23 @@ def render_markdown(report: dict) -> str:
|
|||||||
"the catalog folds this into its area silently; `wikitool move --reconcile` fixes it "
|
"the catalog folds this into its area silently; `wikitool move --reconcile` fixes it "
|
||||||
"when the page's type resolves to a shallower directory, otherwise move it up by hand",
|
"when the page's type resolves to a shallower directory, otherwise move it up by hand",
|
||||||
)
|
)
|
||||||
|
_section(
|
||||||
|
lines, f"Collections Past the Shard Threshold (>{SHARD_THRESHOLD}) With No Areas "
|
||||||
|
"- recommendation, not an error",
|
||||||
|
report.get("unsharded_collections", []),
|
||||||
|
lambda i: f"`kb/{i['collection']}/` holds {i['count']} pages in a single table and has no "
|
||||||
|
f"areas, so the per-area shard threshold never fires. Splitting on `{i['field']}` would "
|
||||||
|
"give: "
|
||||||
|
+ ", ".join(f"{d['value']} {d['count']}" for d in i["distribution"])
|
||||||
|
+ f" - all at or under {SHARD_THRESHOLD}. "
|
||||||
|
+ (
|
||||||
|
"The type-spec already declares the `layout:` for those values, so "
|
||||||
|
"`wikitool move --reconcile` and `wikitool index rebuild` are the whole fix"
|
||||||
|
if i.get("layout_declared")
|
||||||
|
else "Declare a `layout:` for those values in the type-spec, then "
|
||||||
|
"`wikitool move --reconcile` and `wikitool index rebuild`"
|
||||||
|
),
|
||||||
|
)
|
||||||
_section(
|
_section(
|
||||||
lines, "Uncovered Raw Files (no source page)", report["uncovered_raw_files"],
|
lines, "Uncovered Raw Files (no source page)", report["uncovered_raw_files"],
|
||||||
lambda i: f"`{i}`",
|
lambda i: f"`{i}`",
|
||||||
@@ -617,6 +720,14 @@ def default_report_path(report: dict) -> Path:
|
|||||||
# and there is no version at which "not under the computed directory" becomes
|
# and there is no version at which "not under the computed directory" becomes
|
||||||
# wrong - only `wikitool move` someone does or does not get to run.
|
# wrong - only `wikitool move` someone does or does not get to run.
|
||||||
#
|
#
|
||||||
|
# `unsharded_collections` is advisory by construction rather than by tolerance:
|
||||||
|
# it does not describe anything that is wrong, only a collection that has grown
|
||||||
|
# past the size at which areas start paying for themselves. Whether to split it
|
||||||
|
# is an authoring decision about how the corpus is organised - the tool can see
|
||||||
|
# that the split would work and say so, and that is the whole of its authority.
|
||||||
|
# Failing on it would also make `lint` red on a corpus that is entirely
|
||||||
|
# self-consistent, which is the state the recommendation is asking to improve.
|
||||||
|
#
|
||||||
# `malformed_edges` and `unbalanced_markers` are hard from the start: neither
|
# `malformed_edges` and `unbalanced_markers` are hard from the start: neither
|
||||||
# describes an unconverted page, only a broken one.
|
# describes an unconverted page, only a broken one.
|
||||||
#
|
#
|
||||||
|
|||||||
@@ -254,7 +254,7 @@ def kb_dir(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path:
|
|||||||
kb = tmp_path / "kb"
|
kb = tmp_path / "kb"
|
||||||
for sub in ("entities/projects", "entities/systems", "entities/tools",
|
for sub in ("entities/projects", "entities/systems", "entities/tools",
|
||||||
"entities/technologies", "entities/people",
|
"entities/technologies", "entities/people",
|
||||||
"concepts", "sources", "comparisons"):
|
"concepts/protocols", "sources", "comparisons"):
|
||||||
(kb / sub).mkdir(parents=True)
|
(kb / sub).mkdir(parents=True)
|
||||||
# The contracts carry a real declaration, because three things now read one:
|
# The contracts carry a real declaration, because three things now read one:
|
||||||
# `docs verify` checks `profile:`/`required_by_stack:`, and `xref add` asks
|
# `docs verify` checks `profile:`/`required_by_stack:`, and `xref add` asks
|
||||||
@@ -302,7 +302,7 @@ def kb_dir(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path:
|
|||||||
"\n# gdeploy\n\n## Description\n\nDeploy tool.\n",
|
"\n# gdeploy\n\n## Description\n\nDeploy tool.\n",
|
||||||
)
|
)
|
||||||
write_page(
|
write_page(
|
||||||
kb / "concepts/Modbus.md",
|
kb / "concepts/protocols/Modbus.md",
|
||||||
{
|
{
|
||||||
"type": "types/concept.md", "concept_type": "protocol",
|
"type": "types/concept.md", "concept_type": "protocol",
|
||||||
"tags": [], "created": "2026-07-25", "modified": "2026-07-25",
|
"tags": [], "created": "2026-07-25", "modified": "2026-07-25",
|
||||||
|
|||||||
@@ -54,7 +54,7 @@ def test_cite_add_command_writes_definition_and_prints_marker(kb_dir, raw_dir, m
|
|||||||
marker_id = cite_id("Source - Aurora")
|
marker_id = cite_id("Source - Aurora")
|
||||||
assert f"[^{marker_id}]" in result.output
|
assert f"[^{marker_id}]" in result.output
|
||||||
|
|
||||||
fm, body = read_page(kb_dir / "concepts/Modbus.md")
|
fm, body = read_page(kb_dir / "concepts/protocols/Modbus.md")
|
||||||
assert "Source - Aurora" in fm["sources"]
|
assert "Source - Aurora" in fm["sources"]
|
||||||
assert f"[^{marker_id}]: [[Source - Aurora]]" in body
|
assert f"[^{marker_id}]: [[Source - Aurora]]" in body
|
||||||
|
|
||||||
@@ -194,11 +194,11 @@ def test_cite_sync_command_over_kb(kb_dir, raw_dir, monkeypatch):
|
|||||||
assert add_result.exit_code == 0, add_result.output
|
assert add_result.exit_code == 0, add_result.output
|
||||||
marker_id = cite_id("Source - Aurora")
|
marker_id = cite_id("Source - Aurora")
|
||||||
|
|
||||||
fm, body = read_page(kb_dir / "concepts/Modbus.md")
|
fm, body = read_page(kb_dir / "concepts/protocols/Modbus.md")
|
||||||
body = body.replace("Industrial protocol.", f"Industrial protocol [^{marker_id}].")
|
body = body.replace("Industrial protocol.", f"Industrial protocol [^{marker_id}].")
|
||||||
from chemenu.frontmatter_io import write_page
|
from chemenu.frontmatter_io import write_page
|
||||||
|
|
||||||
write_page(kb_dir / "concepts/Modbus.md", fm, body)
|
write_page(kb_dir / "concepts/protocols/Modbus.md", fm, body)
|
||||||
|
|
||||||
result = runner.invoke(app, ["cite", "sync", "--all", "--dry-run"])
|
result = runner.invoke(app, ["cite", "sync", "--all", "--dry-run"])
|
||||||
assert result.exit_code == 0, result.output
|
assert result.exit_code == 0, result.output
|
||||||
|
|||||||
@@ -494,7 +494,7 @@ def test_generated_files_are_recognised_wherever_they_sit():
|
|||||||
assert is_generated("kb/provenance.md")
|
assert is_generated("kb/provenance.md")
|
||||||
assert is_generated("kb/concepts/INDEX.md")
|
assert is_generated("kb/concepts/INDEX.md")
|
||||||
assert is_generated("kb/entities/tools/INDEX.md")
|
assert is_generated("kb/entities/tools/INDEX.md")
|
||||||
assert not is_generated("kb/concepts/Modbus.md")
|
assert not is_generated("kb/concepts/protocols/Modbus.md")
|
||||||
|
|
||||||
|
|
||||||
def test_paths_land_in_the_group_a_reviewer_expects():
|
def test_paths_land_in_the_group_a_reviewer_expects():
|
||||||
|
|||||||
@@ -15,14 +15,14 @@ from chemenu.kb_scan import GENERATED_INDEX, iter_kb_pages
|
|||||||
from chemenu.type_resolver import resolver
|
from chemenu.type_resolver import resolver
|
||||||
|
|
||||||
|
|
||||||
def _area_title(subtype: str) -> str:
|
def _area_title(subtype: str, type_path: str = "types/entity.md") -> str:
|
||||||
"""The display title `index rebuild` will use for an entity subtype.
|
"""The display title `index rebuild` will use for a subtype's area.
|
||||||
|
|
||||||
Read from the type-spec rather than written out, because these titles follow
|
Read from the type-spec rather than written out, because these titles follow
|
||||||
the KB language: hard-coding them made translating the wiki fail tests that
|
the KB language: hard-coding them made translating the wiki fail tests that
|
||||||
are not about wording at all.
|
are not about wording at all.
|
||||||
"""
|
"""
|
||||||
return resolver.get_layout("types/entity.md")[subtype]["title"]
|
return resolver.get_layout(type_path)[subtype]["title"]
|
||||||
|
|
||||||
|
|
||||||
@pytest.fixture
|
@pytest.fixture
|
||||||
@@ -116,6 +116,28 @@ def test_area_titles_come_from_the_entity_type_spec_layout(plan, kb_dir):
|
|||||||
assert f"## {_area_title('tool')}" in entities
|
assert f"## {_area_title('tool')}" in entities
|
||||||
|
|
||||||
|
|
||||||
|
def test_anchor_keeps_non_ascii_letters():
|
||||||
|
"""An area title follows the KB language, so it may carry a letter outside
|
||||||
|
`a-z`. Deleting it - which is what the old `[^a-z0-9\\s-]` did - produced a
|
||||||
|
map link (`#ablufe`) that pointed at no heading in the shard it named."""
|
||||||
|
assert _anchor("Abläufe") == "abläufe"
|
||||||
|
assert _anchor("Größere Muster") == "größere-muster"
|
||||||
|
# Punctuation is still dropped and any run of whitespace still collapses to
|
||||||
|
# a single hyphen.
|
||||||
|
assert _anchor("Tools & Utilities (v2)") == "tools-utilities-v2"
|
||||||
|
|
||||||
|
|
||||||
|
def test_area_titles_are_read_from_every_type_spec_not_only_entity(plan, kb_dir):
|
||||||
|
"""Gitea #59: the title lookup used to resolve `entity` by name and read
|
||||||
|
only its `layout:`, so a second type declaring one got `.title()` on its
|
||||||
|
directory name (`Protocols`) instead of the title it declared. The concept
|
||||||
|
areas are the first case; nothing about them is special."""
|
||||||
|
concepts = _shard(plan, kb_dir, "concepts")
|
||||||
|
declared = _area_title("protocol", "types/concept.md")
|
||||||
|
assert f"## {declared}" in concepts
|
||||||
|
assert "## Protocols" not in concepts
|
||||||
|
|
||||||
|
|
||||||
def test_summary_prefers_frontmatter_then_falls_back_to_body(plan, kb_dir):
|
def test_summary_prefers_frontmatter_then_falls_back_to_body(plan, kb_dir):
|
||||||
entities = _shard(plan, kb_dir, "entities")
|
entities = _shard(plan, kb_dir, "entities")
|
||||||
assert "Server hosting DocStore with ZFS storage" in entities # frontmatter
|
assert "Server hosting DocStore with ZFS storage" in entities # frontmatter
|
||||||
|
|||||||
@@ -236,9 +236,133 @@ def test_lint_is_silent_about_pages_directly_in_an_area(kb_dir):
|
|||||||
assert run_lint(kb_dir)["nested_pages"] == []
|
assert run_lint(kb_dir)["nested_pages"] == []
|
||||||
|
|
||||||
|
|
||||||
|
# --- collections that outgrew the shard threshold without areas (Gitea #59) ---
|
||||||
|
|
||||||
|
|
||||||
|
def _fill_collection(kb_dir, collection: str, type_path: str, field: str, distribution: dict):
|
||||||
|
"""Write pages flat into `kb/<collection>/`, `distribution` many per subtype
|
||||||
|
value - the shape a collection is in when nobody ever created an area.
|
||||||
|
|
||||||
|
The precondition is established rather than assumed: the fixture corpus
|
||||||
|
places its concept page in an area (as the real one does now), and a single
|
||||||
|
area is enough to make this finding stand down.
|
||||||
|
"""
|
||||||
|
for existing in sorted((kb_dir / collection).rglob("*.md")):
|
||||||
|
if existing.parent != kb_dir / collection:
|
||||||
|
existing.rename(kb_dir / collection / existing.name)
|
||||||
|
for area in sorted(p for p in (kb_dir / collection).iterdir() if p.is_dir()):
|
||||||
|
area.rmdir()
|
||||||
|
|
||||||
|
n = 0
|
||||||
|
for value, count in distribution.items():
|
||||||
|
for _ in range(count):
|
||||||
|
n += 1
|
||||||
|
write_page(
|
||||||
|
kb_dir / collection / f"page-{n:03d}.md",
|
||||||
|
{
|
||||||
|
"type": type_path, field: value,
|
||||||
|
"created": "2026-08-01", "modified": "2026-08-01",
|
||||||
|
"provenance": "general", "summary": f"Page {n}",
|
||||||
|
},
|
||||||
|
f"\n# page-{n:03d}\n",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_lint_recommends_areas_for_a_collection_past_the_threshold(kb_dir):
|
||||||
|
"""The concepts shape: 80 pages in one table, no areas, and a subtype axis
|
||||||
|
whose largest value (28) lands well under the threshold. Sharding is
|
||||||
|
per-area and automatic, so a collection with no areas never splits however
|
||||||
|
large it grows - the threshold is a dead value until someone makes areas."""
|
||||||
|
_fill_collection(
|
||||||
|
kb_dir, "concepts", "types/concept.md", "concept_type",
|
||||||
|
{"workflow": 28, "architecture": 20, "pattern": 17,
|
||||||
|
"decision": 7, "problem": 5, "protocol": 3},
|
||||||
|
)
|
||||||
|
finding = next(
|
||||||
|
i for i in run_lint(kb_dir)["unsharded_collections"] if i["collection"] == "concepts"
|
||||||
|
)
|
||||||
|
assert finding["field"] == "concept_type"
|
||||||
|
# Largest first, so the reader sees the area that decides whether it helps.
|
||||||
|
assert finding["distribution"][0] == {"value": "workflow", "count": 28}
|
||||||
|
assert {d["value"] for d in finding["distribution"]} == {
|
||||||
|
"workflow", "architecture", "pattern", "decision", "problem", "protocol"
|
||||||
|
}
|
||||||
|
assert finding["count"] == sum(d["count"] for d in finding["distribution"])
|
||||||
|
# `types/concept.md` carries the layout, so only the move is still owed -
|
||||||
|
# the report line says which half of the fix that is.
|
||||||
|
assert finding["layout_declared"] is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_area_recommendation_says_the_layout_is_missing_when_it_is(kb_dir):
|
||||||
|
"""`types/source.md` deliberately declares no `layout:`, so a sources
|
||||||
|
collection that grew past the threshold has nothing for `move` to compute a
|
||||||
|
destination from - the fix starts one step earlier, and the report says so."""
|
||||||
|
_fill_collection(
|
||||||
|
kb_dir, "sources", "types/source.md", "source_type",
|
||||||
|
{"notes": 30, "article": 21},
|
||||||
|
)
|
||||||
|
report = run_lint(kb_dir)
|
||||||
|
finding = next(i for i in report["unsharded_collections"] if i["collection"] == "sources")
|
||||||
|
assert finding["layout_declared"] is False
|
||||||
|
assert "Declare a `layout:`" in render_markdown(report)
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_area_recommendation_is_not_a_failure(kb_dir):
|
||||||
|
"""It reports a collection that has outgrown a layout, not a broken one.
|
||||||
|
A corpus whose only finding is this must stay green, or every instance
|
||||||
|
goes red on the release that shipped the check."""
|
||||||
|
_fill_collection(
|
||||||
|
kb_dir, "concepts", "types/concept.md", "concept_type",
|
||||||
|
{"workflow": 28, "architecture": 24},
|
||||||
|
)
|
||||||
|
report = run_lint(kb_dir)
|
||||||
|
assert report["unsharded_collections"]
|
||||||
|
assert "unsharded_collections" not in HARD_ERROR_KEYS
|
||||||
|
assert not has_hard_errors({**{key: [] for key in HARD_ERROR_KEYS},
|
||||||
|
"unsharded_collections": report["unsharded_collections"]})
|
||||||
|
|
||||||
|
|
||||||
|
def test_lint_is_silent_about_a_collection_under_the_threshold(kb_dir):
|
||||||
|
"""The sources shape: 29 pages, lopsided across `source_type` - and under
|
||||||
|
the threshold anyway, so it never fires. That is what keeps the bad split
|
||||||
|
(one area of 25 plus four splinters) from ever being recommended, without
|
||||||
|
the check needing to know anything about sources."""
|
||||||
|
_fill_collection(
|
||||||
|
kb_dir, "sources", "types/source.md", "source_type",
|
||||||
|
{"notes": 25, "article": 3, "document": 1},
|
||||||
|
)
|
||||||
|
assert run_lint(kb_dir)["unsharded_collections"] == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_lint_is_silent_when_the_split_would_not_help(kb_dir):
|
||||||
|
"""Past the threshold, but 70 of 80 share one subtype value: splitting
|
||||||
|
produces one area still over the threshold plus splinters, which is not an
|
||||||
|
improvement. The second half of the criterion, and the one a
|
||||||
|
threshold-only check would have got wrong."""
|
||||||
|
_fill_collection(
|
||||||
|
kb_dir, "concepts", "types/concept.md", "concept_type",
|
||||||
|
{"workflow": 70, "architecture": 6, "pattern": 5},
|
||||||
|
)
|
||||||
|
assert run_lint(kb_dir)["unsharded_collections"] == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_lint_is_silent_once_the_collection_has_areas(kb_dir):
|
||||||
|
"""After the fix - the pages sit in their areas - the finding goes away,
|
||||||
|
and `index rebuild` shards whatever outgrows the threshold from here."""
|
||||||
|
_fill_collection(
|
||||||
|
kb_dir, "concepts", "types/concept.md", "concept_type",
|
||||||
|
{"workflow": 28, "architecture": 24},
|
||||||
|
)
|
||||||
|
for page in sorted((kb_dir / "concepts").glob("page-*.md")):
|
||||||
|
area = "workflows" if "workflow" in page.read_text(encoding="utf-8") else "architectures"
|
||||||
|
(kb_dir / "concepts" / area).mkdir(exist_ok=True)
|
||||||
|
page.rename(kb_dir / "concepts" / area / page.name)
|
||||||
|
assert run_lint(kb_dir)["unsharded_collections"] == []
|
||||||
|
|
||||||
|
|
||||||
def test_lint_flags_legacy_citation_marker_as_hard_error(kb_dir):
|
def test_lint_flags_legacy_citation_marker_as_hard_error(kb_dir):
|
||||||
write_page(
|
write_page(
|
||||||
kb_dir / "concepts/Modbus.md",
|
kb_dir / "concepts/protocols/Modbus.md",
|
||||||
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
||||||
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7},
|
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7},
|
||||||
"\n# Modbus\n\n## Definition\n\nUses port 502 ^[[Source - Aurora]].\n",
|
"\n# Modbus\n\n## Definition\n\nUses port 502 ^[[Source - Aurora]].\n",
|
||||||
@@ -250,7 +374,7 @@ def test_lint_flags_legacy_citation_marker_as_hard_error(kb_dir):
|
|||||||
|
|
||||||
def test_lint_flags_undefined_footnote_ref_as_hard_error(kb_dir):
|
def test_lint_flags_undefined_footnote_ref_as_hard_error(kb_dir):
|
||||||
write_page(
|
write_page(
|
||||||
kb_dir / "concepts/Modbus.md",
|
kb_dir / "concepts/protocols/Modbus.md",
|
||||||
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
||||||
"modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7},
|
"modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7},
|
||||||
"\n# Modbus\n\n## Definition\n\nUses port 502 [^s-ghost].\n",
|
"\n# Modbus\n\n## Definition\n\nUses port 502 [^s-ghost].\n",
|
||||||
@@ -264,7 +388,7 @@ def test_lint_flags_orphan_footnote_def_as_hard_error(kb_dir):
|
|||||||
cid = cite_id("Source - Aurora")
|
cid = cite_id("Source - Aurora")
|
||||||
block = render_cite_block({cid: ("Source - Aurora", None)})
|
block = render_cite_block({cid: ("Source - Aurora", None)})
|
||||||
write_page(
|
write_page(
|
||||||
kb_dir / "concepts/Modbus.md",
|
kb_dir / "concepts/protocols/Modbus.md",
|
||||||
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
||||||
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7},
|
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7},
|
||||||
f"\n# Modbus\n\n## Definition\n\nIndustrial protocol, no citation here.\n\n{block}",
|
f"\n# Modbus\n\n## Definition\n\nIndustrial protocol, no citation here.\n\n{block}",
|
||||||
@@ -278,7 +402,7 @@ def test_lint_clean_footnote_citation_has_no_hard_errors(kb_dir):
|
|||||||
cid = cite_id("Source - Aurora")
|
cid = cite_id("Source - Aurora")
|
||||||
block = render_cite_block({cid: ("Source - Aurora", None)})
|
block = render_cite_block({cid: ("Source - Aurora", None)})
|
||||||
write_page(
|
write_page(
|
||||||
kb_dir / "concepts/Modbus.md",
|
kb_dir / "concepts/protocols/Modbus.md",
|
||||||
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
{"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
||||||
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7},
|
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7},
|
||||||
f"\n# Modbus\n\n## Definition\n\nUses port 502 [^{cid}].\n\n{block}",
|
f"\n# Modbus\n\n## Definition\n\nUses port 502 [^{cid}].\n\n{block}",
|
||||||
|
|||||||
@@ -43,8 +43,8 @@ def test_page_subdir_falls_back_for_unmapped_subtype():
|
|||||||
|
|
||||||
|
|
||||||
def test_page_subdir_is_none_for_types_without_layout():
|
def test_page_subdir_is_none_for_types_without_layout():
|
||||||
assert _page_subdir(None, "types/concept.md") is None
|
assert _page_subdir(None, "types/source.md") is None
|
||||||
assert _page_subdir("anything", "types/concept.md") is None
|
assert _page_subdir("anything", "types/source.md") is None
|
||||||
|
|
||||||
|
|
||||||
def test_coerce_set_value_uses_declared_schema_type():
|
def test_coerce_set_value_uses_declared_schema_type():
|
||||||
@@ -258,12 +258,17 @@ def test_new_source_rejects_invalid_source_type(monkeypatch, kb_dir):
|
|||||||
assert result.exit_code != 0
|
assert result.exit_code != 0
|
||||||
|
|
||||||
|
|
||||||
def test_new_concept_creates_page(monkeypatch, kb_dir):
|
def test_new_concept_creates_page_in_its_subtype_area(monkeypatch, kb_dir):
|
||||||
|
"""Gitea #59: `types/concept.md` declares a `layout:` now, so a new concept
|
||||||
|
reaches its area with nothing else asked of the author - the same rule that
|
||||||
|
has always placed an entity. Nothing about `new` changed to make this true;
|
||||||
|
the type-spec did."""
|
||||||
result = _invoke_new(monkeypatch, kb_dir, [
|
result = _invoke_new(monkeypatch, kb_dir, [
|
||||||
"new", "concept", "--name", "Event Sourcing", "--set", "concept_type=pattern",
|
"new", "concept", "--name", "Event Sourcing", "--set", "concept_type=pattern",
|
||||||
])
|
])
|
||||||
assert result.exit_code == 0, result.output
|
assert result.exit_code == 0, result.output
|
||||||
assert (kb_dir / "concepts/Event Sourcing.md").exists()
|
assert (kb_dir / "concepts/patterns/Event Sourcing.md").exists()
|
||||||
|
assert not (kb_dir / "concepts/Event Sourcing.md").exists()
|
||||||
|
|
||||||
|
|
||||||
def test_new_concept_rejects_invalid_concept_type(monkeypatch, kb_dir):
|
def test_new_concept_rejects_invalid_concept_type(monkeypatch, kb_dir):
|
||||||
|
|||||||
@@ -310,7 +310,7 @@ def test_citing_pages_via_frontmatter_and_inline(kb_dir, raw_dir):
|
|||||||
)
|
)
|
||||||
refs, block = _footnote_block(("Source - Aurora", None))
|
refs, block = _footnote_block(("Source - Aurora", None))
|
||||||
write_page(
|
write_page(
|
||||||
kb_dir / "concepts/Modbus.md",
|
kb_dir / "concepts/protocols/Modbus.md",
|
||||||
{
|
{
|
||||||
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
||||||
"modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7,
|
"modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7,
|
||||||
@@ -325,7 +325,7 @@ def test_citing_pages_via_frontmatter_and_inline(kb_dir, raw_dir):
|
|||||||
|
|
||||||
def test_page_raw_files_resolves_through_sources_and_inline(kb_dir, raw_dir):
|
def test_page_raw_files_resolves_through_sources_and_inline(kb_dir, raw_dir):
|
||||||
write_page(
|
write_page(
|
||||||
kb_dir / "concepts/Modbus.md",
|
kb_dir / "concepts/protocols/Modbus.md",
|
||||||
{
|
{
|
||||||
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
||||||
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7,
|
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7,
|
||||||
@@ -429,7 +429,7 @@ def test_lint_flags_citation_not_in_frontmatter_sources(kb_dir, raw_dir, monkeyp
|
|||||||
|
|
||||||
refs, block = _footnote_block(("Source - Aurora", None))
|
refs, block = _footnote_block(("Source - Aurora", None))
|
||||||
write_page(
|
write_page(
|
||||||
kb_dir / "concepts/Modbus.md",
|
kb_dir / "concepts/protocols/Modbus.md",
|
||||||
{
|
{
|
||||||
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
||||||
"modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7,
|
"modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7,
|
||||||
@@ -451,7 +451,7 @@ def test_lint_no_drift_when_source_declared(kb_dir, raw_dir, monkeypatch):
|
|||||||
|
|
||||||
refs, block = _footnote_block(("Source - Aurora", None))
|
refs, block = _footnote_block(("Source - Aurora", None))
|
||||||
write_page(
|
write_page(
|
||||||
kb_dir / "concepts/Modbus.md",
|
kb_dir / "concepts/protocols/Modbus.md",
|
||||||
{
|
{
|
||||||
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
"type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25",
|
||||||
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7,
|
"modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7,
|
||||||
|
|||||||
@@ -86,9 +86,34 @@ def test_get_layout_reads_entity_type_specs_own_layout_field():
|
|||||||
assert list(layout) == ["project", "system", "tool", "technology", "person"]
|
assert list(layout) == ["project", "system", "tool", "technology", "person"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_concept_layout_covers_every_declared_concept_type():
|
||||||
|
"""Gitea #59: `kb/concepts/` had no areas, so the catalog's per-area shard
|
||||||
|
threshold could never fire however large it grew. A value missing from the
|
||||||
|
layout would still be *placed* (`subtype_dir` pluralizes the fallback), but
|
||||||
|
into an area with no declared title - so the schema's enum and the layout
|
||||||
|
have to agree, and this is what checks that they do."""
|
||||||
|
layout = resolver.get_layout("types/concept.md")
|
||||||
|
assert layout is not None
|
||||||
|
assert set(layout) == set(resolver.get_enum("types/concept.md", "concept_type"))
|
||||||
|
# `dir` is structural; `title` is display text following the KB language, so
|
||||||
|
# it is checked for presence rather than wording (see the entity test above).
|
||||||
|
assert {key: spec["dir"] for key, spec in layout.items()} == {
|
||||||
|
"architecture": "architectures",
|
||||||
|
"pattern": "patterns",
|
||||||
|
"protocol": "protocols",
|
||||||
|
"workflow": "workflows",
|
||||||
|
"decision": "decisions",
|
||||||
|
"problem": "problems",
|
||||||
|
}
|
||||||
|
assert all(spec.get("title") for spec in layout.values())
|
||||||
|
|
||||||
|
|
||||||
def test_get_layout_is_none_for_types_without_one():
|
def test_get_layout_is_none_for_types_without_one():
|
||||||
|
"""`comparison` has no subtype field at all; `source` has one and
|
||||||
|
deliberately declares no `layout:` anyway - 25 of its 29 pages carry the
|
||||||
|
same `source_type`, so splitting on it would make one area and four
|
||||||
|
splinters (Gitea #59). Having a subtype axis is not a reason to use it."""
|
||||||
assert resolver.get_layout("types/comparison.md") is None
|
assert resolver.get_layout("types/comparison.md") is None
|
||||||
assert resolver.get_layout("types/concept.md") is None
|
|
||||||
assert resolver.get_layout("types/source.md") is None
|
assert resolver.get_layout("types/source.md") is None
|
||||||
|
|
||||||
|
|
||||||
@@ -237,7 +262,7 @@ def test_subtype_dir_falls_back_for_unmapped_subtype():
|
|||||||
|
|
||||||
|
|
||||||
def test_subtype_dir_is_none_without_layout_or_subtype():
|
def test_subtype_dir_is_none_without_layout_or_subtype():
|
||||||
assert resolver.subtype_dir("types/concept.md", "workflow") is None
|
assert resolver.subtype_dir("types/source.md", "notes") is None
|
||||||
assert resolver.subtype_dir("types/entity.md", None) is None
|
assert resolver.subtype_dir("types/entity.md", None) is None
|
||||||
|
|
||||||
|
|
||||||
@@ -247,8 +272,8 @@ def test_compute_target_dir_applies_layout_subdirectory():
|
|||||||
|
|
||||||
|
|
||||||
def test_compute_target_dir_is_flat_for_a_type_without_layout():
|
def test_compute_target_dir_is_flat_for_a_type_without_layout():
|
||||||
target = resolver.compute_target_dir("types/concept.md", {"concept_type": "workflow"})
|
target = resolver.compute_target_dir("types/source.md", {"source_type": "notes"})
|
||||||
assert target == config.KB_DIR / "concepts"
|
assert target == config.KB_DIR / "sources"
|
||||||
|
|
||||||
|
|
||||||
def test_compute_target_dir_resolves_against_repo_root_for_root_repo_types():
|
def test_compute_target_dir_resolves_against_repo_root_for_root_repo_types():
|
||||||
|
|||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user