From 7f74303a00a18bb5ab7ed0ed57eca66930e9c379 Mon Sep 17 00:00:00 2001 From: Torben Nehmer Date: Tue, 8 Sep 2026 10:07:46 +0200 Subject: [PATCH] kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59) Files changed: - CHANGES.md - README.md - VERSION - kb/concepts/Ambient Environment Dependency.md - kb/concepts/Anti-Cramming Heuristic.md - kb/concepts/Audit Trail.md - kb/concepts/BM25.md - kb/concepts/Bulk Operations.md - kb/concepts/CI Integration.md - kb/concepts/COLLECTION.md - kb/concepts/CPPC.md - kb/concepts/Checkpoint Audit.md - kb/concepts/Claude Code Auto Mode.md - kb/concepts/Command Round-Trip Integrity.md - kb/concepts/Confidence Scoring.md - kb/concepts/Consolidation Tiers.md - kb/concepts/Content Quality Control.md - kb/concepts/Context Isolation.md - kb/concepts/Contradiction Resolution.md - kb/concepts/Cross-platform Agent Skills.md - kb/concepts/Crystallization.md - kb/concepts/Delete Rather Than Anonymize.md - kb/concepts/Denylist over Allowlist.md - kb/concepts/Detect-Repair Asymmetry.md - kb/concepts/Diff-Reviewable Agent Edits.md - kb/concepts/Dual Licensing by File Plan.md - kb/concepts/Entity Extraction.md - kb/concepts/Episodic Memory.md - kb/concepts/Event-Driven Automation.md - kb/concepts/Filter on Ingest.md - kb/concepts/Forgetting.md - kb/concepts/Graph Traversal.md - kb/concepts/Green Suite Blind Spot.md - kb/concepts/Hooks.md - kb/concepts/Hybrid Search.md - kb/concepts/INDEX.md - kb/concepts/Implementation Spectrum.md - kb/concepts/Index Scaling.md - kb/concepts/Issue Label Scheme.md - kb/concepts/Iteration and Cost Limits.md - kb/concepts/KB Migration.md - kb/concepts/KB Stack Versioning.md - kb/concepts/Knowledge Compounding.md - kb/concepts/Knowledge Graph.md - kb/concepts/LLM Wiki Pattern.md - kb/concepts/Lint Workflow.md - kb/concepts/MCP-Leseserver.md - kb/concepts/Mass-Update Gate.md - kb/concepts/Memory Lifecycle.md - kb/concepts/Mesh Sync.md - kb/concepts/Modbus.md - kb/concepts/Multi-Agent Collaboration.md - kb/concepts/Naming Convention Conflict.md - kb/concepts/OKF Compatibility.md - kb/concepts/Optional Instance Context File.md - kb/concepts/Personalization Plane.md - kb/concepts/Privacy and Governance.md - kb/concepts/Procedural Memory.md - kb/concepts/Publish-Remote Gate.md - kb/concepts/Quality Scoring.md - kb/concepts/Quality and Self-Correction.md - kb/concepts/RAG.md - kb/concepts/Reciprocal Rank Fusion.md - kb/concepts/SSD TRIM.md - kb/concepts/Scale Ceiling.md - kb/concepts/Self-Healing.md - kb/concepts/Semantic Lint Automation.md - kb/concepts/Semantic Memory.md - kb/concepts/Session Orientation.md - kb/concepts/Shared vs Private.md - kb/concepts/Split Merge Reclassify.md - kb/concepts/Split Threshold.md - kb/concepts/Structural Enforcement over Documented Rule.md - kb/concepts/Stub Threshold.md - kb/concepts/Supersession.md - kb/concepts/Three-Layer Architecture.md - kb/concepts/Token Economics.md - kb/concepts/Typed Relationships.md - kb/concepts/User Management.md - kb/concepts/Vector Search.md - kb/concepts/Work Coordination.md - kb/concepts/Workflow Extraction.md - kb/concepts/Workflow Orchestration.md - kb/concepts/Working Memory.md - kb/concepts/Write-Once Frontmatter Fields.md - kb/concepts/architectures/Consolidation Tiers.md - kb/concepts/architectures/Context Isolation.md - kb/concepts/architectures/Cross-platform Agent Skills.md - kb/concepts/architectures/Episodic Memory.md - kb/concepts/architectures/Hybrid Search.md - kb/concepts/architectures/Implementation Spectrum.md - kb/concepts/architectures/Knowledge Graph.md - kb/concepts/architectures/LLM Wiki Pattern.md - kb/concepts/architectures/MCP-Leseserver.md - kb/concepts/architectures/Memory Lifecycle.md - kb/concepts/architectures/OKF Compatibility.md - kb/concepts/architectures/Optional Instance Context File.md - kb/concepts/architectures/Personalization Plane.md - kb/concepts/architectures/Procedural Memory.md - kb/concepts/architectures/RAG.md - kb/concepts/architectures/Scale Ceiling.md - kb/concepts/architectures/Semantic Memory.md - kb/concepts/architectures/Three-Layer Architecture.md - kb/concepts/architectures/Token Economics.md - kb/concepts/architectures/Working Memory.md - kb/concepts/decisions/Delete Rather Than Anonymize.md - kb/concepts/decisions/Denylist over Allowlist.md - kb/concepts/decisions/Diff-Reviewable Agent Edits.md - kb/concepts/decisions/Dual Licensing by File Plan.md - kb/concepts/decisions/Issue Label Scheme.md - kb/concepts/decisions/KB Stack Versioning.md - kb/concepts/decisions/Structural Enforcement over Documented Rule.md - kb/concepts/patterns/Audit Trail.md - kb/concepts/patterns/BM25.md - kb/concepts/patterns/Command Round-Trip Integrity.md - kb/concepts/patterns/Confidence Scoring.md - kb/concepts/patterns/Contradiction Resolution.md - kb/concepts/patterns/Entity Extraction.md - kb/concepts/patterns/Filter on Ingest.md - kb/concepts/patterns/Forgetting.md - kb/concepts/patterns/Graph Traversal.md - kb/concepts/patterns/Mesh Sync.md - kb/concepts/patterns/Quality Scoring.md - kb/concepts/patterns/Reciprocal Rank Fusion.md - kb/concepts/patterns/Self-Healing.md - kb/concepts/patterns/Shared vs Private.md - kb/concepts/patterns/Typed Relationships.md - kb/concepts/patterns/Vector Search.md - kb/concepts/patterns/Work Coordination.md - kb/concepts/problems/Ambient Environment Dependency.md - kb/concepts/problems/Detect-Repair Asymmetry.md - kb/concepts/problems/Green Suite Blind Spot.md - kb/concepts/problems/Naming Convention Conflict.md - kb/concepts/problems/Write-Once Frontmatter Fields.md - kb/concepts/protocols/CPPC.md - kb/concepts/protocols/Modbus.md - kb/concepts/protocols/SSD TRIM.md - kb/concepts/workflows/Anti-Cramming Heuristic.md - kb/concepts/workflows/Bulk Operations.md - kb/concepts/workflows/CI Integration.md - kb/concepts/workflows/Checkpoint Audit.md - kb/concepts/workflows/Claude Code Auto Mode.md - kb/concepts/workflows/Content Quality Control.md - kb/concepts/workflows/Crystallization.md - kb/concepts/workflows/Event-Driven Automation.md - kb/concepts/workflows/Hooks.md - kb/concepts/workflows/Index Scaling.md - kb/concepts/workflows/Iteration and Cost Limits.md - kb/concepts/workflows/KB Migration.md - kb/concepts/workflows/Knowledge Compounding.md - kb/concepts/workflows/Lint Workflow.md - kb/concepts/workflows/Mass-Update Gate.md - kb/concepts/workflows/Multi-Agent Collaboration.md - kb/concepts/workflows/Privacy and Governance.md - kb/concepts/workflows/Publish-Remote Gate.md - kb/concepts/workflows/Quality and Self-Correction.md - kb/concepts/workflows/Semantic Lint Automation.md - kb/concepts/workflows/Session Orientation.md - kb/concepts/workflows/Split Merge Reclassify.md - kb/concepts/workflows/Split Threshold.md - kb/concepts/workflows/Stub Threshold.md - kb/concepts/workflows/Supersession.md - kb/concepts/workflows/User Management.md - kb/concepts/workflows/Workflow Extraction.md - kb/concepts/workflows/Workflow Orchestration.md - kb/index.md - kb/log.md - tools/CONTRACT.md - tools/README.md - tools/chemenu/catalog.py - tools/chemenu/commands/index_build.py - tools/chemenu/lint_core.py - tools/chemenu/tests/conftest.py - tools/chemenu/tests/test_cite_cmd.py - tools/chemenu/tests/test_git_publish.py - tools/chemenu/tests/test_index_build.py - tools/chemenu/tests/test_lint.py - tools/chemenu/tests/test_new_page.py - tools/chemenu/tests/test_provenance.py - tools/chemenu/tests/test_type_resolver.py - tools/chemenu/tests/test_xref.py - types/concept.md - types/type-spec.md --- CHANGES.md | 78 +++++++++- README.md | 23 ++- VERSION | 2 +- kb/concepts/COLLECTION.md | 27 +++- kb/concepts/INDEX.md | 127 +++++++++------- .../Consolidation Tiers.md | 0 .../{ => architectures}/Context Isolation.md | 0 .../Cross-platform Agent Skills.md | 0 .../{ => architectures}/Episodic Memory.md | 0 .../{ => architectures}/Hybrid Search.md | 0 .../Implementation Spectrum.md | 0 .../{ => architectures}/Knowledge Graph.md | 0 .../{ => architectures}/LLM Wiki Pattern.md | 0 .../{ => architectures}/MCP-Leseserver.md | 0 .../{ => architectures}/Memory Lifecycle.md | 0 .../{ => architectures}/OKF Compatibility.md | 0 .../Optional Instance Context File.md | 0 .../Personalization Plane.md | 0 .../{ => architectures}/Procedural Memory.md | 0 kb/concepts/{ => architectures}/RAG.md | 0 .../{ => architectures}/Scale Ceiling.md | 0 .../{ => architectures}/Semantic Memory.md | 0 .../Three-Layer Architecture.md | 0 .../{ => architectures}/Token Economics.md | 0 .../{ => architectures}/Working Memory.md | 0 .../Delete Rather Than Anonymize.md | 0 .../Denylist over Allowlist.md | 0 .../Diff-Reviewable Agent Edits.md | 0 .../Dual Licensing by File Plan.md | 0 .../{ => decisions}/Issue Label Scheme.md | 0 .../{ => decisions}/KB Stack Versioning.md | 0 ...ctural Enforcement over Documented Rule.md | 0 kb/concepts/{ => patterns}/Audit Trail.md | 0 kb/concepts/{ => patterns}/BM25.md | 0 .../Command Round-Trip Integrity.md | 0 .../{ => patterns}/Confidence Scoring.md | 0 .../Contradiction Resolution.md | 0 .../{ => patterns}/Entity Extraction.md | 0 .../{ => patterns}/Filter on Ingest.md | 0 kb/concepts/{ => patterns}/Forgetting.md | 0 kb/concepts/{ => patterns}/Graph Traversal.md | 0 kb/concepts/{ => patterns}/Mesh Sync.md | 0 kb/concepts/{ => patterns}/Quality Scoring.md | 0 .../{ => patterns}/Reciprocal Rank Fusion.md | 0 kb/concepts/{ => patterns}/Self-Healing.md | 0 .../{ => patterns}/Shared vs Private.md | 0 .../{ => patterns}/Typed Relationships.md | 0 kb/concepts/{ => patterns}/Vector Search.md | 0 .../{ => patterns}/Work Coordination.md | 0 .../Ambient Environment Dependency.md | 0 .../{ => problems}/Detect-Repair Asymmetry.md | 0 .../{ => problems}/Green Suite Blind Spot.md | 0 .../Naming Convention Conflict.md | 0 .../Write-Once Frontmatter Fields.md | 0 kb/concepts/{ => protocols}/CPPC.md | 0 kb/concepts/{ => protocols}/Modbus.md | 0 kb/concepts/{ => protocols}/SSD TRIM.md | 0 .../Anti-Cramming Heuristic.md | 0 .../{ => workflows}/Bulk Operations.md | 0 kb/concepts/{ => workflows}/CI Integration.md | 0 .../{ => workflows}/Checkpoint Audit.md | 0 .../{ => workflows}/Claude Code Auto Mode.md | 0 .../Content Quality Control.md | 0 .../{ => workflows}/Crystallization.md | 0 .../Event-Driven Automation.md | 0 kb/concepts/{ => workflows}/Hooks.md | 0 kb/concepts/{ => workflows}/Index Scaling.md | 0 .../Iteration and Cost Limits.md | 0 kb/concepts/{ => workflows}/KB Migration.md | 0 .../{ => workflows}/Knowledge Compounding.md | 0 kb/concepts/{ => workflows}/Lint Workflow.md | 0 .../{ => workflows}/Mass-Update Gate.md | 0 .../Multi-Agent Collaboration.md | 0 .../{ => workflows}/Privacy and Governance.md | 0 .../{ => workflows}/Publish-Remote Gate.md | 0 .../Quality and Self-Correction.md | 0 .../Semantic Lint Automation.md | 0 .../{ => workflows}/Session Orientation.md | 0 .../{ => workflows}/Split Merge Reclassify.md | 0 .../{ => workflows}/Split Threshold.md | 0 kb/concepts/{ => workflows}/Stub Threshold.md | 0 kb/concepts/{ => workflows}/Supersession.md | 0 .../{ => workflows}/User Management.md | 0 .../{ => workflows}/Workflow Extraction.md | 0 .../{ => workflows}/Workflow Orchestration.md | 0 kb/index.md | 13 +- kb/log.md | 6 + tools/CONTRACT.md | 2 +- tools/README.md | 5 +- tools/chemenu/catalog.py | 140 ++++++++++++++++++ tools/chemenu/commands/index_build.py | 100 ++----------- tools/chemenu/lint_core.py | 111 ++++++++++++++ tools/chemenu/tests/conftest.py | 4 +- tools/chemenu/tests/test_cite_cmd.py | 6 +- tools/chemenu/tests/test_git_publish.py | 2 +- tools/chemenu/tests/test_index_build.py | 28 +++- tools/chemenu/tests/test_lint.py | 132 ++++++++++++++++- tools/chemenu/tests/test_new_page.py | 13 +- tools/chemenu/tests/test_provenance.py | 8 +- tools/chemenu/tests/test_type_resolver.py | 33 ++++- tools/chemenu/tests/test_xref.py | 6 +- types/concept.md | 7 + types/type-spec.md | 10 ++ 103 files changed, 703 insertions(+), 180 deletions(-) rename kb/concepts/{ => architectures}/Consolidation Tiers.md (100%) rename kb/concepts/{ => architectures}/Context Isolation.md (100%) rename kb/concepts/{ => architectures}/Cross-platform Agent Skills.md (100%) rename kb/concepts/{ => architectures}/Episodic Memory.md (100%) rename kb/concepts/{ => architectures}/Hybrid Search.md (100%) rename kb/concepts/{ => architectures}/Implementation Spectrum.md (100%) rename kb/concepts/{ => architectures}/Knowledge Graph.md (100%) rename kb/concepts/{ => architectures}/LLM Wiki Pattern.md (100%) rename kb/concepts/{ => architectures}/MCP-Leseserver.md (100%) rename kb/concepts/{ => architectures}/Memory Lifecycle.md (100%) rename kb/concepts/{ => architectures}/OKF Compatibility.md (100%) rename kb/concepts/{ => architectures}/Optional Instance Context File.md (100%) rename kb/concepts/{ => architectures}/Personalization Plane.md (100%) rename kb/concepts/{ => architectures}/Procedural Memory.md (100%) rename kb/concepts/{ => architectures}/RAG.md (100%) rename kb/concepts/{ => architectures}/Scale Ceiling.md (100%) rename kb/concepts/{ => architectures}/Semantic Memory.md (100%) rename kb/concepts/{ => architectures}/Three-Layer Architecture.md (100%) rename kb/concepts/{ => architectures}/Token Economics.md (100%) rename kb/concepts/{ => architectures}/Working Memory.md (100%) rename kb/concepts/{ => decisions}/Delete Rather Than Anonymize.md (100%) rename kb/concepts/{ => decisions}/Denylist over Allowlist.md (100%) rename kb/concepts/{ => decisions}/Diff-Reviewable Agent Edits.md (100%) rename kb/concepts/{ => decisions}/Dual Licensing by File Plan.md (100%) rename kb/concepts/{ => decisions}/Issue Label Scheme.md (100%) rename kb/concepts/{ => decisions}/KB Stack Versioning.md (100%) rename kb/concepts/{ => decisions}/Structural Enforcement over Documented Rule.md (100%) rename kb/concepts/{ => patterns}/Audit Trail.md (100%) rename kb/concepts/{ => patterns}/BM25.md (100%) rename kb/concepts/{ => patterns}/Command Round-Trip Integrity.md (100%) rename kb/concepts/{ => patterns}/Confidence Scoring.md (100%) rename kb/concepts/{ => patterns}/Contradiction Resolution.md (100%) rename kb/concepts/{ => patterns}/Entity Extraction.md (100%) rename kb/concepts/{ => patterns}/Filter on Ingest.md (100%) rename kb/concepts/{ => patterns}/Forgetting.md (100%) rename kb/concepts/{ => patterns}/Graph Traversal.md (100%) rename kb/concepts/{ => patterns}/Mesh Sync.md (100%) rename kb/concepts/{ => patterns}/Quality Scoring.md (100%) rename kb/concepts/{ => patterns}/Reciprocal Rank Fusion.md (100%) rename kb/concepts/{ => patterns}/Self-Healing.md (100%) rename kb/concepts/{ => patterns}/Shared vs Private.md (100%) rename kb/concepts/{ => patterns}/Typed Relationships.md (100%) rename kb/concepts/{ => patterns}/Vector Search.md (100%) rename kb/concepts/{ => patterns}/Work Coordination.md (100%) rename kb/concepts/{ => problems}/Ambient Environment Dependency.md (100%) rename kb/concepts/{ => problems}/Detect-Repair Asymmetry.md (100%) rename kb/concepts/{ => problems}/Green Suite Blind Spot.md (100%) rename kb/concepts/{ => problems}/Naming Convention Conflict.md (100%) rename kb/concepts/{ => problems}/Write-Once Frontmatter Fields.md (100%) rename kb/concepts/{ => protocols}/CPPC.md (100%) rename kb/concepts/{ => protocols}/Modbus.md (100%) rename kb/concepts/{ => protocols}/SSD TRIM.md (100%) rename kb/concepts/{ => workflows}/Anti-Cramming Heuristic.md (100%) rename kb/concepts/{ => workflows}/Bulk Operations.md (100%) rename kb/concepts/{ => workflows}/CI Integration.md (100%) rename kb/concepts/{ => workflows}/Checkpoint Audit.md (100%) rename kb/concepts/{ => workflows}/Claude Code Auto Mode.md (100%) rename kb/concepts/{ => workflows}/Content Quality Control.md (100%) rename kb/concepts/{ => workflows}/Crystallization.md (100%) rename kb/concepts/{ => workflows}/Event-Driven Automation.md (100%) rename kb/concepts/{ => workflows}/Hooks.md (100%) rename kb/concepts/{ => workflows}/Index Scaling.md (100%) rename kb/concepts/{ => workflows}/Iteration and Cost Limits.md (100%) rename kb/concepts/{ => workflows}/KB Migration.md (100%) rename kb/concepts/{ => workflows}/Knowledge Compounding.md (100%) rename kb/concepts/{ => workflows}/Lint Workflow.md (100%) rename kb/concepts/{ => workflows}/Mass-Update Gate.md (100%) rename kb/concepts/{ => workflows}/Multi-Agent Collaboration.md (100%) rename kb/concepts/{ => workflows}/Privacy and Governance.md (100%) rename kb/concepts/{ => workflows}/Publish-Remote Gate.md (100%) rename kb/concepts/{ => workflows}/Quality and Self-Correction.md (100%) rename kb/concepts/{ => workflows}/Semantic Lint Automation.md (100%) rename kb/concepts/{ => workflows}/Session Orientation.md (100%) rename kb/concepts/{ => workflows}/Split Merge Reclassify.md (100%) rename kb/concepts/{ => workflows}/Split Threshold.md (100%) rename kb/concepts/{ => workflows}/Stub Threshold.md (100%) rename kb/concepts/{ => workflows}/Supersession.md (100%) rename kb/concepts/{ => workflows}/User Management.md (100%) rename kb/concepts/{ => workflows}/Workflow Extraction.md (100%) rename kb/concepts/{ => workflows}/Workflow Orchestration.md (100%) create mode 100644 tools/chemenu/catalog.py diff --git a/CHANGES.md b/CHANGES.md index d34a687..a5efe84 100644 --- a/CHANGES.md +++ b/CHANGES.md @@ -35,7 +35,7 @@ dev-checkout concern - readable here, never shipped as something to parse. --- -## 4.8.0-beta.5 - 2026-09-05 - update entity naming conventions to use singular form for consistency +## 4.8.0-beta.6 - 2026-09-08 - kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59) **Author:** Torben Nehmer @@ -46,6 +46,7 @@ dev-checkout concern - readable here, never shipped as something to parse. - raw accept: incoming/ als abgeleiteter Rohablage-Eingang (schliesst #58) - raw accept: Stem-Eindeutigkeit im Typverzeichnis erzwingen, --replaces als einziger Weg daran vorbei (schliesst #64) - update entity naming conventions to use singular form for consistency +- kb/concepts/ bekommt Areas: layout: fuer concept, Area-Titel aus jedem Type-Spec, Schwellen-Empfehlung im lint (schliesst #59) Das Label `status/incoming` gibt es seit heute in Gitea: der Mensch legt einen @@ -301,6 +302,81 @@ keine Kompatibilitätsfrage, solange sie vor `4.8.0` landet. Schließt #64. +**`kb/concepts/` bekommt Areas** (#59): Sharden ist längst automatisch — +`index_build.SHARD_THRESHOLD = 50`, hergeleitet aus der wikieigenen Seite +`Index Scaling` — aber es passiert **pro Area**, und eine Area legt niemand an. +`kb/concepts/` hatte keine, also war die Schwelle dort ein toter Wert: 80 Seiten +in einer einzigen Tabelle, weit über der eigenen Grenze, ohne dass je etwas +gefeuert hätte. Die Ursache war eine Asymmetrie in den Type-Specs — `entity` +deklarierte ein `layout:`, `concept` nicht, obwohl das Subtype-Feld fertig dalag. + +`types/concept.md` deklariert es jetzt für alle sechs `concept_type`-Werte +(`architectures/`, `patterns/`, `protocols/`, `workflows/`, `decisions/`, +`problems/`). Für `source` bewusst **nicht**: 25 von 29 Seiten sind `notes`, die +Aufteilung ergäbe eine Area und vier Splitter, und `kb/sources/` liegt mit 29 +Seiten ohnehin unter der Schwelle. Ein Subtype-Feld zu haben ist kein Grund, es +als Achse zu benutzen. + +Zwei Dinge im Code, beide Folgen desselben Befunds. `_area_titles()` in +`index_build.py` löste `entity` fest über `find_type_by_name("entity")` auf und +las nur dessen `layout:` — jeder zweite Typ mit einem `layout:` hätte +`.title()`-Namen auf dem Verzeichnisnamen bekommen statt der deklarierten Titel. +Es liest jetzt jedes Type-Spec, und zwar **pro Collection** geschlüsselt, damit +zwei Typen denselben Area-Namen für Verschiedenes benutzen dürfen. Und der neue +`lint`-Befund meldet eine Collection über der Schwelle **ohne** Areas, mit der +Verteilung ihres Subtype-Felds — als Empfehlung, nicht als Failure, und nur +dann, wenn die Aufteilung jede entstehende Area unter die Schwelle drückt. Das +begrenzt sich selbst in beide Richtungen: `kb/comparisons/` mit einer Seite +feuert nie, und die schlechte Aufteilung nach `source_type` unterbleibt von +allein, ohne dass der Check etwas über Sources wüsste. + +Zwei Dinge fielen unterwegs an, die das Issue nicht vorhergesehen hatte. +`lint_core.py` durfte `SHARD_THRESHOLD`/`group_pages` nicht aus +`commands/index_build.py` importieren — `test_api.py` prüft strukturell, dass +`chemenu.api` kein Modul unter `chemenu.commands` lädt, und der Import hätte den +ganzen CLI-Kopf mitgezogen. Die Gruppierung liegt deshalb neu in +`tools/chemenu/catalog.py`, entlang derselben Linie wie `lint_core.py`: +Korpusform hier, Darstellung dort. Und `_anchor()` strich mit `[^a-z0-9\s-]` +jeden Nicht-ASCII-Buchstaben ersatzlos — die Karte verlinkte auf `#ablufe`, +während die Überschrift im Shard `#abläufe` heißt. Vorher fiel das keinem auf, +weil alle Entity-Area-Titel zufällig ASCII sind; `Abläufe` ist der erste, der es +nicht ist. + +Auf dieser Instanz angewendet: `wikitool move --reconcile` hat alle 80 +Concept-Seiten in ihre Area gezogen, `migrate verify --from HEAD` bestätigt +`182 compared, 0 added, 0 removed, 80 moved, 0 findings` — kein Titel, kein +Body, kein Frontmatter-Feld angefasst. `index rebuild` erzeugt sechs Areas +(Abläufe 28, Architekturen 20, Muster 17, Entscheidungen 7, Problemstellungen 5, +Protokolle 3); keine über der Schwelle, also kein eigener Shard, und die +Schwelle wirkt wieder als Schwelle. + +Geändert: `types/concept.md` (`layout:`), `types/type-spec.md` (wann ein +`layout:` sich lohnt), `tools/chemenu/catalog.py` (neu), +`tools/chemenu/commands/index_build.py` (`area_titles`, `_anchor`), +`tools/chemenu/lint_core.py` (`unsharded_collections`), +`tools/chemenu/tests/` (Fixture-Concept liegt jetzt in seiner Area, plus neun +neue Tests), `tools/CONTRACT.md`, `tools/README.md`, `README.md`, +`kb/concepts/COLLECTION.md`, sowie die 80 bewegten Seiten unter `kb/concepts/`. + +**MINOR**, nicht MAJOR: der Umzug ist ein **Angebot**, kein Zwang. Eine +bestehende Instanz, die `move --reconcile` nicht laufen lässt, bleibt +funktionsfähig — `group_pages` liest das Dateisystem, nicht das `layout:`, also +landen flache Bestandsseiten in der Area „All" und neu angelegte in ihrer +eigenen; beides rendert. Der gemischte Zustand meldet sich als `lint`-Befund +*Misplaced Pages*, der seit jeher advisory ist. Und ein Downgrade auf einen +Stack ohne dieses `layout:` funktioniert weiter: die Verzeichnisse bleiben +Verzeichnisse, nur die Anzeigetitel fallen auf `.title()` zurück. Kosmetik, kein +Bruch der Austauschbarkeit in beiden Richtungen. + +Schließt #59. + +**Nachzug an `63b4bb8`:** die Umstellung der Namenskonvention auf `HA Integration` +hatte in `README.md` das Gegenbeispiel verloren — die Zeile las +``Use singular for entities: `HA Integration.md` (not `HA Integration.md`)``, beide +Seiten des „not" identisch, also eine Regel ohne Fall, an dem sie greift. +`kb/CONVENTIONS.md` und `kb/entities/COLLECTION.md` hatten im selben Commit das +korrekte Paar bekommen; `README.md` zieht jetzt mit `HA Integrations.md` nach. + --- ## 4.7.4 - 2026-09-04 - bootstrap.md nennt den session-id-WARN nach frischem Bootstrap explizit als erwartet diff --git a/README.md b/README.md index a766fa4..50e678e 100644 --- a/README.md +++ b/README.md @@ -105,8 +105,14 @@ chemenu/ │ │ ├── tools/ # own INDEX.md once past 50 pages │ │ ├── technologies/ │ │ └── people/ -│ ├── concepts/ # COLLECTION.md - architectures, patterns, protocols -│ ├── sources/ # COLLECTION.md - source summaries +│ ├── concepts/ # COLLECTION.md + INDEX.md + areas below +│ │ ├── architectures/ +│ │ ├── patterns/ +│ │ ├── protocols/ +│ │ ├── workflows/ +│ │ ├── decisions/ +│ │ └── problems/ +│ ├── sources/ # COLLECTION.md - source summaries, no areas by choice │ └── comparisons/ # COLLECTION.md - comparison pages ├── work/ # WORKSHOP: one directory per multi-session run, tracked │ └── CONTRACT.md # Run keys, required files, how a run closes @@ -130,7 +136,16 @@ A directory under `kb/` is a **collection** exactly when it holds a `COLLECTION. subdirectory inside one is an **area** that inherits it, and that is as deep as a page goes - nothing nests below an area, because the generated catalog reads exactly two path segments under `kb/` and would fold a deeper page into the area silently (`kb/CONTRACT.md` § Collections -has the rule; `wikitool lint` reports a violation as a hard error). `COLLECTION.md` never +has the rule; `wikitool lint` reports a violation as a hard error). + +Which areas a collection has is not chosen per page: a type-spec's `layout:` maps its subtype +field onto directories, and `wikitool new` writes the page straight into the one its subtype +names. That is also what makes the catalog's shard threshold do anything - `index rebuild` +splits **per area**, so a collection with no areas keeps one table however large it grows. +`wikitool lint` reports such a collection once it is past the threshold, as a recommendation +rather than an error, together with the split its subtype field would produce; it stays quiet +when the split would not actually help. `kb/sources/` is the worked example of the second case +and deliberately has no areas. `COLLECTION.md` never appears outside `kb/` - the other layers carry a `CONTRACT.md` or a root type-spec instead. A stage may carry both a `README.md` and a `CONTRACT.md`: they have different readers. The README is for humans working *on* that layer, the contract is what binds an agent working *with* it. @@ -247,7 +262,7 @@ Ingest incoming/notes/my-notes.md ### Naming - Use human-readable titles with spaces for files: `Hybrid Search.md`, not kebab-case -- Use singular for entities: `HA Integration.md` (not `HA Integration.md`) +- Use singular for entities: `HA Integration.md` (not `HA Integrations.md`) - Use wikilinks matching the file name exactly: `[[Entity Name]]` - **Titles follow the subject's own established name, not the wiki's language.** `Act Runner` and `GitOps Ownership Model` keep theirs. A title is the only identifier a page has - it also lives diff --git a/VERSION b/VERSION index 5a4f739..16e8505 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -4.8.0-beta.5 +4.8.0-beta.6 diff --git a/kb/concepts/COLLECTION.md b/kb/concepts/COLLECTION.md index 2fc230e..d125d25 100644 --- a/kb/concepts/COLLECTION.md +++ b/kb/concepts/COLLECTION.md @@ -25,7 +25,32 @@ tone, relationship labels, the confidence rubric. Neither is restated here. ## Types offered -`concept` (`tools/wikitool types describe concept`). +`concept` (`tools/wikitool types describe concept`). Das Feld `concept_type:` +wählt die Area: + +| Area | Hält | +|------|------| +| `architectures/` | Aufbau und Struktur: wie ein System geschnitten ist und warum die Schnitte dort liegen | +| `patterns/` | Wiederverwendbare Lösungsformen, die über mehr als einen Gegenstand hinweg gelten | +| `protocols/` | Kommunikationsprotokolle und Standards, in ihrer üblichen Schreibweise benannt | +| `workflows/` | Abläufe und Prozesse, die projektübergreifend wiederkehren | +| `decisions/` | Architektur- und Entwurfsentscheidungen (siehe unten) | +| `problems/` | Wiederkehrende Problemstellungen und ihre Lösungsansätze | + +Das sind Areas, keine Collections: sie erben diesen Contract und tragen keine +eigene `COLLECTION.md`. + +Die Zuordnung trifft niemand von Hand — sie steht als `layout:` in +`types/concept.md`, und `wikitool new` legt eine neue Seite direkt dort ab. +Eine Seite, die anderswo liegt, meldet `wikitool lint` als *misplaced*; +`wikitool move --page ""` bringt sie an ihren berechneten Ort. + +Die Aufteilung ist keine Geschmacksfrage, sondern das, was die Shard-Schwelle +des Katalogs überhaupt wirksam macht: `index rebuild` teilt **pro Area**, und +eine Collection ohne Areas teilt sich nie — mit 80 Seiten in einer einzigen +Tabelle war die Schwelle hier ein toter Wert (Gitea #59). Keine der sechs +Areas liegt derzeit über der Schwelle, also bekommt auch keine einen eigenen +Shard; wächst eine hinein, passiert das ohne Zutun. ## Decisions diff --git a/kb/concepts/INDEX.md b/kb/concepts/INDEX.md index 58e12e3..d5e6cf2 100644 --- a/kb/concepts/INDEX.md +++ b/kb/concepts/INDEX.md @@ -4,88 +4,113 @@ 80 page(s). Regenerated by `wikitool index rebuild`. -## All +## Abläufe | Page | Type | Summary | Last Modified | |------|------|---------|----------------| -| [[Ambient Environment Dependency]] | problem | Fehlerklasse, in der ein Test gruen ist, weil die Maschine zufaellig passt statt weil der Code stimmt - abgegrenzt gegen den Green Suite Blind Spot, belegt an vier Faellen unter Gitea-Issue #8 | 2026-08-31 | | [[Anti-Cramming Heuristic]] | workflow | Regel gegen überladene Seiten: ab dem dritten Absatz zu einem Unterthema eine eigene Seite anlegen | 2026-08-29 | -| [[Audit Trail]] | pattern | Unveränderliches chronologisches Log aller Wiki-Operationen (Ingest, Bearbeitung, Löschung, Abfrage) mit Zeitstempel, Akteur, Ziel und Änderungsbeschreibung. | 2026-08-29 | -| [[BM25]] | pattern | Schlüsselwortbasiertes Retrieval-Verfahren, das über Termfrequenz, inverse Dokumentfrequenz und Stemming exakte oder teilweise Übereinstimmungen findet. | 2026-08-29 | | [[Bulk Operations]] | workflow | Umkehrbare, protokollierte Operationen zum Massenlöschen, Exportieren, Zusammenführen oder Archivieren von Wiki-Inhalten, mit Freigabepflicht und Undo. | 2026-08-29 | | [[Checkpoint Audit]] | workflow | Regelmäßiger Qualitätsrhythmus: Index und Backlinks alle 15 Einträge neu aufbauen, auf 0 neue Artikel prüfen, die 3 meistgeänderten erneut lesen | 2026-08-29 | | [[CI Integration]] | workflow | CI/CD-Hooks vor dem Publish: ci.yml (Push/PR, Stack-Pfade, seit 1.8.1 mit Coverage-Messung ohne Schwelle) und nightly.yml (Zeitplan, schliesst die paths-ignore-Luecke fuer Content-Drift; schedule-Ausloesung seit 2026-09-01 bestaetigt) setzen Quality Gates durch | 2026-09-01 | | [[Claude Code Auto Mode]] | workflow | auto-Berechtigungsmodus von Claude Code: ein Klassifikator genehmigt Aktionen vor der Ausfuehrung statt nachzufragen; die Beschreibung stammt weit ueberwiegend aus zweiter Hand ueber einen Doku-Subagenten | 2026-08-31 | -| [[Command Round-Trip Integrity]] | pattern | Anforderung, dass zwei Befehle auf derselben Datei in jeder Reihenfolge zusammenpassen und jeder erzeugte Zustand einen Gegenbefehl hat - 2026-08-31 in wikitool zweimal verletzt | 2026-08-31 | -| [[Confidence Scoring]] | pattern | Mechanismus, der faktischen Aussagen quantitative Werte nach Quellenzahl, Aktualität, Qualität und Bestätigung zuweist, um gut gestütztes Wissen zu erkennen. | 2026-08-29 | -| [[Consolidation Tiers]] | architecture | Hierarchische Speicherarchitektur, die Informationen durch zunehmend verdichtete Schichten vom Working Memory bis zum Semantic und Procedural Memory befördert. | 2026-08-29 | | [[Content Quality Control]] | workflow | Regeln und Schwellenwerte für die Seitenqualität: Mindestumfang für Stubs, Aufteilungsschwellen und Zielwerte für die Zeilenzahl | 2026-08-29 | -| [[Context Isolation]] | architecture | Grundsatz, für jede Aufgabe nur den jeweils benötigten Kontext zu laden | 2026-08-29 | -| [[Contradiction Resolution]] | pattern | Automatisches Erkennen und Auflösen widersprüchlicher Aussagen anhand von Konfidenz, Aktualität und Autorität der Quelle. | 2026-08-29 | -| [[CPPC]] | protocol | Hardwareschnittstelle Collaborative Processor Performance Control für feingranulares CPU-Power-Management zwischen Betriebssystem und AMD-Prozessor. | 2026-08-29 | -| [[Cross-platform Agent Skills]] | architecture | Architektur fuer Agent-Skills, die ueber mehrere LLM-Werkzeuge hinweg funktionieren; in Chemenu selbst am 2026-08-04 umgesetzt und ueberprueft | 2026-09-01 | | [[Crystallization]] | workflow | Verdichten abgeschlossener Erkundungen, Debugging-Sitzungen und Recherchen zu strukturierten Wiki-Auszügen als eigenständige Wissensquellen. | 2026-08-29 | -| [[Delete Rather Than Anonymize]] | decision | Private Korpusinhalte per Loeschung entfernen statt zu anonymisieren: ein Seitentitel ist der einzige Identifier eines Wikis, Umbenennen ist die volle page-lifecycle-Prozedur je Seite, Loeschen ist ein unterstuetztes Kommando. | 2026-09-01 | -| [[Denylist over Allowlist]] | decision | Entscheidung, schreibbare Felder als Schema minus kurzer Sperrliste zu bestimmen statt als gepflegte Positivliste, weil die Positivliste eine zweite Kopie des Schemas waere | 2026-08-31 | -| [[Detect-Repair Asymmetry]] | problem | Werkzeugluecke, in der ein Check einen Defekt zuverlaessig meldet, aber kein Befehl ihn behebt - womit die Handeditierung der einzige verbleibende Ausweg ist | 2026-08-31 | -| [[Diff-Reviewable Agent Edits]] | decision | Entscheidung, Dateiaenderungen ueber Edit/Write statt ueber Shell-Heredocs zu fahren, weil nur das erste eine pruefbare Diff hinterlaesst | 2026-08-31 | -| [[Dual Licensing by File Plan]] | decision | Ein Repo mit Code- und Inhaltsanteil erhaelt zwei Lizenzen; die Grenze zwischen ihnen ist kein zweiter, gepflegter Pfadkatalog, sondern der ohnehin vorhandene Dateiplan des Distributionswerkzeugs. | 2026-09-01 | -| [[Entity Extraction]] | pattern | Erkennen und Strukturieren von Entities (Personen, Projekte, Bibliotheken, Concepts, Dateien, Entscheidungen, Systeme, Werkzeuge) samt typspezifischer Attribute aus Rohquellen. | 2026-08-29 | -| [[Episodic Memory]] | architecture | Speicherschicht für verdichtete Sitzungszusammenfassungen und Befunde; Brücke zwischen rohem Working Memory und langlebigem Semantic Memory. | 2026-08-29 | | [[Event-Driven Automation]] | workflow | Muster, das automatische Auslöser an Wiki-Lebenszyklusereignisse hängt, um manuellen Pflegeaufwand und das Risiko der Verwahrlosung zu senken. | 2026-08-29 | -| [[Filter on Ingest]] | pattern | Automatisches Erkennen und Entfernen sensibler Daten (API-Schlüssel, Token, Credentials, personenbezogene Daten) vor der Aufnahme ins Wiki, per Regex und ML-Erkennung. | 2026-08-29 | -| [[Forgetting]] | pattern | Muster zur Wissensbindung, das selten abgerufene Fakten schrittweise zurückstuft, modelliert nach der Ebbinghausschen Vergessenskurve. | 2026-08-29 | -| [[Graph Traversal]] | pattern | Verfahren, verbundene Entities im Wissensgraphen über typisierte Beziehungen (uses, depends-on, contradicts, caused) zu finden und strukturelle Fragen zu beantworten. | 2026-08-29 | -| [[Green Suite Blind Spot]] | problem | Defekt, der eine vollstaendig gruene Testsuite ueberlebt, weil nie ein Test das richtige Verhalten behauptet hat - belegt an drei prio/1-2-Defekten (Round-Trip, Zitat-Notation-als-Code, Zitat-Limit) | 2026-08-31 | | [[Hooks]] | workflow | Mechanismus von Event-Listenern, der bei Wiki-Lebenszyklusereignissen wie Quellen-Ingest, Seitenänderung und Sitzungsende automatisch Aktionen auslöst. | 2026-08-29 | -| [[Hybrid Search]] | architecture | Multimodale Suche, die BM25-Schlüsselwortabgleich, Vektor-Embeddings und Graph Traversal verbindet, um Wissensabruf im Wiki skalierbar zu machen. | 2026-08-29 | -| [[Implementation Spectrum]] | architecture | Modularer Einführungspfad für die Funktionen von LLM Wiki v2, vom minimal tragfähigen Wiki bis zur vollen Umsetzung mit Automatisierung und Governance. | 2026-08-29 | | [[Index Scaling]] | workflow | Skalierungsregeln für Indexseiten: Tabellenabschnitte ab 50 Einträgen teilen, ab 200 Seiten _meta/topic-map.md anlegen | 2026-08-29 | -| [[Issue Label Scheme]] | decision | Pflicht-Labelschema fuer das Gitea-Board: vier Achsen (area/kind/prio/size) plus seit 2026-09-04 drei optionale status/-Flags, darunter status/incoming fuer unausgearbeitete Stubs, die die Vier-Achsen-Pflicht aussetzen statt sie zu ergaenzen; die Regel liegt in instructions/dev/, weil sie keine ausgelieferte Instanz erreichen darf | 2026-09-04 | | [[Iteration and Cost Limits]] | workflow | Im Code durchgesetzte Obergrenze von 60 wikitool-Aufrufen je Session, Loop-Breaker bei 3 identischen Wiederholungen, Slot-Erstattung, ein gemessenes Kalibrierungsband, und Retrieval sowie der MCP-Leseserver bleiben ausgenommen | 2026-09-02 | | [[KB Migration]] | workflow | Migration des KB-Inhalts entlang einer geordneten Versionskette; abgegrenzt gegen offene Instanz-Aktionen, die in den doctor-Check gehoeren statt in die Kette | 2026-08-31 | -| [[KB Stack Versioning]] | decision | Semantische Versionierung des Wiki-Stacks: VERSION beschreibt die Maschinerie, Kompatibilitaet (Drop-in-Ersatz) und Inhaltsmigration sind seit 2.5.0 getrennte, unabhaengig geprueft Fragen | 2026-09-02 | | [[Knowledge Compounding]] | workflow | Effekt, bei dem Wissen im Wiki an Wert gewinnt, weil jede neue Quelle an bestehende, untereinander verwiesene Seiten anknüpft und sie ergänzt. | 2026-08-29 | -| [[Knowledge Graph]] | architecture | Typisierte Schicht aus Entities und Beziehungen über den Wiki-Seiten, die eine reichere Wissensdarstellung und graphbasierte Abfragen ermöglicht. | 2026-08-29 | | [[Lint Workflow]] | workflow | Deterministischer Health-Check rund um wikitool lint; seit 1.7.2 maskiert es Code vor dem Notation-Match und zaehlt Zitat-Bloecke statt Zeilen | 2026-09-01 | -| [[LLM Wiki Pattern]] | architecture | Methodik für persönliches Wissensmanagement, bei der ein LLM aus Rohquellen ein dauerhaftes Wiki aufbaut - Wissen wird kompiliert statt per RAG neu hergeleitet. | 2026-08-29 | | [[Mass-Update Gate]] | workflow | Mass-Update Gate: publish endet mit 42 (Freigabe durch den Menschen noetig) ab 10 gezaehlten Dateien; generierte Dateien und work/ werden committet\, aber seit 1.5.0 nicht gezaehlt; freigegeben per --confirm | 2026-09-01 | +| [[Multi-Agent Collaboration]] | workflow | Wissensmanagement mit mehreren Agenten; erweitert das LLM-Wiki-Muster um Mesh Sync, die Trennung von geteiltem und privatem Wissen und leichtgewichtige Arbeitskoordination. | 2026-08-29 | +| [[Privacy and Governance]] | workflow | Rahmenwerk zur Absicherung von Wiki-Inhalten über Datenfilterung beim Ingest, Audit-Trail-Protokollierung und umkehrbare Massenoperationen. | 2026-08-29 | +| [[Publish-Remote Gate]] | workflow | Drittes, im Code durchgesetztes Gate: publish bricht mit Exit 42 ab, wenn die aufgeloeste Push-URL nicht in einer optionalen, gitignoreten Allowlist steht; doctor benennt seit 2026-09-02 den Gate-Zustand statt nur die Dateiexistenz | 2026-09-02 | +| [[Quality and Self-Correction]] | workflow | Automatische Qualitätssicherung für Wikis mit Inhaltsbewertung, Selbstheilung und Widerspruchserkennung. | 2026-08-29 | +| [[Semantic Lint Automation]] | workflow | Maschinelle Heuristiken zur Priorisierung der semantischen Prüfung: veraltete Aussagen, hohe Änderungsdichte und schwache Verlinkung | 2026-08-29 | +| [[Session Orientation]] | workflow | Verbindliche Vorabprüfung, die vor Query- und Update-Operationen einen Kontextbericht erzeugt (Index, jüngste Logs, Umfang) | 2026-08-29 | +| [[Split Merge Reclassify]] | workflow | Eigene Befehle zum Teilen, Zusammenführen und Umklassifizieren von Seiten, mit automatischer Korrektur von Links und Frontmatter | 2026-08-29 | +| [[Split Threshold]] | workflow | Maximale Seitengröße, ab der eine Aufteilung empfohlen wird (Farza: >120-150 Zeilen, Pascalandy: 200 Zeilen) | 2026-08-29 | +| [[Stub Threshold]] | workflow | Mindestumfang, ab dem eine Wiki-Seite nicht mehr als Stub gilt (Farza: ≥3 Sätze oder 15 Zeilen) | 2026-08-29 | +| [[Supersession]] | workflow | Ablösen alten Wissens durch neue, widersprechende Information; gibt dem Wiki eine Versionierung mit ausdrücklicher Verknüpfung und Erhalt der Historie. | 2026-08-29 | +| [[User Management]] | workflow | Linux-Ablauf zum Anlegen, Ändern, Überwachen und Löschen von Benutzerkonten mit useradd, usermod und userdel, samt Gruppenverwaltung und sudoers-Konfiguration. | 2026-08-29 | +| [[Workflow Extraction]] | workflow | Herauslösen von Workflow-Abschnitten aus monolithischer Dokumentation | 2026-09-01 | +| [[Workflow Orchestration]] | workflow | Orchestrierte Einzelbefehle für vollständige Operationen (ingest run, lint run, update run) mit Dry-Run-Vorschau vor dem Schreiben | 2026-08-29 | + +## Architekturen + +| Page | Type | Summary | Last Modified | +|------|------|---------|----------------| +| [[Consolidation Tiers]] | architecture | Hierarchische Speicherarchitektur, die Informationen durch zunehmend verdichtete Schichten vom Working Memory bis zum Semantic und Procedural Memory befördert. | 2026-08-29 | +| [[Context Isolation]] | architecture | Grundsatz, für jede Aufgabe nur den jeweils benötigten Kontext zu laden | 2026-08-29 | +| [[Cross-platform Agent Skills]] | architecture | Architektur fuer Agent-Skills, die ueber mehrere LLM-Werkzeuge hinweg funktionieren; in Chemenu selbst am 2026-08-04 umgesetzt und ueberprueft | 2026-09-01 | +| [[Episodic Memory]] | architecture | Speicherschicht für verdichtete Sitzungszusammenfassungen und Befunde; Brücke zwischen rohem Working Memory und langlebigem Semantic Memory. | 2026-08-29 | +| [[Hybrid Search]] | architecture | Multimodale Suche, die BM25-Schlüsselwortabgleich, Vektor-Embeddings und Graph Traversal verbindet, um Wissensabruf im Wiki skalierbar zu machen. | 2026-08-29 | +| [[Implementation Spectrum]] | architecture | Modularer Einführungspfad für die Funktionen von LLM Wiki v2, vom minimal tragfähigen Wiki bis zur vollen Umsetzung mit Automatisierung und Governance. | 2026-08-29 | +| [[Knowledge Graph]] | architecture | Typisierte Schicht aus Entities und Beziehungen über den Wiki-Seiten, die eine reichere Wissensdarstellung und graphbasierte Abfragen ermöglicht. | 2026-08-29 | +| [[LLM Wiki Pattern]] | architecture | Methodik für persönliches Wissensmanagement, bei der ein LLM aus Rohquellen ein dauerhaftes Wiki aufbaut - Wissen wird kompiliert statt per RAG neu hergeleitet. | 2026-08-29 | | [[MCP-Leseserver]] | architecture | Zweiter Konsument von chemenu ueber MCP: search/types/describe_type/lint/status auf chemenu.api.Corpus, strukturell ohne Schreibpfad, jede Antwort trage einen Commit-Stempel. | 2026-09-02 | | [[Memory Lifecycle]] | architecture | Architektur des Wissenslebenszyklus mit Confidence Scoring, Supersession, Forgetting und Consolidation Tiers zur Pflege von Fakten über die Zeit. | 2026-08-29 | -| [[Mesh Sync]] | pattern | Abgleichsmechanismus, der Beobachtungen paralleler Agenten in ein gemeinsames Wiki überführt; Last-Write-Wins mit Konflikterkennung und manuellem Eingriff. | 2026-08-29 | -| [[Modbus]] | protocol | Industrielles Kommunikationsprotokoll von 1979 zur Anbindung speicherprogrammierbarer Steuerungen und Geräte über serielle oder TCP-Netze. | 2026-08-29 | -| [[Multi-Agent Collaboration]] | workflow | Wissensmanagement mit mehreren Agenten; erweitert das LLM-Wiki-Muster um Mesh Sync, die Trennung von geteiltem und privatem Wissen und leichtgewichtige Arbeitskoordination. | 2026-08-29 | -| [[Naming Convention Conflict]] | problem | Widerspruch zwischen README.md (kebab-case) und AGENTS.md (lesbar mit Leerzeichen), der zu Drift bei der Validierung führt | 2026-08-29 | | [[OKF Compatibility]] | architecture | Optionale Kompatibilität zum Open Knowledge Framework als Export- und Prüfmodus, ohne das interne Modell zu ersetzen | 2026-08-29 | | [[Optional Instance Context File]] | architecture | Muster fuer eine Datei, die eine Instanz ueber ihre Umgebung informiert, ohne Betriebsvoraussetzung zu sein: Health-Check meldet ohne zu scheitern, pro Checkout statt pro Repo | 2026-08-31 | | [[Personalization Plane]] | architecture | Schicht fuer Instanz-Identitaet: USER.md/SOUL.md werden als Template ausgeliefert, im Setup-Interview woertlich befuellt und vom doctor-Check auf fehlend wie unbefuellt geprueft | 2026-08-31 | -| [[Privacy and Governance]] | workflow | Rahmenwerk zur Absicherung von Wiki-Inhalten über Datenfilterung beim Ingest, Audit-Trail-Protokollierung und umkehrbare Massenoperationen. | 2026-08-29 | | [[Procedural Memory]] | architecture | Langlebigste Speicherschicht für Abläufe, Muster, bewährte Vorgehensweisen und Rezepte, gewonnen aus wiederholten semantischen Beobachtungen. | 2026-08-29 | -| [[Publish-Remote Gate]] | workflow | Drittes, im Code durchgesetztes Gate: publish bricht mit Exit 42 ab, wenn die aufgeloeste Push-URL nicht in einer optionalen, gitignoreten Allowlist steht; doctor benennt seit 2026-09-02 den Gate-Zustand statt nur die Dateiexistenz | 2026-09-02 | -| [[Quality and Self-Correction]] | workflow | Automatische Qualitätssicherung für Wikis mit Inhaltsbewertung, Selbstheilung und Widerspruchserkennung. | 2026-08-29 | -| [[Quality Scoring]] | pattern | Quantitative Bewertung aller vom LLM geschriebenen Inhalte nach struktureller Qualität, Vollständigkeit der Quellenangaben, Konsistenz mit dem Wiki und Themenabdeckung. | 2026-08-29 | | [[RAG]] | architecture | Architekturmuster, bei dem LLMs die Generierung um Dokumente aus einer Wissensbasis anreichern. | 2026-08-29 | -| [[Reciprocal Rank Fusion]] | pattern | Verfahren, das Ergebnislisten mehrerer Suchmodalitäten zu einem gemeinsamen Ranking verbindet, ohne Gewichte zwischen den Modalitäten justieren zu müssen. | 2026-08-29 | | [[Scale Ceiling]] | architecture | Punkt, ab dem Wiki-Ansätze mit einem einzigen Kontext qualitativ abfallen | 2026-09-01 | -| [[Self-Healing]] | pattern | Automatisches Beheben von Mängeln, die beim Lint auffallen: verwaiste Seiten, veraltete Aussagen, kaputte Links und Formatverstöße. | 2026-08-29 | -| [[Semantic Lint Automation]] | workflow | Maschinelle Heuristiken zur Priorisierung der semantischen Prüfung: veraltete Aussagen, hohe Änderungsdichte und schwache Verlinkung | 2026-08-29 | | [[Semantic Memory]] | architecture | Langlebige Schicht für sitzungsübergreifend verdichtete Fakten aus mehreren Episoden, mit höherer Konfidenz und stärkerer Verdichtung als episodische Erinnerungen. | 2026-08-29 | -| [[Session Orientation]] | workflow | Verbindliche Vorabprüfung, die vor Query- und Update-Operationen einen Kontextbericht erzeugt (Index, jüngste Logs, Umfang) | 2026-08-29 | -| [[Shared vs Private]] | pattern | Abgrenzung persönlicher Beobachtungen (privat) von Team- und Projektwissen (geteilt), mit Regeln zum Hochstufen geprüften Wissens. | 2026-08-29 | -| [[Split Merge Reclassify]] | workflow | Eigene Befehle zum Teilen, Zusammenführen und Umklassifizieren von Seiten, mit automatischer Korrektur von Links und Frontmatter | 2026-08-29 | -| [[Split Threshold]] | workflow | Maximale Seitengröße, ab der eine Aufteilung empfohlen wird (Farza: >120-150 Zeilen, Pascalandy: 200 Zeilen) | 2026-08-29 | -| [[SSD TRIM]] | protocol | Datenträgerbefehl, mit dem SSDs ungenutzte Blöcke zurückgewinnen - für gleichbleibende Leistung und längere Lebensdauer. | 2026-08-29 | -| [[Structural Enforcement over Documented Rule]] | decision | Entscheidung, eine wiederkehrende Fehlerregel in die Ausfuehrung einzubauen statt sie aufzuschreiben - Rangfolge erzwingen vor melden vor erinnern, belegt an einer Regel, die gelesen wurde und nicht wirkte | 2026-08-31 | -| [[Stub Threshold]] | workflow | Mindestumfang, ab dem eine Wiki-Seite nicht mehr als Stub gilt (Farza: ≥3 Sätze oder 15 Zeilen) | 2026-08-29 | -| [[Supersession]] | workflow | Ablösen alten Wissens durch neue, widersprechende Information; gibt dem Wiki eine Versionierung mit ausdrücklicher Verknüpfung und Erhalt der Historie. | 2026-08-29 | | [[Three-Layer Architecture]] | architecture | Strukturmodell des LLM-Wiki-Musters mit drei Schichten: unveränderliche Rohquellen, vom LLM gepflegtes Wiki und Schemakonfiguration, die Knowledge Compounding trägt. | 2026-08-29 | | [[Token Economics]] | architecture | Kosten- und Effizienzüberlegungen zum Tokenverbrauch von LLMs - die genannten Werte 5-8x/61%/100% sind unbestätigt, keine gesicherten Fakten | 2026-09-01 | +| [[Working Memory]] | architecture | Kurzlebige Speicherschicht für jüngste Beobachtungen und vorläufige Befunde vor der Verdichtung; niedrigste Konfidenz, keine Verdichtung, wird zum Sitzungsende erneuert. | 2026-08-29 | + +## Entscheidungen + +| Page | Type | Summary | Last Modified | +|------|------|---------|----------------| +| [[Delete Rather Than Anonymize]] | decision | Private Korpusinhalte per Loeschung entfernen statt zu anonymisieren: ein Seitentitel ist der einzige Identifier eines Wikis, Umbenennen ist die volle page-lifecycle-Prozedur je Seite, Loeschen ist ein unterstuetztes Kommando. | 2026-09-01 | +| [[Denylist over Allowlist]] | decision | Entscheidung, schreibbare Felder als Schema minus kurzer Sperrliste zu bestimmen statt als gepflegte Positivliste, weil die Positivliste eine zweite Kopie des Schemas waere | 2026-08-31 | +| [[Diff-Reviewable Agent Edits]] | decision | Entscheidung, Dateiaenderungen ueber Edit/Write statt ueber Shell-Heredocs zu fahren, weil nur das erste eine pruefbare Diff hinterlaesst | 2026-08-31 | +| [[Dual Licensing by File Plan]] | decision | Ein Repo mit Code- und Inhaltsanteil erhaelt zwei Lizenzen; die Grenze zwischen ihnen ist kein zweiter, gepflegter Pfadkatalog, sondern der ohnehin vorhandene Dateiplan des Distributionswerkzeugs. | 2026-09-01 | +| [[Issue Label Scheme]] | decision | Pflicht-Labelschema fuer das Gitea-Board: vier Achsen (area/kind/prio/size) plus seit 2026-09-04 drei optionale status/-Flags, darunter status/incoming fuer unausgearbeitete Stubs, die die Vier-Achsen-Pflicht aussetzen statt sie zu ergaenzen; die Regel liegt in instructions/dev/, weil sie keine ausgelieferte Instanz erreichen darf | 2026-09-04 | +| [[KB Stack Versioning]] | decision | Semantische Versionierung des Wiki-Stacks: VERSION beschreibt die Maschinerie, Kompatibilitaet (Drop-in-Ersatz) und Inhaltsmigration sind seit 2.5.0 getrennte, unabhaengig geprueft Fragen | 2026-09-02 | +| [[Structural Enforcement over Documented Rule]] | decision | Entscheidung, eine wiederkehrende Fehlerregel in die Ausfuehrung einzubauen statt sie aufzuschreiben - Rangfolge erzwingen vor melden vor erinnern, belegt an einer Regel, die gelesen wurde und nicht wirkte | 2026-08-31 | + +## Muster + +| Page | Type | Summary | Last Modified | +|------|------|---------|----------------| +| [[Audit Trail]] | pattern | Unveränderliches chronologisches Log aller Wiki-Operationen (Ingest, Bearbeitung, Löschung, Abfrage) mit Zeitstempel, Akteur, Ziel und Änderungsbeschreibung. | 2026-08-29 | +| [[BM25]] | pattern | Schlüsselwortbasiertes Retrieval-Verfahren, das über Termfrequenz, inverse Dokumentfrequenz und Stemming exakte oder teilweise Übereinstimmungen findet. | 2026-08-29 | +| [[Command Round-Trip Integrity]] | pattern | Anforderung, dass zwei Befehle auf derselben Datei in jeder Reihenfolge zusammenpassen und jeder erzeugte Zustand einen Gegenbefehl hat - 2026-08-31 in wikitool zweimal verletzt | 2026-08-31 | +| [[Confidence Scoring]] | pattern | Mechanismus, der faktischen Aussagen quantitative Werte nach Quellenzahl, Aktualität, Qualität und Bestätigung zuweist, um gut gestütztes Wissen zu erkennen. | 2026-08-29 | +| [[Contradiction Resolution]] | pattern | Automatisches Erkennen und Auflösen widersprüchlicher Aussagen anhand von Konfidenz, Aktualität und Autorität der Quelle. | 2026-08-29 | +| [[Entity Extraction]] | pattern | Erkennen und Strukturieren von Entities (Personen, Projekte, Bibliotheken, Concepts, Dateien, Entscheidungen, Systeme, Werkzeuge) samt typspezifischer Attribute aus Rohquellen. | 2026-08-29 | +| [[Filter on Ingest]] | pattern | Automatisches Erkennen und Entfernen sensibler Daten (API-Schlüssel, Token, Credentials, personenbezogene Daten) vor der Aufnahme ins Wiki, per Regex und ML-Erkennung. | 2026-08-29 | +| [[Forgetting]] | pattern | Muster zur Wissensbindung, das selten abgerufene Fakten schrittweise zurückstuft, modelliert nach der Ebbinghausschen Vergessenskurve. | 2026-08-29 | +| [[Graph Traversal]] | pattern | Verfahren, verbundene Entities im Wissensgraphen über typisierte Beziehungen (uses, depends-on, contradicts, caused) zu finden und strukturelle Fragen zu beantworten. | 2026-08-29 | +| [[Mesh Sync]] | pattern | Abgleichsmechanismus, der Beobachtungen paralleler Agenten in ein gemeinsames Wiki überführt; Last-Write-Wins mit Konflikterkennung und manuellem Eingriff. | 2026-08-29 | +| [[Quality Scoring]] | pattern | Quantitative Bewertung aller vom LLM geschriebenen Inhalte nach struktureller Qualität, Vollständigkeit der Quellenangaben, Konsistenz mit dem Wiki und Themenabdeckung. | 2026-08-29 | +| [[Reciprocal Rank Fusion]] | pattern | Verfahren, das Ergebnislisten mehrerer Suchmodalitäten zu einem gemeinsamen Ranking verbindet, ohne Gewichte zwischen den Modalitäten justieren zu müssen. | 2026-08-29 | +| [[Self-Healing]] | pattern | Automatisches Beheben von Mängeln, die beim Lint auffallen: verwaiste Seiten, veraltete Aussagen, kaputte Links und Formatverstöße. | 2026-08-29 | +| [[Shared vs Private]] | pattern | Abgrenzung persönlicher Beobachtungen (privat) von Team- und Projektwissen (geteilt), mit Regeln zum Hochstufen geprüften Wissens. | 2026-08-29 | | [[Typed Relationships]] | pattern | Verwendung semantisch aussagekräftiger Beziehungstypen (uses, depends-on, contradicts, caused, fixed, supersedes, replaces) statt undifferenzierter Wikilinks. | 2026-08-29 | -| [[User Management]] | workflow | Linux-Ablauf zum Anlegen, Ändern, Überwachen und Löschen von Benutzerkonten mit useradd, usermod und userdel, samt Gruppenverwaltung und sudoers-Konfiguration. | 2026-08-29 | | [[Vector Search]] | pattern | Semantische Ähnlichkeitssuche über Embedding-Vektoren, die inhaltlich verwandte Seiten auch ohne exakte Schlüsselwortübereinstimmung findet. | 2026-08-29 | | [[Work Coordination]] | pattern | Leichtgewichtige Erfassung von Aufgabenstatus (in Arbeit, blockiert, erledigt, prüfbedürftig) und Zuweisung, um Doppelarbeit bei mehreren Agenten zu vermeiden. | 2026-08-29 | -| [[Workflow Extraction]] | workflow | Herauslösen von Workflow-Abschnitten aus monolithischer Dokumentation | 2026-09-01 | -| [[Workflow Orchestration]] | workflow | Orchestrierte Einzelbefehle für vollständige Operationen (ingest run, lint run, update run) mit Dry-Run-Vorschau vor dem Schreiben | 2026-08-29 | -| [[Working Memory]] | architecture | Kurzlebige Speicherschicht für jüngste Beobachtungen und vorläufige Befunde vor der Verdichtung; niedrigste Konfidenz, keine Verdichtung, wird zum Sitzungsende erneuert. | 2026-08-29 | + +## Problemstellungen + +| Page | Type | Summary | Last Modified | +|------|------|---------|----------------| +| [[Ambient Environment Dependency]] | problem | Fehlerklasse, in der ein Test gruen ist, weil die Maschine zufaellig passt statt weil der Code stimmt - abgegrenzt gegen den Green Suite Blind Spot, belegt an vier Faellen unter Gitea-Issue #8 | 2026-08-31 | +| [[Detect-Repair Asymmetry]] | problem | Werkzeugluecke, in der ein Check einen Defekt zuverlaessig meldet, aber kein Befehl ihn behebt - womit die Handeditierung der einzige verbleibende Ausweg ist | 2026-08-31 | +| [[Green Suite Blind Spot]] | problem | Defekt, der eine vollstaendig gruene Testsuite ueberlebt, weil nie ein Test das richtige Verhalten behauptet hat - belegt an drei prio/1-2-Defekten (Round-Trip, Zitat-Notation-als-Code, Zitat-Limit) | 2026-08-31 | +| [[Naming Convention Conflict]] | problem | Widerspruch zwischen README.md (kebab-case) und AGENTS.md (lesbar mit Leerzeichen), der zu Drift bei der Validierung führt | 2026-08-29 | | [[Write-Once Frontmatter Fields]] | problem | Defektklasse, in der ein Feld nur beim Anlegen der Seite schreibbar ist und danach unerreichbar bleibt, weil kein Mutationsbefehl es kennt und new nicht idempotent ist | 2026-08-31 | +## Protokolle + +| Page | Type | Summary | Last Modified | +|------|------|---------|----------------| +| [[CPPC]] | protocol | Hardwareschnittstelle Collaborative Processor Performance Control für feingranulares CPU-Power-Management zwischen Betriebssystem und AMD-Prozessor. | 2026-08-29 | +| [[Modbus]] | protocol | Industrielles Kommunikationsprotokoll von 1979 zur Anbindung speicherprogrammierbarer Steuerungen und Geräte über serielle oder TCP-Netze. | 2026-08-29 | +| [[SSD TRIM]] | protocol | Datenträgerbefehl, mit dem SSDs ungenutzte Blöcke zurückgewinnen - für gleichbleibende Leistung und längere Lebensdauer. | 2026-08-29 | + diff --git a/kb/concepts/Consolidation Tiers.md b/kb/concepts/architectures/Consolidation Tiers.md similarity index 100% rename from kb/concepts/Consolidation Tiers.md rename to kb/concepts/architectures/Consolidation Tiers.md diff --git a/kb/concepts/Context Isolation.md b/kb/concepts/architectures/Context Isolation.md similarity index 100% rename from kb/concepts/Context Isolation.md rename to kb/concepts/architectures/Context Isolation.md diff --git a/kb/concepts/Cross-platform Agent Skills.md b/kb/concepts/architectures/Cross-platform Agent Skills.md similarity index 100% rename from kb/concepts/Cross-platform Agent Skills.md rename to kb/concepts/architectures/Cross-platform Agent Skills.md diff --git a/kb/concepts/Episodic Memory.md b/kb/concepts/architectures/Episodic Memory.md similarity index 100% rename from kb/concepts/Episodic Memory.md rename to kb/concepts/architectures/Episodic Memory.md diff --git a/kb/concepts/Hybrid Search.md b/kb/concepts/architectures/Hybrid Search.md similarity index 100% rename from kb/concepts/Hybrid Search.md rename to kb/concepts/architectures/Hybrid Search.md diff --git a/kb/concepts/Implementation Spectrum.md b/kb/concepts/architectures/Implementation Spectrum.md similarity index 100% rename from kb/concepts/Implementation Spectrum.md rename to kb/concepts/architectures/Implementation Spectrum.md diff --git a/kb/concepts/Knowledge Graph.md b/kb/concepts/architectures/Knowledge Graph.md similarity index 100% rename from kb/concepts/Knowledge Graph.md rename to kb/concepts/architectures/Knowledge Graph.md diff --git a/kb/concepts/LLM Wiki Pattern.md b/kb/concepts/architectures/LLM Wiki Pattern.md similarity index 100% rename from kb/concepts/LLM Wiki Pattern.md rename to kb/concepts/architectures/LLM Wiki Pattern.md diff --git a/kb/concepts/MCP-Leseserver.md b/kb/concepts/architectures/MCP-Leseserver.md similarity index 100% rename from kb/concepts/MCP-Leseserver.md rename to kb/concepts/architectures/MCP-Leseserver.md diff --git a/kb/concepts/Memory Lifecycle.md b/kb/concepts/architectures/Memory Lifecycle.md similarity index 100% rename from kb/concepts/Memory Lifecycle.md rename to kb/concepts/architectures/Memory Lifecycle.md diff --git a/kb/concepts/OKF Compatibility.md b/kb/concepts/architectures/OKF Compatibility.md similarity index 100% rename from kb/concepts/OKF Compatibility.md rename to kb/concepts/architectures/OKF Compatibility.md diff --git a/kb/concepts/Optional Instance Context File.md b/kb/concepts/architectures/Optional Instance Context File.md similarity index 100% rename from kb/concepts/Optional Instance Context File.md rename to kb/concepts/architectures/Optional Instance Context File.md diff --git a/kb/concepts/Personalization Plane.md b/kb/concepts/architectures/Personalization Plane.md similarity index 100% rename from kb/concepts/Personalization Plane.md rename to kb/concepts/architectures/Personalization Plane.md diff --git a/kb/concepts/Procedural Memory.md b/kb/concepts/architectures/Procedural Memory.md similarity index 100% rename from kb/concepts/Procedural Memory.md rename to kb/concepts/architectures/Procedural Memory.md diff --git a/kb/concepts/RAG.md b/kb/concepts/architectures/RAG.md similarity index 100% rename from kb/concepts/RAG.md rename to kb/concepts/architectures/RAG.md diff --git a/kb/concepts/Scale Ceiling.md b/kb/concepts/architectures/Scale Ceiling.md similarity index 100% rename from kb/concepts/Scale Ceiling.md rename to kb/concepts/architectures/Scale Ceiling.md diff --git a/kb/concepts/Semantic Memory.md b/kb/concepts/architectures/Semantic Memory.md similarity index 100% rename from kb/concepts/Semantic Memory.md rename to kb/concepts/architectures/Semantic Memory.md diff --git a/kb/concepts/Three-Layer Architecture.md b/kb/concepts/architectures/Three-Layer Architecture.md similarity index 100% rename from kb/concepts/Three-Layer Architecture.md rename to kb/concepts/architectures/Three-Layer Architecture.md diff --git a/kb/concepts/Token Economics.md b/kb/concepts/architectures/Token Economics.md similarity index 100% rename from kb/concepts/Token Economics.md rename to kb/concepts/architectures/Token Economics.md diff --git a/kb/concepts/Working Memory.md b/kb/concepts/architectures/Working Memory.md similarity index 100% rename from kb/concepts/Working Memory.md rename to kb/concepts/architectures/Working Memory.md diff --git a/kb/concepts/Delete Rather Than Anonymize.md b/kb/concepts/decisions/Delete Rather Than Anonymize.md similarity index 100% rename from kb/concepts/Delete Rather Than Anonymize.md rename to kb/concepts/decisions/Delete Rather Than Anonymize.md diff --git a/kb/concepts/Denylist over Allowlist.md b/kb/concepts/decisions/Denylist over Allowlist.md similarity index 100% rename from kb/concepts/Denylist over Allowlist.md rename to kb/concepts/decisions/Denylist over Allowlist.md diff --git a/kb/concepts/Diff-Reviewable Agent Edits.md b/kb/concepts/decisions/Diff-Reviewable Agent Edits.md similarity index 100% rename from kb/concepts/Diff-Reviewable Agent Edits.md rename to kb/concepts/decisions/Diff-Reviewable Agent Edits.md diff --git a/kb/concepts/Dual Licensing by File Plan.md b/kb/concepts/decisions/Dual Licensing by File Plan.md similarity index 100% rename from kb/concepts/Dual Licensing by File Plan.md rename to kb/concepts/decisions/Dual Licensing by File Plan.md diff --git a/kb/concepts/Issue Label Scheme.md b/kb/concepts/decisions/Issue Label Scheme.md similarity index 100% rename from kb/concepts/Issue Label Scheme.md rename to kb/concepts/decisions/Issue Label Scheme.md diff --git a/kb/concepts/KB Stack Versioning.md b/kb/concepts/decisions/KB Stack Versioning.md similarity index 100% rename from kb/concepts/KB Stack Versioning.md rename to kb/concepts/decisions/KB Stack Versioning.md diff --git a/kb/concepts/Structural Enforcement over Documented Rule.md b/kb/concepts/decisions/Structural Enforcement over Documented Rule.md similarity index 100% rename from kb/concepts/Structural Enforcement over Documented Rule.md rename to kb/concepts/decisions/Structural Enforcement over Documented Rule.md diff --git a/kb/concepts/Audit Trail.md b/kb/concepts/patterns/Audit Trail.md similarity index 100% rename from kb/concepts/Audit Trail.md rename to kb/concepts/patterns/Audit Trail.md diff --git a/kb/concepts/BM25.md b/kb/concepts/patterns/BM25.md similarity index 100% rename from kb/concepts/BM25.md rename to kb/concepts/patterns/BM25.md diff --git a/kb/concepts/Command Round-Trip Integrity.md b/kb/concepts/patterns/Command Round-Trip Integrity.md similarity index 100% rename from kb/concepts/Command Round-Trip Integrity.md rename to kb/concepts/patterns/Command Round-Trip Integrity.md diff --git a/kb/concepts/Confidence Scoring.md b/kb/concepts/patterns/Confidence Scoring.md similarity index 100% rename from kb/concepts/Confidence Scoring.md rename to kb/concepts/patterns/Confidence Scoring.md diff --git a/kb/concepts/Contradiction Resolution.md b/kb/concepts/patterns/Contradiction Resolution.md similarity index 100% rename from kb/concepts/Contradiction Resolution.md rename to kb/concepts/patterns/Contradiction Resolution.md diff --git a/kb/concepts/Entity Extraction.md b/kb/concepts/patterns/Entity Extraction.md similarity index 100% rename from kb/concepts/Entity Extraction.md rename to kb/concepts/patterns/Entity Extraction.md diff --git a/kb/concepts/Filter on Ingest.md b/kb/concepts/patterns/Filter on Ingest.md similarity index 100% rename from kb/concepts/Filter on Ingest.md rename to kb/concepts/patterns/Filter on Ingest.md diff --git a/kb/concepts/Forgetting.md b/kb/concepts/patterns/Forgetting.md similarity index 100% rename from kb/concepts/Forgetting.md rename to kb/concepts/patterns/Forgetting.md diff --git a/kb/concepts/Graph Traversal.md b/kb/concepts/patterns/Graph Traversal.md similarity index 100% rename from kb/concepts/Graph Traversal.md rename to kb/concepts/patterns/Graph Traversal.md diff --git a/kb/concepts/Mesh Sync.md b/kb/concepts/patterns/Mesh Sync.md similarity index 100% rename from kb/concepts/Mesh Sync.md rename to kb/concepts/patterns/Mesh Sync.md diff --git a/kb/concepts/Quality Scoring.md b/kb/concepts/patterns/Quality Scoring.md similarity index 100% rename from kb/concepts/Quality Scoring.md rename to kb/concepts/patterns/Quality Scoring.md diff --git a/kb/concepts/Reciprocal Rank Fusion.md b/kb/concepts/patterns/Reciprocal Rank Fusion.md similarity index 100% rename from kb/concepts/Reciprocal Rank Fusion.md rename to kb/concepts/patterns/Reciprocal Rank Fusion.md diff --git a/kb/concepts/Self-Healing.md b/kb/concepts/patterns/Self-Healing.md similarity index 100% rename from kb/concepts/Self-Healing.md rename to kb/concepts/patterns/Self-Healing.md diff --git a/kb/concepts/Shared vs Private.md b/kb/concepts/patterns/Shared vs Private.md similarity index 100% rename from kb/concepts/Shared vs Private.md rename to kb/concepts/patterns/Shared vs Private.md diff --git a/kb/concepts/Typed Relationships.md b/kb/concepts/patterns/Typed Relationships.md similarity index 100% rename from kb/concepts/Typed Relationships.md rename to kb/concepts/patterns/Typed Relationships.md diff --git a/kb/concepts/Vector Search.md b/kb/concepts/patterns/Vector Search.md similarity index 100% rename from kb/concepts/Vector Search.md rename to kb/concepts/patterns/Vector Search.md diff --git a/kb/concepts/Work Coordination.md b/kb/concepts/patterns/Work Coordination.md similarity index 100% rename from kb/concepts/Work Coordination.md rename to kb/concepts/patterns/Work Coordination.md diff --git a/kb/concepts/Ambient Environment Dependency.md b/kb/concepts/problems/Ambient Environment Dependency.md similarity index 100% rename from kb/concepts/Ambient Environment Dependency.md rename to kb/concepts/problems/Ambient Environment Dependency.md diff --git a/kb/concepts/Detect-Repair Asymmetry.md b/kb/concepts/problems/Detect-Repair Asymmetry.md similarity index 100% rename from kb/concepts/Detect-Repair Asymmetry.md rename to kb/concepts/problems/Detect-Repair Asymmetry.md diff --git a/kb/concepts/Green Suite Blind Spot.md b/kb/concepts/problems/Green Suite Blind Spot.md similarity index 100% rename from kb/concepts/Green Suite Blind Spot.md rename to kb/concepts/problems/Green Suite Blind Spot.md diff --git a/kb/concepts/Naming Convention Conflict.md b/kb/concepts/problems/Naming Convention Conflict.md similarity index 100% rename from kb/concepts/Naming Convention Conflict.md rename to kb/concepts/problems/Naming Convention Conflict.md diff --git a/kb/concepts/Write-Once Frontmatter Fields.md b/kb/concepts/problems/Write-Once Frontmatter Fields.md similarity index 100% rename from kb/concepts/Write-Once Frontmatter Fields.md rename to kb/concepts/problems/Write-Once Frontmatter Fields.md diff --git a/kb/concepts/CPPC.md b/kb/concepts/protocols/CPPC.md similarity index 100% rename from kb/concepts/CPPC.md rename to kb/concepts/protocols/CPPC.md diff --git a/kb/concepts/Modbus.md b/kb/concepts/protocols/Modbus.md similarity index 100% rename from kb/concepts/Modbus.md rename to kb/concepts/protocols/Modbus.md diff --git a/kb/concepts/SSD TRIM.md b/kb/concepts/protocols/SSD TRIM.md similarity index 100% rename from kb/concepts/SSD TRIM.md rename to kb/concepts/protocols/SSD TRIM.md diff --git a/kb/concepts/Anti-Cramming Heuristic.md b/kb/concepts/workflows/Anti-Cramming Heuristic.md similarity index 100% rename from kb/concepts/Anti-Cramming Heuristic.md rename to kb/concepts/workflows/Anti-Cramming Heuristic.md diff --git a/kb/concepts/Bulk Operations.md b/kb/concepts/workflows/Bulk Operations.md similarity index 100% rename from kb/concepts/Bulk Operations.md rename to kb/concepts/workflows/Bulk Operations.md diff --git a/kb/concepts/CI Integration.md b/kb/concepts/workflows/CI Integration.md similarity index 100% rename from kb/concepts/CI Integration.md rename to kb/concepts/workflows/CI Integration.md diff --git a/kb/concepts/Checkpoint Audit.md b/kb/concepts/workflows/Checkpoint Audit.md similarity index 100% rename from kb/concepts/Checkpoint Audit.md rename to kb/concepts/workflows/Checkpoint Audit.md diff --git a/kb/concepts/Claude Code Auto Mode.md b/kb/concepts/workflows/Claude Code Auto Mode.md similarity index 100% rename from kb/concepts/Claude Code Auto Mode.md rename to kb/concepts/workflows/Claude Code Auto Mode.md diff --git a/kb/concepts/Content Quality Control.md b/kb/concepts/workflows/Content Quality Control.md similarity index 100% rename from kb/concepts/Content Quality Control.md rename to kb/concepts/workflows/Content Quality Control.md diff --git a/kb/concepts/Crystallization.md b/kb/concepts/workflows/Crystallization.md similarity index 100% rename from kb/concepts/Crystallization.md rename to kb/concepts/workflows/Crystallization.md diff --git a/kb/concepts/Event-Driven Automation.md b/kb/concepts/workflows/Event-Driven Automation.md similarity index 100% rename from kb/concepts/Event-Driven Automation.md rename to kb/concepts/workflows/Event-Driven Automation.md diff --git a/kb/concepts/Hooks.md b/kb/concepts/workflows/Hooks.md similarity index 100% rename from kb/concepts/Hooks.md rename to kb/concepts/workflows/Hooks.md diff --git a/kb/concepts/Index Scaling.md b/kb/concepts/workflows/Index Scaling.md similarity index 100% rename from kb/concepts/Index Scaling.md rename to kb/concepts/workflows/Index Scaling.md diff --git a/kb/concepts/Iteration and Cost Limits.md b/kb/concepts/workflows/Iteration and Cost Limits.md similarity index 100% rename from kb/concepts/Iteration and Cost Limits.md rename to kb/concepts/workflows/Iteration and Cost Limits.md diff --git a/kb/concepts/KB Migration.md b/kb/concepts/workflows/KB Migration.md similarity index 100% rename from kb/concepts/KB Migration.md rename to kb/concepts/workflows/KB Migration.md diff --git a/kb/concepts/Knowledge Compounding.md b/kb/concepts/workflows/Knowledge Compounding.md similarity index 100% rename from kb/concepts/Knowledge Compounding.md rename to kb/concepts/workflows/Knowledge Compounding.md diff --git a/kb/concepts/Lint Workflow.md b/kb/concepts/workflows/Lint Workflow.md similarity index 100% rename from kb/concepts/Lint Workflow.md rename to kb/concepts/workflows/Lint Workflow.md diff --git a/kb/concepts/Mass-Update Gate.md b/kb/concepts/workflows/Mass-Update Gate.md similarity index 100% rename from kb/concepts/Mass-Update Gate.md rename to kb/concepts/workflows/Mass-Update Gate.md diff --git a/kb/concepts/Multi-Agent Collaboration.md b/kb/concepts/workflows/Multi-Agent Collaboration.md similarity index 100% rename from kb/concepts/Multi-Agent Collaboration.md rename to kb/concepts/workflows/Multi-Agent Collaboration.md diff --git a/kb/concepts/Privacy and Governance.md b/kb/concepts/workflows/Privacy and Governance.md similarity index 100% rename from kb/concepts/Privacy and Governance.md rename to kb/concepts/workflows/Privacy and Governance.md diff --git a/kb/concepts/Publish-Remote Gate.md b/kb/concepts/workflows/Publish-Remote Gate.md similarity index 100% rename from kb/concepts/Publish-Remote Gate.md rename to kb/concepts/workflows/Publish-Remote Gate.md diff --git a/kb/concepts/Quality and Self-Correction.md b/kb/concepts/workflows/Quality and Self-Correction.md similarity index 100% rename from kb/concepts/Quality and Self-Correction.md rename to kb/concepts/workflows/Quality and Self-Correction.md diff --git a/kb/concepts/Semantic Lint Automation.md b/kb/concepts/workflows/Semantic Lint Automation.md similarity index 100% rename from kb/concepts/Semantic Lint Automation.md rename to kb/concepts/workflows/Semantic Lint Automation.md diff --git a/kb/concepts/Session Orientation.md b/kb/concepts/workflows/Session Orientation.md similarity index 100% rename from kb/concepts/Session Orientation.md rename to kb/concepts/workflows/Session Orientation.md diff --git a/kb/concepts/Split Merge Reclassify.md b/kb/concepts/workflows/Split Merge Reclassify.md similarity index 100% rename from kb/concepts/Split Merge Reclassify.md rename to kb/concepts/workflows/Split Merge Reclassify.md diff --git a/kb/concepts/Split Threshold.md b/kb/concepts/workflows/Split Threshold.md similarity index 100% rename from kb/concepts/Split Threshold.md rename to kb/concepts/workflows/Split Threshold.md diff --git a/kb/concepts/Stub Threshold.md b/kb/concepts/workflows/Stub Threshold.md similarity index 100% rename from kb/concepts/Stub Threshold.md rename to kb/concepts/workflows/Stub Threshold.md diff --git a/kb/concepts/Supersession.md b/kb/concepts/workflows/Supersession.md similarity index 100% rename from kb/concepts/Supersession.md rename to kb/concepts/workflows/Supersession.md diff --git a/kb/concepts/User Management.md b/kb/concepts/workflows/User Management.md similarity index 100% rename from kb/concepts/User Management.md rename to kb/concepts/workflows/User Management.md diff --git a/kb/concepts/Workflow Extraction.md b/kb/concepts/workflows/Workflow Extraction.md similarity index 100% rename from kb/concepts/Workflow Extraction.md rename to kb/concepts/workflows/Workflow Extraction.md diff --git a/kb/concepts/Workflow Orchestration.md b/kb/concepts/workflows/Workflow Orchestration.md similarity index 100% rename from kb/concepts/Workflow Orchestration.md rename to kb/concepts/workflows/Workflow Orchestration.md diff --git a/kb/index.md b/kb/index.md index 0486354..4b97475 100644 --- a/kb/index.md +++ b/kb/index.md @@ -18,7 +18,7 @@ The page tables live in a generated `INDEX.md` inside each collection, linked be - **Concepts:** 80 - **Entities:** 72 - **Sources:** 29 -- **Last Updated:** 2026-09-04 +- **Last Updated:** 2026-09-08 --- @@ -31,6 +31,17 @@ The page tables live in a generated `INDEX.md` inside each collection, linked be | `entities/` | 72 | [entities/INDEX.md](entities/INDEX.md) | | `sources/` | 29 | [sources/INDEX.md](sources/INDEX.md) | +### concepts/ + +| Area | Pages | Index | +|------|------:|-------| +| Abläufe | 28 | [concepts/INDEX.md#abläufe](concepts/INDEX.md#abläufe) | +| Architekturen | 20 | [concepts/INDEX.md#architekturen](concepts/INDEX.md#architekturen) | +| Entscheidungen | 7 | [concepts/INDEX.md#entscheidungen](concepts/INDEX.md#entscheidungen) | +| Muster | 17 | [concepts/INDEX.md#muster](concepts/INDEX.md#muster) | +| Problemstellungen | 5 | [concepts/INDEX.md#problemstellungen](concepts/INDEX.md#problemstellungen) | +| Protokolle | 3 | [concepts/INDEX.md#protokolle](concepts/INDEX.md#protokolle) | + ### entities/ | Area | Pages | Index | diff --git a/kb/log.md b/kb/log.md index 1e5c21d..29d54e9 100644 --- a/kb/log.md +++ b/kb/log.md @@ -155,3 +155,9 @@ Drittes status/-Flag status/incoming ergaenzt (Gitea #63): status/-Tabelle auf d wikitool move --reconcile hat die drei nach kb/entities/projects/{kfchou,vanillaflava,yugasun}/*.md verschachtelten Seiten (wiki-skills, wiki-skills-vanillaflava, llm-wiki-skills) nach kb/entities/projects/ hochgezogen und die drei geleerten Owner-Verzeichnisse entfernt. index rebuild und sources rebuild-index liefen danach; lint meldet weder misplaced_pages noch nested_pages noch duplicate_titles; migrate verify --from HEAD bestaetigt compared=182, added=0, removed=0, alle drei als moved. --- + +## [2026-09-08] move | kb/concepts/ bekommt Areas: 80 Seiten in ihre concept_type-Verzeichnisse (#59) + +types/concept.md deklariert ein layout: fuer alle sechs concept_type-Werte; wikitool move --reconcile hat daraufhin alle 80 Seiten aus kb/concepts/ in ihre Area gezogen (architectures 20, patterns 17, protocols 3, workflows 28, decisions 7, problems 5). Kein Titel, kein Body, kein Frontmatter-Feld angefasst: migrate verify --from HEAD bestaetigt compared=182, added=0, removed=0, 80 moved, 0 findings. index rebuild erzeugt sechs Areas, keine ueber der Shard-Schwelle von 50, also kein eigener Shard - die Schwelle wirkt wieder als Schwelle statt als toter Wert. lint meldet danach weder misplaced_pages noch den neuen unsharded_collections-Befund. + +--- diff --git a/tools/CONTRACT.md b/tools/CONTRACT.md index 64f5f5b..aa04b3f 100644 --- a/tools/CONTRACT.md +++ b/tools/CONTRACT.md @@ -51,7 +51,7 @@ tools/wikitool --help | `index rebuild [--dry-run]` | Regenerate the catalog from every page's frontmatter: `kb/index.md` becomes a map (statistics, one row per collection and per area, links to the shards) and the page tables are written to a generated `INDEX.md` in each collection. An area past 50 rows gets its own shard. Stale shards from removed collections/areas are deleted in the same pass | | `log append --op ingest\|query\|lint\|create\|update\|delete\|rename\|move --title "..." [--body "..."\|--body-file path]` | Append a formatted entry to `kb/log.md` | | `log status` | Read-only: count `ingest` entries logged since the last `lint` entry - the deterministic trigger behind the Maintenance Schedule's "every 10 sources" full-lint cadence | -| `lint [--json] [--markdown out.md] [--full] [--fail-on-error]` | Structural + provenance checks: broken wikilinks, dangling frontmatter references, orphan pages, index drift, schema gaps, duplicate titles, title mismatches, pages nested more than one directory below their collection (hard - the generated catalog folds these into their area silently rather than merely reading it, see #57), uncovered raw files, broken `raw_files:` refs, raw files claimed by more than one source page, unmarked provenance, citation/frontmatter drift, unbalanced generated-region markers, edges whose label is missing or not authorised by the source collection's `outbound:` (both hard once `kb_version` has reached the release that introduced labelled edges - advisory below it, so a corpus mid-migration is not refused by the check measuring it), `see-also` edges whose reverse direction already carries a specific label (advisory only - redundant rather than wrong, and never migration-gated, since no version turns the redundancy into an error), quote-limit overages (>2 blockquoted lines/page, advisory only). Prints only the sections that found something and always writes the full report to `reports/Lint Report .md` (or `--markdown`), naming the path - `--full` prints everything, `--json` prints the findings and writes nothing | +| `lint [--json] [--markdown out.md] [--full] [--fail-on-error]` | Structural + provenance checks: broken wikilinks, dangling frontmatter references, orphan pages, index drift, schema gaps, duplicate titles, title mismatches, pages nested more than one directory below their collection (hard - the generated catalog folds these into their area silently rather than merely reading it, see #57), uncovered raw files, broken `raw_files:` refs, raw files claimed by more than one source page, unmarked provenance, citation/frontmatter drift, unbalanced generated-region markers, edges whose label is missing or not authorised by the source collection's `outbound:` (both hard once `kb_version` has reached the release that introduced labelled edges - advisory below it, so a corpus mid-migration is not refused by the check measuring it), `see-also` edges whose reverse direction already carries a specific label (advisory only - redundant rather than wrong, and never migration-gated, since no version turns the redundancy into an error), a collection past the catalog's per-area shard threshold that has no areas to shard (advisory only - sharding is automatic but per *area*, so a collection nobody gave areas keeps one table however large it grows, #59; reported with the split its subtype field would produce, and only when that split puts every resulting area at or under the threshold, so a lopsided or small collection stays silent), quote-limit overages (>2 blockquoted lines/page, advisory only). Prints only the sections that found something and always writes the full report to `reports/Lint Report .md` (or `--markdown`), naming the path - `--full` prints everything, `--json` prints the findings and writes nothing | | `search [""] [--field ...] [--kind/--subtype/--collection/--tag ] [--regex] [--limit N] [--sort [-]] [--backend ] [--matches] [--json]` | Find pages in `kb/` without reading the index. Text search runs through a pluggable backend (`rg` today); `--field` predicates are evaluated on frontmatter - `f=v`, `f~substring`, `'f>=v'`, `'f:*'` (present), `'!f'` (absent), repeatable and ANDed. With no text this is a pure structured query. Results carry kind/summary/confidence so a hit can be judged without opening the page. A page whose frontmatter does not parse can match no positive predicate, so it is **named** rather than dropped: `--json` always carries an `unreadable` list of `{path, reason}` (usually empty), and the table form writes the same lines to stderr. `--regex` is applied by `rg` alone, whose engine is linear; the ranking boosts for title and summary are literal-containment only, so a non-literal pattern is ranked by match count. `rg` is killed after 30 s and reported as a failure. Read-only, and **exempt from the Iteration Budget Gate** | | `confidence decay [--apply]` | Recompute every page's derived `confidence` as `confidence_base * (1 - 0.01/month)`, floored at 0.2; dry-run by default | | `confidence init-base [--apply]` | One-time backfill: set `confidence_base` from the current `confidence` on pages that predate the derived-confidence model | diff --git a/tools/README.md b/tools/README.md index a6a1fe8..878deb0 100644 --- a/tools/README.md +++ b/tools/README.md @@ -45,6 +45,7 @@ tools/ conventions.py kb/CONVENTIONS.md: what this instance decided about authoring, as opposed to what the stack enforces ownership.py the stack-vs-instance boundary under a content stage - one predicate, read by `dist_cmd.py` and `commands/upstream_cmd.py` so the two cannot answer it differently type_resolver.py type-spec loading and schema resolution + catalog.py how the corpus groups into collections and areas, and the shard threshold - with no CLI attached lint_core.py the lint checks and the report, with no CLI attached types_core.py type-spec listing/description, with no CLI attached markdown_code.py masks code spans/fences so a page may show wiki notation, not only use it @@ -57,8 +58,8 @@ tools/ ``` **Two consumers, one core.** The CLI is not the only caller any more. The cores -(`search/service.py`, `lint_core.py`, `types_core.py`) hold what decides an -answer and import no `typer` and no `rich`; the modules under `commands/` turn +(`search/service.py`, `lint_core.py`, `types_core.py`, `catalog.py`) hold what +decides an answer and import no `typer` and no `rich`; the modules under `commands/` turn those values into terminal output and those exceptions into exit codes. `api.Corpus` is the in-process entry point over the same functions - it takes a corpus root, returns exactly the structures the `--json` forms print, and diff --git a/tools/chemenu/catalog.py b/tools/chemenu/catalog.py new file mode 100644 index 0000000..26325b7 --- /dev/null +++ b/tools/chemenu/catalog.py @@ -0,0 +1,140 @@ +"""How the corpus groups into collections and areas, with no CLI attached. + +Split out of `commands/index_build.py` for the reason `lint_core.py` gives at +the top of itself: this is a pure function over a corpus directory, and it was +sitting in a module that imports `typer` and `rich`. `lint` needs the same +grouping - it is what answers "does this collection have areas, and is it over +the threshold?" (Gitea #59) - and `chemenu.api`, the read surface, may not +reach a command module at all. Importing it from there would have pulled the +whole CLI head in behind it. + +So the split runs along the same line as lint's: everything that decides *how +the corpus is shaped* lives here; everything that decides *what the catalog +looks like* - the tables, the map, the shard files - stays in +`commands/index_build.py`, which imports from here. +""" +from __future__ import annotations + +from dataclasses import dataclass, field +from pathlib import Path + +from chemenu.kb_collections import iter_kb_collections +from chemenu.page import Page +from chemenu.type_resolver import resolver + +# Rows per area before it is split into its own shard. From the wiki's own +# `Index Scaling` page ("split table sections at >50 entries"), kept as a plain +# number so growth is handled by arithmetic rather than by a judgment call. +SHARD_THRESHOLD = 50 + +# Display title for pages sitting directly in a collection root rather than in +# an area subdirectory. +UNGROUPED_TITLE = "All" + + +@dataclass +class Area: + """One grouping inside a collection: a subdirectory, or the collection root + for pages that sit directly in it.""" + + name: str + title: str + pages: list[Page] = field(default_factory=list) + own_shard: bool = False + + @property + def count(self) -> int: + return len(self.pages) + + +@dataclass +class Collection: + name: str + areas: list[Area] = field(default_factory=list) + + @property + def count(self) -> int: + return sum(area.count for area in self.areas) + + +def area_titles() -> dict[str, dict[str, str]]: + """Display titles per collection: `{collection: {area_dir: title}}`, taken + from each type-spec's own `layout:` rather than a hardcoded map - so a new + subtype names its own section by adding a type-spec, with no code change. + + Every type-spec is read, not just `entity`'s. That hardcoding was the + asymmetry behind Gitea #59: the axis a collection splits along is declared + in `layout:`, and a second type declaring one would have had its areas + titled by `.title()` on the directory name while entity's got their real + names. + + Keyed by collection rather than by directory name alone, because two types + writing into two collections may legitimately use the same area name for + different things (`kb/entities/tools/` and a hypothetical + `kb/concepts/tools/`); a flat map would hand the second one the first's + title. The collection key is the type's `base_dir:`, which is what put the + page in that directory to begin with. + """ + titles: dict[str, dict[str, str]] = {} + for type_path, _frontmatter in resolver.list_type_specs(): + try: + if resolver.get_root(type_path) != "kb": + continue + base_dir = resolver.get_base_dir(type_path) + layout = resolver.get_layout(type_path) + except (ValueError, OSError): + continue + if not base_dir or not layout: + continue + per_collection = titles.setdefault(str(base_dir).strip("/"), {}) + for key, spec in layout.items(): + per_collection.setdefault(spec.get("dir", key), spec.get("title", str(key).title())) + return titles + + +def group_pages(kb_dir: Path, pages: dict[str, Page]) -> list[Collection]: + """Group pages by their physical location: collection directory, then area + subdirectory. + + Location rather than `kind` because a shard lives in the directory it + describes, and the two agree by construction: a type-spec's `base_dir:` is + what put the page there. + """ + titles = area_titles() + grouped: dict[str, dict[str, Area]] = {} + + # Seed from the collections that exist on disk, not only from the ones that + # happen to hold pages: an empty collection is a real (if unfilled) part of + # the wiki, and dropping it from the map would hide it from every reader. + for collection_dir in iter_kb_collections(kb_dir): + grouped.setdefault(collection_dir.name, {}) + + for page in sorted(pages.values(), key=lambda p: p.title.lower()): + try: + parts = page.path.relative_to(kb_dir).parts + except ValueError: # pragma: no cover - pages always live under kb_dir + continue + if len(parts) < 2: + collection_name, area_name = "(kb root)", "" + else: + collection_name = parts[0] + area_name = parts[1] if len(parts) > 2 else "" + areas = grouped.setdefault(collection_name, {}) + area = areas.get(area_name) + if area is None: + title = ( + titles.get(collection_name, {}).get(area_name, area_name.title()) + if area_name + else UNGROUPED_TITLE + ) + area = Area(name=area_name, title=title) + areas[area_name] = area + area.pages.append(page) + + collections = [] + for name in sorted(grouped): + ordered = sorted(grouped[name].values(), key=lambda a: (a.name == "", a.title.lower())) + for area in ordered: + area.own_shard = bool(area.name) and area.count > SHARD_THRESHOLD + collections.append(Collection(name=name, areas=ordered)) + return collections diff --git a/tools/chemenu/commands/index_build.py b/tools/chemenu/commands/index_build.py index 9067d9a..3e14ea6 100644 --- a/tools/chemenu/commands/index_build.py +++ b/tools/chemenu/commands/index_build.py @@ -18,18 +18,16 @@ The map stays small enough to browse; `wikitool search` answers everything else. from __future__ import annotations import re -from dataclasses import dataclass, field from datetime import date from pathlib import Path import typer from chemenu import config +from chemenu.catalog import SHARD_THRESHOLD, Area, Collection, group_pages from chemenu.commands._util import rel_path, success -from chemenu.kb_collections import iter_kb_collections from chemenu.page import Page from chemenu.kb_scan import GENERATED_INDEX, find_nested_pages, load_kb_pages -from chemenu.type_resolver import resolver app = typer.Typer(help="Manage the generated wiki catalog (kb/index.md + per-collection INDEX.md).") @@ -38,15 +36,6 @@ TABLE_SEP = "|------|------|---------|----------------|" SUMMARY_HEADINGS = ("Description", "Definition", "Summary") -# Rows per area before it is split into its own shard. From the wiki's own -# `Index Scaling` page ("split table sections at >50 entries"), kept as a plain -# number so growth is handled by arithmetic rather than by a judgment call. -SHARD_THRESHOLD = 50 - -# Display title for pages sitting directly in a collection root rather than in -# an area subdirectory. -UNGROUPED_TITLE = "All" - DO_NOT_EDIT = "" @@ -85,86 +74,17 @@ def _table(pages: list[Page]) -> list[str]: def _anchor(title: str) -> str: - """GitHub-style heading anchor, so the map can deep-link into a shard.""" - slug = re.sub(r"[^a-z0-9\s-]", "", title.lower()) - return re.sub(r"\s+", "-", slug.strip()) + """GitHub-style heading anchor, so the map can deep-link into a shard. - -@dataclass -class Area: - """One grouping inside a collection: a subdirectory, or the collection root - for pages that sit directly in it.""" - - name: str - title: str - pages: list[Page] = field(default_factory=list) - own_shard: bool = False - - @property - def count(self) -> int: - return len(self.pages) - - -@dataclass -class Collection: - name: str - areas: list[Area] = field(default_factory=list) - - @property - def count(self) -> int: - return sum(area.count for area in self.areas) - - -def _area_titles() -> dict[str, str]: - """Display titles for entity areas, taken from the entity type-spec's own - `layout:` rather than a hardcoded map - so a new subtype names its own - section by adding a type-spec, with no code change.""" - layout = resolver.get_layout(resolver.find_type_by_name("entity")) or {} - return {spec.get("dir", key): spec.get("title", key.title()) for key, spec in layout.items()} - - -def group_pages(kb_dir: Path, pages: dict[str, Page]) -> list[Collection]: - """Group pages by their physical location: collection directory, then area - subdirectory. - - Location rather than `kind` because a shard lives in the directory it - describes, and the two agree by construction: a type-spec's `base_dir:` is - what put the page there. + `\\w` rather than `a-z0-9`, which is not cosmetic: an area title follows the + KB language, and the first non-English one (`Abläufe`) had its umlaut + *deleted* rather than kept, so the map linked at `#ablufe` and the anchor it + was aiming at was `#abläufe`. Every deep link into a shard whose title + carries a non-ASCII letter was silently dead. Nothing surfaced it while the + only areas were entity ones, whose titles happen to be ASCII throughout. """ - titles = _area_titles() - grouped: dict[str, dict[str, Area]] = {} - - # Seed from the collections that exist on disk, not only from the ones that - # happen to hold pages: an empty collection is a real (if unfilled) part of - # the wiki, and dropping it from the map would hide it from every reader. - for collection_dir in iter_kb_collections(kb_dir): - grouped.setdefault(collection_dir.name, {}) - - for page in sorted(pages.values(), key=lambda p: p.title.lower()): - try: - parts = page.path.relative_to(kb_dir).parts - except ValueError: # pragma: no cover - pages always live under kb_dir - continue - if len(parts) < 2: - collection_name, area_name = "(kb root)", "" - else: - collection_name = parts[0] - area_name = parts[1] if len(parts) > 2 else "" - areas = grouped.setdefault(collection_name, {}) - area = areas.get(area_name) - if area is None: - title = titles.get(area_name, area_name.title()) if area_name else UNGROUPED_TITLE - area = Area(name=area_name, title=title) - areas[area_name] = area - area.pages.append(page) - - collections = [] - for name in sorted(grouped): - ordered = sorted(grouped[name].values(), key=lambda a: (a.name == "", a.title.lower())) - for area in ordered: - area.own_shard = bool(area.name) and area.count > SHARD_THRESHOLD - collections.append(Collection(name=name, areas=ordered)) - return collections + slug = re.sub(r"[^\w\s-]", "", title.lower(), flags=re.UNICODE) + return re.sub(r"\s+", "-", slug.strip()) def build_area_shard(area: Area) -> str: diff --git a/tools/chemenu/lint_core.py b/tools/chemenu/lint_core.py index b7d2f38..9323728 100644 --- a/tools/chemenu/lint_core.py +++ b/tools/chemenu/lint_core.py @@ -16,7 +16,10 @@ from __future__ import annotations from datetime import date from pathlib import Path +from collections import Counter + from chemenu import blocks, config, kb_collections, links +from chemenu.catalog import SHARD_THRESHOLD, group_pages from chemenu.frontmatter_io import frontmatter_error from chemenu.markdown_code import strip_code_spans from chemenu.provenance import broken_raw_refs as find_broken_raw_refs @@ -141,6 +144,88 @@ def nested_pages(kb_dir: Path, pages: dict[str, Page]) -> list[dict]: ] +def unsharded_collections(kb_dir: Path, pages: dict[str, Page]) -> list[dict]: + """Collections past the catalog's shard threshold that have no areas to + shard, together with the subtype split that would give them some. + + Sharding is already automatic, and it is per *area*: `index rebuild` hands + an area over `SHARD_THRESHOLD` rows its own `INDEX.md`. Creating an area is + not automatic and nothing ever asked for one - so a collection that never + grew any keeps its whole catalog in a single table, past the threshold, + forever. The threshold is then not a threshold but a dead value (Gitea + #59), and this is the only check that can notice: an ingest sees one + source and cannot see a collection's size, while `lint` sees the corpus and + runs every 10 sources anyway. + + **A recommendation, not a failure** (it is deliberately absent from + `HARD_ERROR_KEYS`), and narrow enough to stay one: it fires only where the + split actually helps - every area it would create, the ungrouped remainder + included, lands at or under the threshold. That self-limits in both + directions. A collection under the threshold never fires, so a small + `kb/comparisons/` is not permanently in the report; and a collection whose + subtype values are lopsided (25 of 29 `source_type: notes`) does not fire + either, because splitting it would produce one area over the threshold and + a handful of splinters. What is left is a finding that appears when a + collection grows into it and is silent when it does not. + """ + findings: list[dict] = [] + for collection in group_pages(kb_dir, pages): + if collection.count <= SHARD_THRESHOLD: + continue + # An area already exists, so the collection has been split once and + # `index rebuild` shards whatever outgrows the threshold from here. + # A page still sitting in the root is `misplaced_pages`' finding, not + # this one. + if any(area.name for area in collection.areas): + continue + + counts: Counter[str] = Counter() + fields: set[str] = set() + # Whether every type writing here already declares the `layout:` that + # turns the subtype into a directory. It decides which half of the fix + # is still owed: without it there is nothing for `move` to compute a + # destination from, with it the move is all that is left. + layouts: set[bool] = set() + for area in collection.areas: + for page in area.pages: + type_path = page.frontmatter.get("type") + if not type_path: + continue + try: + field = resolver.get_subtype_field(type_path, page.path) + layout = resolver.get_layout(type_path, page.path) + except ValueError: + continue + value = page.frontmatter.get(field) if field else None + if not value: + continue + counts[str(value)] += 1 + fields.add(str(field)) + layouts.add(bool(layout)) + + if not counts: + continue + # The pages the subtype cannot place stay in the collection root, so + # they are an area of their own for the purpose of this test. + unplaced = collection.count - sum(counts.values()) + if max([*counts.values(), unplaced]) > SHARD_THRESHOLD: + continue + + findings.append( + { + "collection": collection.name, + "count": collection.count, + "field": ", ".join(sorted(fields)), + "layout_declared": layouts == {True}, + "distribution": [ + {"value": value, "count": count} + for value, count in sorted(counts.items(), key=lambda kv: (-kv[1], kv[0])) + ], + } + ) + return findings + + def run_lint(kb_dir: Path) -> dict: pages = load_kb_pages(kb_dir) duplicate_titles = find_duplicate_title_paths(kb_dir, config.ROOT) @@ -388,6 +473,7 @@ def run_lint(kb_dir: Path) -> dict: "duplicate_titles": duplicate_titles, "misplaced_pages": misplaced, "nested_pages": nested, + "unsharded_collections": unsharded_collections(kb_dir, pages), "uncovered_raw_files": find_uncovered_raw_files(config.RAW_DIR, pages), "broken_raw_refs": find_broken_raw_refs(pages), "duplicate_raw_file_owners": find_duplicate_raw_file_owners(pages), @@ -464,6 +550,23 @@ def render_markdown(report: dict) -> str: "the catalog folds this into its area silently; `wikitool move --reconcile` fixes it " "when the page's type resolves to a shallower directory, otherwise move it up by hand", ) + _section( + lines, f"Collections Past the Shard Threshold (>{SHARD_THRESHOLD}) With No Areas " + "- recommendation, not an error", + report.get("unsharded_collections", []), + lambda i: f"`kb/{i['collection']}/` holds {i['count']} pages in a single table and has no " + f"areas, so the per-area shard threshold never fires. Splitting on `{i['field']}` would " + "give: " + + ", ".join(f"{d['value']} {d['count']}" for d in i["distribution"]) + + f" - all at or under {SHARD_THRESHOLD}. " + + ( + "The type-spec already declares the `layout:` for those values, so " + "`wikitool move --reconcile` and `wikitool index rebuild` are the whole fix" + if i.get("layout_declared") + else "Declare a `layout:` for those values in the type-spec, then " + "`wikitool move --reconcile` and `wikitool index rebuild`" + ), + ) _section( lines, "Uncovered Raw Files (no source page)", report["uncovered_raw_files"], lambda i: f"`{i}`", @@ -617,6 +720,14 @@ def default_report_path(report: dict) -> Path: # and there is no version at which "not under the computed directory" becomes # wrong - only `wikitool move` someone does or does not get to run. # +# `unsharded_collections` is advisory by construction rather than by tolerance: +# it does not describe anything that is wrong, only a collection that has grown +# past the size at which areas start paying for themselves. Whether to split it +# is an authoring decision about how the corpus is organised - the tool can see +# that the split would work and say so, and that is the whole of its authority. +# Failing on it would also make `lint` red on a corpus that is entirely +# self-consistent, which is the state the recommendation is asking to improve. +# # `malformed_edges` and `unbalanced_markers` are hard from the start: neither # describes an unconverted page, only a broken one. # diff --git a/tools/chemenu/tests/conftest.py b/tools/chemenu/tests/conftest.py index 08a4dce..e51e31f 100644 --- a/tools/chemenu/tests/conftest.py +++ b/tools/chemenu/tests/conftest.py @@ -254,7 +254,7 @@ def kb_dir(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path: kb = tmp_path / "kb" for sub in ("entities/projects", "entities/systems", "entities/tools", "entities/technologies", "entities/people", - "concepts", "sources", "comparisons"): + "concepts/protocols", "sources", "comparisons"): (kb / sub).mkdir(parents=True) # The contracts carry a real declaration, because three things now read one: # `docs verify` checks `profile:`/`required_by_stack:`, and `xref add` asks @@ -302,7 +302,7 @@ def kb_dir(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path: "\n# gdeploy\n\n## Description\n\nDeploy tool.\n", ) write_page( - kb / "concepts/Modbus.md", + kb / "concepts/protocols/Modbus.md", { "type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25", "modified": "2026-07-25", diff --git a/tools/chemenu/tests/test_cite_cmd.py b/tools/chemenu/tests/test_cite_cmd.py index 906a364..465b08c 100644 --- a/tools/chemenu/tests/test_cite_cmd.py +++ b/tools/chemenu/tests/test_cite_cmd.py @@ -54,7 +54,7 @@ def test_cite_add_command_writes_definition_and_prints_marker(kb_dir, raw_dir, m marker_id = cite_id("Source - Aurora") assert f"[^{marker_id}]" in result.output - fm, body = read_page(kb_dir / "concepts/Modbus.md") + fm, body = read_page(kb_dir / "concepts/protocols/Modbus.md") assert "Source - Aurora" in fm["sources"] assert f"[^{marker_id}]: [[Source - Aurora]]" in body @@ -194,11 +194,11 @@ def test_cite_sync_command_over_kb(kb_dir, raw_dir, monkeypatch): assert add_result.exit_code == 0, add_result.output marker_id = cite_id("Source - Aurora") - fm, body = read_page(kb_dir / "concepts/Modbus.md") + fm, body = read_page(kb_dir / "concepts/protocols/Modbus.md") body = body.replace("Industrial protocol.", f"Industrial protocol [^{marker_id}].") from chemenu.frontmatter_io import write_page - write_page(kb_dir / "concepts/Modbus.md", fm, body) + write_page(kb_dir / "concepts/protocols/Modbus.md", fm, body) result = runner.invoke(app, ["cite", "sync", "--all", "--dry-run"]) assert result.exit_code == 0, result.output diff --git a/tools/chemenu/tests/test_git_publish.py b/tools/chemenu/tests/test_git_publish.py index 26de025..36b77f9 100644 --- a/tools/chemenu/tests/test_git_publish.py +++ b/tools/chemenu/tests/test_git_publish.py @@ -494,7 +494,7 @@ def test_generated_files_are_recognised_wherever_they_sit(): assert is_generated("kb/provenance.md") assert is_generated("kb/concepts/INDEX.md") assert is_generated("kb/entities/tools/INDEX.md") - assert not is_generated("kb/concepts/Modbus.md") + assert not is_generated("kb/concepts/protocols/Modbus.md") def test_paths_land_in_the_group_a_reviewer_expects(): diff --git a/tools/chemenu/tests/test_index_build.py b/tools/chemenu/tests/test_index_build.py index 2c5fbaf..c91d537 100644 --- a/tools/chemenu/tests/test_index_build.py +++ b/tools/chemenu/tests/test_index_build.py @@ -15,14 +15,14 @@ from chemenu.kb_scan import GENERATED_INDEX, iter_kb_pages from chemenu.type_resolver import resolver -def _area_title(subtype: str) -> str: - """The display title `index rebuild` will use for an entity subtype. +def _area_title(subtype: str, type_path: str = "types/entity.md") -> str: + """The display title `index rebuild` will use for a subtype's area. Read from the type-spec rather than written out, because these titles follow the KB language: hard-coding them made translating the wiki fail tests that are not about wording at all. """ - return resolver.get_layout("types/entity.md")[subtype]["title"] + return resolver.get_layout(type_path)[subtype]["title"] @pytest.fixture @@ -116,6 +116,28 @@ def test_area_titles_come_from_the_entity_type_spec_layout(plan, kb_dir): assert f"## {_area_title('tool')}" in entities +def test_anchor_keeps_non_ascii_letters(): + """An area title follows the KB language, so it may carry a letter outside + `a-z`. Deleting it - which is what the old `[^a-z0-9\\s-]` did - produced a + map link (`#ablufe`) that pointed at no heading in the shard it named.""" + assert _anchor("Abläufe") == "abläufe" + assert _anchor("Größere Muster") == "größere-muster" + # Punctuation is still dropped and any run of whitespace still collapses to + # a single hyphen. + assert _anchor("Tools & Utilities (v2)") == "tools-utilities-v2" + + +def test_area_titles_are_read_from_every_type_spec_not_only_entity(plan, kb_dir): + """Gitea #59: the title lookup used to resolve `entity` by name and read + only its `layout:`, so a second type declaring one got `.title()` on its + directory name (`Protocols`) instead of the title it declared. The concept + areas are the first case; nothing about them is special.""" + concepts = _shard(plan, kb_dir, "concepts") + declared = _area_title("protocol", "types/concept.md") + assert f"## {declared}" in concepts + assert "## Protocols" not in concepts + + def test_summary_prefers_frontmatter_then_falls_back_to_body(plan, kb_dir): entities = _shard(plan, kb_dir, "entities") assert "Server hosting DocStore with ZFS storage" in entities # frontmatter diff --git a/tools/chemenu/tests/test_lint.py b/tools/chemenu/tests/test_lint.py index 67dc92a..f3ee038 100644 --- a/tools/chemenu/tests/test_lint.py +++ b/tools/chemenu/tests/test_lint.py @@ -236,9 +236,133 @@ def test_lint_is_silent_about_pages_directly_in_an_area(kb_dir): assert run_lint(kb_dir)["nested_pages"] == [] +# --- collections that outgrew the shard threshold without areas (Gitea #59) --- + + +def _fill_collection(kb_dir, collection: str, type_path: str, field: str, distribution: dict): + """Write pages flat into `kb//`, `distribution` many per subtype + value - the shape a collection is in when nobody ever created an area. + + The precondition is established rather than assumed: the fixture corpus + places its concept page in an area (as the real one does now), and a single + area is enough to make this finding stand down. + """ + for existing in sorted((kb_dir / collection).rglob("*.md")): + if existing.parent != kb_dir / collection: + existing.rename(kb_dir / collection / existing.name) + for area in sorted(p for p in (kb_dir / collection).iterdir() if p.is_dir()): + area.rmdir() + + n = 0 + for value, count in distribution.items(): + for _ in range(count): + n += 1 + write_page( + kb_dir / collection / f"page-{n:03d}.md", + { + "type": type_path, field: value, + "created": "2026-08-01", "modified": "2026-08-01", + "provenance": "general", "summary": f"Page {n}", + }, + f"\n# page-{n:03d}\n", + ) + + +def test_lint_recommends_areas_for_a_collection_past_the_threshold(kb_dir): + """The concepts shape: 80 pages in one table, no areas, and a subtype axis + whose largest value (28) lands well under the threshold. Sharding is + per-area and automatic, so a collection with no areas never splits however + large it grows - the threshold is a dead value until someone makes areas.""" + _fill_collection( + kb_dir, "concepts", "types/concept.md", "concept_type", + {"workflow": 28, "architecture": 20, "pattern": 17, + "decision": 7, "problem": 5, "protocol": 3}, + ) + finding = next( + i for i in run_lint(kb_dir)["unsharded_collections"] if i["collection"] == "concepts" + ) + assert finding["field"] == "concept_type" + # Largest first, so the reader sees the area that decides whether it helps. + assert finding["distribution"][0] == {"value": "workflow", "count": 28} + assert {d["value"] for d in finding["distribution"]} == { + "workflow", "architecture", "pattern", "decision", "problem", "protocol" + } + assert finding["count"] == sum(d["count"] for d in finding["distribution"]) + # `types/concept.md` carries the layout, so only the move is still owed - + # the report line says which half of the fix that is. + assert finding["layout_declared"] is True + + +def test_the_area_recommendation_says_the_layout_is_missing_when_it_is(kb_dir): + """`types/source.md` deliberately declares no `layout:`, so a sources + collection that grew past the threshold has nothing for `move` to compute a + destination from - the fix starts one step earlier, and the report says so.""" + _fill_collection( + kb_dir, "sources", "types/source.md", "source_type", + {"notes": 30, "article": 21}, + ) + report = run_lint(kb_dir) + finding = next(i for i in report["unsharded_collections"] if i["collection"] == "sources") + assert finding["layout_declared"] is False + assert "Declare a `layout:`" in render_markdown(report) + + +def test_the_area_recommendation_is_not_a_failure(kb_dir): + """It reports a collection that has outgrown a layout, not a broken one. + A corpus whose only finding is this must stay green, or every instance + goes red on the release that shipped the check.""" + _fill_collection( + kb_dir, "concepts", "types/concept.md", "concept_type", + {"workflow": 28, "architecture": 24}, + ) + report = run_lint(kb_dir) + assert report["unsharded_collections"] + assert "unsharded_collections" not in HARD_ERROR_KEYS + assert not has_hard_errors({**{key: [] for key in HARD_ERROR_KEYS}, + "unsharded_collections": report["unsharded_collections"]}) + + +def test_lint_is_silent_about_a_collection_under_the_threshold(kb_dir): + """The sources shape: 29 pages, lopsided across `source_type` - and under + the threshold anyway, so it never fires. That is what keeps the bad split + (one area of 25 plus four splinters) from ever being recommended, without + the check needing to know anything about sources.""" + _fill_collection( + kb_dir, "sources", "types/source.md", "source_type", + {"notes": 25, "article": 3, "document": 1}, + ) + assert run_lint(kb_dir)["unsharded_collections"] == [] + + +def test_lint_is_silent_when_the_split_would_not_help(kb_dir): + """Past the threshold, but 70 of 80 share one subtype value: splitting + produces one area still over the threshold plus splinters, which is not an + improvement. The second half of the criterion, and the one a + threshold-only check would have got wrong.""" + _fill_collection( + kb_dir, "concepts", "types/concept.md", "concept_type", + {"workflow": 70, "architecture": 6, "pattern": 5}, + ) + assert run_lint(kb_dir)["unsharded_collections"] == [] + + +def test_lint_is_silent_once_the_collection_has_areas(kb_dir): + """After the fix - the pages sit in their areas - the finding goes away, + and `index rebuild` shards whatever outgrows the threshold from here.""" + _fill_collection( + kb_dir, "concepts", "types/concept.md", "concept_type", + {"workflow": 28, "architecture": 24}, + ) + for page in sorted((kb_dir / "concepts").glob("page-*.md")): + area = "workflows" if "workflow" in page.read_text(encoding="utf-8") else "architectures" + (kb_dir / "concepts" / area).mkdir(exist_ok=True) + page.rename(kb_dir / "concepts" / area / page.name) + assert run_lint(kb_dir)["unsharded_collections"] == [] + + def test_lint_flags_legacy_citation_marker_as_hard_error(kb_dir): write_page( - kb_dir / "concepts/Modbus.md", + kb_dir / "concepts/protocols/Modbus.md", {"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25", "modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7}, "\n# Modbus\n\n## Definition\n\nUses port 502 ^[[Source - Aurora]].\n", @@ -250,7 +374,7 @@ def test_lint_flags_legacy_citation_marker_as_hard_error(kb_dir): def test_lint_flags_undefined_footnote_ref_as_hard_error(kb_dir): write_page( - kb_dir / "concepts/Modbus.md", + kb_dir / "concepts/protocols/Modbus.md", {"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25", "modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7}, "\n# Modbus\n\n## Definition\n\nUses port 502 [^s-ghost].\n", @@ -264,7 +388,7 @@ def test_lint_flags_orphan_footnote_def_as_hard_error(kb_dir): cid = cite_id("Source - Aurora") block = render_cite_block({cid: ("Source - Aurora", None)}) write_page( - kb_dir / "concepts/Modbus.md", + kb_dir / "concepts/protocols/Modbus.md", {"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25", "modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7}, f"\n# Modbus\n\n## Definition\n\nIndustrial protocol, no citation here.\n\n{block}", @@ -278,7 +402,7 @@ def test_lint_clean_footnote_citation_has_no_hard_errors(kb_dir): cid = cite_id("Source - Aurora") block = render_cite_block({cid: ("Source - Aurora", None)}) write_page( - kb_dir / "concepts/Modbus.md", + kb_dir / "concepts/protocols/Modbus.md", {"type": "types/concept.md", "concept_type": "protocol", "tags": [], "created": "2026-07-25", "modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7}, f"\n# Modbus\n\n## Definition\n\nUses port 502 [^{cid}].\n\n{block}", diff --git a/tools/chemenu/tests/test_new_page.py b/tools/chemenu/tests/test_new_page.py index 69d1d1a..193eb11 100644 --- a/tools/chemenu/tests/test_new_page.py +++ b/tools/chemenu/tests/test_new_page.py @@ -43,8 +43,8 @@ def test_page_subdir_falls_back_for_unmapped_subtype(): def test_page_subdir_is_none_for_types_without_layout(): - assert _page_subdir(None, "types/concept.md") is None - assert _page_subdir("anything", "types/concept.md") is None + assert _page_subdir(None, "types/source.md") is None + assert _page_subdir("anything", "types/source.md") is None def test_coerce_set_value_uses_declared_schema_type(): @@ -258,12 +258,17 @@ def test_new_source_rejects_invalid_source_type(monkeypatch, kb_dir): assert result.exit_code != 0 -def test_new_concept_creates_page(monkeypatch, kb_dir): +def test_new_concept_creates_page_in_its_subtype_area(monkeypatch, kb_dir): + """Gitea #59: `types/concept.md` declares a `layout:` now, so a new concept + reaches its area with nothing else asked of the author - the same rule that + has always placed an entity. Nothing about `new` changed to make this true; + the type-spec did.""" result = _invoke_new(monkeypatch, kb_dir, [ "new", "concept", "--name", "Event Sourcing", "--set", "concept_type=pattern", ]) assert result.exit_code == 0, result.output - assert (kb_dir / "concepts/Event Sourcing.md").exists() + assert (kb_dir / "concepts/patterns/Event Sourcing.md").exists() + assert not (kb_dir / "concepts/Event Sourcing.md").exists() def test_new_concept_rejects_invalid_concept_type(monkeypatch, kb_dir): diff --git a/tools/chemenu/tests/test_provenance.py b/tools/chemenu/tests/test_provenance.py index 3305506..cf73c4d 100644 --- a/tools/chemenu/tests/test_provenance.py +++ b/tools/chemenu/tests/test_provenance.py @@ -310,7 +310,7 @@ def test_citing_pages_via_frontmatter_and_inline(kb_dir, raw_dir): ) refs, block = _footnote_block(("Source - Aurora", None)) write_page( - kb_dir / "concepts/Modbus.md", + kb_dir / "concepts/protocols/Modbus.md", { "type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25", "modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7, @@ -325,7 +325,7 @@ def test_citing_pages_via_frontmatter_and_inline(kb_dir, raw_dir): def test_page_raw_files_resolves_through_sources_and_inline(kb_dir, raw_dir): write_page( - kb_dir / "concepts/Modbus.md", + kb_dir / "concepts/protocols/Modbus.md", { "type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25", "modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7, @@ -429,7 +429,7 @@ def test_lint_flags_citation_not_in_frontmatter_sources(kb_dir, raw_dir, monkeyp refs, block = _footnote_block(("Source - Aurora", None)) write_page( - kb_dir / "concepts/Modbus.md", + kb_dir / "concepts/protocols/Modbus.md", { "type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25", "modified": "2026-07-25", "related": [], "sources": [], "confidence": 0.7, @@ -451,7 +451,7 @@ def test_lint_no_drift_when_source_declared(kb_dir, raw_dir, monkeypatch): refs, block = _footnote_block(("Source - Aurora", None)) write_page( - kb_dir / "concepts/Modbus.md", + kb_dir / "concepts/protocols/Modbus.md", { "type": "concept", "concept_type": "protocol", "tags": [], "created": "2026-07-25", "modified": "2026-07-25", "related": [], "sources": ["Source - Aurora"], "confidence": 0.7, diff --git a/tools/chemenu/tests/test_type_resolver.py b/tools/chemenu/tests/test_type_resolver.py index 6c0806f..8969dc6 100644 --- a/tools/chemenu/tests/test_type_resolver.py +++ b/tools/chemenu/tests/test_type_resolver.py @@ -86,9 +86,34 @@ def test_get_layout_reads_entity_type_specs_own_layout_field(): assert list(layout) == ["project", "system", "tool", "technology", "person"] +def test_concept_layout_covers_every_declared_concept_type(): + """Gitea #59: `kb/concepts/` had no areas, so the catalog's per-area shard + threshold could never fire however large it grew. A value missing from the + layout would still be *placed* (`subtype_dir` pluralizes the fallback), but + into an area with no declared title - so the schema's enum and the layout + have to agree, and this is what checks that they do.""" + layout = resolver.get_layout("types/concept.md") + assert layout is not None + assert set(layout) == set(resolver.get_enum("types/concept.md", "concept_type")) + # `dir` is structural; `title` is display text following the KB language, so + # it is checked for presence rather than wording (see the entity test above). + assert {key: spec["dir"] for key, spec in layout.items()} == { + "architecture": "architectures", + "pattern": "patterns", + "protocol": "protocols", + "workflow": "workflows", + "decision": "decisions", + "problem": "problems", + } + assert all(spec.get("title") for spec in layout.values()) + + def test_get_layout_is_none_for_types_without_one(): + """`comparison` has no subtype field at all; `source` has one and + deliberately declares no `layout:` anyway - 25 of its 29 pages carry the + same `source_type`, so splitting on it would make one area and four + splinters (Gitea #59). Having a subtype axis is not a reason to use it.""" assert resolver.get_layout("types/comparison.md") is None - assert resolver.get_layout("types/concept.md") is None assert resolver.get_layout("types/source.md") is None @@ -237,7 +262,7 @@ def test_subtype_dir_falls_back_for_unmapped_subtype(): def test_subtype_dir_is_none_without_layout_or_subtype(): - assert resolver.subtype_dir("types/concept.md", "workflow") is None + assert resolver.subtype_dir("types/source.md", "notes") is None assert resolver.subtype_dir("types/entity.md", None) is None @@ -247,8 +272,8 @@ def test_compute_target_dir_applies_layout_subdirectory(): def test_compute_target_dir_is_flat_for_a_type_without_layout(): - target = resolver.compute_target_dir("types/concept.md", {"concept_type": "workflow"}) - assert target == config.KB_DIR / "concepts" + target = resolver.compute_target_dir("types/source.md", {"source_type": "notes"}) + assert target == config.KB_DIR / "sources" def test_compute_target_dir_resolves_against_repo_root_for_root_repo_types(): diff --git a/tools/chemenu/tests/test_xref.py b/tools/chemenu/tests/test_xref.py index e9bf4eb..e861bf9 100644 --- a/tools/chemenu/tests/test_xref.py +++ b/tools/chemenu/tests/test_xref.py @@ -121,7 +121,7 @@ def test_xref_add_updates_both_pages_on_disk(kb_dir): config.INDEX_FILE = kb_dir / "index.md" runner = CliRunner() - modbus_before = (kb_dir / "concepts/Modbus.md").read_text(encoding="utf-8") + modbus_before = (kb_dir / "concepts/protocols/Modbus.md").read_text(encoding="utf-8") result = runner.invoke(app, ["xref", "add", "--a", "gdeploy", "--b", "Modbus", "--rel", "uses"]) assert result.exit_code == 0, result.output @@ -131,7 +131,7 @@ def test_xref_add_updates_both_pages_on_disk(kb_dir): # B is not touched at all. Its inbound view is rendered from the graph, so # nothing has to be written there for a reader to find its way back. - assert (kb_dir / "concepts/Modbus.md").read_text(encoding="utf-8") == modbus_before + assert (kb_dir / "concepts/protocols/Modbus.md").read_text(encoding="utf-8") == modbus_before _fm_before, body_before = read_page(kb_dir / "entities/tools/gdeploy.md") link_count_before = body_before.count("[[Modbus]]") @@ -238,7 +238,7 @@ def test_xref_add_dry_run_writes_nothing(kb_dir): config.INDEX_FILE = kb_dir / "index.md" gdeploy_path = kb_dir / "entities/tools/gdeploy.md" - modbus_path = kb_dir / "concepts/Modbus.md" + modbus_path = kb_dir / "concepts/protocols/Modbus.md" gdeploy_before = gdeploy_path.read_text(encoding="utf-8") modbus_before = modbus_path.read_text(encoding="utf-8") diff --git a/types/concept.md b/types/concept.md index bedd4e2..4305754 100644 --- a/types/concept.md +++ b/types/concept.md @@ -6,6 +6,13 @@ schema: types/concept.schema.yaml subtype_field: concept_type base_dir: concepts page_ref_fields: [related, sources] +layout: + architecture: {dir: architectures, title: Architekturen} + pattern: {dir: patterns, title: Muster} + protocol: {dir: protocols, title: Protokolle} + workflow: {dir: workflows, title: Abläufe} + decision: {dir: decisions, title: Entscheidungen} + problem: {dir: problems, title: Problemstellungen} --- # Concept diff --git a/types/type-spec.md b/types/type-spec.md index 5ab2043..3675007 100644 --- a/types/type-spec.md +++ b/types/type-spec.md @@ -99,6 +99,16 @@ nothing can write into the real `kb/`. `root: repo` is opt-in, for types whose a legitimately not knowledge: `instruction` is the worked example - an agent-directed procedure under `instructions/`, outside the `raw/` -> `kb/` provenance rules entirely. +`layout` is what turns a subtype field into areas, and it is optional in both directions on +purpose. Declaring one makes every instance land in a directory named by its subtype and gives +that directory its index title; leaving it off keeps the collection flat, however many subtypes +the schema allows. The choice is whether the split *helps* - the catalog shards per area, so +areas are how a large collection stays readable, but a subtype whose values are lopsided +produces one big area and a handful of splinters and is better left flat (`source` is the worked +example, `concept` the opposite one). `wikitool lint` will say which case a collection is in +once it grows past the shard threshold; a `dir:` must be a single path segment, because an area +is as deep as a page goes. + A type-spec **without `base_dir` is a contract only** and cannot be instantiated; `wikitool new ` refuses it. `lint-report` is the worked example: its artifacts are generated into `reports/`, so it describes a file format without owning a place anything writes