feat: people live on their organization's page until promoted; organization subtype, member-of, broken_anchors lint (#172)
CI / verify (push) Successful in 5m26s
CI / pwsh (push) Successful in 2m8s
Release / release (push) Successful in 34s

Files changed:
- CHANGES.md
- README.md
- VERSION
- instructions/kb-profiles.md
- instructions/link-taxonomy.md
- instructions/page-lifecycle.md
- instructions/wiki-ingest/SKILL.md
- instructions/wiki-lint/SKILL.md
- kb/entities/COLLECTION.md
- kb/entities/INDEX.md
- kb/entities/organizations/E3DC GmbH.md
- kb/entities/people/E3DC GmbH.md
- kb/index.md
- kb/log.md
- tools/CONTRACT.md
- tools/chemenu/commands/lint.py
- tools/chemenu/kb_scan.py
- tools/chemenu/lint_core.py
- tools/chemenu/tests/test_lint.py
- tools/chemenu/tests/test_new_page.py
- tools/chemenu/tests/test_type_resolver.py
- tools/chemenu/tests/test_types_cmd.py
- types/entity.md
- types/entity.organization.md
- types/entity.person.md
- types/entity.schema.yaml

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SnAJ7Z3CpVD3PRbN73QtU2
This commit is contained in:
torbenandClaude Opus 5.5 committed 2026-10-04 10:51:46 +02:00
1 parent f3ccbd86f9
commit 4ec22d376d
25 files changed
+368 -30

No files matched your search

+1
View File
@@ -1124,6 +1124,7 @@ Run structural lint checks against kb/.
- Structural and provenance checks over `kb/`: broken wikilinks, wikilinks wrapped across a line break, dangling frontmatter references, orphan pages, index drift, schema gaps, duplicate titles, title mismatches, uncovered raw files, broken `raw_files:` refs, raw files claimed by more than one source page, unmarked provenance, citation/frontmatter drift, and unbalanced generated-region markers.
- Unportable Titles is a hard finding, and hard at every `kb_version`: a page whose title is not a valid file name on Windows and macOS (forbidden character, reserved name, trailing dot or space), or that collides with another page by case or Unicode normalization. `wikitool rename` is the fix.
- Wrapped Wikilinks is a hard finding: a `[[...]]` with a line break inside it. The link graph reads it as the title it folds to (the break and its indentation become one space), so it is not also a broken link unless that title is missing; the fix is to put it back on one line.
- Broken Anchors is advisory: a `[[Page#Section]]` whose page exists but has no heading the anchor names (any level, compared without case, inline-code backticks or extra whitespace; every segment of a nested `[[Page#A#B]]`). The link still reaches the page, so nothing else reports it - typically a section that was promoted to a page of its own or renamed. A missing page is Broken Wikilinks instead.
- Pages nested more than one directory below their collection are a hard finding - the generated catalog folds these into their area silently rather than merely reading it.
- Edges whose label is missing or not authorised by the source collection's `outbound:` are both hard once `kb_version` has reached the release that introduced labelled edges, and advisory below it.
- Advisory only: `see-also` edges whose reverse direction already carries a specific label - never migration-gated.
+5
View File
@@ -73,6 +73,11 @@ __all__ = [
"graph reads it as the title it folds to (the break and its indentation become one "
"space), so it is not also a broken link unless that title is missing; the fix is to put "
"it back on one line.",
"Broken Anchors is advisory: a `[[Page#Section]]` whose page exists but has no heading "
"the anchor names (any level, compared without case, inline-code backticks or extra "
"whitespace; every segment of a nested `[[Page#A#B]]`). The link still reaches the page, "
"so nothing else reports it - typically a section that was promoted to a page of its "
"own or renamed. A missing page is Broken Wikilinks instead.",
"Pages nested more than one directory below their collection are a hard finding - the "
"generated catalog folds these into their area silently rather than merely reading it.",
"Edges whose label is missing or not authorised by the source collection's `outbound:` "
+56
View File
@@ -144,6 +144,62 @@ def wrapped_wikilinks(body: str) -> list[str]:
]
# `[[Target#Section]]`, `[[Target#Section#Sub|alias]]`: group 1 is the target,
# exactly as `WIKILINK_RE` reads it; group 2 is everything from the first `#` up
# to an alias or the closing brackets. A bare self-anchor `[[#Section]]` names
# no target and is not matched - `WIKILINK_RE` reads no link there either.
ANCHORED_WIKILINK_RE = re.compile(r"\[\[([^\]|#]+)#([^\]|]*)")
# An ATX heading at any level. Optional closing hashes are not part of its text.
_HEADING_LINE_RE = re.compile(r"^#{1,6}[ \t]+(.+?)(?:[ \t]+#+)?[ \t]*$")
def normalize_anchor(text: str) -> str:
"""The form a section anchor and a heading are compared in: inline-code
backticks dropped, whitespace runs folded, case folded.
Deliberately loose. An anchor is written by hand from a heading a reader
saw rendered, so `[[Kunde X#anna müller]]` means the `### Anna Müller`
section; what `broken_anchors` is for is a section that is *gone*, not one
spelled with different capitals.
"""
return " ".join(text.replace("`", "").split()).casefold()
def heading_anchors(body: str) -> set[str]:
"""Every heading in this body, at any level, in `normalize_anchor` form.
Code-aware like `toc.iter_headings`: detection runs on the masked body, so
a `# comment` line inside a fence is not a heading, and the text is read
back from the unmasked line at the same position.
"""
original = body.split("\n")
masked = strip_code_spans(body).split("\n")
anchors: set[str] = set()
for masked_line, original_line in zip(masked, original):
if not masked_line.startswith("#") or not _HEADING_LINE_RE.match(masked_line):
continue
match = _HEADING_LINE_RE.match(original_line)
if match:
anchors.add(normalize_anchor(match.group(1)))
return anchors
def anchored_wikilinks(body: str) -> list[tuple[str, str]]:
"""`(target, anchor)` for every wikilink in this body that names a
section, in order of appearance, code masked out as everywhere else.
The target is normalized like every other reader's; the anchor is returned
as written (`A#B` for a nested one), so a report can quote it - compare it
through `normalize_anchor`, segment by segment.
"""
return [
(normalize_link_target(m.group(1)), m.group(2).strip())
for m in ANCHORED_WIKILINK_RE.finditer(strip_code_spans(body))
if m.group(2).strip()
]
def count_wikilinks(body: str) -> Counter[str]:
"""How often this body links to each page.
+42
View File
@@ -36,12 +36,15 @@ from chemenu.version import Version
from chemenu.kb_scan import (
GENERATED_INDEX,
WIKILINK_RE,
anchored_wikilinks,
build_link_graph,
find_duplicate_title_paths,
find_nested_pages,
heading_anchors,
inbound_links,
iter_kb_pages,
load_kb_pages,
normalize_anchor,
normalize_link_target,
wrapped_wikilinks,
)
@@ -303,6 +306,33 @@ def _repo_relative(path: Path, kb_dir: Path) -> str:
return path.relative_to(kb_dir.parent).as_posix()
def broken_anchors(pages: dict[str, Page]) -> list[dict]:
"""`{"page", "target", "anchor"}` for every `[[Target#Section]]` whose
target page exists but carries no heading the anchor names.
A missing target is `broken_links`' finding and is not repeated here. A
nested anchor `[[T#A#B]]` resolves when every segment is a heading on `T` -
the section path a renderer would walk, without insisting on the nesting.
The case this exists for is a section that went away: a person promoted
from their organization's page to one of their own, a section renamed. The
link still reaches the right page, so nothing else reports it, and it no
longer reaches the part of the page it was written for.
"""
headings: dict[str, set[str]] = {}
findings = []
for title, page in sorted(pages.items()):
for target, anchor in sorted(set(anchored_wikilinks(page.body))):
if target not in pages:
continue
if target not in headings:
headings[target] = heading_anchors(pages[target].body)
segments = [normalize_anchor(s) for s in anchor.split("#") if s.strip()]
if not all(segment in headings[target] for segment in segments):
findings.append({"page": title, "target": target, "anchor": anchor})
return findings
def run_lint(kb_dir: Path) -> dict:
pages = load_kb_pages(kb_dir)
duplicate_titles = find_duplicate_title_paths(kb_dir, config.ROOT)
@@ -550,6 +580,7 @@ def run_lint(kb_dir: Path) -> dict:
"frontmatter_errors": frontmatter_errors,
"broken_links": broken_links,
"wrapped_wikilinks": wrapped_links,
"broken_anchors": broken_anchors(pages),
"orphan_pages": orphan_pages,
"most_linked": most_linked,
"inbound_counts": inbound_counts,
@@ -615,6 +646,11 @@ def render_markdown(report: dict) -> str:
lines, "Wrapped Wikilinks", report.get("wrapped_wikilinks", []),
lambda i: f"[[{i['page']}]] wraps a wikilink across lines - write it on one: [[{i['target']}]]",
)
_section(
lines, "Broken Anchors - advisory, not an error", report.get("broken_anchors", []),
lambda i: f"[[{i['page']}]] links to [[{i['target']}#{i['anchor']}]], but [[{i['target']}]] "
"has no such heading - point the link at the page that section became, or drop the anchor",
)
_section(lines, "Orphan Pages (no inbound links)", report["orphan_pages"], lambda i: f"[[{i}]]")
_section(
lines, f"Most-Linked Pages (top {MOST_LINKED_COUNT} hubs)", report["most_linked"],
@@ -864,6 +900,12 @@ def default_report_path(report: dict) -> Path:
# it needs no migration - so failing a lint run on it would penalise an instance
# that never asked for that platform.
#
# `broken_anchors` is advisory: the link still reaches the page it names, only
# not the section, which is a weaker defect than `broken_links`. And it arrived
# after corpora that may already carry such links - promoting it would turn an
# instance's lint red on the upgrade that shipped it, which no other part of
# that upgrade asked for.
#
# `wrapped_wikilinks` is hard from the start without tightening anything: every
# link it names was a hard `broken_links` finding before the link graph learned
# to fold a wrapped target, so an instance's lint is exactly as red as it was -
+76
View File
@@ -103,6 +103,82 @@ def test_a_wrapped_link_inside_code_is_no_finding(kb_dir):
assert report["broken_links"] == []
def _organization_with_people(kb_dir) -> None:
"""An organization page holding its people as sections - the shape
`broken_anchors` exists for - plus a heading that only appears inside a
fence, which must not count as one."""
write_page(
kb_dir / "entities/organizations/Kunde X.md",
{"type": "types/entity.md", "entity_type": "organization", "tags": [],
"created": "2026-10-04", "modified": "2026-10-04", "related": [], "sources": []},
"\n# Kunde X\n\n## Personen\n\n### Anna Müller\n\nEinkauf.\n\n"
"#### `Bob` Builder ##\n\nBetrieb.\n\n```\n## Fenced\n```\n",
)
def test_an_anchor_naming_an_existing_heading_is_no_finding(kb_dir):
"""Any heading level, any capitalisation, inline-code backticks and closing
hashes ignored, an alias after the anchor, a nested anchor whose every
segment is a heading."""
_organization_with_people(kb_dir)
_gdeploy_saying(
kb_dir,
"[[Kunde X#Anna Müller]], [[Kunde X#anna MÜLLER|Anna]], [[Kunde X#Bob Builder]], "
"[[Kunde X#Personen#Anna Müller]], [[Kunde X#Kunde X]].",
)
report = run_lint(kb_dir)
assert report["broken_anchors"] == []
assert report["broken_links"] == []
def test_an_anchor_naming_no_heading_is_a_finding(kb_dir):
"""The promoted-person case: the page is still there, the section is not.
The anchor is reported as written, not in its normalized form."""
_organization_with_people(kb_dir)
_gdeploy_saying(kb_dir, "Ask [[Kunde X#Carla Neu|Carla]] or [[Kunde X#Personen#Dora]].")
report = run_lint(kb_dir)
assert report["broken_anchors"] == [
{"page": "gdeploy", "target": "Kunde X", "anchor": "Carla Neu"},
{"page": "gdeploy", "target": "Kunde X", "anchor": "Personen#Dora"},
]
assert report["broken_links"] == []
assert "has no such heading" in render_markdown(report)
def test_a_heading_inside_a_fence_is_not_a_section(kb_dir):
_organization_with_people(kb_dir)
_gdeploy_saying(kb_dir, "See [[Kunde X#Fenced]].")
assert run_lint(kb_dir)["broken_anchors"] == [
{"page": "gdeploy", "target": "Kunde X", "anchor": "Fenced"},
]
def test_an_anchor_on_a_missing_page_is_only_a_broken_link(kb_dir):
_gdeploy_saying(kb_dir, "See [[Nonexistent Page#Somewhere]].")
report = run_lint(kb_dir)
assert {"page": "gdeploy", "target": "Nonexistent Page"} in report["broken_links"]
assert report["broken_anchors"] == []
def test_an_anchor_inside_code_and_a_bare_self_anchor_are_no_finding(kb_dir):
_organization_with_people(kb_dir)
_gdeploy_saying(
kb_dir,
"Write `[[Kunde X#Nowhere]]` like this, or:\n\n```\n[[Kunde X#Nowhere]]\n```\n\n"
"Jump to [[#Nowhere]].",
)
assert run_lint(kb_dir)["broken_anchors"] == []
def test_a_broken_anchor_alone_keeps_lint_green():
"""Advisory: the link still reaches its page, and a hard finding would turn
an existing instance's lint red on the upgrade that shipped the check."""
assert "broken_anchors" not in HARD_ERROR_KEYS
report = {key: [] for key in HARD_ERROR_KEYS}
report["broken_anchors"] = [{"page": "a", "target": "b", "anchor": "c"}]
assert not has_hard_errors(report)
def test_lint_fixture_has_no_dangling_frontmatter_refs(kb_dir):
assert run_lint(kb_dir)["dangling_frontmatter_refs"] == []
+13
View File
@@ -1127,6 +1127,19 @@ def test_new_person_scaffolds_the_person_template(monkeypatch, kb_dir):
assert "{" not in body
def test_new_organization_scaffolds_the_organization_template(monkeypatch, kb_dir):
"""An organization is its own subtype with its own area, and its skeleton
carries the section its people live in until they earn a page."""
result = _invoke_new(monkeypatch, kb_dir, [
"new", "entity", "--name", "Kunde X", "--set", "entity_type=organization",
])
assert result.exit_code == 0, result.output
_fm, body = read_page(kb_dir / "entities/organizations/Kunde X.md")
assert body.strip() == _rendered("entity", "entity_type", "organization", "Kunde X")
assert "## Personen" in body
assert "{" not in body
def test_new_decision_scaffolds_the_decision_template(monkeypatch, kb_dir):
result = _invoke_new(monkeypatch, kb_dir, [
"new", "concept", "--name", "Flat Subtype Files", "--set", "concept_type=decision",
+3 -2
View File
@@ -8,7 +8,7 @@ def test_get_enum_returns_schema_declared_values():
"""entity_type's valid values come from entity.schema.yaml's enum - this
is what lets config.py and new_page.py stop hand-maintaining that list."""
values = resolver.get_enum("types/entity.md", "entity_type")
assert values == ["codebase", "system", "tool", "technology", "person"]
assert values == ["codebase", "system", "tool", "technology", "person", "organization"]
def test_get_enum_shared_across_types():
@@ -147,10 +147,11 @@ def test_get_layout_reads_entity_type_specs_own_layout_field():
"tool": "tools",
"technology": "technologies",
"person": "people",
"organization": "organizations",
}
assert all(spec.get("title") for spec in layout.values())
# Order drives kb/index.md section order.
assert list(layout) == ["codebase", "system", "tool", "technology", "person"]
assert list(layout) == ["codebase", "system", "tool", "technology", "person", "organization"]
def test_concept_layout_covers_every_declared_concept_type():
+1 -1
View File
@@ -41,7 +41,7 @@ def test_types_describe_entity_reports_schema_and_body():
fields_by_name = {f["field"]: f for f in data["fields"]}
assert fields_by_name["entity_type"]["required"] is True
assert fields_by_name["entity_type"]["enum"] == [
"codebase", "system", "tool", "technology", "person",
"codebase", "system", "tool", "technology", "person", "organization",
]
assert fields_by_name["tags"]["required"] is False
# The body must carry the page skeleton an authoring LLM works from...