Files
chemenu/tools/chemenu/toc.py
T
torben f3c80747a5
CI / verify (push) Successful in 48s
Release / release (push) Successful in 41s
docs toc/verify: die .template-Form einer Referenzdatei steht im Dateisatz, Version 6.0.1 freigegeben (schliesst #106)
Files changed:
- CHANGES.md
- VERSION
- instructions/dev/doc-pull-through.md
- kb/CONVENTIONS.md.template
- tools/CONTRACT.md
- tools/chemenu/commands/docs_verify.py
- tools/chemenu/tests/test_docs_verify.py
- tools/chemenu/tests/test_toc.py
- tools/chemenu/toc.py
2026-09-16 06:43:36 +02:00

325 lines
15 KiB
Python

"""Generate the table-of-contents region for reference files over 100 lines.
Anthropic's skill-authoring guidance: "For reference files longer than 100
lines, include a table of contents at the top. This ensures Claude can see
the full scope of available information even when previewing with partial
reads." (`codex-skill-creator/SKILL.md:221-222` names the same threshold as
the mitigation for the exact preview mechanic `instructions/CONTRACT.md`
§ "Reference depth" already treats as real for a repo-wide contract reached
at a second hop.)
A hand-maintained TOC is the next drift source the moment a heading changes -
AGENTS.md invariant 1 ("never hand-edit generated files") applies here the
same way it applies to a page's links/footnotes region. So this is a third
generated region beside `xref`'s and `cite`'s, reusing `blocks`' marker
convention (`<!-- wikitool:toc -->` ... `<!-- /wikitool:toc -->`) without
joining `blocks.BLOCKS`: that tuple feeds `xref`, `cite` and the KB-page
`unbalanced_markers` lint check, and none of the target files here (AGENTS.md,
the stage/collection contracts, the flat `instructions/**.md` files) is a
`kb/` page. `version.py`'s `bumps` region inside `CHANGES.md` set this same
precedent first.
**Scope is computed, never a hand-picked list** - the same principle that
governs `wikitool` itself. `target_files()` walks every file-naming category
AGENTS.md's own table calls agent-loaded: `AGENTS.md`, every stage contract,
`kb/CONVENTIONS.md`, every `kb/*/COLLECTION.md`, every flat `instructions/**.md`
file, every `types/*.md` type-spec, and every `docs/` page - each together with
the `<name>.template` it ships as, where one exists. One rule, and
exactly one exception below it - which is the whole point, because a scope
carrying several unexplained absences reads as an accident rather than a
decision, and did: `docs/` and the page type-specs sat outside it for no
recorded reason at all.
`docs/` belongs in for the reason the threshold exists. AGENTS.md § File naming
calls it "Agents and humans | By link, or on explicit request" - agent-loaded,
at a second hop, which is precisely the preview mechanic
`instructions/CONTRACT.md` § "Reference depth" treats as real. The exclusion of
the human docs (`README.md`, `CHANGES.md`, `EVALS.md`, `INSTALL.md`,
`tools/README.md`) rests on a sentence that does not stretch to cover it: those
are "Never loaded by an agent as instruction," while a `docs/` page is not
loaded *as instruction* but is very much loaded.
**The one exception is `SKILL.md`**, and the vendored skill-authoring guidance
is what puts it there rather than a judgment of ours. `commonplace/kb/work/
skill-creator-distillation/sources/claude-code-skill-creator/SKILL.md:88-99`
sets out three loading levels and places the SKILL.md body on the second - "In
context whenever skill triggers (<500 lines ideal)" - while aiming its own TOC
advice at the third, "large reference files (>300 lines)", i.e. bundled
resources. The Codex source agrees by placement: its TOC bullet sits directly
under "Keep references one level deep from SKILL.md. All reference files should
link directly from SKILL.md" (`.../codex-skill-creator/SKILL.md:221-222`).
Neither asks a skill body to carry a table of contents, because neither expects
one to be previewed. `instructions/CONTRACT.md` § "How much reasoning a step
may carry" arrives at the same place from the other side, treating a checklist
read once as the table of contents it replaced. (Our longest skill is 254
lines, so the <500 guidance costs us nothing either.)
A type-spec is loaded whole too - `tools/wikitool types describe` prints its
entire authoring body - but it is *also* read as a file, by whoever edits it,
and that is the reading the threshold is about. So it carries a region like any
other reference file, and `types_cmd` strips the region back out of what
`describe` prints: the command already hands over the whole body, so a
navigation aid into it would be noise in the output and nothing else.
The threshold's own provenance is worth recording, because the two sources
disagree: Codex says 100 lines, Claude Code says 300. This stack took the
stricter number.
"""
from __future__ import annotations
import re
from pathlib import Path
from typing import NamedTuple
from chemenu import blocks, config, kb_collections, markdown_code
REGION_NAME = "toc"
HEADING_TEXT = "Contents"
# The suffix `dist export` re-keys an instance-owned file to, and the one
# `setup-instance.md` adopts away again. Lives here because this module is what
# decides which files are reference material *in both their forms*;
# `docs_verify` imports it rather than keeping a second spelling. `ownership.py`
# keeps its own literal deliberately - it answers a different question (which
# side an upstream merge keeps) over a narrower scope.
TEMPLATE_SUFFIX = ".template"
# The line threshold Anthropic's own guidance names. Measured on the body
# with any existing TOC region stripped out first, so inserting or updating
# the region can never be what pushes a file over the line.
THRESHOLD = 100
# Mirrors docs_verify.STAGE_CONTRACTS - kept as a separate constant rather than
# imported, because `docs_verify` importing `toc` (for the new check) would
# make a `toc` -> `docs_verify` import a cycle. Both lists are the seven stage
# contracts; `docs_verify.STAGE_CONTRACTS`'s own docstring is the place that
# explains what they are.
_STAGE_CONTRACTS = (
"raw/CONTRACT.md",
"kb/CONTRACT.md",
"types/type-spec.md",
"reports/CONTRACT.md",
"work/CONTRACT.md",
"tools/CONTRACT.md",
"instructions/CONTRACT.md",
)
_HEADING_RE = re.compile(r"^(#{2,3})[ \t]+(.+?)[ \t]*$")
_SLUG_DROP_RE = re.compile(r"[^\w\s-]", re.UNICODE)
_SLUG_SPACE_RE = re.compile(r"\s+")
class Heading(NamedTuple):
level: int # 2 or 3
text: str
def target_files() -> list[Path]:
"""Every file the TOC region applies to, root-relative, sorted.
Walking `instructions/` picks up `instructions/dev/` and
`instructions/migrations/` along with the rest: both are still the flat
`instructions/<name>.md` form (`instructions/CONTRACT.md` says the `dev/`
split is orthogonal to Linked/Manual, not a different file shape), so
there is no separate rule for them to fall out of - and no hand-picking
for a future file under either to be missed. `types/` and `docs/` are
walked for the same reason: a type or a rationale page added later is in
scope by construction, not by someone remembering this function.
`types/type-spec.md` arrives twice - once as a stage contract, once from
the `types/` walk - and the set at the bottom is what makes that a
non-issue rather than something to special-case.
**A file in scope carries its `.template` in with it.** An instance-owned
reference file crosses the distribution boundary as `<name>.template` and
is adopted by copying it back (`instructions/setup-instance.md`), so the
template is the same document one step earlier in its life - the shipped
form of a file this scope already covers. Leaving it out meant nothing
maintained it and nothing checked it: `kb/CONVENTIONS.md.template` grew
past the threshold carrying no region at all, and the first thing to notice
was a fresh instance failing `docs verify` at the end of its own setup,
because the adopted copy inherited the gap. `docs toc --apply` writes the
region into the template as into anything else here; the instance's own
later edits stay its business, and its own `docs toc` run answers for them.
"""
files: list[Path] = [config.ROOT / "AGENTS.md"]
files += [config.ROOT / rel for rel in _STAGE_CONTRACTS]
files.append(config.ROOT / "kb" / "CONVENTIONS.md")
files += sorted(
collection / kb_collections.CONTRACT_NAME
for collection in kb_collections.iter_kb_collections()
)
instructions_dir = config.ROOT / "instructions"
if instructions_dir.is_dir():
files += sorted(
path
for path in instructions_dir.rglob("*.md")
if path.name != "SKILL.md"
)
for subdir in ("types", "docs"):
directory = config.ROOT / subdir
if directory.is_dir():
files += sorted(directory.rglob("*.md"))
files += [
template
for template in (
path.with_name(path.name + TEMPLATE_SUFFIX) for path in tuple(files)
)
if template.is_file()
]
return sorted({f for f in files if f.is_file()})
def body_without_region(text: str) -> str:
"""`text` with any existing TOC region removed - what the 100-line
threshold is measured against, so a stale TOC never counts toward keeping
itself around.
Trailing newlines are canonicalized to exactly one (`blocks.strip`'s own
removal path already does this when a region existed; a file with no
region yet is normalized here too), which is what keeps `upsert`
idempotent - without it, a file missing its final newline would insert
its first TOC one way and every later re-run one byte shorter, since the
"remove an existing region" and "there was never one" paths would
otherwise start from differently-terminated strings.
"""
stripped = blocks.strip(text, REGION_NAME)
return stripped.rstrip("\n") + "\n" if stripped else stripped
_REGION_WITH_PADDING_RE = re.compile(
r"\n*"
+ re.escape(blocks.open_marker(REGION_NAME))
+ r".*?"
+ re.escape(blocks.close_marker(REGION_NAME))
+ r"\n*",
re.DOTALL,
)
def strip_region(text: str) -> str:
"""`text` with the generated TOC region removed, for a caller that is
handing the whole body over anyway.
`types describe` is the one such caller: it prints a type-spec's entire
authoring body, so the region's markers and heading list would be noise in
its output rather than a way into anything. Distinct from
`body_without_region`, which exists to *measure* a body against the line
threshold and therefore also canonicalizes the trailing newline.
Takes the surrounding blank lines with it and puts one back, rather than
calling `blocks.strip`: that collapses the padding to a single newline,
which is the right answer when the result is about to be rebuilt from
scratch (`upsert` does exactly that) and the wrong one here, where the
output is printed as-is - it would leave the following `##` heading welded
to the paragraph above it.
"""
return _REGION_WITH_PADDING_RE.sub("\n\n", text)
def needs_toc(text: str) -> bool:
return len(body_without_region(text).splitlines()) > THRESHOLD
def iter_headings(body: str) -> list[Heading]:
"""Every `##`/`###` heading in document order, code-aware.
Detection runs on `markdown_code.strip_code_spans(body)`, which blanks
fenced blocks and inline code spans but preserves line structure - a
template line inside a fence (`# {Type name}`, a `raw/CONTRACT.md`
example path prefixed `#`) no longer starts with a real `#` marker once
masked, so it is not mistaken for a section. The heading *text* is read
back from the corresponding *unmasked* line, not the masked one: masking
also blanks inline code spans, which would otherwise turn a heading like
`` ## `instructions/dev/` `` into blank spaces instead of its real text.
"""
original = body.split("\n")
masked = markdown_code.strip_code_spans(body).split("\n")
headings: list[Heading] = []
for masked_line, original_line in zip(masked, original):
if not masked_line.startswith("#"):
continue
match = _HEADING_RE.match(masked_line)
if not match:
continue
# The masked line only still starts with `#...` for a genuine heading
# (fenced lines are blanked to spaces); recover the text from the
# unmasked line at the same position.
original_match = _HEADING_RE.match(original_line)
if not original_match:
continue
headings.append(Heading(level=len(match.group(1)), text=original_match.group(2).strip()))
return headings
def _slugify(text: str, seen: dict[str, int]) -> str:
"""A GitHub/Gitea-style anchor slug, de-duplicated like both renderers do
for a repeated heading (`foo`, `foo-1`, `foo-2`, ...)."""
base = _SLUG_DROP_RE.sub("", text.strip().lower())
base = _SLUG_SPACE_RE.sub("-", base).strip("-")
count = seen.get(base, 0)
seen[base] = count + 1
return base if count == 0 else f"{base}-{count}"
def render(headings: list[Heading]) -> str:
"""The TOC region ready to place into a body, or `""` for no headings."""
if not headings:
return ""
seen: dict[str, int] = {}
lines = []
for heading in headings:
indent = " " if heading.level == 3 else ""
slug = _slugify(heading.text, seen)
lines.append(f"{indent}- [{heading.text}](#{slug})")
return blocks.render(REGION_NAME, HEADING_TEXT, lines)
def _first_level2_offset(body: str) -> int | None:
"""Character offset of the first real (non-fenced) `## ` heading, or None."""
masked = markdown_code.strip_code_spans(body)
pos = 0
for masked_line in masked.split("\n"):
if masked_line.startswith("## "):
return pos
pos += len(masked_line) + 1
return None
def upsert(text: str) -> str:
"""`text` with its TOC region created or refreshed, or removed if the
file (measured without any existing region) is at or under the threshold.
Always strips first, then always reinserts fresh - never patches an
existing region in place. `blocks.replace`'s in-place path is right for
the links/footnotes regions it was built for (trailing material, appended
at the end when absent) but wrong here twice over: a TOC has to sit "at
the top" per the guidance this exists to satisfy, not at the end: and
`blocks.strip`'s removal collapses the blank lines around a matched
region to a single newline, which would make patching in place drift
from a fresh insertion's spacing on every second run. Always rebuilding
from the stripped body sidesteps both: the same deterministic
`rstrip("\\n") + "\\n\\n"` join runs whether or not a region existed
before, so a second `upsert` on its own output reproduces that output
exactly.
"""
body = body_without_region(text)
if not needs_toc(text):
return body
region = render(iter_headings(body))
if not region:
return body
offset = _first_level2_offset(body)
if offset is None:
return body.rstrip("\n") + "\n\n" + region + "\n"
head, tail = body[:offset], body[offset:]
return head.rstrip("\n") + "\n\n" + region + "\n\n" + tail
def stale_regions(text: str) -> bool:
"""Whether `text`'s current TOC region (if any) no longer matches what
`upsert` would write - drift `docs verify` can catch mechanically."""
return upsert(text) != text