"""Generate the table-of-contents region for reference files over 100 lines. Anthropic's skill-authoring guidance: "For reference files longer than 100 lines, include a table of contents at the top. This ensures Claude can see the full scope of available information even when previewing with partial reads." (`codex-skill-creator/SKILL.md:221-222` names the same threshold as the mitigation for the exact preview mechanic `instructions/CONTRACT.md` § "Reference depth" already treats as real for a repo-wide contract reached at a second hop.) A hand-maintained TOC is the next drift source the moment a heading changes - AGENTS.md invariant 1 ("never hand-edit generated files") applies here the same way it applies to a page's links/footnotes region. So this is a third generated region beside `xref`'s and `cite`'s, reusing `blocks`' marker convention (`` ... ``) without joining `blocks.BLOCKS`: that tuple feeds `xref`, `cite` and the KB-page `unbalanced_markers` lint check, and none of the target files here (AGENTS.md, the stage/collection contracts, the flat `instructions/**.md` files) is a `kb/` page. `version.py`'s `bumps` region inside `CHANGES.md` set this same precedent first. **Scope is computed, never a hand-picked list** - the same principle that governs `wikitool` itself. `target_files()` walks every file-naming category AGENTS.md's own table calls agent-loaded: `AGENTS.md`, every stage contract, `kb/CONVENTIONS.md`, every `kb/*/COLLECTION.md`, every flat `instructions/**.md` file, every `types/*.md` type-spec and guidance file, and every `docs/` page - each together with the `.template` it ships as, where one exists. One rule, and exactly one exception below it - which is the whole point, because a scope carrying several unexplained absences reads as an accident rather than a decision, and did: `docs/` and the page type-specs sat outside it for no recorded reason at all. `docs/` belongs in for the reason the threshold exists. AGENTS.md § File naming calls it "Agents and humans | By link, or on explicit request" - agent-loaded, at a second hop, which is precisely the preview mechanic `instructions/CONTRACT.md` § "Reference depth" treats as real. The exclusion of the human docs (`README.md`, `CHANGES.md`, `EVALS.md`, `INSTALL.md`, `tools/README.md`) rests on a sentence that does not stretch to cover it: those are "Never loaded by an agent as instruction," while a `docs/` page is not loaded *as instruction* but is very much loaded. **The one exception is `SKILL.md`**, and the vendored skill-authoring guidance is what puts it there rather than a judgment of ours. `commonplace/kb/work/ skill-creator-distillation/sources/claude-code-skill-creator/SKILL.md:88-99` sets out three loading levels and places the SKILL.md body on the second - "In context whenever skill triggers (<500 lines ideal)" - while aiming its own TOC advice at the third, "large reference files (>300 lines)", i.e. bundled resources. The Codex source agrees by placement: its TOC bullet sits directly under "Keep references one level deep from SKILL.md. All reference files should link directly from SKILL.md" (`.../codex-skill-creator/SKILL.md:221-222`). Neither asks a skill body to carry a table of contents, because neither expects one to be previewed. `instructions/CONTRACT.md` § "How much reasoning a step may carry" arrives at the same place from the other side, treating a checklist read once as the table of contents it replaced. (Our longest skill is 254 lines, so the <500 guidance costs us nothing either.) A type-spec is loaded whole too - `tools/wikitool types describe` prints its entire authoring body - but it is *also* read as a file, by whoever edits it, and that is the reading the threshold is about. So it carries a region like any other reference file, and `types_cmd` strips the region back out of what `describe` prints: the command already hands over the whole body, so a navigation aid into it would be noise in the output and nothing else. **A subtype template under `types/` is not a reference file at all**, which is why it is out of scope rather than a second exception to the rule above. `types/..md` is the body skeleton `wikitool new` copies into a page (Gitea #117) - a region written into it would be written into every page scaffolded from it, as an author-facing section nobody may edit. So no file `new` reads as a scaffold ever carries a `wikitool:toc` marker, whatever its length; `type_resolver.split_subtype_template_name` decides which files those are, the same way for this scope as for `dist export` and `docs verify`. The threshold's own provenance is worth recording, because the two sources disagree: Codex says 100 lines, Claude Code says 300. This stack took the stricter number. """ from __future__ import annotations import re from pathlib import Path from typing import NamedTuple from chemenu import blocks, config, kb_collections, markdown_code from chemenu.type_resolver import split_subtype_template_name REGION_NAME = "toc" HEADING_TEXT = "Contents" # The suffix `dist export` re-keys an instance-owned file to, and the one # `dist adopt` adopts away again. Lives here because this module is what # decides which files are reference material *in both their forms*; # `docs_verify` imports it rather than keeping a second spelling. `ownership.py` # keeps its own literal deliberately - it answers a different question (which # path a release replaces) over a narrower scope. TEMPLATE_SUFFIX = ".template" # The line threshold Anthropic's own guidance names. Measured on the body # with any existing TOC region stripped out first, so inserting or updating # the region can never be what pushes a file over the line. THRESHOLD = 100 # Mirrors docs_verify.STAGE_CONTRACTS - kept as a separate constant rather than # imported, because `docs_verify` importing `toc` (for the new check) would # make a `toc` -> `docs_verify` import a cycle. Both lists are the seven stage # contracts; `docs_verify.STAGE_CONTRACTS`'s own docstring is the place that # explains what they are. _STAGE_CONTRACTS = ( "raw/CONTRACT.md", "kb/CONTRACT.md", "types/type-spec.md", "reports/CONTRACT.md", "work/CONTRACT.md", "tools/CONTRACT.md", "instructions/CONTRACT.md", ) _HEADING_RE = re.compile(r"^(#{2,3})[ \t]+(.+?)[ \t]*$") _SLUG_DROP_RE = re.compile(r"[^\w\s-]", re.UNICODE) _SLUG_SPACE_RE = re.compile(r"\s+") class Heading(NamedTuple): level: int # 2 or 3 text: str def target_files() -> list[Path]: """Every file the TOC region applies to, root-relative, sorted. Walking `instructions/` picks up `instructions/dev/` and `instructions/migrations/` along with the rest: both are still the flat `instructions/.md` form (`instructions/CONTRACT.md` says the `dev/` split is orthogonal to Linked/Manual, not a different file shape), so there is no separate rule for them to fall out of - and no hand-picking for a future file under either to be missed. `types/` and `docs/` are walked for the same reason: a type or a rationale page added later is in scope by construction, not by someone remembering this function. `types/type-spec.md` arrives twice - once as a stage contract, once from the `types/` walk - and the set at the bottom is what makes that a non-issue rather than something to special-case. **A file in scope carries its `.template` in with it.** An instance-owned reference file crosses the distribution boundary as `.template` and is adopted by copying it back (`instructions/setup-instance.md`), so the template is the same document one step earlier in its life - the shipped form of a file this scope already covers. Leaving it out meant nothing maintained it and nothing checked it: `kb/CONVENTIONS.md.template` grew past the threshold carrying no region at all, and the first thing to notice was a fresh instance failing `docs verify` at the end of its own setup, because the adopted copy inherited the gap. `docs toc --apply` writes the region into the template as into anything else here; the instance's own later edits stay its business, and its own `docs toc` run answers for them. """ files: list[Path] = [config.ROOT / "AGENTS.md"] files += [config.ROOT / rel for rel in _STAGE_CONTRACTS] files.append(config.ROOT / "kb" / "CONVENTIONS.md") files += sorted( collection / kb_collections.CONTRACT_NAME for collection in kb_collections.iter_kb_collections() ) instructions_dir = config.ROOT / "instructions" if instructions_dir.is_dir(): files += sorted( path for path in instructions_dir.rglob("*.md") if path.name != "SKILL.md" ) for subdir in ("types", "docs"): directory = config.ROOT / subdir if directory.is_dir(): files += sorted( path for path in directory.rglob("*.md") if subdir != "types" or split_subtype_template_name(path.name) is None ) files += [ template for template in ( path.with_name(path.name + TEMPLATE_SUFFIX) for path in tuple(files) ) if template.is_file() ] return sorted({f for f in files if f.is_file()}) def body_without_region(text: str) -> str: """`text` with any existing TOC region removed - what the 100-line threshold is measured against, so a stale TOC never counts toward keeping itself around. Trailing newlines are canonicalized to exactly one (`blocks.strip`'s own removal path already does this when a region existed; a file with no region yet is normalized here too), which is what keeps `upsert` idempotent - without it, a file missing its final newline would insert its first TOC one way and every later re-run one byte shorter, since the "remove an existing region" and "there was never one" paths would otherwise start from differently-terminated strings. """ stripped = blocks.strip(text, REGION_NAME) return stripped.rstrip("\n") + "\n" if stripped else stripped _REGION_WITH_PADDING_RE = re.compile( r"\n*" + re.escape(blocks.open_marker(REGION_NAME)) + r".*?" + re.escape(blocks.close_marker(REGION_NAME)) + r"\n*", re.DOTALL, ) def strip_region(text: str) -> str: """`text` with the generated TOC region removed, for a caller that is handing the whole body over anyway. `types describe` is the one such caller: it prints a type-spec's entire authoring body, so the region's markers and heading list would be noise in its output rather than a way into anything. Distinct from `body_without_region`, which exists to *measure* a body against the line threshold and therefore also canonicalizes the trailing newline. Takes the surrounding blank lines with it and puts one back, rather than calling `blocks.strip`: that collapses the padding to a single newline, which is the right answer when the result is about to be rebuilt from scratch (`upsert` does exactly that) and the wrong one here, where the output is printed as-is - it would leave the following `##` heading welded to the paragraph above it. """ return _REGION_WITH_PADDING_RE.sub("\n\n", text) def needs_toc(text: str) -> bool: return len(body_without_region(text).splitlines()) > THRESHOLD def iter_headings(body: str) -> list[Heading]: """Every `##`/`###` heading in document order, code-aware. Detection runs on `markdown_code.strip_code_spans(body)`, which blanks fenced blocks and inline code spans but preserves line structure - a template line inside a fence (`# {Type name}`, a `raw/CONTRACT.md` example path prefixed `#`) no longer starts with a real `#` marker once masked, so it is not mistaken for a section. The heading *text* is read back from the corresponding *unmasked* line, not the masked one: masking also blanks inline code spans, which would otherwise turn a heading like `` ## `instructions/dev/` `` into blank spaces instead of its real text. """ original = body.split("\n") masked = markdown_code.strip_code_spans(body).split("\n") headings: list[Heading] = [] for masked_line, original_line in zip(masked, original): if not masked_line.startswith("#"): continue match = _HEADING_RE.match(masked_line) if not match: continue # The masked line only still starts with `#...` for a genuine heading # (fenced lines are blanked to spaces); recover the text from the # unmasked line at the same position. original_match = _HEADING_RE.match(original_line) if not original_match: continue headings.append(Heading(level=len(match.group(1)), text=original_match.group(2).strip())) return headings def _slugify(text: str, seen: dict[str, int]) -> str: """A GitHub/Gitea-style anchor slug, de-duplicated like both renderers do for a repeated heading (`foo`, `foo-1`, `foo-2`, ...).""" base = _SLUG_DROP_RE.sub("", text.strip().lower()) base = _SLUG_SPACE_RE.sub("-", base).strip("-") count = seen.get(base, 0) seen[base] = count + 1 return base if count == 0 else f"{base}-{count}" def render(headings: list[Heading]) -> str: """The TOC region ready to place into a body, or `""` for no headings.""" if not headings: return "" seen: dict[str, int] = {} lines = [] for heading in headings: indent = " " if heading.level == 3 else "" slug = _slugify(heading.text, seen) lines.append(f"{indent}- [{heading.text}](#{slug})") return blocks.render(REGION_NAME, HEADING_TEXT, lines) def _first_level2_offset(body: str) -> int | None: """Character offset of the first real (non-fenced) `## ` heading, or None.""" masked = markdown_code.strip_code_spans(body) pos = 0 for masked_line in masked.split("\n"): if masked_line.startswith("## "): return pos pos += len(masked_line) + 1 return None def upsert(text: str) -> str: """`text` with its TOC region created or refreshed, or removed if the file (measured without any existing region) is at or under the threshold. Always strips first, then always reinserts fresh - never patches an existing region in place. `blocks.replace`'s in-place path is right for the links/footnotes regions it was built for (trailing material, appended at the end when absent) but wrong here twice over: a TOC has to sit "at the top" per the guidance this exists to satisfy, not at the end: and `blocks.strip`'s removal collapses the blank lines around a matched region to a single newline, which would make patching in place drift from a fresh insertion's spacing on every second run. Always rebuilding from the stripped body sidesteps both: the same deterministic `rstrip("\\n") + "\\n\\n"` join runs whether or not a region existed before, so a second `upsert` on its own output reproduces that output exactly. """ body = body_without_region(text) if not needs_toc(text): return body region = render(iter_headings(body)) if not region: return body offset = _first_level2_offset(body) if offset is None: return body.rstrip("\n") + "\n\n" + region + "\n" head, tail = body[:offset], body[offset:] return head.rstrip("\n") + "\n\n" + region + "\n\n" + tail def stale_regions(text: str) -> bool: """Whether `text`'s current TOC region (if any) no longer matches what `upsert` would write - drift `docs verify` can catch mechanically.""" return upsert(text) != text