Files changed: - AGENTS.md - CHANGES.md - README.md - VERSION - docs/ownership-and-templates.md - instructions/evolve-subtypes.md - instructions/setup-instance.md - instructions/subtype-templates.md - instructions/upgrade-instance.md - tools/CONTRACT.md - tools/chemenu/commands/dist_cmd.py - tools/chemenu/commands/docs_verify.py - tools/chemenu/commands/new_page.py - tools/chemenu/tests/test_dist_cmd.py - tools/chemenu/tests/test_docs_verify.py - tools/chemenu/tests/test_new_page.py - tools/chemenu/tests/test_toc.py - tools/chemenu/tests/test_type_resolver.py - tools/chemenu/toc.py - tools/chemenu/type_resolver.py - types/concept.decision.md - types/entity.guidance.md - types/entity.md - types/entity.person.md - types/type-spec.md Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SnAJ7Z3CpVD3PRbN73QtU2
339 lines
15 KiB
Python
339 lines
15 KiB
Python
"""Generate the table-of-contents region for reference files over 100 lines.
|
|
|
|
Anthropic's skill-authoring guidance: "For reference files longer than 100
|
|
lines, include a table of contents at the top. This ensures Claude can see
|
|
the full scope of available information even when previewing with partial
|
|
reads." (`codex-skill-creator/SKILL.md:221-222` names the same threshold as
|
|
the mitigation for the exact preview mechanic `instructions/CONTRACT.md`
|
|
§ "Reference depth" already treats as real for a repo-wide contract reached
|
|
at a second hop.)
|
|
|
|
A hand-maintained TOC is the next drift source the moment a heading changes -
|
|
AGENTS.md invariant 1 ("never hand-edit generated files") applies here the
|
|
same way it applies to a page's links/footnotes region. So this is a third
|
|
generated region beside `xref`'s and `cite`'s, reusing `blocks`' marker
|
|
convention (`<!-- wikitool:toc -->` ... `<!-- /wikitool:toc -->`) without
|
|
joining `blocks.BLOCKS`: that tuple feeds `xref`, `cite` and the KB-page
|
|
`unbalanced_markers` lint check, and none of the target files here (AGENTS.md,
|
|
the stage/collection contracts, the flat `instructions/**.md` files) is a
|
|
`kb/` page. `version.py`'s `bumps` region inside `CHANGES.md` set this same
|
|
precedent first.
|
|
|
|
**Scope is computed, never a hand-picked list** - the same principle that
|
|
governs `wikitool` itself. `target_files()` walks every file-naming category
|
|
AGENTS.md's own table calls agent-loaded: `AGENTS.md`, every stage contract,
|
|
`kb/CONVENTIONS.md`, every `kb/*/COLLECTION.md`, every flat `instructions/**.md`
|
|
file, every `types/*.md` type-spec and guidance file, and every `docs/` page - each together with
|
|
the `<name>.template` it ships as, where one exists. One rule, and
|
|
exactly one exception below it - which is the whole point, because a scope
|
|
carrying several unexplained absences reads as an accident rather than a
|
|
decision, and did: `docs/` and the page type-specs sat outside it for no
|
|
recorded reason at all.
|
|
|
|
`docs/` belongs in for the reason the threshold exists. AGENTS.md § File naming
|
|
calls it "Agents and humans | By link, or on explicit request" - agent-loaded,
|
|
at a second hop, which is precisely the preview mechanic
|
|
`instructions/CONTRACT.md` § "Reference depth" treats as real. The exclusion of
|
|
the human docs (`README.md`, `CHANGES.md`, `EVALS.md`, `INSTALL.md`,
|
|
`tools/README.md`) rests on a sentence that does not stretch to cover it: those
|
|
are "Never loaded by an agent as instruction," while a `docs/` page is not
|
|
loaded *as instruction* but is very much loaded.
|
|
|
|
**The one exception is `SKILL.md`**, and the vendored skill-authoring guidance
|
|
is what puts it there rather than a judgment of ours. `commonplace/kb/work/
|
|
skill-creator-distillation/sources/claude-code-skill-creator/SKILL.md:88-99`
|
|
sets out three loading levels and places the SKILL.md body on the second - "In
|
|
context whenever skill triggers (<500 lines ideal)" - while aiming its own TOC
|
|
advice at the third, "large reference files (>300 lines)", i.e. bundled
|
|
resources. The Codex source agrees by placement: its TOC bullet sits directly
|
|
under "Keep references one level deep from SKILL.md. All reference files should
|
|
link directly from SKILL.md" (`.../codex-skill-creator/SKILL.md:221-222`).
|
|
Neither asks a skill body to carry a table of contents, because neither expects
|
|
one to be previewed. `instructions/CONTRACT.md` § "How much reasoning a step
|
|
may carry" arrives at the same place from the other side, treating a checklist
|
|
read once as the table of contents it replaced. (Our longest skill is 254
|
|
lines, so the <500 guidance costs us nothing either.)
|
|
|
|
A type-spec is loaded whole too - `tools/wikitool types describe` prints its
|
|
entire authoring body - but it is *also* read as a file, by whoever edits it,
|
|
and that is the reading the threshold is about. So it carries a region like any
|
|
other reference file, and `types_cmd` strips the region back out of what
|
|
`describe` prints: the command already hands over the whole body, so a
|
|
navigation aid into it would be noise in the output and nothing else.
|
|
|
|
**A subtype template under `types/` is not a reference file at all**, which is
|
|
why it is out of scope rather than a second exception to the rule above.
|
|
`types/<stem>.<value>.md` is the body skeleton `wikitool new` copies into a page
|
|
(Gitea #117) - a region written into it would be written into every page
|
|
scaffolded from it, as an author-facing section nobody may edit. So no file
|
|
`new` reads as a scaffold ever carries a `wikitool:toc` marker, whatever its
|
|
length; `type_resolver.split_subtype_template_name` decides which files those
|
|
are, the same way for this scope as for `dist export` and `docs verify`.
|
|
|
|
The threshold's own provenance is worth recording, because the two sources
|
|
disagree: Codex says 100 lines, Claude Code says 300. This stack took the
|
|
stricter number.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import re
|
|
from pathlib import Path
|
|
from typing import NamedTuple
|
|
|
|
from chemenu import blocks, config, kb_collections, markdown_code
|
|
from chemenu.type_resolver import split_subtype_template_name
|
|
|
|
REGION_NAME = "toc"
|
|
HEADING_TEXT = "Contents"
|
|
|
|
# The suffix `dist export` re-keys an instance-owned file to, and the one
|
|
# `dist adopt` adopts away again. Lives here because this module is what
|
|
# decides which files are reference material *in both their forms*;
|
|
# `docs_verify` imports it rather than keeping a second spelling. `ownership.py`
|
|
# keeps its own literal deliberately - it answers a different question (which
|
|
# path a release replaces) over a narrower scope.
|
|
TEMPLATE_SUFFIX = ".template"
|
|
|
|
# The line threshold Anthropic's own guidance names. Measured on the body
|
|
# with any existing TOC region stripped out first, so inserting or updating
|
|
# the region can never be what pushes a file over the line.
|
|
THRESHOLD = 100
|
|
|
|
# Mirrors docs_verify.STAGE_CONTRACTS - kept as a separate constant rather than
|
|
# imported, because `docs_verify` importing `toc` (for the new check) would
|
|
# make a `toc` -> `docs_verify` import a cycle. Both lists are the seven stage
|
|
# contracts; `docs_verify.STAGE_CONTRACTS`'s own docstring is the place that
|
|
# explains what they are.
|
|
_STAGE_CONTRACTS = (
|
|
"raw/CONTRACT.md",
|
|
"kb/CONTRACT.md",
|
|
"types/type-spec.md",
|
|
"reports/CONTRACT.md",
|
|
"work/CONTRACT.md",
|
|
"tools/CONTRACT.md",
|
|
"instructions/CONTRACT.md",
|
|
)
|
|
|
|
_HEADING_RE = re.compile(r"^(#{2,3})[ \t]+(.+?)[ \t]*$")
|
|
_SLUG_DROP_RE = re.compile(r"[^\w\s-]", re.UNICODE)
|
|
_SLUG_SPACE_RE = re.compile(r"\s+")
|
|
|
|
|
|
class Heading(NamedTuple):
|
|
level: int # 2 or 3
|
|
text: str
|
|
|
|
|
|
def target_files() -> list[Path]:
|
|
"""Every file the TOC region applies to, root-relative, sorted.
|
|
|
|
Walking `instructions/` picks up `instructions/dev/` and
|
|
`instructions/migrations/` along with the rest: both are still the flat
|
|
`instructions/<name>.md` form (`instructions/CONTRACT.md` says the `dev/`
|
|
split is orthogonal to Linked/Manual, not a different file shape), so
|
|
there is no separate rule for them to fall out of - and no hand-picking
|
|
for a future file under either to be missed. `types/` and `docs/` are
|
|
walked for the same reason: a type or a rationale page added later is in
|
|
scope by construction, not by someone remembering this function.
|
|
|
|
`types/type-spec.md` arrives twice - once as a stage contract, once from
|
|
the `types/` walk - and the set at the bottom is what makes that a
|
|
non-issue rather than something to special-case.
|
|
|
|
**A file in scope carries its `.template` in with it.** An instance-owned
|
|
reference file crosses the distribution boundary as `<name>.template` and
|
|
is adopted by copying it back (`instructions/setup-instance.md`), so the
|
|
template is the same document one step earlier in its life - the shipped
|
|
form of a file this scope already covers. Leaving it out meant nothing
|
|
maintained it and nothing checked it: `kb/CONVENTIONS.md.template` grew
|
|
past the threshold carrying no region at all, and the first thing to notice
|
|
was a fresh instance failing `docs verify` at the end of its own setup,
|
|
because the adopted copy inherited the gap. `docs toc --apply` writes the
|
|
region into the template as into anything else here; the instance's own
|
|
later edits stay its business, and its own `docs toc` run answers for them.
|
|
"""
|
|
files: list[Path] = [config.ROOT / "AGENTS.md"]
|
|
files += [config.ROOT / rel for rel in _STAGE_CONTRACTS]
|
|
files.append(config.ROOT / "kb" / "CONVENTIONS.md")
|
|
files += sorted(
|
|
collection / kb_collections.CONTRACT_NAME
|
|
for collection in kb_collections.iter_kb_collections()
|
|
)
|
|
instructions_dir = config.ROOT / "instructions"
|
|
if instructions_dir.is_dir():
|
|
files += sorted(
|
|
path
|
|
for path in instructions_dir.rglob("*.md")
|
|
if path.name != "SKILL.md"
|
|
)
|
|
for subdir in ("types", "docs"):
|
|
directory = config.ROOT / subdir
|
|
if directory.is_dir():
|
|
files += sorted(
|
|
path
|
|
for path in directory.rglob("*.md")
|
|
if subdir != "types" or split_subtype_template_name(path.name) is None
|
|
)
|
|
files += [
|
|
template
|
|
for template in (
|
|
path.with_name(path.name + TEMPLATE_SUFFIX) for path in tuple(files)
|
|
)
|
|
if template.is_file()
|
|
]
|
|
return sorted({f for f in files if f.is_file()})
|
|
|
|
|
|
def body_without_region(text: str) -> str:
|
|
"""`text` with any existing TOC region removed - what the 100-line
|
|
threshold is measured against, so a stale TOC never counts toward keeping
|
|
itself around.
|
|
|
|
Trailing newlines are canonicalized to exactly one (`blocks.strip`'s own
|
|
removal path already does this when a region existed; a file with no
|
|
region yet is normalized here too), which is what keeps `upsert`
|
|
idempotent - without it, a file missing its final newline would insert
|
|
its first TOC one way and every later re-run one byte shorter, since the
|
|
"remove an existing region" and "there was never one" paths would
|
|
otherwise start from differently-terminated strings.
|
|
"""
|
|
stripped = blocks.strip(text, REGION_NAME)
|
|
return stripped.rstrip("\n") + "\n" if stripped else stripped
|
|
|
|
|
|
_REGION_WITH_PADDING_RE = re.compile(
|
|
r"\n*"
|
|
+ re.escape(blocks.open_marker(REGION_NAME))
|
|
+ r".*?"
|
|
+ re.escape(blocks.close_marker(REGION_NAME))
|
|
+ r"\n*",
|
|
re.DOTALL,
|
|
)
|
|
|
|
|
|
def strip_region(text: str) -> str:
|
|
"""`text` with the generated TOC region removed, for a caller that is
|
|
handing the whole body over anyway.
|
|
|
|
`types describe` is the one such caller: it prints a type-spec's entire
|
|
authoring body, so the region's markers and heading list would be noise in
|
|
its output rather than a way into anything. Distinct from
|
|
`body_without_region`, which exists to *measure* a body against the line
|
|
threshold and therefore also canonicalizes the trailing newline.
|
|
|
|
Takes the surrounding blank lines with it and puts one back, rather than
|
|
calling `blocks.strip`: that collapses the padding to a single newline,
|
|
which is the right answer when the result is about to be rebuilt from
|
|
scratch (`upsert` does exactly that) and the wrong one here, where the
|
|
output is printed as-is - it would leave the following `##` heading welded
|
|
to the paragraph above it.
|
|
"""
|
|
return _REGION_WITH_PADDING_RE.sub("\n\n", text)
|
|
|
|
|
|
def needs_toc(text: str) -> bool:
|
|
return len(body_without_region(text).splitlines()) > THRESHOLD
|
|
|
|
|
|
def iter_headings(body: str) -> list[Heading]:
|
|
"""Every `##`/`###` heading in document order, code-aware.
|
|
|
|
Detection runs on `markdown_code.strip_code_spans(body)`, which blanks
|
|
fenced blocks and inline code spans but preserves line structure - a
|
|
template line inside a fence (`# {Type name}`, a `raw/CONTRACT.md`
|
|
example path prefixed `#`) no longer starts with a real `#` marker once
|
|
masked, so it is not mistaken for a section. The heading *text* is read
|
|
back from the corresponding *unmasked* line, not the masked one: masking
|
|
also blanks inline code spans, which would otherwise turn a heading like
|
|
`` ## `instructions/dev/` `` into blank spaces instead of its real text.
|
|
"""
|
|
original = body.split("\n")
|
|
masked = markdown_code.strip_code_spans(body).split("\n")
|
|
headings: list[Heading] = []
|
|
for masked_line, original_line in zip(masked, original):
|
|
if not masked_line.startswith("#"):
|
|
continue
|
|
match = _HEADING_RE.match(masked_line)
|
|
if not match:
|
|
continue
|
|
# The masked line only still starts with `#...` for a genuine heading
|
|
# (fenced lines are blanked to spaces); recover the text from the
|
|
# unmasked line at the same position.
|
|
original_match = _HEADING_RE.match(original_line)
|
|
if not original_match:
|
|
continue
|
|
headings.append(Heading(level=len(match.group(1)), text=original_match.group(2).strip()))
|
|
return headings
|
|
|
|
|
|
def _slugify(text: str, seen: dict[str, int]) -> str:
|
|
"""A GitHub/Gitea-style anchor slug, de-duplicated like both renderers do
|
|
for a repeated heading (`foo`, `foo-1`, `foo-2`, ...)."""
|
|
base = _SLUG_DROP_RE.sub("", text.strip().lower())
|
|
base = _SLUG_SPACE_RE.sub("-", base).strip("-")
|
|
count = seen.get(base, 0)
|
|
seen[base] = count + 1
|
|
return base if count == 0 else f"{base}-{count}"
|
|
|
|
|
|
def render(headings: list[Heading]) -> str:
|
|
"""The TOC region ready to place into a body, or `""` for no headings."""
|
|
if not headings:
|
|
return ""
|
|
seen: dict[str, int] = {}
|
|
lines = []
|
|
for heading in headings:
|
|
indent = " " if heading.level == 3 else ""
|
|
slug = _slugify(heading.text, seen)
|
|
lines.append(f"{indent}- [{heading.text}](#{slug})")
|
|
return blocks.render(REGION_NAME, HEADING_TEXT, lines)
|
|
|
|
|
|
def _first_level2_offset(body: str) -> int | None:
|
|
"""Character offset of the first real (non-fenced) `## ` heading, or None."""
|
|
masked = markdown_code.strip_code_spans(body)
|
|
pos = 0
|
|
for masked_line in masked.split("\n"):
|
|
if masked_line.startswith("## "):
|
|
return pos
|
|
pos += len(masked_line) + 1
|
|
return None
|
|
|
|
|
|
def upsert(text: str) -> str:
|
|
"""`text` with its TOC region created or refreshed, or removed if the
|
|
file (measured without any existing region) is at or under the threshold.
|
|
|
|
Always strips first, then always reinserts fresh - never patches an
|
|
existing region in place. `blocks.replace`'s in-place path is right for
|
|
the links/footnotes regions it was built for (trailing material, appended
|
|
at the end when absent) but wrong here twice over: a TOC has to sit "at
|
|
the top" per the guidance this exists to satisfy, not at the end: and
|
|
`blocks.strip`'s removal collapses the blank lines around a matched
|
|
region to a single newline, which would make patching in place drift
|
|
from a fresh insertion's spacing on every second run. Always rebuilding
|
|
from the stripped body sidesteps both: the same deterministic
|
|
`rstrip("\\n") + "\\n\\n"` join runs whether or not a region existed
|
|
before, so a second `upsert` on its own output reproduces that output
|
|
exactly.
|
|
"""
|
|
body = body_without_region(text)
|
|
if not needs_toc(text):
|
|
return body
|
|
|
|
region = render(iter_headings(body))
|
|
if not region:
|
|
return body
|
|
|
|
offset = _first_level2_offset(body)
|
|
if offset is None:
|
|
return body.rstrip("\n") + "\n\n" + region + "\n"
|
|
head, tail = body[:offset], body[offset:]
|
|
return head.rstrip("\n") + "\n\n" + region + "\n\n" + tail
|
|
|
|
|
|
def stale_regions(text: str) -> bool:
|
|
"""Whether `text`'s current TOC region (if any) no longer matches what
|
|
`upsert` would write - drift `docs verify` can catch mechanically."""
|
|
return upsert(text) != text
|