stack: wikitool review - der Wochenrueckblick als Join zur Lesezeit (#125)
CI / verify (push) Successful in 48s
Release / release (push) Successful in 36s

Files changed:
- CHANGES.md
- VERSION
- tools/CONTRACT.md
- tools/chemenu/cli.py
- tools/chemenu/commands/review_cmd.py
- tools/chemenu/commands/run_budget.py
- tools/chemenu/review.py
- tools/chemenu/tests/test_review.py
This commit is contained in:
torben committed 2026-09-19 22:09:18 +02:00
1 parent 1875449b31
commit 80b57e0d01
8 files changed
+751 -5

No files matched your search

+2
View File
@@ -125,6 +125,7 @@ tools/wikitool <command> --help
|---------|---------|
| `lint [--json] [--markdown out.md] [--full] [--fail-on-error]` | Structural + provenance checks: broken wikilinks, dangling frontmatter references, orphan pages, index drift, schema gaps, duplicate titles, title mismatches, pages nested more than one directory below their collection (hard - the generated catalog folds these into their area silently rather than merely reading it), uncovered raw files, broken `raw_files:` refs, raw files claimed by more than one source page, unmarked provenance, citation/frontmatter drift, unbalanced generated-region markers, edges whose label is missing or not authorised by the source collection's `outbound:` (both hard once `kb_version` has reached the release that introduced labelled edges - advisory below it, so a corpus mid-migration is not refused by the check measuring it), `see-also` edges whose reverse direction already carries a specific label (advisory only - redundant rather than wrong, and never migration-gated, since no version turns the redundancy into an error), a collection past the catalog's per-area shard threshold that has no areas to shard (advisory only - sharding is automatic but per *area*, so a collection nobody gave areas keeps one table however large it grows; reported with the split its subtype field would produce, and only when that split puts every resulting area at or under the threshold, so a lopsided or small collection stays silent), source pages sitting in the `unclassified` catalog slot (advisory only - `unclassified` is the visible fallback for a genuinely unclear source, not a defect), quote-limit overages (>2 blockquoted lines/page, advisory only). Prints only the sections that found something and always writes the full report to `reports/Lint Report <date>.md` (or `--markdown`), naming the path - `--full` prints everything, `--json` prints the findings and writes nothing |
| `search ["<text>"] [--field <predicate> ...] [--kind/--subtype/--collection/--tag <v>] [--regex] [--limit N] [--sort [-]<field>] [--backend <name>] [--matches] [--json]` | Find pages in `kb/` without reading the index. Text search runs through a pluggable backend (`rg` today); `--field` predicates are evaluated on frontmatter - `f=v`, `f~substring`, `'f>=v'`, `'f:*'` (present), `'!f'` (absent), repeatable and ANDed. With no text this is a pure structured query. One hit per line, ` | `-separated as `score \| kind/subtype \| title \| path \| summary`, so a hit can be judged without opening the page and then opened without looking it up: **title and path are never truncated** (the title is the identifier `touch`/`xref`/`cite` take), and the summary - the one lossy field, and the only one that may contain the separator - goes last, so splitting on `" \| "` with `maxsplit=4` is unambiguous. Scope is pages: the backend walks `kb/` but drops anything `kb_scan.iter_kb_pages` excludes (the kb-root meta files, every `COLLECTION.md`, every generated `INDEX.md`), which is why a hand-run grep over `kb/` can add none of them but those. `--limit` defaults to 50 (`0` for no limit) and **a truncated result says so** - `50 of 182 result(s)` in the table, `total`/`truncated`/`limit` beside `count` in `--json`, where `count` stays the number of results in the payload; the same default and the same fields are what `api.search` and the MCP `search` tool carry, from one constant. A page whose frontmatter does not parse can match no positive predicate, so it is **named** rather than dropped: `--json` always carries an `unreadable` list of `{path, reason}` (usually empty), and the table form writes the same lines to stderr. `--regex` is applied by `rg` alone, whose engine is linear; the ranking boosts for title and summary are literal-containment only, so a non-literal pattern is ranked by match count. `rg` is killed after 30 s and reported as a failure. Read-only, and **exempt from the Iteration Budget Gate** |
| `review [--json]` | The GTD weekly review: joins the configured task-tracker provider (`chemenu.tasks`) against `kb/gtd/` project pages over the case-normalized project name, at read time, storing nothing - not even a `reports/` file. Five checks: **stalled** (a tracker project with zero open items whose `kb/` page is `state: active` - `dormant`/`completed`/`abandoned` never fire, since those states mean the initiative not having a next action is expected rather than a problem), **waiting-overdue** (a `WAITING` item whose `follow_up_at` is older than `thresholds.stalled_waiting_days`), **unpaged-project** (a tracker project with no matching `kb/` page, older than `thresholds.unpaged_project_weeks`), **no-open-loop** (a `kb/` page `state: active` with no matching tracker project, or one with zero open items - the reverse direction of the unpaged-project join, so a rename on either side surfaces on both), **someday-stale** (a someday/maybe item untouched for longer than `thresholds.someday_stale_months`). Thresholds come from `.wikitool-tasks.json`, never from the schema. Text output is one `[check] project: message` line per finding; `--json` carries the same findings plus `checks_run`/`checks_skipped`/`kb_project_count`/`complete`. No `.wikitool-tasks.json` fails immediately with a clear "no tracker configured" message; a provider that cannot be reached mid-run degrades only the checks that needed the failing call, and the report is never rendered as if it were complete - see its error-contract row. Read-only, and **exempt from the Iteration Budget Gate** |
### Provenance
@@ -339,6 +340,7 @@ is atomic, and whether a retry is safe.
|---------|--------------|---------|--------------|
| `lint` | Only with `--fail-on-error`: hard findings exist | Writes one report file (single atomic write) unless `--json` | Safe to retry freely, but re-run it to re-*measure*, never to re-read: the printed path holds the full report. Exit 1 means "act on the findings", not "the tool is broken" |
| `search` | `rg` is not installed or did not finish within 30 s, a malformed `--field` predicate, an unknown field name, or an unknown `--backend` | Read-only | Fix the argument and retry. A timeout is a pathological pattern or an unresponsive corpus directory, not a slow answer - narrow the query or drop `--regex` rather than retrying it unchanged. An unknown field name is reported with the list of fields that do exist - it is never answered with an empty result, because that would read as "no such pages" |
| `review` | Either no `.wikitool-tasks.json` (or a malformed one) - not yours to fix by retrying unchanged, configure or repair it first - **or** the provider was reachable at config-parse time but a read call failed mid-run, in which case the full report (findings plus which checks ran) is printed first and exit 1 follows, never a silent partial success | Read-only | The two exit-1 causes above need different responses: a config problem needs editing `.wikitool-tasks.json`; an unreachable provider (e.g. the tracker app not running) needs starting it, then a plain retry - the command re-reads everything fresh each time, so nothing here is ever stale to re-fetch |
### Provenance