# Why the pipeline has four stages Chemenu could, in principle, be one directory: drop a file in, ask a question, get an answer computed fresh each time. It isn't built that way. The pipeline in [AGENTS.md](../AGENTS.md#routing) - `raw/` -> `[types/ + tools/]` -> `kb/` -> `reports/`, with `work/` alongside rather than inside it - separates *material* from *meaning* from *byproduct*, and each seam exists because collapsing it costs something specific. ## Why raw material stays untouched [raw/CONTRACT.md](../raw/CONTRACT.md) keeps a source exactly as it arrived. The reasoning is simple once stated: the moment someone "cleans up" or reformats a source on the way in, the thing later claims get checked against is no longer the thing that was actually said. An immutable `raw/` means a citation always resolves to the original, not to somebody's tidied memory of it. It also draws a trust boundary in one place instead of scattering it - everything past `raw/` can be treated as reviewed, because nothing upstream of it silently already was. `incoming/` (Gitea #58) sits entirely on the near side of that boundary: a file waiting there is not yet reviewed and not yet a citation target, so its being gitignored and readable by an ingest session does not weaken anything - the boundary is the promotion into `raw/` itself, not the moment a human happened to drop a file somewhere. ## Why extraction happens once, through a schema [types/type-spec.md](../types/type-spec.md) is what stands between a raw file and a `kb/` page: a type-spec defines what a conforming instance of a page looks like, and the compiler (`tools/wikitool`) applies it. The alternative - every query re-reading and re-interpreting the source on demand - would mean paying the cost of understanding the material every single time, and getting a slightly different answer each time depending on how the question was phrased. Extracting once, against a fixed schema, turns "re-read and re-guess" into "look up what was already compiled." That is the "never re-derive, always compile" principle from [AGENTS.md](../AGENTS.md): understanding a source is expensive and worth doing exactly once, after which it becomes a cheap, stable lookup. ## Why a `kb/` page has to stand on its own [kb/CONTRACT.md](../kb/CONTRACT.md) sets the bar for the compiled layer: a page should answer a future question without sending the reader back to the source it came from. That's the payoff of compiling in the first place - if every answer still bottomed out in "go re-read the raw file," the `kb/` layer would just be a pointer with extra steps, and the cost of extraction would have bought nothing. A page that stands alone is what makes the corpus fast and consistent to query: the work of understanding is already sitting there, done. ## Why `reports/` doesn't need to be maintained [reports/CONTRACT.md](../reports/CONTRACT.md) treats most of what lands in `reports/` - lint output, telemetry traces - as disposable. The structural content of a lint report can be recomputed from the tree at any commit, so keeping an old copy around would just be a second version of something the tool can already answer on demand, and a second copy is exactly the kind of thing that quietly goes stale. Treating it as derived output rather than a fourth thing to maintain means there is nothing there to fall out of sync - regenerating it is cheaper than reconciling it. The one part that genuinely can't be recomputed - the judgment a pass produced - is carried out into `kb/` or `kb/log.md` before the report itself is discarded, which is the distinction between what's recomputable and what isn't. ## Where `work/` fits [work/CONTRACT.md](../work/CONTRACT.md) describes a workshop, not a fifth pipeline stage: a place for the notes, extracts and open decisions of a task that spans more than one session, on its way toward becoming a `kb/` page. It sits beside the raw -> kb -> reports flow rather than inside it - closer in spirit to a desk than to a conveyor belt. ## The shape this produces Four stages, each answering a different question: `raw/` - what was actually said; `types/` + `tools/` - how to turn that into structured understanding; `kb/` - what is now known; `reports/` - what a pass over the corpus noticed in passing. Keeping them separate is what lets each one be trusted for what it is, instead of every layer having to double as all four at once.