--- type: types/type-spec.md name: source description: Structured type for source pages that record and summarize ingested raw material schema: types/source.schema.yaml subtype_field: source_type base_dir: sources title_prefix: "Source - " page_ref_fields: [entities, concepts] capture_fields: [fidelity, authority] layout: transcript: {dir: transcripts, title: Transkripte} analysis: {dir: analyses, title: Analysen} article: {dir: articles, title: Artikel} document: {dir: documents, title: Dokumente} notes: {dir: notes, title: Notizen} tracker: {dir: trackers, title: Tracker-Exporte} unclassified: {dir: unclassified, title: Unklassifiziert} --- # Source `source` is the type for pages that summarize and catalogue ingested raw material. Source pages are the bridge between the `raw/` layer (immutable source files) and the `kb/` layer (compiled knowledge). One source page stands for **one logical source**, which may span several raw files. ## Contents - [When to use](#when-to-use) - [When NOT to use](#when-not-to-use) - [Frontmatter](#frontmatter) - [Authoring guidance](#authoring-guidance) - [Not Extracted](#not-extracted) - [Template](#template) ## When to use - Summarizing a single external article, document or specification - Recording several related notes or meeting records as one source - Documenting an ingested PDF, manual or other document - Capturing information about an image or a diagram ## When NOT to use - For compiled knowledge (use `entity` or `concept`) - For comparative analyses (use `comparison`) - For original wiki content not derived from raw material ## Frontmatter | Field | Required | Use | |---|---:|---| | `type` | Yes | `types/source.md` | | `source_type` | Yes | One of: transcript, analysis, article, document, notes, tracker, unclassified - no default, see below | | `author` | Yes | Originator of the source material | | `raw_files` | Yes | Raw file paths this source **owns** - see "One raw file, one owner" below | | `fidelity` | Yes (in the tool, not in the schema) | How faithful the *capture* is: `verbatim`, `published`, `secondhand`, `nontextual` - a capture field, see below | | `authority` | Yes (in the tool, not in the schema) | What the material may claim about its *subject*: `normative`, `reporting`, `opinion` - a capture field, see below | | `source_url` | No | Origin URL for external sources | | `source_language` | No | ISO 639-1 code of the raw material's language, e.g. `de`, `en`, `fr` | | `date` | Yes | Publication or creation date (YYYY-MM-DD) | | `tags` | No | Navigation tags for categorization | | `entities` | No | Titles of the entities mentioned in this source | | `concepts` | No | Titles of the concepts mentioned in this source | | `summary` | Yes | One-liner for `kb/index.md` | ## Authoring guidance - The title starts with "Source - ", followed by the name of the source - `source_type` has **no default** - `wikitool new source` refuses without an explicit value. Where the category is unclear, set `unclassified` rather than guessing; that is a visible catalog slot with an advisory `lint` finding, not a dumping ground. What separates `analysis` from `document`, and which area holds which value: [kb/sources/COLLECTION.md](../kb/sources/COLLECTION.md) - `raw_files` lists every raw file this source covers (one source page per logical source, not per file) - `fidelity` and `authority` are **capture fields** (`capture_fields:` above): recorded at the drop point and not freely changeable afterwards. `wikitool raw accept --fidelity --authority ` refuses without both; without `--page` it instead prints the ready-made `new source --set fidelity=... --set authority=...` line, which `new source` in turn refuses without both values. `wikitool touch --set fidelity=` only writes while the field is absent - where a value already stands, it refuses and points at `raw accept --replaces` as the one correction path (a corrected capture is a new edition, not an edit). `unknown` is backfill-only: only `wikitool touch` may write it, never `raw accept` or `new source` - the same construction as `source_language` uses for pages that predate this rule. What the values mean and how they differ: [raw/CONTRACT.md](../raw/CONTRACT.md#getting-a-file-in-incoming) - For external articles, always set `source_url` to the origin URL - Set `source_language` to the raw material's language, not the page's - The page is written in the KB language, whatever language the source is in; verbatim passages are quoted in the original (`kb/CONVENTIONS.md` § "Language") - Summarize the key claims in the summary section - Put anything actionable in the action items section - Put deliberate omissions in the not-extracted section - see below - Link the entities and concepts mentioned under related entities/concepts ## Not Extracted The decision that material should *not* be taken over cannot be reconstructed: nothing in the repository can re-derive it, and `sources coverage` only knows whether a raw file is claimed by some source page - never whether anyone decided about its contents. Left unwritten, the same source is renegotiated on every later pass. - Record every deliberate omission with a **reason**, not just a filename. - Mandatory where the ingest ran through `instructions/ingest-large-tree.md` - on both axes: for a tree ingest the section records what was not taken from the tree; for a thematically broad single source, which named subjects got no page of their own, and why. Optional for a single small file - but an empty section still beats a missing one. - Belongs on the source page, not in `kb/log.md`: it is a statement about *this* source, and the log is chronological rather than per-source. ## Template The block below is page material, so it is written in this instance's KB language (`kb/CONVENTIONS.md` `language:`) rather than in the control plane's English - its headings become the headings of every page `wikitool new source` scaffolds. ```markdown # Source: {name} **Autor:** {author} **Datum:** {date} **Raw-Dateien:** {raw_files|join} **Typ:** {source_type|capitalize} ## Zusammenfassung TODO: 2-3 Absätze zu den Kernaussagen des Quellmaterials. ## Kernaussagen - TODO: Kernaussage 1 - TODO: Kernaussage 2 - TODO: Kernaussage 3 ## Aufgaben - [ ] TODO: Aufgabe 1 - [ ] TODO: Aufgabe 2 ## Nicht übernommen - TODO: Was bewusst nicht übernommen wurde, und warum (oder "nichts - die Quelle wurde vollständig erfasst") ## Verwandte Entities {entities|bullets} ## Verwandte Concepts {concepts|bullets} ``` `# Source:` stays as a prefix - it mirrors the `title_prefix` and with it the title the page is linked and cited under. The value behind `**Typ:**` stays the English enum value. When `wikitool cite` adds a citation, the tool-managed footnote block appears at the end of the page; what it is called is the instance's decision in `kb/CONVENTIONS.md` (`sections:`). --- Additional notes: - Source pages are the authoritative catalogue of what raw material has been ingested - **One raw file, one owner.** A raw file appears in exactly one `raw_files:` - that page is responsible for keeping it summarized. Any number of pages may **cite** it via `[^cite-id]`; a citation is reuse, `raw_files:` is a maintenance responsibility. With two claimants it is undefined which page has to be brought up to date when the raw file changes - and then both rot quietly - Source pages make knowledge traceable back to the original raw material - `raw_files:` holds concrete existing file paths, never directories - A `raw_files:` list beyond roughly 15 entries indicates the cut was too coarse - the source should have been split into several source pages via `instructions/ingest-large-tree.md` - `entities:` plus `concepts:` beyond roughly 20 entries is the counterpart on the other axis: not cut too little, but compiled too much at once. Such a source is not split - a raw file has one owner - it needed the extract pass from `instructions/ingest-large-tree.md` § "A broad source is not cut", so that not every named subject gets a page