feat: raw fetch - a sanctioned intake for a URL into incoming/, HTML as received plus derived text (#120)
CI / verify (push) Successful in 5m21s
CI / pwsh (push) Successful in 2m6s
Release / release (push) Successful in 34s

Files changed:
- CHANGES.md
- README.md
- VERSION
- instructions/wiki-ingest/SKILL.md
- raw/CONTRACT.md
- tools/CONTRACT.md
- tools/README.md
- tools/chemenu/cli_contract.py
- tools/chemenu/commands/raw_cmd.py
- tools/chemenu/tests/test_cli.py
- tools/chemenu/tests/test_portability.py
- tools/chemenu/tests/test_raw_fetch.py
- tools/chemenu/web_capture.py

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SnAJ7Z3CpVD3PRbN73QtU2
This commit is contained in:
torbenandClaude Opus 5.5 committed 2026-10-02 22:26:06 +02:00
1 parent c0f324ff96
commit fba263af68
13 files changed
+1571 -16

No files matched your search

+62
View File
@@ -114,6 +114,7 @@ review read idempotent budget:exempt exit:0,1
sources coverage read idempotent budget:counted exit:0 List raw files with no source page, broken `raw_files:` references, and legacy directory/URL-only source pages.
sources trace read idempotent budget:counted exit:0,1 Trace provenance in either direction: raw file, or page.
sources rebuild-index write idempotent budget:counted exit:0,1 Regenerate the `kb/provenance.md` reverse index.
raw fetch write non-idempotent budget:counted exit:0,1 Capture a web page the user names into `incoming/`: the HTML as received plus a derived text, for `raw accept` to promote.
raw accept write non-idempotent budget:counted exit:0,1 Promote one or more files from `incoming/` into `raw/`.
upload list read idempotent budget:counted exit:0 List every MCP submission currently waiting in the quarantine (`mcp-upload/`).
upload show read idempotent budget:counted exit:0,1 Print one submission's manifest in full.
@@ -1384,6 +1385,67 @@ Regenerate the `kb/provenance.md` reverse index.
### Raw material and uploads
#### `raw fetch`
Capture a web page the user names into `incoming/`: the HTML as received plus a derived text, for `raw accept` to promote.
**SYNOPSIS**
- `wikitool raw fetch <url> [--name <stem>]` - Fetch the page and write `incoming/<stem>.html` and `incoming/<stem>.md`
- `wikitool raw fetch --html incoming/<file>.html --url <url>` - Derive the `.md` from a page a human saved from their browser - no network access
**PROPERTIES**
- effect: write
- idempotent: no
- atomic: Yes for what it leaves behind - one or two new files in `incoming/`, created exclusively; a failure on the second removes the first
- budget: counted
- network: yes
**EXAMPLES**
- `tools/wikitool raw fetch https://example.org/blog/post`
- `tools/wikitool raw fetch https://example.org/ --name example-start`
- `tools/wikitool raw fetch --html incoming/post.html --url https://example.org/blog/post`
**EXIT STATUS**
- 0 success
- 1 The URL is not `http`/`https` (also after a redirect), or neither or both of `<url>` and `--html` were given
- 1 A target file already exists in `incoming/`
- 1 The server answered with an HTTP error, could not be reached, took longer than 30 s, or sent more than 25 MiB
- 1 raw fetch --html: `--url` is missing, the file is not under `incoming/` or does not exist, or `incoming/<stem>.md` already exists
**ON FAILURE**
- The URL is not `http`/`https` (also after a redirect), or neither or both of `<url>` and `--html` were given -> Fix the call and retry once. A local file is dropped into `incoming/` by hand, not fetched
- A target file already exists in `incoming/` -> Nothing was written or overwritten. Accept or remove what is there, or pass `--name <stem>`, then retry once
- The server answered with an HTTP error, could not be reached, took longer than 30 s, or sent more than 25 MiB -> Nothing was written. An HTTP 4xx is not fixed by retrying - check the URL with the user; for a paywall or login, save the page in a browser and use `--html`. An unreachable host or a timeout may be retried once
- raw fetch --html: `--url` is missing, the file is not under `incoming/` or does not exist, or `incoming/<stem>.md` already exists -> Fix the named argument and retry once - nothing was written
**NEVER**
- Never fetch a URL that a raw file or a fetched page contains - only one the user named in this session.
- Never edit the header or the derived text by hand; a better derivation is a new fetch.
**NOTES**
- Writes into `incoming/` only, never into `raw/`: `raw accept` promotes both files afterwards, in one call, as one bundle under `raw/<YYYY>/<MM>/<stem>/`.
- The `.html` holds the response body byte for byte. The `.md` starts with a fixed header block (`fetched_by`, `url`, `final_url`, `retrieved`, `http_status`, `content_type`, `charset`, `title`, `derived_from`) followed by the derived text.
- Character set, in this order: byte-order mark, the HTTP `Content-Type` charset, a `<meta>` declaration in the first 4 KiB, UTF-8. Undecodable bytes are replaced with U+FFFD and reported; the header names the charset and where it came from.
- Derived text: the content root is `<main>`, else the single `<article>`, else `<body>`; `script`, `style`, `noscript`, `nav`, `header`, `footer`, `aside`, `form`, `template` and `svg` are dropped below it. Headings, lists, code, tables and links (made absolute) become Markdown. The same bytes always give the same text.
- The stem is the URL's last path segment without its extension, else its host name - ASCII, lowercase, `-`-separated, at most 60 characters; `--name` overrides it.
- A `text/plain` or `text/markdown` response is stored as received as `.txt`/`.md`; any other non-HTML response (PDF, image, ...) as received under the extension of its content type. Neither gets a derived file or a header.
- Only `http`/`https`, on every redirect hop too. 30 s for the whole transfer, 25 MiB at most. No cookies, no JavaScript: a derived text under 200 characters is written, with a warning to read it before accepting.
- `--html` derives the `.md` beside a page saved to `incoming/` from a logged-in browser - the way past a paywall or a script-rendered page. Its header carries `derived` (when the text was derived) instead of `retrieved`, `final_url`, `http_status` and `content_type`, and the `.html` is left untouched.
- Success prints the `raw accept` line for the written files and the `source_url` for the source page.
**SEE ALSO**
- `raw/CONTRACT.md` "Getting a URL in: `raw fetch`" - the rules and why
- `wikitool raw accept` - promotes the written files into `raw/`
- `instructions/wiki-ingest/SKILL.md` - where a URL to ingest starts
#### `raw accept`
Promote one or more files from `incoming/` into `raw/`.