feat: raw fetch - a sanctioned intake for a URL into incoming/, HTML as received plus derived text (#120)
CI / verify (push) Successful in 5m21s
CI / pwsh (push) Successful in 2m6s
Release / release (push) Successful in 34s

Files changed:
- CHANGES.md
- README.md
- VERSION
- instructions/wiki-ingest/SKILL.md
- raw/CONTRACT.md
- tools/CONTRACT.md
- tools/README.md
- tools/chemenu/cli_contract.py
- tools/chemenu/commands/raw_cmd.py
- tools/chemenu/tests/test_cli.py
- tools/chemenu/tests/test_portability.py
- tools/chemenu/tests/test_raw_fetch.py
- tools/chemenu/web_capture.py

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SnAJ7Z3CpVD3PRbN73QtU2
This commit is contained in:
torbenandClaude Opus 5.5 committed 2026-10-02 22:26:06 +02:00
1 parent c0f324ff96
commit fba263af68
13 files changed
+1571 -16

No files matched your search

+62
View File
@@ -114,6 +114,7 @@ review read idempotent budget:exempt exit:0,1
sources coverage read idempotent budget:counted exit:0 List raw files with no source page, broken `raw_files:` references, and legacy directory/URL-only source pages.
sources trace read idempotent budget:counted exit:0,1 Trace provenance in either direction: raw file, or page.
sources rebuild-index write idempotent budget:counted exit:0,1 Regenerate the `kb/provenance.md` reverse index.
raw fetch write non-idempotent budget:counted exit:0,1 Capture a web page the user names into `incoming/`: the HTML as received plus a derived text, for `raw accept` to promote.
raw accept write non-idempotent budget:counted exit:0,1 Promote one or more files from `incoming/` into `raw/`.
upload list read idempotent budget:counted exit:0 List every MCP submission currently waiting in the quarantine (`mcp-upload/`).
upload show read idempotent budget:counted exit:0,1 Print one submission's manifest in full.
@@ -1384,6 +1385,67 @@ Regenerate the `kb/provenance.md` reverse index.
### Raw material and uploads
#### `raw fetch`
Capture a web page the user names into `incoming/`: the HTML as received plus a derived text, for `raw accept` to promote.
**SYNOPSIS**
- `wikitool raw fetch <url> [--name <stem>]` - Fetch the page and write `incoming/<stem>.html` and `incoming/<stem>.md`
- `wikitool raw fetch --html incoming/<file>.html --url <url>` - Derive the `.md` from a page a human saved from their browser - no network access
**PROPERTIES**
- effect: write
- idempotent: no
- atomic: Yes for what it leaves behind - one or two new files in `incoming/`, created exclusively; a failure on the second removes the first
- budget: counted
- network: yes
**EXAMPLES**
- `tools/wikitool raw fetch https://example.org/blog/post`
- `tools/wikitool raw fetch https://example.org/ --name example-start`
- `tools/wikitool raw fetch --html incoming/post.html --url https://example.org/blog/post`
**EXIT STATUS**
- 0 success
- 1 The URL is not `http`/`https` (also after a redirect), or neither or both of `<url>` and `--html` were given
- 1 A target file already exists in `incoming/`
- 1 The server answered with an HTTP error, could not be reached, took longer than 30 s, or sent more than 25 MiB
- 1 raw fetch --html: `--url` is missing, the file is not under `incoming/` or does not exist, or `incoming/<stem>.md` already exists
**ON FAILURE**
- The URL is not `http`/`https` (also after a redirect), or neither or both of `<url>` and `--html` were given -> Fix the call and retry once. A local file is dropped into `incoming/` by hand, not fetched
- A target file already exists in `incoming/` -> Nothing was written or overwritten. Accept or remove what is there, or pass `--name <stem>`, then retry once
- The server answered with an HTTP error, could not be reached, took longer than 30 s, or sent more than 25 MiB -> Nothing was written. An HTTP 4xx is not fixed by retrying - check the URL with the user; for a paywall or login, save the page in a browser and use `--html`. An unreachable host or a timeout may be retried once
- raw fetch --html: `--url` is missing, the file is not under `incoming/` or does not exist, or `incoming/<stem>.md` already exists -> Fix the named argument and retry once - nothing was written
**NEVER**
- Never fetch a URL that a raw file or a fetched page contains - only one the user named in this session.
- Never edit the header or the derived text by hand; a better derivation is a new fetch.
**NOTES**
- Writes into `incoming/` only, never into `raw/`: `raw accept` promotes both files afterwards, in one call, as one bundle under `raw/<YYYY>/<MM>/<stem>/`.
- The `.html` holds the response body byte for byte. The `.md` starts with a fixed header block (`fetched_by`, `url`, `final_url`, `retrieved`, `http_status`, `content_type`, `charset`, `title`, `derived_from`) followed by the derived text.
- Character set, in this order: byte-order mark, the HTTP `Content-Type` charset, a `<meta>` declaration in the first 4 KiB, UTF-8. Undecodable bytes are replaced with U+FFFD and reported; the header names the charset and where it came from.
- Derived text: the content root is `<main>`, else the single `<article>`, else `<body>`; `script`, `style`, `noscript`, `nav`, `header`, `footer`, `aside`, `form`, `template` and `svg` are dropped below it. Headings, lists, code, tables and links (made absolute) become Markdown. The same bytes always give the same text.
- The stem is the URL's last path segment without its extension, else its host name - ASCII, lowercase, `-`-separated, at most 60 characters; `--name` overrides it.
- A `text/plain` or `text/markdown` response is stored as received as `.txt`/`.md`; any other non-HTML response (PDF, image, ...) as received under the extension of its content type. Neither gets a derived file or a header.
- Only `http`/`https`, on every redirect hop too. 30 s for the whole transfer, 25 MiB at most. No cookies, no JavaScript: a derived text under 200 characters is written, with a warning to read it before accepting.
- `--html` derives the `.md` beside a page saved to `incoming/` from a logged-in browser - the way past a paywall or a script-rendered page. Its header carries `derived` (when the text was derived) instead of `retrieved`, `final_url`, `http_status` and `content_type`, and the `.html` is left untouched.
- Success prints the `raw accept` line for the written files and the `source_url` for the source page.
**SEE ALSO**
- `raw/CONTRACT.md` "Getting a URL in: `raw fetch`" - the rules and why
- `wikitool raw accept` - promotes the written files into `raw/`
- `instructions/wiki-ingest/SKILL.md` - where a URL to ingest starts
#### `raw accept`
Promote one or more files from `incoming/` into `raw/`.
+1
View File
@@ -105,6 +105,7 @@ tools/
version.py the stack version: VERSION, the release stamp, the compatibility rule
kb_state.py the KB version (.wikitool-kb.json) and the migration chain
corpus_diff.py invariant comparison of kb/ between two revisions
web_capture.py `raw fetch`'s core: fetch a page, decide its charset, derive Markdown-like text from the HTML - standard library only, deterministic
search/ pluggable search backends, plus service.py - the search core
tasks/ the task-tracker provider layer: protocol.py (TaskReader/TaskWriter), config.py (.wikitool-tasks.json), one module per adapter - no instruction ever learns which provider it is
commands/ one module per command or command group: the terminal adapters
+1 -1
View File
@@ -259,7 +259,7 @@ GROUPS: tuple[tuple[str, tuple[str, ...]], ...] = (
"sources coverage", "sources trace", "sources rebuild-index",
)),
("Raw material and uploads", (
"raw accept", "upload list", "upload show", "upload accept", "upload reject",
"raw fetch", "raw accept", "upload list", "upload show", "upload accept", "upload reject",
)),
("Git", (
"sync", "publish",
+265 -4
View File
@@ -70,15 +70,18 @@ from pathlib import Path
from typing import Optional
import typer
from rich.markup import escape
from chemenu import cli_contract, config
from chemenu import cli_contract, config, web_capture
from chemenu.commands._util import check_path_budget, fail, rel_path, success
from chemenu.errors import ChemenuError
from chemenu.frontmatter_io import write_page
from chemenu.kb_scan import load_kb_pages
from chemenu.provenance import citing_pages, source_pages_by_raw_file, source_raw_files
from chemenu.type_resolver import resolver
from chemenu.version import VersionError, read_version
app = typer.Typer(help="Promote raw material out of incoming/ into raw/.")
app = typer.Typer(help="Bring raw material into incoming/, and promote it from there into raw/.")
# The type this command always promotes into eventually - hardcoded rather
# than derived from --page, because there is no page yet in the common case
@@ -118,8 +121,8 @@ def _validate_under_incoming(path: Path, incoming: Path) -> None:
rel = path.relative_to(incoming)
except ValueError:
fail(
f"{rel_path(path)} is not under incoming/ - `raw accept` only promotes files "
"from there. See raw/CONTRACT.md."
f"{rel_path(path)} is not under incoming/ - `raw accept` and `raw fetch --html` "
"only take files from there. See raw/CONTRACT.md."
)
if len(rel.parts) < 1:
fail(f"incoming/{rel.as_posix()} names no file.")
@@ -658,3 +661,261 @@ def raw_accept_command(
f" --set fidelity={fidelity} --set authority={authority} \\\n"
" --set source_type=<category>"
)
# --- raw fetch ----------------------------------------------------------------
#
# The sanctioned intake for a URL the user names. It ends in `incoming/`, never
# in `raw/`: promoting stays `raw accept`'s job, so `wiki-ingest` keeps its
# order (the commitment question comes before the promotion) and the capture
# fields are asked exactly once, there. The fetch and derivation themselves live
# in `chemenu.web_capture`; this adapter decides file names and what the caller
# is told.
_TEASER_HINT = (
" If the text stops at a teaser (paywall, login, script-rendered page): save the page from a\n"
" logged-in browser to incoming/ and run tools/wikitool raw fetch --html "
"incoming/<file>.html --url {url}"
)
def _user_agent() -> str:
try:
version = str(read_version())
except VersionError:
version = "unknown"
return f"chemenu-wikitool/{version} (raw fetch)"
def _next_lines(files: list[Path], source_url: str) -> str:
paths = " ".join(rel_path(f) for f in files)
return (
" Next (after the commitment question in wiki-ingest):\n"
f" tools/wikitool raw accept --fidelity published --authority <value> {paths}\n"
f" and on the source page: --set source_url={source_url}"
)
def _refuse_existing(targets: list[Path]) -> None:
existing = [t for t in targets if t.exists()]
if existing:
fail(escape(
f"{', '.join(rel_path(t) for t in existing)} already exists in incoming/ - raw fetch "
"never overwrites. Accept or remove the file that is there, or pass --name <stem>."
))
def _write_exclusive(writes: list[tuple[Path, bytes]]) -> None:
"""Create every file or none: `x` mode refuses a file that appeared since
the check above, and a failure part-way removes what this call wrote."""
written: list[Path] = []
try:
for path, data in writes:
path.parent.mkdir(parents=True, exist_ok=True)
with path.open("xb") as handle:
written.append(path)
handle.write(data)
except OSError as exc:
for path in written:
path.unlink(missing_ok=True)
fail(escape(f"Could not write {rel_path(path)}: {exc}. Nothing was left in incoming/."))
def _warn(derivation: web_capture.Derivation, md_path: Path) -> None:
if derivation.decoded.replaced:
typer.echo(
f"WARN {derivation.decoded.replaced} undecodable byte(s) under "
f"{derivation.decoded.charset} were replaced with U+FFFD in {rel_path(md_path)} - "
"the .html beside it keeps the original bytes."
)
if derivation.text_chars < web_capture.SHORT_TEXT_CHARS:
typer.echo(
f"WARN The derived text is only {derivation.text_chars} characters - likely a "
"script-rendered page, a login wall or an empty response. Read "
f"{rel_path(md_path)} before accepting it."
)
def _fetch_html_file(html: Path, source_url: Optional[str]) -> None:
if source_url is None:
fail("--html needs --url <url>: it goes into the header, resolves relative links, and "
"becomes source_url on the source page.")
try:
web_capture.check_url(source_url)
except ChemenuError as exc:
fail(escape(str(exc)))
path = _resolve(html)
_validate_under_incoming(path, _incoming_dir())
if not path.is_file():
fail(f"{rel_path(path)} does not exist or is not a file.")
md_path = path.with_suffix(".md")
if md_path == path:
fail(f"{rel_path(path)} is already a .md file - --html takes the saved HTML page.")
_refuse_existing([md_path])
derivation = web_capture.derive_document(path.read_bytes(), source_url, path.name, url=source_url)
_write_exclusive([(md_path, derivation.markdown)])
_warn(derivation, md_path)
success(escape(
f"Derived {rel_path(md_path)} from {rel_path(path)} (no network access).\n"
+ _next_lines([path, md_path], source_url)
))
@cli_contract.record(cli_contract.CommandRecord(
path="raw fetch",
summary="Capture a web page the user names into `incoming/`: the HTML as received plus a "
"derived text, for `raw accept` to promote.",
synopsis=(
cli_contract.Variant(
usage="raw fetch <url> [--name <stem>]",
notes="Fetch the page and write `incoming/<stem>.html` and `incoming/<stem>.md`",
),
cli_contract.Variant(
usage="raw fetch --html incoming/<file>.html --url <url>",
notes="Derive the `.md` from a page a human saved from their browser - no network "
"access",
),
),
properties=cli_contract.Properties(
effect=cli_contract.Effect.WRITE,
idempotent=cli_contract.Idempotent.NO,
atomic="Yes for what it leaves behind - one or two new files in `incoming/`, created "
"exclusively; a failure on the second removes the first",
budget=cli_contract.Budget.COUNTED,
network=cli_contract.Network.YES,
),
notes=(
"Writes into `incoming/` only, never into `raw/`: `raw accept` promotes both files "
"afterwards, in one call, as one bundle under `raw/<YYYY>/<MM>/<stem>/`.",
"The `.html` holds the response body byte for byte. The `.md` starts with a fixed "
"header block (`fetched_by`, `url`, `final_url`, `retrieved`, `http_status`, "
"`content_type`, `charset`, `title`, `derived_from`) followed by the derived text.",
"Character set, in this order: byte-order mark, the HTTP `Content-Type` charset, a "
"`<meta>` declaration in the first 4 KiB, UTF-8. Undecodable bytes are replaced with "
"U+FFFD and reported; the header names the charset and where it came from.",
"Derived text: the content root is `<main>`, else the single `<article>`, else "
"`<body>`; `script`, `style`, `noscript`, `nav`, `header`, `footer`, `aside`, `form`, "
"`template` and `svg` are dropped below it. Headings, lists, code, tables and links "
"(made absolute) become Markdown. The same bytes always give the same text.",
"The stem is the URL's last path segment without its extension, else its host name - "
"ASCII, lowercase, `-`-separated, at most 60 characters; `--name` overrides it.",
"A `text/plain` or `text/markdown` response is stored as received as `.txt`/`.md`; any "
"other non-HTML response (PDF, image, ...) as received under the extension of its "
"content type. Neither gets a derived file or a header.",
"Only `http`/`https`, on every redirect hop too. 30 s for the whole transfer, 25 MiB at "
"most. No cookies, no JavaScript: a derived text under 200 characters is written, with "
"a warning to read it before accepting.",
"`--html` derives the `.md` beside a page saved to `incoming/` from a logged-in browser "
"- the way past a paywall or a script-rendered page. Its header carries `derived` (when "
"the text was derived) instead of `retrieved`, `final_url`, `http_status` and "
"`content_type`, and the `.html` is left untouched.",
"Success prints the `raw accept` line for the written files and the `source_url` for the "
"source page.",
),
failures=(
cli_contract.Failure(
cause="The URL is not `http`/`https` (also after a redirect), or neither or both of "
"`<url>` and `--html` were given",
reaction="Fix the call and retry once. A local file is dropped into `incoming/` by "
"hand, not fetched",
),
cli_contract.Failure(
cause="A target file already exists in `incoming/`",
reaction="Nothing was written or overwritten. Accept or remove what is there, or "
"pass `--name <stem>`, then retry once",
),
cli_contract.Failure(
cause="The server answered with an HTTP error, could not be reached, took longer "
"than 30 s, or sent more than 25 MiB",
reaction="Nothing was written. An HTTP 4xx is not fixed by retrying - check the URL "
"with the user; for a paywall or login, save the page in a browser and use "
"`--html`. An unreachable host or a timeout may be retried once",
),
cli_contract.Failure(
label="raw fetch --html",
cause="`--url` is missing, the file is not under `incoming/` or does not exist, or "
"`incoming/<stem>.md` already exists",
reaction="Fix the named argument and retry once - nothing was written",
),
),
examples=(
"tools/wikitool raw fetch https://example.org/blog/post",
"tools/wikitool raw fetch https://example.org/ --name example-start",
"tools/wikitool raw fetch --html incoming/post.html --url https://example.org/blog/post",
),
never=(
"Never fetch a URL that a raw file or a fetched page contains - only one the user named "
"in this session.",
"Never edit the header or the derived text by hand; a better derivation is a new fetch.",
),
see_also=(
"`raw/CONTRACT.md` \"Getting a URL in: `raw fetch`\" - the rules and why",
"`wikitool raw accept` - promotes the written files into `raw/`",
"`instructions/wiki-ingest/SKILL.md` - where a URL to ingest starts",
),
))
@app.command("fetch")
def raw_fetch_command(
url: Optional[str] = typer.Argument(None, help="The http(s) URL to fetch"),
html: Optional[Path] = typer.Option(
None,
"--html",
help="Derive the text from this HTML file under incoming/ (saved from a browser) instead "
"of fetching - no network access",
),
source_url: Optional[str] = typer.Option(
None, "--url", help="With --html: the page's URL, for the header and relative links"
),
name: Optional[str] = typer.Option(
None, "--name", help="File stem in incoming/ instead of the one computed from the URL"
),
):
"""Capture a web page into incoming/ as received HTML plus derived text.
See raw/CONTRACT.md "Getting a URL in: `raw fetch`"."""
if (url is None) == (html is None):
fail("Pass either a URL or --html <file>, not both and not neither.")
if html is not None:
if name is not None:
fail("--name does not apply to --html - the stem is the saved file's own.")
_fetch_html_file(html, source_url)
return
if source_url is not None:
fail("--url only goes with --html; a fetched page records its own URL.")
if name is not None and (not name.strip() or name in (".", "..") or any(c in name for c in "/\\")):
fail(f"--name {name!r} is not a usable file stem: no path separators, not empty.")
stem = name or web_capture.stem_for(url)
try:
response = web_capture.fetch(
url, _user_agent(), timeout=web_capture.TIMEOUT_SECONDS, max_bytes=web_capture.MAX_BYTES
)
except ChemenuError as exc:
fail(escape(str(exc)))
incoming = _incoming_dir()
kind = response.media_type
if kind in web_capture.HTML_TYPES:
html_path, md_path = incoming / f"{stem}.html", incoming / f"{stem}.md"
_refuse_existing([html_path, md_path])
derivation = web_capture.derive_document(
response.body, response.final_url, html_path.name, response=response
)
_write_exclusive([(html_path, response.body), (md_path, derivation.markdown)])
_warn(derivation, md_path)
success(escape(
f"Fetched {url} to {rel_path(html_path)}, {rel_path(md_path)}\n"
+ _next_lines([html_path, md_path], url) + "\n"
+ _TEASER_HINT.format(url=url)
))
return
target = incoming / f"{stem}{web_capture.extension_for(response.content_type)}"
_refuse_existing([target])
_write_exclusive([(target, response.body)])
success(escape(
f"Fetched {url} to {rel_path(target)} ({kind or 'no content type'}, stored as received - "
f"no text derived; retrieved {web_capture.utc_now()}).\n"
+ _next_lines([target], url)
))
+4 -1
View File
@@ -338,10 +338,13 @@ def test_network_yes_is_exactly_the_commands_that_can_reach_outside_this_checkou
"""Gitea #144: `network:` means *any* reach outside this checkout - an HTTP call
`wikitool` makes itself, or a git operation against a remote (fetch/ls-remote/push) -
not only the two `version_cmd.py` used to claim exclusivity for. `dist upgrade` is on the list
since `--latest`, which asks the release feed and downloads from it. Pinned as an explicit
since `--latest`, which asks the release feed and downloads from it, and `raw fetch`
since it exists (Gitea #120) - its `--html` form stays offline, which does not turn the
command back to `no`. Pinned as an explicit
set so a command gaining or losing that reach is a deliberate edit here, not a silent
drift between the property and what the command actually does."""
expected = {
"raw fetch",
"sync",
"publish",
"version check",
+4 -3
View File
@@ -26,9 +26,10 @@ from chemenu import filelock
TOOLS_DIR = Path(__file__).resolve().parents[2]
PACKAGE_DIR = TOOLS_DIR / "chemenu"
# Receivers whose `open` is not a text-file open: `os.open` takes flags, and the
# archive modules open binary members.
NOT_TEXT_OPEN = {"os", "tarfile", "zipfile", "gzip"}
# Receivers whose `open` is not a text-file open: `os.open` takes flags, the
# archive modules open binary members, and `opener` is a `urllib`
# `OpenerDirector` (`web_capture.fetch`), whose `open` sends an HTTP request.
NOT_TEXT_OPEN = {"os", "tarfile", "zipfile", "gzip", "opener"}
# The one module allowed to import the platform lock modules.
LOCK_MODULE = PACKAGE_DIR / "filelock.py"
+429
View File
@@ -0,0 +1,429 @@
"""`wikitool raw fetch` - the sanctioned intake for a URL (Gitea #120).
Every fetch goes to a real `http.server` on 127.0.0.1, never to the network:
what is under test is what arrives on the wire and what lands on disk - the
exact bytes, the charset a header did or did not declare, a redirect, a body
that is too large or too slow - and a fetcher stub would have to fake exactly
the parts that matter.
"""
from __future__ import annotations
import datetime
import http.server
import threading
import time
import urllib.request
from pathlib import Path
import pytest
import typer
import yaml
from chemenu import web_capture
from chemenu.commands.raw_cmd import raw_accept_command, raw_fetch_command
ARTICLE = (
b"<!doctype html><html><head><meta charset=\"utf-8\"><title>A Post</title>"
b"<script>var tracking = 'SCRIPT-TEXT';</script><style>.x{}</style></head><body>"
b"<nav><a href=\"/\">NAV-TEXT</a></nav>"
b"<main><h1>Main heading</h1><p>First paragraph with a <a href=\"other\">relative link</a>.</p>"
b"<ul><li>one</li><li>two</li></ul><pre>code line</pre></main>"
b"<aside>ASIDE-TEXT</aside><footer>FOOTER-TEXT</footer></body></html>"
)
class Site:
"""Serves `routes[path] = (status, headers, body)`; a body that is a
callable gets the handler instead, for the slow and the unbounded cases.
Records every request's path and User-Agent."""
def __init__(self) -> None:
self.routes: dict = {}
self.requests: list[tuple[str, str]] = []
outer = self
class Handler(http.server.BaseHTTPRequestHandler):
def do_GET(self): # noqa: N802 - http.server's name
outer.requests.append((self.path, self.headers.get("User-Agent", "")))
status, headers, body = outer.routes.get(self.path, (404, {}, b"not found"))
self.send_response(status)
for key, value in headers.items():
self.send_header(key, value)
if callable(body):
self.end_headers()
body(self)
return
if "Content-Length" not in headers:
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def log_message(self, *args): # silence the test output
pass
self.httpd = http.server.ThreadingHTTPServer(("127.0.0.1", 0), Handler)
self.httpd.daemon_threads = True
self.thread = threading.Thread(
target=self.httpd.serve_forever, kwargs={"poll_interval": 0.01}, daemon=True
)
self.thread.start()
def url(self, path: str) -> str:
return f"http://127.0.0.1:{self.httpd.server_address[1]}{path}"
def html(self, path: str, body: bytes, content_type: str = "text/html; charset=utf-8") -> str:
self.routes[path] = (200, {"Content-Type": content_type}, body)
return self.url(path)
def stop(self) -> None:
self.httpd.shutdown()
self.httpd.server_close()
@pytest.fixture
def site(monkeypatch: pytest.MonkeyPatch):
# A developer's or CI runner's proxy must not see a loopback request.
for var in ("http_proxy", "HTTP_PROXY", "https_proxy", "HTTPS_PROXY", "all_proxy", "ALL_PROXY"):
monkeypatch.delenv(var, raising=False)
server = Site()
yield server
server.stop()
@pytest.fixture
def tree(kb_dir):
"""kb_dir repoints config.ROOT at tmp_path; add raw/ and incoming/."""
root = kb_dir.parent
(root / "raw").mkdir()
(root / "incoming").mkdir(exist_ok=True)
return root
def _fetch(url=None, html=None, source_url=None, name=None):
return raw_fetch_command(url=url, html=html, source_url=source_url, name=name)
def _refused(**kwargs) -> None:
with pytest.raises(typer.Exit) as excinfo:
_fetch(**kwargs)
assert excinfo.value.exit_code == 1
def _snapshot(root: Path) -> dict[str, bytes]:
return {str(p.relative_to(root)): p.read_bytes() for p in sorted(root.rglob("*")) if p.is_file()}
def _header_and_text(md: Path) -> tuple[dict, str]:
content = md.read_text(encoding="utf-8")
assert content.startswith("---\n")
head, _, text = content[4:].partition("\n---\n")
return yaml.safe_load(head), text
# --- the fetched bundle -----------------------------------------------------
def test_fetch_writes_html_and_md_to_incoming_and_leaves_raw_untouched(tree, site):
url = site.html("/blog/post.html", ARTICLE)
raw_before = _snapshot(tree / "raw")
_fetch(url)
assert sorted(p.name for p in (tree / "incoming").iterdir()) == ["post.html", "post.md"]
assert _snapshot(tree / "raw") == raw_before
def test_the_html_is_byte_identical_to_what_the_server_sent(tree, site):
body = ARTICLE + b"\r\n<!-- trailing CRLF and \xc3\xa4 bytes -->\r\n"
_fetch(site.html("/post", body))
assert (tree / "incoming/post.html").read_bytes() == body
def test_header_records_the_capture(tree, site):
url = site.html("/post", ARTICLE)
_fetch(url)
header, _ = _header_and_text(tree / "incoming/post.md")
assert header["fetched_by"] == "wikitool raw fetch"
assert header["url"] == url
assert header["final_url"] == url
assert header["http_status"] == 200
assert header["content_type"] == "text/html; charset=utf-8"
assert header["charset"] == "utf-8 (from: header)"
assert header["title"] == "A Post"
assert header["derived_from"] == "post.html"
# YAML reads the ISO-8601 UTC stamp back as an aware datetime.
retrieved = header["retrieved"]
assert retrieved.tzinfo is not None and retrieved.utcoffset() == datetime.timedelta(0)
assert abs(datetime.datetime.now(datetime.timezone.utc) - retrieved) < datetime.timedelta(minutes=5)
assert "author" not in header
def test_the_request_identifies_the_tool(tree, site):
_fetch(site.html("/post", ARTICLE))
assert site.requests[0][1].startswith("chemenu-wikitool/")
def test_two_fetches_of_the_same_bytes_derive_identical_text(tree, site):
url = site.html("/post", ARTICLE)
_fetch(url, name="first")
_fetch(url, name="second")
def without_retrieved(path: Path) -> list[str]:
lines = path.read_text(encoding="utf-8").splitlines()
return [line.replace("first.html", "X").replace("second.html", "X")
for line in lines if not line.startswith("retrieved:")]
assert without_retrieved(tree / "incoming/first.md") == without_retrieved(tree / "incoming/second.md")
def test_boilerplate_is_dropped_and_main_is_kept_whole(tree, site):
_fetch(site.html("/post", ARTICLE))
_, text = _header_and_text(tree / "incoming/post.md")
for dropped in ("NAV-TEXT", "FOOTER-TEXT", "ASIDE-TEXT", "SCRIPT-TEXT", ".x{}"):
assert dropped not in text
assert "# Main heading" in text
assert "First paragraph with a [relative link](" in text
assert "- one\n- two" in text
assert "```\ncode line\n```" in text
def test_relative_links_resolve_against_the_final_url_after_a_redirect(tree, site):
target = site.html("/articles/post", ARTICLE)
site.routes["/old"] = (301, {"Location": "/articles/post"}, b"")
_fetch(site.url("/old"), name="post")
header, text = _header_and_text(tree / "incoming/post.md")
assert header["url"] == site.url("/old")
assert header["final_url"] == target
assert f"[relative link]({site.url('/articles/other')})" in text
def test_latin1_declared_only_in_meta_is_decoded(tree, site):
body = (
"<html><head><meta http-equiv=\"Content-Type\" content=\"text/html; charset=iso-8859-1\">"
"<title>Umlaute</title></head><body><main><p>Größe, Übermaß, Äpfel</p></main></body></html>"
).encode("iso-8859-1")
_fetch(site.html("/umlaute", body, content_type="text/html"))
header, text = _header_and_text(tree / "incoming/umlaute.md")
assert header["charset"] == "iso-8859-1 (from: meta)"
assert "Größe, Übermaß, Äpfel" in text
assert (tree / "incoming/umlaute.html").read_bytes() == body
def test_the_header_charset_wins_over_meta_and_the_bom_over_both():
meta_says_latin1 = b'<meta charset="iso-8859-1"><p>\xc3\xa4</p>'
assert web_capture.decode(meta_says_latin1, "text/html; charset=utf-8").text.endswith("<p>ä</p>")
decoded = web_capture.decode(b"\xef\xbb\xbf" + meta_says_latin1, "text/html; charset=iso-8859-1")
assert (decoded.charset, decoded.source) == ("utf-8", "bom")
assert web_capture.decode(b"<p>x</p>", None).source == "default"
def test_undecodable_bytes_are_replaced_and_reported(tree, site, capsys):
body = b"<html><body><main><p>broken \xff\xfe here and enough text " + b"x" * 300 + b"</p></main></body></html>"
_fetch(site.html("/broken", body))
_, text = _header_and_text(tree / "incoming/broken.md")
assert "broken �� here" in text
assert "undecodable byte" in capsys.readouterr().out
def test_a_nearly_empty_page_is_written_with_a_warning(tree, site, capsys):
body = b"<html><body><div id=\"app\"></div><script>render()</script></body></html>"
_fetch(site.html("/spa", body))
assert (tree / "incoming/spa.md").is_file()
assert "only 0 characters" in capsys.readouterr().out
def test_the_bundle_is_accepted_into_raw_as_one_source(tree, site):
_fetch(site.html("/blog/post", ARTICLE))
raw_accept_command(
files=[tree / "incoming/post.html", tree / "incoming/post.md"],
fidelity="published", authority="reporting", page=None, replaces=None, dry_run=False,
)
today = datetime.date.today()
bundle = tree / "raw" / f"{today.year:04d}" / f"{today.month:02d}" / "post"
assert sorted(p.name for p in bundle.iterdir()) == ["post.html", "post.md"]
assert (bundle / "post.html").read_bytes() == ARTICLE
# --- refusals ---------------------------------------------------------------
@pytest.mark.parametrize("existing", ["post.html", "post.md"])
def test_an_existing_target_is_never_overwritten(tree, site, existing):
held = tree / "incoming" / existing
held.write_bytes(b"already here")
_refused(url=site.html("/post", ARTICLE))
assert held.read_bytes() == b"already here"
assert [p.name for p in (tree / "incoming").iterdir()] == [existing]
def test_a_file_url_is_refused_without_writing(tree):
_refused(url="file:///etc/passwd")
assert list((tree / "incoming").iterdir()) == []
def test_a_redirect_to_another_scheme_is_refused(tree, site):
site.routes["/sneaky"] = (302, {"Location": "ftp://127.0.0.1/secret"}, b"")
_refused(url=site.url("/sneaky"))
assert list((tree / "incoming").iterdir()) == []
@pytest.mark.parametrize("declared", [True, False])
def test_a_response_over_the_size_limit_is_refused(tree, site, monkeypatch, declared):
monkeypatch.setattr(web_capture, "MAX_BYTES", 1024)
body = b"<p>" + b"x" * 4096 + b"</p>"
if declared:
site.html("/big", body)
else:
def stream(handler): # no Content-Length: HTTP/1.0, the body ends at close
handler.wfile.write(body)
site.routes["/big"] = (200, {"Content-Type": "text/html"}, stream)
_refused(url=site.url("/big"))
assert list((tree / "incoming").iterdir()) == []
def test_a_slow_server_times_out(tree, site, monkeypatch):
monkeypatch.setattr(web_capture, "TIMEOUT_SECONDS", 0.3)
def stall(handler):
handler.wfile.write(b"<p>start")
handler.wfile.flush()
time.sleep(1.5)
site.routes["/slow"] = (200, {"Content-Type": "text/html"}, stall)
_refused(url=site.url("/slow"))
assert list((tree / "incoming").iterdir()) == []
def test_an_http_error_is_refused(tree, site):
_refused(url=site.url("/missing"))
assert list((tree / "incoming").iterdir()) == []
def test_url_and_html_are_exclusive(tree):
page = tree / "incoming/post.html"
page.write_bytes(ARTICLE)
_refused(url="https://example.org/post", html=page, source_url="https://example.org/post")
_refused()
_refused(url="https://example.org/post", source_url="https://example.org/post")
# --- non-HTML responses -----------------------------------------------------
def test_a_pdf_is_stored_byte_identical_without_a_derivation(tree, site):
body = b"%PDF-1.7\n\x00\x01binary\xff"
site.routes["/paper.pdf"] = (200, {"Content-Type": "application/pdf"}, body)
_fetch(site.url("/paper.pdf"))
assert [p.name for p in (tree / "incoming").iterdir()] == ["paper.pdf"]
assert (tree / "incoming/paper.pdf").read_bytes() == body
def test_plain_text_is_stored_as_txt_unchanged(tree, site):
site.routes["/notes"] = (200, {"Content-Type": "text/plain; charset=utf-8"}, b"line one\r\nline two\n")
_fetch(site.url("/notes"))
assert [p.name for p in (tree / "incoming").iterdir()] == ["notes.txt"]
assert (tree / "incoming/notes.txt").read_bytes() == b"line one\r\nline two\n"
# --- --html: a page saved from a browser ------------------------------------
def test_html_derives_only_the_md_without_network_and_matches_a_fetch(tree, site, monkeypatch):
url = site.html("/blog/post", ARTICLE)
_fetch(url, name="online")
saved = tree / "incoming/post.html"
saved.write_bytes(ARTICLE)
def no_network(*args, **kwargs):
raise AssertionError("--html must not touch the network")
monkeypatch.setattr(web_capture, "fetch", no_network)
monkeypatch.setattr(urllib.request, "urlopen", no_network)
monkeypatch.setattr(urllib.request.OpenerDirector, "open", no_network)
_fetch(html=saved, source_url=url)
assert saved.read_bytes() == ARTICLE
assert sorted(p.name for p in (tree / "incoming").iterdir()) == [
"online.html", "online.md", "post.html", "post.md",
]
header, text = _header_and_text(tree / "incoming/post.md")
_, online_text = _header_and_text(tree / "incoming/online.md")
assert text == online_text
assert header["fetched_by"] == "wikitool raw fetch --html"
assert header["url"] == url
assert header["derived_from"] == "post.html"
assert "derived" in header
for absent in ("retrieved", "final_url", "http_status", "content_type"):
assert absent not in header
assert header["charset"] == "utf-8 (from: meta)"
def test_html_outside_incoming_is_refused(tree):
outside = tree / "saved.html"
outside.write_bytes(ARTICLE)
_refused(html=outside, source_url="https://example.org/post")
assert not (tree / "saved.md").exists()
assert list((tree / "incoming").iterdir()) == []
def test_html_without_url_is_refused(tree):
saved = tree / "incoming/post.html"
saved.write_bytes(ARTICLE)
_refused(html=saved)
assert [p.name for p in (tree / "incoming").iterdir()] == ["post.html"]
def test_html_does_not_overwrite_an_existing_md(tree):
saved = tree / "incoming/post.html"
saved.write_bytes(ARTICLE)
(tree / "incoming/post.md").write_text("mine", encoding="utf-8")
_refused(html=saved, source_url="https://example.org/post")
assert (tree / "incoming/post.md").read_text(encoding="utf-8") == "mine"
# --- naming and the derivation itself ---------------------------------------
@pytest.mark.parametrize(
"url, stem",
[
("https://example.org/blog/My%20Post.html", "my-post"),
("https://example.org/blog/über-uns/", "uber-uns"),
("https://example.org/", "example-org"),
("https://example.org/" + "a" * 100, "a" * 60),
],
)
def test_the_stem_comes_from_the_url(url, stem):
assert web_capture.stem_for(url) == stem
def test_name_overrides_the_stem(tree, site):
_fetch(site.html("/post", ARTICLE), name="chosen")
assert sorted(p.name for p in (tree / "incoming").iterdir()) == ["chosen.html", "chosen.md"]
def test_name_with_a_path_separator_is_refused(tree, site):
_refused(url=site.html("/post", ARTICLE), name="../escape")
assert list((tree / "incoming").iterdir()) == []
def test_a_single_article_is_the_root_when_there_is_no_main():
html = "<body><p>OUTSIDE</p><article><p>INSIDE</p></article></body>"
assert web_capture.derive_text(html, "https://example.org/") == "INSIDE"
def test_two_articles_fall_back_to_body():
html = "<body><article><p>ONE</p></article><article><p>TWO</p></article></body>"
assert web_capture.derive_text(html, "https://example.org/") == "ONE\n\nTWO"
def test_a_title_with_a_colon_keeps_the_header_parseable(tree, site):
body = ARTICLE.replace(b"<title>A Post</title>", b"<title>Part 2: the #1 reason</title>")
_fetch(site.html("/post", body))
header, _ = _header_and_text(tree / "incoming/post.md")
assert header["title"] == "Part 2: the #1 reason"
+649
View File
@@ -0,0 +1,649 @@
"""Capturing a web page for `raw/`: fetch the bytes, decide their character
set, and derive a Markdown-like reading text from them - with no CLI attached.
`wikitool raw fetch` is the terminal adapter over this module; it decides where
the files go and what the caller is told. Everything here is deterministic on
purpose: the same bytes always produce the same derived text, so two sessions
capturing the same article produce the same `raw/` bundle rather than two
hand-built variants of it. That is also why there is no readability heuristic
(text density, line numbers): a heuristic tuned once drifts the next time it is
tuned, and the received HTML is kept beside the derivation anyway, so what the
derivation drops is never lost.
Standard library only. A third-party HTML library would be a new entry in
`tools/requirements.txt`, missing from an instance's venv until someone runs
`pip install` after an upgrade - which would make this a breaking change for a
gain the kept HTML already covers.
"""
from __future__ import annotations
import codecs
import datetime
import email.message
import json
import mimetypes
import re
import time
import unicodedata
import urllib.error
import urllib.parse
import urllib.request
from dataclasses import dataclass
from html.parser import HTMLParser
from typing import Optional, Union
from chemenu.errors import BackendError, ValidationError
TIMEOUT_SECONDS = 30.0
MAX_BYTES = 25 * 1024 * 1024
SHORT_TEXT_CHARS = 200
STEM_MAX_CHARS = 60
META_SCAN_BYTES = 4096
ALLOWED_SCHEMES = ("http", "https")
HTML_TYPES = ("text/html", "application/xhtml+xml")
# Content types stored as received, under a fixed extension, without a
# derivation: they are already text a session can read and cite directly.
PLAIN_TEXT_EXTENSIONS = {"text/plain": ".txt", "text/markdown": ".md"}
# --- fetching ---------------------------------------------------------------
@dataclass(frozen=True)
class Response:
url: str
final_url: str
status: int
content_type: str
body: bytes
@property
def media_type(self) -> str:
"""The bare media type, lowercased - `text/html` out of
`text/html; charset=utf-8`; empty when the server sent none."""
return media_type(self.content_type)
def media_type(content_type: str) -> str:
if not content_type.strip():
return ""
message = email.message.Message()
message["content-type"] = content_type
return message.get_content_type().lower()
def check_url(url: str) -> None:
"""Only `http`/`https`. A `file:` URL would be a way into `incoming/` that
bypasses a human dropping the file there; `ftp:` and the rest are not
pages."""
parts = urllib.parse.urlsplit(url)
if parts.scheme.lower() not in ALLOWED_SCHEMES or not parts.netloc:
raise ValidationError(
f"{url} is not an http(s) URL - `raw fetch` fetches web pages only. A local file "
"goes into incoming/ by hand."
)
class _HttpOnlyRedirects(urllib.request.HTTPRedirectHandler):
"""urllib follows a redirect to `ftp:` on its own; the scheme rule above
has to hold for every hop, not only for the URL the caller typed."""
def redirect_request(self, req, fp, code, msg, headers, newurl): # noqa: D102 - urllib's hook
check_url(urllib.parse.urljoin(req.full_url, newurl))
return super().redirect_request(req, fp, code, msg, headers, newurl)
def fetch(
url: str,
user_agent: str,
timeout: float = TIMEOUT_SECONDS,
max_bytes: int = MAX_BYTES,
) -> Response:
"""GET `url` and return the body exactly as received.
No cookies (the opener carries no cookie processor) and no content
decoding - urllib asks for `identity`, so the bytes are the page, not a
compressed transfer of it. `timeout` bounds the whole transfer, not only a
single read, so a server that trickles bytes cannot hold the call open.
Raises `ValidationError` for a URL outside `http`/`https` (also on a
redirect) and `BackendError` for everything the network does: an HTTP
error status, an unreachable host, a timeout, a body over `max_bytes`.
"""
check_url(url)
opener = urllib.request.build_opener(_HttpOnlyRedirects)
request = urllib.request.Request(url, headers={"User-Agent": user_agent, "Accept": "*/*"})
deadline = time.monotonic() + timeout
try:
with opener.open(request, timeout=timeout) as response:
declared = response.headers.get("Content-Length")
if declared and declared.strip().isdigit() and int(declared) > max_bytes:
raise BackendError(_too_large(url, max_bytes))
chunks: list[bytes] = []
size = 0
while True:
if time.monotonic() > deadline:
raise BackendError(f"{url} did not finish within {timeout:g} s.")
chunk = response.read(64 * 1024)
if not chunk:
break
size += len(chunk)
if size > max_bytes:
raise BackendError(_too_large(url, max_bytes))
chunks.append(chunk)
return Response(
url=url,
final_url=response.geturl(),
status=response.status,
content_type=response.headers.get("Content-Type", "") or "",
body=b"".join(chunks),
)
except urllib.error.HTTPError as exc:
raise BackendError(f"{url} answered HTTP {exc.code} {exc.reason}.") from exc
except urllib.error.URLError as exc:
if isinstance(exc.reason, TimeoutError):
raise BackendError(f"{url} did not answer within {timeout:g} s.") from exc
raise BackendError(f"Could not reach {url}: {exc.reason}") from exc
except TimeoutError as exc:
raise BackendError(f"{url} did not answer within {timeout:g} s.") from exc
except OSError as exc:
raise BackendError(f"Could not fetch {url}: {exc}") from exc
def _too_large(url: str, max_bytes: int) -> str:
return f"{url} is larger than {max_bytes // (1024 * 1024)} MiB - nothing was written."
def extension_for(content_type: str) -> str:
"""The file extension a non-HTML response is stored under. Read from
Python's built-in table only - `mimetypes.MimeTypes()` ignores the
machine's `/etc/mime.types`, so the answer does not depend on the host."""
kind = media_type(content_type)
if kind in PLAIN_TEXT_EXTENSIONS:
return PLAIN_TEXT_EXTENSIONS[kind]
if not kind:
return ".bin"
return mimetypes.MimeTypes().guess_extension(kind) or ".bin"
# --- character set ----------------------------------------------------------
_BOMS = (
(codecs.BOM_UTF8, "utf-8"),
(codecs.BOM_UTF32_LE, "utf-32-le"),
(codecs.BOM_UTF32_BE, "utf-32-be"),
(codecs.BOM_UTF16_LE, "utf-16-le"),
(codecs.BOM_UTF16_BE, "utf-16-be"),
)
# Covers both `<meta charset="x">` and
# `<meta http-equiv="Content-Type" content="text/html; charset=x">`.
_META_CHARSET = re.compile(rb"""<meta\b[^>]*?charset\s*=\s*["']?\s*([A-Za-z0-9_.:\-]+)""", re.IGNORECASE)
@dataclass(frozen=True)
class Decoded:
text: str
charset: str
source: str # bom | header | meta | default
replaced: int # how many U+FFFD the decoding introduced
def _known(label: Optional[str]) -> Optional[str]:
if not label:
return None
label = label.strip().lower()
try:
codecs.lookup(label)
except LookupError:
return None
return label
def decode(body: bytes, content_type: Optional[str]) -> Decoded:
"""Decode an HTML body, in a fixed order: byte-order mark, then the
`charset` of the HTTP `Content-Type` (`None` when there was no HTTP
response, as for a page saved from a browser), then a `<meta>` declaration
in the first 4 KiB, then UTF-8. A label Python does not know is skipped
like an absent one. Undecodable bytes are replaced, never dropped, and
counted, so the caller can say so."""
charset, source, start = None, "", 0
for bom, name in _BOMS:
if body.startswith(bom):
charset, source, start = name, "bom", len(bom)
break
if charset is None and content_type:
message = email.message.Message()
message["content-type"] = content_type
charset = _known(message.get_content_charset())
source = "header" if charset else ""
if charset is None:
match = _META_CHARSET.search(body[:META_SCAN_BYTES])
charset = _known(match.group(1).decode("ascii", "replace")) if match else None
source = "meta" if charset else ""
if charset is None:
charset, source = "utf-8", "default"
payload = body[start:]
try:
return Decoded(payload.decode(charset), charset, source, 0)
except UnicodeDecodeError:
text = payload.decode(charset, errors="replace")
before = payload.decode(charset, errors="ignore").count("�")
return Decoded(text, charset, source, text.count("�") - before)
# --- HTML -> text -----------------------------------------------------------
_VOID = frozenset(
"area base br col embed hr img input keygen link meta param source track wbr".split()
)
_DROPPED = frozenset(
"script style noscript nav header footer aside form template svg head title".split()
)
_BLOCK = frozenset(
"""address article blockquote body center dd details dialog div dl dt fieldset
figcaption figure h1 h2 h3 h4 h5 h6 hgroup hr html li main ol p pre section summary
table tbody thead tfoot tr td th caption ul""".split()
)
# An open `<p>` ends where one of these starts - the parser's share of the
# implied end tags real-world HTML relies on.
_CLOSES_P = frozenset(
"""address article blockquote div dl fieldset figure h1 h2 h3 h4 h5 h6 hr main ol p pre
section table ul""".split()
)
_HEADINGS = {f"h{n}": n for n in range(1, 7)}
_WS = re.compile(r"[ \t\n\r\f\v]+")
_BR = "\x00"
class _Element:
__slots__ = ("tag", "attrs", "children", "parent")
def __init__(self, tag: str, attrs: dict[str, str], parent: Optional["_Element"]):
self.tag = tag
self.attrs = attrs
self.children: list[Union["_Element", str]] = []
self.parent = parent
def iter(self):
yield self
for child in self.children:
if isinstance(child, _Element):
yield from child.iter()
def text(self) -> str:
return "".join(c if isinstance(c, str) else c.text() for c in self.children)
class _TreeBuilder(HTMLParser):
"""A forgiving tree: an end tag closes the nearest open element of that
name and is ignored when none is open, void elements never take children,
and `<p>`/`<li>`/`<dt>`/`<dd>`/`<tr>`/`<td>`/`<th>`/`<option>` end where
the next sibling of their kind begins."""
def __init__(self) -> None:
super().__init__(convert_charrefs=True)
self.root = _Element("#document", {}, None)
self.stack = [self.root]
def _close_to(self, tag: str, stop_at: tuple[str, ...] = ()) -> None:
for index in range(len(self.stack) - 1, 0, -1):
name = self.stack[index].tag
if name == tag:
del self.stack[index:]
return
if name in stop_at:
return
def handle_starttag(self, tag, attrs):
if tag in _CLOSES_P:
self._close_to("p", stop_at=("div", "li", "td", "th", "blockquote", "section", "article", "main", "body"))
if tag == "li":
self._close_to("li", stop_at=("ul", "ol"))
elif tag in ("dt", "dd"):
self._close_to("dt", stop_at=("dl",))
self._close_to("dd", stop_at=("dl",))
elif tag == "tr":
self._close_to("tr", stop_at=("table",))
elif tag in ("td", "th"):
self._close_to("td", stop_at=("tr", "table"))
self._close_to("th", stop_at=("tr", "table"))
elif tag == "option":
self._close_to("option", stop_at=("select",))
element = _Element(tag, {k: (v or "") for k, v in attrs}, self.stack[-1])
self.stack[-1].children.append(element)
if tag not in _VOID:
self.stack.append(element)
def handle_startendtag(self, tag, attrs):
element = _Element(tag, {k: (v or "") for k, v in attrs}, self.stack[-1])
self.stack[-1].children.append(element)
def handle_endtag(self, tag):
if tag not in _VOID:
self._close_to(tag)
def handle_data(self, data):
self.stack[-1].children.append(data)
def parse(text: str) -> _Element:
builder = _TreeBuilder()
builder.feed(text)
builder.close()
return builder.root
def title_of(document: _Element) -> str:
for element in document.iter():
if element.tag == "title":
return _WS.sub(" ", element.text()).strip()
return ""
def content_root(document: _Element) -> _Element:
"""`<main>`, else the one `<article>` if there is exactly one, else
`<body>`, else the whole document."""
elements = list(document.iter())
for element in elements:
if element.tag == "main":
return element
articles = [e for e in elements if e.tag == "article"]
if len(articles) == 1:
return articles[0]
for element in elements:
if element.tag == "body":
return element
return document
@dataclass
class _Block:
kind: str # "text" or "list"
text: str
class _Renderer:
def __init__(self, base_url: str) -> None:
self.base_url = base_url
def url(self, href: str) -> str:
return urllib.parse.urljoin(self.base_url, href.strip())
# Blocks -------------------------------------------------------------
def blocks(self, element: _Element) -> list[_Block]:
"""The blocks of a container: runs of inline content become one
paragraph each, block children contribute their own blocks."""
out: list[_Block] = []
inline: list[str] = []
def flush() -> None:
text = _finish_inline("".join(inline))
if text:
out.append(_Block("text", text))
inline.clear()
for child in element.children:
if isinstance(child, str):
inline.append(child)
elif child.tag in _DROPPED:
continue
elif child.tag in _BLOCK:
flush()
out.extend(self.block(child))
else:
inline.append(self.inline(child))
flush()
return out
def block(self, element: _Element) -> list[_Block]:
tag = element.tag
if tag in _HEADINGS:
text = _finish_inline(self.inline_children(element))
return [_Block("text", "#" * _HEADINGS[tag] + " " + text)] if text else []
if tag == "p":
text = _finish_inline(self.inline_children(element))
return [_Block("text", text)] if text else []
if tag == "pre":
return self.pre(element)
if tag in ("ul", "ol"):
return self.list(element)
if tag == "blockquote":
inner = _join(self.blocks(element))
if not inner:
return []
return [_Block("text", "\n".join(f"> {line}" if line else ">" for line in inner.split("\n")))]
if tag == "hr":
return [_Block("text", "---")]
if tag == "table":
return self.table(element)
return self.blocks(element)
def pre(self, element: _Element) -> list[_Block]:
text = element.text().strip("\n")
if not text.strip():
return []
longest = max((len(run) for run in re.findall(r"`+", text)), default=0)
fence = "`" * max(3, longest + 1)
return [_Block("text", f"{fence}\n{text}\n{fence}")]
def list(self, element: _Element) -> list[_Block]:
ordered = element.tag == "ol"
start = element.attrs.get("start", "").strip()
number = int(start) if ordered and start.lstrip("-").isdigit() else 1
lines: list[str] = []
for child in element.children:
if isinstance(child, str) or child.tag in _DROPPED:
continue
if child.tag == "li":
content = self.blocks(child)
elif child.tag in ("ul", "ol"):
# A list nested directly in a list, without its own <li>.
content = self.block(child)
else:
continue
if not content:
continue
marker = f"{number}. " if ordered else "- "
number += 1
body = ""
for index, block in enumerate(content):
if index:
body += "\n" if block.kind == "list" else "\n\n"
body += block.text
indent = " " * len(marker)
lines.append("\n".join(
(marker if i == 0 else (indent if line else "")) + line
for i, line in enumerate(body.split("\n"))
))
return [_Block("list", "\n".join(lines))] if lines else []
def table(self, element: _Element) -> list[_Block]:
rows: list[list[str]] = []
def collect(node: _Element) -> None:
for child in node.children:
if isinstance(child, str) or child.tag in _DROPPED or child.tag == "table":
continue
if child.tag == "tr":
cells = [
_finish_inline(self.inline_children(cell)).replace("\n", " ").replace("|", "\\|")
for cell in child.children
if isinstance(cell, _Element) and cell.tag in ("td", "th")
]
if any(cells):
rows.append(cells)
else:
collect(child)
collect(element)
if not rows:
return []
width = max(len(row) for row in rows)
padded = [row + [""] * (width - len(row)) for row in rows]
lines = ["| " + " | ".join(padded[0]) + " |", "|" + " --- |" * width]
lines += ["| " + " | ".join(row) + " |" for row in padded[1:]]
return [_Block("text", "\n".join(lines))]
# Inline -------------------------------------------------------------
def inline_children(self, element: _Element) -> str:
parts = []
for child in element.children:
if isinstance(child, str):
parts.append(child)
elif child.tag in _DROPPED:
continue
else:
part = self.inline(child)
# A block element met in inline context still separates words.
parts.append(f" {part} " if child.tag in _BLOCK else part)
return "".join(parts)
def inline(self, element: _Element) -> str:
tag = element.tag
if tag in _DROPPED:
return ""
if tag == "br":
return _BR
if tag == "img":
alt = _WS.sub(" ", element.attrs.get("alt", "")).strip()
src = element.attrs.get("src", "").strip()
if not src or src.lower().startswith("data:"):
return alt
return f"![{alt}]({self.url(src)})"
if tag == "code":
text = _WS.sub(" ", element.text()).strip()
if not text:
return ""
longest = max((len(run) for run in re.findall(r"`+", text)), default=0)
ticks = "`" * (longest + 1)
pad = " " if text.startswith("`") or text.endswith("`") else ""
return f"{ticks}{pad}{text}{pad}{ticks}"
if tag == "a":
text = _WS.sub(" ", self.inline_children(element)).strip()
href = element.attrs.get("href", "").strip()
if not href or href.startswith("#") or href.lower().startswith("javascript:"):
return text
if not text:
return ""
return f"[{text}]({self.url(href)})"
return self.inline_children(element)
def _finish_inline(text: str) -> str:
text = _WS.sub(" ", text)
lines = [line.strip() for line in text.split(_BR)]
return "\n".join(lines).strip("\n").strip()
def _join(blocks: list[_Block]) -> str:
return "\n\n".join(block.text for block in blocks)
def derive_text(html: str, base_url: str) -> str:
"""The reading text of an HTML document: the content root's blocks as
Markdown-like text, with links made absolute against `base_url`."""
document = parse(html)
renderer = _Renderer(base_url)
root = content_root(document)
blocks = renderer.block(root) if root.tag != "#document" else renderer.blocks(root)
return _join(blocks)
# --- the capture file -------------------------------------------------------
def utc_now() -> str:
return datetime.datetime.now(datetime.timezone.utc).replace(microsecond=0).strftime("%Y-%m-%dT%H:%M:%SZ")
_PLAIN_SCALAR = re.compile(r"[^\s\-?:,\[\]{}#&*!|>'\"%@`][^\n]*")
def _scalar(value: str) -> str:
"""A YAML scalar for one header value: plain where YAML reads it back
unchanged, double-quoted (a JSON string is valid YAML) otherwise - a page
title with `: ` in it would end the block's parseability."""
if value and _PLAIN_SCALAR.fullmatch(value) and ": " not in value and " #" not in value and value == value.strip():
return value
return json.dumps(value, ensure_ascii=False)
def header(fields: list[tuple[str, str]]) -> str:
lines = ["---"] + [f"{key}: {_scalar(value)}" for key, value in fields] + ["---"]
return "\n".join(lines) + "\n"
@dataclass(frozen=True)
class Derivation:
markdown: bytes # header + derived text, UTF-8, LF
text_chars: int
title: str
decoded: Decoded
def derive_document(
body: bytes,
base_url: str,
html_name: str,
*,
response: Optional[Response] = None,
url: Optional[str] = None,
derived_at: Optional[str] = None,
) -> Derivation:
"""The `.md` beside a captured `.html`: the fixed header, then the derived
text. With `response` it describes a fetch; without, a page a human saved
(`url` and `derived_at` then fill the header instead)."""
decoded = decode(body, response.content_type if response else None)
document_title = title_of(parse(decoded.text))
text = derive_text(decoded.text, base_url)
charset = f"{decoded.charset} (from: {decoded.source})"
if response is not None:
fields = [
("fetched_by", "wikitool raw fetch"),
("url", response.url),
("final_url", response.final_url),
("retrieved", derived_at or utc_now()),
("http_status", str(response.status)),
("content_type", response.content_type),
("charset", charset),
("title", document_title),
("derived_from", html_name),
]
else:
fields = [
("fetched_by", "wikitool raw fetch --html"),
("url", url or ""),
("derived", derived_at or utc_now()),
("charset", charset),
("title", document_title),
("derived_from", html_name),
]
content = header(fields) + "\n" + (text + "\n" if text else "")
return Derivation(content.encode("utf-8"), len(text), document_title, decoded)
# --- naming -----------------------------------------------------------------
def slug(value: str) -> str:
value = unicodedata.normalize("NFKD", value).encode("ascii", "ignore").decode("ascii").lower()
value = re.sub(r"[^a-z0-9]+", "-", value).strip("-")
return value[:STEM_MAX_CHARS].rstrip("-")
def stem_for(url: str) -> str:
"""The last non-empty path segment without its extension, else the host
name - ASCII, lowercase, `-`-separated, at most 60 characters."""
parts = urllib.parse.urlsplit(url)
segments = [s for s in urllib.parse.unquote(parts.path).split("/") if s]
if segments:
last = segments[-1]
base = last.rsplit(".", 1)[0] if "." in last.strip(".") else last
candidate = slug(base)
if candidate:
return candidate
return slug(parts.hostname or "") or "page"