Files
chemenu/instructions/private-instance.md
T
torben 2c4c2b1c7c
CI / verify (push) Failing after 57s
Release / release (push) Successful in 37s
stack: TOC-Pflicht fuer Referenzdateien ueber 100 Zeilen (docs toc); session-setup.md/gates.md nennen die tatsaechliche Budget-Ausnahmeliste (schliesst #73, #76)
Files changed:
- AGENTS.md
- CHANGES.md
- VERSION
- instructions/CONTRACT.md
- instructions/capture-session.md
- instructions/claude-code-model-selection.md
- instructions/dev/issue-tracking.md
- instructions/dev/testing-conventions.md
- instructions/dev/version-parts.md
- instructions/evolve-subtypes.md
- instructions/gates.md
- instructions/german-terminology.md
- instructions/ingest-large-tree.md
- instructions/kb-profiles.md
- instructions/link-taxonomy.md
- instructions/mcp-read-server.md
- instructions/migrate-corpus.md
- instructions/migrations/3.0.0-authoring-conventions.md
- instructions/migrations/4.0.0-link-taxonomy.md
- instructions/private-instance.md
- instructions/session-setup.md
- instructions/setup-instance.md
- kb/CONTRACT.md
- kb/CONVENTIONS.md
- kb/concepts/COLLECTION.md
- raw/CONTRACT.md
- tools/CONTRACT.md
- tools/chemenu/commands/docs_verify.py
- tools/chemenu/commands/instructions_cmd.py
- tools/chemenu/tests/test_docs_verify.py
- tools/chemenu/tests/test_instructions_cmd.py
- tools/chemenu/tests/test_toc.py
- tools/chemenu/toc.py
- types/type-spec.md
2026-09-09 20:38:42 +02:00

213 lines
11 KiB
Markdown

---
type: types/instruction.md
name: private-instance
description: Set up a private working instance as a clone of a public upstream, so stack updates arrive by merge instead of by copying a tarball over the tree.
---
# Set up a private instance against a public upstream
The distribution path in [setup-instance.md](setup-instance.md) builds an instance from a
`dist export` tarball, with no git ancestry in common with the repo it came from. That is the
right shape for someone who only ever *consumes* the stack.
This is the other shape: a private instance that keeps taking stack changes from a public
upstream, and whose own content must never travel back. It costs one safeguard to set up and
saves the whole update procedure afterwards.
**Read this before, not after, the first `publish`.** The gate in step 4 is the thing that makes
the arrangement safe, and adding it later means the window it closes was open in between.
<!-- wikitool:toc -->
## Contents
- [Why a clone rather than a tarball](#why-a-clone-rather-than-a-tarball)
- [Steps](#steps)
- [Taking a stack update](#taking-a-stack-update)
- [Where stack development happens](#where-stack-development-happens)
- [Decision points](#decision-points)
- [Scope](#scope)
<!-- /wikitool:toc -->
## Why a clone rather than a tarball
`INSTALL.md`'s "Eine Instanz aktualisieren" is `cp -r` as an upgrade strategy: copy `tools/`,
`types/`, `instructions/`, `AGENTS.md`, `VERSION` over the existing tree. It has no three-way
merge, so it cannot notice that the receiving instance changed a file, and it has no conflict
surface, so nobody learns when upstream and local both touched the same one. It overwrites
silently.
A clone gets all of that from git. Stack changes land as real merges, with real conflicts where
they conflict.
**What a plain `git merge` does *not* give you is protection from the upstream's content.** The
private `main` deletes the demo corpus once, but that deletion does not make later upstream
changes to those paths go away. Measured, not assumed:
| Upstream does | `git merge upstream/main` does |
|---|---|
| modifies a page you deleted | `CONFLICT (modify/delete)` - and **leaves the upstream version in your working tree**. Resolve it with `git add -A` and the demo page is back. |
| adds a new page | stages it **silently**. No conflict, no prompt, no mention. |
| deletes a page you also deleted | nothing. The only harmless case. |
The middle row is the one that matters, because nothing announces it. An upstream that ships a
demo corpus *and* uses it as a test bed will add pages, and each one arrives in your instance
and starts showing up in your `lint`, your `index`, your `search` and your `confidence decay`.
So the merge has to be scoped. That is the procedure below, and it is not optional.
## Steps
1. **Clone, and name the two remotes for what they are.**
```bash
git clone <private-repo-url> my-wiki
cd my-wiki
git remote add upstream <public-repo-url>
```
`origin` is yours and is the only thing you ever push to. `upstream` is where stack updates
come from and is fetch-only.
2. **Make the fetch-only half fetch-only in git, too.**
```bash
git remote set-url --push upstream no_push
```
git refuses to push to a URL it cannot resolve. This is a convenience, not the safeguard -
step 4 is the safeguard.
3. **Delete the upstream's demo corpus once, on your own `main`.**
Everything under `kb/` and `raw/` that came with the clone is the upstream's content, not
yours. Remove it with `wikitool rm --page` (never `rm -rf`: `rm` de-links each page from the
rest of the wiki, and a plain delete leaves dead wikilinks and broken citations behind), then
`index rebuild`, `sources rebuild-index`, `lint`.
This is a one-time cut. Afterwards the upstream corpus is frozen from your side, which is
what makes later merges content-free.
4. **Arm the Publish-Remote Gate — before the first `publish`.**
```bash
cat > .wikitool-remotes.json <<'EOF'
{ "schema": 1, "allowed_push_urls": ["<your-private-push-url>"] }
EOF
```
Use the URL `git remote get-url --push origin` prints, exactly. `publish` refuses with exit
42 for anything else, and there is no flag that opens it - see [gates.md](gates.md).
The file is gitignored, so it stays with this checkout and never travels to the upstream.
`wikitool doctor` reports whether the gate is armed, and WARNs at more than one remote
without it.
5. **Take away the write credential, if you can.** A token or deploy key for `origin` only,
with no write access to the upstream, is the one control that holds even if everything above
is misconfigured. Belt and braces.
6. **Personalize and bootstrap.** `USER.md`, `SOUL.md` and optionally `ENVIRONMENT.md` are
yours and unrelated to the upstream's - see the Personalization step of
[setup-instance.md](setup-instance.md), then [bootstrap.md](bootstrap.md) for the venv and
the skills.
A clone inherits the upstream's `kb/CONVENTIONS.md` and `kb/*/COLLECTION.md` rather than
templates, because it inherits the upstream's whole tree. They are yours from this point on:
rewrite them if this instance writes its pages differently - the update procedure below
restores them on every merge, so the change sticks. [kb-profiles.md](kb-profiles.md) has the
alternatives.
## Taking a stack update
```bash
tools/wikitool upstream merge --remote upstream --branch main
```
Take the machinery, never the content. This is the command form of the same idea a hand-rolled
merge would need: hold the merge open, force the content stages back to your own state, restore
only the paths that are machinery, and only then let it close. Which paths those are is not a
short literal list any more (see below) - it is `chemenu.ownership.is_stack_owned`, the same
predicate `dist_cmd.py`'s export reads, so a stack change that adds a new machinery path under a
content stage is recognised automatically rather than needing this document edited first.
**What counts as machinery under a content stage**, for readers who want the shape rather than
the code:
| Path | Why it takes the upstream side |
|---|---|
| `<stage>/CONTRACT.md` (`kb/CONTRACT.md`, `raw/CONTRACT.md`, `work/CONTRACT.md`, `reports/CONTRACT.md`) | The stack's own stage contract. Every rule in it is enforced by `wikitool`; an instance never edits it |
| any `*.template` under a content stage (`kb/CONVENTIONS.md.template`, each `kb/<name>/COLLECTION.md.template`, and any later one) | The template your filled file was adopted from. The filled file is yours; the template is the stack's |
Everything else under `kb/`, `raw/`, `work/` and `reports/` is yours, `kb/CONVENTIONS.md` and
each `kb/<name>/COLLECTION.md` included - they bind your corpus, and they are exactly what
`upstream merge` protects.
**Your local, uncommitted-by-design files under those stages survive.** Forcing a content stage
back to your own state removes only what git tracks, never the directory wholesale - which
matters because `reports/` is gitignored apart from its contract, so it holds data that is in no
commit and cannot be recomputed: the telemetry traces `eval score` reads, saved eval reports,
past lint reports. A merge has no business touching any of it, and does not.
The command itself checks its own result the same way `upstream verify` would, immediately
after committing, and refuses loudly - without rolling the commit back - if anything landed
outside a stack-owned path. A refusal there is a bug report, not something to work around by
hand; see [tools/CONTRACT.md](../tools/CONTRACT.md) for the full error contract, including what
a real conflict in `tools/`/`types/`/`instructions/` leaves behind.
Then, as after any stack change: `doctor`, `docs verify`, `instructions verify`, `migrate status`,
`lint`. A `migrate status` with outstanding links means the update crossed a compatibility
boundary - follow [migrate-corpus.md](migrate-corpus.md) before doing anything else.
**Why not just `git merge upstream/main`?** A page the upstream *adds* arrives with no conflict
and no message under a plain merge - measured in the table further up this document. You would
find out when `lint` starts reporting pages you never wrote, if you noticed at all. `upstream
merge` closes exactly that gap: the content stages never see the upstream's version at all.
**Checking a merge you resolved by hand instead** (or auditing a past one): `tools/wikitool
upstream verify --since <rev-before> --until <rev-after>` runs the same check `upstream merge`
runs on itself, without doing the merge.
## Where stack development happens
**In the public repo, not here.** That is not a preference; the stack is built that way. The
development-only half of the instruction layer is pruned from a distribution one-way, with no
command that reconstructs it, so an instance built this way has no tool-development mode to
switch into in the first place.
When a tool bug blocks real content work here - and it will - file the issue against the public
repo (an MCP server or the web UI reaches it from any session; no shared history needed), fix it
there where the tests, `docs verify` and CI's version gate live, and take the fix back with the
merge above. Nothing is lost by the detour: the fix has to pass that CI either way.
## Decision points
- **Merge conflict in `kb/`, `raw/`, `work/` or `reports/`?** Expected, and already handled:
`upstream merge` overwrites those stages with your own afterwards, so the conflict resolves
itself. Never resolve one by hand with `git add -A` in a merge you are running yourself
instead - that is exactly how the upstream version, which git left sitting in your working
tree, gets committed into your instance.
- **`upstream merge` exits 1 after committing?** Read the message: its own postcheck found
content outside a stack-owned path in the commit it just made. The commit is **not** rolled
back - inspect it (`git show`, or `tools/wikitool upstream verify --since <before> --until
HEAD`) and decide by hand whether to revert it, fix forward, or report it as a stack bug. This
should not happen; if it does, `chemenu.ownership.is_stack_owned` disagreed with itself between
the restore and the check, which is exactly what the shared predicate is meant to prevent.
- **Conflict in `tools/`, `types/` or `instructions/`?** You changed the stack locally, which
step "Where stack development happens" says not to do. `upstream merge` leaves the merge open
rather than guessing - take the upstream side for the named paths and re-file the change as an
issue there, or resolve deliberately and finish the commit yourself.
- **...but you changed how *your pages* are written?** That is not a stack change and the rule
above does not apply to it. Language, section headings, naming forms, tone, relationship
labels and the confidence rubric live in `kb/CONVENTIONS.md`, and each collection's authoring
rules in `kb/<name>/COLLECTION.md` - all under `kb/`, all yours, all restored by the merge
procedure rather than overwritten by it. If you find yourself editing `tools/` or `types/` to
change an authoring convention, that is a stack bug: file it, because the split exists
precisely so you do not have to.
## Scope
Not for a first instance with no upstream - that is [setup-instance.md](setup-instance.md). Not
for a fresh clone of a repo you already own and develop in - that is
[bootstrap.md](bootstrap.md). This is specifically the two-remote case, where the cost of a
mistaken push is disclosure rather than inconvenience.