Files
chemenu/instructions/private-instance.md
T
torben b2f7dec122
CI / verify (push) Successful in 43s
Release / release (push) Successful in 38s
fix: private-instance - Demo-Korpus wandert beim Merge mit, Prozedur korrigiert (2.2.1)
Files changed:
- CHANGES.md
- VERSION
- instructions/private-instance.md
2026-09-01 19:03:23 +02:00

172 lines
7.8 KiB
Markdown

---
type: types/instruction.md
name: private-instance
description: Set up a private working instance as a clone of a public upstream, so stack updates arrive by merge instead of by copying a tarball over the tree.
---
# Set up a private instance against a public upstream
The distribution path in [setup-instance.md](setup-instance.md) builds an instance from a
`dist export` tarball, with no git ancestry in common with the repo it came from. That is the
right shape for someone who only ever *consumes* the stack.
This is the other shape: a private instance that keeps taking stack changes from a public
upstream, and whose own content must never travel back. It costs one safeguard to set up and
saves the whole update procedure afterwards.
**Read this before, not after, the first `publish`.** The gate in step 4 is the thing that makes
the arrangement safe, and adding it later means the window it closes was open in between.
## Why a clone rather than a tarball
`INSTALL.md`'s "Eine Instanz aktualisieren" is `cp -r` as an upgrade strategy: copy `tools/`,
`types/`, `instructions/`, `AGENTS.md`, `VERSION` over the existing tree. It has no three-way
merge, so it cannot notice that the receiving instance changed a file, and it has no conflict
surface, so nobody learns when upstream and local both touched the same one. It overwrites
silently.
A clone gets all of that from git. Stack changes land as real merges, with real conflicts where
they conflict.
**What a plain `git merge` does *not* give you is protection from the upstream's content.** The
private `main` deletes the demo corpus once, but that deletion does not make later upstream
changes to those paths go away. Measured, not assumed:
| Upstream does | `git merge upstream/main` does |
|---|---|
| modifies a page you deleted | `CONFLICT (modify/delete)` - and **leaves the upstream version in your working tree**. Resolve it with `git add -A` and the demo page is back. |
| adds a new page | stages it **silently**. No conflict, no prompt, no mention. |
| deletes a page you also deleted | nothing. The only harmless case. |
The middle row is the one that matters, because nothing announces it. An upstream that ships a
demo corpus *and* uses it as a test bed will add pages, and each one arrives in your instance
and starts showing up in your `lint`, your `index`, your `search` and your `confidence decay`.
So the merge has to be scoped. That is the procedure below, and it is not optional.
## Steps
1. **Clone, and name the two remotes for what they are.**
```bash
git clone <private-repo-url> my-wiki
cd my-wiki
git remote add upstream <public-repo-url>
```
`origin` is yours and is the only thing you ever push to. `upstream` is where stack updates
come from and is fetch-only.
2. **Make the fetch-only half fetch-only in git, too.**
```bash
git remote set-url --push upstream no_push
```
git refuses to push to a URL it cannot resolve. This is a convenience, not the safeguard -
step 4 is the safeguard.
3. **Delete the upstream's demo corpus once, on your own `main`.**
Everything under `kb/` and `raw/` that came with the clone is the upstream's content, not
yours. Remove it with `wikitool rm --page` (never `rm -rf`: `rm` de-links each page from the
rest of the wiki, and a plain delete leaves dead wikilinks and broken citations behind), then
`index rebuild`, `sources rebuild-index`, `lint`.
This is a one-time cut. Afterwards the upstream corpus is frozen from your side, which is
what makes later merges content-free.
4. **Arm the Publish-Remote Gate — before the first `publish`.**
```bash
cat > .wikitool-remotes.json <<'EOF'
{ "schema": 1, "allowed_push_urls": ["<your-private-push-url>"] }
EOF
```
Use the URL `git remote get-url --push origin` prints, exactly. `publish` refuses with exit
42 for anything else, and there is no flag that opens it - see [gates.md](gates.md).
The file is gitignored, so it stays with this checkout and never travels to the upstream.
`wikitool doctor` reports whether the gate is armed, and WARNs at more than one remote
without it.
5. **Take away the write credential, if you can.** A token or deploy key for `origin` only,
with no write access to the upstream, is the one control that holds even if everything above
is misconfigured. Belt and braces.
6. **Personalize and bootstrap.** `USER.md`, `SOUL.md` and optionally `ENVIRONMENT.md` are
yours and unrelated to the upstream's - see the Personalization step of
[setup-instance.md](setup-instance.md), then [bootstrap.md](bootstrap.md) for the venv and
the skills.
## Taking a stack update
Take the machinery, never the content. The merge is held open, the content stages are forced
back to your own state, and only then does it close:
```bash
BEFORE=$(git rev-parse HEAD)
git fetch upstream
# --no-commit holds the merge open; it may report conflicts under kb/ or raw/,
# which the next three lines are about to make irrelevant.
git merge --no-commit --no-ff upstream/main || true
# Whatever the merge did to the content stages, undo it. HEAD is still your
# pre-merge commit while the merge is open, so this restores exactly your side.
git rm -rq --cached --ignore-unmatch kb raw
rm -rf kb raw
git checkout HEAD -- kb raw
git commit --no-edit
```
Then **check that it worked**, rather than trusting that it did:
```bash
git diff --name-only $BEFORE HEAD -- kb raw # must print nothing
```
An empty result is the proof that the update touched machinery only. A non-empty one means a
path slipped through - inspect it before going further.
Then, as after any stack change: `doctor`, `docs verify`, `instructions verify`, `migrate status`,
`lint`. A `migrate status` with outstanding links means the update crossed a compatibility
boundary - follow [migrate-corpus.md](migrate-corpus.md) before doing anything else.
**Why not just `git merge upstream/main`?** Because of the table above: a page the upstream
*adds* arrives with no conflict and no message. You would find out when `lint` starts reporting
pages you never wrote - if you noticed at all.
## Where stack development happens
**In the public repo, not here.** That is not a preference; the stack is built that way. The
development-only half of the instruction layer is pruned from a distribution one-way, with no
command that reconstructs it, so an instance built this way has no tool-development mode to
switch into in the first place.
When a tool bug blocks real content work here - and it will - file the issue against the public
repo (an MCP server or the web UI reaches it from any session; no shared history needed), fix it
there where the tests, `docs verify` and CI's version gate live, and take the fix back with the
merge above. Nothing is lost by the detour: the fix has to pass that CI either way.
## Decision points
- **Merge conflict in `kb/` or `raw/`?** Expected, and already handled: the update procedure
above overwrites those stages with your own afterwards, so the conflict resolves itself.
Never resolve one by hand with `git add -A` - that is exactly how the upstream version, which
git left sitting in your working tree, gets committed into your instance.
- **`git diff` after the merge shows something under `kb/` or `raw/`?** Stop. The scoping step
did not take. Do not publish; find out which path came through and where from.
- **Conflict in `tools/`, `types/` or `instructions/`?** You changed the stack locally, which
step "Where stack development happens" says not to do. Take the upstream side and re-file the
change as an issue there.
## Scope
Not for a first instance with no upstream - that is [setup-instance.md](setup-instance.md). Not
for a fresh clone of a repo you already own and develop in - that is
[bootstrap.md](bootstrap.md). This is specifically the two-remote case, where the cost of a
mistaken push is disclosure rather than inconvenience.