@kontextmind/kxm 0.7.51 → 0.7.53

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -11,7 +11,7 @@
11
11
  "name": "kxm",
12
12
  "source": "./plugins/kxm",
13
13
  "description": "Durable workflows, peer agents, and kxm tui",
14
- "version": "0.7.51",
14
+ "version": "0.7.53",
15
15
  "category": "development",
16
16
  "tags": ["kxm", "multi-agent", "workflows", "mcp"]
17
17
  }
package/CHANGELOG.md CHANGED
@@ -6,17 +6,69 @@ All notable user-facing changes are documented here. The project follows [Semant
6
6
 
7
7
  ### Added
8
8
 
9
- - **Terminal component kit package:** `@kontextmind/tui` (`packages/core/tui`, also
10
- exposed as the `@kontextmind/kxm/tui` export) ships the reusable Pi-renderer-based
11
- terminal components in the package layer shape (`src/{types,tui,services,adapters,exports}`,
12
- `tests/{unit,helpers}`), enforced by `test/core/package-layers.test.ts`.
13
- - **Per-package workspace tooling:** nx + Bun workspace wiring (`nx.json`,
14
- `bunfig.toml`, `packages/*/project.json`, `scripts/build-package.mjs`) with
15
- `npm run build:packages|test:packages|check:packages`. Bun is installer and task
16
- runner only; tests, the hub, and the CLI remain on Node.
9
+ - **Per-tenant hosting operations, and a backup that covers the whole state set.**
10
+ `docs/operations.md` gained a *Per-tenant hosted deployment* section (one tenant = one box
11
+ = one hub; loopback-only hub and supervisor; the tenant's proxy owns TLS and the browser
12
+ session; a reverse-proxy contract states what must hold without shipping generated proxy
13
+ config) and a rewritten *Backup and restore* that enumerates tenant state **by root**,
14
+ because the roots are not interchangeable and two of them are easy to mistake for one:
15
+ `$R` the fixed checkout `.kxm` tree (project, model, role, route, price and provenance
16
+ definition, environment, goals, tasks, memory, improvement candidates, skills — none of
17
+ which follow any workspace override); `$D` the workspace directories (config, logs,
18
+ assets, state — moved by `--workspace`/`KXM_WORKSPACE_DIR`, each individually
19
+ overridable, and `--workspace` ignores those overrides); `$W` the workspace state
20
+ directory (hub database, worker routing and recovery manifests, Pi sessions); `$S`
21
+ host-local machine state (`KXM_STATE_HOME`, absolute or rejected: Runtime `registry.db`
22
+ holding the supervisor claim as a registry row, per-project `run-events.db` **and its
23
+ full-filename `.run-prompts.json` sidecar**, repository bindings, `update.yaml`, hub
24
+ binding and credential records); `$C` user configuration; and `$T` federated telemetry,
25
+ resolved independently of `$C` and written by an exporter that currently has no
26
+ production caller. It separates recovery-critical manifests from Pi model histories whose
27
+ backup is an existing policy **choice**, marks what is disposable (PID and claim files,
28
+ `session-brief.json`, the re-generable supervisor token), applies WAL consistency to
29
+ every SQLite store, records each override as part of the backup, and notes that a bound
30
+ member repository's own `.kxm/repo/` files live on that member's filesystem — so copying
31
+ the binding JSON alone is not a restore path for an external member. The previous recipe
32
+ stopped the hub and copied `kxm.db`, which is a hub-only backup: a restore can pass every
33
+ hub check and still lose run history, the prompts that explain it, the project definition,
34
+ and the bindings that make the box reproducible.
17
35
 
18
36
  ### Changed
19
37
 
38
+ - **No schema migration lanes, no legacy stamp tolerance (single-operator tool).** Stepwise
39
+ schema migration lanes are removed from the hub store and the per-project event store, and
40
+ the external-effects store no longer runs an in-place `ALTER TABLE ... ADD COLUMN` whose
41
+ failure it swallowed (that column is declared in the fresh schema, so the lane could only
42
+ ever fire on an older store; the store has no production caller yet, so this closes the
43
+ lane rather than fixing a live upgrade).
44
+ A database stamped behind this build now fails closed with `runtime_schema_outdated` and
45
+ a message that says to delete the state file or re-run `kxm init` — and the refusal never
46
+ advances `user_version`, so the store stays identifiably old (WAL sidecars may still be
47
+ checkpointed by opening the file, so the whole state set remains the backup unit — see
48
+ [`docs/operations.md`](docs/operations.md)). Relabelling a store it refused to open would
49
+ only hide the problem until a query hit a missing column. The coordinator fingerprint no
50
+ longer recomputes over stored authority to forgive rows written before set canonicalisation:
51
+ a stale coordinator is re-bound. Intake tests go from 23 to 22; the two forced-race tests
52
+ now win against ordinary current-format rows, and one refusal test asserts the untouched
53
+ stamp.
54
+
55
+ - **`kxm hub bind` no longer stores a remote URL it cannot authenticate to.** A remote
56
+ binding is a deliberate network decision, so it is now refused when no credential resolves
57
+ (explicit `KXM_AUTH_TOKEN`, or the persisted hub record's admin or project tokens) —
58
+ mirroring the rule the hub already applies to its own listener, which refuses to bind
59
+ beyond loopback without a token. The refusal carries `nextAction` and the same hint string
60
+ in both the JSON payload and the prose line, so `--json` consumers are not left with a bare
61
+ code and no way forward; a stored-but-unusable binding otherwise reads later like a
62
+ network fault and gets debugged as one.
63
+ - **`kxm hub view` and the session brief label the binding `loopback` or `remote`.**
64
+ "Attached across a network" and "attached on this box" looked identical before, and only
65
+ one of them puts a bearer on a wire. `localhost`, `127.0.0.1`, `::1` and `*.localhost` are
66
+ loopback; `0.0.0.0`, LAN addresses and host names are remote. The **bind** guard is
67
+ scoped to remote URLs — a damaged host record must not cost a local operator their start,
68
+ and the first cut of the guard did exactly that. That is a statement about `hub bind`
69
+ only: other client paths resolve credentials whatever the scope, so loopback commands can
70
+ still fail on a malformed record.
71
+
20
72
  - **Naming sweep:** the retired `vnext` naming is gone from file and folder names,
21
73
  symbols, constants, schema `$id` segments, and error codes (`vnext_*` is now
22
74
  `initialization_failed`, `initialization_io_failed`, `wait_failed`); package and
@@ -296,7 +296,7 @@ The current hub command groups are `agent`, `session`, `workflow`, `gate`, `hub`
296
296
  | `kxm gate github watch` | Poll required GitHub checks and post the existing signed signal |
297
297
  | `kxm init` | Create or validate a project; configuration remains project-owned and no package dogfood templates are copied |
298
298
  | `kxm hub view` | Check `/health` and `/ready` |
299
- | `kxm hub bind <url>` | Bind this machine to a running hub |
299
+ | `kxm hub bind <url>` | Bind this machine to a running hub. A **remote** URL is refused unless a credential resolves (`KXM_AUTH_TOKEN`, or the persisted hub record); `kxm hub view` reports the binding as `loopback` or `remote` |
300
300
  | `kxm hub unbind` | Remove this machine's hub binding |
301
301
  | `kxm update --check` / `kxm update --kxm` | Check or apply a kxm operator package update from an npm-global install only. Other install kinds (source checkout, Pi git, Claude marketplace, npm-local, unknown) refuse `--kxm` and skip auto-apply. Source checkouts neither fetch nor nag. Default source is GitHub release tarballs; the release asset `kxm-<v>.tgz` must carry a sha256 digest or install fails closed. `source: npm` is for after the public package exists. Optional per-user `update.yaml` (`kxm.update.v1`, `auto` boolean) under the host state root (`KXM_STATE_HOME` / `%LOCALAPPDATA%\KXM` / macOS Application Support / XDG state) enables auto-apply on `kxm update`. A project `.kxm/update.yaml` is ignored with a warning. Notice also prints on `kxm hub start` (not from source) and on the session widget from cache |
302
302
  | `kxm dash` | Open the read-only SSE observer dashboard; non-TTY output is one ANSI-free snapshot |
@@ -16,7 +16,7 @@ Presence of KXM documents does not activate new behavior.
16
16
  | Long-lived manually started Pi workers | Runtime-managed run-scoped sessions | Existing worker mode remains available during compatibility release |
17
17
  | Shared/off workflow Pi history | `{run, agent, instance, scopeEpoch}` sessions | Never import shared conversation history into a narrower run scope |
18
18
  | Hub-owned workflow state | Home Runtime event log with hub projection | Import completed history as legacy records; active-run cutover requires quiescence |
19
- | SQLite schema v3 `kxm.db` | Runtime registry, per-project event stores, hub registry/project stores | Copy through versioned migration; never mutate the only database in place |
19
+ | SQLite schema v3 `kxm.db` | Runtime registry, per-project event stores, hub registry/project stores | **No in-place schema migration.** A store stamped behind the current build fails closed with `runtime_schema_outdated` and the refusal never advances `user_version` (opening the file may still checkpoint WAL sidecars, so the whole state set is the backup unit); re-init instead. Stepwise lanes were removed 2026-09-20 under the single-operator decision from the hub store, the per-project event store, and the external-effects `ALTER TABLE` add-column; the Runtime registry never carried a stepwise lane (see [implementation-plan.md](../../plans/implementation-plan.md) → Decided) |
20
20
  | Full peer message bodies in hub DB | Summary-first sync events | Existing bodies remain protected legacy data and are not re-emitted automatically |
21
21
  | Project tokens/manual environment auth | Runtime enrollment and scoped credentials | Preserve current mode until enrollment is confirmed; never copy tokens into Git |
22
22
  | `.kxm/config/env.example` | Built-in defaults plus optional scoped env YAML | Import only explicit portable differences; secrets become references |
@@ -26,12 +26,27 @@ Presence of KXM documents does not activate new behavior.
26
26
 
27
27
  ## Compatibility releases and activation
28
28
 
29
+ > **Scope note (2026-09-20).** This matrix documents the Mesh/v0.5 → KXM cutover. KXM has
30
+ > one operator and no external installs, so **schema migration and old-state tolerance are
31
+ > out of scope** and the lanes that existed are gone: stores refuse an older stamp rather
32
+ > than upgrading, the external-effects store no longer adds a column in place, and the
33
+ > coordinator fingerprint no longer recomputes to forgive pre-canonicalisation rows. What
34
+ > remains here describes the **project/content** cutover, which the follow-up cut removes
35
+ > along with `kxm migrate`. Nothing here promises that the runtime will read an old
36
+ > database.
37
+
29
38
  Local Runtime support may ship publicly before hub KXM, but it remains beside
30
39
  existing hub contracts and stores. Old command names are not preserved. A project
31
40
  activates `kxm.*.v1` only by an explicit successful `kxm init`/migration receipt;
32
41
  file presence alone never activates it. Legacy hub runs continue on the legacy
33
42
  engine.
34
43
 
44
+ > **Superseded (2026-09-20).** The single-operator decision removes legacy readers
45
+ > **without** a compatibility release, so this transition-release list is no longer a plan of
46
+ > record. It stays here to name what was given up: no dual-read window, no `mesh_*` shim, and
47
+ > no period where legacy JSON stays loadable while KXM writes YAML. Anything still holding
48
+ > legacy state is refused rather than served from both shapes.
49
+
35
50
  When Phase 8 activates hub KXM, at least one hub transition release provides:
36
51
 
37
52
  - current `mesh_*` peer tools;
@@ -58,6 +73,12 @@ kxm migrate verify
58
73
 
59
74
  `kxm init` invokes the planning flow when it detects legacy state.
60
75
 
76
+ > **Superseded (2026-09-20).** The single-operator decision removed every schema
77
+ > migration lane, so "database/WAL migration … remain later-phase work" below is
78
+ > no longer the plan: there will be none. The `kxm migrate` commands described on
79
+ > this page are deleted in the follow-up cut, and a tree still holding legacy JSON
80
+ > fails closed at load instead of being converted.
81
+
61
82
  **Implementation status (Phase 1 slice):** the commands above are implemented
62
83
  for **configuration migration only** — legacy `agents.json`, `gates.json`, and
63
84
  workflow-definition JSON under `.kxm/config/`. Database/WAL migration,
@@ -120,6 +141,18 @@ installed resource bytes against the receipt. It performs no writes.
120
141
 
121
142
  ## Database migration
122
143
 
144
+ > **Superseded in part (2026-09-20).** What the code guarantees today is narrower than the
145
+ > checklist below and belongs to two different moments:
146
+ >
147
+ > - **Opening a store** verifies `user_version`, refuses a newer-than-known version, and
148
+ > enables WAL. That is implemented.
149
+ > - **Backing a store up** is where checkpointing, `-wal` handling, integrity verification,
150
+ > and hash recording happen — see the whole-state-set recipe in
151
+ > [`docs/operations.md`](../operations.md). Those are **not** properties of an ordinary open.
152
+ >
153
+ > The legacy-record import paragraphs describe a cutover that will not happen: this build
154
+ > migrates no database, and an older stamp is refused outright.
155
+
123
156
  Before any database operation:
124
157
 
125
158
  - verify SQLite `user_version`;
@@ -211,6 +244,11 @@ workflow history.
211
244
 
212
245
  ## Removal gate
213
246
 
247
+ > **Superseded (2026-09-20).** The operator decided removal happens **without** a
248
+ > compatibility release, so the list below no longer gates removal — it is kept only
249
+ > to record why the gate existed. Legacy readers are being deleted now, and old names
250
+ > fail closed by brake rather than alias.
251
+
214
252
  Legacy readers and command aliases are removed only after:
215
253
 
216
254
  - at least one compatibility release;
@@ -144,28 +144,204 @@ Recommended alerts:
144
144
  - waiting-run count or workflow wait timeouts rise beyond the expected external-system latency.
145
145
  - quorum degradation approvals occur outside a declared incident or change window.
146
146
 
147
- ## Backup and restore
148
-
149
- > **Hub-only today, and labelled as such.** The recipe below stops the hub and copies
150
- > `.kxm/state/kxm.db`. That is not the whole tenant state set: the Runtime keeps its own
151
- > `registry.db`, per-project event stores under the user state root, prompt sidecars,
152
- > bindings and configuration. A restore that follows only these steps can bring the hub back
153
- > while losing Runtime history. Queue step **S1** replaces this section with a stopped-state
154
- > procedure covering the full set, and **S5** proves it with one deployed restore before real
155
- > use; until S1 lands, treat this as the hub database only.
156
-
157
- SQLite runs in WAL mode. The safest simple backup is a coordinated copy while the hub is stopped:
158
-
159
- 1. Stop the hub gracefully.
160
- 2. Copy `.kxm/state/kxm.db` to protected backup storage.
161
- 3. Keep the backup with the application version and configuration used to create it.
162
- 4. Restart the hub and confirm `/ready` returns `ok: true`.
147
+ ## Per-tenant hosted deployment
148
+
149
+ One tenant is one machine: one hub process, one Runtime supervisor, one SQLite state set.
150
+ Tenancy is the box, not a table — the hub has no tenant column and no user accounts, and
151
+ `kxm hub bind` still means *this machine's client attaches to that hub URL*. The portal
152
+ (kontextmind/kxmd-portal) is the multi-tenant, multi-user surface; browsers never talk to
153
+ the hub.
154
+
155
+ Topology on the tenant box:
156
+
157
+ ```text
158
+ browser ──HTTPS──▶ reverse proxy + Authentik ──▶ portal (users, sessions, tenant directory)
159
+ │
160
+ │ server-side, loopback
161
+ ▼
162
+ hub 127.0.0.1:7331 · Runtime supervisor (loopback)
163
+ ```
163
164
 
164
- For online backups, use a SQLite-aware backup tool or snapshot the database, `-wal`, and `-shm` files consistently. A plain copy of only `kxm.db` while the service is writing may omit committed WAL data.
165
+ 1. **Service account, not root.** Run the hub and Runtime under a dedicated unprivileged
166
+ account. Workflow session isolation is a routing and cross-run safety mechanism, not a
167
+ sandbox against a hostile same-OS process (see
168
+ [architecture.md](architecture.md)); a model with shell access can reach anything its
169
+ own account can reach, so untrusted workers need separate accounts or containers.
170
+ 2. **Stable paths, declared explicitly** rather than inherited from a home directory:
171
+ `KXM_WORKSPACE_DIR`, `KXM_STATE_DIR`, `KXM_DATA_PATH`, `KXM_LOG_PATH`, and
172
+ `KXM_STATE_HOME` for the machine-level hub credential and binding records. Pin
173
+ `KXM_HOST=127.0.0.1`.
174
+ 3. **Loopback listeners only.** The hub and the supervisor expose no public port; nothing
175
+ is load-balanced across hubs. The hub keeps one writer per database.
176
+ 4. **One project, the slim workflow.** `kxm init` the workspace, then start the hub with
177
+ `kxm hub start` and confirm with `kxm hub view` — the status line reports the binding as
178
+ `loopback` or `remote`, so an operator can see which side of the trust line they are on
179
+ without reading files. `kxm hub bind` refuses a **remote** URL when the machine has no
180
+ credential to authenticate with, because a stored-but-unusable URL later reads as a
181
+ network fault and gets debugged as one.
182
+ 5. **Restart recovery is the existing one:** the PID claim file, dead-claim reclaim and
183
+ graceful `SIGTERM` shutdown described under *PID claims and restart recovery*. Do not
184
+ add a second service manager for the hub; use the tenant's existing one.
185
+
186
+ ### Reverse-proxy contract
187
+
188
+ The tenant's proxy owns TLS and the browser session. KXM ships no proxy configuration,
189
+ because a generated config reads as authoritative while one missing directive silently
190
+ re-opens header forgery. What must hold, whatever the stack:
191
+
192
+ - The hub and supervisor ports are **not** reachable from outside the box.
193
+ - The proxy **strips** client-supplied identity, tenant, agent, caller and `Authorization`
194
+ headers before injecting its own validated values. The hub must never see a
195
+ browser-forged `x-kxm-agent-id`, `x-kxm-caller-id` or bearer.
196
+ - The proxy **never injects the hub admin token on a user's behalf**. That flattens every
197
+ authenticated user in the tenant to hub admin and destroys attribution.
198
+ - Machine credentials stay server-side in the portal process. A browser must not hold,
199
+ echo, or be redirected with a hub bearer.
200
+ - `kxm hub bind` on a remote hub URL therefore requires an explicit credential, and a
201
+ refusal names the fix instead of only the failure.
202
+
203
+ Example (illustrative shape — not generated config, not tested by this repository's CI):
204
+ Authentik's embedded proxy answers a forward-auth subrequest per request; the tenant proxy
205
+ `proxy_cache_bypass`/`auth_request`-style gate allows only the portal's routes and keeps
206
+ `/v1/*` and the supervisor off the public interface entirely.
165
207
 
166
- To restore, stop the hub, preserve the current files for rollback, place the restored database at `.kxm/state/kxm.db` or the configured `KXM_DATA_PATH`, and start the same or newer compatible release. The runtime refuses a database whose schema version is newer than it supports.
208
+ ## Backup and restore
167
209
 
168
- Test restoration periodically. A backup that has never been restored is not a verified recovery path.
210
+ > **This section covers the whole tenant state set, on purpose.** A recipe that copies only
211
+ > `.kxm/state/kxm.db` is a hub-only backup: it silently omits the Runtime registry,
212
+ > per-project event stores, prompt sidecars, bindings and configuration, so a restore that
213
+ > passes every hub check can still lose run history. Verify with a real restore before first
214
+ > hosted use, not after an incident.
215
+
216
+ SQLite runs in WAL mode, so a consistent copy requires a stopped service (or a SQLite-aware
217
+ online tool). Stop the hub and the Runtime supervisor first.
218
+
219
+ **What a tenant backup contains.** Six roots — one fixed to the checkout, one for the
220
+ workspace directories, and four more that can each sit anywhere — and confusing them is how
221
+ a backup goes missing while looking complete:
222
+
223
+ - **`$S`** — host-local machine state: `$KXM_STATE_HOME` **when set**, and it must be an
224
+ absolute path — a relative value is **rejected** with `local_state_root_not_absolute`, not
225
+ redirected. When unset, the default is `~/.local/state/kxm` on Linux (honouring
226
+ `XDG_STATE_HOME`), `~/Library/Application Support/KXM` on macOS, or
227
+ `%LOCALAPPDATA%\KXM` on Windows. The silent case to know about is a relative
228
+ `XDG_STATE_HOME`/`LOCALAPPDATA` **base**: that falls back to the default without error,
229
+ so a backup path derived from it can quietly point somewhere else.
230
+ - **`$R`** — the checkout root. Everything below it is **fixed to the repository and does
231
+ not follow any workspace override**: `$R/.kxm/project.yaml`, `$R/.kxm/config.yaml`,
232
+ `$R/.kxm/agents/`, `$R/.kxm/workflows/`, `$R/.kxm/gates.yaml`, `$R/.kxm/roles/`,
233
+ `$R/.kxm/role-hosts.yaml` (or `.json`), `$R/.kxm/producers.yaml`, `$R/.kxm/roster.yaml`,
234
+ `$R/.kxm/routes.yaml`, `$R/.kxm/prices.yaml`, `$R/.kxm/repo/`,
235
+ `$R/.kxm/template-provenance.yaml`, plus the durable work and learning records
236
+ `$R/.kxm/goals/`, `$R/.kxm/tasks/`, `$R/.kxm/memory/` (with `memory/candidates/`) and
237
+ `$R/.kxm/skills/`. Conflating these with the next root is how a backup omits the project
238
+ definition while believing it copied the project.
239
+ - **`$D`** — the **workspace directories**, resolved from `--workspace` or
240
+ `KXM_WORKSPACE_DIR`, else `$R/.kxm`, relative to `KXM_WORKDIR`/cwd:
241
+ `$D/config`, `$D/logs`, `$D/assets`, `$D/state`. `--workspace` **derives all four** and
242
+ ignores the per-directory variables; otherwise `KXM_CONFIG_DIR`, `KXM_LOGS_DIR`,
243
+ `KXM_ASSETS_DIR` and `KXM_STATE_DIR` override each one independently, and
244
+ `KXM_DATA_PATH`/`KXM_LOG_PATH` move two files again inside that. `$D` therefore **defaults to `$R/.kxm`**, and the two
245
+ move together only when the *workspace* is relocated: `KXM_WORKSPACE_DIR` (or
246
+ `--workspace`) moves `$D` and every default beneath it, while `KXM_CONFIG_DIR`,
247
+ `KXM_LOGS_DIR`, `KXM_ASSETS_DIR`, `KXM_STATE_DIR`, `KXM_DATA_PATH` and `KXM_LOG_PATH`
248
+ move **their own target and nothing else** — `KXM_STATE_DIR=/srv/state` alone leaves `$D`
249
+ at `$R/.kxm` and shifts only `$W`. A backup that assumes one shared location starts
250
+ omitting the other in exactly that case, which is why every row below is labelled as a
251
+ default.
252
+ - **`$W`** — the workspace *state* directory: `KXM_STATE_DIR` when set, else `$D/state`
253
+ (and `--workspace` derives it, ignoring that variable). It holds the
254
+ hub database, worker routing/recovery manifests and Pi sessions.
255
+ - **`$C`** — user configuration: `KXM_USER_CONFIG_DIR`, else `~/.config/kxm`.
256
+ - **`$T`** — federated telemetry output: an explicit global directory joined with
257
+ **`telemetry/`**, else `$XDG_CONFIG_HOME/kxm/telemetry`, else `~/.config/kxm/telemetry`.
258
+ It is built from `XDG_CONFIG_HOME`/`HOME`, **not** from `KXM_USER_CONFIG_DIR`, so `$T` can
259
+ land outside `$C`; and it is a *different file* from local accounting in `$D/logs`.
260
+
261
+ | Path | Contents | Loss means |
262
+ |---|---|---|
263
+ | `$W/kxm.db` (+ `-wal`, `-shm`, or `KXM_DATA_PATH`) | hub store: agents, messages, workflow runs, checkpoints, gate evidence | hub history and delivery state |
264
+ | `$S/runtime/registry.db` | Runtime registry, including the **supervisor identity and claim row** | which projects this Runtime knows; the claim is a registry row — there is no `supervisor.json` |
265
+ | `$S/runtime/projects/<projectKey>/run-events.db` (+ `-wal`/`-shm`) | event-sourced run state, commands, drives, receipts, gate evidence, intake, coordinators, pause control | run history and every receipt that proves it |
266
+ | `$S/runtime/projects/<projectKey>/run-events.db.run-prompts.json` | prompt text; the sidecar name appends to the **full** database filename | the prompts that explain the runs — restoring databases without sidecars is a partial restore |
267
+ | `$S/projects/<control-root-hash>/repository-bindings.json` | host-local member repository paths | member bindings are host state, outside the project tree |
268
+ | `$S/update.yaml` | release/update configuration consumed by the updater | the box reverts to defaults on the next update path |
269
+ | `$W/pi-sessions/<workerKey>/{default,runs/<runId>}/` | Pi model histories | **optional by existing policy** (see *Workflow-specific Pi sessions*): never a system of record — decide and record, do not silently widen scope |
270
+ | `$W/worker-session-binding-<workerKey>.json` (+ `.corrupt-*`), `worker-context-*.json`, `worker-recovery-*.json` | routing and recovery manifests | not optional: these are what make worker routing resumable after a restart |
271
+ | `$R/.kxm/…` project definition: `project.yaml`, `config.yaml`, `agents/`, `models/`, `workflows/`, `gates.yaml`, `roles/`, `role-hosts.yaml` (or `.json`), `producers.yaml`, `roster.yaml`, `routes.yaml`, `prices.yaml`, `repo/`, `project/env.yaml`, `template-provenance.yaml` | project, role, route, price and provenance definition | the tenant stops being reproducible — and a restore without `roster.yaml`/`routes.yaml`/`prices.yaml` comes back with **different admission and cost behaviour** while reporting itself healthy |
272
+ | `$R/.kxm/goals/`, `tasks/`, `memory/` (with `memory/candidates/`), `skills/` (candidate/promoted/rejected, history, patches) | durable work and learning records | open goals/tasks and approved memory disappear |
273
+ | `$R/.kxm/candidates/` — improvement candidate JSON and their diffs, **default only**: `kxm improve report --out-dir` relocates this directory outside every root listed here | the improvement queue itself | proposed fixes nobody was told about |
274
+ | each bound member repository's own `$memberRepo/.kxm/repo/repo.yaml` and `.kxm/repo/env.yaml` | member repository definition and environment | for **externally bound** members, the binding JSON alone is not enough — these files live on the member's own filesystem and need their own backup or an explicit, checked reconstruction prerequisite |
275
+ | `$D/assets/` (default; `KXM_ASSETS_DIR` relocates it) — retrospectives, improvements, artifacts, evidence | exported evidence | provenance and the ability to audit a past decision |
276
+ | `$D/logs/` (default; `KXM_LOGS_DIR` relocates the directory and `KXM_LOG_PATH` the hub log) and `$D/logs/telemetry.jsonl` | operator logs and **local** usage accounting — the spend numbers routing reports read | no local accounting to reconcile against |
277
+ | `KXM_WORKER_LOG_PATH` / `KXM_AGENT_LOG_PATH` targets (defaulting under `$D/logs`) | per-worker lifecycle and raw Pi output | worker diagnostics; **separate overrides, not local accounting** |
278
+ | `$T/model-metrics.jsonl` | **federated** metrics only. Absent almost everywhere: the exporter exists and `telemetry.federated` defaults to `true` in the shipped config, but **no hub or CLI path calls it today**, so absence is the normal state rather than evidence someone opted out. A different file from local accounting, which is `$D/logs/telemetry.jsonl` | cross-machine reporting continuity, and a privacy boundary worth naming: federated records are separate, with `anonymize` defaulting to `true` |
279
+ | `$S/hub-binding.json`, `$S/hub-env.json`, `$C/session.token` | host hub URL, credentials, local session token | a re-bind and a token rotation. **Secrets:** prefer regeneration to shipping them off-box, and never commit them |
280
+ | `$C` global roles/workflows/host configuration | user-level defaults | operator conventions |
281
+
282
+ **Overrides are part of the backup record.** `KXM_WORKSPACE_DIR` (and the `--workspace`
283
+ flag, which additionally **ignores** the per-directory variables) moves every `$D`
284
+ **default** at once; a directory with its own override stays where that variable points.
285
+ Neither moves `$R`, so the fixed project tree must still be backed up from the checkout even
286
+ when the workspace was relocated elsewhere — and copying the whole checkout is what protects
287
+ `$R`'s default locations, which is why an enumerated-paths backup should be re-checked
288
+ against this table whenever a loader grows a file. `KXM_STATE_HOME` moves `$S`
289
+ only if absolute. An explicit telemetry directory is likewise joined with `telemetry/`,
290
+ not used verbatim.
291
+ `KXM_DATA_PATH`, `KXM_STATE_DIR`, `KXM_CONFIG_DIR`, `KXM_ASSETS_DIR`, `KXM_LOGS_DIR`,
292
+ `KXM_LOG_PATH`, `KXM_WORKER_LOG_PATH`, `KXM_AGENT_LOG_PATH` or an explicit telemetry
293
+ directory each relocate one more thing. Record every override **with** the backup, or a
294
+ restore lands somewhere the running service will not look.
295
+
296
+ **Disposable, not backup material:** `session-brief.json`, `update-check.json`,
297
+ `*.error`, PID/claim files such as `hub.pid` and `worker-<key>.pid`, and
298
+ `supervisor.token` (host-local secret, re-generated on start). WAL-consistent copying or
299
+ `VACUUM INTO` applies to **every** SQLite file above, not only the hub database.
300
+
301
+ **Back up (stopped-state recipe):**
302
+
303
+ 1. Stop **both** services, confirm they are down, and **keep them down until the copy
304
+ finishes**. `kxm hub stop` covers the hub and its worker PID claims; the Runtime
305
+ supervisor is a **separate** process owning `registry.db` and the project event stores,
306
+ stopped by `kxm runtime stop` — and that call acknowledges shutdown *initiated*, not
307
+ databases closed. So: verify neither reports live, then suspend whatever would start them
308
+ again — the service manager's auto-restart, hub autostart on login, and any client that
309
+ would reconnect and begin new work (a bound CLI, MCP server or Pi worker restarting a
310
+ supervisor on demand). A manager that respawns the hub halfway through a copy produces a
311
+ backup that is internally inconsistent across files, which is precisely the failure mode
312
+ this recipe is otherwise careful about. Only then copy.
313
+ 2. Copy the whole set above as one tree — **every root**, `$R`, `$D`, `$W`, `$S`, `$C` and
314
+ `$T` — or take `VACUUM INTO` snapshots per database. **Snapshots replace the database
315
+ copies, not the file copy**: configuration, repository bindings, prompt sidecars,
316
+ routing manifests and update configuration are not databases, so a snapshot-only backup
317
+ reproduces exactly the failure this section exists to remove. The hub's own backup path already writes a hashed manifest and records
318
+ a schema version ceiling; keep that manifest with the files.
319
+ 3. Record the package version, configuration revision and schema versions beside the copy.
320
+ A restore that cannot state which release produced it is not a restore path.
321
+ 4. Keep at least one rotation, and bound retention explicitly — run events and prompt
322
+ sidecars grow, and unbounded retention is how a tenant box fills up.
323
+
324
+ **Restore:**
325
+
326
+ 1. Stop the services. Move the current state aside rather than overwriting it.
327
+ 2. Place each file back at its recorded path under the right root — `KXM_DATA_PATH`, the
328
+ Runtime registry and each project event store **with its sidecar**, repository bindings
329
+ under `$S/projects/…`, config, and `$S/runtime/registry.db` so the supervisor claim
330
+ returns with it.
331
+ 3. Start the hub and confirm `/ready`, then `kxm hub view` — including that the reported
332
+ binding scope is what the environment actually is.
333
+ 4. Read back a run and its drive receipt, and confirm prompt text is present. Restoring
334
+ databases without their sidecars leaves runs whose prompts are gone; that is a partial
335
+ restore, not a success.
336
+
337
+ The runtime refuses a database whose schema version is newer than it supports, so
338
+ restore order is: matching-or-newer release, then data. Online backups need a
339
+ SQLite-aware tool or a consistent snapshot of each database with its `-wal` and `-shm`;
340
+ a plain copy of a live `kxm.db` can omit committed WAL data.
341
+
342
+ Test restoration periodically. Routine unattended recovery (automated discovery of every
343
+ Runtime store plus sidecars) is deliberately **not** claimed here: it is a tracked
344
+ post-MVP item, and today this procedure is executed stopped and by hand.
169
345
 
170
346
  ## Upgrade and rollback
171
347
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@kontextmind/kxm",
3
- "version": "0.7.51",
3
+ "version": "0.7.53",
4
4
  "description": "KXM local-first multi-agent orchestration and operator dashboard",
5
5
  "type": "module",
6
6
  "author": "KontextMind",
@@ -2,7 +2,7 @@
2
2
  "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
3
3
  "name": "kxm",
4
4
  "displayName": "KXM",
5
- "version": "0.7.51",
5
+ "version": "0.7.53",
6
6
  "description": "Headless multi-agent orchestration, durable workflows, and a live operator dashboard for Pi and Claude Code",
7
7
  "author": {
8
8
  "name": "KontextMind",
@@ -19657,6 +19657,21 @@ function readHubEnvRecord(env = process.env) {
19657
19657
  ...record.projectTokens !== void 0 ? { projectTokens: record.projectTokens } : {}
19658
19658
  };
19659
19659
  }
19660
+ function hasClientHubCredential(env = process.env, project) {
19661
+ if (env.KXM_AUTH_TOKEN?.trim()) return true;
19662
+ let record;
19663
+ try {
19664
+ record = readHubEnvRecord(env);
19665
+ } catch (error) {
19666
+ throw new HubEnvError(
19667
+ `${error instanceof Error ? error.message : String(error)}; refusing to guess a credential \u2014 repair or remove ${hubEnvFile(env)}`
19668
+ );
19669
+ }
19670
+ if (record?.authToken?.trim()) return true;
19671
+ const tokens = record?.projectTokens ?? {};
19672
+ if (project !== void 0) return typeof tokens[project] === "string" && tokens[project].trim().length > 0;
19673
+ return Object.values(tokens).some((token) => typeof token === "string" && token.trim().length > 0);
19674
+ }
19660
19675
  function resolveClientHubAuthToken(env, project) {
19661
19676
  const envToken = env.KXM_AUTH_TOKEN?.trim();
19662
19677
  if (envToken) return envToken;
@@ -23752,20 +23767,11 @@ function openDatabase(file, description, spec) {
23752
23767
  database.exec(spec.schema);
23753
23768
  database.exec(`PRAGMA user_version = ${spec.version}`);
23754
23769
  } else if (version < spec.version) {
23755
- let currentVersion = version;
23756
- while (currentVersion < spec.version) {
23757
- const step = spec.migrations?.find((m2) => m2.fromVersion === currentVersion);
23758
- if (!step) {
23759
- throw databaseError(
23760
- "runtime_schema_outdated",
23761
- file,
23762
- `${description} schema version ${version} is older than ${spec.version}; no migration lane, backup and restore remain E6`
23763
- );
23764
- }
23765
- step.migrate(database);
23766
- currentVersion = step.toVersion;
23767
- database.exec(`PRAGMA user_version = ${currentVersion}`);
23768
- }
23770
+ throw databaseError(
23771
+ "runtime_schema_outdated",
23772
+ file,
23773
+ `${description} is schema version ${version}; this build requires ${spec.version}. Delete the state file (or re-run \`kxm init\`) to start fresh \u2014 upgrading old state in place is deliberately unsupported`
23774
+ );
23769
23775
  }
23770
23776
  if (spec.tables) {
23771
23777
  verifyExpectedTables(database, file, description, spec.tables);
@@ -26098,6 +26104,17 @@ function validateHubUrl(raw) {
26098
26104
  }
26099
26105
  return parsed.href.replace(/\/$/, "");
26100
26106
  }
26107
+ function hubBindingScope(url) {
26108
+ let host;
26109
+ try {
26110
+ host = new URL(url).hostname.toLowerCase();
26111
+ } catch {
26112
+ return "remote";
26113
+ }
26114
+ if (host === "localhost" || host === "::1" || host === "[::1]" || host.endsWith(".localhost")) return "loopback";
26115
+ const v4 = /^127\.([0-9]{1,3})\.([0-9]{1,3})\.([0-9]{1,3})$/.exec(host);
26116
+ return v4 && [v4[1], v4[2], v4[3]].every((part) => Number(part) <= 255) ? "loopback" : "remote";
26117
+ }
26101
26118
  function isIsoTimestamp(value) {
26102
26119
  if (Number.isNaN(Date.parse(value))) return false;
26103
26120
  return value === new Date(value).toISOString();
@@ -45212,9 +45229,10 @@ var MAX_SESSION_BRIEF_PLANS = 5;
45212
45229
  var SESSION_BRIEF_SCHEMA = "kxm.session-brief.v1";
45213
45230
  var DEFAULT_SESSION_BRIEF_STALE_SECONDS = 5;
45214
45231
  function hubPrefix(hub) {
45215
- if (hub?.state === "on" || hub?.state === void 0 && hub?.online === true) return "kxm hub:on";
45216
- if (hub?.state === "off" || hub?.state === void 0 && hub?.online === false) return "kxm hub:off";
45217
- if (hub?.state === "unknown") return "kxm hub:unknown";
45232
+ const suffix = hub?.scope === "remote" ? "/remote" : "";
45233
+ if (hub?.state === "on" || hub?.state === void 0 && hub?.online === true) return `kxm hub:on${suffix}`;
45234
+ if (hub?.state === "off" || hub?.state === void 0 && hub?.online === false) return `kxm hub:off${suffix}`;
45235
+ if (hub?.state === "unknown") return `kxm hub:unknown${suffix}`;
45218
45236
  return "kxm";
45219
45237
  }
45220
45238
  function truncate(value, width) {
@@ -45848,8 +45866,21 @@ async function refreshKxmUpdateNotice(runtime, config) {
45848
45866
  async function cmdStatus(runtime) {
45849
45867
  const health = await hubGet(`${runtime.serverUrl}/health`, runtime.fetchImpl);
45850
45868
  const ready = await hubGet(`${runtime.serverUrl}/ready`, runtime.fetchImpl);
45851
- const payload = { ok: health.ok && ready.ok, command: "hub view", health: health.body, ready: ready.body };
45852
- print(runtime.io, runtime.json, payload, `hub health=${health.ok} ready=${ready.ok}`);
45869
+ const effectiveScope = hubBindingScope(runtime.serverUrl);
45870
+ const overridden = Boolean(runtime.boundHubUrl && runtime.boundHubUrl !== runtime.serverUrl);
45871
+ const payload = {
45872
+ ok: health.ok && ready.ok,
45873
+ command: "hub view",
45874
+ target: { url: runtime.serverUrl, scope: effectiveScope, ...overridden ? { source: "env" } : {} },
45875
+ health: health.body,
45876
+ ready: ready.body
45877
+ };
45878
+ print(
45879
+ runtime.io,
45880
+ runtime.json,
45881
+ payload,
45882
+ `hub health=${health.ok} ready=${ready.ok} \xB7 ${effectiveScope} hub${overridden ? " (KXM_SERVER_URL)" : ""}`
45883
+ );
45853
45884
  return payload.ok ? 0 : 1;
45854
45885
  }
45855
45886
  async function cmdDash(runtime, options = {}) {
@@ -45934,6 +45965,7 @@ function formatHubBindHealth(health) {
45934
45965
  if (health === "off") return "health=off (nothing answered; run kxm hub start)";
45935
45966
  return "health=unknown (no reply within 300 ms)";
45936
45967
  }
45968
+ var HUB_BIND_UNAUTHENTICATED_HINT = "export KXM_AUTH_TOKEN (or point KXM_STATE_HOME at the hub-env record that already holds one), then re-run; the hub itself requires a token beyond loopback";
45937
45969
  async function cmdHubBind(runtime, rawUrl) {
45938
45970
  let url;
45939
45971
  try {
@@ -45950,14 +45982,64 @@ async function cmdHubBind(runtime, rawUrl) {
45950
45982
  }
45951
45983
  throw error;
45952
45984
  }
45985
+ const scope = hubBindingScope(url);
45986
+ if (scope === "remote") {
45987
+ const bindProject = defaultProjectName(runtime.dirs.workdir, runtime.env) || "project";
45988
+ let credentialReady = false;
45989
+ try {
45990
+ credentialReady = hasClientHubCredential(runtime.env, bindProject);
45991
+ } catch (error) {
45992
+ print(
45993
+ runtime.io,
45994
+ runtime.json,
45995
+ {
45996
+ ok: false,
45997
+ command: "hub bind",
45998
+ error: "hub_credential_unreadable",
45999
+ url,
46000
+ scope,
46001
+ nextAction: "repair_hub_env_record",
46002
+ hint: `${error instanceof Error ? error.message : String(error)}; no binding was written`
46003
+ },
46004
+ `cannot read the hub credential: ${error instanceof Error ? error.message : String(error)}`
46005
+ );
46006
+ return 2;
46007
+ }
46008
+ if (!credentialReady) {
46009
+ print(
46010
+ runtime.io,
46011
+ runtime.json,
46012
+ {
46013
+ ok: false,
46014
+ command: "hub bind",
46015
+ error: "hub_bind_unauthenticated",
46016
+ url,
46017
+ scope,
46018
+ project: bindProject,
46019
+ // The hint belongs in the payload, not only the prose line: under --json the
46020
+ // prose is suppressed, and a refusal that names no next step gets debugged by
46021
+ // reading source.
46022
+ nextAction: "export_kxm_auth_token",
46023
+ hint: `${HUB_BIND_UNAUTHENTICATED_HINT} (needs a token for project ${bindProject})`
46024
+ },
46025
+ `refusing to bind remote hub ${url} with no credential for project ${bindProject}; ${HUB_BIND_UNAUTHENTICATED_HINT}`
46026
+ );
46027
+ return 2;
46028
+ }
46029
+ }
45953
46030
  const file = hubBindingFile(runtime.env);
45954
46031
  if (runtime.dryRun) {
45955
- print(runtime.io, runtime.json, { ok: true, command: "hub bind", dryRun: true, url, file }, `would bind hub ${url}`);
46032
+ print(runtime.io, runtime.json, { ok: true, command: "hub bind", dryRun: true, url, scope, file }, `would bind hub ${url} (${scope})`);
45956
46033
  return 0;
45957
46034
  }
45958
46035
  writeHubBinding({ schema: HUB_BINDING_SCHEMA, url, boundAt: (/* @__PURE__ */ new Date()).toISOString() }, runtime.env);
45959
46036
  const { health, probeMs } = await probeHubHealth(url, runtime.fetchImpl);
45960
- print(runtime.io, runtime.json, { ok: true, command: "hub bind", url, file, health, probeMs }, `bound hub ${url} \xB7 ${formatHubBindHealth(health)}`);
46037
+ print(
46038
+ runtime.io,
46039
+ runtime.json,
46040
+ { ok: true, command: "hub bind", url, scope, file, health, probeMs },
46041
+ `bound hub ${url} \xB7 ${scope} \xB7 ${formatHubBindHealth(health)}${scope === "remote" ? " \xB7 token leaves this machine" : ""}`
46042
+ );
45961
46043
  return 0;
45962
46044
  }
45963
46045
  async function cmdHubUnbind(runtime) {
@@ -46271,7 +46353,8 @@ async function cmdSessionBrief(runtime, options = {}) {
46271
46353
  state: health,
46272
46354
  evidence: health === "unknown" ? "timeout" : "probed",
46273
46355
  online: health === "on",
46274
- url: targetUrl
46356
+ url: targetUrl,
46357
+ scope: hubBindingScope(targetUrl)
46275
46358
  };
46276
46359
  } else {
46277
46360
  hub = { state: "off", evidence: "unconfigured", online: false };
@@ -36775,9 +36775,10 @@ var MAX_SESSION_BRIEF_PLANS = 5;
36775
36775
  var SESSION_BRIEF_SCHEMA = "kxm.session-brief.v1";
36776
36776
  var DEFAULT_SESSION_BRIEF_STALE_SECONDS = 5;
36777
36777
  function hubPrefix(hub) {
36778
- if (hub?.state === "on" || hub?.state === void 0 && hub?.online === true) return "kxm hub:on";
36779
- if (hub?.state === "off" || hub?.state === void 0 && hub?.online === false) return "kxm hub:off";
36780
- if (hub?.state === "unknown") return "kxm hub:unknown";
36778
+ const suffix = hub?.scope === "remote" ? "/remote" : "";
36779
+ if (hub?.state === "on" || hub?.state === void 0 && hub?.online === true) return `kxm hub:on${suffix}`;
36780
+ if (hub?.state === "off" || hub?.state === void 0 && hub?.online === false) return `kxm hub:off${suffix}`;
36781
+ if (hub?.state === "unknown") return `kxm hub:unknown${suffix}`;
36781
36782
  return "kxm";
36782
36783
  }
36783
36784
  function truncate(value, width) {