@kontextmind/kxm 0.7.95 → 0.7.97

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (158) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.kxm/README.md +39 -9
  3. package/CHANGELOG.md +23 -1
  4. package/README.md +147 -257
  5. package/SECURITY.md +21 -12
  6. package/docs/README.md +133 -54
  7. package/docs/adr/ADR-0002-browser-automation-steel-doks.md +24 -18
  8. package/docs/adr/ADR-0003-sqlite-only-store.md +100 -0
  9. package/docs/adr/ADR-0004-edge-identity-authentik.md +99 -0
  10. package/docs/adr/README.md +33 -0
  11. package/docs/concepts/architecture.md +262 -0
  12. package/docs/concepts/data-and-storage.md +194 -0
  13. package/docs/concepts/trust-model.md +153 -0
  14. package/docs/contracts/README.md +22 -14
  15. package/docs/contracts/effects-and-recovery.md +3 -0
  16. package/docs/contracts/migration.md +2 -2
  17. package/docs/contracts/routing.md +6 -5
  18. package/docs/contributing/assignment-runner.md +388 -0
  19. package/docs/contributing/ci-and-release.md +231 -0
  20. package/docs/contributing/development.md +362 -0
  21. package/docs/contributing/harness-routing-internals.md +192 -0
  22. package/docs/{packages.md → contributing/packages.md} +13 -15
  23. package/docs/{skills → contributing}/repo-work-delivery.md +20 -21
  24. package/docs/contributing/test-matrix.md +208 -0
  25. package/docs/{tui-components.md → contributing/tui-components.md} +30 -22
  26. package/docs/contributing/writing-docs.md +340 -0
  27. package/docs/glossary.md +471 -0
  28. package/docs/guides/agent-skills.md +137 -0
  29. package/docs/guides/browser-automation.md +160 -0
  30. package/docs/guides/context-and-memory.md +352 -0
  31. package/docs/guides/continuous-improvement.md +228 -0
  32. package/docs/guides/governed-skills.md +173 -0
  33. package/docs/guides/nous-providers.md +186 -0
  34. package/docs/guides/peer-messaging.md +304 -0
  35. package/docs/guides/pi-workers.md +219 -0
  36. package/docs/guides/provenance-gates.md +313 -0
  37. package/docs/guides/webhook-workflows.md +399 -0
  38. package/docs/kb/how-credentials-retrieved-safely.md +38 -12
  39. package/docs/kb/how-to-capture-and-annotate-section.md +15 -13
  40. package/docs/kb/how-to-connect-playwright-to-steel.md +16 -11
  41. package/docs/kb/how-to-recover-expired-session-or-orphan.md +26 -16
  42. package/docs/kb/how-to-resume-after-mfa.md +19 -11
  43. package/docs/kb/how-to-take-over-session.md +17 -13
  44. package/docs/kb/why-authentication-disappeared.md +22 -14
  45. package/docs/kb/why-automation-opened-different-browser.md +23 -14
  46. package/docs/kb/why-session-viewer-cannot-control.md +13 -12
  47. package/docs/operations/backup-and-restore.md +248 -0
  48. package/docs/operations/deploy.md +307 -0
  49. package/docs/operations/monitoring.md +209 -0
  50. package/docs/operations/runtime-sync.md +192 -0
  51. package/docs/operations/troubleshooting.md +266 -0
  52. package/docs/operations/upgrade.md +124 -0
  53. package/docs/prompts/browser-annotate-feedback.md +7 -7
  54. package/docs/prompts/browser-diagnose-recover.md +11 -10
  55. package/docs/prompts/browser-explore.md +7 -7
  56. package/docs/prompts/browser-repro-fix.md +7 -7
  57. package/docs/prompts/browser-start.md +12 -11
  58. package/docs/prompts/browser-takeover.md +8 -8
  59. package/docs/{cli-reference.md → reference/cli-reference.md} +88 -46
  60. package/docs/{config-reference.md → reference/config-reference.md} +159 -148
  61. package/docs/reference/configuration.md +299 -0
  62. package/docs/reference/harness-routing.md +508 -0
  63. package/docs/reference/http-api.md +203 -0
  64. package/docs/reference/tools.md +370 -0
  65. package/docs/{workflow-guide.md → reference/workflow-catalog.md} +92 -153
  66. package/docs/reference/workflow-definitions.md +286 -0
  67. package/docs/start/first-workflow.md +287 -0
  68. package/docs/start/install.md +146 -0
  69. package/docs/start/quickstart-claude-code.md +405 -0
  70. package/docs/start/quickstart-pi.md +213 -0
  71. package/docs/templates/README.md +78 -73
  72. package/docs/templates/adr.md +13 -13
  73. package/docs/templates/architecture.md +55 -71
  74. package/docs/templates/bug-fix.md +13 -16
  75. package/docs/templates/feature.md +14 -19
  76. package/docs/templates/handoff.md +44 -46
  77. package/docs/templates/postmortem.md +30 -43
  78. package/docs/templates/research.md +15 -20
  79. package/docs/templates/review.md +49 -50
  80. package/docs/templates/runbook.md +38 -30
  81. package/docs/templates/test-plan.md +16 -23
  82. package/docs/templates/test-report.md +14 -17
  83. package/examples/README.md +9 -5
  84. package/examples/provenance-workflow.json +1 -1
  85. package/examples/webhook-workflows/jira-development.json +59 -0
  86. package/examples/webhook-workflows/jira-issue-updated.json +12 -0
  87. package/examples/workflow-signal.ts +4 -5
  88. package/package.json +1 -1
  89. package/packages/core/tui/README.md +1 -1
  90. package/plugins/kxm/.claude-plugin/plugin.json +1 -1
  91. package/plugins/kxm/README.md +31 -32
  92. package/plugins/kxm/dist/claude-hook.js +11 -1
  93. package/plugins/kxm/dist/cli.js +164 -79
  94. package/plugins/kxm/dist/client.js +3 -1
  95. package/plugins/kxm/dist/core.js +11 -1
  96. package/plugins/kxm/dist/extension.js +45 -13
  97. package/plugins/kxm/dist/mcp-server.js +20 -4
  98. package/plugins/kxm/dist/runtime-supervisor.js +1 -3
  99. package/plugins/kxm/dist/runtime.js +18 -4
  100. package/plugins/kxm/dist/server.js +115 -20
  101. package/plugins/kxm/package.json +1 -1
  102. package/plugins/kxm/skills/kxm/references/protocol.md +3 -1
  103. package/plugins/kxm/skills/kxm-browser-auth/SKILL.md +1 -1
  104. package/plugins/kxm/skills/kxm-browser-diagnostics/SKILL.md +5 -5
  105. package/plugins/kxm/skills/kxm-browser-explore/SKILL.md +2 -2
  106. package/plugins/kxm/skills/kxm-browser-session/SKILL.md +10 -13
  107. package/plugins/kxm/skills/kxm-browser-takeover/SKILL.md +1 -1
  108. package/plugins/kxm/skills/kxm-browser-verify/SKILL.md +1 -1
  109. package/plugins/kxm/skills/kxm-context-memory/SKILL.md +13 -4
  110. package/plugins/kxm/skills/kxm-hub-ops/SKILL.md +3 -1
  111. package/plugins/kxm/skills/kxm-mind-setup/SKILL.md +2 -1
  112. package/plugins/kxm/skills/kxm-project-setup/SKILL.md +31 -54
  113. package/plugins/kxm/skills/kxm-projects/SKILL.md +1 -1
  114. package/plugins/kxm/skills/kxm-protocol/SKILL.md +1 -1
  115. package/plugins/kxm/skills/kxm-routing-improve/SKILL.md +15 -7
  116. package/plugins/kxm/skills/kxm-runs/SKILL.md +11 -5
  117. package/plugins/kxm/skills/kxm-session/SKILL.md +1 -1
  118. package/plugins/kxm/skills/kxm-tasks/SKILL.md +9 -7
  119. package/plugins/kxm/skills/kxm-workflow/SKILL.md +10 -2
  120. package/plugins/kxm/src/cli/system.ts +1 -1
  121. package/plugins/kxm/src/cli/workflows.ts +12 -7
  122. package/plugins/kxm/src/cli.ts +22 -8
  123. package/plugins/kxm/src/client.ts +4 -0
  124. package/plugins/kxm/src/commands.ts +23 -1
  125. package/plugins/kxm/src/extension.ts +20 -14
  126. package/plugins/kxm/src/github-watch.ts +8 -5
  127. package/plugins/kxm/src/hub-env.ts +19 -1
  128. package/plugins/kxm/src/hub.ts +105 -21
  129. package/plugins/kxm/src/improve-sources.ts +2 -7
  130. package/plugins/kxm/src/init-guide-setup.ts +1 -1
  131. package/plugins/kxm/src/mcp-server.ts +9 -2
  132. package/plugins/kxm/src/modes.ts +1 -1
  133. package/plugins/kxm/src/runtime-store.ts +23 -0
  134. package/plugins/kxm/src/workflow.ts +70 -1
  135. package/schemas/README.md +1 -1
  136. package/scripts/smoke-multi-pi.mjs +5 -1
  137. package/docs/agent-communication-envelopes-and-gates.md +0 -553
  138. package/docs/agent-skills.md +0 -198
  139. package/docs/architecture.md +0 -245
  140. package/docs/assignment-runner.md +0 -264
  141. package/docs/browser-automation.md +0 -139
  142. package/docs/configuration.md +0 -437
  143. package/docs/continuous-improvement.md +0 -226
  144. package/docs/getting-started.md +0 -277
  145. package/docs/harness-routing.md +0 -616
  146. package/docs/kb/qa-authentik-authentication.md +0 -97
  147. package/docs/kb/qa-extension-install-and-hub-bootstrap.md +0 -85
  148. package/docs/kb/qa-hub-on-a-public-host.md +0 -48
  149. package/docs/kb/qa-sqlite-vs-duckdb.md +0 -35
  150. package/docs/kb/qa-what-the-hub-stores.md +0 -64
  151. package/docs/kxm-handbook.md +0 -1181
  152. package/docs/operations.md +0 -510
  153. package/docs/operator-pi-packages.md +0 -67
  154. package/docs/provenance-gates.md +0 -295
  155. package/docs/skills.md +0 -47
  156. package/docs/test-matrix.md +0 -132
  157. package/docs/troubleshooting.md +0 -322
  158. package/docs/webhook-workflows.md +0 -240
@@ -1,15 +1,23 @@
1
1
  # KXM contract package
2
2
 
3
- > **Status: planned normative contract.** This directory describes the target
4
- > architecture accepted for KXM. Not all commands are implemented.
5
- > Phase 1 (init/migrate/trust) and Phase 2 (Runtime create/recover) have landed
6
- > slices. Phase 3 has D3 S1–S4 and D4 U2a-2 implemented (unreleased); it is not
7
- > only an agent-only simulated loop, and the default/fix driver gate remains
8
- > open. Operator tracking for the KXM rename, `kxm dash`, hub CLI, and
9
- > harness YAML lives in the
10
- > [implementation plan](../../plans/implementation-plan.md#tracking-working-tree-not-a-release).
11
- > For current hub execution behavior, use [Architecture](../architecture.md) and
12
- > [Configuration](../configuration.md).
3
+ > [!IMPORTANT]
4
+ > Planned: these contracts specify KXM's target architecture. Parts ship today
5
+ > and parts do not; the table below gives each page's status, and a page's own
6
+ > status line wins over this summary. For the behavior that ships, read
7
+ > [Architecture](../concepts/architecture.md), the
8
+ > [configuration reference](../reference/config-reference.md) and the
9
+ > [CLI reference](../reference/cli-reference.md).
10
+
11
+ | Contract | Status |
12
+ |---|---|
13
+ | [Architecture decision](architecture.md) | Accepted target. The local Runtime, per-project event stores and Runtime-to-hub sync exist today |
14
+ | [Terminology](terminology.md) | Normative now |
15
+ | [Lifecycles](lifecycles.md) | Partly implemented: the Runtime engine appends run, step, assignment, attempt and effect events |
16
+ | [Effects and recovery](effects-and-recovery.md) | Partly implemented: effect intents and `blocked_uncertain` are recorded; delivery is at-least-once, never exactly-once |
17
+ | [Synchronization](synchronization.md) | Implemented for the default policy; custom policies and on-demand content transfer are not |
18
+ | [Routing](routing.md) | Implemented: routing records, the dated price catalog, `kxm routing report` and `kxm improve` |
19
+ | [Validation](validation.md) | Largely implemented: the restricted YAML loader, schema, reference and semantic checks, and permission diffs |
20
+ | [Migration](migration.md) | Decided: there is no migration path and no `kxm migrate` command; legacy state is refused |
13
21
 
14
22
  KXM is a convention-over-configuration, local-first orchestration and
15
23
  context platform. One local Runtime owns execution; an optional multi-project
@@ -29,7 +37,7 @@ The words **MUST**, **MUST NOT**, **SHOULD**, and **MAY** are normative.
29
37
  | [Routing](routing.md) | Shipped v1 parser/report vs helper telemetry vs planned v2/catalog |
30
38
  | [Validation](validation.md) | Parse, schema, reference, semantic, permission, and snapshot validation |
31
39
  | [Migration](migration.md) | Compatibility from the current environment/JSON/SQLite surfaces |
32
- | [Implementation plan](../../plans/implementation-plan.md) | Ordered implementation and release gates |
40
+ | [Implementation plan](../../plans/implementation-plan.md) (repository only) | Ordered implementation and release gates; not in the npm package |
33
41
  | [Examples](../../examples/project/README.md) | Complete project and workflow fixture |
34
42
 
35
43
  Machine-readable schemas live under [`schemas`](../../schemas).
@@ -53,9 +61,9 @@ custom tags disabled and bounded aliases, depth, scalar size, and document size.
53
61
 
54
62
  ## Compatibility rule
55
63
 
56
- The current v0.5 contracts remain authoritative until a release explicitly
57
- activates a KXM schema. Implementations MUST NOT infer KXM behavior merely
58
- because these documents or examples are present.
64
+ Only schemas the shipped code loads are active, for example
65
+ `kxm.project.v1` and `kxm.workflow.v1`. Implementations MUST NOT infer KXM
66
+ behavior merely because these documents or examples are present.
59
67
 
60
68
  Every persisted KXM resource carries an exact schema identity. Additive
61
69
  changes require a new compatible schema revision; a semantic breaking change
@@ -1,5 +1,8 @@
1
1
  # Effects, idempotency, and recovery
2
2
 
3
+ > [!IMPORTANT]
4
+ > Planned: this contract is the target design. Today only Runtime gate steps record effect intents, observations and settlements, and a gate outcome that cannot be proven blocks the run as `blocked_uncertain`; the effect registry, adapters, receipt queries and leases around external effects are not wired.
5
+
3
6
  KXM provides at-least-once command delivery with effect-aware recovery. It does
4
7
  not claim exactly-once external execution.
5
8
 
@@ -28,8 +28,8 @@ does exist:
28
28
  store needs is declared in its current schema definition, so an empty store is
29
29
  valid without any upgrade step.
30
30
  - **Backups stay whole.** The hub state set is copied as documented in
31
- [`docs/operations.md`](../operations.md); a single `kxm.db` copy is not a
32
- backup.
31
+ [Back up and restore KXM](../operations/backup-and-restore.md); a single `kxm.db`
32
+ copy is not a backup.
33
33
  - **Out-of-range versions refuse without touching the file.** The stamp is read
34
34
  before anything that can modify the store, so a database this build refuses —
35
35
  older or newer — comes back byte-identical: neither its schema nor its
@@ -6,8 +6,9 @@
6
6
  > enforced with fail-closed `costBasis` requirement. Price catalog `.kxm/prices.yaml`
7
7
  > (`kxm.prices.v1`) is implemented, dated, and hashed. `kxm routing report`
8
8
  > is implemented (`plugins/kxm/src/routing.ts`) and ranks routes quality-first,
9
- > then cost per accepted attempt, never ranking unknown cost cheapest and
10
- > reporting metered, unmetered, and unknown populations separately. By default
9
+ > then cost per accepted attempt, ranking any route with an unknown-cost attempt
10
+ > after every route of equal quality that has none, and reporting metered,
11
+ > unmetered, and unknown populations separately. By default
11
12
  > `kxm routing report` and `kxm improve` read two sources: the current project's
12
13
  > Runtime event store (read-only) and then `.kxm/logs/telemetry.jsonl`; `--file`
13
14
  > reads only the named file (see [Readers](#readers-kxm-routing-report-and-kxm-improve)).
@@ -174,7 +175,7 @@ Fields carried on `RoutingRecordV2`:
174
175
  - Outcomes: `verifierOutcome` (`passed` | `warning` | `failed`), `finalOutcome` (`accepted` | `blocked` | `failed` | `pending`), `retries`, optional `transitions`, optional `humanInterventions`, optional `providerMetadata`.
175
176
  - Cost accounting: `costBasis` (`"metered" | "unmetered" | "unknown"`), `costUsd` (required when metered), optional `priceRef`.
176
177
 
177
- The KXM engine settle transaction appends a `routing.attempt.recorded` event carrying the v2 record and refuses to settle without a valid `costBasis`. Attempt dispatch enforces `limits.maxModelCost` against metered cost before invocation (`budget_model_cost`).
178
+ The KXM engine settle transaction appends a `routing.attempt.recorded` event carrying the v2 record and refuses to settle without a valid `costBasis`. Attempt dispatch enforces `limits.maxModelCost` against metered cost before invocation (`budget_model_cost`). No producer records `metered` cost today, so the cap cannot trip on a live run; the separate limit of 100 unmetered or unknown attempts per run (`budget_unmetered_attempts`) is the only cost-side limit that can stop one.
178
179
 
179
180
  What the engine writes on every settled attempt (`producerRoutingRecord` and the
180
181
  failure path in `settleMember`, `plugins/kxm/src/engine.ts`):
@@ -226,7 +227,7 @@ backfilled: they still resolve an outcome, but they group per run.
226
227
  - **Price catalog:** `.kxm/prices.yaml` (`kxm.prices.v1`, dated and hashed) defines input, output, cache-read, cache-write rates, and context tiers for active models. Missing rows or uncataloged models evaluate to `costBasis: "unknown"`.
227
228
  - **Ranked report:** `kxm routing report` (`plugins/kxm/src/routing.ts`, CLI command `kxm routing report`) groups records by `(harness, model, thinking, role)`.
228
229
  - **Ranking order:** Quality first (`verifyPassRate` descending, then `reworkRate` ascending where rework measures back-edge re-entries `transitions > 0`), followed by `costPerAcceptedUsd` ascending.
229
- - **Underquote prevention:** Routes with unknown cost are flagged (`*`) and **never ranked cheapest**, eliminating silent underquoting.
230
+ - **Underquote prevention:** Routes with unknown cost are flagged (`*`). Any unknown-cost attempt makes a route's cost a lower bound, so the route ranks after every route of equal quality that has no unknown-cost attempt. Displayed values are unchanged: a route that mixes unmetered and unknown-cost attempts still shows `$0` per accepted attempt, so read the flag and the population counts before comparing cost.
230
231
  - **Population separation:** Reports metered cost, unmetered attempt counts, unknown-cost attempt counts, and quota-exhausted attempt counts as separate metrics rather than a single misleading total.
231
232
  - **List prices flag:** Supports `--equivalent-list-cost` / `--list-prices` to display estimated list rates for comparison alongside actual recorded spend.
232
233
  - **Rework column:** reads `transitions`, which Runtime records never set, so Runtime rework shows up only as a resolved `reworked` outcome, which the report does not count as a pass.
@@ -263,6 +264,6 @@ store: there is no cross-worktree aggregation.
263
264
  ## Precedence
264
265
 
265
266
  [AGENTS.md](../../AGENTS.md) and
266
- [Tracking](../../plans/implementation-plan.md#tracking-working-tree-not-a-release)
267
+ Tracking (the implementation plan in the repository's `plans/` directory, which is not shipped)
267
268
  win where they differ from historical 2026-09-04 reviews.
268
269
  Issue 86 stays open for later-phase remainder (see Tracking).
@@ -0,0 +1,388 @@
1
+ # Assignment runner
2
+
3
+ The assignment runner, `scripts/assignment-run.mjs`, is how maintainers delegate
4
+ one unit of work on the KXM repository to a coding agent and accept the result
5
+ with proof: a bound manifest, a deterministic witness, and two independent
6
+ critics. This guide covers the loop, the files it writes, and the checks it
7
+ refuses to skip. It is maintainer tooling (tracked internally as issue 127), not
8
+ a KXM product feature.
9
+
10
+ > [!IMPORTANT]
11
+ > The runner is not the Runtime workflow engine (`kxm run`), not a supervised
12
+ > worker (`kxm agent worker`), and not a product dispatch adapter. Its
13
+ > `accepted.json` is a developer record. It is not human approval, hub peer
14
+ > evidence, CI success, or a merge.
15
+
16
+ ## Before you begin
17
+
18
+ - A clean control checkout of this repository whose `HEAD` is an ancestor of
19
+ `origin/main`. The runner loads `.kxm/roster.yaml` only from there.
20
+ - A separate worktree for the writer. `just worktree <unit>` creates one from
21
+ `origin/main`.
22
+ - [`just`](https://github.com/casey/just), plus the harness CLIs the roster
23
+ admits, installed and logged in. `node scripts/kxm.mjs harness list` shows
24
+ which are.
25
+ - A task directory whose final path segment equals the task ID. Every path you
26
+ pass to the runner must be absolute.
27
+
28
+ ## Roles and routes
29
+
30
+ The trusted roster policy in [`.kxm/roster.yaml`](../../.kxm/roster.yaml)
31
+ (`kxm.developer-roster.v1`) admits each route for specific roles and
32
+ permissions:
33
+
34
+ | Role | Admitted route (harness / model) | Vendor | Permission |
35
+ |---|---|---|---|
36
+ | `writer` | `grok` / `grok-4.6`; relief: `pi` / `openrouter/qwen/qwen3-coder-plus` | `xai`; `alibaba` | `edit` |
37
+ | `planner` | `claude` / `fable` | `anthropic` | `read-only` |
38
+ | `reviewer-arch` | `claude` / `fable` | `anthropic` | `read-only` |
39
+ | `reviewer-cli` | `codex` / `gpt-5.6-sol` | `openai` | `read-only` |
40
+
41
+ The policy itself enforces three rules. Critics are read-only. The two
42
+ designated critics have different vendors. No admitted writer shares a vendor
43
+ with a critic. A route outside the lineup, or a permission above its ceiling,
44
+ fails closed with `route_invalid`.
45
+
46
+ Each assignment has a kind, and the kind fixes its role:
47
+
48
+ | Kind | Role | Base | Witness |
49
+ |---|---|---|---|
50
+ | `plan` | `planner` | Clean | `verify` or `validate-ci` |
51
+ | `implement` | `writer` | Clean | `verify` only |
52
+ | `repair` | `writer` | Clean | `verify` only |
53
+ | `review-arch` | `reviewer-arch` | Clean or staged | `verify` or `validate-ci` |
54
+ | `review-cli` | `reviewer-cli` | Clean or staged | `verify` or `validate-ci` |
55
+
56
+ A clean base means `HEAD`, the index and the worktree are identical. Only
57
+ reviews may target a staged index.
58
+
59
+ ## The loop
60
+
61
+ A unit moves from a pinned plan through one writer, a fixed witness and two
62
+ critics to acceptance; a critic's `BLOCK` or a failed witness sends it back
63
+ through a repair.
64
+
65
+ ```mermaid
66
+ flowchart TD
67
+ P[plan-current<br/>pin the plan] -->|plan_ref| A[assign<br/>writer: implement]
68
+ A -->|completion.json| W[witness<br/>npm run verify]
69
+ W -->|passed| RA[review-arch<br/>read-only]
70
+ W -->|passed| RC[review-cli<br/>read-only]
71
+ RA -->|PASS| ACC[accept<br/>accepted.json]
72
+ RC -->|PASS| ACC
73
+ RA -->|BLOCK| R[repair<br/>writer, rework_of]
74
+ RC -->|BLOCK| R
75
+ W -->|failed| R
76
+ R -->|new completion| W
77
+ ```
78
+
79
+ Use this slim loop for daily work and for docs. The 13-step `fix` workflow in
80
+ `examples/project/` is a product fixture, not the developer loop.
81
+
82
+ > [!NOTE]
83
+ > These recipes never load a `.env` file from the working directory. An
84
+ > unreviewed file could otherwise set `NODE_OPTIONS` and run code before the
85
+ > runner validates anything. To use one, pass it explicitly:
86
+ > `just --dotenv-path /abs/.env assign /abs/manifest.json`.
87
+
88
+ ### 1. Pin the current plan
89
+
90
+ Every writer assignment binds to the current plan. Stamp the pointer, or advance
91
+ it with a generation check:
92
+
93
+ ```bash
94
+ just plan-current /abs/task-dir /abs/plan.md <sha256> <base-commit> <expected-generation>
95
+ ```
96
+
97
+ The runner writes `plan-current.json` (`kxm.plan-pointer.v1`) in the task
98
+ directory. It does not copy the plan. When the pointer advances, the previous
99
+ plan and pointer move to `plan-history/generation-<n>.md` and `.json`. A stale
100
+ `<expected-generation>` fails with `history_conflict`.
101
+
102
+ A `plan` or review assignment may instead use a `bootstrap` plan reference with
103
+ a reason, but only while no pointer exists.
104
+
105
+ ### 2. Dispatch the writer
106
+
107
+ Write a closed `kxm.assignment.v1` manifest. Unknown keys are refused.
108
+
109
+ `/abs/tasks/fix-improve-sources/asg-writer-1.json`:
110
+
111
+ ```json
112
+ {
113
+ "schema": "kxm.assignment.v1",
114
+ "task_id": "fix-improve-sources",
115
+ "assignment_id": "asg-writer-1",
116
+ "kind": "implement",
117
+ "harness": "grok",
118
+ "model": "grok-4.6",
119
+ "effort": "medium",
120
+ "permission": "edit",
121
+ "cwd": "/abs/kxm-fix-improve-sources",
122
+ "task_dir": "/abs/tasks/fix-improve-sources",
123
+ "base": { "kind": "clean", "commit": "<40-hex HEAD of cwd>" },
124
+ "plan_ref": { "kind": "current", "path": "/abs/plan.md", "sha256": "<64-hex>" },
125
+ "inputs": [],
126
+ "contract": {
127
+ "boundary": "plugins/kxm/src/improve-sources.ts and its test only",
128
+ "deliverables": ["The fix and one focused test"],
129
+ "witness": { "id": "verify" },
130
+ "deferred": []
131
+ },
132
+ "output_dir": "/abs/tasks/fix-improve-sources/asg-writer-1"
133
+ }
134
+ ```
135
+
136
+ Optional keys are `rework_of`, `timeout_ms` and `max_turns`. Each `inputs`
137
+ entry is `{ "path", "sha256" }`, and the runner checks the hash. Dispatch it:
138
+
139
+ ```bash
140
+ just assign /abs/tasks/fix-improve-sources/asg-writer-1.json
141
+ ```
142
+
143
+ The runner validates the manifest, the route and the base, writes
144
+ `pre-dispatch.json` before it spawns the harness, and records `completion.json`
145
+ (`kxm.assignment-completion.v1`) or `refusal.json` when the harness exits.
146
+
147
+ ### 3. Run the witness
148
+
149
+ Stage everything the writer produced, including rebuilt `dist` and other
150
+ generated files, then run the witness:
151
+
152
+ ```bash
153
+ git -C /abs/kxm-fix-improve-sources add -A
154
+ just witness /abs/tasks/fix-improve-sources/asg-writer-1
155
+ ```
156
+
157
+ The witness refuses with `dirty_baseline` while anything is unstaged or
158
+ untracked. It re-runs the gate named in the manifest's `contract.witness.id`
159
+ without a shell: `npm run verify` for `verify`, `npm run validate:ci` for
160
+ `validate-ci`. Writers always use `verify`. The receipt binds `HEAD` and the
161
+ staged index tree. If the gate leaves the worktree or index different from
162
+ before, for example an unstaged rebuild, the witness fails with
163
+ `candidate_changed`.
164
+
165
+ Each run writes an immutable receipt (`kxm.assignment-witness.v1`) and moves
166
+ the latest pointer:
167
+
168
+ - `witness/history/<receipt-id>.json`, where the ID is `w-<UTC timestamp>`;
169
+ - `witness/history/<receipt-id>/<gate>.log`, the captured gate output;
170
+ - `witness/latest.json` (`kxm.assignment-witness-latest.v1`), which names the
171
+ newest receipt and its sha256.
172
+
173
+ Acceptance reads only the receipt `latest.json` points to, and requires
174
+ `result: "passed"`.
175
+
176
+ ### 4. Dispatch both critics
177
+
178
+ Dispatch one `review-arch` and one `review-cli` assignment against the
179
+ witnessed tree: either the staged index (`base.kind: "staged"` with its
180
+ `index_tree`) or, after you commit it, a clean base. Each critic writes `completion.json` with a
181
+ `critic` verdict of `PASS` or `BLOCK` and the `judged_tree` it reviewed.
182
+
183
+ ### 5. Accept
184
+
185
+ Commit the exact witnessed tree, then bind the commit and both `PASS` records:
186
+
187
+ ```bash
188
+ just accept /abs/tasks/fix-improve-sources <commit-sha> \
189
+ /abs/tasks/fix-improve-sources/asg-writer-1 \
190
+ /abs/tasks/fix-improve-sources/asg-review-arch-1 \
191
+ /abs/tasks/fix-improve-sources/asg-review-cli-1
192
+ ```
193
+
194
+ To record an observed pull request or CI run, call the script directly, since
195
+ the recipe does not pass those flags:
196
+
197
+ ```bash
198
+ node scripts/assignment-run.mjs accept \
199
+ --task-dir /abs/tasks/fix-improve-sources \
200
+ --commit <commit-sha> \
201
+ --record-dir /abs/tasks/fix-improve-sources/asg-writer-1 \
202
+ --critic /abs/tasks/fix-improve-sources/asg-review-arch-1 \
203
+ --critic /abs/tasks/fix-improve-sources/asg-review-cli-1 \
204
+ --observed-pr <pr-id> \
205
+ --observed-ci <ci-id>
206
+ ```
207
+
208
+ `accept` prints JSON and takes no `--json` flag. It checks, in order, that:
209
+
210
+ 1. the trusted roster policy loads and validates, before anything is written;
211
+ 2. the commit exists and its tree equals the witnessed tree;
212
+ 3. the writer record matches the latest passed witness receipt;
213
+ 4. exactly the two designated critics are present, on their admitted routes;
214
+ 5. both critics judged the accepted tree and returned `PASS`;
215
+ 6. no unresolved `BLOCK` for that tree remains in the task directory;
216
+ 7. the writer and both critics come from three different vendors.
217
+
218
+ It then writes `accepted.json` (`kxm.task-accepted.v1`). The file is immutable;
219
+ a second accept in the same task directory fails with `accepted_exists`.
220
+
221
+ ## Repair after a BLOCK or a failed witness
222
+
223
+ The runner never schedules a repair or fails over to another model on its own.
224
+ You decide, then dispatch:
225
+
226
+ 1. Write a new manifest with `"kind": "repair"` and `"rework_of"` set to the
227
+ assignment ID it reworks. The runner checks that the earlier assignment has a
228
+ `completion.json` in the same task directory, with the same task ID.
229
+ 2. Dispatch it with `just assign`, then run `just witness` on the new record.
230
+ 3. Dispatch fresh critics against the new tree, and accept.
231
+
232
+ A `BLOCK` stops acceptance only for the tree it judged. A repair that changes
233
+ the tree leaves the old `BLOCK` behind. If a critic re-reviews the same tree,
234
+ give its new assignment a `rework_of` chain back to the `BLOCK` review; only
235
+ then is that `BLOCK` resolved.
236
+
237
+ ## Records in the task directory
238
+
239
+ ```text
240
+ <task-dir>/
241
+ ├── plan-current.json # current plan pointer (kxm.plan-pointer.v1)
242
+ ├── plan-history/ # superseded plans and pointers, by generation
243
+ ├── accepted.json # acceptance record (kxm.task-accepted.v1), immutable
244
+ └── <assignment-id>/ # one record directory per assignment
245
+ ├── manifest.json # the bound manifest
246
+ ├── prompt.md # the rendered prompt
247
+ ├── output-schema.json # the structured-output schema given to the harness
248
+ ├── pre-dispatch.json # written before spawn (kxm.assignment-dispatch.v1)
249
+ ├── completion.json # kxm.assignment-completion.v1, or refusal.json
250
+ ├── routing-record.json # read by kxm routing report
251
+ ├── telemetry.jsonl # usage, latency and cost
252
+ ├── runner-errors.jsonl # bounded failure codes
253
+ ├── witness/ # latest.json and history/<receipt-id>.json
254
+ ├── attribution/ # latest.json and history/, private notes
255
+ └── cost-observation.json # imported, cost-only
256
+ ```
257
+
258
+ The harness output files (`completion.json`, `routing-record.json`,
259
+ `telemetry.jsonl`) land in the manifest's `output_dir`, which is usually the
260
+ record directory. `completion.json` and `accepted.json` have closed schemas: do
261
+ not add fields to them.
262
+
263
+ ## Notes, costs and reports
264
+
265
+ Record friction or a model regression as a private note, without touching any
266
+ completion:
267
+
268
+ ```bash
269
+ just attribute /abs/task-dir /abs/record-dir <class> /abs/note.txt
270
+ ```
271
+
272
+ The class is `orchestration`, `model`, `environment` or `unclassified`. Each
273
+ note is a hash-linked entry under the record's `attribution/`. Notes never grant
274
+ tools, waive a witness, or count as approval.
275
+
276
+ Import a cost observation for a run whose native telemetry was not captured,
277
+ such as a subscription session:
278
+
279
+ ```bash
280
+ just observe-cost /abs/task-dir /abs/observation.json
281
+ ```
282
+
283
+ The record (`kxm.cost-observation.v1`) is cost-only. It cannot mint witness
284
+ proof or authorize acceptance.
285
+
286
+ Summarize a task's attempts, rework and spend:
287
+
288
+ ```bash
289
+ just change-report /abs/task-dir
290
+ ```
291
+
292
+ The report (`kxm.change-report.v1`) keeps provider-reported spend, list-price
293
+ estimates, unmetered usage and unknown cost apart. A missing value stays
294
+ missing: it is never counted as zero. Failed and interrupted attempts are
295
+ listed, and nothing is ranked.
296
+
297
+ ## Recover an incomplete record
298
+
299
+ If the harness finished but the runner failed to write the routing record or
300
+ the telemetry line, the witness refuses the record with `recording_unresolved`.
301
+ Rebuild both from `completion.json`, then re-run the witness. There is no recipe
302
+ for this step, so call the script directly:
303
+
304
+ ```bash
305
+ node scripts/assignment-run.mjs observe --record-dir /abs/tasks/fix-improve-sources/asg-writer-1
306
+ ```
307
+
308
+ It writes `recording-resolved.json` in the record directory. It never changes
309
+ `completion.json`.
310
+
311
+ ## Transport-only recipes
312
+
313
+ `just impl|plan|review-arch|review-cli` and `just dispatch` send one
314
+ `kxm.harness-request.v1` envelope through `scripts/harness-run.mjs` and print a
315
+ `kxm.harness-result.v2` envelope. They are harness transport only. They mint no
316
+ assignment, witness or acceptance proof, so their output cannot be accepted.
317
+
318
+ ### Preflight refusals
319
+
320
+ Before any auth probe or model call, the helper checks the request. A refusal
321
+ prints a result with `stage: "preflight"` and `errorCode: "preflight_failed"`,
322
+ and exits with status 2. It refuses when:
323
+
324
+ - the envelope is not `kxm.harness-request.v1`, carries an unknown field, or
325
+ omits `harness`, `role`, `model`, `permission` or `prompt_file`; it never
326
+ falls back to CLI defaults;
327
+ - the harness is `kimi`, `gemini` or `deepseek` (unverified) or unknown, or the
328
+ role, permission or model is not in that harness's route table; `grok`
329
+ refuses `read-only`;
330
+ - a Pi model is not `provider/id`, names a braked native vendor (`anthropic`,
331
+ `openai`, `xai`, `moonshot`, `google` or `deepseek`), or uses a provider
332
+ other than `openrouter`, `nous-portal` or `antigravity`; an `antigravity`
333
+ model must be a two-segment Gemini ID, and an aggregator model's vendor
334
+ segment must not be a braked native vendor (`x-ai`, `moonshotai` and
335
+ `google-ai` count as `xai`, `moonshot` and `google`);
336
+ - a Pi `edit` request comes from a role other than `experiment` or `writer`, or
337
+ a Pi writer asks for anything but `openrouter/qwen/qwen3-coder-plus` with
338
+ `edit`;
339
+ - the request asks for hooks or skills, sets `max_cost_usd` for a harness other
340
+ than Claude, or sets `max_turns` for a harness other than Grok;
341
+ - the harness binary is missing, or, on Windows, the launcher is a `.cmd`,
342
+ `.bat`, `.ps1` or extensionless shim rather than an absolute `.exe`
343
+ (`unsupported-launcher`).
344
+
345
+ A failed login probe also exits with status 2, with `stage: "auth"` and
346
+ `errorCode: "auth_failed"`. Failures after the harness starts, such as
347
+ `spawn_failed` or `timed_out`, come back as an ordinary failed result with exit
348
+ status 1.
349
+
350
+ ## Failure codes
351
+
352
+ The runner fails closed with a bounded code from `RUNNER_CODES`. The common
353
+ ones:
354
+
355
+ | Code | Trigger | Fix |
356
+ |---|---|---|
357
+ | `route_invalid` | Route not in the role's lineup, permission above its ceiling, or the roster policy cannot load | Run from a clean control checkout on `origin/main`; check the lineup |
358
+ | `base_invalid` | `base.commit` is not `HEAD`, the tree is dirty, or a writer targets a staged index | Commit or stash elsewhere; writers need a clean base |
359
+ | `plan_ref_invalid` | The plan hash or path does not match `plan-current.json` | Advance the pointer with `just plan-current` |
360
+ | `rework_invalid` | `rework_of` names no completed assignment in this task | Point it at an existing record directory's assignment ID |
361
+ | `witness_failed` | The fixed gate exited non-zero | Fix the failures and run a repair |
362
+ | `dirty_baseline` | Unstaged or untracked changes when the witness starts | Stage the candidate with `git add -A`, then re-witness |
363
+ | `candidate_changed` | The gate left the index or worktree different, or another process changed it | Stage generated files; keep other writers out of the worktree |
364
+ | `commit_tree_mismatch` | The commit's tree is not the witnessed tree | Commit exactly the witnessed tree |
365
+ | `critic_invalid` | A critic is missing, duplicated, on the wrong route, or shares a vendor | Dispatch the two designated critics |
366
+ | `critic_block` | An unresolved `BLOCK` exists for the accepted tree | Repair, or resolve it through a `rework_of` chain |
367
+ | `accepted_exists` | `accepted.json` already exists | Use a new task directory for new work |
368
+
369
+ ## Test coverage is thin
370
+
371
+ The runner's own suites were removed with an earlier refactor. Today the tests
372
+ cover the edges only:
373
+
374
+ - `test/core/harness-run.test.ts` pins every `just` recipe body that calls the
375
+ runner, and checks that every `just <verb>` in this guide is a real recipe.
376
+ - `test/core/roster-policy.test.ts` covers the roster validator.
377
+ - `test/core/policy-draft.test.ts` covers the passive policy-draft schemas.
378
+
379
+ No test runs `run`, `witness` or `accept` end to end. Treat changes to
380
+ `scripts/assignment-run.mjs` as high risk, and check them by running a real
381
+ unit through the loop.
382
+
383
+ ## Related
384
+
385
+ - [Develop KXM](development.md): the commit gate the witness runs
386
+ - [CI and release](ci-and-release.md): what runs after you push
387
+ - [Harness routing](../reference/harness-routing.md): harness and model pairing
388
+ - [Configuration reference](../reference/config-reference.md#kxmrosteryaml-kxmdeveloper-rosterv1): the roster file