@kontextmind/kxm 0.7.91 → 0.7.93

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (92) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.kxm/workflows/default.yaml +1 -1
  3. package/CHANGELOG.md +212 -0
  4. package/README.md +3 -0
  5. package/docs/README.md +3 -0
  6. package/docs/agent-skills.md +123 -60
  7. package/docs/architecture.md +5 -2
  8. package/docs/cli-reference.md +3527 -0
  9. package/docs/config-reference.md +1943 -0
  10. package/docs/configuration.md +30 -4
  11. package/docs/continuous-improvement.md +122 -10
  12. package/docs/contracts/routing.md +95 -11
  13. package/docs/harness-routing.md +616 -0
  14. package/docs/kxm-handbook.md +106 -19
  15. package/docs/templates/README.md +1 -1
  16. package/docs/test-matrix.md +12 -6
  17. package/docs/troubleshooting.md +2 -2
  18. package/examples/project/.kxm/workflows/fix.yaml +1 -1
  19. package/examples/project/.kxm/workflows/improve.yaml +1 -1
  20. package/package.json +1 -1
  21. package/plugins/kxm/.claude-plugin/plugin.json +9 -10
  22. package/plugins/kxm/README.md +238 -56
  23. package/plugins/kxm/dist/claude-hook.js +10083 -0
  24. package/plugins/kxm/dist/cli.js +2487 -1848
  25. package/plugins/kxm/dist/client.js +64 -0
  26. package/plugins/kxm/dist/core.js +102 -9
  27. package/plugins/kxm/dist/extension.js +210 -68
  28. package/plugins/kxm/dist/mcp-server.js +217 -40
  29. package/plugins/kxm/dist/runtime-supervisor.js +1628 -157
  30. package/plugins/kxm/dist/runtime.js +1874 -298
  31. package/plugins/kxm/dist/server.js +416 -82
  32. package/plugins/kxm/package.json +1 -1
  33. package/plugins/kxm/skills/hints.json +1 -1
  34. package/plugins/kxm/skills/kxm/SKILL.md +48 -24
  35. package/plugins/kxm/skills/kxm/references/protocol.md +3 -3
  36. package/plugins/kxm/skills/kxm-context-memory/SKILL.md +67 -21
  37. package/plugins/kxm/skills/kxm-definitions/SKILL.md +9 -0
  38. package/plugins/kxm/skills/kxm-harness-auth/SKILL.md +82 -16
  39. package/plugins/kxm/skills/kxm-harvest/SKILL.md +1 -1
  40. package/plugins/kxm/skills/kxm-hub-ops/SKILL.md +55 -27
  41. package/plugins/kxm/skills/kxm-insights/SKILL.md +1 -1
  42. package/plugins/kxm/skills/kxm-mind/SKILL.md +2 -2
  43. package/plugins/kxm/skills/{kxm-setup → kxm-mind-setup}/SKILL.md +4 -4
  44. package/plugins/kxm/skills/kxm-peer/SKILL.md +68 -93
  45. package/plugins/kxm/skills/kxm-project-setup/SKILL.md +156 -23
  46. package/plugins/kxm/skills/kxm-projects/SKILL.md +1 -1
  47. package/plugins/kxm/skills/kxm-protocol/SKILL.md +1 -1
  48. package/plugins/kxm/skills/kxm-query/SKILL.md +1 -1
  49. package/plugins/kxm/skills/kxm-routing-improve/SKILL.md +74 -15
  50. package/plugins/kxm/skills/kxm-runs/SKILL.md +46 -17
  51. package/plugins/kxm/skills/kxm-session/SKILL.md +64 -36
  52. package/plugins/kxm/skills/kxm-skill-lifecycle/SKILL.md +44 -15
  53. package/plugins/kxm/skills/kxm-tasks/SKILL.md +16 -4
  54. package/plugins/kxm/skills/kxm-triage/SKILL.md +1 -1
  55. package/plugins/kxm/skills/kxm-work/SKILL.md +1 -1
  56. package/plugins/kxm/skills/kxm-workflow/SKILL.md +60 -19
  57. package/plugins/kxm/src/arbiter.ts +67 -22
  58. package/plugins/kxm/src/autocomplete.ts +1 -1
  59. package/plugins/kxm/src/claude-hook.ts +192 -0
  60. package/plugins/kxm/src/cli/project.ts +11 -5
  61. package/plugins/kxm/src/cli/system.ts +85 -13
  62. package/plugins/kxm/src/cli/types.ts +4 -1
  63. package/plugins/kxm/src/cli/workflows.ts +18 -16
  64. package/plugins/kxm/src/cli.ts +23 -13
  65. package/plugins/kxm/src/client.ts +15 -4
  66. package/plugins/kxm/src/commands.ts +19 -9
  67. package/plugins/kxm/src/config.ts +42 -7
  68. package/plugins/kxm/src/context-packet.ts +14 -2
  69. package/plugins/kxm/src/context.ts +16 -5
  70. package/plugins/kxm/src/dispatch-context.ts +286 -0
  71. package/plugins/kxm/src/engine-plan.ts +40 -0
  72. package/plugins/kxm/src/engine.ts +138 -6
  73. package/plugins/kxm/src/hub-env.ts +17 -1
  74. package/plugins/kxm/src/hub.ts +92 -29
  75. package/plugins/kxm/src/improve-sources.ts +228 -0
  76. package/plugins/kxm/src/improve.ts +325 -140
  77. package/plugins/kxm/src/local-snapshot.ts +101 -42
  78. package/plugins/kxm/src/mcp-server.ts +129 -30
  79. package/plugins/kxm/src/memory.ts +43 -20
  80. package/plugins/kxm/src/project-config.ts +25 -0
  81. package/plugins/kxm/src/protocol.ts +11 -0
  82. package/plugins/kxm/src/relevance.ts +138 -0
  83. package/plugins/kxm/src/retrospective.ts +16 -10
  84. package/plugins/kxm/src/runtime-service.ts +8 -1
  85. package/plugins/kxm/src/runtime-supervisor.ts +16 -2
  86. package/plugins/kxm/src/session-token-hint.ts +17 -0
  87. package/plugins/kxm/src/suggest.ts +7 -7
  88. package/plugins/kxm/src/workflow-manager.ts +80 -78
  89. package/plugins/kxm/src/workflow.ts +202 -12
  90. package/scripts/build-runtime.mjs +7 -1
  91. package/scripts/check-generated.mjs +1 -0
  92. package/scripts/emit-codex-artifacts.mjs +1 -1
@@ -19,6 +19,27 @@ because it is in the file: an unknown `hub.autoStart` fails closed to the
19
19
  default, and an unset `defaults.harness` means Pi rather than the first harness
20
20
  in the catalog.
21
21
 
22
+ ### Improvement settings (`improvement.*`)
23
+
24
+ `kxm improve` reads these keys from the project it runs in (and the user scope) and
25
+ reports them in its output. They are normalized field by field whenever the
26
+ configuration loads, so a typo falls back to the default instead of changing what the
27
+ report says:
28
+
29
+ | Key | Default | Accepted values; anything else becomes the default |
30
+ |---|---:|---|
31
+ | `improvement.promotionPolicy` | `manual_pr` | `manual_pr`, `critic_quorum`, or `auto_threshold` |
32
+ | `improvement.telemetryHalfLifeDays` | `14` | A number greater than 0 and at most 3650 |
33
+ | `improvement.autoThreshold.minRuns` | `10` | An integer from 1 to 1,000,000 |
34
+ | `improvement.autoThreshold.minPassRate` | `0.95` | A number from 0 to 1 |
35
+ | `improvement.autoThreshold.minCostSavings` | `0.5` | A number of at least 0 |
36
+
37
+ None of these values can authorize or activate anything. The policy only selects which
38
+ review-readiness rule `kxm improve` reports for each candidate, and every policy ends at
39
+ an operator PR. The half-life weights report rows (`weightedRecurrence`) and never
40
+ decides whether a group is a candidate. See
41
+ [Continuous improvement](continuous-improvement.md#coded-repeats-kxm-improve).
42
+
22
43
  Project identity and repository bindings are separate, Git-tracked files under
23
44
  `.kxm/` (`project.yaml`, `roster.yaml`, `routes.yaml`, `gates.yaml`, `prices.yaml`,
24
45
  `roles/`, `workflows/`). They are configuration reviewed in a PR, not personal
@@ -95,6 +116,8 @@ rewrites it.
95
116
  | `KXM_WEBHOOK_WORKFLOWS` | None | Inline JSON array of signed webhook workflow definitions |
96
117
  | `KXM_WEBHOOK_WORKFLOWS_FILE` | None | Path to the workflow-definition JSON file |
97
118
 
119
+ The hub's structured log (`KXM_LOG_PATH`) records the size of a context request, not its text: `context_packet_assembled` carries `taskChars`, `taskTokens` (distinct words after stopword removal) and `matchedCandidates` beside the selected ids, provenance summary and token estimate, and `context_recall` carries `queryChars`, `queryTokens`, `limit` and `results`. The task and query text appear only in the caller's own response.
120
+
98
121
  The hub refuses a non-loopback bind without `KXM_AUTH_TOKEN`. Use a long random administrative token even when project tokens are configured, because administrative endpoints such as `/metrics` require it outside loopback.
99
122
 
100
123
  When `KXM_AUTH_TOKEN` is unset, `kxm hub start` resolves credentials from the
@@ -190,7 +213,8 @@ read from source; it does not invent defaults.
190
213
  | Agent stale threshold | 30 seconds |
191
214
  | Client request timeout | 15 seconds |
192
215
  | Default message TTL | 24 hours |
193
- | Default `kxm_await` timeout | 30 minutes |
216
+ | `kxm_await` wait | 60 seconds (default and maximum) |
217
+ | Default `kxm_fanout` local wait | 30 minutes |
194
218
  | Default workflow signal wait | 24 hours |
195
219
  | Workflow signal wait range | 1 second to 30 days |
196
220
 
@@ -300,12 +324,14 @@ The current hub command groups are `agent`, `session`, `workflow`, `gate`, `hub`
300
324
  | `kxm dash` | Open the read-only SSE observer dashboard; non-TTY output is one ANSI-free snapshot |
301
325
  | `kxm hub start` | Start the KXM hub in the foreground |
302
326
  | `kxm hub stop` | Request managed hub and worker shutdown |
303
- | `kxm improve` | Bucket `.kxm/logs/telemetry.jsonl` events into a proposed-only improvement report; does not read the workflow journal |
327
+ | `kxm improve` | Propose coded-repeat candidates from routing records: the current project's Runtime event store (read-only) and `.kxm/logs/telemetry.jsonl`, or only the file named by `--file`. Prints the sources it read, writes proposed candidates under `.kxm/candidates/`, and reports promotion readiness under `improvement.promotionPolicy`; nothing is applied and no policy authorizes. Does not read the workflow journal |
304
328
  | `kxm context get \| recall \| state \| episode \| promote \| explain \| wiki-compile \| wiki-lint` | Context operating system: role-aware packets, metadata search, temporal state, episodes, evidence-backed lineage (`explain`), wiki compile/lint |
305
329
  | `kxm skills` | Governed skill candidate lifecycle |
306
- | `kxm memory sync` | Regenerate the read-only project-memory block — authored facts from `.kxm/memory/` only — between `<!-- kxm:memory:start -->` and `<!-- kxm:memory:end -->` in whichever of `AGENTS.md`, `CLAUDE.md` and `GEMINI.md` already exist in the current directory. Text outside the markers is never rewritten; a file without markers gets the block appended. Missing files are reported as `missing` and **not created** — which harness a project uses is its own choice — and with none present sync exits 1 without writing. `--json` reports `updated`, `unchanged` and `missing` |
330
+ | `kxm memory sync` | Regenerate the read-only project-memory block — authored facts from `.kxm/memory/` only — between `<!-- kxm:memory:start -->` and `<!-- kxm:memory:end -->` in whichever of `AGENTS.md`, `CLAUDE.md` and `GEMINI.md` already exist in the current directory. Text outside the markers is never rewritten; a file without markers gets the block appended. Each file must hold exactly one start marker followed by exactly one end marker, or neither: an orphan marker, an end before its start, or a second block makes sync exit 1 naming the file and the problem, and **no file is written**, the well-formed ones included. Missing files are reported as `missing` and **not created** — which harness a project uses is its own choice — and with none present sync exits 1 without writing. `--json` reports `updated`, `unchanged` and `missing` |
331
+
332
+ Telemetry events carry a `target` label: `project` whenever a project or workflow identity is present, and `cli` for unscoped operator behavior. `KXM_IMPROVE_TARGET=cli` or `KXM_IMPROVE_TARGET=project` overrides that label when an event is written. No report reads the label, and `kxm improve` has no `--target` option.
307
333
 
308
- Improvement telemetry is classified as `project` whenever a project or workflow identity is present, and as `cli` for unscoped operator behavior. Set `KXM_IMPROVE_TARGET=cli` or `KXM_IMPROVE_TARGET=project` only when an operator needs to override that generic classification; this changes report bucketing, not workflow state.
334
+ Journal entries recorded with `kxm workflow record` or `kxm_workflow_record` take one of ten categories (`plan`, `decision`, `contradiction`, `error`, `lesson`, `observation`, `hypothesis`, `experiment`, `state-change`, `skill-candidate`). `--stage-id` binds the entry to a stage; the hub derives the attempt, and the area defaults to the stage's declared area. The journal covers hub webhook runs only; a `kxm run` id is refused with `workflow_not_found`.
309
335
 
310
336
  The gate group contains exactly the five implemented gates listed above. Names declared in a workspace `gates.json` that do not map to one of them (for example `quality`, `git-commit`, `jira-fetch`) are records with no runner; there is no `kxm gate run <name>`.
311
337
 
@@ -5,11 +5,16 @@ Every workflow run produces two distinct records:
5
5
  - operational events for service health and delivery;
6
6
  - a structured journal for plans, decisions, contradictions, errors, lessons, observations, hypotheses, experiments, state changes, and skill candidates.
7
7
 
8
- `kxm improve` buckets redacted `.kxm/logs/telemetry.jsonl` events into a
9
- proposed-only report and does not read the workflow journal. Journal capture
10
- uses `kxm_workflow_record`. Terminal runs export a bounded retrospective under
11
- `.kxm/assets/retrospectives`; re-export from durable local state with
12
- `kxm workflow export`.
8
+ The journal and retrospective loop covers hub webhook runs (signed webhooks and
9
+ `kxm workflow start`). Runs started with `kxm run` on the Runtime have no journal or
10
+ retrospective yet: `kxm_workflow_record` answers `workflow_not_found` for a Runtime
11
+ run id. Journal capture uses `kxm_workflow_record`. Terminal runs export a bounded
12
+ retrospective under `.kxm/assets/retrospectives`; re-export from durable local state
13
+ with `kxm workflow export`.
14
+
15
+ `kxm improve` is a separate loop. It reads routing records (the Runtime's settled
16
+ attempts and telemetry), not the journal, and proposes coded repeats; see
17
+ [Coded repeats](#coded-repeats-kxm-improve).
13
18
 
14
19
  Journal entries carry an improvement area, severity, evidence links, relationships to other entries, and (when stage-bound) run/stage/attempt provenance. The design preserves disagreement instead of flattening it into a single final answer.
15
20
 
@@ -40,15 +45,25 @@ Use `kxm_workflow_record` during the run, not only in a final retrospective:
40
45
  - record a `state-change` when an authoritative project fact changes;
41
46
  - record a `skill-candidate` only with verified run/receipt evidence; candidates never become promoted skills without a protected evaluation.
42
47
 
48
+ `kxm_workflow_record` (and `kxm workflow record`) accepts all ten categories above.
49
+ Pass `stageId` to bind an entry to the stage it is about. The hub, never the caller,
50
+ derives the attempt from the stage's state: the current attempt for an in-progress or
51
+ waiting stage, the last attempt consumed for a finished stage, and none for a pending
52
+ stage that has not run. When `stageId` names a stage that declares an area, `area` may
53
+ be omitted and defaults to that area; otherwise `area` is required, and a request with
54
+ neither is refused with `invalid_improvement_area`. A `stageId` that is not part of the
55
+ run is refused with `invalid_journal_relation`. `lesson` and `skill-candidate` entries
56
+ still require evidence references.
57
+
43
58
  ## Governed promotion
44
59
 
45
60
  `skill-candidate`, `hypothesis`, and `experiment` entries participate in a governed lifecycle: `proposed` → `approved` | `rejected` | `quarantined`. Promotion is an append-only, admin-controlled decision (`POST /v1/journal/:id/promotion`) that requires durable evidence references; the author of an entry can never decide its promotion, and terminal states never re-open. Promotion changes the learning lifecycle of an entry — never gates, workflow policy, or permissions.
46
61
 
47
- The native Pi extension automatically records failed tool results while a webhook workflow is active. The hub also records stage warnings/failures, prompt expiry, and premature coordinator settlement. Agents must still record semantic errors such as a false assumption, rejected design, flaky result, or external integration mismatch.
62
+ The native Pi extension automatically records failed tool results while a webhook workflow is active. The hub also records checkpoint warnings and failures, transition-budget exhaustion, signed signal results, external-wait timeouts, prompt expiry, degraded-quorum approvals, and premature coordinator settlement, each bound to its stage and attempt. Agents must still record semantic errors such as a false assumption, rejected design, flaky result, or external integration mismatch.
48
63
 
49
64
  Never put secrets or unnecessary prompt contents in the journal. Evidence should be durable references such as test names, logs, commits, pull requests, Jira issues, check runs, or documentation paths. Failed tools record an allowlisted diagnostic class, not stdout.
50
65
 
51
- Every terminal workflow automatically exports a bounded retrospective under `.kxm/assets/retrospectives`. Re-export one from durable local state with `kxm workflow export <runId>`; `--input <snapshot.json>` remains available for offline imports. Files stay `reviewDecision=proposed` until a human or coordinator records an explicit decision. Export never edits workflow JSON or weakens gates.
66
+ Every terminal workflow automatically exports a bounded retrospective under `.kxm/assets/retrospectives`, and exports it again when a journal entry is recorded or a promotion decided after the run ended. Re-export one from durable local state with `kxm workflow export <runId>`; `--input <snapshot.json>` remains available for offline imports. `recurringErrorClasses` counts `error` entries only. `proposedImprovements` holds up to 12 of the run's error and lesson entries, merged and ranked the same way as the weekly signals below, each with a success measure that names the signal key. Files stay `reviewDecision=proposed` until a human or coordinator records an explicit decision. Export never edits workflow JSON or weakens gates.
52
67
 
53
68
  Runs with peer policies add an optional metadata-only evidence audit while
54
69
  retaining the `pi-mesh.retrospective.v1` schema. It records each requirement's
@@ -68,13 +83,37 @@ The retrospective stage reviews journal entries, groups contributing causes, and
68
83
 
69
84
  ### Weekly
70
85
 
71
- Call `kxm_improvement_report` and review the top errors, contradictions, and lessons in each area. Merge duplicates while retaining source run IDs. Rank candidates using:
86
+ Call `kxm_improvement_report`. Besides the per-area counts it returns `signals`: the
87
+ project's errors, contradictions, lessons and skill candidates, already merged across
88
+ runs and ranked. Review the top signals, check that each merge groups entries that
89
+ belong together, and pick what to trial.
90
+
91
+ - Entries merge when they share a key: first an evidence class (`class:<name>` in the
92
+ entry's evidence), then, for errors, the workflow definition and stage, then the
93
+ summary after redaction and normalization (ids, timestamps, hex strings and numbers
94
+ are folded). Text is redacted before it becomes a key, so a key never carries a raw
95
+ summary. Each signal keeps up to 16 source run IDs and entry IDs.
96
+ - Only errors, open contradictions (not yet resolved by a related decision or lesson),
97
+ lessons, and skill candidates still `proposed` count.
98
+ - `frequency` is the number of distinct runs, counted over the runs the hub still
99
+ retains: terminal runs and their journal are purged 7 days after they end.
72
100
 
73
101
  ```text
74
- priority = frequency × severity × workflow cost × confidence
102
+ priority = frequency × severity weight × workflow cost × confidence
75
103
  ```
76
104
 
77
- Do not let frequency alone dominate security or data-loss risk.
105
+ - The severity weight is 3 for `error`, 2 for `warning` and 1 for `info`, taken from the
106
+ most severe entry in the signal.
107
+ - Workflow cost is the mean number of run attempts (stage attempts plus transitions)
108
+ over the signal's known runs. It is not dollars. When none of the runs is known it
109
+ counts as 1 and the signal reports `costBasis: "unknown"`.
110
+ - Confidence is 0.5 plus half the share of the signal's entries that cite evidence.
111
+
112
+ Security signals (area `security`, or evidence class `invalid_auth`,
113
+ `invalid_identity` or `signal_mismatch`) rank ahead of every priority. Remaining ties
114
+ break on frequency, then key, never on entry ID or insertion order. There is no
115
+ data-loss override yet, because no deterministic data-loss marker exists, so review
116
+ data-loss risk by hand rather than trusting the order.
78
117
 
79
118
  ### Per release
80
119
 
@@ -89,9 +128,82 @@ Select a small improvement batch. For each proposal:
89
128
  7. Adopt, revise, or roll back the proposal.
90
129
  8. Record the decision and result in a subsequent workflow journal.
91
130
 
131
+ ## Coded repeats (`kxm improve`)
132
+
133
+ `kxm improve` (the same as `kxm improve report`) looks for agent steps that a script,
134
+ test or workflow `gate` could do as well as a model. It proposes; it never applies.
135
+
136
+ **Sources.** Inside a KXM project it reads the project's Runtime event store,
137
+ `<state>/runtime/projects/<key>/run-events.db`, and then `.kxm/logs/telemetry.jsonl`.
138
+ The key comes from the checkout's real path, so each checkout and worktree has its own
139
+ store and the report covers only the one it runs in. The store is opened read-only for
140
+ one query over its events table; `kxm improve` never creates, writes or migrates it. A
141
+ telemetry record whose `attemptId` the store already supplied is dropped. `--file
142
+ <path>` reads only that file. Outside a project only telemetry is read. The output lists
143
+ every source with its path, whether it exists, and its counts (`records`,
144
+ `skippedInvalid`, `excludedSimulated`, `undecided`, `duplicatesDropped`). An unreadable
145
+ store stops the command with `improve_source_unreadable` and the path (exit 1).
146
+
147
+ **Identity.** Records group by workflow, step, agent role and ask. Runtime records carry
148
+ the engine-reserved `workflowId` and `askSha256` keys. The ask digest covers the
149
+ workflow, step, step kind, agent, instructions, outcomes and required evidence keys, so
150
+ it is the same for one step across runs and ignores the run, attempt, model and context
151
+ packet. Records without those keys fall back to a workflow definition digest or the run
152
+ ID, and to `rolePromptSha256`. No prompt text is read or stored: `objectiveSha256` is
153
+ the digest of the run's prompt.
154
+
155
+ **Outcomes.** A Runtime record stores `finalOutcome` only when settlement already knows
156
+ it: `blocked` for a back edge, and `failed` for a producer error, an undeclared outcome
157
+ or a failing terminal. Everything else is resolved when the report reads the event log,
158
+ and never written back: a later entry into the same step makes the attempt `reworked`, a
159
+ completed run makes it `accepted`, a failed run makes it `failed`, and a cancelled or
160
+ still-running run leaves it undecided. Undecided records are counted and left out of the
161
+ pass rate. Attempts from simulated drives are excluded and counted.
162
+
163
+ **Candidacy.** A group becomes a coded-repeat candidate only when all three hold:
164
+
165
+ - the same objective (`objectiveSha256`) was decided in at least 2 runs
166
+ (`askRecurrence`); records without an objective digest share one bucket;
167
+ - at least 0.75 of its decided records were accepted (`verifyPassRate`), where an
168
+ attempt superseded by a later retry of the same step in the same run never counts as a
169
+ pass;
170
+ - its step writes no repository (`writesRepository`, from the engine's `stepWrites` key).
171
+
172
+ A group that passes but misses reports `excludedReason`: `writes-repository`, or
173
+ `ask-not-repeated` when its runs asked different objectives. `weightedRecurrence`
174
+ weights each record by `2^(-age / improvement.telemetryHalfLifeDays)` (14 days by
175
+ default; an undated record weighs 1 and is counted in `undatedRecords`). It orders the
176
+ rows and never decides candidacy.
177
+
178
+ **Outputs.** Each candidate is a `kxm.candidate.v1` JSON file and a proposed diff under
179
+ `.kxm/candidates/` (or `--out-dir`); the report is written as
180
+ `kxm.improvement-report.v2` under `<workspace>/assets/improvements/`. `--dry-run` writes
181
+ neither. The candidate kind comes from a Git-reviewed name rule and only picks the
182
+ proposal template: a step or role that verifies, gates, tests, checks or lints proposes a
183
+ `.kxm/gates.yaml` command entry; a planning, review or repro step proposes a governed
184
+ skill, labelled consolidation because it is not a coded step; any other step proposes
185
+ replacing the agent step with a `kind: gate` step plus a `gates.yaml` command entry that
186
+ runs `scripts/<step>.mjs`, which the operator writes. Diffs are proposals with
187
+ placeholder hunks.
188
+
189
+ **Promotion readiness.** `promotion[]` reports, per candidate, `readyForReview` and a
190
+ reason under `improvement.promotionPolicy`:
191
+
192
+ - `manual_pr` (default): always ready; the operator reviews and applies the diff in a PR.
193
+ - `critic_quorum`: ready once two distinct critic receipts are cited. `kxm improve`
194
+ cites none, so every candidate reports not ready under this policy.
195
+ - `auto_threshold`: ready once the group's distinct runs reach
196
+ `improvement.autoThreshold.minRuns`, its accepted share reaches `minPassRate`, and its
197
+ mean recorded cost is at least `minCostSavings` over at least one cost sample. A group
198
+ with no recorded cost is never ready.
199
+
200
+ No policy authorizes anything. Every policy ends at an operator PR, activation is a
201
+ reviewed Git change for a future run, and telemetry cannot grant tools or skip a gate.
202
+
92
203
  ## Governance safeguards
93
204
 
94
205
  - Journal content is evidence, not executable policy.
206
+ - Improvement candidates are proposals. Promotion readiness never authorizes, and no `improvement.*` value activates a candidate.
95
207
  - An agent may propose a gate change but cannot silently weaken a required gate.
96
208
  - A peer-quorum reduction must be declared by policy and explicitly approved by an administrator for the current attempt; record it as a degraded outcome rather than normal success.
97
209
  - Contradictions stay open until evidence resolves them; synthesis must not erase minority risks.
@@ -7,16 +7,20 @@
7
7
  > (`kxm.prices.v1`) is implemented, dated, and hashed. `kxm routing report`
8
8
  > is implemented (`plugins/kxm/src/routing.ts`) and ranks routes quality-first,
9
9
  > then cost per accepted attempt, never ranking unknown cost cheapest and
10
- > reporting metered, unmetered, and unknown populations separately. Dev-helper
11
- > telemetry (`scripts/harness-run.mjs`) and the issue 127 assignment runner
10
+ > reporting metered, unmetered, and unknown populations separately. By default
11
+ > `kxm routing report` and `kxm improve` read two sources: the current project's
12
+ > Runtime event store (read-only) and then `.kxm/logs/telemetry.jsonl`; `--file`
13
+ > reads only the named file (see [Readers](#readers-kxm-routing-report-and-kxm-improve)).
14
+ > Dev-helper telemetry (`scripts/harness-run.mjs`) and the issue 127 assignment runner
12
15
  > (`scripts/assignment-run.mjs`, `just assign`) are implemented developer tools.
13
16
 
14
17
  This document describes what the tree does today versus what Tracking still
15
18
  plans. It does not invent prices or close product enums.
16
19
 
17
- ## Implemented: v1 record
20
+ ## Implemented: v1 record (parse-only)
18
21
 
19
- Schema id: `kxm.routing-record.v1` (`plugins/kxm/src/routing.ts`).
22
+ Schema id: `kxm.routing-record.v1` (`plugins/kxm/src/routing.ts`). v1 is parsed, never
23
+ written by the product: the Runtime writes v2.
20
24
 
21
25
  Always present or defaulted by `parseRoutingRecord`: `schema`,
22
26
  `behavioralHashVersion`, `behavioralSha256`, `skills` (default `[]`),
@@ -36,17 +40,23 @@ retries, outcomes) is outside the hash.
36
40
 
37
41
  **No built-in product producer.** The worker envelope validates a `routing`
38
42
  field when present. Nothing in `plugins/kxm/src` or `scripts/kxm-worker.mjs`
39
- writes a record. Repo records are test-built. External JSONL can be ingested.
43
+ writes a v1 record. The developer assignment runner (`scripts/assignment-run.mjs`)
44
+ writes v1 records with `finalOutcome: "pending"` and a per-assignment
45
+ `rolePromptSha256`, so the improvement report counts them as undecided and never
46
+ groups two assignments as one ask. Other repo records are test-built. External
47
+ JSONL can be ingested.
40
48
 
41
49
  v1 has **no dedicated fields** for harness, provider, latency, cost basis,
42
50
  cache-write tokens, or context occupancy. Bounded `providerMetadata` may
43
51
  carry extra keys (at most 32; values are strings, numbers, or booleans;
44
- `prompt`/`body`/`content`/`message` keys are rejected), but those keys are
45
- **not standardized** and `kxm routing report` does not read them.
52
+ a key containing `prompt`, `body`, `content` or `message`, in any case, is
53
+ rejected), but those keys are **not standardized**. The ranked report reads only
54
+ `providerMetadata.harness` from a v1 record, to label its harness column.
46
55
 
47
- `kxm routing report` reads `telemetry.jsonl`, groups by behavioral hash, sorts
48
- by run count (then hash), and sums missing `costUsd` as **zero**. That silent
49
- underquote is why the report is **not** a ranking source.
56
+ The `configurations` block of `kxm routing report` groups v1 records by
57
+ behavioral hash, sorts by run count (then hash), and sums missing `costUsd` as
58
+ **zero**. That silent underquote is why the block is **not** a ranking source;
59
+ the ranked table described under [report and price catalog](#implemented-report-and-price-catalog) is.
50
60
 
51
61
  ## Implemented: dev helper telemetry
52
62
 
@@ -166,6 +176,51 @@ Fields carried on `RoutingRecordV2`:
166
176
 
167
177
  The KXM engine settle transaction appends a `routing.attempt.recorded` event carrying the v2 record and refuses to settle without a valid `costBasis`. Attempt dispatch enforces `limits.maxModelCost` against metered cost before invocation (`budget_model_cost`).
168
178
 
179
+ What the engine writes on every settled attempt (`producerRoutingRecord` and the
180
+ failure path in `settleMember`, `plugins/kxm/src/engine.ts`):
181
+
182
+ - **Engine-reserved `providerMetadata` keys.** Four keys are written after the
183
+ producer's keys, so a producer key with the same name is dropped and cannot spoof
184
+ them. A producer keeps at most 28 keys of its own, so the record stays within the
185
+ 32-field limit.
186
+
187
+ | Key | Value |
188
+ |---|---|
189
+ | `workflowId` | The compiled workflow's id |
190
+ | `askSha256` | `kxmStepAskSha256`: a digest of the workflow id, step id, step kind, agent id, instructions, declared outcomes and required evidence keys. It is equal across runs for one step and agent, and excludes the run, assignment and attempt ids, the run objective, repositories, model and context packet |
191
+ | `objectiveSha256` | The run's accepted prompt digest (`sha256:<hex>`); the prompt text is never read here |
192
+ | `stepWrites` | `true` when the step has any repository with `write` access |
193
+
194
+ The key names avoid the parser's refusal (`/prompt|body|content|message/i`), which
195
+ is why the ask digest is not called a prompt hash.
196
+ - **`agentRole`** is the producer's value when it supplies one, otherwise the
197
+ dispatched agent id.
198
+ - **Record-time `finalOutcome`** is only ever `blocked` (the declared outcome takes a
199
+ back edge) or `failed` (a producer error, an outcome the step does not declare, or a
200
+ terminal that is not `completed`). A forward edge or a completed terminal leaves it
201
+ unset: acceptance is not known when the attempt settles. The engine never writes
202
+ `accepted` or `pending`.
203
+ - `retries` is the step attempt minus one. Every record from one step attempt shares
204
+ it, including panel members and any assignment retried inside that step attempt, so
205
+ the improvement report's rule that a later retry supersedes an earlier attempt never
206
+ separates them.
207
+
208
+ **Read-time resolution.** Readers of the event store resolve each Runtime attempt's
209
+ outcome in memory and never write it back (`readEngineRoutingRecords`,
210
+ `plugins/kxm/src/improve-sources.ts`). The first rule that applies wins:
211
+
212
+ 1. a record-time `blocked` or `failed` stands;
213
+ 2. a later `step.entered` for the same step makes it `reworked`;
214
+ 3. a `completed` run makes it `accepted`;
215
+ 4. a `failed` run makes it `failed`;
216
+ 5. anything else (a cancelled or still-running run) is undecided and has no
217
+ `finalOutcome`.
218
+
219
+ `reworked` is a read-time value only; it is outside the stored v2 vocabulary.
220
+ Attempts whose `harness` is `driver-simulated` are dropped and counted as
221
+ `excludedSimulated`. Records written before the engine carried these keys are not
222
+ backfilled: they still resolve an outcome, but they group per run.
223
+
169
224
  ## Implemented: report and price catalog
170
225
 
171
226
  - **Price catalog:** `.kxm/prices.yaml` (`kxm.prices.v1`, dated and hashed) defines input, output, cache-read, cache-write rates, and context tiers for active models. Missing rows or uncataloged models evaluate to `costBasis: "unknown"`.
@@ -174,7 +229,36 @@ The KXM engine settle transaction appends a `routing.attempt.recorded` event car
174
229
  - **Underquote prevention:** Routes with unknown cost are flagged (`*`) and **never ranked cheapest**, eliminating silent underquoting.
175
230
  - **Population separation:** Reports metered cost, unmetered attempt counts, unknown-cost attempt counts, and quota-exhausted attempt counts as separate metrics rather than a single misleading total.
176
231
  - **List prices flag:** Supports `--equivalent-list-cost` / `--list-prices` to display estimated list rates for comparison alongside actual recorded spend.
177
- - **Post-MVP:** Dynamic catalog price feeds (`kxm update --models`), budget roll-over, and automated promotion of repeat successes into workflow gates.
232
+ - **Rework column:** reads `transitions`, which Runtime records never set, so Runtime rework shows up only as a resolved `reworked` outcome, which the report does not count as a pass.
233
+ - **Post-MVP:** Dynamic catalog price feeds (`kxm update --models`) and budget roll-over. Coded-repeat candidates from `kxm improve` only propose; activation is a reviewed Git change.
234
+
235
+ ## Readers: `kxm routing report` and `kxm improve`
236
+
237
+ Both commands load routing records the same way (`loadRoutingSources`,
238
+ `plugins/kxm/src/improve-sources.ts`):
239
+
240
+ 1. With `--file <path>`, only that JSONL file is read. It may hold bare v1 or v2
241
+ records, records nested under `routing` or `envelope.routing`, or
242
+ `routing.attempt.recorded` events.
243
+ 2. Otherwise, when the current directory is inside a KXM project, the project's
244
+ Runtime event store is read first:
245
+ `<state root>/runtime/projects/<key>/run-events.db`, where the key is derived from
246
+ the checkout's real path exactly as the Runtime derives it. It is opened read-only
247
+ for one `SELECT` over the `events` table (routing, `step.entered` and
248
+ `run.status_changed` events only); the runs table, run plans and the prompt sidecar
249
+ are never read, and nothing is created, written or migrated. Outcomes are resolved
250
+ as described above.
251
+ 3. Then `telemetry.jsonl` in the workspace logs directory.
252
+
253
+ A v2 record whose `attemptId` an earlier source already supplied is dropped and
254
+ counted as `duplicatesDropped` on the later source. Every source is reported in a
255
+ `sources` array with `kind` (`engine`, `telemetry` or `file`), `path`, `exists`,
256
+ `records` and, for the store, `skippedInvalid`, `excludedSimulated` and `undecided`.
257
+ `kxm improve` prints the sources in text and JSON and adds `projectRoot`;
258
+ `kxm routing report` adds `sources` to its JSON only, and keeps `file` set to the
259
+ telemetry path. A store that exists but cannot be read stops either command with exit
260
+ 1 and `improve_source_unreadable`, naming the path. Each checkout reads only its own
261
+ store: there is no cross-worktree aggregation.
178
262
 
179
263
  ## Precedence
180
264