@kontextmind/kxm 0.7.94 → 0.7.96

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (140) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.kxm/README.md +39 -9
  3. package/CHANGELOG.md +1 -1
  4. package/README.md +147 -257
  5. package/SECURITY.md +21 -12
  6. package/docs/README.md +133 -54
  7. package/docs/adr/ADR-0002-browser-automation-steel-doks.md +24 -18
  8. package/docs/adr/ADR-0003-sqlite-only-store.md +100 -0
  9. package/docs/adr/ADR-0004-edge-identity-authentik.md +99 -0
  10. package/docs/adr/README.md +33 -0
  11. package/docs/concepts/architecture.md +262 -0
  12. package/docs/concepts/data-and-storage.md +194 -0
  13. package/docs/concepts/trust-model.md +152 -0
  14. package/docs/contracts/README.md +22 -14
  15. package/docs/contracts/effects-and-recovery.md +3 -0
  16. package/docs/contracts/migration.md +2 -2
  17. package/docs/contracts/routing.md +6 -5
  18. package/docs/contributing/assignment-runner.md +388 -0
  19. package/docs/contributing/ci-and-release.md +231 -0
  20. package/docs/contributing/development.md +362 -0
  21. package/docs/contributing/harness-routing-internals.md +192 -0
  22. package/docs/{packages.md → contributing/packages.md} +13 -15
  23. package/docs/{skills → contributing}/repo-work-delivery.md +20 -21
  24. package/docs/contributing/test-matrix.md +208 -0
  25. package/docs/{tui-components.md → contributing/tui-components.md} +30 -22
  26. package/docs/contributing/writing-docs.md +340 -0
  27. package/docs/glossary.md +471 -0
  28. package/docs/guides/agent-skills.md +137 -0
  29. package/docs/guides/browser-automation.md +160 -0
  30. package/docs/guides/context-and-memory.md +352 -0
  31. package/docs/guides/continuous-improvement.md +228 -0
  32. package/docs/guides/governed-skills.md +173 -0
  33. package/docs/guides/nous-providers.md +186 -0
  34. package/docs/guides/peer-messaging.md +304 -0
  35. package/docs/guides/pi-workers.md +219 -0
  36. package/docs/guides/provenance-gates.md +313 -0
  37. package/docs/guides/webhook-workflows.md +364 -0
  38. package/docs/kb/how-credentials-retrieved-safely.md +38 -12
  39. package/docs/kb/how-to-capture-and-annotate-section.md +15 -13
  40. package/docs/kb/how-to-connect-playwright-to-steel.md +16 -11
  41. package/docs/kb/how-to-recover-expired-session-or-orphan.md +26 -16
  42. package/docs/kb/how-to-resume-after-mfa.md +19 -11
  43. package/docs/kb/how-to-take-over-session.md +17 -13
  44. package/docs/kb/why-authentication-disappeared.md +22 -14
  45. package/docs/kb/why-automation-opened-different-browser.md +23 -14
  46. package/docs/kb/why-session-viewer-cannot-control.md +13 -12
  47. package/docs/operations/backup-and-restore.md +248 -0
  48. package/docs/operations/deploy.md +307 -0
  49. package/docs/operations/monitoring.md +209 -0
  50. package/docs/operations/runtime-sync.md +192 -0
  51. package/docs/operations/troubleshooting.md +265 -0
  52. package/docs/operations/upgrade.md +124 -0
  53. package/docs/prompts/browser-annotate-feedback.md +7 -7
  54. package/docs/prompts/browser-diagnose-recover.md +11 -10
  55. package/docs/prompts/browser-explore.md +7 -7
  56. package/docs/prompts/browser-repro-fix.md +7 -7
  57. package/docs/prompts/browser-start.md +12 -11
  58. package/docs/prompts/browser-takeover.md +8 -8
  59. package/docs/{cli-reference.md → reference/cli-reference.md} +83 -41
  60. package/docs/{config-reference.md → reference/config-reference.md} +159 -148
  61. package/docs/reference/configuration.md +299 -0
  62. package/docs/reference/harness-routing.md +508 -0
  63. package/docs/reference/http-api.md +203 -0
  64. package/docs/reference/tools.md +370 -0
  65. package/docs/{workflow-guide.md → reference/workflow-catalog.md} +92 -153
  66. package/docs/reference/workflow-definitions.md +286 -0
  67. package/docs/start/first-workflow.md +287 -0
  68. package/docs/start/install.md +146 -0
  69. package/docs/start/quickstart-claude-code.md +405 -0
  70. package/docs/start/quickstart-pi.md +213 -0
  71. package/docs/templates/README.md +78 -73
  72. package/docs/templates/adr.md +13 -13
  73. package/docs/templates/architecture.md +55 -71
  74. package/docs/templates/bug-fix.md +13 -16
  75. package/docs/templates/feature.md +14 -19
  76. package/docs/templates/handoff.md +44 -46
  77. package/docs/templates/postmortem.md +30 -43
  78. package/docs/templates/research.md +15 -20
  79. package/docs/templates/review.md +49 -50
  80. package/docs/templates/runbook.md +38 -30
  81. package/docs/templates/test-plan.md +16 -23
  82. package/docs/templates/test-report.md +14 -17
  83. package/examples/README.md +9 -5
  84. package/examples/provenance-workflow.json +1 -1
  85. package/examples/webhook-workflows/jira-development.json +59 -0
  86. package/examples/webhook-workflows/jira-issue-updated.json +12 -0
  87. package/package.json +2 -2
  88. package/packages/core/tui/README.md +1 -1
  89. package/plugins/kxm/.claude-plugin/plugin.json +1 -1
  90. package/plugins/kxm/README.md +31 -32
  91. package/plugins/kxm/dist/cli.js +5 -5
  92. package/plugins/kxm/dist/mcp-server.js +1 -1
  93. package/plugins/kxm/dist/runtime.js +1 -1
  94. package/plugins/kxm/package.json +1 -1
  95. package/plugins/kxm/skills/kxm/references/protocol.md +3 -1
  96. package/plugins/kxm/skills/kxm-browser-auth/SKILL.md +1 -1
  97. package/plugins/kxm/skills/kxm-browser-diagnostics/SKILL.md +5 -5
  98. package/plugins/kxm/skills/kxm-browser-explore/SKILL.md +2 -2
  99. package/plugins/kxm/skills/kxm-browser-session/SKILL.md +10 -13
  100. package/plugins/kxm/skills/kxm-browser-takeover/SKILL.md +1 -1
  101. package/plugins/kxm/skills/kxm-browser-verify/SKILL.md +1 -1
  102. package/plugins/kxm/skills/kxm-context-memory/SKILL.md +13 -4
  103. package/plugins/kxm/skills/kxm-hub-ops/SKILL.md +3 -1
  104. package/plugins/kxm/skills/kxm-mind-setup/SKILL.md +2 -1
  105. package/plugins/kxm/skills/kxm-project-setup/SKILL.md +31 -54
  106. package/plugins/kxm/skills/kxm-projects/SKILL.md +1 -1
  107. package/plugins/kxm/skills/kxm-protocol/SKILL.md +1 -1
  108. package/plugins/kxm/skills/kxm-routing-improve/SKILL.md +15 -7
  109. package/plugins/kxm/skills/kxm-runs/SKILL.md +11 -5
  110. package/plugins/kxm/skills/kxm-session/SKILL.md +1 -1
  111. package/plugins/kxm/skills/kxm-tasks/SKILL.md +9 -7
  112. package/plugins/kxm/skills/kxm-workflow/SKILL.md +10 -2
  113. package/plugins/kxm/src/cli/system.ts +1 -1
  114. package/plugins/kxm/src/cli.ts +3 -3
  115. package/plugins/kxm/src/init-guide-setup.ts +1 -1
  116. package/plugins/kxm/src/mcp-server.ts +1 -1
  117. package/plugins/kxm/src/modes.ts +1 -1
  118. package/schemas/README.md +1 -1
  119. package/docs/agent-communication-envelopes-and-gates.md +0 -553
  120. package/docs/agent-skills.md +0 -198
  121. package/docs/architecture.md +0 -245
  122. package/docs/assignment-runner.md +0 -264
  123. package/docs/browser-automation.md +0 -139
  124. package/docs/configuration.md +0 -437
  125. package/docs/continuous-improvement.md +0 -226
  126. package/docs/getting-started.md +0 -277
  127. package/docs/harness-routing.md +0 -616
  128. package/docs/kb/qa-authentik-authentication.md +0 -97
  129. package/docs/kb/qa-extension-install-and-hub-bootstrap.md +0 -85
  130. package/docs/kb/qa-hub-on-a-public-host.md +0 -48
  131. package/docs/kb/qa-sqlite-vs-duckdb.md +0 -35
  132. package/docs/kb/qa-what-the-hub-stores.md +0 -64
  133. package/docs/kxm-handbook.md +0 -1181
  134. package/docs/operations.md +0 -510
  135. package/docs/operator-pi-packages.md +0 -67
  136. package/docs/provenance-gates.md +0 -295
  137. package/docs/skills.md +0 -47
  138. package/docs/test-matrix.md +0 -132
  139. package/docs/troubleshooting.md +0 -293
  140. package/docs/webhook-workflows.md +0 -240
@@ -1,293 +0,0 @@
1
- # Troubleshooting
2
-
3
- Start with the smallest boundary: hub health, authentication, registration, peer discovery, then message delivery.
4
-
5
- ## Quick diagnostic sequence
6
-
7
- 1. Confirm the hub terminal still shows `kxm hub listening`.
8
- 2. Request `/health`, then `/ready` to confirm storage access.
9
- 3. Compare the hub URL, token, and project on both agents.
10
- 4. Confirm every agent has a unique name.
11
- 5. Run `/kxm hub` in Pi or call `kxm_list` in Claude.
12
- 6. Inspect hub logs for registration, stale-agent, or server-error events.
13
- 7. If a workflow tool returns `workflow_forbidden`, read `operation`, `assignedCoordinatorName`, and `nextAction`. Do not retry as a peer.
14
-
15
- ## Common problems
16
-
17
- ### A continued Pi session rejects every turn
18
-
19
- If a worker was stopped during `kxm_await`, `--continue` may leave a `tool_use` without `tool_result`. The worker retries once without `--continue` and writes a project-and-agent identity-keyed recovery envelope under `.kxm/state`. Do not paste agent logs into the journal. Keep the same project and agent name so the hub identity and recovery key resume.
20
-
21
- ### A model quota or provider error settles the agent
22
-
23
- KXM waits until Pi has exhausted its own automatic retries. It then keeps the inbound message in `delivered` state, records an allowlisted `quota` or `provider_error` diagnostic without the provider body, and restarts the RPC child. Configure `KXM_WORKER_FALLBACK_MODELS` (or `--fallback-models`) to rotate immediately; otherwise the worker retries after `KXM_WORKER_PROVIDER_RETRY_MS`. Keep continuation enabled so finished peer calls and tool results survive the model switch. Use `--fresh-start`, not `--no-continue`, when only the first launch must avoid old session state.
24
-
25
- ### A worker heartbeat is healthy but one tool never finishes
26
-
27
- Set `KXM_WORKER_TOOL_TIMEOUT_MS` above the longest legitimate tool call. Its 31-minute default intentionally gives a 30-minute `kxm_fanout` wait time to return durable pending handles before supervision intervenes. When that bound is exceeded, the structured worker log records `worker_tool_timeout` with only the allowlisted tool name and diagnostic class, the delivered hub request stays recoverable, and the RPC process is restarted. If the stuck worker was supposed to be read-only, also set `KXM_WORKER_TOOLS=read,grep,find,ls`; prompt wording alone does not remove shell or write capabilities.
28
-
29
- ### A hub or worker PID claim is stale
30
-
31
- Version 0.4.3 prevents a second wrapper from replacing a live hub or worker claim. `kxm hub stop` ignores an invalid, non-running, or ownership-mismatched record rather than guessing. A hub claim whose wrapper PID is dead is reclaimed automatically on the next `kxm hub start`; the wrapper also terminates an orphaned hub server child recorded by a dead wrapper (for example after `SIGKILL`) before reclaiming, and `kxm hub stop` can stop such an orphan directly. If a pre-0.4.3 process left a malformed claim behind, inspect the exact `.pid` JSON and verify that its recorded PID is no longer running; for a hub, also verify the configured port has no listener. Then remove only that exact `.pid` and its recorded `.stop` control file before relaunching once. Worker filenames include a project/agent identity digest and their records include the exact names and generation, so do not substitute a similarly sanitized filename. Never delete the `.kxm/state` directory or SQLite database to clear a claim.
32
-
33
- ### GitHub checks passed but the workflow is still waiting
34
-
35
- The hub does not poll GitHub. Run `kxm gate github watch` with the same `runId`, `stageId`, and `signalKey`. A watcher timeout posts the exact signed `failed` signal, retains bounded check evidence, and exits `4`; it never invents `passed`.
36
-
37
- ### CI jobs stay queued and never start
38
-
39
- Every CI job in `.github/workflows/ci.yml` runs on the `kontextmind-doks` label. If all jobs sit in `queued` with an online, idle runner, the runner lost that custom label (for example after re-registration — the default labels are only `self-hosted`, `Linux`/`Windows`, `X64`). Confirm with `gh api repos/kontextmind/kxm/actions/runners --jq '.runners[] | {name, labels: [.labels[].name]}'`; jobs cannot match on the default `doks` label alone. Re-add the label with `gh api repos/kontextmind/kxm/actions/runners/<id>/labels -X POST --input - <<< '{"labels":["kontextmind-doks"]}'` and jobs are picked up on the next evaluation; if not, push an empty commit to retrigger the run. The Windows runner (`kxm-win-local`) is paused by policy and must not be re-labeled to satisfy Linux jobs.
40
-
41
- ### The hub refuses to start
42
-
43
- **`KXM_PORT must be an integer between 0 and 65535`**
44
-
45
- Set `KXM_PORT` to a valid integer. Remove the variable to use `7331`.
46
-
47
- **`KXM_AUTH_TOKEN is required when binding beyond localhost`**
48
-
49
- Either restore `KXM_HOST=127.0.0.1` or configure a token before using a non-loopback interface.
50
-
51
- **`KXM hub env file is malformed`**
52
-
53
- The persisted credential file (`hub-env.json` under the user state root) failed
54
- validation. It holds only `KXM_AUTH_TOKEN` / `KXM_PROJECT_TOKENS` values in
55
- `kxm.hub-env.v1` schema; fix its JSON or delete it to have kxm generate a
56
- fresh admin token on the next start. To rotate the generated token, delete
57
- the file and run `kxm hub start` again.
58
-
59
- #### Database schema is newer than this runtime supports
60
-
61
- Do not delete or rewrite the database. Start the package version that created it, or upgrade this runtime. Restore the pre-upgrade backup when rolling back.
62
-
63
- #### Address already in use
64
-
65
- Another process owns the port. Stop that process or choose another port, then update every agent's `KXM_SERVER_URL`.
66
-
67
- ### `kxm harness list` says Claude Code is `not_detected` on Windows
68
-
69
- npm installs Claude Code as `claude.cmd` (and an extensionless shim), not
70
- `claude.exe` on `PATH`. The native binary lives next to the shim at
71
- `%AppData%\Roaming\npm\node_modules\@anthropic-ai\claude-code\bin\claude.exe`.
72
- Older probes spawned `claude` without a shell, got `ENOENT` or `EINVAL`, and
73
- reported the CLI missing even when `claude --version` worked in cmd or Git Bash.
74
-
75
- Current `kxm harness list` retries `claude.exe`, then that inner package
76
- `.exe`, then `claude.cmd` on win32. It sets `issues: ["windows_shim"]` only
77
- when the npm shim is what answered. If the entry is still `not_detected`,
78
- confirm `%AppData%\Roaming\npm` is on `PATH` for the same process that runs
79
- `kxm`, then `claude --version` and `claude auth status`. Assignment dispatch
80
- (`just assign` / `harness-run`) follows the inner `claude.exe` (and Pi's
81
- `node.exe` plus `cli.js`) with `shell: false`; unverified `.cmd` launchers
82
- are still refused.
83
-
84
- The same npm-shim miss can appear for `pi` on Windows.
85
-
86
- ### Pi shows `hub:off`
87
-
88
- - Confirm the hub is reachable from the Pi terminal.
89
- - Verify `KXM_AUTH_TOKEN` exactly matches the hub token.
90
- - Check whether a live agent already uses the same name in the same project.
91
- - Restart Pi after changing environment variables.
92
- - For an exact development load, use `pi --no-extensions -e ./plugins/kxm/src/extension.ts`. Add every required provider extension with another `-e`; otherwise Pi discovery is intentionally disabled.
93
- - For long-lived workers, set the reviewed `KXM_WORKER_EXTENSION_PATHS` and `KXM_WORKER_SKILL_PATHS` described in [Configuration](configuration.md#long-lived-worker-settings). Invalid paths fail before supervision instead of entering a restart loop.
94
-
95
- ### Pi update fails looking for `refs/heads/master`
96
-
97
- The KXM default branch is `main`. An older Pi git checkout still tracking
98
- `master` fails with `couldn't find remote ref refs/heads/master`. Remove the
99
- package and reinstall with an explicit ref:
100
-
101
- ```text
102
- pi remove git:github.com/kontextmind/kxm
103
- pi install git:github.com/kontextmind/kxm@main
104
- ```
105
-
106
- ### `kxm --help` prints a former flat command list
107
-
108
- If the installed `kxm --help` prints `validate | status | hub | worker | stop | …`
109
- instead of the current Commander groups, the committed `plugins/kxm/dist/cli.js`
110
- is stale. Run `npm run build` and commit the generated `dist` so the operator
111
- CLI matches source.
112
-
113
- ### `kxm` is not recognized
114
-
115
- `pi install git:github.com/kontextmind/kxm@main` installs the Pi extension
116
- and Agent Skill, not a global operator command. Install the versioned `.tgz`
117
- release asset through the authenticated `gh release download` flow in
118
- [Getting started](getting-started.md#1-install), or run
119
- `node scripts/kxm.mjs` from a clone after `npm ci`. `npx kxm` and a
120
- global `git+https` npm install are not supported installation paths.
121
-
122
- For bash and zsh, `kxm completion install` can add the kxm bin directory to
123
- `PATH` in the shell rc file when it is missing; restart the shell afterwards.
124
-
125
- ### Tab completion is not active
126
-
127
- Run `kxm completion install` for the detected shell, or pass
128
- `--shell bash|zsh|fish` explicitly. The install appends one guarded stanza to
129
- the shell rc file and is idempotent: rerunning never duplicates it. Fish needs
130
- no rc entry because fish auto-loads `~/.config/fish/completions`. After
131
- installing, start a new terminal or `source` the rc file. To inspect without
132
- writing, use `--dry-run`; to suppress the post-`kxm init` offer, set
133
- `KXM_SKIP_COMPLETION_PROMPT=1`.
134
-
135
- ### An expected peer is missing
136
-
137
- The two agents usually have different `KXM_PROJECT` values or one stopped sending heartbeats. Compare settings and check for an `agent_stale` event. Names and projects are case-sensitive for display; live-name uniqueness is case-insensitive.
138
-
139
- ### A request stays `queued`
140
-
141
- The recipient registered but has no active SSE stream. Confirm its process is running and connected. Proxies must disable response buffering for `/v1/events` and allow long-lived connections.
142
-
143
- ### A request stays `delivered`
144
-
145
- The recipient acknowledged it but has not replied. It may still be working, waiting for approval, or blocked. Avoid sending the same request repeatedly. Check the recipient session directly if the wait is unexpected.
146
-
147
- If the work is obsolete, the sender can call `kxm_cancel`. This changes hub state only; it cannot reverse file changes or external effects already performed by the peer.
148
-
149
- ### `kxm_await` times out
150
-
151
- `kxm_await` waits at most 60 seconds; that is both its default and its maximum. A timeout does not end the request. Use `kxm_get` to inspect the state, or `kxm_workflow_wait` for long external work. `cancelled`, `expired`, and `error` are terminal outcomes. Resend only when the task is safe to repeat, and use an idempotency key when retrying after an uncertain network result.
152
-
153
- ### A message disappears after completion
154
-
155
- Terminal records are removed after seven days by default. Increase `KXM_MESSAGE_RETENTION_MS` if operators need a longer diagnostic window. Durable artifacts should live in Git or another system of record.
156
-
157
- ### Claude tools do not appear
158
-
159
- 1. Confirm the marketplace and plugin are installed.
160
- 2. Run `/reload-plugins` or restart Claude Code.
161
- 3. Inspect `/mcp` and verify the `kxm` server connected.
162
- 4. Confirm Node.js 22.19 or newer on the 22.x line, or Node.js 24 or newer, is on the `PATH` used by Claude Code.
163
- 5. Reinstall or update the marketplace if the cached plugin predates the `dist/mcp-server.js` bundle.
164
-
165
- ### Claude does not receive pushed requests
166
-
167
- Ordinary MCP tools and channel delivery are separate. During the research preview, start the community channel explicitly:
168
-
169
- ```text
170
- claude --dangerously-load-development-channels plugin:kxm@kxm
171
- ```
172
-
173
- Accept the trust prompt and check the channel startup notice. Organization policy can still block channels. If pushed delivery remains unavailable, use `kxm_inbox` and `kxm_reply`.
174
-
175
- ### Jira webhook is rejected
176
-
177
- - HTTP 401 means the SHA-256 signature is missing, uses another algorithm, or does not match the raw UTF-8 body. Confirm Jira and `secretEnv` resolve the same secret.
178
- - HTTP 400 usually means the delivery identifier or JSON body is missing.
179
- - HTTP 409 means the configured coordinator has never registered. Start it once with the matching project and name; Jira retries 409 responses.
180
- - HTTP 204 means the event or JSON-path filter did not match, so no workflow was intended.
181
- - HTTP 200 with `duplicate: true` means a provider retry was safely deduplicated.
182
-
183
- ### Long-lived worker keeps restarting
184
-
185
- Inspect the structured `worker_process_error` and `worker_exited` events. Confirm Pi is installed on the service account's `PATH`, the working directory exists, model credentials are available, the package is enabled, and non-interactive project trust was configured intentionally. Set `KXM_PI_COMMAND` to an explicit executable path when service-manager environments have a reduced `PATH`.
186
-
187
- ### A workflow message stays queued while the worker restarts once
188
-
189
- This is normally the safe session-routing handshake. With `--session-isolation workflow`, a message for a different run is deliberately not acknowledged in the current Pi context. Look for `worker_session_routed`; the old child must close before one replacement starts with the run-specific `--session-dir`, after which the same message ID replays and advances to `delivered`.
190
-
191
- If it repeats, inspect `worker_session_request_rejected` and verify:
192
-
193
- - the worker was started through `kxm agent worker` with a valid state directory;
194
- - `KXM_WORKER_SESSION_SCOPE` was not manually set (the supervisor owns it);
195
- - the state directory is writable by only the service account;
196
- - the hub and worker are from the same release; and
197
- - the message has a canonical hub-owned `workflowRunId`, not only a correlation ID.
198
-
199
- Do not manually acknowledge the message, edit the route request, copy a run JSONL into `default`, or launch a second worker with the same identity. Those actions defeat context isolation.
200
-
201
- ### `worker_session_state_recovered` appears
202
-
203
- The binding manifest did not match its bounded schema or exact worker owner. The supervisor renamed it to `worker-session-binding-<workerKey>.json.corrupt-<timestamp>` and started the stable default binding rather than guessing a workflow. Read `kxm_workflow_get` for unfinished stages and inspect queued/delivered message IDs. Preserve the quarantined manifest for diagnosis, then re-drive unfinished work from the hub. Repeated corruption suggests disk, antivirus, concurrent-service, or permission problems; confirm only one supervisor owns the exact project/agent PID claim.
204
-
205
- ### A workflow seems to remember another run
206
-
207
- Confirm the worker log says `"sessionIsolation":"workflow"` and the Pi child has a `runs/<exact-runId>` session directory. Isolation is opt-in for upgrade compatibility, and both the CLI and raw supervisor default to `off`. Restart cleanly with `kxm agent worker ... --session-isolation workflow`. The first isolated start intentionally uses fresh scoped storage because KXM cannot safely infer which session in the former shared Pi directory belonged to this worker. Existing content created in a formerly shared Pi session cannot be automatically separated retroactively; treat authoritative workflow journal/assets as the recovery source and start a fresh run-specific history.
208
-
209
- ### Fanout returns pending before a model replies
210
-
211
- `kxm_fanout.timeoutMs` is a local wait, not the message lifetime. A pending result includes the durable `messageId`, current message status, expiry, and whether the wait timed out or was aborted. Use `kxm_get` to inspect that ID, or repeat the exact fanout with the same correlation ID, idempotency prefix, targets, and content. Do not send a replacement with a new prefix while the original remains pending. Normally omit `ttlMs` for model work so time spent queued behind another request does not prematurely expire it. A pending peer has not contributed review or planning evidence and must not be counted toward a workflow checkpoint.
212
-
213
- ### Workflow cannot advance
214
-
215
- Call `kxm_workflow_get` and use only `currentStage`. A passing checkpoint needs
216
- a keyed, non-empty value for every declared `requiredEvidence` identity; extra
217
- or unrelated keys do not count. Warnings and failures remain active until
218
- corrected, and their evidence is journaled but does not satisfy a later passing
219
- attempt. If attempts are exhausted or the coordinator settles early, the run
220
- becomes failed and its journal records the reason; start a new provider delivery
221
- only after deciding whether repeating external effects is safe.
222
-
223
- For a requirement with `kind: peer-reply`, inspect
224
- `resolvedEvidencePolicies`, `verifiedEvidence`, and the current attempt. An
225
- ordinary evidence string cannot satisfy it. Every eligible agent must have
226
- registered in the workflow project before the run starts, and a passing
227
- checkpoint must cite durable replied message IDs in `evidenceRefs` before those
228
- source messages reach terminal retention.
229
-
230
- Common provenance failures are:
231
-
232
- - `workflow_context_forbidden`: the sender is not the run's assigned coordinator;
233
- - `workflow_context_inactive`: the run or stage is not currently running;
234
- - `workflow_context_attempt_mismatch`: use `stage.attempts + 1` and send fresh work after a retry;
235
- - `workflow_evidence_producer_forbidden`: the target is not in the run's snapshotted eligible set;
236
- - `workflow_evidence_policy_missing` or `workflow_evidence_policy_unresolved`: the requirement has no usable resolved peer policy;
237
- - `workflow_provenance_invalid`: a cited message is missing, pending, ineligible, wrong-direction, or bound to another project, run, stage, requirement, or attempt;
238
- - `workflow_evidence_incomplete`: there are fewer unique verified producers than the effective minimum.
239
-
240
- Multiple replied messages from one peer count once. Correlation IDs and
241
- idempotency prefixes are retry controls, not provenance. Do not replace a
242
- rejected reference with an unscoped send.
243
-
244
- If policy declares a lower `degradation.minProducers`, an operator can inspect
245
- and approve it with `kxm gate --dry-run --json degrade ...` followed by
246
- the same command without `--dry-run`, using the administrative token. Approval
247
- must target the current stage and attempt and does not advance the workflow;
248
- the coordinator must still checkpoint with enough verified references. A
249
- callback, project token, or peer cannot approve degradation.
250
-
251
- ### `kxm gate degrade` returns HTTP 503 `admin_auth_not_configured`
252
-
253
- The hub started without a non-empty `KXM_AUTH_TOKEN`, so no administrative
254
- credential exists for the degradation route. Project tokens deliberately cannot
255
- substitute for it, even when the operator holds every project credential. The
256
- route fails closed and does not create an approval.
257
-
258
- Stop the hub gracefully, set a new high-entropy `KXM_AUTH_TOKEN` in the hub
259
- service, retain the explicit `KXM_PROJECT_TOKENS` mapping for workers, and
260
- restart against the same `.kxm/state/kxm.db`. Give the administrative token
261
- only to the operator terminal, never to agents or callbacks. Read the run again
262
- because the current attempt may have changed, run the exact degradation command
263
- with `--dry-run --json`, and then approve the current stage, requirement, and
264
- attempt without `--dry-run`. A restart does not make an earlier-attempt approval
265
- valid for the new attempt.
266
-
267
- ### External workflow callback is rejected or does not resume
268
-
269
- - HTTP 401 means the callback signature does not match the exact raw body. Use `signalSecretEnv` when configured; the workflow-start secret will not work in that case.
270
- - HTTP 404 means the workflow definition or run ID does not match this hub.
271
- - HTTP 409 with `workflow_not_waiting` means the coordinator did not successfully call `kxm_workflow_wait`, the deadline already failed the run, or a prior signal advanced it.
272
- - HTTP 409 with `workflow_signal_mismatch` means the URL's signal key differs from the active wait. Read the run and use its exact `waiting.signalKey`.
273
- - HTTP 409 with `workflow_signal_context_mismatch` means a supplied `workflow.run`, `workflow.stage`, or `workflow.signal` evidence value disagrees with the route or active wait. Correct it or omit optional context evidence.
274
- - HTTP 400 with `workflow_evidence_incomplete` means a passing callback omitted one or more named requirements. Read `missingRequirements`; extra checks and context fields cannot substitute for them.
275
- - HTTP 400 with `invalid_workflow_evidence` means evidence was not a keyed string object or contained duplicate keys after case/whitespace normalization.
276
- - HTTP 200 with `duplicate: true` is expected after retrying the same provider delivery ID. Do not generate a new ID for the same callback attempt.
277
- - A failed or timed-out callback consumes that wait attempt. Re-enter the wait and start a new `github watch` or `signal` command so its default delivery generation is new; reserve an explicit `--delivery-id` for retries of one unchanged callback body.
278
-
279
- Inspect `workflow_wait_started`, `workflow_signal_received`, and `workflow_wait_timed_out` logs without copying secrets or full callback bodies. If a run timed out, review whether the external action completed before starting a replacement workflow.
280
-
281
- ## Collecting a useful bug report
282
-
283
- Include:
284
-
285
- - operating system and Node.js version;
286
- - Pi or Claude Code version;
287
- - package version or Git commit;
288
- - whether the hub is local or behind a proxy;
289
- - redacted environment values, excluding the token;
290
- - the relevant structured hub events;
291
- - exact reproduction steps and expected behavior.
292
-
293
- Never attach authentication tokens, private prompts, credentials, or unrelated repository contents.
@@ -1,240 +0,0 @@
1
- # Webhook workflows and long-lived agents
2
-
3
- KXM can turn a signed Jira, GitHub, or generic webhook into a durable prompt for a long-lived coordinator. The hub verifies the original request body, deduplicates provider retries, records the workflow before acknowledging it, and queues the prompt even when a previously registered coordinator is temporarily offline.
4
-
5
- ## How the runtime behaves
6
-
7
- ```text
8
- Jira webhook ── HMAC + delivery ID ──> Hub ── durable workflow + message
9
- │
10
- └── Pi coordinator
11
- ├── peer planning/review
12
- ├── checkpoints and retries
13
- ├── evidence journal
14
- ├── external wait ── signed result ──┐
15
- └── final result <── resumed prompt ─┘
16
- ```
17
-
18
- The coordinator must register at least once before a webhook can target it. An unknown target returns HTTP 409, which causes Jira Cloud to retry. A known but offline target retains the queued workflow until it reconnects.
19
-
20
- The hub stores a SHA-256 payload hash and the rendered coordinator prompt, not the complete raw webhook body. Keep prompt templates narrow so they copy only the issue fields the agent needs.
21
-
22
- ## Configure the Jira example
23
-
24
- The included [`jira-development.json`](../.kxm/workflows/default.yaml) workspace configuration models this path:
25
-
26
- 1. Jira issue enters **In Progress**.
27
- 2. Reproduce the defect and create deterministic evidence.
28
- 3. Plan with one agent or three independent strong planners, then synthesize their best ideas.
29
- 4. Review and revise the plan.
30
- 5. Implement with explicit ownership.
31
- 6. Run lint, build/typecheck, security, and Playwright gates.
32
- 7. Reproduce repository and CodeRabbit-style review gates.
33
- 8. Update documentation.
34
- 9. Push and watch required checks with `kxm gate github watch`; warnings and failures retry the same stage for correction until its attempt limit is exhausted.
35
- 10. Merge only when policy and authorization allow it.
36
- 11. Update Jira with links and evidence.
37
- 12. Produce an evidence-backed improvement backlog.
38
-
39
- Load it without storing its secret in the JSON file:
40
-
41
- ```powershell
42
- $env:JIRA_WEBHOOK_SECRET = "replace-with-a-high-entropy-secret"
43
- $env:WORKFLOW_SIGNAL_SECRET = "replace-with-a-separate-callback-secret"
44
- $env:KXM_WEBHOOK_WORKFLOWS_FILE = ".kxm/workflows/default.yaml"
45
- kxm hub start
46
- ```
47
-
48
- Configure Jira to send `jira:issue_updated` to:
49
-
50
- ```text
51
- https://your-kxm-host.example/v1/webhooks/jira-development
52
- ```
53
-
54
- Set the same secret when creating the Jira webhook. The endpoint requires `X-Hub-Signature` using SHA-256 and `X-Atlassian-Webhook-Identifier`. The stable delivery identifier makes Jira retries idempotent. Terminate TLS and restrict ingress before exposing the endpoint beyond a trusted network.
55
-
56
- Webhook authentication authorizes only workflow creation. The Jira-update stage requires a separate authorized Jira tool, MCP server, CLI, or automation callback in the coordinator's harness. Do not place Jira API credentials in the workflow definition or prompt.
57
-
58
- ## Run a long-lived Pi coordinator
59
-
60
- Install the Pi package, then configure a stable identity that matches the workflow target:
61
-
62
- ```powershell
63
- $env:KXM_SERVER_URL = "http://127.0.0.1:7331"
64
- $env:KXM_AUTH_TOKEN = "product-project-token"
65
- $env:KXM_PROJECT = "product"
66
- $env:KXM_AGENT_NAME = "coordinator"
67
- $env:KXM_AGENT_PURPOSE = "Coordinates Jira development workflows and quality gates"
68
- $env:KXM_WORKDIR = "D:\work\product-repository"
69
- kxm agent worker --session-isolation workflow
70
- ```
71
-
72
- Workflow isolation is explicit during the upgrade-compatible release and begins fresh scoped storage on first use. The worker launches Pi in headless RPC mode, keeps stdin open, preserves its active bound session by default, and restarts with bounded exponential backoff. Run the worker itself under the operating system's service manager for boot startup, resource limits, log collection, and crash policy. Set `KXM_WORKER_CONTINUE=false` only when every process restart should create a fresh Pi session.
73
-
74
- ## Workflow definition fields
75
-
76
- | Field | Meaning |
77
- |---|---|
78
- | `id` | URL-safe workflow identifier |
79
- | `source` | `jira`, `github`, or `generic` |
80
- | `project` | Hub project containing the coordinator |
81
- | `target` | Stable coordinator name or durable agent ID |
82
- | `secretEnv` | Environment variable containing the HMAC secret |
83
- | `signalSecretEnv` | Optional separate HMAC secret for external result callbacks |
84
- | `event` | Optional provider event filter |
85
- | `filter.path` / `filter.equals` | Optional exact JSON-path value filter |
86
- | `delivery` | `followUp` or `steer` |
87
- | `ttlMs` | Time allowed for the coordinator prompt |
88
- | `promptTemplate` | Prompt with `{{nested.payload.path}}` substitutions |
89
- | `stages` | Ordered gates with instructions, evidence requirements, and attempt limits |
90
- | `stages[].evidencePolicies` | Optional per-requirement peer provenance and quorum rules |
91
-
92
- Each stage may set `area` to route automatic warnings and failures into `harness`, `gates`, `implementation`, `workflow`, `documentation`, `security`, or `other`. It defaults to `workflow`.
93
-
94
- An `evidencePolicies` key must match one canonical `requiredEvidence` identity.
95
- A `peer-reply` policy declares a `minProducers`, one or more
96
- `eligibleAgents`, and `acceptedStatuses: ["replied"]`. Eligible names or IDs
97
- must already be known in the workflow project. The hub resolves them to stable
98
- producer IDs when the run starts and fails closed if the coordinator is
99
- included or the unique resolved set cannot satisfy the configured minimum.
100
- See [Peer provenance and quorum gates](provenance-gates.md) for the complete
101
- schema and command-first example.
102
-
103
- Use `KXM_WEBHOOK_WORKFLOWS` for inline JSON or `KXM_WEBHOOK_WORKFLOWS_FILE` for a file, never both. Prefer `secretEnv` over a literal `secret`.
104
-
105
- ## Pause for CI, review, merge, or Jira
106
-
107
- A coordinator should not hold an agent turn open while an external system runs for minutes or hours. On the active stage, call `kxm_workflow_wait` with:
108
-
109
- - the run and active stage IDs;
110
- - a stable `signalKey`, such as `github-pr-42-checks`;
111
- - a concise description of the expected result;
112
- - optional local evidence keyed by its declared requirement identity;
113
- - an optional timeout from one second through 30 days; the default is 24 hours.
114
-
115
- The hub changes the run and stage to `waiting`. The coordinator may then settle its current prompt without triggering the premature-settlement failure. If the deadline passes first, the run fails and records a harness error.
116
-
117
- The external system reports its result to:
118
-
119
- ```text
120
- POST /v1/webhooks/:definitionId/runs/:runId/signals/:signalKey
121
- ```
122
-
123
- The JSON body is:
124
-
125
- ```json
126
- {
127
- "status": "passed",
128
- "summary": "All required GitHub checks passed",
129
- "evidence": {
130
- "github.check:ci": "conclusion:success url:https://github.example/org/repo/actions/runs/123"
131
- }
132
- }
133
- ```
134
-
135
- Sign the exact body bytes with SHA-256 HMAC. Supply the signature in `X-Hub-Signature-256` and a stable retry identifier in `X-GitHub-Delivery`, `X-Atlassian-Webhook-Identifier`, or `X-Mesh-Delivery-ID`. Repeating the same delivery ID and body returns minimal receipt metadata instead of checkpointing twice. Reusing a delivery ID for a different signal or body returns HTTP 409.
136
-
137
- Use `signalSecretEnv` so CI and merge reporters do not need the secret that creates new workflows. If it is omitted, callbacks fall back to `secretEnv` for compatibility. A valid callback can checkpoint only the named run's current wait and must match its signal key.
138
-
139
- Context evidence is optional. When a callback supplies `workflow.run`,
140
- `workflow.stage`, or `workflow.signal`, each value must exactly match the route
141
- run, active waiting stage, or route signal key respectively. A mismatch returns
142
- HTTP 409 without advancing the run or recording a delivery receipt. Adapters
143
- that do not need these diagnostic keys may omit them.
144
-
145
- `passed` applies the normal evidence rule and advances or completes the run. `warning` or `failed` consumes an attempt, records an error, and queues a correction prompt when attempts remain. The run, signal receipt, optional journal entry, and optional resume message commit in one SQLite transaction before delivery. A terminal result does not create another prompt. Only the validated summary and evidence are retained; the complete callback body is not stored.
146
-
147
- Callback responses deliberately expose only status, stage, retry, completion, resumption, and duplicate metadata. They never return the workflow record, coordinator prompt, message routing, journal, or evidence. Those remain behind project and agent authentication.
148
-
149
- The repository includes a small callback sender for smoke tests and automation adapters:
150
-
151
- ```powershell
152
- $env:KXM_SERVER_URL = "https://your-hub-host.example"
153
- $env:KXM_WORKFLOW_ID = "jira-development"
154
- $env:KXM_WORKFLOW_SIGNAL_SECRET = "replace-with-the-callback-secret"
155
- $env:KXM_SIGNAL_DELIVERY_ID = "github-check-run-123-attempt-1"
156
- node --experimental-strip-types examples/workflow-signal.ts `
157
- run_123 github-pr-42-checks passed "All required checks passed" `
158
- "github.check:ci=https://github.example/org/repo/actions/runs/123"
159
- ```
160
-
161
- `examples/workflow-signal.ts` reads `KXM_SIGNAL_DELIVERY_ID` and sends it as
162
- `x-kxm-delivery-id`. The CLI equivalent is `kxm gate signal --delivery-id`.
163
-
164
- In a real integration, store the `runId` and `signalKey` in Jira, pull-request metadata, or the external job's inputs when the coordinator starts the wait. Treat them as routing identifiers rather than secrets.
165
-
166
- To watch GitHub checks and post that same signal, use the command-first adapter:
167
-
168
- ```powershell
169
- $env:KXM_WORKFLOW_ID = "jira-development"
170
- $env:KXM_WORKFLOW_SIGNAL_SECRET = "replace-with-the-callback-secret"
171
- $env:GITHUB_TOKEN = "replace-with-a-checks-read-token"
172
- kxm gate github watch --run-id run_123 --stage-id watch --signal-key github-pr-42-checks --repo org/repo --pr 42 --required ci --timeout-ms 3600000
173
- ```
174
-
175
- The watcher binds every result to the exact run, stage, and signal key, requests
176
- up to 100 check runs per GitHub page, and follows every reported page. Each
177
- check is reported as `github.check:<check-name>`; diagnostic context such as
178
- `workflow.run` never satisfies an unrelated requirement. GitHub
179
- `startup_failure` is a failed result. On timeout the watcher posts a signed
180
- `failed` signal with summary `github_watch_timeout`, then exits `4`; it never
181
- invents a `passed` result.
182
-
183
- Each watcher invocation creates a new bounded delivery generation and includes
184
- the pull-request head SHA when GitHub returned one. Transport retries within
185
- that invocation reuse the exact `x-kxm-delivery-id`. After a failed or timed
186
- out result, start a new watcher for the new workflow wait; do not reuse the old
187
- generated ID. Supply `--delivery-id` only when an external supervisor must
188
- retry the same callback attempt with a stable provider identifier. The standalone
189
- `kxm gate signal` command follows the same rule.
190
-
191
- ## Checkpoint contract
192
-
193
- Only the assigned coordinator can read, journal, checkpoint, or wait a run. A
194
- passing checkpoint must provide a non-empty value for every exact
195
- `requiredEvidence` identity. Evidence is a JSON object rather than a list, so
196
- extra GitHub checks or generic context cannot replace an unrelated review,
197
- artifact, or retrospective requirement. Identities are normalized by trimming,
198
- collapsing repeated whitespace, and case-folding; normalized aliases in one
199
- submission are rejected as duplicates.
200
-
201
- When a requirement has a peer policy, caller-authored evidence text cannot
202
- satisfy it. The coordinator must create peer messages with an authorized,
203
- immutable `workflowContext` for the exact run, active stage, canonical
204
- requirement, and current 1-based attempt. A passing checkpoint or wait cites the
205
- resulting durable message IDs in `evidenceRefs`. The hub verifies project,
206
- direction, eligible target, context, correlation, non-empty replied status,
207
- and coherent timestamps, then counts unique producer IDs. Old, pending,
208
- duplicate-producer, coordinator-authored, or cross-context messages do not
209
- count.
210
-
211
- Evidence supplied when entering `waiting` is accumulated with a later passing
212
- callback. `warning` and `failed` evidence is retained in the journal for
213
- diagnosis but intentionally does not satisfy a later passing attempt. Those
214
- results remain on the active stage and return a correction instruction.
215
- Reaching `maxAttempts` fails the run. Settling the coordinator prompt before all
216
- stages pass also fails the run and records a workflow error unless the
217
- coordinator deliberately placed the active stage in `waiting` first.
218
-
219
- If a policy declares `degradation.minProducers`, an operator may use
220
- `kxm gate degrade` with the administrative token to approve that exact
221
- lower minimum for only the current stage attempt. The coordinator, peer agents,
222
- and callback secret cannot authorize degradation. Approval alone never passes
223
- the stage; the coordinator must still provide the required verified references.
224
- The reason, policy minimum, approved minimum, attempt, and any eventual degraded
225
- pass are retained for audit.
226
-
227
- The hub enforces stage order and requirement identity; agents remain responsible
228
- for the truth of submitted evidence. Repository rules, human approvals, and
229
- harness permissions remain authoritative for push, merge, Jira mutation, and
230
- other external effects.
231
-
232
- Peer quorum proves provenance inside the hub project credential boundary. It
233
- does not prove answer quality, truth, distinct underlying models, independent
234
- inference, non-collusion, or human approval.
235
-
236
- ## Platform references
237
-
238
- - [Pi extension lifecycle and message injection](https://pi.dev/docs/latest/extensions)
239
- - [Pi headless RPC mode](https://pi.dev/docs/latest/rpc)
240
- - [Jira Cloud webhook signing and retry behavior](https://developer.atlassian.com/cloud/jira/software/webhooks/)