@kontextmind/kxm 0.7.94 → 0.7.96
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +1 -1
- package/.kxm/README.md +39 -9
- package/CHANGELOG.md +1 -1
- package/README.md +147 -257
- package/SECURITY.md +21 -12
- package/docs/README.md +133 -54
- package/docs/adr/ADR-0002-browser-automation-steel-doks.md +24 -18
- package/docs/adr/ADR-0003-sqlite-only-store.md +100 -0
- package/docs/adr/ADR-0004-edge-identity-authentik.md +99 -0
- package/docs/adr/README.md +33 -0
- package/docs/concepts/architecture.md +262 -0
- package/docs/concepts/data-and-storage.md +194 -0
- package/docs/concepts/trust-model.md +152 -0
- package/docs/contracts/README.md +22 -14
- package/docs/contracts/effects-and-recovery.md +3 -0
- package/docs/contracts/migration.md +2 -2
- package/docs/contracts/routing.md +6 -5
- package/docs/contributing/assignment-runner.md +388 -0
- package/docs/contributing/ci-and-release.md +231 -0
- package/docs/contributing/development.md +362 -0
- package/docs/contributing/harness-routing-internals.md +192 -0
- package/docs/{packages.md → contributing/packages.md} +13 -15
- package/docs/{skills → contributing}/repo-work-delivery.md +20 -21
- package/docs/contributing/test-matrix.md +208 -0
- package/docs/{tui-components.md → contributing/tui-components.md} +30 -22
- package/docs/contributing/writing-docs.md +340 -0
- package/docs/glossary.md +471 -0
- package/docs/guides/agent-skills.md +137 -0
- package/docs/guides/browser-automation.md +160 -0
- package/docs/guides/context-and-memory.md +352 -0
- package/docs/guides/continuous-improvement.md +228 -0
- package/docs/guides/governed-skills.md +173 -0
- package/docs/guides/nous-providers.md +186 -0
- package/docs/guides/peer-messaging.md +304 -0
- package/docs/guides/pi-workers.md +219 -0
- package/docs/guides/provenance-gates.md +313 -0
- package/docs/guides/webhook-workflows.md +364 -0
- package/docs/kb/how-credentials-retrieved-safely.md +38 -12
- package/docs/kb/how-to-capture-and-annotate-section.md +15 -13
- package/docs/kb/how-to-connect-playwright-to-steel.md +16 -11
- package/docs/kb/how-to-recover-expired-session-or-orphan.md +26 -16
- package/docs/kb/how-to-resume-after-mfa.md +19 -11
- package/docs/kb/how-to-take-over-session.md +17 -13
- package/docs/kb/why-authentication-disappeared.md +22 -14
- package/docs/kb/why-automation-opened-different-browser.md +23 -14
- package/docs/kb/why-session-viewer-cannot-control.md +13 -12
- package/docs/operations/backup-and-restore.md +248 -0
- package/docs/operations/deploy.md +307 -0
- package/docs/operations/monitoring.md +209 -0
- package/docs/operations/runtime-sync.md +192 -0
- package/docs/operations/troubleshooting.md +265 -0
- package/docs/operations/upgrade.md +124 -0
- package/docs/prompts/browser-annotate-feedback.md +7 -7
- package/docs/prompts/browser-diagnose-recover.md +11 -10
- package/docs/prompts/browser-explore.md +7 -7
- package/docs/prompts/browser-repro-fix.md +7 -7
- package/docs/prompts/browser-start.md +12 -11
- package/docs/prompts/browser-takeover.md +8 -8
- package/docs/{cli-reference.md → reference/cli-reference.md} +83 -41
- package/docs/{config-reference.md → reference/config-reference.md} +159 -148
- package/docs/reference/configuration.md +299 -0
- package/docs/reference/harness-routing.md +508 -0
- package/docs/reference/http-api.md +203 -0
- package/docs/reference/tools.md +370 -0
- package/docs/{workflow-guide.md → reference/workflow-catalog.md} +92 -153
- package/docs/reference/workflow-definitions.md +286 -0
- package/docs/start/first-workflow.md +287 -0
- package/docs/start/install.md +146 -0
- package/docs/start/quickstart-claude-code.md +405 -0
- package/docs/start/quickstart-pi.md +213 -0
- package/docs/templates/README.md +78 -73
- package/docs/templates/adr.md +13 -13
- package/docs/templates/architecture.md +55 -71
- package/docs/templates/bug-fix.md +13 -16
- package/docs/templates/feature.md +14 -19
- package/docs/templates/handoff.md +44 -46
- package/docs/templates/postmortem.md +30 -43
- package/docs/templates/research.md +15 -20
- package/docs/templates/review.md +49 -50
- package/docs/templates/runbook.md +38 -30
- package/docs/templates/test-plan.md +16 -23
- package/docs/templates/test-report.md +14 -17
- package/examples/README.md +9 -5
- package/examples/provenance-workflow.json +1 -1
- package/examples/webhook-workflows/jira-development.json +59 -0
- package/examples/webhook-workflows/jira-issue-updated.json +12 -0
- package/package.json +2 -2
- package/packages/core/tui/README.md +1 -1
- package/plugins/kxm/.claude-plugin/plugin.json +1 -1
- package/plugins/kxm/README.md +31 -32
- package/plugins/kxm/dist/cli.js +5 -5
- package/plugins/kxm/dist/mcp-server.js +1 -1
- package/plugins/kxm/dist/runtime.js +1 -1
- package/plugins/kxm/package.json +1 -1
- package/plugins/kxm/skills/kxm/references/protocol.md +3 -1
- package/plugins/kxm/skills/kxm-browser-auth/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-browser-diagnostics/SKILL.md +5 -5
- package/plugins/kxm/skills/kxm-browser-explore/SKILL.md +2 -2
- package/plugins/kxm/skills/kxm-browser-session/SKILL.md +10 -13
- package/plugins/kxm/skills/kxm-browser-takeover/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-browser-verify/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-context-memory/SKILL.md +13 -4
- package/plugins/kxm/skills/kxm-hub-ops/SKILL.md +3 -1
- package/plugins/kxm/skills/kxm-mind-setup/SKILL.md +2 -1
- package/plugins/kxm/skills/kxm-project-setup/SKILL.md +31 -54
- package/plugins/kxm/skills/kxm-projects/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-protocol/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-routing-improve/SKILL.md +15 -7
- package/plugins/kxm/skills/kxm-runs/SKILL.md +11 -5
- package/plugins/kxm/skills/kxm-session/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-tasks/SKILL.md +9 -7
- package/plugins/kxm/skills/kxm-workflow/SKILL.md +10 -2
- package/plugins/kxm/src/cli/system.ts +1 -1
- package/plugins/kxm/src/cli.ts +3 -3
- package/plugins/kxm/src/init-guide-setup.ts +1 -1
- package/plugins/kxm/src/mcp-server.ts +1 -1
- package/plugins/kxm/src/modes.ts +1 -1
- package/schemas/README.md +1 -1
- package/docs/agent-communication-envelopes-and-gates.md +0 -553
- package/docs/agent-skills.md +0 -198
- package/docs/architecture.md +0 -245
- package/docs/assignment-runner.md +0 -264
- package/docs/browser-automation.md +0 -139
- package/docs/configuration.md +0 -437
- package/docs/continuous-improvement.md +0 -226
- package/docs/getting-started.md +0 -277
- package/docs/harness-routing.md +0 -616
- package/docs/kb/qa-authentik-authentication.md +0 -97
- package/docs/kb/qa-extension-install-and-hub-bootstrap.md +0 -85
- package/docs/kb/qa-hub-on-a-public-host.md +0 -48
- package/docs/kb/qa-sqlite-vs-duckdb.md +0 -35
- package/docs/kb/qa-what-the-hub-stores.md +0 -64
- package/docs/kxm-handbook.md +0 -1181
- package/docs/operations.md +0 -510
- package/docs/operator-pi-packages.md +0 -67
- package/docs/provenance-gates.md +0 -295
- package/docs/skills.md +0 -47
- package/docs/test-matrix.md +0 -132
- package/docs/troubleshooting.md +0 -293
- package/docs/webhook-workflows.md +0 -240
|
@@ -0,0 +1,304 @@
|
|
|
1
|
+
# Message peer agents
|
|
2
|
+
|
|
3
|
+
Agents that share a KXM [hub](../glossary.md#hub) project can hand work to each other: send one focused request, keep working, and collect the reply when you need it. This guide shows each step in two forms, the MCP tool an agent calls in Claude Code or Pi and the `kxm peer` command an operator or script runs. It also covers delivery modes, safe retries, expiry, and the errors you are most likely to meet.
|
|
4
|
+
|
|
5
|
+
## Before you begin
|
|
6
|
+
|
|
7
|
+
- A running hub, for example one started with `kxm hub start`, with its URL in `KXM_SERVER_URL`.
|
|
8
|
+
- Two or more agents connected to the same hub project with that project's token: the Claude Code plugin ([quick start](../start/quickstart-claude-code.md)), Pi with the KXM package ([quick start](../start/quickstart-pi.md)), or a [supervised Pi worker](pi-workers.md).
|
|
9
|
+
- For the `kxm peer` examples, the `kxm` CLI with `KXM_AUTH_TOKEN` set to the project token (never the admin token), `KXM_PROJECT` set to the project, and a stable `KXM_AGENT_NAME`. [Use the CLI as an agent](#use-the-cli-as-an-agent) explains why the name matters.
|
|
10
|
+
|
|
11
|
+
## How a request moves through the hub
|
|
12
|
+
|
|
13
|
+
A request is a durable message record in the hub's SQLite store that moves from `queued` to exactly one terminal state.
|
|
14
|
+
|
|
15
|
+
```mermaid
|
|
16
|
+
stateDiagram-v2
|
|
17
|
+
state "error (reserved)" as err
|
|
18
|
+
[*] --> queued: sender posts
|
|
19
|
+
queued --> delivered: recipient acknowledges
|
|
20
|
+
queued --> replied: recipient replies
|
|
21
|
+
delivered --> replied: recipient replies
|
|
22
|
+
queued --> cancelled: sender cancels
|
|
23
|
+
delivered --> cancelled: sender cancels
|
|
24
|
+
queued --> expired: TTL elapses
|
|
25
|
+
delivered --> expired: TTL elapses
|
|
26
|
+
replied --> [*]
|
|
27
|
+
cancelled --> [*]
|
|
28
|
+
expired --> [*]
|
|
29
|
+
note right of queued
|
|
30
|
+
Stored in SQLite. Pushed again on every
|
|
31
|
+
reconnect until the recipient acknowledges it.
|
|
32
|
+
end note
|
|
33
|
+
classDef reserved stroke-dasharray:5 5
|
|
34
|
+
class err reserved
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
| State | Meaning |
|
|
38
|
+
|---|---|
|
|
39
|
+
| `queued` | Accepted and stored. Survives hub and agent restarts. The hub pushes it again each time the recipient reconnects, until the recipient acknowledges it. |
|
|
40
|
+
| `delivered` | Acknowledged: by Pi when the model turn starts, by the Claude Code plugin on arrival. It is not pushed again after a reconnect. |
|
|
41
|
+
| `replied`, `cancelled`, `expired` | Terminal. A later reply or cancel is refused. |
|
|
42
|
+
| `error` | Declared in the protocol but never set by the hub. Treat it as reserved. |
|
|
43
|
+
|
|
44
|
+
A recipient can reply while a request is still `queued`. Delivery is at least once, not exactly once, so make any side effect a peer performs safe to repeat.
|
|
45
|
+
|
|
46
|
+
## Discover peers
|
|
47
|
+
|
|
48
|
+
Call `kxm_list` to see the online agents in your project with their purpose, model, host label, and hub-clocked presence (`online`, `stale`, or `offline`). Set `includeOffline` to also list registered agents whose heartbeat lease has lapsed.
|
|
49
|
+
|
|
50
|
+
Tool call (`kxm_list`):
|
|
51
|
+
|
|
52
|
+
```json
|
|
53
|
+
{ "includeOffline": true }
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
CLI:
|
|
57
|
+
|
|
58
|
+
```bash
|
|
59
|
+
kxm peer list --include-offline
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Agent names are unique among live agents in a project. A second live registration with the same name is refused with `duplicate_agent_name`; the Claude Code plugin then registers as `<name>-<pid>` and says so.
|
|
63
|
+
|
|
64
|
+
## Send a request
|
|
65
|
+
|
|
66
|
+
Call `kxm_send` with a target (an agent name, case-insensitive, or an agent ID) and one focused task that states the response you expect. The call returns at once with a durable message ID.
|
|
67
|
+
|
|
68
|
+
Tool call (`kxm_send`):
|
|
69
|
+
|
|
70
|
+
```json
|
|
71
|
+
{
|
|
72
|
+
"target": "reviewer",
|
|
73
|
+
"content": "Review docs/plan.md for correctness and security risks. List blocking issues first.",
|
|
74
|
+
"idempotencyKey": "plan-review-1"
|
|
75
|
+
}
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
CLI:
|
|
79
|
+
|
|
80
|
+
```bash
|
|
81
|
+
kxm peer send reviewer "Review docs/plan.md for correctness and security risks. List blocking issues first." \
|
|
82
|
+
--idempotency-key plan-review-1
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
Expected output:
|
|
86
|
+
|
|
87
|
+
```text
|
|
88
|
+
{
|
|
89
|
+
"messageId": "msg_779f5e0f22e04ac1af6078589a874971",
|
|
90
|
+
"status": "queued",
|
|
91
|
+
"target": "reviewer"
|
|
92
|
+
}
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
| Field | CLI flag | Default | Purpose |
|
|
96
|
+
|---|---|---|---|
|
|
97
|
+
| `delivery` | `--delivery` | `followUp` | How the recipient schedules the work. See [Choose a delivery mode](#choose-a-delivery-mode). |
|
|
98
|
+
| `idempotencyKey` | `--idempotency-key` | none | Makes an exact retry return the original message. |
|
|
99
|
+
| `correlationId` | `--correlation-id` | none | Groups related requests. It carries no authority. |
|
|
100
|
+
| `ttlMs` | `--ttl-ms` | 24 hours | How long the request stays valid, from 1 second to 7 days. |
|
|
101
|
+
| `allowOffline` | `--allow-offline` | `false` | Queues the request for a registered agent that is offline. |
|
|
102
|
+
| `workflowContext` | `--workflow-context` | none | Binds the request to a workflow stage so its reply can count as evidence. See [Peer provenance and quorum gates](provenance-gates.md). |
|
|
103
|
+
|
|
104
|
+
> [!WARNING]
|
|
105
|
+
> The hub stores message bodies as sent, without application-level encryption. Never put credentials or unneeded private data in a request or reply.
|
|
106
|
+
|
|
107
|
+
## Check and wait for the reply
|
|
108
|
+
|
|
109
|
+
Call `kxm_get` to read the current status and any reply without blocking. Call `kxm_await` only when the reply blocks your next step: it waits up to 60 seconds, which is both the default and the maximum. A timed-out wait reports `timed out waiting for <message-id>` and leaves the request untouched, so check it again later with `kxm_get`.
|
|
110
|
+
|
|
111
|
+
Tool calls (`kxm_get`, then `kxm_await`):
|
|
112
|
+
|
|
113
|
+
```json
|
|
114
|
+
{ "messageId": "msg_779f5e0f22e04ac1af6078589a874971" }
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
```json
|
|
118
|
+
{ "messageId": "msg_779f5e0f22e04ac1af6078589a874971", "timeoutMs": 30000 }
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
CLI:
|
|
122
|
+
|
|
123
|
+
```bash
|
|
124
|
+
kxm peer get msg_779f5e0f22e04ac1af6078589a874971
|
|
125
|
+
kxm peer await msg_779f5e0f22e04ac1af6078589a874971 --timeout-ms 30000
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
Only the sender and the recipient can read a message. For work that runs for minutes or hours, do not loop on `kxm_await`: put the work in a [webhook workflow](webhook-workflows.md) and pause it with `kxm_workflow_wait`.
|
|
129
|
+
|
|
130
|
+
## Ask up to three peers at once
|
|
131
|
+
|
|
132
|
+
Call `kxm_fanout` to [fan out](../glossary.md#fanout) the same request independently to one, two, or three peers and collect the replies for comparison. Duplicate names are merged, every request uses `followUp`, and the call waits locally for up to `timeoutMs` (default and maximum 30 minutes).
|
|
133
|
+
|
|
134
|
+
Tool call (`kxm_fanout`):
|
|
135
|
+
|
|
136
|
+
```json
|
|
137
|
+
{
|
|
138
|
+
"targets": ["planner", "reviewer"],
|
|
139
|
+
"content": "Propose a bounded plan for PROD-123 with file ownership and risks.",
|
|
140
|
+
"idempotencyKeyPrefix": "prod-123-plan",
|
|
141
|
+
"timeoutMs": 600000
|
|
142
|
+
}
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
CLI (`--targets` takes space-separated names):
|
|
146
|
+
|
|
147
|
+
```bash
|
|
148
|
+
kxm peer fanout --targets planner reviewer \
|
|
149
|
+
--content "Propose a bounded plan for PROD-123 with file ownership and risks." \
|
|
150
|
+
--idempotency-key-prefix prod-123-plan --timeout-ms 600000
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
Each entry in `responses` is either terminal (`replied`, `cancelled`, `expired`, or `error`) or `pending`. A pending entry means only that your local wait ended:
|
|
154
|
+
|
|
155
|
+
```text
|
|
156
|
+
{
|
|
157
|
+
"responses": [
|
|
158
|
+
{
|
|
159
|
+
"target": "reviewer",
|
|
160
|
+
"messageId": "msg_7aa3182bcadb476fbcaf2858878b172d",
|
|
161
|
+
"status": "pending",
|
|
162
|
+
"messageStatus": "queued",
|
|
163
|
+
"expiresAt": "2026-09-24T18:30:24.688Z",
|
|
164
|
+
"waitStatus": "timed_out"
|
|
165
|
+
}
|
|
166
|
+
]
|
|
167
|
+
}
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
Keep the `messageId` handles and follow up with `kxm_get`, or repeat the exact call: the same prefix, correlation ID, targets, and workflow context resolve to the same messages instead of duplicates. Do not switch to a new prefix while earlier work is pending. Fanout reaches online peers only; an offline target comes back as `error` with `online target not found`.
|
|
171
|
+
|
|
172
|
+
Compare the replies yourself. A reply is technical input from another model, not a verified fact.
|
|
173
|
+
|
|
174
|
+
## Cancel a request
|
|
175
|
+
|
|
176
|
+
Call `kxm_cancel` (or run `kxm peer cancel <message-id>`) to withdraw a `queued` or `delivered` request you sent. The recipient receives a `cancelled` event: Pi drops the request from its queue, and the Claude Code inbox removes it. Cancelling an already-cancelled request returns it unchanged; a replied or expired request cannot be cancelled.
|
|
177
|
+
|
|
178
|
+
Cancellation is not rollback. If the peer already edited files or called an external system, check and undo that work yourself.
|
|
179
|
+
|
|
180
|
+
## Receive and reply
|
|
181
|
+
|
|
182
|
+
How an agent receives requests depends on its harness.
|
|
183
|
+
|
|
184
|
+
**Pi** handles inbound work for you. The KXM extension takes one request at a time (`steer` first, then `followUp`, then `nextTurn`), acknowledges it as the model turn starts, and returns the settled final answer as the reply. A reply longer than 32,000 characters is truncated with a note. You do not call `kxm_reply` in Pi.
|
|
185
|
+
|
|
186
|
+
**Claude Code** has two modes, described fully in the [plugin guide](../../plugins/kxm/README.md#pushed-channel-mode-and-pull-mode):
|
|
187
|
+
|
|
188
|
+
- **Pushed channel mode.** Start Claude Code with the KXM channel. Each request arrives as a `<channel source="kxm" message_id="...">` event; Claude handles it and calls `kxm_reply`.
|
|
189
|
+
- **Pull mode**, the default. Claude calls `kxm_inbox` to list requests that still need a reply, handles one, and calls `kxm_reply`. Nothing arrives on its own, so ask Claude to check the inbox or to poll it with a backoff.
|
|
190
|
+
|
|
191
|
+
```bash
|
|
192
|
+
# Organization-approved channel
|
|
193
|
+
claude --channels plugin:kxm@kxm
|
|
194
|
+
# During the channels research preview
|
|
195
|
+
claude --dangerously-load-development-channels plugin:kxm@kxm
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
Tool call (`kxm_reply`):
|
|
199
|
+
|
|
200
|
+
```json
|
|
201
|
+
{
|
|
202
|
+
"messageId": "msg_779f5e0f22e04ac1af6078589a874971",
|
|
203
|
+
"content": "The plan is sound. Add a test for a zero discount. No blocking issues."
|
|
204
|
+
}
|
|
205
|
+
```
|
|
206
|
+
|
|
207
|
+
CLI (run with the recipient's `KXM_AGENT_NAME`):
|
|
208
|
+
|
|
209
|
+
```bash
|
|
210
|
+
kxm peer reply msg_779f5e0f22e04ac1af6078589a874971 "The plan is sound. Add a test for a zero discount."
|
|
211
|
+
```
|
|
212
|
+
|
|
213
|
+
Only the recipient can reply, and only once. From the CLI, `kxm peer inbox` always returns an empty list, because a one-shot command has no long-running inbox. Use `kxm dash --screen inbox` to see pending requests.
|
|
214
|
+
|
|
215
|
+
## Choose a delivery mode
|
|
216
|
+
|
|
217
|
+
A [delivery mode](../glossary.md#delivery-mode) tells the recipient how to schedule the work.
|
|
218
|
+
|
|
219
|
+
| Mode | Use it for | Effect in Pi |
|
|
220
|
+
|---|---|---|
|
|
221
|
+
| `followUp` | Normal delegation. The default. | Handled after the current work settles. |
|
|
222
|
+
| `steer` | An active blocker that must change course. | Moves ahead of queued work and starts at the next safe turn boundary. It does not abort a running tool call or write. |
|
|
223
|
+
| `nextTurn` | Low-priority context. | Queued after other work, then run as a follow-up turn, because an autonomous worker has no later human prompt to wait for. |
|
|
224
|
+
|
|
225
|
+
In Claude Code the mode arrives as `delivery` in the channel metadata, and the session decides what to do with it. Webhook workflow definitions accept only `followUp` and `steer`.
|
|
226
|
+
|
|
227
|
+
## Make retries safe
|
|
228
|
+
|
|
229
|
+
An **idempotency key** (up to 128 characters) is scoped to the sender. Repeating a send with the same key and identical fields returns the original message instead of a duplicate. Reusing the key with any different field (target, content, delivery, correlation ID, TTL, or workflow context) fails with `idempotency_conflict`. The key is remembered for as long as the hub retains the message.
|
|
230
|
+
|
|
231
|
+
A **correlation ID** (up to 128 characters) groups related requests for your own bookkeeping. It does not deduplicate or authorize anything. When a request carries `workflowContext`, the hub sets the correlation ID to the workflow run ID and refuses any other value.
|
|
232
|
+
|
|
233
|
+
Neither field proves where a reply came from. Workflow evidence uses `workflowContext`, which the hub authorizes and stores; see [Peer provenance and quorum gates](provenance-gates.md).
|
|
234
|
+
|
|
235
|
+
## Limits, expiry, and retention
|
|
236
|
+
|
|
237
|
+
| Limit | Value | Notes |
|
|
238
|
+
|---|---|---|
|
|
239
|
+
| Request or reply content | 32,000 characters | The whole HTTP body is capped at 256 KiB. |
|
|
240
|
+
| Time to live (TTL) | 24 hours by default; 1 second to 7 days per request | Counted from the send, including time spent queued. The hub default is `KXM_MESSAGE_TTL_MS`. |
|
|
241
|
+
| `kxm_await` | 60 seconds | Default and maximum. |
|
|
242
|
+
| `kxm_fanout` local wait | 30 minutes | Default and maximum. |
|
|
243
|
+
| Hop limit | 5 by default; `maxHops` 1 to 20 | See the note below. |
|
|
244
|
+
| Retention | 7 days after a terminal state | Set with `KXM_MESSAGE_RETENTION_MS`. A purged message returns `message not found`. |
|
|
245
|
+
| Rate limit | 600 requests per 60 seconds | Counted per agent, or per client address. Set with `KXM_RATE_LIMIT_MAX` and `KXM_RATE_LIMIT_WINDOW_MS`. |
|
|
246
|
+
|
|
247
|
+
When a request expires, both the sender and the recipient receive an `expired` event.
|
|
248
|
+
|
|
249
|
+
The hop limit refuses a request whose `hops` count has reached `maxHops` (`hop_limit_reached`). Only a client that forwards a request through the [Hub HTTP API reference](../reference/http-api.md) and carries its hop count is affected. `kxm_send`, `kxm_fanout`, and `kxm peer send` always start at hop 0, so the limit does not stop a chain of agents that each send a new request.
|
|
250
|
+
|
|
251
|
+
## Queue work for an offline agent
|
|
252
|
+
|
|
253
|
+
By default a send fails unless the target is online. Set `allowOffline` (`--allow-offline`) to queue the request for an agent that has registered in this project before but is offline now. The request stays `queued` until the agent reconnects or its TTL passes. An unknown name still fails, and `kxm_fanout` has no offline option.
|
|
254
|
+
|
|
255
|
+
## Use the CLI as an agent
|
|
256
|
+
|
|
257
|
+
Each `kxm peer` command registers with the hub as `KXM_AGENT_NAME` (default `cli-<pid>`), runs one call, and unregisters. This has four consequences:
|
|
258
|
+
|
|
259
|
+
1. Set a stable `KXM_AGENT_NAME`. Later `get`, `await`, and `cancel` commands must run as the sender; any other name gets `message is not visible to this agent`.
|
|
260
|
+
2. Between commands the CLI identity is offline. Peers can still reply to it; read the reply with `kxm peer get`. To send work to a CLI identity, use `--allow-offline`.
|
|
261
|
+
3. The name must not be live elsewhere, or registration fails with `duplicate_agent_name`.
|
|
262
|
+
4. `--dry-run` prints the parsed arguments without contacting the hub.
|
|
263
|
+
|
|
264
|
+
On failure the CLI prints the hub's message, not its code, for example `{"ok":false,"error":"command_failed","detail":"cannot send a request to yourself"}`.
|
|
265
|
+
|
|
266
|
+
## Errors worth knowing
|
|
267
|
+
|
|
268
|
+
Tools and the CLI report the message text; the HTTP response also carries the code.
|
|
269
|
+
|
|
270
|
+
| Message | Code | Cause and fix |
|
|
271
|
+
|---|---|---|
|
|
272
|
+
| `online target not found: <name>` | `target_not_found` | The target is offline or unknown. Check `kxm_list`; use `allowOffline` for a registered agent. |
|
|
273
|
+
| `cannot send a request to yourself` | `self_target` | Send to a different agent. |
|
|
274
|
+
| `idempotency key was already used for another request` | `idempotency_conflict` | A different request reused the key. Repeat the original exactly, or use a new key for new work. |
|
|
275
|
+
| `hop limit reached (<hops>/<max>)` | `hop_limit_reached` | A forwarding client reached `maxHops`. Stop the chain. |
|
|
276
|
+
| `ttlMs must be an integer between 1000 and 604800000` | `protocol_error` | Use a TTL from 1 second to 7 days. |
|
|
277
|
+
| `message is not visible to this agent` | `message_forbidden` | Run as the sender or recipient. |
|
|
278
|
+
| `message already has a reply` | `duplicate_reply` | The request is already answered. |
|
|
279
|
+
| `cannot cancel a replied message` | `invalid_message_state` | The request is terminal. Read it with `kxm_get`. |
|
|
280
|
+
| `message not found` | `message_not_found` | Wrong ID, or the message was purged after retention. |
|
|
281
|
+
| `agent name already active in project: <name>` | `duplicate_agent_name` | Another live agent holds the name. Stop it or choose another name. |
|
|
282
|
+
| `request rate limit exceeded` | `rate_limited` | Back off for the `retry-after` seconds. |
|
|
283
|
+
|
|
284
|
+
Errors about `workflowContext` are covered in [Peer provenance and quorum gates](provenance-gates.md).
|
|
285
|
+
|
|
286
|
+
## Troubleshooting
|
|
287
|
+
|
|
288
|
+
| Symptom | Cause | Fix |
|
|
289
|
+
|---|---|---|
|
|
290
|
+
| A request stays `queued` | The recipient is offline, busy with earlier work, or swapping its Pi session for a workflow run. | Check `kxm_list` and the recipient's log. Do not send a duplicate. |
|
|
291
|
+
| A request stays `delivered` | The recipient's turn, tool, or provider call is still running, or a Claude Code session restarted after acknowledging it. | Wait, or cancel and send it again with a new idempotency key. |
|
|
292
|
+
| `kxm_fanout` returns `pending` | The local wait ended before a reply. | Use the returned message IDs with `kxm_get`, or repeat the exact call. |
|
|
293
|
+
| Claude Code never sees requests | Channel mode is off or blocked by policy. | Use `kxm_inbox` and `kxm_reply`. |
|
|
294
|
+
| `kxm peer inbox` is always empty | The CLI has no long-running inbox. | Use `kxm dash --screen inbox`. |
|
|
295
|
+
|
|
296
|
+
For hub-level problems, see [Troubleshoot KXM](../operations/troubleshooting.md).
|
|
297
|
+
|
|
298
|
+
## Next steps
|
|
299
|
+
|
|
300
|
+
- Run long-lived Pi agents that answer requests unattended: [Run supervised Pi workers](pi-workers.md)
|
|
301
|
+
- Require verified replies from named peers before a stage passes: [Peer provenance and quorum gates](provenance-gates.md)
|
|
302
|
+
- Drive multi-stage work from a Jira or GitHub webhook: [Run webhook workflows](webhook-workflows.md)
|
|
303
|
+
- Every tool parameter: [Agent tools reference](../reference/tools.md); every flag: [CLI reference](../reference/cli-reference.md#kxm-peer)
|
|
304
|
+
- How the hub, agents, and stores fit together: [Architecture](../concepts/architecture.md)
|
|
@@ -0,0 +1,219 @@
|
|
|
1
|
+
# Run supervised Pi workers
|
|
2
|
+
|
|
3
|
+
A [worker](../glossary.md#worker) is a long-lived Pi process that stays connected to the KXM hub and answers peer requests and workflow prompts without anyone at the keyboard. `kxm agent worker` starts Pi in headless RPC mode under a small supervisor that restarts it, rotates to fallback models, and can keep each workflow run in its own Pi session. This guide covers starting a worker, choosing models and tools, session isolation, recovery, settings, and stale-worker problems.
|
|
4
|
+
|
|
5
|
+
## Before you begin
|
|
6
|
+
|
|
7
|
+
- The `kxm` CLI, and Pi with the KXM package installed: see [Install KXM](../start/install.md) and the [Pi quick start](../start/quickstart-pi.md). Put `pi` on `PATH` or set `KXM_PI_COMMAND`.
|
|
8
|
+
- A running hub and its URL in `KXM_SERVER_URL`.
|
|
9
|
+
- The hub project's token, which the worker uses as `KXM_AUTH_TOKEN`. Never give a worker the admin token.
|
|
10
|
+
- A checkout of the repository the worker should work in. The worker runs Pi in `KXM_WORKDIR`, or the current directory.
|
|
11
|
+
- A model route that Pi may run. See [Choose models and fallbacks](#choose-models-and-fallbacks).
|
|
12
|
+
|
|
13
|
+
## What the worker does
|
|
14
|
+
|
|
15
|
+
The supervisor launches `pi --mode rpc` with the KXM extension, keeps its input open, and captures its output. The hub stays the only durable queue:
|
|
16
|
+
|
|
17
|
+
- The extension takes one request at a time and acknowledges it only when the model turn starts, so waiting work stays `queued` on the hub and survives a restart.
|
|
18
|
+
- When the turn settles, the extension returns the final answer as the reply. See [Message peer agents](peer-messaging.md) for the message lifecycle.
|
|
19
|
+
- If Pi exits, the supervisor restarts it with exponential backoff from 1 second to 30 seconds. The backoff resets after a child has run for a minute.
|
|
20
|
+
- A PID claim in `.kxm/state` stops a second supervisor for the same project and agent name.
|
|
21
|
+
|
|
22
|
+
The worker does not read models, tools, roles, or ownership from any project file. Pass them as flags or environment variables. Run the worker itself under your service manager (systemd, launchd, or similar) for start at boot, resource limits, and log collection.
|
|
23
|
+
|
|
24
|
+
## Start a worker
|
|
25
|
+
|
|
26
|
+
Set the hub connection and project token, then start one worker per agent name.
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
export KXM_SERVER_URL=http://127.0.0.1:7331
|
|
30
|
+
export KXM_AUTH_TOKEN="replace-with-the-project-token"
|
|
31
|
+
export KXM_WORKDIR=~/work/product
|
|
32
|
+
kxm agent worker --name reviewer --project product \
|
|
33
|
+
--model <pi-model> --tools read,grep,find,ls \
|
|
34
|
+
--session-isolation workflow
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
<details><summary>PowerShell</summary>
|
|
38
|
+
|
|
39
|
+
```powershell
|
|
40
|
+
$env:KXM_SERVER_URL = "http://127.0.0.1:7331"
|
|
41
|
+
$env:KXM_AUTH_TOKEN = "replace-with-the-project-token"
|
|
42
|
+
$env:KXM_WORKDIR = "$HOME\work\product"
|
|
43
|
+
kxm agent worker --name reviewer --project product `
|
|
44
|
+
--model <pi-model> --tools read,grep,find,ls `
|
|
45
|
+
--session-isolation workflow
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
</details>
|
|
49
|
+
|
|
50
|
+
The command runs in the foreground until it is stopped. Add `--dry-run` to check the settings without starting anything:
|
|
51
|
+
|
|
52
|
+
```bash
|
|
53
|
+
kxm agent worker --name reviewer --project product --tools read,grep,find,ls --session-isolation workflow --dry-run
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
Expected output:
|
|
57
|
+
|
|
58
|
+
```text
|
|
59
|
+
would start worker
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
> [!IMPORTANT]
|
|
63
|
+
> Set `KXM_AUTH_TOKEN` to the project token in the worker's environment. When it is unset, the Pi extension falls back to the hub credential persisted on this machine, which is the admin token for a hub that `kxm hub start` or Pi auto-start created.
|
|
64
|
+
|
|
65
|
+
| Flag | Environment variable | Effect |
|
|
66
|
+
|---|---|---|
|
|
67
|
+
| `--name` | `KXM_AGENT_NAME` | Agent name on the hub. Required. |
|
|
68
|
+
| `--project` | `KXM_PROJECT` | Hub project. Required. |
|
|
69
|
+
| `--model` | `KXM_WORKER_MODEL` | Primary model. Pi's default when unset. |
|
|
70
|
+
| `--fallback-models` | `KXM_WORKER_FALLBACK_MODELS` | Up to eight comma-separated fallbacks. Requires a primary model. |
|
|
71
|
+
| `--tools` | `KXM_WORKER_TOOLS` | Comma-separated Pi tool allowlist. |
|
|
72
|
+
| `--session-isolation` | `KXM_WORKER_SESSION_ISOLATION` | `workflow` or `off` (the default). |
|
|
73
|
+
| `--fresh-start` | `KXM_WORKER_INITIAL_CONTINUE=false` | Skip resuming the previous Pi session on first start only. |
|
|
74
|
+
| `--no-continue` | `KXM_WORKER_CONTINUE=false` | Never resume a previous Pi session. |
|
|
75
|
+
|
|
76
|
+
## Choose models and fallbacks
|
|
77
|
+
|
|
78
|
+
`--model` selects the primary Pi model. `--fallback-models` lists up to eight more, tried in order when a provider fails for good. Pi finishes its own transient retries first; then the supervisor closes the session cleanly and restarts on the next model, resuming the same Pi session so completed tool and peer results are kept.
|
|
79
|
+
|
|
80
|
+
When no unused fallback remains, the worker waits `KXM_WORKER_PROVIDER_RETRY_MS` (60 seconds by default) and tries the last model again. It does not return to the primary model until you restart the worker.
|
|
81
|
+
|
|
82
|
+
> [!IMPORTANT]
|
|
83
|
+
> A worker refuses to start on a model whose vendor has its own native harness. Before Pi starts, the worker runs `--model` and every `--fallback-models` entry through the Pi native-vendor brake, and exits 1 with `pi_native_impersonation_blocked` on the first one it refuses: a native provider (`xai/…`), that vendor's own Pi provider (`openai-codex/…`, `kimi-coding/…`), or an aggregator path to it (`openrouter/x-ai/…`). `--dry-run` does not run this check.
|
|
84
|
+
|
|
85
|
+
Pick every model with [Harness routing](../reference/harness-routing.md), for example `openrouter/qwen/qwen3-coder-plus` with the fallback `openrouter/z-ai/glm-5.3-flash`. The brake cannot check a worker started without `--model`, which runs Pi's default model, or a bare model id that lets Pi choose the provider, so name the provider.
|
|
86
|
+
|
|
87
|
+
An agent name is only a label. To use a native-harness model as a peer, connect that harness to the hub under the agent name instead, for example the Claude Code plugin with the name `reviewer-claude`.
|
|
88
|
+
|
|
89
|
+
## Restrict tools
|
|
90
|
+
|
|
91
|
+
`--tools` passes an allowlist to Pi. It limits which tools the model can call, not which files those tools can reach: a worker that keeps `write`, `edit`, or a shell tool can change anything its operating-system user can.
|
|
92
|
+
|
|
93
|
+
- **Read-only reviewer:** `--tools read,grep,find,ls`. A reviewer does not need any `kxm_*` tool to answer; the extension returns its final answer as the reply.
|
|
94
|
+
- **Coordinator of a [webhook workflow](webhook-workflows.md):** add the hub tools the workflow needs, for example `kxm_list`, `kxm_send`, `kxm_fanout`, `kxm_get`, `kxm_await`, `kxm_workflow_get`, `kxm_workflow_checkpoint`, `kxm_workflow_wait`, `kxm_workflow_record`, and `kxm_improvement_report`, plus only the edit and shell tools its stages require.
|
|
95
|
+
|
|
96
|
+
```bash
|
|
97
|
+
kxm agent worker --name coordinator --project product \
|
|
98
|
+
--model <pi-model> --fallback-models <fallback-model> \
|
|
99
|
+
--tools read,grep,find,ls,edit,write,bash,kxm_send,kxm_fanout,kxm_get,kxm_await,kxm_workflow_get,kxm_workflow_checkpoint,kxm_workflow_wait,kxm_workflow_record \
|
|
100
|
+
--session-isolation workflow
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
For stronger boundaries, give each worker its own operating-system user, a read-only worktree, or a container. `KXM_WORKER_EXTENSION_PATHS` and `KXM_WORKER_SKILL_PATHS` load exact extension and skill files instead of Pi's discovery; treat those paths as executable code with the worker's credentials, and never derive them from a webhook or workflow payload.
|
|
104
|
+
|
|
105
|
+
## Isolate sessions per workflow run
|
|
106
|
+
|
|
107
|
+
With `--session-isolation workflow`, one worker keeps a stable default Pi session for ordinary work and a separate Pi session for each hub workflow run, so unrelated histories never mix in one context window.
|
|
108
|
+
|
|
109
|
+
| Message | Pi session |
|
|
110
|
+
|---|---|
|
|
111
|
+
| Ordinary peer or operator request | The default session |
|
|
112
|
+
| Root workflow prompt, signal resume, or wait timeout notice | That run's session |
|
|
113
|
+
| Peer request with an authorized `workflowContext` | That run's session |
|
|
114
|
+
| Request with only a `run_…` correlation ID | The default session. A correlation ID is not authorization. |
|
|
115
|
+
|
|
116
|
+
A switch never interrupts a turn and never runs two Pi processes at once. The next diagram shows a queued message for another run moving the worker to that run's session.
|
|
117
|
+
|
|
118
|
+
```mermaid
|
|
119
|
+
sequenceDiagram
|
|
120
|
+
participant Hub
|
|
121
|
+
participant A as Pi child (default session)
|
|
122
|
+
participant Sup as Worker supervisor
|
|
123
|
+
participant B as Pi child (run_1 session)
|
|
124
|
+
Hub->>A: push msg_1 for workflow run_1
|
|
125
|
+
A->>A: binding differs, leave msg_1 queued
|
|
126
|
+
A->>Sup: write route request (metadata only)
|
|
127
|
+
A-->>Sup: child closes
|
|
128
|
+
Sup->>Sup: update the binding manifest
|
|
129
|
+
Sup->>B: start one child with --session-dir runs/run_1
|
|
130
|
+
B->>Hub: reconnect
|
|
131
|
+
Hub->>B: replay msg_1 (same message ID)
|
|
132
|
+
B->>Hub: acknowledge and run the turn
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
The route request holds identity, supervisor generation, child incarnation, source and destination bindings, the Pi session ID, the message ID, and a timestamp, never a message body. The supervisor rejects a malformed, stale, or wrong-owner request without changing scope, and refuses session directories that are links.
|
|
136
|
+
|
|
137
|
+
Session files live under `.kxm/state/pi-sessions/<worker-key>/default` and `.../runs/<run-id>`, and the active binding in `.kxm/state/worker-session-binding-<worker-key>.json`. After a restart the worker resumes the bound session only if it has Pi history. It keeps up to `KXM_WORKER_MAX_RUN_SESSIONS` run sessions (128 by default) and deletes the least recently used inactive ones beyond that. A corrupt manifest is quarantined with a `.corrupt-<timestamp>` suffix and replaced by the default binding.
|
|
138
|
+
|
|
139
|
+
Isolation is off by default for upgrade compatibility. The first start with `workflow` begins fresh scoped sessions; KXM does not copy the old shared history, because it cannot be attributed to one run safely.
|
|
140
|
+
|
|
141
|
+
> [!NOTE]
|
|
142
|
+
> Pi session files are a convenience, not the record. The hub's workflow run, journal, messages, workspace assets, and Git are the recovery authority. Automatic per-run sessions apply to supervised Pi workers only; for Claude Code, use a separate session or agent name per run.
|
|
143
|
+
|
|
144
|
+
Isolation keeps model context apart. It is not a sandbox: a shell-capable agent can still read files and environment values that its operating-system user can reach.
|
|
145
|
+
|
|
146
|
+
## How the worker recovers
|
|
147
|
+
|
|
148
|
+
| Failure | What the worker does |
|
|
149
|
+
|---|---|
|
|
150
|
+
| Pi exits or crashes | Restarts with backoff. Stops after `KXM_WORKER_MAX_RESTARTS` restarts when set. |
|
|
151
|
+
| Final provider failure | Leaves the request `delivered`, records bounded metadata, and restarts on the next fallback model. |
|
|
152
|
+
| A tool runs past `KXM_WORKER_TOOL_TIMEOUT_MS` | Stops the Pi process tree and resumes the request without changing model. |
|
|
153
|
+
| A delivered request does not start a turn within `KXM_WORKER_ACTIVATION_TIMEOUT_MS` | Requests a restart and keeps the hub claim. |
|
|
154
|
+
| Pi cannot resume a saved session | Retries once with a fresh session and writes a recovery envelope in `.kxm/state`. |
|
|
155
|
+
| A workflow run needs a fresh session | Sends Pi a recovery prompt to call `kxm_workflow_get` and continue the current stage without repeating completed work. |
|
|
156
|
+
|
|
157
|
+
The supervisor writes a structured lifecycle log to `.kxm/logs/kxm-worker-<worker-key>.jsonl` with bounded metadata only. Raw Pi output, which can contain model and tool output, goes to `.kxm/logs/pi-agent-<worker-key>.log`; protect it accordingly.
|
|
158
|
+
|
|
159
|
+
| Log event | Meaning | Action |
|
|
160
|
+
|---|---|---|
|
|
161
|
+
| `worker_session_routed` | Expected swap to another run's session | None, unless it repeats for one message |
|
|
162
|
+
| `worker_session_evicted` | An inactive run session was deleted at the limit | Keep workflow facts in the journal, assets, and Git |
|
|
163
|
+
| `worker_session_state_recovered` | A bad manifest was quarantined | Inspect the `.corrupt-*` file, the hub run, and disk health |
|
|
164
|
+
| `worker_session_request_rejected` | A route request failed identity or schema checks | Check for a version mismatch or a duplicate supervisor |
|
|
165
|
+
| `worker_continue_fallback` | Pi history could not resume; started fresh | Read the recovery journal and the message state |
|
|
166
|
+
| `worker_provider_failure` | A provider failed after Pi's own retries | Check the provider and fallback list |
|
|
167
|
+
| `worker_restart_limit_reached` | `KXM_WORKER_MAX_RESTARTS` was hit | Fix the cause, then start the worker again |
|
|
168
|
+
|
|
169
|
+
## Worker settings
|
|
170
|
+
|
|
171
|
+
These variables configure a worker; flags override them. The complete list, with limits, is in the [environment variable reference](../reference/configuration.md).
|
|
172
|
+
|
|
173
|
+
| Variable | Default | Effect |
|
|
174
|
+
|---|---|---|
|
|
175
|
+
| `KXM_WORKDIR` | Current directory | Repository Pi works in; `.kxm` paths resolve inside it. |
|
|
176
|
+
| `KXM_PI_COMMAND` | `pi` (`pi.cmd` on Windows) | Pi executable. |
|
|
177
|
+
| `KXM_WORKER_MAX_RUN_SESSIONS` | `128` | Run sessions kept per worker, 1 to 1024. |
|
|
178
|
+
| `KXM_WORKER_TOOL_TIMEOUT_MS` | `1860000` | Longest single tool call, 1 second to 24 hours; `0` disables. The default sits one minute above the longest hub wait. |
|
|
179
|
+
| `KXM_WORKER_ACTIVATION_TIMEOUT_MS` | `60000` | Time for a delivered request to start a turn, 1 second to 10 minutes. |
|
|
180
|
+
| `KXM_WORKER_PROVIDER_RETRY_MS` | `60000` | Wait after the last fallback fails, 1 second to 1 hour. |
|
|
181
|
+
| `KXM_WORKER_DRAIN_MS` | `15000` | Graceful stop window before the child is killed. |
|
|
182
|
+
| `KXM_WORKER_MAX_RESTARTS` | Unlimited | Restart ceiling. |
|
|
183
|
+
| `KXM_WORKER_EXTENSION_PATHS` | Pi discovery | Exact extension files, separated by `:` (`;` on Windows). |
|
|
184
|
+
| `KXM_WORKER_SKILL_PATHS` | Pi discovery | Exact skill files or directories, same separator. |
|
|
185
|
+
| `KXM_WORKER_LOG_PATH`, `KXM_AGENT_LOG_PATH` | `.kxm/logs/…` | Lifecycle log and raw Pi log locations. |
|
|
186
|
+
|
|
187
|
+
Variables such as `KXM_WORKER_IDENTITY_KEY` and `KXM_WORKER_GENERATION` are set by the supervisor for its child; do not set them yourself. Restart the worker after changing any setting.
|
|
188
|
+
|
|
189
|
+
## Stop a worker
|
|
190
|
+
|
|
191
|
+
Press Ctrl+C in the worker's terminal, or send it `SIGTERM`. The supervisor gives Pi `KXM_WORKER_DRAIN_MS` to finish, then stops it. `kxm hub stop` and `kxm session stop` stop every managed hub and worker process in the workspace, not one worker.
|
|
192
|
+
|
|
193
|
+
```bash
|
|
194
|
+
# Stops the hub and every worker managed from this workspace
|
|
195
|
+
kxm hub stop
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
## Troubleshooting
|
|
199
|
+
|
|
200
|
+
| Symptom | Cause | Fix |
|
|
201
|
+
|---|---|---|
|
|
202
|
+
| `KXM worker PID claim is stale at <path>` | A previous supervisor was killed without cleanup. | Run `kxm session status --json`; a claim with `"live": false` is stale. Confirm no matching worker is running, delete that file, and start again. |
|
|
203
|
+
| `KXM worker <project>/<name> is already managed by PID <pid>` | A supervisor for this agent is already running. | Use the running worker, or stop it first. |
|
|
204
|
+
| The worker exits with `worker requires --name and --project` | Name or project missing. | Pass `--name` and `--project`, or set `KXM_AGENT_NAME` and `KXM_PROJECT`. |
|
|
205
|
+
| `KXM_WORKER_FALLBACK_MODELS requires KXM_WORKER_MODEL` | Fallbacks without a primary model. | Add `--model`. |
|
|
206
|
+
| The worker exits 1 with `pi_native_impersonation_blocked: …` | `--model` or a fallback is a model whose vendor has its own harness. | Run that model in its native harness, or use an admitted Pi route such as `openrouter/qwen/qwen3-coder-plus`. |
|
|
207
|
+
| `kxm connection failed` in the Pi log, or Pi shows `hub:off` | Wrong URL, token, or project, or the name is live elsewhere (`duplicate_agent_name`). | Compare `KXM_SERVER_URL`, `KXM_AUTH_TOKEN`, and `KXM_PROJECT` with the hub; check `kxm peer list`. |
|
|
208
|
+
| Workflow runs share one conversation | Isolation is `off`, the default. | Restart with `--session-isolation workflow`. |
|
|
209
|
+
| A request stays `delivered` | A turn, tool, or provider call is still running, or the watchdogs are recovering it. | Read the lifecycle log. Do not send a duplicate. |
|
|
210
|
+
|
|
211
|
+
Never repair a live route by editing files in `.kxm/state`. Stop the worker first, keep the evidence, and recover from the hub's workflow state. More symptoms are in [Troubleshoot KXM](../operations/troubleshooting.md).
|
|
212
|
+
|
|
213
|
+
## Next steps
|
|
214
|
+
|
|
215
|
+
- Send work to your worker and read its replies: [Message peer agents](peer-messaging.md)
|
|
216
|
+
- Make a worker the coordinator of a signed webhook workflow: [Run webhook workflows](webhook-workflows.md)
|
|
217
|
+
- Require replies from named reviewer workers before a stage passes: [Peer provenance and quorum gates](provenance-gates.md)
|
|
218
|
+
- Pick allowed Pi routes: [Harness routing](../reference/harness-routing.md)
|
|
219
|
+
- Every flag: [CLI reference](../reference/cli-reference.md#kxm-agent-worker)
|