@kontextmind/kxm 0.7.94 → 0.7.96
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +1 -1
- package/.kxm/README.md +39 -9
- package/CHANGELOG.md +1 -1
- package/README.md +147 -257
- package/SECURITY.md +21 -12
- package/docs/README.md +133 -54
- package/docs/adr/ADR-0002-browser-automation-steel-doks.md +24 -18
- package/docs/adr/ADR-0003-sqlite-only-store.md +100 -0
- package/docs/adr/ADR-0004-edge-identity-authentik.md +99 -0
- package/docs/adr/README.md +33 -0
- package/docs/concepts/architecture.md +262 -0
- package/docs/concepts/data-and-storage.md +194 -0
- package/docs/concepts/trust-model.md +152 -0
- package/docs/contracts/README.md +22 -14
- package/docs/contracts/effects-and-recovery.md +3 -0
- package/docs/contracts/migration.md +2 -2
- package/docs/contracts/routing.md +6 -5
- package/docs/contributing/assignment-runner.md +388 -0
- package/docs/contributing/ci-and-release.md +231 -0
- package/docs/contributing/development.md +362 -0
- package/docs/contributing/harness-routing-internals.md +192 -0
- package/docs/{packages.md → contributing/packages.md} +13 -15
- package/docs/{skills → contributing}/repo-work-delivery.md +20 -21
- package/docs/contributing/test-matrix.md +208 -0
- package/docs/{tui-components.md → contributing/tui-components.md} +30 -22
- package/docs/contributing/writing-docs.md +340 -0
- package/docs/glossary.md +471 -0
- package/docs/guides/agent-skills.md +137 -0
- package/docs/guides/browser-automation.md +160 -0
- package/docs/guides/context-and-memory.md +352 -0
- package/docs/guides/continuous-improvement.md +228 -0
- package/docs/guides/governed-skills.md +173 -0
- package/docs/guides/nous-providers.md +186 -0
- package/docs/guides/peer-messaging.md +304 -0
- package/docs/guides/pi-workers.md +219 -0
- package/docs/guides/provenance-gates.md +313 -0
- package/docs/guides/webhook-workflows.md +364 -0
- package/docs/kb/how-credentials-retrieved-safely.md +38 -12
- package/docs/kb/how-to-capture-and-annotate-section.md +15 -13
- package/docs/kb/how-to-connect-playwright-to-steel.md +16 -11
- package/docs/kb/how-to-recover-expired-session-or-orphan.md +26 -16
- package/docs/kb/how-to-resume-after-mfa.md +19 -11
- package/docs/kb/how-to-take-over-session.md +17 -13
- package/docs/kb/why-authentication-disappeared.md +22 -14
- package/docs/kb/why-automation-opened-different-browser.md +23 -14
- package/docs/kb/why-session-viewer-cannot-control.md +13 -12
- package/docs/operations/backup-and-restore.md +248 -0
- package/docs/operations/deploy.md +307 -0
- package/docs/operations/monitoring.md +209 -0
- package/docs/operations/runtime-sync.md +192 -0
- package/docs/operations/troubleshooting.md +265 -0
- package/docs/operations/upgrade.md +124 -0
- package/docs/prompts/browser-annotate-feedback.md +7 -7
- package/docs/prompts/browser-diagnose-recover.md +11 -10
- package/docs/prompts/browser-explore.md +7 -7
- package/docs/prompts/browser-repro-fix.md +7 -7
- package/docs/prompts/browser-start.md +12 -11
- package/docs/prompts/browser-takeover.md +8 -8
- package/docs/{cli-reference.md → reference/cli-reference.md} +83 -41
- package/docs/{config-reference.md → reference/config-reference.md} +159 -148
- package/docs/reference/configuration.md +299 -0
- package/docs/reference/harness-routing.md +508 -0
- package/docs/reference/http-api.md +203 -0
- package/docs/reference/tools.md +370 -0
- package/docs/{workflow-guide.md → reference/workflow-catalog.md} +92 -153
- package/docs/reference/workflow-definitions.md +286 -0
- package/docs/start/first-workflow.md +287 -0
- package/docs/start/install.md +146 -0
- package/docs/start/quickstart-claude-code.md +405 -0
- package/docs/start/quickstart-pi.md +213 -0
- package/docs/templates/README.md +78 -73
- package/docs/templates/adr.md +13 -13
- package/docs/templates/architecture.md +55 -71
- package/docs/templates/bug-fix.md +13 -16
- package/docs/templates/feature.md +14 -19
- package/docs/templates/handoff.md +44 -46
- package/docs/templates/postmortem.md +30 -43
- package/docs/templates/research.md +15 -20
- package/docs/templates/review.md +49 -50
- package/docs/templates/runbook.md +38 -30
- package/docs/templates/test-plan.md +16 -23
- package/docs/templates/test-report.md +14 -17
- package/examples/README.md +9 -5
- package/examples/provenance-workflow.json +1 -1
- package/examples/webhook-workflows/jira-development.json +59 -0
- package/examples/webhook-workflows/jira-issue-updated.json +12 -0
- package/package.json +2 -2
- package/packages/core/tui/README.md +1 -1
- package/plugins/kxm/.claude-plugin/plugin.json +1 -1
- package/plugins/kxm/README.md +31 -32
- package/plugins/kxm/dist/cli.js +5 -5
- package/plugins/kxm/dist/mcp-server.js +1 -1
- package/plugins/kxm/dist/runtime.js +1 -1
- package/plugins/kxm/package.json +1 -1
- package/plugins/kxm/skills/kxm/references/protocol.md +3 -1
- package/plugins/kxm/skills/kxm-browser-auth/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-browser-diagnostics/SKILL.md +5 -5
- package/plugins/kxm/skills/kxm-browser-explore/SKILL.md +2 -2
- package/plugins/kxm/skills/kxm-browser-session/SKILL.md +10 -13
- package/plugins/kxm/skills/kxm-browser-takeover/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-browser-verify/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-context-memory/SKILL.md +13 -4
- package/plugins/kxm/skills/kxm-hub-ops/SKILL.md +3 -1
- package/plugins/kxm/skills/kxm-mind-setup/SKILL.md +2 -1
- package/plugins/kxm/skills/kxm-project-setup/SKILL.md +31 -54
- package/plugins/kxm/skills/kxm-projects/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-protocol/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-routing-improve/SKILL.md +15 -7
- package/plugins/kxm/skills/kxm-runs/SKILL.md +11 -5
- package/plugins/kxm/skills/kxm-session/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-tasks/SKILL.md +9 -7
- package/plugins/kxm/skills/kxm-workflow/SKILL.md +10 -2
- package/plugins/kxm/src/cli/system.ts +1 -1
- package/plugins/kxm/src/cli.ts +3 -3
- package/plugins/kxm/src/init-guide-setup.ts +1 -1
- package/plugins/kxm/src/mcp-server.ts +1 -1
- package/plugins/kxm/src/modes.ts +1 -1
- package/schemas/README.md +1 -1
- package/docs/agent-communication-envelopes-and-gates.md +0 -553
- package/docs/agent-skills.md +0 -198
- package/docs/architecture.md +0 -245
- package/docs/assignment-runner.md +0 -264
- package/docs/browser-automation.md +0 -139
- package/docs/configuration.md +0 -437
- package/docs/continuous-improvement.md +0 -226
- package/docs/getting-started.md +0 -277
- package/docs/harness-routing.md +0 -616
- package/docs/kb/qa-authentik-authentication.md +0 -97
- package/docs/kb/qa-extension-install-and-hub-bootstrap.md +0 -85
- package/docs/kb/qa-hub-on-a-public-host.md +0 -48
- package/docs/kb/qa-sqlite-vs-duckdb.md +0 -35
- package/docs/kb/qa-what-the-hub-stores.md +0 -64
- package/docs/kxm-handbook.md +0 -1181
- package/docs/operations.md +0 -510
- package/docs/operator-pi-packages.md +0 -67
- package/docs/provenance-gates.md +0 -295
- package/docs/skills.md +0 -47
- package/docs/test-matrix.md +0 -132
- package/docs/troubleshooting.md +0 -293
- package/docs/webhook-workflows.md +0 -240
package/docs/provenance-gates.md
DELETED
|
@@ -1,295 +0,0 @@
|
|
|
1
|
-
# Peer provenance and quorum gates
|
|
2
|
-
|
|
3
|
-
Peer provenance gates let a workflow require replies from a named, snapshotted
|
|
4
|
-
set of agents before a stage can pass. The hub—not the coordinator—derives the
|
|
5
|
-
producer identity, reply status, workflow scope, and content hashes from the
|
|
6
|
-
durable message record.
|
|
7
|
-
|
|
8
|
-
Use this feature when a gate means “two eligible reviewers replied for this
|
|
9
|
-
exact review attempt.” Do not describe it as proof that the reviews are true,
|
|
10
|
-
independent, high quality, or free from collusion.
|
|
11
|
-
|
|
12
|
-
## What is bound and verified
|
|
13
|
-
|
|
14
|
-
At workflow start, every `eligibleAgents` selector is resolved to a durable
|
|
15
|
-
agent ID and name. The resolved set is copied into the run. The start fails
|
|
16
|
-
closed if an agent is unknown, the coordinator is selected, or fewer unique
|
|
17
|
-
agents resolve than `minProducers` requires.
|
|
18
|
-
|
|
19
|
-
For a peer request to count, the assigned coordinator must send it to an
|
|
20
|
-
eligible producer with an exact `workflowContext`:
|
|
21
|
-
|
|
22
|
-
```json
|
|
23
|
-
{
|
|
24
|
-
"runId": "run_123",
|
|
25
|
-
"stageId": "review",
|
|
26
|
-
"requirementKey": "independent peer reviews",
|
|
27
|
-
"attempt": 1
|
|
28
|
-
}
|
|
29
|
-
```
|
|
30
|
-
|
|
31
|
-
The hub authorizes the context against the current run, active stage, canonical
|
|
32
|
-
requirement key, 1-based attempt, coordinator identity, project, and target.
|
|
33
|
-
It stores the canonical `pi-mesh.workflow-message-context.v1` context and binds
|
|
34
|
-
the message correlation ID to the run. A caller cannot turn an ordinary or old
|
|
35
|
-
message into workflow evidence by choosing an idempotency prefix or correlation
|
|
36
|
-
ID.
|
|
37
|
-
|
|
38
|
-
At checkpoint or wait, the coordinator cites message IDs rather than describing
|
|
39
|
-
the replies itself:
|
|
40
|
-
|
|
41
|
-
```json
|
|
42
|
-
{
|
|
43
|
-
"evidenceRefs": {
|
|
44
|
-
"independent peer reviews": {
|
|
45
|
-
"messageIds": ["msg_reviewer_a", "msg_reviewer_b"]
|
|
46
|
-
}
|
|
47
|
-
}
|
|
48
|
-
}
|
|
49
|
-
```
|
|
50
|
-
|
|
51
|
-
The hub accepts only durable `replied` messages with non-empty replies, coherent
|
|
52
|
-
timestamps, the exact immutable context, the same project and coordinator, an
|
|
53
|
-
eligible producer, and the run correlation ID. Quorum counts unique stable
|
|
54
|
-
producer IDs, not messages. Duplicate replies from one producer still count as
|
|
55
|
-
one producer. Missing, pending, cancelled, expired, wrong-attempt, cross-run,
|
|
56
|
-
cross-stage, cross-project, coordinator-authored, and ineligible messages fail
|
|
57
|
-
closed.
|
|
58
|
-
|
|
59
|
-
After verification, the run stores a metadata-only snapshot containing:
|
|
60
|
-
|
|
61
|
-
- message ID and stable producer ID/name;
|
|
62
|
-
- exact run, stage, requirement, and attempt context;
|
|
63
|
-
- `replied` status and lifecycle timestamps;
|
|
64
|
-
- SHA-256 hashes of the request and reply;
|
|
65
|
-
- the verification timestamp.
|
|
66
|
-
|
|
67
|
-
The snapshot contains no request or reply body. It remains with the workflow
|
|
68
|
-
run after the source message reaches the terminal-message retention limit and
|
|
69
|
-
is purged. The hashes show which content the hub observed during verification;
|
|
70
|
-
they do not reveal the content or prove its correctness.
|
|
71
|
-
|
|
72
|
-
Retrospective quorum summaries are attempt-specific. For a passed or failed
|
|
73
|
-
stage they report the terminal attempt; for an active stage they report the
|
|
74
|
-
current attempt. Stored references from an earlier retry never inflate the
|
|
75
|
-
applied producer count.
|
|
76
|
-
|
|
77
|
-
## Configure a policy
|
|
78
|
-
|
|
79
|
-
`evidencePolicies` keys must match canonical entries in `requiredEvidence`.
|
|
80
|
-
Only the `peer-reply` policy and `replied` status are currently supported.
|
|
81
|
-
|
|
82
|
-
| Field | Constraint |
|
|
83
|
-
|---|---|
|
|
84
|
-
| `kind` | Exact value `peer-reply` |
|
|
85
|
-
| `minProducers` | Integer from 1 through 8 and no greater than the selector count |
|
|
86
|
-
| `eligibleAgents` | 1–16 non-empty, case-insensitively unique names or IDs |
|
|
87
|
-
| `acceptedStatuses` | Exact array `["replied"]`; omission uses the same value |
|
|
88
|
-
| `degradation.minProducers` | Optional integer at least 1 and lower than `minProducers` |
|
|
89
|
-
|
|
90
|
-
```json
|
|
91
|
-
{
|
|
92
|
-
"id": "review",
|
|
93
|
-
"label": "Independent review",
|
|
94
|
-
"instructions": "Collect independent reviews, resolve contradictions, and record the decision.",
|
|
95
|
-
"requiredEvidence": [
|
|
96
|
-
"independent peer reviews",
|
|
97
|
-
"coordinator decision"
|
|
98
|
-
],
|
|
99
|
-
"evidencePolicies": {
|
|
100
|
-
"independent peer reviews": {
|
|
101
|
-
"kind": "peer-reply",
|
|
102
|
-
"minProducers": 2,
|
|
103
|
-
"eligibleAgents": ["reviewer-claude", "reviewer-grok"],
|
|
104
|
-
"acceptedStatuses": ["replied"],
|
|
105
|
-
"degradation": {
|
|
106
|
-
"minProducers": 1
|
|
107
|
-
}
|
|
108
|
-
}
|
|
109
|
-
},
|
|
110
|
-
"maxAttempts": 3,
|
|
111
|
-
"area": "gates"
|
|
112
|
-
}
|
|
113
|
-
```
|
|
114
|
-
|
|
115
|
-
Register the coordinator and every eligible peer at least once before starting
|
|
116
|
-
the workflow. Names are resolved once at run creation; later configuration edits
|
|
117
|
-
do not rewrite an active run's eligible-producer snapshot.
|
|
118
|
-
|
|
119
|
-
The complete command-first example is
|
|
120
|
-
[`examples/provenance-workflow.json`](../examples/provenance-workflow.json).
|
|
121
|
-
It uses project `provenance-demo`, coordinator `coordinator`, and the two
|
|
122
|
-
eligible reviewers `reviewer-claude` and `reviewer-grok`.
|
|
123
|
-
|
|
124
|
-
## Execute a strict quorum
|
|
125
|
-
|
|
126
|
-
Start with separate project and administrative credentials:
|
|
127
|
-
|
|
128
|
-
```powershell
|
|
129
|
-
$env:KXM_AUTH_TOKEN = "replace-with-the-admin-token"
|
|
130
|
-
$env:KXM_PROJECT_TOKENS = '{"provenance-demo":"replace-with-the-project-token"}'
|
|
131
|
-
$env:KXM_WEBHOOK_WORKFLOWS_FILE = "examples/provenance-workflow.json"
|
|
132
|
-
kxm hub start
|
|
133
|
-
```
|
|
134
|
-
|
|
135
|
-
Give the coordinator and peer workers only the project token. Start all three
|
|
136
|
-
before starting the workflow:
|
|
137
|
-
|
|
138
|
-
```powershell
|
|
139
|
-
$env:KXM_AUTH_TOKEN = "replace-with-the-project-token"
|
|
140
|
-
kxm agent worker --name coordinator --project provenance-demo --session-isolation workflow
|
|
141
|
-
kxm agent worker --name reviewer-claude --project provenance-demo --model openrouter/qwen/qwen3-coder-plus --session-isolation workflow
|
|
142
|
-
kxm agent worker --name reviewer-grok --project provenance-demo --model openrouter/z-ai/glm-5.3-flash --session-isolation workflow
|
|
143
|
-
```
|
|
144
|
-
|
|
145
|
-
The reviewer ids are names, not routes. Workers are Pi, and Pi may not run a model whose
|
|
146
|
-
vendor has its own harness, so these two independent reviewers use admitted Pi routes from
|
|
147
|
-
different vendors rather than Claude and Grok.
|
|
148
|
-
|
|
149
|
-
In an operator terminal, supply the workflow-start secret and create a run:
|
|
150
|
-
|
|
151
|
-
```powershell
|
|
152
|
-
$env:KXM_WORKFLOW_SECRET = "replace-with-the-workflow-start-secret"
|
|
153
|
-
kxm gate validate --file examples/provenance-workflow.json
|
|
154
|
-
kxm workflow start provenance-review --payload '{"task":{"id":"DEMO-1","summary":"Review the proposed change"}}'
|
|
155
|
-
kxm workflow list
|
|
156
|
-
```
|
|
157
|
-
|
|
158
|
-
The coordinator gets the run ID in its durable prompt. It should read the run,
|
|
159
|
-
calculate the current attempt as `stage.attempts + 1`, and fan out with one
|
|
160
|
-
shared context:
|
|
161
|
-
|
|
162
|
-
```json
|
|
163
|
-
{
|
|
164
|
-
"targets": ["reviewer-claude", "reviewer-grok"],
|
|
165
|
-
"content": "Review this change independently. Return findings with evidence.",
|
|
166
|
-
"idempotencyKeyPrefix": "review-attempt-1",
|
|
167
|
-
"workflowContext": {
|
|
168
|
-
"runId": "run_123",
|
|
169
|
-
"stageId": "review",
|
|
170
|
-
"requirementKey": "independent peer reviews",
|
|
171
|
-
"attempt": 1
|
|
172
|
-
}
|
|
173
|
-
}
|
|
174
|
-
```
|
|
175
|
-
|
|
176
|
-
`idempotencyKeyPrefix` makes an exact transport retry safe; it does not create
|
|
177
|
-
provenance. After both results are `replied`, checkpoint with their returned
|
|
178
|
-
message IDs and ordinary evidence for the non-peer requirement:
|
|
179
|
-
|
|
180
|
-
```json
|
|
181
|
-
{
|
|
182
|
-
"runId": "run_123",
|
|
183
|
-
"stageId": "review",
|
|
184
|
-
"status": "passed",
|
|
185
|
-
"summary": "Compared both reviews and resolved the material contradiction.",
|
|
186
|
-
"evidence": {
|
|
187
|
-
"coordinator decision": "decision:docs/review-decision.md"
|
|
188
|
-
},
|
|
189
|
-
"evidenceRefs": {
|
|
190
|
-
"independent peer reviews": {
|
|
191
|
-
"messageIds": ["msg_reviewer_a", "msg_reviewer_b"]
|
|
192
|
-
}
|
|
193
|
-
}
|
|
194
|
-
}
|
|
195
|
-
```
|
|
196
|
-
|
|
197
|
-
Caller-authored `evidence` strings never satisfy a requirement that declares a
|
|
198
|
-
peer policy. `warning` and `failed` checkpoints cannot submit `evidenceRefs`.
|
|
199
|
-
After either result consumes an attempt, send fresh peer requests with the new
|
|
200
|
-
attempt number; old references cannot satisfy the retry.
|
|
201
|
-
|
|
202
|
-
## Degrade only through an explicit admin decision
|
|
203
|
-
|
|
204
|
-
Degradation is optional and must be declared in the policy before the run
|
|
205
|
-
starts. A peer, coordinator, webhook, and signed callback cannot approve it.
|
|
206
|
-
Only the administrative bearer token can approve the configured lower minimum,
|
|
207
|
-
and only for the current run, active stage, exact requirement, and current
|
|
208
|
-
attempt.
|
|
209
|
-
|
|
210
|
-
Inspect the requested action, then approve it with a non-secret reason:
|
|
211
|
-
|
|
212
|
-
```powershell
|
|
213
|
-
$env:KXM_AUTH_TOKEN = "replace-with-the-admin-token"
|
|
214
|
-
kxm gate --dry-run --json degrade run_123 review `
|
|
215
|
-
--requirement "independent peer reviews" `
|
|
216
|
-
--reason "reviewer-grok provider outage incident-482"
|
|
217
|
-
kxm gate degrade run_123 review `
|
|
218
|
-
--requirement "independent peer reviews" `
|
|
219
|
-
--reason "reviewer-grok provider outage incident-482"
|
|
220
|
-
```
|
|
221
|
-
|
|
222
|
-
Approval does not pass the stage. The coordinator must still submit enough
|
|
223
|
-
verified replies to meet the approved minimum. Repeating the same approval and
|
|
224
|
-
reason is idempotent; changing the reason conflicts. The approval cannot be
|
|
225
|
-
reused by a later attempt.
|
|
226
|
-
|
|
227
|
-
A passed degraded stage is marked as degraded. Its retrospective contains the
|
|
228
|
-
configured and effective minima, eligible-producer snapshot, verified message
|
|
229
|
-
metadata, hashes, timestamps, and the explicit admin approval. Export it with:
|
|
230
|
-
|
|
231
|
-
```powershell
|
|
232
|
-
kxm workflow export run_123
|
|
233
|
-
```
|
|
234
|
-
|
|
235
|
-
The JSON keeps the existing `pi-mesh.retrospective.v1` schema and adds optional
|
|
236
|
-
`evidenceAudit` and `degradedStageIds` fields. Consumers that do not know these
|
|
237
|
-
fields can continue reading the v1 document.
|
|
238
|
-
|
|
239
|
-
## Trust boundary
|
|
240
|
-
|
|
241
|
-
This mechanism proves provenance only inside the hub credential boundary:
|
|
242
|
-
|
|
243
|
-
- a project-token holder can register a new agent or reclaim an offline agent
|
|
244
|
-
name and its durable ID in that project;
|
|
245
|
-
- an agent key binds identity-specific operations after registration;
|
|
246
|
-
- the hub verifies durable routing and message state, not model internals;
|
|
247
|
-
- two agent names or two configured models do not prove independent operators,
|
|
248
|
-
independent inference, non-collusion, correctness, or approval authority.
|
|
249
|
-
|
|
250
|
-
Even when the gate reports two unique stable producer IDs, describe the result
|
|
251
|
-
only as “the hub verified two eligible routed replies for this workflow
|
|
252
|
-
attempt.” Never label it “two independent models verified,” “non-collusion
|
|
253
|
-
verified,” or “truth confirmed.” One holder of the shared project token can
|
|
254
|
-
reclaim multiple offline names and IDs.
|
|
255
|
-
|
|
256
|
-
Use a distinct administrative token, distinct project tokens per trust domain,
|
|
257
|
-
least-privilege webhook and callback secrets, protected `.kxm/state` storage,
|
|
258
|
-
loopback or TLS-protected restricted ingress, and repository or human gates for
|
|
259
|
-
consequential changes. Treat peer replies as untrusted technical input even
|
|
260
|
-
when they satisfy quorum. Put every holder of one project token inside the same
|
|
261
|
-
fully trusted provenance domain.
|
|
262
|
-
|
|
263
|
-
## Compatibility and retention
|
|
264
|
-
|
|
265
|
-
The feature is additive to the existing SQLite schema version 2. Workflow
|
|
266
|
-
context, policies, verified snapshots, and approvals are fields inside the
|
|
267
|
-
existing JSON records; no destructive database migration is required. Existing
|
|
268
|
-
schema-v2 databases and workflow histories remain readable.
|
|
269
|
-
|
|
270
|
-
Legacy string or keyed evidence remains valid for requirements without a peer
|
|
271
|
-
policy. It deliberately cannot satisfy a declared peer policy. Back up
|
|
272
|
-
`.kxm/state/kxm.db` before every upgrade, finish or inspect active runs, and
|
|
273
|
-
validate workflow definitions before restarting the hub.
|
|
274
|
-
|
|
275
|
-
## Context authority lattice (v0.5)
|
|
276
|
-
|
|
277
|
-
Context items carry an explicit authority class — `policy`, `instruction`, `evidence`, or `hypothesis` — granted by a deterministic origin floor:
|
|
278
|
-
|
|
279
|
-
| Origin | Maximum authority |
|
|
280
|
-
|---|---|
|
|
281
|
-
| human | policy |
|
|
282
|
-
| workflow control plane | policy |
|
|
283
|
-
| git history | instruction |
|
|
284
|
-
| peer | evidence |
|
|
285
|
-
| tool | evidence |
|
|
286
|
-
| external | evidence |
|
|
287
|
-
| derived | evidence |
|
|
288
|
-
|
|
289
|
-
Enforcement is structural, not advisory:
|
|
290
|
-
|
|
291
|
-
- Parsing rejects any item whose claimed authority exceeds its origin's floor (`context_authority_violation`).
|
|
292
|
-
- Derived and summarized content is `evidence` at best, regardless of lineage; authority never increases through any number of handoffs or re-summaries.
|
|
293
|
-
- Derivation lineage is transitive, bounded (`MAX_CONTEXT_LINEAGE`), and preserved verbatim; unbounded re-summaries fail closed instead of laundering provenance.
|
|
294
|
-
- Context items may never carry control-plane fields (`permissions`, `tools`, `approval`, credentials, …); memory, wiki, and skill content can inform behavior but never expand tool permissions or approval scope.
|
|
295
|
-
- Superseded and rejected records are excluded from summarization lineage: dead records are not evidence of current truth.
|
package/docs/skills.md
DELETED
|
@@ -1,47 +0,0 @@
|
|
|
1
|
-
# Skill candidate lifecycle
|
|
2
|
-
|
|
3
|
-
KXM turns verified episodes and lessons into reusable Agent Skills through a governed lifecycle. Runtime experience never becomes promoted skill content automatically, and promoted skills never grant tool or permission authority.
|
|
4
|
-
|
|
5
|
-
This page is the governed `kxm skills` lifecycle. For the bundled command-suite skills, see [Agent Skills](agent-skills.md). For converting a repository request into a delivery prompt, see [Repository work delivery](skills/repo-work-delivery.md).
|
|
6
|
-
|
|
7
|
-
## Lifecycle
|
|
8
|
-
|
|
9
|
-
```text
|
|
10
|
-
verified episode(s) → skill candidate → static/provenance review
|
|
11
|
-
→ sandbox execution → protected functional + safety eval
|
|
12
|
-
→ promote / quarantine / reject
|
|
13
|
-
```
|
|
14
|
-
|
|
15
|
-
## Storage
|
|
16
|
-
|
|
17
|
-
```text
|
|
18
|
-
.kxm/skills/
|
|
19
|
-
├── candidates/<id>/SKILL.md + metadata.json
|
|
20
|
-
├── promoted/<id>/SKILL.md + metadata.json
|
|
21
|
-
├── quarantined/<id>/SKILL.md + metadata.json
|
|
22
|
-
└── history/<id>.jsonl
|
|
23
|
-
```
|
|
24
|
-
|
|
25
|
-
Skill IDs are content-addressed (`<slug>.<hash-prefix>`): changed behavior means changed content means a new candidate. Identical resubmissions are rejected as duplicates.
|
|
26
|
-
|
|
27
|
-
## Rules
|
|
28
|
-
|
|
29
|
-
- A candidate must cite at least one source run, journal entry, or evidence receipt.
|
|
30
|
-
- Candidate metadata records explicit cross-model/cross-harness compatibility (`harness`, `models`).
|
|
31
|
-
- Promotion requires passing `static-review`, `sandbox`, `functional`, and `safety` evaluations, durable evidence references, and a decision by someone other than the author.
|
|
32
|
-
- A failed functional or safety evaluation quarantines the candidate automatically; the evaluator records the decision and a human may later reject fully.
|
|
33
|
-
- Rejected and quarantined candidates remain queryable in `history/` for future learning.
|
|
34
|
-
- Promoted skills are hash-pinned and immutable: `kxm skills verify` detects out-of-band edits (`skill_integrity_violation`), and behavior changes require a new candidate/eval cycle (optionally `supersedes`-linked).
|
|
35
|
-
- Skill content is redacted of secret material at creation; evaluation details are redacted and bounded.
|
|
36
|
-
- Optimization evaluations (skillopt/WikiSkill-style) are a gated hook, disabled unless explicitly enabled.
|
|
37
|
-
|
|
38
|
-
## CLI
|
|
39
|
-
|
|
40
|
-
```text
|
|
41
|
-
kxm skills create --file SKILL.md --name <name> --created-by <id> --harness pi --models <models> --run <ids> --journal <ids>
|
|
42
|
-
kxm skills evaluate <id> --kind static-review|sandbox|functional|safety --evaluator <version> [--fail]
|
|
43
|
-
kxm skills promote <id> --decided-by <id> --evidence <refs>
|
|
44
|
-
kxm skills reject <id> --decided-by <id>
|
|
45
|
-
kxm skills list --state candidate|promoted|quarantined|rejected
|
|
46
|
-
kxm skills verify <id> --state promoted
|
|
47
|
-
```
|
package/docs/test-matrix.md
DELETED
|
@@ -1,132 +0,0 @@
|
|
|
1
|
-
# Test matrix
|
|
2
|
-
|
|
3
|
-
The release gate executes every test, measures the core source directly, type-checks strict TypeScript, lints documentation, verifies package versions, validates Claude manifests, rebuilds the generated runtimes, and installs and executes the npm artifact outside the repository.
|
|
4
|
-
|
|
5
|
-
Run the commit gate with `npm run verify`. CI PR legs run `validate:pr` plus
|
|
6
|
-
`check:generated`; pushes to main run `validate:ci`. Plugin validation is a hosted CI
|
|
7
|
-
job.
|
|
8
|
-
|
|
9
|
-
```powershell
|
|
10
|
-
npm run verify
|
|
11
|
-
```
|
|
12
|
-
|
|
13
|
-
## Product features
|
|
14
|
-
|
|
15
|
-
| Feature | Automated evidence |
|
|
16
|
-
|---|---|
|
|
17
|
-
| Portal tenant read composes two labelled authorities: hub runs carry `hub-projection`, Runtime runs carry `runtime-authoritative`; a down hub leaves runtime state readable and vice versa with stable reasons; projection/authoritative disagreements surface as `discrepancies`; binding scope is labelled loopback/remote | `test/core/studio-layout.test.ts` |
|
|
18
|
-
| Structured-result settlement only: outcome words in prose, empty replies, and outcomes outside the step's declared set all fail closed; a declared JSON object and a JSON result block inside prose both settle, and the engine records `outcome_unknown` and terminates `failed` | `test/core/pi-producer.test.ts`, `test/core/engine.test.ts` |
|
|
19
|
-
| Health, readiness, metrics, request IDs, security headers | `test/core/hub-api.test.ts` |
|
|
20
|
-
| Shared and per-project authentication, project isolation | `test/core/hub-api.test.ts` |
|
|
21
|
-
| Registration, discovery, presence, stale detection, identity resumption | `test/core/hub-api.test.ts`, `test/core/hub.test.ts` |
|
|
22
|
-
| SQLite persistence, restart recovery, schema compatibility | `test/core/hub-api.test.ts`, `test/core/store.test.ts` |
|
|
23
|
-
| All delivery modes, message fields, hop limits, and validation | `test/core/hub-api.test.ts`, `test/core/protocol.test.ts` |
|
|
24
|
-
| Queue, acknowledgement, visibility, reply, and authorization | `test/core/hub-api.test.ts`, `test/core/hub.test.ts` |
|
|
25
|
-
| Queued/delivered replay after recipient restart reuses one message record | `test/core/hub.test.ts`, `test/core/extension.test.ts`, `test/core/mcp.test.ts` |
|
|
26
|
-
| Known-offline peer send with `allowOffline` queues, delivers once on resumption, and expires unread by TTL | `test/core/hub-api.test.ts` |
|
|
27
|
-
| Fenced hub leases: CAS acquire/renew/release, monotonic token on takeover only, hub-clocked expiry, and a shared external effect refused at commit under a superseded token | `test/core/hub-api.test.ts`, `test/core/store.test.ts`, `test/core/external-effects.test.ts` |
|
|
28
|
-
| One-to-three-peer fanout, recoverable local timeouts/aborts, exact retries, and partial-error collection | `test/core/client.test.ts`, `test/core/hub-api.test.ts`, `test/core/extension.test.ts`, `test/core/mcp.test.ts` |
|
|
29
|
-
| TTL expiry, sender cancellation, and terminal retention | `test/core/hub-api.test.ts` |
|
|
30
|
-
| Terminal inbound cleanup and next-request activation | `test/core/extension.test.ts`, `test/core/mcp.test.ts` |
|
|
31
|
-
| Exact-retry idempotency and conflicting-key rejection | `test/core/hub-api.test.ts` |
|
|
32
|
-
| Rate limiting and retry guidance | `test/core/hub-api.test.ts` |
|
|
33
|
-
| Redacted structured logs | `test/core/hub-api.test.ts` |
|
|
34
|
-
| Client lifecycle, aborts, timeouts, invalid responses, reconnection | `test/core/client.test.ts` |
|
|
35
|
-
| Pi tools, inbound turns, automatic replies, status command | `test/core/extension.test.ts` |
|
|
36
|
-
| Claude MCP catalog, outbound and inbound tools, channel delivery | `test/core/mcp.test.ts` |
|
|
37
|
-
| Responsive metadata-only TUI, authenticated ops mode, presence-only fallback, observer filtering, key controls, and local body-free projection | `test/core/tui.test.ts`, `test/core/hub-api.test.ts` |
|
|
38
|
-
| Session manifest creation, fail-closed rosters, shared worker/result envelopes, and hub-owned envelope fields | `test/core/session.test.ts`, `test/core/cli.test.ts`, `test/core/envelope.test.ts`, `test/core/envelope-contract.test.ts` |
|
|
39
|
-
| Generic CLI/project telemetry classification, JSONL recovery, and proposed `kxm improve` output | `test/core/telemetry.test.ts`, `test/core/cli.test.ts`, `test/core/improve.test.ts` |
|
|
40
|
-
| `kxm improve` and `kxm routing report` read the project's Runtime event store read-only plus telemetry, resolve each attempt's outcome from the event log, drop simulated attempts and duplicate attempts, flag only same-ask repeats across runs, exclude write steps, and report promotion readiness that never authorizes | `test/core/improve.test.ts` (`kxm improve report resolves Runtime-settled attempts from the event log and flags only same-ask cross-run repeats`), `test/core/cli.test.ts`, `test/core/cli-experience.test.ts`, `test/core/commands-policy.test.ts` |
|
|
41
|
-
| Engine routing records carry the engine-reserved `workflowId`, `askSha256`, `objectiveSha256` and `stepWrites` keys (stable across runs), default `agentRole` to the agent, and record only `blocked` or `failed` at settlement | `test/core/route-admission.test.ts` |
|
|
42
|
-
| Runtime dispatch context: only committed, pinned project memory and hash-verified promoted skills reach a dispatched agent; malformed, drifted or uncommitted content is withheld with a `dispatch_context_*` gap outside the prompt, and the step still completes | `test/core/engine.test.ts` (`dispatch context: agents receive only committed, pinned memory and verified skills; anything else is withheld with a gap and the step still completes`) |
|
|
43
|
-
| Signed Jira webhook verification, filtering, dispatch, and retry deduplication | `test/core/hub-api.test.ts` |
|
|
44
|
-
| Ordered workflow checkpoints, normalized keyed evidence gates, unrelated-volume rejection, and warning/failure retry | `test/core/hub-api.test.ts`, `test/core/workflow.test.ts` |
|
|
45
|
-
| Run-start eligible-producer resolution, immutable workflow context, per-requirement message-reference verification, unique-producer quorum, and replay/cross-context rejection | `test/core/workflow-provenance.test.ts`, `test/core/workflow.test.ts`, `test/core/hub-api.test.ts`, `test/core/client.test.ts`, `test/core/store.test.ts` |
|
|
46
|
-
| Explicit current-attempt admin degradation, configured lower minimum, audit journal, idempotency, and forbidden or stale approvals | `test/core/workflow-provenance.test.ts`, `test/core/cli.test.ts` |
|
|
47
|
-
| Durable external waits, local/callback evidence accumulation, safe settlement, checkpoint/expiry race rejection, minimal signed responses, retry/conflict deduplication, separate secrets, and timeout notification | `test/core/hub-api.test.ts`, `test/core/workflow.test.ts`, `test/core/workflow-provenance.test.ts` |
|
|
48
|
-
| Plans, decisions, contradictions, errors, lessons, and improvement reports | `test/core/hub-api.test.ts`, `test/core/workflow.test.ts` |
|
|
49
|
-
| All ten journal categories and `stageId` through the shared `kxm_workflow_record` tool, hub-derived attempt, stage-default area, stage-bound hub-authored entries, ranked redacted cross-run signals, and retrospectives refreshed by late entries and promotions | `test/core/journal-evolution.test.ts` (`kxm_workflow_record binds stage provenance and the stage's area end to end, and hub-authored entries carry it too`), `test/core/workflow.test.ts`, `test/core/retrospective.test.ts`, `test/core/hub-api.test.ts` |
|
|
50
|
-
| Context packets rank by deterministic task relevance, fill the budget first-fit, deliver every selected item (including evidence) in a packet section, order ties newest first, and report numeric `audit.relevance`; recall ranks by phrase then relevance; hub logs carry sizes, not task or query text | `test/core/arbiter.test.ts` (`arbitrate ranks task-relevant candidates first, delivers every selected item in a packet section, orders ties newest first, and reports relevance`), `test/core/context-surfaces.test.ts` |
|
|
51
|
-
| Safe diagnostic classification and redaction | `test/core/diagnostics.test.ts`, `test/core/extension.test.ts`, `test/core/hub-api.test.ts` |
|
|
52
|
-
| Operator CLI init/validate/export/watch | `test/core/cli.test.ts`, `test/core/github-watch.test.ts` |
|
|
53
|
-
| Local and isolated-global packed npm CLI plus hub runtimes | `test/core/package-install.test.ts` |
|
|
54
|
-
| Required generated runtimes are present, tracked, and match the staged copy after build | `scripts/check-generated.mjs`, `test/core/generated-artifacts.test.ts` |
|
|
55
|
-
| Tag release packs `kxm-<v>.tgz`, fail-closed draft GitHub upload, 404-then-list draft discovery, digest proof, no clobber | `scripts/kxm-release-github.mjs`, `test/core/kxm-release-github.test.ts`, `test/core/ci-contract.test.ts` |
|
|
56
|
-
| Retrospective export snapshots, metadata-only provenance audit, body allowlisting, degradation records, and v1 compatibility | `test/core/retrospective.test.ts` |
|
|
57
|
-
| Interrupted-worker continue fallback, exact run-bound recovery, unbound telemetry isolation, and one-turn durable replay | `test/core/worker.test.ts`, `test/core/recovery.test.ts`, `test/core/extension.test.ts` |
|
|
58
|
-
| Hub-owned workflow affinity; integrated hub→extension→supervisor→replacement replay; pre-ack default/run/cross-run routing; one-child session-dir swapping; stable ordinary context; LRU retention; and corrupt-state/link containment | `test/core/hub-api.test.ts`, `test/core/extension.test.ts`, `test/core/worker.test.ts`, `test/core/cli.test.ts` |
|
|
59
|
-
| Final provider-error retention, built-in retry ordering, metadata-only journaling, bounded fallback exhaustion, oversized-frame classification, and session-preserving restart | `test/core/extension.test.ts`, `test/core/worker.test.ts`, `test/core/diagnostics.test.ts`, `test/core/cli.test.ts` |
|
|
60
|
-
| Tool capability allowlist, watchdog grace, bounded hung-tool recovery, oversized completed-tool cancellation, and race-safe hub/worker ownership claims | `test/core/worker.test.ts`, `test/core/cli.test.ts`, `test/core/server.test.ts` |
|
|
61
|
-
| Exact worker extension/skill sets, discovery isolation, path preflight, multi-path ordering, and Windows argument safety | `test/core/worker.test.ts` |
|
|
62
|
-
| Opt-in real-Pi smoke contract and safe skip paths | `test/core/smoke-real-pi.test.ts`, `test/core/smoke.test.ts` |
|
|
63
|
-
| Durable workflow and journal recovery | `test/core/store.test.ts` |
|
|
64
|
-
| Atomic workflow transition commit and rollback | `test/core/store.test.ts` |
|
|
65
|
-
| Pi and Claude workflow/journal tools, workflow-context sends, and peer-reference checkpoints/waits | `test/core/extension.test.ts`, `test/core/mcp.test.ts` |
|
|
66
|
-
| Package and marketplace version consistency | `scripts/check-versions.mjs` |
|
|
67
|
-
| Planned KXM schemas, restricted YAML fixtures, cross-resource semantics, and sync-safe rejection | `test/core/contracts.test.ts` |
|
|
68
|
-
| Production KXM restricted loader, deterministic bundle hashing, Git discovery, fail-closed semantics, init classification, provenance-tracked atomic creation, exact three-way repair, authority-change blocking, pinned crash resumption, shadow validation, explicit join, Runtime-local bindings, CLI isolation, and idempotence | `test/core/project-config.test.ts`, `test/core/cli.test.ts`, `test/core/package-install.test.ts` |
|
|
69
|
-
| Legacy state is refused, never converted: a tree with `.kxm/config` JSON fails project load with `legacy_state_unsupported`, `kxm init` classifies it `mode: "legacy"` without writing, `kxm migrate` is an unknown command, and an older stamped store is refused without advancing `user_version` | `test/core/project-config.test.ts`, `test/core/cli.test.ts`, `test/core/e6-backup-restore-migrations.test.ts` |
|
|
70
|
-
| Permission-diff trust workflow: structured authority projections, conservative lattice classification (access, network, budgets, quorums, snapshots, secrets, transitions, shapes), prose neutrality, Git base shadowing, CLI diff/check gating, and packed-consumer round trips | `test/core/permission.test.ts`, `test/core/cli.test.ts`, `test/core/contracts.test.ts`, `test/core/package-install.test.ts` |
|
|
71
|
-
| Event-sourced local Runtime: supervisor singleton with stable logical identity, immutable home bindings, append-only per-project event stores, idempotent acceptance/cancel, projection rebuild equivalence, token-authenticated local API, auto-start, SIGKILL crash recovery, offline CLI, packed consumer | `test/core/runtime.test.ts`, `test/core/cli.test.ts`, `test/core/package-install.test.ts` |
|
|
72
|
-
|
|
73
|
-
The CI minimums are 93% lines, 80% branches, and 93% functions across
|
|
74
|
-
`plugins/kxm/src/**/*.ts` (excludes `server.ts` and `mcp-server.ts`). Those
|
|
75
|
-
floors may only ratchet up. The generated MCP runtime is exercised as a child
|
|
76
|
-
process, while the packed CLI and hub are installed in a clean consumer and
|
|
77
|
-
exercised from `node_modules`.
|
|
78
|
-
|
|
79
|
-
## Executable examples and use cases
|
|
80
|
-
|
|
81
|
-
| Scenario | Location | Verification |
|
|
82
|
-
|---|---|---|
|
|
83
|
-
| Self-contained planner/reviewer round trip | `examples/roundtrip.ts` | Executed by `test/core/examples.test.ts` |
|
|
84
|
-
| Long-running deterministic reviewer | `examples/reviewer-agent.ts` | Type-checked and documented |
|
|
85
|
-
| Command-line requester | `examples/requester.ts` | Type-checked and documented |
|
|
86
|
-
| Plan then review | `examples/README.md` | Uses discovery, send, and wait |
|
|
87
|
-
| Separate file ownership | `examples/README.md` | Documents non-overlapping writers |
|
|
88
|
-
| Non-blocking delegation | `examples/README.md` | Uses send, independent work, and get |
|
|
89
|
-
| Obsolete-work cancellation | `examples/README.md` | Uses cancel and states rollback boundary |
|
|
90
|
-
| Safe network retry | `examples/README.md` | Uses stable idempotency keys |
|
|
91
|
-
| Jira issue-to-merge workflow | `.kxm/workflows/default.yaml` | Parsed, type-checked through workflow tests, and exercised end to end with representative configuration |
|
|
92
|
-
| `.kxm` workspace defaults and persisted hub/worker logs | `.kxm/`, `test/core/server.test.ts`, `test/core/worker.test.ts` | Executed with isolated temporary workspaces |
|
|
93
|
-
| Signed external result callback | `examples/workflow-signal.ts` | Type-checked; equivalent signed callback path is exercised end to end in `test/core/hub-api.test.ts` |
|
|
94
|
-
| Peer provenance and optional explicit degradation | `examples/provenance-workflow.json`, `.kxm/workflows/default.yaml`, `docs/provenance-gates.md` | Both definitions are parser-checked in `test/core/examples.test.ts`; adversarial evidence and degradation behavior is automated in `test/core/workflow-provenance.test.ts` |
|
|
95
|
-
| Quorum parser boundaries and definition identity | `plugins/kxm/src/workflow.ts`, `test/core/workflow-quorum.test.ts`, `test/core/workflow-definition-hash.test.ts` | Rejects impossible peer pools, verifies degradation bounds, and proves secret-free semantic hash stamping plus credential-rotation invariance |
|
|
96
|
-
| Artifact existence and containment gate | `plugins/kxm/src/artifacts-exist.ts`, `test/core/artifacts-exist.test.ts` | Non-empty regular files pass; missing, empty, non-file, lexical escape, and real-path escape cases fail closed (host-permitted symlink coverage) |
|
|
97
|
-
| Harness inventory probe | `plugins/kxm/src/harness.ts`, `test/core/harness.test.ts` | Detect/auth/dispatch for the builtin catalog; win32 `.exe` / inner npm-package `.exe` / `.cmd` candidate order for npm shims (issue #168); `windows_shim` issue only when the shim answered; allowlisted `name.cmd` shell spawn only; rejected metacharacter commands; Linux still `not_detected` when the bare command is missing. |
|
|
98
|
-
| Headless harness helper | `scripts/harness-run.mjs`, `justfile`, `test/core/harness-run.test.ts` | Offline auth success/logout/garbage, role/mode/pair/provider refusals before spawn, missing brief/schema fail closed with zero spawn, invocation-cwd relative `prompt_file`/`output_schema` vs `request.cwd` (absolute argv tokens; Claude/Codex stdin matches the brief), Pi JSONL multi-`message_end` sums, Claude auxiliary usage, native error-on-exit-0, timeout/empty payload, sidecar-only stderr/answer/error, shell:false argv metacharacters, win32 `.cmd` rejection, and win32 npm inner `claude.exe` / Pi `node.exe`+`cli.js` unwrap. Result v2 transport vs closed model claims, dispatch-before-spawn, grok/codex isolation flags, `max_turns` validation, obsolete v1 diagnosis without rewrite or unknown-schema echo, partial usage on fail/interrupt, signaled null `exitCode` plus exact `signal`, exit-before-stdio-close drain vs bounded linger, type-closed usage/cost (no object leak or zero-coercion), malformed optional text as run-stage failure with retained spend, stdin/pid-record write failures not completed, bounded timeout settle without descendant-death claims, spawn/write-failure stage and spend, and capability-fixture parser evidence (Codex `--ignore-user-config` is not invented in top-level help; captured help bytes are not rescrubbed). Recipe quoting is covered from the justfile body without a just binary (POSIX `sh` + positional argv; Windows uses the recipe's `node -e` / argv shape). The developer entry points are gated three ways: docs-to-recipe parity (a documented `just` verb in command form, across inline code or a fenced line with an optional `#`, tolerant of stray whitespace, pipe-separated alternations, leading interpreter options and `~~~` fences, reading a shell pipe as one command rather than an alternation, skipping a glob family like `review-*` instead of demanding a recipe, and *not* parsing prose, headings or captured listing output); per-recipe boundaries (**the load-bearing gate parses nothing**: every line that mentions `assignment-run.mjs` must be one of the seven pinned proof bodies and there must be exactly seven, so a header form no parser recognises still cannot hide a call site; column-0 comments are documentation and excluded), plus a normalized-text pass (`\`-continuations folded, comments stripped) and a second pass over `just --dump`, the interpreter's own rendering, skipped with a visible reason without the binary; the recipe name set must equal a pinned list, duplicate headers and duplicate `run :=` bindings are refused, `set allow-duplicate*`, `import`/`mod` in any spelling and `alias` are refused, `run :=` and `dispatch` are asserted exactly, and non-proof bodies are token-checked against raw text because comment-stripping first hid an executable suffix that `--dump` reproduced verbatim. Stated limit: drift protection, not a sandbox — a body can assemble its command at runtime, and anyone who can edit the file can already do what it does. Auto-loading a working-directory dotenv file is refused by a brake on any `set dotenv*` spelling, plus a real-`just` probe whose preload module writes a marker file itself — so the assertion is that injected code executed — paired with controls on the same entry point: explicit `--dotenv-path`, a justfile with the setting re-added, and a real `witness` recipe run both ways. Real just integration is optional and skipped when the binary is absent. No live paid smoke. |
|
|
99
|
-
| Long-lived headless coordinator | `scripts/kxm-worker.mjs` | Restart limits, spawn failure, collision-resistant ownership, exact resource and tool loading, raw-output isolation, bounded RPC framing, bounded drain, hung-tool recovery, provider/model fallback, and `--continue` fallback are automated; the opt-in real-Pi gate verifies two workers, discovery, request/reply, fanout, durable restart/resume, journal, and checkpoint |
|
|
100
|
-
| GitHub check signal adapter | `plugins/kxm/src/github-watch.ts` | Deterministic pagination, conclusion, retry, and per-wait delivery-generation states in `test/core/github-watch.test.ts` |
|
|
101
|
-
| Operator CLI | `scripts/kxm.mjs` | Isolated workspace commands in `test/core/cli.test.ts`; the packed artifact is installed locally and with the documented global `--omit=peer` path by `test/core/package-install.test.ts` |
|
|
102
|
-
| Hub-local session brief | `plugins/kxm/src/session-work.ts`, `test/core/session-work.test.ts`, `test/core/cli.test.ts` | Status line and task/plan lists from hub SQLite without message bodies; `init --hub` is an unknown option; `hub bind` reports on/off/unknown |
|
|
103
|
-
| Native-free package install and Windows `pi.cmd` worker launch | `package.json`, `test/core/store.test.ts`, `test/core/worker.test.ts` | CI runs on Linux at Node 22.19 and Node 24; the Windows `pi.cmd` fixture stays in `test/core/worker.test.ts` and runs locally on Windows or when the paused Windows legs resume. |
|
|
104
|
-
|
|
105
|
-
## Manual release checks
|
|
106
|
-
|
|
107
|
-
Automation cannot prove that a third-party harness UI renders perfectly. Before a release, connect two current Pi sessions, run `/kxm hub`, complete one inbound round trip, install the marketplace plugin in a clean Claude Code profile, and verify `kxm_list`. Exercise preview channel delivery only when the target Claude Code version supports community channels.
|
|
108
|
-
|
|
109
|
-
Create the versioned tarball with `npm pack`, attach it to the matching GitHub
|
|
110
|
-
release, and verify the authenticated `gh release download` plus
|
|
111
|
-
`npm install --global --omit=peer <local-tarball>` path before publishing the
|
|
112
|
-
operator installation instructions. For version `<release-version>`, the required asset is
|
|
113
|
-
`kxm-<release-version>.tgz`.
|
|
114
|
-
|
|
115
|
-
When adding a feature, add executable coverage and update this matrix in the same change. If a behavior can only be verified manually, state why and add it to the release checklist instead of implying automated coverage.
|
|
116
|
-
|
|
117
|
-
## v0.5 context suites
|
|
118
|
-
|
|
119
|
-
| Suite | Covers |
|
|
120
|
-
|---|---|
|
|
121
|
-
| `test/core/context.test.ts` | Context schema round-trips, hostile input, cross-project fail-closed, storage upgrade |
|
|
122
|
-
| `test/core/state.test.ts` | Temporal state lifecycle, asOf queries, supersession, contradictions, restart durability |
|
|
123
|
-
| `test/core/context-authority.test.ts` | Authority grant floor, reserialization escalation, lineage bounds, control-plane smuggling |
|
|
124
|
-
| `test/core/arbiter.test.ts` | Role-aware packet assembly, task-relevance ranking, first-fit budgets, the evidence section, contradiction routing, journal conversion, hub surfaces |
|
|
125
|
-
| `test/core/context-surfaces.test.ts` | CLI and Pi tool parity for the context API |
|
|
126
|
-
| `test/core/journal-evolution.test.ts` | New journal categories, evidence requirements, governed promotion, stage provenance through the shared tool |
|
|
127
|
-
| `test/core/wiki.test.ts` | Wiki compilation determinism, lifecycle preservation, contradiction visibility, lint |
|
|
128
|
-
| `test/core/workflow-transitions.test.ts` | Typed back-edges, budgets, bypass protection, restart recovery |
|
|
129
|
-
| `test/core/fix-workflow.test.ts` | /fix end-to-end, independent repro-review oracle, wrong-seam invalidation, failed self-retry, plan-hash gating, exhaustion |
|
|
130
|
-
| `test/core/skills.test.ts` | Skill candidate lifecycle, quarantine, immutability, CLI |
|
|
131
|
-
| `test/core/routing.test.ts` | Behavioral hash, record parsing, comparisons, routing report |
|
|
132
|
-
| `test/core/improve.test.ts` | Routing-record sources, event-log outcome resolution, same-ask candidacy, candidate files, promotion readiness |
|