@kontextmind/kxm 0.7.95 → 0.7.96
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/marketplace.json +1 -1
- package/.kxm/README.md +39 -9
- package/CHANGELOG.md +1 -1
- package/README.md +147 -257
- package/SECURITY.md +21 -12
- package/docs/README.md +133 -54
- package/docs/adr/ADR-0002-browser-automation-steel-doks.md +24 -18
- package/docs/adr/ADR-0003-sqlite-only-store.md +100 -0
- package/docs/adr/ADR-0004-edge-identity-authentik.md +99 -0
- package/docs/adr/README.md +33 -0
- package/docs/concepts/architecture.md +262 -0
- package/docs/concepts/data-and-storage.md +194 -0
- package/docs/concepts/trust-model.md +152 -0
- package/docs/contracts/README.md +22 -14
- package/docs/contracts/effects-and-recovery.md +3 -0
- package/docs/contracts/migration.md +2 -2
- package/docs/contracts/routing.md +6 -5
- package/docs/contributing/assignment-runner.md +388 -0
- package/docs/contributing/ci-and-release.md +231 -0
- package/docs/contributing/development.md +362 -0
- package/docs/contributing/harness-routing-internals.md +192 -0
- package/docs/{packages.md → contributing/packages.md} +13 -15
- package/docs/{skills → contributing}/repo-work-delivery.md +20 -21
- package/docs/contributing/test-matrix.md +208 -0
- package/docs/{tui-components.md → contributing/tui-components.md} +30 -22
- package/docs/contributing/writing-docs.md +340 -0
- package/docs/glossary.md +471 -0
- package/docs/guides/agent-skills.md +137 -0
- package/docs/guides/browser-automation.md +160 -0
- package/docs/guides/context-and-memory.md +352 -0
- package/docs/guides/continuous-improvement.md +228 -0
- package/docs/guides/governed-skills.md +173 -0
- package/docs/guides/nous-providers.md +186 -0
- package/docs/guides/peer-messaging.md +304 -0
- package/docs/guides/pi-workers.md +219 -0
- package/docs/guides/provenance-gates.md +313 -0
- package/docs/guides/webhook-workflows.md +364 -0
- package/docs/kb/how-credentials-retrieved-safely.md +38 -12
- package/docs/kb/how-to-capture-and-annotate-section.md +15 -13
- package/docs/kb/how-to-connect-playwright-to-steel.md +16 -11
- package/docs/kb/how-to-recover-expired-session-or-orphan.md +26 -16
- package/docs/kb/how-to-resume-after-mfa.md +19 -11
- package/docs/kb/how-to-take-over-session.md +17 -13
- package/docs/kb/why-authentication-disappeared.md +22 -14
- package/docs/kb/why-automation-opened-different-browser.md +23 -14
- package/docs/kb/why-session-viewer-cannot-control.md +13 -12
- package/docs/operations/backup-and-restore.md +248 -0
- package/docs/operations/deploy.md +307 -0
- package/docs/operations/monitoring.md +209 -0
- package/docs/operations/runtime-sync.md +192 -0
- package/docs/operations/troubleshooting.md +265 -0
- package/docs/operations/upgrade.md +124 -0
- package/docs/prompts/browser-annotate-feedback.md +7 -7
- package/docs/prompts/browser-diagnose-recover.md +11 -10
- package/docs/prompts/browser-explore.md +7 -7
- package/docs/prompts/browser-repro-fix.md +7 -7
- package/docs/prompts/browser-start.md +12 -11
- package/docs/prompts/browser-takeover.md +8 -8
- package/docs/{cli-reference.md → reference/cli-reference.md} +83 -41
- package/docs/{config-reference.md → reference/config-reference.md} +159 -148
- package/docs/reference/configuration.md +299 -0
- package/docs/reference/harness-routing.md +508 -0
- package/docs/reference/http-api.md +203 -0
- package/docs/reference/tools.md +370 -0
- package/docs/{workflow-guide.md → reference/workflow-catalog.md} +92 -153
- package/docs/reference/workflow-definitions.md +286 -0
- package/docs/start/first-workflow.md +287 -0
- package/docs/start/install.md +146 -0
- package/docs/start/quickstart-claude-code.md +405 -0
- package/docs/start/quickstart-pi.md +213 -0
- package/docs/templates/README.md +78 -73
- package/docs/templates/adr.md +13 -13
- package/docs/templates/architecture.md +55 -71
- package/docs/templates/bug-fix.md +13 -16
- package/docs/templates/feature.md +14 -19
- package/docs/templates/handoff.md +44 -46
- package/docs/templates/postmortem.md +30 -43
- package/docs/templates/research.md +15 -20
- package/docs/templates/review.md +49 -50
- package/docs/templates/runbook.md +38 -30
- package/docs/templates/test-plan.md +16 -23
- package/docs/templates/test-report.md +14 -17
- package/examples/README.md +9 -5
- package/examples/provenance-workflow.json +1 -1
- package/examples/webhook-workflows/jira-development.json +59 -0
- package/examples/webhook-workflows/jira-issue-updated.json +12 -0
- package/package.json +1 -1
- package/packages/core/tui/README.md +1 -1
- package/plugins/kxm/.claude-plugin/plugin.json +1 -1
- package/plugins/kxm/README.md +31 -32
- package/plugins/kxm/dist/cli.js +5 -5
- package/plugins/kxm/dist/mcp-server.js +1 -1
- package/plugins/kxm/dist/runtime.js +1 -1
- package/plugins/kxm/package.json +1 -1
- package/plugins/kxm/skills/kxm/references/protocol.md +3 -1
- package/plugins/kxm/skills/kxm-browser-auth/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-browser-diagnostics/SKILL.md +5 -5
- package/plugins/kxm/skills/kxm-browser-explore/SKILL.md +2 -2
- package/plugins/kxm/skills/kxm-browser-session/SKILL.md +10 -13
- package/plugins/kxm/skills/kxm-browser-takeover/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-browser-verify/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-context-memory/SKILL.md +13 -4
- package/plugins/kxm/skills/kxm-hub-ops/SKILL.md +3 -1
- package/plugins/kxm/skills/kxm-mind-setup/SKILL.md +2 -1
- package/plugins/kxm/skills/kxm-project-setup/SKILL.md +31 -54
- package/plugins/kxm/skills/kxm-projects/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-protocol/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-routing-improve/SKILL.md +15 -7
- package/plugins/kxm/skills/kxm-runs/SKILL.md +11 -5
- package/plugins/kxm/skills/kxm-session/SKILL.md +1 -1
- package/plugins/kxm/skills/kxm-tasks/SKILL.md +9 -7
- package/plugins/kxm/skills/kxm-workflow/SKILL.md +10 -2
- package/plugins/kxm/src/cli/system.ts +1 -1
- package/plugins/kxm/src/cli.ts +3 -3
- package/plugins/kxm/src/init-guide-setup.ts +1 -1
- package/plugins/kxm/src/mcp-server.ts +1 -1
- package/plugins/kxm/src/modes.ts +1 -1
- package/schemas/README.md +1 -1
- package/docs/agent-communication-envelopes-and-gates.md +0 -553
- package/docs/agent-skills.md +0 -198
- package/docs/architecture.md +0 -245
- package/docs/assignment-runner.md +0 -264
- package/docs/browser-automation.md +0 -139
- package/docs/configuration.md +0 -437
- package/docs/continuous-improvement.md +0 -226
- package/docs/getting-started.md +0 -277
- package/docs/harness-routing.md +0 -616
- package/docs/kb/qa-authentik-authentication.md +0 -97
- package/docs/kb/qa-extension-install-and-hub-bootstrap.md +0 -85
- package/docs/kb/qa-hub-on-a-public-host.md +0 -48
- package/docs/kb/qa-sqlite-vs-duckdb.md +0 -35
- package/docs/kb/qa-what-the-hub-stores.md +0 -64
- package/docs/kxm-handbook.md +0 -1181
- package/docs/operations.md +0 -510
- package/docs/operator-pi-packages.md +0 -67
- package/docs/provenance-gates.md +0 -295
- package/docs/skills.md +0 -47
- package/docs/test-matrix.md +0 -132
- package/docs/troubleshooting.md +0 -322
- package/docs/webhook-workflows.md +0 -240
|
@@ -1,226 +0,0 @@
|
|
|
1
|
-
# Continuous improvement
|
|
2
|
-
|
|
3
|
-
Every workflow run produces two distinct records:
|
|
4
|
-
|
|
5
|
-
- operational events for service health and delivery;
|
|
6
|
-
- a structured journal for plans, decisions, contradictions, errors, lessons, observations, hypotheses, experiments, state changes, and skill candidates.
|
|
7
|
-
|
|
8
|
-
The journal and retrospective loop covers hub webhook runs (signed webhooks and
|
|
9
|
-
`kxm workflow start`). Runs started with `kxm run` on the Runtime have no journal or
|
|
10
|
-
retrospective yet: `kxm_workflow_record` answers `workflow_not_found` for a Runtime
|
|
11
|
-
run id. Journal capture uses `kxm_workflow_record`. Terminal runs export a bounded
|
|
12
|
-
retrospective under `.kxm/assets/retrospectives`; re-export from durable local state
|
|
13
|
-
with `kxm workflow export`.
|
|
14
|
-
|
|
15
|
-
`kxm improve` is a separate loop. It reads routing records (the Runtime's settled
|
|
16
|
-
attempts and telemetry), not the journal, and proposes coded repeats; see
|
|
17
|
-
[Coded repeats](#coded-repeats-kxm-improve).
|
|
18
|
-
|
|
19
|
-
Journal entries carry an improvement area, severity, evidence links, relationships to other entries, and (when stage-bound) run/stage/attempt provenance. The design preserves disagreement instead of flattening it into a single final answer.
|
|
20
|
-
|
|
21
|
-
## Improvement areas
|
|
22
|
-
|
|
23
|
-
| Area | Questions to improve |
|
|
24
|
-
|---|---|
|
|
25
|
-
| Harness | Were agents available, correctly prompted, recovered, and given usable tools? |
|
|
26
|
-
| Gates | Did checks catch defects, produce actionable output, and avoid flaky noise? |
|
|
27
|
-
| Implementation | Were ownership, design choices, migrations, and tests effective? |
|
|
28
|
-
| Workflow | Were stages ordered well, attempts bounded, and handoffs efficient? |
|
|
29
|
-
| Documentation | Could users and operators follow the changed behavior? |
|
|
30
|
-
| Security | Were trust boundaries, credentials, permissions, and side effects controlled? |
|
|
31
|
-
| Other | What important issue does not fit the established taxonomy? |
|
|
32
|
-
|
|
33
|
-
## Required capture behavior
|
|
34
|
-
|
|
35
|
-
Use `kxm_workflow_record` during the run, not only in a final retrospective:
|
|
36
|
-
|
|
37
|
-
- record a `plan` before implementation;
|
|
38
|
-
- record a `decision` with alternatives and why one was chosen;
|
|
39
|
-
- record a `contradiction` when agents, tests, documentation, or observed behavior disagree;
|
|
40
|
-
- record an `error` when a stage, tool, gate, integration, or assumption fails;
|
|
41
|
-
- record a `lesson` only after evidence supports a reusable conclusion (evidence references are mandatory);
|
|
42
|
-
- record an `observation` for notable behavior without a causal claim;
|
|
43
|
-
- record a `hypothesis` as a falsifiable claim, and keep it when rejected — a disproven hypothesis is durable learning;
|
|
44
|
-
- record an `experiment` with its outcome, including failures;
|
|
45
|
-
- record a `state-change` when an authoritative project fact changes;
|
|
46
|
-
- record a `skill-candidate` only with verified run/receipt evidence; candidates never become promoted skills without a protected evaluation.
|
|
47
|
-
|
|
48
|
-
`kxm_workflow_record` (and `kxm workflow record`) accepts all ten categories above.
|
|
49
|
-
Pass `stageId` to bind an entry to the stage it is about. The hub, never the caller,
|
|
50
|
-
derives the attempt from the stage's state: the current attempt for an in-progress or
|
|
51
|
-
waiting stage, the last attempt consumed for a finished stage, and none for a pending
|
|
52
|
-
stage that has not run. When `stageId` names a stage that declares an area, `area` may
|
|
53
|
-
be omitted and defaults to that area; otherwise `area` is required, and a request with
|
|
54
|
-
neither is refused with `invalid_improvement_area`. A `stageId` that is not part of the
|
|
55
|
-
run is refused with `invalid_journal_relation`. `lesson` and `skill-candidate` entries
|
|
56
|
-
still require evidence references.
|
|
57
|
-
|
|
58
|
-
## Governed promotion
|
|
59
|
-
|
|
60
|
-
`skill-candidate`, `hypothesis`, and `experiment` entries participate in a governed lifecycle: `proposed` → `approved` | `rejected` | `quarantined`. Promotion is an append-only, admin-controlled decision (`POST /v1/journal/:id/promotion`) that requires durable evidence references; the author of an entry can never decide its promotion, and terminal states never re-open. Promotion changes the learning lifecycle of an entry — never gates, workflow policy, or permissions.
|
|
61
|
-
|
|
62
|
-
The native Pi extension automatically records failed tool results while a webhook workflow is active. The hub also records checkpoint warnings and failures, transition-budget exhaustion, signed signal results, external-wait timeouts, prompt expiry, degraded-quorum approvals, and premature coordinator settlement, each bound to its stage and attempt. Agents must still record semantic errors such as a false assumption, rejected design, flaky result, or external integration mismatch.
|
|
63
|
-
|
|
64
|
-
Never put secrets or unnecessary prompt contents in the journal. Evidence should be durable references such as test names, logs, commits, pull requests, Jira issues, check runs, or documentation paths. Failed tools record an allowlisted diagnostic class, not stdout.
|
|
65
|
-
|
|
66
|
-
Every terminal workflow automatically exports a bounded retrospective under `.kxm/assets/retrospectives`, and exports it again when a journal entry is recorded or a promotion decided after the run ended. Re-export one from durable local state with `kxm workflow export <runId>`; `--input <snapshot.json>` remains available for offline imports. `recurringErrorClasses` counts `error` entries only. `proposedImprovements` holds up to 12 of the run's error and lesson entries, merged and ranked the same way as the weekly signals below, each with a success measure that names the signal key. Files stay `reviewDecision=proposed` until a human or coordinator records an explicit decision. Export never edits workflow JSON or weakens gates.
|
|
67
|
-
|
|
68
|
-
Runs with peer policies add an optional metadata-only evidence audit while
|
|
69
|
-
retaining the `pi-mesh.retrospective.v1` schema. It records each requirement's
|
|
70
|
-
configured and effective producer minimum, eligible-producer snapshot, verified
|
|
71
|
-
message and producer IDs, immutable workflow context, lifecycle timestamps,
|
|
72
|
-
request/reply hashes, degraded state, and explicit admin approvals for the
|
|
73
|
-
applied attempt. Earlier-attempt replies never inflate the final quorum. It
|
|
74
|
-
never copies peer request or reply bodies. The snapshot remains useful after normal
|
|
75
|
-
message retention purges the source record, but its hashes are provenance
|
|
76
|
-
metadata—not proof that the peer's conclusion was true.
|
|
77
|
-
|
|
78
|
-
## Review cadence
|
|
79
|
-
|
|
80
|
-
### Per run
|
|
81
|
-
|
|
82
|
-
The retrospective stage reviews journal entries, groups contributing causes, and proposes bounded improvements. It must identify an owner or next action and a measurable success condition.
|
|
83
|
-
|
|
84
|
-
### Weekly
|
|
85
|
-
|
|
86
|
-
Call `kxm_improvement_report`. Besides the per-area counts it returns `signals`: the
|
|
87
|
-
project's errors, contradictions, lessons and skill candidates, already merged across
|
|
88
|
-
runs and ranked. Review the top signals, check that each merge groups entries that
|
|
89
|
-
belong together, and pick what to trial.
|
|
90
|
-
|
|
91
|
-
- Entries merge when they share a key: first an evidence class (`class:<name>` in the
|
|
92
|
-
entry's evidence), then, for errors, the workflow definition and stage, then the
|
|
93
|
-
summary after redaction and normalization (ids, timestamps, hex strings and numbers
|
|
94
|
-
are folded). Text is redacted before it becomes a key, so a key never carries a raw
|
|
95
|
-
summary. Each signal keeps up to 16 source run IDs and entry IDs.
|
|
96
|
-
- Only errors, open contradictions (not yet resolved by a related decision or lesson),
|
|
97
|
-
lessons, and skill candidates still `proposed` count.
|
|
98
|
-
- `frequency` is the number of distinct runs, counted over the runs the hub still
|
|
99
|
-
retains: terminal runs and their journal are purged 7 days after they end.
|
|
100
|
-
|
|
101
|
-
```text
|
|
102
|
-
priority = frequency × severity weight × workflow cost × confidence
|
|
103
|
-
```
|
|
104
|
-
|
|
105
|
-
- The severity weight is 3 for `error`, 2 for `warning` and 1 for `info`, taken from the
|
|
106
|
-
most severe entry in the signal.
|
|
107
|
-
- Workflow cost is the mean number of run attempts (stage attempts plus transitions)
|
|
108
|
-
over the signal's known runs. It is not dollars. When none of the runs is known it
|
|
109
|
-
counts as 1 and the signal reports `costBasis: "unknown"`.
|
|
110
|
-
- Confidence is 0.5 plus half the share of the signal's entries that cite evidence.
|
|
111
|
-
|
|
112
|
-
Security signals (area `security`, or evidence class `invalid_auth`,
|
|
113
|
-
`invalid_identity` or `signal_mismatch`) rank ahead of every priority. Remaining ties
|
|
114
|
-
break on frequency, then key, never on entry ID or insertion order. There is no
|
|
115
|
-
data-loss override yet, because no deterministic data-loss marker exists, so review
|
|
116
|
-
data-loss risk by hand rather than trusting the order.
|
|
117
|
-
|
|
118
|
-
### Per release
|
|
119
|
-
|
|
120
|
-
Select a small improvement batch. For each proposal:
|
|
121
|
-
|
|
122
|
-
1. State the observed problem and linked evidence.
|
|
123
|
-
2. Identify whether the change affects the harness, gates, implementation guidance, workflow definition, documentation, or security policy.
|
|
124
|
-
3. Define the expected outcome and a measurable leading indicator.
|
|
125
|
-
4. Add or update tests before changing enforcement.
|
|
126
|
-
5. Trial the change on a bounded workflow or repository.
|
|
127
|
-
6. Compare failure rate, cycle time, manual intervention, and escaped defects with the baseline.
|
|
128
|
-
7. Adopt, revise, or roll back the proposal.
|
|
129
|
-
8. Record the decision and result in a subsequent workflow journal.
|
|
130
|
-
|
|
131
|
-
## Coded repeats (`kxm improve`)
|
|
132
|
-
|
|
133
|
-
`kxm improve` (the same as `kxm improve report`) looks for agent steps that a script,
|
|
134
|
-
test or workflow `gate` could do as well as a model. It proposes; it never applies.
|
|
135
|
-
|
|
136
|
-
**Sources.** Inside a KXM project it reads the project's Runtime event store,
|
|
137
|
-
`<state>/runtime/projects/<key>/run-events.db`, and then `.kxm/logs/telemetry.jsonl`.
|
|
138
|
-
The key comes from the checkout's real path, so each checkout and worktree has its own
|
|
139
|
-
store and the report covers only the one it runs in. The store is opened read-only for
|
|
140
|
-
one query over its events table; `kxm improve` never creates, writes or migrates it. A
|
|
141
|
-
telemetry record whose `attemptId` the store already supplied is dropped. `--file
|
|
142
|
-
<path>` reads only that file. Outside a project only telemetry is read. The output lists
|
|
143
|
-
every source with its path, whether it exists, and its counts (`records`,
|
|
144
|
-
`skippedInvalid`, `excludedSimulated`, `undecided`, `duplicatesDropped`). An unreadable
|
|
145
|
-
store stops the command with `improve_source_unreadable` and the path (exit 1).
|
|
146
|
-
|
|
147
|
-
**Identity.** Records group by workflow, step, agent role and ask. Runtime records carry
|
|
148
|
-
the engine-reserved `workflowId` and `askSha256` keys. The ask digest covers the
|
|
149
|
-
workflow, step, step kind, agent, instructions, outcomes and required evidence keys, so
|
|
150
|
-
it is the same for one step across runs and ignores the run, attempt, model and context
|
|
151
|
-
packet. Records without those keys fall back to a workflow definition digest or the run
|
|
152
|
-
ID, and to `rolePromptSha256`. No prompt text is read or stored: `objectiveSha256` is
|
|
153
|
-
the digest of the run's prompt.
|
|
154
|
-
|
|
155
|
-
**Outcomes.** A Runtime record stores `finalOutcome` only when settlement already knows
|
|
156
|
-
it: `blocked` for a back edge, and `failed` for a producer error, an undeclared outcome
|
|
157
|
-
or a failing terminal. Everything else is resolved when the report reads the event log,
|
|
158
|
-
and never written back: a later entry into the same step makes the attempt `reworked`, a
|
|
159
|
-
completed run makes it `accepted`, a failed run makes it `failed`, and a cancelled or
|
|
160
|
-
still-running run leaves it undecided. Undecided records are counted and left out of the
|
|
161
|
-
pass rate. Attempts from simulated drives are excluded and counted.
|
|
162
|
-
|
|
163
|
-
**Candidacy.** A group becomes a coded-repeat candidate only when all three hold:
|
|
164
|
-
|
|
165
|
-
- the same objective (`objectiveSha256`) was decided in at least 2 runs
|
|
166
|
-
(`askRecurrence`); records without an objective digest share one bucket;
|
|
167
|
-
- at least 0.75 of its decided records were accepted (`verifyPassRate`), where an
|
|
168
|
-
attempt superseded by a later retry of the same step in the same run never counts as a
|
|
169
|
-
pass;
|
|
170
|
-
- its step writes no repository (`writesRepository`, from the engine's `stepWrites` key).
|
|
171
|
-
|
|
172
|
-
A group that passes but misses reports `excludedReason`: `writes-repository`, or
|
|
173
|
-
`ask-not-repeated` when its runs asked different objectives. `weightedRecurrence`
|
|
174
|
-
weights each record by `2^(-age / improvement.telemetryHalfLifeDays)` (14 days by
|
|
175
|
-
default; an undated record weighs 1 and is counted in `undatedRecords`). It orders the
|
|
176
|
-
rows and never decides candidacy.
|
|
177
|
-
|
|
178
|
-
**Outputs.** Each candidate is a `kxm.candidate.v1` JSON file and a proposed diff under
|
|
179
|
-
`.kxm/candidates/` (or `--out-dir`); the report is written as
|
|
180
|
-
`kxm.improvement-report.v2` under `<workspace>/assets/improvements/`. `--dry-run` writes
|
|
181
|
-
neither. The candidate kind comes from a Git-reviewed name rule and only picks the
|
|
182
|
-
proposal template: a step or role that verifies, gates, tests, checks or lints proposes a
|
|
183
|
-
`.kxm/gates.yaml` command entry; a planning, review or repro step proposes a governed
|
|
184
|
-
skill, labelled consolidation because it is not a coded step; any other step proposes
|
|
185
|
-
replacing the agent step with a `kind: gate` step plus a `gates.yaml` command entry that
|
|
186
|
-
runs `scripts/<step>.mjs`, which the operator writes. Diffs are proposals with
|
|
187
|
-
placeholder hunks.
|
|
188
|
-
|
|
189
|
-
**Promotion readiness.** `promotion[]` reports, per candidate, `readyForReview` and a
|
|
190
|
-
reason under `improvement.promotionPolicy`:
|
|
191
|
-
|
|
192
|
-
- `manual_pr` (default): always ready; the operator reviews and applies the diff in a PR.
|
|
193
|
-
- `critic_quorum`: ready once two distinct critic receipts are cited. `kxm improve`
|
|
194
|
-
cites none, so every candidate reports not ready under this policy.
|
|
195
|
-
- `auto_threshold`: ready once the group's distinct runs reach
|
|
196
|
-
`improvement.autoThreshold.minRuns`, its accepted share reaches `minPassRate`, and its
|
|
197
|
-
mean recorded cost is at least `minCostSavings` over at least one cost sample. A group
|
|
198
|
-
with no recorded cost is never ready.
|
|
199
|
-
|
|
200
|
-
No policy authorizes anything. Every policy ends at an operator PR, activation is a
|
|
201
|
-
reviewed Git change for a future run, and telemetry cannot grant tools or skip a gate.
|
|
202
|
-
|
|
203
|
-
## Governance safeguards
|
|
204
|
-
|
|
205
|
-
- Journal content is evidence, not executable policy.
|
|
206
|
-
- Improvement candidates are proposals. Promotion readiness never authorizes, and no `improvement.*` value activates a candidate.
|
|
207
|
-
- An agent may propose a gate change but cannot silently weaken a required gate.
|
|
208
|
-
- A peer-quorum reduction must be declared by policy and explicitly approved by an administrator for the current attempt; record it as a degraded outcome rather than normal success.
|
|
209
|
-
- Contradictions stay open until evidence resolves them; synthesis must not erase minority risks.
|
|
210
|
-
- Changes involving permissions, secrets, merge policy, or external side effects require human or repository-authorized approval.
|
|
211
|
-
- Improvement reports are project-scoped. Protect the SQLite database because journal details may reveal sensitive engineering context.
|
|
212
|
-
- Periodically export accepted decisions and durable lessons into version-controlled documentation; the SQLite journal is an audit source, not the only system of record.
|
|
213
|
-
|
|
214
|
-
## Initial backlog for this workflow
|
|
215
|
-
|
|
216
|
-
Start by measuring and improving these areas:
|
|
217
|
-
|
|
218
|
-
1. Harness availability: coordinator uptime, reconnect count, prompt expiry, and premature settlement.
|
|
219
|
-
2. Planning quality: number of unresolved contradictions and reviewer-found plan defects.
|
|
220
|
-
3. Gate quality: flaky checks, false negatives, rerun count, and time to actionable failure output.
|
|
221
|
-
4. Implementation quality: escaped defects, rollback rate, review churn, and ownership conflicts.
|
|
222
|
-
5. Delivery flow: time in each stage, webhook-to-start latency, and manual intervention count.
|
|
223
|
-
6. Documentation quality: setup failures and changes shipped without updated examples or operations guidance.
|
|
224
|
-
7. Security posture: rejected signatures, secret rotation age, unauthorized action attempts, and dependency findings.
|
|
225
|
-
|
|
226
|
-
The goal is not maximum automation. It is a workflow that becomes more reliable, explainable, and efficient while preserving review and authorization boundaries.
|
package/docs/getting-started.md
DELETED
|
@@ -1,277 +0,0 @@
|
|
|
1
|
-
# Getting started
|
|
2
|
-
|
|
3
|
-
This guide takes you from a clean machine to a successful peer request. Allow about ten minutes once Node.js, Git, and your agent harnesses are installed.
|
|
4
|
-
|
|
5
|
-
## Before you begin
|
|
6
|
-
|
|
7
|
-
You need:
|
|
8
|
-
|
|
9
|
-
- Node.js 22.19 or newer on the 22.x line, or Node.js 24 or newer;
|
|
10
|
-
- Git;
|
|
11
|
-
- GitHub CLI for the command-first release install;
|
|
12
|
-
- Pi for Pi agents;
|
|
13
|
-
- Claude Code only if you want a mixed Pi/Claude pool;
|
|
14
|
-
- access to `kontextmind/kxm` while the repository is private.
|
|
15
|
-
|
|
16
|
-
All agents in one pool must use the same hub URL, project token, and project name. Keep the hub/operator administrative token separate. Every active agent in that project must have a unique name.
|
|
17
|
-
|
|
18
|
-
## 1. Install
|
|
19
|
-
|
|
20
|
-
Pi's Git package installation supplies the extension and Agent Skill but does
|
|
21
|
-
not add `kxm` to `PATH`. Download the packed release through an authenticated
|
|
22
|
-
GitHub CLI session and install that local tarball. Run `gh auth login` first if
|
|
23
|
-
necessary.
|
|
24
|
-
|
|
25
|
-
PowerShell:
|
|
26
|
-
|
|
27
|
-
```powershell
|
|
28
|
-
$version = "<release-version>"
|
|
29
|
-
$asset = "kxm-$version.tgz"
|
|
30
|
-
$releaseDir = Join-Path $PWD ".kxm-release"
|
|
31
|
-
New-Item -ItemType Directory -Force -Path $releaseDir | Out-Null
|
|
32
|
-
gh release download "v$version" --repo kontextmind/kxm --pattern $asset --dir $releaseDir --clobber
|
|
33
|
-
npm install --global --omit=peer (Join-Path $releaseDir $asset)
|
|
34
|
-
pi install git:github.com/kontextmind/kxm@main
|
|
35
|
-
kxm --help
|
|
36
|
-
```
|
|
37
|
-
|
|
38
|
-
Bash:
|
|
39
|
-
|
|
40
|
-
```bash
|
|
41
|
-
version='<release-version>'
|
|
42
|
-
asset="kxm-${version}.tgz"
|
|
43
|
-
mkdir -p .kxm-release
|
|
44
|
-
gh release download "v${version}" --repo kontextmind/kxm \
|
|
45
|
-
--pattern "$asset" --dir .kxm-release --clobber
|
|
46
|
-
npm install --global --omit=peer ".kxm-release/$asset"
|
|
47
|
-
pi install git:github.com/kontextmind/kxm@main
|
|
48
|
-
kxm --help
|
|
49
|
-
```
|
|
50
|
-
|
|
51
|
-
Do not substitute a global `git+https` npm install; the supported global
|
|
52
|
-
operator package is the versioned release tarball. To run from source instead,
|
|
53
|
-
clone the repository, run `npm ci`, and use `node scripts/kxm.mjs` in place
|
|
54
|
-
of `kxm`.
|
|
55
|
-
|
|
56
|
-
## 2. Initialize the project
|
|
57
|
-
|
|
58
|
-
```text
|
|
59
|
-
kxm init
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
`kxm init` never copies the package repository's dogfood roster or workflows into a consumer workspace.
|
|
63
|
-
|
|
64
|
-
When `kxm init` succeeds in an interactive terminal, it offers to install shell
|
|
65
|
-
completion for the detected shell. Accepting writes the completion script
|
|
66
|
-
under the user config directory, appends one idempotent stanza to the shell
|
|
67
|
-
rc file, and, when the kxm bin directory is not already on `PATH`, adds a
|
|
68
|
-
`PATH` export. Declining is safe: run `kxm completion install` later, or set
|
|
69
|
-
`KXM_SKIP_COMPLETION_PROMPT=1` to suppress the offer. Non-interactive,
|
|
70
|
-
`--json`, and `--dry-run` runs never prompt or write shell files.
|
|
71
|
-
|
|
72
|
-
After the completion offer, an interactive `kxm init` also offers to set up
|
|
73
|
-
workflow-guide agents and workflows for the harnesses you have installed and
|
|
74
|
-
authenticated. Accepting lists the software-engineering workflows from
|
|
75
|
-
[`workflow-guide.md`](workflow-guide.md); pick by number or slug (`all` works
|
|
76
|
-
too). kxm resolves each role's first guide candidate whose harness is
|
|
77
|
-
authenticated and writes only current KXM project resources —
|
|
78
|
-
`.kxm/agents/<role>.yaml` (`kxm.agent.v1`) and `.kxm/workflows/<slug>.yaml`
|
|
79
|
-
(`kxm.workflow.v1`). It never writes retired legacy authority (`.kxm/config`,
|
|
80
|
-
retired `.kxm/roster.json`) or the trusted `.kxm/roster.yaml` policy. Roles whose candidates have no authenticated harness are
|
|
81
|
-
reported as skipped, not silently downgraded. Guide candidates are dated
|
|
82
|
-
research — verify them before dispatch. Declining is safe: set
|
|
83
|
-
`KXM_SKIP_GUIDE_SETUP_PROMPT=1` to suppress the offer.
|
|
84
|
-
|
|
85
|
-
## 3. Start the hub in another terminal
|
|
86
|
-
|
|
87
|
-
`kxm hub start` is foreground. Keep that terminal running. The Pi extension
|
|
88
|
-
can also start the hub for you (`hub.autoStart: background`, the default in
|
|
89
|
-
`kxm.config.v1`): on load it reuses a healthy bound hub or a live local claim
|
|
90
|
-
and starts a detached wrapper only when none exists.
|
|
91
|
-
|
|
92
|
-
PowerShell:
|
|
93
|
-
|
|
94
|
-
```powershell
|
|
95
|
-
$env:KXM_AUTH_TOKEN = "replace-with-an-admin-token"
|
|
96
|
-
$env:KXM_PROJECT_TOKENS = '{"demo":"replace-with-a-demo-project-token"}'
|
|
97
|
-
kxm hub start
|
|
98
|
-
```
|
|
99
|
-
|
|
100
|
-
Bash:
|
|
101
|
-
|
|
102
|
-
```bash
|
|
103
|
-
export KXM_AUTH_TOKEN="replace-with-an-admin-token"
|
|
104
|
-
export KXM_PROJECT_TOKENS='{"demo":"replace-with-a-demo-project-token"}'
|
|
105
|
-
kxm hub start
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
A successful start prints:
|
|
109
|
-
|
|
110
|
-
```text
|
|
111
|
-
kxm hub listening at http://127.0.0.1:7331; storage=<workspace>/.kxm/state/kxm.db
|
|
112
|
-
```
|
|
113
|
-
|
|
114
|
-
The default database survives hub restarts and is ignored by Git.
|
|
115
|
-
|
|
116
|
-
## 4. Bind this machine to the hub
|
|
117
|
-
|
|
118
|
-
```text
|
|
119
|
-
kxm hub bind http://127.0.0.1:7331
|
|
120
|
-
```
|
|
121
|
-
|
|
122
|
-
## 5. Confirm the session
|
|
123
|
-
|
|
124
|
-
```text
|
|
125
|
-
kxm session brief
|
|
126
|
-
```
|
|
127
|
-
|
|
128
|
-
## 6. Open Pi and check the hub
|
|
129
|
-
|
|
130
|
-
Set an identity and start the first agent. Do not give agents the administrative token.
|
|
131
|
-
|
|
132
|
-
PowerShell:
|
|
133
|
-
|
|
134
|
-
```powershell
|
|
135
|
-
$env:KXM_SERVER_URL = "http://127.0.0.1:7331"
|
|
136
|
-
$env:KXM_AUTH_TOKEN = "replace-with-a-demo-project-token"
|
|
137
|
-
$env:KXM_PROJECT = "demo"
|
|
138
|
-
$env:KXM_AGENT_NAME = "planner"
|
|
139
|
-
$env:KXM_AGENT_PURPOSE = "Plans work and coordinates handoffs"
|
|
140
|
-
pi
|
|
141
|
-
```
|
|
142
|
-
|
|
143
|
-
Bash:
|
|
144
|
-
|
|
145
|
-
```bash
|
|
146
|
-
export KXM_SERVER_URL=http://127.0.0.1:7331
|
|
147
|
-
export KXM_AUTH_TOKEN="replace-with-a-demo-project-token"
|
|
148
|
-
export KXM_PROJECT=demo
|
|
149
|
-
export KXM_AGENT_NAME=planner
|
|
150
|
-
export KXM_AGENT_PURPOSE="Plans work and coordinates handoffs"
|
|
151
|
-
pi
|
|
152
|
-
```
|
|
153
|
-
|
|
154
|
-
In Pi, run `/kxm hub`. It should show the connected identity and server.
|
|
155
|
-
|
|
156
|
-
Open a second terminal, repeat the settings, and change only the identity:
|
|
157
|
-
|
|
158
|
-
```powershell
|
|
159
|
-
$env:KXM_AGENT_NAME = "reviewer"
|
|
160
|
-
$env:KXM_AGENT_PURPOSE = "Reviews plans and code for correctness risks"
|
|
161
|
-
pi
|
|
162
|
-
```
|
|
163
|
-
|
|
164
|
-
Ask the planner:
|
|
165
|
-
|
|
166
|
-
```text
|
|
167
|
-
Use the kxm skill. List peers, ask reviewer to examine the current
|
|
168
|
-
plan for its three highest correctness risks, and wait for the response.
|
|
169
|
-
```
|
|
170
|
-
|
|
171
|
-
The planner should call `kxm_list`, `kxm_send`, and `kxm_await`. The reviewer receives an agent turn and its settled response returns to the planner.
|
|
172
|
-
|
|
173
|
-
For an executable transport-only demonstration, run `npm run example`. It starts a temporary in-memory hub, completes a planner-to-reviewer round trip, and exits without changing the normal database.
|
|
174
|
-
|
|
175
|
-
## Connect Claude Code
|
|
176
|
-
|
|
177
|
-
Keep the same hub running. Inside Claude Code, add the marketplace and install the plugin:
|
|
178
|
-
|
|
179
|
-
```text
|
|
180
|
-
/plugin marketplace add kontextmind/kxm
|
|
181
|
-
/plugin install kxm
|
|
182
|
-
/reload-plugins
|
|
183
|
-
```
|
|
184
|
-
|
|
185
|
-
Configure these values when prompted:
|
|
186
|
-
|
|
187
|
-
| Setting | Example |
|
|
188
|
-
|---|---|
|
|
189
|
-
| KXM server URL | `http://127.0.0.1:7331` |
|
|
190
|
-
| Authentication token | The `demo` project token, not the administrative token |
|
|
191
|
-
| Agent name | `claude-reviewer` |
|
|
192
|
-
| Agent purpose | `Reviews implementation and tests` |
|
|
193
|
-
| Project | `demo` |
|
|
194
|
-
|
|
195
|
-
Restart Claude Code after configuration. Ask it to use `kxm_list`; the connected Pi agents should appear.
|
|
196
|
-
|
|
197
|
-
### Optional pushed delivery
|
|
198
|
-
|
|
199
|
-
Claude channels can inject an inbound request into a running session. They are currently a research-preview feature, and a community channel must be explicitly trusted at launch:
|
|
200
|
-
|
|
201
|
-
```text
|
|
202
|
-
claude --dangerously-load-development-channels plugin:kxm
|
|
203
|
-
```
|
|
204
|
-
|
|
205
|
-
Review the trust prompt before accepting it. If an organization administrator has approved the plugin through `allowedChannelPlugins`, use:
|
|
206
|
-
|
|
207
|
-
```text
|
|
208
|
-
claude --channels plugin:kxm
|
|
209
|
-
```
|
|
210
|
-
|
|
211
|
-
Without channel mode, Claude can still send requests and receive them by calling `kxm_inbox`, then answer with `kxm_reply`.
|
|
212
|
-
|
|
213
|
-
## Your first useful topology
|
|
214
|
-
|
|
215
|
-
Start with two or three purposeful roles:
|
|
216
|
-
|
|
217
|
-
| Role | Good responsibilities |
|
|
218
|
-
|---|---|
|
|
219
|
-
| Planner | Break down work, define ownership, collect results |
|
|
220
|
-
| Builder | Implement one bounded change |
|
|
221
|
-
| Reviewer | Check correctness, tests, security, or documentation |
|
|
222
|
-
|
|
223
|
-
Avoid assigning two agents to edit the same files in one checkout. Use separate Git worktrees or give one agent write ownership.
|
|
224
|
-
|
|
225
|
-
## Next steps
|
|
226
|
-
|
|
227
|
-
- Use the wiki-ready [KXM Handbook](kxm-handbook.md) for the complete CLI, Pi, Claude, workflow, gate, and recovery reference.
|
|
228
|
-
- Adjust names, project isolation, and network settings in [Configuration](configuration.md).
|
|
229
|
-
- Learn the request lifecycle in [Architecture](architecture.md).
|
|
230
|
-
- Read [Operations](operations.md) before binding beyond localhost.
|
|
231
|
-
- Review the [Test matrix](test-matrix.md) for verified features and example coverage.
|
|
232
|
-
- Start a long-lived coordinator from Jira with [Webhook workflows](webhook-workflows.md).
|
|
233
|
-
- Use [Troubleshooting](troubleshooting.md) if an agent does not appear or a request does not arrive.
|
|
234
|
-
|
|
235
|
-
## Try the v0.5 context features
|
|
236
|
-
|
|
237
|
-
With a hub running (`kxm hub start`):
|
|
238
|
-
|
|
239
|
-
```bash
|
|
240
|
-
# Role-aware context packet for the current project
|
|
241
|
-
kxm context get my-project --role planner --task "plan the CI migration" --budget 8192
|
|
242
|
-
|
|
243
|
-
# Search durable context records (metadata only)
|
|
244
|
-
kxm context recall my-project --query "flaky"
|
|
245
|
-
|
|
246
|
-
# Explain evidence and lineage for one context item
|
|
247
|
-
kxm context explain my-project ctx_item_abc123
|
|
248
|
-
|
|
249
|
-
# Authoritative temporal state (and historical queries)
|
|
250
|
-
kxm context state my-project ci.pipeline
|
|
251
|
-
kxm context state my-project ci.pipeline --as-of 2026-01-15T00:00:00.000Z
|
|
252
|
-
|
|
253
|
-
# Episodic learning from workflow journals
|
|
254
|
-
kxm context episode my-project
|
|
255
|
-
|
|
256
|
-
# Compile and lint the knowledge wiki
|
|
257
|
-
kxm context wiki-compile my-project
|
|
258
|
-
kxm context wiki-lint my-project
|
|
259
|
-
|
|
260
|
-
# Routing telemetry per behavioral configuration
|
|
261
|
-
kxm routing report
|
|
262
|
-
```
|
|
263
|
-
|
|
264
|
-
State changes follow a propose-then-promote flow: agents propose through the
|
|
265
|
-
`kxm_promote` Pi/MCP tool (or the context API), and an operator promotes with
|
|
266
|
-
durable evidence:
|
|
267
|
-
|
|
268
|
-
```bash
|
|
269
|
-
kxm context promote my-project ctx_prop_abc123 --evidence "receipt:run_9/verify"
|
|
270
|
-
```
|
|
271
|
-
|
|
272
|
-
Role-aware packets differ by role: repro agents see prior reproductions and
|
|
273
|
-
incidents; planners see current state and decisions; critics see
|
|
274
|
-
contradictions; implementers see the approved plan and skills; verifiers see
|
|
275
|
-
acceptance evidence. The same requests through Pi (`kxm_context`) or Claude
|
|
276
|
-
Code (MCP) return the same packets — agents never talk to a memory backend
|
|
277
|
-
directly.
|