@joekytc/dsh-swarm 0.3.7 → 0.3.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -19,10 +19,10 @@ dsh-swarm is a DSH plugin that turns one requirement into a strict, evidence-ver
19
19
  Swarm mode turns your main session into a **team lead**: **you state the requirement, it clarifies, plans, confirms, delegates and follows through** — entirely in natural language, no commands to remember.
20
20
 
21
21
  - **No commands to memorize** — just state your requirement; no `/plan:` or `/openspec:` prefixes needed.
22
- - **Automatic intent recognition** — development requests → clarify/plan and build a chain; lessons & retrospectives → persist to memory; group notifications → deliver to WeCom; Q&A / chit-chat → answered directly.
23
- - **Free delivery** — `/sms <intent>` (e.g. "post current progress to the group"): facts are grounded via kanban lookup, then the body is composed per intent and delivered; `-s` or wording like "private chat" targets the DM. Group and private-chat targets must each be exactly one (0 or 2+ targets error out; clean up in dsh-im settings or set `imDelivery.dmTargetId`).
24
- - **Confirmation gate against accidental chains** — after the checklist is saved, a chain is only built once you reply with an explicit affirmative (`确认` / `开干` / `开跑` / `开始` / `go`, etc.); vague replies, topic switches, or edit-only feedback count as *not confirmed*.
25
- - **The lead is read-only** — the main session cannot write/edit repo sources, nor run git mutations (push/commit/checkout…); writing code is done by the executor (D) in an isolated workspace by design.
22
+ - **Automatic intent recognition** — development requests → clarify/plan and build a chain; lessons & retrospectives → persist to memory; group notifications → deliver to WeCom; Q&A / chit-chat → answered directly. Intent is judged by the model, not by a code-level classifier — the confirmation gate below is what stops a misjudged chain.
23
+ - **Free delivery** — `/sms <intent>` (e.g. "post current progress to the group"): facts are grounded via kanban lookup, then the body is composed per intent and delivered; `-s` or wording like "private chat" targets the DM. A bare `/sms` re-sends the latest completed-chain report and `/sms blocked [chainId]` the block notice — those two bodies are rendered by system code from kanban facts, never rewritten by the lead. Group and private-chat targets auto-resolve to the single saved target (0 or 2+ targets error out; clean up in dsh-im settings, or pin `imDelivery.targetId` / `imDelivery.dmTargetId`, which skips the count check — though with `imDelivery.botId` empty the pinned target must still belong to the auto-discovered bot, or it errors out).
24
+ - **Confirmation gate against accidental chains** — after the checklist is saved, the lead is instructed to build a chain only once you reply with an explicit affirmative (`确认` / `开干` / `开跑` / `开始` / `go`, etc.). The gate is judged by the model, not enforced by a code-level check; vague replies, topic switches, or edit-only feedback count as *not confirmed*.
25
+ - **The lead is read-only** — the main session cannot write/edit repo sources, nor run git mutations (push/commit/reset…); a bare `git checkout`/`git switch` of an existing branch is allowed. Writing code is done by the executor (D) in an isolated workspace by design.
26
26
  - **Progress is always actually queried** — ask "how is it going?" anytime and the lead reports from real kanban lookups, never fabricated.
27
27
 
28
28
  ### Why it's designed this way
@@ -33,7 +33,7 @@ Coordinating several agents on one task typically fails in three ways:
33
33
  - **Unverifiable handoffs** — an agent claims "done" with no reproducible evidence, and the next agent builds on sand.
34
34
  - **Silent deadlocks** — an agent stops without finishing and the pipeline hangs, or bad code is merged before anyone reviewed it.
35
35
 
36
- dsh-swarm encodes a *contract* against all three: one machine-enforced responsibility per role;
36
+ dsh-swarm encodes a *contract* against all three: one responsibility per role, enforced by the permission matrix, tool faces and task-body instructions;
37
37
  every handoff must carry structured evidence or the phase will not close; every stall or review
38
38
  failure lands in a visible, recoverable state — with you (the human) as the final trust anchor.
39
39
  It is built correctness-first: deterministic state machines, append-only event sourcing, idempotent
@@ -55,13 +55,21 @@ schedulers, and a red-team test suite that replays the event log and rejects any
55
55
 
56
56
  Prerequisites: a working DSH runtime (`@deepseek-ai/*`), Node.js ≥ 22.19 and npm. Optional: a wiki-vault HTTP service (KB features, see [Configuration](#configuration)).
57
57
 
58
+ From npm (the published tarball ships the built `lib/`):
59
+
60
+ ```bash
61
+ dsh plugin --profile web add @joekytc/dsh-swarm
62
+ ```
63
+
64
+ From a source checkout (rebuild first so `lib/` matches the sources):
65
+
58
66
  ```bash
59
67
  npm install
60
68
  npm run build # tsc -p tsconfig.build.json + client bundle (lib/client.js)
61
- dsh plugin --profile web add @joekytc/dsh-swarm
69
+ dsh plugin --profile web add .
62
70
  ```
63
71
 
64
- > From GitHub source: `dsh plugin --profile web add github:joekytc/dsh-swarm`
72
+ > From GitHub source: `dsh plugin --profile web add github:joekytc/dsh-swarm` — the repository tracks the built `lib/`.
65
73
 
66
74
  ### 2. Switch your main-session preset
67
75
 
@@ -91,14 +99,14 @@ Lead: Chain created (ch_…), live progress on the kanban tab (Conversation →
91
99
  ```
92
100
 
93
101
  - **Kanban**: the third tab of the conversation center (Conversation → Trajectory → Kanban). Click a card for Overview / Trajectory / Handoff / Spec / Comments.
94
- - **Completion**: when a chain completes, the system audits the workspace and (for D chains) automatically merges the feature branch into the spec-declared target branch; if an audit warning fires, confirm ownership in the GUI first.
102
+ - **Completion**: when a chain completes, the system audits the workspace and (for D chains) automatically merges the feature branch into the spec-declared target branch. An audit warning blocks the final wrap-up until you confirm ownership in the GUI — it does not gate the merge.
95
103
  - **Progress**: just ask "how is it going?" — the lead reports from real kanban lookups and relays blocking reasons faithfully.
96
104
 
97
105
  ---
98
106
 
99
107
  ## What it does for you
100
108
 
101
- Six roles, one job each, machine-enforced boundaries — no role creep:
109
+ Six roles, one job each — boundaries enforced by the permission matrix, trimmed tool faces and task-body instructions, so no role creep:
102
110
 
103
111
  | Role | One-line responsibility | What it never does |
104
112
  |---|---|---|
@@ -129,20 +137,40 @@ All keys are optional; schema lives in `src/config.ts`. **Most users only need t
129
137
  |---|---|---|
130
138
  | `storageDir` | `$DSH_HOME/storages/kanban` | Event log (`events.jsonl`), orchestration state, per-task workspaces, `dispatcher.log`. Value must use the unquoted `!!js dshHomePath("storages/kanban")` form — quoting degrades it into a literal string |
131
139
  | `wikiVault.baseUrl` | `''` (empty) | wiki-vault HTTP service for KB reads/writes — required for KB features; set to your own server |
140
+ | `wikiVault.pagePrefix` | `projects/` | Namespace prefix for generated wiki pages |
132
141
  | `roles.models.<role>` | `{}` | Per-role model: `{ provider, model, reasoningEffort?, fallbacks?[] }` |
133
142
  | `roles.models.<role>.reasoningEffort` | `high` | Default reasoning effort for all roles |
134
143
  | `roles.models.<role>.fallbacks` | `[]` | Silent fallback candidates (audited via `[model-fallback]` comment) |
135
144
  | `dispatcher.staleTimeoutSeconds` | `14400` | Heartbeat timeout; running task without heartbeat is reclaimed |
136
145
  | `dispatcher.maxRetries` | `3` | Failure retries before circuit → `blocked(gave_up)` |
137
146
  | `dispatcher.heartbeatIntervalSeconds` | `300` | Watchdog heartbeat period |
138
- | `dispatcher.maxProtocolViolations` | `2` | Protocol-violation guardrail: after this many consecutive violations the next one is final (`gave_up`) |
147
+ | `dispatcher.maxProtocolViolations` | `2` | Protocol-violation guardrail: once consecutive violations reach this many, the next one is final (`gave_up`) |
139
148
  | `dispatcher.maxReworksPerRole` | `{ pt: 3, dt: 3 }` | Max review rework rounds before `review/gave-up` + `[review-final]` |
140
- | `prefixRoutes.plan` | `/plan:` | Command-mode planning prefix |
149
+ | `prefixRoutes.plan` | `/plan:` | Command-mode phase-0 planning prefix |
141
150
  | `prefixRoutes.openspec` | `/openspec:` | Command-mode approve-and-execute prefix |
142
- | `ui.enabled` | `true` | Enable the kanban web tab |
143
- | `ui.contentMinWidth` | `715` | Minimum kanban content width (px) |
144
- | `ui.contentMaxWidth` | `780` | Maximum kanban content width (px) |
151
+ | `prefixRoutes.learning` | `/learning` | Lessons / retrospective prefix |
152
+ | `prefixRoutes.send` | `/sms` | Free-delivery prefix (DM with `-s`) |
153
+ | `memory.enabled` | `true` | Memory recall index; `false` makes `planning_memory_recall` return a disabled notice |
154
+ | `memory.maxIndexEntries` | `8` | Max recalled memory entries (1–20) |
155
+ | `ui.enabled` | `true` | Declared switch; not consumed yet — the tab registers unconditionally |
156
+ | `ui.contentMinWidth` | `715` | Declared lower width bound (px); not consumed by the client yet — the tab follows the host conversation width |
157
+ | `ui.contentMaxWidth` | `780` | Declared upper width bound (px); not consumed by the client yet |
145
158
  | `ui.sseHeartbeatSeconds` | `20` | SSE heartbeat interval |
159
+ | `gates.enabled` | `true` | TDD measurement gate; `false` → silent skip (no event) |
160
+ | `gates.timeoutMs` | `600000` | Per-command gate timeout (ms), SIGKILL at the deadline |
161
+ | `gates.forbidden` | `['rm -rf /', 'git push']` | Command-blacklist substrings (defence in depth) |
162
+ | `evidenceReplay.enabled` | `false` | L3 replay of model-written commands — not a sandbox, see [Issue evidence check](#issue-evidence-check-pr2-opt-in) |
163
+ | `evidenceReplay.timeoutMs` | `600000` | Per-replay command timeout (ms) |
164
+ | `evidenceReplay.allowPrefixes` | `['npx --no-install vitest', 'npm test', 'npm run build', 'npm run typecheck', 'tsc', 'eslint']` | Allowlisted tool prefixes (word-boundary match) |
165
+ | `imDelivery.enabled` | `false` | WeCom delivery via dsh-im (W3 wrap-up / chain blocked / review gave-up) |
166
+ | `imDelivery.botId` | `''` | Empty = auto-discover the only wecom bot |
167
+ | `imDelivery.targetId` | `''` | Empty = auto-discover the only saved group target |
168
+ | `imDelivery.dmTargetId` | `''` | DM target for `/sms -s` |
169
+ | `imDelivery.fallbackBotId` | `''` | Bot used when preset-based matching finds nothing |
170
+ | `reviewEngine.mode` | `delegate` | `delegate` or `managed` — see [Review engine (ocr)](#review-engine-ocr) |
171
+ | `reviewEngine.managed.provider` | `''` | Model-chain provider id used in managed mode |
172
+ | `reviewEngine.managed.model` | `''` | Model-chain model id used in managed mode |
173
+ | `wikiWritePresets` | `['swarm', 'kanban-w', 'ptc']` | Presets allowed to call `wiki_write` (page paths are still whitelisted by namespace) |
146
174
 
147
175
  ---
148
176
 
@@ -150,7 +178,7 @@ All keys are optional; schema lives in `src/config.ts`. **Most users only need t
150
178
 
151
179
  Implementation reviews (the in-chain DT phase and standalone reviews) are powered by
152
180
  [open-code-review](https://open-codereview.ai) (ocr), with two modes switchable in the
153
- web config panel under "Swarm config → Review engine (ocr)":
181
+ web config panel under 「Swarm 配置 → 评审引擎(ocr)」 (Swarm config → Review engine (ocr)):
154
182
 
155
183
  | Mode | How it works | Notes |
156
184
  |---|---|---|
@@ -159,21 +187,23 @@ web config panel under "Swarm config → Review engine (ocr)":
159
187
 
160
188
  ### Install
161
189
 
162
- - When ocr is missing, the config panel shows a red banner — click "Install ocr" for a one-click global install (async, cancellable);
190
+ - When ocr is missing, the config panel shows a red banner — click 「安装 ocr」 (Install ocr) for a one-click global install (async, cancellable);
163
191
  - or run `npm install -g @alibaba-group/open-code-review` in a terminal, then verify with `ocr --version`.
164
192
 
165
193
  ### Standalone review (no chain needed)
166
194
 
167
- 1. Switch the session to "Delivery Reviewer (DT)" at the top of the dsh web UI and just talk;
195
+ 1. Switch the session to 「交付评审官(DT)」 (Delivery Reviewer (DT)) at the top of the dsh web UI and just talk;
168
196
  2. State the review target: a local directory / branch range (from…to) / a single commit / uncommitted workspace diff / a public repo URL (auto-cloned into a temp dir, discarded afterwards);
169
197
  3. The report is first fully output to the conversation;
170
- 4. Only after you confirm is it written to the wiki at `projects/<repo>/reviews/<topic>-<date>/`. Read-only throughout — reviewed code is never modified.
198
+ 4. Only after you confirm is it written to the KB at `projects/<repo>/reviews/<topic>-<date>/` — remote KB mode only; with the default empty `wikiVault.baseUrl` (local mode) no wiki page is written.
199
+
200
+ Read-only in practice: bash/run_code writes and wiki writes outside the reviews namespace are blocked by the guard, while the fs `write`/`edit` tools are not blocked in a standalone DT session.
171
201
 
172
202
  ### Configuration notes
173
203
 
174
- - Mode, provider and model are all chosen on the "Review engine (ocr)" card; the provider/model dropdowns share the same catalog as the model chain;
175
- - After picking, click "Apply to ocr" — the system writes the wiring into ocr's custom config (`dsh-managed`); the API key is resolved from the dsh model config and written into ocr, never shown in plain text in the panel; if resolution fails it degrades gracefully and points you to a manual `ocr config provider` in a terminal;
176
- - When managed is not ready, reviews silently fall back to delegate mode — nothing is blocked.
204
+ - Mode, provider and model are all chosen on the 「评审引擎(ocr)」 card; the provider/model dropdowns share the same catalog as the model chain;
205
+ - After picking, click 「应用到 ocr」 (Apply to ocr) — the system writes the wiring into ocr's custom config (`dsh-managed`); the API key is resolved from the dsh model config and written into ocr, never shown in plain text in the panel; if resolution fails it degrades gracefully and points you to a manual `ocr config provider` in a terminal;
206
+ - When managed is not ready, the ocr tool refuses the managed call and returns delegate guidance; the reviewer switches to delegate mode — nothing is blocked.
177
207
 
178
208
  Official docs: [Installation](https://open-codereview.ai/docs/installation) · [Model configuration](https://open-codereview.ai/docs/configuration) · [Delegate mode](https://open-codereview.ai/docs/delegate)
179
209
 
@@ -182,9 +212,10 @@ Official docs: [Installation](https://open-codereview.ai/docs/installation) · [
182
212
  ## Trust & guardrails (user's view)
183
213
 
184
214
  - **Read-only hard gate for the lead** — in swarm mode, main-session writes to sources and git mutations are blocked by a system gate; if blocked, just let the lead explain — execution is done by the D role.
185
- - **Confirmation gate** — no chain is ever built without your explicit confirmation.
215
+ - **Confirmation gate** — the lead only builds a chain after you reply with an explicit affirmative (enforced by the lead's instructions, not by a code-level check).
186
216
  - **TDD hard gate** — implementations must ship with tests (or an explained skip); reviews machine-verify "tests really ran, and were written first".
187
- - **Human trust anchors** — spec approval, unblock, audit confirmation and chain deletion are human-only; neither the main session nor role agents can create chains or approve specs.
217
+ - **Human trust anchors** — spec approval, unblock, audit confirmation and chain deletion are human-only; role agents cannot approve specs, and a chain is only created from your confirmed routing call (the main session routes as `human`).
218
+ - **Guardrails are constraints, not a sandbox** — PT/DT write guards rely on path and command regexes (reviewers get no git credentials), and review evidence is existence-checked: fields must be present and well-formed, while replaying the commands to prove they ran happens only when `evidenceReplay.enabled` is turned on.
188
219
  - Full mechanics (permission matrix, delivery contract, review chain, rework, failure recovery) live under [Advanced / Developers](#advanced--developers).
189
220
 
190
221
  ---
@@ -195,16 +226,16 @@ Official docs: [Installation](https://open-codereview.ai/docs/installation) · [
195
226
 
196
227
  ### Roles & the execution pipeline (full table)
197
228
 
198
- Six roles are dispatched by the scheduler as one-shot agent sessions (deterministic session id `kbn-<taskId>`, resumed on retry/rework via `resumeSessionId`). Each role-agent session is bound to exactly one task (`boundTaskId`) and gets a trimmed tool face. V is the exception: a chain-scoped orchestrator session (`kbn-v-<chainId>`) with no `boundTaskId`.
229
+ Six roles are dispatched by the scheduler as one-shot agent sessions (deterministic session id `kbn-<taskId>`; a retry resumes that same session — `resumeSessionId` is always null in the current implementation — while a rework task starts its own `kbn-<reworkId>` session). Each role-agent session is bound to exactly one task (`boundTaskId`) and gets a trimmed tool face. V is the exception: a chain-scoped orchestrator session (`kbn-v-<chainId>`) with no `boundTaskId`.
199
230
 
200
231
  | Role | Alias | Responsibility | Tool face (highlights) |
201
232
  |---|---|---|---|
202
233
  | **V** | Orchestrator | Drives the phase machine, creates one card per phase, posts `[blocked-review]` guidance on stalls. Never executes. | `kanban_create` + task tools + spec view |
203
234
  | **P** | Planner | Reads spec + repo facts (incl. read-only self-checks), writes an OpenSpec implementation plan, opts into PT via `pt_decision.needed`. Never executes. | Task tools + spec view, read-only (writes only `openspec/changes/`) |
204
235
  | **PT** | Plan reviewer | Read-only review of P's plan (requirements alignment, completeness, logic). Outputs verdict + issues. | Task tools + spec view, **read-only ToolGuard** |
205
- | **W** | Knowledge officer | W2/W3 KB sync (`w:kb`). Never touches code/git. | Task tools + `wiki_search/read/write` (remote) / `skill`→llm-wiki (local) + read-only spec view |
236
+ | **W** | Knowledge officer | W2/W3 KB sync (`w:kb`). Never touches code/git. | Task tools + `wiki_search/read/write` (remote) / `skill`→llm-wiki (local) + `prefetch_file`/`prefetch_external`/`prefetch_kb` + read-only spec view |
206
237
  | **D** | Executor | The *only* role that writes code: worktree → implement → verify → `[AI-GEN]` commit → push feature branch (merging into the spec-declared target branch is done by the system only after DT passes). | Task tools + wiki read + bash/fs/run_code (full dev) + subagent (spawn/fork/list-agents) + goal |
207
- | **DT** | Implementation reviewer | Empirically verifies D's work (test/build/typecheck/diff/git + open-code-review), writes review page to KB. Read-only against the repo. | Task tools + wiki read/write (review namespace) + bash/fs/run_code, **read-only ToolGuard** |
238
+ | **DT** | Implementation reviewer | Empirically verifies D's work (test/build/typecheck/diff/git + open-code-review), writes review page to KB. Read-only against the repo. | Task tools + wiki read/write (review namespace) + `ocr_review` + bash/fs/run_code, **read-only ToolGuard** |
208
239
 
209
240
  ### Guardrails in detail
210
241
 
@@ -234,41 +265,64 @@ Six roles are dispatched by the scheduler as one-shot agent sessions (determinis
234
265
  | prefetch | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
235
266
  | audit-confirm | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ |
236
267
  | create-rework-task | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅ |
268
+ | reopen-chain | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ |
269
+ | waive-review | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ |
237
270
 
238
271
  Key guarantees (two):
239
272
 
240
- - **The main session cannot execute.** It only gets `kanban_show`/`kanban_list`/
241
- `kanban_comment` + `spec_card_view` + `kanban_route` — never
242
- `kanban_create`/`kanban_complete`/`kanban_block`. Chains/specs are created only
243
- via swarm-mode intents or `/plan:`+`/openspec:`; the GUI observes and mutates task state but never
244
- creates chains or tasks — "who decided to run what" stays explicit and auditable.
273
+ - **The main session cannot execute.** Its kanban tool face is a read-only subset —
274
+ `kanban_show`/`kanban_chain`/`kanban_list`/`kanban_comment` plus the human-recovery pair
275
+ `kanban_reopen_chain`/`kanban_waive_review` — together with `spec_card_view`, `kanban_route`,
276
+ the `planning_*` tools and `sms_send`; never `kanban_create`/`kanban_complete`/`kanban_block`.
277
+ Chains/specs are created only via swarm-mode intents or `/plan:`+`/openspec:`; the GUI observes
278
+ and mutates task state but never creates chains or tasks — "who decided to run what" stays
279
+ explicit and auditable.
245
280
  - **Session binding prevents cross-task escalation** (a W agent bound to task A
246
281
  cannot complete/block task B even though both are W tasks); DT writes are
247
282
  confined to the `projects/<repoSlug>/<chain>/review/` namespace by a ToolGuard on top of
248
- the matrix; and no role agent can approve specs, unblock, or confirm audits —
249
- those are human trust anchors; `system` handles only mechanical bookkeeping.
283
+ the matrix; and no role agent can approve specs, unblock, confirm audits, waive a
284
+ review or reopen a chain — those are human trust anchors; `system` handles only
285
+ mechanical bookkeeping.
250
286
 
251
287
  #### Delivery contract (upstream owes downstream)
252
288
 
253
289
  Each phase's handoff must carry the keys its downstream actually reads
254
- (`src/domain/delivery-contract.ts`). Missing keys block the current role's card
255
- immediately (and the orchestrator never builds a downstream card on a blocked
256
- parent):
290
+ (`src/domain/delivery-contract.ts`). Missing delivery keys block the W/P card
291
+ immediately and admit no human exemption; the D/PT/DT evidence gates instead reject
292
+ `complete` with an error (the card stays running) and are human-exempt. The
293
+ orchestrator never builds a downstream card on a blocked parent:
257
294
 
258
295
  | Card | Required handoff keys |
259
296
  |---|---|
260
- | W2 / W3 (`w:kb`) | `kb_url` + `page_path` |
297
+ | W2 / W3 (`w:kb`) | `kb_url` + `page_path` — non-empty is not enough: with a configured `wikiVault.baseUrl` the `kb_url` must start with it (and must be exactly `''` in local mode), while `page_path` must sit in the allowlisted namespaces (`wiki/**` locally) |
261
298
  | P (`p:openspec`) | `artifacts_path` + `pt_decision` (`needed` boolean required; when `needed: true`, `reason` is required) |
262
- | D (`d:execute`) | `changed_files` + (`commit_hash` or `push`) — `hasDeliveryEvidence`; `branch` (feature branch) is expected for the merge gate, not a hard-complete blocker; `tdd` (`test_files` or `skipped.reason`, XOR) |
299
+ | D (`d:execute`) | `changed_files` + (`commit_hash` or `push`) — `hasDeliveryEvidence`; `branch` (feature branch) is not a delivery key, but when the TDD gate actually runs the tests it checks the branch and bounces the card if it is absent or mismatched; `tdd` (`test_files` or `skipped.reason`, XOR) |
263
300
  | PT / DT | `review_evidence` (schema-valid) — `validateReviewEvidence` |
264
301
 
265
302
  #### TDD hard gate (evidence threshold)
266
303
 
267
304
  D completes only with `tdd` — `test_files` (with `test_first`) or `skipped.reason`
268
- (XOR, `delivery-evidence.ts`). DT's `review_evidence` must carry `tdd`; on a
269
- `pass` verdict the runner must be `vitest` (`test.runner`) and `test_first === true`
270
- must hold (`review-evidence.ts`). This makes "tests actually ran, and were written
271
- first" a machine-checked property rather than a claim.
305
+ (XOR, `delivery-evidence.ts`). DT's `review_evidence` must carry `tdd`; whenever
306
+ `tdd.test_files` are declared the runner must be `vitest` (`test.runner`), and on a
307
+ `pass` verdict `test_first === true` must hold (`review-evidence.ts`). This makes
308
+ "tests actually ran, and were written first" a machine-checked property rather than
309
+ a claim.
310
+
311
+ #### Gate skip alarm & declaration cross-check
312
+
313
+ A declaration that disagrees with the diff no longer passes the gate silently
314
+ (three silent-skip modes remain, emitting no event: `gates.enabled=false`, a handoff
315
+ without `worktree_dir`, and non-`d:execute` cards):
316
+
317
+ - **`task/gate-skipped` event** — a declared `tdd.skipped` is allowed when the diff is pure docs/config, or when the diff cannot be computed (conservative alarm-skip); both leave the event as an audit trail. Otherwise the gate **bounces** the card back (same session, agent fixes and re-completes).
318
+ - **Declaration ↔ reality cross-check** — declaring `test_files` with no test-file change in the diff (stale-test handoff), an empty `test_files`, an invalid path, or a branch mismatch all **bounce** instead of silently skipping.
319
+ - After 3 cumulative bounces the task is **blocked for human review** (`gave_up: gate bounced 3 times`).
320
+ - Gate runs now archive raw output to `<storageDir>/gate-logs/<taskId>.log` (path is referenced in the gate event detail).
321
+ - `review_evidence.lint` must be a structured object (aligning with `build`/`typecheck`); `null`/scalars are rejected.
322
+
323
+ #### Issue evidence check (PR2, opt-in)
324
+
325
+ Reviewer-reported issues (DT cards only — PT issues are never evidence-checked) can carry `evidence = { file, command, exit }` (raw output archive, first line `[exit code: N]`). Verification is three-tier: missing evidence → flagged (`not-provided`; critical/high → `could-not-replay` + needs-human); archive paper-check (zero execution) → `matches`; on mismatch a **replay** re-runs the command — **only if `evidenceReplay.enabled` is turned on**, and only for allowlisted tool prefixes (word-boundary match; `npm run` restricted to fixed script names). Replays execute model-written commands and are **NOT a sandbox**; keep the switch off unless you accept that risk. Results are summarized in a `review/evidence-check` event; `differs` never auto-fails a review — it surfaces to human review.
272
326
 
273
327
  #### Phase-0 planning checklist
274
328
 
@@ -280,9 +334,9 @@ blocks the save, and chain creation mounts the checklist as the `file-prefetch`
280
334
 
281
335
  #### Review quality chain
282
336
 
283
- - After **P** completes, **PT** is created only when P's handoff delivers
284
- `pt_decision.needed = true`; the orchestrator never overrides the decision
285
- (V only creates the card).
337
+ - After **P** completes, **PT** is skipped only when P's handoff delivers
338
+ `pt_decision.needed = false`; `true` or an absent decision creates the PT card
339
+ (fail-safe) — the orchestrator never overrides the decision (V only creates the card).
286
340
  - After **D** completes, a **DT** card is *always* created.
287
341
  - **PT/DT are read-only**: a ToolGuard mechanically denies writes to the repo
288
342
  sources, git mutations, and (for DT) wiki writes outside the review namespace.
@@ -291,9 +345,10 @@ blocks the save, and chain creation mounts the checklist as the `file-prefetch`
291
345
  `review-tool-unavailable` when ocr is missing (reason notes GUI install),
292
346
  without burning retries.
293
347
  - `review_evidence` must pass `validateReviewEvidence` or the review card cannot
294
- complete: PT needs verdict + issues + plan ref; DT additionally needs
295
- test (exit 0 on pass), build/typecheck, lint, non-empty diff, git,
296
- ocr/fallback conclusion, and `tdd`.
348
+ complete: PT needs verdict + issues + plan ref, and a `fail` verdict requires at
349
+ least one unresolved `critical`/`high` issue (otherwise it must be `pass`); DT
350
+ additionally needs test (exit 0 on pass), build/typecheck, lint, non-empty diff,
351
+ git, ocr/fallback conclusion, and `tdd`.
297
352
 
298
353
  #### Rework (review failure)
299
354
 
@@ -304,10 +359,11 @@ A failed review never mutates a `done` card. Instead the system records
304
359
  previous round's issues verbatim in a `## 本轮修复清单` (fix-this-round) section.
305
360
  A fresh review card is then dispatched for the rework.
306
361
  When `reviewAttempt` reaches `maxReworksPerRole` (PT 3 / DT 3), the system records
307
- `review/gave-up` and posts a `[review-final]` evidence-chain comment, and (with IM
362
+ `review/gave-up` and posts a `[review-final]` evidence-chain comment (the same marker
363
+ is reused by the convergence-gate downgrade below), and (with IM
308
364
  delivery enabled) sends a `[评审超限待裁决]` notification with two exit paths:
309
365
  waive the review (`kanban_waive_review` / GUI “豁免评审”) or reopen the chain
310
- (`kanban_reopen_chain` / GUI “人工恢复”). **Convergence gate (2026-09-15)**: on PT
366
+ (`kanban_reopen_chain` / GUI “人工恢复”). **Convergence gate**: on PT
311
367
  rework rounds, once all legacy issues are resolved and no unresolved `critical`
312
368
  remains, a `fail` verdict is downgraded to `pass` (new findings flow downstream as
313
369
  non-blocking suggestions), so the loop always converges.
@@ -330,17 +386,19 @@ Two orthogonal failure paths, both human-recoverable:
330
386
  model candidates (primary + fallbacks, `reasoningEffort: high` default) fall
331
387
  back silently (audited via `[model-fallback]` comment); if *all* candidates fail
332
388
  it blocks `model-unavailable` for the human. A single hanging V wake cannot
333
- stall the scheduler — every dispatch is wrapped in a timeout.
389
+ stall the scheduler — V wakes are wrapped in a 60 s timeout (role-task dispatch
390
+ still awaits the session's `whenIdle`).
334
391
 
335
392
  #### Chain completion: audit gate + merge gate
336
393
 
337
394
  When the mechanical chain-complete rule fires, two gates run in the
338
395
  `chain/completed` hook:
339
396
 
340
- 1. **Completion audit gate**: the `ChainAuditor` cross-checks the chain
341
- workspace for artifacts written outside the known task outputs. Orphaned writes
342
- emit `chain/audit-warning`; the UI shows a warning banner and blocks the final
343
- summary until the human confirms ownership (`chain/audit-confirmed`, human-only).
397
+ 1. **Completion audit gate**: the `ChainAuditor` scans live non-role sessions for
398
+ write-capable tool calls aimed at the chain workspace (`kanban-*` preset subagents
399
+ exempt), then reconciles artifact ownership as a mechanical fallback. Orphaned
400
+ writes emit `chain/audit-warning`; the UI shows a warning banner and blocks the
401
+ final summary until the human confirms ownership (`chain/audit-confirmed`, human-only).
344
402
  2. **Merge gate (post-DT system merge)**: D never merges to the target branch and
345
403
  never pushes it — it only commits to (and optionally pushes) its feature branch,
346
404
  carrying `branch` in its handoff. The target branch is the one declared in the
@@ -349,33 +407,38 @@ When the mechanical chain-complete rule fires, two gates run in the
349
407
  → git merge --no-ff <feature-branch> → git push`. Outcomes are recorded as
350
408
  idempotent comments: `[merge-done]` (with hash), `[merge-skip]` (merge input
351
409
  unresolvable), or `[merge-failed]` (checkout/merge/push failed, e.g. a conflict).
352
- Failures never throw — a bad merge is never performed, which is the safe
353
- direction; humans can repair afterwards.
410
+ Failures never throw — the gate only records `[merge-failed]`; note that a failed
411
+ push can leave the target branch already merged locally, and a conflicting merge
412
+ leaves the worktree mid-merge. Repairing is a human job.
354
413
 
355
414
  ### Event sourcing & domain model
356
415
 
357
416
  Every state change is appended to `<storageDir>/events.jsonl`, one JSON event per
358
- line. The `seq` is assigned by the store (re-read from the file tail on every
359
- append, so concurrent instances never collide). The **trajectory is the event log
360
- itself**; restart replays it to rebuild the board.
417
+ line. The `seq` is assigned by the store (re-read from the file tail on every append
418
+ as the last line's `seq` + 1) — append-only, except that a human chain delete purges
419
+ that chain's lines and renumbers the remaining `seq` values. The **trajectory is the
420
+ event log itself**; restart replays it to rebuild the board.
361
421
 
362
422
  ```jsonc
363
423
  // one line in events.jsonl
364
424
  { "seq": 12, "chainId": "ch_x_...", "taskId": "t_y_...",
365
425
  "kind": "task/completed",
366
- "payload": { "summary": "...", "metadata": { /* handoff evidence */ } },
426
+ "payload": { "summary": "...", "metadata": { /* handoff evidence */ }, "completedAt": 1760000000000 },
367
427
  "author": "w", "at": 1760000000000 }
368
428
  ```
369
429
 
370
- Event families: `chain/*` (created, executing, completed, aborted, root-task-set,
371
- audit-warning, audit-confirmed, title-updated), `spec-card/*` (created, edited,
372
- approved), `task/*` (created, claimed, heartbeat, commented, completed, blocked,
373
- unblocked, failed, archived, renamed), and `review/*` (passed, failed, gave-up).
430
+ Event families actually emitted: `chain/*` (created, executing, completed, blocked,
431
+ reopened, root-task-set, audit-warning, audit-confirmed, title-updated,
432
+ im-delivery-failed), `spec-card/*` (created, edited, approved), `task/*` (created,
433
+ claimed, heartbeat, commented, completed, blocked, unblocked, failed, archived,
434
+ renamed, gate-passed, gate-failed, gate-skipped), and `review/*` (passed, failed,
435
+ gave-up, waived, evidence-check).
374
436
 
375
- Replay is **strict**: the projection applies every event through the state machine
376
- and throws on any illegal transition, so a corrupted or tampered log fails loudly
377
- instead of silently producing an inconsistent board (covered by
378
- `tests/redteam/anti-escalation.test.ts` and `tests/domain/projection.test.ts`).
437
+ Replay is **strict**: the projection replays every transition event through the state
438
+ machine and throws on any illegal transition, so a corrupted or tampered log fails
439
+ loudly instead of silently producing an inconsistent board (non-transition kinds such
440
+ as `task/commented` are recorded as no-ops). Covered by
441
+ `tests/redteam/anti-escalation.test.ts` and `tests/domain/projection.test.ts`.
379
442
 
380
443
  The service emits events through a serialized queue (append-then-publish), and
381
444
  subscribers (SSE) receive every event exactly once in order. UI and dispatcher both
@@ -383,36 +446,38 @@ consume the same persisted events — there is no secondary source of truth.
383
446
 
384
447
  ### Web client (Workflow kanban tab)
385
448
 
386
- A browser-half React tab registered as the third `conversation.view` slot
387
- (`id=kanban`, `order=20`, after Conversation and Trajectory). It registers **no
449
+ A browser-half React tab registered into `conversation.view` (`id=kanban`,
450
+ `order=20`, so it sits after Conversation and Trajectory). It registers **no
388
451
  shell-level overlays, sidebars, or detail panes**.
389
452
 
390
453
  - **Data path**: initial snapshot (`GET /kanban/board`) → SSE stream
391
454
  (`GET /kanban/events?after=<seq>`) → board-store applies events incrementally,
392
455
  deduplicates by `seq`, and re-pulls the full snapshot on any gap. **No business
393
456
  polling.**
394
- - **Layout**: multi-chain vertical rails; fixed content width 715–780 px, full
395
- height; the active chain is expanded, blocked chains always show a warning
396
- summary. In-page rename/delete use a lightweight modal (no shell overlays);
397
- no drag-and-drop, no width memory.
457
+ - **Layout**: multi-chain vertical rails; the width follows the host conversation
458
+ width (`--dsh-chat-content-width`, 780 px fallback), full height; the active chain
459
+ is expanded, blocked chains always show a warning summary. In-page rename/delete use
460
+ a lightweight modal (no shell overlays); no drag-and-drop, no width memory.
398
461
  - **Cards**: compact two-line cards with profile-colored nodes; status lines are
399
462
  green solid (done) / blue solid (current) / gray dashed (pending) / red broken
400
463
  (blocked).
401
464
  - **Detail drawer**: five sections — Overview / Trajectory / Handoff / Spec /
402
465
  Comments; `Esc` or back returns to the list.
403
466
  - **Actions** (`POST /kanban/action`): block / unblock / retry / complete /
404
- archive / comment, plus chain-level `confirm-audit`, `rename` (chain or task),
405
- and `delete` (chain, human-only, double-confirmed in the GUI). Human actions
406
- apply optimistic updates with rollback; the store reconciles against the
407
- authoritative snapshot on any divergence.
467
+ archive / comment / `waive-review`, plus chain-level `confirm-audit`, `rename`
468
+ (chain or task), `reopen-chain` and `delete` (chain, human-only, double-confirmed in
469
+ the GUI). Status actions (block / unblock / complete / archive / retry) apply
470
+ optimistic updates with rollback; the rest wait for the server event, and the store
471
+ re-pulls the authoritative snapshot on any divergence.
408
472
  - **Build**: `npm run build:client` produces `lib/client.js` in the
409
473
  `window.__ModuleLoader__.load()` format (identical convention to `dsh-client-*`).
410
474
  Adding dsh-swarm to a web profile auto-embeds it into `__DSH_BOOT__`.
411
475
 
412
476
  ### Architecture
413
477
 
414
- Five layers, with the domain layer kept **free of any DSH dependency** so it can be
415
- fully unit-tested and replayed in isolation.
478
+ Layers, with the domain layer kept free of **runtime** DSH dependencies (a single
479
+ type-only import of `ObjectJsonSchema` in `prefetch-manifest.ts`) so it can be fully
480
+ unit-tested and replayed in isolation.
416
481
 
417
482
  ```mermaid
418
483
  flowchart TB
@@ -422,7 +487,7 @@ flowchart TB
422
487
  Model["workflow-model: pure view projection"]
423
488
  end
424
489
 
425
- subgraph Domain ["domain/ (pure TS, zero DSH deps)"]
490
+ subgraph Domain ["domain/ (pure TS; no runtime DSH deps)"]
426
491
  ES["event-store (JSONL append-only, monotonic seq)"]
427
492
  SM["state-machine (task/chain/spec transitions)"]
428
493
  PJ["projection (events → BoardState)"]
@@ -452,6 +517,14 @@ flowchart TB
452
517
  WK["wiki-worker (W prefetch worker)"]
453
518
  end
454
519
 
520
+ subgraph Services ["services/"]
521
+ PROVIDER["kanban-provider (gate + evidence wiring)"]
522
+ GATERUN["gate-runner / gate-evidence / evidence-replay"]
523
+ OCRCLI["ocr-cli"]
524
+ IMD["im-delivery"]
525
+ CFG["config-provider"]
526
+ end
527
+
455
528
  subgraph Wiki ["wiki/"]
456
529
  WVC["wiki-vault-client (search/read/write)"]
457
530
  end
@@ -471,6 +544,11 @@ flowchart TB
471
544
  MG --> KS
472
545
  KS --> ES --> PJ --> SM --> PM
473
546
  EC --> KS
547
+ PROVIDER --> KS
548
+ GATERUN --> PROVIDER
549
+ OCRCLI --> TOOLSETS
550
+ IMD --> KS
551
+ CFG --> VORCH
474
552
  ```
475
553
 
476
554
  #### Layer responsibilities
@@ -482,6 +560,10 @@ flowchart TB
482
560
  - **Integration** (`src/tools/`, `src/routes/`) — cordis tools and routes:
483
561
  the role tool faces, main-session tools (`kanban_route` + read-only subset), and
484
562
  the `/kanban/*` HTTP/SSE bridge.
563
+ - **Services** (`src/services/`) — provider wiring (`kanban-provider` installs the gate
564
+ and evidence-check hooks), the gate runner / evidence replay, `ocr-cli`, IM delivery
565
+ and the config provider; the scheduler itself (`src/dispatcher/dispatcher.ts`)
566
+ belongs to the Dispatcher layer.
485
567
  - **Dispatcher** (`src/dispatcher/`) — event wake, phase orchestration, one-shot
486
568
  agent runner (persona preset mounting, model candidate chain, ToolGuard
487
569
  installation), watchdog, chain auditor, and merge gate.
@@ -496,7 +578,7 @@ Quality gates (see `AGENTS.md`):
496
578
 
497
579
  ```bash
498
580
  npm run typecheck # tsc -p tsconfig.json --noEmit (0 errors)
499
- npm test # npx vitest run (all green)
581
+ npm test # vitest run (all green)
500
582
  npm run build # tsc -p tsconfig.build.json + build:client (lib/client.js)
501
583
  ```
502
584
 
@@ -510,39 +592,6 @@ python tests/e2e/gui-check.py --url http://127.0.0.1:3080/
510
592
  > Deploying to a running DSH instance requires a plugin reload/restart; building
511
593
  > alone does not hot-reload the running plugin.
512
594
 
513
- ### Implemented & known limitations
514
-
515
- #### Implemented (v0.1.0)
516
-
517
- - [x] **Swarm mode**: natural-language intent recognition (plan/openspec/learning/send) + confirmation gate + read-only main-session hard gate
518
- - [x] Event-sourced domain + deterministic state machines (red-team replay)
519
- - [x] 6-role phase pipeline with trimmed presets and session-bound permissions
520
- - [x] Delivery contract + review evidence gates + rework lifecycle
521
- - [x] TDD hard gate (D `tdd` handoff + DT `test_first` / `runner=vitest` verification)
522
- - [x] Protocol-violation recovery, heartbeat watchdog, failure circuit
523
- - [x] Chain completion audit gate + human confirm
524
- - [x] Post-DT merge gate (D pushes feature branch only)
525
- - [x] Phase-0 planning checklist + `file-prefetch` attachment
526
- - [x] GUI chain/task rename + chain delete (human-only)
527
- - [x] Model candidate chain with silent fallback + high reasoning effort
528
- - [x] Live SSE kanban tab (Conversation → Trajectory → Kanban)
529
-
530
- #### Known limitations
531
-
532
- - **Swarm-mode intent recognition relies on model self-judgment**: misjudgments are caught by the confirmation gate (no confirmation, no chain), but the risk is non-zero.
533
- - **Write guards are string-heuristic, not hard isolation.** PT/DT ToolGuards
534
- rely on path/command regex and reviewers get no git credentials; a soft
535
- constraint plus audit trail, not a mount-level sandbox.
536
- - **`open-code-review` (ocr) is optional per machine**: when missing, in-chain
537
- reviews block `review-tool-unavailable` before DT starts, with install guidance
538
- (one-click GUI install available) — no retries burned.
539
- - **Review evidence is existence-checked, not replay-proven.** Fields must be
540
- present and well-formed; proving the tests actually ran is not yet supported.
541
- - **Single default wiki-vault host** in the config default — point
542
- `wikiVault.baseUrl` at your deployment.
543
- - **PT creation depends on P's self-reported `pt_decision.needed`** —
544
- system-assisted detection from repo signals is not yet implemented.
545
-
546
595
  ---
547
596
 
548
597
  ## License