@joekytc/dsh-swarm 0.3.7 → 0.3.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +165 -116
- package/README.zh-CN.md +143 -96
- package/client/timeline-model.ts +14 -1
- package/lib/client.js +20 -2
- package/lib/config.d.ts +14 -3
- package/lib/config.js +10 -0
- package/lib/dispatcher/agent-runner.js +3 -7
- package/lib/dispatcher/chain-auditor.d.ts +7 -1
- package/lib/dispatcher/chain-auditor.js +7 -4
- package/lib/dispatcher/context-dedup.d.ts +1 -0
- package/lib/dispatcher/context-dedup.js +24 -0
- package/lib/dispatcher/dispatcher.js +12 -4
- package/lib/dispatcher/session-preset.d.ts +36 -0
- package/lib/dispatcher/session-preset.js +67 -0
- package/lib/domain/evidence-check.d.ts +17 -0
- package/lib/domain/evidence-check.js +44 -0
- package/lib/domain/gate-policy.d.ts +33 -9
- package/lib/domain/gate-policy.js +60 -19
- package/lib/domain/kanban-service.d.ts +27 -4
- package/lib/domain/kanban-service.js +45 -7
- package/lib/domain/review-evidence.js +3 -1
- package/lib/domain/types.d.ts +7 -1
- package/lib/roles/toolsets.d.ts +19 -8
- package/lib/roles/toolsets.js +24 -15
- package/lib/services/evidence-replay.d.ts +28 -0
- package/lib/services/evidence-replay.js +79 -0
- package/lib/services/gate-evidence.d.ts +2 -0
- package/lib/services/gate-evidence.js +21 -0
- package/lib/services/im-delivery.d.ts +4 -0
- package/lib/services/im-delivery.js +31 -3
- package/lib/services/kanban-provider.d.ts +1 -0
- package/lib/services/kanban-provider.js +92 -19
- package/lib/tools/main-session-tools.js +6 -13
- package/package.json +1 -1
- package/personas/persona-dt.md +3 -0
package/README.md
CHANGED
|
@@ -19,10 +19,10 @@ dsh-swarm is a DSH plugin that turns one requirement into a strict, evidence-ver
|
|
|
19
19
|
Swarm mode turns your main session into a **team lead**: **you state the requirement, it clarifies, plans, confirms, delegates and follows through** — entirely in natural language, no commands to remember.
|
|
20
20
|
|
|
21
21
|
- **No commands to memorize** — just state your requirement; no `/plan:` or `/openspec:` prefixes needed.
|
|
22
|
-
- **Automatic intent recognition** — development requests → clarify/plan and build a chain; lessons & retrospectives → persist to memory; group notifications → deliver to WeCom; Q&A / chit-chat → answered directly.
|
|
23
|
-
- **Free delivery** — `/sms <intent>` (e.g. "post current progress to the group"): facts are grounded via kanban lookup, then the body is composed per intent and delivered; `-s` or wording like "private chat" targets the DM. Group and private-chat targets
|
|
24
|
-
- **Confirmation gate against accidental chains** — after the checklist is saved, a chain
|
|
25
|
-
- **The lead is read-only** — the main session cannot write/edit repo sources, nor run git mutations (push/commit/
|
|
22
|
+
- **Automatic intent recognition** — development requests → clarify/plan and build a chain; lessons & retrospectives → persist to memory; group notifications → deliver to WeCom; Q&A / chit-chat → answered directly. Intent is judged by the model, not by a code-level classifier — the confirmation gate below is what stops a misjudged chain.
|
|
23
|
+
- **Free delivery** — `/sms <intent>` (e.g. "post current progress to the group"): facts are grounded via kanban lookup, then the body is composed per intent and delivered; `-s` or wording like "private chat" targets the DM. A bare `/sms` re-sends the latest completed-chain report and `/sms blocked [chainId]` the block notice — those two bodies are rendered by system code from kanban facts, never rewritten by the lead. Group and private-chat targets auto-resolve to the single saved target (0 or 2+ targets error out; clean up in dsh-im settings, or pin `imDelivery.targetId` / `imDelivery.dmTargetId`, which skips the count check — though with `imDelivery.botId` empty the pinned target must still belong to the auto-discovered bot, or it errors out).
|
|
24
|
+
- **Confirmation gate against accidental chains** — after the checklist is saved, the lead is instructed to build a chain only once you reply with an explicit affirmative (`确认` / `开干` / `开跑` / `开始` / `go`, etc.). The gate is judged by the model, not enforced by a code-level check; vague replies, topic switches, or edit-only feedback count as *not confirmed*.
|
|
25
|
+
- **The lead is read-only** — the main session cannot write/edit repo sources, nor run git mutations (push/commit/reset…); a bare `git checkout`/`git switch` of an existing branch is allowed. Writing code is done by the executor (D) in an isolated workspace by design.
|
|
26
26
|
- **Progress is always actually queried** — ask "how is it going?" anytime and the lead reports from real kanban lookups, never fabricated.
|
|
27
27
|
|
|
28
28
|
### Why it's designed this way
|
|
@@ -33,7 +33,7 @@ Coordinating several agents on one task typically fails in three ways:
|
|
|
33
33
|
- **Unverifiable handoffs** — an agent claims "done" with no reproducible evidence, and the next agent builds on sand.
|
|
34
34
|
- **Silent deadlocks** — an agent stops without finishing and the pipeline hangs, or bad code is merged before anyone reviewed it.
|
|
35
35
|
|
|
36
|
-
dsh-swarm encodes a *contract* against all three: one
|
|
36
|
+
dsh-swarm encodes a *contract* against all three: one responsibility per role, enforced by the permission matrix, tool faces and task-body instructions;
|
|
37
37
|
every handoff must carry structured evidence or the phase will not close; every stall or review
|
|
38
38
|
failure lands in a visible, recoverable state — with you (the human) as the final trust anchor.
|
|
39
39
|
It is built correctness-first: deterministic state machines, append-only event sourcing, idempotent
|
|
@@ -55,13 +55,21 @@ schedulers, and a red-team test suite that replays the event log and rejects any
|
|
|
55
55
|
|
|
56
56
|
Prerequisites: a working DSH runtime (`@deepseek-ai/*`), Node.js ≥ 22.19 and npm. Optional: a wiki-vault HTTP service (KB features, see [Configuration](#configuration)).
|
|
57
57
|
|
|
58
|
+
From npm (the published tarball ships the built `lib/`):
|
|
59
|
+
|
|
60
|
+
```bash
|
|
61
|
+
dsh plugin --profile web add @joekytc/dsh-swarm
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
From a source checkout (rebuild first so `lib/` matches the sources):
|
|
65
|
+
|
|
58
66
|
```bash
|
|
59
67
|
npm install
|
|
60
68
|
npm run build # tsc -p tsconfig.build.json + client bundle (lib/client.js)
|
|
61
|
-
dsh plugin --profile web add
|
|
69
|
+
dsh plugin --profile web add .
|
|
62
70
|
```
|
|
63
71
|
|
|
64
|
-
> From GitHub source: `dsh plugin --profile web add github:joekytc/dsh-swarm`
|
|
72
|
+
> From GitHub source: `dsh plugin --profile web add github:joekytc/dsh-swarm` — the repository tracks the built `lib/`.
|
|
65
73
|
|
|
66
74
|
### 2. Switch your main-session preset
|
|
67
75
|
|
|
@@ -91,14 +99,14 @@ Lead: Chain created (ch_…), live progress on the kanban tab (Conversation →
|
|
|
91
99
|
```
|
|
92
100
|
|
|
93
101
|
- **Kanban**: the third tab of the conversation center (Conversation → Trajectory → Kanban). Click a card for Overview / Trajectory / Handoff / Spec / Comments.
|
|
94
|
-
- **Completion**: when a chain completes, the system audits the workspace and (for D chains) automatically merges the feature branch into the spec-declared target branch
|
|
102
|
+
- **Completion**: when a chain completes, the system audits the workspace and (for D chains) automatically merges the feature branch into the spec-declared target branch. An audit warning blocks the final wrap-up until you confirm ownership in the GUI — it does not gate the merge.
|
|
95
103
|
- **Progress**: just ask "how is it going?" — the lead reports from real kanban lookups and relays blocking reasons faithfully.
|
|
96
104
|
|
|
97
105
|
---
|
|
98
106
|
|
|
99
107
|
## What it does for you
|
|
100
108
|
|
|
101
|
-
Six roles, one job each,
|
|
109
|
+
Six roles, one job each — boundaries enforced by the permission matrix, trimmed tool faces and task-body instructions, so no role creep:
|
|
102
110
|
|
|
103
111
|
| Role | One-line responsibility | What it never does |
|
|
104
112
|
|---|---|---|
|
|
@@ -129,20 +137,40 @@ All keys are optional; schema lives in `src/config.ts`. **Most users only need t
|
|
|
129
137
|
|---|---|---|
|
|
130
138
|
| `storageDir` | `$DSH_HOME/storages/kanban` | Event log (`events.jsonl`), orchestration state, per-task workspaces, `dispatcher.log`. Value must use the unquoted `!!js dshHomePath("storages/kanban")` form — quoting degrades it into a literal string |
|
|
131
139
|
| `wikiVault.baseUrl` | `''` (empty) | wiki-vault HTTP service for KB reads/writes — required for KB features; set to your own server |
|
|
140
|
+
| `wikiVault.pagePrefix` | `projects/` | Namespace prefix for generated wiki pages |
|
|
132
141
|
| `roles.models.<role>` | `{}` | Per-role model: `{ provider, model, reasoningEffort?, fallbacks?[] }` |
|
|
133
142
|
| `roles.models.<role>.reasoningEffort` | `high` | Default reasoning effort for all roles |
|
|
134
143
|
| `roles.models.<role>.fallbacks` | `[]` | Silent fallback candidates (audited via `[model-fallback]` comment) |
|
|
135
144
|
| `dispatcher.staleTimeoutSeconds` | `14400` | Heartbeat timeout; running task without heartbeat is reclaimed |
|
|
136
145
|
| `dispatcher.maxRetries` | `3` | Failure retries before circuit → `blocked(gave_up)` |
|
|
137
146
|
| `dispatcher.heartbeatIntervalSeconds` | `300` | Watchdog heartbeat period |
|
|
138
|
-
| `dispatcher.maxProtocolViolations` | `2` | Protocol-violation guardrail:
|
|
147
|
+
| `dispatcher.maxProtocolViolations` | `2` | Protocol-violation guardrail: once consecutive violations reach this many, the next one is final (`gave_up`) |
|
|
139
148
|
| `dispatcher.maxReworksPerRole` | `{ pt: 3, dt: 3 }` | Max review rework rounds before `review/gave-up` + `[review-final]` |
|
|
140
|
-
| `prefixRoutes.plan` | `/plan:` | Command-mode planning prefix |
|
|
149
|
+
| `prefixRoutes.plan` | `/plan:` | Command-mode phase-0 planning prefix |
|
|
141
150
|
| `prefixRoutes.openspec` | `/openspec:` | Command-mode approve-and-execute prefix |
|
|
142
|
-
| `
|
|
143
|
-
| `
|
|
144
|
-
| `
|
|
151
|
+
| `prefixRoutes.learning` | `/learning` | Lessons / retrospective prefix |
|
|
152
|
+
| `prefixRoutes.send` | `/sms` | Free-delivery prefix (DM with `-s`) |
|
|
153
|
+
| `memory.enabled` | `true` | Memory recall index; `false` makes `planning_memory_recall` return a disabled notice |
|
|
154
|
+
| `memory.maxIndexEntries` | `8` | Max recalled memory entries (1–20) |
|
|
155
|
+
| `ui.enabled` | `true` | Declared switch; not consumed yet — the tab registers unconditionally |
|
|
156
|
+
| `ui.contentMinWidth` | `715` | Declared lower width bound (px); not consumed by the client yet — the tab follows the host conversation width |
|
|
157
|
+
| `ui.contentMaxWidth` | `780` | Declared upper width bound (px); not consumed by the client yet |
|
|
145
158
|
| `ui.sseHeartbeatSeconds` | `20` | SSE heartbeat interval |
|
|
159
|
+
| `gates.enabled` | `true` | TDD measurement gate; `false` → silent skip (no event) |
|
|
160
|
+
| `gates.timeoutMs` | `600000` | Per-command gate timeout (ms), SIGKILL at the deadline |
|
|
161
|
+
| `gates.forbidden` | `['rm -rf /', 'git push']` | Command-blacklist substrings (defence in depth) |
|
|
162
|
+
| `evidenceReplay.enabled` | `false` | L3 replay of model-written commands — not a sandbox, see [Issue evidence check](#issue-evidence-check-pr2-opt-in) |
|
|
163
|
+
| `evidenceReplay.timeoutMs` | `600000` | Per-replay command timeout (ms) |
|
|
164
|
+
| `evidenceReplay.allowPrefixes` | `['npx --no-install vitest', 'npm test', 'npm run build', 'npm run typecheck', 'tsc', 'eslint']` | Allowlisted tool prefixes (word-boundary match) |
|
|
165
|
+
| `imDelivery.enabled` | `false` | WeCom delivery via dsh-im (W3 wrap-up / chain blocked / review gave-up) |
|
|
166
|
+
| `imDelivery.botId` | `''` | Empty = auto-discover the only wecom bot |
|
|
167
|
+
| `imDelivery.targetId` | `''` | Empty = auto-discover the only saved group target |
|
|
168
|
+
| `imDelivery.dmTargetId` | `''` | DM target for `/sms -s` |
|
|
169
|
+
| `imDelivery.fallbackBotId` | `''` | Bot used when preset-based matching finds nothing |
|
|
170
|
+
| `reviewEngine.mode` | `delegate` | `delegate` or `managed` — see [Review engine (ocr)](#review-engine-ocr) |
|
|
171
|
+
| `reviewEngine.managed.provider` | `''` | Model-chain provider id used in managed mode |
|
|
172
|
+
| `reviewEngine.managed.model` | `''` | Model-chain model id used in managed mode |
|
|
173
|
+
| `wikiWritePresets` | `['swarm', 'kanban-w', 'ptc']` | Presets allowed to call `wiki_write` (page paths are still whitelisted by namespace) |
|
|
146
174
|
|
|
147
175
|
---
|
|
148
176
|
|
|
@@ -150,7 +178,7 @@ All keys are optional; schema lives in `src/config.ts`. **Most users only need t
|
|
|
150
178
|
|
|
151
179
|
Implementation reviews (the in-chain DT phase and standalone reviews) are powered by
|
|
152
180
|
[open-code-review](https://open-codereview.ai) (ocr), with two modes switchable in the
|
|
153
|
-
web config panel under
|
|
181
|
+
web config panel under 「Swarm 配置 → 评审引擎(ocr)」 (Swarm config → Review engine (ocr)):
|
|
154
182
|
|
|
155
183
|
| Mode | How it works | Notes |
|
|
156
184
|
|---|---|---|
|
|
@@ -159,21 +187,23 @@ web config panel under "Swarm config → Review engine (ocr)":
|
|
|
159
187
|
|
|
160
188
|
### Install
|
|
161
189
|
|
|
162
|
-
- When ocr is missing, the config panel shows a red banner — click
|
|
190
|
+
- When ocr is missing, the config panel shows a red banner — click 「安装 ocr」 (Install ocr) for a one-click global install (async, cancellable);
|
|
163
191
|
- or run `npm install -g @alibaba-group/open-code-review` in a terminal, then verify with `ocr --version`.
|
|
164
192
|
|
|
165
193
|
### Standalone review (no chain needed)
|
|
166
194
|
|
|
167
|
-
1. Switch the session to
|
|
195
|
+
1. Switch the session to 「交付评审官(DT)」 (Delivery Reviewer (DT)) at the top of the dsh web UI and just talk;
|
|
168
196
|
2. State the review target: a local directory / branch range (from…to) / a single commit / uncommitted workspace diff / a public repo URL (auto-cloned into a temp dir, discarded afterwards);
|
|
169
197
|
3. The report is first fully output to the conversation;
|
|
170
|
-
4. Only after you confirm is it written to the
|
|
198
|
+
4. Only after you confirm is it written to the KB at `projects/<repo>/reviews/<topic>-<date>/` — remote KB mode only; with the default empty `wikiVault.baseUrl` (local mode) no wiki page is written.
|
|
199
|
+
|
|
200
|
+
Read-only in practice: bash/run_code writes and wiki writes outside the reviews namespace are blocked by the guard, while the fs `write`/`edit` tools are not blocked in a standalone DT session.
|
|
171
201
|
|
|
172
202
|
### Configuration notes
|
|
173
203
|
|
|
174
|
-
- Mode, provider and model are all chosen on the
|
|
175
|
-
- After picking, click
|
|
176
|
-
- When managed is not ready,
|
|
204
|
+
- Mode, provider and model are all chosen on the 「评审引擎(ocr)」 card; the provider/model dropdowns share the same catalog as the model chain;
|
|
205
|
+
- After picking, click 「应用到 ocr」 (Apply to ocr) — the system writes the wiring into ocr's custom config (`dsh-managed`); the API key is resolved from the dsh model config and written into ocr, never shown in plain text in the panel; if resolution fails it degrades gracefully and points you to a manual `ocr config provider` in a terminal;
|
|
206
|
+
- When managed is not ready, the ocr tool refuses the managed call and returns delegate guidance; the reviewer switches to delegate mode — nothing is blocked.
|
|
177
207
|
|
|
178
208
|
Official docs: [Installation](https://open-codereview.ai/docs/installation) · [Model configuration](https://open-codereview.ai/docs/configuration) · [Delegate mode](https://open-codereview.ai/docs/delegate)
|
|
179
209
|
|
|
@@ -182,9 +212,10 @@ Official docs: [Installation](https://open-codereview.ai/docs/installation) · [
|
|
|
182
212
|
## Trust & guardrails (user's view)
|
|
183
213
|
|
|
184
214
|
- **Read-only hard gate for the lead** — in swarm mode, main-session writes to sources and git mutations are blocked by a system gate; if blocked, just let the lead explain — execution is done by the D role.
|
|
185
|
-
- **Confirmation gate** —
|
|
215
|
+
- **Confirmation gate** — the lead only builds a chain after you reply with an explicit affirmative (enforced by the lead's instructions, not by a code-level check).
|
|
186
216
|
- **TDD hard gate** — implementations must ship with tests (or an explained skip); reviews machine-verify "tests really ran, and were written first".
|
|
187
|
-
- **Human trust anchors** — spec approval, unblock, audit confirmation and chain deletion are human-only;
|
|
217
|
+
- **Human trust anchors** — spec approval, unblock, audit confirmation and chain deletion are human-only; role agents cannot approve specs, and a chain is only created from your confirmed routing call (the main session routes as `human`).
|
|
218
|
+
- **Guardrails are constraints, not a sandbox** — PT/DT write guards rely on path and command regexes (reviewers get no git credentials), and review evidence is existence-checked: fields must be present and well-formed, while replaying the commands to prove they ran happens only when `evidenceReplay.enabled` is turned on.
|
|
188
219
|
- Full mechanics (permission matrix, delivery contract, review chain, rework, failure recovery) live under [Advanced / Developers](#advanced--developers).
|
|
189
220
|
|
|
190
221
|
---
|
|
@@ -195,16 +226,16 @@ Official docs: [Installation](https://open-codereview.ai/docs/installation) · [
|
|
|
195
226
|
|
|
196
227
|
### Roles & the execution pipeline (full table)
|
|
197
228
|
|
|
198
|
-
Six roles are dispatched by the scheduler as one-shot agent sessions (deterministic session id `kbn-<taskId
|
|
229
|
+
Six roles are dispatched by the scheduler as one-shot agent sessions (deterministic session id `kbn-<taskId>`; a retry resumes that same session — `resumeSessionId` is always null in the current implementation — while a rework task starts its own `kbn-<reworkId>` session). Each role-agent session is bound to exactly one task (`boundTaskId`) and gets a trimmed tool face. V is the exception: a chain-scoped orchestrator session (`kbn-v-<chainId>`) with no `boundTaskId`.
|
|
199
230
|
|
|
200
231
|
| Role | Alias | Responsibility | Tool face (highlights) |
|
|
201
232
|
|---|---|---|---|
|
|
202
233
|
| **V** | Orchestrator | Drives the phase machine, creates one card per phase, posts `[blocked-review]` guidance on stalls. Never executes. | `kanban_create` + task tools + spec view |
|
|
203
234
|
| **P** | Planner | Reads spec + repo facts (incl. read-only self-checks), writes an OpenSpec implementation plan, opts into PT via `pt_decision.needed`. Never executes. | Task tools + spec view, read-only (writes only `openspec/changes/`) |
|
|
204
235
|
| **PT** | Plan reviewer | Read-only review of P's plan (requirements alignment, completeness, logic). Outputs verdict + issues. | Task tools + spec view, **read-only ToolGuard** |
|
|
205
|
-
| **W** | Knowledge officer | W2/W3 KB sync (`w:kb`). Never touches code/git. | Task tools + `wiki_search/read/write` (remote) / `skill`→llm-wiki (local) + read-only spec view |
|
|
236
|
+
| **W** | Knowledge officer | W2/W3 KB sync (`w:kb`). Never touches code/git. | Task tools + `wiki_search/read/write` (remote) / `skill`→llm-wiki (local) + `prefetch_file`/`prefetch_external`/`prefetch_kb` + read-only spec view |
|
|
206
237
|
| **D** | Executor | The *only* role that writes code: worktree → implement → verify → `[AI-GEN]` commit → push feature branch (merging into the spec-declared target branch is done by the system only after DT passes). | Task tools + wiki read + bash/fs/run_code (full dev) + subagent (spawn/fork/list-agents) + goal |
|
|
207
|
-
| **DT** | Implementation reviewer | Empirically verifies D's work (test/build/typecheck/diff/git + open-code-review), writes review page to KB. Read-only against the repo. | Task tools + wiki read/write (review namespace) + bash/fs/run_code, **read-only ToolGuard** |
|
|
238
|
+
| **DT** | Implementation reviewer | Empirically verifies D's work (test/build/typecheck/diff/git + open-code-review), writes review page to KB. Read-only against the repo. | Task tools + wiki read/write (review namespace) + `ocr_review` + bash/fs/run_code, **read-only ToolGuard** |
|
|
208
239
|
|
|
209
240
|
### Guardrails in detail
|
|
210
241
|
|
|
@@ -234,41 +265,64 @@ Six roles are dispatched by the scheduler as one-shot agent sessions (determinis
|
|
|
234
265
|
| prefetch | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ |
|
|
235
266
|
| audit-confirm | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ |
|
|
236
267
|
| create-rework-task | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅ |
|
|
268
|
+
| reopen-chain | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ |
|
|
269
|
+
| waive-review | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ |
|
|
237
270
|
|
|
238
271
|
Key guarantees (two):
|
|
239
272
|
|
|
240
|
-
- **The main session cannot execute.**
|
|
241
|
-
`kanban_comment`
|
|
242
|
-
`
|
|
243
|
-
|
|
244
|
-
|
|
273
|
+
- **The main session cannot execute.** Its kanban tool face is a read-only subset —
|
|
274
|
+
`kanban_show`/`kanban_chain`/`kanban_list`/`kanban_comment` plus the human-recovery pair
|
|
275
|
+
`kanban_reopen_chain`/`kanban_waive_review` — together with `spec_card_view`, `kanban_route`,
|
|
276
|
+
the `planning_*` tools and `sms_send`; never `kanban_create`/`kanban_complete`/`kanban_block`.
|
|
277
|
+
Chains/specs are created only via swarm-mode intents or `/plan:`+`/openspec:`; the GUI observes
|
|
278
|
+
and mutates task state but never creates chains or tasks — "who decided to run what" stays
|
|
279
|
+
explicit and auditable.
|
|
245
280
|
- **Session binding prevents cross-task escalation** (a W agent bound to task A
|
|
246
281
|
cannot complete/block task B even though both are W tasks); DT writes are
|
|
247
282
|
confined to the `projects/<repoSlug>/<chain>/review/` namespace by a ToolGuard on top of
|
|
248
|
-
the matrix; and no role agent can approve specs, unblock,
|
|
249
|
-
those are human trust anchors; `system` handles only
|
|
283
|
+
the matrix; and no role agent can approve specs, unblock, confirm audits, waive a
|
|
284
|
+
review or reopen a chain — those are human trust anchors; `system` handles only
|
|
285
|
+
mechanical bookkeeping.
|
|
250
286
|
|
|
251
287
|
#### Delivery contract (upstream owes downstream)
|
|
252
288
|
|
|
253
289
|
Each phase's handoff must carry the keys its downstream actually reads
|
|
254
|
-
(`src/domain/delivery-contract.ts`). Missing keys block the
|
|
255
|
-
immediately
|
|
256
|
-
|
|
290
|
+
(`src/domain/delivery-contract.ts`). Missing delivery keys block the W/P card
|
|
291
|
+
immediately and admit no human exemption; the D/PT/DT evidence gates instead reject
|
|
292
|
+
`complete` with an error (the card stays running) and are human-exempt. The
|
|
293
|
+
orchestrator never builds a downstream card on a blocked parent:
|
|
257
294
|
|
|
258
295
|
| Card | Required handoff keys |
|
|
259
296
|
|---|---|
|
|
260
|
-
| W2 / W3 (`w:kb`) | `kb_url` + `page_path` |
|
|
297
|
+
| W2 / W3 (`w:kb`) | `kb_url` + `page_path` — non-empty is not enough: with a configured `wikiVault.baseUrl` the `kb_url` must start with it (and must be exactly `''` in local mode), while `page_path` must sit in the allowlisted namespaces (`wiki/**` locally) |
|
|
261
298
|
| P (`p:openspec`) | `artifacts_path` + `pt_decision` (`needed` boolean required; when `needed: true`, `reason` is required) |
|
|
262
|
-
| D (`d:execute`) | `changed_files` + (`commit_hash` or `push`) — `hasDeliveryEvidence`; `branch` (feature branch) is
|
|
299
|
+
| D (`d:execute`) | `changed_files` + (`commit_hash` or `push`) — `hasDeliveryEvidence`; `branch` (feature branch) is not a delivery key, but when the TDD gate actually runs the tests it checks the branch and bounces the card if it is absent or mismatched; `tdd` (`test_files` or `skipped.reason`, XOR) |
|
|
263
300
|
| PT / DT | `review_evidence` (schema-valid) — `validateReviewEvidence` |
|
|
264
301
|
|
|
265
302
|
#### TDD hard gate (evidence threshold)
|
|
266
303
|
|
|
267
304
|
D completes only with `tdd` — `test_files` (with `test_first`) or `skipped.reason`
|
|
268
|
-
(XOR, `delivery-evidence.ts`). DT's `review_evidence` must carry `tdd`;
|
|
269
|
-
`
|
|
270
|
-
must hold (`review-evidence.ts`). This makes
|
|
271
|
-
first" a machine-checked property rather than
|
|
305
|
+
(XOR, `delivery-evidence.ts`). DT's `review_evidence` must carry `tdd`; whenever
|
|
306
|
+
`tdd.test_files` are declared the runner must be `vitest` (`test.runner`), and on a
|
|
307
|
+
`pass` verdict `test_first === true` must hold (`review-evidence.ts`). This makes
|
|
308
|
+
"tests actually ran, and were written first" a machine-checked property rather than
|
|
309
|
+
a claim.
|
|
310
|
+
|
|
311
|
+
#### Gate skip alarm & declaration cross-check
|
|
312
|
+
|
|
313
|
+
A declaration that disagrees with the diff no longer passes the gate silently
|
|
314
|
+
(three silent-skip modes remain, emitting no event: `gates.enabled=false`, a handoff
|
|
315
|
+
without `worktree_dir`, and non-`d:execute` cards):
|
|
316
|
+
|
|
317
|
+
- **`task/gate-skipped` event** — a declared `tdd.skipped` is allowed when the diff is pure docs/config, or when the diff cannot be computed (conservative alarm-skip); both leave the event as an audit trail. Otherwise the gate **bounces** the card back (same session, agent fixes and re-completes).
|
|
318
|
+
- **Declaration ↔ reality cross-check** — declaring `test_files` with no test-file change in the diff (stale-test handoff), an empty `test_files`, an invalid path, or a branch mismatch all **bounce** instead of silently skipping.
|
|
319
|
+
- After 3 cumulative bounces the task is **blocked for human review** (`gave_up: gate bounced 3 times`).
|
|
320
|
+
- Gate runs now archive raw output to `<storageDir>/gate-logs/<taskId>.log` (path is referenced in the gate event detail).
|
|
321
|
+
- `review_evidence.lint` must be a structured object (aligning with `build`/`typecheck`); `null`/scalars are rejected.
|
|
322
|
+
|
|
323
|
+
#### Issue evidence check (PR2, opt-in)
|
|
324
|
+
|
|
325
|
+
Reviewer-reported issues (DT cards only — PT issues are never evidence-checked) can carry `evidence = { file, command, exit }` (raw output archive, first line `[exit code: N]`). Verification is three-tier: missing evidence → flagged (`not-provided`; critical/high → `could-not-replay` + needs-human); archive paper-check (zero execution) → `matches`; on mismatch a **replay** re-runs the command — **only if `evidenceReplay.enabled` is turned on**, and only for allowlisted tool prefixes (word-boundary match; `npm run` restricted to fixed script names). Replays execute model-written commands and are **NOT a sandbox**; keep the switch off unless you accept that risk. Results are summarized in a `review/evidence-check` event; `differs` never auto-fails a review — it surfaces to human review.
|
|
272
326
|
|
|
273
327
|
#### Phase-0 planning checklist
|
|
274
328
|
|
|
@@ -280,9 +334,9 @@ blocks the save, and chain creation mounts the checklist as the `file-prefetch`
|
|
|
280
334
|
|
|
281
335
|
#### Review quality chain
|
|
282
336
|
|
|
283
|
-
- After **P** completes, **PT** is
|
|
284
|
-
`pt_decision.needed =
|
|
285
|
-
(V only creates the card).
|
|
337
|
+
- After **P** completes, **PT** is skipped only when P's handoff delivers
|
|
338
|
+
`pt_decision.needed = false`; `true` or an absent decision creates the PT card
|
|
339
|
+
(fail-safe) — the orchestrator never overrides the decision (V only creates the card).
|
|
286
340
|
- After **D** completes, a **DT** card is *always* created.
|
|
287
341
|
- **PT/DT are read-only**: a ToolGuard mechanically denies writes to the repo
|
|
288
342
|
sources, git mutations, and (for DT) wiki writes outside the review namespace.
|
|
@@ -291,9 +345,10 @@ blocks the save, and chain creation mounts the checklist as the `file-prefetch`
|
|
|
291
345
|
`review-tool-unavailable` when ocr is missing (reason notes GUI install),
|
|
292
346
|
without burning retries.
|
|
293
347
|
- `review_evidence` must pass `validateReviewEvidence` or the review card cannot
|
|
294
|
-
complete: PT needs verdict + issues + plan ref
|
|
295
|
-
|
|
296
|
-
|
|
348
|
+
complete: PT needs verdict + issues + plan ref, and a `fail` verdict requires at
|
|
349
|
+
least one unresolved `critical`/`high` issue (otherwise it must be `pass`); DT
|
|
350
|
+
additionally needs test (exit 0 on pass), build/typecheck, lint, non-empty diff,
|
|
351
|
+
git, ocr/fallback conclusion, and `tdd`.
|
|
297
352
|
|
|
298
353
|
#### Rework (review failure)
|
|
299
354
|
|
|
@@ -304,10 +359,11 @@ A failed review never mutates a `done` card. Instead the system records
|
|
|
304
359
|
previous round's issues verbatim in a `## 本轮修复清单` (fix-this-round) section.
|
|
305
360
|
A fresh review card is then dispatched for the rework.
|
|
306
361
|
When `reviewAttempt` reaches `maxReworksPerRole` (PT 3 / DT 3), the system records
|
|
307
|
-
`review/gave-up` and posts a `[review-final]` evidence-chain comment
|
|
362
|
+
`review/gave-up` and posts a `[review-final]` evidence-chain comment (the same marker
|
|
363
|
+
is reused by the convergence-gate downgrade below), and (with IM
|
|
308
364
|
delivery enabled) sends a `[评审超限待裁决]` notification with two exit paths:
|
|
309
365
|
waive the review (`kanban_waive_review` / GUI “豁免评审”) or reopen the chain
|
|
310
|
-
(`kanban_reopen_chain` / GUI “人工恢复”). **Convergence gate
|
|
366
|
+
(`kanban_reopen_chain` / GUI “人工恢复”). **Convergence gate**: on PT
|
|
311
367
|
rework rounds, once all legacy issues are resolved and no unresolved `critical`
|
|
312
368
|
remains, a `fail` verdict is downgraded to `pass` (new findings flow downstream as
|
|
313
369
|
non-blocking suggestions), so the loop always converges.
|
|
@@ -330,17 +386,19 @@ Two orthogonal failure paths, both human-recoverable:
|
|
|
330
386
|
model candidates (primary + fallbacks, `reasoningEffort: high` default) fall
|
|
331
387
|
back silently (audited via `[model-fallback]` comment); if *all* candidates fail
|
|
332
388
|
it blocks `model-unavailable` for the human. A single hanging V wake cannot
|
|
333
|
-
stall the scheduler —
|
|
389
|
+
stall the scheduler — V wakes are wrapped in a 60 s timeout (role-task dispatch
|
|
390
|
+
still awaits the session's `whenIdle`).
|
|
334
391
|
|
|
335
392
|
#### Chain completion: audit gate + merge gate
|
|
336
393
|
|
|
337
394
|
When the mechanical chain-complete rule fires, two gates run in the
|
|
338
395
|
`chain/completed` hook:
|
|
339
396
|
|
|
340
|
-
1. **Completion audit gate**: the `ChainAuditor`
|
|
341
|
-
|
|
342
|
-
|
|
343
|
-
|
|
397
|
+
1. **Completion audit gate**: the `ChainAuditor` scans live non-role sessions for
|
|
398
|
+
write-capable tool calls aimed at the chain workspace (`kanban-*` preset subagents
|
|
399
|
+
exempt), then reconciles artifact ownership as a mechanical fallback. Orphaned
|
|
400
|
+
writes emit `chain/audit-warning`; the UI shows a warning banner and blocks the
|
|
401
|
+
final summary until the human confirms ownership (`chain/audit-confirmed`, human-only).
|
|
344
402
|
2. **Merge gate (post-DT system merge)**: D never merges to the target branch and
|
|
345
403
|
never pushes it — it only commits to (and optionally pushes) its feature branch,
|
|
346
404
|
carrying `branch` in its handoff. The target branch is the one declared in the
|
|
@@ -349,33 +407,38 @@ When the mechanical chain-complete rule fires, two gates run in the
|
|
|
349
407
|
→ git merge --no-ff <feature-branch> → git push`. Outcomes are recorded as
|
|
350
408
|
idempotent comments: `[merge-done]` (with hash), `[merge-skip]` (merge input
|
|
351
409
|
unresolvable), or `[merge-failed]` (checkout/merge/push failed, e.g. a conflict).
|
|
352
|
-
Failures never throw —
|
|
353
|
-
|
|
410
|
+
Failures never throw — the gate only records `[merge-failed]`; note that a failed
|
|
411
|
+
push can leave the target branch already merged locally, and a conflicting merge
|
|
412
|
+
leaves the worktree mid-merge. Repairing is a human job.
|
|
354
413
|
|
|
355
414
|
### Event sourcing & domain model
|
|
356
415
|
|
|
357
416
|
Every state change is appended to `<storageDir>/events.jsonl`, one JSON event per
|
|
358
|
-
line. The `seq` is assigned by the store (re-read from the file tail on every
|
|
359
|
-
|
|
360
|
-
|
|
417
|
+
line. The `seq` is assigned by the store (re-read from the file tail on every append
|
|
418
|
+
as the last line's `seq` + 1) — append-only, except that a human chain delete purges
|
|
419
|
+
that chain's lines and renumbers the remaining `seq` values. The **trajectory is the
|
|
420
|
+
event log itself**; restart replays it to rebuild the board.
|
|
361
421
|
|
|
362
422
|
```jsonc
|
|
363
423
|
// one line in events.jsonl
|
|
364
424
|
{ "seq": 12, "chainId": "ch_x_...", "taskId": "t_y_...",
|
|
365
425
|
"kind": "task/completed",
|
|
366
|
-
"payload": { "summary": "...", "metadata": { /* handoff evidence */ } },
|
|
426
|
+
"payload": { "summary": "...", "metadata": { /* handoff evidence */ }, "completedAt": 1760000000000 },
|
|
367
427
|
"author": "w", "at": 1760000000000 }
|
|
368
428
|
```
|
|
369
429
|
|
|
370
|
-
Event families: `chain/*` (created, executing, completed,
|
|
371
|
-
audit-warning, audit-confirmed, title-updated
|
|
372
|
-
|
|
373
|
-
|
|
430
|
+
Event families actually emitted: `chain/*` (created, executing, completed, blocked,
|
|
431
|
+
reopened, root-task-set, audit-warning, audit-confirmed, title-updated,
|
|
432
|
+
im-delivery-failed), `spec-card/*` (created, edited, approved), `task/*` (created,
|
|
433
|
+
claimed, heartbeat, commented, completed, blocked, unblocked, failed, archived,
|
|
434
|
+
renamed, gate-passed, gate-failed, gate-skipped), and `review/*` (passed, failed,
|
|
435
|
+
gave-up, waived, evidence-check).
|
|
374
436
|
|
|
375
|
-
Replay is **strict**: the projection
|
|
376
|
-
and throws on any illegal transition, so a corrupted or tampered log fails
|
|
377
|
-
instead of silently producing an inconsistent board (
|
|
378
|
-
`
|
|
437
|
+
Replay is **strict**: the projection replays every transition event through the state
|
|
438
|
+
machine and throws on any illegal transition, so a corrupted or tampered log fails
|
|
439
|
+
loudly instead of silently producing an inconsistent board (non-transition kinds such
|
|
440
|
+
as `task/commented` are recorded as no-ops). Covered by
|
|
441
|
+
`tests/redteam/anti-escalation.test.ts` and `tests/domain/projection.test.ts`.
|
|
379
442
|
|
|
380
443
|
The service emits events through a serialized queue (append-then-publish), and
|
|
381
444
|
subscribers (SSE) receive every event exactly once in order. UI and dispatcher both
|
|
@@ -383,36 +446,38 @@ consume the same persisted events — there is no secondary source of truth.
|
|
|
383
446
|
|
|
384
447
|
### Web client (Workflow kanban tab)
|
|
385
448
|
|
|
386
|
-
A browser-half React tab registered
|
|
387
|
-
|
|
449
|
+
A browser-half React tab registered into `conversation.view` (`id=kanban`,
|
|
450
|
+
`order=20`, so it sits after Conversation and Trajectory). It registers **no
|
|
388
451
|
shell-level overlays, sidebars, or detail panes**.
|
|
389
452
|
|
|
390
453
|
- **Data path**: initial snapshot (`GET /kanban/board`) → SSE stream
|
|
391
454
|
(`GET /kanban/events?after=<seq>`) → board-store applies events incrementally,
|
|
392
455
|
deduplicates by `seq`, and re-pulls the full snapshot on any gap. **No business
|
|
393
456
|
polling.**
|
|
394
|
-
- **Layout**: multi-chain vertical rails;
|
|
395
|
-
|
|
396
|
-
summary. In-page rename/delete use
|
|
397
|
-
no drag-and-drop, no width memory.
|
|
457
|
+
- **Layout**: multi-chain vertical rails; the width follows the host conversation
|
|
458
|
+
width (`--dsh-chat-content-width`, 780 px fallback), full height; the active chain
|
|
459
|
+
is expanded, blocked chains always show a warning summary. In-page rename/delete use
|
|
460
|
+
a lightweight modal (no shell overlays); no drag-and-drop, no width memory.
|
|
398
461
|
- **Cards**: compact two-line cards with profile-colored nodes; status lines are
|
|
399
462
|
green solid (done) / blue solid (current) / gray dashed (pending) / red broken
|
|
400
463
|
(blocked).
|
|
401
464
|
- **Detail drawer**: five sections — Overview / Trajectory / Handoff / Spec /
|
|
402
465
|
Comments; `Esc` or back returns to the list.
|
|
403
466
|
- **Actions** (`POST /kanban/action`): block / unblock / retry / complete /
|
|
404
|
-
archive / comment
|
|
405
|
-
and `delete` (chain, human-only, double-confirmed in
|
|
406
|
-
|
|
407
|
-
|
|
467
|
+
archive / comment / `waive-review`, plus chain-level `confirm-audit`, `rename`
|
|
468
|
+
(chain or task), `reopen-chain` and `delete` (chain, human-only, double-confirmed in
|
|
469
|
+
the GUI). Status actions (block / unblock / complete / archive / retry) apply
|
|
470
|
+
optimistic updates with rollback; the rest wait for the server event, and the store
|
|
471
|
+
re-pulls the authoritative snapshot on any divergence.
|
|
408
472
|
- **Build**: `npm run build:client` produces `lib/client.js` in the
|
|
409
473
|
`window.__ModuleLoader__.load()` format (identical convention to `dsh-client-*`).
|
|
410
474
|
Adding dsh-swarm to a web profile auto-embeds it into `__DSH_BOOT__`.
|
|
411
475
|
|
|
412
476
|
### Architecture
|
|
413
477
|
|
|
414
|
-
|
|
415
|
-
|
|
478
|
+
Layers, with the domain layer kept free of **runtime** DSH dependencies (a single
|
|
479
|
+
type-only import of `ObjectJsonSchema` in `prefetch-manifest.ts`) so it can be fully
|
|
480
|
+
unit-tested and replayed in isolation.
|
|
416
481
|
|
|
417
482
|
```mermaid
|
|
418
483
|
flowchart TB
|
|
@@ -422,7 +487,7 @@ flowchart TB
|
|
|
422
487
|
Model["workflow-model: pure view projection"]
|
|
423
488
|
end
|
|
424
489
|
|
|
425
|
-
subgraph Domain ["domain/ (pure TS
|
|
490
|
+
subgraph Domain ["domain/ (pure TS; no runtime DSH deps)"]
|
|
426
491
|
ES["event-store (JSONL append-only, monotonic seq)"]
|
|
427
492
|
SM["state-machine (task/chain/spec transitions)"]
|
|
428
493
|
PJ["projection (events → BoardState)"]
|
|
@@ -452,6 +517,14 @@ flowchart TB
|
|
|
452
517
|
WK["wiki-worker (W prefetch worker)"]
|
|
453
518
|
end
|
|
454
519
|
|
|
520
|
+
subgraph Services ["services/"]
|
|
521
|
+
PROVIDER["kanban-provider (gate + evidence wiring)"]
|
|
522
|
+
GATERUN["gate-runner / gate-evidence / evidence-replay"]
|
|
523
|
+
OCRCLI["ocr-cli"]
|
|
524
|
+
IMD["im-delivery"]
|
|
525
|
+
CFG["config-provider"]
|
|
526
|
+
end
|
|
527
|
+
|
|
455
528
|
subgraph Wiki ["wiki/"]
|
|
456
529
|
WVC["wiki-vault-client (search/read/write)"]
|
|
457
530
|
end
|
|
@@ -471,6 +544,11 @@ flowchart TB
|
|
|
471
544
|
MG --> KS
|
|
472
545
|
KS --> ES --> PJ --> SM --> PM
|
|
473
546
|
EC --> KS
|
|
547
|
+
PROVIDER --> KS
|
|
548
|
+
GATERUN --> PROVIDER
|
|
549
|
+
OCRCLI --> TOOLSETS
|
|
550
|
+
IMD --> KS
|
|
551
|
+
CFG --> VORCH
|
|
474
552
|
```
|
|
475
553
|
|
|
476
554
|
#### Layer responsibilities
|
|
@@ -482,6 +560,10 @@ flowchart TB
|
|
|
482
560
|
- **Integration** (`src/tools/`, `src/routes/`) — cordis tools and routes:
|
|
483
561
|
the role tool faces, main-session tools (`kanban_route` + read-only subset), and
|
|
484
562
|
the `/kanban/*` HTTP/SSE bridge.
|
|
563
|
+
- **Services** (`src/services/`) — provider wiring (`kanban-provider` installs the gate
|
|
564
|
+
and evidence-check hooks), the gate runner / evidence replay, `ocr-cli`, IM delivery
|
|
565
|
+
and the config provider; the scheduler itself (`src/dispatcher/dispatcher.ts`)
|
|
566
|
+
belongs to the Dispatcher layer.
|
|
485
567
|
- **Dispatcher** (`src/dispatcher/`) — event wake, phase orchestration, one-shot
|
|
486
568
|
agent runner (persona preset mounting, model candidate chain, ToolGuard
|
|
487
569
|
installation), watchdog, chain auditor, and merge gate.
|
|
@@ -496,7 +578,7 @@ Quality gates (see `AGENTS.md`):
|
|
|
496
578
|
|
|
497
579
|
```bash
|
|
498
580
|
npm run typecheck # tsc -p tsconfig.json --noEmit (0 errors)
|
|
499
|
-
npm test #
|
|
581
|
+
npm test # vitest run (all green)
|
|
500
582
|
npm run build # tsc -p tsconfig.build.json + build:client (lib/client.js)
|
|
501
583
|
```
|
|
502
584
|
|
|
@@ -510,39 +592,6 @@ python tests/e2e/gui-check.py --url http://127.0.0.1:3080/
|
|
|
510
592
|
> Deploying to a running DSH instance requires a plugin reload/restart; building
|
|
511
593
|
> alone does not hot-reload the running plugin.
|
|
512
594
|
|
|
513
|
-
### Implemented & known limitations
|
|
514
|
-
|
|
515
|
-
#### Implemented (v0.1.0)
|
|
516
|
-
|
|
517
|
-
- [x] **Swarm mode**: natural-language intent recognition (plan/openspec/learning/send) + confirmation gate + read-only main-session hard gate
|
|
518
|
-
- [x] Event-sourced domain + deterministic state machines (red-team replay)
|
|
519
|
-
- [x] 6-role phase pipeline with trimmed presets and session-bound permissions
|
|
520
|
-
- [x] Delivery contract + review evidence gates + rework lifecycle
|
|
521
|
-
- [x] TDD hard gate (D `tdd` handoff + DT `test_first` / `runner=vitest` verification)
|
|
522
|
-
- [x] Protocol-violation recovery, heartbeat watchdog, failure circuit
|
|
523
|
-
- [x] Chain completion audit gate + human confirm
|
|
524
|
-
- [x] Post-DT merge gate (D pushes feature branch only)
|
|
525
|
-
- [x] Phase-0 planning checklist + `file-prefetch` attachment
|
|
526
|
-
- [x] GUI chain/task rename + chain delete (human-only)
|
|
527
|
-
- [x] Model candidate chain with silent fallback + high reasoning effort
|
|
528
|
-
- [x] Live SSE kanban tab (Conversation → Trajectory → Kanban)
|
|
529
|
-
|
|
530
|
-
#### Known limitations
|
|
531
|
-
|
|
532
|
-
- **Swarm-mode intent recognition relies on model self-judgment**: misjudgments are caught by the confirmation gate (no confirmation, no chain), but the risk is non-zero.
|
|
533
|
-
- **Write guards are string-heuristic, not hard isolation.** PT/DT ToolGuards
|
|
534
|
-
rely on path/command regex and reviewers get no git credentials; a soft
|
|
535
|
-
constraint plus audit trail, not a mount-level sandbox.
|
|
536
|
-
- **`open-code-review` (ocr) is optional per machine**: when missing, in-chain
|
|
537
|
-
reviews block `review-tool-unavailable` before DT starts, with install guidance
|
|
538
|
-
(one-click GUI install available) — no retries burned.
|
|
539
|
-
- **Review evidence is existence-checked, not replay-proven.** Fields must be
|
|
540
|
-
present and well-formed; proving the tests actually ran is not yet supported.
|
|
541
|
-
- **Single default wiki-vault host** in the config default — point
|
|
542
|
-
`wikiVault.baseUrl` at your deployment.
|
|
543
|
-
- **PT creation depends on P's self-reported `pt_decision.needed`** —
|
|
544
|
-
system-assisted detection from repo signals is not yet implemented.
|
|
545
|
-
|
|
546
595
|
---
|
|
547
596
|
|
|
548
597
|
## License
|