thachvd-kit 1.0.37 → 1.0.39

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/LICENSE +21 -21
  2. package/README.md +303 -11
  3. package/THIRD_PARTY_NOTICES.md +49 -49
  4. package/bin/cli.js +1791 -1785
  5. package/bin/config.js +164 -0
  6. package/bin/entry.js +11 -1
  7. package/bin/native-skills.js +183 -149
  8. package/bin/spec-doctor.js +252 -0
  9. package/bin/spec-link.js +97 -0
  10. package/bin/spec-recipe.js +74 -0
  11. package/bin/spec-site.js +639 -0
  12. package/bin/spec-state.js +415 -0
  13. package/bin/spec.js +901 -0
  14. package/bin/upgrade.js +303 -303
  15. package/package.json +3 -3
  16. package/skills/finishing-a-development-branch/SKILL.md +240 -240
  17. package/skills/requesting-code-review/code-reviewer.md +198 -198
  18. package/skills/subagent-driven-development/SKILL.md +574 -574
  19. package/skills/subagent-driven-development/implementer-prompt.md +154 -154
  20. package/skills/subagent-driven-development/re-review-prompt.md +115 -115
  21. package/skills/subagent-driven-development/scripts/review-package +53 -53
  22. package/skills/subagent-driven-development/scripts/review-package.js +52 -52
  23. package/skills/subagent-driven-development/scripts/sdd-workspace +82 -82
  24. package/skills/subagent-driven-development/scripts/sdd-workspace-lib.js +62 -62
  25. package/skills/subagent-driven-development/scripts/sdd-workspace.js +15 -15
  26. package/skills/subagent-driven-development/scripts/task-brief +43 -43
  27. package/skills/subagent-driven-development/scripts/task-brief.js +46 -46
  28. package/skills/subagent-driven-development/task-reviewer-prompt.md +207 -207
  29. package/skills/system-discovery/SKILL.md +140 -0
  30. package/skills/system-reverse-engineer/SKILL.md +208 -0
  31. package/skills/system-spec-review/SKILL.md +177 -0
  32. package/skills/upstream.json +30 -30
  33. package/skills/using-git-worktrees/SKILL.md +175 -175
@@ -1,574 +1,574 @@
1
- ---
2
- name: subagent-driven-development
3
- description: Use when executing implementation plans with independent tasks in the current session
4
- ---
5
-
6
- # Subagent-Driven Development
7
-
8
- Execute an implementation plan with mostly independent tasks by dispatching a fresh implementer subagent per task, a task review (spec compliance + code quality) after each, and a broad whole-branch review at the end. This is an alternative executor to `/implement`, not a step that runs after `/implement`.
9
-
10
- **Why subagents:** You delegate tasks to specialized agents with isolated context. By precisely crafting their instructions and context, you ensure they stay focused and succeed at their task. They should never inherit your session's context or history — you construct exactly what they need. This also preserves your own context for coordination work.
11
-
12
- **Core principle:** Fresh subagent per task + task review (spec + quality) + broad final review = high quality, fast iteration
13
-
14
- **Narration:** between tool calls, narrate at most one short line — the
15
- ledger and the tool results carry the record.
16
-
17
- **Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are the four named below, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.
18
-
19
- **Rulings, not stalls.** A running plan does not wait on a human. Conflicts,
20
- ambiguities, plan defects, a cap you would have asked to exceed — decide
21
- them. The spec is the binding authority, the plan is its argument, and your
22
- judgment settles what neither answers. Record every decision in the ledger as
23
- `Ruling: <what you decided> — <why> — <what it costs if wrong>`, and keep
24
- going. A wrong ruling costs rework your human partner can see and undo; a
25
- session parked on a question costs their whole day and buys nothing.
26
-
27
- Four things stop you, and only these: an irreversible or destructive
28
- operation; a security-sensitive action; a side effect outside this worktree
29
- that norms say you ask about first (a merge, a push to a shared branch, a
30
- publish); and a plan so broken that every path forward is a guess. For those,
31
- stop and ask.
32
-
33
- ## When to Use
34
-
35
- ```dot
36
- digraph when_to_use {
37
- "Have implementation plan?" [shape=diamond];
38
- "Tasks mostly independent?" [shape=diamond];
39
- "Partner chose inline, or no subagent tool?" [shape=diamond];
40
- "subagent-driven-development" [shape=box];
41
- "/implement" [shape=box];
42
- "Manual execution or brainstorm first" [shape=box];
43
-
44
- "Have implementation plan?" -> "Tasks mostly independent?" [label="yes"];
45
- "Have implementation plan?" -> "Manual execution or brainstorm first" [label="no"];
46
- "Tasks mostly independent?" -> "Partner chose inline, or no subagent tool?" [label="yes"];
47
- "Tasks mostly independent?" -> "Manual execution or brainstorm first" [label="no - tightly coupled"];
48
- "Partner chose inline, or no subagent tool?" -> "/implement" [label="yes"];
49
- "Partner chose inline, or no subagent tool?" -> "subagent-driven-development" [label="no"];
50
- }
51
- ```
52
-
53
- **vs. Executing Plans (inline):**
54
- - Fresh subagent per task (no context pollution) instead of one context doing every task
55
- - Review after each task (spec compliance + code quality) instead of only at the end
56
- - Costs a fresh context per task and per review; inline costs one context plus one final reviewer
57
- - Both run in this session, share the same plan workspace and ledger, and never pause between tasks
58
-
59
- If the host cannot dispatch subagents, route to `/implement` and do not claim that subagents ran. SDD is not a reason to force a worktree or to run `/implement` twice.
60
-
61
- ## The Process
62
-
63
- ```dot
64
- digraph process {
65
- rankdir=TB;
66
-
67
- subgraph cluster_per_task {
68
- label="Per Task";
69
- "Dispatch implementer subagent (./implementer-prompt.md)" [shape=box];
70
- "Implementer asks questions?" [shape=diamond];
71
- "Answer questions, provide context" [shape=box];
72
- "Implementer implements, tests, commits, self-reviews" [shape=box];
73
- "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" [shape=box];
74
- "Spec ✅ and quality approved?" [shape=diamond];
75
- "Finding conflicts with plan text?" [shape=diamond];
76
- "Rule on the conflict, ledger the ruling" [shape=box];
77
- "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [shape=box];
78
- "Dispatch scoped re-review (./re-review-prompt.md)" [shape=box];
79
- "All findings addressed?" [shape=diamond];
80
- "R = 5?" [shape=diamond];
81
- "Adjudicate each open finding" [shape=box];
82
- "Any load-bearing finding?" [shape=diamond];
83
- "Rule and continue; stop only if every path forward is a guess" [shape=box];
84
- "Park findings in ledger with rulings" [shape=box];
85
- "Append completion to ledger, mark todo complete" [shape=box];
86
- }
87
-
88
- "Setup: worktree, ledger check, read plan, pre-flight review" [shape=box];
89
- "More tasks remain?" [shape=diamond];
90
- "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" [shape=box];
91
- "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" [shape=box];
92
- "Final review clean: delete this plan's workspace" [shape=box];
93
- "Use finishing-a-development-branch" [shape=box style=filled fillcolor=lightgreen];
94
-
95
- "Setup: worktree, ledger check, read plan, pre-flight review" -> "Dispatch implementer subagent (./implementer-prompt.md)";
96
- "Dispatch implementer subagent (./implementer-prompt.md)" -> "Implementer asks questions?";
97
- "Implementer asks questions?" -> "Answer questions, provide context" [label="yes"];
98
- "Answer questions, provide context" -> "Implementer implements, tests, commits, self-reviews";
99
- "Implementer asks questions?" -> "Implementer implements, tests, commits, self-reviews" [label="no"];
100
- "Implementer implements, tests, commits, self-reviews" -> "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)";
101
- "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" -> "Spec ✅ and quality approved?";
102
- "Spec ✅ and quality approved?" -> "Append completion to ledger, mark todo complete" [label="yes"];
103
- "Spec ✅ and quality approved?" -> "Finding conflicts with plan text?" [label="no"];
104
- "Finding conflicts with plan text?" -> "Rule on the conflict, ledger the ruling" [label="yes"];
105
- "Rule on the conflict, ledger the ruling" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model";
106
- "Finding conflicts with plan text?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no"];
107
- "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" -> "Dispatch scoped re-review (./re-review-prompt.md)";
108
- "Dispatch scoped re-review (./re-review-prompt.md)" -> "All findings addressed?";
109
- "All findings addressed?" -> "Append completion to ledger, mark todo complete" [label="yes"];
110
- "All findings addressed?" -> "R = 5?" [label="no"];
111
- "R = 5?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no - next round"];
112
- "R = 5?" -> "Adjudicate each open finding" [label="yes - breaker trips"];
113
- "Adjudicate each open finding" -> "Any load-bearing finding?";
114
- "Any load-bearing finding?" -> "Rule and continue; stop only if every path forward is a guess" [label="yes"];
115
- "Any load-bearing finding?" -> "Park findings in ledger with rulings" [label="no"];
116
- "Park findings in ledger with rulings" -> "Append completion to ledger, mark todo complete";
117
- "Append completion to ledger, mark todo complete" -> "More tasks remain?";
118
- "More tasks remain?" -> "Dispatch implementer subagent (./implementer-prompt.md)" [label="yes"];
119
- "More tasks remain?" -> "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" [label="no"];
120
- "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" -> "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals";
121
- "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" -> "Final review clean: delete this plan's workspace";
122
- "Final review clean: delete this plan's workspace" -> "Use finishing-a-development-branch";
123
- }
124
- ```
125
-
126
- ## Setup
127
-
128
- Determine the intended workspace before implementation. If the user explicitly
129
- wants the current branch/workspace, use it. If already in an isolated or
130
- host-managed worktree, reuse it. If the user requests isolation, use
131
- `using-git-worktrees`. If intent is unspecified, stay in the current workspace
132
- by default; recommend or ask about isolation for large or risky work, but do
133
- not create a worktree solely because SDD is being used. Do not apply a blanket
134
- "never implement on main/master" rule when the user explicitly requests direct
135
- work there.
136
-
137
- Conversation memory does not survive compaction. In real sessions,
138
- controllers that lost their place have re-dispatched entire completed task
139
- sequences — the single most expensive failure observed. Track progress in
140
- a ledger file, not only in todos.
141
-
142
- - Each plan owns a workspace: at skill start, run this skill's
143
- `node scripts/sdd-workspace.js PLAN_FILE` — it prints the plan's git-ignored
144
- directory (under `<repo-root>/.superpowers/sdd/`), home to
145
- every artifact for THIS plan: ledger, briefs, reports, review packages.
146
- Another plan's directory is never yours to read or write.
147
- - Check for this plan's ledger at `<workspace>/progress.md`. If its first
148
- line names your plan file, tasks with a `Task <N>: complete` line are DONE
149
- — do not re-dispatch them; resume at the first task without one. A task
150
- whose last line is a fix round is mid-loop: resume the loop at the next
151
- round. A ledger whose first line names a different plan file — or a stray
152
- ledger at the old flat path `.superpowers/sdd/progress.md` — is another
153
- plan's progress: leave it in place and start your own, fresh.
154
- - Create the ledger with its identity as the first line:
155
- `# SDD ledger — plan: <plan file path>`.
156
- - The ledger is your recovery map: the commits it names exist in git even
157
- when your context no longer remembers creating them. After compaction,
158
- trust the ledger and `git log` over your own recollection.
159
- - `git clean -fdx` will destroy the workspace (it's git-ignored scratch); if
160
- that happens, recover from `git log`.
161
-
162
- Read the plan once, note its context and Global Constraints, and create a
163
- todo per task. If the plan names a Spec, read that too: the spec is the
164
- authority the plan argues from, and conflicts inside the plan resolve
165
- against it. A plan with no reachable spec gets a ledger note saying so —
166
- rulings made without one are provisional.
167
-
168
- Before dispatching Task 1, scan the plan once for conflicts, writing down
169
- what you checked as you check it:
170
-
171
- - tasks that contradict each other or the plan's Global Constraints
172
- - anything the plan explicitly mandates that the review rubric treats as a
173
- defect (a test that asserts nothing, verbatim duplication of a logic block)
174
-
175
- The scan's output is a table, not a verdict. One row for every pair of tasks
176
- that share a file or an interface: the two tasks, what one produces against
177
- what the other consumes, and what you found. One row for every task: whether
178
- its own text agrees with itself — the tests it specifies against the code it
179
- specifies, the files it creates against the files it later touches. "The scan
180
- is clean" without those rows is not a scan you ran.
181
-
182
- Write the table to the ledger. Rule on everything you find before execution
183
- begins — each finding against the plan text that mandates it — and record
184
- each ruling in the ledger. If the scan is clean, proceed without comment.
185
- Rule on each conflict it surfaces — the spec is the binding authority, the
186
- plan is its argument — record the ruling beside its row, and dispatch
187
- Task 1. The review loop remains the net for conflicts that only emerge from
188
- implementation.
189
-
190
- ## Model Selection
191
-
192
- Use the least powerful model that can handle each role to conserve cost and increase speed.
193
-
194
- **Mechanical implementation tasks** (isolated functions, clear specs, 1-2 files): use a fast, cheap model. Most implementation tasks are mechanical when the plan is well-specified.
195
-
196
- **Integration and judgment tasks** (multi-file coordination, pattern matching, debugging): use a standard model.
197
-
198
- **Architecture and design tasks**: use the most capable available model.
199
- The final whole-branch review is one of these — dispatch it on the most
200
- capable available model, not the session default.
201
-
202
- **Review tasks**: choose the model with the same judgment, scaled to the
203
- diff's size, complexity, and risk. A small mechanical diff does not need the
204
- most capable model; a subtle concurrency change does. Scoped re-reviews of
205
- small fix diffs take a cheap-to-mid tier.
206
-
207
- **Fix-loop escalation (rounds 4-5)**: use a model at least one tier above
208
- the implementer that got stuck.
209
-
210
- **Always specify the model explicitly when dispatching a subagent.** An
211
- omitted model inherits your session's model — often the most capable and
212
- most expensive — which silently defeats this section.
213
-
214
- **Turn count beats token price.** Wall-clock and context cost scale with how
215
- many turns a subagent takes, and the cheapest models routinely take 2-3× the
216
- turns on multi-step work — costing more overall. Use a mid-tier model as the
217
- floor for reviewers and for implementers working from prose descriptions.
218
- When the task's plan text contains the complete code to write, the
219
- implementation is transcription plus testing: use the cheapest tier for
220
- that implementer. Single-file mechanical fixes also take the cheapest tier.
221
-
222
- **Task complexity signals (implementation tasks):**
223
- - Touches 1-2 files with a complete spec → cheap model
224
- - Touches multiple files with integration concerns → standard model
225
- - Requires design judgment or broad codebase understanding → most capable model
226
-
227
- ## The Task Loop
228
-
229
- **Batch small same-shape work.** When the plan lists several tasks that are
230
- each a small, independent edit of the same kind — the same one-line fix,
231
- constant change, or field addition repeated across files — do not dispatch
232
- one subagent per task. Compose ONE dispatch brief listing every file and
233
- its change, send the whole batch to a single subagent, and review its diff
234
- as one unit. Reserve one-dispatch-per-task for work that needs its own
235
- judgment, its own tests, or its own review surface.
236
-
237
- Everything you paste into a dispatch prompt — and everything a subagent
238
- prints back — stays resident in your context for the rest of the session
239
- and is re-read on every later turn. Hand artifacts over as files.
240
-
241
- **Waiting on dispatched subagents:** never poll a wait interface with
242
- short timeouts, and never sit in one silent, open-ended wait either.
243
- While you have local work — ledger updates, packaging the next review,
244
- reading reports — keep working; child results arrive on their own.
245
- When you are genuinely idle, wait in bounded stretches (five to ten
246
- minutes, where your platform allows), and between stretches post one
247
- line of status and reconcile your live children: list them, and chase
248
- any that finished without reporting. A bounded stretch keeps nearly
249
- all of a long wait's efficiency while guaranteeing a stuck or lost
250
- child is noticed within minutes, not at the end of the session.
251
-
252
- ### 1. Dispatch the implementer
253
-
254
- Record BASE (`git rev-parse HEAD`) before dispatching — the review package
255
- and fix-round diffs need it.
256
-
257
- - **Task brief:** before dispatching an implementer, run this skill's
258
- `node scripts/task-brief.js PLAN_FILE N` — it extracts the task's full text to a
259
- uniquely named file and prints the path. Compose the dispatch so the
260
- brief stays the single source of
261
- requirements. Your dispatch should contain: (1) one line on where this
262
- task fits in the project; (2) the brief path, introduced as "read this
263
- first — it is your requirements, with the exact values to use verbatim";
264
- (3) interfaces and decisions from earlier tasks that the brief cannot
265
- know; (4) your resolution of any ambiguity you noticed in the brief;
266
- (5) the report-file path and report contract. Exact values (numbers,
267
- magic strings, signatures, test cases) appear only in the brief. Never
268
- make a subagent read the whole plan file.
269
- - **Report file:** name the implementer's report file after the brief
270
- (brief `…/task-N-brief.md` → report `…/task-N-report.md`) and put it in
271
- the dispatch prompt. The implementer writes the full report there and
272
- returns only status, commits, a one-line test summary, and concerns.
273
- - A dispatch prompt describes one task, not the session's history. Do not
274
- paste accumulated prior-task summaries ("state after Tasks 1-3") into
275
- later dispatches — a real session's dispatch hit 42k chars of which 99%
276
- was pasted history. A fresh subagent needs its task, the interfaces it
277
- touches, and the global constraints. Nothing else.
278
- - The dispatch carries the no-subagents contract (it is in the
279
- implementer template): the implementer never dispatches subagents —
280
- not helpers, and never a reviewer. Review arrives from you, after the
281
- report. In real sessions, every reviewer a worker spawned duplicated
282
- the task review the controller dispatched anyway — a full extra
283
- review seat per task.
284
- - If an earlier task parked a finding in the area this task touches, carry
285
- a pointer to that ledger entry in the dispatch.
286
- - Record the implementer's agent identity from the dispatch result —
287
- fix-loop rounds 1-3 resume this agent.
288
- - Never dispatch multiple implementation subagents in parallel (conflicts).
289
-
290
- Template: [implementer-prompt.md](implementer-prompt.md)
291
-
292
- ### 2. Handle the report
293
-
294
- Implementer subagents report one of four statuses. Handle each appropriately:
295
-
296
- **DONE:** Generate the review package (`node scripts/review-package.js PLAN_FILE BASE HEAD`, from this skill's directory — it prints the unique file path it wrote; BASE is the commit you recorded before dispatching the implementer — never `HEAD~1`, which silently drops all but the last commit of a multi-commit task), then dispatch the task reviewer with the printed path.
297
-
298
- **DONE_WITH_CONCERNS:** The implementer completed the work but flagged doubts. Read the concerns before proceeding. If the concerns are about correctness or scope, address them before review. If they're observations (e.g., "this file is getting large"), note them and proceed to review.
299
-
300
- **NEEDS_CONTEXT:** The implementer needs information that wasn't provided. Provide the missing context and re-dispatch.
301
-
302
- **BLOCKED:** The implementer cannot complete the task. Assess the blocker:
303
- 1. If it's a context problem, provide more context and re-dispatch with the same model
304
- 2. If the task requires more reasoning, re-dispatch with a more capable model
305
- 3. If the task is too large, break it into smaller pieces
306
- 4. If the plan itself is wrong, rule on the correction, ledger it, and re-dispatch with the ruling carried in the dispatch
307
-
308
- **Never** ignore an escalation or force the same model to retry without changes. If the implementer said it's stuck, something needs to change.
309
-
310
- If the implementer asks questions — before starting or mid-task — answer
311
- clearly and completely, provide additional context if needed, and don't
312
- rush it into implementation.
313
-
314
- ### 3. Review the task
315
-
316
- Per-task reviews are task-scoped gates. The broad review happens once, at the
317
- final whole-branch review. Never skip the task review, and never accept a
318
- report missing either verdict — spec compliance AND task quality are both
319
- required. Implementer self-review never replaces the task review; both are
320
- needed.
321
-
322
- - Hand the reviewer its diff as a file: run this skill's
323
- `node scripts/review-package.js PLAN_FILE BASE HEAD` and pass the reviewer the file path
324
- it prints (or, without bash: `git log --oneline`, `git diff --stat`,
325
- and `git diff -U10` for the range, redirected to one uniquely named
326
- file). The output never enters your own context, and the reviewer sees
327
- the commit list, stat summary, and full diff with context in one Read
328
- call. Use the BASE you recorded before dispatching the implementer —
329
- never `HEAD~1`, which silently truncates multi-commit tasks. Never
330
- dispatch a task reviewer without a diff file.
331
- - **Reviewer inputs:** the task reviewer gets three paths — the same brief
332
- file, the report file, and the review package — plus the global
333
- constraints that bind the task.
334
- - The global-constraints block you hand the reviewer is its attention
335
- lens. Copy the binding requirements verbatim from the plan's Global
336
- Constraints section or the spec: exact values, exact formats, and the
337
- stated relationships between components ("same layout as X", "matches
338
- Y"). The reviewer's template already carries the process rules (YAGNI,
339
- test hygiene, review method) — the constraints block is for what THIS
340
- project's spec demands.
341
- - Do not add open-ended directives like "check all uses" or "run race tests
342
- if useful" without a concrete, task-specific reason
343
- - Do not ask a reviewer to re-run tests the implementer already ran on the
344
- same code — the implementer's report carries the test evidence
345
- - Do not pre-judge findings for the reviewer — never instruct a reviewer to
346
- ignore or not flag a specific issue. If you believe a finding would be a
347
- false positive, let the reviewer raise it and adjudicate it in the review
348
- loop. If the prompt you are writing contains "do not flag," "don't treat X
349
- as a defect," "at most Minor," or "the plan chose" — stop: you are
350
- pre-judging, usually to spare yourself a review loop.
351
- The task reviewer may report "⚠️ Cannot verify from diff" items — requirements
352
- that live in unchanged code or span tasks. These do not block the rest of the
353
- review, but you must resolve each one yourself before marking the task
354
- complete: you hold the plan and cross-task context the reviewer
355
- lacks. If you confirm an item is a real gap, treat it as a failed spec
356
- review — it enters the fix loop with the other findings.
357
-
358
- Template: [task-reviewer-prompt.md](task-reviewer-prompt.md)
359
-
360
- ### 4. The fix loop
361
-
362
- The loop triggers when the review reports spec ❌, any Critical or Important
363
- finding, or a ⚠️ item you confirmed as a real gap.
364
-
365
- Before the loop starts, two routes leave it immediately:
366
-
367
- - Record Minor findings in the progress ledger as you go
368
- (`Task <N>: minor (deferred): <one-liner>`), and point the final
369
- whole-branch review at that list so it can triage which must be fixed
370
- before merge. A roll-up nobody reads is a silent discard. Minor findings
371
- never enter the loop.
372
- - A finding labeled plan-mandated — or any finding that conflicts with
373
- what the plan's text requires — is yours to rule on: weigh the finding
374
- against the plan text, decide with the spec as the binding authority, and
375
- ledger the ruling before you act on it. Do not dismiss the finding because
376
- the plan mandates it, and do not dispatch a fix that contradicts the plan
377
- without a recorded ruling.
378
- Everything else enters the loop. A fix round is one fix dispatch plus one
379
- scoped re-review. Five rounds maximum per task:
380
-
381
- **Rounds 1-3 — resume the original implementer.** Send it the open findings
382
- verbatim. Its context is intact: it knows the task, the code, and its own
383
- choices. If your harness cannot send another message to a live subagent,
384
- dispatch a fresh implementer carrying the brief path, the report-file path,
385
- and the findings — the report file is the persistent memory either way.
386
-
387
- **Rounds 4-5 — dispatch a fresh implementer on a more capable model** (per
388
- Model Selection), with the brief path, the report-file path, the open
389
- findings, and this framing: "A prior implementer attempted this task
390
- [N] times; you own it now. Read the report file for what was tried." A loop
391
- that survives three resumes usually means the implementer cannot see its
392
- own problem — fresh eyes and a capability bump in one move.
393
-
394
- **Every round, either way:** the implementer fixes, re-runs the tests
395
- covering the amended code, appends its fix report to the same report file,
396
- and returns the short contract. Before re-dispatching the reviewer, confirm
397
- the fix report contains the covering tests, the command run, and the
398
- output; dispatch the re-review once all three are present. Name the
399
- covering test files in the fix message — a one-line fix does not need the
400
- whole suite.
401
-
402
- **The re-review is scoped.** Run `node scripts/review-package.js PLAN_FILE FIX_BASE HEAD`
403
- where FIX_BASE is the head the previous review saw, and dispatch
404
- [re-review-prompt.md](re-review-prompt.md) with the findings list, the
405
- brief, the report file, and the printed diff path. The re-reviewer verdicts
406
- each finding ADDRESSED or NOT ADDRESSED and flags new breakage in the fix
407
- diff only. New Critical/Important breakage in the fix diff joins the open
408
- findings list. Out-of-scope observations go to the ledger as deferred
409
- minors — they never extend the loop.
410
-
411
- **After each round,** append to the ledger:
412
- `Task <N>: fix round <R>/5 (<X> addressed, <Y> open — <finding one-liners>; commits <a7>..<b7>)`
413
-
414
- Never fix findings yourself in the controller session — your context stays
415
- clean for coordination, and controller fixes skip review.
416
-
417
- **The breaker.** When round 5's re-review still leaves findings open, stop
418
- dispatching. Adjudicate each open finding yourself — you hold the plan and
419
- the cross-task context the reviewer lacks:
420
-
421
- - **The reviewer is wrong, or the point is contestable:** park it —
422
- `Task <N>: parked — <finding> — Ruling: <why the code stands>`. The final
423
- review sees both sides.
424
- - **Real, but nothing downstream builds on it:** park it the same way, with
425
- a ruling that says it's real and deferred.
426
- - **Real and load-bearing** — a later task builds on it, or it reveals a
427
- plan defect: rule on the smallest change that unblocks the dependent work,
428
- ledger it as `Task <N>: Ruling: <finding> — <what you decided and why>`,
429
- and carry it into the next task's dispatch. Parking a structural failure
430
- silently lets every dependent task build on it. Stop only when the defect
431
- leaves every path forward a guess.
432
-
433
- Adjudicate only at the cap. Adjudicating earlier to end a loop is
434
- pre-judging with a different name. Every adjudication is a ledger entry —
435
- a silent discard is forbidden.
436
-
437
- ### 5. Complete the task
438
-
439
- When the review comes back clean — or every open finding is parked with a
440
- ruling at the cap — append the completion line to the ledger in the same
441
- message as your other bookkeeping:
442
-
443
- - `Task <N>: complete (commits <base7>..<head7>, review clean)`
444
- - `Task <N>: complete (commits <base7>..<head7>, <K> parked)` after a
445
- tripped breaker
446
-
447
- Then mark the todo complete and move on. Never move to the next task while
448
- the review has open Critical/Important issues that are neither fixed nor
449
- parked-with-ruling at the cap.
450
-
451
- ## Final Review
452
-
453
- The final whole-branch review gets a package too: run
454
- `node scripts/review-package.js PLAN_FILE MERGE_BASE HEAD` (MERGE_BASE = the commit the
455
- branch started from, e.g. `git merge-base main HEAD`) and include the
456
- printed path in the final review dispatch, so the final reviewer reads
457
- one file instead of re-deriving the branch diff with git commands. Dispatch
458
- on the most capable available model (see Model Selection), using
459
- the bundled requesting-code-review support template
460
- [code-reviewer.md](../requesting-code-review/code-reviewer.md). Point it at
461
- the ledger's deferred-minor and parked lines so it can triage which must be
462
- fixed before merge.
463
-
464
- If the final whole-branch review returns findings, dispatch ONE fix subagent
465
- with the complete findings list — not one fixer per finding.
466
- Per-finding fixers each rebuild context and re-run suites; a real
467
- session's final-review fix wave cost more than all its tasks combined.
468
- Then run exactly one scoped re-review of the fix wave
469
- (`node scripts/review-package.js PLAN_FILE FIX_BASE HEAD` over the fix range,
470
- [re-review-prompt.md](re-review-prompt.md)).
471
- Adjudicate any residual findings as in the task loop's breaker: park with
472
- rulings, or rule on the load-bearing ones and ledger what you decided. Only
473
- the four classes above stop you here. There is no second fix wave —
474
- residual load-bearing findings surface to your human partner when
475
- finishing-a-development-branch presents the options.
476
-
477
- ## Finish
478
-
479
- Before you delete anything, collect every ledger line containing `Ruling:` —
480
- preflight rulings, parked findings, breaker adjudications, all of them — into
481
- your final message under "Rulings I made", in the order you made them, each
482
- with what it costs if wrong. The list is exhaustive: if the ledger holds a
483
- ruling, the list holds it. That list is the only place the decisions you
484
- took on your human partner's behalf reach them — they read it and rework
485
- whatever you got wrong. A ruling that dies with the workspace was a decision
486
- made in secret.
487
-
488
- When the final whole-branch review is clean and its fixes are merged,
489
- delete this plan's workspace (`rm -rf <workspace>`) — the git history is
490
- the record now. Sibling directories belong to other plans; leave them
491
- alone.
492
-
493
- Use `finishing-a-development-branch`.
494
-
495
- ## Common Rationalizations
496
-
497
- | Excuse | Reality |
498
- |--------|---------|
499
- | "Close enough on spec compliance" | Reviewer found spec gaps = not done. Fix or hit the cap and adjudicate — those are the only exits. |
500
- | "I'll fix it myself, dispatching is overhead" | Controller fixes pollute your context and skip review. Resume the implementer. |
501
- | "One more round will converge" | Past the cap, rounds don't converge — the failure is structural. Adjudicate and route. |
502
- | "The reviewer will just find something new anyway" | Scoped re-reviews verify fixes; they cannot wander. New findings on untouched code go to the ledger, not the loop. |
503
- | "This finding is obviously wrong, I'll drop it" | You adjudicate only at the cap, and every ruling is a ledger entry. Silent discards are forbidden. |
504
- | "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
505
- | "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. |
506
- | "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. |
507
- | "The implementer spawned its own reviewer — free extra assurance" | It's a duplicate seat reviewing the same diff; the task review is the gate. A worker-spawned reviewer is a defect to flag, not rigor. |
508
-
509
- ## Example Workflow
510
-
511
- ```
512
- You: I'm using Subagent-Driven Development to execute this plan.
513
-
514
- [Setup: worktree verified]
515
- [Read plan file once: docs/superpowers/plans/feature-plan.md]
516
- [Resolve workspace: node scripts/sdd-workspace.js docs/superpowers/plans/feature-plan.md — no ledger inside, fresh start]
517
- [Create todos for all tasks]
518
-
519
- Task 1: Hook installation script
520
-
521
- [Run task-brief for Task 1; dispatch implementer with brief + report paths + context]
522
-
523
- Implementer: "Before I begin - should the hook be installed at user or system level?"
524
-
525
- You: "User level (~/.config/superpowers/hooks/)"
526
-
527
- Implementer: [Later]
528
- - Implemented install-hook command
529
- - Added tests, 5/5 passing
530
- - Self-review: Found I missed --force flag, added it
531
- - Committed
532
-
533
- [Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path]
534
- Task reviewer: Spec ✅ - all requirements met, nothing extra.
535
- Strengths: Good test coverage, clean. Issues: None. Task quality: Approved.
536
-
537
- [Ledger: Task 1: complete (commits a1b2c3d..d4e5f6a, review clean)]
538
-
539
- Task 2: Recovery modes
540
-
541
- [Run task-brief for Task 2; dispatch implementer with brief + report paths + context]
542
-
543
- Implementer: [No questions]
544
- - Added verify/repair modes
545
- - 8/8 tests passing
546
- - Committed
547
-
548
- [Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path]
549
- Task reviewer: Spec ❌:
550
- - Missing: Progress reporting (spec says "report every 100 items")
551
- Issues (Important): Magic number (100)
552
-
553
- [Fix round 1: resume the implementer with both findings]
554
- Implementer: Added progress reporting, extracted PROGRESS_INTERVAL constant.
555
- Re-ran test/recovery.test.js — 10/10 passing. Fix report appended.
556
-
557
- [Run review-package PLAN_FILE FIX_BASE HEAD; dispatch scoped re-review]
558
- Re-reviewer: Missing progress reporting — ADDRESSED (src/recovery.js:41).
559
- Magic number — ADDRESSED (src/recovery.js:7). New breakage: none.
560
- Verdict: all findings addressed.
561
-
562
- [Ledger: Task 2: fix round 1/5 (2 addressed, 0 open; commits d4e5f6a..b7c8d9e)]
563
- [Ledger: Task 2: complete (commits d4e5f6a..b7c8d9e, review clean)]
564
-
565
- ...
566
-
567
- [After all tasks]
568
- [Run review-package PLAN_FILE MERGE_BASE HEAD; dispatch final code-reviewer, most capable model]
569
- Final reviewer: All requirements met. Deferred minors triaged: none block merge.
570
-
571
- [Delete this plan's workspace — the record now lives in git]
572
-
573
- Done! Using finishing-a-development-branch.
574
- ```
1
+ ---
2
+ name: subagent-driven-development
3
+ description: Use when executing implementation plans with independent tasks in the current session
4
+ ---
5
+
6
+ # Subagent-Driven Development
7
+
8
+ Execute an implementation plan with mostly independent tasks by dispatching a fresh implementer subagent per task, a task review (spec compliance + code quality) after each, and a broad whole-branch review at the end. This is an alternative executor to `/implement`, not a step that runs after `/implement`.
9
+
10
+ **Why subagents:** You delegate tasks to specialized agents with isolated context. By precisely crafting their instructions and context, you ensure they stay focused and succeed at their task. They should never inherit your session's context or history — you construct exactly what they need. This also preserves your own context for coordination work.
11
+
12
+ **Core principle:** Fresh subagent per task + task review (spec + quality) + broad final review = high quality, fast iteration
13
+
14
+ **Narration:** between tool calls, narrate at most one short line — the
15
+ ledger and the tool results carry the record.
16
+
17
+ **Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are the four named below, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.
18
+
19
+ **Rulings, not stalls.** A running plan does not wait on a human. Conflicts,
20
+ ambiguities, plan defects, a cap you would have asked to exceed — decide
21
+ them. The spec is the binding authority, the plan is its argument, and your
22
+ judgment settles what neither answers. Record every decision in the ledger as
23
+ `Ruling: <what you decided> — <why> — <what it costs if wrong>`, and keep
24
+ going. A wrong ruling costs rework your human partner can see and undo; a
25
+ session parked on a question costs their whole day and buys nothing.
26
+
27
+ Four things stop you, and only these: an irreversible or destructive
28
+ operation; a security-sensitive action; a side effect outside this worktree
29
+ that norms say you ask about first (a merge, a push to a shared branch, a
30
+ publish); and a plan so broken that every path forward is a guess. For those,
31
+ stop and ask.
32
+
33
+ ## When to Use
34
+
35
+ ```dot
36
+ digraph when_to_use {
37
+ "Have implementation plan?" [shape=diamond];
38
+ "Tasks mostly independent?" [shape=diamond];
39
+ "Partner chose inline, or no subagent tool?" [shape=diamond];
40
+ "subagent-driven-development" [shape=box];
41
+ "/implement" [shape=box];
42
+ "Manual execution or brainstorm first" [shape=box];
43
+
44
+ "Have implementation plan?" -> "Tasks mostly independent?" [label="yes"];
45
+ "Have implementation plan?" -> "Manual execution or brainstorm first" [label="no"];
46
+ "Tasks mostly independent?" -> "Partner chose inline, or no subagent tool?" [label="yes"];
47
+ "Tasks mostly independent?" -> "Manual execution or brainstorm first" [label="no - tightly coupled"];
48
+ "Partner chose inline, or no subagent tool?" -> "/implement" [label="yes"];
49
+ "Partner chose inline, or no subagent tool?" -> "subagent-driven-development" [label="no"];
50
+ }
51
+ ```
52
+
53
+ **vs. Executing Plans (inline):**
54
+ - Fresh subagent per task (no context pollution) instead of one context doing every task
55
+ - Review after each task (spec compliance + code quality) instead of only at the end
56
+ - Costs a fresh context per task and per review; inline costs one context plus one final reviewer
57
+ - Both run in this session, share the same plan workspace and ledger, and never pause between tasks
58
+
59
+ If the host cannot dispatch subagents, route to `/implement` and do not claim that subagents ran. SDD is not a reason to force a worktree or to run `/implement` twice.
60
+
61
+ ## The Process
62
+
63
+ ```dot
64
+ digraph process {
65
+ rankdir=TB;
66
+
67
+ subgraph cluster_per_task {
68
+ label="Per Task";
69
+ "Dispatch implementer subagent (./implementer-prompt.md)" [shape=box];
70
+ "Implementer asks questions?" [shape=diamond];
71
+ "Answer questions, provide context" [shape=box];
72
+ "Implementer implements, tests, commits, self-reviews" [shape=box];
73
+ "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" [shape=box];
74
+ "Spec ✅ and quality approved?" [shape=diamond];
75
+ "Finding conflicts with plan text?" [shape=diamond];
76
+ "Rule on the conflict, ledger the ruling" [shape=box];
77
+ "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [shape=box];
78
+ "Dispatch scoped re-review (./re-review-prompt.md)" [shape=box];
79
+ "All findings addressed?" [shape=diamond];
80
+ "R = 5?" [shape=diamond];
81
+ "Adjudicate each open finding" [shape=box];
82
+ "Any load-bearing finding?" [shape=diamond];
83
+ "Rule and continue; stop only if every path forward is a guess" [shape=box];
84
+ "Park findings in ledger with rulings" [shape=box];
85
+ "Append completion to ledger, mark todo complete" [shape=box];
86
+ }
87
+
88
+ "Setup: worktree, ledger check, read plan, pre-flight review" [shape=box];
89
+ "More tasks remain?" [shape=diamond];
90
+ "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" [shape=box];
91
+ "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" [shape=box];
92
+ "Final review clean: delete this plan's workspace" [shape=box];
93
+ "Use finishing-a-development-branch" [shape=box style=filled fillcolor=lightgreen];
94
+
95
+ "Setup: worktree, ledger check, read plan, pre-flight review" -> "Dispatch implementer subagent (./implementer-prompt.md)";
96
+ "Dispatch implementer subagent (./implementer-prompt.md)" -> "Implementer asks questions?";
97
+ "Implementer asks questions?" -> "Answer questions, provide context" [label="yes"];
98
+ "Answer questions, provide context" -> "Implementer implements, tests, commits, self-reviews";
99
+ "Implementer asks questions?" -> "Implementer implements, tests, commits, self-reviews" [label="no"];
100
+ "Implementer implements, tests, commits, self-reviews" -> "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)";
101
+ "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" -> "Spec ✅ and quality approved?";
102
+ "Spec ✅ and quality approved?" -> "Append completion to ledger, mark todo complete" [label="yes"];
103
+ "Spec ✅ and quality approved?" -> "Finding conflicts with plan text?" [label="no"];
104
+ "Finding conflicts with plan text?" -> "Rule on the conflict, ledger the ruling" [label="yes"];
105
+ "Rule on the conflict, ledger the ruling" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model";
106
+ "Finding conflicts with plan text?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no"];
107
+ "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" -> "Dispatch scoped re-review (./re-review-prompt.md)";
108
+ "Dispatch scoped re-review (./re-review-prompt.md)" -> "All findings addressed?";
109
+ "All findings addressed?" -> "Append completion to ledger, mark todo complete" [label="yes"];
110
+ "All findings addressed?" -> "R = 5?" [label="no"];
111
+ "R = 5?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no - next round"];
112
+ "R = 5?" -> "Adjudicate each open finding" [label="yes - breaker trips"];
113
+ "Adjudicate each open finding" -> "Any load-bearing finding?";
114
+ "Any load-bearing finding?" -> "Rule and continue; stop only if every path forward is a guess" [label="yes"];
115
+ "Any load-bearing finding?" -> "Park findings in ledger with rulings" [label="no"];
116
+ "Park findings in ledger with rulings" -> "Append completion to ledger, mark todo complete";
117
+ "Append completion to ledger, mark todo complete" -> "More tasks remain?";
118
+ "More tasks remain?" -> "Dispatch implementer subagent (./implementer-prompt.md)" [label="yes"];
119
+ "More tasks remain?" -> "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" [label="no"];
120
+ "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" -> "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals";
121
+ "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" -> "Final review clean: delete this plan's workspace";
122
+ "Final review clean: delete this plan's workspace" -> "Use finishing-a-development-branch";
123
+ }
124
+ ```
125
+
126
+ ## Setup
127
+
128
+ Determine the intended workspace before implementation. If the user explicitly
129
+ wants the current branch/workspace, use it. If already in an isolated or
130
+ host-managed worktree, reuse it. If the user requests isolation, use
131
+ `using-git-worktrees`. If intent is unspecified, stay in the current workspace
132
+ by default; recommend or ask about isolation for large or risky work, but do
133
+ not create a worktree solely because SDD is being used. Do not apply a blanket
134
+ "never implement on main/master" rule when the user explicitly requests direct
135
+ work there.
136
+
137
+ Conversation memory does not survive compaction. In real sessions,
138
+ controllers that lost their place have re-dispatched entire completed task
139
+ sequences — the single most expensive failure observed. Track progress in
140
+ a ledger file, not only in todos.
141
+
142
+ - Each plan owns a workspace: at skill start, run this skill's
143
+ `node scripts/sdd-workspace.js PLAN_FILE` — it prints the plan's git-ignored
144
+ directory (under `<repo-root>/.superpowers/sdd/`), home to
145
+ every artifact for THIS plan: ledger, briefs, reports, review packages.
146
+ Another plan's directory is never yours to read or write.
147
+ - Check for this plan's ledger at `<workspace>/progress.md`. If its first
148
+ line names your plan file, tasks with a `Task <N>: complete` line are DONE
149
+ — do not re-dispatch them; resume at the first task without one. A task
150
+ whose last line is a fix round is mid-loop: resume the loop at the next
151
+ round. A ledger whose first line names a different plan file — or a stray
152
+ ledger at the old flat path `.superpowers/sdd/progress.md` — is another
153
+ plan's progress: leave it in place and start your own, fresh.
154
+ - Create the ledger with its identity as the first line:
155
+ `# SDD ledger — plan: <plan file path>`.
156
+ - The ledger is your recovery map: the commits it names exist in git even
157
+ when your context no longer remembers creating them. After compaction,
158
+ trust the ledger and `git log` over your own recollection.
159
+ - `git clean -fdx` will destroy the workspace (it's git-ignored scratch); if
160
+ that happens, recover from `git log`.
161
+
162
+ Read the plan once, note its context and Global Constraints, and create a
163
+ todo per task. If the plan names a Spec, read that too: the spec is the
164
+ authority the plan argues from, and conflicts inside the plan resolve
165
+ against it. A plan with no reachable spec gets a ledger note saying so —
166
+ rulings made without one are provisional.
167
+
168
+ Before dispatching Task 1, scan the plan once for conflicts, writing down
169
+ what you checked as you check it:
170
+
171
+ - tasks that contradict each other or the plan's Global Constraints
172
+ - anything the plan explicitly mandates that the review rubric treats as a
173
+ defect (a test that asserts nothing, verbatim duplication of a logic block)
174
+
175
+ The scan's output is a table, not a verdict. One row for every pair of tasks
176
+ that share a file or an interface: the two tasks, what one produces against
177
+ what the other consumes, and what you found. One row for every task: whether
178
+ its own text agrees with itself — the tests it specifies against the code it
179
+ specifies, the files it creates against the files it later touches. "The scan
180
+ is clean" without those rows is not a scan you ran.
181
+
182
+ Write the table to the ledger. Rule on everything you find before execution
183
+ begins — each finding against the plan text that mandates it — and record
184
+ each ruling in the ledger. If the scan is clean, proceed without comment.
185
+ Rule on each conflict it surfaces — the spec is the binding authority, the
186
+ plan is its argument — record the ruling beside its row, and dispatch
187
+ Task 1. The review loop remains the net for conflicts that only emerge from
188
+ implementation.
189
+
190
+ ## Model Selection
191
+
192
+ Use the least powerful model that can handle each role to conserve cost and increase speed.
193
+
194
+ **Mechanical implementation tasks** (isolated functions, clear specs, 1-2 files): use a fast, cheap model. Most implementation tasks are mechanical when the plan is well-specified.
195
+
196
+ **Integration and judgment tasks** (multi-file coordination, pattern matching, debugging): use a standard model.
197
+
198
+ **Architecture and design tasks**: use the most capable available model.
199
+ The final whole-branch review is one of these — dispatch it on the most
200
+ capable available model, not the session default.
201
+
202
+ **Review tasks**: choose the model with the same judgment, scaled to the
203
+ diff's size, complexity, and risk. A small mechanical diff does not need the
204
+ most capable model; a subtle concurrency change does. Scoped re-reviews of
205
+ small fix diffs take a cheap-to-mid tier.
206
+
207
+ **Fix-loop escalation (rounds 4-5)**: use a model at least one tier above
208
+ the implementer that got stuck.
209
+
210
+ **Always specify the model explicitly when dispatching a subagent.** An
211
+ omitted model inherits your session's model — often the most capable and
212
+ most expensive — which silently defeats this section.
213
+
214
+ **Turn count beats token price.** Wall-clock and context cost scale with how
215
+ many turns a subagent takes, and the cheapest models routinely take 2-3× the
216
+ turns on multi-step work — costing more overall. Use a mid-tier model as the
217
+ floor for reviewers and for implementers working from prose descriptions.
218
+ When the task's plan text contains the complete code to write, the
219
+ implementation is transcription plus testing: use the cheapest tier for
220
+ that implementer. Single-file mechanical fixes also take the cheapest tier.
221
+
222
+ **Task complexity signals (implementation tasks):**
223
+ - Touches 1-2 files with a complete spec → cheap model
224
+ - Touches multiple files with integration concerns → standard model
225
+ - Requires design judgment or broad codebase understanding → most capable model
226
+
227
+ ## The Task Loop
228
+
229
+ **Batch small same-shape work.** When the plan lists several tasks that are
230
+ each a small, independent edit of the same kind — the same one-line fix,
231
+ constant change, or field addition repeated across files — do not dispatch
232
+ one subagent per task. Compose ONE dispatch brief listing every file and
233
+ its change, send the whole batch to a single subagent, and review its diff
234
+ as one unit. Reserve one-dispatch-per-task for work that needs its own
235
+ judgment, its own tests, or its own review surface.
236
+
237
+ Everything you paste into a dispatch prompt — and everything a subagent
238
+ prints back — stays resident in your context for the rest of the session
239
+ and is re-read on every later turn. Hand artifacts over as files.
240
+
241
+ **Waiting on dispatched subagents:** never poll a wait interface with
242
+ short timeouts, and never sit in one silent, open-ended wait either.
243
+ While you have local work — ledger updates, packaging the next review,
244
+ reading reports — keep working; child results arrive on their own.
245
+ When you are genuinely idle, wait in bounded stretches (five to ten
246
+ minutes, where your platform allows), and between stretches post one
247
+ line of status and reconcile your live children: list them, and chase
248
+ any that finished without reporting. A bounded stretch keeps nearly
249
+ all of a long wait's efficiency while guaranteeing a stuck or lost
250
+ child is noticed within minutes, not at the end of the session.
251
+
252
+ ### 1. Dispatch the implementer
253
+
254
+ Record BASE (`git rev-parse HEAD`) before dispatching — the review package
255
+ and fix-round diffs need it.
256
+
257
+ - **Task brief:** before dispatching an implementer, run this skill's
258
+ `node scripts/task-brief.js PLAN_FILE N` — it extracts the task's full text to a
259
+ uniquely named file and prints the path. Compose the dispatch so the
260
+ brief stays the single source of
261
+ requirements. Your dispatch should contain: (1) one line on where this
262
+ task fits in the project; (2) the brief path, introduced as "read this
263
+ first — it is your requirements, with the exact values to use verbatim";
264
+ (3) interfaces and decisions from earlier tasks that the brief cannot
265
+ know; (4) your resolution of any ambiguity you noticed in the brief;
266
+ (5) the report-file path and report contract. Exact values (numbers,
267
+ magic strings, signatures, test cases) appear only in the brief. Never
268
+ make a subagent read the whole plan file.
269
+ - **Report file:** name the implementer's report file after the brief
270
+ (brief `…/task-N-brief.md` → report `…/task-N-report.md`) and put it in
271
+ the dispatch prompt. The implementer writes the full report there and
272
+ returns only status, commits, a one-line test summary, and concerns.
273
+ - A dispatch prompt describes one task, not the session's history. Do not
274
+ paste accumulated prior-task summaries ("state after Tasks 1-3") into
275
+ later dispatches — a real session's dispatch hit 42k chars of which 99%
276
+ was pasted history. A fresh subagent needs its task, the interfaces it
277
+ touches, and the global constraints. Nothing else.
278
+ - The dispatch carries the no-subagents contract (it is in the
279
+ implementer template): the implementer never dispatches subagents —
280
+ not helpers, and never a reviewer. Review arrives from you, after the
281
+ report. In real sessions, every reviewer a worker spawned duplicated
282
+ the task review the controller dispatched anyway — a full extra
283
+ review seat per task.
284
+ - If an earlier task parked a finding in the area this task touches, carry
285
+ a pointer to that ledger entry in the dispatch.
286
+ - Record the implementer's agent identity from the dispatch result —
287
+ fix-loop rounds 1-3 resume this agent.
288
+ - Never dispatch multiple implementation subagents in parallel (conflicts).
289
+
290
+ Template: [implementer-prompt.md](implementer-prompt.md)
291
+
292
+ ### 2. Handle the report
293
+
294
+ Implementer subagents report one of four statuses. Handle each appropriately:
295
+
296
+ **DONE:** Generate the review package (`node scripts/review-package.js PLAN_FILE BASE HEAD`, from this skill's directory — it prints the unique file path it wrote; BASE is the commit you recorded before dispatching the implementer — never `HEAD~1`, which silently drops all but the last commit of a multi-commit task), then dispatch the task reviewer with the printed path.
297
+
298
+ **DONE_WITH_CONCERNS:** The implementer completed the work but flagged doubts. Read the concerns before proceeding. If the concerns are about correctness or scope, address them before review. If they're observations (e.g., "this file is getting large"), note them and proceed to review.
299
+
300
+ **NEEDS_CONTEXT:** The implementer needs information that wasn't provided. Provide the missing context and re-dispatch.
301
+
302
+ **BLOCKED:** The implementer cannot complete the task. Assess the blocker:
303
+ 1. If it's a context problem, provide more context and re-dispatch with the same model
304
+ 2. If the task requires more reasoning, re-dispatch with a more capable model
305
+ 3. If the task is too large, break it into smaller pieces
306
+ 4. If the plan itself is wrong, rule on the correction, ledger it, and re-dispatch with the ruling carried in the dispatch
307
+
308
+ **Never** ignore an escalation or force the same model to retry without changes. If the implementer said it's stuck, something needs to change.
309
+
310
+ If the implementer asks questions — before starting or mid-task — answer
311
+ clearly and completely, provide additional context if needed, and don't
312
+ rush it into implementation.
313
+
314
+ ### 3. Review the task
315
+
316
+ Per-task reviews are task-scoped gates. The broad review happens once, at the
317
+ final whole-branch review. Never skip the task review, and never accept a
318
+ report missing either verdict — spec compliance AND task quality are both
319
+ required. Implementer self-review never replaces the task review; both are
320
+ needed.
321
+
322
+ - Hand the reviewer its diff as a file: run this skill's
323
+ `node scripts/review-package.js PLAN_FILE BASE HEAD` and pass the reviewer the file path
324
+ it prints (or, without bash: `git log --oneline`, `git diff --stat`,
325
+ and `git diff -U10` for the range, redirected to one uniquely named
326
+ file). The output never enters your own context, and the reviewer sees
327
+ the commit list, stat summary, and full diff with context in one Read
328
+ call. Use the BASE you recorded before dispatching the implementer —
329
+ never `HEAD~1`, which silently truncates multi-commit tasks. Never
330
+ dispatch a task reviewer without a diff file.
331
+ - **Reviewer inputs:** the task reviewer gets three paths — the same brief
332
+ file, the report file, and the review package — plus the global
333
+ constraints that bind the task.
334
+ - The global-constraints block you hand the reviewer is its attention
335
+ lens. Copy the binding requirements verbatim from the plan's Global
336
+ Constraints section or the spec: exact values, exact formats, and the
337
+ stated relationships between components ("same layout as X", "matches
338
+ Y"). The reviewer's template already carries the process rules (YAGNI,
339
+ test hygiene, review method) — the constraints block is for what THIS
340
+ project's spec demands.
341
+ - Do not add open-ended directives like "check all uses" or "run race tests
342
+ if useful" without a concrete, task-specific reason
343
+ - Do not ask a reviewer to re-run tests the implementer already ran on the
344
+ same code — the implementer's report carries the test evidence
345
+ - Do not pre-judge findings for the reviewer — never instruct a reviewer to
346
+ ignore or not flag a specific issue. If you believe a finding would be a
347
+ false positive, let the reviewer raise it and adjudicate it in the review
348
+ loop. If the prompt you are writing contains "do not flag," "don't treat X
349
+ as a defect," "at most Minor," or "the plan chose" — stop: you are
350
+ pre-judging, usually to spare yourself a review loop.
351
+ The task reviewer may report "⚠️ Cannot verify from diff" items — requirements
352
+ that live in unchanged code or span tasks. These do not block the rest of the
353
+ review, but you must resolve each one yourself before marking the task
354
+ complete: you hold the plan and cross-task context the reviewer
355
+ lacks. If you confirm an item is a real gap, treat it as a failed spec
356
+ review — it enters the fix loop with the other findings.
357
+
358
+ Template: [task-reviewer-prompt.md](task-reviewer-prompt.md)
359
+
360
+ ### 4. The fix loop
361
+
362
+ The loop triggers when the review reports spec ❌, any Critical or Important
363
+ finding, or a ⚠️ item you confirmed as a real gap.
364
+
365
+ Before the loop starts, two routes leave it immediately:
366
+
367
+ - Record Minor findings in the progress ledger as you go
368
+ (`Task <N>: minor (deferred): <one-liner>`), and point the final
369
+ whole-branch review at that list so it can triage which must be fixed
370
+ before merge. A roll-up nobody reads is a silent discard. Minor findings
371
+ never enter the loop.
372
+ - A finding labeled plan-mandated — or any finding that conflicts with
373
+ what the plan's text requires — is yours to rule on: weigh the finding
374
+ against the plan text, decide with the spec as the binding authority, and
375
+ ledger the ruling before you act on it. Do not dismiss the finding because
376
+ the plan mandates it, and do not dispatch a fix that contradicts the plan
377
+ without a recorded ruling.
378
+ Everything else enters the loop. A fix round is one fix dispatch plus one
379
+ scoped re-review. Five rounds maximum per task:
380
+
381
+ **Rounds 1-3 — resume the original implementer.** Send it the open findings
382
+ verbatim. Its context is intact: it knows the task, the code, and its own
383
+ choices. If your harness cannot send another message to a live subagent,
384
+ dispatch a fresh implementer carrying the brief path, the report-file path,
385
+ and the findings — the report file is the persistent memory either way.
386
+
387
+ **Rounds 4-5 — dispatch a fresh implementer on a more capable model** (per
388
+ Model Selection), with the brief path, the report-file path, the open
389
+ findings, and this framing: "A prior implementer attempted this task
390
+ [N] times; you own it now. Read the report file for what was tried." A loop
391
+ that survives three resumes usually means the implementer cannot see its
392
+ own problem — fresh eyes and a capability bump in one move.
393
+
394
+ **Every round, either way:** the implementer fixes, re-runs the tests
395
+ covering the amended code, appends its fix report to the same report file,
396
+ and returns the short contract. Before re-dispatching the reviewer, confirm
397
+ the fix report contains the covering tests, the command run, and the
398
+ output; dispatch the re-review once all three are present. Name the
399
+ covering test files in the fix message — a one-line fix does not need the
400
+ whole suite.
401
+
402
+ **The re-review is scoped.** Run `node scripts/review-package.js PLAN_FILE FIX_BASE HEAD`
403
+ where FIX_BASE is the head the previous review saw, and dispatch
404
+ [re-review-prompt.md](re-review-prompt.md) with the findings list, the
405
+ brief, the report file, and the printed diff path. The re-reviewer verdicts
406
+ each finding ADDRESSED or NOT ADDRESSED and flags new breakage in the fix
407
+ diff only. New Critical/Important breakage in the fix diff joins the open
408
+ findings list. Out-of-scope observations go to the ledger as deferred
409
+ minors — they never extend the loop.
410
+
411
+ **After each round,** append to the ledger:
412
+ `Task <N>: fix round <R>/5 (<X> addressed, <Y> open — <finding one-liners>; commits <a7>..<b7>)`
413
+
414
+ Never fix findings yourself in the controller session — your context stays
415
+ clean for coordination, and controller fixes skip review.
416
+
417
+ **The breaker.** When round 5's re-review still leaves findings open, stop
418
+ dispatching. Adjudicate each open finding yourself — you hold the plan and
419
+ the cross-task context the reviewer lacks:
420
+
421
+ - **The reviewer is wrong, or the point is contestable:** park it —
422
+ `Task <N>: parked — <finding> — Ruling: <why the code stands>`. The final
423
+ review sees both sides.
424
+ - **Real, but nothing downstream builds on it:** park it the same way, with
425
+ a ruling that says it's real and deferred.
426
+ - **Real and load-bearing** — a later task builds on it, or it reveals a
427
+ plan defect: rule on the smallest change that unblocks the dependent work,
428
+ ledger it as `Task <N>: Ruling: <finding> — <what you decided and why>`,
429
+ and carry it into the next task's dispatch. Parking a structural failure
430
+ silently lets every dependent task build on it. Stop only when the defect
431
+ leaves every path forward a guess.
432
+
433
+ Adjudicate only at the cap. Adjudicating earlier to end a loop is
434
+ pre-judging with a different name. Every adjudication is a ledger entry —
435
+ a silent discard is forbidden.
436
+
437
+ ### 5. Complete the task
438
+
439
+ When the review comes back clean — or every open finding is parked with a
440
+ ruling at the cap — append the completion line to the ledger in the same
441
+ message as your other bookkeeping:
442
+
443
+ - `Task <N>: complete (commits <base7>..<head7>, review clean)`
444
+ - `Task <N>: complete (commits <base7>..<head7>, <K> parked)` after a
445
+ tripped breaker
446
+
447
+ Then mark the todo complete and move on. Never move to the next task while
448
+ the review has open Critical/Important issues that are neither fixed nor
449
+ parked-with-ruling at the cap.
450
+
451
+ ## Final Review
452
+
453
+ The final whole-branch review gets a package too: run
454
+ `node scripts/review-package.js PLAN_FILE MERGE_BASE HEAD` (MERGE_BASE = the commit the
455
+ branch started from, e.g. `git merge-base main HEAD`) and include the
456
+ printed path in the final review dispatch, so the final reviewer reads
457
+ one file instead of re-deriving the branch diff with git commands. Dispatch
458
+ on the most capable available model (see Model Selection), using
459
+ the bundled requesting-code-review support template
460
+ [code-reviewer.md](../requesting-code-review/code-reviewer.md). Point it at
461
+ the ledger's deferred-minor and parked lines so it can triage which must be
462
+ fixed before merge.
463
+
464
+ If the final whole-branch review returns findings, dispatch ONE fix subagent
465
+ with the complete findings list — not one fixer per finding.
466
+ Per-finding fixers each rebuild context and re-run suites; a real
467
+ session's final-review fix wave cost more than all its tasks combined.
468
+ Then run exactly one scoped re-review of the fix wave
469
+ (`node scripts/review-package.js PLAN_FILE FIX_BASE HEAD` over the fix range,
470
+ [re-review-prompt.md](re-review-prompt.md)).
471
+ Adjudicate any residual findings as in the task loop's breaker: park with
472
+ rulings, or rule on the load-bearing ones and ledger what you decided. Only
473
+ the four classes above stop you here. There is no second fix wave —
474
+ residual load-bearing findings surface to your human partner when
475
+ finishing-a-development-branch presents the options.
476
+
477
+ ## Finish
478
+
479
+ Before you delete anything, collect every ledger line containing `Ruling:` —
480
+ preflight rulings, parked findings, breaker adjudications, all of them — into
481
+ your final message under "Rulings I made", in the order you made them, each
482
+ with what it costs if wrong. The list is exhaustive: if the ledger holds a
483
+ ruling, the list holds it. That list is the only place the decisions you
484
+ took on your human partner's behalf reach them — they read it and rework
485
+ whatever you got wrong. A ruling that dies with the workspace was a decision
486
+ made in secret.
487
+
488
+ When the final whole-branch review is clean and its fixes are merged,
489
+ delete this plan's workspace (`rm -rf <workspace>`) — the git history is
490
+ the record now. Sibling directories belong to other plans; leave them
491
+ alone.
492
+
493
+ Use `finishing-a-development-branch`.
494
+
495
+ ## Common Rationalizations
496
+
497
+ | Excuse | Reality |
498
+ |--------|---------|
499
+ | "Close enough on spec compliance" | Reviewer found spec gaps = not done. Fix or hit the cap and adjudicate — those are the only exits. |
500
+ | "I'll fix it myself, dispatching is overhead" | Controller fixes pollute your context and skip review. Resume the implementer. |
501
+ | "One more round will converge" | Past the cap, rounds don't converge — the failure is structural. Adjudicate and route. |
502
+ | "The reviewer will just find something new anyway" | Scoped re-reviews verify fixes; they cannot wander. New findings on untouched code go to the ledger, not the loop. |
503
+ | "This finding is obviously wrong, I'll drop it" | You adjudicate only at the cap, and every ruling is a ledger entry. Silent discards are forbidden. |
504
+ | "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
505
+ | "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. |
506
+ | "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. |
507
+ | "The implementer spawned its own reviewer — free extra assurance" | It's a duplicate seat reviewing the same diff; the task review is the gate. A worker-spawned reviewer is a defect to flag, not rigor. |
508
+
509
+ ## Example Workflow
510
+
511
+ ```
512
+ You: I'm using Subagent-Driven Development to execute this plan.
513
+
514
+ [Setup: worktree verified]
515
+ [Read plan file once: docs/superpowers/plans/feature-plan.md]
516
+ [Resolve workspace: node scripts/sdd-workspace.js docs/superpowers/plans/feature-plan.md — no ledger inside, fresh start]
517
+ [Create todos for all tasks]
518
+
519
+ Task 1: Hook installation script
520
+
521
+ [Run task-brief for Task 1; dispatch implementer with brief + report paths + context]
522
+
523
+ Implementer: "Before I begin - should the hook be installed at user or system level?"
524
+
525
+ You: "User level (~/.config/superpowers/hooks/)"
526
+
527
+ Implementer: [Later]
528
+ - Implemented install-hook command
529
+ - Added tests, 5/5 passing
530
+ - Self-review: Found I missed --force flag, added it
531
+ - Committed
532
+
533
+ [Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path]
534
+ Task reviewer: Spec ✅ - all requirements met, nothing extra.
535
+ Strengths: Good test coverage, clean. Issues: None. Task quality: Approved.
536
+
537
+ [Ledger: Task 1: complete (commits a1b2c3d..d4e5f6a, review clean)]
538
+
539
+ Task 2: Recovery modes
540
+
541
+ [Run task-brief for Task 2; dispatch implementer with brief + report paths + context]
542
+
543
+ Implementer: [No questions]
544
+ - Added verify/repair modes
545
+ - 8/8 tests passing
546
+ - Committed
547
+
548
+ [Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path]
549
+ Task reviewer: Spec ❌:
550
+ - Missing: Progress reporting (spec says "report every 100 items")
551
+ Issues (Important): Magic number (100)
552
+
553
+ [Fix round 1: resume the implementer with both findings]
554
+ Implementer: Added progress reporting, extracted PROGRESS_INTERVAL constant.
555
+ Re-ran test/recovery.test.js — 10/10 passing. Fix report appended.
556
+
557
+ [Run review-package PLAN_FILE FIX_BASE HEAD; dispatch scoped re-review]
558
+ Re-reviewer: Missing progress reporting — ADDRESSED (src/recovery.js:41).
559
+ Magic number — ADDRESSED (src/recovery.js:7). New breakage: none.
560
+ Verdict: all findings addressed.
561
+
562
+ [Ledger: Task 2: fix round 1/5 (2 addressed, 0 open; commits d4e5f6a..b7c8d9e)]
563
+ [Ledger: Task 2: complete (commits d4e5f6a..b7c8d9e, review clean)]
564
+
565
+ ...
566
+
567
+ [After all tasks]
568
+ [Run review-package PLAN_FILE MERGE_BASE HEAD; dispatch final code-reviewer, most capable model]
569
+ Final reviewer: All requirements met. Deferred minors triaged: none block merge.
570
+
571
+ [Delete this plan's workspace — the record now lives in git]
572
+
573
+ Done! Using finishing-a-development-branch.
574
+ ```