superpowers-mcp 6.2.4 → 6.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (27) hide show
  1. package/README.ja.md +26 -3
  2. package/README.ko.md +26 -3
  3. package/README.md +26 -3
  4. package/README.zh-TW.md +26 -3
  5. package/out/server.js +1 -1
  6. package/package.json +1 -1
  7. package/skills/brainstorming/SKILL.md +109 -9
  8. package/skills/finishing-a-development-branch/SKILL.md +30 -0
  9. package/skills/requesting-code-review/SKILL.md +1 -1
  10. package/skills/requesting-code-review/code-reviewer.md +9 -0
  11. package/skills/subagent-driven-development/SKILL.md +101 -29
  12. package/skills/subagent-driven-development/implementer-prompt.md +12 -0
  13. package/skills/subagent-driven-development/re-review-prompt.md +9 -0
  14. package/skills/subagent-driven-development/scripts/review-package +11 -1
  15. package/skills/subagent-driven-development/scripts/review-package.ps1 +13 -1
  16. package/skills/subagent-driven-development/scripts/sdd-workspace +43 -4
  17. package/skills/subagent-driven-development/scripts/sdd-workspace.ps1 +66 -4
  18. package/skills/subagent-driven-development/scripts/task-brief +1 -1
  19. package/skills/subagent-driven-development/scripts/task-brief.ps1 +1 -1
  20. package/skills/subagent-driven-development/task-reviewer-prompt.md +25 -5
  21. package/skills/test-driven-development/SKILL.md +10 -0
  22. package/skills/using-superpowers/SKILL.md +1 -0
  23. package/skills/using-superpowers/references/codex-tools.md +70 -1
  24. package/skills/using-superpowers/references/hermes-tools.md +56 -0
  25. package/skills/writing-plans/SKILL.md +3 -0
  26. package/skills/writing-skills/anthropic-best-practices.md +1 -1
  27. package/skills/writing-skills/render-graphs.js +3 -2
@@ -14,7 +14,21 @@ Execute plan by dispatching a fresh implementer subagent per task, a task review
14
14
  **Narration:** between tool calls, narrate at most one short line — the
15
15
  ledger and the tool results carry the record.
16
16
 
17
- **Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are: BLOCKED status you cannot resolve, ambiguity that genuinely prevents progress, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.
17
+ **Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are the four named below, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.
18
+
19
+ **Rulings, not stalls.** A running plan does not wait on a human. Conflicts,
20
+ ambiguities, plan defects, a cap you would have asked to exceed — decide
21
+ them. The spec is the binding authority, the plan is its argument, and your
22
+ judgment settles what neither answers. Record every decision in the ledger as
23
+ `Ruling: <what you decided> — <why> — <what it costs if wrong>`, and keep
24
+ going. A wrong ruling costs rework your human partner can see and undo; a
25
+ session parked on a question costs their whole day and buys nothing.
26
+
27
+ Four things stop you, and only these: an irreversible or destructive
28
+ operation; a security-sensitive action; a side effect outside this worktree
29
+ that norms say you ask about first (a merge, a push to a shared branch, a
30
+ publish); and a plan so broken that every path forward is a guess. For those,
31
+ stop and ask.
18
32
 
19
33
  ## When to Use
20
34
 
@@ -57,14 +71,14 @@ digraph process {
57
71
  "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" [shape=box];
58
72
  "Spec ✅ and quality approved?" [shape=diamond];
59
73
  "Finding conflicts with plan text?" [shape=diamond];
60
- "Ask human partner which governs" [shape=box];
74
+ "Rule on the conflict, ledger the ruling" [shape=box];
61
75
  "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [shape=box];
62
76
  "Dispatch scoped re-review (./re-review-prompt.md)" [shape=box];
63
77
  "All findings addressed?" [shape=diamond];
64
78
  "R = 5?" [shape=diamond];
65
79
  "Adjudicate each open finding" [shape=box];
66
80
  "Any load-bearing finding?" [shape=diamond];
67
- "STOP: report BLOCKED to human partner" [shape=box];
81
+ "Rule and continue; stop only if every path forward is a guess" [shape=box];
68
82
  "Park findings in ledger with rulings" [shape=box];
69
83
  "Append completion to ledger, mark todo complete" [shape=box];
70
84
  }
@@ -85,8 +99,8 @@ digraph process {
85
99
  "Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" -> "Spec ✅ and quality approved?";
86
100
  "Spec ✅ and quality approved?" -> "Append completion to ledger, mark todo complete" [label="yes"];
87
101
  "Spec ✅ and quality approved?" -> "Finding conflicts with plan text?" [label="no"];
88
- "Finding conflicts with plan text?" -> "Ask human partner which governs" [label="yes"];
89
- "Ask human partner which governs" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model";
102
+ "Finding conflicts with plan text?" -> "Rule on the conflict, ledger the ruling" [label="yes"];
103
+ "Rule on the conflict, ledger the ruling" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model";
90
104
  "Finding conflicts with plan text?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no"];
91
105
  "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" -> "Dispatch scoped re-review (./re-review-prompt.md)";
92
106
  "Dispatch scoped re-review (./re-review-prompt.md)" -> "All findings addressed?";
@@ -95,7 +109,7 @@ digraph process {
95
109
  "R = 5?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no - next round"];
96
110
  "R = 5?" -> "Adjudicate each open finding" [label="yes - breaker trips"];
97
111
  "Adjudicate each open finding" -> "Any load-bearing finding?";
98
- "Any load-bearing finding?" -> "STOP: report BLOCKED to human partner" [label="yes"];
112
+ "Any load-bearing finding?" -> "Rule and continue; stop only if every path forward is a guess" [label="yes"];
99
113
  "Any load-bearing finding?" -> "Park findings in ledger with rulings" [label="no"];
100
114
  "Park findings in ledger with rulings" -> "Append completion to ledger, mark todo complete";
101
115
  "Append completion to ledger, mark todo complete" -> "More tasks remain?";
@@ -140,19 +154,32 @@ a ledger file, not only in todos.
140
154
  that happens, recover from `git log`.
141
155
 
142
156
  Read the plan once, note its context and Global Constraints, and create a
143
- todo per task.
157
+ todo per task. If the plan names a Spec, read that too: the spec is the
158
+ authority the plan argues from, and conflicts inside the plan resolve
159
+ against it. A plan with no reachable spec gets a ledger note saying so —
160
+ rulings made without one are provisional.
144
161
 
145
- Before dispatching Task 1, scan the plan once for conflicts:
162
+ Before dispatching Task 1, scan the plan once for conflicts, writing down
163
+ what you checked as you check it:
146
164
 
147
165
  - tasks that contradict each other or the plan's Global Constraints
148
166
  - anything the plan explicitly mandates that the review rubric treats as a
149
167
  defect (a test that asserts nothing, verbatim duplication of a logic block)
150
168
 
151
- Present everything you find to your human partner as one batched question —
152
- each finding beside the plan text that mandates it, asking which governs —
153
- before execution begins, not one interrupt per discovery mid-plan. If the
154
- scan is clean, proceed without comment. The review loop remains the net for
155
- conflicts that only emerge from implementation.
169
+ The scan's output is a table, not a verdict. One row for every pair of tasks
170
+ that share a file or an interface: the two tasks, what one produces against
171
+ what the other consumes, and what you found. One row for every task: whether
172
+ its own text agrees with itself — the tests it specifies against the code it
173
+ specifies, the files it creates against the files it later touches. "The scan
174
+ is clean" without those rows is not a scan you ran.
175
+
176
+ Write the table to the ledger. Rule on everything you find before execution
177
+ begins — each finding against the plan text that mandates it — and record
178
+ each ruling in the ledger. If the scan is clean, proceed without comment.
179
+ Rule on each conflict it surfaces — the spec is the binding authority, the
180
+ plan is its argument — record the ruling beside its row, and dispatch
181
+ Task 1. The review loop remains the net for conflicts that only emerge from
182
+ implementation.
156
183
 
157
184
  ## Model Selection
158
185
 
@@ -193,17 +220,37 @@ that implementer. Single-file mechanical fixes also take the cheapest tier.
193
220
 
194
221
  ## The Task Loop
195
222
 
223
+ **Batch small same-shape work.** When the plan lists several tasks that are
224
+ each a small, independent edit of the same kind — the same one-line fix,
225
+ constant change, or field addition repeated across files — do not dispatch
226
+ one subagent per task. Compose ONE dispatch brief listing every file and
227
+ its change, send the whole batch to a single subagent, and review its diff
228
+ as one unit. Reserve one-dispatch-per-task for work that needs its own
229
+ judgment, its own tests, or its own review surface.
230
+
196
231
  Everything you paste into a dispatch prompt — and everything a subagent
197
232
  prints back — stays resident in your context for the rest of the session
198
233
  and is re-read on every later turn. Hand artifacts over as files.
199
234
 
235
+ **Waiting on dispatched subagents:** never poll a wait interface with
236
+ short timeouts, and never sit in one silent, open-ended wait either.
237
+ While you have local work — ledger updates, packaging the next review,
238
+ reading reports — keep working; child results arrive on their own.
239
+ When you are genuinely idle, wait in bounded stretches (five to ten
240
+ minutes, where your platform allows), and between stretches post one
241
+ line of status and reconcile your live children: list them, and chase
242
+ any that finished without reporting. A bounded stretch keeps nearly
243
+ all of a long wait's efficiency while guaranteeing a stuck or lost
244
+ child is noticed within minutes, not at the end of the session.
245
+
200
246
  ### 1. Dispatch the implementer
201
247
 
202
248
  Record BASE (`git rev-parse HEAD`) before dispatching — the review package
203
249
  and fix-round diffs need it.
204
250
 
205
251
  - **Task brief:** before dispatching an implementer, run this skill's
206
- `scripts/task-brief PLAN_FILE N` (or `scripts/task-brief.ps1 PLAN_FILE N` on Windows PowerShell) — it extracts the task's full text to a
252
+ `scripts/task-brief PLAN_FILE N` (or `scripts/task-brief.ps1 PLAN_FILE N` on
253
+ Windows PowerShell) — it extracts the task's full text to a
207
254
  uniquely named file and prints the path. Compose the dispatch so the
208
255
  brief stays the single source of
209
256
  requirements. Your dispatch should contain: (1) one line on where this
@@ -223,6 +270,12 @@ and fix-round diffs need it.
223
270
  later dispatches — a real session's dispatch hit 42k chars of which 99%
224
271
  was pasted history. A fresh subagent needs its task, the interfaces it
225
272
  touches, and the global constraints. Nothing else.
273
+ - The dispatch carries the no-subagents contract (it is in the
274
+ implementer template): the implementer never dispatches subagents —
275
+ not helpers, and never a reviewer. Review arrives from you, after the
276
+ report. In real sessions, every reviewer a worker spawned duplicated
277
+ the task review the controller dispatched anyway — a full extra
278
+ review seat per task.
226
279
  - If an earlier task parked a finding in the area this task touches, carry
227
280
  a pointer to that ledger entry in the dispatch.
228
281
  - Record the implementer's agent identity from the dispatch result —
@@ -245,7 +298,7 @@ Implementer subagents report one of four statuses. Handle each appropriately:
245
298
  1. If it's a context problem, provide more context and re-dispatch with the same model
246
299
  2. If the task requires more reasoning, re-dispatch with a more capable model
247
300
  3. If the task is too large, break it into smaller pieces
248
- 4. If the plan itself is wrong, escalate to the human
301
+ 4. If the plan itself is wrong, rule on the correction, ledger it, and re-dispatch with the ruling carried in the dispatch
249
302
 
250
303
  **Never** ignore an escalation or force the same model to retry without changes. If the implementer said it's stuck, something needs to change.
251
304
 
@@ -262,7 +315,9 @@ required. Implementer self-review never replaces the task review; both are
262
315
  needed.
263
316
 
264
317
  - Hand the reviewer its diff as a file: run this skill's
265
- `scripts/review-package PLAN_FILE BASE HEAD` (or `scripts/review-package.ps1 PLAN_FILE BASE HEAD` on Windows PowerShell) and pass the reviewer the file path
318
+ `scripts/review-package PLAN_FILE BASE HEAD` (or
319
+ `scripts/review-package.ps1 PLAN_FILE BASE HEAD` on Windows PowerShell) and
320
+ pass the reviewer the file path
266
321
  it prints (or, without bash: `git log --oneline`, `git diff --stat`,
267
322
  and `git diff -U10` for the range, redirected to one uniquely named
268
323
  file). The output never enters your own context, and the reviewer sees
@@ -312,10 +367,11 @@ Before the loop starts, two routes leave it immediately:
312
367
  before merge. A roll-up nobody reads is a silent discard. Minor findings
313
368
  never enter the loop.
314
369
  - A finding labeled plan-mandated — or any finding that conflicts with
315
- what the plan's text requires — is the human's decision, like any plan
316
- contradiction: present the finding and the plan text, ask which governs.
317
- Do not dismiss the finding because the plan mandates it, and do not
318
- dispatch a fix that contradicts the plan without asking.
370
+ what the plan's text requires — is yours to rule on: weigh the finding
371
+ against the plan text, decide with the spec as the binding authority, and
372
+ ledger the ruling before you act on it. Do not dismiss the finding because
373
+ the plan mandates it, and do not dispatch a fix that contradicts the plan
374
+ without a recorded ruling.
319
375
  Everything else enters the loop. A fix round is one fix dispatch plus one
320
376
  scoped re-review. Five rounds maximum per task:
321
377
 
@@ -361,15 +417,16 @@ dispatching. Adjudicate each open finding yourself — you hold the plan and
361
417
  the cross-task context the reviewer lacks:
362
418
 
363
419
  - **The reviewer is wrong, or the point is contestable:** park it —
364
- `Task <N>: parked — <finding> — ruling: <why the code stands>`. The final
420
+ `Task <N>: parked — <finding> — Ruling: <why the code stands>`. The final
365
421
  review sees both sides.
366
422
  - **Real, but nothing downstream builds on it:** park it the same way, with
367
423
  a ruling that says it's real and deferred.
368
424
  - **Real and load-bearing** — a later task builds on it, or it reveals a
369
- plan defect: STOP. Append `Task <N>: BLOCKED — <reason>` and report to
370
- your human partner with the finding, the plan text it collides with, and
371
- the fix history. Parking a structural failure lets every dependent task
372
- build on it and hands the final review a problem it cannot fix either.
425
+ plan defect: rule on the smallest change that unblocks the dependent work,
426
+ ledger it as `Task <N>: Ruling: <finding> — <what you decided and why>`,
427
+ and carry it into the next task's dispatch. Parking a structural failure
428
+ silently lets every dependent task build on it. Stop only when the defect
429
+ leaves every path forward a guess.
373
430
 
374
431
  Adjudicate only at the cap. Adjudicating earlier to end a loop is
375
432
  pre-judging with a different name. Every adjudication is a ledger entry —
@@ -392,8 +449,10 @@ parked-with-ruling at the cap.
392
449
  ## Final Review
393
450
 
394
451
  The final whole-branch review gets a package too: run
395
- `scripts/review-package PLAN_FILE MERGE_BASE HEAD` (or `scripts/review-package.ps1 PLAN_FILE MERGE_BASE HEAD` on Windows PowerShell; MERGE_BASE = the commit the
396
- branch started from, e.g. `git merge-base main HEAD`) and include the
452
+ `scripts/review-package PLAN_FILE MERGE_BASE HEAD` (or
453
+ `scripts/review-package.ps1 PLAN_FILE MERGE_BASE HEAD` on Windows PowerShell;
454
+ MERGE_BASE = the commit the branch started from, e.g. `git merge-base main HEAD`)
455
+ and include the
397
456
  printed path in the final review dispatch, so the final reviewer reads
398
457
  one file instead of re-deriving the branch diff with git commands. Dispatch
399
458
  on the most capable available model (see Model Selection), using
@@ -407,15 +466,27 @@ with the complete findings list — not one fixer per finding.
407
466
  Per-finding fixers each rebuild context and re-run suites; a real
408
467
  session's final-review fix wave cost more than all its tasks combined.
409
468
  Then run exactly one scoped re-review of the fix wave
410
- (`scripts/review-package PLAN_FILE FIX_BASE HEAD` over the fix range,
469
+ (`scripts/review-package PLAN_FILE FIX_BASE HEAD`, or
470
+ `scripts/review-package.ps1 PLAN_FILE FIX_BASE HEAD` on Windows PowerShell,
471
+ over the fix range,
411
472
  [re-review-prompt.md](re-review-prompt.md)).
412
473
  Adjudicate any residual findings as in the task loop's breaker: park with
413
- rulings, or stop on load-bearing ones. There is no second fix wave —
474
+ rulings, or rule on the load-bearing ones and ledger what you decided. Only
475
+ the four classes above stop you here. There is no second fix wave —
414
476
  residual load-bearing findings surface to your human partner when
415
477
  finishing-a-development-branch presents the options.
416
478
 
417
479
  ## Finish
418
480
 
481
+ Before you delete anything, collect every ledger line containing `Ruling:` —
482
+ preflight rulings, parked findings, breaker adjudications, all of them — into
483
+ your final message under "Rulings I made", in the order you made them, each
484
+ with what it costs if wrong. The list is exhaustive: if the ledger holds a
485
+ ruling, the list holds it. That list is the only place the decisions you
486
+ took on your human partner's behalf reach them — they read it and rework
487
+ whatever you got wrong. A ruling that dies with the workspace was a decision
488
+ made in secret.
489
+
419
490
  When the final whole-branch review is clean and its fixes are merged,
420
491
  delete this plan's workspace (`rm -rf <workspace>`) — the git history is
421
492
  the record now. Sibling directories belong to other plans; leave them
@@ -435,6 +506,7 @@ Use superpowers:finishing-a-development-branch.
435
506
  | "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
436
507
  | "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. |
437
508
  | "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. |
509
+ | "The implementer spawned its own reviewer — free extra assurance" | It's a duplicate seat reviewing the same diff; the task review is the gate. A worker-spawned reviewer is a defect to flag, not rigor. |
438
510
 
439
511
  ## Example Workflow
440
512
 
@@ -47,6 +47,18 @@ Subagent (general-purpose):
47
47
  While iterating, run the focused test for what you're changing; run the
48
48
  full suite once before committing, not after every edit.
49
49
 
50
+ ## You Do Not Dispatch Subagents
51
+
52
+ Do all of this task's work yourself. Never spawn a subagent to
53
+ implement part of the task, and above all never spawn a reviewer to
54
+ check your work. Self-review (below) means reading your own diff.
55
+ Review is the controller's job: after you report, it dispatches a
56
+ fresh reviewer against your diff. A reviewer you spawn duplicates
57
+ that review at full cost, and its approval counts for nothing in
58
+ the process. If you catch yourself thinking "an independent review
59
+ would strengthen my report" — that review is already scheduled.
60
+ Report instead.
61
+
50
62
  ## Code Organization
51
63
 
52
64
  You reason best about code you can hold in context at once, and your edits are more
@@ -43,6 +43,15 @@ Subagent (general-purpose):
43
43
  Your review is read-only on this checkout. Do not mutate the working
44
44
  tree, the index, HEAD, or branch state in any way.
45
45
 
46
+ ## You Do Not Dispatch Subagents
47
+
48
+ Do all of this review yourself. Never spawn a subagent to review part
49
+ of the diff, and never spawn another reviewer for a second opinion.
50
+ This process already provides every review seat the work gets; a
51
+ reviewer you spawn duplicates one of them at full cost, and its
52
+ verdict counts for nothing. If the diff feels too large for one
53
+ pass, review it in passes yourself and say so in your report.
54
+
46
55
  ## Scope
47
56
 
48
57
  Your scope is the findings list and the fix diff. Verdict every finding.
@@ -22,10 +22,20 @@ head=$3
22
22
  git rev-parse --verify --quiet "$base" >/dev/null || { echo "bad BASE: $base" >&2; exit 2; }
23
23
  git rev-parse --verify --quiet "$head" >/dev/null || { echo "bad HEAD: $head" >&2; exit 2; }
24
24
 
25
+ git merge-base --is-ancestor "$base" "$head" || {
26
+ echo "HEAD ($head) is not a descendant of BASE ($base)" >&2
27
+ exit 3
28
+ }
29
+
30
+ if [ "$(git rev-list --count "${base}..${head}")" -eq 0 ]; then
31
+ echo "empty commit range: ${base}..${head}" >&2
32
+ exit 3
33
+ fi
34
+
25
35
  if [ $# -eq 4 ]; then
26
36
  out=$4
27
37
  else
28
- dir=$("$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan")
38
+ dir=$("${BASH:-bash}" "$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan")
29
39
  out="$dir/review-$(git rev-parse --short "$base")..$(git rev-parse --short "$head").diff"
30
40
  fi
31
41
 
@@ -31,6 +31,18 @@ if ($LASTEXITCODE -ne 0) {
31
31
  exit 2
32
32
  }
33
33
 
34
+ & git merge-base --is-ancestor $base $head *> $null
35
+ if ($LASTEXITCODE -ne 0) {
36
+ [Console]::Error.WriteLine("HEAD ($head) is not a descendant of BASE ($base)")
37
+ exit 3
38
+ }
39
+
40
+ $commitCount = (& git rev-list --count "${base}..${head}").Trim()
41
+ if ([int]$commitCount -eq 0) {
42
+ [Console]::Error.WriteLine("empty commit range: ${base}..${head}")
43
+ exit 3
44
+ }
45
+
34
46
  if ($args.Count -eq 4) {
35
47
  $out = $args[3]
36
48
  } else {
@@ -53,7 +65,7 @@ $content.Add("")
53
65
  $content.Add("## Diff")
54
66
  (& git diff -U10 "${base}..${head}") | ForEach-Object { $content.Add($_) }
55
67
 
56
- Set-Content -Path $out -Value $content -Encoding utf8
68
+ Set-Content -LiteralPath $out -Value $content -Encoding utf8
57
69
  $commits = (& git rev-list --count "${base}..${head}").Trim()
58
70
  $bytes = (Get-Item -LiteralPath $out).Length
59
71
  Write-Output "wrote ${out}: $commits commit(s), $bytes bytes"
@@ -3,11 +3,17 @@
3
3
  # short-lived artifacts: task briefs, implementer reports, review packages,
4
4
  # and the progress ledger. Print the plan directory's absolute path.
5
5
  #
6
- # One directory per plan (.superpowers/sdd/<plan-basename>/) so a follow-up
6
+ # One directory per plan (.superpowers/sdd/<plan-slug>/) so a follow-up
7
7
  # plan in the same working tree can never read or overwrite another plan's
8
8
  # artifacts. A stale ledger misread as current progress makes controllers
9
9
  # skip whole task sequences — plan-scoping removes that failure structurally.
10
10
  #
11
+ # Ownership markers (plan-path file in the workspace directory) disambiguate
12
+ # same-basename plans (e.g. docs/alpha/plan.md vs docs/beta/plan.md) by appending
13
+ # the parent-directory name, then a counter. A workspace with no marker
14
+ # predates the marker scheme and is adopted for the current plan so in-flight
15
+ # workspaces keep resolving.
16
+ #
11
17
  # The workspace lives in the working tree (not under .git/) because Claude Code
12
18
  # treats .git/ as a protected path and denies agent writes there — which blocks
13
19
  # an implementer subagent from writing its report file. A self-ignoring
@@ -28,13 +34,46 @@ fi
28
34
  plan=$1
29
35
  [ -f "$plan" ] || { echo "no such plan file: $plan" >&2; exit 2; }
30
36
 
31
- slug=$(basename "$plan" .md)
37
+ slug=$(basename -- "$plan" .md)
32
38
  [ -n "$slug" ] && [ "$slug" != "." ] && [ "$slug" != ".." ] \
33
39
  || { echo "cannot derive a workspace name from: $plan" >&2; exit 2; }
34
40
 
35
41
  root=$(git rev-parse --show-toplevel)
42
+ root=$(CDPATH= cd -- "$root" && pwd -P)
36
43
  base="$root/.superpowers/sdd"
44
+
45
+ # Normalize the plan path (physical directory, so relative/absolute/../
46
+ # spellings of one plan compare equal) and express it as the marker value:
47
+ # repo-relative when the plan lives under the repo root, absolute otherwise.
48
+ plan_dir=$(CDPATH= cd -- "$(dirname -- "$plan")" && pwd -P)
49
+ plan_abs="$plan_dir/$(basename -- "$plan")"
50
+ case "$plan_abs" in
51
+ "$root"/*) plan_id=${plan_abs#"$root"/} ;;
52
+ *) plan_id=$plan_abs ;;
53
+ esac
54
+
55
+ # True when the workspace at $1 is (or becomes) this plan's: an existing
56
+ # marker must name this plan; a missing marker means a new workspace or a
57
+ # pre-marker legacy one, and either way the plan claims it by writing one.
58
+ owns() {
59
+ if [ -e "$1/plan-path" ]; then
60
+ [ "$(cat "$1/plan-path")" = "$plan_id" ]
61
+ else
62
+ mkdir -p "$1"
63
+ printf '%s\n' "$plan_id" > "$1/plan-path"
64
+ fi
65
+ }
66
+
37
67
  dir="$base/$slug"
38
- mkdir -p "$dir"
68
+ if ! owns "$dir"; then
69
+ parent=$(basename -- "$plan_dir")
70
+ dir="$base/$slug-$parent"
71
+ if ! owns "$dir"; then
72
+ n=2
73
+ while ! owns "$base/$slug-$parent-$n"; do n=$((n + 1)); done
74
+ dir="$base/$slug-$parent-$n"
75
+ fi
76
+ fi
77
+
39
78
  printf '*\n' > "$base/.gitignore"
40
- cd "$dir" && pwd
79
+ CDPATH= cd -- "$dir" && pwd
@@ -22,7 +22,7 @@ if (-not (Test-Path -LiteralPath $plan -PathType Leaf)) {
22
22
  exit 2
23
23
  }
24
24
 
25
- $slug = [System.IO.Path]::GetFileName($plan) -replace '\.md$', ''
25
+ $slug = [System.IO.Path]::GetFileName($plan) -creplace '\.md$', ''
26
26
  if ([string]::IsNullOrEmpty($slug) -or $slug -eq "." -or $slug -eq "..") {
27
27
  [Console]::Error.WriteLine("cannot derive a workspace name from: $plan")
28
28
  exit 2
@@ -30,7 +30,69 @@ if ([string]::IsNullOrEmpty($slug) -or $slug -eq "." -or $slug -eq "..") {
30
30
 
31
31
  $root = (& git rev-parse --show-toplevel).Trim()
32
32
  $base = Join-Path $root ".superpowers/sdd"
33
+
34
+ function Get-PhysicalDirectoryPath($path) {
35
+ if (-not (Test-Path -LiteralPath $path)) { return $path }
36
+ if ($IsWindows) {
37
+ return (Resolve-Path -LiteralPath $path).Path
38
+ }
39
+ $orig = Get-Location
40
+ try {
41
+ Set-Location -LiteralPath $path
42
+ $pwdCmd = (Get-Command -Type Application pwd -ErrorAction SilentlyContinue).Source
43
+ if ($pwdCmd) {
44
+ return (& $pwdCmd -P).Trim()
45
+ } else {
46
+ return (Resolve-Path -LiteralPath $path).Path
47
+ }
48
+ } finally {
49
+ Set-Location $orig
50
+ }
51
+ }
52
+
53
+ # Normalize the plan path (physical directory, so relative/absolute/../
54
+ # spellings of one plan compare equal) and express it as the marker value:
55
+ # repo-relative when the plan lives under the repo root, absolute otherwise.
56
+ $planLeaf = Split-Path -Leaf $plan
57
+ $planParent = Split-Path -Parent $plan
58
+ if ([string]::IsNullOrEmpty($planParent)) { $planParent = "." }
59
+
60
+ $planDir = Get-PhysicalDirectoryPath $planParent
61
+ $rootPhys = Get-PhysicalDirectoryPath $root
62
+
63
+ $planAbs = (Join-Path $planDir $planLeaf) -replace '\\', '/'
64
+ $rootNorm = $rootPhys -replace '\\', '/'
65
+
66
+ if ($planAbs.StartsWith($rootNorm + "/", [System.StringComparison]::OrdinalIgnoreCase)) {
67
+ $planId = $planAbs.Substring($rootNorm.Length + 1)
68
+ } else {
69
+ $planId = $planAbs
70
+ }
71
+
72
+ function Test-And-Claim-Workspace($targetDir, $id) {
73
+ $markerPath = Join-Path $targetDir "plan-path"
74
+ if (Test-Path -LiteralPath $markerPath -PathType Leaf) {
75
+ $existingId = (Get-Content -LiteralPath $markerPath -Raw).Trim()
76
+ return ($existingId -eq $id)
77
+ } else {
78
+ New-Item -ItemType Directory -Force -Path $targetDir | Out-Null
79
+ Set-Content -LiteralPath $markerPath -Value $id -Encoding ascii
80
+ return $true
81
+ }
82
+ }
83
+
33
84
  $dir = Join-Path $base $slug
34
- New-Item -ItemType Directory -Force -Path $dir | Out-Null
35
- Set-Content -Path (Join-Path $base ".gitignore") -Value "*" -NoNewline -Encoding ascii
36
- (Resolve-Path $dir).Path
85
+ if (-not (Test-And-Claim-Workspace $dir $planId)) {
86
+ $parent = Split-Path -Leaf $planDir
87
+ $dir = Join-Path $base "$slug-$parent"
88
+ if (-not (Test-And-Claim-Workspace $dir $planId)) {
89
+ $n = 2
90
+ while (-not (Test-And-Claim-Workspace (Join-Path $base "$slug-$parent-$n") $planId)) {
91
+ $n++
92
+ }
93
+ $dir = Join-Path $base "$slug-$parent-$n"
94
+ }
95
+ }
96
+
97
+ Set-Content -LiteralPath (Join-Path $base ".gitignore") -Value "*" -NoNewline -Encoding ascii
98
+ (Resolve-Path -LiteralPath $dir).Path
@@ -21,7 +21,7 @@ n=$2
21
21
  if [ $# -eq 3 ]; then
22
22
  out=$3
23
23
  else
24
- dir=$("$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan")
24
+ dir=$("${BASH:-bash}" "$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan")
25
25
  out="$dir/task-${n}-brief.md"
26
26
  fi
27
27
 
@@ -42,7 +42,7 @@ foreach ($line in [System.IO.File]::ReadLines((Resolve-Path -LiteralPath $plan).
42
42
  }
43
43
  }
44
44
 
45
- Set-Content -Path $out -Value $selected -Encoding utf8
45
+ Set-Content -LiteralPath $out -Value $selected -Encoding utf8
46
46
  if ((-not (Test-Path -LiteralPath $out)) -or ((Get-Item -LiteralPath $out).Length -eq 0)) {
47
47
  [Console]::Error.WriteLine("task $taskNumber not found in $plan (no heading matching 'Task $taskNumber')")
48
48
  exit 3
@@ -52,6 +52,15 @@ Subagent (general-purpose):
52
52
  Your review is read-only on this checkout. Do not mutate the working
53
53
  tree, the index, HEAD, or branch state in any way.
54
54
 
55
+ ## You Do Not Dispatch Subagents
56
+
57
+ Do all of this review yourself. Never spawn a subagent to review part
58
+ of the diff, and never spawn another reviewer for a second opinion.
59
+ This process already provides every review seat the work gets; a
60
+ reviewer you spawn duplicates one of them at full cost, and its
61
+ verdict counts for nothing. If the diff feels too large for one
62
+ pass, review it in passes yourself and say so in your report.
63
+
55
64
  ## Do Not Trust the Report
56
65
 
57
66
  Treat the implementer's report as unverified claims about the code. It
@@ -75,6 +84,13 @@ Subagent (general-purpose):
75
84
  Warnings or other noise in the implementer's reported test output are
76
85
  findings — test output should be pristine.
77
86
 
87
+ Evidence you cannot see is not evidence that doesn't exist. If the
88
+ report or its test evidence looks truncated, or you cannot locate the
89
+ results it claims, re-read the file at its stated path — and if it is
90
+ genuinely missing or garbled, report that as a gap for the controller.
91
+ Re-running the suite to regenerate what you failed to read is not
92
+ verification; illegibility of the evidence is not invalidation of it.
93
+
78
94
  ## Part 1: Spec Compliance
79
95
 
80
96
  Compare the diff against What Was Requested:
@@ -86,6 +102,12 @@ Subagent (general-purpose):
86
102
  - **Misunderstood:** right feature built the wrong way, wrong problem
87
103
  solved
88
104
 
105
+ If the brief lists several files each with its own change (a batched
106
+ dispatch), check the diff against that list file by file: every listed
107
+ file must have its corresponding hunk. A listed file the diff never
108
+ touches is a Missing finding, no matter how clean the rest of the
109
+ batch looks.
110
+
89
111
  If a requirement cannot be verified from this diff alone (it lives in
90
112
  unchanged code or spans tasks), report it as a ⚠️ item instead of
91
113
  broadening your search.
@@ -167,8 +189,7 @@ Subagent (general-purpose):
167
189
 
168
190
  **Placeholders:**
169
191
  - `[MODEL]` — REQUIRED: reviewer model per SKILL.md Model Selection
170
- - `[BRIEF_FILE]` — REQUIRED: the task brief file (`scripts/task-brief PLAN N`,
171
- or `scripts/task-brief.ps1 PLAN N` on Windows PowerShell,
192
+ - `[BRIEF_FILE]` — REQUIRED: the task brief file (`scripts/task-brief PLAN N`, or `scripts/task-brief.ps1 PLAN N` on Windows PowerShell,
172
193
  prints the path; same file the implementer worked from)
173
194
  - `[GLOBAL_CONSTRAINTS]` — the binding requirements copied verbatim from
174
195
  the plan's Global Constraints section or the spec: exact values, formats,
@@ -179,9 +200,8 @@ Subagent (general-purpose):
179
200
  - `[BASE_SHA]` — commit before this task
180
201
  - `[HEAD_SHA]` — current commit
181
202
  - `[DIFF_FILE]` — REQUIRED: the path the controller wrote the review
182
- package to (`scripts/review-package PLAN_FILE BASE HEAD`, or
183
- `scripts/review-package.ps1 PLAN_FILE BASE HEAD` on Windows PowerShell,
184
- prints the unique path it wrote; the package never enters the controller's context)
203
+ package to (`scripts/review-package PLAN_FILE BASE HEAD`, or `scripts/review-package.ps1 PLAN_FILE BASE HEAD` on Windows PowerShell, prints the unique
204
+ path it wrote; the package never enters the controller's context)
185
205
 
186
206
  **Reviewer returns:** Spec Compliance verdict (✅/❌/⚠️), Strengths, Issues
187
207
  (Critical/Important/Minor), Task quality verdict
@@ -182,6 +182,16 @@ Confirm:
182
182
 
183
183
  **Other tests fail?** Fix now.
184
184
 
185
+ **"Other tests" means the project's suite, not just your file.** A
186
+ green run of the test you wrote is not a green suite. Before you call
187
+ the change done, run the project's test command (bare `pytest`,
188
+ `npm test`, `cargo test` — whatever the repo uses) even when your task
189
+ named only one test file. A scope statement in your task bounds the
190
+ deliverable, not your verification. Any failure that run shows —
191
+ including one you didn't cause — goes in your report by name; a red
192
+ test you watched scroll past and didn't mention is a report falsified
193
+ by omission.
194
+
185
195
  ### REFACTOR - Clean Up
186
196
 
187
197
  After green only:
@@ -56,6 +56,7 @@ If your harness appears here, read its reference file for special instructions:
56
56
  - Codex: `references/codex-tools.md`
57
57
  - Pi: `references/pi-tools.md`
58
58
  - Antigravity: `references/antigravity-tools.md`
59
+ - Hermes Agent: `references/hermes-tools.md`
59
60
 
60
61
  ## User Instructions
61
62