superpowers-mcp 6.2.4 → 6.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.ja.md +26 -3
- package/README.ko.md +26 -3
- package/README.md +26 -3
- package/README.zh-TW.md +26 -3
- package/out/server.js +1 -1
- package/package.json +1 -1
- package/skills/brainstorming/SKILL.md +109 -9
- package/skills/finishing-a-development-branch/SKILL.md +30 -0
- package/skills/requesting-code-review/SKILL.md +1 -1
- package/skills/requesting-code-review/code-reviewer.md +9 -0
- package/skills/subagent-driven-development/SKILL.md +101 -29
- package/skills/subagent-driven-development/implementer-prompt.md +12 -0
- package/skills/subagent-driven-development/re-review-prompt.md +9 -0
- package/skills/subagent-driven-development/scripts/review-package +11 -1
- package/skills/subagent-driven-development/scripts/review-package.ps1 +13 -1
- package/skills/subagent-driven-development/scripts/sdd-workspace +43 -4
- package/skills/subagent-driven-development/scripts/sdd-workspace.ps1 +66 -4
- package/skills/subagent-driven-development/scripts/task-brief +1 -1
- package/skills/subagent-driven-development/scripts/task-brief.ps1 +1 -1
- package/skills/subagent-driven-development/task-reviewer-prompt.md +25 -5
- package/skills/test-driven-development/SKILL.md +10 -0
- package/skills/using-superpowers/SKILL.md +1 -0
- package/skills/using-superpowers/references/codex-tools.md +70 -1
- package/skills/using-superpowers/references/hermes-tools.md +56 -0
- package/skills/writing-plans/SKILL.md +3 -0
- package/skills/writing-skills/anthropic-best-practices.md +1 -1
- package/skills/writing-skills/render-graphs.js +3 -2
|
@@ -14,7 +14,21 @@ Execute plan by dispatching a fresh implementer subagent per task, a task review
|
|
|
14
14
|
**Narration:** between tool calls, narrate at most one short line — the
|
|
15
15
|
ledger and the tool results carry the record.
|
|
16
16
|
|
|
17
|
-
**Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are
|
|
17
|
+
**Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are the four named below, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.
|
|
18
|
+
|
|
19
|
+
**Rulings, not stalls.** A running plan does not wait on a human. Conflicts,
|
|
20
|
+
ambiguities, plan defects, a cap you would have asked to exceed — decide
|
|
21
|
+
them. The spec is the binding authority, the plan is its argument, and your
|
|
22
|
+
judgment settles what neither answers. Record every decision in the ledger as
|
|
23
|
+
`Ruling: <what you decided> — <why> — <what it costs if wrong>`, and keep
|
|
24
|
+
going. A wrong ruling costs rework your human partner can see and undo; a
|
|
25
|
+
session parked on a question costs their whole day and buys nothing.
|
|
26
|
+
|
|
27
|
+
Four things stop you, and only these: an irreversible or destructive
|
|
28
|
+
operation; a security-sensitive action; a side effect outside this worktree
|
|
29
|
+
that norms say you ask about first (a merge, a push to a shared branch, a
|
|
30
|
+
publish); and a plan so broken that every path forward is a guess. For those,
|
|
31
|
+
stop and ask.
|
|
18
32
|
|
|
19
33
|
## When to Use
|
|
20
34
|
|
|
@@ -57,14 +71,14 @@ digraph process {
|
|
|
57
71
|
"Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" [shape=box];
|
|
58
72
|
"Spec ✅ and quality approved?" [shape=diamond];
|
|
59
73
|
"Finding conflicts with plan text?" [shape=diamond];
|
|
60
|
-
"
|
|
74
|
+
"Rule on the conflict, ledger the ruling" [shape=box];
|
|
61
75
|
"Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [shape=box];
|
|
62
76
|
"Dispatch scoped re-review (./re-review-prompt.md)" [shape=box];
|
|
63
77
|
"All findings addressed?" [shape=diamond];
|
|
64
78
|
"R = 5?" [shape=diamond];
|
|
65
79
|
"Adjudicate each open finding" [shape=box];
|
|
66
80
|
"Any load-bearing finding?" [shape=diamond];
|
|
67
|
-
"
|
|
81
|
+
"Rule and continue; stop only if every path forward is a guess" [shape=box];
|
|
68
82
|
"Park findings in ledger with rulings" [shape=box];
|
|
69
83
|
"Append completion to ledger, mark todo complete" [shape=box];
|
|
70
84
|
}
|
|
@@ -85,8 +99,8 @@ digraph process {
|
|
|
85
99
|
"Generate review package, dispatch task reviewer (./task-reviewer-prompt.md)" -> "Spec ✅ and quality approved?";
|
|
86
100
|
"Spec ✅ and quality approved?" -> "Append completion to ledger, mark todo complete" [label="yes"];
|
|
87
101
|
"Spec ✅ and quality approved?" -> "Finding conflicts with plan text?" [label="no"];
|
|
88
|
-
"Finding conflicts with plan text?" -> "
|
|
89
|
-
"
|
|
102
|
+
"Finding conflicts with plan text?" -> "Rule on the conflict, ledger the ruling" [label="yes"];
|
|
103
|
+
"Rule on the conflict, ledger the ruling" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model";
|
|
90
104
|
"Finding conflicts with plan text?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no"];
|
|
91
105
|
"Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" -> "Dispatch scoped re-review (./re-review-prompt.md)";
|
|
92
106
|
"Dispatch scoped re-review (./re-review-prompt.md)" -> "All findings addressed?";
|
|
@@ -95,7 +109,7 @@ digraph process {
|
|
|
95
109
|
"R = 5?" -> "Fix round R of 5: R≤3 resume implementer; R≥4 fresh implementer, more capable model" [label="no - next round"];
|
|
96
110
|
"R = 5?" -> "Adjudicate each open finding" [label="yes - breaker trips"];
|
|
97
111
|
"Adjudicate each open finding" -> "Any load-bearing finding?";
|
|
98
|
-
"Any load-bearing finding?" -> "
|
|
112
|
+
"Any load-bearing finding?" -> "Rule and continue; stop only if every path forward is a guess" [label="yes"];
|
|
99
113
|
"Any load-bearing finding?" -> "Park findings in ledger with rulings" [label="no"];
|
|
100
114
|
"Park findings in ledger with rulings" -> "Append completion to ledger, mark todo complete";
|
|
101
115
|
"Append completion to ledger, mark todo complete" -> "More tasks remain?";
|
|
@@ -140,19 +154,32 @@ a ledger file, not only in todos.
|
|
|
140
154
|
that happens, recover from `git log`.
|
|
141
155
|
|
|
142
156
|
Read the plan once, note its context and Global Constraints, and create a
|
|
143
|
-
todo per task.
|
|
157
|
+
todo per task. If the plan names a Spec, read that too: the spec is the
|
|
158
|
+
authority the plan argues from, and conflicts inside the plan resolve
|
|
159
|
+
against it. A plan with no reachable spec gets a ledger note saying so —
|
|
160
|
+
rulings made without one are provisional.
|
|
144
161
|
|
|
145
|
-
Before dispatching Task 1, scan the plan once for conflicts
|
|
162
|
+
Before dispatching Task 1, scan the plan once for conflicts, writing down
|
|
163
|
+
what you checked as you check it:
|
|
146
164
|
|
|
147
165
|
- tasks that contradict each other or the plan's Global Constraints
|
|
148
166
|
- anything the plan explicitly mandates that the review rubric treats as a
|
|
149
167
|
defect (a test that asserts nothing, verbatim duplication of a logic block)
|
|
150
168
|
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
169
|
+
The scan's output is a table, not a verdict. One row for every pair of tasks
|
|
170
|
+
that share a file or an interface: the two tasks, what one produces against
|
|
171
|
+
what the other consumes, and what you found. One row for every task: whether
|
|
172
|
+
its own text agrees with itself — the tests it specifies against the code it
|
|
173
|
+
specifies, the files it creates against the files it later touches. "The scan
|
|
174
|
+
is clean" without those rows is not a scan you ran.
|
|
175
|
+
|
|
176
|
+
Write the table to the ledger. Rule on everything you find before execution
|
|
177
|
+
begins — each finding against the plan text that mandates it — and record
|
|
178
|
+
each ruling in the ledger. If the scan is clean, proceed without comment.
|
|
179
|
+
Rule on each conflict it surfaces — the spec is the binding authority, the
|
|
180
|
+
plan is its argument — record the ruling beside its row, and dispatch
|
|
181
|
+
Task 1. The review loop remains the net for conflicts that only emerge from
|
|
182
|
+
implementation.
|
|
156
183
|
|
|
157
184
|
## Model Selection
|
|
158
185
|
|
|
@@ -193,17 +220,37 @@ that implementer. Single-file mechanical fixes also take the cheapest tier.
|
|
|
193
220
|
|
|
194
221
|
## The Task Loop
|
|
195
222
|
|
|
223
|
+
**Batch small same-shape work.** When the plan lists several tasks that are
|
|
224
|
+
each a small, independent edit of the same kind — the same one-line fix,
|
|
225
|
+
constant change, or field addition repeated across files — do not dispatch
|
|
226
|
+
one subagent per task. Compose ONE dispatch brief listing every file and
|
|
227
|
+
its change, send the whole batch to a single subagent, and review its diff
|
|
228
|
+
as one unit. Reserve one-dispatch-per-task for work that needs its own
|
|
229
|
+
judgment, its own tests, or its own review surface.
|
|
230
|
+
|
|
196
231
|
Everything you paste into a dispatch prompt — and everything a subagent
|
|
197
232
|
prints back — stays resident in your context for the rest of the session
|
|
198
233
|
and is re-read on every later turn. Hand artifacts over as files.
|
|
199
234
|
|
|
235
|
+
**Waiting on dispatched subagents:** never poll a wait interface with
|
|
236
|
+
short timeouts, and never sit in one silent, open-ended wait either.
|
|
237
|
+
While you have local work — ledger updates, packaging the next review,
|
|
238
|
+
reading reports — keep working; child results arrive on their own.
|
|
239
|
+
When you are genuinely idle, wait in bounded stretches (five to ten
|
|
240
|
+
minutes, where your platform allows), and between stretches post one
|
|
241
|
+
line of status and reconcile your live children: list them, and chase
|
|
242
|
+
any that finished without reporting. A bounded stretch keeps nearly
|
|
243
|
+
all of a long wait's efficiency while guaranteeing a stuck or lost
|
|
244
|
+
child is noticed within minutes, not at the end of the session.
|
|
245
|
+
|
|
200
246
|
### 1. Dispatch the implementer
|
|
201
247
|
|
|
202
248
|
Record BASE (`git rev-parse HEAD`) before dispatching — the review package
|
|
203
249
|
and fix-round diffs need it.
|
|
204
250
|
|
|
205
251
|
- **Task brief:** before dispatching an implementer, run this skill's
|
|
206
|
-
`scripts/task-brief PLAN_FILE N` (or `scripts/task-brief.ps1 PLAN_FILE N` on
|
|
252
|
+
`scripts/task-brief PLAN_FILE N` (or `scripts/task-brief.ps1 PLAN_FILE N` on
|
|
253
|
+
Windows PowerShell) — it extracts the task's full text to a
|
|
207
254
|
uniquely named file and prints the path. Compose the dispatch so the
|
|
208
255
|
brief stays the single source of
|
|
209
256
|
requirements. Your dispatch should contain: (1) one line on where this
|
|
@@ -223,6 +270,12 @@ and fix-round diffs need it.
|
|
|
223
270
|
later dispatches — a real session's dispatch hit 42k chars of which 99%
|
|
224
271
|
was pasted history. A fresh subagent needs its task, the interfaces it
|
|
225
272
|
touches, and the global constraints. Nothing else.
|
|
273
|
+
- The dispatch carries the no-subagents contract (it is in the
|
|
274
|
+
implementer template): the implementer never dispatches subagents —
|
|
275
|
+
not helpers, and never a reviewer. Review arrives from you, after the
|
|
276
|
+
report. In real sessions, every reviewer a worker spawned duplicated
|
|
277
|
+
the task review the controller dispatched anyway — a full extra
|
|
278
|
+
review seat per task.
|
|
226
279
|
- If an earlier task parked a finding in the area this task touches, carry
|
|
227
280
|
a pointer to that ledger entry in the dispatch.
|
|
228
281
|
- Record the implementer's agent identity from the dispatch result —
|
|
@@ -245,7 +298,7 @@ Implementer subagents report one of four statuses. Handle each appropriately:
|
|
|
245
298
|
1. If it's a context problem, provide more context and re-dispatch with the same model
|
|
246
299
|
2. If the task requires more reasoning, re-dispatch with a more capable model
|
|
247
300
|
3. If the task is too large, break it into smaller pieces
|
|
248
|
-
4. If the plan itself is wrong,
|
|
301
|
+
4. If the plan itself is wrong, rule on the correction, ledger it, and re-dispatch with the ruling carried in the dispatch
|
|
249
302
|
|
|
250
303
|
**Never** ignore an escalation or force the same model to retry without changes. If the implementer said it's stuck, something needs to change.
|
|
251
304
|
|
|
@@ -262,7 +315,9 @@ required. Implementer self-review never replaces the task review; both are
|
|
|
262
315
|
needed.
|
|
263
316
|
|
|
264
317
|
- Hand the reviewer its diff as a file: run this skill's
|
|
265
|
-
`scripts/review-package PLAN_FILE BASE HEAD` (or
|
|
318
|
+
`scripts/review-package PLAN_FILE BASE HEAD` (or
|
|
319
|
+
`scripts/review-package.ps1 PLAN_FILE BASE HEAD` on Windows PowerShell) and
|
|
320
|
+
pass the reviewer the file path
|
|
266
321
|
it prints (or, without bash: `git log --oneline`, `git diff --stat`,
|
|
267
322
|
and `git diff -U10` for the range, redirected to one uniquely named
|
|
268
323
|
file). The output never enters your own context, and the reviewer sees
|
|
@@ -312,10 +367,11 @@ Before the loop starts, two routes leave it immediately:
|
|
|
312
367
|
before merge. A roll-up nobody reads is a silent discard. Minor findings
|
|
313
368
|
never enter the loop.
|
|
314
369
|
- A finding labeled plan-mandated — or any finding that conflicts with
|
|
315
|
-
what the plan's text requires — is
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
dispatch a fix that contradicts the plan
|
|
370
|
+
what the plan's text requires — is yours to rule on: weigh the finding
|
|
371
|
+
against the plan text, decide with the spec as the binding authority, and
|
|
372
|
+
ledger the ruling before you act on it. Do not dismiss the finding because
|
|
373
|
+
the plan mandates it, and do not dispatch a fix that contradicts the plan
|
|
374
|
+
without a recorded ruling.
|
|
319
375
|
Everything else enters the loop. A fix round is one fix dispatch plus one
|
|
320
376
|
scoped re-review. Five rounds maximum per task:
|
|
321
377
|
|
|
@@ -361,15 +417,16 @@ dispatching. Adjudicate each open finding yourself — you hold the plan and
|
|
|
361
417
|
the cross-task context the reviewer lacks:
|
|
362
418
|
|
|
363
419
|
- **The reviewer is wrong, or the point is contestable:** park it —
|
|
364
|
-
`Task <N>: parked — <finding> —
|
|
420
|
+
`Task <N>: parked — <finding> — Ruling: <why the code stands>`. The final
|
|
365
421
|
review sees both sides.
|
|
366
422
|
- **Real, but nothing downstream builds on it:** park it the same way, with
|
|
367
423
|
a ruling that says it's real and deferred.
|
|
368
424
|
- **Real and load-bearing** — a later task builds on it, or it reveals a
|
|
369
|
-
plan defect:
|
|
370
|
-
|
|
371
|
-
the
|
|
372
|
-
|
|
425
|
+
plan defect: rule on the smallest change that unblocks the dependent work,
|
|
426
|
+
ledger it as `Task <N>: Ruling: <finding> — <what you decided and why>`,
|
|
427
|
+
and carry it into the next task's dispatch. Parking a structural failure
|
|
428
|
+
silently lets every dependent task build on it. Stop only when the defect
|
|
429
|
+
leaves every path forward a guess.
|
|
373
430
|
|
|
374
431
|
Adjudicate only at the cap. Adjudicating earlier to end a loop is
|
|
375
432
|
pre-judging with a different name. Every adjudication is a ledger entry —
|
|
@@ -392,8 +449,10 @@ parked-with-ruling at the cap.
|
|
|
392
449
|
## Final Review
|
|
393
450
|
|
|
394
451
|
The final whole-branch review gets a package too: run
|
|
395
|
-
`scripts/review-package PLAN_FILE MERGE_BASE HEAD` (or
|
|
396
|
-
|
|
452
|
+
`scripts/review-package PLAN_FILE MERGE_BASE HEAD` (or
|
|
453
|
+
`scripts/review-package.ps1 PLAN_FILE MERGE_BASE HEAD` on Windows PowerShell;
|
|
454
|
+
MERGE_BASE = the commit the branch started from, e.g. `git merge-base main HEAD`)
|
|
455
|
+
and include the
|
|
397
456
|
printed path in the final review dispatch, so the final reviewer reads
|
|
398
457
|
one file instead of re-deriving the branch diff with git commands. Dispatch
|
|
399
458
|
on the most capable available model (see Model Selection), using
|
|
@@ -407,15 +466,27 @@ with the complete findings list — not one fixer per finding.
|
|
|
407
466
|
Per-finding fixers each rebuild context and re-run suites; a real
|
|
408
467
|
session's final-review fix wave cost more than all its tasks combined.
|
|
409
468
|
Then run exactly one scoped re-review of the fix wave
|
|
410
|
-
(`scripts/review-package PLAN_FILE FIX_BASE HEAD
|
|
469
|
+
(`scripts/review-package PLAN_FILE FIX_BASE HEAD`, or
|
|
470
|
+
`scripts/review-package.ps1 PLAN_FILE FIX_BASE HEAD` on Windows PowerShell,
|
|
471
|
+
over the fix range,
|
|
411
472
|
[re-review-prompt.md](re-review-prompt.md)).
|
|
412
473
|
Adjudicate any residual findings as in the task loop's breaker: park with
|
|
413
|
-
rulings, or
|
|
474
|
+
rulings, or rule on the load-bearing ones and ledger what you decided. Only
|
|
475
|
+
the four classes above stop you here. There is no second fix wave —
|
|
414
476
|
residual load-bearing findings surface to your human partner when
|
|
415
477
|
finishing-a-development-branch presents the options.
|
|
416
478
|
|
|
417
479
|
## Finish
|
|
418
480
|
|
|
481
|
+
Before you delete anything, collect every ledger line containing `Ruling:` —
|
|
482
|
+
preflight rulings, parked findings, breaker adjudications, all of them — into
|
|
483
|
+
your final message under "Rulings I made", in the order you made them, each
|
|
484
|
+
with what it costs if wrong. The list is exhaustive: if the ledger holds a
|
|
485
|
+
ruling, the list holds it. That list is the only place the decisions you
|
|
486
|
+
took on your human partner's behalf reach them — they read it and rework
|
|
487
|
+
whatever you got wrong. A ruling that dies with the workspace was a decision
|
|
488
|
+
made in secret.
|
|
489
|
+
|
|
419
490
|
When the final whole-branch review is clean and its fixes are merged,
|
|
420
491
|
delete this plan's workspace (`rm -rf <workspace>`) — the git history is
|
|
421
492
|
the record now. Sibling directories belong to other plans; leave them
|
|
@@ -435,6 +506,7 @@ Use superpowers:finishing-a-development-branch.
|
|
|
435
506
|
| "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
|
|
436
507
|
| "Reviews slow the loop down" | The loop without reviews is just unverified churn. Reviews are the loop's brakes and steering. |
|
|
437
508
|
| "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Controllers without one have re-dispatched entire completed task sequences. |
|
|
509
|
+
| "The implementer spawned its own reviewer — free extra assurance" | It's a duplicate seat reviewing the same diff; the task review is the gate. A worker-spawned reviewer is a defect to flag, not rigor. |
|
|
438
510
|
|
|
439
511
|
## Example Workflow
|
|
440
512
|
|
|
@@ -47,6 +47,18 @@ Subagent (general-purpose):
|
|
|
47
47
|
While iterating, run the focused test for what you're changing; run the
|
|
48
48
|
full suite once before committing, not after every edit.
|
|
49
49
|
|
|
50
|
+
## You Do Not Dispatch Subagents
|
|
51
|
+
|
|
52
|
+
Do all of this task's work yourself. Never spawn a subagent to
|
|
53
|
+
implement part of the task, and above all never spawn a reviewer to
|
|
54
|
+
check your work. Self-review (below) means reading your own diff.
|
|
55
|
+
Review is the controller's job: after you report, it dispatches a
|
|
56
|
+
fresh reviewer against your diff. A reviewer you spawn duplicates
|
|
57
|
+
that review at full cost, and its approval counts for nothing in
|
|
58
|
+
the process. If you catch yourself thinking "an independent review
|
|
59
|
+
would strengthen my report" — that review is already scheduled.
|
|
60
|
+
Report instead.
|
|
61
|
+
|
|
50
62
|
## Code Organization
|
|
51
63
|
|
|
52
64
|
You reason best about code you can hold in context at once, and your edits are more
|
|
@@ -43,6 +43,15 @@ Subagent (general-purpose):
|
|
|
43
43
|
Your review is read-only on this checkout. Do not mutate the working
|
|
44
44
|
tree, the index, HEAD, or branch state in any way.
|
|
45
45
|
|
|
46
|
+
## You Do Not Dispatch Subagents
|
|
47
|
+
|
|
48
|
+
Do all of this review yourself. Never spawn a subagent to review part
|
|
49
|
+
of the diff, and never spawn another reviewer for a second opinion.
|
|
50
|
+
This process already provides every review seat the work gets; a
|
|
51
|
+
reviewer you spawn duplicates one of them at full cost, and its
|
|
52
|
+
verdict counts for nothing. If the diff feels too large for one
|
|
53
|
+
pass, review it in passes yourself and say so in your report.
|
|
54
|
+
|
|
46
55
|
## Scope
|
|
47
56
|
|
|
48
57
|
Your scope is the findings list and the fix diff. Verdict every finding.
|
|
@@ -22,10 +22,20 @@ head=$3
|
|
|
22
22
|
git rev-parse --verify --quiet "$base" >/dev/null || { echo "bad BASE: $base" >&2; exit 2; }
|
|
23
23
|
git rev-parse --verify --quiet "$head" >/dev/null || { echo "bad HEAD: $head" >&2; exit 2; }
|
|
24
24
|
|
|
25
|
+
git merge-base --is-ancestor "$base" "$head" || {
|
|
26
|
+
echo "HEAD ($head) is not a descendant of BASE ($base)" >&2
|
|
27
|
+
exit 3
|
|
28
|
+
}
|
|
29
|
+
|
|
30
|
+
if [ "$(git rev-list --count "${base}..${head}")" -eq 0 ]; then
|
|
31
|
+
echo "empty commit range: ${base}..${head}" >&2
|
|
32
|
+
exit 3
|
|
33
|
+
fi
|
|
34
|
+
|
|
25
35
|
if [ $# -eq 4 ]; then
|
|
26
36
|
out=$4
|
|
27
37
|
else
|
|
28
|
-
dir=$("$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan")
|
|
38
|
+
dir=$("${BASH:-bash}" "$(cd "$(dirname "$0")" && pwd)/sdd-workspace" "$plan")
|
|
29
39
|
out="$dir/review-$(git rev-parse --short "$base")..$(git rev-parse --short "$head").diff"
|
|
30
40
|
fi
|
|
31
41
|
|
|
@@ -31,6 +31,18 @@ if ($LASTEXITCODE -ne 0) {
|
|
|
31
31
|
exit 2
|
|
32
32
|
}
|
|
33
33
|
|
|
34
|
+
& git merge-base --is-ancestor $base $head *> $null
|
|
35
|
+
if ($LASTEXITCODE -ne 0) {
|
|
36
|
+
[Console]::Error.WriteLine("HEAD ($head) is not a descendant of BASE ($base)")
|
|
37
|
+
exit 3
|
|
38
|
+
}
|
|
39
|
+
|
|
40
|
+
$commitCount = (& git rev-list --count "${base}..${head}").Trim()
|
|
41
|
+
if ([int]$commitCount -eq 0) {
|
|
42
|
+
[Console]::Error.WriteLine("empty commit range: ${base}..${head}")
|
|
43
|
+
exit 3
|
|
44
|
+
}
|
|
45
|
+
|
|
34
46
|
if ($args.Count -eq 4) {
|
|
35
47
|
$out = $args[3]
|
|
36
48
|
} else {
|
|
@@ -53,7 +65,7 @@ $content.Add("")
|
|
|
53
65
|
$content.Add("## Diff")
|
|
54
66
|
(& git diff -U10 "${base}..${head}") | ForEach-Object { $content.Add($_) }
|
|
55
67
|
|
|
56
|
-
Set-Content -
|
|
68
|
+
Set-Content -LiteralPath $out -Value $content -Encoding utf8
|
|
57
69
|
$commits = (& git rev-list --count "${base}..${head}").Trim()
|
|
58
70
|
$bytes = (Get-Item -LiteralPath $out).Length
|
|
59
71
|
Write-Output "wrote ${out}: $commits commit(s), $bytes bytes"
|
|
@@ -3,11 +3,17 @@
|
|
|
3
3
|
# short-lived artifacts: task briefs, implementer reports, review packages,
|
|
4
4
|
# and the progress ledger. Print the plan directory's absolute path.
|
|
5
5
|
#
|
|
6
|
-
# One directory per plan (.superpowers/sdd/<plan-
|
|
6
|
+
# One directory per plan (.superpowers/sdd/<plan-slug>/) so a follow-up
|
|
7
7
|
# plan in the same working tree can never read or overwrite another plan's
|
|
8
8
|
# artifacts. A stale ledger misread as current progress makes controllers
|
|
9
9
|
# skip whole task sequences — plan-scoping removes that failure structurally.
|
|
10
10
|
#
|
|
11
|
+
# Ownership markers (plan-path file in the workspace directory) disambiguate
|
|
12
|
+
# same-basename plans (e.g. docs/alpha/plan.md vs docs/beta/plan.md) by appending
|
|
13
|
+
# the parent-directory name, then a counter. A workspace with no marker
|
|
14
|
+
# predates the marker scheme and is adopted for the current plan so in-flight
|
|
15
|
+
# workspaces keep resolving.
|
|
16
|
+
#
|
|
11
17
|
# The workspace lives in the working tree (not under .git/) because Claude Code
|
|
12
18
|
# treats .git/ as a protected path and denies agent writes there — which blocks
|
|
13
19
|
# an implementer subagent from writing its report file. A self-ignoring
|
|
@@ -28,13 +34,46 @@ fi
|
|
|
28
34
|
plan=$1
|
|
29
35
|
[ -f "$plan" ] || { echo "no such plan file: $plan" >&2; exit 2; }
|
|
30
36
|
|
|
31
|
-
slug=$(basename "$plan" .md)
|
|
37
|
+
slug=$(basename -- "$plan" .md)
|
|
32
38
|
[ -n "$slug" ] && [ "$slug" != "." ] && [ "$slug" != ".." ] \
|
|
33
39
|
|| { echo "cannot derive a workspace name from: $plan" >&2; exit 2; }
|
|
34
40
|
|
|
35
41
|
root=$(git rev-parse --show-toplevel)
|
|
42
|
+
root=$(CDPATH= cd -- "$root" && pwd -P)
|
|
36
43
|
base="$root/.superpowers/sdd"
|
|
44
|
+
|
|
45
|
+
# Normalize the plan path (physical directory, so relative/absolute/../
|
|
46
|
+
# spellings of one plan compare equal) and express it as the marker value:
|
|
47
|
+
# repo-relative when the plan lives under the repo root, absolute otherwise.
|
|
48
|
+
plan_dir=$(CDPATH= cd -- "$(dirname -- "$plan")" && pwd -P)
|
|
49
|
+
plan_abs="$plan_dir/$(basename -- "$plan")"
|
|
50
|
+
case "$plan_abs" in
|
|
51
|
+
"$root"/*) plan_id=${plan_abs#"$root"/} ;;
|
|
52
|
+
*) plan_id=$plan_abs ;;
|
|
53
|
+
esac
|
|
54
|
+
|
|
55
|
+
# True when the workspace at $1 is (or becomes) this plan's: an existing
|
|
56
|
+
# marker must name this plan; a missing marker means a new workspace or a
|
|
57
|
+
# pre-marker legacy one, and either way the plan claims it by writing one.
|
|
58
|
+
owns() {
|
|
59
|
+
if [ -e "$1/plan-path" ]; then
|
|
60
|
+
[ "$(cat "$1/plan-path")" = "$plan_id" ]
|
|
61
|
+
else
|
|
62
|
+
mkdir -p "$1"
|
|
63
|
+
printf '%s\n' "$plan_id" > "$1/plan-path"
|
|
64
|
+
fi
|
|
65
|
+
}
|
|
66
|
+
|
|
37
67
|
dir="$base/$slug"
|
|
38
|
-
|
|
68
|
+
if ! owns "$dir"; then
|
|
69
|
+
parent=$(basename -- "$plan_dir")
|
|
70
|
+
dir="$base/$slug-$parent"
|
|
71
|
+
if ! owns "$dir"; then
|
|
72
|
+
n=2
|
|
73
|
+
while ! owns "$base/$slug-$parent-$n"; do n=$((n + 1)); done
|
|
74
|
+
dir="$base/$slug-$parent-$n"
|
|
75
|
+
fi
|
|
76
|
+
fi
|
|
77
|
+
|
|
39
78
|
printf '*\n' > "$base/.gitignore"
|
|
40
|
-
cd "$dir" && pwd
|
|
79
|
+
CDPATH= cd -- "$dir" && pwd
|
|
@@ -22,7 +22,7 @@ if (-not (Test-Path -LiteralPath $plan -PathType Leaf)) {
|
|
|
22
22
|
exit 2
|
|
23
23
|
}
|
|
24
24
|
|
|
25
|
-
$slug = [System.IO.Path]::GetFileName($plan) -
|
|
25
|
+
$slug = [System.IO.Path]::GetFileName($plan) -creplace '\.md$', ''
|
|
26
26
|
if ([string]::IsNullOrEmpty($slug) -or $slug -eq "." -or $slug -eq "..") {
|
|
27
27
|
[Console]::Error.WriteLine("cannot derive a workspace name from: $plan")
|
|
28
28
|
exit 2
|
|
@@ -30,7 +30,69 @@ if ([string]::IsNullOrEmpty($slug) -or $slug -eq "." -or $slug -eq "..") {
|
|
|
30
30
|
|
|
31
31
|
$root = (& git rev-parse --show-toplevel).Trim()
|
|
32
32
|
$base = Join-Path $root ".superpowers/sdd"
|
|
33
|
+
|
|
34
|
+
function Get-PhysicalDirectoryPath($path) {
|
|
35
|
+
if (-not (Test-Path -LiteralPath $path)) { return $path }
|
|
36
|
+
if ($IsWindows) {
|
|
37
|
+
return (Resolve-Path -LiteralPath $path).Path
|
|
38
|
+
}
|
|
39
|
+
$orig = Get-Location
|
|
40
|
+
try {
|
|
41
|
+
Set-Location -LiteralPath $path
|
|
42
|
+
$pwdCmd = (Get-Command -Type Application pwd -ErrorAction SilentlyContinue).Source
|
|
43
|
+
if ($pwdCmd) {
|
|
44
|
+
return (& $pwdCmd -P).Trim()
|
|
45
|
+
} else {
|
|
46
|
+
return (Resolve-Path -LiteralPath $path).Path
|
|
47
|
+
}
|
|
48
|
+
} finally {
|
|
49
|
+
Set-Location $orig
|
|
50
|
+
}
|
|
51
|
+
}
|
|
52
|
+
|
|
53
|
+
# Normalize the plan path (physical directory, so relative/absolute/../
|
|
54
|
+
# spellings of one plan compare equal) and express it as the marker value:
|
|
55
|
+
# repo-relative when the plan lives under the repo root, absolute otherwise.
|
|
56
|
+
$planLeaf = Split-Path -Leaf $plan
|
|
57
|
+
$planParent = Split-Path -Parent $plan
|
|
58
|
+
if ([string]::IsNullOrEmpty($planParent)) { $planParent = "." }
|
|
59
|
+
|
|
60
|
+
$planDir = Get-PhysicalDirectoryPath $planParent
|
|
61
|
+
$rootPhys = Get-PhysicalDirectoryPath $root
|
|
62
|
+
|
|
63
|
+
$planAbs = (Join-Path $planDir $planLeaf) -replace '\\', '/'
|
|
64
|
+
$rootNorm = $rootPhys -replace '\\', '/'
|
|
65
|
+
|
|
66
|
+
if ($planAbs.StartsWith($rootNorm + "/", [System.StringComparison]::OrdinalIgnoreCase)) {
|
|
67
|
+
$planId = $planAbs.Substring($rootNorm.Length + 1)
|
|
68
|
+
} else {
|
|
69
|
+
$planId = $planAbs
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
function Test-And-Claim-Workspace($targetDir, $id) {
|
|
73
|
+
$markerPath = Join-Path $targetDir "plan-path"
|
|
74
|
+
if (Test-Path -LiteralPath $markerPath -PathType Leaf) {
|
|
75
|
+
$existingId = (Get-Content -LiteralPath $markerPath -Raw).Trim()
|
|
76
|
+
return ($existingId -eq $id)
|
|
77
|
+
} else {
|
|
78
|
+
New-Item -ItemType Directory -Force -Path $targetDir | Out-Null
|
|
79
|
+
Set-Content -LiteralPath $markerPath -Value $id -Encoding ascii
|
|
80
|
+
return $true
|
|
81
|
+
}
|
|
82
|
+
}
|
|
83
|
+
|
|
33
84
|
$dir = Join-Path $base $slug
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
85
|
+
if (-not (Test-And-Claim-Workspace $dir $planId)) {
|
|
86
|
+
$parent = Split-Path -Leaf $planDir
|
|
87
|
+
$dir = Join-Path $base "$slug-$parent"
|
|
88
|
+
if (-not (Test-And-Claim-Workspace $dir $planId)) {
|
|
89
|
+
$n = 2
|
|
90
|
+
while (-not (Test-And-Claim-Workspace (Join-Path $base "$slug-$parent-$n") $planId)) {
|
|
91
|
+
$n++
|
|
92
|
+
}
|
|
93
|
+
$dir = Join-Path $base "$slug-$parent-$n"
|
|
94
|
+
}
|
|
95
|
+
}
|
|
96
|
+
|
|
97
|
+
Set-Content -LiteralPath (Join-Path $base ".gitignore") -Value "*" -NoNewline -Encoding ascii
|
|
98
|
+
(Resolve-Path -LiteralPath $dir).Path
|
|
@@ -42,7 +42,7 @@ foreach ($line in [System.IO.File]::ReadLines((Resolve-Path -LiteralPath $plan).
|
|
|
42
42
|
}
|
|
43
43
|
}
|
|
44
44
|
|
|
45
|
-
Set-Content -
|
|
45
|
+
Set-Content -LiteralPath $out -Value $selected -Encoding utf8
|
|
46
46
|
if ((-not (Test-Path -LiteralPath $out)) -or ((Get-Item -LiteralPath $out).Length -eq 0)) {
|
|
47
47
|
[Console]::Error.WriteLine("task $taskNumber not found in $plan (no heading matching 'Task $taskNumber')")
|
|
48
48
|
exit 3
|
|
@@ -52,6 +52,15 @@ Subagent (general-purpose):
|
|
|
52
52
|
Your review is read-only on this checkout. Do not mutate the working
|
|
53
53
|
tree, the index, HEAD, or branch state in any way.
|
|
54
54
|
|
|
55
|
+
## You Do Not Dispatch Subagents
|
|
56
|
+
|
|
57
|
+
Do all of this review yourself. Never spawn a subagent to review part
|
|
58
|
+
of the diff, and never spawn another reviewer for a second opinion.
|
|
59
|
+
This process already provides every review seat the work gets; a
|
|
60
|
+
reviewer you spawn duplicates one of them at full cost, and its
|
|
61
|
+
verdict counts for nothing. If the diff feels too large for one
|
|
62
|
+
pass, review it in passes yourself and say so in your report.
|
|
63
|
+
|
|
55
64
|
## Do Not Trust the Report
|
|
56
65
|
|
|
57
66
|
Treat the implementer's report as unverified claims about the code. It
|
|
@@ -75,6 +84,13 @@ Subagent (general-purpose):
|
|
|
75
84
|
Warnings or other noise in the implementer's reported test output are
|
|
76
85
|
findings — test output should be pristine.
|
|
77
86
|
|
|
87
|
+
Evidence you cannot see is not evidence that doesn't exist. If the
|
|
88
|
+
report or its test evidence looks truncated, or you cannot locate the
|
|
89
|
+
results it claims, re-read the file at its stated path — and if it is
|
|
90
|
+
genuinely missing or garbled, report that as a gap for the controller.
|
|
91
|
+
Re-running the suite to regenerate what you failed to read is not
|
|
92
|
+
verification; illegibility of the evidence is not invalidation of it.
|
|
93
|
+
|
|
78
94
|
## Part 1: Spec Compliance
|
|
79
95
|
|
|
80
96
|
Compare the diff against What Was Requested:
|
|
@@ -86,6 +102,12 @@ Subagent (general-purpose):
|
|
|
86
102
|
- **Misunderstood:** right feature built the wrong way, wrong problem
|
|
87
103
|
solved
|
|
88
104
|
|
|
105
|
+
If the brief lists several files each with its own change (a batched
|
|
106
|
+
dispatch), check the diff against that list file by file: every listed
|
|
107
|
+
file must have its corresponding hunk. A listed file the diff never
|
|
108
|
+
touches is a Missing finding, no matter how clean the rest of the
|
|
109
|
+
batch looks.
|
|
110
|
+
|
|
89
111
|
If a requirement cannot be verified from this diff alone (it lives in
|
|
90
112
|
unchanged code or spans tasks), report it as a ⚠️ item instead of
|
|
91
113
|
broadening your search.
|
|
@@ -167,8 +189,7 @@ Subagent (general-purpose):
|
|
|
167
189
|
|
|
168
190
|
**Placeholders:**
|
|
169
191
|
- `[MODEL]` — REQUIRED: reviewer model per SKILL.md Model Selection
|
|
170
|
-
- `[BRIEF_FILE]` — REQUIRED: the task brief file (`scripts/task-brief PLAN N`,
|
|
171
|
-
or `scripts/task-brief.ps1 PLAN N` on Windows PowerShell,
|
|
192
|
+
- `[BRIEF_FILE]` — REQUIRED: the task brief file (`scripts/task-brief PLAN N`, or `scripts/task-brief.ps1 PLAN N` on Windows PowerShell,
|
|
172
193
|
prints the path; same file the implementer worked from)
|
|
173
194
|
- `[GLOBAL_CONSTRAINTS]` — the binding requirements copied verbatim from
|
|
174
195
|
the plan's Global Constraints section or the spec: exact values, formats,
|
|
@@ -179,9 +200,8 @@ Subagent (general-purpose):
|
|
|
179
200
|
- `[BASE_SHA]` — commit before this task
|
|
180
201
|
- `[HEAD_SHA]` — current commit
|
|
181
202
|
- `[DIFF_FILE]` — REQUIRED: the path the controller wrote the review
|
|
182
|
-
package to (`scripts/review-package PLAN_FILE BASE HEAD`, or
|
|
183
|
-
|
|
184
|
-
prints the unique path it wrote; the package never enters the controller's context)
|
|
203
|
+
package to (`scripts/review-package PLAN_FILE BASE HEAD`, or `scripts/review-package.ps1 PLAN_FILE BASE HEAD` on Windows PowerShell, prints the unique
|
|
204
|
+
path it wrote; the package never enters the controller's context)
|
|
185
205
|
|
|
186
206
|
**Reviewer returns:** Spec Compliance verdict (✅/❌/⚠️), Strengths, Issues
|
|
187
207
|
(Critical/Important/Minor), Task quality verdict
|
|
@@ -182,6 +182,16 @@ Confirm:
|
|
|
182
182
|
|
|
183
183
|
**Other tests fail?** Fix now.
|
|
184
184
|
|
|
185
|
+
**"Other tests" means the project's suite, not just your file.** A
|
|
186
|
+
green run of the test you wrote is not a green suite. Before you call
|
|
187
|
+
the change done, run the project's test command (bare `pytest`,
|
|
188
|
+
`npm test`, `cargo test` — whatever the repo uses) even when your task
|
|
189
|
+
named only one test file. A scope statement in your task bounds the
|
|
190
|
+
deliverable, not your verification. Any failure that run shows —
|
|
191
|
+
including one you didn't cause — goes in your report by name; a red
|
|
192
|
+
test you watched scroll past and didn't mention is a report falsified
|
|
193
|
+
by omission.
|
|
194
|
+
|
|
185
195
|
### REFACTOR - Clean Up
|
|
186
196
|
|
|
187
197
|
After green only:
|
|
@@ -56,6 +56,7 @@ If your harness appears here, read its reference file for special instructions:
|
|
|
56
56
|
- Codex: `references/codex-tools.md`
|
|
57
57
|
- Pi: `references/pi-tools.md`
|
|
58
58
|
- Antigravity: `references/antigravity-tools.md`
|
|
59
|
+
- Hermes Agent: `references/hermes-tools.md`
|
|
59
60
|
|
|
60
61
|
## User Instructions
|
|
61
62
|
|