superpowers-mcp 6.3.6 → 6.3.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -28,7 +28,11 @@ For each task:
28
28
  1. Mark as in_progress
29
29
  2. Follow each step exactly (plan has bite-sized steps)
30
30
  3. Run verifications as specified
31
- 4. Mark as completed
31
+ 4. Mark as completed — the todo **and** the plan file: edit it and flip this
32
+ task's steps from `- [ ]` to `- [x]`. The plan is the artifact a human
33
+ reads to see where things stand, and nothing else in this skill writes
34
+ back to it. Leave a step unticked only when its deliverable does not exist
35
+ yet; never tick one you skipped.
32
36
 
33
37
  ### Step 3: Complete Development
34
38
 
@@ -62,3 +66,5 @@ After all tasks complete and verified:
62
66
  - Reference skills when plan says to
63
67
  - Stop when blocked, don't guess
64
68
  - Never start implementation on main/master branch without explicit user consent
69
+ - Commits stay local — no push/pull/fetch unless the plan or your human partner says so. At each checkpoint glance at `git status -sb`: a branch tracking or ahead of a shared branch (main/dev/production) is a stop-and-fix, not a footnote
70
+ - Never rewrite a shared branch (`git push --force*` to main/dev/production) to undo a mistake — a forward-only `git revert` is the only remedy you apply yourself; anything more is your human partner's call
@@ -85,6 +85,14 @@ is theirs.
85
85
 
86
86
  ### Option 1: Merge Locally
87
87
 
88
+ If the branch came from subagent-driven-development, commit the plan ledger's
89
+ deferred findings first: carry every ledger line tagged `minor (deferred)`,
90
+ `parked`, or `Ruling:` into
91
+ `docs/superpowers/follow-ups/<plan-basename>.md` (append under a dated
92
+ heading if the file already exists) and commit it to the branch before the
93
+ merge — subagent-driven-development's Finish rule owns the details, and the
94
+ file must land here so it survives the branch deletion.
95
+
88
96
  ```bash
89
97
  # Get main repo root for CWD safety
90
98
  MAIN_ROOT=$(git -C "$(git rev-parse --git-common-dir)/.." rev-parse --show-toplevel)
@@ -123,6 +131,13 @@ tooling — its CLI if one is available, or the creation URL most forges
123
131
  print when you push — following the repo's PR template and conventions if
124
132
  present, and report the URL to your human partner.
125
133
 
134
+ If the branch came from subagent-driven-development, append the plan
135
+ ledger's deferred findings to the PR description before reporting the URL —
136
+ a "Deferred items" checklist carrying every ledger line tagged `minor
137
+ (deferred)`, `parked`, or `Ruling:`, verbatim. Subagent-driven-development's
138
+ Finish rule owns the details; the checklist must land here, where your human
139
+ partner reviews.
140
+
126
141
  Keep the worktree — your human partner iterates on PR feedback there.
127
142
 
128
143
  ### Option 3: Keep As-Is
@@ -87,8 +87,9 @@ digraph process {
87
87
  "More tasks remain?" [shape=diamond];
88
88
  "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" [shape=box];
89
89
  "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" [shape=box];
90
- "Final review clean: delete this plan's workspace" [shape=box];
91
- "Use superpowers:finishing-a-development-branch" [shape=box style=filled fillcolor=lightgreen];
90
+ "Use superpowers:finishing-a-development-branch" [shape=box];
91
+ "Finish path resolved: export deferred findings to its durable artifact" [shape=box];
92
+ "Delete this plan's workspace (only after the export; keep-as-is keeps it)" [shape=box style=filled fillcolor=lightgreen];
92
93
 
93
94
  "Setup: worktree, ledger check, read plan, pre-flight review" -> "Dispatch implementer subagent (./implementer-prompt.md)";
94
95
  "Dispatch implementer subagent (./implementer-prompt.md)" -> "Implementer asks questions?";
@@ -116,8 +117,9 @@ digraph process {
116
117
  "More tasks remain?" -> "Dispatch implementer subagent (./implementer-prompt.md)" [label="yes"];
117
118
  "More tasks remain?" -> "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" [label="no"];
118
119
  "Dispatch final code reviewer (../requesting-code-review/code-reviewer.md)" -> "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals";
119
- "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" -> "Final review clean: delete this plan's workspace";
120
- "Final review clean: delete this plan's workspace" -> "Use superpowers:finishing-a-development-branch";
120
+ "Final findings? ONE fix dispatch, one scoped re-review, adjudicate residuals" -> "Use superpowers:finishing-a-development-branch";
121
+ "Use superpowers:finishing-a-development-branch" -> "Finish path resolved: export deferred findings to its durable artifact";
122
+ "Finish path resolved: export deferred findings to its durable artifact" -> "Delete this plan's workspace (only after the export; keep-as-is keeps it)";
121
123
  }
122
124
  ```
123
125
 
@@ -142,14 +144,21 @@ a ledger file, not only in todos.
142
144
  line names your plan file, tasks with a `Task <N>: complete` line are DONE
143
145
  — do not re-dispatch them; resume at the first task without one. A task
144
146
  whose last line is a fix round is mid-loop: resume the loop at the next
145
- round. A ledger whose first line names a different plan file — or a stray
147
+ round. A resumed controller also reads the ledger's `## Discoveries`
148
+ section: it holds what completed tasks found that the plan could not
149
+ know, and it is the source for clause (3) of the next dispatch. A ledger
150
+ whose first line names a different plan file — or a stray
146
151
  ledger at the old flat path `.superpowers/sdd/progress.md` — is another
147
152
  plan's progress: leave it in place and start your own, fresh.
148
153
  - Create the ledger with its identity as the first line:
149
154
  `# SDD ledger — plan: <plan file path>`.
150
155
  - The ledger is your recovery map: the commits it names exist in git even
151
156
  when your context no longer remembers creating them. After compaction,
152
- trust the ledger and `git log` over your own recollection.
157
+ trust the ledger and `git log` over your own recollection. The ledger's
158
+ `## Discoveries` section is the same map for knowledge: what earlier tasks
159
+ found that the plan could not know. A resumed controller composes clause
160
+ (3) of every dispatch from that section, never from recollection of a
161
+ context it no longer has.
153
162
  - `git clean -fdx` will destroy the workspace (it's git-ignored scratch); if
154
163
  that happens, recover from `git log`.
155
164
 
@@ -272,9 +281,12 @@ and fix-round diffs need it.
272
281
  task fits in the project; (2) the brief path, introduced as "read this
273
282
  first — it is your requirements, with the exact values to use verbatim";
274
283
  (3) interfaces and decisions from earlier tasks that the brief cannot
275
- know; (4) your resolution of any ambiguity you noticed in the brief;
276
- (5) the report-file path and report contract. Exact values (numbers,
277
- magic strings, signatures, test cases) appear only in the brief. Never
284
+ know — read them from the ledger's Discoveries section, not from your
285
+ memory of past reports: copy the entries the task's interfaces touch, in
286
+ full, and skip the rest; (4) your resolution of any ambiguity you noticed
287
+ in the brief; (5) the report-file path and report contract. Exact values
288
+ (numbers, magic strings, signatures, test cases) appear only in the
289
+ brief. Never
278
290
  make a subagent read the whole plan file. When the brief is a contract
279
291
  (goal, success criteria, interfaces) rather than written-out code, item
280
292
  (3) also carries the elaboration the contract leaves to dispatch time:
@@ -290,7 +302,10 @@ and fix-round diffs need it.
290
302
  paste accumulated prior-task summaries ("state after Tasks 1-3") into
291
303
  later dispatches — a real session's dispatch hit 42k chars of which 99%
292
304
  was pasted history. A fresh subagent needs its task, the interfaces it
293
- touches, and the global constraints. Nothing else.
305
+ touches, and the global constraints. Nothing else. A curated slice copied
306
+ from the Discoveries section is not pasted history — it is clause (3)
307
+ above, kept small enough to read in full because each entry is a delta
308
+ from the plan.
294
309
  - The dispatch carries the no-subagents contract (it is in the
295
310
  implementer template): the implementer never dispatches subagents —
296
311
  not helpers, and never a reviewer. Review arrives from you, after the
@@ -366,9 +381,16 @@ needed.
366
381
  call. Use the BASE you recorded before dispatching the implementer —
367
382
  never `HEAD~1`, which silently truncates multi-commit tasks. Never
368
383
  dispatch a task reviewer without a diff file.
369
- - **Reviewer inputs:** the task reviewer gets three paths — the same brief
370
- file, the report file, and the review package — plus the global
371
- constraints that bind the task.
384
+ - **Reviewer inputs:** the task reviewer gets four paths — the same brief
385
+ file, the report file, the review package, and a review file to write
386
+ to (brief `…/task-N-brief.md` → review `…/task-N-review.md`) — plus
387
+ the global constraints that bind the task.
388
+ - **Review file:** the reviewer writes its full report there and returns
389
+ only verdicts, ⚠️ items, one line per Critical/Important finding, a
390
+ Minor count, and the path. Don't read the review file during the loop —
391
+ the final message is your decision surface; the detail is for fix
392
+ subagents. One exception: when you ledger a task's deferred minors, copy
393
+ each one-liner from the review file's Minor section.
372
394
  - The global-constraints block you hand the reviewer is its attention
373
395
  lens. Copy the binding requirements verbatim from the plan's Global
374
396
  Constraints section or the spec: exact values, exact formats, and the
@@ -402,9 +424,12 @@ finding, or a ⚠️ item you confirmed as a real gap.
402
424
 
403
425
  Before the loop starts, two routes leave it immediately:
404
426
 
405
- - Record Minor findings in the progress ledger as you go
406
- (`Task <N>: minor (deferred): <one-liner>`), and point the final
407
- whole-branch review at that list so it can triage which must be fixed
427
+ - Dispatch fix subagents for Critical and Important findings, passing the
428
+ review file path for the full detail. Record each task's Minor count,
429
+ review file path, and one ledger line per deferred minor copied from the
430
+ review file's Minor section (`Task <N>: minor (deferred): <one-liner>` —
431
+ the deferred-findings export greps these lines), and point the final
432
+ whole-branch review at those files so it can triage which must be fixed
408
433
  before merge. A roll-up nobody reads is a silent discard. Minor findings
409
434
  never enter the loop.
410
435
  - A finding labeled plan-mandated — or any finding that conflicts with
@@ -417,14 +442,16 @@ Everything else enters the loop. A fix round is one fix dispatch plus one
417
442
  scoped re-review. Five rounds maximum per task:
418
443
 
419
444
  **Rounds 1-3 — resume the original implementer.** Send it the open findings
420
- verbatim. Its context is intact: it knows the task, the code, and its own
421
- choices. If your harness cannot send another message to a live subagent,
422
- dispatch a fresh implementer carrying the brief path, the report-file path,
423
- and the findings — the report file is the persistent memory either way.
445
+ verbatim, plus the review file path for the full detail. Its context is
446
+ intact: it knows the task, the code, and its own choices. If your harness
447
+ cannot send another message to a live subagent, dispatch a fresh
448
+ implementer carrying the brief path, the report-file path, the review file
449
+ path, and the findings — the report file is the persistent memory either
450
+ way.
424
451
 
425
452
  **Rounds 4-5 — dispatch a fresh implementer on a more capable model** (per
426
- Model Selection), with the brief path, the report-file path, the open
427
- findings, and this framing: "A prior implementer attempted this task
453
+ Model Selection), with the brief path, the report-file path, the review
454
+ file path, the open findings, and this framing: "A prior implementer attempted this task
428
455
  [N] times; you own it now. Read the report file for what was tried." A loop
429
456
  that survives three resumes usually means the implementer cannot see its
430
457
  own problem — fresh eyes and a capability bump in one move.
@@ -441,7 +468,8 @@ whole suite.
441
468
  (or `scripts/review-package.ps1 PLAN_FILE FIX_BASE HEAD` on Windows PowerShell)
442
469
  where FIX_BASE is the head the previous review saw, and dispatch
443
470
  [re-review-prompt.md](re-review-prompt.md) with the findings list, the
444
- brief, the report file, and the printed diff path. The re-reviewer verdicts
471
+ brief, the report file, the review file to append its verdicts to, and the
472
+ printed diff path. The re-reviewer verdicts
445
473
  each finding ADDRESSED or NOT ADDRESSED and flags new breakage in the fix
446
474
  diff only. New Critical/Important breakage in the fix diff joins the open
447
475
  findings list. Out-of-scope observations go to the ledger as deferred
@@ -483,6 +511,27 @@ message as your other bookkeeping:
483
511
  - `Task <N>: complete (commits <base7>..<head7>, <K> parked)` after a
484
512
  tripped breaker
485
513
 
514
+ Copy the task's Discoveries into the ledger in the same message, from the
515
+ report's "Discoveries for later tasks" field: append a `### Task <N>`
516
+ heading under the ledger's `## Discoveries` section (create the section on
517
+ first use) with each discovery as one line, or `- None` when the field is
518
+ empty. Keep it a delta from the plan: only what a later task needs and the
519
+ plan could not know — corrections to the plan, negative results, resolved
520
+ unknowns, interfaces discovered in the code. If it does not change what a
521
+ later task does, it does not go in.
522
+
523
+ In the same message, **edit the plan file and flip this task's steps from
524
+ `- [ ]` to `- [x]`**. The plan is what a human reads to see where the work
525
+ stands; the ledger is yours. Tie the two to this one event so they cannot
526
+ drift — a plan left at 0 while its tasks are done and deployed reads as
527
+ "nothing happened", and nothing downstream will catch it: the task reviewer
528
+ sees a diff, the re-review sees findings, the final review sees the ledger.
529
+ None of them ever opens the plan.
530
+
531
+ Leave a step unticked only when its deliverable does not exist yet — a step
532
+ that waits on someone else, or on an event that has not happened. Never tick
533
+ a step you skipped.
534
+
486
535
  **On a skeleton-first plan,** write one plan-check line with the
487
536
  completion line. Re-read the remaining tasks against what this task
488
537
  actually established — interfaces as built, environment facts,
@@ -541,13 +590,45 @@ took on your human partner's behalf reach them — they read it and rework
541
590
  whatever you got wrong. A ruling that dies with the workspace was a decision
542
591
  made in secret.
543
592
 
544
- When the final whole-branch review is clean and its fixes are merged,
545
- delete this plan's workspace (`rm -rf <workspace>`) — the git history is
546
- the record now. Sibling directories belong to other plans; leave them
547
- alone.
593
+ Rulings are not the only content that dies with the workspace. Findings you
594
+ chose not to fix are the record of what was *not* done — git history cannot
595
+ carry them, because git records what was done. So export them before
596
+ anything is deleted: grep the progress ledger for its three finding tags
597
+
598
+ ```bash
599
+ grep -E 'Ruling:|^(Task [0-9]+: )?(minor \(deferred\)|parked)' <workspace>/progress.md
600
+ ```
601
+
602
+ and carry every matching line, verbatim, into a durable, human-reachable
603
+ artifact. The chat roll-up does not satisfy this — scrollback, compaction,
604
+ and session end all eat it. Where the artifact lives depends on how
605
+ finishing-a-development-branch resolves (it carries the same obligation from
606
+ its side):
607
+
608
+ - **Option 2 (push and create PR):** append the lines to the PR description
609
+ under a "Deferred items" checklist. That is where your human partner
610
+ reviews, and checkboxes survive the merge.
611
+ - **Option 1 (merge locally):** write them to
612
+ `docs/superpowers/follow-ups/<plan-basename>.md` — append under a dated
613
+ heading if a re-run under the same basename already created the file — and
614
+ commit the file to the branch before the merge so it survives the branch
615
+ deletion.
616
+ - **Option 3 (keep as-is):** the workspace stays, so there is nothing to
617
+ export.
618
+ - **Explicit discard:** your human partner asked to throw the work away, so
619
+ no export is required.
548
620
 
549
621
  Use superpowers:finishing-a-development-branch.
550
622
 
623
+ When the finish path has resolved and its export exists — the checklist in
624
+ the PR description, or the committed follow-ups file — delete this plan's
625
+ workspace (`rm -rf <workspace>`), provided the final whole-branch review was
626
+ clean and its fixes are merged. The export, not git history, is the record
627
+ of the deferred findings. Sibling directories belong to other plans; leave
628
+ them alone. On an explicit discard there is no export, but the workspace is
629
+ deleted together with the branch — a ledger describing thrown-away work is
630
+ a false record.
631
+
551
632
  ## Common Rationalizations
552
633
 
553
634
  | Excuse | Reality |
@@ -586,11 +667,12 @@ Implementer: [Later]
586
667
  - Self-review: Found I missed --force flag, added it
587
668
  - Committed
588
669
 
589
- [Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path]
590
- Task reviewer: Spec ✅ - all requirements met, nothing extra.
591
- Strengths: Good test coverage, clean. Issues: None. Task quality: Approved.
670
+ [Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path + review file]
671
+ Task reviewer: Spec ✅. Quality: Approved. Minor: 0.
672
+ Full report: task-1-review.md
592
673
 
593
674
  [Ledger: Task 1: complete (commits a1b2c3d..d4e5f6a, review clean)]
675
+ [Ledger: Discoveries — Task 1: hooks dir must be created before install; none other]
594
676
 
595
677
  Task 2: Recovery modes
596
678
 
@@ -601,22 +683,23 @@ Implementer: [No questions]
601
683
  - 8/8 tests passing
602
684
  - Committed
603
685
 
604
- [Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path]
686
+ [Run review-package PLAN_FILE BASE HEAD; dispatch task reviewer with the printed path + review file]
605
687
  Task reviewer: Spec ❌:
606
688
  - Missing: Progress reporting (spec says "report every 100 items")
607
- Issues (Important): Magic number (100)
689
+ Important: Magic number (100). Minor: 0. Full report: task-2-review.md
608
690
 
609
- [Fix round 1: resume the implementer with both findings]
691
+ [Fix round 1: resume the implementer with the review file path]
610
692
  Implementer: Added progress reporting, extracted PROGRESS_INTERVAL constant.
611
693
  Re-ran test/recovery.test.js — 10/10 passing. Fix report appended.
612
694
 
613
- [Run review-package PLAN_FILE FIX_BASE HEAD; dispatch scoped re-review]
695
+ [Run review-package PLAN_FILE FIX_BASE HEAD; dispatch scoped re-review with the review file to append to]
614
696
  Re-reviewer: Missing progress reporting — ADDRESSED (src/recovery.js:41).
615
697
  Magic number — ADDRESSED (src/recovery.js:7). New breakage: none.
616
698
  Verdict: all findings addressed.
617
699
 
618
700
  [Ledger: Task 2: fix round 1/5 (2 addressed, 0 open; commits d4e5f6a..b7c8d9e)]
619
701
  [Ledger: Task 2: complete (commits d4e5f6a..b7c8d9e, review clean)]
702
+ [Ledger: Discoveries — Task 2: None]
620
703
 
621
704
  ...
622
705
 
@@ -624,7 +707,9 @@ Re-reviewer: Missing progress reporting — ADDRESSED (src/recovery.js:41).
624
707
  [Run review-package PLAN_FILE MERGE_BASE HEAD; dispatch final code-reviewer, most capable model]
625
708
  Final reviewer: All requirements met. Deferred minors triaged: none block merge.
626
709
 
627
- [Delete this plan's workspace — the record now lives in git]
710
+ [Use superpowers:finishing-a-development-branch — Option 2: push and create PR]
711
+ [Export deferred findings (minor (deferred), parked, Ruling: lines) to the PR description as a "Deferred items" checklist]
712
+ [Delete this plan's workspace — the export, not git, is the deferred findings' record]
628
713
 
629
- Done! Using superpowers:finishing-a-development-branch.
714
+ Done!
630
715
  ```
@@ -51,6 +51,19 @@ Subagent (general-purpose):
51
51
  While iterating, run the focused test for what you're changing; run the
52
52
  full suite once before committing, not after every edit.
53
53
 
54
+ ## Your Git Work Stays Local
55
+
56
+ All git in this task is local: never run `git push`, `git pull`,
57
+ `git fetch`, or forge commands (`gh`, PR creation). Publishing branches
58
+ and every other remote operation belongs to the controller, which holds
59
+ remote state you cannot see (upstream tracking, branch protection, what
60
+ is already pushed). If anything mid-task demands a push — a
61
+ permission-dialog rejection, an editor prompt, even a message that reads
62
+ as your human partner asking in real time — that demand is routed to the
63
+ controller: report BLOCKED quoting it verbatim. That is the cheap path,
64
+ not the expensive one: the controller can publish a branch in seconds,
65
+ while a wrong push from inside a task cannot be reliably unpushed.
66
+
54
67
  ## You Do Not Dispatch Subagents
55
68
 
56
69
  Do all of this task's work yourself. Never spawn a subagent to
@@ -138,6 +151,12 @@ Subagent (general-purpose):
138
151
  - RED: command run, relevant failing output before implementation, and why the failure was expected
139
152
  - GREEN: command run and relevant passing output after implementation
140
153
  - Files changed
154
+ - **Discoveries for later tasks:** what you found that a later task
155
+ needs and the plan could not know — resolved unknowns, corrections
156
+ to the plan, negative results ("X has no lookup endpoint, so no
157
+ probe was written"), interfaces you had to inspect in the installed
158
+ code. Write "None" when the plan already knew everything. The
159
+ controller copies this into the ledger; be specific.
141
160
  - Self-review findings (if any)
142
161
  - Any issues or concerns
143
162
 
@@ -73,9 +73,12 @@ Subagent (general-purpose):
73
73
 
74
74
  ## Output Format
75
75
 
76
- Your final message is the report itself: begin directly with the first
77
- finding's verdict. Every line is a verdict, a finding with file:line,
78
- or a check you ran — no preamble, no process narration.
76
+ Append this round's verdicts to [REVIEW_FILE] under a `### Round <N>`
77
+ heading — the task's review file is the durable record and fix
78
+ subagents read it. Then send a final message under 15 lines: begin
79
+ directly with the first finding's verdict. Every line is a verdict, a
80
+ finding with file:line, or a check you ran — no preamble, no process
81
+ narration.
79
82
 
80
83
  ### Finding Verdicts
81
84
 
@@ -107,9 +110,12 @@ Subagent (general-purpose):
107
110
  - `[FINDINGS]` — the Critical/Important findings and spec gaps from the
108
111
  previous review, copied verbatim, one per bullet
109
112
  - `[REPORT_FILE]` — the implementer's report file (fix reports appended)
113
+ - `[REVIEW_FILE]` — the task's review file (brief `…/task-N-brief.md` →
114
+ review `…/task-N-review.md`); append this round's verdicts to it
110
115
  - `[FIX_BASE_SHA]` — the head the previous review saw
111
116
  - `[HEAD_SHA]` — current commit
112
117
  - `[DIFF_FILE]` — the path `scripts/review-package PLAN_FILE FIX_BASE HEAD` printed
113
118
 
114
119
  **Re-reviewer returns:** per-finding verdicts (ADDRESSED / NOT ADDRESSED),
115
- new breakage in the fix diff, out-of-scope observations, and a round verdict.
120
+ new breakage in the fix diff, out-of-scope observations, and a round verdict —
121
+ appended to `[REVIEW_FILE]`, with a short final message.
@@ -19,6 +19,12 @@ base=$2
19
19
  head=$3
20
20
  [ -f "$plan" ] || { echo "no such plan file: $plan" >&2; exit 2; }
21
21
 
22
+ git rev-parse --git-dir >/dev/null 2>&1 || {
23
+ echo "error: not a git repository — review-package needs the BASE and HEAD commits of the task under review." >&2
24
+ echo "A greenfield plan's first task creates the repo itself: run this from inside the repo once that task has committed." >&2
25
+ exit 2
26
+ }
27
+
22
28
  git rev-parse --verify --quiet "$base" >/dev/null || { echo "bad BASE: $base" >&2; exit 2; }
23
29
  git rev-parse --verify --quiet "$head" >/dev/null || { echo "bad HEAD: $head" >&2; exit 2; }
24
30
 
@@ -19,6 +19,13 @@ if (-not (Test-Path -LiteralPath $plan -PathType Leaf)) {
19
19
  exit 2
20
20
  }
21
21
 
22
+ & git rev-parse --git-dir *> $null
23
+ if ($LASTEXITCODE -ne 0) {
24
+ [Console]::Error.WriteLine("error: not a git repository — review-package needs the BASE and HEAD commits of the task under review.")
25
+ [Console]::Error.WriteLine("A greenfield plan's first task creates the repo itself: run this from inside the repo once that task has committed.")
26
+ exit 2
27
+ }
28
+
22
29
  & git rev-parse --verify --quiet $base *> $null
23
30
  if ($LASTEXITCODE -ne 0) {
24
31
  [Console]::Error.WriteLine("bad BASE: $base")
@@ -23,6 +23,10 @@
23
23
  # Single source of truth for the workspace location, so task-brief and
24
24
  # review-package cannot drift to different directories.
25
25
  #
26
+ # A greenfield plan's first task is often "create the repo", so there is no
27
+ # repo root to resolve yet: fall back to the current directory rather than
28
+ # failing, since that task is exactly the one needing a brief.
29
+ #
26
30
  # Usage: sdd-workspace PLAN_FILE
27
31
  set -euo pipefail
28
32
 
@@ -38,7 +42,10 @@ slug=$(basename -- "$plan" .md)
38
42
  [ -n "$slug" ] && [ "$slug" != "." ] && [ "$slug" != ".." ] \
39
43
  || { echo "cannot derive a workspace name from: $plan" >&2; exit 2; }
40
44
 
41
- root=$(git rev-parse --show-toplevel)
45
+ if ! root=$(git rev-parse --show-toplevel 2>/dev/null); then
46
+ root=$PWD
47
+ echo "sdd-workspace: no git repository yet, using $root/.superpowers/sdd until this plan's first task creates one" >&2
48
+ fi
42
49
  root=$(CDPATH= cd -- "$root" && pwd -P)
43
50
  base="$root/.superpowers/sdd"
44
51
 
@@ -7,6 +7,10 @@
7
7
  # plan in the same working tree can never read or overwrite another plan's
8
8
  # artifacts.
9
9
  #
10
+ # A greenfield plan's first task is often "create the repo", so there is no
11
+ # repo root to resolve yet: fall back to the current directory rather than
12
+ # failing, since that task is exactly the one needing a brief.
13
+ #
10
14
  # Usage: ./sdd-workspace.ps1 PLAN_FILE
11
15
 
12
16
  $ErrorActionPreference = "Stop"
@@ -28,8 +32,21 @@ if ([string]::IsNullOrEmpty($slug) -or $slug -eq "." -or $slug -eq "..") {
28
32
  exit 2
29
33
  }
30
34
 
31
- $root = (& git rev-parse --show-toplevel).Trim()
32
- $base = Join-Path $root ".superpowers/sdd"
35
+ $root = $null
36
+ try {
37
+ # Collect all output without an early-terminating pipeline: a
38
+ # Select-Object -First teardown can leave $LASTEXITCODE stale.
39
+ $rootOutput = @(& git rev-parse --show-toplevel 2>$null)
40
+ if ($LASTEXITCODE -eq 0 -and $rootOutput.Count -gt 0 -and -not [string]::IsNullOrWhiteSpace($rootOutput[0])) {
41
+ $root = ([string]$rootOutput[0]).Trim()
42
+ }
43
+ } catch {
44
+ $root = $null
45
+ }
46
+ if ([string]::IsNullOrWhiteSpace($root)) {
47
+ $root = (Get-Location).Path
48
+ [Console]::Error.WriteLine("sdd-workspace.ps1: no git repository yet, using $root/.superpowers/sdd until this plan's first task creates one")
49
+ }
33
50
 
34
51
  function Get-PhysicalDirectoryPath($path) {
35
52
  if (-not (Test-Path -LiteralPath $path)) { return $path }
@@ -59,6 +76,10 @@ if ([string]::IsNullOrEmpty($planParent)) { $planParent = "." }
59
76
 
60
77
  $planDir = Get-PhysicalDirectoryPath $planParent
61
78
  $rootPhys = Get-PhysicalDirectoryPath $root
79
+ # Build $base from the physical root so the greenfield fallback (cwd, which
80
+ # may be a symlinked path such as macOS /var) prints the same location as
81
+ # the later git-resolved call.
82
+ $base = Join-Path $rootPhys ".superpowers/sdd"
62
83
 
63
84
  $planAbs = (Join-Path $planDir $planLeaf) -replace '\\', '/'
64
85
  $rootNorm = $rootPhys -replace '\\', '/'
@@ -134,13 +134,27 @@ Subagent (general-purpose):
134
134
 
135
135
  Your report should point at evidence: file:line references for every
136
136
  finding and for any check you would otherwise answer with a bare
137
- "yes." A tight report that cites lines gives the controller everything
138
- it needs.
139
-
140
- Your final message is the report itself: begin directly with the
141
- spec-compliance verdict. Every line is a verdict, a finding with
142
- file:line, or a check you ran — no preamble, no process narration,
143
- no closing summary.
137
+ "yes." A tight report that cites lines gives the fix subagent
138
+ everything it needs.
139
+
140
+ ## Report Format
141
+
142
+ Write your full report to [REVIEW_FILE] using the Output Format
143
+ below: verdicts, strengths, every finding with file:line, and any
144
+ check you ran outside the diff. Fix subagents read this file — it is
145
+ the only place your detail survives.
146
+
147
+ Then report back with ONLY (under 15 lines — the detail lives in the
148
+ review file):
149
+ - **Spec:** ✅ | ❌ — if ❌, one line per missing/extra/misunderstood
150
+ item (file:line + summary)
151
+ - **⚠️ Cannot verify from diff:** [items the controller must resolve
152
+ itself — always in this message, never only in the file]
153
+ - **Quality:** Approved | Needs fixes
154
+ - One line per Critical/Important finding: file:line + what's wrong,
155
+ with any plan-mandated finding marked (the human adjudicates those)
156
+ - Minor findings: count only
157
+ - The review file path
144
158
 
145
159
  ## Calibration
146
160
 
@@ -158,7 +172,7 @@ Subagent (general-purpose):
158
172
  Acknowledge what was done well before listing issues — accurate praise
159
173
  helps the implementer trust the rest of the feedback.
160
174
 
161
- ## Output Format
175
+ ## Output Format (the review file)
162
176
 
163
177
  ### Spec Compliance
164
178
 
@@ -202,6 +216,10 @@ Subagent (general-purpose):
202
216
  - `[DIFF_FILE]` — REQUIRED: the path the controller wrote the review
203
217
  package to (`scripts/review-package PLAN_FILE BASE HEAD`, or `scripts/review-package.ps1 PLAN_FILE BASE HEAD` on Windows PowerShell, prints the unique
204
218
  path it wrote; the package never enters the controller's context)
219
+ - `[REVIEW_FILE]` — REQUIRED: the file the reviewer writes its full report
220
+ to; name it after the brief (brief `…/task-N-brief.md` → review
221
+ `…/task-N-review.md`). Re-reviews append to the same file.
205
222
 
206
- **Reviewer returns:** Spec Compliance verdict (✅/❌/⚠️), Strengths, Issues
207
- (Critical/Important/Minor), Task quality verdict
223
+ **Reviewer returns:** full report in `[REVIEW_FILE]`; final message under
224
+ 15 lines — Spec verdict (✅/❌/⚠️), Quality verdict, Critical/Important
225
+ findings one line each, Minor count, review file path
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: systematic-debugging
3
- description: Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes
3
+ description: Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes - including when you say "systematic debug", "debug this properly", or "find the root cause first"; for driving new behavior test-first, use test-driven-development instead
4
4
  ---
5
5
 
6
6
  # Systematic Debugging
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: test-driven-development
3
- description: Use when implementing any feature or bugfix, before writing implementation code
3
+ description: Use when implementing any feature or bugfix, before writing implementation code - including when you say "tdd", "tdd workflow", "write the test first", or "red green refactor"; for diagnosing an existing defect, use systematic-debugging instead
4
4
  ---
5
5
 
6
6
  # Test-Driven Development (TDD)
@@ -68,6 +68,27 @@ digraph tdd_cycle {
68
68
  }
69
69
  ```
70
70
 
71
+ ### Characterization Test for a Behavior-Preserving Refactor
72
+
73
+ This extends the [upstream characterization guidance](writing-good-tests.md#principle-1-name-the-break)
74
+ to behavior-preserving refactors of your own code.
75
+
76
+ Normal RED applies when behavior should change. If observable behavior must
77
+ remain unchanged, establish a characterization guard before refactoring:
78
+
79
+ 1. Name the behavior and a relevant production mutation that should make the
80
+ test fail.
81
+ 2. Write the test and observe the existing behavior pass.
82
+ 3. Require `git diff HEAD -- <production-paths>` to be empty before mutating.
83
+ If those paths have existing changes, stop without discarding them.
84
+ Make the mutation and verify the expected failure. If the test still
85
+ passes, strengthen or replace it and repeat.
86
+ 4. Restore only the mutated production paths from VCS with
87
+ `git restore --source=HEAD --worktree -- <production-paths>`. Require
88
+ `git diff --exit-code HEAD -- <production-paths>` to succeed with an empty
89
+ diff, then verify green with the characterization test retained.
90
+ 5. Refactor while staying green.
91
+
71
92
  ### RED - Write Failing Test
72
93
 
73
94
  Write one minimal test showing what should happen.
@@ -123,7 +144,10 @@ Confirm:
123
144
  - Failure message is expected
124
145
  - Fails because feature missing (not typos)
125
146
 
126
- **Test passes?** You're testing existing behavior. Fix test.
147
+ **Test passes?** If observable behavior must remain unchanged, follow the
148
+ [characterization guard](#characterization-test-for-a-behavior-preserving-refactor)
149
+ before refactoring. If behavior should change, you're testing existing
150
+ behavior. Fix test.
127
151
 
128
152
  **Test errors?** Fix error, re-run until it fails correctly.
129
153
 
@@ -239,7 +263,7 @@ When writing or changing any test, read [writing-good-tests.md](writing-good-tes
239
263
 
240
264
  - Code before test
241
265
  - Test after implementation
242
- - Test passes immediately
266
+ - Test for new or changed behavior passes immediately
243
267
  - Can't explain why test failed
244
268
  - Tests added "later"
245
269
  - Rationalizing "just this once"
@@ -62,6 +62,10 @@ getters, constants, and trivial forwarding earn tests only when they
62
62
  validate, normalize, default, derive, enforce, or cause side effects —
63
63
  otherwise assert the first consumer-visible result that depends on them.
64
64
 
65
+ For a behavior-preserving refactor of your own code, establish the
66
+ [characterization guard](SKILL.md#characterization-test-for-a-behavior-preserving-refactor)
67
+ before changing production structure.
68
+
65
69
  ### Gate Function
66
70
 
67
71
  ```
@@ -156,6 +160,9 @@ process costs maintenance forever.
156
160
 
157
161
  ## The Mutation Check
158
162
 
163
+ For a pre-refactor characterization guard, run the actual mutation and
164
+ VCS restoration procedure in the TDD skill, including its empty-diff check.
165
+
159
166
  Before finishing, mentally mutate the production code; at least one test
160
167
  should fail for each realistic mutation:
161
168