task-pipeline-skill 0.10.0 → 0.17.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (30) hide show
  1. package/CHANGELOG.md +401 -0
  2. package/LICENSE +85 -0
  3. package/README.md +211 -84
  4. package/cursor/rules/task-pipeline.mdc +135 -20
  5. package/package.json +3 -3
  6. package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
  7. package/plugins/task-pipeline/commands/task-pipeline.md +16 -8
  8. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +139 -57
  9. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +41 -27
  10. package/plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json +1 -1
  11. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
  12. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +32 -6
  13. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
  14. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
  15. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +73 -21
  16. package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
  17. package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
  18. package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +169 -0
  19. package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
  20. package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +55 -22
  21. package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
  22. package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
  23. package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
  24. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +184 -48
  25. package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
  26. package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +11 -3
  27. package/plugins/task-pipeline/skills/task-pipeline/templates/adr.md +64 -0
  28. package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +49 -0
  29. package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
  30. package/plugins/task-pipeline/skills/task-pipeline/templates/context.md +87 -0
@@ -3,12 +3,19 @@
3
3
  For each stage: what it does, what to invoke, artifacts, and the **GATE** that
4
4
  must pass before advancing. Each gate is tagged with its **type** — `auto` (the
5
5
  orchestrator verifies the check itself, pass/fail) or `manual` (wait for the
6
- operator's explicit go). These stages (0 intake + 1→9) are the plugin's
6
+ operator's explicit go). These stages (0 intake + 1→10) are the plugin's
7
7
  **example** flow, encoded in `pipeline.example.json` against the universal contract
8
8
  `pipeline.schema.json`; a host project replaces it with its own
9
9
  stages/agents/types (see SKILL.md → *Bring your own skills*).
10
10
 
11
- ## 0 — Intake grill (Fable)
11
+ ## 0 — Intake grill — MANDATORY
12
+ - **Stage 0 is not optional and not skippable.** There is no "small enough task"
13
+ exemption, no "the request was already clear" exemption, no starting stage 1
14
+ "while the operator thinks". The only sanctioned bypass is the
15
+ entry-from-super-ux short-circuit below, and even that still requires a scope
16
+ confirmation and a record of what was adopted vs skipped. A run that reaches
17
+ stage 1 without a committed, operator-confirmed brief is a **failed run** —
18
+ stop and go back.
12
19
  - **Entry-from-super-ux short-circuit (check FIRST).** task-pipeline is often
13
20
  launched *from* super-ux — its `/ux` action menu offers "execute autonomously
14
21
  via the task-pipeline plugin" once the UX chain (and often a
@@ -25,36 +32,40 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
25
32
  - **What (normal entry):** the operator's one-line task is almost never enough to
26
33
  run autonomously. Before anything else, **grill the operator** to expand that
27
34
  one line into a complete, unambiguous brief — resolve every decision branch
28
- up front so stages 1→9 need no further human input beyond the manual gates.
35
+ up front so stages 1→10 need no further human input beyond the manual gates.
29
36
  This is input expansion, not design: turn "make me feature X" into locked
30
37
  answers for scope, users, constraints, data, edge cases, done-criteria.
31
- - **Invoke:** prefer the `grill-me` / `grilling` skill if it resolves; otherwise
32
- run the built-in grill loop:
33
- 1. **One question per turn** — never bundle.
34
- 2. **Give a recommended answer with every question** (+ 1-line rationale);
35
- "what do you think?" is lazy.
36
- 3. **Explore the codebase/docs before asking** — if `grep`/`Read`/context7
37
- answers it, do that instead of spending a turn.
38
- 4. **Walk the decision tree depth-first**; finish a branch before opening
39
- another; ask prerequisite decisions first.
40
- 5. **Reconcile contradictions** immediately; chase dodges ("we'll decide
41
- later" → "what's the latest you can decide and still ship?").
38
+ - **How it runs: [`grill.md`](grill.md)** — the full doctrine, built into this
39
+ skill (nothing to install). In short: one question per turn, a recommended
40
+ answer with each, explore the codebase before asking, depth-first through the
41
+ decision tree, contradictions reconciled on the spot; plus **domain awareness**
42
+ (challenge terms against `CONTEXT.md`, sharpen fuzzy language, stress-test with
43
+ concrete scenarios, cross-reference the code, record ADRs for hard-to-reverse
44
+ calls) and the **autonomy sweep** that pre-resolves every stage-1→10 blocker.
45
+ Deploy authorization has a hard floor there: a standing go counts only when it
46
+ names the target and the preconditions.
42
47
  - **UI early-detect:** one branch of the grill is always "does this touch a
43
48
  user-facing surface (web/mobile/CLI/TUI)?". If yes → surface **super-ux**
44
49
  now (use it if installed; otherwise give the install line — see SKILL.md
45
50
  *Prerequisites*); this arms the stage-3 UX track.
46
51
  - **Artifact:** lock the resolved decisions into a **task brief** committed at
47
52
  `docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` (scope, users/UI verdict,
48
- constraints, assumptions, explicitly-deferred items, done-criteria). Seed it from
53
+ constraints, assumptions, explicitly-deferred items, done-criteria) **plus the
54
+ autonomy sweep's per-stage answers and the model decision**. Seed it from
49
55
  the skill's `templates/brief.md` skeleton — but only when absent, never
50
- overwrite an existing brief. Stages 2–4 build on this brief.
56
+ overwrite an existing brief. Stages 2–4 build on this brief; stages 5–10 read
57
+ its autonomy section instead of asking. Where the session produced them, also:
58
+ an updated `CONTEXT.md` (terms written as they resolved) and any ADRs under
59
+ `docs/adr/` — see `grill.md` → *Domain awareness*.
51
60
  - **GATE (manual):** shared understanding reached — every detected branch has a
52
- recorded answer or an explicit deferral, no open contradictions, and the
53
- operator confirms the brief. Stop when a re-scan surfaces no new branches
54
- (don't grill past diminishing returns; reversible calls can be deferred with a
55
- note). Only then start stage 1.
61
+ recorded answer or an explicit deferral, no open contradictions, **every
62
+ autonomy-sweep row is answered or explicitly marked "stop and ask here"**, the
63
+ **REQ table is written and every row names its check**, the carry-over ledger is
64
+ seeded, the model decision is recorded, and the operator confirms the brief. Stop when a
65
+ re-scan surfaces no new branches (don't grill past diminishing returns;
66
+ reversible calls can be deferred with a note). Only then start stage 1.
56
67
 
57
- ## 1 — Docs study (Fable)
68
+ ## 1 — Docs study
58
69
  - **What:** ground every external library / API / SDK the task touches on the
59
70
  *current* docs, before locking any contract.
60
71
  - **Invoke:** `context7` MCP (`resolve-library-id` → `get-library-docs`, scope by
@@ -63,16 +74,40 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
63
74
  - **GATE (auto):** every contract the design will lock is grounded in fetched docs,
64
75
  not recall. Unresolvable libraries are flagged in the spec.
65
76
 
66
- ## 2 — Brainstorm (Fable)
67
- - **Invoke:** `superpowers:brainstorming`. One question at a time; 2–3 approaches +
68
- a recommendation; design presented in sections.
77
+ ## 2 — Brainstorm + decompose
78
+ - **How it runs: [`brainstorm.md`](brainstorm.md)** — built into this skill. Read
79
+ the brief first (stage 0 already answered scope/constraints/done-criteria), then
80
+ explore the codebase, scope-check for decomposition, one question at a time, 2–3
81
+ approaches with a recommendation, design presented in sections and approved
82
+ section by section. **Hard gate:** no code, no scaffolding, no implementation
83
+ skill before the operator approves the design — including on "obviously simple"
84
+ tasks.
69
85
  - **UI detection (mandatory check):** decide whether the task touches any
70
86
  user-facing surface (web, mobile, CLI, TUI — new feature, new screen/command,
71
87
  or a change to user-visible behavior). Record the verdict; it arms the UX
72
88
  track in stage 3.
73
- - **GATE (manual):** the user approves the design **and** the UI verdict is recorded.
89
+ - **Decomposition (platforms only): [`decomposition.md`](decomposition.md).** If
90
+ the brief describes a platform rather than a change — several independent
91
+ capabilities, several surfaces that could ship separately, REQs no single
92
+ deliverable satisfies — cut it into **modules** before any spec is written, and
93
+ commit the module map (`specs/<topic>-modules.md`): what each module delivers,
94
+ the entities it owns, what it depends on, the contracts it exposes, its REQs and
95
+ its status, in build order with the walking skeleton first. Single-module work
96
+ records `single module: <name>` in the design and moves on — a skipped
97
+ decomposition is a decision, never an omission.
98
+ - **GATE (manual):** the user approves the design, the UI verdict is recorded,
99
+ **every REQ is answered by the design** — a requirement the design doesn't
100
+ address is either covered now or explicitly dropped by the operator, with the
101
+ drop recorded in the carry-over ledger — **and, for a platform, the module map is
102
+ approved**: brick criteria met or excepted in writing, dependency graph acyclic,
103
+ build order topological, every REQ mapped to exactly one module, cross-module
104
+ contracts named with their owner.
74
105
 
75
- ## 3 — Spec (Fable) — with UX track for user-facing tasks
106
+ ## 3 — Spec — with UX track for user-facing tasks
107
+ - **How it runs: [`spec.md`](spec.md)** — built into this skill: the UX-track order,
108
+ what the spec must lock (types, schemas, signatures, file layout, the **Global
109
+ Constraints** block stages 4–5 depend on), the self-review pass and the operator
110
+ review gate.
76
111
  - **UX track (runs FIRST when stage 2 flagged UI; skip entirely otherwise).**
77
112
  Requires the **super-ux** skills. If missing on a UI task → give the install
78
113
  line and stop (see SKILL.md *Prerequisites*: `/plugin marketplace add
@@ -97,49 +132,72 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
97
132
  never rebuild from scratch. If the chain already exists and is validated (e.g.
98
133
  the task entered from super-ux), just verify (linter green) and embed it into
99
134
  the spec; only build the parts that are missing.
100
- - **Spec:** brainstorming writes the design to
101
- `docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md` and commits it. Lock all
135
+ - **Spec:** write the approved design to
136
+ `docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md` and commit it. Lock all
102
137
  shared contracts (types, schemas, signatures, file layout). For UI tasks the
103
138
  spec **embeds the UX layer**: links the validated scenario IDs, the flows and
104
139
  `SCR-` screens, the CJM stages the feature serves, and the UX
105
140
  patterns/principles from super-ux that apply (`best-practices.md`,
106
141
  `ux-design-principles.md`, `component-guidelines.md`).
107
- - **GATE (manual):** spec committed **and** user-reviewed; for UI tasks
142
+ - **GATE (manual):** spec committed **and** user-reviewed; **every section carries
143
+ `covers: REQ-…` and every REQ appears in at least one section**; for UI tasks
108
144
  additionally: the super-ux chain (foundation → flows → screens → scenarios) is
109
145
  designed, validated and approved; scenarios validated in `docs/ux/scenarios.md`;
110
146
  the linter passes; every user-facing spec requirement traces to a scenario ID
111
147
  (or an explicit v1-mode/tiny-project waiver by the operator). No plan (stage 4)
112
148
  starts before this — the chain comes BEFORE interface.
113
149
 
114
- ## 4 — Plan (Fable)
115
- - **Invoke:** `superpowers:writing-plans` →
116
- `docs/superpowers/plans/YYYY-MM-DD-<feature>.md`. Zero-context tasks, exact
117
- paths, TDD steps, DoD each, dependency graph + parallel groups, non-overlapping
118
- file ownership.
119
- - **GATE (auto):** every spec requirement maps to a task; no placeholders; parallel-group
150
+ ## 4 — Plan
151
+ - **How it runs: [`planning.md`](planning.md)** — built into this skill →
152
+ `docs/superpowers/plans/YYYY-MM-DD-<topic>.md` (same slug as the brief and the
153
+ spec). Zero-context tasks, exact
154
+ paths, complete code in every step, TDD steps with expected output, DoD each,
155
+ dependency graph + parallel groups, non-overlapping file ownership, and the
156
+ Global Constraints block copied verbatim from the spec.
157
+ - **GATE (auto):** **set equality — the REQ ids in the brief equal the union of
158
+ `Implements:` across plan tasks.** A non-empty difference fails the gate and is
159
+ reported as the explicit list of dropped requirements; this is the seam where
160
+ scope leaks silently, so the check is mechanical, not a judgement call. Plus:
161
+ every spec requirement maps to a task; no placeholders; names and
162
+ types consistent across tasks; every task carries a verifiable DoD; parallel-group
120
163
  tasks share no files. For UI tasks: every task building user-facing behavior
121
164
  names the scenario ID(s) and `SCR-` screen(s) it implements, and its DoD
122
165
  includes satisfying them **and** updating the affected super-ux layers in the
123
166
  same change (super-ux *same-change* rule).
124
167
 
125
- ## 5 — Dev (Opus)
126
- - **Invoke:** `superpowers:using-git-worktrees` (isolate) →
127
- `superpowers:subagent-driven-development` (or `superpowers:executing-plans`).
128
- TDD per task (failing test → minimal impl → green → commit). Pin subagents to Opus.
129
- - **GATE (auto):** all plan tasks DONE (two-stage review: spec compliance, then code
130
- quality); full test suite green.
168
+ ## 5 — Dev
169
+ - **How it runs: [`build.md`](build.md)** — built into this skill: isolate the
170
+ workspace (native worktree tool first, git fallback, baseline tests), keep a
171
+ ledger under `.task-pipeline/build/<plan>/` so a compacted context can resume,
172
+ then one fresh implementer subagent per task with a file-based brief and report,
173
+ a review after every task ([`review.md`](review.md)), and a five-round fix loop
174
+ with an explicit breaker. TDD per task ([`tdd.md`](tdd.md)): failing test →
175
+ watch it fail → minimal impl → watch it pass → commit. Pin subagents to the
176
+ run's confirmed model (`model-tiering.md`). The plan's parallel groups fan out
177
+ **only** when each implementer gets its own worktree; otherwise sequential.
178
+ - **Integration closes the stage:** sync with the base branch, re-run the full suite
179
+ on the result, land it the project's way (merge, or a PR — outward, so it needs a
180
+ go), remove the worktree. Stages 7–9 act on the integrated result, so a branch the
181
+ operator chose to leave unmerged is recorded as such.
182
+ - **GATE (auto):** all plan tasks DONE (three review verdicts per task: spec
183
+ compliance, **REQ satisfied**, code quality); every finding fixed or parked with a
184
+ ruling; **every parked finding and implementer concern harvested into the
185
+ carry-over ledger** — nothing stays only in the scratch workspace, which is
186
+ deleted; no task left BLOCKED; full test suite green; branch integrated per the brief's policy (or the operator's
187
+ "leave it" recorded).
131
188
 
132
- ## 6 — Tests (Opus)
189
+ ## 6 — Tests
133
190
  - **What:** consolidate test coverage for the change: confirm new functionality
134
191
  has tests (written test-first in stage 5), update/repair existing tests the
135
192
  change touched, and add edge-case + failure-path tests per DoD.
136
- - **Invoke:** the host test runner (see `conventions.md` → *Lint + test*);
137
- `superpowers:test-driven-development` for any uncovered gap.
193
+ - **Invoke:** the host test runner (see `conventions.md` → *Lint + test*); the
194
+ built-in [`tdd.md`](tdd.md) cycle for any uncovered gap — failing test first,
195
+ same as stage 5.
138
196
  - **GATE (auto):** the **full** suite is green (not just the new tests); new/changed code
139
197
  is covered; no `skip`/`xfail` smuggling a red suite past the gate. Never advance
140
198
  to deploy on a red or partial run.
141
199
 
142
- ## 7 — Lint + deploy (host model)
200
+ ## 7 — Lint + deploy
143
201
  - Read host conventions (`conventions.md`): run the linter; fix failures. The suite
144
202
  is already green from stage 6 — re-run it if code changed since. For UI projects,
145
203
  the **super-ux linter** (`python3 docs/ux/lint.py` / `/ux-lint`) is part of lint —
@@ -147,19 +205,97 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
147
205
  if the project defines release automation (`pipeline.json` → `release`, toggle
148
206
  on), that is what "deploy" runs here.
149
207
  - **GATE (manual):** lint clean (host linter **and**, for UI projects, the super-ux
150
- linter) **and** suite green **before** deploy. Deploy is outward → explicit
208
+ linter) **and** suite green **before** deploy, **and no REQ is still `open`** — a
209
+ `partial` ships only with the operator's explicit acceptance. A gap is cheapest to
210
+ close before it ships, and the operator is already present at this gate. Deploy is outward → explicit
151
211
  operator go. Respect deploy-from-main rules if the project mandates them.
152
212
 
153
- ## 8 — Post-deploy (host model)
213
+ ## 8 — Post-deploy
154
214
  - Tail deploy logs / health-check per conventions. Confirm clean boot, no error
155
215
  spike, live subsystems healthy.
156
216
  - **GATE (auto):** clean boot confirmed, or an **honest degradation report** with next
157
217
  steps — never silent success.
158
218
 
159
- ## 9 — Docs + wiki (host model)
219
+ ## 9 — Docs + wiki
160
220
  - Update host module docs / runbooks per the project's self-update rules, in the
161
221
  **same change**. For UI tasks, confirm the super-ux layers were updated in this
162
222
  change and the linter is green (super-ux *same-change* + *no-drift* rules). Then
163
223
  sync knowledge to the wiki (`wiki-update` skill).
164
224
  - **GATE (auto):** docs in sync with code; UI: super-ux layers current + linter
165
225
  green; wiki synced; dangling links fixed.
226
+
227
+ ## 10 — Acceptance
228
+ - **What:** the closing stage — go back to the brief and account for **every**
229
+ requirement. Doctrine: [`acceptance.md`](acceptance.md). Every earlier gate asks
230
+ "is this artifact good?"; none asks "does this still contain everything that was
231
+ asked for?" The loss happens on the seams between stages, and this is where it
232
+ surfaces.
233
+ - **Runs last**, after docs and wiki — those are deliverables too, and a REQ may
234
+ name them.
235
+ - **How it runs:** built in. Read the brief's REQ table, the carry-over ledger in
236
+ full, the plan's task statuses, git log, the final suite output, stage-8 notes and
237
+ stage-9 doc changes (plus `docs/ux/scenarios.md` + `/ux-lint` for UI tasks). Write
238
+ `docs/superpowers/specs/YYYY-MM-DD-<topic>-acceptance.md` — one row per REQ,
239
+ status `verified` / `partial` / `deferred` / `dropped`, each with **evidence** (a
240
+ passing test name, `file:line`, a command and its output, or a scenario ID).
241
+ "Done" without evidence is not done: downgrade to `partial` and say so rather
242
+ than upgrading the claim.
243
+ - Then ask the operator the closing question out loud, list in hand: *here's what
244
+ you asked for, here's what shipped, here's what's deferred and where it lives —
245
+ what's missing?* Ask it even when the table is green; the operator holds context
246
+ the brief never captured, and this is the cheapest moment in the run to hear it.
247
+ - **GATE (manual):** every REQ has a status (none `unknown`); every `verified`
248
+ carries evidence; every `partial` names what's missing and where it's tracked;
249
+ every `deferred`/`dropped` has the operator's agreement and, for `deferred`, a
250
+ tracker entry; no carry-over row left `unresolved`; the operator answers the
251
+ closing question and signs off. Manual by design — an automated check can prove
252
+ the table is well-formed, only the person who asked can confirm it is what they
253
+ asked for.
254
+
255
+ ## The program loop — a platform, one brick at a time
256
+
257
+ When stage 2 produced a **module map** ([`decomposition.md`](decomposition.md)),
258
+ stages 0–2 have run once for the whole platform and the rest of the pipeline runs
259
+ **per module**, in build order:
260
+
261
+ ```
262
+ module N → 3 spec (dossier) → 4 plan → 5 build → 6 tests → 7 lint+deploy
263
+ → 8 post-deploy → 9 docs+wiki → 10 acceptance → map status: done
264
+ → module N+1 (back to 3)
265
+ ```
266
+
267
+ - **No re-grilling, no re-decomposing per module.** New information that changes the
268
+ map goes back to stage 2 as an explicit, operator-approved map revision — never a
269
+ quiet edit mid-module.
270
+ - **Each module's spec is a full dossier** (`spec.md`): architecture, entities,
271
+ contracts in and out, business rules, edge and failure cases, UI/Figma chain when
272
+ it has a surface.
273
+ - **Deploy cadence is the brief's call** (autonomy sweep): per module, or once after
274
+ several. Decide it up front, not per module.
275
+ - **Update the module map's status in the same commit as that module's acceptance.**
276
+ The map is the resume point after a lost context.
277
+ - **Program done** when every row is `done` or `deferred` with an agreed home, the
278
+ cross-module contracts are covered by tests that cross the seam, and a final
279
+ acceptance covers the platform's whole REQ table — not module by module.
280
+
281
+ ## Cross-cutting — the loop guard
282
+
283
+ Any stage can be re-entered and any loop can churn: a pass undoing what an earlier
284
+ pass decided, two shapes alternating, the same file rewritten with no new
285
+ information. [`loop-guard.md`](loop-guard.md) is the detector and the break
286
+ protocol, and it binds every repeating loop here — the stage-5 fix loop, a stage
287
+ re-entered after a failed gate, the program loop above, any audit → fix → audit
288
+ cycle.
289
+
290
+ - **Every repeating pass logs one line per touched file** (`touch: <file> — pass N —
291
+ reason: <finding id / gate item>`) to the run ledger. Detection is mechanical, not
292
+ a feeling, and the ledger is what survives compaction.
293
+ - **Trips on:** revert-oscillation (A→B→A); the same file edited twice for the same
294
+ reason; a finding already ADDRESSED or parked coming back; a stage entered a third
295
+ time for one artifact; two loops editing one file. Hard caps: 5 fix rounds per
296
+ task, 2 re-entries per stage per artifact, 3 passes per module.
297
+ - **On a trip: stop editing.** Name shapes A and B with their evidence, escalate to
298
+ the layer that owns the conflict (rubric → operator → plan → spec → module map),
299
+ re-plan the check as an ordered one-item-per-line checklist, then go through it in
300
+ order, one commit per item. Never settle a higher-layer conflict inside a lower
301
+ loop, and never adjudicate before the cap.
@@ -0,0 +1,110 @@
1
+ # TDD — stages 5 and 6, built in
2
+
3
+ How every task in the build is implemented, and what stage 6 consolidates. Built
4
+ into this skill; nothing to install.
5
+
6
+ > Ported from the `test-driven-development` skill in
7
+ > [obra/superpowers](https://github.com/obra/superpowers) (MIT — see `LICENSE` →
8
+ > *Third-party*), with the stage-6 suite gate added.
9
+
10
+ ## The iron law
11
+
12
+ ```
13
+ NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST
14
+ ```
15
+
16
+ **If you didn't watch the test fail, you don't know it tests the right thing.**
17
+
18
+ Wrote code before the test? Delete it and start from the test. Not "keep it as
19
+ reference", not "adapt it while writing tests", not "look at it once more". Delete
20
+ means delete — code you kept is code the test was written to fit.
21
+
22
+ **Always:** new features, bug fixes, refactors, behavior changes.
23
+ **Exceptions, and only with the operator's say-so:** throwaway prototypes,
24
+ generated code, pure configuration.
25
+
26
+ Thinking "skip TDD just this once"? That thought is the rationalization, not the
27
+ exception.
28
+
29
+ ## Red → green → refactor
30
+
31
+ **RED — write one failing test.** One behavior, a name that describes that
32
+ behavior, real code rather than mocks wherever mocks are avoidable.
33
+
34
+ **Verify RED — run it and watch it fail. Mandatory.** Confirm it *fails* rather
35
+ than *errors*, that the message is the one you expected, and that it fails because
36
+ the feature is missing — not because of a typo or a bad import. A test that passes
37
+ immediately is testing behavior that already exists: fix the test.
38
+
39
+ **GREEN — the simplest code that passes.** No extra options, no "while I'm here"
40
+ refactor of the neighbors, no configuration surface nobody asked for. YAGNI.
41
+
42
+ **Verify GREEN — run it and watch it pass. Mandatory.** The new test passes, the
43
+ other tests still pass, and the output is pristine — no stray warnings or errors.
44
+ Test still fails → fix the code, never the test. Another test broke → fix it now.
45
+
46
+ **REFACTOR — only once green.** Remove duplication, improve names, extract helpers.
47
+ Tests stay green. No new behavior enters here.
48
+
49
+ Then the next failing test.
50
+
51
+ ## Tests that stay honest
52
+
53
+ - **Before writing a test, name the production change that would make it fail.**
54
+ Can't name one? The test asserts nothing useful.
55
+ - **Assert on real behavior, never on mock behavior.** `expect(mock).toHaveBeenCalled()`
56
+ proves the mock works. Understand a dependency's side effects before mocking it —
57
+ a mock that lies is worse than no test.
58
+ - **One thing per test.** An "and" in the name means two tests.
59
+ - **Test-only helpers live in test utilities**, never as extra branches or flags in
60
+ production classes.
61
+ - **Edge cases and failure paths are part of the task**, not a follow-up ticket:
62
+ empty input, boundary values, the network call that fails, the timeout.
63
+
64
+ ## Stage 6 — consolidation and the suite gate
65
+
66
+ Stage 5 wrote the tests task by task. Stage 6 makes the whole thing true:
67
+
68
+ - New functionality has tests (written test-first in stage 5) — fill any gap now,
69
+ the same way: failing test first.
70
+ - Tests the change touched are updated or repaired, not deleted around.
71
+ - Edge-case and failure-path coverage matches each task's DoD.
72
+ - The test command is the one recorded in the brief's autonomy sweep; "green" means
73
+ what the brief says it means (including a known-red baseline, if one was
74
+ recorded).
75
+
76
+ **GATE (auto):** the **full** suite is green — not just the new tests. New and
77
+ changed code is covered. No `skip` / `xfail` / commented-out assertion smuggles a
78
+ red suite past the gate. A partial or red run never advances to deploy; report it
79
+ honestly instead.
80
+
81
+ ## When stuck
82
+
83
+ | Problem | What it means |
84
+ |---|---|
85
+ | Don't know how to test it | Write the API you wish existed, then the assertion. Still stuck → ask the operator. |
86
+ | The test is too complicated | The design is too complicated. Simplify the interface. |
87
+ | Everything has to be mocked | The code is too coupled. Inject dependencies. |
88
+ | Setup is enormous | Extract helpers; if it's still huge, the design is the problem. |
89
+ | Fixing a bug | Write the failing test that reproduces it first. The test proves the fix and prevents the regression. |
90
+
91
+ ## Rationalizations
92
+
93
+ | Excuse | Reality |
94
+ |---|---|
95
+ | "Too simple to test" | Simple code breaks. The test costs 30 seconds. |
96
+ | "I'll test after" | Tests written after pass immediately, which proves nothing. You never watched it fail, so you never proved it can catch the bug. |
97
+ | "Tests after achieve the same thing — spirit, not ritual" | Tests-after answer "what does this do?"; tests-first answer "what should this do?" After-the-fact tests are biased by the code that already exists. |
98
+ | "I already tested it manually" | Ad-hoc, unrepeatable, no record of what was covered. "Worked when I tried it" is not coverage. |
99
+ | "Deleting X hours of code is wasteful" | Sunk cost. The real choice is rewriting with TDD versus keeping code you can't trust. |
100
+ | "Just exploring first" | Fine — throw the exploration away and start with TDD. |
101
+ | "The existing code has no tests either" | You're improving it. Add tests for what you touch. |
102
+ | "TDD will slow me down" | TDD is the fast path: bugs caught before commit, refactors without fear. The shortcut ends in production debugging. |
103
+
104
+ ## Red flags — stop and start over
105
+
106
+ Code before test · test written after implementation · test passed on the first run
107
+ · can't explain why it failed · "tests later" · "just this once" · "keep it as
108
+ reference" · "it's about spirit not ritual" · "this case is different because…"
109
+
110
+ All of them mean the same thing: delete the code, start from the failing test.
@@ -1,12 +1,20 @@
1
1
  # templates
2
2
 
3
- Skeletons task-pipeline seeds into a host project. Only the **brief** is owned by
4
- this plugin (it is the stage-0 intake artifact); the spec and plan skeletons come
5
- from the `superpowers` skills, and the `docs/ux/*` skeletons from `super-ux`.
3
+ Skeletons task-pipeline seeds into a host project. Only the **brief** is a seeded
4
+ file (it is the stage-0 intake artifact). The spec and plan have no skeleton here —
5
+ their required structure is prescribed inline by `references/spec.md` and
6
+ `references/planning.md`; the `docs/ux/*` skeletons come from `super-ux`.
6
7
 
7
8
  | Template | Seeded to | Stage |
8
9
  |---|---|---|
9
10
  | `brief.md` | `docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` | 0 — intake grill |
11
+ | `carryover.md` | `docs/superpowers/specs/YYYY-MM-DD-<topic>-carryover.md` | 0 seeds, all stages append, 10 reads |
12
+ | `context.md` | `CONTEXT.md` at the repo root (or per context) | 0 — grill, domain awareness |
13
+ | `adr.md` | `docs/adr/NNNN-<slug>.md` | 0 — grill, hard-to-reverse decisions |
14
+
15
+ `context.md` and `adr.md` are **format references**, not files to copy wholesale:
16
+ the grill writes `CONTEXT.md` entries and ADRs in their shape, lazily — only once
17
+ there is a resolved term or a decision worth recording.
10
18
 
11
19
  Seeding rule (per the ssheleg canon): create a template copy **only when the
12
20
  target is absent**; never overwrite an existing brief.
@@ -0,0 +1,64 @@
1
+ # ADR — format
2
+
3
+ Architecture Decision Records live in `docs/adr/` with sequential numbering
4
+ (`0001-slug.md`, `0002-slug.md`, …). Scan for the highest existing number and
5
+ increment. Create the directory **lazily** — only when the first ADR is needed.
6
+
7
+ > Adapted from Matt Pocock's `grill-with-docs` (MIT — see the repo LICENSE →
8
+ > *Third-party*).
9
+
10
+ ## Template
11
+
12
+ ```md
13
+ # {Short title of the decision}
14
+
15
+ {1–3 sentences: what the context was, what was decided, and why.}
16
+ ```
17
+
18
+ That's it. An ADR can be a single paragraph. The value is recording *that* a
19
+ decision was made and *why* — not filling out sections.
20
+
21
+ ## Optional sections
22
+
23
+ Only when they add genuine value; most ADRs need none.
24
+
25
+ - **Status** frontmatter (`proposed | accepted | deprecated | superseded by
26
+ ADR-NNNN`) — useful once decisions start getting revisited.
27
+ - **Considered options** — only when the rejected alternatives are worth
28
+ remembering.
29
+ - **Consequences** — only when non-obvious downstream effects need calling out.
30
+
31
+ ## When to write one
32
+
33
+ All three must hold:
34
+
35
+ 1. **Hard to reverse** — changing your mind later carries real cost.
36
+ 2. **Surprising without context** — a future reader will look at the code and
37
+ wonder "why on earth did they do it this way?"
38
+ 3. **A real trade-off** — genuine alternatives existed and one was picked for
39
+ specific reasons.
40
+
41
+ Easy to reverse → skip it, you'll just reverse it. Not surprising → nobody will
42
+ wonder. No real alternative → there's nothing to record beyond "we did the obvious
43
+ thing."
44
+
45
+ ### What qualifies
46
+
47
+ - **Architectural shape.** "We're using a monorepo." "The write model is
48
+ event-sourced; the read model projects into Postgres."
49
+ - **Integration patterns between contexts.** "Ordering and Billing communicate via
50
+ domain events, not synchronous HTTP."
51
+ - **Technology choices carrying lock-in.** Database, message bus, auth provider,
52
+ deployment target — not every library, just the ones that would take a quarter to
53
+ swap out.
54
+ - **Boundary and scope decisions.** "Customer data is owned by the Customer
55
+ context; others reference it by ID only." The explicit no's are as valuable as
56
+ the yes's.
57
+ - **Deliberate deviations from the obvious path.** "Manual SQL instead of an ORM
58
+ because X." Anything a reasonable reader would assume the opposite of — this is
59
+ what stops the next engineer from "fixing" something deliberate.
60
+ - **Constraints invisible in the code.** "No AWS, for compliance." "Sub-200ms
61
+ responses, per the partner API contract."
62
+ - **Rejected alternatives whose rejection is non-obvious.** Considered GraphQL,
63
+ picked REST for subtle reasons → record it, or someone re-proposes GraphQL in six
64
+ months.
@@ -15,6 +15,28 @@
15
15
  - **Out of scope / explicitly deferred:** … (with the reason and, for deferrals,
16
16
  the latest moment the decision can still be made)
17
17
 
18
+ ## Requirements (the REQ spine — every later stage traces to these IDs)
19
+
20
+ Scope above is prose; this is the **addressable** form of it. One row per
21
+ independently verifiable deliverable — not one per sentence. Every row needs a
22
+ named check: **a requirement you can't say how to verify is a badly-stated
23
+ requirement** — split it here, on the grill, not at acceptance.
24
+
25
+ | ID | Requirement | How it's verified | Status |
26
+ |---|---|---|---|
27
+ | REQ-001 | … | test name / `file:line` / command + expected output / `SCN-…` | open |
28
+ | REQ-002 | … | … | open |
29
+
30
+ Status lifecycle, written at three checkpoints only (stage 4, stage 5, stage 10 —
31
+ not continuously): `open` → `planned` → `built` → `verified` \| `partial` \|
32
+ `deferred` \| `dropped`.
33
+
34
+ > **The list is frozen once confirmed.** Adding a requirement mid-run is fine —
35
+ > append it with its source. **Removing or narrowing one needs the operator's
36
+ > explicit agreement**, recorded in the carry-over ledger. Silently restating the
37
+ > task in smaller terms is the failure this table exists to prevent: every gate
38
+ > after it goes green on the shrunken task and nothing reports the loss.
39
+
18
40
  ## Users & context
19
41
 
20
42
  - **Who / for what:** … (personas, the job being done)
@@ -26,6 +48,33 @@
26
48
  |---|---|---|---|
27
49
  | 1 | … | … | … |
28
50
 
51
+ ## Autonomy (the sweep — stages 1→10 read this instead of asking)
52
+
53
+ Every row is either a resolved answer or an explicit **STOP AND ASK**. A blank row
54
+ is not neutral — it is a scheduled interruption.
55
+
56
+ | Stage | Question | Answer |
57
+ |---|---|---|
58
+ | run-wide | Model for this run | … (most capable available unless overridden; per-stage overrides here) |
59
+ | run-wide | Decide autonomously vs escalate to me | … |
60
+ | 1 Docs | External libs/APIs/SDKs in play; any context7 can't resolve → where their docs live | … |
61
+ | 2 Decompose | Platform (several capabilities/surfaces) or one module? If platform — deploy cadence: per module, or once at the end | … |
62
+ | 2–3 Spec | UI verdict (arms super-ux); scenario-tracing waiver, if any | … |
63
+ | 4–5 Dev | Base branch; worktree/branch policy; is `main` off-limits; commit convention; task tracker | … |
64
+ | 5 Integration | How the branch lands — direct merge, PR (who approves), or "leave it, I'll merge"; is parallel fan-out (one worktree per implementer) wanted? | … |
65
+ | 6 Tests | Test command; what "green" means; known-red baseline; coverage expectation | … |
66
+ | 7 Lint | Lint command (incl. `docs/ux/lint.py` for UI projects) | … |
67
+ | 7 Deploy | Target + path; release automation on/off; deploy-from-main rule | … |
68
+ | 7 Deploy | **Authorization** — standing go, or ask every time? | … |
69
+ | 8 Post-deploy | Where logs / health live (app name, endpoint, workflow) | … |
70
+ | 9 Docs+wiki | Which module docs / runbooks this change updates; wiki sync yes/no | … |
71
+ | 10 Acceptance | Who signs off; where deferred REQs get tracked (issue tracker / backlog) | … |
72
+
73
+ > **Deploy authorization has a hard floor.** A standing go counts only if it is
74
+ > **specific** — named target and named preconditions ("staging, once lint and the
75
+ > full suite are green; production always asks"). Vague blanket permission does not
76
+ > authorize an outward, irreversible action; stage 7 stops and asks.
77
+
29
78
  ## Done-criteria
30
79
 
31
80
  - Observable, verifiable conditions that mean "this task is finished".
@@ -0,0 +1,36 @@
1
+ # Carry-over ledger — <topic>
2
+
3
+ > **Append-only.** Any stage may add a row; nobody edits or deletes one. Committed
4
+ > to `docs/superpowers/specs/YYYY-MM-DD-<topic>-carryover.md` beside the brief, and
5
+ > read in full by stage 10 (acceptance).
6
+ >
7
+ > **The rule: deferred out loud is forgotten.** If it isn't written here, it wasn't
8
+ > deferred — it was lost. That covers everything said in passing: "we'll do that
9
+ > later", "good enough for now", a `DONE_WITH_CONCERNS` from an implementer, a
10
+ > reviewer's non-blocking finding, a requirement the operator agreed to drop.
11
+
12
+ | # | Stage | What | Why it isn't done | REQ | Where it lives now |
13
+ |---|---|---|---|---|---|
14
+ | 1 | 5 Dev | XLSX export path | scope call — CSV first | REQ-004 | LIN-483 |
15
+ | 2 | 5 Review | `export.ts` lacks a size guard | minor, non-blocking | — | backlog |
16
+ | 3 | 2 Brainstorm | REQ-007 dropped: bulk export | operator agreed 2026-07-28 | REQ-007 | dropped |
17
+
18
+ ## Columns
19
+
20
+ - **Stage** — where it surfaced, so acceptance knows how far it travelled.
21
+ - **What** — the concrete thing not done. "Error handling" is not an entry;
22
+ "`export.ts` swallows a failed write instead of surfacing it" is.
23
+ - **Why it isn't done** — scope call, blocked, deliberate deferral, out of budget.
24
+ "Forgot" is a legitimate and useful answer here.
25
+ - **REQ** — the requirement it belongs to, or `—` if it's outside the REQ spine.
26
+ - **Where it lives now** — issue id, backlog, `dropped` (with the operator's
27
+ agreement), or `unresolved`. **`unresolved` blocks the stage-10 gate**: an item
28
+ with no home is exactly the thing that gets forgotten, so acceptance refuses to
29
+ close on it.
30
+
31
+ ## Notes
32
+
33
+ - Adding a row costs one line and never blocks a stage — that is the point. The
34
+ ledger is cheap precisely so nobody is tempted to keep it in their head.
35
+ - Rows referencing a REQ feed that REQ's final status: an open carry-over row
36
+ means the requirement is at best `partial`, never `verified`.