task-pipeline-skill 0.10.0 → 0.17.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +401 -0
- package/LICENSE +85 -0
- package/README.md +211 -84
- package/cursor/rules/task-pipeline.mdc +135 -20
- package/package.json +3 -3
- package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
- package/plugins/task-pipeline/commands/task-pipeline.md +16 -8
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +139 -57
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +41 -27
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +32 -6
- package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +73 -21
- package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +169 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +55 -22
- package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +184 -48
- package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +11 -3
- package/plugins/task-pipeline/skills/task-pipeline/templates/adr.md +64 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +49 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/context.md +87 -0
|
@@ -3,12 +3,19 @@
|
|
|
3
3
|
For each stage: what it does, what to invoke, artifacts, and the **GATE** that
|
|
4
4
|
must pass before advancing. Each gate is tagged with its **type** — `auto` (the
|
|
5
5
|
orchestrator verifies the check itself, pass/fail) or `manual` (wait for the
|
|
6
|
-
operator's explicit go). These stages (0 intake + 1→
|
|
6
|
+
operator's explicit go). These stages (0 intake + 1→10) are the plugin's
|
|
7
7
|
**example** flow, encoded in `pipeline.example.json` against the universal contract
|
|
8
8
|
`pipeline.schema.json`; a host project replaces it with its own
|
|
9
9
|
stages/agents/types (see SKILL.md → *Bring your own skills*).
|
|
10
10
|
|
|
11
|
-
## 0 — Intake grill
|
|
11
|
+
## 0 — Intake grill — MANDATORY
|
|
12
|
+
- **Stage 0 is not optional and not skippable.** There is no "small enough task"
|
|
13
|
+
exemption, no "the request was already clear" exemption, no starting stage 1
|
|
14
|
+
"while the operator thinks". The only sanctioned bypass is the
|
|
15
|
+
entry-from-super-ux short-circuit below, and even that still requires a scope
|
|
16
|
+
confirmation and a record of what was adopted vs skipped. A run that reaches
|
|
17
|
+
stage 1 without a committed, operator-confirmed brief is a **failed run** —
|
|
18
|
+
stop and go back.
|
|
12
19
|
- **Entry-from-super-ux short-circuit (check FIRST).** task-pipeline is often
|
|
13
20
|
launched *from* super-ux — its `/ux` action menu offers "execute autonomously
|
|
14
21
|
via the task-pipeline plugin" once the UX chain (and often a
|
|
@@ -25,36 +32,40 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
|
|
|
25
32
|
- **What (normal entry):** the operator's one-line task is almost never enough to
|
|
26
33
|
run autonomously. Before anything else, **grill the operator** to expand that
|
|
27
34
|
one line into a complete, unambiguous brief — resolve every decision branch
|
|
28
|
-
up front so stages 1→
|
|
35
|
+
up front so stages 1→10 need no further human input beyond the manual gates.
|
|
29
36
|
This is input expansion, not design: turn "make me feature X" into locked
|
|
30
37
|
answers for scope, users, constraints, data, edge cases, done-criteria.
|
|
31
|
-
- **
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
5. **Reconcile contradictions** immediately; chase dodges ("we'll decide
|
|
41
|
-
later" → "what's the latest you can decide and still ship?").
|
|
38
|
+
- **How it runs: [`grill.md`](grill.md)** — the full doctrine, built into this
|
|
39
|
+
skill (nothing to install). In short: one question per turn, a recommended
|
|
40
|
+
answer with each, explore the codebase before asking, depth-first through the
|
|
41
|
+
decision tree, contradictions reconciled on the spot; plus **domain awareness**
|
|
42
|
+
(challenge terms against `CONTEXT.md`, sharpen fuzzy language, stress-test with
|
|
43
|
+
concrete scenarios, cross-reference the code, record ADRs for hard-to-reverse
|
|
44
|
+
calls) and the **autonomy sweep** that pre-resolves every stage-1→10 blocker.
|
|
45
|
+
Deploy authorization has a hard floor there: a standing go counts only when it
|
|
46
|
+
names the target and the preconditions.
|
|
42
47
|
- **UI early-detect:** one branch of the grill is always "does this touch a
|
|
43
48
|
user-facing surface (web/mobile/CLI/TUI)?". If yes → surface **super-ux**
|
|
44
49
|
now (use it if installed; otherwise give the install line — see SKILL.md
|
|
45
50
|
*Prerequisites*); this arms the stage-3 UX track.
|
|
46
51
|
- **Artifact:** lock the resolved decisions into a **task brief** committed at
|
|
47
52
|
`docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` (scope, users/UI verdict,
|
|
48
|
-
constraints, assumptions, explicitly-deferred items, done-criteria)
|
|
53
|
+
constraints, assumptions, explicitly-deferred items, done-criteria) **plus the
|
|
54
|
+
autonomy sweep's per-stage answers and the model decision**. Seed it from
|
|
49
55
|
the skill's `templates/brief.md` skeleton — but only when absent, never
|
|
50
|
-
overwrite an existing brief. Stages 2–4 build on this brief
|
|
56
|
+
overwrite an existing brief. Stages 2–4 build on this brief; stages 5–10 read
|
|
57
|
+
its autonomy section instead of asking. Where the session produced them, also:
|
|
58
|
+
an updated `CONTEXT.md` (terms written as they resolved) and any ADRs under
|
|
59
|
+
`docs/adr/` — see `grill.md` → *Domain awareness*.
|
|
51
60
|
- **GATE (manual):** shared understanding reached — every detected branch has a
|
|
52
|
-
recorded answer or an explicit deferral, no open contradictions,
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
61
|
+
recorded answer or an explicit deferral, no open contradictions, **every
|
|
62
|
+
autonomy-sweep row is answered or explicitly marked "stop and ask here"**, the
|
|
63
|
+
**REQ table is written and every row names its check**, the carry-over ledger is
|
|
64
|
+
seeded, the model decision is recorded, and the operator confirms the brief. Stop when a
|
|
65
|
+
re-scan surfaces no new branches (don't grill past diminishing returns;
|
|
66
|
+
reversible calls can be deferred with a note). Only then start stage 1.
|
|
56
67
|
|
|
57
|
-
## 1 — Docs study
|
|
68
|
+
## 1 — Docs study
|
|
58
69
|
- **What:** ground every external library / API / SDK the task touches on the
|
|
59
70
|
*current* docs, before locking any contract.
|
|
60
71
|
- **Invoke:** `context7` MCP (`resolve-library-id` → `get-library-docs`, scope by
|
|
@@ -63,16 +74,40 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
|
|
|
63
74
|
- **GATE (auto):** every contract the design will lock is grounded in fetched docs,
|
|
64
75
|
not recall. Unresolvable libraries are flagged in the spec.
|
|
65
76
|
|
|
66
|
-
## 2 — Brainstorm
|
|
67
|
-
- **
|
|
68
|
-
|
|
77
|
+
## 2 — Brainstorm + decompose
|
|
78
|
+
- **How it runs: [`brainstorm.md`](brainstorm.md)** — built into this skill. Read
|
|
79
|
+
the brief first (stage 0 already answered scope/constraints/done-criteria), then
|
|
80
|
+
explore the codebase, scope-check for decomposition, one question at a time, 2–3
|
|
81
|
+
approaches with a recommendation, design presented in sections and approved
|
|
82
|
+
section by section. **Hard gate:** no code, no scaffolding, no implementation
|
|
83
|
+
skill before the operator approves the design — including on "obviously simple"
|
|
84
|
+
tasks.
|
|
69
85
|
- **UI detection (mandatory check):** decide whether the task touches any
|
|
70
86
|
user-facing surface (web, mobile, CLI, TUI — new feature, new screen/command,
|
|
71
87
|
or a change to user-visible behavior). Record the verdict; it arms the UX
|
|
72
88
|
track in stage 3.
|
|
73
|
-
- **
|
|
89
|
+
- **Decomposition (platforms only): [`decomposition.md`](decomposition.md).** If
|
|
90
|
+
the brief describes a platform rather than a change — several independent
|
|
91
|
+
capabilities, several surfaces that could ship separately, REQs no single
|
|
92
|
+
deliverable satisfies — cut it into **modules** before any spec is written, and
|
|
93
|
+
commit the module map (`specs/<topic>-modules.md`): what each module delivers,
|
|
94
|
+
the entities it owns, what it depends on, the contracts it exposes, its REQs and
|
|
95
|
+
its status, in build order with the walking skeleton first. Single-module work
|
|
96
|
+
records `single module: <name>` in the design and moves on — a skipped
|
|
97
|
+
decomposition is a decision, never an omission.
|
|
98
|
+
- **GATE (manual):** the user approves the design, the UI verdict is recorded,
|
|
99
|
+
**every REQ is answered by the design** — a requirement the design doesn't
|
|
100
|
+
address is either covered now or explicitly dropped by the operator, with the
|
|
101
|
+
drop recorded in the carry-over ledger — **and, for a platform, the module map is
|
|
102
|
+
approved**: brick criteria met or excepted in writing, dependency graph acyclic,
|
|
103
|
+
build order topological, every REQ mapped to exactly one module, cross-module
|
|
104
|
+
contracts named with their owner.
|
|
74
105
|
|
|
75
|
-
## 3 — Spec
|
|
106
|
+
## 3 — Spec — with UX track for user-facing tasks
|
|
107
|
+
- **How it runs: [`spec.md`](spec.md)** — built into this skill: the UX-track order,
|
|
108
|
+
what the spec must lock (types, schemas, signatures, file layout, the **Global
|
|
109
|
+
Constraints** block stages 4–5 depend on), the self-review pass and the operator
|
|
110
|
+
review gate.
|
|
76
111
|
- **UX track (runs FIRST when stage 2 flagged UI; skip entirely otherwise).**
|
|
77
112
|
Requires the **super-ux** skills. If missing on a UI task → give the install
|
|
78
113
|
line and stop (see SKILL.md *Prerequisites*: `/plugin marketplace add
|
|
@@ -97,49 +132,72 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
|
|
|
97
132
|
never rebuild from scratch. If the chain already exists and is validated (e.g.
|
|
98
133
|
the task entered from super-ux), just verify (linter green) and embed it into
|
|
99
134
|
the spec; only build the parts that are missing.
|
|
100
|
-
- **Spec:**
|
|
101
|
-
`docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md` and
|
|
135
|
+
- **Spec:** write the approved design to
|
|
136
|
+
`docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md` and commit it. Lock all
|
|
102
137
|
shared contracts (types, schemas, signatures, file layout). For UI tasks the
|
|
103
138
|
spec **embeds the UX layer**: links the validated scenario IDs, the flows and
|
|
104
139
|
`SCR-` screens, the CJM stages the feature serves, and the UX
|
|
105
140
|
patterns/principles from super-ux that apply (`best-practices.md`,
|
|
106
141
|
`ux-design-principles.md`, `component-guidelines.md`).
|
|
107
|
-
- **GATE (manual):** spec committed **and** user-reviewed;
|
|
142
|
+
- **GATE (manual):** spec committed **and** user-reviewed; **every section carries
|
|
143
|
+
`covers: REQ-…` and every REQ appears in at least one section**; for UI tasks
|
|
108
144
|
additionally: the super-ux chain (foundation → flows → screens → scenarios) is
|
|
109
145
|
designed, validated and approved; scenarios validated in `docs/ux/scenarios.md`;
|
|
110
146
|
the linter passes; every user-facing spec requirement traces to a scenario ID
|
|
111
147
|
(or an explicit v1-mode/tiny-project waiver by the operator). No plan (stage 4)
|
|
112
148
|
starts before this — the chain comes BEFORE interface.
|
|
113
149
|
|
|
114
|
-
## 4 — Plan
|
|
115
|
-
- **
|
|
116
|
-
`docs/superpowers/plans/YYYY-MM-DD-<
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
150
|
+
## 4 — Plan
|
|
151
|
+
- **How it runs: [`planning.md`](planning.md)** — built into this skill →
|
|
152
|
+
`docs/superpowers/plans/YYYY-MM-DD-<topic>.md` (same slug as the brief and the
|
|
153
|
+
spec). Zero-context tasks, exact
|
|
154
|
+
paths, complete code in every step, TDD steps with expected output, DoD each,
|
|
155
|
+
dependency graph + parallel groups, non-overlapping file ownership, and the
|
|
156
|
+
Global Constraints block copied verbatim from the spec.
|
|
157
|
+
- **GATE (auto):** **set equality — the REQ ids in the brief equal the union of
|
|
158
|
+
`Implements:` across plan tasks.** A non-empty difference fails the gate and is
|
|
159
|
+
reported as the explicit list of dropped requirements; this is the seam where
|
|
160
|
+
scope leaks silently, so the check is mechanical, not a judgement call. Plus:
|
|
161
|
+
every spec requirement maps to a task; no placeholders; names and
|
|
162
|
+
types consistent across tasks; every task carries a verifiable DoD; parallel-group
|
|
120
163
|
tasks share no files. For UI tasks: every task building user-facing behavior
|
|
121
164
|
names the scenario ID(s) and `SCR-` screen(s) it implements, and its DoD
|
|
122
165
|
includes satisfying them **and** updating the affected super-ux layers in the
|
|
123
166
|
same change (super-ux *same-change* rule).
|
|
124
167
|
|
|
125
|
-
## 5 — Dev
|
|
126
|
-
- **
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
168
|
+
## 5 — Dev
|
|
169
|
+
- **How it runs: [`build.md`](build.md)** — built into this skill: isolate the
|
|
170
|
+
workspace (native worktree tool first, git fallback, baseline tests), keep a
|
|
171
|
+
ledger under `.task-pipeline/build/<plan>/` so a compacted context can resume,
|
|
172
|
+
then one fresh implementer subagent per task with a file-based brief and report,
|
|
173
|
+
a review after every task ([`review.md`](review.md)), and a five-round fix loop
|
|
174
|
+
with an explicit breaker. TDD per task ([`tdd.md`](tdd.md)): failing test →
|
|
175
|
+
watch it fail → minimal impl → watch it pass → commit. Pin subagents to the
|
|
176
|
+
run's confirmed model (`model-tiering.md`). The plan's parallel groups fan out
|
|
177
|
+
**only** when each implementer gets its own worktree; otherwise sequential.
|
|
178
|
+
- **Integration closes the stage:** sync with the base branch, re-run the full suite
|
|
179
|
+
on the result, land it the project's way (merge, or a PR — outward, so it needs a
|
|
180
|
+
go), remove the worktree. Stages 7–9 act on the integrated result, so a branch the
|
|
181
|
+
operator chose to leave unmerged is recorded as such.
|
|
182
|
+
- **GATE (auto):** all plan tasks DONE (three review verdicts per task: spec
|
|
183
|
+
compliance, **REQ satisfied**, code quality); every finding fixed or parked with a
|
|
184
|
+
ruling; **every parked finding and implementer concern harvested into the
|
|
185
|
+
carry-over ledger** — nothing stays only in the scratch workspace, which is
|
|
186
|
+
deleted; no task left BLOCKED; full test suite green; branch integrated per the brief's policy (or the operator's
|
|
187
|
+
"leave it" recorded).
|
|
131
188
|
|
|
132
|
-
## 6 — Tests
|
|
189
|
+
## 6 — Tests
|
|
133
190
|
- **What:** consolidate test coverage for the change: confirm new functionality
|
|
134
191
|
has tests (written test-first in stage 5), update/repair existing tests the
|
|
135
192
|
change touched, and add edge-case + failure-path tests per DoD.
|
|
136
|
-
- **Invoke:** the host test runner (see `conventions.md` → *Lint + test*);
|
|
137
|
-
|
|
193
|
+
- **Invoke:** the host test runner (see `conventions.md` → *Lint + test*); the
|
|
194
|
+
built-in [`tdd.md`](tdd.md) cycle for any uncovered gap — failing test first,
|
|
195
|
+
same as stage 5.
|
|
138
196
|
- **GATE (auto):** the **full** suite is green (not just the new tests); new/changed code
|
|
139
197
|
is covered; no `skip`/`xfail` smuggling a red suite past the gate. Never advance
|
|
140
198
|
to deploy on a red or partial run.
|
|
141
199
|
|
|
142
|
-
## 7 — Lint + deploy
|
|
200
|
+
## 7 — Lint + deploy
|
|
143
201
|
- Read host conventions (`conventions.md`): run the linter; fix failures. The suite
|
|
144
202
|
is already green from stage 6 — re-run it if code changed since. For UI projects,
|
|
145
203
|
the **super-ux linter** (`python3 docs/ux/lint.py` / `/ux-lint`) is part of lint —
|
|
@@ -147,19 +205,97 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
|
|
|
147
205
|
if the project defines release automation (`pipeline.json` → `release`, toggle
|
|
148
206
|
on), that is what "deploy" runs here.
|
|
149
207
|
- **GATE (manual):** lint clean (host linter **and**, for UI projects, the super-ux
|
|
150
|
-
linter) **and** suite green **before** deploy
|
|
208
|
+
linter) **and** suite green **before** deploy, **and no REQ is still `open`** — a
|
|
209
|
+
`partial` ships only with the operator's explicit acceptance. A gap is cheapest to
|
|
210
|
+
close before it ships, and the operator is already present at this gate. Deploy is outward → explicit
|
|
151
211
|
operator go. Respect deploy-from-main rules if the project mandates them.
|
|
152
212
|
|
|
153
|
-
## 8 — Post-deploy
|
|
213
|
+
## 8 — Post-deploy
|
|
154
214
|
- Tail deploy logs / health-check per conventions. Confirm clean boot, no error
|
|
155
215
|
spike, live subsystems healthy.
|
|
156
216
|
- **GATE (auto):** clean boot confirmed, or an **honest degradation report** with next
|
|
157
217
|
steps — never silent success.
|
|
158
218
|
|
|
159
|
-
## 9 — Docs + wiki
|
|
219
|
+
## 9 — Docs + wiki
|
|
160
220
|
- Update host module docs / runbooks per the project's self-update rules, in the
|
|
161
221
|
**same change**. For UI tasks, confirm the super-ux layers were updated in this
|
|
162
222
|
change and the linter is green (super-ux *same-change* + *no-drift* rules). Then
|
|
163
223
|
sync knowledge to the wiki (`wiki-update` skill).
|
|
164
224
|
- **GATE (auto):** docs in sync with code; UI: super-ux layers current + linter
|
|
165
225
|
green; wiki synced; dangling links fixed.
|
|
226
|
+
|
|
227
|
+
## 10 — Acceptance
|
|
228
|
+
- **What:** the closing stage — go back to the brief and account for **every**
|
|
229
|
+
requirement. Doctrine: [`acceptance.md`](acceptance.md). Every earlier gate asks
|
|
230
|
+
"is this artifact good?"; none asks "does this still contain everything that was
|
|
231
|
+
asked for?" The loss happens on the seams between stages, and this is where it
|
|
232
|
+
surfaces.
|
|
233
|
+
- **Runs last**, after docs and wiki — those are deliverables too, and a REQ may
|
|
234
|
+
name them.
|
|
235
|
+
- **How it runs:** built in. Read the brief's REQ table, the carry-over ledger in
|
|
236
|
+
full, the plan's task statuses, git log, the final suite output, stage-8 notes and
|
|
237
|
+
stage-9 doc changes (plus `docs/ux/scenarios.md` + `/ux-lint` for UI tasks). Write
|
|
238
|
+
`docs/superpowers/specs/YYYY-MM-DD-<topic>-acceptance.md` — one row per REQ,
|
|
239
|
+
status `verified` / `partial` / `deferred` / `dropped`, each with **evidence** (a
|
|
240
|
+
passing test name, `file:line`, a command and its output, or a scenario ID).
|
|
241
|
+
"Done" without evidence is not done: downgrade to `partial` and say so rather
|
|
242
|
+
than upgrading the claim.
|
|
243
|
+
- Then ask the operator the closing question out loud, list in hand: *here's what
|
|
244
|
+
you asked for, here's what shipped, here's what's deferred and where it lives —
|
|
245
|
+
what's missing?* Ask it even when the table is green; the operator holds context
|
|
246
|
+
the brief never captured, and this is the cheapest moment in the run to hear it.
|
|
247
|
+
- **GATE (manual):** every REQ has a status (none `unknown`); every `verified`
|
|
248
|
+
carries evidence; every `partial` names what's missing and where it's tracked;
|
|
249
|
+
every `deferred`/`dropped` has the operator's agreement and, for `deferred`, a
|
|
250
|
+
tracker entry; no carry-over row left `unresolved`; the operator answers the
|
|
251
|
+
closing question and signs off. Manual by design — an automated check can prove
|
|
252
|
+
the table is well-formed, only the person who asked can confirm it is what they
|
|
253
|
+
asked for.
|
|
254
|
+
|
|
255
|
+
## The program loop — a platform, one brick at a time
|
|
256
|
+
|
|
257
|
+
When stage 2 produced a **module map** ([`decomposition.md`](decomposition.md)),
|
|
258
|
+
stages 0–2 have run once for the whole platform and the rest of the pipeline runs
|
|
259
|
+
**per module**, in build order:
|
|
260
|
+
|
|
261
|
+
```
|
|
262
|
+
module N → 3 spec (dossier) → 4 plan → 5 build → 6 tests → 7 lint+deploy
|
|
263
|
+
→ 8 post-deploy → 9 docs+wiki → 10 acceptance → map status: done
|
|
264
|
+
→ module N+1 (back to 3)
|
|
265
|
+
```
|
|
266
|
+
|
|
267
|
+
- **No re-grilling, no re-decomposing per module.** New information that changes the
|
|
268
|
+
map goes back to stage 2 as an explicit, operator-approved map revision — never a
|
|
269
|
+
quiet edit mid-module.
|
|
270
|
+
- **Each module's spec is a full dossier** (`spec.md`): architecture, entities,
|
|
271
|
+
contracts in and out, business rules, edge and failure cases, UI/Figma chain when
|
|
272
|
+
it has a surface.
|
|
273
|
+
- **Deploy cadence is the brief's call** (autonomy sweep): per module, or once after
|
|
274
|
+
several. Decide it up front, not per module.
|
|
275
|
+
- **Update the module map's status in the same commit as that module's acceptance.**
|
|
276
|
+
The map is the resume point after a lost context.
|
|
277
|
+
- **Program done** when every row is `done` or `deferred` with an agreed home, the
|
|
278
|
+
cross-module contracts are covered by tests that cross the seam, and a final
|
|
279
|
+
acceptance covers the platform's whole REQ table — not module by module.
|
|
280
|
+
|
|
281
|
+
## Cross-cutting — the loop guard
|
|
282
|
+
|
|
283
|
+
Any stage can be re-entered and any loop can churn: a pass undoing what an earlier
|
|
284
|
+
pass decided, two shapes alternating, the same file rewritten with no new
|
|
285
|
+
information. [`loop-guard.md`](loop-guard.md) is the detector and the break
|
|
286
|
+
protocol, and it binds every repeating loop here — the stage-5 fix loop, a stage
|
|
287
|
+
re-entered after a failed gate, the program loop above, any audit → fix → audit
|
|
288
|
+
cycle.
|
|
289
|
+
|
|
290
|
+
- **Every repeating pass logs one line per touched file** (`touch: <file> — pass N —
|
|
291
|
+
reason: <finding id / gate item>`) to the run ledger. Detection is mechanical, not
|
|
292
|
+
a feeling, and the ledger is what survives compaction.
|
|
293
|
+
- **Trips on:** revert-oscillation (A→B→A); the same file edited twice for the same
|
|
294
|
+
reason; a finding already ADDRESSED or parked coming back; a stage entered a third
|
|
295
|
+
time for one artifact; two loops editing one file. Hard caps: 5 fix rounds per
|
|
296
|
+
task, 2 re-entries per stage per artifact, 3 passes per module.
|
|
297
|
+
- **On a trip: stop editing.** Name shapes A and B with their evidence, escalate to
|
|
298
|
+
the layer that owns the conflict (rubric → operator → plan → spec → module map),
|
|
299
|
+
re-plan the check as an ordered one-item-per-line checklist, then go through it in
|
|
300
|
+
order, one commit per item. Never settle a higher-layer conflict inside a lower
|
|
301
|
+
loop, and never adjudicate before the cap.
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
# TDD — stages 5 and 6, built in
|
|
2
|
+
|
|
3
|
+
How every task in the build is implemented, and what stage 6 consolidates. Built
|
|
4
|
+
into this skill; nothing to install.
|
|
5
|
+
|
|
6
|
+
> Ported from the `test-driven-development` skill in
|
|
7
|
+
> [obra/superpowers](https://github.com/obra/superpowers) (MIT — see `LICENSE` →
|
|
8
|
+
> *Third-party*), with the stage-6 suite gate added.
|
|
9
|
+
|
|
10
|
+
## The iron law
|
|
11
|
+
|
|
12
|
+
```
|
|
13
|
+
NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
**If you didn't watch the test fail, you don't know it tests the right thing.**
|
|
17
|
+
|
|
18
|
+
Wrote code before the test? Delete it and start from the test. Not "keep it as
|
|
19
|
+
reference", not "adapt it while writing tests", not "look at it once more". Delete
|
|
20
|
+
means delete — code you kept is code the test was written to fit.
|
|
21
|
+
|
|
22
|
+
**Always:** new features, bug fixes, refactors, behavior changes.
|
|
23
|
+
**Exceptions, and only with the operator's say-so:** throwaway prototypes,
|
|
24
|
+
generated code, pure configuration.
|
|
25
|
+
|
|
26
|
+
Thinking "skip TDD just this once"? That thought is the rationalization, not the
|
|
27
|
+
exception.
|
|
28
|
+
|
|
29
|
+
## Red → green → refactor
|
|
30
|
+
|
|
31
|
+
**RED — write one failing test.** One behavior, a name that describes that
|
|
32
|
+
behavior, real code rather than mocks wherever mocks are avoidable.
|
|
33
|
+
|
|
34
|
+
**Verify RED — run it and watch it fail. Mandatory.** Confirm it *fails* rather
|
|
35
|
+
than *errors*, that the message is the one you expected, and that it fails because
|
|
36
|
+
the feature is missing — not because of a typo or a bad import. A test that passes
|
|
37
|
+
immediately is testing behavior that already exists: fix the test.
|
|
38
|
+
|
|
39
|
+
**GREEN — the simplest code that passes.** No extra options, no "while I'm here"
|
|
40
|
+
refactor of the neighbors, no configuration surface nobody asked for. YAGNI.
|
|
41
|
+
|
|
42
|
+
**Verify GREEN — run it and watch it pass. Mandatory.** The new test passes, the
|
|
43
|
+
other tests still pass, and the output is pristine — no stray warnings or errors.
|
|
44
|
+
Test still fails → fix the code, never the test. Another test broke → fix it now.
|
|
45
|
+
|
|
46
|
+
**REFACTOR — only once green.** Remove duplication, improve names, extract helpers.
|
|
47
|
+
Tests stay green. No new behavior enters here.
|
|
48
|
+
|
|
49
|
+
Then the next failing test.
|
|
50
|
+
|
|
51
|
+
## Tests that stay honest
|
|
52
|
+
|
|
53
|
+
- **Before writing a test, name the production change that would make it fail.**
|
|
54
|
+
Can't name one? The test asserts nothing useful.
|
|
55
|
+
- **Assert on real behavior, never on mock behavior.** `expect(mock).toHaveBeenCalled()`
|
|
56
|
+
proves the mock works. Understand a dependency's side effects before mocking it —
|
|
57
|
+
a mock that lies is worse than no test.
|
|
58
|
+
- **One thing per test.** An "and" in the name means two tests.
|
|
59
|
+
- **Test-only helpers live in test utilities**, never as extra branches or flags in
|
|
60
|
+
production classes.
|
|
61
|
+
- **Edge cases and failure paths are part of the task**, not a follow-up ticket:
|
|
62
|
+
empty input, boundary values, the network call that fails, the timeout.
|
|
63
|
+
|
|
64
|
+
## Stage 6 — consolidation and the suite gate
|
|
65
|
+
|
|
66
|
+
Stage 5 wrote the tests task by task. Stage 6 makes the whole thing true:
|
|
67
|
+
|
|
68
|
+
- New functionality has tests (written test-first in stage 5) — fill any gap now,
|
|
69
|
+
the same way: failing test first.
|
|
70
|
+
- Tests the change touched are updated or repaired, not deleted around.
|
|
71
|
+
- Edge-case and failure-path coverage matches each task's DoD.
|
|
72
|
+
- The test command is the one recorded in the brief's autonomy sweep; "green" means
|
|
73
|
+
what the brief says it means (including a known-red baseline, if one was
|
|
74
|
+
recorded).
|
|
75
|
+
|
|
76
|
+
**GATE (auto):** the **full** suite is green — not just the new tests. New and
|
|
77
|
+
changed code is covered. No `skip` / `xfail` / commented-out assertion smuggles a
|
|
78
|
+
red suite past the gate. A partial or red run never advances to deploy; report it
|
|
79
|
+
honestly instead.
|
|
80
|
+
|
|
81
|
+
## When stuck
|
|
82
|
+
|
|
83
|
+
| Problem | What it means |
|
|
84
|
+
|---|---|
|
|
85
|
+
| Don't know how to test it | Write the API you wish existed, then the assertion. Still stuck → ask the operator. |
|
|
86
|
+
| The test is too complicated | The design is too complicated. Simplify the interface. |
|
|
87
|
+
| Everything has to be mocked | The code is too coupled. Inject dependencies. |
|
|
88
|
+
| Setup is enormous | Extract helpers; if it's still huge, the design is the problem. |
|
|
89
|
+
| Fixing a bug | Write the failing test that reproduces it first. The test proves the fix and prevents the regression. |
|
|
90
|
+
|
|
91
|
+
## Rationalizations
|
|
92
|
+
|
|
93
|
+
| Excuse | Reality |
|
|
94
|
+
|---|---|
|
|
95
|
+
| "Too simple to test" | Simple code breaks. The test costs 30 seconds. |
|
|
96
|
+
| "I'll test after" | Tests written after pass immediately, which proves nothing. You never watched it fail, so you never proved it can catch the bug. |
|
|
97
|
+
| "Tests after achieve the same thing — spirit, not ritual" | Tests-after answer "what does this do?"; tests-first answer "what should this do?" After-the-fact tests are biased by the code that already exists. |
|
|
98
|
+
| "I already tested it manually" | Ad-hoc, unrepeatable, no record of what was covered. "Worked when I tried it" is not coverage. |
|
|
99
|
+
| "Deleting X hours of code is wasteful" | Sunk cost. The real choice is rewriting with TDD versus keeping code you can't trust. |
|
|
100
|
+
| "Just exploring first" | Fine — throw the exploration away and start with TDD. |
|
|
101
|
+
| "The existing code has no tests either" | You're improving it. Add tests for what you touch. |
|
|
102
|
+
| "TDD will slow me down" | TDD is the fast path: bugs caught before commit, refactors without fear. The shortcut ends in production debugging. |
|
|
103
|
+
|
|
104
|
+
## Red flags — stop and start over
|
|
105
|
+
|
|
106
|
+
Code before test · test written after implementation · test passed on the first run
|
|
107
|
+
· can't explain why it failed · "tests later" · "just this once" · "keep it as
|
|
108
|
+
reference" · "it's about spirit not ritual" · "this case is different because…"
|
|
109
|
+
|
|
110
|
+
All of them mean the same thing: delete the code, start from the failing test.
|
|
@@ -1,12 +1,20 @@
|
|
|
1
1
|
# templates
|
|
2
2
|
|
|
3
|
-
Skeletons task-pipeline seeds into a host project. Only the **brief** is
|
|
4
|
-
|
|
5
|
-
|
|
3
|
+
Skeletons task-pipeline seeds into a host project. Only the **brief** is a seeded
|
|
4
|
+
file (it is the stage-0 intake artifact). The spec and plan have no skeleton here —
|
|
5
|
+
their required structure is prescribed inline by `references/spec.md` and
|
|
6
|
+
`references/planning.md`; the `docs/ux/*` skeletons come from `super-ux`.
|
|
6
7
|
|
|
7
8
|
| Template | Seeded to | Stage |
|
|
8
9
|
|---|---|---|
|
|
9
10
|
| `brief.md` | `docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` | 0 — intake grill |
|
|
11
|
+
| `carryover.md` | `docs/superpowers/specs/YYYY-MM-DD-<topic>-carryover.md` | 0 seeds, all stages append, 10 reads |
|
|
12
|
+
| `context.md` | `CONTEXT.md` at the repo root (or per context) | 0 — grill, domain awareness |
|
|
13
|
+
| `adr.md` | `docs/adr/NNNN-<slug>.md` | 0 — grill, hard-to-reverse decisions |
|
|
14
|
+
|
|
15
|
+
`context.md` and `adr.md` are **format references**, not files to copy wholesale:
|
|
16
|
+
the grill writes `CONTEXT.md` entries and ADRs in their shape, lazily — only once
|
|
17
|
+
there is a resolved term or a decision worth recording.
|
|
10
18
|
|
|
11
19
|
Seeding rule (per the ssheleg canon): create a template copy **only when the
|
|
12
20
|
target is absent**; never overwrite an existing brief.
|
|
@@ -0,0 +1,64 @@
|
|
|
1
|
+
# ADR — format
|
|
2
|
+
|
|
3
|
+
Architecture Decision Records live in `docs/adr/` with sequential numbering
|
|
4
|
+
(`0001-slug.md`, `0002-slug.md`, …). Scan for the highest existing number and
|
|
5
|
+
increment. Create the directory **lazily** — only when the first ADR is needed.
|
|
6
|
+
|
|
7
|
+
> Adapted from Matt Pocock's `grill-with-docs` (MIT — see the repo LICENSE →
|
|
8
|
+
> *Third-party*).
|
|
9
|
+
|
|
10
|
+
## Template
|
|
11
|
+
|
|
12
|
+
```md
|
|
13
|
+
# {Short title of the decision}
|
|
14
|
+
|
|
15
|
+
{1–3 sentences: what the context was, what was decided, and why.}
|
|
16
|
+
```
|
|
17
|
+
|
|
18
|
+
That's it. An ADR can be a single paragraph. The value is recording *that* a
|
|
19
|
+
decision was made and *why* — not filling out sections.
|
|
20
|
+
|
|
21
|
+
## Optional sections
|
|
22
|
+
|
|
23
|
+
Only when they add genuine value; most ADRs need none.
|
|
24
|
+
|
|
25
|
+
- **Status** frontmatter (`proposed | accepted | deprecated | superseded by
|
|
26
|
+
ADR-NNNN`) — useful once decisions start getting revisited.
|
|
27
|
+
- **Considered options** — only when the rejected alternatives are worth
|
|
28
|
+
remembering.
|
|
29
|
+
- **Consequences** — only when non-obvious downstream effects need calling out.
|
|
30
|
+
|
|
31
|
+
## When to write one
|
|
32
|
+
|
|
33
|
+
All three must hold:
|
|
34
|
+
|
|
35
|
+
1. **Hard to reverse** — changing your mind later carries real cost.
|
|
36
|
+
2. **Surprising without context** — a future reader will look at the code and
|
|
37
|
+
wonder "why on earth did they do it this way?"
|
|
38
|
+
3. **A real trade-off** — genuine alternatives existed and one was picked for
|
|
39
|
+
specific reasons.
|
|
40
|
+
|
|
41
|
+
Easy to reverse → skip it, you'll just reverse it. Not surprising → nobody will
|
|
42
|
+
wonder. No real alternative → there's nothing to record beyond "we did the obvious
|
|
43
|
+
thing."
|
|
44
|
+
|
|
45
|
+
### What qualifies
|
|
46
|
+
|
|
47
|
+
- **Architectural shape.** "We're using a monorepo." "The write model is
|
|
48
|
+
event-sourced; the read model projects into Postgres."
|
|
49
|
+
- **Integration patterns between contexts.** "Ordering and Billing communicate via
|
|
50
|
+
domain events, not synchronous HTTP."
|
|
51
|
+
- **Technology choices carrying lock-in.** Database, message bus, auth provider,
|
|
52
|
+
deployment target — not every library, just the ones that would take a quarter to
|
|
53
|
+
swap out.
|
|
54
|
+
- **Boundary and scope decisions.** "Customer data is owned by the Customer
|
|
55
|
+
context; others reference it by ID only." The explicit no's are as valuable as
|
|
56
|
+
the yes's.
|
|
57
|
+
- **Deliberate deviations from the obvious path.** "Manual SQL instead of an ORM
|
|
58
|
+
because X." Anything a reasonable reader would assume the opposite of — this is
|
|
59
|
+
what stops the next engineer from "fixing" something deliberate.
|
|
60
|
+
- **Constraints invisible in the code.** "No AWS, for compliance." "Sub-200ms
|
|
61
|
+
responses, per the partner API contract."
|
|
62
|
+
- **Rejected alternatives whose rejection is non-obvious.** Considered GraphQL,
|
|
63
|
+
picked REST for subtle reasons → record it, or someone re-proposes GraphQL in six
|
|
64
|
+
months.
|
|
@@ -15,6 +15,28 @@
|
|
|
15
15
|
- **Out of scope / explicitly deferred:** … (with the reason and, for deferrals,
|
|
16
16
|
the latest moment the decision can still be made)
|
|
17
17
|
|
|
18
|
+
## Requirements (the REQ spine — every later stage traces to these IDs)
|
|
19
|
+
|
|
20
|
+
Scope above is prose; this is the **addressable** form of it. One row per
|
|
21
|
+
independently verifiable deliverable — not one per sentence. Every row needs a
|
|
22
|
+
named check: **a requirement you can't say how to verify is a badly-stated
|
|
23
|
+
requirement** — split it here, on the grill, not at acceptance.
|
|
24
|
+
|
|
25
|
+
| ID | Requirement | How it's verified | Status |
|
|
26
|
+
|---|---|---|---|
|
|
27
|
+
| REQ-001 | … | test name / `file:line` / command + expected output / `SCN-…` | open |
|
|
28
|
+
| REQ-002 | … | … | open |
|
|
29
|
+
|
|
30
|
+
Status lifecycle, written at three checkpoints only (stage 4, stage 5, stage 10 —
|
|
31
|
+
not continuously): `open` → `planned` → `built` → `verified` \| `partial` \|
|
|
32
|
+
`deferred` \| `dropped`.
|
|
33
|
+
|
|
34
|
+
> **The list is frozen once confirmed.** Adding a requirement mid-run is fine —
|
|
35
|
+
> append it with its source. **Removing or narrowing one needs the operator's
|
|
36
|
+
> explicit agreement**, recorded in the carry-over ledger. Silently restating the
|
|
37
|
+
> task in smaller terms is the failure this table exists to prevent: every gate
|
|
38
|
+
> after it goes green on the shrunken task and nothing reports the loss.
|
|
39
|
+
|
|
18
40
|
## Users & context
|
|
19
41
|
|
|
20
42
|
- **Who / for what:** … (personas, the job being done)
|
|
@@ -26,6 +48,33 @@
|
|
|
26
48
|
|---|---|---|---|
|
|
27
49
|
| 1 | … | … | … |
|
|
28
50
|
|
|
51
|
+
## Autonomy (the sweep — stages 1→10 read this instead of asking)
|
|
52
|
+
|
|
53
|
+
Every row is either a resolved answer or an explicit **STOP AND ASK**. A blank row
|
|
54
|
+
is not neutral — it is a scheduled interruption.
|
|
55
|
+
|
|
56
|
+
| Stage | Question | Answer |
|
|
57
|
+
|---|---|---|
|
|
58
|
+
| run-wide | Model for this run | … (most capable available unless overridden; per-stage overrides here) |
|
|
59
|
+
| run-wide | Decide autonomously vs escalate to me | … |
|
|
60
|
+
| 1 Docs | External libs/APIs/SDKs in play; any context7 can't resolve → where their docs live | … |
|
|
61
|
+
| 2 Decompose | Platform (several capabilities/surfaces) or one module? If platform — deploy cadence: per module, or once at the end | … |
|
|
62
|
+
| 2–3 Spec | UI verdict (arms super-ux); scenario-tracing waiver, if any | … |
|
|
63
|
+
| 4–5 Dev | Base branch; worktree/branch policy; is `main` off-limits; commit convention; task tracker | … |
|
|
64
|
+
| 5 Integration | How the branch lands — direct merge, PR (who approves), or "leave it, I'll merge"; is parallel fan-out (one worktree per implementer) wanted? | … |
|
|
65
|
+
| 6 Tests | Test command; what "green" means; known-red baseline; coverage expectation | … |
|
|
66
|
+
| 7 Lint | Lint command (incl. `docs/ux/lint.py` for UI projects) | … |
|
|
67
|
+
| 7 Deploy | Target + path; release automation on/off; deploy-from-main rule | … |
|
|
68
|
+
| 7 Deploy | **Authorization** — standing go, or ask every time? | … |
|
|
69
|
+
| 8 Post-deploy | Where logs / health live (app name, endpoint, workflow) | … |
|
|
70
|
+
| 9 Docs+wiki | Which module docs / runbooks this change updates; wiki sync yes/no | … |
|
|
71
|
+
| 10 Acceptance | Who signs off; where deferred REQs get tracked (issue tracker / backlog) | … |
|
|
72
|
+
|
|
73
|
+
> **Deploy authorization has a hard floor.** A standing go counts only if it is
|
|
74
|
+
> **specific** — named target and named preconditions ("staging, once lint and the
|
|
75
|
+
> full suite are green; production always asks"). Vague blanket permission does not
|
|
76
|
+
> authorize an outward, irreversible action; stage 7 stops and asks.
|
|
77
|
+
|
|
29
78
|
## Done-criteria
|
|
30
79
|
|
|
31
80
|
- Observable, verifiable conditions that mean "this task is finished".
|
|
@@ -0,0 +1,36 @@
|
|
|
1
|
+
# Carry-over ledger — <topic>
|
|
2
|
+
|
|
3
|
+
> **Append-only.** Any stage may add a row; nobody edits or deletes one. Committed
|
|
4
|
+
> to `docs/superpowers/specs/YYYY-MM-DD-<topic>-carryover.md` beside the brief, and
|
|
5
|
+
> read in full by stage 10 (acceptance).
|
|
6
|
+
>
|
|
7
|
+
> **The rule: deferred out loud is forgotten.** If it isn't written here, it wasn't
|
|
8
|
+
> deferred — it was lost. That covers everything said in passing: "we'll do that
|
|
9
|
+
> later", "good enough for now", a `DONE_WITH_CONCERNS` from an implementer, a
|
|
10
|
+
> reviewer's non-blocking finding, a requirement the operator agreed to drop.
|
|
11
|
+
|
|
12
|
+
| # | Stage | What | Why it isn't done | REQ | Where it lives now |
|
|
13
|
+
|---|---|---|---|---|---|
|
|
14
|
+
| 1 | 5 Dev | XLSX export path | scope call — CSV first | REQ-004 | LIN-483 |
|
|
15
|
+
| 2 | 5 Review | `export.ts` lacks a size guard | minor, non-blocking | — | backlog |
|
|
16
|
+
| 3 | 2 Brainstorm | REQ-007 dropped: bulk export | operator agreed 2026-07-28 | REQ-007 | dropped |
|
|
17
|
+
|
|
18
|
+
## Columns
|
|
19
|
+
|
|
20
|
+
- **Stage** — where it surfaced, so acceptance knows how far it travelled.
|
|
21
|
+
- **What** — the concrete thing not done. "Error handling" is not an entry;
|
|
22
|
+
"`export.ts` swallows a failed write instead of surfacing it" is.
|
|
23
|
+
- **Why it isn't done** — scope call, blocked, deliberate deferral, out of budget.
|
|
24
|
+
"Forgot" is a legitimate and useful answer here.
|
|
25
|
+
- **REQ** — the requirement it belongs to, or `—` if it's outside the REQ spine.
|
|
26
|
+
- **Where it lives now** — issue id, backlog, `dropped` (with the operator's
|
|
27
|
+
agreement), or `unresolved`. **`unresolved` blocks the stage-10 gate**: an item
|
|
28
|
+
with no home is exactly the thing that gets forgotten, so acceptance refuses to
|
|
29
|
+
close on it.
|
|
30
|
+
|
|
31
|
+
## Notes
|
|
32
|
+
|
|
33
|
+
- Adding a row costs one line and never blocks a stage — that is the point. The
|
|
34
|
+
ledger is cheap precisely so nobody is tempted to keep it in their head.
|
|
35
|
+
- Rows referencing a REQ feed that REQ's final status: an open carry-over row
|
|
36
|
+
means the requirement is at best `partial`, never `verified`.
|