task-pipeline-skill 0.10.0 → 0.17.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (30) hide show
  1. package/CHANGELOG.md +401 -0
  2. package/LICENSE +85 -0
  3. package/README.md +211 -84
  4. package/cursor/rules/task-pipeline.mdc +135 -20
  5. package/package.json +3 -3
  6. package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
  7. package/plugins/task-pipeline/commands/task-pipeline.md +16 -8
  8. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +139 -57
  9. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +41 -27
  10. package/plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json +1 -1
  11. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
  12. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +32 -6
  13. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
  14. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
  15. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +73 -21
  16. package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
  17. package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
  18. package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +169 -0
  19. package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
  20. package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +55 -22
  21. package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
  22. package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
  23. package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
  24. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +184 -48
  25. package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
  26. package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +11 -3
  27. package/plugins/task-pipeline/skills/task-pipeline/templates/adr.md +64 -0
  28. package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +49 -0
  29. package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
  30. package/plugins/task-pipeline/skills/task-pipeline/templates/context.md +87 -0
@@ -1,16 +1,24 @@
1
1
  ---
2
- description: Run a task through task-pipeline — an intake grill that expands the request, then docs → brainstorm → spec → plan → build → tests → deploy → post-deploy → docs/wiki.
2
+ description: Run a task through task-pipeline — an intake grill that expands the request, then docs → brainstorm → spec → plan → build → tests → deploy → post-deploy → docs/wiki → acceptance.
3
3
  argument-hint: <one-line task description>
4
4
  ---
5
5
  Use the `task-pipeline` skill to run the task below through all gated stages —
6
6
  **stage 0 intake grill** → docs study → brainstorm → spec → plan → subagent
7
- build → tests → lint/deploy → post-deploy → docs/wiki. Start with the **intake
8
- grill**: interview the operator one question at a time (with a recommended answer
9
- each, exploring the codebase before asking) until every decision branch is
10
- resolved and the brief is locked — so the rest runs autonomously. For any
11
- user-facing task, recommend/use **super-ux**. Honor every stage gate by its type
12
- (`auto` = verify yourself; `manual` = wait for explicit go) and emit the
13
- per-stage model reminder when the recommended model differs from the current one.
7
+ build → tests → lint/deploy → post-deploy → docs/wiki → **acceptance**. **Every stage's doctrine is
8
+ built into the skill** (`references/{grill,brainstorm,decomposition,spec,planning,build,review,tdd,acceptance,loop-guard}.md`)
9
+ — no companion plugin is required for any of them. The **intake grill is
10
+ mandatory** (`references/grill.md`): interview the
11
+ operator one question at a time (with a recommended answer each, exploring the
12
+ codebase before asking) until every decision branch is resolved, applying the
13
+ grill's **domain awareness** (challenge terms against `CONTEXT.md`, sharpen fuzzy
14
+ language, ADRs for hard-to-reverse calls) and covering the **autonomy sweep** (what
15
+ would otherwise stop stages 1→10: docs sources, branch/tracker policy, test and lint
16
+ commands, deploy target and authorization, log locations, docs/wiki targets) —
17
+ until the brief is locked — including the **REQ table**, the request as an addressable list where every row names how it is verified — so the rest runs autonomously and the final stage can account for all of it. The list is frozen: adding is free, removing needs the operator's agreement. Anything deferred goes into the carry-over ledger the moment it's said. For any user-facing task, recommend/use
18
+ **super-ux**. **If the brief describes a platform rather than a change**, stage 2 also cuts it into modules (`references/decomposition.md`) — module map committed, walking skeleton first, every REQ in exactly one module — and stages 3→10 then run per module, one brick at a time. **If any loop starts undoing an earlier pass** (same file edited twice for the same reason, a closed finding returning, a third entry into one stage), stop and run the loop guard (`references/loop-guard.md`): name both shapes, escalate to the layer that owns the conflict, re-plan the check as an ordered list, then go item by item. Honor every stage gate by its type (`auto` = verify yourself;
19
+ `manual` = wait for explicit go). Confirm the **model once at preflight** —
20
+ recommend the most capable one the environment offers, never a hardcoded id — then
21
+ run the whole pipeline on it without re-asking.
14
22
 
15
23
  Task: $ARGUMENTS
16
24
 
@@ -1,41 +1,59 @@
1
1
  ---
2
2
  name: task-pipeline
3
- description: "Use when running a substantial task through the full end-to-end delivery pipeline — an up-front intake grill that expands the request into a complete brief, then docs study, brainstorm, spec, plan, subagent-driven build, test suite, lint/deploy, post-deploy log check, and docs/wiki sync — as gated stages built on the superpowers skills. Use when the user wants to run a task through the pipeline, asks for the full cycle / полный цикл / прогони по конвейеру, invokes /task-pipeline, or starts any substantial feature, fix, or build that should follow the disciplined cycle rather than ad-hoc coding. Grills the operator first so the rest runs autonomously; recommends super-ux for any user-facing task; reminds which model to switch to per stage; reads host-project conventions for deploy/docs/wiki so it stays project-agnostic."
3
+ description: "Use when running a substantial task through the full end-to-end delivery pipeline — an up-front intake grill that expands the request into a complete brief, then docs study, brainstorm, spec, plan, subagent-driven build, tests, lint/deploy, post-deploy log check, docs/wiki sync and acceptance — as gated stages whose doctrine is built entirely into this skill (no required companion skills). Triggers - 'run this through the pipeline' / 'прогони по конвейеру', 'the full cycle' / 'полный цикл', /task-pipeline, or any substantial feature, fix, or build that should follow the disciplined cycle rather than ad-hoc coding. The intake grill is mandatory - it front-loads every decision, including the per-stage autonomy sweep, so stages 1→10 run without mid-flight questions; recommends super-ux for user-facing work; confirms one model up front (most capable available, never a hardcoded id); reads host-project conventions for deploy/docs/wiki so it stays project-agnostic."
4
4
  ---
5
5
 
6
6
  # task-pipeline
7
7
 
8
- Thin orchestrator. Runs a task through **gated stages**, each built on an
9
- existing skill. Keeps the main thread disciplined: no stage advances until its
10
- gate passes; each stage names the model to use.
8
+ Self-contained orchestrator. Runs a task through **gated stages**, each carrying its
9
+ own built-in doctrine — no companion plugin required. Keeps the main thread
10
+ disciplined: no stage advances until its gate passes; the whole run uses one model,
11
+ confirmed before it starts.
11
12
 
12
13
  **Grill first, then run autonomously.** A one-line task ("make me feature X") is
13
- never enough to finish without a human in the loop. Stage 0 **grills the operator
14
- up front** — a relentless, one-question-at-a-time interview that resolves every
15
- decision branch — and locks the answers into a brief. That front-loads all the
16
- human input so stages 1→9 can run to the end with only the built-in gate
17
- approvals, not mid-flight discovery.
14
+ never enough to finish without a human in the loop. Stage 0 is **mandatory**: a
15
+ relentless, one-question-at-a-time interview that resolves every decision branch
16
+ *and* sweeps stages 1→10 for anything that would stop the run later — then locks
17
+ the answers into a brief. Autonomy is bought there or not at all; every question
18
+ skipped at stage 0 comes back as an interruption at stage 6.
18
19
 
19
20
  **Config contract: [`pipeline.schema.json`](pipeline.schema.json).** A pipeline is
20
21
  a machine-readable config — an ordered list of stages, each with `skills[]` (the
21
22
  skills/agents that run it) and a `gate {type, check}`. The schema is the universal
22
23
  contract; it imposes **no** specific stages, skills, or gate assignments.
23
24
  [`pipeline.example.json`](pipeline.example.json) is a **copy-and-rewrite example**
24
- that encodes this plugin's own default flow (stage 0 intake + the 1→9 stages
25
+ that encodes this plugin's own default flow (stage 0 intake + the 1→10 stages
25
26
  tabled below) and an optional, toggleable `release` block. Any project replaces it
26
27
  wholesale — any number of stages, run by its own skills/agents, with its own gate
27
28
  types (see *Bring your own skills*). Each gate has a **type**: `auto` (the
28
29
  orchestrator verifies the `check` itself, pass/fail) or `manual` (wait for an
29
30
  explicit operator go); which stages are manual is the operator's call.
30
31
 
31
- ## Prerequisite
32
-
33
- Requires the **superpowers** skills. Preflight: confirm `superpowers:brainstorming`,
34
- `superpowers:writing-plans`, `superpowers:subagent-driven-development`,
35
- `superpowers:using-git-worktrees`, `superpowers:test-driven-development` resolve.
36
- If missing → tell the operator to install from **https://github.com/obra/superpowers**
37
- (`/plugin marketplace add obra/superpowers` → `/plugin install superpowers@superpowers`)
38
- and stop.
32
+ ## Prerequisites — none required
33
+
34
+ **Every stage's doctrine ships inside this skill.** There is no required companion
35
+ plugin, nothing to resolve at preflight, no version skew with someone else's repo,
36
+ and no stage that can fail because a dependency is missing:
37
+
38
+ | Stage | Built-in doctrine |
39
+ |---|---|
40
+ | 0 Intake grill | [`references/grill.md`](references/grill.md) |
41
+ | 2 Brainstorm | [`references/brainstorm.md`](references/brainstorm.md) |
42
+ | 2 Decompose (platforms only) | [`references/decomposition.md`](references/decomposition.md) |
43
+ | 3 Spec | [`references/spec.md`](references/spec.md) |
44
+ | 4 Plan | [`references/planning.md`](references/planning.md) |
45
+ | 5 Build (worktree, subagents, fix loop) | [`references/build.md`](references/build.md) + [`references/review.md`](references/review.md) |
46
+ | 5–6 TDD + suite gate | [`references/tdd.md`](references/tdd.md) |
47
+ | 10 Acceptance (REQ close-out) | [`references/acceptance.md`](references/acceptance.md) |
48
+ | any repeating loop | [`references/loop-guard.md`](references/loop-guard.md) |
49
+
50
+ **Optional bridge.** If the operator already runs an equivalent skill set (e.g.
51
+ `superpowers:brainstorming` / `writing-plans` / `subagent-driven-development` /
52
+ `using-git-worktrees` / `test-driven-development`), it can be mapped onto stages
53
+ 2/4/5/6 in `pipeline.json` → `skills[]`. That is a **substitution, never a
54
+ requirement**: the built-in doctrine is normative, the gates in
55
+ `references/stages.md` still govern, and nothing detects, recommends or waits for
56
+ an external provider.
39
57
 
40
58
  **super-ux — recommended for ANY user-facing task.** The moment a task implies a
41
59
  user interface (web / mobile / CLI / TUI — a screen, a command, a visible
@@ -56,63 +74,119 @@ workflow for the WHY→UI→scenario chain (`/ux`, `ux-foundation`, `ux-flows`,
56
74
  (or `npx skills add ssheleg/super-ux`). For UI tasks the spec gate **requires**
57
75
  it — install before stage 3, otherwise stop and ask the operator to install.
58
76
 
59
- **grill-me (optional, enhances stage 0).** If the `grill-me` / `grilling` skill
60
- resolves, stage 0 uses it; otherwise stage 0 runs its own built-in grill loop —
61
- no hard dependency. Install (optional): `npx skills add mattpocock/skills` or the
62
- engineering-advanced-skills marketplace.
77
+ **The grill is built in — no companion skill, nothing to install.** Stage 0 ships
78
+ with this skill: the full doctrine lives in [`references/grill.md`](references/grill.md)
79
+ (interview loop, domain awareness, autonomy sweep, output). It is **mandatory** —
80
+ no "clear enough task" exemption, no starting stage 1 without a committed,
81
+ operator-confirmed brief. The one sanctioned bypass is the entry-from-super-ux
82
+ short-circuit, and even that demands a scope confirmation.
83
+
84
+ It also produces the **REQ spine**: the request as an addressable list of
85
+ requirements, each naming how it will be verified. Stages 3–5 trace to those ids,
86
+ stage 4's gate is a mechanical set-comparison against them, and **stage 10 accounts
87
+ for every one** — which is what turns the pipeline from a funnel into a circle.
88
+
89
+ Two things the grill does beyond clarifying the request:
90
+ - **Domain awareness.** It reads the project's own `CONTEXT.md` / `docs/adr/` and
91
+ holds the operator to them — challenging terms that conflict with the glossary,
92
+ sharpening overloaded words, stress-testing with concrete scenarios, and
93
+ flagging where the code contradicts what was just said. Resolved terms are
94
+ written to `CONTEXT.md` as they land; genuinely hard-to-reverse decisions get an
95
+ ADR.
96
+ - **The autonomy sweep.** It pre-resolves what would otherwise stop stages 1→10
97
+ mid-flight (test/lint/deploy commands, branch policy, log locations, docs
98
+ targets, the model decision, deploy authorization). Autonomy is bought here or
99
+ not at all — an unasked question is a scheduled interruption.
63
100
 
64
101
  ## How to run
65
102
 
66
103
  1. Restate the task in one line. Create a **TaskList: one task per stage, starting
67
104
  with stage 0** (survives context loss; lets you resume). Then run the
68
- **companion preflight** (`references/companion-skills.md`): detect which
69
- companion skills resolve and emit the recommendation block — install the
70
- required/recommended ones (superpowers always; super-ux for UI tasks) before
71
- proceeding.
72
- 2. **Run stage 0 (Intake grill) first** — grill the operator until shared
73
- understanding is reached and the brief is locked (`references/stages.md` → 0).
74
- Do not touch stage 1 before the brief is confirmed. **Entered from super-ux?**
105
+ **companion preflight** (`references/companion-skills.md`): the stage doctrine
106
+ is built in, so this only checks the *optional* companions (super-ux for UI
107
+ tasks, context7, wiki-update) and emits ONE block covering them
108
+ **and the model decision** (`references/model-tiering.md`): recommend
109
+ the most capable model available, let the operator confirm or override, record
110
+ it. Ask once, here.
111
+ 2. **Run stage 0 (Intake grill) — always, no exceptions.** Grill until shared
112
+ understanding is reached, the autonomy sweep is covered, **the REQ table is
113
+ written (one row per independently verifiable deliverable, each naming its
114
+ check)** and the brief is locked
115
+ (`references/stages.md` → 0). Do not touch stage 1 before the brief is
116
+ committed and confirmed. **Entered from super-ux?**
75
117
  (a validated `docs/ux/` chain and/or a `docs/ux/plans/…` fix plan already
76
118
  exists — super-ux's `/ux` hands off here) → don't re-grill or rebuild the UX
77
119
  chain: just check it's OK (`/ux-lint` green), confirm scope in one line, and
78
120
  skip ahead to the first stage with real work (see `references/stages.md` → 0).
79
- 3. Walk stages 1→9. Before each: **model check** (see `references/model-tiering.md`) —
80
- if recommended ≠ current, emit the reminder block and wait for the operator to `/model`.
121
+ 3. Walk stages 1→10 on the model confirmed at preflight. **Don't re-ask about the
122
+ model at every boundary** — only when the operator recorded a per-stage override
123
+ map and the next stage's entry differs (`references/model-tiering.md`).
124
+ **Is the brief a platform rather than a change?** Then stage 2 also cuts it into
125
+ modules (`references/decomposition.md`) and stages 3→10 run **per module** in
126
+ build order, one brick at a time — stages 0–2 run once, and the module map's
127
+ status column is the resume point (`references/stages.md` → *The program loop*).
81
128
  4. Do **not** advance until the stage **gate** passes (`references/stages.md`).
82
129
  Honor the gate **type**: for `auto`, verify the gate's `check` yourself and
83
130
  stop/return on fail; for `manual`, present the result and **wait for the
84
131
  operator's explicit "continue"/go** — an auto gate never substitutes for a
85
132
  required manual approval.
86
- 5. Cross-cutting, every stage: task tracker + conventional commits per host
87
- conventions; worktree isolation for the build; honest degradation (never claim a
88
- failed/skipped step succeeded); outward/irreversible actions (deploy, publish,
89
- repo create) need explicit operator go.
133
+ 5. Cross-cutting, every stage: **answer from the brief's autonomy section rather
134
+ than asking again** — it was grilled precisely so you wouldn't have to;
135
+ **anything deferred, dropped or left half-done goes into the carry-over ledger
136
+ the moment it's said** — deferred out loud is forgotten; **never narrow the task
137
+ silently** — the REQ list is frozen, adding is free, removing needs the
138
+ operator's explicit agreement; **when a loop starts undoing an earlier pass —
139
+ the same file edited twice for the same reason, a closed finding coming back, a
140
+ third entry into one stage — stop and run the loop guard**
141
+ (`references/loop-guard.md`): name the two shapes, escalate to the layer that
142
+ owns the conflict, re-plan the check as an ordered list, then go through it one
143
+ item at a time; task
144
+ tracker + conventional commits per host conventions; worktree isolation for the
145
+ build, integrated back per the brief's branch policy before stage 7; honest
146
+ degradation (never claim a failed/skipped step succeeded);
147
+ outward/irreversible actions (deploy, publish, repo create, opening a PR) need explicit
148
+ operator go — or a **specific** standing authorization recorded in the brief
149
+ (named target + preconditions; a vague "do everything" is not one).
90
150
 
91
151
  ## Stages (detail in `references/stages.md`)
92
152
 
93
- | # | Stage | Model | Invoke | Gate | Type |
94
- |---|---|---|---|---|---|
95
- | 0 | Intake grill | Fable | `grill-me` / `grilling` if present, else built-in grill loop | shared understanding reached; brief locked + confirmed | manual |
96
- | 1 | Docs study | Fable | `context7` (resolve-library-id → get-library-docs) / `context7-docs` | contracts grounded on fetched docs | auto |
97
- | 2 | Brainstorm | Fable | `superpowers:brainstorming` + **UI detection** | design approved; UI verdict recorded | manual |
98
- | 3 | Spec | Fable | **UI → super-ux chain first** (`/ux` → `ux-foundation` CJM → `ux-flows` screens → `ux-scenarios` → `/ux-lint`), then spec `docs/superpowers/specs/…-design.md` | committed + reviewed; UI: chain validated, linter green, scenarios/`SCR-` traced | manual |
99
- | 4 | Plan | Fable | `superpowers:writing-plans` → `docs/superpowers/plans/…md` | parallel-ready, DoD per task | auto |
100
- | 5 | Dev | **Opus** | `superpowers:using-git-worktrees` + `superpowers:subagent-driven-development` (TDD) | tasks DONE, TDD green per task | auto |
101
- | 6 | Tests | **Opus** | host test runner + `superpowers:test-driven-development` | full suite green; new/changed code covered | auto |
102
- | 7 | Lint + deploy | host | host lint → deploy per host convention | lint clean + suite green before deploy; deploy needs go | manual |
103
- | 8 | Post-deploy | host | tail deploy logs / health-check | clean boot or honest degradation report | auto |
104
- | 9 | Docs + wiki | host | host module docs/runbook rules → `wiki-update` | docs synced, wiki synced | auto |
105
-
106
- ## Model reminder (emit at a boundary when recommended ≠ current)
107
-
108
- > ⏸ **Stage N (`<stage>`) recommends `<model>` (`<id>`).** You're on `<current>`.
109
- > Switch: `/model <id>` — then say "continue". *(Reminder only — override if you
110
- > don't have that model.)*
153
+ All stages run on the **one model confirmed at preflight** (default: the most
154
+ capable available — see `references/model-tiering.md`).
155
+
156
+ | # | Stage | Invoke | Gate | Type |
157
+ |---|---|---|---|---|
158
+ | 0 | Intake grill — **mandatory** | built in: [`references/grill.md`](references/grill.md) | shared understanding reached; autonomy sweep covered; brief locked + confirmed | manual |
159
+ | 1 | Docs study | `context7` (resolve-library-id → get-library-docs) / `context7-docs` | contracts grounded on fetched docs | auto |
160
+ | 2 | Brainstorm + decompose | built in: [`references/brainstorm.md`](references/brainstorm.md) + **UI detection** + [`references/decomposition.md`](references/decomposition.md) for platforms | design approved; UI verdict recorded; every REQ answered; platform: module map approved | manual |
161
+ | 3 | Spec | built in: [`references/spec.md`](references/spec.md) — **UI → super-ux chain first** (`/ux` → `ux-foundation` CJM → `ux-flows` screens → `ux-scenarios` → `/ux-lint`), then spec `docs/superpowers/specs/…-design.md` | committed + reviewed; UI: chain validated, linter green, scenarios/`SCR-` traced | manual |
162
+ | 4 | Plan | built in: [`references/planning.md`](references/planning.md) → `docs/superpowers/plans/…md` | parallel-ready, DoD per task | auto |
163
+ | 5 | Dev | built in: [`references/build.md`](references/build.md) (worktree → subagent per task → review loop → integrate) + [`references/tdd.md`](references/tdd.md) | tasks DONE, TDD green per task, branch integrated per the brief | auto |
164
+ | 6 | Tests | host test runner + built-in [`references/tdd.md`](references/tdd.md) | full suite green; new/changed code covered | auto |
165
+ | 7 | Lint + deploy | host lint → deploy per host convention | lint clean + suite green before deploy; deploy needs a go (or the brief's specific standing authorization) | manual |
166
+ | 8 | Post-deploy | tail deploy logs / health-check | clean boot or honest degradation report | auto |
167
+ | 9 | Docs + wiki | host module docs/runbook rules → `wiki-update` | docs synced, wiki synced | auto |
168
+ | 10 | **Acceptance** | built in: [`references/acceptance.md`](references/acceptance.md) | every REQ accounted for with evidence; ledger has no unresolved row; operator signs off | manual |
169
+
170
+ ## Model — ask once, at preflight
171
+
172
+ Default recommendation: **the most capable reasoning model the environment
173
+ offers** (currently the latest Opus generation — read that as a tier, not a
174
+ string). **Never hardcode a model id**: generations ship, tiers get renamed, and
175
+ the operator may be on another provider entirely — resolve the top tier available
176
+ at runtime. Stage configs use provider-agnostic tokens (`default` / `inherit`).
177
+
178
+ > 🧠 **Model for this run:** recommended **`<top tier available>`**. You're on
179
+ > `<current>`. `/model <id>` to switch, or "keep current", or name per-stage
180
+ > overrides. *(Reminder only — if that tier isn't available, say which one you're
181
+ > using and continue.)*
182
+
183
+ Record the answer in the brief; don't re-ask per stage. Stage-5 subagents are
184
+ pinned to the confirmed model automatically. Detail: `references/model-tiering.md`.
111
185
 
112
186
  ## Bring your own skills
113
187
 
114
- The stages above (stage 0 intake + 1→9) are the **example** flow (grill +
115
- superpowers + a super-ux UX track for user-facing tasks + host conventions). A
188
+ The stages above (stage 0 intake + 1→10) are the **example** flow (this skill's
189
+ built-in doctrine + a super-ux UX track for user-facing tasks + host conventions). A
116
190
  host project owns its pipeline: copy `pipeline.example.json` → `pipeline.json`,
117
191
  then define its **own** stages (any count), point each stage's `skills[]` at the
118
192
  skills/agents its environment resolves, set each `gate.type` (`auto`/`manual`) to
@@ -123,9 +197,17 @@ automation is on — `pipeline.schema.json` is the only contract.
123
197
  ## References
124
198
 
125
199
  - `pipeline.schema.json` — the universal pipeline config contract (stages + release)
126
- - `pipeline.example.json` — this plugin's default flow (stage 0 + 1→9) + release, as config
200
+ - `pipeline.example.json` — this plugin's default flow (stage 0 + 1→10) + release, as config
201
+ - `references/grill.md` — the built-in stage-0 grill: loop, domain awareness, autonomy sweep
202
+ - `references/acceptance.md` — the built-in stage-10 close-out: REQ coverage, evidence, sign-off
203
+ - `references/brainstorm.md` — stage 2: design dialogue, approaches, UI detection, hard gate
204
+ - `references/spec.md` — stage 3: UX track order, the spec contract, self-review, review gate
205
+ - `references/planning.md` — stage 4: zero-context plan format, parallel groups, no placeholders
206
+ - `references/build.md` — stage 5: isolation, ledger, subagent task loop, fix loop, final review
207
+ - `references/review.md` — the review rubric, diff packages and the three reviewer prompts
208
+ - `references/tdd.md` — stages 5–6: the iron law, red/green/refactor, the suite gate
127
209
  - `references/stages.md` — per-stage detail + exact gate criteria + gate types
128
210
  - `references/model-tiering.md` — model map, ids, the `/model` reminder mechanic, override
129
- - `references/conventions.md` — how stages 6–9 read the host project's CLAUDE.md
211
+ - `references/conventions.md` — how stages 6–10 read the host project's CLAUDE.md
130
212
  - `references/companion-skills.md` — companion skills, install lines, preflight recommendation
131
213
  - `references/artifacts.md` — the canonical document/artifact layout per stage
@@ -1,27 +1,26 @@
1
1
  {
2
2
  "$schema": "./pipeline.schema.json",
3
3
  "version": 1,
4
- "_note": "EXAMPLE ONLY — copy this file, rename to pipeline.json in your project, and rewrite it. This particular example encodes the plugin's own default flow (an up-front intake grill + superpowers + a super-ux UX track for user-facing tasks); it is NOT a fixed contract. Your project defines its own stages (any count), each executed by your own skills/agents, with your own gate types. The universal contract is pipeline.schema.json; test/validate.py checks this example against it. gate.type: auto = orchestrator verifies the check itself (pass/fail); manual = wait for an explicit operator go. Which stages are manual vs auto is the operator's decision, not the plugin's.",
4
+ "_note": "EXAMPLE ONLY — copy this file, rename to pipeline.json in your project, and rewrite it. This particular example encodes the plugin's own default flow (an up-front intake grill + this skill's own built-in stage doctrine + a super-ux UX track for user-facing tasks); it is NOT a fixed contract. Your project defines its own stages (any count), each executed by your own skills/agents, with your own gate types. Stage models use provider-agnostic tokens ('default' = the model confirmed for the run, 'inherit' = whatever the operator is on) — never hardcode a vendor model id, it goes stale. The universal contract is pipeline.schema.json; test/validate.py checks this example against it. gate.type: auto = orchestrator verifies the check itself (pass/fail); manual = wait for an explicit operator go. Which stages are manual vs auto is the operator's decision, not the plugin's. Any repeating loop in a run (fix loop, a re-entered stage, the per-module program loop) is bound by the loop guard: log every repeat touch, stop on oscillation, escalate to the layer that owns the conflict, then re-check in a planned order.",
5
5
  "stages": [
6
6
  {
7
7
  "id": 0,
8
8
  "state": "intake",
9
9
  "name": "Intake grill",
10
- "model": "claude-fable-5",
10
+ "model": "default",
11
11
  "skills": [
12
- "grill-me",
13
- "grilling"
12
+ "task-pipeline:grill"
14
13
  ],
15
14
  "gate": {
16
15
  "type": "manual",
17
- "check": "grill the operator one question at a time (recommended answer per question; explore codebase/docs before asking) until shared understanding is reached — every decision branch has a recorded answer or explicit deferral, no open contradictions; UI verdict recorded (arms super-ux); decisions locked into a committed task brief the operator confirms before stage 1"
16
+ "check": "MANDATORY stage — never skipped (only sanctioned bypass: the entry-from-super-ux short-circuit). The grill is built into the skill (references/grill.md) — no companion to install. Per its contract: one question at a time, a recommended answer with each, explore the codebase/docs before asking, depth-first, contradictions reconciled; domain awareness applied (terms challenged against CONTEXT.md, ADRs recorded for hard-to-reverse calls). The autonomy sweep is covered — every stage 1-10 has its blockers pre-resolved (docs sources, branch/tracker policy, test + lint commands, deploy target and authorization, log/health locations, docs+wiki targets) or is explicitly marked 'stop and ask here'. UI verdict recorded (arms super-ux); model decision recorded. All of it locked into a committed task brief the operator confirms before stage 1. The REQ table is written — one row per independently verifiable deliverable, each naming how it is verified — and frozen: adding later is free, removing or narrowing needs the operator's explicit agreement. The carry-over ledger is seeded."
18
17
  }
19
18
  },
20
19
  {
21
20
  "id": 1,
22
21
  "state": "docs-study",
23
22
  "name": "Docs study",
24
- "model": "claude-fable-5",
23
+ "model": "default",
25
24
  "skills": [
26
25
  "context7",
27
26
  "context7-docs"
@@ -34,91 +33,93 @@
34
33
  {
35
34
  "id": 2,
36
35
  "state": "brainstorm",
37
- "name": "Brainstorm",
38
- "model": "claude-fable-5",
36
+ "name": "Brainstorm + decompose",
37
+ "model": "default",
39
38
  "skills": [
40
- "superpowers:brainstorming"
39
+ "task-pipeline:brainstorm",
40
+ "task-pipeline:decompose"
41
41
  ],
42
42
  "gate": {
43
43
  "type": "manual",
44
- "check": "the user approves the design AND the UI verdict is recorded (does the task touch a user-facing surface — web/mobile/CLI/TUI? this arms the stage-3 UX track)"
44
+ "check": "the user approves the design AND the UI verdict is recorded (does the task touch a user-facing surface — web/mobile/CLI/TUI? this arms the stage-3 UX track). Every REQ is answered by the design, or explicitly dropped by the operator into the carry-over ledger. For a platform (several independent capabilities or shippable surfaces): the module map specs/<topic>-modules.md is committed and approved — brick criteria met or excepted in writing, dependency graph acyclic, build order topological with the walking skeleton first, every REQ mapped to exactly one module, cross-module contracts named with their owner. Single-module work records 'single module: <name>' instead — a skipped decomposition is a recorded decision, never an omission"
45
45
  }
46
46
  },
47
47
  {
48
48
  "id": 3,
49
49
  "state": "spec",
50
50
  "name": "Spec",
51
- "model": "claude-fable-5",
51
+ "model": "default",
52
52
  "skills": [
53
53
  "super-ux:ux-foundation",
54
54
  "super-ux:ux-flows",
55
55
  "super-ux:ux-scenarios",
56
- "superpowers:brainstorming"
56
+ "task-pipeline:spec"
57
57
  ],
58
58
  "gate": {
59
59
  "type": "manual",
60
- "check": "UX track ran FIRST for user-facing tasks (/ux -> ux-foundation CJM -> ux-flows screens -> ux-scenarios -> /ux-lint green); spec committed and user-reviewed; every user-facing requirement traces to a scenario ID"
60
+ "check": "UX track ran FIRST for user-facing tasks (/ux -> ux-foundation CJM -> ux-flows screens -> ux-scenarios -> /ux-lint green); spec committed and user-reviewed; every user-facing requirement traces to a scenario ID. Every spec section carries covers: REQ-... and every REQ appears in at least one section."
61
61
  }
62
62
  },
63
63
  {
64
64
  "id": 4,
65
65
  "state": "plan",
66
66
  "name": "Plan",
67
- "model": "claude-fable-5",
67
+ "model": "default",
68
68
  "skills": [
69
- "superpowers:writing-plans"
69
+ "task-pipeline:plan"
70
70
  ],
71
71
  "gate": {
72
72
  "type": "auto",
73
- "check": "every spec requirement maps to a task; no placeholders; parallel-group tasks share no files; UI tasks name the scenario ID(s) and SCR- screen(s) they implement in their DoD"
73
+ "check": "SET EQUALITY: the REQ ids in the brief equal the union of Implements: across plan tasks — a non-empty difference fails the gate and is reported as the explicit list of dropped requirements. Plus: every spec requirement maps to a task; no placeholders; names and types consistent across tasks; every task carries a verifiable DoD; parallel-group tasks share no files; UI tasks name the scenario ID(s) and SCR- screen(s) they implement in their DoD"
74
74
  }
75
75
  },
76
76
  {
77
77
  "id": 5,
78
78
  "state": "dev",
79
79
  "name": "Dev",
80
- "model": "claude-opus-5",
80
+ "model": "default",
81
81
  "skills": [
82
- "superpowers:using-git-worktrees",
83
- "superpowers:subagent-driven-development"
82
+ "task-pipeline:build",
83
+ "task-pipeline:review"
84
84
  ],
85
85
  "gate": {
86
86
  "type": "auto",
87
- "check": "all plan tasks DONE (two-stage review: spec compliance, then code quality); full test suite green"
87
+ "check": "all plan tasks DONE — the per-task review returns three verdicts (spec compliance, REQ satisfied, code quality); every finding fixed or parked with a written ruling; no task left BLOCKED; full test suite green; the branch integrated per the brief's branch policy (base synced, suite green on the result, worktree removed) — or the operator's explicit 'leave it unmerged' recorded. Every parked finding and implementer concern is harvested into the carry-over ledger before the scratch workspace is deleted."
88
88
  }
89
89
  },
90
90
  {
91
91
  "id": 6,
92
92
  "state": "tests",
93
93
  "name": "Tests",
94
- "model": "claude-opus-5",
94
+ "model": "default",
95
95
  "skills": [
96
- "superpowers:test-driven-development"
96
+ "host:test-runner",
97
+ "task-pipeline:tdd"
97
98
  ],
98
99
  "gate": {
99
100
  "type": "auto",
100
- "check": "full suite green (not just new tests); new/changed code covered; no skip/xfail smuggling a red suite past the gate"
101
+ "check": "full suite green (not just new tests); new/changed code covered including failure paths; no skip/xfail smuggling a red suite past the gate; tests assert real behavior, not mock behavior"
101
102
  }
102
103
  },
103
104
  {
104
105
  "id": 7,
105
106
  "state": "lint-deploy",
106
107
  "name": "Lint + deploy",
107
- "model": "inherit",
108
+ "model": "default",
108
109
  "skills": [
109
110
  "host:lint",
110
111
  "host:deploy"
111
112
  ],
112
113
  "gate": {
113
114
  "type": "manual",
114
- "check": "lint clean and full suite green before deploy; deploy is outward and needs explicit operator go"
115
+ "check": "lint clean and full suite green before deploy; deploy is outward and needs explicit operator go — or the specific standing authorization recorded in the stage-0 brief (named target + named preconditions; a vague 'do everything' is not one). No REQ is still open; a partial ships only with the operator's explicit acceptance."
115
116
  }
116
117
  },
117
118
  {
118
119
  "id": 8,
119
120
  "state": "post-deploy",
120
121
  "name": "Post-deploy",
121
- "model": "inherit",
122
+ "model": "default",
122
123
  "skills": [
123
124
  "host:health-check"
124
125
  ],
@@ -131,7 +132,7 @@
131
132
  "id": 9,
132
133
  "state": "docs-wiki",
133
134
  "name": "Docs + wiki",
134
- "model": "inherit",
135
+ "model": "default",
135
136
  "skills": [
136
137
  "host:module-docs",
137
138
  "wiki-update"
@@ -140,6 +141,19 @@
140
141
  "type": "auto",
141
142
  "check": "docs in sync with code in the same change; wiki synced; dangling links fixed"
142
143
  }
144
+ },
145
+ {
146
+ "id": 10,
147
+ "state": "acceptance",
148
+ "name": "Acceptance",
149
+ "model": "default",
150
+ "skills": [
151
+ "task-pipeline:acceptance"
152
+ ],
153
+ "gate": {
154
+ "type": "manual",
155
+ "check": "Close the circle: every REQ in the brief has a status (verified / partial / deferred / dropped) — none unknown; every verified carries evidence (a passing test name, file:line, a command and its output, or a scenario ID) — 'done' without evidence is downgraded to partial, not upgraded; every partial names what is missing and where it is tracked; every deferred/dropped has the operator's agreement and, for deferred, a tracker entry; no carry-over row is left unresolved; and the operator answers the closing question — here is what you asked for, here is what shipped, here is what is deferred, what is missing? — and signs off"
156
+ }
143
157
  }
144
158
  ],
145
159
  "release": {
@@ -45,7 +45,7 @@
45
45
  "id": { "type": "integer", "description": "Optional ordinal." },
46
46
  "state": { "type": "string", "minLength": 1, "description": "Unique stable key for the stage." },
47
47
  "name": { "type": "string", "description": "Optional human label." },
48
- "model": { "type": "string", "description": "Optional recommended model id for the stage." },
48
+ "model": { "type": "string", "description": "Optional model for the stage. Prefer a provider-agnostic token over a vendor id, which goes stale as generations ship and may not exist on the operator's provider at all: 'default' = the model confirmed for this run (recommended: the most capable reasoning model the environment offers), 'inherit' = whatever the operator is currently on. A literal id is allowed but treated as an example, not a contract." },
49
49
  "skills": {
50
50
  "type": "array",
51
51
  "minItems": 1,
@@ -0,0 +1,118 @@
1
+ # Acceptance — stage 10, built in
2
+
3
+ The pipeline is a funnel: every gate before this one asks *"is this artifact
4
+ good?"* — is the spec committed, does the plan parallelize, is the suite green.
5
+ None of them asks *"does this still contain everything that was asked for?"*
6
+
7
+ That is this stage's only job: **go back to the brief and account for every
8
+ requirement.** It is what turns the pipeline from a funnel into a circle.
9
+
10
+ ## Why a stage and not a gate
11
+
12
+ The loss this catches doesn't happen inside a stage — it happens **on the seams**.
13
+ Brief → spec → plan → task briefs is four rewrites by a model, and anything not
14
+ carried forward disappears silently because nothing compares the lists. Stage 4's
15
+ gate catches the brief→plan seam mechanically; stage 10 catches everything the
16
+ run itself decided, deferred, or quietly dropped along the way.
17
+
18
+ It runs **last** — after docs and wiki (stage 9), because those are deliverables
19
+ too and a requirement may name them.
20
+
21
+ ## Inputs
22
+
23
+ Read all of them before writing anything:
24
+
25
+ - the brief's **REQ table** (`docs/superpowers/specs/<topic>-brief.md`)
26
+ - the **carry-over ledger** (`…-carryover.md`) — in full, every row
27
+ - the plan and its task statuses
28
+ - git log for the run's branch; the test suite's final output
29
+ - stage 8's post-deploy notes; stage 9's doc/wiki changes
30
+ - for UI tasks: `docs/ux/scenarios.md` statuses and the `/ux-lint` result
31
+
32
+ ## Output — the coverage table
33
+
34
+ Write `docs/superpowers/specs/YYYY-MM-DD-<topic>-acceptance.md`:
35
+
36
+ ```markdown
37
+ # Acceptance — <topic>
38
+
39
+ Run: <branch/commit range> · Date: YYYY-MM-DD
40
+
41
+ | REQ | Requirement | Status | Evidence |
42
+ |---|---|---|---|
43
+ | REQ-001 | CSV export from a report | verified | `test_export_csv` ✓ · `api/export.ts:88` |
44
+ | REQ-002 | Export respects active filters | verified | `test_export_respects_filters` ✓ · SCN-014 PASS |
45
+ | REQ-003 | Button disabled on an empty report | deferred | agreed 2026-07-28 → LIN-482 |
46
+ | REQ-004 | XLSX export | partial | CSV path done; XLSX missing → LIN-483 |
47
+
48
+ ## Carry-over still open
49
+
50
+ - (rows from the ledger whose home is not an issue/backlog/`dropped`)
51
+
52
+ ## What the operator should look at
53
+
54
+ - <anything the run judged, guessed, or deferred that deserves a second opinion>
55
+ ```
56
+
57
+ ### The four statuses
58
+
59
+ | Status | Means | Requires |
60
+ |---|---|---|
61
+ | `verified` | done and demonstrated | **evidence** — a passing test name, `file:line`, command output, or a scenario ID with PASS |
62
+ | `partial` | works for some of what was asked | an explicit list of what's missing + where it's tracked |
63
+ | `deferred` | agreed not to do it now | the operator's agreement **and** a tracker entry |
64
+ | `dropped` | agreed it isn't wanted | the operator's agreement + the reason |
65
+
66
+ There is no fifth status. A requirement nobody can classify is `unknown`, and
67
+ `unknown` fails the gate — that is the whole mechanism.
68
+
69
+ ## Evidence, not assertion
70
+
71
+ **"Done" without evidence is not done.** This is the same rule the review rubric
72
+ and the test-honesty rules apply one level down, raised to the level of intent:
73
+
74
+ - A passing test **name**, not "tests pass".
75
+ - A `file:line`, not "implemented in the export module".
76
+ - A command **and its output**, not "verified manually".
77
+ - For user-facing behavior, the scenario ID and its status.
78
+
79
+ If the evidence for a requirement is "I read the code and it looks right", the
80
+ status is `partial`, not `verified` — say so plainly rather than upgrading it.
81
+
82
+ ## The closing question
83
+
84
+ The table is preparation. The stage exists for the question that follows it, asked
85
+ out loud, with the list in front of the operator:
86
+
87
+ > Here's what you asked for, here's what shipped, here's what's deferred and where
88
+ > it lives. **What's missing?**
89
+
90
+ Ask it even when the table is all green. The operator holds context the brief
91
+ never captured, and this is the cheapest moment in the whole run to hear it. An
92
+ answer here becomes new REQ rows or new ledger entries — not a new argument about
93
+ whether the run was finished.
94
+
95
+ ## GATE (manual)
96
+
97
+ All of:
98
+
99
+ 1. **Every REQ has a status** — none `unknown`, none blank.
100
+ 2. **Every `verified` carries evidence** of the kind above.
101
+ 3. **Every `partial` names what's missing** and where it's tracked.
102
+ 4. **Every `deferred` / `dropped` has the operator's agreement** recorded (in the
103
+ ledger or here) and, for `deferred`, a tracker entry.
104
+ 5. **No carry-over row is left `unresolved`** — every one has a home.
105
+ 6. **The operator answers the closing question** and signs off.
106
+
107
+ Manual by design. An automated check can prove the table is *well-formed*; only
108
+ the person who asked can confirm it is *what they asked for*. Do not let a green
109
+ table substitute for that answer.
110
+
111
+ ## When the answer is "something's missing"
112
+
113
+ Don't argue and don't re-litigate the gates. Add the missing thing as a new REQ
114
+ row (with its check) or a ledger entry, then say plainly what it costs: a fix now,
115
+ or a tracked follow-up. Both are legitimate outcomes of this stage. Closing the
116
+ run with a known gap is fine **if the gap is written down** — closing it with the
117
+ gap only in someone's memory is the failure mode this whole spine exists to
118
+ prevent.