task-pipeline-skill 0.10.0 → 0.17.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (30) hide show
  1. package/CHANGELOG.md +401 -0
  2. package/LICENSE +85 -0
  3. package/README.md +211 -84
  4. package/cursor/rules/task-pipeline.mdc +135 -20
  5. package/package.json +3 -3
  6. package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
  7. package/plugins/task-pipeline/commands/task-pipeline.md +16 -8
  8. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +139 -57
  9. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +41 -27
  10. package/plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json +1 -1
  11. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
  12. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +32 -6
  13. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
  14. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
  15. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +73 -21
  16. package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
  17. package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
  18. package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +169 -0
  19. package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
  20. package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +55 -22
  21. package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
  22. package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
  23. package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
  24. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +184 -48
  25. package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
  26. package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +11 -3
  27. package/plugins/task-pipeline/skills/task-pipeline/templates/adr.md +64 -0
  28. package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +49 -0
  29. package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
  30. package/plugins/task-pipeline/skills/task-pipeline/templates/context.md +87 -0
@@ -1,46 +1,98 @@
1
- # Companion skills — what powers each stage, how to install, what to run
1
+ # Companions — what's built in, what's optional, what to install
2
2
 
3
- task-pipeline is a thin orchestrator; the actual work is done by companion
4
- skills. Preflight-detect each one; if a needed skill doesn't resolve, **give the
5
- operator the install line immediately** and (for required ones) stop until it's
6
- installed. Never silently degrade a required capability.
3
+ **The pipeline's doctrine is entirely built into this skill.** Stages 0, 2, 3, 4, 5,
4
+ 6 and 10 run from `references/*.md` — no companion plugin, no resolution step, no
5
+ fallback path, no version skew, and no failure mode where a stage can't run because
6
+ something isn't installed.
7
+
8
+ What remains is a short list of **optional** companions that make individual stages
9
+ better, plus one that is required only for user-facing work.
10
+
11
+ ## Built in — nothing to install
12
+
13
+ | Stage | Doctrine |
14
+ |---|---|
15
+ | 0 Intake grill | `references/grill.md` |
16
+ | 2 Brainstorm | `references/brainstorm.md` |
17
+ | 2 Decompose (platforms only) | `references/decomposition.md` |
18
+ | 3 Spec | `references/spec.md` |
19
+ | 4 Plan | `references/planning.md` |
20
+ | 5 Build (isolation, subagents, fix loop) | `references/build.md` + `references/review.md` |
21
+ | 5–6 TDD + suite gate | `references/tdd.md` |
22
+ | 10 Acceptance (REQ close-out) | `references/acceptance.md` |
23
+ | any repeating loop | `references/loop-guard.md` |
7
24
 
8
25
  ## The matrix
9
26
 
10
27
  | Skill / tool | Needed for | Required? | Install |
11
28
  |---|---|---|---|
12
- | **superpowers** (`brainstorming`, `writing-plans`, `subagent-driven-development`, `using-git-worktrees`, `test-driven-development`) | stages 2, 4, 5, 6 | **Required** (always) | `/plugin marketplace add obra/superpowers` → `/plugin install superpowers@superpowers` |
13
29
  | **super-ux** (`ux-foundation`, `ux-flows`, `ux-scenarios`, `ux-audit`, `/ux`, `/ux-lint`) | stage 3 UX track | **Required for any user-facing task** | `/plugin marketplace add ssheleg/super-ux` → `/plugin install super-ux@super-ux` (or `npx skills add ssheleg/super-ux`) |
14
- | **grill-me** / **grilling** | stage 0 intake grill | Optional (built-in grill loop is the fallback) | `npx skills add mattpocock/skills`, or the engineering-advanced-skills marketplace |
15
30
  | **context7** (MCP) | stage 1 docs study | Recommended (web-search fallback) | connect the context7 MCP server |
16
31
  | **wiki-update** | stage 9 wiki sync | Optional (skip wiki if absent) | user's wiki skill set |
32
+ | ~~superpowers~~ | — | **Not a dependency.** Stages 2/4/5/6 run on the built-in doctrine above. See *Optional bridge* | — |
33
+ | ~~grill-me / grilling~~ | — | **Not a dependency.** The stage-0 grill is built in (`references/grill.md`) | — |
34
+
35
+ ## Optional bridge — substituting an external skill set
36
+
37
+ An operator who already runs an equivalent skill set may map it onto stages 2/4/5/6
38
+ in their `pipeline.json` → `skills[]`, e.g. `superpowers:brainstorming`,
39
+ `superpowers:writing-plans`, `superpowers:using-git-worktrees`,
40
+ `superpowers:subagent-driven-development`, `superpowers:test-driven-development`.
41
+
42
+ Rules for that bridge:
43
+
44
+ - **It is a substitution, never a requirement.** Nothing detects it, nothing
45
+ recommends it, nothing waits for it, and its absence is never an error.
46
+ - **The gates still govern.** Whatever runs a stage, `stages.md` decides when the
47
+ stage is done.
48
+ - **Never mix providers inside one stage** — either the built-in doctrine runs it or
49
+ the substitute does; interleaving two review loops produces neither.
17
50
 
18
- ## Preflight recommendation (emit before stage 0)
51
+ ## Preflight (emit before stage 0)
19
52
 
20
- At the very start, detect which of the above resolve and print ONE recommendation
21
- block so the operator can arm the full flow before work begins. Example:
53
+ Detect the optional companions and print ONE block — companions **plus the model
54
+ decision** (`model-tiering.md`), so the operator arms the whole run in a single
55
+ exchange:
22
56
 
23
57
  ```
24
- Pipeline companions:
25
- ✓ superpowers — ready
26
- ✗ super-ux — this task looks user-facing; recommended. Install:
58
+ Pipeline companions (stage doctrine is built in — nothing to install for it):
59
+ ✗ super-ux — this task looks user-facing; required for the UX track:
27
60
  /plugin marketplace add ssheleg/super-ux
28
61
  /plugin install super-ux@super-ux
29
62
  ✓ context7 — ready
30
- ✗ grill-me — optional; falling back to the built-in grill loop
31
63
  ✓ wiki-update — ready
32
- Recommend installing the ✗ items marked recommended, then say "continue".
64
+
65
+ 🧠 Model for this run: recommended <top tier available>. You're on <current>.
66
+ /model <id> to switch, or "keep current", or name per-stage overrides.
67
+
68
+ Install the ✗ items you want, answer the model line, then say "continue".
33
69
  ```
34
70
 
35
71
  Rules:
36
- - Only flag **super-ux** as recommended when the task implies a UI (the stage-0
37
- grill decides this; when unsure, flag it — a false positive costs one install).
38
- - **superpowers** missing → stop; it's required for the core stages.
72
+
73
+ - Only flag **super-ux** when the task implies a UI (the stage-0 grill decides;
74
+ when unsure, flag it — a false positive costs one install).
75
+ - **Never gate any stage on an install** except the stage-3 UX track on a UI task.
39
76
  - Optional tools missing → state the fallback, don't block.
40
77
  - Re-detect after the operator installs; don't assume.
78
+ - The model answer goes into the brief. Don't ask again per stage
79
+ (`model-tiering.md` → *Mechanic*).
80
+
81
+ ## Credit
82
+
83
+ The built-in doctrine is **ported, not depended on**:
84
+
85
+ - The stage-0 grill is adapted from Matt Pocock's `grilling` / `grill-with-docs`
86
+ skills (MIT, https://github.com/mattpocock/skills).
87
+ - Stages 2–6 are adapted from the `brainstorming`, `writing-plans`,
88
+ `using-git-worktrees`, `subagent-driven-development`, `test-driven-development`
89
+ and `requesting-code-review` skills in obra/superpowers (MIT,
90
+ https://github.com/obra/superpowers).
91
+
92
+ Both notices live in the repo `LICENSE` → *Third-party*.
41
93
 
42
94
  ## Hand-off the other direction
43
95
 
44
- super-ux's `/ux` menu can hand off *to* this pipeline (its "execute
45
- autonomously" action). When entered that way the UX chain already exists — see
46
- `stages.md` → 0 *Entry-from-super-ux short-circuit*: verify, don't rebuild.
96
+ super-ux's `/ux` menu can hand off *to* this pipeline (its "execute autonomously"
97
+ action). When entered that way the UX chain already exists — see `stages.md` → 0
98
+ *Entry-from-super-ux short-circuit*: verify, don't rebuild.
@@ -1,4 +1,4 @@
1
- # Host conventions (stages 6–9)
1
+ # Host conventions (stages 6–10)
2
2
 
3
3
  The orchestrator is project-agnostic. For tests / lint / deploy / docs / wiki it reads the
4
4
  **host project's `CLAUDE.md` / `AGENTS.md` first**, then falls back to detection.
@@ -33,3 +33,13 @@ found, surface it and **ask** rather than guessing.
33
33
  - Host self-update rules (module docs, runbooks, agent-self cards, etc.) — update
34
34
  in the same change. Wiki: the `wiki-update` skill (resolves the vault via
35
35
  `~/.obsidian-wiki/config`). Fix dangling links.
36
+
37
+ ## Issue tracker (stage 10)
38
+
39
+ Acceptance parks what wasn't delivered: every `deferred` REQ and every unresolved
40
+ carry-over row needs a **home** — an issue, a backlog entry, a ticket id. Read the
41
+ host's convention (`CLAUDE.md` usually names the tracker and the id format; else
42
+ detect: a `.github/ISSUE_TEMPLATE/`, a Linear/Jira reference in recent commits, a
43
+ `TODO.md`). Never invent a tracker, and never close a run on "we'll remember it" —
44
+ if no tracker exists, write the row into the repo's backlog file and say where it
45
+ went.
@@ -0,0 +1,139 @@
1
+ # Decomposition — cutting a platform into bricks
2
+
3
+ A one-feature task goes through the pipeline once. A **platform** — anything whose
4
+ brief describes more than one deliverable, more than one surface, or a system
5
+ rather than a change — must be cut into modules first, and then built one brick at
6
+ a time, each brick carrying its own documentation, spec, plan, build and gates.
7
+
8
+ This runs at the end of **stage 2**, on the approved design, before any spec is
9
+ written. It is skipped — explicitly, in writing — when the work is a single module.
10
+
11
+ ## When it applies
12
+
13
+ Decompose when any of these is true:
14
+
15
+ - the brief names several independent capabilities ("accounts, billing, reporting");
16
+ - the work spans several surfaces (API + web + worker) that could ship separately;
17
+ - the REQ table has requirements that no single deliverable satisfies together;
18
+ - the design's units have their own data and could plausibly be owned by different
19
+ people.
20
+
21
+ Otherwise record one line in the design — `single module: <name>` — and go to
22
+ stage 3. A skipped decomposition is a decision, never an omission.
23
+
24
+ ## How to cut
25
+
26
+ **By capability, not by layer.** "Ordering", "Billing", "Notifications" are
27
+ modules. "Controllers", "Services", "Database" are not: a layer cut forces every
28
+ feature to touch every module, which is the opposite of a brick.
29
+
30
+ A module is a **brick** when all of these hold:
31
+
32
+ 1. **Independently specifiable** — you can write its dossier without deciding
33
+ another module's internals.
34
+ 2. **Independently buildable and testable** — its tests pass without another
35
+ module's implementation present (stubs at the contract are fine).
36
+ 3. **Owns its data** — the entities it is the source of truth for belong to it, and
37
+ nothing else writes them.
38
+ 4. **Talks through declared contracts only** — every cross-module interaction is a
39
+ named API, event or schema, listed in both modules' dossiers.
40
+ 5. **Deliverable on its own** — landing it leaves the system working, even if the
41
+ capability is not yet reachable by users.
42
+
43
+ If a candidate fails (2) or (3), the cut is in the wrong place: either merge it
44
+ into its neighbor or move the disputed data to the module that truly owns it.
45
+
46
+ **Order the bricks:**
47
+
48
+ - **The walking skeleton first.** The first module is the thinnest end-to-end slice
49
+ that proves the architecture — one real path through the system, however small.
50
+ Building three "foundation" modules before anything runs end-to-end hides
51
+ integration risk until the worst possible moment.
52
+ - Then topological order: nothing is built before what it depends on.
53
+ - **No cycles.** A cycle means the cut is wrong. Break it by moving the shared
54
+ concept into its own module, or by turning one direction of the dependency into
55
+ an event the other module subscribes to. Record which you chose and why.
56
+
57
+ ## The module map — the artifact
58
+
59
+ Write `docs/superpowers/specs/YYYY-MM-DD-<topic>-modules.md` and commit it. It is
60
+ the program's spine: every later run reads it, and its status column is how a
61
+ resumed session knows where the program stopped.
62
+
63
+ ```markdown
64
+ # Module map — <platform>
65
+
66
+ Build order is top to bottom. Status: `planned` → `in progress` → `done` |
67
+ `deferred`. One row per module, no exceptions.
68
+
69
+ | # | Module | Delivers | Owns (entities) | Depends on | Contracts exposed | UI? | REQs | Status |
70
+ |---|---|---|---|---|---|---|---|---|
71
+ | 1 | ordering | place and track an order | Order, OrderLine | — | `POST /orders`, `OrderPlaced` event | yes | REQ-001, REQ-004 | planned |
72
+ | 2 | billing | charge for a placed order | Invoice, Payment | ordering | `InvoiceIssued` event | no | REQ-002 | planned |
73
+
74
+ ## Cut rationale
75
+
76
+ <why these seams and not others; what was merged or split, and what a cycle forced>
77
+
78
+ ## Cross-module contracts
79
+
80
+ <one block per contract: owner module, consumer(s), exact shape (schema or
81
+ signature), and the failure behavior when the other side is unavailable>
82
+
83
+ ## Deferred to later modules
84
+
85
+ <capabilities deliberately postponed, with the module that will carry them>
86
+ ```
87
+
88
+ Every REQ from the brief appears in exactly one module's `REQs` cell. A REQ that
89
+ fits nowhere means the map is incomplete; a REQ in two modules means the seam runs
90
+ through a requirement — re-cut or split the REQ.
91
+
92
+ ## GATE (part of stage 2, manual)
93
+
94
+ Together with the design approval:
95
+
96
+ 1. Every module satisfies the brick criteria, or its exception is written down.
97
+ 2. The dependency graph is acyclic and the build order is topological.
98
+ 3. The first module is a walking skeleton, or the reason it isn't is recorded.
99
+ 4. Every REQ maps to exactly one module.
100
+ 5. Cross-module contracts are named (shape can be locked later, in each module's
101
+ spec — but the *existence* and *owner* of each contract is decided here).
102
+ 6. The operator approves the map and the order.
103
+
104
+ ## The program loop — one brick at a time
105
+
106
+ After the map is approved, the pipeline runs **per module**, in build order:
107
+
108
+ ```
109
+ module N → stage 3 (dossier/spec) → 4 plan → 5 build → 6 tests
110
+ → 7 lint + deploy → 8 post-deploy → 9 docs + wiki → 10 acceptance
111
+ → mark module done → module N+1 (back to stage 3)
112
+ ```
113
+
114
+ Rules for the loop:
115
+
116
+ - **Stages 0–2 run once for the platform.** Modules do not re-grill and do not
117
+ re-decompose. New information that changes the map goes back to stage 2
118
+ deliberately, as a map revision with the operator's approval — not as a quiet
119
+ edit mid-module.
120
+ - **Each module's spec is a full dossier** ([`spec.md`](spec.md)): architecture,
121
+ entities, contracts in and out, business rules, edge and failure cases, UI/Figma
122
+ chain when it has a surface.
123
+ - **The contract is the boundary.** A module may stub what a later module will
124
+ provide, but it may not reach into another module's internals; if it needs to,
125
+ the seam is wrong — back to the map.
126
+ - **Deploy cadence is the brief's call** (autonomy sweep): deploy each module as it
127
+ lands, or build several and deploy once. Record it; don't decide it per module.
128
+ - **Update the map's status column as each module closes**, in the same commit as
129
+ that module's acceptance. The map is the resume point after a lost context.
130
+ - **Loop discipline:** a module re-entering the same stage a third time trips the
131
+ loop guard ([`loop-guard.md`](loop-guard.md)) — stop, name the oscillation, and
132
+ fix the layer that owns it instead of iterating.
133
+
134
+ ## Program done
135
+
136
+ The program is finished when every row is `done` or `deferred` with an agreed home,
137
+ the cross-module contracts are exercised by tests that cross the seam (not just
138
+ per-module unit tests), and the final acceptance covers the platform's REQ table as
139
+ a whole — not module by module.
@@ -0,0 +1,169 @@
1
+ # The grill — stage 0, built in
2
+
3
+ The intake grill is **part of this skill**. No companion skill to install, no
4
+ provider to resolve, nothing to fall back to: this file *is* the implementation.
5
+
6
+ Its job is not to design. It is to take a one-line request ("make me feature X")
7
+ and interview it into a brief complete enough that stages 1→10 finish without
8
+ coming back to the operator.
9
+
10
+ > Adapted, with thanks, from Matt Pocock's `grilling` / `grill-with-docs` skills
11
+ > (MIT — see this repo's `LICENSE` → *Third-party*). The domain-awareness
12
+ > half — glossary challenges, `CONTEXT.md`, ADR discipline — comes from there; the
13
+ > autonomy sweep and the brief are this pipeline's.
14
+
15
+ ## The loop
16
+
17
+ Interview the operator relentlessly about every aspect of the task until you reach
18
+ a **shared understanding**. Walk down each branch of the decision tree, resolving
19
+ dependencies between decisions one by one.
20
+
21
+ 1. **One question per turn.** Never bundle. Wait for the answer before the next.
22
+ 2. **Recommend an answer with every question** (+ a one-line rationale). "What do
23
+ you think?" is lazy — you have the codebase in front of you, they don't.
24
+ 3. **If the codebase can answer it, go read the codebase.** Spending the
25
+ operator's turn on something `grep`/`Read`/context7 would have told you is the
26
+ most common way to waste a grill.
27
+ 4. **Depth-first.** Finish a branch before opening another; ask prerequisite
28
+ decisions first, so later answers don't invalidate earlier ones.
29
+ 5. **Reconcile contradictions immediately**, and chase dodges: "we'll decide
30
+ later" → "what's the latest you can decide and still ship?"
31
+ 6. **Cover the autonomy sweep** (below). An unasked question is not neutral — it
32
+ is a scheduled interruption at stage 6.
33
+
34
+ **Stop** when a re-scan surfaces no new branches. Don't grill past diminishing
35
+ returns: genuinely reversible calls can be deferred with a note.
36
+
37
+ ## Domain awareness
38
+
39
+ While exploring the codebase, also look for what the project already says about
40
+ itself — and hold the operator to it.
41
+
42
+ ### Find the existing docs
43
+
44
+ Most repos have a single context:
45
+
46
+ ```
47
+ /
48
+ ├── CONTEXT.md
49
+ ├── docs/adr/
50
+ │ ├── 0001-event-sourced-orders.md
51
+ │ └── 0002-postgres-for-write-model.md
52
+ └── src/
53
+ ```
54
+
55
+ A `CONTEXT-MAP.md` at the root means multiple contexts, and points at where each
56
+ one lives (`src/ordering/CONTEXT.md`, `src/billing/CONTEXT.md`, …), each with its
57
+ own `docs/adr/` alongside the system-wide one. Infer which context the task
58
+ belongs to; if it's genuinely unclear, ask.
59
+
60
+ Create these files **lazily** — only when you have something real to write.
61
+
62
+ ### Techniques during the session
63
+
64
+ - **Challenge against the glossary.** When a term conflicts with `CONTEXT.md`, say
65
+ so on the spot: *"Your glossary defines 'cancellation' as X, but you seem to mean
66
+ Y — which is it?"*
67
+ - **Sharpen fuzzy language.** Vague or overloaded terms get a proposed canonical
68
+ one: *"You're saying 'account' — do you mean the Customer or the User? Those are
69
+ different things."*
70
+ - **Stress-test with concrete scenarios.** Invent specific cases that probe edge
71
+ conditions and force precision about the boundaries between concepts.
72
+ - **Cross-reference with the code.** When the operator states how something works,
73
+ check whether the code agrees, and surface contradictions: *"Your code cancels
74
+ entire Orders, but you just said partial cancellation is possible — which is
75
+ right?"*
76
+ - **Update `CONTEXT.md` inline.** Resolve a term → write it down right then, not in
77
+ a batch at the end. Format: [`templates/context.md`](../templates/context.md).
78
+ Keep it free of implementation detail — only terms a domain expert would
79
+ recognize.
80
+
81
+ ### Offer an ADR sparingly
82
+
83
+ Only when **all three** are true:
84
+
85
+ 1. **Hard to reverse** — changing your mind later carries real cost.
86
+ 2. **Surprising without context** — a future reader will ask "why on earth this
87
+ way?"
88
+ 3. **A real trade-off** — genuine alternatives existed and one was chosen for
89
+ specific reasons.
90
+
91
+ Any one missing → skip it. Format and what qualifies:
92
+ [`templates/adr.md`](../templates/adr.md). ADRs land in `docs/adr/` with sequential
93
+ numbering (scan for the highest number, increment).
94
+
95
+ ## The autonomy sweep
96
+
97
+ Resolving the *task* is not enough. The grill must also pre-resolve everything that
98
+ would otherwise stop stages 1→10 mid-flight. Every row gets an answer **or** an
99
+ explicit "stop and ask me here":
100
+
101
+ | Stage | What to settle up front |
102
+ |---|---|
103
+ | run-wide | the model decision ([`model-tiering.md`](model-tiering.md)); what to decide autonomously vs escalate |
104
+ | 1 Docs | external libs/APIs/SDKs in play; any private ones context7 can't resolve → where their docs live |
105
+ | 2 Decompose | is this a platform (several capabilities/surfaces) or one module? if platform: deploy cadence — per module or once at the end |
106
+ | 2–3 Spec | UI verdict (arms super-ux); any scenario-tracing waiver |
107
+ | 4–5 Dev | base branch; worktree/branch policy; is `main` off-limits; commit convention; task tracker |
108
+ | 5 Integration | how the branch lands (merge / PR + approver / "leave it unmerged"); parallel fan-out wanted (one worktree per implementer)? |
109
+ | 6 Tests | the test command; what "green" means here; known-red baseline; coverage expectation |
110
+ | 7 Lint+deploy | lint command; deploy target and path; release automation on/off; deploy-from-main rule; **deploy authorization** |
111
+ | 8 Post-deploy | where logs / health live (app name, endpoint, workflow) |
112
+ | 9 Docs+wiki | which module docs / runbooks this change updates; wiki sync yes/no |
113
+ | 10 Acceptance | who signs off; where deferred REQs are tracked (issue tracker, backlog) |
114
+
115
+ **Deploy authorization has a hard floor.** Deploy and publish are outward and
116
+ irreversible, so a vague "just do everything" authorizes nothing. A standing
117
+ authorization counts only when it is **specific** — named target, named
118
+ preconditions ("staging once lint and the full suite are green; production always
119
+ asks"). Specific and recorded → it satisfies the stage-7 manual gate. Broader,
120
+ absent or ambiguous → stage 7 stops and asks.
121
+
122
+ ## The REQ spine — the grill's other hard output
123
+
124
+ Prose scope is not checkable. Before the brief is confirmed, the grill must turn
125
+ what was asked into an **addressable list of requirements**, because every later
126
+ stage traces to these IDs and stage 10 accounts for every one of them.
127
+
128
+ | ID | Requirement | How it's verified | Status |
129
+ |---|---|---|---|
130
+ | REQ-001 | … | test name / `file:line` / command + expected output / `SCN-…` | open |
131
+
132
+ Three rules that decide whether the spine is worth anything:
133
+
134
+ 1. **One REQ = one independently verifiable deliverable.** Not one per sentence of
135
+ the request. A small task gets three rows, not thirty — an inflated table is
136
+ ignored, and an ignored table protects nothing.
137
+ 2. **Every row names its check.** *A requirement you can't say how to verify is a
138
+ badly-stated requirement* — split or sharpen it here, during the grill. This is
139
+ the single defence against the failure mode where three vague REQs cover a large
140
+ task and acceptance goes green over half of it.
141
+ 3. **Ask what "finished" means per row, not for the task overall.** "Export works"
142
+ hides five decisions; "exports the currently filtered rows as CSV, verified by
143
+ `test_export_respects_filters`" hides none.
144
+
145
+ **Then freeze it.** Adding a requirement mid-run is fine — append with its source.
146
+ **Removing or narrowing one requires the operator's explicit agreement**, recorded
147
+ in the carry-over ledger. Quietly restating the task in smaller terms is the
148
+ subtlest way to lose it: every gate downstream then passes honestly, on a task
149
+ that shrank without anyone deciding it should.
150
+
151
+ ## Output
152
+
153
+ Everything resolved goes into the **task brief**, seeded from
154
+ [`templates/brief.md`](../templates/brief.md) and committed to
155
+ `docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` — scope, **the REQ table**,
156
+ users, UI verdict, constraints, locked decisions, the autonomy table,
157
+ done-criteria, open assumptions. Seed the template only when the file is absent;
158
+ never overwrite an existing brief.
159
+
160
+ Alongside it, seed the **carry-over ledger** from
161
+ [`templates/carryover.md`](../templates/carryover.md) at
162
+ `…-carryover.md` — append-only, written by every later stage, read in full by
163
+ stage 10. Anything deferred, dropped, or half-done from here on goes there the
164
+ moment it's said: **deferred out loud is forgotten.**
165
+
166
+ Plus, where the session produced them: an updated `CONTEXT.md` and any ADRs, each
167
+ written as the decision landed.
168
+
169
+ The operator confirms the brief. Only then does stage 1 begin.
@@ -0,0 +1,100 @@
1
+ # Loop guard — breaking churn, cross-cutting
2
+
3
+ Any stage that can repeat can also **churn**: a later pass undoing what an earlier
4
+ pass in the same run already did, two shapes alternating, the same file rewritten
5
+ round after round with no new information. Churn looks like progress and consumes
6
+ a run.
7
+
8
+ This file is the detector and the break protocol. It binds every repeating loop in
9
+ the pipeline: the stage-5 fix loop, a stage re-entered after a failed gate, the
10
+ per-module program loop ([`decomposition.md`](decomposition.md)), and any
11
+ audit → fix → audit cycle.
12
+
13
+ ## Bookkeeping — the thing that makes detection mechanical
14
+
15
+ You cannot detect churn from memory, especially after compaction. Every repeating
16
+ pass appends one line to the run's ledger (`.task-pipeline/build/<plan>/progress.md`
17
+ for stage 5; `.task-pipeline/run.md` for stage-level and program-level loops):
18
+
19
+ ```
20
+ touch: <file> — pass <N> (<stage|round|module>) — reason: <finding id / gate item>
21
+ ```
22
+
23
+ One line per file per pass. The reason must name **what forced the edit** — a
24
+ finding id, a failed gate item, an operator instruction. "Cleanup", "polish" and
25
+ "while I was there" are not reasons; they are churn with better manners.
26
+
27
+ ## Detection — any one of these trips the guard
28
+
29
+ 1. **Revert-oscillation.** An edit restores something an earlier pass in this run
30
+ deliberately removed, or re-removes what an earlier pass added. Shape A → B → A.
31
+ 2. **Repeat touch without new information.** The same file is edited in two
32
+ consecutive passes and the second pass's `reason` is the same finding/gate item
33
+ as the first — the fix did not fix it, or the two passes disagree about what
34
+ "fixed" means.
35
+ 3. **Finding resurrection.** A finding whose text (normalized) matches one already
36
+ marked ADDRESSED or parked-with-ruling in this run comes back.
37
+ 4. **Gate ping-pong.** The same stage is re-entered for the third time on the same
38
+ artifact, or two adjacent stages hand work back and forth (spec ⇄ plan,
39
+ plan ⇄ build) more than twice.
40
+ 5. **Cross-loop contradiction.** A pass in one loop edits a file that a *different*
41
+ loop (another task, another module) already closed in this run — two owners for
42
+ one file.
43
+
44
+ Caps that trip the guard by themselves: **5 fix rounds** per task
45
+ ([`build.md`](build.md)), **2 re-entries** per stage per artifact, **3 passes** per
46
+ module in the program loop.
47
+
48
+ ## The break protocol
49
+
50
+ When the guard trips, **stop editing immediately**. Do not dispatch another fix, do
51
+ not "just try one more thing". Then, in this order:
52
+
53
+ 1. **Freeze and name it.** Write the oscillation down in the ledger and to the
54
+ operator: shape **A** vs shape **B**, one line each, plus who is asking for each
55
+ (a finding, the plan's text, the spec, a gate check, an operator instruction) and
56
+ the evidence for each — `file:line`, the failing command, the review verdict.
57
+ 2. **Find the layer that owns the conflict.** Churn almost always means a decision
58
+ is being re-litigated at the wrong altitude:
59
+ - two findings disagree → the **review rubric** decides
60
+ ([`review.md`](review.md)); if it genuinely doesn't, it's a spec question;
61
+ - a finding contradicts the plan → the **operator** decides which governs
62
+ (never dismiss the finding, never fix against the plan silently);
63
+ - the plan contradicts the spec → back to **stage 4** with the evidence;
64
+ - the spec is ambiguous or wrong → back to **stage 3**, and if the ambiguity was
65
+ an unresolved intake question, say so — that is a stage-0 miss worth recording;
66
+ - two modules claim the same file or entity → back to **decomposition**: the cut
67
+ is wrong.
68
+ **Never resolve a higher-layer conflict inside a lower loop.** Patching code to
69
+ satisfy two contradictory requirements is how a run burns its remaining budget.
70
+ 3. **Re-plan the check.** Replace whatever ad-hoc verification was running with an
71
+ explicit ordered checklist: every disputed item, one line each, in dependency
72
+ order, with a single owner and a single verification command per item. Write it
73
+ to the ledger before touching anything.
74
+ 4. **Go in order, one at a time.** Verify item 1 → if it fails, fix only item 1 →
75
+ re-verify only item 1 → commit → item 2. No parallel edits, no bundled fixes, no
76
+ opportunistic cleanup in the same commit. The point is that each change has one
77
+ reason and one proof.
78
+ 5. **Re-check the whole list once** at the end, in the same order. If a later item
79
+ broke an earlier one, that pair is the real conflict — escalate it per step 2
80
+ instead of looping again.
81
+ 6. **Record the ruling.** Ledger line: `loop-guard: <A vs B> — ruling: <what governs
82
+ and why> — items: <N> verified in order`. The final review reads it.
83
+
84
+ ## When to stop and hand back
85
+
86
+ If step 2 lands on "the operator decides", or a cap is hit a second time after a
87
+ re-planned pass, **stop and report BLOCKED** with: the two shapes, the evidence, the
88
+ history of passes, and your recommendation. That is a complete, honest hand-back —
89
+ far cheaper than a third round of the same argument.
90
+
91
+ ## Rationalizations
92
+
93
+ | Excuse | Reality |
94
+ |---|---|
95
+ | "One more pass and it converges" | Two passes with the same reason already proved it doesn't. The disagreement is above the code. |
96
+ | "I'll just revert to what worked" | That is the oscillation, not the exit. Name A and B first. |
97
+ | "The reviewer keeps changing its mind" | Different findings on the same lines mean the requirement is ambiguous. That's a spec question. |
98
+ | "Tidying while I'm in the file" | Untracked edits are what make churn invisible. One reason per change, in the ledger. |
99
+ | "Logging the loop is bureaucracy" | Detection needs a record; after compaction the ledger is the only memory that survives. |
100
+ | "It's faster than escalating" | A run that spends its budget re-deciding a spec question delivers nothing. Escalation costs one message. |
@@ -1,32 +1,65 @@
1
- # Model tiering
1
+ # Model policy
2
2
 
3
- A **reminder**, not a hard block. Not every environment has every model — if you
4
- lack one, keep your current model; the pipeline still runs.
3
+ **One model, confirmed once, before the run starts.** Not a per-stage tier list,
4
+ not a hardcoded vendor id — a single decision the operator makes at preflight and
5
+ the pipeline then honors without nagging.
5
6
 
6
- | Stages | Recommended | id |
7
- |---|---|---|
8
- | 0–4 (intake grill, docs, brainstorm, spec, plan) | Fable 5 | `claude-fable-5` |
9
- | 5–6 (subagent dev, tests) | Opus 5 | `claude-opus-5` |
10
- | 7–9 (lint/deploy, logs, docs) | inherit current | — |
7
+ ## The default
11
8
 
12
- ## Mechanic
9
+ > **Use the most capable reasoning model the environment offers** — at the time of
10
+ > writing that is the **latest Opus generation**, but read that as *"the top tier
11
+ > of whatever you're on"*, not as a specific string.
13
12
 
14
- At each stage boundary compare recommended vs current. If they differ, emit:
13
+ Every stage runs on that model by default. The pipeline is a full delivery cycle:
14
+ the grill has to hear what the operator didn't say, the spec has to lock contracts
15
+ a zero-context implementer will follow, and the build has to hold a plan in its
16
+ head. Downgrading any of those to save tokens costs more in rework than it saves.
15
17
 
16
- > ⏸ **Stage N (`<name>`) recommends `<model>` (`<id>`).** You're on `<current>`.
17
- > Switch: `/model <id>` — then say "continue". *(Reminder only.)*
18
+ ## Never hardcode a model id
18
19
 
19
- ## Why manual
20
+ Model ids go stale — generations ship, tiers get renamed, and the operator may not
21
+ even be on the same provider. So:
20
22
 
21
- A skill runs inside the current context; it **cannot change the main-loop model**.
22
- Only the operator can, via `/model` (or `/fast`). The stages that most benefit from
23
- Fable (0–4) are interactive anyway (stage 0 is a live grill), so the operator is
24
- present to switch.
23
+ - **Resolve at runtime.** Look at what the environment actually offers (`/model`,
24
+ the harness's model list) and pick the top reasoning tier available there.
25
+ - **Treat any id in this repo as an example**, including in `pipeline.example.json`.
26
+ Stage configs use provider-agnostic tokens:
27
+ - `default` — the model confirmed for this run (the recommendation above)
28
+ - `inherit` — whatever the operator is currently on; no recommendation
29
+ - **Another provider is fine.** "Top tier available" is the contract. If the
30
+ environment has no Opus-class model, the best available one is the right answer —
31
+ say which one you settled on and keep going.
25
32
 
26
- Stage 5 spawns subagents; those **are** pinned to Opus by the orchestrator via the
27
- `Agent` / `Workflow` model override — no operator action needed for subagents.
33
+ ## Mechanic — confirm at preflight, then stop asking
28
34
 
29
- ## Override
35
+ Once, as part of the preflight (before stage 0):
30
36
 
31
- Set your own map if your task warrants it (e.g. a heavy design needs Opus at
32
- stage 2). The recommendations are defaults, not rules.
37
+ > 🧠 **Model for this run:** recommended **`<top tier available>`**. You're on
38
+ > `<current>`.
39
+ > Switch with `/model <id>`, or say "keep current" / name another. Per-stage
40
+ > overrides welcome (e.g. a cheaper model for mechanical stages) — say so now and
41
+ > I'll record the map.
42
+
43
+ Record the answer in the stage-0 brief (`Model` row of the autonomy sweep). After
44
+ that:
45
+
46
+ - **Do not re-prompt at every stage boundary.** The decision is made; nagging is
47
+ the thing this replaces.
48
+ - **Re-prompt only** when the operator recorded a *per-stage override map* and the
49
+ next stage's entry differs from the current model — then emit the same block
50
+ scoped to that stage.
51
+ - A skill runs inside the current context and **cannot change the main-loop
52
+ model**; only the operator can, via `/model` (or `/fast`). Preflight is
53
+ interactive anyway, so this costs one exchange.
54
+
55
+ ## Subagents
56
+
57
+ Stage 5 spawns subagents; the orchestrator pins them to the **run's confirmed
58
+ model** via the `Agent` / `Workflow` model override. No operator action needed —
59
+ and no silent downgrade to a cheaper tier.
60
+
61
+ ## Degradation
62
+
63
+ The recommendation is a **reminder, not a block**. If the recommended tier isn't
64
+ available, keep the current model, state plainly which one is in use, and run. The
65
+ pipeline never stalls on a model it can't get.