task-pipeline-skill 0.10.0 → 0.17.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (30) hide show
  1. package/CHANGELOG.md +401 -0
  2. package/LICENSE +85 -0
  3. package/README.md +211 -84
  4. package/cursor/rules/task-pipeline.mdc +135 -20
  5. package/package.json +3 -3
  6. package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
  7. package/plugins/task-pipeline/commands/task-pipeline.md +16 -8
  8. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +139 -57
  9. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +41 -27
  10. package/plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json +1 -1
  11. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
  12. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +32 -6
  13. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
  14. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
  15. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +73 -21
  16. package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
  17. package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
  18. package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +169 -0
  19. package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
  20. package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +55 -22
  21. package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
  22. package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
  23. package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
  24. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +184 -48
  25. package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
  26. package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +11 -3
  27. package/plugins/task-pipeline/skills/task-pipeline/templates/adr.md +64 -0
  28. package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +49 -0
  29. package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
  30. package/plugins/task-pipeline/skills/task-pipeline/templates/context.md +87 -0
@@ -8,11 +8,17 @@ shape.
8
8
  ## In the host project
9
9
 
10
10
  ```
11
+ CONTEXT.md # stage 0 — domain glossary, written inline as terms resolve
11
12
  docs/
13
+ adr/
14
+ NNNN-<slug>.md # stage 0 — ADRs for hard-to-reverse decisions
12
15
  superpowers/
13
16
  specs/
14
17
  YYYY-MM-DD-<topic>-brief.md # stage 0 — locked intake brief (grill output)
15
- YYYY-MM-DD-<topic>-design.md # stage 3 — the spec (locks shared contracts)
18
+ YYYY-MM-DD-<topic>-carryover.md # stage 0 seeds it; EVERY stage appends; stage 10 reads it
19
+ YYYY-MM-DD-<topic>-modules.md # stage 2 — module map + build order (platforms only)
20
+ YYYY-MM-DD-<topic>-design.md # stage 3 — the spec / module dossier (locks shared contracts)
21
+ YYYY-MM-DD-<topic>-acceptance.md # stage 10 — REQ coverage table + evidence
16
22
  plans/
17
23
  YYYY-MM-DD-<topic>.md # stage 4 — the implementation plan
18
24
  ux/ # super-ux, UI tasks only (see companion-skills.md)
@@ -26,15 +32,33 @@ docs/
26
32
  ```
27
33
 
28
34
  Naming: date-prefixed `YYYY-MM-DD-<topic>` slugs, one topic per file, kebab-case.
29
- The three superpowers artifacts share the **same `<topic>` slug** so brief →
30
- design → plan is traceable at a glance.
35
+ Brief, carry-over, design, plan and acceptance share the **same `<topic>` slug**, so the chain is traceable
36
+ at a glance.
37
+
38
+ > The `docs/superpowers/` directory name is this pipeline's historical convention
39
+ > (kept so existing projects don't have to migrate) — **not a dependency on any
40
+ > external skill**. A host project may relocate the root via its `CLAUDE.md`; keep
41
+ > the shape, keep the slugs.
42
+
43
+ Loop-bearing runs also keep a **git-ignored** run ledger at `.task-pipeline/run.md` —
44
+ stage-level and program-level repeat touches, one line each, so the loop guard can
45
+ detect churn after a lost context (see [`loop-guard.md`](loop-guard.md)).
46
+
47
+ Stage 5 also creates a **git-ignored** scratch workspace per plan at
48
+ `.task-pipeline/build/<plan-basename>/` — ledger, task briefs, implementer reports,
49
+ review packages. It is deleted when the final review is clean; git history is the
50
+ record (see `build.md`).
31
51
 
32
52
  ## Stage → artifact map
33
53
 
34
54
  | Stage | Writes | Consumed by |
35
55
  |---|---|---|
36
- | 0 Intake | `specs/<topic>-brief.md` (seed from the skill's `templates/brief.md`) | stages 2–4 |
37
- | 3 Spec | `specs/<topic>-design.md` (+ links `docs/ux/*` for UI) | stage 4 |
56
+ | 0 Intake | `specs/<topic>-brief.md` — incl. the **REQ table** (seed from `templates/brief.md`) | stages 2–5, 7, 10 |
57
+ | 0→10 all | `specs/<topic>-carryover.md` — append-only ledger (seed from `templates/carryover.md`) | stage 10, in full |
58
+ | 10 Acceptance | `specs/<topic>-acceptance.md` — every REQ with a status and evidence | the operator |
59
+ | 0 Grill (domain) | `CONTEXT.md`, `docs/adr/NNNN-<slug>.md` — created **lazily**, only when a term resolves or a decision qualifies | stages 2–4 + the repo |
60
+ | 2 Decompose | `specs/<topic>-modules.md` — module map, build order, contracts, per-module status (platforms only) | stages 3–10, every module's run |
61
+ | 3 Spec | `specs/<topic>-design.md` — module dossier for a decomposed platform (+ links `docs/ux/*` for UI) | stage 4 |
38
62
  | 4 Plan | `plans/<topic>.md` | stage 5 |
39
63
  | 3 UX track | `docs/ux/{foundation,flows,screens,scenarios}.md` | stages 4–9 + `/ux-lint` |
40
64
  | 8 Post-deploy | log/health notes (in the run, not a committed file) | stage 9 |
@@ -51,9 +75,11 @@ plugins/task-pipeline/
51
75
  SKILL.md
52
76
  pipeline.schema.json # generic pipeline contract
53
77
  pipeline.example.json # this plugin's own flow, as config
78
+ references/{grill,brainstorm,decomposition,spec,planning,build,review,tdd,acceptance}.md # built-in stage doctrine
79
+ references/loop-guard.md # cross-cutting: churn detection + break protocol
54
80
  references/{stages,model-tiering,conventions,artifacts,companion-skills}.md
55
81
  cursor/rules/task-pipeline.mdc # Cursor channel (self-contained rule)
56
- plugins/task-pipeline/skills/task-pipeline/templates/brief.md # stage-0 skeleton (ships on every channel)
82
+ plugins/task-pipeline/skills/task-pipeline/templates/{brief,carryover,context,adr}.md # stage-0 skeletons (ship on every channel)
57
83
  bin/task-pipeline.js # npx installer (package task-pipeline-skill)
58
84
  package.json
59
85
  install.sh # POSIX installer
@@ -0,0 +1,106 @@
1
+ # Brainstorm — stage 2, built in
2
+
3
+ The design conversation is **part of this skill**. No companion to install, no
4
+ provider to resolve, nothing to fall back to: this file is the implementation.
5
+
6
+ Stage 0 locked *what* is being built. Stage 2 decides *how*, and stops at an
7
+ approved design — not at code.
8
+
9
+ > Ported, with thanks, from the `brainstorming` skill in
10
+ > [obra/superpowers](https://github.com/obra/superpowers) (MIT — see this repo's
11
+ > `LICENSE` → *Third-party*), rewritten for this pipeline: the brief is the input,
12
+ > the UI verdict is a required output, and the spec write-up moved to stage 3
13
+ > ([`spec.md`](spec.md)).
14
+
15
+ ## The hard gate
16
+
17
+ **No implementation action before the operator approves a design.** No code, no
18
+ scaffolding, no file creation "to see how it'd look", no invoking a
19
+ frontend/backend/build skill. This holds for every task regardless of how simple it
20
+ looks.
21
+
22
+ **"Too simple to need a design" is the trap, not the exception.** A one-function
23
+ utility, a config flip, a copy change — all of them go through this stage. Simple
24
+ tasks are where unexamined assumptions survive longest. The design may be three
25
+ sentences; it still gets presented and approved.
26
+
27
+ ## Input: the brief, not a blank page
28
+
29
+ Read the stage-0 brief first (`…-brief.md`). Everything it locked — scope, users,
30
+ constraints, done-criteria, the autonomy sweep — is **settled**. Re-asking a
31
+ question the grill already answered is the single most common way to waste this
32
+ stage. If the brief and the codebase disagree, that's a contradiction to surface,
33
+ not a question to re-open from scratch.
34
+
35
+ ## The loop
36
+
37
+ 1. **Explore the current state.** Files, module docs, recent commits, the
38
+ conventions the repo already follows. Do this before asking anything.
39
+ 2. **Scope check, early.** If the task actually describes several independent
40
+ subsystems, say so immediately and help decompose it into sub-projects: what the
41
+ independent pieces are, how they relate, what order they get built in. Then
42
+ brainstorm the first one. Each sub-project gets its own spec → plan → build
43
+ cycle. Don't refine details of something that needs splitting first.
44
+ 3. **Questions one at a time.** Never bundle. Multiple choice where it fits, open
45
+ where it doesn't. Purpose, constraints, success criteria — anything the brief
46
+ left at design level.
47
+ 4. **Propose 2–3 approaches with trade-offs**, lead with your recommendation and
48
+ the reason for it. **YAGNI ruthlessly** — strip anything the task doesn't need
49
+ from every option before presenting.
50
+ 5. **Present the design in sections**, each scaled to its complexity (a couple of
51
+ sentences when it's straightforward, up to a few hundred words when it's
52
+ genuinely nuanced). Ask after each section whether it holds. Cover:
53
+ architecture, components, data flow, error handling and degradation, testing.
54
+ 6. **Go back when something doesn't fit.** A revised section beats a design that
55
+ was approved because it was hard to argue with.
56
+
57
+ ## Design for isolation and clarity
58
+
59
+ - Break the system into units with **one clear purpose each**, communicating
60
+ through well-defined interfaces, understandable and testable on their own.
61
+ - For every unit you should be able to answer: what does it do, how is it used,
62
+ what does it depend on?
63
+ - Can a reader understand a unit without reading its internals? Can the internals
64
+ change without breaking consumers? If not, the boundaries need work.
65
+ - Smaller focused files are also what the *implementer* (often a subagent with a
66
+ narrow context) handles reliably. A file growing large is usually a signal it
67
+ does too much.
68
+
69
+ ## Working in an existing codebase
70
+
71
+ - Explore the structure before proposing changes; follow the patterns already
72
+ there.
73
+ - Where existing code genuinely blocks the work — a file that's grown unwieldy,
74
+ tangled responsibilities, an unclear boundary the change has to cross — include
75
+ the targeted improvement in the design, the way a careful developer improves code
76
+ they're working in.
77
+ - Don't propose unrelated refactoring. Anything out of scope goes to the backlog,
78
+ not into this design.
79
+
80
+ ## UI detection — a required output
81
+
82
+ One branch is always: **does this touch a user-facing surface** (web, mobile, CLI,
83
+ TUI — a screen, a command, a visible behavior)? Stage 0 usually answered it; this
84
+ stage confirms it against the design that actually emerged. Record the verdict —
85
+ it arms the stage-3 UX track ([`spec.md`](spec.md) → *UX track*). When it's
86
+ genuinely borderline, record "yes": a false positive costs one extra chain, a false
87
+ negative ships an unspecified interface.
88
+
89
+ ## GATE (manual)
90
+
91
+ The operator approves the design **and** the UI verdict is recorded **and every REQ
92
+ in the brief is answered by the design** — a requirement the design doesn't address
93
+ is either covered now or explicitly dropped by the operator, with the drop written
94
+ into the carry-over ledger. For a platform, the module map
95
+ ([`decomposition.md`](decomposition.md)) is committed and approved as part of this
96
+ same gate. Then, and only then, stage 3 writes it up.
97
+
98
+ ## Rationalizations
99
+
100
+ | Excuse | Reality |
101
+ |---|---|
102
+ | "The brief already says everything" | The brief locks *what*. If it also locked *how*, say so in one line and get the approval anyway — the gate is the point. |
103
+ | "It's a one-line change, design is ceremony" | Then the design is one line. Present it. |
104
+ | "I'll scaffold while they think" | Scaffolding is implementation. The gate is before it, not around it. |
105
+ | "Both approaches are fine, let them pick" | You read the codebase, they didn't. Recommend, then let them override. |
106
+ | "I'll add the extra option now, it's cheap" | YAGNI. Every unused branch is code someone maintains and a test someone writes. |
@@ -0,0 +1,364 @@
1
+ # Build — stage 5, built in
2
+
3
+ Executing the plan: an isolated workspace, one fresh implementer subagent per task,
4
+ a review after each task, a whole-branch review at the end. Built into this skill;
5
+ nothing to install.
6
+
7
+ > Ported from the `using-git-worktrees` and `subagent-driven-development` skills in
8
+ > [obra/superpowers](https://github.com/obra/superpowers) (MIT — see `LICENSE` →
9
+ > *Third-party*), rewritten for this pipeline: no external scripts, the ledger
10
+ > lives under `.task-pipeline/`, and model choice defers to the run's single
11
+ > confirmed model ([`model-tiering.md`](model-tiering.md)).
12
+
13
+ **Why subagents:** each task goes to an agent with a constructed context — its
14
+ brief, its interfaces, the global constraints, nothing else. It never inherits the
15
+ session's history, so it stays focused; your context stays free for coordination.
16
+
17
+ **Continuous execution:** don't check in between tasks. The operator asked for the
18
+ plan to be executed — execute it. Stop only for BLOCKED you can't resolve, a
19
+ genuine ambiguity, or completion. "Should I continue?" between tasks is noise.
20
+
21
+ **Narration:** at most one short line between tool calls. The ledger and the tool
22
+ results are the record.
23
+
24
+ **No subagents available?** (a harness without them, or a plan so small that
25
+ dispatching costs more than it saves) — run the same loop inline: same isolation,
26
+ same ledger, same TDD per task, and after each task review your own diff against
27
+ the rubric in [`review.md`](review.md) before moving on. What changes is who does
28
+ the work; the gates, the artifacts and the review discipline do not. Say plainly
29
+ that the run is inline, since a self-review is weaker evidence than a fresh
30
+ reviewer's.
31
+
32
+ ## 1. Isolation
33
+
34
+ Work never starts on `main`/`master` without the operator's explicit consent
35
+ (the stage-0 brief usually records the branch policy — read it, don't re-ask).
36
+
37
+ **Detect existing isolation first:**
38
+
39
+ ```bash
40
+ GIT_DIR=$(cd "$(git rev-parse --git-dir)" && pwd -P)
41
+ GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" && pwd -P)
42
+ git rev-parse --show-superproject-working-tree # non-empty ⇒ submodule, not a worktree
43
+ ```
44
+
45
+ - `GIT_DIR != GIT_COMMON` and **not** a submodule → you are already in a linked
46
+ worktree. Do not create another. Report the path and branch, go to *Setup*.
47
+ - Otherwise you are in a normal checkout. Honor the brief's worktree preference; if
48
+ none was recorded, ask once before creating one.
49
+
50
+ **Creating one — native tool first.** If the harness offers a worktree tool
51
+ (`EnterWorktree`, a `/worktree` command, a `--worktree` flag), use it: it owns
52
+ placement, branch creation and cleanup. Reaching for raw `git worktree add` when a
53
+ native tool exists creates state the harness can't see or clean up.
54
+
55
+ **Git fallback**, only when there is no native tool:
56
+
57
+ ```bash
58
+ # both scratch roots MUST be ignored before anything is created
59
+ git check-ignore -q .worktrees || printf '.worktrees/\n' >> .gitignore
60
+ git check-ignore -q .task-pipeline || printf '.task-pipeline/\n' >> .gitignore
61
+ git diff --quiet .gitignore || git commit -m "chore: ignore build scratch dirs" .gitignore
62
+ git worktree add ".worktrees/$BRANCH" -b "$BRANCH"
63
+ ```
64
+
65
+ Directory priority: an explicit operator preference → an existing `.worktrees/` →
66
+ an existing `worktrees/` → default `.worktrees/`. An unignored worktree directory
67
+ commits the entire tree into the repo — verify before creating. If creation fails
68
+ on a sandbox permission error, say so plainly and work in place.
69
+
70
+ **Setup + baseline.** Install dependencies the way the project does (`npm install`,
71
+ `cargo build`, `pip install -r requirements.txt`, `poetry install`, `go mod
72
+ download`), then run the test command from the brief's autonomy sweep. A dirty
73
+ baseline makes every later failure ambiguous: report failures and let the operator
74
+ decide whether to proceed.
75
+
76
+ ## 2. Workspace and ledger
77
+
78
+ Conversation memory does not survive compaction. A controller that lost its place
79
+ re-dispatches completed tasks — the most expensive failure this stage has.
80
+ **Track progress in a file, not only in todos.**
81
+
82
+ - Each plan owns a git-ignored workspace: `.task-pipeline/build/<plan-basename>/`
83
+ at the repo root. Everything for THIS plan lives there — ledger, task briefs,
84
+ implementer reports, review packages. Another plan's directory is never yours to
85
+ read or write. `.task-pipeline/` must be git-ignored — the isolation step above
86
+ adds and commits it; if you skipped that step, do it now, in its own commit, so
87
+ scratch files never land in a task's diff.
88
+ - Ledger: `<workspace>/progress.md`, first line = its identity:
89
+ `# build ledger — plan: <plan file path>`.
90
+ - **Resuming:** a task with a `Task <N>: complete` line is DONE — never
91
+ re-dispatch it; resume at the first task without one. A task whose last line is a
92
+ fix round is mid-loop: continue at the next round. A ledger naming a different
93
+ plan belongs to that plan — leave it and start your own.
94
+ - After compaction, trust the ledger and `git log` over your recollection: the
95
+ commits it names exist even when your context no longer remembers them.
96
+
97
+ Read the plan **once**, note its context and Global Constraints, create a todo per
98
+ task.
99
+
100
+ **Pre-flight conflict scan.** Before Task 1, scan the plan for tasks that
101
+ contradict each other or the Global Constraints, and for anything the plan mandates
102
+ that the review rubric ([`review.md`](review.md)) treats as a defect. Present
103
+ everything you find as **one batched question** — each finding beside the plan text
104
+ that mandates it, asking which governs. Clean scan → proceed silently.
105
+
106
+ ## 3. Models
107
+
108
+ **Default: the run's one confirmed model** ([`model-tiering.md`](model-tiering.md))
109
+ for every subagent — implementers, reviewers, fixers. Pin it explicitly on each
110
+ dispatch; an omitted model silently inherits the session's and defeats whatever the
111
+ operator recorded.
112
+
113
+ **Deviate only from the operator's recorded override map.** If the stage-0 brief
114
+ carries per-stage or per-role overrides, apply them: mechanical transcription tasks
115
+ (the plan carries the complete code, 1–2 files) can take a cheaper tier, while
116
+ integration, design and review work stays on the confirmed model. No map recorded →
117
+ no deviation, and never a silent downgrade. Turn count beats token price: the
118
+ cheapest tier routinely takes 2–3× the turns on multi-step work and costs more
119
+ overall.
120
+
121
+ **Two moments deserve more capability than the run's default** — both are
122
+ *recommendations you state out loud*, never silent switches:
123
+
124
+ - **The final whole-branch review.** If the run is on a tier below the most capable
125
+ one available, say so and offer to run this one review there; if the operator
126
+ declines or the tier doesn't exist, run it on the confirmed model and note it.
127
+ - **Fix-loop rounds 4–5.** One tier above the implementer that got stuck, when the
128
+ environment has one and the override map or the operator allows it; otherwise
129
+ say so and rely on fresh eyes alone.
130
+
131
+ ## 4. The task loop
132
+
133
+ Everything you paste into a dispatch prompt — and everything a subagent prints back
134
+ — stays in your context for the rest of the session. **Hand artifacts over as
135
+ files.**
136
+
137
+ ### 4.1 Dispatch the implementer
138
+
139
+ Record `BASE=$(git rev-parse HEAD)` before dispatching; the review package and the
140
+ fix-round diffs need it.
141
+
142
+ **Write the task brief to a file** — extract the task's full text from the plan to
143
+ `<workspace>/task-<N>-brief.md`. The brief is the single source of requirements;
144
+ exact values (numbers, strings, signatures, test cases) live **only** there.
145
+ Include the task's `Implements:` ids **with each REQ's one-line statement quoted
146
+ verbatim** — an implementer who sees only an instruction optimises the
147
+ instruction; one who sees the requirement behind it catches the case the
148
+ instruction didn't cover.
149
+
150
+ The dispatch prompt contains exactly five things:
151
+
152
+ 1. One line on where this task fits in the project.
153
+ 2. The brief path — "read this first; it is your requirements, use its values
154
+ verbatim".
155
+ 3. Interfaces and decisions from earlier tasks the brief can't know.
156
+ 4. Your resolution of any ambiguity you spotted in the brief.
157
+ 5. The report path (`<workspace>/task-<N>-report.md`) and the report contract.
158
+
159
+ Never paste accumulated history ("state after tasks 1–3") into later dispatches.
160
+ Never make a subagent read the whole plan. If an earlier task parked a finding in
161
+ the area this task touches, carry a pointer to that ledger line.
162
+
163
+ Record the implementer's agent identity: fix rounds 1–3 resume it.
164
+
165
+ **Implementer contract** (put this in the prompt):
166
+
167
+ > Read `<brief path>` first — it is your requirements. Work TDD, and no production
168
+ > code exists before a test you **watched fail**: write the failing test → run it
169
+ > and confirm it fails for the right reason → write the minimal code that passes →
170
+ > run it and confirm it passes, with the rest of the suite still green → commit.
171
+ > Assert on real behavior, never on mock behavior. Commit as you go, conventional
172
+ > commits. When done, self-review your diff, then write the full report to
173
+ > `<report path>`:
174
+ > what you built, the files touched, the commits, the test command and its
175
+ > output, decisions you made, anything you're unsure about. Return **only**:
176
+ > status (`DONE` / `DONE_WITH_CONCERNS` / `NEEDS_CONTEXT` / `BLOCKED`), the
177
+ > commit range, a one-line test summary, and your concerns. Ask before starting
178
+ > if anything in the brief is ambiguous — questions are cheaper than rework.
179
+
180
+ ### 4.2 Parallel groups — when fan-out is allowed
181
+
182
+ The plan's parallel groups ([`planning.md`](planning.md)) describe what *may* run
183
+ concurrently. Whether it actually does is this stage's call, and the constraint is
184
+ physical: **two implementers writing one working tree corrupt each other's state.**
185
+
186
+ - **Default: sequential.** One implementer at a time, review after each. Correct for
187
+ every group, and always correct when the tasks are small.
188
+ - **Fan out only when all three hold:** the tasks are in the same group (no
189
+ `depends:` between them), their file ownership is exclusive per the plan, and
190
+ **each implementer gets its own isolated worktree**. Then dispatch them together,
191
+ review each one against its own diff, and integrate the worktrees back to the
192
+ build branch one at a time, running the suite after each merge.
193
+ - **Any conflict on integration** means the plan's file ownership was wrong: stop
194
+ fanning out, finish the group sequentially, and record it in the ledger.
195
+ - Never fan out the fix loop — a task under repair belongs to one implementer.
196
+
197
+ ### 4.3 Handle the report
198
+
199
+ | Status | Action |
200
+ |---|---|
201
+ | `DONE` | Build the review package, dispatch the task review ([`review.md`](review.md)). |
202
+ | `DONE_WITH_CONCERNS` | Read the concerns first. Correctness or scope → resolve before review. Observations ("this file is getting large") → **append to the carry-over ledger**, then proceed. A concern that stays only in the report dies with the workspace. |
203
+ | `NEEDS_CONTEXT` | Supply exactly what's missing, re-dispatch. |
204
+ | `BLOCKED` | Diagnose: missing context → re-dispatch with it; needs more reasoning → a more capable model; too large → split the task; the plan itself is wrong → escalate to the operator. |
205
+
206
+ **Never** ignore an escalation, and never re-dispatch the same model with the same
207
+ prompt after a BLOCKED. If the implementer says it's stuck, something must change.
208
+ If the implementer asks a question — before or mid-task — answer it completely; do
209
+ not rush it into implementation.
210
+
211
+ ### 4.4 Review the task
212
+
213
+ Every task gets a review with **all three** verdicts — spec compliance, **REQ
214
+ satisfied**, and code quality. The implementer's self-review never substitutes for
215
+ it. Rubric, inputs, prompt templates and how to build the diff package:
216
+ [`review.md`](review.md).
217
+
218
+ The REQ verdict is the one the other two can't produce: a task can meet every line
219
+ of its brief and still miss the requirement it was written to deliver. A ❌ there
220
+ enters the fix loop like any Important finding.
221
+
222
+ A review may report **"cannot verify from diff"** items — requirements that live in
223
+ unchanged code or span tasks. They don't block the review, but you resolve each one
224
+ yourself before completing the task; you hold the cross-task context the reviewer
225
+ lacks. A confirmed gap becomes a failed spec review and enters the fix loop.
226
+
227
+ ### 4.5 The fix loop
228
+
229
+ Triggered by: spec ❌, any Critical or Important finding, or a "cannot verify" item
230
+ you confirmed as a real gap.
231
+
232
+ Two routes leave before the loop starts:
233
+
234
+ - **Minor findings** never enter it. Record each in the ledger
235
+ (`Task <N>: minor (deferred): <one-liner>`) and point the final review at that
236
+ list. A roll-up nobody reads is a silent discard.
237
+ - **A finding that conflicts with what the plan mandates** is the operator's
238
+ call: present the finding beside the plan text and ask which governs. Don't
239
+ dismiss the finding because the plan mandated it; don't fix against the plan
240
+ without asking.
241
+
242
+ Everything else loops. One round = one fix dispatch + one scoped re-review.
243
+ **Five rounds maximum per task.**
244
+
245
+ **The loop guard runs alongside the counter** ([`loop-guard.md`](loop-guard.md)):
246
+ log every repeat touch (`touch: <file> — round N — reason: <finding id>`) and trip
247
+ *before* the cap when a fix undoes an earlier fix, when the same file returns for
248
+ the same reason, or when a finding already ADDRESSED reappears. A tripped guard is
249
+ not another round: stop, name the two shapes, escalate to the layer that owns the
250
+ conflict, then re-check in a planned order.
251
+
252
+ - **Rounds 1–3:** resume the original implementer with the open findings verbatim —
253
+ its context is intact. If the harness can't message a live subagent, dispatch a
254
+ fresh one with the brief path, the report path and the findings; the report file
255
+ is the persistent memory either way.
256
+ - **Rounds 4–5:** fresh implementer, one tier up if available, framed as: "a prior
257
+ implementer attempted this task N times; you own it now — read the report file
258
+ for what was tried." A loop that survives three resumes usually means the
259
+ implementer can't see its own problem.
260
+ - **Every round:** the implementer fixes, re-runs the tests covering the amended
261
+ code, appends its fix report to the same report file, returns the short contract.
262
+ Before re-dispatching the reviewer, confirm the fix report names the covering
263
+ tests, the command run and the output.
264
+ - **The re-review is scoped** to the fix diff (`FIX_BASE`..`HEAD`, where `FIX_BASE`
265
+ is the head the previous review saw). It verdicts each finding ADDRESSED / NOT
266
+ ADDRESSED and flags new breakage in the fix diff only. New Critical/Important
267
+ breakage joins the open list; out-of-scope observations go to the ledger as
268
+ deferred minors — they never extend the loop.
269
+ - **Ledger, every round:**
270
+ `Task <N>: fix round <R>/5 (<X> addressed, <Y> open — <one-liners>; commits <a7>..<b7>)`
271
+
272
+ **In a subagent run, never fix findings yourself in the controller session** —
273
+ controller fixes skip review and pollute the context you need for coordination. In a
274
+ declared inline run you do fix them, and you still review the fix diff against the
275
+ rubric before closing the round.
276
+
277
+ **The breaker.** If round 5's re-review still leaves findings open, stop
278
+ dispatching and adjudicate each one yourself:
279
+
280
+ - **Reviewer wrong or the point contestable** → park it:
281
+ `Task <N>: parked — <finding> — ruling: <why the code stands>`.
282
+ - **Real, but nothing downstream builds on it** → park it the same way, with a
283
+ ruling saying it's real and deferred.
284
+ - **Real and load-bearing** (a later task builds on it, or it exposes a plan defect)
285
+ → **STOP**. Append `Task <N>: BLOCKED — <reason>` and report to the operator with
286
+ the finding, the plan text it collides with, and the fix history. Parking a
287
+ structural failure lets every dependent task build on it.
288
+
289
+ Adjudicate **only at the cap**. Adjudicating earlier to end a loop is pre-judging
290
+ with a nicer name. Every adjudication is a ledger line; silent discards are
291
+ forbidden.
292
+
293
+ ### 4.6 Complete the task
294
+
295
+ When the review is clean — or every open finding is parked with a ruling at the cap
296
+ — append:
297
+
298
+ - `Task <N>: complete (commits <base7>..<head7>, review clean)`, or
299
+ - `Task <N>: complete (commits <base7>..<head7>, <K> parked)`
300
+
301
+ Mark the todo complete, move on. Never start the next task while Critical/Important
302
+ findings are neither fixed nor parked-with-ruling at the cap.
303
+
304
+ ## 5. Final whole-branch review
305
+
306
+ After the last task: build a package over `MERGE_BASE`..`HEAD`
307
+ (`git merge-base main HEAD`), dispatch the whole-branch review
308
+ ([`review.md`](review.md) → *Final review*; on the run's model, escalation offered
309
+ out loud per *Models* above), and point it at the
310
+ ledger's deferred-minor and parked lines so it can triage what must be fixed before
311
+ merge.
312
+
313
+ If it returns findings, dispatch **ONE** fix subagent with the complete list — not
314
+ one fixer per finding; per-finding fixers each rebuild context and re-run suites.
315
+ Then exactly **one** scoped re-review of the fix wave. Adjudicate residuals as in
316
+ the breaker: park with rulings, or stop on load-bearing ones. There is no second
317
+ fix wave.
318
+
319
+ ## 6. Integrate, then finish
320
+
321
+ The work is in a worktree on its own branch; stages 7–9 lint, deploy and document
322
+ the **integrated** result. Close that gap here, honoring the branch policy recorded
323
+ in the stage-0 brief:
324
+
325
+ 1. **Sync with the base branch** (rebase or merge, whichever the project uses) and
326
+ re-run the full suite on the result. A branch that was green in isolation and red
327
+ after integration is red — fix it here, not at stage 7.
328
+ 2. **Land it the project's way:** merge into the base branch, or open a PR when the
329
+ project requires review. Opening a PR is outward-facing — do it only with the
330
+ operator's go or the brief's specific standing authorization.
331
+ 3. **Never force-push a shared branch**, and never land on `main` when the brief put
332
+ it off-limits.
333
+ 4. **Remove the worktree** once merged (`git worktree remove <path>`, or the native
334
+ tool that created it), and delete this plan's workspace
335
+ (`rm -rf .task-pipeline/build/<plan-basename>`) — git history is the record now.
336
+ Sibling directories belong to other plans; leave them.
337
+
338
+ If the operator's policy is "leave the branch, I'll merge it myself", stop after
339
+ step 1, say exactly where the branch is and what state it's in, and record that
340
+ stages 7–9 run against an unintegrated branch.
341
+
342
+ ## GATE (auto)
343
+
344
+ All plan tasks DONE with all three review verdicts (spec compliance, REQ satisfied,
345
+ code quality); the full test suite green; every open finding either fixed or parked
346
+ with a ruling; **every parked finding and implementer concern harvested into the
347
+ carry-over ledger** — the workspace is deleted, so nothing may stay only there;
348
+ no task left BLOCKED; the branch integrated per the brief's policy — or the
349
+ operator explicitly told you to leave it, and that is recorded. Verify it yourself;
350
+ a red suite or an unresolved BLOCKED does not advance to stage 6.
351
+
352
+ ## Rationalizations
353
+
354
+ | Excuse | Reality |
355
+ |---|---|
356
+ | "Close enough on spec compliance" | The reviewer found spec gaps ⇒ not done. Fix, or hit the cap and adjudicate. Those are the only exits. |
357
+ | "I'll fix it myself, dispatching is overhead" | In a subagent run, controller fixes skip review and pollute your context — resume the implementer. (Inline runs are the declared exception, and still review the fix diff.) |
358
+ | "One more round will converge" | Past the cap, rounds don't converge — the failure is structural. Adjudicate and route. |
359
+ | "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
360
+ | "This finding is obviously wrong, drop it" | You adjudicate at the cap, in writing. Silent discards are forbidden. |
361
+ | "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Without one, controllers re-run entire completed task sequences. |
362
+ | "Two implementers in parallel will be faster" | One working tree, two writers = corrupted state. Parallel needs one worktree each. |
363
+ | "I'll paste the earlier tasks so it has context" | A fresh subagent needs its task, its interfaces and the constraints. Pasted history is pure cost. |
364
+ | "Stage 7 can merge the branch" | Stage 7 lints and deploys what is integrated. An unmerged branch means lint, deploy and docs all ran against something that is not what ships. |