task-pipeline-skill 0.12.0 → 0.17.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (26) hide show
  1. package/CHANGELOG.md +316 -0
  2. package/LICENSE +47 -0
  3. package/README.md +155 -83
  4. package/cursor/rules/task-pipeline.mdc +88 -15
  5. package/package.json +3 -3
  6. package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
  7. package/plugins/task-pipeline/commands/task-pipeline.md +8 -6
  8. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +80 -33
  9. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +30 -15
  10. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
  11. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +29 -7
  12. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
  13. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
  14. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +61 -29
  15. package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
  16. package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
  17. package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +44 -6
  18. package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
  19. package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
  20. package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
  21. package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
  22. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +151 -26
  23. package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
  24. package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +5 -3
  25. package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +26 -1
  26. package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
@@ -4,7 +4,7 @@ The intake grill is **part of this skill**. No companion skill to install, no
4
4
  provider to resolve, nothing to fall back to: this file *is* the implementation.
5
5
 
6
6
  Its job is not to design. It is to take a one-line request ("make me feature X")
7
- and interview it into a brief complete enough that stages 1→9 finish without
7
+ and interview it into a brief complete enough that stages 1→10 finish without
8
8
  coming back to the operator.
9
9
 
10
10
  > Adapted, with thanks, from Matt Pocock's `grilling` / `grill-with-docs` skills
@@ -95,19 +95,22 @@ numbering (scan for the highest number, increment).
95
95
  ## The autonomy sweep
96
96
 
97
97
  Resolving the *task* is not enough. The grill must also pre-resolve everything that
98
- would otherwise stop stages 1→9 mid-flight. Every row gets an answer **or** an
98
+ would otherwise stop stages 1→10 mid-flight. Every row gets an answer **or** an
99
99
  explicit "stop and ask me here":
100
100
 
101
101
  | Stage | What to settle up front |
102
102
  |---|---|
103
103
  | run-wide | the model decision ([`model-tiering.md`](model-tiering.md)); what to decide autonomously vs escalate |
104
104
  | 1 Docs | external libs/APIs/SDKs in play; any private ones context7 can't resolve → where their docs live |
105
+ | 2 Decompose | is this a platform (several capabilities/surfaces) or one module? if platform: deploy cadence — per module or once at the end |
105
106
  | 2–3 Spec | UI verdict (arms super-ux); any scenario-tracing waiver |
106
107
  | 4–5 Dev | base branch; worktree/branch policy; is `main` off-limits; commit convention; task tracker |
108
+ | 5 Integration | how the branch lands (merge / PR + approver / "leave it unmerged"); parallel fan-out wanted (one worktree per implementer)? |
107
109
  | 6 Tests | the test command; what "green" means here; known-red baseline; coverage expectation |
108
110
  | 7 Lint+deploy | lint command; deploy target and path; release automation on/off; deploy-from-main rule; **deploy authorization** |
109
111
  | 8 Post-deploy | where logs / health live (app name, endpoint, workflow) |
110
112
  | 9 Docs+wiki | which module docs / runbooks this change updates; wiki sync yes/no |
113
+ | 10 Acceptance | who signs off; where deferred REQs are tracked (issue tracker, backlog) |
111
114
 
112
115
  **Deploy authorization has a hard floor.** Deploy and publish are outward and
113
116
  irreversible, so a vague "just do everything" authorizes nothing. A standing
@@ -116,14 +119,49 @@ preconditions ("staging once lint and the full suite are green; production alway
116
119
  asks"). Specific and recorded → it satisfies the stage-7 manual gate. Broader,
117
120
  absent or ambiguous → stage 7 stops and asks.
118
121
 
122
+ ## The REQ spine — the grill's other hard output
123
+
124
+ Prose scope is not checkable. Before the brief is confirmed, the grill must turn
125
+ what was asked into an **addressable list of requirements**, because every later
126
+ stage traces to these IDs and stage 10 accounts for every one of them.
127
+
128
+ | ID | Requirement | How it's verified | Status |
129
+ |---|---|---|---|
130
+ | REQ-001 | … | test name / `file:line` / command + expected output / `SCN-…` | open |
131
+
132
+ Three rules that decide whether the spine is worth anything:
133
+
134
+ 1. **One REQ = one independently verifiable deliverable.** Not one per sentence of
135
+ the request. A small task gets three rows, not thirty — an inflated table is
136
+ ignored, and an ignored table protects nothing.
137
+ 2. **Every row names its check.** *A requirement you can't say how to verify is a
138
+ badly-stated requirement* — split or sharpen it here, during the grill. This is
139
+ the single defence against the failure mode where three vague REQs cover a large
140
+ task and acceptance goes green over half of it.
141
+ 3. **Ask what "finished" means per row, not for the task overall.** "Export works"
142
+ hides five decisions; "exports the currently filtered rows as CSV, verified by
143
+ `test_export_respects_filters`" hides none.
144
+
145
+ **Then freeze it.** Adding a requirement mid-run is fine — append with its source.
146
+ **Removing or narrowing one requires the operator's explicit agreement**, recorded
147
+ in the carry-over ledger. Quietly restating the task in smaller terms is the
148
+ subtlest way to lose it: every gate downstream then passes honestly, on a task
149
+ that shrank without anyone deciding it should.
150
+
119
151
  ## Output
120
152
 
121
153
  Everything resolved goes into the **task brief**, seeded from
122
154
  [`templates/brief.md`](../templates/brief.md) and committed to
123
- `docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` — scope, users, UI verdict,
124
- constraints, locked decisions, the autonomy table, done-criteria, open
125
- assumptions. Seed the template only when the file is absent; never overwrite an
126
- existing brief.
155
+ `docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` — scope, **the REQ table**,
156
+ users, UI verdict, constraints, locked decisions, the autonomy table,
157
+ done-criteria, open assumptions. Seed the template only when the file is absent;
158
+ never overwrite an existing brief.
159
+
160
+ Alongside it, seed the **carry-over ledger** from
161
+ [`templates/carryover.md`](../templates/carryover.md) at
162
+ `…-carryover.md` — append-only, written by every later stage, read in full by
163
+ stage 10. Anything deferred, dropped, or half-done from here on goes there the
164
+ moment it's said: **deferred out loud is forgotten.**
127
165
 
128
166
  Plus, where the session produced them: an updated `CONTEXT.md` and any ADRs, each
129
167
  written as the decision landed.
@@ -0,0 +1,100 @@
1
+ # Loop guard — breaking churn, cross-cutting
2
+
3
+ Any stage that can repeat can also **churn**: a later pass undoing what an earlier
4
+ pass in the same run already did, two shapes alternating, the same file rewritten
5
+ round after round with no new information. Churn looks like progress and consumes
6
+ a run.
7
+
8
+ This file is the detector and the break protocol. It binds every repeating loop in
9
+ the pipeline: the stage-5 fix loop, a stage re-entered after a failed gate, the
10
+ per-module program loop ([`decomposition.md`](decomposition.md)), and any
11
+ audit → fix → audit cycle.
12
+
13
+ ## Bookkeeping — the thing that makes detection mechanical
14
+
15
+ You cannot detect churn from memory, especially after compaction. Every repeating
16
+ pass appends one line to the run's ledger (`.task-pipeline/build/<plan>/progress.md`
17
+ for stage 5; `.task-pipeline/run.md` for stage-level and program-level loops):
18
+
19
+ ```
20
+ touch: <file> — pass <N> (<stage|round|module>) — reason: <finding id / gate item>
21
+ ```
22
+
23
+ One line per file per pass. The reason must name **what forced the edit** — a
24
+ finding id, a failed gate item, an operator instruction. "Cleanup", "polish" and
25
+ "while I was there" are not reasons; they are churn with better manners.
26
+
27
+ ## Detection — any one of these trips the guard
28
+
29
+ 1. **Revert-oscillation.** An edit restores something an earlier pass in this run
30
+ deliberately removed, or re-removes what an earlier pass added. Shape A → B → A.
31
+ 2. **Repeat touch without new information.** The same file is edited in two
32
+ consecutive passes and the second pass's `reason` is the same finding/gate item
33
+ as the first — the fix did not fix it, or the two passes disagree about what
34
+ "fixed" means.
35
+ 3. **Finding resurrection.** A finding whose text (normalized) matches one already
36
+ marked ADDRESSED or parked-with-ruling in this run comes back.
37
+ 4. **Gate ping-pong.** The same stage is re-entered for the third time on the same
38
+ artifact, or two adjacent stages hand work back and forth (spec ⇄ plan,
39
+ plan ⇄ build) more than twice.
40
+ 5. **Cross-loop contradiction.** A pass in one loop edits a file that a *different*
41
+ loop (another task, another module) already closed in this run — two owners for
42
+ one file.
43
+
44
+ Caps that trip the guard by themselves: **5 fix rounds** per task
45
+ ([`build.md`](build.md)), **2 re-entries** per stage per artifact, **3 passes** per
46
+ module in the program loop.
47
+
48
+ ## The break protocol
49
+
50
+ When the guard trips, **stop editing immediately**. Do not dispatch another fix, do
51
+ not "just try one more thing". Then, in this order:
52
+
53
+ 1. **Freeze and name it.** Write the oscillation down in the ledger and to the
54
+ operator: shape **A** vs shape **B**, one line each, plus who is asking for each
55
+ (a finding, the plan's text, the spec, a gate check, an operator instruction) and
56
+ the evidence for each — `file:line`, the failing command, the review verdict.
57
+ 2. **Find the layer that owns the conflict.** Churn almost always means a decision
58
+ is being re-litigated at the wrong altitude:
59
+ - two findings disagree → the **review rubric** decides
60
+ ([`review.md`](review.md)); if it genuinely doesn't, it's a spec question;
61
+ - a finding contradicts the plan → the **operator** decides which governs
62
+ (never dismiss the finding, never fix against the plan silently);
63
+ - the plan contradicts the spec → back to **stage 4** with the evidence;
64
+ - the spec is ambiguous or wrong → back to **stage 3**, and if the ambiguity was
65
+ an unresolved intake question, say so — that is a stage-0 miss worth recording;
66
+ - two modules claim the same file or entity → back to **decomposition**: the cut
67
+ is wrong.
68
+ **Never resolve a higher-layer conflict inside a lower loop.** Patching code to
69
+ satisfy two contradictory requirements is how a run burns its remaining budget.
70
+ 3. **Re-plan the check.** Replace whatever ad-hoc verification was running with an
71
+ explicit ordered checklist: every disputed item, one line each, in dependency
72
+ order, with a single owner and a single verification command per item. Write it
73
+ to the ledger before touching anything.
74
+ 4. **Go in order, one at a time.** Verify item 1 → if it fails, fix only item 1 →
75
+ re-verify only item 1 → commit → item 2. No parallel edits, no bundled fixes, no
76
+ opportunistic cleanup in the same commit. The point is that each change has one
77
+ reason and one proof.
78
+ 5. **Re-check the whole list once** at the end, in the same order. If a later item
79
+ broke an earlier one, that pair is the real conflict — escalate it per step 2
80
+ instead of looping again.
81
+ 6. **Record the ruling.** Ledger line: `loop-guard: <A vs B> — ruling: <what governs
82
+ and why> — items: <N> verified in order`. The final review reads it.
83
+
84
+ ## When to stop and hand back
85
+
86
+ If step 2 lands on "the operator decides", or a cap is hit a second time after a
87
+ re-planned pass, **stop and report BLOCKED** with: the two shapes, the evidence, the
88
+ history of passes, and your recommendation. That is a complete, honest hand-back —
89
+ far cheaper than a third round of the same argument.
90
+
91
+ ## Rationalizations
92
+
93
+ | Excuse | Reality |
94
+ |---|---|
95
+ | "One more pass and it converges" | Two passes with the same reason already proved it doesn't. The disagreement is above the code. |
96
+ | "I'll just revert to what worked" | That is the oscillation, not the exit. Name A and B first. |
97
+ | "The reviewer keeps changing its mind" | Different findings on the same lines mean the requirement is ambiguous. That's a spec question. |
98
+ | "Tidying while I'm in the file" | Untracked edits are what make churn invisible. One reason per change, in the ledger. |
99
+ | "Logging the loop is bureaucracy" | Detection needs a record; after compaction the ledger is the only memory that survives. |
100
+ | "It's faster than escalating" | A run that spends its budget re-deciding a spec question delivers nothing. Escalation costs one message. |
@@ -0,0 +1,193 @@
1
+ # Plan — stage 4, built in
2
+
3
+ Turning the spec into an implementation plan a **zero-context implementer** can
4
+ execute task by task without reading the spec, the chat, or the rest of the plan.
5
+ Built into this skill; nothing to install.
6
+
7
+ > Ported from the `writing-plans` skill in
8
+ > [obra/superpowers](https://github.com/obra/superpowers) (MIT — see `LICENSE` →
9
+ > *Third-party*), extended with the dependency graph, parallel groups and
10
+ > file-ownership rules this pipeline's stage-5 subagent build depends on.
11
+
12
+ ## Audience
13
+
14
+ Assume a skilled developer who knows nothing about this codebase, this domain or
15
+ this toolset, has questionable taste, and will read **only their own task**.
16
+ Everything they need is in that task: exact paths, complete code, exact commands,
17
+ expected output. DRY. YAGNI. TDD. Frequent commits.
18
+
19
+ Path: `docs/superpowers/plans/YYYY-MM-DD-<topic>.md` — same `<topic>` slug as the
20
+ brief and the spec.
21
+
22
+ ## Before writing tasks
23
+
24
+ **Scope check.** If the spec covers several independent subsystems, split it into
25
+ one plan per subsystem; each plan must produce working, testable software on its
26
+ own.
27
+
28
+ **Map the file structure.** List every file that will be created or modified and
29
+ what each one owns. This is where decomposition gets locked in:
30
+
31
+ - One clear responsibility per file; clear boundaries, defined interfaces.
32
+ - Files that change together live together. Split by responsibility, not by
33
+ technical layer.
34
+ - Follow the existing codebase's patterns. Don't unilaterally restructure — but if
35
+ a file you're modifying has grown unwieldy, planning its split is fair.
36
+
37
+ **Draw the dependency graph.** Which task needs what from which. Then group tasks
38
+ into **parallel groups** in topological order, and tag each task
39
+ `depends: [task ids]`.
40
+
41
+ **File ownership is exclusive within a group.** No two tasks in the same parallel
42
+ group write the same file — that is the rule that makes stage-5 fan-out safe.
43
+ Sequential integration/glue tasks sit *between* groups.
44
+
45
+ ## Task right-sizing
46
+
47
+ A task is the smallest unit that carries its own test cycle and is worth a fresh
48
+ reviewer's gate. Fold setup, configuration, scaffolding and docs into the task
49
+ whose deliverable needs them. Split only where a reviewer could meaningfully reject
50
+ one task while approving its neighbor. Every task ends with an independently
51
+ testable deliverable.
52
+
53
+ Each **step** inside a task is one action, 2–5 minutes: write the failing test →
54
+ run it and watch it fail → minimal implementation → run it and watch it pass →
55
+ commit.
56
+
57
+ ## Plan header — required
58
+
59
+ ```markdown
60
+ # <Feature> — implementation plan
61
+
62
+ > **For agentic workers:** execute this plan task-by-task under the task-pipeline
63
+ > stage-5 build doctrine — isolated workspace, one implementer per task, a review
64
+ > with all three verdicts after each (spec compliance, REQ satisfied, code
65
+ > quality). Steps use `- [ ]` checkboxes.
66
+
67
+ **Goal:** <one sentence>
68
+
69
+ **Architecture:** <2–3 sentences>
70
+
71
+ **Tech stack:** <key technologies>
72
+
73
+ **Spec:** docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md
74
+
75
+ ## Global constraints
76
+
77
+ <the spec's project-wide requirements — version floors, dependency limits, naming
78
+ and copy rules, platform requirements — one line each, exact values copied
79
+ verbatim from the spec. Every task's requirements implicitly include this section.>
80
+
81
+ ## Execution order
82
+
83
+ | Group | Tasks | Runs after |
84
+ |---|---|---|
85
+ | A | 1, 2 | — |
86
+ | B | 3 | A |
87
+
88
+ ---
89
+ ```
90
+
91
+ ## Task structure — required
92
+
93
+ ````markdown
94
+ ### Task N: <component>
95
+
96
+ **Depends:** [task ids, or —]
97
+
98
+ **Implements:** REQ-003, REQ-007 — *(the brief's requirement ids this task
99
+ delivers, or `—` for pure glue/infrastructure tasks. Quote each REQ's one-line
100
+ statement under the DoD so the zero-context implementer sees the intent, not just
101
+ the instruction.)*
102
+
103
+ **Files:**
104
+ - Create: `exact/path/to/file.py`
105
+ - Modify: `exact/path/to/existing.py:123-145`
106
+ - Test: `tests/exact/path/to/test_file.py`
107
+
108
+ **Interfaces:**
109
+ - Consumes: <what this task uses from earlier tasks — exact signatures>
110
+ - Produces: <what later tasks rely on — exact names, parameter and return types.
111
+ The implementer sees only this task; this block is how they learn the names
112
+ neighboring tasks use.>
113
+
114
+ **Definition of done:** <observable, verifiable conditions — tests green, behavior
115
+ demonstrated, docs updated in this same change>
116
+
117
+ - [ ] **Step 1: write the failing test**
118
+
119
+ ```python
120
+ def test_specific_behavior():
121
+ assert function(input) == expected
122
+ ```
123
+
124
+ - [ ] **Step 2: run it and confirm it fails**
125
+
126
+ Run: `pytest tests/path/test_file.py::test_specific_behavior -v`
127
+ Expected: FAIL — `NameError: name 'function' is not defined`
128
+
129
+ - [ ] **Step 3: minimal implementation**
130
+
131
+ ```python
132
+ def function(value):
133
+ return expected
134
+ ```
135
+
136
+ - [ ] **Step 4: run it and confirm it passes**
137
+
138
+ Run: `pytest tests/path/test_file.py::test_specific_behavior -v`
139
+ Expected: PASS
140
+
141
+ - [ ] **Step 5: commit**
142
+
143
+ ```bash
144
+ git add tests/path/test_file.py src/path/file.py
145
+ git commit -m "feat: <what changed>"
146
+ ```
147
+ ````
148
+
149
+ For UI tasks, every task that builds user-facing behavior names the **scenario
150
+ ID(s)** and `SCR-` screen(s) it implements, and its DoD includes satisfying them
151
+ **and** updating the affected super-ux layers in the same change.
152
+
153
+ ## No placeholders
154
+
155
+ These are plan failures. Never write them:
156
+
157
+ - "TBD", "TODO", "implement later", "fill in details"
158
+ - "Add appropriate error handling" / "add validation" / "handle edge cases"
159
+ - "Write tests for the above" without the actual test code
160
+ - "Similar to Task N" — repeat the code; tasks get read out of order
161
+ - A step that says what to do without showing how (code steps need code blocks)
162
+ - References to types, functions or methods no task defines
163
+
164
+ ## Self-review — before handing off
165
+
166
+ A checklist you run yourself, inline. No subagent:
167
+
168
+ 1. **REQ coverage — set equality, not a feeling.** Collect every `Implements:` id
169
+ across all tasks and compare it to the brief's REQ table. The two sets must be
170
+ **equal**: a REQ with no task is scope silently lost; an `Implements:` id that
171
+ isn't in the brief is either a typo or work nobody asked for. Print the
172
+ difference and fix it before anything else — this seam is where scope leaks.
173
+ 2. **Spec coverage:** walk each spec requirement. Point at the task that implements
174
+ it. A requirement with no task → add the task.
175
+ 3. **Placeholder scan:** search the plan for every pattern above. Fix.
176
+ 4. **Name and type consistency:** signatures, property names and types used in
177
+ later tasks match what earlier tasks defined. `clearLayers()` in Task 3 and
178
+ `clearFullLayers()` in Task 7 is a bug, not a style difference.
179
+ 5. **Parallel safety:** no two tasks in the same group write the same file; every
180
+ `depends:` points at a task that really produces what's consumed.
181
+ 6. **DoD present and verifiable** on every task.
182
+
183
+ ## GATE (auto)
184
+
185
+ **Set equality first:** the REQ ids in the brief equal the union of `Implements:`
186
+ across the plan's tasks. A non-empty difference fails the gate and is reported as
187
+ the explicit list of dropped (or invented) requirements — this seam is where scope
188
+ leaks, so the check is mechanical, never a judgement call.
189
+
190
+ Then: every spec requirement maps to a task; no placeholders; names and types
191
+ consistent across tasks; parallel-group tasks share no files; each task has a
192
+ verifiable DoD. UI tasks carry their scenario IDs and `SCR-` screens. Verify all of
193
+ it yourself and stop on failure — this gate has no operator in it.
@@ -0,0 +1,173 @@
1
+ # Review — the rubric and the prompts, built in
2
+
3
+ Used by stage 5's per-task reviews, its scoped re-reviews and its final
4
+ whole-branch review ([`build.md`](build.md)). Built into this skill; nothing to
5
+ install.
6
+
7
+ > Ported from the `requesting-code-review` skill and the reviewer prompts in
8
+ > [obra/superpowers](https://github.com/obra/superpowers) (MIT — see `LICENSE` →
9
+ > *Third-party*), condensed into one rubric plus three copy-paste prompts, with the
10
+ > external helper scripts replaced by plain git commands so the doctrine works on
11
+ > any agent.
12
+
13
+ ## The diff package
14
+
15
+ A reviewer never re-derives the diff with a dozen git calls, and the diff never
16
+ enters **your** context. Write it to one file and pass the path (`$WORKSPACE` is
17
+ this plan's git-ignored directory, `.task-pipeline/build/<plan-basename>/` — see
18
+ [`build.md`](build.md)):
19
+
20
+ ```bash
21
+ { git log --oneline "$BASE..$HEAD"
22
+ echo
23
+ git diff --stat "$BASE..$HEAD"
24
+ echo
25
+ git diff -U10 "$BASE..$HEAD"
26
+ } > "$WORKSPACE/review-task-$N-$(git rev-parse --short "$HEAD").md"
27
+ ```
28
+
29
+ `BASE` is the commit you recorded **before** dispatching the implementer — never
30
+ `HEAD~1`, which silently drops every commit but the last of a multi-commit task.
31
+ For a scoped re-review, `BASE` is the head the previous review saw. For the final
32
+ review, `BASE` is `git merge-base main HEAD`.
33
+
34
+ Never dispatch a reviewer without a diff file.
35
+
36
+ ## Reviewer inputs
37
+
38
+ Four things, three of them paths:
39
+
40
+ 1. The **task brief** file — what was required.
41
+ 2. The **implementer report** file — what was done, with the test evidence.
42
+ 3. The **diff package** path.
43
+ 4. The **Global Constraints** that bind this task — copied **verbatim** from the
44
+ plan: exact values, exact formats, stated relationships ("same layout as X",
45
+ "matches Y"). This block is the reviewer's attention lens.
46
+
47
+ ## Controller rules
48
+
49
+ - **Never pre-judge.** Don't tell a reviewer to ignore something, don't cap a
50
+ severity, don't explain why a finding would be a false positive. If your prompt
51
+ contains "don't flag", "at most Minor", or "the plan chose this" — stop: you're
52
+ buying yourself a shorter loop with an unreviewed defect. Let the finding come and
53
+ adjudicate it in the loop.
54
+ - **Don't ask a reviewer to re-run tests** the implementer already ran on the same
55
+ code; the report carries the evidence.
56
+ - **No open-ended directives** ("check all uses", "run race tests if useful")
57
+ without a concrete, task-specific reason.
58
+
59
+ ## The rubric
60
+
61
+ Review in this order; stop reading the diff only when you've covered all of it.
62
+
63
+ 1. **Spec compliance.** Every requirement in the brief: met, partially met, or
64
+ missing. Anything built that the brief did *not* ask for is scope creep — flag
65
+ it, even when it's nice.
66
+ 2. **REQ satisfaction.** Read the task's `Implements:` requirement statements, not
67
+ just its instructions, and judge the diff against **those**. A task can satisfy
68
+ every line of its brief and still miss the requirement it exists to deliver —
69
+ that gap is invisible one level down, which is why it is asked for here.
70
+ 3. **Correctness.** Logic errors, off-by-one, wrong operator, unhandled `None`/nil,
71
+ race conditions, resource leaks, wrong error propagation. State the concrete
72
+ input or state that produces the wrong output — a finding without a failure
73
+ scenario is an opinion.
74
+ 4. **Global constraints.** Exact values, formats and relationships from the
75
+ constraints block. Approximations are failures.
76
+ 5. **Test honesty.** Tests assert on real behavior, not on mocks. No test that
77
+ passes regardless of the production code. No `skip`/`xfail`/commented assertion
78
+ smuggling a red suite past a gate. New behavior has a covering test; the failure
79
+ path has one too.
80
+ 6. **Error handling and degradation.** Every external call (network, DB, file, MCP,
81
+ API) handles failure, and the failure is reported honestly rather than swallowed.
82
+ 7. **Boundaries and clarity.** One responsibility per unit; names that say what the
83
+ thing is; no duplication of a logic block that should be shared; nothing left
84
+ dead.
85
+ 8. **Security.** No secrets in code, logs or fixtures; input validated at the
86
+ boundary; no new injection or path-traversal surface.
87
+ 9. **Docs in the same change.** Module docs, runbooks and (for UI work) the
88
+ super-ux layers updated alongside the code, not deferred.
89
+
90
+ **Severities:**
91
+
92
+ - **Critical** — wrong behavior, data loss, security hole, a red or dishonest test
93
+ suite. Blocks.
94
+ - **Important** — a real defect or spec gap that will bite: missing requirement,
95
+ unhandled failure path, a magic value the constraints named. Blocks.
96
+ - **Minor** — style, naming, a nit with no behavioral consequence. Never blocks;
97
+ goes to the ledger.
98
+ - **⚠️ Cannot verify from diff** — the requirement lives in unchanged code or spans
99
+ tasks. Not a blocker for the reviewer; the controller resolves it.
100
+
101
+ Formatting nits that don't change meaning are not findings. Praise is not a
102
+ finding either.
103
+
104
+ ## Prompt — task review
105
+
106
+ > You are reviewing one task of an implementation plan. Read, in order:
107
+ > `<brief path>` (the requirements), `<report path>` (what the implementer did and
108
+ > the test evidence), `<diff package path>` (commits, stat, full diff).
109
+ >
110
+ > Global constraints binding this task:
111
+ > ```
112
+ > <verbatim block>
113
+ > ```
114
+ >
115
+ > Produce three verdicts, all required:
116
+ > 1. **Spec compliance:** ✅ or ❌. List every requirement as met / partial /
117
+ > missing, and list anything implemented that was not asked for.
118
+ > 2. **REQ satisfied:** ✅ or ❌ per `Implements:` id. The brief quotes each
119
+ > requirement's statement — judge the diff against **that statement**, not
120
+ > against the task's instructions. A task can follow every instruction and still
121
+ > miss the requirement it exists to deliver; say so when it does.
122
+ > 3. **Code quality:** approved or not. Findings only, each as
123
+ > `severity — file:line — the defect — the failure scenario (concrete input or
124
+ > state → wrong result)`. Severities: Critical, Important, Minor. Use
125
+ > `⚠️ cannot verify from diff` for anything you can't judge from the diff alone.
126
+ >
127
+ > Review against the rubric: correctness, global constraints, test honesty (tests
128
+ > assert real behavior, not mocks; no skipped/empty assertions), error handling and
129
+ > honest degradation, boundaries and naming, security and secrets, docs updated in
130
+ > the same change. No praise, no formatting nits that don't change meaning. Do not
131
+ > re-run the tests the report already covers. Return the verdicts and findings as
132
+ > your final message — nothing else.
133
+
134
+ ## Prompt — scoped re-review
135
+
136
+ > A previous review of this task raised the findings below. The implementer has
137
+ > since pushed fixes. Read `<brief path>`, `<report path>` (its fix report is
138
+ > appended at the end) and `<fix diff package path>` — the fix diff **only**.
139
+ >
140
+ > Open findings:
141
+ > ```
142
+ > 1. <finding>
143
+ > 2. <finding>
144
+ > ```
145
+ >
146
+ > For each finding return `ADDRESSED` (with the `file:line` that resolves it) or
147
+ > `NOT ADDRESSED` (with what's still wrong). Then flag any **new** Critical or
148
+ > Important breakage introduced by this fix diff. Out-of-scope observations about
149
+ > code this diff didn't touch: list them separately as deferred minors — they are
150
+ > not part of this verdict. End with: all findings addressed / N still open.
151
+
152
+ ## Prompt — final whole-branch review
153
+
154
+ > Review this entire branch before merge. Read `<diff package path>` (merge-base to
155
+ > HEAD) and `<spec path>`.
156
+ >
157
+ > Findings deferred or parked during implementation:
158
+ > ```
159
+ > <the ledger's minor + parked lines>
160
+ > ```
161
+ >
162
+ > Judge the branch as a whole: does it deliver the spec; do the pieces fit; is
163
+ > anything half-migrated, duplicated across tasks, or left dead; are the tests
164
+ > honest and the suite genuinely green; is error handling consistent; are docs in
165
+ > sync. Triage the deferred/parked list: which of those must be fixed before merge,
166
+ > which can stand and why. Findings only, with severity and a concrete failure
167
+ > scenario each. Return the findings as your final message.
168
+
169
+ Run the final review on the **run's confirmed model** like everything else
170
+ ([`model-tiering.md`](model-tiering.md)). It is the one review that sees the whole
171
+ change, so if the run is on a tier below the most capable one available, say so and
172
+ offer to escalate just this dispatch — a recommendation stated out loud, never a
173
+ silent switch (`build.md` → *Models*).