task-pipeline-skill 0.12.0 → 0.17.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +316 -0
- package/LICENSE +47 -0
- package/README.md +155 -83
- package/cursor/rules/task-pipeline.mdc +88 -15
- package/package.json +3 -3
- package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
- package/plugins/task-pipeline/commands/task-pipeline.md +8 -6
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +80 -33
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +30 -15
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +29 -7
- package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +61 -29
- package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +44 -6
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +151 -26
- package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +5 -3
- package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +26 -1
- package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
|
@@ -4,7 +4,7 @@ The intake grill is **part of this skill**. No companion skill to install, no
|
|
|
4
4
|
provider to resolve, nothing to fall back to: this file *is* the implementation.
|
|
5
5
|
|
|
6
6
|
Its job is not to design. It is to take a one-line request ("make me feature X")
|
|
7
|
-
and interview it into a brief complete enough that stages 1→
|
|
7
|
+
and interview it into a brief complete enough that stages 1→10 finish without
|
|
8
8
|
coming back to the operator.
|
|
9
9
|
|
|
10
10
|
> Adapted, with thanks, from Matt Pocock's `grilling` / `grill-with-docs` skills
|
|
@@ -95,19 +95,22 @@ numbering (scan for the highest number, increment).
|
|
|
95
95
|
## The autonomy sweep
|
|
96
96
|
|
|
97
97
|
Resolving the *task* is not enough. The grill must also pre-resolve everything that
|
|
98
|
-
would otherwise stop stages 1→
|
|
98
|
+
would otherwise stop stages 1→10 mid-flight. Every row gets an answer **or** an
|
|
99
99
|
explicit "stop and ask me here":
|
|
100
100
|
|
|
101
101
|
| Stage | What to settle up front |
|
|
102
102
|
|---|---|
|
|
103
103
|
| run-wide | the model decision ([`model-tiering.md`](model-tiering.md)); what to decide autonomously vs escalate |
|
|
104
104
|
| 1 Docs | external libs/APIs/SDKs in play; any private ones context7 can't resolve → where their docs live |
|
|
105
|
+
| 2 Decompose | is this a platform (several capabilities/surfaces) or one module? if platform: deploy cadence — per module or once at the end |
|
|
105
106
|
| 2–3 Spec | UI verdict (arms super-ux); any scenario-tracing waiver |
|
|
106
107
|
| 4–5 Dev | base branch; worktree/branch policy; is `main` off-limits; commit convention; task tracker |
|
|
108
|
+
| 5 Integration | how the branch lands (merge / PR + approver / "leave it unmerged"); parallel fan-out wanted (one worktree per implementer)? |
|
|
107
109
|
| 6 Tests | the test command; what "green" means here; known-red baseline; coverage expectation |
|
|
108
110
|
| 7 Lint+deploy | lint command; deploy target and path; release automation on/off; deploy-from-main rule; **deploy authorization** |
|
|
109
111
|
| 8 Post-deploy | where logs / health live (app name, endpoint, workflow) |
|
|
110
112
|
| 9 Docs+wiki | which module docs / runbooks this change updates; wiki sync yes/no |
|
|
113
|
+
| 10 Acceptance | who signs off; where deferred REQs are tracked (issue tracker, backlog) |
|
|
111
114
|
|
|
112
115
|
**Deploy authorization has a hard floor.** Deploy and publish are outward and
|
|
113
116
|
irreversible, so a vague "just do everything" authorizes nothing. A standing
|
|
@@ -116,14 +119,49 @@ preconditions ("staging once lint and the full suite are green; production alway
|
|
|
116
119
|
asks"). Specific and recorded → it satisfies the stage-7 manual gate. Broader,
|
|
117
120
|
absent or ambiguous → stage 7 stops and asks.
|
|
118
121
|
|
|
122
|
+
## The REQ spine — the grill's other hard output
|
|
123
|
+
|
|
124
|
+
Prose scope is not checkable. Before the brief is confirmed, the grill must turn
|
|
125
|
+
what was asked into an **addressable list of requirements**, because every later
|
|
126
|
+
stage traces to these IDs and stage 10 accounts for every one of them.
|
|
127
|
+
|
|
128
|
+
| ID | Requirement | How it's verified | Status |
|
|
129
|
+
|---|---|---|---|
|
|
130
|
+
| REQ-001 | … | test name / `file:line` / command + expected output / `SCN-…` | open |
|
|
131
|
+
|
|
132
|
+
Three rules that decide whether the spine is worth anything:
|
|
133
|
+
|
|
134
|
+
1. **One REQ = one independently verifiable deliverable.** Not one per sentence of
|
|
135
|
+
the request. A small task gets three rows, not thirty — an inflated table is
|
|
136
|
+
ignored, and an ignored table protects nothing.
|
|
137
|
+
2. **Every row names its check.** *A requirement you can't say how to verify is a
|
|
138
|
+
badly-stated requirement* — split or sharpen it here, during the grill. This is
|
|
139
|
+
the single defence against the failure mode where three vague REQs cover a large
|
|
140
|
+
task and acceptance goes green over half of it.
|
|
141
|
+
3. **Ask what "finished" means per row, not for the task overall.** "Export works"
|
|
142
|
+
hides five decisions; "exports the currently filtered rows as CSV, verified by
|
|
143
|
+
`test_export_respects_filters`" hides none.
|
|
144
|
+
|
|
145
|
+
**Then freeze it.** Adding a requirement mid-run is fine — append with its source.
|
|
146
|
+
**Removing or narrowing one requires the operator's explicit agreement**, recorded
|
|
147
|
+
in the carry-over ledger. Quietly restating the task in smaller terms is the
|
|
148
|
+
subtlest way to lose it: every gate downstream then passes honestly, on a task
|
|
149
|
+
that shrank without anyone deciding it should.
|
|
150
|
+
|
|
119
151
|
## Output
|
|
120
152
|
|
|
121
153
|
Everything resolved goes into the **task brief**, seeded from
|
|
122
154
|
[`templates/brief.md`](../templates/brief.md) and committed to
|
|
123
|
-
`docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` — scope,
|
|
124
|
-
constraints, locked decisions, the autonomy table,
|
|
125
|
-
assumptions. Seed the template only when the file is absent;
|
|
126
|
-
existing brief.
|
|
155
|
+
`docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` — scope, **the REQ table**,
|
|
156
|
+
users, UI verdict, constraints, locked decisions, the autonomy table,
|
|
157
|
+
done-criteria, open assumptions. Seed the template only when the file is absent;
|
|
158
|
+
never overwrite an existing brief.
|
|
159
|
+
|
|
160
|
+
Alongside it, seed the **carry-over ledger** from
|
|
161
|
+
[`templates/carryover.md`](../templates/carryover.md) at
|
|
162
|
+
`…-carryover.md` — append-only, written by every later stage, read in full by
|
|
163
|
+
stage 10. Anything deferred, dropped, or half-done from here on goes there the
|
|
164
|
+
moment it's said: **deferred out loud is forgotten.**
|
|
127
165
|
|
|
128
166
|
Plus, where the session produced them: an updated `CONTEXT.md` and any ADRs, each
|
|
129
167
|
written as the decision landed.
|
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# Loop guard — breaking churn, cross-cutting
|
|
2
|
+
|
|
3
|
+
Any stage that can repeat can also **churn**: a later pass undoing what an earlier
|
|
4
|
+
pass in the same run already did, two shapes alternating, the same file rewritten
|
|
5
|
+
round after round with no new information. Churn looks like progress and consumes
|
|
6
|
+
a run.
|
|
7
|
+
|
|
8
|
+
This file is the detector and the break protocol. It binds every repeating loop in
|
|
9
|
+
the pipeline: the stage-5 fix loop, a stage re-entered after a failed gate, the
|
|
10
|
+
per-module program loop ([`decomposition.md`](decomposition.md)), and any
|
|
11
|
+
audit → fix → audit cycle.
|
|
12
|
+
|
|
13
|
+
## Bookkeeping — the thing that makes detection mechanical
|
|
14
|
+
|
|
15
|
+
You cannot detect churn from memory, especially after compaction. Every repeating
|
|
16
|
+
pass appends one line to the run's ledger (`.task-pipeline/build/<plan>/progress.md`
|
|
17
|
+
for stage 5; `.task-pipeline/run.md` for stage-level and program-level loops):
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
touch: <file> — pass <N> (<stage|round|module>) — reason: <finding id / gate item>
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
One line per file per pass. The reason must name **what forced the edit** — a
|
|
24
|
+
finding id, a failed gate item, an operator instruction. "Cleanup", "polish" and
|
|
25
|
+
"while I was there" are not reasons; they are churn with better manners.
|
|
26
|
+
|
|
27
|
+
## Detection — any one of these trips the guard
|
|
28
|
+
|
|
29
|
+
1. **Revert-oscillation.** An edit restores something an earlier pass in this run
|
|
30
|
+
deliberately removed, or re-removes what an earlier pass added. Shape A → B → A.
|
|
31
|
+
2. **Repeat touch without new information.** The same file is edited in two
|
|
32
|
+
consecutive passes and the second pass's `reason` is the same finding/gate item
|
|
33
|
+
as the first — the fix did not fix it, or the two passes disagree about what
|
|
34
|
+
"fixed" means.
|
|
35
|
+
3. **Finding resurrection.** A finding whose text (normalized) matches one already
|
|
36
|
+
marked ADDRESSED or parked-with-ruling in this run comes back.
|
|
37
|
+
4. **Gate ping-pong.** The same stage is re-entered for the third time on the same
|
|
38
|
+
artifact, or two adjacent stages hand work back and forth (spec ⇄ plan,
|
|
39
|
+
plan ⇄ build) more than twice.
|
|
40
|
+
5. **Cross-loop contradiction.** A pass in one loop edits a file that a *different*
|
|
41
|
+
loop (another task, another module) already closed in this run — two owners for
|
|
42
|
+
one file.
|
|
43
|
+
|
|
44
|
+
Caps that trip the guard by themselves: **5 fix rounds** per task
|
|
45
|
+
([`build.md`](build.md)), **2 re-entries** per stage per artifact, **3 passes** per
|
|
46
|
+
module in the program loop.
|
|
47
|
+
|
|
48
|
+
## The break protocol
|
|
49
|
+
|
|
50
|
+
When the guard trips, **stop editing immediately**. Do not dispatch another fix, do
|
|
51
|
+
not "just try one more thing". Then, in this order:
|
|
52
|
+
|
|
53
|
+
1. **Freeze and name it.** Write the oscillation down in the ledger and to the
|
|
54
|
+
operator: shape **A** vs shape **B**, one line each, plus who is asking for each
|
|
55
|
+
(a finding, the plan's text, the spec, a gate check, an operator instruction) and
|
|
56
|
+
the evidence for each — `file:line`, the failing command, the review verdict.
|
|
57
|
+
2. **Find the layer that owns the conflict.** Churn almost always means a decision
|
|
58
|
+
is being re-litigated at the wrong altitude:
|
|
59
|
+
- two findings disagree → the **review rubric** decides
|
|
60
|
+
([`review.md`](review.md)); if it genuinely doesn't, it's a spec question;
|
|
61
|
+
- a finding contradicts the plan → the **operator** decides which governs
|
|
62
|
+
(never dismiss the finding, never fix against the plan silently);
|
|
63
|
+
- the plan contradicts the spec → back to **stage 4** with the evidence;
|
|
64
|
+
- the spec is ambiguous or wrong → back to **stage 3**, and if the ambiguity was
|
|
65
|
+
an unresolved intake question, say so — that is a stage-0 miss worth recording;
|
|
66
|
+
- two modules claim the same file or entity → back to **decomposition**: the cut
|
|
67
|
+
is wrong.
|
|
68
|
+
**Never resolve a higher-layer conflict inside a lower loop.** Patching code to
|
|
69
|
+
satisfy two contradictory requirements is how a run burns its remaining budget.
|
|
70
|
+
3. **Re-plan the check.** Replace whatever ad-hoc verification was running with an
|
|
71
|
+
explicit ordered checklist: every disputed item, one line each, in dependency
|
|
72
|
+
order, with a single owner and a single verification command per item. Write it
|
|
73
|
+
to the ledger before touching anything.
|
|
74
|
+
4. **Go in order, one at a time.** Verify item 1 → if it fails, fix only item 1 →
|
|
75
|
+
re-verify only item 1 → commit → item 2. No parallel edits, no bundled fixes, no
|
|
76
|
+
opportunistic cleanup in the same commit. The point is that each change has one
|
|
77
|
+
reason and one proof.
|
|
78
|
+
5. **Re-check the whole list once** at the end, in the same order. If a later item
|
|
79
|
+
broke an earlier one, that pair is the real conflict — escalate it per step 2
|
|
80
|
+
instead of looping again.
|
|
81
|
+
6. **Record the ruling.** Ledger line: `loop-guard: <A vs B> — ruling: <what governs
|
|
82
|
+
and why> — items: <N> verified in order`. The final review reads it.
|
|
83
|
+
|
|
84
|
+
## When to stop and hand back
|
|
85
|
+
|
|
86
|
+
If step 2 lands on "the operator decides", or a cap is hit a second time after a
|
|
87
|
+
re-planned pass, **stop and report BLOCKED** with: the two shapes, the evidence, the
|
|
88
|
+
history of passes, and your recommendation. That is a complete, honest hand-back —
|
|
89
|
+
far cheaper than a third round of the same argument.
|
|
90
|
+
|
|
91
|
+
## Rationalizations
|
|
92
|
+
|
|
93
|
+
| Excuse | Reality |
|
|
94
|
+
|---|---|
|
|
95
|
+
| "One more pass and it converges" | Two passes with the same reason already proved it doesn't. The disagreement is above the code. |
|
|
96
|
+
| "I'll just revert to what worked" | That is the oscillation, not the exit. Name A and B first. |
|
|
97
|
+
| "The reviewer keeps changing its mind" | Different findings on the same lines mean the requirement is ambiguous. That's a spec question. |
|
|
98
|
+
| "Tidying while I'm in the file" | Untracked edits are what make churn invisible. One reason per change, in the ledger. |
|
|
99
|
+
| "Logging the loop is bureaucracy" | Detection needs a record; after compaction the ledger is the only memory that survives. |
|
|
100
|
+
| "It's faster than escalating" | A run that spends its budget re-deciding a spec question delivers nothing. Escalation costs one message. |
|
|
@@ -0,0 +1,193 @@
|
|
|
1
|
+
# Plan — stage 4, built in
|
|
2
|
+
|
|
3
|
+
Turning the spec into an implementation plan a **zero-context implementer** can
|
|
4
|
+
execute task by task without reading the spec, the chat, or the rest of the plan.
|
|
5
|
+
Built into this skill; nothing to install.
|
|
6
|
+
|
|
7
|
+
> Ported from the `writing-plans` skill in
|
|
8
|
+
> [obra/superpowers](https://github.com/obra/superpowers) (MIT — see `LICENSE` →
|
|
9
|
+
> *Third-party*), extended with the dependency graph, parallel groups and
|
|
10
|
+
> file-ownership rules this pipeline's stage-5 subagent build depends on.
|
|
11
|
+
|
|
12
|
+
## Audience
|
|
13
|
+
|
|
14
|
+
Assume a skilled developer who knows nothing about this codebase, this domain or
|
|
15
|
+
this toolset, has questionable taste, and will read **only their own task**.
|
|
16
|
+
Everything they need is in that task: exact paths, complete code, exact commands,
|
|
17
|
+
expected output. DRY. YAGNI. TDD. Frequent commits.
|
|
18
|
+
|
|
19
|
+
Path: `docs/superpowers/plans/YYYY-MM-DD-<topic>.md` — same `<topic>` slug as the
|
|
20
|
+
brief and the spec.
|
|
21
|
+
|
|
22
|
+
## Before writing tasks
|
|
23
|
+
|
|
24
|
+
**Scope check.** If the spec covers several independent subsystems, split it into
|
|
25
|
+
one plan per subsystem; each plan must produce working, testable software on its
|
|
26
|
+
own.
|
|
27
|
+
|
|
28
|
+
**Map the file structure.** List every file that will be created or modified and
|
|
29
|
+
what each one owns. This is where decomposition gets locked in:
|
|
30
|
+
|
|
31
|
+
- One clear responsibility per file; clear boundaries, defined interfaces.
|
|
32
|
+
- Files that change together live together. Split by responsibility, not by
|
|
33
|
+
technical layer.
|
|
34
|
+
- Follow the existing codebase's patterns. Don't unilaterally restructure — but if
|
|
35
|
+
a file you're modifying has grown unwieldy, planning its split is fair.
|
|
36
|
+
|
|
37
|
+
**Draw the dependency graph.** Which task needs what from which. Then group tasks
|
|
38
|
+
into **parallel groups** in topological order, and tag each task
|
|
39
|
+
`depends: [task ids]`.
|
|
40
|
+
|
|
41
|
+
**File ownership is exclusive within a group.** No two tasks in the same parallel
|
|
42
|
+
group write the same file — that is the rule that makes stage-5 fan-out safe.
|
|
43
|
+
Sequential integration/glue tasks sit *between* groups.
|
|
44
|
+
|
|
45
|
+
## Task right-sizing
|
|
46
|
+
|
|
47
|
+
A task is the smallest unit that carries its own test cycle and is worth a fresh
|
|
48
|
+
reviewer's gate. Fold setup, configuration, scaffolding and docs into the task
|
|
49
|
+
whose deliverable needs them. Split only where a reviewer could meaningfully reject
|
|
50
|
+
one task while approving its neighbor. Every task ends with an independently
|
|
51
|
+
testable deliverable.
|
|
52
|
+
|
|
53
|
+
Each **step** inside a task is one action, 2–5 minutes: write the failing test →
|
|
54
|
+
run it and watch it fail → minimal implementation → run it and watch it pass →
|
|
55
|
+
commit.
|
|
56
|
+
|
|
57
|
+
## Plan header — required
|
|
58
|
+
|
|
59
|
+
```markdown
|
|
60
|
+
# <Feature> — implementation plan
|
|
61
|
+
|
|
62
|
+
> **For agentic workers:** execute this plan task-by-task under the task-pipeline
|
|
63
|
+
> stage-5 build doctrine — isolated workspace, one implementer per task, a review
|
|
64
|
+
> with all three verdicts after each (spec compliance, REQ satisfied, code
|
|
65
|
+
> quality). Steps use `- [ ]` checkboxes.
|
|
66
|
+
|
|
67
|
+
**Goal:** <one sentence>
|
|
68
|
+
|
|
69
|
+
**Architecture:** <2–3 sentences>
|
|
70
|
+
|
|
71
|
+
**Tech stack:** <key technologies>
|
|
72
|
+
|
|
73
|
+
**Spec:** docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md
|
|
74
|
+
|
|
75
|
+
## Global constraints
|
|
76
|
+
|
|
77
|
+
<the spec's project-wide requirements — version floors, dependency limits, naming
|
|
78
|
+
and copy rules, platform requirements — one line each, exact values copied
|
|
79
|
+
verbatim from the spec. Every task's requirements implicitly include this section.>
|
|
80
|
+
|
|
81
|
+
## Execution order
|
|
82
|
+
|
|
83
|
+
| Group | Tasks | Runs after |
|
|
84
|
+
|---|---|---|
|
|
85
|
+
| A | 1, 2 | — |
|
|
86
|
+
| B | 3 | A |
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
## Task structure — required
|
|
92
|
+
|
|
93
|
+
````markdown
|
|
94
|
+
### Task N: <component>
|
|
95
|
+
|
|
96
|
+
**Depends:** [task ids, or —]
|
|
97
|
+
|
|
98
|
+
**Implements:** REQ-003, REQ-007 — *(the brief's requirement ids this task
|
|
99
|
+
delivers, or `—` for pure glue/infrastructure tasks. Quote each REQ's one-line
|
|
100
|
+
statement under the DoD so the zero-context implementer sees the intent, not just
|
|
101
|
+
the instruction.)*
|
|
102
|
+
|
|
103
|
+
**Files:**
|
|
104
|
+
- Create: `exact/path/to/file.py`
|
|
105
|
+
- Modify: `exact/path/to/existing.py:123-145`
|
|
106
|
+
- Test: `tests/exact/path/to/test_file.py`
|
|
107
|
+
|
|
108
|
+
**Interfaces:**
|
|
109
|
+
- Consumes: <what this task uses from earlier tasks — exact signatures>
|
|
110
|
+
- Produces: <what later tasks rely on — exact names, parameter and return types.
|
|
111
|
+
The implementer sees only this task; this block is how they learn the names
|
|
112
|
+
neighboring tasks use.>
|
|
113
|
+
|
|
114
|
+
**Definition of done:** <observable, verifiable conditions — tests green, behavior
|
|
115
|
+
demonstrated, docs updated in this same change>
|
|
116
|
+
|
|
117
|
+
- [ ] **Step 1: write the failing test**
|
|
118
|
+
|
|
119
|
+
```python
|
|
120
|
+
def test_specific_behavior():
|
|
121
|
+
assert function(input) == expected
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
- [ ] **Step 2: run it and confirm it fails**
|
|
125
|
+
|
|
126
|
+
Run: `pytest tests/path/test_file.py::test_specific_behavior -v`
|
|
127
|
+
Expected: FAIL — `NameError: name 'function' is not defined`
|
|
128
|
+
|
|
129
|
+
- [ ] **Step 3: minimal implementation**
|
|
130
|
+
|
|
131
|
+
```python
|
|
132
|
+
def function(value):
|
|
133
|
+
return expected
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
- [ ] **Step 4: run it and confirm it passes**
|
|
137
|
+
|
|
138
|
+
Run: `pytest tests/path/test_file.py::test_specific_behavior -v`
|
|
139
|
+
Expected: PASS
|
|
140
|
+
|
|
141
|
+
- [ ] **Step 5: commit**
|
|
142
|
+
|
|
143
|
+
```bash
|
|
144
|
+
git add tests/path/test_file.py src/path/file.py
|
|
145
|
+
git commit -m "feat: <what changed>"
|
|
146
|
+
```
|
|
147
|
+
````
|
|
148
|
+
|
|
149
|
+
For UI tasks, every task that builds user-facing behavior names the **scenario
|
|
150
|
+
ID(s)** and `SCR-` screen(s) it implements, and its DoD includes satisfying them
|
|
151
|
+
**and** updating the affected super-ux layers in the same change.
|
|
152
|
+
|
|
153
|
+
## No placeholders
|
|
154
|
+
|
|
155
|
+
These are plan failures. Never write them:
|
|
156
|
+
|
|
157
|
+
- "TBD", "TODO", "implement later", "fill in details"
|
|
158
|
+
- "Add appropriate error handling" / "add validation" / "handle edge cases"
|
|
159
|
+
- "Write tests for the above" without the actual test code
|
|
160
|
+
- "Similar to Task N" — repeat the code; tasks get read out of order
|
|
161
|
+
- A step that says what to do without showing how (code steps need code blocks)
|
|
162
|
+
- References to types, functions or methods no task defines
|
|
163
|
+
|
|
164
|
+
## Self-review — before handing off
|
|
165
|
+
|
|
166
|
+
A checklist you run yourself, inline. No subagent:
|
|
167
|
+
|
|
168
|
+
1. **REQ coverage — set equality, not a feeling.** Collect every `Implements:` id
|
|
169
|
+
across all tasks and compare it to the brief's REQ table. The two sets must be
|
|
170
|
+
**equal**: a REQ with no task is scope silently lost; an `Implements:` id that
|
|
171
|
+
isn't in the brief is either a typo or work nobody asked for. Print the
|
|
172
|
+
difference and fix it before anything else — this seam is where scope leaks.
|
|
173
|
+
2. **Spec coverage:** walk each spec requirement. Point at the task that implements
|
|
174
|
+
it. A requirement with no task → add the task.
|
|
175
|
+
3. **Placeholder scan:** search the plan for every pattern above. Fix.
|
|
176
|
+
4. **Name and type consistency:** signatures, property names and types used in
|
|
177
|
+
later tasks match what earlier tasks defined. `clearLayers()` in Task 3 and
|
|
178
|
+
`clearFullLayers()` in Task 7 is a bug, not a style difference.
|
|
179
|
+
5. **Parallel safety:** no two tasks in the same group write the same file; every
|
|
180
|
+
`depends:` points at a task that really produces what's consumed.
|
|
181
|
+
6. **DoD present and verifiable** on every task.
|
|
182
|
+
|
|
183
|
+
## GATE (auto)
|
|
184
|
+
|
|
185
|
+
**Set equality first:** the REQ ids in the brief equal the union of `Implements:`
|
|
186
|
+
across the plan's tasks. A non-empty difference fails the gate and is reported as
|
|
187
|
+
the explicit list of dropped (or invented) requirements — this seam is where scope
|
|
188
|
+
leaks, so the check is mechanical, never a judgement call.
|
|
189
|
+
|
|
190
|
+
Then: every spec requirement maps to a task; no placeholders; names and types
|
|
191
|
+
consistent across tasks; parallel-group tasks share no files; each task has a
|
|
192
|
+
verifiable DoD. UI tasks carry their scenario IDs and `SCR-` screens. Verify all of
|
|
193
|
+
it yourself and stop on failure — this gate has no operator in it.
|
|
@@ -0,0 +1,173 @@
|
|
|
1
|
+
# Review — the rubric and the prompts, built in
|
|
2
|
+
|
|
3
|
+
Used by stage 5's per-task reviews, its scoped re-reviews and its final
|
|
4
|
+
whole-branch review ([`build.md`](build.md)). Built into this skill; nothing to
|
|
5
|
+
install.
|
|
6
|
+
|
|
7
|
+
> Ported from the `requesting-code-review` skill and the reviewer prompts in
|
|
8
|
+
> [obra/superpowers](https://github.com/obra/superpowers) (MIT — see `LICENSE` →
|
|
9
|
+
> *Third-party*), condensed into one rubric plus three copy-paste prompts, with the
|
|
10
|
+
> external helper scripts replaced by plain git commands so the doctrine works on
|
|
11
|
+
> any agent.
|
|
12
|
+
|
|
13
|
+
## The diff package
|
|
14
|
+
|
|
15
|
+
A reviewer never re-derives the diff with a dozen git calls, and the diff never
|
|
16
|
+
enters **your** context. Write it to one file and pass the path (`$WORKSPACE` is
|
|
17
|
+
this plan's git-ignored directory, `.task-pipeline/build/<plan-basename>/` — see
|
|
18
|
+
[`build.md`](build.md)):
|
|
19
|
+
|
|
20
|
+
```bash
|
|
21
|
+
{ git log --oneline "$BASE..$HEAD"
|
|
22
|
+
echo
|
|
23
|
+
git diff --stat "$BASE..$HEAD"
|
|
24
|
+
echo
|
|
25
|
+
git diff -U10 "$BASE..$HEAD"
|
|
26
|
+
} > "$WORKSPACE/review-task-$N-$(git rev-parse --short "$HEAD").md"
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
`BASE` is the commit you recorded **before** dispatching the implementer — never
|
|
30
|
+
`HEAD~1`, which silently drops every commit but the last of a multi-commit task.
|
|
31
|
+
For a scoped re-review, `BASE` is the head the previous review saw. For the final
|
|
32
|
+
review, `BASE` is `git merge-base main HEAD`.
|
|
33
|
+
|
|
34
|
+
Never dispatch a reviewer without a diff file.
|
|
35
|
+
|
|
36
|
+
## Reviewer inputs
|
|
37
|
+
|
|
38
|
+
Four things, three of them paths:
|
|
39
|
+
|
|
40
|
+
1. The **task brief** file — what was required.
|
|
41
|
+
2. The **implementer report** file — what was done, with the test evidence.
|
|
42
|
+
3. The **diff package** path.
|
|
43
|
+
4. The **Global Constraints** that bind this task — copied **verbatim** from the
|
|
44
|
+
plan: exact values, exact formats, stated relationships ("same layout as X",
|
|
45
|
+
"matches Y"). This block is the reviewer's attention lens.
|
|
46
|
+
|
|
47
|
+
## Controller rules
|
|
48
|
+
|
|
49
|
+
- **Never pre-judge.** Don't tell a reviewer to ignore something, don't cap a
|
|
50
|
+
severity, don't explain why a finding would be a false positive. If your prompt
|
|
51
|
+
contains "don't flag", "at most Minor", or "the plan chose this" — stop: you're
|
|
52
|
+
buying yourself a shorter loop with an unreviewed defect. Let the finding come and
|
|
53
|
+
adjudicate it in the loop.
|
|
54
|
+
- **Don't ask a reviewer to re-run tests** the implementer already ran on the same
|
|
55
|
+
code; the report carries the evidence.
|
|
56
|
+
- **No open-ended directives** ("check all uses", "run race tests if useful")
|
|
57
|
+
without a concrete, task-specific reason.
|
|
58
|
+
|
|
59
|
+
## The rubric
|
|
60
|
+
|
|
61
|
+
Review in this order; stop reading the diff only when you've covered all of it.
|
|
62
|
+
|
|
63
|
+
1. **Spec compliance.** Every requirement in the brief: met, partially met, or
|
|
64
|
+
missing. Anything built that the brief did *not* ask for is scope creep — flag
|
|
65
|
+
it, even when it's nice.
|
|
66
|
+
2. **REQ satisfaction.** Read the task's `Implements:` requirement statements, not
|
|
67
|
+
just its instructions, and judge the diff against **those**. A task can satisfy
|
|
68
|
+
every line of its brief and still miss the requirement it exists to deliver —
|
|
69
|
+
that gap is invisible one level down, which is why it is asked for here.
|
|
70
|
+
3. **Correctness.** Logic errors, off-by-one, wrong operator, unhandled `None`/nil,
|
|
71
|
+
race conditions, resource leaks, wrong error propagation. State the concrete
|
|
72
|
+
input or state that produces the wrong output — a finding without a failure
|
|
73
|
+
scenario is an opinion.
|
|
74
|
+
4. **Global constraints.** Exact values, formats and relationships from the
|
|
75
|
+
constraints block. Approximations are failures.
|
|
76
|
+
5. **Test honesty.** Tests assert on real behavior, not on mocks. No test that
|
|
77
|
+
passes regardless of the production code. No `skip`/`xfail`/commented assertion
|
|
78
|
+
smuggling a red suite past a gate. New behavior has a covering test; the failure
|
|
79
|
+
path has one too.
|
|
80
|
+
6. **Error handling and degradation.** Every external call (network, DB, file, MCP,
|
|
81
|
+
API) handles failure, and the failure is reported honestly rather than swallowed.
|
|
82
|
+
7. **Boundaries and clarity.** One responsibility per unit; names that say what the
|
|
83
|
+
thing is; no duplication of a logic block that should be shared; nothing left
|
|
84
|
+
dead.
|
|
85
|
+
8. **Security.** No secrets in code, logs or fixtures; input validated at the
|
|
86
|
+
boundary; no new injection or path-traversal surface.
|
|
87
|
+
9. **Docs in the same change.** Module docs, runbooks and (for UI work) the
|
|
88
|
+
super-ux layers updated alongside the code, not deferred.
|
|
89
|
+
|
|
90
|
+
**Severities:**
|
|
91
|
+
|
|
92
|
+
- **Critical** — wrong behavior, data loss, security hole, a red or dishonest test
|
|
93
|
+
suite. Blocks.
|
|
94
|
+
- **Important** — a real defect or spec gap that will bite: missing requirement,
|
|
95
|
+
unhandled failure path, a magic value the constraints named. Blocks.
|
|
96
|
+
- **Minor** — style, naming, a nit with no behavioral consequence. Never blocks;
|
|
97
|
+
goes to the ledger.
|
|
98
|
+
- **⚠️ Cannot verify from diff** — the requirement lives in unchanged code or spans
|
|
99
|
+
tasks. Not a blocker for the reviewer; the controller resolves it.
|
|
100
|
+
|
|
101
|
+
Formatting nits that don't change meaning are not findings. Praise is not a
|
|
102
|
+
finding either.
|
|
103
|
+
|
|
104
|
+
## Prompt — task review
|
|
105
|
+
|
|
106
|
+
> You are reviewing one task of an implementation plan. Read, in order:
|
|
107
|
+
> `<brief path>` (the requirements), `<report path>` (what the implementer did and
|
|
108
|
+
> the test evidence), `<diff package path>` (commits, stat, full diff).
|
|
109
|
+
>
|
|
110
|
+
> Global constraints binding this task:
|
|
111
|
+
> ```
|
|
112
|
+
> <verbatim block>
|
|
113
|
+
> ```
|
|
114
|
+
>
|
|
115
|
+
> Produce three verdicts, all required:
|
|
116
|
+
> 1. **Spec compliance:** ✅ or ❌. List every requirement as met / partial /
|
|
117
|
+
> missing, and list anything implemented that was not asked for.
|
|
118
|
+
> 2. **REQ satisfied:** ✅ or ❌ per `Implements:` id. The brief quotes each
|
|
119
|
+
> requirement's statement — judge the diff against **that statement**, not
|
|
120
|
+
> against the task's instructions. A task can follow every instruction and still
|
|
121
|
+
> miss the requirement it exists to deliver; say so when it does.
|
|
122
|
+
> 3. **Code quality:** approved or not. Findings only, each as
|
|
123
|
+
> `severity — file:line — the defect — the failure scenario (concrete input or
|
|
124
|
+
> state → wrong result)`. Severities: Critical, Important, Minor. Use
|
|
125
|
+
> `⚠️ cannot verify from diff` for anything you can't judge from the diff alone.
|
|
126
|
+
>
|
|
127
|
+
> Review against the rubric: correctness, global constraints, test honesty (tests
|
|
128
|
+
> assert real behavior, not mocks; no skipped/empty assertions), error handling and
|
|
129
|
+
> honest degradation, boundaries and naming, security and secrets, docs updated in
|
|
130
|
+
> the same change. No praise, no formatting nits that don't change meaning. Do not
|
|
131
|
+
> re-run the tests the report already covers. Return the verdicts and findings as
|
|
132
|
+
> your final message — nothing else.
|
|
133
|
+
|
|
134
|
+
## Prompt — scoped re-review
|
|
135
|
+
|
|
136
|
+
> A previous review of this task raised the findings below. The implementer has
|
|
137
|
+
> since pushed fixes. Read `<brief path>`, `<report path>` (its fix report is
|
|
138
|
+
> appended at the end) and `<fix diff package path>` — the fix diff **only**.
|
|
139
|
+
>
|
|
140
|
+
> Open findings:
|
|
141
|
+
> ```
|
|
142
|
+
> 1. <finding>
|
|
143
|
+
> 2. <finding>
|
|
144
|
+
> ```
|
|
145
|
+
>
|
|
146
|
+
> For each finding return `ADDRESSED` (with the `file:line` that resolves it) or
|
|
147
|
+
> `NOT ADDRESSED` (with what's still wrong). Then flag any **new** Critical or
|
|
148
|
+
> Important breakage introduced by this fix diff. Out-of-scope observations about
|
|
149
|
+
> code this diff didn't touch: list them separately as deferred minors — they are
|
|
150
|
+
> not part of this verdict. End with: all findings addressed / N still open.
|
|
151
|
+
|
|
152
|
+
## Prompt — final whole-branch review
|
|
153
|
+
|
|
154
|
+
> Review this entire branch before merge. Read `<diff package path>` (merge-base to
|
|
155
|
+
> HEAD) and `<spec path>`.
|
|
156
|
+
>
|
|
157
|
+
> Findings deferred or parked during implementation:
|
|
158
|
+
> ```
|
|
159
|
+
> <the ledger's minor + parked lines>
|
|
160
|
+
> ```
|
|
161
|
+
>
|
|
162
|
+
> Judge the branch as a whole: does it deliver the spec; do the pieces fit; is
|
|
163
|
+
> anything half-migrated, duplicated across tasks, or left dead; are the tests
|
|
164
|
+
> honest and the suite genuinely green; is error handling consistent; are docs in
|
|
165
|
+
> sync. Triage the deferred/parked list: which of those must be fixed before merge,
|
|
166
|
+
> which can stand and why. Findings only, with severity and a concrete failure
|
|
167
|
+
> scenario each. Return the findings as your final message.
|
|
168
|
+
|
|
169
|
+
Run the final review on the **run's confirmed model** like everything else
|
|
170
|
+
([`model-tiering.md`](model-tiering.md)). It is the one review that sees the whole
|
|
171
|
+
change, so if the run is on a tier below the most capable one available, say so and
|
|
172
|
+
offer to escalate just this dispatch — a recommendation stated out loud, never a
|
|
173
|
+
silent switch (`build.md` → *Models*).
|