task-pipeline-skill 0.10.0 → 0.17.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +401 -0
- package/LICENSE +85 -0
- package/README.md +211 -84
- package/cursor/rules/task-pipeline.mdc +135 -20
- package/package.json +3 -3
- package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
- package/plugins/task-pipeline/commands/task-pipeline.md +16 -8
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +139 -57
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +41 -27
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +32 -6
- package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +73 -21
- package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +169 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +55 -22
- package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +184 -48
- package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +11 -3
- package/plugins/task-pipeline/skills/task-pipeline/templates/adr.md +64 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +49 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/context.md +87 -0
|
@@ -8,11 +8,17 @@ shape.
|
|
|
8
8
|
## In the host project
|
|
9
9
|
|
|
10
10
|
```
|
|
11
|
+
CONTEXT.md # stage 0 — domain glossary, written inline as terms resolve
|
|
11
12
|
docs/
|
|
13
|
+
adr/
|
|
14
|
+
NNNN-<slug>.md # stage 0 — ADRs for hard-to-reverse decisions
|
|
12
15
|
superpowers/
|
|
13
16
|
specs/
|
|
14
17
|
YYYY-MM-DD-<topic>-brief.md # stage 0 — locked intake brief (grill output)
|
|
15
|
-
YYYY-MM-DD-<topic>-
|
|
18
|
+
YYYY-MM-DD-<topic>-carryover.md # stage 0 seeds it; EVERY stage appends; stage 10 reads it
|
|
19
|
+
YYYY-MM-DD-<topic>-modules.md # stage 2 — module map + build order (platforms only)
|
|
20
|
+
YYYY-MM-DD-<topic>-design.md # stage 3 — the spec / module dossier (locks shared contracts)
|
|
21
|
+
YYYY-MM-DD-<topic>-acceptance.md # stage 10 — REQ coverage table + evidence
|
|
16
22
|
plans/
|
|
17
23
|
YYYY-MM-DD-<topic>.md # stage 4 — the implementation plan
|
|
18
24
|
ux/ # super-ux, UI tasks only (see companion-skills.md)
|
|
@@ -26,15 +32,33 @@ docs/
|
|
|
26
32
|
```
|
|
27
33
|
|
|
28
34
|
Naming: date-prefixed `YYYY-MM-DD-<topic>` slugs, one topic per file, kebab-case.
|
|
29
|
-
|
|
30
|
-
|
|
35
|
+
Brief, carry-over, design, plan and acceptance share the **same `<topic>` slug**, so the chain is traceable
|
|
36
|
+
at a glance.
|
|
37
|
+
|
|
38
|
+
> The `docs/superpowers/` directory name is this pipeline's historical convention
|
|
39
|
+
> (kept so existing projects don't have to migrate) — **not a dependency on any
|
|
40
|
+
> external skill**. A host project may relocate the root via its `CLAUDE.md`; keep
|
|
41
|
+
> the shape, keep the slugs.
|
|
42
|
+
|
|
43
|
+
Loop-bearing runs also keep a **git-ignored** run ledger at `.task-pipeline/run.md` —
|
|
44
|
+
stage-level and program-level repeat touches, one line each, so the loop guard can
|
|
45
|
+
detect churn after a lost context (see [`loop-guard.md`](loop-guard.md)).
|
|
46
|
+
|
|
47
|
+
Stage 5 also creates a **git-ignored** scratch workspace per plan at
|
|
48
|
+
`.task-pipeline/build/<plan-basename>/` — ledger, task briefs, implementer reports,
|
|
49
|
+
review packages. It is deleted when the final review is clean; git history is the
|
|
50
|
+
record (see `build.md`).
|
|
31
51
|
|
|
32
52
|
## Stage → artifact map
|
|
33
53
|
|
|
34
54
|
| Stage | Writes | Consumed by |
|
|
35
55
|
|---|---|---|
|
|
36
|
-
| 0 Intake | `specs/<topic>-brief.md` (seed from
|
|
37
|
-
|
|
|
56
|
+
| 0 Intake | `specs/<topic>-brief.md` — incl. the **REQ table** (seed from `templates/brief.md`) | stages 2–5, 7, 10 |
|
|
57
|
+
| 0→10 all | `specs/<topic>-carryover.md` — append-only ledger (seed from `templates/carryover.md`) | stage 10, in full |
|
|
58
|
+
| 10 Acceptance | `specs/<topic>-acceptance.md` — every REQ with a status and evidence | the operator |
|
|
59
|
+
| 0 Grill (domain) | `CONTEXT.md`, `docs/adr/NNNN-<slug>.md` — created **lazily**, only when a term resolves or a decision qualifies | stages 2–4 + the repo |
|
|
60
|
+
| 2 Decompose | `specs/<topic>-modules.md` — module map, build order, contracts, per-module status (platforms only) | stages 3–10, every module's run |
|
|
61
|
+
| 3 Spec | `specs/<topic>-design.md` — module dossier for a decomposed platform (+ links `docs/ux/*` for UI) | stage 4 |
|
|
38
62
|
| 4 Plan | `plans/<topic>.md` | stage 5 |
|
|
39
63
|
| 3 UX track | `docs/ux/{foundation,flows,screens,scenarios}.md` | stages 4–9 + `/ux-lint` |
|
|
40
64
|
| 8 Post-deploy | log/health notes (in the run, not a committed file) | stage 9 |
|
|
@@ -51,9 +75,11 @@ plugins/task-pipeline/
|
|
|
51
75
|
SKILL.md
|
|
52
76
|
pipeline.schema.json # generic pipeline contract
|
|
53
77
|
pipeline.example.json # this plugin's own flow, as config
|
|
78
|
+
references/{grill,brainstorm,decomposition,spec,planning,build,review,tdd,acceptance}.md # built-in stage doctrine
|
|
79
|
+
references/loop-guard.md # cross-cutting: churn detection + break protocol
|
|
54
80
|
references/{stages,model-tiering,conventions,artifacts,companion-skills}.md
|
|
55
81
|
cursor/rules/task-pipeline.mdc # Cursor channel (self-contained rule)
|
|
56
|
-
plugins/task-pipeline/skills/task-pipeline/templates/brief.md # stage-0
|
|
82
|
+
plugins/task-pipeline/skills/task-pipeline/templates/{brief,carryover,context,adr}.md # stage-0 skeletons (ship on every channel)
|
|
57
83
|
bin/task-pipeline.js # npx installer (package task-pipeline-skill)
|
|
58
84
|
package.json
|
|
59
85
|
install.sh # POSIX installer
|
|
@@ -0,0 +1,106 @@
|
|
|
1
|
+
# Brainstorm — stage 2, built in
|
|
2
|
+
|
|
3
|
+
The design conversation is **part of this skill**. No companion to install, no
|
|
4
|
+
provider to resolve, nothing to fall back to: this file is the implementation.
|
|
5
|
+
|
|
6
|
+
Stage 0 locked *what* is being built. Stage 2 decides *how*, and stops at an
|
|
7
|
+
approved design — not at code.
|
|
8
|
+
|
|
9
|
+
> Ported, with thanks, from the `brainstorming` skill in
|
|
10
|
+
> [obra/superpowers](https://github.com/obra/superpowers) (MIT — see this repo's
|
|
11
|
+
> `LICENSE` → *Third-party*), rewritten for this pipeline: the brief is the input,
|
|
12
|
+
> the UI verdict is a required output, and the spec write-up moved to stage 3
|
|
13
|
+
> ([`spec.md`](spec.md)).
|
|
14
|
+
|
|
15
|
+
## The hard gate
|
|
16
|
+
|
|
17
|
+
**No implementation action before the operator approves a design.** No code, no
|
|
18
|
+
scaffolding, no file creation "to see how it'd look", no invoking a
|
|
19
|
+
frontend/backend/build skill. This holds for every task regardless of how simple it
|
|
20
|
+
looks.
|
|
21
|
+
|
|
22
|
+
**"Too simple to need a design" is the trap, not the exception.** A one-function
|
|
23
|
+
utility, a config flip, a copy change — all of them go through this stage. Simple
|
|
24
|
+
tasks are where unexamined assumptions survive longest. The design may be three
|
|
25
|
+
sentences; it still gets presented and approved.
|
|
26
|
+
|
|
27
|
+
## Input: the brief, not a blank page
|
|
28
|
+
|
|
29
|
+
Read the stage-0 brief first (`…-brief.md`). Everything it locked — scope, users,
|
|
30
|
+
constraints, done-criteria, the autonomy sweep — is **settled**. Re-asking a
|
|
31
|
+
question the grill already answered is the single most common way to waste this
|
|
32
|
+
stage. If the brief and the codebase disagree, that's a contradiction to surface,
|
|
33
|
+
not a question to re-open from scratch.
|
|
34
|
+
|
|
35
|
+
## The loop
|
|
36
|
+
|
|
37
|
+
1. **Explore the current state.** Files, module docs, recent commits, the
|
|
38
|
+
conventions the repo already follows. Do this before asking anything.
|
|
39
|
+
2. **Scope check, early.** If the task actually describes several independent
|
|
40
|
+
subsystems, say so immediately and help decompose it into sub-projects: what the
|
|
41
|
+
independent pieces are, how they relate, what order they get built in. Then
|
|
42
|
+
brainstorm the first one. Each sub-project gets its own spec → plan → build
|
|
43
|
+
cycle. Don't refine details of something that needs splitting first.
|
|
44
|
+
3. **Questions one at a time.** Never bundle. Multiple choice where it fits, open
|
|
45
|
+
where it doesn't. Purpose, constraints, success criteria — anything the brief
|
|
46
|
+
left at design level.
|
|
47
|
+
4. **Propose 2–3 approaches with trade-offs**, lead with your recommendation and
|
|
48
|
+
the reason for it. **YAGNI ruthlessly** — strip anything the task doesn't need
|
|
49
|
+
from every option before presenting.
|
|
50
|
+
5. **Present the design in sections**, each scaled to its complexity (a couple of
|
|
51
|
+
sentences when it's straightforward, up to a few hundred words when it's
|
|
52
|
+
genuinely nuanced). Ask after each section whether it holds. Cover:
|
|
53
|
+
architecture, components, data flow, error handling and degradation, testing.
|
|
54
|
+
6. **Go back when something doesn't fit.** A revised section beats a design that
|
|
55
|
+
was approved because it was hard to argue with.
|
|
56
|
+
|
|
57
|
+
## Design for isolation and clarity
|
|
58
|
+
|
|
59
|
+
- Break the system into units with **one clear purpose each**, communicating
|
|
60
|
+
through well-defined interfaces, understandable and testable on their own.
|
|
61
|
+
- For every unit you should be able to answer: what does it do, how is it used,
|
|
62
|
+
what does it depend on?
|
|
63
|
+
- Can a reader understand a unit without reading its internals? Can the internals
|
|
64
|
+
change without breaking consumers? If not, the boundaries need work.
|
|
65
|
+
- Smaller focused files are also what the *implementer* (often a subagent with a
|
|
66
|
+
narrow context) handles reliably. A file growing large is usually a signal it
|
|
67
|
+
does too much.
|
|
68
|
+
|
|
69
|
+
## Working in an existing codebase
|
|
70
|
+
|
|
71
|
+
- Explore the structure before proposing changes; follow the patterns already
|
|
72
|
+
there.
|
|
73
|
+
- Where existing code genuinely blocks the work — a file that's grown unwieldy,
|
|
74
|
+
tangled responsibilities, an unclear boundary the change has to cross — include
|
|
75
|
+
the targeted improvement in the design, the way a careful developer improves code
|
|
76
|
+
they're working in.
|
|
77
|
+
- Don't propose unrelated refactoring. Anything out of scope goes to the backlog,
|
|
78
|
+
not into this design.
|
|
79
|
+
|
|
80
|
+
## UI detection — a required output
|
|
81
|
+
|
|
82
|
+
One branch is always: **does this touch a user-facing surface** (web, mobile, CLI,
|
|
83
|
+
TUI — a screen, a command, a visible behavior)? Stage 0 usually answered it; this
|
|
84
|
+
stage confirms it against the design that actually emerged. Record the verdict —
|
|
85
|
+
it arms the stage-3 UX track ([`spec.md`](spec.md) → *UX track*). When it's
|
|
86
|
+
genuinely borderline, record "yes": a false positive costs one extra chain, a false
|
|
87
|
+
negative ships an unspecified interface.
|
|
88
|
+
|
|
89
|
+
## GATE (manual)
|
|
90
|
+
|
|
91
|
+
The operator approves the design **and** the UI verdict is recorded **and every REQ
|
|
92
|
+
in the brief is answered by the design** — a requirement the design doesn't address
|
|
93
|
+
is either covered now or explicitly dropped by the operator, with the drop written
|
|
94
|
+
into the carry-over ledger. For a platform, the module map
|
|
95
|
+
([`decomposition.md`](decomposition.md)) is committed and approved as part of this
|
|
96
|
+
same gate. Then, and only then, stage 3 writes it up.
|
|
97
|
+
|
|
98
|
+
## Rationalizations
|
|
99
|
+
|
|
100
|
+
| Excuse | Reality |
|
|
101
|
+
|---|---|
|
|
102
|
+
| "The brief already says everything" | The brief locks *what*. If it also locked *how*, say so in one line and get the approval anyway — the gate is the point. |
|
|
103
|
+
| "It's a one-line change, design is ceremony" | Then the design is one line. Present it. |
|
|
104
|
+
| "I'll scaffold while they think" | Scaffolding is implementation. The gate is before it, not around it. |
|
|
105
|
+
| "Both approaches are fine, let them pick" | You read the codebase, they didn't. Recommend, then let them override. |
|
|
106
|
+
| "I'll add the extra option now, it's cheap" | YAGNI. Every unused branch is code someone maintains and a test someone writes. |
|
|
@@ -0,0 +1,364 @@
|
|
|
1
|
+
# Build — stage 5, built in
|
|
2
|
+
|
|
3
|
+
Executing the plan: an isolated workspace, one fresh implementer subagent per task,
|
|
4
|
+
a review after each task, a whole-branch review at the end. Built into this skill;
|
|
5
|
+
nothing to install.
|
|
6
|
+
|
|
7
|
+
> Ported from the `using-git-worktrees` and `subagent-driven-development` skills in
|
|
8
|
+
> [obra/superpowers](https://github.com/obra/superpowers) (MIT — see `LICENSE` →
|
|
9
|
+
> *Third-party*), rewritten for this pipeline: no external scripts, the ledger
|
|
10
|
+
> lives under `.task-pipeline/`, and model choice defers to the run's single
|
|
11
|
+
> confirmed model ([`model-tiering.md`](model-tiering.md)).
|
|
12
|
+
|
|
13
|
+
**Why subagents:** each task goes to an agent with a constructed context — its
|
|
14
|
+
brief, its interfaces, the global constraints, nothing else. It never inherits the
|
|
15
|
+
session's history, so it stays focused; your context stays free for coordination.
|
|
16
|
+
|
|
17
|
+
**Continuous execution:** don't check in between tasks. The operator asked for the
|
|
18
|
+
plan to be executed — execute it. Stop only for BLOCKED you can't resolve, a
|
|
19
|
+
genuine ambiguity, or completion. "Should I continue?" between tasks is noise.
|
|
20
|
+
|
|
21
|
+
**Narration:** at most one short line between tool calls. The ledger and the tool
|
|
22
|
+
results are the record.
|
|
23
|
+
|
|
24
|
+
**No subagents available?** (a harness without them, or a plan so small that
|
|
25
|
+
dispatching costs more than it saves) — run the same loop inline: same isolation,
|
|
26
|
+
same ledger, same TDD per task, and after each task review your own diff against
|
|
27
|
+
the rubric in [`review.md`](review.md) before moving on. What changes is who does
|
|
28
|
+
the work; the gates, the artifacts and the review discipline do not. Say plainly
|
|
29
|
+
that the run is inline, since a self-review is weaker evidence than a fresh
|
|
30
|
+
reviewer's.
|
|
31
|
+
|
|
32
|
+
## 1. Isolation
|
|
33
|
+
|
|
34
|
+
Work never starts on `main`/`master` without the operator's explicit consent
|
|
35
|
+
(the stage-0 brief usually records the branch policy — read it, don't re-ask).
|
|
36
|
+
|
|
37
|
+
**Detect existing isolation first:**
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
GIT_DIR=$(cd "$(git rev-parse --git-dir)" && pwd -P)
|
|
41
|
+
GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" && pwd -P)
|
|
42
|
+
git rev-parse --show-superproject-working-tree # non-empty ⇒ submodule, not a worktree
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
- `GIT_DIR != GIT_COMMON` and **not** a submodule → you are already in a linked
|
|
46
|
+
worktree. Do not create another. Report the path and branch, go to *Setup*.
|
|
47
|
+
- Otherwise you are in a normal checkout. Honor the brief's worktree preference; if
|
|
48
|
+
none was recorded, ask once before creating one.
|
|
49
|
+
|
|
50
|
+
**Creating one — native tool first.** If the harness offers a worktree tool
|
|
51
|
+
(`EnterWorktree`, a `/worktree` command, a `--worktree` flag), use it: it owns
|
|
52
|
+
placement, branch creation and cleanup. Reaching for raw `git worktree add` when a
|
|
53
|
+
native tool exists creates state the harness can't see or clean up.
|
|
54
|
+
|
|
55
|
+
**Git fallback**, only when there is no native tool:
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
# both scratch roots MUST be ignored before anything is created
|
|
59
|
+
git check-ignore -q .worktrees || printf '.worktrees/\n' >> .gitignore
|
|
60
|
+
git check-ignore -q .task-pipeline || printf '.task-pipeline/\n' >> .gitignore
|
|
61
|
+
git diff --quiet .gitignore || git commit -m "chore: ignore build scratch dirs" .gitignore
|
|
62
|
+
git worktree add ".worktrees/$BRANCH" -b "$BRANCH"
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
Directory priority: an explicit operator preference → an existing `.worktrees/` →
|
|
66
|
+
an existing `worktrees/` → default `.worktrees/`. An unignored worktree directory
|
|
67
|
+
commits the entire tree into the repo — verify before creating. If creation fails
|
|
68
|
+
on a sandbox permission error, say so plainly and work in place.
|
|
69
|
+
|
|
70
|
+
**Setup + baseline.** Install dependencies the way the project does (`npm install`,
|
|
71
|
+
`cargo build`, `pip install -r requirements.txt`, `poetry install`, `go mod
|
|
72
|
+
download`), then run the test command from the brief's autonomy sweep. A dirty
|
|
73
|
+
baseline makes every later failure ambiguous: report failures and let the operator
|
|
74
|
+
decide whether to proceed.
|
|
75
|
+
|
|
76
|
+
## 2. Workspace and ledger
|
|
77
|
+
|
|
78
|
+
Conversation memory does not survive compaction. A controller that lost its place
|
|
79
|
+
re-dispatches completed tasks — the most expensive failure this stage has.
|
|
80
|
+
**Track progress in a file, not only in todos.**
|
|
81
|
+
|
|
82
|
+
- Each plan owns a git-ignored workspace: `.task-pipeline/build/<plan-basename>/`
|
|
83
|
+
at the repo root. Everything for THIS plan lives there — ledger, task briefs,
|
|
84
|
+
implementer reports, review packages. Another plan's directory is never yours to
|
|
85
|
+
read or write. `.task-pipeline/` must be git-ignored — the isolation step above
|
|
86
|
+
adds and commits it; if you skipped that step, do it now, in its own commit, so
|
|
87
|
+
scratch files never land in a task's diff.
|
|
88
|
+
- Ledger: `<workspace>/progress.md`, first line = its identity:
|
|
89
|
+
`# build ledger — plan: <plan file path>`.
|
|
90
|
+
- **Resuming:** a task with a `Task <N>: complete` line is DONE — never
|
|
91
|
+
re-dispatch it; resume at the first task without one. A task whose last line is a
|
|
92
|
+
fix round is mid-loop: continue at the next round. A ledger naming a different
|
|
93
|
+
plan belongs to that plan — leave it and start your own.
|
|
94
|
+
- After compaction, trust the ledger and `git log` over your recollection: the
|
|
95
|
+
commits it names exist even when your context no longer remembers them.
|
|
96
|
+
|
|
97
|
+
Read the plan **once**, note its context and Global Constraints, create a todo per
|
|
98
|
+
task.
|
|
99
|
+
|
|
100
|
+
**Pre-flight conflict scan.** Before Task 1, scan the plan for tasks that
|
|
101
|
+
contradict each other or the Global Constraints, and for anything the plan mandates
|
|
102
|
+
that the review rubric ([`review.md`](review.md)) treats as a defect. Present
|
|
103
|
+
everything you find as **one batched question** — each finding beside the plan text
|
|
104
|
+
that mandates it, asking which governs. Clean scan → proceed silently.
|
|
105
|
+
|
|
106
|
+
## 3. Models
|
|
107
|
+
|
|
108
|
+
**Default: the run's one confirmed model** ([`model-tiering.md`](model-tiering.md))
|
|
109
|
+
for every subagent — implementers, reviewers, fixers. Pin it explicitly on each
|
|
110
|
+
dispatch; an omitted model silently inherits the session's and defeats whatever the
|
|
111
|
+
operator recorded.
|
|
112
|
+
|
|
113
|
+
**Deviate only from the operator's recorded override map.** If the stage-0 brief
|
|
114
|
+
carries per-stage or per-role overrides, apply them: mechanical transcription tasks
|
|
115
|
+
(the plan carries the complete code, 1–2 files) can take a cheaper tier, while
|
|
116
|
+
integration, design and review work stays on the confirmed model. No map recorded →
|
|
117
|
+
no deviation, and never a silent downgrade. Turn count beats token price: the
|
|
118
|
+
cheapest tier routinely takes 2–3× the turns on multi-step work and costs more
|
|
119
|
+
overall.
|
|
120
|
+
|
|
121
|
+
**Two moments deserve more capability than the run's default** — both are
|
|
122
|
+
*recommendations you state out loud*, never silent switches:
|
|
123
|
+
|
|
124
|
+
- **The final whole-branch review.** If the run is on a tier below the most capable
|
|
125
|
+
one available, say so and offer to run this one review there; if the operator
|
|
126
|
+
declines or the tier doesn't exist, run it on the confirmed model and note it.
|
|
127
|
+
- **Fix-loop rounds 4–5.** One tier above the implementer that got stuck, when the
|
|
128
|
+
environment has one and the override map or the operator allows it; otherwise
|
|
129
|
+
say so and rely on fresh eyes alone.
|
|
130
|
+
|
|
131
|
+
## 4. The task loop
|
|
132
|
+
|
|
133
|
+
Everything you paste into a dispatch prompt — and everything a subagent prints back
|
|
134
|
+
— stays in your context for the rest of the session. **Hand artifacts over as
|
|
135
|
+
files.**
|
|
136
|
+
|
|
137
|
+
### 4.1 Dispatch the implementer
|
|
138
|
+
|
|
139
|
+
Record `BASE=$(git rev-parse HEAD)` before dispatching; the review package and the
|
|
140
|
+
fix-round diffs need it.
|
|
141
|
+
|
|
142
|
+
**Write the task brief to a file** — extract the task's full text from the plan to
|
|
143
|
+
`<workspace>/task-<N>-brief.md`. The brief is the single source of requirements;
|
|
144
|
+
exact values (numbers, strings, signatures, test cases) live **only** there.
|
|
145
|
+
Include the task's `Implements:` ids **with each REQ's one-line statement quoted
|
|
146
|
+
verbatim** — an implementer who sees only an instruction optimises the
|
|
147
|
+
instruction; one who sees the requirement behind it catches the case the
|
|
148
|
+
instruction didn't cover.
|
|
149
|
+
|
|
150
|
+
The dispatch prompt contains exactly five things:
|
|
151
|
+
|
|
152
|
+
1. One line on where this task fits in the project.
|
|
153
|
+
2. The brief path — "read this first; it is your requirements, use its values
|
|
154
|
+
verbatim".
|
|
155
|
+
3. Interfaces and decisions from earlier tasks the brief can't know.
|
|
156
|
+
4. Your resolution of any ambiguity you spotted in the brief.
|
|
157
|
+
5. The report path (`<workspace>/task-<N>-report.md`) and the report contract.
|
|
158
|
+
|
|
159
|
+
Never paste accumulated history ("state after tasks 1–3") into later dispatches.
|
|
160
|
+
Never make a subagent read the whole plan. If an earlier task parked a finding in
|
|
161
|
+
the area this task touches, carry a pointer to that ledger line.
|
|
162
|
+
|
|
163
|
+
Record the implementer's agent identity: fix rounds 1–3 resume it.
|
|
164
|
+
|
|
165
|
+
**Implementer contract** (put this in the prompt):
|
|
166
|
+
|
|
167
|
+
> Read `<brief path>` first — it is your requirements. Work TDD, and no production
|
|
168
|
+
> code exists before a test you **watched fail**: write the failing test → run it
|
|
169
|
+
> and confirm it fails for the right reason → write the minimal code that passes →
|
|
170
|
+
> run it and confirm it passes, with the rest of the suite still green → commit.
|
|
171
|
+
> Assert on real behavior, never on mock behavior. Commit as you go, conventional
|
|
172
|
+
> commits. When done, self-review your diff, then write the full report to
|
|
173
|
+
> `<report path>`:
|
|
174
|
+
> what you built, the files touched, the commits, the test command and its
|
|
175
|
+
> output, decisions you made, anything you're unsure about. Return **only**:
|
|
176
|
+
> status (`DONE` / `DONE_WITH_CONCERNS` / `NEEDS_CONTEXT` / `BLOCKED`), the
|
|
177
|
+
> commit range, a one-line test summary, and your concerns. Ask before starting
|
|
178
|
+
> if anything in the brief is ambiguous — questions are cheaper than rework.
|
|
179
|
+
|
|
180
|
+
### 4.2 Parallel groups — when fan-out is allowed
|
|
181
|
+
|
|
182
|
+
The plan's parallel groups ([`planning.md`](planning.md)) describe what *may* run
|
|
183
|
+
concurrently. Whether it actually does is this stage's call, and the constraint is
|
|
184
|
+
physical: **two implementers writing one working tree corrupt each other's state.**
|
|
185
|
+
|
|
186
|
+
- **Default: sequential.** One implementer at a time, review after each. Correct for
|
|
187
|
+
every group, and always correct when the tasks are small.
|
|
188
|
+
- **Fan out only when all three hold:** the tasks are in the same group (no
|
|
189
|
+
`depends:` between them), their file ownership is exclusive per the plan, and
|
|
190
|
+
**each implementer gets its own isolated worktree**. Then dispatch them together,
|
|
191
|
+
review each one against its own diff, and integrate the worktrees back to the
|
|
192
|
+
build branch one at a time, running the suite after each merge.
|
|
193
|
+
- **Any conflict on integration** means the plan's file ownership was wrong: stop
|
|
194
|
+
fanning out, finish the group sequentially, and record it in the ledger.
|
|
195
|
+
- Never fan out the fix loop — a task under repair belongs to one implementer.
|
|
196
|
+
|
|
197
|
+
### 4.3 Handle the report
|
|
198
|
+
|
|
199
|
+
| Status | Action |
|
|
200
|
+
|---|---|
|
|
201
|
+
| `DONE` | Build the review package, dispatch the task review ([`review.md`](review.md)). |
|
|
202
|
+
| `DONE_WITH_CONCERNS` | Read the concerns first. Correctness or scope → resolve before review. Observations ("this file is getting large") → **append to the carry-over ledger**, then proceed. A concern that stays only in the report dies with the workspace. |
|
|
203
|
+
| `NEEDS_CONTEXT` | Supply exactly what's missing, re-dispatch. |
|
|
204
|
+
| `BLOCKED` | Diagnose: missing context → re-dispatch with it; needs more reasoning → a more capable model; too large → split the task; the plan itself is wrong → escalate to the operator. |
|
|
205
|
+
|
|
206
|
+
**Never** ignore an escalation, and never re-dispatch the same model with the same
|
|
207
|
+
prompt after a BLOCKED. If the implementer says it's stuck, something must change.
|
|
208
|
+
If the implementer asks a question — before or mid-task — answer it completely; do
|
|
209
|
+
not rush it into implementation.
|
|
210
|
+
|
|
211
|
+
### 4.4 Review the task
|
|
212
|
+
|
|
213
|
+
Every task gets a review with **all three** verdicts — spec compliance, **REQ
|
|
214
|
+
satisfied**, and code quality. The implementer's self-review never substitutes for
|
|
215
|
+
it. Rubric, inputs, prompt templates and how to build the diff package:
|
|
216
|
+
[`review.md`](review.md).
|
|
217
|
+
|
|
218
|
+
The REQ verdict is the one the other two can't produce: a task can meet every line
|
|
219
|
+
of its brief and still miss the requirement it was written to deliver. A ❌ there
|
|
220
|
+
enters the fix loop like any Important finding.
|
|
221
|
+
|
|
222
|
+
A review may report **"cannot verify from diff"** items — requirements that live in
|
|
223
|
+
unchanged code or span tasks. They don't block the review, but you resolve each one
|
|
224
|
+
yourself before completing the task; you hold the cross-task context the reviewer
|
|
225
|
+
lacks. A confirmed gap becomes a failed spec review and enters the fix loop.
|
|
226
|
+
|
|
227
|
+
### 4.5 The fix loop
|
|
228
|
+
|
|
229
|
+
Triggered by: spec ❌, any Critical or Important finding, or a "cannot verify" item
|
|
230
|
+
you confirmed as a real gap.
|
|
231
|
+
|
|
232
|
+
Two routes leave before the loop starts:
|
|
233
|
+
|
|
234
|
+
- **Minor findings** never enter it. Record each in the ledger
|
|
235
|
+
(`Task <N>: minor (deferred): <one-liner>`) and point the final review at that
|
|
236
|
+
list. A roll-up nobody reads is a silent discard.
|
|
237
|
+
- **A finding that conflicts with what the plan mandates** is the operator's
|
|
238
|
+
call: present the finding beside the plan text and ask which governs. Don't
|
|
239
|
+
dismiss the finding because the plan mandated it; don't fix against the plan
|
|
240
|
+
without asking.
|
|
241
|
+
|
|
242
|
+
Everything else loops. One round = one fix dispatch + one scoped re-review.
|
|
243
|
+
**Five rounds maximum per task.**
|
|
244
|
+
|
|
245
|
+
**The loop guard runs alongside the counter** ([`loop-guard.md`](loop-guard.md)):
|
|
246
|
+
log every repeat touch (`touch: <file> — round N — reason: <finding id>`) and trip
|
|
247
|
+
*before* the cap when a fix undoes an earlier fix, when the same file returns for
|
|
248
|
+
the same reason, or when a finding already ADDRESSED reappears. A tripped guard is
|
|
249
|
+
not another round: stop, name the two shapes, escalate to the layer that owns the
|
|
250
|
+
conflict, then re-check in a planned order.
|
|
251
|
+
|
|
252
|
+
- **Rounds 1–3:** resume the original implementer with the open findings verbatim —
|
|
253
|
+
its context is intact. If the harness can't message a live subagent, dispatch a
|
|
254
|
+
fresh one with the brief path, the report path and the findings; the report file
|
|
255
|
+
is the persistent memory either way.
|
|
256
|
+
- **Rounds 4–5:** fresh implementer, one tier up if available, framed as: "a prior
|
|
257
|
+
implementer attempted this task N times; you own it now — read the report file
|
|
258
|
+
for what was tried." A loop that survives three resumes usually means the
|
|
259
|
+
implementer can't see its own problem.
|
|
260
|
+
- **Every round:** the implementer fixes, re-runs the tests covering the amended
|
|
261
|
+
code, appends its fix report to the same report file, returns the short contract.
|
|
262
|
+
Before re-dispatching the reviewer, confirm the fix report names the covering
|
|
263
|
+
tests, the command run and the output.
|
|
264
|
+
- **The re-review is scoped** to the fix diff (`FIX_BASE`..`HEAD`, where `FIX_BASE`
|
|
265
|
+
is the head the previous review saw). It verdicts each finding ADDRESSED / NOT
|
|
266
|
+
ADDRESSED and flags new breakage in the fix diff only. New Critical/Important
|
|
267
|
+
breakage joins the open list; out-of-scope observations go to the ledger as
|
|
268
|
+
deferred minors — they never extend the loop.
|
|
269
|
+
- **Ledger, every round:**
|
|
270
|
+
`Task <N>: fix round <R>/5 (<X> addressed, <Y> open — <one-liners>; commits <a7>..<b7>)`
|
|
271
|
+
|
|
272
|
+
**In a subagent run, never fix findings yourself in the controller session** —
|
|
273
|
+
controller fixes skip review and pollute the context you need for coordination. In a
|
|
274
|
+
declared inline run you do fix them, and you still review the fix diff against the
|
|
275
|
+
rubric before closing the round.
|
|
276
|
+
|
|
277
|
+
**The breaker.** If round 5's re-review still leaves findings open, stop
|
|
278
|
+
dispatching and adjudicate each one yourself:
|
|
279
|
+
|
|
280
|
+
- **Reviewer wrong or the point contestable** → park it:
|
|
281
|
+
`Task <N>: parked — <finding> — ruling: <why the code stands>`.
|
|
282
|
+
- **Real, but nothing downstream builds on it** → park it the same way, with a
|
|
283
|
+
ruling saying it's real and deferred.
|
|
284
|
+
- **Real and load-bearing** (a later task builds on it, or it exposes a plan defect)
|
|
285
|
+
→ **STOP**. Append `Task <N>: BLOCKED — <reason>` and report to the operator with
|
|
286
|
+
the finding, the plan text it collides with, and the fix history. Parking a
|
|
287
|
+
structural failure lets every dependent task build on it.
|
|
288
|
+
|
|
289
|
+
Adjudicate **only at the cap**. Adjudicating earlier to end a loop is pre-judging
|
|
290
|
+
with a nicer name. Every adjudication is a ledger line; silent discards are
|
|
291
|
+
forbidden.
|
|
292
|
+
|
|
293
|
+
### 4.6 Complete the task
|
|
294
|
+
|
|
295
|
+
When the review is clean — or every open finding is parked with a ruling at the cap
|
|
296
|
+
— append:
|
|
297
|
+
|
|
298
|
+
- `Task <N>: complete (commits <base7>..<head7>, review clean)`, or
|
|
299
|
+
- `Task <N>: complete (commits <base7>..<head7>, <K> parked)`
|
|
300
|
+
|
|
301
|
+
Mark the todo complete, move on. Never start the next task while Critical/Important
|
|
302
|
+
findings are neither fixed nor parked-with-ruling at the cap.
|
|
303
|
+
|
|
304
|
+
## 5. Final whole-branch review
|
|
305
|
+
|
|
306
|
+
After the last task: build a package over `MERGE_BASE`..`HEAD`
|
|
307
|
+
(`git merge-base main HEAD`), dispatch the whole-branch review
|
|
308
|
+
([`review.md`](review.md) → *Final review*; on the run's model, escalation offered
|
|
309
|
+
out loud per *Models* above), and point it at the
|
|
310
|
+
ledger's deferred-minor and parked lines so it can triage what must be fixed before
|
|
311
|
+
merge.
|
|
312
|
+
|
|
313
|
+
If it returns findings, dispatch **ONE** fix subagent with the complete list — not
|
|
314
|
+
one fixer per finding; per-finding fixers each rebuild context and re-run suites.
|
|
315
|
+
Then exactly **one** scoped re-review of the fix wave. Adjudicate residuals as in
|
|
316
|
+
the breaker: park with rulings, or stop on load-bearing ones. There is no second
|
|
317
|
+
fix wave.
|
|
318
|
+
|
|
319
|
+
## 6. Integrate, then finish
|
|
320
|
+
|
|
321
|
+
The work is in a worktree on its own branch; stages 7–9 lint, deploy and document
|
|
322
|
+
the **integrated** result. Close that gap here, honoring the branch policy recorded
|
|
323
|
+
in the stage-0 brief:
|
|
324
|
+
|
|
325
|
+
1. **Sync with the base branch** (rebase or merge, whichever the project uses) and
|
|
326
|
+
re-run the full suite on the result. A branch that was green in isolation and red
|
|
327
|
+
after integration is red — fix it here, not at stage 7.
|
|
328
|
+
2. **Land it the project's way:** merge into the base branch, or open a PR when the
|
|
329
|
+
project requires review. Opening a PR is outward-facing — do it only with the
|
|
330
|
+
operator's go or the brief's specific standing authorization.
|
|
331
|
+
3. **Never force-push a shared branch**, and never land on `main` when the brief put
|
|
332
|
+
it off-limits.
|
|
333
|
+
4. **Remove the worktree** once merged (`git worktree remove <path>`, or the native
|
|
334
|
+
tool that created it), and delete this plan's workspace
|
|
335
|
+
(`rm -rf .task-pipeline/build/<plan-basename>`) — git history is the record now.
|
|
336
|
+
Sibling directories belong to other plans; leave them.
|
|
337
|
+
|
|
338
|
+
If the operator's policy is "leave the branch, I'll merge it myself", stop after
|
|
339
|
+
step 1, say exactly where the branch is and what state it's in, and record that
|
|
340
|
+
stages 7–9 run against an unintegrated branch.
|
|
341
|
+
|
|
342
|
+
## GATE (auto)
|
|
343
|
+
|
|
344
|
+
All plan tasks DONE with all three review verdicts (spec compliance, REQ satisfied,
|
|
345
|
+
code quality); the full test suite green; every open finding either fixed or parked
|
|
346
|
+
with a ruling; **every parked finding and implementer concern harvested into the
|
|
347
|
+
carry-over ledger** — the workspace is deleted, so nothing may stay only there;
|
|
348
|
+
no task left BLOCKED; the branch integrated per the brief's policy — or the
|
|
349
|
+
operator explicitly told you to leave it, and that is recorded. Verify it yourself;
|
|
350
|
+
a red suite or an unresolved BLOCKED does not advance to stage 6.
|
|
351
|
+
|
|
352
|
+
## Rationalizations
|
|
353
|
+
|
|
354
|
+
| Excuse | Reality |
|
|
355
|
+
|---|---|
|
|
356
|
+
| "Close enough on spec compliance" | The reviewer found spec gaps ⇒ not done. Fix, or hit the cap and adjudicate. Those are the only exits. |
|
|
357
|
+
| "I'll fix it myself, dispatching is overhead" | In a subagent run, controller fixes skip review and pollute your context — resume the implementer. (Inline runs are the declared exception, and still review the fix diff.) |
|
|
358
|
+
| "One more round will converge" | Past the cap, rounds don't converge — the failure is structural. Adjudicate and route. |
|
|
359
|
+
| "The fix was small, skip the re-review" | Unreviewed fixes are how regressions land. Every round ends with a scoped re-review. |
|
|
360
|
+
| "This finding is obviously wrong, drop it" | You adjudicate at the cap, in writing. Silent discards are forbidden. |
|
|
361
|
+
| "Ledger bookkeeping is overhead" | The ledger is what survives compaction. Without one, controllers re-run entire completed task sequences. |
|
|
362
|
+
| "Two implementers in parallel will be faster" | One working tree, two writers = corrupted state. Parallel needs one worktree each. |
|
|
363
|
+
| "I'll paste the earlier tasks so it has context" | A fresh subagent needs its task, its interfaces and the constraints. Pasted history is pure cost. |
|
|
364
|
+
| "Stage 7 can merge the branch" | Stage 7 lints and deploys what is integrated. An unmerged branch means lint, deploy and docs all ran against something that is not what ships. |
|