task-pipeline-skill 0.10.0 → 0.17.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +401 -0
- package/LICENSE +85 -0
- package/README.md +211 -84
- package/cursor/rules/task-pipeline.mdc +135 -20
- package/package.json +3 -3
- package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
- package/plugins/task-pipeline/commands/task-pipeline.md +16 -8
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +139 -57
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +41 -27
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +32 -6
- package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +73 -21
- package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +169 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +55 -22
- package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +184 -48
- package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +11 -3
- package/plugins/task-pipeline/skills/task-pipeline/templates/adr.md +64 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +49 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/context.md +87 -0
|
@@ -1,46 +1,98 @@
|
|
|
1
|
-
#
|
|
1
|
+
# Companions — what's built in, what's optional, what to install
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
3
|
+
**The pipeline's doctrine is entirely built into this skill.** Stages 0, 2, 3, 4, 5,
|
|
4
|
+
6 and 10 run from `references/*.md` — no companion plugin, no resolution step, no
|
|
5
|
+
fallback path, no version skew, and no failure mode where a stage can't run because
|
|
6
|
+
something isn't installed.
|
|
7
|
+
|
|
8
|
+
What remains is a short list of **optional** companions that make individual stages
|
|
9
|
+
better, plus one that is required only for user-facing work.
|
|
10
|
+
|
|
11
|
+
## Built in — nothing to install
|
|
12
|
+
|
|
13
|
+
| Stage | Doctrine |
|
|
14
|
+
|---|---|
|
|
15
|
+
| 0 Intake grill | `references/grill.md` |
|
|
16
|
+
| 2 Brainstorm | `references/brainstorm.md` |
|
|
17
|
+
| 2 Decompose (platforms only) | `references/decomposition.md` |
|
|
18
|
+
| 3 Spec | `references/spec.md` |
|
|
19
|
+
| 4 Plan | `references/planning.md` |
|
|
20
|
+
| 5 Build (isolation, subagents, fix loop) | `references/build.md` + `references/review.md` |
|
|
21
|
+
| 5–6 TDD + suite gate | `references/tdd.md` |
|
|
22
|
+
| 10 Acceptance (REQ close-out) | `references/acceptance.md` |
|
|
23
|
+
| any repeating loop | `references/loop-guard.md` |
|
|
7
24
|
|
|
8
25
|
## The matrix
|
|
9
26
|
|
|
10
27
|
| Skill / tool | Needed for | Required? | Install |
|
|
11
28
|
|---|---|---|---|
|
|
12
|
-
| **superpowers** (`brainstorming`, `writing-plans`, `subagent-driven-development`, `using-git-worktrees`, `test-driven-development`) | stages 2, 4, 5, 6 | **Required** (always) | `/plugin marketplace add obra/superpowers` → `/plugin install superpowers@superpowers` |
|
|
13
29
|
| **super-ux** (`ux-foundation`, `ux-flows`, `ux-scenarios`, `ux-audit`, `/ux`, `/ux-lint`) | stage 3 UX track | **Required for any user-facing task** | `/plugin marketplace add ssheleg/super-ux` → `/plugin install super-ux@super-ux` (or `npx skills add ssheleg/super-ux`) |
|
|
14
|
-
| **grill-me** / **grilling** | stage 0 intake grill | Optional (built-in grill loop is the fallback) | `npx skills add mattpocock/skills`, or the engineering-advanced-skills marketplace |
|
|
15
30
|
| **context7** (MCP) | stage 1 docs study | Recommended (web-search fallback) | connect the context7 MCP server |
|
|
16
31
|
| **wiki-update** | stage 9 wiki sync | Optional (skip wiki if absent) | user's wiki skill set |
|
|
32
|
+
| ~~superpowers~~ | — | **Not a dependency.** Stages 2/4/5/6 run on the built-in doctrine above. See *Optional bridge* | — |
|
|
33
|
+
| ~~grill-me / grilling~~ | — | **Not a dependency.** The stage-0 grill is built in (`references/grill.md`) | — |
|
|
34
|
+
|
|
35
|
+
## Optional bridge — substituting an external skill set
|
|
36
|
+
|
|
37
|
+
An operator who already runs an equivalent skill set may map it onto stages 2/4/5/6
|
|
38
|
+
in their `pipeline.json` → `skills[]`, e.g. `superpowers:brainstorming`,
|
|
39
|
+
`superpowers:writing-plans`, `superpowers:using-git-worktrees`,
|
|
40
|
+
`superpowers:subagent-driven-development`, `superpowers:test-driven-development`.
|
|
41
|
+
|
|
42
|
+
Rules for that bridge:
|
|
43
|
+
|
|
44
|
+
- **It is a substitution, never a requirement.** Nothing detects it, nothing
|
|
45
|
+
recommends it, nothing waits for it, and its absence is never an error.
|
|
46
|
+
- **The gates still govern.** Whatever runs a stage, `stages.md` decides when the
|
|
47
|
+
stage is done.
|
|
48
|
+
- **Never mix providers inside one stage** — either the built-in doctrine runs it or
|
|
49
|
+
the substitute does; interleaving two review loops produces neither.
|
|
17
50
|
|
|
18
|
-
## Preflight
|
|
51
|
+
## Preflight (emit before stage 0)
|
|
19
52
|
|
|
20
|
-
|
|
21
|
-
|
|
53
|
+
Detect the optional companions and print ONE block — companions **plus the model
|
|
54
|
+
decision** (`model-tiering.md`), so the operator arms the whole run in a single
|
|
55
|
+
exchange:
|
|
22
56
|
|
|
23
57
|
```
|
|
24
|
-
Pipeline companions:
|
|
25
|
-
|
|
26
|
-
✗ super-ux — this task looks user-facing; recommended. Install:
|
|
58
|
+
Pipeline companions (stage doctrine is built in — nothing to install for it):
|
|
59
|
+
✗ super-ux — this task looks user-facing; required for the UX track:
|
|
27
60
|
/plugin marketplace add ssheleg/super-ux
|
|
28
61
|
/plugin install super-ux@super-ux
|
|
29
62
|
✓ context7 — ready
|
|
30
|
-
✗ grill-me — optional; falling back to the built-in grill loop
|
|
31
63
|
✓ wiki-update — ready
|
|
32
|
-
|
|
64
|
+
|
|
65
|
+
🧠 Model for this run: recommended <top tier available>. You're on <current>.
|
|
66
|
+
/model <id> to switch, or "keep current", or name per-stage overrides.
|
|
67
|
+
|
|
68
|
+
Install the ✗ items you want, answer the model line, then say "continue".
|
|
33
69
|
```
|
|
34
70
|
|
|
35
71
|
Rules:
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
72
|
+
|
|
73
|
+
- Only flag **super-ux** when the task implies a UI (the stage-0 grill decides;
|
|
74
|
+
when unsure, flag it — a false positive costs one install).
|
|
75
|
+
- **Never gate any stage on an install** except the stage-3 UX track on a UI task.
|
|
39
76
|
- Optional tools missing → state the fallback, don't block.
|
|
40
77
|
- Re-detect after the operator installs; don't assume.
|
|
78
|
+
- The model answer goes into the brief. Don't ask again per stage
|
|
79
|
+
(`model-tiering.md` → *Mechanic*).
|
|
80
|
+
|
|
81
|
+
## Credit
|
|
82
|
+
|
|
83
|
+
The built-in doctrine is **ported, not depended on**:
|
|
84
|
+
|
|
85
|
+
- The stage-0 grill is adapted from Matt Pocock's `grilling` / `grill-with-docs`
|
|
86
|
+
skills (MIT, https://github.com/mattpocock/skills).
|
|
87
|
+
- Stages 2–6 are adapted from the `brainstorming`, `writing-plans`,
|
|
88
|
+
`using-git-worktrees`, `subagent-driven-development`, `test-driven-development`
|
|
89
|
+
and `requesting-code-review` skills in obra/superpowers (MIT,
|
|
90
|
+
https://github.com/obra/superpowers).
|
|
91
|
+
|
|
92
|
+
Both notices live in the repo `LICENSE` → *Third-party*.
|
|
41
93
|
|
|
42
94
|
## Hand-off the other direction
|
|
43
95
|
|
|
44
|
-
super-ux's `/ux` menu can hand off *to* this pipeline (its "execute
|
|
45
|
-
|
|
46
|
-
|
|
96
|
+
super-ux's `/ux` menu can hand off *to* this pipeline (its "execute autonomously"
|
|
97
|
+
action). When entered that way the UX chain already exists — see `stages.md` → 0
|
|
98
|
+
*Entry-from-super-ux short-circuit*: verify, don't rebuild.
|
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
# Host conventions (stages 6–
|
|
1
|
+
# Host conventions (stages 6–10)
|
|
2
2
|
|
|
3
3
|
The orchestrator is project-agnostic. For tests / lint / deploy / docs / wiki it reads the
|
|
4
4
|
**host project's `CLAUDE.md` / `AGENTS.md` first**, then falls back to detection.
|
|
@@ -33,3 +33,13 @@ found, surface it and **ask** rather than guessing.
|
|
|
33
33
|
- Host self-update rules (module docs, runbooks, agent-self cards, etc.) — update
|
|
34
34
|
in the same change. Wiki: the `wiki-update` skill (resolves the vault via
|
|
35
35
|
`~/.obsidian-wiki/config`). Fix dangling links.
|
|
36
|
+
|
|
37
|
+
## Issue tracker (stage 10)
|
|
38
|
+
|
|
39
|
+
Acceptance parks what wasn't delivered: every `deferred` REQ and every unresolved
|
|
40
|
+
carry-over row needs a **home** — an issue, a backlog entry, a ticket id. Read the
|
|
41
|
+
host's convention (`CLAUDE.md` usually names the tracker and the id format; else
|
|
42
|
+
detect: a `.github/ISSUE_TEMPLATE/`, a Linear/Jira reference in recent commits, a
|
|
43
|
+
`TODO.md`). Never invent a tracker, and never close a run on "we'll remember it" —
|
|
44
|
+
if no tracker exists, write the row into the repo's backlog file and say where it
|
|
45
|
+
went.
|
|
@@ -0,0 +1,139 @@
|
|
|
1
|
+
# Decomposition — cutting a platform into bricks
|
|
2
|
+
|
|
3
|
+
A one-feature task goes through the pipeline once. A **platform** — anything whose
|
|
4
|
+
brief describes more than one deliverable, more than one surface, or a system
|
|
5
|
+
rather than a change — must be cut into modules first, and then built one brick at
|
|
6
|
+
a time, each brick carrying its own documentation, spec, plan, build and gates.
|
|
7
|
+
|
|
8
|
+
This runs at the end of **stage 2**, on the approved design, before any spec is
|
|
9
|
+
written. It is skipped — explicitly, in writing — when the work is a single module.
|
|
10
|
+
|
|
11
|
+
## When it applies
|
|
12
|
+
|
|
13
|
+
Decompose when any of these is true:
|
|
14
|
+
|
|
15
|
+
- the brief names several independent capabilities ("accounts, billing, reporting");
|
|
16
|
+
- the work spans several surfaces (API + web + worker) that could ship separately;
|
|
17
|
+
- the REQ table has requirements that no single deliverable satisfies together;
|
|
18
|
+
- the design's units have their own data and could plausibly be owned by different
|
|
19
|
+
people.
|
|
20
|
+
|
|
21
|
+
Otherwise record one line in the design — `single module: <name>` — and go to
|
|
22
|
+
stage 3. A skipped decomposition is a decision, never an omission.
|
|
23
|
+
|
|
24
|
+
## How to cut
|
|
25
|
+
|
|
26
|
+
**By capability, not by layer.** "Ordering", "Billing", "Notifications" are
|
|
27
|
+
modules. "Controllers", "Services", "Database" are not: a layer cut forces every
|
|
28
|
+
feature to touch every module, which is the opposite of a brick.
|
|
29
|
+
|
|
30
|
+
A module is a **brick** when all of these hold:
|
|
31
|
+
|
|
32
|
+
1. **Independently specifiable** — you can write its dossier without deciding
|
|
33
|
+
another module's internals.
|
|
34
|
+
2. **Independently buildable and testable** — its tests pass without another
|
|
35
|
+
module's implementation present (stubs at the contract are fine).
|
|
36
|
+
3. **Owns its data** — the entities it is the source of truth for belong to it, and
|
|
37
|
+
nothing else writes them.
|
|
38
|
+
4. **Talks through declared contracts only** — every cross-module interaction is a
|
|
39
|
+
named API, event or schema, listed in both modules' dossiers.
|
|
40
|
+
5. **Deliverable on its own** — landing it leaves the system working, even if the
|
|
41
|
+
capability is not yet reachable by users.
|
|
42
|
+
|
|
43
|
+
If a candidate fails (2) or (3), the cut is in the wrong place: either merge it
|
|
44
|
+
into its neighbor or move the disputed data to the module that truly owns it.
|
|
45
|
+
|
|
46
|
+
**Order the bricks:**
|
|
47
|
+
|
|
48
|
+
- **The walking skeleton first.** The first module is the thinnest end-to-end slice
|
|
49
|
+
that proves the architecture — one real path through the system, however small.
|
|
50
|
+
Building three "foundation" modules before anything runs end-to-end hides
|
|
51
|
+
integration risk until the worst possible moment.
|
|
52
|
+
- Then topological order: nothing is built before what it depends on.
|
|
53
|
+
- **No cycles.** A cycle means the cut is wrong. Break it by moving the shared
|
|
54
|
+
concept into its own module, or by turning one direction of the dependency into
|
|
55
|
+
an event the other module subscribes to. Record which you chose and why.
|
|
56
|
+
|
|
57
|
+
## The module map — the artifact
|
|
58
|
+
|
|
59
|
+
Write `docs/superpowers/specs/YYYY-MM-DD-<topic>-modules.md` and commit it. It is
|
|
60
|
+
the program's spine: every later run reads it, and its status column is how a
|
|
61
|
+
resumed session knows where the program stopped.
|
|
62
|
+
|
|
63
|
+
```markdown
|
|
64
|
+
# Module map — <platform>
|
|
65
|
+
|
|
66
|
+
Build order is top to bottom. Status: `planned` → `in progress` → `done` |
|
|
67
|
+
`deferred`. One row per module, no exceptions.
|
|
68
|
+
|
|
69
|
+
| # | Module | Delivers | Owns (entities) | Depends on | Contracts exposed | UI? | REQs | Status |
|
|
70
|
+
|---|---|---|---|---|---|---|---|---|
|
|
71
|
+
| 1 | ordering | place and track an order | Order, OrderLine | — | `POST /orders`, `OrderPlaced` event | yes | REQ-001, REQ-004 | planned |
|
|
72
|
+
| 2 | billing | charge for a placed order | Invoice, Payment | ordering | `InvoiceIssued` event | no | REQ-002 | planned |
|
|
73
|
+
|
|
74
|
+
## Cut rationale
|
|
75
|
+
|
|
76
|
+
<why these seams and not others; what was merged or split, and what a cycle forced>
|
|
77
|
+
|
|
78
|
+
## Cross-module contracts
|
|
79
|
+
|
|
80
|
+
<one block per contract: owner module, consumer(s), exact shape (schema or
|
|
81
|
+
signature), and the failure behavior when the other side is unavailable>
|
|
82
|
+
|
|
83
|
+
## Deferred to later modules
|
|
84
|
+
|
|
85
|
+
<capabilities deliberately postponed, with the module that will carry them>
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
Every REQ from the brief appears in exactly one module's `REQs` cell. A REQ that
|
|
89
|
+
fits nowhere means the map is incomplete; a REQ in two modules means the seam runs
|
|
90
|
+
through a requirement — re-cut or split the REQ.
|
|
91
|
+
|
|
92
|
+
## GATE (part of stage 2, manual)
|
|
93
|
+
|
|
94
|
+
Together with the design approval:
|
|
95
|
+
|
|
96
|
+
1. Every module satisfies the brick criteria, or its exception is written down.
|
|
97
|
+
2. The dependency graph is acyclic and the build order is topological.
|
|
98
|
+
3. The first module is a walking skeleton, or the reason it isn't is recorded.
|
|
99
|
+
4. Every REQ maps to exactly one module.
|
|
100
|
+
5. Cross-module contracts are named (shape can be locked later, in each module's
|
|
101
|
+
spec — but the *existence* and *owner* of each contract is decided here).
|
|
102
|
+
6. The operator approves the map and the order.
|
|
103
|
+
|
|
104
|
+
## The program loop — one brick at a time
|
|
105
|
+
|
|
106
|
+
After the map is approved, the pipeline runs **per module**, in build order:
|
|
107
|
+
|
|
108
|
+
```
|
|
109
|
+
module N → stage 3 (dossier/spec) → 4 plan → 5 build → 6 tests
|
|
110
|
+
→ 7 lint + deploy → 8 post-deploy → 9 docs + wiki → 10 acceptance
|
|
111
|
+
→ mark module done → module N+1 (back to stage 3)
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
Rules for the loop:
|
|
115
|
+
|
|
116
|
+
- **Stages 0–2 run once for the platform.** Modules do not re-grill and do not
|
|
117
|
+
re-decompose. New information that changes the map goes back to stage 2
|
|
118
|
+
deliberately, as a map revision with the operator's approval — not as a quiet
|
|
119
|
+
edit mid-module.
|
|
120
|
+
- **Each module's spec is a full dossier** ([`spec.md`](spec.md)): architecture,
|
|
121
|
+
entities, contracts in and out, business rules, edge and failure cases, UI/Figma
|
|
122
|
+
chain when it has a surface.
|
|
123
|
+
- **The contract is the boundary.** A module may stub what a later module will
|
|
124
|
+
provide, but it may not reach into another module's internals; if it needs to,
|
|
125
|
+
the seam is wrong — back to the map.
|
|
126
|
+
- **Deploy cadence is the brief's call** (autonomy sweep): deploy each module as it
|
|
127
|
+
lands, or build several and deploy once. Record it; don't decide it per module.
|
|
128
|
+
- **Update the map's status column as each module closes**, in the same commit as
|
|
129
|
+
that module's acceptance. The map is the resume point after a lost context.
|
|
130
|
+
- **Loop discipline:** a module re-entering the same stage a third time trips the
|
|
131
|
+
loop guard ([`loop-guard.md`](loop-guard.md)) — stop, name the oscillation, and
|
|
132
|
+
fix the layer that owns it instead of iterating.
|
|
133
|
+
|
|
134
|
+
## Program done
|
|
135
|
+
|
|
136
|
+
The program is finished when every row is `done` or `deferred` with an agreed home,
|
|
137
|
+
the cross-module contracts are exercised by tests that cross the seam (not just
|
|
138
|
+
per-module unit tests), and the final acceptance covers the platform's REQ table as
|
|
139
|
+
a whole — not module by module.
|
|
@@ -0,0 +1,169 @@
|
|
|
1
|
+
# The grill — stage 0, built in
|
|
2
|
+
|
|
3
|
+
The intake grill is **part of this skill**. No companion skill to install, no
|
|
4
|
+
provider to resolve, nothing to fall back to: this file *is* the implementation.
|
|
5
|
+
|
|
6
|
+
Its job is not to design. It is to take a one-line request ("make me feature X")
|
|
7
|
+
and interview it into a brief complete enough that stages 1→10 finish without
|
|
8
|
+
coming back to the operator.
|
|
9
|
+
|
|
10
|
+
> Adapted, with thanks, from Matt Pocock's `grilling` / `grill-with-docs` skills
|
|
11
|
+
> (MIT — see this repo's `LICENSE` → *Third-party*). The domain-awareness
|
|
12
|
+
> half — glossary challenges, `CONTEXT.md`, ADR discipline — comes from there; the
|
|
13
|
+
> autonomy sweep and the brief are this pipeline's.
|
|
14
|
+
|
|
15
|
+
## The loop
|
|
16
|
+
|
|
17
|
+
Interview the operator relentlessly about every aspect of the task until you reach
|
|
18
|
+
a **shared understanding**. Walk down each branch of the decision tree, resolving
|
|
19
|
+
dependencies between decisions one by one.
|
|
20
|
+
|
|
21
|
+
1. **One question per turn.** Never bundle. Wait for the answer before the next.
|
|
22
|
+
2. **Recommend an answer with every question** (+ a one-line rationale). "What do
|
|
23
|
+
you think?" is lazy — you have the codebase in front of you, they don't.
|
|
24
|
+
3. **If the codebase can answer it, go read the codebase.** Spending the
|
|
25
|
+
operator's turn on something `grep`/`Read`/context7 would have told you is the
|
|
26
|
+
most common way to waste a grill.
|
|
27
|
+
4. **Depth-first.** Finish a branch before opening another; ask prerequisite
|
|
28
|
+
decisions first, so later answers don't invalidate earlier ones.
|
|
29
|
+
5. **Reconcile contradictions immediately**, and chase dodges: "we'll decide
|
|
30
|
+
later" → "what's the latest you can decide and still ship?"
|
|
31
|
+
6. **Cover the autonomy sweep** (below). An unasked question is not neutral — it
|
|
32
|
+
is a scheduled interruption at stage 6.
|
|
33
|
+
|
|
34
|
+
**Stop** when a re-scan surfaces no new branches. Don't grill past diminishing
|
|
35
|
+
returns: genuinely reversible calls can be deferred with a note.
|
|
36
|
+
|
|
37
|
+
## Domain awareness
|
|
38
|
+
|
|
39
|
+
While exploring the codebase, also look for what the project already says about
|
|
40
|
+
itself — and hold the operator to it.
|
|
41
|
+
|
|
42
|
+
### Find the existing docs
|
|
43
|
+
|
|
44
|
+
Most repos have a single context:
|
|
45
|
+
|
|
46
|
+
```
|
|
47
|
+
/
|
|
48
|
+
├── CONTEXT.md
|
|
49
|
+
├── docs/adr/
|
|
50
|
+
│ ├── 0001-event-sourced-orders.md
|
|
51
|
+
│ └── 0002-postgres-for-write-model.md
|
|
52
|
+
└── src/
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
A `CONTEXT-MAP.md` at the root means multiple contexts, and points at where each
|
|
56
|
+
one lives (`src/ordering/CONTEXT.md`, `src/billing/CONTEXT.md`, …), each with its
|
|
57
|
+
own `docs/adr/` alongside the system-wide one. Infer which context the task
|
|
58
|
+
belongs to; if it's genuinely unclear, ask.
|
|
59
|
+
|
|
60
|
+
Create these files **lazily** — only when you have something real to write.
|
|
61
|
+
|
|
62
|
+
### Techniques during the session
|
|
63
|
+
|
|
64
|
+
- **Challenge against the glossary.** When a term conflicts with `CONTEXT.md`, say
|
|
65
|
+
so on the spot: *"Your glossary defines 'cancellation' as X, but you seem to mean
|
|
66
|
+
Y — which is it?"*
|
|
67
|
+
- **Sharpen fuzzy language.** Vague or overloaded terms get a proposed canonical
|
|
68
|
+
one: *"You're saying 'account' — do you mean the Customer or the User? Those are
|
|
69
|
+
different things."*
|
|
70
|
+
- **Stress-test with concrete scenarios.** Invent specific cases that probe edge
|
|
71
|
+
conditions and force precision about the boundaries between concepts.
|
|
72
|
+
- **Cross-reference with the code.** When the operator states how something works,
|
|
73
|
+
check whether the code agrees, and surface contradictions: *"Your code cancels
|
|
74
|
+
entire Orders, but you just said partial cancellation is possible — which is
|
|
75
|
+
right?"*
|
|
76
|
+
- **Update `CONTEXT.md` inline.** Resolve a term → write it down right then, not in
|
|
77
|
+
a batch at the end. Format: [`templates/context.md`](../templates/context.md).
|
|
78
|
+
Keep it free of implementation detail — only terms a domain expert would
|
|
79
|
+
recognize.
|
|
80
|
+
|
|
81
|
+
### Offer an ADR sparingly
|
|
82
|
+
|
|
83
|
+
Only when **all three** are true:
|
|
84
|
+
|
|
85
|
+
1. **Hard to reverse** — changing your mind later carries real cost.
|
|
86
|
+
2. **Surprising without context** — a future reader will ask "why on earth this
|
|
87
|
+
way?"
|
|
88
|
+
3. **A real trade-off** — genuine alternatives existed and one was chosen for
|
|
89
|
+
specific reasons.
|
|
90
|
+
|
|
91
|
+
Any one missing → skip it. Format and what qualifies:
|
|
92
|
+
[`templates/adr.md`](../templates/adr.md). ADRs land in `docs/adr/` with sequential
|
|
93
|
+
numbering (scan for the highest number, increment).
|
|
94
|
+
|
|
95
|
+
## The autonomy sweep
|
|
96
|
+
|
|
97
|
+
Resolving the *task* is not enough. The grill must also pre-resolve everything that
|
|
98
|
+
would otherwise stop stages 1→10 mid-flight. Every row gets an answer **or** an
|
|
99
|
+
explicit "stop and ask me here":
|
|
100
|
+
|
|
101
|
+
| Stage | What to settle up front |
|
|
102
|
+
|---|---|
|
|
103
|
+
| run-wide | the model decision ([`model-tiering.md`](model-tiering.md)); what to decide autonomously vs escalate |
|
|
104
|
+
| 1 Docs | external libs/APIs/SDKs in play; any private ones context7 can't resolve → where their docs live |
|
|
105
|
+
| 2 Decompose | is this a platform (several capabilities/surfaces) or one module? if platform: deploy cadence — per module or once at the end |
|
|
106
|
+
| 2–3 Spec | UI verdict (arms super-ux); any scenario-tracing waiver |
|
|
107
|
+
| 4–5 Dev | base branch; worktree/branch policy; is `main` off-limits; commit convention; task tracker |
|
|
108
|
+
| 5 Integration | how the branch lands (merge / PR + approver / "leave it unmerged"); parallel fan-out wanted (one worktree per implementer)? |
|
|
109
|
+
| 6 Tests | the test command; what "green" means here; known-red baseline; coverage expectation |
|
|
110
|
+
| 7 Lint+deploy | lint command; deploy target and path; release automation on/off; deploy-from-main rule; **deploy authorization** |
|
|
111
|
+
| 8 Post-deploy | where logs / health live (app name, endpoint, workflow) |
|
|
112
|
+
| 9 Docs+wiki | which module docs / runbooks this change updates; wiki sync yes/no |
|
|
113
|
+
| 10 Acceptance | who signs off; where deferred REQs are tracked (issue tracker, backlog) |
|
|
114
|
+
|
|
115
|
+
**Deploy authorization has a hard floor.** Deploy and publish are outward and
|
|
116
|
+
irreversible, so a vague "just do everything" authorizes nothing. A standing
|
|
117
|
+
authorization counts only when it is **specific** — named target, named
|
|
118
|
+
preconditions ("staging once lint and the full suite are green; production always
|
|
119
|
+
asks"). Specific and recorded → it satisfies the stage-7 manual gate. Broader,
|
|
120
|
+
absent or ambiguous → stage 7 stops and asks.
|
|
121
|
+
|
|
122
|
+
## The REQ spine — the grill's other hard output
|
|
123
|
+
|
|
124
|
+
Prose scope is not checkable. Before the brief is confirmed, the grill must turn
|
|
125
|
+
what was asked into an **addressable list of requirements**, because every later
|
|
126
|
+
stage traces to these IDs and stage 10 accounts for every one of them.
|
|
127
|
+
|
|
128
|
+
| ID | Requirement | How it's verified | Status |
|
|
129
|
+
|---|---|---|---|
|
|
130
|
+
| REQ-001 | … | test name / `file:line` / command + expected output / `SCN-…` | open |
|
|
131
|
+
|
|
132
|
+
Three rules that decide whether the spine is worth anything:
|
|
133
|
+
|
|
134
|
+
1. **One REQ = one independently verifiable deliverable.** Not one per sentence of
|
|
135
|
+
the request. A small task gets three rows, not thirty — an inflated table is
|
|
136
|
+
ignored, and an ignored table protects nothing.
|
|
137
|
+
2. **Every row names its check.** *A requirement you can't say how to verify is a
|
|
138
|
+
badly-stated requirement* — split or sharpen it here, during the grill. This is
|
|
139
|
+
the single defence against the failure mode where three vague REQs cover a large
|
|
140
|
+
task and acceptance goes green over half of it.
|
|
141
|
+
3. **Ask what "finished" means per row, not for the task overall.** "Export works"
|
|
142
|
+
hides five decisions; "exports the currently filtered rows as CSV, verified by
|
|
143
|
+
`test_export_respects_filters`" hides none.
|
|
144
|
+
|
|
145
|
+
**Then freeze it.** Adding a requirement mid-run is fine — append with its source.
|
|
146
|
+
**Removing or narrowing one requires the operator's explicit agreement**, recorded
|
|
147
|
+
in the carry-over ledger. Quietly restating the task in smaller terms is the
|
|
148
|
+
subtlest way to lose it: every gate downstream then passes honestly, on a task
|
|
149
|
+
that shrank without anyone deciding it should.
|
|
150
|
+
|
|
151
|
+
## Output
|
|
152
|
+
|
|
153
|
+
Everything resolved goes into the **task brief**, seeded from
|
|
154
|
+
[`templates/brief.md`](../templates/brief.md) and committed to
|
|
155
|
+
`docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` — scope, **the REQ table**,
|
|
156
|
+
users, UI verdict, constraints, locked decisions, the autonomy table,
|
|
157
|
+
done-criteria, open assumptions. Seed the template only when the file is absent;
|
|
158
|
+
never overwrite an existing brief.
|
|
159
|
+
|
|
160
|
+
Alongside it, seed the **carry-over ledger** from
|
|
161
|
+
[`templates/carryover.md`](../templates/carryover.md) at
|
|
162
|
+
`…-carryover.md` — append-only, written by every later stage, read in full by
|
|
163
|
+
stage 10. Anything deferred, dropped, or half-done from here on goes there the
|
|
164
|
+
moment it's said: **deferred out loud is forgotten.**
|
|
165
|
+
|
|
166
|
+
Plus, where the session produced them: an updated `CONTEXT.md` and any ADRs, each
|
|
167
|
+
written as the decision landed.
|
|
168
|
+
|
|
169
|
+
The operator confirms the brief. Only then does stage 1 begin.
|
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# Loop guard — breaking churn, cross-cutting
|
|
2
|
+
|
|
3
|
+
Any stage that can repeat can also **churn**: a later pass undoing what an earlier
|
|
4
|
+
pass in the same run already did, two shapes alternating, the same file rewritten
|
|
5
|
+
round after round with no new information. Churn looks like progress and consumes
|
|
6
|
+
a run.
|
|
7
|
+
|
|
8
|
+
This file is the detector and the break protocol. It binds every repeating loop in
|
|
9
|
+
the pipeline: the stage-5 fix loop, a stage re-entered after a failed gate, the
|
|
10
|
+
per-module program loop ([`decomposition.md`](decomposition.md)), and any
|
|
11
|
+
audit → fix → audit cycle.
|
|
12
|
+
|
|
13
|
+
## Bookkeeping — the thing that makes detection mechanical
|
|
14
|
+
|
|
15
|
+
You cannot detect churn from memory, especially after compaction. Every repeating
|
|
16
|
+
pass appends one line to the run's ledger (`.task-pipeline/build/<plan>/progress.md`
|
|
17
|
+
for stage 5; `.task-pipeline/run.md` for stage-level and program-level loops):
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
touch: <file> — pass <N> (<stage|round|module>) — reason: <finding id / gate item>
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
One line per file per pass. The reason must name **what forced the edit** — a
|
|
24
|
+
finding id, a failed gate item, an operator instruction. "Cleanup", "polish" and
|
|
25
|
+
"while I was there" are not reasons; they are churn with better manners.
|
|
26
|
+
|
|
27
|
+
## Detection — any one of these trips the guard
|
|
28
|
+
|
|
29
|
+
1. **Revert-oscillation.** An edit restores something an earlier pass in this run
|
|
30
|
+
deliberately removed, or re-removes what an earlier pass added. Shape A → B → A.
|
|
31
|
+
2. **Repeat touch without new information.** The same file is edited in two
|
|
32
|
+
consecutive passes and the second pass's `reason` is the same finding/gate item
|
|
33
|
+
as the first — the fix did not fix it, or the two passes disagree about what
|
|
34
|
+
"fixed" means.
|
|
35
|
+
3. **Finding resurrection.** A finding whose text (normalized) matches one already
|
|
36
|
+
marked ADDRESSED or parked-with-ruling in this run comes back.
|
|
37
|
+
4. **Gate ping-pong.** The same stage is re-entered for the third time on the same
|
|
38
|
+
artifact, or two adjacent stages hand work back and forth (spec ⇄ plan,
|
|
39
|
+
plan ⇄ build) more than twice.
|
|
40
|
+
5. **Cross-loop contradiction.** A pass in one loop edits a file that a *different*
|
|
41
|
+
loop (another task, another module) already closed in this run — two owners for
|
|
42
|
+
one file.
|
|
43
|
+
|
|
44
|
+
Caps that trip the guard by themselves: **5 fix rounds** per task
|
|
45
|
+
([`build.md`](build.md)), **2 re-entries** per stage per artifact, **3 passes** per
|
|
46
|
+
module in the program loop.
|
|
47
|
+
|
|
48
|
+
## The break protocol
|
|
49
|
+
|
|
50
|
+
When the guard trips, **stop editing immediately**. Do not dispatch another fix, do
|
|
51
|
+
not "just try one more thing". Then, in this order:
|
|
52
|
+
|
|
53
|
+
1. **Freeze and name it.** Write the oscillation down in the ledger and to the
|
|
54
|
+
operator: shape **A** vs shape **B**, one line each, plus who is asking for each
|
|
55
|
+
(a finding, the plan's text, the spec, a gate check, an operator instruction) and
|
|
56
|
+
the evidence for each — `file:line`, the failing command, the review verdict.
|
|
57
|
+
2. **Find the layer that owns the conflict.** Churn almost always means a decision
|
|
58
|
+
is being re-litigated at the wrong altitude:
|
|
59
|
+
- two findings disagree → the **review rubric** decides
|
|
60
|
+
([`review.md`](review.md)); if it genuinely doesn't, it's a spec question;
|
|
61
|
+
- a finding contradicts the plan → the **operator** decides which governs
|
|
62
|
+
(never dismiss the finding, never fix against the plan silently);
|
|
63
|
+
- the plan contradicts the spec → back to **stage 4** with the evidence;
|
|
64
|
+
- the spec is ambiguous or wrong → back to **stage 3**, and if the ambiguity was
|
|
65
|
+
an unresolved intake question, say so — that is a stage-0 miss worth recording;
|
|
66
|
+
- two modules claim the same file or entity → back to **decomposition**: the cut
|
|
67
|
+
is wrong.
|
|
68
|
+
**Never resolve a higher-layer conflict inside a lower loop.** Patching code to
|
|
69
|
+
satisfy two contradictory requirements is how a run burns its remaining budget.
|
|
70
|
+
3. **Re-plan the check.** Replace whatever ad-hoc verification was running with an
|
|
71
|
+
explicit ordered checklist: every disputed item, one line each, in dependency
|
|
72
|
+
order, with a single owner and a single verification command per item. Write it
|
|
73
|
+
to the ledger before touching anything.
|
|
74
|
+
4. **Go in order, one at a time.** Verify item 1 → if it fails, fix only item 1 →
|
|
75
|
+
re-verify only item 1 → commit → item 2. No parallel edits, no bundled fixes, no
|
|
76
|
+
opportunistic cleanup in the same commit. The point is that each change has one
|
|
77
|
+
reason and one proof.
|
|
78
|
+
5. **Re-check the whole list once** at the end, in the same order. If a later item
|
|
79
|
+
broke an earlier one, that pair is the real conflict — escalate it per step 2
|
|
80
|
+
instead of looping again.
|
|
81
|
+
6. **Record the ruling.** Ledger line: `loop-guard: <A vs B> — ruling: <what governs
|
|
82
|
+
and why> — items: <N> verified in order`. The final review reads it.
|
|
83
|
+
|
|
84
|
+
## When to stop and hand back
|
|
85
|
+
|
|
86
|
+
If step 2 lands on "the operator decides", or a cap is hit a second time after a
|
|
87
|
+
re-planned pass, **stop and report BLOCKED** with: the two shapes, the evidence, the
|
|
88
|
+
history of passes, and your recommendation. That is a complete, honest hand-back —
|
|
89
|
+
far cheaper than a third round of the same argument.
|
|
90
|
+
|
|
91
|
+
## Rationalizations
|
|
92
|
+
|
|
93
|
+
| Excuse | Reality |
|
|
94
|
+
|---|---|
|
|
95
|
+
| "One more pass and it converges" | Two passes with the same reason already proved it doesn't. The disagreement is above the code. |
|
|
96
|
+
| "I'll just revert to what worked" | That is the oscillation, not the exit. Name A and B first. |
|
|
97
|
+
| "The reviewer keeps changing its mind" | Different findings on the same lines mean the requirement is ambiguous. That's a spec question. |
|
|
98
|
+
| "Tidying while I'm in the file" | Untracked edits are what make churn invisible. One reason per change, in the ledger. |
|
|
99
|
+
| "Logging the loop is bureaucracy" | Detection needs a record; after compaction the ledger is the only memory that survives. |
|
|
100
|
+
| "It's faster than escalating" | A run that spends its budget re-deciding a spec question delivers nothing. Escalation costs one message. |
|
|
@@ -1,32 +1,65 @@
|
|
|
1
|
-
# Model
|
|
1
|
+
# Model policy
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
3
|
+
**One model, confirmed once, before the run starts.** Not a per-stage tier list,
|
|
4
|
+
not a hardcoded vendor id — a single decision the operator makes at preflight and
|
|
5
|
+
the pipeline then honors without nagging.
|
|
5
6
|
|
|
6
|
-
|
|
7
|
-
|---|---|---|
|
|
8
|
-
| 0–4 (intake grill, docs, brainstorm, spec, plan) | Fable 5 | `claude-fable-5` |
|
|
9
|
-
| 5–6 (subagent dev, tests) | Opus 5 | `claude-opus-5` |
|
|
10
|
-
| 7–9 (lint/deploy, logs, docs) | inherit current | — |
|
|
7
|
+
## The default
|
|
11
8
|
|
|
12
|
-
|
|
9
|
+
> **Use the most capable reasoning model the environment offers** — at the time of
|
|
10
|
+
> writing that is the **latest Opus generation**, but read that as *"the top tier
|
|
11
|
+
> of whatever you're on"*, not as a specific string.
|
|
13
12
|
|
|
14
|
-
|
|
13
|
+
Every stage runs on that model by default. The pipeline is a full delivery cycle:
|
|
14
|
+
the grill has to hear what the operator didn't say, the spec has to lock contracts
|
|
15
|
+
a zero-context implementer will follow, and the build has to hold a plan in its
|
|
16
|
+
head. Downgrading any of those to save tokens costs more in rework than it saves.
|
|
15
17
|
|
|
16
|
-
|
|
17
|
-
> Switch: `/model <id>` — then say "continue". *(Reminder only.)*
|
|
18
|
+
## Never hardcode a model id
|
|
18
19
|
|
|
19
|
-
|
|
20
|
+
Model ids go stale — generations ship, tiers get renamed, and the operator may not
|
|
21
|
+
even be on the same provider. So:
|
|
20
22
|
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
23
|
+
- **Resolve at runtime.** Look at what the environment actually offers (`/model`,
|
|
24
|
+
the harness's model list) and pick the top reasoning tier available there.
|
|
25
|
+
- **Treat any id in this repo as an example**, including in `pipeline.example.json`.
|
|
26
|
+
Stage configs use provider-agnostic tokens:
|
|
27
|
+
- `default` — the model confirmed for this run (the recommendation above)
|
|
28
|
+
- `inherit` — whatever the operator is currently on; no recommendation
|
|
29
|
+
- **Another provider is fine.** "Top tier available" is the contract. If the
|
|
30
|
+
environment has no Opus-class model, the best available one is the right answer —
|
|
31
|
+
say which one you settled on and keep going.
|
|
25
32
|
|
|
26
|
-
|
|
27
|
-
`Agent` / `Workflow` model override — no operator action needed for subagents.
|
|
33
|
+
## Mechanic — confirm at preflight, then stop asking
|
|
28
34
|
|
|
29
|
-
|
|
35
|
+
Once, as part of the preflight (before stage 0):
|
|
30
36
|
|
|
31
|
-
|
|
32
|
-
|
|
37
|
+
> 🧠 **Model for this run:** recommended **`<top tier available>`**. You're on
|
|
38
|
+
> `<current>`.
|
|
39
|
+
> Switch with `/model <id>`, or say "keep current" / name another. Per-stage
|
|
40
|
+
> overrides welcome (e.g. a cheaper model for mechanical stages) — say so now and
|
|
41
|
+
> I'll record the map.
|
|
42
|
+
|
|
43
|
+
Record the answer in the stage-0 brief (`Model` row of the autonomy sweep). After
|
|
44
|
+
that:
|
|
45
|
+
|
|
46
|
+
- **Do not re-prompt at every stage boundary.** The decision is made; nagging is
|
|
47
|
+
the thing this replaces.
|
|
48
|
+
- **Re-prompt only** when the operator recorded a *per-stage override map* and the
|
|
49
|
+
next stage's entry differs from the current model — then emit the same block
|
|
50
|
+
scoped to that stage.
|
|
51
|
+
- A skill runs inside the current context and **cannot change the main-loop
|
|
52
|
+
model**; only the operator can, via `/model` (or `/fast`). Preflight is
|
|
53
|
+
interactive anyway, so this costs one exchange.
|
|
54
|
+
|
|
55
|
+
## Subagents
|
|
56
|
+
|
|
57
|
+
Stage 5 spawns subagents; the orchestrator pins them to the **run's confirmed
|
|
58
|
+
model** via the `Agent` / `Workflow` model override. No operator action needed —
|
|
59
|
+
and no silent downgrade to a cheaper tier.
|
|
60
|
+
|
|
61
|
+
## Degradation
|
|
62
|
+
|
|
63
|
+
The recommendation is a **reminder, not a block**. If the recommended tier isn't
|
|
64
|
+
available, keep the current model, state plainly which one is in use, and run. The
|
|
65
|
+
pipeline never stalls on a model it can't get.
|