task-pipeline-skill 0.17.1 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (24) hide show
  1. package/CHANGELOG.md +244 -0
  2. package/README.md +314 -135
  3. package/cursor/rules/task-pipeline.mdc +91 -16
  4. package/package.json +7 -3
  5. package/plugins/task-pipeline/.claude-plugin/plugin.json +2 -2
  6. package/plugins/task-pipeline/commands/task-pipeline.md +16 -6
  7. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +39 -8
  8. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +9 -4
  9. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +40 -8
  10. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +23 -11
  11. package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +224 -0
  12. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +6 -4
  13. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +8 -1
  14. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +12 -2
  15. package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +17 -3
  16. package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +37 -4
  17. package/plugins/task-pipeline/skills/task-pipeline/references/knowledge-sources.md +159 -0
  18. package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +8 -0
  19. package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +5 -3
  20. package/plugins/task-pipeline/skills/task-pipeline/references/review.md +2 -1
  21. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +73 -11
  22. package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +5 -0
  23. package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +24 -1
  24. package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +23 -0
@@ -28,16 +28,42 @@ recommended tier isn't available, say which one you're using and continue.
28
28
 
29
29
  Never skipped, and nothing to install — the grill is part of this rule. No "the
30
30
  task was already clear" exemption, no starting stage 1 while the user thinks. A
31
- one-line task ("build feature X") is not enough to finish autonomously. Grill the
32
- user up front, then run the rest without mid-flight questions:
31
+ one-line task ("build feature X") is not enough to finish autonomously.
32
+
33
+ **Phase 1 — harvest the sources BEFORE the first question.** Find what the project
34
+ already knows about this task and read it: the code; `CLAUDE.md` / `AGENTS.md`;
35
+ `CONTEXT.md` (or `CONTEXT-MAP.md`) and `docs/adr/`; `docs/` and `docs/ux/`; past
36
+ briefs/plans and their carry-over ledgers; **the knowledge wiki if one is
37
+ installed** — [obsidian-wiki](https://github.com/ar9av/obsidian-wiki), detect
38
+ `~/.obsidian-wiki/config` or a resolving `wiki-query`; and **any other repository
39
+ or hosted doc system the project names as its docs** (read-only, and never a source
40
+ you invented — it counts because the project names it). Query each by *this task's*
41
+ nouns; it is retrieval, not a full read; stop when the terms return nothing new.
42
+ Write a short **source ledger** into the brief — source, what it says about this
43
+ task, how fresh, and whether this run makes it stale. `none found` is a valid row.
44
+ If no wiki is installed, recommend it once and continue:
45
+ `pip install obsidian-wiki` → `obsidian-wiki setup --vault <path>`. It is never a
46
+ gate.
47
+
48
+ **Phase 2 — grill the user against that harvest**, then run the rest without
49
+ mid-flight questions:
33
50
  1. One question per turn — never bundle.
34
51
  2. Give a recommended answer with every question (+ one-line rationale).
35
52
  3. Explore the codebase before asking — if a search/read answers it, do that.
36
- 4. Walk the decision tree depth-first; ask prerequisite decisions first.
37
- 5. Reconcile contradictions; chase dodges ("decide later" → "latest you can decide
53
+ 4. **Validate every answer against the harvest.** When what the user says
54
+ contradicts a doc you read, quote the doc and ask which governs: *"the March ADR
55
+ says X, you just described Y — has it changed?"* The user **outranks every
56
+ document, but only out loud** — an override quoted against its source is a
57
+ recorded decision; an unquoted one is an undetected divergence that every later
58
+ gate will pass over. When two sources disagree: code > host docs/ADRs > wiki >
59
+ memory. Whichever side loses, if it's written down somewhere, log it for the
60
+ stage-9 doc update.
61
+ 5. Walk the decision tree depth-first; ask prerequisite decisions first.
62
+ 6. Reconcile contradictions; chase dodges ("decide later" → "latest you can decide
38
63
  and still ship?").
39
- 6. Run the **autonomy sweep** — resolve now whatever would stop stages 1→10 later:
40
- external libs and where their docs live; UI verdict; base branch, branch policy,
64
+ 7. Run the **autonomy sweep** — resolve now whatever would stop stages 1→10 later:
65
+ external libs and where their docs live; **which doc sources beyond this repo are
66
+ in play and whether stage 9 may write to them**; UI verdict; base branch, branch policy,
41
67
  commit convention, task tracker; the test command and what "green" means; the
42
68
  lint command; the deploy target, release toggle and **deploy authorization**;
43
69
  where logs/health live; which docs and runbooks this change updates. Each item
@@ -127,16 +153,39 @@ not authorize an outward, irreversible action — stage 7 stops and asks.
127
153
  the stage-0 brief.
128
154
  8. **Post-deploy** (auto) — tail logs / health-check; clean boot or an honest
129
155
  degradation report (never silent success).
130
- 9. **Docs + wiki** (auto) — update module docs/runbooks in the SAME change, and
131
- sync the project's knowledge base/wiki if it has one.
132
- 10. **Acceptance** (manual) — the closing stage: go back to the brief and account
133
- for **every** REQ. One row each, status `verified` / `partial` / `deferred` /
134
- `dropped`, and every `verified` carries **evidence** — a passing test name, a
135
- `file:line`, a command and its output. "Done" without evidence is downgraded to
136
- `partial`, never upgraded. Then ask out loud, list in hand: *here's what you
137
- asked for, here's what shipped, here's what's deferred and where it lives —
138
- what's missing?* Ask it even when the table is green. Gate: no REQ `unknown`,
139
- no ledger row without a home, user signs off.
156
+ 9. **Docs + wiki** (auto) — **the phase-1 source ledger is the work list**: every
157
+ source the harvest read gets updated if this run changed or disproved it. Module
158
+ docs and runbooks in the SAME change; the knowledge wiki via `wiki-update` when
159
+ [obsidian-wiki](https://github.com/ar9av/obsidian-wiki) is installed (absent →
160
+ recommend once, never block). Docs in **another repository** are outward:
161
+ propose the edit and get an explicit go, or carry it over with the exact change
162
+ written down. A doc that was worth reading at stage 0 and is wrong now is the
163
+ next run's false premise.
164
+ 10. **Acceptance** (manual) — the closing stage, in two halves.
165
+ **First the ladder walk**, because the REQ table only finds what was named and
166
+ lost: a comparison needs two sides and **an absence has one**. Walk each REQ
167
+ bottom-up through its rungs — recorded decision → spec section → contract *and
168
+ its failure behavior* → plan task with a satisfiable DoD → the change in the
169
+ tree → an **executed** named assertion → the surface a user reaches, and its
170
+ docs — checking the seam between each pair: does the decision reach the spec;
171
+ does the section say what happens when the contract fails; does every contract
172
+ have a task; did the DoD land in the diff; would that test still pass with the
173
+ production code deleted; can a user reach this and does a doc say so; and
174
+ finally, does what shipped satisfy the requirement's own *statement* rather
175
+ than the task's instructions. Order findings **by seam, not by file** — the
176
+ seam tells you which layer of your process leaks. Every absence becomes a new
177
+ REQ row with its check **before** the table is written; appending afterwards is
178
+ how acceptance goes green over a gap. Findings owned by a lower layer go back
179
+ there (spec → stage 3, plan → stage 4).
180
+ **Then the table:** one row per REQ, status `verified` / `partial` /
181
+ `deferred` / `dropped`, and every `verified` carries **evidence** — a passing
182
+ test name, a `file:line`, a command and its output. "Done" without evidence is
183
+ downgraded to `partial`, never upgraded, and **a green from a check nobody has
184
+ watched fail against a planted defect is not evidence at all**. Then ask out
185
+ loud, list in hand: *here's what you asked for, here's what shipped, here's
186
+ what's deferred and where it lives — what's missing?* Ask it even when the
187
+ table is green. Gate: ladder walk ran, no REQ `unknown`, no ledger row without
188
+ a home, user signs off.
140
189
 
141
190
  Cross-cutting: answer from the brief's autonomy section rather than re-asking, log every deferral in the ledger, never narrow the task silently, track
142
191
  tasks, conventional commits, honest degradation (never claim a failed/skipped step
@@ -161,6 +210,32 @@ no opportunistic edits. Re-check the list once in the same order at the end. If
161
210
  trips again after a re-planned pass, stop and hand back with both shapes, the
162
211
  evidence and your recommendation.
163
212
 
213
+ **Audit rules — for loops that *look* rather than edit.** A searching pass doesn't
214
+ oscillate, it **converges**: each pass edits the corpus the next pass reads, so the
215
+ newest edits are the least-reviewed text and are what the next pass finds. Measured
216
+ over seven passes on a real repository, by pass six the audit was mostly repairing
217
+ its own previous pass while the finding count still looked healthy. So:
218
+ - **Count two numbers every pass** — new findings, and findings caused by the last
219
+ pass's own fixes. When the second overtakes the first, the axis is exhausted:
220
+ **rotate the axis, don't look harder.** The axes are orthogonal by construction —
221
+ seams down one deliverable (the ladder above), then invariants *across*
222
+ deliverables (one name, one enum, one owner everywhere), then one class swept end
223
+ to end (every error path, every count, every status vocabulary).
224
+ - **Audit bottom-up.** A missing artefact at a low rung makes everything above it
225
+ meaningless; top-down you polish a surface for a contract that doesn't exist.
226
+ - **A class that repeats twice becomes a check, not a note.** Once is an incident;
227
+ twice is a category, and a category belongs in lint or CI where nobody has to
228
+ remember it. The third instance in a ledger is how a mechanical defect becomes
229
+ permanent.
230
+ - **What can't be fixed now becomes a ratchet, never a TODO** — a named, counted
231
+ set that may only shrink, **printed beside every gate verdict**
232
+ (`carry-over: 4 open (was 6) · unresolved: 0`). A TODO is invisible until someone
233
+ opens the file; a ratchet makes "green" read as *"green, and here is exactly what
234
+ was not looked at"*. If it grew, one sentence says why.
235
+ - **Never trust an unproven check.** Plant the defect, watch the check fail, remove
236
+ it, then trust the green — same law as the failing test, applied to every gate,
237
+ linter and script the run leans on.
238
+
164
239
  ## super-ux for user-facing tasks (recommended)
165
240
 
166
241
  If the task touches any UI (web/mobile/CLI/TUI), the WHY→UI→scenario chain comes
package/package.json CHANGED
@@ -1,10 +1,13 @@
1
1
  {
2
2
  "name": "task-pipeline-skill",
3
- "version": "0.17.1",
4
- "description": "Full-cycle task delivery pipeline orchestrator skill for Claude Code — a mandatory built-in intake grill + 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance) whose doctrine ships inside the skill with no required companion plugin, plus a super-ux UX track for user-facing tasks and toggleable release automation. This package is the installer CLI.",
3
+ "version": "1.1.0",
4
+ "description": "Full-cycle delivery pipeline for coding agents: a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine ships inside the skill — no companion plugin required. This package is the installer CLI.",
5
5
  "bin": {
6
6
  "task-pipeline": "bin/task-pipeline.js"
7
7
  },
8
+ "scripts": {
9
+ "test": "python3 test/validate.py"
10
+ },
8
11
  "files": [
9
12
  "bin",
10
13
  "plugins",
@@ -14,7 +17,8 @@
14
17
  "CHANGELOG.md"
15
18
  ],
16
19
  "repository": "github:ssheleg/task-pipeline",
17
- "homepage": "https://github.com/ssheleg/task-pipeline",
20
+ "homepage": "https://github.com/ssheleg/task-pipeline#readme",
21
+ "bugs": "https://github.com/ssheleg/task-pipeline/issues",
18
22
  "license": "MIT",
19
23
  "author": "ssheleg",
20
24
  "engines": {
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "task-pipeline",
3
- "description": "Self-contained orchestrator that runs a task through a mandatory built-in intake grill + 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a loop guard that breaks churn, one provider-agnostic model confirmed up front, a super-ux UX track for user-facing tasks, and toggleable project-configurable release automation.",
4
- "version": "0.17.1",
3
+ "description": "Runs a substantial task through a mandatory built-in intake grill, then 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a frozen requirement spine that must close with evidence, a loop guard that breaks churn, one provider-agnostic model confirmed up front, and an optional super-ux UX track for user-facing work.",
4
+ "version": "1.1.0",
5
5
  "author": {
6
6
  "name": "ssheleg"
7
7
  },
@@ -5,17 +5,27 @@ argument-hint: <one-line task description>
5
5
  Use the `task-pipeline` skill to run the task below through all gated stages —
6
6
  **stage 0 intake grill** → docs study → brainstorm → spec → plan → subagent
7
7
  build → tests → lint/deploy → post-deploy → docs/wiki → **acceptance**. **Every stage's doctrine is
8
- built into the skill** (`references/{grill,brainstorm,decomposition,spec,planning,build,review,tdd,acceptance,loop-guard}.md`)
9
- — no companion plugin is required for any of them. The **intake grill is
8
+ built into the skill** (`references/{knowledge-sources,grill,brainstorm,decomposition,spec,planning,build,review,tdd,acceptance,loop-guard}.md`)
9
+ — no companion plugin is required for any of them. **Stage 0 opens with the
10
+ knowledge harvest, before the first question** (`references/knowledge-sources.md`):
11
+ pull what the project already knows about this task from the code, `CLAUDE.md`,
12
+ `CONTEXT.md`/ADRs, `docs/` + `docs/ux/`, past pipeline briefs, the **knowledge wiki**
13
+ if one is installed ([obsidian-wiki](https://github.com/ar9av/obsidian-wiki) —
14
+ recommended, never required; detect `~/.obsidian-wiki/config`) and any **other repo
15
+ or hosted doc system the project names as its docs**, then write the **source
16
+ ledger** into the brief. The **intake grill is
10
17
  mandatory** (`references/grill.md`): interview the
11
18
  operator one question at a time (with a recommended answer each, exploring the
12
- codebase before asking) until every decision branch is resolved, applying the
19
+ codebase before asking) until every decision branch is resolved, **validating every
20
+ answer against the harvested sources** — the operator outranks any document, but
21
+ only out loud, and a doc the run proves stale is logged for the stage-9 update —
22
+ applying the
13
23
  grill's **domain awareness** (challenge terms against `CONTEXT.md`, sharpen fuzzy
14
24
  language, ADRs for hard-to-reverse calls) and covering the **autonomy sweep** (what
15
- would otherwise stop stages 1→10: docs sources, branch/tracker policy, test and lint
16
- commands, deploy target and authorization, log locations, docs/wiki targets) —
25
+ would otherwise stop stages 1→10: docs sources incl. doc repos and the wiki, branch/tracker
26
+ policy, test and lint commands, deploy target and authorization, log locations, docs/wiki targets) —
17
27
  until the brief is locked — including the **REQ table**, the request as an addressable list where every row names how it is verified — so the rest runs autonomously and the final stage can account for all of it. The list is frozen: adding is free, removing needs the operator's agreement. Anything deferred goes into the carry-over ledger the moment it's said. For any user-facing task, recommend/use
18
- **super-ux**. **If the brief describes a platform rather than a change**, stage 2 also cuts it into modules (`references/decomposition.md`) — module map committed, walking skeleton first, every REQ in exactly one module — and stages 3→10 then run per module, one brick at a time. **If any loop starts undoing an earlier pass** (same file edited twice for the same reason, a closed finding returning, a third entry into one stage), stop and run the loop guard (`references/loop-guard.md`): name both shapes, escalate to the layer that owns the conflict, re-plan the check as an ordered list, then go item by item. Honor every stage gate by its type (`auto` = verify yourself;
28
+ **super-ux**. **If the brief describes a platform rather than a change**, stage 2 also cuts it into modules (`references/decomposition.md`) — module map committed, walking skeleton first, every REQ in exactly one module — and stages 3→10 then run per module, one brick at a time. **If any loop starts undoing an earlier pass** (same file edited twice for the same reason, a closed finding returning, a third entry into one stage), stop and run the loop guard (`references/loop-guard.md`): name both shapes, escalate to the layer that owns the conflict, re-plan the check as an ordered list, then go item by item. **The closing stage opens with the ladder walk** (`references/audit.md`): the REQ table finds what was named and lost, but a comparison needs two sides and an absence has one — so walk each REQ bottom-up through its rungs (decision → spec section → contract *and its failure behavior* → task → change → executed test → surface/docs), check the seam at each step, order findings by seam rather than by file, and turn every absence into a new REQ row **before** the coverage table is written. A green from a check nobody has watched fail against a planted defect is not evidence; a finding class seen twice becomes a script rather than a third ledger row; and the carry-over ledger's counts are printed beside every gate verdict, so "green" never reads as "verified". If a searching pass starts finding mostly what the previous pass's own fixes broke, the axis is exhausted — rotate it, don't look harder. Honor every stage gate by its type (`auto` = verify yourself;
19
29
  `manual` = wait for explicit go). Confirm the **model once at preflight** —
20
30
  recommend the most capable one the environment offers, never a hardcoded id — then
21
31
  run the whole pipeline on it without re-asking.
@@ -27,7 +27,11 @@ tabled below) and an optional, toggleable `release` block. Any project replaces
27
27
  wholesale — any number of stages, run by its own skills/agents, with its own gate
28
28
  types (see *Bring your own skills*). Each gate has a **type**: `auto` (the
29
29
  orchestrator verifies the `check` itself, pass/fail) or `manual` (wait for an
30
- explicit operator go); which stages are manual is the operator's call.
30
+ explicit operator go); which stages are manual is the operator's call. In the
31
+ example's `skills[]`, `task-pipeline:<name>` denotes this skill's own built-in
32
+ doctrine (`references/<name>.md`) and `host:<name>` denotes the host project's own
33
+ command for that job (`references/conventions.md`); everything else is a real skill
34
+ the environment resolves.
31
35
 
32
36
  ## Prerequisites — none required
33
37
 
@@ -37,6 +41,7 @@ and no stage that can fail because a dependency is missing:
37
41
 
38
42
  | Stage | Built-in doctrine |
39
43
  |---|---|
44
+ | 0 Knowledge harvest (pre-grill) | [`references/knowledge-sources.md`](references/knowledge-sources.md) |
40
45
  | 0 Intake grill | [`references/grill.md`](references/grill.md) |
41
46
  | 2 Brainstorm | [`references/brainstorm.md`](references/brainstorm.md) |
42
47
  | 2 Decompose (platforms only) | [`references/decomposition.md`](references/decomposition.md) |
@@ -45,6 +50,7 @@ and no stage that can fail because a dependency is missing:
45
50
  | 5 Build (worktree, subagents, fix loop) | [`references/build.md`](references/build.md) + [`references/review.md`](references/review.md) |
46
51
  | 5–6 TDD + suite gate | [`references/tdd.md`](references/tdd.md) |
47
52
  | 10 Acceptance (REQ close-out) | [`references/acceptance.md`](references/acceptance.md) |
53
+ | 10 + any audit (what's *missing*) | [`references/audit.md`](references/audit.md) |
48
54
  | any repeating loop | [`references/loop-guard.md`](references/loop-guard.md) |
49
55
 
50
56
  **Optional bridge.** If the operator already runs an equivalent skill set (e.g.
@@ -86,7 +92,19 @@ requirements, each naming how it will be verified. Stages 3–5 trace to those i
86
92
  stage 4's gate is a mechanical set-comparison against them, and **stage 10 accounts
87
93
  for every one** — which is what turns the pipeline from a funnel into a circle.
88
94
 
89
- Two things the grill does beyond clarifying the request:
95
+ **Harvest before you ask.** Stage 0 opens with a **knowledge harvest**
96
+ ([`references/knowledge-sources.md`](references/knowledge-sources.md)), not a
97
+ question: pull what the project already knows about this task from the code,
98
+ `CLAUDE.md`, `CONTEXT.md`/ADRs, `docs/` + `docs/ux/`, past pipeline briefs, the
99
+ **knowledge wiki** if one is installed
100
+ ([obsidian-wiki](https://github.com/ar9av/obsidian-wiki) — recommended, never
101
+ required) and any **other repo or hosted doc system the project names as its
102
+ docs**. Write the source ledger into the brief, then interview *against* it: every
103
+ answer that touches a source is checked against that source, and the operator
104
+ outranks any document — but only out loud, so an override is a recorded decision
105
+ instead of an undetected divergence. The same ledger is stage 9's work list.
106
+
107
+ Three things the grill does beyond clarifying the request:
90
108
  - **Domain awareness.** It reads the project's own `CONTEXT.md` / `docs/adr/` and
91
109
  holds the operator to them — challenging terms that conflict with the glossary,
92
110
  sharpening overloaded words, stress-testing with concrete scenarios, and
@@ -108,8 +126,13 @@ Two things the grill does beyond clarifying the request:
108
126
  **and the model decision** (`references/model-tiering.md`): recommend
109
127
  the most capable model available, let the operator confirm or override, record
110
128
  it. Ask once, here.
111
- 2. **Run stage 0 (Intake grill) — always, no exceptions.** Grill until shared
112
- understanding is reached, the autonomy sweep is covered, **the REQ table is
129
+ 2. **Run stage 0 — always, no exceptions.** It opens with the **knowledge harvest**
130
+ (`references/knowledge-sources.md`): query the project's own sources — repo docs,
131
+ ADRs, `docs/ux/`, past briefs, the wiki if installed, any doc repo the project
132
+ names — for this task's terms, and write the **source ledger** into the brief
133
+ before question one. Then grill until shared
134
+ understanding is reached, **each answer checked against the harvest**, the
135
+ autonomy sweep is covered, **the REQ table is
113
136
  written (one row per independently verifiable deliverable, each naming its
114
137
  check)** and the brief is locked
115
138
  (`references/stages.md` → 0). Do not touch stage 1 before the brief is
@@ -140,7 +163,10 @@ Two things the grill does beyond clarifying the request:
140
163
  third entry into one stage — stop and run the loop guard**
141
164
  (`references/loop-guard.md`): name the two shapes, escalate to the layer that
142
165
  owns the conflict, re-plan the check as an ordered list, then go through it one
143
- item at a time; task
166
+ item at a time; **when a pass is *searching* rather than editing and starts
167
+ finding mostly what the previous pass's own fixes broke, the axis is exhausted —
168
+ rotate it, don't look harder** (`references/audit.md`), and remember that a
169
+ green from a check nobody has watched fail is not evidence; task
144
170
  tracker + conventional commits per host conventions; worktree isolation for the
145
171
  build, integrated back per the brief's branch policy before stage 7; honest
146
172
  degradation (never claim a failed/skipped step succeeded);
@@ -155,7 +181,7 @@ capable available — see `references/model-tiering.md`).
155
181
 
156
182
  | # | Stage | Invoke | Gate | Type |
157
183
  |---|---|---|---|---|
158
- | 0 | Intake grill — **mandatory** | built in: [`references/grill.md`](references/grill.md) | shared understanding reached; autonomy sweep covered; brief locked + confirmed | manual |
184
+ | 0 | Intake grill — **mandatory** | built in: [`references/knowledge-sources.md`](references/knowledge-sources.md) (harvest) → [`references/grill.md`](references/grill.md) (interview) | source ledger written; shared understanding reached; autonomy sweep covered; brief locked + confirmed | manual |
159
185
  | 1 | Docs study | `context7` (resolve-library-id → get-library-docs) / `context7-docs` | contracts grounded on fetched docs | auto |
160
186
  | 2 | Brainstorm + decompose | built in: [`references/brainstorm.md`](references/brainstorm.md) + **UI detection** + [`references/decomposition.md`](references/decomposition.md) for platforms | design approved; UI verdict recorded; every REQ answered; platform: module map approved | manual |
161
187
  | 3 | Spec | built in: [`references/spec.md`](references/spec.md) — **UI → super-ux chain first** (`/ux` → `ux-foundation` CJM → `ux-flows` screens → `ux-scenarios` → `/ux-lint`), then spec `docs/superpowers/specs/…-design.md` | committed + reviewed; UI: chain validated, linter green, scenarios/`SCR-` traced | manual |
@@ -164,8 +190,8 @@ capable available — see `references/model-tiering.md`).
164
190
  | 6 | Tests | host test runner + built-in [`references/tdd.md`](references/tdd.md) | full suite green; new/changed code covered | auto |
165
191
  | 7 | Lint + deploy | host lint → deploy per host convention | lint clean + suite green before deploy; deploy needs a go (or the brief's specific standing authorization) | manual |
166
192
  | 8 | Post-deploy | tail deploy logs / health-check | clean boot or honest degradation report | auto |
167
- | 9 | Docs + wiki | host module docs/runbook rules → `wiki-update` | docs synced, wiki synced | auto |
168
- | 10 | **Acceptance** | built in: [`references/acceptance.md`](references/acceptance.md) | every REQ accounted for with evidence; ledger has no unresolved row; operator signs off | manual |
193
+ | 9 | Docs + wiki | host module docs/runbook rules → `wiki-update` ([obsidian-wiki](https://github.com/ar9av/obsidian-wiki), recommended) | every stale row of the stage-0 source ledger updated; docs synced; wiki synced | auto |
194
+ | 10 | **Acceptance** | built in: [`references/audit.md`](references/audit.md) (ladder walk) → [`references/acceptance.md`](references/acceptance.md) (coverage table) | ladder walk ran, its absences became REQ rows; every REQ accounted for with evidence from a check seen failing once; ledger has no unresolved row; operator signs off | manual |
169
195
 
170
196
  ## Model — ask once, at preflight
171
197
 
@@ -198,8 +224,10 @@ automation is on — `pipeline.schema.json` is the only contract.
198
224
 
199
225
  - `pipeline.schema.json` — the universal pipeline config contract (stages + release)
200
226
  - `pipeline.example.json` — this plugin's default flow (stage 0 + 1→10) + release, as config
227
+ - `references/knowledge-sources.md` — stage-0 phase 1: the source list, the wiki, the ledger, the stage-9 loop-back
201
228
  - `references/grill.md` — the built-in stage-0 grill: loop, domain awareness, autonomy sweep
202
229
  - `references/acceptance.md` — the built-in stage-10 close-out: REQ coverage, evidence, sign-off
230
+ - `references/audit.md` — cross-cutting: the L0→L7 ladder and its seams (what was never written), axis rotation, ratchets, proven checks
203
231
  - `references/brainstorm.md` — stage 2: design dialogue, approaches, UI detection, hard gate
204
232
  - `references/spec.md` — stage 3: UX track order, the spec contract, self-review, review gate
205
233
  - `references/planning.md` — stage 4: zero-context plan format, parallel groups, no placeholders
@@ -211,3 +239,6 @@ automation is on — `pipeline.schema.json` is the only contract.
211
239
  - `references/conventions.md` — how stages 6–10 read the host project's CLAUDE.md
212
240
  - `references/companion-skills.md` — companion skills, install lines, preflight recommendation
213
241
  - `references/artifacts.md` — the canonical document/artifact layout per stage
242
+ - `templates/` — skeletons seeded into the host project: `brief.md` (stage 0),
243
+ `carryover.md` (seeded at 0, appended by every stage, read in full at 10),
244
+ `context.md` and `adr.md` (format references the grill writes lazily)
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "$schema": "./pipeline.schema.json",
3
3
  "version": 1,
4
- "_note": "EXAMPLE ONLY — copy this file, rename to pipeline.json in your project, and rewrite it. This particular example encodes the plugin's own default flow (an up-front intake grill + this skill's own built-in stage doctrine + a super-ux UX track for user-facing tasks); it is NOT a fixed contract. Your project defines its own stages (any count), each executed by your own skills/agents, with your own gate types. Stage models use provider-agnostic tokens ('default' = the model confirmed for the run, 'inherit' = whatever the operator is on) — never hardcode a vendor model id, it goes stale. The universal contract is pipeline.schema.json; test/validate.py checks this example against it. gate.type: auto = orchestrator verifies the check itself (pass/fail); manual = wait for an explicit operator go. Which stages are manual vs auto is the operator's decision, not the plugin's. Any repeating loop in a run (fix loop, a re-entered stage, the per-module program loop) is bound by the loop guard: log every repeat touch, stop on oscillation, escalate to the layer that owns the conflict, then re-check in a planned order.",
4
+ "_note": "EXAMPLE ONLY — copy this file, rename to pipeline.json in your project, and rewrite it. This particular example encodes the plugin's own default flow (an up-front intake grill + this skill's own built-in stage doctrine + a super-ux UX track for user-facing tasks); it is NOT a fixed contract. Reading skills[] in THIS example: a 'task-pipeline:<name>' entry is not an installable skill — it names this skill's own built-in doctrine file (references/<name>.md, e.g. task-pipeline:grill -> references/grill.md); a 'host:<name>' entry is the host project's own command for that job, resolved from its CLAUDE.md (see references/conventions.md); every other entry is a real skill/agent your environment resolves (super-ux:*, context7, wiki-query, wiki-update). In YOUR pipeline.json, put whatever names your environment actually resolves. Your project defines its own stages (any count), each executed by your own skills/agents, with your own gate types. Stage models use provider-agnostic tokens ('default' = the model confirmed for the run, 'inherit' = whatever the operator is on) — never hardcode a vendor model id, it goes stale. The universal contract is pipeline.schema.json; test/validate.py checks this example against it. gate.type: auto = orchestrator verifies the check itself (pass/fail); manual = wait for an explicit operator go. Which stages are manual vs auto is the operator's decision, not the plugin's. Any repeating loop in a run (fix loop, a re-entered stage, the per-module program loop) is bound by the loop guard: log every repeat touch, stop on oscillation, escalate to the layer that owns the conflict, then re-check in a planned order.",
5
5
  "stages": [
6
6
  {
7
7
  "id": 0,
@@ -9,11 +9,13 @@
9
9
  "name": "Intake grill",
10
10
  "model": "default",
11
11
  "skills": [
12
+ "task-pipeline:knowledge-harvest",
13
+ "wiki-query",
12
14
  "task-pipeline:grill"
13
15
  ],
14
16
  "gate": {
15
17
  "type": "manual",
16
- "check": "MANDATORY stage — never skipped (only sanctioned bypass: the entry-from-super-ux short-circuit). The grill is built into the skill (references/grill.md) — no companion to install. Per its contract: one question at a time, a recommended answer with each, explore the codebase/docs before asking, depth-first, contradictions reconciled; domain awareness applied (terms challenged against CONTEXT.md, ADRs recorded for hard-to-reverse calls). The autonomy sweep is covered — every stage 1-10 has its blockers pre-resolved (docs sources, branch/tracker policy, test + lint commands, deploy target and authorization, log/health locations, docs+wiki targets) or is explicitly marked 'stop and ask here'. UI verdict recorded (arms super-ux); model decision recorded. All of it locked into a committed task brief the operator confirms before stage 1. The REQ table is written — one row per independently verifiable deliverable, each naming how it is verified — and frozen: adding later is free, removing or narrowing needs the operator's explicit agreement. The carry-over ledger is seeded."
18
+ "check": "MANDATORY stage — never skipped (only sanctioned bypass: the entry-from-super-ux short-circuit). PHASE 1, before the first question: harvest the knowledge sources (references/knowledge-sources.md) — code, CLAUDE.md/AGENTS.md, CONTEXT.md + docs/adr, docs/ + docs/ux, past pipeline briefs and carry-over ledgers, the knowledge wiki when installed (obsidian-wiki — recommended, never required; detect ~/.obsidian-wiki/config), and any other repo or hosted doc system the project names as its docs — queried by this task's own terms, with the SOURCE LEDGER written into the brief (a row per source consulted, or an explicit 'none found'). PHASE 2, the grill, built into the skill (references/grill.md) — no companion to install. Per its contract: one question at a time, a recommended answer with each, explore the codebase/docs before asking, depth-first, contradictions reconciled; EVERY answer that touches a harvested source is validated against that source — the operator outranks any document, but only out loud, and the losing side is logged for the stage-9 doc update; domain awareness applied (terms challenged against CONTEXT.md, ADRs recorded for hard-to-reverse calls). The autonomy sweep is covered — every stage 1-10 has its blockers pre-resolved (docs sources, branch/tracker policy, test + lint commands, deploy target and authorization, log/health locations, docs+wiki targets) or is explicitly marked 'stop and ask here'. UI verdict recorded (arms super-ux); model decision recorded. All of it locked into a committed task brief the operator confirms before stage 1. The REQ table is written — one row per independently verifiable deliverable, each naming how it is verified — and frozen: adding later is free, removing or narrowing needs the operator's explicit agreement. The carry-over ledger is seeded."
17
19
  }
18
20
  },
19
21
  {
@@ -50,9 +52,11 @@
50
52
  "name": "Spec",
51
53
  "model": "default",
52
54
  "skills": [
55
+ "super-ux:ux",
53
56
  "super-ux:ux-foundation",
54
57
  "super-ux:ux-flows",
55
58
  "super-ux:ux-scenarios",
59
+ "super-ux:ux-lint",
56
60
  "task-pipeline:spec"
57
61
  ],
58
62
  "gate": {
@@ -139,7 +143,7 @@
139
143
  ],
140
144
  "gate": {
141
145
  "type": "auto",
142
- "check": "docs in sync with code in the same change; wiki synced; dangling links fixed"
146
+ "check": "the stage-0 source ledger is the work list — every source the harvest read is updated if this run changed or disproved it; docs in sync with code in the same change; wiki synced via wiki-update when obsidian-wiki is installed (absent → recommended once, never a blocker); docs living in another repository are outward — proposed with an explicit go, or carried over with the exact edit; dangling links fixed"
143
147
  }
144
148
  },
145
149
  {
@@ -148,11 +152,12 @@
148
152
  "name": "Acceptance",
149
153
  "model": "default",
150
154
  "skills": [
155
+ "task-pipeline:audit",
151
156
  "task-pipeline:acceptance"
152
157
  ],
153
158
  "gate": {
154
159
  "type": "manual",
155
- "check": "Close the circle: every REQ in the brief has a status (verified / partial / deferred / dropped) — none unknown; every verified carries evidence (a passing test name, file:line, a command and its output, or a scenario ID) — 'done' without evidence is downgraded to partial, not upgraded; every partial names what is missing and where it is tracked; every deferred/dropped has the operator's agreement and, for deferred, a tracker entry; no carry-over row is left unresolved; and the operator answers the closing question — here is what you asked for, here is what shipped, here is what is deferred, what is missing? — and signs off"
160
+ "check": "Close the circle. FIRST the LADDER WALK (references/audit.md), because the REQ table can only find what was named and lost — a comparison needs two sides and an absence has one: walk each REQ bottom-up through its rungs (decision -> spec section -> contract AND its failure behavior -> plan task -> change -> executed test -> surface/docs), check the seam at each step, order findings BY SEAM not by file, and turn every absence into a new REQ row with its check BEFORE the table is written; findings belonging to a lower layer go back to that layer (spec -> stage 3, plan -> stage 4); record the pass's two counts (new findings vs findings caused by this run's own fixes) so the next pass can tell whether the axis is exhausted. THEN the coverage table: every REQ has a status (verified / partial / deferred / dropped) — none unknown; every verified carries evidence (a passing test name, file:line, a command and its output, or a scenario ID) — 'done' without evidence is downgraded to partial, not upgraded, and a green from a check nobody has watched fail against a planted defect is not evidence at all; every partial names what is missing and where it is tracked; every deferred/dropped has the operator's agreement and, for deferred, a tracker entry; no carry-over row is left unresolved and the ledger's counts are printed beside this verdict, so 'green' never reads as 'verified'; and the operator answers the closing question — here is what you asked for, here is what shipped, here is what is deferred, what is missing? — and signs off"
156
161
  }
157
162
  }
158
163
  ],
@@ -18,10 +18,34 @@ run itself decided, deferred, or quietly dropped along the way.
18
18
  It runs **last** — after docs and wiki (stage 9), because those are deliverables
19
19
  too and a requirement may name them.
20
20
 
21
+ ## First, the ladder walk — what the list itself is missing
22
+
23
+ The REQ table answers *"did everything on the list ship?"*. It cannot answer
24
+ *"should something else have been on the list?"* — a comparison needs two sides,
25
+ and an absence has one.
26
+
27
+ So **before writing the coverage table**, walk the ladder in
28
+ [`audit.md`](audit.md): each REQ bottom-up through its rungs (decision → spec
29
+ section → contract **and its failure behavior** → task → change → executed test →
30
+ surface and docs), checking the seam at each step. It is one pass, scoped to this
31
+ run's deliverables, and it is the only part of the pipeline that can find a gap
32
+ that was never a row.
33
+
34
+ - **An absence found here becomes a new REQ row with its check**, then the table is
35
+ written. The list is frozen against *narrowing*, never against additions
36
+ ([`grill.md`](grill.md) → *The REQ spine*). Writing the table first and appending
37
+ afterwards is how acceptance goes green over a gap.
38
+ - **A finding that belongs to a lower layer goes back to that layer** — spec gaps to
39
+ stage 3, plan gaps to stage 4 — rather than being patched in place at stage 10.
40
+ - **Report the audit's two counts** (new findings; findings caused by this run's own
41
+ fixes) in the ledger. They are what tells the next pass whether the axis is
42
+ exhausted (`audit.md` → *Every pass changes the axis*).
43
+
21
44
  ## Inputs
22
45
 
23
46
  Read all of them before writing anything:
24
47
 
48
+ - the ladder walk's findings (above) — they may have added REQ rows
25
49
  - the brief's **REQ table** (`docs/superpowers/specs/<topic>-brief.md`)
26
50
  - the **carry-over ledger** (`…-carryover.md`) — in full, every row
27
51
  - the plan and its task statuses
@@ -63,8 +87,9 @@ Run: <branch/commit range> · Date: YYYY-MM-DD
63
87
  | `deferred` | agreed not to do it now | the operator's agreement **and** a tracker entry |
64
88
  | `dropped` | agreed it isn't wanted | the operator's agreement + the reason |
65
89
 
66
- There is no fifth status. A requirement nobody can classify is `unknown`, and
67
- `unknown` fails the gate — that is the whole mechanism.
90
+ Those four are the only ways a requirement may close. Anything that fits none of
91
+ them is `unknown`, and **`unknown` fails the gate** — that is the whole mechanism:
92
+ the run cannot end while a requirement is still unclassified.
68
93
 
69
94
  ## Evidence, not assertion
70
95
 
@@ -96,13 +121,20 @@ whether the run was finished.
96
121
 
97
122
  All of:
98
123
 
99
- 1. **Every REQ has a status** — none `unknown`, none blank.
100
- 2. **Every `verified` carries evidence** of the kind above.
101
- 3. **Every `partial` names what's missing** and where it's tracked.
102
- 4. **Every `deferred` / `dropped` has the operator's agreement** recorded (in the
124
+ 1. **The ladder walk ran** ([`audit.md`](audit.md)) — every REQ's rungs checked
125
+ bottom-up, findings ordered by seam, absences turned into REQ rows **before**
126
+ the table was written, and the two pass counts recorded.
127
+ 2. **Every check this gate leans on has been seen failing** at least once against a
128
+ planted defect (`audit.md` → *Exit criterion*). An unproven check's green is not
129
+ evidence.
130
+ 3. **Every REQ has a status** — none `unknown`, none blank.
131
+ 4. **Every `verified` carries evidence** of the kind above.
132
+ 5. **Every `partial` names what's missing** and where it's tracked.
133
+ 6. **Every `deferred` / `dropped` has the operator's agreement** recorded (in the
103
134
  ledger or here) and, for `deferred`, a tracker entry.
104
- 5. **No carry-over row is left `unresolved`** — every one has a home.
105
- 6. **The operator answers the closing question** and signs off.
135
+ 7. **No carry-over row is left `unresolved`** — every one has a home, and the
136
+ ledger's counts are printed with this verdict, not just filed.
137
+ 8. **The operator answers the closing question** and signs off.
106
138
 
107
139
  Manual by design. An automated check can prove the table is *well-formed*; only
108
140
  the person who asked can confirm it is *what they asked for*. Do not let a green
@@ -53,6 +53,7 @@ record (see `build.md`).
53
53
 
54
54
  | Stage | Writes | Consumed by |
55
55
  |---|---|---|
56
+ | 0 Harvest | the brief's **Knowledge sources** ledger — every source consulted, its freshness, whether this run makes it stale | the grill (validation), **stage 9** (the update work list) |
56
57
  | 0 Intake | `specs/<topic>-brief.md` — incl. the **REQ table** (seed from `templates/brief.md`) | stages 2–5, 7, 10 |
57
58
  | 0→10 all | `specs/<topic>-carryover.md` — append-only ledger (seed from `templates/carryover.md`) | stage 10, in full |
58
59
  | 10 Acceptance | `specs/<topic>-acceptance.md` — every REQ with a status and evidence | the operator |
@@ -67,24 +68,35 @@ record (see `build.md`).
67
68
  ## This repo (task-pipeline itself), for reference
68
69
 
69
70
  ```
70
- .claude-plugin/marketplace.json # marketplace manifest
71
+ .claude-plugin/marketplace.json # marketplace manifest
71
72
  plugins/task-pipeline/
72
- .claude-plugin/plugin.json
73
+ .claude-plugin/plugin.json # plugin manifest
73
74
  commands/task-pipeline.md # /task-pipeline
74
75
  skills/task-pipeline/
75
- SKILL.md
76
+ SKILL.md # the orchestrator itself
76
77
  pipeline.schema.json # generic pipeline contract
77
78
  pipeline.example.json # this plugin's own flow, as config
78
- references/{grill,brainstorm,decomposition,spec,planning,build,review,tdd,acceptance}.md # built-in stage doctrine
79
- references/loop-guard.md # cross-cutting: churn detection + break protocol
80
- references/{stages,model-tiering,conventions,artifacts,companion-skills}.md
79
+ references/ # built-in stage doctrine:
80
+ knowledge-sources.md grill.md # stage 0 (harvest, then interview)
81
+ brainstorm.md decomposition.md # stage 2
82
+ spec.md planning.md # stages 3-4
83
+ build.md review.md tdd.md # stages 5-6
84
+ acceptance.md # stage 10
85
+ audit.md # cross-cutting: the ladder + seams
86
+ loop-guard.md # cross-cutting: churn detection
87
+ stages.md model-tiering.md # gates, model policy
88
+ conventions.md artifacts.md # host conventions, this layout
89
+ companion-skills.md # optional companions + preflight
90
+ templates/ # skeletons seeded into a host project
91
+ README.md brief.md carryover.md context.md adr.md
81
92
  cursor/rules/task-pipeline.mdc # Cursor channel (self-contained rule)
82
- plugins/task-pipeline/skills/task-pipeline/templates/{brief,carryover,context,adr}.md # stage-0 skeletons (ship on every channel)
83
93
  bin/task-pipeline.js # npx installer (package task-pipeline-skill)
84
- package.json
85
94
  install.sh # POSIX installer
86
- test/validate.py # structural validator
87
- .github/workflows/{validate,release}.yml # CI + toggleable release
88
- README.md CHANGELOG.md LICENSE
95
+ test/validate.py # structural validator (npm test)
96
+ .github/workflows/{validate,release}.yml # CI + toggleable release automation
97
+ .github/ISSUE_TEMPLATE/ .github/PULL_REQUEST_TEMPLATE.md
98
+ package.json .gitignore
99
+ README.md CHANGELOG.md LICENSE CLAUDE.md
100
+ CONTRIBUTING.md SECURITY.md CODE_OF_CONDUCT.md
89
101
  docs/superpowers/{specs,plans}/ # this repo's own design history
90
102
  ```