task-pipeline-skill 0.10.0 → 0.17.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +401 -0
- package/LICENSE +85 -0
- package/README.md +211 -84
- package/cursor/rules/task-pipeline.mdc +135 -20
- package/package.json +3 -3
- package/plugins/task-pipeline/.claude-plugin/plugin.json +15 -4
- package/plugins/task-pipeline/commands/task-pipeline.md +16 -8
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +139 -57
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +41 -27
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json +1 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +118 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +32 -6
- package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +106 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/build.md +364 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +73 -21
- package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +11 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/decomposition.md +139 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +169 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +100 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/model-tiering.md +55 -22
- package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +193 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/review.md +173 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/spec.md +144 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +184 -48
- package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +110 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/README.md +11 -3
- package/plugins/task-pipeline/skills/task-pipeline/templates/adr.md +64 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +49 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +36 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/context.md +87 -0
package/README.md
CHANGED
|
@@ -5,34 +5,35 @@
|
|
|
5
5
|
[](LICENSE)
|
|
6
6
|
|
|
7
7
|
Full-cycle task delivery pipeline orchestrator for **Claude Code**. One skill that
|
|
8
|
-
runs any substantial task through an up-front **intake grill** + **
|
|
9
|
-
built
|
|
8
|
+
runs any substantial task through an up-front **intake grill** + **10 gated stages** —
|
|
9
|
+
with every stage's doctrine **built in**: no companion plugin required.
|
|
10
10
|
|
|
11
11
|
## What it does
|
|
12
12
|
|
|
13
|
-
`intake grill → docs study → brainstorm → spec → plan → subagent build →
|
|
14
|
-
lint/deploy → post-deploy log check → docs/wiki sync`
|
|
15
|
-
|
|
16
|
-
It **grills you first**: a one-line task ("make me
|
|
17
|
-
question at a time, into a locked brief
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
**manual** (waits for your go).
|
|
21
|
-
|
|
22
|
-
| # | Stage |
|
|
23
|
-
|
|
24
|
-
| 0 | Intake grill
|
|
25
|
-
| 1 | Docs study |
|
|
26
|
-
| 2 | Brainstorm
|
|
27
|
-
| 3 | Spec |
|
|
28
|
-
| 4 | Plan |
|
|
29
|
-
| 5 | Dev |
|
|
30
|
-
| 6 | Tests |
|
|
31
|
-
| 7 | Lint + deploy |
|
|
32
|
-
| 8 | Post-deploy |
|
|
33
|
-
| 9 | Docs + wiki |
|
|
34
|
-
|
|
35
|
-
|
|
13
|
+
`intake grill → docs study → brainstorm + decompose → spec → plan → subagent build →
|
|
14
|
+
tests → lint/deploy → post-deploy log check → docs/wiki sync → acceptance`
|
|
15
|
+
|
|
16
|
+
It **grills you first, always**: stage 0 is mandatory — a one-line task ("make me
|
|
17
|
+
feature X") is expanded, one question at a time, into a locked brief, and the grill
|
|
18
|
+
also sweeps stages 1→10 for anything that would stop the run later. Each stage gates
|
|
19
|
+
the next. Every gate is typed — **auto** (the orchestrator verifies it, pass/fail)
|
|
20
|
+
or **manual** (waits for your go). One model, confirmed before the run starts.
|
|
21
|
+
|
|
22
|
+
| # | Stage | Gate | Type |
|
|
23
|
+
|---|---|---|---|
|
|
24
|
+
| 0 | Intake grill — **mandatory** | shared understanding + autonomy sweep; brief locked | manual |
|
|
25
|
+
| 1 | Docs study | contracts grounded on current docs | auto |
|
|
26
|
+
| 2 | Brainstorm + decompose | design approved; UI verdict recorded; every REQ answered; platform: module map approved | manual |
|
|
27
|
+
| 3 | Spec | committed + reviewed; UI: super-ux chain validated, linter green | manual |
|
|
28
|
+
| 4 | Plan | parallel-ready, DoD per task | auto |
|
|
29
|
+
| 5 | Dev | tasks DONE (three review verdicts each), TDD green per task | auto |
|
|
30
|
+
| 6 | Tests | full suite green, new code covered | auto |
|
|
31
|
+
| 7 | Lint + deploy | lint clean + suite green before deploy | manual |
|
|
32
|
+
| 8 | Post-deploy | clean boot / honest degradation | auto |
|
|
33
|
+
| 9 | Docs + wiki | docs + wiki synced | auto |
|
|
34
|
+
| 10 | **Acceptance** | every REQ accounted for with evidence; operator signs off | manual |
|
|
35
|
+
|
|
36
|
+
These stages (0 intake + 1→10) are the plugin's **example** flow. It's a machine-readable config
|
|
36
37
|
([`pipeline.example.json`](plugins/task-pipeline/skills/task-pipeline/pipeline.example.json))
|
|
37
38
|
against a universal contract
|
|
38
39
|
([`pipeline.schema.json`](plugins/task-pipeline/skills/task-pipeline/pipeline.schema.json)):
|
|
@@ -40,15 +41,140 @@ a host project copies the example to `pipeline.json` and rewrites it with its ow
|
|
|
40
41
|
stages (any count), its own `skills[]`, and its own `auto`/`manual` gate types —
|
|
41
42
|
"bring your own skills". The framework bakes in no fixed stages.
|
|
42
43
|
|
|
43
|
-
##
|
|
44
|
+
## Everything is built in — zero required dependencies
|
|
45
|
+
|
|
46
|
+
The doctrine each stage runs on ships inside the skill. Nothing to install for it,
|
|
47
|
+
nothing to resolve at preflight, no version skew with someone else's repo, and no
|
|
48
|
+
stage that can fail because a plugin is missing:
|
|
49
|
+
|
|
50
|
+
| Stage | Built-in doctrine |
|
|
51
|
+
|---|---|
|
|
52
|
+
| 0 Intake grill | [`references/grill.md`](plugins/task-pipeline/skills/task-pipeline/references/grill.md) — interview loop, domain awareness, autonomy sweep |
|
|
53
|
+
| 2 Brainstorm | [`references/brainstorm.md`](plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md) — approaches, YAGNI, the no-code-before-approval gate |
|
|
54
|
+
| 2 Decompose | [`references/decomposition.md`](plugins/task-pipeline/skills/task-pipeline/references/decomposition.md) — platforms only: brick criteria, module map, build order |
|
|
55
|
+
| 3 Spec | [`references/spec.md`](plugins/task-pipeline/skills/task-pipeline/references/spec.md) — UX-track order, locked contracts, global constraints, self-review |
|
|
56
|
+
| 4 Plan | [`references/planning.md`](plugins/task-pipeline/skills/task-pipeline/references/planning.md) — zero-context tasks, parallel groups, no placeholders |
|
|
57
|
+
| 5 Build | [`references/build.md`](plugins/task-pipeline/skills/task-pipeline/references/build.md) + [`review.md`](plugins/task-pipeline/skills/task-pipeline/references/review.md) — isolation, ledger, subagent loop, review rubric, fix loop |
|
|
58
|
+
| 5–6 TDD | [`references/tdd.md`](plugins/task-pipeline/skills/task-pipeline/references/tdd.md) — the iron law, red/green/refactor, the suite gate |
|
|
59
|
+
| 10 Acceptance | [`references/acceptance.md`](plugins/task-pipeline/skills/task-pipeline/references/acceptance.md) — REQ coverage table, evidence rules, the closing question |
|
|
60
|
+
| any loop | [`references/loop-guard.md`](plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md) — churn detection, caps, the break protocol |
|
|
61
|
+
|
|
62
|
+
**Ported, not depended on.** Stage 0 is adapted from
|
|
63
|
+
[Matt Pocock's `grilling` / `grill-with-docs`](https://github.com/mattpocock/skills)
|
|
64
|
+
and stages 2–6 from the corresponding skills in
|
|
65
|
+
[obra/superpowers](https://github.com/obra/superpowers) — both MIT, both credited in
|
|
66
|
+
[LICENSE](LICENSE) → *Third-party*. Nothing at runtime reaches for either.
|
|
67
|
+
|
|
68
|
+
**Optional bridge:** if you already run an equivalent skill set, map it onto stages
|
|
69
|
+
2/4/5/6 in your `pipeline.json` → `skills[]`. That's a substitution, never a
|
|
70
|
+
requirement — the gates still govern, and nothing detects, recommends or waits for
|
|
71
|
+
an external provider.
|
|
72
|
+
|
|
73
|
+
## Intake grill (stage 0) — mandatory
|
|
44
74
|
|
|
45
75
|
Inspired by [Matt Pocock's grill-me](https://github.com/mattpocock/skills). Before
|
|
46
76
|
any technical work, task-pipeline interviews you relentlessly — one question per
|
|
47
77
|
turn, each with a recommended answer, exploring the codebase before asking — until
|
|
48
|
-
every decision branch is resolved and locked into a **task brief**.
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
78
|
+
every decision branch is resolved and locked into a **task brief**. There is no
|
|
79
|
+
"clear enough task" exemption: no stage-1 work starts without a committed,
|
|
80
|
+
confirmed brief.
|
|
81
|
+
|
|
82
|
+
**Built in — nothing to install.** The full doctrine ships inside the skill
|
|
83
|
+
([`references/grill.md`](plugins/task-pipeline/skills/task-pipeline/references/grill.md)):
|
|
84
|
+
no companion skill, no resolution step, no fallback path, no version skew. Adapted
|
|
85
|
+
from [Matt Pocock's grill-with-docs](https://github.com/mattpocock/skills) (MIT —
|
|
86
|
+
see [LICENSE](LICENSE) → *Third-party*).
|
|
87
|
+
|
|
88
|
+
**Domain awareness.** While exploring, the grill reads the project's own
|
|
89
|
+
`CONTEXT.md` / `docs/adr/` and holds you to them — calling out terms that conflict
|
|
90
|
+
with the glossary, replacing overloaded words with a canonical one, stress-testing
|
|
91
|
+
relationships against concrete edge cases, and surfacing where the code contradicts
|
|
92
|
+
what you just said. Resolved terms are written into `CONTEXT.md` as they land;
|
|
93
|
+
decisions that are hard to reverse, surprising without context **and** the result of
|
|
94
|
+
a real trade-off get an ADR. Both files are created lazily.
|
|
95
|
+
|
|
96
|
+
**Autonomy comes from the sweep.** Beyond the task itself, the grill pre-resolves
|
|
97
|
+
everything that would otherwise interrupt stages 1→10: which external libs need docs,
|
|
98
|
+
branch and task-tracker policy, the test command and what "green" means, the lint
|
|
99
|
+
command, the deploy target and its **authorization**, where logs and health live,
|
|
100
|
+
which docs and runbooks to update, and the model. Each gets an answer or an explicit
|
|
101
|
+
"stop and ask me here" — an unasked question is a scheduled interruption. Deploy
|
|
102
|
+
authorization has a hard floor: a standing go counts only if it names the target and
|
|
103
|
+
the preconditions.
|
|
104
|
+
|
|
105
|
+
## The REQ spine — why nothing falls out the back
|
|
106
|
+
|
|
107
|
+
Every gate before the last one asks *"is this artifact good?"* — none asks *"does
|
|
108
|
+
this still contain everything that was asked for?"* Scope doesn't leak inside a
|
|
109
|
+
stage; it leaks on the **seams**, because brief → spec → plan → task briefs is four
|
|
110
|
+
rewrites and nothing compares the lists.
|
|
111
|
+
|
|
112
|
+
So the grill's second hard output is an addressable **requirement table**: one row
|
|
113
|
+
per independently verifiable deliverable, each naming how it will be verified. A
|
|
114
|
+
requirement you can't say how to verify is a badly-stated requirement — it gets
|
|
115
|
+
split during the grill, not discovered at the end.
|
|
116
|
+
|
|
117
|
+
From there the ids thread through everything:
|
|
118
|
+
|
|
119
|
+
| Where | What it does |
|
|
120
|
+
|---|---|
|
|
121
|
+
| Spec | every section carries `covers: REQ-…` |
|
|
122
|
+
| Plan | every task carries `Implements: REQ-…`; **the gate is set equality** against the brief — a difference is printed as the explicit list of dropped requirements |
|
|
123
|
+
| Build | the implementer's brief quotes the REQ statement verbatim, so it optimises the requirement and not just the instruction |
|
|
124
|
+
| Review | a third verdict beside spec-compliance and code-quality: **does this satisfy its REQ?** |
|
|
125
|
+
| Deploy | no REQ may still be `open`; a `partial` ships only with explicit acceptance |
|
|
126
|
+
| **Acceptance** | every REQ gets `verified` / `partial` / `deferred` / `dropped` — and `verified` requires **evidence**: a passing test name, a `file:line`, a command and its output |
|
|
127
|
+
|
|
128
|
+
Two rules keep it honest. **The list is frozen** — adding mid-run is free, removing
|
|
129
|
+
or narrowing needs your explicit agreement, because silently restating the task
|
|
130
|
+
smaller makes every later gate pass honestly on a shrunken task. And **deferred out
|
|
131
|
+
loud is forgotten** — anything postponed, dropped or half-done goes into an
|
|
132
|
+
append-only carry-over ledger the moment it's said, including implementer concerns
|
|
133
|
+
and non-blocking review findings.
|
|
134
|
+
|
|
135
|
+
Stage 10 closes the circle with the question the pipeline exists to be able to
|
|
136
|
+
answer from a list rather than from memory: *here's what you asked for, here's what
|
|
137
|
+
shipped, here's what's deferred and where it lives — what's missing?*
|
|
138
|
+
|
|
139
|
+
## Platforms — decomposed into bricks, built one at a time
|
|
140
|
+
|
|
141
|
+
A one-feature task runs the pipeline once. A **platform** — several independent
|
|
142
|
+
capabilities, several separately shippable surfaces, requirements no single
|
|
143
|
+
deliverable satisfies — gets cut into modules at stage 2, before any spec is
|
|
144
|
+
written ([`references/decomposition.md`](plugins/task-pipeline/skills/task-pipeline/references/decomposition.md)).
|
|
145
|
+
|
|
146
|
+
Modules are cut **by capability, never by layer** ("Ordering", "Billing" — not
|
|
147
|
+
"Controllers", "Services"), and a candidate is only a brick when it is
|
|
148
|
+
independently specifiable, buildable and testable, owns its own entities, talks to
|
|
149
|
+
its neighbours through declared contracts only, and can land while leaving the
|
|
150
|
+
system working. The committed module map fixes the build order — **walking skeleton
|
|
151
|
+
first**, then topological, no cycles — and every requirement maps to exactly one
|
|
152
|
+
module.
|
|
153
|
+
|
|
154
|
+
Then stages 3→10 run **per module**: dossier → plan → build → tests → deploy →
|
|
155
|
+
post-deploy → docs → acceptance → next brick. Stages 0–2 run once for the platform,
|
|
156
|
+
and the map's status column is what a resumed session reads to know where it
|
|
157
|
+
stopped. Each module's spec is a full dossier: architecture, entities and
|
|
158
|
+
ownership, contracts in and out with their failure behavior, business rules, edge
|
|
159
|
+
and failure cases, UI/Figma chain, limits, open questions.
|
|
160
|
+
|
|
161
|
+
## Loop guard — churn is detected, not endured
|
|
162
|
+
|
|
163
|
+
Any repeating pass can start undoing the previous one: two shapes alternating, the
|
|
164
|
+
same file rewritten round after round, a finding that was closed coming back. That
|
|
165
|
+
looks like progress and consumes a run, so it is
|
|
166
|
+
[detected mechanically](plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md):
|
|
167
|
+
every repeat pass logs one line per touched file with **the reason that forced it**
|
|
168
|
+
— a finding id, a failed gate item. "Cleanup" is not a reason.
|
|
169
|
+
|
|
170
|
+
It trips on revert-oscillation, a file edited twice for the same reason, a
|
|
171
|
+
resurrected finding, a third entry into one stage, or two loops editing one file —
|
|
172
|
+
plus hard caps (5 fix rounds per task, 2 re-entries per stage, 3 passes per module).
|
|
173
|
+
On a trip the run **stops editing**, names shape A and shape B with their evidence,
|
|
174
|
+
escalates to the layer that owns the conflict (rubric → operator → plan → spec →
|
|
175
|
+
module map), re-plans the check as an ordered checklist with one verification
|
|
176
|
+
command per item, and goes through it one at a time. A higher-layer conflict is
|
|
177
|
+
never settled inside a lower loop.
|
|
52
178
|
|
|
53
179
|
## UX track (user-facing tasks) — super-ux recommended
|
|
54
180
|
|
|
@@ -66,12 +192,8 @@ before interface.
|
|
|
66
192
|
|
|
67
193
|
## Prerequisites
|
|
68
194
|
|
|
69
|
-
**
|
|
70
|
-
|
|
71
|
-
```
|
|
72
|
-
/plugin marketplace add obra/superpowers
|
|
73
|
-
/plugin install superpowers@superpowers
|
|
74
|
-
```
|
|
195
|
+
**None for the pipeline itself** — the doctrine for every stage ships inside the
|
|
196
|
+
skill (see *Everything is built in* above).
|
|
75
197
|
|
|
76
198
|
**super-ux** (only for user-facing tasks) — https://github.com/ssheleg/super-ux
|
|
77
199
|
|
|
@@ -138,14 +260,24 @@ on the same Claude Code install yields a duplicate skill).
|
|
|
138
260
|
|
|
139
261
|
## Use
|
|
140
262
|
|
|
141
|
-
Say *"run this through the pipeline"*
|
|
142
|
-
|
|
263
|
+
Say *"run this through the pipeline"* or *"the full cycle"*, or invoke
|
|
264
|
+
`/task-pipeline`. The skill creates a per-stage TaskList and walks the gates.
|
|
265
|
+
Russian phrasings (*"полный цикл"*, *"прогони по конвейеру"*) route the same way.
|
|
266
|
+
|
|
267
|
+
## Model policy
|
|
143
268
|
|
|
144
|
-
|
|
269
|
+
**One model, confirmed once, at preflight.** The default recommendation is *the most
|
|
270
|
+
capable reasoning model the environment offers* — currently the latest Opus
|
|
271
|
+
generation, but that's a **tier, not a string**. Model ids go stale as generations
|
|
272
|
+
ship, and you may be on another provider entirely, so nothing is hardcoded: the
|
|
273
|
+
pipeline resolves the top tier available at runtime and stage configs use
|
|
274
|
+
provider-agnostic tokens (`default` / `inherit`).
|
|
145
275
|
|
|
146
|
-
|
|
147
|
-
can't switch the main-loop model; `/model` is
|
|
148
|
-
pinned to
|
|
276
|
+
You confirm or override it (per-stage overrides welcome) before stage 0 — then it
|
|
277
|
+
**stops asking**. A skill can't switch the main-loop model; `/model` is yours.
|
|
278
|
+
Stage-5 subagents are pinned to the confirmed model automatically. If the
|
|
279
|
+
recommended tier isn't available, the pipeline says which one it's using and
|
|
280
|
+
continues — a reminder, never a block.
|
|
149
281
|
|
|
150
282
|
## Release automation (project-configurable, toggleable)
|
|
151
283
|
|
|
@@ -161,55 +293,50 @@ clean checkout. Copy and adapt it per project; nothing is hardcoded.
|
|
|
161
293
|
|
|
162
294
|
## Companion skills
|
|
163
295
|
|
|
164
|
-
`references/companion-skills.md`
|
|
165
|
-
|
|
166
|
-
install line surfaced on the spot), **
|
|
167
|
-
|
|
168
|
-
|
|
296
|
+
`references/companion-skills.md` separates what's built in (stages 0, 2, 3, 4, 5, 6
|
|
297
|
+
and 10 — nothing to install) from the short optional list: **super-ux** (required only for
|
|
298
|
+
user-facing tasks — install line surfaced on the spot), **context7** (docs stage),
|
|
299
|
+
**wiki-update** (stage 9). A single preflight block prints which are ready, which to
|
|
300
|
+
install, and the model recommendation, so you arm the whole run in one exchange.
|
|
169
301
|
|
|
170
302
|
## Portability
|
|
171
303
|
|
|
172
|
-
Stages 6–
|
|
304
|
+
Stages 6–10 read the host project's `CLAUDE.md` conventions (tests / lint / deploy /
|
|
173
305
|
docs / wiki) with detection fallbacks, so the skill works in any repo. The
|
|
174
306
|
canonical artifact layout each stage writes to is fixed in
|
|
175
307
|
[`references/artifacts.md`](plugins/task-pipeline/skills/task-pipeline/references/artifacts.md).
|
|
176
308
|
|
|
177
|
-
##
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
-
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
Запуск: скажите *«полный цикл»* / *«прогони по конвейеру»* или `/task-pipeline
|
|
210
|
-
<задача>`. Установка — см. раздел Install выше (плагин, `npx skills add
|
|
211
|
-
ssheleg/task-pipeline`, `npx task-pipeline-skill` / `npx
|
|
212
|
-
github:ssheleg/task-pipeline` или `./install.sh`).
|
|
309
|
+
## What this gives you
|
|
310
|
+
|
|
311
|
+
Agents write code well and judge *when to stop asking you things* badly. A
|
|
312
|
+
substantial task turns into twenty interruptions, or into a confident build that
|
|
313
|
+
skipped the tests. `task-pipeline` front-loads every decision into one intake
|
|
314
|
+
conversation, then runs ten gated stages without stopping to check in.
|
|
315
|
+
|
|
316
|
+
- **The intake grill asks what a senior engineer would ask** before anything is
|
|
317
|
+
touched — scope, edge cases, failure modes, rollback, who the user is — so the
|
|
318
|
+
build does not stall halfway through.
|
|
319
|
+
- **Every stage has a gate.** No code before a spec. No deploy before tests. No
|
|
320
|
+
"done" before the post-deploy logs have been read.
|
|
321
|
+
- **Team discipline without a team.** ADRs, a written plan, a real test suite, a
|
|
322
|
+
wiki entry — produced as part of the work, not promised for later.
|
|
323
|
+
- **It adapts to your repo, not the reverse.** Deploy, docs and wiki conventions
|
|
324
|
+
are read from the host project, so nothing is imposed.
|
|
325
|
+
|
|
326
|
+
## Author
|
|
327
|
+
|
|
328
|
+
Built by ssheleg — [sshlg.me](https://sshlg.me)
|
|
329
|
+
|
|
330
|
+
- X / Twitter — [@fuck_this_year](https://x.com/fuck_this_year)
|
|
331
|
+
- Telegram — [@sshlg](https://t.me/sshlg)
|
|
332
|
+
|
|
333
|
+
Part of the [ssheleg skill family](https://github.com/ssheleg/sshlg-skills):
|
|
334
|
+
`super-ux`, `task-pipeline`, `make-skill`, `sheleg-design`, `seo-aeo-audit`.
|
|
335
|
+
One command installs all five for every agent you use:
|
|
336
|
+
|
|
337
|
+
```bash
|
|
338
|
+
npx sshlg-skills install
|
|
339
|
+
```
|
|
213
340
|
|
|
214
341
|
## License
|
|
215
342
|
|
|
@@ -1,18 +1,34 @@
|
|
|
1
1
|
---
|
|
2
|
-
description: "task-pipeline: run a substantial task through a disciplined, gated delivery cycle — an intake grill that expands the request, then docs→brainstorm→spec→plan→build→tests→lint/deploy→post-deploy→docs. Apply when the user asks to run
|
|
2
|
+
description: "task-pipeline: run a substantial task through a disciplined, gated delivery cycle — an intake grill that expands the request into a checkable requirement list, then docs→brainstorm→spec→plan→build→tests→lint/deploy→post-deploy→docs/wiki→acceptance. Apply when the user asks to 'run this through the pipeline' / 'прогони по конвейеру', asks for 'the full cycle' / 'полный цикл', or starts any substantial feature/fix/build."
|
|
3
3
|
alwaysApply: false
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# task-pipeline — gated delivery cycle (Cursor)
|
|
7
7
|
|
|
8
8
|
Self-contained rule (Cursor copies this file into projects — no external links).
|
|
9
|
-
Run a substantial task through an up-front **intake grill +
|
|
9
|
+
Run a substantial task through an up-front **intake grill + 10 gated stages**. No
|
|
10
10
|
stage advances until its **gate** passes. Each gate is `auto` (verify it yourself,
|
|
11
11
|
pass/fail) or `manual` (wait for the user's explicit go).
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
**No companion skill is required.** Every rule below is the doctrine itself — the
|
|
14
|
+
grill, the design gate, the plan format, the build loop and TDD. Nothing to
|
|
15
|
+
install, nothing to resolve, nothing that breaks when a plugin is missing.
|
|
14
16
|
|
|
15
|
-
|
|
17
|
+
## Model — decide once, before you start
|
|
18
|
+
|
|
19
|
+
Recommended: **the most capable reasoning model the environment offers** (at time
|
|
20
|
+
of writing, the latest Opus generation — read that as a tier, not a string). Never
|
|
21
|
+
hardcode a model id: generations ship, tiers get renamed, and the user may be on a
|
|
22
|
+
different provider. Resolve the top tier available at runtime, tell the user what
|
|
23
|
+
you recommend and what they're on, let them confirm or override (per-stage
|
|
24
|
+
overrides welcome), record the answer — then **stop asking at every stage**. If the
|
|
25
|
+
recommended tier isn't available, say which one you're using and continue.
|
|
26
|
+
|
|
27
|
+
## Stage 0 — intake grill (MANDATORY, do this FIRST)
|
|
28
|
+
|
|
29
|
+
Never skipped, and nothing to install — the grill is part of this rule. No "the
|
|
30
|
+
task was already clear" exemption, no starting stage 1 while the user thinks. A
|
|
31
|
+
one-line task ("build feature X") is not enough to finish autonomously. Grill the
|
|
16
32
|
user up front, then run the rest without mid-flight questions:
|
|
17
33
|
1. One question per turn — never bundle.
|
|
18
34
|
2. Give a recommended answer with every question (+ one-line rationale).
|
|
@@ -20,31 +36,130 @@ user up front, then run the rest without mid-flight questions:
|
|
|
20
36
|
4. Walk the decision tree depth-first; ask prerequisite decisions first.
|
|
21
37
|
5. Reconcile contradictions; chase dodges ("decide later" → "latest you can decide
|
|
22
38
|
and still ship?").
|
|
39
|
+
6. Run the **autonomy sweep** — resolve now whatever would stop stages 1→10 later:
|
|
40
|
+
external libs and where their docs live; UI verdict; base branch, branch policy,
|
|
41
|
+
commit convention, task tracker; the test command and what "green" means; the
|
|
42
|
+
lint command; the deploy target, release toggle and **deploy authorization**;
|
|
43
|
+
where logs/health live; which docs and runbooks this change updates. Each item
|
|
44
|
+
gets an answer or an explicit "stop and ask me here" — an unasked question is a
|
|
45
|
+
scheduled interruption.
|
|
46
|
+
|
|
47
|
+
**Domain awareness while you grill.** Look for the project's own docs as you
|
|
48
|
+
explore — a root `CONTEXT.md` (or a `CONTEXT-MAP.md` pointing at per-context ones)
|
|
49
|
+
and `docs/adr/`. Then hold the user to them: call out terms that conflict with the
|
|
50
|
+
glossary ("your glossary defines 'cancellation' as X, you seem to mean Y"), replace
|
|
51
|
+
vague or overloaded words with a canonical one ("'account' — Customer or User?"),
|
|
52
|
+
stress-test relationships with concrete edge-case scenarios, and surface
|
|
53
|
+
contradictions between what the code does and what the user just said. Write
|
|
54
|
+
resolved terms into `CONTEXT.md` right then, not in a batch — one tight sentence
|
|
55
|
+
per term, aliases listed as "avoid", no implementation detail. Create these files
|
|
56
|
+
lazily, only once there's something real to write.
|
|
57
|
+
|
|
58
|
+
Offer an ADR (`docs/adr/NNNN-slug.md`, 1–3 sentences is a complete ADR) only when
|
|
59
|
+
all three hold: hard to reverse, surprising without context, and the result of a
|
|
60
|
+
real trade-off. Any one missing → skip it.
|
|
61
|
+
|
|
62
|
+
**Write the REQ spine before the brief is confirmed.** Turn the request into an
|
|
63
|
+
addressable list — one row per independently verifiable deliverable (not one per
|
|
64
|
+
sentence), each naming **how it is verified**: a test name, a `file:line`, a command
|
|
65
|
+
and its expected output, or a scenario id. *A requirement you can't say how to
|
|
66
|
+
verify is a badly-stated requirement* — split it here. Then freeze the list:
|
|
67
|
+
adding later is free, **removing or narrowing needs the user's explicit
|
|
68
|
+
agreement**. Quietly restating the task in smaller terms is the subtlest way to
|
|
69
|
+
lose it — every later gate then passes honestly on a task that shrank.
|
|
70
|
+
|
|
71
|
+
Seed a **carry-over ledger** next to the brief, append-only. Anything deferred,
|
|
72
|
+
dropped, or left half-done goes in the moment it's said — including implementer
|
|
73
|
+
concerns and non-blocking review findings. **Deferred out loud is forgotten.**
|
|
74
|
+
|
|
23
75
|
Stop when shared understanding is reached (every branch answered or explicitly
|
|
24
|
-
deferred) and lock the answers into a short brief the user confirms. Only then
|
|
76
|
+
deferred) and lock the answers into a short brief the user confirms. Only then
|
|
77
|
+
start. Later stages read that brief instead of asking again.
|
|
25
78
|
|
|
26
|
-
|
|
79
|
+
**Deploy authorization has a hard floor:** a standing go counts only if it is
|
|
80
|
+
specific (named target + named preconditions). A vague "just do everything" does
|
|
81
|
+
not authorize an outward, irreversible action — stage 7 stops and asks.
|
|
82
|
+
|
|
83
|
+
## Stages 1→10
|
|
27
84
|
|
|
28
85
|
1. **Docs study** (auto) — ground every external lib/API on current docs, not recall.
|
|
29
|
-
2. **Brainstorm** (manual) —
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
86
|
+
2. **Brainstorm + decompose** (manual) — read the brief, explore the code, then 2–3
|
|
87
|
+
approaches with a recommendation, YAGNI applied, design presented in sections and
|
|
88
|
+
approved. **Hard gate: no code, no scaffolding before that approval** — including
|
|
89
|
+
on "obviously simple" tasks. Record whether the task touches a user-facing surface
|
|
90
|
+
(UI verdict). **If the brief is a platform, not a change** — several independent
|
|
91
|
+
capabilities or separately shippable surfaces — cut it into modules before any
|
|
92
|
+
spec: modules by capability, never by layer; a module qualifies as a brick only if
|
|
93
|
+
it is independently specifiable, buildable and testable, owns its own entities,
|
|
94
|
+
talks to others through declared contracts only, and can land while leaving the
|
|
95
|
+
system working. Commit a module map (module, delivers, entities owned, depends on,
|
|
96
|
+
contracts exposed, UI?, REQs, status) in build order — the **walking skeleton
|
|
97
|
+
first**, then topologically, no cycles (a cycle means the cut is wrong). Every REQ
|
|
98
|
+
maps to exactly one module. Then stages 3→10 run **per module**, one brick at a
|
|
99
|
+
time; stages 0–2 run once, and the map's status column is the resume point. Single
|
|
100
|
+
module → record "single module: <name>" and move on.
|
|
101
|
+
3. **Spec** (manual) — write and commit the design; lock every shared contract
|
|
102
|
+
(types, schemas, signatures, file layout) plus a **Global Constraints** block of
|
|
103
|
+
exact values every later task inherits. Self-review it for placeholders,
|
|
104
|
+
contradictions, ambiguity and scope before showing it. For UI tasks, run the UX
|
|
105
|
+
chain FIRST (see super-ux below) and embed it.
|
|
106
|
+
4. **Plan** (auto) — every task names the REQ ids it `Implements:`, and the set of those ids must **equal** the brief's REQ set — a difference is scope lost, print it and fix it. Zero-context tasks: exact paths, complete code in every step,
|
|
107
|
+
the exact command and its expected output, TDD steps, DoD each, dependency graph
|
|
108
|
+
+ parallel groups, non-overlapping file ownership. **No placeholders** — no TBD,
|
|
109
|
+
no "add error handling", no "similar to Task N", no undefined types.
|
|
110
|
+
5. **Dev** (auto) — isolate first (existing worktree? native worktree tool? else
|
|
111
|
+
`git worktree` into an ignored dir), baseline tests green, then one task at a
|
|
112
|
+
time: TDD (failing test → watch it fail → minimal impl → watch it pass →
|
|
113
|
+
commit), then a review with **all three** verdicts — spec compliance, **REQ
|
|
114
|
+
satisfied** (judged against the requirement's own statement, not the task's
|
|
115
|
+
instructions), and code quality. Findings loop back to the implementer (max 5
|
|
116
|
+
rounds, then adjudicate in writing); minor findings are recorded, never blocking. Keep a progress ledger in
|
|
117
|
+
a file so a lost context can resume instead of redoing finished work. Tasks in
|
|
118
|
+
the same parallel group run concurrently **only** if each worker has its own
|
|
119
|
+
worktree — otherwise sequential. Close the stage by integrating: sync with the
|
|
120
|
+
base branch, re-run the full suite on the result, merge (or open a PR — that's
|
|
121
|
+
outward, so ask first), remove the worktree.
|
|
37
122
|
6. **Tests** (auto) — full suite green (not just new tests); new/changed code
|
|
38
|
-
covered; no skip/xfail hiding red.
|
|
123
|
+
covered including failure paths; no skip/xfail hiding red. Tests assert real
|
|
124
|
+
behavior, never mock behavior.
|
|
39
125
|
7. **Lint + deploy** (manual) — lint clean AND suite green before deploy; deploy is
|
|
40
|
-
outward → explicit user go
|
|
126
|
+
outward → explicit user go, or the specific standing authorization recorded in
|
|
127
|
+
the stage-0 brief.
|
|
41
128
|
8. **Post-deploy** (auto) — tail logs / health-check; clean boot or an honest
|
|
42
129
|
degradation report (never silent success).
|
|
43
|
-
9. **Docs** (auto) — update module docs/runbooks in the SAME change
|
|
130
|
+
9. **Docs + wiki** (auto) — update module docs/runbooks in the SAME change, and
|
|
131
|
+
sync the project's knowledge base/wiki if it has one.
|
|
132
|
+
10. **Acceptance** (manual) — the closing stage: go back to the brief and account
|
|
133
|
+
for **every** REQ. One row each, status `verified` / `partial` / `deferred` /
|
|
134
|
+
`dropped`, and every `verified` carries **evidence** — a passing test name, a
|
|
135
|
+
`file:line`, a command and its output. "Done" without evidence is downgraded to
|
|
136
|
+
`partial`, never upgraded. Then ask out loud, list in hand: *here's what you
|
|
137
|
+
asked for, here's what shipped, here's what's deferred and where it lives —
|
|
138
|
+
what's missing?* Ask it even when the table is green. Gate: no REQ `unknown`,
|
|
139
|
+
no ledger row without a home, user signs off.
|
|
140
|
+
|
|
141
|
+
Cross-cutting: answer from the brief's autonomy section rather than re-asking, log every deferral in the ledger, never narrow the task silently, track
|
|
142
|
+
tasks, conventional commits, honest degradation (never claim a failed/skipped step
|
|
143
|
+
succeeded), and get explicit go before any outward/irreversible action (deploy,
|
|
144
|
+
publish, repo create).
|
|
44
145
|
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
146
|
+
**Loop guard — break churn instead of iterating.** Every repeating pass (a fix
|
|
147
|
+
round, a re-entered stage, the next module) logs one line per touched file: the
|
|
148
|
+
file, the pass, and the reason that forced it — a finding id or a failed gate item.
|
|
149
|
+
"Cleanup" is not a reason. **Stop immediately** when: an edit undoes what an earlier
|
|
150
|
+
pass in this run deliberately did (A→B→A); the same file is edited twice for the
|
|
151
|
+
same reason; a finding already fixed or parked comes back; a stage is entered a
|
|
152
|
+
third time for one artifact; two loops edit one file. Hard caps: 5 fix rounds per
|
|
153
|
+
task, 2 re-entries per stage, 3 passes per module. On a stop: name shape A and shape
|
|
154
|
+
B with their evidence and who asks for each; escalate to the layer that owns the
|
|
155
|
+
conflict (two findings → the review rubric; finding vs plan → the user; plan vs spec
|
|
156
|
+
→ back to planning; ambiguous spec → back to spec; two modules claiming one file →
|
|
157
|
+
back to the module map) — **never settle a higher-layer conflict inside a lower
|
|
158
|
+
loop**; then re-plan the check as an ordered checklist, one owner and one
|
|
159
|
+
verification command per item, and go through it one at a time, one commit per item,
|
|
160
|
+
no opportunistic edits. Re-check the list once in the same order at the end. If it
|
|
161
|
+
trips again after a re-planned pass, stop and hand back with both shapes, the
|
|
162
|
+
evidence and your recommendation.
|
|
48
163
|
|
|
49
164
|
## super-ux for user-facing tasks (recommended)
|
|
50
165
|
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "task-pipeline-skill",
|
|
3
|
-
"version": "0.
|
|
4
|
-
"description": "Full-cycle task delivery pipeline orchestrator skill for Claude Code
|
|
3
|
+
"version": "0.17.1",
|
|
4
|
+
"description": "Full-cycle task delivery pipeline orchestrator skill for Claude Code — a mandatory built-in intake grill + 10 gated stages (docs, brainstorm+decompose, spec, plan, build, tests, lint/deploy, post-deploy, docs/wiki, acceptance) whose doctrine ships inside the skill with no required companion plugin, plus a super-ux UX track for user-facing tasks and toggleable release automation. This package is the installer CLI.",
|
|
5
5
|
"bin": {
|
|
6
6
|
"task-pipeline": "bin/task-pipeline.js"
|
|
7
7
|
},
|
|
@@ -23,11 +23,11 @@
|
|
|
23
23
|
"keywords": [
|
|
24
24
|
"pipeline",
|
|
25
25
|
"workflow",
|
|
26
|
-
"superpowers",
|
|
27
26
|
"orchestrator",
|
|
28
27
|
"tdd",
|
|
29
28
|
"deploy",
|
|
30
29
|
"spec-driven",
|
|
30
|
+
"self-contained",
|
|
31
31
|
"claude-code",
|
|
32
32
|
"skill"
|
|
33
33
|
]
|
|
@@ -1,10 +1,21 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "task-pipeline",
|
|
3
|
-
"description": "
|
|
4
|
-
"version": "0.
|
|
5
|
-
"author": {
|
|
3
|
+
"description": "Self-contained orchestrator that runs a task through a mandatory built-in intake grill + 10 gated stages (docs, brainstorm+decompose, spec, plan, subagent build, tests, lint/deploy, post-deploy, docs/wiki, acceptance). Every stage's doctrine is built into the skill — no companion plugin required — with typed auto/manual gates, a loop guard that breaks churn, one provider-agnostic model confirmed up front, a super-ux UX track for user-facing tasks, and toggleable project-configurable release automation.",
|
|
4
|
+
"version": "0.17.1",
|
|
5
|
+
"author": {
|
|
6
|
+
"name": "ssheleg"
|
|
7
|
+
},
|
|
6
8
|
"homepage": "https://github.com/ssheleg/task-pipeline",
|
|
7
9
|
"repository": "https://github.com/ssheleg/task-pipeline",
|
|
8
10
|
"license": "MIT",
|
|
9
|
-
"keywords": [
|
|
11
|
+
"keywords": [
|
|
12
|
+
"pipeline",
|
|
13
|
+
"workflow",
|
|
14
|
+
"orchestrator",
|
|
15
|
+
"tdd",
|
|
16
|
+
"deploy",
|
|
17
|
+
"spec-driven",
|
|
18
|
+
"self-contained",
|
|
19
|
+
"claude-code"
|
|
20
|
+
]
|
|
10
21
|
}
|