task-pipeline-skill 0.17.1 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (24) hide show
  1. package/CHANGELOG.md +244 -0
  2. package/README.md +314 -135
  3. package/cursor/rules/task-pipeline.mdc +91 -16
  4. package/package.json +7 -3
  5. package/plugins/task-pipeline/.claude-plugin/plugin.json +2 -2
  6. package/plugins/task-pipeline/commands/task-pipeline.md +16 -6
  7. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +39 -8
  8. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +9 -4
  9. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +40 -8
  10. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +23 -11
  11. package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +224 -0
  12. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +6 -4
  13. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +8 -1
  14. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +12 -2
  15. package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +17 -3
  16. package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +37 -4
  17. package/plugins/task-pipeline/skills/task-pipeline/references/knowledge-sources.md +159 -0
  18. package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +8 -0
  19. package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +5 -3
  20. package/plugins/task-pipeline/skills/task-pipeline/references/review.md +2 -1
  21. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +73 -11
  22. package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +5 -0
  23. package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +24 -1
  24. package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +23 -0
@@ -35,10 +35,24 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
35
35
  up front so stages 1→10 need no further human input beyond the manual gates.
36
36
  This is input expansion, not design: turn "make me feature X" into locked
37
37
  answers for scope, users, constraints, data, edge cases, done-criteria.
38
+ - **Phase 1 — harvest the knowledge sources FIRST**
39
+ ([`knowledge-sources.md`](knowledge-sources.md)). Before the first question:
40
+ query what the project already knows about this task — code, `CLAUDE.md`,
41
+ `CONTEXT.md`/ADRs, `docs/` + `docs/ux/`, past pipeline briefs and carry-over
42
+ ledgers, the **knowledge wiki** if one is installed
43
+ ([obsidian-wiki](https://github.com/ar9av/obsidian-wiki) — recommended,
44
+ never required), and any **other repo or hosted doc system the project names as
45
+ its docs**. Write the **source ledger** into the brief (a row per source, or an
46
+ explicit "none found"). It is retrieval scoped by the task's own nouns, not a
47
+ read of everything — and it is what makes phase 2's answers checkable instead of
48
+ merely confident.
38
49
  - **How it runs: [`grill.md`](grill.md)** — the full doctrine, built into this
39
50
  skill (nothing to install). In short: one question per turn, a recommended
40
51
  answer with each, explore the codebase before asking, depth-first through the
41
- decision tree, contradictions reconciled on the spot; plus **domain awareness**
52
+ decision tree, contradictions reconciled on the spot; **every answer that touches
53
+ a harvested source is checked against it** — the operator outranks any document,
54
+ but only out loud, and the losing side is logged for the stage-9 doc update; plus
55
+ **domain awareness**
42
56
  (challenge terms against `CONTEXT.md`, sharpen fuzzy language, stress-test with
43
57
  concrete scenarios, cross-reference the code, record ADRs for hard-to-reverse
44
58
  calls) and the **autonomy sweep** that pre-resolves every stage-1→10 blocker.
@@ -57,8 +71,11 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
57
71
  its autonomy section instead of asking. Where the session produced them, also:
58
72
  an updated `CONTEXT.md` (terms written as they resolved) and any ADRs under
59
73
  `docs/adr/` — see `grill.md` → *Domain awareness*.
60
- - **GATE (manual):** shared understanding reached — every detected branch has a
61
- recorded answer or an explicit deferral, no open contradictions, **every
74
+ - **GATE (manual):** shared understanding reached — **the source ledger is written
75
+ (every source consulted, or an explicit "none found")**, every detected branch has
76
+ a recorded answer or an explicit deferral, **every answer that contradicted a
77
+ harvested source has a recorded resolution** (which governs, and whether the doc
78
+ is now stale), no open contradictions, **every
62
79
  autonomy-sweep row is answered or explicitly marked "stop and ask here"**, the
63
80
  **REQ table is written and every row names its check**, the carry-over ledger is
64
81
  seeded, the model decision is recorded, and the operator confirms the brief. Stop when a
@@ -123,8 +140,9 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
123
140
  map**: task analysis, user-flow diagrams (branches, error paths), every
124
141
  screen + state with wireframe and (Figma on) a Figma frame link.
125
142
  4. `ux-scenarios` → `docs/ux/scenarios.md` — the **WHAT** (source of truth for
126
- behavior): scenarios validated per the format contract (`scenario-format.md`,
127
- ux-contract v4) — IDs, statuses, `Traces:` to stories/journey stages/flows,
143
+ behavior): scenarios validated against the scenario-format contract super-ux
144
+ itself ships (`scenario-format.md` — read its current version there, never
145
+ pin one here) — IDs, statuses, `Traces:` to stories/journey stages/flows,
128
146
  edge/error states enumerated.
129
147
  5. **Run the super-ux linter** (`/ux-lint` or `python3 docs/ux/lint.py`) — it
130
148
  must pass: no drift, no orphans, no broken traces or stale Figma links.
@@ -217,12 +235,24 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
217
235
  steps — never silent success.
218
236
 
219
237
  ## 9 — Docs + wiki
238
+ - **The stage-0 source ledger is the work list** ([`knowledge-sources.md`](knowledge-sources.md)
239
+ → *Close the loop*): every source the harvest read gets updated if this run
240
+ changed or disproved it. What was worth reading at stage 0 and is wrong now is
241
+ the next run's false premise.
220
242
  - Update host module docs / runbooks per the project's self-update rules, in the
221
243
  **same change**. For UI tasks, confirm the super-ux layers were updated in this
222
- change and the linter is green (super-ux *same-change* + *no-drift* rules). Then
223
- sync knowledge to the wiki (`wiki-update` skill).
224
- - **GATE (auto):** docs in sync with code; UI: super-ux layers current + linter
225
- green; wiki synced; dangling links fixed.
244
+ change and the linter is green (super-ux *same-change* + *no-drift* rules).
245
+ - **Sync the knowledge wiki** — `wiki-update` when
246
+ [obsidian-wiki](https://github.com/ar9av/obsidian-wiki) is installed (detect:
247
+ `~/.obsidian-wiki/config`, or the skill resolves). Not installed → recommend it
248
+ once with its install line and continue; a missing wiki never blocks the gate.
249
+ Distil the knowledge (decisions, seams, why), not a diff summary.
250
+ - **Docs living in another repository** are outward: propose the edit, get an
251
+ explicit go, then open a PR there. No go → the exact edit goes in the carry-over
252
+ ledger.
253
+ - **GATE (auto):** docs in sync with code; every stale row in the source ledger
254
+ either updated or carried over with its edit; UI: super-ux layers current +
255
+ linter green; wiki synced (or absent and recommended once); dangling links fixed.
226
256
 
227
257
  ## 10 — Acceptance
228
258
  - **What:** the closing stage — go back to the brief and account for **every**
@@ -232,6 +262,17 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
232
262
  surfaces.
233
263
  - **Runs last**, after docs and wiki — those are deliverables too, and a REQ may
234
264
  name them.
265
+ - **The ladder walk runs FIRST** ([`audit.md`](audit.md)). The REQ table can only
266
+ find what was named and lost; it cannot find what was never named, because a
267
+ comparison needs two sides and an absence has one. So before the table: walk each
268
+ REQ bottom-up through its rungs (decision → spec section → contract **and its
269
+ failure behavior** → task → change → executed test → surface/docs), check the
270
+ seam at each step, and order the findings **by seam, not by file**. An absence
271
+ becomes a **new REQ row with its check** and *then* the table is written;
272
+ appending after the table is how acceptance goes green over a gap. Findings that
273
+ belong to a lower layer go back to that layer (spec → stage 3, plan → stage 4).
274
+ Record the pass's two counts — new findings, and findings caused by this run's
275
+ own fixes — so the next pass can tell whether the axis is exhausted.
235
276
  - **How it runs:** built in. Read the brief's REQ table, the carry-over ledger in
236
277
  full, the plan's task statuses, git log, the final suite output, stage-8 notes and
237
278
  stage-9 doc changes (plus `docs/ux/scenarios.md` + `/ux-lint` for UI tasks). Write
@@ -244,10 +285,14 @@ stages/agents/types (see SKILL.md → *Bring your own skills*).
244
285
  you asked for, here's what shipped, here's what's deferred and where it lives —
245
286
  what's missing?* Ask it even when the table is green; the operator holds context
246
287
  the brief never captured, and this is the cheapest moment in the run to hear it.
247
- - **GATE (manual):** every REQ has a status (none `unknown`); every `verified`
288
+ - **GATE (manual):** the ladder walk ran and its absences became REQ rows before
289
+ the table was written; **every check this gate leans on has been seen failing
290
+ once against a planted defect** (an unproven check's green is not evidence);
291
+ every REQ has a status (none `unknown`); every `verified`
248
292
  carries evidence; every `partial` names what's missing and where it's tracked;
249
293
  every `deferred`/`dropped` has the operator's agreement and, for `deferred`, a
250
- tracker entry; no carry-over row left `unresolved`; the operator answers the
294
+ tracker entry; no carry-over row left `unresolved`, and the ledger's counts are
295
+ printed with the verdict; the operator answers the
251
296
  closing question and signs off. Manual by design — an automated check can prove
252
297
  the table is well-formed, only the person who asked can confirm it is what they
253
298
  asked for.
@@ -299,3 +344,20 @@ cycle.
299
344
  re-plan the check as an ordered one-item-per-line checklist, then go through it in
300
345
  order, one commit per item. Never settle a higher-layer conflict inside a lower
301
346
  loop, and never adjudicate before the cap.
347
+
348
+ ## Cross-cutting — the audit
349
+
350
+ The loop guard governs loops that **change** things. A loop that **looks** for
351
+ things fails the other way: it converges, spending pass after pass on its own last
352
+ pass's edits while the finding count stays healthy. [`audit.md`](audit.md) is that
353
+ method and that exit — the L0→L7 ladder, the seam questions, the axis-rotation
354
+ crossover, and the rule that a green from a check nobody has watched fail is worth
355
+ nothing.
356
+
357
+ - It runs **at stage 10 before the coverage table** (the only place that can find a
358
+ requirement nobody ever wrote), **per module** in the program loop, and as the
359
+ whole task when the request is itself an audit.
360
+ - **A finding class seen twice becomes a script**, not a third ledger row.
361
+ - **Whatever can't be fixed now becomes a ratchet** — a named, counted set that may
362
+ only shrink, printed beside every gate verdict, so "green" never reads as
363
+ "verified".
@@ -15,6 +15,11 @@ NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST
15
15
 
16
16
  **If you didn't watch the test fail, you don't know it tests the right thing.**
17
17
 
18
+ The same law governs every other check in the run — a gate's `check`, a lint rule,
19
+ a host script, a detector written during an audit. A check nobody has seen fail is
20
+ a decoration that reports success. [`audit.md`](audit.md) → *Exit criterion* is
21
+ this rule raised from one test to the whole pipeline.
22
+
18
23
  Wrote code before the test? Delete it and start from the test. Not "keep it as
19
24
  reference", not "adapt it while writing tests", not "look at it once more". Delete
20
25
  means delete — code you kept is code the test was written to fit.
@@ -9,6 +9,28 @@
9
9
  - **UI verdict:** yes / no — does this touch a user-facing surface (web/mobile/CLI/TUI)?
10
10
  If yes, the stage-3 super-ux UX track is armed.
11
11
 
12
+ ## Knowledge sources (the phase-1 harvest — written BEFORE the first question)
13
+
14
+ What the project already knew about this task, and where it said so. One row per
15
+ source actually consulted; `none found` is a valid, useful row. Stage 9 updates
16
+ this same list — a source worth reading at the start is the next run's false
17
+ premise if the run leaves it wrong.
18
+
19
+ | Source | What it says about this task | Fresh? | Authority | Stale after this run? |
20
+ |---|---|---|---|---|
21
+ | `docs/adr/NNNN-….md` | … | YYYY-MM | decision | no |
22
+ | wiki: `projects/…/concepts/…` | … | YYYY-MM | context | **yes — update at stage 9** |
23
+ | `CLAUDE.md` | test/lint/deploy commands, house rules | current | convention | no |
24
+
25
+ Precedence when two disagree: **code > host docs and ADRs > wiki > memory.** The
26
+ operator outranks every document — but only **out loud**: an override quoted
27
+ against its source is a recorded decision, an unquoted one is an undetected
28
+ divergence.
29
+
30
+ - **Doc repos / hosted doc systems this project names:** … (or `none`)
31
+ - **Knowledge wiki:** installed / not installed
32
+ ([obsidian-wiki](https://github.com/ar9av/obsidian-wiki); recommended, never a gate)
33
+
12
34
  ## Scope
13
35
 
14
36
  - **In scope:** …
@@ -57,6 +79,7 @@ is not neutral — it is a scheduled interruption.
57
79
  |---|---|---|
58
80
  | run-wide | Model for this run | … (most capable available unless overridden; per-stage overrides here) |
59
81
  | run-wide | Decide autonomously vs escalate to me | … |
82
+ | 0 Harvest | Doc sources beyond this repo — other repos, hosted docs, the knowledge wiki; and may stage 9 write to them? | … (another repo is outward: propose + PR, never a direct push) |
60
83
  | 1 Docs | External libs/APIs/SDKs in play; any context7 can't resolve → where their docs live | … |
61
84
  | 2 Decompose | Platform (several capabilities/surfaces) or one module? If platform — deploy cadence: per module, or once at the end | … |
62
85
  | 2–3 Spec | UI verdict (arms super-ux); scenario-tracing waiver, if any | … |
@@ -67,7 +90,7 @@ is not neutral — it is a scheduled interruption.
67
90
  | 7 Deploy | Target + path; release automation on/off; deploy-from-main rule | … |
68
91
  | 7 Deploy | **Authorization** — standing go, or ask every time? | … |
69
92
  | 8 Post-deploy | Where logs / health live (app name, endpoint, workflow) | … |
70
- | 9 Docs+wiki | Which module docs / runbooks this change updates; wiki sync yes/no | … |
93
+ | 9 Docs+wiki | Which module docs / runbooks this change updates; wiki sync yes/no; which stale ledger rows get fixed | … |
71
94
  | 10 Acceptance | Who signs off; where deferred REQs get tracked (issue tracker / backlog) | … |
72
95
 
73
96
  > **Deploy authorization has a hard floor.** A standing go counts only if it is
@@ -28,6 +28,29 @@
28
28
  with no home is exactly the thing that gets forgotten, so acceptance refuses to
29
29
  close on it.
30
30
 
31
+ ## This ledger is a ratchet, not a TODO list
32
+
33
+ A TODO is invisible until somebody opens the file. **A ratchet is a named, counted
34
+ set that may only shrink, and it is printed beside every gate verdict:**
35
+
36
+ ```
37
+ GATE 6 tests: PASS — full suite green (247 tests)
38
+ carry-over: 4 open (was 6) · unresolved: 0 · audit findings deferred: 2
39
+ ```
40
+
41
+ That one line is the whole mechanism. Without it, `PASS` reads as *verified*; with
42
+ it, `PASS` reads as *"green, and here is exactly what was not looked at"* — which
43
+ is the true statement.
44
+
45
+ - **Print the counts at every gate**, not only at stage 10. A number nobody sees
46
+ until the end is a number nobody acts on.
47
+ - **The set may only shrink.** If it grew, the run log gets one sentence saying why.
48
+ A ratchet that grows silently is a TODO with a better name.
49
+ - **A finding class that appears twice stops belonging here** and becomes a check in
50
+ the host's lint or CI ([`audit.md`](../references/audit.md) → *A class that
51
+ repeats twice becomes a gate*). This ledger is for what cannot be automated, not
52
+ for what nobody automated.
53
+
31
54
  ## Notes
32
55
 
33
56
  - Adding a row costs one line and never blocks a stage — that is the point. The