task-pipeline-skill 0.17.1 → 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +244 -0
- package/README.md +314 -135
- package/cursor/rules/task-pipeline.mdc +91 -16
- package/package.json +7 -3
- package/plugins/task-pipeline/.claude-plugin/plugin.json +2 -2
- package/plugins/task-pipeline/commands/task-pipeline.md +16 -6
- package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +39 -8
- package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +9 -4
- package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +40 -8
- package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +23 -11
- package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +224 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +6 -4
- package/plugins/task-pipeline/skills/task-pipeline/references/build.md +8 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +12 -2
- package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +17 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +37 -4
- package/plugins/task-pipeline/skills/task-pipeline/references/knowledge-sources.md +159 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +8 -0
- package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +5 -3
- package/plugins/task-pipeline/skills/task-pipeline/references/review.md +2 -1
- package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +73 -11
- package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +5 -0
- package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +24 -1
- package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +23 -0
|
@@ -0,0 +1,224 @@
|
|
|
1
|
+
# Audit — finding what is missing, cross-cutting
|
|
2
|
+
|
|
3
|
+
Every gate in this pipeline asks *"is this artifact good?"*. Stage 10 asks *"is
|
|
4
|
+
anything from the list lost?"* ([`acceptance.md`](acceptance.md)). **Neither asks
|
|
5
|
+
what should have been on the list and never was.**
|
|
6
|
+
|
|
7
|
+
That gap is not an oversight in the gates. It is structural: a gate compares two
|
|
8
|
+
things, and **a contradiction has two sides while an absence has one.** Comparing
|
|
9
|
+
the spec against the plan finds a requirement that was dropped. It cannot find the
|
|
10
|
+
error path nobody specified, the entity nobody gave an owner, the failure mode
|
|
11
|
+
nobody named — because on both sides of every comparison, it simply isn't there.
|
|
12
|
+
|
|
13
|
+
This file is the method that finds those. It is **cross-cutting**: stage 10 runs it
|
|
14
|
+
before writing the coverage table, the program loop runs it per module, and a task
|
|
15
|
+
whose whole job is "audit X" runs nothing else.
|
|
16
|
+
|
|
17
|
+
## Three things that are easy to confuse
|
|
18
|
+
|
|
19
|
+
| File | Runs when | Answers |
|
|
20
|
+
|---|---|---|
|
|
21
|
+
| [`acceptance.md`](acceptance.md) | stage 10 | did everything **on the list** ship, with evidence? |
|
|
22
|
+
| [`loop-guard.md`](loop-guard.md) | any **editing** loop churns | is this pass undoing the last one? |
|
|
23
|
+
| **this file** | any **audit** pass | what is broken or missing that nobody has compared? |
|
|
24
|
+
|
|
25
|
+
`loop-guard.md` governs loops that *change* things — the fix loop, a re-entered
|
|
26
|
+
stage. Its trip means a decision is being re-litigated at the wrong altitude. This
|
|
27
|
+
file governs loops that *look* for things. Its trip means the axis is exhausted,
|
|
28
|
+
which is a different failure with a different exit. Both can bind one run; they do
|
|
29
|
+
not overlap.
|
|
30
|
+
|
|
31
|
+
## Why "look again, more carefully" stops working
|
|
32
|
+
|
|
33
|
+
The method most audits use is **horizontal**: compare the documents against each
|
|
34
|
+
other, then do it again. It works, and then it fails in a way that is invisible
|
|
35
|
+
from inside it. Measured over seven passes on a production repository:
|
|
36
|
+
|
|
37
|
+
| Pass | Findings | …of which the previous pass's own fixes caused |
|
|
38
|
+
|---|---|---|
|
|
39
|
+
| 4 | 12 | 5 |
|
|
40
|
+
| 5 | 17 | 9 |
|
|
41
|
+
| 6 | 13 | 10 |
|
|
42
|
+
| 7 | 19 | 4 |
|
|
43
|
+
|
|
44
|
+
By pass six the audit was **mostly repairing itself**. Not fatigue — arithmetic.
|
|
45
|
+
Each pass edits the corpus the next pass reads, so the newest edits are always the
|
|
46
|
+
least-reviewed text present, and they are what the next pass finds. The count stays
|
|
47
|
+
healthy while the yield goes to zero.
|
|
48
|
+
|
|
49
|
+
A single **vertical** pass over the same repository — one capability walked down
|
|
50
|
+
through its layers — found nine defect classes those seven passes had been
|
|
51
|
+
structurally **unable** to see. Not missed: unable. Two of them:
|
|
52
|
+
|
|
53
|
+
- **A component that encrypts every object in the system had no key store
|
|
54
|
+
anywhere.** Its sibling's key table had been modelled for weeks. A comparison
|
|
55
|
+
needs two sides; this had one.
|
|
56
|
+
- **An edit appended a new version while the derived index still pointed at the
|
|
57
|
+
old text.** The archive is correct. The index is correct. The answer is correct —
|
|
58
|
+
*against a version nobody is looking at.* Compare archive to archive and index to
|
|
59
|
+
index: both pass. **The defect lives in the seam.**
|
|
60
|
+
|
|
61
|
+
## The ladder
|
|
62
|
+
|
|
63
|
+
The rungs are **layers of one deliverable**; the work is the **seams between
|
|
64
|
+
them**. Each rung's artefact either exists or it does not — that is what makes
|
|
65
|
+
absence findable.
|
|
66
|
+
|
|
67
|
+
| Rung | Layer | The artefact that must exist |
|
|
68
|
+
|---|---|---|
|
|
69
|
+
| **L0** | Requirement | a `REQ-###` row in the brief **with a named check** |
|
|
70
|
+
| **L1** | Decision | the locked decision, ADR or `CONTEXT.md` term this REQ rests on |
|
|
71
|
+
| **L2** | Design | a spec section carrying `covers: REQ-…` |
|
|
72
|
+
| **L3** | Contract | an exact signature or schema · **and its failure behavior** |
|
|
73
|
+
| **L4** | Task | a plan task with `Implements:` and a DoD satisfiable **as written** |
|
|
74
|
+
| **L5** | Change | the commits — the thing actually in the tree |
|
|
75
|
+
| **L6** | Test | an **executed** assertion, by name — never "the tests pass" |
|
|
76
|
+
| **L7** | Surface | what a user reaches: scenario, screen state, CLI output, runbook |
|
|
77
|
+
|
|
78
|
+
**Audit the seams, not the artifacts.** Each rung is internally consistent most of
|
|
79
|
+
the time — that is exactly what the horizontal pass is good at, and it has already
|
|
80
|
+
done it. What survives lives between rungs:
|
|
81
|
+
|
|
82
|
+
| Seam | The question | What absence looks like here |
|
|
83
|
+
|---|---|---|
|
|
84
|
+
| L0→L1 | does the requirement rest on a **recorded** decision? | a REQ whose check implies a choice nobody ever made or wrote down |
|
|
85
|
+
| L1→L2 | did the decision reach the spec? | an ADR or glossary term agreed at the grill that no spec section cites |
|
|
86
|
+
| L2→L3 | does the section name its contract **and what happens when it fails**? | "handles errors" — no code, no shape, no caller-visible reason |
|
|
87
|
+
| L3→L4 | does every contract have a task that builds it? | stage 4's set-equality covers REQ→task; **nothing** covers contract→task |
|
|
88
|
+
| L4→L5 | did the DoD land in the tree? | a DoD line nothing in the diff satisfies, marked done anyway |
|
|
89
|
+
| L5→L6 | is there an executed observable? | "tests pass"; a test that still passes with the production code deleted |
|
|
90
|
+
| L6→L7 | can a user reach it, and does a doc say so? | shipped behavior with no scenario, no `--help` line, no runbook entry |
|
|
91
|
+
| L7→L0 | does the shipped surface satisfy the requirement's **statement**? | it does what the task said and not what the requirement meant |
|
|
92
|
+
|
|
93
|
+
The last seam is stage 10's question, expressed as a seam. When it fails, the run
|
|
94
|
+
did every instruction correctly and delivered the wrong thing.
|
|
95
|
+
|
|
96
|
+
## How one audit pass runs
|
|
97
|
+
|
|
98
|
+
**Scope: one deliverable, all rungs.** One REQ, one module, one capability. Not
|
|
99
|
+
"audit the docs" and not "audit the change" — an unscoped instruction is what
|
|
100
|
+
produces seven converging passes.
|
|
101
|
+
|
|
102
|
+
**Input** is the artifact that already names every rung: the brief's REQ row plus
|
|
103
|
+
the module map row ([`decomposition.md`](decomposition.md)) when there is one. **If
|
|
104
|
+
the input can't supply the rungs, that is the first finding** — do not go looking
|
|
105
|
+
for the layers by hand; record that the spine is missing and fix that first.
|
|
106
|
+
|
|
107
|
+
**Procedure: bottom-up, L0 → L7, running the seam check at each step.**
|
|
108
|
+
|
|
109
|
+
The direction is not taste. A missing artefact at L1 makes everything above it
|
|
110
|
+
meaningless, so top-down you spend the pass polishing a surface for a contract that
|
|
111
|
+
does not exist. Bottom-up, the absence surfaces first and the six findings above it
|
|
112
|
+
collapse into one.
|
|
113
|
+
|
|
114
|
+
**Output: findings ordered by seam, never by file.** A file-ordered list reads as
|
|
115
|
+
noise; a seam-ordered one tells you **which layer of your own process is leaking**,
|
|
116
|
+
which is the thing worth knowing. Each finding carries `file:line`, the artefact
|
|
117
|
+
that is missing, and the minimal fix.
|
|
118
|
+
|
|
119
|
+
**Close through the pipeline, not around it.** A finding that is a genuine gap
|
|
120
|
+
becomes a **new REQ row** (with its check) or a carry-over row — the list is frozen
|
|
121
|
+
against *narrowing*, never against additions ([`grill.md`](grill.md) → *The REQ
|
|
122
|
+
spine*). A finding that contradicts the spec goes back to stage 3; one that
|
|
123
|
+
contradicts the plan goes back to stage 4. Auditing is not a licence to edit
|
|
124
|
+
across layers in place.
|
|
125
|
+
|
|
126
|
+
## Exit criterion — the part usually skipped
|
|
127
|
+
|
|
128
|
+
A deliverable is **not** audited when somebody has read it. It is audited when:
|
|
129
|
+
|
|
130
|
+
1. every rung has its artefact, **and**
|
|
131
|
+
2. **every check you are relying on has fired at least once against a planted
|
|
132
|
+
defect.**
|
|
133
|
+
|
|
134
|
+
**A green result from an unproven check is worth nothing.** This is the iron law of
|
|
135
|
+
[`tdd.md`](tdd.md) — *if you didn't watch it fail, you don't know it tests the
|
|
136
|
+
right thing* — raised from one test to every gate in the run. It applies to the
|
|
137
|
+
stage-4 set-equality check, the host's lint and test commands, the super-ux linter,
|
|
138
|
+
any script the host added, and every check you write during the audit itself.
|
|
139
|
+
|
|
140
|
+
Checks written under time pressure lie in ways that read as success: a predicate
|
|
141
|
+
that inspects the wrong shape and finds nothing; a probe that removes more than it
|
|
142
|
+
adds and reads the shrinkage as a pass; a regex that misses the very word it
|
|
143
|
+
searches for. All three pass loudly. **Plant the defect. Watch the check fail.
|
|
144
|
+
Remove it. Then trust the green.** Record in the ledger that you did.
|
|
145
|
+
|
|
146
|
+
## The three rules that stop this becoming another loop
|
|
147
|
+
|
|
148
|
+
### 1. A class that repeats twice becomes a gate, not a note
|
|
149
|
+
|
|
150
|
+
Once is an incident. **Twice is a category, and a category belongs in a script** —
|
|
151
|
+
the host's lint, its CI, its check runner — where nobody has to remember it.
|
|
152
|
+
|
|
153
|
+
Writing the third instance into the carry-over ledger is how a known, mechanical
|
|
154
|
+
defect class becomes permanent. If the class genuinely cannot be checked
|
|
155
|
+
mechanically, say so in one line and *say why*; that sentence is itself a finding
|
|
156
|
+
worth having.
|
|
157
|
+
|
|
158
|
+
### 2. Every pass changes the axis, not the effort
|
|
159
|
+
|
|
160
|
+
"Look again, more carefully" is what converges. Passes must be **orthogonal by
|
|
161
|
+
construction**:
|
|
162
|
+
|
|
163
|
+
1. **Seams** — one deliverable walked L0→L7 (this file's ladder).
|
|
164
|
+
2. **Invariants across deliverables** — one name, one enum, one owner, one spelling,
|
|
165
|
+
everywhere. This is the horizontal pass, and it is where it belongs.
|
|
166
|
+
3. **One class swept end to end** — every error path, every count, every status
|
|
167
|
+
vocabulary, every timeout, across the whole change at once.
|
|
168
|
+
|
|
169
|
+
**The crossover is measurable, so measure it.** Every pass, count two numbers: new
|
|
170
|
+
findings, and findings caused by the previous pass's own fixes. When the second
|
|
171
|
+
overtakes the first, the axis is exhausted — **rotate it, don't push harder.** Both
|
|
172
|
+
counts go in the ledger; an audit that reports only "found N" cannot see its own
|
|
173
|
+
exhaustion.
|
|
174
|
+
|
|
175
|
+
### 3. What can't be fixed now becomes a ratchet, never a TODO
|
|
176
|
+
|
|
177
|
+
A **ratchet** is a *named, counted set that may only shrink, printed on every
|
|
178
|
+
run*.
|
|
179
|
+
|
|
180
|
+
The carry-over ledger ([`templates/carryover.md`](../templates/carryover.md)) is
|
|
181
|
+
the pipeline's ratchet, and it only works if its count is **printed at every gate
|
|
182
|
+
beside the verdict**:
|
|
183
|
+
|
|
184
|
+
```
|
|
185
|
+
GATE 6 tests: PASS — full suite green (247 tests)
|
|
186
|
+
carry-over: 4 open (was 6) · unresolved: 0 · audit findings deferred: 2
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
The difference from a TODO is not bookkeeping. A TODO is invisible until somebody
|
|
190
|
+
opens the file. A ratchet sits next to the word `PASS` on every single run, so
|
|
191
|
+
**"green" never reads as "verified"** — it reads as *"green, and here is exactly
|
|
192
|
+
what was not looked at."* A ratchet that grew needs a sentence in the run log
|
|
193
|
+
explaining why; a ratchet nobody prints is a TODO with a better name.
|
|
194
|
+
|
|
195
|
+
## When this runs
|
|
196
|
+
|
|
197
|
+
- **Stage 10, before the coverage table.** Acceptance reads the REQ list; the
|
|
198
|
+
ladder walk is what can add to it. Absences found here become new REQ rows with
|
|
199
|
+
their checks, and *then* the table is written — otherwise acceptance closes green
|
|
200
|
+
over a gap that was never a row.
|
|
201
|
+
- **Per module in the program loop** ([`decomposition.md`](decomposition.md)) — one
|
|
202
|
+
brick's ladder, at that brick's acceptance. Cross-module contracts are audited at
|
|
203
|
+
the seam that owns them, not twice.
|
|
204
|
+
- **As the whole task**, when the operator's request *is* an audit. Then stages 3–5
|
|
205
|
+
produce findings and fixes rather than a feature, and the exit criterion above is
|
|
206
|
+
the stage-10 gate.
|
|
207
|
+
- **Never as an eighth "look again" pass.** If the last two passes found mostly
|
|
208
|
+
self-inflicted findings, the answer is rule 2, not another pass.
|
|
209
|
+
|
|
210
|
+
Once both axes are exhausted, the next finding of a known class should be caught by
|
|
211
|
+
a script — and if it cannot be, **that is the finding: write the check.**
|
|
212
|
+
|
|
213
|
+
## Rationalizations
|
|
214
|
+
|
|
215
|
+
| Excuse | Reality |
|
|
216
|
+
|---|---|
|
|
217
|
+
| "The gates all passed, so it's complete" | Gates compare. Nothing that was never written appears on either side of a comparison. |
|
|
218
|
+
| "One more careful pass will catch it" | Measured: by pass six the passes were mostly fixing their own last pass. Rotate the axis. |
|
|
219
|
+
| "I'll audit top-down, the surface is where users are" | A surface built on an absent contract wastes the whole pass. Bottom-up, that absence is finding #1. |
|
|
220
|
+
| "The check is green, that's evidence" | Only if you have seen it red. An unproven check is a decoration that reports success. |
|
|
221
|
+
| "It's a small gap, I'll note it in the ledger" | Second occurrence of a class → it goes in a script. The ledger is for what cannot be automated, not what nobody automated. |
|
|
222
|
+
| "Findings grouped by file are easier to fix" | And impossible to learn from. Group by seam; the seam names which layer of your process leaks. |
|
|
223
|
+
| "The ledger has it, we won't forget" | Only if it is printed beside every verdict. Unprinted, it is a TODO, and TODOs are invisible by construction. |
|
|
224
|
+
| "This is out of scope for the audit" | Then it is a carry-over row with a home, right now. An audit that silently declines findings is worse than none. |
|
|
@@ -37,10 +37,12 @@ not a question to re-open from scratch.
|
|
|
37
37
|
1. **Explore the current state.** Files, module docs, recent commits, the
|
|
38
38
|
conventions the repo already follows. Do this before asking anything.
|
|
39
39
|
2. **Scope check, early.** If the task actually describes several independent
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
40
|
+
capabilities or separately shippable surfaces, say so immediately: that is a
|
|
41
|
+
**platform**, and it gets cut into modules at the end of this stage by
|
|
42
|
+
[`decomposition.md`](decomposition.md), before any spec is written. Brainstorm
|
|
43
|
+
the platform's shape — the pieces, how they relate, what order they land in —
|
|
44
|
+
not the details of one corner; those belong to each module's own stage-3
|
|
45
|
+
dossier. Don't refine something that needs splitting first.
|
|
44
46
|
3. **Questions one at a time.** Never bundle. Multiple choice where it fits, open
|
|
45
47
|
where it doesn't. Purpose, constraints, success criteria — anything the brief
|
|
46
48
|
left at design level.
|
|
@@ -234,6 +234,12 @@ Two routes leave before the loop starts:
|
|
|
234
234
|
- **Minor findings** never enter it. Record each in the ledger
|
|
235
235
|
(`Task <N>: minor (deferred): <one-liner>`) and point the final review at that
|
|
236
236
|
list. A roll-up nobody reads is a silent discard.
|
|
237
|
+
- **A finding class that shows up a second time stops being a finding and becomes a
|
|
238
|
+
check.** Two tasks flagged for the same mechanical defect — the same missing
|
|
239
|
+
failure path, the same magic value, the same naming slip — means every later task
|
|
240
|
+
will produce it too. Add it to the host's lint or check script now, in its own
|
|
241
|
+
commit, instead of writing the third instance into the ledger
|
|
242
|
+
([`audit.md`](audit.md) → *A class that repeats twice becomes a gate*).
|
|
237
243
|
- **A finding that conflicts with what the plan mandates** is the operator's
|
|
238
244
|
call: present the finding beside the plan text and ask which governs. Don't
|
|
239
245
|
dismiss the finding because the plan mandated it; don't fix against the plan
|
|
@@ -304,7 +310,8 @@ findings are neither fixed nor parked-with-ruling at the cap.
|
|
|
304
310
|
## 5. Final whole-branch review
|
|
305
311
|
|
|
306
312
|
After the last task: build a package over `MERGE_BASE`..`HEAD`
|
|
307
|
-
(`git merge-base
|
|
313
|
+
(`git merge-base "$BASE_BRANCH" HEAD`, where `$BASE_BRANCH` is the base recorded in
|
|
314
|
+
the stage-0 brief — never a hardcoded `main`), dispatch the whole-branch review
|
|
308
315
|
([`review.md`](review.md) → *Final review*; on the run's model, escalation offered
|
|
309
316
|
out loud per *Models* above), and point it at the
|
|
310
317
|
ledger's deferred-minor and parked lines so it can triage what must be fixed before
|
|
@@ -12,6 +12,7 @@ better, plus one that is required only for user-facing work.
|
|
|
12
12
|
|
|
13
13
|
| Stage | Doctrine |
|
|
14
14
|
|---|---|
|
|
15
|
+
| 0 Knowledge harvest (pre-grill) | `references/knowledge-sources.md` |
|
|
15
16
|
| 0 Intake grill | `references/grill.md` |
|
|
16
17
|
| 2 Brainstorm | `references/brainstorm.md` |
|
|
17
18
|
| 2 Decompose (platforms only) | `references/decomposition.md` |
|
|
@@ -20,6 +21,7 @@ better, plus one that is required only for user-facing work.
|
|
|
20
21
|
| 5 Build (isolation, subagents, fix loop) | `references/build.md` + `references/review.md` |
|
|
21
22
|
| 5–6 TDD + suite gate | `references/tdd.md` |
|
|
22
23
|
| 10 Acceptance (REQ close-out) | `references/acceptance.md` |
|
|
24
|
+
| 10 + any audit (finding what's missing) | `references/audit.md` |
|
|
23
25
|
| any repeating loop | `references/loop-guard.md` |
|
|
24
26
|
|
|
25
27
|
## The matrix
|
|
@@ -28,7 +30,7 @@ better, plus one that is required only for user-facing work.
|
|
|
28
30
|
|---|---|---|---|
|
|
29
31
|
| **super-ux** (`ux-foundation`, `ux-flows`, `ux-scenarios`, `ux-audit`, `/ux`, `/ux-lint`) | stage 3 UX track | **Required for any user-facing task** | `/plugin marketplace add ssheleg/super-ux` → `/plugin install super-ux@super-ux` (or `npx skills add ssheleg/super-ux`) |
|
|
30
32
|
| **context7** (MCP) | stage 1 docs study | Recommended (web-search fallback) | connect the context7 MCP server |
|
|
31
|
-
| **wiki-
|
|
33
|
+
| **[obsidian-wiki](https://github.com/ar9av/obsidian-wiki)** (`wiki-query`, `wiki-update`) | **stage 0 harvest** (query what's already known) **+ stage 9 sync** | **Recommended** — never a gate; absent → harvest runs on repo docs alone | `pip install obsidian-wiki` → `obsidian-wiki setup --vault /path/to/your/vault` |
|
|
32
34
|
| ~~superpowers~~ | — | **Not a dependency.** Stages 2/4/5/6 run on the built-in doctrine above. See *Optional bridge* | — |
|
|
33
35
|
| ~~grill-me / grilling~~ | — | **Not a dependency.** The stage-0 grill is built in (`references/grill.md`) | — |
|
|
34
36
|
|
|
@@ -60,7 +62,11 @@ Pipeline companions (stage doctrine is built in — nothing to install for it):
|
|
|
60
62
|
/plugin marketplace add ssheleg/super-ux
|
|
61
63
|
/plugin install super-ux@super-ux
|
|
62
64
|
✓ context7 — ready
|
|
63
|
-
|
|
65
|
+
✗ obsidian-wiki — recommended: stage 0 queries it before grilling you,
|
|
66
|
+
stage 9 syncs back what this run learned:
|
|
67
|
+
pip install obsidian-wiki
|
|
68
|
+
obsidian-wiki setup --vault /path/to/your/vault
|
|
69
|
+
(running without it — the harvest uses repo docs only)
|
|
64
70
|
|
|
65
71
|
🧠 Model for this run: recommended <top tier available>. You're on <current>.
|
|
66
72
|
/model <id> to switch, or "keep current", or name per-stage overrides.
|
|
@@ -72,6 +78,10 @@ Rules:
|
|
|
72
78
|
|
|
73
79
|
- Only flag **super-ux** when the task implies a UI (the stage-0 grill decides;
|
|
74
80
|
when unsure, flag it — a false positive costs one install).
|
|
81
|
+
- **obsidian-wiki**: detect via `~/.obsidian-wiki/config` or a resolving
|
|
82
|
+
`wiki-query`/`wiki-update`. Present → say `✓ ready` and use it in the harvest.
|
|
83
|
+
Absent → print the two install lines **once** and continue; never ask twice in a
|
|
84
|
+
run and never block a stage on it ([`knowledge-sources.md`](knowledge-sources.md)).
|
|
75
85
|
- **Never gate any stage on an install** except the stage-3 UX track on a UI task.
|
|
76
86
|
- Optional tools missing → state the fallback, don't block.
|
|
77
87
|
- Re-detect after the operator installs; don't assume.
|
|
@@ -1,7 +1,10 @@
|
|
|
1
|
-
# Host conventions (stages 6–10)
|
|
1
|
+
# Host conventions (stage 0 harvest, stages 6–10)
|
|
2
2
|
|
|
3
3
|
The orchestrator is project-agnostic. For tests / lint / deploy / docs / wiki it reads the
|
|
4
4
|
**host project's `CLAUDE.md` / `AGENTS.md` first**, then falls back to detection.
|
|
5
|
+
The same files are the stage-0 harvest's first stop — they are where a project
|
|
6
|
+
names its doc repos, its knowledge base and its house rules
|
|
7
|
+
([`knowledge-sources.md`](knowledge-sources.md)).
|
|
5
8
|
Prefer explicit host instructions over detection; if a step's convention can't be
|
|
6
9
|
found, surface it and **ask** rather than guessing.
|
|
7
10
|
|
|
@@ -30,9 +33,20 @@ found, surface it and **ask** rather than guessing.
|
|
|
30
33
|
CI: the workflow run. Hit the health endpoint if one is defined.
|
|
31
34
|
|
|
32
35
|
## Docs + wiki
|
|
36
|
+
- **Start from the stage-0 source ledger** ([`knowledge-sources.md`](knowledge-sources.md)):
|
|
37
|
+
the sources the harvest read are the sources this stage updates. Anything the run
|
|
38
|
+
proved stale is already listed there with what's wrong.
|
|
33
39
|
- Host self-update rules (module docs, runbooks, agent-self cards, etc.) — update
|
|
34
|
-
in the same change.
|
|
35
|
-
|
|
40
|
+
in the same change. Fix dangling links.
|
|
41
|
+
- **Wiki:** [obsidian-wiki](https://github.com/ar9av/obsidian-wiki) — the
|
|
42
|
+
`wiki-update` skill (resolves the vault via `~/.obsidian-wiki/config`). Detect it
|
|
43
|
+
the same way the harvest does; if absent, recommend it once
|
|
44
|
+
(`pip install obsidian-wiki` → `obsidian-wiki setup --vault <path>`) and continue.
|
|
45
|
+
A project may of course use a different knowledge base — then its own
|
|
46
|
+
`CLAUDE.md` names the sync command, and that wins.
|
|
47
|
+
- **Docs in another repository** (a docs repo, a submodule, a sibling checkout the
|
|
48
|
+
project names): updating it is **outward** — propose the change, get an explicit
|
|
49
|
+
operator go, open a PR there. Never push to a repo the task didn't name.
|
|
36
50
|
|
|
37
51
|
## Issue tracker (stage 10)
|
|
38
52
|
|
|
@@ -12,7 +12,27 @@ coming back to the operator.
|
|
|
12
12
|
> half — glossary challenges, `CONTEXT.md`, ADR discipline — comes from there; the
|
|
13
13
|
> autonomy sweep and the brief are this pipeline's.
|
|
14
14
|
|
|
15
|
-
##
|
|
15
|
+
## Phase 1 — harvest before you ask
|
|
16
|
+
|
|
17
|
+
**Do not open the interview cold.** Stage 0 begins by finding what the project
|
|
18
|
+
already knows about this task: the code, `CLAUDE.md`, `CONTEXT.md` and the ADRs,
|
|
19
|
+
`docs/` and `docs/ux/`, past pipeline briefs, the **knowledge wiki** when one is
|
|
20
|
+
installed, and any **other repository or hosted doc system the project names as
|
|
21
|
+
its docs**. Full procedure, source order, the wiki's detection and install line,
|
|
22
|
+
and the ledger to write: [`knowledge-sources.md`](knowledge-sources.md).
|
|
23
|
+
|
|
24
|
+
Two things come out of it, both required before question one:
|
|
25
|
+
|
|
26
|
+
- the **source ledger** in the brief — one row per source consulted, what it says
|
|
27
|
+
about this task, and how fresh it is (`no sources found` is a valid row);
|
|
28
|
+
- the list of things you therefore **don't need to ask**, and the specific points
|
|
29
|
+
where a source looks stale or ambiguous — those become the sharpest questions.
|
|
30
|
+
|
|
31
|
+
Everything below runs against that harvest. An answer you can't check against a
|
|
32
|
+
source is a recollection, and the whole loop exists to stop the run from building
|
|
33
|
+
on one.
|
|
34
|
+
|
|
35
|
+
## Phase 2 — the loop
|
|
16
36
|
|
|
17
37
|
Interview the operator relentlessly about every aspect of the task until you reach
|
|
18
38
|
a **shared understanding**. Walk down each branch of the decision tree, resolving
|
|
@@ -73,6 +93,15 @@ Create these files **lazily** — only when you have something real to write.
|
|
|
73
93
|
check whether the code agrees, and surface contradictions: *"Your code cancels
|
|
74
94
|
entire Orders, but you just said partial cancellation is possible — which is
|
|
75
95
|
right?"*
|
|
96
|
+
- **Cross-reference with the harvest — every answer, not just the domain ones.**
|
|
97
|
+
Phase 1 put the ADRs, runbooks and wiki pages in your hands; use them the same
|
|
98
|
+
way: *"The March ADR says orders are written only through the command handler,
|
|
99
|
+
you just described a direct write — has that changed?"* The operator **outranks
|
|
100
|
+
every document**, but only out loud: an override quoted against its source is a
|
|
101
|
+
recorded decision, an unquoted one is an undetected divergence. When two sources
|
|
102
|
+
disagree, precedence is code > host docs/ADRs > wiki > memory, and the loser is
|
|
103
|
+
logged for the stage-9 update ([`knowledge-sources.md`](knowledge-sources.md) →
|
|
104
|
+
*Phase 2*).
|
|
76
105
|
- **Update `CONTEXT.md` inline.** Resolve a term → write it down right then, not in
|
|
77
106
|
a batch at the end. Format: [`templates/context.md`](../templates/context.md).
|
|
78
107
|
Keep it free of implementation detail — only terms a domain expert would
|
|
@@ -101,6 +130,7 @@ explicit "stop and ask me here":
|
|
|
101
130
|
| Stage | What to settle up front |
|
|
102
131
|
|---|---|
|
|
103
132
|
| run-wide | the model decision ([`model-tiering.md`](model-tiering.md)); what to decide autonomously vs escalate |
|
|
133
|
+
| 0 Harvest | doc sources beyond this repo — other repos, hosted doc systems, the knowledge wiki — and whether stage 9 may write to them (another repo is outward: propose + PR, never a direct push) |
|
|
104
134
|
| 1 Docs | external libs/APIs/SDKs in play; any private ones context7 can't resolve → where their docs live |
|
|
105
135
|
| 2 Decompose | is this a platform (several capabilities/surfaces) or one module? if platform: deploy cadence — per module or once at the end |
|
|
106
136
|
| 2–3 Spec | UI verdict (arms super-ux); any scenario-tracing waiver |
|
|
@@ -153,9 +183,12 @@ that shrank without anyone deciding it should.
|
|
|
153
183
|
Everything resolved goes into the **task brief**, seeded from
|
|
154
184
|
[`templates/brief.md`](../templates/brief.md) and committed to
|
|
155
185
|
`docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` — scope, **the REQ table**,
|
|
156
|
-
users, UI verdict, constraints, locked decisions,
|
|
157
|
-
done-criteria, open assumptions. Seed the template only when
|
|
158
|
-
never overwrite an existing brief.
|
|
186
|
+
**the phase-1 source ledger**, users, UI verdict, constraints, locked decisions,
|
|
187
|
+
the autonomy table, done-criteria, open assumptions. Seed the template only when
|
|
188
|
+
the file is absent; never overwrite an existing brief.
|
|
189
|
+
|
|
190
|
+
The ledger is not decoration: **stage 9 updates exactly what stage 0 read**, and
|
|
191
|
+
every doc the grill proved stale is already listed there with what's wrong.
|
|
159
192
|
|
|
160
193
|
Alongside it, seed the **carry-over ledger** from
|
|
161
194
|
[`templates/carryover.md`](../templates/carryover.md) at
|
|
@@ -0,0 +1,159 @@
|
|
|
1
|
+
# Knowledge sources — harvest before the grill, update after the build
|
|
2
|
+
|
|
3
|
+
Stage 0 has two phases. This file is **phase 1**: before the first question is
|
|
4
|
+
asked, find and read what the project already knows about this task. The interview
|
|
5
|
+
([`grill.md`](grill.md)) is phase 2, and it runs *against* what was harvested here.
|
|
6
|
+
|
|
7
|
+
The same source list closes the loop at **stage 9**: what was read at the start is
|
|
8
|
+
what gets updated at the end. A source good enough to answer a question is a source
|
|
9
|
+
that goes stale when the answer changes.
|
|
10
|
+
|
|
11
|
+
## Why this is a phase and not "explore a bit first"
|
|
12
|
+
|
|
13
|
+
An agent that starts asking without harvesting spends the operator's turns on
|
|
14
|
+
questions the project already answered — in an ADR, in a runbook, in a wiki page
|
|
15
|
+
written three months ago by the same person now being asked. That is the expensive
|
|
16
|
+
failure, but not the worst one.
|
|
17
|
+
|
|
18
|
+
The worst one is silent: **the operator misremembers, the agent believes them, and
|
|
19
|
+
the run builds on it.** People answer from memory about systems they wrote a year
|
|
20
|
+
ago. Without the documents in hand you cannot tell a decision from a recollection,
|
|
21
|
+
so every later gate passes honestly on a false premise. Harvesting first is what
|
|
22
|
+
makes the grill's answers *checkable* instead of merely confident.
|
|
23
|
+
|
|
24
|
+
## The sources, in the order to try them
|
|
25
|
+
|
|
26
|
+
| # | Source | How to find it | What it's good for |
|
|
27
|
+
|---|---|---|---|
|
|
28
|
+
| 1 | **The code** | the repo you're in | what actually runs — the tiebreaker |
|
|
29
|
+
| 2 | **Host agent docs** | `CLAUDE.md`, `AGENTS.md`, `.cursor/rules/` | conventions, commands, deploy path, house rules |
|
|
30
|
+
| 3 | **Domain docs** | `CONTEXT.md` / `CONTEXT-MAP.md`, `docs/adr/` | the glossary and the decisions with their reasons |
|
|
31
|
+
| 4 | **Product/UX docs** | `docs/ux/` (super-ux chain), `README`, runbooks | user-facing behavior that is already specified |
|
|
32
|
+
| 5 | **Pipeline history** | `docs/superpowers/specs/`, `plans/`, past `-carryover.md` | what a previous run of this pipeline decided or deferred |
|
|
33
|
+
| 6 | **The knowledge wiki** | see below | distilled cross-project knowledge, prior sessions, why decisions were made |
|
|
34
|
+
| 7 | **Other doc repos the project names** | a docs repo URL or submodule in `CLAUDE.md`/`README`, a sibling checkout, a `docs/` monorepo package | specs, contracts and runbooks that live outside this repo |
|
|
35
|
+
| 8 | **Hosted doc systems the project names** | Notion / Confluence / Google Docs referenced in the project | the same, when the team keeps them there |
|
|
36
|
+
|
|
37
|
+
Rules for the list:
|
|
38
|
+
|
|
39
|
+
- **Never invent a source.** A doc repo is in scope because the project names it,
|
|
40
|
+
not because it plausibly exists. Nothing is cloned or fetched on a guess.
|
|
41
|
+
- **Sources 7–8 are read-only at this stage**, and reading a hosted system needs a
|
|
42
|
+
connected tool — if there's no tool, record the gap and ask the operator to paste
|
|
43
|
+
what matters rather than pretending the source was covered.
|
|
44
|
+
- **The wiki is optional; the harvest is not.** With no wiki and no doc repos, the
|
|
45
|
+
harvest is sources 1–5 and takes two minutes. Skipping it is never the answer.
|
|
46
|
+
|
|
47
|
+
## The knowledge wiki — recommended
|
|
48
|
+
|
|
49
|
+
The wiki this pipeline is built to work with is
|
|
50
|
+
**[obsidian-wiki](https://github.com/ar9av/obsidian-wiki)** (Karpathy's LLM-wiki
|
|
51
|
+
pattern: raw sources → distilled wiki → schema). It is the one source that carries
|
|
52
|
+
*why* across projects and across months, which is exactly what a fresh context lacks.
|
|
53
|
+
|
|
54
|
+
**Detect it** — any of: `~/.obsidian-wiki/config` exists; the `wiki-query` /
|
|
55
|
+
`wiki-update` skills resolve.
|
|
56
|
+
|
|
57
|
+
- **Installed → use it.** Query it during the harvest (`wiki-query`, or the vault's
|
|
58
|
+
`index.md` + a targeted grep when the skill isn't loaded), and sync back at stage
|
|
59
|
+
9 (`wiki-update`).
|
|
60
|
+
- **Not installed → recommend it once, in the preflight block, with the line:**
|
|
61
|
+
|
|
62
|
+
```
|
|
63
|
+
pip install obsidian-wiki
|
|
64
|
+
obsidian-wiki setup --vault /path/to/your/vault
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
Then continue without it. It is a **recommendation, never a gate** — no stage
|
|
68
|
+
blocks on a missing wiki, and the pipeline never nags twice in a run.
|
|
69
|
+
|
|
70
|
+
## How to harvest — retrieval, not reading
|
|
71
|
+
|
|
72
|
+
The harvest is bounded by the *task*, not by the size of the sources. You are not
|
|
73
|
+
reading the wiki; you are asking it about this task.
|
|
74
|
+
|
|
75
|
+
1. **Take the task's nouns** — the entities, the feature name, the subsystem, the
|
|
76
|
+
file paths the operator mentioned — plus their obvious synonyms.
|
|
77
|
+
2. **Query each source with those terms**: `wiki-query` for the wiki; `grep`/`Read`
|
|
78
|
+
for repo docs; the tracker/hosted-doc tool if one is connected.
|
|
79
|
+
3. **Follow one hop, not ten.** A hit that names an ADR, a scenario id or a module
|
|
80
|
+
is worth opening. A page three links away is context, not evidence.
|
|
81
|
+
4. **Stop when the terms stop returning anything new.** Same rule as the interview:
|
|
82
|
+
no grinding past diminishing returns.
|
|
83
|
+
|
|
84
|
+
## Record it — the source ledger
|
|
85
|
+
|
|
86
|
+
Write what you found into the brief's **Knowledge sources** section
|
|
87
|
+
([`templates/brief.md`](../templates/brief.md)) before the first question. One row
|
|
88
|
+
per source actually consulted:
|
|
89
|
+
|
|
90
|
+
| Source | What it says about this task | Fresh? | Authority |
|
|
91
|
+
|---|---|---|---|
|
|
92
|
+
| `docs/adr/0007-single-write-model.md` | orders are written only through the command handler | 2026-03 | decision |
|
|
93
|
+
| wiki: `projects/x/concepts/billing-seams` | why invoicing was split out; the retry rule | 2026-06 | context |
|
|
94
|
+
| `CLAUDE.md` | test = `npm test`, deploy from `main` only | current | convention |
|
|
95
|
+
| (none for the export UI) | — | — | — |
|
|
96
|
+
|
|
97
|
+
The ledger is what makes phase 2 work: during the interview you cite rows from it,
|
|
98
|
+
and at stage 9 you update the same rows. A source consulted but not recorded is a
|
|
99
|
+
source nobody will update.
|
|
100
|
+
|
|
101
|
+
**"No sources found" is a valid, recorded outcome.** Write the row. An empty ledger
|
|
102
|
+
tells the next run that the search happened and came back empty — silence doesn't.
|
|
103
|
+
|
|
104
|
+
## Phase 2 — validate the answers against the harvest
|
|
105
|
+
|
|
106
|
+
This is the payoff, and it belongs to the grill loop
|
|
107
|
+
([`grill.md`](grill.md) → *Domain awareness*). Every operator answer that touches a
|
|
108
|
+
harvested source gets checked against it, on the spot:
|
|
109
|
+
|
|
110
|
+
> "The ADR from March says orders are written only through the command handler —
|
|
111
|
+
> you just described a direct write. Has that changed, or should the export go
|
|
112
|
+
> through the handler?"
|
|
113
|
+
|
|
114
|
+
Three shapes and what to do with each:
|
|
115
|
+
|
|
116
|
+
| The answer… | Do |
|
|
117
|
+
|---|---|
|
|
118
|
+
| **agrees** with the source | nothing — note it, move on |
|
|
119
|
+
| **contradicts** a source | quote the source, name the conflict, ask which governs. The answer is either "the doc is stale" (→ it gets updated at stage 9, log it now) or "I misremembered" (→ the doc stands). Both are cheap here and expensive at stage 6 |
|
|
120
|
+
| **goes beyond** every source | this is new knowledge — it belongs in the brief, and usually in `CONTEXT.md` or an ADR as it lands |
|
|
121
|
+
|
|
122
|
+
**The operator outranks the docs — but only out loud.** A person may overrule any
|
|
123
|
+
document; they may not do it by accident. The point of quoting the source is that
|
|
124
|
+
the override becomes a recorded decision instead of an undetected divergence.
|
|
125
|
+
|
|
126
|
+
**Precedence when two sources disagree with each other:** code > host docs and
|
|
127
|
+
ADRs > the wiki > anyone's memory. The wiki is *distilled* knowledge and can lag
|
|
128
|
+
the repo by months; the code is what runs. A disagreement between them is a grill
|
|
129
|
+
question, never a silent pick — and it is usually a sign the doc is due an update.
|
|
130
|
+
|
|
131
|
+
## Close the loop — stage 9 updates what stage 0 read
|
|
132
|
+
|
|
133
|
+
The ledger is the stage-9 work list. For each row:
|
|
134
|
+
|
|
135
|
+
- **Host repo docs, ADRs, runbooks, `docs/ux/`** — updated in the **same change**,
|
|
136
|
+
per the host's own rules ([`conventions.md`](conventions.md)).
|
|
137
|
+
- **Anything the run proved stale** — including a doc that was "wrong but nobody
|
|
138
|
+
had time": that's why the conflict was logged in phase 2 instead of only being
|
|
139
|
+
resolved verbally.
|
|
140
|
+
- **The wiki** — `wiki-update` syncs what this run learned. Distil the *knowledge*
|
|
141
|
+
(decisions, seams, gotchas, why), never a diff summary.
|
|
142
|
+
- **Another repository's docs** — writing to a repo the operator didn't ask you to
|
|
143
|
+
touch is **outward**: propose the change, get an explicit go, then open a PR
|
|
144
|
+
there. Absent a go, it goes in the carry-over ledger with the exact edit needed.
|
|
145
|
+
|
|
146
|
+
A source that was worth reading at stage 0 and is wrong at stage 9 is the next
|
|
147
|
+
run's false premise. Closing that loop is the whole point of harvesting from a
|
|
148
|
+
written list instead of from whatever the search happened to surface.
|
|
149
|
+
|
|
150
|
+
## Rationalizations
|
|
151
|
+
|
|
152
|
+
| Excuse | Reality |
|
|
153
|
+
|---|---|
|
|
154
|
+
| "I'll just ask them, it's faster" | You'll ask about things a doc already answers, and you'll believe an answer you can't check. Retrieval is cheaper than a turn. |
|
|
155
|
+
| "The wiki's probably stale" | Then say so with the page in hand and get it corrected. "Probably stale" unread is an assumption; read, it's a finding. |
|
|
156
|
+
| "No docs in this repo" | Check `CLAUDE.md` for the repo that has them, and the wiki for the last time anyone touched this. Then record the empty ledger. |
|
|
157
|
+
| "The operator knows their own system" | They do — a year ago, before three other people changed it. That's the exact case where quoting the doc pays. |
|
|
158
|
+
| "Reading the whole wiki costs too much" | The harvest is a query per task noun, not a read. If it feels expensive, you're reading instead of retrieving. |
|
|
159
|
+
| "I'll update the docs at the end from memory" | The ledger exists because the end is exactly when you no longer remember which sources you leaned on. |
|
|
@@ -10,6 +10,14 @@ the pipeline: the stage-5 fix loop, a stage re-entered after a failed gate, the
|
|
|
10
10
|
per-module program loop ([`decomposition.md`](decomposition.md)), and any
|
|
11
11
|
audit → fix → audit cycle.
|
|
12
12
|
|
|
13
|
+
**This file governs loops that *change* things.** A loop that *looks* for things —
|
|
14
|
+
pass after pass over one corpus — fails differently: it does not oscillate, it
|
|
15
|
+
**converges**, quietly spending each pass on the previous pass's own edits while the
|
|
16
|
+
finding count stays healthy. That has its own detector and its own exit (rotate the
|
|
17
|
+
axis, don't push harder): [`audit.md`](audit.md) → *Every pass changes the axis*.
|
|
18
|
+
Both can bind one run. Use this file's trips for edits, that file's crossover for
|
|
19
|
+
searches.
|
|
20
|
+
|
|
13
21
|
## Bookkeeping — the thing that makes detection mechanical
|
|
14
22
|
|
|
15
23
|
You cannot detect churn from memory, especially after compaction. Every repeating
|
|
@@ -21,9 +21,11 @@ brief and the spec.
|
|
|
21
21
|
|
|
22
22
|
## Before writing tasks
|
|
23
23
|
|
|
24
|
-
**Scope check.**
|
|
25
|
-
one
|
|
26
|
-
|
|
24
|
+
**Scope check.** A plan covers exactly **one** spec — for a decomposed platform,
|
|
25
|
+
one module's dossier ([`decomposition.md`](decomposition.md)). If the spec in front
|
|
26
|
+
of you covers several independent subsystems, the decomposition was missed at stage
|
|
27
|
+
2: say so and go back there for a module map, rather than inventing the split here.
|
|
28
|
+
Whatever this plan covers must produce working, testable software on its own.
|
|
27
29
|
|
|
28
30
|
**Map the file structure.** List every file that will be created or modified and
|
|
29
31
|
what each one owns. This is where decomposition gets locked in:
|
|
@@ -29,7 +29,8 @@ this plan's git-ignored directory, `.task-pipeline/build/<plan-basename>/` — s
|
|
|
29
29
|
`BASE` is the commit you recorded **before** dispatching the implementer — never
|
|
30
30
|
`HEAD~1`, which silently drops every commit but the last of a multi-commit task.
|
|
31
31
|
For a scoped re-review, `BASE` is the head the previous review saw. For the final
|
|
32
|
-
review, `BASE` is `git merge-base
|
|
32
|
+
review, `BASE` is `git merge-base "$BASE_BRANCH" HEAD` — the base branch recorded
|
|
33
|
+
in the stage-0 brief, which is not always `main`.
|
|
33
34
|
|
|
34
35
|
Never dispatch a reviewer without a diff file.
|
|
35
36
|
|