task-pipeline-skill 0.17.1 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (24) hide show
  1. package/CHANGELOG.md +244 -0
  2. package/README.md +314 -135
  3. package/cursor/rules/task-pipeline.mdc +91 -16
  4. package/package.json +7 -3
  5. package/plugins/task-pipeline/.claude-plugin/plugin.json +2 -2
  6. package/plugins/task-pipeline/commands/task-pipeline.md +16 -6
  7. package/plugins/task-pipeline/skills/task-pipeline/SKILL.md +39 -8
  8. package/plugins/task-pipeline/skills/task-pipeline/pipeline.example.json +9 -4
  9. package/plugins/task-pipeline/skills/task-pipeline/references/acceptance.md +40 -8
  10. package/plugins/task-pipeline/skills/task-pipeline/references/artifacts.md +23 -11
  11. package/plugins/task-pipeline/skills/task-pipeline/references/audit.md +224 -0
  12. package/plugins/task-pipeline/skills/task-pipeline/references/brainstorm.md +6 -4
  13. package/plugins/task-pipeline/skills/task-pipeline/references/build.md +8 -1
  14. package/plugins/task-pipeline/skills/task-pipeline/references/companion-skills.md +12 -2
  15. package/plugins/task-pipeline/skills/task-pipeline/references/conventions.md +17 -3
  16. package/plugins/task-pipeline/skills/task-pipeline/references/grill.md +37 -4
  17. package/plugins/task-pipeline/skills/task-pipeline/references/knowledge-sources.md +159 -0
  18. package/plugins/task-pipeline/skills/task-pipeline/references/loop-guard.md +8 -0
  19. package/plugins/task-pipeline/skills/task-pipeline/references/planning.md +5 -3
  20. package/plugins/task-pipeline/skills/task-pipeline/references/review.md +2 -1
  21. package/plugins/task-pipeline/skills/task-pipeline/references/stages.md +73 -11
  22. package/plugins/task-pipeline/skills/task-pipeline/references/tdd.md +5 -0
  23. package/plugins/task-pipeline/skills/task-pipeline/templates/brief.md +24 -1
  24. package/plugins/task-pipeline/skills/task-pipeline/templates/carryover.md +23 -0
@@ -0,0 +1,224 @@
1
+ # Audit — finding what is missing, cross-cutting
2
+
3
+ Every gate in this pipeline asks *"is this artifact good?"*. Stage 10 asks *"is
4
+ anything from the list lost?"* ([`acceptance.md`](acceptance.md)). **Neither asks
5
+ what should have been on the list and never was.**
6
+
7
+ That gap is not an oversight in the gates. It is structural: a gate compares two
8
+ things, and **a contradiction has two sides while an absence has one.** Comparing
9
+ the spec against the plan finds a requirement that was dropped. It cannot find the
10
+ error path nobody specified, the entity nobody gave an owner, the failure mode
11
+ nobody named — because on both sides of every comparison, it simply isn't there.
12
+
13
+ This file is the method that finds those. It is **cross-cutting**: stage 10 runs it
14
+ before writing the coverage table, the program loop runs it per module, and a task
15
+ whose whole job is "audit X" runs nothing else.
16
+
17
+ ## Three things that are easy to confuse
18
+
19
+ | File | Runs when | Answers |
20
+ |---|---|---|
21
+ | [`acceptance.md`](acceptance.md) | stage 10 | did everything **on the list** ship, with evidence? |
22
+ | [`loop-guard.md`](loop-guard.md) | any **editing** loop churns | is this pass undoing the last one? |
23
+ | **this file** | any **audit** pass | what is broken or missing that nobody has compared? |
24
+
25
+ `loop-guard.md` governs loops that *change* things — the fix loop, a re-entered
26
+ stage. Its trip means a decision is being re-litigated at the wrong altitude. This
27
+ file governs loops that *look* for things. Its trip means the axis is exhausted,
28
+ which is a different failure with a different exit. Both can bind one run; they do
29
+ not overlap.
30
+
31
+ ## Why "look again, more carefully" stops working
32
+
33
+ The method most audits use is **horizontal**: compare the documents against each
34
+ other, then do it again. It works, and then it fails in a way that is invisible
35
+ from inside it. Measured over seven passes on a production repository:
36
+
37
+ | Pass | Findings | …of which the previous pass's own fixes caused |
38
+ |---|---|---|
39
+ | 4 | 12 | 5 |
40
+ | 5 | 17 | 9 |
41
+ | 6 | 13 | 10 |
42
+ | 7 | 19 | 4 |
43
+
44
+ By pass six the audit was **mostly repairing itself**. Not fatigue — arithmetic.
45
+ Each pass edits the corpus the next pass reads, so the newest edits are always the
46
+ least-reviewed text present, and they are what the next pass finds. The count stays
47
+ healthy while the yield goes to zero.
48
+
49
+ A single **vertical** pass over the same repository — one capability walked down
50
+ through its layers — found nine defect classes those seven passes had been
51
+ structurally **unable** to see. Not missed: unable. Two of them:
52
+
53
+ - **A component that encrypts every object in the system had no key store
54
+ anywhere.** Its sibling's key table had been modelled for weeks. A comparison
55
+ needs two sides; this had one.
56
+ - **An edit appended a new version while the derived index still pointed at the
57
+ old text.** The archive is correct. The index is correct. The answer is correct —
58
+ *against a version nobody is looking at.* Compare archive to archive and index to
59
+ index: both pass. **The defect lives in the seam.**
60
+
61
+ ## The ladder
62
+
63
+ The rungs are **layers of one deliverable**; the work is the **seams between
64
+ them**. Each rung's artefact either exists or it does not — that is what makes
65
+ absence findable.
66
+
67
+ | Rung | Layer | The artefact that must exist |
68
+ |---|---|---|
69
+ | **L0** | Requirement | a `REQ-###` row in the brief **with a named check** |
70
+ | **L1** | Decision | the locked decision, ADR or `CONTEXT.md` term this REQ rests on |
71
+ | **L2** | Design | a spec section carrying `covers: REQ-…` |
72
+ | **L3** | Contract | an exact signature or schema · **and its failure behavior** |
73
+ | **L4** | Task | a plan task with `Implements:` and a DoD satisfiable **as written** |
74
+ | **L5** | Change | the commits — the thing actually in the tree |
75
+ | **L6** | Test | an **executed** assertion, by name — never "the tests pass" |
76
+ | **L7** | Surface | what a user reaches: scenario, screen state, CLI output, runbook |
77
+
78
+ **Audit the seams, not the artifacts.** Each rung is internally consistent most of
79
+ the time — that is exactly what the horizontal pass is good at, and it has already
80
+ done it. What survives lives between rungs:
81
+
82
+ | Seam | The question | What absence looks like here |
83
+ |---|---|---|
84
+ | L0→L1 | does the requirement rest on a **recorded** decision? | a REQ whose check implies a choice nobody ever made or wrote down |
85
+ | L1→L2 | did the decision reach the spec? | an ADR or glossary term agreed at the grill that no spec section cites |
86
+ | L2→L3 | does the section name its contract **and what happens when it fails**? | "handles errors" — no code, no shape, no caller-visible reason |
87
+ | L3→L4 | does every contract have a task that builds it? | stage 4's set-equality covers REQ→task; **nothing** covers contract→task |
88
+ | L4→L5 | did the DoD land in the tree? | a DoD line nothing in the diff satisfies, marked done anyway |
89
+ | L5→L6 | is there an executed observable? | "tests pass"; a test that still passes with the production code deleted |
90
+ | L6→L7 | can a user reach it, and does a doc say so? | shipped behavior with no scenario, no `--help` line, no runbook entry |
91
+ | L7→L0 | does the shipped surface satisfy the requirement's **statement**? | it does what the task said and not what the requirement meant |
92
+
93
+ The last seam is stage 10's question, expressed as a seam. When it fails, the run
94
+ did every instruction correctly and delivered the wrong thing.
95
+
96
+ ## How one audit pass runs
97
+
98
+ **Scope: one deliverable, all rungs.** One REQ, one module, one capability. Not
99
+ "audit the docs" and not "audit the change" — an unscoped instruction is what
100
+ produces seven converging passes.
101
+
102
+ **Input** is the artifact that already names every rung: the brief's REQ row plus
103
+ the module map row ([`decomposition.md`](decomposition.md)) when there is one. **If
104
+ the input can't supply the rungs, that is the first finding** — do not go looking
105
+ for the layers by hand; record that the spine is missing and fix that first.
106
+
107
+ **Procedure: bottom-up, L0 → L7, running the seam check at each step.**
108
+
109
+ The direction is not taste. A missing artefact at L1 makes everything above it
110
+ meaningless, so top-down you spend the pass polishing a surface for a contract that
111
+ does not exist. Bottom-up, the absence surfaces first and the six findings above it
112
+ collapse into one.
113
+
114
+ **Output: findings ordered by seam, never by file.** A file-ordered list reads as
115
+ noise; a seam-ordered one tells you **which layer of your own process is leaking**,
116
+ which is the thing worth knowing. Each finding carries `file:line`, the artefact
117
+ that is missing, and the minimal fix.
118
+
119
+ **Close through the pipeline, not around it.** A finding that is a genuine gap
120
+ becomes a **new REQ row** (with its check) or a carry-over row — the list is frozen
121
+ against *narrowing*, never against additions ([`grill.md`](grill.md) → *The REQ
122
+ spine*). A finding that contradicts the spec goes back to stage 3; one that
123
+ contradicts the plan goes back to stage 4. Auditing is not a licence to edit
124
+ across layers in place.
125
+
126
+ ## Exit criterion — the part usually skipped
127
+
128
+ A deliverable is **not** audited when somebody has read it. It is audited when:
129
+
130
+ 1. every rung has its artefact, **and**
131
+ 2. **every check you are relying on has fired at least once against a planted
132
+ defect.**
133
+
134
+ **A green result from an unproven check is worth nothing.** This is the iron law of
135
+ [`tdd.md`](tdd.md) — *if you didn't watch it fail, you don't know it tests the
136
+ right thing* — raised from one test to every gate in the run. It applies to the
137
+ stage-4 set-equality check, the host's lint and test commands, the super-ux linter,
138
+ any script the host added, and every check you write during the audit itself.
139
+
140
+ Checks written under time pressure lie in ways that read as success: a predicate
141
+ that inspects the wrong shape and finds nothing; a probe that removes more than it
142
+ adds and reads the shrinkage as a pass; a regex that misses the very word it
143
+ searches for. All three pass loudly. **Plant the defect. Watch the check fail.
144
+ Remove it. Then trust the green.** Record in the ledger that you did.
145
+
146
+ ## The three rules that stop this becoming another loop
147
+
148
+ ### 1. A class that repeats twice becomes a gate, not a note
149
+
150
+ Once is an incident. **Twice is a category, and a category belongs in a script** —
151
+ the host's lint, its CI, its check runner — where nobody has to remember it.
152
+
153
+ Writing the third instance into the carry-over ledger is how a known, mechanical
154
+ defect class becomes permanent. If the class genuinely cannot be checked
155
+ mechanically, say so in one line and *say why*; that sentence is itself a finding
156
+ worth having.
157
+
158
+ ### 2. Every pass changes the axis, not the effort
159
+
160
+ "Look again, more carefully" is what converges. Passes must be **orthogonal by
161
+ construction**:
162
+
163
+ 1. **Seams** — one deliverable walked L0→L7 (this file's ladder).
164
+ 2. **Invariants across deliverables** — one name, one enum, one owner, one spelling,
165
+ everywhere. This is the horizontal pass, and it is where it belongs.
166
+ 3. **One class swept end to end** — every error path, every count, every status
167
+ vocabulary, every timeout, across the whole change at once.
168
+
169
+ **The crossover is measurable, so measure it.** Every pass, count two numbers: new
170
+ findings, and findings caused by the previous pass's own fixes. When the second
171
+ overtakes the first, the axis is exhausted — **rotate it, don't push harder.** Both
172
+ counts go in the ledger; an audit that reports only "found N" cannot see its own
173
+ exhaustion.
174
+
175
+ ### 3. What can't be fixed now becomes a ratchet, never a TODO
176
+
177
+ A **ratchet** is a *named, counted set that may only shrink, printed on every
178
+ run*.
179
+
180
+ The carry-over ledger ([`templates/carryover.md`](../templates/carryover.md)) is
181
+ the pipeline's ratchet, and it only works if its count is **printed at every gate
182
+ beside the verdict**:
183
+
184
+ ```
185
+ GATE 6 tests: PASS — full suite green (247 tests)
186
+ carry-over: 4 open (was 6) · unresolved: 0 · audit findings deferred: 2
187
+ ```
188
+
189
+ The difference from a TODO is not bookkeeping. A TODO is invisible until somebody
190
+ opens the file. A ratchet sits next to the word `PASS` on every single run, so
191
+ **"green" never reads as "verified"** — it reads as *"green, and here is exactly
192
+ what was not looked at."* A ratchet that grew needs a sentence in the run log
193
+ explaining why; a ratchet nobody prints is a TODO with a better name.
194
+
195
+ ## When this runs
196
+
197
+ - **Stage 10, before the coverage table.** Acceptance reads the REQ list; the
198
+ ladder walk is what can add to it. Absences found here become new REQ rows with
199
+ their checks, and *then* the table is written — otherwise acceptance closes green
200
+ over a gap that was never a row.
201
+ - **Per module in the program loop** ([`decomposition.md`](decomposition.md)) — one
202
+ brick's ladder, at that brick's acceptance. Cross-module contracts are audited at
203
+ the seam that owns them, not twice.
204
+ - **As the whole task**, when the operator's request *is* an audit. Then stages 3–5
205
+ produce findings and fixes rather than a feature, and the exit criterion above is
206
+ the stage-10 gate.
207
+ - **Never as an eighth "look again" pass.** If the last two passes found mostly
208
+ self-inflicted findings, the answer is rule 2, not another pass.
209
+
210
+ Once both axes are exhausted, the next finding of a known class should be caught by
211
+ a script — and if it cannot be, **that is the finding: write the check.**
212
+
213
+ ## Rationalizations
214
+
215
+ | Excuse | Reality |
216
+ |---|---|
217
+ | "The gates all passed, so it's complete" | Gates compare. Nothing that was never written appears on either side of a comparison. |
218
+ | "One more careful pass will catch it" | Measured: by pass six the passes were mostly fixing their own last pass. Rotate the axis. |
219
+ | "I'll audit top-down, the surface is where users are" | A surface built on an absent contract wastes the whole pass. Bottom-up, that absence is finding #1. |
220
+ | "The check is green, that's evidence" | Only if you have seen it red. An unproven check is a decoration that reports success. |
221
+ | "It's a small gap, I'll note it in the ledger" | Second occurrence of a class → it goes in a script. The ledger is for what cannot be automated, not what nobody automated. |
222
+ | "Findings grouped by file are easier to fix" | And impossible to learn from. Group by seam; the seam names which layer of your process leaks. |
223
+ | "The ledger has it, we won't forget" | Only if it is printed beside every verdict. Unprinted, it is a TODO, and TODOs are invisible by construction. |
224
+ | "This is out of scope for the audit" | Then it is a carry-over row with a home, right now. An audit that silently declines findings is worse than none. |
@@ -37,10 +37,12 @@ not a question to re-open from scratch.
37
37
  1. **Explore the current state.** Files, module docs, recent commits, the
38
38
  conventions the repo already follows. Do this before asking anything.
39
39
  2. **Scope check, early.** If the task actually describes several independent
40
- subsystems, say so immediately and help decompose it into sub-projects: what the
41
- independent pieces are, how they relate, what order they get built in. Then
42
- brainstorm the first one. Each sub-project gets its own spec → plan → build
43
- cycle. Don't refine details of something that needs splitting first.
40
+ capabilities or separately shippable surfaces, say so immediately: that is a
41
+ **platform**, and it gets cut into modules at the end of this stage by
42
+ [`decomposition.md`](decomposition.md), before any spec is written. Brainstorm
43
+ the platform's shape — the pieces, how they relate, what order they land in —
44
+ not the details of one corner; those belong to each module's own stage-3
45
+ dossier. Don't refine something that needs splitting first.
44
46
  3. **Questions one at a time.** Never bundle. Multiple choice where it fits, open
45
47
  where it doesn't. Purpose, constraints, success criteria — anything the brief
46
48
  left at design level.
@@ -234,6 +234,12 @@ Two routes leave before the loop starts:
234
234
  - **Minor findings** never enter it. Record each in the ledger
235
235
  (`Task <N>: minor (deferred): <one-liner>`) and point the final review at that
236
236
  list. A roll-up nobody reads is a silent discard.
237
+ - **A finding class that shows up a second time stops being a finding and becomes a
238
+ check.** Two tasks flagged for the same mechanical defect — the same missing
239
+ failure path, the same magic value, the same naming slip — means every later task
240
+ will produce it too. Add it to the host's lint or check script now, in its own
241
+ commit, instead of writing the third instance into the ledger
242
+ ([`audit.md`](audit.md) → *A class that repeats twice becomes a gate*).
237
243
  - **A finding that conflicts with what the plan mandates** is the operator's
238
244
  call: present the finding beside the plan text and ask which governs. Don't
239
245
  dismiss the finding because the plan mandated it; don't fix against the plan
@@ -304,7 +310,8 @@ findings are neither fixed nor parked-with-ruling at the cap.
304
310
  ## 5. Final whole-branch review
305
311
 
306
312
  After the last task: build a package over `MERGE_BASE`..`HEAD`
307
- (`git merge-base main HEAD`), dispatch the whole-branch review
313
+ (`git merge-base "$BASE_BRANCH" HEAD`, where `$BASE_BRANCH` is the base recorded in
314
+ the stage-0 brief — never a hardcoded `main`), dispatch the whole-branch review
308
315
  ([`review.md`](review.md) → *Final review*; on the run's model, escalation offered
309
316
  out loud per *Models* above), and point it at the
310
317
  ledger's deferred-minor and parked lines so it can triage what must be fixed before
@@ -12,6 +12,7 @@ better, plus one that is required only for user-facing work.
12
12
 
13
13
  | Stage | Doctrine |
14
14
  |---|---|
15
+ | 0 Knowledge harvest (pre-grill) | `references/knowledge-sources.md` |
15
16
  | 0 Intake grill | `references/grill.md` |
16
17
  | 2 Brainstorm | `references/brainstorm.md` |
17
18
  | 2 Decompose (platforms only) | `references/decomposition.md` |
@@ -20,6 +21,7 @@ better, plus one that is required only for user-facing work.
20
21
  | 5 Build (isolation, subagents, fix loop) | `references/build.md` + `references/review.md` |
21
22
  | 5–6 TDD + suite gate | `references/tdd.md` |
22
23
  | 10 Acceptance (REQ close-out) | `references/acceptance.md` |
24
+ | 10 + any audit (finding what's missing) | `references/audit.md` |
23
25
  | any repeating loop | `references/loop-guard.md` |
24
26
 
25
27
  ## The matrix
@@ -28,7 +30,7 @@ better, plus one that is required only for user-facing work.
28
30
  |---|---|---|---|
29
31
  | **super-ux** (`ux-foundation`, `ux-flows`, `ux-scenarios`, `ux-audit`, `/ux`, `/ux-lint`) | stage 3 UX track | **Required for any user-facing task** | `/plugin marketplace add ssheleg/super-ux` → `/plugin install super-ux@super-ux` (or `npx skills add ssheleg/super-ux`) |
30
32
  | **context7** (MCP) | stage 1 docs study | Recommended (web-search fallback) | connect the context7 MCP server |
31
- | **wiki-update** | stage 9 wiki sync | Optional (skip wiki if absent) | user's wiki skill set |
33
+ | **[obsidian-wiki](https://github.com/ar9av/obsidian-wiki)** (`wiki-query`, `wiki-update`) | **stage 0 harvest** (query what's already known) **+ stage 9 sync** | **Recommended** — never a gate; absent → harvest runs on repo docs alone | `pip install obsidian-wiki` → `obsidian-wiki setup --vault /path/to/your/vault` |
32
34
  | ~~superpowers~~ | — | **Not a dependency.** Stages 2/4/5/6 run on the built-in doctrine above. See *Optional bridge* | — |
33
35
  | ~~grill-me / grilling~~ | — | **Not a dependency.** The stage-0 grill is built in (`references/grill.md`) | — |
34
36
 
@@ -60,7 +62,11 @@ Pipeline companions (stage doctrine is built in — nothing to install for it):
60
62
  /plugin marketplace add ssheleg/super-ux
61
63
  /plugin install super-ux@super-ux
62
64
  ✓ context7 — ready
63
- ✓ wiki-update — ready
65
+ ✗ obsidian-wiki — recommended: stage 0 queries it before grilling you,
66
+ stage 9 syncs back what this run learned:
67
+ pip install obsidian-wiki
68
+ obsidian-wiki setup --vault /path/to/your/vault
69
+ (running without it — the harvest uses repo docs only)
64
70
 
65
71
  🧠 Model for this run: recommended <top tier available>. You're on <current>.
66
72
  /model <id> to switch, or "keep current", or name per-stage overrides.
@@ -72,6 +78,10 @@ Rules:
72
78
 
73
79
  - Only flag **super-ux** when the task implies a UI (the stage-0 grill decides;
74
80
  when unsure, flag it — a false positive costs one install).
81
+ - **obsidian-wiki**: detect via `~/.obsidian-wiki/config` or a resolving
82
+ `wiki-query`/`wiki-update`. Present → say `✓ ready` and use it in the harvest.
83
+ Absent → print the two install lines **once** and continue; never ask twice in a
84
+ run and never block a stage on it ([`knowledge-sources.md`](knowledge-sources.md)).
75
85
  - **Never gate any stage on an install** except the stage-3 UX track on a UI task.
76
86
  - Optional tools missing → state the fallback, don't block.
77
87
  - Re-detect after the operator installs; don't assume.
@@ -1,7 +1,10 @@
1
- # Host conventions (stages 6–10)
1
+ # Host conventions (stage 0 harvest, stages 6–10)
2
2
 
3
3
  The orchestrator is project-agnostic. For tests / lint / deploy / docs / wiki it reads the
4
4
  **host project's `CLAUDE.md` / `AGENTS.md` first**, then falls back to detection.
5
+ The same files are the stage-0 harvest's first stop — they are where a project
6
+ names its doc repos, its knowledge base and its house rules
7
+ ([`knowledge-sources.md`](knowledge-sources.md)).
5
8
  Prefer explicit host instructions over detection; if a step's convention can't be
6
9
  found, surface it and **ask** rather than guessing.
7
10
 
@@ -30,9 +33,20 @@ found, surface it and **ask** rather than guessing.
30
33
  CI: the workflow run. Hit the health endpoint if one is defined.
31
34
 
32
35
  ## Docs + wiki
36
+ - **Start from the stage-0 source ledger** ([`knowledge-sources.md`](knowledge-sources.md)):
37
+ the sources the harvest read are the sources this stage updates. Anything the run
38
+ proved stale is already listed there with what's wrong.
33
39
  - Host self-update rules (module docs, runbooks, agent-self cards, etc.) — update
34
- in the same change. Wiki: the `wiki-update` skill (resolves the vault via
35
- `~/.obsidian-wiki/config`). Fix dangling links.
40
+ in the same change. Fix dangling links.
41
+ - **Wiki:** [obsidian-wiki](https://github.com/ar9av/obsidian-wiki) — the
42
+ `wiki-update` skill (resolves the vault via `~/.obsidian-wiki/config`). Detect it
43
+ the same way the harvest does; if absent, recommend it once
44
+ (`pip install obsidian-wiki` → `obsidian-wiki setup --vault <path>`) and continue.
45
+ A project may of course use a different knowledge base — then its own
46
+ `CLAUDE.md` names the sync command, and that wins.
47
+ - **Docs in another repository** (a docs repo, a submodule, a sibling checkout the
48
+ project names): updating it is **outward** — propose the change, get an explicit
49
+ operator go, open a PR there. Never push to a repo the task didn't name.
36
50
 
37
51
  ## Issue tracker (stage 10)
38
52
 
@@ -12,7 +12,27 @@ coming back to the operator.
12
12
  > half — glossary challenges, `CONTEXT.md`, ADR discipline — comes from there; the
13
13
  > autonomy sweep and the brief are this pipeline's.
14
14
 
15
- ## The loop
15
+ ## Phase 1 — harvest before you ask
16
+
17
+ **Do not open the interview cold.** Stage 0 begins by finding what the project
18
+ already knows about this task: the code, `CLAUDE.md`, `CONTEXT.md` and the ADRs,
19
+ `docs/` and `docs/ux/`, past pipeline briefs, the **knowledge wiki** when one is
20
+ installed, and any **other repository or hosted doc system the project names as
21
+ its docs**. Full procedure, source order, the wiki's detection and install line,
22
+ and the ledger to write: [`knowledge-sources.md`](knowledge-sources.md).
23
+
24
+ Two things come out of it, both required before question one:
25
+
26
+ - the **source ledger** in the brief — one row per source consulted, what it says
27
+ about this task, and how fresh it is (`no sources found` is a valid row);
28
+ - the list of things you therefore **don't need to ask**, and the specific points
29
+ where a source looks stale or ambiguous — those become the sharpest questions.
30
+
31
+ Everything below runs against that harvest. An answer you can't check against a
32
+ source is a recollection, and the whole loop exists to stop the run from building
33
+ on one.
34
+
35
+ ## Phase 2 — the loop
16
36
 
17
37
  Interview the operator relentlessly about every aspect of the task until you reach
18
38
  a **shared understanding**. Walk down each branch of the decision tree, resolving
@@ -73,6 +93,15 @@ Create these files **lazily** — only when you have something real to write.
73
93
  check whether the code agrees, and surface contradictions: *"Your code cancels
74
94
  entire Orders, but you just said partial cancellation is possible — which is
75
95
  right?"*
96
+ - **Cross-reference with the harvest — every answer, not just the domain ones.**
97
+ Phase 1 put the ADRs, runbooks and wiki pages in your hands; use them the same
98
+ way: *"The March ADR says orders are written only through the command handler,
99
+ you just described a direct write — has that changed?"* The operator **outranks
100
+ every document**, but only out loud: an override quoted against its source is a
101
+ recorded decision, an unquoted one is an undetected divergence. When two sources
102
+ disagree, precedence is code > host docs/ADRs > wiki > memory, and the loser is
103
+ logged for the stage-9 update ([`knowledge-sources.md`](knowledge-sources.md) →
104
+ *Phase 2*).
76
105
  - **Update `CONTEXT.md` inline.** Resolve a term → write it down right then, not in
77
106
  a batch at the end. Format: [`templates/context.md`](../templates/context.md).
78
107
  Keep it free of implementation detail — only terms a domain expert would
@@ -101,6 +130,7 @@ explicit "stop and ask me here":
101
130
  | Stage | What to settle up front |
102
131
  |---|---|
103
132
  | run-wide | the model decision ([`model-tiering.md`](model-tiering.md)); what to decide autonomously vs escalate |
133
+ | 0 Harvest | doc sources beyond this repo — other repos, hosted doc systems, the knowledge wiki — and whether stage 9 may write to them (another repo is outward: propose + PR, never a direct push) |
104
134
  | 1 Docs | external libs/APIs/SDKs in play; any private ones context7 can't resolve → where their docs live |
105
135
  | 2 Decompose | is this a platform (several capabilities/surfaces) or one module? if platform: deploy cadence — per module or once at the end |
106
136
  | 2–3 Spec | UI verdict (arms super-ux); any scenario-tracing waiver |
@@ -153,9 +183,12 @@ that shrank without anyone deciding it should.
153
183
  Everything resolved goes into the **task brief**, seeded from
154
184
  [`templates/brief.md`](../templates/brief.md) and committed to
155
185
  `docs/superpowers/specs/YYYY-MM-DD-<topic>-brief.md` — scope, **the REQ table**,
156
- users, UI verdict, constraints, locked decisions, the autonomy table,
157
- done-criteria, open assumptions. Seed the template only when the file is absent;
158
- never overwrite an existing brief.
186
+ **the phase-1 source ledger**, users, UI verdict, constraints, locked decisions,
187
+ the autonomy table, done-criteria, open assumptions. Seed the template only when
188
+ the file is absent; never overwrite an existing brief.
189
+
190
+ The ledger is not decoration: **stage 9 updates exactly what stage 0 read**, and
191
+ every doc the grill proved stale is already listed there with what's wrong.
159
192
 
160
193
  Alongside it, seed the **carry-over ledger** from
161
194
  [`templates/carryover.md`](../templates/carryover.md) at
@@ -0,0 +1,159 @@
1
+ # Knowledge sources — harvest before the grill, update after the build
2
+
3
+ Stage 0 has two phases. This file is **phase 1**: before the first question is
4
+ asked, find and read what the project already knows about this task. The interview
5
+ ([`grill.md`](grill.md)) is phase 2, and it runs *against* what was harvested here.
6
+
7
+ The same source list closes the loop at **stage 9**: what was read at the start is
8
+ what gets updated at the end. A source good enough to answer a question is a source
9
+ that goes stale when the answer changes.
10
+
11
+ ## Why this is a phase and not "explore a bit first"
12
+
13
+ An agent that starts asking without harvesting spends the operator's turns on
14
+ questions the project already answered — in an ADR, in a runbook, in a wiki page
15
+ written three months ago by the same person now being asked. That is the expensive
16
+ failure, but not the worst one.
17
+
18
+ The worst one is silent: **the operator misremembers, the agent believes them, and
19
+ the run builds on it.** People answer from memory about systems they wrote a year
20
+ ago. Without the documents in hand you cannot tell a decision from a recollection,
21
+ so every later gate passes honestly on a false premise. Harvesting first is what
22
+ makes the grill's answers *checkable* instead of merely confident.
23
+
24
+ ## The sources, in the order to try them
25
+
26
+ | # | Source | How to find it | What it's good for |
27
+ |---|---|---|---|
28
+ | 1 | **The code** | the repo you're in | what actually runs — the tiebreaker |
29
+ | 2 | **Host agent docs** | `CLAUDE.md`, `AGENTS.md`, `.cursor/rules/` | conventions, commands, deploy path, house rules |
30
+ | 3 | **Domain docs** | `CONTEXT.md` / `CONTEXT-MAP.md`, `docs/adr/` | the glossary and the decisions with their reasons |
31
+ | 4 | **Product/UX docs** | `docs/ux/` (super-ux chain), `README`, runbooks | user-facing behavior that is already specified |
32
+ | 5 | **Pipeline history** | `docs/superpowers/specs/`, `plans/`, past `-carryover.md` | what a previous run of this pipeline decided or deferred |
33
+ | 6 | **The knowledge wiki** | see below | distilled cross-project knowledge, prior sessions, why decisions were made |
34
+ | 7 | **Other doc repos the project names** | a docs repo URL or submodule in `CLAUDE.md`/`README`, a sibling checkout, a `docs/` monorepo package | specs, contracts and runbooks that live outside this repo |
35
+ | 8 | **Hosted doc systems the project names** | Notion / Confluence / Google Docs referenced in the project | the same, when the team keeps them there |
36
+
37
+ Rules for the list:
38
+
39
+ - **Never invent a source.** A doc repo is in scope because the project names it,
40
+ not because it plausibly exists. Nothing is cloned or fetched on a guess.
41
+ - **Sources 7–8 are read-only at this stage**, and reading a hosted system needs a
42
+ connected tool — if there's no tool, record the gap and ask the operator to paste
43
+ what matters rather than pretending the source was covered.
44
+ - **The wiki is optional; the harvest is not.** With no wiki and no doc repos, the
45
+ harvest is sources 1–5 and takes two minutes. Skipping it is never the answer.
46
+
47
+ ## The knowledge wiki — recommended
48
+
49
+ The wiki this pipeline is built to work with is
50
+ **[obsidian-wiki](https://github.com/ar9av/obsidian-wiki)** (Karpathy's LLM-wiki
51
+ pattern: raw sources → distilled wiki → schema). It is the one source that carries
52
+ *why* across projects and across months, which is exactly what a fresh context lacks.
53
+
54
+ **Detect it** — any of: `~/.obsidian-wiki/config` exists; the `wiki-query` /
55
+ `wiki-update` skills resolve.
56
+
57
+ - **Installed → use it.** Query it during the harvest (`wiki-query`, or the vault's
58
+ `index.md` + a targeted grep when the skill isn't loaded), and sync back at stage
59
+ 9 (`wiki-update`).
60
+ - **Not installed → recommend it once, in the preflight block, with the line:**
61
+
62
+ ```
63
+ pip install obsidian-wiki
64
+ obsidian-wiki setup --vault /path/to/your/vault
65
+ ```
66
+
67
+ Then continue without it. It is a **recommendation, never a gate** — no stage
68
+ blocks on a missing wiki, and the pipeline never nags twice in a run.
69
+
70
+ ## How to harvest — retrieval, not reading
71
+
72
+ The harvest is bounded by the *task*, not by the size of the sources. You are not
73
+ reading the wiki; you are asking it about this task.
74
+
75
+ 1. **Take the task's nouns** — the entities, the feature name, the subsystem, the
76
+ file paths the operator mentioned — plus their obvious synonyms.
77
+ 2. **Query each source with those terms**: `wiki-query` for the wiki; `grep`/`Read`
78
+ for repo docs; the tracker/hosted-doc tool if one is connected.
79
+ 3. **Follow one hop, not ten.** A hit that names an ADR, a scenario id or a module
80
+ is worth opening. A page three links away is context, not evidence.
81
+ 4. **Stop when the terms stop returning anything new.** Same rule as the interview:
82
+ no grinding past diminishing returns.
83
+
84
+ ## Record it — the source ledger
85
+
86
+ Write what you found into the brief's **Knowledge sources** section
87
+ ([`templates/brief.md`](../templates/brief.md)) before the first question. One row
88
+ per source actually consulted:
89
+
90
+ | Source | What it says about this task | Fresh? | Authority |
91
+ |---|---|---|---|
92
+ | `docs/adr/0007-single-write-model.md` | orders are written only through the command handler | 2026-03 | decision |
93
+ | wiki: `projects/x/concepts/billing-seams` | why invoicing was split out; the retry rule | 2026-06 | context |
94
+ | `CLAUDE.md` | test = `npm test`, deploy from `main` only | current | convention |
95
+ | (none for the export UI) | — | — | — |
96
+
97
+ The ledger is what makes phase 2 work: during the interview you cite rows from it,
98
+ and at stage 9 you update the same rows. A source consulted but not recorded is a
99
+ source nobody will update.
100
+
101
+ **"No sources found" is a valid, recorded outcome.** Write the row. An empty ledger
102
+ tells the next run that the search happened and came back empty — silence doesn't.
103
+
104
+ ## Phase 2 — validate the answers against the harvest
105
+
106
+ This is the payoff, and it belongs to the grill loop
107
+ ([`grill.md`](grill.md) → *Domain awareness*). Every operator answer that touches a
108
+ harvested source gets checked against it, on the spot:
109
+
110
+ > "The ADR from March says orders are written only through the command handler —
111
+ > you just described a direct write. Has that changed, or should the export go
112
+ > through the handler?"
113
+
114
+ Three shapes and what to do with each:
115
+
116
+ | The answer… | Do |
117
+ |---|---|
118
+ | **agrees** with the source | nothing — note it, move on |
119
+ | **contradicts** a source | quote the source, name the conflict, ask which governs. The answer is either "the doc is stale" (→ it gets updated at stage 9, log it now) or "I misremembered" (→ the doc stands). Both are cheap here and expensive at stage 6 |
120
+ | **goes beyond** every source | this is new knowledge — it belongs in the brief, and usually in `CONTEXT.md` or an ADR as it lands |
121
+
122
+ **The operator outranks the docs — but only out loud.** A person may overrule any
123
+ document; they may not do it by accident. The point of quoting the source is that
124
+ the override becomes a recorded decision instead of an undetected divergence.
125
+
126
+ **Precedence when two sources disagree with each other:** code > host docs and
127
+ ADRs > the wiki > anyone's memory. The wiki is *distilled* knowledge and can lag
128
+ the repo by months; the code is what runs. A disagreement between them is a grill
129
+ question, never a silent pick — and it is usually a sign the doc is due an update.
130
+
131
+ ## Close the loop — stage 9 updates what stage 0 read
132
+
133
+ The ledger is the stage-9 work list. For each row:
134
+
135
+ - **Host repo docs, ADRs, runbooks, `docs/ux/`** — updated in the **same change**,
136
+ per the host's own rules ([`conventions.md`](conventions.md)).
137
+ - **Anything the run proved stale** — including a doc that was "wrong but nobody
138
+ had time": that's why the conflict was logged in phase 2 instead of only being
139
+ resolved verbally.
140
+ - **The wiki** — `wiki-update` syncs what this run learned. Distil the *knowledge*
141
+ (decisions, seams, gotchas, why), never a diff summary.
142
+ - **Another repository's docs** — writing to a repo the operator didn't ask you to
143
+ touch is **outward**: propose the change, get an explicit go, then open a PR
144
+ there. Absent a go, it goes in the carry-over ledger with the exact edit needed.
145
+
146
+ A source that was worth reading at stage 0 and is wrong at stage 9 is the next
147
+ run's false premise. Closing that loop is the whole point of harvesting from a
148
+ written list instead of from whatever the search happened to surface.
149
+
150
+ ## Rationalizations
151
+
152
+ | Excuse | Reality |
153
+ |---|---|
154
+ | "I'll just ask them, it's faster" | You'll ask about things a doc already answers, and you'll believe an answer you can't check. Retrieval is cheaper than a turn. |
155
+ | "The wiki's probably stale" | Then say so with the page in hand and get it corrected. "Probably stale" unread is an assumption; read, it's a finding. |
156
+ | "No docs in this repo" | Check `CLAUDE.md` for the repo that has them, and the wiki for the last time anyone touched this. Then record the empty ledger. |
157
+ | "The operator knows their own system" | They do — a year ago, before three other people changed it. That's the exact case where quoting the doc pays. |
158
+ | "Reading the whole wiki costs too much" | The harvest is a query per task noun, not a read. If it feels expensive, you're reading instead of retrieving. |
159
+ | "I'll update the docs at the end from memory" | The ledger exists because the end is exactly when you no longer remember which sources you leaned on. |
@@ -10,6 +10,14 @@ the pipeline: the stage-5 fix loop, a stage re-entered after a failed gate, the
10
10
  per-module program loop ([`decomposition.md`](decomposition.md)), and any
11
11
  audit → fix → audit cycle.
12
12
 
13
+ **This file governs loops that *change* things.** A loop that *looks* for things —
14
+ pass after pass over one corpus — fails differently: it does not oscillate, it
15
+ **converges**, quietly spending each pass on the previous pass's own edits while the
16
+ finding count stays healthy. That has its own detector and its own exit (rotate the
17
+ axis, don't push harder): [`audit.md`](audit.md) → *Every pass changes the axis*.
18
+ Both can bind one run. Use this file's trips for edits, that file's crossover for
19
+ searches.
20
+
13
21
  ## Bookkeeping — the thing that makes detection mechanical
14
22
 
15
23
  You cannot detect churn from memory, especially after compaction. Every repeating
@@ -21,9 +21,11 @@ brief and the spec.
21
21
 
22
22
  ## Before writing tasks
23
23
 
24
- **Scope check.** If the spec covers several independent subsystems, split it into
25
- one plan per subsystem; each plan must produce working, testable software on its
26
- own.
24
+ **Scope check.** A plan covers exactly **one** spec — for a decomposed platform,
25
+ one module's dossier ([`decomposition.md`](decomposition.md)). If the spec in front
26
+ of you covers several independent subsystems, the decomposition was missed at stage
27
+ 2: say so and go back there for a module map, rather than inventing the split here.
28
+ Whatever this plan covers must produce working, testable software on its own.
27
29
 
28
30
  **Map the file structure.** List every file that will be created or modified and
29
31
  what each one owns. This is where decomposition gets locked in:
@@ -29,7 +29,8 @@ this plan's git-ignored directory, `.task-pipeline/build/<plan-basename>/` — s
29
29
  `BASE` is the commit you recorded **before** dispatching the implementer — never
30
30
  `HEAD~1`, which silently drops every commit but the last of a multi-commit task.
31
31
  For a scoped re-review, `BASE` is the head the previous review saw. For the final
32
- review, `BASE` is `git merge-base main HEAD`.
32
+ review, `BASE` is `git merge-base "$BASE_BRANCH" HEAD` — the base branch recorded
33
+ in the stage-0 brief, which is not always `main`.
33
34
 
34
35
  Never dispatch a reviewer without a diff file.
35
36