@hecer/yoke 1.6.0 → 1.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (103) hide show
  1. package/.claude-plugin/plugin.json +13 -13
  2. package/.codex-plugin/plugin.json +7 -7
  3. package/CHANGELOG.md +294 -288
  4. package/README.md +874 -874
  5. package/TODOS.md +5 -5
  6. package/agents/docs.toml +6 -6
  7. package/agents/implementer.toml +6 -6
  8. package/agents/reviewer.toml +6 -6
  9. package/agents/security.toml +6 -6
  10. package/bench/README.md +86 -86
  11. package/bench/RESULTS.md +35 -35
  12. package/bench/output-compaction.mjs +65 -65
  13. package/bench/result-schema.mjs +12 -12
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -15
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
  17. package/bench/run-matrix.mjs +26 -26
  18. package/bench/run.mjs +106 -106
  19. package/canon/AGENTS.md +30 -30
  20. package/canon/context/DECISIONS.md +4 -4
  21. package/canon/context/GLOSSARY.md +11 -11
  22. package/canon/context/KNOWLEDGE.md +4 -4
  23. package/canon/context/PROJECT.md +15 -15
  24. package/canon/loop/loop-spec.md +65 -65
  25. package/canon/loop/prd.schema.md +43 -43
  26. package/canon/manifest.yaml +59 -59
  27. package/canon/policy/gates.md +7 -7
  28. package/canon/policy/roles.md +9 -9
  29. package/canon/skills/ATTRIBUTION.md +99 -99
  30. package/canon/skills/authoring-prd/SKILL.md +58 -58
  31. package/canon/skills/brainstorming/SKILL.md +164 -164
  32. package/canon/skills/codebase-design/DEEPENING.md +15 -15
  33. package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
  34. package/canon/skills/codebase-design/SKILL.md +39 -39
  35. package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
  36. package/canon/skills/document-release/SKILL.md +302 -302
  37. package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
  38. package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
  39. package/canon/skills/domain-modeling/SKILL.md +35 -35
  40. package/canon/skills/executing-plans/SKILL.md +70 -70
  41. package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
  42. package/canon/skills/health/SKILL.md +177 -177
  43. package/canon/skills/maintaining-context/SKILL.md +34 -34
  44. package/canon/skills/minimal-code/SKILL.md +21 -21
  45. package/canon/skills/no-ai-slop/SKILL.md +103 -103
  46. package/canon/skills/no-ai-slop/eval.md +43 -43
  47. package/canon/skills/plan-ceo-review/SKILL.md +541 -541
  48. package/canon/skills/plan-eng-review/SKILL.md +362 -362
  49. package/canon/skills/receiving-code-review/SKILL.md +213 -213
  50. package/canon/skills/requesting-code-review/SKILL.md +105 -105
  51. package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
  52. package/canon/skills/retro/SKILL.md +397 -397
  53. package/canon/skills/review/SKILL.md +246 -246
  54. package/canon/skills/ship/SKILL.md +691 -691
  55. package/canon/skills/subagent-driven-development/SKILL.md +277 -277
  56. package/canon/skills/systematic-debugging/SKILL.md +296 -296
  57. package/canon/skills/tdd/SKILL.md +371 -371
  58. package/canon/skills/unslop-ui/SKILL.md +34 -34
  59. package/canon/skills/using-git-worktrees/SKILL.md +218 -218
  60. package/canon/skills/verification-before-completion/SKILL.md +139 -139
  61. package/canon/skills/visual-verification/SKILL.md +54 -54
  62. package/canon/skills/workflow/SKILL.md +22 -22
  63. package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
  64. package/canon/skills/writing-for-agents/SKILL.md +42 -42
  65. package/canon/skills/writing-plans/SKILL.md +152 -152
  66. package/canon/skills/writing-skills/SKILL.md +655 -655
  67. package/canon/skills/yoke-retrofit/SKILL.md +26 -26
  68. package/canon/skills/yoke-workflow/SKILL.md +20 -20
  69. package/canon/tools/codex-rtk-hook.mjs +35 -35
  70. package/canon/tools/graphify.md +3 -3
  71. package/canon/tools/playwright-mcp.md +3 -3
  72. package/canon/tools/rtk.md +7 -7
  73. package/canon/tools/serena.md +6 -6
  74. package/dist/agents/process.js +3 -0
  75. package/dist/loop/watchdog.js +1 -1
  76. package/dist/prd/command.js +17 -17
  77. package/dist/retrofit/planners/claude.js +14 -14
  78. package/dist/retrofit/preserve.js +2 -2
  79. package/docs/MIGRATING-TO-1.0.md +33 -33
  80. package/docs/MIGRATING-TO-1.1.md +27 -27
  81. package/docs/MIGRATING-TO-1.4.md +70 -70
  82. package/docs/PUBLISHING.md +91 -91
  83. package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
  84. package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
  85. package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
  86. package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
  87. package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
  88. package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
  89. package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
  90. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
  91. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
  92. package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
  93. package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
  94. package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
  95. package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
  96. package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
  97. package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
  98. package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
  99. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
  100. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
  101. package/gemini-extension.json +6 -6
  102. package/hooks/hooks.json +19 -19
  103. package/package.json +87 -87
@@ -1,537 +1,537 @@
1
- # Gauntlet Quality Loop Implementation Plan
2
-
3
- > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
-
5
- **Goal:** Add bounded and explicitly unbounded comparative-quality iteration, structured repair, external reference judging, production parallel workers, candidate races, ephemeral decomposition, and versioned provider contracts to Yoke without weakening its mechanical gates.
6
-
7
- **Architecture:** Extend the existing `yoke loop` state machine rather than adding a second loop. New focused `src/quality/`, `src/agents/contracts.ts`, and loop-worker modules provide pure contracts and adapters; `runLoop` retains story/commit authority. Deliver the work in independently green phases so the unchanged serial, no-quality path remains releasable throughout.
8
-
9
- **Tech Stack:** TypeScript ESM, Node.js ≥20 standard library, Zod 3, YAML, Vitest 4, Git worktrees, existing Claude/Codex/Gemini CLI adapters, project-local Playwright when visual artifacts are used.
10
-
11
- **Source design:** `docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md`
12
-
13
- ---
14
-
15
- ## File structure
16
-
17
- ### New production modules
18
-
19
- - `src/agents/contracts.ts` — versioned machine envelopes shared by routing, review, quality, decomposition, and telemetry.
20
- - `src/quality/types.ts` — quality policy, reference, candidate, verdict, budget, and outcome types.
21
- - `src/quality/reference.ts` — safe reference validation, acquisition, hashing, and provenance.
22
- - `src/quality/artifacts.ts` — candidate artifact collection and deterministic digesting.
23
- - `src/quality/verdict.ts` — quality verdict file contract and label-swap consistency reduction.
24
- - `src/quality/runner.ts` — read-only critic invocation and evidence persistence.
25
- - `src/quality/repair.ts` — selected-gap extraction and repair prompt/runner construction.
26
- - `src/quality/loop.ts` — bounded/unbounded quality-round controller; no story or commit ownership.
27
- - `src/loop/worker.ts` — execute one isolated story candidate through implementation and gates.
28
- - `src/loop/dispatcher.ts` — provider subprocess workers, claims, cancellation, and merge-queue integration.
29
- - `src/loop/candidates.ts` — same-story candidate fan-out and winner-only selection.
30
- - `src/loop/decomposition.ts` — ephemeral subtask schema, planning, scheduling, and synthesis.
31
-
32
- ### Existing modules to modify
33
-
34
- - `src/review/verdict.ts` — schema version, provenance, and optional actionable finding metadata.
35
- - `src/retrofit/config.ts` — quality defaults and decomposition configuration.
36
- - `src/loop/prd.ts` — optional story quality declaration.
37
- - `src/loop/runner.ts` — repair/decomposition/quality invocations through provider adapters.
38
- - `src/loop/loop.ts` — invoke quality/repair controller while preserving gate and commit order.
39
- - `src/loop/run-command.ts` — resolve quality policy, parallel dispatcher, candidates, and CLI options.
40
- - `src/loop/reporter.ts` — quality/parallel phases and structured status fields.
41
- - `src/loop/claims.ts` — dispatcher/base/worktree/provider/heartbeat metadata.
42
- - `src/loop/parallel.ts`, `src/loop/scheduler.ts`, `src/loop/merge-queue.ts` — production worker and integration semantics.
43
- - `src/routing/router.ts` — consume versioned route contracts and disclose fallback reasons.
44
- - `src/agents/telemetry.ts` — versioned telemetry envelope plus raw evidence retention.
45
- - `src/cli.ts` — additive CLI flags and conflict validation.
46
- - `src/retrofit/planners/shared.ts` and generated templates — runtime ignore paths and config comments.
47
- - `canon/loop/prd.schema.md`, `canon/loop/loop-spec.md`, `canon/skills/authoring-prd/SKILL.md` — user-facing contracts.
48
- - `README.md`, `CHANGELOG.md`, `TODOS.md`, migration docs — truthful shipped behavior.
49
-
50
- ### New test suites
51
-
52
- - `tests/agents/contracts.test.ts`
53
- - `tests/quality/{reference,artifacts,verdict,repair,loop}.test.ts`
54
- - `tests/loop/{worker,dispatcher,candidates,decomposition}.test.ts`
55
- - `tests/loop/gauntlet-cli.integration.test.ts`
56
- - `tests/fixtures/fake-agent.mjs` additions for deterministic provider outputs and process behavior.
57
-
58
- ---
59
-
60
- ## Phase 1: Versioned contracts and backward-compatible schemas
61
-
62
- ### Task 1: Shared machine-result envelopes
63
-
64
- **Files:**
65
- - Create: `src/agents/contracts.ts`
66
- - Create: `tests/agents/contracts.test.ts`
67
-
68
- - [ ] **Step 1: Write failing schema tests** covering a valid envelope, rejection of unknown schema versions, required role/provider/timing fields, optional reported model, usage, and retained raw metadata.
69
- - [ ] **Step 2: Run** `npx vitest run tests/agents/contracts.test.ts` and confirm failure because the module does not exist.
70
- - [ ] **Step 3: Implement** `MachineEnvelopeSchema`, `MachineRoleSchema`, `MachineUsageSchema`, and inferred types. Version 1 accepts roles `route`, `review`, `quality`, `decomposition`, `candidate-selection`, `telemetry`; timing is nonnegative integer milliseconds; `raw` is `z.record(z.unknown()).optional()`.
71
- - [ ] **Step 4: Run** `npx vitest run tests/agents/contracts.test.ts` and confirm pass.
72
- - [ ] **Step 5: Commit** `feat(contracts): add versioned provider result envelopes`.
73
-
74
- ### Task 2: Review verdict provenance and actionable gaps
75
-
76
- **Files:**
77
- - Modify: `src/review/verdict.ts`
78
- - Modify: `src/loop/runner.ts`
79
- - Modify: `tests/review/verdict.test.ts`
80
- - Modify: `tests/loop/runner.test.ts`
81
-
82
- - [ ] **Step 1: Add failing tests** proving legacy `{approved,summary,findings}` remains valid and versioned verdicts accept `schemaVersion:1`, envelope provenance, finding `id`, `actionable`, and `evidence` fields.
83
- - [ ] **Step 2: Add a failing test** for `selectRepairFinding(verdict)` choosing the first actionable blocking finding deterministically by original order and returning `null` for malformed/infrastructure-only rejection.
84
- - [ ] **Step 3: Run** `npx vitest run tests/review/verdict.test.ts tests/loop/runner.test.ts` and confirm the new expectations fail.
85
- - [ ] **Step 4: Extend schemas minimally** with optional fields and export `selectRepairFinding`. Update `formatReviewContract` to request the versioned shape while `readReviewVerdict` still parses legacy files.
86
- - [ ] **Step 5: Run** the two focused suites and confirm pass.
87
- - [ ] **Step 6: Commit** `feat(review): add actionable versioned verdicts`.
88
-
89
- ### Task 3: Quality configuration and PRD declarations
90
-
91
- **Files:**
92
- - Create: `src/quality/types.ts`
93
- - Modify: `src/retrofit/config.ts`
94
- - Modify: `src/loop/prd.ts`
95
- - Modify: `tests/retrofit/config.test.ts`
96
- - Modify: `tests/loop/prd.test.ts`
97
-
98
- - [ ] **Step 1: Write failing config tests** for safe defaults (`enabled:false`, blocking, three rounds, 60 minutes, two consistency checks, two max candidates), invalid nonpositive limits, and backward-compatible configs without `quality`.
99
- - [ ] **Step 2: Write failing PRD tests** for URL/file/command references, candidate kinds, required rubric, path traversal rejection, and old stories without `quality`.
100
- - [ ] **Step 3: Run** `npx vitest run tests/retrofit/config.test.ts tests/loop/prd.test.ts` and confirm failures.
101
- - [ ] **Step 4: Implement focused Zod schemas** in `quality/types.ts`; import them into config and PRD schemas rather than duplicating declarations.
102
- - [ ] **Step 5: Run focused tests**, then `npm run lint`.
103
- - [ ] **Step 6: Commit** `feat(quality): add config and story contracts`.
104
-
105
- **Phase 1 gate:** `npm run lint && npx vitest run tests/agents tests/review tests/retrofit/config.test.ts tests/loop/prd.test.ts`.
106
-
107
- ---
108
-
109
- ## Phase 2: Bounded reviewer-to-repair loop
110
-
111
- ### Task 4: Repair-gap prompt and runner
112
-
113
- **Files:**
114
- - Create: `src/quality/repair.ts`
115
- - Create: `tests/quality/repair.test.ts`
116
- - Modify: `src/loop/runner.ts`
117
-
118
- - [ ] **Step 1: Write failing tests** for a prompt containing story, acceptance, current diff instruction, exactly one gap, evidence, project context, and gate names while excluding praise, prior deliberation, and unrelated findings.
119
- - [ ] **Step 2: Write a failing invocation test** proving repair uses a fresh provider call with workspace-write permissions and the existing watchdog.
120
- - [ ] **Step 3: Run** `npx vitest run tests/quality/repair.test.ts` and confirm failure.
121
- - [ ] **Step 4: Implement** `RepairGap`, `buildRepairPrompt`, and `makeRepairRunner` by reusing provider invocation/watchdog seams from `runner.ts`.
122
- - [ ] **Step 5: Run focused tests** and confirm pass.
123
- - [ ] **Step 6: Commit** `feat(quality): add fresh single-gap repair runner`.
124
-
125
- ### Task 5: Quality budget controller
126
-
127
- **Files:**
128
- - Create: `src/quality/loop.ts`
129
- - Create: `tests/quality/loop.test.ts`
130
-
131
- - [ ] **Step 1: Write failing pure-controller tests** for immediate pass, repair then pass, three-round exhaustion, 60-minute exhaustion using injected clock, unbounded continuation beyond both limits, pause between rounds, and infrastructure failure blocking without repair.
132
- - [ ] **Step 2: Run** `npx vitest run tests/quality/loop.test.ts` and confirm failure.
133
- - [ ] **Step 3: Implement** an injected `runQualityRounds(options)` controller returning `passed`, `blocked`, or `paused`, last gap, rounds, elapsed time, and evidence. It calls injected `gateCandidate`, `judgeCandidate`, and `repairCandidate`; it owns no Git or PRD operations.
134
- - [ ] **Step 4: Run focused tests** and confirm pass.
135
- - [ ] **Step 5: Commit** `feat(quality): add bounded and unbounded repair controller`.
136
-
137
- ### Task 6: Integrate repair rounds into serial story execution
138
-
139
- **Files:**
140
- - Modify: `src/loop/loop.ts`
141
- - Modify: `src/loop/reporter.ts`
142
- - Modify: `tests/loop/loop.test.ts`
143
- - Modify: `tests/loop/reporter.test.ts`
144
-
145
- - [ ] **Step 1: Add failing loop tests** proving a review rejection with an actionable finding repairs in the same isolated worktree, reruns criterion/verify/perf/audit, invokes a fresh reviewer, and commits only after pass.
146
- - [ ] **Step 2: Add failing tests** proving mechanical failure never invokes repair, malformed review blocks, exhaustion blocks with last gap, pause exits at the next round boundary, and existing no-quality behavior invokes each legacy gate exactly once.
147
- - [ ] **Step 3: Add reporter tests** for phases `repairing` and quality round/budget fields in JSON status.
148
- - [ ] **Step 4: Run** `npx vitest run tests/loop/loop.test.ts tests/loop/reporter.test.ts` and confirm failures.
149
- - [ ] **Step 5: Add optional quality/repair dependencies to `LoopOptions`** and thread the pure controller into both isolated and non-isolated candidate paths without moving commit authority.
150
- - [ ] **Step 6: Run focused tests**, then `npm test` to detect serial regressions.
151
- - [ ] **Step 7: Commit** `feat(loop): repair actionable review failures behind gates`.
152
-
153
- **Phase 2 gate:** full build and tests; a deterministic fake-review integration must demonstrate reject → repair → reverify → approve → one commit.
154
-
155
- ---
156
-
157
- ## Phase 3: External references and blind comparison
158
-
159
- ### Task 7: Safe reference acquisition
160
-
161
- **Files:**
162
- - Create: `src/quality/reference.ts`
163
- - Create: `tests/quality/reference.test.ts`
164
-
165
- - [ ] **Step 1: Write failing tests** for file acquisition, SHA-256 provenance, URL redirect/size/content-type limits through an injected fetcher, private-network rejection, explicit local-project URL allowance, command validation, digest mismatch, and safe storage below `.yoke/references/<digest>`.
166
- - [ ] **Step 2: Run** `npx vitest run tests/quality/reference.test.ts` and confirm failure.
167
- - [ ] **Step 3: Implement** pure validation and injected acquisition adapters using Node `crypto`, `fs`, `path`, `URL`, and `fetch`; no new dependency. Store inert bytes plus `provenance.json`. Never return fetched text as prompt instructions.
168
- - [ ] **Step 4: Run focused tests** and confirm pass.
169
- - [ ] **Step 5: Commit** `feat(quality): acquire pinned untrusted references safely`.
170
-
171
- ### Task 8: Candidate artifact collection
172
-
173
- **Files:**
174
- - Create: `src/quality/artifacts.ts`
175
- - Create: `tests/quality/artifacts.test.ts`
176
- - Modify: `src/smoke/command.ts`
177
- - Modify: `tests/smoke/command.test.ts`
178
-
179
- - [ ] **Step 1: Write failing tests** for declared screenshots/files, stable ordered digesting, missing artifacts, traversal rejection, command-output/benchmark capture, and reuse of existing flow-smoke proof files without recapture.
180
- - [ ] **Step 2: Run** the quality artifact and smoke tests and confirm failure.
181
- - [ ] **Step 3: Implement** collection under the story proof root and expose smoke proof metadata sufficient for reuse.
182
- - [ ] **Step 4: Run focused tests** and confirm pass.
183
- - [ ] **Step 5: Commit** `feat(quality): collect comparable candidate artifacts`.
184
-
185
- ### Task 9: Blind label-swap verdict contract
186
-
187
- **Files:**
188
- - Create: `src/quality/verdict.ts`
189
- - Create: `tests/quality/verdict.test.ts`
190
-
191
- - [ ] **Step 1: Write failing tests** for schema validation, random A/B assignment through injected RNG, mandatory swapped second comparison, candidate/candidate consistency, low-confidence rejection, missing evidence, digest mismatch, and reduction to `pass`, `lose`, or `inconsistent`.
192
- - [ ] **Step 2: Run** `npx vitest run tests/quality/verdict.test.ts` and confirm failure.
193
- - [ ] **Step 3: Implement** `QualityVerdictSchema`, `assignBlindLabels`, and `reduceComparisons`. Do not add numeric quality scores.
194
- - [ ] **Step 4: Run focused tests** and confirm pass.
195
- - [ ] **Step 5: Commit** `feat(quality): enforce blind binary judge consistency`.
196
-
197
- ### Task 10: Read-only critic and evidence persistence
198
-
199
- **Files:**
200
- - Create: `src/quality/runner.ts`
201
- - Create: `tests/quality/runner.test.ts`
202
- - Modify: `src/loop/runner.ts`
203
-
204
- - [ ] **Step 1: Write failing tests** proving the critic receives trusted rubric separately from untrusted artifacts, uses read-only permissions, writes a result file, emits two fresh calls for blocking policy, and persists every raw verdict under round-specific proof paths.
205
- - [ ] **Step 2: Run focused tests** and confirm failure.
206
- - [ ] **Step 3: Implement** prompt/result-file transport and provenance envelopes, reusing the provider watchdog. Advisory policy still writes evidence but returns nonblocking outcome.
207
- - [ ] **Step 4: Run focused tests** and confirm pass.
208
- - [ ] **Step 5: Commit** `feat(quality): add read-only comparative critic`.
209
-
210
- ### Task 11: Quality preflight and serial integration
211
-
212
- **Files:**
213
- - Modify: `src/loop/loop.ts`
214
- - Modify: `src/loop/run-command.ts`
215
- - Modify: `src/loop/reporter.ts`
216
- - Modify: `tests/loop/loop.test.ts`
217
- - Create: `tests/loop/quality.integration.test.ts`
218
-
219
- - [ ] **Step 1: Write failing tests** for blocking preflight before implementation, advisory skip, immutable digest recheck, flow-smoke artifact reuse, consistent win, quality loss → repair, and inconsistent verdict blocking.
220
- - [ ] **Step 2: Run focused suites** and confirm failures.
221
- - [ ] **Step 3: Wire reference/artifact/critic adapters** into the existing quality controller and add reporter phases `quality-preflight` and `comparing`.
222
- - [ ] **Step 4: Run focused tests**, then full suite.
223
- - [ ] **Step 5: Commit** `feat(loop): gate declared stories against pinned quality bars`.
224
-
225
- **Phase 3 gate:** fake provider + static fixture demonstrates both label orders, proof provenance, repair on loss, and no commit on inconsistent judging.
226
-
227
- ---
228
-
229
- ## Phase 4: CLI and explicit unbounded mode
230
-
231
- ### Task 12: Parse and validate quality CLI flags
232
-
233
- **Files:**
234
- - Modify: `src/cli.ts`
235
- - Modify: `src/loop/run-command.ts`
236
- - Modify: `tests/cli.test.ts`
237
- - Modify: `tests/loop/loop-cli.integration.test.ts`
238
-
239
- - [ ] **Step 1: Write failing CLI tests** for `--quality`, `--no-quality`, positive rounds/minutes, policy values, `--quality-unbounded`, conflicts with bounded flags, `--candidates`, and existing flag compatibility.
240
- - [ ] **Step 2: Add failing run-command tests** proving unbounded implies quality, is never persisted, preserves timeout/safe permissions, and prints/streams an explicit warning and effective policy.
241
- - [ ] **Step 3: Run focused tests** and confirm failures.
242
- - [ ] **Step 4: Implement additive parsing and `ResolvedQualityPolicy` precedence.** Resolve CLI override → story → project defaults; undeclared stories remain unchanged.
243
- - [ ] **Step 5: Run focused tests** and confirm pass.
244
- - [ ] **Step 6: Commit** `feat(cli): expose bounded and human-brake quality modes`.
245
-
246
- ### Task 13: Status, pause, cleanup, and resume semantics
247
-
248
- **Files:**
249
- - Modify: `src/loop/reporter.ts`
250
- - Modify: `src/loop/cleanup.ts`
251
- - Modify: `src/loop/run-command.ts`
252
- - Modify: `tests/loop/reporter.test.ts`
253
- - Modify: `tests/loop/cleanup.test.ts`
254
- - Modify: `tests/loop/loop-cli.integration.test.ts`
255
-
256
- - [ ] **Step 1: Add failing tests** for unbounded status text/JSON, current round, elapsed time, reference digest, safe-boundary pause, scoped critic/repair PID cleanup, and resume retaining bounded settings but requiring `--quality-unbounded` again.
257
- - [ ] **Step 2: Run focused suites** and confirm failures.
258
- - [ ] **Step 3: Implement fields and scoped cleanup.** Never persist unbounded as project default or hidden resume escalation.
259
- - [ ] **Step 4: Run focused suites** and full loop tests.
260
- - [ ] **Step 5: Commit** `feat(loop): make quality runtime observable and recoverable`.
261
-
262
- **Phase 4 gate:** manually run the CLI against a fake provider in bounded mode, then unbounded mode; create `.yoke/loop.pause` and observe exit code 3 with retained evidence and active safety settings.
263
-
264
- ---
265
-
266
- ## Phase 5: Production parallel story workers
267
-
268
- ### Task 14: Rich claims and worker lifecycle
269
-
270
- **Files:**
271
- - Modify: `src/loop/claims.ts`
272
- - Create: `src/loop/worker.ts`
273
- - Modify: `tests/loop/claims.test.ts`
274
- - Create: `tests/loop/worker.test.ts`
275
-
276
- - [ ] **Step 1: Write failing tests** for dispatcher ID, PID, base commit, worktree, provider/model, heartbeat, stale takeover, owner-only release, cancellation handle, and one candidate's full gate result without PRD mutation.
277
- - [ ] **Step 2: Run focused tests** and confirm failures.
278
- - [ ] **Step 3: Extend claims backward compatibly** and implement an async `runStoryWorker` adapter around existing runner/verifier/reviewer/quality seams.
279
- - [ ] **Step 4: Run focused tests** and confirm pass.
280
- - [ ] **Step 5: Commit** `feat(loop): add owned provider worker lifecycle`.
281
-
282
- ### Task 15: Async provider subprocess execution
283
-
284
- **Files:**
285
- - Modify: `src/agents/providers.ts`
286
- - Modify: `src/loop/runner.ts`
287
- - Modify: `src/loop/watchdog.ts`
288
- - Modify: `tests/agents/providers.test.ts`
289
- - Modify: `tests/loop/runner.test.ts`
290
- - Modify: `tests/loop/watchdog.test.ts`
291
-
292
- - [ ] **Step 1: Write failing fake-CLI tests** for concurrent streaming processes, telemetry, nonzero exit with partial output, idle timeout, cancellation, Windows command shim path, and project-scoped PID recording.
293
- - [ ] **Step 2: Run focused tests** and confirm failures.
294
- - [ ] **Step 3: Add async provider invocation** using `spawn` with argv arrays where supported and the existing Windows wrapper policy. Return a typed handle with completion and cancellation.
295
- - [ ] **Step 4: Run focused tests** and confirm pass on the current platform; keep platform-conditional contract tests for Windows/non-Windows invocation shapes.
296
- - [ ] **Step 5: Commit** `feat(runner): support cancellable async provider workers`.
297
-
298
- ### Task 16: Dispatcher and verified merge queue
299
-
300
- **Files:**
301
- - Create: `src/loop/dispatcher.ts`
302
- - Modify: `src/loop/parallel.ts`
303
- - Modify: `src/loop/merge-queue.ts`
304
- - Modify: `src/loop/scheduler.ts`
305
- - Create: `tests/loop/dispatcher.test.ts`
306
- - Modify: `tests/loop/parallel.test.ts`
307
- - Modify: `tests/loop/merge-queue.test.ts`
308
-
309
- - [ ] **Step 1: Write failing tests** for dependency readiness, area exclusion, max concurrency, heterogeneous agent affinity, FIFO integration, rebase conflict reopen, integrated criterion/verify/perf/audit rerun, pause before launch, and worker crash cleanup.
310
- - [ ] **Step 2: Run focused tests** and confirm failures.
311
- - [ ] **Step 3: Implement dispatcher composition** over existing scheduler/claims/parallel/merge primitives. Move `passes:true` mutation from worker completion to successful serialized integration.
312
- - [ ] **Step 4: Run focused tests** and confirm pass.
313
- - [ ] **Step 5: Commit** `feat(loop): dispatch parallel stories through verified merge queue`.
314
-
315
- ### Task 17: Enable `--parallel=N` end to end
316
-
317
- **Files:**
318
- - Modify: `src/loop/run-command.ts`
319
- - Modify: `src/cli.ts`
320
- - Modify: `src/loop/reporter.ts`
321
- - Create: `tests/loop/parallel-cli.integration.test.ts`
322
-
323
- - [ ] **Step 1: Replace the existing rejection test** with failing integration tests for `N=1`, `N=2`, automatic isolation, claims, progress, conflict reopen, and final completion.
324
- - [ ] **Step 2: Add failure tests** for nonpositive values, unavailable workers, dirty tree, and unsafe cleanup attempts.
325
- - [ ] **Step 3: Run focused integration tests** and confirm the current `--parallel>1` rejection.
326
- - [ ] **Step 4: Wire dispatcher resolution** while leaving serial default untouched.
327
- - [ ] **Step 5: Run focused tests**, full suite, and build.
328
- - [ ] **Step 6: Commit** `feat(cli): enable dependency-aware parallel loops`.
329
-
330
- **Phase 5 gate:** temporary Git repository with two independent stories and one dependent story runs two fake providers concurrently, serializes integration, reruns integrated gates, and finishes with exactly three story commits and clean PRD state.
331
-
332
- ---
333
-
334
- ## Phase 6: Competing candidates
335
-
336
- ### Task 18: Candidate fan-out and winner selection
337
-
338
- **Files:**
339
- - Create: `src/loop/candidates.ts`
340
- - Create: `tests/loop/candidates.test.ts`
341
- - Modify: `src/quality/verdict.ts`
342
-
343
- - [ ] **Step 1: Write failing tests** for common base, max candidate limit, independent worktrees, mechanical filtering, zero-green block, one-green automatic selection, multiple winners requiring blind candidate-vs-candidate comparison, inconsistent selection block, and winner-only result.
344
- - [ ] **Step 2: Run focused tests** and confirm failure.
345
- - [ ] **Step 3: Implement candidate coordinator** using `runStoryWorker`; never merge branches together.
346
- - [ ] **Step 4: Run focused tests** and confirm pass.
347
- - [ ] **Step 5: Commit** `feat(quality): race isolated candidates and select one winner`.
348
-
349
- ### Task 19: Candidate CLI and cleanup integration
350
-
351
- **Files:**
352
- - Modify: `src/loop/run-command.ts`
353
- - Modify: `src/loop/cleanup.ts`
354
- - Modify: `src/loop/reporter.ts`
355
- - Modify: `tests/loop/loop-cli.integration.test.ts`
356
- - Modify: `tests/loop/cleanup.test.ts`
357
-
358
- - [ ] **Step 1: Add failing tests** for `--candidates=N` requiring a quality declaration, candidate/worktree status, winning merge, losing proof metadata, and cleanup after crash/pause.
359
- - [ ] **Step 2: Run focused tests** and confirm failures.
360
- - [ ] **Step 3: Wire candidate coordinator before merge queue** and reporter phase `selecting-candidate`.
361
- - [ ] **Step 4: Run focused tests** and full suite.
362
- - [ ] **Step 5: Commit** `feat(loop): expose winner-take-all candidate races`.
363
-
364
- **Phase 6 gate:** fake providers produce one mechanically red candidate and two green candidates; blind selection integrates only one green branch and cleans all candidate worktrees.
365
-
366
- ---
367
-
368
- ## Phase 7: Ephemeral decomposition
369
-
370
- ### Task 20: Decomposition contract and planner
371
-
372
- **Files:**
373
- - Create: `src/loop/decomposition.ts`
374
- - Create: `tests/loop/decomposition.test.ts`
375
- - Modify: `src/agents/contracts.ts`
376
- - Modify: `src/loop/runner.ts`
377
-
378
- - [ ] **Step 1: Write failing tests** for valid subtask DAGs, duplicate IDs, cycles, unknown dependencies, area conflicts, parent story binding, result-file transport, and invalid-output fallback to whole-story execution.
379
- - [ ] **Step 2: Run focused tests** and confirm failure.
380
- - [ ] **Step 3: Implement schema and read-only planner call** with a versioned envelope. Persist only under `.yoke/work/<story>/`.
381
- - [ ] **Step 4: Run focused tests** and confirm pass.
382
- - [ ] **Step 5: Commit** `feat(loop): plan ephemeral story subtasks`.
383
-
384
- ### Task 21: Subtask scheduling and synthesis
385
-
386
- **Files:**
387
- - Modify: `src/loop/decomposition.ts`
388
- - Modify: `src/loop/worker.ts`
389
- - Modify: `src/loop/reporter.ts`
390
- - Modify: `tests/loop/decomposition.test.ts`
391
- - Modify: `tests/loop/worker.test.ts`
392
-
393
- - [ ] **Step 1: Add failing tests** for dependency/area scheduling, separate nested worktrees, synthesis into one parent candidate, no subtask commits to main, no PRD mutation, parent-only gates, pause, and cleanup.
394
- - [ ] **Step 2: Run focused tests** and confirm failures.
395
- - [ ] **Step 3: Implement bounded subtask execution** reusing parallel scheduling but preserving the parent as the only story/commit unit.
396
- - [ ] **Step 4: Run focused tests** and full suite.
397
- - [ ] **Step 5: Commit** `feat(loop): execute decomposed work beneath one story gate`.
398
-
399
- ### Task 22: Decomposition opt-in and routing policy
400
-
401
- **Files:**
402
- - Modify: `src/retrofit/config.ts`
403
- - Modify: `src/loop/run-command.ts`
404
- - Modify: `src/routing/router.ts`
405
- - Modify: `tests/retrofit/config.test.ts`
406
- - Modify: `tests/routing/router.test.ts`
407
-
408
- - [ ] **Step 1: Add failing tests** for decomposition disabled by default, explicit enable, no controller call for small stories, and routing-selected decomposition for qualified complex stories.
409
- - [ ] **Step 2: Run focused tests** and confirm failures.
410
- - [ ] **Step 3: Implement minimal policy fields and a deterministic small-story fast path.** Add `decomposition.enabled`, `decomposition.maxSubtasks`, and `decomposition.minAcceptanceCriteria` to config; skip the planner when disabled or below the threshold.
411
- - [ ] **Step 4: Run focused tests** and confirm pass.
412
- - [ ] **Step 5: Commit** `feat(routing): opt into cost-aware story decomposition`.
413
-
414
- **Phase 7 gate:** one parent story decomposes into two independent and one dependent fake subtasks, synthesizes one candidate, runs parent gates once, and lands one commit without PRD additions.
415
-
416
- ---
417
-
418
- ## Phase 8: Routing, telemetry, docs, and release evidence
419
-
420
- ### Task 23: Versioned routing and visible fallback
421
-
422
- **Files:**
423
- - Modify: `src/routing/router.ts`
424
- - Modify: `src/agents/contracts.ts`
425
- - Modify: `tests/routing/router.test.ts`
426
-
427
- - [ ] **Step 1: Write failing tests** for native/result-file route envelope, legacy `YOKE_ROUTE` compatibility, malformed output fallback to SELF with `fallbackReason`, and reporter evidence.
428
- - [ ] **Step 2: Run focused tests** and confirm failures.
429
- - [ ] **Step 3: Prefer versioned result transport** while retaining the old marker as an explicitly reported compatibility fallback.
430
- - [ ] **Step 4: Run focused tests** and confirm pass.
431
- - [ ] **Step 5: Commit** `feat(routing): adopt versioned decisions with visible fallback`.
432
-
433
- ### Task 24: Telemetry envelopes and raw evidence
434
-
435
- **Files:**
436
- - Modify: `src/agents/telemetry.ts`
437
- - Modify: `src/loop/reporter.ts`
438
- - Modify: `tests/agents/telemetry.test.ts`
439
- - Modify: `tests/loop/reporter.test.ts`
440
-
441
- - [ ] **Step 1: Write failing tests** for all three providers, unknown-field retention in per-call raw evidence, aggregate known fields, role/worker/candidate/round attribution, missing usage honesty, and cost accumulation.
442
- - [ ] **Step 2: Run focused tests** and confirm failures.
443
- - [ ] **Step 3: Wrap parsed events in machine envelopes** and persist raw evidence without promoting unknown values.
444
- - [ ] **Step 4: Run focused tests** and confirm pass.
445
- - [ ] **Step 5: Commit** `feat(telemetry): attribute quality and parallel provider calls`.
446
-
447
- ### Task 25: Canon, retrofit, and migration contracts
448
-
449
- **Files:**
450
- - Modify: `canon/loop/prd.schema.md`
451
- - Modify: `canon/loop/loop-spec.md`
452
- - Modify: `canon/skills/authoring-prd/SKILL.md`
453
- - Modify: `canon/skills/yoke-workflow/SKILL.md`
454
- - Modify: `src/retrofit/planners/shared.ts`
455
- - Modify: `tests/canon/real-canon.test.ts`
456
- - Modify: `tests/retrofit/retrofit.integration.test.ts`
457
- - Create: `docs/MIGRATING-TO-1.4.md`
458
-
459
- - [ ] **Step 1: Add failing canon/retrofit assertions** for quality schema guidance, runtime ignores, safe defaults, explicit unbounded warning, parallel/candidate/decomposition contracts, and attribution.
460
- - [ ] **Step 2: Run** `npx vitest run tests/canon/real-canon.test.ts tests/retrofit/retrofit.integration.test.ts` and confirm failures.
461
- - [ ] **Step 3: Update canonical docs and generated guidance** with exact YAML/CLI examples and migration behavior.
462
- - [ ] **Step 4: Run focused tests** and `npm run yoke -- validate canon`.
463
- - [ ] **Step 5: Commit** `docs(canon): define comparative quality and parallel execution`.
464
-
465
- ### Task 26: Benchmarks and end-to-end matrix
466
-
467
- **Files:**
468
- - Modify: `bench/result-schema.mjs`
469
- - Modify: `bench/run.mjs`
470
- - Modify: `bench/run-matrix.mjs`
471
- - Add fixture files under: `bench/fixtures/quality-loop/`
472
- - Modify: `bench/README.md`
473
- - Add deterministic integration tests under: `tests/loop/gauntlet-cli.integration.test.ts`
474
-
475
- - [ ] **Step 1: Add failing result-schema tests/checks** for quality rounds, consistency, reference digest, repair count, parallel workers, candidates, decomposition, conflicts, and per-role calls.
476
- - [ ] **Step 2: Build a deterministic fake quality fixture** with hidden acceptance checks and scripted critic losses/wins.
477
- - [ ] **Step 3: Run the fixture before wiring** and confirm expected schema/test failure.
478
- - [ ] **Step 4: Extend benchmark capture and matrix arms:** quality off/on, bounded, unbounded with pause, serial/parallel, candidates one/two.
479
- - [ ] **Step 5: Run deterministic matrix** and confirm all final hidden tests pass; do not claim authenticated provider performance from fake rows.
480
- - [ ] **Step 6: Commit** `bench: measure quality repair and parallel execution`.
481
-
482
- ### Task 27: Product documentation and final release gates
483
-
484
- **Files:**
485
- - Modify: `README.md`
486
- - Modify: `CHANGELOG.md`
487
- - Modify: `TODOS.md`
488
- - Modify: `docs/MIGRATING-TO-1.4.md`
489
- - Modify: release metadata only when a release is explicitly requested.
490
-
491
- - [ ] **Step 1: Update README** with the authoritative gate order, bounded defaults, explicit human-brake command, safety invariants, quality YAML, parallel/candidate/decomposition behavior, and measured caveats.
492
- - [ ] **Step 2: Remove completed TODOs** for provider subprocess wiring and native schemas only after their end-to-end tests are green; retain broader authenticated samples and signed provenance until separately completed.
493
- - [ ] **Step 3: Run final static gates:** `npm run lint`, `npm run build`, `npm test`, `npm run yoke -- validate canon`, `npm run docs:check`, `npm run audit:ci`, `npm run package:check`.
494
- - [ ] **Step 4: Run manual CLI QA** in a temporary Git fixture: help text; invalid flag; bounded reject/repair/pass; quality-unbounded then pause; parallel independent stories; two-candidate winner; decomposition fallback; status/cleanup after killed fake provider.
495
- - [ ] **Step 5: Review `git diff --check`, `git status --short`, and the complete diff.** Verify `.omo/` and unrelated concurrent changes remain untouched.
496
- - [ ] **Step 6: Commit** only if explicitly requested: `feat: add gated comparative quality loops`.
497
-
498
- ---
499
-
500
- ## Dependency graph and parallel work
501
-
502
- ```text
503
- Phase 1 contracts
504
- -> Phase 2 repair loop
505
- -> Phase 3 reference judging
506
- -> Phase 4 CLI/unbounded
507
-
508
- Phase 1 contracts
509
- -> Phase 5 provider parallelism
510
- -> Phase 6 candidate races
511
-
512
- Phase 5 worker lifecycle
513
- -> Phase 7 decomposition
514
-
515
- Phases 2-7
516
- -> Phase 8 telemetry/docs/benchmarks
517
- ```
518
-
519
- Within Phase 1, Tasks 1 and the initial tests for Task 3 can proceed independently. In Phase 3,
520
- reference acquisition and artifact collection are independent until critic integration. In Phase 5,
521
- async provider work can proceed alongside rich claim tests after their shared handle contract is
522
- agreed. All modifications to `loop.ts`, `run-command.ts`, `runner.ts`, and `reporter.ts` should remain
523
- serialized to avoid conflicting edits.
524
-
525
- ## Plan self-review
526
-
527
- - **Spec coverage:** Tasks 4-6 cover repair; 7-11 references/blind judging; 14-17 parallel stories;
528
- 18-19 candidate races; 20-22 decomposition; 1-3 and 23-24 versioned contracts; 12-13 explicit
529
- bounded/unbounded CLI and operational behavior. Security, observability, attribution, benchmarks,
530
- canon, migration, and manual QA are covered in Tasks 7, 10, 13, and 23-27.
531
- - **Compatibility:** Every new schema field is optional; serial and no-quality paths receive explicit
532
- regression tests; unbounded is a per-run explicit escalation.
533
- - **Type consistency:** `MachineEnvelope`, `QualityVerdict`, `RepairGap`, `ResolvedQualityPolicy`, and
534
- worker/candidate/decomposition outcomes are introduced before their consumers.
535
- - **No hidden weakening:** All repairs rerun gates; workers cannot set PRD pass state; merge queue owns
536
- integrated success; subjective verdicts can reject but never override mechanical failure.
537
- - **No placeholders:** Every task has exact files, expected behavior, commands, and a commit boundary.
1
+ # Gauntlet Quality Loop Implementation Plan
2
+
3
+ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
+
5
+ **Goal:** Add bounded and explicitly unbounded comparative-quality iteration, structured repair, external reference judging, production parallel workers, candidate races, ephemeral decomposition, and versioned provider contracts to Yoke without weakening its mechanical gates.
6
+
7
+ **Architecture:** Extend the existing `yoke loop` state machine rather than adding a second loop. New focused `src/quality/`, `src/agents/contracts.ts`, and loop-worker modules provide pure contracts and adapters; `runLoop` retains story/commit authority. Deliver the work in independently green phases so the unchanged serial, no-quality path remains releasable throughout.
8
+
9
+ **Tech Stack:** TypeScript ESM, Node.js ≥20 standard library, Zod 3, YAML, Vitest 4, Git worktrees, existing Claude/Codex/Gemini CLI adapters, project-local Playwright when visual artifacts are used.
10
+
11
+ **Source design:** `docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md`
12
+
13
+ ---
14
+
15
+ ## File structure
16
+
17
+ ### New production modules
18
+
19
+ - `src/agents/contracts.ts` — versioned machine envelopes shared by routing, review, quality, decomposition, and telemetry.
20
+ - `src/quality/types.ts` — quality policy, reference, candidate, verdict, budget, and outcome types.
21
+ - `src/quality/reference.ts` — safe reference validation, acquisition, hashing, and provenance.
22
+ - `src/quality/artifacts.ts` — candidate artifact collection and deterministic digesting.
23
+ - `src/quality/verdict.ts` — quality verdict file contract and label-swap consistency reduction.
24
+ - `src/quality/runner.ts` — read-only critic invocation and evidence persistence.
25
+ - `src/quality/repair.ts` — selected-gap extraction and repair prompt/runner construction.
26
+ - `src/quality/loop.ts` — bounded/unbounded quality-round controller; no story or commit ownership.
27
+ - `src/loop/worker.ts` — execute one isolated story candidate through implementation and gates.
28
+ - `src/loop/dispatcher.ts` — provider subprocess workers, claims, cancellation, and merge-queue integration.
29
+ - `src/loop/candidates.ts` — same-story candidate fan-out and winner-only selection.
30
+ - `src/loop/decomposition.ts` — ephemeral subtask schema, planning, scheduling, and synthesis.
31
+
32
+ ### Existing modules to modify
33
+
34
+ - `src/review/verdict.ts` — schema version, provenance, and optional actionable finding metadata.
35
+ - `src/retrofit/config.ts` — quality defaults and decomposition configuration.
36
+ - `src/loop/prd.ts` — optional story quality declaration.
37
+ - `src/loop/runner.ts` — repair/decomposition/quality invocations through provider adapters.
38
+ - `src/loop/loop.ts` — invoke quality/repair controller while preserving gate and commit order.
39
+ - `src/loop/run-command.ts` — resolve quality policy, parallel dispatcher, candidates, and CLI options.
40
+ - `src/loop/reporter.ts` — quality/parallel phases and structured status fields.
41
+ - `src/loop/claims.ts` — dispatcher/base/worktree/provider/heartbeat metadata.
42
+ - `src/loop/parallel.ts`, `src/loop/scheduler.ts`, `src/loop/merge-queue.ts` — production worker and integration semantics.
43
+ - `src/routing/router.ts` — consume versioned route contracts and disclose fallback reasons.
44
+ - `src/agents/telemetry.ts` — versioned telemetry envelope plus raw evidence retention.
45
+ - `src/cli.ts` — additive CLI flags and conflict validation.
46
+ - `src/retrofit/planners/shared.ts` and generated templates — runtime ignore paths and config comments.
47
+ - `canon/loop/prd.schema.md`, `canon/loop/loop-spec.md`, `canon/skills/authoring-prd/SKILL.md` — user-facing contracts.
48
+ - `README.md`, `CHANGELOG.md`, `TODOS.md`, migration docs — truthful shipped behavior.
49
+
50
+ ### New test suites
51
+
52
+ - `tests/agents/contracts.test.ts`
53
+ - `tests/quality/{reference,artifacts,verdict,repair,loop}.test.ts`
54
+ - `tests/loop/{worker,dispatcher,candidates,decomposition}.test.ts`
55
+ - `tests/loop/gauntlet-cli.integration.test.ts`
56
+ - `tests/fixtures/fake-agent.mjs` additions for deterministic provider outputs and process behavior.
57
+
58
+ ---
59
+
60
+ ## Phase 1: Versioned contracts and backward-compatible schemas
61
+
62
+ ### Task 1: Shared machine-result envelopes
63
+
64
+ **Files:**
65
+ - Create: `src/agents/contracts.ts`
66
+ - Create: `tests/agents/contracts.test.ts`
67
+
68
+ - [ ] **Step 1: Write failing schema tests** covering a valid envelope, rejection of unknown schema versions, required role/provider/timing fields, optional reported model, usage, and retained raw metadata.
69
+ - [ ] **Step 2: Run** `npx vitest run tests/agents/contracts.test.ts` and confirm failure because the module does not exist.
70
+ - [ ] **Step 3: Implement** `MachineEnvelopeSchema`, `MachineRoleSchema`, `MachineUsageSchema`, and inferred types. Version 1 accepts roles `route`, `review`, `quality`, `decomposition`, `candidate-selection`, `telemetry`; timing is nonnegative integer milliseconds; `raw` is `z.record(z.unknown()).optional()`.
71
+ - [ ] **Step 4: Run** `npx vitest run tests/agents/contracts.test.ts` and confirm pass.
72
+ - [ ] **Step 5: Commit** `feat(contracts): add versioned provider result envelopes`.
73
+
74
+ ### Task 2: Review verdict provenance and actionable gaps
75
+
76
+ **Files:**
77
+ - Modify: `src/review/verdict.ts`
78
+ - Modify: `src/loop/runner.ts`
79
+ - Modify: `tests/review/verdict.test.ts`
80
+ - Modify: `tests/loop/runner.test.ts`
81
+
82
+ - [ ] **Step 1: Add failing tests** proving legacy `{approved,summary,findings}` remains valid and versioned verdicts accept `schemaVersion:1`, envelope provenance, finding `id`, `actionable`, and `evidence` fields.
83
+ - [ ] **Step 2: Add a failing test** for `selectRepairFinding(verdict)` choosing the first actionable blocking finding deterministically by original order and returning `null` for malformed/infrastructure-only rejection.
84
+ - [ ] **Step 3: Run** `npx vitest run tests/review/verdict.test.ts tests/loop/runner.test.ts` and confirm the new expectations fail.
85
+ - [ ] **Step 4: Extend schemas minimally** with optional fields and export `selectRepairFinding`. Update `formatReviewContract` to request the versioned shape while `readReviewVerdict` still parses legacy files.
86
+ - [ ] **Step 5: Run** the two focused suites and confirm pass.
87
+ - [ ] **Step 6: Commit** `feat(review): add actionable versioned verdicts`.
88
+
89
+ ### Task 3: Quality configuration and PRD declarations
90
+
91
+ **Files:**
92
+ - Create: `src/quality/types.ts`
93
+ - Modify: `src/retrofit/config.ts`
94
+ - Modify: `src/loop/prd.ts`
95
+ - Modify: `tests/retrofit/config.test.ts`
96
+ - Modify: `tests/loop/prd.test.ts`
97
+
98
+ - [ ] **Step 1: Write failing config tests** for safe defaults (`enabled:false`, blocking, three rounds, 60 minutes, two consistency checks, two max candidates), invalid nonpositive limits, and backward-compatible configs without `quality`.
99
+ - [ ] **Step 2: Write failing PRD tests** for URL/file/command references, candidate kinds, required rubric, path traversal rejection, and old stories without `quality`.
100
+ - [ ] **Step 3: Run** `npx vitest run tests/retrofit/config.test.ts tests/loop/prd.test.ts` and confirm failures.
101
+ - [ ] **Step 4: Implement focused Zod schemas** in `quality/types.ts`; import them into config and PRD schemas rather than duplicating declarations.
102
+ - [ ] **Step 5: Run focused tests**, then `npm run lint`.
103
+ - [ ] **Step 6: Commit** `feat(quality): add config and story contracts`.
104
+
105
+ **Phase 1 gate:** `npm run lint && npx vitest run tests/agents tests/review tests/retrofit/config.test.ts tests/loop/prd.test.ts`.
106
+
107
+ ---
108
+
109
+ ## Phase 2: Bounded reviewer-to-repair loop
110
+
111
+ ### Task 4: Repair-gap prompt and runner
112
+
113
+ **Files:**
114
+ - Create: `src/quality/repair.ts`
115
+ - Create: `tests/quality/repair.test.ts`
116
+ - Modify: `src/loop/runner.ts`
117
+
118
+ - [ ] **Step 1: Write failing tests** for a prompt containing story, acceptance, current diff instruction, exactly one gap, evidence, project context, and gate names while excluding praise, prior deliberation, and unrelated findings.
119
+ - [ ] **Step 2: Write a failing invocation test** proving repair uses a fresh provider call with workspace-write permissions and the existing watchdog.
120
+ - [ ] **Step 3: Run** `npx vitest run tests/quality/repair.test.ts` and confirm failure.
121
+ - [ ] **Step 4: Implement** `RepairGap`, `buildRepairPrompt`, and `makeRepairRunner` by reusing provider invocation/watchdog seams from `runner.ts`.
122
+ - [ ] **Step 5: Run focused tests** and confirm pass.
123
+ - [ ] **Step 6: Commit** `feat(quality): add fresh single-gap repair runner`.
124
+
125
+ ### Task 5: Quality budget controller
126
+
127
+ **Files:**
128
+ - Create: `src/quality/loop.ts`
129
+ - Create: `tests/quality/loop.test.ts`
130
+
131
+ - [ ] **Step 1: Write failing pure-controller tests** for immediate pass, repair then pass, three-round exhaustion, 60-minute exhaustion using injected clock, unbounded continuation beyond both limits, pause between rounds, and infrastructure failure blocking without repair.
132
+ - [ ] **Step 2: Run** `npx vitest run tests/quality/loop.test.ts` and confirm failure.
133
+ - [ ] **Step 3: Implement** an injected `runQualityRounds(options)` controller returning `passed`, `blocked`, or `paused`, last gap, rounds, elapsed time, and evidence. It calls injected `gateCandidate`, `judgeCandidate`, and `repairCandidate`; it owns no Git or PRD operations.
134
+ - [ ] **Step 4: Run focused tests** and confirm pass.
135
+ - [ ] **Step 5: Commit** `feat(quality): add bounded and unbounded repair controller`.
136
+
137
+ ### Task 6: Integrate repair rounds into serial story execution
138
+
139
+ **Files:**
140
+ - Modify: `src/loop/loop.ts`
141
+ - Modify: `src/loop/reporter.ts`
142
+ - Modify: `tests/loop/loop.test.ts`
143
+ - Modify: `tests/loop/reporter.test.ts`
144
+
145
+ - [ ] **Step 1: Add failing loop tests** proving a review rejection with an actionable finding repairs in the same isolated worktree, reruns criterion/verify/perf/audit, invokes a fresh reviewer, and commits only after pass.
146
+ - [ ] **Step 2: Add failing tests** proving mechanical failure never invokes repair, malformed review blocks, exhaustion blocks with last gap, pause exits at the next round boundary, and existing no-quality behavior invokes each legacy gate exactly once.
147
+ - [ ] **Step 3: Add reporter tests** for phases `repairing` and quality round/budget fields in JSON status.
148
+ - [ ] **Step 4: Run** `npx vitest run tests/loop/loop.test.ts tests/loop/reporter.test.ts` and confirm failures.
149
+ - [ ] **Step 5: Add optional quality/repair dependencies to `LoopOptions`** and thread the pure controller into both isolated and non-isolated candidate paths without moving commit authority.
150
+ - [ ] **Step 6: Run focused tests**, then `npm test` to detect serial regressions.
151
+ - [ ] **Step 7: Commit** `feat(loop): repair actionable review failures behind gates`.
152
+
153
+ **Phase 2 gate:** full build and tests; a deterministic fake-review integration must demonstrate reject → repair → reverify → approve → one commit.
154
+
155
+ ---
156
+
157
+ ## Phase 3: External references and blind comparison
158
+
159
+ ### Task 7: Safe reference acquisition
160
+
161
+ **Files:**
162
+ - Create: `src/quality/reference.ts`
163
+ - Create: `tests/quality/reference.test.ts`
164
+
165
+ - [ ] **Step 1: Write failing tests** for file acquisition, SHA-256 provenance, URL redirect/size/content-type limits through an injected fetcher, private-network rejection, explicit local-project URL allowance, command validation, digest mismatch, and safe storage below `.yoke/references/<digest>`.
166
+ - [ ] **Step 2: Run** `npx vitest run tests/quality/reference.test.ts` and confirm failure.
167
+ - [ ] **Step 3: Implement** pure validation and injected acquisition adapters using Node `crypto`, `fs`, `path`, `URL`, and `fetch`; no new dependency. Store inert bytes plus `provenance.json`. Never return fetched text as prompt instructions.
168
+ - [ ] **Step 4: Run focused tests** and confirm pass.
169
+ - [ ] **Step 5: Commit** `feat(quality): acquire pinned untrusted references safely`.
170
+
171
+ ### Task 8: Candidate artifact collection
172
+
173
+ **Files:**
174
+ - Create: `src/quality/artifacts.ts`
175
+ - Create: `tests/quality/artifacts.test.ts`
176
+ - Modify: `src/smoke/command.ts`
177
+ - Modify: `tests/smoke/command.test.ts`
178
+
179
+ - [ ] **Step 1: Write failing tests** for declared screenshots/files, stable ordered digesting, missing artifacts, traversal rejection, command-output/benchmark capture, and reuse of existing flow-smoke proof files without recapture.
180
+ - [ ] **Step 2: Run** the quality artifact and smoke tests and confirm failure.
181
+ - [ ] **Step 3: Implement** collection under the story proof root and expose smoke proof metadata sufficient for reuse.
182
+ - [ ] **Step 4: Run focused tests** and confirm pass.
183
+ - [ ] **Step 5: Commit** `feat(quality): collect comparable candidate artifacts`.
184
+
185
+ ### Task 9: Blind label-swap verdict contract
186
+
187
+ **Files:**
188
+ - Create: `src/quality/verdict.ts`
189
+ - Create: `tests/quality/verdict.test.ts`
190
+
191
+ - [ ] **Step 1: Write failing tests** for schema validation, random A/B assignment through injected RNG, mandatory swapped second comparison, candidate/candidate consistency, low-confidence rejection, missing evidence, digest mismatch, and reduction to `pass`, `lose`, or `inconsistent`.
192
+ - [ ] **Step 2: Run** `npx vitest run tests/quality/verdict.test.ts` and confirm failure.
193
+ - [ ] **Step 3: Implement** `QualityVerdictSchema`, `assignBlindLabels`, and `reduceComparisons`. Do not add numeric quality scores.
194
+ - [ ] **Step 4: Run focused tests** and confirm pass.
195
+ - [ ] **Step 5: Commit** `feat(quality): enforce blind binary judge consistency`.
196
+
197
+ ### Task 10: Read-only critic and evidence persistence
198
+
199
+ **Files:**
200
+ - Create: `src/quality/runner.ts`
201
+ - Create: `tests/quality/runner.test.ts`
202
+ - Modify: `src/loop/runner.ts`
203
+
204
+ - [ ] **Step 1: Write failing tests** proving the critic receives trusted rubric separately from untrusted artifacts, uses read-only permissions, writes a result file, emits two fresh calls for blocking policy, and persists every raw verdict under round-specific proof paths.
205
+ - [ ] **Step 2: Run focused tests** and confirm failure.
206
+ - [ ] **Step 3: Implement** prompt/result-file transport and provenance envelopes, reusing the provider watchdog. Advisory policy still writes evidence but returns nonblocking outcome.
207
+ - [ ] **Step 4: Run focused tests** and confirm pass.
208
+ - [ ] **Step 5: Commit** `feat(quality): add read-only comparative critic`.
209
+
210
+ ### Task 11: Quality preflight and serial integration
211
+
212
+ **Files:**
213
+ - Modify: `src/loop/loop.ts`
214
+ - Modify: `src/loop/run-command.ts`
215
+ - Modify: `src/loop/reporter.ts`
216
+ - Modify: `tests/loop/loop.test.ts`
217
+ - Create: `tests/loop/quality.integration.test.ts`
218
+
219
+ - [ ] **Step 1: Write failing tests** for blocking preflight before implementation, advisory skip, immutable digest recheck, flow-smoke artifact reuse, consistent win, quality loss → repair, and inconsistent verdict blocking.
220
+ - [ ] **Step 2: Run focused suites** and confirm failures.
221
+ - [ ] **Step 3: Wire reference/artifact/critic adapters** into the existing quality controller and add reporter phases `quality-preflight` and `comparing`.
222
+ - [ ] **Step 4: Run focused tests**, then full suite.
223
+ - [ ] **Step 5: Commit** `feat(loop): gate declared stories against pinned quality bars`.
224
+
225
+ **Phase 3 gate:** fake provider + static fixture demonstrates both label orders, proof provenance, repair on loss, and no commit on inconsistent judging.
226
+
227
+ ---
228
+
229
+ ## Phase 4: CLI and explicit unbounded mode
230
+
231
+ ### Task 12: Parse and validate quality CLI flags
232
+
233
+ **Files:**
234
+ - Modify: `src/cli.ts`
235
+ - Modify: `src/loop/run-command.ts`
236
+ - Modify: `tests/cli.test.ts`
237
+ - Modify: `tests/loop/loop-cli.integration.test.ts`
238
+
239
+ - [ ] **Step 1: Write failing CLI tests** for `--quality`, `--no-quality`, positive rounds/minutes, policy values, `--quality-unbounded`, conflicts with bounded flags, `--candidates`, and existing flag compatibility.
240
+ - [ ] **Step 2: Add failing run-command tests** proving unbounded implies quality, is never persisted, preserves timeout/safe permissions, and prints/streams an explicit warning and effective policy.
241
+ - [ ] **Step 3: Run focused tests** and confirm failures.
242
+ - [ ] **Step 4: Implement additive parsing and `ResolvedQualityPolicy` precedence.** Resolve CLI override → story → project defaults; undeclared stories remain unchanged.
243
+ - [ ] **Step 5: Run focused tests** and confirm pass.
244
+ - [ ] **Step 6: Commit** `feat(cli): expose bounded and human-brake quality modes`.
245
+
246
+ ### Task 13: Status, pause, cleanup, and resume semantics
247
+
248
+ **Files:**
249
+ - Modify: `src/loop/reporter.ts`
250
+ - Modify: `src/loop/cleanup.ts`
251
+ - Modify: `src/loop/run-command.ts`
252
+ - Modify: `tests/loop/reporter.test.ts`
253
+ - Modify: `tests/loop/cleanup.test.ts`
254
+ - Modify: `tests/loop/loop-cli.integration.test.ts`
255
+
256
+ - [ ] **Step 1: Add failing tests** for unbounded status text/JSON, current round, elapsed time, reference digest, safe-boundary pause, scoped critic/repair PID cleanup, and resume retaining bounded settings but requiring `--quality-unbounded` again.
257
+ - [ ] **Step 2: Run focused suites** and confirm failures.
258
+ - [ ] **Step 3: Implement fields and scoped cleanup.** Never persist unbounded as project default or hidden resume escalation.
259
+ - [ ] **Step 4: Run focused suites** and full loop tests.
260
+ - [ ] **Step 5: Commit** `feat(loop): make quality runtime observable and recoverable`.
261
+
262
+ **Phase 4 gate:** manually run the CLI against a fake provider in bounded mode, then unbounded mode; create `.yoke/loop.pause` and observe exit code 3 with retained evidence and active safety settings.
263
+
264
+ ---
265
+
266
+ ## Phase 5: Production parallel story workers
267
+
268
+ ### Task 14: Rich claims and worker lifecycle
269
+
270
+ **Files:**
271
+ - Modify: `src/loop/claims.ts`
272
+ - Create: `src/loop/worker.ts`
273
+ - Modify: `tests/loop/claims.test.ts`
274
+ - Create: `tests/loop/worker.test.ts`
275
+
276
+ - [ ] **Step 1: Write failing tests** for dispatcher ID, PID, base commit, worktree, provider/model, heartbeat, stale takeover, owner-only release, cancellation handle, and one candidate's full gate result without PRD mutation.
277
+ - [ ] **Step 2: Run focused tests** and confirm failures.
278
+ - [ ] **Step 3: Extend claims backward compatibly** and implement an async `runStoryWorker` adapter around existing runner/verifier/reviewer/quality seams.
279
+ - [ ] **Step 4: Run focused tests** and confirm pass.
280
+ - [ ] **Step 5: Commit** `feat(loop): add owned provider worker lifecycle`.
281
+
282
+ ### Task 15: Async provider subprocess execution
283
+
284
+ **Files:**
285
+ - Modify: `src/agents/providers.ts`
286
+ - Modify: `src/loop/runner.ts`
287
+ - Modify: `src/loop/watchdog.ts`
288
+ - Modify: `tests/agents/providers.test.ts`
289
+ - Modify: `tests/loop/runner.test.ts`
290
+ - Modify: `tests/loop/watchdog.test.ts`
291
+
292
+ - [ ] **Step 1: Write failing fake-CLI tests** for concurrent streaming processes, telemetry, nonzero exit with partial output, idle timeout, cancellation, Windows command shim path, and project-scoped PID recording.
293
+ - [ ] **Step 2: Run focused tests** and confirm failures.
294
+ - [ ] **Step 3: Add async provider invocation** using `spawn` with argv arrays where supported and the existing Windows wrapper policy. Return a typed handle with completion and cancellation.
295
+ - [ ] **Step 4: Run focused tests** and confirm pass on the current platform; keep platform-conditional contract tests for Windows/non-Windows invocation shapes.
296
+ - [ ] **Step 5: Commit** `feat(runner): support cancellable async provider workers`.
297
+
298
+ ### Task 16: Dispatcher and verified merge queue
299
+
300
+ **Files:**
301
+ - Create: `src/loop/dispatcher.ts`
302
+ - Modify: `src/loop/parallel.ts`
303
+ - Modify: `src/loop/merge-queue.ts`
304
+ - Modify: `src/loop/scheduler.ts`
305
+ - Create: `tests/loop/dispatcher.test.ts`
306
+ - Modify: `tests/loop/parallel.test.ts`
307
+ - Modify: `tests/loop/merge-queue.test.ts`
308
+
309
+ - [ ] **Step 1: Write failing tests** for dependency readiness, area exclusion, max concurrency, heterogeneous agent affinity, FIFO integration, rebase conflict reopen, integrated criterion/verify/perf/audit rerun, pause before launch, and worker crash cleanup.
310
+ - [ ] **Step 2: Run focused tests** and confirm failures.
311
+ - [ ] **Step 3: Implement dispatcher composition** over existing scheduler/claims/parallel/merge primitives. Move `passes:true` mutation from worker completion to successful serialized integration.
312
+ - [ ] **Step 4: Run focused tests** and confirm pass.
313
+ - [ ] **Step 5: Commit** `feat(loop): dispatch parallel stories through verified merge queue`.
314
+
315
+ ### Task 17: Enable `--parallel=N` end to end
316
+
317
+ **Files:**
318
+ - Modify: `src/loop/run-command.ts`
319
+ - Modify: `src/cli.ts`
320
+ - Modify: `src/loop/reporter.ts`
321
+ - Create: `tests/loop/parallel-cli.integration.test.ts`
322
+
323
+ - [ ] **Step 1: Replace the existing rejection test** with failing integration tests for `N=1`, `N=2`, automatic isolation, claims, progress, conflict reopen, and final completion.
324
+ - [ ] **Step 2: Add failure tests** for nonpositive values, unavailable workers, dirty tree, and unsafe cleanup attempts.
325
+ - [ ] **Step 3: Run focused integration tests** and confirm the current `--parallel>1` rejection.
326
+ - [ ] **Step 4: Wire dispatcher resolution** while leaving serial default untouched.
327
+ - [ ] **Step 5: Run focused tests**, full suite, and build.
328
+ - [ ] **Step 6: Commit** `feat(cli): enable dependency-aware parallel loops`.
329
+
330
+ **Phase 5 gate:** temporary Git repository with two independent stories and one dependent story runs two fake providers concurrently, serializes integration, reruns integrated gates, and finishes with exactly three story commits and clean PRD state.
331
+
332
+ ---
333
+
334
+ ## Phase 6: Competing candidates
335
+
336
+ ### Task 18: Candidate fan-out and winner selection
337
+
338
+ **Files:**
339
+ - Create: `src/loop/candidates.ts`
340
+ - Create: `tests/loop/candidates.test.ts`
341
+ - Modify: `src/quality/verdict.ts`
342
+
343
+ - [ ] **Step 1: Write failing tests** for common base, max candidate limit, independent worktrees, mechanical filtering, zero-green block, one-green automatic selection, multiple winners requiring blind candidate-vs-candidate comparison, inconsistent selection block, and winner-only result.
344
+ - [ ] **Step 2: Run focused tests** and confirm failure.
345
+ - [ ] **Step 3: Implement candidate coordinator** using `runStoryWorker`; never merge branches together.
346
+ - [ ] **Step 4: Run focused tests** and confirm pass.
347
+ - [ ] **Step 5: Commit** `feat(quality): race isolated candidates and select one winner`.
348
+
349
+ ### Task 19: Candidate CLI and cleanup integration
350
+
351
+ **Files:**
352
+ - Modify: `src/loop/run-command.ts`
353
+ - Modify: `src/loop/cleanup.ts`
354
+ - Modify: `src/loop/reporter.ts`
355
+ - Modify: `tests/loop/loop-cli.integration.test.ts`
356
+ - Modify: `tests/loop/cleanup.test.ts`
357
+
358
+ - [ ] **Step 1: Add failing tests** for `--candidates=N` requiring a quality declaration, candidate/worktree status, winning merge, losing proof metadata, and cleanup after crash/pause.
359
+ - [ ] **Step 2: Run focused tests** and confirm failures.
360
+ - [ ] **Step 3: Wire candidate coordinator before merge queue** and reporter phase `selecting-candidate`.
361
+ - [ ] **Step 4: Run focused tests** and full suite.
362
+ - [ ] **Step 5: Commit** `feat(loop): expose winner-take-all candidate races`.
363
+
364
+ **Phase 6 gate:** fake providers produce one mechanically red candidate and two green candidates; blind selection integrates only one green branch and cleans all candidate worktrees.
365
+
366
+ ---
367
+
368
+ ## Phase 7: Ephemeral decomposition
369
+
370
+ ### Task 20: Decomposition contract and planner
371
+
372
+ **Files:**
373
+ - Create: `src/loop/decomposition.ts`
374
+ - Create: `tests/loop/decomposition.test.ts`
375
+ - Modify: `src/agents/contracts.ts`
376
+ - Modify: `src/loop/runner.ts`
377
+
378
+ - [ ] **Step 1: Write failing tests** for valid subtask DAGs, duplicate IDs, cycles, unknown dependencies, area conflicts, parent story binding, result-file transport, and invalid-output fallback to whole-story execution.
379
+ - [ ] **Step 2: Run focused tests** and confirm failure.
380
+ - [ ] **Step 3: Implement schema and read-only planner call** with a versioned envelope. Persist only under `.yoke/work/<story>/`.
381
+ - [ ] **Step 4: Run focused tests** and confirm pass.
382
+ - [ ] **Step 5: Commit** `feat(loop): plan ephemeral story subtasks`.
383
+
384
+ ### Task 21: Subtask scheduling and synthesis
385
+
386
+ **Files:**
387
+ - Modify: `src/loop/decomposition.ts`
388
+ - Modify: `src/loop/worker.ts`
389
+ - Modify: `src/loop/reporter.ts`
390
+ - Modify: `tests/loop/decomposition.test.ts`
391
+ - Modify: `tests/loop/worker.test.ts`
392
+
393
+ - [ ] **Step 1: Add failing tests** for dependency/area scheduling, separate nested worktrees, synthesis into one parent candidate, no subtask commits to main, no PRD mutation, parent-only gates, pause, and cleanup.
394
+ - [ ] **Step 2: Run focused tests** and confirm failures.
395
+ - [ ] **Step 3: Implement bounded subtask execution** reusing parallel scheduling but preserving the parent as the only story/commit unit.
396
+ - [ ] **Step 4: Run focused tests** and full suite.
397
+ - [ ] **Step 5: Commit** `feat(loop): execute decomposed work beneath one story gate`.
398
+
399
+ ### Task 22: Decomposition opt-in and routing policy
400
+
401
+ **Files:**
402
+ - Modify: `src/retrofit/config.ts`
403
+ - Modify: `src/loop/run-command.ts`
404
+ - Modify: `src/routing/router.ts`
405
+ - Modify: `tests/retrofit/config.test.ts`
406
+ - Modify: `tests/routing/router.test.ts`
407
+
408
+ - [ ] **Step 1: Add failing tests** for decomposition disabled by default, explicit enable, no controller call for small stories, and routing-selected decomposition for qualified complex stories.
409
+ - [ ] **Step 2: Run focused tests** and confirm failures.
410
+ - [ ] **Step 3: Implement minimal policy fields and a deterministic small-story fast path.** Add `decomposition.enabled`, `decomposition.maxSubtasks`, and `decomposition.minAcceptanceCriteria` to config; skip the planner when disabled or below the threshold.
411
+ - [ ] **Step 4: Run focused tests** and confirm pass.
412
+ - [ ] **Step 5: Commit** `feat(routing): opt into cost-aware story decomposition`.
413
+
414
+ **Phase 7 gate:** one parent story decomposes into two independent and one dependent fake subtasks, synthesizes one candidate, runs parent gates once, and lands one commit without PRD additions.
415
+
416
+ ---
417
+
418
+ ## Phase 8: Routing, telemetry, docs, and release evidence
419
+
420
+ ### Task 23: Versioned routing and visible fallback
421
+
422
+ **Files:**
423
+ - Modify: `src/routing/router.ts`
424
+ - Modify: `src/agents/contracts.ts`
425
+ - Modify: `tests/routing/router.test.ts`
426
+
427
+ - [ ] **Step 1: Write failing tests** for native/result-file route envelope, legacy `YOKE_ROUTE` compatibility, malformed output fallback to SELF with `fallbackReason`, and reporter evidence.
428
+ - [ ] **Step 2: Run focused tests** and confirm failures.
429
+ - [ ] **Step 3: Prefer versioned result transport** while retaining the old marker as an explicitly reported compatibility fallback.
430
+ - [ ] **Step 4: Run focused tests** and confirm pass.
431
+ - [ ] **Step 5: Commit** `feat(routing): adopt versioned decisions with visible fallback`.
432
+
433
+ ### Task 24: Telemetry envelopes and raw evidence
434
+
435
+ **Files:**
436
+ - Modify: `src/agents/telemetry.ts`
437
+ - Modify: `src/loop/reporter.ts`
438
+ - Modify: `tests/agents/telemetry.test.ts`
439
+ - Modify: `tests/loop/reporter.test.ts`
440
+
441
+ - [ ] **Step 1: Write failing tests** for all three providers, unknown-field retention in per-call raw evidence, aggregate known fields, role/worker/candidate/round attribution, missing usage honesty, and cost accumulation.
442
+ - [ ] **Step 2: Run focused tests** and confirm failures.
443
+ - [ ] **Step 3: Wrap parsed events in machine envelopes** and persist raw evidence without promoting unknown values.
444
+ - [ ] **Step 4: Run focused tests** and confirm pass.
445
+ - [ ] **Step 5: Commit** `feat(telemetry): attribute quality and parallel provider calls`.
446
+
447
+ ### Task 25: Canon, retrofit, and migration contracts
448
+
449
+ **Files:**
450
+ - Modify: `canon/loop/prd.schema.md`
451
+ - Modify: `canon/loop/loop-spec.md`
452
+ - Modify: `canon/skills/authoring-prd/SKILL.md`
453
+ - Modify: `canon/skills/yoke-workflow/SKILL.md`
454
+ - Modify: `src/retrofit/planners/shared.ts`
455
+ - Modify: `tests/canon/real-canon.test.ts`
456
+ - Modify: `tests/retrofit/retrofit.integration.test.ts`
457
+ - Create: `docs/MIGRATING-TO-1.4.md`
458
+
459
+ - [ ] **Step 1: Add failing canon/retrofit assertions** for quality schema guidance, runtime ignores, safe defaults, explicit unbounded warning, parallel/candidate/decomposition contracts, and attribution.
460
+ - [ ] **Step 2: Run** `npx vitest run tests/canon/real-canon.test.ts tests/retrofit/retrofit.integration.test.ts` and confirm failures.
461
+ - [ ] **Step 3: Update canonical docs and generated guidance** with exact YAML/CLI examples and migration behavior.
462
+ - [ ] **Step 4: Run focused tests** and `npm run yoke -- validate canon`.
463
+ - [ ] **Step 5: Commit** `docs(canon): define comparative quality and parallel execution`.
464
+
465
+ ### Task 26: Benchmarks and end-to-end matrix
466
+
467
+ **Files:**
468
+ - Modify: `bench/result-schema.mjs`
469
+ - Modify: `bench/run.mjs`
470
+ - Modify: `bench/run-matrix.mjs`
471
+ - Add fixture files under: `bench/fixtures/quality-loop/`
472
+ - Modify: `bench/README.md`
473
+ - Add deterministic integration tests under: `tests/loop/gauntlet-cli.integration.test.ts`
474
+
475
+ - [ ] **Step 1: Add failing result-schema tests/checks** for quality rounds, consistency, reference digest, repair count, parallel workers, candidates, decomposition, conflicts, and per-role calls.
476
+ - [ ] **Step 2: Build a deterministic fake quality fixture** with hidden acceptance checks and scripted critic losses/wins.
477
+ - [ ] **Step 3: Run the fixture before wiring** and confirm expected schema/test failure.
478
+ - [ ] **Step 4: Extend benchmark capture and matrix arms:** quality off/on, bounded, unbounded with pause, serial/parallel, candidates one/two.
479
+ - [ ] **Step 5: Run deterministic matrix** and confirm all final hidden tests pass; do not claim authenticated provider performance from fake rows.
480
+ - [ ] **Step 6: Commit** `bench: measure quality repair and parallel execution`.
481
+
482
+ ### Task 27: Product documentation and final release gates
483
+
484
+ **Files:**
485
+ - Modify: `README.md`
486
+ - Modify: `CHANGELOG.md`
487
+ - Modify: `TODOS.md`
488
+ - Modify: `docs/MIGRATING-TO-1.4.md`
489
+ - Modify: release metadata only when a release is explicitly requested.
490
+
491
+ - [ ] **Step 1: Update README** with the authoritative gate order, bounded defaults, explicit human-brake command, safety invariants, quality YAML, parallel/candidate/decomposition behavior, and measured caveats.
492
+ - [ ] **Step 2: Remove completed TODOs** for provider subprocess wiring and native schemas only after their end-to-end tests are green; retain broader authenticated samples and signed provenance until separately completed.
493
+ - [ ] **Step 3: Run final static gates:** `npm run lint`, `npm run build`, `npm test`, `npm run yoke -- validate canon`, `npm run docs:check`, `npm run audit:ci`, `npm run package:check`.
494
+ - [ ] **Step 4: Run manual CLI QA** in a temporary Git fixture: help text; invalid flag; bounded reject/repair/pass; quality-unbounded then pause; parallel independent stories; two-candidate winner; decomposition fallback; status/cleanup after killed fake provider.
495
+ - [ ] **Step 5: Review `git diff --check`, `git status --short`, and the complete diff.** Verify `.omo/` and unrelated concurrent changes remain untouched.
496
+ - [ ] **Step 6: Commit** only if explicitly requested: `feat: add gated comparative quality loops`.
497
+
498
+ ---
499
+
500
+ ## Dependency graph and parallel work
501
+
502
+ ```text
503
+ Phase 1 contracts
504
+ -> Phase 2 repair loop
505
+ -> Phase 3 reference judging
506
+ -> Phase 4 CLI/unbounded
507
+
508
+ Phase 1 contracts
509
+ -> Phase 5 provider parallelism
510
+ -> Phase 6 candidate races
511
+
512
+ Phase 5 worker lifecycle
513
+ -> Phase 7 decomposition
514
+
515
+ Phases 2-7
516
+ -> Phase 8 telemetry/docs/benchmarks
517
+ ```
518
+
519
+ Within Phase 1, Tasks 1 and the initial tests for Task 3 can proceed independently. In Phase 3,
520
+ reference acquisition and artifact collection are independent until critic integration. In Phase 5,
521
+ async provider work can proceed alongside rich claim tests after their shared handle contract is
522
+ agreed. All modifications to `loop.ts`, `run-command.ts`, `runner.ts`, and `reporter.ts` should remain
523
+ serialized to avoid conflicting edits.
524
+
525
+ ## Plan self-review
526
+
527
+ - **Spec coverage:** Tasks 4-6 cover repair; 7-11 references/blind judging; 14-17 parallel stories;
528
+ 18-19 candidate races; 20-22 decomposition; 1-3 and 23-24 versioned contracts; 12-13 explicit
529
+ bounded/unbounded CLI and operational behavior. Security, observability, attribution, benchmarks,
530
+ canon, migration, and manual QA are covered in Tasks 7, 10, 13, and 23-27.
531
+ - **Compatibility:** Every new schema field is optional; serial and no-quality paths receive explicit
532
+ regression tests; unbounded is a per-run explicit escalation.
533
+ - **Type consistency:** `MachineEnvelope`, `QualityVerdict`, `RepairGap`, `ResolvedQualityPolicy`, and
534
+ worker/candidate/decomposition outcomes are introduced before their consumers.
535
+ - **No hidden weakening:** All repairs rerun gates; workers cannot set PRD pass state; merge queue owns
536
+ integrated success; subjective verdicts can reject but never override mechanical failure.
537
+ - **No placeholders:** Every task has exact files, expected behavior, commands, and a commit boundary.