@hecer/yoke 1.6.0 → 1.6.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +13 -13
- package/.codex-plugin/plugin.json +7 -7
- package/CHANGELOG.md +294 -288
- package/README.md +874 -874
- package/TODOS.md +5 -5
- package/agents/docs.toml +6 -6
- package/agents/implementer.toml +6 -6
- package/agents/reviewer.toml +6 -6
- package/agents/security.toml +6 -6
- package/bench/README.md +86 -86
- package/bench/RESULTS.md +35 -35
- package/bench/output-compaction.mjs +65 -65
- package/bench/result-schema.mjs +12 -12
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
- package/bench/results/codex-unavailable-1785175418318.json +15 -15
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
- package/bench/run-matrix.mjs +26 -26
- package/bench/run.mjs +106 -106
- package/canon/AGENTS.md +30 -30
- package/canon/context/DECISIONS.md +4 -4
- package/canon/context/GLOSSARY.md +11 -11
- package/canon/context/KNOWLEDGE.md +4 -4
- package/canon/context/PROJECT.md +15 -15
- package/canon/loop/loop-spec.md +65 -65
- package/canon/loop/prd.schema.md +43 -43
- package/canon/manifest.yaml +59 -59
- package/canon/policy/gates.md +7 -7
- package/canon/policy/roles.md +9 -9
- package/canon/skills/ATTRIBUTION.md +99 -99
- package/canon/skills/authoring-prd/SKILL.md +58 -58
- package/canon/skills/brainstorming/SKILL.md +164 -164
- package/canon/skills/codebase-design/DEEPENING.md +15 -15
- package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
- package/canon/skills/codebase-design/SKILL.md +39 -39
- package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
- package/canon/skills/document-release/SKILL.md +302 -302
- package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
- package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
- package/canon/skills/domain-modeling/SKILL.md +35 -35
- package/canon/skills/executing-plans/SKILL.md +70 -70
- package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
- package/canon/skills/health/SKILL.md +177 -177
- package/canon/skills/maintaining-context/SKILL.md +34 -34
- package/canon/skills/minimal-code/SKILL.md +21 -21
- package/canon/skills/no-ai-slop/SKILL.md +103 -103
- package/canon/skills/no-ai-slop/eval.md +43 -43
- package/canon/skills/plan-ceo-review/SKILL.md +541 -541
- package/canon/skills/plan-eng-review/SKILL.md +362 -362
- package/canon/skills/receiving-code-review/SKILL.md +213 -213
- package/canon/skills/requesting-code-review/SKILL.md +105 -105
- package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
- package/canon/skills/retro/SKILL.md +397 -397
- package/canon/skills/review/SKILL.md +246 -246
- package/canon/skills/ship/SKILL.md +691 -691
- package/canon/skills/subagent-driven-development/SKILL.md +277 -277
- package/canon/skills/systematic-debugging/SKILL.md +296 -296
- package/canon/skills/tdd/SKILL.md +371 -371
- package/canon/skills/unslop-ui/SKILL.md +34 -34
- package/canon/skills/using-git-worktrees/SKILL.md +218 -218
- package/canon/skills/verification-before-completion/SKILL.md +139 -139
- package/canon/skills/visual-verification/SKILL.md +54 -54
- package/canon/skills/workflow/SKILL.md +22 -22
- package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
- package/canon/skills/writing-for-agents/SKILL.md +42 -42
- package/canon/skills/writing-plans/SKILL.md +152 -152
- package/canon/skills/writing-skills/SKILL.md +655 -655
- package/canon/skills/yoke-retrofit/SKILL.md +26 -26
- package/canon/skills/yoke-workflow/SKILL.md +20 -20
- package/canon/tools/codex-rtk-hook.mjs +35 -35
- package/canon/tools/graphify.md +3 -3
- package/canon/tools/playwright-mcp.md +3 -3
- package/canon/tools/rtk.md +7 -7
- package/canon/tools/serena.md +6 -6
- package/dist/agents/process.js +3 -0
- package/dist/loop/watchdog.js +1 -1
- package/dist/prd/command.js +17 -17
- package/dist/retrofit/planners/claude.js +14 -14
- package/dist/retrofit/preserve.js +2 -2
- package/docs/MIGRATING-TO-1.0.md +33 -33
- package/docs/MIGRATING-TO-1.1.md +27 -27
- package/docs/MIGRATING-TO-1.4.md +70 -70
- package/docs/PUBLISHING.md +91 -91
- package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
- package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
- package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
- package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
- package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
- package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
- package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
- package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
- package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
- package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
- package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
- package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
- package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
- package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
- package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
- package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
- package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
- package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
- package/gemini-extension.json +6 -6
- package/hooks/hooks.json +19 -19
- package/package.json +87 -87
|
@@ -1,537 +1,537 @@
|
|
|
1
|
-
# Gauntlet Quality Loop Implementation Plan
|
|
2
|
-
|
|
3
|
-
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
|
4
|
-
|
|
5
|
-
**Goal:** Add bounded and explicitly unbounded comparative-quality iteration, structured repair, external reference judging, production parallel workers, candidate races, ephemeral decomposition, and versioned provider contracts to Yoke without weakening its mechanical gates.
|
|
6
|
-
|
|
7
|
-
**Architecture:** Extend the existing `yoke loop` state machine rather than adding a second loop. New focused `src/quality/`, `src/agents/contracts.ts`, and loop-worker modules provide pure contracts and adapters; `runLoop` retains story/commit authority. Deliver the work in independently green phases so the unchanged serial, no-quality path remains releasable throughout.
|
|
8
|
-
|
|
9
|
-
**Tech Stack:** TypeScript ESM, Node.js ≥20 standard library, Zod 3, YAML, Vitest 4, Git worktrees, existing Claude/Codex/Gemini CLI adapters, project-local Playwright when visual artifacts are used.
|
|
10
|
-
|
|
11
|
-
**Source design:** `docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md`
|
|
12
|
-
|
|
13
|
-
---
|
|
14
|
-
|
|
15
|
-
## File structure
|
|
16
|
-
|
|
17
|
-
### New production modules
|
|
18
|
-
|
|
19
|
-
- `src/agents/contracts.ts` — versioned machine envelopes shared by routing, review, quality, decomposition, and telemetry.
|
|
20
|
-
- `src/quality/types.ts` — quality policy, reference, candidate, verdict, budget, and outcome types.
|
|
21
|
-
- `src/quality/reference.ts` — safe reference validation, acquisition, hashing, and provenance.
|
|
22
|
-
- `src/quality/artifacts.ts` — candidate artifact collection and deterministic digesting.
|
|
23
|
-
- `src/quality/verdict.ts` — quality verdict file contract and label-swap consistency reduction.
|
|
24
|
-
- `src/quality/runner.ts` — read-only critic invocation and evidence persistence.
|
|
25
|
-
- `src/quality/repair.ts` — selected-gap extraction and repair prompt/runner construction.
|
|
26
|
-
- `src/quality/loop.ts` — bounded/unbounded quality-round controller; no story or commit ownership.
|
|
27
|
-
- `src/loop/worker.ts` — execute one isolated story candidate through implementation and gates.
|
|
28
|
-
- `src/loop/dispatcher.ts` — provider subprocess workers, claims, cancellation, and merge-queue integration.
|
|
29
|
-
- `src/loop/candidates.ts` — same-story candidate fan-out and winner-only selection.
|
|
30
|
-
- `src/loop/decomposition.ts` — ephemeral subtask schema, planning, scheduling, and synthesis.
|
|
31
|
-
|
|
32
|
-
### Existing modules to modify
|
|
33
|
-
|
|
34
|
-
- `src/review/verdict.ts` — schema version, provenance, and optional actionable finding metadata.
|
|
35
|
-
- `src/retrofit/config.ts` — quality defaults and decomposition configuration.
|
|
36
|
-
- `src/loop/prd.ts` — optional story quality declaration.
|
|
37
|
-
- `src/loop/runner.ts` — repair/decomposition/quality invocations through provider adapters.
|
|
38
|
-
- `src/loop/loop.ts` — invoke quality/repair controller while preserving gate and commit order.
|
|
39
|
-
- `src/loop/run-command.ts` — resolve quality policy, parallel dispatcher, candidates, and CLI options.
|
|
40
|
-
- `src/loop/reporter.ts` — quality/parallel phases and structured status fields.
|
|
41
|
-
- `src/loop/claims.ts` — dispatcher/base/worktree/provider/heartbeat metadata.
|
|
42
|
-
- `src/loop/parallel.ts`, `src/loop/scheduler.ts`, `src/loop/merge-queue.ts` — production worker and integration semantics.
|
|
43
|
-
- `src/routing/router.ts` — consume versioned route contracts and disclose fallback reasons.
|
|
44
|
-
- `src/agents/telemetry.ts` — versioned telemetry envelope plus raw evidence retention.
|
|
45
|
-
- `src/cli.ts` — additive CLI flags and conflict validation.
|
|
46
|
-
- `src/retrofit/planners/shared.ts` and generated templates — runtime ignore paths and config comments.
|
|
47
|
-
- `canon/loop/prd.schema.md`, `canon/loop/loop-spec.md`, `canon/skills/authoring-prd/SKILL.md` — user-facing contracts.
|
|
48
|
-
- `README.md`, `CHANGELOG.md`, `TODOS.md`, migration docs — truthful shipped behavior.
|
|
49
|
-
|
|
50
|
-
### New test suites
|
|
51
|
-
|
|
52
|
-
- `tests/agents/contracts.test.ts`
|
|
53
|
-
- `tests/quality/{reference,artifacts,verdict,repair,loop}.test.ts`
|
|
54
|
-
- `tests/loop/{worker,dispatcher,candidates,decomposition}.test.ts`
|
|
55
|
-
- `tests/loop/gauntlet-cli.integration.test.ts`
|
|
56
|
-
- `tests/fixtures/fake-agent.mjs` additions for deterministic provider outputs and process behavior.
|
|
57
|
-
|
|
58
|
-
---
|
|
59
|
-
|
|
60
|
-
## Phase 1: Versioned contracts and backward-compatible schemas
|
|
61
|
-
|
|
62
|
-
### Task 1: Shared machine-result envelopes
|
|
63
|
-
|
|
64
|
-
**Files:**
|
|
65
|
-
- Create: `src/agents/contracts.ts`
|
|
66
|
-
- Create: `tests/agents/contracts.test.ts`
|
|
67
|
-
|
|
68
|
-
- [ ] **Step 1: Write failing schema tests** covering a valid envelope, rejection of unknown schema versions, required role/provider/timing fields, optional reported model, usage, and retained raw metadata.
|
|
69
|
-
- [ ] **Step 2: Run** `npx vitest run tests/agents/contracts.test.ts` and confirm failure because the module does not exist.
|
|
70
|
-
- [ ] **Step 3: Implement** `MachineEnvelopeSchema`, `MachineRoleSchema`, `MachineUsageSchema`, and inferred types. Version 1 accepts roles `route`, `review`, `quality`, `decomposition`, `candidate-selection`, `telemetry`; timing is nonnegative integer milliseconds; `raw` is `z.record(z.unknown()).optional()`.
|
|
71
|
-
- [ ] **Step 4: Run** `npx vitest run tests/agents/contracts.test.ts` and confirm pass.
|
|
72
|
-
- [ ] **Step 5: Commit** `feat(contracts): add versioned provider result envelopes`.
|
|
73
|
-
|
|
74
|
-
### Task 2: Review verdict provenance and actionable gaps
|
|
75
|
-
|
|
76
|
-
**Files:**
|
|
77
|
-
- Modify: `src/review/verdict.ts`
|
|
78
|
-
- Modify: `src/loop/runner.ts`
|
|
79
|
-
- Modify: `tests/review/verdict.test.ts`
|
|
80
|
-
- Modify: `tests/loop/runner.test.ts`
|
|
81
|
-
|
|
82
|
-
- [ ] **Step 1: Add failing tests** proving legacy `{approved,summary,findings}` remains valid and versioned verdicts accept `schemaVersion:1`, envelope provenance, finding `id`, `actionable`, and `evidence` fields.
|
|
83
|
-
- [ ] **Step 2: Add a failing test** for `selectRepairFinding(verdict)` choosing the first actionable blocking finding deterministically by original order and returning `null` for malformed/infrastructure-only rejection.
|
|
84
|
-
- [ ] **Step 3: Run** `npx vitest run tests/review/verdict.test.ts tests/loop/runner.test.ts` and confirm the new expectations fail.
|
|
85
|
-
- [ ] **Step 4: Extend schemas minimally** with optional fields and export `selectRepairFinding`. Update `formatReviewContract` to request the versioned shape while `readReviewVerdict` still parses legacy files.
|
|
86
|
-
- [ ] **Step 5: Run** the two focused suites and confirm pass.
|
|
87
|
-
- [ ] **Step 6: Commit** `feat(review): add actionable versioned verdicts`.
|
|
88
|
-
|
|
89
|
-
### Task 3: Quality configuration and PRD declarations
|
|
90
|
-
|
|
91
|
-
**Files:**
|
|
92
|
-
- Create: `src/quality/types.ts`
|
|
93
|
-
- Modify: `src/retrofit/config.ts`
|
|
94
|
-
- Modify: `src/loop/prd.ts`
|
|
95
|
-
- Modify: `tests/retrofit/config.test.ts`
|
|
96
|
-
- Modify: `tests/loop/prd.test.ts`
|
|
97
|
-
|
|
98
|
-
- [ ] **Step 1: Write failing config tests** for safe defaults (`enabled:false`, blocking, three rounds, 60 minutes, two consistency checks, two max candidates), invalid nonpositive limits, and backward-compatible configs without `quality`.
|
|
99
|
-
- [ ] **Step 2: Write failing PRD tests** for URL/file/command references, candidate kinds, required rubric, path traversal rejection, and old stories without `quality`.
|
|
100
|
-
- [ ] **Step 3: Run** `npx vitest run tests/retrofit/config.test.ts tests/loop/prd.test.ts` and confirm failures.
|
|
101
|
-
- [ ] **Step 4: Implement focused Zod schemas** in `quality/types.ts`; import them into config and PRD schemas rather than duplicating declarations.
|
|
102
|
-
- [ ] **Step 5: Run focused tests**, then `npm run lint`.
|
|
103
|
-
- [ ] **Step 6: Commit** `feat(quality): add config and story contracts`.
|
|
104
|
-
|
|
105
|
-
**Phase 1 gate:** `npm run lint && npx vitest run tests/agents tests/review tests/retrofit/config.test.ts tests/loop/prd.test.ts`.
|
|
106
|
-
|
|
107
|
-
---
|
|
108
|
-
|
|
109
|
-
## Phase 2: Bounded reviewer-to-repair loop
|
|
110
|
-
|
|
111
|
-
### Task 4: Repair-gap prompt and runner
|
|
112
|
-
|
|
113
|
-
**Files:**
|
|
114
|
-
- Create: `src/quality/repair.ts`
|
|
115
|
-
- Create: `tests/quality/repair.test.ts`
|
|
116
|
-
- Modify: `src/loop/runner.ts`
|
|
117
|
-
|
|
118
|
-
- [ ] **Step 1: Write failing tests** for a prompt containing story, acceptance, current diff instruction, exactly one gap, evidence, project context, and gate names while excluding praise, prior deliberation, and unrelated findings.
|
|
119
|
-
- [ ] **Step 2: Write a failing invocation test** proving repair uses a fresh provider call with workspace-write permissions and the existing watchdog.
|
|
120
|
-
- [ ] **Step 3: Run** `npx vitest run tests/quality/repair.test.ts` and confirm failure.
|
|
121
|
-
- [ ] **Step 4: Implement** `RepairGap`, `buildRepairPrompt`, and `makeRepairRunner` by reusing provider invocation/watchdog seams from `runner.ts`.
|
|
122
|
-
- [ ] **Step 5: Run focused tests** and confirm pass.
|
|
123
|
-
- [ ] **Step 6: Commit** `feat(quality): add fresh single-gap repair runner`.
|
|
124
|
-
|
|
125
|
-
### Task 5: Quality budget controller
|
|
126
|
-
|
|
127
|
-
**Files:**
|
|
128
|
-
- Create: `src/quality/loop.ts`
|
|
129
|
-
- Create: `tests/quality/loop.test.ts`
|
|
130
|
-
|
|
131
|
-
- [ ] **Step 1: Write failing pure-controller tests** for immediate pass, repair then pass, three-round exhaustion, 60-minute exhaustion using injected clock, unbounded continuation beyond both limits, pause between rounds, and infrastructure failure blocking without repair.
|
|
132
|
-
- [ ] **Step 2: Run** `npx vitest run tests/quality/loop.test.ts` and confirm failure.
|
|
133
|
-
- [ ] **Step 3: Implement** an injected `runQualityRounds(options)` controller returning `passed`, `blocked`, or `paused`, last gap, rounds, elapsed time, and evidence. It calls injected `gateCandidate`, `judgeCandidate`, and `repairCandidate`; it owns no Git or PRD operations.
|
|
134
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
135
|
-
- [ ] **Step 5: Commit** `feat(quality): add bounded and unbounded repair controller`.
|
|
136
|
-
|
|
137
|
-
### Task 6: Integrate repair rounds into serial story execution
|
|
138
|
-
|
|
139
|
-
**Files:**
|
|
140
|
-
- Modify: `src/loop/loop.ts`
|
|
141
|
-
- Modify: `src/loop/reporter.ts`
|
|
142
|
-
- Modify: `tests/loop/loop.test.ts`
|
|
143
|
-
- Modify: `tests/loop/reporter.test.ts`
|
|
144
|
-
|
|
145
|
-
- [ ] **Step 1: Add failing loop tests** proving a review rejection with an actionable finding repairs in the same isolated worktree, reruns criterion/verify/perf/audit, invokes a fresh reviewer, and commits only after pass.
|
|
146
|
-
- [ ] **Step 2: Add failing tests** proving mechanical failure never invokes repair, malformed review blocks, exhaustion blocks with last gap, pause exits at the next round boundary, and existing no-quality behavior invokes each legacy gate exactly once.
|
|
147
|
-
- [ ] **Step 3: Add reporter tests** for phases `repairing` and quality round/budget fields in JSON status.
|
|
148
|
-
- [ ] **Step 4: Run** `npx vitest run tests/loop/loop.test.ts tests/loop/reporter.test.ts` and confirm failures.
|
|
149
|
-
- [ ] **Step 5: Add optional quality/repair dependencies to `LoopOptions`** and thread the pure controller into both isolated and non-isolated candidate paths without moving commit authority.
|
|
150
|
-
- [ ] **Step 6: Run focused tests**, then `npm test` to detect serial regressions.
|
|
151
|
-
- [ ] **Step 7: Commit** `feat(loop): repair actionable review failures behind gates`.
|
|
152
|
-
|
|
153
|
-
**Phase 2 gate:** full build and tests; a deterministic fake-review integration must demonstrate reject → repair → reverify → approve → one commit.
|
|
154
|
-
|
|
155
|
-
---
|
|
156
|
-
|
|
157
|
-
## Phase 3: External references and blind comparison
|
|
158
|
-
|
|
159
|
-
### Task 7: Safe reference acquisition
|
|
160
|
-
|
|
161
|
-
**Files:**
|
|
162
|
-
- Create: `src/quality/reference.ts`
|
|
163
|
-
- Create: `tests/quality/reference.test.ts`
|
|
164
|
-
|
|
165
|
-
- [ ] **Step 1: Write failing tests** for file acquisition, SHA-256 provenance, URL redirect/size/content-type limits through an injected fetcher, private-network rejection, explicit local-project URL allowance, command validation, digest mismatch, and safe storage below `.yoke/references/<digest>`.
|
|
166
|
-
- [ ] **Step 2: Run** `npx vitest run tests/quality/reference.test.ts` and confirm failure.
|
|
167
|
-
- [ ] **Step 3: Implement** pure validation and injected acquisition adapters using Node `crypto`, `fs`, `path`, `URL`, and `fetch`; no new dependency. Store inert bytes plus `provenance.json`. Never return fetched text as prompt instructions.
|
|
168
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
169
|
-
- [ ] **Step 5: Commit** `feat(quality): acquire pinned untrusted references safely`.
|
|
170
|
-
|
|
171
|
-
### Task 8: Candidate artifact collection
|
|
172
|
-
|
|
173
|
-
**Files:**
|
|
174
|
-
- Create: `src/quality/artifacts.ts`
|
|
175
|
-
- Create: `tests/quality/artifacts.test.ts`
|
|
176
|
-
- Modify: `src/smoke/command.ts`
|
|
177
|
-
- Modify: `tests/smoke/command.test.ts`
|
|
178
|
-
|
|
179
|
-
- [ ] **Step 1: Write failing tests** for declared screenshots/files, stable ordered digesting, missing artifacts, traversal rejection, command-output/benchmark capture, and reuse of existing flow-smoke proof files without recapture.
|
|
180
|
-
- [ ] **Step 2: Run** the quality artifact and smoke tests and confirm failure.
|
|
181
|
-
- [ ] **Step 3: Implement** collection under the story proof root and expose smoke proof metadata sufficient for reuse.
|
|
182
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
183
|
-
- [ ] **Step 5: Commit** `feat(quality): collect comparable candidate artifacts`.
|
|
184
|
-
|
|
185
|
-
### Task 9: Blind label-swap verdict contract
|
|
186
|
-
|
|
187
|
-
**Files:**
|
|
188
|
-
- Create: `src/quality/verdict.ts`
|
|
189
|
-
- Create: `tests/quality/verdict.test.ts`
|
|
190
|
-
|
|
191
|
-
- [ ] **Step 1: Write failing tests** for schema validation, random A/B assignment through injected RNG, mandatory swapped second comparison, candidate/candidate consistency, low-confidence rejection, missing evidence, digest mismatch, and reduction to `pass`, `lose`, or `inconsistent`.
|
|
192
|
-
- [ ] **Step 2: Run** `npx vitest run tests/quality/verdict.test.ts` and confirm failure.
|
|
193
|
-
- [ ] **Step 3: Implement** `QualityVerdictSchema`, `assignBlindLabels`, and `reduceComparisons`. Do not add numeric quality scores.
|
|
194
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
195
|
-
- [ ] **Step 5: Commit** `feat(quality): enforce blind binary judge consistency`.
|
|
196
|
-
|
|
197
|
-
### Task 10: Read-only critic and evidence persistence
|
|
198
|
-
|
|
199
|
-
**Files:**
|
|
200
|
-
- Create: `src/quality/runner.ts`
|
|
201
|
-
- Create: `tests/quality/runner.test.ts`
|
|
202
|
-
- Modify: `src/loop/runner.ts`
|
|
203
|
-
|
|
204
|
-
- [ ] **Step 1: Write failing tests** proving the critic receives trusted rubric separately from untrusted artifacts, uses read-only permissions, writes a result file, emits two fresh calls for blocking policy, and persists every raw verdict under round-specific proof paths.
|
|
205
|
-
- [ ] **Step 2: Run focused tests** and confirm failure.
|
|
206
|
-
- [ ] **Step 3: Implement** prompt/result-file transport and provenance envelopes, reusing the provider watchdog. Advisory policy still writes evidence but returns nonblocking outcome.
|
|
207
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
208
|
-
- [ ] **Step 5: Commit** `feat(quality): add read-only comparative critic`.
|
|
209
|
-
|
|
210
|
-
### Task 11: Quality preflight and serial integration
|
|
211
|
-
|
|
212
|
-
**Files:**
|
|
213
|
-
- Modify: `src/loop/loop.ts`
|
|
214
|
-
- Modify: `src/loop/run-command.ts`
|
|
215
|
-
- Modify: `src/loop/reporter.ts`
|
|
216
|
-
- Modify: `tests/loop/loop.test.ts`
|
|
217
|
-
- Create: `tests/loop/quality.integration.test.ts`
|
|
218
|
-
|
|
219
|
-
- [ ] **Step 1: Write failing tests** for blocking preflight before implementation, advisory skip, immutable digest recheck, flow-smoke artifact reuse, consistent win, quality loss → repair, and inconsistent verdict blocking.
|
|
220
|
-
- [ ] **Step 2: Run focused suites** and confirm failures.
|
|
221
|
-
- [ ] **Step 3: Wire reference/artifact/critic adapters** into the existing quality controller and add reporter phases `quality-preflight` and `comparing`.
|
|
222
|
-
- [ ] **Step 4: Run focused tests**, then full suite.
|
|
223
|
-
- [ ] **Step 5: Commit** `feat(loop): gate declared stories against pinned quality bars`.
|
|
224
|
-
|
|
225
|
-
**Phase 3 gate:** fake provider + static fixture demonstrates both label orders, proof provenance, repair on loss, and no commit on inconsistent judging.
|
|
226
|
-
|
|
227
|
-
---
|
|
228
|
-
|
|
229
|
-
## Phase 4: CLI and explicit unbounded mode
|
|
230
|
-
|
|
231
|
-
### Task 12: Parse and validate quality CLI flags
|
|
232
|
-
|
|
233
|
-
**Files:**
|
|
234
|
-
- Modify: `src/cli.ts`
|
|
235
|
-
- Modify: `src/loop/run-command.ts`
|
|
236
|
-
- Modify: `tests/cli.test.ts`
|
|
237
|
-
- Modify: `tests/loop/loop-cli.integration.test.ts`
|
|
238
|
-
|
|
239
|
-
- [ ] **Step 1: Write failing CLI tests** for `--quality`, `--no-quality`, positive rounds/minutes, policy values, `--quality-unbounded`, conflicts with bounded flags, `--candidates`, and existing flag compatibility.
|
|
240
|
-
- [ ] **Step 2: Add failing run-command tests** proving unbounded implies quality, is never persisted, preserves timeout/safe permissions, and prints/streams an explicit warning and effective policy.
|
|
241
|
-
- [ ] **Step 3: Run focused tests** and confirm failures.
|
|
242
|
-
- [ ] **Step 4: Implement additive parsing and `ResolvedQualityPolicy` precedence.** Resolve CLI override → story → project defaults; undeclared stories remain unchanged.
|
|
243
|
-
- [ ] **Step 5: Run focused tests** and confirm pass.
|
|
244
|
-
- [ ] **Step 6: Commit** `feat(cli): expose bounded and human-brake quality modes`.
|
|
245
|
-
|
|
246
|
-
### Task 13: Status, pause, cleanup, and resume semantics
|
|
247
|
-
|
|
248
|
-
**Files:**
|
|
249
|
-
- Modify: `src/loop/reporter.ts`
|
|
250
|
-
- Modify: `src/loop/cleanup.ts`
|
|
251
|
-
- Modify: `src/loop/run-command.ts`
|
|
252
|
-
- Modify: `tests/loop/reporter.test.ts`
|
|
253
|
-
- Modify: `tests/loop/cleanup.test.ts`
|
|
254
|
-
- Modify: `tests/loop/loop-cli.integration.test.ts`
|
|
255
|
-
|
|
256
|
-
- [ ] **Step 1: Add failing tests** for unbounded status text/JSON, current round, elapsed time, reference digest, safe-boundary pause, scoped critic/repair PID cleanup, and resume retaining bounded settings but requiring `--quality-unbounded` again.
|
|
257
|
-
- [ ] **Step 2: Run focused suites** and confirm failures.
|
|
258
|
-
- [ ] **Step 3: Implement fields and scoped cleanup.** Never persist unbounded as project default or hidden resume escalation.
|
|
259
|
-
- [ ] **Step 4: Run focused suites** and full loop tests.
|
|
260
|
-
- [ ] **Step 5: Commit** `feat(loop): make quality runtime observable and recoverable`.
|
|
261
|
-
|
|
262
|
-
**Phase 4 gate:** manually run the CLI against a fake provider in bounded mode, then unbounded mode; create `.yoke/loop.pause` and observe exit code 3 with retained evidence and active safety settings.
|
|
263
|
-
|
|
264
|
-
---
|
|
265
|
-
|
|
266
|
-
## Phase 5: Production parallel story workers
|
|
267
|
-
|
|
268
|
-
### Task 14: Rich claims and worker lifecycle
|
|
269
|
-
|
|
270
|
-
**Files:**
|
|
271
|
-
- Modify: `src/loop/claims.ts`
|
|
272
|
-
- Create: `src/loop/worker.ts`
|
|
273
|
-
- Modify: `tests/loop/claims.test.ts`
|
|
274
|
-
- Create: `tests/loop/worker.test.ts`
|
|
275
|
-
|
|
276
|
-
- [ ] **Step 1: Write failing tests** for dispatcher ID, PID, base commit, worktree, provider/model, heartbeat, stale takeover, owner-only release, cancellation handle, and one candidate's full gate result without PRD mutation.
|
|
277
|
-
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
278
|
-
- [ ] **Step 3: Extend claims backward compatibly** and implement an async `runStoryWorker` adapter around existing runner/verifier/reviewer/quality seams.
|
|
279
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
280
|
-
- [ ] **Step 5: Commit** `feat(loop): add owned provider worker lifecycle`.
|
|
281
|
-
|
|
282
|
-
### Task 15: Async provider subprocess execution
|
|
283
|
-
|
|
284
|
-
**Files:**
|
|
285
|
-
- Modify: `src/agents/providers.ts`
|
|
286
|
-
- Modify: `src/loop/runner.ts`
|
|
287
|
-
- Modify: `src/loop/watchdog.ts`
|
|
288
|
-
- Modify: `tests/agents/providers.test.ts`
|
|
289
|
-
- Modify: `tests/loop/runner.test.ts`
|
|
290
|
-
- Modify: `tests/loop/watchdog.test.ts`
|
|
291
|
-
|
|
292
|
-
- [ ] **Step 1: Write failing fake-CLI tests** for concurrent streaming processes, telemetry, nonzero exit with partial output, idle timeout, cancellation, Windows command shim path, and project-scoped PID recording.
|
|
293
|
-
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
294
|
-
- [ ] **Step 3: Add async provider invocation** using `spawn` with argv arrays where supported and the existing Windows wrapper policy. Return a typed handle with completion and cancellation.
|
|
295
|
-
- [ ] **Step 4: Run focused tests** and confirm pass on the current platform; keep platform-conditional contract tests for Windows/non-Windows invocation shapes.
|
|
296
|
-
- [ ] **Step 5: Commit** `feat(runner): support cancellable async provider workers`.
|
|
297
|
-
|
|
298
|
-
### Task 16: Dispatcher and verified merge queue
|
|
299
|
-
|
|
300
|
-
**Files:**
|
|
301
|
-
- Create: `src/loop/dispatcher.ts`
|
|
302
|
-
- Modify: `src/loop/parallel.ts`
|
|
303
|
-
- Modify: `src/loop/merge-queue.ts`
|
|
304
|
-
- Modify: `src/loop/scheduler.ts`
|
|
305
|
-
- Create: `tests/loop/dispatcher.test.ts`
|
|
306
|
-
- Modify: `tests/loop/parallel.test.ts`
|
|
307
|
-
- Modify: `tests/loop/merge-queue.test.ts`
|
|
308
|
-
|
|
309
|
-
- [ ] **Step 1: Write failing tests** for dependency readiness, area exclusion, max concurrency, heterogeneous agent affinity, FIFO integration, rebase conflict reopen, integrated criterion/verify/perf/audit rerun, pause before launch, and worker crash cleanup.
|
|
310
|
-
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
311
|
-
- [ ] **Step 3: Implement dispatcher composition** over existing scheduler/claims/parallel/merge primitives. Move `passes:true` mutation from worker completion to successful serialized integration.
|
|
312
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
313
|
-
- [ ] **Step 5: Commit** `feat(loop): dispatch parallel stories through verified merge queue`.
|
|
314
|
-
|
|
315
|
-
### Task 17: Enable `--parallel=N` end to end
|
|
316
|
-
|
|
317
|
-
**Files:**
|
|
318
|
-
- Modify: `src/loop/run-command.ts`
|
|
319
|
-
- Modify: `src/cli.ts`
|
|
320
|
-
- Modify: `src/loop/reporter.ts`
|
|
321
|
-
- Create: `tests/loop/parallel-cli.integration.test.ts`
|
|
322
|
-
|
|
323
|
-
- [ ] **Step 1: Replace the existing rejection test** with failing integration tests for `N=1`, `N=2`, automatic isolation, claims, progress, conflict reopen, and final completion.
|
|
324
|
-
- [ ] **Step 2: Add failure tests** for nonpositive values, unavailable workers, dirty tree, and unsafe cleanup attempts.
|
|
325
|
-
- [ ] **Step 3: Run focused integration tests** and confirm the current `--parallel>1` rejection.
|
|
326
|
-
- [ ] **Step 4: Wire dispatcher resolution** while leaving serial default untouched.
|
|
327
|
-
- [ ] **Step 5: Run focused tests**, full suite, and build.
|
|
328
|
-
- [ ] **Step 6: Commit** `feat(cli): enable dependency-aware parallel loops`.
|
|
329
|
-
|
|
330
|
-
**Phase 5 gate:** temporary Git repository with two independent stories and one dependent story runs two fake providers concurrently, serializes integration, reruns integrated gates, and finishes with exactly three story commits and clean PRD state.
|
|
331
|
-
|
|
332
|
-
---
|
|
333
|
-
|
|
334
|
-
## Phase 6: Competing candidates
|
|
335
|
-
|
|
336
|
-
### Task 18: Candidate fan-out and winner selection
|
|
337
|
-
|
|
338
|
-
**Files:**
|
|
339
|
-
- Create: `src/loop/candidates.ts`
|
|
340
|
-
- Create: `tests/loop/candidates.test.ts`
|
|
341
|
-
- Modify: `src/quality/verdict.ts`
|
|
342
|
-
|
|
343
|
-
- [ ] **Step 1: Write failing tests** for common base, max candidate limit, independent worktrees, mechanical filtering, zero-green block, one-green automatic selection, multiple winners requiring blind candidate-vs-candidate comparison, inconsistent selection block, and winner-only result.
|
|
344
|
-
- [ ] **Step 2: Run focused tests** and confirm failure.
|
|
345
|
-
- [ ] **Step 3: Implement candidate coordinator** using `runStoryWorker`; never merge branches together.
|
|
346
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
347
|
-
- [ ] **Step 5: Commit** `feat(quality): race isolated candidates and select one winner`.
|
|
348
|
-
|
|
349
|
-
### Task 19: Candidate CLI and cleanup integration
|
|
350
|
-
|
|
351
|
-
**Files:**
|
|
352
|
-
- Modify: `src/loop/run-command.ts`
|
|
353
|
-
- Modify: `src/loop/cleanup.ts`
|
|
354
|
-
- Modify: `src/loop/reporter.ts`
|
|
355
|
-
- Modify: `tests/loop/loop-cli.integration.test.ts`
|
|
356
|
-
- Modify: `tests/loop/cleanup.test.ts`
|
|
357
|
-
|
|
358
|
-
- [ ] **Step 1: Add failing tests** for `--candidates=N` requiring a quality declaration, candidate/worktree status, winning merge, losing proof metadata, and cleanup after crash/pause.
|
|
359
|
-
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
360
|
-
- [ ] **Step 3: Wire candidate coordinator before merge queue** and reporter phase `selecting-candidate`.
|
|
361
|
-
- [ ] **Step 4: Run focused tests** and full suite.
|
|
362
|
-
- [ ] **Step 5: Commit** `feat(loop): expose winner-take-all candidate races`.
|
|
363
|
-
|
|
364
|
-
**Phase 6 gate:** fake providers produce one mechanically red candidate and two green candidates; blind selection integrates only one green branch and cleans all candidate worktrees.
|
|
365
|
-
|
|
366
|
-
---
|
|
367
|
-
|
|
368
|
-
## Phase 7: Ephemeral decomposition
|
|
369
|
-
|
|
370
|
-
### Task 20: Decomposition contract and planner
|
|
371
|
-
|
|
372
|
-
**Files:**
|
|
373
|
-
- Create: `src/loop/decomposition.ts`
|
|
374
|
-
- Create: `tests/loop/decomposition.test.ts`
|
|
375
|
-
- Modify: `src/agents/contracts.ts`
|
|
376
|
-
- Modify: `src/loop/runner.ts`
|
|
377
|
-
|
|
378
|
-
- [ ] **Step 1: Write failing tests** for valid subtask DAGs, duplicate IDs, cycles, unknown dependencies, area conflicts, parent story binding, result-file transport, and invalid-output fallback to whole-story execution.
|
|
379
|
-
- [ ] **Step 2: Run focused tests** and confirm failure.
|
|
380
|
-
- [ ] **Step 3: Implement schema and read-only planner call** with a versioned envelope. Persist only under `.yoke/work/<story>/`.
|
|
381
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
382
|
-
- [ ] **Step 5: Commit** `feat(loop): plan ephemeral story subtasks`.
|
|
383
|
-
|
|
384
|
-
### Task 21: Subtask scheduling and synthesis
|
|
385
|
-
|
|
386
|
-
**Files:**
|
|
387
|
-
- Modify: `src/loop/decomposition.ts`
|
|
388
|
-
- Modify: `src/loop/worker.ts`
|
|
389
|
-
- Modify: `src/loop/reporter.ts`
|
|
390
|
-
- Modify: `tests/loop/decomposition.test.ts`
|
|
391
|
-
- Modify: `tests/loop/worker.test.ts`
|
|
392
|
-
|
|
393
|
-
- [ ] **Step 1: Add failing tests** for dependency/area scheduling, separate nested worktrees, synthesis into one parent candidate, no subtask commits to main, no PRD mutation, parent-only gates, pause, and cleanup.
|
|
394
|
-
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
395
|
-
- [ ] **Step 3: Implement bounded subtask execution** reusing parallel scheduling but preserving the parent as the only story/commit unit.
|
|
396
|
-
- [ ] **Step 4: Run focused tests** and full suite.
|
|
397
|
-
- [ ] **Step 5: Commit** `feat(loop): execute decomposed work beneath one story gate`.
|
|
398
|
-
|
|
399
|
-
### Task 22: Decomposition opt-in and routing policy
|
|
400
|
-
|
|
401
|
-
**Files:**
|
|
402
|
-
- Modify: `src/retrofit/config.ts`
|
|
403
|
-
- Modify: `src/loop/run-command.ts`
|
|
404
|
-
- Modify: `src/routing/router.ts`
|
|
405
|
-
- Modify: `tests/retrofit/config.test.ts`
|
|
406
|
-
- Modify: `tests/routing/router.test.ts`
|
|
407
|
-
|
|
408
|
-
- [ ] **Step 1: Add failing tests** for decomposition disabled by default, explicit enable, no controller call for small stories, and routing-selected decomposition for qualified complex stories.
|
|
409
|
-
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
410
|
-
- [ ] **Step 3: Implement minimal policy fields and a deterministic small-story fast path.** Add `decomposition.enabled`, `decomposition.maxSubtasks`, and `decomposition.minAcceptanceCriteria` to config; skip the planner when disabled or below the threshold.
|
|
411
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
412
|
-
- [ ] **Step 5: Commit** `feat(routing): opt into cost-aware story decomposition`.
|
|
413
|
-
|
|
414
|
-
**Phase 7 gate:** one parent story decomposes into two independent and one dependent fake subtasks, synthesizes one candidate, runs parent gates once, and lands one commit without PRD additions.
|
|
415
|
-
|
|
416
|
-
---
|
|
417
|
-
|
|
418
|
-
## Phase 8: Routing, telemetry, docs, and release evidence
|
|
419
|
-
|
|
420
|
-
### Task 23: Versioned routing and visible fallback
|
|
421
|
-
|
|
422
|
-
**Files:**
|
|
423
|
-
- Modify: `src/routing/router.ts`
|
|
424
|
-
- Modify: `src/agents/contracts.ts`
|
|
425
|
-
- Modify: `tests/routing/router.test.ts`
|
|
426
|
-
|
|
427
|
-
- [ ] **Step 1: Write failing tests** for native/result-file route envelope, legacy `YOKE_ROUTE` compatibility, malformed output fallback to SELF with `fallbackReason`, and reporter evidence.
|
|
428
|
-
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
429
|
-
- [ ] **Step 3: Prefer versioned result transport** while retaining the old marker as an explicitly reported compatibility fallback.
|
|
430
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
431
|
-
- [ ] **Step 5: Commit** `feat(routing): adopt versioned decisions with visible fallback`.
|
|
432
|
-
|
|
433
|
-
### Task 24: Telemetry envelopes and raw evidence
|
|
434
|
-
|
|
435
|
-
**Files:**
|
|
436
|
-
- Modify: `src/agents/telemetry.ts`
|
|
437
|
-
- Modify: `src/loop/reporter.ts`
|
|
438
|
-
- Modify: `tests/agents/telemetry.test.ts`
|
|
439
|
-
- Modify: `tests/loop/reporter.test.ts`
|
|
440
|
-
|
|
441
|
-
- [ ] **Step 1: Write failing tests** for all three providers, unknown-field retention in per-call raw evidence, aggregate known fields, role/worker/candidate/round attribution, missing usage honesty, and cost accumulation.
|
|
442
|
-
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
443
|
-
- [ ] **Step 3: Wrap parsed events in machine envelopes** and persist raw evidence without promoting unknown values.
|
|
444
|
-
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
445
|
-
- [ ] **Step 5: Commit** `feat(telemetry): attribute quality and parallel provider calls`.
|
|
446
|
-
|
|
447
|
-
### Task 25: Canon, retrofit, and migration contracts
|
|
448
|
-
|
|
449
|
-
**Files:**
|
|
450
|
-
- Modify: `canon/loop/prd.schema.md`
|
|
451
|
-
- Modify: `canon/loop/loop-spec.md`
|
|
452
|
-
- Modify: `canon/skills/authoring-prd/SKILL.md`
|
|
453
|
-
- Modify: `canon/skills/yoke-workflow/SKILL.md`
|
|
454
|
-
- Modify: `src/retrofit/planners/shared.ts`
|
|
455
|
-
- Modify: `tests/canon/real-canon.test.ts`
|
|
456
|
-
- Modify: `tests/retrofit/retrofit.integration.test.ts`
|
|
457
|
-
- Create: `docs/MIGRATING-TO-1.4.md`
|
|
458
|
-
|
|
459
|
-
- [ ] **Step 1: Add failing canon/retrofit assertions** for quality schema guidance, runtime ignores, safe defaults, explicit unbounded warning, parallel/candidate/decomposition contracts, and attribution.
|
|
460
|
-
- [ ] **Step 2: Run** `npx vitest run tests/canon/real-canon.test.ts tests/retrofit/retrofit.integration.test.ts` and confirm failures.
|
|
461
|
-
- [ ] **Step 3: Update canonical docs and generated guidance** with exact YAML/CLI examples and migration behavior.
|
|
462
|
-
- [ ] **Step 4: Run focused tests** and `npm run yoke -- validate canon`.
|
|
463
|
-
- [ ] **Step 5: Commit** `docs(canon): define comparative quality and parallel execution`.
|
|
464
|
-
|
|
465
|
-
### Task 26: Benchmarks and end-to-end matrix
|
|
466
|
-
|
|
467
|
-
**Files:**
|
|
468
|
-
- Modify: `bench/result-schema.mjs`
|
|
469
|
-
- Modify: `bench/run.mjs`
|
|
470
|
-
- Modify: `bench/run-matrix.mjs`
|
|
471
|
-
- Add fixture files under: `bench/fixtures/quality-loop/`
|
|
472
|
-
- Modify: `bench/README.md`
|
|
473
|
-
- Add deterministic integration tests under: `tests/loop/gauntlet-cli.integration.test.ts`
|
|
474
|
-
|
|
475
|
-
- [ ] **Step 1: Add failing result-schema tests/checks** for quality rounds, consistency, reference digest, repair count, parallel workers, candidates, decomposition, conflicts, and per-role calls.
|
|
476
|
-
- [ ] **Step 2: Build a deterministic fake quality fixture** with hidden acceptance checks and scripted critic losses/wins.
|
|
477
|
-
- [ ] **Step 3: Run the fixture before wiring** and confirm expected schema/test failure.
|
|
478
|
-
- [ ] **Step 4: Extend benchmark capture and matrix arms:** quality off/on, bounded, unbounded with pause, serial/parallel, candidates one/two.
|
|
479
|
-
- [ ] **Step 5: Run deterministic matrix** and confirm all final hidden tests pass; do not claim authenticated provider performance from fake rows.
|
|
480
|
-
- [ ] **Step 6: Commit** `bench: measure quality repair and parallel execution`.
|
|
481
|
-
|
|
482
|
-
### Task 27: Product documentation and final release gates
|
|
483
|
-
|
|
484
|
-
**Files:**
|
|
485
|
-
- Modify: `README.md`
|
|
486
|
-
- Modify: `CHANGELOG.md`
|
|
487
|
-
- Modify: `TODOS.md`
|
|
488
|
-
- Modify: `docs/MIGRATING-TO-1.4.md`
|
|
489
|
-
- Modify: release metadata only when a release is explicitly requested.
|
|
490
|
-
|
|
491
|
-
- [ ] **Step 1: Update README** with the authoritative gate order, bounded defaults, explicit human-brake command, safety invariants, quality YAML, parallel/candidate/decomposition behavior, and measured caveats.
|
|
492
|
-
- [ ] **Step 2: Remove completed TODOs** for provider subprocess wiring and native schemas only after their end-to-end tests are green; retain broader authenticated samples and signed provenance until separately completed.
|
|
493
|
-
- [ ] **Step 3: Run final static gates:** `npm run lint`, `npm run build`, `npm test`, `npm run yoke -- validate canon`, `npm run docs:check`, `npm run audit:ci`, `npm run package:check`.
|
|
494
|
-
- [ ] **Step 4: Run manual CLI QA** in a temporary Git fixture: help text; invalid flag; bounded reject/repair/pass; quality-unbounded then pause; parallel independent stories; two-candidate winner; decomposition fallback; status/cleanup after killed fake provider.
|
|
495
|
-
- [ ] **Step 5: Review `git diff --check`, `git status --short`, and the complete diff.** Verify `.omo/` and unrelated concurrent changes remain untouched.
|
|
496
|
-
- [ ] **Step 6: Commit** only if explicitly requested: `feat: add gated comparative quality loops`.
|
|
497
|
-
|
|
498
|
-
---
|
|
499
|
-
|
|
500
|
-
## Dependency graph and parallel work
|
|
501
|
-
|
|
502
|
-
```text
|
|
503
|
-
Phase 1 contracts
|
|
504
|
-
-> Phase 2 repair loop
|
|
505
|
-
-> Phase 3 reference judging
|
|
506
|
-
-> Phase 4 CLI/unbounded
|
|
507
|
-
|
|
508
|
-
Phase 1 contracts
|
|
509
|
-
-> Phase 5 provider parallelism
|
|
510
|
-
-> Phase 6 candidate races
|
|
511
|
-
|
|
512
|
-
Phase 5 worker lifecycle
|
|
513
|
-
-> Phase 7 decomposition
|
|
514
|
-
|
|
515
|
-
Phases 2-7
|
|
516
|
-
-> Phase 8 telemetry/docs/benchmarks
|
|
517
|
-
```
|
|
518
|
-
|
|
519
|
-
Within Phase 1, Tasks 1 and the initial tests for Task 3 can proceed independently. In Phase 3,
|
|
520
|
-
reference acquisition and artifact collection are independent until critic integration. In Phase 5,
|
|
521
|
-
async provider work can proceed alongside rich claim tests after their shared handle contract is
|
|
522
|
-
agreed. All modifications to `loop.ts`, `run-command.ts`, `runner.ts`, and `reporter.ts` should remain
|
|
523
|
-
serialized to avoid conflicting edits.
|
|
524
|
-
|
|
525
|
-
## Plan self-review
|
|
526
|
-
|
|
527
|
-
- **Spec coverage:** Tasks 4-6 cover repair; 7-11 references/blind judging; 14-17 parallel stories;
|
|
528
|
-
18-19 candidate races; 20-22 decomposition; 1-3 and 23-24 versioned contracts; 12-13 explicit
|
|
529
|
-
bounded/unbounded CLI and operational behavior. Security, observability, attribution, benchmarks,
|
|
530
|
-
canon, migration, and manual QA are covered in Tasks 7, 10, 13, and 23-27.
|
|
531
|
-
- **Compatibility:** Every new schema field is optional; serial and no-quality paths receive explicit
|
|
532
|
-
regression tests; unbounded is a per-run explicit escalation.
|
|
533
|
-
- **Type consistency:** `MachineEnvelope`, `QualityVerdict`, `RepairGap`, `ResolvedQualityPolicy`, and
|
|
534
|
-
worker/candidate/decomposition outcomes are introduced before their consumers.
|
|
535
|
-
- **No hidden weakening:** All repairs rerun gates; workers cannot set PRD pass state; merge queue owns
|
|
536
|
-
integrated success; subjective verdicts can reject but never override mechanical failure.
|
|
537
|
-
- **No placeholders:** Every task has exact files, expected behavior, commands, and a commit boundary.
|
|
1
|
+
# Gauntlet Quality Loop Implementation Plan
|
|
2
|
+
|
|
3
|
+
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
|
4
|
+
|
|
5
|
+
**Goal:** Add bounded and explicitly unbounded comparative-quality iteration, structured repair, external reference judging, production parallel workers, candidate races, ephemeral decomposition, and versioned provider contracts to Yoke without weakening its mechanical gates.
|
|
6
|
+
|
|
7
|
+
**Architecture:** Extend the existing `yoke loop` state machine rather than adding a second loop. New focused `src/quality/`, `src/agents/contracts.ts`, and loop-worker modules provide pure contracts and adapters; `runLoop` retains story/commit authority. Deliver the work in independently green phases so the unchanged serial, no-quality path remains releasable throughout.
|
|
8
|
+
|
|
9
|
+
**Tech Stack:** TypeScript ESM, Node.js ≥20 standard library, Zod 3, YAML, Vitest 4, Git worktrees, existing Claude/Codex/Gemini CLI adapters, project-local Playwright when visual artifacts are used.
|
|
10
|
+
|
|
11
|
+
**Source design:** `docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md`
|
|
12
|
+
|
|
13
|
+
---
|
|
14
|
+
|
|
15
|
+
## File structure
|
|
16
|
+
|
|
17
|
+
### New production modules
|
|
18
|
+
|
|
19
|
+
- `src/agents/contracts.ts` — versioned machine envelopes shared by routing, review, quality, decomposition, and telemetry.
|
|
20
|
+
- `src/quality/types.ts` — quality policy, reference, candidate, verdict, budget, and outcome types.
|
|
21
|
+
- `src/quality/reference.ts` — safe reference validation, acquisition, hashing, and provenance.
|
|
22
|
+
- `src/quality/artifacts.ts` — candidate artifact collection and deterministic digesting.
|
|
23
|
+
- `src/quality/verdict.ts` — quality verdict file contract and label-swap consistency reduction.
|
|
24
|
+
- `src/quality/runner.ts` — read-only critic invocation and evidence persistence.
|
|
25
|
+
- `src/quality/repair.ts` — selected-gap extraction and repair prompt/runner construction.
|
|
26
|
+
- `src/quality/loop.ts` — bounded/unbounded quality-round controller; no story or commit ownership.
|
|
27
|
+
- `src/loop/worker.ts` — execute one isolated story candidate through implementation and gates.
|
|
28
|
+
- `src/loop/dispatcher.ts` — provider subprocess workers, claims, cancellation, and merge-queue integration.
|
|
29
|
+
- `src/loop/candidates.ts` — same-story candidate fan-out and winner-only selection.
|
|
30
|
+
- `src/loop/decomposition.ts` — ephemeral subtask schema, planning, scheduling, and synthesis.
|
|
31
|
+
|
|
32
|
+
### Existing modules to modify
|
|
33
|
+
|
|
34
|
+
- `src/review/verdict.ts` — schema version, provenance, and optional actionable finding metadata.
|
|
35
|
+
- `src/retrofit/config.ts` — quality defaults and decomposition configuration.
|
|
36
|
+
- `src/loop/prd.ts` — optional story quality declaration.
|
|
37
|
+
- `src/loop/runner.ts` — repair/decomposition/quality invocations through provider adapters.
|
|
38
|
+
- `src/loop/loop.ts` — invoke quality/repair controller while preserving gate and commit order.
|
|
39
|
+
- `src/loop/run-command.ts` — resolve quality policy, parallel dispatcher, candidates, and CLI options.
|
|
40
|
+
- `src/loop/reporter.ts` — quality/parallel phases and structured status fields.
|
|
41
|
+
- `src/loop/claims.ts` — dispatcher/base/worktree/provider/heartbeat metadata.
|
|
42
|
+
- `src/loop/parallel.ts`, `src/loop/scheduler.ts`, `src/loop/merge-queue.ts` — production worker and integration semantics.
|
|
43
|
+
- `src/routing/router.ts` — consume versioned route contracts and disclose fallback reasons.
|
|
44
|
+
- `src/agents/telemetry.ts` — versioned telemetry envelope plus raw evidence retention.
|
|
45
|
+
- `src/cli.ts` — additive CLI flags and conflict validation.
|
|
46
|
+
- `src/retrofit/planners/shared.ts` and generated templates — runtime ignore paths and config comments.
|
|
47
|
+
- `canon/loop/prd.schema.md`, `canon/loop/loop-spec.md`, `canon/skills/authoring-prd/SKILL.md` — user-facing contracts.
|
|
48
|
+
- `README.md`, `CHANGELOG.md`, `TODOS.md`, migration docs — truthful shipped behavior.
|
|
49
|
+
|
|
50
|
+
### New test suites
|
|
51
|
+
|
|
52
|
+
- `tests/agents/contracts.test.ts`
|
|
53
|
+
- `tests/quality/{reference,artifacts,verdict,repair,loop}.test.ts`
|
|
54
|
+
- `tests/loop/{worker,dispatcher,candidates,decomposition}.test.ts`
|
|
55
|
+
- `tests/loop/gauntlet-cli.integration.test.ts`
|
|
56
|
+
- `tests/fixtures/fake-agent.mjs` additions for deterministic provider outputs and process behavior.
|
|
57
|
+
|
|
58
|
+
---
|
|
59
|
+
|
|
60
|
+
## Phase 1: Versioned contracts and backward-compatible schemas
|
|
61
|
+
|
|
62
|
+
### Task 1: Shared machine-result envelopes
|
|
63
|
+
|
|
64
|
+
**Files:**
|
|
65
|
+
- Create: `src/agents/contracts.ts`
|
|
66
|
+
- Create: `tests/agents/contracts.test.ts`
|
|
67
|
+
|
|
68
|
+
- [ ] **Step 1: Write failing schema tests** covering a valid envelope, rejection of unknown schema versions, required role/provider/timing fields, optional reported model, usage, and retained raw metadata.
|
|
69
|
+
- [ ] **Step 2: Run** `npx vitest run tests/agents/contracts.test.ts` and confirm failure because the module does not exist.
|
|
70
|
+
- [ ] **Step 3: Implement** `MachineEnvelopeSchema`, `MachineRoleSchema`, `MachineUsageSchema`, and inferred types. Version 1 accepts roles `route`, `review`, `quality`, `decomposition`, `candidate-selection`, `telemetry`; timing is nonnegative integer milliseconds; `raw` is `z.record(z.unknown()).optional()`.
|
|
71
|
+
- [ ] **Step 4: Run** `npx vitest run tests/agents/contracts.test.ts` and confirm pass.
|
|
72
|
+
- [ ] **Step 5: Commit** `feat(contracts): add versioned provider result envelopes`.
|
|
73
|
+
|
|
74
|
+
### Task 2: Review verdict provenance and actionable gaps
|
|
75
|
+
|
|
76
|
+
**Files:**
|
|
77
|
+
- Modify: `src/review/verdict.ts`
|
|
78
|
+
- Modify: `src/loop/runner.ts`
|
|
79
|
+
- Modify: `tests/review/verdict.test.ts`
|
|
80
|
+
- Modify: `tests/loop/runner.test.ts`
|
|
81
|
+
|
|
82
|
+
- [ ] **Step 1: Add failing tests** proving legacy `{approved,summary,findings}` remains valid and versioned verdicts accept `schemaVersion:1`, envelope provenance, finding `id`, `actionable`, and `evidence` fields.
|
|
83
|
+
- [ ] **Step 2: Add a failing test** for `selectRepairFinding(verdict)` choosing the first actionable blocking finding deterministically by original order and returning `null` for malformed/infrastructure-only rejection.
|
|
84
|
+
- [ ] **Step 3: Run** `npx vitest run tests/review/verdict.test.ts tests/loop/runner.test.ts` and confirm the new expectations fail.
|
|
85
|
+
- [ ] **Step 4: Extend schemas minimally** with optional fields and export `selectRepairFinding`. Update `formatReviewContract` to request the versioned shape while `readReviewVerdict` still parses legacy files.
|
|
86
|
+
- [ ] **Step 5: Run** the two focused suites and confirm pass.
|
|
87
|
+
- [ ] **Step 6: Commit** `feat(review): add actionable versioned verdicts`.
|
|
88
|
+
|
|
89
|
+
### Task 3: Quality configuration and PRD declarations
|
|
90
|
+
|
|
91
|
+
**Files:**
|
|
92
|
+
- Create: `src/quality/types.ts`
|
|
93
|
+
- Modify: `src/retrofit/config.ts`
|
|
94
|
+
- Modify: `src/loop/prd.ts`
|
|
95
|
+
- Modify: `tests/retrofit/config.test.ts`
|
|
96
|
+
- Modify: `tests/loop/prd.test.ts`
|
|
97
|
+
|
|
98
|
+
- [ ] **Step 1: Write failing config tests** for safe defaults (`enabled:false`, blocking, three rounds, 60 minutes, two consistency checks, two max candidates), invalid nonpositive limits, and backward-compatible configs without `quality`.
|
|
99
|
+
- [ ] **Step 2: Write failing PRD tests** for URL/file/command references, candidate kinds, required rubric, path traversal rejection, and old stories without `quality`.
|
|
100
|
+
- [ ] **Step 3: Run** `npx vitest run tests/retrofit/config.test.ts tests/loop/prd.test.ts` and confirm failures.
|
|
101
|
+
- [ ] **Step 4: Implement focused Zod schemas** in `quality/types.ts`; import them into config and PRD schemas rather than duplicating declarations.
|
|
102
|
+
- [ ] **Step 5: Run focused tests**, then `npm run lint`.
|
|
103
|
+
- [ ] **Step 6: Commit** `feat(quality): add config and story contracts`.
|
|
104
|
+
|
|
105
|
+
**Phase 1 gate:** `npm run lint && npx vitest run tests/agents tests/review tests/retrofit/config.test.ts tests/loop/prd.test.ts`.
|
|
106
|
+
|
|
107
|
+
---
|
|
108
|
+
|
|
109
|
+
## Phase 2: Bounded reviewer-to-repair loop
|
|
110
|
+
|
|
111
|
+
### Task 4: Repair-gap prompt and runner
|
|
112
|
+
|
|
113
|
+
**Files:**
|
|
114
|
+
- Create: `src/quality/repair.ts`
|
|
115
|
+
- Create: `tests/quality/repair.test.ts`
|
|
116
|
+
- Modify: `src/loop/runner.ts`
|
|
117
|
+
|
|
118
|
+
- [ ] **Step 1: Write failing tests** for a prompt containing story, acceptance, current diff instruction, exactly one gap, evidence, project context, and gate names while excluding praise, prior deliberation, and unrelated findings.
|
|
119
|
+
- [ ] **Step 2: Write a failing invocation test** proving repair uses a fresh provider call with workspace-write permissions and the existing watchdog.
|
|
120
|
+
- [ ] **Step 3: Run** `npx vitest run tests/quality/repair.test.ts` and confirm failure.
|
|
121
|
+
- [ ] **Step 4: Implement** `RepairGap`, `buildRepairPrompt`, and `makeRepairRunner` by reusing provider invocation/watchdog seams from `runner.ts`.
|
|
122
|
+
- [ ] **Step 5: Run focused tests** and confirm pass.
|
|
123
|
+
- [ ] **Step 6: Commit** `feat(quality): add fresh single-gap repair runner`.
|
|
124
|
+
|
|
125
|
+
### Task 5: Quality budget controller
|
|
126
|
+
|
|
127
|
+
**Files:**
|
|
128
|
+
- Create: `src/quality/loop.ts`
|
|
129
|
+
- Create: `tests/quality/loop.test.ts`
|
|
130
|
+
|
|
131
|
+
- [ ] **Step 1: Write failing pure-controller tests** for immediate pass, repair then pass, three-round exhaustion, 60-minute exhaustion using injected clock, unbounded continuation beyond both limits, pause between rounds, and infrastructure failure blocking without repair.
|
|
132
|
+
- [ ] **Step 2: Run** `npx vitest run tests/quality/loop.test.ts` and confirm failure.
|
|
133
|
+
- [ ] **Step 3: Implement** an injected `runQualityRounds(options)` controller returning `passed`, `blocked`, or `paused`, last gap, rounds, elapsed time, and evidence. It calls injected `gateCandidate`, `judgeCandidate`, and `repairCandidate`; it owns no Git or PRD operations.
|
|
134
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
135
|
+
- [ ] **Step 5: Commit** `feat(quality): add bounded and unbounded repair controller`.
|
|
136
|
+
|
|
137
|
+
### Task 6: Integrate repair rounds into serial story execution
|
|
138
|
+
|
|
139
|
+
**Files:**
|
|
140
|
+
- Modify: `src/loop/loop.ts`
|
|
141
|
+
- Modify: `src/loop/reporter.ts`
|
|
142
|
+
- Modify: `tests/loop/loop.test.ts`
|
|
143
|
+
- Modify: `tests/loop/reporter.test.ts`
|
|
144
|
+
|
|
145
|
+
- [ ] **Step 1: Add failing loop tests** proving a review rejection with an actionable finding repairs in the same isolated worktree, reruns criterion/verify/perf/audit, invokes a fresh reviewer, and commits only after pass.
|
|
146
|
+
- [ ] **Step 2: Add failing tests** proving mechanical failure never invokes repair, malformed review blocks, exhaustion blocks with last gap, pause exits at the next round boundary, and existing no-quality behavior invokes each legacy gate exactly once.
|
|
147
|
+
- [ ] **Step 3: Add reporter tests** for phases `repairing` and quality round/budget fields in JSON status.
|
|
148
|
+
- [ ] **Step 4: Run** `npx vitest run tests/loop/loop.test.ts tests/loop/reporter.test.ts` and confirm failures.
|
|
149
|
+
- [ ] **Step 5: Add optional quality/repair dependencies to `LoopOptions`** and thread the pure controller into both isolated and non-isolated candidate paths without moving commit authority.
|
|
150
|
+
- [ ] **Step 6: Run focused tests**, then `npm test` to detect serial regressions.
|
|
151
|
+
- [ ] **Step 7: Commit** `feat(loop): repair actionable review failures behind gates`.
|
|
152
|
+
|
|
153
|
+
**Phase 2 gate:** full build and tests; a deterministic fake-review integration must demonstrate reject → repair → reverify → approve → one commit.
|
|
154
|
+
|
|
155
|
+
---
|
|
156
|
+
|
|
157
|
+
## Phase 3: External references and blind comparison
|
|
158
|
+
|
|
159
|
+
### Task 7: Safe reference acquisition
|
|
160
|
+
|
|
161
|
+
**Files:**
|
|
162
|
+
- Create: `src/quality/reference.ts`
|
|
163
|
+
- Create: `tests/quality/reference.test.ts`
|
|
164
|
+
|
|
165
|
+
- [ ] **Step 1: Write failing tests** for file acquisition, SHA-256 provenance, URL redirect/size/content-type limits through an injected fetcher, private-network rejection, explicit local-project URL allowance, command validation, digest mismatch, and safe storage below `.yoke/references/<digest>`.
|
|
166
|
+
- [ ] **Step 2: Run** `npx vitest run tests/quality/reference.test.ts` and confirm failure.
|
|
167
|
+
- [ ] **Step 3: Implement** pure validation and injected acquisition adapters using Node `crypto`, `fs`, `path`, `URL`, and `fetch`; no new dependency. Store inert bytes plus `provenance.json`. Never return fetched text as prompt instructions.
|
|
168
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
169
|
+
- [ ] **Step 5: Commit** `feat(quality): acquire pinned untrusted references safely`.
|
|
170
|
+
|
|
171
|
+
### Task 8: Candidate artifact collection
|
|
172
|
+
|
|
173
|
+
**Files:**
|
|
174
|
+
- Create: `src/quality/artifacts.ts`
|
|
175
|
+
- Create: `tests/quality/artifacts.test.ts`
|
|
176
|
+
- Modify: `src/smoke/command.ts`
|
|
177
|
+
- Modify: `tests/smoke/command.test.ts`
|
|
178
|
+
|
|
179
|
+
- [ ] **Step 1: Write failing tests** for declared screenshots/files, stable ordered digesting, missing artifacts, traversal rejection, command-output/benchmark capture, and reuse of existing flow-smoke proof files without recapture.
|
|
180
|
+
- [ ] **Step 2: Run** the quality artifact and smoke tests and confirm failure.
|
|
181
|
+
- [ ] **Step 3: Implement** collection under the story proof root and expose smoke proof metadata sufficient for reuse.
|
|
182
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
183
|
+
- [ ] **Step 5: Commit** `feat(quality): collect comparable candidate artifacts`.
|
|
184
|
+
|
|
185
|
+
### Task 9: Blind label-swap verdict contract
|
|
186
|
+
|
|
187
|
+
**Files:**
|
|
188
|
+
- Create: `src/quality/verdict.ts`
|
|
189
|
+
- Create: `tests/quality/verdict.test.ts`
|
|
190
|
+
|
|
191
|
+
- [ ] **Step 1: Write failing tests** for schema validation, random A/B assignment through injected RNG, mandatory swapped second comparison, candidate/candidate consistency, low-confidence rejection, missing evidence, digest mismatch, and reduction to `pass`, `lose`, or `inconsistent`.
|
|
192
|
+
- [ ] **Step 2: Run** `npx vitest run tests/quality/verdict.test.ts` and confirm failure.
|
|
193
|
+
- [ ] **Step 3: Implement** `QualityVerdictSchema`, `assignBlindLabels`, and `reduceComparisons`. Do not add numeric quality scores.
|
|
194
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
195
|
+
- [ ] **Step 5: Commit** `feat(quality): enforce blind binary judge consistency`.
|
|
196
|
+
|
|
197
|
+
### Task 10: Read-only critic and evidence persistence
|
|
198
|
+
|
|
199
|
+
**Files:**
|
|
200
|
+
- Create: `src/quality/runner.ts`
|
|
201
|
+
- Create: `tests/quality/runner.test.ts`
|
|
202
|
+
- Modify: `src/loop/runner.ts`
|
|
203
|
+
|
|
204
|
+
- [ ] **Step 1: Write failing tests** proving the critic receives trusted rubric separately from untrusted artifacts, uses read-only permissions, writes a result file, emits two fresh calls for blocking policy, and persists every raw verdict under round-specific proof paths.
|
|
205
|
+
- [ ] **Step 2: Run focused tests** and confirm failure.
|
|
206
|
+
- [ ] **Step 3: Implement** prompt/result-file transport and provenance envelopes, reusing the provider watchdog. Advisory policy still writes evidence but returns nonblocking outcome.
|
|
207
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
208
|
+
- [ ] **Step 5: Commit** `feat(quality): add read-only comparative critic`.
|
|
209
|
+
|
|
210
|
+
### Task 11: Quality preflight and serial integration
|
|
211
|
+
|
|
212
|
+
**Files:**
|
|
213
|
+
- Modify: `src/loop/loop.ts`
|
|
214
|
+
- Modify: `src/loop/run-command.ts`
|
|
215
|
+
- Modify: `src/loop/reporter.ts`
|
|
216
|
+
- Modify: `tests/loop/loop.test.ts`
|
|
217
|
+
- Create: `tests/loop/quality.integration.test.ts`
|
|
218
|
+
|
|
219
|
+
- [ ] **Step 1: Write failing tests** for blocking preflight before implementation, advisory skip, immutable digest recheck, flow-smoke artifact reuse, consistent win, quality loss → repair, and inconsistent verdict blocking.
|
|
220
|
+
- [ ] **Step 2: Run focused suites** and confirm failures.
|
|
221
|
+
- [ ] **Step 3: Wire reference/artifact/critic adapters** into the existing quality controller and add reporter phases `quality-preflight` and `comparing`.
|
|
222
|
+
- [ ] **Step 4: Run focused tests**, then full suite.
|
|
223
|
+
- [ ] **Step 5: Commit** `feat(loop): gate declared stories against pinned quality bars`.
|
|
224
|
+
|
|
225
|
+
**Phase 3 gate:** fake provider + static fixture demonstrates both label orders, proof provenance, repair on loss, and no commit on inconsistent judging.
|
|
226
|
+
|
|
227
|
+
---
|
|
228
|
+
|
|
229
|
+
## Phase 4: CLI and explicit unbounded mode
|
|
230
|
+
|
|
231
|
+
### Task 12: Parse and validate quality CLI flags
|
|
232
|
+
|
|
233
|
+
**Files:**
|
|
234
|
+
- Modify: `src/cli.ts`
|
|
235
|
+
- Modify: `src/loop/run-command.ts`
|
|
236
|
+
- Modify: `tests/cli.test.ts`
|
|
237
|
+
- Modify: `tests/loop/loop-cli.integration.test.ts`
|
|
238
|
+
|
|
239
|
+
- [ ] **Step 1: Write failing CLI tests** for `--quality`, `--no-quality`, positive rounds/minutes, policy values, `--quality-unbounded`, conflicts with bounded flags, `--candidates`, and existing flag compatibility.
|
|
240
|
+
- [ ] **Step 2: Add failing run-command tests** proving unbounded implies quality, is never persisted, preserves timeout/safe permissions, and prints/streams an explicit warning and effective policy.
|
|
241
|
+
- [ ] **Step 3: Run focused tests** and confirm failures.
|
|
242
|
+
- [ ] **Step 4: Implement additive parsing and `ResolvedQualityPolicy` precedence.** Resolve CLI override → story → project defaults; undeclared stories remain unchanged.
|
|
243
|
+
- [ ] **Step 5: Run focused tests** and confirm pass.
|
|
244
|
+
- [ ] **Step 6: Commit** `feat(cli): expose bounded and human-brake quality modes`.
|
|
245
|
+
|
|
246
|
+
### Task 13: Status, pause, cleanup, and resume semantics
|
|
247
|
+
|
|
248
|
+
**Files:**
|
|
249
|
+
- Modify: `src/loop/reporter.ts`
|
|
250
|
+
- Modify: `src/loop/cleanup.ts`
|
|
251
|
+
- Modify: `src/loop/run-command.ts`
|
|
252
|
+
- Modify: `tests/loop/reporter.test.ts`
|
|
253
|
+
- Modify: `tests/loop/cleanup.test.ts`
|
|
254
|
+
- Modify: `tests/loop/loop-cli.integration.test.ts`
|
|
255
|
+
|
|
256
|
+
- [ ] **Step 1: Add failing tests** for unbounded status text/JSON, current round, elapsed time, reference digest, safe-boundary pause, scoped critic/repair PID cleanup, and resume retaining bounded settings but requiring `--quality-unbounded` again.
|
|
257
|
+
- [ ] **Step 2: Run focused suites** and confirm failures.
|
|
258
|
+
- [ ] **Step 3: Implement fields and scoped cleanup.** Never persist unbounded as project default or hidden resume escalation.
|
|
259
|
+
- [ ] **Step 4: Run focused suites** and full loop tests.
|
|
260
|
+
- [ ] **Step 5: Commit** `feat(loop): make quality runtime observable and recoverable`.
|
|
261
|
+
|
|
262
|
+
**Phase 4 gate:** manually run the CLI against a fake provider in bounded mode, then unbounded mode; create `.yoke/loop.pause` and observe exit code 3 with retained evidence and active safety settings.
|
|
263
|
+
|
|
264
|
+
---
|
|
265
|
+
|
|
266
|
+
## Phase 5: Production parallel story workers
|
|
267
|
+
|
|
268
|
+
### Task 14: Rich claims and worker lifecycle
|
|
269
|
+
|
|
270
|
+
**Files:**
|
|
271
|
+
- Modify: `src/loop/claims.ts`
|
|
272
|
+
- Create: `src/loop/worker.ts`
|
|
273
|
+
- Modify: `tests/loop/claims.test.ts`
|
|
274
|
+
- Create: `tests/loop/worker.test.ts`
|
|
275
|
+
|
|
276
|
+
- [ ] **Step 1: Write failing tests** for dispatcher ID, PID, base commit, worktree, provider/model, heartbeat, stale takeover, owner-only release, cancellation handle, and one candidate's full gate result without PRD mutation.
|
|
277
|
+
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
278
|
+
- [ ] **Step 3: Extend claims backward compatibly** and implement an async `runStoryWorker` adapter around existing runner/verifier/reviewer/quality seams.
|
|
279
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
280
|
+
- [ ] **Step 5: Commit** `feat(loop): add owned provider worker lifecycle`.
|
|
281
|
+
|
|
282
|
+
### Task 15: Async provider subprocess execution
|
|
283
|
+
|
|
284
|
+
**Files:**
|
|
285
|
+
- Modify: `src/agents/providers.ts`
|
|
286
|
+
- Modify: `src/loop/runner.ts`
|
|
287
|
+
- Modify: `src/loop/watchdog.ts`
|
|
288
|
+
- Modify: `tests/agents/providers.test.ts`
|
|
289
|
+
- Modify: `tests/loop/runner.test.ts`
|
|
290
|
+
- Modify: `tests/loop/watchdog.test.ts`
|
|
291
|
+
|
|
292
|
+
- [ ] **Step 1: Write failing fake-CLI tests** for concurrent streaming processes, telemetry, nonzero exit with partial output, idle timeout, cancellation, Windows command shim path, and project-scoped PID recording.
|
|
293
|
+
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
294
|
+
- [ ] **Step 3: Add async provider invocation** using `spawn` with argv arrays where supported and the existing Windows wrapper policy. Return a typed handle with completion and cancellation.
|
|
295
|
+
- [ ] **Step 4: Run focused tests** and confirm pass on the current platform; keep platform-conditional contract tests for Windows/non-Windows invocation shapes.
|
|
296
|
+
- [ ] **Step 5: Commit** `feat(runner): support cancellable async provider workers`.
|
|
297
|
+
|
|
298
|
+
### Task 16: Dispatcher and verified merge queue
|
|
299
|
+
|
|
300
|
+
**Files:**
|
|
301
|
+
- Create: `src/loop/dispatcher.ts`
|
|
302
|
+
- Modify: `src/loop/parallel.ts`
|
|
303
|
+
- Modify: `src/loop/merge-queue.ts`
|
|
304
|
+
- Modify: `src/loop/scheduler.ts`
|
|
305
|
+
- Create: `tests/loop/dispatcher.test.ts`
|
|
306
|
+
- Modify: `tests/loop/parallel.test.ts`
|
|
307
|
+
- Modify: `tests/loop/merge-queue.test.ts`
|
|
308
|
+
|
|
309
|
+
- [ ] **Step 1: Write failing tests** for dependency readiness, area exclusion, max concurrency, heterogeneous agent affinity, FIFO integration, rebase conflict reopen, integrated criterion/verify/perf/audit rerun, pause before launch, and worker crash cleanup.
|
|
310
|
+
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
311
|
+
- [ ] **Step 3: Implement dispatcher composition** over existing scheduler/claims/parallel/merge primitives. Move `passes:true` mutation from worker completion to successful serialized integration.
|
|
312
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
313
|
+
- [ ] **Step 5: Commit** `feat(loop): dispatch parallel stories through verified merge queue`.
|
|
314
|
+
|
|
315
|
+
### Task 17: Enable `--parallel=N` end to end
|
|
316
|
+
|
|
317
|
+
**Files:**
|
|
318
|
+
- Modify: `src/loop/run-command.ts`
|
|
319
|
+
- Modify: `src/cli.ts`
|
|
320
|
+
- Modify: `src/loop/reporter.ts`
|
|
321
|
+
- Create: `tests/loop/parallel-cli.integration.test.ts`
|
|
322
|
+
|
|
323
|
+
- [ ] **Step 1: Replace the existing rejection test** with failing integration tests for `N=1`, `N=2`, automatic isolation, claims, progress, conflict reopen, and final completion.
|
|
324
|
+
- [ ] **Step 2: Add failure tests** for nonpositive values, unavailable workers, dirty tree, and unsafe cleanup attempts.
|
|
325
|
+
- [ ] **Step 3: Run focused integration tests** and confirm the current `--parallel>1` rejection.
|
|
326
|
+
- [ ] **Step 4: Wire dispatcher resolution** while leaving serial default untouched.
|
|
327
|
+
- [ ] **Step 5: Run focused tests**, full suite, and build.
|
|
328
|
+
- [ ] **Step 6: Commit** `feat(cli): enable dependency-aware parallel loops`.
|
|
329
|
+
|
|
330
|
+
**Phase 5 gate:** temporary Git repository with two independent stories and one dependent story runs two fake providers concurrently, serializes integration, reruns integrated gates, and finishes with exactly three story commits and clean PRD state.
|
|
331
|
+
|
|
332
|
+
---
|
|
333
|
+
|
|
334
|
+
## Phase 6: Competing candidates
|
|
335
|
+
|
|
336
|
+
### Task 18: Candidate fan-out and winner selection
|
|
337
|
+
|
|
338
|
+
**Files:**
|
|
339
|
+
- Create: `src/loop/candidates.ts`
|
|
340
|
+
- Create: `tests/loop/candidates.test.ts`
|
|
341
|
+
- Modify: `src/quality/verdict.ts`
|
|
342
|
+
|
|
343
|
+
- [ ] **Step 1: Write failing tests** for common base, max candidate limit, independent worktrees, mechanical filtering, zero-green block, one-green automatic selection, multiple winners requiring blind candidate-vs-candidate comparison, inconsistent selection block, and winner-only result.
|
|
344
|
+
- [ ] **Step 2: Run focused tests** and confirm failure.
|
|
345
|
+
- [ ] **Step 3: Implement candidate coordinator** using `runStoryWorker`; never merge branches together.
|
|
346
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
347
|
+
- [ ] **Step 5: Commit** `feat(quality): race isolated candidates and select one winner`.
|
|
348
|
+
|
|
349
|
+
### Task 19: Candidate CLI and cleanup integration
|
|
350
|
+
|
|
351
|
+
**Files:**
|
|
352
|
+
- Modify: `src/loop/run-command.ts`
|
|
353
|
+
- Modify: `src/loop/cleanup.ts`
|
|
354
|
+
- Modify: `src/loop/reporter.ts`
|
|
355
|
+
- Modify: `tests/loop/loop-cli.integration.test.ts`
|
|
356
|
+
- Modify: `tests/loop/cleanup.test.ts`
|
|
357
|
+
|
|
358
|
+
- [ ] **Step 1: Add failing tests** for `--candidates=N` requiring a quality declaration, candidate/worktree status, winning merge, losing proof metadata, and cleanup after crash/pause.
|
|
359
|
+
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
360
|
+
- [ ] **Step 3: Wire candidate coordinator before merge queue** and reporter phase `selecting-candidate`.
|
|
361
|
+
- [ ] **Step 4: Run focused tests** and full suite.
|
|
362
|
+
- [ ] **Step 5: Commit** `feat(loop): expose winner-take-all candidate races`.
|
|
363
|
+
|
|
364
|
+
**Phase 6 gate:** fake providers produce one mechanically red candidate and two green candidates; blind selection integrates only one green branch and cleans all candidate worktrees.
|
|
365
|
+
|
|
366
|
+
---
|
|
367
|
+
|
|
368
|
+
## Phase 7: Ephemeral decomposition
|
|
369
|
+
|
|
370
|
+
### Task 20: Decomposition contract and planner
|
|
371
|
+
|
|
372
|
+
**Files:**
|
|
373
|
+
- Create: `src/loop/decomposition.ts`
|
|
374
|
+
- Create: `tests/loop/decomposition.test.ts`
|
|
375
|
+
- Modify: `src/agents/contracts.ts`
|
|
376
|
+
- Modify: `src/loop/runner.ts`
|
|
377
|
+
|
|
378
|
+
- [ ] **Step 1: Write failing tests** for valid subtask DAGs, duplicate IDs, cycles, unknown dependencies, area conflicts, parent story binding, result-file transport, and invalid-output fallback to whole-story execution.
|
|
379
|
+
- [ ] **Step 2: Run focused tests** and confirm failure.
|
|
380
|
+
- [ ] **Step 3: Implement schema and read-only planner call** with a versioned envelope. Persist only under `.yoke/work/<story>/`.
|
|
381
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
382
|
+
- [ ] **Step 5: Commit** `feat(loop): plan ephemeral story subtasks`.
|
|
383
|
+
|
|
384
|
+
### Task 21: Subtask scheduling and synthesis
|
|
385
|
+
|
|
386
|
+
**Files:**
|
|
387
|
+
- Modify: `src/loop/decomposition.ts`
|
|
388
|
+
- Modify: `src/loop/worker.ts`
|
|
389
|
+
- Modify: `src/loop/reporter.ts`
|
|
390
|
+
- Modify: `tests/loop/decomposition.test.ts`
|
|
391
|
+
- Modify: `tests/loop/worker.test.ts`
|
|
392
|
+
|
|
393
|
+
- [ ] **Step 1: Add failing tests** for dependency/area scheduling, separate nested worktrees, synthesis into one parent candidate, no subtask commits to main, no PRD mutation, parent-only gates, pause, and cleanup.
|
|
394
|
+
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
395
|
+
- [ ] **Step 3: Implement bounded subtask execution** reusing parallel scheduling but preserving the parent as the only story/commit unit.
|
|
396
|
+
- [ ] **Step 4: Run focused tests** and full suite.
|
|
397
|
+
- [ ] **Step 5: Commit** `feat(loop): execute decomposed work beneath one story gate`.
|
|
398
|
+
|
|
399
|
+
### Task 22: Decomposition opt-in and routing policy
|
|
400
|
+
|
|
401
|
+
**Files:**
|
|
402
|
+
- Modify: `src/retrofit/config.ts`
|
|
403
|
+
- Modify: `src/loop/run-command.ts`
|
|
404
|
+
- Modify: `src/routing/router.ts`
|
|
405
|
+
- Modify: `tests/retrofit/config.test.ts`
|
|
406
|
+
- Modify: `tests/routing/router.test.ts`
|
|
407
|
+
|
|
408
|
+
- [ ] **Step 1: Add failing tests** for decomposition disabled by default, explicit enable, no controller call for small stories, and routing-selected decomposition for qualified complex stories.
|
|
409
|
+
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
410
|
+
- [ ] **Step 3: Implement minimal policy fields and a deterministic small-story fast path.** Add `decomposition.enabled`, `decomposition.maxSubtasks`, and `decomposition.minAcceptanceCriteria` to config; skip the planner when disabled or below the threshold.
|
|
411
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
412
|
+
- [ ] **Step 5: Commit** `feat(routing): opt into cost-aware story decomposition`.
|
|
413
|
+
|
|
414
|
+
**Phase 7 gate:** one parent story decomposes into two independent and one dependent fake subtasks, synthesizes one candidate, runs parent gates once, and lands one commit without PRD additions.
|
|
415
|
+
|
|
416
|
+
---
|
|
417
|
+
|
|
418
|
+
## Phase 8: Routing, telemetry, docs, and release evidence
|
|
419
|
+
|
|
420
|
+
### Task 23: Versioned routing and visible fallback
|
|
421
|
+
|
|
422
|
+
**Files:**
|
|
423
|
+
- Modify: `src/routing/router.ts`
|
|
424
|
+
- Modify: `src/agents/contracts.ts`
|
|
425
|
+
- Modify: `tests/routing/router.test.ts`
|
|
426
|
+
|
|
427
|
+
- [ ] **Step 1: Write failing tests** for native/result-file route envelope, legacy `YOKE_ROUTE` compatibility, malformed output fallback to SELF with `fallbackReason`, and reporter evidence.
|
|
428
|
+
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
429
|
+
- [ ] **Step 3: Prefer versioned result transport** while retaining the old marker as an explicitly reported compatibility fallback.
|
|
430
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
431
|
+
- [ ] **Step 5: Commit** `feat(routing): adopt versioned decisions with visible fallback`.
|
|
432
|
+
|
|
433
|
+
### Task 24: Telemetry envelopes and raw evidence
|
|
434
|
+
|
|
435
|
+
**Files:**
|
|
436
|
+
- Modify: `src/agents/telemetry.ts`
|
|
437
|
+
- Modify: `src/loop/reporter.ts`
|
|
438
|
+
- Modify: `tests/agents/telemetry.test.ts`
|
|
439
|
+
- Modify: `tests/loop/reporter.test.ts`
|
|
440
|
+
|
|
441
|
+
- [ ] **Step 1: Write failing tests** for all three providers, unknown-field retention in per-call raw evidence, aggregate known fields, role/worker/candidate/round attribution, missing usage honesty, and cost accumulation.
|
|
442
|
+
- [ ] **Step 2: Run focused tests** and confirm failures.
|
|
443
|
+
- [ ] **Step 3: Wrap parsed events in machine envelopes** and persist raw evidence without promoting unknown values.
|
|
444
|
+
- [ ] **Step 4: Run focused tests** and confirm pass.
|
|
445
|
+
- [ ] **Step 5: Commit** `feat(telemetry): attribute quality and parallel provider calls`.
|
|
446
|
+
|
|
447
|
+
### Task 25: Canon, retrofit, and migration contracts
|
|
448
|
+
|
|
449
|
+
**Files:**
|
|
450
|
+
- Modify: `canon/loop/prd.schema.md`
|
|
451
|
+
- Modify: `canon/loop/loop-spec.md`
|
|
452
|
+
- Modify: `canon/skills/authoring-prd/SKILL.md`
|
|
453
|
+
- Modify: `canon/skills/yoke-workflow/SKILL.md`
|
|
454
|
+
- Modify: `src/retrofit/planners/shared.ts`
|
|
455
|
+
- Modify: `tests/canon/real-canon.test.ts`
|
|
456
|
+
- Modify: `tests/retrofit/retrofit.integration.test.ts`
|
|
457
|
+
- Create: `docs/MIGRATING-TO-1.4.md`
|
|
458
|
+
|
|
459
|
+
- [ ] **Step 1: Add failing canon/retrofit assertions** for quality schema guidance, runtime ignores, safe defaults, explicit unbounded warning, parallel/candidate/decomposition contracts, and attribution.
|
|
460
|
+
- [ ] **Step 2: Run** `npx vitest run tests/canon/real-canon.test.ts tests/retrofit/retrofit.integration.test.ts` and confirm failures.
|
|
461
|
+
- [ ] **Step 3: Update canonical docs and generated guidance** with exact YAML/CLI examples and migration behavior.
|
|
462
|
+
- [ ] **Step 4: Run focused tests** and `npm run yoke -- validate canon`.
|
|
463
|
+
- [ ] **Step 5: Commit** `docs(canon): define comparative quality and parallel execution`.
|
|
464
|
+
|
|
465
|
+
### Task 26: Benchmarks and end-to-end matrix
|
|
466
|
+
|
|
467
|
+
**Files:**
|
|
468
|
+
- Modify: `bench/result-schema.mjs`
|
|
469
|
+
- Modify: `bench/run.mjs`
|
|
470
|
+
- Modify: `bench/run-matrix.mjs`
|
|
471
|
+
- Add fixture files under: `bench/fixtures/quality-loop/`
|
|
472
|
+
- Modify: `bench/README.md`
|
|
473
|
+
- Add deterministic integration tests under: `tests/loop/gauntlet-cli.integration.test.ts`
|
|
474
|
+
|
|
475
|
+
- [ ] **Step 1: Add failing result-schema tests/checks** for quality rounds, consistency, reference digest, repair count, parallel workers, candidates, decomposition, conflicts, and per-role calls.
|
|
476
|
+
- [ ] **Step 2: Build a deterministic fake quality fixture** with hidden acceptance checks and scripted critic losses/wins.
|
|
477
|
+
- [ ] **Step 3: Run the fixture before wiring** and confirm expected schema/test failure.
|
|
478
|
+
- [ ] **Step 4: Extend benchmark capture and matrix arms:** quality off/on, bounded, unbounded with pause, serial/parallel, candidates one/two.
|
|
479
|
+
- [ ] **Step 5: Run deterministic matrix** and confirm all final hidden tests pass; do not claim authenticated provider performance from fake rows.
|
|
480
|
+
- [ ] **Step 6: Commit** `bench: measure quality repair and parallel execution`.
|
|
481
|
+
|
|
482
|
+
### Task 27: Product documentation and final release gates
|
|
483
|
+
|
|
484
|
+
**Files:**
|
|
485
|
+
- Modify: `README.md`
|
|
486
|
+
- Modify: `CHANGELOG.md`
|
|
487
|
+
- Modify: `TODOS.md`
|
|
488
|
+
- Modify: `docs/MIGRATING-TO-1.4.md`
|
|
489
|
+
- Modify: release metadata only when a release is explicitly requested.
|
|
490
|
+
|
|
491
|
+
- [ ] **Step 1: Update README** with the authoritative gate order, bounded defaults, explicit human-brake command, safety invariants, quality YAML, parallel/candidate/decomposition behavior, and measured caveats.
|
|
492
|
+
- [ ] **Step 2: Remove completed TODOs** for provider subprocess wiring and native schemas only after their end-to-end tests are green; retain broader authenticated samples and signed provenance until separately completed.
|
|
493
|
+
- [ ] **Step 3: Run final static gates:** `npm run lint`, `npm run build`, `npm test`, `npm run yoke -- validate canon`, `npm run docs:check`, `npm run audit:ci`, `npm run package:check`.
|
|
494
|
+
- [ ] **Step 4: Run manual CLI QA** in a temporary Git fixture: help text; invalid flag; bounded reject/repair/pass; quality-unbounded then pause; parallel independent stories; two-candidate winner; decomposition fallback; status/cleanup after killed fake provider.
|
|
495
|
+
- [ ] **Step 5: Review `git diff --check`, `git status --short`, and the complete diff.** Verify `.omo/` and unrelated concurrent changes remain untouched.
|
|
496
|
+
- [ ] **Step 6: Commit** only if explicitly requested: `feat: add gated comparative quality loops`.
|
|
497
|
+
|
|
498
|
+
---
|
|
499
|
+
|
|
500
|
+
## Dependency graph and parallel work
|
|
501
|
+
|
|
502
|
+
```text
|
|
503
|
+
Phase 1 contracts
|
|
504
|
+
-> Phase 2 repair loop
|
|
505
|
+
-> Phase 3 reference judging
|
|
506
|
+
-> Phase 4 CLI/unbounded
|
|
507
|
+
|
|
508
|
+
Phase 1 contracts
|
|
509
|
+
-> Phase 5 provider parallelism
|
|
510
|
+
-> Phase 6 candidate races
|
|
511
|
+
|
|
512
|
+
Phase 5 worker lifecycle
|
|
513
|
+
-> Phase 7 decomposition
|
|
514
|
+
|
|
515
|
+
Phases 2-7
|
|
516
|
+
-> Phase 8 telemetry/docs/benchmarks
|
|
517
|
+
```
|
|
518
|
+
|
|
519
|
+
Within Phase 1, Tasks 1 and the initial tests for Task 3 can proceed independently. In Phase 3,
|
|
520
|
+
reference acquisition and artifact collection are independent until critic integration. In Phase 5,
|
|
521
|
+
async provider work can proceed alongside rich claim tests after their shared handle contract is
|
|
522
|
+
agreed. All modifications to `loop.ts`, `run-command.ts`, `runner.ts`, and `reporter.ts` should remain
|
|
523
|
+
serialized to avoid conflicting edits.
|
|
524
|
+
|
|
525
|
+
## Plan self-review
|
|
526
|
+
|
|
527
|
+
- **Spec coverage:** Tasks 4-6 cover repair; 7-11 references/blind judging; 14-17 parallel stories;
|
|
528
|
+
18-19 candidate races; 20-22 decomposition; 1-3 and 23-24 versioned contracts; 12-13 explicit
|
|
529
|
+
bounded/unbounded CLI and operational behavior. Security, observability, attribution, benchmarks,
|
|
530
|
+
canon, migration, and manual QA are covered in Tasks 7, 10, 13, and 23-27.
|
|
531
|
+
- **Compatibility:** Every new schema field is optional; serial and no-quality paths receive explicit
|
|
532
|
+
regression tests; unbounded is a per-run explicit escalation.
|
|
533
|
+
- **Type consistency:** `MachineEnvelope`, `QualityVerdict`, `RepairGap`, `ResolvedQualityPolicy`, and
|
|
534
|
+
worker/candidate/decomposition outcomes are introduced before their consumers.
|
|
535
|
+
- **No hidden weakening:** All repairs rerun gates; workers cannot set PRD pass state; merge queue owns
|
|
536
|
+
integrated success; subjective verdicts can reject but never override mechanical failure.
|
|
537
|
+
- **No placeholders:** Every task has exact files, expected behavior, commands, and a commit boundary.
|