@hecer/yoke 1.6.0 → 1.6.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +13 -13
- package/.codex-plugin/plugin.json +7 -7
- package/CHANGELOG.md +294 -288
- package/README.md +874 -874
- package/TODOS.md +5 -5
- package/agents/docs.toml +6 -6
- package/agents/implementer.toml +6 -6
- package/agents/reviewer.toml +6 -6
- package/agents/security.toml +6 -6
- package/bench/README.md +86 -86
- package/bench/RESULTS.md +35 -35
- package/bench/output-compaction.mjs +65 -65
- package/bench/result-schema.mjs +12 -12
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
- package/bench/results/codex-unavailable-1785175418318.json +15 -15
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
- package/bench/run-matrix.mjs +26 -26
- package/bench/run.mjs +106 -106
- package/canon/AGENTS.md +30 -30
- package/canon/context/DECISIONS.md +4 -4
- package/canon/context/GLOSSARY.md +11 -11
- package/canon/context/KNOWLEDGE.md +4 -4
- package/canon/context/PROJECT.md +15 -15
- package/canon/loop/loop-spec.md +65 -65
- package/canon/loop/prd.schema.md +43 -43
- package/canon/manifest.yaml +59 -59
- package/canon/policy/gates.md +7 -7
- package/canon/policy/roles.md +9 -9
- package/canon/skills/ATTRIBUTION.md +99 -99
- package/canon/skills/authoring-prd/SKILL.md +58 -58
- package/canon/skills/brainstorming/SKILL.md +164 -164
- package/canon/skills/codebase-design/DEEPENING.md +15 -15
- package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
- package/canon/skills/codebase-design/SKILL.md +39 -39
- package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
- package/canon/skills/document-release/SKILL.md +302 -302
- package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
- package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
- package/canon/skills/domain-modeling/SKILL.md +35 -35
- package/canon/skills/executing-plans/SKILL.md +70 -70
- package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
- package/canon/skills/health/SKILL.md +177 -177
- package/canon/skills/maintaining-context/SKILL.md +34 -34
- package/canon/skills/minimal-code/SKILL.md +21 -21
- package/canon/skills/no-ai-slop/SKILL.md +103 -103
- package/canon/skills/no-ai-slop/eval.md +43 -43
- package/canon/skills/plan-ceo-review/SKILL.md +541 -541
- package/canon/skills/plan-eng-review/SKILL.md +362 -362
- package/canon/skills/receiving-code-review/SKILL.md +213 -213
- package/canon/skills/requesting-code-review/SKILL.md +105 -105
- package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
- package/canon/skills/retro/SKILL.md +397 -397
- package/canon/skills/review/SKILL.md +246 -246
- package/canon/skills/ship/SKILL.md +691 -691
- package/canon/skills/subagent-driven-development/SKILL.md +277 -277
- package/canon/skills/systematic-debugging/SKILL.md +296 -296
- package/canon/skills/tdd/SKILL.md +371 -371
- package/canon/skills/unslop-ui/SKILL.md +34 -34
- package/canon/skills/using-git-worktrees/SKILL.md +218 -218
- package/canon/skills/verification-before-completion/SKILL.md +139 -139
- package/canon/skills/visual-verification/SKILL.md +54 -54
- package/canon/skills/workflow/SKILL.md +22 -22
- package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
- package/canon/skills/writing-for-agents/SKILL.md +42 -42
- package/canon/skills/writing-plans/SKILL.md +152 -152
- package/canon/skills/writing-skills/SKILL.md +655 -655
- package/canon/skills/yoke-retrofit/SKILL.md +26 -26
- package/canon/skills/yoke-workflow/SKILL.md +20 -20
- package/canon/tools/codex-rtk-hook.mjs +35 -35
- package/canon/tools/graphify.md +3 -3
- package/canon/tools/playwright-mcp.md +3 -3
- package/canon/tools/rtk.md +7 -7
- package/canon/tools/serena.md +6 -6
- package/dist/agents/process.js +3 -0
- package/dist/loop/watchdog.js +1 -1
- package/dist/prd/command.js +17 -17
- package/dist/retrofit/planners/claude.js +14 -14
- package/dist/retrofit/preserve.js +2 -2
- package/docs/MIGRATING-TO-1.0.md +33 -33
- package/docs/MIGRATING-TO-1.1.md +27 -27
- package/docs/MIGRATING-TO-1.4.md +70 -70
- package/docs/PUBLISHING.md +91 -91
- package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
- package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
- package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
- package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
- package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
- package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
- package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
- package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
- package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
- package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
- package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
- package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
- package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
- package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
- package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
- package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
- package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
- package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
- package/gemini-extension.json +6 -6
- package/hooks/hooks.json +19 -19
- package/package.json +87 -87
|
@@ -1,422 +1,422 @@
|
|
|
1
|
-
# Gauntlet quality loop for Yoke
|
|
2
|
-
|
|
3
|
-
**Status:** Approved direction, ready for implementation planning
|
|
4
|
-
**Scope:** All six Gauntlet-inspired improvements, including an explicitly enabled unbounded quality mode
|
|
5
|
-
|
|
6
|
-
## Goal
|
|
7
|
-
|
|
8
|
-
Extend Yoke's mechanically gated story loop with comparative quality iteration without weakening
|
|
9
|
-
its existing authority model. Yoke remains the control plane: executable acceptance evidence,
|
|
10
|
-
project verification, performance, audit, independent review, worktree isolation, commit integrity,
|
|
11
|
-
locks, recovery, pause, and watchdog behavior remain authoritative.
|
|
12
|
-
|
|
13
|
-
The new quality layer adds:
|
|
14
|
-
|
|
15
|
-
1. a reviewer-to-repair loop;
|
|
16
|
-
2. optional external quality references;
|
|
17
|
-
3. blind binary comparison rather than drifting scores;
|
|
18
|
-
4. provider-backed parallel story execution;
|
|
19
|
-
5. ephemeral decomposition within a story; and
|
|
20
|
-
6. versioned provider and critic result contracts with provenance.
|
|
21
|
-
|
|
22
|
-
An explicit unbounded mode may remove the quality-loop round and elapsed-time limits. It does not
|
|
23
|
-
disable any mechanical gate or operational safety mechanism. In that mode the human is the only
|
|
24
|
-
quality brake, while process failures and Yoke safety gates can still block or pause the run.
|
|
25
|
-
|
|
26
|
-
## Non-goals
|
|
27
|
-
|
|
28
|
-
- Do not replace the PRD, acceptance criteria, or project verification with model judgment.
|
|
29
|
-
- Do not introduce a second `yoke gauntlet` state machine competing with `yoke loop`.
|
|
30
|
-
- Do not let ephemeral subtasks rewrite the authoritative PRD.
|
|
31
|
-
- Do not run parallel workers in a shared mutable worktree.
|
|
32
|
-
- Do not merge competing implementations of the same story together.
|
|
33
|
-
- Do not make subjective quality checks run by default for stories that do not declare them.
|
|
34
|
-
- Do not copy Gauntlet Loop prompt text. Reimplement the concepts in Yoke's native contracts.
|
|
35
|
-
|
|
36
|
-
## Compatibility
|
|
37
|
-
|
|
38
|
-
- Existing PRDs and `.yoke/config.yaml` files remain valid.
|
|
39
|
-
- Existing serial loop behavior remains the default.
|
|
40
|
-
- Quality iteration is enabled per story or by an explicit run override.
|
|
41
|
-
- `--parallel=1` remains equivalent to isolated serial execution. Values above one become available
|
|
42
|
-
only through the provider-backed dispatcher described below.
|
|
43
|
-
- Existing exit codes retain their meanings. Quality exhaustion and inconsistent judging are
|
|
44
|
-
ordinary blocked states and therefore return the existing blocked exit code.
|
|
45
|
-
- Existing standalone `yoke review` continues to work. Its verdict schema gains optional fields but
|
|
46
|
-
remains backward compatible with current verdict files.
|
|
47
|
-
|
|
48
|
-
## Authority and invariants
|
|
49
|
-
|
|
50
|
-
The gate order for every candidate is:
|
|
51
|
-
|
|
52
|
-
```text
|
|
53
|
-
implement or repair
|
|
54
|
-
-> structured criterion commands
|
|
55
|
-
-> project verify
|
|
56
|
-
-> performance gate when configured
|
|
57
|
-
-> audit gate when configured
|
|
58
|
-
-> quality challenge when enabled
|
|
59
|
-
-> independent code review when enabled
|
|
60
|
-
-> commit and passes:true atomically
|
|
61
|
-
```
|
|
62
|
-
|
|
63
|
-
The following invariants are absolute:
|
|
64
|
-
|
|
65
|
-
1. A quality verdict can reject mechanically green work but can never approve mechanically red work.
|
|
66
|
-
2. Every repair reruns the complete gate sequence from criterion evidence onward.
|
|
67
|
-
3. A fresh critic evaluates every new candidate. Previous critic explanations are not included in
|
|
68
|
-
the next critic prompt; only the repair worker receives the selected gap.
|
|
69
|
-
4. `passes: true` is written only after integrated-tree verification and a successful commit.
|
|
70
|
-
5. Unbounded mode removes only `maxRounds` and `maxMinutes` from the quality iteration. Watchdog,
|
|
71
|
-
pause, cleanup, locks, worktree isolation, provider failures, audit, and all verification remain.
|
|
72
|
-
6. Malformed, absent, inconsistent, or unverifiable quality evidence never becomes a pass.
|
|
73
|
-
|
|
74
|
-
## Configuration and PRD schema
|
|
75
|
-
|
|
76
|
-
### Project defaults
|
|
77
|
-
|
|
78
|
-
`.yoke/config.yaml` gains optional defaults:
|
|
79
|
-
|
|
80
|
-
```yaml
|
|
81
|
-
quality:
|
|
82
|
-
enabled: false
|
|
83
|
-
policy: blocking # blocking | advisory
|
|
84
|
-
maxRounds: 3
|
|
85
|
-
maxMinutes: 60
|
|
86
|
-
consistencyChecks: 2 # normal order plus label swap
|
|
87
|
-
maxParallelCandidates: 2
|
|
88
|
-
critic:
|
|
89
|
-
agent: codex # optional; independent resolution otherwise
|
|
90
|
-
model: gpt-5.6-sol # opaque provider model string
|
|
91
|
-
reasoningEffort: high
|
|
92
|
-
repair:
|
|
93
|
-
agent: claude # optional; story runner otherwise
|
|
94
|
-
```
|
|
95
|
-
|
|
96
|
-
Defaults apply only when a story declares a quality challenge or the run explicitly enables one.
|
|
97
|
-
`enabled: true` means stories with a complete `quality` declaration run it automatically; it does
|
|
98
|
-
not invent references for ordinary stories.
|
|
99
|
-
|
|
100
|
-
### Story declaration
|
|
101
|
-
|
|
102
|
-
`StorySchema` gains an optional `quality` object:
|
|
103
|
-
|
|
104
|
-
```yaml
|
|
105
|
-
- id: STORY-4
|
|
106
|
-
title: Polish the pricing page
|
|
107
|
-
priority: 4
|
|
108
|
-
acceptance: [...]
|
|
109
|
-
passes: false
|
|
110
|
-
quality:
|
|
111
|
-
reference:
|
|
112
|
-
name: Stripe pricing page
|
|
113
|
-
source: https://stripe.com/pricing
|
|
114
|
-
kind: url # url | file | command
|
|
115
|
-
digest: sha256:... # required for blocking file snapshots; acquired for URLs
|
|
116
|
-
candidate:
|
|
117
|
-
kind: screenshots # screenshots | files | command-output | benchmark
|
|
118
|
-
paths:
|
|
119
|
-
- .yoke/proof/STORY-4/desktop.png
|
|
120
|
-
- .yoke/proof/STORY-4/mobile.png
|
|
121
|
-
rubric: Visual hierarchy, clarity, responsive finish, and interaction quality
|
|
122
|
-
policy: blocking # optional override
|
|
123
|
-
```
|
|
124
|
-
|
|
125
|
-
Validation requires a non-empty, named, fetchable, and comparable reference; a concrete candidate
|
|
126
|
-
artifact contract; and a non-empty rubric. Blocking comparisons cannot use mutable live content as
|
|
127
|
-
their only retained evidence. URL references are fetched during preflight, converted to inert
|
|
128
|
-
artifacts, hashed, timestamped, and stored under `.yoke/references/<digest>/`. The runtime artifact
|
|
129
|
-
is ignored by Git by default; the provenance record is written into story proof evidence.
|
|
130
|
-
|
|
131
|
-
`command` sources and candidate commands use argv arrays in the internal schema. The public YAML
|
|
132
|
-
may use a command string only through the same command validation policy as existing proof commands;
|
|
133
|
-
shell control operators are rejected.
|
|
134
|
-
|
|
135
|
-
### CLI overrides
|
|
136
|
-
|
|
137
|
-
`yoke loop run` gains:
|
|
138
|
-
|
|
139
|
-
```text
|
|
140
|
-
--quality enable declared quality challenges for this run
|
|
141
|
-
--no-quality disable them for this run
|
|
142
|
-
--quality-rounds=N override maxRounds
|
|
143
|
-
--quality-minutes=N override maxMinutes
|
|
144
|
-
--quality-unbounded remove quality round/time limits explicitly
|
|
145
|
-
--quality-policy=blocking|advisory
|
|
146
|
-
--parallel=N run independent stories concurrently
|
|
147
|
-
--candidates=N run competing candidates for one quality-enabled story
|
|
148
|
-
```
|
|
149
|
-
|
|
150
|
-
`--quality-unbounded` implies `--quality`, prints a prominent startup warning, and is recorded in
|
|
151
|
-
status/proof provenance. It conflicts with `--quality-rounds` and `--quality-minutes`. It does not
|
|
152
|
-
imply `--unsafe`, does not disable `--timeout`, and does not change `--parallel` or candidate count.
|
|
153
|
-
|
|
154
|
-
## 1. Reviewer-to-repair transition
|
|
155
|
-
|
|
156
|
-
Current reviewer rejection ends the story. The loop instead classifies the failure:
|
|
157
|
-
|
|
158
|
-
- mechanical gate failure: block immediately; no quality repair is attempted;
|
|
159
|
-
- quality challenge loss: select the verdict's single `biggestGap`;
|
|
160
|
-
- code-review rejection: select the highest-severity actionable blocking finding;
|
|
161
|
-
- malformed reviewer output or provider failure: block as infrastructure failure, not repair work.
|
|
162
|
-
|
|
163
|
-
When an actionable gap exists and the quality budget permits another round, Yoke launches a fresh
|
|
164
|
-
repair worker in the same isolated story worktree. Its prompt contains:
|
|
165
|
-
|
|
166
|
-
- story title and acceptance criteria;
|
|
167
|
-
- settled project context;
|
|
168
|
-
- the current diff;
|
|
169
|
-
- exactly one selected gap and its cited evidence;
|
|
170
|
-
- the instruction to make the smallest change that closes that gap;
|
|
171
|
-
- the immutable mechanical gate commands.
|
|
172
|
-
|
|
173
|
-
It does not receive critic praise, hidden labels, earlier critic deliberation, or unrelated findings.
|
|
174
|
-
After repair, criterion evidence is replaced with evidence from the new candidate and all gates rerun.
|
|
175
|
-
|
|
176
|
-
Default exhaustion is three repair rounds or 60 elapsed quality minutes per story, whichever occurs
|
|
177
|
-
first. Exhaustion writes a blocked status naming the remaining gap and the exhausted budget. In
|
|
178
|
-
unbounded mode, quality rounds continue until a consistent pass, human pause/stop, or a non-quality
|
|
179
|
-
failure blocks the story.
|
|
180
|
-
|
|
181
|
-
## 2. External quality references
|
|
182
|
-
|
|
183
|
-
A new `src/quality/` boundary owns reference acquisition, artifact preparation, comparison, and
|
|
184
|
-
evidence. It does not own story selection, commits, or process lifecycle.
|
|
185
|
-
|
|
186
|
-
Reference preflight occurs before implementation for a blocking challenge so an unreachable bar does
|
|
187
|
-
not waste a full agent run. Acquisition records:
|
|
188
|
-
|
|
189
|
-
- declared source and display name;
|
|
190
|
-
- resolved URL or path;
|
|
191
|
-
- acquisition timestamp;
|
|
192
|
-
- media/content type and byte size;
|
|
193
|
-
- SHA-256 digest;
|
|
194
|
-
- tool/adapter version;
|
|
195
|
-
- any viewport, benchmark, or normalization parameters.
|
|
196
|
-
|
|
197
|
-
Fetched content is untrusted data. HTML is rendered or converted to static evidence; scripts and text
|
|
198
|
-
from the reference are never inserted as instructions. Size and content-type limits apply. Critics
|
|
199
|
-
receive artifacts plus the trusted Yoke rubric in separate, explicitly delimited sections and run
|
|
200
|
-
read-only without source-tree write access.
|
|
201
|
-
|
|
202
|
-
For visual stories, existing `flow-smoke` screenshots are reused when paths match the declared
|
|
203
|
-
candidate contract. Quality comparison never triggers a second capture farm. Nonvisual adapters
|
|
204
|
-
compare declared files, command output, or benchmark records.
|
|
205
|
-
|
|
206
|
-
If a required reference cannot be acquired or its digest changes during the run, blocking policy
|
|
207
|
-
blocks before commit. Advisory policy records `skipped` with the exact cause and continues; skipped is
|
|
208
|
-
never reported as passed.
|
|
209
|
-
|
|
210
|
-
## 3. Blind binary comparison
|
|
211
|
-
|
|
212
|
-
Quality verdict schema version 1:
|
|
213
|
-
|
|
214
|
-
```json
|
|
215
|
-
{
|
|
216
|
-
"schemaVersion": 1,
|
|
217
|
-
"verdict": "candidate|reference",
|
|
218
|
-
"biggestGap": "single actionable gap",
|
|
219
|
-
"evidence": ["artifact-relative citation"],
|
|
220
|
-
"confidence": "high|medium|low",
|
|
221
|
-
"labels": { "candidate": "A", "reference": "B" },
|
|
222
|
-
"provenance": {
|
|
223
|
-
"provider": "codex",
|
|
224
|
-
"model": "gpt-5.6-sol",
|
|
225
|
-
"promptVersion": 1,
|
|
226
|
-
"rubricDigest": "sha256:...",
|
|
227
|
-
"referenceDigest": "sha256:...",
|
|
228
|
-
"candidateDigest": "sha256:..."
|
|
229
|
-
}
|
|
230
|
-
}
|
|
231
|
-
```
|
|
232
|
-
|
|
233
|
-
A blocking pass requires two fresh comparisons: one randomized A/B ordering and one label-swapped
|
|
234
|
-
ordering. Both must select the candidate. A disagreement, missing evidence, low-confidence verdict,
|
|
235
|
-
schema failure, or digest mismatch is `inconsistent`, not a pass. It blocks with retained evidence so
|
|
236
|
-
the user can rerun, change reviewer, or downgrade the story policy explicitly.
|
|
237
|
-
|
|
238
|
-
The critic names exactly one biggest remaining gap. Scores out of ten are not part of the contract.
|
|
239
|
-
Advisory mode records the same evidence but cannot prevent a mechanically valid story from landing.
|
|
240
|
-
|
|
241
|
-
## 4. Provider-backed parallel execution
|
|
242
|
-
|
|
243
|
-
The existing scheduler, claims, `runParallelLoop`, and `MergeQueue` become the production CLI path.
|
|
244
|
-
Parallel mode requires isolation.
|
|
245
|
-
|
|
246
|
-
### Independent stories
|
|
247
|
-
|
|
248
|
-
1. The dispatcher owns the existing loop lock.
|
|
249
|
-
2. Ready stories are selected by `needs`; duplicate active `area` values are excluded.
|
|
250
|
-
3. Each worker atomically claims one story and creates a project-owned worktree from the current HEAD.
|
|
251
|
-
4. It runs implementation and all per-story gates, including quality repair rounds.
|
|
252
|
-
5. A green worker enters the serialized merge queue.
|
|
253
|
-
6. The integrator rebases onto current HEAD, reruns criterion/project/perf/audit gates against the
|
|
254
|
-
integrated candidate, then integrates, updates context/PRD, and commits atomically.
|
|
255
|
-
7. A rebase conflict or integrated-tree failure releases the claim and reopens the story with a
|
|
256
|
-
structured reason. It never sets `passes: true`.
|
|
257
|
-
|
|
258
|
-
Provider subprocess workers expose cancellation and watchdog handles to scoped cleanup. Claim files
|
|
259
|
-
gain dispatcher ID, PID, base commit, worktree, provider/model, and heartbeat. Cleanup only removes
|
|
260
|
-
resources owned by the target project and dead dispatcher.
|
|
261
|
-
|
|
262
|
-
### Competing candidates for one story
|
|
263
|
-
|
|
264
|
-
`--candidates=N` creates N worktrees from the same base and runs independent implementations. Each
|
|
265
|
-
must pass mechanical gates. Quality comparison then selects one mechanically green candidate against
|
|
266
|
-
the reference; if more than one beats the reference, a final blind candidate-vs-candidate comparison
|
|
267
|
-
chooses the winner. Only the winning branch enters the merge queue. Losing candidates are retained as
|
|
268
|
-
proof metadata, then their worktrees are cleaned. Candidate branches are never combined.
|
|
269
|
-
|
|
270
|
-
Concurrency and candidate fan-out are separately limited. Defaults remain one story worker and one
|
|
271
|
-
candidate. Provider subprocess wiring ships only when cancellation, crash recovery, conflict,
|
|
272
|
-
integrated reverify, and cost telemetry tests pass end to end.
|
|
273
|
-
|
|
274
|
-
## 5. Ephemeral story decomposition
|
|
275
|
-
|
|
276
|
-
Before implementation, a worker may emit a versioned decomposition artifact:
|
|
277
|
-
|
|
278
|
-
```json
|
|
279
|
-
{
|
|
280
|
-
"schemaVersion": 1,
|
|
281
|
-
"storyId": "STORY-4",
|
|
282
|
-
"subtasks": [
|
|
283
|
-
{ "id": "hero", "goal": "Improve hero hierarchy", "area": "ui", "needs": [] },
|
|
284
|
-
{ "id": "mobile", "goal": "Polish mobile layout", "area": "ui-mobile", "needs": ["hero"] }
|
|
285
|
-
]
|
|
286
|
-
}
|
|
287
|
-
```
|
|
288
|
-
|
|
289
|
-
Subtasks are runtime scheduling hints under `.yoke/work/<story>/`; they are not PRD stories, cannot
|
|
290
|
-
change acceptance criteria, and cannot independently set `passes`. The parent story remains the only
|
|
291
|
-
commit and quality unit. Subtasks may run concurrently only when their declared areas differ and each
|
|
292
|
-
has its own nested worktree/branch or non-overlapping artifact workspace. Their results are synthesized
|
|
293
|
-
into one parent candidate before any gate runs.
|
|
294
|
-
|
|
295
|
-
Invalid decomposition falls back to one whole-story worker. Dynamic decomposition is optional and
|
|
296
|
-
must not add a controller call for small stories unless explicitly enabled or routing selects it.
|
|
297
|
-
|
|
298
|
-
## 6. Versioned provider and critic contracts
|
|
299
|
-
|
|
300
|
-
Provider adapters own every machine result. Add schemas for:
|
|
301
|
-
|
|
302
|
-
- route decisions;
|
|
303
|
-
- decomposition plans;
|
|
304
|
-
- review verdicts and selected repair gaps;
|
|
305
|
-
- quality comparison verdicts;
|
|
306
|
-
- candidate selection;
|
|
307
|
-
- provider telemetry envelopes.
|
|
308
|
-
|
|
309
|
-
Every envelope carries `schemaVersion`, provider, reported model when available, invocation role,
|
|
310
|
-
prompt/contract version, timing, permission profile, and usage fields. Unknown optional telemetry is
|
|
311
|
-
retained in raw per-call evidence but not promoted to aggregate fields without schema support.
|
|
312
|
-
|
|
313
|
-
Native provider output schemas are used when stable and available. Otherwise Yoke's existing
|
|
314
|
-
result-file transport is the compatibility path. Free-form stdout markers such as `YOKE_ROUTE` remain
|
|
315
|
-
temporary backward compatibility only; malformed output visibly falls back to the strong parent and
|
|
316
|
-
records `fallbackReason`, rather than silently disappearing.
|
|
317
|
-
|
|
318
|
-
All status NDJSON and proof evidence include quality round, candidate, critic, reference digest,
|
|
319
|
-
selected gap, verdict consistency, parallel worker, merge outcome, and cumulative token/cost data.
|
|
320
|
-
|
|
321
|
-
## State and observability
|
|
322
|
-
|
|
323
|
-
Runtime artifacts:
|
|
324
|
-
|
|
325
|
-
```text
|
|
326
|
-
.yoke/references/<digest>/... acquired inert reference
|
|
327
|
-
.yoke/work/<story>/decomposition.json ephemeral subtask plan
|
|
328
|
-
.yoke/proof/<story>/quality/round-<n>/ candidate and critic evidence
|
|
329
|
-
.yoke/proof/<story>/quality/summary.json final quality outcome
|
|
330
|
-
.yoke/claims/<story>.json parallel ownership/heartbeat
|
|
331
|
-
.yoke/loop-status.json current round/candidate/worker
|
|
332
|
-
.yoke/loop.log phase transitions and reasons
|
|
333
|
-
```
|
|
334
|
-
|
|
335
|
-
New reporter phases are `decomposing`, `quality-preflight`, `comparing`, `repairing`, `integrating`,
|
|
336
|
-
and `selecting-candidate`. Status includes bounded/unbounded mode, current/maximum quality round,
|
|
337
|
-
elapsed quality time, reference digest, active workers, and candidate count. Logs stay bounded under
|
|
338
|
-
the existing policy; structured proof artifacts are per-story and never inferred from prose.
|
|
339
|
-
|
|
340
|
-
Pause is honored at safe boundaries: between quality rounds and before launching new parallel workers.
|
|
341
|
-
An in-flight provider invocation is allowed to finish unless cleanup explicitly terminates its scoped
|
|
342
|
-
PID tree. In unbounded mode, `yoke loop status` always shows that only a human pause/stop can end
|
|
343
|
-
quality iteration.
|
|
344
|
-
|
|
345
|
-
## Error and recovery map
|
|
346
|
-
|
|
347
|
-
| Failure | Result |
|
|
348
|
-
|---|---|
|
|
349
|
-
| Mechanical criterion/verify/perf/audit failure | Block story; no quality repair |
|
|
350
|
-
| Quality critic says reference wins | Repair next round when budget allows |
|
|
351
|
-
| Code reviewer returns actionable blocking finding | Repair next round when budget allows |
|
|
352
|
-
| Missing/malformed verdict | Block as reviewer infrastructure failure |
|
|
353
|
-
| Label-swap disagreement or low confidence | Block as inconsistent judge evidence |
|
|
354
|
-
| Reference unavailable before work | Block, or advisory skip when explicitly configured |
|
|
355
|
-
| Reference digest changes mid-run | Block and retain both provenance records |
|
|
356
|
-
| Repair budget exhausted | Block with last gap and budget evidence |
|
|
357
|
-
| Human pause during unbounded mode | Finish current safe unit, persist state, exit paused |
|
|
358
|
-
| Provider crash/timeout | Block worker, preserve logs, release dead claim during cleanup |
|
|
359
|
-
| Parallel rebase conflict | Reopen story with conflict reason; no integration |
|
|
360
|
-
| Integrated-tree verification failure | Reopen story; candidate never passes |
|
|
361
|
-
| Candidate fan-out has no mechanically green candidate | Block without subjective comparison |
|
|
362
|
-
|
|
363
|
-
## Security
|
|
364
|
-
|
|
365
|
-
- Critics and reference evaluators run read-only.
|
|
366
|
-
- External reference content is never trusted as instructions.
|
|
367
|
-
- URL fetching uses protocol allowlists, redirect limits, byte limits, and loopback/private-network
|
|
368
|
-
restrictions unless the user explicitly supplies a local URL for a local project.
|
|
369
|
-
- Reference and candidate paths are normalized beneath Yoke-owned roots; traversal is rejected.
|
|
370
|
-
- Commands use validated argv contracts and existing approved-test restrictions where applicable.
|
|
371
|
-
- Unbounded mode requires the explicit CLI flag on every run; it is not persisted as a default.
|
|
372
|
-
- Startup output names effective provider, permissions, quality policy, budgets, and unbounded state.
|
|
373
|
-
- Token/cost growth is observable. Unbounded removes quality stop budgets by explicit request, but it
|
|
374
|
-
does not hide or stop reporting spend.
|
|
375
|
-
|
|
376
|
-
## Testing strategy
|
|
377
|
-
|
|
378
|
-
Development follows red-green-refactor. Add focused suites for:
|
|
379
|
-
|
|
380
|
-
1. Config and PRD compatibility, quality declaration validation, and CLI conflicts.
|
|
381
|
-
2. Repair-loop transitions, complete gate reruns, fresh reviewer use, and single-gap prompts.
|
|
382
|
-
3. Three-round/60-minute exhaustion and explicit unbounded behavior.
|
|
383
|
-
4. Reference acquisition, hashing, mutable-reference detection, inert conversion, path and network
|
|
384
|
-
security, and advisory-vs-blocking failures.
|
|
385
|
-
5. Randomized blind labels, mandatory label swap, inconsistent verdicts, schema failures, and proof
|
|
386
|
-
provenance.
|
|
387
|
-
6. Real provider subprocess workers using deterministic fake CLIs, including cancellation, timeout,
|
|
388
|
-
stale claims, process cleanup, and unavailable providers.
|
|
389
|
-
7. Parallel story scheduling, area exclusion, dependency ordering, rebase conflicts, integrated-tree
|
|
390
|
-
verification, and atomic PRD/context commits in temporary Git repositories.
|
|
391
|
-
8. Competing candidates: common base, mechanical filtering, winner-only merge, and losing cleanup.
|
|
392
|
-
9. Ephemeral decomposition validation, dependency/area scheduling, synthesis, fallback, and proof that
|
|
393
|
-
it cannot mutate PRD acceptance or pass state.
|
|
394
|
-
10. Versioned route/review/quality/telemetry envelopes and visible fallback reasons.
|
|
395
|
-
11. Reporter/NDJSON status for quality rounds, unbounded mode, parallel workers, and merge outcomes.
|
|
396
|
-
12. End-to-end serial, bounded-quality, unbounded-pause, parallel-story, and candidate-race flows.
|
|
397
|
-
|
|
398
|
-
Final verification runs TypeScript build, the complete Vitest suite, canon validation, docs metadata
|
|
399
|
-
check, package dry run, deterministic fake-provider flows, and authenticated opt-in smoke matrices for
|
|
400
|
-
Claude, Codex, and Gemini. Benchmarks compare quality off/on, bounded/unbounded-paused, serial/parallel,
|
|
401
|
-
and one/multiple candidates. Claims remain workload-specific and report unavailable provider evidence
|
|
402
|
-
honestly.
|
|
403
|
-
|
|
404
|
-
## Delivery order
|
|
405
|
-
|
|
406
|
-
1. Versioned result envelopes and review-gap selection.
|
|
407
|
-
2. Bounded reviewer-to-repair loop with status and evidence.
|
|
408
|
-
3. Reference/candidate schemas, acquisition, and blind consistency comparison.
|
|
409
|
-
4. CLI/config integration including explicit unbounded mode.
|
|
410
|
-
5. Provider subprocess workers wired to scheduler, claims, and merge queue.
|
|
411
|
-
6. Competing candidate mode.
|
|
412
|
-
7. Ephemeral decomposition and subtask scheduling.
|
|
413
|
-
8. Cross-provider benchmarks, documentation, migration notes, and release provenance updates.
|
|
414
|
-
|
|
415
|
-
Each stage preserves the serial non-quality path and lands only with its focused tests plus the full
|
|
416
|
-
existing suite green.
|
|
417
|
-
|
|
418
|
-
## Attribution
|
|
419
|
-
|
|
420
|
-
The comparative quality concepts are inspired by Matt Shumer's Gauntlet Loop technique and the
|
|
421
|
-
`robonuggets/gauntlet-loop` packaging. Yoke reimplements the concepts in its own mechanical control
|
|
422
|
-
plane. Documentation credits the inspiration; no CC BY prompt text is copied into runtime assets.
|
|
1
|
+
# Gauntlet quality loop for Yoke
|
|
2
|
+
|
|
3
|
+
**Status:** Approved direction, ready for implementation planning
|
|
4
|
+
**Scope:** All six Gauntlet-inspired improvements, including an explicitly enabled unbounded quality mode
|
|
5
|
+
|
|
6
|
+
## Goal
|
|
7
|
+
|
|
8
|
+
Extend Yoke's mechanically gated story loop with comparative quality iteration without weakening
|
|
9
|
+
its existing authority model. Yoke remains the control plane: executable acceptance evidence,
|
|
10
|
+
project verification, performance, audit, independent review, worktree isolation, commit integrity,
|
|
11
|
+
locks, recovery, pause, and watchdog behavior remain authoritative.
|
|
12
|
+
|
|
13
|
+
The new quality layer adds:
|
|
14
|
+
|
|
15
|
+
1. a reviewer-to-repair loop;
|
|
16
|
+
2. optional external quality references;
|
|
17
|
+
3. blind binary comparison rather than drifting scores;
|
|
18
|
+
4. provider-backed parallel story execution;
|
|
19
|
+
5. ephemeral decomposition within a story; and
|
|
20
|
+
6. versioned provider and critic result contracts with provenance.
|
|
21
|
+
|
|
22
|
+
An explicit unbounded mode may remove the quality-loop round and elapsed-time limits. It does not
|
|
23
|
+
disable any mechanical gate or operational safety mechanism. In that mode the human is the only
|
|
24
|
+
quality brake, while process failures and Yoke safety gates can still block or pause the run.
|
|
25
|
+
|
|
26
|
+
## Non-goals
|
|
27
|
+
|
|
28
|
+
- Do not replace the PRD, acceptance criteria, or project verification with model judgment.
|
|
29
|
+
- Do not introduce a second `yoke gauntlet` state machine competing with `yoke loop`.
|
|
30
|
+
- Do not let ephemeral subtasks rewrite the authoritative PRD.
|
|
31
|
+
- Do not run parallel workers in a shared mutable worktree.
|
|
32
|
+
- Do not merge competing implementations of the same story together.
|
|
33
|
+
- Do not make subjective quality checks run by default for stories that do not declare them.
|
|
34
|
+
- Do not copy Gauntlet Loop prompt text. Reimplement the concepts in Yoke's native contracts.
|
|
35
|
+
|
|
36
|
+
## Compatibility
|
|
37
|
+
|
|
38
|
+
- Existing PRDs and `.yoke/config.yaml` files remain valid.
|
|
39
|
+
- Existing serial loop behavior remains the default.
|
|
40
|
+
- Quality iteration is enabled per story or by an explicit run override.
|
|
41
|
+
- `--parallel=1` remains equivalent to isolated serial execution. Values above one become available
|
|
42
|
+
only through the provider-backed dispatcher described below.
|
|
43
|
+
- Existing exit codes retain their meanings. Quality exhaustion and inconsistent judging are
|
|
44
|
+
ordinary blocked states and therefore return the existing blocked exit code.
|
|
45
|
+
- Existing standalone `yoke review` continues to work. Its verdict schema gains optional fields but
|
|
46
|
+
remains backward compatible with current verdict files.
|
|
47
|
+
|
|
48
|
+
## Authority and invariants
|
|
49
|
+
|
|
50
|
+
The gate order for every candidate is:
|
|
51
|
+
|
|
52
|
+
```text
|
|
53
|
+
implement or repair
|
|
54
|
+
-> structured criterion commands
|
|
55
|
+
-> project verify
|
|
56
|
+
-> performance gate when configured
|
|
57
|
+
-> audit gate when configured
|
|
58
|
+
-> quality challenge when enabled
|
|
59
|
+
-> independent code review when enabled
|
|
60
|
+
-> commit and passes:true atomically
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
The following invariants are absolute:
|
|
64
|
+
|
|
65
|
+
1. A quality verdict can reject mechanically green work but can never approve mechanically red work.
|
|
66
|
+
2. Every repair reruns the complete gate sequence from criterion evidence onward.
|
|
67
|
+
3. A fresh critic evaluates every new candidate. Previous critic explanations are not included in
|
|
68
|
+
the next critic prompt; only the repair worker receives the selected gap.
|
|
69
|
+
4. `passes: true` is written only after integrated-tree verification and a successful commit.
|
|
70
|
+
5. Unbounded mode removes only `maxRounds` and `maxMinutes` from the quality iteration. Watchdog,
|
|
71
|
+
pause, cleanup, locks, worktree isolation, provider failures, audit, and all verification remain.
|
|
72
|
+
6. Malformed, absent, inconsistent, or unverifiable quality evidence never becomes a pass.
|
|
73
|
+
|
|
74
|
+
## Configuration and PRD schema
|
|
75
|
+
|
|
76
|
+
### Project defaults
|
|
77
|
+
|
|
78
|
+
`.yoke/config.yaml` gains optional defaults:
|
|
79
|
+
|
|
80
|
+
```yaml
|
|
81
|
+
quality:
|
|
82
|
+
enabled: false
|
|
83
|
+
policy: blocking # blocking | advisory
|
|
84
|
+
maxRounds: 3
|
|
85
|
+
maxMinutes: 60
|
|
86
|
+
consistencyChecks: 2 # normal order plus label swap
|
|
87
|
+
maxParallelCandidates: 2
|
|
88
|
+
critic:
|
|
89
|
+
agent: codex # optional; independent resolution otherwise
|
|
90
|
+
model: gpt-5.6-sol # opaque provider model string
|
|
91
|
+
reasoningEffort: high
|
|
92
|
+
repair:
|
|
93
|
+
agent: claude # optional; story runner otherwise
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
Defaults apply only when a story declares a quality challenge or the run explicitly enables one.
|
|
97
|
+
`enabled: true` means stories with a complete `quality` declaration run it automatically; it does
|
|
98
|
+
not invent references for ordinary stories.
|
|
99
|
+
|
|
100
|
+
### Story declaration
|
|
101
|
+
|
|
102
|
+
`StorySchema` gains an optional `quality` object:
|
|
103
|
+
|
|
104
|
+
```yaml
|
|
105
|
+
- id: STORY-4
|
|
106
|
+
title: Polish the pricing page
|
|
107
|
+
priority: 4
|
|
108
|
+
acceptance: [...]
|
|
109
|
+
passes: false
|
|
110
|
+
quality:
|
|
111
|
+
reference:
|
|
112
|
+
name: Stripe pricing page
|
|
113
|
+
source: https://stripe.com/pricing
|
|
114
|
+
kind: url # url | file | command
|
|
115
|
+
digest: sha256:... # required for blocking file snapshots; acquired for URLs
|
|
116
|
+
candidate:
|
|
117
|
+
kind: screenshots # screenshots | files | command-output | benchmark
|
|
118
|
+
paths:
|
|
119
|
+
- .yoke/proof/STORY-4/desktop.png
|
|
120
|
+
- .yoke/proof/STORY-4/mobile.png
|
|
121
|
+
rubric: Visual hierarchy, clarity, responsive finish, and interaction quality
|
|
122
|
+
policy: blocking # optional override
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Validation requires a non-empty, named, fetchable, and comparable reference; a concrete candidate
|
|
126
|
+
artifact contract; and a non-empty rubric. Blocking comparisons cannot use mutable live content as
|
|
127
|
+
their only retained evidence. URL references are fetched during preflight, converted to inert
|
|
128
|
+
artifacts, hashed, timestamped, and stored under `.yoke/references/<digest>/`. The runtime artifact
|
|
129
|
+
is ignored by Git by default; the provenance record is written into story proof evidence.
|
|
130
|
+
|
|
131
|
+
`command` sources and candidate commands use argv arrays in the internal schema. The public YAML
|
|
132
|
+
may use a command string only through the same command validation policy as existing proof commands;
|
|
133
|
+
shell control operators are rejected.
|
|
134
|
+
|
|
135
|
+
### CLI overrides
|
|
136
|
+
|
|
137
|
+
`yoke loop run` gains:
|
|
138
|
+
|
|
139
|
+
```text
|
|
140
|
+
--quality enable declared quality challenges for this run
|
|
141
|
+
--no-quality disable them for this run
|
|
142
|
+
--quality-rounds=N override maxRounds
|
|
143
|
+
--quality-minutes=N override maxMinutes
|
|
144
|
+
--quality-unbounded remove quality round/time limits explicitly
|
|
145
|
+
--quality-policy=blocking|advisory
|
|
146
|
+
--parallel=N run independent stories concurrently
|
|
147
|
+
--candidates=N run competing candidates for one quality-enabled story
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
`--quality-unbounded` implies `--quality`, prints a prominent startup warning, and is recorded in
|
|
151
|
+
status/proof provenance. It conflicts with `--quality-rounds` and `--quality-minutes`. It does not
|
|
152
|
+
imply `--unsafe`, does not disable `--timeout`, and does not change `--parallel` or candidate count.
|
|
153
|
+
|
|
154
|
+
## 1. Reviewer-to-repair transition
|
|
155
|
+
|
|
156
|
+
Current reviewer rejection ends the story. The loop instead classifies the failure:
|
|
157
|
+
|
|
158
|
+
- mechanical gate failure: block immediately; no quality repair is attempted;
|
|
159
|
+
- quality challenge loss: select the verdict's single `biggestGap`;
|
|
160
|
+
- code-review rejection: select the highest-severity actionable blocking finding;
|
|
161
|
+
- malformed reviewer output or provider failure: block as infrastructure failure, not repair work.
|
|
162
|
+
|
|
163
|
+
When an actionable gap exists and the quality budget permits another round, Yoke launches a fresh
|
|
164
|
+
repair worker in the same isolated story worktree. Its prompt contains:
|
|
165
|
+
|
|
166
|
+
- story title and acceptance criteria;
|
|
167
|
+
- settled project context;
|
|
168
|
+
- the current diff;
|
|
169
|
+
- exactly one selected gap and its cited evidence;
|
|
170
|
+
- the instruction to make the smallest change that closes that gap;
|
|
171
|
+
- the immutable mechanical gate commands.
|
|
172
|
+
|
|
173
|
+
It does not receive critic praise, hidden labels, earlier critic deliberation, or unrelated findings.
|
|
174
|
+
After repair, criterion evidence is replaced with evidence from the new candidate and all gates rerun.
|
|
175
|
+
|
|
176
|
+
Default exhaustion is three repair rounds or 60 elapsed quality minutes per story, whichever occurs
|
|
177
|
+
first. Exhaustion writes a blocked status naming the remaining gap and the exhausted budget. In
|
|
178
|
+
unbounded mode, quality rounds continue until a consistent pass, human pause/stop, or a non-quality
|
|
179
|
+
failure blocks the story.
|
|
180
|
+
|
|
181
|
+
## 2. External quality references
|
|
182
|
+
|
|
183
|
+
A new `src/quality/` boundary owns reference acquisition, artifact preparation, comparison, and
|
|
184
|
+
evidence. It does not own story selection, commits, or process lifecycle.
|
|
185
|
+
|
|
186
|
+
Reference preflight occurs before implementation for a blocking challenge so an unreachable bar does
|
|
187
|
+
not waste a full agent run. Acquisition records:
|
|
188
|
+
|
|
189
|
+
- declared source and display name;
|
|
190
|
+
- resolved URL or path;
|
|
191
|
+
- acquisition timestamp;
|
|
192
|
+
- media/content type and byte size;
|
|
193
|
+
- SHA-256 digest;
|
|
194
|
+
- tool/adapter version;
|
|
195
|
+
- any viewport, benchmark, or normalization parameters.
|
|
196
|
+
|
|
197
|
+
Fetched content is untrusted data. HTML is rendered or converted to static evidence; scripts and text
|
|
198
|
+
from the reference are never inserted as instructions. Size and content-type limits apply. Critics
|
|
199
|
+
receive artifacts plus the trusted Yoke rubric in separate, explicitly delimited sections and run
|
|
200
|
+
read-only without source-tree write access.
|
|
201
|
+
|
|
202
|
+
For visual stories, existing `flow-smoke` screenshots are reused when paths match the declared
|
|
203
|
+
candidate contract. Quality comparison never triggers a second capture farm. Nonvisual adapters
|
|
204
|
+
compare declared files, command output, or benchmark records.
|
|
205
|
+
|
|
206
|
+
If a required reference cannot be acquired or its digest changes during the run, blocking policy
|
|
207
|
+
blocks before commit. Advisory policy records `skipped` with the exact cause and continues; skipped is
|
|
208
|
+
never reported as passed.
|
|
209
|
+
|
|
210
|
+
## 3. Blind binary comparison
|
|
211
|
+
|
|
212
|
+
Quality verdict schema version 1:
|
|
213
|
+
|
|
214
|
+
```json
|
|
215
|
+
{
|
|
216
|
+
"schemaVersion": 1,
|
|
217
|
+
"verdict": "candidate|reference",
|
|
218
|
+
"biggestGap": "single actionable gap",
|
|
219
|
+
"evidence": ["artifact-relative citation"],
|
|
220
|
+
"confidence": "high|medium|low",
|
|
221
|
+
"labels": { "candidate": "A", "reference": "B" },
|
|
222
|
+
"provenance": {
|
|
223
|
+
"provider": "codex",
|
|
224
|
+
"model": "gpt-5.6-sol",
|
|
225
|
+
"promptVersion": 1,
|
|
226
|
+
"rubricDigest": "sha256:...",
|
|
227
|
+
"referenceDigest": "sha256:...",
|
|
228
|
+
"candidateDigest": "sha256:..."
|
|
229
|
+
}
|
|
230
|
+
}
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
A blocking pass requires two fresh comparisons: one randomized A/B ordering and one label-swapped
|
|
234
|
+
ordering. Both must select the candidate. A disagreement, missing evidence, low-confidence verdict,
|
|
235
|
+
schema failure, or digest mismatch is `inconsistent`, not a pass. It blocks with retained evidence so
|
|
236
|
+
the user can rerun, change reviewer, or downgrade the story policy explicitly.
|
|
237
|
+
|
|
238
|
+
The critic names exactly one biggest remaining gap. Scores out of ten are not part of the contract.
|
|
239
|
+
Advisory mode records the same evidence but cannot prevent a mechanically valid story from landing.
|
|
240
|
+
|
|
241
|
+
## 4. Provider-backed parallel execution
|
|
242
|
+
|
|
243
|
+
The existing scheduler, claims, `runParallelLoop`, and `MergeQueue` become the production CLI path.
|
|
244
|
+
Parallel mode requires isolation.
|
|
245
|
+
|
|
246
|
+
### Independent stories
|
|
247
|
+
|
|
248
|
+
1. The dispatcher owns the existing loop lock.
|
|
249
|
+
2. Ready stories are selected by `needs`; duplicate active `area` values are excluded.
|
|
250
|
+
3. Each worker atomically claims one story and creates a project-owned worktree from the current HEAD.
|
|
251
|
+
4. It runs implementation and all per-story gates, including quality repair rounds.
|
|
252
|
+
5. A green worker enters the serialized merge queue.
|
|
253
|
+
6. The integrator rebases onto current HEAD, reruns criterion/project/perf/audit gates against the
|
|
254
|
+
integrated candidate, then integrates, updates context/PRD, and commits atomically.
|
|
255
|
+
7. A rebase conflict or integrated-tree failure releases the claim and reopens the story with a
|
|
256
|
+
structured reason. It never sets `passes: true`.
|
|
257
|
+
|
|
258
|
+
Provider subprocess workers expose cancellation and watchdog handles to scoped cleanup. Claim files
|
|
259
|
+
gain dispatcher ID, PID, base commit, worktree, provider/model, and heartbeat. Cleanup only removes
|
|
260
|
+
resources owned by the target project and dead dispatcher.
|
|
261
|
+
|
|
262
|
+
### Competing candidates for one story
|
|
263
|
+
|
|
264
|
+
`--candidates=N` creates N worktrees from the same base and runs independent implementations. Each
|
|
265
|
+
must pass mechanical gates. Quality comparison then selects one mechanically green candidate against
|
|
266
|
+
the reference; if more than one beats the reference, a final blind candidate-vs-candidate comparison
|
|
267
|
+
chooses the winner. Only the winning branch enters the merge queue. Losing candidates are retained as
|
|
268
|
+
proof metadata, then their worktrees are cleaned. Candidate branches are never combined.
|
|
269
|
+
|
|
270
|
+
Concurrency and candidate fan-out are separately limited. Defaults remain one story worker and one
|
|
271
|
+
candidate. Provider subprocess wiring ships only when cancellation, crash recovery, conflict,
|
|
272
|
+
integrated reverify, and cost telemetry tests pass end to end.
|
|
273
|
+
|
|
274
|
+
## 5. Ephemeral story decomposition
|
|
275
|
+
|
|
276
|
+
Before implementation, a worker may emit a versioned decomposition artifact:
|
|
277
|
+
|
|
278
|
+
```json
|
|
279
|
+
{
|
|
280
|
+
"schemaVersion": 1,
|
|
281
|
+
"storyId": "STORY-4",
|
|
282
|
+
"subtasks": [
|
|
283
|
+
{ "id": "hero", "goal": "Improve hero hierarchy", "area": "ui", "needs": [] },
|
|
284
|
+
{ "id": "mobile", "goal": "Polish mobile layout", "area": "ui-mobile", "needs": ["hero"] }
|
|
285
|
+
]
|
|
286
|
+
}
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
Subtasks are runtime scheduling hints under `.yoke/work/<story>/`; they are not PRD stories, cannot
|
|
290
|
+
change acceptance criteria, and cannot independently set `passes`. The parent story remains the only
|
|
291
|
+
commit and quality unit. Subtasks may run concurrently only when their declared areas differ and each
|
|
292
|
+
has its own nested worktree/branch or non-overlapping artifact workspace. Their results are synthesized
|
|
293
|
+
into one parent candidate before any gate runs.
|
|
294
|
+
|
|
295
|
+
Invalid decomposition falls back to one whole-story worker. Dynamic decomposition is optional and
|
|
296
|
+
must not add a controller call for small stories unless explicitly enabled or routing selects it.
|
|
297
|
+
|
|
298
|
+
## 6. Versioned provider and critic contracts
|
|
299
|
+
|
|
300
|
+
Provider adapters own every machine result. Add schemas for:
|
|
301
|
+
|
|
302
|
+
- route decisions;
|
|
303
|
+
- decomposition plans;
|
|
304
|
+
- review verdicts and selected repair gaps;
|
|
305
|
+
- quality comparison verdicts;
|
|
306
|
+
- candidate selection;
|
|
307
|
+
- provider telemetry envelopes.
|
|
308
|
+
|
|
309
|
+
Every envelope carries `schemaVersion`, provider, reported model when available, invocation role,
|
|
310
|
+
prompt/contract version, timing, permission profile, and usage fields. Unknown optional telemetry is
|
|
311
|
+
retained in raw per-call evidence but not promoted to aggregate fields without schema support.
|
|
312
|
+
|
|
313
|
+
Native provider output schemas are used when stable and available. Otherwise Yoke's existing
|
|
314
|
+
result-file transport is the compatibility path. Free-form stdout markers such as `YOKE_ROUTE` remain
|
|
315
|
+
temporary backward compatibility only; malformed output visibly falls back to the strong parent and
|
|
316
|
+
records `fallbackReason`, rather than silently disappearing.
|
|
317
|
+
|
|
318
|
+
All status NDJSON and proof evidence include quality round, candidate, critic, reference digest,
|
|
319
|
+
selected gap, verdict consistency, parallel worker, merge outcome, and cumulative token/cost data.
|
|
320
|
+
|
|
321
|
+
## State and observability
|
|
322
|
+
|
|
323
|
+
Runtime artifacts:
|
|
324
|
+
|
|
325
|
+
```text
|
|
326
|
+
.yoke/references/<digest>/... acquired inert reference
|
|
327
|
+
.yoke/work/<story>/decomposition.json ephemeral subtask plan
|
|
328
|
+
.yoke/proof/<story>/quality/round-<n>/ candidate and critic evidence
|
|
329
|
+
.yoke/proof/<story>/quality/summary.json final quality outcome
|
|
330
|
+
.yoke/claims/<story>.json parallel ownership/heartbeat
|
|
331
|
+
.yoke/loop-status.json current round/candidate/worker
|
|
332
|
+
.yoke/loop.log phase transitions and reasons
|
|
333
|
+
```
|
|
334
|
+
|
|
335
|
+
New reporter phases are `decomposing`, `quality-preflight`, `comparing`, `repairing`, `integrating`,
|
|
336
|
+
and `selecting-candidate`. Status includes bounded/unbounded mode, current/maximum quality round,
|
|
337
|
+
elapsed quality time, reference digest, active workers, and candidate count. Logs stay bounded under
|
|
338
|
+
the existing policy; structured proof artifacts are per-story and never inferred from prose.
|
|
339
|
+
|
|
340
|
+
Pause is honored at safe boundaries: between quality rounds and before launching new parallel workers.
|
|
341
|
+
An in-flight provider invocation is allowed to finish unless cleanup explicitly terminates its scoped
|
|
342
|
+
PID tree. In unbounded mode, `yoke loop status` always shows that only a human pause/stop can end
|
|
343
|
+
quality iteration.
|
|
344
|
+
|
|
345
|
+
## Error and recovery map
|
|
346
|
+
|
|
347
|
+
| Failure | Result |
|
|
348
|
+
|---|---|
|
|
349
|
+
| Mechanical criterion/verify/perf/audit failure | Block story; no quality repair |
|
|
350
|
+
| Quality critic says reference wins | Repair next round when budget allows |
|
|
351
|
+
| Code reviewer returns actionable blocking finding | Repair next round when budget allows |
|
|
352
|
+
| Missing/malformed verdict | Block as reviewer infrastructure failure |
|
|
353
|
+
| Label-swap disagreement or low confidence | Block as inconsistent judge evidence |
|
|
354
|
+
| Reference unavailable before work | Block, or advisory skip when explicitly configured |
|
|
355
|
+
| Reference digest changes mid-run | Block and retain both provenance records |
|
|
356
|
+
| Repair budget exhausted | Block with last gap and budget evidence |
|
|
357
|
+
| Human pause during unbounded mode | Finish current safe unit, persist state, exit paused |
|
|
358
|
+
| Provider crash/timeout | Block worker, preserve logs, release dead claim during cleanup |
|
|
359
|
+
| Parallel rebase conflict | Reopen story with conflict reason; no integration |
|
|
360
|
+
| Integrated-tree verification failure | Reopen story; candidate never passes |
|
|
361
|
+
| Candidate fan-out has no mechanically green candidate | Block without subjective comparison |
|
|
362
|
+
|
|
363
|
+
## Security
|
|
364
|
+
|
|
365
|
+
- Critics and reference evaluators run read-only.
|
|
366
|
+
- External reference content is never trusted as instructions.
|
|
367
|
+
- URL fetching uses protocol allowlists, redirect limits, byte limits, and loopback/private-network
|
|
368
|
+
restrictions unless the user explicitly supplies a local URL for a local project.
|
|
369
|
+
- Reference and candidate paths are normalized beneath Yoke-owned roots; traversal is rejected.
|
|
370
|
+
- Commands use validated argv contracts and existing approved-test restrictions where applicable.
|
|
371
|
+
- Unbounded mode requires the explicit CLI flag on every run; it is not persisted as a default.
|
|
372
|
+
- Startup output names effective provider, permissions, quality policy, budgets, and unbounded state.
|
|
373
|
+
- Token/cost growth is observable. Unbounded removes quality stop budgets by explicit request, but it
|
|
374
|
+
does not hide or stop reporting spend.
|
|
375
|
+
|
|
376
|
+
## Testing strategy
|
|
377
|
+
|
|
378
|
+
Development follows red-green-refactor. Add focused suites for:
|
|
379
|
+
|
|
380
|
+
1. Config and PRD compatibility, quality declaration validation, and CLI conflicts.
|
|
381
|
+
2. Repair-loop transitions, complete gate reruns, fresh reviewer use, and single-gap prompts.
|
|
382
|
+
3. Three-round/60-minute exhaustion and explicit unbounded behavior.
|
|
383
|
+
4. Reference acquisition, hashing, mutable-reference detection, inert conversion, path and network
|
|
384
|
+
security, and advisory-vs-blocking failures.
|
|
385
|
+
5. Randomized blind labels, mandatory label swap, inconsistent verdicts, schema failures, and proof
|
|
386
|
+
provenance.
|
|
387
|
+
6. Real provider subprocess workers using deterministic fake CLIs, including cancellation, timeout,
|
|
388
|
+
stale claims, process cleanup, and unavailable providers.
|
|
389
|
+
7. Parallel story scheduling, area exclusion, dependency ordering, rebase conflicts, integrated-tree
|
|
390
|
+
verification, and atomic PRD/context commits in temporary Git repositories.
|
|
391
|
+
8. Competing candidates: common base, mechanical filtering, winner-only merge, and losing cleanup.
|
|
392
|
+
9. Ephemeral decomposition validation, dependency/area scheduling, synthesis, fallback, and proof that
|
|
393
|
+
it cannot mutate PRD acceptance or pass state.
|
|
394
|
+
10. Versioned route/review/quality/telemetry envelopes and visible fallback reasons.
|
|
395
|
+
11. Reporter/NDJSON status for quality rounds, unbounded mode, parallel workers, and merge outcomes.
|
|
396
|
+
12. End-to-end serial, bounded-quality, unbounded-pause, parallel-story, and candidate-race flows.
|
|
397
|
+
|
|
398
|
+
Final verification runs TypeScript build, the complete Vitest suite, canon validation, docs metadata
|
|
399
|
+
check, package dry run, deterministic fake-provider flows, and authenticated opt-in smoke matrices for
|
|
400
|
+
Claude, Codex, and Gemini. Benchmarks compare quality off/on, bounded/unbounded-paused, serial/parallel,
|
|
401
|
+
and one/multiple candidates. Claims remain workload-specific and report unavailable provider evidence
|
|
402
|
+
honestly.
|
|
403
|
+
|
|
404
|
+
## Delivery order
|
|
405
|
+
|
|
406
|
+
1. Versioned result envelopes and review-gap selection.
|
|
407
|
+
2. Bounded reviewer-to-repair loop with status and evidence.
|
|
408
|
+
3. Reference/candidate schemas, acquisition, and blind consistency comparison.
|
|
409
|
+
4. CLI/config integration including explicit unbounded mode.
|
|
410
|
+
5. Provider subprocess workers wired to scheduler, claims, and merge queue.
|
|
411
|
+
6. Competing candidate mode.
|
|
412
|
+
7. Ephemeral decomposition and subtask scheduling.
|
|
413
|
+
8. Cross-provider benchmarks, documentation, migration notes, and release provenance updates.
|
|
414
|
+
|
|
415
|
+
Each stage preserves the serial non-quality path and lands only with its focused tests plus the full
|
|
416
|
+
existing suite green.
|
|
417
|
+
|
|
418
|
+
## Attribution
|
|
419
|
+
|
|
420
|
+
The comparative quality concepts are inspired by Matt Shumer's Gauntlet Loop technique and the
|
|
421
|
+
`robonuggets/gauntlet-loop` packaging. Yoke reimplements the concepts in its own mechanical control
|
|
422
|
+
plane. Documentation credits the inspiration; no CC BY prompt text is copied into runtime assets.
|