@hecer/yoke 1.2.1 → 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (67) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/.codex-plugin/plugin.json +1 -1
  3. package/CHANGELOG.md +66 -24
  4. package/README.md +186 -104
  5. package/TODOS.md +0 -3
  6. package/canon/loop/loop-spec.md +53 -24
  7. package/canon/loop/prd.schema.md +30 -6
  8. package/canon/manifest.yaml +1 -1
  9. package/canon/skills/authoring-prd/SKILL.md +30 -31
  10. package/dist/agents/contracts.js +50 -0
  11. package/dist/agents/process-incarnation.js +15 -0
  12. package/dist/agents/process-record.js +65 -0
  13. package/dist/agents/process-streams.js +40 -0
  14. package/dist/agents/process.js +177 -0
  15. package/dist/agents/providers.js +10 -7
  16. package/dist/agents/telemetry.js +62 -0
  17. package/dist/change/inbox.js +279 -0
  18. package/dist/cli.js +90 -4
  19. package/dist/loop/candidate-boundaries.js +43 -0
  20. package/dist/loop/candidate-cleanup.js +98 -0
  21. package/dist/loop/candidate-contracts.js +1 -0
  22. package/dist/loop/candidate-selection.js +84 -0
  23. package/dist/loop/candidates.js +228 -0
  24. package/dist/loop/claim-lease.js +131 -0
  25. package/dist/loop/claims.js +177 -40
  26. package/dist/loop/cleanup.js +117 -15
  27. package/dist/loop/decision.js +31 -0
  28. package/dist/loop/dispatcher.js +334 -0
  29. package/dist/loop/evidence.js +31 -0
  30. package/dist/loop/gates.js +10 -1
  31. package/dist/loop/loop.js +236 -33
  32. package/dist/loop/merge-queue.js +12 -6
  33. package/dist/loop/parallel-adapters.js +185 -0
  34. package/dist/loop/parallel-command.js +287 -0
  35. package/dist/loop/parallel.js +2 -4
  36. package/dist/loop/prd.js +63 -2
  37. package/dist/loop/reporter.js +86 -5
  38. package/dist/loop/run-command.js +227 -53
  39. package/dist/loop/runner.js +77 -34
  40. package/dist/loop/verify.js +11 -0
  41. package/dist/loop/watchdog.js +67 -8
  42. package/dist/loop/worker-cancellation.js +17 -0
  43. package/dist/loop/worker-cleanup.js +23 -0
  44. package/dist/loop/worker-contracts.js +1 -0
  45. package/dist/loop/worker.js +254 -0
  46. package/dist/prd/command.js +18 -5
  47. package/dist/quality/artifacts.js +59 -0
  48. package/dist/quality/candidate-comparison.js +130 -0
  49. package/dist/quality/command.js +316 -0
  50. package/dist/quality/loop.js +86 -0
  51. package/dist/quality/process-command.js +57 -0
  52. package/dist/quality/reference.js +187 -0
  53. package/dist/quality/repair.js +11 -0
  54. package/dist/quality/runner.js +66 -0
  55. package/dist/quality/types.js +60 -0
  56. package/dist/quality/verdict.js +142 -0
  57. package/dist/retrofit/config.js +13 -4
  58. package/dist/retrofit/gitignore.js +4 -0
  59. package/dist/review/command.js +27 -38
  60. package/dist/review/verdict.js +38 -7
  61. package/dist/routing/router.js +4 -1
  62. package/docs/MIGRATING-TO-1.4.md +70 -0
  63. package/docs/PUBLISHING.md +16 -2
  64. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -0
  65. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -0
  66. package/gemini-extension.json +1 -1
  67. package/package.json +6 -6
@@ -2,7 +2,7 @@
2
2
  "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
3
3
  "name": "yoke",
4
4
  "displayName": "Yoke",
5
- "version": "1.2.1",
5
+ "version": "1.4.0",
6
6
  "description": "Cross-agent coding harness: one curated skill canon (TDD, brainstorming, plans, reviews, shipping, design verification) plus mechanical safety gates and an autonomous loop via the yoke CLI.",
7
7
  "author": { "name": "HECer", "url": "https://github.com/HECer" },
8
8
  "homepage": "https://github.com/HECer/yoke#readme",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "yoke",
3
- "version": "1.2.1",
3
+ "version": "1.4.0",
4
4
  "description": "Cross-agent coding discipline, mechanical gates, and release workflows",
5
5
  "skills": "./canon/skills/",
6
6
  "hooks": "./hooks/hooks.json"
package/CHANGELOG.md CHANGED
@@ -1,27 +1,69 @@
1
- # Changelog
2
-
3
- ## 1.2.1 — 2026-08-02
4
-
5
- ### Fixed
6
- - `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
7
- - Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
8
-
9
- ## 1.2.0 — 2026-08-02
10
-
11
- ### Added
12
- - Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
13
- - A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
14
- - Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
15
-
16
- ### Changed
17
- - Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
18
- - Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
19
-
20
- ### Fixed
21
- - Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
22
- - Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
23
-
24
- ## 1.1.0 — 2026-07-30
1
+ # Changelog
2
+
3
+ ## 1.4.0 — 2026-08-15
4
+
5
+ ### Added
6
+ - `yoke loop run --parallel=N` now executes dependency-ready, non-colliding stories through real provider subprocess workers, isolated worktrees, leased claims, and a FIFO integration queue with fresh integrated-system gates.
7
+ - Reference-driven quality declarations can collect screenshots, files, command output, or benchmark results and run a schema-validated blind critic with bounded repair rounds, elapsed-time limits, blocking or advisory policy, and retained proof.
8
+ - `--candidates=N` can fan out up to five isolated implementations, discard mechanically red candidates, select one green candidate through identity-blind pairwise comparison, and preserve selected/loser evidence before cleanup.
9
+ - Loop status now exposes dispatcher, worker, integrator, candidate lifecycle, worktree, queue, integration, reopen, quality-round, repair-budget, and trusted provider/model provenance data.
10
+
11
+ ### Changed
12
+ - Provider subprocesses use explicit lifecycle contracts and incarnation-aware process records so worker cancellation and cleanup target only the process tree Yoke actually started.
13
+ - Parallel and candidate runs disable adaptive routing, honor story-level provider affinity, latch pause requests across the whole dispatcher, and rerun quality plus review after integration.
14
+ - `yoke loop cleanup` retains Yoke worktrees unless `--remove-worktrees` is explicit, while still reaping recorded orphan runners and stale locks safely.
15
+
16
+ ### Fixed
17
+ - Expired claims, worker crashes, merge conflicts, pause races, and integration failures now release ownership deterministically, retain terminal proof, and reopen stories without leaking worktrees or marking false completion.
18
+ - Quality repair fails closed on malformed critic output, reference drift, provider/model provenance mismatch, candidate identity leakage, unavailable critics, exhausted limits, and mechanically red repairs.
19
+ - The watchdog resolves its TypeScript loader from both source and built npm layouts on Node 20+, and read-only Codex comparisons can run in disposable candidate worktrees without weakening normal repository checks.
20
+
21
+ ### Security
22
+ - Blind comparison requests expose only opaque labels and digests while binding every verdict to the trusted judge provider, model, prompt, rubric, reference, and candidate provenance.
23
+ - Cleanup and cancellation use project-scoped leases, owner tokens, PID birth/incarnation checks, and recorded process handles rather than machine-wide process-name matching.
24
+
25
+ ## 1.3.0 — 2026-08-09
26
+
27
+ ### Added
28
+ - Teams can submit change requests at any time with `yoke change add` and inspect the append-only inbox with `yoke change status`; queued requests are planned at safe loop boundaries by Claude, Codex, or Gemini without interrupting active story work.
29
+ - Acceptance criteria can carry stable IDs and executable verification commands, and Yoke records per-criterion evidence before a story can be marked complete.
30
+ - Projects can configure an integrated completion command that proves the final system flow after all story-level gates pass.
31
+
32
+ ### Changed
33
+ - New projects use strict criterion-proof requirements by default, while existing boolean-only criteria remain readable for compatibility.
34
+ - PRD authoring, schemas, loop guidance, generated configuration, and provider instructions now treat completion as an ephemeral readiness state: new requests create more stories instead of prematurely forcing a release boundary.
35
+
36
+ ### Fixed
37
+ - Broad green test suites can no longer satisfy unrelated structured acceptance criteria: proof commands must use an approved test runner, contain the criterion ID, and avoid shell operators.
38
+ - Change-intake failures block safely; an independent coverage pass rejects omitted requested outcomes, crash recovery retains uncommitted requests, and concurrent PRD edits are detected instead of silently overwriting user work.
39
+ - Integrated completion checks prevent locally finished stories from masking broken cross-component flows such as authentication callbacks or purchase-to-entitlement activation.
40
+
41
+ ### Security
42
+ - Updated the transitive development dependency `nanoid` to a non-vulnerable release so the complete CI dependency audit is clean.
43
+ - Restricted model-authored criterion proof to criterion-targeted test commands and blocked shell control operators and runner-prefix spoofing before host execution.
44
+
45
+ ## 1.2.1 — 2026-08-02
46
+
47
+ ### Fixed
48
+ - `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
49
+ - Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
50
+
51
+ ## 1.2.0 — 2026-08-02
52
+
53
+ ### Added
54
+ - Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
55
+ - A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
56
+ - Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
57
+
58
+ ### Changed
59
+ - Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
60
+ - Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
61
+
62
+ ### Fixed
63
+ - Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
64
+ - Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
65
+
66
+ ## 1.1.0 — 2026-07-30
25
67
 
26
68
  ### Added
27
69
  - Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
package/README.md CHANGED
@@ -2,8 +2,8 @@
2
2
 
3
3
  # 🐂 Yoke
4
4
 
5
- <!-- yoke:version:start -->1.2.1<!-- yoke:version:end -->
6
- <!-- yoke:tests:start -->582<!-- yoke:tests:end -->
5
+ <!-- yoke:version:start -->1.4.0<!-- yoke:version:end -->
6
+ <!-- yoke:tests:start -->928<!-- yoke:tests:end -->
7
7
  <!-- yoke:skills:start -->29<!-- yoke:skills:end -->
8
8
  <!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
9
9
 
@@ -17,7 +17,7 @@
17
17
  [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](#-license)
18
18
  ![Node](https://img.shields.io/badge/node-%E2%89%A520-339933?logo=node.js&logoColor=white)
19
19
  ![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
20
- ![Tests](https://img.shields.io/badge/tests-559%20passing-brightgreen.svg)
20
+ ![Tests](https://img.shields.io/badge/tests-928%20passing-brightgreen.svg)
21
21
  ![Agents](https://img.shields.io/badge/agents-Claude%20%7C%20Codex%20%7C%20Gemini-8A2BE2)
22
22
  ![Built with TDD](https://img.shields.io/badge/built%20with-TDD%20%2B%20review-ff69b4.svg)
23
23
 
@@ -25,7 +25,11 @@
25
25
 
26
26
  </div>
27
27
 
28
- > **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
28
+ > **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. Add `--parallel=N` for dependency-aware workers, or declare a reference and add `--quality` for a bounded critic/repair gauntlet. If any blocking gate is red, nothing is committed. Proof lives in `.yoke/proof/<story>/`.
29
+
30
+ Yoke 1.4 adds opt-in parallel workers and a bounded, reference-driven quality gauntlet without
31
+ changing existing serial loop defaults. See [the 1.4 migration guide](docs/MIGRATING-TO-1.4.md)
32
+ for the new flags, configuration, cleanup behavior, and review-verdict contract.
29
33
 
30
34
  Yoke 1.1 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
31
35
  is explicit; reviews require a schema-valid verdict and a different model unless
@@ -72,7 +76,7 @@ $ ls reading-app/.yoke/proof/STORY-2/
72
76
  home.png list.png # photographic evidence, labelled per story
73
77
  ```
74
78
 
75
- Every claim in that transcript is enforced by code paths with tests behind them — 559 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
79
+ Every claim in that transcript is enforced by code paths with tests behind them — 928 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
76
80
 
77
81
  ## 🚀 Quickstart
78
82
 
@@ -85,9 +89,9 @@ yoke new my-app --idea="a CLI that tracks reading lists"
85
89
  yoke loop on my-app && yoke loop run my-app --isolate
86
90
 
87
91
  # — or retrofit an existing project —
88
- yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
92
+ yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
89
93
  yoke validate canon # sanity-check the canon
90
- yoke loop run /path/to/project --isolate --reviewer=codex --max=20
94
+ yoke loop run /path/to/project --isolate --parallel=3 --reviewer=codex --max=20
91
95
  ```
92
96
 
93
97
  > Requires Node ≥ 20 and git. No global install? `node /path/to/yoke/dist/cli.js …` or `npm --prefix /path/to/yoke run yoke -- …` work too. The MCP tools (rtk, graphify/Serena, Playwright MCP) are wired by Yoke but installed separately — the generated config is a clearly-labelled, adjustable template.
@@ -104,7 +108,7 @@ The canon is also packaged as a Claude Code plugin — the repo is its own marke
104
108
  That gives you all canon skills under the `yoke:` namespace (e.g. `yoke:tdd`, `yoke:review`) inside Claude Code — no retrofit needed. The `yoke` CLI (loop, gates, retrofit for Codex/Gemini) still comes from `npm i -g @hecer/yoke`. Gemini CLI users can likewise `gemini extensions install https://github.com/HECer/yoke`.
105
109
 
106
110
  For Codex, no preinstalled skill is required: run `npx @hecer/yoke setup .` in a terminal, or
107
- ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
111
+ ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
108
112
  `.agents/skills/`, including `yoke-retrofit` and `yoke-workflow`; start a fresh Codex task if an
109
113
  already-open task does not discover newly installed skills. The npm package also contains
110
114
  `.codex-plugin/plugin.json` for Codex plugin hosts.
@@ -138,7 +142,7 @@ Yoke is meant to be operated *by* your coding agent — after a retrofit, the ag
138
142
 
139
143
  > **Watch / unblock** — *"Run `yoke loop status .`. If it says BLOCKED, run the project's verify command, find the root cause, fix it without weakening tests, then continue the loop."*
140
144
 
141
- > ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
145
+ > ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
142
146
 
143
147
  > ⚠️ **Never kill agent processes by name or command-line pattern** (e.g. every process matching `dangerously-skip-permissions`): on a machine running several yoke projects, that takes down the *healthy* runners of the other projects mid-story — they stall and their loops block. `yoke loop cleanup` is the scoped alternative: each watchdog records its pids in the project's `.yoke/runner.pid`, and cleanup kills exactly those recorded trees — nothing else on the machine.
144
148
 
@@ -148,14 +152,15 @@ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`)
148
152
 
149
153
  | Command | What it does | Exit codes |
150
154
  |---|---|---|
151
- | `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
155
+ | `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
152
156
  | `yoke validate [canonDir]` | Validate the canon (schema, frontmatter, templates) | `0` valid · `1` errors |
153
157
  | `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
154
158
  | `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
155
159
  | `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
156
160
  | `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
161
+ | `yoke change add\|status [dir] [--idea=]` | Queue a change at any time; the loop turns it into append-only stories at the next safe boundary | `0` · `1` invalid inbox/request |
157
162
  | `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
158
- | `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` is unlimited by default and `--max=N` creates an intentional batch cap; `decision` shows a critical stop, `answer` records it and resumes, `resume` retries a failed restart with the preserved safety options | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
163
+ | `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` supports `--parallel=N`, bounded reference-driven `--quality`, and blind `--candidates=N` selection; `--max=N` creates an intentional batch cap; `cleanup` retains worktrees unless `--remove-worktrees` is explicit | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
159
164
  | `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
160
165
  | `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
161
166
  | `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
@@ -310,28 +315,35 @@ Opt-in; `yoke setup` recommends enabling it for new installs, while `retrofit` a
310
315
 
311
316
  ```mermaid
312
317
  flowchart LR
313
- A[pick next PRD story] --> B{clean worktree?}
318
+ I[consume queued change<br/>as new stories] --> A[pick next PRD story]
319
+ A --> B{clean worktree?}
314
320
  B -- no --> X[blocked]
315
321
  B -- yes --> C{acceptance<br/>criteria?}
316
322
  C -- no --> X
317
323
  C -- yes --> D[agent implements<br/>one story]
318
- D --> E{tests green?}
324
+ D --> E{suite + criterion<br/>proof green?}
319
325
  E -- no --> X
320
326
  E -- yes --> F{reviewer<br/>approves?}
321
327
  F -- no --> X
322
328
  F -- yes --> G[commit + mark passes:true<br/>+ proof in .yoke/proof/]
323
- G --> A
329
+ G --> I
330
+ I --> H{all stories pass?}
331
+ H -- yes --> J{integrated system<br/>gate green?}
332
+ J -- no --> X
333
+ J -- yes --> K[current backlog ready]
324
334
  ```
325
335
 
326
336
  ```bash
327
337
  yoke loop on . # enable (recorded in .yoke/config.yaml)
328
338
  yoke loop status . # show state + PRD progress
339
+ yoke change add . --idea="Add passkey login" # safe while the loop runs
329
340
  yoke loop run . \
330
341
  --runner=codex \ # implement with Codex…
331
342
  --reviewer=claude \ # …review with Claude (role separation)
332
343
  --isolate \ # each story in a throwaway git worktree
333
- --decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
334
- # Optional: add --max=20 only when this run should stop after a bounded batch.
344
+ --parallel=3 \ # run dependency-ready, non-colliding stories concurrently
345
+ --decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
346
+ # Optional: add --max=20 only when this run should stop after a bounded batch.
335
347
  yoke loop off . # disable
336
348
  ```
337
349
 
@@ -342,11 +354,75 @@ yoke loop off . # disable
342
354
  title: Add a health endpoint
343
355
  priority: 1 # lower = higher priority
344
356
  acceptance: # Definition of Done (required, else blocked)
345
- - GET /health returns 200
357
+ - id: health-returns-200
358
+ text: GET /health returns 200
359
+ verify: [npm run test:health-returns-200]
360
+ - id: health-rejects-post
361
+ text: POST /health returns 405
362
+ verify: [npm run test:health-rejects-post]
346
363
  passes: false # the loop sets this true only on green tests
347
364
  ```
348
365
 
349
- The loop stops when every story is `passes: true`. State lives **outside the model context** — the PRD file plus git — so each iteration is fresh. The PRD is re-read from disk at every story boundary, so new stories appended to `.yoke/prd.yaml` **while the loop is running** are picked up at the next iteration — no restart needed.
366
+ New projects default to `verify.requireCriteria: true`: every story has 2–5 behavioral criteria.
367
+ Each criterion ID must occur in its single, approved test command; shell operators and broad,
368
+ untargeted suites are rejected. Yoke records each result in `.yoke/proof/<story>/evidence.json`. Configure optional
369
+ `completion.command` for integrated journeys such as purchase → entitlement → relaunch or
370
+ magic-link → callback → authenticated app. It runs whenever the current backlog has no open
371
+ stories; this is readiness, not a release.
372
+
373
+ No generic tool can infer whether arbitrary test code perfectly represents product meaning. Yoke
374
+ closes the mechanical false-done paths—targeted evidence, coverage review, clean committed state,
375
+ and integrated journeys—while the project still owns the correctness of its tests and production
376
+ observability.
377
+
378
+ ### Parallel workers and the quality gauntlet
379
+
380
+ `--parallel=N` dispatches dependency-ready stories concurrently. Claims carry leases, workers use
381
+ isolated worktrees, collision areas are serialized, and only a mechanically green candidate enters
382
+ the FIFO integration queue. Integration repeats the project gates against the merged tree; a worker
383
+ success can never bypass a red integrated result. `yoke loop status` reports the dispatcher,
384
+ workers, providers, worktrees, lifecycle, queue, integrations, and reopened stories.
385
+
386
+ Quality is reference-driven and opt-in. Declare what one story should match:
387
+
388
+ ```yaml
389
+ quality:
390
+ reference: { name: approved-home, source: design/home.png, kind: file }
391
+ candidate: { kind: screenshots, paths: [.yoke/proof/STORY-1/home.png] }
392
+ rubric: Match the approved layout, hierarchy, spacing, and states.
393
+ policy: blocking # or advisory
394
+ ```
395
+
396
+ Configure project defaults, then enable the gauntlet for a run:
397
+
398
+ ```yaml
399
+ quality:
400
+ enabled: false # keep opt-in, or make it the project default
401
+ policy: blocking
402
+ maxRounds: 3
403
+ maxMinutes: 60
404
+ consistencyChecks: 2
405
+ maxParallelCandidates: 2
406
+ critic: { agent: codex, model: gpt-5.6-sol } # model required for --candidates
407
+ repair: { agent: claude }
408
+ ```
409
+
410
+ ```bash
411
+ yoke loop run . --quality --quality-rounds=3 --quality-minutes=60
412
+ yoke loop run . --quality --candidates=2 # blind pairwise selection; stories need quality declarations
413
+ ```
414
+
415
+ The critic compares opaque candidate/reference labels, writes schema-validated provenance, and
416
+ cannot modify the project. Blocking findings enter a bounded repair loop and rerun every mechanical
417
+ gate; advisory findings are retained without blocking. `--quality-policy=`, `--no-quality`, and
418
+ `--quality-unbounded` override defaults for one run. Unbounded mode is explicit and warned because
419
+ it removes repair limits, not Yoke's watchdog, isolation, verification, or commit safety.
420
+
421
+ State lives **outside the model context** — the PRD file plus git — so each iteration is fresh.
422
+ Use `yoke change add` at any time. Its ignored append-only inbox is consumed at the next story
423
+ boundary. A separate coverage pass must confirm that every requested outcome maps to behavioral
424
+ criteria before Yoke appends and commits the new stories; existing stories are never rewritten and
425
+ no restart is needed.
350
426
 
351
427
  ### Watching a run
352
428
 
@@ -365,14 +441,16 @@ Every iteration emits token-free, harness-side feedback (Node console + local fi
365
441
  implementing · iteration 20 · 19/45 (42%) · updated 30s ago
366
442
  ~1h44m remaining (Ø 4m/story)
367
443
  ```
444
+ - **Parallel + quality detail** — active workers include provider, candidate ID, worktree,
445
+ lifecycle, phase, quality round, and repair budget; the integrator is shown separately.
368
446
  - **`.yoke/loop.log`** — an append-only timeline of every phase transition.
369
447
  - **`--json`** — machine mode for supervisors: every status write is *also* emitted as one
370
448
  NDJSON line on stdout (`{"type":"status","state":"running","phase":"verifying",…}` — the
371
449
  same shape as `loop-status.json`), the human narrative moves off stdout (the final summary
372
450
  goes to stderr), and a consumer can follow the stream line by line instead of polling the file.
373
- Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
374
- output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
375
- and cost fields when available. Missing values stay absent—Yoke does not estimate them.
451
+ Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
452
+ output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
453
+ and cost fields when available. Missing values stay absent—Yoke does not estimate them.
376
454
 
377
455
  ### Pausing a run
378
456
 
@@ -432,65 +510,65 @@ cannot begin because a provider/reviewer is unavailable or another process owns
432
510
  until a loop actually runs. To intentionally abandon an orphaned or stale private resume state,
433
511
  use `yoke loop resume . --discard`; pending decisions are never deleted by that command. Existing
434
512
  `loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` remain supported as compatibility aliases;
435
- new projects should use `decisionPolicy: auto|critical`.
436
-
437
- ### Adaptive model routing (explicit opt-in)
438
-
439
- `yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
440
- parent remains the strong planner/controller. Before each bounded story it receives only the
441
- story, acceptance criteria, and at most three eligible worker profiles, then returns one
442
- machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
443
- `SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
444
- runs so Yoke does not pay for two orchestration layers.
445
-
446
- **Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
447
- Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
448
- cover invocation and routing behavior for all three providers. The measured performance evidence
449
- below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
450
- until authenticated, repeated in-the-wild runs exist for those providers.
451
-
452
- ```yaml
453
- runner:
454
- agent: codex
455
- model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
456
- reasoningEffort: high
457
- routing:
458
- enabled: true # setup defaults false; setup --routing opts in
459
- strategy: balanced # balanced | cost | speed | quality
460
- maxCandidates: 3
461
- workers:
462
- - id: codex-light
463
- agent: codex
464
- reasoningEffort: low
465
- costTier: medium
466
- capabilities: [exploration, implementation, tests]
467
- - id: claude-fast
468
- agent: claude
469
- model: haiku # rolling alias; omit to use the provider's current default
470
- reasoningEffort: low
471
- costTier: low
472
- capabilities: [mechanical-edits, tests]
473
- - id: gemini-auto
474
- agent: gemini # omitted model means the account's current Auto/default route
475
- costTier: low
476
- capabilities: [large-context, implementation]
477
- ```
478
-
479
- Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
480
- baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
481
- eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
482
- "intelligence score". Candidate model IDs come from project configuration while setup defaults
483
- prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
484
- independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
485
- after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
486
- time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
487
- cannot overwrite a shared registry file.
488
-
489
- Routing is not free: it adds one controller call per story. It is most promising when a bounded
490
- worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
491
- on your own backlog rather than assuming a win.
492
-
493
- ### Performance budgets: efficiency as a gate, not a style
513
+ new projects should use `decisionPolicy: auto|critical`.
514
+
515
+ ### Adaptive model routing (explicit opt-in)
516
+
517
+ `yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
518
+ parent remains the strong planner/controller. Before each bounded story it receives only the
519
+ story, acceptance criteria, and at most three eligible worker profiles, then returns one
520
+ machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
521
+ `SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
522
+ runs so Yoke does not pay for two orchestration layers.
523
+
524
+ **Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
525
+ Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
526
+ cover invocation and routing behavior for all three providers. The measured performance evidence
527
+ below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
528
+ until authenticated, repeated in-the-wild runs exist for those providers.
529
+
530
+ ```yaml
531
+ runner:
532
+ agent: codex
533
+ model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
534
+ reasoningEffort: high
535
+ routing:
536
+ enabled: true # setup defaults false; setup --routing opts in
537
+ strategy: balanced # balanced | cost | speed | quality
538
+ maxCandidates: 3
539
+ workers:
540
+ - id: codex-light
541
+ agent: codex
542
+ reasoningEffort: low
543
+ costTier: medium
544
+ capabilities: [exploration, implementation, tests]
545
+ - id: claude-fast
546
+ agent: claude
547
+ model: haiku # rolling alias; omit to use the provider's current default
548
+ reasoningEffort: low
549
+ costTier: low
550
+ capabilities: [mechanical-edits, tests]
551
+ - id: gemini-auto
552
+ agent: gemini # omitted model means the account's current Auto/default route
553
+ costTier: low
554
+ capabilities: [large-context, implementation]
555
+ ```
556
+
557
+ Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
558
+ baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
559
+ eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
560
+ "intelligence score". Candidate model IDs come from project configuration while setup defaults
561
+ prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
562
+ independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
563
+ after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
564
+ time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
565
+ cannot overwrite a shared registry file.
566
+
567
+ Routing is not free: it adds one controller call per story. It is most promising when a bounded
568
+ worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
569
+ on your own backlog rather than assuming a win.
570
+
571
+ ### Performance budgets: efficiency as a gate, not a style
494
572
 
495
573
  Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
496
574
  be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
@@ -517,12 +595,13 @@ be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two lev
517
595
  The loop trusts **verify**, not the agent's exit code: a story whose tests are green is
518
596
  committed even if the agent process exited non-zero (a common Windows `.cmd`-wrapper ghost).
519
597
  A failing verify is retried up to `verify.retries` times (default 1) so a transient flake
520
- self-heals while a real failure still blocks.
598
+ self-heals while a real failure still blocks. Structured acceptance criteria are then verified
599
+ individually; an unrelated green suite cannot satisfy a criterion without its proof command.
521
600
 
522
601
  `.yoke/loop-status.json`, `.yoke/loop.log`, `.yoke/loop.lock`, its takeover/recovery leases, lock/decision temp files, `.yoke/story-durations.json`,
523
602
  `.yoke/ambiguity.md`, and the critical-decision request/answering files are runtime artifacts;
524
603
  `yoke retrofit` gitignores them (along with
525
- `.yoke/worktrees/`, `.yoke/backup/`, and `.yoke/proof/`) so they never trip the clean-tree gate.
604
+ `.yoke/worktrees/`, `.yoke/backup/`, `.yoke/proof/`, and `.yoke/changes/`) so they never trip the clean-tree gate.
526
605
 
527
606
  ### Single-flight guard + cleanup
528
607
 
@@ -532,10 +611,11 @@ stale takeover is serialized by `.yoke/loop.lock.takeover`. A second invocation
532
611
  `Another loop is already running here (pid …). If that is wrong, run: yoke loop cleanup`. A lock
533
612
  whose holder process is dead is taken over automatically (with a warning).
534
613
 
535
- **`yoke loop cleanup [dir]`** removes what a crashed loop leaves behind: every worktree under
536
- `.yoke/worktrees/` (via `git worktree remove --force` + `prune` — user-created worktrees are
537
- never touched) and a **stale** lock file. A live lock is reported and left alone. Exits `0`
538
- when everything cleaned, `1` if any removal failed. If a machine/process crash leaves the cleanup
614
+ **`yoke loop cleanup [dir]`** reaps only runner process trees recorded by this project and removes
615
+ a stale lock. Yoke-created worktrees are **retained by default** and listed in the output; pass
616
+ `--remove-worktrees` to remove `.yoke/worktrees/*` with `git worktree remove --force` + `prune`.
617
+ User-created worktrees are never touched. A live lock is reported and left alone. Exits `0` when
618
+ cleanup succeeds, `1` if any requested removal fails. If a machine/process crash leaves the cleanup
539
619
  recovery lease itself behind, an operator can run
540
620
  `yoke loop cleanup . --discard-stale-recovery`; Yoke refuses while its recorded PID is alive, and
541
621
  the force flag must not be run concurrently.
@@ -652,24 +732,24 @@ Yoke's guardrails are **mechanical, not advisory** — the loop blocks on a dirt
652
732
  | Best for | rapid exploration / migration / onboarding | systematic refactoring in typed codebases |
653
733
  | Caveat | heuristic edges; static index can go stale | one language server per language |
654
734
 
655
- ## 🪙 Token efficiency
735
+ ## 🪙 Token efficiency
656
736
 
657
737
  Yoke attacks tokens on two complementary surfaces:
658
738
 
659
739
  - **rtk** compresses noisy command/tool output before it enters context (wired as a hook/instruction per agent).
660
- - The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
661
-
662
- A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
663
- completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
664
- controller selected Luna for every bounded implementation story. Including controller overhead,
665
- the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
666
- and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
667
-
668
- The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
669
- controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
670
- emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
671
- inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
672
- [`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
740
+ - The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
741
+
742
+ A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
743
+ completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
744
+ controller selected Luna for every bounded implementation story. Including controller overhead,
745
+ the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
746
+ and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
747
+
748
+ The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
749
+ controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
750
+ emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
751
+ inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
752
+ [`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
673
753
 
674
754
  ## 🧩 Optional companions
675
755
 
@@ -693,8 +773,10 @@ canon/ # the source of truth — harness-agnostic
693
773
  AGENTS.md skills/ policy/ loop/ tools/ manifest.yaml
694
774
  src/
695
775
  canon/ # manifest schema + validator (yoke validate)
776
+ change/ # append-only change inbox · planning · independent coverage review
696
777
  retrofit/ # detect · plan · apply · planners (claude/codex/gemini) · tools
697
778
  loop/ # prd · gates · runner · verify · git/worktree · loop · run-command · lock · cleanup
779
+ quality/ # reference collection · blind critic · bounded repair · candidate comparison
698
780
  new/ # yoke new — greenfield bootstrap
699
781
  prd/ # yoke prd draft|check — idea → stories + lint gate
700
782
  review/ # yoke review — cross-model diff gate
@@ -706,14 +788,14 @@ docs/superpowers/ # the spec and every component's implementation plan
706
788
 
707
789
  ## 🗺️ Roadmap
708
790
 
709
- Yoke 1.1's completed release work moved to the changelog. Remaining, explicitly scoped work
710
- is tracked in [`TODOS.md`](TODOS.md), including provider subprocess wiring for the tested
711
- parallel dispatcher, broader benchmark samples, native output schemas, and release provenance.
791
+ Completed release work lives in the changelog. Remaining, explicitly scoped work is tracked in
792
+ [`TODOS.md`](TODOS.md), including broader benchmark samples, native output schemas, and signed
793
+ release provenance.
712
794
 
713
795
  ## 🧪 Development
714
796
 
715
797
  ```bash
716
- npm test # vitest (582 tests)
798
+ npm test # vitest (928 tests)
717
799
  npm run build # tsc, no emit errors
718
800
  npm run yoke -- validate canon
719
801
  ```
package/TODOS.md CHANGED
@@ -1,8 +1,5 @@
1
1
  # Yoke follow-up work
2
2
 
3
- - Wire the tested async parallel dispatcher to provider subprocess workers. Until then the CLI
4
- rejects `--parallel=N` for `N > 1`; scheduler, claims, and merge queue APIs are available
5
- without claiming a CLI speed-up.
6
3
  - Add provider-native output schemas when all three CLIs expose compatible stable APIs.
7
4
  - Expand benchmark fixtures and collect multiple authenticated samples per provider/model.
8
5
  - Add signed provenance and attestations to npm and GitHub releases.