@hecer/yoke 1.2.1 → 1.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2,7 +2,7 @@
2
2
  "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
3
3
  "name": "yoke",
4
4
  "displayName": "Yoke",
5
- "version": "1.2.1",
5
+ "version": "1.3.0",
6
6
  "description": "Cross-agent coding harness: one curated skill canon (TDD, brainstorming, plans, reviews, shipping, design verification) plus mechanical safety gates and an autonomous loop via the yoke CLI.",
7
7
  "author": { "name": "HECer", "url": "https://github.com/HECer" },
8
8
  "homepage": "https://github.com/HECer/yoke#readme",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "yoke",
3
- "version": "1.2.1",
3
+ "version": "1.3.0",
4
4
  "description": "Cross-agent coding discipline, mechanical gates, and release workflows",
5
5
  "skills": "./canon/skills/",
6
6
  "hooks": "./hooks/hooks.json"
package/CHANGELOG.md CHANGED
@@ -1,27 +1,47 @@
1
- # Changelog
2
-
3
- ## 1.2.1 — 2026-08-02
4
-
5
- ### Fixed
6
- - `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
7
- - Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
8
-
9
- ## 1.2.0 — 2026-08-02
10
-
11
- ### Added
12
- - Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
13
- - A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
14
- - Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
15
-
16
- ### Changed
17
- - Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
18
- - Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
19
-
20
- ### Fixed
21
- - Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
22
- - Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
23
-
24
- ## 1.1.0 — 2026-07-30
1
+ # Changelog
2
+
3
+ ## 1.3.0 — 2026-08-09
4
+
5
+ ### Added
6
+ - Teams can submit change requests at any time with `yoke change add` and inspect the append-only inbox with `yoke change status`; queued requests are planned at safe loop boundaries by Claude, Codex, or Gemini without interrupting active story work.
7
+ - Acceptance criteria can carry stable IDs and executable verification commands, and Yoke records per-criterion evidence before a story can be marked complete.
8
+ - Projects can configure an integrated completion command that proves the final system flow after all story-level gates pass.
9
+
10
+ ### Changed
11
+ - New projects use strict criterion-proof requirements by default, while existing boolean-only criteria remain readable for compatibility.
12
+ - PRD authoring, schemas, loop guidance, generated configuration, and provider instructions now treat completion as an ephemeral readiness state: new requests create more stories instead of prematurely forcing a release boundary.
13
+
14
+ ### Fixed
15
+ - Broad green test suites can no longer satisfy unrelated structured acceptance criteria: proof commands must use an approved test runner, contain the criterion ID, and avoid shell operators.
16
+ - Change-intake failures block safely; an independent coverage pass rejects omitted requested outcomes, crash recovery retains uncommitted requests, and concurrent PRD edits are detected instead of silently overwriting user work.
17
+ - Integrated completion checks prevent locally finished stories from masking broken cross-component flows such as authentication callbacks or purchase-to-entitlement activation.
18
+
19
+ ### Security
20
+ - Updated the transitive development dependency `nanoid` to a non-vulnerable release so the complete CI dependency audit is clean.
21
+ - Restricted model-authored criterion proof to criterion-targeted test commands and blocked shell control operators and runner-prefix spoofing before host execution.
22
+
23
+ ## 1.2.1 — 2026-08-02
24
+
25
+ ### Fixed
26
+ - `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
27
+ - Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
28
+
29
+ ## 1.2.0 — 2026-08-02
30
+
31
+ ### Added
32
+ - Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
33
+ - A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
34
+ - Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
35
+
36
+ ### Changed
37
+ - Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
38
+ - Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
39
+
40
+ ### Fixed
41
+ - Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
42
+ - Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
43
+
44
+ ## 1.1.0 — 2026-07-30
25
45
 
26
46
  ### Added
27
47
  - Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
package/README.md CHANGED
@@ -2,8 +2,8 @@
2
2
 
3
3
  # 🐂 Yoke
4
4
 
5
- <!-- yoke:version:start -->1.2.1<!-- yoke:version:end -->
6
- <!-- yoke:tests:start -->582<!-- yoke:tests:end -->
5
+ <!-- yoke:version:start -->1.3.0<!-- yoke:version:end -->
6
+ <!-- yoke:tests:start -->657<!-- yoke:tests:end -->
7
7
  <!-- yoke:skills:start -->29<!-- yoke:skills:end -->
8
8
  <!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
9
9
 
@@ -17,7 +17,7 @@
17
17
  [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](#-license)
18
18
  ![Node](https://img.shields.io/badge/node-%E2%89%A520-339933?logo=node.js&logoColor=white)
19
19
  ![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
20
- ![Tests](https://img.shields.io/badge/tests-559%20passing-brightgreen.svg)
20
+ ![Tests](https://img.shields.io/badge/tests-657%20passing-brightgreen.svg)
21
21
  ![Agents](https://img.shields.io/badge/agents-Claude%20%7C%20Codex%20%7C%20Gemini-8A2BE2)
22
22
  ![Built with TDD](https://img.shields.io/badge/built%20with-TDD%20%2B%20review-ff69b4.svg)
23
23
 
@@ -25,7 +25,7 @@
25
25
 
26
26
  </div>
27
27
 
28
- > **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
28
+ > **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
29
29
 
30
30
  Yoke 1.1 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
31
31
  is explicit; reviews require a schema-valid verdict and a different model unless
@@ -72,7 +72,7 @@ $ ls reading-app/.yoke/proof/STORY-2/
72
72
  home.png list.png # photographic evidence, labelled per story
73
73
  ```
74
74
 
75
- Every claim in that transcript is enforced by code paths with tests behind them — 559 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
75
+ Every claim in that transcript is enforced by code paths with tests behind them — 657 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
76
76
 
77
77
  ## 🚀 Quickstart
78
78
 
@@ -85,7 +85,7 @@ yoke new my-app --idea="a CLI that tracks reading lists"
85
85
  yoke loop on my-app && yoke loop run my-app --isolate
86
86
 
87
87
  # — or retrofit an existing project —
88
- yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
88
+ yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
89
89
  yoke validate canon # sanity-check the canon
90
90
  yoke loop run /path/to/project --isolate --reviewer=codex --max=20
91
91
  ```
@@ -104,7 +104,7 @@ The canon is also packaged as a Claude Code plugin — the repo is its own marke
104
104
  That gives you all canon skills under the `yoke:` namespace (e.g. `yoke:tdd`, `yoke:review`) inside Claude Code — no retrofit needed. The `yoke` CLI (loop, gates, retrofit for Codex/Gemini) still comes from `npm i -g @hecer/yoke`. Gemini CLI users can likewise `gemini extensions install https://github.com/HECer/yoke`.
105
105
 
106
106
  For Codex, no preinstalled skill is required: run `npx @hecer/yoke setup .` in a terminal, or
107
- ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
107
+ ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
108
108
  `.agents/skills/`, including `yoke-retrofit` and `yoke-workflow`; start a fresh Codex task if an
109
109
  already-open task does not discover newly installed skills. The npm package also contains
110
110
  `.codex-plugin/plugin.json` for Codex plugin hosts.
@@ -138,7 +138,7 @@ Yoke is meant to be operated *by* your coding agent — after a retrofit, the ag
138
138
 
139
139
  > **Watch / unblock** — *"Run `yoke loop status .`. If it says BLOCKED, run the project's verify command, find the root cause, fix it without weakening tests, then continue the loop."*
140
140
 
141
- > ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
141
+ > ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
142
142
 
143
143
  > ⚠️ **Never kill agent processes by name or command-line pattern** (e.g. every process matching `dangerously-skip-permissions`): on a machine running several yoke projects, that takes down the *healthy* runners of the other projects mid-story — they stall and their loops block. `yoke loop cleanup` is the scoped alternative: each watchdog records its pids in the project's `.yoke/runner.pid`, and cleanup kills exactly those recorded trees — nothing else on the machine.
144
144
 
@@ -148,14 +148,15 @@ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`)
148
148
 
149
149
  | Command | What it does | Exit codes |
150
150
  |---|---|---|
151
- | `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
151
+ | `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
152
152
  | `yoke validate [canonDir]` | Validate the canon (schema, frontmatter, templates) | `0` valid · `1` errors |
153
153
  | `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
154
154
  | `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
155
155
  | `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
156
156
  | `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
157
+ | `yoke change add\|status [dir] [--idea=]` | Queue a change at any time; the loop turns it into append-only stories at the next safe boundary | `0` · `1` invalid inbox/request |
157
158
  | `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
158
- | `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` is unlimited by default and `--max=N` creates an intentional batch cap; `decision` shows a critical stop, `answer` records it and resumes, `resume` retries a failed restart with the preserved safety options | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
159
+ | `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` is unlimited by default and `--max=N` creates an intentional batch cap; `decision` shows a critical stop, `answer` records it and resumes, `resume` retries a failed restart with the preserved safety options | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
159
160
  | `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
160
161
  | `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
161
162
  | `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
@@ -310,28 +311,34 @@ Opt-in; `yoke setup` recommends enabling it for new installs, while `retrofit` a
310
311
 
311
312
  ```mermaid
312
313
  flowchart LR
313
- A[pick next PRD story] --> B{clean worktree?}
314
+ I[consume queued change<br/>as new stories] --> A[pick next PRD story]
315
+ A --> B{clean worktree?}
314
316
  B -- no --> X[blocked]
315
317
  B -- yes --> C{acceptance<br/>criteria?}
316
318
  C -- no --> X
317
319
  C -- yes --> D[agent implements<br/>one story]
318
- D --> E{tests green?}
320
+ D --> E{suite + criterion<br/>proof green?}
319
321
  E -- no --> X
320
322
  E -- yes --> F{reviewer<br/>approves?}
321
323
  F -- no --> X
322
324
  F -- yes --> G[commit + mark passes:true<br/>+ proof in .yoke/proof/]
323
- G --> A
325
+ G --> I
326
+ I --> H{all stories pass?}
327
+ H -- yes --> J{integrated system<br/>gate green?}
328
+ J -- no --> X
329
+ J -- yes --> K[current backlog ready]
324
330
  ```
325
331
 
326
332
  ```bash
327
333
  yoke loop on . # enable (recorded in .yoke/config.yaml)
328
334
  yoke loop status . # show state + PRD progress
335
+ yoke change add . --idea="Add passkey login" # safe while the loop runs
329
336
  yoke loop run . \
330
337
  --runner=codex \ # implement with Codex…
331
338
  --reviewer=claude \ # …review with Claude (role separation)
332
339
  --isolate \ # each story in a throwaway git worktree
333
- --decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
334
- # Optional: add --max=20 only when this run should stop after a bounded batch.
340
+ --decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
341
+ # Optional: add --max=20 only when this run should stop after a bounded batch.
335
342
  yoke loop off . # disable
336
343
  ```
337
344
 
@@ -342,11 +349,32 @@ yoke loop off . # disable
342
349
  title: Add a health endpoint
343
350
  priority: 1 # lower = higher priority
344
351
  acceptance: # Definition of Done (required, else blocked)
345
- - GET /health returns 200
352
+ - id: health-returns-200
353
+ text: GET /health returns 200
354
+ verify: [npm run test:health-returns-200]
355
+ - id: health-rejects-post
356
+ text: POST /health returns 405
357
+ verify: [npm run test:health-rejects-post]
346
358
  passes: false # the loop sets this true only on green tests
347
359
  ```
348
360
 
349
- The loop stops when every story is `passes: true`. State lives **outside the model context** — the PRD file plus git — so each iteration is fresh. The PRD is re-read from disk at every story boundary, so new stories appended to `.yoke/prd.yaml` **while the loop is running** are picked up at the next iteration — no restart needed.
361
+ New projects default to `verify.requireCriteria: true`: every story has 2–5 behavioral criteria.
362
+ Each criterion ID must occur in its single, approved test command; shell operators and broad,
363
+ untargeted suites are rejected. Yoke records each result in `.yoke/proof/<story>/evidence.json`. Configure optional
364
+ `completion.command` for integrated journeys such as purchase → entitlement → relaunch or
365
+ magic-link → callback → authenticated app. It runs whenever the current backlog has no open
366
+ stories; this is readiness, not a release.
367
+
368
+ No generic tool can infer whether arbitrary test code perfectly represents product meaning. Yoke
369
+ closes the mechanical false-done paths—targeted evidence, coverage review, clean committed state,
370
+ and integrated journeys—while the project still owns the correctness of its tests and production
371
+ observability.
372
+
373
+ State lives **outside the model context** — the PRD file plus git — so each iteration is fresh.
374
+ Use `yoke change add` at any time. Its ignored append-only inbox is consumed at the next story
375
+ boundary. A separate coverage pass must confirm that every requested outcome maps to behavioral
376
+ criteria before Yoke appends and commits the new stories; existing stories are never rewritten and
377
+ no restart is needed.
350
378
 
351
379
  ### Watching a run
352
380
 
@@ -370,9 +398,9 @@ Every iteration emits token-free, harness-side feedback (Node console + local fi
370
398
  NDJSON line on stdout (`{"type":"status","state":"running","phase":"verifying",…}` — the
371
399
  same shape as `loop-status.json`), the human narrative moves off stdout (the final summary
372
400
  goes to stderr), and a consumer can follow the stream line by line instead of polling the file.
373
- Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
374
- output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
375
- and cost fields when available. Missing values stay absent—Yoke does not estimate them.
401
+ Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
402
+ output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
403
+ and cost fields when available. Missing values stay absent—Yoke does not estimate them.
376
404
 
377
405
  ### Pausing a run
378
406
 
@@ -432,65 +460,65 @@ cannot begin because a provider/reviewer is unavailable or another process owns
432
460
  until a loop actually runs. To intentionally abandon an orphaned or stale private resume state,
433
461
  use `yoke loop resume . --discard`; pending decisions are never deleted by that command. Existing
434
462
  `loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` remain supported as compatibility aliases;
435
- new projects should use `decisionPolicy: auto|critical`.
436
-
437
- ### Adaptive model routing (explicit opt-in)
438
-
439
- `yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
440
- parent remains the strong planner/controller. Before each bounded story it receives only the
441
- story, acceptance criteria, and at most three eligible worker profiles, then returns one
442
- machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
443
- `SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
444
- runs so Yoke does not pay for two orchestration layers.
445
-
446
- **Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
447
- Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
448
- cover invocation and routing behavior for all three providers. The measured performance evidence
449
- below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
450
- until authenticated, repeated in-the-wild runs exist for those providers.
451
-
452
- ```yaml
453
- runner:
454
- agent: codex
455
- model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
456
- reasoningEffort: high
457
- routing:
458
- enabled: true # setup defaults false; setup --routing opts in
459
- strategy: balanced # balanced | cost | speed | quality
460
- maxCandidates: 3
461
- workers:
462
- - id: codex-light
463
- agent: codex
464
- reasoningEffort: low
465
- costTier: medium
466
- capabilities: [exploration, implementation, tests]
467
- - id: claude-fast
468
- agent: claude
469
- model: haiku # rolling alias; omit to use the provider's current default
470
- reasoningEffort: low
471
- costTier: low
472
- capabilities: [mechanical-edits, tests]
473
- - id: gemini-auto
474
- agent: gemini # omitted model means the account's current Auto/default route
475
- costTier: low
476
- capabilities: [large-context, implementation]
477
- ```
478
-
479
- Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
480
- baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
481
- eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
482
- "intelligence score". Candidate model IDs come from project configuration while setup defaults
483
- prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
484
- independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
485
- after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
486
- time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
487
- cannot overwrite a shared registry file.
488
-
489
- Routing is not free: it adds one controller call per story. It is most promising when a bounded
490
- worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
491
- on your own backlog rather than assuming a win.
492
-
493
- ### Performance budgets: efficiency as a gate, not a style
463
+ new projects should use `decisionPolicy: auto|critical`.
464
+
465
+ ### Adaptive model routing (explicit opt-in)
466
+
467
+ `yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
468
+ parent remains the strong planner/controller. Before each bounded story it receives only the
469
+ story, acceptance criteria, and at most three eligible worker profiles, then returns one
470
+ machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
471
+ `SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
472
+ runs so Yoke does not pay for two orchestration layers.
473
+
474
+ **Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
475
+ Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
476
+ cover invocation and routing behavior for all three providers. The measured performance evidence
477
+ below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
478
+ until authenticated, repeated in-the-wild runs exist for those providers.
479
+
480
+ ```yaml
481
+ runner:
482
+ agent: codex
483
+ model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
484
+ reasoningEffort: high
485
+ routing:
486
+ enabled: true # setup defaults false; setup --routing opts in
487
+ strategy: balanced # balanced | cost | speed | quality
488
+ maxCandidates: 3
489
+ workers:
490
+ - id: codex-light
491
+ agent: codex
492
+ reasoningEffort: low
493
+ costTier: medium
494
+ capabilities: [exploration, implementation, tests]
495
+ - id: claude-fast
496
+ agent: claude
497
+ model: haiku # rolling alias; omit to use the provider's current default
498
+ reasoningEffort: low
499
+ costTier: low
500
+ capabilities: [mechanical-edits, tests]
501
+ - id: gemini-auto
502
+ agent: gemini # omitted model means the account's current Auto/default route
503
+ costTier: low
504
+ capabilities: [large-context, implementation]
505
+ ```
506
+
507
+ Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
508
+ baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
509
+ eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
510
+ "intelligence score". Candidate model IDs come from project configuration while setup defaults
511
+ prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
512
+ independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
513
+ after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
514
+ time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
515
+ cannot overwrite a shared registry file.
516
+
517
+ Routing is not free: it adds one controller call per story. It is most promising when a bounded
518
+ worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
519
+ on your own backlog rather than assuming a win.
520
+
521
+ ### Performance budgets: efficiency as a gate, not a style
494
522
 
495
523
  Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
496
524
  be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
@@ -517,12 +545,13 @@ be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two lev
517
545
  The loop trusts **verify**, not the agent's exit code: a story whose tests are green is
518
546
  committed even if the agent process exited non-zero (a common Windows `.cmd`-wrapper ghost).
519
547
  A failing verify is retried up to `verify.retries` times (default 1) so a transient flake
520
- self-heals while a real failure still blocks.
548
+ self-heals while a real failure still blocks. Structured acceptance criteria are then verified
549
+ individually; an unrelated green suite cannot satisfy a criterion without its proof command.
521
550
 
522
551
  `.yoke/loop-status.json`, `.yoke/loop.log`, `.yoke/loop.lock`, its takeover/recovery leases, lock/decision temp files, `.yoke/story-durations.json`,
523
552
  `.yoke/ambiguity.md`, and the critical-decision request/answering files are runtime artifacts;
524
553
  `yoke retrofit` gitignores them (along with
525
- `.yoke/worktrees/`, `.yoke/backup/`, and `.yoke/proof/`) so they never trip the clean-tree gate.
554
+ `.yoke/worktrees/`, `.yoke/backup/`, `.yoke/proof/`, and `.yoke/changes/`) so they never trip the clean-tree gate.
526
555
 
527
556
  ### Single-flight guard + cleanup
528
557
 
@@ -652,24 +681,24 @@ Yoke's guardrails are **mechanical, not advisory** — the loop blocks on a dirt
652
681
  | Best for | rapid exploration / migration / onboarding | systematic refactoring in typed codebases |
653
682
  | Caveat | heuristic edges; static index can go stale | one language server per language |
654
683
 
655
- ## 🪙 Token efficiency
684
+ ## 🪙 Token efficiency
656
685
 
657
686
  Yoke attacks tokens on two complementary surfaces:
658
687
 
659
688
  - **rtk** compresses noisy command/tool output before it enters context (wired as a hook/instruction per agent).
660
- - The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
661
-
662
- A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
663
- completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
664
- controller selected Luna for every bounded implementation story. Including controller overhead,
665
- the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
666
- and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
667
-
668
- The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
669
- controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
670
- emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
671
- inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
672
- [`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
689
+ - The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
690
+
691
+ A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
692
+ completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
693
+ controller selected Luna for every bounded implementation story. Including controller overhead,
694
+ the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
695
+ and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
696
+
697
+ The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
698
+ controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
699
+ emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
700
+ inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
701
+ [`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
673
702
 
674
703
  ## 🧩 Optional companions
675
704
 
@@ -693,6 +722,7 @@ canon/ # the source of truth — harness-agnostic
693
722
  AGENTS.md skills/ policy/ loop/ tools/ manifest.yaml
694
723
  src/
695
724
  canon/ # manifest schema + validator (yoke validate)
725
+ change/ # append-only change inbox · planning · independent coverage review
696
726
  retrofit/ # detect · plan · apply · planners (claude/codex/gemini) · tools
697
727
  loop/ # prd · gates · runner · verify · git/worktree · loop · run-command · lock · cleanup
698
728
  new/ # yoke new — greenfield bootstrap
@@ -713,7 +743,7 @@ parallel dispatcher, broader benchmark samples, native output schemas, and relea
713
743
  ## 🧪 Development
714
744
 
715
745
  ```bash
716
- npm test # vitest (582 tests)
746
+ npm test # vitest (657 tests)
717
747
  npm run build # tsc, no emit errors
718
748
  npm run yoke -- validate canon
719
749
  ```
@@ -1,36 +1,51 @@
1
1
  # Loop Specification (Ralph + GSD)
2
2
 
3
- The autonomous loop is OPTIONAL and toggle-able:
4
-
5
- - `yoke loop on` / `yoke loop off` — enable/disable (recorded in `.yoke/config.yaml`, default off).
6
- - `yoke loop status` — show enabled state + PRD progress.
7
- - `yoke loop run [--max=N] [--isolate] [--decision-policy=auto|critical]` — run the loop (default cap 25 iterations).
8
- - `yoke loop decision` / `yoke loop answer --choice=<id>` — inspect and answer a structured critical stop; answering records a human-owned, decision-file-only commit and resumes by default with the original runner, isolation, review, permission, timeout, and policy settings. `yoke loop resume` retries a restart that could not begin without weakening those settings.
9
-
10
- Pass `--isolate` to run each iteration in a fresh git worktree: the agent works on a throwaway checkout, and only a verified, committed story is fast-forwarded back into the main tree. A failed iteration never touches your working tree. Requires `.yoke/prd.yaml` to be committed to git, since the worktree is a checkout of HEAD.
11
-
12
- Pass `--review` (or `--reviewer=<claude|codex|gemini>` for a different agent) to add a role-separated review step: after the tests pass, an independent reviewer agent must approve the change before the story is committed and marked done. A rejection blocks the story (no commit). The reviewer is a fresh agent pass — the implementer never reviews its own work.
13
-
14
- Pass `--json` for machine mode: each status transition is emitted as one NDJSON line on stdout (the `.yoke/loop-status.json` shape, tagged `"type":"status"`) instead of the human narrative, so a supervisor can consume the stream instead of polling the file.
15
-
16
- When enabled and run, each iteration:
17
-
18
- 1. Pre-dispatch gate: the git worktree must be clean, else `blocked`.
19
- 2. Pick the highest-priority unfinished PRD story (`.yoke/prd.yaml`).
20
- 3. Stop-the-Line gate: the story must have acceptance criteria, else `blocked`.
21
- 4. Run a fresh agent to implement ONE story. Runner precedence is explicit `--runner`, configured `runner.agent`, active agent host, then the first configured agent. The loop refuses to start if that CLI is not installed.
22
- With `decisionPolicy: auto`, routine ambiguity is resolved from the approved plan and project conventions. With `critical`, only high-impact architecture, security/privacy, destructive data, material-cost, compliance, or irreversible choices may produce `.yoke/decision-request.yaml`; the loop validates its bounded single-line fields, unique options, and active story ID, blocks before verify, and preserves it for `yoke loop answer`.
23
- 5. Run the project's verify command (config `verify.command`, or detected `npm test`).
24
- **Verify is the source of truth** — the agent's exit code is advisory, so a spurious
25
- non-zero exit (e.g. a Windows `.cmd` wrapper) cannot block a story whose tests are green.
26
- A failing verify is retried up to `verify.retries` times (default 1) so a transient flake
27
- self-heals; a real failure still fails. Only if verify passes is the story marked
28
- `passes: true`, committed atomically, and a decision logged. If verify fails: `blocked`.
29
- 6. Stop when all stories `passes: true` (`complete`), or the iteration cap is reached (`cap-reached`).
30
-
31
- A supervisor can pause the loop by creating `.yoke/loop.pause`: at the next story boundary (before the next story is selected — the running story always finishes) the loop consumes the file, records `paused` in the status file and log, and exits with code `3`. Running `yoke loop run` again resumes.
32
-
33
- State lives outside the model context: the PRD file + git. The agent runner is pluggable. The PRD is re-read from disk at every story boundary, so stories appended to `.yoke/prd.yaml` mid-run are picked up at the next iteration without a restart.
3
+ The autonomous loop is optional and toggle-able:
4
+
5
+ - `yoke loop on` / `yoke loop off` — enable or disable it in `.yoke/config.yaml`.
6
+ - `yoke loop status` — show enabled state and backlog progress.
7
+ - `yoke loop run [--max=N] [--isolate] [--decision-policy=auto|critical]` — run until the current backlog is green or a gate blocks.
8
+ - `yoke change add --idea="..."` — queue a product change at any time, including while the loop is running.
9
+ - `yoke loop decision` / `yoke loop answer --choice=<id>` — inspect and answer a structured critical stop.
10
+
11
+ Pass `--isolate` to implement each story in a fresh git worktree. Only a verified, committed
12
+ story is fast-forwarded to the main tree. Pass `--review` or `--reviewer=<provider>` to require
13
+ a separate, schema-validated review. Pass `--json` for NDJSON status on stdout.
14
+
15
+ At every story boundary, Yoke consumes at most one queued change. The configured Claude,
16
+ Codex, or Gemini provider may propose only new stories in a separate runtime file. A fresh
17
+ coverage-review pass must account for every distinct requested outcome before Yoke validates
18
+ strict criterion evidence, appends the stories itself, commits only the PRD, and leaves existing
19
+ stories untouched. The request stays pending on any failure or uncovered outcome.
20
+
21
+ For each story:
22
+
23
+ 1. Require a clean git worktree.
24
+ 2. Pick the highest-priority ready unfinished story.
25
+ 3. Stop the line if acceptance is empty. With `verify.requireCriteria: true`, every criterion
26
+ must be structured. Every structured criterion, including in compatible legacy projects,
27
+ must use a single approved test command containing its criterion ID and no shell operators.
28
+ 4. Run a fresh configured provider to implement exactly one story. Under the `critical`
29
+ decision policy, only high-impact architecture, security/privacy, destructive data,
30
+ material cost, compliance, or irreversible choices may pause for a human decision.
31
+ 5. Run every structured criterion's targeted commands and write
32
+ `.yoke/proof/<story>/evidence.json`. Then run project-wide `verify.command` (or detected
33
+ `npm test`). Performance, audit, and independent review gates follow when configured. Any
34
+ failure blocks, and no proof command runs after review.
35
+ 6. Only after all gates pass, mark the story `passes: true`, log the decision, and commit
36
+ atomically. A failed commit restores the PRD state.
37
+ 7. When all current stories pass, run optional `completion.command` against the integrated
38
+ system. Only a green result reports `complete`; otherwise the loop blocks. This readiness
39
+ result is ephemeral, not a release and not a freeze on future changes.
40
+
41
+ A supervisor can pause the loop by creating `.yoke/loop.pause`. The running story finishes;
42
+ the signal is consumed at the next story boundary and the process exits with code `3`.
43
+
44
+ State lives outside model context: PRD, git, and the ignored `.yoke/changes/` inbox. All are
45
+ re-read at story boundaries, so a request queued mid-run becomes additional stories without a
46
+ restart.
34
47
 
35
48
  ## Limitations
36
- - The loop verifies via the project's test command and an optional agent review; it has no formal merge-queue or multi-reviewer quorum.
49
+
50
+ Yoke cannot infer the correct end-to-end journey command. Projects that need integrated
51
+ readiness must configure `completion.command`, for example a Playwright journey suite.
@@ -1,6 +1,6 @@
1
1
  # PRD Schema
2
2
 
3
- The loop is driven by a versioned PRD file. Each story:
3
+ The loop is driven by a continuous PRD backlog. Each story:
4
4
 
5
5
  ```yaml
6
6
  - id: STORY-1
@@ -9,11 +9,35 @@ The loop is driven by a versioned PRD file. Each story:
9
9
  needs: [] # optional dependency IDs; no unknown IDs, self-links, or cycles
10
10
  area: api # optional collision domain for parallel scheduling
11
11
  agent: codex # optional claude|codex|gemini affinity
12
- acceptance: # Definition of Done (required before implementation)
13
- - The endpoint returns 200 for a valid request.
14
- passes: false # set true only when acceptance is met and tests are green
12
+ acceptance:
13
+ - id: valid-request-returns-200
14
+ text: The endpoint returns 200 for a valid request.
15
+ verify: [npm run test:valid-request-returns-200]
16
+ - id: invalid-request-returns-400
17
+ text: The endpoint returns 400 for an invalid request.
18
+ verify: [npm run test:invalid-request-returns-400]
19
+ passes: false # Yoke-owned; true only after all gates pass and the commit lands
15
20
  ```
16
21
 
17
- Stop condition: every story has `passes: true`.
22
+ Each new story has 2–5 acceptance criteria. Every criterion has a stable `id`, observable
23
+ behavioral `text`, and one or more executable `verify` commands. Each entry is one approved test
24
+ command, contains the normalized criterion ID, and contains no shell control operator. Yoke runs and records each criterion separately in
25
+ `.yoke/proof/<story>/evidence.json`; a broad green suite cannot stand in for an untested
26
+ criterion. Legacy string criteria remain readable, but `verify.requireCriteria: true` blocks
27
+ them.
18
28
 
19
- Stories without `needs`, `area`, or `agent` retain the serial pre-1.0 behavior. A story is ready only when every ID in `needs` passes. The scheduler orders ready work by priority, avoids simultaneously active areas, and uses `agent` as an affinity hint.
29
+ The binding is deliberately mechanical, not an oracle for product meaning: Yoke can require a
30
+ criterion-targeted test command and a separate coverage review, but it cannot prove that arbitrary
31
+ test code faithfully models the real world. Critical cross-component behavior therefore also
32
+ belongs in a trusted `completion.command` journey suite.
33
+
34
+ `sourceChange` is an optional Yoke-owned request ID. The change inbox uses it to append new
35
+ stories idempotently; authors normally omit it.
36
+
37
+ Stories without `needs`, `area`, or `agent` retain serial behavior. A story is ready only when
38
+ every ID in `needs` passes. The scheduler orders ready work by priority, avoids simultaneously
39
+ active areas, and uses `agent` as an affinity hint.
40
+
41
+ The backlog is continuous, not a release object. A momentary stop condition is every story
42
+ having `passes: true`; if configured, `completion.command` must then prove the integrated
43
+ system before the loop reports `complete`.
@@ -1,5 +1,5 @@
1
1
  name: yoke-canon
2
- version: 1.1.0
2
+ version: 1.2.0
3
3
  agents: [claude, codex, gemini]
4
4
  skills:
5
5
  - { id: tdd, path: skills/tdd, kind: methodology }