@hecer/yoke 1.1.1 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/.codex-plugin/plugin.json +1 -1
  3. package/CHANGELOG.md +18 -3
  4. package/README.md +87 -21
  5. package/bench/README.md +46 -14
  6. package/bench/RESULTS.md +108 -11
  7. package/bench/analyze-routing-study.mjs +96 -0
  8. package/bench/fixtures/routing-queue/.yoke/config.yaml +30 -0
  9. package/bench/fixtures/routing-queue/.yoke/prd.yaml +24 -0
  10. package/bench/fixtures/routing-queue/bench-verify.mjs +9 -0
  11. package/bench/fixtures/routing-queue/package.json +9 -0
  12. package/bench/fixtures/routing-queue/src/task-queue.mjs +1 -0
  13. package/bench/fixtures/routing-queue/tests/STORY-1.test.mjs +68 -0
  14. package/bench/fixtures/routing-queue/tests/STORY-2.test.mjs +72 -0
  15. package/bench/results/routing-queue-codex-routing-off-2026-08-01T21-44-00.json +46 -0
  16. package/bench/results/routing-queue-codex-routing-on-2026-08-01T21-50-31.json +109 -0
  17. package/bench/results/yoke-codex-study-codex-routing-off-2026-08-02T07-58-56.json +65 -0
  18. package/bench/results/yoke-codex-study-codex-routing-off-2026-08-02T08-08-48.json +65 -0
  19. package/bench/results/yoke-codex-study-codex-routing-off-2026-08-02T08-34-45.json +65 -0
  20. package/bench/results/yoke-codex-study-codex-routing-on-2026-08-02T07-51-59.json +168 -0
  21. package/bench/results/yoke-codex-study-codex-routing-on-2026-08-02T08-18-59.json +168 -0
  22. package/bench/results/yoke-codex-study-codex-routing-on-2026-08-02T08-25-25.json +168 -0
  23. package/bench/results/yoke-large-codex-routing-off-2026-08-01T22-58-01.json +54 -0
  24. package/bench/results/yoke-large-codex-routing-on-2026-08-01T23-09-39.json +133 -0
  25. package/bench/run-large.mjs +188 -0
  26. package/bench/run.mjs +42 -21
  27. package/dist/agents/providers.js +19 -2
  28. package/dist/agents/telemetry.js +46 -7
  29. package/dist/cli.js +7 -5
  30. package/dist/loop/decision.js +3 -1
  31. package/dist/loop/loop.js +10 -0
  32. package/dist/loop/reporter.js +9 -0
  33. package/dist/loop/run-command.js +33 -5
  34. package/dist/loop/runner.js +68 -9
  35. package/dist/retrofit/config.js +21 -0
  36. package/dist/review/command.js +2 -2
  37. package/dist/routing/registry.js +92 -0
  38. package/dist/routing/router.js +195 -0
  39. package/dist/setup/command.js +23 -1
  40. package/gemini-extension.json +1 -1
  41. package/package.json +6 -4
@@ -2,7 +2,7 @@
2
2
  "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
3
3
  "name": "yoke",
4
4
  "displayName": "Yoke",
5
- "version": "1.1.0",
5
+ "version": "1.2.0",
6
6
  "description": "Cross-agent coding harness: one curated skill canon (TDD, brainstorming, plans, reviews, shipping, design verification) plus mechanical safety gates and an autonomous loop via the yoke CLI.",
7
7
  "author": { "name": "HECer", "url": "https://github.com/HECer" },
8
8
  "homepage": "https://github.com/HECer/yoke#readme",
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "yoke",
3
- "version": "1.1.0",
3
+ "version": "1.2.0",
4
4
  "description": "Cross-agent coding discipline, mechanical gates, and release workflows",
5
5
  "skills": "./canon/skills/",
6
6
  "hooks": "./hooks/hooks.json"
package/CHANGELOG.md CHANGED
@@ -1,6 +1,21 @@
1
- # Changelog
2
-
3
- ## 1.1.0 — 2026-07-30
1
+ # Changelog
2
+
3
+ ## 1.2.0 — 2026-08-02
4
+
5
+ ### Added
6
+ - Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
7
+ - A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
8
+ - Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
9
+
10
+ ### Changed
11
+ - Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
12
+ - Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
13
+
14
+ ### Fixed
15
+ - Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
16
+ - Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
17
+
18
+ ## 1.1.0 — 2026-07-30
4
19
 
5
20
  ### Added
6
21
  - Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
package/README.md CHANGED
@@ -2,8 +2,8 @@
2
2
 
3
3
  # 🐂 Yoke
4
4
 
5
- <!-- yoke:version:start -->1.1.1<!-- yoke:version:end -->
6
- <!-- yoke:tests:start -->559<!-- yoke:tests:end -->
5
+ <!-- yoke:version:start -->1.2.0<!-- yoke:version:end -->
6
+ <!-- yoke:tests:start -->582<!-- yoke:tests:end -->
7
7
  <!-- yoke:skills:start -->29<!-- yoke:skills:end -->
8
8
  <!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
9
9
 
@@ -25,7 +25,7 @@
25
25
 
26
26
  </div>
27
27
 
28
- > **TL;DR** — `yoke setup .` asks five questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
28
+ > **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
29
29
 
30
30
  Yoke 1.1 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
31
31
  is explicit; reviews require a schema-valid verdict and a different model unless
@@ -85,7 +85,7 @@ yoke new my-app --idea="a CLI that tracks reading lists"
85
85
  yoke loop on my-app && yoke loop run my-app --isolate
86
86
 
87
87
  # — or retrofit an existing project —
88
- yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions
88
+ yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
89
89
  yoke validate canon # sanity-check the canon
90
90
  yoke loop run /path/to/project --isolate --reviewer=codex --max=20
91
91
  ```
@@ -104,7 +104,7 @@ The canon is also packaged as a Claude Code plugin — the repo is its own marke
104
104
  That gives you all canon skills under the `yoke:` namespace (e.g. `yoke:tdd`, `yoke:review`) inside Claude Code — no retrofit needed. The `yoke` CLI (loop, gates, retrofit for Codex/Gemini) still comes from `npm i -g @hecer/yoke`. Gemini CLI users can likewise `gemini extensions install https://github.com/HECer/yoke`.
105
105
 
106
106
  For Codex, no preinstalled skill is required: run `npx @hecer/yoke setup .` in a terminal, or
107
- ask Codex to run the five-question Yoke setup flow. The retrofit writes native skills to
107
+ ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
108
108
  `.agents/skills/`, including `yoke-retrofit` and `yoke-workflow`; start a fresh Codex task if an
109
109
  already-open task does not discover newly installed skills. The npm package also contains
110
110
  `.codex-plugin/plugin.json` for Codex plugin hosts.
@@ -138,7 +138,7 @@ Yoke is meant to be operated *by* your coding agent — after a retrofit, the ag
138
138
 
139
139
  > **Watch / unblock** — *"Run `yoke loop status .`. If it says BLOCKED, run the project's verify command, find the root cause, fix it without weakening tests, then continue the loop."*
140
140
 
141
- > ⚠️ **Long runs from inside an agent session:** a multi-story `yoke loop run` outlives most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Rules of thumb: run the loop **in the background** (e.g. Claude Code's `run_in_background`), keep batches small (`--max=3..5`), poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
141
+ > ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
142
142
 
143
143
  > ⚠️ **Never kill agent processes by name or command-line pattern** (e.g. every process matching `dangerously-skip-permissions`): on a machine running several yoke projects, that takes down the *healthy* runners of the other projects mid-story — they stall and their loops block. `yoke loop cleanup` is the scoped alternative: each watchdog records its pids in the project's `.yoke/runner.pid`, and cleanup kills exactly those recorded trees — nothing else on the machine.
144
144
 
@@ -148,14 +148,14 @@ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`)
148
148
 
149
149
  | Command | What it does | Exit codes |
150
150
  |---|---|---|
151
- | `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop]` | Shared five-question setup for Claude, Codex, and Gemini; `--yes` applies supplied/default choices non-interactively | `0` · `1` invalid setup |
151
+ | `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
152
152
  | `yoke validate [canonDir]` | Validate the canon (schema, frontmatter, templates) | `0` valid · `1` errors |
153
153
  | `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
154
154
  | `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
155
155
  | `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
156
156
  | `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
157
157
  | `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
158
- | `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `decision` shows a critical stop, `answer` records it and resumes, `resume` retries a failed restart with the preserved safety options; cleanup deletes worktrees only with `--remove-worktrees` | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
158
+ | `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` is unlimited by default and `--max=N` creates an intentional batch cap; `decision` shows a critical stop, `answer` records it and resumes, `resume` retries a failed restart with the preserved safety options | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
159
159
  | `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
160
160
  | `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
161
161
  | `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
@@ -330,8 +330,8 @@ yoke loop run . \
330
330
  --runner=codex \ # implement with Codex…
331
331
  --reviewer=claude \ # …review with Claude (role separation)
332
332
  --isolate \ # each story in a throwaway git worktree
333
- --decision-policy=critical \ # pause only for high-impact decisions; routine choices stay autonomous
334
- --max=20
333
+ --decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
334
+ # Optional: add --max=20 only when this run should stop after a bounded batch.
335
335
  yoke loop off . # disable
336
336
  ```
337
337
 
@@ -370,11 +370,9 @@ Every iteration emits token-free, harness-side feedback (Node console + local fi
370
370
  NDJSON line on stdout (`{"type":"status","state":"running","phase":"verifying",…}` — the
371
371
  same shape as `loop-status.json`), the human narrative moves off stdout (the final summary
372
372
  goes to stderr), and a consumer can follow the stream line by line instead of polling the file.
373
- With a **claude** runner, json mode also switches the agent to `--output-format stream-json`
374
- and accounts its usage: statuses (file + stream) carry a cumulative
375
- `tokens: { inputTokens, outputTokens, model? }` field for the whole run — `model` is the
376
- last model id seen on the stream (e.g. `claude-opus-4-6-20260501`), omitted if the CLI
377
- never reported one.
373
+ Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
374
+ output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
375
+ and cost fields when available. Missing values stay absent—Yoke does not estimate them.
378
376
 
379
377
  ### Pausing a run
380
378
 
@@ -434,9 +432,65 @@ cannot begin because a provider/reviewer is unavailable or another process owns
434
432
  until a loop actually runs. To intentionally abandon an orphaned or stale private resume state,
435
433
  use `yoke loop resume . --discard`; pending decisions are never deleted by that command. Existing
436
434
  `loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` remain supported as compatibility aliases;
437
- new projects should use `decisionPolicy: auto|critical`.
438
-
439
- ### Performance budgets: efficiency as a gate, not a style
435
+ new projects should use `decisionPolicy: auto|critical`.
436
+
437
+ ### Adaptive model routing (explicit opt-in)
438
+
439
+ `yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
440
+ parent remains the strong planner/controller. Before each bounded story it receives only the
441
+ story, acceptance criteria, and at most three eligible worker profiles, then returns one
442
+ machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
443
+ `SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
444
+ runs so Yoke does not pay for two orchestration layers.
445
+
446
+ **Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
447
+ Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
448
+ cover invocation and routing behavior for all three providers. The measured performance evidence
449
+ below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
450
+ until authenticated, repeated in-the-wild runs exist for those providers.
451
+
452
+ ```yaml
453
+ runner:
454
+ agent: codex
455
+ model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
456
+ reasoningEffort: high
457
+ routing:
458
+ enabled: true # setup defaults false; setup --routing opts in
459
+ strategy: balanced # balanced | cost | speed | quality
460
+ maxCandidates: 3
461
+ workers:
462
+ - id: codex-light
463
+ agent: codex
464
+ reasoningEffort: low
465
+ costTier: medium
466
+ capabilities: [exploration, implementation, tests]
467
+ - id: claude-fast
468
+ agent: claude
469
+ model: haiku # rolling alias; omit to use the provider's current default
470
+ reasoningEffort: low
471
+ costTier: low
472
+ capabilities: [mechanical-edits, tests]
473
+ - id: gemini-auto
474
+ agent: gemini # omitted model means the account's current Auto/default route
475
+ costTier: low
476
+ capabilities: [large-context, implementation]
477
+ ```
478
+
479
+ Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
480
+ baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
481
+ eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
482
+ "intelligence score". Candidate model IDs come from project configuration while setup defaults
483
+ prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
484
+ independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
485
+ after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
486
+ time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
487
+ cannot overwrite a shared registry file.
488
+
489
+ Routing is not free: it adds one controller call per story. It is most promising when a bounded
490
+ worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
491
+ on your own backlog rather than assuming a win.
492
+
493
+ ### Performance budgets: efficiency as a gate, not a style
440
494
 
441
495
  Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
442
496
  be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
@@ -598,12 +652,24 @@ Yoke's guardrails are **mechanical, not advisory** — the loop blocks on a dirt
598
652
  | Best for | rapid exploration / migration / onboarding | systematic refactoring in typed codebases |
599
653
  | Caveat | heuristic edges; static index can go stale | one language server per language |
600
654
 
601
- ## 🪙 Token efficiency
655
+ ## 🪙 Token efficiency
602
656
 
603
657
  Yoke attacks tokens on two complementary surfaces:
604
658
 
605
659
  - **rtk** compresses noisy command/tool output before it enters context (wired as a hook/instruction per agent).
606
- - The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
660
+ - The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
661
+
662
+ A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
663
+ completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
664
+ controller selected Luna for every bounded implementation story. Including controller overhead,
665
+ the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
666
+ and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
667
+
668
+ The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
669
+ controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
670
+ emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
671
+ inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
672
+ [`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
607
673
 
608
674
  ## 🧩 Optional companions
609
675
 
@@ -647,7 +713,7 @@ parallel dispatcher, broader benchmark samples, native output schemas, and relea
647
713
  ## 🧪 Development
648
714
 
649
715
  ```bash
650
- npm test # vitest (559 tests)
716
+ npm test # vitest (579 tests)
651
717
  npm run build # tsc, no emit errors
652
718
  npm run yoke -- validate canon
653
719
  ```
package/bench/README.md CHANGED
@@ -1,42 +1,74 @@
1
1
  # Yoke benchmark — tokens · speed · quality
2
2
 
3
- Reproducible cross-runner benchmark for the Yoke loop. One fixed fixture project, the same
4
- PRD for every runner, three measured dimensions:
3
+ Reproducible cross-runner and routing A/B benchmarks for the Yoke loop. Fixed fixture projects,
4
+ the same PRD within each comparison, and three measured dimensions:
5
+
6
+ Adaptive routing is implemented for Claude Code, Codex CLI, and Gemini CLI. The checked-in
7
+ multi-run performance study currently measures Codex only; provider support tests are not treated
8
+ as performance evidence for Claude or Gemini.
5
9
 
6
10
  | Dimension | How it is measured |
7
11
  |---|---|
8
- | **Tokens** | The loop's own token hook (`.yoke/loop-status.json`, claude runner via `--output-format stream-json`; model id included). Gemini/Codex runners do not report usage yet — recorded as `null`, a documented gap. |
12
+ | **Tokens** | The loop's own provider telemetry in `.yoke/loop-status.json`, split into total input, cached input, fresh input (`input − cached`), output, and reasoning tokens when the provider reports them. Adaptive runs preserve one entry per orchestrator/worker call. Missing telemetry is `null`, never estimated. |
9
13
  | **Speed** | Wall-clock, measured by the harness from outside: total run + per-story (from `--json` NDJSON event timestamps). The loop itself stores no durations. |
10
14
  | **Quality** | Objective, not judged by any model: the fixture ships **pre-written tests** the agent never has to write (only satisfy). After the run, each story's test file is executed against the final tree. `srcLoc` (non-empty lines in `src/`) is a code-economy proxy. |
11
15
 
12
- ## The fixture (`fixtures/string-kit`)
16
+ ## The fixture (`fixtures/string-kit`)
13
17
 
14
18
  A dependency-free ESM library with 3 stories (`slugify`, `truncate`, `titleCase`) and 16
15
19
  `node:test` assertions total. `bench-verify.mjs` is cumulative: story N runs the tests of
16
20
  stories 1…N (the loop exports `YOKE_STORY`), so later stories cannot break earlier work; the
17
21
  final quality check runs everything. No npm installs, so results measure the agent — not the
18
- network.
22
+ network.
23
+
24
+ ## The routing fixture (`fixtures/routing-queue`)
25
+
26
+ A dependency-free in-memory priority queue with two cumulative stories and 10 pre-written
27
+ `node:test` cases covering idempotency, priority/FIFO ordering, leases, retry/backoff, dead
28
+ letters, expired lease recovery, filters, and state statistics. It is intentionally substantial
29
+ enough that a lower-cost worker can potentially repay one routing-controller call per story.
19
30
 
20
31
  ## Running it
21
32
 
22
33
  ```bash
23
34
  npm run build
24
- node bench/run.mjs --runner=claude # or gemini / codex
25
- node bench/run-matrix.mjs --label=release-1.0
35
+ node bench/run.mjs --runner=claude # or gemini / codex
36
+ node bench/run.mjs --runner=codex --fixture=routing-queue --routing=off --unsafe --run-root=G:\NN-Developed\Yoke-Testground
37
+ node bench/run.mjs --runner=codex --fixture=routing-queue --routing=on --unsafe --run-root=G:\NN-Developed\Yoke-Testground
38
+ node bench/run-large.mjs --seed=G:\NN-Developed\Yoke-Testground\yoke-large-seed-2026-08-02 --routing=off --run-root=G:\NN-Developed\Yoke-Testground
39
+ node bench/run-large.mjs --seed=G:\NN-Developed\Yoke-Testground\yoke-large-seed-2026-08-02 --routing=on --run-root=G:\NN-Developed\Yoke-Testground
40
+ node bench/run-large.mjs --seed=G:\NN-Developed\Yoke-Testground\yoke-codex-study-seed-2026-08-02 --routing=on --label=codex-only-pair1-on --run-root=G:\NN-Developed\Yoke-Testground
41
+ node bench/run-large.mjs --seed=G:\NN-Developed\Yoke-Testground\yoke-codex-study-seed-2026-08-02 --routing=off --label=codex-only-pair1-off --run-root=G:\NN-Developed\Yoke-Testground
42
+ node bench/analyze-routing-study.mjs
43
+ node bench/run-matrix.mjs --label=release-1.0
26
44
  ```
27
45
 
28
- Each run copies the fixture to `bench/.runs/<runner>-<stamp>` (git-ignored), git-inits it,
29
- drives `yoke loop run --json --max=6 --timeout=10`, and writes a result JSON to
46
+ Each run copies the fixture to `bench/.runs/<fixture>-<runner>-<routing>-<stamp>` (or the
47
+ explicit `--run-root`), git-inits it,
48
+ drives `yoke loop run --json --max=6 --timeout=10`, and writes a result JSON to
30
49
  `bench/results/`. Runs are billed against your own accounts for the agent CLIs involved.
31
50
  The matrix is sequential to avoid cross-provider load distortion. Missing CLIs and authentication
32
- failures are stored as honest `unavailable`/`auth-failed` rows and never presented as quality measurements.
51
+ failures are stored as honest `unavailable`/`auth-failed` rows and never presented as quality measurements.
52
+
53
+ `run-large.mjs` accepts an external full-repository seed, starts each arm with a separate empty
54
+ routing registry, junctions this checkout's dependencies, and replays the seed's original
55
+ acceptance tests under fresh filenames after the agent run. This prevents an agent editing a
56
+ visible test from turning into false benchmark evidence. A seed can provide `bench-acceptance.json`
57
+ to declare its fixture identity and hidden-test files. `analyze-routing-study.mjs` validates and
58
+ aggregates the checked-in three-pair Codex-only study.
33
59
 
34
60
  ## Caveats (read before quoting numbers)
35
61
 
36
- - **N=1 per run.** Agent runs are stochastic; treat single runs as indicative, not
37
- statistically robust. Re-run and compare.
38
- - Model identity matters more than CLI identity: `tokens.model` records what actually served
39
- the run. Different default models per CLI make "claude vs gemini" really "model X vs model Y".
62
+ - Agent runs are stochastic. Older rows are N=1; the 2026-08-02 Codex-only study uses three
63
+ alternating-order pairs per arm. Re-run before setting broad policy defaults.
64
+ - Compare routing on/off only when fixture, parent model/effort, permissions, native-subagent
65
+ policy, and host load are controlled. The checked-in routing fixture uses `runner.bare: true`
66
+ to exclude personal MCP/plugin startup from both sides.
67
+ - Model identity matters more than CLI identity: `tokens.model` records what actually served
68
+ the run. Different default models per CLI make "claude vs gemini" really "model X vs model Y".
69
+ - Codex input telemetry includes cache reads. Compare both total input and fresh input; do not
70
+ treat cached input as equivalent to newly processed input. Dollar cost is reported only when
71
+ the provider emits it—Yoke does not guess prices from a model name.
40
72
  - The fixture is deliberately small (a loop-overhead + basic-competence probe, minutes not
41
73
  hours). It does not measure large-context refactoring, UI work, or long-horizon planning.
42
74
  - Cumulative verify means a story's duration includes fixing any regressions it caused.
package/bench/RESULTS.md CHANGED
@@ -17,9 +17,104 @@ Fixture `string-kit` (3 stories, 16 pre-written assertions). Methodology and cav
17
17
  | 2026-07-10 | gemini | — | ⛔ blocked: CLI not authenticated on the bench machine (headless needs `GEMINI_API_KEY` or configured OAuth) | — | — | — | — | — |
18
18
  | — | codex | — | ⛔ not installed on the bench machine | — | — | — | — | — |
19
19
 
20
- Per-story wall-clock (claude run): STORY-1 72 s · STORY-2 114 s · STORY-3 81 s. Every story
21
- passed verify on the first iteration; the final quality check (all 16 assertions on the final
22
- tree, outside the loop) is green.
20
+ Per-story wall-clock (claude run): STORY-1 72 s · STORY-2 114 s · STORY-3 81 s. Every story
21
+ passed verify on the first iteration; the final quality check (all 16 assertions on the final
22
+ tree, outside the loop) is green.
23
+
24
+ ## Codex-only full-repository routing study (2026-08-02)
25
+
26
+ This replaces the earlier single-run impression with three paired comparisons. Fixture
27
+ `yoke-codex-study@1` is a full Yoke snapshot with 571 existing tests and two coupled telemetry
28
+ implementation stories. Every arm used the same `gpt-5.6-sol` high-effort parent, Bare mode,
29
+ disabled native Codex multi-agent, an isolated empty registry, and the same machine. Routing was
30
+ the only policy change: the routed arm used a low-effort `gpt-5.6-sol` controller and selected
31
+ `gpt-5.6-luna` for all six story executions. Pair order alternated on/off, off/on, on/off.
32
+
33
+ All **6/6 runs**, **12/12 stories**, and **36/36 independently replayed hidden acceptance
34
+ checks** passed. Each story completed in one iteration.
35
+
36
+ | Metric (median of 3) | Routing off: Sol high | Routing on: Sol controller + Luna worker | Delta |
37
+ |---|---:|---:|---:|
38
+ | Wall-clock | 579.139 s (561.860–589.501) | 383.099 s (350.819–523.110) | **−196.040 s (−33.8%)** |
39
+ | Input tokens, total | 2,279,164 | 1,561,371 | **−31.5%** |
40
+ | Cached input tokens | 2,147,584 | 1,436,672 | **−33.1%** |
41
+ | Fresh input (`input − cached`) | 140,130 (131,580–205,142) | 124,699 (110,541–127,507) | **−15,431 (−11.0%)** |
42
+ | Output tokens | 16,008 | 8,080 | **−49.5%** |
43
+ | Reasoning tokens | 6,536 | 1,426 | **−78.2%** |
44
+
45
+ | Pair | Execution order | Wall-clock delta | Fresh-input delta |
46
+ |---|---|---:|---:|
47
+ | 1 | on → off | −31.8% | −39.2% |
48
+ | 2 | off → on | −39.4% | −21.1% |
49
+ | 3 | on → off | −11.3% | −3.1% |
50
+
51
+ The routing controller is included in those numbers. Across the three routed runs its six calls
52
+ used 82.251 s, 114,767 total input tokens (66,048 cached; 48,719 fresh), and 187 output tokens.
53
+ All three pairs improved both wall-clock and fresh input, although the third pair shows meaningful
54
+ latency variance. This is convincing evidence for **bounded, delegable repository work**, not a
55
+ claim that every task should be routed.
56
+
57
+ The Codex CLI did not report dollar cost for these plan-backed runs, so no USD value is invented.
58
+ The evidence supports lower token/compute use; an exact currency saving still depends on the
59
+ account's current Sol/Luna billing. Reproduce the aggregate with
60
+ `node bench/analyze-routing-study.mjs`. Raw valid rows:
61
+ [`pair 1 on`](results/yoke-codex-study-codex-routing-on-2026-08-02T07-51-59.json),
62
+ [`pair 1 off`](results/yoke-codex-study-codex-routing-off-2026-08-02T07-58-56.json),
63
+ [`pair 2 off`](results/yoke-codex-study-codex-routing-off-2026-08-02T08-08-48.json),
64
+ [`pair 2 on`](results/yoke-codex-study-codex-routing-on-2026-08-02T08-18-59.json),
65
+ [`pair 3 on`](results/yoke-codex-study-codex-routing-on-2026-08-02T08-25-25.json), and
66
+ [`pair 3 off`](results/yoke-codex-study-codex-routing-off-2026-08-02T08-34-45.json).
67
+
68
+ ## Controlled adaptive-routing A/B (2026-08-01)
69
+
70
+ Fixture `routing-queue@1`: two cumulative implementation stories, 10 pre-written assertions.
71
+ Both arms used Codex as the `gpt-5.6-sol` high-effort parent, `runner.bare: true`, native Codex
72
+ multi-agent disabled, the same isolated Testground host, unsafe permission mode, and one
73
+ iteration per story. The routed arm chose `codex-light` (provider current/default model,
74
+ low reasoning effort) for both stories; the high-effort parent made the routing decisions.
75
+
76
+ | Routing | Quality | Wall-clock | Input tok | Output tok | Total tok | Story 1 | Story 2 |
77
+ |---|---:|---:|---:|---:|---:|---:|---:|
78
+ | off | 2/2 stories; 10/10 tests | 299.787 s | 579,301 | 9,103 | 588,404 | 173.400 s | 125.693 s |
79
+ | on | 2/2 stories; 10/10 tests | 283.134 s | 542,788 | 5,486 | 548,274 | 148.000 s | 122.288 s |
80
+ | delta | same measured quality | **−16.653 s (−5.6%)** | **−36,513 (−6.3%)** | **−3,617 (−39.7%)** | **−40,130 (−6.8%)** | −25.400 s | −3.405 s |
81
+
82
+ The two read-only controller calls account for 41.332 s and 38,169 tokens of the routed total.
83
+ The worker savings exceeded that overhead in this sample. This is **N=1**, not a universal
84
+ speed/cost claim; stochastic model output, provider pricing, cache accounting, and host/API load
85
+ can move the result. Raw valid rows:
86
+ [`routing off`](results/routing-queue-codex-routing-off-2026-08-01T21-44-00.json) and
87
+ [`routing on`](results/routing-queue-codex-routing-on-2026-08-01T21-50-31.json).
88
+
89
+ ## Full-repository adaptive-routing A/B (2026-08-02)
90
+
91
+ Fixture `yoke-large@1` is a snapshot of Yoke itself: 171 relevant files and 17,798 lines across
92
+ source, tests, and canon. Two coupled stories added a privacy-safe registry status API/CLI and
93
+ then circuit-breaker evidence plus automatic strong-parent fallback. Each story ran TypeScript
94
+ build and the 571-test existing suite. Final trees passed 578 tests (routing off) and 579 tests
95
+ (routing on), plus a separate replay of the **four original acceptance tests copied from the
96
+ immutable seed** after the agents had finished.
97
+
98
+ Both arms used the same `gpt-5.6-sol` high-effort parent, Bare mode, disabled native Codex
99
+ multi-agent, isolated empty registries, unsafe permissions, dependency junction, and sequential
100
+ host load.
101
+
102
+ | Routing | Quality | Wall-clock | Input tok | Output tok | Total tok | Story 1 | Story 2 |
103
+ |---|---:|---:|---:|---:|---:|---:|---:|
104
+ | off | 2/2; build + 578 + 4 hidden pass | 660.843 s | 1,971,273 | 16,850 | 1,988,123 | 333.022 s | 327.176 s |
105
+ | on | 2/2; build + 579 + 4 hidden pass | 640.151 s | 2,035,962 | 15,483 | 2,051,445 | 322.735 s | 303.662 s |
106
+ | delta | same acceptance quality | **−20.692 s (−3.1%)** | **+64,689 (+3.3%)** | −1,367 (−8.1%) | **+63,322 (+3.2%)** | −10.287 s | −23.514 s |
107
+
108
+ The controller chose `SELF` for **both** stories. Its two calls cost 35.381 s and 38,457 tokens;
109
+ no cheaper worker executed. The small wall-clock improvement therefore cannot be attributed to
110
+ worker routing and is within the range where stochastic parent execution/API load is a plausible
111
+ explanation. For this complex architecture workload, balanced routing preserved quality through
112
+ a conservative decision but **did not meet the token-cost goal**. A useful next optimization is a
113
+ zero-token deterministic SELF fast path for clearly high-risk architecture/privacy stories.
114
+
115
+ This is again N=1. Raw valid rows:
116
+ [`routing off`](results/yoke-large-codex-routing-off-2026-08-01T22-58-01.json) and
117
+ [`routing on`](results/yoke-large-codex-routing-on-2026-08-01T23-09-39.json).
23
118
 
24
119
  The 2026-07-27 release matrix produced no quality measurement: Claude and Gemini exited
25
120
  before changing the fixture, and the Codex executable could not be probed in this Windows
@@ -36,11 +131,13 @@ Yoke bugs, both fixed in 0.3.0:
36
131
  `-p --yolo` died with "Not enough arguments following: p". The runner now relies on piped
37
132
  stdin (which selects headless mode) and passes only `--yolo`.
38
133
 
39
- ## Reading the numbers
40
-
41
- - Tokens come from the loop's own hook (claude runner only — gemini/codex reporting is an
42
- open gap, see the multi-agent design doc).
43
- - N=1: indicative, not statistics. Re-run with `node bench/run.mjs --runner=<agent>` and
44
- compare rows.
45
- - ~53 k tokens / ~4.5 minutes for a 3-story micro-backlog is the current price of the full
46
- gate pipeline (fresh headless session per story + verify + atomic commit) on this fixture.
134
+ ## Reading the numbers
135
+
136
+ - Tokens come from provider telemetry when available; missing telemetry is recorded as missing,
137
+ never estimated. New Codex rows separate cache reads from fresh input; adaptive rows retain
138
+ controller/worker/parent call splits.
139
+ - The two older A/B rows are N=1 diagnostics. The Codex-only study above is N=3 per arm with
140
+ alternating pair order; it is stronger evidence but still workload-specific.
141
+ - Fixture size changes the outcome: the small queue delegated to a lower-effort worker and saved
142
+ tokens; the architecture/privacy task stayed on SELF and paid controller overhead; the newer
143
+ bounded full-repository telemetry task delegated consistently and saved fresh-input tokens.
@@ -0,0 +1,96 @@
1
+ #!/usr/bin/env node
2
+ import { readFileSync, readdirSync } from 'node:fs'
3
+ import { dirname, join } from 'node:path'
4
+ import { fileURLToPath } from 'node:url'
5
+
6
+ const resultsDir = join(dirname(fileURLToPath(import.meta.url)), 'results')
7
+ const rows = readdirSync(resultsDir)
8
+ .filter(file => file.startsWith('yoke-codex-study-codex-routing-') && file.endsWith('.json'))
9
+ .map(file => ({ file, ...JSON.parse(readFileSync(join(resultsDir, file), 'utf8')) }))
10
+ .filter(row => /^codex-only-pair[1-3]-(on|off)$/.test(row.sampleLabel))
11
+
12
+ const median = values => {
13
+ const sorted = [...values].sort((a, b) => a - b)
14
+ return sorted[Math.floor(sorted.length / 2)]
15
+ }
16
+ const pct = (from, to) => ((to - from) / from) * 100
17
+ const metric = (group, read) => {
18
+ const values = group.map(read)
19
+ return { values, min: Math.min(...values), median: median(values), max: Math.max(...values) }
20
+ }
21
+
22
+ if (rows.length !== 6) throw new Error(`expected six codex-only study rows, found ${rows.length}`)
23
+ for (const row of rows) {
24
+ if (row.verdict !== 'completed' || !row.finalTestsPass || row.iterations !== 2) throw new Error(`invalid study row: ${row.file}`)
25
+ if (row.finalVerification?.originalAcceptanceTestsExitCode !== 0 || row.finalVerification?.originalAcceptanceTestCount !== 6) {
26
+ throw new Error(`hidden acceptance replay failed: ${row.file}`)
27
+ }
28
+ }
29
+
30
+ const groups = Object.fromEntries(['off', 'on'].map(routing => {
31
+ const group = rows.filter(row => row.routing === routing).sort((a, b) => a.sampleLabel.localeCompare(b.sampleLabel))
32
+ return [routing, {
33
+ runs: group.map(row => row.file),
34
+ wallClockMs: metric(group, row => row.wallClockMs),
35
+ inputTokens: metric(group, row => row.tokenBreakdown.inputTokens),
36
+ cachedInputTokens: metric(group, row => row.tokenBreakdown.cachedInputTokens),
37
+ freshInputTokens: metric(group, row => row.tokenBreakdown.freshInputTokens),
38
+ outputTokens: metric(group, row => row.tokenBreakdown.outputTokens),
39
+ reasoningOutputTokens: metric(group, row => row.tokenBreakdown.reasoningOutputTokens),
40
+ }]
41
+ }))
42
+
43
+ const routed = rows.filter(row => row.routing === 'on')
44
+ const controllerCalls = routed.flatMap(row => row.modelCalls.filter(call => call.role === 'orchestrator'))
45
+ const workerCalls = routed.flatMap(row => row.modelCalls.filter(call => call.role === 'worker'))
46
+ if (workerCalls.length !== 6 || workerCalls.some(call => call.profile !== 'codex-luna' || call.requestedModel !== 'gpt-5.6-luna')) {
47
+ throw new Error('not every routed story executed on codex-luna')
48
+ }
49
+
50
+ const deltas = {}
51
+ for (const key of ['wallClockMs', 'inputTokens', 'cachedInputTokens', 'freshInputTokens', 'outputTokens', 'reasoningOutputTokens']) {
52
+ deltas[key] = {
53
+ absolute: groups.on[key].median - groups.off[key].median,
54
+ percent: pct(groups.off[key].median, groups.on[key].median),
55
+ }
56
+ }
57
+
58
+ const pairs = [1, 2, 3].map(pair => {
59
+ const off = rows.find(row => row.sampleLabel === `codex-only-pair${pair}-off`)
60
+ const on = rows.find(row => row.sampleLabel === `codex-only-pair${pair}-on`)
61
+ return {
62
+ pair,
63
+ order: pair === 2 ? ['off', 'on'] : ['on', 'off'],
64
+ wallClockMs: { off: off.wallClockMs, on: on.wallClockMs, percent: pct(off.wallClockMs, on.wallClockMs) },
65
+ freshInputTokens: { off: off.tokenBreakdown.freshInputTokens, on: on.tokenBreakdown.freshInputTokens, percent: pct(off.tokenBreakdown.freshInputTokens, on.tokenBreakdown.freshInputTokens) },
66
+ }
67
+ })
68
+
69
+ console.log(JSON.stringify({
70
+ schemaVersion: 1,
71
+ fixtureVersion: 'yoke-codex-study@1',
72
+ sampleCountPerArm: 3,
73
+ quality: {
74
+ completedRuns: 6,
75
+ completedStories: 12,
76
+ hiddenAcceptancePasses: 36,
77
+ hiddenAcceptanceChecks: 36,
78
+ iterationsPerStory: 1,
79
+ },
80
+ models: {
81
+ parent: { model: 'gpt-5.6-sol', reasoningEffort: 'high' },
82
+ orchestrator: { model: 'gpt-5.6-sol', reasoningEffort: 'low' },
83
+ worker: { model: 'gpt-5.6-luna', reasoningEffort: 'low' },
84
+ },
85
+ groups,
86
+ medianDeltas: deltas,
87
+ pairs,
88
+ controller: {
89
+ calls: controllerCalls.length,
90
+ totalDurationMs: controllerCalls.reduce((sum, call) => sum + call.durationMs, 0),
91
+ totalInputTokens: controllerCalls.reduce((sum, call) => sum + call.inputTokens, 0),
92
+ totalCachedInputTokens: controllerCalls.reduce((sum, call) => sum + (call.cachedInputTokens ?? 0), 0),
93
+ totalFreshInputTokens: controllerCalls.reduce((sum, call) => sum + call.inputTokens - (call.cachedInputTokens ?? 0), 0),
94
+ totalOutputTokens: controllerCalls.reduce((sum, call) => sum + call.outputTokens, 0),
95
+ },
96
+ }, null, 2))
@@ -0,0 +1,30 @@
1
+ canonVersion: 1.2.0
2
+ agents: [codex, claude, gemini]
3
+ loop:
4
+ enabled: true
5
+ runner:
6
+ agent: codex
7
+ model: gpt-5.6-sol
8
+ reasoningEffort: high
9
+ bare: true
10
+ routing:
11
+ enabled: true
12
+ strategy: balanced
13
+ maxCandidates: 3
14
+ workers:
15
+ - id: codex-light
16
+ agent: codex
17
+ reasoningEffort: low
18
+ costTier: medium
19
+ capabilities: [exploration, implementation, tests]
20
+ - id: claude-fast
21
+ agent: claude
22
+ model: haiku
23
+ costTier: low
24
+ capabilities: [exploration, mechanical-edits, tests]
25
+ - id: gemini-auto
26
+ agent: gemini
27
+ costTier: low
28
+ capabilities: [large-context, implementation, tests]
29
+ verify:
30
+ command: node bench-verify.mjs
@@ -0,0 +1,24 @@
1
+ - id: STORY-1
2
+ title: "Priority task queue with idempotency and leases"
3
+ priority: 1
4
+ acceptance:
5
+ - "Export a TaskQueue class from src/task-queue.mjs; constructor accepts optional clock, maxAttempts, and baseDelayMs"
6
+ - "enqueue({type,payload,priority,idempotencyKey}) validates a non-empty type, defaults priority to 0, and returns an immutable task snapshot"
7
+ - "Repeated non-empty idempotencyKey returns the original task without adding another queue entry"
8
+ - "claim(workerId,{leaseMs}) returns the available queued task with highest priority and FIFO order for ties; it records workerId, leaseUntil, and increments attempts"
9
+ - "complete(taskId,workerId,result) only allows the current lease owner and returns an immutable completed snapshot"
10
+ - "get(taskId) and list() return snapshots that callers cannot use to mutate queue state"
11
+ - "node bench-verify.mjs exits 0"
12
+ passes: false
13
+ - id: STORY-2
14
+ title: "Retry backoff, dead letters, expired leases, and statistics"
15
+ priority: 2
16
+ acceptance:
17
+ - "fail(taskId,workerId,error) only allows the lease owner; before maxAttempts it requeues with exponential delay baseDelayMs * 2^(attempts-1)"
18
+ - "A task is not claimable before availableAt; at maxAttempts failure moves it to dead status"
19
+ - "reapExpired() requeues active tasks whose lease has expired and returns the number requeued"
20
+ - "list({status,type}) filters without exposing mutable internal objects"
21
+ - "stats() returns total and counts for queued, active, completed, and dead tasks"
22
+ - "Invalid task ids, invalid worker ids, and invalid state transitions throw clear errors"
23
+ - "node bench-verify.mjs exits 0"
24
+ passes: false
@@ -0,0 +1,9 @@
1
+ import { spawnSync } from 'node:child_process'
2
+
3
+ const total = 2
4
+ const story = process.env.YOKE_STORY
5
+ const requested = story ? Number(story.split('-')[1]) : total
6
+ const upTo = Number.isFinite(requested) && requested >= 1 && requested <= total ? requested : total
7
+ const files = Array.from({ length: upTo }, (_, index) => `tests/STORY-${index + 1}.test.mjs`)
8
+ const result = spawnSync(process.execPath, ['--test', ...files], { stdio: 'inherit' })
9
+ process.exit(result.status ?? 1)