@hecer/yoke 1.1.1 → 1.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/CHANGELOG.md +24 -3
- package/README.md +87 -21
- package/bench/README.md +46 -14
- package/bench/RESULTS.md +108 -11
- package/bench/analyze-routing-study.mjs +96 -0
- package/bench/fixtures/routing-queue/.yoke/config.yaml +30 -0
- package/bench/fixtures/routing-queue/.yoke/prd.yaml +24 -0
- package/bench/fixtures/routing-queue/bench-verify.mjs +9 -0
- package/bench/fixtures/routing-queue/package.json +9 -0
- package/bench/fixtures/routing-queue/src/task-queue.mjs +1 -0
- package/bench/fixtures/routing-queue/tests/STORY-1.test.mjs +68 -0
- package/bench/fixtures/routing-queue/tests/STORY-2.test.mjs +72 -0
- package/bench/results/routing-queue-codex-routing-off-2026-08-01T21-44-00.json +46 -0
- package/bench/results/routing-queue-codex-routing-on-2026-08-01T21-50-31.json +109 -0
- package/bench/results/yoke-codex-study-codex-routing-off-2026-08-02T07-58-56.json +65 -0
- package/bench/results/yoke-codex-study-codex-routing-off-2026-08-02T08-08-48.json +65 -0
- package/bench/results/yoke-codex-study-codex-routing-off-2026-08-02T08-34-45.json +65 -0
- package/bench/results/yoke-codex-study-codex-routing-on-2026-08-02T07-51-59.json +168 -0
- package/bench/results/yoke-codex-study-codex-routing-on-2026-08-02T08-18-59.json +168 -0
- package/bench/results/yoke-codex-study-codex-routing-on-2026-08-02T08-25-25.json +168 -0
- package/bench/results/yoke-large-codex-routing-off-2026-08-01T22-58-01.json +54 -0
- package/bench/results/yoke-large-codex-routing-on-2026-08-01T23-09-39.json +133 -0
- package/bench/run-large.mjs +188 -0
- package/bench/run.mjs +42 -21
- package/dist/agents/providers.js +19 -2
- package/dist/agents/telemetry.js +46 -7
- package/dist/cli.js +7 -5
- package/dist/loop/decision.js +3 -1
- package/dist/loop/loop.js +10 -0
- package/dist/loop/reporter.js +9 -0
- package/dist/loop/run-command.js +33 -5
- package/dist/loop/runner.js +68 -9
- package/dist/retrofit/config.js +21 -0
- package/dist/review/command.js +2 -2
- package/dist/routing/registry.js +92 -0
- package/dist/routing/router.js +195 -0
- package/dist/setup/command.js +23 -1
- package/gemini-extension.json +1 -1
- package/package.json +6 -4
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
|
|
3
3
|
"name": "yoke",
|
|
4
4
|
"displayName": "Yoke",
|
|
5
|
-
"version": "1.1
|
|
5
|
+
"version": "1.2.1",
|
|
6
6
|
"description": "Cross-agent coding harness: one curated skill canon (TDD, brainstorming, plans, reviews, shipping, design verification) plus mechanical safety gates and an autonomous loop via the yoke CLI.",
|
|
7
7
|
"author": { "name": "HECer", "url": "https://github.com/HECer" },
|
|
8
8
|
"homepage": "https://github.com/HECer/yoke#readme",
|
package/CHANGELOG.md
CHANGED
|
@@ -1,6 +1,27 @@
|
|
|
1
|
-
# Changelog
|
|
2
|
-
|
|
3
|
-
## 1.1
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 1.2.1 — 2026-08-02
|
|
4
|
+
|
|
5
|
+
### Fixed
|
|
6
|
+
- `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
|
|
7
|
+
- Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
|
|
8
|
+
|
|
9
|
+
## 1.2.0 — 2026-08-02
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
- Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
|
|
13
|
+
- A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
|
|
14
|
+
- Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
|
|
15
|
+
|
|
16
|
+
### Changed
|
|
17
|
+
- Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
|
|
18
|
+
- Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
|
|
19
|
+
|
|
20
|
+
### Fixed
|
|
21
|
+
- Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
|
|
22
|
+
- Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
|
|
23
|
+
|
|
24
|
+
## 1.1.0 — 2026-07-30
|
|
4
25
|
|
|
5
26
|
### Added
|
|
6
27
|
- Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
|
package/README.md
CHANGED
|
@@ -2,8 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
# 🐂 Yoke
|
|
4
4
|
|
|
5
|
-
<!-- yoke:version:start -->1.
|
|
6
|
-
<!-- yoke:tests:start -->
|
|
5
|
+
<!-- yoke:version:start -->1.2.1<!-- yoke:version:end -->
|
|
6
|
+
<!-- yoke:tests:start -->582<!-- yoke:tests:end -->
|
|
7
7
|
<!-- yoke:skills:start -->29<!-- yoke:skills:end -->
|
|
8
8
|
<!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
|
|
9
9
|
|
|
@@ -25,7 +25,7 @@
|
|
|
25
25
|
|
|
26
26
|
</div>
|
|
27
27
|
|
|
28
|
-
> **TL;DR** — `yoke setup .` asks
|
|
28
|
+
> **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
|
|
29
29
|
|
|
30
30
|
Yoke 1.1 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
|
|
31
31
|
is explicit; reviews require a schema-valid verdict and a different model unless
|
|
@@ -85,7 +85,7 @@ yoke new my-app --idea="a CLI that tracks reading lists"
|
|
|
85
85
|
yoke loop on my-app && yoke loop run my-app --isolate
|
|
86
86
|
|
|
87
87
|
# — or retrofit an existing project —
|
|
88
|
-
yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions
|
|
88
|
+
yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
|
|
89
89
|
yoke validate canon # sanity-check the canon
|
|
90
90
|
yoke loop run /path/to/project --isolate --reviewer=codex --max=20
|
|
91
91
|
```
|
|
@@ -104,7 +104,7 @@ The canon is also packaged as a Claude Code plugin — the repo is its own marke
|
|
|
104
104
|
That gives you all canon skills under the `yoke:` namespace (e.g. `yoke:tdd`, `yoke:review`) inside Claude Code — no retrofit needed. The `yoke` CLI (loop, gates, retrofit for Codex/Gemini) still comes from `npm i -g @hecer/yoke`. Gemini CLI users can likewise `gemini extensions install https://github.com/HECer/yoke`.
|
|
105
105
|
|
|
106
106
|
For Codex, no preinstalled skill is required: run `npx @hecer/yoke setup .` in a terminal, or
|
|
107
|
-
ask Codex to run the
|
|
107
|
+
ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
|
|
108
108
|
`.agents/skills/`, including `yoke-retrofit` and `yoke-workflow`; start a fresh Codex task if an
|
|
109
109
|
already-open task does not discover newly installed skills. The npm package also contains
|
|
110
110
|
`.codex-plugin/plugin.json` for Codex plugin hosts.
|
|
@@ -138,7 +138,7 @@ Yoke is meant to be operated *by* your coding agent — after a retrofit, the ag
|
|
|
138
138
|
|
|
139
139
|
> **Watch / unblock** — *"Run `yoke loop status .`. If it says BLOCKED, run the project's verify command, find the root cause, fix it without weakening tests, then continue the loop."*
|
|
140
140
|
|
|
141
|
-
> ⚠️ **Long runs from inside an agent session:** a multi-story
|
|
141
|
+
> ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
|
|
142
142
|
|
|
143
143
|
> ⚠️ **Never kill agent processes by name or command-line pattern** (e.g. every process matching `dangerously-skip-permissions`): on a machine running several yoke projects, that takes down the *healthy* runners of the other projects mid-story — they stall and their loops block. `yoke loop cleanup` is the scoped alternative: each watchdog records its pids in the project's `.yoke/runner.pid`, and cleanup kills exactly those recorded trees — nothing else on the machine.
|
|
144
144
|
|
|
@@ -148,14 +148,14 @@ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`)
|
|
|
148
148
|
|
|
149
149
|
| Command | What it does | Exit codes |
|
|
150
150
|
|---|---|---|
|
|
151
|
-
| `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop]` | Shared
|
|
151
|
+
| `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
|
|
152
152
|
| `yoke validate [canonDir]` | Validate the canon (schema, frontmatter, templates) | `0` valid · `1` errors |
|
|
153
153
|
| `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
|
|
154
154
|
| `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
|
|
155
155
|
| `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
|
|
156
156
|
| `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
|
|
157
157
|
| `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
|
|
158
|
-
| `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `decision` shows a critical stop, `answer` records it and resumes, `resume` retries a failed restart with the preserved safety options
|
|
158
|
+
| `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` is unlimited by default and `--max=N` creates an intentional batch cap; `decision` shows a critical stop, `answer` records it and resumes, `resume` retries a failed restart with the preserved safety options | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
|
|
159
159
|
| `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
|
|
160
160
|
| `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
|
|
161
161
|
| `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
|
|
@@ -330,8 +330,8 @@ yoke loop run . \
|
|
|
330
330
|
--runner=codex \ # implement with Codex…
|
|
331
331
|
--reviewer=claude \ # …review with Claude (role separation)
|
|
332
332
|
--isolate \ # each story in a throwaway git worktree
|
|
333
|
-
--decision-policy=critical
|
|
334
|
-
|
|
333
|
+
--decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
|
|
334
|
+
# Optional: add --max=20 only when this run should stop after a bounded batch.
|
|
335
335
|
yoke loop off . # disable
|
|
336
336
|
```
|
|
337
337
|
|
|
@@ -370,11 +370,9 @@ Every iteration emits token-free, harness-side feedback (Node console + local fi
|
|
|
370
370
|
NDJSON line on stdout (`{"type":"status","state":"running","phase":"verifying",…}` — the
|
|
371
371
|
same shape as `loop-status.json`), the human narrative moves off stdout (the final summary
|
|
372
372
|
goes to stderr), and a consumer can follow the stream line by line instead of polling the file.
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
last model id seen on the stream (e.g. `claude-opus-4-6-20260501`), omitted if the CLI
|
|
377
|
-
never reported one.
|
|
373
|
+
Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
|
|
374
|
+
output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
|
|
375
|
+
and cost fields when available. Missing values stay absent—Yoke does not estimate them.
|
|
378
376
|
|
|
379
377
|
### Pausing a run
|
|
380
378
|
|
|
@@ -434,9 +432,65 @@ cannot begin because a provider/reviewer is unavailable or another process owns
|
|
|
434
432
|
until a loop actually runs. To intentionally abandon an orphaned or stale private resume state,
|
|
435
433
|
use `yoke loop resume . --discard`; pending decisions are never deleted by that command. Existing
|
|
436
434
|
`loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` remain supported as compatibility aliases;
|
|
437
|
-
new projects should use `decisionPolicy: auto|critical`.
|
|
438
|
-
|
|
439
|
-
###
|
|
435
|
+
new projects should use `decisionPolicy: auto|critical`.
|
|
436
|
+
|
|
437
|
+
### Adaptive model routing (explicit opt-in)
|
|
438
|
+
|
|
439
|
+
`yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
|
|
440
|
+
parent remains the strong planner/controller. Before each bounded story it receives only the
|
|
441
|
+
story, acceptance criteria, and at most three eligible worker profiles, then returns one
|
|
442
|
+
machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
|
|
443
|
+
`SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
|
|
444
|
+
runs so Yoke does not pay for two orchestration layers.
|
|
445
|
+
|
|
446
|
+
**Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
|
|
447
|
+
Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
|
|
448
|
+
cover invocation and routing behavior for all three providers. The measured performance evidence
|
|
449
|
+
below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
|
|
450
|
+
until authenticated, repeated in-the-wild runs exist for those providers.
|
|
451
|
+
|
|
452
|
+
```yaml
|
|
453
|
+
runner:
|
|
454
|
+
agent: codex
|
|
455
|
+
model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
|
|
456
|
+
reasoningEffort: high
|
|
457
|
+
routing:
|
|
458
|
+
enabled: true # setup defaults false; setup --routing opts in
|
|
459
|
+
strategy: balanced # balanced | cost | speed | quality
|
|
460
|
+
maxCandidates: 3
|
|
461
|
+
workers:
|
|
462
|
+
- id: codex-light
|
|
463
|
+
agent: codex
|
|
464
|
+
reasoningEffort: low
|
|
465
|
+
costTier: medium
|
|
466
|
+
capabilities: [exploration, implementation, tests]
|
|
467
|
+
- id: claude-fast
|
|
468
|
+
agent: claude
|
|
469
|
+
model: haiku # rolling alias; omit to use the provider's current default
|
|
470
|
+
reasoningEffort: low
|
|
471
|
+
costTier: low
|
|
472
|
+
capabilities: [mechanical-edits, tests]
|
|
473
|
+
- id: gemini-auto
|
|
474
|
+
agent: gemini # omitted model means the account's current Auto/default route
|
|
475
|
+
costTier: low
|
|
476
|
+
capabilities: [large-context, implementation]
|
|
477
|
+
```
|
|
478
|
+
|
|
479
|
+
Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
|
|
480
|
+
baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
|
|
481
|
+
eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
|
|
482
|
+
"intelligence score". Candidate model IDs come from project configuration while setup defaults
|
|
483
|
+
prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
|
|
484
|
+
independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
|
|
485
|
+
after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
|
|
486
|
+
time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
|
|
487
|
+
cannot overwrite a shared registry file.
|
|
488
|
+
|
|
489
|
+
Routing is not free: it adds one controller call per story. It is most promising when a bounded
|
|
490
|
+
worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
|
|
491
|
+
on your own backlog rather than assuming a win.
|
|
492
|
+
|
|
493
|
+
### Performance budgets: efficiency as a gate, not a style
|
|
440
494
|
|
|
441
495
|
Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
|
|
442
496
|
be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
|
|
@@ -598,12 +652,24 @@ Yoke's guardrails are **mechanical, not advisory** — the loop blocks on a dirt
|
|
|
598
652
|
| Best for | rapid exploration / migration / onboarding | systematic refactoring in typed codebases |
|
|
599
653
|
| Caveat | heuristic edges; static index can go stale | one language server per language |
|
|
600
654
|
|
|
601
|
-
## 🪙 Token efficiency
|
|
655
|
+
## 🪙 Token efficiency
|
|
602
656
|
|
|
603
657
|
Yoke attacks tokens on two complementary surfaces:
|
|
604
658
|
|
|
605
659
|
- **rtk** compresses noisy command/tool output before it enters context (wired as a hook/instruction per agent).
|
|
606
|
-
- The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
|
|
660
|
+
- The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
|
|
661
|
+
|
|
662
|
+
A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
|
|
663
|
+
completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
|
|
664
|
+
controller selected Luna for every bounded implementation story. Including controller overhead,
|
|
665
|
+
the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
|
|
666
|
+
and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
|
|
667
|
+
|
|
668
|
+
The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
|
|
669
|
+
controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
|
|
670
|
+
emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
|
|
671
|
+
inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
|
|
672
|
+
[`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
|
|
607
673
|
|
|
608
674
|
## 🧩 Optional companions
|
|
609
675
|
|
|
@@ -647,7 +713,7 @@ parallel dispatcher, broader benchmark samples, native output schemas, and relea
|
|
|
647
713
|
## 🧪 Development
|
|
648
714
|
|
|
649
715
|
```bash
|
|
650
|
-
npm test # vitest (
|
|
716
|
+
npm test # vitest (582 tests)
|
|
651
717
|
npm run build # tsc, no emit errors
|
|
652
718
|
npm run yoke -- validate canon
|
|
653
719
|
```
|
package/bench/README.md
CHANGED
|
@@ -1,42 +1,74 @@
|
|
|
1
1
|
# Yoke benchmark — tokens · speed · quality
|
|
2
2
|
|
|
3
|
-
Reproducible cross-runner
|
|
4
|
-
PRD
|
|
3
|
+
Reproducible cross-runner and routing A/B benchmarks for the Yoke loop. Fixed fixture projects,
|
|
4
|
+
the same PRD within each comparison, and three measured dimensions:
|
|
5
|
+
|
|
6
|
+
Adaptive routing is implemented for Claude Code, Codex CLI, and Gemini CLI. The checked-in
|
|
7
|
+
multi-run performance study currently measures Codex only; provider support tests are not treated
|
|
8
|
+
as performance evidence for Claude or Gemini.
|
|
5
9
|
|
|
6
10
|
| Dimension | How it is measured |
|
|
7
11
|
|---|---|
|
|
8
|
-
| **Tokens** | The loop's own
|
|
12
|
+
| **Tokens** | The loop's own provider telemetry in `.yoke/loop-status.json`, split into total input, cached input, fresh input (`input − cached`), output, and reasoning tokens when the provider reports them. Adaptive runs preserve one entry per orchestrator/worker call. Missing telemetry is `null`, never estimated. |
|
|
9
13
|
| **Speed** | Wall-clock, measured by the harness from outside: total run + per-story (from `--json` NDJSON event timestamps). The loop itself stores no durations. |
|
|
10
14
|
| **Quality** | Objective, not judged by any model: the fixture ships **pre-written tests** the agent never has to write (only satisfy). After the run, each story's test file is executed against the final tree. `srcLoc` (non-empty lines in `src/`) is a code-economy proxy. |
|
|
11
15
|
|
|
12
|
-
## The fixture (`fixtures/string-kit`)
|
|
16
|
+
## The fixture (`fixtures/string-kit`)
|
|
13
17
|
|
|
14
18
|
A dependency-free ESM library with 3 stories (`slugify`, `truncate`, `titleCase`) and 16
|
|
15
19
|
`node:test` assertions total. `bench-verify.mjs` is cumulative: story N runs the tests of
|
|
16
20
|
stories 1…N (the loop exports `YOKE_STORY`), so later stories cannot break earlier work; the
|
|
17
21
|
final quality check runs everything. No npm installs, so results measure the agent — not the
|
|
18
|
-
network.
|
|
22
|
+
network.
|
|
23
|
+
|
|
24
|
+
## The routing fixture (`fixtures/routing-queue`)
|
|
25
|
+
|
|
26
|
+
A dependency-free in-memory priority queue with two cumulative stories and 10 pre-written
|
|
27
|
+
`node:test` cases covering idempotency, priority/FIFO ordering, leases, retry/backoff, dead
|
|
28
|
+
letters, expired lease recovery, filters, and state statistics. It is intentionally substantial
|
|
29
|
+
enough that a lower-cost worker can potentially repay one routing-controller call per story.
|
|
19
30
|
|
|
20
31
|
## Running it
|
|
21
32
|
|
|
22
33
|
```bash
|
|
23
34
|
npm run build
|
|
24
|
-
node bench/run.mjs --runner=claude # or gemini / codex
|
|
25
|
-
node bench/run
|
|
35
|
+
node bench/run.mjs --runner=claude # or gemini / codex
|
|
36
|
+
node bench/run.mjs --runner=codex --fixture=routing-queue --routing=off --unsafe --run-root=G:\NN-Developed\Yoke-Testground
|
|
37
|
+
node bench/run.mjs --runner=codex --fixture=routing-queue --routing=on --unsafe --run-root=G:\NN-Developed\Yoke-Testground
|
|
38
|
+
node bench/run-large.mjs --seed=G:\NN-Developed\Yoke-Testground\yoke-large-seed-2026-08-02 --routing=off --run-root=G:\NN-Developed\Yoke-Testground
|
|
39
|
+
node bench/run-large.mjs --seed=G:\NN-Developed\Yoke-Testground\yoke-large-seed-2026-08-02 --routing=on --run-root=G:\NN-Developed\Yoke-Testground
|
|
40
|
+
node bench/run-large.mjs --seed=G:\NN-Developed\Yoke-Testground\yoke-codex-study-seed-2026-08-02 --routing=on --label=codex-only-pair1-on --run-root=G:\NN-Developed\Yoke-Testground
|
|
41
|
+
node bench/run-large.mjs --seed=G:\NN-Developed\Yoke-Testground\yoke-codex-study-seed-2026-08-02 --routing=off --label=codex-only-pair1-off --run-root=G:\NN-Developed\Yoke-Testground
|
|
42
|
+
node bench/analyze-routing-study.mjs
|
|
43
|
+
node bench/run-matrix.mjs --label=release-1.0
|
|
26
44
|
```
|
|
27
45
|
|
|
28
|
-
Each run copies the fixture to `bench/.runs/<runner>-<stamp>` (
|
|
29
|
-
|
|
46
|
+
Each run copies the fixture to `bench/.runs/<fixture>-<runner>-<routing>-<stamp>` (or the
|
|
47
|
+
explicit `--run-root`), git-inits it,
|
|
48
|
+
drives `yoke loop run --json --max=6 --timeout=10`, and writes a result JSON to
|
|
30
49
|
`bench/results/`. Runs are billed against your own accounts for the agent CLIs involved.
|
|
31
50
|
The matrix is sequential to avoid cross-provider load distortion. Missing CLIs and authentication
|
|
32
|
-
failures are stored as honest `unavailable`/`auth-failed` rows and never presented as quality measurements.
|
|
51
|
+
failures are stored as honest `unavailable`/`auth-failed` rows and never presented as quality measurements.
|
|
52
|
+
|
|
53
|
+
`run-large.mjs` accepts an external full-repository seed, starts each arm with a separate empty
|
|
54
|
+
routing registry, junctions this checkout's dependencies, and replays the seed's original
|
|
55
|
+
acceptance tests under fresh filenames after the agent run. This prevents an agent editing a
|
|
56
|
+
visible test from turning into false benchmark evidence. A seed can provide `bench-acceptance.json`
|
|
57
|
+
to declare its fixture identity and hidden-test files. `analyze-routing-study.mjs` validates and
|
|
58
|
+
aggregates the checked-in three-pair Codex-only study.
|
|
33
59
|
|
|
34
60
|
## Caveats (read before quoting numbers)
|
|
35
61
|
|
|
36
|
-
-
|
|
37
|
-
|
|
38
|
-
-
|
|
39
|
-
|
|
62
|
+
- Agent runs are stochastic. Older rows are N=1; the 2026-08-02 Codex-only study uses three
|
|
63
|
+
alternating-order pairs per arm. Re-run before setting broad policy defaults.
|
|
64
|
+
- Compare routing on/off only when fixture, parent model/effort, permissions, native-subagent
|
|
65
|
+
policy, and host load are controlled. The checked-in routing fixture uses `runner.bare: true`
|
|
66
|
+
to exclude personal MCP/plugin startup from both sides.
|
|
67
|
+
- Model identity matters more than CLI identity: `tokens.model` records what actually served
|
|
68
|
+
the run. Different default models per CLI make "claude vs gemini" really "model X vs model Y".
|
|
69
|
+
- Codex input telemetry includes cache reads. Compare both total input and fresh input; do not
|
|
70
|
+
treat cached input as equivalent to newly processed input. Dollar cost is reported only when
|
|
71
|
+
the provider emits it—Yoke does not guess prices from a model name.
|
|
40
72
|
- The fixture is deliberately small (a loop-overhead + basic-competence probe, minutes not
|
|
41
73
|
hours). It does not measure large-context refactoring, UI work, or long-horizon planning.
|
|
42
74
|
- Cumulative verify means a story's duration includes fixing any regressions it caused.
|
package/bench/RESULTS.md
CHANGED
|
@@ -17,9 +17,104 @@ Fixture `string-kit` (3 stories, 16 pre-written assertions). Methodology and cav
|
|
|
17
17
|
| 2026-07-10 | gemini | — | ⛔ blocked: CLI not authenticated on the bench machine (headless needs `GEMINI_API_KEY` or configured OAuth) | — | — | — | — | — |
|
|
18
18
|
| — | codex | — | ⛔ not installed on the bench machine | — | — | — | — | — |
|
|
19
19
|
|
|
20
|
-
Per-story wall-clock (claude run): STORY-1 72 s · STORY-2 114 s · STORY-3 81 s. Every story
|
|
21
|
-
passed verify on the first iteration; the final quality check (all 16 assertions on the final
|
|
22
|
-
tree, outside the loop) is green.
|
|
20
|
+
Per-story wall-clock (claude run): STORY-1 72 s · STORY-2 114 s · STORY-3 81 s. Every story
|
|
21
|
+
passed verify on the first iteration; the final quality check (all 16 assertions on the final
|
|
22
|
+
tree, outside the loop) is green.
|
|
23
|
+
|
|
24
|
+
## Codex-only full-repository routing study (2026-08-02)
|
|
25
|
+
|
|
26
|
+
This replaces the earlier single-run impression with three paired comparisons. Fixture
|
|
27
|
+
`yoke-codex-study@1` is a full Yoke snapshot with 571 existing tests and two coupled telemetry
|
|
28
|
+
implementation stories. Every arm used the same `gpt-5.6-sol` high-effort parent, Bare mode,
|
|
29
|
+
disabled native Codex multi-agent, an isolated empty registry, and the same machine. Routing was
|
|
30
|
+
the only policy change: the routed arm used a low-effort `gpt-5.6-sol` controller and selected
|
|
31
|
+
`gpt-5.6-luna` for all six story executions. Pair order alternated on/off, off/on, on/off.
|
|
32
|
+
|
|
33
|
+
All **6/6 runs**, **12/12 stories**, and **36/36 independently replayed hidden acceptance
|
|
34
|
+
checks** passed. Each story completed in one iteration.
|
|
35
|
+
|
|
36
|
+
| Metric (median of 3) | Routing off: Sol high | Routing on: Sol controller + Luna worker | Delta |
|
|
37
|
+
|---|---:|---:|---:|
|
|
38
|
+
| Wall-clock | 579.139 s (561.860–589.501) | 383.099 s (350.819–523.110) | **−196.040 s (−33.8%)** |
|
|
39
|
+
| Input tokens, total | 2,279,164 | 1,561,371 | **−31.5%** |
|
|
40
|
+
| Cached input tokens | 2,147,584 | 1,436,672 | **−33.1%** |
|
|
41
|
+
| Fresh input (`input − cached`) | 140,130 (131,580–205,142) | 124,699 (110,541–127,507) | **−15,431 (−11.0%)** |
|
|
42
|
+
| Output tokens | 16,008 | 8,080 | **−49.5%** |
|
|
43
|
+
| Reasoning tokens | 6,536 | 1,426 | **−78.2%** |
|
|
44
|
+
|
|
45
|
+
| Pair | Execution order | Wall-clock delta | Fresh-input delta |
|
|
46
|
+
|---|---|---:|---:|
|
|
47
|
+
| 1 | on → off | −31.8% | −39.2% |
|
|
48
|
+
| 2 | off → on | −39.4% | −21.1% |
|
|
49
|
+
| 3 | on → off | −11.3% | −3.1% |
|
|
50
|
+
|
|
51
|
+
The routing controller is included in those numbers. Across the three routed runs its six calls
|
|
52
|
+
used 82.251 s, 114,767 total input tokens (66,048 cached; 48,719 fresh), and 187 output tokens.
|
|
53
|
+
All three pairs improved both wall-clock and fresh input, although the third pair shows meaningful
|
|
54
|
+
latency variance. This is convincing evidence for **bounded, delegable repository work**, not a
|
|
55
|
+
claim that every task should be routed.
|
|
56
|
+
|
|
57
|
+
The Codex CLI did not report dollar cost for these plan-backed runs, so no USD value is invented.
|
|
58
|
+
The evidence supports lower token/compute use; an exact currency saving still depends on the
|
|
59
|
+
account's current Sol/Luna billing. Reproduce the aggregate with
|
|
60
|
+
`node bench/analyze-routing-study.mjs`. Raw valid rows:
|
|
61
|
+
[`pair 1 on`](results/yoke-codex-study-codex-routing-on-2026-08-02T07-51-59.json),
|
|
62
|
+
[`pair 1 off`](results/yoke-codex-study-codex-routing-off-2026-08-02T07-58-56.json),
|
|
63
|
+
[`pair 2 off`](results/yoke-codex-study-codex-routing-off-2026-08-02T08-08-48.json),
|
|
64
|
+
[`pair 2 on`](results/yoke-codex-study-codex-routing-on-2026-08-02T08-18-59.json),
|
|
65
|
+
[`pair 3 on`](results/yoke-codex-study-codex-routing-on-2026-08-02T08-25-25.json), and
|
|
66
|
+
[`pair 3 off`](results/yoke-codex-study-codex-routing-off-2026-08-02T08-34-45.json).
|
|
67
|
+
|
|
68
|
+
## Controlled adaptive-routing A/B (2026-08-01)
|
|
69
|
+
|
|
70
|
+
Fixture `routing-queue@1`: two cumulative implementation stories, 10 pre-written assertions.
|
|
71
|
+
Both arms used Codex as the `gpt-5.6-sol` high-effort parent, `runner.bare: true`, native Codex
|
|
72
|
+
multi-agent disabled, the same isolated Testground host, unsafe permission mode, and one
|
|
73
|
+
iteration per story. The routed arm chose `codex-light` (provider current/default model,
|
|
74
|
+
low reasoning effort) for both stories; the high-effort parent made the routing decisions.
|
|
75
|
+
|
|
76
|
+
| Routing | Quality | Wall-clock | Input tok | Output tok | Total tok | Story 1 | Story 2 |
|
|
77
|
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
|
78
|
+
| off | 2/2 stories; 10/10 tests | 299.787 s | 579,301 | 9,103 | 588,404 | 173.400 s | 125.693 s |
|
|
79
|
+
| on | 2/2 stories; 10/10 tests | 283.134 s | 542,788 | 5,486 | 548,274 | 148.000 s | 122.288 s |
|
|
80
|
+
| delta | same measured quality | **−16.653 s (−5.6%)** | **−36,513 (−6.3%)** | **−3,617 (−39.7%)** | **−40,130 (−6.8%)** | −25.400 s | −3.405 s |
|
|
81
|
+
|
|
82
|
+
The two read-only controller calls account for 41.332 s and 38,169 tokens of the routed total.
|
|
83
|
+
The worker savings exceeded that overhead in this sample. This is **N=1**, not a universal
|
|
84
|
+
speed/cost claim; stochastic model output, provider pricing, cache accounting, and host/API load
|
|
85
|
+
can move the result. Raw valid rows:
|
|
86
|
+
[`routing off`](results/routing-queue-codex-routing-off-2026-08-01T21-44-00.json) and
|
|
87
|
+
[`routing on`](results/routing-queue-codex-routing-on-2026-08-01T21-50-31.json).
|
|
88
|
+
|
|
89
|
+
## Full-repository adaptive-routing A/B (2026-08-02)
|
|
90
|
+
|
|
91
|
+
Fixture `yoke-large@1` is a snapshot of Yoke itself: 171 relevant files and 17,798 lines across
|
|
92
|
+
source, tests, and canon. Two coupled stories added a privacy-safe registry status API/CLI and
|
|
93
|
+
then circuit-breaker evidence plus automatic strong-parent fallback. Each story ran TypeScript
|
|
94
|
+
build and the 571-test existing suite. Final trees passed 578 tests (routing off) and 579 tests
|
|
95
|
+
(routing on), plus a separate replay of the **four original acceptance tests copied from the
|
|
96
|
+
immutable seed** after the agents had finished.
|
|
97
|
+
|
|
98
|
+
Both arms used the same `gpt-5.6-sol` high-effort parent, Bare mode, disabled native Codex
|
|
99
|
+
multi-agent, isolated empty registries, unsafe permissions, dependency junction, and sequential
|
|
100
|
+
host load.
|
|
101
|
+
|
|
102
|
+
| Routing | Quality | Wall-clock | Input tok | Output tok | Total tok | Story 1 | Story 2 |
|
|
103
|
+
|---|---:|---:|---:|---:|---:|---:|---:|
|
|
104
|
+
| off | 2/2; build + 578 + 4 hidden pass | 660.843 s | 1,971,273 | 16,850 | 1,988,123 | 333.022 s | 327.176 s |
|
|
105
|
+
| on | 2/2; build + 579 + 4 hidden pass | 640.151 s | 2,035,962 | 15,483 | 2,051,445 | 322.735 s | 303.662 s |
|
|
106
|
+
| delta | same acceptance quality | **−20.692 s (−3.1%)** | **+64,689 (+3.3%)** | −1,367 (−8.1%) | **+63,322 (+3.2%)** | −10.287 s | −23.514 s |
|
|
107
|
+
|
|
108
|
+
The controller chose `SELF` for **both** stories. Its two calls cost 35.381 s and 38,457 tokens;
|
|
109
|
+
no cheaper worker executed. The small wall-clock improvement therefore cannot be attributed to
|
|
110
|
+
worker routing and is within the range where stochastic parent execution/API load is a plausible
|
|
111
|
+
explanation. For this complex architecture workload, balanced routing preserved quality through
|
|
112
|
+
a conservative decision but **did not meet the token-cost goal**. A useful next optimization is a
|
|
113
|
+
zero-token deterministic SELF fast path for clearly high-risk architecture/privacy stories.
|
|
114
|
+
|
|
115
|
+
This is again N=1. Raw valid rows:
|
|
116
|
+
[`routing off`](results/yoke-large-codex-routing-off-2026-08-01T22-58-01.json) and
|
|
117
|
+
[`routing on`](results/yoke-large-codex-routing-on-2026-08-01T23-09-39.json).
|
|
23
118
|
|
|
24
119
|
The 2026-07-27 release matrix produced no quality measurement: Claude and Gemini exited
|
|
25
120
|
before changing the fixture, and the Codex executable could not be probed in this Windows
|
|
@@ -36,11 +131,13 @@ Yoke bugs, both fixed in 0.3.0:
|
|
|
36
131
|
`-p --yolo` died with "Not enough arguments following: p". The runner now relies on piped
|
|
37
132
|
stdin (which selects headless mode) and passes only `--yolo`.
|
|
38
133
|
|
|
39
|
-
## Reading the numbers
|
|
40
|
-
|
|
41
|
-
- Tokens come from
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
134
|
+
## Reading the numbers
|
|
135
|
+
|
|
136
|
+
- Tokens come from provider telemetry when available; missing telemetry is recorded as missing,
|
|
137
|
+
never estimated. New Codex rows separate cache reads from fresh input; adaptive rows retain
|
|
138
|
+
controller/worker/parent call splits.
|
|
139
|
+
- The two older A/B rows are N=1 diagnostics. The Codex-only study above is N=3 per arm with
|
|
140
|
+
alternating pair order; it is stronger evidence but still workload-specific.
|
|
141
|
+
- Fixture size changes the outcome: the small queue delegated to a lower-effort worker and saved
|
|
142
|
+
tokens; the architecture/privacy task stayed on SELF and paid controller overhead; the newer
|
|
143
|
+
bounded full-repository telemetry task delegated consistently and saved fresh-input tokens.
|
|
@@ -0,0 +1,96 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
import { readFileSync, readdirSync } from 'node:fs'
|
|
3
|
+
import { dirname, join } from 'node:path'
|
|
4
|
+
import { fileURLToPath } from 'node:url'
|
|
5
|
+
|
|
6
|
+
const resultsDir = join(dirname(fileURLToPath(import.meta.url)), 'results')
|
|
7
|
+
const rows = readdirSync(resultsDir)
|
|
8
|
+
.filter(file => file.startsWith('yoke-codex-study-codex-routing-') && file.endsWith('.json'))
|
|
9
|
+
.map(file => ({ file, ...JSON.parse(readFileSync(join(resultsDir, file), 'utf8')) }))
|
|
10
|
+
.filter(row => /^codex-only-pair[1-3]-(on|off)$/.test(row.sampleLabel))
|
|
11
|
+
|
|
12
|
+
const median = values => {
|
|
13
|
+
const sorted = [...values].sort((a, b) => a - b)
|
|
14
|
+
return sorted[Math.floor(sorted.length / 2)]
|
|
15
|
+
}
|
|
16
|
+
const pct = (from, to) => ((to - from) / from) * 100
|
|
17
|
+
const metric = (group, read) => {
|
|
18
|
+
const values = group.map(read)
|
|
19
|
+
return { values, min: Math.min(...values), median: median(values), max: Math.max(...values) }
|
|
20
|
+
}
|
|
21
|
+
|
|
22
|
+
if (rows.length !== 6) throw new Error(`expected six codex-only study rows, found ${rows.length}`)
|
|
23
|
+
for (const row of rows) {
|
|
24
|
+
if (row.verdict !== 'completed' || !row.finalTestsPass || row.iterations !== 2) throw new Error(`invalid study row: ${row.file}`)
|
|
25
|
+
if (row.finalVerification?.originalAcceptanceTestsExitCode !== 0 || row.finalVerification?.originalAcceptanceTestCount !== 6) {
|
|
26
|
+
throw new Error(`hidden acceptance replay failed: ${row.file}`)
|
|
27
|
+
}
|
|
28
|
+
}
|
|
29
|
+
|
|
30
|
+
const groups = Object.fromEntries(['off', 'on'].map(routing => {
|
|
31
|
+
const group = rows.filter(row => row.routing === routing).sort((a, b) => a.sampleLabel.localeCompare(b.sampleLabel))
|
|
32
|
+
return [routing, {
|
|
33
|
+
runs: group.map(row => row.file),
|
|
34
|
+
wallClockMs: metric(group, row => row.wallClockMs),
|
|
35
|
+
inputTokens: metric(group, row => row.tokenBreakdown.inputTokens),
|
|
36
|
+
cachedInputTokens: metric(group, row => row.tokenBreakdown.cachedInputTokens),
|
|
37
|
+
freshInputTokens: metric(group, row => row.tokenBreakdown.freshInputTokens),
|
|
38
|
+
outputTokens: metric(group, row => row.tokenBreakdown.outputTokens),
|
|
39
|
+
reasoningOutputTokens: metric(group, row => row.tokenBreakdown.reasoningOutputTokens),
|
|
40
|
+
}]
|
|
41
|
+
}))
|
|
42
|
+
|
|
43
|
+
const routed = rows.filter(row => row.routing === 'on')
|
|
44
|
+
const controllerCalls = routed.flatMap(row => row.modelCalls.filter(call => call.role === 'orchestrator'))
|
|
45
|
+
const workerCalls = routed.flatMap(row => row.modelCalls.filter(call => call.role === 'worker'))
|
|
46
|
+
if (workerCalls.length !== 6 || workerCalls.some(call => call.profile !== 'codex-luna' || call.requestedModel !== 'gpt-5.6-luna')) {
|
|
47
|
+
throw new Error('not every routed story executed on codex-luna')
|
|
48
|
+
}
|
|
49
|
+
|
|
50
|
+
const deltas = {}
|
|
51
|
+
for (const key of ['wallClockMs', 'inputTokens', 'cachedInputTokens', 'freshInputTokens', 'outputTokens', 'reasoningOutputTokens']) {
|
|
52
|
+
deltas[key] = {
|
|
53
|
+
absolute: groups.on[key].median - groups.off[key].median,
|
|
54
|
+
percent: pct(groups.off[key].median, groups.on[key].median),
|
|
55
|
+
}
|
|
56
|
+
}
|
|
57
|
+
|
|
58
|
+
const pairs = [1, 2, 3].map(pair => {
|
|
59
|
+
const off = rows.find(row => row.sampleLabel === `codex-only-pair${pair}-off`)
|
|
60
|
+
const on = rows.find(row => row.sampleLabel === `codex-only-pair${pair}-on`)
|
|
61
|
+
return {
|
|
62
|
+
pair,
|
|
63
|
+
order: pair === 2 ? ['off', 'on'] : ['on', 'off'],
|
|
64
|
+
wallClockMs: { off: off.wallClockMs, on: on.wallClockMs, percent: pct(off.wallClockMs, on.wallClockMs) },
|
|
65
|
+
freshInputTokens: { off: off.tokenBreakdown.freshInputTokens, on: on.tokenBreakdown.freshInputTokens, percent: pct(off.tokenBreakdown.freshInputTokens, on.tokenBreakdown.freshInputTokens) },
|
|
66
|
+
}
|
|
67
|
+
})
|
|
68
|
+
|
|
69
|
+
console.log(JSON.stringify({
|
|
70
|
+
schemaVersion: 1,
|
|
71
|
+
fixtureVersion: 'yoke-codex-study@1',
|
|
72
|
+
sampleCountPerArm: 3,
|
|
73
|
+
quality: {
|
|
74
|
+
completedRuns: 6,
|
|
75
|
+
completedStories: 12,
|
|
76
|
+
hiddenAcceptancePasses: 36,
|
|
77
|
+
hiddenAcceptanceChecks: 36,
|
|
78
|
+
iterationsPerStory: 1,
|
|
79
|
+
},
|
|
80
|
+
models: {
|
|
81
|
+
parent: { model: 'gpt-5.6-sol', reasoningEffort: 'high' },
|
|
82
|
+
orchestrator: { model: 'gpt-5.6-sol', reasoningEffort: 'low' },
|
|
83
|
+
worker: { model: 'gpt-5.6-luna', reasoningEffort: 'low' },
|
|
84
|
+
},
|
|
85
|
+
groups,
|
|
86
|
+
medianDeltas: deltas,
|
|
87
|
+
pairs,
|
|
88
|
+
controller: {
|
|
89
|
+
calls: controllerCalls.length,
|
|
90
|
+
totalDurationMs: controllerCalls.reduce((sum, call) => sum + call.durationMs, 0),
|
|
91
|
+
totalInputTokens: controllerCalls.reduce((sum, call) => sum + call.inputTokens, 0),
|
|
92
|
+
totalCachedInputTokens: controllerCalls.reduce((sum, call) => sum + (call.cachedInputTokens ?? 0), 0),
|
|
93
|
+
totalFreshInputTokens: controllerCalls.reduce((sum, call) => sum + call.inputTokens - (call.cachedInputTokens ?? 0), 0),
|
|
94
|
+
totalOutputTokens: controllerCalls.reduce((sum, call) => sum + call.outputTokens, 0),
|
|
95
|
+
},
|
|
96
|
+
}, null, 2))
|
|
@@ -0,0 +1,30 @@
|
|
|
1
|
+
canonVersion: 1.2.0
|
|
2
|
+
agents: [codex, claude, gemini]
|
|
3
|
+
loop:
|
|
4
|
+
enabled: true
|
|
5
|
+
runner:
|
|
6
|
+
agent: codex
|
|
7
|
+
model: gpt-5.6-sol
|
|
8
|
+
reasoningEffort: high
|
|
9
|
+
bare: true
|
|
10
|
+
routing:
|
|
11
|
+
enabled: true
|
|
12
|
+
strategy: balanced
|
|
13
|
+
maxCandidates: 3
|
|
14
|
+
workers:
|
|
15
|
+
- id: codex-light
|
|
16
|
+
agent: codex
|
|
17
|
+
reasoningEffort: low
|
|
18
|
+
costTier: medium
|
|
19
|
+
capabilities: [exploration, implementation, tests]
|
|
20
|
+
- id: claude-fast
|
|
21
|
+
agent: claude
|
|
22
|
+
model: haiku
|
|
23
|
+
costTier: low
|
|
24
|
+
capabilities: [exploration, mechanical-edits, tests]
|
|
25
|
+
- id: gemini-auto
|
|
26
|
+
agent: gemini
|
|
27
|
+
costTier: low
|
|
28
|
+
capabilities: [large-context, implementation, tests]
|
|
29
|
+
verify:
|
|
30
|
+
command: node bench-verify.mjs
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
- id: STORY-1
|
|
2
|
+
title: "Priority task queue with idempotency and leases"
|
|
3
|
+
priority: 1
|
|
4
|
+
acceptance:
|
|
5
|
+
- "Export a TaskQueue class from src/task-queue.mjs; constructor accepts optional clock, maxAttempts, and baseDelayMs"
|
|
6
|
+
- "enqueue({type,payload,priority,idempotencyKey}) validates a non-empty type, defaults priority to 0, and returns an immutable task snapshot"
|
|
7
|
+
- "Repeated non-empty idempotencyKey returns the original task without adding another queue entry"
|
|
8
|
+
- "claim(workerId,{leaseMs}) returns the available queued task with highest priority and FIFO order for ties; it records workerId, leaseUntil, and increments attempts"
|
|
9
|
+
- "complete(taskId,workerId,result) only allows the current lease owner and returns an immutable completed snapshot"
|
|
10
|
+
- "get(taskId) and list() return snapshots that callers cannot use to mutate queue state"
|
|
11
|
+
- "node bench-verify.mjs exits 0"
|
|
12
|
+
passes: false
|
|
13
|
+
- id: STORY-2
|
|
14
|
+
title: "Retry backoff, dead letters, expired leases, and statistics"
|
|
15
|
+
priority: 2
|
|
16
|
+
acceptance:
|
|
17
|
+
- "fail(taskId,workerId,error) only allows the lease owner; before maxAttempts it requeues with exponential delay baseDelayMs * 2^(attempts-1)"
|
|
18
|
+
- "A task is not claimable before availableAt; at maxAttempts failure moves it to dead status"
|
|
19
|
+
- "reapExpired() requeues active tasks whose lease has expired and returns the number requeued"
|
|
20
|
+
- "list({status,type}) filters without exposing mutable internal objects"
|
|
21
|
+
- "stats() returns total and counts for queued, active, completed, and dead tasks"
|
|
22
|
+
- "Invalid task ids, invalid worker ids, and invalid state transitions throw clear errors"
|
|
23
|
+
- "node bench-verify.mjs exits 0"
|
|
24
|
+
passes: false
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
import { spawnSync } from 'node:child_process'
|
|
2
|
+
|
|
3
|
+
const total = 2
|
|
4
|
+
const story = process.env.YOKE_STORY
|
|
5
|
+
const requested = story ? Number(story.split('-')[1]) : total
|
|
6
|
+
const upTo = Number.isFinite(requested) && requested >= 1 && requested <= total ? requested : total
|
|
7
|
+
const files = Array.from({ length: upTo }, (_, index) => `tests/STORY-${index + 1}.test.mjs`)
|
|
8
|
+
const result = spawnSync(process.execPath, ['--test', ...files], { stdio: 'inherit' })
|
|
9
|
+
process.exit(result.status ?? 1)
|