@hecer/yoke 1.2.0 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/CHANGELOG.md +44 -18
- package/README.md +126 -96
- package/canon/loop/loop-spec.md +47 -32
- package/canon/loop/prd.schema.md +30 -6
- package/canon/manifest.yaml +1 -1
- package/canon/skills/authoring-prd/SKILL.md +30 -31
- package/dist/change/inbox.js +279 -0
- package/dist/cli.js +35 -1
- package/dist/loop/evidence.js +31 -0
- package/dist/loop/gates.js +10 -1
- package/dist/loop/loop.js +127 -17
- package/dist/loop/prd.js +60 -2
- package/dist/loop/run-command.js +23 -2
- package/dist/loop/runner.js +10 -2
- package/dist/loop/verify.js +11 -0
- package/dist/prd/command.js +18 -5
- package/dist/retrofit/config.js +9 -2
- package/dist/retrofit/gitignore.js +1 -0
- package/dist/routing/router.js +4 -1
- package/gemini-extension.json +1 -1
- package/package.json +6 -6
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
|
|
3
3
|
"name": "yoke",
|
|
4
4
|
"displayName": "Yoke",
|
|
5
|
-
"version": "1.
|
|
5
|
+
"version": "1.3.0",
|
|
6
6
|
"description": "Cross-agent coding harness: one curated skill canon (TDD, brainstorming, plans, reviews, shipping, design verification) plus mechanical safety gates and an autonomous loop via the yoke CLI.",
|
|
7
7
|
"author": { "name": "HECer", "url": "https://github.com/HECer" },
|
|
8
8
|
"homepage": "https://github.com/HECer/yoke#readme",
|
package/CHANGELOG.md
CHANGED
|
@@ -1,21 +1,47 @@
|
|
|
1
|
-
# Changelog
|
|
2
|
-
|
|
3
|
-
## 1.
|
|
4
|
-
|
|
5
|
-
### Added
|
|
6
|
-
-
|
|
7
|
-
-
|
|
8
|
-
-
|
|
9
|
-
|
|
10
|
-
### Changed
|
|
11
|
-
-
|
|
12
|
-
-
|
|
13
|
-
|
|
14
|
-
### Fixed
|
|
15
|
-
-
|
|
16
|
-
-
|
|
17
|
-
|
|
18
|
-
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 1.3.0 — 2026-08-09
|
|
4
|
+
|
|
5
|
+
### Added
|
|
6
|
+
- Teams can submit change requests at any time with `yoke change add` and inspect the append-only inbox with `yoke change status`; queued requests are planned at safe loop boundaries by Claude, Codex, or Gemini without interrupting active story work.
|
|
7
|
+
- Acceptance criteria can carry stable IDs and executable verification commands, and Yoke records per-criterion evidence before a story can be marked complete.
|
|
8
|
+
- Projects can configure an integrated completion command that proves the final system flow after all story-level gates pass.
|
|
9
|
+
|
|
10
|
+
### Changed
|
|
11
|
+
- New projects use strict criterion-proof requirements by default, while existing boolean-only criteria remain readable for compatibility.
|
|
12
|
+
- PRD authoring, schemas, loop guidance, generated configuration, and provider instructions now treat completion as an ephemeral readiness state: new requests create more stories instead of prematurely forcing a release boundary.
|
|
13
|
+
|
|
14
|
+
### Fixed
|
|
15
|
+
- Broad green test suites can no longer satisfy unrelated structured acceptance criteria: proof commands must use an approved test runner, contain the criterion ID, and avoid shell operators.
|
|
16
|
+
- Change-intake failures block safely; an independent coverage pass rejects omitted requested outcomes, crash recovery retains uncommitted requests, and concurrent PRD edits are detected instead of silently overwriting user work.
|
|
17
|
+
- Integrated completion checks prevent locally finished stories from masking broken cross-component flows such as authentication callbacks or purchase-to-entitlement activation.
|
|
18
|
+
|
|
19
|
+
### Security
|
|
20
|
+
- Updated the transitive development dependency `nanoid` to a non-vulnerable release so the complete CI dependency audit is clean.
|
|
21
|
+
- Restricted model-authored criterion proof to criterion-targeted test commands and blocked shell control operators and runner-prefix spoofing before host execution.
|
|
22
|
+
|
|
23
|
+
## 1.2.1 — 2026-08-02
|
|
24
|
+
|
|
25
|
+
### Fixed
|
|
26
|
+
- `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
|
|
27
|
+
- Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
|
|
28
|
+
|
|
29
|
+
## 1.2.0 — 2026-08-02
|
|
30
|
+
|
|
31
|
+
### Added
|
|
32
|
+
- Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
|
|
33
|
+
- A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
|
|
34
|
+
- Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
|
|
35
|
+
|
|
36
|
+
### Changed
|
|
37
|
+
- Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
|
|
38
|
+
- Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
|
|
39
|
+
|
|
40
|
+
### Fixed
|
|
41
|
+
- Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
|
|
42
|
+
- Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
|
|
43
|
+
|
|
44
|
+
## 1.1.0 — 2026-07-30
|
|
19
45
|
|
|
20
46
|
### Added
|
|
21
47
|
- Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
|
package/README.md
CHANGED
|
@@ -2,8 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
# 🐂 Yoke
|
|
4
4
|
|
|
5
|
-
<!-- yoke:version:start -->1.
|
|
6
|
-
<!-- yoke:tests:start -->
|
|
5
|
+
<!-- yoke:version:start -->1.3.0<!-- yoke:version:end -->
|
|
6
|
+
<!-- yoke:tests:start -->657<!-- yoke:tests:end -->
|
|
7
7
|
<!-- yoke:skills:start -->29<!-- yoke:skills:end -->
|
|
8
8
|
<!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
|
|
9
9
|
|
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
[](#-license)
|
|
18
18
|

|
|
19
19
|

|
|
20
|
-

|
|
21
21
|

|
|
22
22
|

|
|
23
23
|
|
|
@@ -25,7 +25,7 @@
|
|
|
25
25
|
|
|
26
26
|
</div>
|
|
27
27
|
|
|
28
|
-
> **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
|
|
28
|
+
> **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
|
|
29
29
|
|
|
30
30
|
Yoke 1.1 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
|
|
31
31
|
is explicit; reviews require a schema-valid verdict and a different model unless
|
|
@@ -72,7 +72,7 @@ $ ls reading-app/.yoke/proof/STORY-2/
|
|
|
72
72
|
home.png list.png # photographic evidence, labelled per story
|
|
73
73
|
```
|
|
74
74
|
|
|
75
|
-
Every claim in that transcript is enforced by code paths with tests behind them —
|
|
75
|
+
Every claim in that transcript is enforced by code paths with tests behind them — 657 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
|
|
76
76
|
|
|
77
77
|
## 🚀 Quickstart
|
|
78
78
|
|
|
@@ -85,7 +85,7 @@ yoke new my-app --idea="a CLI that tracks reading lists"
|
|
|
85
85
|
yoke loop on my-app && yoke loop run my-app --isolate
|
|
86
86
|
|
|
87
87
|
# — or retrofit an existing project —
|
|
88
|
-
yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
|
|
88
|
+
yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
|
|
89
89
|
yoke validate canon # sanity-check the canon
|
|
90
90
|
yoke loop run /path/to/project --isolate --reviewer=codex --max=20
|
|
91
91
|
```
|
|
@@ -104,7 +104,7 @@ The canon is also packaged as a Claude Code plugin — the repo is its own marke
|
|
|
104
104
|
That gives you all canon skills under the `yoke:` namespace (e.g. `yoke:tdd`, `yoke:review`) inside Claude Code — no retrofit needed. The `yoke` CLI (loop, gates, retrofit for Codex/Gemini) still comes from `npm i -g @hecer/yoke`. Gemini CLI users can likewise `gemini extensions install https://github.com/HECer/yoke`.
|
|
105
105
|
|
|
106
106
|
For Codex, no preinstalled skill is required: run `npx @hecer/yoke setup .` in a terminal, or
|
|
107
|
-
ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
|
|
107
|
+
ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
|
|
108
108
|
`.agents/skills/`, including `yoke-retrofit` and `yoke-workflow`; start a fresh Codex task if an
|
|
109
109
|
already-open task does not discover newly installed skills. The npm package also contains
|
|
110
110
|
`.codex-plugin/plugin.json` for Codex plugin hosts.
|
|
@@ -138,7 +138,7 @@ Yoke is meant to be operated *by* your coding agent — after a retrofit, the ag
|
|
|
138
138
|
|
|
139
139
|
> **Watch / unblock** — *"Run `yoke loop status .`. If it says BLOCKED, run the project's verify command, find the root cause, fix it without weakening tests, then continue the loop."*
|
|
140
140
|
|
|
141
|
-
> ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
|
|
141
|
+
> ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
|
|
142
142
|
|
|
143
143
|
> ⚠️ **Never kill agent processes by name or command-line pattern** (e.g. every process matching `dangerously-skip-permissions`): on a machine running several yoke projects, that takes down the *healthy* runners of the other projects mid-story — they stall and their loops block. `yoke loop cleanup` is the scoped alternative: each watchdog records its pids in the project's `.yoke/runner.pid`, and cleanup kills exactly those recorded trees — nothing else on the machine.
|
|
144
144
|
|
|
@@ -148,14 +148,15 @@ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`)
|
|
|
148
148
|
|
|
149
149
|
| Command | What it does | Exit codes |
|
|
150
150
|
|---|---|---|
|
|
151
|
-
| `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
|
|
151
|
+
| `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
|
|
152
152
|
| `yoke validate [canonDir]` | Validate the canon (schema, frontmatter, templates) | `0` valid · `1` errors |
|
|
153
153
|
| `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
|
|
154
154
|
| `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
|
|
155
155
|
| `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
|
|
156
156
|
| `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
|
|
157
|
+
| `yoke change add\|status [dir] [--idea=]` | Queue a change at any time; the loop turns it into append-only stories at the next safe boundary | `0` · `1` invalid inbox/request |
|
|
157
158
|
| `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
|
|
158
|
-
| `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` is unlimited by default and `--max=N` creates an intentional batch cap; `decision` shows a critical stop, `answer` records it and resumes, `resume` retries a failed restart with the preserved safety options | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
|
|
159
|
+
| `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` is unlimited by default and `--max=N` creates an intentional batch cap; `decision` shows a critical stop, `answer` records it and resumes, `resume` retries a failed restart with the preserved safety options | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
|
|
159
160
|
| `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
|
|
160
161
|
| `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
|
|
161
162
|
| `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
|
|
@@ -310,28 +311,34 @@ Opt-in; `yoke setup` recommends enabling it for new installs, while `retrofit` a
|
|
|
310
311
|
|
|
311
312
|
```mermaid
|
|
312
313
|
flowchart LR
|
|
313
|
-
A[pick next PRD story]
|
|
314
|
+
I[consume queued change<br/>as new stories] --> A[pick next PRD story]
|
|
315
|
+
A --> B{clean worktree?}
|
|
314
316
|
B -- no --> X[blocked]
|
|
315
317
|
B -- yes --> C{acceptance<br/>criteria?}
|
|
316
318
|
C -- no --> X
|
|
317
319
|
C -- yes --> D[agent implements<br/>one story]
|
|
318
|
-
D --> E{
|
|
320
|
+
D --> E{suite + criterion<br/>proof green?}
|
|
319
321
|
E -- no --> X
|
|
320
322
|
E -- yes --> F{reviewer<br/>approves?}
|
|
321
323
|
F -- no --> X
|
|
322
324
|
F -- yes --> G[commit + mark passes:true<br/>+ proof in .yoke/proof/]
|
|
323
|
-
G -->
|
|
325
|
+
G --> I
|
|
326
|
+
I --> H{all stories pass?}
|
|
327
|
+
H -- yes --> J{integrated system<br/>gate green?}
|
|
328
|
+
J -- no --> X
|
|
329
|
+
J -- yes --> K[current backlog ready]
|
|
324
330
|
```
|
|
325
331
|
|
|
326
332
|
```bash
|
|
327
333
|
yoke loop on . # enable (recorded in .yoke/config.yaml)
|
|
328
334
|
yoke loop status . # show state + PRD progress
|
|
335
|
+
yoke change add . --idea="Add passkey login" # safe while the loop runs
|
|
329
336
|
yoke loop run . \
|
|
330
337
|
--runner=codex \ # implement with Codex…
|
|
331
338
|
--reviewer=claude \ # …review with Claude (role separation)
|
|
332
339
|
--isolate \ # each story in a throwaway git worktree
|
|
333
|
-
--decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
|
|
334
|
-
# Optional: add --max=20 only when this run should stop after a bounded batch.
|
|
340
|
+
--decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
|
|
341
|
+
# Optional: add --max=20 only when this run should stop after a bounded batch.
|
|
335
342
|
yoke loop off . # disable
|
|
336
343
|
```
|
|
337
344
|
|
|
@@ -342,11 +349,32 @@ yoke loop off . # disable
|
|
|
342
349
|
title: Add a health endpoint
|
|
343
350
|
priority: 1 # lower = higher priority
|
|
344
351
|
acceptance: # Definition of Done (required, else blocked)
|
|
345
|
-
-
|
|
352
|
+
- id: health-returns-200
|
|
353
|
+
text: GET /health returns 200
|
|
354
|
+
verify: [npm run test:health-returns-200]
|
|
355
|
+
- id: health-rejects-post
|
|
356
|
+
text: POST /health returns 405
|
|
357
|
+
verify: [npm run test:health-rejects-post]
|
|
346
358
|
passes: false # the loop sets this true only on green tests
|
|
347
359
|
```
|
|
348
360
|
|
|
349
|
-
|
|
361
|
+
New projects default to `verify.requireCriteria: true`: every story has 2–5 behavioral criteria.
|
|
362
|
+
Each criterion ID must occur in its single, approved test command; shell operators and broad,
|
|
363
|
+
untargeted suites are rejected. Yoke records each result in `.yoke/proof/<story>/evidence.json`. Configure optional
|
|
364
|
+
`completion.command` for integrated journeys such as purchase → entitlement → relaunch or
|
|
365
|
+
magic-link → callback → authenticated app. It runs whenever the current backlog has no open
|
|
366
|
+
stories; this is readiness, not a release.
|
|
367
|
+
|
|
368
|
+
No generic tool can infer whether arbitrary test code perfectly represents product meaning. Yoke
|
|
369
|
+
closes the mechanical false-done paths—targeted evidence, coverage review, clean committed state,
|
|
370
|
+
and integrated journeys—while the project still owns the correctness of its tests and production
|
|
371
|
+
observability.
|
|
372
|
+
|
|
373
|
+
State lives **outside the model context** — the PRD file plus git — so each iteration is fresh.
|
|
374
|
+
Use `yoke change add` at any time. Its ignored append-only inbox is consumed at the next story
|
|
375
|
+
boundary. A separate coverage pass must confirm that every requested outcome maps to behavioral
|
|
376
|
+
criteria before Yoke appends and commits the new stories; existing stories are never rewritten and
|
|
377
|
+
no restart is needed.
|
|
350
378
|
|
|
351
379
|
### Watching a run
|
|
352
380
|
|
|
@@ -370,9 +398,9 @@ Every iteration emits token-free, harness-side feedback (Node console + local fi
|
|
|
370
398
|
NDJSON line on stdout (`{"type":"status","state":"running","phase":"verifying",…}` — the
|
|
371
399
|
same shape as `loop-status.json`), the human narrative moves off stdout (the final summary
|
|
372
400
|
goes to stderr), and a consumer can follow the stream line by line instead of polling the file.
|
|
373
|
-
Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
|
|
374
|
-
output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
|
|
375
|
-
and cost fields when available. Missing values stay absent—Yoke does not estimate them.
|
|
401
|
+
Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
|
|
402
|
+
output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
|
|
403
|
+
and cost fields when available. Missing values stay absent—Yoke does not estimate them.
|
|
376
404
|
|
|
377
405
|
### Pausing a run
|
|
378
406
|
|
|
@@ -432,65 +460,65 @@ cannot begin because a provider/reviewer is unavailable or another process owns
|
|
|
432
460
|
until a loop actually runs. To intentionally abandon an orphaned or stale private resume state,
|
|
433
461
|
use `yoke loop resume . --discard`; pending decisions are never deleted by that command. Existing
|
|
434
462
|
`loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` remain supported as compatibility aliases;
|
|
435
|
-
new projects should use `decisionPolicy: auto|critical`.
|
|
436
|
-
|
|
437
|
-
### Adaptive model routing (explicit opt-in)
|
|
438
|
-
|
|
439
|
-
`yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
|
|
440
|
-
parent remains the strong planner/controller. Before each bounded story it receives only the
|
|
441
|
-
story, acceptance criteria, and at most three eligible worker profiles, then returns one
|
|
442
|
-
machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
|
|
443
|
-
`SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
|
|
444
|
-
runs so Yoke does not pay for two orchestration layers.
|
|
445
|
-
|
|
446
|
-
**Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
|
|
447
|
-
Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
|
|
448
|
-
cover invocation and routing behavior for all three providers. The measured performance evidence
|
|
449
|
-
below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
|
|
450
|
-
until authenticated, repeated in-the-wild runs exist for those providers.
|
|
451
|
-
|
|
452
|
-
```yaml
|
|
453
|
-
runner:
|
|
454
|
-
agent: codex
|
|
455
|
-
model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
|
|
456
|
-
reasoningEffort: high
|
|
457
|
-
routing:
|
|
458
|
-
enabled: true # setup defaults false; setup --routing opts in
|
|
459
|
-
strategy: balanced # balanced | cost | speed | quality
|
|
460
|
-
maxCandidates: 3
|
|
461
|
-
workers:
|
|
462
|
-
- id: codex-light
|
|
463
|
-
agent: codex
|
|
464
|
-
reasoningEffort: low
|
|
465
|
-
costTier: medium
|
|
466
|
-
capabilities: [exploration, implementation, tests]
|
|
467
|
-
- id: claude-fast
|
|
468
|
-
agent: claude
|
|
469
|
-
model: haiku # rolling alias; omit to use the provider's current default
|
|
470
|
-
reasoningEffort: low
|
|
471
|
-
costTier: low
|
|
472
|
-
capabilities: [mechanical-edits, tests]
|
|
473
|
-
- id: gemini-auto
|
|
474
|
-
agent: gemini # omitted model means the account's current Auto/default route
|
|
475
|
-
costTier: low
|
|
476
|
-
capabilities: [large-context, implementation]
|
|
477
|
-
```
|
|
478
|
-
|
|
479
|
-
Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
|
|
480
|
-
baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
|
|
481
|
-
eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
|
|
482
|
-
"intelligence score". Candidate model IDs come from project configuration while setup defaults
|
|
483
|
-
prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
|
|
484
|
-
independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
|
|
485
|
-
after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
|
|
486
|
-
time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
|
|
487
|
-
cannot overwrite a shared registry file.
|
|
488
|
-
|
|
489
|
-
Routing is not free: it adds one controller call per story. It is most promising when a bounded
|
|
490
|
-
worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
|
|
491
|
-
on your own backlog rather than assuming a win.
|
|
492
|
-
|
|
493
|
-
### Performance budgets: efficiency as a gate, not a style
|
|
463
|
+
new projects should use `decisionPolicy: auto|critical`.
|
|
464
|
+
|
|
465
|
+
### Adaptive model routing (explicit opt-in)
|
|
466
|
+
|
|
467
|
+
`yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
|
|
468
|
+
parent remains the strong planner/controller. Before each bounded story it receives only the
|
|
469
|
+
story, acceptance criteria, and at most three eligible worker profiles, then returns one
|
|
470
|
+
machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
|
|
471
|
+
`SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
|
|
472
|
+
runs so Yoke does not pay for two orchestration layers.
|
|
473
|
+
|
|
474
|
+
**Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
|
|
475
|
+
Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
|
|
476
|
+
cover invocation and routing behavior for all three providers. The measured performance evidence
|
|
477
|
+
below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
|
|
478
|
+
until authenticated, repeated in-the-wild runs exist for those providers.
|
|
479
|
+
|
|
480
|
+
```yaml
|
|
481
|
+
runner:
|
|
482
|
+
agent: codex
|
|
483
|
+
model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
|
|
484
|
+
reasoningEffort: high
|
|
485
|
+
routing:
|
|
486
|
+
enabled: true # setup defaults false; setup --routing opts in
|
|
487
|
+
strategy: balanced # balanced | cost | speed | quality
|
|
488
|
+
maxCandidates: 3
|
|
489
|
+
workers:
|
|
490
|
+
- id: codex-light
|
|
491
|
+
agent: codex
|
|
492
|
+
reasoningEffort: low
|
|
493
|
+
costTier: medium
|
|
494
|
+
capabilities: [exploration, implementation, tests]
|
|
495
|
+
- id: claude-fast
|
|
496
|
+
agent: claude
|
|
497
|
+
model: haiku # rolling alias; omit to use the provider's current default
|
|
498
|
+
reasoningEffort: low
|
|
499
|
+
costTier: low
|
|
500
|
+
capabilities: [mechanical-edits, tests]
|
|
501
|
+
- id: gemini-auto
|
|
502
|
+
agent: gemini # omitted model means the account's current Auto/default route
|
|
503
|
+
costTier: low
|
|
504
|
+
capabilities: [large-context, implementation]
|
|
505
|
+
```
|
|
506
|
+
|
|
507
|
+
Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
|
|
508
|
+
baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
|
|
509
|
+
eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
|
|
510
|
+
"intelligence score". Candidate model IDs come from project configuration while setup defaults
|
|
511
|
+
prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
|
|
512
|
+
independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
|
|
513
|
+
after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
|
|
514
|
+
time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
|
|
515
|
+
cannot overwrite a shared registry file.
|
|
516
|
+
|
|
517
|
+
Routing is not free: it adds one controller call per story. It is most promising when a bounded
|
|
518
|
+
worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
|
|
519
|
+
on your own backlog rather than assuming a win.
|
|
520
|
+
|
|
521
|
+
### Performance budgets: efficiency as a gate, not a style
|
|
494
522
|
|
|
495
523
|
Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
|
|
496
524
|
be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
|
|
@@ -517,12 +545,13 @@ be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two lev
|
|
|
517
545
|
The loop trusts **verify**, not the agent's exit code: a story whose tests are green is
|
|
518
546
|
committed even if the agent process exited non-zero (a common Windows `.cmd`-wrapper ghost).
|
|
519
547
|
A failing verify is retried up to `verify.retries` times (default 1) so a transient flake
|
|
520
|
-
self-heals while a real failure still blocks.
|
|
548
|
+
self-heals while a real failure still blocks. Structured acceptance criteria are then verified
|
|
549
|
+
individually; an unrelated green suite cannot satisfy a criterion without its proof command.
|
|
521
550
|
|
|
522
551
|
`.yoke/loop-status.json`, `.yoke/loop.log`, `.yoke/loop.lock`, its takeover/recovery leases, lock/decision temp files, `.yoke/story-durations.json`,
|
|
523
552
|
`.yoke/ambiguity.md`, and the critical-decision request/answering files are runtime artifacts;
|
|
524
553
|
`yoke retrofit` gitignores them (along with
|
|
525
|
-
`.yoke/worktrees/`, `.yoke/backup/`, and `.yoke/
|
|
554
|
+
`.yoke/worktrees/`, `.yoke/backup/`, `.yoke/proof/`, and `.yoke/changes/`) so they never trip the clean-tree gate.
|
|
526
555
|
|
|
527
556
|
### Single-flight guard + cleanup
|
|
528
557
|
|
|
@@ -652,24 +681,24 @@ Yoke's guardrails are **mechanical, not advisory** — the loop blocks on a dirt
|
|
|
652
681
|
| Best for | rapid exploration / migration / onboarding | systematic refactoring in typed codebases |
|
|
653
682
|
| Caveat | heuristic edges; static index can go stale | one language server per language |
|
|
654
683
|
|
|
655
|
-
## 🪙 Token efficiency
|
|
684
|
+
## 🪙 Token efficiency
|
|
656
685
|
|
|
657
686
|
Yoke attacks tokens on two complementary surfaces:
|
|
658
687
|
|
|
659
688
|
- **rtk** compresses noisy command/tool output before it enters context (wired as a hook/instruction per agent).
|
|
660
|
-
- The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
|
|
661
|
-
|
|
662
|
-
A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
|
|
663
|
-
completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
|
|
664
|
-
controller selected Luna for every bounded implementation story. Including controller overhead,
|
|
665
|
-
the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
|
|
666
|
-
and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
|
|
667
|
-
|
|
668
|
-
The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
|
|
669
|
-
controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
|
|
670
|
-
emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
|
|
671
|
-
inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
|
|
672
|
-
[`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
|
|
689
|
+
- The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
|
|
690
|
+
|
|
691
|
+
A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
|
|
692
|
+
completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
|
|
693
|
+
controller selected Luna for every bounded implementation story. Including controller overhead,
|
|
694
|
+
the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
|
|
695
|
+
and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
|
|
696
|
+
|
|
697
|
+
The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
|
|
698
|
+
controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
|
|
699
|
+
emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
|
|
700
|
+
inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
|
|
701
|
+
[`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
|
|
673
702
|
|
|
674
703
|
## 🧩 Optional companions
|
|
675
704
|
|
|
@@ -693,6 +722,7 @@ canon/ # the source of truth — harness-agnostic
|
|
|
693
722
|
AGENTS.md skills/ policy/ loop/ tools/ manifest.yaml
|
|
694
723
|
src/
|
|
695
724
|
canon/ # manifest schema + validator (yoke validate)
|
|
725
|
+
change/ # append-only change inbox · planning · independent coverage review
|
|
696
726
|
retrofit/ # detect · plan · apply · planners (claude/codex/gemini) · tools
|
|
697
727
|
loop/ # prd · gates · runner · verify · git/worktree · loop · run-command · lock · cleanup
|
|
698
728
|
new/ # yoke new — greenfield bootstrap
|
|
@@ -713,7 +743,7 @@ parallel dispatcher, broader benchmark samples, native output schemas, and relea
|
|
|
713
743
|
## 🧪 Development
|
|
714
744
|
|
|
715
745
|
```bash
|
|
716
|
-
npm test # vitest (
|
|
746
|
+
npm test # vitest (657 tests)
|
|
717
747
|
npm run build # tsc, no emit errors
|
|
718
748
|
npm run yoke -- validate canon
|
|
719
749
|
```
|
package/canon/loop/loop-spec.md
CHANGED
|
@@ -1,36 +1,51 @@
|
|
|
1
1
|
# Loop Specification (Ralph + GSD)
|
|
2
2
|
|
|
3
|
-
The autonomous loop is
|
|
4
|
-
|
|
5
|
-
- `yoke loop on` / `yoke loop off` — enable
|
|
6
|
-
- `yoke loop status` — show enabled state
|
|
7
|
-
- `yoke loop run [--max=N] [--isolate] [--decision-policy=auto|critical]` — run the
|
|
8
|
-
- `yoke
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
3
|
+
The autonomous loop is optional and toggle-able:
|
|
4
|
+
|
|
5
|
+
- `yoke loop on` / `yoke loop off` — enable or disable it in `.yoke/config.yaml`.
|
|
6
|
+
- `yoke loop status` — show enabled state and backlog progress.
|
|
7
|
+
- `yoke loop run [--max=N] [--isolate] [--decision-policy=auto|critical]` — run until the current backlog is green or a gate blocks.
|
|
8
|
+
- `yoke change add --idea="..."` — queue a product change at any time, including while the loop is running.
|
|
9
|
+
- `yoke loop decision` / `yoke loop answer --choice=<id>` — inspect and answer a structured critical stop.
|
|
10
|
+
|
|
11
|
+
Pass `--isolate` to implement each story in a fresh git worktree. Only a verified, committed
|
|
12
|
+
story is fast-forwarded to the main tree. Pass `--review` or `--reviewer=<provider>` to require
|
|
13
|
+
a separate, schema-validated review. Pass `--json` for NDJSON status on stdout.
|
|
14
|
+
|
|
15
|
+
At every story boundary, Yoke consumes at most one queued change. The configured Claude,
|
|
16
|
+
Codex, or Gemini provider may propose only new stories in a separate runtime file. A fresh
|
|
17
|
+
coverage-review pass must account for every distinct requested outcome before Yoke validates
|
|
18
|
+
strict criterion evidence, appends the stories itself, commits only the PRD, and leaves existing
|
|
19
|
+
stories untouched. The request stays pending on any failure or uncovered outcome.
|
|
20
|
+
|
|
21
|
+
For each story:
|
|
22
|
+
|
|
23
|
+
1. Require a clean git worktree.
|
|
24
|
+
2. Pick the highest-priority ready unfinished story.
|
|
25
|
+
3. Stop the line if acceptance is empty. With `verify.requireCriteria: true`, every criterion
|
|
26
|
+
must be structured. Every structured criterion, including in compatible legacy projects,
|
|
27
|
+
must use a single approved test command containing its criterion ID and no shell operators.
|
|
28
|
+
4. Run a fresh configured provider to implement exactly one story. Under the `critical`
|
|
29
|
+
decision policy, only high-impact architecture, security/privacy, destructive data,
|
|
30
|
+
material cost, compliance, or irreversible choices may pause for a human decision.
|
|
31
|
+
5. Run every structured criterion's targeted commands and write
|
|
32
|
+
`.yoke/proof/<story>/evidence.json`. Then run project-wide `verify.command` (or detected
|
|
33
|
+
`npm test`). Performance, audit, and independent review gates follow when configured. Any
|
|
34
|
+
failure blocks, and no proof command runs after review.
|
|
35
|
+
6. Only after all gates pass, mark the story `passes: true`, log the decision, and commit
|
|
36
|
+
atomically. A failed commit restores the PRD state.
|
|
37
|
+
7. When all current stories pass, run optional `completion.command` against the integrated
|
|
38
|
+
system. Only a green result reports `complete`; otherwise the loop blocks. This readiness
|
|
39
|
+
result is ephemeral, not a release and not a freeze on future changes.
|
|
40
|
+
|
|
41
|
+
A supervisor can pause the loop by creating `.yoke/loop.pause`. The running story finishes;
|
|
42
|
+
the signal is consumed at the next story boundary and the process exits with code `3`.
|
|
43
|
+
|
|
44
|
+
State lives outside model context: PRD, git, and the ignored `.yoke/changes/` inbox. All are
|
|
45
|
+
re-read at story boundaries, so a request queued mid-run becomes additional stories without a
|
|
46
|
+
restart.
|
|
34
47
|
|
|
35
48
|
## Limitations
|
|
36
|
-
|
|
49
|
+
|
|
50
|
+
Yoke cannot infer the correct end-to-end journey command. Projects that need integrated
|
|
51
|
+
readiness must configure `completion.command`, for example a Playwright journey suite.
|
package/canon/loop/prd.schema.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# PRD Schema
|
|
2
2
|
|
|
3
|
-
The loop is driven by a
|
|
3
|
+
The loop is driven by a continuous PRD backlog. Each story:
|
|
4
4
|
|
|
5
5
|
```yaml
|
|
6
6
|
- id: STORY-1
|
|
@@ -9,11 +9,35 @@ The loop is driven by a versioned PRD file. Each story:
|
|
|
9
9
|
needs: [] # optional dependency IDs; no unknown IDs, self-links, or cycles
|
|
10
10
|
area: api # optional collision domain for parallel scheduling
|
|
11
11
|
agent: codex # optional claude|codex|gemini affinity
|
|
12
|
-
acceptance:
|
|
13
|
-
-
|
|
14
|
-
|
|
12
|
+
acceptance:
|
|
13
|
+
- id: valid-request-returns-200
|
|
14
|
+
text: The endpoint returns 200 for a valid request.
|
|
15
|
+
verify: [npm run test:valid-request-returns-200]
|
|
16
|
+
- id: invalid-request-returns-400
|
|
17
|
+
text: The endpoint returns 400 for an invalid request.
|
|
18
|
+
verify: [npm run test:invalid-request-returns-400]
|
|
19
|
+
passes: false # Yoke-owned; true only after all gates pass and the commit lands
|
|
15
20
|
```
|
|
16
21
|
|
|
17
|
-
|
|
22
|
+
Each new story has 2–5 acceptance criteria. Every criterion has a stable `id`, observable
|
|
23
|
+
behavioral `text`, and one or more executable `verify` commands. Each entry is one approved test
|
|
24
|
+
command, contains the normalized criterion ID, and contains no shell control operator. Yoke runs and records each criterion separately in
|
|
25
|
+
`.yoke/proof/<story>/evidence.json`; a broad green suite cannot stand in for an untested
|
|
26
|
+
criterion. Legacy string criteria remain readable, but `verify.requireCriteria: true` blocks
|
|
27
|
+
them.
|
|
18
28
|
|
|
19
|
-
|
|
29
|
+
The binding is deliberately mechanical, not an oracle for product meaning: Yoke can require a
|
|
30
|
+
criterion-targeted test command and a separate coverage review, but it cannot prove that arbitrary
|
|
31
|
+
test code faithfully models the real world. Critical cross-component behavior therefore also
|
|
32
|
+
belongs in a trusted `completion.command` journey suite.
|
|
33
|
+
|
|
34
|
+
`sourceChange` is an optional Yoke-owned request ID. The change inbox uses it to append new
|
|
35
|
+
stories idempotently; authors normally omit it.
|
|
36
|
+
|
|
37
|
+
Stories without `needs`, `area`, or `agent` retain serial behavior. A story is ready only when
|
|
38
|
+
every ID in `needs` passes. The scheduler orders ready work by priority, avoids simultaneously
|
|
39
|
+
active areas, and uses `agent` as an affinity hint.
|
|
40
|
+
|
|
41
|
+
The backlog is continuous, not a release object. A momentary stop condition is every story
|
|
42
|
+
having `passes: true`; if configured, `completion.command` must then prove the integrated
|
|
43
|
+
system before the loop reports `complete`.
|
package/canon/manifest.yaml
CHANGED