@hecer/yoke 1.5.1 → 1.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (129) hide show
  1. package/.claude-plugin/plugin.json +13 -13
  2. package/.codex-plugin/plugin.json +7 -7
  3. package/CHANGELOG.md +280 -259
  4. package/README.md +855 -834
  5. package/TODOS.md +5 -5
  6. package/agents/docs.toml +6 -6
  7. package/agents/implementer.toml +6 -6
  8. package/agents/reviewer.toml +6 -6
  9. package/agents/security.toml +6 -6
  10. package/bench/README.md +86 -86
  11. package/bench/RESULTS.md +35 -35
  12. package/bench/output-compaction.mjs +65 -65
  13. package/bench/result-schema.mjs +12 -12
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -15
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
  17. package/bench/run-matrix.mjs +26 -26
  18. package/bench/run.mjs +106 -106
  19. package/canon/AGENTS.md +30 -30
  20. package/canon/context/DECISIONS.md +4 -4
  21. package/canon/context/GLOSSARY.md +11 -0
  22. package/canon/context/KNOWLEDGE.md +4 -4
  23. package/canon/context/PROJECT.md +15 -15
  24. package/canon/loop/loop-spec.md +65 -65
  25. package/canon/loop/prd.schema.md +43 -43
  26. package/canon/manifest.yaml +59 -53
  27. package/canon/policy/gates.md +7 -7
  28. package/canon/policy/roles.md +9 -9
  29. package/canon/skills/ATTRIBUTION.md +99 -71
  30. package/canon/skills/authoring-prd/SKILL.md +58 -58
  31. package/canon/skills/brainstorming/SKILL.md +164 -164
  32. package/canon/skills/codebase-design/DEEPENING.md +15 -0
  33. package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -0
  34. package/canon/skills/codebase-design/SKILL.md +39 -0
  35. package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
  36. package/canon/skills/document-release/SKILL.md +302 -297
  37. package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -0
  38. package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -0
  39. package/canon/skills/domain-modeling/SKILL.md +35 -0
  40. package/canon/skills/executing-plans/SKILL.md +70 -70
  41. package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
  42. package/canon/skills/health/SKILL.md +177 -177
  43. package/canon/skills/maintaining-context/SKILL.md +34 -34
  44. package/canon/skills/minimal-code/SKILL.md +21 -21
  45. package/canon/skills/no-ai-slop/SKILL.md +103 -0
  46. package/canon/skills/no-ai-slop/eval.md +43 -0
  47. package/canon/skills/plan-ceo-review/SKILL.md +541 -541
  48. package/canon/skills/plan-eng-review/SKILL.md +362 -362
  49. package/canon/skills/receiving-code-review/SKILL.md +213 -213
  50. package/canon/skills/requesting-code-review/SKILL.md +105 -105
  51. package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -0
  52. package/canon/skills/retro/SKILL.md +397 -397
  53. package/canon/skills/review/SKILL.md +246 -246
  54. package/canon/skills/ship/SKILL.md +691 -691
  55. package/canon/skills/subagent-driven-development/SKILL.md +277 -277
  56. package/canon/skills/systematic-debugging/SKILL.md +296 -296
  57. package/canon/skills/tdd/SKILL.md +371 -371
  58. package/canon/skills/unslop-ui/SKILL.md +34 -34
  59. package/canon/skills/using-git-worktrees/SKILL.md +218 -218
  60. package/canon/skills/verification-before-completion/SKILL.md +139 -139
  61. package/canon/skills/visual-verification/SKILL.md +54 -54
  62. package/canon/skills/workflow/SKILL.md +22 -22
  63. package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -0
  64. package/canon/skills/writing-for-agents/SKILL.md +42 -0
  65. package/canon/skills/writing-plans/SKILL.md +152 -152
  66. package/canon/skills/writing-skills/SKILL.md +655 -655
  67. package/canon/skills/yoke-retrofit/SKILL.md +26 -26
  68. package/canon/skills/yoke-workflow/SKILL.md +20 -20
  69. package/canon/tools/codex-rtk-hook.mjs +35 -35
  70. package/canon/tools/graphify.md +3 -3
  71. package/canon/tools/playwright-mcp.md +3 -3
  72. package/canon/tools/rtk.md +7 -7
  73. package/canon/tools/serena.md +6 -6
  74. package/dist/agents/process.js +3 -0
  75. package/dist/canon/manifest.js +2 -0
  76. package/dist/canon/skill-package.js +113 -0
  77. package/dist/canon/validate.js +16 -1
  78. package/dist/context/command.js +4 -1
  79. package/dist/context/context.js +6 -0
  80. package/dist/loop/dispatcher.js +1 -1
  81. package/dist/loop/loop.js +26 -0
  82. package/dist/loop/parallel-command.js +3 -0
  83. package/dist/loop/run-command.js +11 -0
  84. package/dist/loop/watchdog.js +28 -11
  85. package/dist/loop/worker.js +11 -0
  86. package/dist/prd/command.js +17 -17
  87. package/dist/retrofit/apply.js +22 -7
  88. package/dist/retrofit/command.js +4 -1
  89. package/dist/retrofit/config.js +4 -0
  90. package/dist/retrofit/context-actions.js +1 -1
  91. package/dist/retrofit/detect.js +2 -0
  92. package/dist/retrofit/planners/claude.js +16 -20
  93. package/dist/retrofit/planners/codex.js +3 -7
  94. package/dist/retrofit/planners/gemini.js +11 -1
  95. package/dist/retrofit/preserve.js +2 -2
  96. package/dist/retrofit/report.js +5 -0
  97. package/dist/retrofit/skill-actions.js +66 -0
  98. package/dist/retrofit/ui-detect.js +83 -0
  99. package/dist/scan/gate.js +36 -0
  100. package/docs/MIGRATING-TO-1.0.md +33 -33
  101. package/docs/MIGRATING-TO-1.1.md +27 -27
  102. package/docs/MIGRATING-TO-1.4.md +70 -70
  103. package/docs/PUBLISHING.md +91 -91
  104. package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
  105. package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
  106. package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
  107. package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
  108. package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
  109. package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
  110. package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
  111. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
  112. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
  113. package/docs/superpowers/plans/2026-08-20-automatic-ui-design-gate.md +59 -0
  114. package/docs/superpowers/plans/2026-08-20-capability-skills-and-context.md +51 -0
  115. package/docs/superpowers/plans/2026-08-20-complete-skill-packages-and-invocation.md +59 -0
  116. package/docs/superpowers/plans/2026-08-20-windows-reliability-and-release.md +67 -0
  117. package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
  118. package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
  119. package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
  120. package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
  121. package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
  122. package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
  123. package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
  124. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
  125. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
  126. package/docs/superpowers/specs/2026-08-20-skill-capabilities-and-reliability-design.md +391 -0
  127. package/gemini-extension.json +6 -6
  128. package/hooks/hooks.json +19 -19
  129. package/package.json +84 -84
package/README.md CHANGED
@@ -1,30 +1,30 @@
1
- <div align="center">
2
-
3
- # 🐂 Yoke
4
-
5
- <!-- yoke:version:start -->1.5.1<!-- yoke:version:end -->
6
- <!-- yoke:tests:start -->971<!-- yoke:tests:end -->
7
- <!-- yoke:skills:start -->29<!-- yoke:skills:end -->
8
- <!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
9
-
10
- ### One harness, three agents — and zero trust in "done."
11
-
12
- **Yoke** installs one curated canon of skills, **mechanical safety gates**, and tool wiring into any project — natively for **Claude Code, OpenAI Codex CLI, and Gemini CLI**. Then, when you want it, an opt-in autonomous loop ships your spec story-by-story: tested, cross-model-reviewed, committed — **with a screenshot to prove every story and a video for every failure**.
13
-
14
- [![npm](https://img.shields.io/npm/v/%40hecer%2Fyoke?logo=npm&color=CB3837)](https://www.npmjs.com/package/@hecer/yoke)
15
- [![npm downloads](https://img.shields.io/npm/dm/%40hecer%2Fyoke?logo=npm)](https://www.npmjs.com/package/@hecer/yoke)
16
- [![CI](https://github.com/HECer/yoke/actions/workflows/ci.yml/badge.svg)](https://github.com/HECer/yoke/actions/workflows/ci.yml)
17
- [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](#-license)
18
- ![Node](https://img.shields.io/badge/node-%E2%89%A520-339933?logo=node.js&logoColor=white)
19
- ![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
20
- ![Tests](https://img.shields.io/badge/tests-971%20passing-brightgreen.svg)
21
- ![Agents](https://img.shields.io/badge/agents-Claude%20%7C%20Codex%20%7C%20Gemini-8A2BE2)
22
- ![Built with TDD](https://img.shields.io/badge/built%20with-TDD%20%2B%20review-ff69b4.svg)
23
-
24
- **Install:** [`npm i -g @hecer/yoke`](https://www.npmjs.com/package/@hecer/yoke)
25
-
26
- </div>
27
-
1
+ <div align="center">
2
+
3
+ # 🐂 Yoke
4
+
5
+ <!-- yoke:version:start -->1.6.1<!-- yoke:version:end -->
6
+ <!-- yoke:tests:start -->1019<!-- yoke:tests:end -->
7
+ <!-- yoke:skills:start -->34<!-- yoke:skills:end -->
8
+ <!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
9
+
10
+ ### One harness, three agents — and zero trust in "done."
11
+
12
+ **Yoke** installs one curated canon of skills, **mechanical safety gates**, and tool wiring into any project — natively for **Claude Code, OpenAI Codex CLI, and Gemini CLI**. Then, when you want it, an opt-in autonomous loop ships your spec story-by-story: tested, cross-model-reviewed, committed — **with a screenshot to prove every story and a video for every failure**.
13
+
14
+ [![npm](https://img.shields.io/npm/v/%40hecer%2Fyoke?logo=npm&color=CB3837)](https://www.npmjs.com/package/@hecer/yoke)
15
+ [![npm downloads](https://img.shields.io/npm/dm/%40hecer%2Fyoke?logo=npm)](https://www.npmjs.com/package/@hecer/yoke)
16
+ [![CI](https://github.com/HECer/yoke/actions/workflows/ci.yml/badge.svg)](https://github.com/HECer/yoke/actions/workflows/ci.yml)
17
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](#-license)
18
+ ![Node](https://img.shields.io/badge/node-%E2%89%A520-339933?logo=node.js&logoColor=white)
19
+ ![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
20
+ ![Tests](https://img.shields.io/badge/tests-1019%20passing-brightgreen.svg)
21
+ ![Agents](https://img.shields.io/badge/agents-Claude%20%7C%20Codex%20%7C%20Gemini-8A2BE2)
22
+ ![Built with TDD](https://img.shields.io/badge/built%20with-TDD%20%2B%20review-ff69b4.svg)
23
+
24
+ **Install:** [`npm i -g @hecer/yoke`](https://www.npmjs.com/package/@hecer/yoke)
25
+
26
+ </div>
27
+
28
28
  > **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. Add `--parallel=N` for dependency-aware workers, or declare a reference and add `--quality` for a bounded critic/repair gauntlet. If any blocking gate is red, nothing is committed. Proof lives in `.yoke/proof/<story>/`.
29
29
 
30
30
  Yoke 1.5 keeps failed gate output compact without throwing evidence away: deterministic previews
@@ -33,592 +33,604 @@ in private, content-addressed local artifacts. Existing projects keep their seri
33
33
  safe 2 KiB preview / 8 KiB artifact defaults unless configured otherwise.
34
34
 
35
35
  Yoke 1.4 adds opt-in parallel workers and a bounded, reference-driven quality gauntlet without
36
- changing existing serial loop defaults. See [the 1.4 migration guide](docs/MIGRATING-TO-1.4.md)
37
- for the new flags, configuration, cleanup behavior, and review-verdict contract.
38
-
39
- Yoke 1.1 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
40
- is explicit; reviews require a schema-valid verdict and a different model unless
41
- `--allow-self-review` is explicit; commits enforce the human identity from project config or Git.
42
- See [the 1.1 migration guide](docs/MIGRATING-TO-1.1.md) for setup/decision parity and
43
- [the 1.0 guide](docs/MIGRATING-TO-1.0.md) for the earlier safety-policy changes.
44
-
45
- ---
46
-
47
- ## Why Yoke exists
48
-
49
- Agentic coding in 2026 fails in four well-documented ways. Yoke answers each one **mechanically** — in code, not in a prompt the agent can ignore:
50
-
51
- | The pain | What actually happens | What Yoke does about it |
52
- |---|---|---|
53
- | 🎭 **The verification gap** — *"agent says done, but it isn't"* | Agents submit confidently on 100% of runs while resolving far fewer; "all tests pass" when they were never run ([silent-failures research](https://arxiv.org/pdf/2603.25764)) | The loop trusts **your verify command's exit code**, never the agent's word. A story is `passes: true` only after tests are green, the reviewer approved, and the commit landed — atomically. Plus: **screenshot proofs** per story. |
54
- | 🔀 **Three agents, three configs** | Teams hand-maintain `CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, skills, and MCP wiring separately — copy-paste drift everywhere | **One canon → `yoke retrofit`** generates the idiomatic native artifacts for each agent. Change the canon once, re-retrofit everywhere. |
55
- | 🌀 **Overnight loops going off the rails** | Raw Ralph-loop users "wake up to broken codebases that don't compile" | Yoke is **"Ralph, but with gates"**: clean-worktree gate, acceptance-criteria gate, green-tests gate, review gate, per-story worktree isolation, idle-timeout watchdog, single-flight lock, commit integrity. |
56
- | 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a second model writes a schema-validated pass/fail verdict — chainable into verify, pre-push, or CI. Cross-model review catches what self-review misses. |
57
-
58
- **Who it's for:** anyone driving Claude Code, Codex CLI, or Gemini CLI on real projects — especially if you use more than one, want autonomous runs you can trust, or are tired of "done" meaning "probably". Greenfield (`yoke new`) and brownfield (`yoke retrofit`) both work.
59
-
60
- **Who it's not for:** if you want a chat pair-programmer with no process, you don't need a harness. Yoke is for shipping with discipline.
61
-
62
- ## ⏱️ 60 seconds: idea → tested, photographed software
63
-
64
- ```console
65
- $ yoke new reading-app --idea="a web app that tracks my reading list"
66
- ✓ reading-app bootstrapped. # git repo · harness for all agents · context · PRD drafted from the idea
67
-
68
- $ yoke prd check reading-app
69
- ✓ PRD valid — 8 stories, 0 pass
70
-
71
- $ yoke loop on reading-app
72
- $ yoke loop run reading-app --isolate --review --max=10
73
- ▶ STORY-1 (0/8 · 0%) — implementing… · verifying… · reviewing… · committing…
74
- ✓ STORY-1 done in 3m12s — 1/8 (13%) · ~22m left
75
- ▶ STORY-2 (1/8 · 13%) — implementing… · ~22m left (Ø 3m12s/story)
76
- ✓ STORY-2 done in 2m48s — 2/8 (25%) · ~18m left
77
- ▶ STORY-3 (2/8 · 25%) — implementing… ✘ blocked: story did not verify (tests red)
78
- # nothing was committed. fix, then re-run.
79
-
80
- $ ls reading-app/.yoke/proof/STORY-2/
81
- home.png list.png # photographic evidence, labelled per story
82
- ```
83
-
84
- Every claim in that transcript is enforced by code paths with tests behind them — 971 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
85
-
86
- ## 🚀 Quickstart
87
-
88
- ```bash
89
- npm install -g @hecer/yoke # → global `yoke` on your PATH
90
- # (or from source: git clone https://github.com/HECer/yoke.git && cd yoke && npm install && npm run build && npm link)
91
-
92
- # Greenfield: idea → loop-ready project in one command
93
- yoke new my-app --idea="a CLI that tracks reading lists"
94
- yoke loop on my-app && yoke loop run my-app --isolate
95
-
96
- # — or retrofit an existing project —
97
- yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
98
- yoke validate canon # sanity-check the canon
99
- yoke loop run /path/to/project --isolate --parallel=3 --reviewer=codex --max=20
100
- ```
101
-
102
- > Requires Node ≥ 20 and git. No global install? `node /path/to/yoke/dist/cli.js …` or `npm --prefix /path/to/yoke run yoke -- …` work too. The MCP tools (rtk, graphify/Serena, Playwright MCP) are wired by Yoke but installed separately — the generated config is a clearly-labelled, adjustable template.
103
-
104
- ### Skills before the first setup
105
-
106
- The canon is also packaged as a Claude Code plugin — the repo is its own marketplace:
107
-
108
- ```text
109
- /plugin marketplace add HECer/yoke
110
- /plugin install yoke@yoke
111
- ```
112
-
113
- That gives you all canon skills under the `yoke:` namespace (e.g. `yoke:tdd`, `yoke:review`) inside Claude Code — no retrofit needed. The `yoke` CLI (loop, gates, retrofit for Codex/Gemini) still comes from `npm i -g @hecer/yoke`. Gemini CLI users can likewise `gemini extensions install https://github.com/HECer/yoke`.
114
-
115
- For Codex, no preinstalled skill is required: run `npx @hecer/yoke setup .` in a terminal, or
116
- ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
117
- `.agents/skills/`, including `yoke-retrofit` and `yoke-workflow`; start a fresh Codex task if an
118
- already-open task does not discover newly installed skills. The npm package also contains
119
- `.codex-plugin/plugin.json` for Codex plugin hosts.
120
-
121
- ### Staying up to date
122
-
123
- Yoke checks for new releases npm/gh-style: a **non-blocking background check** (at most once a day, detached, offline-safe) prints a one-line hint when a newer version exists — upgrading itself is always an explicit act:
124
-
125
- ```bash
126
- yoke upgrade # npm install -g @hecer/yoke@latest
127
- ```
128
-
129
- Disable the check with `YOKE_NO_UPDATE_CHECK=1` (it is also silent in CI, `--json` runs, and piped output). Projects that want the loop to self-update can opt in via `.yoke/config.yaml`:
130
-
131
- ```yaml
132
- update:
133
- auto: true # upgrade at loop START only — never mid-run; applies from the next invocation
134
- ```
135
-
136
- Auto-upgrade is deliberately **not** the default: a gate harness shouldn't change itself mid-project, and unreviewed auto-installs are a supply-chain hazard.
137
-
138
- ## 🤖 Driving it through an agent
139
-
140
- Yoke is meant to be operated *by* your coding agent — after a retrofit, the agent has the skills, the safety policy, and the routing, so it knows the methodology. Copy-paste prompts (identical wording works for Claude Code, Codex CLI, and Gemini CLI):
141
-
142
- > **Set it up** — *"Set up Yoke in this project. Ask me the Yoke setup questions one at a time with your recommendation, then run `yoke setup . --yes` with the selected host, agents, code graph, loop, runner, and decision policy. Commit in my configured identity."*
143
-
144
- > **Work the disciplined way** — *"From now on follow the Yoke skills you just installed: brainstorm → spec → plan → TDD → review before merging. Use the `review` skill before any merge."*
145
-
146
- > **Plan, then run autonomously** — *"Use the `yoke-workflow` skill. Ask only the planning questions that materially change the product, write the approved plan and loop-ready stories, then execute every approved story without routine follow-ups. Follow the configured `auto` or `critical` decision policy."*
147
-
148
- > **Watch / unblock** — *"Run `yoke loop status .`. If it says BLOCKED, run the project's verify command, find the root cause, fix it without weakening tests, then continue the loop."*
149
-
150
- > ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
151
-
152
- > ⚠️ **Never kill agent processes by name or command-line pattern** (e.g. every process matching `dangerously-skip-permissions`): on a machine running several yoke projects, that takes down the *healthy* runners of the other projects mid-story — they stall and their loops block. `yoke loop cleanup` is the scoped alternative: each watchdog records its pids in the project's `.yoke/runner.pid`, and cleanup kills exactly those recorded trees — nothing else on the machine.
153
-
154
- ### Agent cheat sheet — every command is an exit-code contract
155
-
156
- Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`) can branch on exit codes without parsing prose.
157
-
158
- | Command | What it does | Exit codes |
159
- |---|---|---|
160
- | `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
161
- | `yoke validate [canonDir]` | Validate the canon (schema, frontmatter, templates) | `0` valid · `1` errors |
162
- | `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
163
- | `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
164
- | `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
165
- | `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
166
- | `yoke change add\|status [dir] [--idea=]` | Queue a change at any time; the loop turns it into append-only stories at the next safe boundary | `0` · `1` invalid inbox/request |
167
- | `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
168
- | `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` supports `--parallel=N`, bounded reference-driven `--quality`, and blind `--candidates=N` selection; `--max=N` creates an intentional batch cap; `cleanup` retains worktrees unless `--remove-worktrees` is explicit | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
169
- | `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
170
- | `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
171
- | `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
172
- | `yoke flow-smoke [dir] [--url=] [--label=]` | Browser gate with screenshot/video proofs | `0` green · `1` failures · `2` not runnable |
173
-
174
- A genuinely hung agent self-terminates after the idle timeout (default 20 min; `--timeout`), and `yoke loop status` shows the live phase or a `⚠ possibly stuck` hint — an autonomous run is never a black box.
175
-
176
- ## ⚖️ How it compares — superpowers · gstack · Yoke
177
-
178
- Three excellent projects, three different jobs. Honest version:
179
-
180
- | | [superpowers](https://github.com/obra/superpowers) (obra) | [gstack](https://github.com/garrytan/gstack) (Garry Tan) | **Yoke** |
181
- |---|---|---|---|
182
- | **What it is** | The canonical *skills methodology*: brainstorm → plan → TDD → review as composable skills | A *software factory* for Claude Code: ~40 role skills (QA, CSO, ship…) + a real Chromium browser layer | A *cross-agent harness*: one canon → native installs, plus a gated autonomous loop |
183
- | **Agents** | Claude Code first | Claude Code + hosts like Codex/Cursor/Kiro — **no Gemini CLI** | **Claude Code, Codex CLI, Gemini CLI** from one source of truth |
184
- | **Enforcement** | Advisory — skills *describe* the discipline; following them is up to the agent | Skill-driven; browser QA is genuinely real | **Mechanical** — gates live in code: clean tree, acceptance criteria, green tests, review verdict, commit integrity |
185
- | **Autonomy** | Interactive sessions | Interactive slash-commands (`/qa`, `/ship`, …) | Opt-in **Ralph loop** with watchdog, worktree isolation, single-flight lock, per-story proofs |
186
- | **Visual QA** | — | **Best-in-class**: live browser daemon (Chromium/CDP) with deep interactive QA | Built-in `flow-smoke` gate: screenshots always, video on failure, labelled per story — lighter, but *enforced* and cross-agent |
187
- | **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves an independent provider and validates a structured verdict, inside or outside the loop |
188
- | **Footprint** | Markdown skills (plugin) | ~230 MB with browser runtime; hourly auto-update | Node CLI + markdown canon; Playwright only if you use flow-smoke, resolved **from your project** |
189
- | **License** | MIT | MIT | MIT |
190
-
191
- **They compose — use all three where they're strongest.** Yoke's canon *ships* the superpowers methodology natively for all three agents (13 skills, [attributed](canon/skills/ATTRIBUTION.md)). And if gstack is installed, `yoke retrofit` detects it and adds a routing note to `CLAUDE.md` telling Claude to prefer gstack's live-browser `/qa`, `/cso`, and ship pipeline for what Yoke deliberately doesn't bundle — no dependency, no conflict, and Codex/Gemini artifacts stay uniform.
192
-
193
- **Choose Yoke when** you run more than one agent, want autonomy you can audit (gates + proofs + logs), or want one place to maintain your team's methodology. **Choose gstack when** you live 100% in Claude Code and want the deepest interactive browser QA. **Choose superpowers when** you want the methodology alone, interactively, in Claude Code — or just use it *through* Yoke.
194
-
195
- ## 🏗️ Architecture
196
-
197
- You curate **one source of truth** — skills, policy, and tool wiring. Yoke generates the **idiomatic, native artifacts** each agent expects, non-destructively, into any repo:
198
-
199
- ```mermaid
200
- flowchart TD
201
- Canon["📦 CANON — single source of truth<br/>skills · policy · loop spec · tool wiring"]
202
- Skill["🛠️ yoke retrofit<br/>detect → plan → apply (backup) → report"]
203
- Canon --> Skill
204
- Skill --> Claude["Claude Code<br/>.claude/skills · .mcp.json · hook"]
205
- Skill --> Codex["Codex CLI<br/>AGENTS.md · config.toml · RTK.md"]
206
- Skill --> Gemini["Gemini CLI<br/>GEMINI.md · commands · settings.json"]
207
- Loop["🤖 yoke loop — autonomous Ralph loop<br/>gates · verify · review · isolation · proofs"]
208
- Claude -. drives .-> Loop
209
- Codex -. drives .-> Loop
210
- Gemini -. drives .-> Loop
211
- ```
212
-
213
- Three layers — **Canon** (`yoke validate`) → **Retrofit** (`yoke retrofit`) → **Loop** (`yoke loop`) — on top of a durable **Context layer** (`yoke context`).
214
-
215
- ### What gets generated per agent
216
-
217
- | Agent | Artifacts |
218
- |---|---|
219
- | **Claude** | `.claude/skills/`, `AGENTS.md`, `CLAUDE.md`, `.mcp.json` (code-graph + Playwright), and an rtk `PreToolUse` hook when WSL is available |
220
- | **Codex** | `.agents/skills/`, `AGENTS.md`, `RTK.md`, `.codex/config.toml`, native hooks, reusable `.codex/agents/*.toml`, and package plugin metadata |
221
- | **Gemini** | `GEMINI.md`, `.gemini/commands/*.toml` (one per skill, full body), `.gemini/settings.json` (MCP + `AGENTS.md` context) |
222
-
223
- > **rtk integration:** Claude receives its PreToolUse hook; Codex receives a native hook adapter around `rtk hook check`; Gemini retains instruction-mode fallback where its CLI has no equivalent command-rewrite lifecycle.
224
-
225
- > **Composes with gstack:** if [gstack](https://github.com/garrytan/gstack) is installed (repo-local or global), `yoke retrofit` adds a short "Composed tools" routing note to **CLAUDE.md only** — telling Claude to prefer gstack's skills for capabilities Yoke doesn't ship (live-browser QA `/qa`, security audit `/cso`, ship/deploy `/ship`). No bundling, no dependency; the note is never written to the Codex or Gemini artifacts.
226
-
227
- > **Your content survives re-retrofits — preserve blocks:** anything you put between
228
- > `<!-- yoke:preserve:start -->` and `<!-- yoke:preserve:end -->` in a generated file is
229
- > carried into the regenerated version on every future `yoke retrofit`. The generated
230
- > `CLAUDE.md` and `GEMINI.md` ship an empty preserve block scaffold — put your project-specific
231
- > instructions (tech stack, workflow, `@`-includes) inside it. Works in any yoke-written file;
232
- > content *outside* the markers is still replaced (and backed up under `.yoke/backup/`).
233
-
234
- ## 🧰 What's in the canon — 29 skills
235
-
236
- `yoke retrofit` installs all of these into each agent natively. Provenance is credited in [`canon/skills/ATTRIBUTION.md`](canon/skills/ATTRIBUTION.md).
237
-
238
- To stop overlapping skills from auto-invoking against each other, `canon/AGENTS.md` carries a **skill routing & precedence** block (methodology before role; one canonical entrypoint per concern — e.g. pre-merge code review is always `review`), emitted into all three agents.
239
-
240
- **Process / methodology** — *superpowers-derived discipline (13)*
241
-
242
- | Skill | What it does |
243
- |---|---|
244
- | `brainstorming` | Explore intent, requirements & design before any creative work |
245
- | `writing-plans` | Turn a spec into a bite-sized, TDD implementation plan |
246
- | `executing-plans` | Execute a written plan in a separate session with review checkpoints |
247
- | `subagent-driven-development` | Run a plan task-by-task: fresh subagent + two-stage review each |
248
- | `tdd` | Write the test first, watch it fail, write minimal code, refactor |
249
- | `systematic-debugging` | Root-cause first — no fix without a confirmed cause |
250
- | `verification-before-completion` | Prove it actually works before claiming done |
251
- | `using-git-worktrees` | Isolated worktrees for safe / parallel work |
252
- | `requesting-code-review` | Request a structured review before merging |
253
- | `receiving-code-review` | Handle review feedback with rigor, not blind agreement |
254
- | `dispatching-parallel-agents` | Fan out 2+ independent tasks concurrently |
255
- | `finishing-a-development-branch` | Merge / PR / cleanup a finished branch |
256
- | `writing-skills` | Author and verify new skills |
257
-
258
- **Roles** — *gstack-derived, de-gstacked to be harness-agnostic (7)*
259
-
260
- | Skill | What it does |
261
- |---|---|
262
- | `plan-eng-review` | Architecture / edge-case review of a *plan* |
263
- | `plan-ceo-review` | Founder-mode scope & ambition review of a plan |
264
- | `review` | Single canonical pre-merge code review — diff safety + engineering quality (architecture, edge cases, tests, performance) |
265
- | `ship` | Ship workflow: tests → review → version → changelog → PR |
266
- | `health` | Code-quality dashboard with a composite score |
267
- | `retro` | Engineering retrospective from commit history |
268
- | `document-release` | Post-ship documentation sync (README / CHANGELOG / …) |
269
-
270
- **Yoke-native** — *authored or adapted for this harness (9)*
271
-
272
- | Skill | What it does |
273
- |---|---|
274
- | `yoke-retrofit` | Set up the Yoke harness in a project (detect → plan → apply) |
275
- | `yoke-workflow` | Provider-neutral planning questions → approved PRD → autonomous stories → critical-decision resume |
276
- | `authoring-prd` | Slice a product idea into loop-ready stories with testable acceptance criteria |
277
- | `minimal-code` | Write the least code that solves the task (YAGNI; ponytail-derived) |
278
- | `performance` | Efficiency as a measured requirement: benchmarks as tests, budgets as gates, optimizations local + documented |
279
- | `maintaining-context` | Keep `.yoke/context/` the durable source of truth (the Context layer) |
280
- | `workflow` | The default order of operations, from idea to deploy |
281
- | `unslop-ui` | Detect & remove AI-slop design tells (purple gradients, neon glow, emoji-icons…) |
282
- | `visual-verification` | Widen verify to design-scan + the built-in `yoke flow-smoke` gate (screenshot proofs; video on failure) |
283
-
284
- ## 🌱 Zero to 100: `yoke new` + `yoke prd`
285
-
286
- Yoke's greenfield entrypoint — one command from idea to loop-ready project:
287
-
288
- ```bash
289
- yoke new my-app --idea="a CLI that tracks reading lists" # scaffold + retrofit + context + PRD
290
- yoke loop on my-app && yoke loop run my-app --isolate # hand it to the loop
291
- ```
292
-
293
- `yoke new <dir>` refuses a non-empty directory (greenfield-only — use `yoke retrofit` for
294
- existing projects), then: creates and `git init`s the directory, writes a minimal scaffold
295
- (`README.md`, `.gitignore`), runs the full **retrofit** (`--agent=` as usual), initialises the
296
- **context layer** (with `--idea` seeded into `PROJECT.md` as the north star), writes a commented
297
- **PRD template** to `.yoke/prd.yaml`, and makes the initial commit — so `--isolate` works from
298
- iteration 1. With `--idea`, it then drafts the PRD from your idea via an agent (`--runner=`,
299
- the configured runner or active host) and commits it as a second commit (`docs: draft PRD from idea`).
300
-
301
- - **Exit codes** — `0` success; `1` usage / non-empty dir / draft failure (the scaffold survives —
302
- retry with `yoke prd draft`); `2` requested draft agent unavailable.
303
-
304
- **`yoke prd draft [dir] --idea="..."`** turns an idea into 5–12 small, independently shippable
305
- stories with testable behavioral acceptance criteria (greenfield STORY-1 scaffolds the project
306
- skeleton + test suite and wires `verify.command`). An existing PRD with stories is never
307
- overwritten without `--force`; the untouched template doesn't trigger the guard. Runs through
308
- the same idle-timeout watchdog as the loop (`--timeout`). If `.yoke/plan.md` exists, its approved
309
- goals, non-goals, constraints, and decisions are injected as settled context instead of being
310
- reopened by the drafting agent.
311
-
312
- **`yoke prd check [dir]`** is the chainable pre-loop lint gate: schema validation plus
313
- duplicate-id, empty-acceptance, unresolved-placeholder, and zero-stories checks. Exits `0` with
314
- `✓ PRD valid — N stories, M pass`, `1` on any violation. The `authoring-prd` canon skill
315
- teaches interactive sessions the same story-slicing discipline.
316
-
317
- ## 🤖 The autonomous loop
318
-
319
- Opt-in; `yoke setup` recommends enabling it for new installs, while `retrofit` alone keeps it off unless requested. Each iteration starts a **fresh agent** and passes through hard gates before anything is committed:
320
-
321
- ```mermaid
322
- flowchart LR
323
- I[consume queued change<br/>as new stories] --> A[pick next PRD story]
324
- A --> B{clean worktree?}
325
- B -- no --> X[blocked]
326
- B -- yes --> C{acceptance<br/>criteria?}
327
- C -- no --> X
328
- C -- yes --> D[agent implements<br/>one story]
329
- D --> E{suite + criterion<br/>proof green?}
330
- E -- no --> X
331
- E -- yes --> F{reviewer<br/>approves?}
332
- F -- no --> X
333
- F -- yes --> G[commit + mark passes:true<br/>+ proof in .yoke/proof/]
334
- G --> I
335
- I --> H{all stories pass?}
336
- H -- yes --> J{integrated system<br/>gate green?}
337
- J -- no --> X
338
- J -- yes --> K[current backlog ready]
339
- ```
340
-
341
- ```bash
342
- yoke loop on . # enable (recorded in .yoke/config.yaml)
343
- yoke loop status . # show state + PRD progress
344
- yoke change add . --idea="Add passkey login" # safe while the loop runs
345
- yoke loop run . \
346
- --runner=codex \ # implement with Codex…
347
- --reviewer=claude \ # …review with Claude (role separation)
348
- --isolate \ # each story in a throwaway git worktree
349
- --parallel=3 \ # run dependency-ready, non-colliding stories concurrently
350
- --decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
351
- # Optional: add --max=20 only when this run should stop after a bounded batch.
352
- yoke loop off . # disable
353
- ```
354
-
355
- **PRD format** (`.yoke/prd.yaml`):
356
-
357
- ```yaml
358
- - id: STORY-1
359
- title: Add a health endpoint
360
- priority: 1 # lower = higher priority
361
- acceptance: # Definition of Done (required, else blocked)
362
- - id: health-returns-200
363
- text: GET /health returns 200
364
- verify: [npm run test:health-returns-200]
365
- - id: health-rejects-post
366
- text: POST /health returns 405
367
- verify: [npm run test:health-rejects-post]
368
- passes: false # the loop sets this true only on green tests
369
- ```
370
-
371
- New projects default to `verify.requireCriteria: true`: every story has 2–5 behavioral criteria.
372
- Each criterion ID must occur in its single, approved test command; shell operators and broad,
373
- untargeted suites are rejected. Yoke records each result in `.yoke/proof/<story>/evidence.json`. Configure optional
374
- `completion.command` for integrated journeys such as purchase → entitlement → relaunch or
375
- magic-link → callback → authenticated app. It runs whenever the current backlog has no open
376
- stories; this is readiness, not a release.
377
-
378
- No generic tool can infer whether arbitrary test code perfectly represents product meaning. Yoke
379
- closes the mechanical false-done paths—targeted evidence, coverage review, clean committed state,
380
- and integrated journeys—while the project still owns the correctness of its tests and production
381
- observability.
382
-
383
- ### Parallel workers and the quality gauntlet
384
-
385
- `--parallel=N` dispatches dependency-ready stories concurrently. Claims carry leases, workers use
386
- isolated worktrees, collision areas are serialized, and only a mechanically green candidate enters
387
- the FIFO integration queue. Integration repeats the project gates against the merged tree; a worker
388
- success can never bypass a red integrated result. `yoke loop status` reports the dispatcher,
389
- workers, providers, worktrees, lifecycle, queue, integrations, and reopened stories.
390
-
391
- Quality is reference-driven and opt-in. Declare what one story should match:
392
-
393
- ```yaml
394
- quality:
395
- reference: { name: approved-home, source: design/home.png, kind: file }
396
- candidate: { kind: screenshots, paths: [.yoke/proof/STORY-1/home.png] }
397
- rubric: Match the approved layout, hierarchy, spacing, and states.
398
- policy: blocking # or advisory
399
- ```
400
-
401
- Configure project defaults, then enable the gauntlet for a run:
402
-
403
- ```yaml
404
- quality:
405
- enabled: false # keep opt-in, or make it the project default
406
- policy: blocking
407
- maxRounds: 3
408
- maxMinutes: 60
409
- consistencyChecks: 2
410
- maxParallelCandidates: 2
411
- critic: { agent: codex, model: gpt-5.6-sol } # model required for --candidates
412
- repair: { agent: claude }
413
- ```
414
-
415
- ```bash
416
- yoke loop run . --quality --quality-rounds=3 --quality-minutes=60
417
- yoke loop run . --quality --candidates=2 # blind pairwise selection; stories need quality declarations
418
- ```
419
-
420
- The critic compares opaque candidate/reference labels, writes schema-validated provenance, and
421
- cannot modify the project. Blocking findings enter a bounded repair loop and rerun every mechanical
422
- gate; advisory findings are retained without blocking. `--quality-policy=`, `--no-quality`, and
423
- `--quality-unbounded` override defaults for one run. Unbounded mode is explicit and warned because
424
- it removes repair limits, not Yoke's watchdog, isolation, verification, or commit safety.
425
-
426
- State lives **outside the model context** — the PRD file plus git — so each iteration is fresh.
427
- Use `yoke change add` at any time. Its ignored append-only inbox is consumed at the next story
428
- boundary. A separate coverage pass must confirm that every requested outcome maps to behavioral
429
- criteria before Yoke appends and commits the new stories; existing stories are never rewritten and
430
- no restart is needed.
431
-
432
- ### Watching a run
433
-
434
- Every iteration emits token-free, harness-side feedback (Node console + local files — **zero agent tokens**):
435
-
436
- - **Live console with progress + ETA** —
437
- `▶ S6 (19/45 · 42%) — implementing… · ~1h44m left (Ø 4m/story)` … `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`.
438
- The estimate uses the **average duration of stories completed in this run** (current
439
- velocity); before the first story lands it falls back to the recorded history of previous
440
- runs (`.yoke/story-durations.json`, last 50 stories, gitignored). No data yet → no estimate,
441
- never a made-up one.
442
- - **`.yoke/loop-status.json`** — the current state (now including `percent` and an `eta`
443
- block); read it any time with `yoke loop status`:
444
- ```
445
- Loop: RUNNING on S6 "Weekly digest"
446
- implementing · iteration 20 · 19/45 (42%) · updated 30s ago
447
- ~1h44m remaining (Ø 4m/story)
448
- ```
449
- - **Parallel + quality detail** — active workers include provider, candidate ID, worktree,
450
- lifecycle, phase, quality round, and repair budget; the integrator is shown separately.
451
- - **`.yoke/loop.log`** — an append-only timeline of every phase transition.
452
- - **`--json`** — machine mode for supervisors: every status write is *also* emitted as one
453
- NDJSON line on stdout (`{"type":"status","state":"running","phase":"verifying",…}` — the
454
- same shape as `loop-status.json`), the human narrative moves off stdout (the final summary
455
- goes to stderr), and a consumer can follow the stream line by line instead of polling the file.
456
- Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
457
- output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
458
- and cost fields when available. Missing values stay absent—Yoke does not estimate them.
459
-
460
- ### Pausing a run
461
-
462
- Drop a **`.yoke/loop.pause`** file (contents irrelevant) while the loop is running and it
463
- stops at the **next story boundary** — the running story still finishes, verifies, and
464
- commits; no story is ever cut off mid-flight. The loop consumes the pause file, writes
465
- `state: "paused"` to `loop-status.json` (log label `paused`), releases the lock, and exits
466
- with code `3`. Resume by simply running `yoke loop run` again.
467
-
468
- A per-iteration **idle timeout** guards against a genuinely hung agent: if the agent produces
469
- **no output at all** for `--timeout` minutes (default 20; `0` disables), the loop kills it
470
- (SIGTERM→SIGKILL) and marks the story blocked. A slow-but-working agent that keeps streaming
471
- output is **never** killed — the output stream *is* the liveness signal. Set a project default
472
- with `loop.timeoutMinutes` in `.yoke/config.yaml`.
473
-
474
- ### Decision policy: autonomous by default, interrupt only when configured
475
-
476
- Planning questions happen before the loop. The provider-neutral `yoke-workflow` skill asks only
477
- questions whose answer materially changes product behavior, scope, architecture, security, data
478
- ownership, external cost, or an irreversible choice. It saves the approved brief in
479
- `.yoke/plan.md`; `yoke prd draft` consumes it, and `yoke prd check` rejects explicit unresolved
480
- placeholders such as `TBD`.
481
-
482
- The unattended loop then follows `loop.decisionPolicy`:
483
-
484
- ```yaml
485
- loop:
486
- enabled: true
487
- decisionPolicy: critical # or auto
488
- runner:
489
- agent: codex # setup chooses the current host by default
490
- ```
491
-
492
- - **`auto` (default):** routine ambiguity and implementation details are resolved using the
493
- approved plan, acceptance criteria, current code, and project conventions. The loop does not
494
- ask follow-up questions.
495
- - **`critical`:** routine choices are still resolved automatically. Only high-impact decisions
496
- involving public architecture, security/privacy, destructive migration or data loss, material
497
- external cost, legal/compliance exposure, or another irreversible choice may pause the story.
498
- The agent writes a schema-validated request; the loop blocks before verify and preserves it as
499
- `.yoke/pending-decision.yaml`.
500
-
501
- Inspect and answer a critical stop:
502
-
503
- ```bash
504
- yoke loop decision .
505
- yoke loop answer . --choice=A --rationale="Matches the existing identity model"
506
- ```
507
-
508
- `answer` validates the choice against the still-open story, appends it to
509
- `.yoke/context/DECISIONS.md`, commits only that file using the configured human identity, clears
510
- the pending request, and resumes the same story with the original runner, isolation, review,
511
- permission, timeout, JSON, decision-policy, and iteration settings intact. Add
512
- `--no-resume` when a supervisor should restart the loop separately. If the automatic restart
513
- cannot begin because a provider/reviewer is unavailable or another process owns the lock, run
514
- `yoke loop resume .`; its request-bound options are retained under Git's private state directory
515
- until a loop actually runs. To intentionally abandon an orphaned or stale private resume state,
516
- use `yoke loop resume . --discard`; pending decisions are never deleted by that command. Existing
517
- `loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` remain supported as compatibility aliases;
518
- new projects should use `decisionPolicy: auto|critical`.
519
-
520
- ### Adaptive model routing (explicit opt-in)
521
-
522
- `yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
523
- parent remains the strong planner/controller. Before each bounded story it receives only the
524
- story, acceptance criteria, and at most three eligible worker profiles, then returns one
525
- machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
526
- `SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
527
- runs so Yoke does not pay for two orchestration layers.
528
-
529
- **Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
530
- Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
531
- cover invocation and routing behavior for all three providers. The measured performance evidence
532
- below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
533
- until authenticated, repeated in-the-wild runs exist for those providers.
534
-
535
- ```yaml
536
- runner:
537
- agent: codex
538
- model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
539
- reasoningEffort: high
540
- routing:
541
- enabled: true # setup defaults false; setup --routing opts in
542
- strategy: balanced # balanced | cost | speed | quality
543
- maxCandidates: 3
544
- workers:
545
- - id: codex-light
546
- agent: codex
547
- reasoningEffort: low
548
- costTier: medium
549
- capabilities: [exploration, implementation, tests]
550
- - id: claude-fast
551
- agent: claude
552
- model: haiku # rolling alias; omit to use the provider's current default
553
- reasoningEffort: low
554
- costTier: low
555
- capabilities: [mechanical-edits, tests]
556
- - id: gemini-auto
557
- agent: gemini # omitted model means the account's current Auto/default route
558
- costTier: low
559
- capabilities: [large-context, implementation]
560
- ```
561
-
562
- Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
563
- baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
564
- eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
565
- "intelligence score". Candidate model IDs come from project configuration while setup defaults
566
- prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
567
- independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
568
- after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
569
- time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
570
- cannot overwrite a shared registry file.
571
-
572
- Routing is not free: it adds one controller call per story. It is most promising when a bounded
573
- worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
574
- on your own backlog rather than assuming a win.
575
-
576
- ### Performance budgets: efficiency as a gate, not a style
577
-
578
- Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
579
- be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
580
-
581
- - **Per story:** write the requirement as a **measurable acceptance criterion**
582
- ("imports 1M rows in < 2s, asserted by the bench test") and let your verify tests measure
583
- it — no new machinery needed.
584
- - **Per project:** wire a benchmark as a standing **perf gate** in `.yoke/config.yaml`:
585
-
586
- ```yaml
587
- perf:
588
- command: node bench/check-budget.mjs # exit 0 = within budget
589
- retries: 1 # benchmarks are noisy; same retry logic as verify
590
- ```
591
-
592
- The loop runs it **after verify** on every story (phase `perf`, with `YOKE_STORY` set); a
593
- red benchmark blocks the story — `story S6 exceeded its performance budget: p95 62ms > budget 50ms` —
594
- no matter how clean the diff was. The implementer prompt names the budget command, so the
595
- agent knows not to trade hot-path efficiency for style and never "simplifies away" an
596
- optimization without re-running the benchmark. The `performance` canon skill carries the
597
- method: profile first, optimize leaves not boundaries, commit benchmarks as tests, version
598
- the *why* of every optimization in `context/DECISIONS.md`.
599
-
600
- ### Artifact-backed gate output: compact context, complete local evidence
601
-
36
+ changing existing serial loop defaults. See [the 1.4 migration guide](docs/MIGRATING-TO-1.4.md)
37
+ for the new flags, configuration, cleanup behavior, and review-verdict contract.
38
+
39
+ Yoke 1.1 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
40
+ is explicit; reviews require a schema-valid verdict and a different model unless
41
+ `--allow-self-review` is explicit; commits enforce the human identity from project config or Git.
42
+ See [the 1.1 migration guide](docs/MIGRATING-TO-1.1.md) for setup/decision parity and
43
+ [the 1.0 guide](docs/MIGRATING-TO-1.0.md) for the earlier safety-policy changes.
44
+
45
+ ---
46
+
47
+ ## Why Yoke exists
48
+
49
+ Agentic coding in 2026 fails in four well-documented ways. Yoke answers each one **mechanically** — in code, not in a prompt the agent can ignore:
50
+
51
+ | The pain | What actually happens | What Yoke does about it |
52
+ |---|---|---|
53
+ | 🎭 **The verification gap** — *"agent says done, but it isn't"* | Agents submit confidently on 100% of runs while resolving far fewer; "all tests pass" when they were never run ([silent-failures research](https://arxiv.org/pdf/2603.25764)) | The loop trusts **your verify command's exit code**, never the agent's word. A story is `passes: true` only after tests are green, the reviewer approved, and the commit landed — atomically. Plus: **screenshot proofs** per story. |
54
+ | 🔀 **Three agents, three configs** | Teams hand-maintain `CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, skills, and MCP wiring separately — copy-paste drift everywhere | **One canon → `yoke retrofit`** generates the idiomatic native artifacts for each agent. Change the canon once, re-retrofit everywhere. |
55
+ | 🌀 **Overnight loops going off the rails** | Raw Ralph-loop users "wake up to broken codebases that don't compile" | Yoke is **"Ralph, but with gates"**: clean-worktree gate, acceptance-criteria gate, green-tests gate, review gate, per-story worktree isolation, idle-timeout watchdog, single-flight lock, commit integrity. |
56
+ | 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a second model writes a schema-validated pass/fail verdict — chainable into verify, pre-push, or CI. Cross-model review catches what self-review misses. |
57
+
58
+ **Who it's for:** anyone driving Claude Code, Codex CLI, or Gemini CLI on real projects — especially if you use more than one, want autonomous runs you can trust, or are tired of "done" meaning "probably". Greenfield (`yoke new`) and brownfield (`yoke retrofit`) both work.
59
+
60
+ **Who it's not for:** if you want a chat pair-programmer with no process, you don't need a harness. Yoke is for shipping with discipline.
61
+
62
+ ## ⏱️ 60 seconds: idea → tested, photographed software
63
+
64
+ ```console
65
+ $ yoke new reading-app --idea="a web app that tracks my reading list"
66
+ ✓ reading-app bootstrapped. # git repo · harness for all agents · context · PRD drafted from the idea
67
+
68
+ $ yoke prd check reading-app
69
+ ✓ PRD valid — 8 stories, 0 pass
70
+
71
+ $ yoke loop on reading-app
72
+ $ yoke loop run reading-app --isolate --review --max=10
73
+ ▶ STORY-1 (0/8 · 0%) — implementing… · verifying… · reviewing… · committing…
74
+ ✓ STORY-1 done in 3m12s — 1/8 (13%) · ~22m left
75
+ ▶ STORY-2 (1/8 · 13%) — implementing… · ~22m left (Ø 3m12s/story)
76
+ ✓ STORY-2 done in 2m48s — 2/8 (25%) · ~18m left
77
+ ▶ STORY-3 (2/8 · 25%) — implementing… ✘ blocked: story did not verify (tests red)
78
+ # nothing was committed. fix, then re-run.
79
+
80
+ $ ls reading-app/.yoke/proof/STORY-2/
81
+ home.png list.png # photographic evidence, labelled per story
82
+ ```
83
+
84
+ Every claim in that transcript is enforced by code paths with tests behind them — 1019 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
85
+
86
+ ## 🚀 Quickstart
87
+
88
+ ```bash
89
+ npm install -g @hecer/yoke # → global `yoke` on your PATH
90
+ # (or from source: git clone https://github.com/HECer/yoke.git && cd yoke && npm install && npm run build && npm link)
91
+
92
+ # Greenfield: idea → loop-ready project in one command
93
+ yoke new my-app --idea="a CLI that tracks reading lists"
94
+ yoke loop on my-app && yoke loop run my-app --isolate
95
+
96
+ # — or retrofit an existing project —
97
+ yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
98
+ yoke validate canon # sanity-check the canon
99
+ yoke loop run /path/to/project --isolate --parallel=3 --reviewer=codex --max=20
100
+ ```
101
+
102
+ > Requires Node ≥ 20 and git. No global install? `node /path/to/yoke/dist/cli.js …` or `npm --prefix /path/to/yoke run yoke -- …` work too. The MCP tools (rtk, graphify/Serena, Playwright MCP) are wired by Yoke but installed separately — the generated config is a clearly-labelled, adjustable template.
103
+
104
+ ### Skills before the first setup
105
+
106
+ The canon is also packaged as a Claude Code plugin — the repo is its own marketplace:
107
+
108
+ ```text
109
+ /plugin marketplace add HECer/yoke
110
+ /plugin install yoke@yoke
111
+ ```
112
+
113
+ That gives you all canon skills under the `yoke:` namespace (e.g. `yoke:tdd`, `yoke:review`) inside Claude Code — no retrofit needed. The `yoke` CLI (loop, gates, retrofit for Codex/Gemini) still comes from `npm i -g @hecer/yoke`. Gemini CLI users can likewise `gemini extensions install https://github.com/HECer/yoke`.
114
+
115
+ For Codex, no preinstalled skill is required: run `npx @hecer/yoke setup .` in a terminal, or
116
+ ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
117
+ `.agents/skills/`, including `yoke-retrofit` and `yoke-workflow`; start a fresh Codex task if an
118
+ already-open task does not discover newly installed skills. The npm package also contains
119
+ `.codex-plugin/plugin.json` for Codex plugin hosts.
120
+
121
+ ### Staying up to date
122
+
123
+ Yoke checks for new releases npm/gh-style: a **non-blocking background check** (at most once a day, detached, offline-safe) prints a one-line hint when a newer version exists — upgrading itself is always an explicit act:
124
+
125
+ ```bash
126
+ yoke upgrade # npm install -g @hecer/yoke@latest
127
+ ```
128
+
129
+ Disable the check with `YOKE_NO_UPDATE_CHECK=1` (it is also silent in CI, `--json` runs, and piped output). Projects that want the loop to self-update can opt in via `.yoke/config.yaml`:
130
+
131
+ ```yaml
132
+ update:
133
+ auto: true # upgrade at loop START only — never mid-run; applies from the next invocation
134
+ ```
135
+
136
+ Auto-upgrade is deliberately **not** the default: a gate harness shouldn't change itself mid-project, and unreviewed auto-installs are a supply-chain hazard.
137
+
138
+ ## 🤖 Driving it through an agent
139
+
140
+ Yoke is meant to be operated *by* your coding agent — after a retrofit, the agent has the skills, the safety policy, and the routing, so it knows the methodology. Copy-paste prompts (identical wording works for Claude Code, Codex CLI, and Gemini CLI):
141
+
142
+ > **Set it up** — *"Set up Yoke in this project. Ask me the Yoke setup questions one at a time with your recommendation, then run `yoke setup . --yes` with the selected host, agents, code graph, loop, runner, and decision policy. Commit in my configured identity."*
143
+
144
+ > **Work the disciplined way** — *"From now on follow the Yoke skills you just installed: brainstorm → spec → plan → TDD → review before merging. Use the `review` skill before any merge."*
145
+
146
+ > **Plan, then run autonomously** — *"Use the `yoke-workflow` skill. Ask only the planning questions that materially change the product, write the approved plan and loop-ready stories, then execute every approved story without routine follow-ups. Follow the configured `auto` or `critical` decision policy."*
147
+
148
+ > **Watch / unblock** — *"Run `yoke loop status .`. If it says BLOCKED, run the project's verify command, find the root cause, fix it without weakening tests, then continue the loop."*
149
+
150
+ > ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
151
+
152
+ > ⚠️ **Never kill agent processes by name or command-line pattern** (e.g. every process matching `dangerously-skip-permissions`): on a machine running several yoke projects, that takes down the *healthy* runners of the other projects mid-story — they stall and their loops block. `yoke loop cleanup` is the scoped alternative: each watchdog records its pids in the project's `.yoke/runner.pid`, and cleanup kills exactly those recorded trees — nothing else on the machine.
153
+
154
+ ### Agent cheat sheet — every command is an exit-code contract
155
+
156
+ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`) can branch on exit codes without parsing prose.
157
+
158
+ | Command | What it does | Exit codes |
159
+ |---|---|---|
160
+ | `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing]` | Shared six-question setup for Claude, Codex, and Gemini; adaptive routing is always an explicit opt-in | `0` · `1` invalid setup |
161
+ | `yoke validate [canonDir]` | Validate the canon (schema, frontmatter, templates) | `0` valid · `1` errors |
162
+ | `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
163
+ | `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
164
+ | `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
165
+ | `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
166
+ | `yoke change add\|status [dir] [--idea=]` | Queue a change at any time; the loop turns it into append-only stories at the next safe boundary | `0` · `1` invalid inbox/request |
167
+ | `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE/GLOSSARY.md`, optional `CONTEXT-MAP.md`) | `0` |
168
+ | `yoke loop on\|off\|status\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run` supports `--parallel=N`, bounded reference-driven `--quality`, and blind `--candidates=N` selection; `--max=N` creates an intentional batch cap; `cleanup` retains worktrees unless `--remove-worktrees` is explicit | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
169
+ | `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
170
+ | `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
171
+ | `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
172
+ | `yoke flow-smoke [dir] [--url=] [--label=]` | Browser gate with screenshot/video proofs | `0` green · `1` failures · `2` not runnable |
173
+
174
+ A genuinely hung agent self-terminates after the idle timeout (default 20 min; `--timeout`), and `yoke loop status` shows the live phase or a `⚠ possibly stuck` hint — an autonomous run is never a black box.
175
+
176
+ ## ⚖️ How it compares — superpowers · gstack · Yoke
177
+
178
+ Three excellent projects, three different jobs. Honest version:
179
+
180
+ | | [superpowers](https://github.com/obra/superpowers) (obra) | [gstack](https://github.com/garrytan/gstack) (Garry Tan) | **Yoke** |
181
+ |---|---|---|---|
182
+ | **What it is** | The canonical *skills methodology*: brainstorm → plan → TDD → review as composable skills | A *software factory* for Claude Code: ~40 role skills (QA, CSO, ship…) + a real Chromium browser layer | A *cross-agent harness*: one canon → native installs, plus a gated autonomous loop |
183
+ | **Agents** | Claude Code first | Claude Code + hosts like Codex/Cursor/Kiro — **no Gemini CLI** | **Claude Code, Codex CLI, Gemini CLI** from one source of truth |
184
+ | **Enforcement** | Advisory — skills *describe* the discipline; following them is up to the agent | Skill-driven; browser QA is genuinely real | **Mechanical** — gates live in code: clean tree, acceptance criteria, green tests, review verdict, commit integrity |
185
+ | **Autonomy** | Interactive sessions | Interactive slash-commands (`/qa`, `/ship`, …) | Opt-in **Ralph loop** with watchdog, worktree isolation, single-flight lock, per-story proofs |
186
+ | **Visual QA** | — | **Best-in-class**: live browser daemon (Chromium/CDP) with deep interactive QA | Built-in `flow-smoke` gate: screenshots always, video on failure, labelled per story — lighter, but *enforced* and cross-agent |
187
+ | **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves an independent provider and validates a structured verdict, inside or outside the loop |
188
+ | **Footprint** | Markdown skills (plugin) | ~230 MB with browser runtime; hourly auto-update | Node CLI + markdown canon; Playwright only if you use flow-smoke, resolved **from your project** |
189
+ | **License** | MIT | MIT | MIT |
190
+
191
+ **They compose — use all three where they're strongest.** Yoke's canon *ships* the superpowers methodology natively for all three agents (13 skills, [attributed](canon/skills/ATTRIBUTION.md)). And if gstack is installed, `yoke retrofit` detects it and adds a routing note to `CLAUDE.md` telling Claude to prefer gstack's live-browser `/qa`, `/cso`, and ship pipeline for what Yoke deliberately doesn't bundle — no dependency, no conflict, and Codex/Gemini artifacts stay uniform.
192
+
193
+ **Choose Yoke when** you run more than one agent, want autonomy you can audit (gates + proofs + logs), or want one place to maintain your team's methodology. **Choose gstack when** you live 100% in Claude Code and want the deepest interactive browser QA. **Choose superpowers when** you want the methodology alone, interactively, in Claude Code — or just use it *through* Yoke.
194
+
195
+ ## 🏗️ Architecture
196
+
197
+ You curate **one source of truth** — skills, policy, and tool wiring. Yoke generates the **idiomatic, native artifacts** each agent expects, non-destructively, into any repo:
198
+
199
+ ```mermaid
200
+ flowchart TD
201
+ Canon["📦 CANON — single source of truth<br/>skills · policy · loop spec · tool wiring"]
202
+ Skill["🛠️ yoke retrofit<br/>detect → plan → apply (backup) → report"]
203
+ Canon --> Skill
204
+ Skill --> Claude["Claude Code<br/>.claude/skills · .mcp.json · hook"]
205
+ Skill --> Codex["Codex CLI<br/>AGENTS.md · config.toml · RTK.md"]
206
+ Skill --> Gemini["Gemini CLI<br/>GEMINI.md · commands · settings.json"]
207
+ Loop["🤖 yoke loop — autonomous Ralph loop<br/>gates · verify · review · isolation · proofs"]
208
+ Claude -. drives .-> Loop
209
+ Codex -. drives .-> Loop
210
+ Gemini -. drives .-> Loop
211
+ ```
212
+
213
+ Three layers — **Canon** (`yoke validate`) → **Retrofit** (`yoke retrofit`) → **Loop** (`yoke loop`) — on top of a durable **Context layer** (`yoke context`).
214
+
215
+ ### What gets generated per agent
216
+
217
+ | Agent | Artifacts |
218
+ |---|---|
219
+ | **Claude** | Complete skill packages under `.claude/skills/` (including referenced resources), `AGENTS.md`, `CLAUDE.md`, `.mcp.json` (code-graph + Playwright), and an rtk `PreToolUse` hook when WSL is available |
220
+ | **Codex** | Complete skill packages under `.agents/skills/`, per-skill implicit-invocation policy, `AGENTS.md`, `RTK.md`, `.codex/config.toml`, native hooks, reusable `.codex/agents/*.toml`, and package plugin metadata |
221
+ | **Gemini** | Complete skill packages under `.gemini/skills/`, an auto-invocation index, `GEMINI.md`, `.gemini/commands/*.toml`, and `.gemini/settings.json` (MCP + `AGENTS.md` context) |
222
+
223
+ > **rtk integration:** Claude receives its PreToolUse hook; Codex receives a native hook adapter around `rtk hook check`; Gemini retains instruction-mode fallback where its CLI has no equivalent command-rewrite lifecycle.
224
+
225
+ > **Composes with gstack:** if [gstack](https://github.com/garrytan/gstack) is installed (repo-local or global), `yoke retrofit` adds a short "Composed tools" routing note to **CLAUDE.md only** — telling Claude to prefer gstack's skills for capabilities Yoke doesn't ship (live-browser QA `/qa`, security audit `/cso`, ship/deploy `/ship`). No bundling, no dependency; the note is never written to the Codex or Gemini artifacts.
226
+
227
+ > **Your content survives re-retrofits — preserve blocks:** anything you put between
228
+ > `<!-- yoke:preserve:start -->` and `<!-- yoke:preserve:end -->` in a generated file is
229
+ > carried into the regenerated version on every future `yoke retrofit`. The generated
230
+ > `CLAUDE.md` and `GEMINI.md` ship an empty preserve block scaffold — put your project-specific
231
+ > instructions (tech stack, workflow, `@`-includes) inside it. Works in any yoke-written file;
232
+ > content *outside* the markers is still replaced (and backed up under `.yoke/backup/`).
233
+
234
+ ## 🧰 What's in the canon — 34 skills
235
+
236
+ `yoke retrofit` installs all of these into each agent natively. Provenance is credited in [`canon/skills/ATTRIBUTION.md`](canon/skills/ATTRIBUTION.md).
237
+
238
+ To stop overlapping skills from auto-invoking against each other, `canon/AGENTS.md` carries a **skill routing & precedence** block (methodology before role; one canonical entrypoint per concern — e.g. pre-merge code review is always `review`), emitted into all three agents.
239
+
240
+ Each manifest entry also declares `invocation: auto|manual`. Retrofit translates that intent into
241
+ the provider's native controls: Claude disables model invocation for manual skills, Codex writes
242
+ `agents/openai.yaml`, and Gemini lists only automatic skills in its generated index. Validation
243
+ rejects conflicting package metadata and broken local Markdown links before anything is installed.
244
+
245
+ **Process / methodology** — *superpowers-derived discipline (13)*
246
+
247
+ | Skill | What it does |
248
+ |---|---|
249
+ | `brainstorming` | Explore intent, requirements & design before any creative work |
250
+ | `writing-plans` | Turn a spec into a bite-sized, TDD implementation plan |
251
+ | `executing-plans` | Execute a written plan in a separate session with review checkpoints |
252
+ | `subagent-driven-development` | Run a plan task-by-task: fresh subagent + two-stage review each |
253
+ | `tdd` | Write the test first, watch it fail, write minimal code, refactor |
254
+ | `systematic-debugging` | Root-cause first — no fix without a confirmed cause |
255
+ | `verification-before-completion` | Prove it actually works before claiming done |
256
+ | `using-git-worktrees` | Isolated worktrees for safe / parallel work |
257
+ | `requesting-code-review` | Request a structured review before merging |
258
+ | `receiving-code-review` | Handle review feedback with rigor, not blind agreement |
259
+ | `dispatching-parallel-agents` | Fan out 2+ independent tasks concurrently |
260
+ | `finishing-a-development-branch` | Merge / PR / cleanup a finished branch |
261
+ | `writing-skills` | Author and verify new skills |
262
+
263
+ **Roles** — *gstack-derived, de-gstacked to be harness-agnostic (7)*
264
+
265
+ | Skill | What it does |
266
+ |---|---|
267
+ | `plan-eng-review` | Architecture / edge-case review of a *plan* |
268
+ | `plan-ceo-review` | Founder-mode scope & ambition review of a plan |
269
+ | `review` | Single canonical pre-merge code review — diff safety + engineering quality (architecture, edge cases, tests, performance) |
270
+ | `ship` | Ship workflow: tests → review → version → changelog → PR |
271
+ | `health` | Code-quality dashboard with a composite score |
272
+ | `retro` | Engineering retrospective from commit history |
273
+ | `document-release` | Post-ship documentation sync (README / CHANGELOG / …) |
274
+
275
+ **Yoke-native** — *authored or adapted for this harness (14)*
276
+
277
+ | Skill | What it does |
278
+ |---|---|
279
+ | `yoke-retrofit` | Set up the Yoke harness in a project (detect → plan → apply) |
280
+ | `yoke-workflow` | Provider-neutral planning questions → approved PRD → autonomous stories → critical-decision resume |
281
+ | `authoring-prd` | Slice a product idea into loop-ready stories with testable acceptance criteria |
282
+ | `minimal-code` | Write the least code that solves the task (YAGNI; ponytail-derived) |
283
+ | `performance` | Efficiency as a measured requirement: benchmarks as tests, budgets as gates, optimizations local + documented |
284
+ | `maintaining-context` | Keep `.yoke/context/` the durable source of truth (the Context layer) |
285
+ | `workflow` | The default order of operations, from idea to deploy |
286
+ | `unslop-ui` | Detect & remove AI-slop design tells (purple gradients, neon glow, emoji-icons…) |
287
+ | `visual-verification` | Widen verify to design-scan + the built-in `yoke flow-smoke` gate (screenshot proofs; video on failure) |
288
+ | `no-ai-slop` | Detect and edit generic AI prose while preserving the author's voice; includes its evaluation rubric |
289
+ | `domain-modeling` | Model boundaries, invariants, vocabulary, context maps, and decision records before implementation |
290
+ | `codebase-design` | Explore architecture, deepen a chosen design, and compare two viable approaches when tradeoffs matter |
291
+ | `resolving-merge-conflicts` | Resolve conflicts by reconstructing intent, then verify the integrated result |
292
+ | `writing-for-agents` | Write compact agent instructions with explicit triggers, constraints, resources, and checks |
293
+
294
+ ## 🌱 Zero to 100: `yoke new` + `yoke prd`
295
+
296
+ Yoke's greenfield entrypoint — one command from idea to loop-ready project:
297
+
298
+ ```bash
299
+ yoke new my-app --idea="a CLI that tracks reading lists" # scaffold + retrofit + context + PRD
300
+ yoke loop on my-app && yoke loop run my-app --isolate # hand it to the loop
301
+ ```
302
+
303
+ `yoke new <dir>` refuses a non-empty directory (greenfield-only — use `yoke retrofit` for
304
+ existing projects), then: creates and `git init`s the directory, writes a minimal scaffold
305
+ (`README.md`, `.gitignore`), runs the full **retrofit** (`--agent=` as usual), initialises the
306
+ **context layer** (with `--idea` seeded into `PROJECT.md` as the north star), writes a commented
307
+ **PRD template** to `.yoke/prd.yaml`, and makes the initial commit — so `--isolate` works from
308
+ iteration 1. With `--idea`, it then drafts the PRD from your idea via an agent (`--runner=`,
309
+ the configured runner or active host) and commits it as a second commit (`docs: draft PRD from idea`).
310
+
311
+ - **Exit codes** — `0` success; `1` usage / non-empty dir / draft failure (the scaffold survives —
312
+ retry with `yoke prd draft`); `2` requested draft agent unavailable.
313
+
314
+ **`yoke prd draft [dir] --idea="..."`** turns an idea into 5–12 small, independently shippable
315
+ stories with testable behavioral acceptance criteria (greenfield STORY-1 scaffolds the project
316
+ skeleton + test suite and wires `verify.command`). An existing PRD with stories is never
317
+ overwritten without `--force`; the untouched template doesn't trigger the guard. Runs through
318
+ the same idle-timeout watchdog as the loop (`--timeout`). If `.yoke/plan.md` exists, its approved
319
+ goals, non-goals, constraints, and decisions are injected as settled context instead of being
320
+ reopened by the drafting agent.
321
+
322
+ **`yoke prd check [dir]`** is the chainable pre-loop lint gate: schema validation plus
323
+ duplicate-id, empty-acceptance, unresolved-placeholder, and zero-stories checks. Exits `0` with
324
+ `✓ PRD valid — N stories, M pass`, `1` on any violation. The `authoring-prd` canon skill
325
+ teaches interactive sessions the same story-slicing discipline.
326
+
327
+ ## 🤖 The autonomous loop
328
+
329
+ Opt-in; `yoke setup` recommends enabling it for new installs, while `retrofit` alone keeps it off unless requested. Each iteration starts a **fresh agent** and passes through hard gates before anything is committed:
330
+
331
+ ```mermaid
332
+ flowchart LR
333
+ I[consume queued change<br/>as new stories] --> A[pick next PRD story]
334
+ A --> B{clean worktree?}
335
+ B -- no --> X[blocked]
336
+ B -- yes --> C{acceptance<br/>criteria?}
337
+ C -- no --> X
338
+ C -- yes --> D[agent implements<br/>one story]
339
+ D --> E{suite + criterion<br/>proof green?}
340
+ E -- no --> X
341
+ E -- yes --> V{UI design<br/>within budget?}
342
+ V -- no --> X
343
+ V -- yes --> F{reviewer<br/>approves?}
344
+ F -- no --> X
345
+ F -- yes --> G[commit + mark passes:true<br/>+ proof in .yoke/proof/]
346
+ G --> I
347
+ I --> H{all stories pass?}
348
+ H -- yes --> J{integrated system<br/>gate green?}
349
+ J -- no --> X
350
+ J -- yes --> K[current backlog ready]
351
+ ```
352
+
353
+ ```bash
354
+ yoke loop on . # enable (recorded in .yoke/config.yaml)
355
+ yoke loop status . # show state + PRD progress
356
+ yoke change add . --idea="Add passkey login" # safe while the loop runs
357
+ yoke loop run . \
358
+ --runner=codex \ # implement with Codex…
359
+ --reviewer=claude \ # …review with Claude (role separation)
360
+ --isolate \ # each story in a throwaway git worktree
361
+ --parallel=3 \ # run dependency-ready, non-colliding stories concurrently
362
+ --decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
363
+ # Optional: add --max=20 only when this run should stop after a bounded batch.
364
+ yoke loop off . # disable
365
+ ```
366
+
367
+ **PRD format** (`.yoke/prd.yaml`):
368
+
369
+ ```yaml
370
+ - id: STORY-1
371
+ title: Add a health endpoint
372
+ priority: 1 # lower = higher priority
373
+ acceptance: # Definition of Done (required, else blocked)
374
+ - id: health-returns-200
375
+ text: GET /health returns 200
376
+ verify: [npm run test:health-returns-200]
377
+ - id: health-rejects-post
378
+ text: POST /health returns 405
379
+ verify: [npm run test:health-rejects-post]
380
+ passes: false # the loop sets this true only on green tests
381
+ ```
382
+
383
+ New projects default to `verify.requireCriteria: true`: every story has 2–5 behavioral criteria.
384
+ Each criterion ID must occur in its single, approved test command; shell operators and broad,
385
+ untargeted suites are rejected. Yoke records each result in `.yoke/proof/<story>/evidence.json`. Configure optional
386
+ `completion.command` for integrated journeys such as purchase → entitlement → relaunch or
387
+ magic-link → callback → authenticated app. It runs whenever the current backlog has no open
388
+ stories; this is readiness, not a release.
389
+
390
+ No generic tool can infer whether arbitrary test code perfectly represents product meaning. Yoke
391
+ closes the mechanical false-done paths—targeted evidence, coverage review, clean committed state,
392
+ and integrated journeys—while the project still owns the correctness of its tests and production
393
+ observability.
394
+
395
+ ### Parallel workers and the quality gauntlet
396
+
397
+ `--parallel=N` dispatches dependency-ready stories concurrently. Claims carry leases, workers use
398
+ isolated worktrees, collision areas are serialized, and only a mechanically green candidate enters
399
+ the FIFO integration queue. Integration repeats the project gates against the merged tree; a worker
400
+ success can never bypass a red integrated result. `yoke loop status` reports the dispatcher,
401
+ workers, providers, worktrees, lifecycle, queue, integrations, and reopened stories.
402
+
403
+ Quality is reference-driven and opt-in. Declare what one story should match:
404
+
405
+ ```yaml
406
+ quality:
407
+ reference: { name: approved-home, source: design/home.png, kind: file }
408
+ candidate: { kind: screenshots, paths: [.yoke/proof/STORY-1/home.png] }
409
+ rubric: Match the approved layout, hierarchy, spacing, and states.
410
+ policy: blocking # or advisory
411
+ ```
412
+
413
+ Configure project defaults, then enable the gauntlet for a run:
414
+
415
+ ```yaml
416
+ quality:
417
+ enabled: false # keep opt-in, or make it the project default
418
+ policy: blocking
419
+ maxRounds: 3
420
+ maxMinutes: 60
421
+ consistencyChecks: 2
422
+ maxParallelCandidates: 2
423
+ critic: { agent: codex, model: gpt-5.6-sol } # model required for --candidates
424
+ repair: { agent: claude }
425
+ ```
426
+
427
+ ```bash
428
+ yoke loop run . --quality --quality-rounds=3 --quality-minutes=60
429
+ yoke loop run . --quality --candidates=2 # blind pairwise selection; stories need quality declarations
430
+ ```
431
+
432
+ The critic compares opaque candidate/reference labels, writes schema-validated provenance, and
433
+ cannot modify the project. Blocking findings enter a bounded repair loop and rerun every mechanical
434
+ gate; advisory findings are retained without blocking. `--quality-policy=`, `--no-quality`, and
435
+ `--quality-unbounded` override defaults for one run. Unbounded mode is explicit and warned because
436
+ it removes repair limits, not Yoke's watchdog, isolation, verification, or commit safety.
437
+
438
+ State lives **outside the model context** — the PRD file plus git — so each iteration is fresh.
439
+ Use `yoke change add` at any time. Its ignored append-only inbox is consumed at the next story
440
+ boundary. A separate coverage pass must confirm that every requested outcome maps to behavioral
441
+ criteria before Yoke appends and commits the new stories; existing stories are never rewritten and
442
+ no restart is needed.
443
+
444
+ ### Watching a run
445
+
446
+ Every iteration emits token-free, harness-side feedback (Node console + local files — **zero agent tokens**):
447
+
448
+ - **Live console with progress + ETA** —
449
+ `▶ S6 (19/45 · 42%) — implementing… · ~1h44m left (Ø 4m/story)` … `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`.
450
+ The estimate uses the **average duration of stories completed in this run** (current
451
+ velocity); before the first story lands it falls back to the recorded history of previous
452
+ runs (`.yoke/story-durations.json`, last 50 stories, gitignored). No data yet → no estimate,
453
+ never a made-up one.
454
+ - **`.yoke/loop-status.json`** — the current state (now including `percent` and an `eta`
455
+ block); read it any time with `yoke loop status`:
456
+ ```
457
+ Loop: RUNNING on S6 "Weekly digest"
458
+ implementing · iteration 20 · 19/45 (42%) · updated 30s ago
459
+ ~1h44m remaining (Ø 4m/story)
460
+ ```
461
+ - **Parallel + quality detail** — active workers include provider, candidate ID, worktree,
462
+ lifecycle, phase, quality round, and repair budget; the integrator is shown separately.
463
+ - **`.yoke/loop.log`** — an append-only timeline of every phase transition.
464
+ - **`--json`** — machine mode for supervisors: every status write is *also* emitted as one
465
+ NDJSON line on stdout (`{"type":"status","state":"running","phase":"verifying",…}` — the
466
+ same shape as `loop-status.json`), the human narrative moves off stdout (the final summary
467
+ goes to stderr), and a consumer can follow the stream line by line instead of polling the file.
468
+ Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
469
+ output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
470
+ and cost fields when available. Missing values stay absent—Yoke does not estimate them.
471
+
472
+ ### Pausing a run
473
+
474
+ Drop a **`.yoke/loop.pause`** file (contents irrelevant) while the loop is running and it
475
+ stops at the **next story boundary** — the running story still finishes, verifies, and
476
+ commits; no story is ever cut off mid-flight. The loop consumes the pause file, writes
477
+ `state: "paused"` to `loop-status.json` (log label `paused`), releases the lock, and exits
478
+ with code `3`. Resume by simply running `yoke loop run` again.
479
+
480
+ A per-iteration **idle timeout** guards against a genuinely hung agent: if the agent produces
481
+ **no output at all** for `--timeout` minutes (default 20; `0` disables), the loop kills it
482
+ (SIGTERM→SIGKILL) and marks the story blocked. A slow-but-working agent that keeps streaming
483
+ output is **never** killed — the output stream *is* the liveness signal. Set a project default
484
+ with `loop.timeoutMinutes` in `.yoke/config.yaml`.
485
+
486
+ ### Decision policy: autonomous by default, interrupt only when configured
487
+
488
+ Planning questions happen before the loop. The provider-neutral `yoke-workflow` skill asks only
489
+ questions whose answer materially changes product behavior, scope, architecture, security, data
490
+ ownership, external cost, or an irreversible choice. It saves the approved brief in
491
+ `.yoke/plan.md`; `yoke prd draft` consumes it, and `yoke prd check` rejects explicit unresolved
492
+ placeholders such as `TBD`.
493
+
494
+ The unattended loop then follows `loop.decisionPolicy`:
495
+
496
+ ```yaml
497
+ loop:
498
+ enabled: true
499
+ decisionPolicy: critical # or auto
500
+ runner:
501
+ agent: codex # setup chooses the current host by default
502
+ ```
503
+
504
+ - **`auto` (default):** routine ambiguity and implementation details are resolved using the
505
+ approved plan, acceptance criteria, current code, and project conventions. The loop does not
506
+ ask follow-up questions.
507
+ - **`critical`:** routine choices are still resolved automatically. Only high-impact decisions
508
+ involving public architecture, security/privacy, destructive migration or data loss, material
509
+ external cost, legal/compliance exposure, or another irreversible choice may pause the story.
510
+ The agent writes a schema-validated request; the loop blocks before verify and preserves it as
511
+ `.yoke/pending-decision.yaml`.
512
+
513
+ Inspect and answer a critical stop:
514
+
515
+ ```bash
516
+ yoke loop decision .
517
+ yoke loop answer . --choice=A --rationale="Matches the existing identity model"
518
+ ```
519
+
520
+ `answer` validates the choice against the still-open story, appends it to
521
+ `.yoke/context/DECISIONS.md`, commits only that file using the configured human identity, clears
522
+ the pending request, and resumes the same story with the original runner, isolation, review,
523
+ permission, timeout, JSON, decision-policy, and iteration settings intact. Add
524
+ `--no-resume` when a supervisor should restart the loop separately. If the automatic restart
525
+ cannot begin because a provider/reviewer is unavailable or another process owns the lock, run
526
+ `yoke loop resume .`; its request-bound options are retained under Git's private state directory
527
+ until a loop actually runs. To intentionally abandon an orphaned or stale private resume state,
528
+ use `yoke loop resume . --discard`; pending decisions are never deleted by that command. Existing
529
+ `loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` remain supported as compatibility aliases;
530
+ new projects should use `decisionPolicy: auto|critical`.
531
+
532
+ ### Adaptive model routing (explicit opt-in)
533
+
534
+ `yoke setup` asks before enabling routing; the default is **off**. When enabled, the selected
535
+ parent remains the strong planner/controller. Before each bounded story it receives only the
536
+ story, acceptance criteria, and at most three eligible worker profiles, then returns one
537
+ machine-readable choice. The worker can be a cheaper/faster Claude, Codex, or Gemini profile;
538
+ `SELF` keeps difficult work on the parent. Provider-native subagents are disabled for these
539
+ runs so Yoke does not pay for two orchestration layers.
540
+
541
+ **Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
542
+ Code, Codex CLI, and Gemini CLI, including mixed-provider worker lists. Internal contract tests
543
+ cover invocation and routing behavior for all three providers. The measured performance evidence
544
+ below is intentionally **Codex-only**; it does not claim equivalent Claude or Gemini savings
545
+ until authenticated, repeated in-the-wild runs exist for those providers.
546
+
547
+ ```yaml
548
+ runner:
549
+ agent: codex
550
+ model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
551
+ reasoningEffort: high
552
+ routing:
553
+ enabled: true # setup defaults false; setup --routing opts in
554
+ strategy: balanced # balanced | cost | speed | quality
555
+ maxCandidates: 3
556
+ workers:
557
+ - id: codex-light
558
+ agent: codex
559
+ reasoningEffort: low
560
+ costTier: medium
561
+ capabilities: [exploration, implementation, tests]
562
+ - id: claude-fast
563
+ agent: claude
564
+ model: haiku # rolling alias; omit to use the provider's current default
565
+ reasoningEffort: low
566
+ costTier: low
567
+ capabilities: [mechanical-edits, tests]
568
+ - id: gemini-auto
569
+ agent: gemini # omitted model means the account's current Auto/default route
570
+ costTier: low
571
+ capabilities: [large-context, implementation]
572
+ ```
573
+
574
+ Use `yoke loop run . --routing` for a one-run opt-in or `--no-routing` for a controlled
575
+ baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
576
+ eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
577
+ "intelligence score". Candidate model IDs come from project configuration while setup defaults
578
+ prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
579
+ independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
580
+ after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
581
+ time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
582
+ cannot overwrite a shared registry file.
583
+
584
+ Routing is not free: it adds one controller call per story. It is most promising when a bounded
585
+ worker saves more than that call costs; tiny stories may be slower. Keep it opt-in and measure it
586
+ on your own backlog rather than assuming a win.
587
+
588
+ ### Performance budgets: efficiency as a gate, not a style
589
+
590
+ Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
591
+ be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
592
+
593
+ - **Per story:** write the requirement as a **measurable acceptance criterion**
594
+ ("imports 1M rows in < 2s, asserted by the bench test") and let your verify tests measure
595
+ it — no new machinery needed.
596
+ - **Per project:** wire a benchmark as a standing **perf gate** in `.yoke/config.yaml`:
597
+
598
+ ```yaml
599
+ perf:
600
+ command: node bench/check-budget.mjs # exit 0 = within budget
601
+ retries: 1 # benchmarks are noisy; same retry logic as verify
602
+ ```
603
+
604
+ The loop runs it **after verify** on every story (phase `perf`, with `YOKE_STORY` set); a
605
+ red benchmark blocks the story — `story S6 exceeded its performance budget: p95 62ms > budget 50ms` —
606
+ no matter how clean the diff was. The implementer prompt names the budget command, so the
607
+ agent knows not to trade hot-path efficiency for style and never "simplifies away" an
608
+ optimization without re-running the benchmark. The `performance` canon skill carries the
609
+ method: profile first, optimize leaves not boundaries, commit benchmarks as tests, version
610
+ the *why* of every optimization in `context/DECISIONS.md`.
611
+
612
+ ### Artifact-backed gate output: compact context, complete local evidence
613
+
602
614
  Failed verify, executable-criterion, performance, configured custom-audit, and completion commands can emit thousands of
603
- low-signal lines. Yoke keeps the model-visible failure summary deterministic and bounded while
604
- preserving large raw stdout/stderr below `.yoke/artifacts/`:
605
-
606
- ```yaml
607
- output:
608
- previewBytes: 2048 # default: maximum compact preview bytes
609
- artifactThresholdBytes: 8192 # default: persist raw output only above this size
610
- ```
611
-
612
- The preview prioritizes errors, warnings, adjacent context, and final test summaries. Above the
613
- artifact threshold it also includes a project-relative path, byte count, and full SHA-256 digest,
614
- for example:
615
-
616
- ```text
617
- [full output: .yoke/artifacts/STORY-4/verify-0123abcd4567.log | 42810 bytes | sha256:0123...]
618
- ```
619
-
620
- An agent can read that ordinary file when the preview is insufficient; nothing is injected into
621
- later stories automatically. Repeated identical failures reuse the same content-addressed path.
615
+ low-signal lines. Yoke keeps the model-visible failure summary deterministic and bounded while
616
+ preserving large raw stdout/stderr below `.yoke/artifacts/`:
617
+
618
+ ```yaml
619
+ output:
620
+ previewBytes: 2048 # default: maximum compact preview bytes
621
+ artifactThresholdBytes: 8192 # default: persist raw output only above this size
622
+ ```
623
+
624
+ The preview prioritizes errors, warnings, adjacent context, and final test summaries. Above the
625
+ artifact threshold it also includes a project-relative path, byte count, and full SHA-256 digest,
626
+ for example:
627
+
628
+ ```text
629
+ [full output: .yoke/artifacts/STORY-4/verify-0123abcd4567.log | 42810 bytes | sha256:0123...]
630
+ ```
631
+
632
+ An agent can read that ordinary file when the preview is insufficient; nothing is injected into
633
+ later stories automatically. Repeated identical failures reuse the same content-addressed path.
622
634
  Successful gate output is discarded as before. This affects only commands executed by Yoke's own
623
635
  gates. It does **not** intercept tool output generated internally by Claude Code, Codex, or Gemini,
624
636
  so benchmark ratios for this feature are not provider-token or billing claims.
@@ -629,225 +641,234 @@ partial evidence as full output.
629
641
 
630
642
  Yoke treats `.yoke/artifacts/` as local, non-committable runtime state and excludes it from its
631
643
  clean-tree and story-commit operations; `yoke retrofit` also adds it to `.gitignore`. Raw command output is intentionally stored
632
- without redaction so it remains valid evidence and may therefore contain credentials, personal
633
- data, or other sensitive text emitted by project commands. Inspect artifacts before sharing them.
634
-
635
- The loop trusts **verify**, not the agent's exit code: a story whose tests are green is
636
- committed even if the agent process exited non-zero (a common Windows `.cmd`-wrapper ghost).
637
- A failing verify is retried up to `verify.retries` times (default 1) so a transient flake
638
- self-heals while a real failure still blocks. Structured acceptance criteria are then verified
639
- individually; an unrelated green suite cannot satisfy a criterion without its proof command.
640
-
641
- `.yoke/loop-status.json`, `.yoke/loop.log`, `.yoke/loop.lock`, its takeover/recovery leases, lock/decision temp files, `.yoke/story-durations.json`,
642
- `.yoke/ambiguity.md`, `.yoke/artifacts/`, and the critical-decision request/answering files are runtime artifacts;
643
- `yoke retrofit` gitignores them (along with
644
- `.yoke/worktrees/`, `.yoke/backup/`, `.yoke/proof/`, and `.yoke/changes/`) so they never trip the clean-tree gate.
645
-
646
- ### Single-flight guard + cleanup
647
-
648
- Two concurrent `yoke loop run`s would race on the PRD and status files, so the loop takes a
649
- **lock** (`.yoke/loop.lock`) for the duration of a run. Complete lock metadata is published atomically;
650
- stale takeover is serialized by `.yoke/loop.lock.takeover`. A second invocation exits `2` with
651
- `Another loop is already running here (pid …). If that is wrong, run: yoke loop cleanup`. A lock
652
- whose holder process is dead is taken over automatically (with a warning).
653
-
654
- **`yoke loop cleanup [dir]`** reaps only runner process trees recorded by this project and removes
655
- a stale lock. Yoke-created worktrees are **retained by default** and listed in the output; pass
656
- `--remove-worktrees` to remove `.yoke/worktrees/*` with `git worktree remove --force` + `prune`.
657
- User-created worktrees are never touched. A live lock is reported and left alone. Exits `0` when
658
- cleanup succeeds, `1` if any requested removal fails. If a machine/process crash leaves the cleanup
659
- recovery lease itself behind, an operator can run
660
- `yoke loop cleanup . --discard-stale-recovery`; Yoke refuses while its recorded PID is alive, and
661
- the force flag must not be run concurrently.
662
-
663
- ## 🔍 Cross-model review (`yoke review`)
664
-
665
- Outside the loop, `yoke review` has a **second** model review your current diff as a
666
- pass/fail gate — the interactive counterpart to the loop's `--review`/`--reviewer`.
667
-
668
- ```bash
669
- yoke review . # review the uncommitted working tree
670
- yoke review . --base=main # review the range main..HEAD instead
671
- yoke review . --reviewer=codex # force a specific reviewer
672
- yoke review . --focus="the auth layer" # steer what it scrutinises
673
- ```
674
-
675
- - **Reviewer resolution** — picks the first available of **codex → gemini → claude**,
676
- preferring a model *other* than the one you drive so the review is genuinely cross-model.
677
- On a Claude-only machine it degrades to a self-review (and says so).
678
- - **Scope** — the uncommitted working tree by default, or a commit range with `--base=<ref>`.
679
- - **Exit-code gate** — exits `0` when the reviewer approves, `1` when it finds a blocking
680
- issue, `2` when no (or an unavailable) reviewer CLI is found. Chain it: `... && yoke review`,
681
- or wire it into a pre-push hook.
682
- - Runs through the same idle-timeout watchdog as the loop (`--timeout`, default 20 min).
683
-
684
- ## 🎨 Visual & design verification — done, with a photo
685
-
686
- Unit tests don't catch a blank page, an unwired route, or generic AI-slop design. Yoke adds three things:
687
-
688
- - **`yoke design-scan [dir]`** — a static scanner for the visual *tells* of AI-generated UIs
689
- (AI-purple gradients, gradient hero text, neon glow, emoji-as-icons, gradient overload). It
690
- scores findings and **exits non-zero over budget** (`--max`, default 4; `--report` to list only),
691
- so it drops straight into your verify pipeline.
692
- - **`yoke flow-smoke [dir]`** — a built-in browser gate with **proof artifacts** (below).
693
- - **`unslop-ui` + `visual-verification` skills** — the design rubric, plus how to compose a verify
694
- pipeline (`types → units → design-scan → flow-smoke`).
695
-
696
- Because the loop trusts **verify as the source of truth**, widening `verify.command` to include the
697
- scanner and the flow-smoke makes visual quality a real gate — not an afterthought.
698
-
699
- *Tell set informed by the MIT-licensed [vibecoded-design-tells](https://github.com/JCarterJohnson/vibecoded-design-tells) research.*
700
-
701
- ### `yoke flow-smoke [dir] [--url=<baseUrl>] [--label=<name>]`
702
-
703
- Configure your key user flows once in `.yoke/config.yaml`:
704
-
705
- ```yaml
706
- smoke:
707
- baseUrl: http://localhost:3000
708
- flows:
709
- - name: home
710
- path: /
711
- landmark: "main h1" # optional CSS selector to wait for
712
- - name: login
713
- path: /login
714
- ```
715
-
716
- For every flow, `yoke flow-smoke` loads the route against the running dev server, waits for the
717
- landmark, and fails on a non-OK response or **any console/page error**. The proof contract:
718
-
719
- - **Screenshots always** — every flow (pass *or* fail) saves `.yoke/proof/<label>/<flow>.png`;
720
- the failure screenshot *is* the evidence.
721
- - **Video only on failure** — each flow is recorded, but the clip is kept only when the flow
722
- goes red (`<flow>.webm`); green runs delete it.
723
- - **Labelled per story** — inside the loop, verify runs with `YOKE_STORY=<story-id>`, so proofs
724
- land in `.yoke/proof/<story-id>/` automatically. Standalone runs use `latest`, or pass
725
- `--label=`. The label dir is wiped per run — evidence is always from the latest run.
726
- - **Exit codes** — `0` all flows green (chain it: `... && yoke design-scan . && yoke flow-smoke .`),
727
- `1` any flow failed, `2` not runnable (no `smoke:` config, or Playwright missing).
728
- - **Playwright comes from the *target project*, never Yoke** —
729
- `npm i -D playwright && npx playwright install chromium` there. Start the dev server before
730
- verify (e.g. via `start-server-and-test`); `--url=` overrides `baseUrl`.
731
-
732
- `.yoke/proof/` is gitignored by the retrofit — proofs are runtime artifacts and never break the
733
- loop's clean-tree gate.
734
-
735
- ## 🧠 Context layer (`.yoke/context/`)
736
-
737
- Yoke keeps durable, cross-session context so a fresh-context agent is never blind:
738
-
739
- - `PROJECT.md` — the north star (goal, constraints, non-goals, success criteria).
740
- - `DECISIONS.md` — an append-only ledger. The loop adds an entry per completed story; you and agents add the *why*.
741
- - `KNOWLEDGE.md` — reusable gotchas and conventions.
742
-
743
- `yoke retrofit` scaffolds these files (non-destructively — your edits are never overwritten).
744
- The loop reads them into every agent + reviewer prompt and logs decisions back on each story's
745
- commit. Decision history is explicitly delimited as untrusted reference data, so stored text is
746
- never treated as fresh instructions. Manage the files directly with `yoke context init` and `yoke context status`. The
747
- `maintaining-context` skill teaches agents to honour the same files during interactive work.
748
-
749
- > Commit `.yoke/context/` to git. The `--isolate` loop runs each iteration in a worktree
750
- > checked out from HEAD, so it only sees committed context.
751
-
752
- ## 🛡️ Safety model
753
-
754
- Yoke's guardrails are **mechanical, not advisory** — the loop blocks on a dirty worktree, missing acceptance criteria, red tests, or a reviewer rejection, and **none of them rely on the agent choosing to behave**.
755
-
756
- - **Commit integrity** — a story is never recorded `passes: true` without a corresponding commit; a failed commit reverts the PRD.
757
- - **Role separation** — the implementer never reviews its own work; `--reviewer` can even be a different agent.
758
- - **Isolation** — with `--isolate`, failed or partial work is discarded with the worktree and never reaches your main tree.
759
- - **Non-destructive retrofit** — existing files are backed up before any change; settings are merged, not replaced.
760
- - **Independent verification** — "done" means *your test command exits 0*, not "the agent said so".
761
- - **Single-flight** — a lock prevents two loops from racing the same repo; `yoke loop cleanup` recovers after crashes.
762
-
763
- ## 🧠 Choose your code-graph
764
-
765
- `yoke retrofit --code-graph=graphify|serena` (default `graphify`, remembered per project). The `yoke-retrofit` skill asks and recommends based on the project.
766
-
767
- | | **graphify** | **Serena** |
768
- |---|---|---|
769
- | Engine | tree-sitter AST + graph | real language servers (LSP) |
770
- | Strength | fast, multimodal (code + PDFs + images) | symbol-exact cross-file refactoring |
771
- | Token efficiency | ~70× reduction on large mixed repos | standard, no index to go stale |
772
- | Best for | rapid exploration / migration / onboarding | systematic refactoring in typed codebases |
773
- | Caveat | heuristic edges; static index can go stale | one language server per language |
774
-
775
- ## 🪙 Token efficiency
776
-
777
- Yoke attacks tokens on two complementary surfaces:
778
-
779
- - **rtk** compresses noisy command/tool output before it enters context (wired as a hook/instruction per agent).
780
- - The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
781
-
782
- A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
783
- completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
784
- controller selected Luna for every bounded implementation story. Including controller overhead,
785
- the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
786
- and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
787
-
788
- The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
789
- controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
790
- emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
791
- inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
792
- [`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
793
-
794
- ## 🧩 Optional companions
795
-
796
- Two external tools pair well with Yoke and are documented (not bundled) in `canon/tools/` — each has its own installer and update cadence, so Yoke wires the boundary instead of vendoring a copy:
797
-
798
- - **[claude-mem](https://github.com/thedotmack/claude-mem)** — persistent cross-session memory for interactive work. Deliberate boundary: the autonomous loop keeps its memory **explicit and versioned** (`context/*.md` + PRD, fresh context per story), so claude-mem's automatic injection stays out of loop runs.
799
- - **[ui-ux-pro-max](https://github.com/nextlevelbuilder/ui-ux-pro-max-skill)** — data-driven design intelligence (styles, palettes, industry rules) for the *generation* side. Yoke's `unslop-ui`, `design-scan`, and `visual-verification` remain the *verification* side: generate with pro-max, gate with Yoke.
800
-
801
- ## 🌱 Why & how it was built
802
-
803
- **The problem.** Coding agents are powerful, but each speaks its own dialect — Claude has skills and hooks, Codex reads `AGENTS.md` and a TOML config, Gemini wants commands and a settings file. Keeping the same skills, safety policy, and tool wiring consistent across all of them means copy-paste drift and three things to maintain. Yoke exists to keep **one** source of truth, generate the right native artifacts for each agent, and let that harness run **autonomously and safely** when you want to hand it a spec and walk away.
804
-
805
- **The inspiration.** Yoke is a synthesis of ideas already proven across the ecosystem: composable-skills methodology ([superpowers](https://github.com/obra/superpowers), [gstack](https://github.com/garrytan/gstack)); the portable [AGENTS.md](https://agents.md/) standard; the *"one source-of-truth → idiomatic per-harness artifacts"* generation pattern ([wshobson/agents](https://github.com/wshobson/agents)); spec-driven autonomous orchestration (GSD); mechanical safety gates and role separation (safe-agentic-workflow); and the **Ralph loop** (Geoff Huntley) — keep handing a *fresh* agent the next task until the spec is done. Token efficiency comes from [rtk](https://github.com/rtk-ai/rtk) and the write-less-code idea behind [ponytail](https://github.com/DietrichGebert/ponytail).
806
-
807
- **How it was built.** Yoke was built the way it's meant to be *used* — agent-driven, incremental, and test-first. The stack was chosen by **researching alternatives first** (which is how `jcodemunch` was dropped for its license and Serena was added as an option). Then every component shipped one small piece at a time through a disciplined loop: **brainstorm → spec → plan → TDD implementation → an independent two-stage review** (does it match the spec? is it well-built?) **→ merge**. Those reviews caught real bugs before they shipped — a Windows `.cmd` spawn failure, a commit-integrity hole, a path-traversal that could delete project data, a resolution bug that broke the CLI's default invocation, a TOML-escaping bug. Yoke was even **dogfooded on its own repo**, which surfaced (and fixed) a genuine Windows bug. Every spec and plan lives in [`docs/superpowers/`](docs/superpowers/).
808
-
809
- ## 🗂️ Project layout
810
-
811
- ```text
812
- canon/ # the source of truth — harness-agnostic
813
- AGENTS.md skills/ policy/ loop/ tools/ manifest.yaml
814
- src/
815
- canon/ # manifest schema + validator (yoke validate)
816
- change/ # append-only change inbox · planning · independent coverage review
817
- retrofit/ # detect · plan · apply · planners (claude/codex/gemini) · tools
818
- loop/ # prd · gates · runner · verify · git/worktree · loop · run-command · lock · cleanup
819
- quality/ # reference collection · blind critic · bounded repair · candidate comparison
820
- new/ # yoke new — greenfield bootstrap
821
- prd/ # yoke prd draft|check — idea → stories + lint gate
822
- review/ # yoke review — cross-model diff gate
823
- smoke/ # yoke flow-smoke — browser gate with screenshot/video proofs
824
- scan/ # yoke design-scan — AI-slop design gate
825
- context/ # the durable context layer
826
- docs/superpowers/ # the spec and every component's implementation plan
827
- ```
828
-
829
- ## 🗺️ Roadmap
830
-
831
- Completed release work lives in the changelog. Remaining, explicitly scoped work is tracked in
832
- [`TODOS.md`](TODOS.md), including broader benchmark samples, native output schemas, and signed
833
- release provenance.
834
-
835
- ## 🧪 Development
836
-
837
- ```bash
838
- npm test # vitest (971 tests)
839
- npm run build # tsc, no emit errors
840
- npm run yoke -- validate canon
841
- ```
842
-
843
- ## 🙏 Credits & inspiration
844
-
845
- Yoke stands on the shoulders of a great ecosystem: methodology ideas from [superpowers](https://github.com/obra/superpowers) and [gstack](https://github.com/garrytan/gstack); the [AGENTS.md](https://agents.md/) standard; the generator pattern from [wshobson/agents](https://github.com/wshobson/agents); the Ralph autonomous-loop pattern; safety-gate thinking from safe-agentic-workflow; and the wired tools [rtk](https://github.com/rtk-ai/rtk), [graphify](https://github.com/safishamsi/graphify), [Serena](https://github.com/oraios/serena), and [Playwright MCP](https://github.com/microsoft/playwright-mcp). The `minimal-code` skill adapts the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.
846
-
847
- ## 📄 License
848
-
849
- MIT — see [`LICENSE`](LICENSE).
850
-
851
- <div align="center">
852
- <sub>Built with a disciplined loop: brainstorm → spec → plan → TDD → two-stage review → merge — and reviewed by a second model, because we don't trust "done" either.</sub>
853
- </div>
644
+ without redaction so it remains valid evidence and may therefore contain credentials, personal
645
+ data, or other sensitive text emitted by project commands. Inspect artifacts before sharing them.
646
+
647
+ The loop trusts **verify**, not the agent's exit code: a story whose tests are green is
648
+ committed even if the agent process exited non-zero (a common Windows `.cmd`-wrapper ghost).
649
+ A failing verify is retried up to `verify.retries` times (default 1) so a transient flake
650
+ self-heals while a real failure still blocks. Structured acceptance criteria are then verified
651
+ individually; an unrelated green suite cannot satisfy a criterion without its proof command.
652
+
653
+ `.yoke/loop-status.json`, `.yoke/loop.log`, `.yoke/loop.lock`, its takeover/recovery leases, lock/decision temp files, `.yoke/story-durations.json`,
654
+ `.yoke/ambiguity.md`, `.yoke/artifacts/`, and the critical-decision request/answering files are runtime artifacts;
655
+ `yoke retrofit` gitignores them (along with
656
+ `.yoke/worktrees/`, `.yoke/backup/`, `.yoke/proof/`, and `.yoke/changes/`) so they never trip the clean-tree gate.
657
+
658
+ ### Single-flight guard + cleanup
659
+
660
+ Two concurrent `yoke loop run`s would race on the PRD and status files, so the loop takes a
661
+ **lock** (`.yoke/loop.lock`) for the duration of a run. Complete lock metadata is published atomically;
662
+ stale takeover is serialized by `.yoke/loop.lock.takeover`. A second invocation exits `2` with
663
+ `Another loop is already running here (pid …). If that is wrong, run: yoke loop cleanup`. A lock
664
+ whose holder process is dead is taken over automatically (with a warning).
665
+
666
+ **`yoke loop cleanup [dir]`** reaps only runner process trees recorded by this project and removes
667
+ a stale lock. Yoke-created worktrees are **retained by default** and listed in the output; pass
668
+ `--remove-worktrees` to remove `.yoke/worktrees/*` with `git worktree remove --force` + `prune`.
669
+ User-created worktrees are never touched. A live lock is reported and left alone. Exits `0` when
670
+ cleanup succeeds, `1` if any requested removal fails. If a machine/process crash leaves the cleanup
671
+ recovery lease itself behind, an operator can run
672
+ `yoke loop cleanup . --discard-stale-recovery`; Yoke refuses while its recorded PID is alive, and
673
+ the force flag must not be run concurrently.
674
+
675
+ ## 🔍 Cross-model review (`yoke review`)
676
+
677
+ Outside the loop, `yoke review` has a **second** model review your current diff as a
678
+ pass/fail gate — the interactive counterpart to the loop's `--review`/`--reviewer`.
679
+
680
+ ```bash
681
+ yoke review . # review the uncommitted working tree
682
+ yoke review . --base=main # review the range main..HEAD instead
683
+ yoke review . --reviewer=codex # force a specific reviewer
684
+ yoke review . --focus="the auth layer" # steer what it scrutinises
685
+ ```
686
+
687
+ - **Reviewer resolution** — picks the first available of **codex → gemini → claude**,
688
+ preferring a model *other* than the one you drive so the review is genuinely cross-model.
689
+ On a Claude-only machine it degrades to a self-review (and says so).
690
+ - **Scope** — the uncommitted working tree by default, or a commit range with `--base=<ref>`.
691
+ - **Exit-code gate** — exits `0` when the reviewer approves, `1` when it finds a blocking
692
+ issue, `2` when no (or an unavailable) reviewer CLI is found. Chain it: `... && yoke review`,
693
+ or wire it into a pre-push hook.
694
+ - Runs through the same idle-timeout watchdog as the loop (`--timeout`, default 20 min).
695
+
696
+ ## 🎨 Visual & design verification — done, with a photo
697
+
698
+ Unit tests don't catch a blank page, an unwired route, or generic AI-slop design. Yoke adds three things:
699
+
700
+ - **`yoke design-scan [dir]`** — a static scanner for the visual *tells* of AI-generated UIs
701
+ (AI-purple gradients, gradient hero text, neon glow, emoji-as-icons, gradient overload). It
702
+ scores findings and **exits non-zero over budget** (`--max`, default 4; `--report` to list only),
703
+ so it drops straight into your verify pipeline.
704
+ - **`yoke flow-smoke [dir]`** — a built-in browser gate with **proof artifacts** (below).
705
+ - **`unslop-ui` + `visual-verification` skills** — the design rubric, plus how to compose a verify
706
+ pipeline (`types → units → design-scan → flow-smoke`).
707
+
708
+ Retrofit adds `design: { mode: auto, max: 4 }` when it detects UI dependencies, UI source files,
709
+ or configured smoke flows. In `auto` mode the loop runs the design scan after functional verify and
710
+ before performance/audit; `on` forces it for any project and `off` disables it. Existing explicit
711
+ settings are preserved. `flow-smoke` remains an explicit project verify step because Yoke cannot
712
+ infer how to start each application's server.
713
+
714
+ `unslop-ui` is the visual-design skill. `no-ai-slop` is separate: it reviews prose for generic AI
715
+ patterns and edits only confirmed problems while preserving meaning and voice.
716
+
717
+ *Tell set informed by the MIT-licensed [vibecoded-design-tells](https://github.com/JCarterJohnson/vibecoded-design-tells) research.*
718
+
719
+ ### `yoke flow-smoke [dir] [--url=<baseUrl>] [--label=<name>]`
720
+
721
+ Configure your key user flows once in `.yoke/config.yaml`:
722
+
723
+ ```yaml
724
+ smoke:
725
+ baseUrl: http://localhost:3000
726
+ flows:
727
+ - name: home
728
+ path: /
729
+ landmark: "main h1" # optional CSS selector to wait for
730
+ - name: login
731
+ path: /login
732
+ ```
733
+
734
+ For every flow, `yoke flow-smoke` loads the route against the running dev server, waits for the
735
+ landmark, and fails on a non-OK response or **any console/page error**. The proof contract:
736
+
737
+ - **Screenshots always** — every flow (pass *or* fail) saves `.yoke/proof/<label>/<flow>.png`;
738
+ the failure screenshot *is* the evidence.
739
+ - **Video only on failure** — each flow is recorded, but the clip is kept only when the flow
740
+ goes red (`<flow>.webm`); green runs delete it.
741
+ - **Labelled per story** — inside the loop, verify runs with `YOKE_STORY=<story-id>`, so proofs
742
+ land in `.yoke/proof/<story-id>/` automatically. Standalone runs use `latest`, or pass
743
+ `--label=`. The label dir is wiped per run — evidence is always from the latest run.
744
+ - **Exit codes** — `0` all flows green (chain it: `... && yoke design-scan . && yoke flow-smoke .`),
745
+ `1` any flow failed, `2` not runnable (no `smoke:` config, or Playwright missing).
746
+ - **Playwright comes from the *target project*, never Yoke** —
747
+ `npm i -D playwright && npx playwright install chromium` there. Start the dev server before
748
+ verify (e.g. via `start-server-and-test`); `--url=` overrides `baseUrl`.
749
+
750
+ `.yoke/proof/` is gitignored by the retrofit — proofs are runtime artifacts and never break the
751
+ loop's clean-tree gate.
752
+
753
+ ## 🧠 Context layer (`.yoke/context/`)
754
+
755
+ Yoke keeps durable, cross-session context so a fresh-context agent is never blind:
756
+
757
+ - `PROJECT.md` — the north star (goal, constraints, non-goals, success criteria).
758
+ - `DECISIONS.md` — an append-only ledger. The loop adds an entry per completed story; you and agents add the *why*.
759
+ - `KNOWLEDGE.md` — reusable gotchas and conventions.
760
+ - `GLOSSARY.md` — the project's canonical terms, meanings, and aliases.
761
+ - `CONTEXT-MAP.md` — optional bounded-context relationships for projects that need domain mapping.
762
+
763
+ `yoke retrofit` scaffolds the four core files non-destructively; your edits are never overwritten.
764
+ It reports `CONTEXT-MAP.md` when the optional file already exists.
765
+ The loop reads them into every agent + reviewer prompt and logs decisions back on each story's
766
+ commit. Decision history is explicitly delimited as untrusted reference data, so stored text is
767
+ never treated as fresh instructions. Manage the files directly with `yoke context init` and `yoke context status`. The
768
+ `maintaining-context` skill teaches agents to honour the same files during interactive work.
769
+
770
+ > Commit `.yoke/context/` to git. The `--isolate` loop runs each iteration in a worktree
771
+ > checked out from HEAD, so it only sees committed context.
772
+
773
+ ## 🛡️ Safety model
774
+
775
+ Yoke's guardrails are **mechanical, not advisory** — the loop blocks on a dirty worktree, missing acceptance criteria, red tests, or a reviewer rejection, and **none of them rely on the agent choosing to behave**.
776
+
777
+ - **Commit integrity** — a story is never recorded `passes: true` without a corresponding commit; a failed commit reverts the PRD.
778
+ - **Role separation** — the implementer never reviews its own work; `--reviewer` can even be a different agent.
779
+ - **Isolation** — with `--isolate`, failed or partial work is discarded with the worktree and never reaches your main tree.
780
+ - **Non-destructive retrofit** — existing files are backed up before any change; settings are merged, not replaced.
781
+ - **Independent verification** — "done" means *your test command exits 0*, not "the agent said so".
782
+ - **Single-flight** — a lock prevents two loops from racing the same repo; `yoke loop cleanup` recovers after crashes.
783
+
784
+ ## 🧠 Choose your code-graph
785
+
786
+ `yoke retrofit --code-graph=graphify|serena` (default `graphify`, remembered per project). The `yoke-retrofit` skill asks and recommends based on the project.
787
+
788
+ | | **graphify** | **Serena** |
789
+ |---|---|---|
790
+ | Engine | tree-sitter AST + graph | real language servers (LSP) |
791
+ | Strength | fast, multimodal (code + PDFs + images) | symbol-exact cross-file refactoring |
792
+ | Token efficiency | ~70× reduction on large mixed repos | standard, no index to go stale |
793
+ | Best for | rapid exploration / migration / onboarding | systematic refactoring in typed codebases |
794
+ | Caveat | heuristic edges; static index can go stale | one language server per language |
795
+
796
+ ## 🪙 Token efficiency
797
+
798
+ Yoke attacks tokens on two complementary surfaces:
799
+
800
+ - **rtk** compresses noisy command/tool output before it enters context (wired as a hook/instruction per agent).
801
+ - The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
802
+
803
+ A Codex-only full-repository study ran three alternating-order pairs per arm. The same Sol parent
804
+ completed all 12 stories and all **36/36 hidden acceptance checks**; with routing enabled, a Sol
805
+ controller selected Luna for every bounded implementation story. Including controller overhead,
806
+ the routed median used **33.8% less wall time, 11.0% less fresh input, 49.5% fewer output tokens,
807
+ and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
808
+
809
+ The boundary matters: an earlier architecture/privacy task correctly stayed on `SELF` and paid
810
+ controller overhead, so routing is an explicit opt-in rather than a universal win. Codex did not
811
+ emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
812
+ inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
813
+ [`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
814
+
815
+ ## 🧩 Optional companions
816
+
817
+ Two external tools pair well with Yoke and are documented (not bundled) in `canon/tools/` — each has its own installer and update cadence, so Yoke wires the boundary instead of vendoring a copy:
818
+
819
+ - **[claude-mem](https://github.com/thedotmack/claude-mem)** — persistent cross-session memory for interactive work. Deliberate boundary: the autonomous loop keeps its memory **explicit and versioned** (`context/*.md` + PRD, fresh context per story), so claude-mem's automatic injection stays out of loop runs.
820
+ - **[ui-ux-pro-max](https://github.com/nextlevelbuilder/ui-ux-pro-max-skill)** — data-driven design intelligence (styles, palettes, industry rules) for the *generation* side. Yoke's `unslop-ui`, `design-scan`, and `visual-verification` remain the *verification* side: generate with pro-max, gate with Yoke.
821
+
822
+ ## 🌱 Why & how it was built
823
+
824
+ **The problem.** Coding agents are powerful, but each speaks its own dialect — Claude has skills and hooks, Codex reads `AGENTS.md` and a TOML config, Gemini wants commands and a settings file. Keeping the same skills, safety policy, and tool wiring consistent across all of them means copy-paste drift and three things to maintain. Yoke exists to keep **one** source of truth, generate the right native artifacts for each agent, and let that harness run **autonomously and safely** when you want to hand it a spec and walk away.
825
+
826
+ **The inspiration.** Yoke is a synthesis of ideas already proven across the ecosystem: composable-skills methodology ([superpowers](https://github.com/obra/superpowers), [gstack](https://github.com/garrytan/gstack)); the portable [AGENTS.md](https://agents.md/) standard; the *"one source-of-truth → idiomatic per-harness artifacts"* generation pattern ([wshobson/agents](https://github.com/wshobson/agents)); spec-driven autonomous orchestration (GSD); mechanical safety gates and role separation (safe-agentic-workflow); and the **Ralph loop** (Geoff Huntley) — keep handing a *fresh* agent the next task until the spec is done. Token efficiency comes from [rtk](https://github.com/rtk-ai/rtk) and the write-less-code idea behind [ponytail](https://github.com/DietrichGebert/ponytail).
827
+
828
+ **How it was built.** Yoke was built the way it's meant to be *used* — agent-driven, incremental, and test-first. The stack was chosen by **researching alternatives first** (which is how `jcodemunch` was dropped for its license and Serena was added as an option). Then every component shipped one small piece at a time through a disciplined loop: **brainstorm → spec → plan → TDD implementation → an independent two-stage review** (does it match the spec? is it well-built?) **→ merge**. Those reviews caught real bugs before they shipped — a Windows `.cmd` spawn failure, a commit-integrity hole, a path-traversal that could delete project data, a resolution bug that broke the CLI's default invocation, a TOML-escaping bug. Yoke was even **dogfooded on its own repo**, which surfaced (and fixed) a genuine Windows bug. Every spec and plan lives in [`docs/superpowers/`](docs/superpowers/).
829
+
830
+ ## 🗂️ Project layout
831
+
832
+ ```text
833
+ canon/ # the source of truth — harness-agnostic
834
+ AGENTS.md skills/ policy/ loop/ tools/ manifest.yaml
835
+ src/
836
+ canon/ # manifest schema + validator (yoke validate)
837
+ change/ # append-only change inbox · planning · independent coverage review
838
+ retrofit/ # detect · plan · apply · planners (claude/codex/gemini) · tools
839
+ loop/ # prd · gates · runner · verify · git/worktree · loop · run-command · lock · cleanup
840
+ quality/ # reference collection · blind critic · bounded repair · candidate comparison
841
+ new/ # yoke new — greenfield bootstrap
842
+ prd/ # yoke prd draft|check — idea → stories + lint gate
843
+ review/ # yoke review — cross-model diff gate
844
+ smoke/ # yoke flow-smoke — browser gate with screenshot/video proofs
845
+ scan/ # yoke design-scan — AI-slop design gate
846
+ context/ # the durable context layer
847
+ docs/superpowers/ # the spec and every component's implementation plan
848
+ ```
849
+
850
+ ## 🗺️ Roadmap
851
+
852
+ Completed release work lives in the changelog. Remaining, explicitly scoped work is tracked in
853
+ [`TODOS.md`](TODOS.md), including broader benchmark samples, native output schemas, and signed
854
+ release provenance.
855
+
856
+ ## 🧪 Development
857
+
858
+ ```bash
859
+ npm test # vitest (1019 tests)
860
+ npm run build # tsc, no emit errors
861
+ npm run yoke -- validate canon
862
+ ```
863
+
864
+ ## 🙏 Credits & inspiration
865
+
866
+ Yoke stands on the shoulders of a great ecosystem: methodology ideas from [superpowers](https://github.com/obra/superpowers) and [gstack](https://github.com/garrytan/gstack); the [AGENTS.md](https://agents.md/) standard; the generator pattern from [wshobson/agents](https://github.com/wshobson/agents); the Ralph autonomous-loop pattern; safety-gate thinking from safe-agentic-workflow; and the wired tools [rtk](https://github.com/rtk-ai/rtk), [graphify](https://github.com/safishamsi/graphify), [Serena](https://github.com/oraios/serena), and [Playwright MCP](https://github.com/microsoft/playwright-mcp). The `minimal-code` skill adapts the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.
867
+
868
+ ## 📄 License
869
+
870
+ MIT — see [`LICENSE`](LICENSE).
871
+
872
+ <div align="center">
873
+ <sub>Built with a disciplined loop: brainstorm → spec → plan → TDD → two-stage review → merge — and reviewed by a second model, because we don't trust "done" either.</sub>
874
+ </div>