@hecer/yoke 1.18.0 → 1.19.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/CHANGELOG.md +19 -0
- package/README.md +83 -910
- package/canon/manifest.yaml +1 -1
- package/dist/agents/process.js +42 -11
- package/dist/agents/providers.js +11 -2
- package/dist/agents/sol-pi-runtime.js +76 -0
- package/dist/code-intelligence/adapters/graphify.js +1 -1
- package/dist/code-intelligence/adapters/mcp.js +25 -4
- package/dist/code-intelligence/adapters/serena.js +1 -1
- package/dist/code-intelligence/coordinator.js +2 -1
- package/dist/code-intelligence/evidence.js +12 -1
- package/dist/code-intelligence/mcp-client.js +2 -2
- package/dist/dashboard/page.js +1 -1
- package/dist/dashboard/panels.js +25 -3
- package/dist/dashboard/server.js +62 -1
- package/dist/retrofit/config.js +28 -0
- package/docs/HARNESSES.md +1 -1
- package/docs/SOL-PI.md +41 -0
- package/docs/code-intelligence-stdio-fix-2026-09-27.md +8 -0
- package/gemini-extension.json +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -1,970 +1,143 @@
|
|
|
1
1
|
<div align="center">
|
|
2
2
|
|
|
3
|
-
<
|
|
3
|
+
<img src="https://raw.githubusercontent.com/HECer/yoke/main/docs/assets/yoke-logo.png" alt="Yoke" width="104">
|
|
4
4
|
|
|
5
|
-
|
|
6
|
-
<!-- yoke:tests:start -->1313<!-- yoke:tests:end -->
|
|
7
|
-
<!-- yoke:skills:start -->34<!-- yoke:skills:end -->
|
|
8
|
-
<!-- yoke:agents:start -->Claude | Codex | Gemini | Qwen | OpenCode | Kilo | Pi | Hermes<!-- yoke:agents:end -->
|
|
9
|
-
|
|
10
|
-
### One harness, eight agents — and zero trust in "done."
|
|
5
|
+
# Yoke
|
|
11
6
|
|
|
12
|
-
|
|
7
|
+
### Autonomous coding across your agents. Proof before done.
|
|
13
8
|
|
|
14
9
|
[](https://www.npmjs.com/package/@hecer/yoke)
|
|
15
|
-
[](https://www.npmjs.com/package/@hecer/yoke)
|
|
16
10
|
[](https://github.com/HECer/yoke/actions/workflows/ci.yml)
|
|
17
|
-
[](
|
|
18
|
-

|
|
19
|
-

|
|
20
|
-

|
|
21
|
-

|
|
22
|
-

|
|
23
|
-
|
|
24
|
-
**Install:** [`npm i -g @hecer/yoke`](https://www.npmjs.com/package/@hecer/yoke)
|
|
25
|
-
|
|
26
|
-
</div>
|
|
27
|
-
|
|
28
|
-
> **TL;DR** — `yoke setup .` asks six questions and installs the native harness for your agent. `yoke new my-app --idea="..."` bootstraps a project and drafts its story backlog. `yoke loop run my-app --isolate --review` then implements it behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. Add `--parallel=N` for dependency-aware workers, or declare a reference and add `--quality` for a bounded critic/repair gauntlet. If any blocking gate is red, nothing is committed. Proof lives in `.yoke/proof/<story>/`.
|
|
29
|
-
|
|
30
|
-
**New in 1.15.0:** federated [Code Intelligence](docs/CODE-INTELLIGENCE.md) composes Graft, Graphify and Serena behind one Yoke-controlled MCP surface, with structural and semantic evidence, content-addressed snapshots, partial-coverage reporting and isolated edit previews. It is opt-in: use `off` for the unchanged legacy path, `shadow` for read-only canaries, or `active` for previews and approved edits. The [1.14.0 dashboard overhaul](docs/DASHBOARD-OVERHAUL.md) and first-class [OpenCode, Kilo and Pi integrations](docs/HARNESSES.md) remain available. See the [changelog](CHANGELOG.md) and [Code Intelligence guide](docs/CODE-INTELLIGENCE.md) for setup and limitations.
|
|
31
|
-
|
|
32
|
-
OpenCode, Kilo and Pi are real CLI integrations, not bundled runtimes or credentials. OpenCode/Kilo use their JSON headless modes and local MCP configuration; Pi uses JSONL and explicit tool allowlists, but has no native MCP, sub-agent or plan layer. Read the [integration guide](docs/HARNESSES.md) before selecting a permission profile.
|
|
33
|
-
|
|
34
|
-
**Fixed in 1.15.1:** multi-turn Pi usage accounting, provider error/progress handling, task-specific time ranges and missing skill resources. Pi has been supported since **1.13.0**; current Pi versions require explicit project trust for project-local skills in headless runs. The [agent/skill hardening audit](docs/AGENT-HARDENING-2026-09-16.md) documents the fixes, migration notes and validation limits.
|
|
35
|
-
|
|
36
|
-
### One dashboard, multiple projects
|
|
11
|
+
[](LICENSE)
|
|
12
|
+

|
|
37
13
|
|
|
38
|
-
|
|
39
|
-
npm install -g @hecer/yoke@latest
|
|
40
|
-
yoke projects add /path/to/frontend
|
|
41
|
-
yoke projects add /path/to/backend
|
|
42
|
-
yoke dashboard --no-register
|
|
43
|
-
```
|
|
14
|
+
<!-- yoke:version:start -->1.19.0<!-- yoke:version:end --> · <!-- yoke:tests:start -->1339<!-- yoke:tests:end --> test cases · <!-- yoke:skills:start -->34<!-- yoke:skills:end --> skills
|
|
44
15
|
|
|
45
|
-
|
|
16
|
+
<!-- yoke:agents:start -->Claude | Codex | Gemini | Qwen | OpenCode | Kilo | Pi | Hermes<!-- yoke:agents:end -->
|
|
46
17
|
|
|
47
|
-
|
|
18
|
+
</div>
|
|
48
19
|
|
|
49
|
-
|
|
20
|
+
Yoke turns a goal into acceptance-tested stories, coordinates one or more coding agents, and commits work only after the configured checks pass. Run a normal backlog to completion, or optionally let Yoke explore for evidence-backed improvements and continue the same verified loop.
|
|
50
21
|
|
|
51
|
-
|
|
22
|
+
## How it works
|
|
52
23
|
|
|
53
|
-
```
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
24
|
+
```mermaid
|
|
25
|
+
flowchart LR
|
|
26
|
+
A[Goal and acceptance criteria] --> B[PRD stories and dependencies]
|
|
27
|
+
B --> C[Agent workers in isolated worktrees]
|
|
28
|
+
C --> D{Acceptance tests and configured gates}
|
|
29
|
+
D -- needs work --> C
|
|
30
|
+
D -- passes --> E[Commit and local proof]
|
|
31
|
+
E -. explore enabled .-> F[Evidence-backed next tasks]
|
|
32
|
+
F --> B
|
|
60
33
|
```
|
|
61
34
|
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
See the [1.7 workflow guide](docs/VERIFIED-PROJECTS.md) for setup, Gemini selection, recovery and budgets. Provider contracts do not establish equal model quality; live comparative savings and calibrated time predictions remain unmeasured. Token budgets apply between provider calls, and browser proofs still require a configured smoke gate.
|
|
65
|
-
|
|
66
|
-
Yoke 1.5 keeps failed gate output compact without throwing evidence away: deterministic previews
|
|
67
|
-
retain actionable failures and final summaries, while large complete stdout/stderr remains available
|
|
68
|
-
in private, content-addressed local artifacts. Existing projects keep their serial behavior and use
|
|
69
|
-
safe 2 KiB preview / 8 KiB artifact defaults unless configured otherwise.
|
|
70
|
-
|
|
71
|
-
Yoke 1.4 introduced opt-in parallel workers and a bounded, reference-driven quality gauntlet.
|
|
72
|
-
Automatic parallelism was introduced in Yoke 1.8.0. This unreleased branch adds a shared cross-project worker pool. See [the 1.4 migration guide](docs/MIGRATING-TO-1.4.md)
|
|
73
|
-
for the new flags, configuration, cleanup behavior, and review-verdict contract.
|
|
35
|
+
Each story carries observable acceptance criteria and targeted test commands. Parallel workers take independent stories; Yoke checks their integrated result before committing. Proof is saved under `.yoke/proof/`.
|
|
74
36
|
|
|
75
|
-
|
|
76
|
-
is explicit; reviews require a schema-valid verdict and a different model unless
|
|
77
|
-
`--allow-self-review` is explicit; commits enforce the human identity from project config or Git.
|
|
78
|
-
See [the 1.1 migration guide](docs/MIGRATING-TO-1.1.md) for setup/decision parity and
|
|
79
|
-
[the 1.0 guide](docs/MIGRATING-TO-1.0.md) for the earlier safety-policy changes.
|
|
37
|
+
## What Yoke adds
|
|
80
38
|
|
|
81
|
-
|
|
39
|
+
| Capability | What it does for you |
|
|
40
|
+
| --- | --- |
|
|
41
|
+
| **Verified completion** | Checks acceptance criteria, the project verify command, and any configured completion, review, quality, or browser gates before accepting work. |
|
|
42
|
+
| **Parallel execution** | Schedules independent stories in isolated worktrees, respects dependencies and write scopes, and coordinates a shared worker limit across projects. |
|
|
43
|
+
| **Long-running autonomy** | Recovers from provider failures and blocked work. Optional exploration discovers new, repository-evidenced tasks after the planned backlog drains. |
|
|
44
|
+
| **A choice of agents** | Uses the native CLI for Claude, Codex, Gemini, Qwen, OpenCode, Kilo, Pi, or Hermes. Install and authenticate the CLI you choose; Yoke does not bundle model runtimes or credentials. |
|
|
45
|
+
| **Project visibility** | A local dashboard shows project status, **Workspace analytics**, **History**, and controls such as **Queue a change**, safe-boundary pause/resume, and operator notes. It runs with the local Yoke process. Dark and light themes are available. |
|
|
46
|
+
| **Optional Pi efficiency** | An opt-in SoL-Pi integration exposes per-project mechanism settings for Pi. It is off by default; benchmark results are not a savings guarantee. |
|
|
82
47
|
|
|
83
|
-
##
|
|
48
|
+
## Quick start
|
|
84
49
|
|
|
85
|
-
|
|
86
|
-
setups use your configured model. To add DeepSeek and Kimi API model profiles:
|
|
50
|
+
Requires Node.js 20+ and Git. Install Yoke and create a project with a draft backlog:
|
|
87
51
|
|
|
88
52
|
```sh
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
Set `DEEPSEEK_API_KEY` and `MOONSHOT_API_KEY` in the environment. Presets include
|
|
93
|
-
DeepSeek V4 Flash/Pro and Kimi K2.6/K2.7 Code/K3, executed through Qwen Code.
|
|
94
|
-
See [setup, permissions, reasoning configuration and validation limits](docs/QWEN-MODEL-SUPPORT.md).
|
|
95
|
-
|
|
96
|
-
## Why Yoke exists
|
|
97
|
-
|
|
98
|
-
Agentic coding in 2026 fails in four well-documented ways. Yoke answers each one **mechanically** — in code, not in a prompt the agent can ignore:
|
|
99
|
-
|
|
100
|
-
| The pain | What actually happens | What Yoke does about it |
|
|
101
|
-
|---|---|---|
|
|
102
|
-
| 🎭 **The verification gap** — *"agent says done, but it isn't"* | A success message can omit untested acceptance criteria. | The loop executes acceptance and project checks. Enabled review and browser gates must pass before a story lands. `yoke check` exposes unmapped outcomes as unverified. |
|
|
103
|
-
| 🔀 **Four agents, four configs** | Teams hand-maintain `CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, skills, and MCP wiring separately — copy-paste drift everywhere | **One canon → `yoke retrofit`** generates the idiomatic native artifacts for each agent. Change the canon once, re-retrofit everywhere. |
|
|
104
|
-
| 🌀 **Overnight loops going off the rails** | Raw Ralph-loop users "wake up to broken codebases that don't compile" | Yoke is **"Ralph, but with gates"**: clean-worktree gate, acceptance-criteria gate, green-tests gate, review gate, per-story worktree isolation, idle-timeout watchdog, single-flight lock, commit integrity. |
|
|
105
|
-
| 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a second model writes a schema-validated pass/fail verdict — chainable into verify, pre-push, or CI. Cross-model review catches what self-review misses. |
|
|
106
|
-
|
|
107
|
-
**Who it's for:** anyone driving Claude Code, Codex CLI, Gemini CLI, Qwen Code, OpenCode, Kilo, Pi or Hermes on real projects — especially if you use more than one, want autonomous runs you can trust, or are tired of "done" meaning "probably". Greenfield (`yoke new`) and brownfield (`yoke retrofit`) both work.
|
|
108
|
-
|
|
109
|
-
**Who it's not for:** if you want a chat pair-programmer with no process, you don't need a harness. Yoke is for shipping with discipline.
|
|
110
|
-
|
|
111
|
-
## Example: idea → verified implementation
|
|
112
|
-
|
|
113
|
-
```console
|
|
114
|
-
$ yoke new reading-app --idea="a web app that tracks my reading list"
|
|
115
|
-
✓ reading-app bootstrapped. # git repo · harness for all agents · context · PRD drafted from the idea
|
|
116
|
-
|
|
117
|
-
$ yoke prd check reading-app
|
|
118
|
-
✓ PRD valid — 8 stories, 0 pass
|
|
119
|
-
|
|
120
|
-
$ yoke loop on reading-app
|
|
121
|
-
$ yoke loop run reading-app --isolate --review --max=10
|
|
122
|
-
▶ STORY-1 (0/8 · 0%) — implementing… · verifying… · reviewing… · committing…
|
|
123
|
-
✓ STORY-1 done in 3m12s — 1/8 (13%) · ~22m left
|
|
124
|
-
▶ STORY-2 (1/8 · 13%) — implementing… · ~22m left (Ø 3m12s/story)
|
|
125
|
-
✓ STORY-2 done in 2m48s — 2/8 (25%) · ~18m left
|
|
126
|
-
▶ STORY-3 (2/8 · 25%) — implementing… ✘ blocked: story did not verify (tests red)
|
|
127
|
-
# nothing was committed. fix, then re-run.
|
|
128
|
-
|
|
129
|
-
$ ls reading-app/.yoke/proof/STORY-2/
|
|
130
|
-
home.png list.png # example when browser smoke is configured
|
|
131
|
-
```
|
|
132
|
-
|
|
133
|
-
This is an illustrative transcript. Actual durations depend on the project, provider, retries and enabled gates; browser proof requires a configured smoke flow. See [how it was built](#-why--how-it-was-built).
|
|
134
|
-
|
|
135
|
-
## 🚀 Quickstart
|
|
136
|
-
|
|
137
|
-
```bash
|
|
138
|
-
npm install -g @hecer/yoke # → global `yoke` on your PATH
|
|
139
|
-
# (or from source: git clone https://github.com/HECer/yoke.git && cd yoke && npm install && npm run build && npm link)
|
|
140
|
-
|
|
141
|
-
# Greenfield: idea → loop-ready project in one command
|
|
142
|
-
yoke new my-app --idea="a CLI that tracks reading lists"
|
|
143
|
-
yoke loop on my-app && yoke loop run my-app --isolate
|
|
144
|
-
|
|
145
|
-
# — or retrofit an existing project —
|
|
146
|
-
yoke setup /path/to/project # interactive: agents, graph, loop, runner, decisions, routing
|
|
147
|
-
yoke validate canon # sanity-check the canon
|
|
148
|
-
yoke loop run /path/to/project --isolate --parallel=3 --reviewer=codex --max=20
|
|
149
|
-
```
|
|
150
|
-
|
|
151
|
-
> Requires Node ≥ 20 and git. No global install? `node /path/to/yoke/dist/cli.js …` or `npm --prefix /path/to/yoke run yoke -- …` work too. The MCP tools (rtk, graphify/Serena, Playwright MCP) are wired by Yoke but installed separately — the generated config is a clearly-labelled, adjustable template.
|
|
152
|
-
|
|
153
|
-
### Skills before the first setup
|
|
154
|
-
|
|
155
|
-
The canon is also packaged as a Claude Code plugin — the repo is its own marketplace:
|
|
156
|
-
|
|
157
|
-
```text
|
|
158
|
-
/plugin marketplace add HECer/yoke
|
|
159
|
-
/plugin install yoke@yoke
|
|
160
|
-
```
|
|
161
|
-
|
|
162
|
-
That gives you all canon skills under the `yoke:` namespace (e.g. `yoke:tdd`, `yoke:review`) inside Claude Code — no retrofit needed. The `yoke` CLI (loop, gates, retrofit for Codex/Gemini/Qwen/OpenCode/Kilo/Pi/Hermes) still comes from `npm i -g @hecer/yoke`. Gemini CLI users can likewise `gemini extensions install https://github.com/HECer/yoke`.
|
|
163
|
-
|
|
164
|
-
For Codex, no preinstalled skill is required: run `npx @hecer/yoke setup .` in a terminal, or
|
|
165
|
-
ask Codex to run the six-question Yoke setup flow. The retrofit writes native skills to
|
|
166
|
-
`.agents/skills/`, including `yoke-retrofit` and `yoke-workflow`; start a fresh Codex task if an
|
|
167
|
-
already-open task does not discover newly installed skills. The npm package also contains
|
|
168
|
-
`.codex-plugin/plugin.json` for Codex plugin hosts.
|
|
169
|
-
|
|
170
|
-
### Staying up to date
|
|
171
|
-
|
|
172
|
-
Yoke checks for new releases npm/gh-style: a **non-blocking background check** (at most once a day, detached, offline-safe) prints a one-line hint when a newer version exists — upgrading itself is always an explicit act:
|
|
173
|
-
|
|
174
|
-
```bash
|
|
175
|
-
yoke upgrade # npm install -g @hecer/yoke@latest
|
|
176
|
-
```
|
|
177
|
-
|
|
178
|
-
Disable the check with `YOKE_NO_UPDATE_CHECK=1` (it is also silent in CI, `--json` runs, and piped output). Projects that want the loop to self-update can opt in via `.yoke/config.yaml`:
|
|
179
|
-
|
|
180
|
-
```yaml
|
|
181
|
-
update:
|
|
182
|
-
auto: true # upgrade at loop START only — never mid-run; applies from the next invocation
|
|
183
|
-
```
|
|
184
|
-
|
|
185
|
-
Auto-upgrade is deliberately **not** the default: a gate harness shouldn't change itself mid-project, and unreviewed auto-installs are a supply-chain hazard.
|
|
186
|
-
|
|
187
|
-
## 🤖 Driving it through an agent
|
|
188
|
-
|
|
189
|
-
Yoke is meant to be operated *by* your coding agent — after a retrofit, the agent has the skills, the safety policy, and the routing, so it knows the methodology. Copy-paste prompts (identical wording works for Claude Code, Codex CLI, Gemini CLI, Qwen Code, OpenCode, Kilo, Pi, and Hermes):
|
|
190
|
-
|
|
191
|
-
> **Set it up** — *"Set up Yoke in this project. Ask me the Yoke setup questions one at a time with your recommendation, then run `yoke setup . --yes` with the selected host, agents, code graph, loop, runner, and decision policy. Commit in my configured identity."*
|
|
192
|
-
|
|
193
|
-
> **Work the disciplined way** — *"From now on follow the Yoke skills you just installed: brainstorm → spec → plan → TDD → review before merging. Use the `review` skill before any merge."*
|
|
194
|
-
|
|
195
|
-
> **Plan, then run autonomously** — *"Use the `yoke-workflow` skill. Ask only the planning questions that materially change the product, write the approved plan and loop-ready stories, then execute every approved story without routine follow-ups. Follow the configured `auto` or `critical` decision policy."*
|
|
196
|
-
|
|
197
|
-
> **Watch / unblock** — *"Run `yoke loop status .`. If it says BLOCKED, run the project's verify command, find the root cause, fix it without weakening tests, then continue the loop."*
|
|
198
|
-
|
|
199
|
-
> ⚠️ **Long runs from inside an agent session:** `yoke loop run` has no story cap by default; it continues until every planned story passes or a gate blocks. A multi-story run can therefore outlive most agents' shell-tool timeouts (Claude Code's Bash tool defaults to 2 minutes). If the outer tool call is killed mid-run, you get a stale lock and possibly half-finished state — which *looks* like a hang. Run the loop **in the background** (e.g. Claude Code's `run_in_background`), use `--max=3..5` only when you intentionally want a bounded batch, poll with `yoke loop status`, and after any interrupted run do `yoke loop cleanup` before the next one. A `running` status with no update for 20+ minutes on a claude runner is worth checking — since 0.5.0 the runner streams continuously, so prolonged true silence is no longer normal.
|
|
200
|
-
|
|
201
|
-
### Optional continuous exploration
|
|
202
|
-
|
|
203
|
-
`yoke loop run . --explore` keeps the supervisor alive after the PRD drains. It periodically scans
|
|
204
|
-
for evidence-backed, testable improvements, adds only bounded tasks that pass strict contract checks,
|
|
205
|
-
then implements them through the same isolated loop gates. `--explore-interval=10` changes the
|
|
206
|
-
default 30-minute rescan interval. Exploration runs indefinitely by default; set `--explore-limit=12h`,
|
|
207
|
-
`--explore-limit=3d`, or `--explore-limit=2w` to stop automatically after hours, days, or weeks. At
|
|
208
|
-
expiry, Yoke pauses at a safe task boundary, lets active workers finish their gates and integration,
|
|
209
|
-
and exits with code `3`; run it again to resume. `yoke loop pause .` also stops at a safe boundary.
|
|
210
|
-
Failed providers or stories are retried with backoff, and live status heartbeats while the supervisor
|
|
211
|
-
waits. Use `--max=N` for an intentional story-attempt cap. See the [continuous exploration guide](docs/CONTINUOUS-EXPLORATION.md)
|
|
212
|
-
for task validation, history compaction, recovery and process-lifetime limits.
|
|
213
|
-
|
|
214
|
-
> ⚠️ **Never kill agent processes by name or command-line pattern** (e.g. every process matching `dangerously-skip-permissions`): on a machine running several yoke projects, that takes down the *healthy* runners of the other projects mid-story — they stall and their loops block. `yoke loop cleanup` is the scoped alternative: each watchdog records its pids in the project's `.yoke/runner.pid`, and cleanup kills exactly those recorded trees — nothing else on the machine.
|
|
215
|
-
|
|
216
|
-
### Agent cheat sheet — every command is an exit-code contract
|
|
217
|
-
|
|
218
|
-
Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`) can branch on exit codes without parsing prose.
|
|
219
|
-
|
|
220
|
-
| Command | What it does | Exit codes |
|
|
221
|
-
|---|---|---|
|
|
222
|
-
| `yoke dashboard [dir] [--no-register] [--port=N]` | Local overview for every registered project; optionally register `dir` first | `0` stopped normally · `1` shutdown failure · `2` unavailable |
|
|
223
|
-
| `yoke projects add\|list\|remove` | Register a project, list registrations or remove a reference by ID | `0` · `2` invalid/unavailable |
|
|
224
|
-
| `yoke check [dir] [--json] [--requirement=] [--protect [--refresh]]` | Execute acceptance checks or explicitly pin their infrastructure | `0` passed/pinned · `1` failed · `2` unverified/unavailable |
|
|
225
|
-
| `yoke goal set\|run\|resume\|pause\|status\|handoff\|budget [dir]` | Durable objectives, provider handoff, protected checks and checkpoint budgets | run/resume: `0` complete · `1` unfinished · `2` unavailable |
|
|
226
|
-
| `yoke setup [dir] [--yes] [--host=] [--agent=] [--runner=] [--code-graph=] [--code-intelligence=off\|shadow\|active] [--decision-policy=] [--loop\|--no-loop] [--routing\|--no-routing] [--model-provider=deepseek,kimi]` | Shared setup for all eight harnesses; optional federated code intelligence and DeepSeek/Kimi API profiles | `0` · `1` invalid setup |
|
|
227
|
-
| `yoke code-intelligence-server [--workspace=] [--mode=off\|shadow\|active]` | Serve the single Yoke-controlled MCP facade for federated code intelligence | `0` · `1` invalid/unavailable |
|
|
228
|
-
| `yoke validate [canonDir]` | Validate the canon (schema, frontmatter, templates) | `0` valid · `1` errors |
|
|
229
|
-
| `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
|
|
230
|
-
| `yoke retrofit [dir] [--agent=claude,codex,gemini,qwen,opencode,kilo,pi,hermes\|all] [--code-graph=graphify\|serena] [--code-intelligence=off\|shadow\|active] [--loop]` | Install/update the harness for the selected agents, non-destructively | `0` |
|
|
231
|
-
| `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
|
|
232
|
-
| `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
|
|
233
|
-
| `yoke change add\|status [dir] [--idea=]` | Queue a change at any time; the loop turns it into append-only stories at the next safe boundary | `0` · `1` invalid inbox/request |
|
|
234
|
-
| `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE/GLOSSARY.md`, optional `CONTEXT-MAP.md`) | `0` |
|
|
235
|
-
| `yoke loop on\|off\|status\|pause\|decision\|answer\|resume\|run\|cleanup [dir]` | Autonomous loop; `run --explore` opts into continuous discovery, implementation and recovery after a backlog drains; `--parallel=N`, bounded reference-driven `--quality`, and blind `--candidates=N` selection remain available; `--max=N` is an intentional cap; `pause` requests a safe-boundary stop | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
|
|
236
|
-
| `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
|
|
237
|
-
| `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
|
|
238
|
-
| `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
|
|
239
|
-
| `yoke flow-smoke [dir] [--url=] [--label=]` | Browser gate with screenshot/video proofs | `0` green · `1` failures · `2` not runnable |
|
|
240
|
-
|
|
241
|
-
A genuinely hung agent self-terminates after the idle timeout (default 20 min; `--timeout`), and `yoke loop status` shows the live phase or a `⚠ possibly stuck` hint — an autonomous run is never a black box.
|
|
242
|
-
|
|
243
|
-
## ⚖️ How it compares — superpowers · gstack · Yoke
|
|
244
|
-
|
|
245
|
-
Three excellent projects, three different jobs. Honest version:
|
|
246
|
-
|
|
247
|
-
| | [superpowers](https://github.com/obra/superpowers) (obra) | [gstack](https://github.com/garrytan/gstack) (Garry Tan) | **Yoke** |
|
|
248
|
-
|---|---|---|---|
|
|
249
|
-
| **What it is** | The canonical *skills methodology*: brainstorm → plan → TDD → review as composable skills | A *software factory* for Claude Code: ~40 role skills (QA, CSO, ship…) + a real Chromium browser layer | A *cross-agent harness*: one canon → native installs, plus a gated autonomous loop |
|
|
250
|
-
| **Agents** | Claude Code first | Claude Code + hosts like Codex/Cursor/Kiro — **no Gemini CLI** | **Claude Code, Codex CLI, Gemini CLI, Qwen Code, OpenCode, Kilo, Pi, Hermes** from one source of truth |
|
|
251
|
-
| **Enforcement** | Advisory — skills *describe* the discipline; following them is up to the agent | Skill-driven; browser QA is genuinely real | **Mechanical** — gates live in code: clean tree, acceptance criteria, green tests, review verdict, commit integrity |
|
|
252
|
-
| **Autonomy** | Interactive sessions | Interactive slash-commands (`/qa`, `/ship`, …) | Opt-in **Ralph loop** with watchdog, worktree isolation, single-flight lock, per-story proofs |
|
|
253
|
-
| **Visual QA** | — | **Best-in-class**: live browser daemon (Chromium/CDP) with deep interactive QA | Built-in `flow-smoke` gate: screenshots always, video on failure, labelled per story — lighter, but *enforced* and cross-agent |
|
|
254
|
-
| **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves an independent provider and validates a structured verdict, inside or outside the loop |
|
|
255
|
-
| **Footprint** | Markdown skills (plugin) | ~230 MB with browser runtime; hourly auto-update | Node CLI + markdown canon; Playwright only if you use flow-smoke, resolved **from your project** |
|
|
256
|
-
| **License** | MIT | MIT | MIT |
|
|
257
|
-
|
|
258
|
-
**They compose — use all three where they're strongest.** Yoke's canon *ships* the superpowers methodology natively for all eight agents (13 skills, [attributed](canon/skills/ATTRIBUTION.md)). And if gstack is installed, `yoke retrofit` detects it and adds a routing note to `CLAUDE.md` telling Claude to prefer gstack's live-browser `/qa`, `/cso`, and ship pipeline for what Yoke deliberately doesn't bundle — no dependency, no conflict, and non-Claude artifacts stay uniform.
|
|
259
|
-
|
|
260
|
-
**Choose Yoke when** you run more than one agent, want autonomy you can audit (gates + proofs + logs), or want one place to maintain your team's methodology. **Choose gstack when** you live 100% in Claude Code and want the deepest interactive browser QA. **Choose superpowers when** you want the methodology alone, interactively, in Claude Code — or just use it *through* Yoke.
|
|
261
|
-
|
|
262
|
-
## 🏗️ Architecture
|
|
263
|
-
|
|
264
|
-
You curate **one source of truth** — skills, policy, and tool wiring. Yoke generates the **idiomatic, native artifacts** each agent expects, non-destructively, into any repo:
|
|
265
|
-
|
|
266
|
-
```mermaid
|
|
267
|
-
flowchart TD
|
|
268
|
-
Canon["📦 CANON — single source of truth<br/>skills · policy · loop spec · tool wiring"]
|
|
269
|
-
Skill["🛠️ yoke retrofit<br/>detect → plan → apply (backup) → report"]
|
|
270
|
-
Canon --> Skill
|
|
271
|
-
Skill --> Claude["Claude Code<br/>.claude/skills · .mcp.json · hook"]
|
|
272
|
-
Skill --> Codex["Codex CLI<br/>AGENTS.md · config.toml · RTK.md"]
|
|
273
|
-
Skill --> Gemini["Gemini CLI<br/>GEMINI.md · commands · settings.json"]
|
|
274
|
-
Skill --> Qwen["Qwen Code<br/>QWEN.md · skills · settings.json"]
|
|
275
|
-
Skill --> OpenCode["OpenCode<br/>AGENTS.md · opencode.json · .opencode/skills"]
|
|
276
|
-
Skill --> Kilo["Kilo<br/>AGENTS.md · kilo.jsonc · .kilo/skills"]
|
|
277
|
-
Skill --> Pi["Pi<br/>AGENTS.md · .pi/settings.json · .pi/skills"]
|
|
278
|
-
Skill --> Hermes["Hermes<br/>AGENTS.md · .hermes/skills · yoke-reviewer.md"]
|
|
279
|
-
Loop["🤖 yoke loop — autonomous Ralph loop<br/>gates · verify · review · isolation · proofs"]
|
|
280
|
-
Claude -. drives .-> Loop
|
|
281
|
-
Codex -. drives .-> Loop
|
|
282
|
-
Gemini -. drives .-> Loop
|
|
283
|
-
Qwen -. drives .-> Loop
|
|
284
|
-
OpenCode -. drives .-> Loop
|
|
285
|
-
Kilo -. drives .-> Loop
|
|
286
|
-
Pi -. drives .-> Loop
|
|
287
|
-
Hermes -. drives .-> Loop
|
|
288
|
-
```
|
|
289
|
-
|
|
290
|
-
Three layers — **Canon** (`yoke validate`) → **Retrofit** (`yoke retrofit`) → **Loop** (`yoke loop`) — on top of a durable **Context layer** (`yoke context`).
|
|
291
|
-
|
|
292
|
-
### What gets generated per agent
|
|
293
|
-
|
|
294
|
-
| Agent | Artifacts |
|
|
295
|
-
|---|---|
|
|
296
|
-
| **Claude** | Complete skill packages under `.claude/skills/` (including referenced resources), `AGENTS.md`, `CLAUDE.md`, `.mcp.json` (code-graph + Playwright), and an rtk `PreToolUse` hook when WSL is available |
|
|
297
|
-
| **Codex** | Complete skill packages under `.agents/skills/`, per-skill implicit-invocation policy, `AGENTS.md`, `RTK.md`, `.codex/config.toml`, native hooks, reusable `.codex/agents/*.toml`, and package plugin metadata |
|
|
298
|
-
| **Gemini** | Complete skill packages under `.gemini/skills/`, an auto-invocation index, `GEMINI.md`, `.gemini/commands/*.toml`, and `.gemini/settings.json` (MCP + `AGENTS.md` context) |
|
|
299
|
-
| **Qwen** | Complete skill packages under `.qwen/skills/`, `QWEN.md`, `.qwen/settings.json`, native invocation restrictions, and an RTK PreToolUse retry guard |
|
|
300
|
-
| **OpenCode** | Complete skill packages under `.opencode/skills/`, shared `AGENTS.md`, merged `opencode.json` (instructions + local MCP), and `.opencode/agents/yoke-reviewer.md` |
|
|
301
|
-
| **Kilo** | Complete skill packages under `.kilo/skills/`, shared `AGENTS.md`, merged `kilo.jsonc` (instructions + local MCP), and `.kilo/agents/yoke-reviewer.md` |
|
|
302
|
-
| **Pi** | Complete skill packages under `.pi/skills/`, shared `AGENTS.md`, and merged `.pi/settings.json`; Pi's tool allowlists are applied at invocation time |
|
|
303
|
-
| **Hermes** | Complete skill packages under `.hermes/skills/`, shared `AGENTS.md`, and `.hermes/agents/yoke-reviewer.md` |
|
|
304
|
-
|
|
305
|
-
> **rtk integration:** Claude receives its PreToolUse hook; Codex receives a native hook adapter around `rtk hook check`; Gemini retains instruction-mode fallback where its CLI has no equivalent command-rewrite lifecycle.
|
|
306
|
-
|
|
307
|
-
> **Composes with gstack:** if [gstack](https://github.com/garrytan/gstack) is installed (repo-local or global), `yoke retrofit` adds a short "Composed tools" routing note to **CLAUDE.md only** — telling Claude to prefer gstack's skills for capabilities Yoke doesn't ship (live-browser QA `/qa`, security audit `/cso`, ship/deploy `/ship`). No bundling, no dependency; the note is never written to the Codex or Gemini artifacts.
|
|
308
|
-
|
|
309
|
-
> **Your content survives re-retrofits — preserve blocks:** anything you put between
|
|
310
|
-
> `<!-- yoke:preserve:start -->` and `<!-- yoke:preserve:end -->` in a generated file is
|
|
311
|
-
> carried into the regenerated version on every future `yoke retrofit`. The generated
|
|
312
|
-
> `CLAUDE.md` and `GEMINI.md` ship an empty preserve block scaffold — put your project-specific
|
|
313
|
-
> instructions (tech stack, workflow, `@`-includes) inside it. Works in any yoke-written file;
|
|
314
|
-
> content *outside* the markers is still replaced (and backed up under `.yoke/backup/`).
|
|
315
|
-
|
|
316
|
-
## 🧰 What's in the canon — 34 skills
|
|
317
|
-
|
|
318
|
-
`yoke retrofit` installs all of these into each selected agent natively. Provenance is credited in [`canon/skills/ATTRIBUTION.md`](canon/skills/ATTRIBUTION.md).
|
|
319
|
-
|
|
320
|
-
To stop overlapping skills from auto-invoking against each other, `canon/AGENTS.md` carries a **skill routing & precedence** block (methodology before role; one canonical entrypoint per concern — e.g. pre-merge code review is always `review`), emitted into all eight agents.
|
|
321
|
-
|
|
322
|
-
Each manifest entry also declares `invocation: auto|manual`. Retrofit translates that intent into
|
|
323
|
-
the provider's native controls: Claude and Qwen disable model invocation for manual skills, Codex writes
|
|
324
|
-
`agents/openai.yaml`, Gemini lists only automatic skills in its generated index, and OpenCode/Kilo/Pi/Hermes
|
|
325
|
-
receive complete project-local skill packages. Validation
|
|
326
|
-
rejects conflicting package metadata and broken local Markdown links before anything is installed.
|
|
327
|
-
|
|
328
|
-
**Process / methodology** — *superpowers-derived discipline (13)*
|
|
329
|
-
|
|
330
|
-
| Skill | What it does |
|
|
331
|
-
|---|---|
|
|
332
|
-
| `brainstorming` | Explore intent, requirements & design before any creative work |
|
|
333
|
-
| `writing-plans` | Turn a spec into a bite-sized, TDD implementation plan |
|
|
334
|
-
| `executing-plans` | Execute a written plan in a separate session with review checkpoints |
|
|
335
|
-
| `subagent-driven-development` | Run a plan task-by-task: fresh subagent + two-stage review each |
|
|
336
|
-
| `tdd` | Write the test first, watch it fail, write minimal code, refactor |
|
|
337
|
-
| `systematic-debugging` | Root-cause first — no fix without a confirmed cause |
|
|
338
|
-
| `verification-before-completion` | Prove it actually works before claiming done |
|
|
339
|
-
| `using-git-worktrees` | Isolated worktrees for safe / parallel work |
|
|
340
|
-
| `requesting-code-review` | Request a structured review before merging |
|
|
341
|
-
| `receiving-code-review` | Handle review feedback with rigor, not blind agreement |
|
|
342
|
-
| `dispatching-parallel-agents` | Fan out 2+ independent tasks concurrently |
|
|
343
|
-
| `finishing-a-development-branch` | Merge / PR / cleanup a finished branch |
|
|
344
|
-
| `writing-skills` | Author and verify new skills |
|
|
345
|
-
|
|
346
|
-
**Roles** — *gstack-derived, de-gstacked to be harness-agnostic (7)*
|
|
347
|
-
|
|
348
|
-
| Skill | What it does |
|
|
349
|
-
|---|---|
|
|
350
|
-
| `plan-eng-review` | Architecture / edge-case review of a *plan* |
|
|
351
|
-
| `plan-ceo-review` | Founder-mode scope & ambition review of a plan |
|
|
352
|
-
| `review` | Single canonical pre-merge code review — diff safety + engineering quality (architecture, edge cases, tests, performance) |
|
|
353
|
-
| `ship` | Ship workflow: tests → review → version → changelog → PR |
|
|
354
|
-
| `health` | Code-quality dashboard with a composite score |
|
|
355
|
-
| `retro` | Engineering retrospective from commit history |
|
|
356
|
-
| `document-release` | Post-ship documentation sync (README / CHANGELOG / …) |
|
|
357
|
-
|
|
358
|
-
**Yoke-native** — *authored or adapted for this harness (14)*
|
|
359
|
-
|
|
360
|
-
| Skill | What it does |
|
|
361
|
-
|---|---|
|
|
362
|
-
| `yoke-retrofit` | Set up the Yoke harness in a project (detect → plan → apply) |
|
|
363
|
-
| `yoke-workflow` | Provider-neutral planning questions → approved PRD → autonomous stories → critical-decision resume |
|
|
364
|
-
| `authoring-prd` | Slice a product idea into loop-ready stories with testable acceptance criteria |
|
|
365
|
-
| `minimal-code` | Write the least code that solves the task (YAGNI; ponytail-derived) |
|
|
366
|
-
| `performance` | Efficiency as a measured requirement: benchmarks as tests, budgets as gates, optimizations local + documented |
|
|
367
|
-
| `maintaining-context` | Keep `.yoke/context/` the durable source of truth (the Context layer) |
|
|
368
|
-
| `workflow` | The default order of operations, from idea to deploy |
|
|
369
|
-
| `unslop-ui` | Detect & remove AI-slop design tells (purple gradients, neon glow, emoji-icons…) |
|
|
370
|
-
| `visual-verification` | Widen verify to design-scan + the built-in `yoke flow-smoke` gate (screenshot proofs; video on failure) |
|
|
371
|
-
| `no-ai-slop` | Detect and edit generic AI prose while preserving the author's voice; includes its evaluation rubric |
|
|
372
|
-
| `domain-modeling` | Model boundaries, invariants, vocabulary, context maps, and decision records before implementation |
|
|
373
|
-
| `codebase-design` | Explore architecture, deepen a chosen design, and compare two viable approaches when tradeoffs matter |
|
|
374
|
-
| `resolving-merge-conflicts` | Resolve conflicts by reconstructing intent, then verify the integrated result |
|
|
375
|
-
| `writing-for-agents` | Write compact agent instructions with explicit triggers, constraints, resources, and checks |
|
|
376
|
-
|
|
377
|
-
## 🌱 Zero to 100: `yoke new` + `yoke prd`
|
|
378
|
-
|
|
379
|
-
Yoke's greenfield entrypoint — one command from idea to loop-ready project:
|
|
380
|
-
|
|
381
|
-
```bash
|
|
382
|
-
yoke new my-app --idea="a CLI that tracks reading lists" # scaffold + retrofit + context + PRD
|
|
383
|
-
yoke loop on my-app && yoke loop run my-app --isolate # hand it to the loop
|
|
53
|
+
npm install --global @hecer/yoke
|
|
54
|
+
yoke new my-app --idea="A reading list app" --agent=codex --runner=codex
|
|
384
55
|
```
|
|
385
56
|
|
|
386
|
-
`
|
|
387
|
-
existing projects), then: creates and `git init`s the directory, writes a minimal scaffold
|
|
388
|
-
(`README.md`, `.gitignore`), runs the full **retrofit** (`--agent=` as usual), initialises the
|
|
389
|
-
**context layer** (with `--idea` seeded into `PROJECT.md` as the north star), writes a commented
|
|
390
|
-
**PRD template** to `.yoke/prd.yaml`, and makes the initial commit — so `--isolate` works from
|
|
391
|
-
iteration 1. With `--idea`, it then drafts the PRD from your idea via an agent (`--runner=`,
|
|
392
|
-
the configured runner or active host) and commits it as a second commit (`docs: draft PRD from idea`).
|
|
57
|
+
Set `verify.command` in `my-app/.yoke/config.yaml` to the project's real test command, then run:
|
|
393
58
|
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
|
|
397
|
-
**`yoke prd draft [dir] --idea="..."`** turns an idea into 5–12 small, independently shippable
|
|
398
|
-
stories with testable behavioral acceptance criteria (greenfield STORY-1 scaffolds the project
|
|
399
|
-
skeleton + test suite and wires `verify.command`). An existing PRD with stories is never
|
|
400
|
-
overwritten without `--force`; the untouched template doesn't trigger the guard. Runs through
|
|
401
|
-
the same idle-timeout watchdog as the loop (`--timeout`). If `.yoke/plan.md` exists, its approved
|
|
402
|
-
goals, non-goals, constraints, and decisions are injected as settled context instead of being
|
|
403
|
-
reopened by the drafting agent.
|
|
404
|
-
|
|
405
|
-
**`yoke prd check [dir]`** is the chainable pre-loop lint gate: schema validation plus
|
|
406
|
-
duplicate-id, empty-acceptance, unresolved-placeholder, and zero-stories checks. Exits `0` with
|
|
407
|
-
`✓ PRD valid — N stories, M pass`, `1` on any violation. The `authoring-prd` canon skill
|
|
408
|
-
teaches interactive sessions the same story-slicing discipline.
|
|
409
|
-
|
|
410
|
-
## 🤖 The autonomous loop
|
|
411
|
-
|
|
412
|
-
Opt-in; `yoke setup` recommends enabling it for new installs, while `retrofit` alone keeps it off unless requested. Each iteration starts a **fresh agent** and passes through hard gates before anything is committed:
|
|
413
|
-
|
|
414
|
-
```mermaid
|
|
415
|
-
flowchart LR
|
|
416
|
-
I[consume queued change<br/>as new stories] --> A[pick next PRD story]
|
|
417
|
-
A --> B{clean worktree?}
|
|
418
|
-
B -- no --> X[blocked]
|
|
419
|
-
B -- yes --> C{acceptance<br/>criteria?}
|
|
420
|
-
C -- no --> X
|
|
421
|
-
C -- yes --> D[agent implements<br/>one story]
|
|
422
|
-
D --> E{suite + criterion<br/>proof green?}
|
|
423
|
-
E -- no --> X
|
|
424
|
-
E -- yes --> V{UI design<br/>within budget?}
|
|
425
|
-
V -- no --> X
|
|
426
|
-
V -- yes --> F{reviewer<br/>approves?}
|
|
427
|
-
F -- no --> X
|
|
428
|
-
F -- yes --> G[commit + mark passes:true<br/>+ proof in .yoke/proof/]
|
|
429
|
-
G --> I
|
|
430
|
-
I --> H{all stories pass?}
|
|
431
|
-
H -- yes --> J{integrated system<br/>gate green?}
|
|
432
|
-
J -- no --> X
|
|
433
|
-
J -- yes --> K[current backlog ready]
|
|
59
|
+
```sh
|
|
60
|
+
yoke loop on my-app
|
|
61
|
+
yoke loop run my-app --isolate --parallel=auto
|
|
434
62
|
```
|
|
435
63
|
|
|
436
|
-
|
|
437
|
-
yoke loop on . # enable (recorded in .yoke/config.yaml)
|
|
438
|
-
yoke loop status . # show state + PRD progress
|
|
439
|
-
yoke change add . --idea="Add passkey login" # safe while the loop runs
|
|
440
|
-
yoke loop run . \
|
|
441
|
-
--runner=codex \ # implement with Codex…
|
|
442
|
-
--reviewer=claude \ # …review with Claude (role separation)
|
|
443
|
-
--isolate \ # each story in a throwaway git worktree
|
|
444
|
-
--parallel=3 \ # run dependency-ready, non-colliding stories concurrently
|
|
445
|
-
--decision-policy=critical # pause only for high-impact decisions; routine choices stay autonomous
|
|
446
|
-
# Optional: add --max=20 only when this run should stop after a bounded batch.
|
|
447
|
-
yoke loop off . # disable
|
|
448
|
-
```
|
|
64
|
+
For an existing project, run `yoke setup .`, set `verify.command`, and draft a backlog with `yoke prd draft . --idea="..."`. Add `--review` when an independent reviewer is configured.
|
|
449
65
|
|
|
450
66
|
**PRD format** (`.yoke/prd.yaml`):
|
|
451
67
|
|
|
452
68
|
```yaml
|
|
453
69
|
- id: STORY-1
|
|
454
70
|
title: Add a health endpoint
|
|
455
|
-
priority: 1
|
|
456
|
-
acceptance:
|
|
71
|
+
priority: 1
|
|
72
|
+
acceptance:
|
|
457
73
|
- id: health-returns-200
|
|
458
74
|
text: GET /health returns 200
|
|
459
75
|
verify: [npm run test:health-returns-200]
|
|
460
76
|
- id: health-rejects-post
|
|
461
77
|
text: POST /health returns 405
|
|
462
78
|
verify: [npm run test:health-rejects-post]
|
|
463
|
-
passes: false
|
|
464
|
-
```
|
|
465
|
-
|
|
466
|
-
New projects default to `verify.requireCriteria: true`: every story has 2–5 behavioral criteria.
|
|
467
|
-
Each criterion ID must occur in its single, approved test command; shell operators and broad,
|
|
468
|
-
untargeted suites are rejected. Yoke records each result in `.yoke/proof/<story>/evidence.json`. Configure optional
|
|
469
|
-
`completion.command` for integrated journeys such as purchase → entitlement → relaunch or
|
|
470
|
-
magic-link → callback → authenticated app. It runs whenever the current backlog has no open
|
|
471
|
-
stories; this is readiness, not a release.
|
|
472
|
-
|
|
473
|
-
No generic tool can infer whether arbitrary test code perfectly represents product meaning. Yoke
|
|
474
|
-
closes the mechanical false-done paths—targeted evidence, coverage review, clean committed state,
|
|
475
|
-
and integrated journeys—while the project still owns the correctness of its tests and production
|
|
476
|
-
observability.
|
|
477
|
-
|
|
478
|
-
### Parallel workers and the quality gauntlet
|
|
479
|
-
|
|
480
|
-
`--parallel=N` dispatches dependency-ready stories concurrently. The default maximum is three
|
|
481
|
-
shared worker units across Yoke projects; set `YOKE_MAX_PARALLEL_WORKERS=1..8` to change it.
|
|
482
|
-
`--parallel` and `loop.parallel` accept 1–8. Candidate races consume one unit per simultaneous
|
|
483
|
-
candidate. Claims carry leases, workers use isolated worktrees, collision areas and write scopes stay
|
|
484
|
-
reserved through integration, and only a mechanically green candidate enters the FIFO integration
|
|
485
|
-
queue. Integration has its own serialized lane, so it does not occupy an implementation slot while
|
|
486
|
-
unrelated work proceeds. Integrated-tree gates remain mandatory. `yoke loop status` reports local
|
|
487
|
-
workers, resource waits, shared capacity, integrations, and reopened stories. See
|
|
488
|
-
[parallel execution and safe task decomposition](docs/parallel-execution.md).
|
|
489
|
-
|
|
490
|
-
Quality is reference-driven and opt-in. Declare what one story should match:
|
|
491
|
-
|
|
492
|
-
```yaml
|
|
493
|
-
quality:
|
|
494
|
-
reference: { name: approved-home, source: design/home.png, kind: file }
|
|
495
|
-
candidate: { kind: screenshots, paths: [.yoke/proof/STORY-1/home.png] }
|
|
496
|
-
rubric: Match the approved layout, hierarchy, spacing, and states.
|
|
497
|
-
policy: blocking # or advisory
|
|
498
|
-
```
|
|
499
|
-
|
|
500
|
-
Configure project defaults, then enable the gauntlet for a run:
|
|
501
|
-
|
|
502
|
-
```yaml
|
|
503
|
-
quality:
|
|
504
|
-
enabled: false # keep opt-in, or make it the project default
|
|
505
|
-
policy: blocking
|
|
506
|
-
maxRounds: 3
|
|
507
|
-
maxMinutes: 60
|
|
508
|
-
consistencyChecks: 2
|
|
509
|
-
maxParallelCandidates: 2
|
|
510
|
-
critic: { agent: codex, model: gpt-5.6-sol } # model required for --candidates
|
|
511
|
-
repair: { agent: claude }
|
|
512
|
-
```
|
|
513
|
-
|
|
514
|
-
```bash
|
|
515
|
-
yoke loop run . --quality --quality-rounds=3 --quality-minutes=60
|
|
516
|
-
yoke loop run . --quality --candidates=2 # blind pairwise selection; stories need quality declarations
|
|
517
|
-
```
|
|
518
|
-
|
|
519
|
-
The critic compares opaque candidate/reference labels, writes schema-validated provenance, and
|
|
520
|
-
cannot modify the project. Blocking findings enter a bounded repair loop and rerun every mechanical
|
|
521
|
-
gate; advisory findings are retained without blocking. `--quality-policy=`, `--no-quality`, and
|
|
522
|
-
`--quality-unbounded` override defaults for one run. Unbounded mode is explicit and warned because
|
|
523
|
-
it removes repair limits, not Yoke's watchdog, isolation, verification, or commit safety.
|
|
524
|
-
|
|
525
|
-
State lives **outside the model context** — the PRD file plus git — so each iteration is fresh.
|
|
526
|
-
Use `yoke change add` at any time. Its ignored append-only inbox is consumed at the next story
|
|
527
|
-
boundary. A separate coverage pass must confirm that every requested outcome maps to behavioral
|
|
528
|
-
criteria before Yoke appends and commits the new stories; existing stories are never rewritten and
|
|
529
|
-
no restart is needed.
|
|
530
|
-
|
|
531
|
-
### Watching a run
|
|
532
|
-
|
|
533
|
-
Every iteration emits token-free, harness-side feedback (Node console + local files — **zero agent tokens**):
|
|
534
|
-
|
|
535
|
-
- **Live console with progress + ETA** —
|
|
536
|
-
`▶ S6 (19/45 · 42%) — implementing… · ~1h44m left (Ø 4m/story)` … `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`.
|
|
537
|
-
The estimate uses the **average duration of stories completed in this run** (current
|
|
538
|
-
velocity); before the first story lands it falls back to the recorded history of previous
|
|
539
|
-
runs (`.yoke/story-durations.json`, last 50 stories, gitignored). No data yet → no estimate,
|
|
540
|
-
never a made-up one.
|
|
541
|
-
- **`.yoke/loop-status.json`** — the current state (now including `percent` and an `eta`
|
|
542
|
-
block); read it any time with `yoke loop status`:
|
|
543
|
-
```
|
|
544
|
-
Loop: RUNNING on S6 "Weekly digest"
|
|
545
|
-
implementing · iteration 20 · 19/45 (42%) · updated 30s ago
|
|
546
|
-
~1h44m remaining (Ø 4m/story)
|
|
547
|
-
```
|
|
548
|
-
- **Parallel + quality detail** — active workers include provider, candidate ID, worktree,
|
|
549
|
-
lifecycle, phase, quality round, and repair budget; the integrator is shown separately.
|
|
550
|
-
- **`.yoke/loop.log`** — an append-only timeline of every phase transition.
|
|
551
|
-
- **`--json`** — machine mode for supervisors: every status write is *also* emitted as one
|
|
552
|
-
NDJSON line on stdout (`{"type":"status","state":"running","phase":"verifying",…}` — the
|
|
553
|
-
same shape as `loop-status.json`), the human narrative moves off stdout (the final summary
|
|
554
|
-
goes to stderr), and a consumer can follow the stream line by line instead of polling the file.
|
|
555
|
-
Provider JSON streams are also accounted: statuses (file + stream) carry cumulative input and
|
|
556
|
-
output tokens for the whole run plus provider-reported cache-read, cache-write, reasoning, model,
|
|
557
|
-
and cost fields when available. Missing values stay absent—Yoke does not estimate them.
|
|
558
|
-
|
|
559
|
-
### Pausing a run
|
|
560
|
-
|
|
561
|
-
Drop a **`.yoke/loop.pause`** file (contents irrelevant) while the loop is running and it
|
|
562
|
-
stops at the **next story boundary** — the running story still finishes, verifies, and
|
|
563
|
-
commits; no story is ever cut off mid-flight. The loop consumes the pause file, writes
|
|
564
|
-
`state: "paused"` to `loop-status.json` (log label `paused`), releases the lock, and exits
|
|
565
|
-
with code `3`. Resume by simply running `yoke loop run` again.
|
|
566
|
-
|
|
567
|
-
A per-iteration **idle timeout** guards against a genuinely hung agent: if the agent produces
|
|
568
|
-
**no output at all** for `--timeout` minutes (default 20; `0` disables), the loop kills it
|
|
569
|
-
(SIGTERM→SIGKILL) and marks the story blocked. A slow-but-working agent that keeps streaming
|
|
570
|
-
output is **never** killed — the output stream *is* the liveness signal. Set a project default
|
|
571
|
-
with `loop.timeoutMinutes` in `.yoke/config.yaml`.
|
|
572
|
-
|
|
573
|
-
### Decision policy: autonomous by default, interrupt only when configured
|
|
574
|
-
|
|
575
|
-
Planning questions happen before the loop. The provider-neutral `yoke-workflow` skill asks only
|
|
576
|
-
questions whose answer materially changes product behavior, scope, architecture, security, data
|
|
577
|
-
ownership, external cost, or an irreversible choice. It saves the approved brief in
|
|
578
|
-
`.yoke/plan.md`; `yoke prd draft` consumes it, and `yoke prd check` rejects explicit unresolved
|
|
579
|
-
placeholders such as `TBD`.
|
|
580
|
-
|
|
581
|
-
The unattended loop then follows `loop.decisionPolicy`:
|
|
582
|
-
|
|
583
|
-
```yaml
|
|
584
|
-
loop:
|
|
585
|
-
enabled: true
|
|
586
|
-
decisionPolicy: critical # or auto
|
|
587
|
-
runner:
|
|
588
|
-
agent: codex # setup chooses the current host by default
|
|
589
|
-
```
|
|
590
|
-
|
|
591
|
-
- **`auto` (default):** routine ambiguity and implementation details are resolved using the
|
|
592
|
-
approved plan, acceptance criteria, current code, and project conventions. The loop does not
|
|
593
|
-
ask follow-up questions.
|
|
594
|
-
- **`critical`:** routine choices are still resolved automatically. Only high-impact decisions
|
|
595
|
-
involving public architecture, security/privacy, destructive migration or data loss, material
|
|
596
|
-
external cost, legal/compliance exposure, or another irreversible choice may pause the story.
|
|
597
|
-
The agent writes a schema-validated request; the loop blocks before verify and preserves it as
|
|
598
|
-
`.yoke/pending-decision.yaml`.
|
|
599
|
-
|
|
600
|
-
Inspect and answer a critical stop:
|
|
601
|
-
|
|
602
|
-
```bash
|
|
603
|
-
yoke loop decision .
|
|
604
|
-
yoke loop answer . --choice=A --rationale="Matches the existing identity model"
|
|
605
|
-
```
|
|
606
|
-
|
|
607
|
-
`answer` validates the choice against the still-open story, appends it to
|
|
608
|
-
`.yoke/context/DECISIONS.md`, commits only that file using the configured human identity, clears
|
|
609
|
-
the pending request, and resumes the same story with the original runner, isolation, review,
|
|
610
|
-
permission, timeout, JSON, decision-policy, and iteration settings intact. Add
|
|
611
|
-
`--no-resume` when a supervisor should restart the loop separately. If the automatic restart
|
|
612
|
-
cannot begin because a provider/reviewer is unavailable or another process owns the lock, run
|
|
613
|
-
`yoke loop resume .`; its request-bound options are retained under Git's private state directory
|
|
614
|
-
until a loop actually runs. To intentionally abandon an orphaned or stale private resume state,
|
|
615
|
-
use `yoke loop resume . --discard`; pending decisions are never deleted by that command. Existing
|
|
616
|
-
`loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` remain supported as compatibility aliases;
|
|
617
|
-
new projects should use `decisionPolicy: auto|critical`.
|
|
618
|
-
|
|
619
|
-
### Adaptive model routing
|
|
620
|
-
|
|
621
|
-
`yoke setup` enables routing by default and preserves explicit opt-outs. Without configured
|
|
622
|
-
worker profiles, automatic execution keeps the selected parent. When enabled, the selected
|
|
623
|
-
parent remains the strong planner/controller. Before each bounded story it receives only the
|
|
624
|
-
story, acceptance criteria, and at most three eligible worker profiles, then returns one
|
|
625
|
-
machine-readable choice. The worker can be a cheaper/faster profile for any configured harness;
|
|
626
|
-
`SELF` keeps difficult work on the parent. Explicit project rules skip the controller.
|
|
627
|
-
Loop runners disable native delegation in Codex, Claude, Gemini, Qwen, OpenCode and Kilo so it cannot multiply
|
|
628
|
-
the Yoke worker budget. Integration retains its execution slot until the candidate lands.
|
|
629
|
-
|
|
630
|
-
**Provider support:** adaptive routing uses Yoke's shared provider adapter and works with Claude
|
|
631
|
-
Code, Codex CLI, Gemini CLI, Qwen Code, OpenCode, Kilo, Pi, and Hermes, including mixed-provider worker lists.
|
|
632
|
-
Internal contract tests cover invocation and routing behavior for all eight providers. The measured
|
|
633
|
-
performance evidence below is intentionally **Codex-only**; it does not claim equivalent savings
|
|
634
|
-
until authenticated, repeated in-the-wild runs exist for each provider.
|
|
635
|
-
|
|
636
|
-
```yaml
|
|
637
|
-
runner:
|
|
638
|
-
agent: codex
|
|
639
|
-
model: gpt-5.6-sol # optional; provider model strings stay opaque to Yoke
|
|
640
|
-
reasoningEffort: high
|
|
641
|
-
routing:
|
|
642
|
-
enabled: true # new setup default; false preserves an explicit opt-out
|
|
643
|
-
strategy: balanced # balanced | cost | speed | quality
|
|
644
|
-
maxCandidates: 3
|
|
645
|
-
workers:
|
|
646
|
-
- id: codex-light
|
|
647
|
-
agent: codex
|
|
648
|
-
reasoningEffort: low
|
|
649
|
-
costTier: medium
|
|
650
|
-
capabilities: [exploration, implementation, tests]
|
|
651
|
-
- id: claude-fast
|
|
652
|
-
agent: claude
|
|
653
|
-
model: haiku # rolling alias; omit to use the provider's current default
|
|
654
|
-
reasoningEffort: low
|
|
655
|
-
costTier: low
|
|
656
|
-
capabilities: [mechanical-edits, tests]
|
|
657
|
-
- id: gemini-auto
|
|
658
|
-
agent: gemini # omitted model means the account's current Auto/default route
|
|
659
|
-
costTier: low
|
|
660
|
-
capabilities: [large-context, implementation]
|
|
79
|
+
passes: false
|
|
661
80
|
```
|
|
662
81
|
|
|
663
|
-
|
|
664
|
-
baseline. Routing control calls are read-only and deliberately tiny; malformed output or no
|
|
665
|
-
eligible worker falls back to `SELF`. Yoke does not ship a universal, fast-aging
|
|
666
|
-
"intelligence score". Candidate model IDs come from project configuration while setup defaults
|
|
667
|
-
prefer rolling aliases or provider Auto/defaults. A per-user registry learns only from Yoke's
|
|
668
|
-
independent verify/performance/audit/review gates, keyed by worker + provider + model/effort and expired
|
|
669
|
-
after 30 days. It stores no prompts, source, or project paths—only a project hash and aggregate
|
|
670
|
-
time/token/outcome evidence. Writes are immutable one-event files, so concurrent Yoke instances
|
|
671
|
-
cannot overwrite a shared registry file.
|
|
672
|
-
|
|
673
|
-
Routing is not free: stories without a matching rule can add a controller call. Measure it
|
|
674
|
-
on your own backlog rather than assuming a win. Routing now also runs within asynchronous
|
|
675
|
-
parallel workers. Automatic parallelism uses up to the shared worker limit when pending tasks declare
|
|
676
|
-
write scopes; the scheduler still serializes dependencies, collision areas, and overlapping scopes.
|
|
677
|
-
Unknown scopes and configured tool actions keep automatic execution serial. Isolation is
|
|
678
|
-
on by default. Explicit `--parallel=N`, `--no-routing` and `--no-isolate` remain available.
|
|
679
|
-
See [execution defaults and dashboard measurement details](docs/VERIFIED-PROJECTS.md#execution-defaults-in-180).
|
|
680
|
-
|
|
681
|
-
### Performance budgets: efficiency as a gate, not a style
|
|
682
|
-
|
|
683
|
-
Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
|
|
684
|
-
be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
|
|
685
|
-
|
|
686
|
-
- **Per story:** write the requirement as a **measurable acceptance criterion**
|
|
687
|
-
("imports 1M rows in < 2s, asserted by the bench test") and let your verify tests measure
|
|
688
|
-
it — no new machinery needed.
|
|
689
|
-
- **Per project:** wire a benchmark as a standing **perf gate** in `.yoke/config.yaml`:
|
|
690
|
-
|
|
691
|
-
```yaml
|
|
692
|
-
perf:
|
|
693
|
-
command: node bench/check-budget.mjs # exit 0 = within budget
|
|
694
|
-
retries: 1 # benchmarks are noisy; same retry logic as verify
|
|
695
|
-
```
|
|
696
|
-
|
|
697
|
-
The loop runs it **after verify** on every story (phase `perf`, with `YOKE_STORY` set); a
|
|
698
|
-
red benchmark blocks the story — `story S6 exceeded its performance budget: p95 62ms > budget 50ms` —
|
|
699
|
-
no matter how clean the diff was. The implementer prompt names the budget command, so the
|
|
700
|
-
agent knows not to trade hot-path efficiency for style and never "simplifies away" an
|
|
701
|
-
optimization without re-running the benchmark. The `performance` canon skill carries the
|
|
702
|
-
method: profile first, optimize leaves not boundaries, commit benchmarks as tests, version
|
|
703
|
-
the *why* of every optimization in `context/DECISIONS.md`.
|
|
704
|
-
|
|
705
|
-
### Artifact-backed gate output: compact context, complete local evidence
|
|
706
|
-
|
|
707
|
-
Failed verify, executable-criterion, performance, configured custom-audit, and completion commands can emit thousands of
|
|
708
|
-
low-signal lines. Yoke keeps the model-visible failure summary deterministic and bounded while
|
|
709
|
-
preserving large raw stdout/stderr below `.yoke/artifacts/`:
|
|
82
|
+
Yoke requires 2–5 behavioral criteria on new stories by default. Each criterion needs an approved test command containing its ID. See the [PRD schema](canon/loop/prd.schema.md).
|
|
710
83
|
|
|
711
|
-
|
|
712
|
-
output:
|
|
713
|
-
previewBytes: 2048 # default: maximum compact preview bytes
|
|
714
|
-
artifactThresholdBytes: 8192 # default: persist raw output only above this size
|
|
715
|
-
```
|
|
84
|
+
## Optional continuous exploration
|
|
716
85
|
|
|
717
|
-
|
|
718
|
-
artifact threshold it also includes a project-relative path, byte count, and full SHA-256 digest,
|
|
719
|
-
for example:
|
|
86
|
+
Exploration is **off by default**. Add `--explore` to keep the supervisor active after the current backlog drains. It proposes bounded work from repository evidence; accepted tasks still pass through the project's normal tests and configured gates.
|
|
720
87
|
|
|
721
|
-
```
|
|
722
|
-
|
|
723
|
-
|
|
724
|
-
|
|
725
|
-
An agent can read that ordinary file when the preview is insufficient; nothing is injected into
|
|
726
|
-
later stories automatically. Repeated identical failures reuse the same content-addressed path.
|
|
727
|
-
Successful gate output is discarded as before. This affects only commands executed by Yoke's own
|
|
728
|
-
gates. It does **not** intercept tool output generated internally by Claude Code, Codex, Gemini, or Qwen,
|
|
729
|
-
so benchmark ratios for this feature are not provider-token or billing claims.
|
|
730
|
-
|
|
731
|
-
Command capture is capped at 16 MiB per stdout/stderr stream. Exceeding that quota fails the gate
|
|
732
|
-
closed and stores the captured prefix with a `[truncated output: ...]` marker; Yoke never labels
|
|
733
|
-
partial evidence as full output.
|
|
734
|
-
|
|
735
|
-
Yoke treats `.yoke/artifacts/` as local, non-committable runtime state and excludes it from its
|
|
736
|
-
clean-tree and story-commit operations; `yoke retrofit` also adds it to `.gitignore`. Raw command output is intentionally stored
|
|
737
|
-
without redaction so it remains valid evidence and may therefore contain credentials, personal
|
|
738
|
-
data, or other sensitive text emitted by project commands. Inspect artifacts before sharing them.
|
|
739
|
-
|
|
740
|
-
The loop trusts **verify**, not the agent's exit code: a story whose tests are green is
|
|
741
|
-
committed even if the agent process exited non-zero (a common Windows `.cmd`-wrapper ghost).
|
|
742
|
-
A failing verify is retried up to `verify.retries` times (default 1) so a transient flake
|
|
743
|
-
self-heals while a real failure still blocks. Structured acceptance criteria are then verified
|
|
744
|
-
individually; an unrelated green suite cannot satisfy a criterion without its proof command.
|
|
745
|
-
|
|
746
|
-
`.yoke/loop-status.json`, `.yoke/loop.log`, `.yoke/loop.lock`, its takeover/recovery leases, lock/decision temp files, `.yoke/story-durations.json`,
|
|
747
|
-
`.yoke/ambiguity.md`, `.yoke/artifacts/`, and the critical-decision request/answering files are runtime artifacts;
|
|
748
|
-
`yoke retrofit` gitignores them (along with
|
|
749
|
-
`.yoke/worktrees/`, `.yoke/backup/`, `.yoke/proof/`, and `.yoke/changes/`) so they never trip the clean-tree gate.
|
|
750
|
-
|
|
751
|
-
### Single-flight guard + cleanup
|
|
752
|
-
|
|
753
|
-
Two concurrent `yoke loop run`s would race on the PRD and status files, so the loop takes a
|
|
754
|
-
**lock** (`.yoke/loop.lock`) for the duration of a run. Complete lock metadata is published atomically;
|
|
755
|
-
stale takeover is serialized by `.yoke/loop.lock.takeover`. A second invocation exits `2` with
|
|
756
|
-
`Another loop is already running here (pid …). If that is wrong, run: yoke loop cleanup`. A lock
|
|
757
|
-
whose holder process is dead is taken over automatically (with a warning).
|
|
758
|
-
|
|
759
|
-
**`yoke loop cleanup [dir]`** reaps only runner process trees recorded by this project and removes
|
|
760
|
-
a stale lock. Yoke-created worktrees are **retained by default** and listed in the output; pass
|
|
761
|
-
`--remove-worktrees` to remove `.yoke/worktrees/*` with `git worktree remove --force` + `prune`.
|
|
762
|
-
User-created worktrees are never touched. A live lock is reported and left alone. Exits `0` when
|
|
763
|
-
cleanup succeeds, `1` if any requested removal fails. If a machine/process crash leaves the cleanup
|
|
764
|
-
recovery lease itself behind, an operator can run
|
|
765
|
-
`yoke loop cleanup . --discard-stale-recovery`; Yoke refuses while its recorded PID is alive, and
|
|
766
|
-
the force flag must not be run concurrently.
|
|
767
|
-
|
|
768
|
-
## 🔍 Cross-model review (`yoke review`)
|
|
769
|
-
|
|
770
|
-
Outside the loop, `yoke review` has a **second** model review your current diff as a
|
|
771
|
-
pass/fail gate — the interactive counterpart to the loop's `--review`/`--reviewer`.
|
|
772
|
-
|
|
773
|
-
```bash
|
|
774
|
-
yoke review . # review the uncommitted working tree
|
|
775
|
-
yoke review . --base=main # review the range main..HEAD instead
|
|
776
|
-
yoke review . --reviewer=codex # force a specific reviewer
|
|
777
|
-
yoke review . --focus="the auth layer" # steer what it scrutinises
|
|
88
|
+
```sh
|
|
89
|
+
yoke loop run my-app --explore --parallel=auto --explore-limit=3d
|
|
90
|
+
yoke loop status my-app
|
|
91
|
+
yoke loop pause my-app
|
|
778
92
|
```
|
|
779
93
|
|
|
780
|
-
-
|
|
781
|
-
preferring a model *other* than the one you drive so the review is genuinely cross-model.
|
|
782
|
-
On a Claude-only machine it degrades to a self-review (and says so).
|
|
783
|
-
- **Scope** — the uncommitted working tree by default, or a commit range with `--base=<ref>`.
|
|
784
|
-
- **Exit-code gate** — exits `0` when the reviewer approves, `1` when it finds a blocking
|
|
785
|
-
issue, `2` when no (or an unavailable) reviewer CLI is found. Chain it: `... && yoke review`,
|
|
786
|
-
or wire it into a pre-push hook.
|
|
787
|
-
- Runs through the same idle-timeout watchdog as the loop (`--timeout`, default 20 min).
|
|
94
|
+
Without `--explore-limit`, exploration has no time limit. A limit such as `12h`, `3d`, or `2w` requests a pause at the next safe boundary; active stories finish their normal gates and integration first, so the process can run past the deadline while that work completes. The supervisor requires an available machine, running process, and usable providers. Read the [continuous exploration guide](docs/CONTINUOUS-EXPLORATION.md) for discovery rules, retries, stop detection, and recovery.
|
|
788
95
|
|
|
789
|
-
##
|
|
96
|
+
## Parallel workers
|
|
790
97
|
|
|
791
|
-
|
|
98
|
+
Use `--parallel=auto` or `--parallel=N` to run independent stories together. The scheduler observes story dependencies, collision areas, and declared write scopes. A shared user-level pool admits up to three worker units by default; `YOKE_MAX_PARALLEL_WORKERS` adjusts that ceiling from 1 to 8. See [parallel execution](docs/parallel-execution.md) for admission, integration, and recovery details.
|
|
792
99
|
|
|
793
|
-
|
|
794
|
-
(AI-purple gradients, gradient hero text, neon glow, emoji-as-icons, gradient overload). It
|
|
795
|
-
scores findings and **exits non-zero over budget** (`--max`, default 4; `--report` to list only),
|
|
796
|
-
so it drops straight into your verify pipeline.
|
|
797
|
-
- **`yoke flow-smoke [dir]`** — a built-in browser gate with **proof artifacts** (below).
|
|
798
|
-
- **`unslop-ui` + `visual-verification` skills** — the design rubric, plus how to compose a verify
|
|
799
|
-
pipeline (`types → units → design-scan → flow-smoke`).
|
|
100
|
+
## Dashboard
|
|
800
101
|
|
|
801
|
-
|
|
802
|
-
or configured smoke flows. In `auto` mode the loop runs the design scan after functional verify and
|
|
803
|
-
before performance/audit; `on` forces it for any project and `off` disables it. Existing explicit
|
|
804
|
-
settings are preserved. `flow-smoke` remains an explicit project verify step because Yoke cannot
|
|
805
|
-
infer how to start each application's server.
|
|
102
|
+
Register projects and start the loopback-only dashboard:
|
|
806
103
|
|
|
807
|
-
|
|
808
|
-
|
|
809
|
-
|
|
810
|
-
*Tell set informed by the MIT-licensed [vibecoded-design-tells](https://github.com/JCarterJohnson/vibecoded-design-tells) research.*
|
|
811
|
-
|
|
812
|
-
### `yoke flow-smoke [dir] [--url=<baseUrl>] [--label=<name>]`
|
|
813
|
-
|
|
814
|
-
Configure your key user flows once in `.yoke/config.yaml`:
|
|
815
|
-
|
|
816
|
-
```yaml
|
|
817
|
-
smoke:
|
|
818
|
-
baseUrl: http://localhost:3000
|
|
819
|
-
flows:
|
|
820
|
-
- name: home
|
|
821
|
-
path: /
|
|
822
|
-
landmark: "main h1" # optional CSS selector to wait for
|
|
823
|
-
- name: login
|
|
824
|
-
path: /login
|
|
104
|
+
```sh
|
|
105
|
+
yoke projects add /path/to/project
|
|
106
|
+
yoke dashboard
|
|
825
107
|
```
|
|
826
108
|
|
|
827
|
-
|
|
828
|
-
landmark, and fails on a non-OK response or **any console/page error**. The proof contract:
|
|
829
|
-
|
|
830
|
-
- **Screenshots always** — every flow (pass *or* fail) saves `.yoke/proof/<label>/<flow>.png`;
|
|
831
|
-
the failure screenshot *is* the evidence.
|
|
832
|
-
- **Video only on failure** — each flow is recorded, but the clip is kept only when the flow
|
|
833
|
-
goes red (`<flow>.webm`); green runs delete it.
|
|
834
|
-
- **Labelled per story** — inside the loop, verify runs with `YOKE_STORY=<story-id>`, so proofs
|
|
835
|
-
land in `.yoke/proof/<story-id>/` automatically. Standalone runs use `latest`, or pass
|
|
836
|
-
`--label=`. The label dir is wiped per run — evidence is always from the latest run.
|
|
837
|
-
- **Exit codes** — `0` all flows green (chain it: `... && yoke design-scan . && yoke flow-smoke .`),
|
|
838
|
-
`1` any flow failed, `2` not runnable (no `smoke:` config, or Playwright missing).
|
|
839
|
-
- **Playwright comes from the *target project*, never Yoke** —
|
|
840
|
-
`npm i -D playwright && npx playwright install chromium` there. Start the dev server before
|
|
841
|
-
verify (e.g. via `start-server-and-test`); `--url=` overrides `baseUrl`.
|
|
842
|
-
|
|
843
|
-
`.yoke/proof/` is gitignored by the retrofit — proofs are runtime artifacts and never break the
|
|
844
|
-
loop's clean-tree gate.
|
|
845
|
-
|
|
846
|
-
## 🧠 Context layer (`.yoke/context/`)
|
|
847
|
-
|
|
848
|
-
Yoke keeps durable, cross-session context so a fresh-context agent is never blind:
|
|
849
|
-
|
|
850
|
-
- `PROJECT.md` — the north star (goal, constraints, non-goals, success criteria).
|
|
851
|
-
- `DECISIONS.md` — an append-only ledger. The loop adds an entry per completed story; you and agents add the *why*.
|
|
852
|
-
- `KNOWLEDGE.md` — reusable gotchas and conventions.
|
|
853
|
-
- `GLOSSARY.md` — the project's canonical terms, meanings, and aliases.
|
|
854
|
-
- `CONTEXT-MAP.md` — optional bounded-context relationships for projects that need domain mapping.
|
|
855
|
-
|
|
856
|
-
`yoke retrofit` scaffolds the four core files non-destructively; your edits are never overwritten.
|
|
857
|
-
It reports `CONTEXT-MAP.md` when the optional file already exists.
|
|
858
|
-
The loop reads them into every agent + reviewer prompt and logs decisions back on each story's
|
|
859
|
-
commit. Decision history is explicitly delimited as untrusted reference data, so stored text is
|
|
860
|
-
never treated as fresh instructions. Manage the files directly with `yoke context init` and `yoke context status`. The
|
|
861
|
-
`maintaining-context` skill teaches agents to honour the same files during interactive work.
|
|
862
|
-
|
|
863
|
-
> Commit `.yoke/context/` to git. The `--isolate` loop runs each iteration in a worktree
|
|
864
|
-
> checked out from HEAD, so it only sees committed context.
|
|
865
|
-
|
|
866
|
-
## 🛡️ Safety model
|
|
867
|
-
|
|
868
|
-
Yoke's guardrails are **mechanical, not advisory** — the loop blocks on a dirty worktree, missing acceptance criteria, red tests, or a reviewer rejection, and **none of them rely on the agent choosing to behave**.
|
|
869
|
-
|
|
870
|
-
- **Commit integrity** — a story is never recorded `passes: true` without a corresponding commit; a failed commit reverts the PRD.
|
|
871
|
-
- **Role separation** — the implementer never reviews its own work; `--reviewer` can even be a different agent.
|
|
872
|
-
- **Isolation** — with `--isolate`, failed or partial work is discarded with the worktree and never reaches your main tree.
|
|
873
|
-
- **Non-destructive retrofit** — existing files are backed up before any change; settings are merged, not replaced.
|
|
874
|
-
- **Independent verification** — "done" means *your test command exits 0*, not "the agent said so".
|
|
875
|
-
- **Single-flight** — a lock prevents two loops from racing the same repo; `yoke loop cleanup` recovers after crashes.
|
|
876
|
-
|
|
877
|
-
## 🧠 Choose your code-graph
|
|
878
|
-
|
|
879
|
-
`yoke retrofit --code-graph=graphify|serena` (default `graphify`, remembered per project) selects the legacy single graph. For complete code intelligence, enable the federated facade with `yoke retrofit --code-intelligence=active` (or `shadow` for a read-only canary). See the [Code Intelligence guide](docs/CODE-INTELLIGENCE.md).
|
|
880
|
-
|
|
881
|
-
| | **graphify** | **Serena** |
|
|
882
|
-
|---|---|---|
|
|
883
|
-
| Engine | tree-sitter AST + graph | real language servers (LSP) |
|
|
884
|
-
| Strength | fast, multimodal (code + PDFs + images) | symbol-exact cross-file refactoring |
|
|
885
|
-
| Token efficiency | ~70× reduction on large mixed repos | standard, no index to go stale |
|
|
886
|
-
| Best for | rapid exploration / migration / onboarding | systematic refactoring in typed codebases |
|
|
887
|
-
| Caveat | heuristic edges; static index can go stale | one language server per language |
|
|
888
|
-
|
|
889
|
-
The federated mode composes both structural and semantic evidence and adds Graphify's architecture/document graph. It exposes one Yoke-controlled MCP surface, content-addressed snapshots, partial-coverage reporting and isolated edit previews; the legacy `codeGraph` setting remains valid and unchanged when code intelligence is `off`.
|
|
109
|
+
The dashboard is a local control room. It does not discover every process or run arbitrary shell commands. See the [dashboard guide](docs/DASHBOARD-EVOLUTION.md) for data coverage and measurement limits.
|
|
890
110
|
|
|
891
|
-
##
|
|
111
|
+
## Agent and feature guides
|
|
892
112
|
|
|
893
|
-
|
|
113
|
+
| Guide | Details |
|
|
114
|
+
| --- | --- |
|
|
115
|
+
| [Harnesses](docs/HARNESSES.md) | CLI setup, permissions, provider/model selection, and telemetry limits |
|
|
116
|
+
| [SoL-Pi for Pi](docs/SOL-PI.md) | Opt-in settings, supported runtime, project trust, data handling, and paper evidence |
|
|
117
|
+
| [Parallel execution](docs/parallel-execution.md) | Worker scheduling, shared capacity, integration, and recovery |
|
|
118
|
+
| [Continuous exploration](docs/CONTINUOUS-EXPLORATION.md) | Autonomous discovery, runtime limits, pause/resume, and stop detection |
|
|
119
|
+
| [Project workflows](docs/VERIFIED-PROJECTS.md) | Setup, verification, goals, and execution defaults |
|
|
120
|
+
| [Code Intelligence](docs/CODE-INTELLIGENCE.md) | Optional Graphify, Serena, and Graft evidence providers |
|
|
121
|
+
| [Dashboard](docs/DASHBOARD-EVOLUTION.md) | Project views, history, controls, and reporting boundaries |
|
|
122
|
+
| [Changelog](CHANGELOG.md) | Release features, behavior changes, and migration notes |
|
|
894
123
|
|
|
895
|
-
|
|
896
|
-
- The **`minimal-code`** skill installs a YAGNI / "lazy senior dev" ladder so agents write the least code that solves the task — fewer output tokens, smaller review surface. *(Adapted from the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.)*
|
|
124
|
+
## Safety and limits
|
|
897
125
|
|
|
898
|
-
A
|
|
899
|
-
|
|
900
|
-
|
|
901
|
-
|
|
902
|
-
and 78.2% fewer reasoning tokens**. All three pairs improved wall time and fresh input.
|
|
126
|
+
- A green result means the configured checks passed; the project remains responsible for meaningful tests and acceptance criteria.
|
|
127
|
+
- Harness permissions differ. Some CLIs do not provide an OS-level sandbox; use Yoke's read-only profile for inspection and review.
|
|
128
|
+
- Exploration filters proposals but cannot guarantee product maturity or model quality. Time limits pause at safe boundaries; they do not kill an active story mid-integration.
|
|
129
|
+
- SoL-Pi is an optional Pi extension. Its pinned upstream lists Node.js 22.19+ and `@earendil-works/pi-coding-agent@0.84.2` as the tested baseline; Yoke does not enforce the Pi version. Pi project trust may be required. See [SoL-Pi limits](docs/SOL-PI.md).
|
|
903
130
|
|
|
904
|
-
|
|
905
|
-
controller overhead, so the routing default is not evidence of universal savings. Codex did not
|
|
906
|
-
emit dollar cost for these plan-backed runs; Yoke reports the measured token breakdown instead of
|
|
907
|
-
inventing a price. Method, ranges, controller cost, caveats, analyzer, and six raw JSON rows are in
|
|
908
|
-
[`bench/RESULTS.md`](bench/RESULTS.md#codex-only-full-repository-routing-study-2026-08-02).
|
|
131
|
+
## Development
|
|
909
132
|
|
|
910
|
-
|
|
911
|
-
|
|
912
|
-
|
|
913
|
-
|
|
914
|
-
|
|
915
|
-
|
|
916
|
-
|
|
917
|
-
|
|
918
|
-
|
|
919
|
-
**The problem.** Coding agents are powerful, but each speaks its own dialect — Claude has skills and hooks, Codex reads `AGENTS.md` and a TOML config, Gemini wants commands and a settings file. Keeping the same skills, safety policy, and tool wiring consistent across all of them means copy-paste drift and three things to maintain. Yoke exists to keep **one** source of truth, generate the right native artifacts for each agent, and let that harness run **autonomously and safely** when you want to hand it a spec and walk away.
|
|
920
|
-
|
|
921
|
-
**The inspiration.** Yoke is a synthesis of ideas already proven across the ecosystem: composable-skills methodology ([superpowers](https://github.com/obra/superpowers), [gstack](https://github.com/garrytan/gstack)); the portable [AGENTS.md](https://agents.md/) standard; the *"one source-of-truth → idiomatic per-harness artifacts"* generation pattern ([wshobson/agents](https://github.com/wshobson/agents)); spec-driven autonomous orchestration (GSD); mechanical safety gates and role separation (safe-agentic-workflow); and the **Ralph loop** (Geoff Huntley) — keep handing a *fresh* agent the next task until the spec is done. Token efficiency comes from [rtk](https://github.com/rtk-ai/rtk) and the write-less-code idea behind [ponytail](https://github.com/DietrichGebert/ponytail).
|
|
922
|
-
|
|
923
|
-
**How it was built.** Yoke was built the way it's meant to be *used* — agent-driven, incremental, and test-first. The stack was chosen by **researching alternatives first** (which is how `jcodemunch` was dropped for its license and Serena was added as an option). Then every component shipped one small piece at a time through a disciplined loop: **brainstorm → spec → plan → TDD implementation → an independent two-stage review** (does it match the spec? is it well-built?) **→ merge**. Those reviews caught real bugs before they shipped — a Windows `.cmd` spawn failure, a commit-integrity hole, a path-traversal that could delete project data, a resolution bug that broke the CLI's default invocation, a TOML-escaping bug. Yoke was even **dogfooded on its own repo**, which surfaced (and fixed) a genuine Windows bug. Every spec and plan lives in [`docs/superpowers/`](docs/superpowers/).
|
|
924
|
-
|
|
925
|
-
## 🗂️ Project layout
|
|
926
|
-
|
|
927
|
-
```text
|
|
928
|
-
canon/ # the source of truth — harness-agnostic
|
|
929
|
-
AGENTS.md skills/ policy/ loop/ tools/ manifest.yaml
|
|
930
|
-
src/
|
|
931
|
-
canon/ # manifest schema + validator (yoke validate)
|
|
932
|
-
code-intelligence/ # federated MCP facade, adapters, snapshots and guarded edits
|
|
933
|
-
change/ # append-only change inbox · planning · independent coverage review
|
|
934
|
-
retrofit/ # detect · plan · apply · planners (all eight harnesses) · tools
|
|
935
|
-
loop/ # prd · gates · runner · verify · git/worktree · loop · run-command · lock · cleanup
|
|
936
|
-
quality/ # reference collection · blind critic · bounded repair · candidate comparison
|
|
937
|
-
new/ # yoke new — greenfield bootstrap
|
|
938
|
-
prd/ # yoke prd draft|check — idea → stories + lint gate
|
|
939
|
-
review/ # yoke review — cross-model diff gate
|
|
940
|
-
smoke/ # yoke flow-smoke — browser gate with screenshot/video proofs
|
|
941
|
-
scan/ # yoke design-scan — AI-slop design gate
|
|
942
|
-
context/ # the durable context layer
|
|
943
|
-
docs/superpowers/ # the spec and every component's implementation plan
|
|
944
|
-
```
|
|
945
|
-
|
|
946
|
-
## 🗺️ Roadmap
|
|
947
|
-
|
|
948
|
-
Completed release work lives in the changelog. Remaining, explicitly scoped work is tracked in
|
|
949
|
-
[`TODOS.md`](TODOS.md), including broader benchmark samples, native output schemas, and signed
|
|
950
|
-
release provenance.
|
|
951
|
-
|
|
952
|
-
## 🧪 Development
|
|
953
|
-
|
|
954
|
-
```bash
|
|
955
|
-
npm test # vitest (1313 tests)
|
|
956
|
-
npm run build # tsc, no emit errors
|
|
957
|
-
npm run yoke -- validate canon
|
|
133
|
+
```sh
|
|
134
|
+
git clone https://github.com/HECer/yoke.git
|
|
135
|
+
cd yoke
|
|
136
|
+
npm ci
|
|
137
|
+
npm run lint
|
|
138
|
+
npm run build
|
|
139
|
+
npm test
|
|
140
|
+
npm run docs:check
|
|
958
141
|
```
|
|
959
142
|
|
|
960
|
-
|
|
961
|
-
|
|
962
|
-
Yoke stands on the shoulders of a great ecosystem: methodology ideas from [superpowers](https://github.com/obra/superpowers) and [gstack](https://github.com/garrytan/gstack); the [AGENTS.md](https://agents.md/) standard; the generator pattern from [wshobson/agents](https://github.com/wshobson/agents); the Ralph autonomous-loop pattern; safety-gate thinking from safe-agentic-workflow; and the wired tools [rtk](https://github.com/rtk-ai/rtk), [graphify](https://github.com/safishamsi/graphify), [Serena](https://github.com/oraios/serena), and [Playwright MCP](https://github.com/microsoft/playwright-mcp). The `minimal-code` skill adapts the MIT-licensed [ponytail](https://github.com/DietrichGebert/ponytail) ruleset.
|
|
963
|
-
|
|
964
|
-
## 📄 License
|
|
965
|
-
|
|
966
|
-
MIT — see [`LICENSE`](LICENSE).
|
|
967
|
-
|
|
968
|
-
<div align="center">
|
|
969
|
-
<sub>Built with a disciplined loop: brainstorm → spec → plan → TDD → two-stage review → merge — and reviewed by a second model, because we don't trust "done" either.</sub>
|
|
970
|
-
</div>
|
|
143
|
+
Yoke is released under the [MIT License](LICENSE).
|