@hecer/yoke 0.8.0 → 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.codex-plugin/plugin.json +7 -0
- package/CHANGELOG.md +169 -130
- package/README.md +61 -21
- package/TODOS.md +8 -0
- package/agents/docs.toml +6 -0
- package/agents/implementer.toml +6 -0
- package/agents/reviewer.toml +6 -0
- package/agents/security.toml +6 -0
- package/bench/README.md +45 -42
- package/bench/RESULTS.md +46 -36
- package/bench/result-schema.mjs +12 -0
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -0
- package/bench/results/codex-unavailable-1785175418318.json +15 -0
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -0
- package/bench/run-matrix.mjs +26 -0
- package/bench/run.mjs +127 -115
- package/canon/loop/prd.schema.md +5 -0
- package/canon/manifest.yaml +3 -1
- package/canon/skills/authoring-prd/SKILL.md +14 -0
- package/canon/skills/performance/SKILL.md +48 -0
- package/canon/skills/ship/SKILL.md +2 -7
- package/canon/tools/codex-rtk-hook.mjs +36 -0
- package/dist/agents/providers.js +23 -0
- package/dist/agents/telemetry.js +30 -0
- package/dist/agents/types.js +1 -0
- package/dist/audit/changes.js +6 -0
- package/dist/audit/command.js +64 -0
- package/dist/audit/dependencies.js +21 -0
- package/dist/audit/secrets.js +16 -0
- package/dist/audit/types.js +1 -0
- package/dist/cli.js +22 -4
- package/dist/loop/claims.js +57 -0
- package/dist/loop/cleanup.js +10 -4
- package/dist/loop/git.js +8 -2
- package/dist/loop/identity.js +27 -0
- package/dist/loop/loop.js +55 -26
- package/dist/loop/merge-queue.js +20 -0
- package/dist/loop/parallel.js +39 -0
- package/dist/loop/prd.js +48 -2
- package/dist/loop/run-command.js +63 -7
- package/dist/loop/runner.js +47 -31
- package/dist/loop/scheduler.js +8 -0
- package/dist/prd/command.js +6 -0
- package/dist/retrofit/config.js +15 -0
- package/dist/retrofit/planners/codex.js +64 -19
- package/dist/review/command.js +52 -12
- package/dist/review/verdict.js +45 -0
- package/docs/MIGRATING-TO-1.0.md +33 -0
- package/docs/superpowers/plans/2026-07-27-yoke-1.0-release.md +205 -0
- package/docs/superpowers/specs/2026-07-27-yoke-1.0-hardening-and-codex-parity-design.md +164 -0
- package/hooks/hooks.json +19 -0
- package/package.json +82 -67
- package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/config.yaml +0 -6
- package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/context/DECISIONS.md +0 -9
- package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/prd.yaml +0 -38
- package/bench/.runs/claude-2026-07-09T22-34-01/bench-verify.mjs +0 -15
- package/bench/.runs/claude-2026-07-09T22-34-01/package.json +0 -9
- package/bench/.runs/claude-2026-07-09T22-34-01/src/index.mjs +0 -48
- package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-1.test.mjs +0 -24
- package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-2.test.mjs +0 -28
- package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-3.test.mjs +0 -25
- package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/config.yaml +0 -6
- package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/prd.yaml +0 -32
- package/bench/.runs/gemini-2026-07-09T22-34-02/bench-verify.mjs +0 -15
- package/bench/.runs/gemini-2026-07-09T22-34-02/package.json +0 -9
- package/bench/.runs/gemini-2026-07-09T22-34-02/src/index.mjs +0 -3
- package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-1.test.mjs +0 -24
- package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-2.test.mjs +0 -28
- package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-3.test.mjs +0 -25
package/CHANGELOG.md
CHANGED
|
@@ -1,130 +1,169 @@
|
|
|
1
|
-
# Changelog
|
|
2
|
-
|
|
3
|
-
## 0.
|
|
4
|
-
|
|
5
|
-
### Added
|
|
6
|
-
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
the
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
`
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
- **
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
(
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 1.0.0 — 2026-07-27
|
|
4
|
+
|
|
5
|
+
### Added
|
|
6
|
+
- Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
|
|
7
|
+
- Safe provider permission profiles and structured cross-provider telemetry.
|
|
8
|
+
- Schema-validated independent review verdicts with explicit self-review opt-in.
|
|
9
|
+
- Human-owned commit identity enforcement; AI co-author trailers default off.
|
|
10
|
+
- `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
|
|
11
|
+
- PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
|
|
12
|
+
- Reproducible cross-runner benchmark schema and matrix launcher.
|
|
13
|
+
|
|
14
|
+
### Changed
|
|
15
|
+
- Dangerous permission bypass is opt-in via `--unsafe`.
|
|
16
|
+
- Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
|
|
17
|
+
- Reviews no longer trust process exit code alone.
|
|
18
|
+
|
|
19
|
+
### Security
|
|
20
|
+
- Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
|
|
21
|
+
|
|
22
|
+
|
|
23
|
+
## 0.9.0 — 2026-07-22
|
|
24
|
+
|
|
25
|
+
### Added
|
|
26
|
+
- **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
|
|
27
|
+
A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
|
|
28
|
+
verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
|
|
29
|
+
`--isolate` mode. A red benchmark blocks the story
|
|
30
|
+
(`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
|
|
31
|
+
gate is configured, the implementer prompt names the budget command so agents keep hot
|
|
32
|
+
paths efficient and never "simplify away" an optimization without re-running the benchmark.
|
|
33
|
+
- **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
|
|
34
|
+
clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
|
|
35
|
+
project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
|
|
36
|
+
tests, the *why* of every optimization versioned so future agents don't clean fast code
|
|
37
|
+
back to slow.
|
|
38
|
+
- **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
|
|
39
|
+
numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
|
|
40
|
+
round — a criterion still needing a decision is not loop-ready.
|
|
41
|
+
|
|
42
|
+
## 0.8.0 — 2026-07-20
|
|
43
|
+
|
|
44
|
+
### Added
|
|
45
|
+
- **Live progress + ETA.** Story completions are now first-class events: the console shows
|
|
46
|
+
`✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
|
|
47
|
+
`yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
|
|
48
|
+
durations of stories completed **in this run** (current velocity) and falls back to the
|
|
49
|
+
persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
|
|
50
|
+
No data → no estimate, never an invented one.
|
|
51
|
+
- **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
|
|
52
|
+
prompt now always forbids asking questions (a loop run has nobody to answer). Default
|
|
53
|
+
`resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
|
|
54
|
+
loop never stops. Opt-in `abort`: the agent writes its open questions to
|
|
55
|
+
`.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
|
|
56
|
+
story would otherwise pass on pre-existing green tests), and blocks with the question as
|
|
57
|
+
the reason. Companion principle: clarifying questions belong in the planning round, before
|
|
58
|
+
the loop starts.
|
|
59
|
+
|
|
60
|
+
## 0.7.0 — 2026-07-17
|
|
61
|
+
|
|
62
|
+
### Added
|
|
63
|
+
- **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
|
|
64
|
+
version hint (npm/gh-style): a detached background refresher caches the registry's
|
|
65
|
+
latest at most once a day; when it is newer, a one-line stderr hint suggests
|
|
66
|
+
`yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
|
|
67
|
+
`--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
|
|
68
|
+
- **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
|
|
69
|
+
loop START only — never mid-run; the running process finishes on its version and
|
|
70
|
+
the upgrade applies from the next invocation. Deliberately NOT the default:
|
|
71
|
+
a gate harness must not change itself mid-project (determinism), and unreviewed
|
|
72
|
+
auto-installs are a supply-chain hazard.
|
|
73
|
+
|
|
74
|
+
## 0.6.0 — 2026-07-17
|
|
75
|
+
|
|
76
|
+
### Added
|
|
77
|
+
- **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
|
|
78
|
+
`.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
|
|
79
|
+
`yoke loop cleanup` kills exactly those recorded process trees — and only while no
|
|
80
|
+
live loop holds the lock. Background: without a scoped mechanism, users and agents
|
|
81
|
+
resorted to machine-wide pattern kills (every process matching
|
|
82
|
+
`dangerously-skip-permissions`), which took down *healthy* runners of other projects
|
|
83
|
+
mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
|
|
84
|
+
safe path. `.yoke/runner.pid` is gitignored by retrofit.
|
|
85
|
+
|
|
86
|
+
## 0.5.0 — 2026-07-17
|
|
87
|
+
|
|
88
|
+
Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
|
|
89
|
+
(orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
|
|
90
|
+
idle window):
|
|
91
|
+
|
|
92
|
+
### Fixed
|
|
93
|
+
- **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
|
|
94
|
+
Previously it killed only the spawned shell (`shell: true`), orphaning the actual
|
|
95
|
+
agent process — which kept writing to the worktree (dirty-tree blocks, failing
|
|
96
|
+
worktree removal) and kept burning API tokens. Observed in the field as ~10
|
|
97
|
+
zombie `claude.exe` per machine plus surviving dev servers.
|
|
98
|
+
- **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
|
|
99
|
+
the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
|
|
100
|
+
processes and killed them at exactly the idle timeout — while the user saw dead air.
|
|
101
|
+
The stream doubles as liveness; token usage is now reported on every run (not just
|
|
102
|
+
`--json` mode).
|
|
103
|
+
|
|
104
|
+
### Changed
|
|
105
|
+
- README: operating notes for driving the loop from inside an agent session
|
|
106
|
+
(background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
|
|
107
|
+
outer shell-tool timeouts killing a foreground `yoke loop run` were the third
|
|
108
|
+
observed "hang" pattern.
|
|
109
|
+
|
|
110
|
+
## 0.4.0 — 2026-07-17
|
|
111
|
+
|
|
112
|
+
### Added
|
|
113
|
+
- **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
|
|
114
|
+
scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
|
|
115
|
+
documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
|
|
116
|
+
bounded final messages (cuts output-token waste). Review prompts now ground verdicts
|
|
117
|
+
in observed evidence only and keep them brief.
|
|
118
|
+
|
|
119
|
+
### Fixed
|
|
120
|
+
- `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
|
|
121
|
+
`git add -A` story commit swept the pause control file into history in
|
|
122
|
+
un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
|
|
123
|
+
blocked the resume run — the loop locked itself out.
|
|
124
|
+
|
|
125
|
+
> Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
|
|
126
|
+
> was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
|
|
127
|
+
|
|
128
|
+
## 0.3.0 — 2026-07-10
|
|
129
|
+
|
|
130
|
+
### Added
|
|
131
|
+
- **Claude Code plugin packaging** — the repo is now its own plugin marketplace
|
|
132
|
+
(`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
|
|
133
|
+
then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
|
|
134
|
+
- **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
|
|
135
|
+
installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
|
|
136
|
+
daily-crawled extensions gallery.
|
|
137
|
+
- **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
|
|
138
|
+
quality) with a fixed fixture, pre-written objective tests, and committed result data.
|
|
139
|
+
- **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
|
|
140
|
+
sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
|
|
141
|
+
(design generation paired with Yoke's design verification gates).
|
|
142
|
+
- **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
|
|
143
|
+
stories across parallel workers (`needs` dependency field, claim files, merge queue,
|
|
144
|
+
heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
|
|
145
|
+
|
|
146
|
+
### Fixed
|
|
147
|
+
- Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
|
|
148
|
+
Windows, so the loop misreported an installed `gemini` as "not found on PATH"
|
|
149
|
+
(found by the new benchmark harness).
|
|
150
|
+
- Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
|
|
151
|
+
requires a value after `-p` and errored with "Not enough arguments following: p".
|
|
152
|
+
Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
|
|
153
|
+
(also found by the benchmark harness).
|
|
154
|
+
|
|
155
|
+
### Changed
|
|
156
|
+
- README: npm install is now the primary quickstart path; documented plugin/extension
|
|
157
|
+
installs and optional companions.
|
|
158
|
+
- npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
|
|
159
|
+
`docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
|
|
160
|
+
|
|
161
|
+
## 0.2.0 — 2026-07-09
|
|
162
|
+
|
|
163
|
+
- First npm release as `@hecer/yoke`.
|
|
164
|
+
- Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
|
|
165
|
+
token-usage + model-id reporting for the claude runner.
|
|
166
|
+
- `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
|
|
167
|
+
`yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
|
|
168
|
+
- Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
|
|
169
|
+
loop with worktree isolation, watchdog, single-flight lock, commit integrity.
|
package/README.md
CHANGED
|
@@ -2,22 +2,36 @@
|
|
|
2
2
|
|
|
3
3
|
# 🐂 Yoke
|
|
4
4
|
|
|
5
|
+
<!-- yoke:version:start -->1.0.0<!-- yoke:version:end -->
|
|
6
|
+
<!-- yoke:tests:start -->500<!-- yoke:tests:end -->
|
|
7
|
+
<!-- yoke:skills:start -->28<!-- yoke:skills:end -->
|
|
8
|
+
<!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
|
|
9
|
+
|
|
5
10
|
### One harness, three agents — and zero trust in "done."
|
|
6
11
|
|
|
7
12
|
**Yoke** installs one curated canon of skills, **mechanical safety gates**, and tool wiring into any project — natively for **Claude Code, OpenAI Codex CLI, and Gemini CLI**. Then, when you want it, an opt-in autonomous loop ships your spec story-by-story: tested, cross-model-reviewed, committed — **with a screenshot to prove every story and a video for every failure**.
|
|
8
13
|
|
|
14
|
+
[](https://www.npmjs.com/package/@hecer/yoke)
|
|
15
|
+
[](https://www.npmjs.com/package/@hecer/yoke)
|
|
9
16
|
[](https://github.com/HECer/yoke/actions/workflows/ci.yml)
|
|
10
17
|
[](#-license)
|
|
11
18
|

|
|
12
19
|

|
|
13
|
-

|
|
14
21
|

|
|
15
22
|

|
|
16
23
|
|
|
24
|
+
**Install:** [`npm i -g @hecer/yoke`](https://www.npmjs.com/package/@hecer/yoke)
|
|
25
|
+
|
|
17
26
|
</div>
|
|
18
27
|
|
|
19
28
|
> **TL;DR** — `yoke new my-app --idea="..."` scaffolds a git repo, installs the harness for all three agents, and drafts a story backlog from your idea. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
|
|
20
29
|
|
|
30
|
+
Yoke 1.0 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
|
|
31
|
+
is explicit; reviews require a schema-valid verdict and a different model unless
|
|
32
|
+
`--allow-self-review` is explicit; commits enforce the human identity from project config or Git.
|
|
33
|
+
See [the 1.0 migration guide](docs/MIGRATING-TO-1.0.md).
|
|
34
|
+
|
|
21
35
|
---
|
|
22
36
|
|
|
23
37
|
## Why Yoke exists
|
|
@@ -29,7 +43,7 @@ Agentic coding in 2026 fails in four well-documented ways. Yoke answers each one
|
|
|
29
43
|
| 🎭 **The verification gap** — *"agent says done, but it isn't"* | Agents submit confidently on 100% of runs while resolving far fewer; "all tests pass" when they were never run ([silent-failures research](https://arxiv.org/pdf/2603.25764)) | The loop trusts **your verify command's exit code**, never the agent's word. A story is `passes: true` only after tests are green, the reviewer approved, and the commit landed — atomically. Plus: **screenshot proofs** per story. |
|
|
30
44
|
| 🔀 **Three agents, three configs** | Teams hand-maintain `CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, skills, and MCP wiring separately — copy-paste drift everywhere | **One canon → `yoke retrofit`** generates the idiomatic native artifacts for each agent. Change the canon once, re-retrofit everywhere. |
|
|
31
45
|
| 🌀 **Overnight loops going off the rails** | Raw Ralph-loop users "wake up to broken codebases that don't compile" | Yoke is **"Ralph, but with gates"**: clean-worktree gate, acceptance-criteria gate, green-tests gate, review gate, per-story worktree isolation, idle-timeout watchdog, single-flight lock, commit integrity. |
|
|
32
|
-
| 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a
|
|
46
|
+
| 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a second model writes a schema-validated pass/fail verdict — chainable into verify, pre-push, or CI. Cross-model review catches what self-review misses. |
|
|
33
47
|
|
|
34
48
|
**Who it's for:** anyone driving Claude Code, Codex CLI, or Gemini CLI on real projects — especially if you use more than one, want autonomous runs you can trust, or are tired of "done" meaning "probably". Greenfield (`yoke new`) and brownfield (`yoke retrofit`) both work.
|
|
35
49
|
|
|
@@ -46,16 +60,18 @@ $ yoke prd check reading-app
|
|
|
46
60
|
|
|
47
61
|
$ yoke loop on reading-app
|
|
48
62
|
$ yoke loop run reading-app --isolate --review --max=10
|
|
49
|
-
▶ STORY-1 (0/8) — implementing… · verifying… · reviewing…
|
|
50
|
-
|
|
51
|
-
▶ STORY-
|
|
63
|
+
▶ STORY-1 (0/8 · 0%) — implementing… · verifying… · reviewing… · committing…
|
|
64
|
+
✓ STORY-1 done in 3m12s — 1/8 (13%) · ~22m left
|
|
65
|
+
▶ STORY-2 (1/8 · 13%) — implementing… · ~22m left (Ø 3m12s/story)
|
|
66
|
+
✓ STORY-2 done in 2m48s — 2/8 (25%) · ~18m left
|
|
67
|
+
▶ STORY-3 (2/8 · 25%) — implementing… ✘ blocked: story did not verify (tests red)
|
|
52
68
|
# nothing was committed. fix, then re-run.
|
|
53
69
|
|
|
54
70
|
$ ls reading-app/.yoke/proof/STORY-2/
|
|
55
71
|
home.png list.png # photographic evidence, labelled per story
|
|
56
72
|
```
|
|
57
73
|
|
|
58
|
-
Every claim in that transcript is enforced by code paths with tests behind them —
|
|
74
|
+
Every claim in that transcript is enforced by code paths with tests behind them — 500 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
|
|
59
75
|
|
|
60
76
|
## 🚀 Quickstart
|
|
61
77
|
|
|
@@ -130,10 +146,11 @@ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`)
|
|
|
130
146
|
| `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
|
|
131
147
|
| `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
|
|
132
148
|
| `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
|
|
133
|
-
| `yoke prd check [dir]` | PRD lint gate (schema, duplicate ids,
|
|
149
|
+
| `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
|
|
134
150
|
| `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
|
|
135
|
-
| `yoke loop on\|off\|status\|run\|cleanup [dir]` |
|
|
136
|
-
| `yoke review [dir] [--reviewer=] [--base=] [--focus=]` |
|
|
151
|
+
| `yoke loop on\|off\|status\|run\|cleanup [dir]` | Autonomous loop; cleanup deletes worktrees only with `--remove-worktrees` | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
|
|
152
|
+
| `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
|
|
153
|
+
| `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
|
|
137
154
|
| `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
|
|
138
155
|
| `yoke flow-smoke [dir] [--url=] [--label=]` | Browser gate with screenshot/video proofs | `0` green · `1` failures · `2` not runnable |
|
|
139
156
|
|
|
@@ -150,7 +167,7 @@ Three excellent projects, three different jobs. Honest version:
|
|
|
150
167
|
| **Enforcement** | Advisory — skills *describe* the discipline; following them is up to the agent | Skill-driven; browser QA is genuinely real | **Mechanical** — gates live in code: clean tree, acceptance criteria, green tests, review verdict, commit integrity |
|
|
151
168
|
| **Autonomy** | Interactive sessions | Interactive slash-commands (`/qa`, `/ship`, …) | Opt-in **Ralph loop** with watchdog, worktree isolation, single-flight lock, per-story proofs |
|
|
152
169
|
| **Visual QA** | — | **Best-in-class**: live browser daemon (Chromium/CDP) with deep interactive QA | Built-in `flow-smoke` gate: screenshots always, video on failure, labelled per story — lighter, but *enforced* and cross-agent |
|
|
153
|
-
| **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves
|
|
170
|
+
| **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves an independent provider and validates a structured verdict, inside or outside the loop |
|
|
154
171
|
| **Footprint** | Markdown skills (plugin) | ~230 MB with browser runtime; hourly auto-update | Node CLI + markdown canon; Playwright only if you use flow-smoke, resolved **from your project** |
|
|
155
172
|
| **License** | MIT | MIT | MIT |
|
|
156
173
|
|
|
@@ -183,10 +200,10 @@ Three layers — **Canon** (`yoke validate`) → **Retrofit** (`yoke retrofit`)
|
|
|
183
200
|
| Agent | Artifacts |
|
|
184
201
|
|---|---|
|
|
185
202
|
| **Claude** | `.claude/skills/`, `AGENTS.md`, `CLAUDE.md`, `.mcp.json` (code-graph + Playwright), and an rtk `PreToolUse` hook when WSL is available |
|
|
186
|
-
| **Codex** | `AGENTS.md`
|
|
203
|
+
| **Codex** | `.agents/skills/`, `AGENTS.md`, `RTK.md`, `.codex/config.toml`, native hooks, reusable `.codex/agents/*.toml`, and package plugin metadata |
|
|
187
204
|
| **Gemini** | `GEMINI.md`, `.gemini/commands/*.toml` (one per skill, full body), `.gemini/settings.json` (MCP + `AGENTS.md` context) |
|
|
188
205
|
|
|
189
|
-
> **rtk
|
|
206
|
+
> **rtk integration:** Claude receives its PreToolUse hook; Codex receives a native hook adapter around `rtk hook check`; Gemini retains instruction-mode fallback where its CLI has no equivalent command-rewrite lifecycle.
|
|
190
207
|
|
|
191
208
|
> **Composes with gstack:** if [gstack](https://github.com/garrytan/gstack) is installed (repo-local or global), `yoke retrofit` adds a short "Composed tools" routing note to **CLAUDE.md only** — telling Claude to prefer gstack's skills for capabilities Yoke doesn't ship (live-browser QA `/qa`, security audit `/cso`, ship/deploy `/ship`). No bundling, no dependency; the note is never written to the Codex or Gemini artifacts.
|
|
192
209
|
|
|
@@ -197,7 +214,7 @@ Three layers — **Canon** (`yoke validate`) → **Retrofit** (`yoke retrofit`)
|
|
|
197
214
|
> instructions (tech stack, workflow, `@`-includes) inside it. Works in any yoke-written file;
|
|
198
215
|
> content *outside* the markers is still replaced (and backed up under `.yoke/backup/`).
|
|
199
216
|
|
|
200
|
-
## 🧰 What's in the canon —
|
|
217
|
+
## 🧰 What's in the canon — 28 skills
|
|
201
218
|
|
|
202
219
|
`yoke retrofit` installs all of these into each agent natively. Provenance is credited in [`canon/skills/ATTRIBUTION.md`](canon/skills/ATTRIBUTION.md).
|
|
203
220
|
|
|
@@ -233,13 +250,14 @@ To stop overlapping skills from auto-invoking against each other, `canon/AGENTS.
|
|
|
233
250
|
| `retro` | Engineering retrospective from commit history |
|
|
234
251
|
| `document-release` | Post-ship documentation sync (README / CHANGELOG / …) |
|
|
235
252
|
|
|
236
|
-
**Yoke-native** — *authored or adapted for this harness (
|
|
253
|
+
**Yoke-native** — *authored or adapted for this harness (8)*
|
|
237
254
|
|
|
238
255
|
| Skill | What it does |
|
|
239
256
|
|---|---|
|
|
240
257
|
| `yoke-retrofit` | Set up the Yoke harness in a project (detect → plan → apply) |
|
|
241
258
|
| `authoring-prd` | Slice a product idea into loop-ready stories with testable acceptance criteria |
|
|
242
259
|
| `minimal-code` | Write the least code that solves the task (YAGNI; ponytail-derived) |
|
|
260
|
+
| `performance` | Efficiency as a measured requirement: benchmarks as tests, budgets as gates, optimizations local + documented |
|
|
243
261
|
| `maintaining-context` | Keep `.yoke/context/` the durable source of truth (the Context layer) |
|
|
244
262
|
| `workflow` | The default order of operations, from idea to deploy |
|
|
245
263
|
| `unslop-ui` | Detect & remove AI-slop design tells (purple gradients, neon glow, emoji-icons…) |
|
|
@@ -302,6 +320,7 @@ yoke loop run . \
|
|
|
302
320
|
--runner=codex \ # implement with Codex…
|
|
303
321
|
--reviewer=claude \ # …review with Claude (role separation)
|
|
304
322
|
--isolate \ # each story in a throwaway git worktree
|
|
323
|
+
--on-ambiguity=abort \ # strict: stop on undecidable criteria instead of guessing
|
|
305
324
|
--max=20
|
|
306
325
|
yoke loop off . # disable
|
|
307
326
|
```
|
|
@@ -381,6 +400,30 @@ Enable strict mode per run with `yoke loop run . --on-ambiguity=abort` or per pr
|
|
|
381
400
|
put every clarifying question into the PRD **before** the loop starts (`yoke prd draft`
|
|
382
401
|
criteria must be testable and decision-free).
|
|
383
402
|
|
|
403
|
+
### Performance budgets: efficiency as a gate, not a style
|
|
404
|
+
|
|
405
|
+
Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
|
|
406
|
+
be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
|
|
407
|
+
|
|
408
|
+
- **Per story:** write the requirement as a **measurable acceptance criterion**
|
|
409
|
+
("imports 1M rows in < 2s, asserted by the bench test") and let your verify tests measure
|
|
410
|
+
it — no new machinery needed.
|
|
411
|
+
- **Per project:** wire a benchmark as a standing **perf gate** in `.yoke/config.yaml`:
|
|
412
|
+
|
|
413
|
+
```yaml
|
|
414
|
+
perf:
|
|
415
|
+
command: node bench/check-budget.mjs # exit 0 = within budget
|
|
416
|
+
retries: 1 # benchmarks are noisy; same retry logic as verify
|
|
417
|
+
```
|
|
418
|
+
|
|
419
|
+
The loop runs it **after verify** on every story (phase `perf`, with `YOKE_STORY` set); a
|
|
420
|
+
red benchmark blocks the story — `story S6 exceeded its performance budget: p95 62ms > budget 50ms` —
|
|
421
|
+
no matter how clean the diff was. The implementer prompt names the budget command, so the
|
|
422
|
+
agent knows not to trade hot-path efficiency for style and never "simplifies away" an
|
|
423
|
+
optimization without re-running the benchmark. The `performance` canon skill carries the
|
|
424
|
+
method: profile first, optimize leaves not boundaries, commit benchmarks as tests, version
|
|
425
|
+
the *why* of every optimization in `context/DECISIONS.md`.
|
|
426
|
+
|
|
384
427
|
The loop trusts **verify**, not the agent's exit code: a story whose tests are green is
|
|
385
428
|
committed even if the agent process exited non-zero (a common Windows `.cmd`-wrapper ghost).
|
|
386
429
|
A failing verify is retried up to `verify.retries` times (default 1) so a transient flake
|
|
@@ -555,17 +598,14 @@ docs/superpowers/ # the spec and every component's implementation plan
|
|
|
555
598
|
|
|
556
599
|
## 🗺️ Roadmap
|
|
557
600
|
|
|
558
|
-
|
|
559
|
-
|
|
560
|
-
|
|
561
|
-
- **Multi-reviewer quorum** — N independent reviewers with distinct lenses (correctness / security / acceptance).
|
|
562
|
-
- **Merge queue** — re-test against the latest main before integrating, for parallel/multi-agent loops.
|
|
563
|
-
- **More agents** — the canon→retrofit pattern generalises; OpenCode and Copilot CLI are natural next targets.
|
|
601
|
+
Yoke 1.0's completed release work moved to the changelog. Remaining, explicitly scoped work
|
|
602
|
+
is tracked in [`TODOS.md`](TODOS.md), including provider subprocess wiring for the tested
|
|
603
|
+
parallel dispatcher, broader benchmark samples, native output schemas, and release provenance.
|
|
564
604
|
|
|
565
605
|
## 🧪 Development
|
|
566
606
|
|
|
567
607
|
```bash
|
|
568
|
-
npm test # vitest (
|
|
608
|
+
npm test # vitest (500 tests)
|
|
569
609
|
npm run build # tsc, no emit errors
|
|
570
610
|
npm run yoke -- validate canon
|
|
571
611
|
```
|
package/TODOS.md
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
# Yoke follow-up work
|
|
2
|
+
|
|
3
|
+
- Wire the tested async parallel dispatcher to provider subprocess workers. Until then the CLI
|
|
4
|
+
rejects `--parallel=N` for `N > 1`; scheduler, claims, and merge queue APIs are available
|
|
5
|
+
without claiming a CLI speed-up.
|
|
6
|
+
- Add provider-native output schemas when all three CLIs expose compatible stable APIs.
|
|
7
|
+
- Expand benchmark fixtures and collect multiple authenticated samples per provider/model.
|
|
8
|
+
- Add signed provenance and attestations to npm and GitHub releases.
|
package/agents/docs.toml
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
name = "docs"
|
|
2
|
+
description = "Documentation specialist for release and API consistency."
|
|
3
|
+
sandbox_mode = "workspace-write"
|
|
4
|
+
developer_instructions = """
|
|
5
|
+
Update only documentation required by the assigned change. Verify commands and version references against the repository.
|
|
6
|
+
"""
|
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
name = "implementer"
|
|
2
|
+
description = "Implementation specialist for one scoped story."
|
|
3
|
+
sandbox_mode = "workspace-write"
|
|
4
|
+
developer_instructions = """
|
|
5
|
+
Implement only the assigned scope. Use tests first, run verification, and do not review or commit your own work.
|
|
6
|
+
"""
|