@hecer/yoke 0.9.0 → 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.codex-plugin/plugin.json +7 -0
- package/CHANGELOG.md +169 -149
- package/README.md +24 -16
- package/TODOS.md +8 -0
- package/agents/docs.toml +6 -0
- package/agents/implementer.toml +6 -0
- package/agents/reviewer.toml +6 -0
- package/agents/security.toml +6 -0
- package/bench/README.md +45 -42
- package/bench/RESULTS.md +46 -36
- package/bench/result-schema.mjs +12 -0
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -0
- package/bench/results/codex-unavailable-1785175418318.json +15 -0
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -0
- package/bench/run-matrix.mjs +26 -0
- package/bench/run.mjs +127 -115
- package/canon/loop/prd.schema.md +5 -0
- package/canon/manifest.yaml +1 -1
- package/canon/skills/authoring-prd/SKILL.md +6 -0
- package/canon/skills/ship/SKILL.md +2 -7
- package/canon/tools/codex-rtk-hook.mjs +36 -0
- package/dist/agents/providers.js +23 -0
- package/dist/agents/telemetry.js +30 -0
- package/dist/agents/types.js +1 -0
- package/dist/audit/changes.js +6 -0
- package/dist/audit/command.js +64 -0
- package/dist/audit/dependencies.js +21 -0
- package/dist/audit/secrets.js +16 -0
- package/dist/audit/types.js +1 -0
- package/dist/cli.js +22 -4
- package/dist/loop/claims.js +57 -0
- package/dist/loop/cleanup.js +10 -4
- package/dist/loop/git.js +8 -2
- package/dist/loop/identity.js +27 -0
- package/dist/loop/loop.js +20 -2
- package/dist/loop/merge-queue.js +20 -0
- package/dist/loop/parallel.js +39 -0
- package/dist/loop/prd.js +48 -2
- package/dist/loop/run-command.js +51 -6
- package/dist/loop/runner.js +41 -29
- package/dist/loop/scheduler.js +8 -0
- package/dist/prd/command.js +6 -0
- package/dist/retrofit/config.js +12 -0
- package/dist/retrofit/planners/codex.js +64 -19
- package/dist/review/command.js +52 -12
- package/dist/review/verdict.js +45 -0
- package/docs/MIGRATING-TO-1.0.md +33 -0
- package/docs/superpowers/plans/2026-07-27-yoke-1.0-release.md +205 -0
- package/docs/superpowers/specs/2026-07-27-yoke-1.0-hardening-and-codex-parity-design.md +164 -0
- package/hooks/hooks.json +19 -0
- package/package.json +82 -67
- package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/config.yaml +0 -6
- package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/context/DECISIONS.md +0 -9
- package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/prd.yaml +0 -38
- package/bench/.runs/claude-2026-07-09T22-34-01/bench-verify.mjs +0 -15
- package/bench/.runs/claude-2026-07-09T22-34-01/package.json +0 -9
- package/bench/.runs/claude-2026-07-09T22-34-01/src/index.mjs +0 -48
- package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-1.test.mjs +0 -24
- package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-2.test.mjs +0 -28
- package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-3.test.mjs +0 -25
- package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/config.yaml +0 -6
- package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/prd.yaml +0 -32
- package/bench/.runs/gemini-2026-07-09T22-34-02/bench-verify.mjs +0 -15
- package/bench/.runs/gemini-2026-07-09T22-34-02/package.json +0 -9
- package/bench/.runs/gemini-2026-07-09T22-34-02/src/index.mjs +0 -3
- package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-1.test.mjs +0 -24
- package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-2.test.mjs +0 -28
- package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-3.test.mjs +0 -25
package/CHANGELOG.md
CHANGED
|
@@ -1,149 +1,169 @@
|
|
|
1
|
-
# Changelog
|
|
2
|
-
|
|
3
|
-
## 0.
|
|
4
|
-
|
|
5
|
-
### Added
|
|
6
|
-
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
###
|
|
93
|
-
- **
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 1.0.0 — 2026-07-27
|
|
4
|
+
|
|
5
|
+
### Added
|
|
6
|
+
- Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
|
|
7
|
+
- Safe provider permission profiles and structured cross-provider telemetry.
|
|
8
|
+
- Schema-validated independent review verdicts with explicit self-review opt-in.
|
|
9
|
+
- Human-owned commit identity enforcement; AI co-author trailers default off.
|
|
10
|
+
- `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
|
|
11
|
+
- PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
|
|
12
|
+
- Reproducible cross-runner benchmark schema and matrix launcher.
|
|
13
|
+
|
|
14
|
+
### Changed
|
|
15
|
+
- Dangerous permission bypass is opt-in via `--unsafe`.
|
|
16
|
+
- Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
|
|
17
|
+
- Reviews no longer trust process exit code alone.
|
|
18
|
+
|
|
19
|
+
### Security
|
|
20
|
+
- Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
|
|
21
|
+
|
|
22
|
+
|
|
23
|
+
## 0.9.0 — 2026-07-22
|
|
24
|
+
|
|
25
|
+
### Added
|
|
26
|
+
- **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
|
|
27
|
+
A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
|
|
28
|
+
verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
|
|
29
|
+
`--isolate` mode. A red benchmark blocks the story
|
|
30
|
+
(`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
|
|
31
|
+
gate is configured, the implementer prompt names the budget command so agents keep hot
|
|
32
|
+
paths efficient and never "simplify away" an optimization without re-running the benchmark.
|
|
33
|
+
- **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
|
|
34
|
+
clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
|
|
35
|
+
project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
|
|
36
|
+
tests, the *why* of every optimization versioned so future agents don't clean fast code
|
|
37
|
+
back to slow.
|
|
38
|
+
- **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
|
|
39
|
+
numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
|
|
40
|
+
round — a criterion still needing a decision is not loop-ready.
|
|
41
|
+
|
|
42
|
+
## 0.8.0 — 2026-07-20
|
|
43
|
+
|
|
44
|
+
### Added
|
|
45
|
+
- **Live progress + ETA.** Story completions are now first-class events: the console shows
|
|
46
|
+
`✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
|
|
47
|
+
`yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
|
|
48
|
+
durations of stories completed **in this run** (current velocity) and falls back to the
|
|
49
|
+
persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
|
|
50
|
+
No data → no estimate, never an invented one.
|
|
51
|
+
- **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
|
|
52
|
+
prompt now always forbids asking questions (a loop run has nobody to answer). Default
|
|
53
|
+
`resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
|
|
54
|
+
loop never stops. Opt-in `abort`: the agent writes its open questions to
|
|
55
|
+
`.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
|
|
56
|
+
story would otherwise pass on pre-existing green tests), and blocks with the question as
|
|
57
|
+
the reason. Companion principle: clarifying questions belong in the planning round, before
|
|
58
|
+
the loop starts.
|
|
59
|
+
|
|
60
|
+
## 0.7.0 — 2026-07-17
|
|
61
|
+
|
|
62
|
+
### Added
|
|
63
|
+
- **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
|
|
64
|
+
version hint (npm/gh-style): a detached background refresher caches the registry's
|
|
65
|
+
latest at most once a day; when it is newer, a one-line stderr hint suggests
|
|
66
|
+
`yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
|
|
67
|
+
`--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
|
|
68
|
+
- **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
|
|
69
|
+
loop START only — never mid-run; the running process finishes on its version and
|
|
70
|
+
the upgrade applies from the next invocation. Deliberately NOT the default:
|
|
71
|
+
a gate harness must not change itself mid-project (determinism), and unreviewed
|
|
72
|
+
auto-installs are a supply-chain hazard.
|
|
73
|
+
|
|
74
|
+
## 0.6.0 — 2026-07-17
|
|
75
|
+
|
|
76
|
+
### Added
|
|
77
|
+
- **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
|
|
78
|
+
`.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
|
|
79
|
+
`yoke loop cleanup` kills exactly those recorded process trees — and only while no
|
|
80
|
+
live loop holds the lock. Background: without a scoped mechanism, users and agents
|
|
81
|
+
resorted to machine-wide pattern kills (every process matching
|
|
82
|
+
`dangerously-skip-permissions`), which took down *healthy* runners of other projects
|
|
83
|
+
mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
|
|
84
|
+
safe path. `.yoke/runner.pid` is gitignored by retrofit.
|
|
85
|
+
|
|
86
|
+
## 0.5.0 — 2026-07-17
|
|
87
|
+
|
|
88
|
+
Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
|
|
89
|
+
(orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
|
|
90
|
+
idle window):
|
|
91
|
+
|
|
92
|
+
### Fixed
|
|
93
|
+
- **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
|
|
94
|
+
Previously it killed only the spawned shell (`shell: true`), orphaning the actual
|
|
95
|
+
agent process — which kept writing to the worktree (dirty-tree blocks, failing
|
|
96
|
+
worktree removal) and kept burning API tokens. Observed in the field as ~10
|
|
97
|
+
zombie `claude.exe` per machine plus surviving dev servers.
|
|
98
|
+
- **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
|
|
99
|
+
the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
|
|
100
|
+
processes and killed them at exactly the idle timeout — while the user saw dead air.
|
|
101
|
+
The stream doubles as liveness; token usage is now reported on every run (not just
|
|
102
|
+
`--json` mode).
|
|
103
|
+
|
|
104
|
+
### Changed
|
|
105
|
+
- README: operating notes for driving the loop from inside an agent session
|
|
106
|
+
(background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
|
|
107
|
+
outer shell-tool timeouts killing a foreground `yoke loop run` were the third
|
|
108
|
+
observed "hang" pattern.
|
|
109
|
+
|
|
110
|
+
## 0.4.0 — 2026-07-17
|
|
111
|
+
|
|
112
|
+
### Added
|
|
113
|
+
- **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
|
|
114
|
+
scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
|
|
115
|
+
documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
|
|
116
|
+
bounded final messages (cuts output-token waste). Review prompts now ground verdicts
|
|
117
|
+
in observed evidence only and keep them brief.
|
|
118
|
+
|
|
119
|
+
### Fixed
|
|
120
|
+
- `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
|
|
121
|
+
`git add -A` story commit swept the pause control file into history in
|
|
122
|
+
un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
|
|
123
|
+
blocked the resume run — the loop locked itself out.
|
|
124
|
+
|
|
125
|
+
> Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
|
|
126
|
+
> was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
|
|
127
|
+
|
|
128
|
+
## 0.3.0 — 2026-07-10
|
|
129
|
+
|
|
130
|
+
### Added
|
|
131
|
+
- **Claude Code plugin packaging** — the repo is now its own plugin marketplace
|
|
132
|
+
(`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
|
|
133
|
+
then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
|
|
134
|
+
- **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
|
|
135
|
+
installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
|
|
136
|
+
daily-crawled extensions gallery.
|
|
137
|
+
- **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
|
|
138
|
+
quality) with a fixed fixture, pre-written objective tests, and committed result data.
|
|
139
|
+
- **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
|
|
140
|
+
sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
|
|
141
|
+
(design generation paired with Yoke's design verification gates).
|
|
142
|
+
- **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
|
|
143
|
+
stories across parallel workers (`needs` dependency field, claim files, merge queue,
|
|
144
|
+
heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
|
|
145
|
+
|
|
146
|
+
### Fixed
|
|
147
|
+
- Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
|
|
148
|
+
Windows, so the loop misreported an installed `gemini` as "not found on PATH"
|
|
149
|
+
(found by the new benchmark harness).
|
|
150
|
+
- Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
|
|
151
|
+
requires a value after `-p` and errored with "Not enough arguments following: p".
|
|
152
|
+
Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
|
|
153
|
+
(also found by the benchmark harness).
|
|
154
|
+
|
|
155
|
+
### Changed
|
|
156
|
+
- README: npm install is now the primary quickstart path; documented plugin/extension
|
|
157
|
+
installs and optional companions.
|
|
158
|
+
- npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
|
|
159
|
+
`docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
|
|
160
|
+
|
|
161
|
+
## 0.2.0 — 2026-07-09
|
|
162
|
+
|
|
163
|
+
- First npm release as `@hecer/yoke`.
|
|
164
|
+
- Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
|
|
165
|
+
token-usage + model-id reporting for the claude runner.
|
|
166
|
+
- `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
|
|
167
|
+
`yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
|
|
168
|
+
- Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
|
|
169
|
+
loop with worktree isolation, watchdog, single-flight lock, commit integrity.
|
package/README.md
CHANGED
|
@@ -2,6 +2,11 @@
|
|
|
2
2
|
|
|
3
3
|
# 🐂 Yoke
|
|
4
4
|
|
|
5
|
+
<!-- yoke:version:start -->1.0.0<!-- yoke:version:end -->
|
|
6
|
+
<!-- yoke:tests:start -->500<!-- yoke:tests:end -->
|
|
7
|
+
<!-- yoke:skills:start -->28<!-- yoke:skills:end -->
|
|
8
|
+
<!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
|
|
9
|
+
|
|
5
10
|
### One harness, three agents — and zero trust in "done."
|
|
6
11
|
|
|
7
12
|
**Yoke** installs one curated canon of skills, **mechanical safety gates**, and tool wiring into any project — natively for **Claude Code, OpenAI Codex CLI, and Gemini CLI**. Then, when you want it, an opt-in autonomous loop ships your spec story-by-story: tested, cross-model-reviewed, committed — **with a screenshot to prove every story and a video for every failure**.
|
|
@@ -12,7 +17,7 @@
|
|
|
12
17
|
[](#-license)
|
|
13
18
|

|
|
14
19
|

|
|
15
|
-

|
|
16
21
|

|
|
17
22
|

|
|
18
23
|
|
|
@@ -22,6 +27,11 @@
|
|
|
22
27
|
|
|
23
28
|
> **TL;DR** — `yoke new my-app --idea="..."` scaffolds a git repo, installs the harness for all three agents, and drafts a story backlog from your idea. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
|
|
24
29
|
|
|
30
|
+
Yoke 1.0 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
|
|
31
|
+
is explicit; reviews require a schema-valid verdict and a different model unless
|
|
32
|
+
`--allow-self-review` is explicit; commits enforce the human identity from project config or Git.
|
|
33
|
+
See [the 1.0 migration guide](docs/MIGRATING-TO-1.0.md).
|
|
34
|
+
|
|
25
35
|
---
|
|
26
36
|
|
|
27
37
|
## Why Yoke exists
|
|
@@ -33,7 +43,7 @@ Agentic coding in 2026 fails in four well-documented ways. Yoke answers each one
|
|
|
33
43
|
| 🎭 **The verification gap** — *"agent says done, but it isn't"* | Agents submit confidently on 100% of runs while resolving far fewer; "all tests pass" when they were never run ([silent-failures research](https://arxiv.org/pdf/2603.25764)) | The loop trusts **your verify command's exit code**, never the agent's word. A story is `passes: true` only after tests are green, the reviewer approved, and the commit landed — atomically. Plus: **screenshot proofs** per story. |
|
|
34
44
|
| 🔀 **Three agents, three configs** | Teams hand-maintain `CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, skills, and MCP wiring separately — copy-paste drift everywhere | **One canon → `yoke retrofit`** generates the idiomatic native artifacts for each agent. Change the canon once, re-retrofit everywhere. |
|
|
35
45
|
| 🌀 **Overnight loops going off the rails** | Raw Ralph-loop users "wake up to broken codebases that don't compile" | Yoke is **"Ralph, but with gates"**: clean-worktree gate, acceptance-criteria gate, green-tests gate, review gate, per-story worktree isolation, idle-timeout watchdog, single-flight lock, commit integrity. |
|
|
36
|
-
| 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a
|
|
46
|
+
| 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a second model writes a schema-validated pass/fail verdict — chainable into verify, pre-push, or CI. Cross-model review catches what self-review misses. |
|
|
37
47
|
|
|
38
48
|
**Who it's for:** anyone driving Claude Code, Codex CLI, or Gemini CLI on real projects — especially if you use more than one, want autonomous runs you can trust, or are tired of "done" meaning "probably". Greenfield (`yoke new`) and brownfield (`yoke retrofit`) both work.
|
|
39
49
|
|
|
@@ -61,7 +71,7 @@ $ ls reading-app/.yoke/proof/STORY-2/
|
|
|
61
71
|
home.png list.png # photographic evidence, labelled per story
|
|
62
72
|
```
|
|
63
73
|
|
|
64
|
-
Every claim in that transcript is enforced by code paths with tests behind them —
|
|
74
|
+
Every claim in that transcript is enforced by code paths with tests behind them — 500 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
|
|
65
75
|
|
|
66
76
|
## 🚀 Quickstart
|
|
67
77
|
|
|
@@ -136,10 +146,11 @@ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`)
|
|
|
136
146
|
| `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
|
|
137
147
|
| `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
|
|
138
148
|
| `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
|
|
139
|
-
| `yoke prd check [dir]` | PRD lint gate (schema, duplicate ids,
|
|
149
|
+
| `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
|
|
140
150
|
| `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
|
|
141
|
-
| `yoke loop on\|off\|status\|run\|cleanup [dir]` |
|
|
142
|
-
| `yoke review [dir] [--reviewer=] [--base=] [--focus=]` |
|
|
151
|
+
| `yoke loop on\|off\|status\|run\|cleanup [dir]` | Autonomous loop; cleanup deletes worktrees only with `--remove-worktrees` | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
|
|
152
|
+
| `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
|
|
153
|
+
| `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
|
|
143
154
|
| `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
|
|
144
155
|
| `yoke flow-smoke [dir] [--url=] [--label=]` | Browser gate with screenshot/video proofs | `0` green · `1` failures · `2` not runnable |
|
|
145
156
|
|
|
@@ -156,7 +167,7 @@ Three excellent projects, three different jobs. Honest version:
|
|
|
156
167
|
| **Enforcement** | Advisory — skills *describe* the discipline; following them is up to the agent | Skill-driven; browser QA is genuinely real | **Mechanical** — gates live in code: clean tree, acceptance criteria, green tests, review verdict, commit integrity |
|
|
157
168
|
| **Autonomy** | Interactive sessions | Interactive slash-commands (`/qa`, `/ship`, …) | Opt-in **Ralph loop** with watchdog, worktree isolation, single-flight lock, per-story proofs |
|
|
158
169
|
| **Visual QA** | — | **Best-in-class**: live browser daemon (Chromium/CDP) with deep interactive QA | Built-in `flow-smoke` gate: screenshots always, video on failure, labelled per story — lighter, but *enforced* and cross-agent |
|
|
159
|
-
| **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves
|
|
170
|
+
| **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves an independent provider and validates a structured verdict, inside or outside the loop |
|
|
160
171
|
| **Footprint** | Markdown skills (plugin) | ~230 MB with browser runtime; hourly auto-update | Node CLI + markdown canon; Playwright only if you use flow-smoke, resolved **from your project** |
|
|
161
172
|
| **License** | MIT | MIT | MIT |
|
|
162
173
|
|
|
@@ -189,10 +200,10 @@ Three layers — **Canon** (`yoke validate`) → **Retrofit** (`yoke retrofit`)
|
|
|
189
200
|
| Agent | Artifacts |
|
|
190
201
|
|---|---|
|
|
191
202
|
| **Claude** | `.claude/skills/`, `AGENTS.md`, `CLAUDE.md`, `.mcp.json` (code-graph + Playwright), and an rtk `PreToolUse` hook when WSL is available |
|
|
192
|
-
| **Codex** | `AGENTS.md`
|
|
203
|
+
| **Codex** | `.agents/skills/`, `AGENTS.md`, `RTK.md`, `.codex/config.toml`, native hooks, reusable `.codex/agents/*.toml`, and package plugin metadata |
|
|
193
204
|
| **Gemini** | `GEMINI.md`, `.gemini/commands/*.toml` (one per skill, full body), `.gemini/settings.json` (MCP + `AGENTS.md` context) |
|
|
194
205
|
|
|
195
|
-
> **rtk
|
|
206
|
+
> **rtk integration:** Claude receives its PreToolUse hook; Codex receives a native hook adapter around `rtk hook check`; Gemini retains instruction-mode fallback where its CLI has no equivalent command-rewrite lifecycle.
|
|
196
207
|
|
|
197
208
|
> **Composes with gstack:** if [gstack](https://github.com/garrytan/gstack) is installed (repo-local or global), `yoke retrofit` adds a short "Composed tools" routing note to **CLAUDE.md only** — telling Claude to prefer gstack's skills for capabilities Yoke doesn't ship (live-browser QA `/qa`, security audit `/cso`, ship/deploy `/ship`). No bundling, no dependency; the note is never written to the Codex or Gemini artifacts.
|
|
198
209
|
|
|
@@ -587,17 +598,14 @@ docs/superpowers/ # the spec and every component's implementation plan
|
|
|
587
598
|
|
|
588
599
|
## 🗺️ Roadmap
|
|
589
600
|
|
|
590
|
-
|
|
591
|
-
|
|
592
|
-
|
|
593
|
-
- **Multi-reviewer quorum** — N independent reviewers with distinct lenses (correctness / security / acceptance).
|
|
594
|
-
- **Merge queue** — re-test against the latest main before integrating, for parallel/multi-agent loops.
|
|
595
|
-
- **More agents** — the canon→retrofit pattern generalises; OpenCode and Copilot CLI are natural next targets.
|
|
601
|
+
Yoke 1.0's completed release work moved to the changelog. Remaining, explicitly scoped work
|
|
602
|
+
is tracked in [`TODOS.md`](TODOS.md), including provider subprocess wiring for the tested
|
|
603
|
+
parallel dispatcher, broader benchmark samples, native output schemas, and release provenance.
|
|
596
604
|
|
|
597
605
|
## 🧪 Development
|
|
598
606
|
|
|
599
607
|
```bash
|
|
600
|
-
npm test # vitest (
|
|
608
|
+
npm test # vitest (500 tests)
|
|
601
609
|
npm run build # tsc, no emit errors
|
|
602
610
|
npm run yoke -- validate canon
|
|
603
611
|
```
|
package/TODOS.md
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
# Yoke follow-up work
|
|
2
|
+
|
|
3
|
+
- Wire the tested async parallel dispatcher to provider subprocess workers. Until then the CLI
|
|
4
|
+
rejects `--parallel=N` for `N > 1`; scheduler, claims, and merge queue APIs are available
|
|
5
|
+
without claiming a CLI speed-up.
|
|
6
|
+
- Add provider-native output schemas when all three CLIs expose compatible stable APIs.
|
|
7
|
+
- Expand benchmark fixtures and collect multiple authenticated samples per provider/model.
|
|
8
|
+
- Add signed provenance and attestations to npm and GitHub releases.
|
package/agents/docs.toml
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
name = "docs"
|
|
2
|
+
description = "Documentation specialist for release and API consistency."
|
|
3
|
+
sandbox_mode = "workspace-write"
|
|
4
|
+
developer_instructions = """
|
|
5
|
+
Update only documentation required by the assigned change. Verify commands and version references against the repository.
|
|
6
|
+
"""
|
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
name = "implementer"
|
|
2
|
+
description = "Implementation specialist for one scoped story."
|
|
3
|
+
sandbox_mode = "workspace-write"
|
|
4
|
+
developer_instructions = """
|
|
5
|
+
Implement only the assigned scope. Use tests first, run verification, and do not review or commit your own work.
|
|
6
|
+
"""
|
package/bench/README.md
CHANGED
|
@@ -1,42 +1,45 @@
|
|
|
1
|
-
# Yoke benchmark — tokens · speed · quality
|
|
2
|
-
|
|
3
|
-
Reproducible cross-runner benchmark for the Yoke loop. One fixed fixture project, the same
|
|
4
|
-
PRD for every runner, three measured dimensions:
|
|
5
|
-
|
|
6
|
-
| Dimension | How it is measured |
|
|
7
|
-
|---|---|
|
|
8
|
-
| **Tokens** | The loop's own token hook (`.yoke/loop-status.json`, claude runner via `--output-format stream-json`; model id included). Gemini/Codex runners do not report usage yet — recorded as `null`, a documented gap. |
|
|
9
|
-
| **Speed** | Wall-clock, measured by the harness from outside: total run + per-story (from `--json` NDJSON event timestamps). The loop itself stores no durations. |
|
|
10
|
-
| **Quality** | Objective, not judged by any model: the fixture ships **pre-written tests** the agent never has to write (only satisfy). After the run, each story's test file is executed against the final tree. `srcLoc` (non-empty lines in `src/`) is a code-economy proxy. |
|
|
11
|
-
|
|
12
|
-
## The fixture (`fixtures/string-kit`)
|
|
13
|
-
|
|
14
|
-
A dependency-free ESM library with 3 stories (`slugify`, `truncate`, `titleCase`) and 16
|
|
15
|
-
`node:test` assertions total. `bench-verify.mjs` is cumulative: story N runs the tests of
|
|
16
|
-
stories 1…N (the loop exports `YOKE_STORY`), so later stories cannot break earlier work; the
|
|
17
|
-
final quality check runs everything. No npm installs, so results measure the agent — not the
|
|
18
|
-
network.
|
|
19
|
-
|
|
20
|
-
## Running it
|
|
21
|
-
|
|
22
|
-
```bash
|
|
23
|
-
npm run build
|
|
24
|
-
node bench/run.mjs --runner=claude # or gemini / codex
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
`
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
1
|
+
# Yoke benchmark — tokens · speed · quality
|
|
2
|
+
|
|
3
|
+
Reproducible cross-runner benchmark for the Yoke loop. One fixed fixture project, the same
|
|
4
|
+
PRD for every runner, three measured dimensions:
|
|
5
|
+
|
|
6
|
+
| Dimension | How it is measured |
|
|
7
|
+
|---|---|
|
|
8
|
+
| **Tokens** | The loop's own token hook (`.yoke/loop-status.json`, claude runner via `--output-format stream-json`; model id included). Gemini/Codex runners do not report usage yet — recorded as `null`, a documented gap. |
|
|
9
|
+
| **Speed** | Wall-clock, measured by the harness from outside: total run + per-story (from `--json` NDJSON event timestamps). The loop itself stores no durations. |
|
|
10
|
+
| **Quality** | Objective, not judged by any model: the fixture ships **pre-written tests** the agent never has to write (only satisfy). After the run, each story's test file is executed against the final tree. `srcLoc` (non-empty lines in `src/`) is a code-economy proxy. |
|
|
11
|
+
|
|
12
|
+
## The fixture (`fixtures/string-kit`)
|
|
13
|
+
|
|
14
|
+
A dependency-free ESM library with 3 stories (`slugify`, `truncate`, `titleCase`) and 16
|
|
15
|
+
`node:test` assertions total. `bench-verify.mjs` is cumulative: story N runs the tests of
|
|
16
|
+
stories 1…N (the loop exports `YOKE_STORY`), so later stories cannot break earlier work; the
|
|
17
|
+
final quality check runs everything. No npm installs, so results measure the agent — not the
|
|
18
|
+
network.
|
|
19
|
+
|
|
20
|
+
## Running it
|
|
21
|
+
|
|
22
|
+
```bash
|
|
23
|
+
npm run build
|
|
24
|
+
node bench/run.mjs --runner=claude # or gemini / codex
|
|
25
|
+
node bench/run-matrix.mjs --label=release-1.0
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
Each run copies the fixture to `bench/.runs/<runner>-<stamp>` (git-ignored), git-inits it,
|
|
29
|
+
drives `yoke loop run --json --max=6 --timeout=10`, and writes a result JSON to
|
|
30
|
+
`bench/results/`. Runs are billed against your own accounts for the agent CLIs involved.
|
|
31
|
+
The matrix is sequential to avoid cross-provider load distortion. Missing CLIs and authentication
|
|
32
|
+
failures are stored as honest `unavailable`/`auth-failed` rows and never presented as quality measurements.
|
|
33
|
+
|
|
34
|
+
## Caveats (read before quoting numbers)
|
|
35
|
+
|
|
36
|
+
- **N=1 per run.** Agent runs are stochastic; treat single runs as indicative, not
|
|
37
|
+
statistically robust. Re-run and compare.
|
|
38
|
+
- Model identity matters more than CLI identity: `tokens.model` records what actually served
|
|
39
|
+
the run. Different default models per CLI make "claude vs gemini" really "model X vs model Y".
|
|
40
|
+
- The fixture is deliberately small (a loop-overhead + basic-competence probe, minutes not
|
|
41
|
+
hours). It does not measure large-context refactoring, UI work, or long-horizon planning.
|
|
42
|
+
- Cumulative verify means a story's duration includes fixing any regressions it caused.
|
|
43
|
+
|
|
44
|
+
Results live in [`results/`](results/) — one JSON per run, summarized in
|
|
45
|
+
[`RESULTS.md`](RESULTS.md).
|