@hecer/yoke 0.9.0 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (68) hide show
  1. package/.codex-plugin/plugin.json +7 -0
  2. package/CHANGELOG.md +169 -149
  3. package/README.md +24 -16
  4. package/TODOS.md +8 -0
  5. package/agents/docs.toml +6 -0
  6. package/agents/implementer.toml +6 -0
  7. package/agents/reviewer.toml +6 -0
  8. package/agents/security.toml +6 -0
  9. package/bench/README.md +45 -42
  10. package/bench/RESULTS.md +46 -36
  11. package/bench/result-schema.mjs +12 -0
  12. package/bench/results/claude-2026-07-27T18-03-26.json +50 -0
  13. package/bench/results/codex-unavailable-1785175418318.json +15 -0
  14. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -0
  15. package/bench/run-matrix.mjs +26 -0
  16. package/bench/run.mjs +127 -115
  17. package/canon/loop/prd.schema.md +5 -0
  18. package/canon/manifest.yaml +1 -1
  19. package/canon/skills/authoring-prd/SKILL.md +6 -0
  20. package/canon/skills/ship/SKILL.md +2 -7
  21. package/canon/tools/codex-rtk-hook.mjs +36 -0
  22. package/dist/agents/providers.js +23 -0
  23. package/dist/agents/telemetry.js +30 -0
  24. package/dist/agents/types.js +1 -0
  25. package/dist/audit/changes.js +6 -0
  26. package/dist/audit/command.js +64 -0
  27. package/dist/audit/dependencies.js +21 -0
  28. package/dist/audit/secrets.js +16 -0
  29. package/dist/audit/types.js +1 -0
  30. package/dist/cli.js +22 -4
  31. package/dist/loop/claims.js +57 -0
  32. package/dist/loop/cleanup.js +10 -4
  33. package/dist/loop/git.js +8 -2
  34. package/dist/loop/identity.js +27 -0
  35. package/dist/loop/loop.js +20 -2
  36. package/dist/loop/merge-queue.js +20 -0
  37. package/dist/loop/parallel.js +39 -0
  38. package/dist/loop/prd.js +48 -2
  39. package/dist/loop/run-command.js +51 -6
  40. package/dist/loop/runner.js +41 -29
  41. package/dist/loop/scheduler.js +8 -0
  42. package/dist/prd/command.js +6 -0
  43. package/dist/retrofit/config.js +12 -0
  44. package/dist/retrofit/planners/codex.js +64 -19
  45. package/dist/review/command.js +52 -12
  46. package/dist/review/verdict.js +45 -0
  47. package/docs/MIGRATING-TO-1.0.md +33 -0
  48. package/docs/superpowers/plans/2026-07-27-yoke-1.0-release.md +205 -0
  49. package/docs/superpowers/specs/2026-07-27-yoke-1.0-hardening-and-codex-parity-design.md +164 -0
  50. package/hooks/hooks.json +19 -0
  51. package/package.json +82 -67
  52. package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/config.yaml +0 -6
  53. package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/context/DECISIONS.md +0 -9
  54. package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/prd.yaml +0 -38
  55. package/bench/.runs/claude-2026-07-09T22-34-01/bench-verify.mjs +0 -15
  56. package/bench/.runs/claude-2026-07-09T22-34-01/package.json +0 -9
  57. package/bench/.runs/claude-2026-07-09T22-34-01/src/index.mjs +0 -48
  58. package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-1.test.mjs +0 -24
  59. package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-2.test.mjs +0 -28
  60. package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-3.test.mjs +0 -25
  61. package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/config.yaml +0 -6
  62. package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/prd.yaml +0 -32
  63. package/bench/.runs/gemini-2026-07-09T22-34-02/bench-verify.mjs +0 -15
  64. package/bench/.runs/gemini-2026-07-09T22-34-02/package.json +0 -9
  65. package/bench/.runs/gemini-2026-07-09T22-34-02/src/index.mjs +0 -3
  66. package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-1.test.mjs +0 -24
  67. package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-2.test.mjs +0 -28
  68. package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-3.test.mjs +0 -25
@@ -0,0 +1,7 @@
1
+ {
2
+ "name": "yoke",
3
+ "version": "1.0.0",
4
+ "description": "Cross-agent coding discipline, mechanical gates, and release workflows",
5
+ "skills": "./canon/skills/",
6
+ "hooks": "./hooks/hooks.json"
7
+ }
package/CHANGELOG.md CHANGED
@@ -1,149 +1,169 @@
1
- # Changelog
2
-
3
- ## 0.9.0 — 2026-07-22
4
-
5
- ### Added
6
- - **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
7
- A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
8
- verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
9
- `--isolate` mode. A red benchmark blocks the story
10
- (`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
11
- gate is configured, the implementer prompt names the budget command so agents keep hot
12
- paths efficient and never "simplify away" an optimization without re-running the benchmark.
13
- - **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
14
- clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
15
- project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
16
- tests, the *why* of every optimization versioned so future agents don't clean fast code
17
- back to slow.
18
- - **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
19
- numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
20
- round — a criterion still needing a decision is not loop-ready.
21
-
22
- ## 0.8.0 — 2026-07-20
23
-
24
- ### Added
25
- - **Live progress + ETA.** Story completions are now first-class events: the console shows
26
- `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
27
- `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
28
- durations of stories completed **in this run** (current velocity) and falls back to the
29
- persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
30
- No data → no estimate, never an invented one.
31
- - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
32
- prompt now always forbids asking questions (a loop run has nobody to answer). Default
33
- `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
34
- loop never stops. Opt-in `abort`: the agent writes its open questions to
35
- `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
36
- story would otherwise pass on pre-existing green tests), and blocks with the question as
37
- the reason. Companion principle: clarifying questions belong in the planning round, before
38
- the loop starts.
39
-
40
- ## 0.7.0 — 2026-07-17
41
-
42
- ### Added
43
- - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
44
- version hint (npm/gh-style): a detached background refresher caches the registry's
45
- latest at most once a day; when it is newer, a one-line stderr hint suggests
46
- `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
47
- `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
48
- - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
49
- loop START only — never mid-run; the running process finishes on its version and
50
- the upgrade applies from the next invocation. Deliberately NOT the default:
51
- a gate harness must not change itself mid-project (determinism), and unreviewed
52
- auto-installs are a supply-chain hazard.
53
-
54
- ## 0.6.0 — 2026-07-17
55
-
56
- ### Added
57
- - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
58
- `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
59
- `yoke loop cleanup` kills exactly those recorded process trees — and only while no
60
- live loop holds the lock. Background: without a scoped mechanism, users and agents
61
- resorted to machine-wide pattern kills (every process matching
62
- `dangerously-skip-permissions`), which took down *healthy* runners of other projects
63
- mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
64
- safe path. `.yoke/runner.pid` is gitignored by retrofit.
65
-
66
- ## 0.5.0 — 2026-07-17
67
-
68
- Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
69
- (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
70
- idle window):
71
-
72
- ### Fixed
73
- - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
74
- Previously it killed only the spawned shell (`shell: true`), orphaning the actual
75
- agent process — which kept writing to the worktree (dirty-tree blocks, failing
76
- worktree removal) and kept burning API tokens. Observed in the field as ~10
77
- zombie `claude.exe` per machine plus surviving dev servers.
78
- - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
79
- the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
80
- processes and killed them at exactly the idle timeout — while the user saw dead air.
81
- The stream doubles as liveness; token usage is now reported on every run (not just
82
- `--json` mode).
83
-
84
- ### Changed
85
- - README: operating notes for driving the loop from inside an agent session
86
- (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
87
- outer shell-tool timeouts killing a foreground `yoke loop run` were the third
88
- observed "hang" pattern.
89
-
90
- ## 0.4.0 — 2026-07-17
91
-
92
- ### Added
93
- - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
94
- scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
95
- documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
96
- bounded final messages (cuts output-token waste). Review prompts now ground verdicts
97
- in observed evidence only and keep them brief.
98
-
99
- ### Fixed
100
- - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
101
- `git add -A` story commit swept the pause control file into history in
102
- un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
103
- blocked the resume run — the loop locked itself out.
104
-
105
- > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
106
- > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
107
-
108
- ## 0.3.0 — 2026-07-10
109
-
110
- ### Added
111
- - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
112
- (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
113
- then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
114
- - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
115
- installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
116
- daily-crawled extensions gallery.
117
- - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
118
- quality) with a fixed fixture, pre-written objective tests, and committed result data.
119
- - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
120
- sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
121
- (design generation paired with Yoke's design verification gates).
122
- - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
123
- stories across parallel workers (`needs` dependency field, claim files, merge queue,
124
- heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
125
-
126
- ### Fixed
127
- - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
128
- Windows, so the loop misreported an installed `gemini` as "not found on PATH"
129
- (found by the new benchmark harness).
130
- - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
131
- requires a value after `-p` and errored with "Not enough arguments following: p".
132
- Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
133
- (also found by the benchmark harness).
134
-
135
- ### Changed
136
- - README: npm install is now the primary quickstart path; documented plugin/extension
137
- installs and optional companions.
138
- - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
139
- `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
140
-
141
- ## 0.2.0 — 2026-07-09
142
-
143
- - First npm release as `@hecer/yoke`.
144
- - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
145
- token-usage + model-id reporting for the claude runner.
146
- - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
147
- `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
148
- - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
149
- loop with worktree isolation, watchdog, single-flight lock, commit integrity.
1
+ # Changelog
2
+
3
+ ## 1.0.0 — 2026-07-27
4
+
5
+ ### Added
6
+ - Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
7
+ - Safe provider permission profiles and structured cross-provider telemetry.
8
+ - Schema-validated independent review verdicts with explicit self-review opt-in.
9
+ - Human-owned commit identity enforcement; AI co-author trailers default off.
10
+ - `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
11
+ - PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
12
+ - Reproducible cross-runner benchmark schema and matrix launcher.
13
+
14
+ ### Changed
15
+ - Dangerous permission bypass is opt-in via `--unsafe`.
16
+ - Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
17
+ - Reviews no longer trust process exit code alone.
18
+
19
+ ### Security
20
+ - Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
21
+
22
+
23
+ ## 0.9.0 — 2026-07-22
24
+
25
+ ### Added
26
+ - **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
27
+ A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
28
+ verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
29
+ `--isolate` mode. A red benchmark blocks the story
30
+ (`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
31
+ gate is configured, the implementer prompt names the budget command so agents keep hot
32
+ paths efficient and never "simplify away" an optimization without re-running the benchmark.
33
+ - **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
34
+ clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
35
+ project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
36
+ tests, the *why* of every optimization versioned so future agents don't clean fast code
37
+ back to slow.
38
+ - **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
39
+ numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
40
+ round — a criterion still needing a decision is not loop-ready.
41
+
42
+ ## 0.8.0 — 2026-07-20
43
+
44
+ ### Added
45
+ - **Live progress + ETA.** Story completions are now first-class events: the console shows
46
+ `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
47
+ `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
48
+ durations of stories completed **in this run** (current velocity) and falls back to the
49
+ persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
50
+ No data → no estimate, never an invented one.
51
+ - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
52
+ prompt now always forbids asking questions (a loop run has nobody to answer). Default
53
+ `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
54
+ loop never stops. Opt-in `abort`: the agent writes its open questions to
55
+ `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
56
+ story would otherwise pass on pre-existing green tests), and blocks with the question as
57
+ the reason. Companion principle: clarifying questions belong in the planning round, before
58
+ the loop starts.
59
+
60
+ ## 0.7.0 — 2026-07-17
61
+
62
+ ### Added
63
+ - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
64
+ version hint (npm/gh-style): a detached background refresher caches the registry's
65
+ latest at most once a day; when it is newer, a one-line stderr hint suggests
66
+ `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
67
+ `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
68
+ - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
69
+ loop START only — never mid-run; the running process finishes on its version and
70
+ the upgrade applies from the next invocation. Deliberately NOT the default:
71
+ a gate harness must not change itself mid-project (determinism), and unreviewed
72
+ auto-installs are a supply-chain hazard.
73
+
74
+ ## 0.6.0 — 2026-07-17
75
+
76
+ ### Added
77
+ - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
78
+ `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
79
+ `yoke loop cleanup` kills exactly those recorded process trees — and only while no
80
+ live loop holds the lock. Background: without a scoped mechanism, users and agents
81
+ resorted to machine-wide pattern kills (every process matching
82
+ `dangerously-skip-permissions`), which took down *healthy* runners of other projects
83
+ mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
84
+ safe path. `.yoke/runner.pid` is gitignored by retrofit.
85
+
86
+ ## 0.5.0 — 2026-07-17
87
+
88
+ Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
89
+ (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
90
+ idle window):
91
+
92
+ ### Fixed
93
+ - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
94
+ Previously it killed only the spawned shell (`shell: true`), orphaning the actual
95
+ agent process — which kept writing to the worktree (dirty-tree blocks, failing
96
+ worktree removal) and kept burning API tokens. Observed in the field as ~10
97
+ zombie `claude.exe` per machine plus surviving dev servers.
98
+ - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
99
+ the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
100
+ processes and killed them at exactly the idle timeout — while the user saw dead air.
101
+ The stream doubles as liveness; token usage is now reported on every run (not just
102
+ `--json` mode).
103
+
104
+ ### Changed
105
+ - README: operating notes for driving the loop from inside an agent session
106
+ (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
107
+ outer shell-tool timeouts killing a foreground `yoke loop run` were the third
108
+ observed "hang" pattern.
109
+
110
+ ## 0.4.0 — 2026-07-17
111
+
112
+ ### Added
113
+ - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
114
+ scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
115
+ documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
116
+ bounded final messages (cuts output-token waste). Review prompts now ground verdicts
117
+ in observed evidence only and keep them brief.
118
+
119
+ ### Fixed
120
+ - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
121
+ `git add -A` story commit swept the pause control file into history in
122
+ un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
123
+ blocked the resume run — the loop locked itself out.
124
+
125
+ > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
126
+ > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
127
+
128
+ ## 0.3.0 — 2026-07-10
129
+
130
+ ### Added
131
+ - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
132
+ (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
133
+ then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
134
+ - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
135
+ installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
136
+ daily-crawled extensions gallery.
137
+ - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
138
+ quality) with a fixed fixture, pre-written objective tests, and committed result data.
139
+ - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
140
+ sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
141
+ (design generation paired with Yoke's design verification gates).
142
+ - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
143
+ stories across parallel workers (`needs` dependency field, claim files, merge queue,
144
+ heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
145
+
146
+ ### Fixed
147
+ - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
148
+ Windows, so the loop misreported an installed `gemini` as "not found on PATH"
149
+ (found by the new benchmark harness).
150
+ - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
151
+ requires a value after `-p` and errored with "Not enough arguments following: p".
152
+ Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
153
+ (also found by the benchmark harness).
154
+
155
+ ### Changed
156
+ - README: npm install is now the primary quickstart path; documented plugin/extension
157
+ installs and optional companions.
158
+ - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
159
+ `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
160
+
161
+ ## 0.2.0 — 2026-07-09
162
+
163
+ - First npm release as `@hecer/yoke`.
164
+ - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
165
+ token-usage + model-id reporting for the claude runner.
166
+ - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
167
+ `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
168
+ - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
169
+ loop with worktree isolation, watchdog, single-flight lock, commit integrity.
package/README.md CHANGED
@@ -2,6 +2,11 @@
2
2
 
3
3
  # 🐂 Yoke
4
4
 
5
+ <!-- yoke:version:start -->1.0.0<!-- yoke:version:end -->
6
+ <!-- yoke:tests:start -->500<!-- yoke:tests:end -->
7
+ <!-- yoke:skills:start -->28<!-- yoke:skills:end -->
8
+ <!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
9
+
5
10
  ### One harness, three agents — and zero trust in "done."
6
11
 
7
12
  **Yoke** installs one curated canon of skills, **mechanical safety gates**, and tool wiring into any project — natively for **Claude Code, OpenAI Codex CLI, and Gemini CLI**. Then, when you want it, an opt-in autonomous loop ships your spec story-by-story: tested, cross-model-reviewed, committed — **with a screenshot to prove every story and a video for every failure**.
@@ -12,7 +17,7 @@
12
17
  [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](#-license)
13
18
  ![Node](https://img.shields.io/badge/node-%E2%89%A520-339933?logo=node.js&logoColor=white)
14
19
  ![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
15
- ![Tests](https://img.shields.io/badge/tests-439%20passing-brightgreen.svg)
20
+ ![Tests](https://img.shields.io/badge/tests-500%20passing-brightgreen.svg)
16
21
  ![Agents](https://img.shields.io/badge/agents-Claude%20%7C%20Codex%20%7C%20Gemini-8A2BE2)
17
22
  ![Built with TDD](https://img.shields.io/badge/built%20with-TDD%20%2B%20review-ff69b4.svg)
18
23
 
@@ -22,6 +27,11 @@
22
27
 
23
28
  > **TL;DR** — `yoke new my-app --idea="..."` scaffolds a git repo, installs the harness for all three agents, and drafts a story backlog from your idea. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
24
29
 
30
+ Yoke 1.0 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
31
+ is explicit; reviews require a schema-valid verdict and a different model unless
32
+ `--allow-self-review` is explicit; commits enforce the human identity from project config or Git.
33
+ See [the 1.0 migration guide](docs/MIGRATING-TO-1.0.md).
34
+
25
35
  ---
26
36
 
27
37
  ## Why Yoke exists
@@ -33,7 +43,7 @@ Agentic coding in 2026 fails in four well-documented ways. Yoke answers each one
33
43
  | 🎭 **The verification gap** — *"agent says done, but it isn't"* | Agents submit confidently on 100% of runs while resolving far fewer; "all tests pass" when they were never run ([silent-failures research](https://arxiv.org/pdf/2603.25764)) | The loop trusts **your verify command's exit code**, never the agent's word. A story is `passes: true` only after tests are green, the reviewer approved, and the commit landed — atomically. Plus: **screenshot proofs** per story. |
34
44
  | 🔀 **Three agents, three configs** | Teams hand-maintain `CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, skills, and MCP wiring separately — copy-paste drift everywhere | **One canon → `yoke retrofit`** generates the idiomatic native artifacts for each agent. Change the canon once, re-retrofit everywhere. |
35
45
  | 🌀 **Overnight loops going off the rails** | Raw Ralph-loop users "wake up to broken codebases that don't compile" | Yoke is **"Ralph, but with gates"**: clean-worktree gate, acceptance-criteria gate, green-tests gate, review gate, per-story worktree isolation, idle-timeout watchdog, single-flight lock, commit integrity. |
36
- | 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a *second* model reviews the diff as a pass/fail exit-code gate — chainable into verify, pre-push, or CI. Cross-model review measurably catches what self-review misses. |
46
+ | 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a second model writes a schema-validated pass/fail verdict — chainable into verify, pre-push, or CI. Cross-model review catches what self-review misses. |
37
47
 
38
48
  **Who it's for:** anyone driving Claude Code, Codex CLI, or Gemini CLI on real projects — especially if you use more than one, want autonomous runs you can trust, or are tired of "done" meaning "probably". Greenfield (`yoke new`) and brownfield (`yoke retrofit`) both work.
39
49
 
@@ -61,7 +71,7 @@ $ ls reading-app/.yoke/proof/STORY-2/
61
71
  home.png list.png # photographic evidence, labelled per story
62
72
  ```
63
73
 
64
- Every claim in that transcript is enforced by code paths with tests behind them — 439 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
74
+ Every claim in that transcript is enforced by code paths with tests behind them — 500 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
65
75
 
66
76
  ## 🚀 Quickstart
67
77
 
@@ -136,10 +146,11 @@ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`)
136
146
  | `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
137
147
  | `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
138
148
  | `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
139
- | `yoke prd check [dir]` | PRD lint gate (schema, duplicate ids, empty acceptance) | `0` valid · `1` violations |
149
+ | `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
140
150
  | `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
141
- | `yoke loop on\|off\|status\|run\|cleanup [dir]` | The autonomous loop (see below) | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
142
- | `yoke review [dir] [--reviewer=] [--base=] [--focus=]` | A **second model** reviews your diff | `0` approved · `1` findings · `2` no reviewer CLI |
151
+ | `yoke loop on\|off\|status\|run\|cleanup [dir]` | Autonomous loop; cleanup deletes worktrees only with `--remove-worktrees` | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
152
+ | `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
153
+ | `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
143
154
  | `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
144
155
  | `yoke flow-smoke [dir] [--url=] [--label=]` | Browser gate with screenshot/video proofs | `0` green · `1` failures · `2` not runnable |
145
156
 
@@ -156,7 +167,7 @@ Three excellent projects, three different jobs. Honest version:
156
167
  | **Enforcement** | Advisory — skills *describe* the discipline; following them is up to the agent | Skill-driven; browser QA is genuinely real | **Mechanical** — gates live in code: clean tree, acceptance criteria, green tests, review verdict, commit integrity |
157
168
  | **Autonomy** | Interactive sessions | Interactive slash-commands (`/qa`, `/ship`, …) | Opt-in **Ralph loop** with watchdog, worktree isolation, single-flight lock, per-story proofs |
158
169
  | **Visual QA** | — | **Best-in-class**: live browser daemon (Chromium/CDP) with deep interactive QA | Built-in `flow-smoke` gate: screenshots always, video on failure, labelled per story — lighter, but *enforced* and cross-agent |
159
- | **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves **codex → gemini → claude**, exit-code gate, works in and outside the loop |
170
+ | **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves an independent provider and validates a structured verdict, inside or outside the loop |
160
171
  | **Footprint** | Markdown skills (plugin) | ~230 MB with browser runtime; hourly auto-update | Node CLI + markdown canon; Playwright only if you use flow-smoke, resolved **from your project** |
161
172
  | **License** | MIT | MIT | MIT |
162
173
 
@@ -189,10 +200,10 @@ Three layers — **Canon** (`yoke validate`) → **Retrofit** (`yoke retrofit`)
189
200
  | Agent | Artifacts |
190
201
  |---|---|
191
202
  | **Claude** | `.claude/skills/`, `AGENTS.md`, `CLAUDE.md`, `.mcp.json` (code-graph + Playwright), and an rtk `PreToolUse` hook when WSL is available |
192
- | **Codex** | `AGENTS.md` (native), `.codex/config.toml` (MCP servers), `RTK.md` |
203
+ | **Codex** | `.agents/skills/`, `AGENTS.md`, `RTK.md`, `.codex/config.toml`, native hooks, reusable `.codex/agents/*.toml`, and package plugin metadata |
193
204
  | **Gemini** | `GEMINI.md`, `.gemini/commands/*.toml` (one per skill, full body), `.gemini/settings.json` (MCP + `AGENTS.md` context) |
194
205
 
195
- > **rtk asymmetry, handled:** Claude can rewrite commands transparently via a hook (needs WSL on Windows); Codex and Gemini have no such hook, so they get an instruction to prefix commands with `rtk` instead.
206
+ > **rtk integration:** Claude receives its PreToolUse hook; Codex receives a native hook adapter around `rtk hook check`; Gemini retains instruction-mode fallback where its CLI has no equivalent command-rewrite lifecycle.
196
207
 
197
208
  > **Composes with gstack:** if [gstack](https://github.com/garrytan/gstack) is installed (repo-local or global), `yoke retrofit` adds a short "Composed tools" routing note to **CLAUDE.md only** — telling Claude to prefer gstack's skills for capabilities Yoke doesn't ship (live-browser QA `/qa`, security audit `/cso`, ship/deploy `/ship`). No bundling, no dependency; the note is never written to the Codex or Gemini artifacts.
198
209
 
@@ -587,17 +598,14 @@ docs/superpowers/ # the spec and every component's implementation plan
587
598
 
588
599
  ## 🗺️ Roadmap
589
600
 
590
- - **npm publish** (`@hecer/yoke`) — one-liner `npx` install (package prepared).
591
- - **Security gate** — a `cso`-style audit skill + `yoke audit` (deps, secrets, diff surface).
592
- - **Token-budget gate** — per-story budget with abort; cost transparency for loop runs.
593
- - **Multi-reviewer quorum** — N independent reviewers with distinct lenses (correctness / security / acceptance).
594
- - **Merge queue** — re-test against the latest main before integrating, for parallel/multi-agent loops.
595
- - **More agents** — the canon→retrofit pattern generalises; OpenCode and Copilot CLI are natural next targets.
601
+ Yoke 1.0's completed release work moved to the changelog. Remaining, explicitly scoped work
602
+ is tracked in [`TODOS.md`](TODOS.md), including provider subprocess wiring for the tested
603
+ parallel dispatcher, broader benchmark samples, native output schemas, and release provenance.
596
604
 
597
605
  ## 🧪 Development
598
606
 
599
607
  ```bash
600
- npm test # vitest (322 tests)
608
+ npm test # vitest (500 tests)
601
609
  npm run build # tsc, no emit errors
602
610
  npm run yoke -- validate canon
603
611
  ```
package/TODOS.md ADDED
@@ -0,0 +1,8 @@
1
+ # Yoke follow-up work
2
+
3
+ - Wire the tested async parallel dispatcher to provider subprocess workers. Until then the CLI
4
+ rejects `--parallel=N` for `N > 1`; scheduler, claims, and merge queue APIs are available
5
+ without claiming a CLI speed-up.
6
+ - Add provider-native output schemas when all three CLIs expose compatible stable APIs.
7
+ - Expand benchmark fixtures and collect multiple authenticated samples per provider/model.
8
+ - Add signed provenance and attestations to npm and GitHub releases.
@@ -0,0 +1,6 @@
1
+ name = "docs"
2
+ description = "Documentation specialist for release and API consistency."
3
+ sandbox_mode = "workspace-write"
4
+ developer_instructions = """
5
+ Update only documentation required by the assigned change. Verify commands and version references against the repository.
6
+ """
@@ -0,0 +1,6 @@
1
+ name = "implementer"
2
+ description = "Implementation specialist for one scoped story."
3
+ sandbox_mode = "workspace-write"
4
+ developer_instructions = """
5
+ Implement only the assigned scope. Use tests first, run verification, and do not review or commit your own work.
6
+ """
@@ -0,0 +1,6 @@
1
+ name = "reviewer"
2
+ description = "Read-only reviewer for correctness and acceptance criteria."
3
+ sandbox_mode = "read-only"
4
+ developer_instructions = """
5
+ Review observed diffs and test evidence. Do not modify files. Return only findings grounded in evidence.
6
+ """
@@ -0,0 +1,6 @@
1
+ name = "security"
2
+ description = "Read-only security reviewer for changed code."
3
+ sandbox_mode = "read-only"
4
+ developer_instructions = """
5
+ Inspect changed code for exploitable security regressions. Do not modify files and avoid speculative findings.
6
+ """
package/bench/README.md CHANGED
@@ -1,42 +1,45 @@
1
- # Yoke benchmark — tokens · speed · quality
2
-
3
- Reproducible cross-runner benchmark for the Yoke loop. One fixed fixture project, the same
4
- PRD for every runner, three measured dimensions:
5
-
6
- | Dimension | How it is measured |
7
- |---|---|
8
- | **Tokens** | The loop's own token hook (`.yoke/loop-status.json`, claude runner via `--output-format stream-json`; model id included). Gemini/Codex runners do not report usage yet — recorded as `null`, a documented gap. |
9
- | **Speed** | Wall-clock, measured by the harness from outside: total run + per-story (from `--json` NDJSON event timestamps). The loop itself stores no durations. |
10
- | **Quality** | Objective, not judged by any model: the fixture ships **pre-written tests** the agent never has to write (only satisfy). After the run, each story's test file is executed against the final tree. `srcLoc` (non-empty lines in `src/`) is a code-economy proxy. |
11
-
12
- ## The fixture (`fixtures/string-kit`)
13
-
14
- A dependency-free ESM library with 3 stories (`slugify`, `truncate`, `titleCase`) and 16
15
- `node:test` assertions total. `bench-verify.mjs` is cumulative: story N runs the tests of
16
- stories 1…N (the loop exports `YOKE_STORY`), so later stories cannot break earlier work; the
17
- final quality check runs everything. No npm installs, so results measure the agent — not the
18
- network.
19
-
20
- ## Running it
21
-
22
- ```bash
23
- npm run build
24
- node bench/run.mjs --runner=claude # or gemini / codex
25
- ```
26
-
27
- Each run copies the fixture to `bench/.runs/<runner>-<stamp>` (git-ignored), git-inits it,
28
- drives `yoke loop run --json --max=6 --timeout=10`, and writes a result JSON to
29
- `bench/results/`. Runs are billed against your own accounts for the agent CLIs involved.
30
-
31
- ## Caveats (read before quoting numbers)
32
-
33
- - **N=1 per run.** Agent runs are stochastic; treat single runs as indicative, not
34
- statistically robust. Re-run and compare.
35
- - Model identity matters more than CLI identity: `tokens.model` records what actually served
36
- the run. Different default models per CLI make "claude vs gemini" really "model X vs model Y".
37
- - The fixture is deliberately small (a loop-overhead + basic-competence probe, minutes not
38
- hours). It does not measure large-context refactoring, UI work, or long-horizon planning.
39
- - Cumulative verify means a story's duration includes fixing any regressions it caused.
40
-
41
- Results live in [`results/`](results/) — one JSON per run, summarized in
42
- [`RESULTS.md`](RESULTS.md).
1
+ # Yoke benchmark — tokens · speed · quality
2
+
3
+ Reproducible cross-runner benchmark for the Yoke loop. One fixed fixture project, the same
4
+ PRD for every runner, three measured dimensions:
5
+
6
+ | Dimension | How it is measured |
7
+ |---|---|
8
+ | **Tokens** | The loop's own token hook (`.yoke/loop-status.json`, claude runner via `--output-format stream-json`; model id included). Gemini/Codex runners do not report usage yet — recorded as `null`, a documented gap. |
9
+ | **Speed** | Wall-clock, measured by the harness from outside: total run + per-story (from `--json` NDJSON event timestamps). The loop itself stores no durations. |
10
+ | **Quality** | Objective, not judged by any model: the fixture ships **pre-written tests** the agent never has to write (only satisfy). After the run, each story's test file is executed against the final tree. `srcLoc` (non-empty lines in `src/`) is a code-economy proxy. |
11
+
12
+ ## The fixture (`fixtures/string-kit`)
13
+
14
+ A dependency-free ESM library with 3 stories (`slugify`, `truncate`, `titleCase`) and 16
15
+ `node:test` assertions total. `bench-verify.mjs` is cumulative: story N runs the tests of
16
+ stories 1…N (the loop exports `YOKE_STORY`), so later stories cannot break earlier work; the
17
+ final quality check runs everything. No npm installs, so results measure the agent — not the
18
+ network.
19
+
20
+ ## Running it
21
+
22
+ ```bash
23
+ npm run build
24
+ node bench/run.mjs --runner=claude # or gemini / codex
25
+ node bench/run-matrix.mjs --label=release-1.0
26
+ ```
27
+
28
+ Each run copies the fixture to `bench/.runs/<runner>-<stamp>` (git-ignored), git-inits it,
29
+ drives `yoke loop run --json --max=6 --timeout=10`, and writes a result JSON to
30
+ `bench/results/`. Runs are billed against your own accounts for the agent CLIs involved.
31
+ The matrix is sequential to avoid cross-provider load distortion. Missing CLIs and authentication
32
+ failures are stored as honest `unavailable`/`auth-failed` rows and never presented as quality measurements.
33
+
34
+ ## Caveats (read before quoting numbers)
35
+
36
+ - **N=1 per run.** Agent runs are stochastic; treat single runs as indicative, not
37
+ statistically robust. Re-run and compare.
38
+ - Model identity matters more than CLI identity: `tokens.model` records what actually served
39
+ the run. Different default models per CLI make "claude vs gemini" really "model X vs model Y".
40
+ - The fixture is deliberately small (a loop-overhead + basic-competence probe, minutes not
41
+ hours). It does not measure large-context refactoring, UI work, or long-horizon planning.
42
+ - Cumulative verify means a story's duration includes fixing any regressions it caused.
43
+
44
+ Results live in [`results/`](results/) — one JSON per run, summarized in
45
+ [`RESULTS.md`](RESULTS.md).