@hecer/yoke 0.8.0 → 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (69) hide show
  1. package/.codex-plugin/plugin.json +7 -0
  2. package/CHANGELOG.md +169 -130
  3. package/README.md +61 -21
  4. package/TODOS.md +8 -0
  5. package/agents/docs.toml +6 -0
  6. package/agents/implementer.toml +6 -0
  7. package/agents/reviewer.toml +6 -0
  8. package/agents/security.toml +6 -0
  9. package/bench/README.md +45 -42
  10. package/bench/RESULTS.md +46 -36
  11. package/bench/result-schema.mjs +12 -0
  12. package/bench/results/claude-2026-07-27T18-03-26.json +50 -0
  13. package/bench/results/codex-unavailable-1785175418318.json +15 -0
  14. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -0
  15. package/bench/run-matrix.mjs +26 -0
  16. package/bench/run.mjs +127 -115
  17. package/canon/loop/prd.schema.md +5 -0
  18. package/canon/manifest.yaml +3 -1
  19. package/canon/skills/authoring-prd/SKILL.md +14 -0
  20. package/canon/skills/performance/SKILL.md +48 -0
  21. package/canon/skills/ship/SKILL.md +2 -7
  22. package/canon/tools/codex-rtk-hook.mjs +36 -0
  23. package/dist/agents/providers.js +23 -0
  24. package/dist/agents/telemetry.js +30 -0
  25. package/dist/agents/types.js +1 -0
  26. package/dist/audit/changes.js +6 -0
  27. package/dist/audit/command.js +64 -0
  28. package/dist/audit/dependencies.js +21 -0
  29. package/dist/audit/secrets.js +16 -0
  30. package/dist/audit/types.js +1 -0
  31. package/dist/cli.js +22 -4
  32. package/dist/loop/claims.js +57 -0
  33. package/dist/loop/cleanup.js +10 -4
  34. package/dist/loop/git.js +8 -2
  35. package/dist/loop/identity.js +27 -0
  36. package/dist/loop/loop.js +55 -26
  37. package/dist/loop/merge-queue.js +20 -0
  38. package/dist/loop/parallel.js +39 -0
  39. package/dist/loop/prd.js +48 -2
  40. package/dist/loop/run-command.js +63 -7
  41. package/dist/loop/runner.js +47 -31
  42. package/dist/loop/scheduler.js +8 -0
  43. package/dist/prd/command.js +6 -0
  44. package/dist/retrofit/config.js +15 -0
  45. package/dist/retrofit/planners/codex.js +64 -19
  46. package/dist/review/command.js +52 -12
  47. package/dist/review/verdict.js +45 -0
  48. package/docs/MIGRATING-TO-1.0.md +33 -0
  49. package/docs/superpowers/plans/2026-07-27-yoke-1.0-release.md +205 -0
  50. package/docs/superpowers/specs/2026-07-27-yoke-1.0-hardening-and-codex-parity-design.md +164 -0
  51. package/hooks/hooks.json +19 -0
  52. package/package.json +82 -67
  53. package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/config.yaml +0 -6
  54. package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/context/DECISIONS.md +0 -9
  55. package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/prd.yaml +0 -38
  56. package/bench/.runs/claude-2026-07-09T22-34-01/bench-verify.mjs +0 -15
  57. package/bench/.runs/claude-2026-07-09T22-34-01/package.json +0 -9
  58. package/bench/.runs/claude-2026-07-09T22-34-01/src/index.mjs +0 -48
  59. package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-1.test.mjs +0 -24
  60. package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-2.test.mjs +0 -28
  61. package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-3.test.mjs +0 -25
  62. package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/config.yaml +0 -6
  63. package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/prd.yaml +0 -32
  64. package/bench/.runs/gemini-2026-07-09T22-34-02/bench-verify.mjs +0 -15
  65. package/bench/.runs/gemini-2026-07-09T22-34-02/package.json +0 -9
  66. package/bench/.runs/gemini-2026-07-09T22-34-02/src/index.mjs +0 -3
  67. package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-1.test.mjs +0 -24
  68. package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-2.test.mjs +0 -28
  69. package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-3.test.mjs +0 -25
@@ -0,0 +1,7 @@
1
+ {
2
+ "name": "yoke",
3
+ "version": "1.0.0",
4
+ "description": "Cross-agent coding discipline, mechanical gates, and release workflows",
5
+ "skills": "./canon/skills/",
6
+ "hooks": "./hooks/hooks.json"
7
+ }
package/CHANGELOG.md CHANGED
@@ -1,130 +1,169 @@
1
- # Changelog
2
-
3
- ## 0.8.0 — 2026-07-20
4
-
5
- ### Added
6
- - **Live progress + ETA.** Story completions are now first-class events: the console shows
7
- `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
8
- `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
9
- durations of stories completed **in this run** (current velocity) and falls back to the
10
- persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
11
- No data → no estimate, never an invented one.
12
- - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
13
- prompt now always forbids asking questions (a loop run has nobody to answer). Default
14
- `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
15
- loop never stops. Opt-in `abort`: the agent writes its open questions to
16
- `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
17
- story would otherwise pass on pre-existing green tests), and blocks with the question as
18
- the reason. Companion principle: clarifying questions belong in the planning round, before
19
- the loop starts.
20
-
21
- ## 0.7.0 — 2026-07-17
22
-
23
- ### Added
24
- - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
25
- version hint (npm/gh-style): a detached background refresher caches the registry's
26
- latest at most once a day; when it is newer, a one-line stderr hint suggests
27
- `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
28
- `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
29
- - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
30
- loop START only — never mid-run; the running process finishes on its version and
31
- the upgrade applies from the next invocation. Deliberately NOT the default:
32
- a gate harness must not change itself mid-project (determinism), and unreviewed
33
- auto-installs are a supply-chain hazard.
34
-
35
- ## 0.6.0 — 2026-07-17
36
-
37
- ### Added
38
- - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
39
- `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
40
- `yoke loop cleanup` kills exactly those recorded process trees — and only while no
41
- live loop holds the lock. Background: without a scoped mechanism, users and agents
42
- resorted to machine-wide pattern kills (every process matching
43
- `dangerously-skip-permissions`), which took down *healthy* runners of other projects
44
- mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
45
- safe path. `.yoke/runner.pid` is gitignored by retrofit.
46
-
47
- ## 0.5.0 — 2026-07-17
48
-
49
- Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
50
- (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
51
- idle window):
52
-
53
- ### Fixed
54
- - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
55
- Previously it killed only the spawned shell (`shell: true`), orphaning the actual
56
- agent process — which kept writing to the worktree (dirty-tree blocks, failing
57
- worktree removal) and kept burning API tokens. Observed in the field as ~10
58
- zombie `claude.exe` per machine plus surviving dev servers.
59
- - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
60
- the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
61
- processes and killed them at exactly the idle timeout — while the user saw dead air.
62
- The stream doubles as liveness; token usage is now reported on every run (not just
63
- `--json` mode).
64
-
65
- ### Changed
66
- - README: operating notes for driving the loop from inside an agent session
67
- (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
68
- outer shell-tool timeouts killing a foreground `yoke loop run` were the third
69
- observed "hang" pattern.
70
-
71
- ## 0.4.0 — 2026-07-17
72
-
73
- ### Added
74
- - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
75
- scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
76
- documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
77
- bounded final messages (cuts output-token waste). Review prompts now ground verdicts
78
- in observed evidence only and keep them brief.
79
-
80
- ### Fixed
81
- - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
82
- `git add -A` story commit swept the pause control file into history in
83
- un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
84
- blocked the resume run — the loop locked itself out.
85
-
86
- > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
87
- > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
88
-
89
- ## 0.3.0 — 2026-07-10
90
-
91
- ### Added
92
- - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
93
- (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
94
- then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
95
- - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
96
- installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
97
- daily-crawled extensions gallery.
98
- - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
99
- quality) with a fixed fixture, pre-written objective tests, and committed result data.
100
- - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
101
- sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
102
- (design generation paired with Yoke's design verification gates).
103
- - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
104
- stories across parallel workers (`needs` dependency field, claim files, merge queue,
105
- heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
106
-
107
- ### Fixed
108
- - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
109
- Windows, so the loop misreported an installed `gemini` as "not found on PATH"
110
- (found by the new benchmark harness).
111
- - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
112
- requires a value after `-p` and errored with "Not enough arguments following: p".
113
- Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
114
- (also found by the benchmark harness).
115
-
116
- ### Changed
117
- - README: npm install is now the primary quickstart path; documented plugin/extension
118
- installs and optional companions.
119
- - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
120
- `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
121
-
122
- ## 0.2.0 — 2026-07-09
123
-
124
- - First npm release as `@hecer/yoke`.
125
- - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
126
- token-usage + model-id reporting for the claude runner.
127
- - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
128
- `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
129
- - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
130
- loop with worktree isolation, watchdog, single-flight lock, commit integrity.
1
+ # Changelog
2
+
3
+ ## 1.0.0 — 2026-07-27
4
+
5
+ ### Added
6
+ - Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
7
+ - Safe provider permission profiles and structured cross-provider telemetry.
8
+ - Schema-validated independent review verdicts with explicit self-review opt-in.
9
+ - Human-owned commit identity enforcement; AI co-author trailers default off.
10
+ - `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
11
+ - PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
12
+ - Reproducible cross-runner benchmark schema and matrix launcher.
13
+
14
+ ### Changed
15
+ - Dangerous permission bypass is opt-in via `--unsafe`.
16
+ - Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
17
+ - Reviews no longer trust process exit code alone.
18
+
19
+ ### Security
20
+ - Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
21
+
22
+
23
+ ## 0.9.0 — 2026-07-22
24
+
25
+ ### Added
26
+ - **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
27
+ A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
28
+ verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
29
+ `--isolate` mode. A red benchmark blocks the story
30
+ (`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
31
+ gate is configured, the implementer prompt names the budget command so agents keep hot
32
+ paths efficient and never "simplify away" an optimization without re-running the benchmark.
33
+ - **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
34
+ clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
35
+ project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
36
+ tests, the *why* of every optimization versioned so future agents don't clean fast code
37
+ back to slow.
38
+ - **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
39
+ numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
40
+ round — a criterion still needing a decision is not loop-ready.
41
+
42
+ ## 0.8.0 — 2026-07-20
43
+
44
+ ### Added
45
+ - **Live progress + ETA.** Story completions are now first-class events: the console shows
46
+ `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
47
+ `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
48
+ durations of stories completed **in this run** (current velocity) and falls back to the
49
+ persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
50
+ No data → no estimate, never an invented one.
51
+ - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
52
+ prompt now always forbids asking questions (a loop run has nobody to answer). Default
53
+ `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
54
+ loop never stops. Opt-in `abort`: the agent writes its open questions to
55
+ `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
56
+ story would otherwise pass on pre-existing green tests), and blocks with the question as
57
+ the reason. Companion principle: clarifying questions belong in the planning round, before
58
+ the loop starts.
59
+
60
+ ## 0.7.0 — 2026-07-17
61
+
62
+ ### Added
63
+ - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
64
+ version hint (npm/gh-style): a detached background refresher caches the registry's
65
+ latest at most once a day; when it is newer, a one-line stderr hint suggests
66
+ `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
67
+ `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
68
+ - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
69
+ loop START only — never mid-run; the running process finishes on its version and
70
+ the upgrade applies from the next invocation. Deliberately NOT the default:
71
+ a gate harness must not change itself mid-project (determinism), and unreviewed
72
+ auto-installs are a supply-chain hazard.
73
+
74
+ ## 0.6.0 — 2026-07-17
75
+
76
+ ### Added
77
+ - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
78
+ `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
79
+ `yoke loop cleanup` kills exactly those recorded process trees — and only while no
80
+ live loop holds the lock. Background: without a scoped mechanism, users and agents
81
+ resorted to machine-wide pattern kills (every process matching
82
+ `dangerously-skip-permissions`), which took down *healthy* runners of other projects
83
+ mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
84
+ safe path. `.yoke/runner.pid` is gitignored by retrofit.
85
+
86
+ ## 0.5.0 — 2026-07-17
87
+
88
+ Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
89
+ (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
90
+ idle window):
91
+
92
+ ### Fixed
93
+ - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
94
+ Previously it killed only the spawned shell (`shell: true`), orphaning the actual
95
+ agent process — which kept writing to the worktree (dirty-tree blocks, failing
96
+ worktree removal) and kept burning API tokens. Observed in the field as ~10
97
+ zombie `claude.exe` per machine plus surviving dev servers.
98
+ - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
99
+ the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
100
+ processes and killed them at exactly the idle timeout — while the user saw dead air.
101
+ The stream doubles as liveness; token usage is now reported on every run (not just
102
+ `--json` mode).
103
+
104
+ ### Changed
105
+ - README: operating notes for driving the loop from inside an agent session
106
+ (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
107
+ outer shell-tool timeouts killing a foreground `yoke loop run` were the third
108
+ observed "hang" pattern.
109
+
110
+ ## 0.4.0 — 2026-07-17
111
+
112
+ ### Added
113
+ - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
114
+ scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
115
+ documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
116
+ bounded final messages (cuts output-token waste). Review prompts now ground verdicts
117
+ in observed evidence only and keep them brief.
118
+
119
+ ### Fixed
120
+ - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
121
+ `git add -A` story commit swept the pause control file into history in
122
+ un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
123
+ blocked the resume run — the loop locked itself out.
124
+
125
+ > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
126
+ > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
127
+
128
+ ## 0.3.0 — 2026-07-10
129
+
130
+ ### Added
131
+ - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
132
+ (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
133
+ then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
134
+ - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
135
+ installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
136
+ daily-crawled extensions gallery.
137
+ - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
138
+ quality) with a fixed fixture, pre-written objective tests, and committed result data.
139
+ - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
140
+ sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
141
+ (design generation paired with Yoke's design verification gates).
142
+ - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
143
+ stories across parallel workers (`needs` dependency field, claim files, merge queue,
144
+ heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
145
+
146
+ ### Fixed
147
+ - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
148
+ Windows, so the loop misreported an installed `gemini` as "not found on PATH"
149
+ (found by the new benchmark harness).
150
+ - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
151
+ requires a value after `-p` and errored with "Not enough arguments following: p".
152
+ Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
153
+ (also found by the benchmark harness).
154
+
155
+ ### Changed
156
+ - README: npm install is now the primary quickstart path; documented plugin/extension
157
+ installs and optional companions.
158
+ - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
159
+ `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
160
+
161
+ ## 0.2.0 — 2026-07-09
162
+
163
+ - First npm release as `@hecer/yoke`.
164
+ - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
165
+ token-usage + model-id reporting for the claude runner.
166
+ - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
167
+ `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
168
+ - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
169
+ loop with worktree isolation, watchdog, single-flight lock, commit integrity.
package/README.md CHANGED
@@ -2,22 +2,36 @@
2
2
 
3
3
  # 🐂 Yoke
4
4
 
5
+ <!-- yoke:version:start -->1.0.0<!-- yoke:version:end -->
6
+ <!-- yoke:tests:start -->500<!-- yoke:tests:end -->
7
+ <!-- yoke:skills:start -->28<!-- yoke:skills:end -->
8
+ <!-- yoke:agents:start -->Claude | Codex | Gemini<!-- yoke:agents:end -->
9
+
5
10
  ### One harness, three agents — and zero trust in "done."
6
11
 
7
12
  **Yoke** installs one curated canon of skills, **mechanical safety gates**, and tool wiring into any project — natively for **Claude Code, OpenAI Codex CLI, and Gemini CLI**. Then, when you want it, an opt-in autonomous loop ships your spec story-by-story: tested, cross-model-reviewed, committed — **with a screenshot to prove every story and a video for every failure**.
8
13
 
14
+ [![npm](https://img.shields.io/npm/v/%40hecer%2Fyoke?logo=npm&color=CB3837)](https://www.npmjs.com/package/@hecer/yoke)
15
+ [![npm downloads](https://img.shields.io/npm/dm/%40hecer%2Fyoke?logo=npm)](https://www.npmjs.com/package/@hecer/yoke)
9
16
  [![CI](https://github.com/HECer/yoke/actions/workflows/ci.yml/badge.svg)](https://github.com/HECer/yoke/actions/workflows/ci.yml)
10
17
  [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](#-license)
11
18
  ![Node](https://img.shields.io/badge/node-%E2%89%A520-339933?logo=node.js&logoColor=white)
12
19
  ![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
13
- ![Tests](https://img.shields.io/badge/tests-322%20passing-brightgreen.svg)
20
+ ![Tests](https://img.shields.io/badge/tests-500%20passing-brightgreen.svg)
14
21
  ![Agents](https://img.shields.io/badge/agents-Claude%20%7C%20Codex%20%7C%20Gemini-8A2BE2)
15
22
  ![Built with TDD](https://img.shields.io/badge/built%20with-TDD%20%2B%20review-ff69b4.svg)
16
23
 
24
+ **Install:** [`npm i -g @hecer/yoke`](https://www.npmjs.com/package/@hecer/yoke)
25
+
17
26
  </div>
18
27
 
19
28
  > **TL;DR** — `yoke new my-app --idea="..."` scaffolds a git repo, installs the harness for all three agents, and drafts a story backlog from your idea. `yoke loop run my-app --isolate --review` then implements it story by story behind hard gates: **clean tree → acceptance criteria → your real tests green → an independent model approves → commit**. If any gate is red, nothing is committed. When a story is done, there's a photo of it in `.yoke/proof/<story>/`.
20
29
 
30
+ Yoke 1.0 is safe-by-default: provider CLIs use autonomous sandbox profiles unless `--unsafe`
31
+ is explicit; reviews require a schema-valid verdict and a different model unless
32
+ `--allow-self-review` is explicit; commits enforce the human identity from project config or Git.
33
+ See [the 1.0 migration guide](docs/MIGRATING-TO-1.0.md).
34
+
21
35
  ---
22
36
 
23
37
  ## Why Yoke exists
@@ -29,7 +43,7 @@ Agentic coding in 2026 fails in four well-documented ways. Yoke answers each one
29
43
  | 🎭 **The verification gap** — *"agent says done, but it isn't"* | Agents submit confidently on 100% of runs while resolving far fewer; "all tests pass" when they were never run ([silent-failures research](https://arxiv.org/pdf/2603.25764)) | The loop trusts **your verify command's exit code**, never the agent's word. A story is `passes: true` only after tests are green, the reviewer approved, and the commit landed — atomically. Plus: **screenshot proofs** per story. |
30
44
  | 🔀 **Three agents, three configs** | Teams hand-maintain `CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, skills, and MCP wiring separately — copy-paste drift everywhere | **One canon → `yoke retrofit`** generates the idiomatic native artifacts for each agent. Change the canon once, re-retrofit everywhere. |
31
45
  | 🌀 **Overnight loops going off the rails** | Raw Ralph-loop users "wake up to broken codebases that don't compile" | Yoke is **"Ralph, but with gates"**: clean-worktree gate, acceptance-criteria gate, green-tests gate, review gate, per-story worktree isolation, idle-timeout watchdog, single-flight lock, commit integrity. |
32
- | 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a *second* model reviews the diff as a pass/fail exit-code gate — chainable into verify, pre-push, or CI. Cross-model review measurably catches what self-review misses. |
46
+ | 😵 **Review fatigue** | AI adoption nearly doubles PR volume and review time; humans start skimming | **`yoke review`**: a second model writes a schema-validated pass/fail verdict — chainable into verify, pre-push, or CI. Cross-model review catches what self-review misses. |
33
47
 
34
48
  **Who it's for:** anyone driving Claude Code, Codex CLI, or Gemini CLI on real projects — especially if you use more than one, want autonomous runs you can trust, or are tired of "done" meaning "probably". Greenfield (`yoke new`) and brownfield (`yoke retrofit`) both work.
35
49
 
@@ -46,16 +60,18 @@ $ yoke prd check reading-app
46
60
 
47
61
  $ yoke loop on reading-app
48
62
  $ yoke loop run reading-app --isolate --review --max=10
49
- ▶ STORY-1 (0/8) — implementing… · verifying… · reviewing… ✔ committed → 1/8
50
- ▶ STORY-2 (1/8) — implementing… · verifying… ✔ committed → 2/8
51
- ▶ STORY-3 (2/8) — implementing… · verifying… ✘ blocked: story did not verify (tests red)
63
+ ▶ STORY-1 (0/8 · 0%) — implementing… · verifying… · reviewing… · committing…
64
+ ✓ STORY-1 done in 3m12s — 1/8 (13%) · ~22m left
65
+ ▶ STORY-2 (1/8 · 13%) — implementing… · ~22m left (Ø 3m12s/story)
66
+ ✓ STORY-2 done in 2m48s — 2/8 (25%) · ~18m left
67
+ ▶ STORY-3 (2/8 · 25%) — implementing… ✘ blocked: story did not verify (tests red)
52
68
  # nothing was committed. fix, then re-run.
53
69
 
54
70
  $ ls reading-app/.yoke/proof/STORY-2/
55
71
  home.png list.png # photographic evidence, labelled per story
56
72
  ```
57
73
 
58
- Every claim in that transcript is enforced by code paths with tests behind them — 322 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
74
+ Every claim in that transcript is enforced by code paths with tests behind them — 500 of them, and this repo was built by its own loop and gates ([how it was built](#-why--how-it-was-built)).
59
75
 
60
76
  ## 🚀 Quickstart
61
77
 
@@ -130,10 +146,11 @@ Yoke's CLI is deterministic and chainable by design: an agent (or a shell `&&`)
130
146
  | `yoke new <dir> [--idea=] [--agent=] [--runner=] [--loop]` | Greenfield bootstrap: git init → scaffold → retrofit → context → PRD (drafted from `--idea`) → committed | `0` · `1` usage / non-empty dir / draft failed (scaffold survives) · `2` draft agent unavailable |
131
147
  | `yoke retrofit [dir] [--agent=claude,codex,gemini\|all] [--code-graph=graphify\|serena] [--loop]` | Install/update the harness, non-destructively | `0` |
132
148
  | `yoke prd draft [dir] --idea= [--runner=] [--force]` | Idea → 5–12 stories with testable acceptance criteria | `0` · `1` invalid/guarded · `2` agent unavailable |
133
- | `yoke prd check [dir]` | PRD lint gate (schema, duplicate ids, empty acceptance) | `0` valid · `1` violations |
149
+ | `yoke prd check [dir]` | PRD lint gate (schema, dependencies, cycles, duplicate ids, acceptance) | `0` valid · `1` violations |
134
150
  | `yoke context init\|status [dir]` | Durable context layer (`PROJECT/DECISIONS/KNOWLEDGE.md`) | `0` |
135
- | `yoke loop on\|off\|status\|run\|cleanup [dir]` | The autonomous loop (see below) | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
136
- | `yoke review [dir] [--reviewer=] [--base=] [--focus=]` | A **second model** reviews your diff | `0` approved · `1` findings · `2` no reviewer CLI |
151
+ | `yoke loop on\|off\|status\|run\|cleanup [dir]` | Autonomous loop; cleanup deletes worktrees only with `--remove-worktrees` | run: `0` complete · `1` blocked/cap · `2` not runnable / already locked · `3` paused |
152
+ | `yoke review [dir] [--reviewer=] [--base=] [--focus=] [--json] [--allow-self-review]` | An independent model writes a schema-valid verdict | `0` approved · `1` findings/invalid verdict · `2` no independent reviewer |
153
+ | `yoke audit [dir] [--json]` | Dependency, high-confidence secret, and sensitive-change audit | `0` green · `1` blocking findings · `2` not runnable |
137
154
  | `yoke design-scan [dir] [--max=N] [--report]` | Static AI-slop design gate | `0` within budget · `1` over |
138
155
  | `yoke flow-smoke [dir] [--url=] [--label=]` | Browser gate with screenshot/video proofs | `0` green · `1` failures · `2` not runnable |
139
156
 
@@ -150,7 +167,7 @@ Three excellent projects, three different jobs. Honest version:
150
167
  | **Enforcement** | Advisory — skills *describe* the discipline; following them is up to the agent | Skill-driven; browser QA is genuinely real | **Mechanical** — gates live in code: clean tree, acceptance criteria, green tests, review verdict, commit integrity |
151
168
  | **Autonomy** | Interactive sessions | Interactive slash-commands (`/qa`, `/ship`, …) | Opt-in **Ralph loop** with watchdog, worktree isolation, single-flight lock, per-story proofs |
152
169
  | **Visual QA** | — | **Best-in-class**: live browser daemon (Chromium/CDP) with deep interactive QA | Built-in `flow-smoke` gate: screenshots always, video on failure, labelled per story — lighter, but *enforced* and cross-agent |
153
- | **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves **codex → gemini → claude**, exit-code gate, works in and outside the loop |
170
+ | **Cross-model review** | — | `/codex` second opinion (Codex-only direction) | `yoke review` — resolves an independent provider and validates a structured verdict, inside or outside the loop |
154
171
  | **Footprint** | Markdown skills (plugin) | ~230 MB with browser runtime; hourly auto-update | Node CLI + markdown canon; Playwright only if you use flow-smoke, resolved **from your project** |
155
172
  | **License** | MIT | MIT | MIT |
156
173
 
@@ -183,10 +200,10 @@ Three layers — **Canon** (`yoke validate`) → **Retrofit** (`yoke retrofit`)
183
200
  | Agent | Artifacts |
184
201
  |---|---|
185
202
  | **Claude** | `.claude/skills/`, `AGENTS.md`, `CLAUDE.md`, `.mcp.json` (code-graph + Playwright), and an rtk `PreToolUse` hook when WSL is available |
186
- | **Codex** | `AGENTS.md` (native), `.codex/config.toml` (MCP servers), `RTK.md` |
203
+ | **Codex** | `.agents/skills/`, `AGENTS.md`, `RTK.md`, `.codex/config.toml`, native hooks, reusable `.codex/agents/*.toml`, and package plugin metadata |
187
204
  | **Gemini** | `GEMINI.md`, `.gemini/commands/*.toml` (one per skill, full body), `.gemini/settings.json` (MCP + `AGENTS.md` context) |
188
205
 
189
- > **rtk asymmetry, handled:** Claude can rewrite commands transparently via a hook (needs WSL on Windows); Codex and Gemini have no such hook, so they get an instruction to prefix commands with `rtk` instead.
206
+ > **rtk integration:** Claude receives its PreToolUse hook; Codex receives a native hook adapter around `rtk hook check`; Gemini retains instruction-mode fallback where its CLI has no equivalent command-rewrite lifecycle.
190
207
 
191
208
  > **Composes with gstack:** if [gstack](https://github.com/garrytan/gstack) is installed (repo-local or global), `yoke retrofit` adds a short "Composed tools" routing note to **CLAUDE.md only** — telling Claude to prefer gstack's skills for capabilities Yoke doesn't ship (live-browser QA `/qa`, security audit `/cso`, ship/deploy `/ship`). No bundling, no dependency; the note is never written to the Codex or Gemini artifacts.
192
209
 
@@ -197,7 +214,7 @@ Three layers — **Canon** (`yoke validate`) → **Retrofit** (`yoke retrofit`)
197
214
  > instructions (tech stack, workflow, `@`-includes) inside it. Works in any yoke-written file;
198
215
  > content *outside* the markers is still replaced (and backed up under `.yoke/backup/`).
199
216
 
200
- ## 🧰 What's in the canon — 27 skills
217
+ ## 🧰 What's in the canon — 28 skills
201
218
 
202
219
  `yoke retrofit` installs all of these into each agent natively. Provenance is credited in [`canon/skills/ATTRIBUTION.md`](canon/skills/ATTRIBUTION.md).
203
220
 
@@ -233,13 +250,14 @@ To stop overlapping skills from auto-invoking against each other, `canon/AGENTS.
233
250
  | `retro` | Engineering retrospective from commit history |
234
251
  | `document-release` | Post-ship documentation sync (README / CHANGELOG / …) |
235
252
 
236
- **Yoke-native** — *authored or adapted for this harness (7)*
253
+ **Yoke-native** — *authored or adapted for this harness (8)*
237
254
 
238
255
  | Skill | What it does |
239
256
  |---|---|
240
257
  | `yoke-retrofit` | Set up the Yoke harness in a project (detect → plan → apply) |
241
258
  | `authoring-prd` | Slice a product idea into loop-ready stories with testable acceptance criteria |
242
259
  | `minimal-code` | Write the least code that solves the task (YAGNI; ponytail-derived) |
260
+ | `performance` | Efficiency as a measured requirement: benchmarks as tests, budgets as gates, optimizations local + documented |
243
261
  | `maintaining-context` | Keep `.yoke/context/` the durable source of truth (the Context layer) |
244
262
  | `workflow` | The default order of operations, from idea to deploy |
245
263
  | `unslop-ui` | Detect & remove AI-slop design tells (purple gradients, neon glow, emoji-icons…) |
@@ -302,6 +320,7 @@ yoke loop run . \
302
320
  --runner=codex \ # implement with Codex…
303
321
  --reviewer=claude \ # …review with Claude (role separation)
304
322
  --isolate \ # each story in a throwaway git worktree
323
+ --on-ambiguity=abort \ # strict: stop on undecidable criteria instead of guessing
305
324
  --max=20
306
325
  yoke loop off . # disable
307
326
  ```
@@ -381,6 +400,30 @@ Enable strict mode per run with `yoke loop run . --on-ambiguity=abort` or per pr
381
400
  put every clarifying question into the PRD **before** the loop starts (`yoke prd draft`
382
401
  criteria must be testable and decision-free).
383
402
 
403
+ ### Performance budgets: efficiency as a gate, not a style
404
+
405
+ Clean code is the default (the `minimal-code` skill) — but when efficiency matters, "should
406
+ be fast" is a vibe the loop cannot enforce. Yoke makes it mechanical, at two levels:
407
+
408
+ - **Per story:** write the requirement as a **measurable acceptance criterion**
409
+ ("imports 1M rows in < 2s, asserted by the bench test") and let your verify tests measure
410
+ it — no new machinery needed.
411
+ - **Per project:** wire a benchmark as a standing **perf gate** in `.yoke/config.yaml`:
412
+
413
+ ```yaml
414
+ perf:
415
+ command: node bench/check-budget.mjs # exit 0 = within budget
416
+ retries: 1 # benchmarks are noisy; same retry logic as verify
417
+ ```
418
+
419
+ The loop runs it **after verify** on every story (phase `perf`, with `YOKE_STORY` set); a
420
+ red benchmark blocks the story — `story S6 exceeded its performance budget: p95 62ms > budget 50ms` —
421
+ no matter how clean the diff was. The implementer prompt names the budget command, so the
422
+ agent knows not to trade hot-path efficiency for style and never "simplifies away" an
423
+ optimization without re-running the benchmark. The `performance` canon skill carries the
424
+ method: profile first, optimize leaves not boundaries, commit benchmarks as tests, version
425
+ the *why* of every optimization in `context/DECISIONS.md`.
426
+
384
427
  The loop trusts **verify**, not the agent's exit code: a story whose tests are green is
385
428
  committed even if the agent process exited non-zero (a common Windows `.cmd`-wrapper ghost).
386
429
  A failing verify is retried up to `verify.retries` times (default 1) so a transient flake
@@ -555,17 +598,14 @@ docs/superpowers/ # the spec and every component's implementation plan
555
598
 
556
599
  ## 🗺️ Roadmap
557
600
 
558
- - **npm publish** (`@hecer/yoke`) — one-liner `npx` install (package prepared).
559
- - **Security gate** — a `cso`-style audit skill + `yoke audit` (deps, secrets, diff surface).
560
- - **Token-budget gate** — per-story budget with abort; cost transparency for loop runs.
561
- - **Multi-reviewer quorum** — N independent reviewers with distinct lenses (correctness / security / acceptance).
562
- - **Merge queue** — re-test against the latest main before integrating, for parallel/multi-agent loops.
563
- - **More agents** — the canon→retrofit pattern generalises; OpenCode and Copilot CLI are natural next targets.
601
+ Yoke 1.0's completed release work moved to the changelog. Remaining, explicitly scoped work
602
+ is tracked in [`TODOS.md`](TODOS.md), including provider subprocess wiring for the tested
603
+ parallel dispatcher, broader benchmark samples, native output schemas, and release provenance.
564
604
 
565
605
  ## 🧪 Development
566
606
 
567
607
  ```bash
568
- npm test # vitest (322 tests)
608
+ npm test # vitest (500 tests)
569
609
  npm run build # tsc, no emit errors
570
610
  npm run yoke -- validate canon
571
611
  ```
package/TODOS.md ADDED
@@ -0,0 +1,8 @@
1
+ # Yoke follow-up work
2
+
3
+ - Wire the tested async parallel dispatcher to provider subprocess workers. Until then the CLI
4
+ rejects `--parallel=N` for `N > 1`; scheduler, claims, and merge queue APIs are available
5
+ without claiming a CLI speed-up.
6
+ - Add provider-native output schemas when all three CLIs expose compatible stable APIs.
7
+ - Expand benchmark fixtures and collect multiple authenticated samples per provider/model.
8
+ - Add signed provenance and attestations to npm and GitHub releases.
@@ -0,0 +1,6 @@
1
+ name = "docs"
2
+ description = "Documentation specialist for release and API consistency."
3
+ sandbox_mode = "workspace-write"
4
+ developer_instructions = """
5
+ Update only documentation required by the assigned change. Verify commands and version references against the repository.
6
+ """
@@ -0,0 +1,6 @@
1
+ name = "implementer"
2
+ description = "Implementation specialist for one scoped story."
3
+ sandbox_mode = "workspace-write"
4
+ developer_instructions = """
5
+ Implement only the assigned scope. Use tests first, run verification, and do not review or commit your own work.
6
+ """
@@ -0,0 +1,6 @@
1
+ name = "reviewer"
2
+ description = "Read-only reviewer for correctness and acceptance criteria."
3
+ sandbox_mode = "read-only"
4
+ developer_instructions = """
5
+ Review observed diffs and test evidence. Do not modify files. Return only findings grounded in evidence.
6
+ """
@@ -0,0 +1,6 @@
1
+ name = "security"
2
+ description = "Read-only security reviewer for changed code."
3
+ sandbox_mode = "read-only"
4
+ developer_instructions = """
5
+ Inspect changed code for exploitable security regressions. Do not modify files and avoid speculative findings.
6
+ """