@hecer/yoke 0.9.0 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (86) hide show
  1. package/.claude-plugin/marketplace.json +18 -0
  2. package/.claude-plugin/plugin.json +13 -0
  3. package/.codex-plugin/plugin.json +7 -0
  4. package/CHANGELOG.md +192 -149
  5. package/README.md +101 -51
  6. package/TODOS.md +8 -0
  7. package/agents/docs.toml +6 -0
  8. package/agents/implementer.toml +6 -0
  9. package/agents/reviewer.toml +6 -0
  10. package/agents/security.toml +6 -0
  11. package/bench/README.md +45 -42
  12. package/bench/RESULTS.md +46 -36
  13. package/bench/result-schema.mjs +12 -0
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -0
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -0
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -0
  17. package/bench/run-matrix.mjs +26 -0
  18. package/bench/run.mjs +127 -115
  19. package/canon/AGENTS.md +2 -0
  20. package/canon/loop/loop-spec.md +4 -2
  21. package/canon/loop/prd.schema.md +5 -0
  22. package/canon/manifest.yaml +2 -1
  23. package/canon/skills/authoring-prd/SKILL.md +10 -3
  24. package/canon/skills/ship/SKILL.md +2 -7
  25. package/canon/skills/workflow/SKILL.md +4 -0
  26. package/canon/skills/yoke-retrofit/SKILL.md +18 -11
  27. package/canon/skills/yoke-workflow/SKILL.md +20 -0
  28. package/canon/tools/codex-rtk-hook.mjs +36 -0
  29. package/dist/agents/host.js +26 -0
  30. package/dist/agents/providers.js +23 -0
  31. package/dist/agents/telemetry.js +30 -0
  32. package/dist/agents/types.js +1 -0
  33. package/dist/audit/changes.js +6 -0
  34. package/dist/audit/command.js +64 -0
  35. package/dist/audit/dependencies.js +21 -0
  36. package/dist/audit/secrets.js +16 -0
  37. package/dist/audit/types.js +1 -0
  38. package/dist/cli.js +189 -6
  39. package/dist/context/context.js +15 -2
  40. package/dist/loop/claims.js +57 -0
  41. package/dist/loop/cleanup.js +98 -27
  42. package/dist/loop/decision.js +517 -0
  43. package/dist/loop/git.js +31 -2
  44. package/dist/loop/identity.js +27 -0
  45. package/dist/loop/lock.js +104 -13
  46. package/dist/loop/loop.js +49 -2
  47. package/dist/loop/merge-queue.js +20 -0
  48. package/dist/loop/parallel.js +39 -0
  49. package/dist/loop/prd.js +48 -2
  50. package/dist/loop/run-command.js +118 -12
  51. package/dist/loop/runner.js +48 -30
  52. package/dist/loop/scheduler.js +8 -0
  53. package/dist/prd/command.js +30 -21
  54. package/dist/retrofit/command.js +3 -2
  55. package/dist/retrofit/config.js +16 -0
  56. package/dist/retrofit/gitignore.js +8 -0
  57. package/dist/retrofit/planners/codex.js +64 -19
  58. package/dist/retrofit/report.js +1 -1
  59. package/dist/review/command.js +52 -12
  60. package/dist/review/verdict.js +45 -0
  61. package/dist/setup/command.js +82 -0
  62. package/docs/MIGRATING-TO-1.0.md +33 -0
  63. package/docs/MIGRATING-TO-1.1.md +27 -0
  64. package/docs/PUBLISHING.md +77 -41
  65. package/docs/superpowers/plans/2026-07-27-yoke-1.0-release.md +205 -0
  66. package/docs/superpowers/specs/2026-07-27-yoke-1.0-hardening-and-codex-parity-design.md +164 -0
  67. package/gemini-extension.json +6 -0
  68. package/hooks/hooks.json +19 -0
  69. package/package.json +84 -67
  70. package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/config.yaml +0 -6
  71. package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/context/DECISIONS.md +0 -9
  72. package/bench/.runs/claude-2026-07-09T22-34-01/.yoke/prd.yaml +0 -38
  73. package/bench/.runs/claude-2026-07-09T22-34-01/bench-verify.mjs +0 -15
  74. package/bench/.runs/claude-2026-07-09T22-34-01/package.json +0 -9
  75. package/bench/.runs/claude-2026-07-09T22-34-01/src/index.mjs +0 -48
  76. package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-1.test.mjs +0 -24
  77. package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-2.test.mjs +0 -28
  78. package/bench/.runs/claude-2026-07-09T22-34-01/tests/STORY-3.test.mjs +0 -25
  79. package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/config.yaml +0 -6
  80. package/bench/.runs/gemini-2026-07-09T22-34-02/.yoke/prd.yaml +0 -32
  81. package/bench/.runs/gemini-2026-07-09T22-34-02/bench-verify.mjs +0 -15
  82. package/bench/.runs/gemini-2026-07-09T22-34-02/package.json +0 -9
  83. package/bench/.runs/gemini-2026-07-09T22-34-02/src/index.mjs +0 -3
  84. package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-1.test.mjs +0 -24
  85. package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-2.test.mjs +0 -28
  86. package/bench/.runs/gemini-2026-07-09T22-34-02/tests/STORY-3.test.mjs +0 -25
@@ -0,0 +1,18 @@
1
+ {
2
+ "name": "yoke",
3
+ "owner": { "name": "HECer" },
4
+ "metadata": { "description": "Yoke — cross-agent coding harness for Claude Code, Codex CLI, and Gemini CLI" },
5
+ "plugins": [
6
+ {
7
+ "name": "yoke",
8
+ "source": "./",
9
+ "description": "One curated skill canon plus mechanical safety gates and an autonomous loop (yoke CLI: npm i -g @hecer/yoke).",
10
+ "author": { "name": "HECer" },
11
+ "homepage": "https://github.com/HECer/yoke",
12
+ "repository": "https://github.com/HECer/yoke",
13
+ "license": "MIT",
14
+ "keywords": ["harness", "cross-agent", "tdd", "code-review", "codex", "gemini-cli"],
15
+ "category": "development"
16
+ }
17
+ ]
18
+ }
@@ -0,0 +1,13 @@
1
+ {
2
+ "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
3
+ "name": "yoke",
4
+ "displayName": "Yoke",
5
+ "version": "1.1.0",
6
+ "description": "Cross-agent coding harness: one curated skill canon (TDD, brainstorming, plans, reviews, shipping, design verification) plus mechanical safety gates and an autonomous loop via the yoke CLI.",
7
+ "author": { "name": "HECer", "url": "https://github.com/HECer" },
8
+ "homepage": "https://github.com/HECer/yoke#readme",
9
+ "repository": "https://github.com/HECer/yoke",
10
+ "license": "MIT",
11
+ "keywords": ["harness", "cross-agent", "tdd", "code-review", "autonomous-loop", "codex", "gemini-cli"],
12
+ "skills": "./canon/skills/"
13
+ }
@@ -0,0 +1,7 @@
1
+ {
2
+ "name": "yoke",
3
+ "version": "1.1.0",
4
+ "description": "Cross-agent coding discipline, mechanical gates, and release workflows",
5
+ "skills": "./canon/skills/",
6
+ "hooks": "./hooks/hooks.json"
7
+ }
package/CHANGELOG.md CHANGED
@@ -1,149 +1,192 @@
1
- # Changelog
2
-
3
- ## 0.9.0 — 2026-07-22
4
-
5
- ### Added
6
- - **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
7
- A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
8
- verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
9
- `--isolate` mode. A red benchmark blocks the story
10
- (`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
11
- gate is configured, the implementer prompt names the budget command so agents keep hot
12
- paths efficient and never "simplify away" an optimization without re-running the benchmark.
13
- - **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
14
- clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
15
- project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
16
- tests, the *why* of every optimization versioned so future agents don't clean fast code
17
- back to slow.
18
- - **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
19
- numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
20
- round — a criterion still needing a decision is not loop-ready.
21
-
22
- ## 0.8.0 — 2026-07-20
23
-
24
- ### Added
25
- - **Live progress + ETA.** Story completions are now first-class events: the console shows
26
- `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
27
- `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
28
- durations of stories completed **in this run** (current velocity) and falls back to the
29
- persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
30
- No data → no estimate, never an invented one.
31
- - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
32
- prompt now always forbids asking questions (a loop run has nobody to answer). Default
33
- `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
34
- loop never stops. Opt-in `abort`: the agent writes its open questions to
35
- `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
36
- story would otherwise pass on pre-existing green tests), and blocks with the question as
37
- the reason. Companion principle: clarifying questions belong in the planning round, before
38
- the loop starts.
39
-
40
- ## 0.7.0 — 2026-07-17
41
-
42
- ### Added
43
- - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
44
- version hint (npm/gh-style): a detached background refresher caches the registry's
45
- latest at most once a day; when it is newer, a one-line stderr hint suggests
46
- `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
47
- `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
48
- - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
49
- loop START only — never mid-run; the running process finishes on its version and
50
- the upgrade applies from the next invocation. Deliberately NOT the default:
51
- a gate harness must not change itself mid-project (determinism), and unreviewed
52
- auto-installs are a supply-chain hazard.
53
-
54
- ## 0.6.0 — 2026-07-17
55
-
56
- ### Added
57
- - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
58
- `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
59
- `yoke loop cleanup` kills exactly those recorded process trees — and only while no
60
- live loop holds the lock. Background: without a scoped mechanism, users and agents
61
- resorted to machine-wide pattern kills (every process matching
62
- `dangerously-skip-permissions`), which took down *healthy* runners of other projects
63
- mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
64
- safe path. `.yoke/runner.pid` is gitignored by retrofit.
65
-
66
- ## 0.5.0 — 2026-07-17
67
-
68
- Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
69
- (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
70
- idle window):
71
-
72
- ### Fixed
73
- - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
74
- Previously it killed only the spawned shell (`shell: true`), orphaning the actual
75
- agent process — which kept writing to the worktree (dirty-tree blocks, failing
76
- worktree removal) and kept burning API tokens. Observed in the field as ~10
77
- zombie `claude.exe` per machine plus surviving dev servers.
78
- - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
79
- the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
80
- processes and killed them at exactly the idle timeout — while the user saw dead air.
81
- The stream doubles as liveness; token usage is now reported on every run (not just
82
- `--json` mode).
83
-
84
- ### Changed
85
- - README: operating notes for driving the loop from inside an agent session
86
- (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
87
- outer shell-tool timeouts killing a foreground `yoke loop run` were the third
88
- observed "hang" pattern.
89
-
90
- ## 0.4.0 — 2026-07-17
91
-
92
- ### Added
93
- - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
94
- scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
95
- documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
96
- bounded final messages (cuts output-token waste). Review prompts now ground verdicts
97
- in observed evidence only and keep them brief.
98
-
99
- ### Fixed
100
- - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
101
- `git add -A` story commit swept the pause control file into history in
102
- un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
103
- blocked the resume run — the loop locked itself out.
104
-
105
- > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
106
- > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
107
-
108
- ## 0.3.0 — 2026-07-10
109
-
110
- ### Added
111
- - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
112
- (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
113
- then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
114
- - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
115
- installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
116
- daily-crawled extensions gallery.
117
- - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
118
- quality) with a fixed fixture, pre-written objective tests, and committed result data.
119
- - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
120
- sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
121
- (design generation paired with Yoke's design verification gates).
122
- - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
123
- stories across parallel workers (`needs` dependency field, claim files, merge queue,
124
- heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
125
-
126
- ### Fixed
127
- - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
128
- Windows, so the loop misreported an installed `gemini` as "not found on PATH"
129
- (found by the new benchmark harness).
130
- - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
131
- requires a value after `-p` and errored with "Not enough arguments following: p".
132
- Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
133
- (also found by the benchmark harness).
134
-
135
- ### Changed
136
- - README: npm install is now the primary quickstart path; documented plugin/extension
137
- installs and optional companions.
138
- - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
139
- `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
140
-
141
- ## 0.2.0 — 2026-07-09
142
-
143
- - First npm release as `@hecer/yoke`.
144
- - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
145
- token-usage + model-id reporting for the claude runner.
146
- - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
147
- `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
148
- - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
149
- loop with worktree isolation, watchdog, single-flight lock, commit integrity.
1
+ # Changelog
2
+
3
+ ## 1.1.0 — 2026-07-30
4
+
5
+ ### Added
6
+ - Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
7
+ - Persisted default runner selection and `auto|critical` loop decision policies.
8
+ - Provider-neutral `yoke-workflow` skill for planning questions, approved-plan PRD handoff, autonomous story execution, and critical-decision resume.
9
+ - Structured critical-decision requests with `yoke loop decision` and `yoke loop answer`; answers are validated, committed under the configured human identity, and resume the same story.
10
+ - Approved `.yoke/plan.md` context in PRD drafting and a lint gate for unresolved planning placeholders.
11
+
12
+ ### Fixed
13
+ - Retrofit and loop on/off now preserve timeout, decision, runner, and permission settings.
14
+ - Empty projects prefer the active agent host instead of silently installing Claude artifacts in Codex.
15
+ - Loop and PRD runner selection now prefer an explicit flag, then the configured runner, then the active host.
16
+ - Retrofit reports no longer label every provider as Claude Code.
17
+ - Critical-decision resumes retain isolation, review, runner, permissions, timeout, JSON, policy, and iteration settings instead of falling back to an unreviewed default run.
18
+ - Decision answers use an atomic owner-token lock and recoverable request journal, are checked against the active PRD story, bounded as untrusted data, and committed path-by-path so unrelated edits cannot enter the human-owned commit.
19
+ - Decision recovery now binds the exact selected answer to its commit, rolls back only its own interrupted context append, namespaces resume state per project/worktree, and serializes cleanup with loop startup.
20
+ - Active agent session markers now outrank globally configured provider home directories, and setup rejects partially invalid agent lists.
21
+
22
+ ### Changed
23
+ - New setups enable the loop by default and choose `decisionPolicy: auto`; the wizard can select `critical` or disable the loop.
24
+ - Legacy `loop.onAmbiguity` and `--on-ambiguity` remain compatibility aliases.
25
+
26
+ ## 1.0.0 — 2026-07-27
27
+
28
+ ### Added
29
+ - Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
30
+ - Safe provider permission profiles and structured cross-provider telemetry.
31
+ - Schema-validated independent review verdicts with explicit self-review opt-in.
32
+ - Human-owned commit identity enforcement; AI co-author trailers default off.
33
+ - `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
34
+ - PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
35
+ - Reproducible cross-runner benchmark schema and matrix launcher.
36
+
37
+ ### Changed
38
+ - Dangerous permission bypass is opt-in via `--unsafe`.
39
+ - Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
40
+ - Reviews no longer trust process exit code alone.
41
+
42
+ ### Security
43
+ - Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
44
+
45
+
46
+ ## 0.9.0 — 2026-07-22
47
+
48
+ ### Added
49
+ - **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
50
+ A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
51
+ verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
52
+ `--isolate` mode. A red benchmark blocks the story
53
+ (`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
54
+ gate is configured, the implementer prompt names the budget command so agents keep hot
55
+ paths efficient and never "simplify away" an optimization without re-running the benchmark.
56
+ - **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
57
+ clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
58
+ project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
59
+ tests, the *why* of every optimization versioned so future agents don't clean fast code
60
+ back to slow.
61
+ - **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
62
+ numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
63
+ round — a criterion still needing a decision is not loop-ready.
64
+
65
+ ## 0.8.0 — 2026-07-20
66
+
67
+ ### Added
68
+ - **Live progress + ETA.** Story completions are now first-class events: the console shows
69
+ `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
70
+ `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
71
+ durations of stories completed **in this run** (current velocity) and falls back to the
72
+ persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
73
+ No data → no estimate, never an invented one.
74
+ - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
75
+ prompt now always forbids asking questions (a loop run has nobody to answer). Default
76
+ `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
77
+ loop never stops. Opt-in `abort`: the agent writes its open questions to
78
+ `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
79
+ story would otherwise pass on pre-existing green tests), and blocks with the question as
80
+ the reason. Companion principle: clarifying questions belong in the planning round, before
81
+ the loop starts.
82
+
83
+ ## 0.7.0 — 2026-07-17
84
+
85
+ ### Added
86
+ - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
87
+ version hint (npm/gh-style): a detached background refresher caches the registry's
88
+ latest at most once a day; when it is newer, a one-line stderr hint suggests
89
+ `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
90
+ `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
91
+ - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
92
+ loop START only — never mid-run; the running process finishes on its version and
93
+ the upgrade applies from the next invocation. Deliberately NOT the default:
94
+ a gate harness must not change itself mid-project (determinism), and unreviewed
95
+ auto-installs are a supply-chain hazard.
96
+
97
+ ## 0.6.0 — 2026-07-17
98
+
99
+ ### Added
100
+ - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
101
+ `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
102
+ `yoke loop cleanup` kills exactly those recorded process trees — and only while no
103
+ live loop holds the lock. Background: without a scoped mechanism, users and agents
104
+ resorted to machine-wide pattern kills (every process matching
105
+ `dangerously-skip-permissions`), which took down *healthy* runners of other projects
106
+ mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
107
+ safe path. `.yoke/runner.pid` is gitignored by retrofit.
108
+
109
+ ## 0.5.0 — 2026-07-17
110
+
111
+ Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
112
+ (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
113
+ idle window):
114
+
115
+ ### Fixed
116
+ - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
117
+ Previously it killed only the spawned shell (`shell: true`), orphaning the actual
118
+ agent process — which kept writing to the worktree (dirty-tree blocks, failing
119
+ worktree removal) and kept burning API tokens. Observed in the field as ~10
120
+ zombie `claude.exe` per machine plus surviving dev servers.
121
+ - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
122
+ the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
123
+ processes and killed them at exactly the idle timeout — while the user saw dead air.
124
+ The stream doubles as liveness; token usage is now reported on every run (not just
125
+ `--json` mode).
126
+
127
+ ### Changed
128
+ - README: operating notes for driving the loop from inside an agent session
129
+ (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
130
+ outer shell-tool timeouts killing a foreground `yoke loop run` were the third
131
+ observed "hang" pattern.
132
+
133
+ ## 0.4.0 — 2026-07-17
134
+
135
+ ### Added
136
+ - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
137
+ scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
138
+ documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
139
+ bounded final messages (cuts output-token waste). Review prompts now ground verdicts
140
+ in observed evidence only and keep them brief.
141
+
142
+ ### Fixed
143
+ - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
144
+ `git add -A` story commit swept the pause control file into history in
145
+ un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
146
+ blocked the resume run — the loop locked itself out.
147
+
148
+ > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
149
+ > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
150
+
151
+ ## 0.3.0 — 2026-07-10
152
+
153
+ ### Added
154
+ - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
155
+ (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
156
+ then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
157
+ - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
158
+ installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
159
+ daily-crawled extensions gallery.
160
+ - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
161
+ quality) with a fixed fixture, pre-written objective tests, and committed result data.
162
+ - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
163
+ sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
164
+ (design generation paired with Yoke's design verification gates).
165
+ - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
166
+ stories across parallel workers (`needs` dependency field, claim files, merge queue,
167
+ heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
168
+
169
+ ### Fixed
170
+ - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
171
+ Windows, so the loop misreported an installed `gemini` as "not found on PATH"
172
+ (found by the new benchmark harness).
173
+ - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
174
+ requires a value after `-p` and errored with "Not enough arguments following: p".
175
+ Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
176
+ (also found by the benchmark harness).
177
+
178
+ ### Changed
179
+ - README: npm install is now the primary quickstart path; documented plugin/extension
180
+ installs and optional companions.
181
+ - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
182
+ `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
183
+
184
+ ## 0.2.0 — 2026-07-09
185
+
186
+ - First npm release as `@hecer/yoke`.
187
+ - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
188
+ token-usage + model-id reporting for the claude runner.
189
+ - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
190
+ `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
191
+ - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
192
+ loop with worktree isolation, watchdog, single-flight lock, commit integrity.