@hecer/yoke 1.6.0 → 1.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (103) hide show
  1. package/.claude-plugin/plugin.json +13 -13
  2. package/.codex-plugin/plugin.json +7 -7
  3. package/CHANGELOG.md +294 -288
  4. package/README.md +874 -874
  5. package/TODOS.md +5 -5
  6. package/agents/docs.toml +6 -6
  7. package/agents/implementer.toml +6 -6
  8. package/agents/reviewer.toml +6 -6
  9. package/agents/security.toml +6 -6
  10. package/bench/README.md +86 -86
  11. package/bench/RESULTS.md +35 -35
  12. package/bench/output-compaction.mjs +65 -65
  13. package/bench/result-schema.mjs +12 -12
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -15
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
  17. package/bench/run-matrix.mjs +26 -26
  18. package/bench/run.mjs +106 -106
  19. package/canon/AGENTS.md +30 -30
  20. package/canon/context/DECISIONS.md +4 -4
  21. package/canon/context/GLOSSARY.md +11 -11
  22. package/canon/context/KNOWLEDGE.md +4 -4
  23. package/canon/context/PROJECT.md +15 -15
  24. package/canon/loop/loop-spec.md +65 -65
  25. package/canon/loop/prd.schema.md +43 -43
  26. package/canon/manifest.yaml +59 -59
  27. package/canon/policy/gates.md +7 -7
  28. package/canon/policy/roles.md +9 -9
  29. package/canon/skills/ATTRIBUTION.md +99 -99
  30. package/canon/skills/authoring-prd/SKILL.md +58 -58
  31. package/canon/skills/brainstorming/SKILL.md +164 -164
  32. package/canon/skills/codebase-design/DEEPENING.md +15 -15
  33. package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
  34. package/canon/skills/codebase-design/SKILL.md +39 -39
  35. package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
  36. package/canon/skills/document-release/SKILL.md +302 -302
  37. package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
  38. package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
  39. package/canon/skills/domain-modeling/SKILL.md +35 -35
  40. package/canon/skills/executing-plans/SKILL.md +70 -70
  41. package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
  42. package/canon/skills/health/SKILL.md +177 -177
  43. package/canon/skills/maintaining-context/SKILL.md +34 -34
  44. package/canon/skills/minimal-code/SKILL.md +21 -21
  45. package/canon/skills/no-ai-slop/SKILL.md +103 -103
  46. package/canon/skills/no-ai-slop/eval.md +43 -43
  47. package/canon/skills/plan-ceo-review/SKILL.md +541 -541
  48. package/canon/skills/plan-eng-review/SKILL.md +362 -362
  49. package/canon/skills/receiving-code-review/SKILL.md +213 -213
  50. package/canon/skills/requesting-code-review/SKILL.md +105 -105
  51. package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
  52. package/canon/skills/retro/SKILL.md +397 -397
  53. package/canon/skills/review/SKILL.md +246 -246
  54. package/canon/skills/ship/SKILL.md +691 -691
  55. package/canon/skills/subagent-driven-development/SKILL.md +277 -277
  56. package/canon/skills/systematic-debugging/SKILL.md +296 -296
  57. package/canon/skills/tdd/SKILL.md +371 -371
  58. package/canon/skills/unslop-ui/SKILL.md +34 -34
  59. package/canon/skills/using-git-worktrees/SKILL.md +218 -218
  60. package/canon/skills/verification-before-completion/SKILL.md +139 -139
  61. package/canon/skills/visual-verification/SKILL.md +54 -54
  62. package/canon/skills/workflow/SKILL.md +22 -22
  63. package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
  64. package/canon/skills/writing-for-agents/SKILL.md +42 -42
  65. package/canon/skills/writing-plans/SKILL.md +152 -152
  66. package/canon/skills/writing-skills/SKILL.md +655 -655
  67. package/canon/skills/yoke-retrofit/SKILL.md +26 -26
  68. package/canon/skills/yoke-workflow/SKILL.md +20 -20
  69. package/canon/tools/codex-rtk-hook.mjs +35 -35
  70. package/canon/tools/graphify.md +3 -3
  71. package/canon/tools/playwright-mcp.md +3 -3
  72. package/canon/tools/rtk.md +7 -7
  73. package/canon/tools/serena.md +6 -6
  74. package/dist/agents/process.js +3 -0
  75. package/dist/loop/watchdog.js +1 -1
  76. package/dist/prd/command.js +17 -17
  77. package/dist/retrofit/planners/claude.js +14 -14
  78. package/dist/retrofit/preserve.js +2 -2
  79. package/docs/MIGRATING-TO-1.0.md +33 -33
  80. package/docs/MIGRATING-TO-1.1.md +27 -27
  81. package/docs/MIGRATING-TO-1.4.md +70 -70
  82. package/docs/PUBLISHING.md +91 -91
  83. package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
  84. package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
  85. package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
  86. package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
  87. package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
  88. package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
  89. package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
  90. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
  91. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
  92. package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
  93. package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
  94. package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
  95. package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
  96. package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
  97. package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
  98. package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
  99. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
  100. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
  101. package/gemini-extension.json +6 -6
  102. package/hooks/hooks.json +19 -19
  103. package/package.json +87 -87
package/CHANGELOG.md CHANGED
@@ -1,288 +1,294 @@
1
- # Changelog
2
-
3
- ## Unreleased
4
-
5
- ## 1.6.0 — 2026-08-20
6
-
7
- ### Added
8
- - The canon now includes `no-ai-slop`, `domain-modeling`, `codebase-design`, `resolving-merge-conflicts`, and `writing-for-agents`, with their supporting templates and evaluation material. Source adaptations are credited in `canon/skills/ATTRIBUTION.md`.
9
- - UI projects receive an automatic design-quality gate with configurable `design.mode` (`off`, `auto`, or `on`) and score budget. Detection uses package dependencies, UI source files, and configured smoke flows.
10
- - Durable context now includes a project glossary and can expose an optional bounded-context map.
11
-
12
- ### Changed
13
- - Retrofit installs complete skill packages for Claude, Codex, and Gemini instead of copying only `SKILL.md`. Local resources, binary files, and executable bits are preserved where supported.
14
- - Every canon skill declares `invocation: auto|manual`; retrofit maps that policy to Claude frontmatter, Codex `agents/openai.yaml`, and Gemini's generated automatic-skill index.
15
- - Canon validation rejects missing local Markdown resources, unsafe package entries, symlinks, and provider metadata that conflicts with the manifest.
16
-
17
- ### Fixed
18
- - Windows provider cleanup now treats a successful `taskkill` as a request and confirms that the recorded PID has stopped before deleting its ownership record.
19
-
20
- ## 1.5.1 — 2026-08-17
21
-
22
- ### Fixed
23
- - Serena MCP configurations generated by Yoke no longer open the local web dashboard automatically on every startup.
24
-
25
- ## 1.5.0 — 2026-08-16
26
-
27
- ### Added
28
- - Failed verify, executable-criterion, performance, configured custom-audit, and completion gates now produce deterministic byte-bounded previews and preserve large complete stdout/stderr in content-addressed `.yoke/artifacts/` files with SHA-256 references.
29
- - Projects can tune `output.previewBytes` and `output.artifactThresholdBytes`; existing projects use backward-compatible 2 KiB/8 KiB defaults.
30
- - A deterministic local benchmark verifies signal retention, preview bounds, compression measurement, and artifact digest round-trips without making provider-token claims.
31
-
32
- ### Security
33
- - Output artifact paths sanitize story identifiers, stay below a project-local root, use user-only file modes where supported, and are excluded from Yoke's clean-tree and story-commit operations even in upgraded projects. Raw artifacts are never injected automatically and documentation warns that project commands may emit secrets or personal data.
34
- - Gate command capture is capped at 16 MiB per stdout/stderr stream. Quota overflow fails closed and labels retained evidence as truncated instead of risking unbounded memory or claiming partial output is complete.
35
-
36
- ## 1.4.0 — 2026-08-15
37
-
38
- ### Added
39
- - `yoke loop run --parallel=N` now executes dependency-ready, non-colliding stories through real provider subprocess workers, isolated worktrees, leased claims, and a FIFO integration queue with fresh integrated-system gates.
40
- - Reference-driven quality declarations can collect screenshots, files, command output, or benchmark results and run a schema-validated blind critic with bounded repair rounds, elapsed-time limits, blocking or advisory policy, and retained proof.
41
- - `--candidates=N` can fan out up to five isolated implementations, discard mechanically red candidates, select one green candidate through identity-blind pairwise comparison, and preserve selected/loser evidence before cleanup.
42
- - Loop status now exposes dispatcher, worker, integrator, candidate lifecycle, worktree, queue, integration, reopen, quality-round, repair-budget, and trusted provider/model provenance data.
43
-
44
- ### Changed
45
- - Provider subprocesses use explicit lifecycle contracts and incarnation-aware process records so worker cancellation and cleanup target only the process tree Yoke actually started.
46
- - Parallel and candidate runs disable adaptive routing, honor story-level provider affinity, latch pause requests across the whole dispatcher, and rerun quality plus review after integration.
47
- - `yoke loop cleanup` retains Yoke worktrees unless `--remove-worktrees` is explicit, while still reaping recorded orphan runners and stale locks safely.
48
-
49
- ### Fixed
50
- - Expired claims, worker crashes, merge conflicts, pause races, and integration failures now release ownership deterministically, retain terminal proof, and reopen stories without leaking worktrees or marking false completion.
51
- - Quality repair fails closed on malformed critic output, reference drift, provider/model provenance mismatch, candidate identity leakage, unavailable critics, exhausted limits, and mechanically red repairs.
52
- - The watchdog resolves its TypeScript loader from both source and built npm layouts on Node 20+, and read-only Codex comparisons can run in disposable candidate worktrees without weakening normal repository checks.
53
-
54
- ### Security
55
- - Blind comparison requests expose only opaque labels and digests while binding every verdict to the trusted judge provider, model, prompt, rubric, reference, and candidate provenance.
56
- - Cleanup and cancellation use project-scoped leases, owner tokens, PID birth/incarnation checks, and recorded process handles rather than machine-wide process-name matching.
57
-
58
- ## 1.3.0 — 2026-08-09
59
-
60
- ### Added
61
- - Teams can submit change requests at any time with `yoke change add` and inspect the append-only inbox with `yoke change status`; queued requests are planned at safe loop boundaries by Claude, Codex, or Gemini without interrupting active story work.
62
- - Acceptance criteria can carry stable IDs and executable verification commands, and Yoke records per-criterion evidence before a story can be marked complete.
63
- - Projects can configure an integrated completion command that proves the final system flow after all story-level gates pass.
64
-
65
- ### Changed
66
- - New projects use strict criterion-proof requirements by default, while existing boolean-only criteria remain readable for compatibility.
67
- - PRD authoring, schemas, loop guidance, generated configuration, and provider instructions now treat completion as an ephemeral readiness state: new requests create more stories instead of prematurely forcing a release boundary.
68
-
69
- ### Fixed
70
- - Broad green test suites can no longer satisfy unrelated structured acceptance criteria: proof commands must use an approved test runner, contain the criterion ID, and avoid shell operators.
71
- - Change-intake failures block safely; an independent coverage pass rejects omitted requested outcomes, crash recovery retains uncommitted requests, and concurrent PRD edits are detected instead of silently overwriting user work.
72
- - Integrated completion checks prevent locally finished stories from masking broken cross-component flows such as authentication callbacks or purchase-to-entitlement activation.
73
-
74
- ### Security
75
- - Updated the transitive development dependency `nanoid` to a non-vulnerable release so the complete CI dependency audit is clean.
76
- - Restricted model-authored criterion proof to criterion-targeted test commands and blocked shell control operators and runner-prefix spoofing before host execution.
77
-
78
- ## 1.2.1 — 2026-08-02
79
-
80
- ### Fixed
81
- - `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
82
- - Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
83
-
84
- ## 1.2.0 — 2026-08-02
85
-
86
- ### Added
87
- - Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
88
- - A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
89
- - Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
90
-
91
- ### Changed
92
- - Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
93
- - Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
94
-
95
- ### Fixed
96
- - Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
97
- - Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
98
-
99
- ## 1.1.0 — 2026-07-30
100
-
101
- ### Added
102
- - Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
103
- - Persisted default runner selection and `auto|critical` loop decision policies.
104
- - Provider-neutral `yoke-workflow` skill for planning questions, approved-plan PRD handoff, autonomous story execution, and critical-decision resume.
105
- - Structured critical-decision requests with `yoke loop decision` and `yoke loop answer`; answers are validated, committed under the configured human identity, and resume the same story.
106
- - Approved `.yoke/plan.md` context in PRD drafting and a lint gate for unresolved planning placeholders.
107
-
108
- ### Fixed
109
- - Retrofit and loop on/off now preserve timeout, decision, runner, and permission settings.
110
- - Empty projects prefer the active agent host instead of silently installing Claude artifacts in Codex.
111
- - Loop and PRD runner selection now prefer an explicit flag, then the configured runner, then the active host.
112
- - Retrofit reports no longer label every provider as Claude Code.
113
- - Critical-decision resumes retain isolation, review, runner, permissions, timeout, JSON, policy, and iteration settings instead of falling back to an unreviewed default run.
114
- - Decision answers use an atomic owner-token lock and recoverable request journal, are checked against the active PRD story, bounded as untrusted data, and committed path-by-path so unrelated edits cannot enter the human-owned commit.
115
- - Decision recovery now binds the exact selected answer to its commit, rolls back only its own interrupted context append, namespaces resume state per project/worktree, and serializes cleanup with loop startup.
116
- - Active agent session markers now outrank globally configured provider home directories, and setup rejects partially invalid agent lists.
117
-
118
- ### Changed
119
- - New setups enable the loop by default and choose `decisionPolicy: auto`; the wizard can select `critical` or disable the loop.
120
- - Legacy `loop.onAmbiguity` and `--on-ambiguity` remain compatibility aliases.
121
-
122
- ## 1.0.0 — 2026-07-27
123
-
124
- ### Added
125
- - Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
126
- - Safe provider permission profiles and structured cross-provider telemetry.
127
- - Schema-validated independent review verdicts with explicit self-review opt-in.
128
- - Human-owned commit identity enforcement; AI co-author trailers default off.
129
- - `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
130
- - PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
131
- - Reproducible cross-runner benchmark schema and matrix launcher.
132
-
133
- ### Changed
134
- - Dangerous permission bypass is opt-in via `--unsafe`.
135
- - Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
136
- - Reviews no longer trust process exit code alone.
137
-
138
- ### Security
139
- - Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
140
-
141
-
142
- ## 0.9.0 — 2026-07-22
143
-
144
- ### Added
145
- - **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
146
- A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
147
- verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
148
- `--isolate` mode. A red benchmark blocks the story
149
- (`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
150
- gate is configured, the implementer prompt names the budget command so agents keep hot
151
- paths efficient and never "simplify away" an optimization without re-running the benchmark.
152
- - **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
153
- clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
154
- project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
155
- tests, the *why* of every optimization versioned so future agents don't clean fast code
156
- back to slow.
157
- - **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
158
- numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
159
- round — a criterion still needing a decision is not loop-ready.
160
-
161
- ## 0.8.0 — 2026-07-20
162
-
163
- ### Added
164
- - **Live progress + ETA.** Story completions are now first-class events: the console shows
165
- `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
166
- `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
167
- durations of stories completed **in this run** (current velocity) and falls back to the
168
- persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
169
- No data → no estimate, never an invented one.
170
- - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
171
- prompt now always forbids asking questions (a loop run has nobody to answer). Default
172
- `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
173
- loop never stops. Opt-in `abort`: the agent writes its open questions to
174
- `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
175
- story would otherwise pass on pre-existing green tests), and blocks with the question as
176
- the reason. Companion principle: clarifying questions belong in the planning round, before
177
- the loop starts.
178
-
179
- ## 0.7.0 — 2026-07-17
180
-
181
- ### Added
182
- - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
183
- version hint (npm/gh-style): a detached background refresher caches the registry's
184
- latest at most once a day; when it is newer, a one-line stderr hint suggests
185
- `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
186
- `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
187
- - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
188
- loop START only — never mid-run; the running process finishes on its version and
189
- the upgrade applies from the next invocation. Deliberately NOT the default:
190
- a gate harness must not change itself mid-project (determinism), and unreviewed
191
- auto-installs are a supply-chain hazard.
192
-
193
- ## 0.6.0 — 2026-07-17
194
-
195
- ### Added
196
- - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
197
- `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
198
- `yoke loop cleanup` kills exactly those recorded process trees — and only while no
199
- live loop holds the lock. Background: without a scoped mechanism, users and agents
200
- resorted to machine-wide pattern kills (every process matching
201
- `dangerously-skip-permissions`), which took down *healthy* runners of other projects
202
- mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
203
- safe path. `.yoke/runner.pid` is gitignored by retrofit.
204
-
205
- ## 0.5.0 — 2026-07-17
206
-
207
- Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
208
- (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
209
- idle window):
210
-
211
- ### Fixed
212
- - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
213
- Previously it killed only the spawned shell (`shell: true`), orphaning the actual
214
- agent process — which kept writing to the worktree (dirty-tree blocks, failing
215
- worktree removal) and kept burning API tokens. Observed in the field as ~10
216
- zombie `claude.exe` per machine plus surviving dev servers.
217
- - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
218
- the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
219
- processes and killed them at exactly the idle timeout — while the user saw dead air.
220
- The stream doubles as liveness; token usage is now reported on every run (not just
221
- `--json` mode).
222
-
223
- ### Changed
224
- - README: operating notes for driving the loop from inside an agent session
225
- (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
226
- outer shell-tool timeouts killing a foreground `yoke loop run` were the third
227
- observed "hang" pattern.
228
-
229
- ## 0.4.0 — 2026-07-17
230
-
231
- ### Added
232
- - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
233
- scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
234
- documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
235
- bounded final messages (cuts output-token waste). Review prompts now ground verdicts
236
- in observed evidence only and keep them brief.
237
-
238
- ### Fixed
239
- - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
240
- `git add -A` story commit swept the pause control file into history in
241
- un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
242
- blocked the resume run — the loop locked itself out.
243
-
244
- > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
245
- > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
246
-
247
- ## 0.3.0 — 2026-07-10
248
-
249
- ### Added
250
- - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
251
- (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
252
- then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
253
- - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
254
- installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
255
- daily-crawled extensions gallery.
256
- - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
257
- quality) with a fixed fixture, pre-written objective tests, and committed result data.
258
- - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
259
- sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
260
- (design generation paired with Yoke's design verification gates).
261
- - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
262
- stories across parallel workers (`needs` dependency field, claim files, merge queue,
263
- heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
264
-
265
- ### Fixed
266
- - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
267
- Windows, so the loop misreported an installed `gemini` as "not found on PATH"
268
- (found by the new benchmark harness).
269
- - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
270
- requires a value after `-p` and errored with "Not enough arguments following: p".
271
- Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
272
- (also found by the benchmark harness).
273
-
274
- ### Changed
275
- - README: npm install is now the primary quickstart path; documented plugin/extension
276
- installs and optional companions.
277
- - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
278
- `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
279
-
280
- ## 0.2.0 — 2026-07-09
281
-
282
- - First npm release as `@hecer/yoke`.
283
- - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
284
- token-usage + model-id reporting for the claude runner.
285
- - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
286
- `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
287
- - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
288
- loop with worktree isolation, watchdog, single-flight lock, commit integrity.
1
+ # Changelog
2
+
3
+ ## Unreleased
4
+
5
+ ## 1.6.1 — 2026-08-21
6
+
7
+ ### Fixed
8
+ - Release metadata now counts platform-conditional tests consistently on Windows and Ubuntu, so `docs:check` no longer fails after an otherwise green cross-platform test matrix.
9
+ - Provider cleanup now rechecks termination after the child closes, preventing stale ownership records when Windows reports process exit asynchronously.
10
+
11
+ ## 1.6.0 — 2026-08-20
12
+
13
+ ### Added
14
+ - The canon now includes `no-ai-slop`, `domain-modeling`, `codebase-design`, `resolving-merge-conflicts`, and `writing-for-agents`, with their supporting templates and evaluation material. Source adaptations are credited in `canon/skills/ATTRIBUTION.md`.
15
+ - UI projects receive an automatic design-quality gate with configurable `design.mode` (`off`, `auto`, or `on`) and score budget. Detection uses package dependencies, UI source files, and configured smoke flows.
16
+ - Durable context now includes a project glossary and can expose an optional bounded-context map.
17
+
18
+ ### Changed
19
+ - Retrofit installs complete skill packages for Claude, Codex, and Gemini instead of copying only `SKILL.md`. Local resources, binary files, and executable bits are preserved where supported.
20
+ - Every canon skill declares `invocation: auto|manual`; retrofit maps that policy to Claude frontmatter, Codex `agents/openai.yaml`, and Gemini's generated automatic-skill index.
21
+ - Canon validation rejects missing local Markdown resources, unsafe package entries, symlinks, and provider metadata that conflicts with the manifest.
22
+
23
+ ### Fixed
24
+ - Windows provider cleanup now treats a successful `taskkill` as a request and confirms that the recorded PID has stopped before deleting its ownership record.
25
+
26
+ ## 1.5.1 — 2026-08-17
27
+
28
+ ### Fixed
29
+ - Serena MCP configurations generated by Yoke no longer open the local web dashboard automatically on every startup.
30
+
31
+ ## 1.5.0 — 2026-08-16
32
+
33
+ ### Added
34
+ - Failed verify, executable-criterion, performance, configured custom-audit, and completion gates now produce deterministic byte-bounded previews and preserve large complete stdout/stderr in content-addressed `.yoke/artifacts/` files with SHA-256 references.
35
+ - Projects can tune `output.previewBytes` and `output.artifactThresholdBytes`; existing projects use backward-compatible 2 KiB/8 KiB defaults.
36
+ - A deterministic local benchmark verifies signal retention, preview bounds, compression measurement, and artifact digest round-trips without making provider-token claims.
37
+
38
+ ### Security
39
+ - Output artifact paths sanitize story identifiers, stay below a project-local root, use user-only file modes where supported, and are excluded from Yoke's clean-tree and story-commit operations even in upgraded projects. Raw artifacts are never injected automatically and documentation warns that project commands may emit secrets or personal data.
40
+ - Gate command capture is capped at 16 MiB per stdout/stderr stream. Quota overflow fails closed and labels retained evidence as truncated instead of risking unbounded memory or claiming partial output is complete.
41
+
42
+ ## 1.4.0 — 2026-08-15
43
+
44
+ ### Added
45
+ - `yoke loop run --parallel=N` now executes dependency-ready, non-colliding stories through real provider subprocess workers, isolated worktrees, leased claims, and a FIFO integration queue with fresh integrated-system gates.
46
+ - Reference-driven quality declarations can collect screenshots, files, command output, or benchmark results and run a schema-validated blind critic with bounded repair rounds, elapsed-time limits, blocking or advisory policy, and retained proof.
47
+ - `--candidates=N` can fan out up to five isolated implementations, discard mechanically red candidates, select one green candidate through identity-blind pairwise comparison, and preserve selected/loser evidence before cleanup.
48
+ - Loop status now exposes dispatcher, worker, integrator, candidate lifecycle, worktree, queue, integration, reopen, quality-round, repair-budget, and trusted provider/model provenance data.
49
+
50
+ ### Changed
51
+ - Provider subprocesses use explicit lifecycle contracts and incarnation-aware process records so worker cancellation and cleanup target only the process tree Yoke actually started.
52
+ - Parallel and candidate runs disable adaptive routing, honor story-level provider affinity, latch pause requests across the whole dispatcher, and rerun quality plus review after integration.
53
+ - `yoke loop cleanup` retains Yoke worktrees unless `--remove-worktrees` is explicit, while still reaping recorded orphan runners and stale locks safely.
54
+
55
+ ### Fixed
56
+ - Expired claims, worker crashes, merge conflicts, pause races, and integration failures now release ownership deterministically, retain terminal proof, and reopen stories without leaking worktrees or marking false completion.
57
+ - Quality repair fails closed on malformed critic output, reference drift, provider/model provenance mismatch, candidate identity leakage, unavailable critics, exhausted limits, and mechanically red repairs.
58
+ - The watchdog resolves its TypeScript loader from both source and built npm layouts on Node 20+, and read-only Codex comparisons can run in disposable candidate worktrees without weakening normal repository checks.
59
+
60
+ ### Security
61
+ - Blind comparison requests expose only opaque labels and digests while binding every verdict to the trusted judge provider, model, prompt, rubric, reference, and candidate provenance.
62
+ - Cleanup and cancellation use project-scoped leases, owner tokens, PID birth/incarnation checks, and recorded process handles rather than machine-wide process-name matching.
63
+
64
+ ## 1.3.0 — 2026-08-09
65
+
66
+ ### Added
67
+ - Teams can submit change requests at any time with `yoke change add` and inspect the append-only inbox with `yoke change status`; queued requests are planned at safe loop boundaries by Claude, Codex, or Gemini without interrupting active story work.
68
+ - Acceptance criteria can carry stable IDs and executable verification commands, and Yoke records per-criterion evidence before a story can be marked complete.
69
+ - Projects can configure an integrated completion command that proves the final system flow after all story-level gates pass.
70
+
71
+ ### Changed
72
+ - New projects use strict criterion-proof requirements by default, while existing boolean-only criteria remain readable for compatibility.
73
+ - PRD authoring, schemas, loop guidance, generated configuration, and provider instructions now treat completion as an ephemeral readiness state: new requests create more stories instead of prematurely forcing a release boundary.
74
+
75
+ ### Fixed
76
+ - Broad green test suites can no longer satisfy unrelated structured acceptance criteria: proof commands must use an approved test runner, contain the criterion ID, and avoid shell operators.
77
+ - Change-intake failures block safely; an independent coverage pass rejects omitted requested outcomes, crash recovery retains uncommitted requests, and concurrent PRD edits are detected instead of silently overwriting user work.
78
+ - Integrated completion checks prevent locally finished stories from masking broken cross-component flows such as authentication callbacks or purchase-to-entitlement activation.
79
+
80
+ ### Security
81
+ - Updated the transitive development dependency `nanoid` to a non-vulnerable release so the complete CI dependency audit is clean.
82
+ - Restricted model-authored criterion proof to criterion-targeted test commands and blocked shell control operators and runner-prefix spoofing before host execution.
83
+
84
+ ## 1.2.1 — 2026-08-02
85
+
86
+ ### Fixed
87
+ - `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
88
+ - Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
89
+
90
+ ## 1.2.0 — 2026-08-02
91
+
92
+ ### Added
93
+ - Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
94
+ - A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
95
+ - Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
96
+
97
+ ### Changed
98
+ - Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
99
+ - Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
100
+
101
+ ### Fixed
102
+ - Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
103
+ - Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
104
+
105
+ ## 1.1.0 — 2026-07-30
106
+
107
+ ### Added
108
+ - Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
109
+ - Persisted default runner selection and `auto|critical` loop decision policies.
110
+ - Provider-neutral `yoke-workflow` skill for planning questions, approved-plan PRD handoff, autonomous story execution, and critical-decision resume.
111
+ - Structured critical-decision requests with `yoke loop decision` and `yoke loop answer`; answers are validated, committed under the configured human identity, and resume the same story.
112
+ - Approved `.yoke/plan.md` context in PRD drafting and a lint gate for unresolved planning placeholders.
113
+
114
+ ### Fixed
115
+ - Retrofit and loop on/off now preserve timeout, decision, runner, and permission settings.
116
+ - Empty projects prefer the active agent host instead of silently installing Claude artifacts in Codex.
117
+ - Loop and PRD runner selection now prefer an explicit flag, then the configured runner, then the active host.
118
+ - Retrofit reports no longer label every provider as Claude Code.
119
+ - Critical-decision resumes retain isolation, review, runner, permissions, timeout, JSON, policy, and iteration settings instead of falling back to an unreviewed default run.
120
+ - Decision answers use an atomic owner-token lock and recoverable request journal, are checked against the active PRD story, bounded as untrusted data, and committed path-by-path so unrelated edits cannot enter the human-owned commit.
121
+ - Decision recovery now binds the exact selected answer to its commit, rolls back only its own interrupted context append, namespaces resume state per project/worktree, and serializes cleanup with loop startup.
122
+ - Active agent session markers now outrank globally configured provider home directories, and setup rejects partially invalid agent lists.
123
+
124
+ ### Changed
125
+ - New setups enable the loop by default and choose `decisionPolicy: auto`; the wizard can select `critical` or disable the loop.
126
+ - Legacy `loop.onAmbiguity` and `--on-ambiguity` remain compatibility aliases.
127
+
128
+ ## 1.0.0 — 2026-07-27
129
+
130
+ ### Added
131
+ - Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
132
+ - Safe provider permission profiles and structured cross-provider telemetry.
133
+ - Schema-validated independent review verdicts with explicit self-review opt-in.
134
+ - Human-owned commit identity enforcement; AI co-author trailers default off.
135
+ - `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
136
+ - PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
137
+ - Reproducible cross-runner benchmark schema and matrix launcher.
138
+
139
+ ### Changed
140
+ - Dangerous permission bypass is opt-in via `--unsafe`.
141
+ - Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
142
+ - Reviews no longer trust process exit code alone.
143
+
144
+ ### Security
145
+ - Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
146
+
147
+
148
+ ## 0.9.0 — 2026-07-22
149
+
150
+ ### Added
151
+ - **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
152
+ A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
153
+ verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
154
+ `--isolate` mode. A red benchmark blocks the story
155
+ (`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
156
+ gate is configured, the implementer prompt names the budget command so agents keep hot
157
+ paths efficient and never "simplify away" an optimization without re-running the benchmark.
158
+ - **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
159
+ clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
160
+ project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
161
+ tests, the *why* of every optimization versioned so future agents don't clean fast code
162
+ back to slow.
163
+ - **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
164
+ numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
165
+ round — a criterion still needing a decision is not loop-ready.
166
+
167
+ ## 0.8.0 — 2026-07-20
168
+
169
+ ### Added
170
+ - **Live progress + ETA.** Story completions are now first-class events: the console shows
171
+ `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
172
+ `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
173
+ durations of stories completed **in this run** (current velocity) and falls back to the
174
+ persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
175
+ No data → no estimate, never an invented one.
176
+ - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
177
+ prompt now always forbids asking questions (a loop run has nobody to answer). Default
178
+ `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
179
+ loop never stops. Opt-in `abort`: the agent writes its open questions to
180
+ `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
181
+ story would otherwise pass on pre-existing green tests), and blocks with the question as
182
+ the reason. Companion principle: clarifying questions belong in the planning round, before
183
+ the loop starts.
184
+
185
+ ## 0.7.0 — 2026-07-17
186
+
187
+ ### Added
188
+ - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
189
+ version hint (npm/gh-style): a detached background refresher caches the registry's
190
+ latest at most once a day; when it is newer, a one-line stderr hint suggests
191
+ `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
192
+ `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
193
+ - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
194
+ loop START only — never mid-run; the running process finishes on its version and
195
+ the upgrade applies from the next invocation. Deliberately NOT the default:
196
+ a gate harness must not change itself mid-project (determinism), and unreviewed
197
+ auto-installs are a supply-chain hazard.
198
+
199
+ ## 0.6.0 — 2026-07-17
200
+
201
+ ### Added
202
+ - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
203
+ `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
204
+ `yoke loop cleanup` kills exactly those recorded process trees — and only while no
205
+ live loop holds the lock. Background: without a scoped mechanism, users and agents
206
+ resorted to machine-wide pattern kills (every process matching
207
+ `dangerously-skip-permissions`), which took down *healthy* runners of other projects
208
+ mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
209
+ safe path. `.yoke/runner.pid` is gitignored by retrofit.
210
+
211
+ ## 0.5.0 — 2026-07-17
212
+
213
+ Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
214
+ (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
215
+ idle window):
216
+
217
+ ### Fixed
218
+ - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
219
+ Previously it killed only the spawned shell (`shell: true`), orphaning the actual
220
+ agent process — which kept writing to the worktree (dirty-tree blocks, failing
221
+ worktree removal) and kept burning API tokens. Observed in the field as ~10
222
+ zombie `claude.exe` per machine plus surviving dev servers.
223
+ - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
224
+ the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
225
+ processes and killed them at exactly the idle timeout — while the user saw dead air.
226
+ The stream doubles as liveness; token usage is now reported on every run (not just
227
+ `--json` mode).
228
+
229
+ ### Changed
230
+ - README: operating notes for driving the loop from inside an agent session
231
+ (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
232
+ outer shell-tool timeouts killing a foreground `yoke loop run` were the third
233
+ observed "hang" pattern.
234
+
235
+ ## 0.4.0 — 2026-07-17
236
+
237
+ ### Added
238
+ - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
239
+ scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
240
+ documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
241
+ bounded final messages (cuts output-token waste). Review prompts now ground verdicts
242
+ in observed evidence only and keep them brief.
243
+
244
+ ### Fixed
245
+ - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
246
+ `git add -A` story commit swept the pause control file into history in
247
+ un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
248
+ blocked the resume run — the loop locked itself out.
249
+
250
+ > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
251
+ > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
252
+
253
+ ## 0.3.0 — 2026-07-10
254
+
255
+ ### Added
256
+ - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
257
+ (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
258
+ then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
259
+ - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
260
+ installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
261
+ daily-crawled extensions gallery.
262
+ - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
263
+ quality) with a fixed fixture, pre-written objective tests, and committed result data.
264
+ - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
265
+ sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
266
+ (design generation paired with Yoke's design verification gates).
267
+ - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
268
+ stories across parallel workers (`needs` dependency field, claim files, merge queue,
269
+ heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
270
+
271
+ ### Fixed
272
+ - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
273
+ Windows, so the loop misreported an installed `gemini` as "not found on PATH"
274
+ (found by the new benchmark harness).
275
+ - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
276
+ requires a value after `-p` and errored with "Not enough arguments following: p".
277
+ Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
278
+ (also found by the benchmark harness).
279
+
280
+ ### Changed
281
+ - README: npm install is now the primary quickstart path; documented plugin/extension
282
+ installs and optional companions.
283
+ - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
284
+ `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
285
+
286
+ ## 0.2.0 — 2026-07-09
287
+
288
+ - First npm release as `@hecer/yoke`.
289
+ - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
290
+ token-usage + model-id reporting for the claude runner.
291
+ - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
292
+ `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
293
+ - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
294
+ loop with worktree isolation, watchdog, single-flight lock, commit integrity.