@hecer/yoke 1.11.0 → 1.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (128) hide show
  1. package/.claude-plugin/plugin.json +13 -13
  2. package/.codex-plugin/plugin.json +7 -7
  3. package/CHANGELOG.md +416 -398
  4. package/README.md +931 -915
  5. package/TODOS.md +5 -5
  6. package/agents/docs.toml +6 -6
  7. package/agents/implementer.toml +6 -6
  8. package/agents/reviewer.toml +6 -6
  9. package/agents/security.toml +6 -6
  10. package/bench/README.md +86 -86
  11. package/bench/RESULTS.md +35 -35
  12. package/bench/output-compaction.mjs +65 -65
  13. package/bench/result-schema.mjs +12 -12
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -15
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
  17. package/bench/run-matrix.mjs +26 -26
  18. package/bench/run.mjs +106 -106
  19. package/canon/AGENTS.md +30 -30
  20. package/canon/context/DECISIONS.md +4 -4
  21. package/canon/context/GLOSSARY.md +11 -11
  22. package/canon/context/KNOWLEDGE.md +4 -4
  23. package/canon/context/PROJECT.md +15 -15
  24. package/canon/loop/loop-spec.md +65 -65
  25. package/canon/loop/prd.schema.md +41 -41
  26. package/canon/manifest.yaml +59 -59
  27. package/canon/policy/gates.md +7 -7
  28. package/canon/policy/roles.md +9 -9
  29. package/canon/skills/ATTRIBUTION.md +99 -99
  30. package/canon/skills/authoring-prd/SKILL.md +56 -56
  31. package/canon/skills/brainstorming/SKILL.md +164 -164
  32. package/canon/skills/codebase-design/DEEPENING.md +15 -15
  33. package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
  34. package/canon/skills/codebase-design/SKILL.md +39 -39
  35. package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
  36. package/canon/skills/document-release/SKILL.md +302 -302
  37. package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
  38. package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
  39. package/canon/skills/domain-modeling/SKILL.md +35 -35
  40. package/canon/skills/executing-plans/SKILL.md +70 -70
  41. package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
  42. package/canon/skills/health/SKILL.md +177 -177
  43. package/canon/skills/maintaining-context/SKILL.md +34 -34
  44. package/canon/skills/minimal-code/SKILL.md +21 -21
  45. package/canon/skills/no-ai-slop/SKILL.md +103 -103
  46. package/canon/skills/no-ai-slop/eval.md +43 -43
  47. package/canon/skills/plan-ceo-review/SKILL.md +541 -541
  48. package/canon/skills/plan-eng-review/SKILL.md +362 -362
  49. package/canon/skills/receiving-code-review/SKILL.md +213 -213
  50. package/canon/skills/requesting-code-review/SKILL.md +105 -105
  51. package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
  52. package/canon/skills/retro/SKILL.md +397 -397
  53. package/canon/skills/review/SKILL.md +246 -246
  54. package/canon/skills/ship/SKILL.md +691 -691
  55. package/canon/skills/subagent-driven-development/SKILL.md +277 -277
  56. package/canon/skills/systematic-debugging/SKILL.md +296 -296
  57. package/canon/skills/tdd/SKILL.md +371 -371
  58. package/canon/skills/unslop-ui/SKILL.md +34 -34
  59. package/canon/skills/using-git-worktrees/SKILL.md +218 -218
  60. package/canon/skills/verification-before-completion/SKILL.md +139 -139
  61. package/canon/skills/visual-verification/SKILL.md +54 -54
  62. package/canon/skills/workflow/SKILL.md +22 -22
  63. package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
  64. package/canon/skills/writing-for-agents/SKILL.md +42 -42
  65. package/canon/skills/writing-plans/SKILL.md +152 -152
  66. package/canon/skills/writing-skills/SKILL.md +655 -655
  67. package/canon/skills/yoke-retrofit/SKILL.md +26 -26
  68. package/canon/skills/yoke-workflow/SKILL.md +20 -20
  69. package/canon/tools/codex-rtk-hook.mjs +35 -35
  70. package/canon/tools/gemini-rtk-hook.mjs +25 -25
  71. package/canon/tools/graphify.md +3 -3
  72. package/canon/tools/playwright-mcp.md +3 -3
  73. package/canon/tools/qwen-rtk-hook.mjs +25 -0
  74. package/canon/tools/rtk.md +7 -7
  75. package/canon/tools/serena.md +6 -6
  76. package/dist/agents/host.js +1 -1
  77. package/dist/agents/providers.js +18 -5
  78. package/dist/agents/telemetry.js +35 -36
  79. package/dist/cli.js +18 -10
  80. package/dist/dashboard/page.js +122 -122
  81. package/dist/dashboard/panels.js +91 -91
  82. package/dist/loop/run-command.js +3 -3
  83. package/dist/prd/command.js +17 -17
  84. package/dist/retrofit/apply.js +8 -1
  85. package/dist/retrofit/config.js +1 -1
  86. package/dist/retrofit/detect.js +2 -0
  87. package/dist/retrofit/planners/claude.js +14 -14
  88. package/dist/retrofit/planners/qwen.js +3 -3
  89. package/dist/retrofit/preserve.js +2 -2
  90. package/dist/retrofit/qwen-settings.js +17 -0
  91. package/dist/retrofit/skill-actions.js +1 -1
  92. package/dist/setup/command.js +22 -8
  93. package/dist/setup/model-presets.js +48 -0
  94. package/docs/CAPABILITY-ROUTING.md +51 -51
  95. package/docs/DASHBOARD-EVOLUTION.md +33 -33
  96. package/docs/MIGRATING-TO-1.0.md +33 -33
  97. package/docs/MIGRATING-TO-1.1.md +27 -27
  98. package/docs/MIGRATING-TO-1.4.md +70 -70
  99. package/docs/PRODUCT-DIRECTION-2026-09-05.md +210 -210
  100. package/docs/PUBLISHING.md +114 -114
  101. package/docs/QWEN-MODEL-SUPPORT.md +142 -0
  102. package/docs/VERIFIED-PROJECTS-VALIDATION.md +29 -29
  103. package/docs/VERIFIED-PROJECTS.md +167 -167
  104. package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
  105. package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
  106. package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
  107. package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
  108. package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
  109. package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
  110. package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
  111. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
  112. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
  113. package/docs/superpowers/plans/2026-09-05-verified-projects.md +83 -83
  114. package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
  115. package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
  116. package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
  117. package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
  118. package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
  119. package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
  120. package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
  121. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
  122. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
  123. package/gemini-extension.json +6 -6
  124. package/hooks/hooks.json +19 -19
  125. package/package.json +87 -87
  126. package/dist/dashboard/discovery.js +0 -73
  127. package/docs/community-outreach-2026-08-20.md +0 -85
  128. package/docs/launch-copy-2026-08-21.md +0 -193
package/CHANGELOG.md CHANGED
@@ -1,398 +1,416 @@
1
- # Changelog
2
-
3
- ## 1.11.0 — 2026-09-07
4
-
5
- ### Added
6
- - Add Qwen Code (Alibaba) as fourth supported provider alongside Claude, Codex and Gemini.
7
- - Detect Qwen host environment via `QWEN_CLI` and `QWEN_CLI_HOME` environment variables.
8
- - Add Qwen routing workers with four capability tiers: `qwen-turbo-latest` (light), `qwen3-coder-plus` (standard/strong), `qwen3-235b-a22b` (frontier).
9
- - Parse Qwen streaming telemetry for token usage and model reporting.
10
- - Support Qwen in all CLI commands: `setup`, `loop`, `review`, `prd draft`, `prd assess`, `goal run`.
11
-
12
- ### Changed
13
- - Update project description from "three agents" to "four agents" to reflect Qwen support.
14
- - Extend review resolution order to include Qwen for cross-model reviews.
15
- - Update setup prompts to offer Qwen as agent and runner option.
16
-
17
- ### Migration and validation limits
18
- - Existing configurations remain compatible. New setups can select Qwen as agent/runner.
19
- - Qwen CLI uses Gemini-style arguments (`--approval-mode`, `--output-format stream-json`). `bare`, `reasoningEffort` and `nativeMultiAgent` selections are not supported (like Gemini).
20
- - Qwen routing profiles are configurable hypotheses, not authenticated benchmarks. Update installed packages and restart the dashboard/runner.
21
-
22
- ## 1.10.0 — 2026-09-06
23
-
24
- ### Added
25
- - Adopt the Yoke wordmark in the GitHub and npm README.
26
- - Prepare task assessments in one bounded planning call with `yoke prd assess`; draft and inbox planning bind assessments to requirements, upstream dependencies and the approved brief.
27
- - Separate `planning.agent`/`model`/`reasoningEffort` from execution settings. New setups require prepared assessments and block missing worker profiles; existing configurations retain on-demand planning and parent fallback.
28
- - Add selective reassessment, planning input limits and `routing.maxTier` to bound automatic model selection and escalation.
29
- - Add attention-first project search and status filters, with separate goal/loop states and explicit stale activity warnings.
30
- - Restore dashboard views and UTC period controls through URL links, browser history and refresh; cancel obsolete requests and bound project comparison concurrency.
31
- - Compare recorded tokens, accepted tasks and reported costs against the preceding equal-duration period without inventing missing measurements or percentages from zero baselines.
32
-
33
- ### Fixed
34
- - Fix Windows safe-mode Codex execution by screening Store PowerShell/aliases and probing a native shell under the same sandbox before model work; keep permissions intact.
35
- - Launch Windows provider executables/npm entry points with literal argv; bound preflight and provider lifetime, recognize streamed infrastructure failures, retain failed-worktree/process evidence, and guard against unconfirmed termination before another worker starts.
36
- - Separate provider liveness, supervisor heartbeat, output and successful-tool progress in status/dashboard; label backlog percentages and preserve explicit bare startup through capability routing. See [Windows runner validation](docs/WINDOWS-RUNNER-VALIDATION.md).
37
- - Keep local loop locks and dashboard status out of implementation commits and clean-worktree checks.
38
- - Stage implementation files safely when runtime history directories are already ignored, including literal filenames and tracked deletions, for serial commits and parallel candidate snapshots.
39
-
40
- ### Migration and validation limits
41
- - Existing routing configurations retain on-demand assessment and parent fallback. Use `yoke prd assess` to prepare task packages; configure `planning.agent`, `planning.model`, `planning.reasoningEffort` and `routing.maxTier` to separate planning from bounded execution. New setups require prepared assessments and block missing profiles.
42
- - Update installed packages and restart the dashboard/runner to use the new code. Existing running processes are not upgraded or instrumented retroactively. Windows preflight preserves sandbox permissions and changes only the provider environment.
43
- - Dashboard comparisons report recorded measurements, not reconstructed history or calibrated savings. Tokens per minute are consumption rates, not generation speed. Live routed Windows validation covered Codex on the documented machine; cross-provider performance and universal Windows compatibility are not established. See the linked dashboard, batch-planning and runner validation records.
44
-
45
- ## 1.9.0 — 2026-09-06
46
-
47
- ### Added
48
- - Add capability-based routing with persisted task assessments, explicit model/effort tiers, role eligibility and conservative use of independent task-class outcomes.
49
- - Keep planning on the start model; reuse assessments across attempts/worktrees and invalidate them when task requirements change.
50
- - Add bounded repair and tier escalation after mechanical gate failures, retaining the patch and forwarding failure evidence. Infrastructure failures do not trigger capability escalation.
51
- - Apply task-based profiles to reviews, quality critics/repairs and goal execution; display implementation selection reasons and next escalation in the dashboard.
52
- - Add setup options `--routing-strategy=capability` and `--routing-preset` for explicit migration. Preserve existing strategies and custom profiles by default.
53
-
54
- ### Validation limits
55
- - Initial profile tiers are configurable hypotheses, not authenticated model benchmarks, price estimates or calibrated success probabilities. See [capability routing](docs/CAPABILITY-ROUTING.md) for defaults and bounds.
56
-
57
- ## 1.8.0 — 2026-09-06
58
-
59
- ### Added
60
- - Add persistent local measurement history and dashboard views for current work, usage/time and results, including UTC day/week/month filters, model history, project comparisons and consumption charts.
61
- - Display current tasks, worker phases, integration progress and status age, with automatic refresh of the current-work view.
62
- - Record explicit acceptances and show measured tokens and time per acceptance. Attribute available reviewer, critic and repair usage; distinguish actual reported models, unknown calls and partial costs.
63
-
64
- ### Changed
65
- - Enable routing in new setups and automatically select up to three parallel workers when all pending tasks declare write scopes. Dependencies and overlapping scopes still constrain dispatch. Preserve explicit opt-outs and use isolated worktrees by default.
66
- - Share routing decisions between synchronous and asynchronous runners; support routed parallel workers, stable recovery history and explicit provider affinity.
67
- - Reserve execution capacity through integration and disable native delegation for loop providers; preserve Gemini system policy in a temporary bounded-execution configuration.
68
- - Require a dated changelog entry, synchronized version metadata and verified release checks for every new version in the project instructions.
69
-
70
- ### Fixed
71
- - Avoid conflicting Codex sandbox arguments and prevent Codex-only options from leaking into Gemini workers.
72
- - Keep compact measurement history after recent activity expires, deduplicate archived events, and report incomplete history instead of treating missing usage as zero.
73
- - Preserve reviewer telemetry and worker/model attribution across parallel execution and recovery.
74
-
75
- ### Migration and validation limits
76
- - Use `--parallel=N` to choose a worker limit, `--parallel=auto` for automatic selection, `--no-routing` to opt out of routing, and `--no-isolate` to opt out of default isolation. Explicit existing configuration remains authoritative. Unknown write scopes, tool actions and worktree recovery select serial execution in auto mode.
77
- - Automatic routing needs configured profiles; otherwise it keeps the selected parent provider. Explicit `--routing` without profiles reports a configuration error.
78
- - Missing historical usage cannot be reconstructed. Tokens per minute describe interval or summed call consumption, not measured generation speed. Live authenticated provider benchmarks, resource-adaptive concurrency and calibrated time/cost predictions are not established by this release.
79
-
80
- ## 1.7.0 — 2026-09-05
81
-
82
- ### Added
83
- - Independent `yoke check` with executable acceptance mapping, protected test infrastructure and content-bound evidence.
84
- - Durable project goals, provider handoff, checkpoint budgets, interruption accounting and project-scoped recovery.
85
- - Local project registry and loopback dashboard with goals, task estimates, evidence, consumption and pause controls.
86
- - Explicit routing rules with persisted gate-driven escalation; bounded tool actions without model calls.
87
- - Task-aware context packets, advisory write scopes, dependency-depth scheduling and empirical time ranges with prediction error records.
88
-
89
- ### Fixed
90
- - Failed serial isolated worktrees are retained and can be explicitly resumed against their original target and PRD.
91
- - Reviewer fingerprints include untracked contents and acceptance inputs; unsupported nested repository identity fails closed.
92
- - Gemini always emits streaming telemetry, preserves model identity, honors aggregate token aliases and rejects unsupported selections. Its RTK hook merges native settings without a shell dependency.
93
- - Runtime evidence is excluded from Git gates and commits in existing projects. Partial usage and costs remain visibly incomplete.
94
- - Protected acceptance is checked after verification and repair, and linked goal state cannot overwrite unrelated files through pause.
95
-
96
- ### Validation limits
97
- - Live authenticated provider comparisons and calibrated development-time/cost estimates are not established by the automated tests. Browser proof and independent model review still require explicit configuration.
98
-
99
- ## 1.6.2 — 2026-09-02
100
-
101
- ### Added
102
- - GitHub Releases can now publish `@hecer/yoke` through npm trusted publishing with short-lived OIDC credentials and automatic provenance, without a long-lived npm token.
103
-
104
- ### Fixed
105
- - Codex safe-mode invocations now use the supported `workspace-write` sandbox with `--approve-for-me`; the removed `--full-auto` flag no longer blocks current Codex CLI releases.
106
- - Windows provider cleanup rechecks termination after process close and accepts an already-absent process as successfully cleaned up, avoiding stale ownership records and unnecessary watchdog waits.
107
- - Successful stale-loop cleanup removes obsolete runtime status, so `yoke loop status` no longer reports a dead run as `RUNNING`.
108
-
109
- ## 1.6.1 — 2026-08-21
110
-
111
- ### Fixed
112
- - Release metadata counts platform-conditional tests consistently on Windows and Ubuntu.
113
- - Provider cleanup confirms termination after child close, preventing stale ownership records when Windows reports process exit asynchronously.
114
-
115
- ## 1.6.0 — 2026-08-20
116
-
117
- ### Added
118
- - The canon now includes `no-ai-slop`, `domain-modeling`, `codebase-design`, `resolving-merge-conflicts`, and `writing-for-agents`, with their supporting templates and evaluation material. Source adaptations are credited in `canon/skills/ATTRIBUTION.md`.
119
- - UI projects receive an automatic design-quality gate with configurable `design.mode` (`off`, `auto`, or `on`) and score budget. Detection uses package dependencies, UI source files, and configured smoke flows.
120
- - Durable context now includes a project glossary and can expose an optional bounded-context map.
121
-
122
- ### Changed
123
- - Retrofit installs complete skill packages for Claude, Codex, and Gemini instead of copying only `SKILL.md`. Local resources, binary files, and executable bits are preserved where supported.
124
- - Every canon skill declares `invocation: auto|manual`; retrofit maps that policy to Claude frontmatter, Codex `agents/openai.yaml`, and Gemini's generated automatic-skill index.
125
- - Canon validation rejects missing local Markdown resources, unsafe package entries, symlinks, and provider metadata that conflicts with the manifest.
126
-
127
- ### Fixed
128
- - Windows provider cleanup now treats a successful `taskkill` as a request and confirms that the recorded PID has stopped before deleting its ownership record.
129
-
130
- ## 1.5.1 — 2026-08-17
131
-
132
- ### Fixed
133
- - Serena MCP configurations generated by Yoke no longer open the local web dashboard automatically on every startup.
134
-
135
- ## 1.5.0 — 2026-08-16
136
-
137
- ### Added
138
- - Failed verify, executable-criterion, performance, configured custom-audit, and completion gates now produce deterministic byte-bounded previews and preserve large complete stdout/stderr in content-addressed `.yoke/artifacts/` files with SHA-256 references.
139
- - Projects can tune `output.previewBytes` and `output.artifactThresholdBytes`; existing projects use backward-compatible 2 KiB/8 KiB defaults.
140
- - A deterministic local benchmark verifies signal retention, preview bounds, compression measurement, and artifact digest round-trips without making provider-token claims.
141
-
142
- ### Security
143
- - Output artifact paths sanitize story identifiers, stay below a project-local root, use user-only file modes where supported, and are excluded from Yoke's clean-tree and story-commit operations even in upgraded projects. Raw artifacts are never injected automatically and documentation warns that project commands may emit secrets or personal data.
144
- - Gate command capture is capped at 16 MiB per stdout/stderr stream. Quota overflow fails closed and labels retained evidence as truncated instead of risking unbounded memory or claiming partial output is complete.
145
-
146
- ## 1.4.0 — 2026-08-15
147
-
148
- ### Added
149
- - `yoke loop run --parallel=N` now executes dependency-ready, non-colliding stories through real provider subprocess workers, isolated worktrees, leased claims, and a FIFO integration queue with fresh integrated-system gates.
150
- - Reference-driven quality declarations can collect screenshots, files, command output, or benchmark results and run a schema-validated blind critic with bounded repair rounds, elapsed-time limits, blocking or advisory policy, and retained proof.
151
- - `--candidates=N` can fan out up to five isolated implementations, discard mechanically red candidates, select one green candidate through identity-blind pairwise comparison, and preserve selected/loser evidence before cleanup.
152
- - Loop status now exposes dispatcher, worker, integrator, candidate lifecycle, worktree, queue, integration, reopen, quality-round, repair-budget, and trusted provider/model provenance data.
153
-
154
- ### Changed
155
- - Provider subprocesses use explicit lifecycle contracts and incarnation-aware process records so worker cancellation and cleanup target only the process tree Yoke actually started.
156
- - Parallel and candidate runs disable adaptive routing, honor story-level provider affinity, latch pause requests across the whole dispatcher, and rerun quality plus review after integration.
157
- - `yoke loop cleanup` retains Yoke worktrees unless `--remove-worktrees` is explicit, while still reaping recorded orphan runners and stale locks safely.
158
-
159
- ### Fixed
160
- - Expired claims, worker crashes, merge conflicts, pause races, and integration failures now release ownership deterministically, retain terminal proof, and reopen stories without leaking worktrees or marking false completion.
161
- - Quality repair fails closed on malformed critic output, reference drift, provider/model provenance mismatch, candidate identity leakage, unavailable critics, exhausted limits, and mechanically red repairs.
162
- - The watchdog resolves its TypeScript loader from both source and built npm layouts on Node 20+, and read-only Codex comparisons can run in disposable candidate worktrees without weakening normal repository checks.
163
-
164
- ### Security
165
- - Blind comparison requests expose only opaque labels and digests while binding every verdict to the trusted judge provider, model, prompt, rubric, reference, and candidate provenance.
166
- - Cleanup and cancellation use project-scoped leases, owner tokens, PID birth/incarnation checks, and recorded process handles rather than machine-wide process-name matching.
167
-
168
- ## 1.3.0 — 2026-08-09
169
-
170
- ### Added
171
- - Teams can submit change requests at any time with `yoke change add` and inspect the append-only inbox with `yoke change status`; queued requests are planned at safe loop boundaries by Claude, Codex, or Gemini without interrupting active story work.
172
- - Acceptance criteria can carry stable IDs and executable verification commands, and Yoke records per-criterion evidence before a story can be marked complete.
173
- - Projects can configure an integrated completion command that proves the final system flow after all story-level gates pass.
174
-
175
- ### Changed
176
- - New projects use strict criterion-proof requirements by default, while existing boolean-only criteria remain readable for compatibility.
177
- - PRD authoring, schemas, loop guidance, generated configuration, and provider instructions now treat completion as an ephemeral readiness state: new requests create more stories instead of prematurely forcing a release boundary.
178
-
179
- ### Fixed
180
- - Broad green test suites can no longer satisfy unrelated structured acceptance criteria: proof commands must use an approved test runner, contain the criterion ID, and avoid shell operators.
181
- - Change-intake failures block safely; an independent coverage pass rejects omitted requested outcomes, crash recovery retains uncommitted requests, and concurrent PRD edits are detected instead of silently overwriting user work.
182
- - Integrated completion checks prevent locally finished stories from masking broken cross-component flows such as authentication callbacks or purchase-to-entitlement activation.
183
-
184
- ### Security
185
- - Updated the transitive development dependency `nanoid` to a non-vulnerable release so the complete CI dependency audit is clean.
186
- - Restricted model-authored criterion proof to criterion-targeted test commands and blocked shell control operators and runner-prefix spoofing before host execution.
187
-
188
- ## 1.2.1 — 2026-08-02
189
-
190
- ### Fixed
191
- - `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
192
- - Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
193
-
194
- ## 1.2.0 — 2026-08-02
195
-
196
- ### Added
197
- - Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
198
- - A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
199
- - Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
200
-
201
- ### Changed
202
- - Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
203
- - Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
204
-
205
- ### Fixed
206
- - Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
207
- - Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
208
-
209
- ## 1.1.0 — 2026-07-30
210
-
211
- ### Added
212
- - Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
213
- - Persisted default runner selection and `auto|critical` loop decision policies.
214
- - Provider-neutral `yoke-workflow` skill for planning questions, approved-plan PRD handoff, autonomous story execution, and critical-decision resume.
215
- - Structured critical-decision requests with `yoke loop decision` and `yoke loop answer`; answers are validated, committed under the configured human identity, and resume the same story.
216
- - Approved `.yoke/plan.md` context in PRD drafting and a lint gate for unresolved planning placeholders.
217
-
218
- ### Fixed
219
- - Retrofit and loop on/off now preserve timeout, decision, runner, and permission settings.
220
- - Empty projects prefer the active agent host instead of silently installing Claude artifacts in Codex.
221
- - Loop and PRD runner selection now prefer an explicit flag, then the configured runner, then the active host.
222
- - Retrofit reports no longer label every provider as Claude Code.
223
- - Critical-decision resumes retain isolation, review, runner, permissions, timeout, JSON, policy, and iteration settings instead of falling back to an unreviewed default run.
224
- - Decision answers use an atomic owner-token lock and recoverable request journal, are checked against the active PRD story, bounded as untrusted data, and committed path-by-path so unrelated edits cannot enter the human-owned commit.
225
- - Decision recovery now binds the exact selected answer to its commit, rolls back only its own interrupted context append, namespaces resume state per project/worktree, and serializes cleanup with loop startup.
226
- - Active agent session markers now outrank globally configured provider home directories, and setup rejects partially invalid agent lists.
227
-
228
- ### Changed
229
- - New setups enable the loop by default and choose `decisionPolicy: auto`; the wizard can select `critical` or disable the loop.
230
- - Legacy `loop.onAmbiguity` and `--on-ambiguity` remain compatibility aliases.
231
-
232
- ## 1.0.0 — 2026-07-27
233
-
234
- ### Added
235
- - Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
236
- - Safe provider permission profiles and structured cross-provider telemetry.
237
- - Schema-validated independent review verdicts with explicit self-review opt-in.
238
- - Human-owned commit identity enforcement; AI co-author trailers default off.
239
- - `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
240
- - PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
241
- - Reproducible cross-runner benchmark schema and matrix launcher.
242
-
243
- ### Changed
244
- - Dangerous permission bypass is opt-in via `--unsafe`.
245
- - Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
246
- - Reviews no longer trust process exit code alone.
247
-
248
- ### Security
249
- - Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
250
-
251
-
252
- ## 0.9.0 — 2026-07-22
253
-
254
- ### Added
255
- - **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
256
- A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
257
- verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
258
- `--isolate` mode. A red benchmark blocks the story
259
- (`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
260
- gate is configured, the implementer prompt names the budget command so agents keep hot
261
- paths efficient and never "simplify away" an optimization without re-running the benchmark.
262
- - **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
263
- clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
264
- project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
265
- tests, the *why* of every optimization versioned so future agents don't clean fast code
266
- back to slow.
267
- - **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
268
- numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
269
- round — a criterion still needing a decision is not loop-ready.
270
-
271
- ## 0.8.0 — 2026-07-20
272
-
273
- ### Added
274
- - **Live progress + ETA.** Story completions are now first-class events: the console shows
275
- `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
276
- `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
277
- durations of stories completed **in this run** (current velocity) and falls back to the
278
- persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
279
- No data → no estimate, never an invented one.
280
- - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
281
- prompt now always forbids asking questions (a loop run has nobody to answer). Default
282
- `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
283
- loop never stops. Opt-in `abort`: the agent writes its open questions to
284
- `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
285
- story would otherwise pass on pre-existing green tests), and blocks with the question as
286
- the reason. Companion principle: clarifying questions belong in the planning round, before
287
- the loop starts.
288
-
289
- ## 0.7.0 — 2026-07-17
290
-
291
- ### Added
292
- - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
293
- version hint (npm/gh-style): a detached background refresher caches the registry's
294
- latest at most once a day; when it is newer, a one-line stderr hint suggests
295
- `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
296
- `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
297
- - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
298
- loop START only — never mid-run; the running process finishes on its version and
299
- the upgrade applies from the next invocation. Deliberately NOT the default:
300
- a gate harness must not change itself mid-project (determinism), and unreviewed
301
- auto-installs are a supply-chain hazard.
302
-
303
- ## 0.6.0 — 2026-07-17
304
-
305
- ### Added
306
- - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
307
- `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
308
- `yoke loop cleanup` kills exactly those recorded process trees — and only while no
309
- live loop holds the lock. Background: without a scoped mechanism, users and agents
310
- resorted to machine-wide pattern kills (every process matching
311
- `dangerously-skip-permissions`), which took down *healthy* runners of other projects
312
- mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
313
- safe path. `.yoke/runner.pid` is gitignored by retrofit.
314
-
315
- ## 0.5.0 — 2026-07-17
316
-
317
- Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
318
- (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
319
- idle window):
320
-
321
- ### Fixed
322
- - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
323
- Previously it killed only the spawned shell (`shell: true`), orphaning the actual
324
- agent process — which kept writing to the worktree (dirty-tree blocks, failing
325
- worktree removal) and kept burning API tokens. Observed in the field as ~10
326
- zombie `claude.exe` per machine plus surviving dev servers.
327
- - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
328
- the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
329
- processes and killed them at exactly the idle timeout — while the user saw dead air.
330
- The stream doubles as liveness; token usage is now reported on every run (not just
331
- `--json` mode).
332
-
333
- ### Changed
334
- - README: operating notes for driving the loop from inside an agent session
335
- (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
336
- outer shell-tool timeouts killing a foreground `yoke loop run` were the third
337
- observed "hang" pattern.
338
-
339
- ## 0.4.0 — 2026-07-17
340
-
341
- ### Added
342
- - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
343
- scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
344
- documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
345
- bounded final messages (cuts output-token waste). Review prompts now ground verdicts
346
- in observed evidence only and keep them brief.
347
-
348
- ### Fixed
349
- - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
350
- `git add -A` story commit swept the pause control file into history in
351
- un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
352
- blocked the resume run — the loop locked itself out.
353
-
354
- > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
355
- > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
356
-
357
- ## 0.3.0 — 2026-07-10
358
-
359
- ### Added
360
- - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
361
- (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
362
- then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
363
- - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
364
- installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
365
- daily-crawled extensions gallery.
366
- - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
367
- quality) with a fixed fixture, pre-written objective tests, and committed result data.
368
- - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
369
- sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
370
- (design generation paired with Yoke's design verification gates).
371
- - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
372
- stories across parallel workers (`needs` dependency field, claim files, merge queue,
373
- heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
374
-
375
- ### Fixed
376
- - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
377
- Windows, so the loop misreported an installed `gemini` as "not found on PATH"
378
- (found by the new benchmark harness).
379
- - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
380
- requires a value after `-p` and errored with "Not enough arguments following: p".
381
- Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
382
- (also found by the benchmark harness).
383
-
384
- ### Changed
385
- - README: npm install is now the primary quickstart path; documented plugin/extension
386
- installs and optional companions.
387
- - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
388
- `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
389
-
390
- ## 0.2.0 — 2026-07-09
391
-
392
- - First npm release as `@hecer/yoke`.
393
- - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
394
- token-usage + model-id reporting for the claude runner.
395
- - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
396
- `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
397
- - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
398
- loop with worktree isolation, watchdog, single-flight lock, commit integrity.
1
+ # Changelog
2
+
3
+ ## 1.12.0 — 2026-09-08
4
+
5
+ ### Fixed
6
+ - Parse Qwen Code's native assistant/result/structured-result output and cumulative native model statistics, ignoring nested subagent results and refusing terminal-error verdicts.
7
+ - Use Qwen's native headless approval flags, retain the sandbox in safe mode and explicitly permit shell tools there; exclude native delegation tools when Yoke owns concurrency.
8
+ - Detect native Qwen session/project markers and include Qwen in automatic loop review selection and CLI review validation.
9
+ - Install manual Qwen skills with native invocation restrictions and a supported RTK PreToolUse retry guard; back up settings and remove only Yoke's obsolete BeforeTool hook on retrofit.
10
+ - Accept up to 32 routing profiles so an all-agent setup remains valid.
11
+
12
+ ### Added
13
+ - Opt-in `setup --model-provider=deepseek,kimi` with DeepSeek V4 Flash/Pro and Kimi K2.6/K2.7 Code/K3 API configurations, environment key references, preserved custom settings and routing profiles.
14
+ - Explicit Qwen `PROTOCOL::MODEL` selectors mapped to native authentication/model arguments without changing global login settings.
15
+
16
+ ### Migration and validation limits
17
+ - Fresh Qwen setups now use one standard profile with the user's configured model. Existing workers remain unchanged unless reset with `--routing-preset`; configure stronger profiles for stronger assessments.
18
+ - API presets require separately configured credentials and are not model-quality or cost benchmarks. Read [Qwen model support](docs/QWEN-MODEL-SUPPORT.md) for exact behavior and migration.
19
+ - Tested with regression fixtures and real Qwen Code 0.23.0 against a synthetic local tool-calling server; no authenticated provider benchmarks or production sandbox validation. This release targets npm package 1.12.0; publication is triggered by the published GitHub release and verified separately.
20
+
21
+ ## 1.11.0 — 2026-09-07
22
+
23
+ ### Added
24
+ - Add Qwen Code (Alibaba) as fourth supported provider alongside Claude, Codex and Gemini.
25
+ - Detect Qwen host environment via `QWEN_CLI` and `QWEN_CLI_HOME` environment variables.
26
+ - Add Qwen routing workers with four capability tiers: `qwen-turbo-latest` (light), `qwen3-coder-plus` (standard/strong), `qwen3-235b-a22b` (frontier).
27
+ - Parse Qwen streaming telemetry for token usage and model reporting.
28
+ - Support Qwen in all CLI commands: `setup`, `loop`, `review`, `prd draft`, `prd assess`, `goal run`.
29
+
30
+ ### Changed
31
+ - Update project description from "three agents" to "four agents" to reflect Qwen support.
32
+ - Extend review resolution order to include Qwen for cross-model reviews.
33
+ - Update setup prompts to offer Qwen as agent and runner option.
34
+
35
+ ### Migration and validation limits
36
+ - Existing configurations remain compatible. New setups can select Qwen as agent/runner.
37
+ - Qwen CLI uses Gemini-style arguments (`--approval-mode`, `--output-format stream-json`). `bare`, `reasoningEffort` and `nativeMultiAgent` selections are not supported (like Gemini).
38
+ - Qwen routing profiles are configurable hypotheses, not authenticated benchmarks. Update installed packages and restart the dashboard/runner.
39
+
40
+ ## 1.10.0 — 2026-09-06
41
+
42
+ ### Added
43
+ - Adopt the Yoke wordmark in the GitHub and npm README.
44
+ - Prepare task assessments in one bounded planning call with `yoke prd assess`; draft and inbox planning bind assessments to requirements, upstream dependencies and the approved brief.
45
+ - Separate `planning.agent`/`model`/`reasoningEffort` from execution settings. New setups require prepared assessments and block missing worker profiles; existing configurations retain on-demand planning and parent fallback.
46
+ - Add selective reassessment, planning input limits and `routing.maxTier` to bound automatic model selection and escalation.
47
+ - Add attention-first project search and status filters, with separate goal/loop states and explicit stale activity warnings.
48
+ - Restore dashboard views and UTC period controls through URL links, browser history and refresh; cancel obsolete requests and bound project comparison concurrency.
49
+ - Compare recorded tokens, accepted tasks and reported costs against the preceding equal-duration period without inventing missing measurements or percentages from zero baselines.
50
+
51
+ ### Fixed
52
+ - Fix Windows safe-mode Codex execution by screening Store PowerShell/aliases and probing a native shell under the same sandbox before model work; keep permissions intact.
53
+ - Launch Windows provider executables/npm entry points with literal argv; bound preflight and provider lifetime, recognize streamed infrastructure failures, retain failed-worktree/process evidence, and guard against unconfirmed termination before another worker starts.
54
+ - Separate provider liveness, supervisor heartbeat, output and successful-tool progress in status/dashboard; label backlog percentages and preserve explicit bare startup through capability routing. See [Windows runner validation](docs/WINDOWS-RUNNER-VALIDATION.md).
55
+ - Keep local loop locks and dashboard status out of implementation commits and clean-worktree checks.
56
+ - Stage implementation files safely when runtime history directories are already ignored, including literal filenames and tracked deletions, for serial commits and parallel candidate snapshots.
57
+
58
+ ### Migration and validation limits
59
+ - Existing routing configurations retain on-demand assessment and parent fallback. Use `yoke prd assess` to prepare task packages; configure `planning.agent`, `planning.model`, `planning.reasoningEffort` and `routing.maxTier` to separate planning from bounded execution. New setups require prepared assessments and block missing profiles.
60
+ - Update installed packages and restart the dashboard/runner to use the new code. Existing running processes are not upgraded or instrumented retroactively. Windows preflight preserves sandbox permissions and changes only the provider environment.
61
+ - Dashboard comparisons report recorded measurements, not reconstructed history or calibrated savings. Tokens per minute are consumption rates, not generation speed. Live routed Windows validation covered Codex on the documented machine; cross-provider performance and universal Windows compatibility are not established. See the linked dashboard, batch-planning and runner validation records.
62
+
63
+ ## 1.9.0 — 2026-09-06
64
+
65
+ ### Added
66
+ - Add capability-based routing with persisted task assessments, explicit model/effort tiers, role eligibility and conservative use of independent task-class outcomes.
67
+ - Keep planning on the start model; reuse assessments across attempts/worktrees and invalidate them when task requirements change.
68
+ - Add bounded repair and tier escalation after mechanical gate failures, retaining the patch and forwarding failure evidence. Infrastructure failures do not trigger capability escalation.
69
+ - Apply task-based profiles to reviews, quality critics/repairs and goal execution; display implementation selection reasons and next escalation in the dashboard.
70
+ - Add setup options `--routing-strategy=capability` and `--routing-preset` for explicit migration. Preserve existing strategies and custom profiles by default.
71
+
72
+ ### Validation limits
73
+ - Initial profile tiers are configurable hypotheses, not authenticated model benchmarks, price estimates or calibrated success probabilities. See [capability routing](docs/CAPABILITY-ROUTING.md) for defaults and bounds.
74
+
75
+ ## 1.8.0 — 2026-09-06
76
+
77
+ ### Added
78
+ - Add persistent local measurement history and dashboard views for current work, usage/time and results, including UTC day/week/month filters, model history, project comparisons and consumption charts.
79
+ - Display current tasks, worker phases, integration progress and status age, with automatic refresh of the current-work view.
80
+ - Record explicit acceptances and show measured tokens and time per acceptance. Attribute available reviewer, critic and repair usage; distinguish actual reported models, unknown calls and partial costs.
81
+
82
+ ### Changed
83
+ - Enable routing in new setups and automatically select up to three parallel workers when all pending tasks declare write scopes. Dependencies and overlapping scopes still constrain dispatch. Preserve explicit opt-outs and use isolated worktrees by default.
84
+ - Share routing decisions between synchronous and asynchronous runners; support routed parallel workers, stable recovery history and explicit provider affinity.
85
+ - Reserve execution capacity through integration and disable native delegation for loop providers; preserve Gemini system policy in a temporary bounded-execution configuration.
86
+ - Require a dated changelog entry, synchronized version metadata and verified release checks for every new version in the project instructions.
87
+
88
+ ### Fixed
89
+ - Avoid conflicting Codex sandbox arguments and prevent Codex-only options from leaking into Gemini workers.
90
+ - Keep compact measurement history after recent activity expires, deduplicate archived events, and report incomplete history instead of treating missing usage as zero.
91
+ - Preserve reviewer telemetry and worker/model attribution across parallel execution and recovery.
92
+
93
+ ### Migration and validation limits
94
+ - Use `--parallel=N` to choose a worker limit, `--parallel=auto` for automatic selection, `--no-routing` to opt out of routing, and `--no-isolate` to opt out of default isolation. Explicit existing configuration remains authoritative. Unknown write scopes, tool actions and worktree recovery select serial execution in auto mode.
95
+ - Automatic routing needs configured profiles; otherwise it keeps the selected parent provider. Explicit `--routing` without profiles reports a configuration error.
96
+ - Missing historical usage cannot be reconstructed. Tokens per minute describe interval or summed call consumption, not measured generation speed. Live authenticated provider benchmarks, resource-adaptive concurrency and calibrated time/cost predictions are not established by this release.
97
+
98
+ ## 1.7.0 — 2026-09-05
99
+
100
+ ### Added
101
+ - Independent `yoke check` with executable acceptance mapping, protected test infrastructure and content-bound evidence.
102
+ - Durable project goals, provider handoff, checkpoint budgets, interruption accounting and project-scoped recovery.
103
+ - Local project registry and loopback dashboard with goals, task estimates, evidence, consumption and pause controls.
104
+ - Explicit routing rules with persisted gate-driven escalation; bounded tool actions without model calls.
105
+ - Task-aware context packets, advisory write scopes, dependency-depth scheduling and empirical time ranges with prediction error records.
106
+
107
+ ### Fixed
108
+ - Failed serial isolated worktrees are retained and can be explicitly resumed against their original target and PRD.
109
+ - Reviewer fingerprints include untracked contents and acceptance inputs; unsupported nested repository identity fails closed.
110
+ - Gemini always emits streaming telemetry, preserves model identity, honors aggregate token aliases and rejects unsupported selections. Its RTK hook merges native settings without a shell dependency.
111
+ - Runtime evidence is excluded from Git gates and commits in existing projects. Partial usage and costs remain visibly incomplete.
112
+ - Protected acceptance is checked after verification and repair, and linked goal state cannot overwrite unrelated files through pause.
113
+
114
+ ### Validation limits
115
+ - Live authenticated provider comparisons and calibrated development-time/cost estimates are not established by the automated tests. Browser proof and independent model review still require explicit configuration.
116
+
117
+ ## 1.6.2 — 2026-09-02
118
+
119
+ ### Added
120
+ - GitHub Releases can now publish `@hecer/yoke` through npm trusted publishing with short-lived OIDC credentials and automatic provenance, without a long-lived npm token.
121
+
122
+ ### Fixed
123
+ - Codex safe-mode invocations now use the supported `workspace-write` sandbox with `--approve-for-me`; the removed `--full-auto` flag no longer blocks current Codex CLI releases.
124
+ - Windows provider cleanup rechecks termination after process close and accepts an already-absent process as successfully cleaned up, avoiding stale ownership records and unnecessary watchdog waits.
125
+ - Successful stale-loop cleanup removes obsolete runtime status, so `yoke loop status` no longer reports a dead run as `RUNNING`.
126
+
127
+ ## 1.6.1 — 2026-08-21
128
+
129
+ ### Fixed
130
+ - Release metadata counts platform-conditional tests consistently on Windows and Ubuntu.
131
+ - Provider cleanup confirms termination after child close, preventing stale ownership records when Windows reports process exit asynchronously.
132
+
133
+ ## 1.6.0 — 2026-08-20
134
+
135
+ ### Added
136
+ - The canon now includes `no-ai-slop`, `domain-modeling`, `codebase-design`, `resolving-merge-conflicts`, and `writing-for-agents`, with their supporting templates and evaluation material. Source adaptations are credited in `canon/skills/ATTRIBUTION.md`.
137
+ - UI projects receive an automatic design-quality gate with configurable `design.mode` (`off`, `auto`, or `on`) and score budget. Detection uses package dependencies, UI source files, and configured smoke flows.
138
+ - Durable context now includes a project glossary and can expose an optional bounded-context map.
139
+
140
+ ### Changed
141
+ - Retrofit installs complete skill packages for Claude, Codex, and Gemini instead of copying only `SKILL.md`. Local resources, binary files, and executable bits are preserved where supported.
142
+ - Every canon skill declares `invocation: auto|manual`; retrofit maps that policy to Claude frontmatter, Codex `agents/openai.yaml`, and Gemini's generated automatic-skill index.
143
+ - Canon validation rejects missing local Markdown resources, unsafe package entries, symlinks, and provider metadata that conflicts with the manifest.
144
+
145
+ ### Fixed
146
+ - Windows provider cleanup now treats a successful `taskkill` as a request and confirms that the recorded PID has stopped before deleting its ownership record.
147
+
148
+ ## 1.5.1 — 2026-08-17
149
+
150
+ ### Fixed
151
+ - Serena MCP configurations generated by Yoke no longer open the local web dashboard automatically on every startup.
152
+
153
+ ## 1.5.0 — 2026-08-16
154
+
155
+ ### Added
156
+ - Failed verify, executable-criterion, performance, configured custom-audit, and completion gates now produce deterministic byte-bounded previews and preserve large complete stdout/stderr in content-addressed `.yoke/artifacts/` files with SHA-256 references.
157
+ - Projects can tune `output.previewBytes` and `output.artifactThresholdBytes`; existing projects use backward-compatible 2 KiB/8 KiB defaults.
158
+ - A deterministic local benchmark verifies signal retention, preview bounds, compression measurement, and artifact digest round-trips without making provider-token claims.
159
+
160
+ ### Security
161
+ - Output artifact paths sanitize story identifiers, stay below a project-local root, use user-only file modes where supported, and are excluded from Yoke's clean-tree and story-commit operations even in upgraded projects. Raw artifacts are never injected automatically and documentation warns that project commands may emit secrets or personal data.
162
+ - Gate command capture is capped at 16 MiB per stdout/stderr stream. Quota overflow fails closed and labels retained evidence as truncated instead of risking unbounded memory or claiming partial output is complete.
163
+
164
+ ## 1.4.0 — 2026-08-15
165
+
166
+ ### Added
167
+ - `yoke loop run --parallel=N` now executes dependency-ready, non-colliding stories through real provider subprocess workers, isolated worktrees, leased claims, and a FIFO integration queue with fresh integrated-system gates.
168
+ - Reference-driven quality declarations can collect screenshots, files, command output, or benchmark results and run a schema-validated blind critic with bounded repair rounds, elapsed-time limits, blocking or advisory policy, and retained proof.
169
+ - `--candidates=N` can fan out up to five isolated implementations, discard mechanically red candidates, select one green candidate through identity-blind pairwise comparison, and preserve selected/loser evidence before cleanup.
170
+ - Loop status now exposes dispatcher, worker, integrator, candidate lifecycle, worktree, queue, integration, reopen, quality-round, repair-budget, and trusted provider/model provenance data.
171
+
172
+ ### Changed
173
+ - Provider subprocesses use explicit lifecycle contracts and incarnation-aware process records so worker cancellation and cleanup target only the process tree Yoke actually started.
174
+ - Parallel and candidate runs disable adaptive routing, honor story-level provider affinity, latch pause requests across the whole dispatcher, and rerun quality plus review after integration.
175
+ - `yoke loop cleanup` retains Yoke worktrees unless `--remove-worktrees` is explicit, while still reaping recorded orphan runners and stale locks safely.
176
+
177
+ ### Fixed
178
+ - Expired claims, worker crashes, merge conflicts, pause races, and integration failures now release ownership deterministically, retain terminal proof, and reopen stories without leaking worktrees or marking false completion.
179
+ - Quality repair fails closed on malformed critic output, reference drift, provider/model provenance mismatch, candidate identity leakage, unavailable critics, exhausted limits, and mechanically red repairs.
180
+ - The watchdog resolves its TypeScript loader from both source and built npm layouts on Node 20+, and read-only Codex comparisons can run in disposable candidate worktrees without weakening normal repository checks.
181
+
182
+ ### Security
183
+ - Blind comparison requests expose only opaque labels and digests while binding every verdict to the trusted judge provider, model, prompt, rubric, reference, and candidate provenance.
184
+ - Cleanup and cancellation use project-scoped leases, owner tokens, PID birth/incarnation checks, and recorded process handles rather than machine-wide process-name matching.
185
+
186
+ ## 1.3.0 — 2026-08-09
187
+
188
+ ### Added
189
+ - Teams can submit change requests at any time with `yoke change add` and inspect the append-only inbox with `yoke change status`; queued requests are planned at safe loop boundaries by Claude, Codex, or Gemini without interrupting active story work.
190
+ - Acceptance criteria can carry stable IDs and executable verification commands, and Yoke records per-criterion evidence before a story can be marked complete.
191
+ - Projects can configure an integrated completion command that proves the final system flow after all story-level gates pass.
192
+
193
+ ### Changed
194
+ - New projects use strict criterion-proof requirements by default, while existing boolean-only criteria remain readable for compatibility.
195
+ - PRD authoring, schemas, loop guidance, generated configuration, and provider instructions now treat completion as an ephemeral readiness state: new requests create more stories instead of prematurely forcing a release boundary.
196
+
197
+ ### Fixed
198
+ - Broad green test suites can no longer satisfy unrelated structured acceptance criteria: proof commands must use an approved test runner, contain the criterion ID, and avoid shell operators.
199
+ - Change-intake failures block safely; an independent coverage pass rejects omitted requested outcomes, crash recovery retains uncommitted requests, and concurrent PRD edits are detected instead of silently overwriting user work.
200
+ - Integrated completion checks prevent locally finished stories from masking broken cross-component flows such as authentication callbacks or purchase-to-entitlement activation.
201
+
202
+ ### Security
203
+ - Updated the transitive development dependency `nanoid` to a non-vulnerable release so the complete CI dependency audit is clean.
204
+ - Restricted model-authored criterion proof to criterion-targeted test commands and blocked shell control operators and runner-prefix spoofing before host execution.
205
+
206
+ ## 1.2.1 — 2026-08-02
207
+
208
+ ### Fixed
209
+ - `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
210
+ - Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
211
+
212
+ ## 1.2.0 — 2026-08-02
213
+
214
+ ### Added
215
+ - Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
216
+ - A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
217
+ - Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
218
+
219
+ ### Changed
220
+ - Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
221
+ - Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
222
+
223
+ ### Fixed
224
+ - Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
225
+ - Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
226
+
227
+ ## 1.1.0 — 2026-07-30
228
+
229
+ ### Added
230
+ - Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
231
+ - Persisted default runner selection and `auto|critical` loop decision policies.
232
+ - Provider-neutral `yoke-workflow` skill for planning questions, approved-plan PRD handoff, autonomous story execution, and critical-decision resume.
233
+ - Structured critical-decision requests with `yoke loop decision` and `yoke loop answer`; answers are validated, committed under the configured human identity, and resume the same story.
234
+ - Approved `.yoke/plan.md` context in PRD drafting and a lint gate for unresolved planning placeholders.
235
+
236
+ ### Fixed
237
+ - Retrofit and loop on/off now preserve timeout, decision, runner, and permission settings.
238
+ - Empty projects prefer the active agent host instead of silently installing Claude artifacts in Codex.
239
+ - Loop and PRD runner selection now prefer an explicit flag, then the configured runner, then the active host.
240
+ - Retrofit reports no longer label every provider as Claude Code.
241
+ - Critical-decision resumes retain isolation, review, runner, permissions, timeout, JSON, policy, and iteration settings instead of falling back to an unreviewed default run.
242
+ - Decision answers use an atomic owner-token lock and recoverable request journal, are checked against the active PRD story, bounded as untrusted data, and committed path-by-path so unrelated edits cannot enter the human-owned commit.
243
+ - Decision recovery now binds the exact selected answer to its commit, rolls back only its own interrupted context append, namespaces resume state per project/worktree, and serializes cleanup with loop startup.
244
+ - Active agent session markers now outrank globally configured provider home directories, and setup rejects partially invalid agent lists.
245
+
246
+ ### Changed
247
+ - New setups enable the loop by default and choose `decisionPolicy: auto`; the wizard can select `critical` or disable the loop.
248
+ - Legacy `loop.onAmbiguity` and `--on-ambiguity` remain compatibility aliases.
249
+
250
+ ## 1.0.0 — 2026-07-27
251
+
252
+ ### Added
253
+ - Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
254
+ - Safe provider permission profiles and structured cross-provider telemetry.
255
+ - Schema-validated independent review verdicts with explicit self-review opt-in.
256
+ - Human-owned commit identity enforcement; AI co-author trailers default off.
257
+ - `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
258
+ - PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
259
+ - Reproducible cross-runner benchmark schema and matrix launcher.
260
+
261
+ ### Changed
262
+ - Dangerous permission bypass is opt-in via `--unsafe`.
263
+ - Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
264
+ - Reviews no longer trust process exit code alone.
265
+
266
+ ### Security
267
+ - Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
268
+
269
+
270
+ ## 0.9.0 — 2026-07-22
271
+
272
+ ### Added
273
+ - **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
274
+ A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
275
+ verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
276
+ `--isolate` mode. A red benchmark blocks the story
277
+ (`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
278
+ gate is configured, the implementer prompt names the budget command so agents keep hot
279
+ paths efficient and never "simplify away" an optimization without re-running the benchmark.
280
+ - **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
281
+ clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
282
+ project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
283
+ tests, the *why* of every optimization versioned so future agents don't clean fast code
284
+ back to slow.
285
+ - **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
286
+ numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
287
+ round — a criterion still needing a decision is not loop-ready.
288
+
289
+ ## 0.8.0 — 2026-07-20
290
+
291
+ ### Added
292
+ - **Live progress + ETA.** Story completions are now first-class events: the console shows
293
+ `✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
294
+ `yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
295
+ durations of stories completed **in this run** (current velocity) and falls back to the
296
+ persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
297
+ No data → no estimate, never an invented one.
298
+ - **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
299
+ prompt now always forbids asking questions (a loop run has nobody to answer). Default
300
+ `resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
301
+ loop never stops. Opt-in `abort`: the agent writes its open questions to
302
+ `.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
303
+ story would otherwise pass on pre-existing green tests), and blocks with the question as
304
+ the reason. Companion principle: clarifying questions belong in the planning round, before
305
+ the loop starts.
306
+
307
+ ## 0.7.0 — 2026-07-17
308
+
309
+ ### Added
310
+ - **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
311
+ version hint (npm/gh-style): a detached background refresher caches the registry's
312
+ latest at most once a day; when it is newer, a one-line stderr hint suggests
313
+ `yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
314
+ `--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
315
+ - **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
316
+ loop START only — never mid-run; the running process finishes on its version and
317
+ the upgrade applies from the next invocation. Deliberately NOT the default:
318
+ a gate harness must not change itself mid-project (determinism), and unreviewed
319
+ auto-installs are a supply-chain hazard.
320
+
321
+ ## 0.6.0 — 2026-07-17
322
+
323
+ ### Added
324
+ - **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
325
+ `.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
326
+ `yoke loop cleanup` kills exactly those recorded process trees — and only while no
327
+ live loop holds the lock. Background: without a scoped mechanism, users and agents
328
+ resorted to machine-wide pattern kills (every process matching
329
+ `dangerously-skip-permissions`), which took down *healthy* runners of other projects
330
+ mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
331
+ safe path. `.yoke/runner.pid` is gitignored by retrofit.
332
+
333
+ ## 0.5.0 — 2026-07-17
334
+
335
+ Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
336
+ (orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
337
+ idle window):
338
+
339
+ ### Fixed
340
+ - **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
341
+ Previously it killed only the spawned shell (`shell: true`), orphaning the actual
342
+ agent process — which kept writing to the worktree (dirty-tree blocks, failing
343
+ worktree removal) and kept burning API tokens. Observed in the field as ~10
344
+ zombie `claude.exe` per machine plus surviving dev servers.
345
+ - **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
346
+ the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
347
+ processes and killed them at exactly the idle timeout — while the user saw dead air.
348
+ The stream doubles as liveness; token usage is now reported on every run (not just
349
+ `--json` mode).
350
+
351
+ ### Changed
352
+ - README: operating notes for driving the loop from inside an agent session
353
+ (background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
354
+ outer shell-tool timeouts killing a foreground `yoke loop run` were the third
355
+ observed "hang" pattern.
356
+
357
+ ## 0.4.0 — 2026-07-17
358
+
359
+ ### Added
360
+ - **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
361
+ scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
362
+ documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
363
+ bounded final messages (cuts output-token waste). Review prompts now ground verdicts
364
+ in observed evidence only and keep them brief.
365
+
366
+ ### Fixed
367
+ - `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
368
+ `git add -A` story commit swept the pause control file into history in
369
+ un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
370
+ blocked the resume run — the loop locked itself out.
371
+
372
+ > Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
373
+ > was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
374
+
375
+ ## 0.3.0 — 2026-07-10
376
+
377
+ ### Added
378
+ - **Claude Code plugin packaging** — the repo is now its own plugin marketplace
379
+ (`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
380
+ then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
381
+ - **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
382
+ installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
383
+ daily-crawled extensions gallery.
384
+ - **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
385
+ quality) with a fixed fixture, pre-written objective tests, and committed result data.
386
+ - **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
387
+ sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
388
+ (design generation paired with Yoke's design verification gates).
389
+ - **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
390
+ stories across parallel workers (`needs` dependency field, claim files, merge queue,
391
+ heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
392
+
393
+ ### Fixed
394
+ - Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
395
+ Windows, so the loop misreported an installed `gemini` as "not found on PATH"
396
+ (found by the new benchmark harness).
397
+ - Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
398
+ requires a value after `-p` and errored with "Not enough arguments following: p".
399
+ Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
400
+ (also found by the benchmark harness).
401
+
402
+ ### Changed
403
+ - README: npm install is now the primary quickstart path; documented plugin/extension
404
+ installs and optional companions.
405
+ - npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
406
+ `docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
407
+
408
+ ## 0.2.0 — 2026-07-09
409
+
410
+ - First npm release as `@hecer/yoke`.
411
+ - Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
412
+ token-usage + model-id reporting for the claude runner.
413
+ - `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
414
+ `yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
415
+ - Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
416
+ loop with worktree isolation, watchdog, single-flight lock, commit integrity.