@hecer/yoke 1.5.1 → 1.6.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +13 -13
- package/.codex-plugin/plugin.json +7 -7
- package/CHANGELOG.md +280 -259
- package/README.md +855 -834
- package/TODOS.md +5 -5
- package/agents/docs.toml +6 -6
- package/agents/implementer.toml +6 -6
- package/agents/reviewer.toml +6 -6
- package/agents/security.toml +6 -6
- package/bench/README.md +86 -86
- package/bench/RESULTS.md +35 -35
- package/bench/output-compaction.mjs +65 -65
- package/bench/result-schema.mjs +12 -12
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
- package/bench/results/codex-unavailable-1785175418318.json +15 -15
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
- package/bench/run-matrix.mjs +26 -26
- package/bench/run.mjs +106 -106
- package/canon/AGENTS.md +30 -30
- package/canon/context/DECISIONS.md +4 -4
- package/canon/context/GLOSSARY.md +11 -0
- package/canon/context/KNOWLEDGE.md +4 -4
- package/canon/context/PROJECT.md +15 -15
- package/canon/loop/loop-spec.md +65 -65
- package/canon/loop/prd.schema.md +43 -43
- package/canon/manifest.yaml +59 -53
- package/canon/policy/gates.md +7 -7
- package/canon/policy/roles.md +9 -9
- package/canon/skills/ATTRIBUTION.md +99 -71
- package/canon/skills/authoring-prd/SKILL.md +58 -58
- package/canon/skills/brainstorming/SKILL.md +164 -164
- package/canon/skills/codebase-design/DEEPENING.md +15 -0
- package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -0
- package/canon/skills/codebase-design/SKILL.md +39 -0
- package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
- package/canon/skills/document-release/SKILL.md +302 -297
- package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -0
- package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -0
- package/canon/skills/domain-modeling/SKILL.md +35 -0
- package/canon/skills/executing-plans/SKILL.md +70 -70
- package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
- package/canon/skills/health/SKILL.md +177 -177
- package/canon/skills/maintaining-context/SKILL.md +34 -34
- package/canon/skills/minimal-code/SKILL.md +21 -21
- package/canon/skills/no-ai-slop/SKILL.md +103 -0
- package/canon/skills/no-ai-slop/eval.md +43 -0
- package/canon/skills/plan-ceo-review/SKILL.md +541 -541
- package/canon/skills/plan-eng-review/SKILL.md +362 -362
- package/canon/skills/receiving-code-review/SKILL.md +213 -213
- package/canon/skills/requesting-code-review/SKILL.md +105 -105
- package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -0
- package/canon/skills/retro/SKILL.md +397 -397
- package/canon/skills/review/SKILL.md +246 -246
- package/canon/skills/ship/SKILL.md +691 -691
- package/canon/skills/subagent-driven-development/SKILL.md +277 -277
- package/canon/skills/systematic-debugging/SKILL.md +296 -296
- package/canon/skills/tdd/SKILL.md +371 -371
- package/canon/skills/unslop-ui/SKILL.md +34 -34
- package/canon/skills/using-git-worktrees/SKILL.md +218 -218
- package/canon/skills/verification-before-completion/SKILL.md +139 -139
- package/canon/skills/visual-verification/SKILL.md +54 -54
- package/canon/skills/workflow/SKILL.md +22 -22
- package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -0
- package/canon/skills/writing-for-agents/SKILL.md +42 -0
- package/canon/skills/writing-plans/SKILL.md +152 -152
- package/canon/skills/writing-skills/SKILL.md +655 -655
- package/canon/skills/yoke-retrofit/SKILL.md +26 -26
- package/canon/skills/yoke-workflow/SKILL.md +20 -20
- package/canon/tools/codex-rtk-hook.mjs +35 -35
- package/canon/tools/graphify.md +3 -3
- package/canon/tools/playwright-mcp.md +3 -3
- package/canon/tools/rtk.md +7 -7
- package/canon/tools/serena.md +6 -6
- package/dist/agents/process.js +3 -0
- package/dist/canon/manifest.js +2 -0
- package/dist/canon/skill-package.js +113 -0
- package/dist/canon/validate.js +16 -1
- package/dist/context/command.js +4 -1
- package/dist/context/context.js +6 -0
- package/dist/loop/dispatcher.js +1 -1
- package/dist/loop/loop.js +26 -0
- package/dist/loop/parallel-command.js +3 -0
- package/dist/loop/run-command.js +11 -0
- package/dist/loop/watchdog.js +28 -11
- package/dist/loop/worker.js +11 -0
- package/dist/prd/command.js +17 -17
- package/dist/retrofit/apply.js +22 -7
- package/dist/retrofit/command.js +4 -1
- package/dist/retrofit/config.js +4 -0
- package/dist/retrofit/context-actions.js +1 -1
- package/dist/retrofit/detect.js +2 -0
- package/dist/retrofit/planners/claude.js +16 -20
- package/dist/retrofit/planners/codex.js +3 -7
- package/dist/retrofit/planners/gemini.js +11 -1
- package/dist/retrofit/preserve.js +2 -2
- package/dist/retrofit/report.js +5 -0
- package/dist/retrofit/skill-actions.js +66 -0
- package/dist/retrofit/ui-detect.js +83 -0
- package/dist/scan/gate.js +36 -0
- package/docs/MIGRATING-TO-1.0.md +33 -33
- package/docs/MIGRATING-TO-1.1.md +27 -27
- package/docs/MIGRATING-TO-1.4.md +70 -70
- package/docs/PUBLISHING.md +91 -91
- package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
- package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
- package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
- package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
- package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
- package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
- package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
- package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
- package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
- package/docs/superpowers/plans/2026-08-20-automatic-ui-design-gate.md +59 -0
- package/docs/superpowers/plans/2026-08-20-capability-skills-and-context.md +51 -0
- package/docs/superpowers/plans/2026-08-20-complete-skill-packages-and-invocation.md +59 -0
- package/docs/superpowers/plans/2026-08-20-windows-reliability-and-release.md +67 -0
- package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
- package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
- package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
- package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
- package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
- package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
- package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
- package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
- package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
- package/docs/superpowers/specs/2026-08-20-skill-capabilities-and-reliability-design.md +391 -0
- package/gemini-extension.json +6 -6
- package/hooks/hooks.json +19 -19
- package/package.json +84 -84
|
@@ -1,13 +1,13 @@
|
|
|
1
|
-
{
|
|
2
|
-
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
|
|
3
|
-
"name": "yoke",
|
|
4
|
-
"displayName": "Yoke",
|
|
5
|
-
"version": "1.
|
|
6
|
-
"description": "Cross-agent coding harness: one curated skill canon (TDD, brainstorming, plans, reviews, shipping, design verification) plus mechanical safety gates and an autonomous loop via the yoke CLI.",
|
|
7
|
-
"author": { "name": "HECer", "url": "https://github.com/HECer" },
|
|
8
|
-
"homepage": "https://github.com/HECer/yoke#readme",
|
|
9
|
-
"repository": "https://github.com/HECer/yoke",
|
|
10
|
-
"license": "MIT",
|
|
11
|
-
"keywords": ["harness", "cross-agent", "tdd", "code-review", "autonomous-loop", "codex", "gemini-cli"],
|
|
12
|
-
"skills": "./canon/skills/"
|
|
13
|
-
}
|
|
1
|
+
{
|
|
2
|
+
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
|
|
3
|
+
"name": "yoke",
|
|
4
|
+
"displayName": "Yoke",
|
|
5
|
+
"version": "1.6.1",
|
|
6
|
+
"description": "Cross-agent coding harness: one curated skill canon (TDD, brainstorming, plans, reviews, shipping, design verification) plus mechanical safety gates and an autonomous loop via the yoke CLI.",
|
|
7
|
+
"author": { "name": "HECer", "url": "https://github.com/HECer" },
|
|
8
|
+
"homepage": "https://github.com/HECer/yoke#readme",
|
|
9
|
+
"repository": "https://github.com/HECer/yoke",
|
|
10
|
+
"license": "MIT",
|
|
11
|
+
"keywords": ["harness", "cross-agent", "tdd", "code-review", "autonomous-loop", "codex", "gemini-cli"],
|
|
12
|
+
"skills": "./canon/skills/"
|
|
13
|
+
}
|
|
@@ -1,7 +1,7 @@
|
|
|
1
|
-
{
|
|
2
|
-
"name": "yoke",
|
|
3
|
-
"version": "1.
|
|
4
|
-
"description": "Cross-agent coding discipline, mechanical gates, and release workflows",
|
|
5
|
-
"skills": "./canon/skills/",
|
|
6
|
-
"hooks": "./hooks/hooks.json"
|
|
7
|
-
}
|
|
1
|
+
{
|
|
2
|
+
"name": "yoke",
|
|
3
|
+
"version": "1.6.1",
|
|
4
|
+
"description": "Cross-agent coding discipline, mechanical gates, and release workflows",
|
|
5
|
+
"skills": "./canon/skills/",
|
|
6
|
+
"hooks": "./hooks/hooks.json"
|
|
7
|
+
}
|
package/CHANGELOG.md
CHANGED
|
@@ -1,273 +1,294 @@
|
|
|
1
|
-
# Changelog
|
|
2
|
-
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
3
|
## Unreleased
|
|
4
4
|
|
|
5
|
+
## 1.6.1 — 2026-08-21
|
|
6
|
+
|
|
7
|
+
### Fixed
|
|
8
|
+
- Release metadata now counts platform-conditional tests consistently on Windows and Ubuntu, so `docs:check` no longer fails after an otherwise green cross-platform test matrix.
|
|
9
|
+
- Provider cleanup now rechecks termination after the child closes, preventing stale ownership records when Windows reports process exit asynchronously.
|
|
10
|
+
|
|
11
|
+
## 1.6.0 — 2026-08-20
|
|
12
|
+
|
|
13
|
+
### Added
|
|
14
|
+
- The canon now includes `no-ai-slop`, `domain-modeling`, `codebase-design`, `resolving-merge-conflicts`, and `writing-for-agents`, with their supporting templates and evaluation material. Source adaptations are credited in `canon/skills/ATTRIBUTION.md`.
|
|
15
|
+
- UI projects receive an automatic design-quality gate with configurable `design.mode` (`off`, `auto`, or `on`) and score budget. Detection uses package dependencies, UI source files, and configured smoke flows.
|
|
16
|
+
- Durable context now includes a project glossary and can expose an optional bounded-context map.
|
|
17
|
+
|
|
18
|
+
### Changed
|
|
19
|
+
- Retrofit installs complete skill packages for Claude, Codex, and Gemini instead of copying only `SKILL.md`. Local resources, binary files, and executable bits are preserved where supported.
|
|
20
|
+
- Every canon skill declares `invocation: auto|manual`; retrofit maps that policy to Claude frontmatter, Codex `agents/openai.yaml`, and Gemini's generated automatic-skill index.
|
|
21
|
+
- Canon validation rejects missing local Markdown resources, unsafe package entries, symlinks, and provider metadata that conflicts with the manifest.
|
|
22
|
+
|
|
23
|
+
### Fixed
|
|
24
|
+
- Windows provider cleanup now treats a successful `taskkill` as a request and confirms that the recorded PID has stopped before deleting its ownership record.
|
|
25
|
+
|
|
5
26
|
## 1.5.1 — 2026-08-17
|
|
6
27
|
|
|
7
28
|
### Fixed
|
|
8
29
|
- Serena MCP configurations generated by Yoke no longer open the local web dashboard automatically on every startup.
|
|
9
30
|
|
|
10
31
|
## 1.5.0 — 2026-08-16
|
|
11
|
-
|
|
12
|
-
### Added
|
|
32
|
+
|
|
33
|
+
### Added
|
|
13
34
|
- Failed verify, executable-criterion, performance, configured custom-audit, and completion gates now produce deterministic byte-bounded previews and preserve large complete stdout/stderr in content-addressed `.yoke/artifacts/` files with SHA-256 references.
|
|
14
|
-
- Projects can tune `output.previewBytes` and `output.artifactThresholdBytes`; existing projects use backward-compatible 2 KiB/8 KiB defaults.
|
|
15
|
-
- A deterministic local benchmark verifies signal retention, preview bounds, compression measurement, and artifact digest round-trips without making provider-token claims.
|
|
16
|
-
|
|
35
|
+
- Projects can tune `output.previewBytes` and `output.artifactThresholdBytes`; existing projects use backward-compatible 2 KiB/8 KiB defaults.
|
|
36
|
+
- A deterministic local benchmark verifies signal retention, preview bounds, compression measurement, and artifact digest round-trips without making provider-token claims.
|
|
37
|
+
|
|
17
38
|
### Security
|
|
18
39
|
- Output artifact paths sanitize story identifiers, stay below a project-local root, use user-only file modes where supported, and are excluded from Yoke's clean-tree and story-commit operations even in upgraded projects. Raw artifacts are never injected automatically and documentation warns that project commands may emit secrets or personal data.
|
|
19
40
|
- Gate command capture is capped at 16 MiB per stdout/stderr stream. Quota overflow fails closed and labels retained evidence as truncated instead of risking unbounded memory or claiming partial output is complete.
|
|
20
41
|
|
|
21
42
|
## 1.4.0 — 2026-08-15
|
|
22
|
-
|
|
23
|
-
### Added
|
|
24
|
-
- `yoke loop run --parallel=N` now executes dependency-ready, non-colliding stories through real provider subprocess workers, isolated worktrees, leased claims, and a FIFO integration queue with fresh integrated-system gates.
|
|
25
|
-
- Reference-driven quality declarations can collect screenshots, files, command output, or benchmark results and run a schema-validated blind critic with bounded repair rounds, elapsed-time limits, blocking or advisory policy, and retained proof.
|
|
26
|
-
- `--candidates=N` can fan out up to five isolated implementations, discard mechanically red candidates, select one green candidate through identity-blind pairwise comparison, and preserve selected/loser evidence before cleanup.
|
|
27
|
-
- Loop status now exposes dispatcher, worker, integrator, candidate lifecycle, worktree, queue, integration, reopen, quality-round, repair-budget, and trusted provider/model provenance data.
|
|
28
|
-
|
|
29
|
-
### Changed
|
|
30
|
-
- Provider subprocesses use explicit lifecycle contracts and incarnation-aware process records so worker cancellation and cleanup target only the process tree Yoke actually started.
|
|
31
|
-
- Parallel and candidate runs disable adaptive routing, honor story-level provider affinity, latch pause requests across the whole dispatcher, and rerun quality plus review after integration.
|
|
32
|
-
- `yoke loop cleanup` retains Yoke worktrees unless `--remove-worktrees` is explicit, while still reaping recorded orphan runners and stale locks safely.
|
|
33
|
-
|
|
34
|
-
### Fixed
|
|
35
|
-
- Expired claims, worker crashes, merge conflicts, pause races, and integration failures now release ownership deterministically, retain terminal proof, and reopen stories without leaking worktrees or marking false completion.
|
|
36
|
-
- Quality repair fails closed on malformed critic output, reference drift, provider/model provenance mismatch, candidate identity leakage, unavailable critics, exhausted limits, and mechanically red repairs.
|
|
37
|
-
- The watchdog resolves its TypeScript loader from both source and built npm layouts on Node 20+, and read-only Codex comparisons can run in disposable candidate worktrees without weakening normal repository checks.
|
|
38
|
-
|
|
39
|
-
### Security
|
|
40
|
-
- Blind comparison requests expose only opaque labels and digests while binding every verdict to the trusted judge provider, model, prompt, rubric, reference, and candidate provenance.
|
|
41
|
-
- Cleanup and cancellation use project-scoped leases, owner tokens, PID birth/incarnation checks, and recorded process handles rather than machine-wide process-name matching.
|
|
42
|
-
|
|
43
|
-
## 1.3.0 — 2026-08-09
|
|
44
|
-
|
|
45
|
-
### Added
|
|
46
|
-
- Teams can submit change requests at any time with `yoke change add` and inspect the append-only inbox with `yoke change status`; queued requests are planned at safe loop boundaries by Claude, Codex, or Gemini without interrupting active story work.
|
|
47
|
-
- Acceptance criteria can carry stable IDs and executable verification commands, and Yoke records per-criterion evidence before a story can be marked complete.
|
|
48
|
-
- Projects can configure an integrated completion command that proves the final system flow after all story-level gates pass.
|
|
49
|
-
|
|
50
|
-
### Changed
|
|
51
|
-
- New projects use strict criterion-proof requirements by default, while existing boolean-only criteria remain readable for compatibility.
|
|
52
|
-
- PRD authoring, schemas, loop guidance, generated configuration, and provider instructions now treat completion as an ephemeral readiness state: new requests create more stories instead of prematurely forcing a release boundary.
|
|
53
|
-
|
|
54
|
-
### Fixed
|
|
55
|
-
- Broad green test suites can no longer satisfy unrelated structured acceptance criteria: proof commands must use an approved test runner, contain the criterion ID, and avoid shell operators.
|
|
56
|
-
- Change-intake failures block safely; an independent coverage pass rejects omitted requested outcomes, crash recovery retains uncommitted requests, and concurrent PRD edits are detected instead of silently overwriting user work.
|
|
57
|
-
- Integrated completion checks prevent locally finished stories from masking broken cross-component flows such as authentication callbacks or purchase-to-entitlement activation.
|
|
58
|
-
|
|
59
|
-
### Security
|
|
60
|
-
- Updated the transitive development dependency `nanoid` to a non-vulnerable release so the complete CI dependency audit is clean.
|
|
61
|
-
- Restricted model-authored criterion proof to criterion-targeted test commands and blocked shell control operators and runner-prefix spoofing before host execution.
|
|
62
|
-
|
|
63
|
-
## 1.2.1 — 2026-08-02
|
|
64
|
-
|
|
65
|
-
### Fixed
|
|
66
|
-
- `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
|
|
67
|
-
- Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
|
|
68
|
-
|
|
69
|
-
## 1.2.0 — 2026-08-02
|
|
70
|
-
|
|
71
|
-
### Added
|
|
72
|
-
- Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
|
|
73
|
-
- A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
|
|
74
|
-
- Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
|
|
75
|
-
|
|
76
|
-
### Changed
|
|
77
|
-
- Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
|
|
78
|
-
- Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
|
|
79
|
-
|
|
80
|
-
### Fixed
|
|
81
|
-
- Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
|
|
82
|
-
- Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
|
|
83
|
-
|
|
84
|
-
## 1.1.0 — 2026-07-30
|
|
85
|
-
|
|
86
|
-
### Added
|
|
87
|
-
- Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
|
|
88
|
-
- Persisted default runner selection and `auto|critical` loop decision policies.
|
|
89
|
-
- Provider-neutral `yoke-workflow` skill for planning questions, approved-plan PRD handoff, autonomous story execution, and critical-decision resume.
|
|
90
|
-
- Structured critical-decision requests with `yoke loop decision` and `yoke loop answer`; answers are validated, committed under the configured human identity, and resume the same story.
|
|
91
|
-
- Approved `.yoke/plan.md` context in PRD drafting and a lint gate for unresolved planning placeholders.
|
|
92
|
-
|
|
93
|
-
### Fixed
|
|
94
|
-
- Retrofit and loop on/off now preserve timeout, decision, runner, and permission settings.
|
|
95
|
-
- Empty projects prefer the active agent host instead of silently installing Claude artifacts in Codex.
|
|
96
|
-
- Loop and PRD runner selection now prefer an explicit flag, then the configured runner, then the active host.
|
|
97
|
-
- Retrofit reports no longer label every provider as Claude Code.
|
|
98
|
-
- Critical-decision resumes retain isolation, review, runner, permissions, timeout, JSON, policy, and iteration settings instead of falling back to an unreviewed default run.
|
|
99
|
-
- Decision answers use an atomic owner-token lock and recoverable request journal, are checked against the active PRD story, bounded as untrusted data, and committed path-by-path so unrelated edits cannot enter the human-owned commit.
|
|
100
|
-
- Decision recovery now binds the exact selected answer to its commit, rolls back only its own interrupted context append, namespaces resume state per project/worktree, and serializes cleanup with loop startup.
|
|
101
|
-
- Active agent session markers now outrank globally configured provider home directories, and setup rejects partially invalid agent lists.
|
|
102
|
-
|
|
103
|
-
### Changed
|
|
104
|
-
- New setups enable the loop by default and choose `decisionPolicy: auto`; the wizard can select `critical` or disable the loop.
|
|
105
|
-
- Legacy `loop.onAmbiguity` and `--on-ambiguity` remain compatibility aliases.
|
|
106
|
-
|
|
107
|
-
## 1.0.0 — 2026-07-27
|
|
108
|
-
|
|
109
|
-
### Added
|
|
110
|
-
- Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
|
|
111
|
-
- Safe provider permission profiles and structured cross-provider telemetry.
|
|
112
|
-
- Schema-validated independent review verdicts with explicit self-review opt-in.
|
|
113
|
-
- Human-owned commit identity enforcement; AI co-author trailers default off.
|
|
114
|
-
- `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
|
|
115
|
-
- PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
|
|
116
|
-
- Reproducible cross-runner benchmark schema and matrix launcher.
|
|
117
|
-
|
|
118
|
-
### Changed
|
|
119
|
-
- Dangerous permission bypass is opt-in via `--unsafe`.
|
|
120
|
-
- Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
|
|
121
|
-
- Reviews no longer trust process exit code alone.
|
|
122
|
-
|
|
123
|
-
### Security
|
|
124
|
-
- Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
## 0.9.0 — 2026-07-22
|
|
128
|
-
|
|
129
|
-
### Added
|
|
130
|
-
- **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
|
|
131
|
-
A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
|
|
132
|
-
verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
|
|
133
|
-
`--isolate` mode. A red benchmark blocks the story
|
|
134
|
-
(`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
|
|
135
|
-
gate is configured, the implementer prompt names the budget command so agents keep hot
|
|
136
|
-
paths efficient and never "simplify away" an optimization without re-running the benchmark.
|
|
137
|
-
- **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
|
|
138
|
-
clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
|
|
139
|
-
project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
|
|
140
|
-
tests, the *why* of every optimization versioned so future agents don't clean fast code
|
|
141
|
-
back to slow.
|
|
142
|
-
- **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
|
|
143
|
-
numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
|
|
144
|
-
round — a criterion still needing a decision is not loop-ready.
|
|
145
|
-
|
|
146
|
-
## 0.8.0 — 2026-07-20
|
|
147
|
-
|
|
148
|
-
### Added
|
|
149
|
-
- **Live progress + ETA.** Story completions are now first-class events: the console shows
|
|
150
|
-
`✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
|
|
151
|
-
`yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
|
|
152
|
-
durations of stories completed **in this run** (current velocity) and falls back to the
|
|
153
|
-
persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
|
|
154
|
-
No data → no estimate, never an invented one.
|
|
155
|
-
- **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
|
|
156
|
-
prompt now always forbids asking questions (a loop run has nobody to answer). Default
|
|
157
|
-
`resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
|
|
158
|
-
loop never stops. Opt-in `abort`: the agent writes its open questions to
|
|
159
|
-
`.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
|
|
160
|
-
story would otherwise pass on pre-existing green tests), and blocks with the question as
|
|
161
|
-
the reason. Companion principle: clarifying questions belong in the planning round, before
|
|
162
|
-
the loop starts.
|
|
163
|
-
|
|
164
|
-
## 0.7.0 — 2026-07-17
|
|
165
|
-
|
|
166
|
-
### Added
|
|
167
|
-
- **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
|
|
168
|
-
version hint (npm/gh-style): a detached background refresher caches the registry's
|
|
169
|
-
latest at most once a day; when it is newer, a one-line stderr hint suggests
|
|
170
|
-
`yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
|
|
171
|
-
`--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
|
|
172
|
-
- **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
|
|
173
|
-
loop START only — never mid-run; the running process finishes on its version and
|
|
174
|
-
the upgrade applies from the next invocation. Deliberately NOT the default:
|
|
175
|
-
a gate harness must not change itself mid-project (determinism), and unreviewed
|
|
176
|
-
auto-installs are a supply-chain hazard.
|
|
177
|
-
|
|
178
|
-
## 0.6.0 — 2026-07-17
|
|
179
|
-
|
|
180
|
-
### Added
|
|
181
|
-
- **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
|
|
182
|
-
`.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
|
|
183
|
-
`yoke loop cleanup` kills exactly those recorded process trees — and only while no
|
|
184
|
-
live loop holds the lock. Background: without a scoped mechanism, users and agents
|
|
185
|
-
resorted to machine-wide pattern kills (every process matching
|
|
186
|
-
`dangerously-skip-permissions`), which took down *healthy* runners of other projects
|
|
187
|
-
mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
|
|
188
|
-
safe path. `.yoke/runner.pid` is gitignored by retrofit.
|
|
189
|
-
|
|
190
|
-
## 0.5.0 — 2026-07-17
|
|
191
|
-
|
|
192
|
-
Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
|
|
193
|
-
(orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
|
|
194
|
-
idle window):
|
|
195
|
-
|
|
196
|
-
### Fixed
|
|
197
|
-
- **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
|
|
198
|
-
Previously it killed only the spawned shell (`shell: true`), orphaning the actual
|
|
199
|
-
agent process — which kept writing to the worktree (dirty-tree blocks, failing
|
|
200
|
-
worktree removal) and kept burning API tokens. Observed in the field as ~10
|
|
201
|
-
zombie `claude.exe` per machine plus surviving dev servers.
|
|
202
|
-
- **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
|
|
203
|
-
the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
|
|
204
|
-
processes and killed them at exactly the idle timeout — while the user saw dead air.
|
|
205
|
-
The stream doubles as liveness; token usage is now reported on every run (not just
|
|
206
|
-
`--json` mode).
|
|
207
|
-
|
|
208
|
-
### Changed
|
|
209
|
-
- README: operating notes for driving the loop from inside an agent session
|
|
210
|
-
(background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
|
|
211
|
-
outer shell-tool timeouts killing a foreground `yoke loop run` were the third
|
|
212
|
-
observed "hang" pattern.
|
|
213
|
-
|
|
214
|
-
## 0.4.0 — 2026-07-17
|
|
215
|
-
|
|
216
|
-
### Added
|
|
217
|
-
- **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
|
|
218
|
-
scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
|
|
219
|
-
documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
|
|
220
|
-
bounded final messages (cuts output-token waste). Review prompts now ground verdicts
|
|
221
|
-
in observed evidence only and keep them brief.
|
|
222
|
-
|
|
223
|
-
### Fixed
|
|
224
|
-
- `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
|
|
225
|
-
`git add -A` story commit swept the pause control file into history in
|
|
226
|
-
un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
|
|
227
|
-
blocked the resume run — the loop locked itself out.
|
|
228
|
-
|
|
229
|
-
> Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
|
|
230
|
-
> was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
|
|
231
|
-
|
|
232
|
-
## 0.3.0 — 2026-07-10
|
|
233
|
-
|
|
234
|
-
### Added
|
|
235
|
-
- **Claude Code plugin packaging** — the repo is now its own plugin marketplace
|
|
236
|
-
(`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
|
|
237
|
-
then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
|
|
238
|
-
- **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
|
|
239
|
-
installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
|
|
240
|
-
daily-crawled extensions gallery.
|
|
241
|
-
- **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
|
|
242
|
-
quality) with a fixed fixture, pre-written objective tests, and committed result data.
|
|
243
|
-
- **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
|
|
244
|
-
sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
|
|
245
|
-
(design generation paired with Yoke's design verification gates).
|
|
246
|
-
- **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
|
|
247
|
-
stories across parallel workers (`needs` dependency field, claim files, merge queue,
|
|
248
|
-
heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
|
|
249
|
-
|
|
250
|
-
### Fixed
|
|
251
|
-
- Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
|
|
252
|
-
Windows, so the loop misreported an installed `gemini` as "not found on PATH"
|
|
253
|
-
(found by the new benchmark harness).
|
|
254
|
-
- Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
|
|
255
|
-
requires a value after `-p` and errored with "Not enough arguments following: p".
|
|
256
|
-
Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
|
|
257
|
-
(also found by the benchmark harness).
|
|
258
|
-
|
|
259
|
-
### Changed
|
|
260
|
-
- README: npm install is now the primary quickstart path; documented plugin/extension
|
|
261
|
-
installs and optional companions.
|
|
262
|
-
- npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
|
|
263
|
-
`docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
|
|
264
|
-
|
|
265
|
-
## 0.2.0 — 2026-07-09
|
|
266
|
-
|
|
267
|
-
- First npm release as `@hecer/yoke`.
|
|
268
|
-
- Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
|
|
269
|
-
token-usage + model-id reporting for the claude runner.
|
|
270
|
-
- `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
|
|
271
|
-
`yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
|
|
272
|
-
- Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
|
|
273
|
-
loop with worktree isolation, watchdog, single-flight lock, commit integrity.
|
|
43
|
+
|
|
44
|
+
### Added
|
|
45
|
+
- `yoke loop run --parallel=N` now executes dependency-ready, non-colliding stories through real provider subprocess workers, isolated worktrees, leased claims, and a FIFO integration queue with fresh integrated-system gates.
|
|
46
|
+
- Reference-driven quality declarations can collect screenshots, files, command output, or benchmark results and run a schema-validated blind critic with bounded repair rounds, elapsed-time limits, blocking or advisory policy, and retained proof.
|
|
47
|
+
- `--candidates=N` can fan out up to five isolated implementations, discard mechanically red candidates, select one green candidate through identity-blind pairwise comparison, and preserve selected/loser evidence before cleanup.
|
|
48
|
+
- Loop status now exposes dispatcher, worker, integrator, candidate lifecycle, worktree, queue, integration, reopen, quality-round, repair-budget, and trusted provider/model provenance data.
|
|
49
|
+
|
|
50
|
+
### Changed
|
|
51
|
+
- Provider subprocesses use explicit lifecycle contracts and incarnation-aware process records so worker cancellation and cleanup target only the process tree Yoke actually started.
|
|
52
|
+
- Parallel and candidate runs disable adaptive routing, honor story-level provider affinity, latch pause requests across the whole dispatcher, and rerun quality plus review after integration.
|
|
53
|
+
- `yoke loop cleanup` retains Yoke worktrees unless `--remove-worktrees` is explicit, while still reaping recorded orphan runners and stale locks safely.
|
|
54
|
+
|
|
55
|
+
### Fixed
|
|
56
|
+
- Expired claims, worker crashes, merge conflicts, pause races, and integration failures now release ownership deterministically, retain terminal proof, and reopen stories without leaking worktrees or marking false completion.
|
|
57
|
+
- Quality repair fails closed on malformed critic output, reference drift, provider/model provenance mismatch, candidate identity leakage, unavailable critics, exhausted limits, and mechanically red repairs.
|
|
58
|
+
- The watchdog resolves its TypeScript loader from both source and built npm layouts on Node 20+, and read-only Codex comparisons can run in disposable candidate worktrees without weakening normal repository checks.
|
|
59
|
+
|
|
60
|
+
### Security
|
|
61
|
+
- Blind comparison requests expose only opaque labels and digests while binding every verdict to the trusted judge provider, model, prompt, rubric, reference, and candidate provenance.
|
|
62
|
+
- Cleanup and cancellation use project-scoped leases, owner tokens, PID birth/incarnation checks, and recorded process handles rather than machine-wide process-name matching.
|
|
63
|
+
|
|
64
|
+
## 1.3.0 — 2026-08-09
|
|
65
|
+
|
|
66
|
+
### Added
|
|
67
|
+
- Teams can submit change requests at any time with `yoke change add` and inspect the append-only inbox with `yoke change status`; queued requests are planned at safe loop boundaries by Claude, Codex, or Gemini without interrupting active story work.
|
|
68
|
+
- Acceptance criteria can carry stable IDs and executable verification commands, and Yoke records per-criterion evidence before a story can be marked complete.
|
|
69
|
+
- Projects can configure an integrated completion command that proves the final system flow after all story-level gates pass.
|
|
70
|
+
|
|
71
|
+
### Changed
|
|
72
|
+
- New projects use strict criterion-proof requirements by default, while existing boolean-only criteria remain readable for compatibility.
|
|
73
|
+
- PRD authoring, schemas, loop guidance, generated configuration, and provider instructions now treat completion as an ephemeral readiness state: new requests create more stories instead of prematurely forcing a release boundary.
|
|
74
|
+
|
|
75
|
+
### Fixed
|
|
76
|
+
- Broad green test suites can no longer satisfy unrelated structured acceptance criteria: proof commands must use an approved test runner, contain the criterion ID, and avoid shell operators.
|
|
77
|
+
- Change-intake failures block safely; an independent coverage pass rejects omitted requested outcomes, crash recovery retains uncommitted requests, and concurrent PRD edits are detected instead of silently overwriting user work.
|
|
78
|
+
- Integrated completion checks prevent locally finished stories from masking broken cross-component flows such as authentication callbacks or purchase-to-entitlement activation.
|
|
79
|
+
|
|
80
|
+
### Security
|
|
81
|
+
- Updated the transitive development dependency `nanoid` to a non-vulnerable release so the complete CI dependency audit is clean.
|
|
82
|
+
- Restricted model-authored criterion proof to criterion-targeted test commands and blocked shell control operators and runner-prefix spoofing before host execution.
|
|
83
|
+
|
|
84
|
+
## 1.2.1 — 2026-08-02
|
|
85
|
+
|
|
86
|
+
### Fixed
|
|
87
|
+
- `yoke loop run` now executes every remaining planned story by default instead of stopping after an implicit 25-iteration batch. Use `--max=N` only when an intentional bounded batch is wanted.
|
|
88
|
+
- Unlimited runs remain unlimited after a critical-decision answer/resume cycle, while explicit caps remain preserved and accept positive integers only.
|
|
89
|
+
|
|
90
|
+
## 1.2.0 — 2026-08-02
|
|
91
|
+
|
|
92
|
+
### Added
|
|
93
|
+
- Opt-in adaptive model routing lets a strong parent orchestrate each bounded story while Yoke selects an available Claude, Codex, or Gemini worker by quality, speed, cost, or balanced strategy.
|
|
94
|
+
- A concurrency-safe local evidence registry learns from independent verification results without sharing mutable state between simultaneous Yoke processes.
|
|
95
|
+
- Routing telemetry, reproducible benchmark fixtures, and analysis tooling make worker selection, token use, timing, and gate outcomes auditable.
|
|
96
|
+
|
|
97
|
+
### Changed
|
|
98
|
+
- Setup keeps routing disabled unless explicitly enabled and validates provider/model worker pools before execution.
|
|
99
|
+
- Internal provider contract tests cover Claude Code, Codex CLI, and Gemini CLI through the shared adapter. Codex-only authenticated trials completed all 12 stories and 36 hidden checks; median routed runs used 11.0% fewer fresh input tokens, 49.5% fewer output tokens, 78.2% fewer reasoning tokens, and 33.8% less wall time than routing off.
|
|
100
|
+
|
|
101
|
+
### Fixed
|
|
102
|
+
- Reviewer prompts now permit their required verdict file while continuing to forbid project changes.
|
|
103
|
+
- Failed reviewer subprocesses retain bounded provider stderr in loop status, exposing authentication, quota, sandbox, and startup failures instead of a generic command error.
|
|
104
|
+
|
|
105
|
+
## 1.1.0 — 2026-07-30
|
|
106
|
+
|
|
107
|
+
### Added
|
|
108
|
+
- Shared five-question `yoke setup` wizard with provider-aware defaults for Claude, Codex, and Gemini.
|
|
109
|
+
- Persisted default runner selection and `auto|critical` loop decision policies.
|
|
110
|
+
- Provider-neutral `yoke-workflow` skill for planning questions, approved-plan PRD handoff, autonomous story execution, and critical-decision resume.
|
|
111
|
+
- Structured critical-decision requests with `yoke loop decision` and `yoke loop answer`; answers are validated, committed under the configured human identity, and resume the same story.
|
|
112
|
+
- Approved `.yoke/plan.md` context in PRD drafting and a lint gate for unresolved planning placeholders.
|
|
113
|
+
|
|
114
|
+
### Fixed
|
|
115
|
+
- Retrofit and loop on/off now preserve timeout, decision, runner, and permission settings.
|
|
116
|
+
- Empty projects prefer the active agent host instead of silently installing Claude artifacts in Codex.
|
|
117
|
+
- Loop and PRD runner selection now prefer an explicit flag, then the configured runner, then the active host.
|
|
118
|
+
- Retrofit reports no longer label every provider as Claude Code.
|
|
119
|
+
- Critical-decision resumes retain isolation, review, runner, permissions, timeout, JSON, policy, and iteration settings instead of falling back to an unreviewed default run.
|
|
120
|
+
- Decision answers use an atomic owner-token lock and recoverable request journal, are checked against the active PRD story, bounded as untrusted data, and committed path-by-path so unrelated edits cannot enter the human-owned commit.
|
|
121
|
+
- Decision recovery now binds the exact selected answer to its commit, rolls back only its own interrupted context append, namespaces resume state per project/worktree, and serializes cleanup with loop startup.
|
|
122
|
+
- Active agent session markers now outrank globally configured provider home directories, and setup rejects partially invalid agent lists.
|
|
123
|
+
|
|
124
|
+
### Changed
|
|
125
|
+
- New setups enable the loop by default and choose `decisionPolicy: auto`; the wizard can select `critical` or disable the loop.
|
|
126
|
+
- Legacy `loop.onAmbiguity` and `--on-ambiguity` remain compatibility aliases.
|
|
127
|
+
|
|
128
|
+
## 1.0.0 — 2026-07-27
|
|
129
|
+
|
|
130
|
+
### Added
|
|
131
|
+
- Native Codex skills, project config, hooks, reusable agents, and plugin metadata.
|
|
132
|
+
- Safe provider permission profiles and structured cross-provider telemetry.
|
|
133
|
+
- Schema-validated independent review verdicts with explicit self-review opt-in.
|
|
134
|
+
- Human-owned commit identity enforcement; AI co-author trailers default off.
|
|
135
|
+
- `yoke audit` dependency, secret, and sensitive-diff gate with versioned suppressions.
|
|
136
|
+
- PRD dependency graphs, collision areas, agent affinity, claims, FIFO merge queue, and bounded async dispatcher APIs.
|
|
137
|
+
- Reproducible cross-runner benchmark schema and matrix launcher.
|
|
138
|
+
|
|
139
|
+
### Changed
|
|
140
|
+
- Dangerous permission bypass is opt-in via `--unsafe`.
|
|
141
|
+
- Worktree cleanup is non-destructive unless `--remove-worktrees` is passed.
|
|
142
|
+
- Reviews no longer trust process exit code alone.
|
|
143
|
+
|
|
144
|
+
### Security
|
|
145
|
+
- Vitest upgraded to 4.1.10; the dependency tree reports zero known vulnerabilities.
|
|
146
|
+
|
|
147
|
+
|
|
148
|
+
## 0.9.0 — 2026-07-22
|
|
149
|
+
|
|
150
|
+
### Added
|
|
151
|
+
- **Performance budget gate** (`perf.command` in `.yoke/config.yaml`, optional `perf.retries`).
|
|
152
|
+
A benchmark command with the same contract as verify (exit 0 = within budget) runs **after
|
|
153
|
+
verify** on every story — new loop phase `perf`, `YOKE_STORY` exposed, worktree-aware in
|
|
154
|
+
`--isolate` mode. A red benchmark blocks the story
|
|
155
|
+
(`story S6 exceeded its performance budget: …`) no matter how clean the diff is. When the
|
|
156
|
+
gate is configured, the implementer prompt names the budget command so agents keep hot
|
|
157
|
+
paths efficient and never "simplify away" an optimization without re-running the benchmark.
|
|
158
|
+
- **`performance` canon skill** (28 skills now): efficiency as a measured requirement —
|
|
159
|
+
clean-by-default with the decision ladder (minimal-code → measurable acceptance criterion →
|
|
160
|
+
project perf gate), profile-first, optimize leaves not boundaries, benchmarks committed as
|
|
161
|
+
tests, the *why* of every optimization versioned so future agents don't clean fast code
|
|
162
|
+
back to slow.
|
|
163
|
+
- **`authoring-prd` guidance**: performance requirements belong in acceptance criteria as
|
|
164
|
+
numbers ("imports 1M rows in < 2s"), and every clarifying question belongs in the planning
|
|
165
|
+
round — a criterion still needing a decision is not loop-ready.
|
|
166
|
+
|
|
167
|
+
## 0.8.0 — 2026-07-20
|
|
168
|
+
|
|
169
|
+
### Added
|
|
170
|
+
- **Live progress + ETA.** Story completions are now first-class events: the console shows
|
|
171
|
+
`✓ S6 done in 4m28s — 20/45 (44%) · ~1h40m left`, every status (file, NDJSON stream,
|
|
172
|
+
`yoke loop status`) carries `percent` and an `eta` block. The estimate averages the
|
|
173
|
+
durations of stories completed **in this run** (current velocity) and falls back to the
|
|
174
|
+
persisted history of previous runs (`.yoke/story-durations.json`, last 50, gitignored).
|
|
175
|
+
No data → no estimate, never an invented one.
|
|
176
|
+
- **Ambiguity policy** (`loop.onAmbiguity` / `--on-ambiguity=<resolve|abort>`). The runner
|
|
177
|
+
prompt now always forbids asking questions (a loop run has nobody to answer). Default
|
|
178
|
+
`resolve`: the agent settles ambiguous criteria itself, states the interpretation, and the
|
|
179
|
+
loop never stops. Opt-in `abort`: the agent writes its open questions to
|
|
180
|
+
`.yoke/ambiguity.md` and stops; the loop consumes the file, skips verify (an unimplemented
|
|
181
|
+
story would otherwise pass on pre-existing green tests), and blocks with the question as
|
|
182
|
+
the reason. Companion principle: clarifying questions belong in the planning round, before
|
|
183
|
+
the loop starts.
|
|
184
|
+
|
|
185
|
+
## 0.7.0 — 2026-07-17
|
|
186
|
+
|
|
187
|
+
### Added
|
|
188
|
+
- **Update check + `yoke upgrade`.** Every CLI invocation ends with a non-blocking
|
|
189
|
+
version hint (npm/gh-style): a detached background refresher caches the registry's
|
|
190
|
+
latest at most once a day; when it is newer, a one-line stderr hint suggests
|
|
191
|
+
`yoke upgrade` (which runs `npm install -g @hecer/yoke@latest`). Silent in CI,
|
|
192
|
+
`--json` runs, non-TTY pipes, and under `YOKE_NO_UPDATE_CHECK=1`.
|
|
193
|
+
- **Opt-in auto-upgrade** (`update.auto: true` in `.yoke/config.yaml`): evaluated at
|
|
194
|
+
loop START only — never mid-run; the running process finishes on its version and
|
|
195
|
+
the upgrade applies from the next invocation. Deliberately NOT the default:
|
|
196
|
+
a gate harness must not change itself mid-project (determinism), and unreviewed
|
|
197
|
+
auto-installs are a supply-chain hazard.
|
|
198
|
+
|
|
199
|
+
## 0.6.0 — 2026-07-17
|
|
200
|
+
|
|
201
|
+
### Added
|
|
202
|
+
- **Project-scoped orphan reaping.** The watchdog now records its pids in the project's
|
|
203
|
+
`.yoke/runner.pid` (main dir and per-story worktrees; removed on clean exit), and
|
|
204
|
+
`yoke loop cleanup` kills exactly those recorded process trees — and only while no
|
|
205
|
+
live loop holds the lock. Background: without a scoped mechanism, users and agents
|
|
206
|
+
resorted to machine-wide pattern kills (every process matching
|
|
207
|
+
`dangerously-skip-permissions`), which took down *healthy* runners of other projects
|
|
208
|
+
mid-story and stalled their loops. Never kill by pattern; `yoke loop cleanup` is the
|
|
209
|
+
safe path. `.yoke/runner.pid` is gitignored by retrofit.
|
|
210
|
+
|
|
211
|
+
## 0.5.0 — 2026-07-17
|
|
212
|
+
|
|
213
|
+
Root-cause fixes for the two "yoke keeps hanging" failure modes observed in the field
|
|
214
|
+
(orphaned `claude.exe` runners piling up, healthy long stories dying at exactly the
|
|
215
|
+
idle window):
|
|
216
|
+
|
|
217
|
+
### Fixed
|
|
218
|
+
- **Watchdog now kills the whole process tree on Windows** (`taskkill /T /F`).
|
|
219
|
+
Previously it killed only the spawned shell (`shell: true`), orphaning the actual
|
|
220
|
+
agent process — which kept writing to the worktree (dirty-tree blocks, failing
|
|
221
|
+
worktree removal) and kept burning API tokens. Observed in the field as ~10
|
|
222
|
+
zombie `claude.exe` per machine plus surviving dev servers.
|
|
223
|
+
- **Claude runner always runs in stream-json mode.** Plain `-p` prints nothing until
|
|
224
|
+
the run finishes, so the idle watchdog mistook healthy >20-minute stories for dead
|
|
225
|
+
processes and killed them at exactly the idle timeout — while the user saw dead air.
|
|
226
|
+
The stream doubles as liveness; token usage is now reported on every run (not just
|
|
227
|
+
`--json` mode).
|
|
228
|
+
|
|
229
|
+
### Changed
|
|
230
|
+
- README: operating notes for driving the loop from inside an agent session
|
|
231
|
+
(background execution, small `--max` batches, `yoke loop cleanup` after interrupts) —
|
|
232
|
+
outer shell-tool timeouts killing a foreground `yoke loop run` were the third
|
|
233
|
+
observed "hang" pattern.
|
|
234
|
+
|
|
235
|
+
## 0.4.0 — 2026-07-17
|
|
236
|
+
|
|
237
|
+
### Added
|
|
238
|
+
- **Hardened runner prompts** — distilled agent-harness patterns for headless runs:
|
|
239
|
+
scope discipline (nothing beyond the story), no unsolicited summary/plan/analysis
|
|
240
|
+
documents, root-cause fixes instead of gate bypasses, faithful outcome reporting,
|
|
241
|
+
bounded final messages (cuts output-token waste). Review prompts now ground verdicts
|
|
242
|
+
in observed evidence only and keep them brief.
|
|
243
|
+
|
|
244
|
+
### Fixed
|
|
245
|
+
- `.yoke/loop.pause` is now gitignored by retrofit. Previously the loop's own
|
|
246
|
+
`git add -A` story commit swept the pause control file into history in
|
|
247
|
+
un-retrofitted targets; removing it dirtied the tree and the clean-tree gate
|
|
248
|
+
blocked the resume run — the loop locked itself out.
|
|
249
|
+
|
|
250
|
+
> Note: 0.3.0 was tagged and released on GitHub but never reached npm (2FA re-login
|
|
251
|
+
> was pending), so for npm users 0.4.0 is the first release with the 0.3.0 changes below.
|
|
252
|
+
|
|
253
|
+
## 0.3.0 — 2026-07-10
|
|
254
|
+
|
|
255
|
+
### Added
|
|
256
|
+
- **Claude Code plugin packaging** — the repo is now its own plugin marketplace
|
|
257
|
+
(`.claude-plugin/plugin.json` + `marketplace.json`): `/plugin marketplace add HECer/yoke`,
|
|
258
|
+
then `/plugin install yoke@yoke` installs the full canon under the `yoke:` skill namespace.
|
|
259
|
+
- **Gemini CLI extension manifest** (`gemini-extension.json` + `GEMINI-EXTENSION.md`) —
|
|
260
|
+
installable via `gemini extensions install https://github.com/HECer/yoke`, listed in the
|
|
261
|
+
daily-crawled extensions gallery.
|
|
262
|
+
- **Benchmark harness** (`bench/`) — reproducible cross-runner benchmark (tokens · speed ·
|
|
263
|
+
quality) with a fixed fixture, pre-written objective tests, and committed result data.
|
|
264
|
+
- **Companion tool docs** — `canon/tools/claude-mem.md` (persistent memory; interactive
|
|
265
|
+
sessions only, explicitly kept out of loop runs) and `canon/tools/ui-ux-pro-max.md`
|
|
266
|
+
(design generation paired with Yoke's design verification gates).
|
|
267
|
+
- **Multi-agent parallel loop design** — evaluation + phased design for distributing PRD
|
|
268
|
+
stories across parallel workers (`needs` dependency field, claim files, merge queue,
|
|
269
|
+
heterogeneous cross-agent dispatch): `docs/superpowers/specs/2026-07-10-multi-agent-parallel-loop-design.md`.
|
|
270
|
+
|
|
271
|
+
### Fixed
|
|
272
|
+
- Agent-availability probe timeout raised 5s → 20s: Gemini CLI cold-starts in ~6s on
|
|
273
|
+
Windows, so the loop misreported an installed `gemini` as "not found on PATH"
|
|
274
|
+
(found by the new benchmark harness).
|
|
275
|
+
- Gemini runner invocation: dropped the bare `-p` flag — current Gemini CLI (0.33+)
|
|
276
|
+
requires a value after `-p` and errored with "Not enough arguments following: p".
|
|
277
|
+
Piped stdin selects headless mode by itself, so the runner now passes only `--yolo`
|
|
278
|
+
(also found by the benchmark harness).
|
|
279
|
+
|
|
280
|
+
### Changed
|
|
281
|
+
- README: npm install is now the primary quickstart path; documented plugin/extension
|
|
282
|
+
installs and optional companions.
|
|
283
|
+
- npm package now ships `CHANGELOG.md`, `bench/` (harness + result data), and
|
|
284
|
+
`docs/superpowers/` (all specs and plans, including the multi-agent parallel loop design).
|
|
285
|
+
|
|
286
|
+
## 0.2.0 — 2026-07-09
|
|
287
|
+
|
|
288
|
+
- First npm release as `@hecer/yoke`.
|
|
289
|
+
- Hyperflow integration surface: `yoke loop run --json` NDJSON stream, pause signal,
|
|
290
|
+
token-usage + model-id reporting for the claude runner.
|
|
291
|
+
- `yoke new` greenfield bootstrap, `yoke prd draft`, cross-model `yoke review`,
|
|
292
|
+
`yoke flow-smoke` browser gate with proof artifacts, `yoke design-scan`.
|
|
293
|
+
- Retrofit planners for Claude Code, Codex CLI, Gemini CLI; canon of 26 skills;
|
|
294
|
+
loop with worktree isolation, watchdog, single-flight lock, commit integrity.
|