@hecer/yoke 1.9.0 → 1.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (153) hide show
  1. package/.claude-plugin/plugin.json +13 -13
  2. package/.codex-plugin/plugin.json +7 -7
  3. package/CHANGELOG.md +398 -358
  4. package/README.md +915 -913
  5. package/TODOS.md +5 -5
  6. package/agents/docs.toml +6 -6
  7. package/agents/implementer.toml +6 -6
  8. package/agents/reviewer.toml +6 -6
  9. package/agents/security.toml +6 -6
  10. package/bench/README.md +86 -86
  11. package/bench/RESULTS.md +35 -35
  12. package/bench/output-compaction.mjs +65 -65
  13. package/bench/result-schema.mjs +12 -12
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -15
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
  17. package/bench/run-matrix.mjs +26 -26
  18. package/bench/run.mjs +106 -106
  19. package/canon/AGENTS.md +30 -30
  20. package/canon/context/DECISIONS.md +4 -4
  21. package/canon/context/GLOSSARY.md +11 -11
  22. package/canon/context/KNOWLEDGE.md +4 -4
  23. package/canon/context/PROJECT.md +15 -15
  24. package/canon/loop/loop-spec.md +65 -65
  25. package/canon/loop/prd.schema.md +46 -40
  26. package/canon/manifest.yaml +59 -59
  27. package/canon/policy/gates.md +7 -7
  28. package/canon/policy/roles.md +9 -9
  29. package/canon/skills/ATTRIBUTION.md +99 -99
  30. package/canon/skills/authoring-prd/SKILL.md +56 -56
  31. package/canon/skills/brainstorming/SKILL.md +164 -164
  32. package/canon/skills/codebase-design/DEEPENING.md +15 -15
  33. package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
  34. package/canon/skills/codebase-design/SKILL.md +39 -39
  35. package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
  36. package/canon/skills/document-release/SKILL.md +302 -302
  37. package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
  38. package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
  39. package/canon/skills/domain-modeling/SKILL.md +35 -35
  40. package/canon/skills/executing-plans/SKILL.md +70 -70
  41. package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
  42. package/canon/skills/health/SKILL.md +177 -177
  43. package/canon/skills/maintaining-context/SKILL.md +34 -34
  44. package/canon/skills/minimal-code/SKILL.md +21 -21
  45. package/canon/skills/no-ai-slop/SKILL.md +103 -103
  46. package/canon/skills/no-ai-slop/eval.md +43 -43
  47. package/canon/skills/plan-ceo-review/SKILL.md +541 -541
  48. package/canon/skills/plan-eng-review/SKILL.md +362 -362
  49. package/canon/skills/receiving-code-review/SKILL.md +213 -213
  50. package/canon/skills/requesting-code-review/SKILL.md +105 -105
  51. package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
  52. package/canon/skills/retro/SKILL.md +397 -397
  53. package/canon/skills/review/SKILL.md +246 -246
  54. package/canon/skills/ship/SKILL.md +691 -691
  55. package/canon/skills/subagent-driven-development/SKILL.md +277 -277
  56. package/canon/skills/systematic-debugging/SKILL.md +296 -296
  57. package/canon/skills/tdd/SKILL.md +371 -371
  58. package/canon/skills/unslop-ui/SKILL.md +34 -34
  59. package/canon/skills/using-git-worktrees/SKILL.md +218 -218
  60. package/canon/skills/verification-before-completion/SKILL.md +139 -139
  61. package/canon/skills/visual-verification/SKILL.md +54 -54
  62. package/canon/skills/workflow/SKILL.md +22 -22
  63. package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
  64. package/canon/skills/writing-for-agents/SKILL.md +42 -42
  65. package/canon/skills/writing-plans/SKILL.md +152 -152
  66. package/canon/skills/writing-skills/SKILL.md +655 -655
  67. package/canon/skills/yoke-retrofit/SKILL.md +26 -26
  68. package/canon/skills/yoke-workflow/SKILL.md +20 -20
  69. package/canon/tools/codex-rtk-hook.mjs +35 -35
  70. package/canon/tools/gemini-rtk-hook.mjs +25 -25
  71. package/canon/tools/graphify.md +3 -3
  72. package/canon/tools/playwright-mcp.md +3 -3
  73. package/canon/tools/rtk.md +7 -7
  74. package/canon/tools/serena.md +6 -6
  75. package/dist/agents/contracts.js +1 -1
  76. package/dist/agents/host.js +4 -0
  77. package/dist/agents/process-incarnation.js +1 -1
  78. package/dist/agents/process.js +74 -6
  79. package/dist/agents/providers.js +13 -0
  80. package/dist/agents/supervision.js +153 -0
  81. package/dist/agents/telemetry.js +33 -0
  82. package/dist/agents/windows-launch.js +80 -0
  83. package/dist/canon/manifest.js +1 -1
  84. package/dist/change/inbox.js +21 -5
  85. package/dist/cli.js +19 -10
  86. package/dist/dashboard/discovery.js +73 -0
  87. package/dist/dashboard/page.js +122 -28
  88. package/dist/dashboard/panels.js +91 -15
  89. package/dist/goals/command.js +4 -2
  90. package/dist/loop/claims.js +1 -1
  91. package/dist/loop/decision.js +2 -2
  92. package/dist/loop/git.js +12 -4
  93. package/dist/loop/loop.js +8 -4
  94. package/dist/loop/parallel-adapters.js +2 -3
  95. package/dist/loop/parallel-command.js +5 -0
  96. package/dist/loop/prd.js +3 -1
  97. package/dist/loop/reporter.js +4 -1
  98. package/dist/loop/run-command.js +11 -2
  99. package/dist/loop/runner.js +22 -26
  100. package/dist/loop/watchdog.js +87 -11
  101. package/dist/loop/worker.js +5 -3
  102. package/dist/prd/assess.js +145 -0
  103. package/dist/prd/command.js +76 -38
  104. package/dist/quality/types.js +1 -1
  105. package/dist/retrofit/config.js +11 -0
  106. package/dist/retrofit/plan.js +2 -0
  107. package/dist/retrofit/planners/claude.js +14 -14
  108. package/dist/retrofit/planners/qwen.js +73 -0
  109. package/dist/retrofit/preserve.js +2 -2
  110. package/dist/retrofit/skill-actions.js +1 -0
  111. package/dist/review/command.js +1 -1
  112. package/dist/routing/assessment.js +1 -1
  113. package/dist/routing/capability.js +25 -13
  114. package/dist/routing/contracts.js +60 -0
  115. package/dist/routing/planning.js +12 -0
  116. package/dist/routing/router.js +51 -16
  117. package/dist/setup/command.js +11 -3
  118. package/docs/BATCH-PLANNING-VALIDATION.md +67 -0
  119. package/docs/CAPABILITY-ROUTING.md +78 -50
  120. package/docs/DASHBOARD-EVOLUTION.md +33 -0
  121. package/docs/MIGRATING-TO-1.0.md +33 -33
  122. package/docs/MIGRATING-TO-1.1.md +27 -27
  123. package/docs/MIGRATING-TO-1.4.md +70 -70
  124. package/docs/PRODUCT-DIRECTION-2026-09-05.md +218 -200
  125. package/docs/PUBLISHING.md +114 -114
  126. package/docs/VERIFIED-PROJECTS-VALIDATION.md +29 -29
  127. package/docs/VERIFIED-PROJECTS.md +167 -167
  128. package/docs/WINDOWS-RUNNER-VALIDATION.md +104 -0
  129. package/docs/assets/yoke-logo.png +0 -0
  130. package/docs/community-outreach-2026-08-20.md +85 -0
  131. package/docs/launch-copy-2026-08-21.md +193 -0
  132. package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
  133. package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
  134. package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
  135. package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
  136. package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
  137. package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
  138. package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
  139. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
  140. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
  141. package/docs/superpowers/plans/2026-09-05-verified-projects.md +83 -83
  142. package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
  143. package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
  144. package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
  145. package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
  146. package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
  147. package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
  148. package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
  149. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
  150. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
  151. package/gemini-extension.json +6 -6
  152. package/hooks/hooks.json +19 -19
  153. package/package.json +87 -87
@@ -1,167 +1,167 @@
1
- # Verified projects, goals and measured execution
2
-
3
- Available in Yoke 1.7.0. These additions do not require a cloud account or replace your provider CLI. Model availability, reasoning controls, structured output and telemetry depend on the selected provider; unknown usage is never a measured zero. Live authenticated provider comparison and calibrated time/cost benchmarks remain separate validation work.
4
-
5
- ## Check an existing project
6
-
7
- ```sh
8
- yoke check .
9
- yoke check . --json
10
- yoke check . --requirement="Guest checkout completes"
11
- ```
12
-
13
- Checks run without retrofit. Exit codes: `0` passed, `1` failed, `2` unverified or unavailable. A free-text requirement stays unverified until you map its behavior to an executable acceptance contract. Passing a project suite does not establish arbitrary product correctness.
14
-
15
- Create `.yoke/acceptance.yaml` with tests that inspect application behavior:
16
-
17
- ```yaml
18
- version: 1
19
- protected:
20
- - tests/acceptance/checkout.test.mjs
21
- - tests/acceptance/helpers.mjs
22
- criteria:
23
- - id: guest-checkout
24
- text: A guest can complete checkout and receives one order confirmation.
25
- commands:
26
- - node --test tests/acceptance/checkout.test.mjs
27
- ```
28
-
29
- List every local test helper, fixture and configuration file whose modification could weaken this contract. Paths must resolve inside the project. `yoke check . --protect` pins the manifest, listed files and existing package manifests/lockfiles outside the worker checkout. Autonomous goals require explicit protected test infrastructure and establish this pin automatically. Protection is change detection within the local execution model, not an OS security boundary against a hostile process using your account. Review the contract before running a goal.
30
-
31
- After intentionally editing protected tests, use `yoke check . --protect --refresh` to approve the new baseline. Ordinary checks and retries never refresh it. Source, acceptance and configuration identity are checked before and after verification; changed inputs invalidate evidence. Gitlinks/submodule directories currently fail closed because recursive identity is not implemented. Evidence is stored in `.yoke/checks/<id>.json`; it describes that checked snapshot and is not a permanent success certificate.
32
-
33
- ## Durable goals across providers
34
-
35
- ```sh
36
- yoke goal set . --objective="Finish guest checkout" --attempts=3 --minutes=30
37
- yoke goal run . --runner=codex
38
- yoke goal resume . --runner=claude --model=<installed-model-id>
39
- yoke goal resume . --runner=gemini --model=<installed-model-id>
40
- yoke goal status .
41
- yoke goal handoff .
42
- yoke goal pause .
43
- ```
44
-
45
- Goals keep objective, attempts, failure context and check IDs in `.yoke/goal.json`. They share the story-loop lock. Each implementation attempt is followed by independent executable acceptance. A completed goal is checked again on a later run. Failed work stays in the project; goal execution itself does not commit or publish it. Goal handoff is readable context for native agent goal facilities; Yoke does not invent or call an undocumented native goal API.
46
-
47
- `--minutes` limits cumulative agent execution time. Verification is measured separately and commands retain their verification timeout. `--tokens=N` is a checkpoint budget: it prevents another attempt when measured consumption is exhausted or unknown, but cannot promise a hard token cap within a single provider call. An interrupted attempt is persisted before dispatch; recovery accounts for it conservatively and reconciles recorded provider processes before continuing. Unknown process ownership blocks execution. Pause takes effect at an attempt boundary; it does not instantly kill an in-flight agent.
48
-
49
- Explicitly extend total budgets without deleting history:
50
-
51
- ```sh
52
- yoke goal budget . --attempts=5 --minutes=60
53
- yoke goal budget . --tokens=200000
54
- # Only if you deliberately want to remove the checkpoint token limit:
55
- yoke goal budget . --clear-token-budget
56
- ```
57
-
58
- Unknown historical token usage cannot be made known by increasing its limit. Inspect interrupted work and keep that limitation visible.
59
-
60
- ## Recover an isolated story
61
-
62
- Failed or paused isolated story worktrees are retained. Resume the same story with:
63
-
64
- ```sh
65
- yoke loop run . --isolate --resume-worktree --parallel=1 --candidates=1
66
- ```
67
-
68
- Yoke validates the registered worktree, repository, original target commit and PRD digest. Changed target/PRD state is refused; inspect retained edits and reconcile deliberately. This flag is for serial isolated recovery; parallel workers retain their existing coordinator lifecycle. `yoke loop cleanup` remains an explicit cleanup operation; inspect its removal flags before discarding unfinished work. Existing projects should rerun retrofit to add ignore entries for checks, events and goals.
69
-
70
- ## Spend fewer model calls
71
-
72
- Add explicit project rules to `.yoke/config.yaml`:
73
-
74
- ```yaml
75
- routing:
76
- enabled: true
77
- strategy: cost
78
- maxCandidates: 2
79
- workers:
80
- - id: fast
81
- agent: claude
82
- model: <your-fast-model-id>
83
- costTier: low
84
- capabilities: [docs, tests]
85
- - id: strong
86
- agent: codex
87
- model: <your-strong-model-id>
88
- costTier: high
89
- capabilities: [implementation]
90
- rules:
91
- - area: docs
92
- worker: fast
93
- escalateTo: strong
94
- ```
95
-
96
- The first matching rule bypasses the routing controller. Its optional area and story ID selectors both have to match when supplied. Independent gate failure escalates the next attempt, including after loop restart; an unavailable or unknown worker falls back to the parent. Without an explicit escalation target the parent handles the failed rule. Other stories retain adaptive routing. Gate-driven routing observations are local and bounded when read. These are measured outcomes, not guarantees that a named inexpensive model can handle every task.
97
-
98
- For deterministic operations, configure exact executable arguments:
99
-
100
- ```yaml
101
- actions:
102
- - storyId: regenerate-types
103
- file: node
104
- args: [scripts/generate-types.mjs]
105
- timeoutMs: 60000
106
- ```
107
-
108
- Actions use no model, run without a shell, have bounded output/time and still pass the ordinary story gates before commit. They currently require serial execution with one candidate. On Windows use an executable such as `node` and its script path; shell-only `.cmd` wrappers are not automatically enabled. Commands are project-authored configuration, never free-form model-generated commands. Projects consisting entirely of configured actions do not need a model CLI unless they enable a model-based review or planning workflow.
109
-
110
- Implementation/review prompts now select task-relevant project context deterministically within a 6,000-character budget. A stable project/glossary prefix precedes ranked historical references with source paths and content hashes. Excerpts point back to complete files. This does not claim a fixed token count or guaranteed provider cache hit. Mandatory verification always reruns; no broad result cache or selective-test bypass was introduced.
111
-
112
- ## Parallelism and time estimates
113
-
114
- Stories can declare relative file/directory scopes:
115
-
116
- ```yaml
117
- - id: api-types
118
- title: Update API types
119
- priority: 1
120
- area: api
121
- writes: [src/api, tests/api]
122
- needs: []
123
- acceptance: ["Replace with executable structured criteria in strict projects"]
124
- passes: false
125
- ```
126
-
127
- Overlapping declared scopes cannot run simultaneously, including while integration is pending. Dependencies, areas, explicit priorities and concurrency limits also apply. Equal-priority work that unlocks a longer dependency chain is scheduled first. Declarations are advisory scheduling input, not filesystem write enforcement; absent declarations preserve previous behavior. Final integrated gates remain mandatory.
128
-
129
- Versioned local events record status, phase duration, attempts and available usage. Retention is bounded to 1,000 events. Estimates combine historical and current durations, preserve failed-attempt cost and expose sample counts and empirical ranges. Schedule estimates simulate dependency, scope and concurrency constraints. Predictions and later errors are attached to completed attempt events so accuracy can be evaluated. Ranges describe observed data; they are not calibrated probabilities or exact deadlines. No history means no justified time estimate.
130
-
131
- ## Local project dashboard
132
-
133
- ### Execution defaults in 1.8.0
134
-
135
- New setups enable routing, `loop.parallel: auto` and `loop.isolate: true`. Existing explicit settings remain authoritative. At execution time, automatic routing uses configured profiles; without profiles it keeps the selected parent. Explicit `--routing` without profiles still reports a configuration error. Routing rules bypass the controller and can escalate following failed independent gates, including across worktrees and restarts. Routing now also runs asynchronously inside parallel workers. An explicit task provider affinity takes precedence over routing.
136
-
137
- Automatic parallelism allows at most three Yoke workers when every pending task declares nonempty write scopes. Dependencies and overlapping scopes still constrain dispatch. Unknown scopes, configured tool actions and worktree recovery select serial execution. Use `--parallel=N`, `--parallel=auto`, `--no-routing` or `--no-isolate` to override defaults. Serial worktrees retain failed work and require deliberate recovery; the default does not discard an existing recovery tree. Quality repair budgets and opt-in competing candidates retain their existing policies.
138
-
139
- Integration retains an execution slot until its candidate lands. Yoke loop runners disable native delegation so it cannot multiply the default worker budget: Codex disables multi_agent; Claude disallows Agent, Task, TeamCreate and SendMessage; Gemini uses a separate temporary system-settings copy that disables experimental agents and its always-on investigator/help overrides. Existing Gemini system policy and system-default paths are preserved; unreadable or malformed policy blocks launch. Original settings are never overwritten. The temporary copy is removed on normal exit; forced process termination may leave a private temporary directory. Explicit competing candidate counts remain a separate opt-in workload.
140
-
141
- ### Dashboard views
142
-
143
- - **Now:** reported task and worker activity, requested provider/model, elapsed worker time, phase, integration, blockers, status age, backlog and empirical remaining-time ranges. The view refreshes every five seconds while visible and not being operated with keyboard focus. A stale status is marked rather than asserted to be live.
144
- - **Usage & time:** last 24 hours, 7/30/90/365 days or custom dates, grouped by UTC day, Monday-based week or month. Displays reported input/output and cache categories, cost coverage, consumption charts, per-model time buckets, per-task usage and summed phase/call durations. The overview also compares projects for the same period.
145
- - **Results:** recorded acceptances, ended attempts, explicitly successful outcomes, repair phases, rule-driven escalations and recorded tokens/time per acceptance. Saved acceptance evidence and goal attempts remain available; their timestamps may fall outside the statistics period.
146
-
147
- The short activity list still retains at most 1,000 events. Compact measurements now also persist under `.yoke/history/YYYY-MM-DD/` independently of that retention and across runs. They contain identifiers, model/provider evidence and measurements, not prompts or full status snapshots. Runtime history is excluded from Yoke commits and added to new project ignore rules. Available reviewer, quality critic and repair usage is recorded separately; absent usage is counted as unknown. Recent and archived records are deduplicated by event ID.
148
-
149
- Dates and buckets use UTC. The custom end date is inclusive; the API uses an exclusive upper timestamp. Usage belongs to the time the provider reports it, and durations to their end time. Tokens per elapsed minute divide recorded input/output by the full selected interval. Tokens per call minute use summed reported call duration, which can overlap across workers. Neither is measured generation speed. Cache categories are shown separately without adding them to input again.
150
-
151
- Historical activity that was never recorded or already expired cannot be reconstructed. Queue/human waiting time that lacks measurements remains unknown; phase and attempt durations are summed worker time, not automatically elapsed project time. Queries allow at most 366 days and bounded reads (8 MiB per shard, 32 MiB overall, 50,000 records); skipped, malformed or oversized history is reported as incomplete. History currently requires local disk retention rather than automatic monthly compaction. Unknown model identity is never replaced by a requested model, and missing charges are not estimated from token counts.
152
-
153
- ```sh
154
- yoke projects add /path/to/project
155
- yoke projects list
156
- yoke dashboard .
157
- yoke dashboard --no-register --port=4100
158
- yoke projects remove <opaque-project-id>
159
- ```
160
-
161
- Open the printed loopback URL. The dashboard displays project goals, backlog, worker/status data, evidence, blockers, events, usage availability and schedule ranges. It can request a goal pause through the same service as the CLI. Start, resume and budget changes remain explicit CLI operations. Removing registration only removes the registry reference.
162
-
163
- The HTTP service binds to `127.0.0.1`, validates Host/Origin, requires a per-session token for pause, sends a restrictive CSP and renders project strings as text. It does not serve arbitrary local files or allow remote registration. Missing/corrupt/oversized project data is shown as unavailable. Shared project registration and protected acceptance live under `YOKE_STATE_DIR` or `~/.yoke/state`; existing routing history retains its `YOKE_REGISTRY_DIR` location. The dashboard is local, not a hosted multi-user service.
164
-
165
- ## Remaining validation and product work
166
-
167
- Live Codex/Claude/Gemini comparisons, measured competitive development-time/cost studies and estimate calibration require representative real projects and authenticated runs. Equal adapter contracts do not imply equal model capability. Cloud/team access, automatic selective test reuse, web-based start/resume controls and commercial rollout are not included in this local foundation. The complete saved direction is in [PRODUCT-DIRECTION-2026-09-05.md](PRODUCT-DIRECTION-2026-09-05.md).
1
+ # Verified projects, goals and measured execution
2
+
3
+ Available in Yoke 1.7.0. These additions do not require a cloud account or replace your provider CLI. Model availability, reasoning controls, structured output and telemetry depend on the selected provider; unknown usage is never a measured zero. Live authenticated provider comparison and calibrated time/cost benchmarks remain separate validation work.
4
+
5
+ ## Check an existing project
6
+
7
+ ```sh
8
+ yoke check .
9
+ yoke check . --json
10
+ yoke check . --requirement="Guest checkout completes"
11
+ ```
12
+
13
+ Checks run without retrofit. Exit codes: `0` passed, `1` failed, `2` unverified or unavailable. A free-text requirement stays unverified until you map its behavior to an executable acceptance contract. Passing a project suite does not establish arbitrary product correctness.
14
+
15
+ Create `.yoke/acceptance.yaml` with tests that inspect application behavior:
16
+
17
+ ```yaml
18
+ version: 1
19
+ protected:
20
+ - tests/acceptance/checkout.test.mjs
21
+ - tests/acceptance/helpers.mjs
22
+ criteria:
23
+ - id: guest-checkout
24
+ text: A guest can complete checkout and receives one order confirmation.
25
+ commands:
26
+ - node --test tests/acceptance/checkout.test.mjs
27
+ ```
28
+
29
+ List every local test helper, fixture and configuration file whose modification could weaken this contract. Paths must resolve inside the project. `yoke check . --protect` pins the manifest, listed files and existing package manifests/lockfiles outside the worker checkout. Autonomous goals require explicit protected test infrastructure and establish this pin automatically. Protection is change detection within the local execution model, not an OS security boundary against a hostile process using your account. Review the contract before running a goal.
30
+
31
+ After intentionally editing protected tests, use `yoke check . --protect --refresh` to approve the new baseline. Ordinary checks and retries never refresh it. Source, acceptance and configuration identity are checked before and after verification; changed inputs invalidate evidence. Gitlinks/submodule directories currently fail closed because recursive identity is not implemented. Evidence is stored in `.yoke/checks/<id>.json`; it describes that checked snapshot and is not a permanent success certificate.
32
+
33
+ ## Durable goals across providers
34
+
35
+ ```sh
36
+ yoke goal set . --objective="Finish guest checkout" --attempts=3 --minutes=30
37
+ yoke goal run . --runner=codex
38
+ yoke goal resume . --runner=claude --model=<installed-model-id>
39
+ yoke goal resume . --runner=gemini --model=<installed-model-id>
40
+ yoke goal status .
41
+ yoke goal handoff .
42
+ yoke goal pause .
43
+ ```
44
+
45
+ Goals keep objective, attempts, failure context and check IDs in `.yoke/goal.json`. They share the story-loop lock. Each implementation attempt is followed by independent executable acceptance. A completed goal is checked again on a later run. Failed work stays in the project; goal execution itself does not commit or publish it. Goal handoff is readable context for native agent goal facilities; Yoke does not invent or call an undocumented native goal API.
46
+
47
+ `--minutes` limits cumulative agent execution time. Verification is measured separately and commands retain their verification timeout. `--tokens=N` is a checkpoint budget: it prevents another attempt when measured consumption is exhausted or unknown, but cannot promise a hard token cap within a single provider call. An interrupted attempt is persisted before dispatch; recovery accounts for it conservatively and reconciles recorded provider processes before continuing. Unknown process ownership blocks execution. Pause takes effect at an attempt boundary; it does not instantly kill an in-flight agent.
48
+
49
+ Explicitly extend total budgets without deleting history:
50
+
51
+ ```sh
52
+ yoke goal budget . --attempts=5 --minutes=60
53
+ yoke goal budget . --tokens=200000
54
+ # Only if you deliberately want to remove the checkpoint token limit:
55
+ yoke goal budget . --clear-token-budget
56
+ ```
57
+
58
+ Unknown historical token usage cannot be made known by increasing its limit. Inspect interrupted work and keep that limitation visible.
59
+
60
+ ## Recover an isolated story
61
+
62
+ Failed or paused isolated story worktrees are retained. Resume the same story with:
63
+
64
+ ```sh
65
+ yoke loop run . --isolate --resume-worktree --parallel=1 --candidates=1
66
+ ```
67
+
68
+ Yoke validates the registered worktree, repository, original target commit and PRD digest. Changed target/PRD state is refused; inspect retained edits and reconcile deliberately. This flag is for serial isolated recovery; parallel workers retain their existing coordinator lifecycle. `yoke loop cleanup` remains an explicit cleanup operation; inspect its removal flags before discarding unfinished work. Existing projects should rerun retrofit to add ignore entries for checks, events and goals.
69
+
70
+ ## Spend fewer model calls
71
+
72
+ Add explicit project rules to `.yoke/config.yaml`:
73
+
74
+ ```yaml
75
+ routing:
76
+ enabled: true
77
+ strategy: cost
78
+ maxCandidates: 2
79
+ workers:
80
+ - id: fast
81
+ agent: claude
82
+ model: <your-fast-model-id>
83
+ costTier: low
84
+ capabilities: [docs, tests]
85
+ - id: strong
86
+ agent: codex
87
+ model: <your-strong-model-id>
88
+ costTier: high
89
+ capabilities: [implementation]
90
+ rules:
91
+ - area: docs
92
+ worker: fast
93
+ escalateTo: strong
94
+ ```
95
+
96
+ The first matching rule bypasses the routing controller. Its optional area and story ID selectors both have to match when supplied. Independent gate failure escalates the next attempt, including after loop restart; an unavailable or unknown worker falls back to the parent. Without an explicit escalation target the parent handles the failed rule. Other stories retain adaptive routing. Gate-driven routing observations are local and bounded when read. These are measured outcomes, not guarantees that a named inexpensive model can handle every task.
97
+
98
+ For deterministic operations, configure exact executable arguments:
99
+
100
+ ```yaml
101
+ actions:
102
+ - storyId: regenerate-types
103
+ file: node
104
+ args: [scripts/generate-types.mjs]
105
+ timeoutMs: 60000
106
+ ```
107
+
108
+ Actions use no model, run without a shell, have bounded output/time and still pass the ordinary story gates before commit. They currently require serial execution with one candidate. On Windows use an executable such as `node` and its script path; shell-only `.cmd` wrappers are not automatically enabled. Commands are project-authored configuration, never free-form model-generated commands. Projects consisting entirely of configured actions do not need a model CLI unless they enable a model-based review or planning workflow.
109
+
110
+ Implementation/review prompts now select task-relevant project context deterministically within a 6,000-character budget. A stable project/glossary prefix precedes ranked historical references with source paths and content hashes. Excerpts point back to complete files. This does not claim a fixed token count or guaranteed provider cache hit. Mandatory verification always reruns; no broad result cache or selective-test bypass was introduced.
111
+
112
+ ## Parallelism and time estimates
113
+
114
+ Stories can declare relative file/directory scopes:
115
+
116
+ ```yaml
117
+ - id: api-types
118
+ title: Update API types
119
+ priority: 1
120
+ area: api
121
+ writes: [src/api, tests/api]
122
+ needs: []
123
+ acceptance: ["Replace with executable structured criteria in strict projects"]
124
+ passes: false
125
+ ```
126
+
127
+ Overlapping declared scopes cannot run simultaneously, including while integration is pending. Dependencies, areas, explicit priorities and concurrency limits also apply. Equal-priority work that unlocks a longer dependency chain is scheduled first. Declarations are advisory scheduling input, not filesystem write enforcement; absent declarations preserve previous behavior. Final integrated gates remain mandatory.
128
+
129
+ Versioned local events record status, phase duration, attempts and available usage. Retention is bounded to 1,000 events. Estimates combine historical and current durations, preserve failed-attempt cost and expose sample counts and empirical ranges. Schedule estimates simulate dependency, scope and concurrency constraints. Predictions and later errors are attached to completed attempt events so accuracy can be evaluated. Ranges describe observed data; they are not calibrated probabilities or exact deadlines. No history means no justified time estimate.
130
+
131
+ ## Local project dashboard
132
+
133
+ ### Execution defaults in 1.8.0
134
+
135
+ New setups enable routing, `loop.parallel: auto` and `loop.isolate: true`. Existing explicit settings remain authoritative. At execution time, automatic routing uses configured profiles; without profiles it keeps the selected parent. Explicit `--routing` without profiles still reports a configuration error. Routing rules bypass the controller and can escalate following failed independent gates, including across worktrees and restarts. Routing now also runs asynchronously inside parallel workers. An explicit task provider affinity takes precedence over routing.
136
+
137
+ Automatic parallelism allows at most three Yoke workers when every pending task declares nonempty write scopes. Dependencies and overlapping scopes still constrain dispatch. Unknown scopes, configured tool actions and worktree recovery select serial execution. Use `--parallel=N`, `--parallel=auto`, `--no-routing` or `--no-isolate` to override defaults. Serial worktrees retain failed work and require deliberate recovery; the default does not discard an existing recovery tree. Quality repair budgets and opt-in competing candidates retain their existing policies.
138
+
139
+ Integration retains an execution slot until its candidate lands. Yoke loop runners disable native delegation so it cannot multiply the default worker budget: Codex disables multi_agent; Claude disallows Agent, Task, TeamCreate and SendMessage; Gemini uses a separate temporary system-settings copy that disables experimental agents and its always-on investigator/help overrides. Existing Gemini system policy and system-default paths are preserved; unreadable or malformed policy blocks launch. Original settings are never overwritten. The temporary copy is removed on normal exit; forced process termination may leave a private temporary directory. Explicit competing candidate counts remain a separate opt-in workload.
140
+
141
+ ### Dashboard views
142
+
143
+ - **Now:** reported task and worker activity, requested provider/model, elapsed worker time, phase, integration, blockers, status age, backlog and empirical remaining-time ranges. The view refreshes every five seconds while visible and not being operated with keyboard focus. A stale status is marked rather than asserted to be live.
144
+ - **Usage & time:** last 24 hours, 7/30/90/365 days or custom dates, grouped by UTC day, Monday-based week or month. Displays reported input/output and cache categories, cost coverage, consumption charts, per-model time buckets, per-task usage and summed phase/call durations. The overview also compares projects for the same period.
145
+ - **Results:** recorded acceptances, ended attempts, explicitly successful outcomes, repair phases, rule-driven escalations and recorded tokens/time per acceptance. Saved acceptance evidence and goal attempts remain available; their timestamps may fall outside the statistics period.
146
+
147
+ The short activity list still retains at most 1,000 events. Compact measurements now also persist under `.yoke/history/YYYY-MM-DD/` independently of that retention and across runs. They contain identifiers, model/provider evidence and measurements, not prompts or full status snapshots. Runtime history is excluded from Yoke commits and added to new project ignore rules. Available reviewer, quality critic and repair usage is recorded separately; absent usage is counted as unknown. Recent and archived records are deduplicated by event ID.
148
+
149
+ Dates and buckets use UTC. The custom end date is inclusive; the API uses an exclusive upper timestamp. Usage belongs to the time the provider reports it, and durations to their end time. Tokens per elapsed minute divide recorded input/output by the full selected interval. Tokens per call minute use summed reported call duration, which can overlap across workers. Neither is measured generation speed. Cache categories are shown separately without adding them to input again.
150
+
151
+ Historical activity that was never recorded or already expired cannot be reconstructed. Queue/human waiting time that lacks measurements remains unknown; phase and attempt durations are summed worker time, not automatically elapsed project time. Queries allow at most 366 days and bounded reads (8 MiB per shard, 32 MiB overall, 50,000 records); skipped, malformed or oversized history is reported as incomplete. History currently requires local disk retention rather than automatic monthly compaction. Unknown model identity is never replaced by a requested model, and missing charges are not estimated from token counts.
152
+
153
+ ```sh
154
+ yoke projects add /path/to/project
155
+ yoke projects list
156
+ yoke dashboard .
157
+ yoke dashboard --no-register --port=4100
158
+ yoke projects remove <opaque-project-id>
159
+ ```
160
+
161
+ Open the printed loopback URL. The dashboard displays project goals, backlog, worker/status data, evidence, blockers, events, usage availability and schedule ranges. It can request a goal pause through the same service as the CLI. Start, resume and budget changes remain explicit CLI operations. Removing registration only removes the registry reference.
162
+
163
+ The HTTP service binds to `127.0.0.1`, validates Host/Origin, requires a per-session token for pause, sends a restrictive CSP and renders project strings as text. It does not serve arbitrary local files or allow remote registration. Missing/corrupt/oversized project data is shown as unavailable. Shared project registration and protected acceptance live under `YOKE_STATE_DIR` or `~/.yoke/state`; existing routing history retains its `YOKE_REGISTRY_DIR` location. The dashboard is local, not a hosted multi-user service.
164
+
165
+ ## Remaining validation and product work
166
+
167
+ Live Codex/Claude/Gemini comparisons, measured competitive development-time/cost studies and estimate calibration require representative real projects and authenticated runs. Equal adapter contracts do not imply equal model capability. Cloud/team access, automatic selective test reuse, web-based start/resume controls and commercial rollout are not included in this local foundation. The complete saved direction is in [PRODUCT-DIRECTION-2026-09-05.md](PRODUCT-DIRECTION-2026-09-05.md).
@@ -0,0 +1,104 @@
1
+ # Windows runner correction — 2026-09-06
2
+
3
+ AI-assisted implementation and validation record for issue #5. These changes are
4
+ included in 1.10.0; installed 1.9.0 packages and existing processes must be updated/restarted to use them.
5
+
6
+ ## Reproduction and correction
7
+
8
+ On this machine, Codex 0.153.4 running the Microsoft Store PowerShell executable
9
+ under `codex sandbox -P :workspace` reproduced `CreateProcessAsUserW failed:
10
+ -1073283067`. The System32 Windows PowerShell executable succeeded under the
11
+ same sandbox profile. A first end-to-end probe exposed a second entry point:
12
+ the WindowsApps `pwsh.exe` app-execution alias failed with access denied.
13
+
14
+ Yoke now filters Store PowerShell directories and their aliases from the provider's
15
+ own PATH, prefers an available native PowerShell executable and tests it in the
16
+ actual isolated working directory before starting the model. The same PATH is
17
+ passed to Codex and its shell environment. This changes neither machine/user PATH
18
+ nor sandbox privileges. Native PowerShell 7 is preferred; Windows PowerShell is
19
+ the fallback when no native `pwsh.exe` is available. Projects requiring PowerShell
20
+ 7 should install a native distribution and make it available on PATH.
21
+
22
+ The preflight uses the installed CLI's built-in `:workspace` or `:read-only`
23
+ permission profile. An unsupported CLI/profile or failing shell blocks execution
24
+ with an actionable message. No permission escalation or unsandboxed retry occurs.
25
+ The preflight itself has a 30-second execution deadline, followed by bounded
26
+ process-tree cleanup; an outer 90-second guard bounds an unresponsive supervisor.
27
+ Environment values, prompts and raw authentication diagnostics are not stored in
28
+ the supervision record.
29
+
30
+ Windows npm launchers are resolved to their JavaScript entry point and invoked
31
+ with a native Node argv array; executable providers are launched directly. Unknown
32
+ batch launcher formats are rejected. This removes the provider/watchdog use of
33
+ `shell: true` and preserves spaces, quotes, percent signs and shell metacharacters
34
+ as literal arguments.
35
+
36
+ ## Supervision and failure behavior
37
+
38
+ Both serial and parallel implementation providers have independent output,
39
+ successful-tool/edit progress and overall timers. Defaults:
40
+
41
+ ```yaml
42
+ loop:
43
+ timeoutMinutes: 20 # output inactivity; existing setting
44
+ progressTimeoutMinutes: 20 # no observed successful tool result or edit
45
+ maxCallMinutes: 30 # overall provider call
46
+ ```
47
+
48
+ The two new settings accept positive values up to 1,440 minutes. Disabling the
49
+ old output timeout does not disable the other bounds. Successful tool/edit events
50
+ are evidence of activity, not proof that the product requirement has been met.
51
+
52
+ Codex structured failed-tool events and its actual `exec_command failed:
53
+ CreateProcess` stderr diagnostic stop the worker immediately. Terminal provider
54
+ authentication errors are separate from optional MCP startup warnings; the latter
55
+ alone do not fail the task. Failed infrastructure cannot produce a candidate or
56
+ story success merely because existing tests happen to pass, and does not train
57
+ capability escalation. Explicit `bare` startup settings now survive capability
58
+ profile selection.
59
+
60
+ Process records under `.yoke/supervision/` expose PID, process identity, supervisor
61
+ heartbeat, current attempt, last output, last successful tool/edit and terminal reason. Status CLI
62
+ and dashboard show these separately. Backlog ratios are labeled as backlog, not
63
+ overall product completion. Live identity checks distinguish an existing PID from
64
+ the recorded process. Tree termination checks that identity before reaping;
65
+ unconfirmed termination retains ownership/evidence and blocks a subsequent worker.
66
+ No retry/restart is triggered merely by an observer timing out. Failed worktrees
67
+ remain available for inspection.
68
+
69
+ ## Validation
70
+
71
+ - Model-free reproduction: Store executable failed with the reported code;
72
+ native System32 PowerShell passed in the same permission profile.
73
+ - Both safe and read-only preflights passed with a 30,707 UTF-16-character inherited
74
+ environment in a path containing spaces. This simulates a large environment;
75
+ it does not reconstruct every variable in the original Visual Studio session.
76
+ - Actual Yoke isolated safe-mode run: `codex-light`, requested `gpt-5.6-luna` at low
77
+ effort; task started 18:49:44 UTC and completed 18:52:52 UTC. Real shell commands,
78
+ file creation, two independent acceptance commands, verification and commit
79
+ integration succeeded. One model call, no assessment call. Recorded input
80
+ 345,821 tokens (321,024 cached subset), output 2,353; monetary cost and actual
81
+ reported model identity remain unknown.
82
+ - Regression tests cover literal argv, packaged aliases, streaming failure,
83
+ optional MCP diagnostics, total timeout despite heartbeat output, retained
84
+ redacted proof, no new worker after unconfirmed termination, and an actual
85
+ child failure propagated through capability routing and the worker gate.
86
+ - Full suite: 1,170 passed, two skipped across 128 files at the recorded full-run
87
+ checkpoint. Final focused validation: 119 passed across five files, plus a
88
+ successful build and documentation consistency check.
89
+
90
+ The original DeviceLane process was observed alive. It was not restarted or
91
+ terminated, its retained worktree was not modified, and its application task was
92
+ not claimed complete. Issue #5 was not closed or externally commented on. The fix
93
+ must be installed before it can supervise a newly started DeviceLane run; it cannot
94
+ retrofit supervision into an already running old process.
95
+
96
+ Primary implementation reference consulted:
97
+ [OpenAI shell detection](https://github.com/openai/codex/blob/main/codex-rs/shell-command/src/shell_detect.rs).
98
+ The local executable probes above, rather than assumptions about upstream release
99
+ contents, establish the behavior observed here.
100
+
101
+ Read-only provenance audit: supported scan complete, no C2PA located; verification,
102
+ signer trust and Markdown metadata privacy unknown. The audit states: "No conforming
103
+ verifier was supplied." and "Keyed model-level watermarks cannot be checked without
104
+ the provider's key." No authorship inference or watermark removal was performed.
Binary file
@@ -0,0 +1,85 @@
1
+ # Yoke community outreach — 2026-08-20
2
+
3
+ These posts are intentionally feedback-first. They disclose the creator relationship, avoid vote requests, and link directly to the free MIT-licensed project.
4
+
5
+ ## Hacker News — Show HN
6
+
7
+ **Title**
8
+
9
+ Show HN: Yoke – safety-gated autonomous coding loops for Claude, Codex, and Gemini
10
+
11
+ **URL**
12
+
13
+ https://github.com/HECer/yoke
14
+
15
+ **First comment**
16
+
17
+ Hi HN — I built Yoke after repeatedly seeing the same failure mode in long-running coding-agent loops: the agent says a task is done, but the tests were not actually run, the working tree is inconsistent, or a later iteration breaks earlier work.
18
+
19
+ Yoke is an MIT-licensed TypeScript CLI that puts mechanical gates around those loops. A story only passes after the configured verification command succeeds, review approves it, and the commit lands. Stories run in isolated worktrees, and optional Playwright evidence can attach screenshots to acceptance criteria. It generates native project instructions for Claude Code, Codex CLI, and Gemini CLI from one canon rather than asking all three tools to interpret the same generic prompt.
20
+
21
+ You can try it without signing up:
22
+
23
+ npm i -g @hecer/yoke
24
+ yoke new my-app
25
+
26
+ I would especially value feedback from people who already run Ralph-style or overnight agent loops: which gate feels essential, which feels like ceremony, and what would stop you from trying this on a real repository?
27
+
28
+ I am the creator and will be around to answer technical questions.
29
+
30
+ ## Reddit — r/ClaudeCodeTLDR weekly showcase (English)
31
+
32
+ I built **Yoke**, a free MIT-licensed harness for running Claude Code, Codex CLI, or Gemini CLI in autonomous coding loops with mechanical safety gates.
33
+
34
+ The problem I wanted to solve was the gap between an agent saying “done” and the repository actually being in a verified state. Yoke uses isolated worktrees and only marks a story complete after the configured tests pass, review approves it, and the commit lands. It can also require Playwright screenshot evidence for visual acceptance criteria.
35
+
36
+ Claude Code was both a target runtime and part of the development/review workflow; the project also generates native instructions for Codex and Gemini from the same canon.
37
+
38
+ It is free to use, requires no account, and the quickstart is:
39
+
40
+ npm i -g @hecer/yoke
41
+ yoke new my-app
42
+
43
+ Repo: https://github.com/HECer/yoke
44
+
45
+ I’m the creator. I’d love blunt feedback from people who already use long-running Claude Code loops: would these gates make you trust an overnight run more, or does the workflow look too heavy? If you try it, which part is confusing first?
46
+
47
+ ## Reddit — r/ChatGPTCoding self-promotion thread (English)
48
+
49
+ I built **Yoke**, an MIT-licensed CLI for people using Claude Code, Codex CLI, or Gemini CLI for longer autonomous coding runs.
50
+
51
+ Instead of trusting the agent’s “done” message, Yoke gates each story on a real verification command, review approval, and an atomic commit. It also isolates stories in worktrees and can collect Playwright screenshots as acceptance evidence.
52
+
53
+ Install and try it without an account:
54
+
55
+ npm i -g @hecer/yoke
56
+ yoke new my-app
57
+
58
+ https://github.com/HECer/yoke
59
+
60
+ I’m the creator. I’m mainly looking for honest feedback: if you already run coding agents for more than one task at a time, what would make you try this, and what looks like unnecessary process?
61
+
62
+ ## German-language developer communities (draft for communities that explicitly permit project showcases)
63
+
64
+ **Titel**
65
+
66
+ Feedback gesucht: Yoke – abgesicherte autonome Coding-Loops für Claude, Codex und Gemini
67
+
68
+ **Text**
69
+
70
+ Ich habe **Yoke** gebaut, ein kostenloses, MIT-lizenziertes CLI für längere autonome Coding-Läufe mit Claude Code, Codex CLI oder Gemini CLI.
71
+
72
+ Der Auslöser war die Lücke zwischen „der Agent sagt, er sei fertig“ und einem tatsächlich verifizierten Repository. Yoke markiert eine Story erst dann als erledigt, wenn der konfigurierte Test-/Verify-Befehl erfolgreich war, ein Review zugestimmt hat und der Commit sauber gelandet ist. Einzelne Stories laufen isoliert in Worktrees; für visuelle Akzeptanzkriterien können Playwright-Screenshots als Nachweis verlangt werden.
73
+
74
+ Ausprobieren ohne Account:
75
+
76
+ npm i -g @hecer/yoke
77
+ yoke new meine-app
78
+
79
+ Repo: https://github.com/HECer/yoke
80
+
81
+ Ich bin der Entwickler des Projekts. Mich interessiert vor allem ehrliches Feedback von Leuten, die Coding-Agents länger oder autonom laufen lassen: Würden solche Gates euer Vertrauen erhöhen, oder wirkt der Ablauf zu schwergewichtig? Was wäre beim ersten Ausprobieren vermutlich die größte Hürde?
82
+
83
+ ## Disclosure note
84
+
85
+ The outreach copy in this file was drafted with AI assistance on the creator's instructions. Add a human-readable AI-assistance disclosure wherever a community's rules require it; do not infer human authorship from the absence of machine-readable provenance.