@hecer/yoke 1.9.0 → 1.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (153) hide show
  1. package/.claude-plugin/plugin.json +13 -13
  2. package/.codex-plugin/plugin.json +7 -7
  3. package/CHANGELOG.md +398 -358
  4. package/README.md +915 -913
  5. package/TODOS.md +5 -5
  6. package/agents/docs.toml +6 -6
  7. package/agents/implementer.toml +6 -6
  8. package/agents/reviewer.toml +6 -6
  9. package/agents/security.toml +6 -6
  10. package/bench/README.md +86 -86
  11. package/bench/RESULTS.md +35 -35
  12. package/bench/output-compaction.mjs +65 -65
  13. package/bench/result-schema.mjs +12 -12
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -15
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
  17. package/bench/run-matrix.mjs +26 -26
  18. package/bench/run.mjs +106 -106
  19. package/canon/AGENTS.md +30 -30
  20. package/canon/context/DECISIONS.md +4 -4
  21. package/canon/context/GLOSSARY.md +11 -11
  22. package/canon/context/KNOWLEDGE.md +4 -4
  23. package/canon/context/PROJECT.md +15 -15
  24. package/canon/loop/loop-spec.md +65 -65
  25. package/canon/loop/prd.schema.md +46 -40
  26. package/canon/manifest.yaml +59 -59
  27. package/canon/policy/gates.md +7 -7
  28. package/canon/policy/roles.md +9 -9
  29. package/canon/skills/ATTRIBUTION.md +99 -99
  30. package/canon/skills/authoring-prd/SKILL.md +56 -56
  31. package/canon/skills/brainstorming/SKILL.md +164 -164
  32. package/canon/skills/codebase-design/DEEPENING.md +15 -15
  33. package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
  34. package/canon/skills/codebase-design/SKILL.md +39 -39
  35. package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
  36. package/canon/skills/document-release/SKILL.md +302 -302
  37. package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
  38. package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
  39. package/canon/skills/domain-modeling/SKILL.md +35 -35
  40. package/canon/skills/executing-plans/SKILL.md +70 -70
  41. package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
  42. package/canon/skills/health/SKILL.md +177 -177
  43. package/canon/skills/maintaining-context/SKILL.md +34 -34
  44. package/canon/skills/minimal-code/SKILL.md +21 -21
  45. package/canon/skills/no-ai-slop/SKILL.md +103 -103
  46. package/canon/skills/no-ai-slop/eval.md +43 -43
  47. package/canon/skills/plan-ceo-review/SKILL.md +541 -541
  48. package/canon/skills/plan-eng-review/SKILL.md +362 -362
  49. package/canon/skills/receiving-code-review/SKILL.md +213 -213
  50. package/canon/skills/requesting-code-review/SKILL.md +105 -105
  51. package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
  52. package/canon/skills/retro/SKILL.md +397 -397
  53. package/canon/skills/review/SKILL.md +246 -246
  54. package/canon/skills/ship/SKILL.md +691 -691
  55. package/canon/skills/subagent-driven-development/SKILL.md +277 -277
  56. package/canon/skills/systematic-debugging/SKILL.md +296 -296
  57. package/canon/skills/tdd/SKILL.md +371 -371
  58. package/canon/skills/unslop-ui/SKILL.md +34 -34
  59. package/canon/skills/using-git-worktrees/SKILL.md +218 -218
  60. package/canon/skills/verification-before-completion/SKILL.md +139 -139
  61. package/canon/skills/visual-verification/SKILL.md +54 -54
  62. package/canon/skills/workflow/SKILL.md +22 -22
  63. package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
  64. package/canon/skills/writing-for-agents/SKILL.md +42 -42
  65. package/canon/skills/writing-plans/SKILL.md +152 -152
  66. package/canon/skills/writing-skills/SKILL.md +655 -655
  67. package/canon/skills/yoke-retrofit/SKILL.md +26 -26
  68. package/canon/skills/yoke-workflow/SKILL.md +20 -20
  69. package/canon/tools/codex-rtk-hook.mjs +35 -35
  70. package/canon/tools/gemini-rtk-hook.mjs +25 -25
  71. package/canon/tools/graphify.md +3 -3
  72. package/canon/tools/playwright-mcp.md +3 -3
  73. package/canon/tools/rtk.md +7 -7
  74. package/canon/tools/serena.md +6 -6
  75. package/dist/agents/contracts.js +1 -1
  76. package/dist/agents/host.js +4 -0
  77. package/dist/agents/process-incarnation.js +1 -1
  78. package/dist/agents/process.js +74 -6
  79. package/dist/agents/providers.js +13 -0
  80. package/dist/agents/supervision.js +153 -0
  81. package/dist/agents/telemetry.js +33 -0
  82. package/dist/agents/windows-launch.js +80 -0
  83. package/dist/canon/manifest.js +1 -1
  84. package/dist/change/inbox.js +21 -5
  85. package/dist/cli.js +19 -10
  86. package/dist/dashboard/discovery.js +73 -0
  87. package/dist/dashboard/page.js +122 -28
  88. package/dist/dashboard/panels.js +91 -15
  89. package/dist/goals/command.js +4 -2
  90. package/dist/loop/claims.js +1 -1
  91. package/dist/loop/decision.js +2 -2
  92. package/dist/loop/git.js +12 -4
  93. package/dist/loop/loop.js +8 -4
  94. package/dist/loop/parallel-adapters.js +2 -3
  95. package/dist/loop/parallel-command.js +5 -0
  96. package/dist/loop/prd.js +3 -1
  97. package/dist/loop/reporter.js +4 -1
  98. package/dist/loop/run-command.js +11 -2
  99. package/dist/loop/runner.js +22 -26
  100. package/dist/loop/watchdog.js +87 -11
  101. package/dist/loop/worker.js +5 -3
  102. package/dist/prd/assess.js +145 -0
  103. package/dist/prd/command.js +76 -38
  104. package/dist/quality/types.js +1 -1
  105. package/dist/retrofit/config.js +11 -0
  106. package/dist/retrofit/plan.js +2 -0
  107. package/dist/retrofit/planners/claude.js +14 -14
  108. package/dist/retrofit/planners/qwen.js +73 -0
  109. package/dist/retrofit/preserve.js +2 -2
  110. package/dist/retrofit/skill-actions.js +1 -0
  111. package/dist/review/command.js +1 -1
  112. package/dist/routing/assessment.js +1 -1
  113. package/dist/routing/capability.js +25 -13
  114. package/dist/routing/contracts.js +60 -0
  115. package/dist/routing/planning.js +12 -0
  116. package/dist/routing/router.js +51 -16
  117. package/dist/setup/command.js +11 -3
  118. package/docs/BATCH-PLANNING-VALIDATION.md +67 -0
  119. package/docs/CAPABILITY-ROUTING.md +78 -50
  120. package/docs/DASHBOARD-EVOLUTION.md +33 -0
  121. package/docs/MIGRATING-TO-1.0.md +33 -33
  122. package/docs/MIGRATING-TO-1.1.md +27 -27
  123. package/docs/MIGRATING-TO-1.4.md +70 -70
  124. package/docs/PRODUCT-DIRECTION-2026-09-05.md +218 -200
  125. package/docs/PUBLISHING.md +114 -114
  126. package/docs/VERIFIED-PROJECTS-VALIDATION.md +29 -29
  127. package/docs/VERIFIED-PROJECTS.md +167 -167
  128. package/docs/WINDOWS-RUNNER-VALIDATION.md +104 -0
  129. package/docs/assets/yoke-logo.png +0 -0
  130. package/docs/community-outreach-2026-08-20.md +85 -0
  131. package/docs/launch-copy-2026-08-21.md +193 -0
  132. package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
  133. package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
  134. package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
  135. package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
  136. package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
  137. package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
  138. package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
  139. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
  140. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
  141. package/docs/superpowers/plans/2026-09-05-verified-projects.md +83 -83
  142. package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
  143. package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
  144. package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
  145. package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
  146. package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
  147. package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
  148. package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
  149. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
  150. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
  151. package/gemini-extension.json +6 -6
  152. package/hooks/hooks.json +19 -19
  153. package/package.json +87 -87
@@ -0,0 +1,193 @@
1
+ # Yoke launch copy
2
+
3
+ ## 1. Show HN
4
+
5
+ ### Title
6
+
7
+ Show HN: Yoke – gated coding-agent loops for Claude, Codex, and Gemini
8
+
9
+ ### URL
10
+
11
+ https://github.com/HECer/yoke
12
+
13
+ ### First comment
14
+
15
+ I built Yoke after seeing coding agents report that a task was finished when the repository told a different story: tests had not run, the working tree contained unrelated changes, or a later iteration had broken earlier work.
16
+
17
+ Yoke is an MIT-licensed TypeScript CLI for running coding-agent work as a sequence of verifiable stories. It supports Claude Code, Codex CLI, and Gemini CLI.
18
+
19
+ The runner, rather than the agent, decides whether a story passed. It checks the configured verification command's exit code, requires review approval, and confirms that the commit landed. Stories run in isolated Git worktrees. For visual acceptance criteria, a run can require Playwright screenshots or video as evidence.
20
+
21
+ Yoke also generates native project instructions for each supported agent from one versioned skill canon. That part is Markdown. The state transitions, worktree lifecycle, process supervision, gates, and run state are executable TypeScript.
22
+
23
+ The closest comparison I know is WorkOS Case. Both projects treat the harness as the reliability boundary. Case is focused on turning GitHub or Linear issues into reviewed pull requests. Yoke includes greenfield and retrofit setup, PRD-to-story execution, and portable project configuration across three coding agents.
24
+
25
+ You can try it without an account:
26
+
27
+ npm install -g @hecer/yoke
28
+ yoke new my-app
29
+
30
+ I am looking for two kinds of feedback:
31
+
32
+ 1. Which gate would you remove because it adds more process than safety?
33
+ 2. What failure from a real autonomous run would Yoke still miss?
34
+
35
+ I am the maintainer and will be around to answer questions.
36
+
37
+ ## 2. DEV Community (`#showdev`)
38
+
39
+ ### Title
40
+
41
+ I moved coding-agent verification out of the prompt
42
+
43
+ ### Tags
44
+
45
+ `showdev`, `ai`, `opensource`, `typescript`
46
+
47
+ ### Article
48
+
49
+ A coding agent can agree to run the tests and still fail to run them. It can also run the wrong command, overlook a dirty working tree, or declare success before the result is committed.
50
+
51
+ Adding another sentence to the prompt does not change who controls the completion decision. The same model doing the work still decides whether it followed the instruction.
52
+
53
+ I built Yoke to put that decision in a TypeScript runner.
54
+
55
+ Yoke executes work as a sequence of stories. For each story, the runner creates an isolated Git worktree, starts a fresh agent session, runs the configured verification command, requests a separate review, and checks that the resulting commit landed. A story is not marked as passed because the agent says it is done.
56
+
57
+ For user-interface work, the acceptance criteria can require Playwright screenshots or video. The evidence belongs to the story instead of living in a chat transcript.
58
+
59
+ Yoke supports Claude Code, Codex CLI, and Gemini CLI. A versioned canon generates the native instructions each agent expects, so a project can keep one methodology without pretending that all three tools use identical configuration formats.
60
+
61
+ Some of Yoke is Markdown. The skills, policies, and project context should be readable and editable. The enforcement layer is code: process supervision, worktree isolation, persisted run state, verification exit codes, review gates, and commit checks.
62
+
63
+ The closest existing project is WorkOS Case. Case uses a deterministic TypeScript pipeline to turn issues into reviewed pull requests with evidence. Yoke covers a somewhat different workflow: creating or retrofitting a project, converting a PRD into ordered stories, and running those stories across Claude, Codex, or Gemini. The projects share more philosophy than I realized when I first described Yoke, and I now call that out directly.
64
+
65
+ Yoke is MIT-licensed and does not require an account.
66
+
67
+ npm install -g @hecer/yoke
68
+ yoke new my-app
69
+
70
+ Repository: https://github.com/HECer/yoke
71
+
72
+ I would value examples of failures that prompts and skills do not handle reliably. I am also interested in the opposite feedback: which Yoke gate feels unnecessary in normal development?
73
+
74
+ Disclosure: I used AI assistance while drafting and editing this article. The product description and technical claims were checked against the repository by the maintainer.
75
+
76
+ ## 3. DevHunt
77
+
78
+ ### Name
79
+
80
+ Yoke
81
+
82
+ ### Tagline
83
+
84
+ Run coding agents with tests, review, and commit gates outside the model.
85
+
86
+ ### Short description
87
+
88
+ Yoke is an MIT-licensed TypeScript CLI for autonomous coding loops with Claude Code, Codex CLI, and Gemini CLI. It runs stories in isolated Git worktrees and only records a pass after the configured verification command succeeds, review approves the change, and the commit lands. Visual stories can require Playwright screenshot or video evidence.
89
+
90
+ ### Maker comment
91
+
92
+ I built Yoke because instructions such as “run the tests before finishing” leave the completion decision with the same agent doing the work. Yoke moves that decision into an executable runner.
93
+
94
+ I would like feedback from developers who already run agents unattended: which gate earns its cost, and what would prevent you from using Yoke on an existing repository?
95
+
96
+ Source: https://github.com/HECer/yoke
97
+
98
+ ## 4. Uneed
99
+
100
+ ### Title
101
+
102
+ Yoke
103
+
104
+ ### Tagline
105
+
106
+ A safety-gated coding harness for Claude, Codex, and Gemini.
107
+
108
+ ### Description
109
+
110
+ Yoke runs autonomous coding work as verifiable stories. Each story gets an isolated Git worktree and must pass the project's real verification command, an independent review, and a commit check. User-interface stories can also require Playwright screenshots or video.
111
+
112
+ The CLI can create a new project or retrofit an existing one. It generates native instructions for Claude Code, Codex CLI, and Gemini CLI from one versioned canon.
113
+
114
+ Yoke is open source under the MIT license and requires no account.
115
+
116
+ ### Launch comment
117
+
118
+ I am the maintainer. I am looking for concrete failure cases from people who use coding agents for long or unattended runs. If Yoke would not catch one of yours, please tell me what happened.
119
+
120
+ https://github.com/HECer/yoke
121
+
122
+ ## 5. Product Hunt
123
+
124
+ ### Name
125
+
126
+ Yoke
127
+
128
+ ### Tagline
129
+
130
+ Mechanical safety gates for autonomous coding agents
131
+
132
+ ### Description
133
+
134
+ Run Claude Code, Codex CLI, or Gemini CLI through isolated worktrees, real test commands, independent review, commit checks, and optional visual evidence. Open source and local-first.
135
+
136
+ ### First maker comment
137
+
138
+ Hi Product Hunt,
139
+
140
+ I built Yoke for developers who let coding agents work through more than one task at a time.
141
+
142
+ The failure I wanted to address was straightforward: an agent could say that tests passed without giving the repository a reliable completion boundary. Yoke makes the runner responsible for that boundary. A story passes only after the configured verification command exits successfully, review approves the change, and the commit lands.
143
+
144
+ Yoke supports Claude Code, Codex CLI, and Gemini CLI. It can bootstrap a new project or retrofit an existing repository, and it keeps stories isolated in Git worktrees.
145
+
146
+ The project is MIT-licensed, runs locally, and requires no account:
147
+
148
+ https://github.com/HECer/yoke
149
+
150
+ I would appreciate feedback from anyone already running coding agents unattended. What evidence do you need before trusting the result?
151
+
152
+ ## 6. Peerlist Launchpad
153
+
154
+ ### Project summary
155
+
156
+ Yoke is an open-source TypeScript CLI for running Claude Code, Codex CLI, and Gemini CLI with completion gates outside the model.
157
+
158
+ It breaks a PRD into stories, runs each story in an isolated Git worktree, and checks the project's verification command, review result, and final commit before recording a pass. Visual acceptance criteria can require Playwright screenshots or video.
159
+
160
+ I built it after finding that prompt-level instructions were useful guidance but a weak enforcement boundary for unattended work. The readable methodology lives in versioned Markdown; the enforcement lives in the runner.
161
+
162
+ Try it:
163
+
164
+ npm install -g @hecer/yoke
165
+ yoke new my-app
166
+
167
+ Source: https://github.com/HECer/yoke
168
+
169
+ I am looking for developers willing to try it on an existing repository and report the first confusing or unnecessary step.
170
+
171
+ ## 7. Indie Hackers
172
+
173
+ ### Title
174
+
175
+ I built a coding-agent harness, then learned that “it is just Markdown” was the right criticism to answer
176
+
177
+ ### Post
178
+
179
+ I recently shared Yoke, an open-source harness for autonomous coding agents, and received a blunt question: why not encode the workflow in one skill or even one sentence?
180
+
181
+ That criticism exposed a problem in how I described the project. The interesting part is not the workflow advice. Agents can read instructions telling them to plan, test, review, and commit.
182
+
183
+ The product decision is who controls the pass condition.
184
+
185
+ Yoke puts that condition in a TypeScript runner. It creates isolated worktrees, executes the configured verification command, requests review, checks the commit, and persists story state. The agent produces the change, but it cannot mark its own story as passed merely by saying the work is complete.
186
+
187
+ The skills and policies are Markdown because users should be able to read and change them. If someone only needs those instructions, a skill is enough and Yoke is unnecessary.
188
+
189
+ I also learned that WorkOS Case already takes a closely related approach. Case focuses on moving an issue through a multi-agent pipeline into a reviewed pull request. Yoke focuses on creating or retrofitting projects, turning PRDs into stories, and running them across Claude Code, Codex CLI, or Gemini CLI. I now describe that overlap directly instead of pretending the category is empty.
190
+
191
+ The project is MIT-licensed: https://github.com/HECer/yoke
192
+
193
+ My question for other builders is about scope. Would you keep the product narrow around enforced story completion, or is cross-agent project setup a meaningful part of the value?