@hecer/yoke 1.9.0 → 1.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (153) hide show
  1. package/.claude-plugin/plugin.json +13 -13
  2. package/.codex-plugin/plugin.json +7 -7
  3. package/CHANGELOG.md +398 -358
  4. package/README.md +915 -913
  5. package/TODOS.md +5 -5
  6. package/agents/docs.toml +6 -6
  7. package/agents/implementer.toml +6 -6
  8. package/agents/reviewer.toml +6 -6
  9. package/agents/security.toml +6 -6
  10. package/bench/README.md +86 -86
  11. package/bench/RESULTS.md +35 -35
  12. package/bench/output-compaction.mjs +65 -65
  13. package/bench/result-schema.mjs +12 -12
  14. package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
  15. package/bench/results/codex-unavailable-1785175418318.json +15 -15
  16. package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
  17. package/bench/run-matrix.mjs +26 -26
  18. package/bench/run.mjs +106 -106
  19. package/canon/AGENTS.md +30 -30
  20. package/canon/context/DECISIONS.md +4 -4
  21. package/canon/context/GLOSSARY.md +11 -11
  22. package/canon/context/KNOWLEDGE.md +4 -4
  23. package/canon/context/PROJECT.md +15 -15
  24. package/canon/loop/loop-spec.md +65 -65
  25. package/canon/loop/prd.schema.md +46 -40
  26. package/canon/manifest.yaml +59 -59
  27. package/canon/policy/gates.md +7 -7
  28. package/canon/policy/roles.md +9 -9
  29. package/canon/skills/ATTRIBUTION.md +99 -99
  30. package/canon/skills/authoring-prd/SKILL.md +56 -56
  31. package/canon/skills/brainstorming/SKILL.md +164 -164
  32. package/canon/skills/codebase-design/DEEPENING.md +15 -15
  33. package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
  34. package/canon/skills/codebase-design/SKILL.md +39 -39
  35. package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
  36. package/canon/skills/document-release/SKILL.md +302 -302
  37. package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
  38. package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
  39. package/canon/skills/domain-modeling/SKILL.md +35 -35
  40. package/canon/skills/executing-plans/SKILL.md +70 -70
  41. package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
  42. package/canon/skills/health/SKILL.md +177 -177
  43. package/canon/skills/maintaining-context/SKILL.md +34 -34
  44. package/canon/skills/minimal-code/SKILL.md +21 -21
  45. package/canon/skills/no-ai-slop/SKILL.md +103 -103
  46. package/canon/skills/no-ai-slop/eval.md +43 -43
  47. package/canon/skills/plan-ceo-review/SKILL.md +541 -541
  48. package/canon/skills/plan-eng-review/SKILL.md +362 -362
  49. package/canon/skills/receiving-code-review/SKILL.md +213 -213
  50. package/canon/skills/requesting-code-review/SKILL.md +105 -105
  51. package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
  52. package/canon/skills/retro/SKILL.md +397 -397
  53. package/canon/skills/review/SKILL.md +246 -246
  54. package/canon/skills/ship/SKILL.md +691 -691
  55. package/canon/skills/subagent-driven-development/SKILL.md +277 -277
  56. package/canon/skills/systematic-debugging/SKILL.md +296 -296
  57. package/canon/skills/tdd/SKILL.md +371 -371
  58. package/canon/skills/unslop-ui/SKILL.md +34 -34
  59. package/canon/skills/using-git-worktrees/SKILL.md +218 -218
  60. package/canon/skills/verification-before-completion/SKILL.md +139 -139
  61. package/canon/skills/visual-verification/SKILL.md +54 -54
  62. package/canon/skills/workflow/SKILL.md +22 -22
  63. package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
  64. package/canon/skills/writing-for-agents/SKILL.md +42 -42
  65. package/canon/skills/writing-plans/SKILL.md +152 -152
  66. package/canon/skills/writing-skills/SKILL.md +655 -655
  67. package/canon/skills/yoke-retrofit/SKILL.md +26 -26
  68. package/canon/skills/yoke-workflow/SKILL.md +20 -20
  69. package/canon/tools/codex-rtk-hook.mjs +35 -35
  70. package/canon/tools/gemini-rtk-hook.mjs +25 -25
  71. package/canon/tools/graphify.md +3 -3
  72. package/canon/tools/playwright-mcp.md +3 -3
  73. package/canon/tools/rtk.md +7 -7
  74. package/canon/tools/serena.md +6 -6
  75. package/dist/agents/contracts.js +1 -1
  76. package/dist/agents/host.js +4 -0
  77. package/dist/agents/process-incarnation.js +1 -1
  78. package/dist/agents/process.js +74 -6
  79. package/dist/agents/providers.js +13 -0
  80. package/dist/agents/supervision.js +153 -0
  81. package/dist/agents/telemetry.js +33 -0
  82. package/dist/agents/windows-launch.js +80 -0
  83. package/dist/canon/manifest.js +1 -1
  84. package/dist/change/inbox.js +21 -5
  85. package/dist/cli.js +19 -10
  86. package/dist/dashboard/discovery.js +73 -0
  87. package/dist/dashboard/page.js +122 -28
  88. package/dist/dashboard/panels.js +91 -15
  89. package/dist/goals/command.js +4 -2
  90. package/dist/loop/claims.js +1 -1
  91. package/dist/loop/decision.js +2 -2
  92. package/dist/loop/git.js +12 -4
  93. package/dist/loop/loop.js +8 -4
  94. package/dist/loop/parallel-adapters.js +2 -3
  95. package/dist/loop/parallel-command.js +5 -0
  96. package/dist/loop/prd.js +3 -1
  97. package/dist/loop/reporter.js +4 -1
  98. package/dist/loop/run-command.js +11 -2
  99. package/dist/loop/runner.js +22 -26
  100. package/dist/loop/watchdog.js +87 -11
  101. package/dist/loop/worker.js +5 -3
  102. package/dist/prd/assess.js +145 -0
  103. package/dist/prd/command.js +76 -38
  104. package/dist/quality/types.js +1 -1
  105. package/dist/retrofit/config.js +11 -0
  106. package/dist/retrofit/plan.js +2 -0
  107. package/dist/retrofit/planners/claude.js +14 -14
  108. package/dist/retrofit/planners/qwen.js +73 -0
  109. package/dist/retrofit/preserve.js +2 -2
  110. package/dist/retrofit/skill-actions.js +1 -0
  111. package/dist/review/command.js +1 -1
  112. package/dist/routing/assessment.js +1 -1
  113. package/dist/routing/capability.js +25 -13
  114. package/dist/routing/contracts.js +60 -0
  115. package/dist/routing/planning.js +12 -0
  116. package/dist/routing/router.js +51 -16
  117. package/dist/setup/command.js +11 -3
  118. package/docs/BATCH-PLANNING-VALIDATION.md +67 -0
  119. package/docs/CAPABILITY-ROUTING.md +78 -50
  120. package/docs/DASHBOARD-EVOLUTION.md +33 -0
  121. package/docs/MIGRATING-TO-1.0.md +33 -33
  122. package/docs/MIGRATING-TO-1.1.md +27 -27
  123. package/docs/MIGRATING-TO-1.4.md +70 -70
  124. package/docs/PRODUCT-DIRECTION-2026-09-05.md +218 -200
  125. package/docs/PUBLISHING.md +114 -114
  126. package/docs/VERIFIED-PROJECTS-VALIDATION.md +29 -29
  127. package/docs/VERIFIED-PROJECTS.md +167 -167
  128. package/docs/WINDOWS-RUNNER-VALIDATION.md +104 -0
  129. package/docs/assets/yoke-logo.png +0 -0
  130. package/docs/community-outreach-2026-08-20.md +85 -0
  131. package/docs/launch-copy-2026-08-21.md +193 -0
  132. package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
  133. package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
  134. package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
  135. package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
  136. package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
  137. package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
  138. package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
  139. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
  140. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
  141. package/docs/superpowers/plans/2026-09-05-verified-projects.md +83 -83
  142. package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
  143. package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
  144. package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
  145. package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
  146. package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
  147. package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
  148. package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
  149. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
  150. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
  151. package/gemini-extension.json +6 -6
  152. package/hooks/hooks.json +19 -19
  153. package/package.json +87 -87
@@ -1,8 +1,9 @@
1
1
  import { buildWatchdogInvocation, makeRunner, runCapturedAgent, runnerInvocation, contextBlockFor, } from '../loop/runner.js';
2
- import { isAcceptanceCriterion } from '../loop/prd.js';
2
+ import { isAcceptanceCriterion, criterionCommandProblem } from '../loop/prd.js';
3
3
  import { historyForWorkers, projectHash, readRoutingObservations, recordRoutingObservation, storyHash } from './registry.js';
4
- import { assessmentInstructions, assessmentKey, parseAssessment } from './assessment.js';
5
- import { chooseCapability, readAssessment, saveAssessment } from './capability.js';
4
+ import { assessmentInstructions, parseAssessment, tiers } from './assessment.js';
5
+ import { chooseCapability, readAssessment, saveAssessment, routingAssessmentKey, knownInfrastructureFailure } from './capability.js';
6
+ import { readPlanningFile } from './contracts.js';
6
7
  const costRank = { low: 0, medium: 1, high: 2 };
7
8
  export function rankWorkers(workers, strategy, maxCandidates) {
8
9
  const history = historyForWorkers(workers);
@@ -144,43 +145,73 @@ function routingSteps(options) {
144
145
  selection,
145
146
  }));
146
147
  return function* (ctx) {
148
+ const blocked = (summary) => ({ success: false, summary, routing: { blocked: true, recordOutcome: () => undefined } });
149
+ if (options.strategy === 'capability' && options.assessmentPolicy === 'prepared') {
150
+ try {
151
+ const criteria = ctx.story.acceptance;
152
+ if (criteria.length < 2 || criteria.length > 5 || criteria.some(c => !isAcceptanceCriterion(c) || criterionCommandProblem(c)))
153
+ return blocked('Prepared routing requires 2-5 executable acceptance criteria');
154
+ if (!readAssessment(options.projectRoot ?? ctx.targetDir, ctx.story, true))
155
+ return blocked('Task assessment is missing or stale. Run yoke prd assess before execution.');
156
+ }
157
+ catch (error) {
158
+ return blocked(`Cannot read prepared assessment: ${error.message}`);
159
+ }
160
+ }
147
161
  if (options.strategy === 'capability' && !options.rules?.some(rule => (!rule.area || rule.area === ctx.story.area) && (!rule.storyId || rule.storyId === ctx.story.id))) {
148
162
  const root = options.projectRoot ?? ctx.targetDir;
149
- let assessment = readAssessment(root, ctx.story);
163
+ let assessment;
164
+ try {
165
+ assessment = readAssessment(root, ctx.story, options.assessmentPolicy === 'prepared');
166
+ }
167
+ catch (error) {
168
+ return blocked(`Cannot read assessment: ${error.message}`);
169
+ }
150
170
  let planning;
151
171
  const calls = [];
152
172
  if (!assessment) {
173
+ const inputKey = routingAssessmentKey(root, ctx.story);
153
174
  const prompt = [assessmentInstructions, 'Use the supplied task contract and project context to produce a bounded plan. Do not implement or change files.',
175
+ 'Approved planning brief:', readPlanningFile(root, '.yoke/plan.md', 80_000) ?? '',
154
176
  contextBlockFor(ctx.targetDir, ctx.story), JSON.stringify(ctx.story), 'Return exactly one line: YOKE_ASSESS {"taskClass":"implementation","difficulty":"medium","uncertainty":"low","risk":"low","scope":"low","testability":"high","reason":"evidence","approach":"steps and tests"}'].join('\n');
155
- const selection = { ...options.parentSelection, nativeMultiAgent: false };
177
+ const planner = options.planner?.agent ?? options.parent;
178
+ if (!available(planner))
179
+ return blocked('Configured planning provider is unavailable');
180
+ const selection = { ...(options.planner?.selection ?? options.parentSelection), nativeMultiAgent: false };
156
181
  const started = now();
157
- planning = yield () => options.captureRoute ? options.captureRoute(options.parent, ctx, prompt, selection)
158
- : runCapturedAgent(options.parent, buildWatchdogInvocation(runnerInvocation(options.parent, prompt, ctx.targetDir, true, 'read-only', selection), options.idleTimeoutMs ?? 0));
159
- calls.push(callUsage('orchestrator', options.parent, selection, planning.tokens, now() - started));
182
+ planning = yield () => options.captureRoute ? options.captureRoute(planner, ctx, prompt, selection)
183
+ : runCapturedAgent(planner, buildWatchdogInvocation(runnerInvocation(planner, prompt, ctx.targetDir, true, 'read-only', selection), options.idleTimeoutMs ?? 0));
184
+ calls.push(callUsage('orchestrator', planner, selection, planning.tokens, now() - started));
185
+ if (routingAssessmentKey(root, ctx.story) !== inputKey)
186
+ return { ...blocked('Planning inputs changed during assessment; retry planning with the current contract'), tokens: aggregateCalls(calls) };
160
187
  assessment = planning.success ? parseAssessment(planning.output) : undefined;
161
188
  if (assessment)
162
- saveAssessment(root, ctx.story, assessment, { provider: options.parent, model: planning.tokens?.model ?? selection.model });
189
+ saveAssessment(root, ctx.story, assessment, { provider: planner, model: planning.tokens?.model ?? selection.model });
163
190
  }
164
191
  if (!assessment)
165
192
  return { success: false, summary: 'Routing assessment unavailable or invalid; implementation was not started', tokens: aggregateCalls(calls), routing: { recordOutcome: () => undefined, blocked: true } };
166
- const choice = chooseCapability({ root, story: ctx.story, assessment, workers: eligibleWorkers, parent: options.parent, parentSelection: options.parentSelection, maxAttempts: options.maxAttempts });
193
+ const choice = chooseCapability({ root, story: ctx.story, assessment, workers: eligibleWorkers, parent: options.parent, parentSelection: options.parentSelection, maxAttempts: options.maxAttempts, fallback: options.fallback, maxTier: options.maxTier });
167
194
  options.onDecision?.(ctx.story.id, { profile: choice.worker?.id ?? 'SELF', provider: choice.provider, model: choice.selection.model, reasoningEffort: choice.selection.reasoningEffort, reason: choice.reason, next: choice.next, assessment });
195
+ if (choice.blocked)
196
+ return { ...blocked(choice.reason), tokens: aggregateCalls(calls) };
168
197
  if (choice.exhausted)
169
198
  return { success: false, summary: 'Routing attempt budget exhausted; replan this task before retrying', tokens: aggregateCalls(calls), routing: { recordOutcome: () => undefined, blocked: true } };
170
199
  const started = now();
171
- const result = yield () => makeWorker(choice.provider, choice.selection)({ ...ctx, story: { ...ctx.story, assessment } });
200
+ const result = yield () => makeWorker(choice.provider, choice.selection)({ ...ctx, attempt: choice.failures + 1, story: { ...ctx.story, assessment } });
172
201
  calls.push(callUsage(choice.worker ? 'worker' : 'parent', choice.provider, choice.selection, result.tokens, now() - started, choice.worker?.id ?? 'SELF'));
173
202
  let recorded = false;
203
+ const infrastructureFailure = result.infrastructureFailure || (!result.success && knownInfrastructureFailure(result.summary));
174
204
  return { ...result, summary: `route=${choice.worker?.id ?? 'SELF'} (${choice.reason}); ${result.summary}`,
205
+ ...(infrastructureFailure ? { success: false, infrastructureFailure: true } : {}),
175
206
  tokens: { ...aggregateCalls(calls), storyId: ctx.story.id, escalated: choice.failures > 1 },
176
- routing: { canRetry: !result.infrastructureFailure && choice.failures + 1 < (options.maxAttempts ?? 5), recordOutcome: (verified, failureKind) => {
207
+ routing: { blocked: infrastructureFailure || undefined, canRetry: !infrastructureFailure && choice.failures + 1 < (options.maxAttempts ?? 5), recordOutcome: (verified, failureKind) => {
177
208
  if (recorded)
178
209
  return;
179
210
  recorded = true;
180
- recordRoutingObservation({ projectHash: projectHash(root), storyHash: storyHash(projectHash(root), ctx.story.id), assessmentKey: assessmentKey(ctx.story), taskClass: assessment.taskClass, requiredTier: choice.requiredTier,
211
+ recordRoutingObservation({ projectHash: projectHash(root), storyHash: storyHash(projectHash(root), ctx.story.id), assessmentKey: routingAssessmentKey(root, ctx.story), taskClass: assessment.taskClass, requiredTier: choice.requiredTier,
181
212
  role: 'implementation', strategy: 'capability', selected: choice.worker?.id ?? 'SELF', provider: choice.provider, requestedModel: choice.selection.model, requestedReasoningEffort: choice.selection.reasoningEffort,
182
- actualModel: result.tokens?.model, orchestratorProvider: options.parent, orchestratorModel: options.parentSelection?.model, orchestratorDurationMs: calls.filter(c => c.role === 'orchestrator').reduce((s, c) => s + c.durationMs, 0), workerDurationMs: calls[calls.length - 1].durationMs,
183
- processSuccess: result.success, verificationSuccess: verified, failureKind: failureKind ?? (result.infrastructureFailure ? 'infrastructure' : 'implementation'), usageAvailable: result.tokens !== undefined && result.tokens.measurementComplete !== false,
213
+ actualModel: result.tokens?.model, orchestratorProvider: options.planner?.agent ?? options.parent, orchestratorModel: (options.planner?.selection ?? options.parentSelection)?.model, orchestratorDurationMs: calls.filter(c => c.role === 'orchestrator').reduce((s, c) => s + c.durationMs, 0), workerDurationMs: calls[calls.length - 1].durationMs,
214
+ processSuccess: result.success, verificationSuccess: infrastructureFailure ? false : verified, failureKind: infrastructureFailure ? 'infrastructure' : failureKind ?? 'implementation', usageAvailable: result.tokens !== undefined && result.tokens.measurementComplete !== false,
184
215
  inputTokens: result.tokens?.inputTokens ?? 0, outputTokens: result.tokens?.outputTokens ?? 0, totalCostUsd: result.tokens?.totalCostUsd });
185
216
  } } };
186
217
  }
@@ -196,6 +227,8 @@ function routingSteps(options) {
196
227
  const ruleWorker = rule && failedStories.has(ctx.story.id) ? rule.escalateTo ?? 'SELF' : rule?.worker;
197
228
  const candidates = rule ? eligibleWorkers : rankWorkers(eligibleWorkers, options.strategy, options.maxCandidates);
198
229
  if (candidates.length === 0) {
230
+ if (options.fallback === 'block' || options.maxTier)
231
+ return blocked('No eligible routing profiles; parent fallback is disabled');
199
232
  return yield () => makeWorker(options.parent, options.parentSelection ?? {})(ctx);
200
233
  }
201
234
  const prompt = buildRoutingPrompt(ctx, candidates, options.strategy);
@@ -210,9 +243,11 @@ function routingSteps(options) {
210
243
  : routeRun.success ? parseRouteDecision(routeRun.output, candidates.map(worker => worker.id)) : null;
211
244
  const selected = decision?.worker ?? 'SELF';
212
245
  const worker = selected === 'SELF' ? undefined : candidates.find(candidate => candidate.id === selected);
246
+ if ((!worker && (options.fallback === 'block' || options.maxTier)) || (options.maxTier && (!worker?.tier || tiers.indexOf(worker.tier) > tiers.indexOf(options.maxTier))))
247
+ return blocked('Selected routing profile exceeds configured limits; execution blocked');
213
248
  const provider = worker?.agent ?? options.parent;
214
249
  const selection = worker
215
- ? { model: worker.model, reasoningEffort: worker.reasoningEffort, nativeMultiAgent: false, ...(provider !== 'gemini' ? { bare: options.parentSelection?.bare } : {}) }
250
+ ? { model: worker.model, reasoningEffort: worker.reasoningEffort, nativeMultiAgent: false, ...(provider !== 'gemini' && provider !== 'qwen' ? { bare: options.parentSelection?.bare } : {}) }
216
251
  : { ...(options.parentSelection ?? {}), nativeMultiAgent: false };
217
252
  const workerStarted = now();
218
253
  const result = yield () => makeWorker(provider, selection)(ctx);
@@ -4,7 +4,7 @@ import { detectHostAgent } from '../agents/host.js';
4
4
  import { loadConfig, saveConfig } from '../retrofit/config.js';
5
5
  import { detectProject } from '../retrofit/detect.js';
6
6
  import { runRetrofit } from '../retrofit/command.js';
7
- const ALL_AGENTS = ['claude', 'codex', 'gemini'];
7
+ const ALL_AGENTS = ['claude', 'codex', 'gemini', 'qwen'];
8
8
  export function defaultRoutingWorkers(agents) {
9
9
  const workers = {
10
10
  claude: [
@@ -25,6 +25,12 @@ export function defaultRoutingWorkers(agents) {
25
25
  { id: 'gemini-strong', agent: 'gemini', model: 'gemini-2.5-pro', tier: 'strong', costTier: 'medium', capabilities: ['debugging'] },
26
26
  { id: 'gemini-frontier', agent: 'gemini', model: 'gemini-2.5-pro', tier: 'frontier', costTier: 'high', capabilities: ['architecture'] },
27
27
  ],
28
+ qwen: [
29
+ { id: 'qwen-light', agent: 'qwen', model: 'qwen-turbo-latest', tier: 'light', costTier: 'low', capabilities: ['mechanical', 'tests'] },
30
+ { id: 'qwen-standard', agent: 'qwen', model: 'qwen3-coder-plus', tier: 'standard', costTier: 'medium', capabilities: ['implementation'] },
31
+ { id: 'qwen-strong', agent: 'qwen', model: 'qwen3-coder-plus', tier: 'strong', costTier: 'medium', capabilities: ['debugging'] },
32
+ { id: 'qwen-frontier', agent: 'qwen', model: 'qwen3-235b-a22b', tier: 'frontier', costTier: 'high', capabilities: ['architecture'] },
33
+ ],
28
34
  };
29
35
  return agents.flatMap(agent => workers[agent]);
30
36
  }
@@ -75,12 +81,12 @@ export async function runSetup(targetDir, opts = {}) {
75
81
  let decisionPolicy = defaultPolicy;
76
82
  let routing = defaultRouting;
77
83
  if (interactive && ask) {
78
- agents = parseAgents(await ask(`Agents [${defaultAgents.join(',')}] (claude,codex,gemini|all): `), defaultAgents);
84
+ agents = parseAgents(await ask(`Agents [${defaultAgents.join(',')}] (claude,codex,gemini,qwen|all): `), defaultAgents);
79
85
  const graphAnswer = (await ask(`Code graph [${defaultGraph}] (graphify|serena): `)).trim().toLowerCase();
80
86
  if (graphAnswer === 'graphify' || graphAnswer === 'serena')
81
87
  codeGraph = graphAnswer;
82
88
  loop = yes(await ask(`Enable autonomous loop? [${defaultLoop ? 'yes' : 'no'}]: `), defaultLoop);
83
- const runnerAnswer = (await ask(`Default runner [${runner}] (claude|codex|gemini): `)).trim().toLowerCase();
89
+ const runnerAnswer = (await ask(`Default runner [${runner}] (claude|codex|gemini|qwen): `)).trim().toLowerCase();
84
90
  if (ALL_AGENTS.includes(runnerAnswer))
85
91
  runner = runnerAnswer;
86
92
  const policyAnswer = (await ask(`Decision mode [${decisionPolicy}] (auto|critical): `)).trim().toLowerCase();
@@ -104,6 +110,8 @@ export async function runSetup(targetDir, opts = {}) {
104
110
  enabled: routing,
105
111
  strategy: opts.routingStrategy ?? config.routing?.strategy ?? 'capability',
106
112
  maxCandidates: config.routing?.maxCandidates ?? 3,
113
+ assessmentPolicy: existing?.routing?.assessmentPolicy ?? (existing ? 'on-demand' : 'prepared'),
114
+ fallback: existing?.routing?.fallback ?? (existing ? 'parent' : 'block'),
107
115
  ...(config.routing?.orchestrator ? { orchestrator: config.routing.orchestrator } : {}),
108
116
  workers: existingWorkers.length > 0 && !opts.routingPreset ? existingWorkers : defaultRoutingWorkers(agents),
109
117
  };
@@ -0,0 +1,67 @@
1
+ # Batch planning validation — 2026-09-06
2
+
3
+ AI-assisted implementation record for the local development build after 1.9.0.
4
+ No new version has been published by this task.
5
+
6
+ ## Implemented
7
+
8
+ - One bounded assessment call for a selected package, with exact output IDs,
9
+ executable-criteria checks, project locking and atomic PRD replacement.
10
+ - Separate planning provider/model/effort, complete draft/inbox assessments,
11
+ stale-contract detection including upstream requirements and the approved brief.
12
+ - Prepared routing and blocked fallback in new setups; existing settings remain
13
+ compatible. Automatic tier ceilings block insufficient configurations.
14
+ - Capability routing recognizes reported Windows process-creation/authentication
15
+ failures and failed gate evidence as infrastructure, stopping repair without
16
+ adding model-quality failures.
17
+
18
+ ## Actual Yoke run
19
+
20
+ The current compiled development CLI ran in a separate `Yoke-batch` checkout with
21
+ capability routing, prepared assessment, blocked fallback, a strong tier ceiling
22
+ and one serial worker. The assigned task was limited to batch-command tests.
23
+ Yoke chose `codex-standard`, requested `gpt-5.6-terra` at medium effort, verified
24
+ the criterion tests and committed the result. The main checkout received only
25
+ the reviewed test file; one assertion was strengthened during review.
26
+
27
+ Recorded start: 2026-09-06T17:45:54.729Z. Terminal state: complete at
28
+ 17:50:41.108Z, one backlog task accepted. One measured worker call and no
29
+ orchestrator/assessment call were recorded. Input: 546,591 tokens, including
30
+ 496,384 cached input tokens; output: 6,294 tokens. Reported monetary cost and
31
+ actual model identity are unknown. This is no cost benchmark or measured saving.
32
+
33
+ This run explicitly used unsafe permissions. It does not validate the Windows
34
+ safe-mode execution reported in issue #5.
35
+
36
+ ## Issue #5 boundary
37
+
38
+ **Follow-up:** the user subsequently requested completion of the runner correction.
39
+ The reproduction, safe-mode end-to-end result and supervision changes are recorded
40
+ in [WINDOWS-RUNNER-VALIDATION.md](WINDOWS-RUNNER-VALIDATION.md). The following paragraph
41
+ describes the earlier batch-planning checkpoint.
42
+
43
+ The user-provided runner handoff and [issue #5](https://github.com/HECer/yoke/issues/5)
44
+ were read. The infrastructure-classification correction is partial. The original
45
+ safe-mode shell failure has not been reproduced or diagnosed here. In particular,
46
+ a preflight under the actual sandbox identity, streamed tool-error handling when
47
+ a provider exits zero, separate heartbeat/useful-progress reporting, total process
48
+ budgets and argument-safe Windows launch regressions remain open. Optional MCP
49
+ startup warnings alone are not treated as proof of a task failure. No affected
50
+ DeviceLane worktree was restarted or cleaned, and the issue was not closed.
51
+
52
+ ## Evidence scope
53
+
54
+ Automated tests cover batch call suppression, invalid/partial/duplicate output,
55
+ concurrent PRD edits, contract invalidation, planner/worker separation, prepared
56
+ dispatch, tier limits and infrastructure repair suppression. Injected planner
57
+ responses establish command behavior; they are not authenticated provider parity
58
+ benchmarks. The full suite executed 1,166 tests in 127 files: 1,163 passed, two
59
+ were skipped and one new test fixture lacked a required configuration field.
60
+ After correcting that fixture, all 11 tests in its file passed. Subsequent
61
+ focused routing checks cover the final input-freshness guard. Planning hashes
62
+ establish input freshness, not provenance or correctness.
63
+
64
+ Read-only provenance scan: no C2PA located, supported scan complete, verification,
65
+ signer trust and Markdown metadata privacy unknown. The audit's limits are:
66
+ "No conforming verifier was supplied." and "Keyed model-level watermarks cannot
67
+ be checked without the provider's key." No authorship inference or mark removal.
@@ -1,54 +1,82 @@
1
- # Routing by task requirements
2
-
3
- Available in Yoke 1.9.0.
4
-
5
- New setups use `routing.strategy: capability`. Existing explicit strategies and profiles remain unchanged. To opt an existing project into capability routing with its current profiles:
6
-
7
- ```sh
8
- yoke setup . --yes --routing --routing-strategy=capability
9
- ```
10
-
11
- Give each existing worker a `tier: light|standard|strong|frontier`. Profiles without a tier remain usable with legacy strategies but are not candidates for capability selection. To explicitly replace worker profiles with the supplied provider presets, add `--routing-preset`. This replaces customized worker profiles; omit it to retain them.
12
-
13
- ## Planning and selection
14
-
15
- The configured start provider/model plans new change requests. The planner supplies an `assessment` with the task class, difficulty, uncertainty, risk, scope, testability, rationale and implementation approach. Existing tasks without an assessment receive one read-only assessment call using the start model. This call does not use a cheaper orchestration override. Its result is cached under `.yoke/routing/`, keyed by the task contract. Changing the contract invalidates the cached assessment; toggling `passes` does not.
16
-
17
- An assessment is a planning judgment, not a measured success probability. High testability means executable checks can detect an incorrect implementation. High uncertainty, architecture work or high risk require the frontier tier; difficult or broadly coupled work requires strong; routine implementation requires standard. Light is reserved for clear, low-risk mechanical work with strong checks. Weak testability raises the minimum tier. Reviews and critics have a standard minimum even for light tasks.
1
+ # Routing by task requirements
2
+
3
+ Capability routing is available in Yoke 1.9.0. Batch preparation, separate planning settings and routing limits described below are local, unreleased additions.
4
+
5
+ New setups use `routing.strategy: capability`. Existing explicit strategies and profiles remain unchanged. To opt an existing project into capability routing with its current profiles:
6
+
7
+ ```sh
8
+ yoke setup . --yes --routing --routing-strategy=capability
9
+ ```
10
+
11
+ Give each existing worker a `tier: light|standard|strong|frontier`. Profiles without a tier remain usable with legacy strategies but are not candidates for capability selection. To explicitly replace worker profiles with the supplied provider presets, add `--routing-preset`. This replaces customized worker profiles; omit it to retain them.
12
+
13
+ ## Planning and selection
14
+
15
+ The start provider/model remains the planning default. Optional `planning.agent`, `planning.model` and `planning.reasoningEffort` select a separate planner without changing the execution model. Draft and change-inbox planning request complete assessments in the same pass that creates the tasks. The inbox still performs its separate coverage review.
16
+
17
+ New setups use `routing.assessmentPolicy: prepared` and `routing.fallback: block`. Before dispatch, each unfinished task must have 2–5 executable criteria and a current assessment. No per-task planning call runs in this mode. Existing configurations retain `on-demand` and `parent` unless explicitly changed; on-demand routing makes a read-only planning call for an unassessed task and caches its result.
18
+
19
+ Run `yoke prd assess .` to assess all missing or stale unfinished tasks together. One invocation makes at most one read-only planner call, with exact task IDs and validated output, then atomically replaces the PRD. Invalid, incomplete or duplicate output leaves the PRD unchanged; concurrent input edits are preserved and the result is rejected. A current package uses zero model calls. `--story=ID` selects one unfinished task; `--reassess` also includes already-current assessments. The default package limit is 20 tasks (`planning.maxTasks`, range 1–50) and the prompt limit is 60,000 characters; oversized input is rejected before calling the provider.
20
+
21
+ Yoke writes `assessmentFor` bindings over requirements, declared write scope/provider, transitive dependency contracts and `.yoke/plan.md`. Changes invalidate affected unfinished tasks; progress or priority changes do not. A changed brief invalidates all unfinished tasks. These hashes detect stale input, not authorship or semantic correctness. Source-code changes outside these contracts require explicit reassessment when relevant.
22
+
23
+ Example settings alongside the project's existing worker profiles:
18
24
 
19
25
  ```yaml
20
- assessment:
21
- taskClass: implementation
22
- difficulty: medium
23
- uncertainty: low
24
- risk: low
25
- scope: low
26
- testability: high
27
- reason: Existing handler pattern and executable contract tests
28
- approach: Extend the handler, cover the boundary cases, run contract tests
26
+ runner:
27
+ agent: codex
28
+ model: gpt-5.6-terra
29
+ planning:
30
+ agent: codex
31
+ model: gpt-6-astra
32
+ reasoningEffort: high
33
+ maxTasks: 20
34
+ routing:
35
+ enabled: true
36
+ strategy: capability
37
+ assessmentPolicy: prepared
38
+ fallback: block
39
+ maxTier: strong
40
+ # Keep the existing workers list here.
29
41
  ```
30
42
 
31
- Yoke chooses an eligible profile at or above the required tier, then compares declared cost tiers. Optional `roles: [implementation, reviewer, critic, repair]` limits a profile's uses. Task `agent` affinity restricts implementation to that provider. Explicit routing rules and explicit quality role models retain precedence. A missing suitable profile falls back to the start model (or the explicitly bound provider's default) and labels the fallback; it does not prove that the fallback has sufficient capability. An invalid assessment blocks implementation.
32
-
33
- ## Initial profiles
34
-
35
- | Tier | Codex | Claude | Gemini |
36
- | --- | --- | --- | --- |
37
- | light | gpt-5.6-luna, low | haiku | gemini-2.5-flash |
38
- | standard | gpt-5.6-terra, medium | sonnet | gemini-2.5-pro |
39
- | strong | gpt-5.6-sol, high | sonnet, high effort | gemini-2.5-pro |
40
- | frontier | gpt-6-astra, high | opus | gemini-2.5-pro |
41
-
42
- These are editable starting hypotheses, not measured equivalences or price claims. The Codex names follow the requested profile family. Account access is not established by finding an installed CLI. Gemini uses documented explicit model IDs and receives no unsupported reasoning-effort parameter. Several Gemini tiers deliberately share Pro; moving between those tiers alone is not a stronger-model transition. Adjust the presets to the models available to your account. Claude aliases can resolve to different concrete models over time. Provider-reported model identity remains separate from requested identity.
43
-
44
- Provider references: [Claude model configuration](https://code.claude.com/docs/en/model-config), [Gemini model selection](https://geminicli.com/docs/cli/model/).
45
-
46
- ## Repair, escalation and evidence
47
-
48
- After an independent mechanical failure, capability routing permits one targeted repair at the initial tier, then raises the required tier on further failures. Attempts retain the current worktree and receive the previous gate findings. Every returned candidate still passes the normal acceptance, protection, quality, review and integration gates. Critical decisions, pause/cancellation and protected-acceptance violations stop retries. Provider process failures are classified conservatively as infrastructure; they do not count as evidence that a stronger model is needed.
49
-
50
- `routing.maxAttempts` limits implementation calls per unchanged task contract (default 5, configurable 1–8). The initial tier imposes an additional bound: light at most 5, standard 4, strong 3, frontier 2. An exhausted task blocks and requires a revised plan. These are inner implementation attempts; the outer loop's iteration count still counts task dispatches. Existing quality repair rounds and time limits remain separate bounds, and quality repairs can raise their profile tier by round. Goal execution keeps its existing global attempt, time and token budgets.
51
-
52
- Routing observations record task class, required tier, requested and reported models, effort, independent result, duration and available consumption. Selection considers matching project/task-class/tier history from the last 30 days within the bounded registry read. At least ten matching observations are required before an observed success rate below 80% excludes a profile. History is scoped to the concrete reported model to avoid mixing changed aliases. This is a conservative exclusion rule; it does not lower the planner's safety floor or claim calibrated probabilities. Missing usage remains unknown. Financial optimization and cross-provider performance require authenticated benchmarks.
53
-
54
- The dashboard's Now view shows the last recorded implementation profile, requested model/effort, rationale and next escalation tier. Usage & time retains reported model and role consumption, including assessment calls. Cached planning has no new model-call charge. Routing state is local runtime data and excluded from Yoke story commits.
43
+ `maxTier` limits automatic execution and escalation, including routing-rule selections. If a task needs frontier while the ceiling is strong, it blocks; Yoke does not lower the required capability. Planning itself may still use Astra. Explicit quality-role model overrides retain precedence. Goal execution retains its own protected manifest and budgets, uses the configured planner on demand, and honors routing fallback/tier limits; PRD preparation policy does not apply to synthetic goal tasks.
44
+
45
+ An assessment is a planning judgment, not a measured success probability. High testability means executable checks can detect an incorrect implementation. High uncertainty, architecture work or high risk require the frontier tier; difficult or broadly coupled work requires strong; routine implementation requires standard. Light is reserved for clear, low-risk mechanical work with strong checks. Weak testability raises the minimum tier. Reviews and critics have a standard minimum even for light tasks.
46
+
47
+ ```yaml
48
+ assessment:
49
+ taskClass: implementation
50
+ difficulty: medium
51
+ uncertainty: low
52
+ risk: low
53
+ scope: low
54
+ testability: high
55
+ reason: Existing handler pattern and executable contract tests
56
+ approach: Extend the handler, cover the boundary cases, run contract tests
57
+ ```
58
+
59
+ Yoke chooses an eligible profile at or above the required tier, then compares declared cost tiers. Optional `roles: [implementation, reviewer, critic, repair]` limits a profile's uses. Task `agent` affinity restricts implementation to that provider. Explicit routing rules and explicit quality role models retain precedence. With legacy `fallback: parent` and no tier ceiling, a missing suitable profile falls back to the start model (or the explicitly bound provider's default) and labels the fallback; it does not prove sufficient capability. `fallback: block` or a configured tier ceiling prevents that fallback. An invalid assessment blocks implementation.
60
+
61
+ ## Initial profiles
62
+
63
+ | Tier | Codex | Claude | Gemini |
64
+ | --- | --- | --- | --- |
65
+ | light | gpt-5.6-luna, low | haiku | gemini-2.5-flash |
66
+ | standard | gpt-5.6-terra, medium | sonnet | gemini-2.5-pro |
67
+ | strong | gpt-5.6-sol, high | sonnet, high effort | gemini-2.5-pro |
68
+ | frontier | gpt-6-astra, high | opus | gemini-2.5-pro |
69
+
70
+ These are editable starting hypotheses, not measured equivalences or price claims. The Codex names follow the requested profile family. Account access is not established by finding an installed CLI. Gemini uses documented explicit model IDs and receives no unsupported reasoning-effort parameter. Several Gemini tiers deliberately share Pro; moving between those tiers alone is not a stronger-model transition. Adjust the presets to the models available to your account. Claude aliases can resolve to different concrete models over time. Provider-reported model identity remains separate from requested identity.
71
+
72
+ Provider references: [Claude model configuration](https://code.claude.com/docs/en/model-config), [Gemini model selection](https://geminicli.com/docs/cli/model/).
73
+
74
+ ## Repair, escalation and evidence
75
+
76
+ After an independent mechanical failure, capability routing permits one targeted repair at the initial tier, then raises the required tier on further failures. Attempts retain the current worktree and receive the previous gate findings. Every returned candidate still passes the normal acceptance, protection, quality, review and integration gates. Critical decisions, pause/cancellation and protected-acceptance violations stop retries. Provider process failures are classified conservatively as infrastructure; they do not count as evidence that a stronger model is needed.
77
+
78
+ `routing.maxAttempts` limits implementation calls per unchanged task contract (default 5, configurable 1–8). The initial tier imposes an additional bound: light at most 5, standard 4, strong 3, frontier 2. An exhausted task blocks and requires a revised plan. These are inner implementation attempts; the outer loop's iteration count still counts task dispatches. Existing quality repair rounds and time limits remain separate bounds, and quality repairs can raise their profile tier by round. Goal execution keeps its existing global attempt, time and token budgets.
79
+
80
+ Routing observations record task class, required tier, requested and reported models, effort, independent result, duration and available consumption. Selection considers matching project/task-class/tier history from the last 30 days within the bounded registry read. At least ten matching observations are required before an observed success rate below 80% excludes a profile. History is scoped to the concrete reported model to avoid mixing changed aliases. This is a conservative exclusion rule; it does not lower the planner's safety floor or claim calibrated probabilities. Missing usage remains unknown. Financial optimization and cross-provider performance require authenticated benchmarks.
81
+
82
+ The dashboard's Now view shows the last recorded implementation profile, requested model/effort, rationale and next escalation tier. Usage & time retains reported model and role consumption, including assessment calls. Cached planning has no new model-call charge. Routing state is local runtime data and excluded from Yoke story commits.
@@ -0,0 +1,33 @@
1
+ # Dashboard evolution
2
+
3
+ The local dashboard is an actionable workspace for registered Yoke projects. It reads the same saved project, loop, goal, acceptance, and measurement data as the CLI. Project-controlled text is rendered through `textContent`, and the server remains bound to the loopback interface with same-origin authorization for pause requests.
4
+
5
+ ## Overview
6
+
7
+ Project status is derived from both the saved goal and the latest loop report. A blocked or running loop is not hidden by a completed goal. When goal and loop states differ, the card displays both. Blocked, failed, paused, unavailable, and stale active reports need attention; those projects are ordered before active projects, followed by the remaining projects.
8
+
9
+ Cards also show the reported current task and saved blocker reason when available, so the overview explains why a project needs attention before opening it.
10
+
11
+ An active loop report more than 20 minutes old is labeled **unconfirmed**. This means Yoke has an old active report, not evidence that the process is still live. The overview can be searched by project name, canonical path, or goal objective and filtered to All, Active, or Needs attention. A no-match state explains the result and provides a clear action that resets both search and filter.
12
+
13
+ ## Durable navigation
14
+
15
+ The URL hash stores the current screen, project, project tab, period, UTC grouping, and complete custom date range. Supported screens are the overview, workspace comparison, and project detail. Supported project tabs are Now, Usage & time, and Results; periods are 1, 7, 30, 90, or 365 days; groupings are day, week, or month. Custom dates must be real ISO calendar dates in chronological order and cover at most 366 inclusive days.
16
+
17
+ Invalid hash state returns to the overview with the 30-day/day defaults. Browser back and forward, a page reload, and Refresh restore the validated state. Refresh reloads data without resetting the selected view or controls. Starting any navigation aborts earlier fetches and changes a request generation, so an older response cannot replace the current screen.
18
+
19
+ The workspace comparison schedules at most three project analytics requests at once. If navigation changes, in-flight fetches are aborted and no additional obsolete project requests are scheduled. All projects and individual project links remain available in the navigation while viewing the comparison.
20
+
21
+ ## Usage comparisons
22
+
23
+ Usage & time compares the selected period with the immediately preceding period of exactly the same duration. A single time boundary is captured before either analytics request is made. The comparison covers recorded input plus output tokens, recorded acceptance events, and reported cost.
24
+
25
+ When the preceding value is zero, the dashboard describes no change or new recorded activity instead of calculating an infinite percentage. A cost percentage is shown only when both periods have fully measured cost. Otherwise the dashboard says the percentage is unavailable. Current and previous missing-usage coverage is displayed from calls with unknown usage and unmeasured attempts. Recorded portions remain recorded portions; missing tokens or charges are not estimated as zero.
26
+
27
+ Charts and their tables include periods that contain recorded events. Empty UTC calendar buckets are omitted and the chart explains that a gap means no recorded activity, not a measured zero. Detailed provider/model/role, model timeline, task, and phase tables remain below the comparison.
28
+
29
+ ## Design findings and limitations
30
+
31
+ Goal state and loop state describe different durable facts and need independent presentation. Freshness is also separate from state: a saved `running` value can become unconfirmed without being rewritten. Navigation state belongs in the URL because the dashboard has several independently useful views and time controls. Analytics fan-out needs cancellation as well as a concurrency bound because registered workspaces may contain many projects.
32
+
33
+ The dashboard combines local saved evidence with process-identity checks for supervised providers (see [Windows runner validation](WINDOWS-RUNNER-VALIDATION.md)). Unverified process identity remains unknown. It does not reconstruct activity that predates retained measurements, estimate missing provider usage or prices, or turn requested model names into proof of models used. Parallel call durations can overlap, so summed call time is consumption intensity rather than generation speed. The 20-minute freshness threshold is a presentation rule, not a process-health guarantee. No provider performance benchmark is implied by these views.
@@ -1,33 +1,33 @@
1
- # Migrating to Yoke 1.0
2
-
3
- Yoke 1.0 changes unsafe implicit behavior into explicit policy.
4
-
5
- ## Runner permissions
6
-
7
- The default is `runner.permissions: safe`. Automation that intentionally requires a full
8
- sandbox bypass must pass `--unsafe` or configure `runner.permissions: unsafe`. Use
9
- `read-only` for planning and probing.
10
-
11
- ## Reviews
12
-
13
- Reviewers write `.yoke/review-verdict.json`; Yoke validates and consumes it. The reviewer must
14
- differ from the implementer unless `--allow-self-review` is explicit. CI can add `--json`.
15
-
16
- ## Commit ownership
17
-
18
- Yoke resolves identity before implementation. Configure it when Git has no identity:
19
-
20
- ```yaml
21
- commit:
22
- authorName: HECer
23
- authorEmail: hec_er@web.de
24
- allowCoAuthors: false
25
- ```
26
-
27
- ## PRDs, audit, and cleanup
28
-
29
- Existing PRDs remain valid. Optional `needs`, `area`, and `agent` fields add dependencies,
30
- collision domains, and affinity. Enable the story audit gate with `audit.enabled: true` and
31
- version suppressions with `suppressionsVersion: 1`.
32
-
33
- `yoke loop cleanup` now reports retained worktrees. Add `--remove-worktrees` for deletion.
1
+ # Migrating to Yoke 1.0
2
+
3
+ Yoke 1.0 changes unsafe implicit behavior into explicit policy.
4
+
5
+ ## Runner permissions
6
+
7
+ The default is `runner.permissions: safe`. Automation that intentionally requires a full
8
+ sandbox bypass must pass `--unsafe` or configure `runner.permissions: unsafe`. Use
9
+ `read-only` for planning and probing.
10
+
11
+ ## Reviews
12
+
13
+ Reviewers write `.yoke/review-verdict.json`; Yoke validates and consumes it. The reviewer must
14
+ differ from the implementer unless `--allow-self-review` is explicit. CI can add `--json`.
15
+
16
+ ## Commit ownership
17
+
18
+ Yoke resolves identity before implementation. Configure it when Git has no identity:
19
+
20
+ ```yaml
21
+ commit:
22
+ authorName: HECer
23
+ authorEmail: hec_er@web.de
24
+ allowCoAuthors: false
25
+ ```
26
+
27
+ ## PRDs, audit, and cleanup
28
+
29
+ Existing PRDs remain valid. Optional `needs`, `area`, and `agent` fields add dependencies,
30
+ collision domains, and affinity. Enable the story audit gate with `audit.enabled: true` and
31
+ version suppressions with `suppressionsVersion: 1`.
32
+
33
+ `yoke loop cleanup` now reports retained worktrees. Add `--remove-worktrees` for deletion.
@@ -1,27 +1,27 @@
1
- # Migrating to Yoke 1.1
2
-
3
- Yoke 1.1 makes Claude, Codex, and Gemini share the same setup, planning, runner-selection, and decision behavior.
4
-
5
- Run the setup wizard once in an existing project:
6
-
7
- ```bash
8
- npx @hecer/yoke@1.1.0 setup .
9
- ```
10
-
11
- It preserves existing Yoke configuration and asks for target agents, code graph, loop state, default runner, and decision mode. A non-interactive agent can apply explicit choices with `--yes`.
12
-
13
- The new config fields are optional and backward-compatible:
14
-
15
- ```yaml
16
- runner:
17
- agent: codex
18
- permissions: safe
19
- loop:
20
- enabled: true
21
- timeoutMinutes: 30
22
- decisionPolicy: critical # auto | critical
23
- ```
24
-
25
- `loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` still work. Prefer `decisionPolicy` for new projects: `auto` resolves routine choices without asking; `critical` pauses only for high-impact decisions and continues after `yoke loop answer`. Resume restores the original isolation, review, runner, permissions, timeout, and policy flags instead of silently weakening the run. If the restart cannot begin, `yoke loop resume` retries from the request-bound state stored under Git's private state directory.
26
-
27
- Restart an already-open Codex task after retrofit if it does not discover the newly generated `.agents/skills/yoke-workflow/SKILL.md`.
1
+ # Migrating to Yoke 1.1
2
+
3
+ Yoke 1.1 makes Claude, Codex, and Gemini share the same setup, planning, runner-selection, and decision behavior.
4
+
5
+ Run the setup wizard once in an existing project:
6
+
7
+ ```bash
8
+ npx @hecer/yoke@1.1.0 setup .
9
+ ```
10
+
11
+ It preserves existing Yoke configuration and asks for target agents, code graph, loop state, default runner, and decision mode. A non-interactive agent can apply explicit choices with `--yes`.
12
+
13
+ The new config fields are optional and backward-compatible:
14
+
15
+ ```yaml
16
+ runner:
17
+ agent: codex
18
+ permissions: safe
19
+ loop:
20
+ enabled: true
21
+ timeoutMinutes: 30
22
+ decisionPolicy: critical # auto | critical
23
+ ```
24
+
25
+ `loop.onAmbiguity: resolve|abort` and `--on-ambiguity=` still work. Prefer `decisionPolicy` for new projects: `auto` resolves routine choices without asking; `critical` pauses only for high-impact decisions and continues after `yoke loop answer`. Resume restores the original isolation, review, runner, permissions, timeout, and policy flags instead of silently weakening the run. If the restart cannot begin, `yoke loop resume` retries from the request-bound state stored under Git's private state directory.
26
+
27
+ Restart an already-open Codex task after retrofit if it does not discover the newly generated `.agents/skills/yoke-workflow/SKILL.md`.