@hecer/yoke 1.3.0 → 1.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (69) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/.codex-plugin/plugin.json +1 -1
  3. package/CHANGELOG.md +33 -0
  4. package/README.md +108 -16
  5. package/TODOS.md +0 -3
  6. package/bench/README.md +55 -46
  7. package/bench/output-compaction.mjs +65 -0
  8. package/canon/loop/loop-spec.md +22 -8
  9. package/canon/manifest.yaml +1 -1
  10. package/dist/agents/contracts.js +50 -0
  11. package/dist/agents/process-incarnation.js +15 -0
  12. package/dist/agents/process-record.js +65 -0
  13. package/dist/agents/process-streams.js +40 -0
  14. package/dist/agents/process.js +177 -0
  15. package/dist/agents/providers.js +10 -7
  16. package/dist/agents/telemetry.js +62 -0
  17. package/dist/audit/command.js +13 -5
  18. package/dist/cli.js +55 -3
  19. package/dist/loop/candidate-boundaries.js +43 -0
  20. package/dist/loop/candidate-cleanup.js +98 -0
  21. package/dist/loop/candidate-contracts.js +1 -0
  22. package/dist/loop/candidate-selection.js +84 -0
  23. package/dist/loop/candidates.js +228 -0
  24. package/dist/loop/claim-lease.js +131 -0
  25. package/dist/loop/claims.js +177 -40
  26. package/dist/loop/cleanup.js +117 -15
  27. package/dist/loop/decision.js +31 -0
  28. package/dist/loop/dispatcher.js +334 -0
  29. package/dist/loop/git.js +6 -4
  30. package/dist/loop/loop.js +109 -16
  31. package/dist/loop/merge-queue.js +12 -6
  32. package/dist/loop/parallel-adapters.js +185 -0
  33. package/dist/loop/parallel-command.js +287 -0
  34. package/dist/loop/parallel.js +2 -4
  35. package/dist/loop/prd.js +4 -1
  36. package/dist/loop/reporter.js +86 -5
  37. package/dist/loop/run-command.js +216 -58
  38. package/dist/loop/runner.js +67 -32
  39. package/dist/loop/verify.js +65 -11
  40. package/dist/loop/watchdog.js +67 -8
  41. package/dist/loop/worker-cancellation.js +17 -0
  42. package/dist/loop/worker-cleanup.js +23 -0
  43. package/dist/loop/worker-contracts.js +1 -0
  44. package/dist/loop/worker.js +254 -0
  45. package/dist/output/artifact.js +63 -0
  46. package/dist/output/compact.js +192 -0
  47. package/dist/output/types.js +4 -0
  48. package/dist/quality/artifacts.js +59 -0
  49. package/dist/quality/candidate-comparison.js +130 -0
  50. package/dist/quality/command.js +316 -0
  51. package/dist/quality/loop.js +86 -0
  52. package/dist/quality/process-command.js +57 -0
  53. package/dist/quality/reference.js +187 -0
  54. package/dist/quality/repair.js +11 -0
  55. package/dist/quality/runner.js +66 -0
  56. package/dist/quality/types.js +60 -0
  57. package/dist/quality/verdict.js +142 -0
  58. package/dist/retrofit/config.js +26 -2
  59. package/dist/retrofit/gitignore.js +4 -0
  60. package/dist/review/command.js +27 -38
  61. package/dist/review/verdict.js +38 -7
  62. package/docs/MIGRATING-TO-1.4.md +70 -0
  63. package/docs/PUBLISHING.md +16 -2
  64. package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -0
  65. package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -0
  66. package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -0
  67. package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +181 -0
  68. package/gemini-extension.json +1 -1
  69. package/package.json +4 -3
@@ -0,0 +1,422 @@
1
+ # Gauntlet quality loop for Yoke
2
+
3
+ **Status:** Approved direction, ready for implementation planning
4
+ **Scope:** All six Gauntlet-inspired improvements, including an explicitly enabled unbounded quality mode
5
+
6
+ ## Goal
7
+
8
+ Extend Yoke's mechanically gated story loop with comparative quality iteration without weakening
9
+ its existing authority model. Yoke remains the control plane: executable acceptance evidence,
10
+ project verification, performance, audit, independent review, worktree isolation, commit integrity,
11
+ locks, recovery, pause, and watchdog behavior remain authoritative.
12
+
13
+ The new quality layer adds:
14
+
15
+ 1. a reviewer-to-repair loop;
16
+ 2. optional external quality references;
17
+ 3. blind binary comparison rather than drifting scores;
18
+ 4. provider-backed parallel story execution;
19
+ 5. ephemeral decomposition within a story; and
20
+ 6. versioned provider and critic result contracts with provenance.
21
+
22
+ An explicit unbounded mode may remove the quality-loop round and elapsed-time limits. It does not
23
+ disable any mechanical gate or operational safety mechanism. In that mode the human is the only
24
+ quality brake, while process failures and Yoke safety gates can still block or pause the run.
25
+
26
+ ## Non-goals
27
+
28
+ - Do not replace the PRD, acceptance criteria, or project verification with model judgment.
29
+ - Do not introduce a second `yoke gauntlet` state machine competing with `yoke loop`.
30
+ - Do not let ephemeral subtasks rewrite the authoritative PRD.
31
+ - Do not run parallel workers in a shared mutable worktree.
32
+ - Do not merge competing implementations of the same story together.
33
+ - Do not make subjective quality checks run by default for stories that do not declare them.
34
+ - Do not copy Gauntlet Loop prompt text. Reimplement the concepts in Yoke's native contracts.
35
+
36
+ ## Compatibility
37
+
38
+ - Existing PRDs and `.yoke/config.yaml` files remain valid.
39
+ - Existing serial loop behavior remains the default.
40
+ - Quality iteration is enabled per story or by an explicit run override.
41
+ - `--parallel=1` remains equivalent to isolated serial execution. Values above one become available
42
+ only through the provider-backed dispatcher described below.
43
+ - Existing exit codes retain their meanings. Quality exhaustion and inconsistent judging are
44
+ ordinary blocked states and therefore return the existing blocked exit code.
45
+ - Existing standalone `yoke review` continues to work. Its verdict schema gains optional fields but
46
+ remains backward compatible with current verdict files.
47
+
48
+ ## Authority and invariants
49
+
50
+ The gate order for every candidate is:
51
+
52
+ ```text
53
+ implement or repair
54
+ -> structured criterion commands
55
+ -> project verify
56
+ -> performance gate when configured
57
+ -> audit gate when configured
58
+ -> quality challenge when enabled
59
+ -> independent code review when enabled
60
+ -> commit and passes:true atomically
61
+ ```
62
+
63
+ The following invariants are absolute:
64
+
65
+ 1. A quality verdict can reject mechanically green work but can never approve mechanically red work.
66
+ 2. Every repair reruns the complete gate sequence from criterion evidence onward.
67
+ 3. A fresh critic evaluates every new candidate. Previous critic explanations are not included in
68
+ the next critic prompt; only the repair worker receives the selected gap.
69
+ 4. `passes: true` is written only after integrated-tree verification and a successful commit.
70
+ 5. Unbounded mode removes only `maxRounds` and `maxMinutes` from the quality iteration. Watchdog,
71
+ pause, cleanup, locks, worktree isolation, provider failures, audit, and all verification remain.
72
+ 6. Malformed, absent, inconsistent, or unverifiable quality evidence never becomes a pass.
73
+
74
+ ## Configuration and PRD schema
75
+
76
+ ### Project defaults
77
+
78
+ `.yoke/config.yaml` gains optional defaults:
79
+
80
+ ```yaml
81
+ quality:
82
+ enabled: false
83
+ policy: blocking # blocking | advisory
84
+ maxRounds: 3
85
+ maxMinutes: 60
86
+ consistencyChecks: 2 # normal order plus label swap
87
+ maxParallelCandidates: 2
88
+ critic:
89
+ agent: codex # optional; independent resolution otherwise
90
+ model: gpt-5.6-sol # opaque provider model string
91
+ reasoningEffort: high
92
+ repair:
93
+ agent: claude # optional; story runner otherwise
94
+ ```
95
+
96
+ Defaults apply only when a story declares a quality challenge or the run explicitly enables one.
97
+ `enabled: true` means stories with a complete `quality` declaration run it automatically; it does
98
+ not invent references for ordinary stories.
99
+
100
+ ### Story declaration
101
+
102
+ `StorySchema` gains an optional `quality` object:
103
+
104
+ ```yaml
105
+ - id: STORY-4
106
+ title: Polish the pricing page
107
+ priority: 4
108
+ acceptance: [...]
109
+ passes: false
110
+ quality:
111
+ reference:
112
+ name: Stripe pricing page
113
+ source: https://stripe.com/pricing
114
+ kind: url # url | file | command
115
+ digest: sha256:... # required for blocking file snapshots; acquired for URLs
116
+ candidate:
117
+ kind: screenshots # screenshots | files | command-output | benchmark
118
+ paths:
119
+ - .yoke/proof/STORY-4/desktop.png
120
+ - .yoke/proof/STORY-4/mobile.png
121
+ rubric: Visual hierarchy, clarity, responsive finish, and interaction quality
122
+ policy: blocking # optional override
123
+ ```
124
+
125
+ Validation requires a non-empty, named, fetchable, and comparable reference; a concrete candidate
126
+ artifact contract; and a non-empty rubric. Blocking comparisons cannot use mutable live content as
127
+ their only retained evidence. URL references are fetched during preflight, converted to inert
128
+ artifacts, hashed, timestamped, and stored under `.yoke/references/<digest>/`. The runtime artifact
129
+ is ignored by Git by default; the provenance record is written into story proof evidence.
130
+
131
+ `command` sources and candidate commands use argv arrays in the internal schema. The public YAML
132
+ may use a command string only through the same command validation policy as existing proof commands;
133
+ shell control operators are rejected.
134
+
135
+ ### CLI overrides
136
+
137
+ `yoke loop run` gains:
138
+
139
+ ```text
140
+ --quality enable declared quality challenges for this run
141
+ --no-quality disable them for this run
142
+ --quality-rounds=N override maxRounds
143
+ --quality-minutes=N override maxMinutes
144
+ --quality-unbounded remove quality round/time limits explicitly
145
+ --quality-policy=blocking|advisory
146
+ --parallel=N run independent stories concurrently
147
+ --candidates=N run competing candidates for one quality-enabled story
148
+ ```
149
+
150
+ `--quality-unbounded` implies `--quality`, prints a prominent startup warning, and is recorded in
151
+ status/proof provenance. It conflicts with `--quality-rounds` and `--quality-minutes`. It does not
152
+ imply `--unsafe`, does not disable `--timeout`, and does not change `--parallel` or candidate count.
153
+
154
+ ## 1. Reviewer-to-repair transition
155
+
156
+ Current reviewer rejection ends the story. The loop instead classifies the failure:
157
+
158
+ - mechanical gate failure: block immediately; no quality repair is attempted;
159
+ - quality challenge loss: select the verdict's single `biggestGap`;
160
+ - code-review rejection: select the highest-severity actionable blocking finding;
161
+ - malformed reviewer output or provider failure: block as infrastructure failure, not repair work.
162
+
163
+ When an actionable gap exists and the quality budget permits another round, Yoke launches a fresh
164
+ repair worker in the same isolated story worktree. Its prompt contains:
165
+
166
+ - story title and acceptance criteria;
167
+ - settled project context;
168
+ - the current diff;
169
+ - exactly one selected gap and its cited evidence;
170
+ - the instruction to make the smallest change that closes that gap;
171
+ - the immutable mechanical gate commands.
172
+
173
+ It does not receive critic praise, hidden labels, earlier critic deliberation, or unrelated findings.
174
+ After repair, criterion evidence is replaced with evidence from the new candidate and all gates rerun.
175
+
176
+ Default exhaustion is three repair rounds or 60 elapsed quality minutes per story, whichever occurs
177
+ first. Exhaustion writes a blocked status naming the remaining gap and the exhausted budget. In
178
+ unbounded mode, quality rounds continue until a consistent pass, human pause/stop, or a non-quality
179
+ failure blocks the story.
180
+
181
+ ## 2. External quality references
182
+
183
+ A new `src/quality/` boundary owns reference acquisition, artifact preparation, comparison, and
184
+ evidence. It does not own story selection, commits, or process lifecycle.
185
+
186
+ Reference preflight occurs before implementation for a blocking challenge so an unreachable bar does
187
+ not waste a full agent run. Acquisition records:
188
+
189
+ - declared source and display name;
190
+ - resolved URL or path;
191
+ - acquisition timestamp;
192
+ - media/content type and byte size;
193
+ - SHA-256 digest;
194
+ - tool/adapter version;
195
+ - any viewport, benchmark, or normalization parameters.
196
+
197
+ Fetched content is untrusted data. HTML is rendered or converted to static evidence; scripts and text
198
+ from the reference are never inserted as instructions. Size and content-type limits apply. Critics
199
+ receive artifacts plus the trusted Yoke rubric in separate, explicitly delimited sections and run
200
+ read-only without source-tree write access.
201
+
202
+ For visual stories, existing `flow-smoke` screenshots are reused when paths match the declared
203
+ candidate contract. Quality comparison never triggers a second capture farm. Nonvisual adapters
204
+ compare declared files, command output, or benchmark records.
205
+
206
+ If a required reference cannot be acquired or its digest changes during the run, blocking policy
207
+ blocks before commit. Advisory policy records `skipped` with the exact cause and continues; skipped is
208
+ never reported as passed.
209
+
210
+ ## 3. Blind binary comparison
211
+
212
+ Quality verdict schema version 1:
213
+
214
+ ```json
215
+ {
216
+ "schemaVersion": 1,
217
+ "verdict": "candidate|reference",
218
+ "biggestGap": "single actionable gap",
219
+ "evidence": ["artifact-relative citation"],
220
+ "confidence": "high|medium|low",
221
+ "labels": { "candidate": "A", "reference": "B" },
222
+ "provenance": {
223
+ "provider": "codex",
224
+ "model": "gpt-5.6-sol",
225
+ "promptVersion": 1,
226
+ "rubricDigest": "sha256:...",
227
+ "referenceDigest": "sha256:...",
228
+ "candidateDigest": "sha256:..."
229
+ }
230
+ }
231
+ ```
232
+
233
+ A blocking pass requires two fresh comparisons: one randomized A/B ordering and one label-swapped
234
+ ordering. Both must select the candidate. A disagreement, missing evidence, low-confidence verdict,
235
+ schema failure, or digest mismatch is `inconsistent`, not a pass. It blocks with retained evidence so
236
+ the user can rerun, change reviewer, or downgrade the story policy explicitly.
237
+
238
+ The critic names exactly one biggest remaining gap. Scores out of ten are not part of the contract.
239
+ Advisory mode records the same evidence but cannot prevent a mechanically valid story from landing.
240
+
241
+ ## 4. Provider-backed parallel execution
242
+
243
+ The existing scheduler, claims, `runParallelLoop`, and `MergeQueue` become the production CLI path.
244
+ Parallel mode requires isolation.
245
+
246
+ ### Independent stories
247
+
248
+ 1. The dispatcher owns the existing loop lock.
249
+ 2. Ready stories are selected by `needs`; duplicate active `area` values are excluded.
250
+ 3. Each worker atomically claims one story and creates a project-owned worktree from the current HEAD.
251
+ 4. It runs implementation and all per-story gates, including quality repair rounds.
252
+ 5. A green worker enters the serialized merge queue.
253
+ 6. The integrator rebases onto current HEAD, reruns criterion/project/perf/audit gates against the
254
+ integrated candidate, then integrates, updates context/PRD, and commits atomically.
255
+ 7. A rebase conflict or integrated-tree failure releases the claim and reopens the story with a
256
+ structured reason. It never sets `passes: true`.
257
+
258
+ Provider subprocess workers expose cancellation and watchdog handles to scoped cleanup. Claim files
259
+ gain dispatcher ID, PID, base commit, worktree, provider/model, and heartbeat. Cleanup only removes
260
+ resources owned by the target project and dead dispatcher.
261
+
262
+ ### Competing candidates for one story
263
+
264
+ `--candidates=N` creates N worktrees from the same base and runs independent implementations. Each
265
+ must pass mechanical gates. Quality comparison then selects one mechanically green candidate against
266
+ the reference; if more than one beats the reference, a final blind candidate-vs-candidate comparison
267
+ chooses the winner. Only the winning branch enters the merge queue. Losing candidates are retained as
268
+ proof metadata, then their worktrees are cleaned. Candidate branches are never combined.
269
+
270
+ Concurrency and candidate fan-out are separately limited. Defaults remain one story worker and one
271
+ candidate. Provider subprocess wiring ships only when cancellation, crash recovery, conflict,
272
+ integrated reverify, and cost telemetry tests pass end to end.
273
+
274
+ ## 5. Ephemeral story decomposition
275
+
276
+ Before implementation, a worker may emit a versioned decomposition artifact:
277
+
278
+ ```json
279
+ {
280
+ "schemaVersion": 1,
281
+ "storyId": "STORY-4",
282
+ "subtasks": [
283
+ { "id": "hero", "goal": "Improve hero hierarchy", "area": "ui", "needs": [] },
284
+ { "id": "mobile", "goal": "Polish mobile layout", "area": "ui-mobile", "needs": ["hero"] }
285
+ ]
286
+ }
287
+ ```
288
+
289
+ Subtasks are runtime scheduling hints under `.yoke/work/<story>/`; they are not PRD stories, cannot
290
+ change acceptance criteria, and cannot independently set `passes`. The parent story remains the only
291
+ commit and quality unit. Subtasks may run concurrently only when their declared areas differ and each
292
+ has its own nested worktree/branch or non-overlapping artifact workspace. Their results are synthesized
293
+ into one parent candidate before any gate runs.
294
+
295
+ Invalid decomposition falls back to one whole-story worker. Dynamic decomposition is optional and
296
+ must not add a controller call for small stories unless explicitly enabled or routing selects it.
297
+
298
+ ## 6. Versioned provider and critic contracts
299
+
300
+ Provider adapters own every machine result. Add schemas for:
301
+
302
+ - route decisions;
303
+ - decomposition plans;
304
+ - review verdicts and selected repair gaps;
305
+ - quality comparison verdicts;
306
+ - candidate selection;
307
+ - provider telemetry envelopes.
308
+
309
+ Every envelope carries `schemaVersion`, provider, reported model when available, invocation role,
310
+ prompt/contract version, timing, permission profile, and usage fields. Unknown optional telemetry is
311
+ retained in raw per-call evidence but not promoted to aggregate fields without schema support.
312
+
313
+ Native provider output schemas are used when stable and available. Otherwise Yoke's existing
314
+ result-file transport is the compatibility path. Free-form stdout markers such as `YOKE_ROUTE` remain
315
+ temporary backward compatibility only; malformed output visibly falls back to the strong parent and
316
+ records `fallbackReason`, rather than silently disappearing.
317
+
318
+ All status NDJSON and proof evidence include quality round, candidate, critic, reference digest,
319
+ selected gap, verdict consistency, parallel worker, merge outcome, and cumulative token/cost data.
320
+
321
+ ## State and observability
322
+
323
+ Runtime artifacts:
324
+
325
+ ```text
326
+ .yoke/references/<digest>/... acquired inert reference
327
+ .yoke/work/<story>/decomposition.json ephemeral subtask plan
328
+ .yoke/proof/<story>/quality/round-<n>/ candidate and critic evidence
329
+ .yoke/proof/<story>/quality/summary.json final quality outcome
330
+ .yoke/claims/<story>.json parallel ownership/heartbeat
331
+ .yoke/loop-status.json current round/candidate/worker
332
+ .yoke/loop.log phase transitions and reasons
333
+ ```
334
+
335
+ New reporter phases are `decomposing`, `quality-preflight`, `comparing`, `repairing`, `integrating`,
336
+ and `selecting-candidate`. Status includes bounded/unbounded mode, current/maximum quality round,
337
+ elapsed quality time, reference digest, active workers, and candidate count. Logs stay bounded under
338
+ the existing policy; structured proof artifacts are per-story and never inferred from prose.
339
+
340
+ Pause is honored at safe boundaries: between quality rounds and before launching new parallel workers.
341
+ An in-flight provider invocation is allowed to finish unless cleanup explicitly terminates its scoped
342
+ PID tree. In unbounded mode, `yoke loop status` always shows that only a human pause/stop can end
343
+ quality iteration.
344
+
345
+ ## Error and recovery map
346
+
347
+ | Failure | Result |
348
+ |---|---|
349
+ | Mechanical criterion/verify/perf/audit failure | Block story; no quality repair |
350
+ | Quality critic says reference wins | Repair next round when budget allows |
351
+ | Code reviewer returns actionable blocking finding | Repair next round when budget allows |
352
+ | Missing/malformed verdict | Block as reviewer infrastructure failure |
353
+ | Label-swap disagreement or low confidence | Block as inconsistent judge evidence |
354
+ | Reference unavailable before work | Block, or advisory skip when explicitly configured |
355
+ | Reference digest changes mid-run | Block and retain both provenance records |
356
+ | Repair budget exhausted | Block with last gap and budget evidence |
357
+ | Human pause during unbounded mode | Finish current safe unit, persist state, exit paused |
358
+ | Provider crash/timeout | Block worker, preserve logs, release dead claim during cleanup |
359
+ | Parallel rebase conflict | Reopen story with conflict reason; no integration |
360
+ | Integrated-tree verification failure | Reopen story; candidate never passes |
361
+ | Candidate fan-out has no mechanically green candidate | Block without subjective comparison |
362
+
363
+ ## Security
364
+
365
+ - Critics and reference evaluators run read-only.
366
+ - External reference content is never trusted as instructions.
367
+ - URL fetching uses protocol allowlists, redirect limits, byte limits, and loopback/private-network
368
+ restrictions unless the user explicitly supplies a local URL for a local project.
369
+ - Reference and candidate paths are normalized beneath Yoke-owned roots; traversal is rejected.
370
+ - Commands use validated argv contracts and existing approved-test restrictions where applicable.
371
+ - Unbounded mode requires the explicit CLI flag on every run; it is not persisted as a default.
372
+ - Startup output names effective provider, permissions, quality policy, budgets, and unbounded state.
373
+ - Token/cost growth is observable. Unbounded removes quality stop budgets by explicit request, but it
374
+ does not hide or stop reporting spend.
375
+
376
+ ## Testing strategy
377
+
378
+ Development follows red-green-refactor. Add focused suites for:
379
+
380
+ 1. Config and PRD compatibility, quality declaration validation, and CLI conflicts.
381
+ 2. Repair-loop transitions, complete gate reruns, fresh reviewer use, and single-gap prompts.
382
+ 3. Three-round/60-minute exhaustion and explicit unbounded behavior.
383
+ 4. Reference acquisition, hashing, mutable-reference detection, inert conversion, path and network
384
+ security, and advisory-vs-blocking failures.
385
+ 5. Randomized blind labels, mandatory label swap, inconsistent verdicts, schema failures, and proof
386
+ provenance.
387
+ 6. Real provider subprocess workers using deterministic fake CLIs, including cancellation, timeout,
388
+ stale claims, process cleanup, and unavailable providers.
389
+ 7. Parallel story scheduling, area exclusion, dependency ordering, rebase conflicts, integrated-tree
390
+ verification, and atomic PRD/context commits in temporary Git repositories.
391
+ 8. Competing candidates: common base, mechanical filtering, winner-only merge, and losing cleanup.
392
+ 9. Ephemeral decomposition validation, dependency/area scheduling, synthesis, fallback, and proof that
393
+ it cannot mutate PRD acceptance or pass state.
394
+ 10. Versioned route/review/quality/telemetry envelopes and visible fallback reasons.
395
+ 11. Reporter/NDJSON status for quality rounds, unbounded mode, parallel workers, and merge outcomes.
396
+ 12. End-to-end serial, bounded-quality, unbounded-pause, parallel-story, and candidate-race flows.
397
+
398
+ Final verification runs TypeScript build, the complete Vitest suite, canon validation, docs metadata
399
+ check, package dry run, deterministic fake-provider flows, and authenticated opt-in smoke matrices for
400
+ Claude, Codex, and Gemini. Benchmarks compare quality off/on, bounded/unbounded-paused, serial/parallel,
401
+ and one/multiple candidates. Claims remain workload-specific and report unavailable provider evidence
402
+ honestly.
403
+
404
+ ## Delivery order
405
+
406
+ 1. Versioned result envelopes and review-gap selection.
407
+ 2. Bounded reviewer-to-repair loop with status and evidence.
408
+ 3. Reference/candidate schemas, acquisition, and blind consistency comparison.
409
+ 4. CLI/config integration including explicit unbounded mode.
410
+ 5. Provider subprocess workers wired to scheduler, claims, and merge queue.
411
+ 6. Competing candidate mode.
412
+ 7. Ephemeral decomposition and subtask scheduling.
413
+ 8. Cross-provider benchmarks, documentation, migration notes, and release provenance updates.
414
+
415
+ Each stage preserves the serial non-quality path and lands only with its focused tests plus the full
416
+ existing suite green.
417
+
418
+ ## Attribution
419
+
420
+ The comparative quality concepts are inspired by Matt Shumer's Gauntlet Loop technique and the
421
+ `robonuggets/gauntlet-loop` packaging. Yoke reimplements the concepts in its own mechanical control
422
+ plane. Documentation credits the inspiration; no CC BY prompt text is copied into runtime assets.
@@ -0,0 +1,181 @@
1
+ # Artifact-backed output compaction design
2
+
3
+ **Date:** 2026-08-16
4
+ **Status:** implemented, hardened, and verified on `main`
5
+ **Target:** Yoke 1.5.0
6
+
7
+ ## Problem
8
+
9
+ Yoke already reduces shell noise through RTK and keeps loop prompts intentionally small. It does not, however, retain complete evidence from failed verify, criterion, performance, or completion commands. `commandVerifier` currently chooses stderr or stdout and keeps only the final five lines. That is token-cheap, but it can discard the first error, structured summaries, and the stdout half of a mixed failure.
10
+
11
+ Aphrodite demonstrates a useful pattern: keep a compact, type-aware preview in model-visible context and place the complete output behind a stable reference. Embedding Aphrodite itself is not a good fit because its primary integration is Hermes-specific, while Yoke must behave consistently across Claude Code, Codex, and Gemini and preserve its explicit fresh-context model.
12
+
13
+ ## Goals
14
+
15
+ - Produce short, deterministic failure previews that retain the most actionable error and warning lines.
16
+ - Preserve complete stdout and stderr as a local artifact when output exceeds the inline budget.
17
+ - Put only the preview and a verifiable artifact reference into loop evidence and repair/review context.
18
+ - Use the same implementation for verify, executable acceptance criteria, performance, audit, and completion gates.
19
+ - Remain provider-independent and add no runtime dependency, daemon, database, proxy, or network request.
20
+ - Make savings and correctness measurable with Yoke's existing tests and benchmark schema.
21
+
22
+ ## Non-goals
23
+
24
+ - Intercepting tool calls occurring inside Claude Code, Codex, or Gemini. Yoke cannot transparently alter those provider-internal streams.
25
+ - Replacing RTK, provider prompt caching, or Yoke's versioned context files.
26
+ - Semantic summarization by another model.
27
+ - Persisting interactive conversation memory or automatically injecting historical artifacts into later stories.
28
+ - Claiming Aphrodite's published compression ratios for Yoke.
29
+
30
+ ## Chosen approach
31
+
32
+ Implement a small Yoke-native output subsystem with two isolated units:
33
+
34
+ 1. `compactCommandOutput` is a pure deterministic function. It normalizes ANSI/control noise, classifies high-signal lines, removes repeated lines, and constructs a byte-bounded preview.
35
+ 2. `writeOutputArtifact` stores the unmodified captured stdout and stderr in `.yoke/artifacts/` under a content-addressed filename and returns a relative path, byte count, and SHA-256 digest.
36
+
37
+ The gate runner combines both units. Small failures remain inline and create no artifact. Large failures include a compact preview followed by an artifact marker. Successful gates retain today's one-line summary and do not persist their output because success logs do not feed repair decisions.
38
+
39
+ This approach is preferred over:
40
+
41
+ - **Embedding Aphrodite:** stronger generic CCR machinery, but Hermes-oriented and operationally disproportionate for Yoke.
42
+ - **Adding a Chat Completions proxy:** could theoretically observe more traffic, but would couple Yoke to provider protocols, credentials, streaming semantics, and tool-call formats.
43
+ - **Only documenting Aphrodite as a companion:** zero maintenance, but provides no consistent behavior for Yoke's three supported providers and does not fix Yoke's current loss of gate evidence.
44
+
45
+ ## Configuration
46
+
47
+ The feature is configured under an optional `output` block:
48
+
49
+ ```yaml
50
+ output:
51
+ previewBytes: 2048
52
+ artifactThresholdBytes: 8192
53
+ ```
54
+
55
+ Defaults apply when the block is absent:
56
+
57
+ - `previewBytes`: 2,048 bytes
58
+ - `artifactThresholdBytes`: 8,192 bytes
59
+
60
+ Both values are positive integers. `artifactThresholdBytes` must be greater than or equal to `previewBytes`. Existing configurations remain valid.
61
+
62
+ Artifact persistence is automatic only after the threshold is crossed. The artifacts directory is added to Yoke's managed `.gitignore` block. Files are created with user-only permissions where the platform supports POSIX modes. Yoke does not redact or transform the stored raw evidence, and documentation must therefore state that command output can contain secrets and must not be published blindly.
63
+
64
+ ## Classification and preview rules
65
+
66
+ The compactor operates line-by-line without parsing project-specific formats:
67
+
68
+ 1. Strip ANSI escape sequences and disallowed control characters from the preview only.
69
+ 2. Mark case-insensitive error signals as highest priority: `error`, `failed`, `failure`, `fatal`, `panic`, `exception`, `traceback`, compiler error codes, and test failure markers.
70
+ 3. Mark warnings as second priority: `warn`, `warning`, and deprecation notices.
71
+ 4. Retain bounded context immediately around the first high-priority lines.
72
+ 5. Retain the final non-empty lines because many test runners put totals and exit summaries at the end.
73
+ 6. Remove duplicate preview lines while preserving their first selected order.
74
+ 7. Enforce the preview byte budget at UTF-8 boundaries and append a deterministic omission line containing original line and byte counts.
75
+
76
+ The raw artifact contains the exact captured stdout and stderr with explicit stream headings.
77
+ Preview cleanup must never modify the artifact. Capture is bounded at 16 MiB per stream; if that
78
+ quota is exceeded, the command fails closed and the retained prefix is explicitly marked truncated.
79
+
80
+ ## Artifact identity and layout
81
+
82
+ Artifacts use this layout:
83
+
84
+ ```text
85
+ .yoke/artifacts/<story-or-session>/<phase>-<sha256-prefix>.log
86
+ ```
87
+
88
+ - `story-or-session` is the sanitized `YOKE_STORY` value, falling back to `session`.
89
+ - `phase` is one of `criterion`, `verify`, `perf`, `audit`, or `completion`.
90
+ - The digest is computed from the complete artifact bytes; the filename uses a readable prefix while the marker records the full SHA-256.
91
+ - Repeating the same failure for the same story and phase resolves to the same path and content rather than generating timestamp noise.
92
+ - Paths returned to summaries are project-relative and use `/` separators so evidence is portable across Windows and Unix output.
93
+
94
+ The marker format is human-readable rather than a proprietary retrieval protocol:
95
+
96
+ ```text
97
+ [full output: .yoke/artifacts/STORY/verify-0123abcd.log | 42,810 bytes | sha256:0123...]
98
+ ```
99
+
100
+ Agents can retrieve it with their normal file-reading tools when the preview is insufficient.
101
+
102
+ ## Data flow
103
+
104
+ 1. A Yoke gate executes a configured command with stdout and stderr captured.
105
+ 2. On exit zero, Yoke discards captured bytes and returns the existing compact success summary.
106
+ 3. On timeout or non-zero exit, Yoke combines both streams with labels.
107
+ Capture quota overflow follows the same failure path but marks the retained prefix as truncated.
108
+ 4. The pure compactor creates the bounded preview.
109
+ 5. If the combined raw output crosses `artifactThresholdBytes`, Yoke writes the content-addressed artifact.
110
+ 6. `VerifyResult.summary` carries the command, timeout state, preview, and optional artifact marker.
111
+ 7. Existing worker evidence, status reporting, quality repair, and retry logic transport that bounded summary unchanged.
112
+
113
+ ## Failure handling
114
+
115
+ - An artifact write failure must not hide the original gate failure or crash the loop. The summary keeps the compact preview and adds a bounded `artifact unavailable` reason.
116
+ - Invalid configuration is rejected by the existing Zod configuration boundary before any command runs.
117
+ - Empty command output produces the current command-only failure summary.
118
+ - stdout and stderr are both retained even when one is empty.
119
+ - Retried identical failures reuse the same artifact; a changed failure gets a different digest.
120
+ - Artifact paths are built from fixed directories and sanitized labels. User-controlled story IDs never become raw path segments.
121
+ - stdout and stderr capture is capped at 16 MiB per stream. Overflow fails closed, retains only the
122
+ bounded prefix, and uses a `truncated output` marker rather than a `full output` marker.
123
+
124
+ ## Security and privacy
125
+
126
+ - `.yoke/artifacts/` is runtime state and must be gitignored by retrofit/setup.
127
+ - Artifacts never leave the local project through Yoke.
128
+ - No artifact content is injected automatically into prompts; only the bounded preview and reference are transported.
129
+ - The full digest lets reviewers verify that retrieved evidence matches the retained artifact bytes.
130
+ - Raw logs may contain credentials or personal data emitted by project commands. The README must warn users to inspect artifacts before sharing them.
131
+ - Absence of provenance metadata or detectable text markers must never be treated as evidence of human authorship.
132
+
133
+ ## Testing strategy
134
+
135
+ Implementation follows red-green-refactor cycles:
136
+
137
+ - Pure compactor tests cover error prioritization, warning/context selection, duplicate removal, ANSI cleanup, UTF-8 byte bounds, tail summaries, empty input, and deterministic output.
138
+ - Artifact tests cover content fidelity, SHA-256 identity, stable paths, story sanitization, directory creation, and project-relative markers.
139
+ - Verifier integration tests prove that small failures remain inline, large failures create retrievable artifacts, mixed stdout/stderr is retained, successful commands create no artifacts, timeouts remain labelled, capture overflow fails closed as truncated, and artifact-write failures preserve the gate result.
140
+ - Configuration tests cover defaults, valid overrides, and the threshold invariant.
141
+ - Retrofit tests require `.yoke/artifacts/` in the managed ignore set.
142
+ - Existing loop, parallel-worker, quality, and retry suites must stay green.
143
+
144
+ ## Measurement and release claims
145
+
146
+ Add a deterministic benchmark fixture that emits a large noisy failure containing an early actionable error and a final test summary. Record:
147
+
148
+ - raw byte and approximate token counts,
149
+ - preview byte and approximate token counts,
150
+ - compression ratio,
151
+ - whether the early error and final summary survived,
152
+ - whether the artifact digest round-trips.
153
+
154
+ The benchmark is for the Yoke-visible gate summary only. Release notes must not imply savings for provider-internal tool usage and must not reuse Aphrodite's reported ratios.
155
+
156
+ ## Documentation and compatibility
157
+
158
+ - Add an output-compaction section to the README describing defaults, configuration, retrieval, privacy, and scope limitations.
159
+ - Add a changelog entry under the next unreleased section, without changing the already published `1.4.0` package version during implementation.
160
+ - Existing `Verifier` callers remain source-compatible. New options are threaded from the loaded Yoke config by `runLoopCommand`.
161
+ - No migration is required; projects receive the new gitignore line on their next setup/retrofit run.
162
+
163
+ ## Acceptance criteria
164
+
165
+ 1. A failed gate whose combined stdout/stderr is at most 8 KiB returns a deterministic preview and writes no artifact under default configuration.
166
+ 2. A failed gate larger than 8 KiB but within the 16 MiB-per-stream capture quota writes the complete combined output below `.yoke/artifacts/` and returns a preview of at most 2 KiB plus a relative path, byte count, and full SHA-256 digest. Quota overflow fails closed and labels the bounded retained prefix as truncated.
167
+ 3. The preview retains an early error and a final runner summary for the benchmark fixture.
168
+ 4. Successful gates retain current summaries and write no artifacts.
169
+ 5. Artifact failures do not change a gate's pass/fail result and cannot remove its inline preview.
170
+ 6. `.yoke/artifacts/` is managed runtime state and is gitignored.
171
+ 7. Configuration is validated and remains backward-compatible when `output` is absent.
172
+ 8. The full test suite, lint, build, documentation check, package dry-run, audit, and benchmark verification complete successfully before release readiness is claimed.
173
+
174
+ ## Implementation outcome
175
+
176
+ The implementation covers verify, executable-criterion, performance, completion, and configured
177
+ custom-audit commands. Yoke's structured built-in audit findings remain bounded by their existing
178
+ finding schema. The deterministic `gate-output-v1` fixture measured 26,699 raw bytes versus 470
179
+ bytes for the preview plus artifact reference (56.81×) while retaining its early compiler error,
180
+ final test summary, and SHA-256 round-trip. This is fixture-specific Yoke gate evidence, not a
181
+ provider-token or billing claim.
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "yoke",
3
- "version": "1.3.0",
3
+ "version": "1.5.0",
4
4
  "description": "Cross-agent coding harness: curated skill canon, mechanical safety gates, autonomous loop with proof artifacts. CLI: npm i -g @hecer/yoke",
5
5
  "contextFileName": "GEMINI-EXTENSION.md"
6
6
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@hecer/yoke",
3
- "version": "1.3.0",
3
+ "version": "1.5.0",
4
4
  "description": "One harness, three agents, zero trust in \"done\" — cross-agent coding harness for Claude Code, Codex CLI, and Gemini CLI: one skill canon, mechanical safety gates, an autonomous loop with screenshot/video proofs.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -14,8 +14,9 @@
14
14
  "gemini-extension.json",
15
15
  "agents",
16
16
  "hooks",
17
- "bench/README.md",
18
- "bench/RESULTS.md",
17
+ "bench/README.md",
18
+ "bench/output-compaction.mjs",
19
+ "bench/RESULTS.md",
19
20
  "bench/result-schema.mjs",
20
21
  "bench/run.mjs",
21
22
  "bench/run-large.mjs",