@hecer/yoke 1.3.0 → 1.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/CHANGELOG.md +33 -0
- package/README.md +108 -16
- package/TODOS.md +0 -3
- package/bench/README.md +55 -46
- package/bench/output-compaction.mjs +65 -0
- package/canon/loop/loop-spec.md +22 -8
- package/canon/manifest.yaml +1 -1
- package/dist/agents/contracts.js +50 -0
- package/dist/agents/process-incarnation.js +15 -0
- package/dist/agents/process-record.js +65 -0
- package/dist/agents/process-streams.js +40 -0
- package/dist/agents/process.js +177 -0
- package/dist/agents/providers.js +10 -7
- package/dist/agents/telemetry.js +62 -0
- package/dist/audit/command.js +13 -5
- package/dist/cli.js +55 -3
- package/dist/loop/candidate-boundaries.js +43 -0
- package/dist/loop/candidate-cleanup.js +98 -0
- package/dist/loop/candidate-contracts.js +1 -0
- package/dist/loop/candidate-selection.js +84 -0
- package/dist/loop/candidates.js +228 -0
- package/dist/loop/claim-lease.js +131 -0
- package/dist/loop/claims.js +177 -40
- package/dist/loop/cleanup.js +117 -15
- package/dist/loop/decision.js +31 -0
- package/dist/loop/dispatcher.js +334 -0
- package/dist/loop/git.js +6 -4
- package/dist/loop/loop.js +109 -16
- package/dist/loop/merge-queue.js +12 -6
- package/dist/loop/parallel-adapters.js +185 -0
- package/dist/loop/parallel-command.js +287 -0
- package/dist/loop/parallel.js +2 -4
- package/dist/loop/prd.js +4 -1
- package/dist/loop/reporter.js +86 -5
- package/dist/loop/run-command.js +216 -58
- package/dist/loop/runner.js +67 -32
- package/dist/loop/verify.js +65 -11
- package/dist/loop/watchdog.js +67 -8
- package/dist/loop/worker-cancellation.js +17 -0
- package/dist/loop/worker-cleanup.js +23 -0
- package/dist/loop/worker-contracts.js +1 -0
- package/dist/loop/worker.js +254 -0
- package/dist/output/artifact.js +63 -0
- package/dist/output/compact.js +192 -0
- package/dist/output/types.js +4 -0
- package/dist/quality/artifacts.js +59 -0
- package/dist/quality/candidate-comparison.js +130 -0
- package/dist/quality/command.js +316 -0
- package/dist/quality/loop.js +86 -0
- package/dist/quality/process-command.js +57 -0
- package/dist/quality/reference.js +187 -0
- package/dist/quality/repair.js +11 -0
- package/dist/quality/runner.js +66 -0
- package/dist/quality/types.js +60 -0
- package/dist/quality/verdict.js +142 -0
- package/dist/retrofit/config.js +26 -2
- package/dist/retrofit/gitignore.js +4 -0
- package/dist/review/command.js +27 -38
- package/dist/review/verdict.js +38 -7
- package/docs/MIGRATING-TO-1.4.md +70 -0
- package/docs/PUBLISHING.md +16 -2
- package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -0
- package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -0
- package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -0
- package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +181 -0
- package/gemini-extension.json +1 -1
- package/package.json +4 -3
|
@@ -0,0 +1,422 @@
|
|
|
1
|
+
# Gauntlet quality loop for Yoke
|
|
2
|
+
|
|
3
|
+
**Status:** Approved direction, ready for implementation planning
|
|
4
|
+
**Scope:** All six Gauntlet-inspired improvements, including an explicitly enabled unbounded quality mode
|
|
5
|
+
|
|
6
|
+
## Goal
|
|
7
|
+
|
|
8
|
+
Extend Yoke's mechanically gated story loop with comparative quality iteration without weakening
|
|
9
|
+
its existing authority model. Yoke remains the control plane: executable acceptance evidence,
|
|
10
|
+
project verification, performance, audit, independent review, worktree isolation, commit integrity,
|
|
11
|
+
locks, recovery, pause, and watchdog behavior remain authoritative.
|
|
12
|
+
|
|
13
|
+
The new quality layer adds:
|
|
14
|
+
|
|
15
|
+
1. a reviewer-to-repair loop;
|
|
16
|
+
2. optional external quality references;
|
|
17
|
+
3. blind binary comparison rather than drifting scores;
|
|
18
|
+
4. provider-backed parallel story execution;
|
|
19
|
+
5. ephemeral decomposition within a story; and
|
|
20
|
+
6. versioned provider and critic result contracts with provenance.
|
|
21
|
+
|
|
22
|
+
An explicit unbounded mode may remove the quality-loop round and elapsed-time limits. It does not
|
|
23
|
+
disable any mechanical gate or operational safety mechanism. In that mode the human is the only
|
|
24
|
+
quality brake, while process failures and Yoke safety gates can still block or pause the run.
|
|
25
|
+
|
|
26
|
+
## Non-goals
|
|
27
|
+
|
|
28
|
+
- Do not replace the PRD, acceptance criteria, or project verification with model judgment.
|
|
29
|
+
- Do not introduce a second `yoke gauntlet` state machine competing with `yoke loop`.
|
|
30
|
+
- Do not let ephemeral subtasks rewrite the authoritative PRD.
|
|
31
|
+
- Do not run parallel workers in a shared mutable worktree.
|
|
32
|
+
- Do not merge competing implementations of the same story together.
|
|
33
|
+
- Do not make subjective quality checks run by default for stories that do not declare them.
|
|
34
|
+
- Do not copy Gauntlet Loop prompt text. Reimplement the concepts in Yoke's native contracts.
|
|
35
|
+
|
|
36
|
+
## Compatibility
|
|
37
|
+
|
|
38
|
+
- Existing PRDs and `.yoke/config.yaml` files remain valid.
|
|
39
|
+
- Existing serial loop behavior remains the default.
|
|
40
|
+
- Quality iteration is enabled per story or by an explicit run override.
|
|
41
|
+
- `--parallel=1` remains equivalent to isolated serial execution. Values above one become available
|
|
42
|
+
only through the provider-backed dispatcher described below.
|
|
43
|
+
- Existing exit codes retain their meanings. Quality exhaustion and inconsistent judging are
|
|
44
|
+
ordinary blocked states and therefore return the existing blocked exit code.
|
|
45
|
+
- Existing standalone `yoke review` continues to work. Its verdict schema gains optional fields but
|
|
46
|
+
remains backward compatible with current verdict files.
|
|
47
|
+
|
|
48
|
+
## Authority and invariants
|
|
49
|
+
|
|
50
|
+
The gate order for every candidate is:
|
|
51
|
+
|
|
52
|
+
```text
|
|
53
|
+
implement or repair
|
|
54
|
+
-> structured criterion commands
|
|
55
|
+
-> project verify
|
|
56
|
+
-> performance gate when configured
|
|
57
|
+
-> audit gate when configured
|
|
58
|
+
-> quality challenge when enabled
|
|
59
|
+
-> independent code review when enabled
|
|
60
|
+
-> commit and passes:true atomically
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
The following invariants are absolute:
|
|
64
|
+
|
|
65
|
+
1. A quality verdict can reject mechanically green work but can never approve mechanically red work.
|
|
66
|
+
2. Every repair reruns the complete gate sequence from criterion evidence onward.
|
|
67
|
+
3. A fresh critic evaluates every new candidate. Previous critic explanations are not included in
|
|
68
|
+
the next critic prompt; only the repair worker receives the selected gap.
|
|
69
|
+
4. `passes: true` is written only after integrated-tree verification and a successful commit.
|
|
70
|
+
5. Unbounded mode removes only `maxRounds` and `maxMinutes` from the quality iteration. Watchdog,
|
|
71
|
+
pause, cleanup, locks, worktree isolation, provider failures, audit, and all verification remain.
|
|
72
|
+
6. Malformed, absent, inconsistent, or unverifiable quality evidence never becomes a pass.
|
|
73
|
+
|
|
74
|
+
## Configuration and PRD schema
|
|
75
|
+
|
|
76
|
+
### Project defaults
|
|
77
|
+
|
|
78
|
+
`.yoke/config.yaml` gains optional defaults:
|
|
79
|
+
|
|
80
|
+
```yaml
|
|
81
|
+
quality:
|
|
82
|
+
enabled: false
|
|
83
|
+
policy: blocking # blocking | advisory
|
|
84
|
+
maxRounds: 3
|
|
85
|
+
maxMinutes: 60
|
|
86
|
+
consistencyChecks: 2 # normal order plus label swap
|
|
87
|
+
maxParallelCandidates: 2
|
|
88
|
+
critic:
|
|
89
|
+
agent: codex # optional; independent resolution otherwise
|
|
90
|
+
model: gpt-5.6-sol # opaque provider model string
|
|
91
|
+
reasoningEffort: high
|
|
92
|
+
repair:
|
|
93
|
+
agent: claude # optional; story runner otherwise
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
Defaults apply only when a story declares a quality challenge or the run explicitly enables one.
|
|
97
|
+
`enabled: true` means stories with a complete `quality` declaration run it automatically; it does
|
|
98
|
+
not invent references for ordinary stories.
|
|
99
|
+
|
|
100
|
+
### Story declaration
|
|
101
|
+
|
|
102
|
+
`StorySchema` gains an optional `quality` object:
|
|
103
|
+
|
|
104
|
+
```yaml
|
|
105
|
+
- id: STORY-4
|
|
106
|
+
title: Polish the pricing page
|
|
107
|
+
priority: 4
|
|
108
|
+
acceptance: [...]
|
|
109
|
+
passes: false
|
|
110
|
+
quality:
|
|
111
|
+
reference:
|
|
112
|
+
name: Stripe pricing page
|
|
113
|
+
source: https://stripe.com/pricing
|
|
114
|
+
kind: url # url | file | command
|
|
115
|
+
digest: sha256:... # required for blocking file snapshots; acquired for URLs
|
|
116
|
+
candidate:
|
|
117
|
+
kind: screenshots # screenshots | files | command-output | benchmark
|
|
118
|
+
paths:
|
|
119
|
+
- .yoke/proof/STORY-4/desktop.png
|
|
120
|
+
- .yoke/proof/STORY-4/mobile.png
|
|
121
|
+
rubric: Visual hierarchy, clarity, responsive finish, and interaction quality
|
|
122
|
+
policy: blocking # optional override
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Validation requires a non-empty, named, fetchable, and comparable reference; a concrete candidate
|
|
126
|
+
artifact contract; and a non-empty rubric. Blocking comparisons cannot use mutable live content as
|
|
127
|
+
their only retained evidence. URL references are fetched during preflight, converted to inert
|
|
128
|
+
artifacts, hashed, timestamped, and stored under `.yoke/references/<digest>/`. The runtime artifact
|
|
129
|
+
is ignored by Git by default; the provenance record is written into story proof evidence.
|
|
130
|
+
|
|
131
|
+
`command` sources and candidate commands use argv arrays in the internal schema. The public YAML
|
|
132
|
+
may use a command string only through the same command validation policy as existing proof commands;
|
|
133
|
+
shell control operators are rejected.
|
|
134
|
+
|
|
135
|
+
### CLI overrides
|
|
136
|
+
|
|
137
|
+
`yoke loop run` gains:
|
|
138
|
+
|
|
139
|
+
```text
|
|
140
|
+
--quality enable declared quality challenges for this run
|
|
141
|
+
--no-quality disable them for this run
|
|
142
|
+
--quality-rounds=N override maxRounds
|
|
143
|
+
--quality-minutes=N override maxMinutes
|
|
144
|
+
--quality-unbounded remove quality round/time limits explicitly
|
|
145
|
+
--quality-policy=blocking|advisory
|
|
146
|
+
--parallel=N run independent stories concurrently
|
|
147
|
+
--candidates=N run competing candidates for one quality-enabled story
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
`--quality-unbounded` implies `--quality`, prints a prominent startup warning, and is recorded in
|
|
151
|
+
status/proof provenance. It conflicts with `--quality-rounds` and `--quality-minutes`. It does not
|
|
152
|
+
imply `--unsafe`, does not disable `--timeout`, and does not change `--parallel` or candidate count.
|
|
153
|
+
|
|
154
|
+
## 1. Reviewer-to-repair transition
|
|
155
|
+
|
|
156
|
+
Current reviewer rejection ends the story. The loop instead classifies the failure:
|
|
157
|
+
|
|
158
|
+
- mechanical gate failure: block immediately; no quality repair is attempted;
|
|
159
|
+
- quality challenge loss: select the verdict's single `biggestGap`;
|
|
160
|
+
- code-review rejection: select the highest-severity actionable blocking finding;
|
|
161
|
+
- malformed reviewer output or provider failure: block as infrastructure failure, not repair work.
|
|
162
|
+
|
|
163
|
+
When an actionable gap exists and the quality budget permits another round, Yoke launches a fresh
|
|
164
|
+
repair worker in the same isolated story worktree. Its prompt contains:
|
|
165
|
+
|
|
166
|
+
- story title and acceptance criteria;
|
|
167
|
+
- settled project context;
|
|
168
|
+
- the current diff;
|
|
169
|
+
- exactly one selected gap and its cited evidence;
|
|
170
|
+
- the instruction to make the smallest change that closes that gap;
|
|
171
|
+
- the immutable mechanical gate commands.
|
|
172
|
+
|
|
173
|
+
It does not receive critic praise, hidden labels, earlier critic deliberation, or unrelated findings.
|
|
174
|
+
After repair, criterion evidence is replaced with evidence from the new candidate and all gates rerun.
|
|
175
|
+
|
|
176
|
+
Default exhaustion is three repair rounds or 60 elapsed quality minutes per story, whichever occurs
|
|
177
|
+
first. Exhaustion writes a blocked status naming the remaining gap and the exhausted budget. In
|
|
178
|
+
unbounded mode, quality rounds continue until a consistent pass, human pause/stop, or a non-quality
|
|
179
|
+
failure blocks the story.
|
|
180
|
+
|
|
181
|
+
## 2. External quality references
|
|
182
|
+
|
|
183
|
+
A new `src/quality/` boundary owns reference acquisition, artifact preparation, comparison, and
|
|
184
|
+
evidence. It does not own story selection, commits, or process lifecycle.
|
|
185
|
+
|
|
186
|
+
Reference preflight occurs before implementation for a blocking challenge so an unreachable bar does
|
|
187
|
+
not waste a full agent run. Acquisition records:
|
|
188
|
+
|
|
189
|
+
- declared source and display name;
|
|
190
|
+
- resolved URL or path;
|
|
191
|
+
- acquisition timestamp;
|
|
192
|
+
- media/content type and byte size;
|
|
193
|
+
- SHA-256 digest;
|
|
194
|
+
- tool/adapter version;
|
|
195
|
+
- any viewport, benchmark, or normalization parameters.
|
|
196
|
+
|
|
197
|
+
Fetched content is untrusted data. HTML is rendered or converted to static evidence; scripts and text
|
|
198
|
+
from the reference are never inserted as instructions. Size and content-type limits apply. Critics
|
|
199
|
+
receive artifacts plus the trusted Yoke rubric in separate, explicitly delimited sections and run
|
|
200
|
+
read-only without source-tree write access.
|
|
201
|
+
|
|
202
|
+
For visual stories, existing `flow-smoke` screenshots are reused when paths match the declared
|
|
203
|
+
candidate contract. Quality comparison never triggers a second capture farm. Nonvisual adapters
|
|
204
|
+
compare declared files, command output, or benchmark records.
|
|
205
|
+
|
|
206
|
+
If a required reference cannot be acquired or its digest changes during the run, blocking policy
|
|
207
|
+
blocks before commit. Advisory policy records `skipped` with the exact cause and continues; skipped is
|
|
208
|
+
never reported as passed.
|
|
209
|
+
|
|
210
|
+
## 3. Blind binary comparison
|
|
211
|
+
|
|
212
|
+
Quality verdict schema version 1:
|
|
213
|
+
|
|
214
|
+
```json
|
|
215
|
+
{
|
|
216
|
+
"schemaVersion": 1,
|
|
217
|
+
"verdict": "candidate|reference",
|
|
218
|
+
"biggestGap": "single actionable gap",
|
|
219
|
+
"evidence": ["artifact-relative citation"],
|
|
220
|
+
"confidence": "high|medium|low",
|
|
221
|
+
"labels": { "candidate": "A", "reference": "B" },
|
|
222
|
+
"provenance": {
|
|
223
|
+
"provider": "codex",
|
|
224
|
+
"model": "gpt-5.6-sol",
|
|
225
|
+
"promptVersion": 1,
|
|
226
|
+
"rubricDigest": "sha256:...",
|
|
227
|
+
"referenceDigest": "sha256:...",
|
|
228
|
+
"candidateDigest": "sha256:..."
|
|
229
|
+
}
|
|
230
|
+
}
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
A blocking pass requires two fresh comparisons: one randomized A/B ordering and one label-swapped
|
|
234
|
+
ordering. Both must select the candidate. A disagreement, missing evidence, low-confidence verdict,
|
|
235
|
+
schema failure, or digest mismatch is `inconsistent`, not a pass. It blocks with retained evidence so
|
|
236
|
+
the user can rerun, change reviewer, or downgrade the story policy explicitly.
|
|
237
|
+
|
|
238
|
+
The critic names exactly one biggest remaining gap. Scores out of ten are not part of the contract.
|
|
239
|
+
Advisory mode records the same evidence but cannot prevent a mechanically valid story from landing.
|
|
240
|
+
|
|
241
|
+
## 4. Provider-backed parallel execution
|
|
242
|
+
|
|
243
|
+
The existing scheduler, claims, `runParallelLoop`, and `MergeQueue` become the production CLI path.
|
|
244
|
+
Parallel mode requires isolation.
|
|
245
|
+
|
|
246
|
+
### Independent stories
|
|
247
|
+
|
|
248
|
+
1. The dispatcher owns the existing loop lock.
|
|
249
|
+
2. Ready stories are selected by `needs`; duplicate active `area` values are excluded.
|
|
250
|
+
3. Each worker atomically claims one story and creates a project-owned worktree from the current HEAD.
|
|
251
|
+
4. It runs implementation and all per-story gates, including quality repair rounds.
|
|
252
|
+
5. A green worker enters the serialized merge queue.
|
|
253
|
+
6. The integrator rebases onto current HEAD, reruns criterion/project/perf/audit gates against the
|
|
254
|
+
integrated candidate, then integrates, updates context/PRD, and commits atomically.
|
|
255
|
+
7. A rebase conflict or integrated-tree failure releases the claim and reopens the story with a
|
|
256
|
+
structured reason. It never sets `passes: true`.
|
|
257
|
+
|
|
258
|
+
Provider subprocess workers expose cancellation and watchdog handles to scoped cleanup. Claim files
|
|
259
|
+
gain dispatcher ID, PID, base commit, worktree, provider/model, and heartbeat. Cleanup only removes
|
|
260
|
+
resources owned by the target project and dead dispatcher.
|
|
261
|
+
|
|
262
|
+
### Competing candidates for one story
|
|
263
|
+
|
|
264
|
+
`--candidates=N` creates N worktrees from the same base and runs independent implementations. Each
|
|
265
|
+
must pass mechanical gates. Quality comparison then selects one mechanically green candidate against
|
|
266
|
+
the reference; if more than one beats the reference, a final blind candidate-vs-candidate comparison
|
|
267
|
+
chooses the winner. Only the winning branch enters the merge queue. Losing candidates are retained as
|
|
268
|
+
proof metadata, then their worktrees are cleaned. Candidate branches are never combined.
|
|
269
|
+
|
|
270
|
+
Concurrency and candidate fan-out are separately limited. Defaults remain one story worker and one
|
|
271
|
+
candidate. Provider subprocess wiring ships only when cancellation, crash recovery, conflict,
|
|
272
|
+
integrated reverify, and cost telemetry tests pass end to end.
|
|
273
|
+
|
|
274
|
+
## 5. Ephemeral story decomposition
|
|
275
|
+
|
|
276
|
+
Before implementation, a worker may emit a versioned decomposition artifact:
|
|
277
|
+
|
|
278
|
+
```json
|
|
279
|
+
{
|
|
280
|
+
"schemaVersion": 1,
|
|
281
|
+
"storyId": "STORY-4",
|
|
282
|
+
"subtasks": [
|
|
283
|
+
{ "id": "hero", "goal": "Improve hero hierarchy", "area": "ui", "needs": [] },
|
|
284
|
+
{ "id": "mobile", "goal": "Polish mobile layout", "area": "ui-mobile", "needs": ["hero"] }
|
|
285
|
+
]
|
|
286
|
+
}
|
|
287
|
+
```
|
|
288
|
+
|
|
289
|
+
Subtasks are runtime scheduling hints under `.yoke/work/<story>/`; they are not PRD stories, cannot
|
|
290
|
+
change acceptance criteria, and cannot independently set `passes`. The parent story remains the only
|
|
291
|
+
commit and quality unit. Subtasks may run concurrently only when their declared areas differ and each
|
|
292
|
+
has its own nested worktree/branch or non-overlapping artifact workspace. Their results are synthesized
|
|
293
|
+
into one parent candidate before any gate runs.
|
|
294
|
+
|
|
295
|
+
Invalid decomposition falls back to one whole-story worker. Dynamic decomposition is optional and
|
|
296
|
+
must not add a controller call for small stories unless explicitly enabled or routing selects it.
|
|
297
|
+
|
|
298
|
+
## 6. Versioned provider and critic contracts
|
|
299
|
+
|
|
300
|
+
Provider adapters own every machine result. Add schemas for:
|
|
301
|
+
|
|
302
|
+
- route decisions;
|
|
303
|
+
- decomposition plans;
|
|
304
|
+
- review verdicts and selected repair gaps;
|
|
305
|
+
- quality comparison verdicts;
|
|
306
|
+
- candidate selection;
|
|
307
|
+
- provider telemetry envelopes.
|
|
308
|
+
|
|
309
|
+
Every envelope carries `schemaVersion`, provider, reported model when available, invocation role,
|
|
310
|
+
prompt/contract version, timing, permission profile, and usage fields. Unknown optional telemetry is
|
|
311
|
+
retained in raw per-call evidence but not promoted to aggregate fields without schema support.
|
|
312
|
+
|
|
313
|
+
Native provider output schemas are used when stable and available. Otherwise Yoke's existing
|
|
314
|
+
result-file transport is the compatibility path. Free-form stdout markers such as `YOKE_ROUTE` remain
|
|
315
|
+
temporary backward compatibility only; malformed output visibly falls back to the strong parent and
|
|
316
|
+
records `fallbackReason`, rather than silently disappearing.
|
|
317
|
+
|
|
318
|
+
All status NDJSON and proof evidence include quality round, candidate, critic, reference digest,
|
|
319
|
+
selected gap, verdict consistency, parallel worker, merge outcome, and cumulative token/cost data.
|
|
320
|
+
|
|
321
|
+
## State and observability
|
|
322
|
+
|
|
323
|
+
Runtime artifacts:
|
|
324
|
+
|
|
325
|
+
```text
|
|
326
|
+
.yoke/references/<digest>/... acquired inert reference
|
|
327
|
+
.yoke/work/<story>/decomposition.json ephemeral subtask plan
|
|
328
|
+
.yoke/proof/<story>/quality/round-<n>/ candidate and critic evidence
|
|
329
|
+
.yoke/proof/<story>/quality/summary.json final quality outcome
|
|
330
|
+
.yoke/claims/<story>.json parallel ownership/heartbeat
|
|
331
|
+
.yoke/loop-status.json current round/candidate/worker
|
|
332
|
+
.yoke/loop.log phase transitions and reasons
|
|
333
|
+
```
|
|
334
|
+
|
|
335
|
+
New reporter phases are `decomposing`, `quality-preflight`, `comparing`, `repairing`, `integrating`,
|
|
336
|
+
and `selecting-candidate`. Status includes bounded/unbounded mode, current/maximum quality round,
|
|
337
|
+
elapsed quality time, reference digest, active workers, and candidate count. Logs stay bounded under
|
|
338
|
+
the existing policy; structured proof artifacts are per-story and never inferred from prose.
|
|
339
|
+
|
|
340
|
+
Pause is honored at safe boundaries: between quality rounds and before launching new parallel workers.
|
|
341
|
+
An in-flight provider invocation is allowed to finish unless cleanup explicitly terminates its scoped
|
|
342
|
+
PID tree. In unbounded mode, `yoke loop status` always shows that only a human pause/stop can end
|
|
343
|
+
quality iteration.
|
|
344
|
+
|
|
345
|
+
## Error and recovery map
|
|
346
|
+
|
|
347
|
+
| Failure | Result |
|
|
348
|
+
|---|---|
|
|
349
|
+
| Mechanical criterion/verify/perf/audit failure | Block story; no quality repair |
|
|
350
|
+
| Quality critic says reference wins | Repair next round when budget allows |
|
|
351
|
+
| Code reviewer returns actionable blocking finding | Repair next round when budget allows |
|
|
352
|
+
| Missing/malformed verdict | Block as reviewer infrastructure failure |
|
|
353
|
+
| Label-swap disagreement or low confidence | Block as inconsistent judge evidence |
|
|
354
|
+
| Reference unavailable before work | Block, or advisory skip when explicitly configured |
|
|
355
|
+
| Reference digest changes mid-run | Block and retain both provenance records |
|
|
356
|
+
| Repair budget exhausted | Block with last gap and budget evidence |
|
|
357
|
+
| Human pause during unbounded mode | Finish current safe unit, persist state, exit paused |
|
|
358
|
+
| Provider crash/timeout | Block worker, preserve logs, release dead claim during cleanup |
|
|
359
|
+
| Parallel rebase conflict | Reopen story with conflict reason; no integration |
|
|
360
|
+
| Integrated-tree verification failure | Reopen story; candidate never passes |
|
|
361
|
+
| Candidate fan-out has no mechanically green candidate | Block without subjective comparison |
|
|
362
|
+
|
|
363
|
+
## Security
|
|
364
|
+
|
|
365
|
+
- Critics and reference evaluators run read-only.
|
|
366
|
+
- External reference content is never trusted as instructions.
|
|
367
|
+
- URL fetching uses protocol allowlists, redirect limits, byte limits, and loopback/private-network
|
|
368
|
+
restrictions unless the user explicitly supplies a local URL for a local project.
|
|
369
|
+
- Reference and candidate paths are normalized beneath Yoke-owned roots; traversal is rejected.
|
|
370
|
+
- Commands use validated argv contracts and existing approved-test restrictions where applicable.
|
|
371
|
+
- Unbounded mode requires the explicit CLI flag on every run; it is not persisted as a default.
|
|
372
|
+
- Startup output names effective provider, permissions, quality policy, budgets, and unbounded state.
|
|
373
|
+
- Token/cost growth is observable. Unbounded removes quality stop budgets by explicit request, but it
|
|
374
|
+
does not hide or stop reporting spend.
|
|
375
|
+
|
|
376
|
+
## Testing strategy
|
|
377
|
+
|
|
378
|
+
Development follows red-green-refactor. Add focused suites for:
|
|
379
|
+
|
|
380
|
+
1. Config and PRD compatibility, quality declaration validation, and CLI conflicts.
|
|
381
|
+
2. Repair-loop transitions, complete gate reruns, fresh reviewer use, and single-gap prompts.
|
|
382
|
+
3. Three-round/60-minute exhaustion and explicit unbounded behavior.
|
|
383
|
+
4. Reference acquisition, hashing, mutable-reference detection, inert conversion, path and network
|
|
384
|
+
security, and advisory-vs-blocking failures.
|
|
385
|
+
5. Randomized blind labels, mandatory label swap, inconsistent verdicts, schema failures, and proof
|
|
386
|
+
provenance.
|
|
387
|
+
6. Real provider subprocess workers using deterministic fake CLIs, including cancellation, timeout,
|
|
388
|
+
stale claims, process cleanup, and unavailable providers.
|
|
389
|
+
7. Parallel story scheduling, area exclusion, dependency ordering, rebase conflicts, integrated-tree
|
|
390
|
+
verification, and atomic PRD/context commits in temporary Git repositories.
|
|
391
|
+
8. Competing candidates: common base, mechanical filtering, winner-only merge, and losing cleanup.
|
|
392
|
+
9. Ephemeral decomposition validation, dependency/area scheduling, synthesis, fallback, and proof that
|
|
393
|
+
it cannot mutate PRD acceptance or pass state.
|
|
394
|
+
10. Versioned route/review/quality/telemetry envelopes and visible fallback reasons.
|
|
395
|
+
11. Reporter/NDJSON status for quality rounds, unbounded mode, parallel workers, and merge outcomes.
|
|
396
|
+
12. End-to-end serial, bounded-quality, unbounded-pause, parallel-story, and candidate-race flows.
|
|
397
|
+
|
|
398
|
+
Final verification runs TypeScript build, the complete Vitest suite, canon validation, docs metadata
|
|
399
|
+
check, package dry run, deterministic fake-provider flows, and authenticated opt-in smoke matrices for
|
|
400
|
+
Claude, Codex, and Gemini. Benchmarks compare quality off/on, bounded/unbounded-paused, serial/parallel,
|
|
401
|
+
and one/multiple candidates. Claims remain workload-specific and report unavailable provider evidence
|
|
402
|
+
honestly.
|
|
403
|
+
|
|
404
|
+
## Delivery order
|
|
405
|
+
|
|
406
|
+
1. Versioned result envelopes and review-gap selection.
|
|
407
|
+
2. Bounded reviewer-to-repair loop with status and evidence.
|
|
408
|
+
3. Reference/candidate schemas, acquisition, and blind consistency comparison.
|
|
409
|
+
4. CLI/config integration including explicit unbounded mode.
|
|
410
|
+
5. Provider subprocess workers wired to scheduler, claims, and merge queue.
|
|
411
|
+
6. Competing candidate mode.
|
|
412
|
+
7. Ephemeral decomposition and subtask scheduling.
|
|
413
|
+
8. Cross-provider benchmarks, documentation, migration notes, and release provenance updates.
|
|
414
|
+
|
|
415
|
+
Each stage preserves the serial non-quality path and lands only with its focused tests plus the full
|
|
416
|
+
existing suite green.
|
|
417
|
+
|
|
418
|
+
## Attribution
|
|
419
|
+
|
|
420
|
+
The comparative quality concepts are inspired by Matt Shumer's Gauntlet Loop technique and the
|
|
421
|
+
`robonuggets/gauntlet-loop` packaging. Yoke reimplements the concepts in its own mechanical control
|
|
422
|
+
plane. Documentation credits the inspiration; no CC BY prompt text is copied into runtime assets.
|
|
@@ -0,0 +1,181 @@
|
|
|
1
|
+
# Artifact-backed output compaction design
|
|
2
|
+
|
|
3
|
+
**Date:** 2026-08-16
|
|
4
|
+
**Status:** implemented, hardened, and verified on `main`
|
|
5
|
+
**Target:** Yoke 1.5.0
|
|
6
|
+
|
|
7
|
+
## Problem
|
|
8
|
+
|
|
9
|
+
Yoke already reduces shell noise through RTK and keeps loop prompts intentionally small. It does not, however, retain complete evidence from failed verify, criterion, performance, or completion commands. `commandVerifier` currently chooses stderr or stdout and keeps only the final five lines. That is token-cheap, but it can discard the first error, structured summaries, and the stdout half of a mixed failure.
|
|
10
|
+
|
|
11
|
+
Aphrodite demonstrates a useful pattern: keep a compact, type-aware preview in model-visible context and place the complete output behind a stable reference. Embedding Aphrodite itself is not a good fit because its primary integration is Hermes-specific, while Yoke must behave consistently across Claude Code, Codex, and Gemini and preserve its explicit fresh-context model.
|
|
12
|
+
|
|
13
|
+
## Goals
|
|
14
|
+
|
|
15
|
+
- Produce short, deterministic failure previews that retain the most actionable error and warning lines.
|
|
16
|
+
- Preserve complete stdout and stderr as a local artifact when output exceeds the inline budget.
|
|
17
|
+
- Put only the preview and a verifiable artifact reference into loop evidence and repair/review context.
|
|
18
|
+
- Use the same implementation for verify, executable acceptance criteria, performance, audit, and completion gates.
|
|
19
|
+
- Remain provider-independent and add no runtime dependency, daemon, database, proxy, or network request.
|
|
20
|
+
- Make savings and correctness measurable with Yoke's existing tests and benchmark schema.
|
|
21
|
+
|
|
22
|
+
## Non-goals
|
|
23
|
+
|
|
24
|
+
- Intercepting tool calls occurring inside Claude Code, Codex, or Gemini. Yoke cannot transparently alter those provider-internal streams.
|
|
25
|
+
- Replacing RTK, provider prompt caching, or Yoke's versioned context files.
|
|
26
|
+
- Semantic summarization by another model.
|
|
27
|
+
- Persisting interactive conversation memory or automatically injecting historical artifacts into later stories.
|
|
28
|
+
- Claiming Aphrodite's published compression ratios for Yoke.
|
|
29
|
+
|
|
30
|
+
## Chosen approach
|
|
31
|
+
|
|
32
|
+
Implement a small Yoke-native output subsystem with two isolated units:
|
|
33
|
+
|
|
34
|
+
1. `compactCommandOutput` is a pure deterministic function. It normalizes ANSI/control noise, classifies high-signal lines, removes repeated lines, and constructs a byte-bounded preview.
|
|
35
|
+
2. `writeOutputArtifact` stores the unmodified captured stdout and stderr in `.yoke/artifacts/` under a content-addressed filename and returns a relative path, byte count, and SHA-256 digest.
|
|
36
|
+
|
|
37
|
+
The gate runner combines both units. Small failures remain inline and create no artifact. Large failures include a compact preview followed by an artifact marker. Successful gates retain today's one-line summary and do not persist their output because success logs do not feed repair decisions.
|
|
38
|
+
|
|
39
|
+
This approach is preferred over:
|
|
40
|
+
|
|
41
|
+
- **Embedding Aphrodite:** stronger generic CCR machinery, but Hermes-oriented and operationally disproportionate for Yoke.
|
|
42
|
+
- **Adding a Chat Completions proxy:** could theoretically observe more traffic, but would couple Yoke to provider protocols, credentials, streaming semantics, and tool-call formats.
|
|
43
|
+
- **Only documenting Aphrodite as a companion:** zero maintenance, but provides no consistent behavior for Yoke's three supported providers and does not fix Yoke's current loss of gate evidence.
|
|
44
|
+
|
|
45
|
+
## Configuration
|
|
46
|
+
|
|
47
|
+
The feature is configured under an optional `output` block:
|
|
48
|
+
|
|
49
|
+
```yaml
|
|
50
|
+
output:
|
|
51
|
+
previewBytes: 2048
|
|
52
|
+
artifactThresholdBytes: 8192
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
Defaults apply when the block is absent:
|
|
56
|
+
|
|
57
|
+
- `previewBytes`: 2,048 bytes
|
|
58
|
+
- `artifactThresholdBytes`: 8,192 bytes
|
|
59
|
+
|
|
60
|
+
Both values are positive integers. `artifactThresholdBytes` must be greater than or equal to `previewBytes`. Existing configurations remain valid.
|
|
61
|
+
|
|
62
|
+
Artifact persistence is automatic only after the threshold is crossed. The artifacts directory is added to Yoke's managed `.gitignore` block. Files are created with user-only permissions where the platform supports POSIX modes. Yoke does not redact or transform the stored raw evidence, and documentation must therefore state that command output can contain secrets and must not be published blindly.
|
|
63
|
+
|
|
64
|
+
## Classification and preview rules
|
|
65
|
+
|
|
66
|
+
The compactor operates line-by-line without parsing project-specific formats:
|
|
67
|
+
|
|
68
|
+
1. Strip ANSI escape sequences and disallowed control characters from the preview only.
|
|
69
|
+
2. Mark case-insensitive error signals as highest priority: `error`, `failed`, `failure`, `fatal`, `panic`, `exception`, `traceback`, compiler error codes, and test failure markers.
|
|
70
|
+
3. Mark warnings as second priority: `warn`, `warning`, and deprecation notices.
|
|
71
|
+
4. Retain bounded context immediately around the first high-priority lines.
|
|
72
|
+
5. Retain the final non-empty lines because many test runners put totals and exit summaries at the end.
|
|
73
|
+
6. Remove duplicate preview lines while preserving their first selected order.
|
|
74
|
+
7. Enforce the preview byte budget at UTF-8 boundaries and append a deterministic omission line containing original line and byte counts.
|
|
75
|
+
|
|
76
|
+
The raw artifact contains the exact captured stdout and stderr with explicit stream headings.
|
|
77
|
+
Preview cleanup must never modify the artifact. Capture is bounded at 16 MiB per stream; if that
|
|
78
|
+
quota is exceeded, the command fails closed and the retained prefix is explicitly marked truncated.
|
|
79
|
+
|
|
80
|
+
## Artifact identity and layout
|
|
81
|
+
|
|
82
|
+
Artifacts use this layout:
|
|
83
|
+
|
|
84
|
+
```text
|
|
85
|
+
.yoke/artifacts/<story-or-session>/<phase>-<sha256-prefix>.log
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
- `story-or-session` is the sanitized `YOKE_STORY` value, falling back to `session`.
|
|
89
|
+
- `phase` is one of `criterion`, `verify`, `perf`, `audit`, or `completion`.
|
|
90
|
+
- The digest is computed from the complete artifact bytes; the filename uses a readable prefix while the marker records the full SHA-256.
|
|
91
|
+
- Repeating the same failure for the same story and phase resolves to the same path and content rather than generating timestamp noise.
|
|
92
|
+
- Paths returned to summaries are project-relative and use `/` separators so evidence is portable across Windows and Unix output.
|
|
93
|
+
|
|
94
|
+
The marker format is human-readable rather than a proprietary retrieval protocol:
|
|
95
|
+
|
|
96
|
+
```text
|
|
97
|
+
[full output: .yoke/artifacts/STORY/verify-0123abcd.log | 42,810 bytes | sha256:0123...]
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
Agents can retrieve it with their normal file-reading tools when the preview is insufficient.
|
|
101
|
+
|
|
102
|
+
## Data flow
|
|
103
|
+
|
|
104
|
+
1. A Yoke gate executes a configured command with stdout and stderr captured.
|
|
105
|
+
2. On exit zero, Yoke discards captured bytes and returns the existing compact success summary.
|
|
106
|
+
3. On timeout or non-zero exit, Yoke combines both streams with labels.
|
|
107
|
+
Capture quota overflow follows the same failure path but marks the retained prefix as truncated.
|
|
108
|
+
4. The pure compactor creates the bounded preview.
|
|
109
|
+
5. If the combined raw output crosses `artifactThresholdBytes`, Yoke writes the content-addressed artifact.
|
|
110
|
+
6. `VerifyResult.summary` carries the command, timeout state, preview, and optional artifact marker.
|
|
111
|
+
7. Existing worker evidence, status reporting, quality repair, and retry logic transport that bounded summary unchanged.
|
|
112
|
+
|
|
113
|
+
## Failure handling
|
|
114
|
+
|
|
115
|
+
- An artifact write failure must not hide the original gate failure or crash the loop. The summary keeps the compact preview and adds a bounded `artifact unavailable` reason.
|
|
116
|
+
- Invalid configuration is rejected by the existing Zod configuration boundary before any command runs.
|
|
117
|
+
- Empty command output produces the current command-only failure summary.
|
|
118
|
+
- stdout and stderr are both retained even when one is empty.
|
|
119
|
+
- Retried identical failures reuse the same artifact; a changed failure gets a different digest.
|
|
120
|
+
- Artifact paths are built from fixed directories and sanitized labels. User-controlled story IDs never become raw path segments.
|
|
121
|
+
- stdout and stderr capture is capped at 16 MiB per stream. Overflow fails closed, retains only the
|
|
122
|
+
bounded prefix, and uses a `truncated output` marker rather than a `full output` marker.
|
|
123
|
+
|
|
124
|
+
## Security and privacy
|
|
125
|
+
|
|
126
|
+
- `.yoke/artifacts/` is runtime state and must be gitignored by retrofit/setup.
|
|
127
|
+
- Artifacts never leave the local project through Yoke.
|
|
128
|
+
- No artifact content is injected automatically into prompts; only the bounded preview and reference are transported.
|
|
129
|
+
- The full digest lets reviewers verify that retrieved evidence matches the retained artifact bytes.
|
|
130
|
+
- Raw logs may contain credentials or personal data emitted by project commands. The README must warn users to inspect artifacts before sharing them.
|
|
131
|
+
- Absence of provenance metadata or detectable text markers must never be treated as evidence of human authorship.
|
|
132
|
+
|
|
133
|
+
## Testing strategy
|
|
134
|
+
|
|
135
|
+
Implementation follows red-green-refactor cycles:
|
|
136
|
+
|
|
137
|
+
- Pure compactor tests cover error prioritization, warning/context selection, duplicate removal, ANSI cleanup, UTF-8 byte bounds, tail summaries, empty input, and deterministic output.
|
|
138
|
+
- Artifact tests cover content fidelity, SHA-256 identity, stable paths, story sanitization, directory creation, and project-relative markers.
|
|
139
|
+
- Verifier integration tests prove that small failures remain inline, large failures create retrievable artifacts, mixed stdout/stderr is retained, successful commands create no artifacts, timeouts remain labelled, capture overflow fails closed as truncated, and artifact-write failures preserve the gate result.
|
|
140
|
+
- Configuration tests cover defaults, valid overrides, and the threshold invariant.
|
|
141
|
+
- Retrofit tests require `.yoke/artifacts/` in the managed ignore set.
|
|
142
|
+
- Existing loop, parallel-worker, quality, and retry suites must stay green.
|
|
143
|
+
|
|
144
|
+
## Measurement and release claims
|
|
145
|
+
|
|
146
|
+
Add a deterministic benchmark fixture that emits a large noisy failure containing an early actionable error and a final test summary. Record:
|
|
147
|
+
|
|
148
|
+
- raw byte and approximate token counts,
|
|
149
|
+
- preview byte and approximate token counts,
|
|
150
|
+
- compression ratio,
|
|
151
|
+
- whether the early error and final summary survived,
|
|
152
|
+
- whether the artifact digest round-trips.
|
|
153
|
+
|
|
154
|
+
The benchmark is for the Yoke-visible gate summary only. Release notes must not imply savings for provider-internal tool usage and must not reuse Aphrodite's reported ratios.
|
|
155
|
+
|
|
156
|
+
## Documentation and compatibility
|
|
157
|
+
|
|
158
|
+
- Add an output-compaction section to the README describing defaults, configuration, retrieval, privacy, and scope limitations.
|
|
159
|
+
- Add a changelog entry under the next unreleased section, without changing the already published `1.4.0` package version during implementation.
|
|
160
|
+
- Existing `Verifier` callers remain source-compatible. New options are threaded from the loaded Yoke config by `runLoopCommand`.
|
|
161
|
+
- No migration is required; projects receive the new gitignore line on their next setup/retrofit run.
|
|
162
|
+
|
|
163
|
+
## Acceptance criteria
|
|
164
|
+
|
|
165
|
+
1. A failed gate whose combined stdout/stderr is at most 8 KiB returns a deterministic preview and writes no artifact under default configuration.
|
|
166
|
+
2. A failed gate larger than 8 KiB but within the 16 MiB-per-stream capture quota writes the complete combined output below `.yoke/artifacts/` and returns a preview of at most 2 KiB plus a relative path, byte count, and full SHA-256 digest. Quota overflow fails closed and labels the bounded retained prefix as truncated.
|
|
167
|
+
3. The preview retains an early error and a final runner summary for the benchmark fixture.
|
|
168
|
+
4. Successful gates retain current summaries and write no artifacts.
|
|
169
|
+
5. Artifact failures do not change a gate's pass/fail result and cannot remove its inline preview.
|
|
170
|
+
6. `.yoke/artifacts/` is managed runtime state and is gitignored.
|
|
171
|
+
7. Configuration is validated and remains backward-compatible when `output` is absent.
|
|
172
|
+
8. The full test suite, lint, build, documentation check, package dry-run, audit, and benchmark verification complete successfully before release readiness is claimed.
|
|
173
|
+
|
|
174
|
+
## Implementation outcome
|
|
175
|
+
|
|
176
|
+
The implementation covers verify, executable-criterion, performance, completion, and configured
|
|
177
|
+
custom-audit commands. Yoke's structured built-in audit findings remain bounded by their existing
|
|
178
|
+
finding schema. The deterministic `gate-output-v1` fixture measured 26,699 raw bytes versus 470
|
|
179
|
+
bytes for the preview plus artifact reference (56.81×) while retaining its early compiler error,
|
|
180
|
+
final test summary, and SHA-256 round-trip. This is fixture-specific Yoke gate evidence, not a
|
|
181
|
+
provider-token or billing claim.
|
package/gemini-extension.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "yoke",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.5.0",
|
|
4
4
|
"description": "Cross-agent coding harness: curated skill canon, mechanical safety gates, autonomous loop with proof artifacts. CLI: npm i -g @hecer/yoke",
|
|
5
5
|
"contextFileName": "GEMINI-EXTENSION.md"
|
|
6
6
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@hecer/yoke",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.5.0",
|
|
4
4
|
"description": "One harness, three agents, zero trust in \"done\" — cross-agent coding harness for Claude Code, Codex CLI, and Gemini CLI: one skill canon, mechanical safety gates, an autonomous loop with screenshot/video proofs.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -14,8 +14,9 @@
|
|
|
14
14
|
"gemini-extension.json",
|
|
15
15
|
"agents",
|
|
16
16
|
"hooks",
|
|
17
|
-
"bench/README.md",
|
|
18
|
-
"bench/
|
|
17
|
+
"bench/README.md",
|
|
18
|
+
"bench/output-compaction.mjs",
|
|
19
|
+
"bench/RESULTS.md",
|
|
19
20
|
"bench/result-schema.mjs",
|
|
20
21
|
"bench/run.mjs",
|
|
21
22
|
"bench/run-large.mjs",
|