muse-crew 0.13.2 → 0.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/decisions/AGENTS.md +94 -0
- package/docs/decisions/publish-path.md +1280 -0
- package/docs/decisions/qa-reproduce.md +486 -0
- package/docs/decisions/workflow-core.md +234 -0
- package/docs/publish-unknown-recovery.md +113 -0
- package/docs/visual-verdict.md +6 -7
- package/lib/AGENTS.md +3 -1
- package/lib/classify-publish-absence.js +451 -0
- package/lib/compose-evidence-caption.js +36 -5
- package/lib/crew-api.js +584 -22
- package/lib/crew-release.sh +51 -9
- package/lib/publish-content.js +154 -0
- package/lib/retry-publish.js +370 -0
- package/lib/verify-publish.js +101 -84
- package/package.json +2 -2
- package/seed/cron-body-template.md +57 -2
- package/workflows/AGENTS.md +3 -1
- package/workflows/bugfix.js +301 -592
- package/workflows/chore.js +292 -526
- package/workflows/standard.js +294 -550
- package/workflows/upgrade.js +1 -1
|
@@ -0,0 +1,486 @@
|
|
|
1
|
+
# Decision history: qa reproduce
|
|
2
|
+
|
|
3
|
+
Relocated from workflow source comments during H5 (2026-09-18). The workflows keep only the relied-upon invariant inline; the full decision history lives here.
|
|
4
|
+
|
|
5
|
+
<a id="reproduce-wrong-layer-guard"></a>
|
|
6
|
+
## Reproduce wrong-layer guard
|
|
7
|
+
|
|
8
|
+
Invariant: for engine/docs tasks the experiential see-act loop is forbidden; the workflow checks the repro dir mechanically for frame artifacts, attempt-scoped; no baseline is inconclusive fail-closed.
|
|
9
|
+
|
|
10
|
+
Applies to: standard, bugfix, chore.
|
|
11
|
+
|
|
12
|
+
```
|
|
13
|
+
// Reproduce wrong-layer guard (2026-09-16): for engine/docs tasks the
|
|
14
|
+
// experiential see-act loop is forbidden — a reproducer that drives the
|
|
15
|
+
// browser burns the run and emits a wrong-layer verdict (three tasks
|
|
16
|
+
// parked on exactly this in the 2026-09-17 UTC shift). The workflow
|
|
17
|
+
// checks the repro evidence dir mechanically for see-act frame artifacts
|
|
18
|
+
// (the driver's archive naming: <n>-<action>-<desktop|mobile>.png).
|
|
19
|
+
// Attempt-scoped (2026-09-17): the after-listing is diffed against the
|
|
20
|
+
// baseline taken before this phase attempt's dispatch loop — only frames
|
|
21
|
+
// CREATED BY the current attempt count as a violation; stale frames from
|
|
22
|
+
// earlier attempts appear in both snapshots and are excluded. Without a
|
|
23
|
+
// baseline the guard is inconclusive and fails closed for retry (passing
|
|
24
|
+
// could route a genuinely wrong-layer verdict to Map).
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
<a id="baseline-evidence-capture"></a>
|
|
28
|
+
## Baseline evidence capture
|
|
29
|
+
|
|
30
|
+
Invariant: capture baseline evidence for experiential tasks before the QA agent runs; the baseline is the control.
|
|
31
|
+
|
|
32
|
+
Applies to: standard.
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
// ── Capture: baseline evidence for experiential tasks ─────────────
|
|
36
|
+
// Hazel's QA capture pass runs right after Triage, before Map, for tasks
|
|
37
|
+
// Sage flagged experiential. The capture itself is parent-driven (the
|
|
38
|
+
// inspection handoff arrives at the root agent, outside this script), so
|
|
39
|
+
// when no baseline evidence is recorded yet the script logs a note event
|
|
40
|
+
// and parks with the exact parent protocol + resume path. Never fails the
|
|
41
|
+
// task over missing evidence: after two requests, baseline:none is
|
|
42
|
+
// recorded and final QA judges on the rubric alone.
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Applies to: bugfix.
|
|
46
|
+
|
|
47
|
+
```
|
|
48
|
+
// ── Capture: baseline evidence for experiential tasks ─────────────
|
|
49
|
+
// Hazel's QA capture pass runs right after Triage, before Map, for tasks
|
|
50
|
+
// Sage flagged experiential. The capture itself is parent-driven (the
|
|
51
|
+
// inspection handoff arrives at the root agent, outside this script), so
|
|
52
|
+
// when no baseline evidence is recorded yet the script logs a note event
|
|
53
|
+
// and parks with the exact parent protocol + resume path. Never fails the
|
|
54
|
+
// task over missing evidence: after two requests, baseline:none is
|
|
55
|
+
// recorded and final QA judges on the rubric alone.
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
Applies to: chore.
|
|
59
|
+
|
|
60
|
+
```
|
|
61
|
+
// ── Capture: baseline evidence for experiential tasks ─────────────
|
|
62
|
+
// Hazel's QA capture pass runs right after Triage, before Map, for tasks
|
|
63
|
+
// Sage flagged experiential. The capture itself is parent-driven (the
|
|
64
|
+
// inspection handoff arrives at the root agent, outside this script), so
|
|
65
|
+
// when no baseline evidence is recorded yet the script logs a note event
|
|
66
|
+
// and parks with the exact parent protocol + resume path. Never fails the
|
|
67
|
+
// task over missing evidence: after two requests, baseline:none is
|
|
68
|
+
// recorded and the task continues without baseline comparison.
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
|
|
72
|
+
<a id="pre-publish-pin-guard"></a>
|
|
73
|
+
## Pre-Publish pin guard
|
|
74
|
+
|
|
75
|
+
Invariant: the pin step materializes the required scripts before Publish; the agent cannot self-certify.
|
|
76
|
+
|
|
77
|
+
Applies to: standard, bugfix, chore.
|
|
78
|
+
|
|
79
|
+
```
|
|
80
|
+
// The seven basenames the pin step must materialize — asserted mechanically
|
|
81
|
+
// by workflow code from the verbatim listing, never from agent prose.
|
|
82
|
+
// COMPUTE_DIFF is the deterministic publish-diff computer (room #14,
|
|
83
|
+
// 2026-09-17): the diff is computed by this script, never ferried as an
|
|
84
|
+
// agent JSON string. Pinned like the other publish-critical modules so a
|
|
85
|
+
// mid-run release swap cannot change it under the workflow.
|
|
86
|
+
// CLASSIFY_SURFACE is the surface classifier (room #15, 2026-09-18):
|
|
87
|
+
// crew-api.js statically imports it, so the pin must carry it — a pin
|
|
88
|
+
// without it kills every claim with ERR_MODULE_NOT_FOUND.
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
<a id="qa-verdict-gate"></a>
|
|
92
|
+
## QA verdict.json closeout gate
|
|
93
|
+
|
|
94
|
+
Invariant: the prose VERDICT: line drives the closeout gate via extractVerdict; a bare FAIL with an all-positive report is a defect.
|
|
95
|
+
|
|
96
|
+
Applies to: standard, bugfix, chore.
|
|
97
|
+
|
|
98
|
+
```
|
|
99
|
+
// QA verdict.json closeout gate (task 30dceb78): the prose VERDICT: line is
|
|
100
|
+
// the routing signal, but verdict.json is the reason-carrying record the
|
|
101
|
+
// workflow actually reads. A FAIL verdict must carry a machine-readable
|
|
102
|
+
// reason; the workflow refuses to route to rework on an unreasoned or
|
|
103
|
+
// contradictory verdict. The deterministic cross-checker
|
|
104
|
+
// (lib/read-ooda-verdict.js) runs against the prose verdict: a missing,
|
|
105
|
+
// corrupt, contradictory, or reason-less record fails the phase for retry
|
|
106
|
+
// — the dispatcher re-runs QA at the same step under its
|
|
107
|
+
// consecutive-failure cap — instead of routing to rework. Without this
|
|
108
|
+
// gate, a bare VERDICT: FAIL with an all-positive report (canary
|
|
109
|
+
// 2026-09-15, task 1d692d91) rebuilt nothing and parked at Publish on an
|
|
110
|
+
// unobserved artifact build. The gate applies only to the experiential QA
|
|
111
|
+
// path (qaArtifact / qaTerminal), the paths that write verdict.json.
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
<a id="experiential-loop-guard"></a>
|
|
115
|
+
## QA experiential-loop guard
|
|
116
|
+
|
|
117
|
+
Invariant: a PASS verdict with missing experiential evidence never completes — parks fail-closed with unattributable_reason.
|
|
118
|
+
|
|
119
|
+
Applies to: standard.
|
|
120
|
+
|
|
121
|
+
```
|
|
122
|
+
// QA experiential-loop guard (clean-room defect 2026-09-16): Hazel's verdict.json
|
|
123
|
+
// is honest about missing experiential evidence, but the closeout treated a PASS
|
|
124
|
+
// as terminal done even when the experiential loop never ran (playwright-core
|
|
125
|
+
// was unresolvable from the release layout — the dependency lived in the
|
|
126
|
+
// npm install dir, severed from the crew home). A PASS verdict with missing
|
|
127
|
+
// experiential evidence must never be terminal: the task parks fail-closed
|
|
128
|
+
// with unattributable_reason=qa-visual-loop-unavailable (artifact surface)
|
|
129
|
+
// or qa-terminal-loop-unavailable (terminal surface) instead of
|
|
130
|
+
// transitioning to done. Code, not prompt text: the check reads the
|
|
131
|
+
// machine-readable verdict via lib/read-ooda-verdict.js, which reports
|
|
132
|
+
// visual_loop_unavailable from the OODA log's NOT POSSIBLE browser steps
|
|
133
|
+
// and terminal_loop_unavailable from NOT POSSIBLE terminal steps, plus the
|
|
134
|
+
// verdict's missing_evidence tool-unavailability notes.
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
Applies to: bugfix.
|
|
138
|
+
|
|
139
|
+
```
|
|
140
|
+
// QA experiential-loop guard (clean-room defect 2026-09-16): Hazel's
|
|
141
|
+
// verdict.json is honest about missing experiential evidence, but the
|
|
142
|
+
// closeout treated a PASS as terminal done even when the experiential
|
|
143
|
+
// loop never ran. A PASS verdict with missing experiential evidence must
|
|
144
|
+
// never be terminal: the task parks fail-closed with
|
|
145
|
+
// unattributable_reason=qa-visual-loop-unavailable (artifact surface)
|
|
146
|
+
// or qa-terminal-loop-unavailable (terminal surface) instead of
|
|
147
|
+
// transitioning to done. Code, not prompt text: the loop-availability
|
|
148
|
+
// flags come from lib/read-ooda-verdict.js (the OODA log's NOT POSSIBLE
|
|
149
|
+
// steps and the verdict's missing_evidence tool-unavailability notes),
|
|
150
|
+
// already parsed into qaVerdictGate above.
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
|
|
154
|
+
<a id="reproduce-baseline"></a>
|
|
155
|
+
## Reproduce wrong-layer guard baseline
|
|
156
|
+
|
|
157
|
+
Invariant: snapshot the repro evidence dir before the Reproduce agent runs; only frames created by the current attempt count; null baseline is inconclusive fail-closed.
|
|
158
|
+
|
|
159
|
+
Applies to: standard, bugfix, chore.
|
|
160
|
+
|
|
161
|
+
```
|
|
162
|
+
// Reproduce wrong-layer guard baseline (2026-09-17): snapshot the repro
|
|
163
|
+
// evidence dir listing BEFORE the Reproduce work agent runs. The guard
|
|
164
|
+
// after the run diffs against this baseline — only frames created by the
|
|
165
|
+
// current attempt count as a wrong-layer violation; stale frames from
|
|
166
|
+
// pre-fix attempts appear in both snapshots and are excluded (set diff,
|
|
167
|
+
// no wall-clock, deterministic). One snapshot per phase attempt covers
|
|
168
|
+
// all transport retries inside the dispatch loop. null = baseline
|
|
169
|
+
// unavailable; the guard then fails closed for retry.
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
<a id="visual-verdict-ownership"></a>
|
|
173
|
+
## Visual verdict ownership
|
|
174
|
+
|
|
175
|
+
Invariant: Hazel owns the visual verdict experientially (see-act loop + verdict.json); her prose VERDICT drives passed; the old parent note gate is deleted.
|
|
176
|
+
|
|
177
|
+
Applies to: standard, bugfix, chore.
|
|
178
|
+
|
|
179
|
+
```
|
|
180
|
+
// Visual verdict ownership (2026-09-15): Hazel owns the visual verdict
|
|
181
|
+
// experientially — the QA instructions above have her drive the see-act
|
|
182
|
+
// loop herself and record verdict.json + the append-only verdicts.jsonl.
|
|
183
|
+
// Her prose VERDICT: line drives `passed` via extractVerdict; a FAIL
|
|
184
|
+
// bounces to Build through the standard Review/QA rejection path below.
|
|
185
|
+
// The old parent-recorded note gate is deleted.
|
|
186
|
+
```
|
|
187
|
+
|
|
188
|
+
<a id="baseline-evidence-status"></a>
|
|
189
|
+
## Baseline evidence status
|
|
190
|
+
|
|
191
|
+
Invariant: reads the task's note events for baseline status.
|
|
192
|
+
|
|
193
|
+
Applies to: standard, bugfix, chore.
|
|
194
|
+
|
|
195
|
+
```
|
|
196
|
+
// Baseline evidence status: reads the task's note events for the exact
|
|
197
|
+
// protocol prefixes (explicit state, never English matching). Returns
|
|
198
|
+
// { baseline_found, baseline_kind, baseline_refs, requested_count, evidence_count }.
|
|
199
|
+
// found = any note starting exactly "baseline: captured" or "baseline:
|
|
200
|
+
// none" (kind/refs come from the LATEST such message); requested_count =
|
|
201
|
+
// the number of notes starting exactly "baseline: requested";
|
|
202
|
+
// evidence_count = the number of notes starting exactly "baseline: captured" or
|
|
203
|
+
// "baseline: none" (the per-attempt evidence chain; state-derived, no wall clock). Each call
|
|
204
|
+
// uses a fresh key: the evidence changes between calls (the parent logs
|
|
205
|
+
// the capture while this run is parked), so a cached replay would lie.
|
|
206
|
+
```
|
|
207
|
+
|
|
208
|
+
<a id="pin-basenames"></a>
|
|
209
|
+
## Pin basenames
|
|
210
|
+
|
|
211
|
+
Invariant: the pin step materializes the required scripts.
|
|
212
|
+
|
|
213
|
+
Applies to: standard, bugfix, chore.
|
|
214
|
+
|
|
215
|
+
```
|
|
216
|
+
// The seven basenames the pin step must materialize — asserted mechanically
|
|
217
|
+
// by workflow code from the verbatim listing, never from agent prose.
|
|
218
|
+
// COMPUTE_DIFF is the deterministic publish-diff computer (room #14,
|
|
219
|
+
// 2026-09-17): the diff is computed by this script, never ferried as an
|
|
220
|
+
// agent JSON string. Pinned like the other publish-critical modules so a
|
|
221
|
+
// mid-run release swap cannot change it under the workflow.
|
|
222
|
+
// CLASSIFY_SURFACE is the surface classifier (room #15, 2026-09-18):
|
|
223
|
+
// crew-api.js statically imports it, so the pin must carry it — a pin
|
|
224
|
+
// without it kills every claim with ERR_MODULE_NOT_FOUND.
|
|
225
|
+
```
|
|
226
|
+
|
|
227
|
+
<a id="visual-protocol-avail"></a>
|
|
228
|
+
## Visual protocol availability
|
|
229
|
+
|
|
230
|
+
Invariant: the workflow parks if the visual protocol is unavailable.
|
|
231
|
+
|
|
232
|
+
Applies to: standard.
|
|
233
|
+
|
|
234
|
+
```
|
|
235
|
+
// Visual verdict protocol availability — the workflow parks for parent-run
|
|
236
|
+
// baseline capture and visual verdict ONLY when the protocol is fully
|
|
237
|
+
// shipped. The protocol requires docs/visual-verdict.md in the release AND
|
|
238
|
+
// the parent-side capture tooling (task b309a97d, "QA owns the visual
|
|
239
|
+
// verdict"). Until both exist, the parks would deadlock waiting for a
|
|
240
|
+
// parent who cannot fulfill them.
|
|
241
|
+
// Effective value for this run, resolved by the dispatcher from the
|
|
242
|
+
// project's visual_protocol setting (null=inherits crew default=off).
|
|
243
|
+
// Manual launches without the arg default to off (previous behavior).
|
|
244
|
+
```
|
|
245
|
+
|
|
246
|
+
Applies to: bugfix.
|
|
247
|
+
|
|
248
|
+
```
|
|
249
|
+
// Visual verdict protocol availability — the workflow parks for parent-run
|
|
250
|
+
// baseline capture and visual verdict ONLY when the protocol is fully
|
|
251
|
+
// shipped. The protocol requires docs/visual-verdict.md in the release AND
|
|
252
|
+
// the parent-side capture tooling (task b309a97d, "QA owns the visual
|
|
253
|
+
// verdict"). Until both exist, the parks would deadlock waiting for a
|
|
254
|
+
// parent who cannot fulfill them.
|
|
255
|
+
// Effective value for this run, resolved by the dispatcher from the
|
|
256
|
+
// project's visual_protocol setting (null=inherits crew default=off).
|
|
257
|
+
// Manual launches without the arg default to off (previous behavior).
|
|
258
|
+
```
|
|
259
|
+
|
|
260
|
+
Applies to: chore.
|
|
261
|
+
|
|
262
|
+
```
|
|
263
|
+
// Visual verdict protocol availability — the workflow parks for parent-run
|
|
264
|
+
// baseline capture ONLY when the protocol is fully shipped. The protocol
|
|
265
|
+
// requires docs/visual-verdict.md in the release AND the parent-side
|
|
266
|
+
// capture tooling (task b309a97d, "QA owns the visual verdict"). Until both
|
|
267
|
+
// exist, the park would deadlock waiting for a parent who cannot fulfill
|
|
268
|
+
// it.
|
|
269
|
+
// Effective value for this run, resolved by the dispatcher from the
|
|
270
|
+
// project's visual_protocol setting (null=inherits crew default=off).
|
|
271
|
+
// Manual launches without the arg default to off (previous behavior).
|
|
272
|
+
```
|
|
273
|
+
|
|
274
|
+
|
|
275
|
+
<a id="capture-baseline"></a>
|
|
276
|
+
## Capture baseline
|
|
277
|
+
|
|
278
|
+
Invariant: baseline evidence for experiential tasks is captured before work begins.
|
|
279
|
+
|
|
280
|
+
Applies to: standard.
|
|
281
|
+
|
|
282
|
+
```
|
|
283
|
+
// ── Capture: baseline evidence for experiential tasks ─────────────
|
|
284
|
+
// Hazel's QA capture pass runs right after Triage, before Map, for tasks
|
|
285
|
+
// Sage flagged experiential. The capture itself is parent-driven (the
|
|
286
|
+
// inspection handoff arrives at the root agent, outside this script), so
|
|
287
|
+
// when no baseline evidence is recorded yet the script logs a note event
|
|
288
|
+
// and parks with the exact parent protocol + resume path. Never fails the
|
|
289
|
+
// task over missing evidence: after two requests, baseline:none is
|
|
290
|
+
// recorded and final QA judges on the rubric alone.
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
Applies to: bugfix.
|
|
294
|
+
|
|
295
|
+
```
|
|
296
|
+
// ── Capture: baseline evidence for experiential tasks ─────────────
|
|
297
|
+
// Hazel's QA capture pass runs right after Triage, before Map, for tasks
|
|
298
|
+
// Sage flagged experiential. The capture itself is parent-driven (the
|
|
299
|
+
// inspection handoff arrives at the root agent, outside this script), so
|
|
300
|
+
// when no baseline evidence is recorded yet the script logs a note event
|
|
301
|
+
// and parks with the exact parent protocol + resume path. Never fails the
|
|
302
|
+
// task over missing evidence: after two requests, baseline:none is
|
|
303
|
+
// recorded and final QA judges on the rubric alone.
|
|
304
|
+
```
|
|
305
|
+
|
|
306
|
+
Applies to: chore.
|
|
307
|
+
|
|
308
|
+
```
|
|
309
|
+
// ── Capture: baseline evidence for experiential tasks ─────────────
|
|
310
|
+
// Hazel's QA capture pass runs right after Triage, before Map, for tasks
|
|
311
|
+
// Sage flagged experiential. The capture itself is parent-driven (the
|
|
312
|
+
// inspection handoff arrives at the root agent, outside this script), so
|
|
313
|
+
// when no baseline evidence is recorded yet the script logs a note event
|
|
314
|
+
// and parks with the exact parent protocol + resume path. Never fails the
|
|
315
|
+
// task over missing evidence: after two requests, baseline:none is
|
|
316
|
+
// recorded and the task continues without baseline comparison.
|
|
317
|
+
```
|
|
318
|
+
|
|
319
|
+
|
|
320
|
+
<a id="wrong-layer-guard"></a>
|
|
321
|
+
## Wrong-layer guard
|
|
322
|
+
|
|
323
|
+
Invariant: reproduce must run at the correct layer.
|
|
324
|
+
|
|
325
|
+
Applies to: bugfix.
|
|
326
|
+
|
|
327
|
+
```
|
|
328
|
+
// Reproduce wrong-layer guard baseline (2026-09-17): snapshot the repro
|
|
329
|
+
// evidence dir listing BEFORE the Reproduce work agent runs. The guard
|
|
330
|
+
// after the run diffs against this baseline — only frames created by the
|
|
331
|
+
// current attempt count as a wrong-layer violation; stale frames from
|
|
332
|
+
// pre-fix attempts appear in both snapshots and are excluded (set diff,
|
|
333
|
+
// no wall-clock, deterministic). One snapshot per phase attempt covers
|
|
334
|
+
// all transport retries inside the dispatch loop. null = baseline
|
|
335
|
+
// unavailable; the guard then fails closed for retry.
|
|
336
|
+
```
|
|
337
|
+
|
|
338
|
+
<a id="parent-reader-removed"></a>
|
|
339
|
+
## Parent reader removed
|
|
340
|
+
|
|
341
|
+
Invariant: the parent visual-verdict reader was removed; Hazel owns visual QA.
|
|
342
|
+
|
|
343
|
+
Applies to: bugfix.
|
|
344
|
+
|
|
345
|
+
```
|
|
346
|
+
// The parent visual-verdict reader was removed 2026-09-15: Hazel (the QA
|
|
347
|
+
// work agent) now owns the visual verdict experientially — she drives the
|
|
348
|
+
// now owns the visual verdict experientially — she drives the see-act loop
|
|
349
|
+
// herself and records verdict.json + the append-only verdicts.jsonl. The old
|
|
350
|
+
// see-act loop herself and records verdict.json + the append-only
|
|
351
|
+
// verdicts.jsonl. The old parent note gate is obsolete and has been
|
|
352
|
+
// deleted from the QA closeout below. (See docs/visual-verdict.md.)
|
|
353
|
+
// from the QA closeout below. (See docs/visual-verdict.md.)
|
|
354
|
+
```
|
|
355
|
+
|
|
356
|
+
<a id="surface-routing"></a>
|
|
357
|
+
## Surface routing
|
|
358
|
+
|
|
359
|
+
Invariant: the user-facing surface determines experiential QA routing.
|
|
360
|
+
|
|
361
|
+
Applies to: standard, bugfix.
|
|
362
|
+
|
|
363
|
+
```
|
|
364
|
+
// User-facing surface for experiential QA routing: 'artifact' (a rendered
|
|
365
|
+
// web UI Hazel drives with the see-act browser loop) | 'terminal' (a CLI
|
|
366
|
+
// Hazel drives herself, keeping attempt-scoped transcripts) | null
|
|
367
|
+
// (unclassified — no experiential QA). environment_type is the canonical
|
|
368
|
+
// UX-surface axis; deploy_type names the deployment target, but
|
|
369
|
+
// deploy_type === "artifact" remains a legacy artifact-surface signal so
|
|
370
|
+
// pre-field projects keep today's experiential QA (the migration does not
|
|
371
|
+
// backfill the column).
|
|
372
|
+
```
|
|
373
|
+
|
|
374
|
+
<a id="worktree-confinement"></a>
|
|
375
|
+
## Worktree confinement
|
|
376
|
+
|
|
377
|
+
Invariant: the Build agent must declare its worktree.
|
|
378
|
+
|
|
379
|
+
Applies to: standard, bugfix, chore.
|
|
380
|
+
|
|
381
|
+
```
|
|
382
|
+
// Worktree confinement: the Build agent must declare the exact worktree
|
|
383
|
+
// path it built in on a `worktree:` marker line. The workflow compares it
|
|
384
|
+
// against WORKTREE_HINT mechanically (exact string match) — never by
|
|
385
|
+
// reading agent prose. This closes the hole where a builder whose prepare
|
|
386
|
+
// failed freelanced into a different checkout (canary, 2026-09-11): the
|
|
387
|
+
// honest-but-confused case fails here, and a fabricated path is caught one
|
|
388
|
+
// phase later when Review's inspect finds no commits in the configured repo.
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
<a id="terminal-qa"></a>
|
|
392
|
+
## Terminal QA
|
|
393
|
+
|
|
394
|
+
Invariant: Hazel drives CLI transcripts for terminal surfaces.
|
|
395
|
+
|
|
396
|
+
Applies to: standard, bugfix.
|
|
397
|
+
|
|
398
|
+
```
|
|
399
|
+
// Terminal-surface experiential QA: Hazel drives the CLI herself —
|
|
400
|
+
// the terminal counterpart to the see-act loop above. Same OODA
|
|
401
|
+
// discipline (append-ooda-step with --action terminal and a transcript
|
|
402
|
+
// per step; verdict via write-ooda-verdict), judged against the
|
|
403
|
+
// shared bar resolved via UX_DOCTRINE_PATH (the terminal doctrine page here). Transcripts are delivered;
|
|
404
|
+
// screenshots are never invented for terminal work.
|
|
405
|
+
```
|
|
406
|
+
|
|
407
|
+
<a id="layer-resolution"></a>
|
|
408
|
+
## Layer resolution
|
|
409
|
+
|
|
410
|
+
Invariant: the task's layer determines the QA surface.
|
|
411
|
+
|
|
412
|
+
Applies to: bugfix.
|
|
413
|
+
|
|
414
|
+
```
|
|
415
|
+
// Layer resolution: the task's layer is "artifact", "engine", or "docs" —
|
|
416
|
+
// the machine-read "layer:" marker Sage's Triage report ends with. Unknown
|
|
417
|
+
// (missing/garbled line, failed lookup) degrades to "artifact": today's
|
|
418
|
+
// single-strategy behavior, never a park. Mirrors resolveExperiential()
|
|
419
|
+
// (same cache shape, same Triage-notes re-read); the layer flag is captured
|
|
420
|
+
// at Triage closeout (bugLayer) so the common path needs no extra agent call.
|
|
421
|
+
```
|
|
422
|
+
|
|
423
|
+
<a id="ux-doctrine-page"></a>
|
|
424
|
+
## UX doctrine page
|
|
425
|
+
|
|
426
|
+
Invariant: the shared UX bar for this run's surface.
|
|
427
|
+
|
|
428
|
+
Applies to: standard, bugfix, chore.
|
|
429
|
+
|
|
430
|
+
```
|
|
431
|
+
// UX doctrine page: the shared UX bar for this run's surface, resolved
|
|
432
|
+
// mechanically — every phase prompt reads UX_DOCTRINE_PATH, never a
|
|
433
|
+
// hardcoded filename. Canonical map: lib/ux-doctrine.js (mirrored here as a
|
|
434
|
+
// one-liner because the workflow runtime's relative-import support is
|
|
435
|
+
// unverified; tests pin the mirror). Null on unclassified surfaces: no
|
|
436
|
+
// shared page, and prompts say so instead of naming the wrong one.
|
|
437
|
+
```
|
|
438
|
+
|
|
439
|
+
<a id="visual-ownership"></a>
|
|
440
|
+
## Visual verdict ownership
|
|
441
|
+
|
|
442
|
+
Invariant: Hazel owns the visual verdict.
|
|
443
|
+
|
|
444
|
+
Applies to: bugfix.
|
|
445
|
+
|
|
446
|
+
```
|
|
447
|
+
// Visual verdict ownership (2026-09-15): Hazel owns the visual verdict
|
|
448
|
+
// experientially — the QA instructions above have her drive the see-act
|
|
449
|
+
// loop herself and record verdict.json + the append-only verdicts.jsonl.
|
|
450
|
+
// Her prose VERDICT: line drives `passed` via extractVerdict; a FAIL
|
|
451
|
+
// bounces to Build through the standard Review/QA rejection path below.
|
|
452
|
+
// The old parent-recorded note gate is deleted.
|
|
453
|
+
```
|
|
454
|
+
|
|
455
|
+
<a id="stale-decision-guard"></a>
|
|
456
|
+
## Stale-decision guard
|
|
457
|
+
|
|
458
|
+
Invariant: a stale baseline decision parks fail-closed.
|
|
459
|
+
|
|
460
|
+
Applies to: standard, bugfix, chore.
|
|
461
|
+
|
|
462
|
+
```
|
|
463
|
+
// Stale-decision guard: a "baseline: none (visual protocol unavailable)"
|
|
464
|
+
// note is only durable while the protocol is unavailable. When
|
|
465
|
+
// VISUAL_PROTOCOL_AVAILABLE is true, that old decision no longer
|
|
466
|
+
// stands — fall through to the request path for a fresh capture
|
|
467
|
+
// attempt. Exact-string trim comparison against the workflow's own
|
|
468
|
+
// written message (explicit state, never English matching).
|
|
469
|
+
```
|
|
470
|
+
|
|
471
|
+
<a id="pin-lifecycle"></a>
|
|
472
|
+
## Pin lifecycle
|
|
473
|
+
|
|
474
|
+
Invariant: snapshot the lifecycle scripts in the pin.
|
|
475
|
+
|
|
476
|
+
Applies to: standard, bugfix, chore.
|
|
477
|
+
|
|
478
|
+
```
|
|
479
|
+
// pinLifecycle(key) — snapshot the lifecycle scripts into RUN_LIB and return
|
|
480
|
+
// the verbatim `ls -1` listing so WORKFLOW CODE asserts the six pinned
|
|
481
|
+
// basenames; the agent cannot self-certify. (The pin step was the one place
|
|
482
|
+
// the workflows trusted agent prose: task 24be1cd6 walked to Publish on an
|
|
483
|
+
// empty pin dir.) Byte-identical across standard/bugfix/chore — pinned by
|
|
484
|
+
// tests/pin-location.test.js.
|
|
485
|
+
```
|
|
486
|
+
|