muse-crew 0.13.1 → 0.13.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,486 @@
1
+ # Decision history: qa reproduce
2
+
3
+ Relocated from workflow source comments during H5 (2026-09-18). The workflows keep only the relied-upon invariant inline; the full decision history lives here.
4
+
5
+ <a id="reproduce-wrong-layer-guard"></a>
6
+ ## Reproduce wrong-layer guard
7
+
8
+ Invariant: for engine/docs tasks the experiential see-act loop is forbidden; the workflow checks the repro dir mechanically for frame artifacts, attempt-scoped; no baseline is inconclusive fail-closed.
9
+
10
+ Applies to: standard, bugfix, chore.
11
+
12
+ ```
13
+ // Reproduce wrong-layer guard (2026-09-16): for engine/docs tasks the
14
+ // experiential see-act loop is forbidden — a reproducer that drives the
15
+ // browser burns the run and emits a wrong-layer verdict (three tasks
16
+ // parked on exactly this in the 2026-09-17 UTC shift). The workflow
17
+ // checks the repro evidence dir mechanically for see-act frame artifacts
18
+ // (the driver's archive naming: <n>-<action>-<desktop|mobile>.png).
19
+ // Attempt-scoped (2026-09-17): the after-listing is diffed against the
20
+ // baseline taken before this phase attempt's dispatch loop — only frames
21
+ // CREATED BY the current attempt count as a violation; stale frames from
22
+ // earlier attempts appear in both snapshots and are excluded. Without a
23
+ // baseline the guard is inconclusive and fails closed for retry (passing
24
+ // could route a genuinely wrong-layer verdict to Map).
25
+ ```
26
+
27
+ <a id="baseline-evidence-capture"></a>
28
+ ## Baseline evidence capture
29
+
30
+ Invariant: capture baseline evidence for experiential tasks before the QA agent runs; the baseline is the control.
31
+
32
+ Applies to: standard.
33
+
34
+ ```
35
+ // ── Capture: baseline evidence for experiential tasks ─────────────
36
+ // Hazel's QA capture pass runs right after Triage, before Map, for tasks
37
+ // Sage flagged experiential. The capture itself is parent-driven (the
38
+ // inspection handoff arrives at the root agent, outside this script), so
39
+ // when no baseline evidence is recorded yet the script logs a note event
40
+ // and parks with the exact parent protocol + resume path. Never fails the
41
+ // task over missing evidence: after two requests, baseline:none is
42
+ // recorded and final QA judges on the rubric alone.
43
+ ```
44
+
45
+ Applies to: bugfix.
46
+
47
+ ```
48
+ // ── Capture: baseline evidence for experiential tasks ─────────────
49
+ // Hazel's QA capture pass runs right after Triage, before Map, for tasks
50
+ // Sage flagged experiential. The capture itself is parent-driven (the
51
+ // inspection handoff arrives at the root agent, outside this script), so
52
+ // when no baseline evidence is recorded yet the script logs a note event
53
+ // and parks with the exact parent protocol + resume path. Never fails the
54
+ // task over missing evidence: after two requests, baseline:none is
55
+ // recorded and final QA judges on the rubric alone.
56
+ ```
57
+
58
+ Applies to: chore.
59
+
60
+ ```
61
+ // ── Capture: baseline evidence for experiential tasks ─────────────
62
+ // Hazel's QA capture pass runs right after Triage, before Map, for tasks
63
+ // Sage flagged experiential. The capture itself is parent-driven (the
64
+ // inspection handoff arrives at the root agent, outside this script), so
65
+ // when no baseline evidence is recorded yet the script logs a note event
66
+ // and parks with the exact parent protocol + resume path. Never fails the
67
+ // task over missing evidence: after two requests, baseline:none is
68
+ // recorded and the task continues without baseline comparison.
69
+ ```
70
+
71
+
72
+ <a id="pre-publish-pin-guard"></a>
73
+ ## Pre-Publish pin guard
74
+
75
+ Invariant: the pin step materializes the required scripts before Publish; the agent cannot self-certify.
76
+
77
+ Applies to: standard, bugfix, chore.
78
+
79
+ ```
80
+ // The seven basenames the pin step must materialize — asserted mechanically
81
+ // by workflow code from the verbatim listing, never from agent prose.
82
+ // COMPUTE_DIFF is the deterministic publish-diff computer (room #14,
83
+ // 2026-09-17): the diff is computed by this script, never ferried as an
84
+ // agent JSON string. Pinned like the other publish-critical modules so a
85
+ // mid-run release swap cannot change it under the workflow.
86
+ // CLASSIFY_SURFACE is the surface classifier (room #15, 2026-09-18):
87
+ // crew-api.js statically imports it, so the pin must carry it — a pin
88
+ // without it kills every claim with ERR_MODULE_NOT_FOUND.
89
+ ```
90
+
91
+ <a id="qa-verdict-gate"></a>
92
+ ## QA verdict.json closeout gate
93
+
94
+ Invariant: the prose VERDICT: line drives the closeout gate via extractVerdict; a bare FAIL with an all-positive report is a defect.
95
+
96
+ Applies to: standard, bugfix, chore.
97
+
98
+ ```
99
+ // QA verdict.json closeout gate (task 30dceb78): the prose VERDICT: line is
100
+ // the routing signal, but verdict.json is the reason-carrying record the
101
+ // workflow actually reads. A FAIL verdict must carry a machine-readable
102
+ // reason; the workflow refuses to route to rework on an unreasoned or
103
+ // contradictory verdict. The deterministic cross-checker
104
+ // (lib/read-ooda-verdict.js) runs against the prose verdict: a missing,
105
+ // corrupt, contradictory, or reason-less record fails the phase for retry
106
+ // — the dispatcher re-runs QA at the same step under its
107
+ // consecutive-failure cap — instead of routing to rework. Without this
108
+ // gate, a bare VERDICT: FAIL with an all-positive report (canary
109
+ // 2026-09-15, task 1d692d91) rebuilt nothing and parked at Publish on an
110
+ // unobserved artifact build. The gate applies only to the experiential QA
111
+ // path (qaArtifact / qaTerminal), the paths that write verdict.json.
112
+ ```
113
+
114
+ <a id="experiential-loop-guard"></a>
115
+ ## QA experiential-loop guard
116
+
117
+ Invariant: a PASS verdict with missing experiential evidence never completes — parks fail-closed with unattributable_reason.
118
+
119
+ Applies to: standard.
120
+
121
+ ```
122
+ // QA experiential-loop guard (clean-room defect 2026-09-16): Hazel's verdict.json
123
+ // is honest about missing experiential evidence, but the closeout treated a PASS
124
+ // as terminal done even when the experiential loop never ran (playwright-core
125
+ // was unresolvable from the release layout — the dependency lived in the
126
+ // npm install dir, severed from the crew home). A PASS verdict with missing
127
+ // experiential evidence must never be terminal: the task parks fail-closed
128
+ // with unattributable_reason=qa-visual-loop-unavailable (artifact surface)
129
+ // or qa-terminal-loop-unavailable (terminal surface) instead of
130
+ // transitioning to done. Code, not prompt text: the check reads the
131
+ // machine-readable verdict via lib/read-ooda-verdict.js, which reports
132
+ // visual_loop_unavailable from the OODA log's NOT POSSIBLE browser steps
133
+ // and terminal_loop_unavailable from NOT POSSIBLE terminal steps, plus the
134
+ // verdict's missing_evidence tool-unavailability notes.
135
+ ```
136
+
137
+ Applies to: bugfix.
138
+
139
+ ```
140
+ // QA experiential-loop guard (clean-room defect 2026-09-16): Hazel's
141
+ // verdict.json is honest about missing experiential evidence, but the
142
+ // closeout treated a PASS as terminal done even when the experiential
143
+ // loop never ran. A PASS verdict with missing experiential evidence must
144
+ // never be terminal: the task parks fail-closed with
145
+ // unattributable_reason=qa-visual-loop-unavailable (artifact surface)
146
+ // or qa-terminal-loop-unavailable (terminal surface) instead of
147
+ // transitioning to done. Code, not prompt text: the loop-availability
148
+ // flags come from lib/read-ooda-verdict.js (the OODA log's NOT POSSIBLE
149
+ // steps and the verdict's missing_evidence tool-unavailability notes),
150
+ // already parsed into qaVerdictGate above.
151
+ ```
152
+
153
+
154
+ <a id="reproduce-baseline"></a>
155
+ ## Reproduce wrong-layer guard baseline
156
+
157
+ Invariant: snapshot the repro evidence dir before the Reproduce agent runs; only frames created by the current attempt count; null baseline is inconclusive fail-closed.
158
+
159
+ Applies to: standard, bugfix, chore.
160
+
161
+ ```
162
+ // Reproduce wrong-layer guard baseline (2026-09-17): snapshot the repro
163
+ // evidence dir listing BEFORE the Reproduce work agent runs. The guard
164
+ // after the run diffs against this baseline — only frames created by the
165
+ // current attempt count as a wrong-layer violation; stale frames from
166
+ // pre-fix attempts appear in both snapshots and are excluded (set diff,
167
+ // no wall-clock, deterministic). One snapshot per phase attempt covers
168
+ // all transport retries inside the dispatch loop. null = baseline
169
+ // unavailable; the guard then fails closed for retry.
170
+ ```
171
+
172
+ <a id="visual-verdict-ownership"></a>
173
+ ## Visual verdict ownership
174
+
175
+ Invariant: Hazel owns the visual verdict experientially (see-act loop + verdict.json); her prose VERDICT drives passed; the old parent note gate is deleted.
176
+
177
+ Applies to: standard, bugfix, chore.
178
+
179
+ ```
180
+ // Visual verdict ownership (2026-09-15): Hazel owns the visual verdict
181
+ // experientially — the QA instructions above have her drive the see-act
182
+ // loop herself and record verdict.json + the append-only verdicts.jsonl.
183
+ // Her prose VERDICT: line drives `passed` via extractVerdict; a FAIL
184
+ // bounces to Build through the standard Review/QA rejection path below.
185
+ // The old parent-recorded note gate is deleted.
186
+ ```
187
+
188
+ <a id="baseline-evidence-status"></a>
189
+ ## Baseline evidence status
190
+
191
+ Invariant: reads the task's note events for baseline status.
192
+
193
+ Applies to: standard, bugfix, chore.
194
+
195
+ ```
196
+ // Baseline evidence status: reads the task's note events for the exact
197
+ // protocol prefixes (explicit state, never English matching). Returns
198
+ // { baseline_found, baseline_kind, baseline_refs, requested_count, evidence_count }.
199
+ // found = any note starting exactly "baseline: captured" or "baseline:
200
+ // none" (kind/refs come from the LATEST such message); requested_count =
201
+ // the number of notes starting exactly "baseline: requested";
202
+ // evidence_count = the number of notes starting exactly "baseline: captured" or
203
+ // "baseline: none" (the per-attempt evidence chain; state-derived, no wall clock). Each call
204
+ // uses a fresh key: the evidence changes between calls (the parent logs
205
+ // the capture while this run is parked), so a cached replay would lie.
206
+ ```
207
+
208
+ <a id="pin-basenames"></a>
209
+ ## Pin basenames
210
+
211
+ Invariant: the pin step materializes the required scripts.
212
+
213
+ Applies to: standard, bugfix, chore.
214
+
215
+ ```
216
+ // The seven basenames the pin step must materialize — asserted mechanically
217
+ // by workflow code from the verbatim listing, never from agent prose.
218
+ // COMPUTE_DIFF is the deterministic publish-diff computer (room #14,
219
+ // 2026-09-17): the diff is computed by this script, never ferried as an
220
+ // agent JSON string. Pinned like the other publish-critical modules so a
221
+ // mid-run release swap cannot change it under the workflow.
222
+ // CLASSIFY_SURFACE is the surface classifier (room #15, 2026-09-18):
223
+ // crew-api.js statically imports it, so the pin must carry it — a pin
224
+ // without it kills every claim with ERR_MODULE_NOT_FOUND.
225
+ ```
226
+
227
+ <a id="visual-protocol-avail"></a>
228
+ ## Visual protocol availability
229
+
230
+ Invariant: the workflow parks if the visual protocol is unavailable.
231
+
232
+ Applies to: standard.
233
+
234
+ ```
235
+ // Visual verdict protocol availability — the workflow parks for parent-run
236
+ // baseline capture and visual verdict ONLY when the protocol is fully
237
+ // shipped. The protocol requires docs/visual-verdict.md in the release AND
238
+ // the parent-side capture tooling (task b309a97d, "QA owns the visual
239
+ // verdict"). Until both exist, the parks would deadlock waiting for a
240
+ // parent who cannot fulfill them.
241
+ // Effective value for this run, resolved by the dispatcher from the
242
+ // project's visual_protocol setting (null=inherits crew default=off).
243
+ // Manual launches without the arg default to off (previous behavior).
244
+ ```
245
+
246
+ Applies to: bugfix.
247
+
248
+ ```
249
+ // Visual verdict protocol availability — the workflow parks for parent-run
250
+ // baseline capture and visual verdict ONLY when the protocol is fully
251
+ // shipped. The protocol requires docs/visual-verdict.md in the release AND
252
+ // the parent-side capture tooling (task b309a97d, "QA owns the visual
253
+ // verdict"). Until both exist, the parks would deadlock waiting for a
254
+ // parent who cannot fulfill them.
255
+ // Effective value for this run, resolved by the dispatcher from the
256
+ // project's visual_protocol setting (null=inherits crew default=off).
257
+ // Manual launches without the arg default to off (previous behavior).
258
+ ```
259
+
260
+ Applies to: chore.
261
+
262
+ ```
263
+ // Visual verdict protocol availability — the workflow parks for parent-run
264
+ // baseline capture ONLY when the protocol is fully shipped. The protocol
265
+ // requires docs/visual-verdict.md in the release AND the parent-side
266
+ // capture tooling (task b309a97d, "QA owns the visual verdict"). Until both
267
+ // exist, the park would deadlock waiting for a parent who cannot fulfill
268
+ // it.
269
+ // Effective value for this run, resolved by the dispatcher from the
270
+ // project's visual_protocol setting (null=inherits crew default=off).
271
+ // Manual launches without the arg default to off (previous behavior).
272
+ ```
273
+
274
+
275
+ <a id="capture-baseline"></a>
276
+ ## Capture baseline
277
+
278
+ Invariant: baseline evidence for experiential tasks is captured before work begins.
279
+
280
+ Applies to: standard.
281
+
282
+ ```
283
+ // ── Capture: baseline evidence for experiential tasks ─────────────
284
+ // Hazel's QA capture pass runs right after Triage, before Map, for tasks
285
+ // Sage flagged experiential. The capture itself is parent-driven (the
286
+ // inspection handoff arrives at the root agent, outside this script), so
287
+ // when no baseline evidence is recorded yet the script logs a note event
288
+ // and parks with the exact parent protocol + resume path. Never fails the
289
+ // task over missing evidence: after two requests, baseline:none is
290
+ // recorded and final QA judges on the rubric alone.
291
+ ```
292
+
293
+ Applies to: bugfix.
294
+
295
+ ```
296
+ // ── Capture: baseline evidence for experiential tasks ─────────────
297
+ // Hazel's QA capture pass runs right after Triage, before Map, for tasks
298
+ // Sage flagged experiential. The capture itself is parent-driven (the
299
+ // inspection handoff arrives at the root agent, outside this script), so
300
+ // when no baseline evidence is recorded yet the script logs a note event
301
+ // and parks with the exact parent protocol + resume path. Never fails the
302
+ // task over missing evidence: after two requests, baseline:none is
303
+ // recorded and final QA judges on the rubric alone.
304
+ ```
305
+
306
+ Applies to: chore.
307
+
308
+ ```
309
+ // ── Capture: baseline evidence for experiential tasks ─────────────
310
+ // Hazel's QA capture pass runs right after Triage, before Map, for tasks
311
+ // Sage flagged experiential. The capture itself is parent-driven (the
312
+ // inspection handoff arrives at the root agent, outside this script), so
313
+ // when no baseline evidence is recorded yet the script logs a note event
314
+ // and parks with the exact parent protocol + resume path. Never fails the
315
+ // task over missing evidence: after two requests, baseline:none is
316
+ // recorded and the task continues without baseline comparison.
317
+ ```
318
+
319
+
320
+ <a id="wrong-layer-guard"></a>
321
+ ## Wrong-layer guard
322
+
323
+ Invariant: reproduce must run at the correct layer.
324
+
325
+ Applies to: bugfix.
326
+
327
+ ```
328
+ // Reproduce wrong-layer guard baseline (2026-09-17): snapshot the repro
329
+ // evidence dir listing BEFORE the Reproduce work agent runs. The guard
330
+ // after the run diffs against this baseline — only frames created by the
331
+ // current attempt count as a wrong-layer violation; stale frames from
332
+ // pre-fix attempts appear in both snapshots and are excluded (set diff,
333
+ // no wall-clock, deterministic). One snapshot per phase attempt covers
334
+ // all transport retries inside the dispatch loop. null = baseline
335
+ // unavailable; the guard then fails closed for retry.
336
+ ```
337
+
338
+ <a id="parent-reader-removed"></a>
339
+ ## Parent reader removed
340
+
341
+ Invariant: the parent visual-verdict reader was removed; Hazel owns visual QA.
342
+
343
+ Applies to: bugfix.
344
+
345
+ ```
346
+ // The parent visual-verdict reader was removed 2026-09-15: Hazel (the QA
347
+ // work agent) now owns the visual verdict experientially — she drives the
348
+ // now owns the visual verdict experientially — she drives the see-act loop
349
+ // herself and records verdict.json + the append-only verdicts.jsonl. The old
350
+ // see-act loop herself and records verdict.json + the append-only
351
+ // verdicts.jsonl. The old parent note gate is obsolete and has been
352
+ // deleted from the QA closeout below. (See docs/visual-verdict.md.)
353
+ // from the QA closeout below. (See docs/visual-verdict.md.)
354
+ ```
355
+
356
+ <a id="surface-routing"></a>
357
+ ## Surface routing
358
+
359
+ Invariant: the user-facing surface determines experiential QA routing.
360
+
361
+ Applies to: standard, bugfix.
362
+
363
+ ```
364
+ // User-facing surface for experiential QA routing: 'artifact' (a rendered
365
+ // web UI Hazel drives with the see-act browser loop) | 'terminal' (a CLI
366
+ // Hazel drives herself, keeping attempt-scoped transcripts) | null
367
+ // (unclassified — no experiential QA). environment_type is the canonical
368
+ // UX-surface axis; deploy_type names the deployment target, but
369
+ // deploy_type === "artifact" remains a legacy artifact-surface signal so
370
+ // pre-field projects keep today's experiential QA (the migration does not
371
+ // backfill the column).
372
+ ```
373
+
374
+ <a id="worktree-confinement"></a>
375
+ ## Worktree confinement
376
+
377
+ Invariant: the Build agent must declare its worktree.
378
+
379
+ Applies to: standard, bugfix, chore.
380
+
381
+ ```
382
+ // Worktree confinement: the Build agent must declare the exact worktree
383
+ // path it built in on a `worktree:` marker line. The workflow compares it
384
+ // against WORKTREE_HINT mechanically (exact string match) — never by
385
+ // reading agent prose. This closes the hole where a builder whose prepare
386
+ // failed freelanced into a different checkout (canary, 2026-09-11): the
387
+ // honest-but-confused case fails here, and a fabricated path is caught one
388
+ // phase later when Review's inspect finds no commits in the configured repo.
389
+ ```
390
+
391
+ <a id="terminal-qa"></a>
392
+ ## Terminal QA
393
+
394
+ Invariant: Hazel drives CLI transcripts for terminal surfaces.
395
+
396
+ Applies to: standard, bugfix.
397
+
398
+ ```
399
+ // Terminal-surface experiential QA: Hazel drives the CLI herself —
400
+ // the terminal counterpart to the see-act loop above. Same OODA
401
+ // discipline (append-ooda-step with --action terminal and a transcript
402
+ // per step; verdict via write-ooda-verdict), judged against the
403
+ // shared bar resolved via UX_DOCTRINE_PATH (the terminal doctrine page here). Transcripts are delivered;
404
+ // screenshots are never invented for terminal work.
405
+ ```
406
+
407
+ <a id="layer-resolution"></a>
408
+ ## Layer resolution
409
+
410
+ Invariant: the task's layer determines the QA surface.
411
+
412
+ Applies to: bugfix.
413
+
414
+ ```
415
+ // Layer resolution: the task's layer is "artifact", "engine", or "docs" —
416
+ // the machine-read "layer:" marker Sage's Triage report ends with. Unknown
417
+ // (missing/garbled line, failed lookup) degrades to "artifact": today's
418
+ // single-strategy behavior, never a park. Mirrors resolveExperiential()
419
+ // (same cache shape, same Triage-notes re-read); the layer flag is captured
420
+ // at Triage closeout (bugLayer) so the common path needs no extra agent call.
421
+ ```
422
+
423
+ <a id="ux-doctrine-page"></a>
424
+ ## UX doctrine page
425
+
426
+ Invariant: the shared UX bar for this run's surface.
427
+
428
+ Applies to: standard, bugfix, chore.
429
+
430
+ ```
431
+ // UX doctrine page: the shared UX bar for this run's surface, resolved
432
+ // mechanically — every phase prompt reads UX_DOCTRINE_PATH, never a
433
+ // hardcoded filename. Canonical map: lib/ux-doctrine.js (mirrored here as a
434
+ // one-liner because the workflow runtime's relative-import support is
435
+ // unverified; tests pin the mirror). Null on unclassified surfaces: no
436
+ // shared page, and prompts say so instead of naming the wrong one.
437
+ ```
438
+
439
+ <a id="visual-ownership"></a>
440
+ ## Visual verdict ownership
441
+
442
+ Invariant: Hazel owns the visual verdict.
443
+
444
+ Applies to: bugfix.
445
+
446
+ ```
447
+ // Visual verdict ownership (2026-09-15): Hazel owns the visual verdict
448
+ // experientially — the QA instructions above have her drive the see-act
449
+ // loop herself and record verdict.json + the append-only verdicts.jsonl.
450
+ // Her prose VERDICT: line drives `passed` via extractVerdict; a FAIL
451
+ // bounces to Build through the standard Review/QA rejection path below.
452
+ // The old parent-recorded note gate is deleted.
453
+ ```
454
+
455
+ <a id="stale-decision-guard"></a>
456
+ ## Stale-decision guard
457
+
458
+ Invariant: a stale baseline decision parks fail-closed.
459
+
460
+ Applies to: standard, bugfix, chore.
461
+
462
+ ```
463
+ // Stale-decision guard: a "baseline: none (visual protocol unavailable)"
464
+ // note is only durable while the protocol is unavailable. When
465
+ // VISUAL_PROTOCOL_AVAILABLE is true, that old decision no longer
466
+ // stands — fall through to the request path for a fresh capture
467
+ // attempt. Exact-string trim comparison against the workflow's own
468
+ // written message (explicit state, never English matching).
469
+ ```
470
+
471
+ <a id="pin-lifecycle"></a>
472
+ ## Pin lifecycle
473
+
474
+ Invariant: snapshot the lifecycle scripts in the pin.
475
+
476
+ Applies to: standard, bugfix, chore.
477
+
478
+ ```
479
+ // pinLifecycle(key) — snapshot the lifecycle scripts into RUN_LIB and return
480
+ // the verbatim `ls -1` listing so WORKFLOW CODE asserts the six pinned
481
+ // basenames; the agent cannot self-certify. (The pin step was the one place
482
+ // the workflows trusted agent prose: task 24be1cd6 walked to Publish on an
483
+ // empty pin dir.) Byte-identical across standard/bugfix/chore — pinned by
484
+ // tests/pin-location.test.js.
485
+ ```
486
+