muse-crew 0.13.2 → 0.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/decisions/AGENTS.md +94 -0
- package/docs/decisions/publish-path.md +1280 -0
- package/docs/decisions/qa-reproduce.md +486 -0
- package/docs/decisions/workflow-core.md +234 -0
- package/docs/publish-unknown-recovery.md +113 -0
- package/docs/visual-verdict.md +6 -7
- package/lib/AGENTS.md +3 -1
- package/lib/classify-publish-absence.js +451 -0
- package/lib/compose-evidence-caption.js +36 -5
- package/lib/crew-api.js +584 -22
- package/lib/crew-release.sh +51 -9
- package/lib/publish-content.js +154 -0
- package/lib/retry-publish.js +370 -0
- package/lib/verify-publish.js +101 -84
- package/package.json +2 -2
- package/seed/cron-body-template.md +57 -2
- package/workflows/AGENTS.md +3 -1
- package/workflows/bugfix.js +301 -592
- package/workflows/chore.js +292 -526
- package/workflows/standard.js +294 -550
- package/workflows/upgrade.js +1 -1
|
@@ -0,0 +1,234 @@
|
|
|
1
|
+
# Decision history: workflow core
|
|
2
|
+
|
|
3
|
+
Relocated from workflow source comments during H5 (2026-09-18). The workflows keep only the relied-upon invariant inline; the full decision history lives here.
|
|
4
|
+
|
|
5
|
+
<a id="verdict-reask"></a>
|
|
6
|
+
## Verdict re-ask
|
|
7
|
+
|
|
8
|
+
Invariant: a verdict-step report that fails extractVerdict gets up to two bounded re-ask calls; the re-ask agent transcribes, never decides; exhaustion keeps fail-closed behavior.
|
|
9
|
+
|
|
10
|
+
Applies to: standard, bugfix, chore.
|
|
11
|
+
|
|
12
|
+
```
|
|
13
|
+
// Verdict re-ask (bug cd18ccc2): a verdict-step report that fails
|
|
14
|
+
// extractVerdict is not failed immediately. Stochastic verdict-line
|
|
15
|
+
// non-compliance (the agent did the work but omitted or garbled the VERDICT
|
|
16
|
+
// line) gets up to two bounded follow-up agent() calls whose only job is to
|
|
17
|
+
// read the preserved report and emit exactly one VERDICT line. The verdict
|
|
18
|
+
// is still extracted mechanically by extractVerdict — the re-ask agent
|
|
19
|
+
// transcribes, never decides the phase outcome. Each attempt uses a fresh
|
|
20
|
+
// stable-key suffix so a cached failure can never replay deterministically.
|
|
21
|
+
// Exhaustion keeps the existing fail-closed behavior. This is structure, not
|
|
22
|
+
// prompt hardening: no instruction text was stern-ified to get here.
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
<a id="park-contract"></a>
|
|
26
|
+
## Park contract + terminal cleanup
|
|
27
|
+
|
|
28
|
+
Invariant: parking is one atomic parktask action; a failed park reports failed (retryable); terminal cleanup releases the lock and reclaims only fully-merged work.
|
|
29
|
+
|
|
30
|
+
Applies to: standard, bugfix, chore.
|
|
31
|
+
|
|
32
|
+
```
|
|
33
|
+
// Park the task for human attention and end the run. "blocked" is never
|
|
34
|
+
// manually authored — the dashboard derives it mechanically from unmet
|
|
35
|
+
// dependencies — so a workflow outcome that needs a human parks the task
|
|
36
|
+
// instead. Parking is one atomic dashboard action (parktask): the parked
|
|
37
|
+
// state and the explanatory note land in one transaction, never half.
|
|
38
|
+
// The dispatcher skips parked tasks; a human moving parked→todo
|
|
39
|
+
// mechanically resets the retry counters. Returns the workflow result
|
|
40
|
+
// envelope the launcher sees. If the park call itself fails, the run
|
|
41
|
+
// reports "failed" (retryable) so the next tick re-attempts the park —
|
|
42
|
+
// a lost park is never reported as parked.
|
|
43
|
+
// Terminal cleanup: the run's last act at every park/fail boundary. A run
|
|
44
|
+
// that parks or fails must not leak its worktree, branch, or merge lock.
|
|
45
|
+
// The lifecycle's terminal-cleanup releases the lock unconditionally and
|
|
46
|
+
// reclaims the worktree+branch ONLY when the task branch is fully merged
|
|
47
|
+
// into main (then it is redundant); unmerged work is preserved for the
|
|
48
|
+
// human by design. Fire-and-forget with one bounded retry — the merge-lock
|
|
49
|
+
// lease expiry and the orphan sweep are the backstop for a dead transport.
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
<a id="closeout-envelope"></a>
|
|
53
|
+
## Closeout transport envelope
|
|
54
|
+
|
|
55
|
+
Invariant: the work agent returns the runtime's native transport envelope with no schema; the verdict is extracted mechanically by extractVerdict — never by an agent.
|
|
56
|
+
|
|
57
|
+
Applies to: standard, bugfix, chore.
|
|
58
|
+
|
|
59
|
+
```
|
|
60
|
+
// Work agent returns the runtime's native envelope {"status": "ok",
|
|
61
|
+
// "result": "<prose>"} with no schema. The runtime requires JSON output;
|
|
62
|
+
// the envelope is its own documented shape, so there is nothing for the
|
|
63
|
+
// agent to improvise. The workflow receives the prose report as a plain
|
|
64
|
+
// string. The verdict is still extracted deterministically from the report
|
|
65
|
+
// text by extractVerdict below — never by an agent.
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
<a id="extract-verdict-contract"></a>
|
|
69
|
+
## extractVerdict contract
|
|
70
|
+
|
|
71
|
+
Invariant: the verdict is the LAST VERDICT: PASS/FAIL line in the report; extraction is mechanical and deterministic.
|
|
72
|
+
|
|
73
|
+
Applies to: standard, bugfix, chore.
|
|
74
|
+
|
|
75
|
+
```
|
|
76
|
+
// The verdict is the LAST VERDICT: PASS/FAIL in the report (contract: end
|
|
77
|
+
// your report with the verdict). This ignores literal VERDICT strings echoed
|
|
78
|
+
// from the worker's instructions (which contain quoted examples). Fail
|
|
79
|
+
// closed if: no verdict found, the last verdict is not in the trailing 100
|
|
80
|
+
// chars (verdict must be at the end), or conflicting verdicts appear in the
|
|
81
|
+
// trailing 200 chars. Word boundary prevents "PASSING" matching as PASS.
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
<a id="rework-attempt-key"></a>
|
|
85
|
+
## Rework attempt key
|
|
86
|
+
|
|
87
|
+
Invariant: rework attempts are keyed deterministically; the count is a parameter because chore names it differently.
|
|
88
|
+
|
|
89
|
+
Applies to: standard, bugfix, chore.
|
|
90
|
+
|
|
91
|
+
```
|
|
92
|
+
// replay-key scoping for bug 1b8bb875 — runtime keys agent() calls by explicit key per workflow process; on in-process rework a re-executed phase must mint fresh keys; uses the same -r<N> suffix as the phase loop; pure function of inputs, no clock, no randomness (determinism contract). The count is a parameter because chore.js names the counter reworkCount while standard.js/bugfix.js name it totalReworkCount.
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
<a id="ux-doctrine-mirror"></a>
|
|
96
|
+
## UX doctrine mirror
|
|
97
|
+
|
|
98
|
+
Invariant: workflows mirror the UX doctrine map inline (relative-import support unverified); prompts read UX_DOCTRINE_PATH.
|
|
99
|
+
|
|
100
|
+
Applies to: standard, bugfix, chore.
|
|
101
|
+
|
|
102
|
+
```
|
|
103
|
+
// UX doctrine page: the shared UX bar for this run's surface, resolved
|
|
104
|
+
// mechanically — every phase prompt reads UX_DOCTRINE_PATH, never a
|
|
105
|
+
// hardcoded filename. Canonical map: lib/ux-doctrine.js (mirrored here as a
|
|
106
|
+
// one-liner because the workflow runtime's relative-import support is
|
|
107
|
+
// unverified; tests pin the mirror). Null on unclassified surfaces: no
|
|
108
|
+
// shared page, and prompts say so instead of naming the wrong one.
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
<a id="replay-key-scoping"></a>
|
|
112
|
+
## Replay-key scoping
|
|
113
|
+
|
|
114
|
+
Invariant: runtime keys agent() calls by explicit key per workflow process; on in-process rework a re-executed phase mints fresh keys with the -r<N> suffix.
|
|
115
|
+
|
|
116
|
+
Applies to: standard, bugfix, chore.
|
|
117
|
+
|
|
118
|
+
```
|
|
119
|
+
// replay-key scoping for bug 1b8bb875 — runtime keys agent() calls by explicit key per workflow process; on in-process rework a re-executed phase must mint fresh keys; uses the same -r<N> suffix as the phase loop; pure function of inputs, no clock, no randomness (determinism contract). The count is a parameter because chore.js names the counter reworkCount while standard.js/bugfix.js name it totalReworkCount.
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
<a id="transport-retry"></a>
|
|
123
|
+
## Transport retry
|
|
124
|
+
|
|
125
|
+
Invariant: the work-agent agent() call can throw even when the agent did the work; retry boundedly with fresh keys before failing closed; re-entry is safe.
|
|
126
|
+
|
|
127
|
+
Applies to: standard, bugfix, chore.
|
|
128
|
+
|
|
129
|
+
```
|
|
130
|
+
// Transport retry: the work-agent agent() call can throw even when the agent
|
|
131
|
+
// did the work. Stochastic envelope non-compliance (bare prose instead of
|
|
132
|
+
// the native {"status":"ok","result":"..."} envelope) trips the runtime's
|
|
133
|
+
// JSON-candidate heuristic when the prose contains a {...}-looking
|
|
134
|
+
// substring — canary 39457ee9's QA report quoted the change's own
|
|
135
|
+
// {/* ... */} JSX comment, the runtime tried to parse it as JSON, threw,
|
|
136
|
+
// and the workflow discarded a complete VERDICT: PASS report as "no output".
|
|
137
|
+
// The verdict re-ask covers an unreadable verdict inside a RECEIVED report;
|
|
138
|
+
// this covers the report never arriving. The assignment is retried boundedly
|
|
139
|
+
// with fresh keys (never a cached replay) before failing closed. Re-entry is
|
|
140
|
+
// safe: lifecycle scripts answer REUSED for existing worktrees/branches, the
|
|
141
|
+
// retry trailer tells the agent to check existing state first and report
|
|
142
|
+
// rather than duplicate completed side effects, and the rework path already
|
|
143
|
+
// re-runs Build after rejection — Build re-entry is an established pattern.
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
<a id="honest-classification"></a>
|
|
147
|
+
## Honest classification
|
|
148
|
+
|
|
149
|
+
Invariant: classify work-agent calls that yield no usable output honestly.
|
|
150
|
+
|
|
151
|
+
Applies to: standard, bugfix, chore.
|
|
152
|
+
|
|
153
|
+
```
|
|
154
|
+
// Honest classification of a work-agent call that yielded no usable
|
|
155
|
+
// report, with the per-attempt evidence preserved in the session notes.
|
|
156
|
+
// Two distinct cases:
|
|
157
|
+
// - agent() THREW: the runtime discarded output it could not machine-read
|
|
158
|
+
// (e.g. prose tripping the JSON-candidate heuristic). The raw output is
|
|
159
|
+
// gone — the workflow never received it — so the surviving error text is
|
|
160
|
+
// recorded here instead. (An earlier comment claimed the raw output was
|
|
161
|
+
// "preserved in the run record"; that was false — nothing preserved it —
|
|
162
|
+
// and the claim is removed.)
|
|
163
|
+
// - agent() RETURNED EMPTY without throwing: the child produced nothing
|
|
164
|
+
// usable. Each attempt's outcome is the evidence.
|
|
165
|
+
// attempts: [{threw, error, outcome}, ...], in order.
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
<a id="oneshot-recovery"></a>
|
|
169
|
+
## One-shot recovery
|
|
170
|
+
|
|
171
|
+
Invariant: the dispatcher sets input for one-shot recovery routing.
|
|
172
|
+
|
|
173
|
+
Applies to: standard, bugfix, chore.
|
|
174
|
+
|
|
175
|
+
```
|
|
176
|
+
// One-shot recovery routing: the dispatcher sets inputs.next_phase when it
|
|
177
|
+
// routes this run via an explicit recover-task redirect. The value is
|
|
178
|
+
// consumed (cleared) atomically by the successful self-claim below:
|
|
179
|
+
// claim-task takes expected_next_phase and clears the matching next_phase in
|
|
180
|
+
// the same transaction as the winning session insert, so no platform death
|
|
181
|
+
// can slip between claim and consumption and replay the routing. A stale or
|
|
182
|
+
// superseded routing survives — only an exact match clears.
|
|
183
|
+
// what the dispatcher routed on.
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
<a id="discarded-reason"></a>
|
|
187
|
+
## Discarded reason
|
|
188
|
+
|
|
189
|
+
Invariant: the runtime threw the output away; reason is discarded.
|
|
190
|
+
|
|
191
|
+
Applies to: standard, bugfix, chore.
|
|
192
|
+
|
|
193
|
+
```
|
|
194
|
+
// reason: "discarded" (the runtime threw the output away - it could not be
|
|
195
|
+
// machine-read), "empty" (agent() returned without throwing but produced
|
|
196
|
+
// nothing usable), "no-tools" (the worker's TOOL CHECK reported
|
|
197
|
+
// artifact_tools: missing), or "no-transport" (the worker's TOOL CHECK
|
|
198
|
+
// reported shell_transport: unavailable).
|
|
199
|
+
// The trailer tells the retry what to expect, not just to try again.
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
|
|
203
|
+
<a id="h5-probe-blocked"></a>
|
|
204
|
+
## H5 import probe: blocked, extraction refused (2026-09-18)
|
|
205
|
+
|
|
206
|
+
Invariant: workflow scripts carry no imports; shared logic is mirrored inline, never imported.
|
|
207
|
+
|
|
208
|
+
Applies to: standard, bugfix, chore.
|
|
209
|
+
|
|
210
|
+
```
|
|
211
|
+
# H5 (2026-09-18): the Phase 1 plan was probe-gated on the workflow runtime
|
|
212
|
+
# resolving relative imports of a sibling lib/ module. The probe could not
|
|
213
|
+
# be run: no platform launch capability exists in the build environment
|
|
214
|
+
# (workflow_launch is a platform capability, not a shell command; no local
|
|
215
|
+
# launcher exists), and a local node ESM check would prove nothing about the
|
|
216
|
+
# platform runtime — so it was refused, not faked. Outcome: PROBE-BLOCKED.
|
|
217
|
+
# Extraction without the probe would ship an unverified import into a
|
|
218
|
+
# launch-gated script — a worse defect than the size defect. The failover
|
|
219
|
+
# was measured in-file subtraction: delete dead buildVisualCapturePlan and
|
|
220
|
+
# relocate decision-history narration to docs/decisions/ (this file),
|
|
221
|
+
# keeping inline only the relied-upon invariant. The import question stays
|
|
222
|
+
# an OPEN PLATFORM FINDING; helper extraction (and the Phase 2 publish-block
|
|
223
|
+
# unification) is sequenced follow-up once the runtime's module support is
|
|
224
|
+
# proven by a real launch.
|
|
225
|
+
# Probe outcome, recorded 2026-09-18: the helper-extraction probe was
|
|
226
|
+
# refused as unverified — workflow scripts have no module system
|
|
227
|
+
# (typeof require === "undefined" on the platform runtime), so extraction
|
|
228
|
+
# would have shipped an unverified mechanism into a launch-gated script.
|
|
229
|
+
# The relocation-to-docs/decisions/ failover above was used instead.
|
|
230
|
+
# Sequencing: the next publish-path change must be preceded by the Phase 2
|
|
231
|
+
# publish-phase unification (or equivalent subtraction). bugfix.js sits at
|
|
232
|
+
# 244,349 bytes — 1,411 bytes of headroom under the 245,760-byte target —
|
|
233
|
+
# so an ad-hoc cut under gate pressure is not a plan.
|
|
234
|
+
```
|
|
@@ -0,0 +1,113 @@
|
|
|
1
|
+
# Publish-unknown recovery (blocker 15, 2026-09-18)
|
|
2
|
+
|
|
3
|
+
When a standard/bugfix Publish parks with "Publish outcome unknown", the
|
|
4
|
+
artifact-edit trigger went out fire-and-forget and no receipt came back —
|
|
5
|
+
async was planned for, receipt-less was not. The unknown-recovery loop
|
|
6
|
+
closes that gap without re-issuing blindly.
|
|
7
|
+
|
|
8
|
+
## The note is the state machine
|
|
9
|
+
|
|
10
|
+
Recovery state lives in the task's `note` events, keyed on machine-written
|
|
11
|
+
`publish: <transition>` markers. Deterministic code (`lib/crew-api.js`) owns
|
|
12
|
+
every transition; the cron tick (Step 4.4) is only the ferry between the
|
|
13
|
+
deterministic steps. The initial park note is the workflow's
|
|
14
|
+
"Publish outcome unknown …" note (no `publish:` marker — the scan matches it
|
|
15
|
+
explicitly); every transition after that is machine-written.
|
|
16
|
+
|
|
17
|
+
State diagram (latest `publish:` note wins; history is the retry budget):
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
"Publish outcome unknown" park
|
|
21
|
+
(< 30m) ──waiting──> scan: waiting (leave alone)
|
|
22
|
+
(>= 30m) ──scan──> publish: unknown-recovery-claimed <expiry> (1h lease)
|
|
23
|
+
│
|
|
24
|
+
├─ tick runs lib/classify-publish-absence.js, then
|
|
25
|
+
│ record-unknown-classification (CAS on the claim expiry):
|
|
26
|
+
│
|
|
27
|
+
├─ verified ──────────> publish: verification-requested <commit>
|
|
28
|
+
│ (content IS present; Step 4.5 verifies;
|
|
29
|
+
│ ledger: unknown-resolved)
|
|
30
|
+
├─ provably-dropped ──> publish: dropped <commit>
|
|
31
|
+
│ (queues the retry protocol, next tick)
|
|
32
|
+
├─ applied-not-built / ambiguous ──> publish: ambiguous <commit>
|
|
33
|
+
│ (terminal; ledger: unknown-classified)
|
|
34
|
+
├─ superseded ────────> publish: superseded <commit> (terminal)
|
|
35
|
+
└─ deferred ──────────> (no note; the claim expires; the next scan
|
|
36
|
+
re-claims and re-classifies)
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
Retry protocol (Step 4.4, for `publish: dropped` with a free budget):
|
|
40
|
+
|
|
41
|
+
```
|
|
42
|
+
publish: dropped
|
|
43
|
+
── tick: HEAD == commit? no ──> publish: retry-superseded (terminal)
|
|
44
|
+
── tick: acquire merge lock publish-retry:<task> (600s); held ──> stop, retry next tick
|
|
45
|
+
── tick: re-read content in the lock (classifier)
|
|
46
|
+
├─ not provably-dropped ──> record-retry-recheck routes it
|
|
47
|
+
│ (verified → verification-requested; ambiguous → terminal;
|
|
48
|
+
│ superseded → terminal; deferred → no-op)
|
|
49
|
+
└─ provably-dropped ──> publish: retry-intended <commit> <ts>
|
|
50
|
+
── trigger (same child shape as the first attempt)
|
|
51
|
+
├─ ARTIFACT_EDIT_REFUSED ──> publish: retry-refused (terminal)
|
|
52
|
+
└─ no refusal ──> publish: retry-issued <commit>
|
|
53
|
+
──> publish: verification-requested <commit> not-before=<ts+20m>
|
|
54
|
+
(Step 4.5 skips not-before entries until the window passes)
|
|
55
|
+
──> one -retry1 ledger entry (best-effort)
|
|
56
|
+
── release the merge lock (every path)
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Crash recovery (fail closed, never re-trigger blind):
|
|
60
|
+
|
|
61
|
+
- `publish: retry-intended` without `retry-issued` → the next scan mirrors
|
|
62
|
+
`verification-requested` with `not-before=intended+20m`. The trigger may
|
|
63
|
+
or may not have gone out; content verification is the arbiter.
|
|
64
|
+
- `publish: retry-issued` without a mirrored request → the next scan writes
|
|
65
|
+
the missing mirror.
|
|
66
|
+
- Exactly one retry per task, enforced from note history: the scan emits
|
|
67
|
+
`retry_due` only when no `publish: retry-issued` exists in the task's
|
|
68
|
+
history — including a retry for an earlier unknown attempt on a reworked
|
|
69
|
+
task. When the budget is spent, the drop is terminal: `publish: ambiguous`.
|
|
70
|
+
|
|
71
|
+
## Commands
|
|
72
|
+
|
|
73
|
+
- `scan-publish-unknown` — the cron scan (Step 4.5 of the old numbering).
|
|
74
|
+
Returns `{ waiting, due, retry_due, mirrored, skipped }`. `due` entries
|
|
75
|
+
ferry the classifier inputs: `commit`, `trigger_ts` (from the `submitted`
|
|
76
|
+
ledger entry — never a time window), `park_ts`, `slug`, `repo_path`,
|
|
77
|
+
`base` (provenance `source_commit`, else the empty tree), and
|
|
78
|
+
`claim_expiry` for the record CAS. `retry_due` entries ferry the retry
|
|
79
|
+
inputs (`commit`, `attempt`, `slug`, `repo_path`, `base`, `ledger_path`).
|
|
80
|
+
- `record-unknown-classification --json '{task_id, claim_expiry, decision}'`
|
|
81
|
+
— routes the classifier's decision; CAS on the claim expiry (a stale tick
|
|
82
|
+
records nothing).
|
|
83
|
+
- `record-retry-recheck --json '{task_id, decision}'` — routes the retry
|
|
84
|
+
protocol's in-lock content re-read; CAS on latest being `publish: dropped`.
|
|
85
|
+
- `resolve-publish-unknown` — the manual one-shot for a single parked task
|
|
86
|
+
(unchanged; the audit-window contract, not the classifier).
|
|
87
|
+
|
|
88
|
+
## Classifier (lib/classify-publish-absence.js)
|
|
89
|
+
|
|
90
|
+
Decides, from the task's repo and the platform's on-disk state, whether the
|
|
91
|
+
dropped edit is proven:
|
|
92
|
+
|
|
93
|
+
- `verified` — the content IS present (the platform applied it; the receipt
|
|
94
|
+
was the only thing lost). Never re-issue.
|
|
95
|
+
- `provably-dropped` — the old source is live AND the manifest shows no
|
|
96
|
+
build since the trigger. Only this decision may retry.
|
|
97
|
+
- `applied-not-built` — the new source is live but no build ran (the
|
|
98
|
+
trigger reached the platform but the build didn't). Ambiguous outcome,
|
|
99
|
+
terminal: a retry would double-apply.
|
|
100
|
+
- `ambiguous` — the content check is inconclusive. Never retry blind.
|
|
101
|
+
- `deferred` — not yet quiesced; re-check next tick.
|
|
102
|
+
- `superseded` — HEAD moved past the attempt's commit. Terminal.
|
|
103
|
+
|
|
104
|
+
The classifier's verified path requires the CURRENT manifest to be a new
|
|
105
|
+
build identity: `built_at` advanced past the trigger AND `content_sha256`
|
|
106
|
+
differs from the pre-trigger baseline the workflow snapshots into the
|
|
107
|
+
submitted ledger entry (design §1.9 — defeats a replayed manifest). When
|
|
108
|
+
the baseline is unavailable the classifier falls back to the time-based
|
|
109
|
+
advance check and notes it; the Step 4.5 verifier (`lib/verify-publish.js`)
|
|
110
|
+
is the strict gate and fails closed without a baseline before stamping.
|
|
111
|
+
Build-in-flight is checked twice bracketing the content read via manifest
|
|
112
|
+
state change (no mtime heuristics), and the suite carries add-only and
|
|
113
|
+
removal-only regression fixtures.
|
package/docs/visual-verdict.md
CHANGED
|
@@ -98,13 +98,12 @@ Triggered by the Capture phase when the task is experiential and no
|
|
|
98
98
|
baseline evidence is recorded yet. The workflow logs
|
|
99
99
|
`baseline: requested (attempt N)` and parks; the parent then:
|
|
100
100
|
|
|
101
|
-
1. Run the deterministic capture plan
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
`capture_targets:` line; if none, fall back to the task description.
|
|
101
|
+
1. Run the deterministic capture plan — desktop 1440x900 top, desktop
|
|
102
|
+
1440x900 target-centered, mobile 390x844 target-centered, hover,
|
|
103
|
+
keyboard focus, active/pressed where applicable, plus console error
|
|
104
|
+
count, the ARIA tree of the target region, and any horizontal overflow.
|
|
105
|
+
Capture targets come from the Map step's `capture_targets:` line; if
|
|
106
|
+
none, fall back to the task description.
|
|
108
107
|
2. Save the captures under `$CREW_HOME/task-evidence/<task-id>/baseline/`.
|
|
109
108
|
3. Log the task note event:
|
|
110
109
|
`baseline: captured <space-separated evidence refs>`
|
package/lib/AGENTS.md
CHANGED
|
@@ -24,8 +24,10 @@ Shell scripts for the crew's infrastructure. Called by workflow scripts, cron, a
|
|
|
24
24
|
- `ux-doctrine.js` — UX-surface doctrine page resolution (2026-09-17): the canonical map from `environment_type` to the crew's shared UX bar (`artifact` → `docs/artifact-ux.md`, `terminal` → `docs/terminal-ux.md`; null/unknown → no page). Pure and deterministic: `doctrinePage(env)`, `doctrinePath(crewHome, env)`, `doctrinePageExists(crewHome, env)`; CLI `--page <env>` / `--path <crewHome> <env>`. Workflows mirror the map inline (one line — the workflow runtime's relative-import support is unverified) and tests pin the mirror against this file.
|
|
25
25
|
- `serve-artifact.js` — local server for a built TS space for experiential QA (2026-09-14): serves `<space-dir>/client/dist` statically and dispatches POST `*/actions` to the compiled server actions with a locally-built Ctx. Prints `READY port=<n>` then serves until killed. Read-only w.r.t. the space directory. Fidelity: the served client and action handlers are the artifact's own built code; the Ctx is locally built (privileged handlers run from the space's own `server/dist/privileged.js` when present; blobs are stored in a per-run temp dir and served back at `/__blobs/<key>`); environment is inherited from the caller. It is not the hosted runtime — tasks that cannot be judged under it must report `NOT POSSIBLE: <reason>`.
|
|
26
26
|
- `readback-disk.js` — deterministic publish content sensor (2026-09-16): reads the on-disk tree the artifact is built/served from and emits the machine-readable findings block (`FILE:`/`ADDED:`/`REMOVED:`/`END_FILE`) that `verify-publish.js` judges. The primary sensor — the LLM-inspector path (`build-readback-request.js`) is manual-fallback only since `artifact_inspect` was removed by the platform 2026-09-14.
|
|
27
|
+
- `publish-content.js` — shared ESM content-primitives for publish verification (2026-09-18, blocker 15): diff parsing (`parseDiff`), findings parsing (`parseFindings`), old-tree occurrence counting (`makeOldCounter`), and the discriminating-line / collision-exemption logic (`discriminatingLines`). Unifies `verify-publish.js` and the unknown-recovery classifier on one judgment so the two paths can never disagree about what a diff proves.
|
|
28
|
+
- `classify-publish-absence.js` — deterministic six-way classifier for publish-parked UNKNOWN outcomes (2026-09-18, blocker 15): decides from durable signals only — the pre-trigger manifest baseline (captured by the workflow into the submitted ledger entry; design §1.9) vs the current manifest's `built_at`/`content_sha256`, `git diff <base> <commit>` discriminating lines against the on-disk source tree, and HEAD vs the publish commit. Outcomes: `provably-dropped` (source shows pre-edit state, manifest NOT advanced past the trigger, HEAD == commit, aged past quiesce — retry once), `verified` (manifest advanced past the trigger AND content_sha256 differs from the pre-trigger baseline — a new build identity, not a replayed manifest; falls back to the time-based advance check with a note when the baseline is absent, and the Step 4.5 verifier fails closed without a baseline), `applied-not-built` (platform build-emission failure; no retry — the 2026-09-12 re-trigger hazard), `ambiguous` (any inconclusive shape — no retry by design), `deferred` (build in flight — manifest changed during the content read, or built within the settle window — or park below quiesce; not a verdict, retry later), `superseded` (HEAD != commit — never retry the old commit). Retry budget is consumed by the classification itself, never by the edit attempt. Never reads the wall clock except for recovery timing; addition-only and removal-only diffs are vacuously satisfied on their empty side.
|
|
27
29
|
- `build-readback-request.js` — builds the LLM-inspector read-back `verbatim_request` from the merge commit's diff (2026-09-14): carries the merged diff as the expected change and asks for an independent read of the artifact's actual source. Retained as the manual fallback; the deterministic `readback-disk.js` is the primary sensor.
|
|
28
|
-
- `verify-publish.js` — mechanical publish verification judge (2026-09-14/16): certifies the read-back findings block against `git diff` (strict `FILE:`/`ADDED:`/`REMOVED:`/`END_FILE` parsing, every added line PRESENT / every removed line ABSENT, HEAD==commit supersession check) and only then stamps provenance. Binary files, mode-only changes, and fully-colliding added hunks fail closed as `unverifiable-content` (2026-09-16, critic findings 1/5) — they can never vacuously stamp. Content-mismatch, unreadable-result, superseded, and stamp failures exit 1 with `publish: verification-failed` and no stamp.
|
|
30
|
+
- `verify-publish.js` — mechanical publish verification judge (2026-09-14/16; shared primitives 2026-09-18): certifies the read-back findings block against `git diff` (strict `FILE:`/`ADDED:`/`REMOVED:`/`END_FILE` parsing, every discriminating added line PRESENT / every discriminating removed line ABSENT, HEAD==commit supersession check), then the design §1.9 manifest-freshness gate (current manifest `built_at` advanced past the trigger AND `content_sha256` differs from the workflow's pre-trigger baseline in the submitted ledger entry — a new build identity, not a replayed manifest; missing baseline fails closed), and only then stamps provenance. Diff parsing, findings parsing, and the collision-exemption rules come from the shared `lib/publish-content.js` (the unknown-recovery classifier's own judgment — one definition, never two). Binary files, mode-only changes, and fully-colliding added hunks fail closed as `unverifiable-content` (2026-09-16, critic findings 1/5) — they can never vacuously stamp. Content-mismatch, unreadable-result, superseded, and stamp failures exit 1 with `publish: verification-failed` and no stamp.
|
|
29
31
|
- `update-watch.js` — deterministic automatic update watcher (2026-09-16, zero deps): `node update-watch.js --crew-home <path>` (missing arg → usage, exit 2; every other path exits 0). Watches the public npm registry (`npm view muse-crew version` pinned to `https://registry.npmjs.org/`) vs `crew-release.sh current` and files a `workflow: "upgrade"` task with `source: npm@<version>` when policy (`auto_update_crew`, `update_channel`) and channel gating allow; watches `git ls-remote origin HEAD` on the first `deploy_type=artifact` project vs `$CREW_HOME/.update-watch.json` and files a `workflow: "chore"` task carrying the mechanical dashboard-upgrade journey. Reads the `.crew-version` compatibility anchor at the new ref via `git fetch` + `git show <sha>:.crew-version` (never the working tree) and orders dashboard-led: a declared newer crew files the crew upgrade task FIRST and the dashboard task notes it follows the crew upgrade (declaration bypasses `update_channel`, not the `auto_update_crew=false` opt-out); a declared older crew skips the dashboard leg entirely as a human decision; a missing/invalid/unfetchable anchor fails open to the dashboard leg as today. Idempotency via the same state file (records at file time); check failures log to `$CREW_HOME/update-watch.log` and are never thrown. Safety: only files tasks — never deploys, never touches the artifact/config/scheduler. Run by the daily `crew-update-watch` cron through the `current` symlink (latest release); deliberately NOT in the lib-pinning `PIN_BASENAMES`.
|
|
30
32
|
- `gitignore.js` — deterministic .gitignore management for crew-owned paths (2026-09-17): the crew touches exactly one user-owned file outside `.orchestration/` — the repo's `.gitignore`. `ensureGitignoreEntries(repoPath, entries)` creates the file when missing, appends missing entries (exact line match, no duplicates), preserves existing content byte-for-byte, and is idempotent. `describeGitignoreChange(repoPath, entries)` renders the exact diff for the setup consent conversation. Crew-owned entries: `.worktrees/`, `.orchestration/user/`. CLI: `--repo <path> [--dry-run]`.
|
|
31
33
|
- `repo-orchestration.js` — repository-local `.orchestration/` scaffold (2026-09-17): `scaffoldRepoOrchestration(repoPath, crewRepoPath)` creates `$REPO/.orchestration/{workflows,identities,phases,user}/`, seeds workflows/identities/phases from the crew repo's platform defaults with no-clobber semantics (existing project customizations never overwritten), and writes a README in `user/` explaining it's for local config. Idempotent. CLI: `--repo <path> --crew-repo <path>`.
|