muse-crew 0.13.1 → 0.13.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/decisions/AGENTS.md +94 -0
- package/docs/decisions/publish-path.md +1280 -0
- package/docs/decisions/qa-reproduce.md +486 -0
- package/docs/decisions/workflow-core.md +234 -0
- package/docs/guide.md +3 -3
- package/docs/visual-verdict.md +6 -7
- package/lib/compose-evidence-caption.js +36 -5
- package/lib/crew-api.js +8 -1
- package/lib/crew-release.sh +51 -9
- package/package.json +2 -2
- package/seed/crons.json +0 -2
- package/workflows/AGENTS.md +3 -1
- package/workflows/bugfix.js +271 -583
- package/workflows/chore.js +261 -516
- package/workflows/crew-init.js +6 -6
- package/workflows/standard.js +264 -541
|
@@ -0,0 +1,234 @@
|
|
|
1
|
+
# Decision history: workflow core
|
|
2
|
+
|
|
3
|
+
Relocated from workflow source comments during H5 (2026-09-18). The workflows keep only the relied-upon invariant inline; the full decision history lives here.
|
|
4
|
+
|
|
5
|
+
<a id="verdict-reask"></a>
|
|
6
|
+
## Verdict re-ask
|
|
7
|
+
|
|
8
|
+
Invariant: a verdict-step report that fails extractVerdict gets up to two bounded re-ask calls; the re-ask agent transcribes, never decides; exhaustion keeps fail-closed behavior.
|
|
9
|
+
|
|
10
|
+
Applies to: standard, bugfix, chore.
|
|
11
|
+
|
|
12
|
+
```
|
|
13
|
+
// Verdict re-ask (bug cd18ccc2): a verdict-step report that fails
|
|
14
|
+
// extractVerdict is not failed immediately. Stochastic verdict-line
|
|
15
|
+
// non-compliance (the agent did the work but omitted or garbled the VERDICT
|
|
16
|
+
// line) gets up to two bounded follow-up agent() calls whose only job is to
|
|
17
|
+
// read the preserved report and emit exactly one VERDICT line. The verdict
|
|
18
|
+
// is still extracted mechanically by extractVerdict — the re-ask agent
|
|
19
|
+
// transcribes, never decides the phase outcome. Each attempt uses a fresh
|
|
20
|
+
// stable-key suffix so a cached failure can never replay deterministically.
|
|
21
|
+
// Exhaustion keeps the existing fail-closed behavior. This is structure, not
|
|
22
|
+
// prompt hardening: no instruction text was stern-ified to get here.
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
<a id="park-contract"></a>
|
|
26
|
+
## Park contract + terminal cleanup
|
|
27
|
+
|
|
28
|
+
Invariant: parking is one atomic parktask action; a failed park reports failed (retryable); terminal cleanup releases the lock and reclaims only fully-merged work.
|
|
29
|
+
|
|
30
|
+
Applies to: standard, bugfix, chore.
|
|
31
|
+
|
|
32
|
+
```
|
|
33
|
+
// Park the task for human attention and end the run. "blocked" is never
|
|
34
|
+
// manually authored — the dashboard derives it mechanically from unmet
|
|
35
|
+
// dependencies — so a workflow outcome that needs a human parks the task
|
|
36
|
+
// instead. Parking is one atomic dashboard action (parktask): the parked
|
|
37
|
+
// state and the explanatory note land in one transaction, never half.
|
|
38
|
+
// The dispatcher skips parked tasks; a human moving parked→todo
|
|
39
|
+
// mechanically resets the retry counters. Returns the workflow result
|
|
40
|
+
// envelope the launcher sees. If the park call itself fails, the run
|
|
41
|
+
// reports "failed" (retryable) so the next tick re-attempts the park —
|
|
42
|
+
// a lost park is never reported as parked.
|
|
43
|
+
// Terminal cleanup: the run's last act at every park/fail boundary. A run
|
|
44
|
+
// that parks or fails must not leak its worktree, branch, or merge lock.
|
|
45
|
+
// The lifecycle's terminal-cleanup releases the lock unconditionally and
|
|
46
|
+
// reclaims the worktree+branch ONLY when the task branch is fully merged
|
|
47
|
+
// into main (then it is redundant); unmerged work is preserved for the
|
|
48
|
+
// human by design. Fire-and-forget with one bounded retry — the merge-lock
|
|
49
|
+
// lease expiry and the orphan sweep are the backstop for a dead transport.
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
<a id="closeout-envelope"></a>
|
|
53
|
+
## Closeout transport envelope
|
|
54
|
+
|
|
55
|
+
Invariant: the work agent returns the runtime's native transport envelope with no schema; the verdict is extracted mechanically by extractVerdict — never by an agent.
|
|
56
|
+
|
|
57
|
+
Applies to: standard, bugfix, chore.
|
|
58
|
+
|
|
59
|
+
```
|
|
60
|
+
// Work agent returns the runtime's native envelope {"status": "ok",
|
|
61
|
+
// "result": "<prose>"} with no schema. The runtime requires JSON output;
|
|
62
|
+
// the envelope is its own documented shape, so there is nothing for the
|
|
63
|
+
// agent to improvise. The workflow receives the prose report as a plain
|
|
64
|
+
// string. The verdict is still extracted deterministically from the report
|
|
65
|
+
// text by extractVerdict below — never by an agent.
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
<a id="extract-verdict-contract"></a>
|
|
69
|
+
## extractVerdict contract
|
|
70
|
+
|
|
71
|
+
Invariant: the verdict is the LAST VERDICT: PASS/FAIL line in the report; extraction is mechanical and deterministic.
|
|
72
|
+
|
|
73
|
+
Applies to: standard, bugfix, chore.
|
|
74
|
+
|
|
75
|
+
```
|
|
76
|
+
// The verdict is the LAST VERDICT: PASS/FAIL in the report (contract: end
|
|
77
|
+
// your report with the verdict). This ignores literal VERDICT strings echoed
|
|
78
|
+
// from the worker's instructions (which contain quoted examples). Fail
|
|
79
|
+
// closed if: no verdict found, the last verdict is not in the trailing 100
|
|
80
|
+
// chars (verdict must be at the end), or conflicting verdicts appear in the
|
|
81
|
+
// trailing 200 chars. Word boundary prevents "PASSING" matching as PASS.
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
<a id="rework-attempt-key"></a>
|
|
85
|
+
## Rework attempt key
|
|
86
|
+
|
|
87
|
+
Invariant: rework attempts are keyed deterministically; the count is a parameter because chore names it differently.
|
|
88
|
+
|
|
89
|
+
Applies to: standard, bugfix, chore.
|
|
90
|
+
|
|
91
|
+
```
|
|
92
|
+
// replay-key scoping for bug 1b8bb875 — runtime keys agent() calls by explicit key per workflow process; on in-process rework a re-executed phase must mint fresh keys; uses the same -r<N> suffix as the phase loop; pure function of inputs, no clock, no randomness (determinism contract). The count is a parameter because chore.js names the counter reworkCount while standard.js/bugfix.js name it totalReworkCount.
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
<a id="ux-doctrine-mirror"></a>
|
|
96
|
+
## UX doctrine mirror
|
|
97
|
+
|
|
98
|
+
Invariant: workflows mirror the UX doctrine map inline (relative-import support unverified); prompts read UX_DOCTRINE_PATH.
|
|
99
|
+
|
|
100
|
+
Applies to: standard, bugfix, chore.
|
|
101
|
+
|
|
102
|
+
```
|
|
103
|
+
// UX doctrine page: the shared UX bar for this run's surface, resolved
|
|
104
|
+
// mechanically — every phase prompt reads UX_DOCTRINE_PATH, never a
|
|
105
|
+
// hardcoded filename. Canonical map: lib/ux-doctrine.js (mirrored here as a
|
|
106
|
+
// one-liner because the workflow runtime's relative-import support is
|
|
107
|
+
// unverified; tests pin the mirror). Null on unclassified surfaces: no
|
|
108
|
+
// shared page, and prompts say so instead of naming the wrong one.
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
<a id="replay-key-scoping"></a>
|
|
112
|
+
## Replay-key scoping
|
|
113
|
+
|
|
114
|
+
Invariant: runtime keys agent() calls by explicit key per workflow process; on in-process rework a re-executed phase mints fresh keys with the -r<N> suffix.
|
|
115
|
+
|
|
116
|
+
Applies to: standard, bugfix, chore.
|
|
117
|
+
|
|
118
|
+
```
|
|
119
|
+
// replay-key scoping for bug 1b8bb875 — runtime keys agent() calls by explicit key per workflow process; on in-process rework a re-executed phase must mint fresh keys; uses the same -r<N> suffix as the phase loop; pure function of inputs, no clock, no randomness (determinism contract). The count is a parameter because chore.js names the counter reworkCount while standard.js/bugfix.js name it totalReworkCount.
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
<a id="transport-retry"></a>
|
|
123
|
+
## Transport retry
|
|
124
|
+
|
|
125
|
+
Invariant: the work-agent agent() call can throw even when the agent did the work; retry boundedly with fresh keys before failing closed; re-entry is safe.
|
|
126
|
+
|
|
127
|
+
Applies to: standard, bugfix, chore.
|
|
128
|
+
|
|
129
|
+
```
|
|
130
|
+
// Transport retry: the work-agent agent() call can throw even when the agent
|
|
131
|
+
// did the work. Stochastic envelope non-compliance (bare prose instead of
|
|
132
|
+
// the native {"status":"ok","result":"..."} envelope) trips the runtime's
|
|
133
|
+
// JSON-candidate heuristic when the prose contains a {...}-looking
|
|
134
|
+
// substring — canary 39457ee9's QA report quoted the change's own
|
|
135
|
+
// {/* ... */} JSX comment, the runtime tried to parse it as JSON, threw,
|
|
136
|
+
// and the workflow discarded a complete VERDICT: PASS report as "no output".
|
|
137
|
+
// The verdict re-ask covers an unreadable verdict inside a RECEIVED report;
|
|
138
|
+
// this covers the report never arriving. The assignment is retried boundedly
|
|
139
|
+
// with fresh keys (never a cached replay) before failing closed. Re-entry is
|
|
140
|
+
// safe: lifecycle scripts answer REUSED for existing worktrees/branches, the
|
|
141
|
+
// retry trailer tells the agent to check existing state first and report
|
|
142
|
+
// rather than duplicate completed side effects, and the rework path already
|
|
143
|
+
// re-runs Build after rejection — Build re-entry is an established pattern.
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
<a id="honest-classification"></a>
|
|
147
|
+
## Honest classification
|
|
148
|
+
|
|
149
|
+
Invariant: classify work-agent calls that yield no usable output honestly.
|
|
150
|
+
|
|
151
|
+
Applies to: standard, bugfix, chore.
|
|
152
|
+
|
|
153
|
+
```
|
|
154
|
+
// Honest classification of a work-agent call that yielded no usable
|
|
155
|
+
// report, with the per-attempt evidence preserved in the session notes.
|
|
156
|
+
// Two distinct cases:
|
|
157
|
+
// - agent() THREW: the runtime discarded output it could not machine-read
|
|
158
|
+
// (e.g. prose tripping the JSON-candidate heuristic). The raw output is
|
|
159
|
+
// gone — the workflow never received it — so the surviving error text is
|
|
160
|
+
// recorded here instead. (An earlier comment claimed the raw output was
|
|
161
|
+
// "preserved in the run record"; that was false — nothing preserved it —
|
|
162
|
+
// and the claim is removed.)
|
|
163
|
+
// - agent() RETURNED EMPTY without throwing: the child produced nothing
|
|
164
|
+
// usable. Each attempt's outcome is the evidence.
|
|
165
|
+
// attempts: [{threw, error, outcome}, ...], in order.
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
<a id="oneshot-recovery"></a>
|
|
169
|
+
## One-shot recovery
|
|
170
|
+
|
|
171
|
+
Invariant: the dispatcher sets input for one-shot recovery routing.
|
|
172
|
+
|
|
173
|
+
Applies to: standard, bugfix, chore.
|
|
174
|
+
|
|
175
|
+
```
|
|
176
|
+
// One-shot recovery routing: the dispatcher sets inputs.next_phase when it
|
|
177
|
+
// routes this run via an explicit recover-task redirect. The value is
|
|
178
|
+
// consumed (cleared) atomically by the successful self-claim below:
|
|
179
|
+
// claim-task takes expected_next_phase and clears the matching next_phase in
|
|
180
|
+
// the same transaction as the winning session insert, so no platform death
|
|
181
|
+
// can slip between claim and consumption and replay the routing. A stale or
|
|
182
|
+
// superseded routing survives — only an exact match clears.
|
|
183
|
+
// what the dispatcher routed on.
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
<a id="discarded-reason"></a>
|
|
187
|
+
## Discarded reason
|
|
188
|
+
|
|
189
|
+
Invariant: the runtime threw the output away; reason is discarded.
|
|
190
|
+
|
|
191
|
+
Applies to: standard, bugfix, chore.
|
|
192
|
+
|
|
193
|
+
```
|
|
194
|
+
// reason: "discarded" (the runtime threw the output away - it could not be
|
|
195
|
+
// machine-read), "empty" (agent() returned without throwing but produced
|
|
196
|
+
// nothing usable), "no-tools" (the worker's TOOL CHECK reported
|
|
197
|
+
// artifact_tools: missing), or "no-transport" (the worker's TOOL CHECK
|
|
198
|
+
// reported shell_transport: unavailable).
|
|
199
|
+
// The trailer tells the retry what to expect, not just to try again.
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
|
|
203
|
+
<a id="h5-probe-blocked"></a>
|
|
204
|
+
## H5 import probe: blocked, extraction refused (2026-09-18)
|
|
205
|
+
|
|
206
|
+
Invariant: workflow scripts carry no imports; shared logic is mirrored inline, never imported.
|
|
207
|
+
|
|
208
|
+
Applies to: standard, bugfix, chore.
|
|
209
|
+
|
|
210
|
+
```
|
|
211
|
+
# H5 (2026-09-18): the Phase 1 plan was probe-gated on the workflow runtime
|
|
212
|
+
# resolving relative imports of a sibling lib/ module. The probe could not
|
|
213
|
+
# be run: no platform launch capability exists in the build environment
|
|
214
|
+
# (workflow_launch is a platform capability, not a shell command; no local
|
|
215
|
+
# launcher exists), and a local node ESM check would prove nothing about the
|
|
216
|
+
# platform runtime — so it was refused, not faked. Outcome: PROBE-BLOCKED.
|
|
217
|
+
# Extraction without the probe would ship an unverified import into a
|
|
218
|
+
# launch-gated script — a worse defect than the size defect. The failover
|
|
219
|
+
# was measured in-file subtraction: delete dead buildVisualCapturePlan and
|
|
220
|
+
# relocate decision-history narration to docs/decisions/ (this file),
|
|
221
|
+
# keeping inline only the relied-upon invariant. The import question stays
|
|
222
|
+
# an OPEN PLATFORM FINDING; helper extraction (and the Phase 2 publish-block
|
|
223
|
+
# unification) is sequenced follow-up once the runtime's module support is
|
|
224
|
+
# proven by a real launch.
|
|
225
|
+
# Probe outcome, recorded 2026-09-18: the helper-extraction probe was
|
|
226
|
+
# refused as unverified — workflow scripts have no module system
|
|
227
|
+
# (typeof require === "undefined" on the platform runtime), so extraction
|
|
228
|
+
# would have shipped an unverified mechanism into a launch-gated script.
|
|
229
|
+
# The relocation-to-docs/decisions/ failover above was used instead.
|
|
230
|
+
# Sequencing: the next publish-path change must be preceded by the Phase 2
|
|
231
|
+
# publish-phase unification (or equivalent subtraction). bugfix.js sits at
|
|
232
|
+
# 244,349 bytes — 1,411 bytes of headroom under the 245,760-byte target —
|
|
233
|
+
# so an ad-hoc cut under gate pressure is not a plan.
|
|
234
|
+
```
|
package/docs/guide.md
CHANGED
|
@@ -17,7 +17,7 @@ This is the full setup and operations reference. If you're new, start with the [
|
|
|
17
17
|
- A **release** — the first immutable snapshot of the crew's runtime code.
|
|
18
18
|
- An **`.orchestration/` directory** in the crew home with identities, personas, workflow docs, and feedback conventions.
|
|
19
19
|
- A **project registration** — the task service registered as its own first project.
|
|
20
|
-
- **Cron jobs** — the scheduler state declared in `seed/crons.json`: a polling loop (every 15 minutes via Muse's scheduling, even when nobody's in the conversation). The scheduler identity is dashboard-independent —
|
|
20
|
+
- **Cron jobs** — the scheduler state declared in `seed/crons.json`: a polling loop (every 15 minutes via Muse's scheduling, even when nobody's in the conversation). The scheduler identity is dashboard-independent — chosen once at first init (no owner; the platform rejects `cli:` owners) — so deleting the dashboard artifact does not stop the crew. Removing a crew entirely is `crew-uninstall`'s job.
|
|
21
21
|
|
|
22
22
|
The agent running in the main chat receives the dispatcher's claims and launches each task workflow. Workflows can't launch workflows, so this handoff is structural.
|
|
23
23
|
|
|
@@ -63,7 +63,7 @@ Optional arguments:
|
|
|
63
63
|
- `dashboardName` — display name for the project registration (default: `"Muse Crew"`).
|
|
64
64
|
- `crewName` — the human's chosen name for the crew. Stored as plain text at `$CREW_HOME/crew-name` and returned in the init summary. The setup conversation should always ask for one; init won't fail without it.
|
|
65
65
|
- `cronIds` — manifest id → live id map for parallel instances on one account (default: `{}`). The manifest's ids are used as-is unless overridden, e.g. `{"crew-poll": "crew-poll-canary"}`.
|
|
66
|
-
- `cronOwner` — explicit owner for the created cron jobs. Defaults to
|
|
66
|
+
- `cronOwner` — explicit owner for the created cron jobs. Defaults to null (omitted): the instance id is chosen at first init (the crew-home basename) and kept from `.cron-registry.json` on re-runs, so re-init never renames the crew's jobs. The platform rejects `cli:` owners.
|
|
67
67
|
- `autoUpdateCrew` — automatic crew upgrades on/off (default: `true`). Persists as the `auto_update_crew` config key; `false` turns the update watcher's crew check off.
|
|
68
68
|
- `autoUpdateDashboard` — automatic dashboard upgrades on/off (default: `true`). Persists as the `auto_update_dashboard` config key.
|
|
69
69
|
- `updateChannel` — `latest` (default) or `patch`: narrows crew upgrades to patch releases. Persists as the `update_channel` config key; any other value fails the init.
|
|
@@ -79,7 +79,7 @@ Init runs five phases, each idempotent — re-running converges anything that dr
|
|
|
79
79
|
|
|
80
80
|
3. **Project registration** — registers the task service as a project in its own database via `createproject`, with `repo_path` set to the validated `dashboardRepoPath`. The dashboard becomes its own first project, so the crew can work on the dashboard itself. Re-running init is self-healing: a project already registered with the correct path is left alone, while one with a different path (e.g. from an older init) is repaired via `update-project`. Skipped entirely in CLI-only mode (no `dashboardSlug`) — create projects with `crew-api.js create-project` instead.
|
|
81
81
|
|
|
82
|
-
4. **Crons** — reads the manifest at `seed/crons.json` and makes the scheduler match it: missing jobs are created from the entry (title, enabled, mode, schedule,
|
|
82
|
+
4. **Crons** — reads the manifest at `seed/crons.json` and makes the scheduler match it: missing jobs are created from the entry (title, enabled, mode, schedule, timeout_secs, and the body built from the entry's template with `crewHome` substituted in — owner is omitted); existing jobs are viewed and converged — drifted fields are updated, unchanged jobs are left alone. The scheduler identity (`instanceId`) is chosen once at first init and never renamed by re-runs: the instance id comes from the existing `.cron-registry.json` when present, otherwise the crew-home basename. One exception: init never touches `enabled` on an existing job. `enabled` is a creation-time default only — a disabled job is a deliberate human decision, and re-init must not silently resurrect it. Jobs removed from the manifest are left alone; deletion is a human decision.
|
|
83
83
|
|
|
84
84
|
5. **Update policy** — persists the automatic-update policy (`auto_update_crew`, `auto_update_dashboard`, `update_channel`) via the Crew API's `update-config`, with the re-init semantics above: explicit inputs win, existing values survive, defaults apply only on first init. Seeds `$CREW_HOME/.update-watch.json` as `{}` only when it does not exist (the watcher's idempotency record is never overwritten).
|
|
85
85
|
|
package/docs/visual-verdict.md
CHANGED
|
@@ -98,13 +98,12 @@ Triggered by the Capture phase when the task is experiential and no
|
|
|
98
98
|
baseline evidence is recorded yet. The workflow logs
|
|
99
99
|
`baseline: requested (attempt N)` and parks; the parent then:
|
|
100
100
|
|
|
101
|
-
1. Run the deterministic capture plan
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
`capture_targets:` line; if none, fall back to the task description.
|
|
101
|
+
1. Run the deterministic capture plan — desktop 1440x900 top, desktop
|
|
102
|
+
1440x900 target-centered, mobile 390x844 target-centered, hover,
|
|
103
|
+
keyboard focus, active/pressed where applicable, plus console error
|
|
104
|
+
count, the ARIA tree of the target region, and any horizontal overflow.
|
|
105
|
+
Capture targets come from the Map step's `capture_targets:` line; if
|
|
106
|
+
none, fall back to the task description.
|
|
108
107
|
2. Save the captures under `$CREW_HOME/task-evidence/<task-id>/baseline/`.
|
|
109
108
|
3. Log the task note event:
|
|
110
109
|
`baseline: captured <space-separated evidence refs>`
|
|
@@ -22,6 +22,14 @@
|
|
|
22
22
|
// Missing deploy_slug, missing dir, or missing PNGs gets the honest
|
|
23
23
|
// no-captures line; terminal surfaces are named, not assumed.
|
|
24
24
|
//
|
|
25
|
+
// Freshness gate (2026-09-18, room #17): seeing the files is not enough —
|
|
26
|
+
// the composer may claim captures ONLY when both PNGs exist AND both file
|
|
27
|
+
// mtimes are strictly newer than task.created_at. That is all the gate
|
|
28
|
+
// proves: captures from an earlier attempt on the same task are also newer
|
|
29
|
+
// than creation, so the gate cannot tell them apart from this run's
|
|
30
|
+
// captures — it excludes only captures predating the task. Missing or
|
|
31
|
+
// unparseable created_at fails closed: the captures are never claimed.
|
|
32
|
+
//
|
|
25
33
|
// Usage:
|
|
26
34
|
// node lib/compose-evidence-caption.js --crew-home <home> \
|
|
27
35
|
// --task-id <id> --audit-dir <resolved-dir-name>
|
|
@@ -93,17 +101,40 @@ const project = (state.projects || []).find((p) => p.id === task.project);
|
|
|
93
101
|
// Capture guard (2026-09-17, room #14): resolve the audit dir the audit
|
|
94
102
|
// harness actually writes and verify the captures exist before claiming
|
|
95
103
|
// them. A bare --audit-dir name is never trusted on its own.
|
|
96
|
-
|
|
97
|
-
|
|
104
|
+
// Parse task.created_at (ISO-8601 UTC from the API, e.g. "...T...Z") to
|
|
105
|
+
// epoch ms. Missing/unparseable values return NaN and the caller fails
|
|
106
|
+
// closed: captures are never claimed without a trustworthy creation time.
|
|
107
|
+
// Note: only the space-form legacy timestamp is normalized to UTC; a
|
|
108
|
+
// T-form WITHOUT a zone marker parses as local time. Neither is a
|
|
109
|
+
// documented input shape, and the machine runs UTC — no live impact.
|
|
110
|
+
function parseCreatedAtMs(s) {
|
|
111
|
+
if (typeof s !== "string" || !s) return NaN;
|
|
112
|
+
// Legacy sqlite rows may carry "YYYY-MM-DD HH:MM:SS" (UTC, no zone
|
|
113
|
+
// marker), which V8 parses as LOCAL time — a timezone-sized skew that
|
|
114
|
+
// could fail the gate open. Normalize the space form to UTC.
|
|
115
|
+
const iso = s.includes("T") ? s : s.replace(" ", "T") + "Z";
|
|
116
|
+
const ms = Date.parse(iso);
|
|
117
|
+
return Number.isFinite(ms) ? ms : NaN;
|
|
118
|
+
}
|
|
119
|
+
// Strictly newer: an mtime exactly equal to created_at is stale — the
|
|
120
|
+
// capture cannot predate the task it evidences, but "same instant" proves
|
|
121
|
+
// nothing. (Only "newer than task creation" is proven: captures from an
|
|
122
|
+
// earlier attempt on the same task pass this check too.)
|
|
123
|
+
function isFreshFile(p, createdAtMs) {
|
|
124
|
+
try {
|
|
125
|
+
const st = fs.statSync(p);
|
|
126
|
+
return st.isFile() && st.mtimeMs > createdAtMs;
|
|
127
|
+
} catch { return false; }
|
|
98
128
|
}
|
|
99
129
|
let auditPath = null;
|
|
100
130
|
if (project && project.deploy_slug) {
|
|
101
131
|
auditPath = path.join(os.homedir(), "workspace", "ts-spaces",
|
|
102
132
|
project.deploy_slug, "audits", auditDir);
|
|
103
133
|
}
|
|
104
|
-
const
|
|
105
|
-
|
|
106
|
-
|
|
134
|
+
const createdAtMs = parseCreatedAtMs(task.created_at);
|
|
135
|
+
const hasCaptures = !!(auditPath && Number.isFinite(createdAtMs) &&
|
|
136
|
+
isFreshFile(path.join(auditPath, "screenshot.png"), createdAtMs) &&
|
|
137
|
+
isFreshFile(path.join(auditPath, "screenshot-mobile.png"), createdAtMs));
|
|
107
138
|
|
|
108
139
|
let events = [];
|
|
109
140
|
try {
|
package/lib/crew-api.js
CHANGED
|
@@ -1675,6 +1675,13 @@ commands["resolve-publish-unknown"] = (db, args, ctx) => {
|
|
|
1675
1675
|
if (evidenceDirs.length === 0) {
|
|
1676
1676
|
return { resolved: false, reason: "no audit build completed inside the publish window — still unknown" };
|
|
1677
1677
|
}
|
|
1678
|
+
// (N3) Ambiguity guard: multiple builds in the recovery window have the
|
|
1679
|
+
// same attribution ambiguity as the workflow's immediate verdict — no
|
|
1680
|
+
// single dir can be attributed to this attempt, so resolution stays
|
|
1681
|
+
// unknown. Only when exactly one dir remains is evidenceDirs[0] safe.
|
|
1682
|
+
if (evidenceDirs.length > 1) {
|
|
1683
|
+
return { resolved: false, reason: `audit-dir ambiguity: ${evidenceDirs.length} audit dirs inside the publish window — still unknown` };
|
|
1684
|
+
}
|
|
1678
1685
|
const evidenceDir = evidenceDirs[0];
|
|
1679
1686
|
|
|
1680
1687
|
// Corroboration, observation only: the audit harness's own verdict.
|
|
@@ -1700,7 +1707,7 @@ commands["resolve-publish-unknown"] = (db, args, ctx) => {
|
|
|
1700
1707
|
db.prepare(
|
|
1701
1708
|
"INSERT INTO events (id, type, task_id, identity, message, timestamp) VALUES (?, 'note', ?, NULL, ?, ?)"
|
|
1702
1709
|
).run(uuid(), taskId,
|
|
1703
|
-
`Parked: publish: verification-requested ${latest.commit} (build agent_id unobserved) — recovered from publish-unknown via resolve-publish-unknown; durable evidence ${evidenceDir}.
|
|
1710
|
+
`Parked: publish: verification-requested ${latest.commit} (build agent_id unobserved) — recovered from publish-unknown via resolve-publish-unknown; durable evidence ${evidenceDir}. waiting on the manual read-back in docs/publish-verification.md.`,
|
|
1704
1711
|
tsLater);
|
|
1705
1712
|
db.prepare("UPDATE tasks SET state = 'parked', updated_at = ? WHERE id = ?").run(tsLater, taskId);
|
|
1706
1713
|
db.exec("COMMIT");
|
package/lib/crew-release.sh
CHANGED
|
@@ -73,9 +73,12 @@ cmd_init() {
|
|
|
73
73
|
echo "CURRENT: $cur"
|
|
74
74
|
}
|
|
75
75
|
|
|
76
|
-
# ── workflow
|
|
77
|
-
# Every workflow script must parse
|
|
76
|
+
# ── workflow validation ─────────────────────────────────────────────
|
|
77
|
+
# Every workflow script must (a) parse under the workflow-loader emulation
|
|
78
|
+
# and (b) fit the launch budget before a release can install.
|
|
79
|
+
# Exit codes: 10 = parse check failed, 20 = size check failed (0 = pass).
|
|
78
80
|
_validate_workflows() {
|
|
81
|
+
# --- Parse gate (restored 2026-09-18, finding B2) ---
|
|
79
82
|
# The workflow runtime accepts top-level `export`, `return`, and `await`
|
|
80
83
|
# (it wraps scripts in an async function), so neither `node --check` on the
|
|
81
84
|
# raw .js (vacuous for ESM — exits 0 even on blatant syntax errors) nor a
|
|
@@ -83,8 +86,23 @@ _validate_workflows() {
|
|
|
83
86
|
# Emulate the runtime instead: strip `export`, wrap the script in an async
|
|
84
87
|
# function, then node --check the result. Tokenizer errors (e.g. an
|
|
85
88
|
# unterminated string literal) still fail under the wrap.
|
|
89
|
+
#
|
|
90
|
+
# HONESTY NOTE: this is a provisional emulation, not a verification. The
|
|
91
|
+
# platform's actual loading mechanism was never verified (probe blocked,
|
|
92
|
+
# 2026-09-18), so this check cannot prove how the platform loads the
|
|
93
|
+
# workflows. What it DOES prove: the file is syntactically well-formed JS
|
|
94
|
+
# when read the way the runtime claims to read it (strip export, run as one
|
|
95
|
+
# async function body). What it does NOT prove: that the platform actually
|
|
96
|
+
# loads it that way, or that the code runs correctly. The export-strip +
|
|
97
|
+
# async-wrap transform exists precisely because raw `node --check` on ESM
|
|
98
|
+
# `.js` is unreliable — it silently exits 0 on some malformed input — so
|
|
99
|
+
# the transform is what makes the check actually parse.
|
|
100
|
+
# A failure here is a signal to INSPECT THE CODE, never to delete the
|
|
101
|
+
# check: this gate was once removed on a misdiagnosis ("the check
|
|
102
|
+
# false-positives"), and the workflows it rejected turned out to be
|
|
103
|
+
# genuinely broken (2026-09-18, Gate 1 H2 — a true positive).
|
|
86
104
|
local dir="$1"
|
|
87
|
-
local f tmp base
|
|
105
|
+
local f tmp base size
|
|
88
106
|
tmp="$(mktemp -d)"
|
|
89
107
|
for f in "$dir"/workflows/*.js; do
|
|
90
108
|
[ -f "$f" ] || continue
|
|
@@ -96,13 +114,28 @@ _validate_workflows() {
|
|
|
96
114
|
} > "$tmp/$base.js"
|
|
97
115
|
if ! node --check "$tmp/$base.js" 2>"$tmp/$base.err"; then
|
|
98
116
|
cat "$tmp/$base.err" >&2
|
|
99
|
-
echo "VALIDATION FAILED: $f does not parse" >&2
|
|
117
|
+
echo "VALIDATION FAILED [parse]: $f does not parse under the workflow-loader emulation" >&2
|
|
100
118
|
rm -rf "$tmp"
|
|
101
|
-
return
|
|
119
|
+
return 10
|
|
102
120
|
fi
|
|
103
121
|
done
|
|
104
122
|
rm -rf "$tmp"
|
|
105
|
-
|
|
123
|
+
|
|
124
|
+
# --- Size gate (H5, 2026-09-18) ---
|
|
125
|
+
# The platform reads the full workflow file on launch; the platform hard
|
|
126
|
+
# maximum is 262144 bytes (256 KiB) and 245760 bytes (240 KiB) is the
|
|
127
|
+
# project's safety target. The release fails closed if any launchable
|
|
128
|
+
# workflow hits the safety target.
|
|
129
|
+
for f in "$dir"/workflows/*.js; do
|
|
130
|
+
[ -f "$f" ] || continue
|
|
131
|
+
size="$(wc -c < "$f")"
|
|
132
|
+
if [ "$size" -ge 245760 ]; then
|
|
133
|
+
echo "VALIDATION FAILED [size]: $f is $size bytes (>= 245760 byte launch budget)" >&2
|
|
134
|
+
return 20
|
|
135
|
+
fi
|
|
136
|
+
done
|
|
137
|
+
|
|
138
|
+
echo "VALIDATED: workflow scripts parse and fit the launch budget"
|
|
106
139
|
}
|
|
107
140
|
|
|
108
141
|
# ── workflow doc sync ───────────────────────────────────────────────
|
|
@@ -255,10 +288,19 @@ cmd_deploy() {
|
|
|
255
288
|
rm -rf "$staging_dir"
|
|
256
289
|
die "release $hash rejected: registry build failed"
|
|
257
290
|
fi
|
|
258
|
-
# Gate: refuse to install a release whose workflow scripts
|
|
259
|
-
|
|
291
|
+
# Gate: refuse to install a release whose workflow scripts fail the
|
|
292
|
+
# parse or size checks. _validate_workflows returns 10 on a parse
|
|
293
|
+
# failure, 20 on a size-budget failure; the rejection names the check
|
|
294
|
+
# that failed so the message is never misleading.
|
|
295
|
+
_vw_status=0
|
|
296
|
+
_validate_workflows "$staging_dir" || _vw_status=$?
|
|
297
|
+
if [ "$_vw_status" -ne 0 ]; then
|
|
260
298
|
rm -rf "$staging_dir"
|
|
261
|
-
|
|
299
|
+
case "$_vw_status" in
|
|
300
|
+
10) die "release $hash rejected: workflow parse check failed" ;;
|
|
301
|
+
20) die "release $hash rejected: workflow size-budget check failed" ;;
|
|
302
|
+
*) die "release $hash rejected: workflow validation failed (unexpected exit $_vw_status)" ;;
|
|
303
|
+
esac
|
|
262
304
|
fi
|
|
263
305
|
# Atomic rename into place
|
|
264
306
|
mv "$staging_dir" "$release_dir"
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "muse-crew",
|
|
3
|
-
"version": "0.13.
|
|
3
|
+
"version": "0.13.3",
|
|
4
4
|
"description": "Opinionated orchestration for Muse \u2014 workflows, identities, and tooling for autonomous software development.",
|
|
5
5
|
"license": "UNLICENSED",
|
|
6
6
|
"private": false,
|
|
@@ -29,4 +29,4 @@
|
|
|
29
29
|
"dependencies": {
|
|
30
30
|
"playwright-core": "1.63.0"
|
|
31
31
|
}
|
|
32
|
-
}
|
|
32
|
+
}
|
package/seed/crons.json
CHANGED
|
@@ -5,7 +5,6 @@
|
|
|
5
5
|
"enabled": true,
|
|
6
6
|
"id": "crew-poll",
|
|
7
7
|
"mode": "task",
|
|
8
|
-
"owner": "cli:{instanceId}",
|
|
9
8
|
"schedule": {
|
|
10
9
|
"every": "15m",
|
|
11
10
|
"kind": "interval"
|
|
@@ -18,7 +17,6 @@
|
|
|
18
17
|
"enabled": true,
|
|
19
18
|
"id": "crew-update-watch",
|
|
20
19
|
"mode": "task",
|
|
21
|
-
"owner": "cli:{instanceId}",
|
|
22
20
|
"schedule": {
|
|
23
21
|
"every": "24h",
|
|
24
22
|
"kind": "interval"
|
package/workflows/AGENTS.md
CHANGED
|
@@ -4,7 +4,9 @@ Executable Muse workflow scripts (JavaScript). These are what the workflow runti
|
|
|
4
4
|
|
|
5
5
|
- `crew-dispatch.js` — reads the board, recommends eligible tasks, returns structured launch records (the launched workflow self-claims; the dispatcher never writes claims). Skips tasks on quiesced projects and on projects with no repo_path configured.
|
|
6
6
|
|
|
7
|
-
All four workflow scripts share byte-identical transport helpers (`workRetryKey`, `buildTransportRetryTrailer`, `describeWorkAgentFailure
|
|
7
|
+
All four workflow scripts share byte-identical transport helpers (`workRetryKey`, `buildTransportRetryTrailer`, `describeWorkAgentFailure` — pinned by `tests/closeout.test.js`). The transport-retry loop treats a worker report naming the missing artifact tool namespace (bug 3472bf36, a per-launch platform flake) as a retryable attempt with a fresh launch rather than accepting a useless report.
|
|
8
|
+
|
|
9
|
+
**Growth budget (2026-09-18):** the platform reads the full workflow file on launch; the platform hard maximum is 262144 bytes (256 KiB) and the project's safety target is 245760 bytes (240 KiB). Any change adding net bytes to a workflow already over 240 KiB must subtract ≥ the addition in the same change — see `tests/workflow-size.test.js` and the release-time gate in `lib/crew-release.sh` `_validate_workflows`. Decision-history narration lives in `docs/decisions/`; workflows keep only the relied-upon invariant inline. See `docs/decisions/workflow-core.md#h5-probe-blocked` for why helper extraction was refused and decision history was relocated instead. Target-vs-cap (2026-09-18): the release gate intentionally enforces the 245760-byte target, not just the 262144-byte cap — safe because target ≤ cap, stricter by design.
|
|
8
10
|
- `crew-init.js` — sets up a new crew instance: Gate 0 validates crewHome is workspace-contained and a valid git repo, and (in dashboard mode) validates dashboardRepoPath is workspace-contained and an existing git repo (never auto-created); bootstraps the release system; scaffolds orchestration folders; registers the first project with repo_path set to the validated dashboard repo — skipped in CLI-only mode (no dashboardSlug/dashboardRepoPath), where projects are created later via crew-api.js; creates/converges cron jobs from the declarative manifest (scheduler identity is dashboard-independent: first init uses the crew-home basename, re-runs keep the instanceId recorded in `.cron-registry.json`; owner `cli:<instanceId>` — deleting a dashboard artifact never stops the crew). Idempotent. The crew owns its state via the Crew API (lib/crew-api.js); the dashboard is an optional client and is never a dependency.
|
|
9
11
|
- `standard.js` — default task workflow: Triage → Capture → Map → Build → Review → Integrate → Publish → QA
|
|
10
12
|
- `bugfix.js` — adds Capture after Triage, then Reproduce
|