muse-crew 0.13.1 → 0.13.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,234 @@
1
+ # Decision history: workflow core
2
+
3
+ Relocated from workflow source comments during H5 (2026-09-18). The workflows keep only the relied-upon invariant inline; the full decision history lives here.
4
+
5
+ <a id="verdict-reask"></a>
6
+ ## Verdict re-ask
7
+
8
+ Invariant: a verdict-step report that fails extractVerdict gets up to two bounded re-ask calls; the re-ask agent transcribes, never decides; exhaustion keeps fail-closed behavior.
9
+
10
+ Applies to: standard, bugfix, chore.
11
+
12
+ ```
13
+ // Verdict re-ask (bug cd18ccc2): a verdict-step report that fails
14
+ // extractVerdict is not failed immediately. Stochastic verdict-line
15
+ // non-compliance (the agent did the work but omitted or garbled the VERDICT
16
+ // line) gets up to two bounded follow-up agent() calls whose only job is to
17
+ // read the preserved report and emit exactly one VERDICT line. The verdict
18
+ // is still extracted mechanically by extractVerdict — the re-ask agent
19
+ // transcribes, never decides the phase outcome. Each attempt uses a fresh
20
+ // stable-key suffix so a cached failure can never replay deterministically.
21
+ // Exhaustion keeps the existing fail-closed behavior. This is structure, not
22
+ // prompt hardening: no instruction text was stern-ified to get here.
23
+ ```
24
+
25
+ <a id="park-contract"></a>
26
+ ## Park contract + terminal cleanup
27
+
28
+ Invariant: parking is one atomic parktask action; a failed park reports failed (retryable); terminal cleanup releases the lock and reclaims only fully-merged work.
29
+
30
+ Applies to: standard, bugfix, chore.
31
+
32
+ ```
33
+ // Park the task for human attention and end the run. "blocked" is never
34
+ // manually authored — the dashboard derives it mechanically from unmet
35
+ // dependencies — so a workflow outcome that needs a human parks the task
36
+ // instead. Parking is one atomic dashboard action (parktask): the parked
37
+ // state and the explanatory note land in one transaction, never half.
38
+ // The dispatcher skips parked tasks; a human moving parked→todo
39
+ // mechanically resets the retry counters. Returns the workflow result
40
+ // envelope the launcher sees. If the park call itself fails, the run
41
+ // reports "failed" (retryable) so the next tick re-attempts the park —
42
+ // a lost park is never reported as parked.
43
+ // Terminal cleanup: the run's last act at every park/fail boundary. A run
44
+ // that parks or fails must not leak its worktree, branch, or merge lock.
45
+ // The lifecycle's terminal-cleanup releases the lock unconditionally and
46
+ // reclaims the worktree+branch ONLY when the task branch is fully merged
47
+ // into main (then it is redundant); unmerged work is preserved for the
48
+ // human by design. Fire-and-forget with one bounded retry — the merge-lock
49
+ // lease expiry and the orphan sweep are the backstop for a dead transport.
50
+ ```
51
+
52
+ <a id="closeout-envelope"></a>
53
+ ## Closeout transport envelope
54
+
55
+ Invariant: the work agent returns the runtime's native transport envelope with no schema; the verdict is extracted mechanically by extractVerdict — never by an agent.
56
+
57
+ Applies to: standard, bugfix, chore.
58
+
59
+ ```
60
+ // Work agent returns the runtime's native envelope {"status": "ok",
61
+ // "result": "<prose>"} with no schema. The runtime requires JSON output;
62
+ // the envelope is its own documented shape, so there is nothing for the
63
+ // agent to improvise. The workflow receives the prose report as a plain
64
+ // string. The verdict is still extracted deterministically from the report
65
+ // text by extractVerdict below — never by an agent.
66
+ ```
67
+
68
+ <a id="extract-verdict-contract"></a>
69
+ ## extractVerdict contract
70
+
71
+ Invariant: the verdict is the LAST VERDICT: PASS/FAIL line in the report; extraction is mechanical and deterministic.
72
+
73
+ Applies to: standard, bugfix, chore.
74
+
75
+ ```
76
+ // The verdict is the LAST VERDICT: PASS/FAIL in the report (contract: end
77
+ // your report with the verdict). This ignores literal VERDICT strings echoed
78
+ // from the worker's instructions (which contain quoted examples). Fail
79
+ // closed if: no verdict found, the last verdict is not in the trailing 100
80
+ // chars (verdict must be at the end), or conflicting verdicts appear in the
81
+ // trailing 200 chars. Word boundary prevents "PASSING" matching as PASS.
82
+ ```
83
+
84
+ <a id="rework-attempt-key"></a>
85
+ ## Rework attempt key
86
+
87
+ Invariant: rework attempts are keyed deterministically; the count is a parameter because chore names it differently.
88
+
89
+ Applies to: standard, bugfix, chore.
90
+
91
+ ```
92
+ // replay-key scoping for bug 1b8bb875 — runtime keys agent() calls by explicit key per workflow process; on in-process rework a re-executed phase must mint fresh keys; uses the same -r<N> suffix as the phase loop; pure function of inputs, no clock, no randomness (determinism contract). The count is a parameter because chore.js names the counter reworkCount while standard.js/bugfix.js name it totalReworkCount.
93
+ ```
94
+
95
+ <a id="ux-doctrine-mirror"></a>
96
+ ## UX doctrine mirror
97
+
98
+ Invariant: workflows mirror the UX doctrine map inline (relative-import support unverified); prompts read UX_DOCTRINE_PATH.
99
+
100
+ Applies to: standard, bugfix, chore.
101
+
102
+ ```
103
+ // UX doctrine page: the shared UX bar for this run's surface, resolved
104
+ // mechanically — every phase prompt reads UX_DOCTRINE_PATH, never a
105
+ // hardcoded filename. Canonical map: lib/ux-doctrine.js (mirrored here as a
106
+ // one-liner because the workflow runtime's relative-import support is
107
+ // unverified; tests pin the mirror). Null on unclassified surfaces: no
108
+ // shared page, and prompts say so instead of naming the wrong one.
109
+ ```
110
+
111
+ <a id="replay-key-scoping"></a>
112
+ ## Replay-key scoping
113
+
114
+ Invariant: runtime keys agent() calls by explicit key per workflow process; on in-process rework a re-executed phase mints fresh keys with the -r<N> suffix.
115
+
116
+ Applies to: standard, bugfix, chore.
117
+
118
+ ```
119
+ // replay-key scoping for bug 1b8bb875 — runtime keys agent() calls by explicit key per workflow process; on in-process rework a re-executed phase must mint fresh keys; uses the same -r<N> suffix as the phase loop; pure function of inputs, no clock, no randomness (determinism contract). The count is a parameter because chore.js names the counter reworkCount while standard.js/bugfix.js name it totalReworkCount.
120
+ ```
121
+
122
+ <a id="transport-retry"></a>
123
+ ## Transport retry
124
+
125
+ Invariant: the work-agent agent() call can throw even when the agent did the work; retry boundedly with fresh keys before failing closed; re-entry is safe.
126
+
127
+ Applies to: standard, bugfix, chore.
128
+
129
+ ```
130
+ // Transport retry: the work-agent agent() call can throw even when the agent
131
+ // did the work. Stochastic envelope non-compliance (bare prose instead of
132
+ // the native {"status":"ok","result":"..."} envelope) trips the runtime's
133
+ // JSON-candidate heuristic when the prose contains a {...}-looking
134
+ // substring — canary 39457ee9's QA report quoted the change's own
135
+ // {/* ... */} JSX comment, the runtime tried to parse it as JSON, threw,
136
+ // and the workflow discarded a complete VERDICT: PASS report as "no output".
137
+ // The verdict re-ask covers an unreadable verdict inside a RECEIVED report;
138
+ // this covers the report never arriving. The assignment is retried boundedly
139
+ // with fresh keys (never a cached replay) before failing closed. Re-entry is
140
+ // safe: lifecycle scripts answer REUSED for existing worktrees/branches, the
141
+ // retry trailer tells the agent to check existing state first and report
142
+ // rather than duplicate completed side effects, and the rework path already
143
+ // re-runs Build after rejection — Build re-entry is an established pattern.
144
+ ```
145
+
146
+ <a id="honest-classification"></a>
147
+ ## Honest classification
148
+
149
+ Invariant: classify work-agent calls that yield no usable output honestly.
150
+
151
+ Applies to: standard, bugfix, chore.
152
+
153
+ ```
154
+ // Honest classification of a work-agent call that yielded no usable
155
+ // report, with the per-attempt evidence preserved in the session notes.
156
+ // Two distinct cases:
157
+ // - agent() THREW: the runtime discarded output it could not machine-read
158
+ // (e.g. prose tripping the JSON-candidate heuristic). The raw output is
159
+ // gone — the workflow never received it — so the surviving error text is
160
+ // recorded here instead. (An earlier comment claimed the raw output was
161
+ // "preserved in the run record"; that was false — nothing preserved it —
162
+ // and the claim is removed.)
163
+ // - agent() RETURNED EMPTY without throwing: the child produced nothing
164
+ // usable. Each attempt's outcome is the evidence.
165
+ // attempts: [{threw, error, outcome}, ...], in order.
166
+ ```
167
+
168
+ <a id="oneshot-recovery"></a>
169
+ ## One-shot recovery
170
+
171
+ Invariant: the dispatcher sets input for one-shot recovery routing.
172
+
173
+ Applies to: standard, bugfix, chore.
174
+
175
+ ```
176
+ // One-shot recovery routing: the dispatcher sets inputs.next_phase when it
177
+ // routes this run via an explicit recover-task redirect. The value is
178
+ // consumed (cleared) atomically by the successful self-claim below:
179
+ // claim-task takes expected_next_phase and clears the matching next_phase in
180
+ // the same transaction as the winning session insert, so no platform death
181
+ // can slip between claim and consumption and replay the routing. A stale or
182
+ // superseded routing survives — only an exact match clears.
183
+ // what the dispatcher routed on.
184
+ ```
185
+
186
+ <a id="discarded-reason"></a>
187
+ ## Discarded reason
188
+
189
+ Invariant: the runtime threw the output away; reason is discarded.
190
+
191
+ Applies to: standard, bugfix, chore.
192
+
193
+ ```
194
+ // reason: "discarded" (the runtime threw the output away - it could not be
195
+ // machine-read), "empty" (agent() returned without throwing but produced
196
+ // nothing usable), "no-tools" (the worker's TOOL CHECK reported
197
+ // artifact_tools: missing), or "no-transport" (the worker's TOOL CHECK
198
+ // reported shell_transport: unavailable).
199
+ // The trailer tells the retry what to expect, not just to try again.
200
+ ```
201
+
202
+
203
+ <a id="h5-probe-blocked"></a>
204
+ ## H5 import probe: blocked, extraction refused (2026-09-18)
205
+
206
+ Invariant: workflow scripts carry no imports; shared logic is mirrored inline, never imported.
207
+
208
+ Applies to: standard, bugfix, chore.
209
+
210
+ ```
211
+ # H5 (2026-09-18): the Phase 1 plan was probe-gated on the workflow runtime
212
+ # resolving relative imports of a sibling lib/ module. The probe could not
213
+ # be run: no platform launch capability exists in the build environment
214
+ # (workflow_launch is a platform capability, not a shell command; no local
215
+ # launcher exists), and a local node ESM check would prove nothing about the
216
+ # platform runtime — so it was refused, not faked. Outcome: PROBE-BLOCKED.
217
+ # Extraction without the probe would ship an unverified import into a
218
+ # launch-gated script — a worse defect than the size defect. The failover
219
+ # was measured in-file subtraction: delete dead buildVisualCapturePlan and
220
+ # relocate decision-history narration to docs/decisions/ (this file),
221
+ # keeping inline only the relied-upon invariant. The import question stays
222
+ # an OPEN PLATFORM FINDING; helper extraction (and the Phase 2 publish-block
223
+ # unification) is sequenced follow-up once the runtime's module support is
224
+ # proven by a real launch.
225
+ # Probe outcome, recorded 2026-09-18: the helper-extraction probe was
226
+ # refused as unverified — workflow scripts have no module system
227
+ # (typeof require === "undefined" on the platform runtime), so extraction
228
+ # would have shipped an unverified mechanism into a launch-gated script.
229
+ # The relocation-to-docs/decisions/ failover above was used instead.
230
+ # Sequencing: the next publish-path change must be preceded by the Phase 2
231
+ # publish-phase unification (or equivalent subtraction). bugfix.js sits at
232
+ # 244,349 bytes — 1,411 bytes of headroom under the 245,760-byte target —
233
+ # so an ad-hoc cut under gate pressure is not a plan.
234
+ ```
package/docs/guide.md CHANGED
@@ -17,7 +17,7 @@ This is the full setup and operations reference. If you're new, start with the [
17
17
  - A **release** — the first immutable snapshot of the crew's runtime code.
18
18
  - An **`.orchestration/` directory** in the crew home with identities, personas, workflow docs, and feedback conventions.
19
19
  - A **project registration** — the task service registered as its own first project.
20
- - **Cron jobs** — the scheduler state declared in `seed/crons.json`: a polling loop (every 15 minutes via Muse's scheduling, even when nobody's in the conversation). The scheduler identity is dashboard-independent — owner `cli:<instance-id>`, chosen once at first init — so deleting the dashboard artifact does not stop the crew. Removing a crew entirely is `crew-uninstall`'s job.
20
+ - **Cron jobs** — the scheduler state declared in `seed/crons.json`: a polling loop (every 15 minutes via Muse's scheduling, even when nobody's in the conversation). The scheduler identity is dashboard-independent — chosen once at first init (no owner; the platform rejects `cli:` owners) — so deleting the dashboard artifact does not stop the crew. Removing a crew entirely is `crew-uninstall`'s job.
21
21
 
22
22
  The agent running in the main chat receives the dispatcher's claims and launches each task workflow. Workflows can't launch workflows, so this handoff is structural.
23
23
 
@@ -63,7 +63,7 @@ Optional arguments:
63
63
  - `dashboardName` — display name for the project registration (default: `"Muse Crew"`).
64
64
  - `crewName` — the human's chosen name for the crew. Stored as plain text at `$CREW_HOME/crew-name` and returned in the init summary. The setup conversation should always ask for one; init won't fail without it.
65
65
  - `cronIds` — manifest id → live id map for parallel instances on one account (default: `{}`). The manifest's ids are used as-is unless overridden, e.g. `{"crew-poll": "crew-poll-canary"}`.
66
- - `cronOwner` — explicit owner for the created cron jobs. Defaults to `cli:<instance-id>`: the instance id is chosen at first init (the crew-home basename) and kept from `.cron-registry.json` on re-runs, so re-init never renames the crew's jobs.
66
+ - `cronOwner` — explicit owner for the created cron jobs. Defaults to null (omitted): the instance id is chosen at first init (the crew-home basename) and kept from `.cron-registry.json` on re-runs, so re-init never renames the crew's jobs. The platform rejects `cli:` owners.
67
67
  - `autoUpdateCrew` — automatic crew upgrades on/off (default: `true`). Persists as the `auto_update_crew` config key; `false` turns the update watcher's crew check off.
68
68
  - `autoUpdateDashboard` — automatic dashboard upgrades on/off (default: `true`). Persists as the `auto_update_dashboard` config key.
69
69
  - `updateChannel` — `latest` (default) or `patch`: narrows crew upgrades to patch releases. Persists as the `update_channel` config key; any other value fails the init.
@@ -79,7 +79,7 @@ Init runs five phases, each idempotent — re-running converges anything that dr
79
79
 
80
80
  3. **Project registration** — registers the task service as a project in its own database via `createproject`, with `repo_path` set to the validated `dashboardRepoPath`. The dashboard becomes its own first project, so the crew can work on the dashboard itself. Re-running init is self-healing: a project already registered with the correct path is left alone, while one with a different path (e.g. from an older init) is repaired via `update-project`. Skipped entirely in CLI-only mode (no `dashboardSlug`) — create projects with `crew-api.js create-project` instead.
81
81
 
82
- 4. **Crons** — reads the manifest at `seed/crons.json` and makes the scheduler match it: missing jobs are created from the entry (title, enabled, mode, schedule, owner, timeout_secs, and the body built from the entry's template with `crewHome` substituted in); existing jobs are viewed and converged — drifted fields are updated, unchanged jobs are left alone. The scheduler identity (`instanceId`, `cli:<instanceId>` owner) is chosen once at first init and never renamed by re-runs: the instance id comes from the existing `.cron-registry.json` when present, otherwise the crew-home basename. One exception: init never touches `enabled` on an existing job. `enabled` is a creation-time default only — a disabled job is a deliberate human decision, and re-init must not silently resurrect it. Jobs removed from the manifest are left alone; deletion is a human decision.
82
+ 4. **Crons** — reads the manifest at `seed/crons.json` and makes the scheduler match it: missing jobs are created from the entry (title, enabled, mode, schedule, timeout_secs, and the body built from the entry's template with `crewHome` substituted in — owner is omitted); existing jobs are viewed and converged — drifted fields are updated, unchanged jobs are left alone. The scheduler identity (`instanceId`) is chosen once at first init and never renamed by re-runs: the instance id comes from the existing `.cron-registry.json` when present, otherwise the crew-home basename. One exception: init never touches `enabled` on an existing job. `enabled` is a creation-time default only — a disabled job is a deliberate human decision, and re-init must not silently resurrect it. Jobs removed from the manifest are left alone; deletion is a human decision.
83
83
 
84
84
  5. **Update policy** — persists the automatic-update policy (`auto_update_crew`, `auto_update_dashboard`, `update_channel`) via the Crew API's `update-config`, with the re-init semantics above: explicit inputs win, existing values survive, defaults apply only on first init. Seeds `$CREW_HOME/.update-watch.json` as `{}` only when it does not exist (the watcher's idempotency record is never overwritten).
85
85
 
@@ -98,13 +98,12 @@ Triggered by the Capture phase when the task is experiential and no
98
98
  baseline evidence is recorded yet. The workflow logs
99
99
  `baseline: requested (attempt N)` and parks; the parent then:
100
100
 
101
- 1. Run the deterministic capture plan:
102
- `buildVisualCapturePlan(taskTitle, taskDescription, "baseline", captureTargets)`
103
- — desktop 1440x900 top, desktop 1440x900 target-centered, mobile 390x844
104
- target-centered, hover, keyboard focus, active/pressed where applicable,
105
- plus console error count, the ARIA tree of the target region, and any
106
- horizontal overflow. Capture targets come from the Map step's
107
- `capture_targets:` line; if none, fall back to the task description.
101
+ 1. Run the deterministic capture plan — desktop 1440x900 top, desktop
102
+ 1440x900 target-centered, mobile 390x844 target-centered, hover,
103
+ keyboard focus, active/pressed where applicable, plus console error
104
+ count, the ARIA tree of the target region, and any horizontal overflow.
105
+ Capture targets come from the Map step's `capture_targets:` line; if
106
+ none, fall back to the task description.
108
107
  2. Save the captures under `$CREW_HOME/task-evidence/<task-id>/baseline/`.
109
108
  3. Log the task note event:
110
109
  `baseline: captured <space-separated evidence refs>`
@@ -22,6 +22,14 @@
22
22
  // Missing deploy_slug, missing dir, or missing PNGs gets the honest
23
23
  // no-captures line; terminal surfaces are named, not assumed.
24
24
  //
25
+ // Freshness gate (2026-09-18, room #17): seeing the files is not enough —
26
+ // the composer may claim captures ONLY when both PNGs exist AND both file
27
+ // mtimes are strictly newer than task.created_at. That is all the gate
28
+ // proves: captures from an earlier attempt on the same task are also newer
29
+ // than creation, so the gate cannot tell them apart from this run's
30
+ // captures — it excludes only captures predating the task. Missing or
31
+ // unparseable created_at fails closed: the captures are never claimed.
32
+ //
25
33
  // Usage:
26
34
  // node lib/compose-evidence-caption.js --crew-home <home> \
27
35
  // --task-id <id> --audit-dir <resolved-dir-name>
@@ -93,17 +101,40 @@ const project = (state.projects || []).find((p) => p.id === task.project);
93
101
  // Capture guard (2026-09-17, room #14): resolve the audit dir the audit
94
102
  // harness actually writes and verify the captures exist before claiming
95
103
  // them. A bare --audit-dir name is never trusted on its own.
96
- function isFile(p) {
97
- try { return fs.statSync(p).isFile(); } catch { return false; }
104
+ // Parse task.created_at (ISO-8601 UTC from the API, e.g. "...T...Z") to
105
+ // epoch ms. Missing/unparseable values return NaN and the caller fails
106
+ // closed: captures are never claimed without a trustworthy creation time.
107
+ // Note: only the space-form legacy timestamp is normalized to UTC; a
108
+ // T-form WITHOUT a zone marker parses as local time. Neither is a
109
+ // documented input shape, and the machine runs UTC — no live impact.
110
+ function parseCreatedAtMs(s) {
111
+ if (typeof s !== "string" || !s) return NaN;
112
+ // Legacy sqlite rows may carry "YYYY-MM-DD HH:MM:SS" (UTC, no zone
113
+ // marker), which V8 parses as LOCAL time — a timezone-sized skew that
114
+ // could fail the gate open. Normalize the space form to UTC.
115
+ const iso = s.includes("T") ? s : s.replace(" ", "T") + "Z";
116
+ const ms = Date.parse(iso);
117
+ return Number.isFinite(ms) ? ms : NaN;
118
+ }
119
+ // Strictly newer: an mtime exactly equal to created_at is stale — the
120
+ // capture cannot predate the task it evidences, but "same instant" proves
121
+ // nothing. (Only "newer than task creation" is proven: captures from an
122
+ // earlier attempt on the same task pass this check too.)
123
+ function isFreshFile(p, createdAtMs) {
124
+ try {
125
+ const st = fs.statSync(p);
126
+ return st.isFile() && st.mtimeMs > createdAtMs;
127
+ } catch { return false; }
98
128
  }
99
129
  let auditPath = null;
100
130
  if (project && project.deploy_slug) {
101
131
  auditPath = path.join(os.homedir(), "workspace", "ts-spaces",
102
132
  project.deploy_slug, "audits", auditDir);
103
133
  }
104
- const hasCaptures = !!(auditPath &&
105
- isFile(path.join(auditPath, "screenshot.png")) &&
106
- isFile(path.join(auditPath, "screenshot-mobile.png")));
134
+ const createdAtMs = parseCreatedAtMs(task.created_at);
135
+ const hasCaptures = !!(auditPath && Number.isFinite(createdAtMs) &&
136
+ isFreshFile(path.join(auditPath, "screenshot.png"), createdAtMs) &&
137
+ isFreshFile(path.join(auditPath, "screenshot-mobile.png"), createdAtMs));
107
138
 
108
139
  let events = [];
109
140
  try {
package/lib/crew-api.js CHANGED
@@ -1675,6 +1675,13 @@ commands["resolve-publish-unknown"] = (db, args, ctx) => {
1675
1675
  if (evidenceDirs.length === 0) {
1676
1676
  return { resolved: false, reason: "no audit build completed inside the publish window — still unknown" };
1677
1677
  }
1678
+ // (N3) Ambiguity guard: multiple builds in the recovery window have the
1679
+ // same attribution ambiguity as the workflow's immediate verdict — no
1680
+ // single dir can be attributed to this attempt, so resolution stays
1681
+ // unknown. Only when exactly one dir remains is evidenceDirs[0] safe.
1682
+ if (evidenceDirs.length > 1) {
1683
+ return { resolved: false, reason: `audit-dir ambiguity: ${evidenceDirs.length} audit dirs inside the publish window — still unknown` };
1684
+ }
1678
1685
  const evidenceDir = evidenceDirs[0];
1679
1686
 
1680
1687
  // Corroboration, observation only: the audit harness's own verdict.
@@ -1700,7 +1707,7 @@ commands["resolve-publish-unknown"] = (db, args, ctx) => {
1700
1707
  db.prepare(
1701
1708
  "INSERT INTO events (id, type, task_id, identity, message, timestamp) VALUES (?, 'note', ?, NULL, ?, ?)"
1702
1709
  ).run(uuid(), taskId,
1703
- `Parked: publish: verification-requested ${latest.commit} (build agent_id unobserved) — recovered from publish-unknown via resolve-publish-unknown; durable evidence ${evidenceDir}. Parent: run docs/publish-verification.md.`,
1710
+ `Parked: publish: verification-requested ${latest.commit} (build agent_id unobserved) — recovered from publish-unknown via resolve-publish-unknown; durable evidence ${evidenceDir}. waiting on the manual read-back in docs/publish-verification.md.`,
1704
1711
  tsLater);
1705
1712
  db.prepare("UPDATE tasks SET state = 'parked', updated_at = ? WHERE id = ?").run(tsLater, taskId);
1706
1713
  db.exec("COMMIT");
@@ -73,9 +73,12 @@ cmd_init() {
73
73
  echo "CURRENT: $cur"
74
74
  }
75
75
 
76
- # ── workflow syntax validation ──────────────────────────────────────
77
- # Every workflow script must parse before a release can install.
76
+ # ── workflow validation ─────────────────────────────────────────────
77
+ # Every workflow script must (a) parse under the workflow-loader emulation
78
+ # and (b) fit the launch budget before a release can install.
79
+ # Exit codes: 10 = parse check failed, 20 = size check failed (0 = pass).
78
80
  _validate_workflows() {
81
+ # --- Parse gate (restored 2026-09-18, finding B2) ---
79
82
  # The workflow runtime accepts top-level `export`, `return`, and `await`
80
83
  # (it wraps scripts in an async function), so neither `node --check` on the
81
84
  # raw .js (vacuous for ESM — exits 0 even on blatant syntax errors) nor a
@@ -83,8 +86,23 @@ _validate_workflows() {
83
86
  # Emulate the runtime instead: strip `export`, wrap the script in an async
84
87
  # function, then node --check the result. Tokenizer errors (e.g. an
85
88
  # unterminated string literal) still fail under the wrap.
89
+ #
90
+ # HONESTY NOTE: this is a provisional emulation, not a verification. The
91
+ # platform's actual loading mechanism was never verified (probe blocked,
92
+ # 2026-09-18), so this check cannot prove how the platform loads the
93
+ # workflows. What it DOES prove: the file is syntactically well-formed JS
94
+ # when read the way the runtime claims to read it (strip export, run as one
95
+ # async function body). What it does NOT prove: that the platform actually
96
+ # loads it that way, or that the code runs correctly. The export-strip +
97
+ # async-wrap transform exists precisely because raw `node --check` on ESM
98
+ # `.js` is unreliable — it silently exits 0 on some malformed input — so
99
+ # the transform is what makes the check actually parse.
100
+ # A failure here is a signal to INSPECT THE CODE, never to delete the
101
+ # check: this gate was once removed on a misdiagnosis ("the check
102
+ # false-positives"), and the workflows it rejected turned out to be
103
+ # genuinely broken (2026-09-18, Gate 1 H2 — a true positive).
86
104
  local dir="$1"
87
- local f tmp base
105
+ local f tmp base size
88
106
  tmp="$(mktemp -d)"
89
107
  for f in "$dir"/workflows/*.js; do
90
108
  [ -f "$f" ] || continue
@@ -96,13 +114,28 @@ _validate_workflows() {
96
114
  } > "$tmp/$base.js"
97
115
  if ! node --check "$tmp/$base.js" 2>"$tmp/$base.err"; then
98
116
  cat "$tmp/$base.err" >&2
99
- echo "VALIDATION FAILED: $f does not parse" >&2
117
+ echo "VALIDATION FAILED [parse]: $f does not parse under the workflow-loader emulation" >&2
100
118
  rm -rf "$tmp"
101
- return 1
119
+ return 10
102
120
  fi
103
121
  done
104
122
  rm -rf "$tmp"
105
- echo "VALIDATED: workflow scripts parse"
123
+
124
+ # --- Size gate (H5, 2026-09-18) ---
125
+ # The platform reads the full workflow file on launch; the platform hard
126
+ # maximum is 262144 bytes (256 KiB) and 245760 bytes (240 KiB) is the
127
+ # project's safety target. The release fails closed if any launchable
128
+ # workflow hits the safety target.
129
+ for f in "$dir"/workflows/*.js; do
130
+ [ -f "$f" ] || continue
131
+ size="$(wc -c < "$f")"
132
+ if [ "$size" -ge 245760 ]; then
133
+ echo "VALIDATION FAILED [size]: $f is $size bytes (>= 245760 byte launch budget)" >&2
134
+ return 20
135
+ fi
136
+ done
137
+
138
+ echo "VALIDATED: workflow scripts parse and fit the launch budget"
106
139
  }
107
140
 
108
141
  # ── workflow doc sync ───────────────────────────────────────────────
@@ -255,10 +288,19 @@ cmd_deploy() {
255
288
  rm -rf "$staging_dir"
256
289
  die "release $hash rejected: registry build failed"
257
290
  fi
258
- # Gate: refuse to install a release whose workflow scripts don't parse.
259
- if ! _validate_workflows "$staging_dir"; then
291
+ # Gate: refuse to install a release whose workflow scripts fail the
292
+ # parse or size checks. _validate_workflows returns 10 on a parse
293
+ # failure, 20 on a size-budget failure; the rejection names the check
294
+ # that failed so the message is never misleading.
295
+ _vw_status=0
296
+ _validate_workflows "$staging_dir" || _vw_status=$?
297
+ if [ "$_vw_status" -ne 0 ]; then
260
298
  rm -rf "$staging_dir"
261
- die "release $hash rejected: workflow syntax validation failed"
299
+ case "$_vw_status" in
300
+ 10) die "release $hash rejected: workflow parse check failed" ;;
301
+ 20) die "release $hash rejected: workflow size-budget check failed" ;;
302
+ *) die "release $hash rejected: workflow validation failed (unexpected exit $_vw_status)" ;;
303
+ esac
262
304
  fi
263
305
  # Atomic rename into place
264
306
  mv "$staging_dir" "$release_dir"
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "muse-crew",
3
- "version": "0.13.1",
3
+ "version": "0.13.3",
4
4
  "description": "Opinionated orchestration for Muse \u2014 workflows, identities, and tooling for autonomous software development.",
5
5
  "license": "UNLICENSED",
6
6
  "private": false,
@@ -29,4 +29,4 @@
29
29
  "dependencies": {
30
30
  "playwright-core": "1.63.0"
31
31
  }
32
- }
32
+ }
package/seed/crons.json CHANGED
@@ -5,7 +5,6 @@
5
5
  "enabled": true,
6
6
  "id": "crew-poll",
7
7
  "mode": "task",
8
- "owner": "cli:{instanceId}",
9
8
  "schedule": {
10
9
  "every": "15m",
11
10
  "kind": "interval"
@@ -18,7 +17,6 @@
18
17
  "enabled": true,
19
18
  "id": "crew-update-watch",
20
19
  "mode": "task",
21
- "owner": "cli:{instanceId}",
22
20
  "schedule": {
23
21
  "every": "24h",
24
22
  "kind": "interval"
@@ -4,7 +4,9 @@ Executable Muse workflow scripts (JavaScript). These are what the workflow runti
4
4
 
5
5
  - `crew-dispatch.js` — reads the board, recommends eligible tasks, returns structured launch records (the launched workflow self-claims; the dispatcher never writes claims). Skips tasks on quiesced projects and on projects with no repo_path configured.
6
6
 
7
- All four workflow scripts share byte-identical transport helpers (`workRetryKey`, `buildTransportRetryTrailer`, `describeWorkAgentFailure`, `workerMissingArtifactTools` — pinned by `tests/closeout.test.js`). The transport-retry loop treats a worker report naming the missing artifact tool namespace (bug 3472bf36, a per-launch platform flake) as a retryable attempt with a fresh launch rather than accepting a useless report.
7
+ All four workflow scripts share byte-identical transport helpers (`workRetryKey`, `buildTransportRetryTrailer`, `describeWorkAgentFailure` — pinned by `tests/closeout.test.js`). The transport-retry loop treats a worker report naming the missing artifact tool namespace (bug 3472bf36, a per-launch platform flake) as a retryable attempt with a fresh launch rather than accepting a useless report.
8
+
9
+ **Growth budget (2026-09-18):** the platform reads the full workflow file on launch; the platform hard maximum is 262144 bytes (256 KiB) and the project's safety target is 245760 bytes (240 KiB). Any change adding net bytes to a workflow already over 240 KiB must subtract ≥ the addition in the same change — see `tests/workflow-size.test.js` and the release-time gate in `lib/crew-release.sh` `_validate_workflows`. Decision-history narration lives in `docs/decisions/`; workflows keep only the relied-upon invariant inline. See `docs/decisions/workflow-core.md#h5-probe-blocked` for why helper extraction was refused and decision history was relocated instead. Target-vs-cap (2026-09-18): the release gate intentionally enforces the 245760-byte target, not just the 262144-byte cap — safe because target ≤ cap, stricter by design.
8
10
  - `crew-init.js` — sets up a new crew instance: Gate 0 validates crewHome is workspace-contained and a valid git repo, and (in dashboard mode) validates dashboardRepoPath is workspace-contained and an existing git repo (never auto-created); bootstraps the release system; scaffolds orchestration folders; registers the first project with repo_path set to the validated dashboard repo — skipped in CLI-only mode (no dashboardSlug/dashboardRepoPath), where projects are created later via crew-api.js; creates/converges cron jobs from the declarative manifest (scheduler identity is dashboard-independent: first init uses the crew-home basename, re-runs keep the instanceId recorded in `.cron-registry.json`; owner `cli:<instanceId>` — deleting a dashboard artifact never stops the crew). Idempotent. The crew owns its state via the Crew API (lib/crew-api.js); the dashboard is an optional client and is never a dependency.
9
11
  - `standard.js` — default task workflow: Triage → Capture → Map → Build → Review → Integrate → Publish → QA
10
12
  - `bugfix.js` — adds Capture after Triage, then Reproduce