muse-crew 0.16.0 → 0.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/API.md CHANGED
@@ -133,6 +133,20 @@ Records a builder's refusal report (0.14.6). Mechanically requires the EXACT ver
133
133
 
134
134
  Returns `{ "ok": true, "publish_attempt" }`. The ack scan consumes refusal reports (terminal `publish: publish-refused`); positive acknowledgements come only from the on-disk version receipt.
135
135
 
136
+ ### `record-worker-run`
137
+
138
+ Records a worker-layer execution in the `worker_runs` ledger (Piece 1, 2026-09-26). Identity is minted SQLite-side: a `start` call (no `id`) inserts a `running` row and returns its integer id; a `finish` call (`id` + terminal `status`) completes it. Terminal rows are immutable (R-2) — a finish against a non-running row is a silent no-op returning `{ "ok": true, "noop": true }`. `executor` is fixed at `'worker'` by the schema default.
139
+
140
+ | Field | Type | Required | Notes |
141
+ |-------|------|----------|-------|
142
+ | `id` | integer | on finish | SQLite-minted on start |
143
+ | `task_id` | uuid or null | no | Optional correlation; null for dispatch runs |
144
+ | `phase` | string | yes | e.g. `dispatch` |
145
+ | `status` | `running` \| `completed` \| `failed` | yes | Start must pass `running`; finish must pass a terminal status |
146
+ | `error` | string or null | no | Failure detail on `failed` |
147
+
148
+ Bad transitions (start with an `id`, finish without one, finish with `running`, missing `phase`) throw usage errors (exit 2). Returns `{ "ok": true, "id" }` on start and `{ "ok": true, "finished": true }` on finish. Orphan note: a dispatcher killed between start and finish leaves a `running` row — readers treat a `running` row with no live dispatcher as stale, never as success.
149
+
136
150
  ### `record-version-ack`
137
151
 
138
152
  Stamps a version acknowledgement (0.14.6). The version must name an issued attempt for the task/commit.
package/lib/AGENTS.md CHANGED
@@ -10,6 +10,9 @@ Shipped library: ESM JavaScript CLIs and import-safe modules, shell scripts for
10
10
 
11
11
  - `build-registry.js` — deterministic extractor that generates `workflows/registry.json` (workflow step registry) from the workflow files' `meta` blocks at release time; invoked by `crew-release.sh` deploy
12
12
  - `crew-api.js` — the crew-owned task-service API (dependency inversion, 2026-09-11): a zero-dependency Node CLI implementing the API.md contract against `$CREW_HOME/crew-state.db` (schema in `schema.sql`). Workflows call it through their agents' shell; the dashboard delegates to it. All state-machine invariants live as CHECK constraints in the schema, never in client prose. Includes the `record-phase` composite (session + event in one transaction; 2026-09-21 optionally carries a structured Review verdict payload — validated at the API boundary, written to the `verdicts` table in the same transaction, session-keyed first-write-wins idempotency — plus the `get-verdicts` read surface) and a one-time `migrate` import from a dashboard app.db. Active-release resolution is split in two (room #15): `resolveActiveReleaseName` (symlink-only — writers like the initial-provenance stamp record the active release without proving they are it) and `resolveActiveRelease` (symlink + self-path cross-check — verifiers like `scan-ack-pending` refuse to stamp claims when the running code isn't the active release's own). Provenance is per-project (2026-09-18, room #15 blocker 8): nullable `provenance_*` columns on the projects row, a single `stampProvenance()` writer, `set-provenance`/`get-provenance` require `project_id` (no silent global fallback), and a watermarked openDb backfill that attributes the legacy `config.provenance.*` triple to exactly-one ancestor match — never fabricated, otherwise deferred.
13
+ - `crew-dispatch-worker.js` — worker-layer port of the sandboxed dispatch dispatcher (Piece 1, 2026-09-26): full-privilege Node, reads the board through crew-api.js via execFile (argv only, no shell), reproduces the eligibility / retry / reservation / playtest / tick-release / dispatch-decision logic, and tags every claim `executor: "worker"`. In Piece 1 it runs `--read-only` as the tick's shadow BEFORE the authoritative sandboxed dispatcher — both read the same board state; the shadow never mutates it, never launches, never writes the tick-release or dispatch-decision logs. Appends its own shadow evidence line to `$CREW_HOME/.dispatch-shadow.jsonl` (machine-written, no agent transcription). Claims are recommendations only: `executor` names the layer that computed the recommendation, not the layer that launched it — in Piece 1 no `executor: "worker"` claim is ever launched.
14
+ - `compare-dispatch-shadow.js` — mechanical shadow verdict (Piece 1, 2026-09-26): pairs each `.dispatch-shadow.jsonl` line with the authoritative `.dispatch-decisions.jsonl` line for the same tick (`decision.tick_seq === shadow.tick_seq_before + 1` — the shadow runs first, the sandbox dispatcher then writes its tick-release line and its decision) and compares the sorted `task_id|workflow|step` claim triples, ignoring `executor`. Emits `SHADOW_MATCH` / `SHADOW_DIVERGE` (names the differing triples) / `SHADOW_UNPAIRED` (decision missing — the authoritative dispatcher died) and appends the verdict to `$CREW_HOME/.dispatch-shadow-verdicts.jsonl`; idempotent (already-verdict shadows are skipped). Exit 0 on a verdict — divergence is evidence for the cutover review, not a tick failure; exit 2 on usage/IO errors.
15
+ - `spawn-boundary.js` — fail-closed spawn-bounds validator (Piece 1, 2026-09-26): import-safe ESM, `validateSpawnBounds()` requires prompt/schema co-location, cwd inside the workdir, rejection of cwd inside crew state/observer paths or crewHome, and rejection of sensitive env-key patterns; throws `BOUNDS_REJECTED`. Bare module execution exits 0 (release entry-gate contract). The production spawn/request wrapper is Piece-3 work; this is only the validator.
13
16
  - `schema.sql` — the crew-owned state schema: projects, tasks, poll_state, config, agent_sessions, events. Vocabularies enforced by CHECK constraints; `rejected` is a valid event type (the 2026-09-11 crash was a stored session whose event was rejected). Column names match the historical dashboard tables for a verbatim migration.
14
17
  - `crew-release.sh` — immutable release manager: deploy, rollback, prune
15
18
  - `merge-lock.sh` — serialized merge lock for concurrent agents: time-based holder lease (bug 2fc8f52f — an unexpired lease is held regardless of process liveness; only an expired lease may be broken). Requires both `CREW_REPO` and `CREW_HOME` (fail closed: BLOCKED, exit 2 when either is unset). Lock file is key=value: task_id, opaque holder identity (never a PID), acquired_at epoch, lease_seconds (default 600, override via MERGE_LOCK_LEASE_SECONDS). acquire/refresh/release/status/force-release; holder-only refresh and release; every op appends to $CREW_HOME/.merge-lock.log. Stale-lease breaks serialize on a sidecar `$LOCK_FILE.flock` with an in-critical-section lease re-read (R-B1, 2026-09-21); refresh rewrites via temp-file + atomic rename so readers never see a torn file, and takes the same sidecar flock with an identity-only re-check inside the critical section — a reclaim always changes `task_id`, so identity alone closes the clobber (no expiry check on refresh: a long build that outran the lease legitimately revives its lock); release takes the same flock with the identity re-check inside (review pass 2, 2026-09-21) so a release can never `rm` a reclaimer's fresh lock.
@@ -0,0 +1,150 @@
1
+ #!/usr/bin/env node
2
+ // lib/compare-dispatch-shadow.js — shadow verdict for the Piece-1 sandbox exit.
3
+ //
4
+ // Pairs the worker dispatcher's read-only shadow evidence
5
+ // ($CREW_HOME/.dispatch-shadow.jsonl) with the authoritative sandbox
6
+ // dispatcher's decision ($CREW_HOME/.dispatch-decisions.jsonl) and emits a
7
+ // mechanical verdict. No agent transcribes claims; both sides are
8
+ // machine-written JSONL.
9
+ //
10
+ // Pairing: shadow.tick_seq_before + 1 === decision.tick_seq. The shadow runs
11
+ // first in the tick (before the sandbox dispatcher writes its tick-release
12
+ // line and its decision), so the authoritative decision for the same board
13
+ // state is the first decision line with tick_seq exactly one past the
14
+ // shadow's observed count.
15
+ //
16
+ // Match predicate: the sorted claim triples task_id|workflow|step are
17
+ // identical. The executor tag is excluded by construction (both sides are
18
+ // normalized to the triple before comparing). Reservation outcomes are NOT
19
+ // compared — the shadow skips reservations; a sandbox reservation failure
20
+ // surfaces as a DIVERGE with the differing triples named, which is honest
21
+ // evidence, not a false alarm.
22
+ //
23
+ // Usage: node compare-dispatch-shadow.js --crew-home <path>
24
+ // Prints one of SHADOW_MATCH / SHADOW_DIVERGE / SHADOW_UNPAIRED /
25
+ // SHADOW_NO_EVIDENCE and appends the verdict to
26
+ // $CREW_HOME/.dispatch-shadow-verdicts.jsonl. Exit 0 on a verdict (even a
27
+ // diverge — divergence is evidence, not a failure); exit 2 on usage or IO
28
+ // errors.
29
+
30
+ import { readFileSync, appendFileSync } from "node:fs";
31
+ import { join } from "node:path";
32
+
33
+ function usage() {
34
+ console.log("Usage: node compare-dispatch-shadow.js --crew-home <path>");
35
+ console.log("");
36
+ console.log("Pairs the latest shadow evidence line with the authoritative");
37
+ console.log("dispatch decision for the same tick and emits a verdict.");
38
+ process.exit(0);
39
+ }
40
+
41
+ const argv = process.argv.slice(2);
42
+ if (argv.includes("--help") || argv.includes("-h")) usage();
43
+ let crewHome = null;
44
+ for (let i = 0; i < argv.length; i++) {
45
+ if (argv[i] === "--crew-home" && i + 1 < argv.length) crewHome = argv[++i];
46
+ }
47
+ if (!crewHome) {
48
+ console.error("compare-dispatch-shadow: --crew-home <path> is required");
49
+ process.exit(2);
50
+ }
51
+
52
+ function readJsonl(file) {
53
+ let lines = [];
54
+ try {
55
+ lines = readFileSync(file, "utf8").split("\n");
56
+ } catch (e) {
57
+ return [];
58
+ }
59
+ const out = [];
60
+ for (const l of lines) {
61
+ if (l.trim() === "") continue;
62
+ try { out.push(JSON.parse(l)); } catch (e) { /* skip corrupt lines loudly below */ }
63
+ }
64
+ return out;
65
+ }
66
+
67
+ function normalizeClaims(list) {
68
+ return (list || [])
69
+ .map((c) => `${c.task_id}|${c.workflow}|${c.step}`)
70
+ .sort();
71
+ }
72
+
73
+ function emit(verdict) {
74
+ const file = join(crewHome, ".dispatch-shadow-verdicts.jsonl");
75
+ let seq = 0;
76
+ try {
77
+ const lines = readFileSync(file, "utf8").split("\n");
78
+ for (const l of lines) { if (l.trim() !== "") seq++; }
79
+ } catch (e) { /* missing file */ }
80
+ const line = JSON.stringify({ seq: seq + 1, ...verdict });
81
+ try {
82
+ appendFileSync(file, line + "\n");
83
+ } catch (e) {
84
+ console.error("compare-dispatch-shadow: verdict write failed: " + e.message);
85
+ process.exit(2);
86
+ }
87
+ return { seq: seq + 1, ...verdict };
88
+ }
89
+
90
+ const shadows = readJsonl(join(crewHome, ".dispatch-shadow.jsonl")).filter((l) => l.mode === "shadow");
91
+ if (shadows.length === 0) {
92
+ console.log("SHADOW_NO_EVIDENCE no shadow lines yet");
93
+ process.exit(0);
94
+ }
95
+ const shadow = shadows[shadows.length - 1];
96
+
97
+ const decisions = readJsonl(join(crewHome, ".dispatch-decisions.jsonl"));
98
+ const verdicts = readJsonl(join(crewHome, ".dispatch-shadow-verdicts.jsonl"));
99
+ const done = new Set(verdicts.map((v) => v.shadow_seq));
100
+ const pending = shadows.filter((s) => !done.has(s.seq));
101
+
102
+ if (pending.length === 0) {
103
+ console.log(`SHADOW_CURRENT all ${shadows.length} shadow line(s) already have verdicts`);
104
+ process.exit(0);
105
+ }
106
+
107
+ let matched = 0, diverged = 0, unpaired = 0;
108
+ for (const shadow of pending) {
109
+ const wantTickSeq = (shadow.tick_seq_before || 0) + 1;
110
+ const decision = decisions.find((d) => d.tick_seq === wantTickSeq);
111
+
112
+ if (!decision) {
113
+ emit({ verdict: "unpaired", shadow_seq: shadow.seq, want_tick_seq: wantTickSeq });
114
+ console.log(`SHADOW_UNPAIRED shadow_seq=${shadow.seq} want_tick_seq=${wantTickSeq} (authoritative decision not yet written)`);
115
+ unpaired++;
116
+ continue;
117
+ }
118
+
119
+ const shadowClaims = normalizeClaims(shadow.claims);
120
+ const decisionClaims = normalizeClaims(decision.launched);
121
+ const isMatch =
122
+ shadowClaims.length === decisionClaims.length &&
123
+ shadowClaims.every((c, i) => c === decisionClaims[i]);
124
+
125
+ const onlyShadow = shadowClaims.filter((c) => !decisionClaims.includes(c));
126
+ const onlyDecision = decisionClaims.filter((c) => !shadowClaims.includes(c));
127
+
128
+ emit({
129
+ verdict: isMatch ? "match" : "diverge",
130
+ shadow_seq: shadow.seq,
131
+ decision_seq: decision.seq,
132
+ tick_seq: decision.tick_seq,
133
+ shadow_claims: shadowClaims,
134
+ decision_claims: decisionClaims,
135
+ only_shadow: onlyShadow,
136
+ only_decision: onlyDecision,
137
+ shadow_partial: !!shadow.partial,
138
+ decision_partial: !!decision.partial,
139
+ });
140
+
141
+ if (isMatch) {
142
+ console.log(`SHADOW_MATCH shadow_seq=${shadow.seq} decision_seq=${decision.seq} claims=${shadowClaims.length}`);
143
+ matched++;
144
+ } else {
145
+ console.log(`SHADOW_DIVERGE shadow_seq=${shadow.seq} decision_seq=${decision.seq} only_shadow=[${onlyShadow.join(",")}] only_decision=[${onlyDecision.join(",")}]`);
146
+ diverged++;
147
+ }
148
+ }
149
+ console.log(`SHADOW_SUMMARY matched=${matched} diverged=${diverged} unpaired=${unpaired}`);
150
+ process.exit(0);
package/lib/crew-api.js CHANGED
@@ -1314,6 +1314,37 @@ commands["record-run-end"] = (db, args) => {
1314
1314
  return { ok: true, run_id: args.run_id, status };
1315
1315
  };
1316
1316
 
1317
+ // record-worker-run: upsert a worker-layer orchestration run (Piece 1,
1318
+ // 2026-09-26). The worker layer's counterpart to record-run-start/end.
1319
+ // Args: { id?, phase, task_id?, status, error? }
1320
+ // - status "running" without id: INSERTs a row (id assigned by SQLite —
1321
+ // worker JS never mints identity, see G3), returns { ok, id }.
1322
+ // - status "completed"|"failed" with id: UPDATEs that row's status,
1323
+ // error, and ended_at (only from 'running'; terminal rows are
1324
+ // immutable, same R-2 rule as record-run-end), returns { ok, id }.
1325
+ // - Anything else is a usage error. task_id is optional (the dispatcher
1326
+ // run is not tied to one task).
1327
+ commands["record-worker-run"] = (db, args) => {
1328
+ if (!args.phase) throw usageError("phase is required.");
1329
+ const status = args.status || "running";
1330
+ if (!["running", "completed", "failed"].includes(status)) {
1331
+ throw usageError("status must be running, completed, or failed.");
1332
+ }
1333
+ if (status === "running") {
1334
+ if (args.id !== undefined && args.id !== null) throw usageError("id must not be set when starting a run.");
1335
+ const row = db.prepare(
1336
+ `INSERT INTO worker_runs (task_id, phase, status) VALUES (?, ?, 'running') RETURNING id`
1337
+ ).get(args.task_id || null, args.phase);
1338
+ return { ok: true, id: row.id, status: "running" };
1339
+ }
1340
+ if (args.id === undefined || args.id === null) throw usageError("id is required to finish a run.");
1341
+ db.prepare(
1342
+ `UPDATE worker_runs SET status = ?, error = ?, ended_at = strftime('%Y-%m-%dT%H:%M:%fZ', 'now')
1343
+ WHERE id = ? AND status = 'running'`
1344
+ ).run(status, args.error || null, args.id);
1345
+ return { ok: true, id: args.id, status };
1346
+ };
1347
+
1317
1348
  commands["get-run-timeline"] = (db, args) => {
1318
1349
  if (!args.run_id && !args.task_id) throw usageError("run_id or task_id is required.");
1319
1350
  let runs;