muse-crew 0.14.12 → 0.15.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/API.md CHANGED
@@ -489,6 +489,8 @@ Register a new project.
489
489
  | `quiesced` | boolean | no | Start paused; defaults to false |
490
490
  | `visual_protocol` | boolean or null | no | Tri-state: `null` = inherit crew default (off), `true` = enable for this project, `false` = explicit off. Defaults to null. |
491
491
  | `environment_type` | `artifact` · `terminal` · null | no | User-facing surface for experiential QA routing: `artifact` = a rendered web UI (Hazel drives it with the see-act browser loop), `terminal` = a CLI (Hazel drives it herself, keeping transcripts). `null` = unclassified: no experiential QA. `deploy_type` names the deployment target, but `deploy_type: "artifact"` remains a legacy artifact-surface signal so pre-field projects keep today's experiential QA (the migration does not backfill the column); on conflicting config artifact wins. Omitted means auto-classify from `repo_path` (see surface classification below); an explicit value, including explicit `null`, always wins. |
492
+ | `versioned_build` | boolean | no | Opt-in to the QA staleness gate (issue #3): when true, Integrate bumps `BUILD_NUMBER` in `version_file` after each merge (separate commit, pushed), and QA verifies the tested bundle's footer `build <n>` is at/past the recorded version before testing. Defaults to false (unversioned); the migration does not backfill. |
493
+ | `version_file` | string or null | no | Repo-relative path of the TypeScript file exporting `BUILD_NUMBER` (e.g. `client/src/buildNumber.ts`). Only meaningful when `versioned_build` is true; defaults to `client/src/buildNumber.ts` when omitted. |
492
494
 
493
495
  **Surface classification (repo-derived defaults):** when `repo_path` is provided and `environment_type` is omitted, `create-project` mechanically classifies the repo at registration via `lib/classify-surface.js` — `space.json` present → `artifact`; else a non-empty `package.json` `bin` → `terminal`; otherwise unclassified (`null`). The classifier never throws: bad paths and malformed files classify as unclassified. Classification only fills omitted fields — an explicitly supplied `environment_type` (including explicit `null`) always wins, and deploy defaults follow the *effective* surface, so an explicit `"terminal"` never triggers artifact deploy defaults.
494
496
 
@@ -6,6 +6,12 @@ history lives here. One canonical section per `<a id>` anchor — a critic
6
6
  re-verified that every `docs/decisions/*.md#anchor` reference in
7
7
  `workflows/*.js` resolves (45 unique references).
8
8
 
9
+ ## dispatch-decision-log.md — Decision history: dispatch decision log (hoverboat)
10
+
11
+ - `#checksum-vs-hoverboat` — The ferry-corruption problem and the two candidates
12
+ - `#why-hoverboat-won` — Availability, trust, subtraction, primary-source rationale
13
+ - `#what-was-built` — Declaration shape, skip counters, observer ingestion, retirements
14
+
9
15
  ## publish-path.md — Decision history: publish path
10
16
 
11
17
  - `#fire-and-forget-trigger` — Fire-and-forget trigger + workflow-owned observation
@@ -0,0 +1,69 @@
1
+ # Dispatch decision log (hoverboat) — 2026-09-24
2
+
3
+ <a id="checksum-vs-hoverboat"></a>
4
+ ## The problem
5
+
6
+ The Gate 2 observer attributed dispatch decisions by ferrying platform
7
+ `scheduler.job_runs` rows through an LLM paste (tick worker: `muse.db`
8
+ result → heredoc → `record_platform.sh --dispatch` → `dispatch_sample`
9
+ rows → `attribute_dispatch.sh` → `dispatch_decision` rows). The paste
10
+ mutated plausible values: a 2026-09-24 consistency scan found 275
11
+ same-`run_id` contradictory `dispatch_sample` rows and 65 derived
12
+ `dispatch_decision` rows based on corrupt data (5.7% of samples
13
+ quarantined; 63.6% of ticks touched). The older `quarantine_suspects.sh`
14
+ handled near-neighbor UUID mutations, not plausible value mutation.
15
+
16
+ ## The two candidates
17
+
18
+ **Checksum ferry** (containment): keep the ferry; the worker checksums the
19
+ raw tool-result JSON before the paste and `record_platform.sh` verifies
20
+ after. Corrupt pastes become loud absences instead of silent corruption.
21
+
22
+ **Hoverboat** (local declaration): the dispatcher writes its decision to
23
+ `$CREW_HOME/.dispatch-decisions.jsonl` in its own bytes (deterministic JS,
24
+ same courier shape as the proven §4b tick-release writer); the observer
25
+ reads the file directly and verifies launch declarations mechanically
26
+ against crew DB effects (`platform_run_tasks`, `dispatch_reservations`).
27
+
28
+ <a id="why-hoverboat-won"></a>
29
+ ## Why hoverboat won
30
+
31
+ 1. **Availability.** Corruption touched ~63.6% of ticks. Checksum converts
32
+ corruption into absence, but the entry demo needs ten consecutive ticks
33
+ each recording a decision — absence fails the demo as surely as
34
+ corruption. Hoverboat's channel has no LLM re-emission at all.
35
+ 2. **Trust.** The checksum would be computed by the same worker agent that
36
+ corrupts the paste — asking the corrupting agent to honestly report its
37
+ own corruption. Hoverboat's declaration is composed by deterministic
38
+ workflow JS and checked against independent DB effects.
39
+ 3. **Subtraction.** Checksum keeps the ferry, the rolling 25h re-record
40
+ window, and the quarantine machinery permanently busy. Hoverboat deletes
41
+ the sample→decision attribution entirely; the platform ferry remains
42
+ only as platform-health telemetry.
43
+ 4. **The crew already knows.** Having the platform tell the observer what
44
+ the crew decided, through an LLM paste, when the crew can write it
45
+ directly, is the ferry. The declaration is the primary source; the DB
46
+ effects are the independent verification.
47
+
48
+ Rejected alternative considered: a deterministic Crew API action for the
49
+ write. It would not remove the agent from the path (workflows reach the
50
+ CLI only through `agent()` shell calls), so it adds API surface for no
51
+ integrity gain. The embedded courier is the proven seam.
52
+
53
+ <a id="what-was-built"></a>
54
+ ## What was built
55
+
56
+ - `workflows/crew-dispatch.js` §4c: per-tick declaration
57
+ `{seq, tick_seq, release, decision, launched[], completed[], parked[],
58
+ board{seen,eligible}, skipped{reason:count sparse}, partial}`.
59
+ `tick_seq` couples each line to the `.tick-releases.jsonl` line for the
60
+ same poll — a tick-release line with no decision line means the tick died
61
+ after poll-ack, visibly.
62
+ - Per-reason ineligibility counters (`countSkipped`) at every skip site in
63
+ the eligibility loop + the simultaneity-limit skip, so stand-downs are
64
+ auditable without re-deriving eligibility.
65
+ - Observer: `ingest_decisions.sh` replaces `attribute_dispatch.sh`;
66
+ `snapshot.sh` dispatch section rewritten around decision rows.
67
+ - Retired: sample→decision attribution, `state/attributed.json` watermark.
68
+ Kept: `record_platform.sh --dispatch` as platform-health telemetry,
69
+ `quarantine_suspects.sh` for the health samples' ID mutations.
package/lib/crew-api.js CHANGED
@@ -73,6 +73,21 @@ function validateEnvironmentType(value) {
73
73
  if (value === "artifact" || value === "terminal") return value;
74
74
  throw usageError("environment_type must be 'artifact', 'terminal', or null.");
75
75
  }
76
+ // versioned_build validation: boolean. undefined means "not provided"
77
+ // (create: defaults to false; update: leaves the field alone). Coerces
78
+ // truthy/falsy to 1/0 for the INTEGER column.
79
+ function validateVersionedBuild(value) {
80
+ if (value === undefined) return undefined;
81
+ return value ? 1 : 0;
82
+ }
83
+ // version_file validation: string or null. undefined means "not provided"
84
+ // (create: defaults to null; update: leaves the field alone). Empty string
85
+ // is rejected — use null to clear.
86
+ function validateVersionFile(value) {
87
+ if (value === undefined || value === null) return null;
88
+ if (typeof value === "string" && value.length > 0) return value;
89
+ throw usageError("version_file must be a non-empty string or null.");
90
+ }
76
91
 
77
92
  // already_merged_sha validation (room #16 blocker 11, 2026-09-18): the
78
93
  // workflow-verified already-merged sha as structured session control state.
@@ -315,6 +330,24 @@ function openDb(crewHome) {
315
330
  if (!/duplicate column name/i.test(e.message)) throw e;
316
331
  }
317
332
  }
333
+ // Migration: add versioned_build + version_file columns if missing (QA
334
+ // staleness gate, emojimanegg1/muse-crew#3, 2026-09-24). Same idempotent
335
+ // PRAGMA-check pattern; versioned_build defaults 0 (unversioned by
336
+ // design, no backfill), version_file stays NULL.
337
+ if (!projectCols.some((c) => c.name === "versioned_build")) {
338
+ try {
339
+ db.exec("ALTER TABLE projects ADD COLUMN versioned_build INTEGER NOT NULL DEFAULT 0 CHECK (versioned_build IN (0, 1))");
340
+ } catch (e) {
341
+ if (!/duplicate column name/i.test(e.message)) throw e;
342
+ }
343
+ }
344
+ if (!projectCols.some((c) => c.name === "version_file")) {
345
+ try {
346
+ db.exec("ALTER TABLE projects ADD COLUMN version_file TEXT");
347
+ } catch (e) {
348
+ if (!/duplicate column name/i.test(e.message)) throw e;
349
+ }
350
+ }
318
351
  // Migration: per-project provenance columns (room #15 blocker 8,
319
352
  // 2026-09-18). Same idempotent PRAGMA-check pattern; ADD COLUMN without
320
353
  // a default leaves existing rows NULL, which is "no provenance yet" by
@@ -446,6 +479,9 @@ function mapProject(row) {
446
479
  // environment_type: 'artifact' | 'terminal' | null (unclassified).
447
480
  // Null stays null — no coercion, no default.
448
481
  environment_type: row.environment_type == null ? null : row.environment_type,
482
+ // versioned_build: 0/1 INTEGER → boolean. version_file: string or null.
483
+ versioned_build: !!row.versioned_build,
484
+ version_file: row.version_file == null ? null : row.version_file,
449
485
  created_at: row.created_at,
450
486
  updated_at: row.updated_at,
451
487
  };
@@ -3074,6 +3110,8 @@ commands["create-project"] = (db, args, ctx) => {
3074
3110
  quiesced: args.quiesced ? 1 : 0,
3075
3111
  visual_protocol: visualProtocol == null ? null : (visualProtocol ? 1 : 0),
3076
3112
  environment_type: environmentType,
3113
+ versioned_build: validateVersionedBuild(args.versioned_build) ?? 0,
3114
+ version_file: validateVersionFile(args.version_file),
3077
3115
  created_at: timestamp, updated_at: timestamp,
3078
3116
  };
3079
3117
  if (!Number.isInteger(row.simultaneity) || row.simultaneity < 1 || row.simultaneity > 100) {
@@ -3081,9 +3119,9 @@ commands["create-project"] = (db, args, ctx) => {
3081
3119
  }
3082
3120
  db.prepare(
3083
3121
  `INSERT INTO projects (id, display_name, repo_path, deploy_type, deploy_slug,
3084
- description, simultaneity, quiesced, visual_protocol, environment_type, created_at, updated_at)
3122
+ description, simultaneity, quiesced, visual_protocol, environment_type, versioned_build, version_file, created_at, updated_at)
3085
3123
  VALUES (@id, @display_name, @repo_path, @deploy_type, @deploy_slug,
3086
- @description, @simultaneity, @quiesced, @visual_protocol, @environment_type, @created_at, @updated_at)`).run(row);
3124
+ @description, @simultaneity, @quiesced, @visual_protocol, @environment_type, @versioned_build, @version_file, @created_at, @updated_at)`).run(row);
3087
3125
  // A deferred backfill (zero/ambiguous matches at the last look) gets one
3088
3126
  // more evaluation now that the match set changed: a new project whose
3089
3127
  // repo descends from the legacy commit inherits the attribution instead
@@ -3142,6 +3180,15 @@ commands["update-project"] = (db, args) => {
3142
3180
  if (args.environment_type !== undefined) {
3143
3181
  patch.environment_type = validateEnvironmentType(args.environment_type);
3144
3182
  }
3183
+ // versioned_build / version_file are NOT covered by the context-change guard:
3184
+ // like environment_type, they are read once from launch args at dispatch
3185
+ // time, so a mid-run change only affects future launches.
3186
+ if (args.versioned_build !== undefined) {
3187
+ patch.versioned_build = validateVersionedBuild(args.versioned_build) ?? 0;
3188
+ }
3189
+ if (args.version_file !== undefined) {
3190
+ patch.version_file = validateVersionFile(args.version_file);
3191
+ }
3145
3192
 
3146
3193
  const contextChanged =
3147
3194
  (args.repo_path !== undefined && args.repo_path !== current.repo_path) ||
package/lib/schema.sql CHANGED
@@ -46,6 +46,13 @@ CREATE TABLE IF NOT EXISTS projects (
46
46
  provenance_crew_release TEXT,
47
47
  provenance_published_at TEXT,
48
48
  provenance_task_id TEXT,
49
+ -- versioned_build: the QA staleness gate (emojimanegg1/muse-crew#3). When 1,
50
+ -- Integrate bumps the version file and QA verifies the tested bundle is at
51
+ -- or past the required version. version_file is the repo-relative path to
52
+ -- the TS file exporting BUILD_NUMBER. Both nullable/0 by design: existing
53
+ -- rows stay unversioned (no backfill).
54
+ versioned_build INTEGER NOT NULL DEFAULT 0 CHECK (versioned_build IN (0, 1)),
55
+ version_file TEXT,
49
56
  created_at TEXT NOT NULL,
50
57
  updated_at TEXT NOT NULL
51
58
  );
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "muse-crew",
3
- "version": "0.14.12",
3
+ "version": "0.15.0",
4
4
  "description": "Opinionated orchestration for Muse — workflows, identities, and tooling for autonomous software development.",
5
5
  "license": "UNLICENSED",
6
6
  "private": false,
@@ -563,6 +563,18 @@ function decideIntegrateRetry(o) {
563
563
  if (rec.malformed) return "park";
564
564
  return ancestor ? "skip-to-publish" : "proceed";
565
565
  }
566
+ // QA version gate (issue #3): S=footer build, R=required. S>=R proves the
567
+ // bundle is at/past the merge (single-branch: version order = containment).
568
+ // Pure — pinned by tests/version-gate.test.js. Null = unreadable/missing.
569
+ function decideVersionGate(S, R) {
570
+ if (!(R > 0)) return { proceed: false, reason: "qa-version-missing",
571
+ remedy: "no build version recorded — re-run Integrate or record dashboard-version." };
572
+ if (!(S > 0)) return { proceed: false, reason: "qa-version-unreadable",
573
+ remedy: "no readable build number in bundle footer — rebuild from versioned source." };
574
+ if (S >= R) return { proceed: true };
575
+ return { proceed: false, reason: "qa-bundle-stale",
576
+ remedy: "bundle build " + S + " behind required " + R + " — builder did not rebuild from current source." };
577
+ }
566
578
  // On a dispatcher retry resumed at Integrate, the run-local releaseDecision
567
579
  // is null (Build doesn't re-run). Hydrate it from the merge record so the
568
580
  // npm Publish gate doesn't park a publishable merge. Pure — pinned by
@@ -1355,10 +1367,26 @@ while (i < STEPS.length) {
1355
1367
  // verdicts.jsonl). No parent verdict gate remains.
1356
1368
  var qaArtifact = false;
1357
1369
  var qaTerminal = false;
1370
+ var qaRequiredVersion = null;
1358
1371
  if (step.name === "QA") {
1359
1372
  var qaExp = (await resolveExperiential()) === "yes";
1360
1373
  qaArtifact = qaExp && SURFACE_ARTIFACT;
1361
1374
  qaTerminal = qaExp && SURFACE_TERMINAL;
1375
+ if (qaArtifact && projectConfig.versioned_build === true) {
1376
+ var vfile = projectConfig.version_file || "client/src/buildNumber.ts";
1377
+ var vgPre = await agent(
1378
+ "Return ONLY JSON. 1. Run, return stdout verbatim:\n" + crewCmd("get-events", { task_id: taskId }) + "\n" +
1379
+ "R = integer after `dashboard-version: ` in newest matching note. 2. If none, run, return stdout verbatim:\n" + crewCmd("get-provenance", { project_id: LAUNCH_PROJECT_ID }) + "\n" +
1380
+ "R = BUILD_NUMBER from: cd " + REPO_PATH + " && git show <source_commit>:" + vfile + " | grep -o 'BUILD_NUMBER = [0-9]*'. If unknown: {\"ok\":false,\"park\":\"qa-version-missing: no build version recorded — re-run Integrate or record dashboard-version.\"}. " +
1381
+ "3. cd " + REPO_PATH + " && git fetch origin && V=$(grep -o 'BUILD_NUMBER = [0-9]*' " + vfile + " | grep -o '[0-9]*'|head -1); if [ $V -lt R ]; then git pull --ff-only; V=$(grep -o 'BUILD_NUMBER = [0-9]*' " + vfile + " | grep -o '[0-9]*'|head -1); fi; " +
1382
+ "if [ $V -lt R ]: {\"ok\":false,\"park\":\"qa-bundle-stale: checkout \" + V + \" < required \" + R + \" — sync past the merge, re-run QA.\"}. 4. Else {\"ok\":true,\"R\":R}.",
1383
+ { key: "qa-version-pre-" + taskId, label: "QA version-gate pre-check",
1384
+ schema: { type: "object", required: ["ok"],
1385
+ properties: { ok: { type: "boolean" }, R: { type: "integer" }, park: { type: "string" } } } }
1386
+ );
1387
+ if (!vgPre.ok) return await parkTask(vgPre.park);
1388
+ qaRequiredVersion = vgPre.R;
1389
+ }
1362
1390
  }
1363
1391
 
1364
1392
  // Reproduce-layer routing (2026-09-16): Sage's Triage classifies the bug's
@@ -1687,6 +1715,14 @@ while (i < STEPS.length) {
1687
1715
  "R5-PUSH (manual R5 resolution only — the normal path pushed inline). Run: " + LIFECYCLE_ENV + LIFECYCLE + " push-target " + taskId + "\n" +
1688
1716
  "PUSHED — report the merged hash (detached prints PUSHED: origin/main (refspec HEAD:main)), VERDICT: PASS; PUSH_SKIPPED — no push attempted (no record + no lock); NO_REMOTE_PUSH — no remote; VERDICT: PASS; ERROR or CONFLICT — report it, VERDICT: FAIL.\n" +
1689
1717
  "NEVER force-push.\n\n" +
1718
+ (projectConfig.versioned_build === true ?
1719
+ "VERSION BUMP (issue #3): after MERGED+PUSHED, bump so QA can prove the bundle contains this fix — QA parks without it.\n" +
1720
+ "1. cd " + REPO_PATH + " && git fetch origin; BR=<branch-from-integration-target>; F=" + (projectConfig.version_file || "client/src/buildNumber.ts") + "\n" +
1721
+ "2. N=$(git show origin/$BR:$F | grep -o 'BUILD_NUMBER = [0-9]*' | grep -o '[0-9]*'); if missing/unparsable: VERDICT: FAIL.\n" +
1722
+ "3. Edit $F: `export const BUILD_NUMBER = $N;` → `export const BUILD_NUMBER = $((N+1));` (keep header comment).\n" +
1723
+ "4. git add $F && git commit -m \"build-number: $((N+1)) - QA version gate\" (SEPARATE commit, never amend).\n" +
1724
+ "5. git push origin $BR; if rejected retry 3x (fetch, re-read N, re-bump, re-commit, push). NEVER force-push. Push MUST succeed or VERDICT: FAIL.\n" +
1725
+ "6. Run, return stdout verbatim:\n" + crewCmd("log-event", { task_id: taskId, type: "note", identity: step.identity, message: "dashboard-version: <new> — QA must test a bundle built from source at or after the merge that recorded this (build <new> or later)." }) + "\n(substitute <new>).\n\n" : "") +
1690
1726
  "Report what happened at each step, ending with exactly one line: VERDICT: PASS or VERDICT: FAIL.";
1691
1727
 
1692
1728
  } else if (step.name === "Publish") {
@@ -2140,6 +2176,8 @@ while (i < STEPS.length) {
2140
2176
  "c3. Start with: SEE_ACT_ARCHIVE_DIR=" + crewHome + "/task-evidence/" + taskId + "/postchange/ node " + crewHome + "/current/lib/see-act.js --url http://localhost:<N>/ aria — read the JSON, log the step. Then: SEE_ACT_ARCHIVE_DIR=" + crewHome + "/task-evidence/" + taskId + "/postchange/ node " + crewHome + "/current/lib/see-act.js --url http://localhost:<N>/ shot — READ the screenshot, log the step. Act on what you see: click, scroll, type, then re-observe, logging each step. Prefer aria (cheap text) to find controls; screenshot when the view changes and for your final verdict frames (one desktop, one mobile). If a click exits non-zero, do NOT retry the same ref blindly: re-run aria first (refs go stale between invocations), then click the fresh ref exactly once. If it still fails, log the failure and move on — a flaky control is a finding, not a loop.\n" +
2141
2177
  "d. Reach: with the session protocol, any flow reachable by N in-page actions is drivable — open the dialog, then confirm it, then judge the result. Without a session (one-shot invocations), anything reachable by (navigate, one action) is testable and sequences needing prior in-page state are not — use a session for those. Report NOT POSSIBLE only when the tooling itself fails (session-start exits 3): a flow you could not reach is not NOT POSSIBLE — name the exact step that stopped you in verdict.json's missing evidence and continue with the mechanical checks.\n" +
2142
2178
  "e. Judge as a user against the task description: is the reported bug fixed AND is nothing else visibly broken? Look for broken layout, overflow, missing or wrong content, stale data, and console errors. Compare against the task's expected behavior, never against source code (you are code-blind). Every frame you captured is already archived under " + crewHome + "/task-evidence/" + taskId + "/postchange/ and indexed in ooda-log.jsonl. A frame you did not read is not evidence. Loading, error, or blank frames never pass. If you cannot complete the loop, say exactly which steps are missing — unknown is not PASS.\n" +
2179
+ (projectConfig.versioned_build === true ?
2180
+ "e2. VERSION (issue #3): footer shows `build <n>` — report `footer_build: <n>` on its own line, or `footer_build: unreadable`. Required — the gate cannot pass without it.\n" : "") +
2143
2181
  "f. Kill ONLY the server you started: pkill -f 'serve-artifact[.]js.*--tag " + taskId + "-qa' — never another task's server. (The [.] keeps pkill from matching its own command line.) Do not leave it running.\n" +
2144
2182
  "Then continue with the mechanical checks below. Your VERDICT covers both the visual and the mechanical checks.\n\n" +
2145
2183
  "MECHANICAL CHECKS:\n" +
@@ -2622,6 +2660,20 @@ while (i < STEPS.length) {
2622
2660
  // reports FAIL, it stands — finding attribution informs follow-up
2623
2661
  // filing only.
2624
2662
  if (step.name === "QA") {
2663
+ if (projectConfig.versioned_build === true && qaArtifact && qaRequiredVersion !== null) {
2664
+ var vgM = /footer_build:\s*(\d+|unreadable)/i.exec(workerText || "");
2665
+ var vgS = (vgM && /^\d+$/.test(vgM[1])) ? parseInt(vgM[1], 10) : null;
2666
+ var vgD = decideVersionGate(vgS, qaRequiredVersion);
2667
+ if (!vgD.proceed) {
2668
+ return await parkTask(vgD.reason + ": " + vgD.remedy + " [built=" + vgS + " required=" + qaRequiredVersion + "]");
2669
+ }
2670
+ await agent(
2671
+ "Run in shell and return the stdout verbatim:\n" + crewCmd("log-event", {
2672
+ task_id: taskId, type: "note", identity: step.identity,
2673
+ message: "version-check: built=" + vgS + " >= required=" + qaRequiredVersion + " → testing now" }),
2674
+ { key: "record-version-check-" + taskId, label: "Recording version-gate pass" }
2675
+ );
2676
+ }
2625
2677
  var contentFindings = extractContentFindings(workerText);
2626
2678
  if (!contentFindings.ok) {
2627
2679
  // Unknown attribution fails closed: a missing or malformed
@@ -476,7 +476,13 @@ for (var pi = 0; pi < projects.length; pi++) {
476
476
  // The user-facing surface for experiential QA routing (artifact |
477
477
  // terminal | null=unclassified). Carried alongside deploy_type — it is a
478
478
  // separate axis, not a redeclaration of the deployment target.
479
- environment_type: proj.environment_type || null
479
+ environment_type: proj.environment_type || null,
480
+ // QA staleness gate (emojimanegg1/muse-crew#3): versioned_build opts the
481
+ // project into the Integrate version bump + QA bundle-freshness gate;
482
+ // version_file is the repo-relative path of the TS file exporting
483
+ // BUILD_NUMBER (null when unversioned).
484
+ versioned_build: !!proj.versioned_build,
485
+ version_file: proj.version_file || null
480
486
  };
481
487
  }
482
488
 
@@ -561,6 +567,12 @@ const eligible = [];
561
567
  const retryCandidates = [];
562
568
  const rejectionCandidates = [];
563
569
 
570
+ // Dispatch-decision log (2026-09-24): per-reason ineligibility counters for
571
+ // the .dispatch-decisions.jsonl declaration. Sparse — only non-zero keys are
572
+ // emitted. Additive only; the eligibility decisions below are unchanged.
573
+ var decisionSkipped = {};
574
+ function countSkipped(reason) { decisionSkipped[reason] = (decisionSkipped[reason] || 0) + 1; }
575
+
564
576
  // ── Retry cap ────────────────────────────────────────────────────────
565
577
  // Symphony owns phase-redispatch policy (coordination layer): the spec
566
578
  // defines backoff but no attempt cap ("implementation-defined"), so the
@@ -591,12 +603,13 @@ var MAX_CONSECUTIVE_REJECTIONS = configInt(config, "maxConsecutiveRejections", 2
591
603
 
592
604
  for (var t = 0; t < allTasks.length; t++) {
593
605
  var task = allTasks[t];
594
- if (task.blocked) continue;
606
+ if (task.blocked) { countSkipped("blocked"); continue; }
595
607
 
596
608
  // Skip tasks with an active dispatch reservation: the worker launched a
597
609
  // workflow for this task on a previous tick, but the workflow has not yet
598
610
  // self-claimed (claims can take 15+ minutes for cron-launched runs).
599
611
  if (reservedTaskIds.has(task.id)) {
612
+ countSkipped("reserved");
600
613
  log("Skipped \"" + task.title + "\" — active dispatch reservation (workflow launched, claim pending)");
601
614
  continue;
602
615
  }
@@ -604,6 +617,7 @@ for (var t = 0; t < allTasks.length; t++) {
604
617
  // Skip tasks from quiesced projects
605
618
  var taskProject = task.project || DEFAULT_PROJECT;
606
619
  if (quiescedProjects[taskProject]) {
620
+ countSkipped("quiesced");
607
621
  log("Skipped \"" + task.title + "\" — project " + taskProject + " is quiesced");
608
622
  continue;
609
623
  }
@@ -614,6 +628,7 @@ for (var t = 0; t < allTasks.length; t++) {
614
628
  // task never moves. Set repo_path via updateproject to re-enable.
615
629
  var taskProjCfg = PROJECTS[taskProject];
616
630
  if (!taskProjCfg || !taskProjCfg.repo_path) {
631
+ countSkipped("no_repo");
617
632
  log("Skipped \"" + task.title + "\" — project " + taskProject + " has no repo_path configured");
618
633
  continue;
619
634
  }
@@ -633,12 +648,14 @@ for (var t = 0; t < allTasks.length; t++) {
633
648
  // (clears) next_phase on its successful self-claim, exactly once.
634
649
  if (task.next_phase) {
635
650
  if (task.state !== "todo" && task.state !== "in_progress") {
651
+ countSkipped("next_phase_state");
636
652
  log("Skipped \"" + task.title + "\" — next_phase \"" + task.next_phase + "\" set but state is " + task.state + "; left set for inspection");
637
653
  continue;
638
654
  }
639
- if (latest && latest.status === "running") continue; // work in flight
655
+ if (latest && latest.status === "running") { countSkipped("in_flight"); continue; } // work in flight
640
656
  var npIdx = steps.indexOf(task.next_phase);
641
657
  if (npIdx < 0) {
658
+ countSkipped("next_phase_bad_step");
642
659
  log("Skipped \"" + task.title + "\" — next_phase \"" + task.next_phase + "\" not in " + workflow + " step registry; left set for a corrected recover-task");
643
660
  continue;
644
661
  }
@@ -654,19 +671,20 @@ for (var t = 0; t < allTasks.length; t++) {
654
671
  continue;
655
672
  }
656
673
 
657
- if (task.state !== "in_progress") continue;
674
+ if (task.state !== "in_progress") { countSkipped("terminal_state"); continue; }
658
675
 
659
676
  if (!latest) {
660
677
  eligible.push({ task: task, startStep: 0, reason: "no_session", workflow: workflow });
661
678
  continue;
662
679
  }
663
680
 
664
- if (latest.status === "running") continue; // work in flight
681
+ if (latest.status === "running") { countSkipped("in_flight"); continue; } // work in flight
665
682
 
666
683
  if (latest.status === "completed") {
667
684
  var stepName = latest.step || "";
668
685
  var stepIndex = steps.indexOf(stepName);
669
686
  if (stepIndex < 0) {
687
+ countSkipped("unknown_step");
670
688
  log("Skipped \"" + task.title + "\" — step \"" + stepName + "\" not in " + workflow + " step registry");
671
689
  continue;
672
690
  }
@@ -695,6 +713,7 @@ for (var t = 0; t < allTasks.length; t++) {
695
713
  var failedStep = latest.step || "";
696
714
  var retryIdx = steps.indexOf(failedStep);
697
715
  if (retryIdx < 0) {
716
+ countSkipped("unknown_step");
698
717
  log("Skipped \"" + task.title + "\" — step \"" + failedStep + "\" not in " + workflow + " step registry");
699
718
  continue;
700
719
  }
@@ -756,6 +775,7 @@ if (rejectionCandidates.length > 0) {
756
775
  // the task stays retryable and is re-parked next tick, so this is
757
776
  // fail-safe. The dashboard stamps retry_reset_at on parked→todo, which
758
777
  // restarts both counters mechanically.
778
+ var parkedTaskIds = [];
759
779
  if (parkJobs.length > 0) {
760
780
  var parkSteps = [];
761
781
  for (var pji = 0; pji < parkJobs.length; pji++) {
@@ -775,6 +795,7 @@ if (parkJobs.length > 0) {
775
795
  { key: "park-batch", label: "Parking " + parkJobs.length + " task(s) at retry cap" }
776
796
  );
777
797
  log("Parked " + parkJobs.length + " task(s)");
798
+ parkedTaskIds = parkJobs.map(function(pj) { return pj.task.id; });
778
799
  } catch (parkErr) {
779
800
  log("WARNING: park batch failed (" + (parkErr.message || String(parkErr)).slice(0, 200) + ") — tasks remain retryable");
780
801
  }
@@ -959,6 +980,7 @@ for (var ei = 0; ei < eligible.length; ei++) {
959
980
  toProcess.push(eitem);
960
981
  inFlightByProject[ep] = current + 1;
961
982
  } else {
983
+ countSkipped("at_limit");
962
984
  log("Skipped \"" + eitem.task.title + "\" — project " + ep + " at simultaneity limit (" + limit + ")");
963
985
  }
964
986
  }
@@ -1164,6 +1186,63 @@ if (recommended.length > 0) {
1164
1186
  recommended = acquired;
1165
1187
  }
1166
1188
 
1189
+ // ── 4c. Dispatch-decision log ──────────────────────────────────
1190
+ // Hoverboat (2026-09-24): the dispatcher declares its own decision, in its
1191
+ // own bytes, on crew-home disk — one append-only line per completed tick in
1192
+ // $CREW_HOME/.dispatch-decisions.jsonl. The observer reads the file
1193
+ // directly (no LLM re-emission, no platform ferry), so the plausible-value
1194
+ // mutation that poisoned the old dispatch_sample attribution cannot occur.
1195
+ // Launch declarations are verified mechanically against crew DB effects
1196
+ // (platform_run_tasks / dispatch_reservations); a stand-down is the
1197
+ // dispatcher's explicit declaration with board context, auditable without
1198
+ // re-deriving eligibility. tick_seq couples each line to the .tick-releases
1199
+ // line for the same poll — a tick-release line with no decision line means
1200
+ // the tick died after poll-ack and is visible as such.
1201
+ // seq is the non-empty line count + 1 of this file (file order is the
1202
+ // proof — no wall-clock calls, see tests/determinism.test.js).
1203
+ // Fire-and-forget: the script swallows its own errors and exits 0, the
1204
+ // agent call carries no schema, and the whole call is wrapped in
1205
+ // try/catch — a failed write can never fail the tick. The script echoes
1206
+ // the appended line (same rooms #12–#14 empty-result contract as §4b).
1207
+ var dispatchDecisionScript = [
1208
+ "var crewHome=process.argv[1];",
1209
+ "var decisionJson=process.argv[2];",
1210
+ "try{",
1211
+ "var fs=require(\"fs\"),path=require(\"path\");",
1212
+ "var decision=JSON.parse(decisionJson);",
1213
+ "if(!decision||typeof decision.decision!==\"string\"||!Array.isArray(decision.launched)){throw new Error(\"bad decision payload\");}",
1214
+ "var release=path.basename(path.resolve(crewHome,fs.readlinkSync(path.join(crewHome,\"current\"))));",
1215
+ "var file=path.join(crewHome,\".dispatch-decisions.jsonl\");",
1216
+ "var count=0;",
1217
+ "try{var lines=fs.readFileSync(file,\"utf8\").split(\"\\n\");for(var i=0;i<lines.length;i++){if(lines[i].trim()!==\"\"){count++;}}}catch(e){}",
1218
+ "var tickSeq=0;",
1219
+ "try{var tlines=fs.readFileSync(path.join(crewHome,\".tick-releases.jsonl\"),\"utf8\").split(\"\\n\");for(var j=0;j<tlines.length;j++){if(tlines[j].trim()!==\"\"){tickSeq++;}}}catch(e){}",
1220
+ "var line=JSON.stringify({seq:count+1,tick_seq:tickSeq,release:release,decision:decision.decision,launched:decision.launched,completed:decision.completed||[],parked:decision.parked||[],board:decision.board||null,skipped:decision.skipped||{},partial:!!decision.partial});",
1221
+ "fs.appendFileSync(file,line+\"\\n\");",
1222
+ "process.stdout.write(line+\"\\n\");",
1223
+ "}catch(e){}",
1224
+ "process.exit(0);"
1225
+ ].join("");
1226
+ var launchedRecs = recommended.map(function(r) { return { task_id: r.task_id, workflow: r.workflow, step: r.step }; });
1227
+ var decisionObj = {
1228
+ decision: launchedRecs.length > 0 ? "launch" : "stand_down",
1229
+ launched: launchedRecs,
1230
+ completed: completed.map(function(r) { return r.task_id; }),
1231
+ parked: parkedTaskIds,
1232
+ board: { seen: allTasks.length, eligible: eligible.length },
1233
+ skipped: decisionSkipped,
1234
+ partial: partial
1235
+ };
1236
+ var dispatchDecisionCmd = "node -e '" + dispatchDecisionScript + "' '" + crewHome.replace(/'/g, "'\\''") + "' '" + JSON.stringify(decisionObj).replace(/'/g, "'\\''") + "'";
1237
+ try {
1238
+ await agent(
1239
+ "Record this tick's dispatch decision.\nRun in shell and return the stdout verbatim:\n" + dispatchDecisionCmd,
1240
+ { key: "dispatch-decision", label: "Recording dispatch decision" }
1241
+ );
1242
+ } catch (e) {
1243
+ log("WARNING: dispatch-decision log write failed (tick continues): " + e.message);
1244
+ }
1245
+
1167
1246
  var msg = "Dispatch complete.";
1168
1247
  if (recommended.length > 0) {
1169
1248
  msg += " Recommended: " + recommended.map(function(r) { return r.workflow + "/" + r.step + " for " + r.task_id; }).join(", ") + ".";
@@ -571,6 +571,18 @@ function decideIntegrateRetry(o) {
571
571
  if (rec.malformed) return "park";
572
572
  return ancestor ? "skip-to-publish" : "proceed";
573
573
  }
574
+ // QA version gate (issue #3): S=footer build, R=required. S>=R proves the
575
+ // bundle is at/past the merge (single-branch: version order = containment).
576
+ // Pure — pinned by tests/version-gate.test.js. Null = unreadable/missing.
577
+ function decideVersionGate(S, R) {
578
+ if (!(R > 0)) return { proceed: false, reason: "qa-version-missing",
579
+ remedy: "no build version recorded — re-run Integrate or record dashboard-version." };
580
+ if (!(S > 0)) return { proceed: false, reason: "qa-version-unreadable",
581
+ remedy: "no readable build number in bundle footer — rebuild from versioned source." };
582
+ if (S >= R) return { proceed: true };
583
+ return { proceed: false, reason: "qa-bundle-stale",
584
+ remedy: "bundle build " + S + " behind required " + R + " — builder did not rebuild from current source." };
585
+ }
574
586
  // On a dispatcher retry resumed at Integrate, the run-local releaseDecision
575
587
  // is null (Build doesn't re-run). Hydrate it from the merge record so the
576
588
  // npm Publish gate doesn't park a publishable merge. Pure — pinned by
@@ -1300,6 +1312,22 @@ while (i < STEPS.length) {
1300
1312
  var qaExperiential = false;
1301
1313
  if (step.name === "QA") {
1302
1314
  qaExperiential = (await resolveExperiential()) === "yes" && SURFACE_CLASSIFIED;
1315
+ var qaRequiredVersion = null;
1316
+ if (qaExperiential && SURFACE_ARTIFACT && projectConfig.versioned_build === true) {
1317
+ var vfileStd = projectConfig.version_file || "client/src/buildNumber.ts";
1318
+ var vgPreStd = await agent(
1319
+ "Return ONLY JSON. 1. Run, return stdout verbatim:\n" + crewCmd("get-events", { task_id: taskId }) + "\n" +
1320
+ "R = integer after `dashboard-version: ` in newest matching note. 2. If none, run, return stdout verbatim:\n" + crewCmd("get-provenance", { project_id: LAUNCH_PROJECT_ID }) + "\n" +
1321
+ "R = BUILD_NUMBER from: cd " + REPO_PATH + " && git show <source_commit>:" + vfileStd + " | grep -o 'BUILD_NUMBER = [0-9]*'. If unknown: {\"ok\":false,\"park\":\"qa-version-missing: no build version recorded — re-run Integrate or record dashboard-version.\"}. " +
1322
+ "3. cd " + REPO_PATH + " && git fetch origin && V=$(grep -o 'BUILD_NUMBER = [0-9]*' " + vfileStd + " | grep -o '[0-9]*'|head -1); if [ $V -lt R ]; then git pull --ff-only; V=$(grep -o 'BUILD_NUMBER = [0-9]*' " + vfileStd + " | grep -o '[0-9]*'|head -1); fi; " +
1323
+ "if [ $V -lt R ]: {\"ok\":false,\"park\":\"qa-bundle-stale: checkout \" + V + \" < required \" + R + \" — sync past the merge, re-run QA.\"}. 4. Else {\"ok\":true,\"R\":R}.",
1324
+ { key: "qa-version-pre-" + taskId, label: "QA version-gate pre-check",
1325
+ schema: { type: "object", required: ["ok"],
1326
+ properties: { ok: { type: "boolean" }, R: { type: "integer" }, park: { type: "string" } } } }
1327
+ );
1328
+ if (!vgPreStd.ok) return await parkTask(vgPreStd.park);
1329
+ qaRequiredVersion = vgPreStd.R;
1330
+ }
1303
1331
  }
1304
1332
 
1305
1333
  // Step-specific instructions
@@ -1576,6 +1604,14 @@ while (i < STEPS.length) {
1576
1604
  "R5-PUSH (manual R5 resolution only — the normal path pushed inline). Run: " + LIFECYCLE_ENV + LIFECYCLE + " push-target " + taskId + "\n" +
1577
1605
  "PUSHED — report the merged hash (detached prints PUSHED: origin/main (refspec HEAD:main)), VERDICT: PASS; PUSH_SKIPPED — no push attempted (no record + no lock); NO_REMOTE_PUSH — no remote; VERDICT: PASS; ERROR or CONFLICT — report it, VERDICT: FAIL.\n" +
1578
1606
  "NEVER force-push.\n\n" +
1607
+ (projectConfig.versioned_build === true ?
1608
+ "VERSION BUMP (issue #3): after MERGED+PUSHED, bump so QA can prove the bundle contains this fix — QA parks without it.\n" +
1609
+ "1. cd " + REPO_PATH + " && git fetch origin; BR=<branch-from-integration-target>; F=" + (projectConfig.version_file || "client/src/buildNumber.ts") + "\n" +
1610
+ "2. N=$(git show origin/$BR:$F | grep -o 'BUILD_NUMBER = [0-9]*' | grep -o '[0-9]*'); if missing/unparsable: VERDICT: FAIL.\n" +
1611
+ "3. Edit $F: `export const BUILD_NUMBER = $N;` → `export const BUILD_NUMBER = $((N+1));` (keep header comment).\n" +
1612
+ "4. git add $F && git commit -m \"build-number: $((N+1)) - QA version gate\" (SEPARATE commit, never amend).\n" +
1613
+ "5. git push origin $BR; if rejected retry 3x (fetch, re-read N, re-bump, re-commit, push). NEVER force-push. Push MUST succeed or VERDICT: FAIL.\n" +
1614
+ "6. Run, return stdout verbatim:\n" + crewCmd("log-event", { task_id: taskId, type: "note", identity: step.identity, message: "dashboard-version: <new> — QA must test a bundle built from source at or after the merge that recorded this (build <new> or later)." }) + "\n(substitute <new>).\n\n" : "") +
1579
1615
  "Report what happened at each step, ending with exactly one line: VERDICT: PASS or VERDICT: FAIL.";
1580
1616
 
1581
1617
  } else if (step.name === "Publish") {
@@ -2029,6 +2065,8 @@ while (i < STEPS.length) {
2029
2065
  "c3. Start with: SEE_ACT_ARCHIVE_DIR=" + crewHome + "/task-evidence/" + taskId + "/postchange/ node " + crewHome + "/current/lib/see-act.js --url http://localhost:<N>/ aria — read the JSON, log the step. Then: SEE_ACT_ARCHIVE_DIR=" + crewHome + "/task-evidence/" + taskId + "/postchange/ node " + crewHome + "/current/lib/see-act.js --url http://localhost:<N>/ shot — READ the screenshot, log the step. Act on what you see: click, scroll, type, then re-observe, logging each step. Prefer aria (cheap text) to find controls; screenshot when the view changes and for your final verdict frames (one desktop, one mobile). If a click exits non-zero, do NOT retry the same ref blindly: re-run aria first (refs go stale between invocations), then click the fresh ref exactly once. If it still fails, log the failure and move on — a flaky control is a finding, not a loop.\n" +
2030
2066
  "d. Reach: with the session protocol, any flow reachable by N in-page actions is drivable — open the deck, then click Study, then judge the study view. Without a session (one-shot invocations), anything reachable by (navigate, one action) is testable and sequences needing prior in-page state are not — use a session for those. Report NOT POSSIBLE only when the tooling itself fails (session-start exits 3): a flow you could not reach is not NOT POSSIBLE — name the exact step that stopped you in verdict.json's missing evidence and judge what you did reach.\n" +
2031
2067
  "e. Judge as a user against the task description: does the change render correctly? Look for broken layout, overflow, missing or wrong content, stale data, and console errors. Compare against the task's expected behavior, never against source code (you are code-blind). Every frame you captured is already archived under " + crewHome + "/task-evidence/" + taskId + "/postchange/ and indexed in ooda-log.jsonl. A frame you did not read is not evidence. Loading, error, or blank frames never pass. If you cannot complete the loop, say exactly which steps are missing — unknown is not PASS.\n" +
2068
+ (projectConfig.versioned_build === true ?
2069
+ "e2. VERSION (issue #3): footer shows `build <n>` — report `footer_build: <n>` on its own line, or `footer_build: unreadable`. Required — the gate cannot pass without it.\n" : "") +
2032
2070
  "f. Kill ONLY the server you started: pkill -f 'serve-artifact[.]js.*--tag " + taskId + "-qa' — never another task's server. (The [.] keeps pkill from matching its own command line.) Do not leave it running.\n" +
2033
2071
  "Then continue with the mechanical checks below. Your VERDICT covers both the visual and the mechanical checks.\n\n" +
2034
2072
  "STEP 2: Verify data integrity via the crew API.\n" +
@@ -2374,6 +2412,20 @@ while (i < STEPS.length) {
2374
2412
  // reports FAIL, it stands — finding attribution informs follow-up
2375
2413
  // filing only.
2376
2414
  if (step.name === "QA") {
2415
+ if (projectConfig.versioned_build === true && SURFACE_ARTIFACT && qaRequiredVersion !== null) {
2416
+ var vgM = /footer_build:\s*(\d+|unreadable)/i.exec(workerText || "");
2417
+ var vgS = (vgM && /^\d+$/.test(vgM[1])) ? parseInt(vgM[1], 10) : null;
2418
+ var vgD = decideVersionGate(vgS, qaRequiredVersion);
2419
+ if (!vgD.proceed) {
2420
+ return await parkTask(vgD.reason + ": " + vgD.remedy + " [built=" + vgS + " required=" + qaRequiredVersion + "]");
2421
+ }
2422
+ await agent(
2423
+ "Run in shell and return the stdout verbatim:\n" + crewCmd("log-event", {
2424
+ task_id: taskId, type: "note", identity: step.identity,
2425
+ message: "version-check: built=" + vgS + " >= required=" + qaRequiredVersion + " → testing now" }),
2426
+ { key: "record-version-check-" + taskId, label: "Recording version-gate pass" }
2427
+ );
2428
+ }
2377
2429
  var contentFindings = extractContentFindings(workerText);
2378
2430
  if (!contentFindings.ok) {
2379
2431
  // Unknown attribution fails closed: a missing or malformed