muse-crew 0.14.12 → 0.15.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/API.md +2 -0
- package/docs/decisions/AGENTS.md +6 -0
- package/docs/decisions/dispatch-decision-log.md +69 -0
- package/lib/crew-api.js +49 -2
- package/lib/schema.sql +7 -0
- package/package.json +1 -1
- package/workflows/bugfix.js +52 -0
- package/workflows/crew-dispatch.js +84 -5
- package/workflows/standard.js +52 -0
package/API.md
CHANGED
|
@@ -489,6 +489,8 @@ Register a new project.
|
|
|
489
489
|
| `quiesced` | boolean | no | Start paused; defaults to false |
|
|
490
490
|
| `visual_protocol` | boolean or null | no | Tri-state: `null` = inherit crew default (off), `true` = enable for this project, `false` = explicit off. Defaults to null. |
|
|
491
491
|
| `environment_type` | `artifact` · `terminal` · null | no | User-facing surface for experiential QA routing: `artifact` = a rendered web UI (Hazel drives it with the see-act browser loop), `terminal` = a CLI (Hazel drives it herself, keeping transcripts). `null` = unclassified: no experiential QA. `deploy_type` names the deployment target, but `deploy_type: "artifact"` remains a legacy artifact-surface signal so pre-field projects keep today's experiential QA (the migration does not backfill the column); on conflicting config artifact wins. Omitted means auto-classify from `repo_path` (see surface classification below); an explicit value, including explicit `null`, always wins. |
|
|
492
|
+
| `versioned_build` | boolean | no | Opt-in to the QA staleness gate (issue #3): when true, Integrate bumps `BUILD_NUMBER` in `version_file` after each merge (separate commit, pushed), and QA verifies the tested bundle's footer `build <n>` is at/past the recorded version before testing. Defaults to false (unversioned); the migration does not backfill. |
|
|
493
|
+
| `version_file` | string or null | no | Repo-relative path of the TypeScript file exporting `BUILD_NUMBER` (e.g. `client/src/buildNumber.ts`). Only meaningful when `versioned_build` is true; defaults to `client/src/buildNumber.ts` when omitted. |
|
|
492
494
|
|
|
493
495
|
**Surface classification (repo-derived defaults):** when `repo_path` is provided and `environment_type` is omitted, `create-project` mechanically classifies the repo at registration via `lib/classify-surface.js` — `space.json` present → `artifact`; else a non-empty `package.json` `bin` → `terminal`; otherwise unclassified (`null`). The classifier never throws: bad paths and malformed files classify as unclassified. Classification only fills omitted fields — an explicitly supplied `environment_type` (including explicit `null`) always wins, and deploy defaults follow the *effective* surface, so an explicit `"terminal"` never triggers artifact deploy defaults.
|
|
494
496
|
|
package/docs/decisions/AGENTS.md
CHANGED
|
@@ -6,6 +6,12 @@ history lives here. One canonical section per `<a id>` anchor — a critic
|
|
|
6
6
|
re-verified that every `docs/decisions/*.md#anchor` reference in
|
|
7
7
|
`workflows/*.js` resolves (45 unique references).
|
|
8
8
|
|
|
9
|
+
## dispatch-decision-log.md — Decision history: dispatch decision log (hoverboat)
|
|
10
|
+
|
|
11
|
+
- `#checksum-vs-hoverboat` — The ferry-corruption problem and the two candidates
|
|
12
|
+
- `#why-hoverboat-won` — Availability, trust, subtraction, primary-source rationale
|
|
13
|
+
- `#what-was-built` — Declaration shape, skip counters, observer ingestion, retirements
|
|
14
|
+
|
|
9
15
|
## publish-path.md — Decision history: publish path
|
|
10
16
|
|
|
11
17
|
- `#fire-and-forget-trigger` — Fire-and-forget trigger + workflow-owned observation
|
|
@@ -0,0 +1,69 @@
|
|
|
1
|
+
# Dispatch decision log (hoverboat) — 2026-09-24
|
|
2
|
+
|
|
3
|
+
<a id="checksum-vs-hoverboat"></a>
|
|
4
|
+
## The problem
|
|
5
|
+
|
|
6
|
+
The Gate 2 observer attributed dispatch decisions by ferrying platform
|
|
7
|
+
`scheduler.job_runs` rows through an LLM paste (tick worker: `muse.db`
|
|
8
|
+
result → heredoc → `record_platform.sh --dispatch` → `dispatch_sample`
|
|
9
|
+
rows → `attribute_dispatch.sh` → `dispatch_decision` rows). The paste
|
|
10
|
+
mutated plausible values: a 2026-09-24 consistency scan found 275
|
|
11
|
+
same-`run_id` contradictory `dispatch_sample` rows and 65 derived
|
|
12
|
+
`dispatch_decision` rows based on corrupt data (5.7% of samples
|
|
13
|
+
quarantined; 63.6% of ticks touched). The older `quarantine_suspects.sh`
|
|
14
|
+
handled near-neighbor UUID mutations, not plausible value mutation.
|
|
15
|
+
|
|
16
|
+
## The two candidates
|
|
17
|
+
|
|
18
|
+
**Checksum ferry** (containment): keep the ferry; the worker checksums the
|
|
19
|
+
raw tool-result JSON before the paste and `record_platform.sh` verifies
|
|
20
|
+
after. Corrupt pastes become loud absences instead of silent corruption.
|
|
21
|
+
|
|
22
|
+
**Hoverboat** (local declaration): the dispatcher writes its decision to
|
|
23
|
+
`$CREW_HOME/.dispatch-decisions.jsonl` in its own bytes (deterministic JS,
|
|
24
|
+
same courier shape as the proven §4b tick-release writer); the observer
|
|
25
|
+
reads the file directly and verifies launch declarations mechanically
|
|
26
|
+
against crew DB effects (`platform_run_tasks`, `dispatch_reservations`).
|
|
27
|
+
|
|
28
|
+
<a id="why-hoverboat-won"></a>
|
|
29
|
+
## Why hoverboat won
|
|
30
|
+
|
|
31
|
+
1. **Availability.** Corruption touched ~63.6% of ticks. Checksum converts
|
|
32
|
+
corruption into absence, but the entry demo needs ten consecutive ticks
|
|
33
|
+
each recording a decision — absence fails the demo as surely as
|
|
34
|
+
corruption. Hoverboat's channel has no LLM re-emission at all.
|
|
35
|
+
2. **Trust.** The checksum would be computed by the same worker agent that
|
|
36
|
+
corrupts the paste — asking the corrupting agent to honestly report its
|
|
37
|
+
own corruption. Hoverboat's declaration is composed by deterministic
|
|
38
|
+
workflow JS and checked against independent DB effects.
|
|
39
|
+
3. **Subtraction.** Checksum keeps the ferry, the rolling 25h re-record
|
|
40
|
+
window, and the quarantine machinery permanently busy. Hoverboat deletes
|
|
41
|
+
the sample→decision attribution entirely; the platform ferry remains
|
|
42
|
+
only as platform-health telemetry.
|
|
43
|
+
4. **The crew already knows.** Having the platform tell the observer what
|
|
44
|
+
the crew decided, through an LLM paste, when the crew can write it
|
|
45
|
+
directly, is the ferry. The declaration is the primary source; the DB
|
|
46
|
+
effects are the independent verification.
|
|
47
|
+
|
|
48
|
+
Rejected alternative considered: a deterministic Crew API action for the
|
|
49
|
+
write. It would not remove the agent from the path (workflows reach the
|
|
50
|
+
CLI only through `agent()` shell calls), so it adds API surface for no
|
|
51
|
+
integrity gain. The embedded courier is the proven seam.
|
|
52
|
+
|
|
53
|
+
<a id="what-was-built"></a>
|
|
54
|
+
## What was built
|
|
55
|
+
|
|
56
|
+
- `workflows/crew-dispatch.js` §4c: per-tick declaration
|
|
57
|
+
`{seq, tick_seq, release, decision, launched[], completed[], parked[],
|
|
58
|
+
board{seen,eligible}, skipped{reason:count sparse}, partial}`.
|
|
59
|
+
`tick_seq` couples each line to the `.tick-releases.jsonl` line for the
|
|
60
|
+
same poll — a tick-release line with no decision line means the tick died
|
|
61
|
+
after poll-ack, visibly.
|
|
62
|
+
- Per-reason ineligibility counters (`countSkipped`) at every skip site in
|
|
63
|
+
the eligibility loop + the simultaneity-limit skip, so stand-downs are
|
|
64
|
+
auditable without re-deriving eligibility.
|
|
65
|
+
- Observer: `ingest_decisions.sh` replaces `attribute_dispatch.sh`;
|
|
66
|
+
`snapshot.sh` dispatch section rewritten around decision rows.
|
|
67
|
+
- Retired: sample→decision attribution, `state/attributed.json` watermark.
|
|
68
|
+
Kept: `record_platform.sh --dispatch` as platform-health telemetry,
|
|
69
|
+
`quarantine_suspects.sh` for the health samples' ID mutations.
|
package/lib/crew-api.js
CHANGED
|
@@ -73,6 +73,21 @@ function validateEnvironmentType(value) {
|
|
|
73
73
|
if (value === "artifact" || value === "terminal") return value;
|
|
74
74
|
throw usageError("environment_type must be 'artifact', 'terminal', or null.");
|
|
75
75
|
}
|
|
76
|
+
// versioned_build validation: boolean. undefined means "not provided"
|
|
77
|
+
// (create: defaults to false; update: leaves the field alone). Coerces
|
|
78
|
+
// truthy/falsy to 1/0 for the INTEGER column.
|
|
79
|
+
function validateVersionedBuild(value) {
|
|
80
|
+
if (value === undefined) return undefined;
|
|
81
|
+
return value ? 1 : 0;
|
|
82
|
+
}
|
|
83
|
+
// version_file validation: string or null. undefined means "not provided"
|
|
84
|
+
// (create: defaults to null; update: leaves the field alone). Empty string
|
|
85
|
+
// is rejected — use null to clear.
|
|
86
|
+
function validateVersionFile(value) {
|
|
87
|
+
if (value === undefined || value === null) return null;
|
|
88
|
+
if (typeof value === "string" && value.length > 0) return value;
|
|
89
|
+
throw usageError("version_file must be a non-empty string or null.");
|
|
90
|
+
}
|
|
76
91
|
|
|
77
92
|
// already_merged_sha validation (room #16 blocker 11, 2026-09-18): the
|
|
78
93
|
// workflow-verified already-merged sha as structured session control state.
|
|
@@ -315,6 +330,24 @@ function openDb(crewHome) {
|
|
|
315
330
|
if (!/duplicate column name/i.test(e.message)) throw e;
|
|
316
331
|
}
|
|
317
332
|
}
|
|
333
|
+
// Migration: add versioned_build + version_file columns if missing (QA
|
|
334
|
+
// staleness gate, emojimanegg1/muse-crew#3, 2026-09-24). Same idempotent
|
|
335
|
+
// PRAGMA-check pattern; versioned_build defaults 0 (unversioned by
|
|
336
|
+
// design, no backfill), version_file stays NULL.
|
|
337
|
+
if (!projectCols.some((c) => c.name === "versioned_build")) {
|
|
338
|
+
try {
|
|
339
|
+
db.exec("ALTER TABLE projects ADD COLUMN versioned_build INTEGER NOT NULL DEFAULT 0 CHECK (versioned_build IN (0, 1))");
|
|
340
|
+
} catch (e) {
|
|
341
|
+
if (!/duplicate column name/i.test(e.message)) throw e;
|
|
342
|
+
}
|
|
343
|
+
}
|
|
344
|
+
if (!projectCols.some((c) => c.name === "version_file")) {
|
|
345
|
+
try {
|
|
346
|
+
db.exec("ALTER TABLE projects ADD COLUMN version_file TEXT");
|
|
347
|
+
} catch (e) {
|
|
348
|
+
if (!/duplicate column name/i.test(e.message)) throw e;
|
|
349
|
+
}
|
|
350
|
+
}
|
|
318
351
|
// Migration: per-project provenance columns (room #15 blocker 8,
|
|
319
352
|
// 2026-09-18). Same idempotent PRAGMA-check pattern; ADD COLUMN without
|
|
320
353
|
// a default leaves existing rows NULL, which is "no provenance yet" by
|
|
@@ -446,6 +479,9 @@ function mapProject(row) {
|
|
|
446
479
|
// environment_type: 'artifact' | 'terminal' | null (unclassified).
|
|
447
480
|
// Null stays null — no coercion, no default.
|
|
448
481
|
environment_type: row.environment_type == null ? null : row.environment_type,
|
|
482
|
+
// versioned_build: 0/1 INTEGER → boolean. version_file: string or null.
|
|
483
|
+
versioned_build: !!row.versioned_build,
|
|
484
|
+
version_file: row.version_file == null ? null : row.version_file,
|
|
449
485
|
created_at: row.created_at,
|
|
450
486
|
updated_at: row.updated_at,
|
|
451
487
|
};
|
|
@@ -3074,6 +3110,8 @@ commands["create-project"] = (db, args, ctx) => {
|
|
|
3074
3110
|
quiesced: args.quiesced ? 1 : 0,
|
|
3075
3111
|
visual_protocol: visualProtocol == null ? null : (visualProtocol ? 1 : 0),
|
|
3076
3112
|
environment_type: environmentType,
|
|
3113
|
+
versioned_build: validateVersionedBuild(args.versioned_build) ?? 0,
|
|
3114
|
+
version_file: validateVersionFile(args.version_file),
|
|
3077
3115
|
created_at: timestamp, updated_at: timestamp,
|
|
3078
3116
|
};
|
|
3079
3117
|
if (!Number.isInteger(row.simultaneity) || row.simultaneity < 1 || row.simultaneity > 100) {
|
|
@@ -3081,9 +3119,9 @@ commands["create-project"] = (db, args, ctx) => {
|
|
|
3081
3119
|
}
|
|
3082
3120
|
db.prepare(
|
|
3083
3121
|
`INSERT INTO projects (id, display_name, repo_path, deploy_type, deploy_slug,
|
|
3084
|
-
description, simultaneity, quiesced, visual_protocol, environment_type, created_at, updated_at)
|
|
3122
|
+
description, simultaneity, quiesced, visual_protocol, environment_type, versioned_build, version_file, created_at, updated_at)
|
|
3085
3123
|
VALUES (@id, @display_name, @repo_path, @deploy_type, @deploy_slug,
|
|
3086
|
-
@description, @simultaneity, @quiesced, @visual_protocol, @environment_type, @created_at, @updated_at)`).run(row);
|
|
3124
|
+
@description, @simultaneity, @quiesced, @visual_protocol, @environment_type, @versioned_build, @version_file, @created_at, @updated_at)`).run(row);
|
|
3087
3125
|
// A deferred backfill (zero/ambiguous matches at the last look) gets one
|
|
3088
3126
|
// more evaluation now that the match set changed: a new project whose
|
|
3089
3127
|
// repo descends from the legacy commit inherits the attribution instead
|
|
@@ -3142,6 +3180,15 @@ commands["update-project"] = (db, args) => {
|
|
|
3142
3180
|
if (args.environment_type !== undefined) {
|
|
3143
3181
|
patch.environment_type = validateEnvironmentType(args.environment_type);
|
|
3144
3182
|
}
|
|
3183
|
+
// versioned_build / version_file are NOT covered by the context-change guard:
|
|
3184
|
+
// like environment_type, they are read once from launch args at dispatch
|
|
3185
|
+
// time, so a mid-run change only affects future launches.
|
|
3186
|
+
if (args.versioned_build !== undefined) {
|
|
3187
|
+
patch.versioned_build = validateVersionedBuild(args.versioned_build) ?? 0;
|
|
3188
|
+
}
|
|
3189
|
+
if (args.version_file !== undefined) {
|
|
3190
|
+
patch.version_file = validateVersionFile(args.version_file);
|
|
3191
|
+
}
|
|
3145
3192
|
|
|
3146
3193
|
const contextChanged =
|
|
3147
3194
|
(args.repo_path !== undefined && args.repo_path !== current.repo_path) ||
|
package/lib/schema.sql
CHANGED
|
@@ -46,6 +46,13 @@ CREATE TABLE IF NOT EXISTS projects (
|
|
|
46
46
|
provenance_crew_release TEXT,
|
|
47
47
|
provenance_published_at TEXT,
|
|
48
48
|
provenance_task_id TEXT,
|
|
49
|
+
-- versioned_build: the QA staleness gate (emojimanegg1/muse-crew#3). When 1,
|
|
50
|
+
-- Integrate bumps the version file and QA verifies the tested bundle is at
|
|
51
|
+
-- or past the required version. version_file is the repo-relative path to
|
|
52
|
+
-- the TS file exporting BUILD_NUMBER. Both nullable/0 by design: existing
|
|
53
|
+
-- rows stay unversioned (no backfill).
|
|
54
|
+
versioned_build INTEGER NOT NULL DEFAULT 0 CHECK (versioned_build IN (0, 1)),
|
|
55
|
+
version_file TEXT,
|
|
49
56
|
created_at TEXT NOT NULL,
|
|
50
57
|
updated_at TEXT NOT NULL
|
|
51
58
|
);
|
package/package.json
CHANGED
package/workflows/bugfix.js
CHANGED
|
@@ -563,6 +563,18 @@ function decideIntegrateRetry(o) {
|
|
|
563
563
|
if (rec.malformed) return "park";
|
|
564
564
|
return ancestor ? "skip-to-publish" : "proceed";
|
|
565
565
|
}
|
|
566
|
+
// QA version gate (issue #3): S=footer build, R=required. S>=R proves the
|
|
567
|
+
// bundle is at/past the merge (single-branch: version order = containment).
|
|
568
|
+
// Pure — pinned by tests/version-gate.test.js. Null = unreadable/missing.
|
|
569
|
+
function decideVersionGate(S, R) {
|
|
570
|
+
if (!(R > 0)) return { proceed: false, reason: "qa-version-missing",
|
|
571
|
+
remedy: "no build version recorded — re-run Integrate or record dashboard-version." };
|
|
572
|
+
if (!(S > 0)) return { proceed: false, reason: "qa-version-unreadable",
|
|
573
|
+
remedy: "no readable build number in bundle footer — rebuild from versioned source." };
|
|
574
|
+
if (S >= R) return { proceed: true };
|
|
575
|
+
return { proceed: false, reason: "qa-bundle-stale",
|
|
576
|
+
remedy: "bundle build " + S + " behind required " + R + " — builder did not rebuild from current source." };
|
|
577
|
+
}
|
|
566
578
|
// On a dispatcher retry resumed at Integrate, the run-local releaseDecision
|
|
567
579
|
// is null (Build doesn't re-run). Hydrate it from the merge record so the
|
|
568
580
|
// npm Publish gate doesn't park a publishable merge. Pure — pinned by
|
|
@@ -1355,10 +1367,26 @@ while (i < STEPS.length) {
|
|
|
1355
1367
|
// verdicts.jsonl). No parent verdict gate remains.
|
|
1356
1368
|
var qaArtifact = false;
|
|
1357
1369
|
var qaTerminal = false;
|
|
1370
|
+
var qaRequiredVersion = null;
|
|
1358
1371
|
if (step.name === "QA") {
|
|
1359
1372
|
var qaExp = (await resolveExperiential()) === "yes";
|
|
1360
1373
|
qaArtifact = qaExp && SURFACE_ARTIFACT;
|
|
1361
1374
|
qaTerminal = qaExp && SURFACE_TERMINAL;
|
|
1375
|
+
if (qaArtifact && projectConfig.versioned_build === true) {
|
|
1376
|
+
var vfile = projectConfig.version_file || "client/src/buildNumber.ts";
|
|
1377
|
+
var vgPre = await agent(
|
|
1378
|
+
"Return ONLY JSON. 1. Run, return stdout verbatim:\n" + crewCmd("get-events", { task_id: taskId }) + "\n" +
|
|
1379
|
+
"R = integer after `dashboard-version: ` in newest matching note. 2. If none, run, return stdout verbatim:\n" + crewCmd("get-provenance", { project_id: LAUNCH_PROJECT_ID }) + "\n" +
|
|
1380
|
+
"R = BUILD_NUMBER from: cd " + REPO_PATH + " && git show <source_commit>:" + vfile + " | grep -o 'BUILD_NUMBER = [0-9]*'. If unknown: {\"ok\":false,\"park\":\"qa-version-missing: no build version recorded — re-run Integrate or record dashboard-version.\"}. " +
|
|
1381
|
+
"3. cd " + REPO_PATH + " && git fetch origin && V=$(grep -o 'BUILD_NUMBER = [0-9]*' " + vfile + " | grep -o '[0-9]*'|head -1); if [ $V -lt R ]; then git pull --ff-only; V=$(grep -o 'BUILD_NUMBER = [0-9]*' " + vfile + " | grep -o '[0-9]*'|head -1); fi; " +
|
|
1382
|
+
"if [ $V -lt R ]: {\"ok\":false,\"park\":\"qa-bundle-stale: checkout \" + V + \" < required \" + R + \" — sync past the merge, re-run QA.\"}. 4. Else {\"ok\":true,\"R\":R}.",
|
|
1383
|
+
{ key: "qa-version-pre-" + taskId, label: "QA version-gate pre-check",
|
|
1384
|
+
schema: { type: "object", required: ["ok"],
|
|
1385
|
+
properties: { ok: { type: "boolean" }, R: { type: "integer" }, park: { type: "string" } } } }
|
|
1386
|
+
);
|
|
1387
|
+
if (!vgPre.ok) return await parkTask(vgPre.park);
|
|
1388
|
+
qaRequiredVersion = vgPre.R;
|
|
1389
|
+
}
|
|
1362
1390
|
}
|
|
1363
1391
|
|
|
1364
1392
|
// Reproduce-layer routing (2026-09-16): Sage's Triage classifies the bug's
|
|
@@ -1687,6 +1715,14 @@ while (i < STEPS.length) {
|
|
|
1687
1715
|
"R5-PUSH (manual R5 resolution only — the normal path pushed inline). Run: " + LIFECYCLE_ENV + LIFECYCLE + " push-target " + taskId + "\n" +
|
|
1688
1716
|
"PUSHED — report the merged hash (detached prints PUSHED: origin/main (refspec HEAD:main)), VERDICT: PASS; PUSH_SKIPPED — no push attempted (no record + no lock); NO_REMOTE_PUSH — no remote; VERDICT: PASS; ERROR or CONFLICT — report it, VERDICT: FAIL.\n" +
|
|
1689
1717
|
"NEVER force-push.\n\n" +
|
|
1718
|
+
(projectConfig.versioned_build === true ?
|
|
1719
|
+
"VERSION BUMP (issue #3): after MERGED+PUSHED, bump so QA can prove the bundle contains this fix — QA parks without it.\n" +
|
|
1720
|
+
"1. cd " + REPO_PATH + " && git fetch origin; BR=<branch-from-integration-target>; F=" + (projectConfig.version_file || "client/src/buildNumber.ts") + "\n" +
|
|
1721
|
+
"2. N=$(git show origin/$BR:$F | grep -o 'BUILD_NUMBER = [0-9]*' | grep -o '[0-9]*'); if missing/unparsable: VERDICT: FAIL.\n" +
|
|
1722
|
+
"3. Edit $F: `export const BUILD_NUMBER = $N;` → `export const BUILD_NUMBER = $((N+1));` (keep header comment).\n" +
|
|
1723
|
+
"4. git add $F && git commit -m \"build-number: $((N+1)) - QA version gate\" (SEPARATE commit, never amend).\n" +
|
|
1724
|
+
"5. git push origin $BR; if rejected retry 3x (fetch, re-read N, re-bump, re-commit, push). NEVER force-push. Push MUST succeed or VERDICT: FAIL.\n" +
|
|
1725
|
+
"6. Run, return stdout verbatim:\n" + crewCmd("log-event", { task_id: taskId, type: "note", identity: step.identity, message: "dashboard-version: <new> — QA must test a bundle built from source at or after the merge that recorded this (build <new> or later)." }) + "\n(substitute <new>).\n\n" : "") +
|
|
1690
1726
|
"Report what happened at each step, ending with exactly one line: VERDICT: PASS or VERDICT: FAIL.";
|
|
1691
1727
|
|
|
1692
1728
|
} else if (step.name === "Publish") {
|
|
@@ -2140,6 +2176,8 @@ while (i < STEPS.length) {
|
|
|
2140
2176
|
"c3. Start with: SEE_ACT_ARCHIVE_DIR=" + crewHome + "/task-evidence/" + taskId + "/postchange/ node " + crewHome + "/current/lib/see-act.js --url http://localhost:<N>/ aria — read the JSON, log the step. Then: SEE_ACT_ARCHIVE_DIR=" + crewHome + "/task-evidence/" + taskId + "/postchange/ node " + crewHome + "/current/lib/see-act.js --url http://localhost:<N>/ shot — READ the screenshot, log the step. Act on what you see: click, scroll, type, then re-observe, logging each step. Prefer aria (cheap text) to find controls; screenshot when the view changes and for your final verdict frames (one desktop, one mobile). If a click exits non-zero, do NOT retry the same ref blindly: re-run aria first (refs go stale between invocations), then click the fresh ref exactly once. If it still fails, log the failure and move on — a flaky control is a finding, not a loop.\n" +
|
|
2141
2177
|
"d. Reach: with the session protocol, any flow reachable by N in-page actions is drivable — open the dialog, then confirm it, then judge the result. Without a session (one-shot invocations), anything reachable by (navigate, one action) is testable and sequences needing prior in-page state are not — use a session for those. Report NOT POSSIBLE only when the tooling itself fails (session-start exits 3): a flow you could not reach is not NOT POSSIBLE — name the exact step that stopped you in verdict.json's missing evidence and continue with the mechanical checks.\n" +
|
|
2142
2178
|
"e. Judge as a user against the task description: is the reported bug fixed AND is nothing else visibly broken? Look for broken layout, overflow, missing or wrong content, stale data, and console errors. Compare against the task's expected behavior, never against source code (you are code-blind). Every frame you captured is already archived under " + crewHome + "/task-evidence/" + taskId + "/postchange/ and indexed in ooda-log.jsonl. A frame you did not read is not evidence. Loading, error, or blank frames never pass. If you cannot complete the loop, say exactly which steps are missing — unknown is not PASS.\n" +
|
|
2179
|
+
(projectConfig.versioned_build === true ?
|
|
2180
|
+
"e2. VERSION (issue #3): footer shows `build <n>` — report `footer_build: <n>` on its own line, or `footer_build: unreadable`. Required — the gate cannot pass without it.\n" : "") +
|
|
2143
2181
|
"f. Kill ONLY the server you started: pkill -f 'serve-artifact[.]js.*--tag " + taskId + "-qa' — never another task's server. (The [.] keeps pkill from matching its own command line.) Do not leave it running.\n" +
|
|
2144
2182
|
"Then continue with the mechanical checks below. Your VERDICT covers both the visual and the mechanical checks.\n\n" +
|
|
2145
2183
|
"MECHANICAL CHECKS:\n" +
|
|
@@ -2622,6 +2660,20 @@ while (i < STEPS.length) {
|
|
|
2622
2660
|
// reports FAIL, it stands — finding attribution informs follow-up
|
|
2623
2661
|
// filing only.
|
|
2624
2662
|
if (step.name === "QA") {
|
|
2663
|
+
if (projectConfig.versioned_build === true && qaArtifact && qaRequiredVersion !== null) {
|
|
2664
|
+
var vgM = /footer_build:\s*(\d+|unreadable)/i.exec(workerText || "");
|
|
2665
|
+
var vgS = (vgM && /^\d+$/.test(vgM[1])) ? parseInt(vgM[1], 10) : null;
|
|
2666
|
+
var vgD = decideVersionGate(vgS, qaRequiredVersion);
|
|
2667
|
+
if (!vgD.proceed) {
|
|
2668
|
+
return await parkTask(vgD.reason + ": " + vgD.remedy + " [built=" + vgS + " required=" + qaRequiredVersion + "]");
|
|
2669
|
+
}
|
|
2670
|
+
await agent(
|
|
2671
|
+
"Run in shell and return the stdout verbatim:\n" + crewCmd("log-event", {
|
|
2672
|
+
task_id: taskId, type: "note", identity: step.identity,
|
|
2673
|
+
message: "version-check: built=" + vgS + " >= required=" + qaRequiredVersion + " → testing now" }),
|
|
2674
|
+
{ key: "record-version-check-" + taskId, label: "Recording version-gate pass" }
|
|
2675
|
+
);
|
|
2676
|
+
}
|
|
2625
2677
|
var contentFindings = extractContentFindings(workerText);
|
|
2626
2678
|
if (!contentFindings.ok) {
|
|
2627
2679
|
// Unknown attribution fails closed: a missing or malformed
|
|
@@ -476,7 +476,13 @@ for (var pi = 0; pi < projects.length; pi++) {
|
|
|
476
476
|
// The user-facing surface for experiential QA routing (artifact |
|
|
477
477
|
// terminal | null=unclassified). Carried alongside deploy_type — it is a
|
|
478
478
|
// separate axis, not a redeclaration of the deployment target.
|
|
479
|
-
environment_type: proj.environment_type || null
|
|
479
|
+
environment_type: proj.environment_type || null,
|
|
480
|
+
// QA staleness gate (emojimanegg1/muse-crew#3): versioned_build opts the
|
|
481
|
+
// project into the Integrate version bump + QA bundle-freshness gate;
|
|
482
|
+
// version_file is the repo-relative path of the TS file exporting
|
|
483
|
+
// BUILD_NUMBER (null when unversioned).
|
|
484
|
+
versioned_build: !!proj.versioned_build,
|
|
485
|
+
version_file: proj.version_file || null
|
|
480
486
|
};
|
|
481
487
|
}
|
|
482
488
|
|
|
@@ -561,6 +567,12 @@ const eligible = [];
|
|
|
561
567
|
const retryCandidates = [];
|
|
562
568
|
const rejectionCandidates = [];
|
|
563
569
|
|
|
570
|
+
// Dispatch-decision log (2026-09-24): per-reason ineligibility counters for
|
|
571
|
+
// the .dispatch-decisions.jsonl declaration. Sparse — only non-zero keys are
|
|
572
|
+
// emitted. Additive only; the eligibility decisions below are unchanged.
|
|
573
|
+
var decisionSkipped = {};
|
|
574
|
+
function countSkipped(reason) { decisionSkipped[reason] = (decisionSkipped[reason] || 0) + 1; }
|
|
575
|
+
|
|
564
576
|
// ── Retry cap ────────────────────────────────────────────────────────
|
|
565
577
|
// Symphony owns phase-redispatch policy (coordination layer): the spec
|
|
566
578
|
// defines backoff but no attempt cap ("implementation-defined"), so the
|
|
@@ -591,12 +603,13 @@ var MAX_CONSECUTIVE_REJECTIONS = configInt(config, "maxConsecutiveRejections", 2
|
|
|
591
603
|
|
|
592
604
|
for (var t = 0; t < allTasks.length; t++) {
|
|
593
605
|
var task = allTasks[t];
|
|
594
|
-
if (task.blocked) continue;
|
|
606
|
+
if (task.blocked) { countSkipped("blocked"); continue; }
|
|
595
607
|
|
|
596
608
|
// Skip tasks with an active dispatch reservation: the worker launched a
|
|
597
609
|
// workflow for this task on a previous tick, but the workflow has not yet
|
|
598
610
|
// self-claimed (claims can take 15+ minutes for cron-launched runs).
|
|
599
611
|
if (reservedTaskIds.has(task.id)) {
|
|
612
|
+
countSkipped("reserved");
|
|
600
613
|
log("Skipped \"" + task.title + "\" — active dispatch reservation (workflow launched, claim pending)");
|
|
601
614
|
continue;
|
|
602
615
|
}
|
|
@@ -604,6 +617,7 @@ for (var t = 0; t < allTasks.length; t++) {
|
|
|
604
617
|
// Skip tasks from quiesced projects
|
|
605
618
|
var taskProject = task.project || DEFAULT_PROJECT;
|
|
606
619
|
if (quiescedProjects[taskProject]) {
|
|
620
|
+
countSkipped("quiesced");
|
|
607
621
|
log("Skipped \"" + task.title + "\" — project " + taskProject + " is quiesced");
|
|
608
622
|
continue;
|
|
609
623
|
}
|
|
@@ -614,6 +628,7 @@ for (var t = 0; t < allTasks.length; t++) {
|
|
|
614
628
|
// task never moves. Set repo_path via updateproject to re-enable.
|
|
615
629
|
var taskProjCfg = PROJECTS[taskProject];
|
|
616
630
|
if (!taskProjCfg || !taskProjCfg.repo_path) {
|
|
631
|
+
countSkipped("no_repo");
|
|
617
632
|
log("Skipped \"" + task.title + "\" — project " + taskProject + " has no repo_path configured");
|
|
618
633
|
continue;
|
|
619
634
|
}
|
|
@@ -633,12 +648,14 @@ for (var t = 0; t < allTasks.length; t++) {
|
|
|
633
648
|
// (clears) next_phase on its successful self-claim, exactly once.
|
|
634
649
|
if (task.next_phase) {
|
|
635
650
|
if (task.state !== "todo" && task.state !== "in_progress") {
|
|
651
|
+
countSkipped("next_phase_state");
|
|
636
652
|
log("Skipped \"" + task.title + "\" — next_phase \"" + task.next_phase + "\" set but state is " + task.state + "; left set for inspection");
|
|
637
653
|
continue;
|
|
638
654
|
}
|
|
639
|
-
if (latest && latest.status === "running") continue; // work in flight
|
|
655
|
+
if (latest && latest.status === "running") { countSkipped("in_flight"); continue; } // work in flight
|
|
640
656
|
var npIdx = steps.indexOf(task.next_phase);
|
|
641
657
|
if (npIdx < 0) {
|
|
658
|
+
countSkipped("next_phase_bad_step");
|
|
642
659
|
log("Skipped \"" + task.title + "\" — next_phase \"" + task.next_phase + "\" not in " + workflow + " step registry; left set for a corrected recover-task");
|
|
643
660
|
continue;
|
|
644
661
|
}
|
|
@@ -654,19 +671,20 @@ for (var t = 0; t < allTasks.length; t++) {
|
|
|
654
671
|
continue;
|
|
655
672
|
}
|
|
656
673
|
|
|
657
|
-
if (task.state !== "in_progress") continue;
|
|
674
|
+
if (task.state !== "in_progress") { countSkipped("terminal_state"); continue; }
|
|
658
675
|
|
|
659
676
|
if (!latest) {
|
|
660
677
|
eligible.push({ task: task, startStep: 0, reason: "no_session", workflow: workflow });
|
|
661
678
|
continue;
|
|
662
679
|
}
|
|
663
680
|
|
|
664
|
-
if (latest.status === "running") continue; // work in flight
|
|
681
|
+
if (latest.status === "running") { countSkipped("in_flight"); continue; } // work in flight
|
|
665
682
|
|
|
666
683
|
if (latest.status === "completed") {
|
|
667
684
|
var stepName = latest.step || "";
|
|
668
685
|
var stepIndex = steps.indexOf(stepName);
|
|
669
686
|
if (stepIndex < 0) {
|
|
687
|
+
countSkipped("unknown_step");
|
|
670
688
|
log("Skipped \"" + task.title + "\" — step \"" + stepName + "\" not in " + workflow + " step registry");
|
|
671
689
|
continue;
|
|
672
690
|
}
|
|
@@ -695,6 +713,7 @@ for (var t = 0; t < allTasks.length; t++) {
|
|
|
695
713
|
var failedStep = latest.step || "";
|
|
696
714
|
var retryIdx = steps.indexOf(failedStep);
|
|
697
715
|
if (retryIdx < 0) {
|
|
716
|
+
countSkipped("unknown_step");
|
|
698
717
|
log("Skipped \"" + task.title + "\" — step \"" + failedStep + "\" not in " + workflow + " step registry");
|
|
699
718
|
continue;
|
|
700
719
|
}
|
|
@@ -756,6 +775,7 @@ if (rejectionCandidates.length > 0) {
|
|
|
756
775
|
// the task stays retryable and is re-parked next tick, so this is
|
|
757
776
|
// fail-safe. The dashboard stamps retry_reset_at on parked→todo, which
|
|
758
777
|
// restarts both counters mechanically.
|
|
778
|
+
var parkedTaskIds = [];
|
|
759
779
|
if (parkJobs.length > 0) {
|
|
760
780
|
var parkSteps = [];
|
|
761
781
|
for (var pji = 0; pji < parkJobs.length; pji++) {
|
|
@@ -775,6 +795,7 @@ if (parkJobs.length > 0) {
|
|
|
775
795
|
{ key: "park-batch", label: "Parking " + parkJobs.length + " task(s) at retry cap" }
|
|
776
796
|
);
|
|
777
797
|
log("Parked " + parkJobs.length + " task(s)");
|
|
798
|
+
parkedTaskIds = parkJobs.map(function(pj) { return pj.task.id; });
|
|
778
799
|
} catch (parkErr) {
|
|
779
800
|
log("WARNING: park batch failed (" + (parkErr.message || String(parkErr)).slice(0, 200) + ") — tasks remain retryable");
|
|
780
801
|
}
|
|
@@ -959,6 +980,7 @@ for (var ei = 0; ei < eligible.length; ei++) {
|
|
|
959
980
|
toProcess.push(eitem);
|
|
960
981
|
inFlightByProject[ep] = current + 1;
|
|
961
982
|
} else {
|
|
983
|
+
countSkipped("at_limit");
|
|
962
984
|
log("Skipped \"" + eitem.task.title + "\" — project " + ep + " at simultaneity limit (" + limit + ")");
|
|
963
985
|
}
|
|
964
986
|
}
|
|
@@ -1164,6 +1186,63 @@ if (recommended.length > 0) {
|
|
|
1164
1186
|
recommended = acquired;
|
|
1165
1187
|
}
|
|
1166
1188
|
|
|
1189
|
+
// ── 4c. Dispatch-decision log ──────────────────────────────────
|
|
1190
|
+
// Hoverboat (2026-09-24): the dispatcher declares its own decision, in its
|
|
1191
|
+
// own bytes, on crew-home disk — one append-only line per completed tick in
|
|
1192
|
+
// $CREW_HOME/.dispatch-decisions.jsonl. The observer reads the file
|
|
1193
|
+
// directly (no LLM re-emission, no platform ferry), so the plausible-value
|
|
1194
|
+
// mutation that poisoned the old dispatch_sample attribution cannot occur.
|
|
1195
|
+
// Launch declarations are verified mechanically against crew DB effects
|
|
1196
|
+
// (platform_run_tasks / dispatch_reservations); a stand-down is the
|
|
1197
|
+
// dispatcher's explicit declaration with board context, auditable without
|
|
1198
|
+
// re-deriving eligibility. tick_seq couples each line to the .tick-releases
|
|
1199
|
+
// line for the same poll — a tick-release line with no decision line means
|
|
1200
|
+
// the tick died after poll-ack and is visible as such.
|
|
1201
|
+
// seq is the non-empty line count + 1 of this file (file order is the
|
|
1202
|
+
// proof — no wall-clock calls, see tests/determinism.test.js).
|
|
1203
|
+
// Fire-and-forget: the script swallows its own errors and exits 0, the
|
|
1204
|
+
// agent call carries no schema, and the whole call is wrapped in
|
|
1205
|
+
// try/catch — a failed write can never fail the tick. The script echoes
|
|
1206
|
+
// the appended line (same rooms #12–#14 empty-result contract as §4b).
|
|
1207
|
+
var dispatchDecisionScript = [
|
|
1208
|
+
"var crewHome=process.argv[1];",
|
|
1209
|
+
"var decisionJson=process.argv[2];",
|
|
1210
|
+
"try{",
|
|
1211
|
+
"var fs=require(\"fs\"),path=require(\"path\");",
|
|
1212
|
+
"var decision=JSON.parse(decisionJson);",
|
|
1213
|
+
"if(!decision||typeof decision.decision!==\"string\"||!Array.isArray(decision.launched)){throw new Error(\"bad decision payload\");}",
|
|
1214
|
+
"var release=path.basename(path.resolve(crewHome,fs.readlinkSync(path.join(crewHome,\"current\"))));",
|
|
1215
|
+
"var file=path.join(crewHome,\".dispatch-decisions.jsonl\");",
|
|
1216
|
+
"var count=0;",
|
|
1217
|
+
"try{var lines=fs.readFileSync(file,\"utf8\").split(\"\\n\");for(var i=0;i<lines.length;i++){if(lines[i].trim()!==\"\"){count++;}}}catch(e){}",
|
|
1218
|
+
"var tickSeq=0;",
|
|
1219
|
+
"try{var tlines=fs.readFileSync(path.join(crewHome,\".tick-releases.jsonl\"),\"utf8\").split(\"\\n\");for(var j=0;j<tlines.length;j++){if(tlines[j].trim()!==\"\"){tickSeq++;}}}catch(e){}",
|
|
1220
|
+
"var line=JSON.stringify({seq:count+1,tick_seq:tickSeq,release:release,decision:decision.decision,launched:decision.launched,completed:decision.completed||[],parked:decision.parked||[],board:decision.board||null,skipped:decision.skipped||{},partial:!!decision.partial});",
|
|
1221
|
+
"fs.appendFileSync(file,line+\"\\n\");",
|
|
1222
|
+
"process.stdout.write(line+\"\\n\");",
|
|
1223
|
+
"}catch(e){}",
|
|
1224
|
+
"process.exit(0);"
|
|
1225
|
+
].join("");
|
|
1226
|
+
var launchedRecs = recommended.map(function(r) { return { task_id: r.task_id, workflow: r.workflow, step: r.step }; });
|
|
1227
|
+
var decisionObj = {
|
|
1228
|
+
decision: launchedRecs.length > 0 ? "launch" : "stand_down",
|
|
1229
|
+
launched: launchedRecs,
|
|
1230
|
+
completed: completed.map(function(r) { return r.task_id; }),
|
|
1231
|
+
parked: parkedTaskIds,
|
|
1232
|
+
board: { seen: allTasks.length, eligible: eligible.length },
|
|
1233
|
+
skipped: decisionSkipped,
|
|
1234
|
+
partial: partial
|
|
1235
|
+
};
|
|
1236
|
+
var dispatchDecisionCmd = "node -e '" + dispatchDecisionScript + "' '" + crewHome.replace(/'/g, "'\\''") + "' '" + JSON.stringify(decisionObj).replace(/'/g, "'\\''") + "'";
|
|
1237
|
+
try {
|
|
1238
|
+
await agent(
|
|
1239
|
+
"Record this tick's dispatch decision.\nRun in shell and return the stdout verbatim:\n" + dispatchDecisionCmd,
|
|
1240
|
+
{ key: "dispatch-decision", label: "Recording dispatch decision" }
|
|
1241
|
+
);
|
|
1242
|
+
} catch (e) {
|
|
1243
|
+
log("WARNING: dispatch-decision log write failed (tick continues): " + e.message);
|
|
1244
|
+
}
|
|
1245
|
+
|
|
1167
1246
|
var msg = "Dispatch complete.";
|
|
1168
1247
|
if (recommended.length > 0) {
|
|
1169
1248
|
msg += " Recommended: " + recommended.map(function(r) { return r.workflow + "/" + r.step + " for " + r.task_id; }).join(", ") + ".";
|
package/workflows/standard.js
CHANGED
|
@@ -571,6 +571,18 @@ function decideIntegrateRetry(o) {
|
|
|
571
571
|
if (rec.malformed) return "park";
|
|
572
572
|
return ancestor ? "skip-to-publish" : "proceed";
|
|
573
573
|
}
|
|
574
|
+
// QA version gate (issue #3): S=footer build, R=required. S>=R proves the
|
|
575
|
+
// bundle is at/past the merge (single-branch: version order = containment).
|
|
576
|
+
// Pure — pinned by tests/version-gate.test.js. Null = unreadable/missing.
|
|
577
|
+
function decideVersionGate(S, R) {
|
|
578
|
+
if (!(R > 0)) return { proceed: false, reason: "qa-version-missing",
|
|
579
|
+
remedy: "no build version recorded — re-run Integrate or record dashboard-version." };
|
|
580
|
+
if (!(S > 0)) return { proceed: false, reason: "qa-version-unreadable",
|
|
581
|
+
remedy: "no readable build number in bundle footer — rebuild from versioned source." };
|
|
582
|
+
if (S >= R) return { proceed: true };
|
|
583
|
+
return { proceed: false, reason: "qa-bundle-stale",
|
|
584
|
+
remedy: "bundle build " + S + " behind required " + R + " — builder did not rebuild from current source." };
|
|
585
|
+
}
|
|
574
586
|
// On a dispatcher retry resumed at Integrate, the run-local releaseDecision
|
|
575
587
|
// is null (Build doesn't re-run). Hydrate it from the merge record so the
|
|
576
588
|
// npm Publish gate doesn't park a publishable merge. Pure — pinned by
|
|
@@ -1300,6 +1312,22 @@ while (i < STEPS.length) {
|
|
|
1300
1312
|
var qaExperiential = false;
|
|
1301
1313
|
if (step.name === "QA") {
|
|
1302
1314
|
qaExperiential = (await resolveExperiential()) === "yes" && SURFACE_CLASSIFIED;
|
|
1315
|
+
var qaRequiredVersion = null;
|
|
1316
|
+
if (qaExperiential && SURFACE_ARTIFACT && projectConfig.versioned_build === true) {
|
|
1317
|
+
var vfileStd = projectConfig.version_file || "client/src/buildNumber.ts";
|
|
1318
|
+
var vgPreStd = await agent(
|
|
1319
|
+
"Return ONLY JSON. 1. Run, return stdout verbatim:\n" + crewCmd("get-events", { task_id: taskId }) + "\n" +
|
|
1320
|
+
"R = integer after `dashboard-version: ` in newest matching note. 2. If none, run, return stdout verbatim:\n" + crewCmd("get-provenance", { project_id: LAUNCH_PROJECT_ID }) + "\n" +
|
|
1321
|
+
"R = BUILD_NUMBER from: cd " + REPO_PATH + " && git show <source_commit>:" + vfileStd + " | grep -o 'BUILD_NUMBER = [0-9]*'. If unknown: {\"ok\":false,\"park\":\"qa-version-missing: no build version recorded — re-run Integrate or record dashboard-version.\"}. " +
|
|
1322
|
+
"3. cd " + REPO_PATH + " && git fetch origin && V=$(grep -o 'BUILD_NUMBER = [0-9]*' " + vfileStd + " | grep -o '[0-9]*'|head -1); if [ $V -lt R ]; then git pull --ff-only; V=$(grep -o 'BUILD_NUMBER = [0-9]*' " + vfileStd + " | grep -o '[0-9]*'|head -1); fi; " +
|
|
1323
|
+
"if [ $V -lt R ]: {\"ok\":false,\"park\":\"qa-bundle-stale: checkout \" + V + \" < required \" + R + \" — sync past the merge, re-run QA.\"}. 4. Else {\"ok\":true,\"R\":R}.",
|
|
1324
|
+
{ key: "qa-version-pre-" + taskId, label: "QA version-gate pre-check",
|
|
1325
|
+
schema: { type: "object", required: ["ok"],
|
|
1326
|
+
properties: { ok: { type: "boolean" }, R: { type: "integer" }, park: { type: "string" } } } }
|
|
1327
|
+
);
|
|
1328
|
+
if (!vgPreStd.ok) return await parkTask(vgPreStd.park);
|
|
1329
|
+
qaRequiredVersion = vgPreStd.R;
|
|
1330
|
+
}
|
|
1303
1331
|
}
|
|
1304
1332
|
|
|
1305
1333
|
// Step-specific instructions
|
|
@@ -1576,6 +1604,14 @@ while (i < STEPS.length) {
|
|
|
1576
1604
|
"R5-PUSH (manual R5 resolution only — the normal path pushed inline). Run: " + LIFECYCLE_ENV + LIFECYCLE + " push-target " + taskId + "\n" +
|
|
1577
1605
|
"PUSHED — report the merged hash (detached prints PUSHED: origin/main (refspec HEAD:main)), VERDICT: PASS; PUSH_SKIPPED — no push attempted (no record + no lock); NO_REMOTE_PUSH — no remote; VERDICT: PASS; ERROR or CONFLICT — report it, VERDICT: FAIL.\n" +
|
|
1578
1606
|
"NEVER force-push.\n\n" +
|
|
1607
|
+
(projectConfig.versioned_build === true ?
|
|
1608
|
+
"VERSION BUMP (issue #3): after MERGED+PUSHED, bump so QA can prove the bundle contains this fix — QA parks without it.\n" +
|
|
1609
|
+
"1. cd " + REPO_PATH + " && git fetch origin; BR=<branch-from-integration-target>; F=" + (projectConfig.version_file || "client/src/buildNumber.ts") + "\n" +
|
|
1610
|
+
"2. N=$(git show origin/$BR:$F | grep -o 'BUILD_NUMBER = [0-9]*' | grep -o '[0-9]*'); if missing/unparsable: VERDICT: FAIL.\n" +
|
|
1611
|
+
"3. Edit $F: `export const BUILD_NUMBER = $N;` → `export const BUILD_NUMBER = $((N+1));` (keep header comment).\n" +
|
|
1612
|
+
"4. git add $F && git commit -m \"build-number: $((N+1)) - QA version gate\" (SEPARATE commit, never amend).\n" +
|
|
1613
|
+
"5. git push origin $BR; if rejected retry 3x (fetch, re-read N, re-bump, re-commit, push). NEVER force-push. Push MUST succeed or VERDICT: FAIL.\n" +
|
|
1614
|
+
"6. Run, return stdout verbatim:\n" + crewCmd("log-event", { task_id: taskId, type: "note", identity: step.identity, message: "dashboard-version: <new> — QA must test a bundle built from source at or after the merge that recorded this (build <new> or later)." }) + "\n(substitute <new>).\n\n" : "") +
|
|
1579
1615
|
"Report what happened at each step, ending with exactly one line: VERDICT: PASS or VERDICT: FAIL.";
|
|
1580
1616
|
|
|
1581
1617
|
} else if (step.name === "Publish") {
|
|
@@ -2029,6 +2065,8 @@ while (i < STEPS.length) {
|
|
|
2029
2065
|
"c3. Start with: SEE_ACT_ARCHIVE_DIR=" + crewHome + "/task-evidence/" + taskId + "/postchange/ node " + crewHome + "/current/lib/see-act.js --url http://localhost:<N>/ aria — read the JSON, log the step. Then: SEE_ACT_ARCHIVE_DIR=" + crewHome + "/task-evidence/" + taskId + "/postchange/ node " + crewHome + "/current/lib/see-act.js --url http://localhost:<N>/ shot — READ the screenshot, log the step. Act on what you see: click, scroll, type, then re-observe, logging each step. Prefer aria (cheap text) to find controls; screenshot when the view changes and for your final verdict frames (one desktop, one mobile). If a click exits non-zero, do NOT retry the same ref blindly: re-run aria first (refs go stale between invocations), then click the fresh ref exactly once. If it still fails, log the failure and move on — a flaky control is a finding, not a loop.\n" +
|
|
2030
2066
|
"d. Reach: with the session protocol, any flow reachable by N in-page actions is drivable — open the deck, then click Study, then judge the study view. Without a session (one-shot invocations), anything reachable by (navigate, one action) is testable and sequences needing prior in-page state are not — use a session for those. Report NOT POSSIBLE only when the tooling itself fails (session-start exits 3): a flow you could not reach is not NOT POSSIBLE — name the exact step that stopped you in verdict.json's missing evidence and judge what you did reach.\n" +
|
|
2031
2067
|
"e. Judge as a user against the task description: does the change render correctly? Look for broken layout, overflow, missing or wrong content, stale data, and console errors. Compare against the task's expected behavior, never against source code (you are code-blind). Every frame you captured is already archived under " + crewHome + "/task-evidence/" + taskId + "/postchange/ and indexed in ooda-log.jsonl. A frame you did not read is not evidence. Loading, error, or blank frames never pass. If you cannot complete the loop, say exactly which steps are missing — unknown is not PASS.\n" +
|
|
2068
|
+
(projectConfig.versioned_build === true ?
|
|
2069
|
+
"e2. VERSION (issue #3): footer shows `build <n>` — report `footer_build: <n>` on its own line, or `footer_build: unreadable`. Required — the gate cannot pass without it.\n" : "") +
|
|
2032
2070
|
"f. Kill ONLY the server you started: pkill -f 'serve-artifact[.]js.*--tag " + taskId + "-qa' — never another task's server. (The [.] keeps pkill from matching its own command line.) Do not leave it running.\n" +
|
|
2033
2071
|
"Then continue with the mechanical checks below. Your VERDICT covers both the visual and the mechanical checks.\n\n" +
|
|
2034
2072
|
"STEP 2: Verify data integrity via the crew API.\n" +
|
|
@@ -2374,6 +2412,20 @@ while (i < STEPS.length) {
|
|
|
2374
2412
|
// reports FAIL, it stands — finding attribution informs follow-up
|
|
2375
2413
|
// filing only.
|
|
2376
2414
|
if (step.name === "QA") {
|
|
2415
|
+
if (projectConfig.versioned_build === true && SURFACE_ARTIFACT && qaRequiredVersion !== null) {
|
|
2416
|
+
var vgM = /footer_build:\s*(\d+|unreadable)/i.exec(workerText || "");
|
|
2417
|
+
var vgS = (vgM && /^\d+$/.test(vgM[1])) ? parseInt(vgM[1], 10) : null;
|
|
2418
|
+
var vgD = decideVersionGate(vgS, qaRequiredVersion);
|
|
2419
|
+
if (!vgD.proceed) {
|
|
2420
|
+
return await parkTask(vgD.reason + ": " + vgD.remedy + " [built=" + vgS + " required=" + qaRequiredVersion + "]");
|
|
2421
|
+
}
|
|
2422
|
+
await agent(
|
|
2423
|
+
"Run in shell and return the stdout verbatim:\n" + crewCmd("log-event", {
|
|
2424
|
+
task_id: taskId, type: "note", identity: step.identity,
|
|
2425
|
+
message: "version-check: built=" + vgS + " >= required=" + qaRequiredVersion + " → testing now" }),
|
|
2426
|
+
{ key: "record-version-check-" + taskId, label: "Recording version-gate pass" }
|
|
2427
|
+
);
|
|
2428
|
+
}
|
|
2377
2429
|
var contentFindings = extractContentFindings(workerText);
|
|
2378
2430
|
if (!contentFindings.ok) {
|
|
2379
2431
|
// Unknown attribution fails closed: a missing or malformed
|