muse-crew 0.14.4 → 0.14.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +1 -1
- package/API.md +61 -7
- package/docs/decisions/AGENTS.md +3 -0
- package/docs/decisions/publish-path.md +244 -1
- package/docs/guide.md +42 -6
- package/docs/publish-unknown-recovery.md +28 -113
- package/docs/publish-verification.md +148 -445
- package/docs/release-integrity.md +5 -3
- package/docs/reviews/critic-0145.md +106 -0
- package/lib/AGENTS.md +4 -8
- package/lib/advance-publish-base.js +3 -3
- package/lib/compose-evidence-caption.js +1 -2
- package/lib/compute-publish-diff.js +113 -15
- package/lib/crew-api.js +855 -725
- package/lib/publish-note-vocabulary.js +137 -24
- package/lib/schema.sql +19 -0
- package/package.json +1 -1
- package/seed/AGENTS.md +1 -0
- package/seed/cron-body-ack-scan.md +44 -0
- package/seed/cron-body-template.md +65 -65
- package/seed/crons.json +12 -0
- package/workflows/bugfix.js +146 -740
- package/workflows/chore.js +146 -740
- package/workflows/crew-dispatch.js +36 -0
- package/workflows/crew-init.js +23 -0
- package/workflows/crew-uninstall.js +7 -4
- package/workflows/standard.js +146 -763
- package/lib/build-readback-request.js +0 -140
- package/lib/classify-publish-absence.js +0 -462
- package/lib/publish-content.js +0 -154
- package/lib/readback-disk.js +0 -195
- package/lib/retry-publish.js +0 -405
- package/lib/verify-publish.js +0 -414
|
@@ -1,44 +1,157 @@
|
|
|
1
|
-
//
|
|
2
|
-
//
|
|
3
|
-
// parked. Readers (scan-publish-unknown in lib/crew-api.js) consult this;
|
|
4
|
-
// writers assert against it (lib/verify-publish.js's terminal() fails loud
|
|
5
|
-
// before emitting a verb that is not in the registry). Add a note here when
|
|
6
|
-
// a new terminal verb is introduced — the scan recognizes it with no other
|
|
7
|
-
// change.
|
|
1
|
+
// Publish-note vocabulary registry (D7, 2026-09-19; 0.14.6 version-acknowledgement
|
|
2
|
+
// rewrite, 2026-09-20). Explicit state, not prose.
|
|
8
3
|
//
|
|
9
|
-
//
|
|
10
|
-
//
|
|
11
|
-
//
|
|
12
|
-
//
|
|
13
|
-
//
|
|
14
|
-
//
|
|
15
|
-
//
|
|
4
|
+
// 0.14.6 replaced the content-verdict/readback architecture with Eric's
|
|
5
|
+
// publication contract: exact version acknowledgement is the sole positive
|
|
6
|
+
// completion criterion ("If we hear that the artifact acknowledges our
|
|
7
|
+
// version, that's it. We don't verify against content."). Provenance now
|
|
8
|
+
// certifies that an issuance request was ACKNOWLEDGED — never that bytes
|
|
9
|
+
// matched.
|
|
10
|
+
//
|
|
11
|
+
// Two registries:
|
|
12
|
+
// TERMINAL_PUBLISH_NOTES — every `publish: …` note the ack/intent scans
|
|
13
|
+
// treat as terminally parked. Readers consult this; a recognized
|
|
14
|
+
// terminal note is skipped as terminal with its meaning, never as
|
|
15
|
+
// unrecognized. Legacy verbs from the retired content-verdict machine
|
|
16
|
+
// stay recognized (marked below) so old note history never becomes
|
|
17
|
+
// "unrecognized" — the new machine never writes them.
|
|
18
|
+
// TRANSITIONAL_PUBLISH_NOTES — every non-terminal state-carrying
|
|
19
|
+
// `publish: …` note the 0.14.6 machine writes.
|
|
20
|
+
//
|
|
21
|
+
// Writers are split in two (2026-09-20 REVIEW): deterministic code writes
|
|
22
|
+
// through writeGuardedPublishNote (asserts CODE_WRITABLE_PUBLISH_NOTES —
|
|
23
|
+
// transitional + the four current terminals); tick prose writes only
|
|
24
|
+
// through the `record-publish-note` crew-api command (asserts
|
|
25
|
+
// WRITABLE_PUBLISH_NOTES — transitional ONLY, so prose can never mint a
|
|
26
|
+
// terminal state). A third prose path, `log-event` with type "note", is
|
|
27
|
+
// mechanically rejected for any `publish:` verb — notes are advisory, but
|
|
28
|
+
// only the closed writer may mint state-carrying ones.
|
|
29
|
+
//
|
|
30
|
+
// The one gap: the workflow's `publish: publish-requested` park note is
|
|
31
|
+
// written by the workflow's own parkTask plumbing, not through crew-api —
|
|
32
|
+
// the intent scan validates its shape (commit + attempt) against the intent
|
|
33
|
+
// ledger entry when claiming, which is the mechanical guard for that path.
|
|
34
|
+
//
|
|
35
|
+
// Add a note here when a new verb is introduced — readers recognize it with
|
|
36
|
+
// no other change, and writers cannot emit it without registering it.
|
|
16
37
|
//
|
|
17
38
|
// ESM, no shebang, no side effects on import (import-safe module — bare
|
|
18
39
|
// `node publish-note-vocabulary.js` exits 0).
|
|
40
|
+
//
|
|
41
|
+
// The four 0.14.6 terminals the machine's own deterministic code writes
|
|
42
|
+
// (never tick prose — see the writer split above).
|
|
43
|
+
export const CODE_TERMINAL_PUBLISH_NOTES = [
|
|
44
|
+
// The artifact acknowledged the attempt's exact version on disk.
|
|
45
|
+
// Provenance is stamped; the task re-queues. (2026-09-20 REVIEW: disk
|
|
46
|
+
// is the sole positive evidence — the builder-report path was circular.)
|
|
47
|
+
"publish: version-acknowledged",
|
|
48
|
+
// The platform explicitly refused the tick worker's directly-issued edit.
|
|
49
|
+
// Honest terminal state — no re-issue; parked for human attention.
|
|
50
|
+
"publish: publish-refused",
|
|
51
|
+
// The 2-hour acknowledgement budget from first issuance is exhausted with
|
|
52
|
+
// no acknowledgement (unobserved within budget — the edit may still have
|
|
53
|
+
// landed). Parked for human attention.
|
|
54
|
+
"publish: version-timeout",
|
|
55
|
+
// One-party intent terminal: the diff for the committed intent cannot be
|
|
56
|
+
// (re)generated byte-identically — the staged file is missing or corrupt,
|
|
57
|
+
// or the base..commit range cannot be regenerated. Never re-claim, never
|
|
58
|
+
// loop; parked for human attention.
|
|
59
|
+
"publish: publish-unissuable",
|
|
60
|
+
];
|
|
19
61
|
export const TERMINAL_PUBLISH_NOTES = [
|
|
62
|
+
...CODE_TERMINAL_PUBLISH_NOTES,
|
|
63
|
+
// Legacy terminals (retired content-verdict/unknown-recovery machine):
|
|
64
|
+
// recognized so old note history skips as terminal, never written anew.
|
|
20
65
|
"publish: ambiguous",
|
|
21
66
|
"publish: retry-superseded",
|
|
22
67
|
"publish: retry-refused",
|
|
23
68
|
"publish: verified",
|
|
24
69
|
"publish: superseded",
|
|
25
|
-
// Verification pipeline's terminal content verdict (verify-publish.js).
|
|
26
|
-
// The task stays parked for human attention; unknown-recovery never
|
|
27
|
-
// re-enters it. Added 2026-09-19 (Room #23 J1/J3: was mislogged as
|
|
28
|
-
// unrecognized-publish-note).
|
|
29
70
|
"publish: verification-failed",
|
|
71
|
+
"publish: intent-unverifiable",
|
|
30
72
|
];
|
|
31
73
|
export const TERMINAL_NOTE_MEANINGS = {
|
|
32
|
-
"publish:
|
|
33
|
-
"publish:
|
|
34
|
-
"publish:
|
|
35
|
-
"publish:
|
|
36
|
-
"publish:
|
|
37
|
-
"publish:
|
|
74
|
+
"publish: version-acknowledged": "the artifact acknowledged this attempt's exact version — provenance stamped; task re-queued (terminal for the scans)",
|
|
75
|
+
"publish: publish-refused": "the platform refused the tick worker's directly-issued edit (or a builder report carried the version + an explicit refusal) — parked for human attention",
|
|
76
|
+
"publish: version-timeout": "the 2-hour acknowledgement budget from first issuance expired with no version acknowledgement (unobserved within budget — the edit may still have landed; the outcome is unknown) — parked for human attention",
|
|
77
|
+
"publish: publish-unissuable": "one-party intent terminal: the intent diff cannot be (re)generated byte-identically; parked for human attention",
|
|
78
|
+
"publish: ambiguous": "legacy (retired unknown-recovery): could not prove the drop; parked for human attention",
|
|
79
|
+
"publish: retry-superseded": "legacy (retired retry protocol): superseded; parked for human attention",
|
|
80
|
+
"publish: retry-refused": "legacy (retired retry protocol): platform refused the re-issued edit; parked for human attention",
|
|
81
|
+
"publish: verified": "legacy (retired content-verdict pipeline): publish verified and stamped; task re-queued",
|
|
82
|
+
"publish: superseded": "legacy (retired unknown-recovery): HEAD moved past the attempt; recovery aborted; parked for human attention",
|
|
83
|
+
"publish: verification-failed": "legacy (retired content-verdict pipeline): terminal content verdict — parked for human attention",
|
|
84
|
+
"publish: intent-unverifiable": "legacy (retired manifest-freshness path): re-claimed intent with no manifest baseline; parked for human attention",
|
|
38
85
|
};
|
|
86
|
+
// 0.14.6 transitional notes — every non-terminal state-carrying note the
|
|
87
|
+
// machine writes. The ack scan (scan-ack-pending) is atomic code and needs
|
|
88
|
+
// no claim note of its own.
|
|
89
|
+
export const TRANSITIONAL_PUBLISH_NOTES = [
|
|
90
|
+
// The workflow parked at Publish with a checksummed publish-intent ledger
|
|
91
|
+
// entry (carries the per-attempt version). The intent scan claims these.
|
|
92
|
+
"publish: publish-requested",
|
|
93
|
+
// The intent scan claimed this tick's issuance lease (1 hour).
|
|
94
|
+
"publish: publish-intent-claimed",
|
|
95
|
+
// The intent's staged diff was computed against a superseded base —
|
|
96
|
+
// re-queued for the workflow's Publish to re-prepare. Never issued.
|
|
97
|
+
"publish: publish-base-stale",
|
|
98
|
+
// The tick worker issued the edit and recorded it (issuer-stamped
|
|
99
|
+
// submitted ledger entry). The ack scan owns the task from here.
|
|
100
|
+
"publish: edit-issued",
|
|
101
|
+
];
|
|
102
|
+
// The prose-writable set (0.14.6 §1.5, closed publish-note writer): the ONLY
|
|
103
|
+
// verbs tick prose may mint, via the `record-publish-note` crew-api command.
|
|
104
|
+
// Transitional notes only — prose can never mint a terminal state. Legacy
|
|
105
|
+
// terminals are recognized (read-side, so old note history still skips as
|
|
106
|
+
// terminal) but NOT writable — history stays readable; nothing new may be
|
|
107
|
+
// written in the old tongue.
|
|
108
|
+
export const WRITABLE_PUBLISH_NOTES = [
|
|
109
|
+
...TRANSITIONAL_PUBLISH_NOTES,
|
|
110
|
+
];
|
|
111
|
+
// The code-writable set: deterministic code (the scans and the guarded
|
|
112
|
+
// record-* commands) writes through writeGuardedPublishNote, which asserts
|
|
113
|
+
// here — transitional notes plus the four current terminals the machine
|
|
114
|
+
// itself transitions to. Legacy terminals are never written anew.
|
|
115
|
+
export const CODE_WRITABLE_PUBLISH_NOTES = [
|
|
116
|
+
...TRANSITIONAL_PUBLISH_NOTES,
|
|
117
|
+
...CODE_TERMINAL_PUBLISH_NOTES,
|
|
118
|
+
];
|
|
119
|
+
function extractPublishVerb(message) {
|
|
120
|
+
const m = /publish:\s*([a-z0-9-]+)/.exec(String(message || ""));
|
|
121
|
+
return m ? `publish: ${m[1]}` : null;
|
|
122
|
+
}
|
|
39
123
|
export function matchTerminalPublishNote(message) {
|
|
40
124
|
for (const note of TERMINAL_PUBLISH_NOTES) {
|
|
41
125
|
if (message.includes(note)) return note; // same substring semantics the scan already uses
|
|
42
126
|
}
|
|
43
127
|
return null;
|
|
44
128
|
}
|
|
129
|
+
// Prose-side guard (0.14.6 §1.5, closed publish-note writer): the
|
|
130
|
+
// `record-publish-note` command asserts here. Returns the matched
|
|
131
|
+
// "publish: <verb>". Throws (usage) on anything else — prose cannot mint
|
|
132
|
+
// new states, and cannot mint terminal states.
|
|
133
|
+
export function assertWritablePublishNote(message) {
|
|
134
|
+
const verb = extractPublishVerb(message);
|
|
135
|
+
if (!verb || !WRITABLE_PUBLISH_NOTES.includes(verb)) {
|
|
136
|
+
throw new Error(
|
|
137
|
+
`unregistered publish-note verb ${verb === null ? "(none found)" : `'${verb}'`} — ` +
|
|
138
|
+
`tick prose may only mint transitional publish notes (lib/publish-note-vocabulary.js); ` +
|
|
139
|
+
`terminal states are written by deterministic code, never prose. Legacy terminals are recognized but not writable.`
|
|
140
|
+
);
|
|
141
|
+
}
|
|
142
|
+
return verb;
|
|
143
|
+
}
|
|
144
|
+
// Code-side guard: writeGuardedPublishNote asserts here. Returns the
|
|
145
|
+
// matched "publish: <verb>". Throws on anything else — a registry
|
|
146
|
+
// violation is a caller bug (usage).
|
|
147
|
+
export function assertCodeWritablePublishNote(message) {
|
|
148
|
+
const verb = extractPublishVerb(message);
|
|
149
|
+
if (!verb || !CODE_WRITABLE_PUBLISH_NOTES.includes(verb)) {
|
|
150
|
+
throw new Error(
|
|
151
|
+
`unregistered publish-note verb ${verb === null ? "(none found)" : `'${verb}'`} — ` +
|
|
152
|
+
`state-carrying publish notes are a closed set (lib/publish-note-vocabulary.js); ` +
|
|
153
|
+
`register the verb before writing it. Legacy terminals are recognized but never written anew.`
|
|
154
|
+
);
|
|
155
|
+
}
|
|
156
|
+
return verb;
|
|
157
|
+
}
|
package/lib/schema.sql
CHANGED
|
@@ -205,3 +205,22 @@ CREATE TABLE IF NOT EXISTS platform_run_tasks (
|
|
|
205
205
|
linked_at TEXT NOT NULL DEFAULT (datetime('now'))
|
|
206
206
|
);
|
|
207
207
|
CREATE INDEX IF NOT EXISTS platform_run_tasks_task_id_idx ON platform_run_tasks(task_id);
|
|
208
|
+
|
|
209
|
+
-- Builder reports (0.14.6, 2026-09-20). The tick worker records the artifact
|
|
210
|
+
-- platform's report text verbatim after each directly-issued edit via the
|
|
211
|
+
-- `record-builder-report` command. The command mechanically rejects unless
|
|
212
|
+
-- the report text contains the exact per-attempt version, so every row here is
|
|
213
|
+
-- a version-keyed refusal the ack scan can evaluate: outcome 'refused' + a
|
|
214
|
+
-- version matching an issued attempt = terminal publish-refused. (The old
|
|
215
|
+
-- 'acknowledged' outcome was cut 2026-09-20 REVIEW as circular — disk is the
|
|
216
|
+
-- sole positive evidence.) Rows are evidence, never deleted by the scans.
|
|
217
|
+
CREATE TABLE IF NOT EXISTS builder_reports (
|
|
218
|
+
id TEXT PRIMARY KEY,
|
|
219
|
+
task_id TEXT NOT NULL REFERENCES tasks(id) ON DELETE CASCADE,
|
|
220
|
+
version TEXT NOT NULL,
|
|
221
|
+
outcome TEXT NOT NULL CHECK (outcome IN ('acknowledged', 'refused')),
|
|
222
|
+
report_text TEXT NOT NULL,
|
|
223
|
+
created_at TEXT NOT NULL DEFAULT (datetime('now'))
|
|
224
|
+
);
|
|
225
|
+
CREATE INDEX IF NOT EXISTS builder_reports_task_id_idx ON builder_reports(task_id);
|
|
226
|
+
CREATE INDEX IF NOT EXISTS builder_reports_version_idx ON builder_reports(version);
|
package/package.json
CHANGED
package/seed/AGENTS.md
CHANGED
|
@@ -5,6 +5,7 @@ Init source data. Everything `crew-init.js` reads when setting up a new crew ins
|
|
|
5
5
|
- `crons.json` — the declarative cron manifest; crew-init creates/updates each entry idempotently; `enabled` is a creation-time default only
|
|
6
6
|
- `cron-body-template.md` — template for creating the dispatch cron job
|
|
7
7
|
- `cron-body-update-watch.md` — template for the daily `crew-update-watch` cron job: runs `lib/update-watch.js --crew-home {crewHome}` verbatim and reports stdout
|
|
8
|
+
- `cron-body-ack-scan.md` — template for the 5-minute `crew-ack-scan` cron job: runs the observation-only `scan-ack-pending` pass so publish verdicts land quickly. It never issues artifact edits — issuance stays in the 15-minute poll tick's claim-and-issue path.
|
|
8
9
|
- `posture.md` — PM posture that shapes the user's assistant for crew interaction
|
|
9
10
|
- `feedback/` — feedback feature placeholder
|
|
10
11
|
- `workflows/` — human-readable workflow definitions, copied to `$CREW_HOME/.orchestration/workflows/` during init
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
# Muse Crew acknowledgement scan (fast observation tick)
|
|
2
|
+
|
|
3
|
+
You are the ack-scan worker for the Muse Crew at {crewHome}. This is an
|
|
4
|
+
observation-only tick: it runs the version acknowledgement scan every 5
|
|
5
|
+
minutes so publish verdicts land quickly. It NEVER issues artifact edits —
|
|
6
|
+
issuance belongs to the 15-minute poll tick's Step 4.4/4.4b claim-and-issue
|
|
7
|
+
path.
|
|
8
|
+
|
|
9
|
+
1. Run the scan verbatim:
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
node {crewHome}/lib/crew-api.js --crew-home {crewHome} scan-ack-pending
|
|
13
|
+
```
|
|
14
|
+
|
|
15
|
+
The command cross-checks that the running release is the active one
|
|
16
|
+
(`resolveActiveRelease`); if it reports a release mismatch, do NOT
|
|
17
|
+
proceed — log the mismatch and stop.
|
|
18
|
+
|
|
19
|
+
Failure behavior (mechanical): if the scan exits nonzero, report its
|
|
20
|
+
stderr verbatim as the complete result of this tick, take NO other
|
|
21
|
+
action, do NOT rerun the scan, do NOT invoke any other Crew API
|
|
22
|
+
command, and stop. A release mismatch ("not the active release") is a
|
|
23
|
+
stop, not a retry. A nonzero exit is never retried inside this tick —
|
|
24
|
+
the next 5-minute tick runs the scan again on its own.
|
|
25
|
+
|
|
26
|
+
2. Log every list the scan returns: `acknowledged`, `refused`, `reissued`,
|
|
27
|
+
`timeouts`, `skipped`.
|
|
28
|
+
- `acknowledged` / `refused` / `timeouts` are fully handled by code — take
|
|
29
|
+
NO action.
|
|
30
|
+
- `reissued` entries are fresh intents staged for the poll tick —
|
|
31
|
+
immediately claimable, no backoff. This worker does NOT claim or
|
|
32
|
+
issue them; the poll tick's `scan-publish-intent` +
|
|
33
|
+
`record-intent-issuance` path does that.
|
|
34
|
+
- `skipped` with `window-open` means the attempt's 30-minute window from
|
|
35
|
+
latest issuance is still open — no action. The scan runs every 5 minutes
|
|
36
|
+
so the scan's own latency never eats the window: with 6 scans per
|
|
37
|
+
window, the verdict lands within one cadence of the ack or the expiry.
|
|
38
|
+
|
|
39
|
+
3. Verdicts are evidence-only. The exact per-attempt version string on the
|
|
40
|
+
artifact's disk is the sole positive signal; a `scan-ack-pending` result
|
|
41
|
+
that names `acknowledged` IS the verdict. Never stamp provenance from
|
|
42
|
+
prose, never infer a verdict from timing, and never issue a publish edit
|
|
43
|
+
yourself — a publish edit is always the tick worker's directly-issued
|
|
44
|
+
`artifact_edit` against a claimed intent, never an agent-issued one.
|
|
@@ -2,6 +2,16 @@
|
|
|
2
2
|
|
|
3
3
|
You are the dispatch trigger for Muse Crew. Run the authoritative dispatcher workflow and launch its claims autonomously.
|
|
4
4
|
|
|
5
|
+
### Known blockers — read first, do not re-derive or re-report
|
|
6
|
+
|
|
7
|
+
The failures below are settled and recorded in the Gate 1 OODA state. The authoritative list lives at `~/workspace/goals/gate-1-onboarding-canary-passes/hidden_files/ooda/known-blockers.md` — if this copy ever disagrees with that file, the file wins. Their fixes are already queued in the main repo BUILD. If you observe one, log exactly one line — `known-blocker: <id> — <one-line observation>` — and move on. Do NOT write root-cause analysis and do NOT write "follow-up needed"; the follow-up is already owned.
|
|
8
|
+
|
|
9
|
+
- **Blocker 22 — wrong publication caller (settled 2026-09-20):** workflow children cannot call `artifact_edit` (the builder requires a parent-conversation session id they don't have). The 0.14.x trigger/observation path is dead by architecture; `submitted` writes from it are phantom. Fix in BUILD: one-party worker-owned artifact publication. (0.14.6: publication is a handoff with strong suggestions, not deterministic patch application. Each attempt carries a fresh deterministic version — `sha256("publish-version:v1:" + task_id + ":" + commit + ":" + attempt)` — synthesized into the staged diff as `.crew-publish-version`. Acknowledgement is the exact version string echoed, nothing else — Eric's contract: "If we hear that the artifact acknowledges our version, that's it. We don't verify against content.")
|
|
10
|
+
- **Blocker 23 — retry harness cannot access Crew state (settled 2026-09-20):** `lib/retry-publish.js` imported `better-sqlite3` (not shipped; the release uses `node:sqlite`) and targeted `crew.db` instead of the real `crew-state.db`. Every retry prepare returned `db-unavailable` and failed closed before any lock, edit, or journal write. Fixed in BUILD (0.14.5): ported to `node:sqlite` + `$CREW_HOME/crew-state.db`, the missing `CREW_REPO` for the merge lock supplied, and the trigger step converted from a spawned child to the session-carrying tick worker's direct `artifact_edit` call (blocker 22). (0.14.6: the retry harness, the publish-absence classifier, and the content-verification chain are RETIRED — subsumed into the single version attempt loop: a 30-minute window per attempt with immediate reissue (no backoff), a 2-hour total budget from first issuance, and the ack scan as the sole verdict. There is no separate retry queue and no content verifier.)
|
|
11
|
+
- **Blocker 28 — Step 0 sees completed runs too (settled 2026-09-20):** a COMPLETED platform run can still contain FAILED `agent()` calls — the old Step 0 query (failed runs only) was blind to it. The Step 0 check now also sweeps completed runs in the last 15 minutes with failed agent() calls, through the same `record-platform-failure` / `retry-platform-failure` path. If you observe a failed agent() call inside a completed run, log `known-blocker: 28 — <one line>` and move on.
|
|
12
|
+
|
|
13
|
+
(Maintainer note: this snapshot is refreshed from known-blockers.md at release time. The poller always treats the file as authoritative.)
|
|
14
|
+
|
|
5
15
|
### Steps
|
|
6
16
|
|
|
7
17
|
0. **Check for platform workflow failures (opacity killer):** The platform records workflow run failures in `runtime.workflow_runs` — these are invisible in the crew DB unless you check. A workflow that dies on a fatal `agent()` error (e.g. "subagent bootstrap is no longer authorized") leaves its task stranded with no explanation.
|
|
@@ -18,6 +28,15 @@ You are the dispatch trigger for Muse Crew. Run the authoritative dispatcher wor
|
|
|
18
28
|
`node {crewHome}/lib/crew-api.js --crew-home {crewHome} record-platform-failure --json '{"platform_run_id": "<run_id>", "error_message": "<error>", "platform_created_at": "<created_at>"}'`
|
|
19
29
|
- Then retry each failure (clears stale reservation, re-queues task, or parks after 3 attempts):
|
|
20
30
|
`node {crewHome}/lib/crew-api.js --crew-home {crewHome} retry-platform-failure --json '{"platform_run_id": "<run_id>"}'`
|
|
31
|
+
- **Also** sweep COMPLETED runs containing failed agent() calls (2026-09-20, blocker 28 — the failed-runs query above is blind to a run that completed with a dead child):
|
|
32
|
+
```sql
|
|
33
|
+
SELECT w.run_id, w.created_at, c.error
|
|
34
|
+
FROM runtime.workflow_runs w
|
|
35
|
+
JOIN runtime.workflow_agent_calls c ON c.run_id = w.run_id AND c.status = 'failed'
|
|
36
|
+
WHERE w.created_at > now() - interval '15 minutes' AND w.status = 'completed'
|
|
37
|
+
ORDER BY w.created_at DESC
|
|
38
|
+
```
|
|
39
|
+
- Route each completed-run failure through the same two commands — `record-platform-failure`, then `retry-platform-failure` — exactly as above.
|
|
21
40
|
- Log the results. This ensures transient platform failures don't strand tasks.
|
|
22
41
|
|
|
23
42
|
1. **Load tools:** Call tool_search_load_tool_namespace with paths ["workflow_launch"].
|
|
@@ -35,74 +54,55 @@ You are the dispatch trigger for Muse Crew. Run the authoritative dispatcher wor
|
|
|
35
54
|
- If the acknowledge fails (no reservation exists), DO NOT LAUNCH — the dispatcher did not acquire this task. This is a safety invariant.
|
|
36
55
|
- The launched workflow self-claims the task and clears the reservation as its first actions. If the task was already claimed or is done, the claim fails closed and the run stands down quietly — this is the mechanical duplicate protection, not an error.
|
|
37
56
|
|
|
38
|
-
If the dispatcher returned no claims or the claims array is empty, log NO_DISPATCH and CONTINUE to Step 4.4 — do NOT exit.
|
|
39
|
-
|
|
40
|
-
4.4. **
|
|
41
|
-
- Scan (code): `node {crewHome}/lib/crew-api.js --crew-home {crewHome} scan-publish-
|
|
42
|
-
|
|
43
|
-
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
`
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
Spawn ONE child (subagent) whose first instruction loads the artifact namespace (`tool_search.load_tool_namespace` with paths `["artifact"]`), with this exact brief — the same shape as the workflow's first attempt:
|
|
61
|
-
"The change to apply is the unified diff in the file \"<diff_path>\" (sha256 <diff_sha256>).
|
|
62
|
-
1. Verify the file: run sha256sum on it. If the printed hash is not exactly <diff_sha256>, STOP and end your turn — do not call artifact_edit.
|
|
63
|
-
2. Read the file's full content.
|
|
64
|
-
3. Call artifact_edit with slug \"<slug>\" and verbatim_request:
|
|
65
|
-
'Apply the following change to your source tree, then rebuild and deploy.
|
|
66
|
-
|
|
67
|
-
UNIFIED DIFF (relative to your source tree):
|
|
57
|
+
If the dispatcher returned no claims or the claims array is empty, log NO_DISPATCH and CONTINUE to Step 4.4 — do NOT exit. Intent scan (4.4), ack scan (4.5) and evidence (6) run independently of dispatch claims. A no-claims tick must still issue parked publish intents, scan for acknowledgements, and deliver evidence. (Fixed 2026-09-14: the old "exit on NO_DISPATCH" skipped publication permanently.)
|
|
58
|
+
|
|
59
|
+
4.4. **Publish intent scan (0.14.6):** the workflow parks at Publish with a version-carrying publish-intent ledger entry; issuance belongs to the tick. Deterministic code owns every transition; you are the ferry (scan → issue → record). Run this BEFORE Step 4.5 so an issued edit's ack window is open to the acknowledgement scan on a later tick. Log every list the scan returns.
|
|
60
|
+
- Scan (code): `node {crewHome}/lib/crew-api.js --crew-home {crewHome} scan-publish-intent`
|
|
61
|
+
This atomically claims parked publish intents for this tick (1-hour lease, so a second tick cannot double-issue). It returns `{ intent: [...], skipped: [...], requeued: [...] }`. Each intent entry carries `task_id`, `commit`, `attempt`, `slug`, `repo_path`, `project_id`, `ledger_path`, `base`, `diff_path`, `diff_sha256`, `version` (`<commit>:<attempt>`), `files` (changed paths), `claimed_at`, `claim_expiry`, `reclaimed` (true when a previous tick died mid-issuance). Copy the entry's fields verbatim — never re-derive them.
|
|
62
|
+
- `skipped` = log only, take no action. `requeued` (the intent was superseded — a newer commit restaged and the task is back in the workflow) = log only, take no action.
|
|
63
|
+
- For each entry in `intent`: issue it in Step 4.4b. Issuance is one-party and version-bound; the suggestion handoff and the record protocol live there.
|
|
64
|
+
|
|
65
|
+
4.4b. **One-party intent issuance (0.14.6, blocker 22):** publication is a handoff with strong suggestions — byte-for-byte control is not the contract. The staged diff already contains the synthesized protocol receipt `.crew-publish-version` (content = this attempt's version + newline); EXACT acknowledgement of the version is the sole positive completion criterion (Eric's contract). Issuance belongs to YOU, the session-carrying tick worker, calling `artifact_edit` directly in this turn. Never spawn a child for it (blocker 22: children cannot reach `artifact_edit`).
|
|
66
|
+
- For each intent entry from the 4.4 scan (copy the entry's `task_id`, `commit`, `attempt`, `version`, and `claim_expiry` verbatim — never re-derive them):
|
|
67
|
+
1. **Verify the staged diff (code):** `sha256sum "<diff_path>"` must print exactly `diff_sha256`. If the file is missing, regenerate deterministically — the version pair regenerates the identical version and identical bytes — and re-verify:
|
|
68
|
+
`node {crewHome}/lib/compute-publish-diff.js --repo-path "<repo_path>" --base <base> --commit <commit> --task-id "<task_id>" --attempt <attempt> --out "<diff_path>"`
|
|
69
|
+
then re-run the sha256sum. If the regeneration exits non-zero, or the hash still does not match after regeneration, the diff is UNISSUABLE — record it terminally (code, never prose):
|
|
70
|
+
`node {crewHome}/lib/crew-api.js --crew-home {crewHome} record-intent-unissuable --json '{"task_id": "<task_id>", "claim_expiry": "<claim_expiry>", "reason": "<missing-staged-diff|diff-unregenerable|hash-mismatch-after-regeneration>"}'`
|
|
71
|
+
(copy `claim_expiry` verbatim from the scan entry — it is the compare-and-swap). This writes the terminal `publish: publish-unissuable` note and an issuer-stamped ledger entry — parked for human attention. Never re-claim, never loop on an hourly reclaim; a diff that cannot be checksummed never becomes an edit.
|
|
72
|
+
2. **Re-claim check — the version receipt decides (code):** when the entry's `reclaimed` is true, a previous tick died mid-issuance and may have issued the edit without recording it. Read the artifact's on-disk version receipt `~/workspace/ts-spaces/<slug>/.crew-publish-version-<task_id>` (trim whitespace) and compare it to the entry's `version`. If it EQUALS the dead attempt's version → the edit landed: call `record-version-ack --json '{"task_id":"<task_id>","version":"<entry version>","via":"disk"}'` and NEVER re-issue. If it differs or the file is unreadable → the edit provably did not land (the receipt is the only positive signal): proceed to issuance. Never decide this from prose, from `artifact_status`, or from an inspector's summary.
|
|
73
|
+
3. **Issue the edit directly — in this turn, never in a child.** Load the artifact namespace in your own turn (`tool_search.load_tool_namespace` with paths `["artifact"]`) and call `artifact_edit` with the entry's `slug` and this suggestion handoff (paste the verified staged diff file's full content):
|
|
74
|
+
'Adopt the following change into your source tree, then rebuild and redeploy.
|
|
75
|
+
|
|
76
|
+
CHANGE: task <task_id>, commit <first 12 hex of commit>. Files: <comma-separated files list from the entry's `files`>.
|
|
77
|
+
|
|
78
|
+
SUGGESTED PATCH (staged unified diff, relative to your source tree — treat as a strong suggestion, apply with judgment):
|
|
68
79
|
```diff
|
|
69
|
-
<the full content of the verified file, pasted
|
|
80
|
+
<the full content of the verified staged file, pasted>
|
|
70
81
|
```
|
|
71
82
|
|
|
72
|
-
|
|
73
|
-
-
|
|
74
|
-
-
|
|
75
|
-
-
|
|
76
|
-
|
|
77
|
-
-
|
|
78
|
-
-
|
|
79
|
-
|
|
80
|
-
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
-
|
|
94
|
-
|
|
95
|
-
4.5. **Parent publish verification (docs/publish-verification.md):** The publisher parks instead of stamping provenance; the parent — this tick, the live root agent — verifies content and stamps. Deterministic code detects, reads back, and certifies; you are only the ferry between the deterministic steps (scan → sensor → verifier).
|
|
96
|
-
- Scan (code): `node {crewHome}/lib/crew-api.js --crew-home {crewHome} scan-verification-pending`
|
|
97
|
-
This atomically claims each verification-pending task (1-hour lease, so a second tick cannot double-verify) and reconciles verified-but-still-parked tasks to `in_progress`. It returns `{ to_verify: [...], reconciled: [...] }`. Log both lists. If the scan exits 2 (e.g. the active release cannot be resolved), log the error loudly and continue — do NOT work around it.
|
|
98
|
-
- **Read-back (deterministic sensor, 2026-09-15):** no platform inspection tool is needed — `lib/readback-disk.js` reads the platform's on-disk working copy of the artifact source (`~/workspace/ts-spaces/<slug>/`) and emits the machine-readable findings block the verifier parses. For each entry in `to_verify`, resolve the base via `get-provenance` (`source_commit`; the empty-tree sha `4b825dc642cb6eb9a060e54bf8d69288fbee4904` when nothing is stamped yet — a first publish):
|
|
99
|
-
`node {crewHome}/lib/crew-api.js --crew-home {crewHome} get-provenance --json '{"project_id": "<project_id>"}'`
|
|
100
|
-
(use the entry's `project_id` verbatim), then run:
|
|
101
|
-
`node {crewHome}/lib/readback-disk.js --repo-path "<repo_path>" --commit <commit> --base <base> --slug "<deploy_slug>" --task-id <task_id> > /tmp/readback-<task_id>.txt 2> /tmp/readback-<task_id>.err`
|
|
102
|
-
Use the entry's `repo_path` and `deploy_slug` verbatim. If the sensor exits 0, the result file holds the findings block — hand it to the verify step below. If it exits non-zero, do NOT save or use stdout: log `publish: verification-procedural-error <commit> <first line of the .err file>` and leave the task parked — the next tick retries. A sensor failure is procedural (the read could not be performed), never a content verdict. Do NOT judge content yourself, and do NOT stamp provenance.
|
|
103
|
-
- **Verify (code):** `node {crewHome}/lib/verify-publish.js --crew-home {crewHome} --task-id <task_id> --commit <commit> --base <base> --repo-path "<repo_path>" --slug "<deploy_slug>" --crew-release <crew_release> --project-id <project_id> --inspection-id <task_id>-disk --result-file /tmp/readback-<task_id>.txt`
|
|
104
|
-
Pass `--build-agent-id <id>` from the entry's `build_agent_id` when it is present. Pass `--attempt <attempt>` from the entry's `attempt` when it is present (2026-09-18, blocker 4: the verifier binds the ledger entry by exact (task_id, commit, attempt) — without the attempt, a retry's trigger is indistinguishable from the original's). Use the SAME `<base>` the sensor ran with and the entry's `project_id` verbatim. The verifier parses the findings, compares mechanically against the base..commit diff, checks supersession, and stamps only on a match. Its terminal verdicts (`publish: verified` → task re-queued; `publish: verification-failed` → stays parked) are final — log them and continue.
|
|
105
|
-
- Never stamp provenance from prose. Never infer a verdict from an inspector's summary text. The verify script's machine-checked comparison is the only certification.
|
|
83
|
+
PROTOCOL (the one hard requirement):
|
|
84
|
+
- One file in the patch is `.crew-publish-version-<task_id>` — a protocol receipt, not product code. Apply it like any other file.
|
|
85
|
+
- When you report the result, quote the exact 64-character version `<version>` verbatim in your report text. If you cannot or will not apply this change, say REFUSED and quote the version anyway.
|
|
86
|
+
- Do not make any other source changes.'
|
|
87
|
+
The diff and rules are suggestions now, not byte-exact commands (this also sidesteps the platform newline-escaping defect on verbatim requests — the receipt file is what matters, not byte fidelity).
|
|
88
|
+
Edit-request contract: call artifact_edit exactly once. Never retry the edit yourself. Tool unavailable (the artifact namespace would not load), a timeout, a transport error, or any otherwise ambiguous result is inconclusive: record NOTHING for this task, log the ambiguity in one line, and let the claim expire; the next scan re-claims and step 2's receipt check decides. An unrecorded outcome is the only honest record of an unknown one.
|
|
89
|
+
After the call returns, inspect the result text. (2026-09-20 REVIEW) Do NOT record an "acknowledged" builder report — the positive builder-report path was circular (the worker knows the version, so the report proves nothing) and is removed; the ack scan reads the version receipt from the artifact's disk as the sole positive evidence. If the result is an explicit refusal → `node {crewHome}/lib/crew-api.js --crew-home {crewHome} record-builder-report --json '{"task_id":"<task_id>","version":"<version>","outcome":"refused","report_text":"<refusal text verbatim>"}'` (the command MECHANICALLY rejects unless `report_text` contains the exact version — if the refusal text lacks the version, do NOT call record-builder-report; just proceed to the record step with `refused`).
|
|
90
|
+
4. **Record the outcome (code):**
|
|
91
|
+
`node {crewHome}/lib/crew-api.js --crew-home {crewHome} record-intent-issuance --json '{"task_id": "<task_id>", "claim_expiry": "<claim_expiry>", "edit_outcome": "<accepted|refused>", "refusal_text": "<first line>"}'`
|
|
92
|
+
(omit `refusal_text` unless refused). Copy `claim_expiry` verbatim from the scan entry — it is the compare-and-swap. If the command returns `recorded: false`, log the reason and stop: another tick owns the task. `accepted` writes the issuer-stamped `submitted` ledger entry and mirrors `publish: edit-issued <commit> attempt=<n> version=<commit>:<attempt>` (the 4.5 ack scan claims it on a later tick); `refused` writes `rejected` and the terminal `publish: publish-refused` note. There is no outcome for an inconclusive call (step 3) — do not call record-intent-issuance at all when the edit's fate is unknown.
|
|
93
|
+
- Prose never mints `publish:` notes directly. Any tick-authored publish note goes through exactly one path (code asserts the verb against the closed registry, rejects unknown verbs):
|
|
94
|
+
`node {crewHome}/lib/crew-api.js --crew-home {crewHome} record-publish-note --json '{"task_id":"<task_id>","note":"<full note text>"}'`
|
|
95
|
+
- Crash rules (fail closed, never blind re-issue): if this tick dies between the artifact_edit call and the record step, the next scan re-claims with `reclaimed: true` and step 2's version-receipt check decides between ack (`record-version-ack` with `via: "disk"`) and a fresh issuance. The issuance contract is at-most-one RECORDED issuance per attempt; a duplicate physical issuance across a lease expiry is content-identical (the same checksummed staged diff) and bounded by the ack scan's 2-hour budget.
|
|
96
|
+
|
|
97
|
+
4.5. **Acknowledgement scan (version, 0.14.6):** the version acknowledgement is the entire publish verdict — Eric's contract: "If we hear that the artifact acknowledges our version, that's it. We don't verify against content." No content verification, no manifest comparison, no read-back. A dedicated 5-minute `crew-ack-scan` cron runs the same scan on an observation-only tick (it never issues edits) so verdicts land quickly; this tick ALSO runs the scan AFTER 4.4b so an issued edit's verdict path is synchronous. Deterministic code evaluates the evidence and owns every transition; you only log.
|
|
98
|
+
- Scan (code): `node {crewHome}/lib/crew-api.js --crew-home {crewHome} scan-ack-pending`
|
|
99
|
+
For each issued attempt inside its 30-minute window, code checks the on-disk version receipt at `~/workspace/ts-spaces/<slug>/.crew-publish-version-<task_id>` (the SOLE positive evidence — 2026-09-20 REVIEW removed the circular builder-report positive path) plus durably recorded builder reports for explicit refusals. It returns `{ acknowledged: [...], refused: [...], reissued: [...], timeouts: [...], skipped: [...] }`. Log every list.
|
|
100
|
+
- `acknowledged` = fully handled by code — take NO action, just log. The version was acknowledged (exact string match on disk). Code stamps provenance and writes the terminal `publish: version-acknowledged` note; the task is requeued. Any acknowledgement for the same task/commit counts — including a late earlier attempt's version — the receipt is what matters, not which attempt was latest.
|
|
101
|
+
- `refused` = fully handled by code — take NO action, just log. A recorded builder report echoed the version with outcome `refused`. Code writes the terminal `publish: publish-refused` note. An explicit refusal is never re-issued.
|
|
102
|
+
- `timeouts` = fully handled by code — take NO action, just log. The 2-hour total budget from first issuance is exhausted with no acknowledgement; code writes the terminal `publish: version-timeout` note — parked for human attention. `timeouts` is terminal-budget exhaustion only — a window expiry is a `reissued`, never a timeout.
|
|
103
|
+
- `reissued` = fresh intents for the next attempt — a 30-minute window expired with no acknowledgement, so code re-issued with a fresh version (by derivation), immediately claimable (no backoff). This tick does NOT issue them; the NEXT tick's `scan-publish-intent` claims them. Log only.
|
|
104
|
+
- `skipped` = window still open or evidence absent — log only, take no action.
|
|
105
|
+
- Never stamp provenance from prose. Never infer a verdict from an inspector's summary text. The exact version string is the sole positive signal.
|
|
106
106
|
|
|
107
107
|
5. **Monitor launched workflows until terminal (stay-alive — 2026-09-13):** The platform ties async workflow `agent()` authorization to the launcher's lifetime: if THIS tick ends while a workflow is still running, the workflow's next `agent()` call fails with "subagent bootstrap is no longer authorized" / "subagent reservation owner is terminal". Prevention beats recovery here, so this tick is configured with a 90-minute execution timeout (`timeout_secs: 5400` in seed/crons.json) and you MUST stay alive until every launched run reaches a terminal state. Do not exit early while a launched run is still `running` — your death is what kills it.
|
|
108
108
|
- For each launched run_id, poll its status every ~2 minutes via muse.db:
|
package/seed/crons.json
CHANGED
|
@@ -23,6 +23,18 @@
|
|
|
23
23
|
},
|
|
24
24
|
"timeout_secs": 600,
|
|
25
25
|
"title": "Muse Crew automatic update watcher"
|
|
26
|
+
},
|
|
27
|
+
{
|
|
28
|
+
"body_template": "cron-body-ack-scan.md",
|
|
29
|
+
"enabled": true,
|
|
30
|
+
"id": "crew-ack-scan",
|
|
31
|
+
"mode": "task",
|
|
32
|
+
"schedule": {
|
|
33
|
+
"every": "5m",
|
|
34
|
+
"kind": "interval"
|
|
35
|
+
},
|
|
36
|
+
"timeout_secs": 600,
|
|
37
|
+
"title": "Muse Crew acknowledgement scan"
|
|
26
38
|
}
|
|
27
39
|
],
|
|
28
40
|
"version": 1
|