shapeup-sdlc 3.9.2 → 3.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.9.2",
4
+ "version": "3.11.0",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -13,7 +13,7 @@ Skills and commands are named short throughout this file; every one of them reso
13
13
  - **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
14
14
  - GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
15
15
  - Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
16
- - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
16
+ - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it. A relaunch that resumes a run which *did* close reopens it on the record: the earlier close is kept under `prior_closes`, the close fields clear, and the run's next ending is recorded as its own — so a run that aborted, was resumed and then shipped reads `shipped` with the abort beside it, never an abort under a `shipped` status. Opening a run over one that closed says it is closed, not open.
17
17
  - **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
18
18
  - The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
19
19
 
@@ -36,12 +36,12 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
36
36
  | Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
37
37
  | Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
38
38
  | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
39
- | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
39
+ | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade |
40
40
  | FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
41
41
 
42
42
  ✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
43
43
 
44
- ✧ **Two refusals, and they answer different questions.** Opening a run refuses outright when the worker skills are not on disk — the roster comes from the schema that defines it, so it cannot drift from a hand-kept list. That alone passes green on the two states that actually happen (installed but disabled, or a different version loaded), so the run's first act is one live canary dispatch, and the evidence is the hook layer's rather than the sub-agent's account of it. A run that cannot reach its workers stops before it spends anything, instead of reporting phases complete while none of the shipped craft was applied.
44
+ ✧ **Two refusals, and they answer different questions.** Opening a run refuses outright when the worker skills are not on disk — the roster comes from the schema that defines it, so it cannot drift from a hand-kept list. That alone passes green on the two states that actually happen (installed but disabled, or a different version loaded), so the run's first act is one live canary dispatch, and the evidence is the hook layer's rather than the sub-agent's account of it. A run that cannot reach its workers stops before it spends anything, instead of reporting phases complete while none of the shipped craft was applied. The same Preflight then runs the profile's `build_probe` and `launch_probe` once, from a sub-agent through the session's own grant, and writes nothing: a missing package install, a probe outside the grant or an absent device is reported before planning is paid for, not at the first round gate. It warns rather than refuses — a baseline can be red for a reason the feature is meant to fix — and whether a permission exists is judged by that run, never by reading a settings file.
45
45
 
46
46
 
47
47
  ### QA Edge Hunt (`/qa-edge-hunter`, post-PASS, pre-ship)
@@ -44,7 +44,7 @@ import { writeActiveOrder } from "./probe/resume.mjs";
44
44
  import { greenVerdict } from "./probe/t0.mjs";
45
45
  import { attemptEvidence, readReceipts } from "./probe/attempts.mjs";
46
46
  import { readLegs } from "./probe/leg.mjs";
47
- import { latestRoundBuild } from "./verify/build.mjs";
47
+ import { latestRoundBuild, latestRoundBuildFile, launchProbeFor } from "./verify/build.mjs";
48
48
  // The SAME matcher the sandbox hook enforces with. "Is this cited file inside this scope's
49
49
  // substrate" has to mean exactly what the guard means, or a bug is addressed to a scope that is
50
50
  // then denied the write that fixes it.
@@ -748,6 +748,42 @@ export function t0ArtifactsFor(cwd, slug, round) {
748
748
  return { artifacts, missing };
749
749
  }
750
750
 
751
+ // --- the launch evidence the judge grades `[ui]` rows against ---------------------------------
752
+ //
753
+ // WHY THE KERNEL DERIVES IT. A `[ui]` criterion is graded on the RUNNING app, and the evaluator's
754
+ // contract for getting one is `payload.run_cmd`; absent, an orchestrated evaluator escalates rather
755
+ // than guess. The ledger's `run_cmd` is what the round build gate runs FIRST, as the build — on a
756
+ // toolchain where building and launching are different acts it is a build, and the ledger carries
757
+ // none at all when nobody pinned one. Meanwhile the gate had already run the project's launch probe
758
+ // (install, start, assert the first screen) moments earlier, recorded its output, and left the
759
+ // app up. None of that reached the order: the gate's artifact was forwarded only when RED, as the
760
+ // next round's bug list, so a green launch was proven and then unavailable to the one reader who
761
+ // needed it. Every `[ui]` row was graded "no evidence on the running app" over a build that launched.
762
+ //
763
+ // Two fields, both derived from disk, both optional, both absent when there is nothing to say —
764
+ // the same non-regression rule as `t0_artifacts`: a project that declares no launch probe compiles
765
+ // exactly the order it always did.
766
+
767
+ /**
768
+ * What an evaluate order should carry about launching the app.
769
+ *
770
+ * @param {string} cwd - Project root.
771
+ * @param {string} slug - Feature slug.
772
+ * @param {number} [round] - The round being evaluated. Omitted, the newest gate artifact of any
773
+ * round — a standalone evaluation has no round.
774
+ * @returns {{build_gate?: string, launch_cmd?: string}} `build_gate`: repo-relative path of this
775
+ * run's newest round build gate artifact (each step's exit and output tail). `launch_cmd`: the
776
+ * profile's launch probe. A key is omitted, never null, when its source is absent.
777
+ */
778
+ export function launchEvidenceFor(cwd, slug, round) {
779
+ const out = {};
780
+ const gate = latestRoundBuildFile(cwd, slug, round);
781
+ if (gate) out.build_gate = relLocal(slug, "build", basename(gate.path));
782
+ const cmd = launchProbeFor(cwd, slug);
783
+ if (cmd) out.launch_cmd = cmd;
784
+ return out;
785
+ }
786
+
751
787
  /**
752
788
  * Assemble a WorkOrder envelope. Pure given its inputs — the CLI wrapper does the disk reads.
753
789
  * @param {object} opts - The order inputs (destructured):
@@ -1093,6 +1129,13 @@ export async function cli(rawArgv) {
1093
1129
  `${missing.join(", ")}; the evaluator has nothing to cite for ${missing.length === 1 ? "that scope" : "those scopes"}`);
1094
1130
  }
1095
1131
  }
1132
+ // The launch evidence, for every lane (see launchEvidenceFor). Same rule as above: an explicit
1133
+ // `--payload` value outranks the derivation.
1134
+ if (operation === "evaluate") {
1135
+ const ev = launchEvidenceFor(cwd, slug, round);
1136
+ if (ev.build_gate !== undefined && payloadExtra.build_gate === undefined) payloadExtra.build_gate = ev.build_gate;
1137
+ if (ev.launch_cmd !== undefined && payloadExtra.launch_cmd === undefined) payloadExtra.launch_cmd = ev.launch_cmd;
1138
+ }
1096
1139
 
1097
1140
  const order = compileOrder({
1098
1141
  slug, worker, operation, round, attempt, scope, tasks, decisions, digestedErrors, trialHistory, bugs,
@@ -75,12 +75,12 @@ import { join, dirname, resolve, relative, sep } from "node:path";
75
75
  import { createHash } from "node:crypto";
76
76
  import { decideLane, treeSize } from "./fit.mjs";
77
77
  import { runArgs } from "../lib/argv.mjs";
78
- import { uncoerce } from "../lib/contract.mjs";
78
+ import { uncoerce, splitFrontmatter } from "../lib/contract.mjs";
79
79
  import { deriveSnapshot } from "../reduce/snapshot.mjs";
80
80
  import { mintRunId } from "../lib/paths.mjs";
81
81
  import {
82
82
  localRoot, activeScope, activeOrder, globLocal, globShared, ordersDir, resultsDir,
83
- workflowsStage, globWorkflowsStage, sharedRoot, shapingDir, breadboard as stagedBreadboard,
83
+ workflowsStage, globWorkflowsStage, sharedRoot, shapingDir, breadboard as stagedBreadboard, harnessRun,
84
84
  } from "../lib/paths.mjs";
85
85
  import { parseBreadboard, hasBreadboardTables, idCounts } from "../lib/breadboard.mjs";
86
86
  import { resolveWorkers } from "../verify/skills.mjs";
@@ -603,8 +603,22 @@ export function cli(rawArgv) {
603
603
  if (existsSync(receiptPath) && !args.force) {
604
604
  let resume = null;
605
605
  try { resume = deriveSnapshot(cwd); } catch { /* a broken run must still produce the refusal */ }
606
+ // A receipt says a run EXISTS, not that it is open: a terminal close leaves the receipt in place.
607
+ // Calling a closed run "open" hid the one fact that changes what resuming it means.
608
+ let closed = null;
609
+ try {
610
+ const fm = splitFrontmatter(readFileSync(harnessRun(cwd, slug), "utf8")).meta || {};
611
+ if (fm.closed_status && fm.closed_status !== "~") closed = { status: fm.closed_status, at: fm.closed_at, cause: fm.close_cause };
612
+ } catch { /* no ledger — the receipt alone decides, as before */ }
606
613
  fail(3, [
607
- `✋ init-run: a run is ALREADY OPEN — receipt exists at ${receiptPath}.`,
614
+ closed
615
+ ? `✋ init-run: this run is CLOSED as "${closed.status}" at ${closed.at ?? "?"} — receipt exists at ${receiptPath}.`
616
+ : `✋ init-run: a run is ALREADY OPEN — receipt exists at ${receiptPath}.`,
617
+ ...(closed ? [
618
+ ` Its close: ${closed.cause && closed.cause !== "~" ? closed.cause : "no cause recorded"}`,
619
+ " Resuming it REOPENS it: the first phase that runs moves the ledger off the closed state, keeps",
620
+ " this close under prior_closes, and the run's next terminal close is recorded as its own.",
621
+ ] : []),
608
622
  "",
609
623
  "Do NOT re-initialise and do NOT restart the pipeline from phase 1. Re-opening would discard",
610
624
  "the round history the circuit breaker counts against, and the board, ledger and receipt below",
@@ -557,14 +557,53 @@ export function setRunStatus(cwd, slug, status) {
557
557
  if (!/^status:.*$/m.test(body)) {
558
558
  return { ok: false, path: p, status, reason: `harness-run.md carries no "status:" line to replace — the ledger's frontmatter is malformed (references/protocol.md)` };
559
559
  }
560
+ // A CLOSED RUN THAT MOVES AGAIN IS REOPENED, AND SAYS SO. A relaunch resumes the same run by
561
+ // design, so a run closed `aborted` can be carried on to a later ship. Its close was written once
562
+ // and the once-only guard in closeRun refuses to flip an outcome, which is right for a close
563
+ // nobody took back — but here somebody did, by resuming it. Leaving the old close in place made
564
+ // the ledger read `status: shipped` over `closed_status: aborted`, and every reader of the close
565
+ // reported an abort for a run that shipped. The move to a live status is the one act that takes a
566
+ // close back, so it is recorded here: the prior close joins `prior_closes`, the three close
567
+ // fields return to `~`, and the next terminal close is written as the run's own. Nothing is lost
568
+ // and nothing is flipped silently — the abort stays on the record beside the outcome that
569
+ // replaced it.
570
+ const fm = parseFrontmatter(body);
571
+ const priorStatus = fm.closed_status && fm.closed_status !== "~" ? String(fm.closed_status) : null;
572
+ let next = body.replace(/^status:.*$/m, `status: ${status}`);
573
+ let reopened = null;
574
+ if (priorStatus && TERMINAL_STATUSES.includes(priorStatus) && !TERMINAL_STATUSES.includes(status)) {
575
+ const priorAt = fm.closed_at && fm.closed_at !== "~" ? String(fm.closed_at) : null;
576
+ const priorCause = fm.close_cause && fm.close_cause !== "~" ? String(fm.close_cause) : null;
577
+ const reopenedAt = new Date().toISOString();
578
+ reopened = { closed_status: priorStatus, closed_at: priorAt, cause: priorCause, reopened_at: reopenedAt };
579
+ const entry = `${priorStatus} at ${priorAt ?? "?"} (reopened ${reopenedAt}): ${priorCause ?? "no cause recorded"}`;
580
+ const history = fm.prior_closes && fm.prior_closes !== "~" ? `${fm.prior_closes} | ${entry}` : entry;
581
+ const historyLine = `prior_closes: ${uncoerce(history)}`;
582
+ next = next
583
+ .replace(/^closed_at:.*$/m, "closed_at: ~")
584
+ .replace(/^closed_status:.*$/m, "closed_status: ~")
585
+ .replace(/^close_cause:.*$/m, "close_cause: ~");
586
+ next = /^prior_closes:.*$/m.test(next)
587
+ ? next.replace(/^prior_closes:.*$/m, historyLine)
588
+ : next.replace(/^closed_at:.*$/m, (m) => `${m}\n${historyLine}`);
589
+ }
560
590
  try {
561
- writeFileSync(p, body.replace(/^status:.*$/m, `status: ${status}`));
591
+ writeFileSync(p, next);
562
592
  } catch (e) {
563
593
  return { ok: false, path: p, status, reason: `could not write the ledger: ${e.message}` };
564
594
  }
565
- const after = parseFrontmatter(readFileSync(p, "utf8")).status;
566
- if (after !== status) {
567
- return { ok: false, path: p, status, reason: `wrote "status: ${status}" but the ledger reads "${after}" — the write did not take` };
595
+ const after = parseFrontmatter(readFileSync(p, "utf8"));
596
+ if (after.status !== status) {
597
+ return { ok: false, path: p, status, reason: `wrote "status: ${status}" but the ledger reads "${after.status}" — the write did not take` };
598
+ }
599
+ if (reopened) {
600
+ // The breadcrumb names a run that is OVER; this one no longer is. Removed only when it names
601
+ // this run, so another run's close is left alone.
602
+ try {
603
+ const crumb = JSON.parse(readFileSync(lastRun(cwd), "utf8"));
604
+ if (crumb?.run_id && crumb.run_id === readRunId(cwd, slug)) rmSync(lastRun(cwd), { force: true });
605
+ } catch { /* no breadcrumb, or unreadable — nothing to retire */ }
606
+ return { ok: true, path: p, status, reopened, decision: "reopened" };
568
607
  }
569
608
  return { ok: true, path: p, status };
570
609
  }
@@ -1,8 +1,16 @@
1
1
  // probe t0 — "has this scope already reached T0-green in this round?"
2
2
  //
3
3
  // CONTRACT. A bounded, read-only query over the verdict artifacts on disk. Prints
4
- // `{green, path, round, scope_id}` on stdout; exits 0 when green, 1 when not, 2 on a bad argv.
5
- // Writes nothing.
4
+ // `{green, path, sha256?, round, scope_id}` on stdout; exits 0 when green, 1 when not, 2 on a bad
5
+ // argv. Writes nothing.
6
+ //
7
+ // WHY IT PRINTS A DIGEST. A verdict on a scoped spec must cite a T0 artifact with the sha256 of the
8
+ // file as it is on disk now, and the judge is asked to obtain that itself rather than accept one it
9
+ // was handed. Some sessions have no shell hasher they may run, and a grant that covers this entry
10
+ // point and not `shasum` is the ordinary case, not an oversight. The digest here is computed from
11
+ // the bytes at the moment of the call, with the same reading the ingest step re-checks it against,
12
+ // so a judge that asks for it is asking the file, not a caller. `sha256` is present only when the
13
+ // verdict is green and readable; absent — never null — otherwise.
6
14
  //
7
15
  // WHY IT IS A SUBCOMMAND AND NOT AN INLINE SNIPPET. The control plane has no filesystem of its
8
16
  // own, so the alternative is a `node -e` blob assembled inside the workflow script. Such a blob is
@@ -14,6 +22,7 @@
14
22
  // the only durable evidence of that is the verdict artifact the evaluator is required to cite.
15
23
 
16
24
  import { existsSync, readdirSync, readFileSync } from "node:fs";
25
+ import { createHash } from "node:crypto";
17
26
  import { join, resolve } from "node:path";
18
27
  import { runArgs } from "../lib/argv.mjs";
19
28
  import { verdictsDir, readRunId } from "../lib/paths.mjs";
@@ -89,6 +98,18 @@ export function cli(rawArgv) {
89
98
  const args = runArgs(ARGV_SPEC, rawArgv);
90
99
  const cwd = resolve(args.cwd || process.cwd());
91
100
  const { green, path } = greenVerdict(cwd, args.slug, args.scope, args.round);
92
- console.log(JSON.stringify({ green, path, scope_id: args.scope, round: args.round }));
101
+ const sha256 = green ? digestOf(path) : null;
102
+ console.log(JSON.stringify({ green, path, ...(sha256 ? { sha256 } : {}), scope_id: args.scope, round: args.round }));
93
103
  process.exit(green ? 0 : 1);
94
104
  }
105
+
106
+ /**
107
+ * The sha256 of a verdict artifact, read as `reduce ingest` re-reads it (utf8 text, then hashed) so
108
+ * the two can never disagree about the same bytes.
109
+ * @param {string} path - Absolute path of the verdict file.
110
+ * @returns {(string|null)} Lowercase hex digest, or null when the file cannot be read.
111
+ */
112
+ function digestOf(path) {
113
+ try { return createHash("sha256").update(readFileSync(path, "utf8")).digest("hex"); }
114
+ catch { return null; }
115
+ }
@@ -293,7 +293,15 @@ export function derive({ cwd, slug, appetiteHours = null }) {
293
293
  }
294
294
 
295
295
  /**
296
- * Persist derived `unlocks` into task frontmatter — the ONE write this script makes.
296
+ * Persist derived `unlocks` into task frontmatter — and give a task with NO status the one status a
297
+ * task on a freshly written board can have.
298
+ *
299
+ * `unlocks` is derived, so writing it is safe by construction. `status: todo` is the other field a
300
+ * board is written with, and it is added only where the line is absent — an existing status is never
301
+ * touched, so a task in progress or done is never reset. A regenerated board once came back with
302
+ * `depends_on` and neither field on every task; spec-lint failed all of them at L1b and an unattended
303
+ * run had nobody to add a line the worker's own template already prescribes.
304
+ *
297
305
  * @param {{_tasks:Array<object>, unlocks:Object<string,string[]>}} report - A {@link derive} report.
298
306
  * @returns {string[]} The ids of task files actually rewritten (unchanged files are skipped).
299
307
  * Side effect: writes those task files.
@@ -305,7 +313,8 @@ export function writeUnlocks(report) {
305
313
  const fmMatch = t.body.match(/^---\r?\n([\s\S]*?)\r?\n---/);
306
314
  if (!fmMatch) continue;
307
315
  const fm = fmMatch[1];
308
- const next = /^unlocks:.*$/m.test(fm) ? fm.replace(/^unlocks:.*$/m, `unlocks: ${want}`) : `${fm}\nunlocks: ${want}`;
316
+ let next = /^unlocks:.*$/m.test(fm) ? fm.replace(/^unlocks:.*$/m, `unlocks: ${want}`) : `${fm}\nunlocks: ${want}`;
317
+ if (!/^status:/m.test(next)) next = `${next}\nstatus: todo`;
309
318
  if (next !== fm) {
310
319
  writeFileSync(t.file, t.body.replace(fm, next));
311
320
  written.push(t.id);
@@ -32,7 +32,7 @@ import { runArgs } from "../lib/argv.mjs";
32
32
  import {
33
33
  report as reportPath, tasksDir, verdictsDir, trials, evaluationDir, qaDir,
34
34
  roundLedger, discoveryLedger, receipt as receiptPath, harnessRun, relShared,
35
- activeOrder, runArgsPath, readReceipt, runIdFromReceipt,
35
+ activeOrder, runArgsPath, readReceipt, runIdFromReceipt, hammerCensus,
36
36
  } from "../lib/paths.mjs";
37
37
  import { readTrials } from "../verify/t0.mjs";
38
38
  import { ratchetReport } from "../probe/stats.mjs";
@@ -191,6 +191,31 @@ export function section(md, heading) {
191
191
  return body || null;
192
192
  }
193
193
 
194
+ /**
195
+ * The report's opening line when the ship is not a passing one — null for a PASS.
196
+ *
197
+ * @param {string} verdict - The run's final verdict (`PASS`, `FAIL`, `not-evaluated`, …).
198
+ * @param {({verdict?:string, cut_list?:Array}|null)} census - GATE H's census as data, or null.
199
+ * @returns {(string|null)} One markdown line naming the verdict, and the cut list that cleared it
200
+ * when a census is on record; null when the verdict is PASS.
201
+ */
202
+ export function shipHeadline(verdict, census) {
203
+ if (verdict === "PASS") return null;
204
+ const cuts = Array.isArray(census?.cut_list) ? census.cut_list.length : null;
205
+ const why = census
206
+ ? `GATE H's baseline comparison cleared it${census.verdict ? ` (census: ${census.verdict})` : ""}` +
207
+ (cuts != null ? ` with ${cuts} item${cuts === 1 ? "" : "s"} cut` : "")
208
+ : "no GATE H census was on record when it shipped";
209
+ // A FAIL is a judgement that some criterion failed; anything else here means no judgement was
210
+ // made at all, and the line must not claim criteria failed that nobody graded.
211
+ if (verdict === "FAIL") {
212
+ return `> **Shipped with a FAIL verdict.** Not every criterion passed; ${why}. Read the failed ` +
213
+ "criteria and the cut list below before treating this feature as done.";
214
+ }
215
+ return `> **Shipped without a passing verdict (${verdict}).** No criterion was graded as passing; ${why}. ` +
216
+ "Nothing below is evidence that this feature works.";
217
+ }
218
+
194
219
  /**
195
220
  * Assemble the report. Pure given its inputs, so the structural tests can assert its shape
196
221
  * without a filesystem.
@@ -201,6 +226,7 @@ export function buildReport(facts) {
201
226
  const {
202
227
  slug, at, verdict, qa, rounds, roundsJudged, board, t0, artifacts, ratchet,
203
228
  evalCriteria, evalBugs, qaFindings, decisions, discovered, intakeSha, leftovers, requirements,
229
+ census = null,
204
230
  } = facts;
205
231
 
206
232
  const L = [];
@@ -208,6 +234,12 @@ export function buildReport(facts) {
208
234
  `verdict: ${verdict}`, `rounds_used: ${rounds ?? "~"}`, `rounds_judged: ${roundsJudged ?? "~"}`, `qa: ${qa}`,
209
235
  `intake_sha256: ${intakeSha ?? "~"}`, "---", "");
210
236
  L.push(`# ${slug} — ship report`, "");
237
+ // A SHIP IS NOT A PASS, AND THE FIRST LINE SAYS WHICH ONE THIS IS. A run whose verdict FAILed can
238
+ // still ship — GATE H compares against the baseline, not the ideal, and a cut list can clear it —
239
+ // and its status then reads `shipped` like any other. The verdict sat in a table cell below the
240
+ // fold, so a reader who stopped at the title took a shipped FAIL for a passing build.
241
+ const headline = shipHeadline(verdict, census);
242
+ if (headline) L.push(headline, "");
211
243
  L.push("Frozen at GATE L4. Every figure below is derived from run artifacts on disk — the trial",
212
244
  "ledger, the verdict artifacts, the board — never from a summary of the run.", "");
213
245
 
@@ -361,6 +393,7 @@ export function generate({ cwd, slug, verdict, qa }) {
361
393
  slug,
362
394
  at: today(),
363
395
  verdict: verdict || run.final_verdict || "not-evaluated",
396
+ census: (() => { try { return JSON.parse(readIf(hammerCensus(cwd, slug)) || "null"); } catch { return null; } })(),
364
397
  qa: qa || (huntReport ? "run" : "skipped"),
365
398
  rounds: derivedRounds.rounds_used,
366
399
  roundsJudged: derivedRounds.rounds_judged,
@@ -301,6 +301,13 @@
301
301
  "via": "discovered_tasks[]",
302
302
  "note": "a red gate's output, digested — surfaces as payload.bugs on the next round's orders"
303
303
  },
304
+ {
305
+ "from": "WorkOrder",
306
+ "to": "RoundBuildVerdict",
307
+ "cardinality": "N:1",
308
+ "via": "payload.build_gate",
309
+ "note": "an evaluate order names this run's newest gate artifact for its round, green or red, so the judge knows the build ran and the app launched before grading a [ui] row; payload.launch_cmd carries the profile's launch_probe beside it"
310
+ },
304
311
  {
305
312
  "from": "RoundBuildVerdict",
306
313
  "to": "HillShard",
@@ -353,6 +360,8 @@
353
360
  "feature",
354
361
  "dimensions",
355
362
  "run_cmd",
363
+ "launch_cmd",
364
+ "build_gate",
356
365
  "t0_artifacts",
357
366
  "browser",
358
367
  "tasks"
@@ -2416,6 +2425,14 @@
2416
2425
  "type": "string",
2417
2426
  "description": "spec-evaluator: how to start the running app. Absent orchestrated → ESCALATE, never guess."
2418
2427
  },
2428
+ "launch_cmd": {
2429
+ "type": "string",
2430
+ "description": "spec-evaluator: the project profile's launch probe — installs the built artifact, starts it and asserts the first screen. Derived by `harness compile` from project-profile.md, absent when the profile declares none. Where `run_cmd` is only a build, this is how the app is brought up; a non-zero exit is a finding, not a reason to guess another way."
2431
+ },
2432
+ "build_gate": {
2433
+ "type": "string",
2434
+ "description": "spec-evaluator: this run's newest round build gate artifact (build/r<N>-t<T>.json) — each step's exit code and output tail, the record that the build ran and the app launched. Derived by `harness compile`, absent when the gate never ran. Read it before grading a [ui] row NO EVIDENCE."
2435
+ },
2419
2436
  "t0_artifacts": {
2420
2437
  "type": "array",
2421
2438
  "items": {
@@ -220,6 +220,62 @@ export function latestRoundBuild(cwd, slug, round) {
220
220
  return all.length ? all[all.length - 1].body : null;
221
221
  }
222
222
 
223
+ /**
224
+ * The gate artifact an evaluate order should name — the newest one THIS run wrote, with the file it
225
+ * lives in.
226
+ *
227
+ * {@link latestRoundBuild} answers "what did the gate say about round N" and hands back a body; an
228
+ * order has to point at a file, so the path travels with it here. Scoped to the current run when a
229
+ * receipt is readable: a prior run over the same slug leaves its own `r1-t1.json` on disk, and the
230
+ * judge must not be pointed at another run's launch evidence. Scoped to one round when given; with
231
+ * no round it is the newest of any round — a standalone evaluation has none, the same reading
232
+ * `t0ArtifactsFor` uses for the T0 verdicts.
233
+ *
234
+ * @param {string} cwd - Project root.
235
+ * @param {string} slug - Feature slug.
236
+ * @param {number} [round] - The round being evaluated.
237
+ * @returns {({path:string, body:object}|null)} null when the gate never ran, or nothing readable
238
+ * belongs to this run. Never throws: an order must compile without it.
239
+ */
240
+ export function latestRoundBuildFile(cwd, slug, round) {
241
+ let files;
242
+ const dir = roundBuildDir(cwd, slug);
243
+ try { files = readdirSync(dir); } catch { return null; }
244
+ const runId = readRunId(cwd, slug);
245
+ let best = null;
246
+ for (const f of files) {
247
+ const m = f.match(/^r(\d+)-t(\d+)\.json$/);
248
+ if (!m) continue;
249
+ const r = Number(m[1]), t = Number(m[2]);
250
+ if (round != null && r !== Number(round)) continue;
251
+ let body;
252
+ try { body = JSON.parse(readFileSync(join(dir, f), "utf8")); } catch { continue; }
253
+ if (runId && body?.run_id && body.run_id !== runId) continue;
254
+ if (!best || r > best.r || (r === best.r && t > best.t)) best = { r, t, path: join(dir, f), body };
255
+ }
256
+ return best && { path: best.path, body: best.body };
257
+ }
258
+
259
+ /**
260
+ * The command the project declared for bringing the built app up — `launch_probe` in the committed
261
+ * profile: install the artifact, start it, assert the first screen. Read on its own rather than
262
+ * through {@link declaredSteps}, which also reads the ledger and computes warnings an order has no
263
+ * use for.
264
+ *
265
+ * @param {string} cwd - Project root.
266
+ * @param {string} slug - Feature slug.
267
+ * @returns {(string|null)} The command, or null when the profile is absent or declares none.
268
+ * Never throws.
269
+ */
270
+ export function launchProbeFor(cwd, slug) {
271
+ try {
272
+ const pp = projectProfile(cwd, slug);
273
+ if (!existsSync(pp)) return null;
274
+ const v = readContract(pp, PROJECT_PROFILE)?.contract?.launch_probe;
275
+ return typeof v === "string" && v.trim() ? v.trim() : null;
276
+ } catch { return null; }
277
+ }
278
+
223
279
  /**
224
280
  * The rounds whose latest gate artifact is red — the set `reduce hill` subtracts from.
225
281
  * @param {string} cwd - Project root.
@@ -275,13 +331,40 @@ export function writeRoundBuild(cwd, slug, round, body) {
275
331
  }
276
332
 
277
333
  export const ARGV_SPEC = {
278
- usage: "harness.mjs verify build --slug <slug> --round <N> [--cwd <dir>]",
334
+ usage: "harness.mjs verify build --slug <slug> (--round <N> | --preflight) [--cwd <dir>]",
279
335
  _: { arity: 0, max: 0, name: "(no positional operands)" },
280
336
  slug: { type: "str", required: true },
281
- round: { type: "int", min: 1, required: true },
337
+ round: { type: "int", min: 1 },
338
+ preflight: { type: "flag" },
282
339
  cwd: { type: "path" },
283
340
  };
284
341
 
342
+ /**
343
+ * Classify a preflight: the declared steps run, nothing is written.
344
+ *
345
+ * WHY A PREFLIGHT, AND WHY IT WRITES NOTHING. Every environment fault a run has died of was
346
+ * discoverable in seconds and discovered after most of an hour: a checkout missing its package
347
+ * install or its local SDK pointer fails the build in about a second, a probe outside the session's
348
+ * grant is refused, a device that is not attached leaves every `[ui]` row ungraded. Each surfaced
349
+ * only at the first round gate, after planning had been paid for. The same commands, run once before
350
+ * anything is dispatched, turn that into one report up front. It is not a round, so it writes no
351
+ * round artifact: a red preflight read back as round 0's gate would reach round 1's orders as bugs,
352
+ * handing a feature worker an environment fault it cannot fix.
353
+ *
354
+ * A step that exits 2 is a probe that COULD NOT RUN — the fixture convention, e.g. no device
355
+ * attached — which is a gap in the environment rather than a red build, and is reported apart.
356
+ *
357
+ * @param {{overall:string, steps:Array<{kind:string, exit?:number, pass?:boolean, skipped?:boolean}>}} gate
358
+ * What {@link runGate} returned.
359
+ * @returns {{status:("green"|"red"|"cannot-run"), failed_step:(string|null)}} `cannot-run` when the
360
+ * first failing step exited 2.
361
+ */
362
+ export function classifyPreflight(gate) {
363
+ const failing = gate.steps.find((s) => !s.skipped && !s.pass);
364
+ if (!failing) return { status: "green", failed_step: null };
365
+ return { status: failing.exit === 2 ? "cannot-run" : "red", failed_step: failing.kind };
366
+ }
367
+
285
368
  /**
286
369
  * Run the round build gate and write its artifact.
287
370
  *
@@ -292,9 +375,29 @@ export async function cli(rawArgv) {
292
375
  const args = runArgs(ARGV_SPEC, rawArgv);
293
376
  const cwd = resolve(args.cwd || process.cwd());
294
377
  const { slug, round } = args;
378
+ if (Boolean(args.preflight) === (round != null)) {
379
+ console.error("verify build: pass exactly one of --round <N> (the round gate, writes its artifact) or --preflight (runs the same steps, writes nothing)");
380
+ process.exit(2);
381
+ }
295
382
  const { archetype, steps, warnings } = declaredSteps(cwd, slug);
296
383
  for (const w of warnings) console.error(`build-gate: ${w}`);
297
384
 
385
+ if (args.preflight) {
386
+ if (steps.length === 0) {
387
+ console.log(JSON.stringify({ preflight: true, status: "undeclared", steps: [], warnings }, null, 2));
388
+ process.exit(3);
389
+ }
390
+ const gate = runGate(steps, cwd);
391
+ const { status, failed_step } = classifyPreflight(gate);
392
+ const failing = gate.steps.find((s) => s.kind === failed_step);
393
+ console.log(JSON.stringify({
394
+ preflight: true, status, archetype, warnings,
395
+ steps: gate.steps.map((s) => (s.skipped ? { kind: s.kind, skipped: true } : { kind: s.kind, exit: s.exit, pass: s.pass })),
396
+ ...(failing ? { failed_step, stderr_tail: `${failing.stdout_tail || ""}\n${failing.stderr_tail || ""}`.trim().slice(-1200) } : {}),
397
+ }, null, 2));
398
+ process.exit(status === "green" ? 0 : status === "cannot-run" ? 4 : 1);
399
+ }
400
+
298
401
  if (steps.length === 0) {
299
402
  console.log(JSON.stringify({ round, overall: "skipped", steps: [], warnings,
300
403
  reason: "nothing declared — no run_cmd in the run ledger, no build_probe or launch_probe in project-profile.md" }, null, 2));
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.9.2",
3
+ "version": "3.11.0",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -36,7 +36,9 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
36
36
  | `payload.spec_folder` | The committed grading truth: `usecases/` + `domain-model.md` (+ `contracts/`, `scope-summary.md`, `_index.md`). No `usecases/` → HARD STOP, nothing to grade against |
37
37
  | `payload.feature` | Feature slug — scopes the probe and names the report |
38
38
  | `payload.dimensions[]` | The active dimension set (the caller resolved precedence). Absent → `[spec-conformance]` + the auto-enable rules below |
39
- | `payload.run_cmd` | How to start the running app. Absent standalone → ask; absent orchestrated → ESCALATE, do not guess |
39
+ | `payload.run_cmd` | How to start the running app. On a stack where building and launching are different acts it is only the build — prefer `payload.launch_cmd` when the order carries one. Absent standalone → ask; absent orchestrated with no `launch_cmd` either → ESCALATE, do not guess |
40
+ | `payload.launch_cmd` | The project profile's launch probe: installs the built artifact, starts it and asserts the first screen. This is how the app is brought up for `[ui]` probing. A non-zero exit is a finding to cite, not a reason to try another way. Absent → the profile declares none; fall back to `run_cmd` |
41
+ | `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
40
42
  | `payload.t0_artifacts[]` | Per-scope T0 verdict paths for this round (scoped specs), compiled from each scope's green verdict. An artifact listed but missing/red on disk, or a scoped spec with none listed → the round is NOT gradeable: return `status: failed` with the reason, naming the scope, as your FIRST deviation — a structural precondition, not a criterion |
41
43
  | `payload.browser` | `cli` (default, ~4x cheaper) \| `mcp` \| `none` |
42
44
  | `payload.tasks[]` | Traceability only (which UCs a task claims): NEVER a grading source — the committed UC text is the criterion, a paraphrase mismatch is a finding |
@@ -94,7 +96,10 @@ Done-when statements; `_index.md` Non-Go list. Which UCs are in scope comes from
94
96
  contract triplet + Non-Go). Overall PASS only if ALL active dimensions pass — the halo
95
97
  effect is banned; a strong dimension never lifts a failing one.
96
98
  - **T0 citation (scoped specs).** Recompute each cited artifact's sha256 from disk — never
97
- trust a handed hash. A verdict on a scoped spec without a T0 citation is structurally
99
+ trust a handed hash. Use any hasher you may run; when none is available to you,
100
+ `harness probe t0 --slug <slug> --scope <scope_id> --round <r>` prints the digest of that
101
+ scope's green verdict as the file is at that moment. A digest from any other place — a ledger
102
+ line, an earlier report, an error message — is a handed hash, however correct it looks. A verdict on a scoped spec without a T0 citation is structurally
98
103
  invalid, regardless of how convincing your own probing looked; generator prose ("tests
99
104
  pass", "verified locally") is never admissible evidence.
100
105
  - **When a criterion names a command, run THAT command.** Not the one that works, not the
@@ -39,9 +39,9 @@ derived RESUME STATE (slug, status, round, board counts) — read it and go stra
39
39
  this session's memory. `--force` re-opens deliberately and discards the round history the breaker
40
40
  counts.
41
41
 
42
- **If this command comes back "requires approval", stop and say so.** These scripts ship with the
43
- plugin and need a one-time permission grant (`npx shapeup-sdlc init` writes it). Do not route
44
- around it, and do not silently hand-build the feature instead.
42
+ **If this command comes back "requires approval", stop and say so** — the scripts need a one-time grant
43
+ (`npx shapeup-sdlc init`); never route around it or hand-build the feature. Know a grant by running the
44
+ command, never by reading `.claude/settings.json` — Preflight runs the project's probes and reports.
45
45
 
46
46
  **Language gate (delegated to `translator`, not this skill):** before Step 1 opens the run, dispatch
47
47
  an Agent (model: exec) calling `Skill(shapeup-sdlc-plugin:translator) --check` on the pitch *and* its
@@ -477,6 +477,13 @@ Read back: stdout JSON — {overall: green|red, steps[], warnings[], failed_step
477
477
  A `mobile` profile with no launch_probe is warned about on stderr every round, and so is
478
478
  every scope none of whose fixtures invoke the tool run_cmd builds with — advisory, because a
479
479
  green T0 from such fixtures is not evidence the scope compiles.
480
+ Preflight form — the same steps, run once before anything is dispatched, WRITING NOTHING:
481
+ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify build --slug <slug> --preflight
482
+ stdout JSON {preflight: true, status: green|red|cannot-run|undeclared, steps[], failed_step?,
483
+ stderr_tail?}. Exit 0 green · 1 a step red · 4 a probe could not run (it exited 2 — no
484
+ device, no artifact) · 3 nothing declared. No round artifact: a red preflight read back as
485
+ round 0's gate would reach round 1's orders as bugs. The workflow warns on 1 and 4, never
486
+ aborts.
480
487
  Consequences, both mechanical and both read off the artifact, never off this prose:
481
488
  red → EVAL is not dispatched this round; `harness compile` turns each failing step into a
482
489
  `payload.bugs` entry for round N+1, addressed to the scope whose substrate holds the
@@ -490,6 +497,9 @@ compile-order --operation evaluate --slug <slug> --worker spec-evaluator --round
490
497
  --payload '{"dimensions": ["spec-conformance"], "run_cmd": "<cmd>"}'
491
498
  t0_artifacts is compiled from each scope's green T0 verdict for round <r> — pass it only to
492
499
  override. A scope with no green verdict is named on stderr: the judge has nothing to cite for it.
500
+ build_gate and launch_cmd are compiled too: this run's newest round build gate artifact, and the
501
+ profile's launch_probe. Neither is passed here — a `[ui]` row is graded on the running app, and
502
+ these are how the judge finds out the app was launched and how to bring it up again.
493
503
  Invoke via Agent (model: eval), ONCE, after GATE L2:
494
504
  Skill(shapeup-sdlc-plugin:spec-evaluator) --order <path>
495
505
  Effect: one feature-level pass over the running app against all AC + Done-when; writes
@@ -548,7 +558,7 @@ Read back: the proposed cut list + verdict (SHIP now | SHIP after fixing ship-bl
548
558
  | `harness-run.md` | **tech lead (sole writer)** | tech lead (round ledger + Hill + run-state), PO (audit) |
549
559
  | `scopes/<scope-id>.md` | `scope-architect` (sole writer) | tech lead (substrate/sequence), sandbox hook (write-whitelist), compile-order (inlined into orders) |
550
560
  | `t0/verdicts/r<N>-a<M>-t<T>.json` | `harness verify t0` (skill-local, mechanical — not a worker) | spec-evaluator (required citation), tech lead (hill derivation), compile-order (digested errors) |
551
- | `build/r<N>-t<T>.json` (the round build gate) | `harness verify build` (mechanical — run_cmd + build_probe + launch_probe, once per round before EVAL) | tech lead (GATE L2 block), `harness reduce hill` (a red round moves no dot), compile-order (`payload.bugs` for round N+1) |
561
+ | `build/r<N>-t<T>.json` (the round build gate) | `harness verify build` (mechanical — run_cmd + build_probe + launch_probe, once per round before EVAL) | tech lead (GATE L2 block), `harness reduce hill` (a red round moves no dot), compile-order (`payload.bugs` for round N+1 when red; `payload.build_gate` on the same round's evaluate order) |
552
562
  | `t0/trials.jsonl` (the ratchet ledger, append-only, `baseline_trial` as the parent link) | `harness verify t0` (one row per attempt: score, status, delta, tree_ref) | compile-order (`trial_history` into the next order), ship-report (T0 + Ratchet sections), `harness probe stats --ratchet` |
553
563
  | `round-ledger.md` | **tech lead (sole writer)** | compile-order (decisions into every order), PO (audit) |
554
564
  | `hill/<scope-id>.yml` + `hill-chart.md` | **tech lead (sole writer)** | PO ("status without asking"), scope-hammer (H0 census) |
@@ -54,7 +54,7 @@ export const meta = {
54
54
  name: "shapeup-run",
55
55
  description: "BUILD-phase pipeline: ORIENT → ANALYZE → WIRE → MAP SCOPES → rounds of BUILD/EVAL → QA → GATE H → ship. Gates resolve by the kernel's exit code; every dispatch is WorkOrder in / WorkResult out; the fast-forward is derived from artifacts on disk.",
56
56
  phases: [
57
- { title: "Preflight", detail: "one canary dispatch — can this session resolve a worker skill" },
57
+ { title: "Preflight", detail: "one canary dispatch — can this session resolve a worker skill; then the project's build and launch probes, run once and recorded nowhere" },
58
58
  { title: "Orient" }, { title: "Analyze" }, { title: "Wire" }, { title: "MapScopes" },
59
59
  { title: "Build" }, { title: "Eval" }, { title: "Refute" }, { title: "QA" }, { title: "Ship" },
60
60
  ],
@@ -1117,6 +1117,11 @@ async function setRunStatus(status, phaseName) {
1117
1117
  `artifacts, not from this field), but the snapshot and the the ship report's census hook will read ` +
1118
1118
  `this run as unfinished.`);
1119
1119
  stateWarnings.push(`status="${status}" did not take: ${why}`);
1120
+ } else if (r.decision === "reopened") {
1121
+ // The ledger now carries the earlier close under `prior_closes`; say so where a reader of the
1122
+ // run's own log will see it, because resuming a closed run changes what its close will say.
1123
+ log(`RUN STATE — this run had been closed; moving it to "${status}" reopened it. The earlier close is ` +
1124
+ `kept in the ledger under prior_closes, and the run's next terminal close is recorded as its own.`);
1120
1125
  }
1121
1126
  }
1122
1127
 
@@ -1235,6 +1240,33 @@ if (!canary.ok) {
1235
1240
  `(${canary.detail || `exit ${canary.exit_code}`})`));
1236
1241
  }
1237
1242
 
1243
+ // ENVIRONMENT CANARY — do the project's own build and launch probes run in THIS session?
1244
+ //
1245
+ // The skill canary proves the workers resolve; it says nothing about whether the commands the round
1246
+ // gate will run can. A missing package install or local SDK pointer fails the build in a second, a
1247
+ // probe outside the grant is refused, a device that is not attached leaves every `[ui]` row
1248
+ // ungraded — and each used to surface only at the first round gate, after planning was paid for.
1249
+ // `verify build --preflight` runs the same declared steps from a sub-agent, through the same grant,
1250
+ // and writes nothing. It WARNS rather than aborts: a baseline can be red for a reason the feature is
1251
+ // meant to fix, and an unattended run should still say so up front rather than stop.
1252
+ //
1253
+ // The evidence is this command's exit code. Whether a grant "looks present" in a settings file is
1254
+ // not evidence of anything: a grant can arrive by other routes, and one in the file can still be
1255
+ // refused above it.
1256
+ const envCanary = await cmd(`verify build --slug ${slug} --preflight`, "Preflight", "canary-env");
1257
+ if (envCanary.exit_code === 1 || envCanary.exit_code === 4) {
1258
+ const kind = envCanary.exit_code === 4 ? "a probe could not run (exit 2 — typically no device attached, or a missing artifact)" : "a declared step is red";
1259
+ const msg = `ENVIRONMENT — preflight: ${kind}${envCanary.detail ? `: ${envCanary.detail}` : ""}. The same step runs in every ` +
1260
+ `round's build gate, so fix the environment now if it is not the feature's to fix.`;
1261
+ log(msg);
1262
+ stateWarnings.push(msg);
1263
+ } else if (envCanary.exit_code !== 0 && envCanary.exit_code !== 3) {
1264
+ const msg = `ENVIRONMENT — preflight did not run (exit ${envCanary.exit_code}${envCanary.detail ? `: ${envCanary.detail}` : ""}); ` +
1265
+ `nothing is known about the build or launch probes until the first round gate.`;
1266
+ log(msg);
1267
+ stateWarnings.push(msg);
1268
+ }
1269
+
1238
1270
  // GATE L0.9b's launch record must exist before anything past Preflight dispatches; see
1239
1271
  // requireLaunchRecord()'s own banner for why this cannot be left to Step 2's prose alone.
1240
1272
  const launchRecordAbort = await requireLaunchRecord();
@@ -1374,6 +1406,10 @@ if (!rs.has_spec_tree) {
1374
1406
  "carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder.",
1375
1407
  });
1376
1408
  if (b.__failed) return await withWarnings(diedAt("ANALYZE", b));
1409
+ // The board a worker writes carries `depends_on`; its inverse (`unlocks`) and a new task's
1410
+ // `status: todo` are mechanical, so the kernel writes them rather than trusting every regeneration
1411
+ // to remember — a board that came back without both failed spec-lint on every task at L1b.
1412
+ await advisory(`reduce board --slug ${slug} --write`, "Analyze", "board:derive");
1377
1413
  const post = await requirePhase("ANALYZE", "analyze", "Analyze", "board");
1378
1414
  if (post) return await withWarnings(post);
1379
1415
  await advisory(`reduce graph --slug ${slug}`, "Analyze", "graph:board");
@@ -1837,6 +1873,9 @@ while (verdict !== "pass" && round <= maxRounds) {
1837
1873
  model: evalModel, round,
1838
1874
  // No `t0_artifacts` here, deliberately: `harness compile` derives them from the round's green
1839
1875
  // T0 verdicts on disk, for every lane — this script could only name paths it was told about.
1876
+ // The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
1877
+ // artifact and the project profile, so the judge gets the launch evidence without this script
1878
+ // having to carry it.
1840
1879
  payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
1841
1880
  extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself.",
1842
1881
  });