shapeup-sdlc 3.10.0 → 3.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.10.0",
4
+ "version": "3.12.0",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -13,7 +13,7 @@ Skills and commands are named short throughout this file; every one of them reso
13
13
  - **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
14
14
  - GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
15
15
  - Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
16
- - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
16
+ - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it. A relaunch that resumes a run which *did* close reopens it on the record: the earlier close is kept under `prior_closes`, the close fields clear, and the run's next ending is recorded as its own — so a run that aborted, was resumed and then shipped reads `shipped` with the abort beside it, never an abort under a `shipped` status. Opening a run over one that closed says it is closed, not open.
17
17
  - **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
18
18
  - The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
19
19
 
@@ -41,7 +41,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
41
41
 
42
42
  ✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
43
43
 
44
- ✧ **Two refusals, and they answer different questions.** Opening a run refuses outright when the worker skills are not on disk — the roster comes from the schema that defines it, so it cannot drift from a hand-kept list. That alone passes green on the two states that actually happen (installed but disabled, or a different version loaded), so the run's first act is one live canary dispatch, and the evidence is the hook layer's rather than the sub-agent's account of it. A run that cannot reach its workers stops before it spends anything, instead of reporting phases complete while none of the shipped craft was applied.
44
+ ✧ **Two refusals, and they answer different questions.** Opening a run refuses outright when the worker skills are not on disk — the roster comes from the schema that defines it, so it cannot drift from a hand-kept list. That alone passes green on the two states that actually happen (installed but disabled, or a different version loaded), so the run's first act is one live canary dispatch, and the evidence is the hook layer's rather than the sub-agent's account of it. A run that cannot reach its workers stops before it spends anything, instead of reporting phases complete while none of the shipped craft was applied. The same Preflight then runs the profile's `build_probe` and `launch_probe` once, from a sub-agent through the session's own grant, and writes nothing: a missing package install, a probe outside the grant or an absent device is reported before planning is paid for, not at the first round gate. It warns rather than refuses — a baseline can be red for a reason the feature is meant to fix — and whether a permission exists is judged by that run, never by reading a settings file.
45
45
 
46
46
 
47
47
  ### QA Edge Hunt (`/qa-edge-hunter`, post-PASS, pre-ship)
package/SECURITY.md CHANGED
@@ -58,7 +58,10 @@ test against machines you don't own.
58
58
  6. **Every hook decision is recorded.** `hooks/lib/decision.mjs` is the only exit path a hook
59
59
  has, so allow, deny, block and error each leave a row in `.shapeup/decisions.jsonl`. An
60
60
  inert hook and a permitting hook are therefore distinguishable — which matters, because
61
- "exit 0, no output" is what both used to look like.
61
+ "exit 0, no output" is what both used to look like. A permitted Bash call's row names the
62
+ programs the command runs, by basename (`hdc`, `curl | jq`) — never the arguments, which is
63
+ where a secret, a token or a private path would be. A denied call's row keeps the first 200
64
+ characters of the command, as it always has, because a denial must be reviewable.
62
65
 
63
66
  If you find any of these to be false, that is a vulnerability — report it as claim #ⁿ.
64
67
 
@@ -96,6 +96,32 @@ function tokens(segment) {
96
96
  .map((t) => t.replace(/^['"]|['"]$/g, ""));
97
97
  }
98
98
 
99
+ /**
100
+ * The programs a command runs — the executable of each `&&`/`;`/`|` segment, by basename, and
101
+ * nothing else.
102
+ *
103
+ * WHY AN ALLOW ROW NAMES THE PROGRAM. Every permitted Bash call used to be recorded with no subject,
104
+ * so "did this worker ever call the device tool, and was it refused?" had no answer in the ledger:
105
+ * a sub-agent that never tried and one whose call was stopped above this hook left the same rows.
106
+ * The program names answer it. Arguments are never recorded — they are where a secret, a token or a
107
+ * private path would be — and a basename says which tool ran without saying where it lives.
108
+ *
109
+ * @param {string} command - The raw Bash command.
110
+ * @returns {(string|null)} Program basenames joined by " | ", in order, duplicates kept; null when
111
+ * no segment yields one.
112
+ */
113
+ export function programsOf(command) {
114
+ if (!command || typeof command !== "string") return null;
115
+ const names = [];
116
+ for (const segment of command.split(/\s*(?:\|\||&&|;|\||\n)\s*/).filter(Boolean)) {
117
+ const first = commandTokens(segment)[0];
118
+ if (!first) continue;
119
+ const base = first.replace(/^["']|["']$/g, "").split("/").pop();
120
+ if (base && !/^[({]$/.test(base)) names.push(base.slice(0, 64));
121
+ }
122
+ return names.length ? names.join(" | ") : null;
123
+ }
124
+
99
125
  /** Strip leading env assignments and privilege/no-op wrappers to find the real command. */
100
126
  function commandTokens(segment) {
101
127
  const ts = tokens(segment);
@@ -217,8 +243,9 @@ async function main() {
217
243
  const raw = await readStdin();
218
244
  let p;
219
245
  /** Fail-open, with the reason on the record (hooks/lib/decision.mjs). */
220
- const defer = (reason, rule) => settle({
246
+ const defer = (reason, rule, subject = null) => settle({
221
247
  verdict: "allow", event: "PreToolUse", tool: p?.tool_name ?? null, cwd: p?.cwd, reason, rule,
248
+ ...(subject ? { subject } : {}),
222
249
  });
223
250
  try { p = JSON.parse(raw || "{}"); }
224
251
  catch (e) { settle({ verdict: "error", event: "PreToolUse", reason: `unparseable payload: ${e.message}` }); }
@@ -258,7 +285,7 @@ async function main() {
258
285
  if (p.tool_name === "Bash") {
259
286
  const command = p.tool_input?.command || "";
260
287
  const verdict = classifyCommand(command, overrides);
261
- if (!verdict.deny) defer("command matched no destructive rule — inspected and permitted", "bash-clean");
288
+ if (!verdict.deny) defer("command matched no destructive rule — inspected and permitted", "bash-clean", programsOf(command));
262
289
  if (commandOverridden(command, overrides)) {
263
290
  // Exercised override: allowed, but never invisible.
264
291
  logPathology(metricsPath, {
@@ -75,12 +75,12 @@ import { join, dirname, resolve, relative, sep } from "node:path";
75
75
  import { createHash } from "node:crypto";
76
76
  import { decideLane, treeSize } from "./fit.mjs";
77
77
  import { runArgs } from "../lib/argv.mjs";
78
- import { uncoerce } from "../lib/contract.mjs";
78
+ import { uncoerce, splitFrontmatter } from "../lib/contract.mjs";
79
79
  import { deriveSnapshot } from "../reduce/snapshot.mjs";
80
80
  import { mintRunId } from "../lib/paths.mjs";
81
81
  import {
82
82
  localRoot, activeScope, activeOrder, globLocal, globShared, ordersDir, resultsDir,
83
- workflowsStage, globWorkflowsStage, sharedRoot, shapingDir, breadboard as stagedBreadboard,
83
+ workflowsStage, globWorkflowsStage, sharedRoot, shapingDir, breadboard as stagedBreadboard, harnessRun,
84
84
  } from "../lib/paths.mjs";
85
85
  import { parseBreadboard, hasBreadboardTables, idCounts } from "../lib/breadboard.mjs";
86
86
  import { resolveWorkers } from "../verify/skills.mjs";
@@ -603,8 +603,22 @@ export function cli(rawArgv) {
603
603
  if (existsSync(receiptPath) && !args.force) {
604
604
  let resume = null;
605
605
  try { resume = deriveSnapshot(cwd); } catch { /* a broken run must still produce the refusal */ }
606
+ // A receipt says a run EXISTS, not that it is open: a terminal close leaves the receipt in place.
607
+ // Calling a closed run "open" hid the one fact that changes what resuming it means.
608
+ let closed = null;
609
+ try {
610
+ const fm = splitFrontmatter(readFileSync(harnessRun(cwd, slug), "utf8")).meta || {};
611
+ if (fm.closed_status && fm.closed_status !== "~") closed = { status: fm.closed_status, at: fm.closed_at, cause: fm.close_cause };
612
+ } catch { /* no ledger — the receipt alone decides, as before */ }
606
613
  fail(3, [
607
- `✋ init-run: a run is ALREADY OPEN — receipt exists at ${receiptPath}.`,
614
+ closed
615
+ ? `✋ init-run: this run is CLOSED as "${closed.status}" at ${closed.at ?? "?"} — receipt exists at ${receiptPath}.`
616
+ : `✋ init-run: a run is ALREADY OPEN — receipt exists at ${receiptPath}.`,
617
+ ...(closed ? [
618
+ ` Its close: ${closed.cause && closed.cause !== "~" ? closed.cause : "no cause recorded"}`,
619
+ " Resuming it REOPENS it: the first phase that runs moves the ledger off the closed state, keeps",
620
+ " this close under prior_closes, and the run's next terminal close is recorded as its own.",
621
+ ] : []),
608
622
  "",
609
623
  "Do NOT re-initialise and do NOT restart the pipeline from phase 1. Re-opening would discard",
610
624
  "the round history the circuit breaker counts against, and the board, ledger and receipt below",
@@ -33,6 +33,12 @@ const PATTERNS = [
33
33
  // ("ERROR in the build pipeline", "ERROR in test suite failed to run") is left unmatched
34
34
  // instead of handing back a fabricated file.
35
35
  { re: /^(?:ERROR|WARNING)\s+in\s+(\.{1,2}\/[^\s:]*|[^\s:]+\.[A-Za-z0-9]{1,10})\b/i, kind: "compiler-diagnostic" },
36
+ // A test that FAILED BY NAME, with no file:line: "FAIL TS-05-05 step 4: no text 'Bread' on screen"
37
+ // or jest's "FAIL src/cart.test.js". Runners that drive an app from outside it (a device flow, an
38
+ // end-to-end script) report a case this way and nothing else, and the line is the whole signal —
39
+ // dropping it handed the next attempt an empty error list over a red fixture. The name is kept as
40
+ // the file only when it looks like a path; an id like TS-05-05 is not one.
41
+ { re: /^(?:FAIL|FAILED)\s+(\S+)(?:\s+.*)?$/, kind: "named-test-failure" },
36
42
  // Generic "Error: message" line followed later by a stack — capture the message alone.
37
43
  { re: /^\s*(?:Error|TypeError|ReferenceError|AssertionError)\s*:\s*(.+)$/, kind: "error-message" },
38
44
  ];
@@ -71,8 +77,9 @@ export function digest(rawText) {
71
77
  pendingMessage = coreMessage(m[1]);
72
78
  continue; // wait for the stack frame that follows to get a file:line
73
79
  }
74
- const file = m[1]?.trim();
75
- const lineNo = m[2] ? Number(m[2]) : null;
80
+ const named = kind === "named-test-failure";
81
+ const file = named ? (/[\\/]|\.[A-Za-z0-9]{1,10}$/.test(m[1]) ? m[1] : null) : m[1]?.trim();
82
+ const lineNo = !named && m[2] ? Number(m[2]) : null;
76
83
  triples.push({
77
84
  file: file || null,
78
85
  line: lineNo,
@@ -557,14 +557,53 @@ export function setRunStatus(cwd, slug, status) {
557
557
  if (!/^status:.*$/m.test(body)) {
558
558
  return { ok: false, path: p, status, reason: `harness-run.md carries no "status:" line to replace — the ledger's frontmatter is malformed (references/protocol.md)` };
559
559
  }
560
+ // A CLOSED RUN THAT MOVES AGAIN IS REOPENED, AND SAYS SO. A relaunch resumes the same run by
561
+ // design, so a run closed `aborted` can be carried on to a later ship. Its close was written once
562
+ // and the once-only guard in closeRun refuses to flip an outcome, which is right for a close
563
+ // nobody took back — but here somebody did, by resuming it. Leaving the old close in place made
564
+ // the ledger read `status: shipped` over `closed_status: aborted`, and every reader of the close
565
+ // reported an abort for a run that shipped. The move to a live status is the one act that takes a
566
+ // close back, so it is recorded here: the prior close joins `prior_closes`, the three close
567
+ // fields return to `~`, and the next terminal close is written as the run's own. Nothing is lost
568
+ // and nothing is flipped silently — the abort stays on the record beside the outcome that
569
+ // replaced it.
570
+ const fm = parseFrontmatter(body);
571
+ const priorStatus = fm.closed_status && fm.closed_status !== "~" ? String(fm.closed_status) : null;
572
+ let next = body.replace(/^status:.*$/m, `status: ${status}`);
573
+ let reopened = null;
574
+ if (priorStatus && TERMINAL_STATUSES.includes(priorStatus) && !TERMINAL_STATUSES.includes(status)) {
575
+ const priorAt = fm.closed_at && fm.closed_at !== "~" ? String(fm.closed_at) : null;
576
+ const priorCause = fm.close_cause && fm.close_cause !== "~" ? String(fm.close_cause) : null;
577
+ const reopenedAt = new Date().toISOString();
578
+ reopened = { closed_status: priorStatus, closed_at: priorAt, cause: priorCause, reopened_at: reopenedAt };
579
+ const entry = `${priorStatus} at ${priorAt ?? "?"} (reopened ${reopenedAt}): ${priorCause ?? "no cause recorded"}`;
580
+ const history = fm.prior_closes && fm.prior_closes !== "~" ? `${fm.prior_closes} | ${entry}` : entry;
581
+ const historyLine = `prior_closes: ${uncoerce(history)}`;
582
+ next = next
583
+ .replace(/^closed_at:.*$/m, "closed_at: ~")
584
+ .replace(/^closed_status:.*$/m, "closed_status: ~")
585
+ .replace(/^close_cause:.*$/m, "close_cause: ~");
586
+ next = /^prior_closes:.*$/m.test(next)
587
+ ? next.replace(/^prior_closes:.*$/m, historyLine)
588
+ : next.replace(/^closed_at:.*$/m, (m) => `${m}\n${historyLine}`);
589
+ }
560
590
  try {
561
- writeFileSync(p, body.replace(/^status:.*$/m, `status: ${status}`));
591
+ writeFileSync(p, next);
562
592
  } catch (e) {
563
593
  return { ok: false, path: p, status, reason: `could not write the ledger: ${e.message}` };
564
594
  }
565
- const after = parseFrontmatter(readFileSync(p, "utf8")).status;
566
- if (after !== status) {
567
- return { ok: false, path: p, status, reason: `wrote "status: ${status}" but the ledger reads "${after}" — the write did not take` };
595
+ const after = parseFrontmatter(readFileSync(p, "utf8"));
596
+ if (after.status !== status) {
597
+ return { ok: false, path: p, status, reason: `wrote "status: ${status}" but the ledger reads "${after.status}" — the write did not take` };
598
+ }
599
+ if (reopened) {
600
+ // The breadcrumb names a run that is OVER; this one no longer is. Removed only when it names
601
+ // this run, so another run's close is left alone.
602
+ try {
603
+ const crumb = JSON.parse(readFileSync(lastRun(cwd), "utf8"));
604
+ if (crumb?.run_id && crumb.run_id === readRunId(cwd, slug)) rmSync(lastRun(cwd), { force: true });
605
+ } catch { /* no breadcrumb, or unreadable — nothing to retire */ }
606
+ return { ok: true, path: p, status, reopened, decision: "reopened" };
568
607
  }
569
608
  return { ok: true, path: p, status };
570
609
  }
@@ -293,7 +293,15 @@ export function derive({ cwd, slug, appetiteHours = null }) {
293
293
  }
294
294
 
295
295
  /**
296
- * Persist derived `unlocks` into task frontmatter — the ONE write this script makes.
296
+ * Persist derived `unlocks` into task frontmatter — and give a task with NO status the one status a
297
+ * task on a freshly written board can have.
298
+ *
299
+ * `unlocks` is derived, so writing it is safe by construction. `status: todo` is the other field a
300
+ * board is written with, and it is added only where the line is absent — an existing status is never
301
+ * touched, so a task in progress or done is never reset. A regenerated board once came back with
302
+ * `depends_on` and neither field on every task; spec-lint failed all of them at L1b and an unattended
303
+ * run had nobody to add a line the worker's own template already prescribes.
304
+ *
297
305
  * @param {{_tasks:Array<object>, unlocks:Object<string,string[]>}} report - A {@link derive} report.
298
306
  * @returns {string[]} The ids of task files actually rewritten (unchanged files are skipped).
299
307
  * Side effect: writes those task files.
@@ -305,7 +313,8 @@ export function writeUnlocks(report) {
305
313
  const fmMatch = t.body.match(/^---\r?\n([\s\S]*?)\r?\n---/);
306
314
  if (!fmMatch) continue;
307
315
  const fm = fmMatch[1];
308
- const next = /^unlocks:.*$/m.test(fm) ? fm.replace(/^unlocks:.*$/m, `unlocks: ${want}`) : `${fm}\nunlocks: ${want}`;
316
+ let next = /^unlocks:.*$/m.test(fm) ? fm.replace(/^unlocks:.*$/m, `unlocks: ${want}`) : `${fm}\nunlocks: ${want}`;
317
+ if (!/^status:/m.test(next)) next = `${next}\nstatus: todo`;
309
318
  if (next !== fm) {
310
319
  writeFileSync(t.file, t.body.replace(fm, next));
311
320
  written.push(t.id);
@@ -32,7 +32,7 @@ import { runArgs } from "../lib/argv.mjs";
32
32
  import {
33
33
  report as reportPath, tasksDir, verdictsDir, trials, evaluationDir, qaDir,
34
34
  roundLedger, discoveryLedger, receipt as receiptPath, harnessRun, relShared,
35
- activeOrder, runArgsPath, readReceipt, runIdFromReceipt,
35
+ activeOrder, runArgsPath, readReceipt, runIdFromReceipt, hammerCensus,
36
36
  } from "../lib/paths.mjs";
37
37
  import { readTrials } from "../verify/t0.mjs";
38
38
  import { ratchetReport } from "../probe/stats.mjs";
@@ -191,6 +191,31 @@ export function section(md, heading) {
191
191
  return body || null;
192
192
  }
193
193
 
194
+ /**
195
+ * The report's opening line when the ship is not a passing one — null for a PASS.
196
+ *
197
+ * @param {string} verdict - The run's final verdict (`PASS`, `FAIL`, `not-evaluated`, …).
198
+ * @param {({verdict?:string, cut_list?:Array}|null)} census - GATE H's census as data, or null.
199
+ * @returns {(string|null)} One markdown line naming the verdict, and the cut list that cleared it
200
+ * when a census is on record; null when the verdict is PASS.
201
+ */
202
+ export function shipHeadline(verdict, census) {
203
+ if (verdict === "PASS") return null;
204
+ const cuts = Array.isArray(census?.cut_list) ? census.cut_list.length : null;
205
+ const why = census
206
+ ? `GATE H's baseline comparison cleared it${census.verdict ? ` (census: ${census.verdict})` : ""}` +
207
+ (cuts != null ? ` with ${cuts} item${cuts === 1 ? "" : "s"} cut` : "")
208
+ : "no GATE H census was on record when it shipped";
209
+ // A FAIL is a judgement that some criterion failed; anything else here means no judgement was
210
+ // made at all, and the line must not claim criteria failed that nobody graded.
211
+ if (verdict === "FAIL") {
212
+ return `> **Shipped with a FAIL verdict.** Not every criterion passed; ${why}. Read the failed ` +
213
+ "criteria and the cut list below before treating this feature as done.";
214
+ }
215
+ return `> **Shipped without a passing verdict (${verdict}).** No criterion was graded as passing; ${why}. ` +
216
+ "Nothing below is evidence that this feature works.";
217
+ }
218
+
194
219
  /**
195
220
  * Assemble the report. Pure given its inputs, so the structural tests can assert its shape
196
221
  * without a filesystem.
@@ -201,6 +226,7 @@ export function buildReport(facts) {
201
226
  const {
202
227
  slug, at, verdict, qa, rounds, roundsJudged, board, t0, artifacts, ratchet,
203
228
  evalCriteria, evalBugs, qaFindings, decisions, discovered, intakeSha, leftovers, requirements,
229
+ census = null,
204
230
  } = facts;
205
231
 
206
232
  const L = [];
@@ -208,6 +234,12 @@ export function buildReport(facts) {
208
234
  `verdict: ${verdict}`, `rounds_used: ${rounds ?? "~"}`, `rounds_judged: ${roundsJudged ?? "~"}`, `qa: ${qa}`,
209
235
  `intake_sha256: ${intakeSha ?? "~"}`, "---", "");
210
236
  L.push(`# ${slug} — ship report`, "");
237
+ // A SHIP IS NOT A PASS, AND THE FIRST LINE SAYS WHICH ONE THIS IS. A run whose verdict FAILed can
238
+ // still ship — GATE H compares against the baseline, not the ideal, and a cut list can clear it —
239
+ // and its status then reads `shipped` like any other. The verdict sat in a table cell below the
240
+ // fold, so a reader who stopped at the title took a shipped FAIL for a passing build.
241
+ const headline = shipHeadline(verdict, census);
242
+ if (headline) L.push(headline, "");
211
243
  L.push("Frozen at GATE L4. Every figure below is derived from run artifacts on disk — the trial",
212
244
  "ledger, the verdict artifacts, the board — never from a summary of the run.", "");
213
245
 
@@ -361,6 +393,7 @@ export function generate({ cwd, slug, verdict, qa }) {
361
393
  slug,
362
394
  at: today(),
363
395
  verdict: verdict || run.final_verdict || "not-evaluated",
396
+ census: (() => { try { return JSON.parse(readIf(hammerCensus(cwd, slug)) || "null"); } catch { return null; } })(),
364
397
  qa: qa || (huntReport ? "run" : "skipped"),
365
398
  rounds: derivedRounds.rounds_used,
366
399
  roundsJudged: derivedRounds.rounds_judged,
@@ -331,13 +331,40 @@ export function writeRoundBuild(cwd, slug, round, body) {
331
331
  }
332
332
 
333
333
  export const ARGV_SPEC = {
334
- usage: "harness.mjs verify build --slug <slug> --round <N> [--cwd <dir>]",
334
+ usage: "harness.mjs verify build --slug <slug> (--round <N> | --preflight) [--cwd <dir>]",
335
335
  _: { arity: 0, max: 0, name: "(no positional operands)" },
336
336
  slug: { type: "str", required: true },
337
- round: { type: "int", min: 1, required: true },
337
+ round: { type: "int", min: 1 },
338
+ preflight: { type: "flag" },
338
339
  cwd: { type: "path" },
339
340
  };
340
341
 
342
+ /**
343
+ * Classify a preflight: the declared steps run, nothing is written.
344
+ *
345
+ * WHY A PREFLIGHT, AND WHY IT WRITES NOTHING. Every environment fault a run has died of was
346
+ * discoverable in seconds and discovered after most of an hour: a checkout missing its package
347
+ * install or its local SDK pointer fails the build in about a second, a probe outside the session's
348
+ * grant is refused, a device that is not attached leaves every `[ui]` row ungraded. Each surfaced
349
+ * only at the first round gate, after planning had been paid for. The same commands, run once before
350
+ * anything is dispatched, turn that into one report up front. It is not a round, so it writes no
351
+ * round artifact: a red preflight read back as round 0's gate would reach round 1's orders as bugs,
352
+ * handing a feature worker an environment fault it cannot fix.
353
+ *
354
+ * A step that exits 2 is a probe that COULD NOT RUN — the fixture convention, e.g. no device
355
+ * attached — which is a gap in the environment rather than a red build, and is reported apart.
356
+ *
357
+ * @param {{overall:string, steps:Array<{kind:string, exit?:number, pass?:boolean, skipped?:boolean}>}} gate
358
+ * What {@link runGate} returned.
359
+ * @returns {{status:("green"|"red"|"cannot-run"), failed_step:(string|null)}} `cannot-run` when the
360
+ * first failing step exited 2.
361
+ */
362
+ export function classifyPreflight(gate) {
363
+ const failing = gate.steps.find((s) => !s.skipped && !s.pass);
364
+ if (!failing) return { status: "green", failed_step: null };
365
+ return { status: failing.exit === 2 ? "cannot-run" : "red", failed_step: failing.kind };
366
+ }
367
+
341
368
  /**
342
369
  * Run the round build gate and write its artifact.
343
370
  *
@@ -348,9 +375,29 @@ export async function cli(rawArgv) {
348
375
  const args = runArgs(ARGV_SPEC, rawArgv);
349
376
  const cwd = resolve(args.cwd || process.cwd());
350
377
  const { slug, round } = args;
378
+ if (Boolean(args.preflight) === (round != null)) {
379
+ console.error("verify build: pass exactly one of --round <N> (the round gate, writes its artifact) or --preflight (runs the same steps, writes nothing)");
380
+ process.exit(2);
381
+ }
351
382
  const { archetype, steps, warnings } = declaredSteps(cwd, slug);
352
383
  for (const w of warnings) console.error(`build-gate: ${w}`);
353
384
 
385
+ if (args.preflight) {
386
+ if (steps.length === 0) {
387
+ console.log(JSON.stringify({ preflight: true, status: "undeclared", steps: [], warnings }, null, 2));
388
+ process.exit(3);
389
+ }
390
+ const gate = runGate(steps, cwd);
391
+ const { status, failed_step } = classifyPreflight(gate);
392
+ const failing = gate.steps.find((s) => s.kind === failed_step);
393
+ console.log(JSON.stringify({
394
+ preflight: true, status, archetype, warnings,
395
+ steps: gate.steps.map((s) => (s.skipped ? { kind: s.kind, skipped: true } : { kind: s.kind, exit: s.exit, pass: s.pass })),
396
+ ...(failing ? { failed_step, stderr_tail: `${failing.stdout_tail || ""}\n${failing.stderr_tail || ""}`.trim().slice(-1200) } : {}),
397
+ }, null, 2));
398
+ process.exit(status === "green" ? 0 : status === "cannot-run" ? 4 : 1);
399
+ }
400
+
354
401
  if (steps.length === 0) {
355
402
  console.log(JSON.stringify({ round, overall: "skipped", steps: [], warnings,
356
403
  reason: "nothing declared — no run_cmd in the run ledger, no build_probe or launch_probe in project-profile.md" }, null, 2));
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.10.0",
3
+ "version": "3.12.0",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -83,6 +83,11 @@ Done-when statements; `_index.md` Non-Go list. Which UCs are in scope comes from
83
83
  freeze through the judge). Ugly-but-correct PASSes; pretty-but-wrong-`data-state` FAILs.
84
84
  - `[data]`: query the DB/storage, capture actual state.
85
85
  - Contract work: send real requests, compare field-by-field.
86
+ - **A fixture that names a row is evidence for that row.** When a T0 artifact you cite, or the
87
+ build gate, carries output that names a Test Surface row by id — `PASS TS-05-05`, or
88
+ `FAIL TS-05-05 step 4: …` — grade that row on it: a named PASS confirms, a named FAIL is a
89
+ FAIL whose bug is that line. This is how a device row is evidenced when you cannot drive the
90
+ app yourself; it is never a reason to skip driving it when you can.
86
91
  - No evidence collected = recorded "NO EVIDENCE" → FAILs at verdict.
87
92
 
88
93
  **VERDICT.**
@@ -25,7 +25,7 @@ rely on (anything absent = **unknown**; never invent it):
25
25
  | `payload.scope_contract` | The active scope: `affordance_manifest`, `e2e_verification_fixtures`, topology |
26
26
  | `substrate.allowed` / `substrate.shared` | The ONLY globs you may write. A needed file outside them → ESCALATE, never a write (a sandbox hook blocks it anyway) |
27
27
  | `payload.decisions[]` | Adjudicated answers from prior escalations — binding precedent, apply them |
28
- | `payload.digested_errors[]` | `{file, line, core_message}` triples from the previous attempt's failed verification — your starting bug list |
28
+ | `payload.digested_errors[]` | `{file, line, core_message}` triples from the previous attempt's failed verification (a test that failed by name with no file:line — `FAIL TS-05-05 step 4: …` — arrives with `file: null` and the whole line as its message) — your starting bug list |
29
29
  | `payload.trial_history[]` | Up to 8 prior attempts on this scope, oldest first, CROSSING the round boundary: `{score, status, delta, digest}`. `status: "reverted"` is a change that was tried and made things WORSE — do not re-propose it. `status: "kept"` with a still-red score is the tree you are building ON, not a failure to undo. Absent on the first attempt |
30
30
  | `payload.verify.test_cmd` | The command that verifies your work. No test_cmd → command-verifiable ACs still need *some* observable check; say what you used |
31
31
  | `payload.kb_rules_path` | Team guidelines (read if the file exists) — steering, never spec; conflict → the AC wins, note it in `deviations` |
@@ -39,9 +39,9 @@ derived RESUME STATE (slug, status, round, board counts) — read it and go stra
39
39
  this session's memory. `--force` re-opens deliberately and discards the round history the breaker
40
40
  counts.
41
41
 
42
- **If this command comes back "requires approval", stop and say so.** These scripts ship with the
43
- plugin and need a one-time permission grant (`npx shapeup-sdlc init` writes it). Do not route
44
- around it, and do not silently hand-build the feature instead.
42
+ **If this command comes back "requires approval", stop and say so** — the scripts need a one-time grant
43
+ (`npx shapeup-sdlc init`); never route around it or hand-build the feature. Know a grant by running the
44
+ command, never by reading `.claude/settings.json` — Preflight runs the project's probes and reports.
45
45
 
46
46
  **Language gate (delegated to `translator`, not this skill):** before Step 1 opens the run, dispatch
47
47
  an Agent (model: exec) calling `Skill(shapeup-sdlc-plugin:translator) --check` on the pitch *and* its
@@ -477,6 +477,13 @@ Read back: stdout JSON — {overall: green|red, steps[], warnings[], failed_step
477
477
  A `mobile` profile with no launch_probe is warned about on stderr every round, and so is
478
478
  every scope none of whose fixtures invoke the tool run_cmd builds with — advisory, because a
479
479
  green T0 from such fixtures is not evidence the scope compiles.
480
+ Preflight form — the same steps, run once before anything is dispatched, WRITING NOTHING:
481
+ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify build --slug <slug> --preflight
482
+ stdout JSON {preflight: true, status: green|red|cannot-run|undeclared, steps[], failed_step?,
483
+ stderr_tail?}. Exit 0 green · 1 a step red · 4 a probe could not run (it exited 2 — no
484
+ device, no artifact) · 3 nothing declared. No round artifact: a red preflight read back as
485
+ round 0's gate would reach round 1's orders as bugs. The workflow warns on 1 and 4, never
486
+ aborts.
480
487
  Consequences, both mechanical and both read off the artifact, never off this prose:
481
488
  red → EVAL is not dispatched this round; `harness compile` turns each failing step into a
482
489
  `payload.bugs` entry for round N+1, addressed to the scope whose substrate holds the
@@ -54,7 +54,7 @@ export const meta = {
54
54
  name: "shapeup-run",
55
55
  description: "BUILD-phase pipeline: ORIENT → ANALYZE → WIRE → MAP SCOPES → rounds of BUILD/EVAL → QA → GATE H → ship. Gates resolve by the kernel's exit code; every dispatch is WorkOrder in / WorkResult out; the fast-forward is derived from artifacts on disk.",
56
56
  phases: [
57
- { title: "Preflight", detail: "one canary dispatch — can this session resolve a worker skill" },
57
+ { title: "Preflight", detail: "one canary dispatch — can this session resolve a worker skill; then the project's build and launch probes, run once and recorded nowhere" },
58
58
  { title: "Orient" }, { title: "Analyze" }, { title: "Wire" }, { title: "MapScopes" },
59
59
  { title: "Build" }, { title: "Eval" }, { title: "Refute" }, { title: "QA" }, { title: "Ship" },
60
60
  ],
@@ -1117,6 +1117,11 @@ async function setRunStatus(status, phaseName) {
1117
1117
  `artifacts, not from this field), but the snapshot and the the ship report's census hook will read ` +
1118
1118
  `this run as unfinished.`);
1119
1119
  stateWarnings.push(`status="${status}" did not take: ${why}`);
1120
+ } else if (r.decision === "reopened") {
1121
+ // The ledger now carries the earlier close under `prior_closes`; say so where a reader of the
1122
+ // run's own log will see it, because resuming a closed run changes what its close will say.
1123
+ log(`RUN STATE — this run had been closed; moving it to "${status}" reopened it. The earlier close is ` +
1124
+ `kept in the ledger under prior_closes, and the run's next terminal close is recorded as its own.`);
1120
1125
  }
1121
1126
  }
1122
1127
 
@@ -1235,6 +1240,33 @@ if (!canary.ok) {
1235
1240
  `(${canary.detail || `exit ${canary.exit_code}`})`));
1236
1241
  }
1237
1242
 
1243
+ // ENVIRONMENT CANARY — do the project's own build and launch probes run in THIS session?
1244
+ //
1245
+ // The skill canary proves the workers resolve; it says nothing about whether the commands the round
1246
+ // gate will run can. A missing package install or local SDK pointer fails the build in a second, a
1247
+ // probe outside the grant is refused, a device that is not attached leaves every `[ui]` row
1248
+ // ungraded — and each used to surface only at the first round gate, after planning was paid for.
1249
+ // `verify build --preflight` runs the same declared steps from a sub-agent, through the same grant,
1250
+ // and writes nothing. It WARNS rather than aborts: a baseline can be red for a reason the feature is
1251
+ // meant to fix, and an unattended run should still say so up front rather than stop.
1252
+ //
1253
+ // The evidence is this command's exit code. Whether a grant "looks present" in a settings file is
1254
+ // not evidence of anything: a grant can arrive by other routes, and one in the file can still be
1255
+ // refused above it.
1256
+ const envCanary = await cmd(`verify build --slug ${slug} --preflight`, "Preflight", "canary-env");
1257
+ if (envCanary.exit_code === 1 || envCanary.exit_code === 4) {
1258
+ const kind = envCanary.exit_code === 4 ? "a probe could not run (exit 2 — typically no device attached, or a missing artifact)" : "a declared step is red";
1259
+ const msg = `ENVIRONMENT — preflight: ${kind}${envCanary.detail ? `: ${envCanary.detail}` : ""}. The same step runs in every ` +
1260
+ `round's build gate, so fix the environment now if it is not the feature's to fix.`;
1261
+ log(msg);
1262
+ stateWarnings.push(msg);
1263
+ } else if (envCanary.exit_code !== 0 && envCanary.exit_code !== 3) {
1264
+ const msg = `ENVIRONMENT — preflight did not run (exit ${envCanary.exit_code}${envCanary.detail ? `: ${envCanary.detail}` : ""}); ` +
1265
+ `nothing is known about the build or launch probes until the first round gate.`;
1266
+ log(msg);
1267
+ stateWarnings.push(msg);
1268
+ }
1269
+
1238
1270
  // GATE L0.9b's launch record must exist before anything past Preflight dispatches; see
1239
1271
  // requireLaunchRecord()'s own banner for why this cannot be left to Step 2's prose alone.
1240
1272
  const launchRecordAbort = await requireLaunchRecord();
@@ -1374,6 +1406,10 @@ if (!rs.has_spec_tree) {
1374
1406
  "carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder.",
1375
1407
  });
1376
1408
  if (b.__failed) return await withWarnings(diedAt("ANALYZE", b));
1409
+ // The board a worker writes carries `depends_on`; its inverse (`unlocks`) and a new task's
1410
+ // `status: todo` are mechanical, so the kernel writes them rather than trusting every regeneration
1411
+ // to remember — a board that came back without both failed spec-lint on every task at L1b.
1412
+ await advisory(`reduce board --slug ${slug} --write`, "Analyze", "board:derive");
1377
1413
  const post = await requirePhase("ANALYZE", "analyze", "Analyze", "board");
1378
1414
  if (post) return await withWarnings(post);
1379
1415
  await advisory(`reduce graph --slug ${slug}`, "Analyze", "graph:board");