shapeup-sdlc 3.17.1 → 3.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.17.1",
4
+ "version": "3.18.0",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -7,13 +7,13 @@ A three-phase Shape Up loop orchestrated by `/tech-lead`. Invariants live in the
7
7
 
8
8
  Skills and commands are named short throughout this file; every one of them resolves only under the plugin's namespace — `Skill(shapeup-sdlc-plugin:orient)`, `/shapeup-sdlc-plugin:ship`. A bare name is not a typo you get a warning for: an unknown skill name is rejected upstream of the hook layer, so the dispatch fires no hook, leaves no decision row, and is answered by improvisation.
9
9
 
10
- - Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
10
+ - Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing the result of a live order your own dispatch did not take, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
11
11
  - **The tier-direction rule is refused at the write, not argued at L1b.** A committed file cannot carry a `.shapeup/` path or a `TASK-NNN` id — boards renumber per machine and the run trace is gitignored, so either one resolves only where it was written. Spec-lint has always red'd this at Board Review; the same rule now also refuses the write itself, naming the token, because four producers hit it and one of them hit it twice in one session — a worker carries no lesson across a dispatch, and by L1b the phase that wrote the line is over. The knowledge base is outside the rule: those files are instructions, not references anyone resolves.
12
12
  - Attested, not assumed: a dispatch leaves a receipt naming the skill that ran, and ingest refuses an orchestrated result that has none. A dispatch that fails is answered by the sub-agent improvising the craft, which every other check accepts — so "the artifact exists" is not evidence that the shipped skill produced it. `--no-receipt-check` is the way through when the receipt channel itself fails.
13
13
  - **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
14
14
  - GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
15
15
  - Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
16
- - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it. A relaunch that resumes a run which *did* close reopens it on the record: the earlier close is kept under `prior_closes`, the close fields clear, and the run's next ending is recorded as its own — so a run that aborted, was resumed and then shipped reads `shipped` with the abort beside it, never an abort under a `shipped` status. Opening a run over one that closed says it is closed, not open.
16
+ - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it. A relaunch that resumes a run which *did* close reopens it on the record: the earlier close is kept under `prior_closes`, the close fields clear, the run pointer the close retired comes back (so the resumed run's hook decisions carry its key again), and the run's next ending is recorded as its own — so a run that aborted, was resumed and then shipped reads `shipped` with the abort beside it, never an abort under a `shipped` status. Opening a run over one that closed says it is closed, not open.
17
17
  - **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
18
18
  - The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
19
19
 
@@ -36,7 +36,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
36
36
  | Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
37
37
  | Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
38
38
  | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
39
- | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade. A PASS names every Test Surface row as its own criterion: one that grades the surface as a group is refused, and the judge is sent back once with the reason. A scope that did not go green in the round has its rows graded FAIL, and the order names such scopes, so one red scope never makes the whole round ungradeable |
39
+ | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; the order names the dimensions it grades — the set the PO named at L0, or else the set the spec calls for (a Test Surface turns on test-surface-conformance, Invariants turn on completeness, and the test-surface and suite checks are always on), and L4 lists the ones left out. A PASS that grades no criterion under a dimension its order named is refused, and the judge is sent back once; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade. A PASS names every Test Surface row as its own criterion: one that grades the surface as a group is refused, and the judge is sent back once with the reason. A scope that did not go green in the round has its rows graded FAIL, and the order names such scopes, so one red scope never makes the whole round ungradeable |
40
40
  | FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC. Every criterion the verdict graded FAIL reaches the round, whether or not the judge filed a bug for it, addressed to the scope that owns its use case — a round with a FAIL verdict and nothing to fix is how a run stalls. A row whose check the scope rewrote between a failing trial and a passing one is listed for the judge, who reads it against the row before the pass counts |
41
41
 
42
42
  ✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
@@ -62,7 +62,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
62
62
  - **Ledger = single source of truth** — every discovery flow writes only its own section.
63
63
  - **QA is a level-up, not a gate** — `--no-qa` skips it; circuit breaker outranks the Hunter.
64
64
  - **Role separation** — Evaluator grades, task-executor fixes, QA discovers.
65
- - **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
65
+ - **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. A verdict that anchors no criterion at all, beside a board whose ACs carry `covers:`, is sent back to the judge once. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
66
66
  - **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts, never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
67
67
  - **Envelope port (v1.0)** — every dispatch is WorkOrder in / WorkResult out; shared state has exactly one writer (the ingest step); malformed envelopes are hook-denied. Workers: stateless, craft-only, pipeline-blind.
68
68
 
package/README.md CHANGED
@@ -235,7 +235,8 @@ the layer that carries it, and the three layers here fail differently:
235
235
  an evaluation or a QA leg that never returned a result no longer holds the board. `frozen` is
236
236
  checked first and outranks everything, across every live contract — including the carve-out that
237
237
  otherwise keeps the active feature's own run trace writable, so a path a live order froze stays
238
- frozen wherever it lives.
238
+ frozen wherever it lives. An order's own result answers to the sub-agent that took its
239
+ dispatch, so with two legs live neither can write the other's.
239
240
  - `PreToolUse` (`Edit|Write|MultiEdit`) — **`hooks/tier-guard.mjs` refuses a committed-tier write
240
241
  whose content names a path into the local run trace or a board id.** Same rule spec-lint reds at
241
242
  GATE L1b, asked at the moment of writing: four different producers wrote a committed file their
package/SECURITY.md CHANGED
@@ -75,7 +75,7 @@ sitting, and reading them is the recommended review.
75
75
  | [`safety-spine.mjs`](hooks/safety-spine.mjs) | PreToolUse (`Bash\|Read\|Write\|Edit\|MultiEdit`) | The proposed command/path; `.shapeup/safety-overrides.json` | Yes — provably destructive ops only: `rm -rf` on unrecoverable targets, `git push --force` / push to main, `git reset --hard`, `git clean -fdx`, `DROP TABLE`/`TRUNCATE`, reads of `.env`/keys/cloud credentials, and any write to its own overrides file | Never blocks an unmatched command; `--force-with-lease` stays allowed |
76
76
  | [`gate-intake.mjs`](hooks/gate-intake.mjs) | PreToolUse (`Skill`) | The `tech-lead` dispatch's own arguments | Yes — an orchestrator dispatch carrying no resolvable intake (no pitch, spec, resume or requirement text) | Fails open on `--order` and on any ambiguous arg shape |
77
77
  | [`harness verify envelope`](kernel/verify/envelope.mjs) | PreToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the JSON schemas | Yes — a worker dispatch whose order file is missing or schema-invalid | Never gates a dispatch that carries no `--order` (standalone skill use stays free) |
78
- | [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path, resolved through the filesystem before it is compared — symlinks followed, the unborn tail re-attached to its real ancestor, case folded where the filesystem folds case — so a spelling variant of a frozen file is the frozen file; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
78
+ | [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path, resolved through the filesystem before it is compared — symlinks followed, the unborn tail re-attached to its real ancestor, case folded where the filesystem folds case — so a spelling variant of a frozen file is the frozen file; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out. An order's own paths (its result first) are writable only by the sub-agent its dispatch receipt names, so one live leg cannot write another's result; a write with no sub-agent id, or an order with no receipt, keeps the path-only rule | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
79
79
  | [`tier-guard.mjs`](hooks/tier-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the text the call is about to write (`content`, `new_string`, or any edit in a `MultiEdit` batch) | Yes — a write into `shapeup/<slug>/` whose content names a `.shapeup/` path or a `TASK-NNN` board id: the same tier-direction rule spec-lint reds at GATE L1b, refused at the moment of writing, with the offending token quoted back | No-op outside `shapeup/<slug>/`, on file forms the lint does not scan, and on `shapeup/knowledge-base/` (instructions to a worker, not references). Reads nothing off disk and writes only its decision row; it never inspects a file it is not being asked to write |
80
80
  | [`dispatch-receipt.mjs`](hooks/dispatch-receipt.mjs) | PostToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the tool result's own report of which skill ran | **No — it has no deny path at all.** It records that the shipped skill ran, so `harness reduce ingest` can refuse a result no dispatch produced | Never writes an attestation for a result that does not name a resolved skill; never fails the call it observes (every write is inside `try`/`catch`) |
81
81
  | [`gate-zerowork.mjs`](hooks/gate-zerowork.mjs) | Stop | Run receipts on disk; the session transcript; the decision ledger | **Yes — the one blocking hook.** Returns `decision:"block"` when the session dispatched the orchestrator and produced no run receipt | Defers the moment any receipt exists; `stop_hook_active` caps it at one block per stop chain |
@@ -88,7 +88,7 @@
88
88
  import { readFileSync, existsSync, appendFileSync, mkdirSync, readdirSync, statSync, realpathSync, lstatSync, readlinkSync } from "node:fs";
89
89
  import { resolve, join, relative, dirname, basename, sep } from "node:path";
90
90
  import { isMain } from "../kernel/lib/argv.mjs";
91
- import { LOCAL, SHARED, activeOrder, ordersDir, resultsDir, metricsShard } from "../kernel/lib/paths.mjs";
91
+ import { LOCAL, SHARED, activeOrder, ordersDir, resultsDir, metricsShard, dispatchReceipts } from "../kernel/lib/paths.mjs";
92
92
  import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
93
93
 
94
94
  // --- tiny glob matcher: supports *, **, ? — enough for substrate globs, zero dependencies ---
@@ -262,6 +262,41 @@ function liveOrders(cwd, slug) {
262
262
  return unanswered.filter((o) => !pastItsPhase(o, stamps));
263
263
  }
264
264
 
265
+ /**
266
+ * The agent that took each live order's dispatch, from the receipts the dispatch hook wrote.
267
+ *
268
+ * WHO WRITES, not only where. An order's own paths — its WorkResult above all — are carved out of
269
+ * the frozen run trace, and with two build legs live the carve-out answered for either of them: a
270
+ * leg could write its sibling's result. The host names the sub-agent behind every tool call
271
+ * (`agent_id`), and the dispatch receipt recorded the same id when that sub-agent invoked the
272
+ * worker skill, so the skill's writes carry the id of the receipt that aimed it. Only a receipt
273
+ * taken after the order was compiled counts, and the newest wins, so a re-dispatch of the same
274
+ * order binds to the agent that took it this time.
275
+ *
276
+ * @param {string} cwd - Project root.
277
+ * @param {string} slug - The run named by the pointer.
278
+ * @param {object[]} orders - The live orders.
279
+ * @returns {Map<string,string>} order_id → agent_id; an order with no attributable receipt is absent.
280
+ */
281
+ export function dispatchAgents(cwd, slug, orders) {
282
+ const out = new Map();
283
+ let text = "";
284
+ try { text = readFileSync(dispatchReceipts(cwd, slug), "utf8"); } catch { return out; }
285
+ const compiled = new Map(orders.map((o) => [o.order_id, Date.parse(o.compiled_at ?? "")]));
286
+ const newest = new Map();
287
+ for (const line of text.split("\n")) {
288
+ let r; try { r = JSON.parse(line); } catch { continue; }
289
+ if (!r?.order_id || !compiled.has(r.order_id) || typeof r.agent_id !== "string" || !r.agent_id) continue;
290
+ const at = Date.parse(r.at ?? "");
291
+ const c = compiled.get(r.order_id);
292
+ if (Number.isNaN(at) || (!Number.isNaN(c) && at < c)) continue;
293
+ const prev = newest.get(r.order_id);
294
+ if (!prev || at >= prev.at) newest.set(r.order_id, { at, agent: r.agent_id });
295
+ }
296
+ for (const [id, v] of newest) out.set(id, v.agent);
297
+ return out;
298
+ }
299
+
265
300
  function extractPaths(toolInput) {
266
301
  const paths = [];
267
302
  if (toolInput?.file_path) paths.push(toolInput.file_path);
@@ -337,6 +372,7 @@ async function main() {
337
372
  })).filter((c) => c.allowed.length || c.appendOnly.length || c.frozen.length);
338
373
 
339
374
  if (contracts.length === 0) defer("no live order declares write/append/frozen boundaries", "no-whitelist");
375
+ const agents = dispatchAgents(root, active.slug, withSubstrate);
340
376
 
341
377
  const targetPaths = extractPaths(p.tool_input);
342
378
  if (targetPaths.length === 0) defer("no writable path in the tool input", "no-target");
@@ -370,7 +406,17 @@ async function main() {
370
406
  // wherever it lives.
371
407
  // `own` first, and only against the contract that declared it: another live order's exception
372
408
  // never licenses this write. A path no contract claims as its own falls through to the freeze.
373
- if (contracts.some((c) => matchesAny(rel, c.own, fold))) continue;
409
+ // And only from the dispatch that took that order. A write from a sub-agent is refused when every
410
+ // live order claiming the path was taken by a different sub-agent; a write with no agent id (the
411
+ // operator's own session) or an order with no attributable receipt keeps today's permit.
412
+ const owners = contracts.filter((c) => matchesAny(rel, c.own, fold));
413
+ if (owners.length) {
414
+ const writer = typeof p.agent_id === "string" && p.agent_id ? p.agent_id : null;
415
+ if (!writer || owners.some((c) => !agents.has(c.order_id) || agents.get(c.order_id) === writer)) continue;
416
+ violations.push(rel);
417
+ blockReasons.push(`${rel} belongs to ${owners.map((c) => c.order_id).join(", ")}, whose dispatch is another agent's`);
418
+ continue;
419
+ }
374
420
 
375
421
  const freezer = contracts.find((c) => matchesAny(rel, c.frozen, fold));
376
422
  if (freezer) {
@@ -409,7 +455,10 @@ async function main() {
409
455
  // direction. Those files belong to the orchestrator, whose write window is a phase boundary —
410
456
  // no dispatch in flight — and never the middle of somebody else's dispatch.
411
457
  const committed = violations.filter((v) => (fold ? v.split(/[\\/]/)[0].toLowerCase() === SHARED.toLowerCase() : v.split(/[\\/]/)[0] === SHARED));
412
- const hint = committed.length === violations.length
458
+ const notOwn = blockReasons.every((r) => r.endsWith("whose dispatch is another agent's"));
459
+ const hint = notOwn
460
+ ? "Each leg writes only the paths of the order it was dispatched with. Return your own result, never a sibling's."
461
+ : committed.length === violations.length
413
462
  ? `${SHARED}/ is committed tier: these belong to the orchestrator, not to a worker substrate. `
414
463
  + "Write them at a phase boundary, with no dispatch in flight — do not widen an order to reach one."
415
464
  : "If this write legitimately crosses scopes, the order's substrate needs to be expanded (e.g. via ba --remap).";
@@ -430,14 +479,14 @@ async function main() {
430
479
  // FROZE the path and a write refused because no contract covers it are different facts with
431
480
  // different remedies, and a single rule string cannot tell the reader which one happened —
432
481
  // which is how a frozen declaration can stop being enforced without a single row moving.
433
- rule: frozenHits === violations.length ? "frozen" : "outside-substrate",
482
+ rule: frozenHits === violations.length ? "frozen" : notOwn ? "not-own-dispatch" : "outside-substrate",
434
483
  reason: `${violations.length} write(s) rejected by substrate boundaries: ${blockReasons.join("; ")}`,
435
484
  payload: {
436
485
  hookSpecificOutput: {
437
486
  hookEventName: "PreToolUse",
438
487
  permissionDecision: "deny",
439
488
  permissionDecisionReason:
440
- `Sandbox guard (PA3) — no live order's substrate covers these writes:\n` +
489
+ `Sandbox guard (PA3) — ${notOwn ? "these paths belong to another live dispatch" : "no live order's substrate covers these writes"}:\n` +
441
490
  `${blockReasons.join("\n")}\n` +
442
491
  hint,
443
492
  },
@@ -921,6 +921,50 @@ export function launchEvidenceFor(cwd, slug, round) {
921
921
  return out;
922
922
  }
923
923
 
924
+ // --- the dimensions the judge grades ------------------------------------------------------------
925
+ //
926
+ // WHY THE KERNEL RESOLVES THEM. The judge's own rule is that an explicit `dimensions` list overrides
927
+ // its auto-enable, and every orchestrated order used to carry one: the ledger's default
928
+ // `[spec-conformance]`, which nobody chose. So a spec whose every use case carried a Test Surface
929
+ // was never graded by `test-surface-conformance`, `completeness` never ran over Invariants, and
930
+ // `tdd-surface` — always-on in the judge's registry — was off in every orchestrated run. Resolving
931
+ // the set here, from the spec, with the judge's rules, makes it a fact on the order rather than a
932
+ // default: the order file is the record of what was graded, and GATE L4 reads it back.
933
+
934
+ /** The shipped dimensions, in the judge's load order. Security and performance ship disabled. */
935
+ export const SHIPPED_DIMENSIONS = ["spec-conformance", "tdd-surface", "integration", "completeness",
936
+ "test-surface-conformance", "security", "performance"];
937
+
938
+ /**
939
+ * The dimension set an evaluate order grades when the run named none — the judge registry's rules,
940
+ * applied to the spec on disk.
941
+ *
942
+ * @param {string} cwd - Project root.
943
+ * @param {string} slug - Feature slug (reads the board for task variants).
944
+ * @param {(string|null)} specDir - The spec folder, relative to `cwd`; null → no use cases read.
945
+ * @returns {string[]} `spec-conformance` and `tdd-surface` always; `integration` when a board task is
946
+ * a `.be`/`.e2e` variant; `completeness` when a use case has `## Invariants`; `test-surface-conformance`
947
+ * when one has `## Test Surface`. In {@link SHIPPED_DIMENSIONS} order.
948
+ */
949
+ export function resolveDimensions(cwd, slug, specDir) {
950
+ const on = new Set(["spec-conformance", "tdd-surface"]);
951
+ let tasks = [];
952
+ try { tasks = readBoard(cwd, slug); } catch { /* an unreadable board names no variant */ }
953
+ if (tasks.some((t) => /\.(be|e2e)$/i.test(String(t.id || "")))) on.add("integration");
954
+ if (specDir) {
955
+ const dir = join(resolve(cwd, specDir), "usecases");
956
+ let files = [];
957
+ try { files = readdirSync(dir).filter((f) => /^UC-.*\.md$/i.test(f)); } catch { /* no use cases */ }
958
+ for (const f of files) {
959
+ let body = "";
960
+ try { body = readFileSync(join(dir, f), "utf8"); } catch { continue; }
961
+ if (/^##\s+Invariants\b/m.test(body)) on.add("completeness");
962
+ if (/^##\s+Test Surface\b/m.test(body)) on.add("test-surface-conformance");
963
+ }
964
+ }
965
+ return SHIPPED_DIMENSIONS.filter((d) => on.has(d));
966
+ }
967
+
924
968
  /**
925
969
  * Assemble a WorkOrder envelope. Pure given its inputs — the CLI wrapper does the disk reads.
926
970
  * @param {object} opts - The order inputs (destructured):
@@ -1278,6 +1322,11 @@ export async function cli(rawArgv) {
1278
1322
  `${missing.join(", ")}; the evaluator has nothing to cite for ${missing.length === 1 ? "that scope" : "those scopes"}`);
1279
1323
  }
1280
1324
  }
1325
+ // The graded dimensions (see resolveDimensions). A set the PO named at L0.5 travels in `--payload`
1326
+ // and wins; an absent or empty list is resolved from the spec, never left to the judge's default.
1327
+ if (operation === "evaluate" && !(Array.isArray(payloadExtra.dimensions) && payloadExtra.dimensions.length)) {
1328
+ payloadExtra.dimensions = resolveDimensions(cwd, slug, payloadExtra.spec_folder || specDir);
1329
+ }
1281
1330
  // The launch evidence, for every lane (see launchEvidenceFor). Same rule as above: an explicit
1282
1331
  // `--payload` value outranks the derivation.
1283
1332
  if (operation === "evaluate" && payloadExtra.revised_checks === undefined) {
@@ -52,7 +52,7 @@
52
52
  // --max-rounds N outer circuit breaker (default: 3)
53
53
  // --attempts N inner per-scope T0 budget (default: 5)
54
54
  // --spec-folder SHARED spec deliverable path (default: shapeup/<slug>/spec/)
55
- // --dimensions comma-separated eval dimensions (default: spec-conformance)
55
+ // --dimensions comma-separated eval dimensions (default: auto — resolved from the spec per EVAL order)
56
56
  // --gate-answers path | preset name (see `harness gate`; recorded, not read)
57
57
  // --breadboard path to the pitch's breadboard (default: found — see resolveBreadboard())
58
58
  // --wall-clock-budget N deadline breaker, seconds (off by default; see `harness verify budget`)
@@ -91,10 +91,13 @@ const AUTO_LEVELS = new Set(["interactive", "auto", "unattended"]);
91
91
  const LENSES = new Set(["lite", "standard", "cross-context"]);
92
92
 
93
93
  /**
94
- * The eval dimension set when the caller names none. Kept to the base correctness dimension so an
95
- * unconfigured run behaves exactly as it did before this flag existed.
94
+ * The eval dimension set when the caller names none: `auto`. The spec does not exist yet when a run
95
+ * opens, so the set cannot be resolved here; each evaluate order resolves it from the spec on disk
96
+ * (`harness compile`), with the judge's own auto-enable rules. A fixed default in this place made
97
+ * every orchestrated run grade `spec-conformance` alone, because an explicit list switches the
98
+ * judge's auto-enable off.
96
99
  */
97
- export const DEFAULT_DIMENSIONS = ["spec-conformance"];
100
+ export const DEFAULT_DIMENSIONS = "auto";
98
101
 
99
102
  /**
100
103
  * Parse `--dimensions` into the set written to the ledger. Shape-validated only, NOT checked against
@@ -103,14 +106,14 @@ export const DEFAULT_DIMENSIONS = ["spec-conformance"];
103
106
  * unreachable. An id with no file behind it is skipped-with-a-warning at dimension resolution, which
104
107
  * is where that check belongs and where it can actually see the files.
105
108
  *
106
- * @param {(string|null|undefined)} raw - The comma-separated flag value; absent → the default set.
107
- * @returns {string[]} Trimmed, de-duplicated ids in the caller's order.
109
+ * @param {(string|null|undefined)} raw - The comma-separated flag value; absent → `auto`.
110
+ * @returns {(string[]|string)} Trimmed, de-duplicated ids in the caller's order, or `auto`.
108
111
  * @throws {Error} If the list is empty or an entry is not a kebab-case id.
109
112
  */
110
113
  export function parseDimensions(raw) {
111
- if (raw === null || raw === undefined) return [...DEFAULT_DIMENSIONS];
114
+ if (raw === null || raw === undefined) return DEFAULT_DIMENSIONS;
112
115
  const ids = String(raw).split(",").map((s) => s.trim()).filter(Boolean);
113
- if (!ids.length) throw new Error("--dimensions: empty list — omit the flag to use the default [spec-conformance]");
116
+ if (!ids.length) throw new Error("--dimensions: empty list — omit the flag to resolve the set from the spec");
114
117
  for (const id of ids) {
115
118
  if (!/^[a-z][a-z0-9]*(-[a-z0-9]+)*$/.test(id)) {
116
119
  throw new Error(`--dimensions: "${id}" is not a dimension id (kebab-case, e.g. spec-conformance, tdd-surface, integration)`);
@@ -35,7 +35,9 @@ import { existsSync, readFileSync, readdirSync, realpathSync } from "node:fs";
35
35
  import { join, resolve, resolve as resolvePath, sep } from "node:path";
36
36
  import { createHash } from "node:crypto";
37
37
  import { runArgs } from "../lib/argv.mjs";
38
- import { resultsDir, scopesDir, readRunId, verdictsDir, specDir } from "../lib/paths.mjs";
38
+ import { resultsDir, ordersDir, scopesDir, readRunId, verdictsDir, specDir } from "../lib/paths.mjs";
39
+ import { readBoard } from "../compile.mjs";
40
+ import { coveringAcs } from "./requirements.mjs";
39
41
 
40
42
  /** Longest `reason` reported. A deviation is prose written by a worker and can run to paragraphs. */
41
43
  const REASON_MAX = 400;
@@ -253,6 +255,59 @@ export function coverageProblem(cwd, slug, verdict) {
253
255
  `each row is its own criterion, its id in the criterion or its traces_to; ungraded: ${shown}`;
254
256
  }
255
257
 
258
+ /**
259
+ * Whether a verdict anchors any criterion to a requirement when the board says which ones it could.
260
+ *
261
+ * `traces_to` is copied from the `(covers: REQ-…)` clauses of the acceptance criteria a criterion
262
+ * grades, and the requirements matrix is joined along it. Two judges on the same spec and the same
263
+ * instruction differed: one anchored every row and the matrix read 5/5, the next anchored none and it
264
+ * read 0/5 over the same passes. Held only to the all-empty case: which requirement a criterion
265
+ * traces to is the judge's reading, but none at all, beside a board that covers some, is a copy
266
+ * step skipped.
267
+ *
268
+ * @param {string} cwd - Project root.
269
+ * @param {string} slug - Feature slug.
270
+ * @param {object} verdict - The WorkResult's `verdict`.
271
+ * @returns {(string|null)} A reason phrased for the judge, or null.
272
+ */
273
+ export function tracesProblem(cwd, slug, verdict) {
274
+ if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
275
+ const criteria = Array.isArray(verdict.criteria) ? verdict.criteria : [];
276
+ if (!criteria.length) return null;
277
+ if (criteria.some((c) => Array.isArray(c?.traces_to) && c.traces_to.length)) return null;
278
+ let covered = 0;
279
+ try { covered = coveringAcs(readBoard(cwd, slug)).size; } catch { return null; }
280
+ if (!covered) return null;
281
+ return `the verdict anchors none of its ${criteria.length} criteria to a requirement while the board's acceptance ` +
282
+ `criteria cover ${covered} — copy each graded AC's (covers: REQ-…) clause into that criterion's traces_to`;
283
+ }
284
+
285
+ /**
286
+ * A PASS that grades nothing under a dimension its order named.
287
+ *
288
+ * The order names the dimensions the judge grades, and GATE L4 prints them back as what "PASS"
289
+ * covers. Measured: an order named four, the PASS graded criteria under three, and L4 still listed
290
+ * the fourth as evaluated. A PASS must grade at least one criterion under every dimension its order
291
+ * named; a FAIL already stops the ship, so it is not held to this.
292
+ *
293
+ * @param {string} cwd - Project root.
294
+ * @param {string} slug - Feature slug.
295
+ * @param {object} verdict - The WorkResult's `verdict`.
296
+ * @param {{round:(number|null)}} [opts] - The EVAL round whose order names the set; null → no check.
297
+ * @returns {(string|null)} Why the verdict is refused, or null.
298
+ */
299
+ export function dimensionsProblem(cwd, slug, verdict, { round = null } = {}) {
300
+ if (verdict?.overall !== "PASS" || round == null) return null;
301
+ const named = orderDimensions(cwd, slug, round);
302
+ if (!named?.length) return null;
303
+ const criteria = Array.isArray(verdict.criteria) ? verdict.criteria : [];
304
+ const graded = new Set(criteria.map((c) => c?.dimension).filter(Boolean));
305
+ const missing = named.filter((d) => !graded.has(d));
306
+ if (!missing.length) return null;
307
+ return `the PASS grades no criterion under ${missing.join(", ")}, which the order names — grade each named ` +
308
+ "dimension's criteria, or FAIL the one you cannot grade with the reason";
309
+ }
310
+
256
311
  export function citationProblem(cwd, slug, verdict, { round = null } = {}) {
257
312
  if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
258
313
  if (!isScoped(cwd, slug)) return null;
@@ -297,7 +352,8 @@ export function evalVerdict(cwd, slug, round) {
297
352
  ? `the evaluator returned ${status || "no status"}: ${first}`
298
353
  : `status ${status || "unknown"} with no PASS/FAIL verdict`), status);
299
354
  }
300
- const problem = verdictProblem(v) || coverageProblem(cwd, slug, v) || citationProblem(cwd, slug, v, { round });
355
+ const problem = verdictProblem(v) || coverageProblem(cwd, slug, v) || tracesProblem(cwd, slug, v)
356
+ || dimensionsProblem(cwd, slug, v, { round }) || citationProblem(cwd, slug, v, { round });
301
357
  if (problem) return unfit(problem, status, overall);
302
358
  return {
303
359
  found: true,
@@ -309,6 +365,22 @@ export function evalVerdict(cwd, slug, round) {
309
365
  };
310
366
  }
311
367
 
368
+ /**
369
+ * The dimensions round N's evaluate order named — the set the judge was asked to grade.
370
+ *
371
+ * @param {string} cwd - Project root.
372
+ * @param {string} slug - Feature slug.
373
+ * @param {number} round - The EVAL round (`evaluate-r<N>.json`).
374
+ * @returns {(string[]|null)} The order's `payload.dimensions`; null when the order or the list is absent.
375
+ */
376
+ export function orderDimensions(cwd, slug, round) {
377
+ try {
378
+ const o = JSON.parse(readFileSync(join(ordersDir(cwd, slug), `evaluate-r${round}.json`), "utf8"));
379
+ const d = o?.payload?.dimensions;
380
+ return Array.isArray(d) && d.every((x) => typeof x === "string") ? d : null;
381
+ } catch { return null; }
382
+ }
383
+
312
384
  export const ARGV_SPEC = {
313
385
  usage: "harness.mjs probe eval --slug <slug> --round N [--cwd <dir>]",
314
386
  _: { arity: 0, max: 0, name: "(no positional operands)" },
@@ -327,6 +399,7 @@ export function cli(rawArgv) {
327
399
  const args = runArgs(ARGV_SPEC, rawArgv);
328
400
  const cwd = resolve(args.cwd || process.cwd());
329
401
  const { found, overall, status, reason, bug_count, report_path } = evalVerdict(cwd, args.slug, args.round);
330
- console.log(JSON.stringify({ ok: found, overall, bug_count, report_path, round: args.round, status, reason }));
402
+ const dimensions = orderDimensions(cwd, args.slug, args.round);
403
+ console.log(JSON.stringify({ ok: found, overall, bug_count, report_path, round: args.round, status, reason, dimensions }));
331
404
  process.exit(found ? 0 : 1);
332
405
  }
@@ -477,11 +477,16 @@ export function deriveResumeState(cwd, slug, { pluginRoot = null } = {}) {
477
477
  breadboard_source: readReceipt(receipt(cwd, slug))?.breadboard?.source ?? null,
478
478
  spec_folder: hr.spec_folder || null,
479
479
  status: hr.status || null,
480
+ // The terminal status of a close nobody has taken back yet. A relaunch reopens the run before its
481
+ // first gate on this fact, so no gate a relaunch crosses is recorded inside a closed window.
482
+ closed_status: hr.closed_status ? String(hr.closed_status) : null,
480
483
  lens: hr.lens || null,
481
484
  stack: hr.stack || null,
482
485
  run_cmd: hr.run_cmd || null,
483
486
  app_url: hr.app_url || null,
484
- eval_dimensions: Array.isArray(hr.eval_dimensions) ? hr.eval_dimensions : ["spec-conformance"],
487
+ // A list the PO named at L0.5, or [] — `auto`, or a ledger with no line — which the evaluate order
488
+ // resolves from the spec at compile time.
489
+ eval_dimensions: Array.isArray(hr.eval_dimensions) ? hr.eval_dimensions : [],
485
490
  orient_dir: `.shapeup/${slug}/orient/`,
486
491
  has_orient_artifacts: hasOrientArtifacts(cwd, slug),
487
492
  has_spec_tree: hasSpecTree(cwd, slug, hr.spec_folder || null),
@@ -597,6 +602,16 @@ export function setRunStatus(cwd, slug, status) {
597
602
  return { ok: false, path: p, status, reason: `wrote "status: ${status}" but the ledger reads "${after.status}" — the write did not take` };
598
603
  }
599
604
  if (reopened) {
605
+ // THE RUN POINTER COMES BACK WITH THE RUN. The close retired it, and nothing but `init run` wrote
606
+ // it, so a resumed run ran to its next close with no pointer: every hook row it produced carried
607
+ // no run key, and the wall-clock budget had no run to read. Restored only when absent, so a
608
+ // pointer another run holds is left alone.
609
+ try {
610
+ if (!existsSync(activeScope(cwd))) {
611
+ mkdirSync(dirname(activeScope(cwd)), { recursive: true });
612
+ writeFileSync(activeScope(cwd), JSON.stringify({ slug, started_at: fm.started_at ?? null, reopened_at: reopened.reopened_at }, null, 2) + "\n");
613
+ }
614
+ } catch { /* the reopen stands; an unkeyed row is degraded, not wrong */ }
600
615
  // The breadcrumb names a run that is OVER; this one no longer is. Removed only when it names
601
616
  // this run, so another run's close is left alone.
602
617
  try {
@@ -656,6 +671,15 @@ export function deriveLedgerFacts(cwd, slug) {
656
671
  evalRows.sort((a, b) => a.round - b.round);
657
672
  const finalVerdict = evalRows.length ? evalRows[evalRows.length - 1].overall : "not-evaluated";
658
673
  const decisions = [];
674
+ // WHICH LAUNCH TOOK EACH DECISION. A relaunch re-crosses the gates it fast-forwards, and the table
675
+ // listed each of them twice with nothing to tell a second launch from a double sign-off. Every
676
+ // reopen leaves its timestamp in `prior_closes`; a row taken after the n-th one is launch n+1's.
677
+ let reopens = [];
678
+ try {
679
+ const pc = parseFrontmatter(readFileSync(harnessRun(cwd, slug), "utf8")).prior_closes;
680
+ reopens = [...String(pc ?? "").matchAll(/\(reopened ([^)]+)\)/g)].map((m) => m[1]).sort();
681
+ } catch { /* no ledger, no reopen */ }
682
+ const launchOf = (at) => 1 + reopens.filter((t) => typeof at === "string" && at > t).length;
659
683
  const gp = gates(cwd, slug);
660
684
  if (existsSync(gp)) {
661
685
  for (const line of readFileSync(gp, "utf8").split("\n")) {
@@ -663,7 +687,7 @@ export function deriveLedgerFacts(cwd, slug) {
663
687
  try {
664
688
  const g = JSON.parse(line);
665
689
  // Gate rows carry the run key; a row from an earlier run over the same slug is its history.
666
- if (!runId || !g?.run_id || g.run_id === runId) decisions.push(g);
690
+ if (!runId || !g?.run_id || g.run_id === runId) decisions.push(reopens.length ? { ...g, launch: launchOf(g.at) } : g);
667
691
  } catch { /* a torn line proves nothing */ }
668
692
  }
669
693
  }
@@ -733,7 +757,7 @@ function writeCloseLines(body, { status, closedAt, cause, derived = null }) {
733
757
  if (/^final_verdict:.*$/m.test(out)) out = out.replace(/^final_verdict:.*$/m, `final_verdict: ${derived.final_verdict}`);
734
758
  if (typeof derived.rounds_used === "number" && /^rounds_used:.*$/m.test(out)) out = out.replace(/^rounds_used:.*$/m, `rounds_used: ${derived.rounds_used}`);
735
759
  out = rewriteTable(out, /(\| Phase \| Round \| Result \| Duration \| Notes \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, derived.roundRows, "| Init");
736
- const decisionRows = derived.decisions.map((g) => `| ${g.gate ?? "?"} | ${g.decision ?? g.status ?? "?"} | ${g.source ?? "?"} | ${String(g.note ?? "").replace(/\|/g, "\\|").replace(/\s+/g, " ").slice(0, 160)} |`);
760
+ const decisionRows = derived.decisions.map((g) => `| ${g.gate ?? "?"} | ${g.decision ?? g.status ?? "?"} | ${g.source ?? "?"} | ${g.launch > 1 ? `launch ${g.launch} (after a reopen) — ` : ""}${String(g.note ?? "").replace(/\|/g, "\\|").replace(/\s+/g, " ").slice(0, 160)} |`);
737
761
  out = rewriteTable(out, /(\| Gate \| Decision \| Source \| Note \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, decisionRows);
738
762
  }
739
763
  return out;
@@ -37,7 +37,7 @@ import { fileURLToPath } from "node:url";
37
37
  import { validate } from "../verify/envelope.mjs";
38
38
  import { runArgs } from "../lib/argv.mjs";
39
39
  import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId, qaDir } from "../lib/paths.mjs";
40
- import { citationProblem, coverageProblem, verdictProblem } from "../probe/eval.mjs";
40
+ import { citationProblem, coverageProblem, dimensionsProblem, tracesProblem, verdictProblem } from "../probe/eval.mjs";
41
41
  import { huntProblem } from "../probe/hunt.mjs";
42
42
 
43
43
  const HERE = dirname(fileURLToPath(import.meta.url));
@@ -702,6 +702,7 @@ export async function cli(rawArgv) {
702
702
  const evalSlug = String(result.order_id).split("/")[0];
703
703
  const evalRound = Number((String(result.order_id).match(/-r(\d+)$/) || [])[1]) || null;
704
704
  const problem = verdictProblem(result.verdict) || coverageProblem(cwd, evalSlug, result.verdict)
705
+ || tracesProblem(cwd, evalSlug, result.verdict) || dimensionsProblem(cwd, evalSlug, result.verdict, { round: evalRound })
705
706
  || citationProblem(cwd, evalSlug, result.verdict, { round: evalRound });
706
707
  if (problem) {
707
708
  console.error(`ingest-result: result refused — ${problem}.`);
@@ -113,7 +113,7 @@ export function runRow({ receipt, ledger = null, runId = null, rounds = null })
113
113
  max_rounds: num(c.max_rounds),
114
114
  attempt_budget: num(c.attempt_budget),
115
115
  wall_clock_budget_s: num(c.wall_clock_budget_s),
116
- eval_dimensions: Array.isArray(c.eval_dimensions) ? c.eval_dimensions.join(" ") : null,
116
+ eval_dimensions: Array.isArray(c.eval_dimensions) ? c.eval_dimensions.join(" ") : (typeof c.eval_dimensions === "string" ? c.eval_dimensions : null),
117
117
  // Copied from the ledger, never re-derived: the run's own status line is the harness's answer,
118
118
  // and a read plane that recomputed it would be asserting a second one.
119
119
  status: fm.status ?? null,
@@ -2423,7 +2423,7 @@
2423
2423
  "items": {
2424
2424
  "type": "string"
2425
2425
  },
2426
- "description": "spec-evaluator: the active dimension set (caller-resolved precedence). Absent → [spec-conformance] + auto-enable rules."
2426
+ "description": "spec-evaluator: the active dimension set. harness compile fills it on every evaluate order: the run's named set, or the set resolved from the spec with the judge's auto-enable rules. Absent only on a standalone order."
2427
2427
  },
2428
2428
  "run_cmd": {
2429
2429
  "type": "string",
@@ -2586,6 +2586,13 @@
2586
2586
  ],
2587
2587
  "description": "The ledger's stored status, reported for the readers that legitimately hold a MID_RUN set over it (harness reduce snapshot, and the ship report's census). NOT a resume predicate: no phase decision in shapeup-run.js may branch on this field."
2588
2588
  },
2589
+ "closed_status": {
2590
+ "type": [
2591
+ "string",
2592
+ "null"
2593
+ ],
2594
+ "description": "The ledger's closed_status: the terminal status of a close nobody has taken back, null on a run that is open. A relaunch reads it once, to reopen the run before its first gate resolves. Not a phase predicate either."
2595
+ },
2589
2596
  "lens": {
2590
2597
  "type": [
2591
2598
  "string",
@@ -2868,6 +2875,13 @@
2868
2875
  "rounds_used": {
2869
2876
  "type": "integer"
2870
2877
  },
2878
+ "dims_evaluated": {
2879
+ "type": "array",
2880
+ "items": {
2881
+ "type": "string"
2882
+ },
2883
+ "description": "The dimensions the last evaluate order named: the set the PO chose at L0.5, or the one resolved from the spec."
2884
+ },
2871
2885
  "dims_not_evaluated": {
2872
2886
  "type": "array",
2873
2887
  "items": {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.17.1",
3
+ "version": "3.18.0",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -37,7 +37,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
37
37
  |---|---|
38
38
  | `payload.spec_folder` | The committed grading truth: `usecases/` + `domain-model.md` (+ `contracts/`, `scope-summary.md`, `_index.md`). No `usecases/` → HARD STOP, nothing to grade against |
39
39
  | `payload.feature` | Feature slug — scopes the probe and names the report |
40
- | `payload.dimensions[]` | The active dimension set (the caller resolved precedence). Absent → `[spec-conformance]` + the auto-enable rules below |
40
+ | `payload.dimensions[]` | The active dimension set. An orchestrated order always carries it: the set the PO named, or the one resolved from the spec with the auto-enable rules below. Grade exactly this set, at least one criterion under each (a PASS that leaves a named dimension ungraded is refused and sent back once). Absent (standalone only) → apply the rules below yourself |
41
41
  | `payload.run_cmd` | How to start the running app. On a stack where building and launching are different acts it is only the build — prefer `payload.launch_cmd` when the order carries one. Absent standalone → ask; absent orchestrated with no `launch_cmd` either → ESCALATE, do not guess |
42
42
  | `payload.launch_cmd` | The project profile's launch probe: installs the built artifact, starts it and asserts the first screen. This is how the app is brought up for `[ui]` probing. A non-zero exit is a finding to cite, not a reason to try another way. Absent → the profile declares none; fall back to `run_cmd` |
43
43
  | `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
@@ -52,7 +52,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
52
52
  (Steps, Error Cases, Invariants, Test Surface) and `domain-model.md` — never against a task
53
53
  file's own AC paraphrase. Task boards are LOCAL, regenerable bookkeeping the judge never touches.
54
54
 
55
- **Dimension resolution (craft, kept).** Base `[spec-conformance]` + always-on `tdd-surface` +
55
+ **Dimension resolution (standalone; an orchestrated order arrives resolved).** Base `[spec-conformance]` + always-on `tdd-surface` +
56
56
  `integration` (`.be`/`.e2e`); auto-enable `completeness` when any UC has `## Invariants`,
57
57
  `test-surface-conformance` when any UC has `## Test Surface`; an explicit `dimensions[]` list
58
58
  overrides. Each active dimension's file must satisfy `references/dimension-contract.md` — a
@@ -188,7 +188,7 @@ clause yields no anchor — leave the array empty rather than guessing, and neve
188
188
  supply one. This changes nothing you grade: the anchor is a navigation path, never a grading input,
189
189
  and a criterion passes or fails on its evidence exactly as before. It matters downstream because
190
190
  the requirement matrix at GATE L4 and the census at GATE H are projected from these anchors; a
191
- verdict that drops them grades the build and says nothing about what the pitch asked for.
191
+ verdict that drops them grades the build and says nothing about what the pitch asked for. Ingest refuses a verdict that anchors no criterion at all while the board's ACs carry `covers:` clauses, and you are sent back once.
192
192
 
193
193
  **Every FAIL criterion's `evidence` MUST carry a `file:line` locator** — schema-enforced, not
194
194
  advice: the envelope is validated against `work-result.schema.json` at ingest and a locatorless
@@ -16,14 +16,11 @@ what `harness init run` writes. Working around the harness is not an exemption:
16
16
  switch this gate off and no longer does. Loading these instructions is not running them.
17
17
 
18
18
  **Step 1 — open the run.** Write the requirement to a file first, then pass the path — a
19
- multi-line requirement inlined into a shell argument is where this step goes wrong (measured: six
20
- turns fighting shell quoting):
19
+ multi-line requirement inlined into a shell argument is where this step goes wrong. Keep the command
20
+ on ONE line: a grant matches a single-line command, and a `\`-continued one comes back "requires approval":
21
21
 
22
22
  ```bash
23
- node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" init run \
24
- --slug <slug-from-the-request> --intake-file <path/to/the/requirement.md> \
25
- --auto-level <interactive|auto|unattended> \
26
- [--dimensions <a,b>] [--gate-answers <ci|guarded|path.json>] [--wall-clock-budget <seconds>] [--max-rounds 3] [--breadboard <path>]
23
+ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" init run --slug <slug-from-the-request> --intake-file <path/to/the/requirement.md> --auto-level <interactive|auto|unattended> [--dimensions <a,b>] [--gate-answers <ci|guarded|path.json>] [--wall-clock-budget <seconds>] [--max-rounds 3] [--breadboard <path>]
27
24
  ```
28
25
 
29
26
  **After a compaction, or in a fresh session over an open run, re-derive before you act.** One
@@ -102,7 +99,7 @@ failed, launch from the install path and have the operator `/add-dir` the plugin
102
99
  | `status` | What the workflow is telling you | What you do |
103
100
  |---|---|---|
104
101
  | `paused` | A gate resolved "ask" — `paused_at` names it, `block` is composed and ready | Emit `block` **verbatim** (never re-summarise it — that is the paraphrase channel this design exists to close). Put it to the PO, get a decision. Write it to `.shapeup/<slug>/gate-answers.json` (`{"version":1,"preset":"custom","answers":{"<paused_at>":{"decision":"<answer>"}}}`, merging with any prior gate's answer already there). **Relaunch the SAME `Workflow` call, same `args`** — the fast-forward re-derives from disk and re-dispatches nothing already done (verify: `orders/` minus `results/` is empty before it proceeds) |
105
- | `aborted` | A gate resolved "abort", or a hard stop (spec-lint red, scope-hammer CANNOT SHIP) | Report `aborted_at` + `reason` to the PO. Do not relaunch without a human decision — `--force` on `harness init run` if truly restarting |
102
+ | `aborted` | A gate resolved "abort", or a hard stop (spec-lint red, scope-hammer CANNOT SHIP) | Report `aborted_at` + `reason` to the PO. Do not relaunch without a human decision. A relaunch RESUMES the closed run (it reopens, keeping the abort under `prior_closes`); `--force` discards its history and only the PO may ask for it |
106
103
  | `gate_h` | A circuit breaker tripped (`breaker`: outer \| inner \| deadline) — `green_scopes` shipped nothing, `hammer_proposals` needs a census | Dispatch a fresh Agent (model: exec): `Skill(shapeup-sdlc-plugin:scope-hammer) --slug <slug> --breaker <breaker> [--scope <id>]` for the census + cut list, put the PO's decision to `references/gates.md` GATE H, then close out via Step 4 below |
107
104
  | `shipped` | The board's final round passed EVAL, QA ran, GATE H accepted the cut list, `report` names the frozen `shapeup/<slug>/REPORT.md` | Go straight to Step 4 |
108
105
 
@@ -120,7 +117,7 @@ now and RESOLVES the gate (`references/gates.md` GATE L4). Either way, emit:
120
117
  ⏸ GATE L4 — Ship Sign-Off
121
118
  Feature : [slug] — [SHIPPED (deployed) | BUILT & VERIFIED — deploy pending (PO)]
122
119
  Rounds : [rounds_used]
123
- Verdict : [verdict] (dims: [spec-conformance]; not evaluated: [dims_not_evaluated])
120
+ Verdict : [verdict] (dims: [dims_evaluated]; not evaluated: [dims_not_evaluated])
124
121
  QA : [qa_findings] findings | skipped
125
122
  Ledger : harness-run.md
126
123
  ```
@@ -60,10 +60,12 @@ Collect (explicit — never inferred):
60
60
  spec_folder = shapeup/<slug>/spec/ (the deliverable arg passed to ba/eval/exec)
61
61
  L0.3 lens: lite | standard | cross-context (passed to planner at step 8)
62
62
  L0.4 stack hint (e.g. "pnpm, Next 16 web :3000") — aims orient's code-surface sweeps + run commands
63
- L0.5 eval dimensions: default [spec-conformance]; only add if user asks. An added dimension
64
- must reach `harness init run --dimensions <a,b>` — it is recorded in the ledger's
65
- `eval_dimensions:` line and every EVAL order is compiled from there, so a set agreed
66
- in conversation and not passed to the flag grades nothing. Shipped ids:
63
+ L0.5 eval dimensions: default `auto` — each EVAL order resolves the set from the spec
64
+ (spec-conformance + tdd-surface always; integration for .be/.e2e tasks; completeness
65
+ when a UC has Invariants; test-surface-conformance when one has a Test Surface).
66
+ Name a set only if the user asks. It must reach `harness init run --dimensions <a,b>`
67
+ — it is recorded in the ledger's `eval_dimensions:` line and replaces the resolved
68
+ set, so a set agreed in conversation and not passed to the flag grades nothing. Shipped ids:
67
69
  spec-conformance, tdd-surface, integration, completeness, test-surface-conformance
68
70
  (security + performance ship disabled). Whatever is left out is reported at L4 as
69
71
  `dims_not_evaluated` — "shipped" never silently means "verified for all".
@@ -161,7 +163,7 @@ Feature : [slug] (kicked-off pitch: [path])
161
163
  Intake lang : [English | translated via /translator → <name>.en.md]
162
164
  Appetite : [~1 week | ~2 weeks | ~6 weeks | ⚠️ missing — scope uncapped]
163
165
  Spec folder : [path] (lens: [lite|standard])
164
- Eval dims : [spec-conformance] max_rounds: [N, appetite-informed] auto: [interactive|auto|unattended]
166
+ Eval dims : [auto | the named set] max_rounds: [N, appetite-informed] auto: [interactive|auto|unattended]
165
167
  Run commands : [web: ... | api: ... | mobile: ...] (run_cmd → the round build gate, every round before EVAL)
166
168
  Build gate : build_probe [set | —] launch_probe [set | — ⚠ mobile: the install/launch risk has no owner]
167
169
  Model matrix : orch=[model] exec=[model] eval=[model] qa=[model] digester=[script|sonnet] (source: [flags|settings.local|settings.json|default])
@@ -555,7 +557,7 @@ S.7 Export the run's records → one keyed dataset, before the trace is superse
555
557
  ⏸ GATE L4 — Ship Sign-Off
556
558
  Feature : [slug] — [SHIPPED (deployed) | BUILT & VERIFIED — deploy pending (PO)]
557
559
  Rounds : [r] (build+eval cycles)
558
- Verdict : PASS (dims: [spec-conformance]; not evaluated: [security, performance])
560
+ Verdict : PASS (dims: [dims_evaluated]; not evaluated: [dims_not_evaluated])
559
561
  QA : [hunt done — N findings, M promoted+fixed, rest ~ | skipped (--no-qa) | n/a (pre-QA spec)]
560
562
  Requirements: [15/17 PASS · 1 CUT (PO) · 1 no evidence (REQ-12 ← R12) | n/a (no registry)]
561
563
  Ledger : harness-run.md
@@ -692,7 +692,7 @@ type: harness-run
692
692
  feature: [slug]
693
693
  spec_folder: [path to SHARED spec deliverable, e.g. shapeup/<slug>/spec/]
694
694
  lens: lite | standard | cross-context
695
- eval_dimensions: [spec-conformance] # the set from GATE L0.5 (init-run --dimensions); every EVAL order is compiled from THIS line
695
+ eval_dimensions: auto # or the list from GATE L0.5 (init-run --dimensions); `auto` → each EVAL order resolves the set from the spec
696
696
  max_rounds: 3
697
697
  auto_level: interactive | auto | unattended
698
698
  status: orienting | mapping | building | evaluating | shipped | escalated | aborted
@@ -411,7 +411,7 @@ const RESUME = {
411
411
  properties: {
412
412
  intake_path: nullable("string"), spec_folder: nullable("string"), orient_dir: nullable("string"),
413
413
  breadboard_path: nullable("string"), breadboard_source: nullable("string"),
414
- project_profile_path: nullable("string"), status: nullable("string"),
414
+ project_profile_path: nullable("string"), status: nullable("string"), closed_status: nullable("string"),
415
415
  lens: nullable("string"), stack: nullable("string"),
416
416
  run_cmd: nullable("string"), app_url: nullable("string"),
417
417
  eval_dimensions: { type: "array", items: { type: "string" } },
@@ -612,6 +612,8 @@ const EVAL_VERDICT = {
612
612
  bug_count: nullable("integer"),
613
613
  report_path: nullable("string"),
614
614
  round: { type: "integer" },
615
+ // The dimensions round N's evaluate order named — what "PASS" actually covers.
616
+ dimensions: { type: ["array", "null"], items: { type: "string" } },
615
617
  // Why the round holds no verdict it may act on — the evaluator's own first deviation when it
616
618
  // refused, or what is structurally wrong with the verdict it returned. Null when `ok`.
617
619
  status: nullable("string"),
@@ -1360,8 +1362,19 @@ if (!rs) {
1360
1362
  return await withWarnings(aborted("probe", "the fast-forward derivation returned no state — refusing to re-dispatch a run that may already be in progress"));
1361
1363
  }
1362
1364
 
1365
+ // A RELAUNCH OVER A CLOSED RUN REOPENS IT FIRST. The first status write is what takes a close back,
1366
+ // and a launch that fast-forwards every planning phase made that write only at BUILD, so the gates
1367
+ // it crossed on the way were recorded against a run whose ledger still read closed. Bookkeeping,
1368
+ // not a phase decision: the phases below still branch on artifacts alone.
1369
+ if (rs.closed_status) await setRunStatus("orienting", "Orient");
1370
+
1363
1371
  const specFolder = rs.spec_folder || `shapeup/${slug}/spec/`;
1364
- const evalDims = rs.eval_dimensions?.length ? rs.eval_dimensions : ["spec-conformance"];
1372
+ // A set the PO named at L0.5, or null: `harness compile` then resolves it from the spec for each
1373
+ // evaluate order. A fixed fallback here would travel as an explicit list and switch off the judge's
1374
+ // own auto-enable, which is how every run once graded spec-conformance alone.
1375
+ const evalDims = rs.eval_dimensions?.length ? rs.eval_dimensions : null;
1376
+ // What the judge was actually asked to grade, read back off the newest evaluate order.
1377
+ let gradedDims = evalDims || [];
1365
1378
 
1366
1379
  // ---- ORIENT + GATE L1a ----------------------------------------------------------------------
1367
1380
  let spikedArea = "~", spikeResult = "~", riskiest = [];
@@ -1691,6 +1704,7 @@ if (lastEval) {
1691
1704
  const prior = await query(`probe eval --slug ${slug} --round ${lastEval}`, EVAL_VERDICT, "Eval", `ff:eval-r${lastEval}`);
1692
1705
  if (prior?.ok && prior.overall === "PASS") {
1693
1706
  verdict = "pass";
1707
+ if (Array.isArray(prior.dimensions) && prior.dimensions.length) gradedDims = prior.dimensions;
1694
1708
  round = lastEval;
1695
1709
  log(`EVAL — round ${lastEval} already returned PASS on disk, fast-forwarding past the build/eval loop`);
1696
1710
  }
@@ -1949,7 +1963,7 @@ while (verdict !== "pass" && round <= maxRounds) {
1949
1963
  // The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
1950
1964
  // artifact and the project profile, so the judge gets the launch evidence without this script
1951
1965
  // having to carry it.
1952
- payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
1966
+ payload: { ...(evalDims ? { dimensions: evalDims } : {}), run_cmd: rs.run_cmd, round },
1953
1967
  extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself." + note,
1954
1968
  }),
1955
1969
  // The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
@@ -1968,6 +1982,7 @@ while (verdict !== "pass" && round <= maxRounds) {
1968
1982
  return await withWarnings(diedAt("L3", { __failed: `verdict:r${round}: no verdict this round can act on — ${ev.reason || `status ${ev.status || "unknown"}`}` }));
1969
1983
  }
1970
1984
  verdict = ev.overall === "PASS" ? "pass" : "fail";
1985
+ if (Array.isArray(ev.dimensions) && ev.dimensions.length) gradedDims = ev.dimensions;
1971
1986
  findings = e.findings || [];
1972
1987
  }
1973
1988
 
@@ -2088,7 +2103,8 @@ return await withWarnings({
2088
2103
  // human answering that gate the run was graded when it never was.
2089
2104
  verdict,
2090
2105
  rounds_used: round,
2091
- dims_not_evaluated: ALL_DIMS.filter((d) => !evalDims.includes(d)),
2106
+ dims_evaluated: gradedDims,
2107
+ dims_not_evaluated: ALL_DIMS.filter((d) => !gradedDims.includes(d)),
2092
2108
  qa_findings: qaFindings,
2093
2109
  qa: qaState,
2094
2110
  report: REPORT_PATH,