shapeup-sdlc 3.7.1 → 3.7.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.7.1",
4
+ "version": "3.7.3",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -10,6 +10,7 @@ Skills and commands are named short throughout this file; every one of them reso
10
10
  - Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
11
11
  - **The tier-direction rule is refused at the write, not argued at L1b.** A committed file cannot carry a `.shapeup/` path or a `TASK-NNN` id — boards renumber per machine and the run trace is gitignored, so either one resolves only where it was written. Spec-lint has always red'd this at Board Review; the same rule now also refuses the write itself, naming the token, because four producers hit it and one of them hit it twice in one session — a worker carries no lesson across a dispatch, and by L1b the phase that wrote the line is over. The knowledge base is outside the rule: those files are instructions, not references anyone resolves.
12
12
  - Attested, not assumed: a dispatch leaves a receipt naming the skill that ran, and ingest refuses an orchestrated result that has none. A dispatch that fails is answered by the sub-agent improvising the craft, which every other check accepts — so "the artifact exists" is not evidence that the shipped skill produced it. `--no-receipt-check` is the way through when the receipt channel itself fails.
13
+ - **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
13
14
  - GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
14
15
  - Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
15
16
  - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
@@ -34,7 +35,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
34
35
  | Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
35
36
  | Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
36
37
  | Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
37
- | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe + seesaw ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
38
+ | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦; the seesaw regression arm is declared and not yet wired), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
38
39
  | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
39
40
  | FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
40
41
 
@@ -62,7 +63,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
62
63
  - **QA is a level-up, not a gate** — `--no-qa` skips it; circuit breaker outranks the Hunter.
63
64
  - **Role separation** — Evaluator grades, task-executor fixes, QA discovers.
64
65
  - **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
65
- - **Hill phase is mechanical ✦** — derived only from T0/T1/seesaw artifacts, never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
66
+ - **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts (the seesaw arm is declared, not yet wired), never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
66
67
  - **Envelope port (v1.0)** — every dispatch is WorkOrder in / WorkResult out; shared state has exactly one writer (the ingest step); malformed envelopes are hook-denied. Workers: stateless, craft-only, pipeline-blind.
67
68
 
68
69
  ## Setup & Execution
@@ -81,11 +82,11 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
81
82
  through the chained launch, until the run can resume unattended again.
82
83
  - Two storage tiers (ADR-0001): COMMITTED `shapeup/<slug>/` (shaping, spec, scopes, wiring-map, project-profile, requirements, hill, `REPORT.md` frozen at L4) vs GITIGNORED `.shapeup/` (board, orders/results, T0/eval/QA artifacts, ledgers, metrics, gate answers, and the run scripts staged for launch).
83
84
  - **The run launches from a copy inside your project, and it has to.** The Workflow tool loads a script only from a directory the session may already read; the plugin installs outside your project, so naming the shipped path is refused before the run begins and no permission rule repairs it — the grant authorises the tool, not what it may read. Opening a run therefore re-copies the run scripts to `.shapeup/workflows/` and reports the path the launch names. A run in flight keeps the copy it started with: an upgrade reaches the next run, not the current round.
84
- - **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — leaves an order still open by construction, and the pointer is what decides whether that still fences: the fence is enforced only while the pointer exists on disk, so a close that removes it releases the fence whatever the order set says, and restoring the pointer flips it straight back to denying. **Every terminal close retires it** (3.7.1 onward), not only a ship — an escalated or aborted run stops fencing the checkout the moment it closes, and a close that cannot remove a pointer says so in its own return rather than failing. Retiring is not answering, though, and the difference is the whole reason there are still two levers: the pointer says a run is in flight, the missing result says nobody came back, and only the first stops being true at a close. An order the run left unanswered stays genuinely unresolved — un-fenced, but not done — until `init run --force` writes a synthetic result for it, which is also what lets the attempt census keep telling work nobody did apart from work that failed. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
85
- - Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows, agent-call journal rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded — and so does **every other terminal ending**: a run that aborts, escalates or stops at GATE H exports on its close, because the runs whose records are worth most are the ones that never reach the Ship phase. The export is advisory at every one of them: a failed export degrades the trace and is reported in the run's state warnings, but never turns a close into a non-close. A WorkResult carries no `run_id` and reaches it through `order_id`.
85
+ - **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — leaves an order still open by construction, and the pointer is what decides whether that still fences: the fence is enforced only while the pointer exists on disk, so a close that removes it releases the fence whatever the order set says, and restoring the pointer flips it straight back to denying. **Every terminal close retires it** (3.7.1 onward), not only a ship — an escalated or aborted run stops fencing the checkout the moment it closes, and a close that cannot remove a pointer says so in its own return rather than failing. Retiring the pointer does not orphan what follows: the close leaves a breadcrumb naming the run it ended, so the hook decisions taken in the census and ship phase after it still carry that run's key, marked as taken over a closed run. Retiring is not answering, though, and the difference is the whole reason there are still two levers: the pointer says a run is in flight, the missing result says nobody came back, and only the first stops being true at a close. An order the run left unanswered stays genuinely unresolved — un-fenced, but not done — until `init run --force` writes a synthetic result for it, which is also what lets the attempt census keep telling work nobody did apart from work that failed. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
86
+ - Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded — and so does **every other terminal ending**: a run that aborts, escalates or stops at GATE H exports on its close, because the runs whose records are worth most are the ones that never reach the Ship phase. The export is advisory at every one of them: a failed export degrades the trace and is reported in the run's state warnings, but never turns a close into a non-close. A WorkResult carries no `run_id` and reaches it through `order_id`.
86
87
  - Every run projects a **run graph** — `.shapeup/<slug>/graph.jsonl`, append-only, written only by
87
- `reduce graph`. Two families kept separate: work lineage (Run, Order, Result, Verdict, Trial,
88
- GateDecision) and domain (Scope, UseCase, Requirement, Seam). It is derived from the artifacts,
88
+ `reduce graph`. Two families kept separate: work lineage (Run, Order, Result, Leg, Verdict,
89
+ Trial, GateDecision) and domain (Scope, UseCase, Requirement, Seam). It is derived from the artifacts,
89
90
  never authored, so it can be deleted and rebuilt — and a run recorded before it existed backfills
90
91
  on first touch. `--subgraph run` is the fast-forward as one bounded query; `--trace <node>` walks
91
92
  a verdict back to the objective, the plan, the source, the execution record and the gate that
package/README.md CHANGED
@@ -41,9 +41,17 @@ runs over unfinished tasks rather than denying it. The board is local to the mac
41
41
  harness — see [ADR-0001](docs/design/adr/0001-consumer-file-organization.md).)
42
42
 
43
43
  **2. Progress is measured, not claimed.** A scope counts as built only when `t0-verify` runs
44
- its fixtures, a DB probe, and the seesaw, and writes an artifact to disk. The evaluator must
45
- cite that artifact and re-hashes it itself; hill phase is derived from those facts, so no
46
- worker can self-report confidence.
44
+ its fixtures and its DB probe and writes an artifact to disk — with each command's exit code,
45
+ its captured output, and whether it ran at all. The evaluator must cite that artifact, and the
46
+ hill phase is derived from artifacts rather than from a worker's own account of its progress.
47
+ Two limits, stated here because the point of this section is that a claim without a mechanism
48
+ behind it is the thing this harness exists to prevent: the **seesaw** regression arm is declared
49
+ and not yet wired (no run writes its registry — wiring it is an open Betting Table decision), and
50
+ the citation **re-hash** the kernel performs proves self-consistency, not provenance — the digest
51
+ and the cited verdict are checked, the scope, round and run the artifact belongs to are not. A T0
52
+ artifact is also evidence about the machine that produced it: it records the tree and nothing
53
+ about the toolchain or caches the commands resolved through. All three are open items in
54
+ `shapeup/knowledge-base/harness-defects.md`, not shipped guarantees.
47
55
  → *Prevents: "done" asserted with nothing behind it.*
48
56
 
49
57
  **3. Parallel work can't corrupt shared state.** Each scope gets a write-whitelist of files
@@ -122,8 +130,8 @@ rest of this README after this table and nothing will be a surprise.
122
130
  |---|---|
123
131
  | **board** | The round's task list. "Green" means every task is done. GATE L2's hook reads this before an evaluation and warns if it is not green. |
124
132
  | **round** | One build → evaluate cycle. A FAIL verdict starts round *r+1*. |
125
- | **T0** | The smoke test a scope must pass before it counts as built: its fixtures + a DB probe + the seesaw. Writes an artifact to disk that the evaluator must cite. |
126
- | **seesaw** | The part of T0 that re-runs *other* scopes' fixtures — so a regression is never mistaken for progress. |
133
+ | **T0** | The smoke test a scope must pass before it counts as built: its fixtures + a DB probe (the seesaw arm is declared and not yet wired — see §2 above). Writes an artifact to disk that the evaluator must cite. |
134
+ | **seesaw** | The regression arm of T0, meant to re-run *other* scopes' fixtures so a regression is never mistaken for progress. Declared, not yet wired: no run writes its registry, so no run has executed it. |
127
135
  | **substrate** | The exact list of files one dispatch is allowed to write, stamped into its work order. A hook blocks anything outside it — and anything the order marks frozen. |
128
136
  | **scope contract** | The file defining one vertical slice: its substrate, its fixtures, its affordances. |
129
137
  | **affordance** | The thing a user can actually click, type or call. UI is graded on affordances, not on looks. |
@@ -178,7 +186,7 @@ every arm is skipped when its artifact is absent, so older specs are unaffected.
178
186
  | QA (post-PASS) | `qa-edge-hunter` | v1.1 | Exploratory edge hunt on the running app through six fixed lenses, charting edges *outside* what the evaluator probed. Findings go to the ledger as `~`; never blocks ship. |
179
187
  | Stop (11) | `scope-hammer` | v0.1 | GATE H: must-have census → baseline comparison (never vs. the ideal) → cut list + ship verdict. Handles the normal stop and both circuit-breaker triggers. |
180
188
  | Retro (post-L4) | `coach` | — | RLHF for the harness: turns raw PO/TL feedback at Ship Sign-off into per-skill guidelines under committed `shapeup/knowledge-base/<skill>.md`, read back by six coachable workers on their next run and by `tech-lead` at GATE L0 (workflow guidance, never a gate answer). `--scan` seeds the same files from the project on disk before the first run; `--research <stack>` seeds them from the platform's official documentation when the project has nothing to scan, and cross-checks a scan's rules when it has. GATE COACH-1 asks the PO which skill owns each rule — never assumes; mechanism defects are filed to the harness-defect register instead. |
181
- | Orchestrator | `tech-lead` | v1.0 | Owns the run end-to-end: PLAN once → BUILD all tasks → EVAL once per round, looping on FAIL. Three-level circuit breaker (rounds / T0 attempts / wall clock), T0/seesaw-verified build rounds, mechanical hill derivation. Sole writer of run-state. |
189
+ | Orchestrator | `tech-lead` | v1.0 | Owns the run end-to-end: PLAN once → BUILD all tasks → EVAL once per round, looping on FAIL. Three-level circuit breaker (rounds / T0 attempts / wall clock), T0-verified build rounds, mechanical hill derivation. Sole writer of run-state. |
182
190
 
183
191
  ### Commands
184
192
 
@@ -299,8 +307,10 @@ These hold across the harness and are the reason it stays predictable:
299
307
  of blocking the round. An opt-in third breaker bounds the **wall clock**, because the other two
300
308
  count events and neither can notice a single round running for half an hour — tripping it routes
301
309
  to GATE H, so a run out of time ships what is green instead of being killed and shipping nothing.
302
- - **Hill phase is mechanical, never self-reported** — derived only from T0/T1/seesaw facts, closing
303
- the self-reported-confidence risk outright.
310
+ - **Hill phase is mechanical, never self-reported** — derived only from T0/T1 facts (the seesaw
311
+ arm is declared, not yet wired), closing the self-reported-confidence risk. A scope with no
312
+ discovery ledger derives no phase rather than a solved one, so absence no longer reads as
313
+ progress on that arm.
304
314
  - **One writer per shared file** — every board/ledger/verdict write goes through
305
315
  `harness reduce ingest`; workers return data and never touch shared state.
306
316
  - **Traceability is oracle-checked, opt-in** — `harness verify trace` verifies covers-closure and
package/SECURITY.md CHANGED
@@ -45,7 +45,7 @@ test against machines you don't own.
45
45
  (override channel fails closed), and every exercised override is logged. The same principle
46
46
  covers `.shapeup/active-order`, which `sandbox-guard` reads to find the run whose orders fence
47
47
  a worker's writes: it sits outside the run-trace carve-out, so a worker cannot repoint its own
48
- sandbox. Every hook resolves that pointer, the decision ledger and the substrate globs against
48
+ sandbox. The `last-run` breadcrumb a terminal close leaves beside it is read for one purpose — keying a decision row to the run that just ended, marked `run_closed` — and arms nothing: liveness is derived from the order set and the pointer, never from that file. Every hook resolves that pointer, the decision ledger and the substrate globs against
49
49
  the project root it finds above the tool call's working directory (a run pointer, the committed
50
50
  tier, or a git boundary) — a worker that `cd`s into a sub-folder is fenced exactly as one at the
51
51
  top, and its receipts land in the same ledger.
@@ -72,7 +72,7 @@ sitting, and reading them is the recommended review.
72
72
  | [`safety-spine.mjs`](hooks/safety-spine.mjs) | PreToolUse (`Bash\|Read\|Write\|Edit\|MultiEdit`) | The proposed command/path; `.shapeup/safety-overrides.json` | Yes — provably destructive ops only: `rm -rf` on unrecoverable targets, `git push --force` / push to main, `git reset --hard`, `git clean -fdx`, `DROP TABLE`/`TRUNCATE`, reads of `.env`/keys/cloud credentials, and any write to its own overrides file | Never blocks an unmatched command; `--force-with-lease` stays allowed |
73
73
  | [`gate-intake.mjs`](hooks/gate-intake.mjs) | PreToolUse (`Skill`) | The `tech-lead` dispatch's own arguments | Yes — an orchestrator dispatch carrying no resolvable intake (no pitch, spec, resume or requirement text) | Fails open on `--order` and on any ambiguous arg shape |
74
74
  | [`harness verify envelope`](kernel/verify/envelope.mjs) | PreToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the JSON schemas | Yes — a worker dispatch whose order file is missing or schema-invalid | Never gates a dispatch that carries no `--order` (standalone skill use stays free) |
75
- | [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
75
+ | [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path, resolved through the filesystem before it is compared — symlinks followed, the unborn tail re-attached to its real ancestor, case folded where the filesystem folds case — so a spelling variant of a frozen file is the frozen file; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
76
76
  | [`tier-guard.mjs`](hooks/tier-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the text the call is about to write (`content`, `new_string`, or any edit in a `MultiEdit` batch) | Yes — a write into `shapeup/<slug>/` whose content names a `.shapeup/` path or a `TASK-NNN` board id: the same tier-direction rule spec-lint reds at GATE L1b, refused at the moment of writing, with the offending token quoted back | No-op outside `shapeup/<slug>/`, on file forms the lint does not scan, and on `shapeup/knowledge-base/` (instructions to a worker, not references). Reads nothing off disk and writes only its decision row; it never inspects a file it is not being asked to write |
77
77
  | [`dispatch-receipt.mjs`](hooks/dispatch-receipt.mjs) | PostToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the tool result's own report of which skill ran | **No — it has no deny path at all.** It records that the shipped skill ran, so `harness reduce ingest` can refuse a result no dispatch produced | Never writes an attestation for a result that does not name a resolved skill; never fails the call it observes (every write is inside `try`/`catch`) |
78
78
  | [`gate-zerowork.mjs`](hooks/gate-zerowork.mjs) | Stop | Run receipts on disk; the session transcript; the decision ledger | **Yes — the one blocking hook.** Returns `decision:"block"` when the session dispatched the orchestrator and produced no run receipt | Defers the moment any receipt exists; `stop_hook_active` caps it at one block per stop chain |
@@ -49,7 +49,7 @@
49
49
  import { appendFileSync, mkdirSync, existsSync, statSync } from "node:fs";
50
50
  import { dirname, resolve } from "node:path";
51
51
  import { decisions, activeScope, sharedDir } from "../../kernel/lib/paths.mjs";
52
- import { resolveRunId } from "../../kernel/lib/paths.mjs";
52
+ import { resolveRun } from "../../kernel/lib/paths.mjs";
53
53
 
54
54
  /**
55
55
  * The project root a hook should file under, from wherever the tool call happened to fire.
@@ -221,7 +221,15 @@ export async function runHook(name, fn) {
221
221
  // run, and recording that is what lets the export tier partition ambient decisions from run
222
222
  // ones. Resolution reads two small files and swallows every error — a receipt must never be
223
223
  // able to fail a tool call.
224
- run_id: (() => { try { return resolveRunId(projectRoot(d.cwd || process.cwd())); } catch { return null; } })(),
224
+ // A row written after the run's terminal close still keys to that run — the close leaves a
225
+ // breadcrumb for exactly this stretch — and says so, because a decision taken over a closed run
226
+ // is a different fact from one taken inside it.
227
+ ...(() => {
228
+ try {
229
+ const { run_id, source } = resolveRun(projectRoot(d.cwd || process.cwd()));
230
+ return source === "closed" ? { run_id, run_closed: true } : { run_id };
231
+ } catch { return { run_id: null }; }
232
+ })(),
225
233
  event: d.event ?? null,
226
234
  tool: d.tool ?? null,
227
235
  subject: d.subject ?? null,
@@ -85,8 +85,8 @@
85
85
  // Contract: PreToolUse stdin JSON { tool_name, tool_input:{file_path | edits[].file_path}, cwd }.
86
86
  // Deny via { hookSpecificOutput: { hookEventName, permissionDecision:"deny", permissionDecisionReason } }.
87
87
 
88
- import { readFileSync, existsSync, appendFileSync, mkdirSync, readdirSync, statSync } from "node:fs";
89
- import { resolve, join, relative, dirname, sep } from "node:path";
88
+ import { readFileSync, existsSync, appendFileSync, mkdirSync, readdirSync, statSync, realpathSync, lstatSync, readlinkSync } from "node:fs";
89
+ import { resolve, join, relative, dirname, basename, sep } from "node:path";
90
90
  import { isMain } from "../kernel/lib/argv.mjs";
91
91
  import { LOCAL, SHARED, activeOrder, ordersDir, resultsDir, metricsShard } from "../kernel/lib/paths.mjs";
92
92
  import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
@@ -115,8 +115,69 @@ export function globToRegExp(glob) {
115
115
  return new RegExp(`^${re}$`);
116
116
  }
117
117
 
118
- export function matchesAny(relPath, globs) {
119
- return (globs || []).some((g) => globToRegExp(g).test(relPath));
118
+ /**
119
+ * Does `relPath` match any of `globs`? With `fold` set, both sides are compared case-folded — the
120
+ * caller passes it when the filesystem underneath does not distinguish case, so that a glob
121
+ * declared as `receipts/**` also covers the spelling `Receipts/**`, which on that filesystem is the
122
+ * same directory.
123
+ */
124
+ export function matchesAny(relPath, globs, fold = false) {
125
+ const path = fold ? relPath.toLowerCase() : relPath;
126
+ return (globs || []).some((g) => globToRegExp(fold ? g.toLowerCase() : g).test(path));
127
+ }
128
+
129
+ // --- the path a write actually lands on ---------------------------------------------------------
130
+ //
131
+ // A substrate glob is a SPELLING; what it protects is a FILE. Measured against a real compiled
132
+ // order: `.shapeup/<slug>/receipts/dispatch.jsonl` was denied as frozen, `Receipts/dispatch.jsonl`
133
+ // was permitted, and on this machine's case-insensitive filesystem the second spelling overwrote
134
+ // the first file. A symlink was the same hole from another side: `src/x -> ../.shapeup/<slug>/legs.jsonl`
135
+ // sits inside an allowed `src/**` and resolves `..` without following the link, so the write went
136
+ // through. No list of globs can close either — case, links, Unicode forms are all further spellings
137
+ // of one target — so the comparison is made on the resolved real path instead, and case-folded when
138
+ // the filesystem itself folds case. The globs stay exactly as the compiler wrote them.
139
+
140
+ /**
141
+ * Resolve `abs` through the filesystem: symlinks in any existing ancestor are followed (a dangling
142
+ * link is followed by reading it), and the part that does not exist yet is re-attached to the real
143
+ * path of the deepest ancestor that does. A path nothing on disk can resolve is returned as given.
144
+ *
145
+ * @param {string} abs - Absolute path as the tool named it.
146
+ * @param {number} [depth] - Recursion guard for link chains.
147
+ * @returns {string} The path the write would actually reach.
148
+ */
149
+ export function realPathOf(abs, depth = 0) {
150
+ if (depth > 40) return abs;
151
+ let cur = abs;
152
+ const tail = [];
153
+ for (let i = 0; i < 128; i++) {
154
+ try { return join(realpathSync.native(cur), ...[...tail].reverse()); } catch { /* not yet resolvable as a whole */ }
155
+ try {
156
+ if (lstatSync(cur).isSymbolicLink()) {
157
+ const target = resolve(dirname(cur), readlinkSync(cur));
158
+ return realPathOf(join(target, ...[...tail].reverse()), depth + 1);
159
+ }
160
+ } catch { /* does not exist at all — climb */ }
161
+ const parent = dirname(cur);
162
+ if (parent === cur) return abs;
163
+ tail.push(basename(cur));
164
+ cur = parent;
165
+ }
166
+ return abs;
167
+ }
168
+
169
+ /**
170
+ * Does the filesystem under `dir` fold case? Decided by asking it: the directory's own real path,
171
+ * with every letter's case swapped, exists and resolves back to the same real path only where the
172
+ * filesystem does not distinguish case. A path with no letters to swap answers "no".
173
+ *
174
+ * @param {string} dir - An existing directory (the project root).
175
+ * @returns {boolean} True on a case-insensitive filesystem.
176
+ */
177
+ export function fsFoldsCase(dir) {
178
+ const swapped = dir.replace(/[a-zA-Z]/g, (c) => (c === c.toLowerCase() ? c.toUpperCase() : c.toLowerCase()));
179
+ if (swapped === dir) return false;
180
+ try { return existsSync(swapped) && realpathSync.native(swapped) === realpathSync.native(dir); } catch { return false; }
120
181
  }
121
182
 
122
183
  function readJSON(p) {
@@ -281,9 +342,15 @@ async function main() {
281
342
  const blockReasons = [];
282
343
  let frozenHits = 0;
283
344
 
345
+ // Compare where the write lands, not how it was spelled (see `realPathOf`). The root is resolved
346
+ // the same way so a checkout reached through a linked directory still yields a relative path.
347
+ const realRoot = realPathOf(resolve(root));
348
+ const fold = fsFoldsCase(realRoot);
349
+ const under = (rel, prefix) => (fold ? rel.toLowerCase().startsWith(prefix.toLowerCase()) : rel.startsWith(prefix));
350
+
284
351
  for (const raw of targetPaths) {
285
- const abs = resolve(cwd, raw);
286
- const rel = relative(root, abs);
352
+ const abs = realPathOf(resolve(cwd, raw));
353
+ const rel = relative(realRoot, abs);
287
354
 
288
355
  // Frozen takes absolute precedence, and it is checked across EVERY live contract: a path one
289
356
  // scope froze stays frozen while another scope is in flight, which is the whole point of
@@ -296,7 +363,7 @@ async function main() {
296
363
  // a planner is graded against — so the compiler emitted a declaration with no enforcer, which is
297
364
  // the exact state this hook exists to end. A path a live contract freezes is a violation
298
365
  // wherever it lives.
299
- const freezer = contracts.find((c) => matchesAny(rel, c.frozen));
366
+ const freezer = contracts.find((c) => matchesAny(rel, c.frozen, fold));
300
367
  if (freezer) {
301
368
  violations.push(rel);
302
369
  frozenHits++;
@@ -304,11 +371,11 @@ async function main() {
304
371
  continue;
305
372
  }
306
373
 
307
- if (rel.startsWith(runTracePrefix)) continue;
374
+ if (under(rel, runTracePrefix)) continue;
308
375
 
309
- if (contracts.some((c) => matchesAny(rel, c.allowed))) continue; // inside a live contract
376
+ if (contracts.some((c) => matchesAny(rel, c.allowed, fold))) continue; // inside a live contract
310
377
 
311
- if (contracts.some((c) => matchesAny(rel, c.appendOnly))) {
378
+ if (contracts.some((c) => matchesAny(rel, c.appendOnly, fold))) {
312
379
  if (p.tool_name === "Write") {
313
380
  violations.push(rel);
314
381
  blockReasons.push(`${rel} is append-only (Write overwrites, use Edit)`);
@@ -332,7 +399,7 @@ async function main() {
332
399
  // spec artifacts, so widening a substrate to reach one is the wrong move in a plausible-looking
333
400
  // direction. Those files belong to the orchestrator, whose write window is a phase boundary —
334
401
  // no dispatch in flight — and never the middle of somebody else's dispatch.
335
- const committed = violations.filter((v) => v.split(/[\\/]/)[0] === SHARED);
402
+ const committed = violations.filter((v) => (fold ? v.split(/[\\/]/)[0].toLowerCase() === SHARED.toLowerCase() : v.split(/[\\/]/)[0] === SHARED));
336
403
  const hint = committed.length === violations.length
337
404
  ? `${SHARED}/ is committed tier: these belong to the orchestrator, not to a worker substrate. `
338
405
  + "Write them at a phase boundary, with no dispatch in flight — do not widen an order to reach one."
@@ -218,6 +218,32 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
218
218
  // in its `allowed` globs, so only a `frozen` entry denies the write. `translate` is the one
219
219
  // operation that legitimately rewrites a pitch, and it writes the COMMITTED copy, not this one.
220
220
  const FROZEN_INTAKE = [`${local}/intake.md`, `${local}/breadboard.md`];
221
+ // THE LEG'S OWN ATTESTATION. A dispatch receipt, a leg-completion row and a T0 verdict exist to
222
+ // answer "did a worker actually run, and what did it measure" from evidence the leg being judged
223
+ // does not control — the whole reason to prefer them over a compiled order, which is writable by
224
+ // that same hand. None of the three is ever produced by this assistant's own edit tool in normal
225
+ // operation: a receipt is stamped by the dispatch hook reacting to the call itself, a
226
+ // leg-completion row by the step that reads a landed result, and a verdict by the command that
227
+ // runs the fixtures. A build leg reaching for `Edit`/`Write` on any of them is not doing its job
228
+ // by another route — it is handing itself the grade — so freezing them costs a well-behaved leg
229
+ // nothing and removes the one channel through which it could attest work it did not do. The
230
+ // carve-out below still covers everything else under this root — the doer's own task board and
231
+ // discovery ledger — because neither lives under any of these three.
232
+ //
233
+ // THE RESULT ENVELOPE IS DELIBERATELY NOT ON THIS LIST, and the reasoning is worth keeping because
234
+ // it looks like it belongs. It is the leg's own claim about its own work, so on the argument above
235
+ // it is the first thing you would freeze. But a result is not evidence ABOUT the leg, it is the
236
+ // leg's PRODUCT — the other half of the envelope port, and the thing that answers the order. The
237
+ // order is unanswered at that moment by construction, so freezing the path would deny every build
238
+ // leg its documented last step, every time, on the first round: measured end to end against a
239
+ // compiled order, with the denial telling the worker to widen a substrate that cannot lift a
240
+ // frozen entry. Nothing else writes an ordinary result either, so there is no fallback. The census
241
+ // that reads it is already built for this: a result alone attests nothing, and can only turn an
242
+ // ALREADY-receipted attempt into a spent one, which spends the forger's own budget. What that does
243
+ // not cover is a leg forging a SIBLING's result, which is a real hole with a different fix — the
244
+ // glob form here cannot say "every result except this order's own", so closing it needs a
245
+ // mechanism rather than one more entry on this list.
246
+ const FROZEN_ATTESTATION = [`${local}/receipts/**`, `${local}/legs.jsonl`, `${local}/t0/verdicts/**`];
221
247
  switch (operation) {
222
248
  case "execute": case "fix": case "spike":
223
249
  // Build legs are the widest window on FROZEN_INTAKE, not an exemption from it: they are the
@@ -228,7 +254,7 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
228
254
  return {
229
255
  allowed: [...(scope?.allowed_file_substrate || []), `${local}/spikes/**`],
230
256
  shared: scope?.shared_substrate || [],
231
- frozen: [...FROZEN_INTAKE],
257
+ frozen: [...FROZEN_INTAKE, ...FROZEN_ATTESTATION],
232
258
  };
233
259
  case "analyze":
234
260
  return { allowed: [`${spec}/**`, `${local}/**`], frozen: [...FROZEN_INTAKE] };
package/kernel/gate.mjs CHANGED
@@ -110,7 +110,7 @@ export const PRESETS = {
110
110
  "L1a": { decision: "proceed", note: "Orient review — advisory read." },
111
111
  "L1a.5": { decision: "proceed", note: "Wiring review — checked by trace-lint." },
112
112
  "L1b": { decision: "ask", note: "Board review is where scope is actually decided. Not pre-approvable." },
113
- "L2": { decision: "proceed", note: "Board-green is verified by hook, not by opinion." },
113
+ "L2": { decision: "proceed", note: "The board facts travel in the gate block itself — green_scopes and hammer_proposals — so a preset answering here is not answering blind. Note there is no board check behind this: the L2 hook was retired into the gate block in v2.0." },
114
114
  "L3": { decision: "loop", max_rounds: 3, note: "Loop on FAIL; the breaker ends it." },
115
115
  "QA": { decision: "run" },
116
116
  "H": { decision: "ask", note: "The cut list changes what ships." },
@@ -293,6 +293,19 @@ export const activeScope = (cwd) => join(localDir(cwd), "active-scope");
293
293
  */
294
294
  export const activeOrder = (cwd) => join(localDir(cwd), "active-order");
295
295
 
296
+ /**
297
+ * The run a terminal close just ended — `{slug, run_id, closed_at, closed_status}`, written by the
298
+ * close in the same act that retires the run pointers, and superseded by the next close.
299
+ *
300
+ * Why it exists: retiring the pointers is right (a finished run must stop fencing the checkout),
301
+ * but the close lands BEFORE the stretch that decides the most — the scope-hammer census, the ship
302
+ * phase, the export — and every hook decision in that stretch resolved its run through the pointer
303
+ * that was just removed. Measured on a live run: every decision row after `closed_at` carried
304
+ * `run_id: null`, the census dispatch included. The breadcrumb keeps that join without keeping the
305
+ * fence: it names a run, it never arms anything, and a row keyed through it says so.
306
+ */
307
+ export const lastRun = (cwd) => join(localDir(cwd), "last-run");
308
+
296
309
  /**
297
310
  * Where the run scripts are staged for launch, inside the project.
298
311
  *
@@ -531,9 +544,31 @@ export function readRunId(cwd, slug) {
531
544
  * @returns {(string|null)} The run key, or null when no run is active or readable.
532
545
  */
533
546
  export function resolveRunId(cwd, slug = null) {
534
- if (slug) return readRunId(cwd, slug);
547
+ return resolveRun(cwd, slug).run_id;
548
+ }
549
+
550
+ /**
551
+ * The run key AND where it came from — the shape a ledger row needs when the distinction matters.
552
+ *
553
+ * Resolution order: the slug the caller knows (`slug`); else the `active-scope` pointer a live run
554
+ * publishes (`pointer`); else the breadcrumb the last terminal close left (`closed`), so the census
555
+ * and ship phase that follow a close still key to the run they belong to; else `null`, which is
556
+ * "no run" and a real answer. A key resolved through the breadcrumb is a key to a run that is
557
+ * OVER: the caller records that alongside it rather than presenting the two the same way.
558
+ *
559
+ * @param {string} cwd - Project root.
560
+ * @param {(string|null)} [slug=null] - Feature slug when the caller already knows it.
561
+ * @returns {{run_id:(string|null), source:("slug"|"pointer"|"closed"|null)}}
562
+ */
563
+ export function resolveRun(cwd, slug = null) {
564
+ if (slug) return { run_id: readRunId(cwd, slug), source: "slug" };
535
565
  try {
536
566
  const ptr = JSON.parse(readFileSync(activeScope(cwd), "utf8"));
537
- return ptr?.slug ? readRunId(cwd, ptr.slug) : null;
538
- } catch { return null; }
567
+ if (ptr?.slug) return { run_id: readRunId(cwd, ptr.slug), source: "pointer" };
568
+ } catch { /* no live run — fall through to the breadcrumb */ }
569
+ try {
570
+ const crumb = JSON.parse(readFileSync(lastRun(cwd), "utf8"));
571
+ if (typeof crumb?.run_id === "string" && crumb.run_id) return { run_id: crumb.run_id, source: "closed" };
572
+ } catch { /* no close on record either */ }
573
+ return { run_id: null, source: null };
539
574
  }
@@ -33,6 +33,7 @@
33
33
 
34
34
  import { existsSync, readFileSync, readdirSync } from "node:fs";
35
35
  import { join, resolve } from "node:path";
36
+ import { createHash } from "node:crypto";
36
37
  import { runArgs } from "../lib/argv.mjs";
37
38
  import { resultsDir, scopesDir } from "../lib/paths.mjs";
38
39
 
@@ -59,6 +60,58 @@ export function isScoped(cwd, slug) {
59
60
  catch { return false; }
60
61
  }
61
62
 
63
+ /**
64
+ * Why one T0 citation does not resolve, or null when the kernel can positively confirm it does.
65
+ *
66
+ * Re-checks the two facts a citation can lie about and the schema promises are enforced
67
+ * (`T0Citation`: "the evaluator RECOMPUTES sha256 from disk — a handed hash is never trusted"):
68
+ * that the bytes at `path` hash to the cited `sha256`, and that what those bytes actually say is a
69
+ * green T0 verdict — a PASS/FAIL cannot ride on an artifact that itself recorded red. NOT checked:
70
+ * whether `path` is one of the order's own `payload.t0_artifacts` (see the note on
71
+ * {@link citationProblem} for why that is left to the operator rather than enforced here).
72
+ *
73
+ * FAILS OPEN ON "CANNOT TELL", CLOSED ON "PROVEN WRONG" — and the two are not the same fact. A read
74
+ * that fails for a reason that says something DEFINITE about what is (or isn't) at `path` is proof,
75
+ * same as a hash mismatch: `ENOENT` (nothing there) and `EISDIR` (a directory, never a file — T0
76
+ * verdicts are always files, per `writeArtifact`) both mean no such artifact was ever produced, so
77
+ * both refuse. `verify t0` writes verdict artifacts immutably and never deletes one
78
+ * (`writeArtifact`'s `wx` flag — see kernel/verify/t0.mjs), so that absence or shape mismatch is a
79
+ * positive fact about the citation, not a guess. Anything else a read can fail with — permission
80
+ * denied, a symlink loop, a transient I/O error — says nothing about the citation's honesty, only
81
+ * that THIS MACHINE could not check it just now, so it is not refused on that ground alone: the
82
+ * fail-open discipline this repo's guards already use elsewhere for an unproven bad state.
83
+ *
84
+ * @param {string} cwd - Project root.
85
+ * @param {object} citation - One `T0Citation` (`scope_id`, `path`, `sha256`). Read defensively:
86
+ * `reduce ingest` only reaches this after the enclosing WorkResult passed schema validation, but
87
+ * `probe eval` reaches it over a result file it merely `JSON.parse`s, so a malformed citation must
88
+ * fail this check rather than throw.
89
+ * @returns {(string|null)} A reason phrased for an operator, or null when the file at `path` exists,
90
+ * hashes to the cited `sha256`, and its own `overall` reads "green".
91
+ */
92
+ function unresolvedCitation(cwd, citation) {
93
+ const rel = typeof citation?.path === "string" ? citation.path : "";
94
+ if (!rel) return "names no artifact path";
95
+ let text;
96
+ try {
97
+ text = readFileSync(resolve(cwd, rel), "utf8");
98
+ } catch (e) {
99
+ if (e.code === "ENOENT") return `cites ${rel}, which does not exist on disk`;
100
+ if (e.code === "EISDIR") return `cites ${rel}, which is a directory, not a T0 verdict file`;
101
+ // EACCES, ELOOP, EIO, EMFILE… — this machine failing to look, not evidence against the
102
+ // citation, so it is not refused on that ground.
103
+ return null;
104
+ }
105
+ const actual = createHash("sha256").update(text).digest("hex");
106
+ const claimed = typeof citation.sha256 === "string" ? citation.sha256.toLowerCase() : "";
107
+ if (actual !== claimed) return `cites ${rel} with sha256 ${citation.sha256}, but the file on disk hashes to ${actual}`;
108
+ let body;
109
+ try { body = JSON.parse(text); }
110
+ catch { return `cites ${rel}, whose bytes match the hash but do not read as a T0 verdict`; }
111
+ if (body?.overall !== "green") return `cites ${rel}, whose own verdict is "${body?.overall ?? "unknown"}", not green`;
112
+ return null;
113
+ }
114
+
62
115
  /**
63
116
  * Why a verdict cannot stand as its round's judgement on T0 grounds, or null when it can.
64
117
  *
@@ -68,21 +121,40 @@ export function isScoped(cwd, slug) {
68
121
  * and branched on like any other. It is checked here so the round loop, the resume derivation, the
69
122
  * hill and ingest all refuse the same verdict for the same reason.
70
123
  *
71
- * PRESENCE, NOT HASHES. The evaluator re-hashes what it cites; a slip transcribing a digest is not
72
- * evidence the verdict is wrong, and refusing a round over one would cost a whole re-evaluation.
124
+ * RESOLVES, NOT JUST PRESENT. A citation naming a path and a hash used to be taken on faith: any
125
+ * non-empty `t0_citations[]` passed, whatever it pointed at. Measured live: a PASS citing a path
126
+ * that does not exist on disk, and a PASS citing a real artifact whose own verdict was red, both
127
+ * ingested clean — presence stood in for a re-hash the schema had already promised. Each citation is
128
+ * now re-checked through {@link unresolvedCitation}.
129
+ *
130
+ * WHAT THIS DOES NOT CHECK: whether a citation is drawn from the order's own `payload.t0_artifacts`
131
+ * list. That would need the compiled order, which this function's callers do not equally have —
132
+ * `reduce ingest` holds it, but `probe eval` (and the round-loop/resume/hill readers behind it)
133
+ * knows only (cwd, slug, round), and a scope's T0 attempt can legitimately go green again LATER than
134
+ * whatever list was frozen at compile time (`greenVerdict` already treats "newest green" as
135
+ * authoritative for exactly this reason — see kernel/probe/t0.mjs). Enforcing membership only where
136
+ * the order happens to be on hand would let one channel refuse a citation the other accepts, for
137
+ * evidence that may simply be fresher than the order — worse than leaving it unenforced.
73
138
  *
74
139
  * @param {string} cwd - Project root.
75
140
  * @param {string} slug - Feature slug.
76
141
  * @param {object} verdict - The WorkResult's `verdict` block.
77
- * @returns {(string|null)} The problem, phrased for an operator; null for a cited verdict, an
78
- * unscoped spec, or a block with no PASS/FAIL in it (there is no judgement to invalidate).
142
+ * @returns {(string|null)} The problem, phrased for an operator; null for a verdict whose every
143
+ * citation resolves, an unscoped spec, or a block with no PASS/FAIL in it (there is no judgement
144
+ * to invalidate).
79
145
  */
80
146
  export function citationProblem(cwd, slug, verdict) {
81
147
  if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
82
- if (Array.isArray(verdict.t0_citations) && verdict.t0_citations.length) return null;
83
148
  if (!isScoped(cwd, slug)) return null;
84
- return `the ${verdict.overall} verdict cites no T0 artifact, and a verdict on a scoped spec must ` +
85
- "cite the T0 verdict it re-hashed (the order lists them under payload.t0_artifacts)";
149
+ if (!Array.isArray(verdict.t0_citations) || !verdict.t0_citations.length) {
150
+ return `the ${verdict.overall} verdict cites no T0 artifact, and a verdict on a scoped spec must ` +
151
+ "cite the T0 verdict it re-hashed (the order lists them under payload.t0_artifacts)";
152
+ }
153
+ for (const citation of verdict.t0_citations) {
154
+ const reason = unresolvedCitation(cwd, citation);
155
+ if (reason) return `the ${verdict.overall} verdict ${reason} — a T0 citation is re-hashed from disk, never taken on the handed word`;
156
+ }
157
+ return null;
86
158
  }
87
159
 
88
160
  /**