shapeup-sdlc 3.7.2 → 3.7.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.7.2",
4
+ "version": "3.7.3",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -10,6 +10,7 @@ Skills and commands are named short throughout this file; every one of them reso
10
10
  - Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
11
11
  - **The tier-direction rule is refused at the write, not argued at L1b.** A committed file cannot carry a `.shapeup/` path or a `TASK-NNN` id — boards renumber per machine and the run trace is gitignored, so either one resolves only where it was written. Spec-lint has always red'd this at Board Review; the same rule now also refuses the write itself, naming the token, because four producers hit it and one of them hit it twice in one session — a worker carries no lesson across a dispatch, and by L1b the phase that wrote the line is over. The knowledge base is outside the rule: those files are instructions, not references anyone resolves.
12
12
  - Attested, not assumed: a dispatch leaves a receipt naming the skill that ran, and ingest refuses an orchestrated result that has none. A dispatch that fails is answered by the sub-agent improvising the craft, which every other check accepts — so "the artifact exists" is not evidence that the shipped skill produced it. `--no-receipt-check` is the way through when the receipt channel itself fails.
13
+ - **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
13
14
  - GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
14
15
  - Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
15
16
  - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
@@ -34,7 +35,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
34
35
  | Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
35
36
  | Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
36
37
  | Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
37
- | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe + seesaw ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
38
+ | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦; the seesaw regression arm is declared and not yet wired), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
38
39
  | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
39
40
  | FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
40
41
 
@@ -62,7 +63,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
62
63
  - **QA is a level-up, not a gate** — `--no-qa` skips it; circuit breaker outranks the Hunter.
63
64
  - **Role separation** — Evaluator grades, task-executor fixes, QA discovers.
64
65
  - **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
65
- - **Hill phase is mechanical ✦** — derived only from T0/T1/seesaw artifacts, never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
66
+ - **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts (the seesaw arm is declared, not yet wired), never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
66
67
  - **Envelope port (v1.0)** — every dispatch is WorkOrder in / WorkResult out; shared state has exactly one writer (the ingest step); malformed envelopes are hook-denied. Workers: stateless, craft-only, pipeline-blind.
67
68
 
68
69
  ## Setup & Execution
@@ -81,11 +82,11 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
81
82
  through the chained launch, until the run can resume unattended again.
82
83
  - Two storage tiers (ADR-0001): COMMITTED `shapeup/<slug>/` (shaping, spec, scopes, wiring-map, project-profile, requirements, hill, `REPORT.md` frozen at L4) vs GITIGNORED `.shapeup/` (board, orders/results, T0/eval/QA artifacts, ledgers, metrics, gate answers, and the run scripts staged for launch).
83
84
  - **The run launches from a copy inside your project, and it has to.** The Workflow tool loads a script only from a directory the session may already read; the plugin installs outside your project, so naming the shipped path is refused before the run begins and no permission rule repairs it — the grant authorises the tool, not what it may read. Opening a run therefore re-copies the run scripts to `.shapeup/workflows/` and reports the path the launch names. A run in flight keeps the copy it started with: an upgrade reaches the next run, not the current round.
84
- - **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — leaves an order still open by construction, and the pointer is what decides whether that still fences: the fence is enforced only while the pointer exists on disk, so a close that removes it releases the fence whatever the order set says, and restoring the pointer flips it straight back to denying. **Every terminal close retires it** (3.7.1 onward), not only a ship — an escalated or aborted run stops fencing the checkout the moment it closes, and a close that cannot remove a pointer says so in its own return rather than failing. Retiring is not answering, though, and the difference is the whole reason there are still two levers: the pointer says a run is in flight, the missing result says nobody came back, and only the first stops being true at a close. An order the run left unanswered stays genuinely unresolved — un-fenced, but not done — until `init run --force` writes a synthetic result for it, which is also what lets the attempt census keep telling work nobody did apart from work that failed. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
85
- - Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows, agent-call journal rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded — and so does **every other terminal ending**: a run that aborts, escalates or stops at GATE H exports on its close, because the runs whose records are worth most are the ones that never reach the Ship phase. The export is advisory at every one of them: a failed export degrades the trace and is reported in the run's state warnings, but never turns a close into a non-close. A WorkResult carries no `run_id` and reaches it through `order_id`.
85
+ - **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — leaves an order still open by construction, and the pointer is what decides whether that still fences: the fence is enforced only while the pointer exists on disk, so a close that removes it releases the fence whatever the order set says, and restoring the pointer flips it straight back to denying. **Every terminal close retires it** (3.7.1 onward), not only a ship — an escalated or aborted run stops fencing the checkout the moment it closes, and a close that cannot remove a pointer says so in its own return rather than failing. Retiring the pointer does not orphan what follows: the close leaves a breadcrumb naming the run it ended, so the hook decisions taken in the census and ship phase after it still carry that run's key, marked as taken over a closed run. Retiring is not answering, though, and the difference is the whole reason there are still two levers: the pointer says a run is in flight, the missing result says nobody came back, and only the first stops being true at a close. An order the run left unanswered stays genuinely unresolved — un-fenced, but not done — until `init run --force` writes a synthetic result for it, which is also what lets the attempt census keep telling work nobody did apart from work that failed. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
86
+ - Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded — and so does **every other terminal ending**: a run that aborts, escalates or stops at GATE H exports on its close, because the runs whose records are worth most are the ones that never reach the Ship phase. The export is advisory at every one of them: a failed export degrades the trace and is reported in the run's state warnings, but never turns a close into a non-close. A WorkResult carries no `run_id` and reaches it through `order_id`.
86
87
  - Every run projects a **run graph** — `.shapeup/<slug>/graph.jsonl`, append-only, written only by
87
- `reduce graph`. Two families kept separate: work lineage (Run, Order, Result, Verdict, Trial,
88
- GateDecision) and domain (Scope, UseCase, Requirement, Seam). It is derived from the artifacts,
88
+ `reduce graph`. Two families kept separate: work lineage (Run, Order, Result, Leg, Verdict,
89
+ Trial, GateDecision) and domain (Scope, UseCase, Requirement, Seam). It is derived from the artifacts,
89
90
  never authored, so it can be deleted and rebuilt — and a run recorded before it existed backfills
90
91
  on first touch. `--subgraph run` is the fast-forward as one bounded query; `--trace <node>` walks
91
92
  a verdict back to the objective, the plan, the source, the execution record and the gate that
package/README.md CHANGED
@@ -47,8 +47,11 @@ hill phase is derived from artifacts rather than from a worker's own account of
47
47
  Two limits, stated here because the point of this section is that a claim without a mechanism
48
48
  behind it is the thing this harness exists to prevent: the **seesaw** regression arm is declared
49
49
  and not yet wired (no run writes its registry — wiring it is an open Betting Table decision), and
50
- nothing in the runtime **re-hashes** the citation the evaluator is instructed to re-hash. Both
51
- are open items in `shapeup/knowledge-base/harness-defects.md`, not shipped guarantees.
50
+ the citation **re-hash** the kernel performs proves self-consistency, not provenance — the digest
51
+ and the cited verdict are checked, the scope, round and run the artifact belongs to are not. A T0
52
+ artifact is also evidence about the machine that produced it: it records the tree and nothing
53
+ about the toolchain or caches the commands resolved through. All three are open items in
54
+ `shapeup/knowledge-base/harness-defects.md`, not shipped guarantees.
52
55
  → *Prevents: "done" asserted with nothing behind it.*
53
56
 
54
57
  **3. Parallel work can't corrupt shared state.** Each scope gets a write-whitelist of files
@@ -127,8 +130,8 @@ rest of this README after this table and nothing will be a surprise.
127
130
  |---|---|
128
131
  | **board** | The round's task list. "Green" means every task is done. GATE L2's hook reads this before an evaluation and warns if it is not green. |
129
132
  | **round** | One build → evaluate cycle. A FAIL verdict starts round *r+1*. |
130
- | **T0** | The smoke test a scope must pass before it counts as built: its fixtures + a DB probe + the seesaw. Writes an artifact to disk that the evaluator must cite. |
131
- | **seesaw** | The part of T0 that re-runs *other* scopes' fixtures — so a regression is never mistaken for progress. |
133
+ | **T0** | The smoke test a scope must pass before it counts as built: its fixtures + a DB probe (the seesaw arm is declared and not yet wired — see §2 above). Writes an artifact to disk that the evaluator must cite. |
134
+ | **seesaw** | The regression arm of T0, meant to re-run *other* scopes' fixtures so a regression is never mistaken for progress. Declared, not yet wired: no run writes its registry, so no run has executed it. |
132
135
  | **substrate** | The exact list of files one dispatch is allowed to write, stamped into its work order. A hook blocks anything outside it — and anything the order marks frozen. |
133
136
  | **scope contract** | The file defining one vertical slice: its substrate, its fixtures, its affordances. |
134
137
  | **affordance** | The thing a user can actually click, type or call. UI is graded on affordances, not on looks. |
@@ -183,7 +186,7 @@ every arm is skipped when its artifact is absent, so older specs are unaffected.
183
186
  | QA (post-PASS) | `qa-edge-hunter` | v1.1 | Exploratory edge hunt on the running app through six fixed lenses, charting edges *outside* what the evaluator probed. Findings go to the ledger as `~`; never blocks ship. |
184
187
  | Stop (11) | `scope-hammer` | v0.1 | GATE H: must-have census → baseline comparison (never vs. the ideal) → cut list + ship verdict. Handles the normal stop and both circuit-breaker triggers. |
185
188
  | Retro (post-L4) | `coach` | — | RLHF for the harness: turns raw PO/TL feedback at Ship Sign-off into per-skill guidelines under committed `shapeup/knowledge-base/<skill>.md`, read back by six coachable workers on their next run and by `tech-lead` at GATE L0 (workflow guidance, never a gate answer). `--scan` seeds the same files from the project on disk before the first run; `--research <stack>` seeds them from the platform's official documentation when the project has nothing to scan, and cross-checks a scan's rules when it has. GATE COACH-1 asks the PO which skill owns each rule — never assumes; mechanism defects are filed to the harness-defect register instead. |
186
- | Orchestrator | `tech-lead` | v1.0 | Owns the run end-to-end: PLAN once → BUILD all tasks → EVAL once per round, looping on FAIL. Three-level circuit breaker (rounds / T0 attempts / wall clock), T0/seesaw-verified build rounds, mechanical hill derivation. Sole writer of run-state. |
189
+ | Orchestrator | `tech-lead` | v1.0 | Owns the run end-to-end: PLAN once → BUILD all tasks → EVAL once per round, looping on FAIL. Three-level circuit breaker (rounds / T0 attempts / wall clock), T0-verified build rounds, mechanical hill derivation. Sole writer of run-state. |
187
190
 
188
191
  ### Commands
189
192
 
@@ -304,10 +307,10 @@ These hold across the harness and are the reason it stays predictable:
304
307
  of blocking the round. An opt-in third breaker bounds the **wall clock**, because the other two
305
308
  count events and neither can notice a single round running for half an hour — tripping it routes
306
309
  to GATE H, so a run out of time ships what is green instead of being killed and shipping nothing.
307
- - **Hill phase is mechanical, never self-reported** — derived only from T0/T1/seesaw facts, closing
308
- the self-reported-confidence risk — with one gap on record: a scope with no discovery
309
- ledger derives the same phase as one whose unknowns are all closed, so absence still reads
310
- as progress on that one arm.
310
+ - **Hill phase is mechanical, never self-reported** — derived only from T0/T1 facts (the seesaw
311
+ arm is declared, not yet wired), closing the self-reported-confidence risk. A scope with no
312
+ discovery ledger derives no phase rather than a solved one, so absence no longer reads as
313
+ progress on that arm.
311
314
  - **One writer per shared file** — every board/ledger/verdict write goes through
312
315
  `harness reduce ingest`; workers return data and never touch shared state.
313
316
  - **Traceability is oracle-checked, opt-in** — `harness verify trace` verifies covers-closure and
package/SECURITY.md CHANGED
@@ -45,7 +45,7 @@ test against machines you don't own.
45
45
  (override channel fails closed), and every exercised override is logged. The same principle
46
46
  covers `.shapeup/active-order`, which `sandbox-guard` reads to find the run whose orders fence
47
47
  a worker's writes: it sits outside the run-trace carve-out, so a worker cannot repoint its own
48
- sandbox. Every hook resolves that pointer, the decision ledger and the substrate globs against
48
+ sandbox. The `last-run` breadcrumb a terminal close leaves beside it is read for one purpose — keying a decision row to the run that just ended, marked `run_closed` — and arms nothing: liveness is derived from the order set and the pointer, never from that file. Every hook resolves that pointer, the decision ledger and the substrate globs against
49
49
  the project root it finds above the tool call's working directory (a run pointer, the committed
50
50
  tier, or a git boundary) — a worker that `cd`s into a sub-folder is fenced exactly as one at the
51
51
  top, and its receipts land in the same ledger.
@@ -72,7 +72,7 @@ sitting, and reading them is the recommended review.
72
72
  | [`safety-spine.mjs`](hooks/safety-spine.mjs) | PreToolUse (`Bash\|Read\|Write\|Edit\|MultiEdit`) | The proposed command/path; `.shapeup/safety-overrides.json` | Yes — provably destructive ops only: `rm -rf` on unrecoverable targets, `git push --force` / push to main, `git reset --hard`, `git clean -fdx`, `DROP TABLE`/`TRUNCATE`, reads of `.env`/keys/cloud credentials, and any write to its own overrides file | Never blocks an unmatched command; `--force-with-lease` stays allowed |
73
73
  | [`gate-intake.mjs`](hooks/gate-intake.mjs) | PreToolUse (`Skill`) | The `tech-lead` dispatch's own arguments | Yes — an orchestrator dispatch carrying no resolvable intake (no pitch, spec, resume or requirement text) | Fails open on `--order` and on any ambiguous arg shape |
74
74
  | [`harness verify envelope`](kernel/verify/envelope.mjs) | PreToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the JSON schemas | Yes — a worker dispatch whose order file is missing or schema-invalid | Never gates a dispatch that carries no `--order` (standalone skill use stays free) |
75
- | [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
75
+ | [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path, resolved through the filesystem before it is compared — symlinks followed, the unborn tail re-attached to its real ancestor, case folded where the filesystem folds case — so a spelling variant of a frozen file is the frozen file; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
76
76
  | [`tier-guard.mjs`](hooks/tier-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the text the call is about to write (`content`, `new_string`, or any edit in a `MultiEdit` batch) | Yes — a write into `shapeup/<slug>/` whose content names a `.shapeup/` path or a `TASK-NNN` board id: the same tier-direction rule spec-lint reds at GATE L1b, refused at the moment of writing, with the offending token quoted back | No-op outside `shapeup/<slug>/`, on file forms the lint does not scan, and on `shapeup/knowledge-base/` (instructions to a worker, not references). Reads nothing off disk and writes only its decision row; it never inspects a file it is not being asked to write |
77
77
  | [`dispatch-receipt.mjs`](hooks/dispatch-receipt.mjs) | PostToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the tool result's own report of which skill ran | **No — it has no deny path at all.** It records that the shipped skill ran, so `harness reduce ingest` can refuse a result no dispatch produced | Never writes an attestation for a result that does not name a resolved skill; never fails the call it observes (every write is inside `try`/`catch`) |
78
78
  | [`gate-zerowork.mjs`](hooks/gate-zerowork.mjs) | Stop | Run receipts on disk; the session transcript; the decision ledger | **Yes — the one blocking hook.** Returns `decision:"block"` when the session dispatched the orchestrator and produced no run receipt | Defers the moment any receipt exists; `stop_hook_active` caps it at one block per stop chain |
@@ -49,7 +49,7 @@
49
49
  import { appendFileSync, mkdirSync, existsSync, statSync } from "node:fs";
50
50
  import { dirname, resolve } from "node:path";
51
51
  import { decisions, activeScope, sharedDir } from "../../kernel/lib/paths.mjs";
52
- import { resolveRunId } from "../../kernel/lib/paths.mjs";
52
+ import { resolveRun } from "../../kernel/lib/paths.mjs";
53
53
 
54
54
  /**
55
55
  * The project root a hook should file under, from wherever the tool call happened to fire.
@@ -221,7 +221,15 @@ export async function runHook(name, fn) {
221
221
  // run, and recording that is what lets the export tier partition ambient decisions from run
222
222
  // ones. Resolution reads two small files and swallows every error — a receipt must never be
223
223
  // able to fail a tool call.
224
- run_id: (() => { try { return resolveRunId(projectRoot(d.cwd || process.cwd())); } catch { return null; } })(),
224
+ // A row written after the run's terminal close still keys to that run — the close leaves a
225
+ // breadcrumb for exactly this stretch — and says so, because a decision taken over a closed run
226
+ // is a different fact from one taken inside it.
227
+ ...(() => {
228
+ try {
229
+ const { run_id, source } = resolveRun(projectRoot(d.cwd || process.cwd()));
230
+ return source === "closed" ? { run_id, run_closed: true } : { run_id };
231
+ } catch { return { run_id: null }; }
232
+ })(),
225
233
  event: d.event ?? null,
226
234
  tool: d.tool ?? null,
227
235
  subject: d.subject ?? null,
@@ -85,8 +85,8 @@
85
85
  // Contract: PreToolUse stdin JSON { tool_name, tool_input:{file_path | edits[].file_path}, cwd }.
86
86
  // Deny via { hookSpecificOutput: { hookEventName, permissionDecision:"deny", permissionDecisionReason } }.
87
87
 
88
- import { readFileSync, existsSync, appendFileSync, mkdirSync, readdirSync, statSync } from "node:fs";
89
- import { resolve, join, relative, dirname, sep } from "node:path";
88
+ import { readFileSync, existsSync, appendFileSync, mkdirSync, readdirSync, statSync, realpathSync, lstatSync, readlinkSync } from "node:fs";
89
+ import { resolve, join, relative, dirname, basename, sep } from "node:path";
90
90
  import { isMain } from "../kernel/lib/argv.mjs";
91
91
  import { LOCAL, SHARED, activeOrder, ordersDir, resultsDir, metricsShard } from "../kernel/lib/paths.mjs";
92
92
  import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
@@ -115,8 +115,69 @@ export function globToRegExp(glob) {
115
115
  return new RegExp(`^${re}$`);
116
116
  }
117
117
 
118
- export function matchesAny(relPath, globs) {
119
- return (globs || []).some((g) => globToRegExp(g).test(relPath));
118
+ /**
119
+ * Does `relPath` match any of `globs`? With `fold` set, both sides are compared case-folded — the
120
+ * caller passes it when the filesystem underneath does not distinguish case, so that a glob
121
+ * declared as `receipts/**` also covers the spelling `Receipts/**`, which on that filesystem is the
122
+ * same directory.
123
+ */
124
+ export function matchesAny(relPath, globs, fold = false) {
125
+ const path = fold ? relPath.toLowerCase() : relPath;
126
+ return (globs || []).some((g) => globToRegExp(fold ? g.toLowerCase() : g).test(path));
127
+ }
128
+
129
+ // --- the path a write actually lands on ---------------------------------------------------------
130
+ //
131
+ // A substrate glob is a SPELLING; what it protects is a FILE. Measured against a real compiled
132
+ // order: `.shapeup/<slug>/receipts/dispatch.jsonl` was denied as frozen, `Receipts/dispatch.jsonl`
133
+ // was permitted, and on this machine's case-insensitive filesystem the second spelling overwrote
134
+ // the first file. A symlink was the same hole from another side: `src/x -> ../.shapeup/<slug>/legs.jsonl`
135
+ // sits inside an allowed `src/**` and resolves `..` without following the link, so the write went
136
+ // through. No list of globs can close either — case, links, Unicode forms are all further spellings
137
+ // of one target — so the comparison is made on the resolved real path instead, and case-folded when
138
+ // the filesystem itself folds case. The globs stay exactly as the compiler wrote them.
139
+
140
+ /**
141
+ * Resolve `abs` through the filesystem: symlinks in any existing ancestor are followed (a dangling
142
+ * link is followed by reading it), and the part that does not exist yet is re-attached to the real
143
+ * path of the deepest ancestor that does. A path nothing on disk can resolve is returned as given.
144
+ *
145
+ * @param {string} abs - Absolute path as the tool named it.
146
+ * @param {number} [depth] - Recursion guard for link chains.
147
+ * @returns {string} The path the write would actually reach.
148
+ */
149
+ export function realPathOf(abs, depth = 0) {
150
+ if (depth > 40) return abs;
151
+ let cur = abs;
152
+ const tail = [];
153
+ for (let i = 0; i < 128; i++) {
154
+ try { return join(realpathSync.native(cur), ...[...tail].reverse()); } catch { /* not yet resolvable as a whole */ }
155
+ try {
156
+ if (lstatSync(cur).isSymbolicLink()) {
157
+ const target = resolve(dirname(cur), readlinkSync(cur));
158
+ return realPathOf(join(target, ...[...tail].reverse()), depth + 1);
159
+ }
160
+ } catch { /* does not exist at all — climb */ }
161
+ const parent = dirname(cur);
162
+ if (parent === cur) return abs;
163
+ tail.push(basename(cur));
164
+ cur = parent;
165
+ }
166
+ return abs;
167
+ }
168
+
169
+ /**
170
+ * Does the filesystem under `dir` fold case? Decided by asking it: the directory's own real path,
171
+ * with every letter's case swapped, exists and resolves back to the same real path only where the
172
+ * filesystem does not distinguish case. A path with no letters to swap answers "no".
173
+ *
174
+ * @param {string} dir - An existing directory (the project root).
175
+ * @returns {boolean} True on a case-insensitive filesystem.
176
+ */
177
+ export function fsFoldsCase(dir) {
178
+ const swapped = dir.replace(/[a-zA-Z]/g, (c) => (c === c.toLowerCase() ? c.toUpperCase() : c.toLowerCase()));
179
+ if (swapped === dir) return false;
180
+ try { return existsSync(swapped) && realpathSync.native(swapped) === realpathSync.native(dir); } catch { return false; }
120
181
  }
121
182
 
122
183
  function readJSON(p) {
@@ -281,9 +342,15 @@ async function main() {
281
342
  const blockReasons = [];
282
343
  let frozenHits = 0;
283
344
 
345
+ // Compare where the write lands, not how it was spelled (see `realPathOf`). The root is resolved
346
+ // the same way so a checkout reached through a linked directory still yields a relative path.
347
+ const realRoot = realPathOf(resolve(root));
348
+ const fold = fsFoldsCase(realRoot);
349
+ const under = (rel, prefix) => (fold ? rel.toLowerCase().startsWith(prefix.toLowerCase()) : rel.startsWith(prefix));
350
+
284
351
  for (const raw of targetPaths) {
285
- const abs = resolve(cwd, raw);
286
- const rel = relative(root, abs);
352
+ const abs = realPathOf(resolve(cwd, raw));
353
+ const rel = relative(realRoot, abs);
287
354
 
288
355
  // Frozen takes absolute precedence, and it is checked across EVERY live contract: a path one
289
356
  // scope froze stays frozen while another scope is in flight, which is the whole point of
@@ -296,7 +363,7 @@ async function main() {
296
363
  // a planner is graded against — so the compiler emitted a declaration with no enforcer, which is
297
364
  // the exact state this hook exists to end. A path a live contract freezes is a violation
298
365
  // wherever it lives.
299
- const freezer = contracts.find((c) => matchesAny(rel, c.frozen));
366
+ const freezer = contracts.find((c) => matchesAny(rel, c.frozen, fold));
300
367
  if (freezer) {
301
368
  violations.push(rel);
302
369
  frozenHits++;
@@ -304,11 +371,11 @@ async function main() {
304
371
  continue;
305
372
  }
306
373
 
307
- if (rel.startsWith(runTracePrefix)) continue;
374
+ if (under(rel, runTracePrefix)) continue;
308
375
 
309
- if (contracts.some((c) => matchesAny(rel, c.allowed))) continue; // inside a live contract
376
+ if (contracts.some((c) => matchesAny(rel, c.allowed, fold))) continue; // inside a live contract
310
377
 
311
- if (contracts.some((c) => matchesAny(rel, c.appendOnly))) {
378
+ if (contracts.some((c) => matchesAny(rel, c.appendOnly, fold))) {
312
379
  if (p.tool_name === "Write") {
313
380
  violations.push(rel);
314
381
  blockReasons.push(`${rel} is append-only (Write overwrites, use Edit)`);
@@ -332,7 +399,7 @@ async function main() {
332
399
  // spec artifacts, so widening a substrate to reach one is the wrong move in a plausible-looking
333
400
  // direction. Those files belong to the orchestrator, whose write window is a phase boundary —
334
401
  // no dispatch in flight — and never the middle of somebody else's dispatch.
335
- const committed = violations.filter((v) => v.split(/[\\/]/)[0] === SHARED);
402
+ const committed = violations.filter((v) => (fold ? v.split(/[\\/]/)[0].toLowerCase() === SHARED.toLowerCase() : v.split(/[\\/]/)[0] === SHARED));
336
403
  const hint = committed.length === violations.length
337
404
  ? `${SHARED}/ is committed tier: these belong to the orchestrator, not to a worker substrate. `
338
405
  + "Write them at a phase boundary, with no dispatch in flight — do not widen an order to reach one."
@@ -293,6 +293,19 @@ export const activeScope = (cwd) => join(localDir(cwd), "active-scope");
293
293
  */
294
294
  export const activeOrder = (cwd) => join(localDir(cwd), "active-order");
295
295
 
296
+ /**
297
+ * The run a terminal close just ended — `{slug, run_id, closed_at, closed_status}`, written by the
298
+ * close in the same act that retires the run pointers, and superseded by the next close.
299
+ *
300
+ * Why it exists: retiring the pointers is right (a finished run must stop fencing the checkout),
301
+ * but the close lands BEFORE the stretch that decides the most — the scope-hammer census, the ship
302
+ * phase, the export — and every hook decision in that stretch resolved its run through the pointer
303
+ * that was just removed. Measured on a live run: every decision row after `closed_at` carried
304
+ * `run_id: null`, the census dispatch included. The breadcrumb keeps that join without keeping the
305
+ * fence: it names a run, it never arms anything, and a row keyed through it says so.
306
+ */
307
+ export const lastRun = (cwd) => join(localDir(cwd), "last-run");
308
+
296
309
  /**
297
310
  * Where the run scripts are staged for launch, inside the project.
298
311
  *
@@ -531,9 +544,31 @@ export function readRunId(cwd, slug) {
531
544
  * @returns {(string|null)} The run key, or null when no run is active or readable.
532
545
  */
533
546
  export function resolveRunId(cwd, slug = null) {
534
- if (slug) return readRunId(cwd, slug);
547
+ return resolveRun(cwd, slug).run_id;
548
+ }
549
+
550
+ /**
551
+ * The run key AND where it came from — the shape a ledger row needs when the distinction matters.
552
+ *
553
+ * Resolution order: the slug the caller knows (`slug`); else the `active-scope` pointer a live run
554
+ * publishes (`pointer`); else the breadcrumb the last terminal close left (`closed`), so the census
555
+ * and ship phase that follow a close still key to the run they belong to; else `null`, which is
556
+ * "no run" and a real answer. A key resolved through the breadcrumb is a key to a run that is
557
+ * OVER: the caller records that alongside it rather than presenting the two the same way.
558
+ *
559
+ * @param {string} cwd - Project root.
560
+ * @param {(string|null)} [slug=null] - Feature slug when the caller already knows it.
561
+ * @returns {{run_id:(string|null), source:("slug"|"pointer"|"closed"|null)}}
562
+ */
563
+ export function resolveRun(cwd, slug = null) {
564
+ if (slug) return { run_id: readRunId(cwd, slug), source: "slug" };
535
565
  try {
536
566
  const ptr = JSON.parse(readFileSync(activeScope(cwd), "utf8"));
537
- return ptr?.slug ? readRunId(cwd, ptr.slug) : null;
538
- } catch { return null; }
567
+ if (ptr?.slug) return { run_id: readRunId(cwd, ptr.slug), source: "pointer" };
568
+ } catch { /* no live run — fall through to the breadcrumb */ }
569
+ try {
570
+ const crumb = JSON.parse(readFileSync(lastRun(cwd), "utf8"));
571
+ if (typeof crumb?.run_id === "string" && crumb.run_id) return { run_id: crumb.run_id, source: "closed" };
572
+ } catch { /* no close on record either */ }
573
+ return { run_id: null, source: null };
539
574
  }
@@ -59,6 +59,65 @@ export function readLegs(path) {
59
59
  * result exists that was never ingested; a scope with no results at all reports `closed: false`
60
60
  * with an empty `orders` list, so a caller can tell "nothing ran" from "nothing was applied".
61
61
  */
62
+ /**
63
+ * Every order the run compiled, with whether its result landed and whether the single writer
64
+ * applied it — planning orders (`analyze.json`) and build orders (`alpha-r1-a1.json`) alike.
65
+ *
66
+ * Why the generic form exists: the per-scope reading below was the only reader of the leg ledger
67
+ * in the run loop, and it was asked in one place, behind the checks that decide a scope is green.
68
+ * Measured on a live run: five dispatches, five results, three leg rows — and the planning phase
69
+ * whose result nobody applied walked on, because its post-condition checked the artifact the
70
+ * worker wrote directly and never asked whether the writer ran. This is the question, asked of any
71
+ * order by name.
72
+ *
73
+ * @param {string} cwd - Project root.
74
+ * @param {string} slug - Feature slug.
75
+ * @returns {{order:string, name:string, order_id:(string|null), has_result:boolean, applied:boolean}[]}
76
+ */
77
+ export function legsOf(cwd, slug) {
78
+ const oDir = ordersDir(cwd, slug);
79
+ const rDir = resultsDir(cwd, slug);
80
+ const ingested = new Set(readLegs(legLedger(cwd, slug))
81
+ .filter((r) => r.ingested_at)
82
+ .map((r) => String(r.order_id)));
83
+ const out = [];
84
+ for (const f of (existsSync(oDir) ? readdirSync(oDir) : []).filter((x) => x.endsWith(".json")).sort()) {
85
+ let orderId = null;
86
+ try { orderId = JSON.parse(readFileSync(join(oDir, f), "utf8")).order_id ?? null; } catch { /* unreadable — reported as unapplied */ }
87
+ out.push({
88
+ order: join(oDir, f),
89
+ name: f.replace(/\.json$/, ""),
90
+ order_id: orderId,
91
+ has_result: existsSync(join(rDir, f)),
92
+ applied: orderId !== null && ingested.has(orderId),
93
+ });
94
+ }
95
+ return out;
96
+ }
97
+
98
+ /**
99
+ * One order by its file stem — `analyze`, `wire`, `alpha-r1-a1` — and whether its leg closed.
100
+ * @param {string} cwd - Project root.
101
+ * @param {string} slug - Feature slug.
102
+ * @param {string} name - The order file's stem.
103
+ * @returns {{closed:boolean, found:boolean, order:(string|null), order_id:(string|null), has_result:boolean, applied:boolean}}
104
+ */
105
+ export function orderLegState(cwd, slug, name) {
106
+ const o = legsOf(cwd, slug).find((x) => x.name === name);
107
+ if (!o) return { closed: false, found: false, order: null, order_id: null, has_result: false, applied: false };
108
+ return { closed: o.applied, found: true, ...o };
109
+ }
110
+
111
+ /**
112
+ * The results on disk that no leg row applied — finished work the board never saw.
113
+ * @param {string} cwd - Project root.
114
+ * @param {string} slug - Feature slug.
115
+ * @returns {{order:string, name:string, order_id:(string|null)}[]}
116
+ */
117
+ export function openLegs(cwd, slug) {
118
+ return legsOf(cwd, slug).filter((o) => o.has_result && !o.applied);
119
+ }
120
+
62
121
  export function legState(cwd, slug, scopeId, round) {
63
122
  const oDir = ordersDir(cwd, slug);
64
123
  const rDir = resultsDir(cwd, slug);
@@ -89,11 +148,13 @@ export function legState(cwd, slug, scopeId, round) {
89
148
  }
90
149
 
91
150
  export const ARGV_SPEC = {
92
- usage: "harness.mjs probe leg --slug <slug> --scope <scope-id> --round N [--cwd <dir>]",
151
+ usage: "harness.mjs probe leg --slug <slug> (--scope <scope-id> --round N | --order <stem> | --open) [--cwd <dir>]",
93
152
  _: { arity: 0, max: 0, name: "(no positional operands)" },
94
153
  slug: { type: "str", required: true },
95
- scope: { type: "str", required: true },
96
- round: { type: "int", min: 1, required: true },
154
+ scope: { type: "str" },
155
+ round: { type: "int", min: 1 },
156
+ order: { type: "str" },
157
+ open: { type: "flag" },
97
158
  cwd: { type: "path" },
98
159
  };
99
160
 
@@ -107,6 +168,22 @@ export const ARGV_SPEC = {
107
168
  export function cli(rawArgv) {
108
169
  const args = runArgs(ARGV_SPEC, rawArgv);
109
170
  const cwd = resolve(args.cwd || process.cwd());
171
+ // `--open`: every result nothing applied, across the whole run, any phase. Exit 0 when none.
172
+ if (args.open) {
173
+ const open = openLegs(cwd, args.slug);
174
+ console.log(JSON.stringify({ closed: open.length === 0, open_total: open.length, open: open.map((o) => o.order), open_ids: open.map((o) => o.order_id) }));
175
+ process.exit(open.length === 0 ? 0 : 1);
176
+ }
177
+ // `--order <stem>`: one order by file stem (`analyze`, `alpha-r1-a1`). Exit 0 when its leg closed.
178
+ if (args.order) {
179
+ const s = orderLegState(cwd, args.slug, args.order);
180
+ console.log(JSON.stringify(s));
181
+ process.exit(s.closed ? 0 : 1);
182
+ }
183
+ if (!args.scope || !args.round) {
184
+ console.error("probe leg: pass --scope <id> --round N, or --order <stem>, or --open");
185
+ process.exit(2);
186
+ }
110
187
  const s = legState(cwd, args.slug, args.scope, args.round);
111
188
  console.log(JSON.stringify({
112
189
  closed: s.closed,
@@ -61,11 +61,7 @@ import { dirname, join, resolve } from "node:path";
61
61
  import { runArgs } from "../lib/argv.mjs";
62
62
  import { splitFrontmatter, uncoerce } from "../lib/contract.mjs";
63
63
  import { globToRegExp } from "../verify/spec.mjs";
64
- import {
65
- intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir,
66
- orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements,
67
- exportRunDir,
68
- } from "../lib/paths.mjs";
64
+ import { intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir, orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements, exportRunDir, lastRun, readRunId } from "../lib/paths.mjs";
69
65
  import { evalVerdict } from "./eval.mjs";
70
66
  import { collectRun, writeRun } from "../report/export.mjs";
71
67
 
@@ -709,6 +705,15 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
709
705
  */
710
706
  const finishClose = (result) => {
711
707
  const warning = shouldExport ? exportOnClose(cwd, slug) : null;
708
+ // The breadcrumb goes down before the pointers come up: the hooks that fire next — the
709
+ // scope-hammer census, the ship phase — resolve their run through whichever of the two exists,
710
+ // and there must be no instant in which neither does. Best-effort like the rest of this close.
711
+ try {
712
+ mkdirSync(dirname(lastRun(cwd)), { recursive: true });
713
+ writeFileSync(lastRun(cwd), JSON.stringify({
714
+ slug, run_id: readRunId(cwd, slug), closed_status: status, closed_at: new Date().toISOString(),
715
+ }) + "\n");
716
+ } catch { /* a missing breadcrumb costs the post-close rows their key, never the close */ }
712
717
  const stuck = [];
713
718
  for (const pointer of [activeOrder(cwd), activeScope(cwd)]) {
714
719
  try { rmSync(pointer, { force: true }); } catch { /* fall through to the check below */ }
@@ -26,11 +26,11 @@
26
26
  // appended again and the LAST line wins on read, so the file is a log and the projection is a fold.
27
27
 
28
28
  import { existsSync, readdirSync, readFileSync, appendFileSync, mkdirSync } from "node:fs";
29
+ import { readLegs } from "../probe/leg.mjs";
29
30
  import { join, dirname, basename, resolve } from "node:path";
30
31
  import { runArgs } from "../lib/argv.mjs";
31
32
  import {
32
- localRoot, receipt as receiptPath, ordersDir, resultsDir, verdictsDir, trials as trialsPath,
33
- gates as gatesPath, scopesDir, usecasesDir, requirements as requirementsPath, wiringMap as wiringMapPath,
33
+ localRoot, receipt as receiptPath, ordersDir, resultsDir, verdictsDir, trials as trialsPath, gates as gatesPath, scopesDir, usecasesDir, requirements as requirementsPath, wiringMap as wiringMapPath, legLedger,
34
34
  } from "../lib/paths.mjs";
35
35
  import { readAllContracts, readContract, ucId, reqId, SCOPE_CONTRACT, WIRING_MAP } from "../lib/contract.mjs";
36
36
  import { runIdFromReceipt } from "../lib/paths.mjs";
@@ -39,11 +39,11 @@ import { runIdFromReceipt } from "../lib/paths.mjs";
39
39
  export const graphPath = (cwd, slug) => join(localRoot(cwd, slug), "graph.jsonl");
40
40
 
41
41
  /** Node types, by family. A type outside these sets is a bug, not an extension point. */
42
- export const WORK_NODES = ["Run", "Order", "Result", "Verdict", "Trial", "GateDecision"];
42
+ export const WORK_NODES = ["Run", "Order", "Result", "Leg", "Verdict", "Trial", "GateDecision"];
43
43
  export const DOMAIN_NODES = ["Scope", "UseCase", "Requirement", "Seam"];
44
44
 
45
45
  /** Edge types. Each names a direction that is meaningful to read backwards. */
46
- export const EDGES = ["PRODUCED", "EVALUATES", "SUPERSEDES", "COVERS", "DEPENDS_ON", "DERIVED_FROM", "IMPLEMENTS"];
46
+ export const EDGES = ["PRODUCED", "INGESTED", "EVALUATES", "SUPERSEDES", "COVERS", "DEPENDS_ON", "DERIVED_FROM", "IMPLEMENTS"];
47
47
 
48
48
  /**
49
49
  * Read the graph as a log and fold it into nodes and edges.
@@ -148,6 +148,22 @@ export function project(cwd, slug) {
148
148
  }
149
149
  }
150
150
 
151
+ // Leg-completion rows — the record that separates "the result landed" from "the single writer
152
+ // applied it". It was in neither this graph nor the export, so a run could show five orders
153
+ // producing five results and nobody could see that only three were ever read. A Result with no
154
+ // INGESTED edge leaving it is finished work the board never saw, and `--subgraph run` names it.
155
+ for (const l of readLegs(legLedger(cwd, slug))) {
156
+ if (!l?.order_id) continue;
157
+ const id = `leg:${l.order_id}`;
158
+ node(id, "Leg", {
159
+ order_id: l.order_id, worker: l.worker ?? null, operation: l.operation ?? null,
160
+ scope_id: l.scope_id ?? null, round: l.round ?? null, attempt: l.attempt ?? null,
161
+ dispatched_at: l.dispatched_at ?? null, ingested_at: l.ingested_at ?? null,
162
+ attested: l.attested ?? null, run_id: l.run_id ?? runId ?? null,
163
+ });
164
+ edge(`result:${l.order_id}`, "INGESTED", id);
165
+ }
166
+
151
167
  // T0 verdicts — the artifact the evaluator is required to cite, and the reason the lineage half
152
168
  // of this graph is worth having: a verdict node is the anchor of every "show me the evidence".
153
169
  const vDir = verdictsDir(cwd, slug);
@@ -370,6 +386,7 @@ export function runSubgraph(cwd, slug) {
370
386
  const of = (t) => [...nodes.values()].filter((n) => n.t === t);
371
387
  const orders = of("Order"), results = of("Result"), verdicts = of("Verdict");
372
388
  const resultIds = new Set(results.map((r) => r.order_id));
389
+ const ingestedFrom = new Set([...edges.values()].filter((e) => e.t === "INGESTED").map((e) => e.from));
373
390
  const greenByRound = {};
374
391
  for (const v of verdicts) {
375
392
  if (v.overall !== "green" || v.round == null || !v.scope_id) continue;
@@ -384,6 +401,8 @@ export function runSubgraph(cwd, slug) {
384
401
  seams: of("Seam").map((s) => s.seam).sort(),
385
402
  orders: orders.length,
386
403
  pending_orders: orders.filter((o) => !resultIds.has(o.order_id)).map((o) => o.order_id).sort(),
404
+ // Results the single writer never applied — a leg that came back and was not read.
405
+ unapplied_results: results.filter((r) => !ingestedFrom.has(`result:${r.order_id}`)).map((r) => r.order_id).sort(),
387
406
  rounds_with_green: Object.keys(greenByRound).map(Number).sort((a, b) => a - b),
388
407
  green_scopes_by_round: Object.fromEntries(Object.entries(greenByRound).map(([r, s]) => [r, [...s].sort()])),
389
408
  trials: of("Trial").length,
@@ -45,7 +45,9 @@ import { readFileSync, writeFileSync, readdirSync, mkdirSync, existsSync, statSy
45
45
  import { join, resolve } from "node:path";
46
46
  import { runArgs } from "../lib/argv.mjs";
47
47
  import { splitFrontmatter } from "../lib/contract.mjs";
48
- import { runIdFromReceipt, readReceipt } from "../lib/paths.mjs";
48
+ import {
49
+ runIdFromReceipt, readReceipt, legLedger,
50
+ } from "../lib/paths.mjs";
49
51
  import { TABLES, runRow, dispatchFacts } from "./facts.mjs";
50
52
  import { deriveRounds } from "../probe/rounds.mjs";
51
53
  import {
@@ -260,6 +262,9 @@ export function collectRun(cwd, slug) {
260
262
  // The decision that crossed each gate, and the round build gate's own artifact.
261
263
  gate_decision: readJsonl(gatesPath(cwd, slug), t).map((g) => gateDecisionRow(g, runId)),
262
264
  build_gate: readJsonDir(roundBuildDir(cwd, slug), t).map((a) => buildGateRow(a, runId)),
265
+ leg: readJsonl(legLedger(cwd, slug), t)
266
+ .filter((r) => !runId || !r?.run_id || r.run_id === runId)
267
+ .map((r) => ({ ...r, run_id: r.run_id ?? runId ?? null })),
263
268
  },
264
269
  defects: { records_skipped: t.skipped },
265
270
  };
@@ -32,6 +32,9 @@ export const TABLES = [
32
32
  // has to. `build_gate` is the round build gate's own artifact (kernel/verify/build.mjs), on the
33
33
  // same terms: it ends a round exactly as EVAL does, and had no table either.
34
34
  "gate_decision", "build_gate",
35
+ // The leg-completion ledger — one row per order the single writer applied. Without it the
36
+ // warehouse could join an order to its result and never say whether anyone read the result.
37
+ "leg",
35
38
  ];
36
39
 
37
40
  /** Coerce anything to a finite number, or null. Keeps `0` and rejects `NaN`/`""`/undefined. */
@@ -33,6 +33,10 @@
33
33
  "type": null,
34
34
  "gap": "a gate decision row — emitted by reduce graph, never typed here"
35
35
  },
36
+ "Leg": {
37
+ "type": null,
38
+ "gap": "a leg-completion row (reduce ingest's record that a result was applied) — emitted by reduce graph, never typed here"
39
+ },
36
40
  "Scope": {
37
41
  "type": "ScopeContract"
38
42
  },
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.7.2",
3
+ "version": "3.7.3",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -41,7 +41,8 @@
41
41
  // { status: "shipped", verdict, rounds_used, dims_not_evaluated, qa_findings, report }
42
42
  // { status: "paused", paused_at, block, valid_decisions, context }
43
43
  // { status: "aborted", aborted_at, reason }
44
- // { status: "gate_h", breaker: "outer"|"inner"|"deadline", hammer_proposals, green_scopes }
44
+ // { status: "gate_h", breaker: "outer"|"attempt_budget"|"none"|"deadline", hammer_proposals, green_scopes,
45
+ // tripped_scopes?, unapplied_results? }
45
46
 
46
47
  // meta must be a PURE LITERAL — the runtime parses it statically, before the body ever runs, and
47
48
  // rejects the whole script on anything it has to evaluate. A `+`-joined description is a
@@ -471,6 +472,32 @@ const T0CHECK = {
471
472
  required: ["green"],
472
473
  };
473
474
 
475
+ /** `probe leg --order` — did one named order's result reach the single writer? Any phase. */
476
+ const ORDERLEG = {
477
+ type: "object",
478
+ properties: {
479
+ closed: { type: "boolean" },
480
+ found: { type: "boolean" },
481
+ order: nullable("string"),
482
+ has_result: { type: "boolean" },
483
+ applied: { type: "boolean" },
484
+ },
485
+ required: ["closed", "found", "has_result", "applied"],
486
+ };
487
+
488
+ /** `probe attempts` — the attested census: what was spent, and whether the breaker really tripped. */
489
+ const ATTEMPTS = {
490
+ type: "object",
491
+ properties: {
492
+ scope_id: { type: "string" },
493
+ spent: { type: "integer" },
494
+ in_flight: { type: "integer" },
495
+ green: { type: "boolean" },
496
+ tripped: { type: "boolean" },
497
+ },
498
+ required: ["scope_id", "spent", "tripped"],
499
+ };
500
+
474
501
  /** `probe leg` — did this scope's result reach the board, or is it finished work nothing applied? */
475
502
  const LEGCHECK = {
476
503
  type: "object",
@@ -817,6 +844,42 @@ async function crossGate(gateId, phaseName, validDecisions, ctx) {
817
844
  const attest = (phaseKey, phaseName, label) =>
818
845
  cmd(`probe resume --slug ${slug} --require ${phaseKey}`, phaseName, label);
819
846
 
847
+ /**
848
+ * The other half of a phase post-condition: the single writer ran.
849
+ *
850
+ * The artifact check asks whether the worker wrote its product. It cannot ask whether the
851
+ * WorkResult was applied — the leg row, the discoveries, the board — because the product is written
852
+ * by the worker directly and the envelope is applied by `reduce ingest`, a separate act the leg's
853
+ * script names as its last step. Measured on a live run: a planning phase landed a result naming
854
+ * seventeen artifacts and no leg row, the artifact check passed, and the run walked on with the
855
+ * result's discoveries never reaching the ledger. So this asks the leg ledger by order name. A
856
+ * result on disk that nothing applied is ingested here — the same repair the build round makes,
857
+ * for the same reason: the single writer is the invariant, not which step invokes it — and if it
858
+ * still is not applied after that, the run stops with a cause naming the writer that did not run,
859
+ * rather than folding the gap into a phase that "completed".
860
+ *
861
+ * @param {string} gate - The gate name to report the abort under.
862
+ * @param {string} phaseKey - The phase, which is also its order's file stem.
863
+ * @param {string} phaseName - Progress group.
864
+ * @returns {Promise<(object|null)>} An aborted RunReturn, or null when the leg closed (or the
865
+ * question could not be asked — a probe that did not run proves nothing, and is logged as such).
866
+ */
867
+ async function requireLeg(gate, phaseKey, phaseName) {
868
+ const ask = () => query(`probe leg --slug ${slug} --order "${phaseKey}"`, ORDERLEG, phaseName, `legcheck:${phaseKey}`);
869
+ let leg = await ask();
870
+ if (!leg || !leg.found) { log(`${gate} — could not ask the leg ledger about "${phaseKey}" (probe returned ${leg ? "no order" : "nothing"}); proceeding on the artifact alone.`); return null; }
871
+ if (!leg.has_result || leg.applied) return null;
872
+ log(`${gate} — "${phaseKey}" came back with a result nothing applied (no leg row). Ingesting it here: ${leg.order}.`);
873
+ await advisory(`reduce ingest --order "${leg.order}"`, phaseName, `late-ingest:${phaseKey}`);
874
+ leg = await ask();
875
+ if (leg?.applied) return null;
876
+ return aborted(gate,
877
+ `${gate}: the single writer did not run for "${phaseKey}" — its WorkResult is on disk and no leg row ` +
878
+ `records it being applied, and a late \`reduce ingest\` did not take either. The phase's product exists; ` +
879
+ `what it discovered and reported never reached the ledger or the board. Read the result, run ` +
880
+ `\`reduce ingest --order "${leg?.order ?? "<its order>"}"\` by hand to see why it refuses, then relaunch.`);
881
+ }
882
+
820
883
  /**
821
884
  * The phase post-condition: the artifact is on disk, or the run stops here.
822
885
  *
@@ -830,7 +893,7 @@ const attest = (phaseKey, phaseName, label) =>
830
893
  */
831
894
  async function requirePhase(gate, phaseKey, phaseName) {
832
895
  const r = await attest(phaseKey, phaseName, `require:${phaseKey}`);
833
- if (r.exit_code === 0) return null;
896
+ if (r.exit_code === 0) return await requireLeg(gate, phaseKey, phaseName);
834
897
  // Exit 6 is `probe resume --require`'s OWN documented code for "the artifact really is not on
835
898
  // disk" (kernel/probe/resume.mjs banner). Any other value — including -1, the courier's sentinel
836
899
  // for a tool call that never ran — is not that predicate answering "no"; it is the predicate never
@@ -1001,7 +1064,11 @@ async function closeIfTerminal(ret) {
1001
1064
  const cause = ret.status === "aborted"
1002
1065
  ? `${ret.aborted_at || "?"}: ${ret.reason || "no reason recorded"}`
1003
1066
  : ret.status === "gate_h"
1004
- ? `breaker=${ret.breaker ?? "?"} green_scopes=${Array.isArray(ret.green_scopes) ? ret.green_scopes.length : "?"} hammer_proposals=${Array.isArray(ret.hammer_proposals) ? ret.hammer_proposals.length : "?"}`
1067
+ ? (ret.stalled ? `stalled=${ret.stalled} ` : "")
1068
+ + `breaker=${ret.breaker ?? "?"} green_scopes=${Array.isArray(ret.green_scopes) ? ret.green_scopes.length : "?"} hammer_proposals=${Array.isArray(ret.hammer_proposals) ? ret.hammer_proposals.length : "?"}`
1069
+ + (Array.isArray(ret.tripped_scopes) ? ` tripped_scopes=${ret.tripped_scopes.length}` : "")
1070
+ // A result the single writer never applied is named at the close, not folded into "not green".
1071
+ + (Array.isArray(ret.unapplied_results) && ret.unapplied_results.length ? ` unapplied_results=${ret.unapplied_results.length}` : "")
1005
1072
  : `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}`;
1006
1073
  // `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
1007
1074
  // a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
@@ -1368,6 +1435,7 @@ let verdict = null;
1368
1435
  // other two. The cure is the same each time and it is not a bigger variable: re-derive the fact.
1369
1436
  const allGreen = [];
1370
1437
  const allHammer = [];
1438
+ const allUnapplied = []; // results on disk the single writer never applied, run-wide
1371
1439
  // OUTSIDE the loop, because its whole purpose is to cross a round boundary: round r's verdict is
1372
1440
  // what round r+1 has to act on. Declared inside, it was in the temporal dead zone at the BUILD that
1373
1441
  // needed it — a runtime error no static check can see, since nothing but a real second round ever
@@ -1408,7 +1476,7 @@ while (verdict !== "pass" && round <= maxRounds) {
1408
1476
  const budget = await cmd(`verify budget --slug ${slug} --strict`, "Build", `budget:r${round}`);
1409
1477
  if (budget.exit_code === 6) {
1410
1478
  await advisory(`reduce hill --slug ${slug}`, "Build", "hill-derive");
1411
- return await withWarnings({ status: "gate_h", breaker: "deadline", hammer_proposals: allHammer, green_scopes: allGreen });
1479
+ return await withWarnings({ status: "gate_h", breaker: "deadline", unapplied_results: allUnapplied, hammer_proposals: allHammer, green_scopes: allGreen });
1412
1480
  }
1413
1481
 
1414
1482
  log(`BUILD round ${round} — ${scopes.length} scope(s), up to ${maxParallelScopes} at once, attempt budget ${attemptBudget}`);
@@ -1463,8 +1531,22 @@ while (verdict !== "pass" && round <= maxRounds) {
1463
1531
  : { scope_id: s.scope_id, pending: true }), // not green yet → stage 2 builds it
1464
1532
  async (pre, s) => (pre?.pending ? buildScope(s, round) : pre),
1465
1533
  async (res, s) => {
1466
- if (!res || res.__failed) return res;
1467
- if (!res.green) return res;
1534
+ if (!res) return res;
1535
+ // THE SINGLE WRITER IS ASKED FIRST, WHATEVER THE LEG SAID. This question used to sit behind
1536
+ // the green checks below, and the state it exists to catch — a result on disk that nothing
1537
+ // applied — was reachable only for a scope that was already fully green. Measured live: a
1538
+ // leg reported green with no T0 verdict at all, the T0 re-read correctly returned "not
1539
+ // green", the round returned before this line, and the result — six tasks, two discoveries
1540
+ // — was never read by anyone. A dead leg (`__failed`) can have left a result too. So every
1541
+ // settled scope is asked, and what it answers travels on the result to the close, where an
1542
+ // unapplied result is named rather than folded into "not green". Application stays gated
1543
+ // on the green checks: ingest ticks acceptance boxes, and a result T0 never measured must
1544
+ // not mark work green — the repair below is for a scope that IS green.
1545
+ const applied = await query(`probe leg --slug ${slug} --scope ${s.scope_id} --round ${round}`,
1546
+ LEGCHECK, "Build", `legcheck:${s.scope_id}-r${round}`);
1547
+ const unapplied = applied?.closed ? [] : (applied?.unapplied || []);
1548
+ if (res.__failed) return unapplied.length ? { ...res, unapplied } : res;
1549
+ if (!res.green) return unapplied.length ? { ...res, unapplied } : res;
1468
1550
  // THE T0 RE-READ IS SKIPPED FOR A RESUMED SCOPE; THE LEG CHECK BELOW IS NOT.
1469
1551
  //
1470
1552
  // `resumed` means the graph already reported this scope green for this round, so re-reading
@@ -1481,7 +1563,7 @@ while (verdict !== "pass" && round <= maxRounds) {
1481
1563
  if (!confirmed?.green) {
1482
1564
  log(`BUILD r${round} — ${s.scope_id} reported green but no T0 verdict is on disk for this ` +
1483
1565
  `round; treating it as not green (the evaluator cites that artifact, and it is not there).`);
1484
- return { ...res, green: false, reason: "reported green with no T0 verdict artifact on disk" };
1566
+ return { ...res, green: false, reason: "reported green with no T0 verdict artifact on disk", ...(unapplied.length ? { unapplied } : {}) };
1485
1567
  }
1486
1568
  }
1487
1569
  // AND ITS RESULT HAS TO HAVE REACHED THE BOARD. A green T0 says the worker's fixtures ran and
@@ -1494,9 +1576,7 @@ while (verdict !== "pass" && round <= maxRounds) {
1494
1576
  //
1495
1577
  // The evidence is the leg-completion row, because `reduce ingest` writes it: its presence
1496
1578
  // proves the writer ran, and it is not something the leg can assert about itself.
1497
- const applied = await query(`probe leg --slug ${slug} --scope ${s.scope_id} --round ${round}`,
1498
- LEGCHECK, "Build", `legcheck:${s.scope_id}-r${round}`);
1499
- for (const orderPath of (applied?.closed ? [] : applied?.unapplied || [])) {
1579
+ for (const orderPath of unapplied) {
1500
1580
  // INGESTED HERE RATHER THAN FAILED. The result is on disk and valid — re-running the leg
1501
1581
  // would pay a whole attempt again for work already done. Only the single writer writes
1502
1582
  // shared state, and that writer is this command; which step invokes it is not the invariant.
@@ -1524,6 +1604,10 @@ while (verdict !== "pass" && round <= maxRounds) {
1524
1604
 
1525
1605
  for (const [i, res] of settled.entries()) {
1526
1606
  const scopeId = buildOrder[i].scope_id;
1607
+ for (const p of (res?.unapplied || [])) {
1608
+ if (!allUnapplied.includes(p)) allUnapplied.push(p);
1609
+ log(`BUILD r${round} — ${scopeId} has a result on disk nothing applied: ${p}`);
1610
+ }
1527
1611
  // A dead builder is a SPENT ATTEMPT, not a dead run: the scope goes to GATE H's census and
1528
1612
  // the round continues. Killing the run here would discard every other scope's green work.
1529
1613
  if (!res || res.__failed) {
@@ -1541,10 +1625,26 @@ while (verdict !== "pass" && round <= maxRounds) {
1541
1625
  for (const sid of roundGreen) { const i = allHammer.indexOf(sid); if (i !== -1) allHammer.splice(i, 1); }
1542
1626
  for (const sid of roundHammer) if (!allHammer.includes(sid) && !allGreen.includes(sid)) allHammer.push(sid);
1543
1627
 
1544
- // INNER breaker: nothing green and something queued → GATE H. The census is scope-hammer's job.
1628
+ // NOTHING GREEN AND SOMETHING QUEUED → GATE H. This used to return the literal `inner` as its breaker, and the
1629
+ // protocol's INNER breaker is the per-scope attempt budget, which "queues a proposal, never blocks
1630
+ // the round" — so the close named a breaker the attested census flatly denied (one attempt spent
1631
+ // of five, `tripped: false`), and the operator was told a scope had exhausted its attempts after
1632
+ // it used one. The word is now earned: `probe attempts` — the one derivation built so the census
1633
+ // and the breaker cannot disagree — is asked for every queued scope, and the return names
1634
+ // `attempt_budget` only for the scopes it says tripped, `none` when the round simply stalled.
1545
1635
  if (roundGreen.length === 0 && roundHammer.length > 0) {
1636
+ const tripped = [];
1637
+ for (const sid of roundHammer) {
1638
+ const census = await query(`probe attempts --slug ${slug} --scope "${sid}" --round ${round} --attempt-budget ${attemptBudget}`,
1639
+ ATTEMPTS, "Build", `census:${sid}-r${round}`);
1640
+ if (census?.tripped) tripped.push(sid);
1641
+ }
1546
1642
  await advisory(`reduce hill --slug ${slug}`, "Build", "hill-derive");
1547
- return await withWarnings({ status: "gate_h", breaker: "inner", hammer_proposals: allHammer, green_scopes: allGreen });
1643
+ return await withWarnings({
1644
+ status: "gate_h", breaker: tripped.length ? "attempt_budget" : "none", stalled: "no_green",
1645
+ tripped_scopes: tripped, unapplied_results: allUnapplied,
1646
+ hammer_proposals: allHammer, green_scopes: allGreen,
1647
+ });
1548
1648
  }
1549
1649
 
1550
1650
  // ---- ROUND BUILD GATE — the feature builds and launches, measured before anyone is asked --------
@@ -1655,13 +1755,13 @@ while (verdict !== "pass" && round <= maxRounds) {
1655
1755
  // SHIP" once EVAL is skipped, never spends another round waiting on a verdict nobody is producing.
1656
1756
  if (verdict === "pass" || verdict === "not-evaluated") break; // → QA → GATE H → ship
1657
1757
  if (g3.decision === "stop" || round >= maxRounds) {
1658
- return await withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
1758
+ return await withWarnings({ status: "gate_h", breaker: "outer", unapplied_results: allUnapplied, hammer_proposals: allHammer, green_scopes: allGreen });
1659
1759
  }
1660
1760
  round += 1;
1661
1761
  }
1662
1762
 
1663
1763
  if (verdict !== "pass" && verdict !== "not-evaluated") {
1664
- return await withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
1764
+ return await withWarnings({ status: "gate_h", breaker: "outer", unapplied_results: allUnapplied, hammer_proposals: allHammer, green_scopes: allGreen });
1665
1765
  }
1666
1766
 
1667
1767
  // ---- QA (post-PASS, pre-ship) — a level-up, never a gate. `--no-qa` answers it "skip". --------