shapeup-sdlc 3.7.0 → 3.7.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.7.0",
4
+ "version": "3.7.2",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -7,7 +7,8 @@ A three-phase Shape Up loop orchestrated by `/tech-lead`. Invariants live in the
7
7
 
8
8
  Skills and commands are named short throughout this file; every one of them resolves only under the plugin's namespace — `Skill(shapeup-sdlc-plugin:orient)`, `/shapeup-sdlc-plugin:ship`. A bare name is not a typo you get a warning for: an unknown skill name is rejected upstream of the hook layer, so the dispatch fires no hook, leaves no decision row, and is answered by improvisation.
9
9
 
10
- - Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, stopping a run with no receipt (the run's first act writes one).
10
+ - Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
11
+ - **The tier-direction rule is refused at the write, not argued at L1b.** A committed file cannot carry a `.shapeup/` path or a `TASK-NNN` id — boards renumber per machine and the run trace is gitignored, so either one resolves only where it was written. Spec-lint has always red'd this at Board Review; the same rule now also refuses the write itself, naming the token, because four producers hit it and one of them hit it twice in one session — a worker carries no lesson across a dispatch, and by L1b the phase that wrote the line is over. The knowledge base is outside the rule: those files are instructions, not references anyone resolves.
11
12
  - Attested, not assumed: a dispatch leaves a receipt naming the skill that ran, and ingest refuses an orchestrated result that has none. A dispatch that fails is answered by the sub-agent improvising the craft, which every other check accepts — so "the artifact exists" is not evidence that the shipped skill produced it. `--no-receipt-check` is the way through when the receipt channel itself fails.
12
13
  - GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
13
14
  - Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
@@ -80,7 +81,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
80
81
  through the chained launch, until the run can resume unattended again.
81
82
  - Two storage tiers (ADR-0001): COMMITTED `shapeup/<slug>/` (shaping, spec, scopes, wiring-map, project-profile, requirements, hill, `REPORT.md` frozen at L4) vs GITIGNORED `.shapeup/` (board, orders/results, T0/eval/QA artifacts, ledgers, metrics, gate answers, and the run scripts staged for launch).
82
83
  - **The run launches from a copy inside your project, and it has to.** The Workflow tool loads a script only from a directory the session may already read; the plugin installs outside your project, so naming the shipped path is refused before the run begins and no permission rule repairs it — the grant authorises the tool, not what it may read. Opening a run therefore re-copies the run scripts to `.shapeup/workflows/` and reports the path the launch names. A run in flight keeps the copy it started with: an upgrade reaches the next run, not the current round.
83
- - **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — can leave one still open, and while the run's pointer is still on disk the fence holds on that order exactly as if the run were live. The pointer is what actually switches it off, though: the fence is enforced only while that pointer exists on disk, so a close that removes it releases the fence whatever the order set still says — restoring the pointer flips it straight back to denying. A **ship** close does not itself answer what it leaves outstanding — a ship close can retire the pointer over an order that never got a result — so it is not special because nothing is left unanswered; it is special only because retiring the pointer is the one lever every close needs pulled, and `reduce ship` pulls it for you. Whichever way a run closes, an order it leaves unanswered stays genuinely unresolved — not merely un-fenced — until `init run --force` runs: it writes a synthetic result for every order the closed run left unanswered, so the fence lifts without waiting on a worker that is never coming back. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
84
+ - **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — leaves an order still open by construction, and the pointer is what decides whether that still fences: the fence is enforced only while the pointer exists on disk, so a close that removes it releases the fence whatever the order set says, and restoring the pointer flips it straight back to denying. **Every terminal close retires it** (3.7.1 onward), not only a ship — an escalated or aborted run stops fencing the checkout the moment it closes, and a close that cannot remove a pointer says so in its own return rather than failing. Retiring is not answering, though, and the difference is the whole reason there are still two levers: the pointer says a run is in flight, the missing result says nobody came back, and only the first stops being true at a close. An order the run left unanswered stays genuinely unresolved — un-fenced, but not done — until `init run --force` writes a synthetic result for it, which is also what lets the attempt census keep telling work nobody did apart from work that failed. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
84
85
  - Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows, agent-call journal rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded — and so does **every other terminal ending**: a run that aborts, escalates or stops at GATE H exports on its close, because the runs whose records are worth most are the ones that never reach the Ship phase. The export is advisory at every one of them: a failed export degrades the trace and is reported in the run's state warnings, but never turns a close into a non-close. A WorkResult carries no `run_id` and reaches it through `order_id`.
85
86
  - Every run projects a **run graph** — `.shapeup/<slug>/graph.jsonl`, append-only, written only by
86
87
  `reduce graph`. Two families kept separate: work lineage (Run, Order, Result, Verdict, Trial,
package/README.md CHANGED
@@ -41,9 +41,14 @@ runs over unfinished tasks rather than denying it. The board is local to the mac
41
41
  harness — see [ADR-0001](docs/design/adr/0001-consumer-file-organization.md).)
42
42
 
43
43
  **2. Progress is measured, not claimed.** A scope counts as built only when `t0-verify` runs
44
- its fixtures, a DB probe, and the seesaw, and writes an artifact to disk. The evaluator must
45
- cite that artifact and re-hashes it itself; hill phase is derived from those facts, so no
46
- worker can self-report confidence.
44
+ its fixtures and its DB probe and writes an artifact to disk — with each command's exit code,
45
+ its captured output, and whether it ran at all. The evaluator must cite that artifact, and the
46
+ hill phase is derived from artifacts rather than from a worker's own account of its progress.
47
+ Two limits, stated here because the point of this section is that a claim without a mechanism
48
+ behind it is the thing this harness exists to prevent: the **seesaw** regression arm is declared
49
+ and not yet wired (no run writes its registry — wiring it is an open Betting Table decision), and
50
+ nothing in the runtime **re-hashes** the citation the evaluator is instructed to re-hash. Both
51
+ are open items in `shapeup/knowledge-base/harness-defects.md`, not shipped guarantees.
47
52
  → *Prevents: "done" asserted with nothing behind it.*
48
53
 
49
54
  **3. Parallel work can't corrupt shared state.** Each scope gets a write-whitelist of files
@@ -210,7 +215,7 @@ the layer that carries it, and the three layers here fail differently:
210
215
  | **Runtime** — the kernel and the run script | When the run goes through the harness. A schema rejection or a non-zero exit stops the step. | A lane that never calls the kernel is never checked — which is why the hooks below cover the doors, not the steps. |
211
216
  | **Advisory** — a report section | When somebody reads the artifact. | Silently, if nobody does. It is a cleanup list, never a verdict. |
212
217
 
213
- **Four walls.** These are hooks because nothing in the runtime can substitute for them:
218
+ **Five walls.** These are hooks because nothing in the runtime can substitute for them:
214
219
 
215
220
  - `PreToolUse` (`Skill`) — **`hooks/gate-intake.mjs` denies a `tech-lead` dispatch that carries no
216
221
  pitch, no spec folder, and no requirement text.** Observed, not theorized: when the requirement
@@ -227,6 +232,13 @@ the layer that carries it, and the three layers here fail differently:
227
232
  checked first and outranks everything, across every live contract — including the carve-out that
228
233
  otherwise keeps the active feature's own run trace writable, so a path a live order froze stays
229
234
  frozen wherever it lives.
235
+ - `PreToolUse` (`Edit|Write|MultiEdit`) — **`hooks/tier-guard.mjs` refuses a committed-tier write
236
+ whose content names a path into the local run trace or a board id.** Same rule spec-lint reds at
237
+ GATE L1b, asked at the moment of writing: four different producers wrote a committed file their
238
+ own lint then rejected, and one of them did it twice in one session, an hour apart, after fixing
239
+ the first occurrence itself. A worker carries no lesson across a dispatch, so the remedy is a
240
+ refusal the writer can act on while it still knows what it meant to say. The knowledge base is
241
+ outside it — those files are instructions, not references a reader has to resolve.
230
242
  - `PreToolUse` (`Bash|Read|Write|Edit|MultiEdit`) — **`hooks/safety-spine.mjs` denies destructive
231
243
  commands** (`rm -rf` on unrecoverable targets, force-push/push-to-main, `git reset --hard`,
232
244
  `DROP TABLE`) and secret-file reads. A machine guard, not a pipeline guard; the escape hatch is
@@ -265,7 +277,7 @@ the layer that carries it, and the three layers here fail differently:
265
277
  | `anti-rationalization` (claims the facts contradict) | The ship report's census, derived from the board and the T0 artifacts | The facts are in an artifact a teammate finds on `git pull`, not in a transcript nobody re-reads. |
266
278
  | `slop-cleaner` (TODO/`console.log` leftovers) | The ship report's **Leftovers** section | Same scan, same added-lines-only rule; it lands somewhere checkable. |
267
279
 
268
- **Nothing load-bearing depends on permission mode.** The four walls plus the zero-work gate run
280
+ **Nothing load-bearing depends on permission mode.** The five walls plus the zero-work gate run
269
281
  under every mode. The kernel needs a grant to be *invoked* — two Bash lines `npx shapeup-sdlc init`
270
282
  writes — but a session that never gets that grant is a session that cannot run the pipeline at all,
271
283
  not one that runs it unguarded.
@@ -293,7 +305,9 @@ These hold across the harness and are the reason it stays predictable:
293
305
  count events and neither can notice a single round running for half an hour — tripping it routes
294
306
  to GATE H, so a run out of time ships what is green instead of being killed and shipping nothing.
295
307
  - **Hill phase is mechanical, never self-reported** — derived only from T0/T1/seesaw facts, closing
296
- the self-reported-confidence risk outright.
308
+ the self-reported-confidence risk — with one gap on record: a scope with no discovery
309
+ ledger derives the same phase as one whose unknowns are all closed, so absence still reads
310
+ as progress on that one arm.
297
311
  - **One writer per shared file** — every board/ledger/verdict write goes through
298
312
  `harness reduce ingest`; workers return data and never touch shared state.
299
313
  - **Traceability is oracle-checked, opt-in** — `harness verify trace` verifies covers-closure and
@@ -338,8 +352,8 @@ kernel/{verify,reduce,probe,init,report}/ # its subcommands, plus compile and
338
352
  kernel/lib/ # argv (the typed CLI boundary), paths (+ the run key), contract (shape)
339
353
  kernel/schemas/ # the envelope port: WorkOrder, WorkResult, domain registry
340
354
  commands/*.md # slash commands (/ship + the 9 phase commands)
341
- hooks/ # hooks.json + the four walls: safety-spine, gate-intake, sandbox-guard
342
- # (PreToolUse) + gate-zerowork (Stop, the one blocking hook)
355
+ hooks/ # hooks.json + the five walls: safety-spine, gate-intake, sandbox-guard,
356
+ # tier-guard (PreToolUse) + gate-zerowork (Stop, the one blocking hook)
343
357
  # + dispatch-receipt (PostToolUse, denies nothing, attests which skill ran)
344
358
  # + lib/decision.mjs (every hook records allow / deny / error)
345
359
  oracles/ # the evaluation-contract oracle registry (test · snapshot · http · process)
package/SECURITY.md CHANGED
@@ -1,7 +1,7 @@
1
1
  # Security
2
2
 
3
- This plugin installs **seven hook entries (six Node scripts + one `echo`)**: four in a
4
- `PreToolUse` position, **all four of which can deny a tool call**, one `PostToolUse` hook that has no
3
+ This plugin installs **eight hook entries (seven Node scripts + one `echo`)**: five in a
4
+ `PreToolUse` position, **all five of which can deny a tool call**, one `PostToolUse` hook that has no
5
5
  deny path at all, and one `Stop`-position hook that can block a session from ending. That is the
6
6
  product — and it is also exactly the kind of surface a careful reviewer should want spelled out
7
7
  before installing. This page is that spelling-out.
@@ -73,6 +73,7 @@ sitting, and reading them is the recommended review.
73
73
  | [`gate-intake.mjs`](hooks/gate-intake.mjs) | PreToolUse (`Skill`) | The `tech-lead` dispatch's own arguments | Yes — an orchestrator dispatch carrying no resolvable intake (no pitch, spec, resume or requirement text) | Fails open on `--order` and on any ambiguous arg shape |
74
74
  | [`harness verify envelope`](kernel/verify/envelope.mjs) | PreToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the JSON schemas | Yes — a worker dispatch whose order file is missing or schema-invalid | Never gates a dispatch that carries no `--order` (standalone skill use stays free) |
75
75
  | [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
76
+ | [`tier-guard.mjs`](hooks/tier-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the text the call is about to write (`content`, `new_string`, or any edit in a `MultiEdit` batch) | Yes — a write into `shapeup/<slug>/` whose content names a `.shapeup/` path or a `TASK-NNN` board id: the same tier-direction rule spec-lint reds at GATE L1b, refused at the moment of writing, with the offending token quoted back | No-op outside `shapeup/<slug>/`, on file forms the lint does not scan, and on `shapeup/knowledge-base/` (instructions to a worker, not references). Reads nothing off disk and writes only its decision row; it never inspects a file it is not being asked to write |
76
77
  | [`dispatch-receipt.mjs`](hooks/dispatch-receipt.mjs) | PostToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the tool result's own report of which skill ran | **No — it has no deny path at all.** It records that the shipped skill ran, so `harness reduce ingest` can refuse a result no dispatch produced | Never writes an attestation for a result that does not name a resolved skill; never fails the call it observes (every write is inside `try`/`catch`) |
77
78
  | [`gate-zerowork.mjs`](hooks/gate-zerowork.mjs) | Stop | Run receipts on disk; the session transcript; the decision ledger | **Yes — the one blocking hook.** Returns `decision:"block"` when the session dispatched the orchestrator and produced no run receipt | Defers the moment any receipt exists; `stop_hook_active` caps it at one block per stop chain |
78
79
 
package/hooks/hooks.json CHANGED
@@ -50,6 +50,16 @@
50
50
  "timeout": 10
51
51
  }
52
52
  ]
53
+ },
54
+ {
55
+ "matcher": "Edit|Write|MultiEdit",
56
+ "hooks": [
57
+ {
58
+ "type": "command",
59
+ "command": "node \"${CLAUDE_PLUGIN_ROOT}/hooks/tier-guard.mjs\"",
60
+ "timeout": 10
61
+ }
62
+ ]
53
63
  }
54
64
  ],
55
65
  "PostToolUse": [
@@ -0,0 +1,162 @@
1
+ #!/usr/bin/env node
2
+ // Tier guard — PreToolUse hook: the committed tier's write boundary.
3
+ //
4
+ // Refuses an Edit/Write/MultiEdit into `shapeup/<slug>/` whose CONTENT carries a reference the
5
+ // committed tier cannot hold — a path into the gitignored run trace, or a machine-local board id.
6
+ // It is the same rule spec-lint's TIER-DIRECTION already enforces, asked at a different moment.
7
+ //
8
+ // WHY A SECOND ENFORCEMENT POINT FOR A RULE THAT WAS NEVER IN DOUBT. Four different producers wrote
9
+ // committed files this harness's own lint then reds: the requirements registry, the ship report, the
10
+ // project profile, the coverage clauses. The lint was right every time and caught every one of them
11
+ // — at GATE L1b, a whole phase after the sentence was written, addressed to a worker that no longer
12
+ // holds the context needed to rephrase it. The run pays for the round trip, and the fix lands in
13
+ // whatever wording the next reader guesses at.
14
+ //
15
+ // TEACHING DOES NOT CLOSE IT, and that is measured rather than assumed. One of the four recurred
16
+ // INSIDE A SINGLE SESSION, an hour apart, on a different line of the same file, after that same
17
+ // author had fixed the first occurrence and watched its own lint go green. A worker carries no
18
+ // lesson across a dispatch. So the remedy has to be a precondition the model cannot talk past, and
19
+ // it has to speak AT THE WRITE — the one moment the writer still knows what it meant to say.
20
+ //
21
+ // THE PREDICATE IS IMPORTED, NEVER RESTATED. `tierLeaks` is the lint's own scanner
22
+ // (kernel/verify/spec.mjs), so this hook cannot become narrower than the rule it fronts — which
23
+ // would let the defect through to L1b exactly as before — nor wider, which is how a guard earns
24
+ // being switched off. The structural suite executes both over one corpus and requires agreement.
25
+ //
26
+ // Design (deliberately conservative, the shape `sandbox-guard.mjs` argues for):
27
+ // • Fail-OPEN on everything it cannot positively prove: not a write tool, no resolvable path, a
28
+ // path outside `shapeup/<slug>/`, a file form the lint does not scan, or a payload carrying no
29
+ // content to read. A guard that blocks legitimate work gets disabled, and a disabled guard
30
+ // enforces nothing.
31
+ // • `shapeup/knowledge-base/` is OUTSIDE by construction, exactly as it is outside the lint's
32
+ // walk: those files are instructions read by a worker at runtime, not references a reader is
33
+ // expected to resolve, and the defect register they hold has every reason to quote a local path.
34
+ // • Fail-CLOSED only on a leak the lint would red, with the file, the line, the offending token,
35
+ // why it is wrong and what to write instead — all in the denial, because a denial the writer
36
+ // cannot act on immediately just becomes the L1b round trip with extra steps.
37
+ //
38
+ // WHAT IT DOES NOT COVER, stated because the gap is the same one the substrate fence has: this is a
39
+ // PreToolUse hook, so it sees this assistant's own edit path and nothing else. A kernel subcommand
40
+ // writing a committed file, a shell heredoc, or any other editor writes straight through it. Those
41
+ // producers are answered where they are written; this closes the channel the four measured
42
+ // recurrences actually came through.
43
+ //
44
+ // Contract: PreToolUse stdin JSON { tool_name, tool_input:{file_path, content | new_string |
45
+ // edits[]}, cwd }. Deny via { hookSpecificOutput: { hookEventName, permissionDecision:"deny",
46
+ // permissionDecisionReason } }.
47
+
48
+ import { resolve, relative, sep } from "node:path";
49
+ import { isMain } from "../kernel/lib/argv.mjs";
50
+ import { SHARED } from "../kernel/lib/paths.mjs";
51
+ import { SCANNED, tierLeaks } from "../kernel/verify/spec.mjs";
52
+ import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
53
+
54
+ /** Committed-tier trees that hold instructions rather than references — outside the lint's walk. */
55
+ const NOT_A_SLUG = new Set(["knowledge-base"]);
56
+
57
+ /**
58
+ * The (path, content) pairs one write tool call is about to commit to disk.
59
+ *
60
+ * THE THREE WRITE TOOLS SPELL CONTENT THREE WAYS, and a guard that reads only `content` is blind to
61
+ * the Edit that adds the same line — which is the form the measured recurrence took, since the file
62
+ * already existed by then.
63
+ *
64
+ * @param {object} toolInput - The PreToolUse `tool_input` block.
65
+ * @returns {Array<{path:string, content:string}>} Pairs with a usable path and string content.
66
+ */
67
+ export function extractWrites(toolInput) {
68
+ const out = [];
69
+ const base = toolInput?.file_path;
70
+ if (typeof toolInput?.content === "string" && base) out.push({ path: base, content: toolInput.content });
71
+ if (typeof toolInput?.new_string === "string" && base) out.push({ path: base, content: toolInput.new_string });
72
+ if (Array.isArray(toolInput?.edits)) {
73
+ for (const e of toolInput.edits) {
74
+ const p = e?.file_path || base;
75
+ if (p && typeof e?.new_string === "string") out.push({ path: p, content: e.new_string });
76
+ }
77
+ }
78
+ return out;
79
+ }
80
+
81
+ /**
82
+ * Is this path inside the committed tier the lint walks — `shapeup/<slug>/…`, scanned form?
83
+ *
84
+ * @param {string} rel - Path relative to the project root, as the lint would name it.
85
+ * @returns {boolean} True when a leak here is a leak the lint would red.
86
+ */
87
+ export function inCommittedTier(rel) {
88
+ if (!rel || rel.startsWith("..") || rel.startsWith(sep)) return false;
89
+ const parts = rel.split(/[\\/]/);
90
+ // `shapeup/<slug>/<file>` — the lint walks a slug's tree, so a file at the tier root belongs to
91
+ // no run and is left alone, and `knowledge-base` is a sibling of the slugs, not one of them.
92
+ if (parts[0] !== SHARED || parts.length < 3 || NOT_A_SLUG.has(parts[1])) return false;
93
+ return SCANNED.test(rel);
94
+ }
95
+
96
+ async function main() {
97
+ await runHook("tier-guard", async () => {
98
+ const raw = await readStdin();
99
+ let p;
100
+ /** Fail-open, with the reason on the record (hooks/lib/decision.mjs). */
101
+ const defer = (reason, rule) => settle({
102
+ verdict: "allow", event: "PreToolUse", tool: p?.tool_name ?? null, cwd: p?.cwd, reason, rule,
103
+ });
104
+ try { p = JSON.parse(raw || "{}"); }
105
+ catch (e) { settle({ verdict: "error", event: "PreToolUse", reason: `unparseable payload: ${e.message}` }); }
106
+
107
+ if (!["Edit", "Write", "MultiEdit"].includes(p.tool_name)) {
108
+ defer(`${p.tool_name ?? "no tool_name"} is not a write tool — out of scope`, "not-write-tool");
109
+ }
110
+
111
+ // The shell's cwd is where the call fired; the tier lives at the project root. Same split the
112
+ // substrate fence has to make, and for the same reason: a worker that `cd`s into a sub-folder
113
+ // must not thereby leave the tier unguarded. Raw tool paths still resolve against the shell.
114
+ const cwd = p.cwd || process.cwd();
115
+ const root = projectRoot(cwd);
116
+
117
+ const writes = extractWrites(p.tool_input);
118
+ if (writes.length === 0) defer("no readable path+content pair in the tool input", "no-content");
119
+
120
+ const inTier = writes
121
+ .map((w) => ({ ...w, rel: relative(root, resolve(cwd, w.path)) }))
122
+ .filter((w) => inCommittedTier(w.rel));
123
+ if (inTier.length === 0) defer(`${writes.length} write(s), none inside ${SHARED}/<slug>/ in a scanned form`, "not-committed-tier");
124
+
125
+ const blocked = [];
126
+ for (const w of inTier) {
127
+ // THE TOKEN IS QUOTED BACK, which the lint's own message does not do and does not need to —
128
+ // it reports against a file on disk the reader can open at that line. Here the line does not
129
+ // exist yet, so "line 1 points into the local tier" leaves the writer hunting through a
130
+ // fragment it is holding in its head. Naming the exact string is the difference between a
131
+ // denial that is acted on and one that is retried verbatim.
132
+ for (const leak of tierLeaks(w.content)) blocked.push(`${w.rel}:${leak.line} — \`${leak.token}\` ${leak.detail}`);
133
+ }
134
+
135
+ if (blocked.length === 0) {
136
+ defer(`${inTier.length} committed-tier write(s) carry no tier-direction leak — permitted`, "tier-clean");
137
+ }
138
+
139
+ return {
140
+ verdict: "deny", event: "PreToolUse", tool: p.tool_name, subject: inTier[0].rel, cwd: root,
141
+ rule: "tier-direction",
142
+ reason: `${blocked.length} tier-direction leak(s) refused at the write boundary: ${blocked.join("; ")}`,
143
+ payload: {
144
+ hookSpecificOutput: {
145
+ hookEventName: "PreToolUse",
146
+ permissionDecision: "deny",
147
+ permissionDecisionReason:
148
+ "Tier guard (TIER-DIRECTION) — this write would put a reference into the committed tier that " +
149
+ "cannot survive the trip to another machine:\n" +
150
+ `${blocked.join("\n")}\n` +
151
+ `Refused here rather than at GATE L1b, where spec-lint reds the same file after the phase is over. ` +
152
+ "Rephrase the line now: cite the committed artifact, the use case or the scope_id, or describe the " +
153
+ "tier without naming a path.",
154
+ },
155
+ },
156
+ };
157
+ });
158
+ }
159
+
160
+ if (isMain(import.meta.url)) {
161
+ main();
162
+ }
@@ -218,6 +218,32 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
218
218
  // in its `allowed` globs, so only a `frozen` entry denies the write. `translate` is the one
219
219
  // operation that legitimately rewrites a pitch, and it writes the COMMITTED copy, not this one.
220
220
  const FROZEN_INTAKE = [`${local}/intake.md`, `${local}/breadboard.md`];
221
+ // THE LEG'S OWN ATTESTATION. A dispatch receipt, a leg-completion row and a T0 verdict exist to
222
+ // answer "did a worker actually run, and what did it measure" from evidence the leg being judged
223
+ // does not control — the whole reason to prefer them over a compiled order, which is writable by
224
+ // that same hand. None of the three is ever produced by this assistant's own edit tool in normal
225
+ // operation: a receipt is stamped by the dispatch hook reacting to the call itself, a
226
+ // leg-completion row by the step that reads a landed result, and a verdict by the command that
227
+ // runs the fixtures. A build leg reaching for `Edit`/`Write` on any of them is not doing its job
228
+ // by another route — it is handing itself the grade — so freezing them costs a well-behaved leg
229
+ // nothing and removes the one channel through which it could attest work it did not do. The
230
+ // carve-out below still covers everything else under this root — the doer's own task board and
231
+ // discovery ledger — because neither lives under any of these three.
232
+ //
233
+ // THE RESULT ENVELOPE IS DELIBERATELY NOT ON THIS LIST, and the reasoning is worth keeping because
234
+ // it looks like it belongs. It is the leg's own claim about its own work, so on the argument above
235
+ // it is the first thing you would freeze. But a result is not evidence ABOUT the leg, it is the
236
+ // leg's PRODUCT — the other half of the envelope port, and the thing that answers the order. The
237
+ // order is unanswered at that moment by construction, so freezing the path would deny every build
238
+ // leg its documented last step, every time, on the first round: measured end to end against a
239
+ // compiled order, with the denial telling the worker to widen a substrate that cannot lift a
240
+ // frozen entry. Nothing else writes an ordinary result either, so there is no fallback. The census
241
+ // that reads it is already built for this: a result alone attests nothing, and can only turn an
242
+ // ALREADY-receipted attempt into a spent one, which spends the forger's own budget. What that does
243
+ // not cover is a leg forging a SIBLING's result, which is a real hole with a different fix — the
244
+ // glob form here cannot say "every result except this order's own", so closing it needs a
245
+ // mechanism rather than one more entry on this list.
246
+ const FROZEN_ATTESTATION = [`${local}/receipts/**`, `${local}/legs.jsonl`, `${local}/t0/verdicts/**`];
221
247
  switch (operation) {
222
248
  case "execute": case "fix": case "spike":
223
249
  // Build legs are the widest window on FROZEN_INTAKE, not an exemption from it: they are the
@@ -228,7 +254,7 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
228
254
  return {
229
255
  allowed: [...(scope?.allowed_file_substrate || []), `${local}/spikes/**`],
230
256
  shared: scope?.shared_substrate || [],
231
- frozen: [...FROZEN_INTAKE],
257
+ frozen: [...FROZEN_INTAKE, ...FROZEN_ATTESTATION],
232
258
  };
233
259
  case "analyze":
234
260
  return { allowed: [`${spec}/**`, `${local}/**`], frozen: [...FROZEN_INTAKE] };
@@ -1007,7 +1033,21 @@ export async function cli(rawArgv) {
1007
1033
  // than the flailing it detects. It advises; the orchestrator queues the GATE H proposal.
1008
1034
  if (scope?.scope_id) {
1009
1035
  const k = Number(scope.no_progress_k ?? payloadExtra.no_progress_k ?? 2);
1010
- const st = stagnation(allTrials.filter((t) => t.scope_id === scope.scope_id), k);
1036
+ // SCOPED TO THIS RUN, for the same reason the attempt census is. `t0/trials.jsonl` is per-slug
1037
+ // and append-only, so a streak survives the run that produced it. Measured on a consumer: two
1038
+ // non-kept trials from earlier runs — one of them graded against an order no worker was ever
1039
+ // dispatched for — read as a stagnation streak that every later run of that scope tripped on,
1040
+ // seven hours after the tree they graded had been fixed. The breaker escalated on evidence that
1041
+ // predated its own fix, and no attempt of the tripping run had been dispatched at all.
1042
+ //
1043
+ // A trial with no run key counts for no run; an unresolvable current run matches nothing. Both
1044
+ // leave the breaker un-tripped, which is the fail-open direction this repo's guards take when
1045
+ // the bad state cannot be positively proven.
1046
+ const myRunId = readRunId(cwd, slug);
1047
+ const st = stagnation(
1048
+ allTrials.filter((t) => t.scope_id === scope.scope_id && myRunId != null && t.run_id === myRunId),
1049
+ k,
1050
+ );
1011
1051
  if (st.stagnant) {
1012
1052
  console.error(JSON.stringify({
1013
1053
  breaker: "stagnation", scope_id: scope.scope_id, streak: st.streak, no_progress_k: st.k,
package/kernel/gate.mjs CHANGED
@@ -53,7 +53,7 @@
53
53
  import { readFileSync, writeFileSync, appendFileSync, existsSync, mkdirSync } from "node:fs";
54
54
  import { join, dirname } from "node:path";
55
55
  import { runArgs } from "./lib/argv.mjs";
56
- import { gateAnswerCandidates, gates as gatesPath, LOCAL } from "./lib/paths.mjs";
56
+ import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir } from "./lib/paths.mjs";
57
57
 
58
58
  export const GATE_IDS = ["L0", "L1a", "L1a.5", "L1b", "L2", "L3", "QA", "H", "L4", "COACH-1"];
59
59
 
@@ -110,7 +110,7 @@ export const PRESETS = {
110
110
  "L1a": { decision: "proceed", note: "Orient review — advisory read." },
111
111
  "L1a.5": { decision: "proceed", note: "Wiring review — checked by trace-lint." },
112
112
  "L1b": { decision: "ask", note: "Board review is where scope is actually decided. Not pre-approvable." },
113
- "L2": { decision: "proceed", note: "Board-green is verified by hook, not by opinion." },
113
+ "L2": { decision: "proceed", note: "The board facts travel in the gate block itself — green_scopes and hammer_proposals — so a preset answering here is not answering blind. Note there is no board check behind this: the L2 hook was retired into the gate block in v2.0." },
114
114
  "L3": { decision: "loop", max_rounds: 3, note: "Loop on FAIL; the breaker ends it." },
115
115
  "QA": { decision: "run" },
116
116
  "H": { decision: "ask", note: "The cut list changes what ships." },
@@ -168,6 +168,71 @@ export function validate(set) {
168
168
  * Returns { gate, decision, source, note, status } where status ∈ ok | ask | abort.
169
169
  * Pure: the CLI turns `status` into the exit code, nothing here exits.
170
170
  */
171
+ /** The three verdicts `scope-hammer` may return (shapeup-run.js's HAMMER schema). */
172
+ export const HAMMER_VERDICTS = ["ship-now", "ship-after-fixes", "cannot-ship"];
173
+
174
+ /**
175
+ * The census verdict on disk for one run, or null when no census has run.
176
+ *
177
+ * Derived from the hammer's own WorkResult rather than taken from a caller: the verdict is the
178
+ * product of a dispatch, and a gate that accepts it as an argument accepts whatever the caller
179
+ * believes. `null` is a real answer here — "scope-hammer has not run" — and is the state a run that
180
+ * returned `gate_h` from a circuit breaker is in, because the breaker returns long before the
181
+ * hammer is dispatched.
182
+ *
183
+ * @param {string} cwd - Project root.
184
+ * @param {string} slug - Feature slug.
185
+ * @returns {(string|null)} The verdict, or null when there is no readable census.
186
+ */
187
+ export function censusVerdict(cwd, slug) {
188
+ if (!slug) return null;
189
+ try {
190
+ const p = join(resultsDir(cwd, slug), "hammer.json");
191
+ if (!existsSync(p)) return null;
192
+ const r = JSON.parse(readFileSync(p, "utf8"));
193
+ const v = r?.payload?.verdict ?? r?.verdict ?? null;
194
+ return HAMMER_VERDICTS.includes(v) ? v : null;
195
+ } catch { return null; }
196
+ }
197
+
198
+ /**
199
+ * Narrow a resolved gate answer to what the run's own evidence supports.
200
+ *
201
+ * THE RULE, and it is about answers rather than callers. An answer set chooses among the answers a
202
+ * gate ALLOWS; it can never supply the evidence that makes one allowed. `ship` at L4 asserts the
203
+ * census cleared the run, so `ship` is only on the menu when a census exists and did clear it.
204
+ *
205
+ * Measured on a consumer: a run whose dispatch receipts carried `orient` and `task-executor` and
206
+ * nothing else — `scope-hammer` never dispatched — recorded `H → accept-cut-list` and `L4 → ship`
207
+ * from the `ci` preset, while its own ledger read `status: escalated`, `final_verdict: ~`, every
208
+ * requirement at `no evidence`. `shapeup-run.js` does guard this, but only after the hammer
209
+ * dispatch, so a run that returns `gate_h` from a breaker reaches L4 by a route the guard does not
210
+ * cover. A check on one call site is not a property.
211
+ *
212
+ * Absence of a census disqualifies `ship` on its own: "no one looked" and "someone looked and it
213
+ * was fine" are different facts, and only the second warrants a ship.
214
+ *
215
+ * @param {object} r - A resolved gate result from {@link resolve}.
216
+ * @param {string} cwd - Project root.
217
+ * @param {(string|null)} slug - Feature slug, when the gate is being resolved inside a run.
218
+ * @returns {object} The result, or an `ask` carrying why `ship` was not available.
219
+ */
220
+ export function narrowToEvidence(r, cwd, slug) {
221
+ if (r?.gate !== "L4" || r?.status !== "ok" || r?.decision !== "ship") return r;
222
+ const verdict = censusVerdict(cwd, slug);
223
+ if (verdict === "ship-now" || verdict === "ship-after-fixes") return r;
224
+ const why = verdict === "cannot-ship"
225
+ ? "scope-hammer's census returned CANNOT SHIP"
226
+ : "scope-hammer has not run, so no census exists";
227
+ return {
228
+ gate: r.gate, status: "ask", source: r.source, decision: "ask", note: r.note,
229
+ refused: "ship", census: verdict,
230
+ reason: `GATE L4 cannot be answered "ship": ${why}. An answer set chooses among the answers a ` +
231
+ `gate allows; it cannot supply the evidence that makes one allowed. Put the block to the ` +
232
+ `PO, or run the census first.`,
233
+ };
234
+ }
235
+
171
236
  export function resolve(set, gate, source) {
172
237
  if (!GATE_IDS.includes(gate)) {
173
238
  return { gate, status: "error", reason: `unknown gate "${gate}" — known: ${GATE_IDS.join(", ")}` };
@@ -362,7 +427,7 @@ export function cli(rawArgv) {
362
427
  const gate = args.resolve ?? null;
363
428
  if (!gate) die("nothing to do — pass --init, --list, --verify, or --resolve <gate-id>");
364
429
 
365
- const r = resolve(found.set, gate, found.source);
430
+ const r = narrowToEvidence(resolve(found.set, gate, found.source), cwd, args.slug ?? null);
366
431
  if (r.status === "error") die(r.reason);
367
432
  // A gate with no `--slug` (e.g. `--file` used ad hoc, outside any run) has nowhere to file a
368
433
  // per-run ledger row — the same reasoning `resolveRunId` uses for "no run is active": absence is
@@ -30,7 +30,7 @@
30
30
  import { existsSync, readdirSync, readFileSync } from "node:fs";
31
31
  import { join, resolve } from "node:path";
32
32
  import { runArgs } from "../lib/argv.mjs";
33
- import { dispatchReceipts, legLedger, resultsDir } from "../lib/paths.mjs";
33
+ import { dispatchReceipts, legLedger, resultsDir, readRunId } from "../lib/paths.mjs";
34
34
  import { readLegs } from "./leg.mjs";
35
35
  import { greenVerdict } from "./t0.mjs";
36
36
 
@@ -64,10 +64,26 @@ export function readReceipts(path) {
64
64
  * @param {object[]} legs - Pre-read `legs.jsonl` rows.
65
65
  * @returns {{orderId:string, hasReceipt:boolean, hasResult:boolean, hasLeg:boolean, state:("unattested"|"in-flight"|"spent")}}
66
66
  */
67
- export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, legs) {
67
+ export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, legs, runId = null) {
68
68
  const orderId = `${slug}/${scopeId}-r${round}-a${attempt}`;
69
- const hasReceipt = receipts.some((r) => r?.order_id === orderId);
70
- const hasLeg = legs.some((r) => r?.order_id === orderId);
69
+ // SCOPED TO THIS RUN, and that qualifier is the whole correction. `receipts/dispatch.jsonl` and
70
+ // `legs.jsonl` are per-slug and append-only, so they accumulate across every run of a pitch,
71
+ // while `order_id` repeats — the run key is the only thing that separates two runs of one
72
+ // feature. Matching on `order_id` alone answered "was this attempt spent?" with a PREVIOUS run's
73
+ // receipt: measured on a consumer, a run that dispatched nothing was told its first attempt was
74
+ // spent, complete with leg and result, by rows two launches old.
75
+ //
76
+ // A row carrying no run key belongs to NO run rather than to this one, and an unresolvable
77
+ // current run (no receipt on disk) matches nothing. Both directions under-count rather than
78
+ // over-count, which is the safe way to be wrong here: an under-count leaves a breaker un-tripped
79
+ // and the round continues, where an over-count stops work that was never done.
80
+ const mine = (r) => r?.order_id === orderId && runId != null && r?.run_id === runId;
81
+ const hasReceipt = receipts.some(mine);
82
+ const hasLeg = legs.some(mine);
83
+ // A WorkResult carries no run key and reaches one only through its `order_id`, which repeats —
84
+ // so this stays a file check and is deliberately NOT sufficient on its own. It can only turn an
85
+ // already run-scoped receipt into `spent`; a result left behind by an earlier run cannot attest
86
+ // an attempt this run never dispatched.
71
87
  const hasResult = existsSync(join(resultsDir(cwd, slug), `${scopeId}-r${round}-a${attempt}.json`));
72
88
  const state = !hasReceipt ? "unattested" : (hasResult || hasLeg) ? "spent" : "in-flight";
73
89
  return { orderId, hasReceipt, hasResult, hasLeg, state };
@@ -91,10 +107,11 @@ export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, le
91
107
  export function scopeAttempts(cwd, slug, scopeId, round, attemptBudget) {
92
108
  const receipts = readReceipts(dispatchReceipts(cwd, slug));
93
109
  const legs = readLegs(legLedger(cwd, slug));
110
+ const runId = readRunId(cwd, slug);
94
111
  const attempts = [];
95
112
  let spent = 0, inFlight = 0, unattested = 0;
96
113
  for (let a = 1; a <= attemptBudget; a++) {
97
- const ev = attemptEvidence(cwd, slug, scopeId, round, a, receipts, legs);
114
+ const ev = attemptEvidence(cwd, slug, scopeId, round, a, receipts, legs, runId);
98
115
  if (ev.state === "spent") spent++;
99
116
  else if (ev.state === "in-flight") inFlight++;
100
117
  else unattested++;