shapeup-sdlc 3.7.0 → 3.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.7.0",
4
+ "version": "3.7.1",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -7,7 +7,8 @@ A three-phase Shape Up loop orchestrated by `/tech-lead`. Invariants live in the
7
7
 
8
8
  Skills and commands are named short throughout this file; every one of them resolves only under the plugin's namespace — `Skill(shapeup-sdlc-plugin:orient)`, `/shapeup-sdlc-plugin:ship`. A bare name is not a typo you get a warning for: an unknown skill name is rejected upstream of the hook layer, so the dispatch fires no hook, leaves no decision row, and is answered by improvisation.
9
9
 
10
- - Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, stopping a run with no receipt (the run's first act writes one).
10
+ - Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
11
+ - **The tier-direction rule is refused at the write, not argued at L1b.** A committed file cannot carry a `.shapeup/` path or a `TASK-NNN` id — boards renumber per machine and the run trace is gitignored, so either one resolves only where it was written. Spec-lint has always red'd this at Board Review; the same rule now also refuses the write itself, naming the token, because four producers hit it and one of them hit it twice in one session — a worker carries no lesson across a dispatch, and by L1b the phase that wrote the line is over. The knowledge base is outside the rule: those files are instructions, not references anyone resolves.
11
12
  - Attested, not assumed: a dispatch leaves a receipt naming the skill that ran, and ingest refuses an orchestrated result that has none. A dispatch that fails is answered by the sub-agent improvising the craft, which every other check accepts — so "the artifact exists" is not evidence that the shipped skill produced it. `--no-receipt-check` is the way through when the receipt channel itself fails.
12
13
  - GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
13
14
  - Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
@@ -80,7 +81,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
80
81
  through the chained launch, until the run can resume unattended again.
81
82
  - Two storage tiers (ADR-0001): COMMITTED `shapeup/<slug>/` (shaping, spec, scopes, wiring-map, project-profile, requirements, hill, `REPORT.md` frozen at L4) vs GITIGNORED `.shapeup/` (board, orders/results, T0/eval/QA artifacts, ledgers, metrics, gate answers, and the run scripts staged for launch).
82
83
  - **The run launches from a copy inside your project, and it has to.** The Workflow tool loads a script only from a directory the session may already read; the plugin installs outside your project, so naming the shipped path is refused before the run begins and no permission rule repairs it — the grant authorises the tool, not what it may read. Opening a run therefore re-copies the run scripts to `.shapeup/workflows/` and reports the path the launch names. A run in flight keeps the copy it started with: an upgrade reaches the next run, not the current round.
83
- - **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — can leave one still open, and while the run's pointer is still on disk the fence holds on that order exactly as if the run were live. The pointer is what actually switches it off, though: the fence is enforced only while that pointer exists on disk, so a close that removes it releases the fence whatever the order set still says — restoring the pointer flips it straight back to denying. A **ship** close does not itself answer what it leaves outstanding — a ship close can retire the pointer over an order that never got a result — so it is not special because nothing is left unanswered; it is special only because retiring the pointer is the one lever every close needs pulled, and `reduce ship` pulls it for you. Whichever way a run closes, an order it leaves unanswered stays genuinely unresolved — not merely un-fenced — until `init run --force` runs: it writes a synthetic result for every order the closed run left unanswered, so the fence lifts without waiting on a worker that is never coming back. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
84
+ - **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — leaves an order still open by construction, and the pointer is what decides whether that still fences: the fence is enforced only while the pointer exists on disk, so a close that removes it releases the fence whatever the order set says, and restoring the pointer flips it straight back to denying. **Every terminal close retires it** (3.7.1 onward), not only a ship — an escalated or aborted run stops fencing the checkout the moment it closes, and a close that cannot remove a pointer says so in its own return rather than failing. Retiring is not answering, though, and the difference is the whole reason there are still two levers: the pointer says a run is in flight, the missing result says nobody came back, and only the first stops being true at a close. An order the run left unanswered stays genuinely unresolved — un-fenced, but not done — until `init run --force` writes a synthetic result for it, which is also what lets the attempt census keep telling work nobody did apart from work that failed. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
84
85
  - Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows, agent-call journal rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded — and so does **every other terminal ending**: a run that aborts, escalates or stops at GATE H exports on its close, because the runs whose records are worth most are the ones that never reach the Ship phase. The export is advisory at every one of them: a failed export degrades the trace and is reported in the run's state warnings, but never turns a close into a non-close. A WorkResult carries no `run_id` and reaches it through `order_id`.
85
86
  - Every run projects a **run graph** — `.shapeup/<slug>/graph.jsonl`, append-only, written only by
86
87
  `reduce graph`. Two families kept separate: work lineage (Run, Order, Result, Verdict, Trial,
package/README.md CHANGED
@@ -210,7 +210,7 @@ the layer that carries it, and the three layers here fail differently:
210
210
  | **Runtime** — the kernel and the run script | When the run goes through the harness. A schema rejection or a non-zero exit stops the step. | A lane that never calls the kernel is never checked — which is why the hooks below cover the doors, not the steps. |
211
211
  | **Advisory** — a report section | When somebody reads the artifact. | Silently, if nobody does. It is a cleanup list, never a verdict. |
212
212
 
213
- **Four walls.** These are hooks because nothing in the runtime can substitute for them:
213
+ **Five walls.** These are hooks because nothing in the runtime can substitute for them:
214
214
 
215
215
  - `PreToolUse` (`Skill`) — **`hooks/gate-intake.mjs` denies a `tech-lead` dispatch that carries no
216
216
  pitch, no spec folder, and no requirement text.** Observed, not theorized: when the requirement
@@ -227,6 +227,13 @@ the layer that carries it, and the three layers here fail differently:
227
227
  checked first and outranks everything, across every live contract — including the carve-out that
228
228
  otherwise keeps the active feature's own run trace writable, so a path a live order froze stays
229
229
  frozen wherever it lives.
230
+ - `PreToolUse` (`Edit|Write|MultiEdit`) — **`hooks/tier-guard.mjs` refuses a committed-tier write
231
+ whose content names a path into the local run trace or a board id.** Same rule spec-lint reds at
232
+ GATE L1b, asked at the moment of writing: four different producers wrote a committed file their
233
+ own lint then rejected, and one of them did it twice in one session, an hour apart, after fixing
234
+ the first occurrence itself. A worker carries no lesson across a dispatch, so the remedy is a
235
+ refusal the writer can act on while it still knows what it meant to say. The knowledge base is
236
+ outside it — those files are instructions, not references a reader has to resolve.
230
237
  - `PreToolUse` (`Bash|Read|Write|Edit|MultiEdit`) — **`hooks/safety-spine.mjs` denies destructive
231
238
  commands** (`rm -rf` on unrecoverable targets, force-push/push-to-main, `git reset --hard`,
232
239
  `DROP TABLE`) and secret-file reads. A machine guard, not a pipeline guard; the escape hatch is
@@ -265,7 +272,7 @@ the layer that carries it, and the three layers here fail differently:
265
272
  | `anti-rationalization` (claims the facts contradict) | The ship report's census, derived from the board and the T0 artifacts | The facts are in an artifact a teammate finds on `git pull`, not in a transcript nobody re-reads. |
266
273
  | `slop-cleaner` (TODO/`console.log` leftovers) | The ship report's **Leftovers** section | Same scan, same added-lines-only rule; it lands somewhere checkable. |
267
274
 
268
- **Nothing load-bearing depends on permission mode.** The four walls plus the zero-work gate run
275
+ **Nothing load-bearing depends on permission mode.** The five walls plus the zero-work gate run
269
276
  under every mode. The kernel needs a grant to be *invoked* — two Bash lines `npx shapeup-sdlc init`
270
277
  writes — but a session that never gets that grant is a session that cannot run the pipeline at all,
271
278
  not one that runs it unguarded.
@@ -338,8 +345,8 @@ kernel/{verify,reduce,probe,init,report}/ # its subcommands, plus compile and
338
345
  kernel/lib/ # argv (the typed CLI boundary), paths (+ the run key), contract (shape)
339
346
  kernel/schemas/ # the envelope port: WorkOrder, WorkResult, domain registry
340
347
  commands/*.md # slash commands (/ship + the 9 phase commands)
341
- hooks/ # hooks.json + the four walls: safety-spine, gate-intake, sandbox-guard
342
- # (PreToolUse) + gate-zerowork (Stop, the one blocking hook)
348
+ hooks/ # hooks.json + the five walls: safety-spine, gate-intake, sandbox-guard,
349
+ # tier-guard (PreToolUse) + gate-zerowork (Stop, the one blocking hook)
343
350
  # + dispatch-receipt (PostToolUse, denies nothing, attests which skill ran)
344
351
  # + lib/decision.mjs (every hook records allow / deny / error)
345
352
  oracles/ # the evaluation-contract oracle registry (test · snapshot · http · process)
package/SECURITY.md CHANGED
@@ -1,7 +1,7 @@
1
1
  # Security
2
2
 
3
- This plugin installs **seven hook entries (six Node scripts + one `echo`)**: four in a
4
- `PreToolUse` position, **all four of which can deny a tool call**, one `PostToolUse` hook that has no
3
+ This plugin installs **eight hook entries (seven Node scripts + one `echo`)**: five in a
4
+ `PreToolUse` position, **all five of which can deny a tool call**, one `PostToolUse` hook that has no
5
5
  deny path at all, and one `Stop`-position hook that can block a session from ending. That is the
6
6
  product — and it is also exactly the kind of surface a careful reviewer should want spelled out
7
7
  before installing. This page is that spelling-out.
@@ -73,6 +73,7 @@ sitting, and reading them is the recommended review.
73
73
  | [`gate-intake.mjs`](hooks/gate-intake.mjs) | PreToolUse (`Skill`) | The `tech-lead` dispatch's own arguments | Yes — an orchestrator dispatch carrying no resolvable intake (no pitch, spec, resume or requirement text) | Fails open on `--order` and on any ambiguous arg shape |
74
74
  | [`harness verify envelope`](kernel/verify/envelope.mjs) | PreToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the JSON schemas | Yes — a worker dispatch whose order file is missing or schema-invalid | Never gates a dispatch that carries no `--order` (standalone skill use stays free) |
75
75
  | [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
76
+ | [`tier-guard.mjs`](hooks/tier-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the text the call is about to write (`content`, `new_string`, or any edit in a `MultiEdit` batch) | Yes — a write into `shapeup/<slug>/` whose content names a `.shapeup/` path or a `TASK-NNN` board id: the same tier-direction rule spec-lint reds at GATE L1b, refused at the moment of writing, with the offending token quoted back | No-op outside `shapeup/<slug>/`, on file forms the lint does not scan, and on `shapeup/knowledge-base/` (instructions to a worker, not references). Reads nothing off disk and writes only its decision row; it never inspects a file it is not being asked to write |
76
77
  | [`dispatch-receipt.mjs`](hooks/dispatch-receipt.mjs) | PostToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the tool result's own report of which skill ran | **No — it has no deny path at all.** It records that the shipped skill ran, so `harness reduce ingest` can refuse a result no dispatch produced | Never writes an attestation for a result that does not name a resolved skill; never fails the call it observes (every write is inside `try`/`catch`) |
77
78
  | [`gate-zerowork.mjs`](hooks/gate-zerowork.mjs) | Stop | Run receipts on disk; the session transcript; the decision ledger | **Yes — the one blocking hook.** Returns `decision:"block"` when the session dispatched the orchestrator and produced no run receipt | Defers the moment any receipt exists; `stop_hook_active` caps it at one block per stop chain |
78
79
 
package/hooks/hooks.json CHANGED
@@ -50,6 +50,16 @@
50
50
  "timeout": 10
51
51
  }
52
52
  ]
53
+ },
54
+ {
55
+ "matcher": "Edit|Write|MultiEdit",
56
+ "hooks": [
57
+ {
58
+ "type": "command",
59
+ "command": "node \"${CLAUDE_PLUGIN_ROOT}/hooks/tier-guard.mjs\"",
60
+ "timeout": 10
61
+ }
62
+ ]
53
63
  }
54
64
  ],
55
65
  "PostToolUse": [
@@ -0,0 +1,162 @@
1
+ #!/usr/bin/env node
2
+ // Tier guard — PreToolUse hook: the committed tier's write boundary.
3
+ //
4
+ // Refuses an Edit/Write/MultiEdit into `shapeup/<slug>/` whose CONTENT carries a reference the
5
+ // committed tier cannot hold — a path into the gitignored run trace, or a machine-local board id.
6
+ // It is the same rule spec-lint's TIER-DIRECTION already enforces, asked at a different moment.
7
+ //
8
+ // WHY A SECOND ENFORCEMENT POINT FOR A RULE THAT WAS NEVER IN DOUBT. Four different producers wrote
9
+ // committed files this harness's own lint then reds: the requirements registry, the ship report, the
10
+ // project profile, the coverage clauses. The lint was right every time and caught every one of them
11
+ // — at GATE L1b, a whole phase after the sentence was written, addressed to a worker that no longer
12
+ // holds the context needed to rephrase it. The run pays for the round trip, and the fix lands in
13
+ // whatever wording the next reader guesses at.
14
+ //
15
+ // TEACHING DOES NOT CLOSE IT, and that is measured rather than assumed. One of the four recurred
16
+ // INSIDE A SINGLE SESSION, an hour apart, on a different line of the same file, after that same
17
+ // author had fixed the first occurrence and watched its own lint go green. A worker carries no
18
+ // lesson across a dispatch. So the remedy has to be a precondition the model cannot talk past, and
19
+ // it has to speak AT THE WRITE — the one moment the writer still knows what it meant to say.
20
+ //
21
+ // THE PREDICATE IS IMPORTED, NEVER RESTATED. `tierLeaks` is the lint's own scanner
22
+ // (kernel/verify/spec.mjs), so this hook cannot become narrower than the rule it fronts — which
23
+ // would let the defect through to L1b exactly as before — nor wider, which is how a guard earns
24
+ // being switched off. The structural suite executes both over one corpus and requires agreement.
25
+ //
26
+ // Design (deliberately conservative, the shape `sandbox-guard.mjs` argues for):
27
+ // • Fail-OPEN on everything it cannot positively prove: not a write tool, no resolvable path, a
28
+ // path outside `shapeup/<slug>/`, a file form the lint does not scan, or a payload carrying no
29
+ // content to read. A guard that blocks legitimate work gets disabled, and a disabled guard
30
+ // enforces nothing.
31
+ // • `shapeup/knowledge-base/` is OUTSIDE by construction, exactly as it is outside the lint's
32
+ // walk: those files are instructions read by a worker at runtime, not references a reader is
33
+ // expected to resolve, and the defect register they hold has every reason to quote a local path.
34
+ // • Fail-CLOSED only on a leak the lint would red, with the file, the line, the offending token,
35
+ // why it is wrong and what to write instead — all in the denial, because a denial the writer
36
+ // cannot act on immediately just becomes the L1b round trip with extra steps.
37
+ //
38
+ // WHAT IT DOES NOT COVER, stated because the gap is the same one the substrate fence has: this is a
39
+ // PreToolUse hook, so it sees this assistant's own edit path and nothing else. A kernel subcommand
40
+ // writing a committed file, a shell heredoc, or any other editor writes straight through it. Those
41
+ // producers are answered where they are written; this closes the channel the four measured
42
+ // recurrences actually came through.
43
+ //
44
+ // Contract: PreToolUse stdin JSON { tool_name, tool_input:{file_path, content | new_string |
45
+ // edits[]}, cwd }. Deny via { hookSpecificOutput: { hookEventName, permissionDecision:"deny",
46
+ // permissionDecisionReason } }.
47
+
48
+ import { resolve, relative, sep } from "node:path";
49
+ import { isMain } from "../kernel/lib/argv.mjs";
50
+ import { SHARED } from "../kernel/lib/paths.mjs";
51
+ import { SCANNED, tierLeaks } from "../kernel/verify/spec.mjs";
52
+ import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
53
+
54
+ /** Committed-tier trees that hold instructions rather than references — outside the lint's walk. */
55
+ const NOT_A_SLUG = new Set(["knowledge-base"]);
56
+
57
+ /**
58
+ * The (path, content) pairs one write tool call is about to commit to disk.
59
+ *
60
+ * THE THREE WRITE TOOLS SPELL CONTENT THREE WAYS, and a guard that reads only `content` is blind to
61
+ * the Edit that adds the same line — which is the form the measured recurrence took, since the file
62
+ * already existed by then.
63
+ *
64
+ * @param {object} toolInput - The PreToolUse `tool_input` block.
65
+ * @returns {Array<{path:string, content:string}>} Pairs with a usable path and string content.
66
+ */
67
+ export function extractWrites(toolInput) {
68
+ const out = [];
69
+ const base = toolInput?.file_path;
70
+ if (typeof toolInput?.content === "string" && base) out.push({ path: base, content: toolInput.content });
71
+ if (typeof toolInput?.new_string === "string" && base) out.push({ path: base, content: toolInput.new_string });
72
+ if (Array.isArray(toolInput?.edits)) {
73
+ for (const e of toolInput.edits) {
74
+ const p = e?.file_path || base;
75
+ if (p && typeof e?.new_string === "string") out.push({ path: p, content: e.new_string });
76
+ }
77
+ }
78
+ return out;
79
+ }
80
+
81
+ /**
82
+ * Is this path inside the committed tier the lint walks — `shapeup/<slug>/…`, scanned form?
83
+ *
84
+ * @param {string} rel - Path relative to the project root, as the lint would name it.
85
+ * @returns {boolean} True when a leak here is a leak the lint would red.
86
+ */
87
+ export function inCommittedTier(rel) {
88
+ if (!rel || rel.startsWith("..") || rel.startsWith(sep)) return false;
89
+ const parts = rel.split(/[\\/]/);
90
+ // `shapeup/<slug>/<file>` — the lint walks a slug's tree, so a file at the tier root belongs to
91
+ // no run and is left alone, and `knowledge-base` is a sibling of the slugs, not one of them.
92
+ if (parts[0] !== SHARED || parts.length < 3 || NOT_A_SLUG.has(parts[1])) return false;
93
+ return SCANNED.test(rel);
94
+ }
95
+
96
+ async function main() {
97
+ await runHook("tier-guard", async () => {
98
+ const raw = await readStdin();
99
+ let p;
100
+ /** Fail-open, with the reason on the record (hooks/lib/decision.mjs). */
101
+ const defer = (reason, rule) => settle({
102
+ verdict: "allow", event: "PreToolUse", tool: p?.tool_name ?? null, cwd: p?.cwd, reason, rule,
103
+ });
104
+ try { p = JSON.parse(raw || "{}"); }
105
+ catch (e) { settle({ verdict: "error", event: "PreToolUse", reason: `unparseable payload: ${e.message}` }); }
106
+
107
+ if (!["Edit", "Write", "MultiEdit"].includes(p.tool_name)) {
108
+ defer(`${p.tool_name ?? "no tool_name"} is not a write tool — out of scope`, "not-write-tool");
109
+ }
110
+
111
+ // The shell's cwd is where the call fired; the tier lives at the project root. Same split the
112
+ // substrate fence has to make, and for the same reason: a worker that `cd`s into a sub-folder
113
+ // must not thereby leave the tier unguarded. Raw tool paths still resolve against the shell.
114
+ const cwd = p.cwd || process.cwd();
115
+ const root = projectRoot(cwd);
116
+
117
+ const writes = extractWrites(p.tool_input);
118
+ if (writes.length === 0) defer("no readable path+content pair in the tool input", "no-content");
119
+
120
+ const inTier = writes
121
+ .map((w) => ({ ...w, rel: relative(root, resolve(cwd, w.path)) }))
122
+ .filter((w) => inCommittedTier(w.rel));
123
+ if (inTier.length === 0) defer(`${writes.length} write(s), none inside ${SHARED}/<slug>/ in a scanned form`, "not-committed-tier");
124
+
125
+ const blocked = [];
126
+ for (const w of inTier) {
127
+ // THE TOKEN IS QUOTED BACK, which the lint's own message does not do and does not need to —
128
+ // it reports against a file on disk the reader can open at that line. Here the line does not
129
+ // exist yet, so "line 1 points into the local tier" leaves the writer hunting through a
130
+ // fragment it is holding in its head. Naming the exact string is the difference between a
131
+ // denial that is acted on and one that is retried verbatim.
132
+ for (const leak of tierLeaks(w.content)) blocked.push(`${w.rel}:${leak.line} — \`${leak.token}\` ${leak.detail}`);
133
+ }
134
+
135
+ if (blocked.length === 0) {
136
+ defer(`${inTier.length} committed-tier write(s) carry no tier-direction leak — permitted`, "tier-clean");
137
+ }
138
+
139
+ return {
140
+ verdict: "deny", event: "PreToolUse", tool: p.tool_name, subject: inTier[0].rel, cwd: root,
141
+ rule: "tier-direction",
142
+ reason: `${blocked.length} tier-direction leak(s) refused at the write boundary: ${blocked.join("; ")}`,
143
+ payload: {
144
+ hookSpecificOutput: {
145
+ hookEventName: "PreToolUse",
146
+ permissionDecision: "deny",
147
+ permissionDecisionReason:
148
+ "Tier guard (TIER-DIRECTION) — this write would put a reference into the committed tier that " +
149
+ "cannot survive the trip to another machine:\n" +
150
+ `${blocked.join("\n")}\n` +
151
+ `Refused here rather than at GATE L1b, where spec-lint reds the same file after the phase is over. ` +
152
+ "Rephrase the line now: cite the committed artifact, the use case or the scope_id, or describe the " +
153
+ "tier without naming a path.",
154
+ },
155
+ },
156
+ };
157
+ });
158
+ }
159
+
160
+ if (isMain(import.meta.url)) {
161
+ main();
162
+ }
@@ -1007,7 +1007,21 @@ export async function cli(rawArgv) {
1007
1007
  // than the flailing it detects. It advises; the orchestrator queues the GATE H proposal.
1008
1008
  if (scope?.scope_id) {
1009
1009
  const k = Number(scope.no_progress_k ?? payloadExtra.no_progress_k ?? 2);
1010
- const st = stagnation(allTrials.filter((t) => t.scope_id === scope.scope_id), k);
1010
+ // SCOPED TO THIS RUN, for the same reason the attempt census is. `t0/trials.jsonl` is per-slug
1011
+ // and append-only, so a streak survives the run that produced it. Measured on a consumer: two
1012
+ // non-kept trials from earlier runs — one of them graded against an order no worker was ever
1013
+ // dispatched for — read as a stagnation streak that every later run of that scope tripped on,
1014
+ // seven hours after the tree they graded had been fixed. The breaker escalated on evidence that
1015
+ // predated its own fix, and no attempt of the tripping run had been dispatched at all.
1016
+ //
1017
+ // A trial with no run key counts for no run; an unresolvable current run matches nothing. Both
1018
+ // leave the breaker un-tripped, which is the fail-open direction this repo's guards take when
1019
+ // the bad state cannot be positively proven.
1020
+ const myRunId = readRunId(cwd, slug);
1021
+ const st = stagnation(
1022
+ allTrials.filter((t) => t.scope_id === scope.scope_id && myRunId != null && t.run_id === myRunId),
1023
+ k,
1024
+ );
1011
1025
  if (st.stagnant) {
1012
1026
  console.error(JSON.stringify({
1013
1027
  breaker: "stagnation", scope_id: scope.scope_id, streak: st.streak, no_progress_k: st.k,
package/kernel/gate.mjs CHANGED
@@ -53,7 +53,7 @@
53
53
  import { readFileSync, writeFileSync, appendFileSync, existsSync, mkdirSync } from "node:fs";
54
54
  import { join, dirname } from "node:path";
55
55
  import { runArgs } from "./lib/argv.mjs";
56
- import { gateAnswerCandidates, gates as gatesPath, LOCAL } from "./lib/paths.mjs";
56
+ import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir } from "./lib/paths.mjs";
57
57
 
58
58
  export const GATE_IDS = ["L0", "L1a", "L1a.5", "L1b", "L2", "L3", "QA", "H", "L4", "COACH-1"];
59
59
 
@@ -168,6 +168,71 @@ export function validate(set) {
168
168
  * Returns { gate, decision, source, note, status } where status ∈ ok | ask | abort.
169
169
  * Pure: the CLI turns `status` into the exit code, nothing here exits.
170
170
  */
171
+ /** The three verdicts `scope-hammer` may return (shapeup-run.js's HAMMER schema). */
172
+ export const HAMMER_VERDICTS = ["ship-now", "ship-after-fixes", "cannot-ship"];
173
+
174
+ /**
175
+ * The census verdict on disk for one run, or null when no census has run.
176
+ *
177
+ * Derived from the hammer's own WorkResult rather than taken from a caller: the verdict is the
178
+ * product of a dispatch, and a gate that accepts it as an argument accepts whatever the caller
179
+ * believes. `null` is a real answer here — "scope-hammer has not run" — and is the state a run that
180
+ * returned `gate_h` from a circuit breaker is in, because the breaker returns long before the
181
+ * hammer is dispatched.
182
+ *
183
+ * @param {string} cwd - Project root.
184
+ * @param {string} slug - Feature slug.
185
+ * @returns {(string|null)} The verdict, or null when there is no readable census.
186
+ */
187
+ export function censusVerdict(cwd, slug) {
188
+ if (!slug) return null;
189
+ try {
190
+ const p = join(resultsDir(cwd, slug), "hammer.json");
191
+ if (!existsSync(p)) return null;
192
+ const r = JSON.parse(readFileSync(p, "utf8"));
193
+ const v = r?.payload?.verdict ?? r?.verdict ?? null;
194
+ return HAMMER_VERDICTS.includes(v) ? v : null;
195
+ } catch { return null; }
196
+ }
197
+
198
+ /**
199
+ * Narrow a resolved gate answer to what the run's own evidence supports.
200
+ *
201
+ * THE RULE, and it is about answers rather than callers. An answer set chooses among the answers a
202
+ * gate ALLOWS; it can never supply the evidence that makes one allowed. `ship` at L4 asserts the
203
+ * census cleared the run, so `ship` is only on the menu when a census exists and did clear it.
204
+ *
205
+ * Measured on a consumer: a run whose dispatch receipts carried `orient` and `task-executor` and
206
+ * nothing else — `scope-hammer` never dispatched — recorded `H → accept-cut-list` and `L4 → ship`
207
+ * from the `ci` preset, while its own ledger read `status: escalated`, `final_verdict: ~`, every
208
+ * requirement at `no evidence`. `shapeup-run.js` does guard this, but only after the hammer
209
+ * dispatch, so a run that returns `gate_h` from a breaker reaches L4 by a route the guard does not
210
+ * cover. A check on one call site is not a property.
211
+ *
212
+ * Absence of a census disqualifies `ship` on its own: "no one looked" and "someone looked and it
213
+ * was fine" are different facts, and only the second warrants a ship.
214
+ *
215
+ * @param {object} r - A resolved gate result from {@link resolve}.
216
+ * @param {string} cwd - Project root.
217
+ * @param {(string|null)} slug - Feature slug, when the gate is being resolved inside a run.
218
+ * @returns {object} The result, or an `ask` carrying why `ship` was not available.
219
+ */
220
+ export function narrowToEvidence(r, cwd, slug) {
221
+ if (r?.gate !== "L4" || r?.status !== "ok" || r?.decision !== "ship") return r;
222
+ const verdict = censusVerdict(cwd, slug);
223
+ if (verdict === "ship-now" || verdict === "ship-after-fixes") return r;
224
+ const why = verdict === "cannot-ship"
225
+ ? "scope-hammer's census returned CANNOT SHIP"
226
+ : "scope-hammer has not run, so no census exists";
227
+ return {
228
+ gate: r.gate, status: "ask", source: r.source, decision: "ask", note: r.note,
229
+ refused: "ship", census: verdict,
230
+ reason: `GATE L4 cannot be answered "ship": ${why}. An answer set chooses among the answers a ` +
231
+ `gate allows; it cannot supply the evidence that makes one allowed. Put the block to the ` +
232
+ `PO, or run the census first.`,
233
+ };
234
+ }
235
+
171
236
  export function resolve(set, gate, source) {
172
237
  if (!GATE_IDS.includes(gate)) {
173
238
  return { gate, status: "error", reason: `unknown gate "${gate}" — known: ${GATE_IDS.join(", ")}` };
@@ -362,7 +427,7 @@ export function cli(rawArgv) {
362
427
  const gate = args.resolve ?? null;
363
428
  if (!gate) die("nothing to do — pass --init, --list, --verify, or --resolve <gate-id>");
364
429
 
365
- const r = resolve(found.set, gate, found.source);
430
+ const r = narrowToEvidence(resolve(found.set, gate, found.source), cwd, args.slug ?? null);
366
431
  if (r.status === "error") die(r.reason);
367
432
  // A gate with no `--slug` (e.g. `--file` used ad hoc, outside any run) has nowhere to file a
368
433
  // per-run ledger row — the same reasoning `resolveRunId` uses for "no run is active": absence is
@@ -30,7 +30,7 @@
30
30
  import { existsSync, readdirSync, readFileSync } from "node:fs";
31
31
  import { join, resolve } from "node:path";
32
32
  import { runArgs } from "../lib/argv.mjs";
33
- import { dispatchReceipts, legLedger, resultsDir } from "../lib/paths.mjs";
33
+ import { dispatchReceipts, legLedger, resultsDir, readRunId } from "../lib/paths.mjs";
34
34
  import { readLegs } from "./leg.mjs";
35
35
  import { greenVerdict } from "./t0.mjs";
36
36
 
@@ -64,10 +64,26 @@ export function readReceipts(path) {
64
64
  * @param {object[]} legs - Pre-read `legs.jsonl` rows.
65
65
  * @returns {{orderId:string, hasReceipt:boolean, hasResult:boolean, hasLeg:boolean, state:("unattested"|"in-flight"|"spent")}}
66
66
  */
67
- export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, legs) {
67
+ export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, legs, runId = null) {
68
68
  const orderId = `${slug}/${scopeId}-r${round}-a${attempt}`;
69
- const hasReceipt = receipts.some((r) => r?.order_id === orderId);
70
- const hasLeg = legs.some((r) => r?.order_id === orderId);
69
+ // SCOPED TO THIS RUN, and that qualifier is the whole correction. `receipts/dispatch.jsonl` and
70
+ // `legs.jsonl` are per-slug and append-only, so they accumulate across every run of a pitch,
71
+ // while `order_id` repeats — the run key is the only thing that separates two runs of one
72
+ // feature. Matching on `order_id` alone answered "was this attempt spent?" with a PREVIOUS run's
73
+ // receipt: measured on a consumer, a run that dispatched nothing was told its first attempt was
74
+ // spent, complete with leg and result, by rows two launches old.
75
+ //
76
+ // A row carrying no run key belongs to NO run rather than to this one, and an unresolvable
77
+ // current run (no receipt on disk) matches nothing. Both directions under-count rather than
78
+ // over-count, which is the safe way to be wrong here: an under-count leaves a breaker un-tripped
79
+ // and the round continues, where an over-count stops work that was never done.
80
+ const mine = (r) => r?.order_id === orderId && runId != null && r?.run_id === runId;
81
+ const hasReceipt = receipts.some(mine);
82
+ const hasLeg = legs.some(mine);
83
+ // A WorkResult carries no run key and reaches one only through its `order_id`, which repeats —
84
+ // so this stays a file check and is deliberately NOT sufficient on its own. It can only turn an
85
+ // already run-scoped receipt into `spent`; a result left behind by an earlier run cannot attest
86
+ // an attempt this run never dispatched.
71
87
  const hasResult = existsSync(join(resultsDir(cwd, slug), `${scopeId}-r${round}-a${attempt}.json`));
72
88
  const state = !hasReceipt ? "unattested" : (hasResult || hasLeg) ? "spent" : "in-flight";
73
89
  return { orderId, hasReceipt, hasResult, hasLeg, state };
@@ -91,10 +107,11 @@ export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, le
91
107
  export function scopeAttempts(cwd, slug, scopeId, round, attemptBudget) {
92
108
  const receipts = readReceipts(dispatchReceipts(cwd, slug));
93
109
  const legs = readLegs(legLedger(cwd, slug));
110
+ const runId = readRunId(cwd, slug);
94
111
  const attempts = [];
95
112
  let spent = 0, inFlight = 0, unattested = 0;
96
113
  for (let a = 1; a <= attemptBudget; a++) {
97
- const ev = attemptEvidence(cwd, slug, scopeId, round, a, receipts, legs);
114
+ const ev = attemptEvidence(cwd, slug, scopeId, round, a, receipts, legs, runId);
98
115
  if (ev.state === "spent") spent++;
99
116
  else if (ev.state === "in-flight") inFlight++;
100
117
  else unattested++;
@@ -56,14 +56,14 @@
56
56
  // Exit: 0 ok · 2 malformed argv (nothing ran) · 3 the target the operation needs is not on disk ·
57
57
  // 6 the required phase's artifact is NOT on disk (the phase did not complete).
58
58
 
59
- import { existsSync, readdirSync, readFileSync, writeFileSync, mkdirSync } from "node:fs";
59
+ import { existsSync, readdirSync, readFileSync, writeFileSync, mkdirSync, rmSync } from "node:fs";
60
60
  import { dirname, join, resolve } from "node:path";
61
61
  import { runArgs } from "../lib/argv.mjs";
62
62
  import { splitFrontmatter, uncoerce } from "../lib/contract.mjs";
63
63
  import { globToRegExp } from "../verify/spec.mjs";
64
64
  import {
65
65
  intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir,
66
- orientDir, activeOrder, usecasesDir, breadboard, receipt, readReceipt, requirements,
66
+ orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements,
67
67
  exportRunDir,
68
68
  } from "../lib/paths.mjs";
69
69
  import { evalVerdict } from "./eval.mjs";
@@ -679,10 +679,46 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
679
679
  // The Ship phase (`shapeup-run.js`) already exports a shipped run itself, before this call ever
680
680
  // runs — Stage 2 adds the endings that wrote nothing, and leaves that path untouched.
681
681
  const shouldExport = withExport && status !== "shipped";
682
- const withExportWarning = (result) => {
683
- if (!shouldExport) return result;
684
- const warning = exportOnClose(cwd, slug);
685
- return warning ? { ...result, export_warning: warning } : result;
682
+
683
+ /**
684
+ * Everything a close owes the checkout once the ledger line is written: export the run's
685
+ * records, then retire its pointers.
686
+ *
687
+ * The export runs FIRST, but not because it has to: `exportOnClose` is handed the slug and keys
688
+ * its output by the receipt's `run_id`, so it never reads the pointers this retires. (The bare
689
+ * `reduce export` CLI does read `active-scope` — only to work out which run the operator meant
690
+ * when they named none.) Reporting before teardown is a defensive default, not a correctness
691
+ * requirement, and it is recorded as such so the next reader does not defend an ordering that
692
+ * carries nothing. Measured: swapping the two leaves every check green.
693
+ *
694
+ * WHY RETIRE AT ALL. The substrate fence is enforced while an order is compiled and unanswered,
695
+ * and a run that ends any way other than shipping leaves exactly that by construction. Until this
696
+ * ran, the fence outlived the run: after a close that exited 0 and recorded everything, an
697
+ * ordinary write anywhere in the project was still denied, with no dispatch in flight. The
698
+ * operator's obvious remedy did not help either — `init run --force` is documented as "abandon
699
+ * the open run and start over", never as "release a stuck fence".
700
+ *
701
+ * AND RETIRING IS NOT ANSWERING. The pointer says "a run is in flight"; the order's missing
702
+ * result says "nobody came back". Only the first is untrue after a close. The abandoned order
703
+ * stays unanswered, the attempt census still sees nothing spent on it, and `init run --force`
704
+ * remains the one thing that writes a synthetic result — because a close that quietly claimed the
705
+ * work was answered would spend an attempt budget on work nobody did.
706
+ *
707
+ * Best-effort, like the export: a pointer that cannot be removed degrades the close and is
708
+ * reported in its return, but never turns a close into a non-close.
709
+ */
710
+ const finishClose = (result) => {
711
+ const warning = shouldExport ? exportOnClose(cwd, slug) : null;
712
+ const stuck = [];
713
+ for (const pointer of [activeOrder(cwd), activeScope(cwd)]) {
714
+ try { rmSync(pointer, { force: true }); } catch { /* fall through to the check below */ }
715
+ if (existsSync(pointer)) stuck.push(pointer);
716
+ }
717
+ return {
718
+ ...result,
719
+ ...(warning ? { export_warning: warning } : {}),
720
+ ...(stuck.length ? { pointer_warning: `could not retire ${stuck.join(", ")} — the substrate fence may still deny writes until it is removed by hand` } : {}),
721
+ };
686
722
  };
687
723
 
688
724
  // Truncated, not elided: a cause this long has already done its job in the run's own log — the
@@ -699,7 +735,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
699
735
  if (priorClosedStatus && priorClosedAt) {
700
736
  if (priorClosedStatus === status && normCause === priorCause) {
701
737
  // The identical fact, restated — a retried or duplicated call costs nothing.
702
- return withExportWarning({ ok: true, path: p, status, closed_at: priorClosedAt, cause: priorCause, decision: "idempotent", reason: `already closed as "${status}" at ${priorClosedAt} — idempotent no-op` });
738
+ return finishClose({ ok: true, path: p, status, closed_at: priorClosedAt, cause: priorCause, decision: "idempotent", reason: `already closed as "${status}" at ${priorClosedAt} — idempotent no-op` });
703
739
  }
704
740
  if (priorClosedStatus !== status) {
705
741
  // A DIFFERENT terminal status over an already-closed run — refused outright, the original
@@ -724,7 +760,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
724
760
  if (afterSup.status !== status || !afterSup.closed_at || afterSup.closed_at === "~") {
725
761
  return { ok: false, path: p, status, reason: `wrote the superseding close but the ledger reads back status="${afterSup.status}" closed_at="${afterSup.closed_at}" — the write did not take` };
726
762
  }
727
- return withExportWarning({
763
+ return finishClose({
728
764
  ok: true, path: p, status, closed_at: afterSup.closed_at, cause: afterSup.close_cause ?? null,
729
765
  superseded: true, decision: "superseded", prior_cause: priorCause, prior_closed_at: priorClosedAt,
730
766
  });
@@ -739,7 +775,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
739
775
  if (after.status !== status || !after.closed_at || after.closed_at === "~") {
740
776
  return { ok: false, path: p, status, reason: `wrote the close but the ledger reads back status="${after.status}" closed_at="${after.closed_at}" — the write did not take` };
741
777
  }
742
- return withExportWarning({ ok: true, path: p, status, closed_at: after.closed_at, cause: after.close_cause ?? null, decision: "closed" });
778
+ return finishClose({ ok: true, path: p, status, closed_at: after.closed_at, cause: after.close_cause ?? null, decision: "closed" });
743
779
  }
744
780
 
745
781
  /**
@@ -78,7 +78,7 @@ export function frontmatter(text) {
78
78
  */
79
79
  export function boardCensus(cwd, slug) {
80
80
  const dir = tasksDir(cwd, slug);
81
- const out = { total: 0, done: 0, unfinished: [] };
81
+ const out = { total: 0, done: 0, unfinished: [], anchors: {} };
82
82
  if (!existsSync(dir)) return out;
83
83
  for (const f of readdirSync(dir)) {
84
84
  if (!/^TASK-[\w.-]+\.md$/i.test(f)) continue;
@@ -86,6 +86,12 @@ export function boardCensus(cwd, slug) {
86
86
  const fm = frontmatter(body);
87
87
  const id = fm.id || f.replace(/\.md$/, "");
88
88
  out.total++;
89
+ // The COMMITTED anchor for this board id. `use_case_refs` is the tier-direction rule's own
90
+ // sanctioned direction (LOCAL names SHARED), and it is what the frozen report cites instead of
91
+ // the id — boards renumber per machine, use cases do not.
92
+ const ucs = String(fm.use_case_refs ?? "").replace(/^\[|\]$/g, "")
93
+ .split(",").map((x) => x.trim()).filter(Boolean);
94
+ out.anchors[id] = ucs;
89
95
  if (fm.status === "done") out.done++;
90
96
  else out.unfinished.push(id);
91
97
  }
@@ -93,6 +99,31 @@ export function boardCensus(cwd, slug) {
93
99
  return out;
94
100
  }
95
101
 
102
+ /**
103
+ * Replace every board id in free prose with its committed anchor.
104
+ *
105
+ * THE WRITE BOUNDARY, not the column, and that is the whole point. Board ids reached the frozen
106
+ * report three different ways — the unfinished-task callout, the covering-AC column's own prefix,
107
+ * and INSIDE acceptance-criterion prose a planner wrote ("given the seeded todos (TASK-006)"). The
108
+ * third is upstream free text, so a fix that only changes what the columns interpolate still
109
+ * commits a file the next run's spec-lint reds. Everything written into the committed report passes
110
+ * through here.
111
+ *
112
+ * An id with no resolvable use case becomes a neutral phrase rather than the id: the report loses a
113
+ * pointer that never resolved off this machine anyway, and keeps the sentence around it.
114
+ *
115
+ * @param {*} text - Any value destined for the committed report.
116
+ * @param {Record<string, string[]>} anchors - Board id → its `use_case_refs`.
117
+ * @returns {string} The text with every `TASK-…` replaced by a stable anchor.
118
+ */
119
+ export function deboard(text, anchors = {}) {
120
+ return String(text ?? "").replace(/\bTASK-[A-Za-z0-9][\w.-]*/g, (id) => {
121
+ const ucs = anchors[id];
122
+ if (ucs && ucs.length) return ucs.join("/");
123
+ return "a board task";
124
+ });
125
+ }
126
+
96
127
  /**
97
128
  * Per-scope T0 outcome, reduced from the trial ledger.
98
129
  *
@@ -194,7 +225,10 @@ export function buildReport(facts) {
194
225
  L.push("");
195
226
 
196
227
  if (board.unfinished.length) {
197
- L.push(`> **${board.unfinished.length} task(s) did not finish:** ${board.unfinished.join(", ")}.`,
228
+ // Anchored, never enumerated by board id: the ids renumber per machine, and a committed file
229
+ // carrying one reds the NEXT run of this pitch at L1b.
230
+ const unfinishedAnchors = [...new Set(board.unfinished.flatMap((id) => board.anchors?.[id] ?? []))];
231
+ L.push(`> **${board.unfinished.length} task(s) did not finish**${unfinishedAnchors.length ? ` — use cases: ${unfinishedAnchors.join(", ")}` : ""}.`,
198
232
  "> The verdict above grades what was built, not what was planned.", "");
199
233
  }
200
234
 
@@ -241,7 +275,9 @@ export function buildReport(facts) {
241
275
  const cell = (s) => String(s).replace(/\|/g, "\\|");
242
276
  L.push("| REQ | source | evidence | covering AC | criterion | T0 |", "|---|---|---|---|---|---|");
243
277
  for (const r of requirements.rows) {
244
- const ac = r.covering_acs.length ? `${r.covering_acs[0].task_id}: ${r.covering_acs[0].ac}${r.covering_acs.length > 1 ? ` (+${r.covering_acs.length - 1})` : ""}` : "—";
278
+ const ac = r.covering_acs.length
279
+ ? `${deboard(r.covering_acs[0].task_id, facts.board?.anchors)}: ${deboard(r.covering_acs[0].ac, facts.board?.anchors)}${r.covering_acs.length > 1 ? ` (+${r.covering_acs.length - 1})` : ""}`
280
+ : "—";
245
281
  const crit = r.criteria.length ? `${r.criteria[0].criterion}${r.criteria.length > 1 ? ` (+${r.criteria.length - 1})` : ""} → ${r.criteria.map((c) => c.verdict).join(",")}` : "—";
246
282
  const t0h = r.t0.length ? r.t0.map((h) => String(h).slice(0, 12)).join(", ") : "—";
247
283
  L.push(`| ${r.id} | ${cell(r.source || "—")} | ${r.evidence} | ${cell(ac)} | ${cell(crit)} | ${t0h} |`);
@@ -1215,7 +1215,7 @@
1215
1215
  }
1216
1216
  },
1217
1217
  "CommandResult": {
1218
- "description": "One executed command's outcome inside a T0 artifact — produced by actually running the command; no agent can fabricate it.",
1218
+ "description": "One executed command's outcome inside a T0 artifact — produced by actually running the command; no agent can fabricate it. It carries the EVIDENCE, not only the score: `exit` maps a command that never started onto the same 1 a real failure returns, so without `error` a refused or timed-out command and a broken build are the same record everywhere downstream, and without the output a verdict asserting `exit 1` is an assertion nobody can check.",
1219
1219
  "x-tier": "EMBEDDED",
1220
1220
  "type": "object",
1221
1221
  "properties": {
@@ -1223,10 +1223,23 @@
1223
1223
  "type": "string"
1224
1224
  },
1225
1225
  "exit": {
1226
- "type": "integer"
1226
+ "type": "integer",
1227
+ "description": "The process exit code, or 1 when the command never produced one. Not a discriminator on its own — read `error` to tell a crash from a failure."
1227
1228
  },
1228
1229
  "pass": {
1229
1230
  "type": "boolean"
1231
+ },
1232
+ "error": {
1233
+ "type": "string",
1234
+ "description": "Present ONLY when the command did not run to completion — a spawn failure, a maxBuffer overflow, or the 10-minute timeout. Its presence is the fact the ratchet grades as `crash` (tree restored, never counted as a reverted attempt); its absence means the command ran and the exit code is its own."
1235
+ },
1236
+ "stdout_tail": {
1237
+ "type": "string",
1238
+ "description": "The last 4000 characters of stdout, omitted when nothing was printed. Kept for passing commands too: a fixture that exits 0 having run zero tests is the false green this layer exists to catch. A truncated tail carries a leading `[truncated: kept the last N of M characters]` line, so a partial stream never reads as a complete one."
1239
+ },
1240
+ "stderr_tail": {
1241
+ "type": "string",
1242
+ "description": "The last 4000 characters of stderr, same bound and same truncation marker as `stdout_tail`."
1230
1243
  }
1231
1244
  }
1232
1245
  },
@@ -235,11 +235,56 @@ export function lintScopes(scopes, repoFiles) {
235
235
  }
236
236
 
237
237
  /** Text forms worth scanning; anything else in a committed tree is not a reference carrier. */
238
- const SCANNED = /\.(md|markdown|yml|yaml|json|txt)$/i;
238
+ export const SCANNED = /\.(md|markdown|yml|yaml|json|txt)$/i;
239
239
 
240
240
  /** A machine-local board id. Strict on purpose: a committed tree has no reason to carry one at all. */
241
241
  const TASK_ID = /\bTASK-[A-Za-z0-9][\w.-]*/;
242
242
 
243
+ /**
244
+ * The tier-direction violations one piece of text carries — the whole rule, on a string.
245
+ *
246
+ * EXPORTED BECAUSE IT HAS TWO ENFORCEMENT POINTS NOW, and they must not be allowed to drift.
247
+ * `lintCommittedTier` below walks files at GATE L1b; `hooks/tier-guard.mjs` refuses the same text
248
+ * at the moment a tool writes it, minutes earlier, while the writer still holds the context needed
249
+ * to rephrase. Four producers wrote committed files this rule reds and none of them learned from
250
+ * the lint, because by the time it speaks the dispatch that wrote the line is over. A second
251
+ * enforcement point is only worth having if it enforces the SAME predicate, so both call this.
252
+ *
253
+ * The detail strings are the message the writer reads, so they carry the remedy, not just the
254
+ * verdict: a board id resolves on the machine that wrote it and nowhere else, and a path into the
255
+ * gitignored tier dangles on every clone.
256
+ *
257
+ * @param {string} text - The file body, or the fragment a tool is about to write.
258
+ * @returns {Array<{kind:("board-id"|"local-path"), token:string, line:number, detail:string}>}
259
+ * One entry per offending line and form, in file order; [] when clean.
260
+ */
261
+ export function tierLeaks(text) {
262
+ // Built from the LOCAL constant, never a literal — the storage roots have exactly one home.
263
+ const esc = LOCAL.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
264
+ // `\S+` where the walk's own rule is `\S`: the same LINES match either way (a path with one
265
+ // non-space character after the slash has at least one), and the longer form yields the token to
266
+ // quote back at the writer. Trailing punctuation the prose wrapped it in is trimmed off the
267
+ // quote only — never off the test.
268
+ const localPath = new RegExp(`${esc}/\\S+`);
269
+ const leaks = [];
270
+ String(text ?? "").split(/\r?\n/).forEach((line, i) => {
271
+ const task = line.match(TASK_ID);
272
+ if (task) {
273
+ leaks.push({ kind: "board-id", token: task[0], line: i + 1, detail:
274
+ `names ${task[0]} — a committed file cannot carry a board id. Boards live in ${LOCAL}/ ` +
275
+ "(gitignored) and renumber on every regeneration, so this resolves on the machine that wrote it " +
276
+ "and nowhere else. Cite the use case or the scope_id, which are stable." });
277
+ }
278
+ const path = line.match(localPath);
279
+ if (path) {
280
+ leaks.push({ kind: "local-path", token: path[0].replace(/[`)\]},.;:'"]+$/, ""), line: i + 1, detail:
281
+ `points into ${LOCAL}/ — a committed file cannot reference the gitignored tier; the path ` +
282
+ "dangles on every other clone. Name the committed artifact, or describe the tier without a path." });
283
+ }
284
+ });
285
+ return leaks;
286
+ }
287
+
243
288
  /**
244
289
  * Lint the WHOLE committed tree for references into the gitignored tier.
245
290
  *
@@ -265,28 +310,15 @@ const TASK_ID = /\bTASK-[A-Za-z0-9][\w.-]*/;
265
310
  export function lintCommittedTier({ cwd, slug }) {
266
311
  const root = sharedRoot(cwd, slug);
267
312
  if (!existsSync(root)) return [];
268
- // Built from the LOCAL constant, never a literal — the storage roots have exactly one home.
269
- const localPath = new RegExp(`${LOCAL.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")}/\\S`);
270
313
  const findings = [];
271
314
  for (const rel of walkFiles(root)) {
272
315
  if (!SCANNED.test(rel)) continue;
273
- let lines;
274
- try { lines = readFileSync(join(root, rel), "utf8").split(/\r?\n/); } catch { continue; }
275
- lines.forEach((line, i) => {
276
- const at = `${relative(cwd, join(root, rel))}:${i + 1}`;
277
- const task = line.match(TASK_ID);
278
- if (task) {
279
- findings.push({ rule: "TIER-DIRECTION", level: "red", detail:
280
- `${at} names ${task[0]} — a committed file cannot carry a board id. Boards live in ${LOCAL}/ ` +
281
- "(gitignored) and renumber on every regeneration, so this resolves on the machine that wrote it " +
282
- "and nowhere else. Cite the use case or the scope_id, which are stable." });
283
- }
284
- if (localPath.test(line)) {
285
- findings.push({ rule: "TIER-DIRECTION", level: "red", detail:
286
- `${at} points into ${LOCAL}/ — a committed file cannot reference the gitignored tier; the path ` +
287
- "dangles on every other clone. Name the committed artifact, or describe the tier without a path." });
288
- }
289
- });
316
+ let text;
317
+ try { text = readFileSync(join(root, rel), "utf8"); } catch { continue; }
318
+ for (const leak of tierLeaks(text)) {
319
+ findings.push({ rule: "TIER-DIRECTION", level: "red",
320
+ detail: `${relative(cwd, join(root, rel))}:${leak.line} ${leak.detail}` });
321
+ }
290
322
  }
291
323
  return findings;
292
324
  }
@@ -79,6 +79,69 @@ function runCommand(cmd, cwd) {
79
79
  return { cmd, exit: r.status ?? 1, pass: r.status === 0, stdout, stderr, ...(error ? { error } : {}) };
80
80
  }
81
81
 
82
+ /**
83
+ * How much of each stream the persisted record keeps, per command.
84
+ *
85
+ * A bound rather than the whole stream, because one chatty fixture would otherwise make every
86
+ * reader of the run trace pay for it — and a bound at the END rather than the start, because that
87
+ * is where a stack trace, an assertion diff and a test summary all land.
88
+ */
89
+ export const EVIDENCE_TAIL_CHARS = 4000;
90
+
91
+ /**
92
+ * The last `limit` characters of a stream, marked when anything was dropped.
93
+ *
94
+ * THE MARKER IS NOT DECORATION. A truncated tail that does not say so is a partial stream that
95
+ * reads as a complete one, which is the same class of defect as the one this whole change fixes:
96
+ * a record that overstates what it actually holds.
97
+ *
98
+ * @param {(string|null|undefined)} text - The captured stream.
99
+ * @param {number} [limit] - Characters to keep.
100
+ * @returns {(string|null)} The kept tail, or null when the stream was empty (the field is then
101
+ * omitted rather than stored as "", so "nothing was printed" stays visibly different from
102
+ * "nothing was kept").
103
+ */
104
+ export function boundedTail(text, limit = EVIDENCE_TAIL_CHARS) {
105
+ const s = String(text ?? "");
106
+ if (!s) return null;
107
+ if (s.length <= limit) return s;
108
+ return `[truncated: kept the last ${limit} of ${s.length} characters]\n${s.slice(-limit)}`;
109
+ }
110
+
111
+ /**
112
+ * One command's outcome as the verdict artifact stores it — the evidence, not just the score.
113
+ *
114
+ * WHAT THIS FIXES. The artifact used to keep `{cmd, exit, pass}`, and `runCommand` maps a command
115
+ * that never started onto `exit: 1` — the same number a genuine failure returns. So a refused or
116
+ * timed-out command and a broken build were the SAME RECORD everywhere downstream: the digest, the
117
+ * hill, the report, the evaluator's citation and anyone reading the trace back afterwards. The
118
+ * kernel computed the difference (`error`, which the ratchet grades as `crash`) and discarded it
119
+ * one line later.
120
+ *
121
+ * `exit` is deliberately left as it is. Mapping a crash to some other number would change what the
122
+ * ratchet compares and what every existing reader parses; the distinction travels in `error`, which
123
+ * is the field that actually means "this never ran", and which the crash branch already reads.
124
+ *
125
+ * Output is kept for PASSING commands too, and that direction is not an afterthought: a fixture
126
+ * that exits 0 having run zero tests is the false green this evidence layer exists to catch, and
127
+ * its stdout is the only place that shows.
128
+ *
129
+ * @param {({cmd:string, exit:number, pass:boolean, stdout?:string, stderr?:string, error?:string}|null)} r
130
+ * A `runCommand` result, or null when no command was declared.
131
+ * @returns {(object|null)} The record to persist; null passes through unchanged.
132
+ */
133
+ export function commandEvidence(r) {
134
+ if (!r) return null;
135
+ const stdout = boundedTail(r.stdout);
136
+ const stderr = boundedTail(r.stderr);
137
+ return {
138
+ cmd: r.cmd, exit: r.exit, pass: r.pass,
139
+ ...(r.error ? { error: r.error } : {}),
140
+ ...(stdout ? { stdout_tail: stdout } : {}),
141
+ ...(stderr ? { stderr_tail: stderr } : {}),
142
+ };
143
+ }
144
+
82
145
  /**
83
146
  * Run every e2e fixture command for a scope.
84
147
  * @param {string[]} fixtures - Fixture command lines (null/empty → no commands).
@@ -513,8 +576,10 @@ export async function cli(rawArgv) {
513
576
  const { path, sha256: hash, trial } = writeArtifact(outDir, round, attempt, {
514
577
  ...(runId ? { run_id: runId } : {}),
515
578
  scope_id: contract.scope_id,
516
- fixtures: fixtures.results.map(({ cmd, exit, pass }) => ({ cmd, exit, pass })),
517
- db_probe: dbProbe && { cmd: dbProbe.cmd, exit: dbProbe.exit, pass: dbProbe.pass },
579
+ // The evidence, not just the score — see `commandEvidence` for what the three-field record
580
+ // could not tell apart, and why `exit` still reads the way it always did.
581
+ fixtures: fixtures.results.map((r) => commandEvidence(r)),
582
+ db_probe: commandEvidence(dbProbe),
518
583
  seesaw,
519
584
  ...verdict,
520
585
  score: s,
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.7.0",
3
+ "version": "3.7.1",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -281,6 +281,10 @@ Always use wikilinks (double brackets), never relative paths like `../domain-mod
281
281
  (`"extracted from .shapeup/<slug>/intake.md"`) reds for the same reason a task
282
282
  wikilink does: the path dangles on every other clone. Cite the committed pitch or
283
283
  shaping doc instead, or describe the run tier without a path.
284
+ - **You will be stopped at the write, not at the gate.** The same rule is enforced as a
285
+ refusal on the Write/Edit itself, quoting the offending token back. There is nothing to
286
+ appeal: rephrase the line and write it again. Waiting for spec-lint to tell you at Board
287
+ Review costs the whole phase, which is what it used to cost every time.
284
288
  - `[[tasks/...]]` wikilinks are valid only inside LOCAL documents (task files, the board,
285
289
  EVAL reports), where they resolve against the LOCAL root (`.shapeup/<slug>/`);
286
290
  every wikilink in a SHARED doc stays `spec_folder`-relative.
@@ -26,7 +26,8 @@ scales down, its *verification floor* does not.
26
26
  ingest-result dispatch. Tiny never means "just edit the file inline".
27
27
  - **T0 verification.** A tiny change still proves itself by running — never by claim. If there
28
28
  is no runnable check at all, that is a fit-check failure, not a reason to skip T0.
29
- - **The safety-spine and sandbox hooks.** Machine guards do not scale down.
29
+ - **The machine guards — safety spine, substrate sandbox, tier guard.** They do not scale down,
30
+ and a tiny lane is where a committed file is most likely to be written by hand.
30
31
  - **The discovery ledger.** `lane: tiny` is recorded, so a later reader knows exactly what was
31
32
  NOT checked (no EVAL verdict, no QA charter, no wiring assertion).
32
33