shapeup-sdlc 3.2.0 → 3.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.2.0",
4
+ "version": "3.4.0",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -31,8 +31,8 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
31
31
  | Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases) |
32
32
  | Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
33
33
  | Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚ |
34
- | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe + seesaw ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt |
35
- | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order): spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
34
+ | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe + seesaw ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
35
+ | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
36
36
  | FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
37
37
 
38
38
  ✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
@@ -44,9 +44,9 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
44
44
  **Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`.
45
45
 
46
46
  ### Ship & Triage
47
- - **SHIP S.0 / GATE H** — `/scope-hammer`: census (QA findings + discovered ledger + attempt-budget proposals ✦) → baseline comparison (never the ideal) → cut list; TL/PO promotes selected items only.
47
+ - **SHIP S.0 / GATE H** — `/scope-hammer`: census (QA findings + discovered ledger + attempt-budget proposals ✦; every "no scope owns X" cites `probe owner`, which derives ownership from the contracts) → baseline comparison (never the ideal) → cut list; TL/PO promotes selected items only.
48
48
  - ⏸ **L4** — Ship Sign-off (shows QA status ★).
49
- - **Coach retro** — L4 feedback → `/coach`; GATE COACH-1 asks the PO which skill owns each rule (never assumes) → committed `shapeup/knowledge-base/<skill>.md` (team inherits on pull). Coachable: `/task-executor`, `/ba-pitch-analyzer`, `/qa-edge-hunter`; `/spec-evaluator` is not (single judge). Mechanism defects file to `knowledge-base/harness-defects.md` as Betting Table raw ideas, never worker steering.
49
+ - **Coach retro** — L4 feedback → `/coach`; GATE COACH-1 asks the PO which skill owns each rule (never assumes) → committed `shapeup/knowledge-base/<skill>.md` (team inherits on pull). Coachable: `/task-executor`, `/ba-pitch-analyzer`, `/qa-edge-hunter`, `/orient`, `/scope-architect`, `/solution-architect`, and `/tech-lead` (workflow guidance read at L0, plus suggested L0 values it confirms before pinning); `/spec-evaluator` is not (single judge) and neither is `/scope-hammer` (its census cites `probe owner`). **Guidance never decides a gate**: a rule may add a question or a check to a gate block, never an answer, a skip or a wider substrate. `/retro --scan` seeds the same files from the project on disk before the first run, and `/retro --research <stack>` from the platform's official documentation when there is nothing on disk yet (a source, never a verification: nothing it reads runs until the tech lead pins it at L0), optional, every rule confirmed at COACH-1. Mechanism defects file to `knowledge-base/harness-defects.md` as Betting Table raw ideas, never worker steering.
50
50
  - Post-fix: `eval --single-pass` → remaining `~` + new feedback → new raw idea.
51
51
 
52
52
  ### Discovered Tasks
@@ -58,7 +58,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
58
58
  - **Ledger = single source of truth** — every discovery flow writes only its own section.
59
59
  - **QA is a level-up, not a gate** — `--no-qa` skips it; circuit breaker outranks the Hunter.
60
60
  - **Role separation** — Evaluator grades, task-executor fixes, QA discovers.
61
- - **Hill phase is mechanical ✦** — derived only from T0/T1/seesaw artifacts, never self-reported; the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
61
+ - **Hill phase is mechanical ✦** — derived only from T0/T1/seesaw artifacts, never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
62
62
  - **Envelope port (v1.0)** — every dispatch is WorkOrder in / WorkResult out; shared state has exactly one writer (the ingest step); malformed envelopes are hook-denied. Workers: stateless, craft-only, pipeline-blind.
63
63
 
64
64
  ## Setup & Execution
@@ -79,5 +79,6 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
79
79
  - `/hill-chart` (skill `hill-chart`, not a pipeline worker — invoked directly, like `shapeup`) renders both the committed hill shards (`shapeup/<slug>/hill/<scope-id>.yml`, the mechanical phase from the invariant above) and the local run graph as one dashboard: a portfolio card per pitch, and per-pitch a Hill Chart, an attention list, a scope board, round history, and the run graph one click deeper. A pitch whose local run trace was cleaned up after shipping still renders — marked Archived — from its committed hill shards alone.
80
80
  - Contracts: markdown on disk, JSON on the wire; a single library reads/writes the file form.
81
81
  - Never hard-code a storage root — generated paths resolve through the shared path resolver.
82
+ - Hooks file under the project root they find above the shell's working directory (a run pointer, the committed tier, or a git boundary), never under the folder a worker happened to `cd` into — so a sub-folder shell neither splits the decision ledger nor slips the substrate fence. `probe stats --hooks` lists any stray ledger it still finds.
82
83
  - The traceability oracle emits `.shapeup/<slug>/trace/report.json` from the spine artifacts.
83
84
  <!-- HARNESS_END -->
package/README.md CHANGED
@@ -177,7 +177,7 @@ every arm is skipped when its artifact is absent, so older specs are unaffected.
177
177
  | Evaluate (GATE L3) | `spec-evaluator` | v1.0 | The single judge (pure worker). Verifies spec-conformance, TDD surface, and integration against the running app — skeptical, files `file:line` bugs, runs exactly once per build round. Requires a T0 artifact citation, grades UI affordance-only; verdict + refuted boxes return as data. |
178
178
  | QA (post-PASS) | `qa-edge-hunter` | v1.1 | Exploratory edge hunt on the running app through six fixed lenses, charting edges *outside* what the evaluator probed. Findings go to the ledger as `~`; never blocks ship. |
179
179
  | Stop (11) | `scope-hammer` | v0.1 | GATE H: must-have census → baseline comparison (never vs. the ideal) → cut list + ship verdict. Handles the normal stop and both circuit-breaker triggers. |
180
- | Retro (post-L4) | `coach` | — | RLHF for the harness: turns raw PO/TL feedback at Ship Sign-off into per-skill guidelines under committed `shapeup/knowledge-base/<skill>.md`, read back by `task-executor` / `ba-pitch-analyzer` / `qa-edge-hunter` on their next run. GATE COACH-1 asks the PO which skill owns each rule — never assumes; mechanism defects are filed to the harness-defect register instead. |
180
+ | Retro (post-L4) | `coach` | — | RLHF for the harness: turns raw PO/TL feedback at Ship Sign-off into per-skill guidelines under committed `shapeup/knowledge-base/<skill>.md`, read back by six coachable workers on their next run and by `tech-lead` at GATE L0 (workflow guidance, never a gate answer). `--scan` seeds the same files from the project on disk before the first run; `--research <stack>` seeds them from the platform's official documentation when the project has nothing to scan, and cross-checks a scan's rules when it has. GATE COACH-1 asks the PO which skill owns each rule — never assumes; mechanism defects are filed to the harness-defect register instead. |
181
181
  | Orchestrator | `tech-lead` | v1.0 | Owns the run end-to-end: PLAN once → BUILD all tasks → EVAL once per round, looping on FAIL. Three-level circuit breaker (rounds / T0 attempts / wall clock), T0/seesaw-verified build rounds, mechanical hill derivation. Sole writer of run-state. |
182
182
 
183
183
  ### Commands
package/SECURITY.md CHANGED
@@ -45,7 +45,10 @@ test against machines you don't own.
45
45
  (override channel fails closed), and every exercised override is logged. The same principle
46
46
  covers `.shapeup/active-order`, which `sandbox-guard` reads to find the run whose orders fence
47
47
  a worker's writes: it sits outside the run-trace carve-out, so a worker cannot repoint its own
48
- sandbox.
48
+ sandbox. Every hook resolves that pointer, the decision ledger and the substrate globs against
49
+ the project root it finds above the tool call's working directory (a run pointer, the committed
50
+ tier, or a git boundary) — a worker that `cd`s into a sub-folder is fenced exactly as one at the
51
+ top, and its receipts land in the same ledger.
49
52
  5. **Exactly one hook can block, and only on a mechanical absence.** `gate-zerowork` returns
50
53
  `decision: "block"` in one state: the session dispatched the orchestrator and left no run
51
54
  receipt on disk. It makes no judgement about quality — it reports that there is no work to
package/bin/init.mjs CHANGED
@@ -181,6 +181,9 @@ for (const [srcRel, note] of [
181
181
 
182
182
  console.log("\n✅ Harness installation and scaffolding completed.");
183
183
  console.log(" Next: open a Claude Code session in this directory and run /ship \"<your idea>\".");
184
+ console.log(" Optional: /retro --scan first seeds the team knowledge base from this project's");
185
+ console.log(" build files, or /retro --research \"<stack>\" seeds it from the platform's official");
186
+ console.log(" docs when there is nothing on disk yet (every drafted rule is confirmed by you).");
184
187
  console.log(" ([ui] evaluation needs a browser — `npx playwright install chromium` — but only");
185
188
  console.log(" when a run actually reaches a [ui] criterion; nothing else requires it.)");
186
189
 
package/commands/retro.md CHANGED
@@ -5,9 +5,26 @@ Use the **coach** skill on $ARGUMENTS.
5
5
 
6
6
  Turns raw PO/TL feedback (usually from the L4 Ship gate) into per-skill guideline files under
7
7
  committed `shapeup/knowledge-base/<skill>.md`, which the coachable skills read back at
8
- the top of their next run.
8
+ the top of their next run, and which the tech lead reads at GATE L0 for workflow guidance.
9
+
10
+ `/retro --scan` runs the coach's **scan** operation instead: it reads the project on disk (build
11
+ and toolchain files, CI config, CLAUDE.md, README) and drafts the same kind of guidelines from
12
+ it, before the first feature or after the toolchain changes. Optional, and never on the scan's
13
+ own authority — every drafted rule goes through GATE COACH-1 like feedback.
14
+
15
+ `/retro --research <stack>` runs the **research** operation: for a project with nothing on disk
16
+ yet, it drafts the same kind of guidelines from the platform's official documentation only —
17
+ build, launch, test, package manager, lint, in that order of leverage — each rule cited with url,
18
+ version and fetch date; on a project that has been scanned, it cross-checks every scan rule
19
+ against the documentation and reports each as confirmed, contradicted or unknown. The stack is
20
+ required and never guessed. Research is a source, not a verification: nothing it reads runs, and
21
+ the tech lead still pins every suggested command at GATE L0 before the kernel executes it. A
22
+ later `--scan`, once the first feature has created a toolchain, retires the research rules it
23
+ confirms with disk evidence.
9
24
 
10
25
  Two rules the skill enforces and this command must not soften: GATE COACH-1 **asks** the PO
11
26
  which skill owns each rule — it never assumes; and feedback whose root cause is the mechanism
12
27
  itself (a gate, hook, or contract defect) is categorized `harness-defect` and filed to the
13
- defect register as a raw idea for the Betting Table, never as worker steering.
28
+ defect register as a raw idea for the Betting Table, never as worker steering. A third holds
29
+ for every category: guidance never decides a gate — a rule may add a question or a check to a
30
+ gate block, never an answer.
@@ -49,7 +49,7 @@ import { appendFileSync, mkdirSync, readFileSync } from "node:fs";
49
49
  import { resolve, dirname } from "node:path";
50
50
  import { isMain } from "../kernel/lib/argv.mjs";
51
51
  import { dispatchReceipts } from "../kernel/lib/paths.mjs";
52
- import { runHook, readStdin, settle } from "./lib/decision.mjs";
52
+ import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
53
53
 
54
54
  /**
55
55
  * The `--order` matcher, character-for-character the one the PreToolUse order gate uses
@@ -173,7 +173,10 @@ export async function main() {
173
173
  defer("dispatch result names no resolved skill — nothing to attest", "no-skill-named");
174
174
  }
175
175
 
176
+ // The order path is relative to the shell; the receipt file is relative to the project root
177
+ // (see `projectRoot`) — the same ledger whichever folder the dispatch was issued from.
176
178
  const cwd = p.cwd || process.cwd();
179
+ const root = projectRoot(cwd);
177
180
  const orderPath = resolve(cwd, cited);
178
181
  let order;
179
182
  try { order = JSON.parse(readFileSync(orderPath, "utf8")); }
@@ -181,9 +184,9 @@ export async function main() {
181
184
  if (!order?.order_id) defer("order carries no order_id — nothing to key a receipt by", "order-unkeyed");
182
185
 
183
186
  const row = receiptRow(order, skillInvoked, p);
184
- const written = writeReceipt(row, cwd);
187
+ const written = writeReceipt(row, root);
185
188
  return {
186
- verdict: "allow", event: "PostToolUse", tool: p.tool_name, cwd, subject: row.order_id,
189
+ verdict: "allow", event: "PostToolUse", tool: p.tool_name, cwd: root, subject: row.order_id,
187
190
  rule: written ? "receipt-written" : "receipt-write-failed",
188
191
  reason: written
189
192
  ? `dispatch receipt: ${row.order_id} ran ${skillInvoked} (declared ${row.worker_declared}, ok=${row.dispatch_ok})`
@@ -57,7 +57,7 @@ import { readFileSync, readdirSync, existsSync, statSync } from "node:fs";
57
57
  import { join } from "node:path";
58
58
  import { isMain } from "../kernel/lib/argv.mjs";
59
59
  import { localDir, globLocal, globWorkflowsStage } from "../kernel/lib/paths.mjs";
60
- import { runHook, readStdin, settle, decisionsPath } from "./lib/decision.mjs";
60
+ import { runHook, readStdin, settle, decisionsPath, projectRoot } from "./lib/decision.mjs";
61
61
 
62
62
  const MAX_TRANSCRIPT_BYTES = 20 * 1024 * 1024;
63
63
 
@@ -287,7 +287,10 @@ async function main() {
287
287
  // (read-only cwd, missing node) would be held open forever.
288
288
  if (p.stop_hook_active) defer("stop_hook_active — at most one block per stop chain", "loop-guard");
289
289
 
290
- const cwd = p.cwd || process.cwd();
290
+ // Receipts and the decision ledger sit at the project root, wherever the session's shell ended
291
+ // up (see `projectRoot`); a run started from the top must not read as "never started" because
292
+ // the last command `cd`ed somewhere.
293
+ const cwd = projectRoot(p.cwd || process.cwd());
291
294
  const events = readEvents(p.transcript_path);
292
295
  // no transcript → no facts → fail open
293
296
  if (!events || events.length === 0) defer("no readable transcript — no facts to assert", "no-transcript");
@@ -46,11 +46,52 @@
46
46
  // failed tool call: a receipt that can break a run would get the whole layer disabled, which is
47
47
  // the exact outcome this file exists to prevent. Every write here is inside a try/catch.
48
48
 
49
- import { appendFileSync, mkdirSync } from "node:fs";
50
- import { dirname } from "node:path";
51
- import { decisions } from "../../kernel/lib/paths.mjs";
49
+ import { appendFileSync, mkdirSync, existsSync, statSync } from "node:fs";
50
+ import { dirname, resolve } from "node:path";
51
+ import { decisions, activeScope, sharedDir } from "../../kernel/lib/paths.mjs";
52
52
  import { resolveRunId } from "../../kernel/lib/paths.mjs";
53
53
 
54
+ /**
55
+ * The project root a hook should file under, from wherever the tool call happened to fire.
56
+ *
57
+ * THE HOOK PAYLOAD'S `cwd` FOLLOWS THE SHELL, NOT THE PROJECT. A worker that `cd`s into a
58
+ * sub-folder — a mobile app's module directory, a package in a monorepo, even the run trace
59
+ * itself while it inspects an artifact — fires every later hook with that folder as `cwd`. Read
60
+ * as the project root, that started a fresh `.shapeup/decisions.jsonl` in the sub-folder (nine of
61
+ * them on one measured run, one inside the committed tier and two inside the run trace), left
62
+ * every row there with `run_id: null` because the active-scope pointer was not beside it, and —
63
+ * the part that matters more than a split audit log — made `sandbox-guard` fail open on every
64
+ * write from that shell, since the active-order pointer it fences from was not beside it either.
65
+ *
66
+ * So the root is FOUND, not assumed: walk up from `cwd` to the nearest ancestor that carries a
67
+ * run pointer, the committed tier, or a git boundary. The order is the order of specificity — a
68
+ * live run outranks a repo boundary, so a project nested inside a larger repository still files
69
+ * under its own root — and a bare `.shapeup/` directory is deliberately NOT a marker, because the
70
+ * stray ledgers this fixes are exactly what would create one.
71
+ *
72
+ * Fail-open: no marker anywhere up the tree returns `cwd` unchanged, which is the pre-fix
73
+ * behaviour. Never throws.
74
+ *
75
+ * @param {string} cwd - Where the hook fired (`payload.cwd`, else the process cwd).
76
+ * @returns {string} The nearest project root at or above `cwd`, or `cwd` itself.
77
+ */
78
+ export function projectRoot(cwd) {
79
+ let dir;
80
+ try { dir = resolve(cwd || process.cwd()); } catch { return cwd; }
81
+ const isDir = (p) => { try { return statSync(p).isDirectory(); } catch { return false; } };
82
+ for (let i = 0; i < 64; i++) {
83
+ try {
84
+ if (existsSync(activeScope(dir))) return dir;
85
+ if (isDir(sharedDir(dir))) return dir;
86
+ if (existsSync(resolve(dir, ".git"))) return dir;
87
+ } catch { /* unreadable ancestor — keep climbing */ }
88
+ const parent = dirname(dir);
89
+ if (parent === dir) break;
90
+ dir = parent;
91
+ }
92
+ return cwd;
93
+ }
94
+
54
95
  /**
55
96
  * Where the receipts land.
56
97
  *
@@ -60,12 +101,14 @@ import { resolveRunId } from "../../kernel/lib/paths.mjs";
60
101
  * they were evaluations from a real run. A measurement instrument that its own test suite
61
102
  * contaminates is not an instrument.
62
103
  *
63
- * @param {string} [cwd] - Project root; defaults to the process cwd.
104
+ * @param {string} [cwd] - Where the hook fired; defaults to the process cwd. Resolved to the
105
+ * project root through {@link projectRoot} — a hook fired from a sub-folder files under the
106
+ * same ledger as one fired from the top.
64
107
  * @returns {string} The ledger path — `SHAPEUP_DECISIONS_PATH` when set, else the LOCAL root's
65
108
  * `decisions.jsonl`, resolved through `lib/paths.mjs`.
66
109
  */
67
110
  export function decisionsPath(cwd) {
68
- return process.env.SHAPEUP_DECISIONS_PATH || decisions(cwd || process.cwd());
111
+ return process.env.SHAPEUP_DECISIONS_PATH || decisions(projectRoot(cwd || process.cwd()));
69
112
  }
70
113
 
71
114
  /**
@@ -178,7 +221,7 @@ export async function runHook(name, fn) {
178
221
  // run, and recording that is what lets the export tier partition ambient decisions from run
179
222
  // ones. Resolution reads two small files and swallows every error — a receipt must never be
180
223
  // able to fail a tool call.
181
- run_id: (() => { try { return resolveRunId(d.cwd || process.cwd()); } catch { return null; } })(),
224
+ run_id: (() => { try { return resolveRunId(projectRoot(d.cwd || process.cwd())); } catch { return null; } })(),
182
225
  event: d.event ?? null,
183
226
  tool: d.tool ?? null,
184
227
  subject: d.subject ?? null,
@@ -36,7 +36,7 @@ import { globToRegExp, logPathology } from "./sandbox-guard.mjs";
36
36
  import { isMain } from "../kernel/lib/argv.mjs";
37
37
  import { LOCAL, safetyOverrides, metricsShard } from "../kernel/lib/paths.mjs";
38
38
 
39
- import { runHook, readStdin, settle } from "./lib/decision.mjs";
39
+ import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
40
40
 
41
41
  // --- overrides ---------------------------------------------------------------
42
42
 
@@ -225,9 +225,12 @@ async function main() {
225
225
 
226
226
  if (!HOOK_TOOLS.has(p.tool_name)) defer(`${p.tool_name ?? "no tool_name"} is not a guarded tool — out of scope`);
227
227
 
228
+ // The envelope and the telemetry shard live at the project root; the shell may be anywhere below
229
+ // it (see `projectRoot`). A relative tool path still means "relative to the shell".
228
230
  const cwd = p.cwd || process.cwd();
229
- const overrides = loadOverrides(cwd);
230
- const metricsPath = metricsShard(cwd);
231
+ const root = projectRoot(cwd);
232
+ const overrides = loadOverrides(root);
233
+ const metricsPath = metricsShard(root);
231
234
 
232
235
  const deny = (category, reason, detail) => {
233
236
  logPathology(metricsPath, {
@@ -240,7 +243,7 @@ async function main() {
240
243
  ...detail,
241
244
  });
242
245
  settle({
243
- verdict: "deny", event: "PreToolUse", tool: p.tool_name, cwd, rule: category,
246
+ verdict: "deny", event: "PreToolUse", tool: p.tool_name, cwd: root, rule: category,
244
247
  subject: detail?.path ?? detail?.command ?? null, reason,
245
248
  payload: {
246
249
  hookSpecificOutput: {
@@ -282,7 +285,7 @@ async function main() {
282
285
  }
283
286
 
284
287
  // Write | Edit | MultiEdit — only the self-protect rule; substrates stay sandbox-guard's job.
285
- const overridesAbs = resolve(safetyOverrides(cwd));
288
+ const overridesAbs = resolve(safetyOverrides(root));
286
289
  const hit = extractPaths(p.tool_input).find((raw) => resolve(cwd, raw) === overridesAbs);
287
290
  if (hit) {
288
291
  deny("self-protect", `The safety-overrides file is human-authored only — the session must never widen (or remove) its own safety envelope. Ask the PO to edit ${LOCAL}/safety-overrides.json.`, { path: hit });
@@ -74,7 +74,7 @@ import { readFileSync, existsSync, appendFileSync, mkdirSync, readdirSync, statS
74
74
  import { resolve, join, relative, dirname, sep } from "node:path";
75
75
  import { isMain } from "../kernel/lib/argv.mjs";
76
76
  import { LOCAL, SHARED, activeOrder, ordersDir, resultsDir, metricsShard } from "../kernel/lib/paths.mjs";
77
- import { runHook, readStdin, settle } from "./lib/decision.mjs";
77
+ import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
78
78
 
79
79
  // --- tiny glob matcher: supports *, **, ? — enough for substrate globs, zero dependencies ---
80
80
  export function globToRegExp(glob) {
@@ -184,8 +184,15 @@ async function main() {
184
184
  defer(`${p.tool_name ?? "no tool_name"} is not a write tool — out of scope`);
185
185
  }
186
186
 
187
+ // WHERE THE SHELL IS versus WHERE THE PROJECT IS. `p.cwd` follows the worker's shell — a leg that
188
+ // `cd`s into a sub-folder fires this hook from there — and the pointer, the order set and the
189
+ // substrate globs all live at the project root. Read from the sub-folder, the pointer was simply
190
+ // absent and this guard deferred at `no-round` on every write from that shell: a substrate fence
191
+ // that switches off whenever the worker changes directory. Raw tool paths still resolve against
192
+ // the shell's own cwd, because that is what a relative path in the tool input means.
187
193
  const cwd = p.cwd || process.cwd();
188
- const activeOrderPath = activeOrder(cwd);
194
+ const root = projectRoot(cwd);
195
+ const activeOrderPath = activeOrder(root);
189
196
  if (!existsSync(activeOrderPath)) defer("no active-order pointer — no tracked task running", "no-round");
190
197
 
191
198
  const active = readJSON(activeOrderPath);
@@ -202,7 +209,7 @@ async function main() {
202
209
  // Live = compiled and not yet answered. An order whose result is on disk has finished; leaving it
203
210
  // in the candidate set would keep a finished scope's substrate open for the rest of the run — and
204
211
  // leaving the POINTER's own order in unconditionally kept a finished RUN's substrate open forever.
205
- const orders = liveOrders(cwd, active.slug);
212
+ const orders = liveOrders(root, active.slug);
206
213
  if (orders.length === 0) defer(`no live order for ${active.slug}`, "no-order");
207
214
 
208
215
  const withSubstrate = orders.filter((o) => o.substrate);
@@ -220,14 +227,14 @@ async function main() {
220
227
  const targetPaths = extractPaths(p.tool_input);
221
228
  if (targetPaths.length === 0) defer("no writable path in the tool input", "no-target");
222
229
 
223
- const metricsPath = metricsShard(cwd);
230
+ const metricsPath = metricsShard(root);
224
231
  const runTracePrefix = join(LOCAL, active.slug) + sep;
225
232
  const violations = [];
226
233
  const blockReasons = [];
227
234
 
228
235
  for (const raw of targetPaths) {
229
236
  const abs = resolve(cwd, raw);
230
- const rel = relative(cwd, abs);
237
+ const rel = relative(root, abs);
231
238
  if (rel.startsWith(runTracePrefix)) continue;
232
239
 
233
240
  // Frozen takes absolute precedence, and it is checked across EVERY live contract: a path one
@@ -283,7 +290,7 @@ async function main() {
283
290
  });
284
291
 
285
292
  return {
286
- verdict: "deny", event: "PreToolUse", tool: p.tool_name, subject: active.order_path, cwd,
293
+ verdict: "deny", event: "PreToolUse", tool: p.tool_name, subject: active.order_path, cwd: root,
287
294
  rule: "outside-substrate",
288
295
  reason: `${violations.length} write(s) rejected by substrate boundaries: ${blockReasons.join("; ")}`,
289
296
  payload: {
@@ -26,7 +26,7 @@
26
26
  // pretty-printed envelope, colocated so audits can read it). Prints the path on stdout.
27
27
 
28
28
  import { readFileSync, writeFileSync, mkdirSync, existsSync, readdirSync } from "node:fs";
29
- import { resolve, join, dirname, basename } from "node:path";
29
+ import { resolve, join, dirname, basename, relative, sep } from "node:path";
30
30
  import { fileURLToPath } from "node:url";
31
31
  import { validate } from "./verify/envelope.mjs";
32
32
  import { readTrials } from "./verify/t0.mjs";
@@ -41,11 +41,22 @@ import {
41
41
  import { readContract, readAllContracts, tasksForScope, SCOPE_CONTRACT } from "./lib/contract.mjs";
42
42
  import { writeActiveOrder } from "./probe/resume.mjs";
43
43
  import { greenVerdict } from "./probe/t0.mjs";
44
+ import { latestRoundBuild } from "./verify/build.mjs";
44
45
  // The SAME matcher the sandbox hook enforces with. "Is this cited file inside this scope's
45
46
  // substrate" has to mean exactly what the guard means, or a bug is addressed to a scope that is
46
47
  // then denied the write that fixes it.
47
48
  import { matchesAny } from "../hooks/sandbox-guard.mjs";
48
49
 
50
+ /**
51
+ * Workers that read a coaching file. Kept beside the compile step because this is the only place
52
+ * the file is handed over: a worker not in this set never sees `payload.kb_rules_path`, however
53
+ * many rules the coach files for it. Mirrored by the coach skill's category list (structural test).
54
+ */
55
+ export const COACHABLE = new Set([
56
+ "task-executor", "ba-pitch-analyzer", "qa-edge-hunter",
57
+ "orient", "scope-architect", "solution-architect",
58
+ ]);
59
+
49
60
  const HERE = dirname(fileURLToPath(import.meta.url));
50
61
  const ORDER_SCHEMA = JSON.parse(readFileSync(resolve(HERE, "./../skills/tech-lead/schemas/work-order.schema.json"), "utf8"));
51
62
 
@@ -167,13 +178,20 @@ export const OP_OWNER = {
167
178
  wire: "solution-architect", evaluate: "spec-evaluator", orient: "orient",
168
179
  hunt: "qa-edge-hunter", translate: "translator",
169
180
  hammer: "scope-hammer", coach: "coach",
181
+ // `scan` is the coach reading the project instead of L4 feedback, and `research` the coach
182
+ // reading the platform's official documentation instead of the project — a project with nothing
183
+ // on disk has nothing to scan. Same worker, same write surface, same categorization gate. An
184
+ // operation the schema enumerates but this table does not route compiles no order at all, so the
185
+ // suite checks the two the other way round as well.
186
+ scan: "coach",
187
+ research: "coach",
170
188
  };
171
189
 
172
190
  /**
173
191
  * Resolve the write-contract (sandbox substrate) for an operation — one whitelist template per
174
192
  * operation, so mode/flag differences are enforced by the sandbox hook reading the order's substrate, not trusted to prose.
175
193
  * @param {string} operation - The order's operation (execute|fix|spike|analyze|reconcile|
176
- * retrofit-surface|coverage|map-scopes|wire|evaluate|orient|hunt|translate|hammer|coach).
194
+ * retrofit-surface|coverage|map-scopes|wire|evaluate|orient|hunt|translate|hammer|coach|scan|research).
177
195
  * @param {{slug?:string, specDir?:string, scope?:object}} [ctx] - slug (names LOCAL/SHARED roots),
178
196
  * specDir (overrides the default spec path), scope (contract supplying allowed/shared substrates).
179
197
  * @returns {{allowed:string[], shared?:string[], frozen?:string[], append_only?:string[]}} The
@@ -236,6 +254,13 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
236
254
  case "hammer":
237
255
  return { allowed: [globShared(slug, "REPORT.md"), `${local}/reports/**`] };
238
256
  case "coach":
257
+ case "scan":
258
+ case "research":
259
+ // `scan` seeds the same files from the project on disk and `research` from the platform's
260
+ // official documentation, instead of from L4 feedback; all three write only the knowledge
261
+ // base, and the profile they suggest probes for stays the tech lead's to write (single
262
+ // writer of the committed tier). Research is a source, not a verification: what it reads
263
+ // is still a claim until the tech lead pins it at L0 and the kernel runs it.
239
264
  return { allowed: [relKnowledgeBase("*")] };
240
265
  default:
241
266
  return { allowed: [`${local}/**`] };
@@ -379,6 +404,80 @@ export function verdictBugs(cwd, slug, round) {
379
404
  return v.bugs.filter((b) => !refuted.has(String(b?.id)) && !refuted.has(String(b?.criterion)));
380
405
  }
381
406
 
407
+ /**
408
+ * The previous round's RED BUILD GATE, as bug entries the fix round can act on.
409
+ *
410
+ * THE JUDGE IS NOT THE ONLY SOURCE OF A FAIL. `verify build` runs the feature's build and its
411
+ * launch probe once per round, before EVAL, and a red gate ends the round with no verdict — there is
412
+ * nothing to grade in an app that does not compile or start. But a round that ends without a
413
+ * verdict left the next round with no `payload.bugs`, so it compiled as a plain `execute`, every
414
+ * worker re-ran its fixtures, found them green (they never tested the build) and reported done —
415
+ * the identical loop the ledgered verdict was threaded through {@link verdictBugs} to break.
416
+ *
417
+ * So a failing step becomes a bug: the criterion is the command that must exit 0, the locator is
418
+ * every file the tool's own output names that exists in the tree, and the evidence is the output's
419
+ * tail. Addressed by {@link bugsForScope} exactly like a judged defect — the scope whose substrate
420
+ * holds the cited file gets it, and a failure citing no file goes to every scope marked `unowned`.
421
+ *
422
+ * @param {string} cwd - Project root.
423
+ * @param {string} slug - Feature slug.
424
+ * @param {number} [round] - The round being compiled; round 1 has no predecessor.
425
+ * @returns {Array<object>} One entry per failing step of the previous round's latest gate
426
+ * artifact; empty when the gate did not run, was green, or this is round 1.
427
+ */
428
+ export function buildBugs(cwd, slug, round) {
429
+ if (!round || round < 2) return [];
430
+ const gate = latestRoundBuild(cwd, slug, round - 1);
431
+ if (!gate || gate.overall !== "red") return [];
432
+ const out = [];
433
+ for (const step of gate.steps || []) {
434
+ if (step.skipped || step.pass) continue;
435
+ const text = `${step.stdout_tail || ""}\n${step.stderr_tail || ""}`;
436
+ const files = outputPaths(text, cwd);
437
+ out.push({
438
+ id: `BUILD-r${round - 1}-${step.kind}`,
439
+ severity: "blocker",
440
+ source: "verify build",
441
+ criterion: `${step.kind} exits 0: ${step.cmd}`,
442
+ location: files.join(", "),
443
+ expected: "exit 0",
444
+ actual: step.error ? `did not run: ${step.error}` : `exit ${step.exit}`,
445
+ evidence: (step.stderr_tail || step.stdout_tail || "").slice(-2000),
446
+ digest: Array.isArray(gate.discovered_tasks) ? gate.discovered_tasks.slice(0, 8) : [],
447
+ });
448
+ }
449
+ return out;
450
+ }
451
+
452
+ /**
453
+ * The project files a build or launch log names, repo-relative and de-duplicated.
454
+ *
455
+ * Wider than {@link bugLocations} in one way and narrower in another: it accepts absolute paths
456
+ * (compilers print them) and strips the project root off; and it keeps only paths that EXIST under
457
+ * the project, because a tool's own stack frames and cache paths look exactly like file citations
458
+ * and would address the bug to nobody. First-seen order.
459
+ *
460
+ * @param {string} text - Raw log text.
461
+ * @param {string} cwd - Project root.
462
+ * @returns {string[]} Repo-relative POSIX paths.
463
+ */
464
+ export function outputPaths(text, cwd) {
465
+ const out = [];
466
+ const root = resolve(cwd);
467
+ for (const m of String(text || "").matchAll(/(?:^|[\s(,\[:"'`])((?:\/|\.\/)?[\w@][\w.\/-]*\.[A-Za-z][A-Za-z0-9]{0,5})(?::\d+)?/gm)) {
468
+ let p = m[1].replace(/^\.\//, "");
469
+ if (p.startsWith("/")) {
470
+ const rel = relative(root, p);
471
+ if (!rel || rel.startsWith("..")) continue;
472
+ p = rel.split(sep).join("/");
473
+ }
474
+ if (out.includes(p)) continue;
475
+ if (!existsSync(join(root, p))) continue;
476
+ out.push(p);
477
+ }
478
+ return out;
479
+ }
480
+
382
481
  /**
383
482
  * Every repo-relative file a bug is cited against.
384
483
  *
@@ -543,7 +642,7 @@ export function t0ArtifactsFor(cwd, slug, round) {
543
642
  * @param {string} [opts.compiledAt] - ISO compile time; omitted rather than invented.
544
643
  * @returns {object} A WorkOrder: {schema_version, order_id ("<slug>/<suffix>"), run_id?,
545
644
  * compiled_at?, worker, mode, operation?, interaction?, substrate (from {@link substrateFor}),
546
- * payload{…}}. A coachable worker also gets payload.kb_rules_path. Not validated here — the CLI
645
+ * payload{…}}. A coachable worker ({@link COACHABLE}) also gets payload.kb_rules_path. Not validated here — the CLI
547
646
  * validates before writing.
548
647
  */
549
648
  export function compileOrder({
@@ -620,8 +719,12 @@ export function compileOrder({
620
719
  ...(payloadExtra || {}),
621
720
  },
622
721
  };
623
- const kbByWorker = { "task-executor": "task-executor", "ba-pitch-analyzer": "ba-pitch-analyzer", "qa-edge-hunter": "qa-edge-hunter" };
624
- if (kbByWorker[worker]) order.payload.kb_rules_path = relKnowledgeBase(kbByWorker[worker]);
722
+ // The coachable set — every worker that reads `shapeup/knowledge-base/<worker>.md` at the top
723
+ // of its run. Guidance only: a rule here steers craft and never resolves, skips or reorders a
724
+ // gate, and never widens a substrate (the sandbox hook reads the order, not the KB). The judge
725
+ // (spec-evaluator) and the census (scope-hammer) are excluded on purpose: coaching the judge
726
+ // makes a second grader, and the hammer's ownership claims must come from `probe owner`.
727
+ if (COACHABLE.has(worker)) order.payload.kb_rules_path = relKnowledgeBase(worker);
625
728
  return order;
626
729
  }
627
730
 
@@ -695,7 +798,10 @@ export async function cli(rawArgv) {
695
798
  // The fix round's inbound evidence. Derived here, from the ledgered verdict, for every lane —
696
799
  // the workflow, `--tiny`, the prose round loop and a standalone `/build` all compile through
697
800
  // this line, and none of them can pass a payload to a build order (see the banner above).
698
- const bugs = scope ? bugsForScope(verdictBugs(cwd, slug, round), scope.scope_id, scopeSubstrates(cwd, slug)) : [];
801
+ // Two sources, one channel: the judge's cited defects and the build gate's failing steps.
802
+ const bugs = scope
803
+ ? bugsForScope([...verdictBugs(cwd, slug, round), ...buildBugs(cwd, slug, round)], scope.scope_id, scopeSubstrates(cwd, slug))
804
+ : [];
699
805
 
700
806
  // A ROUND CARRYING CITED DEFECTS IS A `fix`, AND THE ORDER HAS TO SAY SO.
701
807
  //
@@ -21,14 +21,17 @@
21
21
  //
22
22
  // verify t0 · budget · envelope · Measured, not claimed. A model verifying itself is
23
23
  // trace · spec · skills · claiming; these read artifacts and re-hash them.
24
- // dispatch `skills` reads the roster off disk; `dispatch` reads
24
+ // dispatch · build `skills` reads the roster off disk; `dispatch` reads
25
25
  // the hook layer's evidence that a skill really resolved
26
26
  // in this session — the half a file check cannot answer.
27
+ // `build` runs the feature's build and launch probes once
28
+ // per round before EVAL — T0 is per scope and proves
29
+ // nothing about whether the whole compiles or starts.
27
30
  // reduce ingest · hill · snapshot · Single writer. Shared state has exactly one author.
28
31
  // ship · board · verdict · graph
29
32
  // gate An answer file with a source, not a vibe.
30
33
  // probe resume · t0 · stats · digest · Read-only queries over run state. `concurrency`
31
- // concurrency · leg · eval answers how many legs ran at once and what the
34
+ // concurrency · leg · eval · owner answers how many legs ran at once and what the
32
35
  // fan-out bought, and refuses a figure the record set
33
36
  // cannot support rather than printing a plausible one.
34
37
  // `leg` answers whether a scope's work reached the
@@ -39,7 +42,10 @@
39
42
  // `eval` answers what an EVAL round's WorkResult
40
43
  // actually said, mechanically — the round-loop branch
41
44
  // reads this instead of trusting a dispatching agent's
42
- // own end-of-turn summary of its own verdict.
45
+ // own end-of-turn summary of its own verdict. `owner`
46
+ // answers which scope may write a path, elected from
47
+ // the contracts — so a census cites it instead of
48
+ // asserting ownership from memory.
43
49
  // init run · fit Opens a run, or refuses it (exit 3).
44
50
  // report export Projects the run's records as fact tables.
45
51
  // compile The WorkOrder: schema-valid or nothing is dispatched.
@@ -64,7 +70,7 @@ export const ROUTES = {
64
70
  verify: {
65
71
  t0: "./verify/t0.mjs", budget: "./verify/budget.mjs", envelope: "./verify/envelope.mjs",
66
72
  trace: "./verify/trace.mjs", spec: "./verify/spec.mjs", skills: "./verify/skills.mjs",
67
- dispatch: "./verify/dispatch.mjs",
73
+ dispatch: "./verify/dispatch.mjs", build: "./verify/build.mjs",
68
74
  },
69
75
  reduce: {
70
76
  ingest: "./reduce/ingest.mjs", hill: "./reduce/hill.mjs", snapshot: "./reduce/snapshot.mjs",
@@ -74,7 +80,7 @@ export const ROUTES = {
74
80
  probe: {
75
81
  resume: "./probe/resume.mjs", t0: "./probe/t0.mjs", stats: "./probe/stats.mjs",
76
82
  digest: "./probe/digest.mjs", concurrency: "./probe/concurrency.mjs",
77
- leg: "./probe/leg.mjs", eval: "./probe/eval.mjs",
83
+ leg: "./probe/leg.mjs", eval: "./probe/eval.mjs", owner: "./probe/owner.mjs",
78
84
  },
79
85
  init: { run: "./init/run.mjs", fit: "./init/fit.mjs" },
80
86
  report: { export: "./report/export.mjs", _default: "export" },