shapeup-sdlc 3.6.0 → 3.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,152 @@
1
+ // probe attempts — "how many of this scope's attempts are ATTESTED work, not writable artifacts?"
2
+ //
3
+ // CONTRACT. A bounded, read-only query over one scope's attested channels for one round: dispatch
4
+ // receipts (`receipts/dispatch.jsonl`), leg-completion rows (`legs.jsonl`) and WorkResults
5
+ // (`results/`). Prints `{scope_id, round, attempt_budget, spent, in_flight, unattested, green,
6
+ // tripped, attempts}` on stdout; exits 0 when the breaker holds, 1 when it has tripped, 2 on a bad
7
+ // argv. Writes nothing.
8
+ //
9
+ // WHY IT EXISTS. Measured on a real run: attempt 2 of a scope was compiled and T0-verified
10
+ // several minutes BEFORE attempt 1 ingested — while attempt 1 was still in flight. No worker was
11
+ // ever dispatched for attempt 2:
12
+ // `receipts/dispatch.jsonl`, `legs.jsonl` and `results/` carried no row for it. A derivation keyed
13
+ // off the order set (`orders/`) or the T0 verdict set (`t0/verdicts/`) alone counts a compiled
14
+ // order or a green trial as a spent attempt regardless of whether a worker ever ran — both are
15
+ // WRITABLE by the very leg whose exhaustion is being judged. This module counts an attempt as
16
+ // SPENT only when a dispatch receipt attests it started AND either a leg-completion row or a
17
+ // WorkResult on disk attests it closed. Neither channel alone is enough: a receipt with no result
18
+ // is a leg still in flight (open, not spent — the breaker must not trip on unanswered work), and a
19
+ // result or verdict with no receipt is unattested — no worker ran, which is the shape that
20
+ // produced this module.
21
+ //
22
+ // WHY THE SAME FUNCTION SERVES THE BREAKER AND THE CENSUS. Before this module, the round loop's
23
+ // inner breaker and `scope-hammer`'s GATE H0 census read different evidence for the same question
24
+ // — the loop trusted the worker's own self-reported `attempts_used`/`breaker` (schema-shaped, not
25
+ // re-verified), the census read `t0/verdicts/*.json` directly — and disagreed on the measured run.
26
+ // One shared, attested derivation is what makes "the census and the breaker agree" true by
27
+ // construction rather than by coincidence: `skills/scope-hammer/SKILL.md` cites this same probe
28
+ // the way it already cites `probe owner` for ownership claims.
29
+
30
+ import { existsSync, readdirSync, readFileSync } from "node:fs";
31
+ import { join, resolve } from "node:path";
32
+ import { runArgs } from "../lib/argv.mjs";
33
+ import { dispatchReceipts, legLedger, resultsDir, readRunId } from "../lib/paths.mjs";
34
+ import { readLegs } from "./leg.mjs";
35
+ import { greenVerdict } from "./t0.mjs";
36
+
37
+ /**
38
+ * Every dispatch-receipt row on disk, tolerant of a torn last line (mirrors {@link readLegs}).
39
+ * @param {string} path - `receipts/dispatch.jsonl`.
40
+ * @returns {object[]} Parsed rows; an unparsable line is skipped rather than fatal.
41
+ */
42
+ export function readReceipts(path) {
43
+ if (!existsSync(path)) return [];
44
+ return readFileSync(path, "utf8").split("\n").filter(Boolean)
45
+ .map((l) => { try { return JSON.parse(l); } catch { return null; } })
46
+ .filter(Boolean);
47
+ }
48
+
49
+ /**
50
+ * One attempt's evidence across the three attested channels, and the state it derives to.
51
+ *
52
+ * `unattested` (no receipt at all) counts as though the attempt never happened — the order and any
53
+ * T0 verdict may still be sitting on disk, written by something other than a dispatched worker, and
54
+ * neither is asked here. `in-flight` (a receipt, but no leg row and no result) is a real dispatch
55
+ * whose leg has not yet closed: open, not spent. `spent` is closed either way a leg closes — via
56
+ * `reduce ingest`'s own leg-completion row, or a WorkResult already on disk pending ingest.
57
+ *
58
+ * @param {string} cwd - Project root.
59
+ * @param {string} slug - Feature slug.
60
+ * @param {string} scopeId - Scope contract id (the order id's own address, e.g. `shell-r1-a2`).
61
+ * @param {number} round - Build round.
62
+ * @param {number} attempt - Attempt number within the round.
63
+ * @param {object[]} receipts - Pre-read `receipts/dispatch.jsonl` rows (avoids re-reading per attempt).
64
+ * @param {object[]} legs - Pre-read `legs.jsonl` rows.
65
+ * @returns {{orderId:string, hasReceipt:boolean, hasResult:boolean, hasLeg:boolean, state:("unattested"|"in-flight"|"spent")}}
66
+ */
67
+ export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, legs, runId = null) {
68
+ const orderId = `${slug}/${scopeId}-r${round}-a${attempt}`;
69
+ // SCOPED TO THIS RUN, and that qualifier is the whole correction. `receipts/dispatch.jsonl` and
70
+ // `legs.jsonl` are per-slug and append-only, so they accumulate across every run of a pitch,
71
+ // while `order_id` repeats — the run key is the only thing that separates two runs of one
72
+ // feature. Matching on `order_id` alone answered "was this attempt spent?" with a PREVIOUS run's
73
+ // receipt: measured on a consumer, a run that dispatched nothing was told its first attempt was
74
+ // spent, complete with leg and result, by rows two launches old.
75
+ //
76
+ // A row carrying no run key belongs to NO run rather than to this one, and an unresolvable
77
+ // current run (no receipt on disk) matches nothing. Both directions under-count rather than
78
+ // over-count, which is the safe way to be wrong here: an under-count leaves a breaker un-tripped
79
+ // and the round continues, where an over-count stops work that was never done.
80
+ const mine = (r) => r?.order_id === orderId && runId != null && r?.run_id === runId;
81
+ const hasReceipt = receipts.some(mine);
82
+ const hasLeg = legs.some(mine);
83
+ // A WorkResult carries no run key and reaches one only through its `order_id`, which repeats —
84
+ // so this stays a file check and is deliberately NOT sufficient on its own. It can only turn an
85
+ // already run-scoped receipt into `spent`; a result left behind by an earlier run cannot attest
86
+ // an attempt this run never dispatched.
87
+ const hasResult = existsSync(join(resultsDir(cwd, slug), `${scopeId}-r${round}-a${attempt}.json`));
88
+ const state = !hasReceipt ? "unattested" : (hasResult || hasLeg) ? "spent" : "in-flight";
89
+ return { orderId, hasReceipt, hasResult, hasLeg, state };
90
+ }
91
+
92
+ /**
93
+ * A scope's attempt census for one round, derived ONLY from attested channels — the function both
94
+ * the round loop's inner breaker and scope-hammer's GATE H0 census call, so they cannot drift apart
95
+ * again the way a measured run found them.
96
+ *
97
+ * @param {string} cwd - Project root.
98
+ * @param {string} slug - Feature slug.
99
+ * @param {string} scopeId - Scope contract id.
100
+ * @param {number} round - Build round.
101
+ * @param {number} attemptBudget - The scope's configured `attempt_budget`.
102
+ * @returns {{scope_id:string, round:number, attempt_budget:number, spent:number, in_flight:number,
103
+ * unattested:number, green:boolean, tripped:boolean, attempts:object[]}} `tripped` is true only
104
+ * when every attempt within budget is genuinely SPENT and none produced a green T0 — an attempt
105
+ * still in flight holds the breaker open regardless of how many slots are nominally used.
106
+ */
107
+ export function scopeAttempts(cwd, slug, scopeId, round, attemptBudget) {
108
+ const receipts = readReceipts(dispatchReceipts(cwd, slug));
109
+ const legs = readLegs(legLedger(cwd, slug));
110
+ const runId = readRunId(cwd, slug);
111
+ const attempts = [];
112
+ let spent = 0, inFlight = 0, unattested = 0;
113
+ for (let a = 1; a <= attemptBudget; a++) {
114
+ const ev = attemptEvidence(cwd, slug, scopeId, round, a, receipts, legs, runId);
115
+ if (ev.state === "spent") spent++;
116
+ else if (ev.state === "in-flight") inFlight++;
117
+ else unattested++;
118
+ attempts.push({ attempt: a, ...ev });
119
+ }
120
+ const { green } = greenVerdict(cwd, slug, scopeId, round);
121
+ return {
122
+ scope_id: scopeId, round, attempt_budget: attemptBudget,
123
+ spent, in_flight: inFlight, unattested, green,
124
+ tripped: !green && spent >= attemptBudget,
125
+ attempts,
126
+ };
127
+ }
128
+
129
+ export const ARGV_SPEC = {
130
+ usage: "harness.mjs probe attempts --slug <slug> --scope <scope-id> --round N --attempt-budget N [--cwd <dir>]",
131
+ _: { arity: 0, max: 0, name: "(no positional operands)" },
132
+ slug: { type: "str", required: true },
133
+ scope: { type: "str", required: true },
134
+ round: { type: "int", min: 1, required: true },
135
+ "attempt-budget": { type: "int", min: 1, required: true },
136
+ cwd: { type: "path" },
137
+ };
138
+
139
+ /**
140
+ * Report one scope's attested attempt census for one round.
141
+ *
142
+ * @param {string[]} rawArgv - The subcommand's own arguments (harness.mjs strips the verb words).
143
+ * @returns {void} Exits 0 when the breaker holds, 1 when it has tripped — the shape a caller (or a
144
+ * census citing this row) can branch on without parsing prose.
145
+ */
146
+ export function cli(rawArgv) {
147
+ const args = runArgs(ARGV_SPEC, rawArgv);
148
+ const cwd = resolve(args.cwd || process.cwd());
149
+ const r = scopeAttempts(cwd, args.slug, args.scope, args.round, args.attemptBudget);
150
+ console.log(JSON.stringify(r));
151
+ process.exit(r.tripped ? 1 : 0);
152
+ }
@@ -56,16 +56,18 @@
56
56
  // Exit: 0 ok · 2 malformed argv (nothing ran) · 3 the target the operation needs is not on disk ·
57
57
  // 6 the required phase's artifact is NOT on disk (the phase did not complete).
58
58
 
59
- import { existsSync, readdirSync, readFileSync, writeFileSync, mkdirSync } from "node:fs";
59
+ import { existsSync, readdirSync, readFileSync, writeFileSync, mkdirSync, rmSync } from "node:fs";
60
60
  import { dirname, join, resolve } from "node:path";
61
61
  import { runArgs } from "../lib/argv.mjs";
62
62
  import { splitFrontmatter, uncoerce } from "../lib/contract.mjs";
63
63
  import { globToRegExp } from "../verify/spec.mjs";
64
64
  import {
65
65
  intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir,
66
- orientDir, activeOrder, usecasesDir, breadboard, receipt, readReceipt, requirements,
66
+ orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements,
67
+ exportRunDir,
67
68
  } from "../lib/paths.mjs";
68
69
  import { evalVerdict } from "./eval.mjs";
70
+ import { collectRun, writeRun } from "../report/export.mjs";
69
71
 
70
72
  /** The run-state values `references/protocol.md` (Part 4 — State) defines. A typo'd status is a rejection,
71
73
  * not a write — the whole point of this file is that a write nobody validates is a write nobody
@@ -81,6 +83,66 @@ export const RUN_STATUSES = ["orienting", "mapping", "building", "evaluating", "
81
83
  */
82
84
  export const TERMINAL_STATUSES = ["shipped", "aborted", "escalated"];
83
85
 
86
+ /**
87
+ * The RunReturn union (`kernel/schemas/domain.schema.json` `$defs/RunReturn.properties.status.enum`)
88
+ * mapped to what closing the run means for each arm — the derivation `closeIfTerminal`
89
+ * (`skills/tech-lead/workflows/shapeup-run.js`) now reads instead of a hand-typed
90
+ * `status !== "aborted" && status !== "shipped"` pair that referenced {@link TERMINAL_STATUSES}
91
+ * zero times and so could not see when a new arm went unhandled.
92
+ *
93
+ * A lookup answers one of three ways, and the distinction is load-bearing for what this map must
94
+ * catch: an arm ABSENT from this object (never listed as a key) returns `undefined` — an arm the
95
+ * schema carries that nobody has mapped, which a caller must treat as a defect, never as "fine to
96
+ * skip". An arm mapped to a terminal status closes the run as that status. An arm mapped to `null`
97
+ * is EXPLICITLY non-terminal — `paused` resumes on relaunch and `ok` is one inner round finishing,
98
+ * not the run — so it is a key with a falsy value, not an omission a reader could mistake for "not
99
+ * decided yet".
100
+ *
101
+ * `gate_h` → `escalated`: the breaker that tripped (`outer`/`inner`/`deadline`) travels in
102
+ * `close_cause`, never as a new member of {@link TERMINAL_STATUSES} or {@link RUN_STATUSES} — a
103
+ * circuit breaker tripping is not a new way a run ends, it is the reason an existing one
104
+ * (`escalated`) fires this time.
105
+ *
106
+ * @type {Object<string, (string|null)>}
107
+ */
108
+ export const RUN_RETURN_CLOSE = {
109
+ shipped: "shipped",
110
+ aborted: "aborted",
111
+ gate_h: "escalated",
112
+ paused: null,
113
+ ok: null,
114
+ };
115
+
116
+ /**
117
+ * Resolve one RunReturn arm to a close outcome and, when the arm is terminal, perform the close —
118
+ * the single call `closeIfTerminal` makes instead of deciding locally which arms are terminal.
119
+ *
120
+ * @param {string} cwd - Project root.
121
+ * @param {string} slug - Feature slug.
122
+ * @param {string} arm - A `RunReturn.status` value.
123
+ * @param {(string|null)} [cause] - Why the run ended there; only used when `arm` is terminal.
124
+ * @param {boolean} [withExport] - Forwarded to {@link closeRun} — defaults true. The one caller
125
+ * that ever passes `false` is a test fixture proving the export assertion is real — a check that
126
+ * cannot fail did not happen; production call sites never set this.
127
+ * @returns {({ok:true, arm:string, terminal:false, reason:string} |
128
+ * {ok:false, arm:string, reason:string} |
129
+ * ({ok:boolean, arm:string, terminal:true} & ReturnType<typeof closeRun>))} `terminal:false` when
130
+ * `arm` maps to `null` (nothing closed, not an error). `ok:false` with no `terminal` field when
131
+ * `arm` is not a key of {@link RUN_RETURN_CLOSE} at all — a schema arm this map has not been
132
+ * taught, which must never be silently treated as non-terminal. Otherwise the {@link closeRun}
133
+ * outcome, tagged with the arm that produced it.
134
+ */
135
+ export function closeArm(cwd, slug, arm, cause = null, withExport = true) {
136
+ if (!Object.hasOwn(RUN_RETURN_CLOSE, arm)) {
137
+ return { ok: false, arm, reason: `closeArm: "${arm}" is not a RunReturn arm this kernel maps — known arms: ${Object.keys(RUN_RETURN_CLOSE).join(", ")}` };
138
+ }
139
+ const status = RUN_RETURN_CLOSE[arm];
140
+ if (!status) {
141
+ return { ok: true, arm, terminal: false, reason: `"${arm}" is explicitly non-terminal — no close` };
142
+ }
143
+ return { ...closeRun(cwd, slug, { status, cause, withExport }), arm, terminal: true };
144
+ }
145
+
84
146
  /** ORIENT's four artifacts (skills/orient/SKILL.md §Outputs): three by exact name, plus a spike
85
147
  * whose filename carries the area it spiked (`spike-<area>.md`, or `spike-not-needed.md` when
86
148
  * the risk scan came back rank 0 — both count, because both are ORIENT having finished). */
@@ -498,6 +560,40 @@ function writeCloseLines(body, { status, closedAt, cause }) {
498
560
  return out;
499
561
  }
500
562
 
563
+ /**
564
+ * Export a just-closed run's own records into fact tables, best-effort
565
+ * (`report export`, `kernel/report/export.mjs`). {@link closeRun} calls this for every terminal
566
+ * status except `"shipped"`: the Ship phase (`skills/tech-lead/workflows/shapeup-run.js`) already
567
+ * calls `report export` itself, several lines before this file's own close-out runs, and that call
568
+ * site is left untouched on purpose — the shipped path's export stays byte-comparable to what it
569
+ * wrote before Stage 2. `aborted`, `escalated` and `gate_h` (which closes as `escalated`, see
570
+ * {@link RUN_RETURN_CLOSE}) never reached that call site at all, because it sits inside a phase
571
+ * those endings never enter — so a run ending any of them left its whole trace (orders, results,
572
+ * T0 verdicts, hook decisions, `graph.jsonl`) in the gitignored LOCAL tier with nothing durable
573
+ * surviving the next `init run`'s wipe. This is the read that was missing for those endings, run at
574
+ * the one point every one of them passes through: the close itself.
575
+ *
576
+ * FAIL-OPEN, the hook discipline (CLAUDE.md) extended to a write that is not a hook: an export that
577
+ * cannot write — a blocked or missing exports directory, a full disk — must never turn a close that
578
+ * DID take into one that looks like it did not. The three facts {@link closeRun} just wrote
579
+ * (`closed_status`/`close_cause`/`closed_at`) are never touched by this function; a failure here is
580
+ * handed back to the caller as a warning string, never thrown.
581
+ *
582
+ * @param {string} cwd - Project root.
583
+ * @param {string} slug - Feature slug.
584
+ * @returns {(string|null)} A one-line warning when the export did not complete; `null` on success.
585
+ */
586
+ function exportOnClose(cwd, slug) {
587
+ try {
588
+ const collected = collectRun(cwd, slug);
589
+ if (!collected) return `export on close: no readable receipt for "${slug}" — nothing to export`;
590
+ writeRun(collected, exportRunDir(cwd, collected.run_id ?? slug));
591
+ return null;
592
+ } catch (e) {
593
+ return `export on close: ${e.message}`;
594
+ }
595
+ }
596
+
501
597
  /**
502
598
  * Close the run: a terminal status, its cause, and a close timestamp, written together in ONE
503
599
  * pass.
@@ -548,20 +644,26 @@ function writeCloseLines(body, { status, closedAt, cause }) {
548
644
  *
549
645
  * @param {string} cwd - Project root.
550
646
  * @param {string} slug - Feature slug.
551
- * @param {{status:string, cause:(string|null)}} o - The terminal status (one of
552
- * {@link TERMINAL_STATUSES}) and why the run ended there. `cause` travels through `uncoerce` (the
553
- * one dialect `harness-run.md`'s frontmatter is read and written in), so free prose — quotes and
554
- * colons included — round-trips as one frontmatter line; an embedded newline is collapsed to a
555
- * space first, because this dialect is line-based and could not carry one either way.
647
+ * @param {{status:string, cause:(string|null), withExport?:boolean}} o - The terminal status (one
648
+ * of {@link TERMINAL_STATUSES}) and why the run ended there. `cause` travels through `uncoerce`
649
+ * (the one dialect `harness-run.md`'s frontmatter is read and written in), so free prose — quotes
650
+ * and colons included — round-trips as one frontmatter line; an embedded newline is collapsed to
651
+ * a space first, because this dialect is line-based and could not carry one either way.
652
+ * `withExport` (default true) gates {@link exportOnClose} — every status except `"shipped"` runs
653
+ * it on a successful close; `false` exists only for a test fixture proving the export assertion
654
+ * is real, never for a production call site.
556
655
  * @returns {{ok:boolean, path:string, status:string, closed_at?:string, cause?:(string|null),
557
656
  * reason?:string, closed_status?:string, superseded?:boolean, decision?:string,
558
- * prior_cause?:(string|null), prior_closed_at?:string}} Outcome. A refused overwrite (already
559
- * closed with a DIFFERENT terminal status) carries `closed_status`/`closed_at`/`cause` naming what
560
- * is actually on disk. A successful supersede (same status, different cause) carries
561
- * `superseded:true`, `decision:"superseded"` (the one-token signal the courier boundary in
562
- * `shapeup-run.js` relays verbatim — see its own `cmd()` banner) and the prior close it folded in.
657
+ * prior_cause?:(string|null), prior_closed_at?:string, export_warning?:string}} Outcome. A refused
658
+ * overwrite (already closed with a DIFFERENT terminal status) carries
659
+ * `closed_status`/`closed_at`/`cause` naming what is actually on disk. A successful supersede
660
+ * (same status, different cause) carries `superseded:true`, `decision:"superseded"` (the
661
+ * one-token signal the courier boundary in `shapeup-run.js` relays verbatim — see its own
662
+ * `cmd()` banner) and the prior close it folded in. Any successful, non-`"shipped"` close carries
663
+ * `export_warning` when {@link exportOnClose} could not write the run's fact tables — the close
664
+ * itself still stands; this is advisory only.
563
665
  */
564
- export function closeRun(cwd, slug, { status, cause = null } = {}) {
666
+ export function closeRun(cwd, slug, { status, cause = null, withExport = true } = {}) {
565
667
  const p = harnessRun(cwd, slug);
566
668
  if (!TERMINAL_STATUSES.includes(status)) {
567
669
  return { ok: false, path: p, status, reason: `closeRun: "${status}" is not terminal — expected one of ${TERMINAL_STATUSES.join(" | ")}` };
@@ -574,6 +676,51 @@ export function closeRun(cwd, slug, { status, cause = null } = {}) {
574
676
  return { ok: false, path: p, status, reason: `harness-run.md carries no "status:"/"closed_at:" line to replace — the ledger's frontmatter is malformed (references/protocol.md)` };
575
677
  }
576
678
 
679
+ // The Ship phase (`shapeup-run.js`) already exports a shipped run itself, before this call ever
680
+ // runs — Stage 2 adds the endings that wrote nothing, and leaves that path untouched.
681
+ const shouldExport = withExport && status !== "shipped";
682
+
683
+ /**
684
+ * Everything a close owes the checkout once the ledger line is written: export the run's
685
+ * records, then retire its pointers.
686
+ *
687
+ * The export runs FIRST, but not because it has to: `exportOnClose` is handed the slug and keys
688
+ * its output by the receipt's `run_id`, so it never reads the pointers this retires. (The bare
689
+ * `reduce export` CLI does read `active-scope` — only to work out which run the operator meant
690
+ * when they named none.) Reporting before teardown is a defensive default, not a correctness
691
+ * requirement, and it is recorded as such so the next reader does not defend an ordering that
692
+ * carries nothing. Measured: swapping the two leaves every check green.
693
+ *
694
+ * WHY RETIRE AT ALL. The substrate fence is enforced while an order is compiled and unanswered,
695
+ * and a run that ends any way other than shipping leaves exactly that by construction. Until this
696
+ * ran, the fence outlived the run: after a close that exited 0 and recorded everything, an
697
+ * ordinary write anywhere in the project was still denied, with no dispatch in flight. The
698
+ * operator's obvious remedy did not help either — `init run --force` is documented as "abandon
699
+ * the open run and start over", never as "release a stuck fence".
700
+ *
701
+ * AND RETIRING IS NOT ANSWERING. The pointer says "a run is in flight"; the order's missing
702
+ * result says "nobody came back". Only the first is untrue after a close. The abandoned order
703
+ * stays unanswered, the attempt census still sees nothing spent on it, and `init run --force`
704
+ * remains the one thing that writes a synthetic result — because a close that quietly claimed the
705
+ * work was answered would spend an attempt budget on work nobody did.
706
+ *
707
+ * Best-effort, like the export: a pointer that cannot be removed degrades the close and is
708
+ * reported in its return, but never turns a close into a non-close.
709
+ */
710
+ const finishClose = (result) => {
711
+ const warning = shouldExport ? exportOnClose(cwd, slug) : null;
712
+ const stuck = [];
713
+ for (const pointer of [activeOrder(cwd), activeScope(cwd)]) {
714
+ try { rmSync(pointer, { force: true }); } catch { /* fall through to the check below */ }
715
+ if (existsSync(pointer)) stuck.push(pointer);
716
+ }
717
+ return {
718
+ ...result,
719
+ ...(warning ? { export_warning: warning } : {}),
720
+ ...(stuck.length ? { pointer_warning: `could not retire ${stuck.join(", ")} — the substrate fence may still deny writes until it is removed by hand` } : {}),
721
+ };
722
+ };
723
+
577
724
  // Truncated, not elided: a cause this long has already done its job in the run's own log — the
578
725
  // ledger line is a pointer back to it, not the full transcript. Newlines are collapsed to spaces
579
726
  // FIRST — this dialect is line-based, so a raw embedded newline would split one field into a value
@@ -588,7 +735,7 @@ export function closeRun(cwd, slug, { status, cause = null } = {}) {
588
735
  if (priorClosedStatus && priorClosedAt) {
589
736
  if (priorClosedStatus === status && normCause === priorCause) {
590
737
  // The identical fact, restated — a retried or duplicated call costs nothing.
591
- return { ok: true, path: p, status, closed_at: priorClosedAt, cause: priorCause, decision: "idempotent", reason: `already closed as "${status}" at ${priorClosedAt} — idempotent no-op` };
738
+ return finishClose({ ok: true, path: p, status, closed_at: priorClosedAt, cause: priorCause, decision: "idempotent", reason: `already closed as "${status}" at ${priorClosedAt} — idempotent no-op` });
592
739
  }
593
740
  if (priorClosedStatus !== status) {
594
741
  // A DIFFERENT terminal status over an already-closed run — refused outright, the original
@@ -613,10 +760,10 @@ export function closeRun(cwd, slug, { status, cause = null } = {}) {
613
760
  if (afterSup.status !== status || !afterSup.closed_at || afterSup.closed_at === "~") {
614
761
  return { ok: false, path: p, status, reason: `wrote the superseding close but the ledger reads back status="${afterSup.status}" closed_at="${afterSup.closed_at}" — the write did not take` };
615
762
  }
616
- return {
763
+ return finishClose({
617
764
  ok: true, path: p, status, closed_at: afterSup.closed_at, cause: afterSup.close_cause ?? null,
618
765
  superseded: true, decision: "superseded", prior_cause: priorCause, prior_closed_at: priorClosedAt,
619
- };
766
+ });
620
767
  }
621
768
 
622
769
  const closedAt = new Date().toISOString();
@@ -628,7 +775,7 @@ export function closeRun(cwd, slug, { status, cause = null } = {}) {
628
775
  if (after.status !== status || !after.closed_at || after.closed_at === "~") {
629
776
  return { ok: false, path: p, status, reason: `wrote the close but the ledger reads back status="${after.status}" closed_at="${after.closed_at}" — the write did not take` };
630
777
  }
631
- return { ok: true, path: p, status, closed_at: after.closed_at, cause: after.close_cause ?? null, decision: "closed" };
778
+ return finishClose({ ok: true, path: p, status, closed_at: after.closed_at, cause: after.close_cause ?? null, decision: "closed" });
632
779
  }
633
780
 
634
781
  /**
@@ -658,7 +805,8 @@ export function writeActiveOrder(cwd, slug, orderPath) {
658
805
  /** The typed argv contract (see `./lib/argv.mjs`). */
659
806
  export const ARGV_SPEC = {
660
807
  usage: "harness.mjs probe resume --slug <slug> [--cwd <dir>] " +
661
- "[--require <phase> | --set-status <status> | --set-active-order <path> | --close <status> [--cause <text>]]",
808
+ "[--require <phase> | --set-status <status> | --set-active-order <path> | " +
809
+ "--close <status> | --close-arm <RunReturn.status> [--cause <text>]]",
662
810
  _: { arity: 0, max: 0, name: "(no positional operands)" },
663
811
  slug: { type: "str", required: true },
664
812
  cwd: { type: "path" },
@@ -667,6 +815,9 @@ export const ARGV_SPEC = {
667
815
  "set-active-order": { type: "str" },
668
816
  // The one call site that stamps a terminal status, its cause and closed_at together.
669
817
  close: { type: "enum", values: TERMINAL_STATUSES },
818
+ // Not `--close`: the caller (shapeup-run.js's closeIfTerminal) hands over a RunReturn arm, never
819
+ // a status it decided was terminal itself — RUN_RETURN_CLOSE/closeArm above make that call.
820
+ "close-arm": { type: "str" },
670
821
  cause: { type: "str" },
671
822
  };
672
823
 
@@ -681,13 +832,13 @@ export function cli(rawArgv) {
681
832
  const args = runArgs(ARGV_SPEC, rawArgv);
682
833
  const cwd = args.cwd || process.cwd();
683
834
 
684
- const ops = [args.require && "--require", args.setStatus && "--set-status", args.setActiveOrder && "--set-active-order", args.close && "--close"].filter(Boolean);
835
+ const ops = [args.require && "--require", args.setStatus && "--set-status", args.setActiveOrder && "--set-active-order", args.close && "--close", args.closeArm && "--close-arm"].filter(Boolean);
685
836
  if (ops.length > 1) {
686
837
  process.stderr.write(JSON.stringify({ error: "conflicting_flags", flags: ops, expected: "one operation per invocation" }) + "\n");
687
838
  process.exit(2);
688
839
  }
689
- if (args.cause !== undefined && !args.close) {
690
- process.stderr.write(JSON.stringify({ error: "conflicting_flags", flags: ["--cause"], expected: "--cause is only meaningful with --close" }) + "\n");
840
+ if (args.cause !== undefined && !args.close && !args.closeArm) {
841
+ process.stderr.write(JSON.stringify({ error: "conflicting_flags", flags: ["--cause"], expected: "--cause is only meaningful with --close or --close-arm" }) + "\n");
691
842
  process.exit(2);
692
843
  }
693
844
 
@@ -724,6 +875,17 @@ export function cli(rawArgv) {
724
875
  process.exit(r.ok ? 0 : 3);
725
876
  }
726
877
 
878
+ // The arm-derived close: the caller hands a RunReturn arm and this kernel decides — via
879
+ // RUN_RETURN_CLOSE/closeArm above — whether it is terminal and, if so, what status it closes as.
880
+ // Exit 0 for both a real close AND a correctly-declined non-terminal arm (`terminal:false`) —
881
+ // neither is an error the caller (shapeup-run.js's closeIfTerminal) should treat as failed; only
882
+ // an arm this map does not recognize at all, or a close `closeRun` itself refuses, exits non-zero.
883
+ if (args.closeArm) {
884
+ const r = closeArm(cwd, args.slug, args.closeArm, args.cause ?? null);
885
+ console.log(JSON.stringify(r));
886
+ process.exit(r.ok ? 0 : 3);
887
+ }
888
+
727
889
  console.log(JSON.stringify(deriveResumeState(cwd, args.slug)));
728
890
  process.exit(0);
729
891
  }
@@ -78,7 +78,7 @@ export function frontmatter(text) {
78
78
  */
79
79
  export function boardCensus(cwd, slug) {
80
80
  const dir = tasksDir(cwd, slug);
81
- const out = { total: 0, done: 0, unfinished: [] };
81
+ const out = { total: 0, done: 0, unfinished: [], anchors: {} };
82
82
  if (!existsSync(dir)) return out;
83
83
  for (const f of readdirSync(dir)) {
84
84
  if (!/^TASK-[\w.-]+\.md$/i.test(f)) continue;
@@ -86,6 +86,12 @@ export function boardCensus(cwd, slug) {
86
86
  const fm = frontmatter(body);
87
87
  const id = fm.id || f.replace(/\.md$/, "");
88
88
  out.total++;
89
+ // The COMMITTED anchor for this board id. `use_case_refs` is the tier-direction rule's own
90
+ // sanctioned direction (LOCAL names SHARED), and it is what the frozen report cites instead of
91
+ // the id — boards renumber per machine, use cases do not.
92
+ const ucs = String(fm.use_case_refs ?? "").replace(/^\[|\]$/g, "")
93
+ .split(",").map((x) => x.trim()).filter(Boolean);
94
+ out.anchors[id] = ucs;
89
95
  if (fm.status === "done") out.done++;
90
96
  else out.unfinished.push(id);
91
97
  }
@@ -93,6 +99,31 @@ export function boardCensus(cwd, slug) {
93
99
  return out;
94
100
  }
95
101
 
102
+ /**
103
+ * Replace every board id in free prose with its committed anchor.
104
+ *
105
+ * THE WRITE BOUNDARY, not the column, and that is the whole point. Board ids reached the frozen
106
+ * report three different ways — the unfinished-task callout, the covering-AC column's own prefix,
107
+ * and INSIDE acceptance-criterion prose a planner wrote ("given the seeded todos (TASK-006)"). The
108
+ * third is upstream free text, so a fix that only changes what the columns interpolate still
109
+ * commits a file the next run's spec-lint reds. Everything written into the committed report passes
110
+ * through here.
111
+ *
112
+ * An id with no resolvable use case becomes a neutral phrase rather than the id: the report loses a
113
+ * pointer that never resolved off this machine anyway, and keeps the sentence around it.
114
+ *
115
+ * @param {*} text - Any value destined for the committed report.
116
+ * @param {Record<string, string[]>} anchors - Board id → its `use_case_refs`.
117
+ * @returns {string} The text with every `TASK-…` replaced by a stable anchor.
118
+ */
119
+ export function deboard(text, anchors = {}) {
120
+ return String(text ?? "").replace(/\bTASK-[A-Za-z0-9][\w.-]*/g, (id) => {
121
+ const ucs = anchors[id];
122
+ if (ucs && ucs.length) return ucs.join("/");
123
+ return "a board task";
124
+ });
125
+ }
126
+
96
127
  /**
97
128
  * Per-scope T0 outcome, reduced from the trial ledger.
98
129
  *
@@ -194,7 +225,10 @@ export function buildReport(facts) {
194
225
  L.push("");
195
226
 
196
227
  if (board.unfinished.length) {
197
- L.push(`> **${board.unfinished.length} task(s) did not finish:** ${board.unfinished.join(", ")}.`,
228
+ // Anchored, never enumerated by board id: the ids renumber per machine, and a committed file
229
+ // carrying one reds the NEXT run of this pitch at L1b.
230
+ const unfinishedAnchors = [...new Set(board.unfinished.flatMap((id) => board.anchors?.[id] ?? []))];
231
+ L.push(`> **${board.unfinished.length} task(s) did not finish**${unfinishedAnchors.length ? ` — use cases: ${unfinishedAnchors.join(", ")}` : ""}.`,
198
232
  "> The verdict above grades what was built, not what was planned.", "");
199
233
  }
200
234
 
@@ -241,7 +275,9 @@ export function buildReport(facts) {
241
275
  const cell = (s) => String(s).replace(/\|/g, "\\|");
242
276
  L.push("| REQ | source | evidence | covering AC | criterion | T0 |", "|---|---|---|---|---|---|");
243
277
  for (const r of requirements.rows) {
244
- const ac = r.covering_acs.length ? `${r.covering_acs[0].task_id}: ${r.covering_acs[0].ac}${r.covering_acs.length > 1 ? ` (+${r.covering_acs.length - 1})` : ""}` : "—";
278
+ const ac = r.covering_acs.length
279
+ ? `${deboard(r.covering_acs[0].task_id, facts.board?.anchors)}: ${deboard(r.covering_acs[0].ac, facts.board?.anchors)}${r.covering_acs.length > 1 ? ` (+${r.covering_acs.length - 1})` : ""}`
280
+ : "—";
245
281
  const crit = r.criteria.length ? `${r.criteria[0].criterion}${r.criteria.length > 1 ? ` (+${r.criteria.length - 1})` : ""} → ${r.criteria.map((c) => c.verdict).join(",")}` : "—";
246
282
  const t0h = r.t0.length ? r.t0.map((h) => String(h).slice(0, 12)).join(", ") : "—";
247
283
  L.push(`| ${r.id} | ${cell(r.source || "—")} | ${r.evidence} | ${cell(ac)} | ${cell(crit)} | ${t0h} |`);
@@ -1215,7 +1215,7 @@
1215
1215
  }
1216
1216
  },
1217
1217
  "CommandResult": {
1218
- "description": "One executed command's outcome inside a T0 artifact — produced by actually running the command; no agent can fabricate it.",
1218
+ "description": "One executed command's outcome inside a T0 artifact — produced by actually running the command; no agent can fabricate it. It carries the EVIDENCE, not only the score: `exit` maps a command that never started onto the same 1 a real failure returns, so without `error` a refused or timed-out command and a broken build are the same record everywhere downstream, and without the output a verdict asserting `exit 1` is an assertion nobody can check.",
1219
1219
  "x-tier": "EMBEDDED",
1220
1220
  "type": "object",
1221
1221
  "properties": {
@@ -1223,10 +1223,23 @@
1223
1223
  "type": "string"
1224
1224
  },
1225
1225
  "exit": {
1226
- "type": "integer"
1226
+ "type": "integer",
1227
+ "description": "The process exit code, or 1 when the command never produced one. Not a discriminator on its own — read `error` to tell a crash from a failure."
1227
1228
  },
1228
1229
  "pass": {
1229
1230
  "type": "boolean"
1231
+ },
1232
+ "error": {
1233
+ "type": "string",
1234
+ "description": "Present ONLY when the command did not run to completion — a spawn failure, a maxBuffer overflow, or the 10-minute timeout. Its presence is the fact the ratchet grades as `crash` (tree restored, never counted as a reverted attempt); its absence means the command ran and the exit code is its own."
1235
+ },
1236
+ "stdout_tail": {
1237
+ "type": "string",
1238
+ "description": "The last 4000 characters of stdout, omitted when nothing was printed. Kept for passing commands too: a fixture that exits 0 having run zero tests is the false green this layer exists to catch. A truncated tail carries a leading `[truncated: kept the last N of M characters]` line, so a partial stream never reads as a complete one."
1239
+ },
1240
+ "stderr_tail": {
1241
+ "type": "string",
1242
+ "description": "The last 4000 characters of stderr, same bound and same truncation marker as `stdout_tail`."
1230
1243
  }
1231
1244
  }
1232
1245
  },