shapeup-sdlc 3.7.5 → 3.7.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.7.5",
4
+ "version": "3.7.7",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -13,7 +13,7 @@ Skills and commands are named short throughout this file; every one of them reso
13
13
  - **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
14
14
  - GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
15
15
  - Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
16
- - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
16
+ - **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
17
17
  - **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
18
18
  - The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
19
19
 
@@ -246,7 +246,10 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
246
246
  // not cover is a leg forging a SIBLING's result, which is a real hole with a different fix — the
247
247
  // glob form here cannot say "every result except this order's own", so closing it needs a
248
248
  // mechanism rather than one more entry on this list.
249
- const FROZEN_ATTESTATION = [`${local}/receipts/**`, `${local}/legs.jsonl`, `${local}/t0/verdicts/**`];
249
+ // The board index is ingest's projection of the task results, not the doer's bookkeeping — the
250
+ // task files are. The executor's contract already forbids editing it; the freeze makes that a
251
+ // denial rather than a rule, on the same terms as the attestation channels.
252
+ const FROZEN_ATTESTATION = [`${local}/receipts/**`, `${local}/legs.jsonl`, `${local}/t0/verdicts/**`, `${local}/tasks/_index.md`];
250
253
  switch (operation) {
251
254
  case "execute": case "fix": case "spike":
252
255
  // Build legs are the widest window on FROZEN_INTAKE, not an exemption from it: they are the
package/kernel/gate.mjs CHANGED
@@ -54,7 +54,7 @@ import { readFileSync, writeFileSync, appendFileSync, existsSync, mkdirSync } fr
54
54
  import { parseBoard } from "./reduce/board.mjs";
55
55
  import { join, dirname } from "node:path";
56
56
  import { runArgs } from "./lib/argv.mjs";
57
- import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir, tasksDir, hammerCensus } from "./lib/paths.mjs";
57
+ import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir, tasksDir, hammerCensus, readRunId } from "./lib/paths.mjs";
58
58
 
59
59
  export const GATE_IDS = ["L0", "L1a", "L1a.5", "L1b", "L2", "L3", "QA", "H", "L4", "COACH-1"];
60
60
 
@@ -465,7 +465,11 @@ export function cli(rawArgv) {
465
465
  // per-run ledger row — the same reasoning `resolveRunId` uses for "no run is active": absence is
466
466
  // the correct answer, not an error, so the write is skipped rather than guessing a location.
467
467
  if (args.slug) {
468
+ // THE ROW CARRIES THE RUN KEY. It did not, and the export stamped the current run's key onto
469
+ // every row it found — a prior run's sign-off became this run's in the one table that answers
470
+ // "was this ship signed off". Driven on a two-run fixture before it was fixed.
468
471
  appendGateLedger(cwd, args.slug, {
472
+ at: new Date().toISOString(), run_id: readRunId(cwd, args.slug),
469
473
  gate: r.gate, status: r.status, decision: r.decision ?? null,
470
474
  source: r.source ?? found.source, note: r.note ?? r.reason ?? null,
471
475
  round: args.round ?? null,
@@ -35,7 +35,7 @@ import { existsSync, readFileSync, readdirSync } from "node:fs";
35
35
  import { join, resolve } from "node:path";
36
36
  import { createHash } from "node:crypto";
37
37
  import { runArgs } from "../lib/argv.mjs";
38
- import { resultsDir, scopesDir } from "../lib/paths.mjs";
38
+ import { resultsDir, scopesDir, readRunId } from "../lib/paths.mjs";
39
39
 
40
40
  /** Longest `reason` reported. A deviation is prose written by a worker and can run to paragraphs. */
41
41
  const REASON_MAX = 400;
@@ -89,7 +89,7 @@ export function isScoped(cwd, slug) {
89
89
  * @returns {(string|null)} A reason phrased for an operator, or null when the file at `path` exists,
90
90
  * hashes to the cited `sha256`, and its own `overall` reads "green".
91
91
  */
92
- function unresolvedCitation(cwd, citation) {
92
+ function unresolvedCitation(cwd, citation, { round = null, runId = null } = {}) {
93
93
  const rel = typeof citation?.path === "string" ? citation.path : "";
94
94
  if (!rel) return "names no artifact path";
95
95
  let text;
@@ -109,6 +109,20 @@ function unresolvedCitation(cwd, citation) {
109
109
  try { body = JSON.parse(text); }
110
110
  catch { return `cites ${rel}, whose bytes match the hash but do not read as a T0 verdict`; }
111
111
  if (body?.overall !== "green") return `cites ${rel}, whose own verdict is "${body?.overall ?? "unknown"}", not green`;
112
+ // THE ARTIFACT HAS TO BE THE ONE THE CITATION SAYS IT IS. A re-hash proves the bytes are the
113
+ // file's; it says nothing about whose verdict the file holds. A PASS citing scope alpha's green
114
+ // artifact while declaring scope beta, or a prior round's, or a prior run's over the same slug,
115
+ // passed the digest check unremarked. The citation's own required `scope_id`, the round being
116
+ // judged and the run's key are compared to what the artifact records about itself.
117
+ if (typeof citation.scope_id === "string" && body?.scope_id && body.scope_id !== citation.scope_id) {
118
+ return `cites ${rel} for scope "${citation.scope_id}", but the artifact records scope "${body.scope_id}"`;
119
+ }
120
+ if (round != null && typeof body?.round === "number" && body.round !== round) {
121
+ return `cites ${rel}, a round ${body.round} artifact, as evidence for round ${round}`;
122
+ }
123
+ if (runId && body?.run_id && body.run_id !== runId) {
124
+ return `cites ${rel}, an artifact of run ${body.run_id}, as evidence for run ${runId}`;
125
+ }
112
126
  return null;
113
127
  }
114
128
 
@@ -143,15 +157,49 @@ function unresolvedCitation(cwd, citation) {
143
157
  * citation resolves, an unscoped spec, or a block with no PASS/FAIL in it (there is no judgement
144
158
  * to invalidate).
145
159
  */
146
- export function citationProblem(cwd, slug, verdict) {
160
+ /**
161
+ * Why a verdict cannot stand on its own criteria, or null when it can.
162
+ *
163
+ * `overall` is the judge's field, and nothing recomputed it from the criteria the judge graded: a
164
+ * PASS over a failing criterion, or over no criterion at all, validated and ingested, and the
165
+ * round loop branched on it. The evaluator's own first rule is that absence of evidence is a FAIL,
166
+ * so a PASS criterion with no evidence is no evidence either. Recomputed here, on ingest and on
167
+ * read alike: PASS means every graded criterion passed with evidence and at least one was graded;
168
+ * FAIL means at least one graded criterion failed.
169
+ *
170
+ * @param {object} verdict - The WorkResult's `verdict`.
171
+ * @returns {(string|null)} A reason phrased for an operator, or null.
172
+ */
173
+ export function verdictProblem(verdict) {
174
+ const overall = verdict?.overall;
175
+ if (overall !== "PASS" && overall !== "FAIL") return null;
176
+ const criteria = Array.isArray(verdict.criteria) ? verdict.criteria : [];
177
+ const fails = criteria.filter((c) => c?.verdict === "FAIL");
178
+ const passes = criteria.filter((c) => c?.verdict === "PASS");
179
+ const other = criteria.length - fails.length - passes.length;
180
+ if (overall === "PASS") {
181
+ if (criteria.length === 0) return "the PASS verdict grades no criterion at all — a PASS with no evidence is a claim";
182
+ if (fails.length) return `the verdict says PASS while ${fails.length} of its ${criteria.length} criteria read FAIL — overall is derived from the criteria, never declared over them`;
183
+ if (other) return `the verdict says PASS while ${other} of its criteria carry no PASS/FAIL verdict`;
184
+ const bare = passes.filter((c) => !(typeof c?.evidence === "string" && c.evidence.trim()));
185
+ if (bare.length) return `the PASS verdict has ${bare.length} criterion(s) marked PASS with no evidence — absence of evidence is a FAIL by the evaluator's own first rule`;
186
+ return null;
187
+ }
188
+ if (criteria.length === 0) return "the FAIL verdict grades no criterion at all — a FAIL must name what failed";
189
+ if (!fails.length) return `the verdict says FAIL while every one of its ${criteria.length} graded criteria reads PASS — a FAIL must cite the criterion it failed`;
190
+ return null;
191
+ }
192
+
193
+ export function citationProblem(cwd, slug, verdict, { round = null } = {}) {
147
194
  if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
148
195
  if (!isScoped(cwd, slug)) return null;
149
196
  if (!Array.isArray(verdict.t0_citations) || !verdict.t0_citations.length) {
150
197
  return `the ${verdict.overall} verdict cites no T0 artifact, and a verdict on a scoped spec must ` +
151
198
  "cite the T0 verdict it re-hashed (the order lists them under payload.t0_artifacts)";
152
199
  }
200
+ const runId = readRunId(cwd, slug);
153
201
  for (const citation of verdict.t0_citations) {
154
- const reason = unresolvedCitation(cwd, citation);
202
+ const reason = unresolvedCitation(cwd, citation, { round, runId });
155
203
  if (reason) return `the ${verdict.overall} verdict ${reason} — a T0 citation is re-hashed from disk, never taken on the handed word`;
156
204
  }
157
205
  return null;
@@ -185,7 +233,7 @@ export function evalVerdict(cwd, slug, round) {
185
233
  ? `the evaluator returned ${status || "no status"}: ${first}`
186
234
  : `status ${status || "unknown"} with no PASS/FAIL verdict`), status);
187
235
  }
188
- const problem = citationProblem(cwd, slug, v);
236
+ const problem = verdictProblem(v) || citationProblem(cwd, slug, v, { round });
189
237
  if (problem) return unfit(problem, status, overall);
190
238
  return {
191
239
  found: true,
@@ -27,7 +27,7 @@
27
27
  import { existsSync, readdirSync, readFileSync } from "node:fs";
28
28
  import { join, resolve } from "node:path";
29
29
  import { runArgs } from "../lib/argv.mjs";
30
- import { ordersDir, resultsDir, legLedger } from "../lib/paths.mjs";
30
+ import { ordersDir, resultsDir, legLedger, dispatchReceipts, readRunId } from "../lib/paths.mjs";
31
31
 
32
32
  /** `<scope>-r<round>-a<attempt>.json` — the only address a build order is written under. */
33
33
  const BUILD_ORDER = /^(.+)-r(\d+)-a(\d+)\.json$/;
@@ -80,6 +80,12 @@ export function legsOf(cwd, slug) {
80
80
  const ingested = new Set(readLegs(legLedger(cwd, slug))
81
81
  .filter((r) => r.ingested_at)
82
82
  .map((r) => String(r.order_id)));
83
+ // Dispatch receipts are the hook layer's record that a Skill was invoked against an order —
84
+ // scoped to this run, since receipts over one slug accumulate across runs and order ids repeat.
85
+ const runId = readRunId(cwd, slug);
86
+ const receipted = new Set(readLegs(dispatchReceipts(cwd, slug))
87
+ .filter((r) => r?.order_id && (!runId || !r.run_id || r.run_id === runId))
88
+ .map((r) => String(r.order_id)));
83
89
  const out = [];
84
90
  for (const f of (existsSync(oDir) ? readdirSync(oDir) : []).filter((x) => x.endsWith(".json")).sort()) {
85
91
  let orderId = null;
@@ -88,6 +94,7 @@ export function legsOf(cwd, slug) {
88
94
  order: join(oDir, f),
89
95
  name: f.replace(/\.json$/, ""),
90
96
  order_id: orderId,
97
+ has_receipt: orderId !== null && receipted.has(orderId),
91
98
  has_result: existsSync(join(rDir, f)),
92
99
  applied: orderId !== null && ingested.has(orderId),
93
100
  });
@@ -104,10 +111,25 @@ export function legsOf(cwd, slug) {
104
111
  */
105
112
  export function orderLegState(cwd, slug, name) {
106
113
  const o = legsOf(cwd, slug).find((x) => x.name === name);
107
- if (!o) return { closed: false, found: false, order: null, order_id: null, has_result: false, applied: false };
114
+ if (!o) return { closed: false, found: false, order: null, order_id: null, has_receipt: false, has_result: false, applied: false };
108
115
  return { closed: o.applied, found: true, ...o };
109
116
  }
110
117
 
118
+ /**
119
+ * Orders that were dispatched — a receipt says a Skill ran against them — and never answered:
120
+ * no WorkResult on disk. Measured on a live run: the first dispatch wrote its four artifacts and
121
+ * no envelope, every later phase read the artifacts, and the run closed `shipped` with the order
122
+ * still open by construction — the export's own dispatch row said `answered: false` and nothing
123
+ * upstream of it had noticed. Distinct from {@link openLegs}: that is work nobody applied; this is
124
+ * a leg that never came back.
125
+ * @param {string} cwd - Project root.
126
+ * @param {string} slug - Feature slug.
127
+ * @returns {{order:string, name:string, order_id:(string|null)}[]}
128
+ */
129
+ export function unansweredOrders(cwd, slug) {
130
+ return legsOf(cwd, slug).filter((o) => o.has_receipt && !o.has_result);
131
+ }
132
+
111
133
  /**
112
134
  * The results on disk that no leg row applied — finished work the board never saw.
113
135
  * @param {string} cwd - Project root.
@@ -171,8 +193,13 @@ export function cli(rawArgv) {
171
193
  // `--open`: every result nothing applied, across the whole run, any phase. Exit 0 when none.
172
194
  if (args.open) {
173
195
  const open = openLegs(cwd, args.slug);
174
- console.log(JSON.stringify({ closed: open.length === 0, open_total: open.length, open: open.map((o) => o.order), open_ids: open.map((o) => o.order_id) }));
175
- process.exit(open.length === 0 ? 0 : 1);
196
+ const unanswered = unansweredOrders(cwd, args.slug);
197
+ console.log(JSON.stringify({
198
+ closed: open.length === 0 && unanswered.length === 0,
199
+ open_total: open.length, open: open.map((o) => o.order), open_ids: open.map((o) => o.order_id),
200
+ unanswered_total: unanswered.length, unanswered: unanswered.map((o) => o.order), unanswered_ids: unanswered.map((o) => o.order_id),
201
+ }));
202
+ process.exit(open.length === 0 && unanswered.length === 0 ? 0 : 1);
176
203
  }
177
204
  // `--order <stem>`: one order by file stem (`analyze`, `alpha-r1-a1`). Exit 0 when its leg closed.
178
205
  if (args.order) {
@@ -62,8 +62,9 @@ import { runArgs } from "../lib/argv.mjs";
62
62
  import { splitFrontmatter, uncoerce } from "../lib/contract.mjs";
63
63
  import { globToRegExp } from "../verify/spec.mjs";
64
64
  import { parseBoard } from "../reduce/board.mjs";
65
- import { intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir, orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements, exportRunDir, lastRun, readRunId, tasksDir } from "../lib/paths.mjs";
65
+ import { intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir, orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements, exportRunDir, lastRun, readRunId, tasksDir, gates, verdictsDir, roundBuildDir } from "../lib/paths.mjs";
66
66
  import { evalVerdict } from "./eval.mjs";
67
+ import { deriveRounds } from "./rounds.mjs";
67
68
  import { collectRun, writeRun } from "../report/export.mjs";
68
69
 
69
70
  /** The run-state values `references/protocol.md` (Part 4 — State) defines. A typo'd status is a rejection,
@@ -560,7 +561,83 @@ export function setRunStatus(cwd, slug, status) {
560
561
  * already normalized prose (newlines collapsed, truncated) — this function only `uncoerce`s it.
561
562
  * @returns {string} The rewritten text.
562
563
  */
563
- function writeCloseLines(body, { status, closedAt, cause }) {
564
+ /**
565
+ * What the ledger's front matter and tables should say at a close, derived from the run's own
566
+ * records — the same ones the export reads. The counters (`final_verdict`, `rounds_used`) and the
567
+ * Rounds/Decisions tables used to be hand-shaped placeholders nothing filled: a run closed
568
+ * `shipped` beside `final_verdict: ~`, `rounds_used: 0`, an empty Decisions table and seven rows in
569
+ * `gates.jsonl`. A reader who trusted the counters concluded no round ran. Measured live.
570
+ *
571
+ * @param {string} cwd - Project root.
572
+ * @param {string} slug - Feature slug.
573
+ * @returns {{final_verdict:string, rounds_used:(number|null), decisions:object[], roundRows:string[]}}
574
+ */
575
+ export function deriveLedgerFacts(cwd, slug) {
576
+ const readJson = (p) => { try { return JSON.parse(readFileSync(p, "utf8")); } catch { return null; } };
577
+ const evalRows = [];
578
+ const rDir = resultsDir(cwd, slug);
579
+ if (existsSync(rDir)) {
580
+ for (const f of readdirSync(rDir)) {
581
+ const m = f.match(/^evaluate-r(\d+)\.json$/);
582
+ if (!m) continue;
583
+ const r = readJson(join(rDir, f));
584
+ const overall = r?.verdict?.overall;
585
+ if (overall) evalRows.push({ round: Number(m[1]), overall: String(overall), criteria: Array.isArray(r.verdict.criteria) ? r.verdict.criteria : [] });
586
+ }
587
+ }
588
+ evalRows.sort((a, b) => a.round - b.round);
589
+ const finalVerdict = evalRows.length ? evalRows[evalRows.length - 1].overall : "not-evaluated";
590
+ const decisions = [];
591
+ const gp = gates(cwd, slug);
592
+ if (existsSync(gp)) {
593
+ for (const line of readFileSync(gp, "utf8").split("\n")) {
594
+ if (!line.trim()) continue;
595
+ try { decisions.push(JSON.parse(line)); } catch { /* a torn line proves nothing */ }
596
+ }
597
+ }
598
+ const t0ByRound = {};
599
+ const vDir = verdictsDir(cwd, slug);
600
+ if (existsSync(vDir)) {
601
+ for (const f of readdirSync(vDir).filter((x) => x.endsWith(".json")).sort()) {
602
+ const v = readJson(join(vDir, f));
603
+ if (typeof v?.round !== "number") continue;
604
+ const t = (t0ByRound[v.round] ||= { green: 0, red: 0 });
605
+ if (v.overall === "green") t.green++; else t.red++;
606
+ }
607
+ }
608
+ const gateByRound = {};
609
+ const bDir = roundBuildDir(cwd, slug);
610
+ if (existsSync(bDir)) {
611
+ for (const f of readdirSync(bDir).sort()) {
612
+ const m = f.match(/^r(\d+)-t\d+\.json$/);
613
+ if (m) gateByRound[Number(m[1])] = readJson(join(bDir, f))?.overall ?? "?";
614
+ }
615
+ }
616
+ const roundNums = [...new Set([...Object.keys(t0ByRound), ...Object.keys(gateByRound), ...evalRows.map((e) => e.round)].map(Number))].sort((a, b) => a - b);
617
+ const roundRows = [];
618
+ for (const r of roundNums) {
619
+ const t0 = t0ByRound[r], bg = gateByRound[r], ev = evalRows.find((e) => e.round === r);
620
+ roundRows.push(`| Build | ${r} | ${t0 ? `T0 ${t0.green} green / ${t0.red} red` : "no T0 verdict"} | — | round build gate: ${bg ?? "not run"} |`);
621
+ if (ev) {
622
+ const p = ev.criteria.filter((c) => c?.verdict === "PASS").length, f = ev.criteria.filter((c) => c?.verdict === "FAIL").length;
623
+ roundRows.push(`| Eval | ${r} | ${ev.overall} | — | ${p} PASS / ${f} FAIL criteria |`);
624
+ }
625
+ }
626
+ let roundsUsed = null;
627
+ try { roundsUsed = deriveRounds(cwd, slug, null)?.rounds_used ?? null; } catch { roundsUsed = null; }
628
+ if (roundsUsed == null && roundNums.length) roundsUsed = roundNums[roundNums.length - 1];
629
+ return { final_verdict: finalVerdict, rounds_used: roundsUsed, decisions, roundRows };
630
+ }
631
+
632
+ /** Replace the rows of one markdown table (identified by its header line) with `rows`, keeping the header, the separator and any row the caller marks as kept. */
633
+ function rewriteTable(body, headerRe, rows, keepRow = null) {
634
+ return body.replace(headerRe, (m, header, sep, oldRows) => {
635
+ const kept = keepRow && oldRows ? oldRows.split("\n").filter((l) => l.startsWith(keepRow)) : [];
636
+ return header + sep + [...kept, ...rows].map((l) => `${l}\n`).join("");
637
+ });
638
+ }
639
+
640
+ function writeCloseLines(body, { status, closedAt, cause, derived = null }) {
564
641
  const causeLine = `close_cause: ${uncoerce(cause || null)}`;
565
642
  const closedStatusLine = `closed_status: ${status}`;
566
643
  let out = body
@@ -574,6 +651,15 @@ function writeCloseLines(body, { status, closedAt, cause }) {
574
651
  out = /^close_cause:.*$/m.test(out)
575
652
  ? out.replace(/^close_cause:.*$/m, causeLine)
576
653
  : out.replace(/^closed_at:.*$/m, (m) => `${m}\n${causeLine}`);
654
+ if (derived) {
655
+ // THE COUNTERS AND TABLES ARE DERIVED, IN THE SAME PASS AS THE CLOSE LINE. Written from the
656
+ // run's own records so they cannot disagree with the export that reads the same records.
657
+ if (/^final_verdict:.*$/m.test(out)) out = out.replace(/^final_verdict:.*$/m, `final_verdict: ${derived.final_verdict}`);
658
+ if (typeof derived.rounds_used === "number" && /^rounds_used:.*$/m.test(out)) out = out.replace(/^rounds_used:.*$/m, `rounds_used: ${derived.rounds_used}`);
659
+ out = rewriteTable(out, /(\| Phase \| Round \| Result \| Duration \| Notes \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, derived.roundRows, "| Init");
660
+ const decisionRows = derived.decisions.map((g) => `| ${g.gate ?? "?"} | ${g.decision ?? g.status ?? "?"} | ${g.source ?? "?"} | ${String(g.note ?? "").replace(/\|/g, "\\|").replace(/\s+/g, " ").slice(0, 160)} |`);
661
+ out = rewriteTable(out, /(\| Gate \| Decision \| Source \| Note \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, decisionRows);
662
+ }
577
663
  return out;
578
664
  }
579
665
 
@@ -696,6 +782,9 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
696
782
  // The Ship phase (`shapeup-run.js`) already exports a shipped run itself, before this call ever
697
783
  // runs — Stage 2 adds the endings that wrote nothing, and leaves that path untouched.
698
784
  const shouldExport = withExport && status !== "shipped";
785
+ // Derived once, from the run's own records, and written with the close line (see deriveLedgerFacts).
786
+ let derived = null;
787
+ try { derived = deriveLedgerFacts(cwd, slug); } catch { derived = null; }
699
788
 
700
789
  /**
701
790
  * Everything a close owes the checkout once the ledger line is written: export the run's
@@ -778,7 +867,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
778
867
  // new line rather than lost, and the return says so explicitly.
779
868
  const closedAt = new Date().toISOString();
780
869
  const foldedCause = `${normCause || "no reason recorded"} — supersedes an earlier close recorded ${priorClosedAt} (cause: ${JSON.stringify(priorCause)})`.slice(0, 4000);
781
- body = writeCloseLines(body, { status, closedAt, cause: foldedCause });
870
+ body = writeCloseLines(body, { status, closedAt, cause: foldedCause, derived });
782
871
  try { writeFileSync(p, body); } catch (e) {
783
872
  return { ok: false, path: p, status, reason: `could not write the ledger: ${e.message}` };
784
873
  }
@@ -793,7 +882,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
793
882
  }
794
883
 
795
884
  const closedAt = new Date().toISOString();
796
- body = writeCloseLines(body, { status, closedAt, cause: normCause });
885
+ body = writeCloseLines(body, { status, closedAt, cause: normCause, derived });
797
886
  try { writeFileSync(p, body); } catch (e) {
798
887
  return { ok: false, path: p, status, reason: `could not write the ledger: ${e.message}` };
799
888
  }
@@ -22,7 +22,7 @@
22
22
 
23
23
  import { existsSync, readdirSync, readFileSync } from "node:fs";
24
24
  import { join } from "node:path";
25
- import { ordersDir, verdictsDir, roundBuildDir, resultsDir } from "../lib/paths.mjs";
25
+ import { ordersDir, verdictsDir, roundBuildDir, resultsDir, readRunId } from "../lib/paths.mjs";
26
26
 
27
27
  /** Parse a JSON file, returning null rather than throwing — every reader here is best-effort. */
28
28
  function readJson(p) {
@@ -59,12 +59,22 @@ const maxOf = (nums) => (nums.length ? Math.max(...nums) : null);
59
59
  * @returns {{rounds_used:*, rounds_judged:(number|null)}} Both counts.
60
60
  */
61
61
  export function deriveRounds(cwd, slug, fallback) {
62
+ // THIS RUN'S RECORDS ONLY. Orders, verdicts and build gates over one slug accumulate across runs
63
+ // and carry their run key; walking the directories unfiltered handed a run that had dispatched
64
+ // nothing a prior run's round count — into the committed report. A record with no key at all was
65
+ // written before the key existed and is kept; one with a different key is another run's.
66
+ const runId = readRunId(cwd, slug);
67
+ const mine = (rec) => !runId || !rec?.run_id || rec.run_id === runId;
62
68
  const orderRounds = [];
63
69
  const oDir = ordersDir(cwd, slug);
70
+ const orderOf = {};
64
71
  if (existsSync(oDir)) {
65
72
  for (const f of readdirSync(oDir)) {
66
73
  if (!f.endsWith(".json")) continue;
67
- const r = orderRound(readJson(join(oDir, f))?.order_id);
74
+ const o = readJson(join(oDir, f));
75
+ orderOf[f] = o;
76
+ if (!mine(o)) continue;
77
+ const r = orderRound(o?.order_id);
68
78
  if (r !== null) orderRounds.push(r);
69
79
  }
70
80
  }
@@ -74,7 +84,7 @@ export function deriveRounds(cwd, slug, fallback) {
74
84
  if (existsSync(vDir)) {
75
85
  for (const f of readdirSync(vDir).filter((x) => x.endsWith(".json"))) {
76
86
  const v = readJson(join(vDir, f));
77
- if (typeof v?.round === "number") verdictRounds.push(v.round);
87
+ if (mine(v) && typeof v?.round === "number") verdictRounds.push(v.round);
78
88
  }
79
89
  }
80
90
 
@@ -83,7 +93,7 @@ export function deriveRounds(cwd, slug, fallback) {
83
93
  if (existsSync(bDir)) {
84
94
  for (const f of readdirSync(bDir)) {
85
95
  const m = f.match(/^r(\d+)-t\d+\.json$/);
86
- if (m) buildGateRounds.push(Number(m[1]));
96
+ if (m && mine(readJson(join(bDir, f)))) buildGateRounds.push(Number(m[1]));
87
97
  }
88
98
  }
89
99
 
@@ -92,7 +102,8 @@ export function deriveRounds(cwd, slug, fallback) {
92
102
  if (existsSync(rDir)) {
93
103
  for (const f of readdirSync(rDir)) {
94
104
  const m = f.match(/^evaluate-r(\d+)\.json$/);
95
- if (m) evalRounds.push(Number(m[1]));
105
+ // A WorkResult carries no run key; it reaches one through its order of the same name.
106
+ if (m && mine(orderOf[f] ?? readJson(join(oDir, f)))) evalRounds.push(Number(m[1]));
96
107
  }
97
108
  }
98
109
 
@@ -37,7 +37,7 @@ import { fileURLToPath } from "node:url";
37
37
  import { validate } from "../verify/envelope.mjs";
38
38
  import { runArgs } from "../lib/argv.mjs";
39
39
  import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId } from "../lib/paths.mjs";
40
- import { citationProblem } from "../probe/eval.mjs";
40
+ import { citationProblem, verdictProblem } from "../probe/eval.mjs";
41
41
 
42
42
  const HERE = dirname(fileURLToPath(import.meta.url));
43
43
  const RESULT_SCHEMA = JSON.parse(readFileSync(resolve(HERE, "../schemas/work-result.schema.json"), "utf8"));
@@ -187,11 +187,48 @@ export function setCheckbox(body, ac, checked) {
187
187
  * @param {boolean} done - When true, rewrite the row's status emoji/word to done; false is a no-op.
188
188
  * @returns {string} The board text with the matching row updated (unchanged when no row matches).
189
189
  */
190
- export function updateBoardRow(indexBody, taskId, done) {
190
+ /**
191
+ * Is this index line the row FOR `taskId` — its first cell — rather than a row that merely
192
+ * mentions it? A task id appears in other rows' `Depends On` column, and matching by substring
193
+ * anywhere in the line flipped every dependent of a finished task to done along with it. Measured
194
+ * on a live run: the executor reported one task `skipped`, its task file still said `ready`, and
195
+ * the index showed it ✅ because the task it depended on had just been ticked. The census read the
196
+ * index.
197
+ * @param {string} line - One line of `tasks/_index.md`.
198
+ * @param {string} taskId - `TASK-NNN`.
199
+ * @returns {boolean}
200
+ */
201
+ export function rowIs(line, taskId) {
202
+ const id = taskId.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
203
+ return new RegExp(`^\\|\\s*(?:\\[\\[)?${id}(?:\\\\\\|${id}\\]\\])?\\s*\\|`).test(line);
204
+ }
205
+
206
+ /**
207
+ * The status cell for one TaskResult status, or null when the index should not change.
208
+ * @param {(string|boolean)} status - A TaskResult status (`true` is accepted as `done` for callers written before skipped existed).
209
+ * @returns {({icon:string, word:string}|null)} The icon and word the index row should carry.
210
+ */
211
+ export function boardCell(status) {
212
+ if (status === true || status === "done") return { icon: "✅", word: "done" };
213
+ if (status === "skipped") return { icon: "⏭", word: "skipped" };
214
+ if (status === "partial" || status === "failed") return { icon: "🔄", word: "in-progress" };
215
+ return null;
216
+ }
217
+
218
+ /**
219
+ * Rewrite one task's status cell in `tasks/_index.md` — the row whose id cell is `taskId`, never a
220
+ * row that merely mentions it in `Depends On`.
221
+ * @param {string} indexBody - The index file's text.
222
+ * @param {string} taskId - `TASK-NNN`.
223
+ * @param {(string|boolean)} status - A TaskResult status; statuses the index does not render leave it unchanged.
224
+ * @returns {string} The rewritten index text.
225
+ */
226
+ export function updateBoardRow(indexBody, taskId, status) {
227
+ const cell = boardCell(status);
228
+ if (!cell) return indexBody;
191
229
  return indexBody.split(/\r?\n/).map((line) => {
192
- if (!line.includes(taskId) || !line.includes("|")) return line;
193
- if (done) return line.replace(/⬜|🔄|⏳|🚫/g, "✅").replace(/\b(ready|in-progress|blocked)\b/gi, "done");
194
- return line;
230
+ if (!rowIs(line, taskId)) return line;
231
+ return line.replace(/⬜|🔄|⏳|🚫|✅|⏭/g, cell.icon).replace(/\b(ready|in-progress|blocked|done|skipped)\b/gi, cell.word);
195
232
  }).join("\n");
196
233
  }
197
234
 
@@ -240,14 +277,18 @@ function applyResultLocked(result, { cwd, slug }) {
240
277
  body = setFrontmatter(body, "completed_at", today());
241
278
  } else if (tr.status === "partial" || tr.status === "failed") {
242
279
  body = setFrontmatter(body, "status", "in-progress");
280
+ } else if (tr.status === "skipped") {
281
+ // Rendered as what it is. A skipped task used to leave its file at `ready` and, through the
282
+ // substring match above, could show ✅ on the index — the census read the index.
283
+ body = setFrontmatter(body, "status", "skipped");
243
284
  }
244
285
  // Execution Log (append; the checkbox list must never disagree with it).
245
286
  const logLines = (tr.ac_results || []).map((a) => `- ${a.ac}: ${a.result}${a.evidence ? ` (${a.evidence})` : ""}`).join("\n");
246
287
  body += `\n\n## Execution Log — ${today()} (${result.order_id})\n- executor: ${result.worker || "task-executor"} via ingest-result\n- status: ${tr.status}\n${logLines}${tr.notes ? `\n- notes: ${tr.notes}` : ""}\n`;
247
288
  writeFileSync(path, body);
248
289
  summary.tasks_updated.push(tr.task_id);
249
- if (tr.status === "done" && existsSync(boardIndex)) {
250
- writeFileSync(boardIndex, updateBoardRow(readFileSync(boardIndex, "utf8"), tr.task_id, true));
290
+ if (existsSync(boardIndex) && boardCell(tr.status)) {
291
+ writeFileSync(boardIndex, updateBoardRow(readFileSync(boardIndex, "utf8"), tr.task_id, tr.status));
251
292
  }
252
293
  }
253
294
 
@@ -271,7 +312,7 @@ function applyResultLocked(result, { cwd, slug }) {
271
312
  summary.unblocked.push(t.id);
272
313
  if (existsSync(boardIndex)) {
273
314
  const idx = readFileSync(boardIndex, "utf8").split(/\r?\n/).map((line) =>
274
- line.includes(t.id) && line.includes("|")
315
+ rowIs(line, t.id)
275
316
  ? line.replace(/🚫|⏳/g, "⬜").replace(/\bblocked\b/gi, "ready")
276
317
  : line).join("\n");
277
318
  writeFileSync(boardIndex, idx);
@@ -656,7 +697,8 @@ export async function cli(rawArgv) {
656
697
  // on (see `citationProblem`). `probe eval` refuses it to the round loop; refusing it here as well
657
698
  // keeps the verdict ledger from recording a verdict the loop will never branch on.
658
699
  if (result.verdict) {
659
- const problem = citationProblem(cwd, String(result.order_id).split("/")[0], result.verdict);
700
+ const evalRound = Number((String(result.order_id).match(/-r(\d+)$/) || [])[1]) || null;
701
+ const problem = verdictProblem(result.verdict) || citationProblem(cwd, String(result.order_id).split("/")[0], result.verdict, { round: evalRound });
660
702
  if (problem) {
661
703
  console.error(`ingest-result: result refused — ${problem}.`);
662
704
  console.error(` The round stays open: re-dispatch the evaluator against its order, which lists`);
@@ -329,7 +329,13 @@ export function buildReport(facts) {
329
329
  "*Run state (board, orders, results, T0 artifacts, evaluation and QA reports) stays in the",
330
330
  "gitignored local tier (ADR-0001). This report",
331
331
  "is the frozen conclusion of it.*", "");
332
- return L.join("\n");
332
+ // THE WHOLE REPORT, ONCE. Board ids were anchored in three places and leaked through a fourth: a
333
+ // discovery-ledger entry copied verbatim into "Discovered, not built" carried `[TASK-005]`, this
334
+ // kernel wrote it straight to the committed tier past the hook that guards the model's edits, and
335
+ // the next run's L1b lint red'd the file the previous run had frozen. A committed report may not
336
+ // name a board id anywhere, so the rule is applied to the finished text rather than section by
337
+ // section.
338
+ return deboard(L.join("\n"), board.anchors);
333
339
  }
334
340
 
335
341
  /**
@@ -173,12 +173,14 @@ function criterionRows(dir, runId, t) {
173
173
  * (see `appendGateLedger`), so this is a pass-through with a stamped `run_id` fallback rather than
174
174
  * a re-derivation: two readers of "what did this gate decide" must not compute the answer twice.
175
175
  * @param {object} g - One parsed line of `gates.jsonl`.
176
- * @param {(string|null)} runId - Run key for a row written before it carried its own.
176
+ * @param {(string|null)} runId - The run being exported (unused for attribution — the row's own key is the only one exported).
177
177
  * @returns {object} A flat `gate_decision` row.
178
178
  */
179
179
  function gateDecisionRow(g, runId) {
180
180
  return {
181
- run_id: g?.run_id ?? runId ?? null,
181
+ // The row's own key, never the current run's stamped on: a row that carries no key was written
182
+ // before the ledger did, and is exported as unattributed rather than claimed.
183
+ run_id: g?.run_id ?? null,
182
184
  gate: g?.gate ?? null,
183
185
  decision: g?.decision ?? null,
184
186
  status: g?.status ?? null,
@@ -260,7 +262,9 @@ export function collectRun(cwd, slug) {
260
262
  criterion_verdict: criterionRows(evaluationDir(cwd, slug), runId, t),
261
263
  hook_decision,
262
264
  // The decision that crossed each gate, and the round build gate's own artifact.
263
- gate_decision: readJsonl(gatesPath(cwd, slug), t).map((g) => gateDecisionRow(g, runId)),
265
+ // Scoped to the run, like hook_decision one line up: gate rows over one slug accumulate across
266
+ // runs, and a prior run's L4 exported under this run's key is a fabricated sign-off.
267
+ gate_decision: readJsonl(gatesPath(cwd, slug), t).filter((g) => runId && g?.run_id === runId).map((g) => gateDecisionRow(g, runId)),
264
268
  build_gate: readJsonDir(roundBuildDir(cwd, slug), t).map((a) => buildGateRow(a, runId)),
265
269
  leg: readJsonl(legLedger(cwd, slug), t)
266
270
  .filter((r) => !runId || !r?.run_id || r.run_id === runId)
@@ -40,7 +40,7 @@ import { resolve, join, dirname, relative, isAbsolute } from "node:path";
40
40
  import { readBoard } from "../compile.mjs";
41
41
  import { runArgs } from "../lib/argv.mjs";
42
42
  import { sharedRoot, traceDir, relLocal } from "../lib/paths.mjs";
43
- import { readContract, unreadableReason, LEGACY_LAYOUT, WIRING_MAP, PROJECT_PROFILE } from "../lib/contract.mjs";
43
+ import { readContract, unreadableReason, LEGACY_LAYOUT, WIRING_MAP, PROJECT_PROFILE, reqId } from "../lib/contract.mjs";
44
44
 
45
45
  // --- requirements.md registry parser -----------------------------------------
46
46
  // A committed markdown table: | REQ-id | clause (verbatim) | source | status | note |
@@ -87,7 +87,10 @@ export function coveredReqIds(board) {
87
87
  for (const task of board) {
88
88
  for (const ac of task.acceptance_criteria || []) {
89
89
  const covers = typeof ac === "object" && Array.isArray(ac.covers) ? ac.covers : [];
90
- for (const id of covers) if (/^REQ-\d+$/.test(id)) covered.add(id);
90
+ // ONE KEY SPACE. `R-2`, `[[REQ-5]]` and `req-4` are the same clause spelled three ways, and the
91
+ // sibling rule accepts all of them; testing the raw string here counted every one as nothing, so
92
+ // an author told at L1b to cover a requirement with an AC — which they had — stayed red.
93
+ for (const raw of covers) { const id = reqId(raw); if (/^REQ-\d+$/.test(id)) covered.add(id); }
91
94
  }
92
95
  }
93
96
  return covered;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.7.5",
3
+ "version": "3.7.7",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -482,6 +482,7 @@ const ORDERLEG = {
482
482
  closed: { type: "boolean" },
483
483
  found: { type: "boolean" },
484
484
  order: nullable("string"),
485
+ has_receipt: { type: "boolean" },
485
486
  has_result: { type: "boolean" },
486
487
  applied: { type: "boolean" },
487
488
  },
@@ -936,7 +937,18 @@ async function requireLeg(gate, phaseKey, phaseName, orderStem = phaseKey) {
936
937
  const ask = () => query(`probe leg --slug ${slug} --order "${orderStem}"`, ORDERLEG, phaseName, `legcheck:${orderStem}`);
937
938
  let leg = await ask();
938
939
  if (!leg || !leg.found) { log(`${gate} — could not ask the leg ledger about "${phaseKey}" (probe returned ${leg ? "no order" : "nothing"}); proceeding on the artifact alone.`); return null; }
939
- if (!leg.has_result || leg.applied) return null;
940
+ if (!leg.has_result) {
941
+ // The artifact is on disk and the envelope never came back. Measured live: the first dispatch of
942
+ // a run wrote its four artifacts and no WorkResult, every later phase read the artifacts, and
943
+ // the run closed `shipped` over an order still open by construction. The phase stands on its
944
+ // artifact — that is what the post-condition checks — but the order is named at the close, and
945
+ // the next run over this slug will find it unanswered rather than be surprised by it.
946
+ log(`${gate} — "${phaseKey}" ${leg.has_receipt ? "was dispatched (receipt on disk) and" : "has an order but no receipt, and"} never answered: no WorkResult at all. ` +
947
+ `Its artifacts are on disk and the phase proceeds on them; the order stays unanswered and is named at the close.`);
948
+ if (leg.order && !unansweredOrders.includes(leg.order)) unansweredOrders.push(leg.order);
949
+ return null;
950
+ }
951
+ if (leg.applied) return null;
940
952
  log(`${gate} — "${phaseKey}" came back with a result nothing applied (no leg row). Ingesting it here: ${leg.order}.`);
941
953
  await advisory(`reduce ingest --order "${leg.order}"`, phaseName, `late-ingest:${phaseKey}`);
942
954
  leg = await ask();
@@ -1084,6 +1096,9 @@ async function requireLaunchRecord() {
1084
1096
  // it warns and continues — and the warning travels in the RunReturn, because a headless stdout
1085
1097
  // carries only the final message and a diagnostic on a channel nobody reads is not a diagnostic.
1086
1098
  const stateWarnings = [];
1099
+ // Orders a receipt says were dispatched and no WorkResult ever answered — a leg that did its craft
1100
+ // and never came back. Named at every close, whatever the close; never folded into "complete".
1101
+ const unansweredOrders = [];
1087
1102
  async function setRunStatus(status, phaseName) {
1088
1103
  const r = await cmd(`probe resume --slug ${slug} --set-status ${status}`, phaseName, `status:${status}`);
1089
1104
  if (!r.ok) {
@@ -1143,7 +1158,8 @@ async function closeIfTerminal(ret) {
1143
1158
  // `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
1144
1159
  // a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
1145
1160
  // stay silent for it exactly as they did when this file's own guard returned early.
1146
- const r = await cmd(`probe resume --slug ${slug} --close-arm ${ret.status} --cause "${causeArg(cause)}"`, "Ship", `close:${ret.status}`);
1161
+ const unanswered = unansweredOrders.length ? ` unanswered_orders=${unansweredOrders.length}` : "";
1162
+ const r = await cmd(`probe resume --slug ${slug} --close-arm ${ret.status} --cause "${causeArg(cause + unanswered)}"`, "Ship", `close:${ret.status}`);
1147
1163
  if (!r.ok) {
1148
1164
  const why = (r.detail || `exit ${r.exit_code}`).trim();
1149
1165
  log(`RUN STATE — close(${ret.status}) did not take: ${why}. This return's own status and reason still ` +