shapeup-sdlc 3.7.5 → 3.7.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +1 -1
- package/kernel/compile.mjs +4 -1
- package/kernel/gate.mjs +5 -1
- package/kernel/probe/eval.mjs +53 -5
- package/kernel/probe/leg.mjs +31 -4
- package/kernel/probe/resume.mjs +93 -4
- package/kernel/probe/rounds.mjs +16 -5
- package/kernel/reduce/ingest.mjs +51 -9
- package/kernel/reduce/ship.mjs +7 -1
- package/kernel/report/export.mjs +7 -3
- package/kernel/verify/trace.mjs +5 -2
- package/package.json +1 -1
- package/skills/tech-lead/workflows/shapeup-run.js +18 -2
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.7.
|
|
4
|
+
"version": "3.7.7",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -13,7 +13,7 @@ Skills and commands are named short throughout this file; every one of them reso
|
|
|
13
13
|
- **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
|
|
14
14
|
- GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
|
|
15
15
|
- Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
|
|
16
|
-
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
|
|
16
|
+
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
|
|
17
17
|
- **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
|
|
18
18
|
- The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
|
|
19
19
|
|
package/kernel/compile.mjs
CHANGED
|
@@ -246,7 +246,10 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
|
|
|
246
246
|
// not cover is a leg forging a SIBLING's result, which is a real hole with a different fix — the
|
|
247
247
|
// glob form here cannot say "every result except this order's own", so closing it needs a
|
|
248
248
|
// mechanism rather than one more entry on this list.
|
|
249
|
-
|
|
249
|
+
// The board index is ingest's projection of the task results, not the doer's bookkeeping — the
|
|
250
|
+
// task files are. The executor's contract already forbids editing it; the freeze makes that a
|
|
251
|
+
// denial rather than a rule, on the same terms as the attestation channels.
|
|
252
|
+
const FROZEN_ATTESTATION = [`${local}/receipts/**`, `${local}/legs.jsonl`, `${local}/t0/verdicts/**`, `${local}/tasks/_index.md`];
|
|
250
253
|
switch (operation) {
|
|
251
254
|
case "execute": case "fix": case "spike":
|
|
252
255
|
// Build legs are the widest window on FROZEN_INTAKE, not an exemption from it: they are the
|
package/kernel/gate.mjs
CHANGED
|
@@ -54,7 +54,7 @@ import { readFileSync, writeFileSync, appendFileSync, existsSync, mkdirSync } fr
|
|
|
54
54
|
import { parseBoard } from "./reduce/board.mjs";
|
|
55
55
|
import { join, dirname } from "node:path";
|
|
56
56
|
import { runArgs } from "./lib/argv.mjs";
|
|
57
|
-
import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir, tasksDir, hammerCensus } from "./lib/paths.mjs";
|
|
57
|
+
import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir, tasksDir, hammerCensus, readRunId } from "./lib/paths.mjs";
|
|
58
58
|
|
|
59
59
|
export const GATE_IDS = ["L0", "L1a", "L1a.5", "L1b", "L2", "L3", "QA", "H", "L4", "COACH-1"];
|
|
60
60
|
|
|
@@ -465,7 +465,11 @@ export function cli(rawArgv) {
|
|
|
465
465
|
// per-run ledger row — the same reasoning `resolveRunId` uses for "no run is active": absence is
|
|
466
466
|
// the correct answer, not an error, so the write is skipped rather than guessing a location.
|
|
467
467
|
if (args.slug) {
|
|
468
|
+
// THE ROW CARRIES THE RUN KEY. It did not, and the export stamped the current run's key onto
|
|
469
|
+
// every row it found — a prior run's sign-off became this run's in the one table that answers
|
|
470
|
+
// "was this ship signed off". Driven on a two-run fixture before it was fixed.
|
|
468
471
|
appendGateLedger(cwd, args.slug, {
|
|
472
|
+
at: new Date().toISOString(), run_id: readRunId(cwd, args.slug),
|
|
469
473
|
gate: r.gate, status: r.status, decision: r.decision ?? null,
|
|
470
474
|
source: r.source ?? found.source, note: r.note ?? r.reason ?? null,
|
|
471
475
|
round: args.round ?? null,
|
package/kernel/probe/eval.mjs
CHANGED
|
@@ -35,7 +35,7 @@ import { existsSync, readFileSync, readdirSync } from "node:fs";
|
|
|
35
35
|
import { join, resolve } from "node:path";
|
|
36
36
|
import { createHash } from "node:crypto";
|
|
37
37
|
import { runArgs } from "../lib/argv.mjs";
|
|
38
|
-
import { resultsDir, scopesDir } from "../lib/paths.mjs";
|
|
38
|
+
import { resultsDir, scopesDir, readRunId } from "../lib/paths.mjs";
|
|
39
39
|
|
|
40
40
|
/** Longest `reason` reported. A deviation is prose written by a worker and can run to paragraphs. */
|
|
41
41
|
const REASON_MAX = 400;
|
|
@@ -89,7 +89,7 @@ export function isScoped(cwd, slug) {
|
|
|
89
89
|
* @returns {(string|null)} A reason phrased for an operator, or null when the file at `path` exists,
|
|
90
90
|
* hashes to the cited `sha256`, and its own `overall` reads "green".
|
|
91
91
|
*/
|
|
92
|
-
function unresolvedCitation(cwd, citation) {
|
|
92
|
+
function unresolvedCitation(cwd, citation, { round = null, runId = null } = {}) {
|
|
93
93
|
const rel = typeof citation?.path === "string" ? citation.path : "";
|
|
94
94
|
if (!rel) return "names no artifact path";
|
|
95
95
|
let text;
|
|
@@ -109,6 +109,20 @@ function unresolvedCitation(cwd, citation) {
|
|
|
109
109
|
try { body = JSON.parse(text); }
|
|
110
110
|
catch { return `cites ${rel}, whose bytes match the hash but do not read as a T0 verdict`; }
|
|
111
111
|
if (body?.overall !== "green") return `cites ${rel}, whose own verdict is "${body?.overall ?? "unknown"}", not green`;
|
|
112
|
+
// THE ARTIFACT HAS TO BE THE ONE THE CITATION SAYS IT IS. A re-hash proves the bytes are the
|
|
113
|
+
// file's; it says nothing about whose verdict the file holds. A PASS citing scope alpha's green
|
|
114
|
+
// artifact while declaring scope beta, or a prior round's, or a prior run's over the same slug,
|
|
115
|
+
// passed the digest check unremarked. The citation's own required `scope_id`, the round being
|
|
116
|
+
// judged and the run's key are compared to what the artifact records about itself.
|
|
117
|
+
if (typeof citation.scope_id === "string" && body?.scope_id && body.scope_id !== citation.scope_id) {
|
|
118
|
+
return `cites ${rel} for scope "${citation.scope_id}", but the artifact records scope "${body.scope_id}"`;
|
|
119
|
+
}
|
|
120
|
+
if (round != null && typeof body?.round === "number" && body.round !== round) {
|
|
121
|
+
return `cites ${rel}, a round ${body.round} artifact, as evidence for round ${round}`;
|
|
122
|
+
}
|
|
123
|
+
if (runId && body?.run_id && body.run_id !== runId) {
|
|
124
|
+
return `cites ${rel}, an artifact of run ${body.run_id}, as evidence for run ${runId}`;
|
|
125
|
+
}
|
|
112
126
|
return null;
|
|
113
127
|
}
|
|
114
128
|
|
|
@@ -143,15 +157,49 @@ function unresolvedCitation(cwd, citation) {
|
|
|
143
157
|
* citation resolves, an unscoped spec, or a block with no PASS/FAIL in it (there is no judgement
|
|
144
158
|
* to invalidate).
|
|
145
159
|
*/
|
|
146
|
-
|
|
160
|
+
/**
|
|
161
|
+
* Why a verdict cannot stand on its own criteria, or null when it can.
|
|
162
|
+
*
|
|
163
|
+
* `overall` is the judge's field, and nothing recomputed it from the criteria the judge graded: a
|
|
164
|
+
* PASS over a failing criterion, or over no criterion at all, validated and ingested, and the
|
|
165
|
+
* round loop branched on it. The evaluator's own first rule is that absence of evidence is a FAIL,
|
|
166
|
+
* so a PASS criterion with no evidence is no evidence either. Recomputed here, on ingest and on
|
|
167
|
+
* read alike: PASS means every graded criterion passed with evidence and at least one was graded;
|
|
168
|
+
* FAIL means at least one graded criterion failed.
|
|
169
|
+
*
|
|
170
|
+
* @param {object} verdict - The WorkResult's `verdict`.
|
|
171
|
+
* @returns {(string|null)} A reason phrased for an operator, or null.
|
|
172
|
+
*/
|
|
173
|
+
export function verdictProblem(verdict) {
|
|
174
|
+
const overall = verdict?.overall;
|
|
175
|
+
if (overall !== "PASS" && overall !== "FAIL") return null;
|
|
176
|
+
const criteria = Array.isArray(verdict.criteria) ? verdict.criteria : [];
|
|
177
|
+
const fails = criteria.filter((c) => c?.verdict === "FAIL");
|
|
178
|
+
const passes = criteria.filter((c) => c?.verdict === "PASS");
|
|
179
|
+
const other = criteria.length - fails.length - passes.length;
|
|
180
|
+
if (overall === "PASS") {
|
|
181
|
+
if (criteria.length === 0) return "the PASS verdict grades no criterion at all — a PASS with no evidence is a claim";
|
|
182
|
+
if (fails.length) return `the verdict says PASS while ${fails.length} of its ${criteria.length} criteria read FAIL — overall is derived from the criteria, never declared over them`;
|
|
183
|
+
if (other) return `the verdict says PASS while ${other} of its criteria carry no PASS/FAIL verdict`;
|
|
184
|
+
const bare = passes.filter((c) => !(typeof c?.evidence === "string" && c.evidence.trim()));
|
|
185
|
+
if (bare.length) return `the PASS verdict has ${bare.length} criterion(s) marked PASS with no evidence — absence of evidence is a FAIL by the evaluator's own first rule`;
|
|
186
|
+
return null;
|
|
187
|
+
}
|
|
188
|
+
if (criteria.length === 0) return "the FAIL verdict grades no criterion at all — a FAIL must name what failed";
|
|
189
|
+
if (!fails.length) return `the verdict says FAIL while every one of its ${criteria.length} graded criteria reads PASS — a FAIL must cite the criterion it failed`;
|
|
190
|
+
return null;
|
|
191
|
+
}
|
|
192
|
+
|
|
193
|
+
export function citationProblem(cwd, slug, verdict, { round = null } = {}) {
|
|
147
194
|
if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
|
|
148
195
|
if (!isScoped(cwd, slug)) return null;
|
|
149
196
|
if (!Array.isArray(verdict.t0_citations) || !verdict.t0_citations.length) {
|
|
150
197
|
return `the ${verdict.overall} verdict cites no T0 artifact, and a verdict on a scoped spec must ` +
|
|
151
198
|
"cite the T0 verdict it re-hashed (the order lists them under payload.t0_artifacts)";
|
|
152
199
|
}
|
|
200
|
+
const runId = readRunId(cwd, slug);
|
|
153
201
|
for (const citation of verdict.t0_citations) {
|
|
154
|
-
const reason = unresolvedCitation(cwd, citation);
|
|
202
|
+
const reason = unresolvedCitation(cwd, citation, { round, runId });
|
|
155
203
|
if (reason) return `the ${verdict.overall} verdict ${reason} — a T0 citation is re-hashed from disk, never taken on the handed word`;
|
|
156
204
|
}
|
|
157
205
|
return null;
|
|
@@ -185,7 +233,7 @@ export function evalVerdict(cwd, slug, round) {
|
|
|
185
233
|
? `the evaluator returned ${status || "no status"}: ${first}`
|
|
186
234
|
: `status ${status || "unknown"} with no PASS/FAIL verdict`), status);
|
|
187
235
|
}
|
|
188
|
-
const problem = citationProblem(cwd, slug, v);
|
|
236
|
+
const problem = verdictProblem(v) || citationProblem(cwd, slug, v, { round });
|
|
189
237
|
if (problem) return unfit(problem, status, overall);
|
|
190
238
|
return {
|
|
191
239
|
found: true,
|
package/kernel/probe/leg.mjs
CHANGED
|
@@ -27,7 +27,7 @@
|
|
|
27
27
|
import { existsSync, readdirSync, readFileSync } from "node:fs";
|
|
28
28
|
import { join, resolve } from "node:path";
|
|
29
29
|
import { runArgs } from "../lib/argv.mjs";
|
|
30
|
-
import { ordersDir, resultsDir, legLedger } from "../lib/paths.mjs";
|
|
30
|
+
import { ordersDir, resultsDir, legLedger, dispatchReceipts, readRunId } from "../lib/paths.mjs";
|
|
31
31
|
|
|
32
32
|
/** `<scope>-r<round>-a<attempt>.json` — the only address a build order is written under. */
|
|
33
33
|
const BUILD_ORDER = /^(.+)-r(\d+)-a(\d+)\.json$/;
|
|
@@ -80,6 +80,12 @@ export function legsOf(cwd, slug) {
|
|
|
80
80
|
const ingested = new Set(readLegs(legLedger(cwd, slug))
|
|
81
81
|
.filter((r) => r.ingested_at)
|
|
82
82
|
.map((r) => String(r.order_id)));
|
|
83
|
+
// Dispatch receipts are the hook layer's record that a Skill was invoked against an order —
|
|
84
|
+
// scoped to this run, since receipts over one slug accumulate across runs and order ids repeat.
|
|
85
|
+
const runId = readRunId(cwd, slug);
|
|
86
|
+
const receipted = new Set(readLegs(dispatchReceipts(cwd, slug))
|
|
87
|
+
.filter((r) => r?.order_id && (!runId || !r.run_id || r.run_id === runId))
|
|
88
|
+
.map((r) => String(r.order_id)));
|
|
83
89
|
const out = [];
|
|
84
90
|
for (const f of (existsSync(oDir) ? readdirSync(oDir) : []).filter((x) => x.endsWith(".json")).sort()) {
|
|
85
91
|
let orderId = null;
|
|
@@ -88,6 +94,7 @@ export function legsOf(cwd, slug) {
|
|
|
88
94
|
order: join(oDir, f),
|
|
89
95
|
name: f.replace(/\.json$/, ""),
|
|
90
96
|
order_id: orderId,
|
|
97
|
+
has_receipt: orderId !== null && receipted.has(orderId),
|
|
91
98
|
has_result: existsSync(join(rDir, f)),
|
|
92
99
|
applied: orderId !== null && ingested.has(orderId),
|
|
93
100
|
});
|
|
@@ -104,10 +111,25 @@ export function legsOf(cwd, slug) {
|
|
|
104
111
|
*/
|
|
105
112
|
export function orderLegState(cwd, slug, name) {
|
|
106
113
|
const o = legsOf(cwd, slug).find((x) => x.name === name);
|
|
107
|
-
if (!o) return { closed: false, found: false, order: null, order_id: null, has_result: false, applied: false };
|
|
114
|
+
if (!o) return { closed: false, found: false, order: null, order_id: null, has_receipt: false, has_result: false, applied: false };
|
|
108
115
|
return { closed: o.applied, found: true, ...o };
|
|
109
116
|
}
|
|
110
117
|
|
|
118
|
+
/**
|
|
119
|
+
* Orders that were dispatched — a receipt says a Skill ran against them — and never answered:
|
|
120
|
+
* no WorkResult on disk. Measured on a live run: the first dispatch wrote its four artifacts and
|
|
121
|
+
* no envelope, every later phase read the artifacts, and the run closed `shipped` with the order
|
|
122
|
+
* still open by construction — the export's own dispatch row said `answered: false` and nothing
|
|
123
|
+
* upstream of it had noticed. Distinct from {@link openLegs}: that is work nobody applied; this is
|
|
124
|
+
* a leg that never came back.
|
|
125
|
+
* @param {string} cwd - Project root.
|
|
126
|
+
* @param {string} slug - Feature slug.
|
|
127
|
+
* @returns {{order:string, name:string, order_id:(string|null)}[]}
|
|
128
|
+
*/
|
|
129
|
+
export function unansweredOrders(cwd, slug) {
|
|
130
|
+
return legsOf(cwd, slug).filter((o) => o.has_receipt && !o.has_result);
|
|
131
|
+
}
|
|
132
|
+
|
|
111
133
|
/**
|
|
112
134
|
* The results on disk that no leg row applied — finished work the board never saw.
|
|
113
135
|
* @param {string} cwd - Project root.
|
|
@@ -171,8 +193,13 @@ export function cli(rawArgv) {
|
|
|
171
193
|
// `--open`: every result nothing applied, across the whole run, any phase. Exit 0 when none.
|
|
172
194
|
if (args.open) {
|
|
173
195
|
const open = openLegs(cwd, args.slug);
|
|
174
|
-
|
|
175
|
-
|
|
196
|
+
const unanswered = unansweredOrders(cwd, args.slug);
|
|
197
|
+
console.log(JSON.stringify({
|
|
198
|
+
closed: open.length === 0 && unanswered.length === 0,
|
|
199
|
+
open_total: open.length, open: open.map((o) => o.order), open_ids: open.map((o) => o.order_id),
|
|
200
|
+
unanswered_total: unanswered.length, unanswered: unanswered.map((o) => o.order), unanswered_ids: unanswered.map((o) => o.order_id),
|
|
201
|
+
}));
|
|
202
|
+
process.exit(open.length === 0 && unanswered.length === 0 ? 0 : 1);
|
|
176
203
|
}
|
|
177
204
|
// `--order <stem>`: one order by file stem (`analyze`, `alpha-r1-a1`). Exit 0 when its leg closed.
|
|
178
205
|
if (args.order) {
|
package/kernel/probe/resume.mjs
CHANGED
|
@@ -62,8 +62,9 @@ import { runArgs } from "../lib/argv.mjs";
|
|
|
62
62
|
import { splitFrontmatter, uncoerce } from "../lib/contract.mjs";
|
|
63
63
|
import { globToRegExp } from "../verify/spec.mjs";
|
|
64
64
|
import { parseBoard } from "../reduce/board.mjs";
|
|
65
|
-
import { intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir, orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements, exportRunDir, lastRun, readRunId, tasksDir } from "../lib/paths.mjs";
|
|
65
|
+
import { intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir, orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements, exportRunDir, lastRun, readRunId, tasksDir, gates, verdictsDir, roundBuildDir } from "../lib/paths.mjs";
|
|
66
66
|
import { evalVerdict } from "./eval.mjs";
|
|
67
|
+
import { deriveRounds } from "./rounds.mjs";
|
|
67
68
|
import { collectRun, writeRun } from "../report/export.mjs";
|
|
68
69
|
|
|
69
70
|
/** The run-state values `references/protocol.md` (Part 4 — State) defines. A typo'd status is a rejection,
|
|
@@ -560,7 +561,83 @@ export function setRunStatus(cwd, slug, status) {
|
|
|
560
561
|
* already normalized prose (newlines collapsed, truncated) — this function only `uncoerce`s it.
|
|
561
562
|
* @returns {string} The rewritten text.
|
|
562
563
|
*/
|
|
563
|
-
|
|
564
|
+
/**
|
|
565
|
+
* What the ledger's front matter and tables should say at a close, derived from the run's own
|
|
566
|
+
* records — the same ones the export reads. The counters (`final_verdict`, `rounds_used`) and the
|
|
567
|
+
* Rounds/Decisions tables used to be hand-shaped placeholders nothing filled: a run closed
|
|
568
|
+
* `shipped` beside `final_verdict: ~`, `rounds_used: 0`, an empty Decisions table and seven rows in
|
|
569
|
+
* `gates.jsonl`. A reader who trusted the counters concluded no round ran. Measured live.
|
|
570
|
+
*
|
|
571
|
+
* @param {string} cwd - Project root.
|
|
572
|
+
* @param {string} slug - Feature slug.
|
|
573
|
+
* @returns {{final_verdict:string, rounds_used:(number|null), decisions:object[], roundRows:string[]}}
|
|
574
|
+
*/
|
|
575
|
+
export function deriveLedgerFacts(cwd, slug) {
|
|
576
|
+
const readJson = (p) => { try { return JSON.parse(readFileSync(p, "utf8")); } catch { return null; } };
|
|
577
|
+
const evalRows = [];
|
|
578
|
+
const rDir = resultsDir(cwd, slug);
|
|
579
|
+
if (existsSync(rDir)) {
|
|
580
|
+
for (const f of readdirSync(rDir)) {
|
|
581
|
+
const m = f.match(/^evaluate-r(\d+)\.json$/);
|
|
582
|
+
if (!m) continue;
|
|
583
|
+
const r = readJson(join(rDir, f));
|
|
584
|
+
const overall = r?.verdict?.overall;
|
|
585
|
+
if (overall) evalRows.push({ round: Number(m[1]), overall: String(overall), criteria: Array.isArray(r.verdict.criteria) ? r.verdict.criteria : [] });
|
|
586
|
+
}
|
|
587
|
+
}
|
|
588
|
+
evalRows.sort((a, b) => a.round - b.round);
|
|
589
|
+
const finalVerdict = evalRows.length ? evalRows[evalRows.length - 1].overall : "not-evaluated";
|
|
590
|
+
const decisions = [];
|
|
591
|
+
const gp = gates(cwd, slug);
|
|
592
|
+
if (existsSync(gp)) {
|
|
593
|
+
for (const line of readFileSync(gp, "utf8").split("\n")) {
|
|
594
|
+
if (!line.trim()) continue;
|
|
595
|
+
try { decisions.push(JSON.parse(line)); } catch { /* a torn line proves nothing */ }
|
|
596
|
+
}
|
|
597
|
+
}
|
|
598
|
+
const t0ByRound = {};
|
|
599
|
+
const vDir = verdictsDir(cwd, slug);
|
|
600
|
+
if (existsSync(vDir)) {
|
|
601
|
+
for (const f of readdirSync(vDir).filter((x) => x.endsWith(".json")).sort()) {
|
|
602
|
+
const v = readJson(join(vDir, f));
|
|
603
|
+
if (typeof v?.round !== "number") continue;
|
|
604
|
+
const t = (t0ByRound[v.round] ||= { green: 0, red: 0 });
|
|
605
|
+
if (v.overall === "green") t.green++; else t.red++;
|
|
606
|
+
}
|
|
607
|
+
}
|
|
608
|
+
const gateByRound = {};
|
|
609
|
+
const bDir = roundBuildDir(cwd, slug);
|
|
610
|
+
if (existsSync(bDir)) {
|
|
611
|
+
for (const f of readdirSync(bDir).sort()) {
|
|
612
|
+
const m = f.match(/^r(\d+)-t\d+\.json$/);
|
|
613
|
+
if (m) gateByRound[Number(m[1])] = readJson(join(bDir, f))?.overall ?? "?";
|
|
614
|
+
}
|
|
615
|
+
}
|
|
616
|
+
const roundNums = [...new Set([...Object.keys(t0ByRound), ...Object.keys(gateByRound), ...evalRows.map((e) => e.round)].map(Number))].sort((a, b) => a - b);
|
|
617
|
+
const roundRows = [];
|
|
618
|
+
for (const r of roundNums) {
|
|
619
|
+
const t0 = t0ByRound[r], bg = gateByRound[r], ev = evalRows.find((e) => e.round === r);
|
|
620
|
+
roundRows.push(`| Build | ${r} | ${t0 ? `T0 ${t0.green} green / ${t0.red} red` : "no T0 verdict"} | — | round build gate: ${bg ?? "not run"} |`);
|
|
621
|
+
if (ev) {
|
|
622
|
+
const p = ev.criteria.filter((c) => c?.verdict === "PASS").length, f = ev.criteria.filter((c) => c?.verdict === "FAIL").length;
|
|
623
|
+
roundRows.push(`| Eval | ${r} | ${ev.overall} | — | ${p} PASS / ${f} FAIL criteria |`);
|
|
624
|
+
}
|
|
625
|
+
}
|
|
626
|
+
let roundsUsed = null;
|
|
627
|
+
try { roundsUsed = deriveRounds(cwd, slug, null)?.rounds_used ?? null; } catch { roundsUsed = null; }
|
|
628
|
+
if (roundsUsed == null && roundNums.length) roundsUsed = roundNums[roundNums.length - 1];
|
|
629
|
+
return { final_verdict: finalVerdict, rounds_used: roundsUsed, decisions, roundRows };
|
|
630
|
+
}
|
|
631
|
+
|
|
632
|
+
/** Replace the rows of one markdown table (identified by its header line) with `rows`, keeping the header, the separator and any row the caller marks as kept. */
|
|
633
|
+
function rewriteTable(body, headerRe, rows, keepRow = null) {
|
|
634
|
+
return body.replace(headerRe, (m, header, sep, oldRows) => {
|
|
635
|
+
const kept = keepRow && oldRows ? oldRows.split("\n").filter((l) => l.startsWith(keepRow)) : [];
|
|
636
|
+
return header + sep + [...kept, ...rows].map((l) => `${l}\n`).join("");
|
|
637
|
+
});
|
|
638
|
+
}
|
|
639
|
+
|
|
640
|
+
function writeCloseLines(body, { status, closedAt, cause, derived = null }) {
|
|
564
641
|
const causeLine = `close_cause: ${uncoerce(cause || null)}`;
|
|
565
642
|
const closedStatusLine = `closed_status: ${status}`;
|
|
566
643
|
let out = body
|
|
@@ -574,6 +651,15 @@ function writeCloseLines(body, { status, closedAt, cause }) {
|
|
|
574
651
|
out = /^close_cause:.*$/m.test(out)
|
|
575
652
|
? out.replace(/^close_cause:.*$/m, causeLine)
|
|
576
653
|
: out.replace(/^closed_at:.*$/m, (m) => `${m}\n${causeLine}`);
|
|
654
|
+
if (derived) {
|
|
655
|
+
// THE COUNTERS AND TABLES ARE DERIVED, IN THE SAME PASS AS THE CLOSE LINE. Written from the
|
|
656
|
+
// run's own records so they cannot disagree with the export that reads the same records.
|
|
657
|
+
if (/^final_verdict:.*$/m.test(out)) out = out.replace(/^final_verdict:.*$/m, `final_verdict: ${derived.final_verdict}`);
|
|
658
|
+
if (typeof derived.rounds_used === "number" && /^rounds_used:.*$/m.test(out)) out = out.replace(/^rounds_used:.*$/m, `rounds_used: ${derived.rounds_used}`);
|
|
659
|
+
out = rewriteTable(out, /(\| Phase \| Round \| Result \| Duration \| Notes \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, derived.roundRows, "| Init");
|
|
660
|
+
const decisionRows = derived.decisions.map((g) => `| ${g.gate ?? "?"} | ${g.decision ?? g.status ?? "?"} | ${g.source ?? "?"} | ${String(g.note ?? "").replace(/\|/g, "\\|").replace(/\s+/g, " ").slice(0, 160)} |`);
|
|
661
|
+
out = rewriteTable(out, /(\| Gate \| Decision \| Source \| Note \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, decisionRows);
|
|
662
|
+
}
|
|
577
663
|
return out;
|
|
578
664
|
}
|
|
579
665
|
|
|
@@ -696,6 +782,9 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
696
782
|
// The Ship phase (`shapeup-run.js`) already exports a shipped run itself, before this call ever
|
|
697
783
|
// runs — Stage 2 adds the endings that wrote nothing, and leaves that path untouched.
|
|
698
784
|
const shouldExport = withExport && status !== "shipped";
|
|
785
|
+
// Derived once, from the run's own records, and written with the close line (see deriveLedgerFacts).
|
|
786
|
+
let derived = null;
|
|
787
|
+
try { derived = deriveLedgerFacts(cwd, slug); } catch { derived = null; }
|
|
699
788
|
|
|
700
789
|
/**
|
|
701
790
|
* Everything a close owes the checkout once the ledger line is written: export the run's
|
|
@@ -778,7 +867,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
778
867
|
// new line rather than lost, and the return says so explicitly.
|
|
779
868
|
const closedAt = new Date().toISOString();
|
|
780
869
|
const foldedCause = `${normCause || "no reason recorded"} — supersedes an earlier close recorded ${priorClosedAt} (cause: ${JSON.stringify(priorCause)})`.slice(0, 4000);
|
|
781
|
-
body = writeCloseLines(body, { status, closedAt, cause: foldedCause });
|
|
870
|
+
body = writeCloseLines(body, { status, closedAt, cause: foldedCause, derived });
|
|
782
871
|
try { writeFileSync(p, body); } catch (e) {
|
|
783
872
|
return { ok: false, path: p, status, reason: `could not write the ledger: ${e.message}` };
|
|
784
873
|
}
|
|
@@ -793,7 +882,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
793
882
|
}
|
|
794
883
|
|
|
795
884
|
const closedAt = new Date().toISOString();
|
|
796
|
-
body = writeCloseLines(body, { status, closedAt, cause: normCause });
|
|
885
|
+
body = writeCloseLines(body, { status, closedAt, cause: normCause, derived });
|
|
797
886
|
try { writeFileSync(p, body); } catch (e) {
|
|
798
887
|
return { ok: false, path: p, status, reason: `could not write the ledger: ${e.message}` };
|
|
799
888
|
}
|
package/kernel/probe/rounds.mjs
CHANGED
|
@@ -22,7 +22,7 @@
|
|
|
22
22
|
|
|
23
23
|
import { existsSync, readdirSync, readFileSync } from "node:fs";
|
|
24
24
|
import { join } from "node:path";
|
|
25
|
-
import { ordersDir, verdictsDir, roundBuildDir, resultsDir } from "../lib/paths.mjs";
|
|
25
|
+
import { ordersDir, verdictsDir, roundBuildDir, resultsDir, readRunId } from "../lib/paths.mjs";
|
|
26
26
|
|
|
27
27
|
/** Parse a JSON file, returning null rather than throwing — every reader here is best-effort. */
|
|
28
28
|
function readJson(p) {
|
|
@@ -59,12 +59,22 @@ const maxOf = (nums) => (nums.length ? Math.max(...nums) : null);
|
|
|
59
59
|
* @returns {{rounds_used:*, rounds_judged:(number|null)}} Both counts.
|
|
60
60
|
*/
|
|
61
61
|
export function deriveRounds(cwd, slug, fallback) {
|
|
62
|
+
// THIS RUN'S RECORDS ONLY. Orders, verdicts and build gates over one slug accumulate across runs
|
|
63
|
+
// and carry their run key; walking the directories unfiltered handed a run that had dispatched
|
|
64
|
+
// nothing a prior run's round count — into the committed report. A record with no key at all was
|
|
65
|
+
// written before the key existed and is kept; one with a different key is another run's.
|
|
66
|
+
const runId = readRunId(cwd, slug);
|
|
67
|
+
const mine = (rec) => !runId || !rec?.run_id || rec.run_id === runId;
|
|
62
68
|
const orderRounds = [];
|
|
63
69
|
const oDir = ordersDir(cwd, slug);
|
|
70
|
+
const orderOf = {};
|
|
64
71
|
if (existsSync(oDir)) {
|
|
65
72
|
for (const f of readdirSync(oDir)) {
|
|
66
73
|
if (!f.endsWith(".json")) continue;
|
|
67
|
-
const
|
|
74
|
+
const o = readJson(join(oDir, f));
|
|
75
|
+
orderOf[f] = o;
|
|
76
|
+
if (!mine(o)) continue;
|
|
77
|
+
const r = orderRound(o?.order_id);
|
|
68
78
|
if (r !== null) orderRounds.push(r);
|
|
69
79
|
}
|
|
70
80
|
}
|
|
@@ -74,7 +84,7 @@ export function deriveRounds(cwd, slug, fallback) {
|
|
|
74
84
|
if (existsSync(vDir)) {
|
|
75
85
|
for (const f of readdirSync(vDir).filter((x) => x.endsWith(".json"))) {
|
|
76
86
|
const v = readJson(join(vDir, f));
|
|
77
|
-
if (typeof v?.round === "number") verdictRounds.push(v.round);
|
|
87
|
+
if (mine(v) && typeof v?.round === "number") verdictRounds.push(v.round);
|
|
78
88
|
}
|
|
79
89
|
}
|
|
80
90
|
|
|
@@ -83,7 +93,7 @@ export function deriveRounds(cwd, slug, fallback) {
|
|
|
83
93
|
if (existsSync(bDir)) {
|
|
84
94
|
for (const f of readdirSync(bDir)) {
|
|
85
95
|
const m = f.match(/^r(\d+)-t\d+\.json$/);
|
|
86
|
-
if (m) buildGateRounds.push(Number(m[1]));
|
|
96
|
+
if (m && mine(readJson(join(bDir, f)))) buildGateRounds.push(Number(m[1]));
|
|
87
97
|
}
|
|
88
98
|
}
|
|
89
99
|
|
|
@@ -92,7 +102,8 @@ export function deriveRounds(cwd, slug, fallback) {
|
|
|
92
102
|
if (existsSync(rDir)) {
|
|
93
103
|
for (const f of readdirSync(rDir)) {
|
|
94
104
|
const m = f.match(/^evaluate-r(\d+)\.json$/);
|
|
95
|
-
|
|
105
|
+
// A WorkResult carries no run key; it reaches one through its order of the same name.
|
|
106
|
+
if (m && mine(orderOf[f] ?? readJson(join(oDir, f)))) evalRounds.push(Number(m[1]));
|
|
96
107
|
}
|
|
97
108
|
}
|
|
98
109
|
|
package/kernel/reduce/ingest.mjs
CHANGED
|
@@ -37,7 +37,7 @@ import { fileURLToPath } from "node:url";
|
|
|
37
37
|
import { validate } from "../verify/envelope.mjs";
|
|
38
38
|
import { runArgs } from "../lib/argv.mjs";
|
|
39
39
|
import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId } from "../lib/paths.mjs";
|
|
40
|
-
import { citationProblem } from "../probe/eval.mjs";
|
|
40
|
+
import { citationProblem, verdictProblem } from "../probe/eval.mjs";
|
|
41
41
|
|
|
42
42
|
const HERE = dirname(fileURLToPath(import.meta.url));
|
|
43
43
|
const RESULT_SCHEMA = JSON.parse(readFileSync(resolve(HERE, "../schemas/work-result.schema.json"), "utf8"));
|
|
@@ -187,11 +187,48 @@ export function setCheckbox(body, ac, checked) {
|
|
|
187
187
|
* @param {boolean} done - When true, rewrite the row's status emoji/word to done; false is a no-op.
|
|
188
188
|
* @returns {string} The board text with the matching row updated (unchanged when no row matches).
|
|
189
189
|
*/
|
|
190
|
-
|
|
190
|
+
/**
|
|
191
|
+
* Is this index line the row FOR `taskId` — its first cell — rather than a row that merely
|
|
192
|
+
* mentions it? A task id appears in other rows' `Depends On` column, and matching by substring
|
|
193
|
+
* anywhere in the line flipped every dependent of a finished task to done along with it. Measured
|
|
194
|
+
* on a live run: the executor reported one task `skipped`, its task file still said `ready`, and
|
|
195
|
+
* the index showed it ✅ because the task it depended on had just been ticked. The census read the
|
|
196
|
+
* index.
|
|
197
|
+
* @param {string} line - One line of `tasks/_index.md`.
|
|
198
|
+
* @param {string} taskId - `TASK-NNN`.
|
|
199
|
+
* @returns {boolean}
|
|
200
|
+
*/
|
|
201
|
+
export function rowIs(line, taskId) {
|
|
202
|
+
const id = taskId.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
|
203
|
+
return new RegExp(`^\\|\\s*(?:\\[\\[)?${id}(?:\\\\\\|${id}\\]\\])?\\s*\\|`).test(line);
|
|
204
|
+
}
|
|
205
|
+
|
|
206
|
+
/**
|
|
207
|
+
* The status cell for one TaskResult status, or null when the index should not change.
|
|
208
|
+
* @param {(string|boolean)} status - A TaskResult status (`true` is accepted as `done` for callers written before skipped existed).
|
|
209
|
+
* @returns {({icon:string, word:string}|null)} The icon and word the index row should carry.
|
|
210
|
+
*/
|
|
211
|
+
export function boardCell(status) {
|
|
212
|
+
if (status === true || status === "done") return { icon: "✅", word: "done" };
|
|
213
|
+
if (status === "skipped") return { icon: "⏭", word: "skipped" };
|
|
214
|
+
if (status === "partial" || status === "failed") return { icon: "🔄", word: "in-progress" };
|
|
215
|
+
return null;
|
|
216
|
+
}
|
|
217
|
+
|
|
218
|
+
/**
|
|
219
|
+
* Rewrite one task's status cell in `tasks/_index.md` — the row whose id cell is `taskId`, never a
|
|
220
|
+
* row that merely mentions it in `Depends On`.
|
|
221
|
+
* @param {string} indexBody - The index file's text.
|
|
222
|
+
* @param {string} taskId - `TASK-NNN`.
|
|
223
|
+
* @param {(string|boolean)} status - A TaskResult status; statuses the index does not render leave it unchanged.
|
|
224
|
+
* @returns {string} The rewritten index text.
|
|
225
|
+
*/
|
|
226
|
+
export function updateBoardRow(indexBody, taskId, status) {
|
|
227
|
+
const cell = boardCell(status);
|
|
228
|
+
if (!cell) return indexBody;
|
|
191
229
|
return indexBody.split(/\r?\n/).map((line) => {
|
|
192
|
-
if (!line
|
|
193
|
-
|
|
194
|
-
return line;
|
|
230
|
+
if (!rowIs(line, taskId)) return line;
|
|
231
|
+
return line.replace(/⬜|🔄|⏳|🚫|✅|⏭/g, cell.icon).replace(/\b(ready|in-progress|blocked|done|skipped)\b/gi, cell.word);
|
|
195
232
|
}).join("\n");
|
|
196
233
|
}
|
|
197
234
|
|
|
@@ -240,14 +277,18 @@ function applyResultLocked(result, { cwd, slug }) {
|
|
|
240
277
|
body = setFrontmatter(body, "completed_at", today());
|
|
241
278
|
} else if (tr.status === "partial" || tr.status === "failed") {
|
|
242
279
|
body = setFrontmatter(body, "status", "in-progress");
|
|
280
|
+
} else if (tr.status === "skipped") {
|
|
281
|
+
// Rendered as what it is. A skipped task used to leave its file at `ready` and, through the
|
|
282
|
+
// substring match above, could show ✅ on the index — the census read the index.
|
|
283
|
+
body = setFrontmatter(body, "status", "skipped");
|
|
243
284
|
}
|
|
244
285
|
// Execution Log (append; the checkbox list must never disagree with it).
|
|
245
286
|
const logLines = (tr.ac_results || []).map((a) => `- ${a.ac}: ${a.result}${a.evidence ? ` (${a.evidence})` : ""}`).join("\n");
|
|
246
287
|
body += `\n\n## Execution Log — ${today()} (${result.order_id})\n- executor: ${result.worker || "task-executor"} via ingest-result\n- status: ${tr.status}\n${logLines}${tr.notes ? `\n- notes: ${tr.notes}` : ""}\n`;
|
|
247
288
|
writeFileSync(path, body);
|
|
248
289
|
summary.tasks_updated.push(tr.task_id);
|
|
249
|
-
if (
|
|
250
|
-
writeFileSync(boardIndex, updateBoardRow(readFileSync(boardIndex, "utf8"), tr.task_id,
|
|
290
|
+
if (existsSync(boardIndex) && boardCell(tr.status)) {
|
|
291
|
+
writeFileSync(boardIndex, updateBoardRow(readFileSync(boardIndex, "utf8"), tr.task_id, tr.status));
|
|
251
292
|
}
|
|
252
293
|
}
|
|
253
294
|
|
|
@@ -271,7 +312,7 @@ function applyResultLocked(result, { cwd, slug }) {
|
|
|
271
312
|
summary.unblocked.push(t.id);
|
|
272
313
|
if (existsSync(boardIndex)) {
|
|
273
314
|
const idx = readFileSync(boardIndex, "utf8").split(/\r?\n/).map((line) =>
|
|
274
|
-
line
|
|
315
|
+
rowIs(line, t.id)
|
|
275
316
|
? line.replace(/🚫|⏳/g, "⬜").replace(/\bblocked\b/gi, "ready")
|
|
276
317
|
: line).join("\n");
|
|
277
318
|
writeFileSync(boardIndex, idx);
|
|
@@ -656,7 +697,8 @@ export async function cli(rawArgv) {
|
|
|
656
697
|
// on (see `citationProblem`). `probe eval` refuses it to the round loop; refusing it here as well
|
|
657
698
|
// keeps the verdict ledger from recording a verdict the loop will never branch on.
|
|
658
699
|
if (result.verdict) {
|
|
659
|
-
const
|
|
700
|
+
const evalRound = Number((String(result.order_id).match(/-r(\d+)$/) || [])[1]) || null;
|
|
701
|
+
const problem = verdictProblem(result.verdict) || citationProblem(cwd, String(result.order_id).split("/")[0], result.verdict, { round: evalRound });
|
|
660
702
|
if (problem) {
|
|
661
703
|
console.error(`ingest-result: result refused — ${problem}.`);
|
|
662
704
|
console.error(` The round stays open: re-dispatch the evaluator against its order, which lists`);
|
package/kernel/reduce/ship.mjs
CHANGED
|
@@ -329,7 +329,13 @@ export function buildReport(facts) {
|
|
|
329
329
|
"*Run state (board, orders, results, T0 artifacts, evaluation and QA reports) stays in the",
|
|
330
330
|
"gitignored local tier (ADR-0001). This report",
|
|
331
331
|
"is the frozen conclusion of it.*", "");
|
|
332
|
-
|
|
332
|
+
// THE WHOLE REPORT, ONCE. Board ids were anchored in three places and leaked through a fourth: a
|
|
333
|
+
// discovery-ledger entry copied verbatim into "Discovered, not built" carried `[TASK-005]`, this
|
|
334
|
+
// kernel wrote it straight to the committed tier past the hook that guards the model's edits, and
|
|
335
|
+
// the next run's L1b lint red'd the file the previous run had frozen. A committed report may not
|
|
336
|
+
// name a board id anywhere, so the rule is applied to the finished text rather than section by
|
|
337
|
+
// section.
|
|
338
|
+
return deboard(L.join("\n"), board.anchors);
|
|
333
339
|
}
|
|
334
340
|
|
|
335
341
|
/**
|
package/kernel/report/export.mjs
CHANGED
|
@@ -173,12 +173,14 @@ function criterionRows(dir, runId, t) {
|
|
|
173
173
|
* (see `appendGateLedger`), so this is a pass-through with a stamped `run_id` fallback rather than
|
|
174
174
|
* a re-derivation: two readers of "what did this gate decide" must not compute the answer twice.
|
|
175
175
|
* @param {object} g - One parsed line of `gates.jsonl`.
|
|
176
|
-
* @param {(string|null)} runId -
|
|
176
|
+
* @param {(string|null)} runId - The run being exported (unused for attribution — the row's own key is the only one exported).
|
|
177
177
|
* @returns {object} A flat `gate_decision` row.
|
|
178
178
|
*/
|
|
179
179
|
function gateDecisionRow(g, runId) {
|
|
180
180
|
return {
|
|
181
|
-
|
|
181
|
+
// The row's own key, never the current run's stamped on: a row that carries no key was written
|
|
182
|
+
// before the ledger did, and is exported as unattributed rather than claimed.
|
|
183
|
+
run_id: g?.run_id ?? null,
|
|
182
184
|
gate: g?.gate ?? null,
|
|
183
185
|
decision: g?.decision ?? null,
|
|
184
186
|
status: g?.status ?? null,
|
|
@@ -260,7 +262,9 @@ export function collectRun(cwd, slug) {
|
|
|
260
262
|
criterion_verdict: criterionRows(evaluationDir(cwd, slug), runId, t),
|
|
261
263
|
hook_decision,
|
|
262
264
|
// The decision that crossed each gate, and the round build gate's own artifact.
|
|
263
|
-
|
|
265
|
+
// Scoped to the run, like hook_decision one line up: gate rows over one slug accumulate across
|
|
266
|
+
// runs, and a prior run's L4 exported under this run's key is a fabricated sign-off.
|
|
267
|
+
gate_decision: readJsonl(gatesPath(cwd, slug), t).filter((g) => runId && g?.run_id === runId).map((g) => gateDecisionRow(g, runId)),
|
|
264
268
|
build_gate: readJsonDir(roundBuildDir(cwd, slug), t).map((a) => buildGateRow(a, runId)),
|
|
265
269
|
leg: readJsonl(legLedger(cwd, slug), t)
|
|
266
270
|
.filter((r) => !runId || !r?.run_id || r.run_id === runId)
|
package/kernel/verify/trace.mjs
CHANGED
|
@@ -40,7 +40,7 @@ import { resolve, join, dirname, relative, isAbsolute } from "node:path";
|
|
|
40
40
|
import { readBoard } from "../compile.mjs";
|
|
41
41
|
import { runArgs } from "../lib/argv.mjs";
|
|
42
42
|
import { sharedRoot, traceDir, relLocal } from "../lib/paths.mjs";
|
|
43
|
-
import { readContract, unreadableReason, LEGACY_LAYOUT, WIRING_MAP, PROJECT_PROFILE } from "../lib/contract.mjs";
|
|
43
|
+
import { readContract, unreadableReason, LEGACY_LAYOUT, WIRING_MAP, PROJECT_PROFILE, reqId } from "../lib/contract.mjs";
|
|
44
44
|
|
|
45
45
|
// --- requirements.md registry parser -----------------------------------------
|
|
46
46
|
// A committed markdown table: | REQ-id | clause (verbatim) | source | status | note |
|
|
@@ -87,7 +87,10 @@ export function coveredReqIds(board) {
|
|
|
87
87
|
for (const task of board) {
|
|
88
88
|
for (const ac of task.acceptance_criteria || []) {
|
|
89
89
|
const covers = typeof ac === "object" && Array.isArray(ac.covers) ? ac.covers : [];
|
|
90
|
-
|
|
90
|
+
// ONE KEY SPACE. `R-2`, `[[REQ-5]]` and `req-4` are the same clause spelled three ways, and the
|
|
91
|
+
// sibling rule accepts all of them; testing the raw string here counted every one as nothing, so
|
|
92
|
+
// an author told at L1b to cover a requirement with an AC — which they had — stayed red.
|
|
93
|
+
for (const raw of covers) { const id = reqId(raw); if (/^REQ-\d+$/.test(id)) covered.add(id); }
|
|
91
94
|
}
|
|
92
95
|
}
|
|
93
96
|
return covered;
|
package/package.json
CHANGED
|
@@ -482,6 +482,7 @@ const ORDERLEG = {
|
|
|
482
482
|
closed: { type: "boolean" },
|
|
483
483
|
found: { type: "boolean" },
|
|
484
484
|
order: nullable("string"),
|
|
485
|
+
has_receipt: { type: "boolean" },
|
|
485
486
|
has_result: { type: "boolean" },
|
|
486
487
|
applied: { type: "boolean" },
|
|
487
488
|
},
|
|
@@ -936,7 +937,18 @@ async function requireLeg(gate, phaseKey, phaseName, orderStem = phaseKey) {
|
|
|
936
937
|
const ask = () => query(`probe leg --slug ${slug} --order "${orderStem}"`, ORDERLEG, phaseName, `legcheck:${orderStem}`);
|
|
937
938
|
let leg = await ask();
|
|
938
939
|
if (!leg || !leg.found) { log(`${gate} — could not ask the leg ledger about "${phaseKey}" (probe returned ${leg ? "no order" : "nothing"}); proceeding on the artifact alone.`); return null; }
|
|
939
|
-
if (!leg.has_result
|
|
940
|
+
if (!leg.has_result) {
|
|
941
|
+
// The artifact is on disk and the envelope never came back. Measured live: the first dispatch of
|
|
942
|
+
// a run wrote its four artifacts and no WorkResult, every later phase read the artifacts, and
|
|
943
|
+
// the run closed `shipped` over an order still open by construction. The phase stands on its
|
|
944
|
+
// artifact — that is what the post-condition checks — but the order is named at the close, and
|
|
945
|
+
// the next run over this slug will find it unanswered rather than be surprised by it.
|
|
946
|
+
log(`${gate} — "${phaseKey}" ${leg.has_receipt ? "was dispatched (receipt on disk) and" : "has an order but no receipt, and"} never answered: no WorkResult at all. ` +
|
|
947
|
+
`Its artifacts are on disk and the phase proceeds on them; the order stays unanswered and is named at the close.`);
|
|
948
|
+
if (leg.order && !unansweredOrders.includes(leg.order)) unansweredOrders.push(leg.order);
|
|
949
|
+
return null;
|
|
950
|
+
}
|
|
951
|
+
if (leg.applied) return null;
|
|
940
952
|
log(`${gate} — "${phaseKey}" came back with a result nothing applied (no leg row). Ingesting it here: ${leg.order}.`);
|
|
941
953
|
await advisory(`reduce ingest --order "${leg.order}"`, phaseName, `late-ingest:${phaseKey}`);
|
|
942
954
|
leg = await ask();
|
|
@@ -1084,6 +1096,9 @@ async function requireLaunchRecord() {
|
|
|
1084
1096
|
// it warns and continues — and the warning travels in the RunReturn, because a headless stdout
|
|
1085
1097
|
// carries only the final message and a diagnostic on a channel nobody reads is not a diagnostic.
|
|
1086
1098
|
const stateWarnings = [];
|
|
1099
|
+
// Orders a receipt says were dispatched and no WorkResult ever answered — a leg that did its craft
|
|
1100
|
+
// and never came back. Named at every close, whatever the close; never folded into "complete".
|
|
1101
|
+
const unansweredOrders = [];
|
|
1087
1102
|
async function setRunStatus(status, phaseName) {
|
|
1088
1103
|
const r = await cmd(`probe resume --slug ${slug} --set-status ${status}`, phaseName, `status:${status}`);
|
|
1089
1104
|
if (!r.ok) {
|
|
@@ -1143,7 +1158,8 @@ async function closeIfTerminal(ret) {
|
|
|
1143
1158
|
// `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
|
|
1144
1159
|
// a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
|
|
1145
1160
|
// stay silent for it exactly as they did when this file's own guard returned early.
|
|
1146
|
-
const
|
|
1161
|
+
const unanswered = unansweredOrders.length ? ` unanswered_orders=${unansweredOrders.length}` : "";
|
|
1162
|
+
const r = await cmd(`probe resume --slug ${slug} --close-arm ${ret.status} --cause "${causeArg(cause + unanswered)}"`, "Ship", `close:${ret.status}`);
|
|
1147
1163
|
if (!r.ok) {
|
|
1148
1164
|
const why = (r.detail || `exit ${r.exit_code}`).trim();
|
|
1149
1165
|
log(`RUN STATE — close(${ret.status}) did not take: ${why}. This return's own status and reason still ` +
|