shapeup-sdlc 3.7.1 → 3.7.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/README.md +11 -4
- package/kernel/compile.mjs +27 -1
- package/kernel/gate.mjs +1 -1
- package/kernel/probe/eval.mjs +79 -7
- package/kernel/reduce/hill.mjs +200 -25
- package/kernel/reduce/ship.mjs +43 -1
- package/kernel/schemas/domain.schema.json +3 -2
- package/package.json +1 -1
- package/skills/hill-chart/SKILL.md +10 -6
- package/skills/tech-lead/references/gates.md +1 -1
- package/skills/tech-lead/workflows/shapeup-run.js +19 -5
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.7.
|
|
4
|
+
"version": "3.7.2",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/README.md
CHANGED
|
@@ -41,9 +41,14 @@ runs over unfinished tasks rather than denying it. The board is local to the mac
|
|
|
41
41
|
harness — see [ADR-0001](docs/design/adr/0001-consumer-file-organization.md).)
|
|
42
42
|
|
|
43
43
|
**2. Progress is measured, not claimed.** A scope counts as built only when `t0-verify` runs
|
|
44
|
-
its fixtures
|
|
45
|
-
|
|
46
|
-
worker
|
|
44
|
+
its fixtures and its DB probe and writes an artifact to disk — with each command's exit code,
|
|
45
|
+
its captured output, and whether it ran at all. The evaluator must cite that artifact, and the
|
|
46
|
+
hill phase is derived from artifacts rather than from a worker's own account of its progress.
|
|
47
|
+
Two limits, stated here because the point of this section is that a claim without a mechanism
|
|
48
|
+
behind it is the thing this harness exists to prevent: the **seesaw** regression arm is declared
|
|
49
|
+
and not yet wired (no run writes its registry — wiring it is an open Betting Table decision), and
|
|
50
|
+
nothing in the runtime **re-hashes** the citation the evaluator is instructed to re-hash. Both
|
|
51
|
+
are open items in `shapeup/knowledge-base/harness-defects.md`, not shipped guarantees.
|
|
47
52
|
→ *Prevents: "done" asserted with nothing behind it.*
|
|
48
53
|
|
|
49
54
|
**3. Parallel work can't corrupt shared state.** Each scope gets a write-whitelist of files
|
|
@@ -300,7 +305,9 @@ These hold across the harness and are the reason it stays predictable:
|
|
|
300
305
|
count events and neither can notice a single round running for half an hour — tripping it routes
|
|
301
306
|
to GATE H, so a run out of time ships what is green instead of being killed and shipping nothing.
|
|
302
307
|
- **Hill phase is mechanical, never self-reported** — derived only from T0/T1/seesaw facts, closing
|
|
303
|
-
the self-reported-confidence risk
|
|
308
|
+
the self-reported-confidence risk — with one gap on record: a scope with no discovery
|
|
309
|
+
ledger derives the same phase as one whose unknowns are all closed, so absence still reads
|
|
310
|
+
as progress on that one arm.
|
|
304
311
|
- **One writer per shared file** — every board/ledger/verdict write goes through
|
|
305
312
|
`harness reduce ingest`; workers return data and never touch shared state.
|
|
306
313
|
- **Traceability is oracle-checked, opt-in** — `harness verify trace` verifies covers-closure and
|
package/kernel/compile.mjs
CHANGED
|
@@ -218,6 +218,32 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
|
|
|
218
218
|
// in its `allowed` globs, so only a `frozen` entry denies the write. `translate` is the one
|
|
219
219
|
// operation that legitimately rewrites a pitch, and it writes the COMMITTED copy, not this one.
|
|
220
220
|
const FROZEN_INTAKE = [`${local}/intake.md`, `${local}/breadboard.md`];
|
|
221
|
+
// THE LEG'S OWN ATTESTATION. A dispatch receipt, a leg-completion row and a T0 verdict exist to
|
|
222
|
+
// answer "did a worker actually run, and what did it measure" from evidence the leg being judged
|
|
223
|
+
// does not control — the whole reason to prefer them over a compiled order, which is writable by
|
|
224
|
+
// that same hand. None of the three is ever produced by this assistant's own edit tool in normal
|
|
225
|
+
// operation: a receipt is stamped by the dispatch hook reacting to the call itself, a
|
|
226
|
+
// leg-completion row by the step that reads a landed result, and a verdict by the command that
|
|
227
|
+
// runs the fixtures. A build leg reaching for `Edit`/`Write` on any of them is not doing its job
|
|
228
|
+
// by another route — it is handing itself the grade — so freezing them costs a well-behaved leg
|
|
229
|
+
// nothing and removes the one channel through which it could attest work it did not do. The
|
|
230
|
+
// carve-out below still covers everything else under this root — the doer's own task board and
|
|
231
|
+
// discovery ledger — because neither lives under any of these three.
|
|
232
|
+
//
|
|
233
|
+
// THE RESULT ENVELOPE IS DELIBERATELY NOT ON THIS LIST, and the reasoning is worth keeping because
|
|
234
|
+
// it looks like it belongs. It is the leg's own claim about its own work, so on the argument above
|
|
235
|
+
// it is the first thing you would freeze. But a result is not evidence ABOUT the leg, it is the
|
|
236
|
+
// leg's PRODUCT — the other half of the envelope port, and the thing that answers the order. The
|
|
237
|
+
// order is unanswered at that moment by construction, so freezing the path would deny every build
|
|
238
|
+
// leg its documented last step, every time, on the first round: measured end to end against a
|
|
239
|
+
// compiled order, with the denial telling the worker to widen a substrate that cannot lift a
|
|
240
|
+
// frozen entry. Nothing else writes an ordinary result either, so there is no fallback. The census
|
|
241
|
+
// that reads it is already built for this: a result alone attests nothing, and can only turn an
|
|
242
|
+
// ALREADY-receipted attempt into a spent one, which spends the forger's own budget. What that does
|
|
243
|
+
// not cover is a leg forging a SIBLING's result, which is a real hole with a different fix — the
|
|
244
|
+
// glob form here cannot say "every result except this order's own", so closing it needs a
|
|
245
|
+
// mechanism rather than one more entry on this list.
|
|
246
|
+
const FROZEN_ATTESTATION = [`${local}/receipts/**`, `${local}/legs.jsonl`, `${local}/t0/verdicts/**`];
|
|
221
247
|
switch (operation) {
|
|
222
248
|
case "execute": case "fix": case "spike":
|
|
223
249
|
// Build legs are the widest window on FROZEN_INTAKE, not an exemption from it: they are the
|
|
@@ -228,7 +254,7 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
|
|
|
228
254
|
return {
|
|
229
255
|
allowed: [...(scope?.allowed_file_substrate || []), `${local}/spikes/**`],
|
|
230
256
|
shared: scope?.shared_substrate || [],
|
|
231
|
-
frozen: [...FROZEN_INTAKE],
|
|
257
|
+
frozen: [...FROZEN_INTAKE, ...FROZEN_ATTESTATION],
|
|
232
258
|
};
|
|
233
259
|
case "analyze":
|
|
234
260
|
return { allowed: [`${spec}/**`, `${local}/**`], frozen: [...FROZEN_INTAKE] };
|
package/kernel/gate.mjs
CHANGED
|
@@ -110,7 +110,7 @@ export const PRESETS = {
|
|
|
110
110
|
"L1a": { decision: "proceed", note: "Orient review — advisory read." },
|
|
111
111
|
"L1a.5": { decision: "proceed", note: "Wiring review — checked by trace-lint." },
|
|
112
112
|
"L1b": { decision: "ask", note: "Board review is where scope is actually decided. Not pre-approvable." },
|
|
113
|
-
"L2": { decision: "proceed", note: "
|
|
113
|
+
"L2": { decision: "proceed", note: "The board facts travel in the gate block itself — green_scopes and hammer_proposals — so a preset answering here is not answering blind. Note there is no board check behind this: the L2 hook was retired into the gate block in v2.0." },
|
|
114
114
|
"L3": { decision: "loop", max_rounds: 3, note: "Loop on FAIL; the breaker ends it." },
|
|
115
115
|
"QA": { decision: "run" },
|
|
116
116
|
"H": { decision: "ask", note: "The cut list changes what ships." },
|
package/kernel/probe/eval.mjs
CHANGED
|
@@ -33,6 +33,7 @@
|
|
|
33
33
|
|
|
34
34
|
import { existsSync, readFileSync, readdirSync } from "node:fs";
|
|
35
35
|
import { join, resolve } from "node:path";
|
|
36
|
+
import { createHash } from "node:crypto";
|
|
36
37
|
import { runArgs } from "../lib/argv.mjs";
|
|
37
38
|
import { resultsDir, scopesDir } from "../lib/paths.mjs";
|
|
38
39
|
|
|
@@ -59,6 +60,58 @@ export function isScoped(cwd, slug) {
|
|
|
59
60
|
catch { return false; }
|
|
60
61
|
}
|
|
61
62
|
|
|
63
|
+
/**
|
|
64
|
+
* Why one T0 citation does not resolve, or null when the kernel can positively confirm it does.
|
|
65
|
+
*
|
|
66
|
+
* Re-checks the two facts a citation can lie about and the schema promises are enforced
|
|
67
|
+
* (`T0Citation`: "the evaluator RECOMPUTES sha256 from disk — a handed hash is never trusted"):
|
|
68
|
+
* that the bytes at `path` hash to the cited `sha256`, and that what those bytes actually say is a
|
|
69
|
+
* green T0 verdict — a PASS/FAIL cannot ride on an artifact that itself recorded red. NOT checked:
|
|
70
|
+
* whether `path` is one of the order's own `payload.t0_artifacts` (see the note on
|
|
71
|
+
* {@link citationProblem} for why that is left to the operator rather than enforced here).
|
|
72
|
+
*
|
|
73
|
+
* FAILS OPEN ON "CANNOT TELL", CLOSED ON "PROVEN WRONG" — and the two are not the same fact. A read
|
|
74
|
+
* that fails for a reason that says something DEFINITE about what is (or isn't) at `path` is proof,
|
|
75
|
+
* same as a hash mismatch: `ENOENT` (nothing there) and `EISDIR` (a directory, never a file — T0
|
|
76
|
+
* verdicts are always files, per `writeArtifact`) both mean no such artifact was ever produced, so
|
|
77
|
+
* both refuse. `verify t0` writes verdict artifacts immutably and never deletes one
|
|
78
|
+
* (`writeArtifact`'s `wx` flag — see kernel/verify/t0.mjs), so that absence or shape mismatch is a
|
|
79
|
+
* positive fact about the citation, not a guess. Anything else a read can fail with — permission
|
|
80
|
+
* denied, a symlink loop, a transient I/O error — says nothing about the citation's honesty, only
|
|
81
|
+
* that THIS MACHINE could not check it just now, so it is not refused on that ground alone: the
|
|
82
|
+
* fail-open discipline this repo's guards already use elsewhere for an unproven bad state.
|
|
83
|
+
*
|
|
84
|
+
* @param {string} cwd - Project root.
|
|
85
|
+
* @param {object} citation - One `T0Citation` (`scope_id`, `path`, `sha256`). Read defensively:
|
|
86
|
+
* `reduce ingest` only reaches this after the enclosing WorkResult passed schema validation, but
|
|
87
|
+
* `probe eval` reaches it over a result file it merely `JSON.parse`s, so a malformed citation must
|
|
88
|
+
* fail this check rather than throw.
|
|
89
|
+
* @returns {(string|null)} A reason phrased for an operator, or null when the file at `path` exists,
|
|
90
|
+
* hashes to the cited `sha256`, and its own `overall` reads "green".
|
|
91
|
+
*/
|
|
92
|
+
function unresolvedCitation(cwd, citation) {
|
|
93
|
+
const rel = typeof citation?.path === "string" ? citation.path : "";
|
|
94
|
+
if (!rel) return "names no artifact path";
|
|
95
|
+
let text;
|
|
96
|
+
try {
|
|
97
|
+
text = readFileSync(resolve(cwd, rel), "utf8");
|
|
98
|
+
} catch (e) {
|
|
99
|
+
if (e.code === "ENOENT") return `cites ${rel}, which does not exist on disk`;
|
|
100
|
+
if (e.code === "EISDIR") return `cites ${rel}, which is a directory, not a T0 verdict file`;
|
|
101
|
+
// EACCES, ELOOP, EIO, EMFILE… — this machine failing to look, not evidence against the
|
|
102
|
+
// citation, so it is not refused on that ground.
|
|
103
|
+
return null;
|
|
104
|
+
}
|
|
105
|
+
const actual = createHash("sha256").update(text).digest("hex");
|
|
106
|
+
const claimed = typeof citation.sha256 === "string" ? citation.sha256.toLowerCase() : "";
|
|
107
|
+
if (actual !== claimed) return `cites ${rel} with sha256 ${citation.sha256}, but the file on disk hashes to ${actual}`;
|
|
108
|
+
let body;
|
|
109
|
+
try { body = JSON.parse(text); }
|
|
110
|
+
catch { return `cites ${rel}, whose bytes match the hash but do not read as a T0 verdict`; }
|
|
111
|
+
if (body?.overall !== "green") return `cites ${rel}, whose own verdict is "${body?.overall ?? "unknown"}", not green`;
|
|
112
|
+
return null;
|
|
113
|
+
}
|
|
114
|
+
|
|
62
115
|
/**
|
|
63
116
|
* Why a verdict cannot stand as its round's judgement on T0 grounds, or null when it can.
|
|
64
117
|
*
|
|
@@ -68,21 +121,40 @@ export function isScoped(cwd, slug) {
|
|
|
68
121
|
* and branched on like any other. It is checked here so the round loop, the resume derivation, the
|
|
69
122
|
* hill and ingest all refuse the same verdict for the same reason.
|
|
70
123
|
*
|
|
71
|
-
*
|
|
72
|
-
*
|
|
124
|
+
* RESOLVES, NOT JUST PRESENT. A citation naming a path and a hash used to be taken on faith: any
|
|
125
|
+
* non-empty `t0_citations[]` passed, whatever it pointed at. Measured live: a PASS citing a path
|
|
126
|
+
* that does not exist on disk, and a PASS citing a real artifact whose own verdict was red, both
|
|
127
|
+
* ingested clean — presence stood in for a re-hash the schema had already promised. Each citation is
|
|
128
|
+
* now re-checked through {@link unresolvedCitation}.
|
|
129
|
+
*
|
|
130
|
+
* WHAT THIS DOES NOT CHECK: whether a citation is drawn from the order's own `payload.t0_artifacts`
|
|
131
|
+
* list. That would need the compiled order, which this function's callers do not equally have —
|
|
132
|
+
* `reduce ingest` holds it, but `probe eval` (and the round-loop/resume/hill readers behind it)
|
|
133
|
+
* knows only (cwd, slug, round), and a scope's T0 attempt can legitimately go green again LATER than
|
|
134
|
+
* whatever list was frozen at compile time (`greenVerdict` already treats "newest green" as
|
|
135
|
+
* authoritative for exactly this reason — see kernel/probe/t0.mjs). Enforcing membership only where
|
|
136
|
+
* the order happens to be on hand would let one channel refuse a citation the other accepts, for
|
|
137
|
+
* evidence that may simply be fresher than the order — worse than leaving it unenforced.
|
|
73
138
|
*
|
|
74
139
|
* @param {string} cwd - Project root.
|
|
75
140
|
* @param {string} slug - Feature slug.
|
|
76
141
|
* @param {object} verdict - The WorkResult's `verdict` block.
|
|
77
|
-
* @returns {(string|null)} The problem, phrased for an operator; null for a
|
|
78
|
-
* unscoped spec, or a block with no PASS/FAIL in it (there is no judgement
|
|
142
|
+
* @returns {(string|null)} The problem, phrased for an operator; null for a verdict whose every
|
|
143
|
+
* citation resolves, an unscoped spec, or a block with no PASS/FAIL in it (there is no judgement
|
|
144
|
+
* to invalidate).
|
|
79
145
|
*/
|
|
80
146
|
export function citationProblem(cwd, slug, verdict) {
|
|
81
147
|
if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
|
|
82
|
-
if (Array.isArray(verdict.t0_citations) && verdict.t0_citations.length) return null;
|
|
83
148
|
if (!isScoped(cwd, slug)) return null;
|
|
84
|
-
|
|
85
|
-
|
|
149
|
+
if (!Array.isArray(verdict.t0_citations) || !verdict.t0_citations.length) {
|
|
150
|
+
return `the ${verdict.overall} verdict cites no T0 artifact, and a verdict on a scoped spec must ` +
|
|
151
|
+
"cite the T0 verdict it re-hashed (the order lists them under payload.t0_artifacts)";
|
|
152
|
+
}
|
|
153
|
+
for (const citation of verdict.t0_citations) {
|
|
154
|
+
const reason = unresolvedCitation(cwd, citation);
|
|
155
|
+
if (reason) return `the ${verdict.overall} verdict ${reason} — a T0 citation is re-hashed from disk, never taken on the handed word`;
|
|
156
|
+
}
|
|
157
|
+
return null;
|
|
86
158
|
}
|
|
87
159
|
|
|
88
160
|
/**
|
package/kernel/reduce/hill.mjs
CHANGED
|
@@ -6,24 +6,158 @@
|
|
|
6
6
|
import { readFileSync, writeFileSync, existsSync, readdirSync, mkdirSync } from "node:fs";
|
|
7
7
|
import { resolve, join } from "node:path";
|
|
8
8
|
import { runArgs } from "../lib/argv.mjs";
|
|
9
|
-
import { scopesDir, hillDir, verdictsDir, resultsDir, discoveryLedger } from "../lib/paths.mjs";
|
|
9
|
+
import { scopesDir, hillDir, verdictsDir, resultsDir, discoveryLedger, localRoot } from "../lib/paths.mjs";
|
|
10
10
|
import { readAllContracts, SCOPE_CONTRACT } from "../lib/contract.mjs";
|
|
11
11
|
import { evalVerdict } from "../probe/eval.mjs";
|
|
12
12
|
import { redBuildRounds } from "../verify/build.mjs";
|
|
13
13
|
|
|
14
|
+
// ---------------------------------------------------------------------------------------------
|
|
15
|
+
// THE DISCOVERY LEDGER, AND WHY ITS ABSENCE IS NOT A ZERO.
|
|
16
|
+
//
|
|
17
|
+
// The ledger arm is the only thing that promotes a scope from UPHILL_UNKNOWN to UPHILL_SOLVED —
|
|
18
|
+
// "the open questions are closed". It used to be read as `scopeUnknowns[id] || 0`, which made
|
|
19
|
+
// "nobody has filed a ledger yet" and "every unknown is closed" the same value, and that value
|
|
20
|
+
// selects the FLATTERING phase. An absent value and a real one must not share a signature; when
|
|
21
|
+
// they do, the run reports progress it has no evidence for. So the count is `null` — not `0` —
|
|
22
|
+
// whenever the ledger could not actually be read and understood, and `null` promotes nothing.
|
|
23
|
+
//
|
|
24
|
+
// THE HEADING IS THE ATTRIBUTION KEY, and it has to match what the ingest step really writes:
|
|
25
|
+
// `## Discovered — <order_id> (<date>)`, where `order_id` is `<slug>/<suffix>`. Two ways that
|
|
26
|
+
// parse used to fail silently, both ending in the same false zero:
|
|
27
|
+
// * It required a COLON between the slug and the suffix. The order-envelope schema pins
|
|
28
|
+
// `order_id` to a slug, a SLASH, and a suffix drawn from `[a-z0-9.-]` — a colon cannot appear
|
|
29
|
+
// in a schema-valid order id at all, so the scope was never captured from any real ledger and
|
|
30
|
+
// every scope on every project read zero unknowns regardless of what was open.
|
|
31
|
+
// * It matched the em dash only, so a heading typed with a plain hyphen contributed nothing.
|
|
32
|
+
// Both are fixed by matching the dash as a class and splitting the order id on its "/", and by
|
|
33
|
+
// resolving the scope against the CONTRACTS ON DISK rather than guessing at the suffix's shape: a
|
|
34
|
+
// build suffix is `<scope>-r<N>-a<M>` and a scoped operation's is `<operation>-<scope>[-r<N>]`, so
|
|
35
|
+
// a pattern that tried to carve the scope out positionally captured the round suffix along with it.
|
|
36
|
+
// ---------------------------------------------------------------------------------------------
|
|
37
|
+
|
|
38
|
+
/** A ledger block heading, naming the order it files: `## Discovered — <slug>/<suffix> (<date>)`. */
|
|
39
|
+
const LEDGER_HEADING = /^##\s+Discovered\s+[—–-]\s+(\S+)/;
|
|
40
|
+
|
|
41
|
+
/**
|
|
42
|
+
* Index the scopes by the filename-safe form `harness compile` puts in an order suffix, so a
|
|
43
|
+
* heading is matched against ids that actually exist rather than parsed speculatively.
|
|
44
|
+
*
|
|
45
|
+
* @param {Array<object>} scopes - Parsed scope contracts.
|
|
46
|
+
* @returns {Map<string,string>} Compile's suffix form → the contract's own `scope_id`.
|
|
47
|
+
*/
|
|
48
|
+
function scopeSuffixIndex(scopes) {
|
|
49
|
+
const ix = new Map();
|
|
50
|
+
for (const s of scopes) {
|
|
51
|
+
const id = String(s?.scope_id || "");
|
|
52
|
+
if (!id) continue;
|
|
53
|
+
// Mirrors `harness compile`'s own normalisation of a scope id into an order suffix.
|
|
54
|
+
const key = id.toLowerCase().replace(/[^a-z0-9.-]/g, "-").replace(/^[^a-z0-9]+/, "");
|
|
55
|
+
if (key) ix.set(key, id);
|
|
56
|
+
}
|
|
57
|
+
return ix;
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
/**
|
|
61
|
+
* The scope a ledger heading's order belongs to, or `null` when it belongs to none.
|
|
62
|
+
*
|
|
63
|
+
* Operation-level dispatches (orient, analyze, wire, evaluate, hunt, hammer) carry no scope in
|
|
64
|
+
* their order id by construction, so "no scope" is a real and common answer here, not a parse
|
|
65
|
+
* failure — their rows are feature-wide and are deliberately credited to nobody.
|
|
66
|
+
*
|
|
67
|
+
* @param {string} orderId - The order id as the heading names it (`<slug>/<suffix>`).
|
|
68
|
+
* @param {Map<string,string>} ix - The index from `scopeSuffixIndex`.
|
|
69
|
+
* @returns {string|null} The owning contract's `scope_id`, or null.
|
|
70
|
+
*/
|
|
71
|
+
function scopeOfHeading(orderId, ix) {
|
|
72
|
+
const slash = orderId.indexOf("/");
|
|
73
|
+
const suffix = slash === -1 ? orderId : orderId.slice(slash + 1);
|
|
74
|
+
const core = suffix.replace(/-r\d+(?:-a\d+)?$/, "");
|
|
75
|
+
// Longest match wins, so a project holding both `pages` and `pages-admin` attributes each
|
|
76
|
+
// heading to the scope it actually names rather than to whichever was indexed first.
|
|
77
|
+
let best = null;
|
|
78
|
+
let bestLen = -1;
|
|
79
|
+
for (const [key, id] of ix) {
|
|
80
|
+
if ((core === key || core.endsWith(`-${key}`)) && key.length > bestLen) { best = id; bestLen = key.length; }
|
|
81
|
+
}
|
|
82
|
+
return best;
|
|
83
|
+
}
|
|
84
|
+
|
|
85
|
+
/**
|
|
86
|
+
* Open unknowns (`~` rows) per scope, read from the discovery ledger.
|
|
87
|
+
*
|
|
88
|
+
* @param {string} ledgerPath - Path to the run's discovery ledger.
|
|
89
|
+
* @param {Array<object>} scopes - Parsed scope contracts, used to attribute each block.
|
|
90
|
+
* @returns {Object<string,number>|null} Counts per `scope_id` — a scope with a block and no open
|
|
91
|
+
* row is a genuine `0`. `null` means the ledger was NOT read: it is absent, unreadable, or
|
|
92
|
+
* nothing in it could be attributed to a scope. `null` is never treated as zero, because "we
|
|
93
|
+
* understood none of this file" is not evidence that nothing is open.
|
|
94
|
+
*/
|
|
95
|
+
function ledgerUnknowns(ledgerPath, scopes) {
|
|
96
|
+
if (!existsSync(ledgerPath)) return null;
|
|
97
|
+
let text;
|
|
98
|
+
try { text = readFileSync(ledgerPath, "utf8"); } catch { return null; }
|
|
99
|
+
|
|
100
|
+
const ix = scopeSuffixIndex(scopes);
|
|
101
|
+
const counts = {};
|
|
102
|
+
let headings = 0;
|
|
103
|
+
let attributed = 0;
|
|
104
|
+
let current = null;
|
|
105
|
+
|
|
106
|
+
for (const line of text.split("\n")) {
|
|
107
|
+
if (line.startsWith("## ")) {
|
|
108
|
+
// Reset on EVERY heading, matched or not. Carrying the previous block's scope across an
|
|
109
|
+
// unrecognised heading would credit one scope's open rows to another.
|
|
110
|
+
current = null;
|
|
111
|
+
const m = line.match(LEDGER_HEADING);
|
|
112
|
+
if (!m) continue;
|
|
113
|
+
headings++;
|
|
114
|
+
const id = scopeOfHeading(m[1], ix);
|
|
115
|
+
if (id) { attributed++; current = id; counts[id] = counts[id] || 0; }
|
|
116
|
+
continue;
|
|
117
|
+
}
|
|
118
|
+
if (current && line.startsWith("~ ")) counts[current]++;
|
|
119
|
+
}
|
|
120
|
+
|
|
121
|
+
// A ledger whose blocks we could not attribute to a single scope tells us nothing per scope.
|
|
122
|
+
// Reporting that as all-zero is the same false signature the colon-vs-slash bug produced.
|
|
123
|
+
if (headings === 0 || attributed === 0) return null;
|
|
124
|
+
return counts;
|
|
125
|
+
}
|
|
126
|
+
|
|
127
|
+
/**
|
|
128
|
+
* The phase currently recorded in a committed hill shard.
|
|
129
|
+
*
|
|
130
|
+
* @param {string} hDir - The committed hill directory.
|
|
131
|
+
* @param {string} id - Scope id.
|
|
132
|
+
* @returns {string|null} The recorded phase, or null when no shard exists for that scope.
|
|
133
|
+
*/
|
|
134
|
+
function committedPhase(hDir, id) {
|
|
135
|
+
const p = join(hDir, `${id}.yml`);
|
|
136
|
+
if (!existsSync(p)) return null;
|
|
137
|
+
try { return readFileSync(p, "utf8").match(/^phase:\s*(\S+)/m)?.[1] ?? null; } catch { return null; }
|
|
138
|
+
}
|
|
139
|
+
|
|
14
140
|
/**
|
|
15
141
|
* Derive and write the hill phase for all scopes mechanically based on T0, T1, and ledger facts.
|
|
16
142
|
*
|
|
17
143
|
* The derived phase follows these progression rules (facts move dots, not authors):
|
|
18
|
-
* - UPHILL_UNKNOWN: open unknowns > 0 in the ledger for this scope
|
|
19
|
-
*
|
|
144
|
+
* - UPHILL_UNKNOWN: open unknowns > 0 in the ledger for this scope — and the floor the scope sits
|
|
145
|
+
* at whenever the ledger has not answered at all, which is where every run legitimately begins
|
|
146
|
+
* - UPHILL_SOLVED: the ledger was read and reports zero open unknowns, no T0-green yet
|
|
20
147
|
* - DOWNHILL_EXECUTION: ≥1 T0-green in a round whose build gate is not red; T1/seesaw pending
|
|
21
148
|
* - FINISHED: T1 PASS ∧ seesaw green
|
|
22
149
|
*
|
|
23
150
|
* @param {string} cwd - The project root directory.
|
|
24
151
|
* @param {string} slug - The feature slug being built.
|
|
25
|
-
* @returns {Array<{scope_id: string, phase: string, changed: boolean
|
|
26
|
-
*
|
|
152
|
+
* @returns {Array<{scope_id: string, phase: (string|null), changed: boolean, derived: boolean,
|
|
153
|
+
* unknowns: (number|null), reason?: string}>} A report of all scopes processed. `derived` says
|
|
154
|
+
* whether the phase was computed from run evidence at all: when it is false the phase is
|
|
155
|
+
* whatever the committed shard already records (or null when there is none), nothing was
|
|
156
|
+
* written, and `reason` names why. `unknowns` is the ledger count behind the phase, or null
|
|
157
|
+
* when the ledger could not be read — the two are deliberately distinguishable in the output as
|
|
158
|
+
* well as in the derivation.
|
|
159
|
+
* Side effects: writes to `shapeup/<slug>/hill/<scope-id>.yml` for each scope, EXCEPT on the
|
|
160
|
+
* refusal path below, which writes nothing at all.
|
|
27
161
|
*/
|
|
28
162
|
export function deriveHill(cwd, slug) {
|
|
29
163
|
const scopes = readAllContracts(scopesDir(cwd, slug), SCOPE_CONTRACT).map((x) => x.contract);
|
|
@@ -31,6 +165,44 @@ export function deriveHill(cwd, slug) {
|
|
|
31
165
|
const ledgerPath = discoveryLedger(cwd, slug);
|
|
32
166
|
const hDir = hillDir(cwd, slug);
|
|
33
167
|
|
|
168
|
+
// -------------------------------------------------------------------------------------------
|
|
169
|
+
// A DERIVATION THAT CANNOT SEE THE RUN TRACE MUST NOT WRITE THE DELIVERABLE.
|
|
170
|
+
//
|
|
171
|
+
// This function reads one tier and writes the other. Every input below — T0 verdicts, the EVAL
|
|
172
|
+
// results, the round build gates, the discovery ledger — lives in the gitignored run trace,
|
|
173
|
+
// while the shards it writes are committed and outlive it: after a ship the run trace is cleaned
|
|
174
|
+
// up and the shards are the ONLY surviving record of where the work got to. So when the run
|
|
175
|
+
// trace for this slug is not on disk, every input is provably absent, and anything derived from
|
|
176
|
+
// that is derived from nothing.
|
|
177
|
+
//
|
|
178
|
+
// This is not a hypothetical. Pulling a branch mid-run is a supported state: the puller gets the
|
|
179
|
+
// committed spec, scopes and shards and no run trace of their own. Their first launch used to
|
|
180
|
+
// re-derive every scope from the empty set and overwrite the shards with the result — losing
|
|
181
|
+
// committed history rather than misreporting it.
|
|
182
|
+
//
|
|
183
|
+
// WHY REFUSE RATHER THAN WRITE AN "UNKNOWN" PHASE. Writing anything here destroys the record
|
|
184
|
+
// just as thoroughly; a shard that says "I could not look" has still replaced the one that said
|
|
185
|
+
// FINISHED, and the phase enum is a committed data format that a reader parses as current
|
|
186
|
+
// status. Not writing already expresses "no opinion" exactly, and it needs no new enum value.
|
|
187
|
+
//
|
|
188
|
+
// WHY THE CONDITION IS THE TIER'S EXISTENCE AND NOT "the phase would go down". Moving a dot
|
|
189
|
+
// backwards is correct and must stay possible — this is a pure function of the artifacts present
|
|
190
|
+
// and reports what they currently support, in both directions. A guard phrased as "never lower a
|
|
191
|
+
// phase" would quietly turn a derived value into a high-water mark, which is a different defect
|
|
192
|
+
// wearing this one's clothes. The condition is the narrowest one that is positively provable:
|
|
193
|
+
// the tier holding every input is not there.
|
|
194
|
+
// -------------------------------------------------------------------------------------------
|
|
195
|
+
if (!existsSync(localRoot(cwd, slug))) {
|
|
196
|
+
return scopes.map((s) => ({
|
|
197
|
+
scope_id: s.scope_id,
|
|
198
|
+
phase: committedPhase(hDir, s.scope_id),
|
|
199
|
+
changed: false,
|
|
200
|
+
derived: false,
|
|
201
|
+
unknowns: null,
|
|
202
|
+
reason: "local-run-trace-absent",
|
|
203
|
+
}));
|
|
204
|
+
}
|
|
205
|
+
|
|
34
206
|
if (!existsSync(hDir)) mkdirSync(hDir, { recursive: true });
|
|
35
207
|
|
|
36
208
|
// 1. Check if T1 Evaluation passed — the LATEST evaluate round's verdict, read the same way
|
|
@@ -95,28 +267,20 @@ export function deriveHill(cwd, slug) {
|
|
|
95
267
|
}
|
|
96
268
|
}
|
|
97
269
|
|
|
98
|
-
// 3. Ledger unknowns per scope
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
let currentScope = null;
|
|
103
|
-
for (const line of lines) {
|
|
104
|
-
const m = line.match(/^## Discovered — .*?:([\w.-]+)-a\d+/);
|
|
105
|
-
if (m) {
|
|
106
|
-
currentScope = m[1];
|
|
107
|
-
}
|
|
108
|
-
if (currentScope && line.startsWith("~ ")) {
|
|
109
|
-
scopeUnknowns[currentScope] = (scopeUnknowns[currentScope] || 0) + 1;
|
|
110
|
-
}
|
|
111
|
-
}
|
|
112
|
-
}
|
|
113
|
-
|
|
270
|
+
// 3. Ledger unknowns per scope — `null` for every scope when the ledger itself was not readable
|
|
271
|
+
// or nothing in it named a scope (see `ledgerUnknowns`). Only a real count can promote.
|
|
272
|
+
const scopeUnknowns = ledgerUnknowns(ledgerPath, scopes);
|
|
273
|
+
|
|
114
274
|
const report = [];
|
|
115
275
|
for (const s of scopes) {
|
|
116
276
|
const id = s.scope_id;
|
|
117
277
|
const t0 = t0Facts[id] || { hasGreen: false, seesawGreen: false };
|
|
118
|
-
|
|
119
|
-
|
|
278
|
+
// `null` = the ledger did not answer; a number = it did. `|| 0` collapsed the two.
|
|
279
|
+
const unknowns = scopeUnknowns === null ? null : (scopeUnknowns[id] || 0);
|
|
280
|
+
|
|
281
|
+
// UPHILL_UNKNOWN is the floor, and the honest answer whenever nothing has promoted a scope off
|
|
282
|
+
// it — including before Orient has filed anything, which is where every run legitimately
|
|
283
|
+
// starts. Only an ANSWERED count of zero promotes to UPHILL_SOLVED; `null` never does.
|
|
120
284
|
let phase = "UPHILL_UNKNOWN";
|
|
121
285
|
if (t1Pass && t0.hasGreen && t0.seesawGreen) {
|
|
122
286
|
phase = "FINISHED";
|
|
@@ -125,7 +289,7 @@ export function deriveHill(cwd, slug) {
|
|
|
125
289
|
} else if (unknowns === 0) {
|
|
126
290
|
phase = "UPHILL_SOLVED";
|
|
127
291
|
}
|
|
128
|
-
|
|
292
|
+
|
|
129
293
|
const yaml = `scope_id: ${id}\nphase: ${phase}\n`;
|
|
130
294
|
const out = join(hDir, `${id}.yml`);
|
|
131
295
|
let changed = false;
|
|
@@ -133,7 +297,7 @@ export function deriveHill(cwd, slug) {
|
|
|
133
297
|
writeFileSync(out, yaml);
|
|
134
298
|
changed = true;
|
|
135
299
|
}
|
|
136
|
-
report.push({ scope_id: id, phase, changed });
|
|
300
|
+
report.push({ scope_id: id, phase, changed, derived: true, unknowns });
|
|
137
301
|
}
|
|
138
302
|
return report;
|
|
139
303
|
}
|
|
@@ -156,5 +320,16 @@ export async function cli(rawArgv) {
|
|
|
156
320
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
157
321
|
const cwd = resolve(args.cwd || process.cwd());
|
|
158
322
|
const report = deriveHill(cwd, args.slug);
|
|
323
|
+
// A refusal that is visible only as a missing write reads exactly like a derivation that
|
|
324
|
+
// happened to agree with what was already on disk, so say it out loud. It is not an error —
|
|
325
|
+
// the caller runs this advisorily several times a run, and declining to derive from nothing is
|
|
326
|
+
// the correct outcome, not a failure — so the exit code stays 0 and the warning goes to stderr.
|
|
327
|
+
const underived = report.filter((r) => r.derived === false);
|
|
328
|
+
if (underived.length) {
|
|
329
|
+
console.error(
|
|
330
|
+
`hill: derived nothing for ${underived.length} scope(s) (${underived[0].reason}) — ` +
|
|
331
|
+
`the run trace this phase is derived from is not on disk, so the committed shards were left as they are.`,
|
|
332
|
+
);
|
|
333
|
+
}
|
|
159
334
|
console.log(JSON.stringify(report, null, 2));
|
|
160
335
|
}
|
package/kernel/reduce/ship.mjs
CHANGED
|
@@ -32,7 +32,7 @@ import { runArgs } from "../lib/argv.mjs";
|
|
|
32
32
|
import {
|
|
33
33
|
report as reportPath, tasksDir, verdictsDir, trials, evaluationDir, qaDir,
|
|
34
34
|
roundLedger, discoveryLedger, receipt as receiptPath, harnessRun, relShared,
|
|
35
|
-
activeOrder,
|
|
35
|
+
activeOrder, runArgsPath, readReceipt, runIdFromReceipt,
|
|
36
36
|
} from "../lib/paths.mjs";
|
|
37
37
|
import { readTrials } from "../verify/t0.mjs";
|
|
38
38
|
import { ratchetReport } from "../probe/stats.mjs";
|
|
@@ -398,9 +398,51 @@ export const ARGV_SPEC = {
|
|
|
398
398
|
* @returns {(Promise<void>|void)} Settles when the subcommand has written its output; most paths
|
|
399
399
|
* call `process.exit()` with the subcommand's documented code rather than returning.
|
|
400
400
|
*/
|
|
401
|
+
/**
|
|
402
|
+
* Did THIS run skip evaluation? Read from the run's own recorded arguments, and believed only when
|
|
403
|
+
* the record can be shown to belong to this run.
|
|
404
|
+
*
|
|
405
|
+
* WHY THE KERNEL ASKS AT ALL. A run launched with `--no-eval` verified nothing, and the protocol
|
|
406
|
+
* has always said such a run ships as `not-evaluated`, "recorded plainly — never silently
|
|
407
|
+
* upgraded". That promise lived entirely inside the orchestrator's control flow, where one
|
|
408
|
+
* assignment downstream of the branch restores the old behaviour with every check still green —
|
|
409
|
+
* measured, not supposed. The gate block tells the human the truth either way; the committed report
|
|
410
|
+
* a teammate inherits on `git pull` is the artifact that was lying, so the refusal belongs at the
|
|
411
|
+
* writer of that artifact, where a future orchestrator edit cannot reach it.
|
|
412
|
+
*
|
|
413
|
+
* POSITIVELY PROVEN OR NOT AT ALL. The record is written at launch, later than the receipt that
|
|
414
|
+
* mints the run key, so a freshly opened run can still find the PREVIOUS run's record on disk. A
|
|
415
|
+
* stale flag would refuse a ship that verified everything — a worse failure than the one this
|
|
416
|
+
* closes. The flag therefore counts only when the record names the same run the receipt does;
|
|
417
|
+
* a missing, unreadable or differently-keyed record proves nothing and permits.
|
|
418
|
+
*
|
|
419
|
+
* @param {string} cwd - Project root.
|
|
420
|
+
* @param {string} slug - Feature slug.
|
|
421
|
+
* @returns {boolean} True only when this run's own record says evaluation was skipped.
|
|
422
|
+
*/
|
|
423
|
+
function evalWasSkipped(cwd, slug) {
|
|
424
|
+
try {
|
|
425
|
+
const record = JSON.parse(readFileSync(runArgsPath(cwd, slug), "utf8"));
|
|
426
|
+
if (record?.noEval !== true) return false;
|
|
427
|
+
const mine = runIdFromReceipt(readReceipt(receiptPath(cwd, slug)));
|
|
428
|
+
return Boolean(mine) && record.runId === mine;
|
|
429
|
+
} catch { return false; }
|
|
430
|
+
}
|
|
431
|
+
|
|
432
|
+
/** The verdict values that assert the feature was graded and passed. */
|
|
433
|
+
const PASSING = new Set(["PASS", "pass"]);
|
|
434
|
+
|
|
401
435
|
export async function cli(rawArgv) {
|
|
402
436
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
403
437
|
const cwd = args.cwd || process.cwd();
|
|
438
|
+
if (PASSING.has(String(args.verdict ?? "")) && evalWasSkipped(cwd, args.slug)) {
|
|
439
|
+
console.error(
|
|
440
|
+
"✋ reduce ship: this run was launched with --no-eval, so nothing graded it — refusing to freeze a report " +
|
|
441
|
+
`that says ${args.verdict}. Ship it as --verdict not-evaluated, which the report, the ledger and the ` +
|
|
442
|
+
"sign-off block all carry.",
|
|
443
|
+
);
|
|
444
|
+
process.exit(3);
|
|
445
|
+
}
|
|
404
446
|
const { markdown, path } = generate({ cwd, slug: args.slug, verdict: args.verdict, qa: args.qa });
|
|
405
447
|
if (args.stdout) {
|
|
406
448
|
process.stdout.write(markdown);
|
|
@@ -2846,9 +2846,10 @@
|
|
|
2846
2846
|
"type": "string",
|
|
2847
2847
|
"enum": [
|
|
2848
2848
|
"pass",
|
|
2849
|
-
"fail"
|
|
2849
|
+
"fail",
|
|
2850
|
+
"not-evaluated"
|
|
2850
2851
|
],
|
|
2851
|
-
"description": "shipped | ok: the round or run's EVAL verdict, lowercased from spec-evaluator's PASS|FAIL."
|
|
2852
|
+
"description": "shipped | ok: the round or run's EVAL verdict, lowercased from spec-evaluator's PASS|FAIL — or `not-evaluated` when a `shipped` return closed a `--no-eval` run, which reaches SHIP without ever dispatching the judge. Recorded plainly, never upgraded to `pass` (references/protocol.md)."
|
|
2852
2853
|
},
|
|
2853
2854
|
"rounds_used": {
|
|
2854
2855
|
"type": "integer"
|
package/package.json
CHANGED
|
@@ -48,12 +48,16 @@ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" reduce graph --slug <slug>
|
|
|
48
48
|
```
|
|
49
49
|
|
|
50
50
|
**For a committed-only slug (`hasCommitted && !hasLocal`), NEVER call `reduce hill`.**
|
|
51
|
-
`deriveHill()` folds whatever T0 verdicts
|
|
52
|
-
none
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
51
|
+
`deriveHill()` folds whatever T0 verdicts, evaluation results and ledger rows exist in the LOCAL
|
|
52
|
+
run trace; a committed-only pitch has none, because that trace was cleaned up after shipping.
|
|
53
|
+
|
|
54
|
+
The runtime no longer takes that silence for an answer: a derivation whose run trace is absent
|
|
55
|
+
declines to write anything and reports every scope as underived (`derived: false`), so the
|
|
56
|
+
committed shards survive a call made in error. Treat that as a backstop, not a licence — an
|
|
57
|
+
underived report is not a refresh, and rendering it as one would show a pitch's true `FINISHED`
|
|
58
|
+
as freshly confirmed when nothing confirmed it. Read `shapeup/<slug>/hill/*.yml` for those slugs
|
|
59
|
+
exactly as committed, and render them with the archived state the template already implements
|
|
60
|
+
(see below). Same reasoning applies to `reduce graph` — there is no local trace to append from.
|
|
57
61
|
|
|
58
62
|
## Reading the hill shards
|
|
59
63
|
|
|
@@ -389,7 +389,7 @@ feature) and the failing step is compiled into round r+1's orders as `payload.bu
|
|
|
389
389
|
means L0 pinned no run command and the profile names no probe — the round proceeds over an
|
|
390
390
|
unproven build, and the block says so.
|
|
391
391
|
|
|
392
|
-
Under `--interactive` / `--auto`, the
|
|
392
|
+
Under `--interactive` / `--auto`, the gate block carries the board facts — `green_scopes`, `hammer_proposals` — and requires explicit PO approval to proceed. Under `--unattended` the answer set resolves it. **There is no hook behind this gate**: the `PreToolUse` warning it used to carry was retired into the block in v2.0, so what an unfinished board costs you here is a human reading the numbers, not a machine refusing the call.
|
|
393
393
|
|
|
394
394
|
---
|
|
395
395
|
|
|
@@ -1596,7 +1596,11 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1596
1596
|
`feature; this one does not build or launch. Round ${round + 1} fixes the gate's failing step.`);
|
|
1597
1597
|
} else if (args.noEval) {
|
|
1598
1598
|
log("EVAL — skipped (--no-eval)");
|
|
1599
|
-
|
|
1599
|
+
// references/protocol.md's own words for this, twice: a run --no-eval ships records
|
|
1600
|
+
// `not-evaluated` — "recorded plainly — never silently upgraded [to `pass`]". Nothing verified
|
|
1601
|
+
// the feature beyond task-executor's own per-AC evidence checks, and the report this run
|
|
1602
|
+
// freezes at GATE L4 has to say that as plainly as the gate block already does.
|
|
1603
|
+
verdict = "not-evaluated";
|
|
1600
1604
|
} else {
|
|
1601
1605
|
const e = await worker({
|
|
1602
1606
|
skill: "spec-evaluator", operation: "evaluate", schema: EVAL, phase: "Eval", label: `eval:r${round}`,
|
|
@@ -1647,14 +1651,16 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1647
1651
|
const g3 = await crossGate("L3", "Eval", ["loop", "stop", "ask"], { round, verdict, build_gate: buildGate });
|
|
1648
1652
|
if (g3.stop) return await withWarnings(g3.stop);
|
|
1649
1653
|
|
|
1650
|
-
|
|
1654
|
+
// "not-evaluated" (--no-eval) ships exactly like "pass" — protocol.md: the run "goes straight to
|
|
1655
|
+
// SHIP" once EVAL is skipped, never spends another round waiting on a verdict nobody is producing.
|
|
1656
|
+
if (verdict === "pass" || verdict === "not-evaluated") break; // → QA → GATE H → ship
|
|
1651
1657
|
if (g3.decision === "stop" || round >= maxRounds) {
|
|
1652
1658
|
return await withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
|
|
1653
1659
|
}
|
|
1654
1660
|
round += 1;
|
|
1655
1661
|
}
|
|
1656
1662
|
|
|
1657
|
-
if (verdict !== "pass") {
|
|
1663
|
+
if (verdict !== "pass" && verdict !== "not-evaluated") {
|
|
1658
1664
|
return await withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
|
|
1659
1665
|
}
|
|
1660
1666
|
|
|
@@ -1692,7 +1698,12 @@ if (h.verdict === "cannot-ship") {
|
|
|
1692
1698
|
if (g.stop) return await withWarnings(g.stop);
|
|
1693
1699
|
}
|
|
1694
1700
|
|
|
1695
|
-
|
|
1701
|
+
// The frozen report's own verdict line has to say what actually happened — a `--no-eval` run
|
|
1702
|
+
// verified nothing beyond task-executor's per-AC checks, and `reduce ship` already accepts
|
|
1703
|
+
// "not-evaluated" as a real verdict (its own usage string, and `generate()`'s no-artifact
|
|
1704
|
+
// default). Hardcoding PASS here is exactly the silent upgrade protocol.md's Rules forbid.
|
|
1705
|
+
const shipVerdict = verdict === "not-evaluated" ? "not-evaluated" : "PASS";
|
|
1706
|
+
const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${qaRan ? "run" : "skipped"}`, "Ship", "ship-report");
|
|
1696
1707
|
await advisory(`report export --slug ${slug}`, "Ship", "export-run");
|
|
1697
1708
|
// The run's own concurrency, printed once where the records are complete and before the next run
|
|
1698
1709
|
// supersedes the trace. It is a projection over `receipts/dispatch.jsonl` and `legs.jsonl`, so it
|
|
@@ -1707,7 +1718,10 @@ const ALL_DIMS = ["spec-conformance", "tdd-surface", "integration", "completenes
|
|
|
1707
1718
|
|
|
1708
1719
|
return await withWarnings({
|
|
1709
1720
|
status: "shipped",
|
|
1710
|
-
verdict
|
|
1721
|
+
// The real verdict this run reached — "pass" or, over a --no-eval run, "not-evaluated". GATE L4's
|
|
1722
|
+
// own sign-off block reads this field verbatim (SKILL.md Step 4); hardcoding "pass" here told a
|
|
1723
|
+
// human answering that gate the run was graded when it never was.
|
|
1724
|
+
verdict,
|
|
1711
1725
|
rounds_used: round,
|
|
1712
1726
|
dims_not_evaluated: ALL_DIMS.filter((d) => !evalDims.includes(d)),
|
|
1713
1727
|
qa_findings: qaFindings,
|