shapeup-sdlc 3.7.1 → 3.7.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.7.1",
4
+ "version": "3.7.2",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/README.md CHANGED
@@ -41,9 +41,14 @@ runs over unfinished tasks rather than denying it. The board is local to the mac
41
41
  harness — see [ADR-0001](docs/design/adr/0001-consumer-file-organization.md).)
42
42
 
43
43
  **2. Progress is measured, not claimed.** A scope counts as built only when `t0-verify` runs
44
- its fixtures, a DB probe, and the seesaw, and writes an artifact to disk. The evaluator must
45
- cite that artifact and re-hashes it itself; hill phase is derived from those facts, so no
46
- worker can self-report confidence.
44
+ its fixtures and its DB probe and writes an artifact to disk — with each command's exit code,
45
+ its captured output, and whether it ran at all. The evaluator must cite that artifact, and the
46
+ hill phase is derived from artifacts rather than from a worker's own account of its progress.
47
+ Two limits, stated here because the point of this section is that a claim without a mechanism
48
+ behind it is the thing this harness exists to prevent: the **seesaw** regression arm is declared
49
+ and not yet wired (no run writes its registry — wiring it is an open Betting Table decision), and
50
+ nothing in the runtime **re-hashes** the citation the evaluator is instructed to re-hash. Both
51
+ are open items in `shapeup/knowledge-base/harness-defects.md`, not shipped guarantees.
47
52
  → *Prevents: "done" asserted with nothing behind it.*
48
53
 
49
54
  **3. Parallel work can't corrupt shared state.** Each scope gets a write-whitelist of files
@@ -300,7 +305,9 @@ These hold across the harness and are the reason it stays predictable:
300
305
  count events and neither can notice a single round running for half an hour — tripping it routes
301
306
  to GATE H, so a run out of time ships what is green instead of being killed and shipping nothing.
302
307
  - **Hill phase is mechanical, never self-reported** — derived only from T0/T1/seesaw facts, closing
303
- the self-reported-confidence risk outright.
308
+ the self-reported-confidence risk — with one gap on record: a scope with no discovery
309
+ ledger derives the same phase as one whose unknowns are all closed, so absence still reads
310
+ as progress on that one arm.
304
311
  - **One writer per shared file** — every board/ledger/verdict write goes through
305
312
  `harness reduce ingest`; workers return data and never touch shared state.
306
313
  - **Traceability is oracle-checked, opt-in** — `harness verify trace` verifies covers-closure and
@@ -218,6 +218,32 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
218
218
  // in its `allowed` globs, so only a `frozen` entry denies the write. `translate` is the one
219
219
  // operation that legitimately rewrites a pitch, and it writes the COMMITTED copy, not this one.
220
220
  const FROZEN_INTAKE = [`${local}/intake.md`, `${local}/breadboard.md`];
221
+ // THE LEG'S OWN ATTESTATION. A dispatch receipt, a leg-completion row and a T0 verdict exist to
222
+ // answer "did a worker actually run, and what did it measure" from evidence the leg being judged
223
+ // does not control — the whole reason to prefer them over a compiled order, which is writable by
224
+ // that same hand. None of the three is ever produced by this assistant's own edit tool in normal
225
+ // operation: a receipt is stamped by the dispatch hook reacting to the call itself, a
226
+ // leg-completion row by the step that reads a landed result, and a verdict by the command that
227
+ // runs the fixtures. A build leg reaching for `Edit`/`Write` on any of them is not doing its job
228
+ // by another route — it is handing itself the grade — so freezing them costs a well-behaved leg
229
+ // nothing and removes the one channel through which it could attest work it did not do. The
230
+ // carve-out below still covers everything else under this root — the doer's own task board and
231
+ // discovery ledger — because neither lives under any of these three.
232
+ //
233
+ // THE RESULT ENVELOPE IS DELIBERATELY NOT ON THIS LIST, and the reasoning is worth keeping because
234
+ // it looks like it belongs. It is the leg's own claim about its own work, so on the argument above
235
+ // it is the first thing you would freeze. But a result is not evidence ABOUT the leg, it is the
236
+ // leg's PRODUCT — the other half of the envelope port, and the thing that answers the order. The
237
+ // order is unanswered at that moment by construction, so freezing the path would deny every build
238
+ // leg its documented last step, every time, on the first round: measured end to end against a
239
+ // compiled order, with the denial telling the worker to widen a substrate that cannot lift a
240
+ // frozen entry. Nothing else writes an ordinary result either, so there is no fallback. The census
241
+ // that reads it is already built for this: a result alone attests nothing, and can only turn an
242
+ // ALREADY-receipted attempt into a spent one, which spends the forger's own budget. What that does
243
+ // not cover is a leg forging a SIBLING's result, which is a real hole with a different fix — the
244
+ // glob form here cannot say "every result except this order's own", so closing it needs a
245
+ // mechanism rather than one more entry on this list.
246
+ const FROZEN_ATTESTATION = [`${local}/receipts/**`, `${local}/legs.jsonl`, `${local}/t0/verdicts/**`];
221
247
  switch (operation) {
222
248
  case "execute": case "fix": case "spike":
223
249
  // Build legs are the widest window on FROZEN_INTAKE, not an exemption from it: they are the
@@ -228,7 +254,7 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
228
254
  return {
229
255
  allowed: [...(scope?.allowed_file_substrate || []), `${local}/spikes/**`],
230
256
  shared: scope?.shared_substrate || [],
231
- frozen: [...FROZEN_INTAKE],
257
+ frozen: [...FROZEN_INTAKE, ...FROZEN_ATTESTATION],
232
258
  };
233
259
  case "analyze":
234
260
  return { allowed: [`${spec}/**`, `${local}/**`], frozen: [...FROZEN_INTAKE] };
package/kernel/gate.mjs CHANGED
@@ -110,7 +110,7 @@ export const PRESETS = {
110
110
  "L1a": { decision: "proceed", note: "Orient review — advisory read." },
111
111
  "L1a.5": { decision: "proceed", note: "Wiring review — checked by trace-lint." },
112
112
  "L1b": { decision: "ask", note: "Board review is where scope is actually decided. Not pre-approvable." },
113
- "L2": { decision: "proceed", note: "Board-green is verified by hook, not by opinion." },
113
+ "L2": { decision: "proceed", note: "The board facts travel in the gate block itself — green_scopes and hammer_proposals — so a preset answering here is not answering blind. Note there is no board check behind this: the L2 hook was retired into the gate block in v2.0." },
114
114
  "L3": { decision: "loop", max_rounds: 3, note: "Loop on FAIL; the breaker ends it." },
115
115
  "QA": { decision: "run" },
116
116
  "H": { decision: "ask", note: "The cut list changes what ships." },
@@ -33,6 +33,7 @@
33
33
 
34
34
  import { existsSync, readFileSync, readdirSync } from "node:fs";
35
35
  import { join, resolve } from "node:path";
36
+ import { createHash } from "node:crypto";
36
37
  import { runArgs } from "../lib/argv.mjs";
37
38
  import { resultsDir, scopesDir } from "../lib/paths.mjs";
38
39
 
@@ -59,6 +60,58 @@ export function isScoped(cwd, slug) {
59
60
  catch { return false; }
60
61
  }
61
62
 
63
+ /**
64
+ * Why one T0 citation does not resolve, or null when the kernel can positively confirm it does.
65
+ *
66
+ * Re-checks the two facts a citation can lie about and the schema promises are enforced
67
+ * (`T0Citation`: "the evaluator RECOMPUTES sha256 from disk — a handed hash is never trusted"):
68
+ * that the bytes at `path` hash to the cited `sha256`, and that what those bytes actually say is a
69
+ * green T0 verdict — a PASS/FAIL cannot ride on an artifact that itself recorded red. NOT checked:
70
+ * whether `path` is one of the order's own `payload.t0_artifacts` (see the note on
71
+ * {@link citationProblem} for why that is left to the operator rather than enforced here).
72
+ *
73
+ * FAILS OPEN ON "CANNOT TELL", CLOSED ON "PROVEN WRONG" — and the two are not the same fact. A read
74
+ * that fails for a reason that says something DEFINITE about what is (or isn't) at `path` is proof,
75
+ * same as a hash mismatch: `ENOENT` (nothing there) and `EISDIR` (a directory, never a file — T0
76
+ * verdicts are always files, per `writeArtifact`) both mean no such artifact was ever produced, so
77
+ * both refuse. `verify t0` writes verdict artifacts immutably and never deletes one
78
+ * (`writeArtifact`'s `wx` flag — see kernel/verify/t0.mjs), so that absence or shape mismatch is a
79
+ * positive fact about the citation, not a guess. Anything else a read can fail with — permission
80
+ * denied, a symlink loop, a transient I/O error — says nothing about the citation's honesty, only
81
+ * that THIS MACHINE could not check it just now, so it is not refused on that ground alone: the
82
+ * fail-open discipline this repo's guards already use elsewhere for an unproven bad state.
83
+ *
84
+ * @param {string} cwd - Project root.
85
+ * @param {object} citation - One `T0Citation` (`scope_id`, `path`, `sha256`). Read defensively:
86
+ * `reduce ingest` only reaches this after the enclosing WorkResult passed schema validation, but
87
+ * `probe eval` reaches it over a result file it merely `JSON.parse`s, so a malformed citation must
88
+ * fail this check rather than throw.
89
+ * @returns {(string|null)} A reason phrased for an operator, or null when the file at `path` exists,
90
+ * hashes to the cited `sha256`, and its own `overall` reads "green".
91
+ */
92
+ function unresolvedCitation(cwd, citation) {
93
+ const rel = typeof citation?.path === "string" ? citation.path : "";
94
+ if (!rel) return "names no artifact path";
95
+ let text;
96
+ try {
97
+ text = readFileSync(resolve(cwd, rel), "utf8");
98
+ } catch (e) {
99
+ if (e.code === "ENOENT") return `cites ${rel}, which does not exist on disk`;
100
+ if (e.code === "EISDIR") return `cites ${rel}, which is a directory, not a T0 verdict file`;
101
+ // EACCES, ELOOP, EIO, EMFILE… — this machine failing to look, not evidence against the
102
+ // citation, so it is not refused on that ground.
103
+ return null;
104
+ }
105
+ const actual = createHash("sha256").update(text).digest("hex");
106
+ const claimed = typeof citation.sha256 === "string" ? citation.sha256.toLowerCase() : "";
107
+ if (actual !== claimed) return `cites ${rel} with sha256 ${citation.sha256}, but the file on disk hashes to ${actual}`;
108
+ let body;
109
+ try { body = JSON.parse(text); }
110
+ catch { return `cites ${rel}, whose bytes match the hash but do not read as a T0 verdict`; }
111
+ if (body?.overall !== "green") return `cites ${rel}, whose own verdict is "${body?.overall ?? "unknown"}", not green`;
112
+ return null;
113
+ }
114
+
62
115
  /**
63
116
  * Why a verdict cannot stand as its round's judgement on T0 grounds, or null when it can.
64
117
  *
@@ -68,21 +121,40 @@ export function isScoped(cwd, slug) {
68
121
  * and branched on like any other. It is checked here so the round loop, the resume derivation, the
69
122
  * hill and ingest all refuse the same verdict for the same reason.
70
123
  *
71
- * PRESENCE, NOT HASHES. The evaluator re-hashes what it cites; a slip transcribing a digest is not
72
- * evidence the verdict is wrong, and refusing a round over one would cost a whole re-evaluation.
124
+ * RESOLVES, NOT JUST PRESENT. A citation naming a path and a hash used to be taken on faith: any
125
+ * non-empty `t0_citations[]` passed, whatever it pointed at. Measured live: a PASS citing a path
126
+ * that does not exist on disk, and a PASS citing a real artifact whose own verdict was red, both
127
+ * ingested clean — presence stood in for a re-hash the schema had already promised. Each citation is
128
+ * now re-checked through {@link unresolvedCitation}.
129
+ *
130
+ * WHAT THIS DOES NOT CHECK: whether a citation is drawn from the order's own `payload.t0_artifacts`
131
+ * list. That would need the compiled order, which this function's callers do not equally have —
132
+ * `reduce ingest` holds it, but `probe eval` (and the round-loop/resume/hill readers behind it)
133
+ * knows only (cwd, slug, round), and a scope's T0 attempt can legitimately go green again LATER than
134
+ * whatever list was frozen at compile time (`greenVerdict` already treats "newest green" as
135
+ * authoritative for exactly this reason — see kernel/probe/t0.mjs). Enforcing membership only where
136
+ * the order happens to be on hand would let one channel refuse a citation the other accepts, for
137
+ * evidence that may simply be fresher than the order — worse than leaving it unenforced.
73
138
  *
74
139
  * @param {string} cwd - Project root.
75
140
  * @param {string} slug - Feature slug.
76
141
  * @param {object} verdict - The WorkResult's `verdict` block.
77
- * @returns {(string|null)} The problem, phrased for an operator; null for a cited verdict, an
78
- * unscoped spec, or a block with no PASS/FAIL in it (there is no judgement to invalidate).
142
+ * @returns {(string|null)} The problem, phrased for an operator; null for a verdict whose every
143
+ * citation resolves, an unscoped spec, or a block with no PASS/FAIL in it (there is no judgement
144
+ * to invalidate).
79
145
  */
80
146
  export function citationProblem(cwd, slug, verdict) {
81
147
  if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
82
- if (Array.isArray(verdict.t0_citations) && verdict.t0_citations.length) return null;
83
148
  if (!isScoped(cwd, slug)) return null;
84
- return `the ${verdict.overall} verdict cites no T0 artifact, and a verdict on a scoped spec must ` +
85
- "cite the T0 verdict it re-hashed (the order lists them under payload.t0_artifacts)";
149
+ if (!Array.isArray(verdict.t0_citations) || !verdict.t0_citations.length) {
150
+ return `the ${verdict.overall} verdict cites no T0 artifact, and a verdict on a scoped spec must ` +
151
+ "cite the T0 verdict it re-hashed (the order lists them under payload.t0_artifacts)";
152
+ }
153
+ for (const citation of verdict.t0_citations) {
154
+ const reason = unresolvedCitation(cwd, citation);
155
+ if (reason) return `the ${verdict.overall} verdict ${reason} — a T0 citation is re-hashed from disk, never taken on the handed word`;
156
+ }
157
+ return null;
86
158
  }
87
159
 
88
160
  /**
@@ -6,24 +6,158 @@
6
6
  import { readFileSync, writeFileSync, existsSync, readdirSync, mkdirSync } from "node:fs";
7
7
  import { resolve, join } from "node:path";
8
8
  import { runArgs } from "../lib/argv.mjs";
9
- import { scopesDir, hillDir, verdictsDir, resultsDir, discoveryLedger } from "../lib/paths.mjs";
9
+ import { scopesDir, hillDir, verdictsDir, resultsDir, discoveryLedger, localRoot } from "../lib/paths.mjs";
10
10
  import { readAllContracts, SCOPE_CONTRACT } from "../lib/contract.mjs";
11
11
  import { evalVerdict } from "../probe/eval.mjs";
12
12
  import { redBuildRounds } from "../verify/build.mjs";
13
13
 
14
+ // ---------------------------------------------------------------------------------------------
15
+ // THE DISCOVERY LEDGER, AND WHY ITS ABSENCE IS NOT A ZERO.
16
+ //
17
+ // The ledger arm is the only thing that promotes a scope from UPHILL_UNKNOWN to UPHILL_SOLVED —
18
+ // "the open questions are closed". It used to be read as `scopeUnknowns[id] || 0`, which made
19
+ // "nobody has filed a ledger yet" and "every unknown is closed" the same value, and that value
20
+ // selects the FLATTERING phase. An absent value and a real one must not share a signature; when
21
+ // they do, the run reports progress it has no evidence for. So the count is `null` — not `0` —
22
+ // whenever the ledger could not actually be read and understood, and `null` promotes nothing.
23
+ //
24
+ // THE HEADING IS THE ATTRIBUTION KEY, and it has to match what the ingest step really writes:
25
+ // `## Discovered — <order_id> (<date>)`, where `order_id` is `<slug>/<suffix>`. Two ways that
26
+ // parse used to fail silently, both ending in the same false zero:
27
+ // * It required a COLON between the slug and the suffix. The order-envelope schema pins
28
+ // `order_id` to a slug, a SLASH, and a suffix drawn from `[a-z0-9.-]` — a colon cannot appear
29
+ // in a schema-valid order id at all, so the scope was never captured from any real ledger and
30
+ // every scope on every project read zero unknowns regardless of what was open.
31
+ // * It matched the em dash only, so a heading typed with a plain hyphen contributed nothing.
32
+ // Both are fixed by matching the dash as a class and splitting the order id on its "/", and by
33
+ // resolving the scope against the CONTRACTS ON DISK rather than guessing at the suffix's shape: a
34
+ // build suffix is `<scope>-r<N>-a<M>` and a scoped operation's is `<operation>-<scope>[-r<N>]`, so
35
+ // a pattern that tried to carve the scope out positionally captured the round suffix along with it.
36
+ // ---------------------------------------------------------------------------------------------
37
+
38
+ /** A ledger block heading, naming the order it files: `## Discovered — <slug>/<suffix> (<date>)`. */
39
+ const LEDGER_HEADING = /^##\s+Discovered\s+[—–-]\s+(\S+)/;
40
+
41
+ /**
42
+ * Index the scopes by the filename-safe form `harness compile` puts in an order suffix, so a
43
+ * heading is matched against ids that actually exist rather than parsed speculatively.
44
+ *
45
+ * @param {Array<object>} scopes - Parsed scope contracts.
46
+ * @returns {Map<string,string>} Compile's suffix form → the contract's own `scope_id`.
47
+ */
48
+ function scopeSuffixIndex(scopes) {
49
+ const ix = new Map();
50
+ for (const s of scopes) {
51
+ const id = String(s?.scope_id || "");
52
+ if (!id) continue;
53
+ // Mirrors `harness compile`'s own normalisation of a scope id into an order suffix.
54
+ const key = id.toLowerCase().replace(/[^a-z0-9.-]/g, "-").replace(/^[^a-z0-9]+/, "");
55
+ if (key) ix.set(key, id);
56
+ }
57
+ return ix;
58
+ }
59
+
60
+ /**
61
+ * The scope a ledger heading's order belongs to, or `null` when it belongs to none.
62
+ *
63
+ * Operation-level dispatches (orient, analyze, wire, evaluate, hunt, hammer) carry no scope in
64
+ * their order id by construction, so "no scope" is a real and common answer here, not a parse
65
+ * failure — their rows are feature-wide and are deliberately credited to nobody.
66
+ *
67
+ * @param {string} orderId - The order id as the heading names it (`<slug>/<suffix>`).
68
+ * @param {Map<string,string>} ix - The index from `scopeSuffixIndex`.
69
+ * @returns {string|null} The owning contract's `scope_id`, or null.
70
+ */
71
+ function scopeOfHeading(orderId, ix) {
72
+ const slash = orderId.indexOf("/");
73
+ const suffix = slash === -1 ? orderId : orderId.slice(slash + 1);
74
+ const core = suffix.replace(/-r\d+(?:-a\d+)?$/, "");
75
+ // Longest match wins, so a project holding both `pages` and `pages-admin` attributes each
76
+ // heading to the scope it actually names rather than to whichever was indexed first.
77
+ let best = null;
78
+ let bestLen = -1;
79
+ for (const [key, id] of ix) {
80
+ if ((core === key || core.endsWith(`-${key}`)) && key.length > bestLen) { best = id; bestLen = key.length; }
81
+ }
82
+ return best;
83
+ }
84
+
85
+ /**
86
+ * Open unknowns (`~` rows) per scope, read from the discovery ledger.
87
+ *
88
+ * @param {string} ledgerPath - Path to the run's discovery ledger.
89
+ * @param {Array<object>} scopes - Parsed scope contracts, used to attribute each block.
90
+ * @returns {Object<string,number>|null} Counts per `scope_id` — a scope with a block and no open
91
+ * row is a genuine `0`. `null` means the ledger was NOT read: it is absent, unreadable, or
92
+ * nothing in it could be attributed to a scope. `null` is never treated as zero, because "we
93
+ * understood none of this file" is not evidence that nothing is open.
94
+ */
95
+ function ledgerUnknowns(ledgerPath, scopes) {
96
+ if (!existsSync(ledgerPath)) return null;
97
+ let text;
98
+ try { text = readFileSync(ledgerPath, "utf8"); } catch { return null; }
99
+
100
+ const ix = scopeSuffixIndex(scopes);
101
+ const counts = {};
102
+ let headings = 0;
103
+ let attributed = 0;
104
+ let current = null;
105
+
106
+ for (const line of text.split("\n")) {
107
+ if (line.startsWith("## ")) {
108
+ // Reset on EVERY heading, matched or not. Carrying the previous block's scope across an
109
+ // unrecognised heading would credit one scope's open rows to another.
110
+ current = null;
111
+ const m = line.match(LEDGER_HEADING);
112
+ if (!m) continue;
113
+ headings++;
114
+ const id = scopeOfHeading(m[1], ix);
115
+ if (id) { attributed++; current = id; counts[id] = counts[id] || 0; }
116
+ continue;
117
+ }
118
+ if (current && line.startsWith("~ ")) counts[current]++;
119
+ }
120
+
121
+ // A ledger whose blocks we could not attribute to a single scope tells us nothing per scope.
122
+ // Reporting that as all-zero is the same false signature the colon-vs-slash bug produced.
123
+ if (headings === 0 || attributed === 0) return null;
124
+ return counts;
125
+ }
126
+
127
+ /**
128
+ * The phase currently recorded in a committed hill shard.
129
+ *
130
+ * @param {string} hDir - The committed hill directory.
131
+ * @param {string} id - Scope id.
132
+ * @returns {string|null} The recorded phase, or null when no shard exists for that scope.
133
+ */
134
+ function committedPhase(hDir, id) {
135
+ const p = join(hDir, `${id}.yml`);
136
+ if (!existsSync(p)) return null;
137
+ try { return readFileSync(p, "utf8").match(/^phase:\s*(\S+)/m)?.[1] ?? null; } catch { return null; }
138
+ }
139
+
14
140
  /**
15
141
  * Derive and write the hill phase for all scopes mechanically based on T0, T1, and ledger facts.
16
142
  *
17
143
  * The derived phase follows these progression rules (facts move dots, not authors):
18
- * - UPHILL_UNKNOWN: open unknowns > 0 in the ledger for this scope
19
- * - UPHILL_SOLVED: unknowns 0, no T0-green yet
144
+ * - UPHILL_UNKNOWN: open unknowns > 0 in the ledger for this scope — and the floor the scope sits
145
+ * at whenever the ledger has not answered at all, which is where every run legitimately begins
146
+ * - UPHILL_SOLVED: the ledger was read and reports zero open unknowns, no T0-green yet
20
147
  * - DOWNHILL_EXECUTION: ≥1 T0-green in a round whose build gate is not red; T1/seesaw pending
21
148
  * - FINISHED: T1 PASS ∧ seesaw green
22
149
  *
23
150
  * @param {string} cwd - The project root directory.
24
151
  * @param {string} slug - The feature slug being built.
25
- * @returns {Array<{scope_id: string, phase: string, changed: boolean}>} A report of all scopes processed, their derived phase, and whether the hill shard on disk was modified.
26
- * Side effects: writes to `shapeup/<slug>/hill/<scope-id>.yml` for each scope.
152
+ * @returns {Array<{scope_id: string, phase: (string|null), changed: boolean, derived: boolean,
153
+ * unknowns: (number|null), reason?: string}>} A report of all scopes processed. `derived` says
154
+ * whether the phase was computed from run evidence at all: when it is false the phase is
155
+ * whatever the committed shard already records (or null when there is none), nothing was
156
+ * written, and `reason` names why. `unknowns` is the ledger count behind the phase, or null
157
+ * when the ledger could not be read — the two are deliberately distinguishable in the output as
158
+ * well as in the derivation.
159
+ * Side effects: writes to `shapeup/<slug>/hill/<scope-id>.yml` for each scope, EXCEPT on the
160
+ * refusal path below, which writes nothing at all.
27
161
  */
28
162
  export function deriveHill(cwd, slug) {
29
163
  const scopes = readAllContracts(scopesDir(cwd, slug), SCOPE_CONTRACT).map((x) => x.contract);
@@ -31,6 +165,44 @@ export function deriveHill(cwd, slug) {
31
165
  const ledgerPath = discoveryLedger(cwd, slug);
32
166
  const hDir = hillDir(cwd, slug);
33
167
 
168
+ // -------------------------------------------------------------------------------------------
169
+ // A DERIVATION THAT CANNOT SEE THE RUN TRACE MUST NOT WRITE THE DELIVERABLE.
170
+ //
171
+ // This function reads one tier and writes the other. Every input below — T0 verdicts, the EVAL
172
+ // results, the round build gates, the discovery ledger — lives in the gitignored run trace,
173
+ // while the shards it writes are committed and outlive it: after a ship the run trace is cleaned
174
+ // up and the shards are the ONLY surviving record of where the work got to. So when the run
175
+ // trace for this slug is not on disk, every input is provably absent, and anything derived from
176
+ // that is derived from nothing.
177
+ //
178
+ // This is not a hypothetical. Pulling a branch mid-run is a supported state: the puller gets the
179
+ // committed spec, scopes and shards and no run trace of their own. Their first launch used to
180
+ // re-derive every scope from the empty set and overwrite the shards with the result — losing
181
+ // committed history rather than misreporting it.
182
+ //
183
+ // WHY REFUSE RATHER THAN WRITE AN "UNKNOWN" PHASE. Writing anything here destroys the record
184
+ // just as thoroughly; a shard that says "I could not look" has still replaced the one that said
185
+ // FINISHED, and the phase enum is a committed data format that a reader parses as current
186
+ // status. Not writing already expresses "no opinion" exactly, and it needs no new enum value.
187
+ //
188
+ // WHY THE CONDITION IS THE TIER'S EXISTENCE AND NOT "the phase would go down". Moving a dot
189
+ // backwards is correct and must stay possible — this is a pure function of the artifacts present
190
+ // and reports what they currently support, in both directions. A guard phrased as "never lower a
191
+ // phase" would quietly turn a derived value into a high-water mark, which is a different defect
192
+ // wearing this one's clothes. The condition is the narrowest one that is positively provable:
193
+ // the tier holding every input is not there.
194
+ // -------------------------------------------------------------------------------------------
195
+ if (!existsSync(localRoot(cwd, slug))) {
196
+ return scopes.map((s) => ({
197
+ scope_id: s.scope_id,
198
+ phase: committedPhase(hDir, s.scope_id),
199
+ changed: false,
200
+ derived: false,
201
+ unknowns: null,
202
+ reason: "local-run-trace-absent",
203
+ }));
204
+ }
205
+
34
206
  if (!existsSync(hDir)) mkdirSync(hDir, { recursive: true });
35
207
 
36
208
  // 1. Check if T1 Evaluation passed — the LATEST evaluate round's verdict, read the same way
@@ -95,28 +267,20 @@ export function deriveHill(cwd, slug) {
95
267
  }
96
268
  }
97
269
 
98
- // 3. Ledger unknowns per scope
99
- const scopeUnknowns = {};
100
- if (existsSync(ledgerPath)) {
101
- const lines = readFileSync(ledgerPath, "utf8").split("\n");
102
- let currentScope = null;
103
- for (const line of lines) {
104
- const m = line.match(/^## Discovered — .*?:([\w.-]+)-a\d+/);
105
- if (m) {
106
- currentScope = m[1];
107
- }
108
- if (currentScope && line.startsWith("~ ")) {
109
- scopeUnknowns[currentScope] = (scopeUnknowns[currentScope] || 0) + 1;
110
- }
111
- }
112
- }
113
-
270
+ // 3. Ledger unknowns per scope — `null` for every scope when the ledger itself was not readable
271
+ // or nothing in it named a scope (see `ledgerUnknowns`). Only a real count can promote.
272
+ const scopeUnknowns = ledgerUnknowns(ledgerPath, scopes);
273
+
114
274
  const report = [];
115
275
  for (const s of scopes) {
116
276
  const id = s.scope_id;
117
277
  const t0 = t0Facts[id] || { hasGreen: false, seesawGreen: false };
118
- const unknowns = scopeUnknowns[id] || 0;
119
-
278
+ // `null` = the ledger did not answer; a number = it did. `|| 0` collapsed the two.
279
+ const unknowns = scopeUnknowns === null ? null : (scopeUnknowns[id] || 0);
280
+
281
+ // UPHILL_UNKNOWN is the floor, and the honest answer whenever nothing has promoted a scope off
282
+ // it — including before Orient has filed anything, which is where every run legitimately
283
+ // starts. Only an ANSWERED count of zero promotes to UPHILL_SOLVED; `null` never does.
120
284
  let phase = "UPHILL_UNKNOWN";
121
285
  if (t1Pass && t0.hasGreen && t0.seesawGreen) {
122
286
  phase = "FINISHED";
@@ -125,7 +289,7 @@ export function deriveHill(cwd, slug) {
125
289
  } else if (unknowns === 0) {
126
290
  phase = "UPHILL_SOLVED";
127
291
  }
128
-
292
+
129
293
  const yaml = `scope_id: ${id}\nphase: ${phase}\n`;
130
294
  const out = join(hDir, `${id}.yml`);
131
295
  let changed = false;
@@ -133,7 +297,7 @@ export function deriveHill(cwd, slug) {
133
297
  writeFileSync(out, yaml);
134
298
  changed = true;
135
299
  }
136
- report.push({ scope_id: id, phase, changed });
300
+ report.push({ scope_id: id, phase, changed, derived: true, unknowns });
137
301
  }
138
302
  return report;
139
303
  }
@@ -156,5 +320,16 @@ export async function cli(rawArgv) {
156
320
  const args = runArgs(ARGV_SPEC, rawArgv);
157
321
  const cwd = resolve(args.cwd || process.cwd());
158
322
  const report = deriveHill(cwd, args.slug);
323
+ // A refusal that is visible only as a missing write reads exactly like a derivation that
324
+ // happened to agree with what was already on disk, so say it out loud. It is not an error —
325
+ // the caller runs this advisorily several times a run, and declining to derive from nothing is
326
+ // the correct outcome, not a failure — so the exit code stays 0 and the warning goes to stderr.
327
+ const underived = report.filter((r) => r.derived === false);
328
+ if (underived.length) {
329
+ console.error(
330
+ `hill: derived nothing for ${underived.length} scope(s) (${underived[0].reason}) — ` +
331
+ `the run trace this phase is derived from is not on disk, so the committed shards were left as they are.`,
332
+ );
333
+ }
159
334
  console.log(JSON.stringify(report, null, 2));
160
335
  }
@@ -32,7 +32,7 @@ import { runArgs } from "../lib/argv.mjs";
32
32
  import {
33
33
  report as reportPath, tasksDir, verdictsDir, trials, evaluationDir, qaDir,
34
34
  roundLedger, discoveryLedger, receipt as receiptPath, harnessRun, relShared,
35
- activeOrder,
35
+ activeOrder, runArgsPath, readReceipt, runIdFromReceipt,
36
36
  } from "../lib/paths.mjs";
37
37
  import { readTrials } from "../verify/t0.mjs";
38
38
  import { ratchetReport } from "../probe/stats.mjs";
@@ -398,9 +398,51 @@ export const ARGV_SPEC = {
398
398
  * @returns {(Promise<void>|void)} Settles when the subcommand has written its output; most paths
399
399
  * call `process.exit()` with the subcommand's documented code rather than returning.
400
400
  */
401
+ /**
402
+ * Did THIS run skip evaluation? Read from the run's own recorded arguments, and believed only when
403
+ * the record can be shown to belong to this run.
404
+ *
405
+ * WHY THE KERNEL ASKS AT ALL. A run launched with `--no-eval` verified nothing, and the protocol
406
+ * has always said such a run ships as `not-evaluated`, "recorded plainly — never silently
407
+ * upgraded". That promise lived entirely inside the orchestrator's control flow, where one
408
+ * assignment downstream of the branch restores the old behaviour with every check still green —
409
+ * measured, not supposed. The gate block tells the human the truth either way; the committed report
410
+ * a teammate inherits on `git pull` is the artifact that was lying, so the refusal belongs at the
411
+ * writer of that artifact, where a future orchestrator edit cannot reach it.
412
+ *
413
+ * POSITIVELY PROVEN OR NOT AT ALL. The record is written at launch, later than the receipt that
414
+ * mints the run key, so a freshly opened run can still find the PREVIOUS run's record on disk. A
415
+ * stale flag would refuse a ship that verified everything — a worse failure than the one this
416
+ * closes. The flag therefore counts only when the record names the same run the receipt does;
417
+ * a missing, unreadable or differently-keyed record proves nothing and permits.
418
+ *
419
+ * @param {string} cwd - Project root.
420
+ * @param {string} slug - Feature slug.
421
+ * @returns {boolean} True only when this run's own record says evaluation was skipped.
422
+ */
423
+ function evalWasSkipped(cwd, slug) {
424
+ try {
425
+ const record = JSON.parse(readFileSync(runArgsPath(cwd, slug), "utf8"));
426
+ if (record?.noEval !== true) return false;
427
+ const mine = runIdFromReceipt(readReceipt(receiptPath(cwd, slug)));
428
+ return Boolean(mine) && record.runId === mine;
429
+ } catch { return false; }
430
+ }
431
+
432
+ /** The verdict values that assert the feature was graded and passed. */
433
+ const PASSING = new Set(["PASS", "pass"]);
434
+
401
435
  export async function cli(rawArgv) {
402
436
  const args = runArgs(ARGV_SPEC, rawArgv);
403
437
  const cwd = args.cwd || process.cwd();
438
+ if (PASSING.has(String(args.verdict ?? "")) && evalWasSkipped(cwd, args.slug)) {
439
+ console.error(
440
+ "✋ reduce ship: this run was launched with --no-eval, so nothing graded it — refusing to freeze a report " +
441
+ `that says ${args.verdict}. Ship it as --verdict not-evaluated, which the report, the ledger and the ` +
442
+ "sign-off block all carry.",
443
+ );
444
+ process.exit(3);
445
+ }
404
446
  const { markdown, path } = generate({ cwd, slug: args.slug, verdict: args.verdict, qa: args.qa });
405
447
  if (args.stdout) {
406
448
  process.stdout.write(markdown);
@@ -2846,9 +2846,10 @@
2846
2846
  "type": "string",
2847
2847
  "enum": [
2848
2848
  "pass",
2849
- "fail"
2849
+ "fail",
2850
+ "not-evaluated"
2850
2851
  ],
2851
- "description": "shipped | ok: the round or run's EVAL verdict, lowercased from spec-evaluator's PASS|FAIL."
2852
+ "description": "shipped | ok: the round or run's EVAL verdict, lowercased from spec-evaluator's PASS|FAIL — or `not-evaluated` when a `shipped` return closed a `--no-eval` run, which reaches SHIP without ever dispatching the judge. Recorded plainly, never upgraded to `pass` (references/protocol.md)."
2852
2853
  },
2853
2854
  "rounds_used": {
2854
2855
  "type": "integer"
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.7.1",
3
+ "version": "3.7.2",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -48,12 +48,16 @@ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" reduce graph --slug <slug>
48
48
  ```
49
49
 
50
50
  **For a committed-only slug (`hasCommitted && !hasLocal`), NEVER call `reduce hill`.**
51
- `deriveHill()` folds whatever T0 verdicts currently exist on disk; a committed-only pitch has
52
- none (the local trace was cleaned up after shipping), so re-running it would silently regress a
53
- true historical `FINISHED` down to a fabricated `UPHILL_SOLVED` — reading absence of evidence as
54
- evidence of absence. Read `shapeup/<slug>/hill/*.yml` for those slugs exactly as committed, and
55
- render them with the archived state the template already implements (see below). Same reasoning
56
- applies to `reduce graph` — there is no local trace to append from.
51
+ `deriveHill()` folds whatever T0 verdicts, evaluation results and ledger rows exist in the LOCAL
52
+ run trace; a committed-only pitch has none, because that trace was cleaned up after shipping.
53
+
54
+ The runtime no longer takes that silence for an answer: a derivation whose run trace is absent
55
+ declines to write anything and reports every scope as underived (`derived: false`), so the
56
+ committed shards survive a call made in error. Treat that as a backstop, not a licence — an
57
+ underived report is not a refresh, and rendering it as one would show a pitch's true `FINISHED`
58
+ as freshly confirmed when nothing confirmed it. Read `shapeup/<slug>/hill/*.yml` for those slugs
59
+ exactly as committed, and render them with the archived state the template already implements
60
+ (see below). Same reasoning applies to `reduce graph` — there is no local trace to append from.
57
61
 
58
62
  ## Reading the hill shards
59
63
 
@@ -389,7 +389,7 @@ feature) and the failing step is compiled into round r+1's orders as `payload.bu
389
389
  means L0 pinned no run command and the profile names no probe — the round proceeds over an
390
390
  unproven build, and the block says so.
391
391
 
392
- Under `--interactive` / `--auto`, the hook warns if the board is not truly green (advisory) and requires explicit PO approval to proceed. Under `--unattended`, it automatically aborts on a red board or proceeds on a green one.
392
+ Under `--interactive` / `--auto`, the gate block carries the board facts — `green_scopes`, `hammer_proposals` — and requires explicit PO approval to proceed. Under `--unattended` the answer set resolves it. **There is no hook behind this gate**: the `PreToolUse` warning it used to carry was retired into the block in v2.0, so what an unfinished board costs you here is a human reading the numbers, not a machine refusing the call.
393
393
 
394
394
  ---
395
395
 
@@ -1596,7 +1596,11 @@ while (verdict !== "pass" && round <= maxRounds) {
1596
1596
  `feature; this one does not build or launch. Round ${round + 1} fixes the gate's failing step.`);
1597
1597
  } else if (args.noEval) {
1598
1598
  log("EVAL — skipped (--no-eval)");
1599
- verdict = "pass";
1599
+ // references/protocol.md's own words for this, twice: a run --no-eval ships records
1600
+ // `not-evaluated` — "recorded plainly — never silently upgraded [to `pass`]". Nothing verified
1601
+ // the feature beyond task-executor's own per-AC evidence checks, and the report this run
1602
+ // freezes at GATE L4 has to say that as plainly as the gate block already does.
1603
+ verdict = "not-evaluated";
1600
1604
  } else {
1601
1605
  const e = await worker({
1602
1606
  skill: "spec-evaluator", operation: "evaluate", schema: EVAL, phase: "Eval", label: `eval:r${round}`,
@@ -1647,14 +1651,16 @@ while (verdict !== "pass" && round <= maxRounds) {
1647
1651
  const g3 = await crossGate("L3", "Eval", ["loop", "stop", "ask"], { round, verdict, build_gate: buildGate });
1648
1652
  if (g3.stop) return await withWarnings(g3.stop);
1649
1653
 
1650
- if (verdict === "pass") break; // → QA → GATE H → ship
1654
+ // "not-evaluated" (--no-eval) ships exactly like "pass" — protocol.md: the run "goes straight to
1655
+ // SHIP" once EVAL is skipped, never spends another round waiting on a verdict nobody is producing.
1656
+ if (verdict === "pass" || verdict === "not-evaluated") break; // → QA → GATE H → ship
1651
1657
  if (g3.decision === "stop" || round >= maxRounds) {
1652
1658
  return await withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
1653
1659
  }
1654
1660
  round += 1;
1655
1661
  }
1656
1662
 
1657
- if (verdict !== "pass") {
1663
+ if (verdict !== "pass" && verdict !== "not-evaluated") {
1658
1664
  return await withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
1659
1665
  }
1660
1666
 
@@ -1692,7 +1698,12 @@ if (h.verdict === "cannot-ship") {
1692
1698
  if (g.stop) return await withWarnings(g.stop);
1693
1699
  }
1694
1700
 
1695
- const ship = await cmd(`reduce ship --slug ${slug} --verdict PASS --qa ${qaRan ? "run" : "skipped"}`, "Ship", "ship-report");
1701
+ // The frozen report's own verdict line has to say what actually happened — a `--no-eval` run
1702
+ // verified nothing beyond task-executor's per-AC checks, and `reduce ship` already accepts
1703
+ // "not-evaluated" as a real verdict (its own usage string, and `generate()`'s no-artifact
1704
+ // default). Hardcoding PASS here is exactly the silent upgrade protocol.md's Rules forbid.
1705
+ const shipVerdict = verdict === "not-evaluated" ? "not-evaluated" : "PASS";
1706
+ const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${qaRan ? "run" : "skipped"}`, "Ship", "ship-report");
1696
1707
  await advisory(`report export --slug ${slug}`, "Ship", "export-run");
1697
1708
  // The run's own concurrency, printed once where the records are complete and before the next run
1698
1709
  // supersedes the trace. It is a projection over `receipts/dispatch.jsonl` and `legs.jsonl`, so it
@@ -1707,7 +1718,10 @@ const ALL_DIMS = ["spec-conformance", "tdd-surface", "integration", "completenes
1707
1718
 
1708
1719
  return await withWarnings({
1709
1720
  status: "shipped",
1710
- verdict: "pass",
1721
+ // The real verdict this run reached — "pass" or, over a --no-eval run, "not-evaluated". GATE L4's
1722
+ // own sign-off block reads this field verbatim (SKILL.md Step 4); hardcoding "pass" here told a
1723
+ // human answering that gate the run was graded when it never was.
1724
+ verdict,
1711
1725
  rounds_used: round,
1712
1726
  dims_not_evaluated: ALL_DIMS.filter((d) => !evalDims.includes(d)),
1713
1727
  qa_findings: qaFindings,