shapeup-sdlc 3.7.9 → 3.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +2 -2
- package/README.md +8 -9
- package/hooks/sandbox-guard.mjs +9 -0
- package/kernel/compile.mjs +67 -15
- package/kernel/lib/argv.mjs +4 -4
- package/kernel/lib/paths.mjs +0 -2
- package/kernel/probe/eval.mjs +19 -5
- package/kernel/reduce/hill.mjs +17 -39
- package/kernel/report/export.mjs +0 -3
- package/kernel/schemas/domain.schema.json +16 -93
- package/kernel/verify/env.mjs +25 -7
- package/kernel/verify/ratchet-tree.mjs +1 -1
- package/kernel/verify/spec.mjs +70 -0
- package/kernel/verify/t0.mjs +35 -78
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/assets/templates/synthesis.tmpl.md +5 -0
- package/skills/ba-pitch-analyzer/references/task-generation.md +1 -1
- package/skills/hill-chart/assets/dashboard.template.html +1 -1
- package/skills/scope-architect/SKILL.md +3 -3
- package/skills/tech-lead/references/gates.md +1 -1
- package/skills/tech-lead/references/protocol.md +6 -8
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.8.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -35,7 +35,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
35
35
|
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
|
-
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe
|
|
38
|
+
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
39
39
|
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
|
|
40
40
|
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
|
|
41
41
|
|
|
@@ -63,7 +63,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
|
|
|
63
63
|
- **QA is a level-up, not a gate** — `--no-qa` skips it; circuit breaker outranks the Hunter.
|
|
64
64
|
- **Role separation** — Evaluator grades, task-executor fixes, QA discovers.
|
|
65
65
|
- **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
|
|
66
|
-
- **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts
|
|
66
|
+
- **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts, never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
|
|
67
67
|
- **Envelope port (v1.0)** — every dispatch is WorkOrder in / WorkResult out; shared state has exactly one writer (the ingest step); malformed envelopes are hook-denied. Workers: stateless, craft-only, pipeline-blind.
|
|
68
68
|
|
|
69
69
|
## Setup & Execution
|
package/README.md
CHANGED
|
@@ -45,11 +45,11 @@ its fixtures and its DB probe and writes an artifact to disk — with each comma
|
|
|
45
45
|
its captured output, and whether it ran at all. The evaluator must cite that artifact, and the
|
|
46
46
|
hill phase is derived from artifacts rather than from a worker's own account of its progress.
|
|
47
47
|
Two limits, stated here because the point of this section is that a claim without a mechanism
|
|
48
|
-
behind it is the thing this harness exists to prevent:
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
48
|
+
behind it is the thing this harness exists to prevent: a citation is re-hashed from disk, checked
|
|
49
|
+
against the scope, round and run the artifact records, and refused unless it names a file this
|
|
50
|
+
run's own verifier wrote — what that does NOT prove is that a judge with a dishonest hand could
|
|
51
|
+
not have arranged the artifact first, which is why the judge's own substrate freezes the verdicts
|
|
52
|
+
directory. A T0 artifact is
|
|
53
53
|
also evidence about the machine that produced it, and now says so: each verdict carries where it
|
|
54
54
|
ran — the absolute path, the git tree, the resolved toolchain, lockfile digests, declared cache
|
|
55
55
|
directories and a digest over an allowlist of environment values — so a disagreeing re-run can be
|
|
@@ -132,8 +132,7 @@ rest of this README after this table and nothing will be a surprise.
|
|
|
132
132
|
|---|---|
|
|
133
133
|
| **board** | The round's task list. "Green" means every task is done. GATE L2's hook reads this before an evaluation and warns if it is not green. |
|
|
134
134
|
| **round** | One build → evaluate cycle. A FAIL verdict starts round *r+1*. |
|
|
135
|
-
| **T0** | The smoke test a scope must pass before it counts as built: its fixtures
|
|
136
|
-
| **seesaw** | The regression arm of T0, meant to re-run *other* scopes' fixtures so a regression is never mistaken for progress. Declared, not yet wired: no run writes its registry, so no run has executed it. |
|
|
135
|
+
| **T0** | The smoke test a scope must pass before it counts as built: its fixtures and a DB probe. Writes an artifact to disk that the evaluator must cite. |
|
|
137
136
|
| **substrate** | The exact list of files one dispatch is allowed to write, stamped into its work order. A hook blocks anything outside it — and anything the order marks frozen. |
|
|
138
137
|
| **scope contract** | The file defining one vertical slice: its substrate, its fixtures, its affordances. |
|
|
139
138
|
| **affordance** | The thing a user can actually click, type or call. UI is graded on affordances, not on looks. |
|
|
@@ -309,8 +308,8 @@ These hold across the harness and are the reason it stays predictable:
|
|
|
309
308
|
of blocking the round. An opt-in third breaker bounds the **wall clock**, because the other two
|
|
310
309
|
count events and neither can notice a single round running for half an hour — tripping it routes
|
|
311
310
|
to GATE H, so a run out of time ships what is green instead of being killed and shipping nothing.
|
|
312
|
-
- **Hill phase is mechanical, never self-reported** — derived only from T0/T1 facts
|
|
313
|
-
|
|
311
|
+
- **Hill phase is mechanical, never self-reported** — derived only from T0/T1 facts, closing the
|
|
312
|
+
self-reported-confidence risk. A scope with no
|
|
314
313
|
discovery ledger derives no phase rather than a solved one, so absence no longer reads as
|
|
315
314
|
progress on that arm.
|
|
316
315
|
- **One writer per shared file** — every board/ledger/verdict write goes through
|
package/hooks/sandbox-guard.mjs
CHANGED
|
@@ -329,6 +329,11 @@ async function main() {
|
|
|
329
329
|
allowed: [...(o.substrate.allowed || []), ...(o.substrate.shared || [])],
|
|
330
330
|
appendOnly: o.substrate.append_only || [],
|
|
331
331
|
frozen: o.substrate.frozen || [],
|
|
332
|
+
// The one exception to "frozen outranks everything", and it is the compiler's to grant, never
|
|
333
|
+
// the worker's to request: the paths THIS order authors, named from its own identity. A build
|
|
334
|
+
// leg's substrate freezes the whole run trace, so without this its own WorkResult — its
|
|
335
|
+
// documented last step — would be denied along with every channel it must not touch.
|
|
336
|
+
own: o.substrate.own || [],
|
|
332
337
|
})).filter((c) => c.allowed.length || c.appendOnly.length || c.frozen.length);
|
|
333
338
|
|
|
334
339
|
if (contracts.length === 0) defer("no live order declares write/append/frozen boundaries", "no-whitelist");
|
|
@@ -363,6 +368,10 @@ async function main() {
|
|
|
363
368
|
// a planner is graded against — so the compiler emitted a declaration with no enforcer, which is
|
|
364
369
|
// the exact state this hook exists to end. A path a live contract freezes is a violation
|
|
365
370
|
// wherever it lives.
|
|
371
|
+
// `own` first, and only against the contract that declared it: another live order's exception
|
|
372
|
+
// never licenses this write. A path no contract claims as its own falls through to the freeze.
|
|
373
|
+
if (contracts.some((c) => matchesAny(rel, c.own, fold))) continue;
|
|
374
|
+
|
|
366
375
|
const freezer = contracts.find((c) => matchesAny(rel, c.frozen, fold));
|
|
367
376
|
if (freezer) {
|
|
368
377
|
violations.push(rel);
|
package/kernel/compile.mjs
CHANGED
|
@@ -211,13 +211,37 @@ export const OP_OWNER = {
|
|
|
211
211
|
* operation, so mode/flag differences are enforced by the sandbox hook reading the order's substrate, not trusted to prose.
|
|
212
212
|
* @param {string} operation - The order's operation (execute|fix|spike|analyze|reconcile|
|
|
213
213
|
* retrofit-surface|coverage|map-scopes|wire|evaluate|orient|hunt|translate|hammer|coach|scan|research).
|
|
214
|
-
* @param {{slug?:string, specDir?:string, scope?:object}} [ctx] - slug (names
|
|
215
|
-
* specDir (overrides the default spec path), scope (contract supplying
|
|
216
|
-
*
|
|
217
|
-
*
|
|
218
|
-
*
|
|
214
|
+
* @param {{slug?:string, specDir?:string, scope?:object, ownStem?:string}} [ctx] - slug (names
|
|
215
|
+
* LOCAL/SHARED roots), specDir (overrides the default spec path), scope (contract supplying
|
|
216
|
+
* allowed/shared substrates), ownStem (this order's own file stem, so a build leg may write its
|
|
217
|
+
* own WorkResult and no one else's).
|
|
218
|
+
* @returns {{allowed:string[], shared?:string[], frozen?:string[], append_only?:string[], own?:string[]}}
|
|
219
|
+
* The substrate contract: globs the worker may write (`allowed`), shared-write globs, read-only
|
|
220
|
+
* `frozen` globs, `append_only` globs, and `own` — the paths this order may write DESPITE a
|
|
221
|
+
* broader freeze, derived by the compiler from the order's own identity and never requested.
|
|
222
|
+
* An unknown operation returns a LOCAL-only default.
|
|
219
223
|
*/
|
|
220
|
-
export function substrateFor(operation,
|
|
224
|
+
export function substrateFor(operation, ctx = {}) {
|
|
225
|
+
// EVERY ORDER NAMES THE RESULT IT ANSWERS WITH. A worker used to infer that path from its order's
|
|
226
|
+
// own filename — the one thing in the envelope that was convention rather than contract — so the
|
|
227
|
+
// one file every dispatch must write was the one the order did not mention. It is `own` for every
|
|
228
|
+
// operation now: the compiler derives it from the order's identity, which is also what keeps a leg
|
|
229
|
+
// from writing somebody else's.
|
|
230
|
+
const base = substrateTemplate(operation, ctx);
|
|
231
|
+
if (!ctx.ownStem) return base;
|
|
232
|
+
const ownResult = `${globLocal(ctx.slug)}/results/${ctx.ownStem}.json`;
|
|
233
|
+
return { ...base, own: [...new Set([...(base.own || []), ownResult])] };
|
|
234
|
+
}
|
|
235
|
+
|
|
236
|
+
/**
|
|
237
|
+
* The per-operation template {@link substrateFor} builds on.
|
|
238
|
+
*
|
|
239
|
+
* @param {string} operation - The order's operation.
|
|
240
|
+
* @param {object} [ctx] - As {@link substrateFor}: slug, specDir, scope, ownStem.
|
|
241
|
+
* @returns {{allowed:string[], shared?:string[], frozen?:string[], append_only?:string[], own?:string[]}}
|
|
242
|
+
* The operation's write contract before the result path every order names is merged in.
|
|
243
|
+
*/
|
|
244
|
+
function substrateTemplate(operation, { slug, specDir, scope, ownStem = null } = {}) {
|
|
221
245
|
const local = globLocal(slug);
|
|
222
246
|
const spec = specDir || globShared(slug, "spec");
|
|
223
247
|
const scopesDir = globShared(slug, "scopes");
|
|
@@ -266,15 +290,30 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
|
|
|
266
290
|
const FROZEN_ATTESTATION = [`${local}/receipts/**`, `${local}/legs.jsonl`, `${local}/t0/verdicts/**`, `${local}/tasks/_index.md`];
|
|
267
291
|
switch (operation) {
|
|
268
292
|
case "execute": case "fix": case "spike":
|
|
269
|
-
//
|
|
270
|
-
//
|
|
271
|
-
//
|
|
272
|
-
//
|
|
273
|
-
//
|
|
293
|
+
// THE RUN TRACE IS THE KERNEL'S, EXCEPT WHAT THIS LEG AUTHORS. The freeze used to be a list of
|
|
294
|
+
// channels a defect had named — the staged pitch, the receipts, the leg ledger, the T0
|
|
295
|
+
// verdicts, the board index — and each review found more of the same class: the trial ledger,
|
|
296
|
+
// the gate ledger, the round build gates, the graph, the run args, the run ledger itself. A
|
|
297
|
+
// list that grows one defect at a time is not a boundary. So the boundary is inverted here:
|
|
298
|
+
// everything under the run trace is frozen for a build leg, and `own` carries the short,
|
|
299
|
+
// derived list of what such a leg actually writes. Everything else under there is written by
|
|
300
|
+
// the kernel in its own process or by the hook layer, and neither goes through this guard —
|
|
301
|
+
// so freezing it costs no legitimate write. `${local}/**` subsumes FROZEN_INTAKE and
|
|
302
|
+
// FROZEN_ATTESTATION for this operation, and a check asserts that rather than restating them.
|
|
274
303
|
return {
|
|
275
304
|
allowed: [...(scope?.allowed_file_substrate || []), `${local}/spikes/**`],
|
|
276
305
|
shared: scope?.shared_substrate || [],
|
|
277
|
-
frozen: [
|
|
306
|
+
frozen: [`${local}/**`],
|
|
307
|
+
own: [
|
|
308
|
+
// The result this order answers with, and no sibling's: a leg that can write another
|
|
309
|
+
// leg's WorkResult can report work nobody did, and ingest would have no way to tell.
|
|
310
|
+
// With no stem the caller is asking about the operation in general, not about one order,
|
|
311
|
+
// and the honest answer is the whole directory rather than a guess at which file.
|
|
312
|
+
...(ownStem ? [`${local}/results/${ownStem}.json`] : [`${local}/results/**`]),
|
|
313
|
+
`${local}/tasks/TASK-*.md`,
|
|
314
|
+
`${local}/discovery/**`,
|
|
315
|
+
`${local}/spikes/**`,
|
|
316
|
+
],
|
|
278
317
|
};
|
|
279
318
|
case "analyze":
|
|
280
319
|
return { allowed: [`${spec}/**`, `${local}/**`], frozen: [...FROZEN_INTAKE] };
|
|
@@ -315,9 +354,22 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
|
|
|
315
354
|
frozen: [...FROZEN_SPEC_CORE, ...FROZEN_INTAKE, `${scopesDir}/**`, globShared(slug, "project-profile.md")],
|
|
316
355
|
};
|
|
317
356
|
case "evaluate":
|
|
318
|
-
|
|
357
|
+
// THE JUDGE MAY NOT WRITE THE EVIDENCE IT CITES. Its substrate froze the spec, the pitch and
|
|
358
|
+
// the board — and not `t0/verdicts/**`, so a judge could write a green verdict artifact into
|
|
359
|
+
// the canonical directory under the run-trace carve-out and then cite it, correctly hashed.
|
|
360
|
+
// Same inversion as a build leg: the run trace is the kernel's, and `own` names what the
|
|
361
|
+
// judge authors — its report and the envelope that answers its order.
|
|
362
|
+
return {
|
|
363
|
+
allowed: [],
|
|
364
|
+
frozen: [`${spec}/**`, `${local}/**`],
|
|
365
|
+
own: [`${local}/evaluation/**`, ...(ownStem ? [`${local}/results/${ownStem}.json`] : [`${local}/results/**`])],
|
|
366
|
+
};
|
|
319
367
|
case "hunt":
|
|
320
|
-
return {
|
|
368
|
+
return {
|
|
369
|
+
allowed: [],
|
|
370
|
+
frozen: [`${spec}/**`, `${local}/**`],
|
|
371
|
+
own: [`${local}/qa/**`, ...(ownStem ? [`${local}/results/${ownStem}.json`] : [`${local}/results/**`])],
|
|
372
|
+
};
|
|
321
373
|
case "orient":
|
|
322
374
|
return { allowed: [`${local}/orient/**`], frozen: [`${spec}/**`] };
|
|
323
375
|
case "translate":
|
|
@@ -782,7 +834,7 @@ export function compileOrder({
|
|
|
782
834
|
mode,
|
|
783
835
|
...(operation ? { operation } : {}),
|
|
784
836
|
...(interaction ? { interaction } : {}),
|
|
785
|
-
substrate: substrateFor(operation, { slug, specDir, scope }),
|
|
837
|
+
substrate: substrateFor(operation, { slug, specDir, scope, ownStem: suffix }),
|
|
786
838
|
payload: {
|
|
787
839
|
...(scope ? { scope_contract: scope } : {}),
|
|
788
840
|
...(tasks?.length ? { tasks } : {}),
|
package/kernel/lib/argv.mjs
CHANGED
|
@@ -8,11 +8,11 @@
|
|
|
8
8
|
// const SPEC = {
|
|
9
9
|
// _: { arity: 1, name: "scope-contract.json" },
|
|
10
10
|
// round: { type: "int", min: 1, required: true },
|
|
11
|
-
// "no-
|
|
11
|
+
// "no-ratchet": { type: "flag" },
|
|
12
12
|
// };
|
|
13
|
-
// const args = runArgs(SPEC, argv); // args.round, args.
|
|
13
|
+
// const args = runArgs(SPEC, argv); // args.round, args.noRatchet, args._
|
|
14
14
|
//
|
|
15
|
-
// Flag names reach the caller camelCased (`--no-
|
|
15
|
+
// Flag names reach the caller camelCased (`--no-ratchet` → `noRatchet`). Unknown flags are rejected
|
|
16
16
|
// rather than swallowed as positionals: a typo'd `--rounds 2` landing in `_` is the same defect
|
|
17
17
|
// wearing a different hat. Untyped coercion is the failure this guards — `Number(undefined)` is
|
|
18
18
|
// `NaN`, `??` does not catch `NaN`, and a verdict written to `r NaN-a1.json` with exit 0 is
|
|
@@ -35,7 +35,7 @@ export class ArgvError extends Error {
|
|
|
35
35
|
}
|
|
36
36
|
}
|
|
37
37
|
|
|
38
|
-
/** `--
|
|
38
|
+
/** `--attempt-budget` → `attemptBudget`. */
|
|
39
39
|
function camel(name) {
|
|
40
40
|
return name.replace(/-([a-z0-9])/g, (_, c) => c.toUpperCase());
|
|
41
41
|
}
|
package/kernel/lib/paths.mjs
CHANGED
|
@@ -216,8 +216,6 @@ export const trials = (cwd, slug) => join(t0Dir(cwd, slug), "trials.jsonl");
|
|
|
216
216
|
* {@link decisions}, one small file with one writer.
|
|
217
217
|
*/
|
|
218
218
|
export const gates = (cwd, slug) => join(localRoot(cwd, slug), "gates.jsonl");
|
|
219
|
-
/** Finished-scope fixture registry for the seesaw regression check. */
|
|
220
|
-
export const seesawRegistry = (cwd, slug) => join(localRoot(cwd, slug), "seesaw", "registry.json");
|
|
221
219
|
/**
|
|
222
220
|
* The round build gate's verdicts — one immutable artifact per gate run, `r<N>-t<T>.json`.
|
|
223
221
|
*
|
package/kernel/probe/eval.mjs
CHANGED
|
@@ -31,11 +31,11 @@
|
|
|
31
31
|
// cause as its first deviation. A bare `ok: false` reached the operator as a sub-agent that died
|
|
32
32
|
// after retries, while the one sentence naming the actual cause sat in a file nobody was pointed at.
|
|
33
33
|
|
|
34
|
-
import { existsSync, readFileSync, readdirSync } from "node:fs";
|
|
35
|
-
import { join, resolve } from "node:path";
|
|
34
|
+
import { existsSync, readFileSync, readdirSync, realpathSync } from "node:fs";
|
|
35
|
+
import { join, resolve, resolve as resolvePath, sep } from "node:path";
|
|
36
36
|
import { createHash } from "node:crypto";
|
|
37
37
|
import { runArgs } from "../lib/argv.mjs";
|
|
38
|
-
import { resultsDir, scopesDir, readRunId } from "../lib/paths.mjs";
|
|
38
|
+
import { resultsDir, scopesDir, readRunId, verdictsDir } from "../lib/paths.mjs";
|
|
39
39
|
|
|
40
40
|
/** Longest `reason` reported. A deviation is prose written by a worker and can run to paragraphs. */
|
|
41
41
|
const REASON_MAX = 400;
|
|
@@ -89,7 +89,7 @@ export function isScoped(cwd, slug) {
|
|
|
89
89
|
* @returns {(string|null)} A reason phrased for an operator, or null when the file at `path` exists,
|
|
90
90
|
* hashes to the cited `sha256`, and its own `overall` reads "green".
|
|
91
91
|
*/
|
|
92
|
-
function unresolvedCitation(cwd, citation, { round = null, runId = null } = {}) {
|
|
92
|
+
function unresolvedCitation(cwd, citation, { round = null, runId = null, verdicts = null } = {}) {
|
|
93
93
|
const rel = typeof citation?.path === "string" ? citation.path : "";
|
|
94
94
|
if (!rel) return "names no artifact path";
|
|
95
95
|
let text;
|
|
@@ -109,6 +109,19 @@ function unresolvedCitation(cwd, citation, { round = null, runId = null } = {})
|
|
|
109
109
|
try { body = JSON.parse(text); }
|
|
110
110
|
catch { return `cites ${rel}, whose bytes match the hash but do not read as a T0 verdict`; }
|
|
111
111
|
if (body?.overall !== "green") return `cites ${rel}, whose own verdict is "${body?.overall ?? "unknown"}", not green`;
|
|
112
|
+
// INSIDE THIS RUN'S OWN VERDICTS DIRECTORY, resolved — asked of a file that exists, so "does not
|
|
113
|
+
// exist" and "resolves somewhere else" stay different answers. Accepted before this check: a
|
|
114
|
+
// green artifact in the source tree, one outside the project reached by `../`, one by absolute
|
|
115
|
+
// path, and a symlink in the verdicts directory pointing at a forged file. Every one hashed
|
|
116
|
+
// correctly, because a digest says the bytes are the file's and nothing about which file it
|
|
117
|
+
// should have been.
|
|
118
|
+
if (verdicts) {
|
|
119
|
+
const real = (() => { try { return realpathSync.native(resolvePath(cwd, rel)); } catch { return resolvePath(cwd, rel); } })();
|
|
120
|
+
const home = (() => { try { return realpathSync.native(verdicts); } catch { return verdicts; } })();
|
|
121
|
+
if (!real.startsWith(home.endsWith(sep) ? home : home + sep)) {
|
|
122
|
+
return `cites ${rel}, which resolves outside this run's own verdicts directory (${home}) — a verdict cites what this run's own verifier wrote, not a file the judge can reach`;
|
|
123
|
+
}
|
|
124
|
+
}
|
|
112
125
|
// THE ARTIFACT HAS TO BE THE ONE THE CITATION SAYS IT IS. A re-hash proves the bytes are the
|
|
113
126
|
// file's; it says nothing about whose verdict the file holds. A PASS citing scope alpha's green
|
|
114
127
|
// artifact while declaring scope beta, or a prior round's, or a prior run's over the same slug,
|
|
@@ -198,8 +211,9 @@ export function citationProblem(cwd, slug, verdict, { round = null } = {}) {
|
|
|
198
211
|
"cite the T0 verdict it re-hashed (the order lists them under payload.t0_artifacts)";
|
|
199
212
|
}
|
|
200
213
|
const runId = readRunId(cwd, slug);
|
|
214
|
+
const verdicts = verdictsDir(cwd, slug);
|
|
201
215
|
for (const citation of verdict.t0_citations) {
|
|
202
|
-
const reason = unresolvedCitation(cwd, citation, { round, runId });
|
|
216
|
+
const reason = unresolvedCitation(cwd, citation, { round, runId, verdicts });
|
|
203
217
|
if (reason) return `the ${verdict.overall} verdict ${reason} — a T0 citation is re-hashed from disk, never taken on the handed word`;
|
|
204
218
|
}
|
|
205
219
|
return null;
|
package/kernel/reduce/hill.mjs
CHANGED
|
@@ -144,8 +144,8 @@ function committedPhase(hDir, id) {
|
|
|
144
144
|
* - UPHILL_UNKNOWN: open unknowns > 0 in the ledger for this scope — and the floor the scope sits
|
|
145
145
|
* at whenever the ledger has not answered at all, which is where every run legitimately begins
|
|
146
146
|
* - UPHILL_SOLVED: the ledger was read and reports zero open unknowns, no T0-green yet
|
|
147
|
-
* - DOWNHILL_EXECUTION: ≥1 T0-green in a round whose build gate is not red; T1
|
|
148
|
-
* - FINISHED: T1 PASS ∧
|
|
147
|
+
* - DOWNHILL_EXECUTION: ≥1 T0-green in a round whose build gate is not red; T1 pending
|
|
148
|
+
* - FINISHED: T1 PASS ∧ a T0-green that a red round did not invalidate
|
|
149
149
|
*
|
|
150
150
|
* @param {string} cwd - The project root directory.
|
|
151
151
|
* @param {string} slug - The feature slug being built.
|
|
@@ -230,18 +230,16 @@ export function deriveHill(cwd, slug) {
|
|
|
230
230
|
}
|
|
231
231
|
}
|
|
232
232
|
|
|
233
|
-
// 2. T0 facts per scope: has it achieved a green overall verdict
|
|
233
|
+
// 2. T0 facts per scope: has it achieved a green overall verdict this run?
|
|
234
234
|
//
|
|
235
|
-
//
|
|
236
|
-
//
|
|
237
|
-
//
|
|
238
|
-
//
|
|
239
|
-
//
|
|
240
|
-
//
|
|
241
|
-
// verdict counting exactly as before.
|
|
235
|
+
// FINISHED USED TO WAIT ON AN ARM THAT NEVER RAN. The phase required `seesaw.ran && seesaw.pass`,
|
|
236
|
+
// nothing in the codebase ever wrote the registry that arm read, and so no scope in any recorded
|
|
237
|
+
// run reached FINISHED — 38 committed shards on the live consumer, not one of them. The arm was
|
|
238
|
+
// removed in 3.8.0 by decision; the precondition goes with it, and the top phase is reachable
|
|
239
|
+
// again on the evidence that does exist: T1 passed, and a T0 green from a round the build gate
|
|
240
|
+
// did not red.
|
|
242
241
|
const redRounds = redBuildRounds(cwd, slug);
|
|
243
242
|
const t0Facts = {};
|
|
244
|
-
// This run's verdicts only — a prior run's green over the same slug moved this run's dot.
|
|
245
243
|
const hillRunId = readRunId(cwd, slug);
|
|
246
244
|
if (existsSync(vDir)) {
|
|
247
245
|
for (const f of readdirSync(vDir)) {
|
|
@@ -249,41 +247,21 @@ export function deriveHill(cwd, slug) {
|
|
|
249
247
|
try {
|
|
250
248
|
const b = JSON.parse(readFileSync(join(vDir, f), "utf8"));
|
|
251
249
|
if (hillRunId && b.run_id && b.run_id !== hillRunId) continue;
|
|
252
|
-
if (!t0Facts[b.scope_id]) t0Facts[b.scope_id] = { hasGreen: false
|
|
253
|
-
if (b.overall === "green" && !redRounds.has(Number(b.round)))
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
// inferring it from a false `regression` flag. That inference was vacuously true on every
|
|
258
|
-
// green T0 whether or not a seesaw check ever ran: nothing in this codebase currently
|
|
259
|
-
// passes `--seesaw-registry` to `verify t0`, so `seesaw.ran` is always `false` today and
|
|
260
|
-
// `regression` is always `false` too — "not asked" was being read as "clean," letting a
|
|
261
|
-
// scope reach FINISHED on a regression check that had never executed.
|
|
262
|
-
//
|
|
263
|
-
// Betting Table decision (Phase 3.5 / S4): wiring the seesaw registry for real is a
|
|
264
|
-
// genuine feature with a real running cost (re-running every finished scope's fixtures
|
|
265
|
-
// on every later attempt) and is out of proportion to a certification-gap fix. Deferred,
|
|
266
|
-
// not silently dropped — a scope with no registry wired simply cannot reach FINISHED via
|
|
267
|
-
// this path today, which is the honest state of the system: this check was never really
|
|
268
|
-
// gating FINISHED before either.
|
|
269
|
-
if (b.seesaw?.ran && b.seesaw?.pass) {
|
|
270
|
-
t0Facts[b.scope_id].seesawGreen = true;
|
|
271
|
-
}
|
|
272
|
-
}
|
|
273
|
-
} catch (e) {
|
|
274
|
-
// ignore parse errors
|
|
250
|
+
if (!t0Facts[b.scope_id]) t0Facts[b.scope_id] = { hasGreen: false };
|
|
251
|
+
if (b.overall === "green" && !redRounds.has(Number(b.round))) t0Facts[b.scope_id].hasGreen = true;
|
|
252
|
+
} catch {
|
|
253
|
+
// A torn or unreadable verdict proves nothing about the scope — skip it rather than let it
|
|
254
|
+
// decide a phase.
|
|
275
255
|
}
|
|
276
256
|
}
|
|
277
257
|
}
|
|
278
|
-
|
|
279
|
-
// 3. Ledger unknowns per scope — `null` for every scope when the ledger itself was not readable
|
|
280
|
-
// or nothing in it named a scope (see `ledgerUnknowns`). Only a real count can promote.
|
|
258
|
+
|
|
281
259
|
const scopeUnknowns = ledgerUnknowns(ledgerPath, scopes);
|
|
282
260
|
|
|
283
261
|
const report = [];
|
|
284
262
|
for (const s of scopes) {
|
|
285
263
|
const id = s.scope_id;
|
|
286
|
-
const t0 = t0Facts[id] || { hasGreen: false
|
|
264
|
+
const t0 = t0Facts[id] || { hasGreen: false };
|
|
287
265
|
// `null` = the ledger did not answer; a number = it did. `|| 0` collapsed the two.
|
|
288
266
|
const unknowns = scopeUnknowns === null ? null : (scopeUnknowns[id] || 0);
|
|
289
267
|
|
|
@@ -291,7 +269,7 @@ export function deriveHill(cwd, slug) {
|
|
|
291
269
|
// it — including before Orient has filed anything, which is where every run legitimately
|
|
292
270
|
// starts. Only an ANSWERED count of zero promotes to UPHILL_SOLVED; `null` never does.
|
|
293
271
|
let phase = "UPHILL_UNKNOWN";
|
|
294
|
-
if (t1Pass && t0.hasGreen
|
|
272
|
+
if (t1Pass && t0.hasGreen) {
|
|
295
273
|
phase = "FINISHED";
|
|
296
274
|
} else if (t0.hasGreen) {
|
|
297
275
|
phase = "DOWNHILL_EXECUTION";
|
package/kernel/report/export.mjs
CHANGED
|
@@ -129,7 +129,6 @@ function t0Row(a, runId) {
|
|
|
129
129
|
regression: a?.regression ?? null,
|
|
130
130
|
fixtures_green: a?.fixtures_green ?? null,
|
|
131
131
|
db_probe_green: a?.db_probe_green ?? null,
|
|
132
|
-
seesaw_green: a?.seesaw_green ?? null,
|
|
133
132
|
// One field a reader compares, and the block itself stays in the artifact for a human to diff:
|
|
134
133
|
// two rows with the same tree and different env digests are two machines, not a regression.
|
|
135
134
|
env_sha256: a?.env?.env_sha256 ?? null,
|
|
@@ -137,8 +136,6 @@ function t0Row(a, runId) {
|
|
|
137
136
|
tree_dirty: a?.env?.tree?.dirty ?? null,
|
|
138
137
|
fixtures_total: fixtures.length,
|
|
139
138
|
fixtures_passed: fixtures.filter((f) => f?.pass === true).length,
|
|
140
|
-
seesaw_ran: a?.seesaw?.ran ?? null,
|
|
141
|
-
seesaw_failing: Array.isArray(a?.seesaw?.failing) ? a.seesaw.failing.length : null,
|
|
142
139
|
discovered_tasks: Array.isArray(a?.discovered_tasks) ? a.discovered_tasks.length : 0,
|
|
143
140
|
};
|
|
144
141
|
}
|
|
@@ -236,7 +236,7 @@
|
|
|
236
236
|
"to": "HillShard",
|
|
237
237
|
"cardinality": "1:0..1",
|
|
238
238
|
"via": "scope_id",
|
|
239
|
-
"note": "phase derived from T0/T1
|
|
239
|
+
"note": "phase derived from T0/T1 facts, never authored"
|
|
240
240
|
},
|
|
241
241
|
{
|
|
242
242
|
"from": "ScopeContract",
|
|
@@ -252,13 +252,6 @@
|
|
|
252
252
|
"via": "fixtures[] + db_probe",
|
|
253
253
|
"note": "produced by actually running the commands — no agent can fabricate them"
|
|
254
254
|
},
|
|
255
|
-
{
|
|
256
|
-
"from": "T0Artifact",
|
|
257
|
-
"to": "SeesawCheck",
|
|
258
|
-
"cardinality": "1:1",
|
|
259
|
-
"via": "seesaw",
|
|
260
|
-
"note": "the regression half of T0 — every FINISHED scope's fixtures re-run"
|
|
261
|
-
},
|
|
262
255
|
{
|
|
263
256
|
"from": "T0Artifact",
|
|
264
257
|
"to": "AegisTriple",
|
|
@@ -287,13 +280,6 @@
|
|
|
287
280
|
"via": "payload.trial_history[]",
|
|
288
281
|
"note": "inspect(): the loop reads its own recent history, across the round boundary"
|
|
289
282
|
},
|
|
290
|
-
{
|
|
291
|
-
"from": "SeesawRegistry",
|
|
292
|
-
"to": "ScopeContract",
|
|
293
|
-
"cardinality": "1:N",
|
|
294
|
-
"via": "scopes[].scope_id",
|
|
295
|
-
"note": "every FINISHED scope's fixtures re-run on each later attempt"
|
|
296
|
-
},
|
|
297
283
|
{
|
|
298
284
|
"from": "UseCase",
|
|
299
285
|
"to": "Seam",
|
|
@@ -528,7 +514,7 @@
|
|
|
528
514
|
"items": {
|
|
529
515
|
"type": "string"
|
|
530
516
|
},
|
|
531
|
-
"description": "Globs ≥2 scopes both touch —
|
|
517
|
+
"description": "Globs ≥2 scopes both touch — an edit here is read-modify-write, so two scopes that both declare one never build at the same time."
|
|
532
518
|
},
|
|
533
519
|
"append_only": {
|
|
534
520
|
"type": "array",
|
|
@@ -542,7 +528,14 @@
|
|
|
542
528
|
"items": {
|
|
543
529
|
"type": "string"
|
|
544
530
|
},
|
|
545
|
-
"description": "Explicitly untouchable paths (spec core: domain-model, UC Steps, contracts, ux-behavior)."
|
|
531
|
+
"description": "Explicitly untouchable paths (spec core: domain-model, UC Steps, contracts, ux-behavior). A build leg's order freezes the whole run trace here and carves out what it authors through `own`: a freeze that lists the channels a defect happened to name grows one defect at a time and is not a boundary."
|
|
532
|
+
},
|
|
533
|
+
"own": {
|
|
534
|
+
"type": "array",
|
|
535
|
+
"items": {
|
|
536
|
+
"type": "string"
|
|
537
|
+
},
|
|
538
|
+
"description": "Paths this order may write DESPITE a broader freeze — the compiler's grant, derived from the order's own identity and never requested by a worker. For a build leg: its own WorkResult (no sibling's — a leg that can write another's can report work nobody did), its task files, the discovery ledger, its spikes. Checked BEFORE `frozen`, and only against the contract that declared it, so another live order's exception never licenses this write."
|
|
546
539
|
}
|
|
547
540
|
}
|
|
548
541
|
},
|
|
@@ -837,7 +830,7 @@
|
|
|
837
830
|
"items": {
|
|
838
831
|
"type": "string"
|
|
839
832
|
},
|
|
840
|
-
"description": "Files ≥2 scopes both touch — must be declared in BOTH scopes' lists (spec-lint DISJOINT)
|
|
833
|
+
"description": "Files ≥2 scopes both touch — must be declared in BOTH scopes' lists (spec-lint DISJOINT); declaring one costs concurrency, since two scopes that share a path build one at a time."
|
|
841
834
|
},
|
|
842
835
|
"affordance_manifest": {
|
|
843
836
|
"type": "array",
|
|
@@ -865,7 +858,7 @@
|
|
|
865
858
|
"DOWNHILL_EXECUTION",
|
|
866
859
|
"FINISHED"
|
|
867
860
|
],
|
|
868
|
-
"description": "ALWAYS authored as UPHILL_UNKNOWN — self-reported confidence is the risk this closes. The live phase is DERIVED from T0/T1
|
|
861
|
+
"description": "ALWAYS authored as UPHILL_UNKNOWN — self-reported confidence is the risk this closes. The live phase is DERIVED from T0/T1 facts into hill/<scope-id>.yml — facts move dots, not authors."
|
|
869
862
|
},
|
|
870
863
|
"superseded_by": {
|
|
871
864
|
"type": "array",
|
|
@@ -1247,35 +1240,8 @@
|
|
|
1247
1240
|
}
|
|
1248
1241
|
}
|
|
1249
1242
|
},
|
|
1250
|
-
|
|
1251
|
-
"description": "The
|
|
1252
|
-
"x-tier": "EMBEDDED",
|
|
1253
|
-
"type": "object",
|
|
1254
|
-
"properties": {
|
|
1255
|
-
"ran": {
|
|
1256
|
-
"type": "boolean",
|
|
1257
|
-
"description": "false when no registry was passed or the attempt was already red (don't seesaw a red attempt)."
|
|
1258
|
-
},
|
|
1259
|
-
"pass": {
|
|
1260
|
-
"type": "boolean"
|
|
1261
|
-
},
|
|
1262
|
-
"scopes_checked": {
|
|
1263
|
-
"type": "array",
|
|
1264
|
-
"items": {
|
|
1265
|
-
"type": "string"
|
|
1266
|
-
}
|
|
1267
|
-
},
|
|
1268
|
-
"failing": {
|
|
1269
|
-
"type": "array",
|
|
1270
|
-
"items": {
|
|
1271
|
-
"type": "string"
|
|
1272
|
-
},
|
|
1273
|
-
"description": "scope_ids whose fixtures broke."
|
|
1274
|
-
}
|
|
1275
|
-
}
|
|
1276
|
-
},
|
|
1277
|
-
"T0Artifact": {
|
|
1278
|
-
"description": "The mechanical verification verdict for one build attempt — the evidence layer under the LLM judge. Written by harness verify t0 from actually running the scope's fixtures + DB probe + seesaw; zero LLM tokens. spec-evaluator must cite it (sha256) on scoped specs. Red artifacts carry AEGIS triples that become the next attempt's digested_errors.",
|
|
1243
|
+
"T0Artifact": {
|
|
1244
|
+
"description": "The mechanical verification verdict for one build attempt — the evidence layer under the LLM judge. Written by harness verify t0 from actually running the scope's fixtures and DB probe; zero LLM tokens. spec-evaluator must cite it (sha256) on scoped specs. Red artifacts carry AEGIS triples that become the next attempt's digested_errors.",
|
|
1279
1245
|
"x-tier": "LOCAL",
|
|
1280
1246
|
"x-location": ".shapeup/<slug>/t0/verdicts/r<N>-a<M>-t<T>.json (schema_version 2; the unsuffixed r<N>-a<M>.json of schema_version 1 is still readable)",
|
|
1281
1247
|
"x-writer": "harness verify t0",
|
|
@@ -1320,9 +1286,6 @@
|
|
|
1320
1286
|
"db_probe": {
|
|
1321
1287
|
"$ref": "#/$defs/CommandResult"
|
|
1322
1288
|
},
|
|
1323
|
-
"seesaw": {
|
|
1324
|
-
"$ref": "#/$defs/SeesawCheck"
|
|
1325
|
-
},
|
|
1326
1289
|
"fixtures_green": {
|
|
1327
1290
|
"type": "boolean"
|
|
1328
1291
|
},
|
|
@@ -1330,9 +1293,6 @@
|
|
|
1330
1293
|
"type": "boolean",
|
|
1331
1294
|
"description": "true when no probe declared (null probe never counts as failure)."
|
|
1332
1295
|
},
|
|
1333
|
-
"seesaw_green": {
|
|
1334
|
-
"type": "boolean"
|
|
1335
|
-
},
|
|
1336
1296
|
"overall": {
|
|
1337
1297
|
"type": "string",
|
|
1338
1298
|
"enum": [
|
|
@@ -1340,10 +1300,6 @@
|
|
|
1340
1300
|
"red"
|
|
1341
1301
|
]
|
|
1342
1302
|
},
|
|
1343
|
-
"regression": {
|
|
1344
|
-
"type": "boolean",
|
|
1345
|
-
"description": "fixtures+db green but seesaw red — the rollback+retry case."
|
|
1346
|
-
},
|
|
1347
1303
|
"score": {
|
|
1348
1304
|
"$ref": "#/$defs/T0Score",
|
|
1349
1305
|
"description": "schema_version 2+: the comparable outcome vector better() ranks. A reduce over fixtures[] — no new measurement."
|
|
@@ -1377,11 +1333,6 @@
|
|
|
1377
1333
|
"x-tier": "EMBEDDED",
|
|
1378
1334
|
"type": "object",
|
|
1379
1335
|
"properties": {
|
|
1380
|
-
"regressions": {
|
|
1381
|
-
"type": "integer",
|
|
1382
|
-
"minimum": 0,
|
|
1383
|
-
"description": "Previously-FINISHED scopes now failing (seesaw). Dominates: breaking a shipped scope is never an improvement."
|
|
1384
|
-
},
|
|
1385
1336
|
"fixtures_passed": {
|
|
1386
1337
|
"type": "integer",
|
|
1387
1338
|
"minimum": 0
|
|
@@ -1405,7 +1356,6 @@
|
|
|
1405
1356
|
}
|
|
1406
1357
|
},
|
|
1407
1358
|
"required": [
|
|
1408
|
-
"regressions",
|
|
1409
1359
|
"fixtures_passed",
|
|
1410
1360
|
"fixtures_total"
|
|
1411
1361
|
]
|
|
@@ -1612,34 +1562,7 @@
|
|
|
1612
1562
|
}
|
|
1613
1563
|
}
|
|
1614
1564
|
},
|
|
1615
|
-
|
|
1616
|
-
"description": "The fixture registry of every FINISHED scope — what seesawCheck re-runs on each later attempt so a new scope cannot silently break a shipped one — a regression mistaken for progress is the pathology the seesaw exists for.",
|
|
1617
|
-
"x-tier": "LOCAL",
|
|
1618
|
-
"x-location": ".shapeup/<slug>/seesaw/registry.json",
|
|
1619
|
-
"x-writer": "tech-lead (when a scope reaches FINISHED)",
|
|
1620
|
-
"x-readers": "harness verify t0",
|
|
1621
|
-
"type": "object",
|
|
1622
|
-
"properties": {
|
|
1623
|
-
"scopes": {
|
|
1624
|
-
"type": "array",
|
|
1625
|
-
"items": {
|
|
1626
|
-
"type": "object",
|
|
1627
|
-
"properties": {
|
|
1628
|
-
"scope_id": {
|
|
1629
|
-
"type": "string"
|
|
1630
|
-
},
|
|
1631
|
-
"fixtures": {
|
|
1632
|
-
"type": "array",
|
|
1633
|
-
"items": {
|
|
1634
|
-
"type": "string"
|
|
1635
|
-
}
|
|
1636
|
-
}
|
|
1637
|
-
}
|
|
1638
|
-
}
|
|
1639
|
-
}
|
|
1640
|
-
}
|
|
1641
|
-
},
|
|
1642
|
-
"VerdictLedgerLine": {
|
|
1565
|
+
"VerdictLedgerLine": {
|
|
1643
1566
|
"description": "One appended line of judge history (JSONL). Never rewritten — flips across runs are detected here and force confidence low. run auto-increments per append batch.",
|
|
1644
1567
|
"x-tier": "LOCAL",
|
|
1645
1568
|
"x-location": ".shapeup/<slug>/evaluation/.verdicts-<target>.jsonl",
|
|
@@ -1683,7 +1606,7 @@
|
|
|
1683
1606
|
}
|
|
1684
1607
|
},
|
|
1685
1608
|
"HillShard": {
|
|
1686
|
-
"description": "One scope's hill position — DERIVED, never self-reported: UPHILL_UNKNOWN (open unknowns > 0) → UPHILL_SOLVED (unknowns 0, no T0-green yet) → DOWNHILL_EXECUTION (≥1 T0-green; T1
|
|
1609
|
+
"description": "One scope's hill position — DERIVED, never self-reported: UPHILL_UNKNOWN (open unknowns > 0) → UPHILL_SOLVED (unknowns 0, no T0-green yet) → DOWNHILL_EXECUTION (≥1 T0-green from a round that built; T1 pending) → FINISHED (T1 PASS ∧ that T0-green). Single-writer = whoever holds that scope's branch. Progress is reported by hill position, never task counts.",
|
|
1687
1610
|
"x-tier": "SHARED",
|
|
1688
1611
|
"x-location": "shapeup/<slug>/hill/<scope-id>.yml",
|
|
1689
1612
|
"x-writer": "tech-lead (GATE L2 derivation)",
|
package/kernel/verify/env.mjs
CHANGED
|
@@ -52,6 +52,16 @@ const UNSET = "<unset>";
|
|
|
52
52
|
* @param {string} cmd - A shell command line.
|
|
53
53
|
* @returns {string[]} Invoked tokens, in order, without duplicates.
|
|
54
54
|
*/
|
|
55
|
+
export function cdTargets(cmd) {
|
|
56
|
+
const out = [];
|
|
57
|
+
for (const seg of String(cmd || "").split(/&&|;|\|\|/).map((s) => s.trim()).filter(Boolean)) {
|
|
58
|
+
const m = seg.match(/^cd\s+(?:"([^"]+)"|'([^']+)'|(\S+))/);
|
|
59
|
+
const dir = m && (m[1] || m[2] || m[3]);
|
|
60
|
+
if (dir && !dir.startsWith("-") && !out.includes(dir)) out.push(dir);
|
|
61
|
+
}
|
|
62
|
+
return out;
|
|
63
|
+
}
|
|
64
|
+
|
|
55
65
|
export function invokedTokens(cmd) {
|
|
56
66
|
const out = [];
|
|
57
67
|
for (const seg of String(cmd || "").split(/&&|;|\|\|/).map((s) => s.trim()).filter(Boolean)) {
|
|
@@ -96,7 +106,8 @@ function treeState(cwd) {
|
|
|
96
106
|
* path, which is exactly the mechanism that made one tree build three ways.
|
|
97
107
|
*
|
|
98
108
|
* `null` means the profile declared nothing — NOT that there are none. "Not asked" and "none" are
|
|
99
|
-
* different facts, and collapsing them is the mistake the
|
|
109
|
+
* different facts, and collapsing them is the mistake that kept the hill's top phase shut for the
|
|
110
|
+
* life of the seesaw arm: "not asked" was recorded the same way as "nothing wrong".
|
|
100
111
|
*
|
|
101
112
|
* @param {(string|null)} profilePath - `shapeup/<slug>/project-profile.md`, when the caller knows it.
|
|
102
113
|
* @returns {(object[]|null)} One entry per declared cache, or null when none is declared.
|
|
@@ -162,12 +173,19 @@ export function environmentFingerprint(rawCwd, { commands = [], profilePath = nu
|
|
|
162
173
|
cwd,
|
|
163
174
|
tree: treeState(cwd),
|
|
164
175
|
toolchain: tokens.map((bin) => ({ bin, path: resolveBin(bin, cwd) })),
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
176
|
+
// The root AND wherever the commands actually run. Measured on a consumer whose fixtures are
|
|
177
|
+
// `cd app && …`: the lockfile that decides what the build resolves lives in `app/`, and a scan
|
|
178
|
+
// of the project root alone recorded an empty list beside a build whose dependencies were the
|
|
179
|
+
// whole question.
|
|
180
|
+
lockfiles: [...new Set(["", ...commands.flatMap((c) => cdTargets(c))])]
|
|
181
|
+
.flatMap((sub) => LOCKFILES
|
|
182
|
+
.map((f) => (sub ? `${sub.replace(/\/+$/, "")}/${f}` : f))
|
|
183
|
+
.filter((rel) => existsSync(join(cwd, rel)))
|
|
184
|
+
.map((rel) => {
|
|
185
|
+
try { return { file: rel, sha256: sha256(readFileSync(join(cwd, rel))) }; }
|
|
186
|
+
catch { return { file: rel, sha256: null }; }
|
|
187
|
+
}))
|
|
188
|
+
.filter((l, i, all) => all.findIndex((x) => x.file === l.file) === i),
|
|
171
189
|
caches: declaredCaches(profilePath),
|
|
172
190
|
env: {
|
|
173
191
|
allowlist: ENV_ALLOWLIST,
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
//
|
|
5
5
|
// The attempt loop branched a red T0 two ways, and only one of them reverted anything:
|
|
6
6
|
//
|
|
7
|
-
// • a
|
|
7
|
+
// • a trial that scored worse than the incumbent → `git stash push -u`;
|
|
8
8
|
// • a red on the scope's OWN fixtures → "loop to the next attempt", and no revert at all.
|
|
9
9
|
//
|
|
10
10
|
// So the failing tree stayed on the branch, and attempt N+1's fresh, zero-memory subagent began
|
package/kernel/verify/spec.mjs
CHANGED
|
@@ -422,6 +422,9 @@ export function lintScopeAnchors({ scopes, specDir: specRoot, reqIds = null, tas
|
|
|
422
422
|
return findings;
|
|
423
423
|
}
|
|
424
424
|
|
|
425
|
+
/** Registry sources that name the pitch's own out-of-scope section, in the spellings pitches use. */
|
|
426
|
+
const NOGO_SOURCE = /\bno[-\s]?gos?\b|\bnon[-\s]?goals?\b|\bout[-\s]of[-\s]scope\b|\bwill not build\b/i;
|
|
427
|
+
|
|
425
428
|
/**
|
|
426
429
|
* REQ-UNCOVERED — a live requirement that nothing in the plan reaches.
|
|
427
430
|
*
|
|
@@ -451,6 +454,57 @@ export function lintScopeAnchors({ scopes, specDir: specRoot, reqIds = null, tas
|
|
|
451
454
|
* @returns {Array<{rule:string, level:("red"|"warn"), scope:string, detail:string}>} One red per
|
|
452
455
|
* uncovered live requirement; [] when every one is graded, claimed or cut.
|
|
453
456
|
*/
|
|
457
|
+
/**
|
|
458
|
+
* REQ-NARRATED — a committed spec file stating the requirement-coverage verdict as fact.
|
|
459
|
+
*
|
|
460
|
+
* The Health Dashboard's `Coverage` row is about USE CASES and tasks, derived by inverting each
|
|
461
|
+
* task's `use_case_refs` over the local board. Measured on a consumer, a worker filled its Signal
|
|
462
|
+
* cell with a different claim entirely — *"every registered non-CUT REQ-id (REQ-1 … REQ-7) reaches
|
|
463
|
+
* an AC carrying `(covers: REQ-…)`"*, with a 🟢 beside it — while a grep for `covers:` across the
|
|
464
|
+
* whole spec folder returned that sentence and nothing else. Not one acceptance criterion carried
|
|
465
|
+
* the clause, and the run's own derived report said `0/11 PASS`.
|
|
466
|
+
*
|
|
467
|
+
* `AGENTS.md` names the invariant this breaks: the requirements matrix is a projection, never a
|
|
468
|
+
* verdict, derived from files for one named run and never narrated. The rule is the narrow,
|
|
469
|
+
* checkable form of it — a dashboard Coverage row in a committed file may not name a REQ id — and
|
|
470
|
+
* it cannot fire on the legitimate signal, which counts use cases and tasks.
|
|
471
|
+
*
|
|
472
|
+
* @param {{cwd:string, slug:string}} opts - Working root and feature slug.
|
|
473
|
+
* @returns {object[]} Findings, one per offending line.
|
|
474
|
+
*/
|
|
475
|
+
export function lintNarratedCoverage({ cwd, slug }) {
|
|
476
|
+
const findings = [];
|
|
477
|
+
const dir = join(sharedRoot(cwd, slug), "spec");
|
|
478
|
+
let files;
|
|
479
|
+
try { files = readdirSync(dir).filter((f) => f.endsWith(".md")); } catch { return findings; }
|
|
480
|
+
for (const f of files) {
|
|
481
|
+
let lines;
|
|
482
|
+
try { lines = readFileSync(join(dir, f), "utf8").split(/\r?\n/); } catch { continue; }
|
|
483
|
+
lines.forEach((line, i) => {
|
|
484
|
+
if (!/^\|\s*Coverage\s*\|/i.test(line.trim())) return;
|
|
485
|
+
const named = [...line.matchAll(/\bREQ-\d+/g)].map((m) => m[0]);
|
|
486
|
+
if (!named.length) return;
|
|
487
|
+
findings.push({ rule: "REQ-NARRATED", level: "red", scope: `${f}:${i + 1}`, detail:
|
|
488
|
+
`${f}:${i + 1} states the requirement-coverage verdict in a committed file, naming ${named.slice(0, 3).join(", ")}` +
|
|
489
|
+
`${named.length > 3 ? ` (+${named.length - 3})` : ""}. That row is the UC × Task indicator; the ` +
|
|
490
|
+
"REQ → AC → criterion → verdict state is a projection derived per run (probe requirements), " +
|
|
491
|
+
"never a claim a committed artifact may make — a reader who checks the file finds corroboration " +
|
|
492
|
+
"for something no run measured. Say what the use cases and tasks show, and leave the requirement " +
|
|
493
|
+
"matrix to the run that derives it." });
|
|
494
|
+
});
|
|
495
|
+
}
|
|
496
|
+
return findings;
|
|
497
|
+
}
|
|
498
|
+
|
|
499
|
+
/**
|
|
500
|
+
* REQ-UNCOVERED and REQ-NOGO — the registry's two ways of being wrong about what ships.
|
|
501
|
+
*
|
|
502
|
+
* @param {{clauses:object[], board:object[], scopes:object[]}} opts - The parsed registry, the
|
|
503
|
+
* board `readBoard` produced (its acceptance criteria carry the covers clauses), and the scope
|
|
504
|
+
* contracts.
|
|
505
|
+
* @returns {object[]} Findings, most specific first: a no-go registered as covered is reported as
|
|
506
|
+
* itself rather than as the coverage gap it inevitably becomes.
|
|
507
|
+
*/
|
|
454
508
|
export function lintRequirementCoverage({ clauses = [], board = [], scopes = [] }) {
|
|
455
509
|
const findings = [];
|
|
456
510
|
const graded = coveredReqIds(board);
|
|
@@ -459,6 +513,21 @@ export function lintRequirementCoverage({ clauses = [], board = [], scopes = []
|
|
|
459
513
|
const claimed = new Set();
|
|
460
514
|
for (const s of scopes) for (const r of s.covers || []) claimed.add(reqId(r).toUpperCase());
|
|
461
515
|
for (const c of clauses) {
|
|
516
|
+
// A NO-GO IS A CONSTRAINT, NOT A DELIVERABLE, and marking one `covered` asserts something that
|
|
517
|
+
// cannot be true: nothing grades "do not build a settings screen". Measured on a consumer — a
|
|
518
|
+
// coverage dispatch lifted seven clauses out of the pitch's No-gos section, registered each as
|
|
519
|
+
// covered, and L1b then refused the run with seven REQ-UNCOVERED findings, correctly and
|
|
520
|
+
// unavoidably. Reported here as itself, so the operator reads one cause instead of seven
|
|
521
|
+
// symptoms, and named before REQ-UNCOVERED can fire on the same row.
|
|
522
|
+
if (c.status === "covered" && NOGO_SOURCE.test(c.source || "")) {
|
|
523
|
+
findings.push({ rule: "REQ-NOGO", level: "red", scope: c.id, detail:
|
|
524
|
+
`${c.id} ← ${c.source} registers a NO-GO as a covered requirement — "${(c.clause || "").slice(0, 60)}". ` +
|
|
525
|
+
"A no-go is a constraint the shape deliberately does not build, so no acceptance criterion can " +
|
|
526
|
+
"grade it and nothing downstream can ever turn it green. Mark it CUT (PO-approved) in " +
|
|
527
|
+
"requirements.md — the family that already means deliberately-not-built — or drop the row and give " +
|
|
528
|
+
"the breach a Test Surface row (TS-NOGO-NN) instead, which is the channel that does grade one." });
|
|
529
|
+
continue;
|
|
530
|
+
}
|
|
462
531
|
if (c.status !== "covered") continue; // CUT (PO-approved) — an answer on the record, not a gap
|
|
463
532
|
const id = c.id.toUpperCase();
|
|
464
533
|
if (graded.has(c.id) || claimed.has(id)) continue;
|
|
@@ -851,6 +920,7 @@ export function lint({ cwd, slug }) {
|
|
|
851
920
|
// `readBoard`, not the `tasks` above: only the compile-order parser carries acceptance_criteria.
|
|
852
921
|
...lintRequirementCoverage({ clauses: reqClauses, board: readBoard(cwd, slug), scopes }),
|
|
853
922
|
...lintCommittedTier({ cwd, slug }),
|
|
923
|
+
...lintNarratedCoverage({ cwd, slug }),
|
|
854
924
|
...lintStructure({ specDir: specRoot, tasks, intakeContent }),
|
|
855
925
|
...(() => {
|
|
856
926
|
const bbText = runBreadboard(cwd, slug, intakeContent);
|
package/kernel/verify/t0.mjs
CHANGED
|
@@ -1,8 +1,7 @@
|
|
|
1
1
|
#!/usr/bin/env node
|
|
2
2
|
// T0 mechanical verification layer.
|
|
3
3
|
//
|
|
4
|
-
// Runs a scope's e2e fixtures
|
|
5
|
-
// regression check (re-runs every FINISHED scope's fixtures from the registry). Writes one
|
|
4
|
+
// Runs a scope's e2e fixtures and its DB probe (zero LLM tokens). Writes one
|
|
6
5
|
// verdict artifact per attempt that spec-evaluator (T1) must cite; a verdict without it is
|
|
7
6
|
// structurally invalid. No agent can fabricate this file's contents
|
|
8
7
|
// because it is produced by actually running the commands.
|
|
@@ -30,7 +29,7 @@
|
|
|
30
29
|
//
|
|
31
30
|
// Usage:
|
|
32
31
|
// node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify t0 <scope-contract.json> \
|
|
33
|
-
// --round N --attempt M [--cwd <dir>] [--out <dir>]
|
|
32
|
+
// --round N --attempt M [--cwd <dir>] [--out <dir>]
|
|
34
33
|
// [--no-ratchet]
|
|
35
34
|
//
|
|
36
35
|
// Exit code: 0 = overall green, 1 = overall red (mirrors the oracle convention), 2 = bad argv.
|
|
@@ -185,71 +184,45 @@ export function runDbProbe(dbProbeCmd, cwd) {
|
|
|
185
184
|
}
|
|
186
185
|
|
|
187
186
|
/**
|
|
188
|
-
*
|
|
189
|
-
*
|
|
190
|
-
*
|
|
191
|
-
*
|
|
192
|
-
*
|
|
193
|
-
*
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
} catch {
|
|
203
|
-
return { ran: false, pass: true, scopes_checked: [], failing: [], error: "registry unparsable" };
|
|
204
|
-
}
|
|
205
|
-
const scopes = registry.scopes || [];
|
|
206
|
-
const failing = [];
|
|
207
|
-
for (const s of scopes) {
|
|
208
|
-
const { pass } = runFixtures(s.fixtures, cwd);
|
|
209
|
-
if (!pass) failing.push(s.scope_id);
|
|
210
|
-
}
|
|
211
|
-
return { ran: true, pass: failing.length === 0, scopes_checked: scopes.map((s) => s.scope_id), failing };
|
|
212
|
-
}
|
|
213
|
-
|
|
214
|
-
/**
|
|
215
|
-
* Combine fixtures + DB probe + seesaw into the overall T0 verdict.
|
|
216
|
-
* @param {{fixtures:{pass:boolean}, dbProbe:({pass:boolean}|null),
|
|
217
|
-
* seesaw:{ran:boolean,pass:boolean}}} parts - The three sub-results.
|
|
218
|
-
* @returns {{fixtures_green:boolean, db_probe_green:boolean, seesaw_green:boolean,
|
|
219
|
-
* overall:("green"|"red"), regression:boolean}} Per-arm greens, the overall verdict (green iff
|
|
220
|
-
* all three), and `regression` = fixtures+db green but seesaw red (the rollback-and-retry case).
|
|
187
|
+
* Combine the fixtures and the DB probe into the overall T0 verdict.
|
|
188
|
+
*
|
|
189
|
+
* THE SEESAW ARM IS GONE (3.8.0), by a Betting Table decision rather than by neglect. It was
|
|
190
|
+
* declared in the schema, the docs and this function, and nothing ever wrote the registry it read,
|
|
191
|
+
* so it never ran once in any recorded run — while its absence held the hill's top phase shut:
|
|
192
|
+
* FINISHED required `seesaw.ran && seesaw.pass`, and the 38 committed hill shards across the live
|
|
193
|
+
* consumer's features contain no FINISHED at all. A cross-scope regression is still caught by the
|
|
194
|
+
* round build gate, which builds and launches the whole feature once per round; what the arm would
|
|
195
|
+
* have added is attribution and an earlier signal, at the price of re-running every finished
|
|
196
|
+
* scope's fixtures on every attempt — minutes per attempt on an eighteen-scope feature.
|
|
197
|
+
*
|
|
198
|
+
* @param {{fixtures:{pass:boolean}, dbProbe:({pass:boolean}|null)}} parts - The two sub-results.
|
|
199
|
+
* @returns {{fixtures_green:boolean, db_probe_green:boolean, overall:("green"|"red")}} Per-arm
|
|
200
|
+
* greens and the overall verdict, green iff both.
|
|
221
201
|
*/
|
|
222
|
-
export function computeVerdict({ fixtures, dbProbe
|
|
202
|
+
export function computeVerdict({ fixtures, dbProbe }) {
|
|
223
203
|
const fixturesGreen = fixtures.pass;
|
|
224
204
|
const dbGreen = dbProbe === null || dbProbe.pass;
|
|
225
|
-
const seesawGreen = !seesaw.ran || seesaw.pass;
|
|
226
205
|
return {
|
|
227
206
|
fixtures_green: fixturesGreen,
|
|
228
207
|
db_probe_green: dbGreen,
|
|
229
|
-
|
|
230
|
-
overall: fixturesGreen && dbGreen && seesawGreen ? "green" : "red",
|
|
231
|
-
// A regression is specifically fixtures/db green but seesaw red — the case that should
|
|
232
|
-
// trigger rollback+retry (spec §3.5) rather than "go fix the new scope's own bug".
|
|
233
|
-
regression: fixturesGreen && dbGreen && !seesawGreen,
|
|
208
|
+
overall: fixturesGreen && dbGreen ? "green" : "red",
|
|
234
209
|
};
|
|
235
210
|
}
|
|
236
211
|
|
|
237
212
|
/**
|
|
238
|
-
* The comparable T0 outcome — a VECTOR, not a float, because the
|
|
213
|
+
* The comparable T0 outcome — a VECTOR, not a float, because the arms are not fungible.
|
|
239
214
|
*
|
|
240
215
|
* Every number here is a reduce over data `writeArtifact` already persists (`fixtures:
|
|
241
216
|
* [{cmd, exit, pass}]`). Nothing new is measured; a number that has always been on disk is
|
|
242
217
|
* finally counted.
|
|
243
218
|
*
|
|
244
|
-
* @param {{fixtures:{results:Array<{pass:boolean}>}, dbProbe:({pass:boolean}|null)
|
|
245
|
-
*
|
|
246
|
-
* @returns {{
|
|
247
|
-
*
|
|
248
|
-
* is never a failure — only an absence.
|
|
219
|
+
* @param {{fixtures:{results:Array<{pass:boolean}>}, dbProbe:({pass:boolean}|null)}} parts - The
|
|
220
|
+
* two T0 sub-results.
|
|
221
|
+
* @returns {{fixtures_passed:number, fixtures_total:number, db_probe:(0|1|null)}} The score vector.
|
|
222
|
+
* `db_probe` is null when no probe is declared, which is never a failure — only an absence.
|
|
249
223
|
*/
|
|
250
|
-
export function score({ fixtures, dbProbe
|
|
224
|
+
export function score({ fixtures, dbProbe }) {
|
|
251
225
|
return {
|
|
252
|
-
regressions: seesaw?.ran ? (seesaw.failing || []).length : 0,
|
|
253
226
|
fixtures_passed: fixtures.results.filter((r) => r.pass).length,
|
|
254
227
|
fixtures_total: fixtures.results.length,
|
|
255
228
|
db_probe: dbProbe === null || dbProbe === undefined ? null : (dbProbe.pass ? 1 : 0),
|
|
@@ -262,23 +235,19 @@ export function score({ fixtures, dbProbe, seesaw }) {
|
|
|
262
235
|
* Three decisions worth defending:
|
|
263
236
|
* • A TIE IS NOT BETTER. A tie that counted as an improvement would make a sawtooth look like a
|
|
264
237
|
* ratchet, and the whole point of the Day-1 measurement is to tell those two apart.
|
|
265
|
-
* • REGRESSIONS DOMINATE. Breaking a previously-finished scope is never an improvement, whatever
|
|
266
|
-
* the new scope's fixtures did. This is what lets the old seesaw branch collapse into the
|
|
267
|
-
* general rule rather than needing a special case.
|
|
268
238
|
* • DIFFERENT `fixtures_total` IS INCOMPARABLE, not worse. A re-slice changes the
|
|
269
239
|
* denominator; comparing across it is a category error, so the ratchet treats it as a baseline
|
|
270
240
|
* reset (`rebased`) rather than issuing a false verdict.
|
|
271
241
|
*
|
|
272
|
-
* @param {{
|
|
242
|
+
* @param {{fixtures_passed:number, fixtures_total:number,
|
|
273
243
|
* db_probe:(0|1|null)}} next - The candidate score.
|
|
274
|
-
* @param {({
|
|
244
|
+
* @param {({fixtures_passed:number, fixtures_total:number,
|
|
275
245
|
* db_probe:(0|1|null)}|null)} current - The incumbent score, or null for the first trial.
|
|
276
246
|
* @returns {(boolean|null)} true = strictly better · false = not better · null = incomparable.
|
|
277
247
|
*/
|
|
278
248
|
export function better(next, current) {
|
|
279
249
|
if (current === null || current === undefined) return true; // baseline
|
|
280
250
|
if (next.fixtures_total !== current.fixtures_total) return null; // the contract changed
|
|
281
|
-
if (next.regressions !== current.regressions) return next.regressions < current.regressions;
|
|
282
251
|
if (next.fixtures_passed !== current.fixtures_passed) return next.fixtures_passed > current.fixtures_passed;
|
|
283
252
|
if (next.db_probe !== current.db_probe) return (next.db_probe ?? 0) > (current.db_probe ?? 0);
|
|
284
253
|
// EVERY COMPONENT TIES. What that means depends entirely on whether the incumbent was green.
|
|
@@ -305,17 +274,17 @@ export function better(next, current) {
|
|
|
305
274
|
}
|
|
306
275
|
|
|
307
276
|
/**
|
|
308
|
-
* Is this score a clean pass — every fixture passing, none of them absent
|
|
277
|
+
* Is this score a clean pass — every fixture passing, and none of them absent?
|
|
309
278
|
*
|
|
310
279
|
* `fixtures_total > 0` is load-bearing: a scope with no fixtures has nothing to be green ABOUT, and
|
|
311
280
|
* treating its empty score as a pass is the same absence-reads-as-success mistake `runFixtures`
|
|
312
281
|
* made one function above.
|
|
313
282
|
*
|
|
314
|
-
* @param {{
|
|
283
|
+
* @param {{fixtures_passed:number, fixtures_total:number}} s - A trial score.
|
|
315
284
|
* @returns {boolean} True when the score represents a real, complete pass.
|
|
316
285
|
*/
|
|
317
286
|
function isGreenScore(s) {
|
|
318
|
-
return s.
|
|
287
|
+
return s.fixtures_total > 0 && s.fixtures_passed === s.fixtures_total;
|
|
319
288
|
}
|
|
320
289
|
|
|
321
290
|
/**
|
|
@@ -343,7 +312,7 @@ export function decideStatus(verdict, crashed) {
|
|
|
343
312
|
* Human-readable one-line summary of a score change, for the trial row's `delta` field.
|
|
344
313
|
* @param {object} next - The candidate score.
|
|
345
314
|
* @param {(object|null)} current - The incumbent score, or null.
|
|
346
|
-
* @returns {string} e.g. "+2 fixtures", "
|
|
315
|
+
* @returns {string} e.g. "+2 fixtures", "-1 db_probe", "baseline", "no change".
|
|
347
316
|
*/
|
|
348
317
|
export function describeDelta(next, current) {
|
|
349
318
|
if (!current) return "baseline";
|
|
@@ -351,10 +320,8 @@ export function describeDelta(next, current) {
|
|
|
351
320
|
return `denominator ${current.fixtures_total} → ${next.fixtures_total}`;
|
|
352
321
|
}
|
|
353
322
|
const parts = [];
|
|
354
|
-
const dr = next.regressions - current.regressions;
|
|
355
323
|
const df = next.fixtures_passed - current.fixtures_passed;
|
|
356
324
|
const dp = (next.db_probe ?? 0) - (current.db_probe ?? 0);
|
|
357
|
-
if (dr) parts.push(`${dr > 0 ? "+" : ""}${dr} regression${Math.abs(dr) === 1 ? "" : "s"}`);
|
|
358
325
|
if (df) parts.push(`${df > 0 ? "+" : ""}${df} fixture${Math.abs(df) === 1 ? "" : "s"}`);
|
|
359
326
|
if (dp) parts.push(`${dp > 0 ? "+" : ""}${dp} db_probe`);
|
|
360
327
|
return parts.length ? parts.join(", ") : "no change";
|
|
@@ -449,7 +416,7 @@ function sha256(text) {
|
|
|
449
416
|
* every superseded object remains addressable).
|
|
450
417
|
*
|
|
451
418
|
* WHAT THIS REPLACED, and why the remedy is `wx` rather than a guard. The address used to be
|
|
452
|
-
* `r<round>-a<attempt>.json`, written with a bare `writeFileSync` — and on a
|
|
419
|
+
* `r<round>-a<attempt>.json`, written with a bare `writeFileSync` — and on a revert-and-retry the
|
|
453
420
|
* protocol says stash, then RETRY THIS ATTEMPT, same attempt number. The address had no term for
|
|
454
421
|
* the retry, so the artifact recording the regression was silently replaced by the one recording
|
|
455
422
|
* the recovery, at the same path. Reproduced against the shipped script: two runs at
|
|
@@ -497,14 +464,12 @@ export function writeArtifact(outDir, round, attempt, verdictBody) {
|
|
|
497
464
|
/** The typed argv contract (see `./lib/argv.mjs`). */
|
|
498
465
|
export const ARGV_SPEC = {
|
|
499
466
|
usage: "harness.mjs verify t0 <scope-contract.json> --round N --attempt M [--cwd <dir>] [--out <dir>] " +
|
|
500
|
-
"[--
|
|
467
|
+
"[--no-ratchet]",
|
|
501
468
|
_: { arity: 1, max: 1, name: "scope-contract.json" },
|
|
502
469
|
round: { type: "int", min: 1, required: true },
|
|
503
470
|
attempt: { type: "int", min: 1, required: true },
|
|
504
471
|
cwd: { type: "path" },
|
|
505
472
|
out: { type: "path" },
|
|
506
|
-
"seesaw-registry": { type: "path" },
|
|
507
|
-
"no-seesaw": { type: "flag" },
|
|
508
473
|
"no-ratchet": { type: "flag" },
|
|
509
474
|
};
|
|
510
475
|
|
|
@@ -547,14 +512,7 @@ export async function cli(rawArgv) {
|
|
|
547
512
|
|
|
548
513
|
const fixtures = runFixtures(contract.e2e_verification_fixtures, cwd);
|
|
549
514
|
const dbProbe = runDbProbe(contract.db_probe, cwd);
|
|
550
|
-
|
|
551
|
-
// standalone CLI use without it simply skips the seesaw check rather than guessing a path.
|
|
552
|
-
const seesawRegistry = args.noSeesaw ? null : args.seesawRegistry || null;
|
|
553
|
-
const seesaw = args.noSeesaw || fixtures.pass === false
|
|
554
|
-
? { ran: false, pass: true, scopes_checked: [], failing: [] } // don't seesaw on an already-red attempt
|
|
555
|
-
: seesawCheck(seesawRegistry, cwd);
|
|
556
|
-
|
|
557
|
-
const verdict = computeVerdict({ fixtures, dbProbe, seesaw });
|
|
515
|
+
const verdict = computeVerdict({ fixtures, dbProbe });
|
|
558
516
|
const discovered = verdict.overall === "red" ? digestFailures({ fixtures, dbProbe }) : [];
|
|
559
517
|
|
|
560
518
|
// ---- the ratchet ---------------------------------------------------------------------
|
|
@@ -564,7 +522,7 @@ export async function cli(rawArgv) {
|
|
|
564
522
|
const trialsPath = join(outDir, "t0", "trials.jsonl");
|
|
565
523
|
const priorTrials = readTrials(trialsPath).filter((t) => t.scope_id === contract.scope_id);
|
|
566
524
|
const baseline = [...priorTrials].reverse().find((t) => t.status === "kept" || t.status === "rebased") || null;
|
|
567
|
-
const s = score({ fixtures, dbProbe
|
|
525
|
+
const s = score({ fixtures, dbProbe });
|
|
568
526
|
const verdictBetter = better(s, baseline ? baseline.score : null);
|
|
569
527
|
const crashed = fixtures.results.some((r) => r.error) || !!dbProbe?.error;
|
|
570
528
|
const { status, action } = decideStatus(verdictBetter, crashed);
|
|
@@ -589,7 +547,6 @@ export async function cli(rawArgv) {
|
|
|
589
547
|
// could not tell apart, and why `exit` still reads the way it always did.
|
|
590
548
|
fixtures: fixtures.results.map((r) => commandEvidence(r)),
|
|
591
549
|
db_probe: commandEvidence(dbProbe),
|
|
592
|
-
seesaw,
|
|
593
550
|
...verdict,
|
|
594
551
|
score: s,
|
|
595
552
|
discovered_tasks: discovered,
|
|
@@ -645,7 +602,7 @@ export async function cli(rawArgv) {
|
|
|
645
602
|
appendTrial(trialsPath, row);
|
|
646
603
|
|
|
647
604
|
console.log(JSON.stringify({
|
|
648
|
-
path, sha256: hash, trial, overall: verdict.overall,
|
|
605
|
+
path, sha256: hash, trial, overall: verdict.overall,
|
|
649
606
|
score: s, status, baseline_trial: row.baseline_trial, delta: row.delta,
|
|
650
607
|
tree_ref: row.tree_ref ?? keptRef(contract.scope_id),
|
|
651
608
|
}, null, 2));
|
package/package.json
CHANGED
|
@@ -26,6 +26,11 @@ depends_on:
|
|
|
26
26
|
|
|
27
27
|
| Indicator | Status | Signal |
|
|
28
28
|
|-----------|--------|--------|
|
|
29
|
+
<!-- Coverage here is USE CASES × TASKS, derived by inverting each task's use_case_refs over the
|
|
30
|
+
local board. It is NOT requirement coverage: whether every REQ-id reaches an acceptance
|
|
31
|
+
criterion that a judge graded is a projection derived per run (`probe requirements`), and a
|
|
32
|
+
committed file that states it is corroborating something no run measured. Name REQ ids in this
|
|
33
|
+
row and spec-lint reds it (REQ-NARRATED). Count use cases and tasks; say nothing about REQ. -->
|
|
29
34
|
| Coverage | COVERAGE_STATUS | COVERAGE_SIGNAL |
|
|
30
35
|
| Risk | RISK_STATUS | RISK_SIGNAL |
|
|
31
36
|
| Dependency | DEPENDENCY_STATUS | DEPENDENCY_SIGNAL |
|
|
@@ -614,7 +614,7 @@ or entirely `apps/api/**` with no cross-layer flow is the PA1 failure mode — r
|
|
|
614
614
|
}
|
|
615
615
|
```
|
|
616
616
|
`hill_phase` is always written `UPHILL_UNKNOWN` at generation time — it is derived later from
|
|
617
|
-
mechanical T0/T1
|
|
617
|
+
mechanical T0/T1 facts, never declared by `ba`. `superseded_by` stays
|
|
618
618
|
`null` until a scope-architect `map-scopes` order retires this contract in favor of its replacements.
|
|
619
619
|
|
|
620
620
|
**PA2 size lint:** a scope whose `allowed_file_substrate` glob set resolves to more than ~15
|
|
@@ -916,7 +916,7 @@
|
|
|
916
916
|
html += '</div>';
|
|
917
917
|
|
|
918
918
|
// ---- 2. hill chart (headline) ----
|
|
919
|
-
html += '<div class="section"><h3>Hill Chart</h3><div class="sub">Where each scope sits between figuring it out and getting it done — mechanical, derived only from T0/T1
|
|
919
|
+
html += '<div class="section"><h3>Hill Chart</h3><div class="sub">Where each scope sits between figuring it out and getting it done — mechanical, derived only from T0/T1 evidence, never self-reported. Position is how derisked a scope has ever been; dot color is its current-round health, so a scope can sit downhill and still show red if its latest attempt regressed.</div>';
|
|
920
920
|
html += '<div class="section-card"><div class="hill-wrap hero">';
|
|
921
921
|
if (pitch.hillAvailable) {
|
|
922
922
|
html += renderHillSVG(pitch.scopeStatus, 'hero');
|
|
@@ -77,8 +77,8 @@ the ship report's census table.
|
|
|
77
77
|
write-whitelist; wrong here =
|
|
78
78
|
a legitimate ESCALATE later
|
|
79
79
|
shared_substrate[] — files ≥2 scopes both touch;
|
|
80
|
-
|
|
81
|
-
|
|
80
|
+
declaring one costs concurrency:
|
|
81
|
+
they build one at a time
|
|
82
82
|
affordance_manifest — from ux-behavior.md state
|
|
83
83
|
tables: every interactive
|
|
84
84
|
element as {test_id, role} +
|
|
@@ -93,7 +93,7 @@ the ship report's census table.
|
|
|
93
93
|
TBD and flag it, never invent
|
|
94
94
|
a fixture for unbuilt behavior
|
|
95
95
|
hill_phase: "UPHILL_UNKNOWN" — ALWAYS; phase is derived from
|
|
96
|
-
T0/T1
|
|
96
|
+
T0/T1 facts later,
|
|
97
97
|
never authored
|
|
98
98
|
4 LINT node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify spec --slug <slug>
|
|
99
99
|
→ PA1 (directory alignment), PA2 (>~15 files), DISJOINT (undeclared overlap),
|
|
@@ -90,7 +90,7 @@ Collect (explicit — never inferred):
|
|
|
90
90
|
attempts a single scope gets inside one round before its attempt loop trips and
|
|
91
91
|
queues a hammer PROPOSAL for GATE H rather than blocking the round. Only meaningful
|
|
92
92
|
when the spec folder has scope contracts; a spec with none skips the attempt loop
|
|
93
|
-
entirely and BUILD behaves exactly as in v0.2.6 (task-executor --next, no T0
|
|
93
|
+
entirely and BUILD behaves exactly as in v0.2.6 (task-executor --next, no T0).
|
|
94
94
|
no_progress_k (v1.5): the STAGNATION term of the same inner breaker. Default 2 — the
|
|
95
95
|
number of consecutive non-`kept` trials after which a scope ends early and queues the
|
|
96
96
|
same GATE H proposal. attempt_budget counts ATTEMPTS and cannot see that the last two
|
|
@@ -205,7 +205,7 @@ ingest-result <result> → board/ledger writes
|
|
|
205
205
|
fails, and the run ABORTS naming the phase. Resolve it
|
|
206
206
|
yourself and record the answer in round-ledger.md, which
|
|
207
207
|
the NEXT attempt's fresh context reads back.
|
|
208
|
-
harness verify t0 → fixtures + DB probe
|
|
208
|
+
harness verify t0 → fixtures + DB probe, then scores the
|
|
209
209
|
attempt against the baseline trial and snapshots or
|
|
210
210
|
restores the tree. Branch on `status` from its stdout
|
|
211
211
|
JSON — the tree action has ALREADY happened:
|
|
@@ -217,9 +217,8 @@ harness verify t0 → fixtures + DB probe + (on green) sees
|
|
|
217
217
|
KEEPS, because a spec-conformance fix cannot raise a score that is already at
|
|
218
218
|
full marks, and reverting it would discard exactly the work a fix round exists
|
|
219
219
|
to do. Tree already restored from the last kept
|
|
220
|
-
snapshot. Subsumes the retired stash-and-retry branch:
|
|
221
|
-
|
|
222
|
-
is why seesaw runs before anything is declared green.
|
|
220
|
+
snapshot. Subsumes the retired stash-and-retry branch: an attempt that scores
|
|
221
|
+
worse than the incumbent reverts through this same rule.
|
|
223
222
|
rather than the code. Tree kept, baseline reset. Not a verdict, not a failure.
|
|
224
223
|
crash a fixture command failed to spawn or timed out; tree restored. Fix the fixture,
|
|
225
224
|
not the code.
|
|
@@ -446,9 +445,8 @@ Do: verify the phase's artifact by hand before trusting either diagnosis. If it
|
|
|
446
445
|
```
|
|
447
446
|
Invoke via Bash directly — NOT an Agent, this is deterministic tooling, not a worker:
|
|
448
447
|
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify t0 shapeup/<slug>/scopes/<scope-id>.md
|
|
449
|
-
--round <N> --attempt <M>
|
|
450
|
-
Effect: runs the scope's e2e fixtures
|
|
451
|
-
over every FINISHED scope's fixtures. Writes the verdict artifact spec-evaluator's
|
|
448
|
+
--round <N> --attempt <M>
|
|
449
|
+
Effect: runs the scope's e2e fixtures and DB probe. Writes the verdict artifact spec-evaluator's
|
|
452
450
|
T0-citation rule will require a citation to, appends one row to t0/trials.jsonl, and — this
|
|
453
451
|
is the ratchet — scores the attempt against the last kept trial and snapshots or
|
|
454
452
|
restores the working tree ITSELF. Zero LLM tokens — deterministic tooling, not a
|
|
@@ -592,7 +590,7 @@ guarantee lives in the script and, where noted, in a hook.
|
|
|
592
590
|
| Three-level circuit breaker: attempt_budget (inner, per scope) nests inside round_budget (outer), with an opt-in wall_clock_budget_s deadline | An exhausted scope queues a GATE H hammer proposal, it never blocks the round; only round_budget hitting 0 stops the whole run; the deadline breaker (checked every round boundary in `shapeup-run.js`) routes to GATE H so a run out of clock still ships what is green instead of being killed from outside |
|
|
593
591
|
| The tech lead never hand-edits a scope contract | scope-architect is its sole writer (single-writer-per-file) |
|
|
594
592
|
| Substrate-disjointness + PA1/PA2 lints are re-asserted at GATE L1b (harness verify spec) even when scope-architect already checked them | A human may have hand-approved past a 🔴 at the architect's checkpoint; `shapeup-run.js` runs spec-lint itself, in code, before resolving L1b |
|
|
595
|
-
| Hill phase is read from mechanical facts (T0/T1
|
|
593
|
+
| Hill phase is read from mechanical facts (T0/T1), never declared by a worker | Closes the self-reported-confidence risk outright — facts move dots, not authors |
|
|
596
594
|
| GATE H is delegated to scope-hammer, never adjudicated inline by the tech lead | Keeps the orchestrator thin; census/baseline-comparison/cut-list logic has one owner |
|
|
597
595
|
|
|
598
596
|
---
|