shapeup-sdlc 3.7.12 → 3.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +2 -2
- package/README.md +8 -9
- package/kernel/lib/argv.mjs +4 -4
- package/kernel/lib/paths.mjs +0 -2
- package/kernel/reduce/hill.mjs +17 -39
- package/kernel/report/export.mjs +0 -3
- package/kernel/schemas/domain.schema.json +8 -92
- package/kernel/verify/env.mjs +2 -1
- package/kernel/verify/ratchet-tree.mjs +1 -1
- package/kernel/verify/t0.mjs +35 -87
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/references/task-generation.md +1 -1
- package/skills/hill-chart/assets/dashboard.template.html +1 -1
- package/skills/scope-architect/SKILL.md +3 -3
- package/skills/tech-lead/references/gates.md +1 -1
- package/skills/tech-lead/references/protocol.md +6 -8
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.8.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -35,7 +35,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
35
35
|
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
|
-
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe
|
|
38
|
+
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
39
39
|
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
|
|
40
40
|
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
|
|
41
41
|
|
|
@@ -63,7 +63,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
|
|
|
63
63
|
- **QA is a level-up, not a gate** — `--no-qa` skips it; circuit breaker outranks the Hunter.
|
|
64
64
|
- **Role separation** — Evaluator grades, task-executor fixes, QA discovers.
|
|
65
65
|
- **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
|
|
66
|
-
- **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts
|
|
66
|
+
- **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts, never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
|
|
67
67
|
- **Envelope port (v1.0)** — every dispatch is WorkOrder in / WorkResult out; shared state has exactly one writer (the ingest step); malformed envelopes are hook-denied. Workers: stateless, craft-only, pipeline-blind.
|
|
68
68
|
|
|
69
69
|
## Setup & Execution
|
package/README.md
CHANGED
|
@@ -45,11 +45,11 @@ its fixtures and its DB probe and writes an artifact to disk — with each comma
|
|
|
45
45
|
its captured output, and whether it ran at all. The evaluator must cite that artifact, and the
|
|
46
46
|
hill phase is derived from artifacts rather than from a worker's own account of its progress.
|
|
47
47
|
Two limits, stated here because the point of this section is that a claim without a mechanism
|
|
48
|
-
behind it is the thing this harness exists to prevent:
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
48
|
+
behind it is the thing this harness exists to prevent: a citation is re-hashed from disk, checked
|
|
49
|
+
against the scope, round and run the artifact records, and refused unless it names a file this
|
|
50
|
+
run's own verifier wrote — what that does NOT prove is that a judge with a dishonest hand could
|
|
51
|
+
not have arranged the artifact first, which is why the judge's own substrate freezes the verdicts
|
|
52
|
+
directory. A T0 artifact is
|
|
53
53
|
also evidence about the machine that produced it, and now says so: each verdict carries where it
|
|
54
54
|
ran — the absolute path, the git tree, the resolved toolchain, lockfile digests, declared cache
|
|
55
55
|
directories and a digest over an allowlist of environment values — so a disagreeing re-run can be
|
|
@@ -132,8 +132,7 @@ rest of this README after this table and nothing will be a surprise.
|
|
|
132
132
|
|---|---|
|
|
133
133
|
| **board** | The round's task list. "Green" means every task is done. GATE L2's hook reads this before an evaluation and warns if it is not green. |
|
|
134
134
|
| **round** | One build → evaluate cycle. A FAIL verdict starts round *r+1*. |
|
|
135
|
-
| **T0** | The smoke test a scope must pass before it counts as built: its fixtures
|
|
136
|
-
| **seesaw** | The regression arm of T0, meant to re-run *other* scopes' fixtures so a regression is never mistaken for progress. Declared, not yet wired: no run writes its registry, so no run has executed it. |
|
|
135
|
+
| **T0** | The smoke test a scope must pass before it counts as built: its fixtures and a DB probe. Writes an artifact to disk that the evaluator must cite. |
|
|
137
136
|
| **substrate** | The exact list of files one dispatch is allowed to write, stamped into its work order. A hook blocks anything outside it — and anything the order marks frozen. |
|
|
138
137
|
| **scope contract** | The file defining one vertical slice: its substrate, its fixtures, its affordances. |
|
|
139
138
|
| **affordance** | The thing a user can actually click, type or call. UI is graded on affordances, not on looks. |
|
|
@@ -309,8 +308,8 @@ These hold across the harness and are the reason it stays predictable:
|
|
|
309
308
|
of blocking the round. An opt-in third breaker bounds the **wall clock**, because the other two
|
|
310
309
|
count events and neither can notice a single round running for half an hour — tripping it routes
|
|
311
310
|
to GATE H, so a run out of time ships what is green instead of being killed and shipping nothing.
|
|
312
|
-
- **Hill phase is mechanical, never self-reported** — derived only from T0/T1 facts
|
|
313
|
-
|
|
311
|
+
- **Hill phase is mechanical, never self-reported** — derived only from T0/T1 facts, closing the
|
|
312
|
+
self-reported-confidence risk. A scope with no
|
|
314
313
|
discovery ledger derives no phase rather than a solved one, so absence no longer reads as
|
|
315
314
|
progress on that arm.
|
|
316
315
|
- **One writer per shared file** — every board/ledger/verdict write goes through
|
package/kernel/lib/argv.mjs
CHANGED
|
@@ -8,11 +8,11 @@
|
|
|
8
8
|
// const SPEC = {
|
|
9
9
|
// _: { arity: 1, name: "scope-contract.json" },
|
|
10
10
|
// round: { type: "int", min: 1, required: true },
|
|
11
|
-
// "no-
|
|
11
|
+
// "no-ratchet": { type: "flag" },
|
|
12
12
|
// };
|
|
13
|
-
// const args = runArgs(SPEC, argv); // args.round, args.
|
|
13
|
+
// const args = runArgs(SPEC, argv); // args.round, args.noRatchet, args._
|
|
14
14
|
//
|
|
15
|
-
// Flag names reach the caller camelCased (`--no-
|
|
15
|
+
// Flag names reach the caller camelCased (`--no-ratchet` → `noRatchet`). Unknown flags are rejected
|
|
16
16
|
// rather than swallowed as positionals: a typo'd `--rounds 2` landing in `_` is the same defect
|
|
17
17
|
// wearing a different hat. Untyped coercion is the failure this guards — `Number(undefined)` is
|
|
18
18
|
// `NaN`, `??` does not catch `NaN`, and a verdict written to `r NaN-a1.json` with exit 0 is
|
|
@@ -35,7 +35,7 @@ export class ArgvError extends Error {
|
|
|
35
35
|
}
|
|
36
36
|
}
|
|
37
37
|
|
|
38
|
-
/** `--
|
|
38
|
+
/** `--attempt-budget` → `attemptBudget`. */
|
|
39
39
|
function camel(name) {
|
|
40
40
|
return name.replace(/-([a-z0-9])/g, (_, c) => c.toUpperCase());
|
|
41
41
|
}
|
package/kernel/lib/paths.mjs
CHANGED
|
@@ -216,8 +216,6 @@ export const trials = (cwd, slug) => join(t0Dir(cwd, slug), "trials.jsonl");
|
|
|
216
216
|
* {@link decisions}, one small file with one writer.
|
|
217
217
|
*/
|
|
218
218
|
export const gates = (cwd, slug) => join(localRoot(cwd, slug), "gates.jsonl");
|
|
219
|
-
/** Finished-scope fixture registry for the seesaw regression check. */
|
|
220
|
-
export const seesawRegistry = (cwd, slug) => join(localRoot(cwd, slug), "seesaw", "registry.json");
|
|
221
219
|
/**
|
|
222
220
|
* The round build gate's verdicts — one immutable artifact per gate run, `r<N>-t<T>.json`.
|
|
223
221
|
*
|
package/kernel/reduce/hill.mjs
CHANGED
|
@@ -144,8 +144,8 @@ function committedPhase(hDir, id) {
|
|
|
144
144
|
* - UPHILL_UNKNOWN: open unknowns > 0 in the ledger for this scope — and the floor the scope sits
|
|
145
145
|
* at whenever the ledger has not answered at all, which is where every run legitimately begins
|
|
146
146
|
* - UPHILL_SOLVED: the ledger was read and reports zero open unknowns, no T0-green yet
|
|
147
|
-
* - DOWNHILL_EXECUTION: ≥1 T0-green in a round whose build gate is not red; T1
|
|
148
|
-
* - FINISHED: T1 PASS ∧
|
|
147
|
+
* - DOWNHILL_EXECUTION: ≥1 T0-green in a round whose build gate is not red; T1 pending
|
|
148
|
+
* - FINISHED: T1 PASS ∧ a T0-green that a red round did not invalidate
|
|
149
149
|
*
|
|
150
150
|
* @param {string} cwd - The project root directory.
|
|
151
151
|
* @param {string} slug - The feature slug being built.
|
|
@@ -230,18 +230,16 @@ export function deriveHill(cwd, slug) {
|
|
|
230
230
|
}
|
|
231
231
|
}
|
|
232
232
|
|
|
233
|
-
// 2. T0 facts per scope: has it achieved a green overall verdict
|
|
233
|
+
// 2. T0 facts per scope: has it achieved a green overall verdict this run?
|
|
234
234
|
//
|
|
235
|
-
//
|
|
236
|
-
//
|
|
237
|
-
//
|
|
238
|
-
//
|
|
239
|
-
//
|
|
240
|
-
//
|
|
241
|
-
// verdict counting exactly as before.
|
|
235
|
+
// FINISHED USED TO WAIT ON AN ARM THAT NEVER RAN. The phase required `seesaw.ran && seesaw.pass`,
|
|
236
|
+
// nothing in the codebase ever wrote the registry that arm read, and so no scope in any recorded
|
|
237
|
+
// run reached FINISHED — 38 committed shards on the live consumer, not one of them. The arm was
|
|
238
|
+
// removed in 3.8.0 by decision; the precondition goes with it, and the top phase is reachable
|
|
239
|
+
// again on the evidence that does exist: T1 passed, and a T0 green from a round the build gate
|
|
240
|
+
// did not red.
|
|
242
241
|
const redRounds = redBuildRounds(cwd, slug);
|
|
243
242
|
const t0Facts = {};
|
|
244
|
-
// This run's verdicts only — a prior run's green over the same slug moved this run's dot.
|
|
245
243
|
const hillRunId = readRunId(cwd, slug);
|
|
246
244
|
if (existsSync(vDir)) {
|
|
247
245
|
for (const f of readdirSync(vDir)) {
|
|
@@ -249,41 +247,21 @@ export function deriveHill(cwd, slug) {
|
|
|
249
247
|
try {
|
|
250
248
|
const b = JSON.parse(readFileSync(join(vDir, f), "utf8"));
|
|
251
249
|
if (hillRunId && b.run_id && b.run_id !== hillRunId) continue;
|
|
252
|
-
if (!t0Facts[b.scope_id]) t0Facts[b.scope_id] = { hasGreen: false
|
|
253
|
-
if (b.overall === "green" && !redRounds.has(Number(b.round)))
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
// inferring it from a false `regression` flag. That inference was vacuously true on every
|
|
258
|
-
// green T0 whether or not a seesaw check ever ran: nothing in this codebase currently
|
|
259
|
-
// passes `--seesaw-registry` to `verify t0`, so `seesaw.ran` is always `false` today and
|
|
260
|
-
// `regression` is always `false` too — "not asked" was being read as "clean," letting a
|
|
261
|
-
// scope reach FINISHED on a regression check that had never executed.
|
|
262
|
-
//
|
|
263
|
-
// Betting Table decision (Phase 3.5 / S4): wiring the seesaw registry for real is a
|
|
264
|
-
// genuine feature with a real running cost (re-running every finished scope's fixtures
|
|
265
|
-
// on every later attempt) and is out of proportion to a certification-gap fix. Deferred,
|
|
266
|
-
// not silently dropped — a scope with no registry wired simply cannot reach FINISHED via
|
|
267
|
-
// this path today, which is the honest state of the system: this check was never really
|
|
268
|
-
// gating FINISHED before either.
|
|
269
|
-
if (b.seesaw?.ran && b.seesaw?.pass) {
|
|
270
|
-
t0Facts[b.scope_id].seesawGreen = true;
|
|
271
|
-
}
|
|
272
|
-
}
|
|
273
|
-
} catch (e) {
|
|
274
|
-
// ignore parse errors
|
|
250
|
+
if (!t0Facts[b.scope_id]) t0Facts[b.scope_id] = { hasGreen: false };
|
|
251
|
+
if (b.overall === "green" && !redRounds.has(Number(b.round))) t0Facts[b.scope_id].hasGreen = true;
|
|
252
|
+
} catch {
|
|
253
|
+
// A torn or unreadable verdict proves nothing about the scope — skip it rather than let it
|
|
254
|
+
// decide a phase.
|
|
275
255
|
}
|
|
276
256
|
}
|
|
277
257
|
}
|
|
278
|
-
|
|
279
|
-
// 3. Ledger unknowns per scope — `null` for every scope when the ledger itself was not readable
|
|
280
|
-
// or nothing in it named a scope (see `ledgerUnknowns`). Only a real count can promote.
|
|
258
|
+
|
|
281
259
|
const scopeUnknowns = ledgerUnknowns(ledgerPath, scopes);
|
|
282
260
|
|
|
283
261
|
const report = [];
|
|
284
262
|
for (const s of scopes) {
|
|
285
263
|
const id = s.scope_id;
|
|
286
|
-
const t0 = t0Facts[id] || { hasGreen: false
|
|
264
|
+
const t0 = t0Facts[id] || { hasGreen: false };
|
|
287
265
|
// `null` = the ledger did not answer; a number = it did. `|| 0` collapsed the two.
|
|
288
266
|
const unknowns = scopeUnknowns === null ? null : (scopeUnknowns[id] || 0);
|
|
289
267
|
|
|
@@ -291,7 +269,7 @@ export function deriveHill(cwd, slug) {
|
|
|
291
269
|
// it — including before Orient has filed anything, which is where every run legitimately
|
|
292
270
|
// starts. Only an ANSWERED count of zero promotes to UPHILL_SOLVED; `null` never does.
|
|
293
271
|
let phase = "UPHILL_UNKNOWN";
|
|
294
|
-
if (t1Pass && t0.hasGreen
|
|
272
|
+
if (t1Pass && t0.hasGreen) {
|
|
295
273
|
phase = "FINISHED";
|
|
296
274
|
} else if (t0.hasGreen) {
|
|
297
275
|
phase = "DOWNHILL_EXECUTION";
|
package/kernel/report/export.mjs
CHANGED
|
@@ -129,7 +129,6 @@ function t0Row(a, runId) {
|
|
|
129
129
|
regression: a?.regression ?? null,
|
|
130
130
|
fixtures_green: a?.fixtures_green ?? null,
|
|
131
131
|
db_probe_green: a?.db_probe_green ?? null,
|
|
132
|
-
seesaw_green: a?.seesaw_green ?? null,
|
|
133
132
|
// One field a reader compares, and the block itself stays in the artifact for a human to diff:
|
|
134
133
|
// two rows with the same tree and different env digests are two machines, not a regression.
|
|
135
134
|
env_sha256: a?.env?.env_sha256 ?? null,
|
|
@@ -137,8 +136,6 @@ function t0Row(a, runId) {
|
|
|
137
136
|
tree_dirty: a?.env?.tree?.dirty ?? null,
|
|
138
137
|
fixtures_total: fixtures.length,
|
|
139
138
|
fixtures_passed: fixtures.filter((f) => f?.pass === true).length,
|
|
140
|
-
seesaw_ran: a?.seesaw?.ran ?? null,
|
|
141
|
-
seesaw_failing: Array.isArray(a?.seesaw?.failing) ? a.seesaw.failing.length : null,
|
|
142
139
|
discovered_tasks: Array.isArray(a?.discovered_tasks) ? a.discovered_tasks.length : 0,
|
|
143
140
|
};
|
|
144
141
|
}
|
|
@@ -236,7 +236,7 @@
|
|
|
236
236
|
"to": "HillShard",
|
|
237
237
|
"cardinality": "1:0..1",
|
|
238
238
|
"via": "scope_id",
|
|
239
|
-
"note": "phase derived from T0/T1
|
|
239
|
+
"note": "phase derived from T0/T1 facts, never authored"
|
|
240
240
|
},
|
|
241
241
|
{
|
|
242
242
|
"from": "ScopeContract",
|
|
@@ -252,13 +252,6 @@
|
|
|
252
252
|
"via": "fixtures[] + db_probe",
|
|
253
253
|
"note": "produced by actually running the commands — no agent can fabricate them"
|
|
254
254
|
},
|
|
255
|
-
{
|
|
256
|
-
"from": "T0Artifact",
|
|
257
|
-
"to": "SeesawCheck",
|
|
258
|
-
"cardinality": "1:1",
|
|
259
|
-
"via": "seesaw",
|
|
260
|
-
"note": "the regression half of T0 — every FINISHED scope's fixtures re-run"
|
|
261
|
-
},
|
|
262
255
|
{
|
|
263
256
|
"from": "T0Artifact",
|
|
264
257
|
"to": "AegisTriple",
|
|
@@ -287,13 +280,6 @@
|
|
|
287
280
|
"via": "payload.trial_history[]",
|
|
288
281
|
"note": "inspect(): the loop reads its own recent history, across the round boundary"
|
|
289
282
|
},
|
|
290
|
-
{
|
|
291
|
-
"from": "SeesawRegistry",
|
|
292
|
-
"to": "ScopeContract",
|
|
293
|
-
"cardinality": "1:N",
|
|
294
|
-
"via": "scopes[].scope_id",
|
|
295
|
-
"note": "every FINISHED scope's fixtures re-run on each later attempt"
|
|
296
|
-
},
|
|
297
283
|
{
|
|
298
284
|
"from": "UseCase",
|
|
299
285
|
"to": "Seam",
|
|
@@ -528,7 +514,7 @@
|
|
|
528
514
|
"items": {
|
|
529
515
|
"type": "string"
|
|
530
516
|
},
|
|
531
|
-
"description": "Globs ≥2 scopes both touch —
|
|
517
|
+
"description": "Globs ≥2 scopes both touch — an edit here is read-modify-write, so two scopes that both declare one never build at the same time."
|
|
532
518
|
},
|
|
533
519
|
"append_only": {
|
|
534
520
|
"type": "array",
|
|
@@ -844,7 +830,7 @@
|
|
|
844
830
|
"items": {
|
|
845
831
|
"type": "string"
|
|
846
832
|
},
|
|
847
|
-
"description": "Files ≥2 scopes both touch — must be declared in BOTH scopes' lists (spec-lint DISJOINT)
|
|
833
|
+
"description": "Files ≥2 scopes both touch — must be declared in BOTH scopes' lists (spec-lint DISJOINT); declaring one costs concurrency, since two scopes that share a path build one at a time."
|
|
848
834
|
},
|
|
849
835
|
"affordance_manifest": {
|
|
850
836
|
"type": "array",
|
|
@@ -872,7 +858,7 @@
|
|
|
872
858
|
"DOWNHILL_EXECUTION",
|
|
873
859
|
"FINISHED"
|
|
874
860
|
],
|
|
875
|
-
"description": "ALWAYS authored as UPHILL_UNKNOWN — self-reported confidence is the risk this closes. The live phase is DERIVED from T0/T1
|
|
861
|
+
"description": "ALWAYS authored as UPHILL_UNKNOWN — self-reported confidence is the risk this closes. The live phase is DERIVED from T0/T1 facts into hill/<scope-id>.yml — facts move dots, not authors."
|
|
876
862
|
},
|
|
877
863
|
"superseded_by": {
|
|
878
864
|
"type": "array",
|
|
@@ -1254,35 +1240,8 @@
|
|
|
1254
1240
|
}
|
|
1255
1241
|
}
|
|
1256
1242
|
},
|
|
1257
|
-
|
|
1258
|
-
"description": "The
|
|
1259
|
-
"x-tier": "EMBEDDED",
|
|
1260
|
-
"type": "object",
|
|
1261
|
-
"properties": {
|
|
1262
|
-
"ran": {
|
|
1263
|
-
"type": "boolean",
|
|
1264
|
-
"description": "false when no registry was passed or the attempt was already red (don't seesaw a red attempt)."
|
|
1265
|
-
},
|
|
1266
|
-
"pass": {
|
|
1267
|
-
"type": "boolean"
|
|
1268
|
-
},
|
|
1269
|
-
"scopes_checked": {
|
|
1270
|
-
"type": "array",
|
|
1271
|
-
"items": {
|
|
1272
|
-
"type": "string"
|
|
1273
|
-
}
|
|
1274
|
-
},
|
|
1275
|
-
"failing": {
|
|
1276
|
-
"type": "array",
|
|
1277
|
-
"items": {
|
|
1278
|
-
"type": "string"
|
|
1279
|
-
},
|
|
1280
|
-
"description": "scope_ids whose fixtures broke."
|
|
1281
|
-
}
|
|
1282
|
-
}
|
|
1283
|
-
},
|
|
1284
|
-
"T0Artifact": {
|
|
1285
|
-
"description": "The mechanical verification verdict for one build attempt — the evidence layer under the LLM judge. Written by harness verify t0 from actually running the scope's fixtures + DB probe + seesaw; zero LLM tokens. spec-evaluator must cite it (sha256) on scoped specs. Red artifacts carry AEGIS triples that become the next attempt's digested_errors.",
|
|
1243
|
+
"T0Artifact": {
|
|
1244
|
+
"description": "The mechanical verification verdict for one build attempt — the evidence layer under the LLM judge. Written by harness verify t0 from actually running the scope's fixtures and DB probe; zero LLM tokens. spec-evaluator must cite it (sha256) on scoped specs. Red artifacts carry AEGIS triples that become the next attempt's digested_errors.",
|
|
1286
1245
|
"x-tier": "LOCAL",
|
|
1287
1246
|
"x-location": ".shapeup/<slug>/t0/verdicts/r<N>-a<M>-t<T>.json (schema_version 2; the unsuffixed r<N>-a<M>.json of schema_version 1 is still readable)",
|
|
1288
1247
|
"x-writer": "harness verify t0",
|
|
@@ -1327,9 +1286,6 @@
|
|
|
1327
1286
|
"db_probe": {
|
|
1328
1287
|
"$ref": "#/$defs/CommandResult"
|
|
1329
1288
|
},
|
|
1330
|
-
"seesaw": {
|
|
1331
|
-
"$ref": "#/$defs/SeesawCheck"
|
|
1332
|
-
},
|
|
1333
1289
|
"fixtures_green": {
|
|
1334
1290
|
"type": "boolean"
|
|
1335
1291
|
},
|
|
@@ -1337,9 +1293,6 @@
|
|
|
1337
1293
|
"type": "boolean",
|
|
1338
1294
|
"description": "true when no probe declared (null probe never counts as failure)."
|
|
1339
1295
|
},
|
|
1340
|
-
"seesaw_green": {
|
|
1341
|
-
"type": "boolean"
|
|
1342
|
-
},
|
|
1343
1296
|
"overall": {
|
|
1344
1297
|
"type": "string",
|
|
1345
1298
|
"enum": [
|
|
@@ -1347,10 +1300,6 @@
|
|
|
1347
1300
|
"red"
|
|
1348
1301
|
]
|
|
1349
1302
|
},
|
|
1350
|
-
"regression": {
|
|
1351
|
-
"type": "boolean",
|
|
1352
|
-
"description": "fixtures+db green but seesaw red — the rollback+retry case."
|
|
1353
|
-
},
|
|
1354
1303
|
"score": {
|
|
1355
1304
|
"$ref": "#/$defs/T0Score",
|
|
1356
1305
|
"description": "schema_version 2+: the comparable outcome vector better() ranks. A reduce over fixtures[] — no new measurement."
|
|
@@ -1384,11 +1333,6 @@
|
|
|
1384
1333
|
"x-tier": "EMBEDDED",
|
|
1385
1334
|
"type": "object",
|
|
1386
1335
|
"properties": {
|
|
1387
|
-
"regressions": {
|
|
1388
|
-
"type": "integer",
|
|
1389
|
-
"minimum": 0,
|
|
1390
|
-
"description": "Previously-FINISHED scopes now failing (seesaw). Dominates: breaking a shipped scope is never an improvement."
|
|
1391
|
-
},
|
|
1392
1336
|
"fixtures_passed": {
|
|
1393
1337
|
"type": "integer",
|
|
1394
1338
|
"minimum": 0
|
|
@@ -1412,7 +1356,6 @@
|
|
|
1412
1356
|
}
|
|
1413
1357
|
},
|
|
1414
1358
|
"required": [
|
|
1415
|
-
"regressions",
|
|
1416
1359
|
"fixtures_passed",
|
|
1417
1360
|
"fixtures_total"
|
|
1418
1361
|
]
|
|
@@ -1619,34 +1562,7 @@
|
|
|
1619
1562
|
}
|
|
1620
1563
|
}
|
|
1621
1564
|
},
|
|
1622
|
-
|
|
1623
|
-
"description": "The fixture registry of every FINISHED scope — what seesawCheck re-runs on each later attempt so a new scope cannot silently break a shipped one — a regression mistaken for progress is the pathology the seesaw exists for.",
|
|
1624
|
-
"x-tier": "LOCAL",
|
|
1625
|
-
"x-location": ".shapeup/<slug>/seesaw/registry.json",
|
|
1626
|
-
"x-writer": "tech-lead (when a scope reaches FINISHED)",
|
|
1627
|
-
"x-readers": "harness verify t0",
|
|
1628
|
-
"type": "object",
|
|
1629
|
-
"properties": {
|
|
1630
|
-
"scopes": {
|
|
1631
|
-
"type": "array",
|
|
1632
|
-
"items": {
|
|
1633
|
-
"type": "object",
|
|
1634
|
-
"properties": {
|
|
1635
|
-
"scope_id": {
|
|
1636
|
-
"type": "string"
|
|
1637
|
-
},
|
|
1638
|
-
"fixtures": {
|
|
1639
|
-
"type": "array",
|
|
1640
|
-
"items": {
|
|
1641
|
-
"type": "string"
|
|
1642
|
-
}
|
|
1643
|
-
}
|
|
1644
|
-
}
|
|
1645
|
-
}
|
|
1646
|
-
}
|
|
1647
|
-
}
|
|
1648
|
-
},
|
|
1649
|
-
"VerdictLedgerLine": {
|
|
1565
|
+
"VerdictLedgerLine": {
|
|
1650
1566
|
"description": "One appended line of judge history (JSONL). Never rewritten — flips across runs are detected here and force confidence low. run auto-increments per append batch.",
|
|
1651
1567
|
"x-tier": "LOCAL",
|
|
1652
1568
|
"x-location": ".shapeup/<slug>/evaluation/.verdicts-<target>.jsonl",
|
|
@@ -1690,7 +1606,7 @@
|
|
|
1690
1606
|
}
|
|
1691
1607
|
},
|
|
1692
1608
|
"HillShard": {
|
|
1693
|
-
"description": "One scope's hill position — DERIVED, never self-reported: UPHILL_UNKNOWN (open unknowns > 0) → UPHILL_SOLVED (unknowns 0, no T0-green yet) → DOWNHILL_EXECUTION (≥1 T0-green; T1
|
|
1609
|
+
"description": "One scope's hill position — DERIVED, never self-reported: UPHILL_UNKNOWN (open unknowns > 0) → UPHILL_SOLVED (unknowns 0, no T0-green yet) → DOWNHILL_EXECUTION (≥1 T0-green from a round that built; T1 pending) → FINISHED (T1 PASS ∧ that T0-green). Single-writer = whoever holds that scope's branch. Progress is reported by hill position, never task counts.",
|
|
1694
1610
|
"x-tier": "SHARED",
|
|
1695
1611
|
"x-location": "shapeup/<slug>/hill/<scope-id>.yml",
|
|
1696
1612
|
"x-writer": "tech-lead (GATE L2 derivation)",
|
package/kernel/verify/env.mjs
CHANGED
|
@@ -106,7 +106,8 @@ function treeState(cwd) {
|
|
|
106
106
|
* path, which is exactly the mechanism that made one tree build three ways.
|
|
107
107
|
*
|
|
108
108
|
* `null` means the profile declared nothing — NOT that there are none. "Not asked" and "none" are
|
|
109
|
-
* different facts, and collapsing them is the mistake the
|
|
109
|
+
* different facts, and collapsing them is the mistake that kept the hill's top phase shut for the
|
|
110
|
+
* life of the seesaw arm: "not asked" was recorded the same way as "nothing wrong".
|
|
110
111
|
*
|
|
111
112
|
* @param {(string|null)} profilePath - `shapeup/<slug>/project-profile.md`, when the caller knows it.
|
|
112
113
|
* @returns {(object[]|null)} One entry per declared cache, or null when none is declared.
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
//
|
|
5
5
|
// The attempt loop branched a red T0 two ways, and only one of them reverted anything:
|
|
6
6
|
//
|
|
7
|
-
// • a
|
|
7
|
+
// • a trial that scored worse than the incumbent → `git stash push -u`;
|
|
8
8
|
// • a red on the scope's OWN fixtures → "loop to the next attempt", and no revert at all.
|
|
9
9
|
//
|
|
10
10
|
// So the failing tree stayed on the branch, and attempt N+1's fresh, zero-memory subagent began
|
package/kernel/verify/t0.mjs
CHANGED
|
@@ -1,8 +1,7 @@
|
|
|
1
1
|
#!/usr/bin/env node
|
|
2
2
|
// T0 mechanical verification layer.
|
|
3
3
|
//
|
|
4
|
-
// Runs a scope's e2e fixtures
|
|
5
|
-
// regression check (re-runs every FINISHED scope's fixtures from the registry). Writes one
|
|
4
|
+
// Runs a scope's e2e fixtures and its DB probe (zero LLM tokens). Writes one
|
|
6
5
|
// verdict artifact per attempt that spec-evaluator (T1) must cite; a verdict without it is
|
|
7
6
|
// structurally invalid. No agent can fabricate this file's contents
|
|
8
7
|
// because it is produced by actually running the commands.
|
|
@@ -30,7 +29,7 @@
|
|
|
30
29
|
//
|
|
31
30
|
// Usage:
|
|
32
31
|
// node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify t0 <scope-contract.json> \
|
|
33
|
-
// --round N --attempt M [--cwd <dir>] [--out <dir>]
|
|
32
|
+
// --round N --attempt M [--cwd <dir>] [--out <dir>]
|
|
34
33
|
// [--no-ratchet]
|
|
35
34
|
//
|
|
36
35
|
// Exit code: 0 = overall green, 1 = overall red (mirrors the oracle convention), 2 = bad argv.
|
|
@@ -185,80 +184,45 @@ export function runDbProbe(dbProbeCmd, cwd) {
|
|
|
185
184
|
}
|
|
186
185
|
|
|
187
186
|
/**
|
|
188
|
-
*
|
|
189
|
-
*
|
|
190
|
-
*
|
|
191
|
-
*
|
|
192
|
-
*
|
|
193
|
-
*
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
}
|
|
203
|
-
let registry;
|
|
204
|
-
try {
|
|
205
|
-
registry = JSON.parse(readFileSync(registryPath, "utf8"));
|
|
206
|
-
} catch {
|
|
207
|
-
return { ran: false, pass: null, scopes_checked: [], failing: [], error: "registry unparsable" };
|
|
208
|
-
}
|
|
209
|
-
const scopes = registry.scopes || [];
|
|
210
|
-
const failing = [];
|
|
211
|
-
for (const s of scopes) {
|
|
212
|
-
const { pass } = runFixtures(s.fixtures, cwd);
|
|
213
|
-
if (!pass) failing.push(s.scope_id);
|
|
214
|
-
}
|
|
215
|
-
return { ran: true, pass: failing.length === 0, scopes_checked: scopes.map((s) => s.scope_id), failing };
|
|
216
|
-
}
|
|
217
|
-
|
|
218
|
-
/**
|
|
219
|
-
* Combine fixtures + DB probe + seesaw into the overall T0 verdict.
|
|
220
|
-
* @param {{fixtures:{pass:boolean}, dbProbe:({pass:boolean}|null),
|
|
221
|
-
* seesaw:{ran:boolean,pass:boolean}}} parts - The three sub-results.
|
|
222
|
-
* @returns {{fixtures_green:boolean, db_probe_green:boolean, seesaw_green:boolean,
|
|
223
|
-
* overall:("green"|"red"), regression:boolean}} Per-arm greens, the overall verdict (green iff
|
|
224
|
-
* all three), and `regression` = fixtures+db green but seesaw red (the rollback-and-retry case).
|
|
187
|
+
* Combine the fixtures and the DB probe into the overall T0 verdict.
|
|
188
|
+
*
|
|
189
|
+
* THE SEESAW ARM IS GONE (3.8.0), by a Betting Table decision rather than by neglect. It was
|
|
190
|
+
* declared in the schema, the docs and this function, and nothing ever wrote the registry it read,
|
|
191
|
+
* so it never ran once in any recorded run — while its absence held the hill's top phase shut:
|
|
192
|
+
* FINISHED required `seesaw.ran && seesaw.pass`, and the 38 committed hill shards across the live
|
|
193
|
+
* consumer's features contain no FINISHED at all. A cross-scope regression is still caught by the
|
|
194
|
+
* round build gate, which builds and launches the whole feature once per round; what the arm would
|
|
195
|
+
* have added is attribution and an earlier signal, at the price of re-running every finished
|
|
196
|
+
* scope's fixtures on every attempt — minutes per attempt on an eighteen-scope feature.
|
|
197
|
+
*
|
|
198
|
+
* @param {{fixtures:{pass:boolean}, dbProbe:({pass:boolean}|null)}} parts - The two sub-results.
|
|
199
|
+
* @returns {{fixtures_green:boolean, db_probe_green:boolean, overall:("green"|"red")}} Per-arm
|
|
200
|
+
* greens and the overall verdict, green iff both.
|
|
225
201
|
*/
|
|
226
|
-
export function computeVerdict({ fixtures, dbProbe
|
|
202
|
+
export function computeVerdict({ fixtures, dbProbe }) {
|
|
227
203
|
const fixturesGreen = fixtures.pass;
|
|
228
204
|
const dbGreen = dbProbe === null || dbProbe.pass;
|
|
229
|
-
// A seesaw that did not run does not hold the verdict red — the arm is declared and unwired, and
|
|
230
|
-
// blocking every build on it would be a different defect. It does not make it green either: the
|
|
231
|
-
// hill requires `ran && pass` before a scope may reach FINISHED, and `seesaw_green` here means
|
|
232
|
-
// "nothing this check found is wrong", which is true of a check that found nothing because it
|
|
233
|
-
// never looked.
|
|
234
|
-
const seesawGreen = seesaw.ran ? seesaw.pass === true : true;
|
|
235
205
|
return {
|
|
236
206
|
fixtures_green: fixturesGreen,
|
|
237
207
|
db_probe_green: dbGreen,
|
|
238
|
-
|
|
239
|
-
overall: fixturesGreen && dbGreen && seesawGreen ? "green" : "red",
|
|
240
|
-
// A regression is specifically fixtures/db green but seesaw red — the case that should
|
|
241
|
-
// trigger rollback+retry (spec §3.5) rather than "go fix the new scope's own bug".
|
|
242
|
-
regression: fixturesGreen && dbGreen && !seesawGreen,
|
|
208
|
+
overall: fixturesGreen && dbGreen ? "green" : "red",
|
|
243
209
|
};
|
|
244
210
|
}
|
|
245
211
|
|
|
246
212
|
/**
|
|
247
|
-
* The comparable T0 outcome — a VECTOR, not a float, because the
|
|
213
|
+
* The comparable T0 outcome — a VECTOR, not a float, because the arms are not fungible.
|
|
248
214
|
*
|
|
249
215
|
* Every number here is a reduce over data `writeArtifact` already persists (`fixtures:
|
|
250
216
|
* [{cmd, exit, pass}]`). Nothing new is measured; a number that has always been on disk is
|
|
251
217
|
* finally counted.
|
|
252
218
|
*
|
|
253
|
-
* @param {{fixtures:{results:Array<{pass:boolean}>}, dbProbe:({pass:boolean}|null)
|
|
254
|
-
*
|
|
255
|
-
* @returns {{
|
|
256
|
-
*
|
|
257
|
-
* is never a failure — only an absence.
|
|
219
|
+
* @param {{fixtures:{results:Array<{pass:boolean}>}, dbProbe:({pass:boolean}|null)}} parts - The
|
|
220
|
+
* two T0 sub-results.
|
|
221
|
+
* @returns {{fixtures_passed:number, fixtures_total:number, db_probe:(0|1|null)}} The score vector.
|
|
222
|
+
* `db_probe` is null when no probe is declared, which is never a failure — only an absence.
|
|
258
223
|
*/
|
|
259
|
-
export function score({ fixtures, dbProbe
|
|
224
|
+
export function score({ fixtures, dbProbe }) {
|
|
260
225
|
return {
|
|
261
|
-
regressions: seesaw?.ran ? (seesaw.failing || []).length : 0,
|
|
262
226
|
fixtures_passed: fixtures.results.filter((r) => r.pass).length,
|
|
263
227
|
fixtures_total: fixtures.results.length,
|
|
264
228
|
db_probe: dbProbe === null || dbProbe === undefined ? null : (dbProbe.pass ? 1 : 0),
|
|
@@ -271,23 +235,19 @@ export function score({ fixtures, dbProbe, seesaw }) {
|
|
|
271
235
|
* Three decisions worth defending:
|
|
272
236
|
* • A TIE IS NOT BETTER. A tie that counted as an improvement would make a sawtooth look like a
|
|
273
237
|
* ratchet, and the whole point of the Day-1 measurement is to tell those two apart.
|
|
274
|
-
* • REGRESSIONS DOMINATE. Breaking a previously-finished scope is never an improvement, whatever
|
|
275
|
-
* the new scope's fixtures did. This is what lets the old seesaw branch collapse into the
|
|
276
|
-
* general rule rather than needing a special case.
|
|
277
238
|
* • DIFFERENT `fixtures_total` IS INCOMPARABLE, not worse. A re-slice changes the
|
|
278
239
|
* denominator; comparing across it is a category error, so the ratchet treats it as a baseline
|
|
279
240
|
* reset (`rebased`) rather than issuing a false verdict.
|
|
280
241
|
*
|
|
281
|
-
* @param {{
|
|
242
|
+
* @param {{fixtures_passed:number, fixtures_total:number,
|
|
282
243
|
* db_probe:(0|1|null)}} next - The candidate score.
|
|
283
|
-
* @param {({
|
|
244
|
+
* @param {({fixtures_passed:number, fixtures_total:number,
|
|
284
245
|
* db_probe:(0|1|null)}|null)} current - The incumbent score, or null for the first trial.
|
|
285
246
|
* @returns {(boolean|null)} true = strictly better · false = not better · null = incomparable.
|
|
286
247
|
*/
|
|
287
248
|
export function better(next, current) {
|
|
288
249
|
if (current === null || current === undefined) return true; // baseline
|
|
289
250
|
if (next.fixtures_total !== current.fixtures_total) return null; // the contract changed
|
|
290
|
-
if (next.regressions !== current.regressions) return next.regressions < current.regressions;
|
|
291
251
|
if (next.fixtures_passed !== current.fixtures_passed) return next.fixtures_passed > current.fixtures_passed;
|
|
292
252
|
if (next.db_probe !== current.db_probe) return (next.db_probe ?? 0) > (current.db_probe ?? 0);
|
|
293
253
|
// EVERY COMPONENT TIES. What that means depends entirely on whether the incumbent was green.
|
|
@@ -314,17 +274,17 @@ export function better(next, current) {
|
|
|
314
274
|
}
|
|
315
275
|
|
|
316
276
|
/**
|
|
317
|
-
* Is this score a clean pass — every fixture passing, none of them absent
|
|
277
|
+
* Is this score a clean pass — every fixture passing, and none of them absent?
|
|
318
278
|
*
|
|
319
279
|
* `fixtures_total > 0` is load-bearing: a scope with no fixtures has nothing to be green ABOUT, and
|
|
320
280
|
* treating its empty score as a pass is the same absence-reads-as-success mistake `runFixtures`
|
|
321
281
|
* made one function above.
|
|
322
282
|
*
|
|
323
|
-
* @param {{
|
|
283
|
+
* @param {{fixtures_passed:number, fixtures_total:number}} s - A trial score.
|
|
324
284
|
* @returns {boolean} True when the score represents a real, complete pass.
|
|
325
285
|
*/
|
|
326
286
|
function isGreenScore(s) {
|
|
327
|
-
return s.
|
|
287
|
+
return s.fixtures_total > 0 && s.fixtures_passed === s.fixtures_total;
|
|
328
288
|
}
|
|
329
289
|
|
|
330
290
|
/**
|
|
@@ -352,7 +312,7 @@ export function decideStatus(verdict, crashed) {
|
|
|
352
312
|
* Human-readable one-line summary of a score change, for the trial row's `delta` field.
|
|
353
313
|
* @param {object} next - The candidate score.
|
|
354
314
|
* @param {(object|null)} current - The incumbent score, or null.
|
|
355
|
-
* @returns {string} e.g. "+2 fixtures", "
|
|
315
|
+
* @returns {string} e.g. "+2 fixtures", "-1 db_probe", "baseline", "no change".
|
|
356
316
|
*/
|
|
357
317
|
export function describeDelta(next, current) {
|
|
358
318
|
if (!current) return "baseline";
|
|
@@ -360,10 +320,8 @@ export function describeDelta(next, current) {
|
|
|
360
320
|
return `denominator ${current.fixtures_total} → ${next.fixtures_total}`;
|
|
361
321
|
}
|
|
362
322
|
const parts = [];
|
|
363
|
-
const dr = next.regressions - current.regressions;
|
|
364
323
|
const df = next.fixtures_passed - current.fixtures_passed;
|
|
365
324
|
const dp = (next.db_probe ?? 0) - (current.db_probe ?? 0);
|
|
366
|
-
if (dr) parts.push(`${dr > 0 ? "+" : ""}${dr} regression${Math.abs(dr) === 1 ? "" : "s"}`);
|
|
367
325
|
if (df) parts.push(`${df > 0 ? "+" : ""}${df} fixture${Math.abs(df) === 1 ? "" : "s"}`);
|
|
368
326
|
if (dp) parts.push(`${dp > 0 ? "+" : ""}${dp} db_probe`);
|
|
369
327
|
return parts.length ? parts.join(", ") : "no change";
|
|
@@ -458,7 +416,7 @@ function sha256(text) {
|
|
|
458
416
|
* every superseded object remains addressable).
|
|
459
417
|
*
|
|
460
418
|
* WHAT THIS REPLACED, and why the remedy is `wx` rather than a guard. The address used to be
|
|
461
|
-
* `r<round>-a<attempt>.json`, written with a bare `writeFileSync` — and on a
|
|
419
|
+
* `r<round>-a<attempt>.json`, written with a bare `writeFileSync` — and on a revert-and-retry the
|
|
462
420
|
* protocol says stash, then RETRY THIS ATTEMPT, same attempt number. The address had no term for
|
|
463
421
|
* the retry, so the artifact recording the regression was silently replaced by the one recording
|
|
464
422
|
* the recovery, at the same path. Reproduced against the shipped script: two runs at
|
|
@@ -506,14 +464,12 @@ export function writeArtifact(outDir, round, attempt, verdictBody) {
|
|
|
506
464
|
/** The typed argv contract (see `./lib/argv.mjs`). */
|
|
507
465
|
export const ARGV_SPEC = {
|
|
508
466
|
usage: "harness.mjs verify t0 <scope-contract.json> --round N --attempt M [--cwd <dir>] [--out <dir>] " +
|
|
509
|
-
"[--
|
|
467
|
+
"[--no-ratchet]",
|
|
510
468
|
_: { arity: 1, max: 1, name: "scope-contract.json" },
|
|
511
469
|
round: { type: "int", min: 1, required: true },
|
|
512
470
|
attempt: { type: "int", min: 1, required: true },
|
|
513
471
|
cwd: { type: "path" },
|
|
514
472
|
out: { type: "path" },
|
|
515
|
-
"seesaw-registry": { type: "path" },
|
|
516
|
-
"no-seesaw": { type: "flag" },
|
|
517
473
|
"no-ratchet": { type: "flag" },
|
|
518
474
|
};
|
|
519
475
|
|
|
@@ -556,14 +512,7 @@ export async function cli(rawArgv) {
|
|
|
556
512
|
|
|
557
513
|
const fixtures = runFixtures(contract.e2e_verification_fixtures, cwd);
|
|
558
514
|
const dbProbe = runDbProbe(contract.db_probe, cwd);
|
|
559
|
-
|
|
560
|
-
// standalone CLI use without it simply skips the seesaw check rather than guessing a path.
|
|
561
|
-
const seesawRegistry = args.noSeesaw ? null : args.seesawRegistry || null;
|
|
562
|
-
const seesaw = args.noSeesaw || fixtures.pass === false
|
|
563
|
-
? { ran: false, pass: true, scopes_checked: [], failing: [] } // don't seesaw on an already-red attempt
|
|
564
|
-
: seesawCheck(seesawRegistry, cwd);
|
|
565
|
-
|
|
566
|
-
const verdict = computeVerdict({ fixtures, dbProbe, seesaw });
|
|
515
|
+
const verdict = computeVerdict({ fixtures, dbProbe });
|
|
567
516
|
const discovered = verdict.overall === "red" ? digestFailures({ fixtures, dbProbe }) : [];
|
|
568
517
|
|
|
569
518
|
// ---- the ratchet ---------------------------------------------------------------------
|
|
@@ -573,7 +522,7 @@ export async function cli(rawArgv) {
|
|
|
573
522
|
const trialsPath = join(outDir, "t0", "trials.jsonl");
|
|
574
523
|
const priorTrials = readTrials(trialsPath).filter((t) => t.scope_id === contract.scope_id);
|
|
575
524
|
const baseline = [...priorTrials].reverse().find((t) => t.status === "kept" || t.status === "rebased") || null;
|
|
576
|
-
const s = score({ fixtures, dbProbe
|
|
525
|
+
const s = score({ fixtures, dbProbe });
|
|
577
526
|
const verdictBetter = better(s, baseline ? baseline.score : null);
|
|
578
527
|
const crashed = fixtures.results.some((r) => r.error) || !!dbProbe?.error;
|
|
579
528
|
const { status, action } = decideStatus(verdictBetter, crashed);
|
|
@@ -598,7 +547,6 @@ export async function cli(rawArgv) {
|
|
|
598
547
|
// could not tell apart, and why `exit` still reads the way it always did.
|
|
599
548
|
fixtures: fixtures.results.map((r) => commandEvidence(r)),
|
|
600
549
|
db_probe: commandEvidence(dbProbe),
|
|
601
|
-
seesaw,
|
|
602
550
|
...verdict,
|
|
603
551
|
score: s,
|
|
604
552
|
discovered_tasks: discovered,
|
|
@@ -654,7 +602,7 @@ export async function cli(rawArgv) {
|
|
|
654
602
|
appendTrial(trialsPath, row);
|
|
655
603
|
|
|
656
604
|
console.log(JSON.stringify({
|
|
657
|
-
path, sha256: hash, trial, overall: verdict.overall,
|
|
605
|
+
path, sha256: hash, trial, overall: verdict.overall,
|
|
658
606
|
score: s, status, baseline_trial: row.baseline_trial, delta: row.delta,
|
|
659
607
|
tree_ref: row.tree_ref ?? keptRef(contract.scope_id),
|
|
660
608
|
}, null, 2));
|
package/package.json
CHANGED
|
@@ -614,7 +614,7 @@ or entirely `apps/api/**` with no cross-layer flow is the PA1 failure mode — r
|
|
|
614
614
|
}
|
|
615
615
|
```
|
|
616
616
|
`hill_phase` is always written `UPHILL_UNKNOWN` at generation time — it is derived later from
|
|
617
|
-
mechanical T0/T1
|
|
617
|
+
mechanical T0/T1 facts, never declared by `ba`. `superseded_by` stays
|
|
618
618
|
`null` until a scope-architect `map-scopes` order retires this contract in favor of its replacements.
|
|
619
619
|
|
|
620
620
|
**PA2 size lint:** a scope whose `allowed_file_substrate` glob set resolves to more than ~15
|
|
@@ -916,7 +916,7 @@
|
|
|
916
916
|
html += '</div>';
|
|
917
917
|
|
|
918
918
|
// ---- 2. hill chart (headline) ----
|
|
919
|
-
html += '<div class="section"><h3>Hill Chart</h3><div class="sub">Where each scope sits between figuring it out and getting it done — mechanical, derived only from T0/T1
|
|
919
|
+
html += '<div class="section"><h3>Hill Chart</h3><div class="sub">Where each scope sits between figuring it out and getting it done — mechanical, derived only from T0/T1 evidence, never self-reported. Position is how derisked a scope has ever been; dot color is its current-round health, so a scope can sit downhill and still show red if its latest attempt regressed.</div>';
|
|
920
920
|
html += '<div class="section-card"><div class="hill-wrap hero">';
|
|
921
921
|
if (pitch.hillAvailable) {
|
|
922
922
|
html += renderHillSVG(pitch.scopeStatus, 'hero');
|
|
@@ -77,8 +77,8 @@ the ship report's census table.
|
|
|
77
77
|
write-whitelist; wrong here =
|
|
78
78
|
a legitimate ESCALATE later
|
|
79
79
|
shared_substrate[] — files ≥2 scopes both touch;
|
|
80
|
-
|
|
81
|
-
|
|
80
|
+
declaring one costs concurrency:
|
|
81
|
+
they build one at a time
|
|
82
82
|
affordance_manifest — from ux-behavior.md state
|
|
83
83
|
tables: every interactive
|
|
84
84
|
element as {test_id, role} +
|
|
@@ -93,7 +93,7 @@ the ship report's census table.
|
|
|
93
93
|
TBD and flag it, never invent
|
|
94
94
|
a fixture for unbuilt behavior
|
|
95
95
|
hill_phase: "UPHILL_UNKNOWN" — ALWAYS; phase is derived from
|
|
96
|
-
T0/T1
|
|
96
|
+
T0/T1 facts later,
|
|
97
97
|
never authored
|
|
98
98
|
4 LINT node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify spec --slug <slug>
|
|
99
99
|
→ PA1 (directory alignment), PA2 (>~15 files), DISJOINT (undeclared overlap),
|
|
@@ -90,7 +90,7 @@ Collect (explicit — never inferred):
|
|
|
90
90
|
attempts a single scope gets inside one round before its attempt loop trips and
|
|
91
91
|
queues a hammer PROPOSAL for GATE H rather than blocking the round. Only meaningful
|
|
92
92
|
when the spec folder has scope contracts; a spec with none skips the attempt loop
|
|
93
|
-
entirely and BUILD behaves exactly as in v0.2.6 (task-executor --next, no T0
|
|
93
|
+
entirely and BUILD behaves exactly as in v0.2.6 (task-executor --next, no T0).
|
|
94
94
|
no_progress_k (v1.5): the STAGNATION term of the same inner breaker. Default 2 — the
|
|
95
95
|
number of consecutive non-`kept` trials after which a scope ends early and queues the
|
|
96
96
|
same GATE H proposal. attempt_budget counts ATTEMPTS and cannot see that the last two
|
|
@@ -205,7 +205,7 @@ ingest-result <result> → board/ledger writes
|
|
|
205
205
|
fails, and the run ABORTS naming the phase. Resolve it
|
|
206
206
|
yourself and record the answer in round-ledger.md, which
|
|
207
207
|
the NEXT attempt's fresh context reads back.
|
|
208
|
-
harness verify t0 → fixtures + DB probe
|
|
208
|
+
harness verify t0 → fixtures + DB probe, then scores the
|
|
209
209
|
attempt against the baseline trial and snapshots or
|
|
210
210
|
restores the tree. Branch on `status` from its stdout
|
|
211
211
|
JSON — the tree action has ALREADY happened:
|
|
@@ -217,9 +217,8 @@ harness verify t0 → fixtures + DB probe + (on green) sees
|
|
|
217
217
|
KEEPS, because a spec-conformance fix cannot raise a score that is already at
|
|
218
218
|
full marks, and reverting it would discard exactly the work a fix round exists
|
|
219
219
|
to do. Tree already restored from the last kept
|
|
220
|
-
snapshot. Subsumes the retired stash-and-retry branch:
|
|
221
|
-
|
|
222
|
-
is why seesaw runs before anything is declared green.
|
|
220
|
+
snapshot. Subsumes the retired stash-and-retry branch: an attempt that scores
|
|
221
|
+
worse than the incumbent reverts through this same rule.
|
|
223
222
|
rather than the code. Tree kept, baseline reset. Not a verdict, not a failure.
|
|
224
223
|
crash a fixture command failed to spawn or timed out; tree restored. Fix the fixture,
|
|
225
224
|
not the code.
|
|
@@ -446,9 +445,8 @@ Do: verify the phase's artifact by hand before trusting either diagnosis. If it
|
|
|
446
445
|
```
|
|
447
446
|
Invoke via Bash directly — NOT an Agent, this is deterministic tooling, not a worker:
|
|
448
447
|
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify t0 shapeup/<slug>/scopes/<scope-id>.md
|
|
449
|
-
--round <N> --attempt <M>
|
|
450
|
-
Effect: runs the scope's e2e fixtures
|
|
451
|
-
over every FINISHED scope's fixtures. Writes the verdict artifact spec-evaluator's
|
|
448
|
+
--round <N> --attempt <M>
|
|
449
|
+
Effect: runs the scope's e2e fixtures and DB probe. Writes the verdict artifact spec-evaluator's
|
|
452
450
|
T0-citation rule will require a citation to, appends one row to t0/trials.jsonl, and — this
|
|
453
451
|
is the ratchet — scores the attempt against the last kept trial and snapshots or
|
|
454
452
|
restores the working tree ITSELF. Zero LLM tokens — deterministic tooling, not a
|
|
@@ -592,7 +590,7 @@ guarantee lives in the script and, where noted, in a hook.
|
|
|
592
590
|
| Three-level circuit breaker: attempt_budget (inner, per scope) nests inside round_budget (outer), with an opt-in wall_clock_budget_s deadline | An exhausted scope queues a GATE H hammer proposal, it never blocks the round; only round_budget hitting 0 stops the whole run; the deadline breaker (checked every round boundary in `shapeup-run.js`) routes to GATE H so a run out of clock still ships what is green instead of being killed from outside |
|
|
593
591
|
| The tech lead never hand-edits a scope contract | scope-architect is its sole writer (single-writer-per-file) |
|
|
594
592
|
| Substrate-disjointness + PA1/PA2 lints are re-asserted at GATE L1b (harness verify spec) even when scope-architect already checked them | A human may have hand-approved past a 🔴 at the architect's checkpoint; `shapeup-run.js` runs spec-lint itself, in code, before resolving L1b |
|
|
595
|
-
| Hill phase is read from mechanical facts (T0/T1
|
|
593
|
+
| Hill phase is read from mechanical facts (T0/T1), never declared by a worker | Closes the self-reported-confidence risk outright — facts move dots, not authors |
|
|
596
594
|
| GATE H is delegated to scope-hammer, never adjudicated inline by the tech lead | Keeps the orchestrator thin; census/baseline-comparison/cut-list logic has one owner |
|
|
597
595
|
|
|
598
596
|
---
|