shapeup-sdlc 3.7.12 → 3.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +2 -2
- package/README.md +13 -11
- package/kernel/lib/argv.mjs +4 -4
- package/kernel/lib/paths.mjs +0 -2
- package/kernel/reduce/hill.mjs +17 -39
- package/kernel/report/export.mjs +0 -3
- package/kernel/schemas/domain.schema.json +13 -92
- package/kernel/verify/env.mjs +2 -1
- package/kernel/verify/ratchet-tree.mjs +1 -1
- package/kernel/verify/t0.mjs +35 -87
- package/kernel/verify/trace.mjs +230 -30
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/references/task-generation.md +1 -1
- package/skills/hill-chart/assets/dashboard.template.html +1 -1
- package/skills/scope-architect/SKILL.md +3 -3
- package/skills/tech-lead/references/gates.md +20 -6
- package/skills/tech-lead/references/protocol.md +6 -8
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.9.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -35,7 +35,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
35
35
|
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
|
-
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe
|
|
38
|
+
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
39
39
|
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
|
|
40
40
|
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
|
|
41
41
|
|
|
@@ -63,7 +63,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
|
|
|
63
63
|
- **QA is a level-up, not a gate** — `--no-qa` skips it; circuit breaker outranks the Hunter.
|
|
64
64
|
- **Role separation** — Evaluator grades, task-executor fixes, QA discovers.
|
|
65
65
|
- **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
|
|
66
|
-
- **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts
|
|
66
|
+
- **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts, never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
|
|
67
67
|
- **Envelope port (v1.0)** — every dispatch is WorkOrder in / WorkResult out; shared state has exactly one writer (the ingest step); malformed envelopes are hook-denied. Workers: stateless, craft-only, pipeline-blind.
|
|
68
68
|
|
|
69
69
|
## Setup & Execution
|
package/README.md
CHANGED
|
@@ -45,11 +45,11 @@ its fixtures and its DB probe and writes an artifact to disk — with each comma
|
|
|
45
45
|
its captured output, and whether it ran at all. The evaluator must cite that artifact, and the
|
|
46
46
|
hill phase is derived from artifacts rather than from a worker's own account of its progress.
|
|
47
47
|
Two limits, stated here because the point of this section is that a claim without a mechanism
|
|
48
|
-
behind it is the thing this harness exists to prevent:
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
48
|
+
behind it is the thing this harness exists to prevent: a citation is re-hashed from disk, checked
|
|
49
|
+
against the scope, round and run the artifact records, and refused unless it names a file this
|
|
50
|
+
run's own verifier wrote — what that does NOT prove is that a judge with a dishonest hand could
|
|
51
|
+
not have arranged the artifact first, which is why the judge's own substrate freezes the verdicts
|
|
52
|
+
directory. A T0 artifact is
|
|
53
53
|
also evidence about the machine that produced it, and now says so: each verdict carries where it
|
|
54
54
|
ran — the absolute path, the git tree, the resolved toolchain, lockfile digests, declared cache
|
|
55
55
|
directories and a digest over an allowlist of environment values — so a disagreeing re-run can be
|
|
@@ -132,15 +132,14 @@ rest of this README after this table and nothing will be a surprise.
|
|
|
132
132
|
|---|---|
|
|
133
133
|
| **board** | The round's task list. "Green" means every task is done. GATE L2's hook reads this before an evaluation and warns if it is not green. |
|
|
134
134
|
| **round** | One build → evaluate cycle. A FAIL verdict starts round *r+1*. |
|
|
135
|
-
| **T0** | The smoke test a scope must pass before it counts as built: its fixtures
|
|
136
|
-
| **seesaw** | The regression arm of T0, meant to re-run *other* scopes' fixtures so a regression is never mistaken for progress. Declared, not yet wired: no run writes its registry, so no run has executed it. |
|
|
135
|
+
| **T0** | The smoke test a scope must pass before it counts as built: its fixtures and a DB probe. Writes an artifact to disk that the evaluator must cite. |
|
|
137
136
|
| **substrate** | The exact list of files one dispatch is allowed to write, stamped into its work order. A hook blocks anything outside it — and anything the order marks frozen. |
|
|
138
137
|
| **scope contract** | The file defining one vertical slice: its substrate, its fixtures, its affordances. |
|
|
139
138
|
| **affordance** | The thing a user can actually click, type or call. UI is graded on affordances, not on looks. |
|
|
140
139
|
| **hill / hill phase** | How much of a scope is still *unknown* versus merely *unfinished*. Derived from T0 facts — never self-reported. |
|
|
141
140
|
| **gate (L0–L4)** | A numbered checkpoint in a run. Most pause for you; GATE L2 is the one a hook observes and reports on. |
|
|
142
141
|
| **covers-closure** | Every requirement clause has at least one task claiming to cover it. Nothing silently drops. |
|
|
143
|
-
| **wiring reachability** | Every engine has a call site reachable from the app's real entry point. Catches "built, but never wired up". |
|
|
142
|
+
| **wiring reachability** | Every engine has a call site reachable from the app's real entry point. Catches "built, but never wired up". Reports itself unchecked, with a reason, when the import walk cannot be rooted — an unfollowable import, or no reachable engine to control it. |
|
|
144
143
|
| **discovery ledger** | The one file everything found mid-run gets written to, so nothing is lost between rounds. |
|
|
145
144
|
|
|
146
145
|
A longer version, including the internals, is in [docs/glossary.md](docs/glossary.md).
|
|
@@ -309,15 +308,18 @@ These hold across the harness and are the reason it stays predictable:
|
|
|
309
308
|
of blocking the round. An opt-in third breaker bounds the **wall clock**, because the other two
|
|
310
309
|
count events and neither can notice a single round running for half an hour — tripping it routes
|
|
311
310
|
to GATE H, so a run out of time ships what is green instead of being killed and shipping nothing.
|
|
312
|
-
- **Hill phase is mechanical, never self-reported** — derived only from T0/T1 facts
|
|
313
|
-
|
|
311
|
+
- **Hill phase is mechanical, never self-reported** — derived only from T0/T1 facts, closing the
|
|
312
|
+
self-reported-confidence risk. A scope with no
|
|
314
313
|
discovery ledger derives no phase rather than a solved one, so absence no longer reads as
|
|
315
314
|
progress on that arm.
|
|
316
315
|
- **One writer per shared file** — every board/ledger/verdict write goes through
|
|
317
316
|
`harness reduce ingest`; workers return data and never touch shared state.
|
|
318
317
|
- **Traceability is oracle-checked, opt-in** — `harness verify trace` verifies covers-closure and
|
|
319
318
|
wiring reachability from the committed spine artifacts; it ships advisory (warn-only) and every
|
|
320
|
-
arm is skipped when its artifact is absent, so older specs are non-regressed.
|
|
319
|
+
arm is skipped when its artifact is absent, so older specs are non-regressed. Reachability also
|
|
320
|
+
skips when it cannot root its walk — an import it cannot follow, or no reachable engine to
|
|
321
|
+
control the result — because "every engine is orphaned" and "this is the wrong entry point" are
|
|
322
|
+
the same evidence, and a check that cannot tell them apart must say so rather than pick.
|
|
321
323
|
|
|
322
324
|
## Known rough edges
|
|
323
325
|
|
package/kernel/lib/argv.mjs
CHANGED
|
@@ -8,11 +8,11 @@
|
|
|
8
8
|
// const SPEC = {
|
|
9
9
|
// _: { arity: 1, name: "scope-contract.json" },
|
|
10
10
|
// round: { type: "int", min: 1, required: true },
|
|
11
|
-
// "no-
|
|
11
|
+
// "no-ratchet": { type: "flag" },
|
|
12
12
|
// };
|
|
13
|
-
// const args = runArgs(SPEC, argv); // args.round, args.
|
|
13
|
+
// const args = runArgs(SPEC, argv); // args.round, args.noRatchet, args._
|
|
14
14
|
//
|
|
15
|
-
// Flag names reach the caller camelCased (`--no-
|
|
15
|
+
// Flag names reach the caller camelCased (`--no-ratchet` → `noRatchet`). Unknown flags are rejected
|
|
16
16
|
// rather than swallowed as positionals: a typo'd `--rounds 2` landing in `_` is the same defect
|
|
17
17
|
// wearing a different hat. Untyped coercion is the failure this guards — `Number(undefined)` is
|
|
18
18
|
// `NaN`, `??` does not catch `NaN`, and a verdict written to `r NaN-a1.json` with exit 0 is
|
|
@@ -35,7 +35,7 @@ export class ArgvError extends Error {
|
|
|
35
35
|
}
|
|
36
36
|
}
|
|
37
37
|
|
|
38
|
-
/** `--
|
|
38
|
+
/** `--attempt-budget` → `attemptBudget`. */
|
|
39
39
|
function camel(name) {
|
|
40
40
|
return name.replace(/-([a-z0-9])/g, (_, c) => c.toUpperCase());
|
|
41
41
|
}
|
package/kernel/lib/paths.mjs
CHANGED
|
@@ -216,8 +216,6 @@ export const trials = (cwd, slug) => join(t0Dir(cwd, slug), "trials.jsonl");
|
|
|
216
216
|
* {@link decisions}, one small file with one writer.
|
|
217
217
|
*/
|
|
218
218
|
export const gates = (cwd, slug) => join(localRoot(cwd, slug), "gates.jsonl");
|
|
219
|
-
/** Finished-scope fixture registry for the seesaw regression check. */
|
|
220
|
-
export const seesawRegistry = (cwd, slug) => join(localRoot(cwd, slug), "seesaw", "registry.json");
|
|
221
219
|
/**
|
|
222
220
|
* The round build gate's verdicts — one immutable artifact per gate run, `r<N>-t<T>.json`.
|
|
223
221
|
*
|
package/kernel/reduce/hill.mjs
CHANGED
|
@@ -144,8 +144,8 @@ function committedPhase(hDir, id) {
|
|
|
144
144
|
* - UPHILL_UNKNOWN: open unknowns > 0 in the ledger for this scope — and the floor the scope sits
|
|
145
145
|
* at whenever the ledger has not answered at all, which is where every run legitimately begins
|
|
146
146
|
* - UPHILL_SOLVED: the ledger was read and reports zero open unknowns, no T0-green yet
|
|
147
|
-
* - DOWNHILL_EXECUTION: ≥1 T0-green in a round whose build gate is not red; T1
|
|
148
|
-
* - FINISHED: T1 PASS ∧
|
|
147
|
+
* - DOWNHILL_EXECUTION: ≥1 T0-green in a round whose build gate is not red; T1 pending
|
|
148
|
+
* - FINISHED: T1 PASS ∧ a T0-green that a red round did not invalidate
|
|
149
149
|
*
|
|
150
150
|
* @param {string} cwd - The project root directory.
|
|
151
151
|
* @param {string} slug - The feature slug being built.
|
|
@@ -230,18 +230,16 @@ export function deriveHill(cwd, slug) {
|
|
|
230
230
|
}
|
|
231
231
|
}
|
|
232
232
|
|
|
233
|
-
// 2. T0 facts per scope: has it achieved a green overall verdict
|
|
233
|
+
// 2. T0 facts per scope: has it achieved a green overall verdict this run?
|
|
234
234
|
//
|
|
235
|
-
//
|
|
236
|
-
//
|
|
237
|
-
//
|
|
238
|
-
//
|
|
239
|
-
//
|
|
240
|
-
//
|
|
241
|
-
// verdict counting exactly as before.
|
|
235
|
+
// FINISHED USED TO WAIT ON AN ARM THAT NEVER RAN. The phase required `seesaw.ran && seesaw.pass`,
|
|
236
|
+
// nothing in the codebase ever wrote the registry that arm read, and so no scope in any recorded
|
|
237
|
+
// run reached FINISHED — 38 committed shards on the live consumer, not one of them. The arm was
|
|
238
|
+
// removed in 3.8.0 by decision; the precondition goes with it, and the top phase is reachable
|
|
239
|
+
// again on the evidence that does exist: T1 passed, and a T0 green from a round the build gate
|
|
240
|
+
// did not red.
|
|
242
241
|
const redRounds = redBuildRounds(cwd, slug);
|
|
243
242
|
const t0Facts = {};
|
|
244
|
-
// This run's verdicts only — a prior run's green over the same slug moved this run's dot.
|
|
245
243
|
const hillRunId = readRunId(cwd, slug);
|
|
246
244
|
if (existsSync(vDir)) {
|
|
247
245
|
for (const f of readdirSync(vDir)) {
|
|
@@ -249,41 +247,21 @@ export function deriveHill(cwd, slug) {
|
|
|
249
247
|
try {
|
|
250
248
|
const b = JSON.parse(readFileSync(join(vDir, f), "utf8"));
|
|
251
249
|
if (hillRunId && b.run_id && b.run_id !== hillRunId) continue;
|
|
252
|
-
if (!t0Facts[b.scope_id]) t0Facts[b.scope_id] = { hasGreen: false
|
|
253
|
-
if (b.overall === "green" && !redRounds.has(Number(b.round)))
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
// inferring it from a false `regression` flag. That inference was vacuously true on every
|
|
258
|
-
// green T0 whether or not a seesaw check ever ran: nothing in this codebase currently
|
|
259
|
-
// passes `--seesaw-registry` to `verify t0`, so `seesaw.ran` is always `false` today and
|
|
260
|
-
// `regression` is always `false` too — "not asked" was being read as "clean," letting a
|
|
261
|
-
// scope reach FINISHED on a regression check that had never executed.
|
|
262
|
-
//
|
|
263
|
-
// Betting Table decision (Phase 3.5 / S4): wiring the seesaw registry for real is a
|
|
264
|
-
// genuine feature with a real running cost (re-running every finished scope's fixtures
|
|
265
|
-
// on every later attempt) and is out of proportion to a certification-gap fix. Deferred,
|
|
266
|
-
// not silently dropped — a scope with no registry wired simply cannot reach FINISHED via
|
|
267
|
-
// this path today, which is the honest state of the system: this check was never really
|
|
268
|
-
// gating FINISHED before either.
|
|
269
|
-
if (b.seesaw?.ran && b.seesaw?.pass) {
|
|
270
|
-
t0Facts[b.scope_id].seesawGreen = true;
|
|
271
|
-
}
|
|
272
|
-
}
|
|
273
|
-
} catch (e) {
|
|
274
|
-
// ignore parse errors
|
|
250
|
+
if (!t0Facts[b.scope_id]) t0Facts[b.scope_id] = { hasGreen: false };
|
|
251
|
+
if (b.overall === "green" && !redRounds.has(Number(b.round))) t0Facts[b.scope_id].hasGreen = true;
|
|
252
|
+
} catch {
|
|
253
|
+
// A torn or unreadable verdict proves nothing about the scope — skip it rather than let it
|
|
254
|
+
// decide a phase.
|
|
275
255
|
}
|
|
276
256
|
}
|
|
277
257
|
}
|
|
278
|
-
|
|
279
|
-
// 3. Ledger unknowns per scope — `null` for every scope when the ledger itself was not readable
|
|
280
|
-
// or nothing in it named a scope (see `ledgerUnknowns`). Only a real count can promote.
|
|
258
|
+
|
|
281
259
|
const scopeUnknowns = ledgerUnknowns(ledgerPath, scopes);
|
|
282
260
|
|
|
283
261
|
const report = [];
|
|
284
262
|
for (const s of scopes) {
|
|
285
263
|
const id = s.scope_id;
|
|
286
|
-
const t0 = t0Facts[id] || { hasGreen: false
|
|
264
|
+
const t0 = t0Facts[id] || { hasGreen: false };
|
|
287
265
|
// `null` = the ledger did not answer; a number = it did. `|| 0` collapsed the two.
|
|
288
266
|
const unknowns = scopeUnknowns === null ? null : (scopeUnknowns[id] || 0);
|
|
289
267
|
|
|
@@ -291,7 +269,7 @@ export function deriveHill(cwd, slug) {
|
|
|
291
269
|
// it — including before Orient has filed anything, which is where every run legitimately
|
|
292
270
|
// starts. Only an ANSWERED count of zero promotes to UPHILL_SOLVED; `null` never does.
|
|
293
271
|
let phase = "UPHILL_UNKNOWN";
|
|
294
|
-
if (t1Pass && t0.hasGreen
|
|
272
|
+
if (t1Pass && t0.hasGreen) {
|
|
295
273
|
phase = "FINISHED";
|
|
296
274
|
} else if (t0.hasGreen) {
|
|
297
275
|
phase = "DOWNHILL_EXECUTION";
|
package/kernel/report/export.mjs
CHANGED
|
@@ -129,7 +129,6 @@ function t0Row(a, runId) {
|
|
|
129
129
|
regression: a?.regression ?? null,
|
|
130
130
|
fixtures_green: a?.fixtures_green ?? null,
|
|
131
131
|
db_probe_green: a?.db_probe_green ?? null,
|
|
132
|
-
seesaw_green: a?.seesaw_green ?? null,
|
|
133
132
|
// One field a reader compares, and the block itself stays in the artifact for a human to diff:
|
|
134
133
|
// two rows with the same tree and different env digests are two machines, not a regression.
|
|
135
134
|
env_sha256: a?.env?.env_sha256 ?? null,
|
|
@@ -137,8 +136,6 @@ function t0Row(a, runId) {
|
|
|
137
136
|
tree_dirty: a?.env?.tree?.dirty ?? null,
|
|
138
137
|
fixtures_total: fixtures.length,
|
|
139
138
|
fixtures_passed: fixtures.filter((f) => f?.pass === true).length,
|
|
140
|
-
seesaw_ran: a?.seesaw?.ran ?? null,
|
|
141
|
-
seesaw_failing: Array.isArray(a?.seesaw?.failing) ? a.seesaw.failing.length : null,
|
|
142
139
|
discovered_tasks: Array.isArray(a?.discovered_tasks) ? a.discovered_tasks.length : 0,
|
|
143
140
|
};
|
|
144
141
|
}
|
|
@@ -236,7 +236,7 @@
|
|
|
236
236
|
"to": "HillShard",
|
|
237
237
|
"cardinality": "1:0..1",
|
|
238
238
|
"via": "scope_id",
|
|
239
|
-
"note": "phase derived from T0/T1
|
|
239
|
+
"note": "phase derived from T0/T1 facts, never authored"
|
|
240
240
|
},
|
|
241
241
|
{
|
|
242
242
|
"from": "ScopeContract",
|
|
@@ -252,13 +252,6 @@
|
|
|
252
252
|
"via": "fixtures[] + db_probe",
|
|
253
253
|
"note": "produced by actually running the commands — no agent can fabricate them"
|
|
254
254
|
},
|
|
255
|
-
{
|
|
256
|
-
"from": "T0Artifact",
|
|
257
|
-
"to": "SeesawCheck",
|
|
258
|
-
"cardinality": "1:1",
|
|
259
|
-
"via": "seesaw",
|
|
260
|
-
"note": "the regression half of T0 — every FINISHED scope's fixtures re-run"
|
|
261
|
-
},
|
|
262
255
|
{
|
|
263
256
|
"from": "T0Artifact",
|
|
264
257
|
"to": "AegisTriple",
|
|
@@ -287,13 +280,6 @@
|
|
|
287
280
|
"via": "payload.trial_history[]",
|
|
288
281
|
"note": "inspect(): the loop reads its own recent history, across the round boundary"
|
|
289
282
|
},
|
|
290
|
-
{
|
|
291
|
-
"from": "SeesawRegistry",
|
|
292
|
-
"to": "ScopeContract",
|
|
293
|
-
"cardinality": "1:N",
|
|
294
|
-
"via": "scopes[].scope_id",
|
|
295
|
-
"note": "every FINISHED scope's fixtures re-run on each later attempt"
|
|
296
|
-
},
|
|
297
283
|
{
|
|
298
284
|
"from": "UseCase",
|
|
299
285
|
"to": "Seam",
|
|
@@ -528,7 +514,7 @@
|
|
|
528
514
|
"items": {
|
|
529
515
|
"type": "string"
|
|
530
516
|
},
|
|
531
|
-
"description": "Globs ≥2 scopes both touch —
|
|
517
|
+
"description": "Globs ≥2 scopes both touch — an edit here is read-modify-write, so two scopes that both declare one never build at the same time."
|
|
532
518
|
},
|
|
533
519
|
"append_only": {
|
|
534
520
|
"type": "array",
|
|
@@ -844,7 +830,7 @@
|
|
|
844
830
|
"items": {
|
|
845
831
|
"type": "string"
|
|
846
832
|
},
|
|
847
|
-
"description": "Files ≥2 scopes both touch — must be declared in BOTH scopes' lists (spec-lint DISJOINT)
|
|
833
|
+
"description": "Files ≥2 scopes both touch — must be declared in BOTH scopes' lists (spec-lint DISJOINT); declaring one costs concurrency, since two scopes that share a path build one at a time."
|
|
848
834
|
},
|
|
849
835
|
"affordance_manifest": {
|
|
850
836
|
"type": "array",
|
|
@@ -872,7 +858,7 @@
|
|
|
872
858
|
"DOWNHILL_EXECUTION",
|
|
873
859
|
"FINISHED"
|
|
874
860
|
],
|
|
875
|
-
"description": "ALWAYS authored as UPHILL_UNKNOWN — self-reported confidence is the risk this closes. The live phase is DERIVED from T0/T1
|
|
861
|
+
"description": "ALWAYS authored as UPHILL_UNKNOWN — self-reported confidence is the risk this closes. The live phase is DERIVED from T0/T1 facts into hill/<scope-id>.yml — facts move dots, not authors."
|
|
876
862
|
},
|
|
877
863
|
"superseded_by": {
|
|
878
864
|
"type": "array",
|
|
@@ -1254,35 +1240,8 @@
|
|
|
1254
1240
|
}
|
|
1255
1241
|
}
|
|
1256
1242
|
},
|
|
1257
|
-
|
|
1258
|
-
"description": "The
|
|
1259
|
-
"x-tier": "EMBEDDED",
|
|
1260
|
-
"type": "object",
|
|
1261
|
-
"properties": {
|
|
1262
|
-
"ran": {
|
|
1263
|
-
"type": "boolean",
|
|
1264
|
-
"description": "false when no registry was passed or the attempt was already red (don't seesaw a red attempt)."
|
|
1265
|
-
},
|
|
1266
|
-
"pass": {
|
|
1267
|
-
"type": "boolean"
|
|
1268
|
-
},
|
|
1269
|
-
"scopes_checked": {
|
|
1270
|
-
"type": "array",
|
|
1271
|
-
"items": {
|
|
1272
|
-
"type": "string"
|
|
1273
|
-
}
|
|
1274
|
-
},
|
|
1275
|
-
"failing": {
|
|
1276
|
-
"type": "array",
|
|
1277
|
-
"items": {
|
|
1278
|
-
"type": "string"
|
|
1279
|
-
},
|
|
1280
|
-
"description": "scope_ids whose fixtures broke."
|
|
1281
|
-
}
|
|
1282
|
-
}
|
|
1283
|
-
},
|
|
1284
|
-
"T0Artifact": {
|
|
1285
|
-
"description": "The mechanical verification verdict for one build attempt — the evidence layer under the LLM judge. Written by harness verify t0 from actually running the scope's fixtures + DB probe + seesaw; zero LLM tokens. spec-evaluator must cite it (sha256) on scoped specs. Red artifacts carry AEGIS triples that become the next attempt's digested_errors.",
|
|
1243
|
+
"T0Artifact": {
|
|
1244
|
+
"description": "The mechanical verification verdict for one build attempt — the evidence layer under the LLM judge. Written by harness verify t0 from actually running the scope's fixtures and DB probe; zero LLM tokens. spec-evaluator must cite it (sha256) on scoped specs. Red artifacts carry AEGIS triples that become the next attempt's digested_errors.",
|
|
1286
1245
|
"x-tier": "LOCAL",
|
|
1287
1246
|
"x-location": ".shapeup/<slug>/t0/verdicts/r<N>-a<M>-t<T>.json (schema_version 2; the unsuffixed r<N>-a<M>.json of schema_version 1 is still readable)",
|
|
1288
1247
|
"x-writer": "harness verify t0",
|
|
@@ -1327,9 +1286,6 @@
|
|
|
1327
1286
|
"db_probe": {
|
|
1328
1287
|
"$ref": "#/$defs/CommandResult"
|
|
1329
1288
|
},
|
|
1330
|
-
"seesaw": {
|
|
1331
|
-
"$ref": "#/$defs/SeesawCheck"
|
|
1332
|
-
},
|
|
1333
1289
|
"fixtures_green": {
|
|
1334
1290
|
"type": "boolean"
|
|
1335
1291
|
},
|
|
@@ -1337,9 +1293,6 @@
|
|
|
1337
1293
|
"type": "boolean",
|
|
1338
1294
|
"description": "true when no probe declared (null probe never counts as failure)."
|
|
1339
1295
|
},
|
|
1340
|
-
"seesaw_green": {
|
|
1341
|
-
"type": "boolean"
|
|
1342
|
-
},
|
|
1343
1296
|
"overall": {
|
|
1344
1297
|
"type": "string",
|
|
1345
1298
|
"enum": [
|
|
@@ -1347,10 +1300,6 @@
|
|
|
1347
1300
|
"red"
|
|
1348
1301
|
]
|
|
1349
1302
|
},
|
|
1350
|
-
"regression": {
|
|
1351
|
-
"type": "boolean",
|
|
1352
|
-
"description": "fixtures+db green but seesaw red — the rollback+retry case."
|
|
1353
|
-
},
|
|
1354
1303
|
"score": {
|
|
1355
1304
|
"$ref": "#/$defs/T0Score",
|
|
1356
1305
|
"description": "schema_version 2+: the comparable outcome vector better() ranks. A reduce over fixtures[] — no new measurement."
|
|
@@ -1384,11 +1333,6 @@
|
|
|
1384
1333
|
"x-tier": "EMBEDDED",
|
|
1385
1334
|
"type": "object",
|
|
1386
1335
|
"properties": {
|
|
1387
|
-
"regressions": {
|
|
1388
|
-
"type": "integer",
|
|
1389
|
-
"minimum": 0,
|
|
1390
|
-
"description": "Previously-FINISHED scopes now failing (seesaw). Dominates: breaking a shipped scope is never an improvement."
|
|
1391
|
-
},
|
|
1392
1336
|
"fixtures_passed": {
|
|
1393
1337
|
"type": "integer",
|
|
1394
1338
|
"minimum": 0
|
|
@@ -1412,7 +1356,6 @@
|
|
|
1412
1356
|
}
|
|
1413
1357
|
},
|
|
1414
1358
|
"required": [
|
|
1415
|
-
"regressions",
|
|
1416
1359
|
"fixtures_passed",
|
|
1417
1360
|
"fixtures_total"
|
|
1418
1361
|
]
|
|
@@ -1619,34 +1562,7 @@
|
|
|
1619
1562
|
}
|
|
1620
1563
|
}
|
|
1621
1564
|
},
|
|
1622
|
-
|
|
1623
|
-
"description": "The fixture registry of every FINISHED scope — what seesawCheck re-runs on each later attempt so a new scope cannot silently break a shipped one — a regression mistaken for progress is the pathology the seesaw exists for.",
|
|
1624
|
-
"x-tier": "LOCAL",
|
|
1625
|
-
"x-location": ".shapeup/<slug>/seesaw/registry.json",
|
|
1626
|
-
"x-writer": "tech-lead (when a scope reaches FINISHED)",
|
|
1627
|
-
"x-readers": "harness verify t0",
|
|
1628
|
-
"type": "object",
|
|
1629
|
-
"properties": {
|
|
1630
|
-
"scopes": {
|
|
1631
|
-
"type": "array",
|
|
1632
|
-
"items": {
|
|
1633
|
-
"type": "object",
|
|
1634
|
-
"properties": {
|
|
1635
|
-
"scope_id": {
|
|
1636
|
-
"type": "string"
|
|
1637
|
-
},
|
|
1638
|
-
"fixtures": {
|
|
1639
|
-
"type": "array",
|
|
1640
|
-
"items": {
|
|
1641
|
-
"type": "string"
|
|
1642
|
-
}
|
|
1643
|
-
}
|
|
1644
|
-
}
|
|
1645
|
-
}
|
|
1646
|
-
}
|
|
1647
|
-
}
|
|
1648
|
-
},
|
|
1649
|
-
"VerdictLedgerLine": {
|
|
1565
|
+
"VerdictLedgerLine": {
|
|
1650
1566
|
"description": "One appended line of judge history (JSONL). Never rewritten — flips across runs are detected here and force confidence low. run auto-increments per append batch.",
|
|
1651
1567
|
"x-tier": "LOCAL",
|
|
1652
1568
|
"x-location": ".shapeup/<slug>/evaluation/.verdicts-<target>.jsonl",
|
|
@@ -1690,7 +1606,7 @@
|
|
|
1690
1606
|
}
|
|
1691
1607
|
},
|
|
1692
1608
|
"HillShard": {
|
|
1693
|
-
"description": "One scope's hill position — DERIVED, never self-reported: UPHILL_UNKNOWN (open unknowns > 0) → UPHILL_SOLVED (unknowns 0, no T0-green yet) → DOWNHILL_EXECUTION (≥1 T0-green; T1
|
|
1609
|
+
"description": "One scope's hill position — DERIVED, never self-reported: UPHILL_UNKNOWN (open unknowns > 0) → UPHILL_SOLVED (unknowns 0, no T0-green yet) → DOWNHILL_EXECUTION (≥1 T0-green from a round that built; T1 pending) → FINISHED (T1 PASS ∧ that T0-green). Single-writer = whoever holds that scope's branch. Progress is reported by hill position, never task counts.",
|
|
1694
1610
|
"x-tier": "SHARED",
|
|
1695
1611
|
"x-location": "shapeup/<slug>/hill/<scope-id>.yml",
|
|
1696
1612
|
"x-writer": "tech-lead (GATE L2 derivation)",
|
|
@@ -2371,6 +2287,11 @@
|
|
|
2371
2287
|
"type": "string",
|
|
2372
2288
|
"description": "Optional context on the archetype/entry-point choice."
|
|
2373
2289
|
},
|
|
2290
|
+
"source_extensions": {
|
|
2291
|
+
"type": "array",
|
|
2292
|
+
"items": { "type": "string" },
|
|
2293
|
+
"description": "OPTIONAL — the file extensions this project's modules use, for the import walk that reachability runs. Absent means the JS/TS family plus the entry point's own suffix, which is right for most stacks and wrong for any stack whose engines end in something the entry point does not. Declare it when they differ (e.g. [\".ets\"] for ArkTS, [\".vue\", \".ts\"] for a Vue app): a walk that cannot follow an edge reports itself unchecked rather than reporting every engine orphaned, so the cost of leaving this absent is a skipped arm, never a false red."
|
|
2294
|
+
},
|
|
2374
2295
|
"build_probe": {
|
|
2375
2296
|
"type": "string",
|
|
2376
2297
|
"description": "OPTIONAL — a command that asserts the BUILT ARTIFACT, not the build's exit code, and exits 0 only when it holds. Exists because a green build is not proof the feature compiled: some toolchains compile only the files reachable from an entry point, so a scope's new files can sit outside the compiled set while the build stays green (measured: an app package holding 3 compiled files, 58 errors once the rest became reachable). Archetype-specific by construction — e.g. 'the compiled source map lists every file under each scope's substrate'. Run by harness verify build after run_cmd, once per round before EVAL; absent = no such step, never a failure."
|
package/kernel/verify/env.mjs
CHANGED
|
@@ -106,7 +106,8 @@ function treeState(cwd) {
|
|
|
106
106
|
* path, which is exactly the mechanism that made one tree build three ways.
|
|
107
107
|
*
|
|
108
108
|
* `null` means the profile declared nothing — NOT that there are none. "Not asked" and "none" are
|
|
109
|
-
* different facts, and collapsing them is the mistake the
|
|
109
|
+
* different facts, and collapsing them is the mistake that kept the hill's top phase shut for the
|
|
110
|
+
* life of the seesaw arm: "not asked" was recorded the same way as "nothing wrong".
|
|
110
111
|
*
|
|
111
112
|
* @param {(string|null)} profilePath - `shapeup/<slug>/project-profile.md`, when the caller knows it.
|
|
112
113
|
* @returns {(object[]|null)} One entry per declared cache, or null when none is declared.
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
//
|
|
5
5
|
// The attempt loop branched a red T0 two ways, and only one of them reverted anything:
|
|
6
6
|
//
|
|
7
|
-
// • a
|
|
7
|
+
// • a trial that scored worse than the incumbent → `git stash push -u`;
|
|
8
8
|
// • a red on the scope's OWN fixtures → "loop to the next attempt", and no revert at all.
|
|
9
9
|
//
|
|
10
10
|
// So the failing tree stayed on the branch, and attempt N+1's fresh, zero-memory subagent began
|