shapeup-sdlc 3.9.2 → 3.10.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +1 -1
- package/kernel/compile.mjs +44 -1
- package/kernel/probe/t0.mjs +24 -3
- package/kernel/schemas/domain.schema.json +17 -0
- package/kernel/verify/build.mjs +56 -0
- package/package.json +1 -1
- package/skills/spec-evaluator/SKILL.md +7 -2
- package/skills/tech-lead/references/protocol.md +4 -1
- package/skills/tech-lead/workflows/shapeup-run.js +3 -0
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.10.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -36,7 +36,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
38
|
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
39
|
-
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
|
|
39
|
+
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade |
|
|
40
40
|
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
|
|
41
41
|
|
|
42
42
|
✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
|
package/kernel/compile.mjs
CHANGED
|
@@ -44,7 +44,7 @@ import { writeActiveOrder } from "./probe/resume.mjs";
|
|
|
44
44
|
import { greenVerdict } from "./probe/t0.mjs";
|
|
45
45
|
import { attemptEvidence, readReceipts } from "./probe/attempts.mjs";
|
|
46
46
|
import { readLegs } from "./probe/leg.mjs";
|
|
47
|
-
import { latestRoundBuild } from "./verify/build.mjs";
|
|
47
|
+
import { latestRoundBuild, latestRoundBuildFile, launchProbeFor } from "./verify/build.mjs";
|
|
48
48
|
// The SAME matcher the sandbox hook enforces with. "Is this cited file inside this scope's
|
|
49
49
|
// substrate" has to mean exactly what the guard means, or a bug is addressed to a scope that is
|
|
50
50
|
// then denied the write that fixes it.
|
|
@@ -748,6 +748,42 @@ export function t0ArtifactsFor(cwd, slug, round) {
|
|
|
748
748
|
return { artifacts, missing };
|
|
749
749
|
}
|
|
750
750
|
|
|
751
|
+
// --- the launch evidence the judge grades `[ui]` rows against ---------------------------------
|
|
752
|
+
//
|
|
753
|
+
// WHY THE KERNEL DERIVES IT. A `[ui]` criterion is graded on the RUNNING app, and the evaluator's
|
|
754
|
+
// contract for getting one is `payload.run_cmd`; absent, an orchestrated evaluator escalates rather
|
|
755
|
+
// than guess. The ledger's `run_cmd` is what the round build gate runs FIRST, as the build — on a
|
|
756
|
+
// toolchain where building and launching are different acts it is a build, and the ledger carries
|
|
757
|
+
// none at all when nobody pinned one. Meanwhile the gate had already run the project's launch probe
|
|
758
|
+
// (install, start, assert the first screen) moments earlier, recorded its output, and left the
|
|
759
|
+
// app up. None of that reached the order: the gate's artifact was forwarded only when RED, as the
|
|
760
|
+
// next round's bug list, so a green launch was proven and then unavailable to the one reader who
|
|
761
|
+
// needed it. Every `[ui]` row was graded "no evidence on the running app" over a build that launched.
|
|
762
|
+
//
|
|
763
|
+
// Two fields, both derived from disk, both optional, both absent when there is nothing to say —
|
|
764
|
+
// the same non-regression rule as `t0_artifacts`: a project that declares no launch probe compiles
|
|
765
|
+
// exactly the order it always did.
|
|
766
|
+
|
|
767
|
+
/**
|
|
768
|
+
* What an evaluate order should carry about launching the app.
|
|
769
|
+
*
|
|
770
|
+
* @param {string} cwd - Project root.
|
|
771
|
+
* @param {string} slug - Feature slug.
|
|
772
|
+
* @param {number} [round] - The round being evaluated. Omitted, the newest gate artifact of any
|
|
773
|
+
* round — a standalone evaluation has no round.
|
|
774
|
+
* @returns {{build_gate?: string, launch_cmd?: string}} `build_gate`: repo-relative path of this
|
|
775
|
+
* run's newest round build gate artifact (each step's exit and output tail). `launch_cmd`: the
|
|
776
|
+
* profile's launch probe. A key is omitted, never null, when its source is absent.
|
|
777
|
+
*/
|
|
778
|
+
export function launchEvidenceFor(cwd, slug, round) {
|
|
779
|
+
const out = {};
|
|
780
|
+
const gate = latestRoundBuildFile(cwd, slug, round);
|
|
781
|
+
if (gate) out.build_gate = relLocal(slug, "build", basename(gate.path));
|
|
782
|
+
const cmd = launchProbeFor(cwd, slug);
|
|
783
|
+
if (cmd) out.launch_cmd = cmd;
|
|
784
|
+
return out;
|
|
785
|
+
}
|
|
786
|
+
|
|
751
787
|
/**
|
|
752
788
|
* Assemble a WorkOrder envelope. Pure given its inputs — the CLI wrapper does the disk reads.
|
|
753
789
|
* @param {object} opts - The order inputs (destructured):
|
|
@@ -1093,6 +1129,13 @@ export async function cli(rawArgv) {
|
|
|
1093
1129
|
`${missing.join(", ")}; the evaluator has nothing to cite for ${missing.length === 1 ? "that scope" : "those scopes"}`);
|
|
1094
1130
|
}
|
|
1095
1131
|
}
|
|
1132
|
+
// The launch evidence, for every lane (see launchEvidenceFor). Same rule as above: an explicit
|
|
1133
|
+
// `--payload` value outranks the derivation.
|
|
1134
|
+
if (operation === "evaluate") {
|
|
1135
|
+
const ev = launchEvidenceFor(cwd, slug, round);
|
|
1136
|
+
if (ev.build_gate !== undefined && payloadExtra.build_gate === undefined) payloadExtra.build_gate = ev.build_gate;
|
|
1137
|
+
if (ev.launch_cmd !== undefined && payloadExtra.launch_cmd === undefined) payloadExtra.launch_cmd = ev.launch_cmd;
|
|
1138
|
+
}
|
|
1096
1139
|
|
|
1097
1140
|
const order = compileOrder({
|
|
1098
1141
|
slug, worker, operation, round, attempt, scope, tasks, decisions, digestedErrors, trialHistory, bugs,
|
package/kernel/probe/t0.mjs
CHANGED
|
@@ -1,8 +1,16 @@
|
|
|
1
1
|
// probe t0 — "has this scope already reached T0-green in this round?"
|
|
2
2
|
//
|
|
3
3
|
// CONTRACT. A bounded, read-only query over the verdict artifacts on disk. Prints
|
|
4
|
-
// `{green, path, round, scope_id}` on stdout; exits 0 when green, 1 when not, 2 on a bad
|
|
5
|
-
// Writes nothing.
|
|
4
|
+
// `{green, path, sha256?, round, scope_id}` on stdout; exits 0 when green, 1 when not, 2 on a bad
|
|
5
|
+
// argv. Writes nothing.
|
|
6
|
+
//
|
|
7
|
+
// WHY IT PRINTS A DIGEST. A verdict on a scoped spec must cite a T0 artifact with the sha256 of the
|
|
8
|
+
// file as it is on disk now, and the judge is asked to obtain that itself rather than accept one it
|
|
9
|
+
// was handed. Some sessions have no shell hasher they may run, and a grant that covers this entry
|
|
10
|
+
// point and not `shasum` is the ordinary case, not an oversight. The digest here is computed from
|
|
11
|
+
// the bytes at the moment of the call, with the same reading the ingest step re-checks it against,
|
|
12
|
+
// so a judge that asks for it is asking the file, not a caller. `sha256` is present only when the
|
|
13
|
+
// verdict is green and readable; absent — never null — otherwise.
|
|
6
14
|
//
|
|
7
15
|
// WHY IT IS A SUBCOMMAND AND NOT AN INLINE SNIPPET. The control plane has no filesystem of its
|
|
8
16
|
// own, so the alternative is a `node -e` blob assembled inside the workflow script. Such a blob is
|
|
@@ -14,6 +22,7 @@
|
|
|
14
22
|
// the only durable evidence of that is the verdict artifact the evaluator is required to cite.
|
|
15
23
|
|
|
16
24
|
import { existsSync, readdirSync, readFileSync } from "node:fs";
|
|
25
|
+
import { createHash } from "node:crypto";
|
|
17
26
|
import { join, resolve } from "node:path";
|
|
18
27
|
import { runArgs } from "../lib/argv.mjs";
|
|
19
28
|
import { verdictsDir, readRunId } from "../lib/paths.mjs";
|
|
@@ -89,6 +98,18 @@ export function cli(rawArgv) {
|
|
|
89
98
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
90
99
|
const cwd = resolve(args.cwd || process.cwd());
|
|
91
100
|
const { green, path } = greenVerdict(cwd, args.slug, args.scope, args.round);
|
|
92
|
-
|
|
101
|
+
const sha256 = green ? digestOf(path) : null;
|
|
102
|
+
console.log(JSON.stringify({ green, path, ...(sha256 ? { sha256 } : {}), scope_id: args.scope, round: args.round }));
|
|
93
103
|
process.exit(green ? 0 : 1);
|
|
94
104
|
}
|
|
105
|
+
|
|
106
|
+
/**
|
|
107
|
+
* The sha256 of a verdict artifact, read as `reduce ingest` re-reads it (utf8 text, then hashed) so
|
|
108
|
+
* the two can never disagree about the same bytes.
|
|
109
|
+
* @param {string} path - Absolute path of the verdict file.
|
|
110
|
+
* @returns {(string|null)} Lowercase hex digest, or null when the file cannot be read.
|
|
111
|
+
*/
|
|
112
|
+
function digestOf(path) {
|
|
113
|
+
try { return createHash("sha256").update(readFileSync(path, "utf8")).digest("hex"); }
|
|
114
|
+
catch { return null; }
|
|
115
|
+
}
|
|
@@ -301,6 +301,13 @@
|
|
|
301
301
|
"via": "discovered_tasks[]",
|
|
302
302
|
"note": "a red gate's output, digested — surfaces as payload.bugs on the next round's orders"
|
|
303
303
|
},
|
|
304
|
+
{
|
|
305
|
+
"from": "WorkOrder",
|
|
306
|
+
"to": "RoundBuildVerdict",
|
|
307
|
+
"cardinality": "N:1",
|
|
308
|
+
"via": "payload.build_gate",
|
|
309
|
+
"note": "an evaluate order names this run's newest gate artifact for its round, green or red, so the judge knows the build ran and the app launched before grading a [ui] row; payload.launch_cmd carries the profile's launch_probe beside it"
|
|
310
|
+
},
|
|
304
311
|
{
|
|
305
312
|
"from": "RoundBuildVerdict",
|
|
306
313
|
"to": "HillShard",
|
|
@@ -353,6 +360,8 @@
|
|
|
353
360
|
"feature",
|
|
354
361
|
"dimensions",
|
|
355
362
|
"run_cmd",
|
|
363
|
+
"launch_cmd",
|
|
364
|
+
"build_gate",
|
|
356
365
|
"t0_artifacts",
|
|
357
366
|
"browser",
|
|
358
367
|
"tasks"
|
|
@@ -2416,6 +2425,14 @@
|
|
|
2416
2425
|
"type": "string",
|
|
2417
2426
|
"description": "spec-evaluator: how to start the running app. Absent orchestrated → ESCALATE, never guess."
|
|
2418
2427
|
},
|
|
2428
|
+
"launch_cmd": {
|
|
2429
|
+
"type": "string",
|
|
2430
|
+
"description": "spec-evaluator: the project profile's launch probe — installs the built artifact, starts it and asserts the first screen. Derived by `harness compile` from project-profile.md, absent when the profile declares none. Where `run_cmd` is only a build, this is how the app is brought up; a non-zero exit is a finding, not a reason to guess another way."
|
|
2431
|
+
},
|
|
2432
|
+
"build_gate": {
|
|
2433
|
+
"type": "string",
|
|
2434
|
+
"description": "spec-evaluator: this run's newest round build gate artifact (build/r<N>-t<T>.json) — each step's exit code and output tail, the record that the build ran and the app launched. Derived by `harness compile`, absent when the gate never ran. Read it before grading a [ui] row NO EVIDENCE."
|
|
2435
|
+
},
|
|
2419
2436
|
"t0_artifacts": {
|
|
2420
2437
|
"type": "array",
|
|
2421
2438
|
"items": {
|
package/kernel/verify/build.mjs
CHANGED
|
@@ -220,6 +220,62 @@ export function latestRoundBuild(cwd, slug, round) {
|
|
|
220
220
|
return all.length ? all[all.length - 1].body : null;
|
|
221
221
|
}
|
|
222
222
|
|
|
223
|
+
/**
|
|
224
|
+
* The gate artifact an evaluate order should name — the newest one THIS run wrote, with the file it
|
|
225
|
+
* lives in.
|
|
226
|
+
*
|
|
227
|
+
* {@link latestRoundBuild} answers "what did the gate say about round N" and hands back a body; an
|
|
228
|
+
* order has to point at a file, so the path travels with it here. Scoped to the current run when a
|
|
229
|
+
* receipt is readable: a prior run over the same slug leaves its own `r1-t1.json` on disk, and the
|
|
230
|
+
* judge must not be pointed at another run's launch evidence. Scoped to one round when given; with
|
|
231
|
+
* no round it is the newest of any round — a standalone evaluation has none, the same reading
|
|
232
|
+
* `t0ArtifactsFor` uses for the T0 verdicts.
|
|
233
|
+
*
|
|
234
|
+
* @param {string} cwd - Project root.
|
|
235
|
+
* @param {string} slug - Feature slug.
|
|
236
|
+
* @param {number} [round] - The round being evaluated.
|
|
237
|
+
* @returns {({path:string, body:object}|null)} null when the gate never ran, or nothing readable
|
|
238
|
+
* belongs to this run. Never throws: an order must compile without it.
|
|
239
|
+
*/
|
|
240
|
+
export function latestRoundBuildFile(cwd, slug, round) {
|
|
241
|
+
let files;
|
|
242
|
+
const dir = roundBuildDir(cwd, slug);
|
|
243
|
+
try { files = readdirSync(dir); } catch { return null; }
|
|
244
|
+
const runId = readRunId(cwd, slug);
|
|
245
|
+
let best = null;
|
|
246
|
+
for (const f of files) {
|
|
247
|
+
const m = f.match(/^r(\d+)-t(\d+)\.json$/);
|
|
248
|
+
if (!m) continue;
|
|
249
|
+
const r = Number(m[1]), t = Number(m[2]);
|
|
250
|
+
if (round != null && r !== Number(round)) continue;
|
|
251
|
+
let body;
|
|
252
|
+
try { body = JSON.parse(readFileSync(join(dir, f), "utf8")); } catch { continue; }
|
|
253
|
+
if (runId && body?.run_id && body.run_id !== runId) continue;
|
|
254
|
+
if (!best || r > best.r || (r === best.r && t > best.t)) best = { r, t, path: join(dir, f), body };
|
|
255
|
+
}
|
|
256
|
+
return best && { path: best.path, body: best.body };
|
|
257
|
+
}
|
|
258
|
+
|
|
259
|
+
/**
|
|
260
|
+
* The command the project declared for bringing the built app up — `launch_probe` in the committed
|
|
261
|
+
* profile: install the artifact, start it, assert the first screen. Read on its own rather than
|
|
262
|
+
* through {@link declaredSteps}, which also reads the ledger and computes warnings an order has no
|
|
263
|
+
* use for.
|
|
264
|
+
*
|
|
265
|
+
* @param {string} cwd - Project root.
|
|
266
|
+
* @param {string} slug - Feature slug.
|
|
267
|
+
* @returns {(string|null)} The command, or null when the profile is absent or declares none.
|
|
268
|
+
* Never throws.
|
|
269
|
+
*/
|
|
270
|
+
export function launchProbeFor(cwd, slug) {
|
|
271
|
+
try {
|
|
272
|
+
const pp = projectProfile(cwd, slug);
|
|
273
|
+
if (!existsSync(pp)) return null;
|
|
274
|
+
const v = readContract(pp, PROJECT_PROFILE)?.contract?.launch_probe;
|
|
275
|
+
return typeof v === "string" && v.trim() ? v.trim() : null;
|
|
276
|
+
} catch { return null; }
|
|
277
|
+
}
|
|
278
|
+
|
|
223
279
|
/**
|
|
224
280
|
* The rounds whose latest gate artifact is red — the set `reduce hill` subtracts from.
|
|
225
281
|
* @param {string} cwd - Project root.
|
package/package.json
CHANGED
|
@@ -36,7 +36,9 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
|
|
|
36
36
|
| `payload.spec_folder` | The committed grading truth: `usecases/` + `domain-model.md` (+ `contracts/`, `scope-summary.md`, `_index.md`). No `usecases/` → HARD STOP, nothing to grade against |
|
|
37
37
|
| `payload.feature` | Feature slug — scopes the probe and names the report |
|
|
38
38
|
| `payload.dimensions[]` | The active dimension set (the caller resolved precedence). Absent → `[spec-conformance]` + the auto-enable rules below |
|
|
39
|
-
| `payload.run_cmd` | How to start the running app. Absent standalone → ask; absent orchestrated → ESCALATE, do not guess |
|
|
39
|
+
| `payload.run_cmd` | How to start the running app. On a stack where building and launching are different acts it is only the build — prefer `payload.launch_cmd` when the order carries one. Absent standalone → ask; absent orchestrated with no `launch_cmd` either → ESCALATE, do not guess |
|
|
40
|
+
| `payload.launch_cmd` | The project profile's launch probe: installs the built artifact, starts it and asserts the first screen. This is how the app is brought up for `[ui]` probing. A non-zero exit is a finding to cite, not a reason to try another way. Absent → the profile declares none; fall back to `run_cmd` |
|
|
41
|
+
| `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
|
|
40
42
|
| `payload.t0_artifacts[]` | Per-scope T0 verdict paths for this round (scoped specs), compiled from each scope's green verdict. An artifact listed but missing/red on disk, or a scoped spec with none listed → the round is NOT gradeable: return `status: failed` with the reason, naming the scope, as your FIRST deviation — a structural precondition, not a criterion |
|
|
41
43
|
| `payload.browser` | `cli` (default, ~4x cheaper) \| `mcp` \| `none` |
|
|
42
44
|
| `payload.tasks[]` | Traceability only (which UCs a task claims): NEVER a grading source — the committed UC text is the criterion, a paraphrase mismatch is a finding |
|
|
@@ -94,7 +96,10 @@ Done-when statements; `_index.md` Non-Go list. Which UCs are in scope comes from
|
|
|
94
96
|
contract triplet + Non-Go). Overall PASS only if ALL active dimensions pass — the halo
|
|
95
97
|
effect is banned; a strong dimension never lifts a failing one.
|
|
96
98
|
- **T0 citation (scoped specs).** Recompute each cited artifact's sha256 from disk — never
|
|
97
|
-
trust a handed hash.
|
|
99
|
+
trust a handed hash. Use any hasher you may run; when none is available to you,
|
|
100
|
+
`harness probe t0 --slug <slug> --scope <scope_id> --round <r>` prints the digest of that
|
|
101
|
+
scope's green verdict as the file is at that moment. A digest from any other place — a ledger
|
|
102
|
+
line, an earlier report, an error message — is a handed hash, however correct it looks. A verdict on a scoped spec without a T0 citation is structurally
|
|
98
103
|
invalid, regardless of how convincing your own probing looked; generator prose ("tests
|
|
99
104
|
pass", "verified locally") is never admissible evidence.
|
|
100
105
|
- **When a criterion names a command, run THAT command.** Not the one that works, not the
|
|
@@ -490,6 +490,9 @@ compile-order --operation evaluate --slug <slug> --worker spec-evaluator --round
|
|
|
490
490
|
--payload '{"dimensions": ["spec-conformance"], "run_cmd": "<cmd>"}'
|
|
491
491
|
t0_artifacts is compiled from each scope's green T0 verdict for round <r> — pass it only to
|
|
492
492
|
override. A scope with no green verdict is named on stderr: the judge has nothing to cite for it.
|
|
493
|
+
build_gate and launch_cmd are compiled too: this run's newest round build gate artifact, and the
|
|
494
|
+
profile's launch_probe. Neither is passed here — a `[ui]` row is graded on the running app, and
|
|
495
|
+
these are how the judge finds out the app was launched and how to bring it up again.
|
|
493
496
|
Invoke via Agent (model: eval), ONCE, after GATE L2:
|
|
494
497
|
Skill(shapeup-sdlc-plugin:spec-evaluator) --order <path>
|
|
495
498
|
Effect: one feature-level pass over the running app against all AC + Done-when; writes
|
|
@@ -548,7 +551,7 @@ Read back: the proposed cut list + verdict (SHIP now | SHIP after fixing ship-bl
|
|
|
548
551
|
| `harness-run.md` | **tech lead (sole writer)** | tech lead (round ledger + Hill + run-state), PO (audit) |
|
|
549
552
|
| `scopes/<scope-id>.md` | `scope-architect` (sole writer) | tech lead (substrate/sequence), sandbox hook (write-whitelist), compile-order (inlined into orders) |
|
|
550
553
|
| `t0/verdicts/r<N>-a<M>-t<T>.json` | `harness verify t0` (skill-local, mechanical — not a worker) | spec-evaluator (required citation), tech lead (hill derivation), compile-order (digested errors) |
|
|
551
|
-
| `build/r<N>-t<T>.json` (the round build gate) | `harness verify build` (mechanical — run_cmd + build_probe + launch_probe, once per round before EVAL) | tech lead (GATE L2 block), `harness reduce hill` (a red round moves no dot), compile-order (`payload.bugs` for round N+1) |
|
|
554
|
+
| `build/r<N>-t<T>.json` (the round build gate) | `harness verify build` (mechanical — run_cmd + build_probe + launch_probe, once per round before EVAL) | tech lead (GATE L2 block), `harness reduce hill` (a red round moves no dot), compile-order (`payload.bugs` for round N+1 when red; `payload.build_gate` on the same round's evaluate order) |
|
|
552
555
|
| `t0/trials.jsonl` (the ratchet ledger, append-only, `baseline_trial` as the parent link) | `harness verify t0` (one row per attempt: score, status, delta, tree_ref) | compile-order (`trial_history` into the next order), ship-report (T0 + Ratchet sections), `harness probe stats --ratchet` |
|
|
553
556
|
| `round-ledger.md` | **tech lead (sole writer)** | compile-order (decisions into every order), PO (audit) |
|
|
554
557
|
| `hill/<scope-id>.yml` + `hill-chart.md` | **tech lead (sole writer)** | PO ("status without asking"), scope-hammer (H0 census) |
|
|
@@ -1837,6 +1837,9 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1837
1837
|
model: evalModel, round,
|
|
1838
1838
|
// No `t0_artifacts` here, deliberately: `harness compile` derives them from the round's green
|
|
1839
1839
|
// T0 verdicts on disk, for every lane — this script could only name paths it was told about.
|
|
1840
|
+
// The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
|
|
1841
|
+
// artifact and the project profile, so the judge gets the launch evidence without this script
|
|
1842
|
+
// having to carry it.
|
|
1840
1843
|
payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
|
|
1841
1844
|
extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself.",
|
|
1842
1845
|
});
|