shapeup-sdlc 3.13.0 → 3.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +1 -1
- package/kernel/compile.mjs +89 -2
- package/kernel/probe/digest.mjs +10 -0
- package/kernel/schemas/domain.schema.json +14 -0
- package/kernel/verify/t0.mjs +105 -2
- package/package.json +1 -1
- package/skills/spec-evaluator/SKILL.md +1 -0
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.14.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -37,7 +37,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
38
|
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
39
39
|
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade |
|
|
40
|
-
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
|
|
40
|
+
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC. Every criterion the verdict graded FAIL reaches the round, whether or not the judge filed a bug for it, addressed to the scope that owns its use case — a round with a FAIL verdict and nothing to fix is how a run stalls. A row whose check the scope rewrote between a failing trial and a passing one is listed for the judge, who reads it against the row before the pass counts |
|
|
41
41
|
|
|
42
42
|
✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
|
|
43
43
|
|
package/kernel/compile.mjs
CHANGED
|
@@ -37,7 +37,7 @@ import { readRunId, dispatchReceipts, legLedger, readReceipt, receipt } from "./
|
|
|
37
37
|
// --spec-overridden directory, and the import is the convention-derived default.
|
|
38
38
|
import {
|
|
39
39
|
tasksDir, specDir as defaultSpecDir, roundLedger, trials, verdictsDir, ordersDir,
|
|
40
|
-
relShared, relLocal, globLocal, globShared, relKnowledgeBase, resultsDir, scopesDir,
|
|
40
|
+
relShared, relLocal, globLocal, globShared, relKnowledgeBase, resultsDir, scopesDir, localRoot,
|
|
41
41
|
} from "./lib/paths.mjs";
|
|
42
42
|
import { readContract, readAllContracts, tasksForScope, SCOPE_CONTRACT, reqId } from "./lib/contract.mjs";
|
|
43
43
|
import { writeActiveOrder } from "./probe/resume.mjs";
|
|
@@ -527,6 +527,57 @@ export function verdictBugs(cwd, slug, round) {
|
|
|
527
527
|
return v.bugs.filter((b) => !refuted.has(String(b?.id)) && !refuted.has(String(b?.criterion)));
|
|
528
528
|
}
|
|
529
529
|
|
|
530
|
+
/**
|
|
531
|
+
* The previous round's FAILED CRITERIA that the judge filed no bug for, as bug entries.
|
|
532
|
+
*
|
|
533
|
+
* A verdict carries two lists: the criteria it graded and the bugs it filed. Only the second
|
|
534
|
+
* reached the next round, so a criterion graded FAIL — "no evidence: no check names this row", a
|
|
535
|
+
* device row the judge could not observe — with no matching bug handed the fix round nothing to do.
|
|
536
|
+
* Measured: three fix rounds in a row compiled with no bugs over a FAIL verdict, and the run ended
|
|
537
|
+
* where it started. Every failed criterion is a thing the round must address, so each one without a
|
|
538
|
+
* bug becomes one, addressed to the scope whose contract lists the use case the criterion names.
|
|
539
|
+
*
|
|
540
|
+
* @param {string} cwd - Project root.
|
|
541
|
+
* @param {string} slug - Feature slug.
|
|
542
|
+
* @param {number} [round] - The round being compiled.
|
|
543
|
+
* @returns {Array<object>} `{criterion, severity, expected, actual, source:"criterion", scope_id?}`
|
|
544
|
+
* per uncovered FAIL criterion; [] for round 1, a PASS, or an unreadable result.
|
|
545
|
+
*/
|
|
546
|
+
export function criteriaBugs(cwd, slug, round) {
|
|
547
|
+
if (!round || round < 2) return [];
|
|
548
|
+
const p = join(resultsDir(cwd, slug), `evaluate-r${round - 1}.json`);
|
|
549
|
+
if (!existsSync(p)) return [];
|
|
550
|
+
let v;
|
|
551
|
+
try { v = JSON.parse(readFileSync(p, "utf8"))?.verdict; } catch { return []; }
|
|
552
|
+
if (v?.overall !== "FAIL" || !Array.isArray(v.criteria)) return [];
|
|
553
|
+
const filed = (Array.isArray(v.bugs) ? v.bugs : []).map((b) => String(b?.criterion ?? ""));
|
|
554
|
+
const refuted = new Set((Array.isArray(v.refuted) ? v.refuted : [])
|
|
555
|
+
.flatMap((r) => [r?.id, r?.ac_id, r?.criterion, typeof r === "string" ? r : null]).filter(Boolean).map(String));
|
|
556
|
+
const owners = new Map();
|
|
557
|
+
for (const { contract, id } of readAllContracts(scopesDir(cwd, slug))) {
|
|
558
|
+
const sid = contract?.scope_id || id;
|
|
559
|
+
for (const uc of Array.isArray(contract?.use_cases) ? contract.use_cases : []) if (!owners.has(uc)) owners.set(uc, sid);
|
|
560
|
+
}
|
|
561
|
+
const out = [];
|
|
562
|
+
for (const c of v.criteria) {
|
|
563
|
+
if (c?.verdict !== "FAIL" || typeof c.criterion !== "string") continue;
|
|
564
|
+
const name = c.criterion;
|
|
565
|
+
if (refuted.has(name)) continue;
|
|
566
|
+
if (filed.some((f) => f && (f === name || f.includes(name) || name.includes(f)))) continue;
|
|
567
|
+
const uc = (name.match(/\bUC-[A-Za-z0-9_-]+/) || [])[0];
|
|
568
|
+
const owner = uc ? owners.get(uc) : undefined;
|
|
569
|
+
out.push({
|
|
570
|
+
criterion: name,
|
|
571
|
+
severity: "major",
|
|
572
|
+
expected: "the criterion is met, with evidence the judge can cite",
|
|
573
|
+
actual: String(c.evidence ?? "graded FAIL").slice(0, 400),
|
|
574
|
+
source: "criterion",
|
|
575
|
+
...(owner ? { scope_id: owner } : {}),
|
|
576
|
+
});
|
|
577
|
+
}
|
|
578
|
+
return out;
|
|
579
|
+
}
|
|
580
|
+
|
|
530
581
|
/**
|
|
531
582
|
* The previous round's RED BUILD GATE, as bug entries the fix round can act on.
|
|
532
583
|
*
|
|
@@ -748,6 +799,38 @@ export function t0ArtifactsFor(cwd, slug, round) {
|
|
|
748
799
|
return { artifacts, missing };
|
|
749
800
|
}
|
|
750
801
|
|
|
802
|
+
// --- checks a scope rewrote between a failing trial and a passing one --------------------------
|
|
803
|
+
|
|
804
|
+
/**
|
|
805
|
+
* Every revised check any trial of the round recorded, with the scope it belongs to.
|
|
806
|
+
*
|
|
807
|
+
* Read across all of the round's trials rather than off the cited verdict alone: a check rewritten
|
|
808
|
+
* on attempt 2 and re-run unchanged on attempt 3 leaves no mark on attempt 3's verdict, which is the
|
|
809
|
+
* one the judge cites.
|
|
810
|
+
*
|
|
811
|
+
* @param {string} cwd - Project root.
|
|
812
|
+
* @param {string} slug - Feature slug.
|
|
813
|
+
* @param {number} [round] - The round being evaluated; omitted, every round.
|
|
814
|
+
* @returns {Array<{scope_id:string, id:string, file:string}>} Deduplicated by scope, id and file.
|
|
815
|
+
*/
|
|
816
|
+
export function revisedChecksFor(cwd, slug, round) {
|
|
817
|
+
const seen = new Set();
|
|
818
|
+
const out = [];
|
|
819
|
+
for (const tr of readTrials(trials(cwd, slug))) {
|
|
820
|
+
if (round != null && tr.round !== round) continue;
|
|
821
|
+
if (!tr.artifact) continue;
|
|
822
|
+
let v;
|
|
823
|
+
try { v = JSON.parse(readFileSync(join(localRoot(cwd, slug), tr.artifact), "utf8")); } catch { continue; }
|
|
824
|
+
for (const r of Array.isArray(v?.revised_checks) ? v.revised_checks : []) {
|
|
825
|
+
const key = `${tr.scope_id}\u0000${r.id}\u0000${r.file}`;
|
|
826
|
+
if (seen.has(key)) continue;
|
|
827
|
+
seen.add(key);
|
|
828
|
+
out.push({ scope_id: tr.scope_id, id: r.id, file: r.file });
|
|
829
|
+
}
|
|
830
|
+
}
|
|
831
|
+
return out;
|
|
832
|
+
}
|
|
833
|
+
|
|
751
834
|
// --- the launch evidence the judge grades `[ui]` rows against ---------------------------------
|
|
752
835
|
//
|
|
753
836
|
// WHY THE KERNEL DERIVES IT. A `[ui]` criterion is graded on the RUNNING app, and the evaluator's
|
|
@@ -1016,7 +1099,7 @@ export async function cli(rawArgv) {
|
|
|
1016
1099
|
// this line, and none of them can pass a payload to a build order (see the banner above).
|
|
1017
1100
|
// Two sources, one channel: the judge's cited defects and the build gate's failing steps.
|
|
1018
1101
|
const bugs = scope
|
|
1019
|
-
? bugsForScope([...verdictBugs(cwd, slug, round), ...buildBugs(cwd, slug, round)], scope.scope_id, scopeSubstrates(cwd, slug))
|
|
1102
|
+
? bugsForScope([...verdictBugs(cwd, slug, round), ...criteriaBugs(cwd, slug, round), ...buildBugs(cwd, slug, round)], scope.scope_id, scopeSubstrates(cwd, slug))
|
|
1020
1103
|
: [];
|
|
1021
1104
|
|
|
1022
1105
|
// A ROUND CARRYING CITED DEFECTS IS A `fix`, AND THE ORDER HAS TO SAY SO.
|
|
@@ -1138,6 +1221,10 @@ export async function cli(rawArgv) {
|
|
|
1138
1221
|
}
|
|
1139
1222
|
// The launch evidence, for every lane (see launchEvidenceFor). Same rule as above: an explicit
|
|
1140
1223
|
// `--payload` value outranks the derivation.
|
|
1224
|
+
if (operation === "evaluate" && payloadExtra.revised_checks === undefined) {
|
|
1225
|
+
const revised = revisedChecksFor(cwd, slug, round);
|
|
1226
|
+
if (revised.length) payloadExtra.revised_checks = revised;
|
|
1227
|
+
}
|
|
1141
1228
|
if (operation === "evaluate") {
|
|
1142
1229
|
const ev = launchEvidenceFor(cwd, slug, round);
|
|
1143
1230
|
if (ev.build_gate !== undefined && payloadExtra.build_gate === undefined) payloadExtra.build_gate = ev.build_gate;
|
package/kernel/probe/digest.mjs
CHANGED
|
@@ -33,6 +33,11 @@ const PATTERNS = [
|
|
|
33
33
|
// ("ERROR in the build pipeline", "ERROR in test suite failed to run") is left unmatched
|
|
34
34
|
// instead of handing back a fabricated file.
|
|
35
35
|
{ re: /^(?:ERROR|WARNING)\s+in\s+(\.{1,2}\/[^\s:]*|[^\s:]+\.[A-Za-z0-9]{1,10})\b/i, kind: "compiler-diagnostic" },
|
|
36
|
+
// hvigor / ArkTS compiler: the message and the location arrive on ONE line,
|
|
37
|
+
// "Error Message: Expected 5 arguments, but got 3. At File: /abs/path/Foo.test.ets:75:33"
|
|
38
|
+
// — the commonest failure on that toolchain, and one no other pattern here anchored, so every red
|
|
39
|
+
// compile handed the next attempt an empty error list.
|
|
40
|
+
{ re: /^\s*Error Message:\s*(.+?)\s+At File:\s*(.+?):(\d+):\d+\s*$/, kind: "arkts-compiler" },
|
|
36
41
|
// A test that FAILED BY NAME, with no file:line: "FAIL TS-05-05 step 4: no text 'Bread' on screen"
|
|
37
42
|
// or jest's "FAIL src/cart.test.js". Runners that drive an app from outside it (a device flow, an
|
|
38
43
|
// end-to-end script) report a case this way and nothing else, and the line is the whole signal —
|
|
@@ -77,6 +82,11 @@ export function digest(rawText) {
|
|
|
77
82
|
pendingMessage = coreMessage(m[1]);
|
|
78
83
|
continue; // wait for the stack frame that follows to get a file:line
|
|
79
84
|
}
|
|
85
|
+
if (kind === "arkts-compiler") {
|
|
86
|
+
triples.push({ file: m[2].trim(), line: Number(m[3]), core_message: coreMessage(m[1]), kind });
|
|
87
|
+
pendingMessage = null;
|
|
88
|
+
break;
|
|
89
|
+
}
|
|
80
90
|
const named = kind === "named-test-failure";
|
|
81
91
|
const file = named ? (/[\\/]|\.[A-Za-z0-9]{1,10}$/.test(m[1]) ? m[1] : null) : m[1]?.trim();
|
|
82
92
|
const lineNo = !named && m[2] ? Number(m[2]) : null;
|
|
@@ -362,6 +362,7 @@
|
|
|
362
362
|
"run_cmd",
|
|
363
363
|
"launch_cmd",
|
|
364
364
|
"build_gate",
|
|
365
|
+
"revised_checks",
|
|
365
366
|
"t0_artifacts",
|
|
366
367
|
"browser",
|
|
367
368
|
"tasks"
|
|
@@ -2429,6 +2430,19 @@
|
|
|
2429
2430
|
"type": "string",
|
|
2430
2431
|
"description": "spec-evaluator: the project profile's launch probe — installs the built artifact, starts it and asserts the first screen. Derived by `harness compile` from project-profile.md, absent when the profile declares none. Where `run_cmd` is only a build, this is how the app is brought up; a non-zero exit is a finding, not a reason to guess another way."
|
|
2431
2432
|
},
|
|
2433
|
+
"revised_checks": {
|
|
2434
|
+
"type": "array",
|
|
2435
|
+
"items": {
|
|
2436
|
+
"type": "object",
|
|
2437
|
+
"properties": {
|
|
2438
|
+
"scope_id": { "type": "string" },
|
|
2439
|
+
"id": { "type": "string" },
|
|
2440
|
+
"file": { "type": "string" }
|
|
2441
|
+
},
|
|
2442
|
+
"required": ["id", "file"]
|
|
2443
|
+
},
|
|
2444
|
+
"description": "spec-evaluator: rows that FAILed in one of the round's T0 trials and PASS in a later one whose own check file (named by the row id) changed in between. Derived by harness compile. Each must be read against its row before its PASS counts; absent when none."
|
|
2445
|
+
},
|
|
2432
2446
|
"build_gate": {
|
|
2433
2447
|
"type": "string",
|
|
2434
2448
|
"description": "spec-evaluator: this run's newest round build gate artifact (build/r<N>-t<T>.json) — each step's exit code and output tail, the record that the build ran and the app launched. Derived by `harness compile`, absent when the gate never ran. Read it before grading a [ui] row NO EVIDENCE."
|
package/kernel/verify/t0.mjs
CHANGED
|
@@ -34,8 +34,8 @@
|
|
|
34
34
|
//
|
|
35
35
|
// Exit code: 0 = overall green, 1 = overall red (mirrors the oracle convention), 2 = bad argv.
|
|
36
36
|
|
|
37
|
-
import { readFileSync, writeFileSync, appendFileSync, mkdirSync, existsSync, readdirSync } from "node:fs";
|
|
38
|
-
import { join, dirname } from "node:path";
|
|
37
|
+
import { readFileSync, writeFileSync, appendFileSync, mkdirSync, existsSync, readdirSync, statSync } from "node:fs";
|
|
38
|
+
import { join, dirname, basename, extname, resolve as resolvePath, relative } from "node:path";
|
|
39
39
|
import { spawnSync } from "node:child_process";
|
|
40
40
|
import { createHash } from "node:crypto";
|
|
41
41
|
import { digest } from "../probe/digest.mjs";
|
|
@@ -389,6 +389,97 @@ export function nextTrialNo(dir, round, attempt) {
|
|
|
389
389
|
return max + 1;
|
|
390
390
|
}
|
|
391
391
|
|
|
392
|
+
// --- a check rewritten to pass ----------------------------------------------------------------
|
|
393
|
+
//
|
|
394
|
+
// A fixture that names Test Surface rows (`PASS TS-02-04`, `FAIL TS-02-04 step 10: …`) reads its
|
|
395
|
+
// checks from files the scope itself writes. Measured on a live run: a row FAILed on attempt 1 —
|
|
396
|
+
// the check renamed a list and found the old name still on screen — and went green on attempt 2
|
|
397
|
+
// because the check had been rewritten to stop renaming, while the code was unchanged. Nothing
|
|
398
|
+
// recorded that the evidence had moved. So each verdict records the checks it read, by digest, and
|
|
399
|
+
// the named results they printed; a row that failed in this scope's previous trial and passes now,
|
|
400
|
+
// whose own check file changed in between, is a REVISED check — a pass the judge must read against
|
|
401
|
+
// its row before it can count. Recorded, never refused: rewriting a check that overreached is
|
|
402
|
+
// legitimate, and only a person or the judge can tell which it was.
|
|
403
|
+
|
|
404
|
+
/**
|
|
405
|
+
* The files a fixture's arguments name, with their digests — a named directory contributes its
|
|
406
|
+
* direct children. Paths resolve against the project root and are recorded relative to it.
|
|
407
|
+
*
|
|
408
|
+
* @param {string[]} commands - The fixture command lines.
|
|
409
|
+
* @param {string} cwd - Project root.
|
|
410
|
+
* @returns {Object<string,string>} Relative path → sha256 of its bytes. Empty when no argument names
|
|
411
|
+
* an existing path.
|
|
412
|
+
*/
|
|
413
|
+
export function checkFiles(commands, cwd) {
|
|
414
|
+
const out = {};
|
|
415
|
+
/**
|
|
416
|
+
* Record one file's digest under its project-relative path.
|
|
417
|
+
* @param {string} abs - Absolute path of the file.
|
|
418
|
+
* @returns {void} Nothing; an unreadable file is skipped, since it is not evidence.
|
|
419
|
+
*/
|
|
420
|
+
const add = (abs) => {
|
|
421
|
+
try { out[relative(cwd, abs)] = createHash("sha256").update(readFileSync(abs)).digest("hex"); } catch { /* unreadable — not evidence */ }
|
|
422
|
+
};
|
|
423
|
+
for (const cmd of commands || []) {
|
|
424
|
+
const tokens = String(cmd).match(/"[^"]*"|'[^']*'|\S+/g) || [];
|
|
425
|
+
for (const raw of tokens.slice(1)) {
|
|
426
|
+
const tok = raw.replace(/^["']|["']$/g, "");
|
|
427
|
+
if (!tok || tok.startsWith("-")) continue;
|
|
428
|
+
const abs = resolvePath(cwd, tok);
|
|
429
|
+
let st; try { st = statSync(abs); } catch { continue; }
|
|
430
|
+
if (st.isFile()) add(abs);
|
|
431
|
+
else if (st.isDirectory()) {
|
|
432
|
+
for (const f of readdirSync(abs).sort()) {
|
|
433
|
+
const child = join(abs, f);
|
|
434
|
+
try { if (statSync(child).isFile()) add(child); } catch { /* skip */ }
|
|
435
|
+
}
|
|
436
|
+
}
|
|
437
|
+
}
|
|
438
|
+
}
|
|
439
|
+
return out;
|
|
440
|
+
}
|
|
441
|
+
|
|
442
|
+
/**
|
|
443
|
+
* The rows a fixture's output names, and how each came out.
|
|
444
|
+
*
|
|
445
|
+
* @param {Array<{stdout?:string}>} results - Fixture results with their full stdout.
|
|
446
|
+
* @returns {Object<string,("PASS"|"FAIL")>} Row id → its result; a FAIL anywhere wins over a PASS.
|
|
447
|
+
*/
|
|
448
|
+
export function namedResults(results) {
|
|
449
|
+
const out = {};
|
|
450
|
+
for (const r of results || []) {
|
|
451
|
+
for (const line of String(r?.stdout || "").split(/\r?\n/)) {
|
|
452
|
+
const m = line.match(/^(PASS|FAIL)\s+(\S+)/);
|
|
453
|
+
if (!m) continue;
|
|
454
|
+
if (m[1] === "FAIL" || out[m[2]] !== "FAIL") out[m[2]] = m[1];
|
|
455
|
+
}
|
|
456
|
+
}
|
|
457
|
+
return out;
|
|
458
|
+
}
|
|
459
|
+
|
|
460
|
+
/**
|
|
461
|
+
* Rows that failed in the previous trial and pass now while their own check file changed.
|
|
462
|
+
*
|
|
463
|
+
* A row's check is the file whose name, without extension, is the row id — the convention a
|
|
464
|
+
* per-row check follows. A row with no such file is not reported: nothing can be said about it.
|
|
465
|
+
*
|
|
466
|
+
* @param {({check_files?:Object<string,string>, named_results?:Object<string,string>}|null)} prev
|
|
467
|
+
* The scope's previous verdict in this round, or null.
|
|
468
|
+
* @param {{check_files:Object<string,string>, named_results:Object<string,string>}} curr - This one.
|
|
469
|
+
* @returns {Array<{id:string, file:string}>} One entry per revised check.
|
|
470
|
+
*/
|
|
471
|
+
export function revisedChecks(prev, curr) {
|
|
472
|
+
if (!prev?.named_results || !prev?.check_files) return [];
|
|
473
|
+
const out = [];
|
|
474
|
+
for (const [id, now] of Object.entries(curr.named_results || {})) {
|
|
475
|
+
if (now !== "PASS" || prev.named_results[id] !== "FAIL") continue;
|
|
476
|
+
const file = Object.keys(curr.check_files || {}).find((f) => basename(f, extname(f)) === id);
|
|
477
|
+
if (!file) continue;
|
|
478
|
+
if (prev.check_files[file] !== curr.check_files[file]) out.push({ id, file });
|
|
479
|
+
}
|
|
480
|
+
return out;
|
|
481
|
+
}
|
|
482
|
+
|
|
392
483
|
/**
|
|
393
484
|
* Distill every failing command's output into AEGIS {file,line,core_message} triples.
|
|
394
485
|
* @param {{fixtures:{results:Array<{pass:boolean,stdout:string,stderr:string}>},
|
|
@@ -514,6 +605,10 @@ export async function cli(rawArgv) {
|
|
|
514
605
|
const dbProbe = runDbProbe(contract.db_probe, cwd);
|
|
515
606
|
const verdict = computeVerdict({ fixtures, dbProbe });
|
|
516
607
|
const discovered = verdict.overall === "red" ? digestFailures({ fixtures, dbProbe }) : [];
|
|
608
|
+
const checks = {
|
|
609
|
+
check_files: checkFiles(fixtures.results.map((r) => r.cmd), cwd),
|
|
610
|
+
named_results: namedResults(fixtures.results),
|
|
611
|
+
};
|
|
517
612
|
|
|
518
613
|
// ---- the ratchet ---------------------------------------------------------------------
|
|
519
614
|
// `current` is the incumbent: the score of the most recent trial whose TREE is the one on disk
|
|
@@ -526,6 +621,11 @@ export async function cli(rawArgv) {
|
|
|
526
621
|
const verdictBetter = better(s, baseline ? baseline.score : null);
|
|
527
622
|
const crashed = fixtures.results.some((r) => r.error) || !!dbProbe?.error;
|
|
528
623
|
const { status, action } = decideStatus(verdictBetter, crashed);
|
|
624
|
+
// The scope's previous trial in this round, read back for the checks it recorded.
|
|
625
|
+
const prevTrial = [...priorTrials].reverse().find((tr) => tr.round === round);
|
|
626
|
+
let prevVerdict = null;
|
|
627
|
+
if (prevTrial?.artifact) { try { prevVerdict = JSON.parse(readFileSync(join(outDir, prevTrial.artifact), "utf8")); } catch { /* none */ } }
|
|
628
|
+
const revised = revisedChecks(prevVerdict, checks);
|
|
529
629
|
|
|
530
630
|
// The run key, read from the receipt that lives in the run root this script was pointed at.
|
|
531
631
|
// `--out` IS that root, so identity comes from the receipt rather than from parsing a slug back
|
|
@@ -550,6 +650,9 @@ export async function cli(rawArgv) {
|
|
|
550
650
|
...verdict,
|
|
551
651
|
score: s,
|
|
552
652
|
discovered_tasks: discovered,
|
|
653
|
+
...(Object.keys(checks.check_files).length ? { check_files: checks.check_files } : {}),
|
|
654
|
+
...(Object.keys(checks.named_results).length ? { named_results: checks.named_results } : {}),
|
|
655
|
+
...(revised.length ? { revised_checks: revised } : {}),
|
|
553
656
|
});
|
|
554
657
|
|
|
555
658
|
// The tree operation. `--no-ratchet` leaves the working tree exactly as the attempt left it —
|
package/package.json
CHANGED
|
@@ -41,6 +41,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
|
|
|
41
41
|
| `payload.run_cmd` | How to start the running app. On a stack where building and launching are different acts it is only the build — prefer `payload.launch_cmd` when the order carries one. Absent standalone → ask; absent orchestrated with no `launch_cmd` either → ESCALATE, do not guess |
|
|
42
42
|
| `payload.launch_cmd` | The project profile's launch probe: installs the built artifact, starts it and asserts the first screen. This is how the app is brought up for `[ui]` probing. A non-zero exit is a finding to cite, not a reason to try another way. Absent → the profile declares none; fall back to `run_cmd` |
|
|
43
43
|
| `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
|
|
44
|
+
| `payload.revised_checks[]` | Rows that FAILed in one trial of this round and PASS in a later one whose own check file changed in between — a pass obtained by editing the check. Read each listed file against its row's Expect before you let its PASS count; a check that no longer asserts what the row asks makes that row a FAIL, and the bug names the check. Absent → no check was rewritten |
|
|
44
45
|
| `payload.t0_artifacts[]` | Per-scope T0 verdict paths for this round (scoped specs), compiled from each scope's green verdict. An artifact listed but missing/red on disk, or a scoped spec with none listed → the round is NOT gradeable: return `status: failed` with the reason, naming the scope, as your FIRST deviation — a structural precondition, not a criterion |
|
|
45
46
|
| `payload.browser` | `cli` (default, ~4x cheaper) \| `mcp` \| `none` |
|
|
46
47
|
| `payload.tasks[]` | Traceability only (which UCs a task claims): NEVER a grading source — the committed UC text is the criterion, a paraphrase mismatch is a finding |
|