shapeup-sdlc 3.16.1 → 3.17.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +3 -3
- package/kernel/compile.mjs +5 -0
- package/kernel/harness.mjs +3 -2
- package/kernel/probe/hunt.mjs +104 -0
- package/kernel/probe/requirements.mjs +33 -1
- package/kernel/reduce/ingest.mjs +20 -1
- package/kernel/reduce/ship.mjs +2 -1
- package/kernel/schemas/domain.schema.json +11 -0
- package/package.json +1 -1
- package/skills/qa-edge-hunter/SKILL.md +3 -1
- package/skills/spec-evaluator/SKILL.md +1 -0
- package/skills/tech-lead/workflows/shapeup-run.js +116 -41
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.17.1",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -32,11 +32,11 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
32
32
|
| Kick-off | ⏸ **L0** — Intake & Config (L0.8 model/budget matrix) + worker roster ✧ | `/translator` if non-English |
|
|
33
33
|
| Orient (Scout) | ⏸ **L1a** — Orient Review | `/orient` |
|
|
34
34
|
| Requirements | — (reviewed at L1b) | `/ba-pitch-analyzer` (`coverage`): the pitch's clauses → committed `requirements.md`, one atomic clause per `REQ-<n>` row naming the pitch clause it came from, ids assigned once and frozen; dispatched once, ahead of Analyze because the acceptance criteria are what cite its ids. A registry already on disk is not re-dispatched |
|
|
35
|
-
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
|
|
35
|
+
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along. A board with no such clause beside a registry sends its writer back once, and one that still has none continues with a warning |
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
38
|
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
39
|
-
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade. A PASS names every Test Surface row as its own criterion: one that grades the surface as a group is refused, and the judge is sent back once with the reason |
|
|
39
|
+
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade. A PASS names every Test Surface row as its own criterion: one that grades the surface as a group is refused, and the judge is sent back once with the reason. A scope that did not go green in the round has its rows graded FAIL, and the order names such scopes, so one red scope never makes the whole round ungradeable |
|
|
40
40
|
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC. Every criterion the verdict graded FAIL reaches the round, whether or not the judge filed a bug for it, addressed to the scope that owns its use case — a round with a FAIL verdict and nothing to fix is how a run stalls. A row whose check the scope rewrote between a failing trial and a passing one is listed for the judge, who reads it against the row before the pass counts |
|
|
41
41
|
|
|
42
42
|
✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
|
|
@@ -45,7 +45,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
45
45
|
|
|
46
46
|
|
|
47
47
|
### QA Edge Hunt (`/qa-edge-hunter`, post-PASS, pre-ship)
|
|
48
|
-
**Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`.
|
|
48
|
+
**Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`. A hunt that reached the app and ran no charter is sent back once; if it still ran none, the report and the run's close say `not-hunted`, never that QA ran.
|
|
49
49
|
|
|
50
50
|
### Ship & Triage
|
|
51
51
|
- **SHIP S.0 / GATE H** — `/scope-hammer`: census (QA findings + discovered ledger + attempt-budget proposals ✦; every "no scope owns X" cites `probe owner`, which derives ownership from the contracts) → baseline comparison (never the ideal) → cut list; TL/PO promotes selected items only. The census is written as data beside the report, and GATE L4 reads that file — and nothing else — before any answer set may say `ship`. A breaker does not end the run: the run itself dispatches the census, crosses GATE H, writes the ship report with the verdict as it is, crosses L4, and only then closes — `shipped` with the cut list when the census clears it, `escalated` naming the census when it does not.
|
package/kernel/compile.mjs
CHANGED
|
@@ -1267,6 +1267,11 @@ export async function cli(rawArgv) {
|
|
|
1267
1267
|
if (operation === "evaluate" && payloadExtra.t0_artifacts === undefined) {
|
|
1268
1268
|
const { artifacts, missing } = t0ArtifactsFor(cwd, slug, round);
|
|
1269
1269
|
if (artifacts.length) payloadExtra.t0_artifacts = artifacts;
|
|
1270
|
+
// A scope that did not go green leaves its rows ungraded, not the round: the judge refused a whole
|
|
1271
|
+
// round over one such scope and the run aborted at L3, before the next round or GATE H's census
|
|
1272
|
+
// could act on it. With at least one citation the round is gradeable, and the order names the
|
|
1273
|
+
// scopes whose rows are FAILs for want of a green T0.
|
|
1274
|
+
if (artifacts.length && missing.length) payloadExtra.scopes_without_t0 = missing;
|
|
1270
1275
|
// On stderr, never stdout: stdout is the order path the caller consumes.
|
|
1271
1276
|
if (missing.length) {
|
|
1272
1277
|
console.error(`compile-order: warning — no green T0 verdict${round ? ` in round ${round}` : ""} for ` +
|
package/kernel/harness.mjs
CHANGED
|
@@ -32,7 +32,8 @@
|
|
|
32
32
|
// gate An answer file with a source, not a vibe.
|
|
33
33
|
// probe resume · t0 · stats · digest · Read-only queries over run state. `concurrency`
|
|
34
34
|
// concurrency · leg · eval · answers how many legs ran at once and what the
|
|
35
|
-
// owner · requirements · attempts
|
|
35
|
+
// owner · requirements · attempts ·
|
|
36
|
+
// hunt
|
|
36
37
|
// fan-out bought, and refuses a figure the record set
|
|
37
38
|
// cannot support rather than printing a plausible one.
|
|
38
39
|
// `leg` answers whether a scope's work reached the
|
|
@@ -92,7 +93,7 @@ export const ROUTES = {
|
|
|
92
93
|
resume: "./probe/resume.mjs", t0: "./probe/t0.mjs", stats: "./probe/stats.mjs",
|
|
93
94
|
digest: "./probe/digest.mjs", concurrency: "./probe/concurrency.mjs",
|
|
94
95
|
leg: "./probe/leg.mjs", eval: "./probe/eval.mjs", owner: "./probe/owner.mjs",
|
|
95
|
-
requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs",
|
|
96
|
+
requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs", hunt: "./probe/hunt.mjs",
|
|
96
97
|
},
|
|
97
98
|
init: { run: "./init/run.mjs", fit: "./init/fit.mjs", "run-args": "./init/run-args.mjs" },
|
|
98
99
|
report: { export: "./report/export.mjs", _default: "export" },
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// probe hunt — "did the QA hunt hunt anything?"
|
|
3
|
+
//
|
|
4
|
+
// CONTRACT. A bounded, read-only query over the hunt's WorkResult, its order and its report.
|
|
5
|
+
// Prints `{ok, status, qa, charters_run, findings, reason}` on stdout; exits 0 when the hunt's
|
|
6
|
+
// result is one the run may accept, 1 when it is not, and 2 on a bad argv. Writes nothing.
|
|
7
|
+
//
|
|
8
|
+
// WHY THIS EXISTS. A hunter that reached the app, swept the fault log and then drafted no charter
|
|
9
|
+
// still returned `done`, and a `done` with no findings reads as an app with nothing wrong. The hunt
|
|
10
|
+
// report's charter count says what actually happened. A `done` hunt over an app its order could
|
|
11
|
+
// reach, with no charter run, is refused here and on ingest, and the run sends the hunter back once
|
|
12
|
+
// — the same shape as a verdict that grades nothing by name.
|
|
13
|
+
|
|
14
|
+
import { existsSync, readFileSync } from "node:fs";
|
|
15
|
+
import { join, resolve } from "node:path";
|
|
16
|
+
import { runArgs } from "../lib/argv.mjs";
|
|
17
|
+
import { resultsDir, ordersDir, qaDir } from "../lib/paths.mjs";
|
|
18
|
+
|
|
19
|
+
/**
|
|
20
|
+
* The number of charters a hunt report says it ran.
|
|
21
|
+
*
|
|
22
|
+
* @param {(string|null)} text - The hunt report's text.
|
|
23
|
+
* @returns {(number|null)} The run count from its `charters: <run>/<approved>` line, or null when
|
|
24
|
+
* there is no report or no such line.
|
|
25
|
+
*/
|
|
26
|
+
export function chartersRun(text) {
|
|
27
|
+
const m = String(text ?? "").match(/^charters:\s*(\d+)/m);
|
|
28
|
+
return m ? Number(m[1]) : null;
|
|
29
|
+
}
|
|
30
|
+
|
|
31
|
+
/**
|
|
32
|
+
* Whether a hunt result may be accepted as it stands.
|
|
33
|
+
*
|
|
34
|
+
* Held to it only when the result says `done` and the order gave the hunter a way into the app
|
|
35
|
+
* (`launch_cmd` or `app_url`): a hunt that returns `failed` says it could not hunt, which is honest,
|
|
36
|
+
* and a hunt with no way in cannot be asked to drive anything.
|
|
37
|
+
*
|
|
38
|
+
* @param {object} result - The hunt WorkResult.
|
|
39
|
+
* @param {object} payload - The hunt order's payload.
|
|
40
|
+
* @param {(string|null)} report - The hunt report's text, or null.
|
|
41
|
+
* @returns {(string|null)} A reason phrased for the hunter, or null.
|
|
42
|
+
*/
|
|
43
|
+
export function huntProblem(result, payload, report) {
|
|
44
|
+
if (result?.status !== "done") return null;
|
|
45
|
+
if (!payload?.launch_cmd && !payload?.app_url) return null;
|
|
46
|
+
const n = chartersRun(report);
|
|
47
|
+
if (n === null) return "the hunt returned done with no hunt report carrying a `charters:` line";
|
|
48
|
+
if (n > 0) return null;
|
|
49
|
+
return "the hunt returned done over an app its order could reach, having run 0 charters — " +
|
|
50
|
+
"draft at least one charter where the EVAL left territory uncovered and hunt it, or return failed saying why the app could not be driven";
|
|
51
|
+
}
|
|
52
|
+
|
|
53
|
+
/**
|
|
54
|
+
* Read a JSON file, or null.
|
|
55
|
+
* @param {string} p - Path.
|
|
56
|
+
* @returns {(object|null)} The parsed document, or null when absent or unreadable.
|
|
57
|
+
*/
|
|
58
|
+
const readJson = (p) => { try { return JSON.parse(readFileSync(p, "utf8")); } catch { return null; } };
|
|
59
|
+
|
|
60
|
+
/**
|
|
61
|
+
* What the hunt on disk amounts to.
|
|
62
|
+
*
|
|
63
|
+
* @param {string} cwd - Project root.
|
|
64
|
+
* @param {string} slug - Feature slug.
|
|
65
|
+
* @returns {{ok:boolean, status:(string|null), qa:string, charters_run:(number|null), findings:number, reason:(string|null)}}
|
|
66
|
+
* `qa` is `run` (charters ran), `not-hunted` (the hunt returned, no charter ran), or `missing`.
|
|
67
|
+
*/
|
|
68
|
+
export function huntOutcome(cwd, slug) {
|
|
69
|
+
const result = readJson(join(resultsDir(cwd, slug), "hunt.json"));
|
|
70
|
+
if (!result) return { ok: false, status: null, qa: "missing", charters_run: null, findings: 0, reason: "no hunt result on disk" };
|
|
71
|
+
const payload = readJson(join(ordersDir(cwd, slug), "hunt.json"))?.payload || {};
|
|
72
|
+
const reportPath = join(qaDir(cwd, slug), "hunt-report.md");
|
|
73
|
+
const report = existsSync(reportPath) ? readFileSync(reportPath, "utf8") : null;
|
|
74
|
+
const n = chartersRun(report);
|
|
75
|
+
const reason = huntProblem(result, payload, report);
|
|
76
|
+
return {
|
|
77
|
+
ok: !reason,
|
|
78
|
+
status: typeof result.status === "string" ? result.status : null,
|
|
79
|
+
qa: n ? "run" : "not-hunted",
|
|
80
|
+
charters_run: n,
|
|
81
|
+
findings: Array.isArray(result.discoveries) ? result.discoveries.length : 0,
|
|
82
|
+
reason,
|
|
83
|
+
};
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
export const ARGV_SPEC = {
|
|
87
|
+
usage: "harness.mjs probe hunt --slug <slug> [--cwd <dir>]",
|
|
88
|
+
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
89
|
+
slug: { type: "str", required: true },
|
|
90
|
+
cwd: { type: "path" },
|
|
91
|
+
};
|
|
92
|
+
|
|
93
|
+
/**
|
|
94
|
+
* Report what the hunt on disk amounts to.
|
|
95
|
+
*
|
|
96
|
+
* @param {string[]} rawArgv - The subcommand's own arguments (harness.mjs strips the verb words).
|
|
97
|
+
* @returns {void} Exits 0 when the hunt result may be accepted, 1 when it may not.
|
|
98
|
+
*/
|
|
99
|
+
export function cli(rawArgv) {
|
|
100
|
+
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
101
|
+
const out = huntOutcome(resolve(args.cwd || process.cwd()), args.slug);
|
|
102
|
+
console.log(JSON.stringify(out));
|
|
103
|
+
process.exit(out.ok ? 0 : 1);
|
|
104
|
+
}
|
|
@@ -269,12 +269,38 @@ export function renderTable(r) {
|
|
|
269
269
|
return out.join("\n");
|
|
270
270
|
}
|
|
271
271
|
|
|
272
|
+
/**
|
|
273
|
+
* Whether a board that has tasks carries any `covers:` clause while a requirements registry exists.
|
|
274
|
+
*
|
|
275
|
+
* Two bars answer "is this requirement planned": L1b is satisfied by a scope contract's own covers
|
|
276
|
+
* list, while the matrix reads only the acceptance criteria's `(covers: REQ-…)` clauses. A board
|
|
277
|
+
* regenerated with none crossed every gate, and the matrix then read "no evidence" under every
|
|
278
|
+
* requirement of a run whose criteria had passed. Measured on two consecutive runs of one pitch:
|
|
279
|
+
* thirty clauses on one board, zero on the next, the same instruction both times.
|
|
280
|
+
*
|
|
281
|
+
* @param {string} cwd - Project root.
|
|
282
|
+
* @param {string} slug - Feature slug.
|
|
283
|
+
* @returns {(string|null)} A reason phrased for the board's writer, or null when there is no
|
|
284
|
+
* registry, no board yet, or at least one clause.
|
|
285
|
+
*/
|
|
286
|
+
export function boardCoversProblem(cwd, slug) {
|
|
287
|
+
const clauses = parseRequirements(readIf(requirementsFile(cwd, slug)) || "");
|
|
288
|
+
if (!clauses.length) return null;
|
|
289
|
+
const board = readBoard(cwd, slug);
|
|
290
|
+
if (!Array.isArray(board) || !board.length) return null;
|
|
291
|
+
if (coveringAcs(board).size) return null;
|
|
292
|
+
return `the board's ${board.length} tasks carry no \`(covers: REQ-…)\` clause while the registry holds ${clauses.length} ` +
|
|
293
|
+
"requirements — every acceptance criterion that grades a requirement carries its covers clause, or the requirements " +
|
|
294
|
+
"matrix reads no evidence for any of them";
|
|
295
|
+
}
|
|
296
|
+
|
|
272
297
|
export const ARGV_SPEC = {
|
|
273
|
-
usage: "harness.mjs probe requirements --slug <slug> [--run-id <id>] [--format json|table] [--cwd <dir>]",
|
|
298
|
+
usage: "harness.mjs probe requirements --slug <slug> [--run-id <id>] [--format json|table] [--board-check] [--cwd <dir>]",
|
|
274
299
|
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
275
300
|
slug: { type: "str", required: true },
|
|
276
301
|
"run-id": { type: "str" },
|
|
277
302
|
format: { type: "enum", values: ["json", "table"], default: "json" },
|
|
303
|
+
"board-check": { type: "flag" },
|
|
278
304
|
cwd: { type: "path" },
|
|
279
305
|
};
|
|
280
306
|
|
|
@@ -288,6 +314,12 @@ export const ARGV_SPEC = {
|
|
|
288
314
|
export async function cli(rawArgv) {
|
|
289
315
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
290
316
|
const cwd = resolve(args.cwd || process.cwd());
|
|
317
|
+
if (args.boardCheck) {
|
|
318
|
+
// `--board-check` answers one question for the run's send-back: may this board stand?
|
|
319
|
+
const reason = boardCoversProblem(cwd, args.slug);
|
|
320
|
+
console.log(JSON.stringify({ ok: !reason, reason }));
|
|
321
|
+
process.exit(reason ? 1 : 0);
|
|
322
|
+
}
|
|
291
323
|
const report = projectRequirements({ cwd, slug: args.slug, runId: args.runId ?? undefined });
|
|
292
324
|
console.log(args.format === "table" ? renderTable(report) : JSON.stringify(report, null, 2));
|
|
293
325
|
process.exit(0);
|
package/kernel/reduce/ingest.mjs
CHANGED
|
@@ -36,8 +36,9 @@ import { resolve, join, dirname, basename } from "node:path";
|
|
|
36
36
|
import { fileURLToPath } from "node:url";
|
|
37
37
|
import { validate } from "../verify/envelope.mjs";
|
|
38
38
|
import { runArgs } from "../lib/argv.mjs";
|
|
39
|
-
import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId } from "../lib/paths.mjs";
|
|
39
|
+
import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId, qaDir } from "../lib/paths.mjs";
|
|
40
40
|
import { citationProblem, coverageProblem, verdictProblem } from "../probe/eval.mjs";
|
|
41
|
+
import { huntProblem } from "../probe/hunt.mjs";
|
|
41
42
|
|
|
42
43
|
const HERE = dirname(fileURLToPath(import.meta.url));
|
|
43
44
|
const RESULT_SCHEMA = JSON.parse(readFileSync(resolve(HERE, "../schemas/work-result.schema.json"), "utf8"));
|
|
@@ -710,6 +711,24 @@ export async function cli(rawArgv) {
|
|
|
710
711
|
}
|
|
711
712
|
}
|
|
712
713
|
|
|
714
|
+
// --- Hunt gate: a `done` hunt over a reachable app ran at least one charter ---------------------
|
|
715
|
+
// A hunter that reached the app and drafted nothing returned `done`, and a `done` with no findings
|
|
716
|
+
// reads as a clean app. See `huntProblem`; `probe hunt` refuses the same result to the run.
|
|
717
|
+
if (result.worker === "qa-edge-hunter") {
|
|
718
|
+
const reportRel = (Array.isArray(result.artifacts) ? result.artifacts : []).find((a) => /hunt-report\.md$/.test(String(a)));
|
|
719
|
+
const reportAbs = reportRel ? (String(reportRel).startsWith("/") ? reportRel : join(cwd, reportRel)) : null;
|
|
720
|
+
let report = null;
|
|
721
|
+
for (const p of [reportAbs, join(qaDir(cwd, String(result.order_id).split("/")[0]), "hunt-report.md")]) {
|
|
722
|
+
if (!report && p && existsSync(p)) report = readFileSync(p, "utf8");
|
|
723
|
+
}
|
|
724
|
+
const problem = huntProblem(result, order.payload || {}, report);
|
|
725
|
+
if (problem) {
|
|
726
|
+
console.error(`ingest-result: result refused — ${problem}.`);
|
|
727
|
+
console.error(` Re-dispatch the hunter against its order. Nothing was written.`);
|
|
728
|
+
process.exit(1);
|
|
729
|
+
}
|
|
730
|
+
}
|
|
731
|
+
|
|
713
732
|
// Resolved for EVERY order, not only the gated ones: the attesting receipt is this leg's start,
|
|
714
733
|
// and a standalone or `--no-receipt-check` ingest still deserves a truthful timing row rather
|
|
715
734
|
// than one silently falling back to the order's re-writable `compiled_at`.
|
package/kernel/reduce/ship.mjs
CHANGED
|
@@ -39,6 +39,7 @@ import { ratchetReport } from "../probe/stats.mjs";
|
|
|
39
39
|
import { projectRequirements, summaryLine } from "../probe/requirements.mjs";
|
|
40
40
|
import { deriveRounds } from "../probe/rounds.mjs";
|
|
41
41
|
import { collectDiff, scanDiff, summarize } from "./leftovers.mjs";
|
|
42
|
+
import { chartersRun } from "../probe/hunt.mjs";
|
|
42
43
|
|
|
43
44
|
/** @returns {string} Today as `YYYY-MM-DD` (UTC). */
|
|
44
45
|
const today = () => new Date().toISOString().slice(0, 10);
|
|
@@ -392,7 +393,7 @@ export function buildReport(facts) {
|
|
|
392
393
|
*/
|
|
393
394
|
export function qaStatus(passed, huntReport) {
|
|
394
395
|
if (passed === "skipped" || (!passed && !huntReport)) return passed || "skipped";
|
|
395
|
-
if (huntReport &&
|
|
396
|
+
if (huntReport && chartersRun(huntReport) === 0) return "not-hunted";
|
|
396
397
|
return passed || "run";
|
|
397
398
|
}
|
|
398
399
|
|
|
@@ -363,6 +363,7 @@
|
|
|
363
363
|
"launch_cmd",
|
|
364
364
|
"build_gate",
|
|
365
365
|
"revised_checks",
|
|
366
|
+
"scopes_without_t0",
|
|
366
367
|
"t0_artifacts",
|
|
367
368
|
"browser",
|
|
368
369
|
"tasks"
|
|
@@ -2445,6 +2446,11 @@
|
|
|
2445
2446
|
},
|
|
2446
2447
|
"description": "spec-evaluator: rows that FAILed in one of the round's T0 trials and PASS in a later one whose own check file (named by the row id) changed in between. Derived by harness compile. Each must be read against its row before its PASS counts; absent when none."
|
|
2447
2448
|
},
|
|
2449
|
+
"scopes_without_t0": {
|
|
2450
|
+
"type": "array",
|
|
2451
|
+
"items": { "type": "string" },
|
|
2452
|
+
"description": "spec-evaluator: scopes whose contract exists but which hold no green T0 verdict for the round, while at least one other scope does. Derived by harness compile. The round is still gradeable: each row such a scope owns is a FAIL citing the scope contract, never a reason to refuse the round. Absent when every scope is green."
|
|
2453
|
+
},
|
|
2448
2454
|
"build_gate": {
|
|
2449
2455
|
"type": "string",
|
|
2450
2456
|
"description": "spec-evaluator, qa-edge-hunter: this run's newest round build gate artifact (build/r<N>-t<T>.json) — each step's exit code and output tail, the record that the build ran and the app launched. Derived by `harness compile`, absent when the gate never ran. Read it before grading a [ui] row NO EVIDENCE."
|
|
@@ -2871,6 +2877,11 @@
|
|
|
2871
2877
|
"qa_findings": {
|
|
2872
2878
|
"type": "integer"
|
|
2873
2879
|
},
|
|
2880
|
+
"qa": {
|
|
2881
|
+
"type": "string",
|
|
2882
|
+
"enum": ["run", "not-hunted", "skipped"],
|
|
2883
|
+
"description": "What QA amounted to: run (at least one charter ran), not-hunted (the hunt returned and ran no charter), skipped (--no-qa or the QA gate answered skip). qa_findings is 0 in the last two for different reasons."
|
|
2884
|
+
},
|
|
2874
2885
|
"report": {
|
|
2875
2886
|
"type": "string",
|
|
2876
2887
|
"description": "shipped: shapeup/<slug>/REPORT.md path."
|
package/package.json
CHANGED
|
@@ -76,7 +76,9 @@ HARD (any miss → STOP, report which):
|
|
|
76
76
|
✅ deliverable reachable: one real request at `app_url`, or `launch_cmd` run and exit 0, or one
|
|
77
77
|
real invocation of the entry point (not a ping, not a guess that a tool is missing). When the
|
|
78
78
|
order carries `launch_cmd`, run it before concluding anything; the report names the command
|
|
79
|
-
and its exit. A hunt that reached nothing returns `status: failed`, never `done`.
|
|
79
|
+
and its exit. A hunt that reached nothing returns `status: failed`, never `done`. A hunt that
|
|
80
|
+
reached the app runs at least one charter: ingest refuses a `done` whose report says
|
|
81
|
+
`charters: 0/…` while the order carried `launch_cmd` or `app_url`.
|
|
80
82
|
✅ EVAL-FEATURE-<slug>.md exists with verdict: PASS
|
|
81
83
|
✅ if discovery/ledger.md exists: ledger.feature == <feature> (read-only context check —
|
|
82
84
|
a missing ledger is fine; ingest creates it when your findings land)
|
|
@@ -43,6 +43,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
|
|
|
43
43
|
| `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
|
|
44
44
|
| `payload.revised_checks[]` | Rows that FAILed in one trial of this round and PASS in a later one whose own check file changed in between — a pass obtained by editing the check. Read each listed file against its row's Expect before you let its PASS count; a check that no longer asserts what the row asks makes that row a FAIL, and the bug names the check. Absent → no check was rewritten |
|
|
45
45
|
| `payload.t0_artifacts[]` | Per-scope T0 verdict paths for this round (scoped specs), compiled from each scope's green verdict. An artifact listed but missing/red on disk, or a scoped spec with none listed → the round is NOT gradeable: return `status: failed` with the reason, naming the scope, as your FIRST deviation — a structural precondition, not a criterion |
|
|
46
|
+
| `payload.scopes_without_t0[]` | Scopes with no green T0 this round while others have one. The round IS gradeable: grade every row such a scope owns as FAIL, evidence "no green T0 for scope <id>" with the scope contract as its locator (`…/scopes/<id>.md:1`). Never refuse the round for them — a refusal ends the run before the next round or GATE H can act. Absent → every scope is green |
|
|
46
47
|
| `payload.browser` | `cli` (default, ~4x cheaper) \| `mcp` \| `none` |
|
|
47
48
|
| `payload.tasks[]` | Traceability only (which UCs a task claims): NEVER a grading source — the committed UC text is the criterion, a paraphrase mismatch is a finding |
|
|
48
49
|
| `substrate.allowed` | Your only write surface: `.shapeup/<slug>/evaluation/**` (the report + evidence) |
|
|
@@ -38,7 +38,7 @@
|
|
|
38
38
|
// maxParallelScopes (default 4).
|
|
39
39
|
//
|
|
40
40
|
// return — RunReturn (domain.schema.json $defs/RunReturn), the full union:
|
|
41
|
-
// { status: "shipped", verdict, rounds_used, dims_not_evaluated, qa_findings, report }
|
|
41
|
+
// { status: "shipped", verdict, rounds_used, dims_not_evaluated, qa_findings, qa, report }
|
|
42
42
|
// { status: "paused", paused_at, block, valid_decisions, context }
|
|
43
43
|
// { status: "aborted", aborted_at, reason }
|
|
44
44
|
// { status: "gate_h", breaker: "outer"|"attempt_budget"|"none"|"deadline", hammer_proposals, green_scopes,
|
|
@@ -632,6 +632,22 @@ const QA_REPORT = {
|
|
|
632
632
|
required: ["ok", "findings_count"],
|
|
633
633
|
};
|
|
634
634
|
|
|
635
|
+
const BOARD_CHECK = {
|
|
636
|
+
type: "object",
|
|
637
|
+
properties: { ok: { type: "boolean" }, reason: { type: ["string", "null"] } },
|
|
638
|
+
required: ["ok"],
|
|
639
|
+
};
|
|
640
|
+
|
|
641
|
+
const HUNT_OUTCOME = {
|
|
642
|
+
type: "object",
|
|
643
|
+
properties: {
|
|
644
|
+
ok: { type: "boolean" }, status: { type: ["string", "null"] },
|
|
645
|
+
qa: { type: "string", enum: ["run", "not-hunted", "missing"] },
|
|
646
|
+
charters_run: { type: ["integer", "null"] }, findings: { type: "integer" }, reason: { type: ["string", "null"] },
|
|
647
|
+
},
|
|
648
|
+
required: ["ok", "qa"],
|
|
649
|
+
};
|
|
650
|
+
|
|
635
651
|
const HAMMER = {
|
|
636
652
|
type: "object",
|
|
637
653
|
properties: {
|
|
@@ -856,6 +872,58 @@ async function worker({ skill, operation, payload, schema, phase: phaseName, lab
|
|
|
856
872
|
return (r && typeof r === "object") ? r : nullFail(label);
|
|
857
873
|
}
|
|
858
874
|
|
|
875
|
+
/**
|
|
876
|
+
* Dispatch a worker, check its result mechanically, and send it back ONCE with the refusal.
|
|
877
|
+
*
|
|
878
|
+
* A worker's result can be well-formed and still not what the run asked for — a verdict that grades
|
|
879
|
+
* the surface as a group, a hunt that reached the app and ran no charter. The kernel says so; this
|
|
880
|
+
* turns that into one more dispatch carrying the kernel's own reason, and only a second refusal
|
|
881
|
+
* stands. Measured: a rule in the skill made the judge grade row by row on its first dispatch, and
|
|
882
|
+
* the refusal is what holds when the model takes the loose path anyway.
|
|
883
|
+
*
|
|
884
|
+
* @param {function(string): Promise<object>} dispatch - Dispatches the worker; takes a note to append.
|
|
885
|
+
* @param {function(string): Promise<(object|null)>} check - Runs the kernel's check; takes a label suffix.
|
|
886
|
+
* @param {function(object): boolean} refused - Whether a check's answer is a correctable refusal.
|
|
887
|
+
* @param {string} what - What is being checked, for the log line.
|
|
888
|
+
* @returns {Promise<{r: object, v: (object|null)}>} The last dispatch and the last check.
|
|
889
|
+
*/
|
|
890
|
+
async function sendBackOnce(dispatch, check, refused, what) {
|
|
891
|
+
let r = await dispatch("");
|
|
892
|
+
if (r.__failed) return { r, v: null };
|
|
893
|
+
let v = await check("");
|
|
894
|
+
if (v && refused(v)) {
|
|
895
|
+
log(`${what} — refused (${v.reason}); sent back once`);
|
|
896
|
+
r = await dispatch(` Your previous result was refused: ${v.reason}. Correct that and return again.`);
|
|
897
|
+
if (r.__failed) return { r, v: null };
|
|
898
|
+
v = await check(":again");
|
|
899
|
+
}
|
|
900
|
+
return { r, v };
|
|
901
|
+
}
|
|
902
|
+
|
|
903
|
+
/**
|
|
904
|
+
* Dispatch a board writer (analyze, or the board-only regeneration) and send it back once when the
|
|
905
|
+
* board carries no `covers:` clause while a requirements registry exists.
|
|
906
|
+
*
|
|
907
|
+
* Advisory past the second dispatch: the requirements matrix never blocks a ship, so a board that
|
|
908
|
+
* still carries none continues with a state warning naming what the matrix will read, rather than
|
|
909
|
+
* aborting a run whose plan is otherwise sound.
|
|
910
|
+
*
|
|
911
|
+
* @param {function(string): Promise<object>} dispatch - The board writer's dispatch; takes a note.
|
|
912
|
+
* @param {string} what - The operation, for the log and the warning.
|
|
913
|
+
* @returns {Promise<{r: object, v: (object|null)}>} The last dispatch and the last check.
|
|
914
|
+
*/
|
|
915
|
+
async function boardSentBack(dispatch, what) {
|
|
916
|
+
const out = await sendBackOnce(dispatch,
|
|
917
|
+
(suffix) => query(`probe requirements --slug ${slug} --board-check`, BOARD_CHECK, "Analyze", `board-covers${suffix}`),
|
|
918
|
+
(v) => !v.ok && !!v.reason, `ANALYZE ${what}`);
|
|
919
|
+
if (!out.r.__failed && out.v && !out.v.ok && out.v.reason) {
|
|
920
|
+
const msg = `ANALYZE ${what}: ${out.v.reason} (after one send-back; the requirements matrix will read no evidence)`;
|
|
921
|
+
log(msg);
|
|
922
|
+
stateWarnings.push(msg);
|
|
923
|
+
}
|
|
924
|
+
return out;
|
|
925
|
+
}
|
|
926
|
+
|
|
859
927
|
// ---------------------------------------------------------------------------------------------
|
|
860
928
|
// GATES — the kernel's exit-code convention, unchanged: 0 cross · 4 pause · 5 abort. The resolved
|
|
861
929
|
// decision travels in `decision`, copied verbatim out of the kernel's own JSON.
|
|
@@ -1172,7 +1240,7 @@ async function closeIfTerminal(ret) {
|
|
|
1172
1240
|
// A result the single writer never applied is named at the close, not folded into "not green".
|
|
1173
1241
|
+ (Array.isArray(ret.unapplied_results) && ret.unapplied_results.length ? ` unapplied_results=${ret.unapplied_results.length}` : "")
|
|
1174
1242
|
+ (ret.census ? ` census=${ret.census}` : "") + (ret.l4 ? ` l4=${ret.l4}` : "") + (ret.ship_report ? ` ship_report=${ret.ship_report}` : "")
|
|
1175
|
-
: `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}`
|
|
1243
|
+
: `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}` + (ret.qa ? ` qa=${ret.qa}` : "")
|
|
1176
1244
|
+ (ret.after ? ` after=${ret.after} breaker=${ret.breaker ?? "?"} census=${ret.census ?? "?"} cut_list=${Array.isArray(ret.cut_list) ? ret.cut_list.length : "?"}` : "");
|
|
1177
1245
|
// `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
|
|
1178
1246
|
// a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
|
|
@@ -1384,11 +1452,11 @@ if (!rs.has_spec_tree) {
|
|
|
1384
1452
|
// single run since the cutover. It is advisory, so nothing stopped; the ledger simply stayed on
|
|
1385
1453
|
// the previous phase and the snapshot under-reported where the run had got to.
|
|
1386
1454
|
await setRunStatus("mapping", "Analyze");
|
|
1387
|
-
const a = await worker({
|
|
1388
|
-
skill: "ba-pitch-analyzer", operation: "analyze", schema: PHASE_OK, phase: "Analyze", label: "analyze",
|
|
1455
|
+
const { r: a } = await boardSentBack((note) => worker({
|
|
1456
|
+
skill: "ba-pitch-analyzer", operation: "analyze", schema: PHASE_OK, phase: "Analyze", label: note ? "analyze:again" : "analyze",
|
|
1389
1457
|
payload: { pitch: rs.intake_path, breadboard: rs.breadboard_path, spec_folder: specFolder, feature: slug, lens: rs.lens, orient_dir: rs.orient_dir },
|
|
1390
|
-
extra: "Write the spec tree and the board from the orient artifacts — do not re-scan the code.",
|
|
1391
|
-
});
|
|
1458
|
+
extra: "Write the spec tree and the board from the orient artifacts — do not re-scan the code." + note,
|
|
1459
|
+
}), "analyze");
|
|
1392
1460
|
if (a.__failed) return await withWarnings(diedAt("ANALYZE", a));
|
|
1393
1461
|
const post = await requirePhase("ANALYZE", "analyze", "Analyze");
|
|
1394
1462
|
if (post) return await withWarnings(post);
|
|
@@ -1402,13 +1470,13 @@ if (!rs.has_spec_tree) {
|
|
|
1402
1470
|
// operation regenerates it from the tree without re-deriving the tree.
|
|
1403
1471
|
log(`ANALYZE — spec tree on disk, no board: dispatching the board-only operation (slug ${slug})`);
|
|
1404
1472
|
await setRunStatus("mapping", "Analyze");
|
|
1405
|
-
const b = await worker({
|
|
1406
|
-
skill: "ba-pitch-analyzer", operation: "board", schema: PHASE_OK, phase: "Analyze", label: "board",
|
|
1473
|
+
const { r: b } = await boardSentBack((note) => worker({
|
|
1474
|
+
skill: "ba-pitch-analyzer", operation: "board", schema: PHASE_OK, phase: "Analyze", label: note ? "board:again" : "board",
|
|
1407
1475
|
payload: { spec_folder: specFolder, feature: slug, lens: rs.lens },
|
|
1408
1476
|
extra: "The spec tree is committed and FROZEN for this dispatch. Regenerate the per-machine board " +
|
|
1409
1477
|
"under the run's tasks/ directory from the use cases on disk — every acceptance criterion " +
|
|
1410
|
-
"carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder.",
|
|
1411
|
-
});
|
|
1478
|
+
"carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder." + note,
|
|
1479
|
+
}), "board");
|
|
1412
1480
|
if (b.__failed) return await withWarnings(diedAt("ANALYZE", b));
|
|
1413
1481
|
// The board a worker writes carries `depends_on`; its inverse (`unlocks`) and a new task's
|
|
1414
1482
|
// `status: todo` are mechanical, so the kernel writes them rather than trusting every regeneration
|
|
@@ -1872,31 +1940,26 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1872
1940
|
// freezes at GATE L4 has to say that as plainly as the gate block already does.
|
|
1873
1941
|
verdict = "not-evaluated";
|
|
1874
1942
|
} else {
|
|
1875
|
-
const
|
|
1876
|
-
|
|
1877
|
-
|
|
1878
|
-
|
|
1879
|
-
|
|
1880
|
-
|
|
1881
|
-
|
|
1882
|
-
|
|
1883
|
-
|
|
1884
|
-
|
|
1885
|
-
|
|
1886
|
-
|
|
1943
|
+
const { r: e, v: ev } = await sendBackOnce(
|
|
1944
|
+
(note) => worker({
|
|
1945
|
+
skill: "spec-evaluator", operation: "evaluate", schema: EVAL, phase: "Eval",
|
|
1946
|
+
label: note ? `eval:r${round}:again` : `eval:r${round}`, model: evalModel, round,
|
|
1947
|
+
// No `t0_artifacts` here, deliberately: `harness compile` derives them from the round's green
|
|
1948
|
+
// T0 verdicts on disk, for every lane — this script could only name paths it was told about.
|
|
1949
|
+
// The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
|
|
1950
|
+
// artifact and the project profile, so the judge gets the launch evidence without this script
|
|
1951
|
+
// having to carry it.
|
|
1952
|
+
payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
|
|
1953
|
+
extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself." + note,
|
|
1954
|
+
}),
|
|
1955
|
+
// The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
|
|
1956
|
+
// agent's own summary of it (`e.overall`) — see EVAL_VERDICT's comment for why.
|
|
1957
|
+
(suffix) => query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}${suffix}`),
|
|
1958
|
+
// A verdict the kernel refused — grouped rows, a missing citation — is correctable, not a dead judge.
|
|
1959
|
+
(v) => !v.ok && !!v.overall && !!v.reason,
|
|
1960
|
+
`EVAL r${round}`,
|
|
1961
|
+
);
|
|
1887
1962
|
if (e.__failed) return await withWarnings(diedAt("L3", e));
|
|
1888
|
-
// The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
|
|
1889
|
-
// agent's own summary of it (`e.overall`) — see EVAL_VERDICT's comment for why.
|
|
1890
|
-
let ev = await query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}`);
|
|
1891
|
-
// A verdict the kernel refused — a PASS that grades the surface as a group, a missing citation —
|
|
1892
|
-
// is a correctable answer, not a dead judge: the judge is sent back once with the refusal's own
|
|
1893
|
-
// words, and only a second refusal ends the round.
|
|
1894
|
-
if (ev && !ev.ok && ev.overall && ev.reason) {
|
|
1895
|
-
log(`EVAL r${round} — verdict refused (${ev.reason}); the judge is sent back once`);
|
|
1896
|
-
e = await evalOnce(`eval:r${round}:again`, ` Your previous verdict for this round was refused: ${ev.reason}. Grade again and correct that.`);
|
|
1897
|
-
if (e.__failed) return await withWarnings(diedAt("L3", e));
|
|
1898
|
-
ev = await query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}:again`);
|
|
1899
|
-
}
|
|
1900
1963
|
if (!ev) return await withWarnings(diedAt("L3", nullFail(`verdict:r${round}`)));
|
|
1901
1964
|
// A round with no verdict to act on is NOT a dead worker. An evaluator that refused the round
|
|
1902
1965
|
// wrote a result saying why, and `probe eval` carries it as `reason`; reported as "died after
|
|
@@ -1953,15 +2016,26 @@ let qaFindings = 0;
|
|
|
1953
2016
|
const qaG = await crossGate("QA", "QA", ["run", "skip", "ask"], { round, verdict });
|
|
1954
2017
|
if (qaG.stop) return await withWarnings(qaG.stop);
|
|
1955
2018
|
const qaRan = !args.noQa && qaG.decision === "run";
|
|
2019
|
+
// What QA amounted to, from the hunt's own record: `run` only when a charter ran. The QA gate's `run`
|
|
2020
|
+
// is the decision to dispatch, and a dispatched hunt that drove nothing is `not-hunted`.
|
|
2021
|
+
let qaState = "skipped";
|
|
1956
2022
|
if (qaRan) {
|
|
1957
|
-
const q = await
|
|
1958
|
-
|
|
1959
|
-
|
|
1960
|
-
|
|
1961
|
-
|
|
2023
|
+
const { r: q, v: hv } = await sendBackOnce(
|
|
2024
|
+
(note) => worker({
|
|
2025
|
+
skill: "qa-edge-hunter", operation: "hunt", schema: QA_REPORT, phase: "QA", label: note ? "hunt:again" : "hunt", model: qaModel,
|
|
2026
|
+
payload: { feature: slug, spec_folder: specFolder, app_url: rs.app_url, round },
|
|
2027
|
+
extra: "Exploratory hunt over the shipped feature. No verdict and no score — findings only, each with a repro." + note,
|
|
2028
|
+
}),
|
|
2029
|
+
(suffix) => query(`probe hunt --slug ${slug}`, HUNT_OUTCOME, "QA", `hunt-outcome${suffix}`),
|
|
2030
|
+
(v) => !v.ok && v.status === "done" && !!v.reason,
|
|
2031
|
+
"QA hunt",
|
|
2032
|
+
);
|
|
1962
2033
|
// QA is a level-up: losing its worker costs the findings, not the run.
|
|
1963
2034
|
if (q.__failed) log(`QA — the hunt lost its worker: ${q.__failed}. Shipping without QA findings.`);
|
|
1964
|
-
else
|
|
2035
|
+
else {
|
|
2036
|
+
qaFindings = hv && Number.isInteger(hv.findings) ? hv.findings : q.findings_count;
|
|
2037
|
+
qaState = hv?.qa === "run" ? "run" : "not-hunted";
|
|
2038
|
+
}
|
|
1965
2039
|
}
|
|
1966
2040
|
|
|
1967
2041
|
// ---- GATE H — delegated to scope-hammer (census, baseline comparison, cut list) ----------------
|
|
@@ -1986,7 +2060,7 @@ if (h.verdict === "cannot-ship") {
|
|
|
1986
2060
|
// "not-evaluated" as a real verdict (its own usage string, and `generate()`'s no-artifact
|
|
1987
2061
|
// default). Hardcoding PASS here is exactly the silent upgrade protocol.md's Rules forbid.
|
|
1988
2062
|
const shipVerdict = verdict === "not-evaluated" ? "not-evaluated" : "PASS";
|
|
1989
|
-
const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${
|
|
2063
|
+
const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${qaState}`, "Ship", "ship-report");
|
|
1990
2064
|
// GATE L4 HAS A CALL SITE. It used to be a line of prose in the tech lead's references — "resolve
|
|
1991
2065
|
// the gate itself before any of the above" — and a run reached its ship decision with no ledger row
|
|
1992
2066
|
// for it. The resolver narrows `ship` on the census artifact the hammer just wrote.
|
|
@@ -2016,6 +2090,7 @@ return await withWarnings({
|
|
|
2016
2090
|
rounds_used: round,
|
|
2017
2091
|
dims_not_evaluated: ALL_DIMS.filter((d) => !evalDims.includes(d)),
|
|
2018
2092
|
qa_findings: qaFindings,
|
|
2093
|
+
qa: qaState,
|
|
2019
2094
|
report: REPORT_PATH,
|
|
2020
2095
|
});
|
|
2021
2096
|
|