shapeup-sdlc 3.16.1 → 3.17.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.16.1",
4
+ "version": "3.17.0",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -32,7 +32,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
32
32
  | Kick-off | ⏸ **L0** — Intake & Config (L0.8 model/budget matrix) + worker roster ✧ | `/translator` if non-English |
33
33
  | Orient (Scout) | ⏸ **L1a** — Orient Review | `/orient` |
34
34
  | Requirements | — (reviewed at L1b) | `/ba-pitch-analyzer` (`coverage`): the pitch's clauses → committed `requirements.md`, one atomic clause per `REQ-<n>` row naming the pitch clause it came from, ids assigned once and frozen; dispatched once, ahead of Analyze because the acceptance criteria are what cite its ids. A registry already on disk is not re-dispatched |
35
- | Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
35
+ | Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along. A board with no such clause beside a registry sends its writer back once, and one that still has none continues with a warning |
36
36
  | Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
37
37
  | Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
38
38
  | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
@@ -45,7 +45,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
45
45
 
46
46
 
47
47
  ### QA Edge Hunt (`/qa-edge-hunter`, post-PASS, pre-ship)
48
- **Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`.
48
+ **Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`. A hunt that reached the app and ran no charter is sent back once; if it still ran none, the report and the run's close say `not-hunted`, never that QA ran.
49
49
 
50
50
  ### Ship & Triage
51
51
  - **SHIP S.0 / GATE H** — `/scope-hammer`: census (QA findings + discovered ledger + attempt-budget proposals ✦; every "no scope owns X" cites `probe owner`, which derives ownership from the contracts) → baseline comparison (never the ideal) → cut list; TL/PO promotes selected items only. The census is written as data beside the report, and GATE L4 reads that file — and nothing else — before any answer set may say `ship`. A breaker does not end the run: the run itself dispatches the census, crosses GATE H, writes the ship report with the verdict as it is, crosses L4, and only then closes — `shipped` with the cut list when the census clears it, `escalated` naming the census when it does not.
@@ -32,7 +32,8 @@
32
32
  // gate An answer file with a source, not a vibe.
33
33
  // probe resume · t0 · stats · digest · Read-only queries over run state. `concurrency`
34
34
  // concurrency · leg · eval · answers how many legs ran at once and what the
35
- // owner · requirements · attempts
35
+ // owner · requirements · attempts ·
36
+ // hunt
36
37
  // fan-out bought, and refuses a figure the record set
37
38
  // cannot support rather than printing a plausible one.
38
39
  // `leg` answers whether a scope's work reached the
@@ -92,7 +93,7 @@ export const ROUTES = {
92
93
  resume: "./probe/resume.mjs", t0: "./probe/t0.mjs", stats: "./probe/stats.mjs",
93
94
  digest: "./probe/digest.mjs", concurrency: "./probe/concurrency.mjs",
94
95
  leg: "./probe/leg.mjs", eval: "./probe/eval.mjs", owner: "./probe/owner.mjs",
95
- requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs",
96
+ requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs", hunt: "./probe/hunt.mjs",
96
97
  },
97
98
  init: { run: "./init/run.mjs", fit: "./init/fit.mjs", "run-args": "./init/run-args.mjs" },
98
99
  report: { export: "./report/export.mjs", _default: "export" },
@@ -0,0 +1,104 @@
1
+ #!/usr/bin/env node
2
+ // probe hunt — "did the QA hunt hunt anything?"
3
+ //
4
+ // CONTRACT. A bounded, read-only query over the hunt's WorkResult, its order and its report.
5
+ // Prints `{ok, status, qa, charters_run, findings, reason}` on stdout; exits 0 when the hunt's
6
+ // result is one the run may accept, 1 when it is not, and 2 on a bad argv. Writes nothing.
7
+ //
8
+ // WHY THIS EXISTS. A hunter that reached the app, swept the fault log and then drafted no charter
9
+ // still returned `done`, and a `done` with no findings reads as an app with nothing wrong. The hunt
10
+ // report's charter count says what actually happened. A `done` hunt over an app its order could
11
+ // reach, with no charter run, is refused here and on ingest, and the run sends the hunter back once
12
+ // — the same shape as a verdict that grades nothing by name.
13
+
14
+ import { existsSync, readFileSync } from "node:fs";
15
+ import { join, resolve } from "node:path";
16
+ import { runArgs } from "../lib/argv.mjs";
17
+ import { resultsDir, ordersDir, qaDir } from "../lib/paths.mjs";
18
+
19
+ /**
20
+ * The number of charters a hunt report says it ran.
21
+ *
22
+ * @param {(string|null)} text - The hunt report's text.
23
+ * @returns {(number|null)} The run count from its `charters: <run>/<approved>` line, or null when
24
+ * there is no report or no such line.
25
+ */
26
+ export function chartersRun(text) {
27
+ const m = String(text ?? "").match(/^charters:\s*(\d+)/m);
28
+ return m ? Number(m[1]) : null;
29
+ }
30
+
31
+ /**
32
+ * Whether a hunt result may be accepted as it stands.
33
+ *
34
+ * Held to it only when the result says `done` and the order gave the hunter a way into the app
35
+ * (`launch_cmd` or `app_url`): a hunt that returns `failed` says it could not hunt, which is honest,
36
+ * and a hunt with no way in cannot be asked to drive anything.
37
+ *
38
+ * @param {object} result - The hunt WorkResult.
39
+ * @param {object} payload - The hunt order's payload.
40
+ * @param {(string|null)} report - The hunt report's text, or null.
41
+ * @returns {(string|null)} A reason phrased for the hunter, or null.
42
+ */
43
+ export function huntProblem(result, payload, report) {
44
+ if (result?.status !== "done") return null;
45
+ if (!payload?.launch_cmd && !payload?.app_url) return null;
46
+ const n = chartersRun(report);
47
+ if (n === null) return "the hunt returned done with no hunt report carrying a `charters:` line";
48
+ if (n > 0) return null;
49
+ return "the hunt returned done over an app its order could reach, having run 0 charters — " +
50
+ "draft at least one charter where the EVAL left territory uncovered and hunt it, or return failed saying why the app could not be driven";
51
+ }
52
+
53
+ /**
54
+ * Read a JSON file, or null.
55
+ * @param {string} p - Path.
56
+ * @returns {(object|null)} The parsed document, or null when absent or unreadable.
57
+ */
58
+ const readJson = (p) => { try { return JSON.parse(readFileSync(p, "utf8")); } catch { return null; } };
59
+
60
+ /**
61
+ * What the hunt on disk amounts to.
62
+ *
63
+ * @param {string} cwd - Project root.
64
+ * @param {string} slug - Feature slug.
65
+ * @returns {{ok:boolean, status:(string|null), qa:string, charters_run:(number|null), findings:number, reason:(string|null)}}
66
+ * `qa` is `run` (charters ran), `not-hunted` (the hunt returned, no charter ran), or `missing`.
67
+ */
68
+ export function huntOutcome(cwd, slug) {
69
+ const result = readJson(join(resultsDir(cwd, slug), "hunt.json"));
70
+ if (!result) return { ok: false, status: null, qa: "missing", charters_run: null, findings: 0, reason: "no hunt result on disk" };
71
+ const payload = readJson(join(ordersDir(cwd, slug), "hunt.json"))?.payload || {};
72
+ const reportPath = join(qaDir(cwd, slug), "hunt-report.md");
73
+ const report = existsSync(reportPath) ? readFileSync(reportPath, "utf8") : null;
74
+ const n = chartersRun(report);
75
+ const reason = huntProblem(result, payload, report);
76
+ return {
77
+ ok: !reason,
78
+ status: typeof result.status === "string" ? result.status : null,
79
+ qa: n ? "run" : "not-hunted",
80
+ charters_run: n,
81
+ findings: Array.isArray(result.discoveries) ? result.discoveries.length : 0,
82
+ reason,
83
+ };
84
+ }
85
+
86
+ export const ARGV_SPEC = {
87
+ usage: "harness.mjs probe hunt --slug <slug> [--cwd <dir>]",
88
+ _: { arity: 0, max: 0, name: "(no positional operands)" },
89
+ slug: { type: "str", required: true },
90
+ cwd: { type: "path" },
91
+ };
92
+
93
+ /**
94
+ * Report what the hunt on disk amounts to.
95
+ *
96
+ * @param {string[]} rawArgv - The subcommand's own arguments (harness.mjs strips the verb words).
97
+ * @returns {void} Exits 0 when the hunt result may be accepted, 1 when it may not.
98
+ */
99
+ export function cli(rawArgv) {
100
+ const args = runArgs(ARGV_SPEC, rawArgv);
101
+ const out = huntOutcome(resolve(args.cwd || process.cwd()), args.slug);
102
+ console.log(JSON.stringify(out));
103
+ process.exit(out.ok ? 0 : 1);
104
+ }
@@ -269,12 +269,38 @@ export function renderTable(r) {
269
269
  return out.join("\n");
270
270
  }
271
271
 
272
+ /**
273
+ * Whether a board that has tasks carries any `covers:` clause while a requirements registry exists.
274
+ *
275
+ * Two bars answer "is this requirement planned": L1b is satisfied by a scope contract's own covers
276
+ * list, while the matrix reads only the acceptance criteria's `(covers: REQ-…)` clauses. A board
277
+ * regenerated with none crossed every gate, and the matrix then read "no evidence" under every
278
+ * requirement of a run whose criteria had passed. Measured on two consecutive runs of one pitch:
279
+ * thirty clauses on one board, zero on the next, the same instruction both times.
280
+ *
281
+ * @param {string} cwd - Project root.
282
+ * @param {string} slug - Feature slug.
283
+ * @returns {(string|null)} A reason phrased for the board's writer, or null when there is no
284
+ * registry, no board yet, or at least one clause.
285
+ */
286
+ export function boardCoversProblem(cwd, slug) {
287
+ const clauses = parseRequirements(readIf(requirementsFile(cwd, slug)) || "");
288
+ if (!clauses.length) return null;
289
+ const board = readBoard(cwd, slug);
290
+ if (!Array.isArray(board) || !board.length) return null;
291
+ if (coveringAcs(board).size) return null;
292
+ return `the board's ${board.length} tasks carry no \`(covers: REQ-…)\` clause while the registry holds ${clauses.length} ` +
293
+ "requirements — every acceptance criterion that grades a requirement carries its covers clause, or the requirements " +
294
+ "matrix reads no evidence for any of them";
295
+ }
296
+
272
297
  export const ARGV_SPEC = {
273
- usage: "harness.mjs probe requirements --slug <slug> [--run-id <id>] [--format json|table] [--cwd <dir>]",
298
+ usage: "harness.mjs probe requirements --slug <slug> [--run-id <id>] [--format json|table] [--board-check] [--cwd <dir>]",
274
299
  _: { arity: 0, max: 0, name: "(no positional operands)" },
275
300
  slug: { type: "str", required: true },
276
301
  "run-id": { type: "str" },
277
302
  format: { type: "enum", values: ["json", "table"], default: "json" },
303
+ "board-check": { type: "flag" },
278
304
  cwd: { type: "path" },
279
305
  };
280
306
 
@@ -288,6 +314,12 @@ export const ARGV_SPEC = {
288
314
  export async function cli(rawArgv) {
289
315
  const args = runArgs(ARGV_SPEC, rawArgv);
290
316
  const cwd = resolve(args.cwd || process.cwd());
317
+ if (args.boardCheck) {
318
+ // `--board-check` answers one question for the run's send-back: may this board stand?
319
+ const reason = boardCoversProblem(cwd, args.slug);
320
+ console.log(JSON.stringify({ ok: !reason, reason }));
321
+ process.exit(reason ? 1 : 0);
322
+ }
291
323
  const report = projectRequirements({ cwd, slug: args.slug, runId: args.runId ?? undefined });
292
324
  console.log(args.format === "table" ? renderTable(report) : JSON.stringify(report, null, 2));
293
325
  process.exit(0);
@@ -36,8 +36,9 @@ import { resolve, join, dirname, basename } from "node:path";
36
36
  import { fileURLToPath } from "node:url";
37
37
  import { validate } from "../verify/envelope.mjs";
38
38
  import { runArgs } from "../lib/argv.mjs";
39
- import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId } from "../lib/paths.mjs";
39
+ import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId, qaDir } from "../lib/paths.mjs";
40
40
  import { citationProblem, coverageProblem, verdictProblem } from "../probe/eval.mjs";
41
+ import { huntProblem } from "../probe/hunt.mjs";
41
42
 
42
43
  const HERE = dirname(fileURLToPath(import.meta.url));
43
44
  const RESULT_SCHEMA = JSON.parse(readFileSync(resolve(HERE, "../schemas/work-result.schema.json"), "utf8"));
@@ -710,6 +711,24 @@ export async function cli(rawArgv) {
710
711
  }
711
712
  }
712
713
 
714
+ // --- Hunt gate: a `done` hunt over a reachable app ran at least one charter ---------------------
715
+ // A hunter that reached the app and drafted nothing returned `done`, and a `done` with no findings
716
+ // reads as a clean app. See `huntProblem`; `probe hunt` refuses the same result to the run.
717
+ if (result.worker === "qa-edge-hunter") {
718
+ const reportRel = (Array.isArray(result.artifacts) ? result.artifacts : []).find((a) => /hunt-report\.md$/.test(String(a)));
719
+ const reportAbs = reportRel ? (String(reportRel).startsWith("/") ? reportRel : join(cwd, reportRel)) : null;
720
+ let report = null;
721
+ for (const p of [reportAbs, join(qaDir(cwd, String(result.order_id).split("/")[0]), "hunt-report.md")]) {
722
+ if (!report && p && existsSync(p)) report = readFileSync(p, "utf8");
723
+ }
724
+ const problem = huntProblem(result, order.payload || {}, report);
725
+ if (problem) {
726
+ console.error(`ingest-result: result refused — ${problem}.`);
727
+ console.error(` Re-dispatch the hunter against its order. Nothing was written.`);
728
+ process.exit(1);
729
+ }
730
+ }
731
+
713
732
  // Resolved for EVERY order, not only the gated ones: the attesting receipt is this leg's start,
714
733
  // and a standalone or `--no-receipt-check` ingest still deserves a truthful timing row rather
715
734
  // than one silently falling back to the order's re-writable `compiled_at`.
@@ -39,6 +39,7 @@ import { ratchetReport } from "../probe/stats.mjs";
39
39
  import { projectRequirements, summaryLine } from "../probe/requirements.mjs";
40
40
  import { deriveRounds } from "../probe/rounds.mjs";
41
41
  import { collectDiff, scanDiff, summarize } from "./leftovers.mjs";
42
+ import { chartersRun } from "../probe/hunt.mjs";
42
43
 
43
44
  /** @returns {string} Today as `YYYY-MM-DD` (UTC). */
44
45
  const today = () => new Date().toISOString().slice(0, 10);
@@ -392,7 +393,7 @@ export function buildReport(facts) {
392
393
  */
393
394
  export function qaStatus(passed, huntReport) {
394
395
  if (passed === "skipped" || (!passed && !huntReport)) return passed || "skipped";
395
- if (huntReport && /^charters:\s*0\s*\//m.test(huntReport)) return "not-hunted";
396
+ if (huntReport && chartersRun(huntReport) === 0) return "not-hunted";
396
397
  return passed || "run";
397
398
  }
398
399
 
@@ -2871,6 +2871,11 @@
2871
2871
  "qa_findings": {
2872
2872
  "type": "integer"
2873
2873
  },
2874
+ "qa": {
2875
+ "type": "string",
2876
+ "enum": ["run", "not-hunted", "skipped"],
2877
+ "description": "What QA amounted to: run (at least one charter ran), not-hunted (the hunt returned and ran no charter), skipped (--no-qa or the QA gate answered skip). qa_findings is 0 in the last two for different reasons."
2878
+ },
2874
2879
  "report": {
2875
2880
  "type": "string",
2876
2881
  "description": "shipped: shapeup/<slug>/REPORT.md path."
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.16.1",
3
+ "version": "3.17.0",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -76,7 +76,9 @@ HARD (any miss → STOP, report which):
76
76
  ✅ deliverable reachable: one real request at `app_url`, or `launch_cmd` run and exit 0, or one
77
77
  real invocation of the entry point (not a ping, not a guess that a tool is missing). When the
78
78
  order carries `launch_cmd`, run it before concluding anything; the report names the command
79
- and its exit. A hunt that reached nothing returns `status: failed`, never `done`.
79
+ and its exit. A hunt that reached nothing returns `status: failed`, never `done`. A hunt that
80
+ reached the app runs at least one charter: ingest refuses a `done` whose report says
81
+ `charters: 0/…` while the order carried `launch_cmd` or `app_url`.
80
82
  ✅ EVAL-FEATURE-<slug>.md exists with verdict: PASS
81
83
  ✅ if discovery/ledger.md exists: ledger.feature == <feature> (read-only context check —
82
84
  a missing ledger is fine; ingest creates it when your findings land)
@@ -38,7 +38,7 @@
38
38
  // maxParallelScopes (default 4).
39
39
  //
40
40
  // return — RunReturn (domain.schema.json $defs/RunReturn), the full union:
41
- // { status: "shipped", verdict, rounds_used, dims_not_evaluated, qa_findings, report }
41
+ // { status: "shipped", verdict, rounds_used, dims_not_evaluated, qa_findings, qa, report }
42
42
  // { status: "paused", paused_at, block, valid_decisions, context }
43
43
  // { status: "aborted", aborted_at, reason }
44
44
  // { status: "gate_h", breaker: "outer"|"attempt_budget"|"none"|"deadline", hammer_proposals, green_scopes,
@@ -632,6 +632,22 @@ const QA_REPORT = {
632
632
  required: ["ok", "findings_count"],
633
633
  };
634
634
 
635
+ const BOARD_CHECK = {
636
+ type: "object",
637
+ properties: { ok: { type: "boolean" }, reason: { type: ["string", "null"] } },
638
+ required: ["ok"],
639
+ };
640
+
641
+ const HUNT_OUTCOME = {
642
+ type: "object",
643
+ properties: {
644
+ ok: { type: "boolean" }, status: { type: ["string", "null"] },
645
+ qa: { type: "string", enum: ["run", "not-hunted", "missing"] },
646
+ charters_run: { type: ["integer", "null"] }, findings: { type: "integer" }, reason: { type: ["string", "null"] },
647
+ },
648
+ required: ["ok", "qa"],
649
+ };
650
+
635
651
  const HAMMER = {
636
652
  type: "object",
637
653
  properties: {
@@ -856,6 +872,58 @@ async function worker({ skill, operation, payload, schema, phase: phaseName, lab
856
872
  return (r && typeof r === "object") ? r : nullFail(label);
857
873
  }
858
874
 
875
+ /**
876
+ * Dispatch a worker, check its result mechanically, and send it back ONCE with the refusal.
877
+ *
878
+ * A worker's result can be well-formed and still not what the run asked for — a verdict that grades
879
+ * the surface as a group, a hunt that reached the app and ran no charter. The kernel says so; this
880
+ * turns that into one more dispatch carrying the kernel's own reason, and only a second refusal
881
+ * stands. Measured: a rule in the skill made the judge grade row by row on its first dispatch, and
882
+ * the refusal is what holds when the model takes the loose path anyway.
883
+ *
884
+ * @param {function(string): Promise<object>} dispatch - Dispatches the worker; takes a note to append.
885
+ * @param {function(string): Promise<(object|null)>} check - Runs the kernel's check; takes a label suffix.
886
+ * @param {function(object): boolean} refused - Whether a check's answer is a correctable refusal.
887
+ * @param {string} what - What is being checked, for the log line.
888
+ * @returns {Promise<{r: object, v: (object|null)}>} The last dispatch and the last check.
889
+ */
890
+ async function sendBackOnce(dispatch, check, refused, what) {
891
+ let r = await dispatch("");
892
+ if (r.__failed) return { r, v: null };
893
+ let v = await check("");
894
+ if (v && refused(v)) {
895
+ log(`${what} — refused (${v.reason}); sent back once`);
896
+ r = await dispatch(` Your previous result was refused: ${v.reason}. Correct that and return again.`);
897
+ if (r.__failed) return { r, v: null };
898
+ v = await check(":again");
899
+ }
900
+ return { r, v };
901
+ }
902
+
903
+ /**
904
+ * Dispatch a board writer (analyze, or the board-only regeneration) and send it back once when the
905
+ * board carries no `covers:` clause while a requirements registry exists.
906
+ *
907
+ * Advisory past the second dispatch: the requirements matrix never blocks a ship, so a board that
908
+ * still carries none continues with a state warning naming what the matrix will read, rather than
909
+ * aborting a run whose plan is otherwise sound.
910
+ *
911
+ * @param {function(string): Promise<object>} dispatch - The board writer's dispatch; takes a note.
912
+ * @param {string} what - The operation, for the log and the warning.
913
+ * @returns {Promise<{r: object, v: (object|null)}>} The last dispatch and the last check.
914
+ */
915
+ async function boardSentBack(dispatch, what) {
916
+ const out = await sendBackOnce(dispatch,
917
+ (suffix) => query(`probe requirements --slug ${slug} --board-check`, BOARD_CHECK, "Analyze", `board-covers${suffix}`),
918
+ (v) => !v.ok && !!v.reason, `ANALYZE ${what}`);
919
+ if (!out.r.__failed && out.v && !out.v.ok && out.v.reason) {
920
+ const msg = `ANALYZE ${what}: ${out.v.reason} (after one send-back; the requirements matrix will read no evidence)`;
921
+ log(msg);
922
+ stateWarnings.push(msg);
923
+ }
924
+ return out;
925
+ }
926
+
859
927
  // ---------------------------------------------------------------------------------------------
860
928
  // GATES — the kernel's exit-code convention, unchanged: 0 cross · 4 pause · 5 abort. The resolved
861
929
  // decision travels in `decision`, copied verbatim out of the kernel's own JSON.
@@ -1172,7 +1240,7 @@ async function closeIfTerminal(ret) {
1172
1240
  // A result the single writer never applied is named at the close, not folded into "not green".
1173
1241
  + (Array.isArray(ret.unapplied_results) && ret.unapplied_results.length ? ` unapplied_results=${ret.unapplied_results.length}` : "")
1174
1242
  + (ret.census ? ` census=${ret.census}` : "") + (ret.l4 ? ` l4=${ret.l4}` : "") + (ret.ship_report ? ` ship_report=${ret.ship_report}` : "")
1175
- : `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}`
1243
+ : `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}` + (ret.qa ? ` qa=${ret.qa}` : "")
1176
1244
  + (ret.after ? ` after=${ret.after} breaker=${ret.breaker ?? "?"} census=${ret.census ?? "?"} cut_list=${Array.isArray(ret.cut_list) ? ret.cut_list.length : "?"}` : "");
1177
1245
  // `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
1178
1246
  // a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
@@ -1384,11 +1452,11 @@ if (!rs.has_spec_tree) {
1384
1452
  // single run since the cutover. It is advisory, so nothing stopped; the ledger simply stayed on
1385
1453
  // the previous phase and the snapshot under-reported where the run had got to.
1386
1454
  await setRunStatus("mapping", "Analyze");
1387
- const a = await worker({
1388
- skill: "ba-pitch-analyzer", operation: "analyze", schema: PHASE_OK, phase: "Analyze", label: "analyze",
1455
+ const { r: a } = await boardSentBack((note) => worker({
1456
+ skill: "ba-pitch-analyzer", operation: "analyze", schema: PHASE_OK, phase: "Analyze", label: note ? "analyze:again" : "analyze",
1389
1457
  payload: { pitch: rs.intake_path, breadboard: rs.breadboard_path, spec_folder: specFolder, feature: slug, lens: rs.lens, orient_dir: rs.orient_dir },
1390
- extra: "Write the spec tree and the board from the orient artifacts — do not re-scan the code.",
1391
- });
1458
+ extra: "Write the spec tree and the board from the orient artifacts — do not re-scan the code." + note,
1459
+ }), "analyze");
1392
1460
  if (a.__failed) return await withWarnings(diedAt("ANALYZE", a));
1393
1461
  const post = await requirePhase("ANALYZE", "analyze", "Analyze");
1394
1462
  if (post) return await withWarnings(post);
@@ -1402,13 +1470,13 @@ if (!rs.has_spec_tree) {
1402
1470
  // operation regenerates it from the tree without re-deriving the tree.
1403
1471
  log(`ANALYZE — spec tree on disk, no board: dispatching the board-only operation (slug ${slug})`);
1404
1472
  await setRunStatus("mapping", "Analyze");
1405
- const b = await worker({
1406
- skill: "ba-pitch-analyzer", operation: "board", schema: PHASE_OK, phase: "Analyze", label: "board",
1473
+ const { r: b } = await boardSentBack((note) => worker({
1474
+ skill: "ba-pitch-analyzer", operation: "board", schema: PHASE_OK, phase: "Analyze", label: note ? "board:again" : "board",
1407
1475
  payload: { spec_folder: specFolder, feature: slug, lens: rs.lens },
1408
1476
  extra: "The spec tree is committed and FROZEN for this dispatch. Regenerate the per-machine board " +
1409
1477
  "under the run's tasks/ directory from the use cases on disk — every acceptance criterion " +
1410
- "carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder.",
1411
- });
1478
+ "carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder." + note,
1479
+ }), "board");
1412
1480
  if (b.__failed) return await withWarnings(diedAt("ANALYZE", b));
1413
1481
  // The board a worker writes carries `depends_on`; its inverse (`unlocks`) and a new task's
1414
1482
  // `status: todo` are mechanical, so the kernel writes them rather than trusting every regeneration
@@ -1872,31 +1940,26 @@ while (verdict !== "pass" && round <= maxRounds) {
1872
1940
  // freezes at GATE L4 has to say that as plainly as the gate block already does.
1873
1941
  verdict = "not-evaluated";
1874
1942
  } else {
1875
- const evalOnce = (label, extraNote = "") => worker({
1876
- skill: "spec-evaluator", operation: "evaluate", schema: EVAL, phase: "Eval", label,
1877
- model: evalModel, round,
1878
- // No `t0_artifacts` here, deliberately: `harness compile` derives them from the round's green
1879
- // T0 verdicts on disk, for every lane — this script could only name paths it was told about.
1880
- // The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
1881
- // artifact and the project profile, so the judge gets the launch evidence without this script
1882
- // having to carry it.
1883
- payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
1884
- extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself." + extraNote,
1885
- });
1886
- let e = await evalOnce(`eval:r${round}`);
1943
+ const { r: e, v: ev } = await sendBackOnce(
1944
+ (note) => worker({
1945
+ skill: "spec-evaluator", operation: "evaluate", schema: EVAL, phase: "Eval",
1946
+ label: note ? `eval:r${round}:again` : `eval:r${round}`, model: evalModel, round,
1947
+ // No `t0_artifacts` here, deliberately: `harness compile` derives them from the round's green
1948
+ // T0 verdicts on disk, for every lane — this script could only name paths it was told about.
1949
+ // The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
1950
+ // artifact and the project profile, so the judge gets the launch evidence without this script
1951
+ // having to carry it.
1952
+ payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
1953
+ extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself." + note,
1954
+ }),
1955
+ // The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
1956
+ // agent's own summary of it (`e.overall`) — see EVAL_VERDICT's comment for why.
1957
+ (suffix) => query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}${suffix}`),
1958
+ // A verdict the kernel refused — grouped rows, a missing citation — is correctable, not a dead judge.
1959
+ (v) => !v.ok && !!v.overall && !!v.reason,
1960
+ `EVAL r${round}`,
1961
+ );
1887
1962
  if (e.__failed) return await withWarnings(diedAt("L3", e));
1888
- // The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
1889
- // agent's own summary of it (`e.overall`) — see EVAL_VERDICT's comment for why.
1890
- let ev = await query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}`);
1891
- // A verdict the kernel refused — a PASS that grades the surface as a group, a missing citation —
1892
- // is a correctable answer, not a dead judge: the judge is sent back once with the refusal's own
1893
- // words, and only a second refusal ends the round.
1894
- if (ev && !ev.ok && ev.overall && ev.reason) {
1895
- log(`EVAL r${round} — verdict refused (${ev.reason}); the judge is sent back once`);
1896
- e = await evalOnce(`eval:r${round}:again`, ` Your previous verdict for this round was refused: ${ev.reason}. Grade again and correct that.`);
1897
- if (e.__failed) return await withWarnings(diedAt("L3", e));
1898
- ev = await query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}:again`);
1899
- }
1900
1963
  if (!ev) return await withWarnings(diedAt("L3", nullFail(`verdict:r${round}`)));
1901
1964
  // A round with no verdict to act on is NOT a dead worker. An evaluator that refused the round
1902
1965
  // wrote a result saying why, and `probe eval` carries it as `reason`; reported as "died after
@@ -1953,15 +2016,26 @@ let qaFindings = 0;
1953
2016
  const qaG = await crossGate("QA", "QA", ["run", "skip", "ask"], { round, verdict });
1954
2017
  if (qaG.stop) return await withWarnings(qaG.stop);
1955
2018
  const qaRan = !args.noQa && qaG.decision === "run";
2019
+ // What QA amounted to, from the hunt's own record: `run` only when a charter ran. The QA gate's `run`
2020
+ // is the decision to dispatch, and a dispatched hunt that drove nothing is `not-hunted`.
2021
+ let qaState = "skipped";
1956
2022
  if (qaRan) {
1957
- const q = await worker({
1958
- skill: "qa-edge-hunter", operation: "hunt", schema: QA_REPORT, phase: "QA", label: "hunt", model: qaModel,
1959
- payload: { feature: slug, spec_folder: specFolder, app_url: rs.app_url, round },
1960
- extra: "Exploratory hunt over the shipped feature. No verdict and no score — findings only, each with a repro.",
1961
- });
2023
+ const { r: q, v: hv } = await sendBackOnce(
2024
+ (note) => worker({
2025
+ skill: "qa-edge-hunter", operation: "hunt", schema: QA_REPORT, phase: "QA", label: note ? "hunt:again" : "hunt", model: qaModel,
2026
+ payload: { feature: slug, spec_folder: specFolder, app_url: rs.app_url, round },
2027
+ extra: "Exploratory hunt over the shipped feature. No verdict and no score — findings only, each with a repro." + note,
2028
+ }),
2029
+ (suffix) => query(`probe hunt --slug ${slug}`, HUNT_OUTCOME, "QA", `hunt-outcome${suffix}`),
2030
+ (v) => !v.ok && v.status === "done" && !!v.reason,
2031
+ "QA hunt",
2032
+ );
1962
2033
  // QA is a level-up: losing its worker costs the findings, not the run.
1963
2034
  if (q.__failed) log(`QA — the hunt lost its worker: ${q.__failed}. Shipping without QA findings.`);
1964
- else qaFindings = q.findings_count;
2035
+ else {
2036
+ qaFindings = hv && Number.isInteger(hv.findings) ? hv.findings : q.findings_count;
2037
+ qaState = hv?.qa === "run" ? "run" : "not-hunted";
2038
+ }
1965
2039
  }
1966
2040
 
1967
2041
  // ---- GATE H — delegated to scope-hammer (census, baseline comparison, cut list) ----------------
@@ -1986,7 +2060,7 @@ if (h.verdict === "cannot-ship") {
1986
2060
  // "not-evaluated" as a real verdict (its own usage string, and `generate()`'s no-artifact
1987
2061
  // default). Hardcoding PASS here is exactly the silent upgrade protocol.md's Rules forbid.
1988
2062
  const shipVerdict = verdict === "not-evaluated" ? "not-evaluated" : "PASS";
1989
- const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${qaRan ? "run" : "skipped"}`, "Ship", "ship-report");
2063
+ const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${qaState}`, "Ship", "ship-report");
1990
2064
  // GATE L4 HAS A CALL SITE. It used to be a line of prose in the tech lead's references — "resolve
1991
2065
  // the gate itself before any of the above" — and a run reached its ship decision with no ledger row
1992
2066
  // for it. The resolver narrows `ship` on the census artifact the hammer just wrote.
@@ -2016,6 +2090,7 @@ return await withWarnings({
2016
2090
  rounds_used: round,
2017
2091
  dims_not_evaluated: ALL_DIMS.filter((d) => !evalDims.includes(d)),
2018
2092
  qa_findings: qaFindings,
2093
+ qa: qaState,
2019
2094
  report: REPORT_PATH,
2020
2095
  });
2021
2096