shapeup-sdlc 3.16.0 → 3.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +2 -2
- package/commands/ship.md +4 -2
- package/kernel/gate.mjs +33 -1
- package/kernel/harness.mjs +3 -2
- package/kernel/probe/hunt.mjs +104 -0
- package/kernel/probe/requirements.mjs +33 -1
- package/kernel/reduce/ingest.mjs +20 -1
- package/kernel/reduce/ship.mjs +20 -8
- package/kernel/schemas/domain.schema.json +5 -0
- package/package.json +1 -1
- package/skills/qa-edge-hunter/SKILL.md +3 -1
- package/skills/tech-lead/SKILL.md +4 -4
- package/skills/tech-lead/references/gates.md +3 -1
- package/skills/tech-lead/workflows/shapeup-run.js +122 -43
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.17.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -32,7 +32,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
32
32
|
| Kick-off | ⏸ **L0** — Intake & Config (L0.8 model/budget matrix) + worker roster ✧ | `/translator` if non-English |
|
|
33
33
|
| Orient (Scout) | ⏸ **L1a** — Orient Review | `/orient` |
|
|
34
34
|
| Requirements | — (reviewed at L1b) | `/ba-pitch-analyzer` (`coverage`): the pitch's clauses → committed `requirements.md`, one atomic clause per `REQ-<n>` row naming the pitch clause it came from, ids assigned once and frozen; dispatched once, ahead of Analyze because the acceptance criteria are what cite its ids. A registry already on disk is not re-dispatched |
|
|
35
|
-
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
|
|
35
|
+
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along. A board with no such clause beside a registry sends its writer back once, and one that still has none continues with a warning |
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
38
|
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
@@ -45,7 +45,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
45
45
|
|
|
46
46
|
|
|
47
47
|
### QA Edge Hunt (`/qa-edge-hunter`, post-PASS, pre-ship)
|
|
48
|
-
**Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`.
|
|
48
|
+
**Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`. A hunt that reached the app and ran no charter is sent back once; if it still ran none, the report and the run's close say `not-hunted`, never that QA ran.
|
|
49
49
|
|
|
50
50
|
### Ship & Triage
|
|
51
51
|
- **SHIP S.0 / GATE H** — `/scope-hammer`: census (QA findings + discovered ledger + attempt-budget proposals ✦; every "no scope owns X" cites `probe owner`, which derives ownership from the contracts) → baseline comparison (never the ideal) → cut list; TL/PO promotes selected items only. The census is written as data beside the report, and GATE L4 reads that file — and nothing else — before any answer set may say `ship`. A breaker does not end the run: the run itself dispatches the census, crosses GATE H, writes the ship report with the verdict as it is, crosses L4, and only then closes — `shipped` with the cut list when the census clears it, `escalated` naming the census when it does not.
|
package/commands/ship.md
CHANGED
|
@@ -99,6 +99,8 @@ Additional flags, pass through to `tech-lead` only when the user names them:
|
|
|
99
99
|
- `--rounds N` → override the outer circuit breaker (build+eval cycles, default 3).
|
|
100
100
|
- `--attempts N` → override the inner circuit breaker (per-scope T0 attempts, default 5;
|
|
101
101
|
no-op on specs without scope contracts).
|
|
102
|
-
- `--
|
|
102
|
+
- `--exec-model / --eval-model / --qa-model <name>` → override GATE L0.8's
|
|
103
103
|
resolved model matrix for this run only (highest precedence over `.claude/settings.local.json`
|
|
104
|
-
/ `.claude/settings.json` / skill defaults).
|
|
104
|
+
/ `.claude/settings.json` / skill defaults). `--orch-model` is shown in the L0 block and changes
|
|
105
|
+
nothing: the orchestrator is the session already running this command, so its model is the one
|
|
106
|
+
that session was started with.
|
package/kernel/gate.mjs
CHANGED
|
@@ -54,7 +54,7 @@ import { readFileSync, writeFileSync, appendFileSync, existsSync, mkdirSync } fr
|
|
|
54
54
|
import { parseBoard } from "./reduce/board.mjs";
|
|
55
55
|
import { join, dirname } from "node:path";
|
|
56
56
|
import { runArgs } from "./lib/argv.mjs";
|
|
57
|
-
import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir, tasksDir, hammerCensus, readRunId } from "./lib/paths.mjs";
|
|
57
|
+
import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir, tasksDir, hammerCensus, readRunId, harnessRun } from "./lib/paths.mjs";
|
|
58
58
|
|
|
59
59
|
export const GATE_IDS = ["L0", "L1a", "L1a.5", "L1b", "L2", "L3", "QA", "H", "L4", "COACH-1"];
|
|
60
60
|
|
|
@@ -327,6 +327,34 @@ export function appendGateLedger(cwd, slug, row) {
|
|
|
327
327
|
} catch { return false; }
|
|
328
328
|
}
|
|
329
329
|
|
|
330
|
+
/**
|
|
331
|
+
* This run's earlier ship sign-off, if it already has one.
|
|
332
|
+
*
|
|
333
|
+
* A launched run crosses L4 itself and then closes, and the orchestrator's own closing step resolved
|
|
334
|
+
* it again afterwards — a second `ship` row, after the close, for a decision already taken. Once the
|
|
335
|
+
* run is closed, its decided L4 row is returned instead of resolving again. An open run still
|
|
336
|
+
* re-resolves (the census it reads may have changed), a reopened run has no close, and a paused or
|
|
337
|
+
* aborted row is not a sign-off.
|
|
338
|
+
*
|
|
339
|
+
* @param {string} cwd - Project root.
|
|
340
|
+
* @param {string} slug - Feature slug.
|
|
341
|
+
* @param {string|null} runId - The run the resolve belongs to.
|
|
342
|
+
* @returns {(object|null)} The earlier L4 row, or null.
|
|
343
|
+
*/
|
|
344
|
+
export function priorSignOff(cwd, slug, runId) {
|
|
345
|
+
if (!runId) return null;
|
|
346
|
+
let ledger = "";
|
|
347
|
+
try { ledger = readFileSync(harnessRun(cwd, slug), "utf8"); } catch { return null; }
|
|
348
|
+
if (!/^closed_status:[ \t]*[^\s~]/m.test(ledger)) return null;
|
|
349
|
+
let text = "";
|
|
350
|
+
try { text = readFileSync(gatesPath(cwd, slug), "utf8"); } catch { return null; }
|
|
351
|
+
for (const line of text.split("\n")) {
|
|
352
|
+
let row; try { row = JSON.parse(line); } catch { continue; }
|
|
353
|
+
if (row?.gate === "L4" && row.run_id === runId && row.status === "ok") return row;
|
|
354
|
+
}
|
|
355
|
+
return null;
|
|
356
|
+
}
|
|
357
|
+
|
|
330
358
|
/** Gates this lane will actually hit — used by --verify to catch a set that stalls halfway. */
|
|
331
359
|
export function requiredGates({ autoLevel = "unattended", tiny = false, qa = true } = {}) {
|
|
332
360
|
if (tiny) return ["L0", "L4"];
|
|
@@ -483,6 +511,10 @@ export function cli(rawArgv) {
|
|
|
483
511
|
const gate = args.resolve ?? null;
|
|
484
512
|
if (!gate) die("nothing to do — pass --init, --list, --verify, or --resolve <gate-id>");
|
|
485
513
|
|
|
514
|
+
if (gate === "L4" && args.slug) {
|
|
515
|
+
const prior = priorSignOff(cwd, args.slug, readRunId(cwd, args.slug));
|
|
516
|
+
if (prior) out({ ok: true, gate: "L4", status: prior.status, decision: prior.decision, source: prior.source, already_resolved_at: prior.at }, 0);
|
|
517
|
+
}
|
|
486
518
|
const r = narrowToEvidence(resolve(found.set, gate, found.source), cwd, args.slug ?? null);
|
|
487
519
|
if (r.status === "error") die(r.reason);
|
|
488
520
|
// A gate with no `--slug` (e.g. `--file` used ad hoc, outside any run) has nowhere to file a
|
package/kernel/harness.mjs
CHANGED
|
@@ -32,7 +32,8 @@
|
|
|
32
32
|
// gate An answer file with a source, not a vibe.
|
|
33
33
|
// probe resume · t0 · stats · digest · Read-only queries over run state. `concurrency`
|
|
34
34
|
// concurrency · leg · eval · answers how many legs ran at once and what the
|
|
35
|
-
// owner · requirements · attempts
|
|
35
|
+
// owner · requirements · attempts ·
|
|
36
|
+
// hunt
|
|
36
37
|
// fan-out bought, and refuses a figure the record set
|
|
37
38
|
// cannot support rather than printing a plausible one.
|
|
38
39
|
// `leg` answers whether a scope's work reached the
|
|
@@ -92,7 +93,7 @@ export const ROUTES = {
|
|
|
92
93
|
resume: "./probe/resume.mjs", t0: "./probe/t0.mjs", stats: "./probe/stats.mjs",
|
|
93
94
|
digest: "./probe/digest.mjs", concurrency: "./probe/concurrency.mjs",
|
|
94
95
|
leg: "./probe/leg.mjs", eval: "./probe/eval.mjs", owner: "./probe/owner.mjs",
|
|
95
|
-
requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs",
|
|
96
|
+
requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs", hunt: "./probe/hunt.mjs",
|
|
96
97
|
},
|
|
97
98
|
init: { run: "./init/run.mjs", fit: "./init/fit.mjs", "run-args": "./init/run-args.mjs" },
|
|
98
99
|
report: { export: "./report/export.mjs", _default: "export" },
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// probe hunt — "did the QA hunt hunt anything?"
|
|
3
|
+
//
|
|
4
|
+
// CONTRACT. A bounded, read-only query over the hunt's WorkResult, its order and its report.
|
|
5
|
+
// Prints `{ok, status, qa, charters_run, findings, reason}` on stdout; exits 0 when the hunt's
|
|
6
|
+
// result is one the run may accept, 1 when it is not, and 2 on a bad argv. Writes nothing.
|
|
7
|
+
//
|
|
8
|
+
// WHY THIS EXISTS. A hunter that reached the app, swept the fault log and then drafted no charter
|
|
9
|
+
// still returned `done`, and a `done` with no findings reads as an app with nothing wrong. The hunt
|
|
10
|
+
// report's charter count says what actually happened. A `done` hunt over an app its order could
|
|
11
|
+
// reach, with no charter run, is refused here and on ingest, and the run sends the hunter back once
|
|
12
|
+
// — the same shape as a verdict that grades nothing by name.
|
|
13
|
+
|
|
14
|
+
import { existsSync, readFileSync } from "node:fs";
|
|
15
|
+
import { join, resolve } from "node:path";
|
|
16
|
+
import { runArgs } from "../lib/argv.mjs";
|
|
17
|
+
import { resultsDir, ordersDir, qaDir } from "../lib/paths.mjs";
|
|
18
|
+
|
|
19
|
+
/**
|
|
20
|
+
* The number of charters a hunt report says it ran.
|
|
21
|
+
*
|
|
22
|
+
* @param {(string|null)} text - The hunt report's text.
|
|
23
|
+
* @returns {(number|null)} The run count from its `charters: <run>/<approved>` line, or null when
|
|
24
|
+
* there is no report or no such line.
|
|
25
|
+
*/
|
|
26
|
+
export function chartersRun(text) {
|
|
27
|
+
const m = String(text ?? "").match(/^charters:\s*(\d+)/m);
|
|
28
|
+
return m ? Number(m[1]) : null;
|
|
29
|
+
}
|
|
30
|
+
|
|
31
|
+
/**
|
|
32
|
+
* Whether a hunt result may be accepted as it stands.
|
|
33
|
+
*
|
|
34
|
+
* Held to it only when the result says `done` and the order gave the hunter a way into the app
|
|
35
|
+
* (`launch_cmd` or `app_url`): a hunt that returns `failed` says it could not hunt, which is honest,
|
|
36
|
+
* and a hunt with no way in cannot be asked to drive anything.
|
|
37
|
+
*
|
|
38
|
+
* @param {object} result - The hunt WorkResult.
|
|
39
|
+
* @param {object} payload - The hunt order's payload.
|
|
40
|
+
* @param {(string|null)} report - The hunt report's text, or null.
|
|
41
|
+
* @returns {(string|null)} A reason phrased for the hunter, or null.
|
|
42
|
+
*/
|
|
43
|
+
export function huntProblem(result, payload, report) {
|
|
44
|
+
if (result?.status !== "done") return null;
|
|
45
|
+
if (!payload?.launch_cmd && !payload?.app_url) return null;
|
|
46
|
+
const n = chartersRun(report);
|
|
47
|
+
if (n === null) return "the hunt returned done with no hunt report carrying a `charters:` line";
|
|
48
|
+
if (n > 0) return null;
|
|
49
|
+
return "the hunt returned done over an app its order could reach, having run 0 charters — " +
|
|
50
|
+
"draft at least one charter where the EVAL left territory uncovered and hunt it, or return failed saying why the app could not be driven";
|
|
51
|
+
}
|
|
52
|
+
|
|
53
|
+
/**
|
|
54
|
+
* Read a JSON file, or null.
|
|
55
|
+
* @param {string} p - Path.
|
|
56
|
+
* @returns {(object|null)} The parsed document, or null when absent or unreadable.
|
|
57
|
+
*/
|
|
58
|
+
const readJson = (p) => { try { return JSON.parse(readFileSync(p, "utf8")); } catch { return null; } };
|
|
59
|
+
|
|
60
|
+
/**
|
|
61
|
+
* What the hunt on disk amounts to.
|
|
62
|
+
*
|
|
63
|
+
* @param {string} cwd - Project root.
|
|
64
|
+
* @param {string} slug - Feature slug.
|
|
65
|
+
* @returns {{ok:boolean, status:(string|null), qa:string, charters_run:(number|null), findings:number, reason:(string|null)}}
|
|
66
|
+
* `qa` is `run` (charters ran), `not-hunted` (the hunt returned, no charter ran), or `missing`.
|
|
67
|
+
*/
|
|
68
|
+
export function huntOutcome(cwd, slug) {
|
|
69
|
+
const result = readJson(join(resultsDir(cwd, slug), "hunt.json"));
|
|
70
|
+
if (!result) return { ok: false, status: null, qa: "missing", charters_run: null, findings: 0, reason: "no hunt result on disk" };
|
|
71
|
+
const payload = readJson(join(ordersDir(cwd, slug), "hunt.json"))?.payload || {};
|
|
72
|
+
const reportPath = join(qaDir(cwd, slug), "hunt-report.md");
|
|
73
|
+
const report = existsSync(reportPath) ? readFileSync(reportPath, "utf8") : null;
|
|
74
|
+
const n = chartersRun(report);
|
|
75
|
+
const reason = huntProblem(result, payload, report);
|
|
76
|
+
return {
|
|
77
|
+
ok: !reason,
|
|
78
|
+
status: typeof result.status === "string" ? result.status : null,
|
|
79
|
+
qa: n ? "run" : "not-hunted",
|
|
80
|
+
charters_run: n,
|
|
81
|
+
findings: Array.isArray(result.discoveries) ? result.discoveries.length : 0,
|
|
82
|
+
reason,
|
|
83
|
+
};
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
export const ARGV_SPEC = {
|
|
87
|
+
usage: "harness.mjs probe hunt --slug <slug> [--cwd <dir>]",
|
|
88
|
+
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
89
|
+
slug: { type: "str", required: true },
|
|
90
|
+
cwd: { type: "path" },
|
|
91
|
+
};
|
|
92
|
+
|
|
93
|
+
/**
|
|
94
|
+
* Report what the hunt on disk amounts to.
|
|
95
|
+
*
|
|
96
|
+
* @param {string[]} rawArgv - The subcommand's own arguments (harness.mjs strips the verb words).
|
|
97
|
+
* @returns {void} Exits 0 when the hunt result may be accepted, 1 when it may not.
|
|
98
|
+
*/
|
|
99
|
+
export function cli(rawArgv) {
|
|
100
|
+
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
101
|
+
const out = huntOutcome(resolve(args.cwd || process.cwd()), args.slug);
|
|
102
|
+
console.log(JSON.stringify(out));
|
|
103
|
+
process.exit(out.ok ? 0 : 1);
|
|
104
|
+
}
|
|
@@ -269,12 +269,38 @@ export function renderTable(r) {
|
|
|
269
269
|
return out.join("\n");
|
|
270
270
|
}
|
|
271
271
|
|
|
272
|
+
/**
|
|
273
|
+
* Whether a board that has tasks carries any `covers:` clause while a requirements registry exists.
|
|
274
|
+
*
|
|
275
|
+
* Two bars answer "is this requirement planned": L1b is satisfied by a scope contract's own covers
|
|
276
|
+
* list, while the matrix reads only the acceptance criteria's `(covers: REQ-…)` clauses. A board
|
|
277
|
+
* regenerated with none crossed every gate, and the matrix then read "no evidence" under every
|
|
278
|
+
* requirement of a run whose criteria had passed. Measured on two consecutive runs of one pitch:
|
|
279
|
+
* thirty clauses on one board, zero on the next, the same instruction both times.
|
|
280
|
+
*
|
|
281
|
+
* @param {string} cwd - Project root.
|
|
282
|
+
* @param {string} slug - Feature slug.
|
|
283
|
+
* @returns {(string|null)} A reason phrased for the board's writer, or null when there is no
|
|
284
|
+
* registry, no board yet, or at least one clause.
|
|
285
|
+
*/
|
|
286
|
+
export function boardCoversProblem(cwd, slug) {
|
|
287
|
+
const clauses = parseRequirements(readIf(requirementsFile(cwd, slug)) || "");
|
|
288
|
+
if (!clauses.length) return null;
|
|
289
|
+
const board = readBoard(cwd, slug);
|
|
290
|
+
if (!Array.isArray(board) || !board.length) return null;
|
|
291
|
+
if (coveringAcs(board).size) return null;
|
|
292
|
+
return `the board's ${board.length} tasks carry no \`(covers: REQ-…)\` clause while the registry holds ${clauses.length} ` +
|
|
293
|
+
"requirements — every acceptance criterion that grades a requirement carries its covers clause, or the requirements " +
|
|
294
|
+
"matrix reads no evidence for any of them";
|
|
295
|
+
}
|
|
296
|
+
|
|
272
297
|
export const ARGV_SPEC = {
|
|
273
|
-
usage: "harness.mjs probe requirements --slug <slug> [--run-id <id>] [--format json|table] [--cwd <dir>]",
|
|
298
|
+
usage: "harness.mjs probe requirements --slug <slug> [--run-id <id>] [--format json|table] [--board-check] [--cwd <dir>]",
|
|
274
299
|
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
275
300
|
slug: { type: "str", required: true },
|
|
276
301
|
"run-id": { type: "str" },
|
|
277
302
|
format: { type: "enum", values: ["json", "table"], default: "json" },
|
|
303
|
+
"board-check": { type: "flag" },
|
|
278
304
|
cwd: { type: "path" },
|
|
279
305
|
};
|
|
280
306
|
|
|
@@ -288,6 +314,12 @@ export const ARGV_SPEC = {
|
|
|
288
314
|
export async function cli(rawArgv) {
|
|
289
315
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
290
316
|
const cwd = resolve(args.cwd || process.cwd());
|
|
317
|
+
if (args.boardCheck) {
|
|
318
|
+
// `--board-check` answers one question for the run's send-back: may this board stand?
|
|
319
|
+
const reason = boardCoversProblem(cwd, args.slug);
|
|
320
|
+
console.log(JSON.stringify({ ok: !reason, reason }));
|
|
321
|
+
process.exit(reason ? 1 : 0);
|
|
322
|
+
}
|
|
291
323
|
const report = projectRequirements({ cwd, slug: args.slug, runId: args.runId ?? undefined });
|
|
292
324
|
console.log(args.format === "table" ? renderTable(report) : JSON.stringify(report, null, 2));
|
|
293
325
|
process.exit(0);
|
package/kernel/reduce/ingest.mjs
CHANGED
|
@@ -36,8 +36,9 @@ import { resolve, join, dirname, basename } from "node:path";
|
|
|
36
36
|
import { fileURLToPath } from "node:url";
|
|
37
37
|
import { validate } from "../verify/envelope.mjs";
|
|
38
38
|
import { runArgs } from "../lib/argv.mjs";
|
|
39
|
-
import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId } from "../lib/paths.mjs";
|
|
39
|
+
import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId, qaDir } from "../lib/paths.mjs";
|
|
40
40
|
import { citationProblem, coverageProblem, verdictProblem } from "../probe/eval.mjs";
|
|
41
|
+
import { huntProblem } from "../probe/hunt.mjs";
|
|
41
42
|
|
|
42
43
|
const HERE = dirname(fileURLToPath(import.meta.url));
|
|
43
44
|
const RESULT_SCHEMA = JSON.parse(readFileSync(resolve(HERE, "../schemas/work-result.schema.json"), "utf8"));
|
|
@@ -710,6 +711,24 @@ export async function cli(rawArgv) {
|
|
|
710
711
|
}
|
|
711
712
|
}
|
|
712
713
|
|
|
714
|
+
// --- Hunt gate: a `done` hunt over a reachable app ran at least one charter ---------------------
|
|
715
|
+
// A hunter that reached the app and drafted nothing returned `done`, and a `done` with no findings
|
|
716
|
+
// reads as a clean app. See `huntProblem`; `probe hunt` refuses the same result to the run.
|
|
717
|
+
if (result.worker === "qa-edge-hunter") {
|
|
718
|
+
const reportRel = (Array.isArray(result.artifacts) ? result.artifacts : []).find((a) => /hunt-report\.md$/.test(String(a)));
|
|
719
|
+
const reportAbs = reportRel ? (String(reportRel).startsWith("/") ? reportRel : join(cwd, reportRel)) : null;
|
|
720
|
+
let report = null;
|
|
721
|
+
for (const p of [reportAbs, join(qaDir(cwd, String(result.order_id).split("/")[0]), "hunt-report.md")]) {
|
|
722
|
+
if (!report && p && existsSync(p)) report = readFileSync(p, "utf8");
|
|
723
|
+
}
|
|
724
|
+
const problem = huntProblem(result, order.payload || {}, report);
|
|
725
|
+
if (problem) {
|
|
726
|
+
console.error(`ingest-result: result refused — ${problem}.`);
|
|
727
|
+
console.error(` Re-dispatch the hunter against its order. Nothing was written.`);
|
|
728
|
+
process.exit(1);
|
|
729
|
+
}
|
|
730
|
+
}
|
|
731
|
+
|
|
713
732
|
// Resolved for EVERY order, not only the gated ones: the attesting receipt is this leg's start,
|
|
714
733
|
// and a standalone or `--no-receipt-check` ingest still deserves a truthful timing row rather
|
|
715
734
|
// than one silently falling back to the order's re-writable `compiled_at`.
|
package/kernel/reduce/ship.mjs
CHANGED
|
@@ -32,13 +32,14 @@ import { runArgs } from "../lib/argv.mjs";
|
|
|
32
32
|
import {
|
|
33
33
|
report as reportPath, tasksDir, verdictsDir, trials, evaluationDir, qaDir,
|
|
34
34
|
roundLedger, discoveryLedger, receipt as receiptPath, harnessRun, relShared,
|
|
35
|
-
activeOrder, runArgsPath, readReceipt, runIdFromReceipt, hammerCensus,
|
|
35
|
+
activeOrder, runArgsPath, readReceipt, runIdFromReceipt, hammerCensus, LOCAL,
|
|
36
36
|
} from "../lib/paths.mjs";
|
|
37
37
|
import { readTrials } from "../verify/t0.mjs";
|
|
38
38
|
import { ratchetReport } from "../probe/stats.mjs";
|
|
39
39
|
import { projectRequirements, summaryLine } from "../probe/requirements.mjs";
|
|
40
40
|
import { deriveRounds } from "../probe/rounds.mjs";
|
|
41
41
|
import { collectDiff, scanDiff, summarize } from "./leftovers.mjs";
|
|
42
|
+
import { chartersRun } from "../probe/hunt.mjs";
|
|
42
43
|
|
|
43
44
|
/** @returns {string} Today as `YYYY-MM-DD` (UTC). */
|
|
44
45
|
const today = () => new Date().toISOString().slice(0, 10);
|
|
@@ -112,16 +113,27 @@ export function boardCensus(cwd, slug) {
|
|
|
112
113
|
* An id with no resolvable use case becomes a neutral phrase rather than the id: the report loses a
|
|
113
114
|
* pointer that never resolved off this machine anyway, and keeps the sentence around it.
|
|
114
115
|
*
|
|
116
|
+
* A path into the run trace is the same leak by another spelling: the QA findings section quotes
|
|
117
|
+
* the hunt report, which points at the discovery ledger by its local path, and the next run's
|
|
118
|
+
* spec-lint refused the report for it. The path becomes "the run trace", and the sentence stays.
|
|
119
|
+
*
|
|
115
120
|
* @param {*} text - Any value destined for the committed report.
|
|
116
121
|
* @param {Record<string, string[]>} anchors - Board id → its `use_case_refs`.
|
|
117
|
-
* @returns {string} The text with every `TASK-…` replaced by a stable anchor
|
|
122
|
+
* @returns {string} The text with every `TASK-…` replaced by a stable anchor and every run-trace
|
|
123
|
+
* path by a phrase.
|
|
118
124
|
*/
|
|
119
125
|
export function deboard(text, anchors = {}) {
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
126
|
+
const esc = LOCAL.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
|
127
|
+
return String(text ?? "")
|
|
128
|
+
.replace(/\bTASK-[A-Za-z0-9][\w.-]*/g, (id) => {
|
|
129
|
+
const ucs = anchors[id];
|
|
130
|
+
if (ucs && ucs.length) return ucs.join("/");
|
|
131
|
+
return "a board task";
|
|
132
|
+
})
|
|
133
|
+
.replace(new RegExp(`\`?${esc}/[^\\s\`]+\`?`, "g"), (m) => {
|
|
134
|
+
const tail = (m.replace(/`/g, "").match(/[)\]},.;:'"]+$/) || [""])[0];
|
|
135
|
+
return `the run trace${tail}`;
|
|
136
|
+
});
|
|
125
137
|
}
|
|
126
138
|
|
|
127
139
|
/**
|
|
@@ -381,7 +393,7 @@ export function buildReport(facts) {
|
|
|
381
393
|
*/
|
|
382
394
|
export function qaStatus(passed, huntReport) {
|
|
383
395
|
if (passed === "skipped" || (!passed && !huntReport)) return passed || "skipped";
|
|
384
|
-
if (huntReport &&
|
|
396
|
+
if (huntReport && chartersRun(huntReport) === 0) return "not-hunted";
|
|
385
397
|
return passed || "run";
|
|
386
398
|
}
|
|
387
399
|
|
|
@@ -2871,6 +2871,11 @@
|
|
|
2871
2871
|
"qa_findings": {
|
|
2872
2872
|
"type": "integer"
|
|
2873
2873
|
},
|
|
2874
|
+
"qa": {
|
|
2875
|
+
"type": "string",
|
|
2876
|
+
"enum": ["run", "not-hunted", "skipped"],
|
|
2877
|
+
"description": "What QA amounted to: run (at least one charter ran), not-hunted (the hunt returned and ran no charter), skipped (--no-qa or the QA gate answered skip). qa_findings is 0 in the last two for different reasons."
|
|
2878
|
+
},
|
|
2874
2879
|
"report": {
|
|
2875
2880
|
"type": "string",
|
|
2876
2881
|
"description": "shipped: shapeup/<slug>/REPORT.md path."
|
package/package.json
CHANGED
|
@@ -76,7 +76,9 @@ HARD (any miss → STOP, report which):
|
|
|
76
76
|
✅ deliverable reachable: one real request at `app_url`, or `launch_cmd` run and exit 0, or one
|
|
77
77
|
real invocation of the entry point (not a ping, not a guess that a tool is missing). When the
|
|
78
78
|
order carries `launch_cmd`, run it before concluding anything; the report names the command
|
|
79
|
-
and its exit. A hunt that reached nothing returns `status: failed`, never `done`.
|
|
79
|
+
and its exit. A hunt that reached nothing returns `status: failed`, never `done`. A hunt that
|
|
80
|
+
reached the app runs at least one charter: ingest refuses a `done` whose report says
|
|
81
|
+
`charters: 0/…` while the order carried `launch_cmd` or `app_url`.
|
|
80
82
|
✅ EVAL-FEATURE-<slug>.md exists with verdict: PASS
|
|
81
83
|
✅ if discovery/ledger.md exists: ledger.feature == <feature> (read-only context check —
|
|
82
84
|
a missing ledger is fine; ingest creates it when your findings land)
|
|
@@ -110,11 +110,11 @@ failed, launch from the install path and have the operator `/add-dir` the plugin
|
|
|
110
110
|
did not actually receive from the PO — an unattended lane with no answer for a gate is meant to
|
|
111
111
|
`abort` (see `harness gate`'s `on_missing`), not silently proceed.
|
|
112
112
|
|
|
113
|
-
## Step 4 — GATE L4 — Ship Sign-Off
|
|
113
|
+
## Step 4 — GATE L4 — Ship Sign-Off
|
|
114
114
|
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
115
|
+
A `shipped`/`escalated` return already crossed L4 in the run and froze `REPORT.md`: never resolve it
|
|
116
|
+
again (a row after the close is a second sign-off). A return that did not reach L4 freezes the report
|
|
117
|
+
now and RESOLVES the gate (`references/gates.md` GATE L4). Either way, emit:
|
|
118
118
|
|
|
119
119
|
```
|
|
120
120
|
⏸ GATE L4 — Ship Sign-Off
|
|
@@ -571,7 +571,9 @@ On confirm:
|
|
|
571
571
|
- If the PO provides substantive feedback (not just 'y' or empty) → automatically delegate via Agent (model: exec — see references/protocol.md "Invocation mechanism"): Skill(shapeup-sdlc-plugin:coach) with the provided feedback for RLHF. The coach runs its own GATE COACH-1 to have the PO categorize each rule, then files it under the responsible skill in `shapeup/knowledge-base/<skill>.md` (committed → team-shared). Coachable: `task-executor`, `ba-pitch-analyzer`, `qa-edge-hunter`, `orient`, `scope-architect`, `solution-architect` (each reads its own file at the top of its next run) and `tech-lead` (workflow guidance, read at the next GATE L0). Guidance never decides a gate: a filed rule may add a question or a check to a gate block, never an answer. The tech lead does not categorize the feedback itself — that is the coach's gate, by design (no assumptions).
|
|
572
572
|
- Then output → `✅ [slug] [shipped & deployed | built & verified, deploy pending] — [r] rounds, verdict PASS.`
|
|
573
573
|
|
|
574
|
-
**
|
|
574
|
+
**A launched run already crossed L4 — never resolve it again after a `shipped` or `escalated`
|
|
575
|
+
return; a second row after the close is a second sign-off.** In the prose lane, **resolve the gate
|
|
576
|
+
itself before any of the above** — this is the decision that shipped the run,
|
|
575
577
|
and without it the trace holds no record of that decision at all: `node
|
|
576
578
|
"${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" gate --resolve L4 --slug <slug>
|
|
577
579
|
[--file <path>|--preset <name>]`. Exit 0 (`decision=ship|hold`) — render the block above and close
|
|
@@ -38,7 +38,7 @@
|
|
|
38
38
|
// maxParallelScopes (default 4).
|
|
39
39
|
//
|
|
40
40
|
// return — RunReturn (domain.schema.json $defs/RunReturn), the full union:
|
|
41
|
-
// { status: "shipped", verdict, rounds_used, dims_not_evaluated, qa_findings, report }
|
|
41
|
+
// { status: "shipped", verdict, rounds_used, dims_not_evaluated, qa_findings, qa, report }
|
|
42
42
|
// { status: "paused", paused_at, block, valid_decisions, context }
|
|
43
43
|
// { status: "aborted", aborted_at, reason }
|
|
44
44
|
// { status: "gate_h", breaker: "outer"|"attempt_budget"|"none"|"deadline", hammer_proposals, green_scopes,
|
|
@@ -112,6 +112,10 @@ const argProblems = validateArgs(args);
|
|
|
112
112
|
if (argProblems.length) return { status: "aborted", aborted_at: "args", reason: argProblems.join("; ") };
|
|
113
113
|
|
|
114
114
|
const slug = args.slug;
|
|
115
|
+
// The report's path, never the ship command's one-line `detail`: that is a sub-agent's sentence, and
|
|
116
|
+
// it reached the RunReturn's `report` field in place of the path the orchestrator opens. The
|
|
117
|
+
// committed root is a fixed name, so the path is known without reading anything.
|
|
118
|
+
const REPORT_PATH = `shapeup/${slug}/REPORT.md`;
|
|
115
119
|
const KERNEL = `${args.pluginRoot}/kernel/harness.mjs`;
|
|
116
120
|
const execModel = args.models.exec;
|
|
117
121
|
const evalModel = args.models.eval;
|
|
@@ -628,6 +632,22 @@ const QA_REPORT = {
|
|
|
628
632
|
required: ["ok", "findings_count"],
|
|
629
633
|
};
|
|
630
634
|
|
|
635
|
+
const BOARD_CHECK = {
|
|
636
|
+
type: "object",
|
|
637
|
+
properties: { ok: { type: "boolean" }, reason: { type: ["string", "null"] } },
|
|
638
|
+
required: ["ok"],
|
|
639
|
+
};
|
|
640
|
+
|
|
641
|
+
const HUNT_OUTCOME = {
|
|
642
|
+
type: "object",
|
|
643
|
+
properties: {
|
|
644
|
+
ok: { type: "boolean" }, status: { type: ["string", "null"] },
|
|
645
|
+
qa: { type: "string", enum: ["run", "not-hunted", "missing"] },
|
|
646
|
+
charters_run: { type: ["integer", "null"] }, findings: { type: "integer" }, reason: { type: ["string", "null"] },
|
|
647
|
+
},
|
|
648
|
+
required: ["ok", "qa"],
|
|
649
|
+
};
|
|
650
|
+
|
|
631
651
|
const HAMMER = {
|
|
632
652
|
type: "object",
|
|
633
653
|
properties: {
|
|
@@ -697,7 +717,7 @@ async function settleAtGateH(ret) {
|
|
|
697
717
|
status: "shipped", verdict, rounds_used: round, after: "gate_h", breaker: ret.breaker ?? null,
|
|
698
718
|
census: h.verdict, cut_list: h.cut_list, green_scopes: ret.green_scopes,
|
|
699
719
|
unapplied_results: ret.unapplied_results || [], qa_findings: 0,
|
|
700
|
-
report:
|
|
720
|
+
report: REPORT_PATH,
|
|
701
721
|
});
|
|
702
722
|
}
|
|
703
723
|
|
|
@@ -852,6 +872,58 @@ async function worker({ skill, operation, payload, schema, phase: phaseName, lab
|
|
|
852
872
|
return (r && typeof r === "object") ? r : nullFail(label);
|
|
853
873
|
}
|
|
854
874
|
|
|
875
|
+
/**
|
|
876
|
+
* Dispatch a worker, check its result mechanically, and send it back ONCE with the refusal.
|
|
877
|
+
*
|
|
878
|
+
* A worker's result can be well-formed and still not what the run asked for — a verdict that grades
|
|
879
|
+
* the surface as a group, a hunt that reached the app and ran no charter. The kernel says so; this
|
|
880
|
+
* turns that into one more dispatch carrying the kernel's own reason, and only a second refusal
|
|
881
|
+
* stands. Measured: a rule in the skill made the judge grade row by row on its first dispatch, and
|
|
882
|
+
* the refusal is what holds when the model takes the loose path anyway.
|
|
883
|
+
*
|
|
884
|
+
* @param {function(string): Promise<object>} dispatch - Dispatches the worker; takes a note to append.
|
|
885
|
+
* @param {function(string): Promise<(object|null)>} check - Runs the kernel's check; takes a label suffix.
|
|
886
|
+
* @param {function(object): boolean} refused - Whether a check's answer is a correctable refusal.
|
|
887
|
+
* @param {string} what - What is being checked, for the log line.
|
|
888
|
+
* @returns {Promise<{r: object, v: (object|null)}>} The last dispatch and the last check.
|
|
889
|
+
*/
|
|
890
|
+
async function sendBackOnce(dispatch, check, refused, what) {
|
|
891
|
+
let r = await dispatch("");
|
|
892
|
+
if (r.__failed) return { r, v: null };
|
|
893
|
+
let v = await check("");
|
|
894
|
+
if (v && refused(v)) {
|
|
895
|
+
log(`${what} — refused (${v.reason}); sent back once`);
|
|
896
|
+
r = await dispatch(` Your previous result was refused: ${v.reason}. Correct that and return again.`);
|
|
897
|
+
if (r.__failed) return { r, v: null };
|
|
898
|
+
v = await check(":again");
|
|
899
|
+
}
|
|
900
|
+
return { r, v };
|
|
901
|
+
}
|
|
902
|
+
|
|
903
|
+
/**
|
|
904
|
+
* Dispatch a board writer (analyze, or the board-only regeneration) and send it back once when the
|
|
905
|
+
* board carries no `covers:` clause while a requirements registry exists.
|
|
906
|
+
*
|
|
907
|
+
* Advisory past the second dispatch: the requirements matrix never blocks a ship, so a board that
|
|
908
|
+
* still carries none continues with a state warning naming what the matrix will read, rather than
|
|
909
|
+
* aborting a run whose plan is otherwise sound.
|
|
910
|
+
*
|
|
911
|
+
* @param {function(string): Promise<object>} dispatch - The board writer's dispatch; takes a note.
|
|
912
|
+
* @param {string} what - The operation, for the log and the warning.
|
|
913
|
+
* @returns {Promise<{r: object, v: (object|null)}>} The last dispatch and the last check.
|
|
914
|
+
*/
|
|
915
|
+
async function boardSentBack(dispatch, what) {
|
|
916
|
+
const out = await sendBackOnce(dispatch,
|
|
917
|
+
(suffix) => query(`probe requirements --slug ${slug} --board-check`, BOARD_CHECK, "Analyze", `board-covers${suffix}`),
|
|
918
|
+
(v) => !v.ok && !!v.reason, `ANALYZE ${what}`);
|
|
919
|
+
if (!out.r.__failed && out.v && !out.v.ok && out.v.reason) {
|
|
920
|
+
const msg = `ANALYZE ${what}: ${out.v.reason} (after one send-back; the requirements matrix will read no evidence)`;
|
|
921
|
+
log(msg);
|
|
922
|
+
stateWarnings.push(msg);
|
|
923
|
+
}
|
|
924
|
+
return out;
|
|
925
|
+
}
|
|
926
|
+
|
|
855
927
|
// ---------------------------------------------------------------------------------------------
|
|
856
928
|
// GATES — the kernel's exit-code convention, unchanged: 0 cross · 4 pause · 5 abort. The resolved
|
|
857
929
|
// decision travels in `decision`, copied verbatim out of the kernel's own JSON.
|
|
@@ -1168,7 +1240,7 @@ async function closeIfTerminal(ret) {
|
|
|
1168
1240
|
// A result the single writer never applied is named at the close, not folded into "not green".
|
|
1169
1241
|
+ (Array.isArray(ret.unapplied_results) && ret.unapplied_results.length ? ` unapplied_results=${ret.unapplied_results.length}` : "")
|
|
1170
1242
|
+ (ret.census ? ` census=${ret.census}` : "") + (ret.l4 ? ` l4=${ret.l4}` : "") + (ret.ship_report ? ` ship_report=${ret.ship_report}` : "")
|
|
1171
|
-
: `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}`
|
|
1243
|
+
: `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}` + (ret.qa ? ` qa=${ret.qa}` : "")
|
|
1172
1244
|
+ (ret.after ? ` after=${ret.after} breaker=${ret.breaker ?? "?"} census=${ret.census ?? "?"} cut_list=${Array.isArray(ret.cut_list) ? ret.cut_list.length : "?"}` : "");
|
|
1173
1245
|
// `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
|
|
1174
1246
|
// a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
|
|
@@ -1380,11 +1452,11 @@ if (!rs.has_spec_tree) {
|
|
|
1380
1452
|
// single run since the cutover. It is advisory, so nothing stopped; the ledger simply stayed on
|
|
1381
1453
|
// the previous phase and the snapshot under-reported where the run had got to.
|
|
1382
1454
|
await setRunStatus("mapping", "Analyze");
|
|
1383
|
-
const a = await worker({
|
|
1384
|
-
skill: "ba-pitch-analyzer", operation: "analyze", schema: PHASE_OK, phase: "Analyze", label: "analyze",
|
|
1455
|
+
const { r: a } = await boardSentBack((note) => worker({
|
|
1456
|
+
skill: "ba-pitch-analyzer", operation: "analyze", schema: PHASE_OK, phase: "Analyze", label: note ? "analyze:again" : "analyze",
|
|
1385
1457
|
payload: { pitch: rs.intake_path, breadboard: rs.breadboard_path, spec_folder: specFolder, feature: slug, lens: rs.lens, orient_dir: rs.orient_dir },
|
|
1386
|
-
extra: "Write the spec tree and the board from the orient artifacts — do not re-scan the code.",
|
|
1387
|
-
});
|
|
1458
|
+
extra: "Write the spec tree and the board from the orient artifacts — do not re-scan the code." + note,
|
|
1459
|
+
}), "analyze");
|
|
1388
1460
|
if (a.__failed) return await withWarnings(diedAt("ANALYZE", a));
|
|
1389
1461
|
const post = await requirePhase("ANALYZE", "analyze", "Analyze");
|
|
1390
1462
|
if (post) return await withWarnings(post);
|
|
@@ -1398,13 +1470,13 @@ if (!rs.has_spec_tree) {
|
|
|
1398
1470
|
// operation regenerates it from the tree without re-deriving the tree.
|
|
1399
1471
|
log(`ANALYZE — spec tree on disk, no board: dispatching the board-only operation (slug ${slug})`);
|
|
1400
1472
|
await setRunStatus("mapping", "Analyze");
|
|
1401
|
-
const b = await worker({
|
|
1402
|
-
skill: "ba-pitch-analyzer", operation: "board", schema: PHASE_OK, phase: "Analyze", label: "board",
|
|
1473
|
+
const { r: b } = await boardSentBack((note) => worker({
|
|
1474
|
+
skill: "ba-pitch-analyzer", operation: "board", schema: PHASE_OK, phase: "Analyze", label: note ? "board:again" : "board",
|
|
1403
1475
|
payload: { spec_folder: specFolder, feature: slug, lens: rs.lens },
|
|
1404
1476
|
extra: "The spec tree is committed and FROZEN for this dispatch. Regenerate the per-machine board " +
|
|
1405
1477
|
"under the run's tasks/ directory from the use cases on disk — every acceptance criterion " +
|
|
1406
|
-
"carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder.",
|
|
1407
|
-
});
|
|
1478
|
+
"carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder." + note,
|
|
1479
|
+
}), "board");
|
|
1408
1480
|
if (b.__failed) return await withWarnings(diedAt("ANALYZE", b));
|
|
1409
1481
|
// The board a worker writes carries `depends_on`; its inverse (`unlocks`) and a new task's
|
|
1410
1482
|
// `status: todo` are mechanical, so the kernel writes them rather than trusting every regeneration
|
|
@@ -1868,31 +1940,26 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1868
1940
|
// freezes at GATE L4 has to say that as plainly as the gate block already does.
|
|
1869
1941
|
verdict = "not-evaluated";
|
|
1870
1942
|
} else {
|
|
1871
|
-
const
|
|
1872
|
-
|
|
1873
|
-
|
|
1874
|
-
|
|
1875
|
-
|
|
1876
|
-
|
|
1877
|
-
|
|
1878
|
-
|
|
1879
|
-
|
|
1880
|
-
|
|
1881
|
-
|
|
1882
|
-
|
|
1943
|
+
const { r: e, v: ev } = await sendBackOnce(
|
|
1944
|
+
(note) => worker({
|
|
1945
|
+
skill: "spec-evaluator", operation: "evaluate", schema: EVAL, phase: "Eval",
|
|
1946
|
+
label: note ? `eval:r${round}:again` : `eval:r${round}`, model: evalModel, round,
|
|
1947
|
+
// No `t0_artifacts` here, deliberately: `harness compile` derives them from the round's green
|
|
1948
|
+
// T0 verdicts on disk, for every lane — this script could only name paths it was told about.
|
|
1949
|
+
// The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
|
|
1950
|
+
// artifact and the project profile, so the judge gets the launch evidence without this script
|
|
1951
|
+
// having to carry it.
|
|
1952
|
+
payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
|
|
1953
|
+
extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself." + note,
|
|
1954
|
+
}),
|
|
1955
|
+
// The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
|
|
1956
|
+
// agent's own summary of it (`e.overall`) — see EVAL_VERDICT's comment for why.
|
|
1957
|
+
(suffix) => query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}${suffix}`),
|
|
1958
|
+
// A verdict the kernel refused — grouped rows, a missing citation — is correctable, not a dead judge.
|
|
1959
|
+
(v) => !v.ok && !!v.overall && !!v.reason,
|
|
1960
|
+
`EVAL r${round}`,
|
|
1961
|
+
);
|
|
1883
1962
|
if (e.__failed) return await withWarnings(diedAt("L3", e));
|
|
1884
|
-
// The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
|
|
1885
|
-
// agent's own summary of it (`e.overall`) — see EVAL_VERDICT's comment for why.
|
|
1886
|
-
let ev = await query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}`);
|
|
1887
|
-
// A verdict the kernel refused — a PASS that grades the surface as a group, a missing citation —
|
|
1888
|
-
// is a correctable answer, not a dead judge: the judge is sent back once with the refusal's own
|
|
1889
|
-
// words, and only a second refusal ends the round.
|
|
1890
|
-
if (ev && !ev.ok && ev.overall && ev.reason) {
|
|
1891
|
-
log(`EVAL r${round} — verdict refused (${ev.reason}); the judge is sent back once`);
|
|
1892
|
-
e = await evalOnce(`eval:r${round}:again`, ` Your previous verdict for this round was refused: ${ev.reason}. Grade again and correct that.`);
|
|
1893
|
-
if (e.__failed) return await withWarnings(diedAt("L3", e));
|
|
1894
|
-
ev = await query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}:again`);
|
|
1895
|
-
}
|
|
1896
1963
|
if (!ev) return await withWarnings(diedAt("L3", nullFail(`verdict:r${round}`)));
|
|
1897
1964
|
// A round with no verdict to act on is NOT a dead worker. An evaluator that refused the round
|
|
1898
1965
|
// wrote a result saying why, and `probe eval` carries it as `reason`; reported as "died after
|
|
@@ -1949,15 +2016,26 @@ let qaFindings = 0;
|
|
|
1949
2016
|
const qaG = await crossGate("QA", "QA", ["run", "skip", "ask"], { round, verdict });
|
|
1950
2017
|
if (qaG.stop) return await withWarnings(qaG.stop);
|
|
1951
2018
|
const qaRan = !args.noQa && qaG.decision === "run";
|
|
2019
|
+
// What QA amounted to, from the hunt's own record: `run` only when a charter ran. The QA gate's `run`
|
|
2020
|
+
// is the decision to dispatch, and a dispatched hunt that drove nothing is `not-hunted`.
|
|
2021
|
+
let qaState = "skipped";
|
|
1952
2022
|
if (qaRan) {
|
|
1953
|
-
const q = await
|
|
1954
|
-
|
|
1955
|
-
|
|
1956
|
-
|
|
1957
|
-
|
|
2023
|
+
const { r: q, v: hv } = await sendBackOnce(
|
|
2024
|
+
(note) => worker({
|
|
2025
|
+
skill: "qa-edge-hunter", operation: "hunt", schema: QA_REPORT, phase: "QA", label: note ? "hunt:again" : "hunt", model: qaModel,
|
|
2026
|
+
payload: { feature: slug, spec_folder: specFolder, app_url: rs.app_url, round },
|
|
2027
|
+
extra: "Exploratory hunt over the shipped feature. No verdict and no score — findings only, each with a repro." + note,
|
|
2028
|
+
}),
|
|
2029
|
+
(suffix) => query(`probe hunt --slug ${slug}`, HUNT_OUTCOME, "QA", `hunt-outcome${suffix}`),
|
|
2030
|
+
(v) => !v.ok && v.status === "done" && !!v.reason,
|
|
2031
|
+
"QA hunt",
|
|
2032
|
+
);
|
|
1958
2033
|
// QA is a level-up: losing its worker costs the findings, not the run.
|
|
1959
2034
|
if (q.__failed) log(`QA — the hunt lost its worker: ${q.__failed}. Shipping without QA findings.`);
|
|
1960
|
-
else
|
|
2035
|
+
else {
|
|
2036
|
+
qaFindings = hv && Number.isInteger(hv.findings) ? hv.findings : q.findings_count;
|
|
2037
|
+
qaState = hv?.qa === "run" ? "run" : "not-hunted";
|
|
2038
|
+
}
|
|
1961
2039
|
}
|
|
1962
2040
|
|
|
1963
2041
|
// ---- GATE H — delegated to scope-hammer (census, baseline comparison, cut list) ----------------
|
|
@@ -1982,7 +2060,7 @@ if (h.verdict === "cannot-ship") {
|
|
|
1982
2060
|
// "not-evaluated" as a real verdict (its own usage string, and `generate()`'s no-artifact
|
|
1983
2061
|
// default). Hardcoding PASS here is exactly the silent upgrade protocol.md's Rules forbid.
|
|
1984
2062
|
const shipVerdict = verdict === "not-evaluated" ? "not-evaluated" : "PASS";
|
|
1985
|
-
const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${
|
|
2063
|
+
const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${qaState}`, "Ship", "ship-report");
|
|
1986
2064
|
// GATE L4 HAS A CALL SITE. It used to be a line of prose in the tech lead's references — "resolve
|
|
1987
2065
|
// the gate itself before any of the above" — and a run reached its ship decision with no ledger row
|
|
1988
2066
|
// for it. The resolver narrows `ship` on the census artifact the hammer just wrote.
|
|
@@ -2012,7 +2090,8 @@ return await withWarnings({
|
|
|
2012
2090
|
rounds_used: round,
|
|
2013
2091
|
dims_not_evaluated: ALL_DIMS.filter((d) => !evalDims.includes(d)),
|
|
2014
2092
|
qa_findings: qaFindings,
|
|
2015
|
-
|
|
2093
|
+
qa: qaState,
|
|
2094
|
+
report: REPORT_PATH,
|
|
2016
2095
|
});
|
|
2017
2096
|
|
|
2018
2097
|
// =============================================================================================
|