shapeup-sdlc 3.16.1 → 3.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +2 -2
- package/kernel/harness.mjs +3 -2
- package/kernel/probe/hunt.mjs +104 -0
- package/kernel/probe/requirements.mjs +33 -1
- package/kernel/reduce/ingest.mjs +20 -1
- package/kernel/reduce/ship.mjs +2 -1
- package/kernel/schemas/domain.schema.json +5 -0
- package/package.json +1 -1
- package/skills/qa-edge-hunter/SKILL.md +3 -1
- package/skills/tech-lead/workflows/shapeup-run.js +116 -41
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.17.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -32,7 +32,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
32
32
|
| Kick-off | ⏸ **L0** — Intake & Config (L0.8 model/budget matrix) + worker roster ✧ | `/translator` if non-English |
|
|
33
33
|
| Orient (Scout) | ⏸ **L1a** — Orient Review | `/orient` |
|
|
34
34
|
| Requirements | — (reviewed at L1b) | `/ba-pitch-analyzer` (`coverage`): the pitch's clauses → committed `requirements.md`, one atomic clause per `REQ-<n>` row naming the pitch clause it came from, ids assigned once and frozen; dispatched once, ahead of Analyze because the acceptance criteria are what cite its ids. A registry already on disk is not re-dispatched |
|
|
35
|
-
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
|
|
35
|
+
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along. A board with no such clause beside a registry sends its writer back once, and one that still has none continues with a warning |
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
38
|
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
@@ -45,7 +45,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
45
45
|
|
|
46
46
|
|
|
47
47
|
### QA Edge Hunt (`/qa-edge-hunter`, post-PASS, pre-ship)
|
|
48
|
-
**Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`.
|
|
48
|
+
**Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`. A hunt that reached the app and ran no charter is sent back once; if it still ran none, the report and the run's close say `not-hunted`, never that QA ran.
|
|
49
49
|
|
|
50
50
|
### Ship & Triage
|
|
51
51
|
- **SHIP S.0 / GATE H** — `/scope-hammer`: census (QA findings + discovered ledger + attempt-budget proposals ✦; every "no scope owns X" cites `probe owner`, which derives ownership from the contracts) → baseline comparison (never the ideal) → cut list; TL/PO promotes selected items only. The census is written as data beside the report, and GATE L4 reads that file — and nothing else — before any answer set may say `ship`. A breaker does not end the run: the run itself dispatches the census, crosses GATE H, writes the ship report with the verdict as it is, crosses L4, and only then closes — `shipped` with the cut list when the census clears it, `escalated` naming the census when it does not.
|
package/kernel/harness.mjs
CHANGED
|
@@ -32,7 +32,8 @@
|
|
|
32
32
|
// gate An answer file with a source, not a vibe.
|
|
33
33
|
// probe resume · t0 · stats · digest · Read-only queries over run state. `concurrency`
|
|
34
34
|
// concurrency · leg · eval · answers how many legs ran at once and what the
|
|
35
|
-
// owner · requirements · attempts
|
|
35
|
+
// owner · requirements · attempts ·
|
|
36
|
+
// hunt
|
|
36
37
|
// fan-out bought, and refuses a figure the record set
|
|
37
38
|
// cannot support rather than printing a plausible one.
|
|
38
39
|
// `leg` answers whether a scope's work reached the
|
|
@@ -92,7 +93,7 @@ export const ROUTES = {
|
|
|
92
93
|
resume: "./probe/resume.mjs", t0: "./probe/t0.mjs", stats: "./probe/stats.mjs",
|
|
93
94
|
digest: "./probe/digest.mjs", concurrency: "./probe/concurrency.mjs",
|
|
94
95
|
leg: "./probe/leg.mjs", eval: "./probe/eval.mjs", owner: "./probe/owner.mjs",
|
|
95
|
-
requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs",
|
|
96
|
+
requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs", hunt: "./probe/hunt.mjs",
|
|
96
97
|
},
|
|
97
98
|
init: { run: "./init/run.mjs", fit: "./init/fit.mjs", "run-args": "./init/run-args.mjs" },
|
|
98
99
|
report: { export: "./report/export.mjs", _default: "export" },
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// probe hunt — "did the QA hunt hunt anything?"
|
|
3
|
+
//
|
|
4
|
+
// CONTRACT. A bounded, read-only query over the hunt's WorkResult, its order and its report.
|
|
5
|
+
// Prints `{ok, status, qa, charters_run, findings, reason}` on stdout; exits 0 when the hunt's
|
|
6
|
+
// result is one the run may accept, 1 when it is not, and 2 on a bad argv. Writes nothing.
|
|
7
|
+
//
|
|
8
|
+
// WHY THIS EXISTS. A hunter that reached the app, swept the fault log and then drafted no charter
|
|
9
|
+
// still returned `done`, and a `done` with no findings reads as an app with nothing wrong. The hunt
|
|
10
|
+
// report's charter count says what actually happened. A `done` hunt over an app its order could
|
|
11
|
+
// reach, with no charter run, is refused here and on ingest, and the run sends the hunter back once
|
|
12
|
+
// — the same shape as a verdict that grades nothing by name.
|
|
13
|
+
|
|
14
|
+
import { existsSync, readFileSync } from "node:fs";
|
|
15
|
+
import { join, resolve } from "node:path";
|
|
16
|
+
import { runArgs } from "../lib/argv.mjs";
|
|
17
|
+
import { resultsDir, ordersDir, qaDir } from "../lib/paths.mjs";
|
|
18
|
+
|
|
19
|
+
/**
|
|
20
|
+
* The number of charters a hunt report says it ran.
|
|
21
|
+
*
|
|
22
|
+
* @param {(string|null)} text - The hunt report's text.
|
|
23
|
+
* @returns {(number|null)} The run count from its `charters: <run>/<approved>` line, or null when
|
|
24
|
+
* there is no report or no such line.
|
|
25
|
+
*/
|
|
26
|
+
export function chartersRun(text) {
|
|
27
|
+
const m = String(text ?? "").match(/^charters:\s*(\d+)/m);
|
|
28
|
+
return m ? Number(m[1]) : null;
|
|
29
|
+
}
|
|
30
|
+
|
|
31
|
+
/**
|
|
32
|
+
* Whether a hunt result may be accepted as it stands.
|
|
33
|
+
*
|
|
34
|
+
* Held to it only when the result says `done` and the order gave the hunter a way into the app
|
|
35
|
+
* (`launch_cmd` or `app_url`): a hunt that returns `failed` says it could not hunt, which is honest,
|
|
36
|
+
* and a hunt with no way in cannot be asked to drive anything.
|
|
37
|
+
*
|
|
38
|
+
* @param {object} result - The hunt WorkResult.
|
|
39
|
+
* @param {object} payload - The hunt order's payload.
|
|
40
|
+
* @param {(string|null)} report - The hunt report's text, or null.
|
|
41
|
+
* @returns {(string|null)} A reason phrased for the hunter, or null.
|
|
42
|
+
*/
|
|
43
|
+
export function huntProblem(result, payload, report) {
|
|
44
|
+
if (result?.status !== "done") return null;
|
|
45
|
+
if (!payload?.launch_cmd && !payload?.app_url) return null;
|
|
46
|
+
const n = chartersRun(report);
|
|
47
|
+
if (n === null) return "the hunt returned done with no hunt report carrying a `charters:` line";
|
|
48
|
+
if (n > 0) return null;
|
|
49
|
+
return "the hunt returned done over an app its order could reach, having run 0 charters — " +
|
|
50
|
+
"draft at least one charter where the EVAL left territory uncovered and hunt it, or return failed saying why the app could not be driven";
|
|
51
|
+
}
|
|
52
|
+
|
|
53
|
+
/**
|
|
54
|
+
* Read a JSON file, or null.
|
|
55
|
+
* @param {string} p - Path.
|
|
56
|
+
* @returns {(object|null)} The parsed document, or null when absent or unreadable.
|
|
57
|
+
*/
|
|
58
|
+
const readJson = (p) => { try { return JSON.parse(readFileSync(p, "utf8")); } catch { return null; } };
|
|
59
|
+
|
|
60
|
+
/**
|
|
61
|
+
* What the hunt on disk amounts to.
|
|
62
|
+
*
|
|
63
|
+
* @param {string} cwd - Project root.
|
|
64
|
+
* @param {string} slug - Feature slug.
|
|
65
|
+
* @returns {{ok:boolean, status:(string|null), qa:string, charters_run:(number|null), findings:number, reason:(string|null)}}
|
|
66
|
+
* `qa` is `run` (charters ran), `not-hunted` (the hunt returned, no charter ran), or `missing`.
|
|
67
|
+
*/
|
|
68
|
+
export function huntOutcome(cwd, slug) {
|
|
69
|
+
const result = readJson(join(resultsDir(cwd, slug), "hunt.json"));
|
|
70
|
+
if (!result) return { ok: false, status: null, qa: "missing", charters_run: null, findings: 0, reason: "no hunt result on disk" };
|
|
71
|
+
const payload = readJson(join(ordersDir(cwd, slug), "hunt.json"))?.payload || {};
|
|
72
|
+
const reportPath = join(qaDir(cwd, slug), "hunt-report.md");
|
|
73
|
+
const report = existsSync(reportPath) ? readFileSync(reportPath, "utf8") : null;
|
|
74
|
+
const n = chartersRun(report);
|
|
75
|
+
const reason = huntProblem(result, payload, report);
|
|
76
|
+
return {
|
|
77
|
+
ok: !reason,
|
|
78
|
+
status: typeof result.status === "string" ? result.status : null,
|
|
79
|
+
qa: n ? "run" : "not-hunted",
|
|
80
|
+
charters_run: n,
|
|
81
|
+
findings: Array.isArray(result.discoveries) ? result.discoveries.length : 0,
|
|
82
|
+
reason,
|
|
83
|
+
};
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
export const ARGV_SPEC = {
|
|
87
|
+
usage: "harness.mjs probe hunt --slug <slug> [--cwd <dir>]",
|
|
88
|
+
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
89
|
+
slug: { type: "str", required: true },
|
|
90
|
+
cwd: { type: "path" },
|
|
91
|
+
};
|
|
92
|
+
|
|
93
|
+
/**
|
|
94
|
+
* Report what the hunt on disk amounts to.
|
|
95
|
+
*
|
|
96
|
+
* @param {string[]} rawArgv - The subcommand's own arguments (harness.mjs strips the verb words).
|
|
97
|
+
* @returns {void} Exits 0 when the hunt result may be accepted, 1 when it may not.
|
|
98
|
+
*/
|
|
99
|
+
export function cli(rawArgv) {
|
|
100
|
+
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
101
|
+
const out = huntOutcome(resolve(args.cwd || process.cwd()), args.slug);
|
|
102
|
+
console.log(JSON.stringify(out));
|
|
103
|
+
process.exit(out.ok ? 0 : 1);
|
|
104
|
+
}
|
|
@@ -269,12 +269,38 @@ export function renderTable(r) {
|
|
|
269
269
|
return out.join("\n");
|
|
270
270
|
}
|
|
271
271
|
|
|
272
|
+
/**
|
|
273
|
+
* Whether a board that has tasks carries any `covers:` clause while a requirements registry exists.
|
|
274
|
+
*
|
|
275
|
+
* Two bars answer "is this requirement planned": L1b is satisfied by a scope contract's own covers
|
|
276
|
+
* list, while the matrix reads only the acceptance criteria's `(covers: REQ-…)` clauses. A board
|
|
277
|
+
* regenerated with none crossed every gate, and the matrix then read "no evidence" under every
|
|
278
|
+
* requirement of a run whose criteria had passed. Measured on two consecutive runs of one pitch:
|
|
279
|
+
* thirty clauses on one board, zero on the next, the same instruction both times.
|
|
280
|
+
*
|
|
281
|
+
* @param {string} cwd - Project root.
|
|
282
|
+
* @param {string} slug - Feature slug.
|
|
283
|
+
* @returns {(string|null)} A reason phrased for the board's writer, or null when there is no
|
|
284
|
+
* registry, no board yet, or at least one clause.
|
|
285
|
+
*/
|
|
286
|
+
export function boardCoversProblem(cwd, slug) {
|
|
287
|
+
const clauses = parseRequirements(readIf(requirementsFile(cwd, slug)) || "");
|
|
288
|
+
if (!clauses.length) return null;
|
|
289
|
+
const board = readBoard(cwd, slug);
|
|
290
|
+
if (!Array.isArray(board) || !board.length) return null;
|
|
291
|
+
if (coveringAcs(board).size) return null;
|
|
292
|
+
return `the board's ${board.length} tasks carry no \`(covers: REQ-…)\` clause while the registry holds ${clauses.length} ` +
|
|
293
|
+
"requirements — every acceptance criterion that grades a requirement carries its covers clause, or the requirements " +
|
|
294
|
+
"matrix reads no evidence for any of them";
|
|
295
|
+
}
|
|
296
|
+
|
|
272
297
|
export const ARGV_SPEC = {
|
|
273
|
-
usage: "harness.mjs probe requirements --slug <slug> [--run-id <id>] [--format json|table] [--cwd <dir>]",
|
|
298
|
+
usage: "harness.mjs probe requirements --slug <slug> [--run-id <id>] [--format json|table] [--board-check] [--cwd <dir>]",
|
|
274
299
|
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
275
300
|
slug: { type: "str", required: true },
|
|
276
301
|
"run-id": { type: "str" },
|
|
277
302
|
format: { type: "enum", values: ["json", "table"], default: "json" },
|
|
303
|
+
"board-check": { type: "flag" },
|
|
278
304
|
cwd: { type: "path" },
|
|
279
305
|
};
|
|
280
306
|
|
|
@@ -288,6 +314,12 @@ export const ARGV_SPEC = {
|
|
|
288
314
|
export async function cli(rawArgv) {
|
|
289
315
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
290
316
|
const cwd = resolve(args.cwd || process.cwd());
|
|
317
|
+
if (args.boardCheck) {
|
|
318
|
+
// `--board-check` answers one question for the run's send-back: may this board stand?
|
|
319
|
+
const reason = boardCoversProblem(cwd, args.slug);
|
|
320
|
+
console.log(JSON.stringify({ ok: !reason, reason }));
|
|
321
|
+
process.exit(reason ? 1 : 0);
|
|
322
|
+
}
|
|
291
323
|
const report = projectRequirements({ cwd, slug: args.slug, runId: args.runId ?? undefined });
|
|
292
324
|
console.log(args.format === "table" ? renderTable(report) : JSON.stringify(report, null, 2));
|
|
293
325
|
process.exit(0);
|
package/kernel/reduce/ingest.mjs
CHANGED
|
@@ -36,8 +36,9 @@ import { resolve, join, dirname, basename } from "node:path";
|
|
|
36
36
|
import { fileURLToPath } from "node:url";
|
|
37
37
|
import { validate } from "../verify/envelope.mjs";
|
|
38
38
|
import { runArgs } from "../lib/argv.mjs";
|
|
39
|
-
import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId } from "../lib/paths.mjs";
|
|
39
|
+
import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId, qaDir } from "../lib/paths.mjs";
|
|
40
40
|
import { citationProblem, coverageProblem, verdictProblem } from "../probe/eval.mjs";
|
|
41
|
+
import { huntProblem } from "../probe/hunt.mjs";
|
|
41
42
|
|
|
42
43
|
const HERE = dirname(fileURLToPath(import.meta.url));
|
|
43
44
|
const RESULT_SCHEMA = JSON.parse(readFileSync(resolve(HERE, "../schemas/work-result.schema.json"), "utf8"));
|
|
@@ -710,6 +711,24 @@ export async function cli(rawArgv) {
|
|
|
710
711
|
}
|
|
711
712
|
}
|
|
712
713
|
|
|
714
|
+
// --- Hunt gate: a `done` hunt over a reachable app ran at least one charter ---------------------
|
|
715
|
+
// A hunter that reached the app and drafted nothing returned `done`, and a `done` with no findings
|
|
716
|
+
// reads as a clean app. See `huntProblem`; `probe hunt` refuses the same result to the run.
|
|
717
|
+
if (result.worker === "qa-edge-hunter") {
|
|
718
|
+
const reportRel = (Array.isArray(result.artifacts) ? result.artifacts : []).find((a) => /hunt-report\.md$/.test(String(a)));
|
|
719
|
+
const reportAbs = reportRel ? (String(reportRel).startsWith("/") ? reportRel : join(cwd, reportRel)) : null;
|
|
720
|
+
let report = null;
|
|
721
|
+
for (const p of [reportAbs, join(qaDir(cwd, String(result.order_id).split("/")[0]), "hunt-report.md")]) {
|
|
722
|
+
if (!report && p && existsSync(p)) report = readFileSync(p, "utf8");
|
|
723
|
+
}
|
|
724
|
+
const problem = huntProblem(result, order.payload || {}, report);
|
|
725
|
+
if (problem) {
|
|
726
|
+
console.error(`ingest-result: result refused — ${problem}.`);
|
|
727
|
+
console.error(` Re-dispatch the hunter against its order. Nothing was written.`);
|
|
728
|
+
process.exit(1);
|
|
729
|
+
}
|
|
730
|
+
}
|
|
731
|
+
|
|
713
732
|
// Resolved for EVERY order, not only the gated ones: the attesting receipt is this leg's start,
|
|
714
733
|
// and a standalone or `--no-receipt-check` ingest still deserves a truthful timing row rather
|
|
715
734
|
// than one silently falling back to the order's re-writable `compiled_at`.
|
package/kernel/reduce/ship.mjs
CHANGED
|
@@ -39,6 +39,7 @@ import { ratchetReport } from "../probe/stats.mjs";
|
|
|
39
39
|
import { projectRequirements, summaryLine } from "../probe/requirements.mjs";
|
|
40
40
|
import { deriveRounds } from "../probe/rounds.mjs";
|
|
41
41
|
import { collectDiff, scanDiff, summarize } from "./leftovers.mjs";
|
|
42
|
+
import { chartersRun } from "../probe/hunt.mjs";
|
|
42
43
|
|
|
43
44
|
/** @returns {string} Today as `YYYY-MM-DD` (UTC). */
|
|
44
45
|
const today = () => new Date().toISOString().slice(0, 10);
|
|
@@ -392,7 +393,7 @@ export function buildReport(facts) {
|
|
|
392
393
|
*/
|
|
393
394
|
export function qaStatus(passed, huntReport) {
|
|
394
395
|
if (passed === "skipped" || (!passed && !huntReport)) return passed || "skipped";
|
|
395
|
-
if (huntReport &&
|
|
396
|
+
if (huntReport && chartersRun(huntReport) === 0) return "not-hunted";
|
|
396
397
|
return passed || "run";
|
|
397
398
|
}
|
|
398
399
|
|
|
@@ -2871,6 +2871,11 @@
|
|
|
2871
2871
|
"qa_findings": {
|
|
2872
2872
|
"type": "integer"
|
|
2873
2873
|
},
|
|
2874
|
+
"qa": {
|
|
2875
|
+
"type": "string",
|
|
2876
|
+
"enum": ["run", "not-hunted", "skipped"],
|
|
2877
|
+
"description": "What QA amounted to: run (at least one charter ran), not-hunted (the hunt returned and ran no charter), skipped (--no-qa or the QA gate answered skip). qa_findings is 0 in the last two for different reasons."
|
|
2878
|
+
},
|
|
2874
2879
|
"report": {
|
|
2875
2880
|
"type": "string",
|
|
2876
2881
|
"description": "shipped: shapeup/<slug>/REPORT.md path."
|
package/package.json
CHANGED
|
@@ -76,7 +76,9 @@ HARD (any miss → STOP, report which):
|
|
|
76
76
|
✅ deliverable reachable: one real request at `app_url`, or `launch_cmd` run and exit 0, or one
|
|
77
77
|
real invocation of the entry point (not a ping, not a guess that a tool is missing). When the
|
|
78
78
|
order carries `launch_cmd`, run it before concluding anything; the report names the command
|
|
79
|
-
and its exit. A hunt that reached nothing returns `status: failed`, never `done`.
|
|
79
|
+
and its exit. A hunt that reached nothing returns `status: failed`, never `done`. A hunt that
|
|
80
|
+
reached the app runs at least one charter: ingest refuses a `done` whose report says
|
|
81
|
+
`charters: 0/…` while the order carried `launch_cmd` or `app_url`.
|
|
80
82
|
✅ EVAL-FEATURE-<slug>.md exists with verdict: PASS
|
|
81
83
|
✅ if discovery/ledger.md exists: ledger.feature == <feature> (read-only context check —
|
|
82
84
|
a missing ledger is fine; ingest creates it when your findings land)
|
|
@@ -38,7 +38,7 @@
|
|
|
38
38
|
// maxParallelScopes (default 4).
|
|
39
39
|
//
|
|
40
40
|
// return — RunReturn (domain.schema.json $defs/RunReturn), the full union:
|
|
41
|
-
// { status: "shipped", verdict, rounds_used, dims_not_evaluated, qa_findings, report }
|
|
41
|
+
// { status: "shipped", verdict, rounds_used, dims_not_evaluated, qa_findings, qa, report }
|
|
42
42
|
// { status: "paused", paused_at, block, valid_decisions, context }
|
|
43
43
|
// { status: "aborted", aborted_at, reason }
|
|
44
44
|
// { status: "gate_h", breaker: "outer"|"attempt_budget"|"none"|"deadline", hammer_proposals, green_scopes,
|
|
@@ -632,6 +632,22 @@ const QA_REPORT = {
|
|
|
632
632
|
required: ["ok", "findings_count"],
|
|
633
633
|
};
|
|
634
634
|
|
|
635
|
+
const BOARD_CHECK = {
|
|
636
|
+
type: "object",
|
|
637
|
+
properties: { ok: { type: "boolean" }, reason: { type: ["string", "null"] } },
|
|
638
|
+
required: ["ok"],
|
|
639
|
+
};
|
|
640
|
+
|
|
641
|
+
const HUNT_OUTCOME = {
|
|
642
|
+
type: "object",
|
|
643
|
+
properties: {
|
|
644
|
+
ok: { type: "boolean" }, status: { type: ["string", "null"] },
|
|
645
|
+
qa: { type: "string", enum: ["run", "not-hunted", "missing"] },
|
|
646
|
+
charters_run: { type: ["integer", "null"] }, findings: { type: "integer" }, reason: { type: ["string", "null"] },
|
|
647
|
+
},
|
|
648
|
+
required: ["ok", "qa"],
|
|
649
|
+
};
|
|
650
|
+
|
|
635
651
|
const HAMMER = {
|
|
636
652
|
type: "object",
|
|
637
653
|
properties: {
|
|
@@ -856,6 +872,58 @@ async function worker({ skill, operation, payload, schema, phase: phaseName, lab
|
|
|
856
872
|
return (r && typeof r === "object") ? r : nullFail(label);
|
|
857
873
|
}
|
|
858
874
|
|
|
875
|
+
/**
|
|
876
|
+
* Dispatch a worker, check its result mechanically, and send it back ONCE with the refusal.
|
|
877
|
+
*
|
|
878
|
+
* A worker's result can be well-formed and still not what the run asked for — a verdict that grades
|
|
879
|
+
* the surface as a group, a hunt that reached the app and ran no charter. The kernel says so; this
|
|
880
|
+
* turns that into one more dispatch carrying the kernel's own reason, and only a second refusal
|
|
881
|
+
* stands. Measured: a rule in the skill made the judge grade row by row on its first dispatch, and
|
|
882
|
+
* the refusal is what holds when the model takes the loose path anyway.
|
|
883
|
+
*
|
|
884
|
+
* @param {function(string): Promise<object>} dispatch - Dispatches the worker; takes a note to append.
|
|
885
|
+
* @param {function(string): Promise<(object|null)>} check - Runs the kernel's check; takes a label suffix.
|
|
886
|
+
* @param {function(object): boolean} refused - Whether a check's answer is a correctable refusal.
|
|
887
|
+
* @param {string} what - What is being checked, for the log line.
|
|
888
|
+
* @returns {Promise<{r: object, v: (object|null)}>} The last dispatch and the last check.
|
|
889
|
+
*/
|
|
890
|
+
async function sendBackOnce(dispatch, check, refused, what) {
|
|
891
|
+
let r = await dispatch("");
|
|
892
|
+
if (r.__failed) return { r, v: null };
|
|
893
|
+
let v = await check("");
|
|
894
|
+
if (v && refused(v)) {
|
|
895
|
+
log(`${what} — refused (${v.reason}); sent back once`);
|
|
896
|
+
r = await dispatch(` Your previous result was refused: ${v.reason}. Correct that and return again.`);
|
|
897
|
+
if (r.__failed) return { r, v: null };
|
|
898
|
+
v = await check(":again");
|
|
899
|
+
}
|
|
900
|
+
return { r, v };
|
|
901
|
+
}
|
|
902
|
+
|
|
903
|
+
/**
|
|
904
|
+
* Dispatch a board writer (analyze, or the board-only regeneration) and send it back once when the
|
|
905
|
+
* board carries no `covers:` clause while a requirements registry exists.
|
|
906
|
+
*
|
|
907
|
+
* Advisory past the second dispatch: the requirements matrix never blocks a ship, so a board that
|
|
908
|
+
* still carries none continues with a state warning naming what the matrix will read, rather than
|
|
909
|
+
* aborting a run whose plan is otherwise sound.
|
|
910
|
+
*
|
|
911
|
+
* @param {function(string): Promise<object>} dispatch - The board writer's dispatch; takes a note.
|
|
912
|
+
* @param {string} what - The operation, for the log and the warning.
|
|
913
|
+
* @returns {Promise<{r: object, v: (object|null)}>} The last dispatch and the last check.
|
|
914
|
+
*/
|
|
915
|
+
async function boardSentBack(dispatch, what) {
|
|
916
|
+
const out = await sendBackOnce(dispatch,
|
|
917
|
+
(suffix) => query(`probe requirements --slug ${slug} --board-check`, BOARD_CHECK, "Analyze", `board-covers${suffix}`),
|
|
918
|
+
(v) => !v.ok && !!v.reason, `ANALYZE ${what}`);
|
|
919
|
+
if (!out.r.__failed && out.v && !out.v.ok && out.v.reason) {
|
|
920
|
+
const msg = `ANALYZE ${what}: ${out.v.reason} (after one send-back; the requirements matrix will read no evidence)`;
|
|
921
|
+
log(msg);
|
|
922
|
+
stateWarnings.push(msg);
|
|
923
|
+
}
|
|
924
|
+
return out;
|
|
925
|
+
}
|
|
926
|
+
|
|
859
927
|
// ---------------------------------------------------------------------------------------------
|
|
860
928
|
// GATES — the kernel's exit-code convention, unchanged: 0 cross · 4 pause · 5 abort. The resolved
|
|
861
929
|
// decision travels in `decision`, copied verbatim out of the kernel's own JSON.
|
|
@@ -1172,7 +1240,7 @@ async function closeIfTerminal(ret) {
|
|
|
1172
1240
|
// A result the single writer never applied is named at the close, not folded into "not green".
|
|
1173
1241
|
+ (Array.isArray(ret.unapplied_results) && ret.unapplied_results.length ? ` unapplied_results=${ret.unapplied_results.length}` : "")
|
|
1174
1242
|
+ (ret.census ? ` census=${ret.census}` : "") + (ret.l4 ? ` l4=${ret.l4}` : "") + (ret.ship_report ? ` ship_report=${ret.ship_report}` : "")
|
|
1175
|
-
: `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}`
|
|
1243
|
+
: `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}` + (ret.qa ? ` qa=${ret.qa}` : "")
|
|
1176
1244
|
+ (ret.after ? ` after=${ret.after} breaker=${ret.breaker ?? "?"} census=${ret.census ?? "?"} cut_list=${Array.isArray(ret.cut_list) ? ret.cut_list.length : "?"}` : "");
|
|
1177
1245
|
// `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
|
|
1178
1246
|
// a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
|
|
@@ -1384,11 +1452,11 @@ if (!rs.has_spec_tree) {
|
|
|
1384
1452
|
// single run since the cutover. It is advisory, so nothing stopped; the ledger simply stayed on
|
|
1385
1453
|
// the previous phase and the snapshot under-reported where the run had got to.
|
|
1386
1454
|
await setRunStatus("mapping", "Analyze");
|
|
1387
|
-
const a = await worker({
|
|
1388
|
-
skill: "ba-pitch-analyzer", operation: "analyze", schema: PHASE_OK, phase: "Analyze", label: "analyze",
|
|
1455
|
+
const { r: a } = await boardSentBack((note) => worker({
|
|
1456
|
+
skill: "ba-pitch-analyzer", operation: "analyze", schema: PHASE_OK, phase: "Analyze", label: note ? "analyze:again" : "analyze",
|
|
1389
1457
|
payload: { pitch: rs.intake_path, breadboard: rs.breadboard_path, spec_folder: specFolder, feature: slug, lens: rs.lens, orient_dir: rs.orient_dir },
|
|
1390
|
-
extra: "Write the spec tree and the board from the orient artifacts — do not re-scan the code.",
|
|
1391
|
-
});
|
|
1458
|
+
extra: "Write the spec tree and the board from the orient artifacts — do not re-scan the code." + note,
|
|
1459
|
+
}), "analyze");
|
|
1392
1460
|
if (a.__failed) return await withWarnings(diedAt("ANALYZE", a));
|
|
1393
1461
|
const post = await requirePhase("ANALYZE", "analyze", "Analyze");
|
|
1394
1462
|
if (post) return await withWarnings(post);
|
|
@@ -1402,13 +1470,13 @@ if (!rs.has_spec_tree) {
|
|
|
1402
1470
|
// operation regenerates it from the tree without re-deriving the tree.
|
|
1403
1471
|
log(`ANALYZE — spec tree on disk, no board: dispatching the board-only operation (slug ${slug})`);
|
|
1404
1472
|
await setRunStatus("mapping", "Analyze");
|
|
1405
|
-
const b = await worker({
|
|
1406
|
-
skill: "ba-pitch-analyzer", operation: "board", schema: PHASE_OK, phase: "Analyze", label: "board",
|
|
1473
|
+
const { r: b } = await boardSentBack((note) => worker({
|
|
1474
|
+
skill: "ba-pitch-analyzer", operation: "board", schema: PHASE_OK, phase: "Analyze", label: note ? "board:again" : "board",
|
|
1407
1475
|
payload: { spec_folder: specFolder, feature: slug, lens: rs.lens },
|
|
1408
1476
|
extra: "The spec tree is committed and FROZEN for this dispatch. Regenerate the per-machine board " +
|
|
1409
1477
|
"under the run's tasks/ directory from the use cases on disk — every acceptance criterion " +
|
|
1410
|
-
"carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder.",
|
|
1411
|
-
});
|
|
1478
|
+
"carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder." + note,
|
|
1479
|
+
}), "board");
|
|
1412
1480
|
if (b.__failed) return await withWarnings(diedAt("ANALYZE", b));
|
|
1413
1481
|
// The board a worker writes carries `depends_on`; its inverse (`unlocks`) and a new task's
|
|
1414
1482
|
// `status: todo` are mechanical, so the kernel writes them rather than trusting every regeneration
|
|
@@ -1872,31 +1940,26 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1872
1940
|
// freezes at GATE L4 has to say that as plainly as the gate block already does.
|
|
1873
1941
|
verdict = "not-evaluated";
|
|
1874
1942
|
} else {
|
|
1875
|
-
const
|
|
1876
|
-
|
|
1877
|
-
|
|
1878
|
-
|
|
1879
|
-
|
|
1880
|
-
|
|
1881
|
-
|
|
1882
|
-
|
|
1883
|
-
|
|
1884
|
-
|
|
1885
|
-
|
|
1886
|
-
|
|
1943
|
+
const { r: e, v: ev } = await sendBackOnce(
|
|
1944
|
+
(note) => worker({
|
|
1945
|
+
skill: "spec-evaluator", operation: "evaluate", schema: EVAL, phase: "Eval",
|
|
1946
|
+
label: note ? `eval:r${round}:again` : `eval:r${round}`, model: evalModel, round,
|
|
1947
|
+
// No `t0_artifacts` here, deliberately: `harness compile` derives them from the round's green
|
|
1948
|
+
// T0 verdicts on disk, for every lane — this script could only name paths it was told about.
|
|
1949
|
+
// The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
|
|
1950
|
+
// artifact and the project profile, so the judge gets the launch evidence without this script
|
|
1951
|
+
// having to carry it.
|
|
1952
|
+
payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
|
|
1953
|
+
extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself." + note,
|
|
1954
|
+
}),
|
|
1955
|
+
// The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
|
|
1956
|
+
// agent's own summary of it (`e.overall`) — see EVAL_VERDICT's comment for why.
|
|
1957
|
+
(suffix) => query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}${suffix}`),
|
|
1958
|
+
// A verdict the kernel refused — grouped rows, a missing citation — is correctable, not a dead judge.
|
|
1959
|
+
(v) => !v.ok && !!v.overall && !!v.reason,
|
|
1960
|
+
`EVAL r${round}`,
|
|
1961
|
+
);
|
|
1887
1962
|
if (e.__failed) return await withWarnings(diedAt("L3", e));
|
|
1888
|
-
// The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
|
|
1889
|
-
// agent's own summary of it (`e.overall`) — see EVAL_VERDICT's comment for why.
|
|
1890
|
-
let ev = await query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}`);
|
|
1891
|
-
// A verdict the kernel refused — a PASS that grades the surface as a group, a missing citation —
|
|
1892
|
-
// is a correctable answer, not a dead judge: the judge is sent back once with the refusal's own
|
|
1893
|
-
// words, and only a second refusal ends the round.
|
|
1894
|
-
if (ev && !ev.ok && ev.overall && ev.reason) {
|
|
1895
|
-
log(`EVAL r${round} — verdict refused (${ev.reason}); the judge is sent back once`);
|
|
1896
|
-
e = await evalOnce(`eval:r${round}:again`, ` Your previous verdict for this round was refused: ${ev.reason}. Grade again and correct that.`);
|
|
1897
|
-
if (e.__failed) return await withWarnings(diedAt("L3", e));
|
|
1898
|
-
ev = await query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}:again`);
|
|
1899
|
-
}
|
|
1900
1963
|
if (!ev) return await withWarnings(diedAt("L3", nullFail(`verdict:r${round}`)));
|
|
1901
1964
|
// A round with no verdict to act on is NOT a dead worker. An evaluator that refused the round
|
|
1902
1965
|
// wrote a result saying why, and `probe eval` carries it as `reason`; reported as "died after
|
|
@@ -1953,15 +2016,26 @@ let qaFindings = 0;
|
|
|
1953
2016
|
const qaG = await crossGate("QA", "QA", ["run", "skip", "ask"], { round, verdict });
|
|
1954
2017
|
if (qaG.stop) return await withWarnings(qaG.stop);
|
|
1955
2018
|
const qaRan = !args.noQa && qaG.decision === "run";
|
|
2019
|
+
// What QA amounted to, from the hunt's own record: `run` only when a charter ran. The QA gate's `run`
|
|
2020
|
+
// is the decision to dispatch, and a dispatched hunt that drove nothing is `not-hunted`.
|
|
2021
|
+
let qaState = "skipped";
|
|
1956
2022
|
if (qaRan) {
|
|
1957
|
-
const q = await
|
|
1958
|
-
|
|
1959
|
-
|
|
1960
|
-
|
|
1961
|
-
|
|
2023
|
+
const { r: q, v: hv } = await sendBackOnce(
|
|
2024
|
+
(note) => worker({
|
|
2025
|
+
skill: "qa-edge-hunter", operation: "hunt", schema: QA_REPORT, phase: "QA", label: note ? "hunt:again" : "hunt", model: qaModel,
|
|
2026
|
+
payload: { feature: slug, spec_folder: specFolder, app_url: rs.app_url, round },
|
|
2027
|
+
extra: "Exploratory hunt over the shipped feature. No verdict and no score — findings only, each with a repro." + note,
|
|
2028
|
+
}),
|
|
2029
|
+
(suffix) => query(`probe hunt --slug ${slug}`, HUNT_OUTCOME, "QA", `hunt-outcome${suffix}`),
|
|
2030
|
+
(v) => !v.ok && v.status === "done" && !!v.reason,
|
|
2031
|
+
"QA hunt",
|
|
2032
|
+
);
|
|
1962
2033
|
// QA is a level-up: losing its worker costs the findings, not the run.
|
|
1963
2034
|
if (q.__failed) log(`QA — the hunt lost its worker: ${q.__failed}. Shipping without QA findings.`);
|
|
1964
|
-
else
|
|
2035
|
+
else {
|
|
2036
|
+
qaFindings = hv && Number.isInteger(hv.findings) ? hv.findings : q.findings_count;
|
|
2037
|
+
qaState = hv?.qa === "run" ? "run" : "not-hunted";
|
|
2038
|
+
}
|
|
1965
2039
|
}
|
|
1966
2040
|
|
|
1967
2041
|
// ---- GATE H — delegated to scope-hammer (census, baseline comparison, cut list) ----------------
|
|
@@ -1986,7 +2060,7 @@ if (h.verdict === "cannot-ship") {
|
|
|
1986
2060
|
// "not-evaluated" as a real verdict (its own usage string, and `generate()`'s no-artifact
|
|
1987
2061
|
// default). Hardcoding PASS here is exactly the silent upgrade protocol.md's Rules forbid.
|
|
1988
2062
|
const shipVerdict = verdict === "not-evaluated" ? "not-evaluated" : "PASS";
|
|
1989
|
-
const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${
|
|
2063
|
+
const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${qaState}`, "Ship", "ship-report");
|
|
1990
2064
|
// GATE L4 HAS A CALL SITE. It used to be a line of prose in the tech lead's references — "resolve
|
|
1991
2065
|
// the gate itself before any of the above" — and a run reached its ship decision with no ledger row
|
|
1992
2066
|
// for it. The resolver narrows `ship` on the census artifact the hammer just wrote.
|
|
@@ -2016,6 +2090,7 @@ return await withWarnings({
|
|
|
2016
2090
|
rounds_used: round,
|
|
2017
2091
|
dims_not_evaluated: ALL_DIMS.filter((d) => !evalDims.includes(d)),
|
|
2018
2092
|
qa_findings: qaFindings,
|
|
2093
|
+
qa: qaState,
|
|
2019
2094
|
report: REPORT_PATH,
|
|
2020
2095
|
});
|
|
2021
2096
|
|