shapeup-sdlc 3.17.1 → 3.18.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +4 -4
- package/README.md +2 -1
- package/SECURITY.md +1 -1
- package/hooks/sandbox-guard.mjs +54 -5
- package/kernel/compile.mjs +49 -0
- package/kernel/init/run.mjs +11 -8
- package/kernel/probe/eval.mjs +76 -3
- package/kernel/probe/resume.mjs +27 -3
- package/kernel/reduce/ingest.mjs +2 -1
- package/kernel/report/facts.mjs +1 -1
- package/kernel/schemas/domain.schema.json +15 -1
- package/package.json +1 -1
- package/skills/spec-evaluator/SKILL.md +3 -3
- package/skills/tech-lead/SKILL.md +5 -8
- package/skills/tech-lead/references/gates.md +8 -6
- package/skills/tech-lead/references/protocol.md +1 -1
- package/skills/tech-lead/workflows/shapeup-run.js +20 -4
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.18.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -7,13 +7,13 @@ A three-phase Shape Up loop orchestrated by `/tech-lead`. Invariants live in the
|
|
|
7
7
|
|
|
8
8
|
Skills and commands are named short throughout this file; every one of them resolves only under the plugin's namespace — `Skill(shapeup-sdlc-plugin:orient)`, `/shapeup-sdlc-plugin:ship`. A bare name is not a typo you get a warning for: an unknown skill name is rejected upstream of the hook layer, so the dispatch fires no hook, leaves no decision row, and is answered by improvisation.
|
|
9
9
|
|
|
10
|
-
- Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
|
|
10
|
+
- Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing the result of a live order your own dispatch did not take, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
|
|
11
11
|
- **The tier-direction rule is refused at the write, not argued at L1b.** A committed file cannot carry a `.shapeup/` path or a `TASK-NNN` id — boards renumber per machine and the run trace is gitignored, so either one resolves only where it was written. Spec-lint has always red'd this at Board Review; the same rule now also refuses the write itself, naming the token, because four producers hit it and one of them hit it twice in one session — a worker carries no lesson across a dispatch, and by L1b the phase that wrote the line is over. The knowledge base is outside the rule: those files are instructions, not references anyone resolves.
|
|
12
12
|
- Attested, not assumed: a dispatch leaves a receipt naming the skill that ran, and ingest refuses an orchestrated result that has none. A dispatch that fails is answered by the sub-agent improvising the craft, which every other check accepts — so "the artifact exists" is not evidence that the shipped skill produced it. `--no-receipt-check` is the way through when the receipt channel itself fails.
|
|
13
13
|
- **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
|
|
14
14
|
- GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
|
|
15
15
|
- Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
|
|
16
|
-
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it. A relaunch that resumes a run which *did* close reopens it on the record: the earlier close is kept under `prior_closes`, the close fields clear, and the run's next ending is recorded as its own — so a run that aborted, was resumed and then shipped reads `shipped` with the abort beside it, never an abort under a `shipped` status. Opening a run over one that closed says it is closed, not open.
|
|
16
|
+
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it. A relaunch that resumes a run which *did* close reopens it on the record: the earlier close is kept under `prior_closes`, the close fields clear, the run pointer the close retired comes back (so the resumed run's hook decisions carry its key again), and the run's next ending is recorded as its own — so a run that aborted, was resumed and then shipped reads `shipped` with the abort beside it, never an abort under a `shipped` status. Opening a run over one that closed says it is closed, not open.
|
|
17
17
|
- **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
|
|
18
18
|
- The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
|
|
19
19
|
|
|
@@ -36,7 +36,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
38
|
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
39
|
-
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade. A PASS names every Test Surface row as its own criterion: one that grades the surface as a group is refused, and the judge is sent back once with the reason. A scope that did not go green in the round has its rows graded FAIL, and the order names such scopes, so one red scope never makes the whole round ungradeable |
|
|
39
|
+
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; the order names the dimensions it grades — the set the PO named at L0, or else the set the spec calls for (a Test Surface turns on test-surface-conformance, Invariants turn on completeness, and the test-surface and suite checks are always on), and L4 lists the ones left out. A PASS that grades no criterion under a dimension its order named is refused, and the judge is sent back once; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade. A PASS names every Test Surface row as its own criterion: one that grades the surface as a group is refused, and the judge is sent back once with the reason. A scope that did not go green in the round has its rows graded FAIL, and the order names such scopes, so one red scope never makes the whole round ungradeable |
|
|
40
40
|
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC. Every criterion the verdict graded FAIL reaches the round, whether or not the judge filed a bug for it, addressed to the scope that owns its use case — a round with a FAIL verdict and nothing to fix is how a run stalls. A row whose check the scope rewrote between a failing trial and a passing one is listed for the judge, who reads it against the row before the pass counts |
|
|
41
41
|
|
|
42
42
|
✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
|
|
@@ -62,7 +62,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
|
|
|
62
62
|
- **Ledger = single source of truth** — every discovery flow writes only its own section.
|
|
63
63
|
- **QA is a level-up, not a gate** — `--no-qa` skips it; circuit breaker outranks the Hunter.
|
|
64
64
|
- **Role separation** — Evaluator grades, task-executor fixes, QA discovers.
|
|
65
|
-
- **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
|
|
65
|
+
- **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. A verdict that anchors no criterion at all, beside a board whose ACs carry `covers:`, is sent back to the judge once. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
|
|
66
66
|
- **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts, never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
|
|
67
67
|
- **Envelope port (v1.0)** — every dispatch is WorkOrder in / WorkResult out; shared state has exactly one writer (the ingest step); malformed envelopes are hook-denied. Workers: stateless, craft-only, pipeline-blind.
|
|
68
68
|
|
package/README.md
CHANGED
|
@@ -235,7 +235,8 @@ the layer that carries it, and the three layers here fail differently:
|
|
|
235
235
|
an evaluation or a QA leg that never returned a result no longer holds the board. `frozen` is
|
|
236
236
|
checked first and outranks everything, across every live contract — including the carve-out that
|
|
237
237
|
otherwise keeps the active feature's own run trace writable, so a path a live order froze stays
|
|
238
|
-
frozen wherever it lives.
|
|
238
|
+
frozen wherever it lives. An order's own result answers to the sub-agent that took its
|
|
239
|
+
dispatch, so with two legs live neither can write the other's.
|
|
239
240
|
- `PreToolUse` (`Edit|Write|MultiEdit`) — **`hooks/tier-guard.mjs` refuses a committed-tier write
|
|
240
241
|
whose content names a path into the local run trace or a board id.** Same rule spec-lint reds at
|
|
241
242
|
GATE L1b, asked at the moment of writing: four different producers wrote a committed file their
|
package/SECURITY.md
CHANGED
|
@@ -75,7 +75,7 @@ sitting, and reading them is the recommended review.
|
|
|
75
75
|
| [`safety-spine.mjs`](hooks/safety-spine.mjs) | PreToolUse (`Bash\|Read\|Write\|Edit\|MultiEdit`) | The proposed command/path; `.shapeup/safety-overrides.json` | Yes — provably destructive ops only: `rm -rf` on unrecoverable targets, `git push --force` / push to main, `git reset --hard`, `git clean -fdx`, `DROP TABLE`/`TRUNCATE`, reads of `.env`/keys/cloud credentials, and any write to its own overrides file | Never blocks an unmatched command; `--force-with-lease` stays allowed |
|
|
76
76
|
| [`gate-intake.mjs`](hooks/gate-intake.mjs) | PreToolUse (`Skill`) | The `tech-lead` dispatch's own arguments | Yes — an orchestrator dispatch carrying no resolvable intake (no pitch, spec, resume or requirement text) | Fails open on `--order` and on any ambiguous arg shape |
|
|
77
77
|
| [`harness verify envelope`](kernel/verify/envelope.mjs) | PreToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the JSON schemas | Yes — a worker dispatch whose order file is missing or schema-invalid | Never gates a dispatch that carries no `--order` (standalone skill use stays free) |
|
|
78
|
-
| [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path, resolved through the filesystem before it is compared — symlinks followed, the unborn tail re-attached to its real ancestor, case folded where the filesystem folds case — so a spelling variant of a frozen file is the frozen file; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
|
|
78
|
+
| [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path, resolved through the filesystem before it is compared — symlinks followed, the unborn tail re-attached to its real ancestor, case folded where the filesystem folds case — so a spelling variant of a frozen file is the frozen file; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out. An order's own paths (its result first) are writable only by the sub-agent its dispatch receipt names, so one live leg cannot write another's result; a write with no sub-agent id, or an order with no receipt, keeps the path-only rule | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
|
|
79
79
|
| [`tier-guard.mjs`](hooks/tier-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the text the call is about to write (`content`, `new_string`, or any edit in a `MultiEdit` batch) | Yes — a write into `shapeup/<slug>/` whose content names a `.shapeup/` path or a `TASK-NNN` board id: the same tier-direction rule spec-lint reds at GATE L1b, refused at the moment of writing, with the offending token quoted back | No-op outside `shapeup/<slug>/`, on file forms the lint does not scan, and on `shapeup/knowledge-base/` (instructions to a worker, not references). Reads nothing off disk and writes only its decision row; it never inspects a file it is not being asked to write |
|
|
80
80
|
| [`dispatch-receipt.mjs`](hooks/dispatch-receipt.mjs) | PostToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the tool result's own report of which skill ran | **No — it has no deny path at all.** It records that the shipped skill ran, so `harness reduce ingest` can refuse a result no dispatch produced | Never writes an attestation for a result that does not name a resolved skill; never fails the call it observes (every write is inside `try`/`catch`) |
|
|
81
81
|
| [`gate-zerowork.mjs`](hooks/gate-zerowork.mjs) | Stop | Run receipts on disk; the session transcript; the decision ledger | **Yes — the one blocking hook.** Returns `decision:"block"` when the session dispatched the orchestrator and produced no run receipt | Defers the moment any receipt exists; `stop_hook_active` caps it at one block per stop chain |
|
package/hooks/sandbox-guard.mjs
CHANGED
|
@@ -88,7 +88,7 @@
|
|
|
88
88
|
import { readFileSync, existsSync, appendFileSync, mkdirSync, readdirSync, statSync, realpathSync, lstatSync, readlinkSync } from "node:fs";
|
|
89
89
|
import { resolve, join, relative, dirname, basename, sep } from "node:path";
|
|
90
90
|
import { isMain } from "../kernel/lib/argv.mjs";
|
|
91
|
-
import { LOCAL, SHARED, activeOrder, ordersDir, resultsDir, metricsShard } from "../kernel/lib/paths.mjs";
|
|
91
|
+
import { LOCAL, SHARED, activeOrder, ordersDir, resultsDir, metricsShard, dispatchReceipts } from "../kernel/lib/paths.mjs";
|
|
92
92
|
import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
|
|
93
93
|
|
|
94
94
|
// --- tiny glob matcher: supports *, **, ? — enough for substrate globs, zero dependencies ---
|
|
@@ -262,6 +262,41 @@ function liveOrders(cwd, slug) {
|
|
|
262
262
|
return unanswered.filter((o) => !pastItsPhase(o, stamps));
|
|
263
263
|
}
|
|
264
264
|
|
|
265
|
+
/**
|
|
266
|
+
* The agent that took each live order's dispatch, from the receipts the dispatch hook wrote.
|
|
267
|
+
*
|
|
268
|
+
* WHO WRITES, not only where. An order's own paths — its WorkResult above all — are carved out of
|
|
269
|
+
* the frozen run trace, and with two build legs live the carve-out answered for either of them: a
|
|
270
|
+
* leg could write its sibling's result. The host names the sub-agent behind every tool call
|
|
271
|
+
* (`agent_id`), and the dispatch receipt recorded the same id when that sub-agent invoked the
|
|
272
|
+
* worker skill, so the skill's writes carry the id of the receipt that aimed it. Only a receipt
|
|
273
|
+
* taken after the order was compiled counts, and the newest wins, so a re-dispatch of the same
|
|
274
|
+
* order binds to the agent that took it this time.
|
|
275
|
+
*
|
|
276
|
+
* @param {string} cwd - Project root.
|
|
277
|
+
* @param {string} slug - The run named by the pointer.
|
|
278
|
+
* @param {object[]} orders - The live orders.
|
|
279
|
+
* @returns {Map<string,string>} order_id → agent_id; an order with no attributable receipt is absent.
|
|
280
|
+
*/
|
|
281
|
+
export function dispatchAgents(cwd, slug, orders) {
|
|
282
|
+
const out = new Map();
|
|
283
|
+
let text = "";
|
|
284
|
+
try { text = readFileSync(dispatchReceipts(cwd, slug), "utf8"); } catch { return out; }
|
|
285
|
+
const compiled = new Map(orders.map((o) => [o.order_id, Date.parse(o.compiled_at ?? "")]));
|
|
286
|
+
const newest = new Map();
|
|
287
|
+
for (const line of text.split("\n")) {
|
|
288
|
+
let r; try { r = JSON.parse(line); } catch { continue; }
|
|
289
|
+
if (!r?.order_id || !compiled.has(r.order_id) || typeof r.agent_id !== "string" || !r.agent_id) continue;
|
|
290
|
+
const at = Date.parse(r.at ?? "");
|
|
291
|
+
const c = compiled.get(r.order_id);
|
|
292
|
+
if (Number.isNaN(at) || (!Number.isNaN(c) && at < c)) continue;
|
|
293
|
+
const prev = newest.get(r.order_id);
|
|
294
|
+
if (!prev || at >= prev.at) newest.set(r.order_id, { at, agent: r.agent_id });
|
|
295
|
+
}
|
|
296
|
+
for (const [id, v] of newest) out.set(id, v.agent);
|
|
297
|
+
return out;
|
|
298
|
+
}
|
|
299
|
+
|
|
265
300
|
function extractPaths(toolInput) {
|
|
266
301
|
const paths = [];
|
|
267
302
|
if (toolInput?.file_path) paths.push(toolInput.file_path);
|
|
@@ -337,6 +372,7 @@ async function main() {
|
|
|
337
372
|
})).filter((c) => c.allowed.length || c.appendOnly.length || c.frozen.length);
|
|
338
373
|
|
|
339
374
|
if (contracts.length === 0) defer("no live order declares write/append/frozen boundaries", "no-whitelist");
|
|
375
|
+
const agents = dispatchAgents(root, active.slug, withSubstrate);
|
|
340
376
|
|
|
341
377
|
const targetPaths = extractPaths(p.tool_input);
|
|
342
378
|
if (targetPaths.length === 0) defer("no writable path in the tool input", "no-target");
|
|
@@ -370,7 +406,17 @@ async function main() {
|
|
|
370
406
|
// wherever it lives.
|
|
371
407
|
// `own` first, and only against the contract that declared it: another live order's exception
|
|
372
408
|
// never licenses this write. A path no contract claims as its own falls through to the freeze.
|
|
373
|
-
|
|
409
|
+
// And only from the dispatch that took that order. A write from a sub-agent is refused when every
|
|
410
|
+
// live order claiming the path was taken by a different sub-agent; a write with no agent id (the
|
|
411
|
+
// operator's own session) or an order with no attributable receipt keeps today's permit.
|
|
412
|
+
const owners = contracts.filter((c) => matchesAny(rel, c.own, fold));
|
|
413
|
+
if (owners.length) {
|
|
414
|
+
const writer = typeof p.agent_id === "string" && p.agent_id ? p.agent_id : null;
|
|
415
|
+
if (!writer || owners.some((c) => !agents.has(c.order_id) || agents.get(c.order_id) === writer)) continue;
|
|
416
|
+
violations.push(rel);
|
|
417
|
+
blockReasons.push(`${rel} belongs to ${owners.map((c) => c.order_id).join(", ")}, whose dispatch is another agent's`);
|
|
418
|
+
continue;
|
|
419
|
+
}
|
|
374
420
|
|
|
375
421
|
const freezer = contracts.find((c) => matchesAny(rel, c.frozen, fold));
|
|
376
422
|
if (freezer) {
|
|
@@ -409,7 +455,10 @@ async function main() {
|
|
|
409
455
|
// direction. Those files belong to the orchestrator, whose write window is a phase boundary —
|
|
410
456
|
// no dispatch in flight — and never the middle of somebody else's dispatch.
|
|
411
457
|
const committed = violations.filter((v) => (fold ? v.split(/[\\/]/)[0].toLowerCase() === SHARED.toLowerCase() : v.split(/[\\/]/)[0] === SHARED));
|
|
412
|
-
const
|
|
458
|
+
const notOwn = blockReasons.every((r) => r.endsWith("whose dispatch is another agent's"));
|
|
459
|
+
const hint = notOwn
|
|
460
|
+
? "Each leg writes only the paths of the order it was dispatched with. Return your own result, never a sibling's."
|
|
461
|
+
: committed.length === violations.length
|
|
413
462
|
? `${SHARED}/ is committed tier: these belong to the orchestrator, not to a worker substrate. `
|
|
414
463
|
+ "Write them at a phase boundary, with no dispatch in flight — do not widen an order to reach one."
|
|
415
464
|
: "If this write legitimately crosses scopes, the order's substrate needs to be expanded (e.g. via ba --remap).";
|
|
@@ -430,14 +479,14 @@ async function main() {
|
|
|
430
479
|
// FROZE the path and a write refused because no contract covers it are different facts with
|
|
431
480
|
// different remedies, and a single rule string cannot tell the reader which one happened —
|
|
432
481
|
// which is how a frozen declaration can stop being enforced without a single row moving.
|
|
433
|
-
rule: frozenHits === violations.length ? "frozen" : "outside-substrate",
|
|
482
|
+
rule: frozenHits === violations.length ? "frozen" : notOwn ? "not-own-dispatch" : "outside-substrate",
|
|
434
483
|
reason: `${violations.length} write(s) rejected by substrate boundaries: ${blockReasons.join("; ")}`,
|
|
435
484
|
payload: {
|
|
436
485
|
hookSpecificOutput: {
|
|
437
486
|
hookEventName: "PreToolUse",
|
|
438
487
|
permissionDecision: "deny",
|
|
439
488
|
permissionDecisionReason:
|
|
440
|
-
`Sandbox guard (PA3) — no live order's substrate covers these writes:\n` +
|
|
489
|
+
`Sandbox guard (PA3) — ${notOwn ? "these paths belong to another live dispatch" : "no live order's substrate covers these writes"}:\n` +
|
|
441
490
|
`${blockReasons.join("\n")}\n` +
|
|
442
491
|
hint,
|
|
443
492
|
},
|
package/kernel/compile.mjs
CHANGED
|
@@ -921,6 +921,50 @@ export function launchEvidenceFor(cwd, slug, round) {
|
|
|
921
921
|
return out;
|
|
922
922
|
}
|
|
923
923
|
|
|
924
|
+
// --- the dimensions the judge grades ------------------------------------------------------------
|
|
925
|
+
//
|
|
926
|
+
// WHY THE KERNEL RESOLVES THEM. The judge's own rule is that an explicit `dimensions` list overrides
|
|
927
|
+
// its auto-enable, and every orchestrated order used to carry one: the ledger's default
|
|
928
|
+
// `[spec-conformance]`, which nobody chose. So a spec whose every use case carried a Test Surface
|
|
929
|
+
// was never graded by `test-surface-conformance`, `completeness` never ran over Invariants, and
|
|
930
|
+
// `tdd-surface` — always-on in the judge's registry — was off in every orchestrated run. Resolving
|
|
931
|
+
// the set here, from the spec, with the judge's rules, makes it a fact on the order rather than a
|
|
932
|
+
// default: the order file is the record of what was graded, and GATE L4 reads it back.
|
|
933
|
+
|
|
934
|
+
/** The shipped dimensions, in the judge's load order. Security and performance ship disabled. */
|
|
935
|
+
export const SHIPPED_DIMENSIONS = ["spec-conformance", "tdd-surface", "integration", "completeness",
|
|
936
|
+
"test-surface-conformance", "security", "performance"];
|
|
937
|
+
|
|
938
|
+
/**
|
|
939
|
+
* The dimension set an evaluate order grades when the run named none — the judge registry's rules,
|
|
940
|
+
* applied to the spec on disk.
|
|
941
|
+
*
|
|
942
|
+
* @param {string} cwd - Project root.
|
|
943
|
+
* @param {string} slug - Feature slug (reads the board for task variants).
|
|
944
|
+
* @param {(string|null)} specDir - The spec folder, relative to `cwd`; null → no use cases read.
|
|
945
|
+
* @returns {string[]} `spec-conformance` and `tdd-surface` always; `integration` when a board task is
|
|
946
|
+
* a `.be`/`.e2e` variant; `completeness` when a use case has `## Invariants`; `test-surface-conformance`
|
|
947
|
+
* when one has `## Test Surface`. In {@link SHIPPED_DIMENSIONS} order.
|
|
948
|
+
*/
|
|
949
|
+
export function resolveDimensions(cwd, slug, specDir) {
|
|
950
|
+
const on = new Set(["spec-conformance", "tdd-surface"]);
|
|
951
|
+
let tasks = [];
|
|
952
|
+
try { tasks = readBoard(cwd, slug); } catch { /* an unreadable board names no variant */ }
|
|
953
|
+
if (tasks.some((t) => /\.(be|e2e)$/i.test(String(t.id || "")))) on.add("integration");
|
|
954
|
+
if (specDir) {
|
|
955
|
+
const dir = join(resolve(cwd, specDir), "usecases");
|
|
956
|
+
let files = [];
|
|
957
|
+
try { files = readdirSync(dir).filter((f) => /^UC-.*\.md$/i.test(f)); } catch { /* no use cases */ }
|
|
958
|
+
for (const f of files) {
|
|
959
|
+
let body = "";
|
|
960
|
+
try { body = readFileSync(join(dir, f), "utf8"); } catch { continue; }
|
|
961
|
+
if (/^##\s+Invariants\b/m.test(body)) on.add("completeness");
|
|
962
|
+
if (/^##\s+Test Surface\b/m.test(body)) on.add("test-surface-conformance");
|
|
963
|
+
}
|
|
964
|
+
}
|
|
965
|
+
return SHIPPED_DIMENSIONS.filter((d) => on.has(d));
|
|
966
|
+
}
|
|
967
|
+
|
|
924
968
|
/**
|
|
925
969
|
* Assemble a WorkOrder envelope. Pure given its inputs — the CLI wrapper does the disk reads.
|
|
926
970
|
* @param {object} opts - The order inputs (destructured):
|
|
@@ -1278,6 +1322,11 @@ export async function cli(rawArgv) {
|
|
|
1278
1322
|
`${missing.join(", ")}; the evaluator has nothing to cite for ${missing.length === 1 ? "that scope" : "those scopes"}`);
|
|
1279
1323
|
}
|
|
1280
1324
|
}
|
|
1325
|
+
// The graded dimensions (see resolveDimensions). A set the PO named at L0.5 travels in `--payload`
|
|
1326
|
+
// and wins; an absent or empty list is resolved from the spec, never left to the judge's default.
|
|
1327
|
+
if (operation === "evaluate" && !(Array.isArray(payloadExtra.dimensions) && payloadExtra.dimensions.length)) {
|
|
1328
|
+
payloadExtra.dimensions = resolveDimensions(cwd, slug, payloadExtra.spec_folder || specDir);
|
|
1329
|
+
}
|
|
1281
1330
|
// The launch evidence, for every lane (see launchEvidenceFor). Same rule as above: an explicit
|
|
1282
1331
|
// `--payload` value outranks the derivation.
|
|
1283
1332
|
if (operation === "evaluate" && payloadExtra.revised_checks === undefined) {
|
package/kernel/init/run.mjs
CHANGED
|
@@ -52,7 +52,7 @@
|
|
|
52
52
|
// --max-rounds N outer circuit breaker (default: 3)
|
|
53
53
|
// --attempts N inner per-scope T0 budget (default: 5)
|
|
54
54
|
// --spec-folder SHARED spec deliverable path (default: shapeup/<slug>/spec/)
|
|
55
|
-
// --dimensions comma-separated eval dimensions (default: spec
|
|
55
|
+
// --dimensions comma-separated eval dimensions (default: auto — resolved from the spec per EVAL order)
|
|
56
56
|
// --gate-answers path | preset name (see `harness gate`; recorded, not read)
|
|
57
57
|
// --breadboard path to the pitch's breadboard (default: found — see resolveBreadboard())
|
|
58
58
|
// --wall-clock-budget N deadline breaker, seconds (off by default; see `harness verify budget`)
|
|
@@ -91,10 +91,13 @@ const AUTO_LEVELS = new Set(["interactive", "auto", "unattended"]);
|
|
|
91
91
|
const LENSES = new Set(["lite", "standard", "cross-context"]);
|
|
92
92
|
|
|
93
93
|
/**
|
|
94
|
-
* The eval dimension set when the caller names none
|
|
95
|
-
*
|
|
94
|
+
* The eval dimension set when the caller names none: `auto`. The spec does not exist yet when a run
|
|
95
|
+
* opens, so the set cannot be resolved here; each evaluate order resolves it from the spec on disk
|
|
96
|
+
* (`harness compile`), with the judge's own auto-enable rules. A fixed default in this place made
|
|
97
|
+
* every orchestrated run grade `spec-conformance` alone, because an explicit list switches the
|
|
98
|
+
* judge's auto-enable off.
|
|
96
99
|
*/
|
|
97
|
-
export const DEFAULT_DIMENSIONS =
|
|
100
|
+
export const DEFAULT_DIMENSIONS = "auto";
|
|
98
101
|
|
|
99
102
|
/**
|
|
100
103
|
* Parse `--dimensions` into the set written to the ledger. Shape-validated only, NOT checked against
|
|
@@ -103,14 +106,14 @@ export const DEFAULT_DIMENSIONS = ["spec-conformance"];
|
|
|
103
106
|
* unreachable. An id with no file behind it is skipped-with-a-warning at dimension resolution, which
|
|
104
107
|
* is where that check belongs and where it can actually see the files.
|
|
105
108
|
*
|
|
106
|
-
* @param {(string|null|undefined)} raw - The comma-separated flag value; absent →
|
|
107
|
-
* @returns {string[]} Trimmed, de-duplicated ids in the caller's order
|
|
109
|
+
* @param {(string|null|undefined)} raw - The comma-separated flag value; absent → `auto`.
|
|
110
|
+
* @returns {(string[]|string)} Trimmed, de-duplicated ids in the caller's order, or `auto`.
|
|
108
111
|
* @throws {Error} If the list is empty or an entry is not a kebab-case id.
|
|
109
112
|
*/
|
|
110
113
|
export function parseDimensions(raw) {
|
|
111
|
-
if (raw === null || raw === undefined) return
|
|
114
|
+
if (raw === null || raw === undefined) return DEFAULT_DIMENSIONS;
|
|
112
115
|
const ids = String(raw).split(",").map((s) => s.trim()).filter(Boolean);
|
|
113
|
-
if (!ids.length) throw new Error("--dimensions: empty list — omit the flag to
|
|
116
|
+
if (!ids.length) throw new Error("--dimensions: empty list — omit the flag to resolve the set from the spec");
|
|
114
117
|
for (const id of ids) {
|
|
115
118
|
if (!/^[a-z][a-z0-9]*(-[a-z0-9]+)*$/.test(id)) {
|
|
116
119
|
throw new Error(`--dimensions: "${id}" is not a dimension id (kebab-case, e.g. spec-conformance, tdd-surface, integration)`);
|
package/kernel/probe/eval.mjs
CHANGED
|
@@ -35,7 +35,9 @@ import { existsSync, readFileSync, readdirSync, realpathSync } from "node:fs";
|
|
|
35
35
|
import { join, resolve, resolve as resolvePath, sep } from "node:path";
|
|
36
36
|
import { createHash } from "node:crypto";
|
|
37
37
|
import { runArgs } from "../lib/argv.mjs";
|
|
38
|
-
import { resultsDir, scopesDir, readRunId, verdictsDir, specDir } from "../lib/paths.mjs";
|
|
38
|
+
import { resultsDir, ordersDir, scopesDir, readRunId, verdictsDir, specDir } from "../lib/paths.mjs";
|
|
39
|
+
import { readBoard } from "../compile.mjs";
|
|
40
|
+
import { coveringAcs } from "./requirements.mjs";
|
|
39
41
|
|
|
40
42
|
/** Longest `reason` reported. A deviation is prose written by a worker and can run to paragraphs. */
|
|
41
43
|
const REASON_MAX = 400;
|
|
@@ -253,6 +255,59 @@ export function coverageProblem(cwd, slug, verdict) {
|
|
|
253
255
|
`each row is its own criterion, its id in the criterion or its traces_to; ungraded: ${shown}`;
|
|
254
256
|
}
|
|
255
257
|
|
|
258
|
+
/**
|
|
259
|
+
* Whether a verdict anchors any criterion to a requirement when the board says which ones it could.
|
|
260
|
+
*
|
|
261
|
+
* `traces_to` is copied from the `(covers: REQ-…)` clauses of the acceptance criteria a criterion
|
|
262
|
+
* grades, and the requirements matrix is joined along it. Two judges on the same spec and the same
|
|
263
|
+
* instruction differed: one anchored every row and the matrix read 5/5, the next anchored none and it
|
|
264
|
+
* read 0/5 over the same passes. Held only to the all-empty case: which requirement a criterion
|
|
265
|
+
* traces to is the judge's reading, but none at all, beside a board that covers some, is a copy
|
|
266
|
+
* step skipped.
|
|
267
|
+
*
|
|
268
|
+
* @param {string} cwd - Project root.
|
|
269
|
+
* @param {string} slug - Feature slug.
|
|
270
|
+
* @param {object} verdict - The WorkResult's `verdict`.
|
|
271
|
+
* @returns {(string|null)} A reason phrased for the judge, or null.
|
|
272
|
+
*/
|
|
273
|
+
export function tracesProblem(cwd, slug, verdict) {
|
|
274
|
+
if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
|
|
275
|
+
const criteria = Array.isArray(verdict.criteria) ? verdict.criteria : [];
|
|
276
|
+
if (!criteria.length) return null;
|
|
277
|
+
if (criteria.some((c) => Array.isArray(c?.traces_to) && c.traces_to.length)) return null;
|
|
278
|
+
let covered = 0;
|
|
279
|
+
try { covered = coveringAcs(readBoard(cwd, slug)).size; } catch { return null; }
|
|
280
|
+
if (!covered) return null;
|
|
281
|
+
return `the verdict anchors none of its ${criteria.length} criteria to a requirement while the board's acceptance ` +
|
|
282
|
+
`criteria cover ${covered} — copy each graded AC's (covers: REQ-…) clause into that criterion's traces_to`;
|
|
283
|
+
}
|
|
284
|
+
|
|
285
|
+
/**
|
|
286
|
+
* A PASS that grades nothing under a dimension its order named.
|
|
287
|
+
*
|
|
288
|
+
* The order names the dimensions the judge grades, and GATE L4 prints them back as what "PASS"
|
|
289
|
+
* covers. Measured: an order named four, the PASS graded criteria under three, and L4 still listed
|
|
290
|
+
* the fourth as evaluated. A PASS must grade at least one criterion under every dimension its order
|
|
291
|
+
* named; a FAIL already stops the ship, so it is not held to this.
|
|
292
|
+
*
|
|
293
|
+
* @param {string} cwd - Project root.
|
|
294
|
+
* @param {string} slug - Feature slug.
|
|
295
|
+
* @param {object} verdict - The WorkResult's `verdict`.
|
|
296
|
+
* @param {{round:(number|null)}} [opts] - The EVAL round whose order names the set; null → no check.
|
|
297
|
+
* @returns {(string|null)} Why the verdict is refused, or null.
|
|
298
|
+
*/
|
|
299
|
+
export function dimensionsProblem(cwd, slug, verdict, { round = null } = {}) {
|
|
300
|
+
if (verdict?.overall !== "PASS" || round == null) return null;
|
|
301
|
+
const named = orderDimensions(cwd, slug, round);
|
|
302
|
+
if (!named?.length) return null;
|
|
303
|
+
const criteria = Array.isArray(verdict.criteria) ? verdict.criteria : [];
|
|
304
|
+
const graded = new Set(criteria.map((c) => c?.dimension).filter(Boolean));
|
|
305
|
+
const missing = named.filter((d) => !graded.has(d));
|
|
306
|
+
if (!missing.length) return null;
|
|
307
|
+
return `the PASS grades no criterion under ${missing.join(", ")}, which the order names — grade each named ` +
|
|
308
|
+
"dimension's criteria, or FAIL the one you cannot grade with the reason";
|
|
309
|
+
}
|
|
310
|
+
|
|
256
311
|
export function citationProblem(cwd, slug, verdict, { round = null } = {}) {
|
|
257
312
|
if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
|
|
258
313
|
if (!isScoped(cwd, slug)) return null;
|
|
@@ -297,7 +352,8 @@ export function evalVerdict(cwd, slug, round) {
|
|
|
297
352
|
? `the evaluator returned ${status || "no status"}: ${first}`
|
|
298
353
|
: `status ${status || "unknown"} with no PASS/FAIL verdict`), status);
|
|
299
354
|
}
|
|
300
|
-
const problem = verdictProblem(v) || coverageProblem(cwd, slug, v) ||
|
|
355
|
+
const problem = verdictProblem(v) || coverageProblem(cwd, slug, v) || tracesProblem(cwd, slug, v)
|
|
356
|
+
|| dimensionsProblem(cwd, slug, v, { round }) || citationProblem(cwd, slug, v, { round });
|
|
301
357
|
if (problem) return unfit(problem, status, overall);
|
|
302
358
|
return {
|
|
303
359
|
found: true,
|
|
@@ -309,6 +365,22 @@ export function evalVerdict(cwd, slug, round) {
|
|
|
309
365
|
};
|
|
310
366
|
}
|
|
311
367
|
|
|
368
|
+
/**
|
|
369
|
+
* The dimensions round N's evaluate order named — the set the judge was asked to grade.
|
|
370
|
+
*
|
|
371
|
+
* @param {string} cwd - Project root.
|
|
372
|
+
* @param {string} slug - Feature slug.
|
|
373
|
+
* @param {number} round - The EVAL round (`evaluate-r<N>.json`).
|
|
374
|
+
* @returns {(string[]|null)} The order's `payload.dimensions`; null when the order or the list is absent.
|
|
375
|
+
*/
|
|
376
|
+
export function orderDimensions(cwd, slug, round) {
|
|
377
|
+
try {
|
|
378
|
+
const o = JSON.parse(readFileSync(join(ordersDir(cwd, slug), `evaluate-r${round}.json`), "utf8"));
|
|
379
|
+
const d = o?.payload?.dimensions;
|
|
380
|
+
return Array.isArray(d) && d.every((x) => typeof x === "string") ? d : null;
|
|
381
|
+
} catch { return null; }
|
|
382
|
+
}
|
|
383
|
+
|
|
312
384
|
export const ARGV_SPEC = {
|
|
313
385
|
usage: "harness.mjs probe eval --slug <slug> --round N [--cwd <dir>]",
|
|
314
386
|
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
@@ -327,6 +399,7 @@ export function cli(rawArgv) {
|
|
|
327
399
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
328
400
|
const cwd = resolve(args.cwd || process.cwd());
|
|
329
401
|
const { found, overall, status, reason, bug_count, report_path } = evalVerdict(cwd, args.slug, args.round);
|
|
330
|
-
|
|
402
|
+
const dimensions = orderDimensions(cwd, args.slug, args.round);
|
|
403
|
+
console.log(JSON.stringify({ ok: found, overall, bug_count, report_path, round: args.round, status, reason, dimensions }));
|
|
331
404
|
process.exit(found ? 0 : 1);
|
|
332
405
|
}
|
package/kernel/probe/resume.mjs
CHANGED
|
@@ -477,11 +477,16 @@ export function deriveResumeState(cwd, slug, { pluginRoot = null } = {}) {
|
|
|
477
477
|
breadboard_source: readReceipt(receipt(cwd, slug))?.breadboard?.source ?? null,
|
|
478
478
|
spec_folder: hr.spec_folder || null,
|
|
479
479
|
status: hr.status || null,
|
|
480
|
+
// The terminal status of a close nobody has taken back yet. A relaunch reopens the run before its
|
|
481
|
+
// first gate on this fact, so no gate a relaunch crosses is recorded inside a closed window.
|
|
482
|
+
closed_status: hr.closed_status ? String(hr.closed_status) : null,
|
|
480
483
|
lens: hr.lens || null,
|
|
481
484
|
stack: hr.stack || null,
|
|
482
485
|
run_cmd: hr.run_cmd || null,
|
|
483
486
|
app_url: hr.app_url || null,
|
|
484
|
-
|
|
487
|
+
// A list the PO named at L0.5, or [] — `auto`, or a ledger with no line — which the evaluate order
|
|
488
|
+
// resolves from the spec at compile time.
|
|
489
|
+
eval_dimensions: Array.isArray(hr.eval_dimensions) ? hr.eval_dimensions : [],
|
|
485
490
|
orient_dir: `.shapeup/${slug}/orient/`,
|
|
486
491
|
has_orient_artifacts: hasOrientArtifacts(cwd, slug),
|
|
487
492
|
has_spec_tree: hasSpecTree(cwd, slug, hr.spec_folder || null),
|
|
@@ -597,6 +602,16 @@ export function setRunStatus(cwd, slug, status) {
|
|
|
597
602
|
return { ok: false, path: p, status, reason: `wrote "status: ${status}" but the ledger reads "${after.status}" — the write did not take` };
|
|
598
603
|
}
|
|
599
604
|
if (reopened) {
|
|
605
|
+
// THE RUN POINTER COMES BACK WITH THE RUN. The close retired it, and nothing but `init run` wrote
|
|
606
|
+
// it, so a resumed run ran to its next close with no pointer: every hook row it produced carried
|
|
607
|
+
// no run key, and the wall-clock budget had no run to read. Restored only when absent, so a
|
|
608
|
+
// pointer another run holds is left alone.
|
|
609
|
+
try {
|
|
610
|
+
if (!existsSync(activeScope(cwd))) {
|
|
611
|
+
mkdirSync(dirname(activeScope(cwd)), { recursive: true });
|
|
612
|
+
writeFileSync(activeScope(cwd), JSON.stringify({ slug, started_at: fm.started_at ?? null, reopened_at: reopened.reopened_at }, null, 2) + "\n");
|
|
613
|
+
}
|
|
614
|
+
} catch { /* the reopen stands; an unkeyed row is degraded, not wrong */ }
|
|
600
615
|
// The breadcrumb names a run that is OVER; this one no longer is. Removed only when it names
|
|
601
616
|
// this run, so another run's close is left alone.
|
|
602
617
|
try {
|
|
@@ -656,6 +671,15 @@ export function deriveLedgerFacts(cwd, slug) {
|
|
|
656
671
|
evalRows.sort((a, b) => a.round - b.round);
|
|
657
672
|
const finalVerdict = evalRows.length ? evalRows[evalRows.length - 1].overall : "not-evaluated";
|
|
658
673
|
const decisions = [];
|
|
674
|
+
// WHICH LAUNCH TOOK EACH DECISION. A relaunch re-crosses the gates it fast-forwards, and the table
|
|
675
|
+
// listed each of them twice with nothing to tell a second launch from a double sign-off. Every
|
|
676
|
+
// reopen leaves its timestamp in `prior_closes`; a row taken after the n-th one is launch n+1's.
|
|
677
|
+
let reopens = [];
|
|
678
|
+
try {
|
|
679
|
+
const pc = parseFrontmatter(readFileSync(harnessRun(cwd, slug), "utf8")).prior_closes;
|
|
680
|
+
reopens = [...String(pc ?? "").matchAll(/\(reopened ([^)]+)\)/g)].map((m) => m[1]).sort();
|
|
681
|
+
} catch { /* no ledger, no reopen */ }
|
|
682
|
+
const launchOf = (at) => 1 + reopens.filter((t) => typeof at === "string" && at > t).length;
|
|
659
683
|
const gp = gates(cwd, slug);
|
|
660
684
|
if (existsSync(gp)) {
|
|
661
685
|
for (const line of readFileSync(gp, "utf8").split("\n")) {
|
|
@@ -663,7 +687,7 @@ export function deriveLedgerFacts(cwd, slug) {
|
|
|
663
687
|
try {
|
|
664
688
|
const g = JSON.parse(line);
|
|
665
689
|
// Gate rows carry the run key; a row from an earlier run over the same slug is its history.
|
|
666
|
-
if (!runId || !g?.run_id || g.run_id === runId) decisions.push(g);
|
|
690
|
+
if (!runId || !g?.run_id || g.run_id === runId) decisions.push(reopens.length ? { ...g, launch: launchOf(g.at) } : g);
|
|
667
691
|
} catch { /* a torn line proves nothing */ }
|
|
668
692
|
}
|
|
669
693
|
}
|
|
@@ -733,7 +757,7 @@ function writeCloseLines(body, { status, closedAt, cause, derived = null }) {
|
|
|
733
757
|
if (/^final_verdict:.*$/m.test(out)) out = out.replace(/^final_verdict:.*$/m, `final_verdict: ${derived.final_verdict}`);
|
|
734
758
|
if (typeof derived.rounds_used === "number" && /^rounds_used:.*$/m.test(out)) out = out.replace(/^rounds_used:.*$/m, `rounds_used: ${derived.rounds_used}`);
|
|
735
759
|
out = rewriteTable(out, /(\| Phase \| Round \| Result \| Duration \| Notes \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, derived.roundRows, "| Init");
|
|
736
|
-
const decisionRows = derived.decisions.map((g) => `| ${g.gate ?? "?"} | ${g.decision ?? g.status ?? "?"} | ${g.source ?? "?"} | ${String(g.note ?? "").replace(/\|/g, "\\|").replace(/\s+/g, " ").slice(0, 160)} |`);
|
|
760
|
+
const decisionRows = derived.decisions.map((g) => `| ${g.gate ?? "?"} | ${g.decision ?? g.status ?? "?"} | ${g.source ?? "?"} | ${g.launch > 1 ? `launch ${g.launch} (after a reopen) — ` : ""}${String(g.note ?? "").replace(/\|/g, "\\|").replace(/\s+/g, " ").slice(0, 160)} |`);
|
|
737
761
|
out = rewriteTable(out, /(\| Gate \| Decision \| Source \| Note \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, decisionRows);
|
|
738
762
|
}
|
|
739
763
|
return out;
|
package/kernel/reduce/ingest.mjs
CHANGED
|
@@ -37,7 +37,7 @@ import { fileURLToPath } from "node:url";
|
|
|
37
37
|
import { validate } from "../verify/envelope.mjs";
|
|
38
38
|
import { runArgs } from "../lib/argv.mjs";
|
|
39
39
|
import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId, qaDir } from "../lib/paths.mjs";
|
|
40
|
-
import { citationProblem, coverageProblem, verdictProblem } from "../probe/eval.mjs";
|
|
40
|
+
import { citationProblem, coverageProblem, dimensionsProblem, tracesProblem, verdictProblem } from "../probe/eval.mjs";
|
|
41
41
|
import { huntProblem } from "../probe/hunt.mjs";
|
|
42
42
|
|
|
43
43
|
const HERE = dirname(fileURLToPath(import.meta.url));
|
|
@@ -702,6 +702,7 @@ export async function cli(rawArgv) {
|
|
|
702
702
|
const evalSlug = String(result.order_id).split("/")[0];
|
|
703
703
|
const evalRound = Number((String(result.order_id).match(/-r(\d+)$/) || [])[1]) || null;
|
|
704
704
|
const problem = verdictProblem(result.verdict) || coverageProblem(cwd, evalSlug, result.verdict)
|
|
705
|
+
|| tracesProblem(cwd, evalSlug, result.verdict) || dimensionsProblem(cwd, evalSlug, result.verdict, { round: evalRound })
|
|
705
706
|
|| citationProblem(cwd, evalSlug, result.verdict, { round: evalRound });
|
|
706
707
|
if (problem) {
|
|
707
708
|
console.error(`ingest-result: result refused — ${problem}.`);
|
package/kernel/report/facts.mjs
CHANGED
|
@@ -113,7 +113,7 @@ export function runRow({ receipt, ledger = null, runId = null, rounds = null })
|
|
|
113
113
|
max_rounds: num(c.max_rounds),
|
|
114
114
|
attempt_budget: num(c.attempt_budget),
|
|
115
115
|
wall_clock_budget_s: num(c.wall_clock_budget_s),
|
|
116
|
-
eval_dimensions: Array.isArray(c.eval_dimensions) ? c.eval_dimensions.join(" ") : null,
|
|
116
|
+
eval_dimensions: Array.isArray(c.eval_dimensions) ? c.eval_dimensions.join(" ") : (typeof c.eval_dimensions === "string" ? c.eval_dimensions : null),
|
|
117
117
|
// Copied from the ledger, never re-derived: the run's own status line is the harness's answer,
|
|
118
118
|
// and a read plane that recomputed it would be asserting a second one.
|
|
119
119
|
status: fm.status ?? null,
|
|
@@ -2423,7 +2423,7 @@
|
|
|
2423
2423
|
"items": {
|
|
2424
2424
|
"type": "string"
|
|
2425
2425
|
},
|
|
2426
|
-
"description": "spec-evaluator: the active dimension set
|
|
2426
|
+
"description": "spec-evaluator: the active dimension set. harness compile fills it on every evaluate order: the run's named set, or the set resolved from the spec with the judge's auto-enable rules. Absent only on a standalone order."
|
|
2427
2427
|
},
|
|
2428
2428
|
"run_cmd": {
|
|
2429
2429
|
"type": "string",
|
|
@@ -2586,6 +2586,13 @@
|
|
|
2586
2586
|
],
|
|
2587
2587
|
"description": "The ledger's stored status, reported for the readers that legitimately hold a MID_RUN set over it (harness reduce snapshot, and the ship report's census). NOT a resume predicate: no phase decision in shapeup-run.js may branch on this field."
|
|
2588
2588
|
},
|
|
2589
|
+
"closed_status": {
|
|
2590
|
+
"type": [
|
|
2591
|
+
"string",
|
|
2592
|
+
"null"
|
|
2593
|
+
],
|
|
2594
|
+
"description": "The ledger's closed_status: the terminal status of a close nobody has taken back, null on a run that is open. A relaunch reads it once, to reopen the run before its first gate resolves. Not a phase predicate either."
|
|
2595
|
+
},
|
|
2589
2596
|
"lens": {
|
|
2590
2597
|
"type": [
|
|
2591
2598
|
"string",
|
|
@@ -2868,6 +2875,13 @@
|
|
|
2868
2875
|
"rounds_used": {
|
|
2869
2876
|
"type": "integer"
|
|
2870
2877
|
},
|
|
2878
|
+
"dims_evaluated": {
|
|
2879
|
+
"type": "array",
|
|
2880
|
+
"items": {
|
|
2881
|
+
"type": "string"
|
|
2882
|
+
},
|
|
2883
|
+
"description": "The dimensions the last evaluate order named: the set the PO chose at L0.5, or the one resolved from the spec."
|
|
2884
|
+
},
|
|
2871
2885
|
"dims_not_evaluated": {
|
|
2872
2886
|
"type": "array",
|
|
2873
2887
|
"items": {
|
package/package.json
CHANGED
|
@@ -37,7 +37,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
|
|
|
37
37
|
|---|---|
|
|
38
38
|
| `payload.spec_folder` | The committed grading truth: `usecases/` + `domain-model.md` (+ `contracts/`, `scope-summary.md`, `_index.md`). No `usecases/` → HARD STOP, nothing to grade against |
|
|
39
39
|
| `payload.feature` | Feature slug — scopes the probe and names the report |
|
|
40
|
-
| `payload.dimensions[]` | The active dimension set
|
|
40
|
+
| `payload.dimensions[]` | The active dimension set. An orchestrated order always carries it: the set the PO named, or the one resolved from the spec with the auto-enable rules below. Grade exactly this set, at least one criterion under each (a PASS that leaves a named dimension ungraded is refused and sent back once). Absent (standalone only) → apply the rules below yourself |
|
|
41
41
|
| `payload.run_cmd` | How to start the running app. On a stack where building and launching are different acts it is only the build — prefer `payload.launch_cmd` when the order carries one. Absent standalone → ask; absent orchestrated with no `launch_cmd` either → ESCALATE, do not guess |
|
|
42
42
|
| `payload.launch_cmd` | The project profile's launch probe: installs the built artifact, starts it and asserts the first screen. This is how the app is brought up for `[ui]` probing. A non-zero exit is a finding to cite, not a reason to try another way. Absent → the profile declares none; fall back to `run_cmd` |
|
|
43
43
|
| `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
|
|
@@ -52,7 +52,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
|
|
|
52
52
|
(Steps, Error Cases, Invariants, Test Surface) and `domain-model.md` — never against a task
|
|
53
53
|
file's own AC paraphrase. Task boards are LOCAL, regenerable bookkeeping the judge never touches.
|
|
54
54
|
|
|
55
|
-
**Dimension resolution (
|
|
55
|
+
**Dimension resolution (standalone; an orchestrated order arrives resolved).** Base `[spec-conformance]` + always-on `tdd-surface` +
|
|
56
56
|
`integration` (`.be`/`.e2e`); auto-enable `completeness` when any UC has `## Invariants`,
|
|
57
57
|
`test-surface-conformance` when any UC has `## Test Surface`; an explicit `dimensions[]` list
|
|
58
58
|
overrides. Each active dimension's file must satisfy `references/dimension-contract.md` — a
|
|
@@ -188,7 +188,7 @@ clause yields no anchor — leave the array empty rather than guessing, and neve
|
|
|
188
188
|
supply one. This changes nothing you grade: the anchor is a navigation path, never a grading input,
|
|
189
189
|
and a criterion passes or fails on its evidence exactly as before. It matters downstream because
|
|
190
190
|
the requirement matrix at GATE L4 and the census at GATE H are projected from these anchors; a
|
|
191
|
-
verdict that drops them grades the build and says nothing about what the pitch asked for.
|
|
191
|
+
verdict that drops them grades the build and says nothing about what the pitch asked for. Ingest refuses a verdict that anchors no criterion at all while the board's ACs carry `covers:` clauses, and you are sent back once.
|
|
192
192
|
|
|
193
193
|
**Every FAIL criterion's `evidence` MUST carry a `file:line` locator** — schema-enforced, not
|
|
194
194
|
advice: the envelope is validated against `work-result.schema.json` at ingest and a locatorless
|
|
@@ -16,14 +16,11 @@ what `harness init run` writes. Working around the harness is not an exemption:
|
|
|
16
16
|
switch this gate off and no longer does. Loading these instructions is not running them.
|
|
17
17
|
|
|
18
18
|
**Step 1 — open the run.** Write the requirement to a file first, then pass the path — a
|
|
19
|
-
multi-line requirement inlined into a shell argument is where this step goes wrong
|
|
20
|
-
|
|
19
|
+
multi-line requirement inlined into a shell argument is where this step goes wrong. Keep the command
|
|
20
|
+
on ONE line: a grant matches a single-line command, and a `\`-continued one comes back "requires approval":
|
|
21
21
|
|
|
22
22
|
```bash
|
|
23
|
-
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" init run
|
|
24
|
-
--slug <slug-from-the-request> --intake-file <path/to/the/requirement.md> \
|
|
25
|
-
--auto-level <interactive|auto|unattended> \
|
|
26
|
-
[--dimensions <a,b>] [--gate-answers <ci|guarded|path.json>] [--wall-clock-budget <seconds>] [--max-rounds 3] [--breadboard <path>]
|
|
23
|
+
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" init run --slug <slug-from-the-request> --intake-file <path/to/the/requirement.md> --auto-level <interactive|auto|unattended> [--dimensions <a,b>] [--gate-answers <ci|guarded|path.json>] [--wall-clock-budget <seconds>] [--max-rounds 3] [--breadboard <path>]
|
|
27
24
|
```
|
|
28
25
|
|
|
29
26
|
**After a compaction, or in a fresh session over an open run, re-derive before you act.** One
|
|
@@ -102,7 +99,7 @@ failed, launch from the install path and have the operator `/add-dir` the plugin
|
|
|
102
99
|
| `status` | What the workflow is telling you | What you do |
|
|
103
100
|
|---|---|---|
|
|
104
101
|
| `paused` | A gate resolved "ask" — `paused_at` names it, `block` is composed and ready | Emit `block` **verbatim** (never re-summarise it — that is the paraphrase channel this design exists to close). Put it to the PO, get a decision. Write it to `.shapeup/<slug>/gate-answers.json` (`{"version":1,"preset":"custom","answers":{"<paused_at>":{"decision":"<answer>"}}}`, merging with any prior gate's answer already there). **Relaunch the SAME `Workflow` call, same `args`** — the fast-forward re-derives from disk and re-dispatches nothing already done (verify: `orders/` minus `results/` is empty before it proceeds) |
|
|
105
|
-
| `aborted` | A gate resolved "abort", or a hard stop (spec-lint red, scope-hammer CANNOT SHIP) | Report `aborted_at` + `reason` to the PO. Do not relaunch without a human decision
|
|
102
|
+
| `aborted` | A gate resolved "abort", or a hard stop (spec-lint red, scope-hammer CANNOT SHIP) | Report `aborted_at` + `reason` to the PO. Do not relaunch without a human decision. A relaunch RESUMES the closed run (it reopens, keeping the abort under `prior_closes`); `--force` discards its history and only the PO may ask for it |
|
|
106
103
|
| `gate_h` | A circuit breaker tripped (`breaker`: outer \| inner \| deadline) — `green_scopes` shipped nothing, `hammer_proposals` needs a census | Dispatch a fresh Agent (model: exec): `Skill(shapeup-sdlc-plugin:scope-hammer) --slug <slug> --breaker <breaker> [--scope <id>]` for the census + cut list, put the PO's decision to `references/gates.md` GATE H, then close out via Step 4 below |
|
|
107
104
|
| `shipped` | The board's final round passed EVAL, QA ran, GATE H accepted the cut list, `report` names the frozen `shapeup/<slug>/REPORT.md` | Go straight to Step 4 |
|
|
108
105
|
|
|
@@ -120,7 +117,7 @@ now and RESOLVES the gate (`references/gates.md` GATE L4). Either way, emit:
|
|
|
120
117
|
⏸ GATE L4 — Ship Sign-Off
|
|
121
118
|
Feature : [slug] — [SHIPPED (deployed) | BUILT & VERIFIED — deploy pending (PO)]
|
|
122
119
|
Rounds : [rounds_used]
|
|
123
|
-
Verdict : [verdict] (dims: [
|
|
120
|
+
Verdict : [verdict] (dims: [dims_evaluated]; not evaluated: [dims_not_evaluated])
|
|
124
121
|
QA : [qa_findings] findings | skipped
|
|
125
122
|
Ledger : harness-run.md
|
|
126
123
|
```
|
|
@@ -60,10 +60,12 @@ Collect (explicit — never inferred):
|
|
|
60
60
|
spec_folder = shapeup/<slug>/spec/ (the deliverable arg passed to ba/eval/exec)
|
|
61
61
|
L0.3 lens: lite | standard | cross-context (passed to planner at step 8)
|
|
62
62
|
L0.4 stack hint (e.g. "pnpm, Next 16 web :3000") — aims orient's code-surface sweeps + run commands
|
|
63
|
-
L0.5 eval dimensions: default
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
63
|
+
L0.5 eval dimensions: default `auto` — each EVAL order resolves the set from the spec
|
|
64
|
+
(spec-conformance + tdd-surface always; integration for .be/.e2e tasks; completeness
|
|
65
|
+
when a UC has Invariants; test-surface-conformance when one has a Test Surface).
|
|
66
|
+
Name a set only if the user asks. It must reach `harness init run --dimensions <a,b>`
|
|
67
|
+
— it is recorded in the ledger's `eval_dimensions:` line and replaces the resolved
|
|
68
|
+
set, so a set agreed in conversation and not passed to the flag grades nothing. Shipped ids:
|
|
67
69
|
spec-conformance, tdd-surface, integration, completeness, test-surface-conformance
|
|
68
70
|
(security + performance ship disabled). Whatever is left out is reported at L4 as
|
|
69
71
|
`dims_not_evaluated` — "shipped" never silently means "verified for all".
|
|
@@ -161,7 +163,7 @@ Feature : [slug] (kicked-off pitch: [path])
|
|
|
161
163
|
Intake lang : [English | translated via /translator → <name>.en.md]
|
|
162
164
|
Appetite : [~1 week | ~2 weeks | ~6 weeks | ⚠️ missing — scope uncapped]
|
|
163
165
|
Spec folder : [path] (lens: [lite|standard])
|
|
164
|
-
Eval dims : [
|
|
166
|
+
Eval dims : [auto | the named set] max_rounds: [N, appetite-informed] auto: [interactive|auto|unattended]
|
|
165
167
|
Run commands : [web: ... | api: ... | mobile: ...] (run_cmd → the round build gate, every round before EVAL)
|
|
166
168
|
Build gate : build_probe [set | —] launch_probe [set | — ⚠ mobile: the install/launch risk has no owner]
|
|
167
169
|
Model matrix : orch=[model] exec=[model] eval=[model] qa=[model] digester=[script|sonnet] (source: [flags|settings.local|settings.json|default])
|
|
@@ -555,7 +557,7 @@ S.7 Export the run's records → one keyed dataset, before the trace is superse
|
|
|
555
557
|
⏸ GATE L4 — Ship Sign-Off
|
|
556
558
|
Feature : [slug] — [SHIPPED (deployed) | BUILT & VERIFIED — deploy pending (PO)]
|
|
557
559
|
Rounds : [r] (build+eval cycles)
|
|
558
|
-
Verdict : PASS (dims: [
|
|
560
|
+
Verdict : PASS (dims: [dims_evaluated]; not evaluated: [dims_not_evaluated])
|
|
559
561
|
QA : [hunt done — N findings, M promoted+fixed, rest ~ | skipped (--no-qa) | n/a (pre-QA spec)]
|
|
560
562
|
Requirements: [15/17 PASS · 1 CUT (PO) · 1 no evidence (REQ-12 ← R12) | n/a (no registry)]
|
|
561
563
|
Ledger : harness-run.md
|
|
@@ -692,7 +692,7 @@ type: harness-run
|
|
|
692
692
|
feature: [slug]
|
|
693
693
|
spec_folder: [path to SHARED spec deliverable, e.g. shapeup/<slug>/spec/]
|
|
694
694
|
lens: lite | standard | cross-context
|
|
695
|
-
eval_dimensions:
|
|
695
|
+
eval_dimensions: auto # or the list from GATE L0.5 (init-run --dimensions); `auto` → each EVAL order resolves the set from the spec
|
|
696
696
|
max_rounds: 3
|
|
697
697
|
auto_level: interactive | auto | unattended
|
|
698
698
|
status: orienting | mapping | building | evaluating | shipped | escalated | aborted
|
|
@@ -411,7 +411,7 @@ const RESUME = {
|
|
|
411
411
|
properties: {
|
|
412
412
|
intake_path: nullable("string"), spec_folder: nullable("string"), orient_dir: nullable("string"),
|
|
413
413
|
breadboard_path: nullable("string"), breadboard_source: nullable("string"),
|
|
414
|
-
project_profile_path: nullable("string"), status: nullable("string"),
|
|
414
|
+
project_profile_path: nullable("string"), status: nullable("string"), closed_status: nullable("string"),
|
|
415
415
|
lens: nullable("string"), stack: nullable("string"),
|
|
416
416
|
run_cmd: nullable("string"), app_url: nullable("string"),
|
|
417
417
|
eval_dimensions: { type: "array", items: { type: "string" } },
|
|
@@ -612,6 +612,8 @@ const EVAL_VERDICT = {
|
|
|
612
612
|
bug_count: nullable("integer"),
|
|
613
613
|
report_path: nullable("string"),
|
|
614
614
|
round: { type: "integer" },
|
|
615
|
+
// The dimensions round N's evaluate order named — what "PASS" actually covers.
|
|
616
|
+
dimensions: { type: ["array", "null"], items: { type: "string" } },
|
|
615
617
|
// Why the round holds no verdict it may act on — the evaluator's own first deviation when it
|
|
616
618
|
// refused, or what is structurally wrong with the verdict it returned. Null when `ok`.
|
|
617
619
|
status: nullable("string"),
|
|
@@ -1360,8 +1362,19 @@ if (!rs) {
|
|
|
1360
1362
|
return await withWarnings(aborted("probe", "the fast-forward derivation returned no state — refusing to re-dispatch a run that may already be in progress"));
|
|
1361
1363
|
}
|
|
1362
1364
|
|
|
1365
|
+
// A RELAUNCH OVER A CLOSED RUN REOPENS IT FIRST. The first status write is what takes a close back,
|
|
1366
|
+
// and a launch that fast-forwards every planning phase made that write only at BUILD, so the gates
|
|
1367
|
+
// it crossed on the way were recorded against a run whose ledger still read closed. Bookkeeping,
|
|
1368
|
+
// not a phase decision: the phases below still branch on artifacts alone.
|
|
1369
|
+
if (rs.closed_status) await setRunStatus("orienting", "Orient");
|
|
1370
|
+
|
|
1363
1371
|
const specFolder = rs.spec_folder || `shapeup/${slug}/spec/`;
|
|
1364
|
-
|
|
1372
|
+
// A set the PO named at L0.5, or null: `harness compile` then resolves it from the spec for each
|
|
1373
|
+
// evaluate order. A fixed fallback here would travel as an explicit list and switch off the judge's
|
|
1374
|
+
// own auto-enable, which is how every run once graded spec-conformance alone.
|
|
1375
|
+
const evalDims = rs.eval_dimensions?.length ? rs.eval_dimensions : null;
|
|
1376
|
+
// What the judge was actually asked to grade, read back off the newest evaluate order.
|
|
1377
|
+
let gradedDims = evalDims || [];
|
|
1365
1378
|
|
|
1366
1379
|
// ---- ORIENT + GATE L1a ----------------------------------------------------------------------
|
|
1367
1380
|
let spikedArea = "~", spikeResult = "~", riskiest = [];
|
|
@@ -1691,6 +1704,7 @@ if (lastEval) {
|
|
|
1691
1704
|
const prior = await query(`probe eval --slug ${slug} --round ${lastEval}`, EVAL_VERDICT, "Eval", `ff:eval-r${lastEval}`);
|
|
1692
1705
|
if (prior?.ok && prior.overall === "PASS") {
|
|
1693
1706
|
verdict = "pass";
|
|
1707
|
+
if (Array.isArray(prior.dimensions) && prior.dimensions.length) gradedDims = prior.dimensions;
|
|
1694
1708
|
round = lastEval;
|
|
1695
1709
|
log(`EVAL — round ${lastEval} already returned PASS on disk, fast-forwarding past the build/eval loop`);
|
|
1696
1710
|
}
|
|
@@ -1949,7 +1963,7 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1949
1963
|
// The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
|
|
1950
1964
|
// artifact and the project profile, so the judge gets the launch evidence without this script
|
|
1951
1965
|
// having to carry it.
|
|
1952
|
-
payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
|
|
1966
|
+
payload: { ...(evalDims ? { dimensions: evalDims } : {}), run_cmd: rs.run_cmd, round },
|
|
1953
1967
|
extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself." + note,
|
|
1954
1968
|
}),
|
|
1955
1969
|
// The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
|
|
@@ -1968,6 +1982,7 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1968
1982
|
return await withWarnings(diedAt("L3", { __failed: `verdict:r${round}: no verdict this round can act on — ${ev.reason || `status ${ev.status || "unknown"}`}` }));
|
|
1969
1983
|
}
|
|
1970
1984
|
verdict = ev.overall === "PASS" ? "pass" : "fail";
|
|
1985
|
+
if (Array.isArray(ev.dimensions) && ev.dimensions.length) gradedDims = ev.dimensions;
|
|
1971
1986
|
findings = e.findings || [];
|
|
1972
1987
|
}
|
|
1973
1988
|
|
|
@@ -2088,7 +2103,8 @@ return await withWarnings({
|
|
|
2088
2103
|
// human answering that gate the run was graded when it never was.
|
|
2089
2104
|
verdict,
|
|
2090
2105
|
rounds_used: round,
|
|
2091
|
-
|
|
2106
|
+
dims_evaluated: gradedDims,
|
|
2107
|
+
dims_not_evaluated: ALL_DIMS.filter((d) => !gradedDims.includes(d)),
|
|
2092
2108
|
qa_findings: qaFindings,
|
|
2093
2109
|
qa: qaState,
|
|
2094
2110
|
report: REPORT_PATH,
|