shapeup-sdlc 3.9.2 → 3.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +3 -3
- package/kernel/compile.mjs +44 -1
- package/kernel/init/run.mjs +17 -3
- package/kernel/probe/resume.mjs +43 -4
- package/kernel/probe/t0.mjs +24 -3
- package/kernel/reduce/board.mjs +11 -2
- package/kernel/reduce/ship.mjs +34 -1
- package/kernel/schemas/domain.schema.json +17 -0
- package/kernel/verify/build.mjs +105 -2
- package/package.json +1 -1
- package/skills/spec-evaluator/SKILL.md +7 -2
- package/skills/tech-lead/SKILL.md +3 -3
- package/skills/tech-lead/references/protocol.md +11 -1
- package/skills/tech-lead/workflows/shapeup-run.js +40 -1
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.11.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -13,7 +13,7 @@ Skills and commands are named short throughout this file; every one of them reso
|
|
|
13
13
|
- **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
|
|
14
14
|
- GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
|
|
15
15
|
- Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
|
|
16
|
-
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
|
|
16
|
+
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it. A relaunch that resumes a run which *did* close reopens it on the record: the earlier close is kept under `prior_closes`, the close fields clear, and the run's next ending is recorded as its own — so a run that aborted, was resumed and then shipped reads `shipped` with the abort beside it, never an abort under a `shipped` status. Opening a run over one that closed says it is closed, not open.
|
|
17
17
|
- **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
|
|
18
18
|
- The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
|
|
19
19
|
|
|
@@ -36,12 +36,12 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
38
|
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
39
|
-
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest |
|
|
39
|
+
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade |
|
|
40
40
|
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
|
|
41
41
|
|
|
42
42
|
✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
|
|
43
43
|
|
|
44
|
-
✧ **Two refusals, and they answer different questions.** Opening a run refuses outright when the worker skills are not on disk — the roster comes from the schema that defines it, so it cannot drift from a hand-kept list. That alone passes green on the two states that actually happen (installed but disabled, or a different version loaded), so the run's first act is one live canary dispatch, and the evidence is the hook layer's rather than the sub-agent's account of it. A run that cannot reach its workers stops before it spends anything, instead of reporting phases complete while none of the shipped craft was applied.
|
|
44
|
+
✧ **Two refusals, and they answer different questions.** Opening a run refuses outright when the worker skills are not on disk — the roster comes from the schema that defines it, so it cannot drift from a hand-kept list. That alone passes green on the two states that actually happen (installed but disabled, or a different version loaded), so the run's first act is one live canary dispatch, and the evidence is the hook layer's rather than the sub-agent's account of it. A run that cannot reach its workers stops before it spends anything, instead of reporting phases complete while none of the shipped craft was applied. The same Preflight then runs the profile's `build_probe` and `launch_probe` once, from a sub-agent through the session's own grant, and writes nothing: a missing package install, a probe outside the grant or an absent device is reported before planning is paid for, not at the first round gate. It warns rather than refuses — a baseline can be red for a reason the feature is meant to fix — and whether a permission exists is judged by that run, never by reading a settings file.
|
|
45
45
|
|
|
46
46
|
|
|
47
47
|
### QA Edge Hunt (`/qa-edge-hunter`, post-PASS, pre-ship)
|
package/kernel/compile.mjs
CHANGED
|
@@ -44,7 +44,7 @@ import { writeActiveOrder } from "./probe/resume.mjs";
|
|
|
44
44
|
import { greenVerdict } from "./probe/t0.mjs";
|
|
45
45
|
import { attemptEvidence, readReceipts } from "./probe/attempts.mjs";
|
|
46
46
|
import { readLegs } from "./probe/leg.mjs";
|
|
47
|
-
import { latestRoundBuild } from "./verify/build.mjs";
|
|
47
|
+
import { latestRoundBuild, latestRoundBuildFile, launchProbeFor } from "./verify/build.mjs";
|
|
48
48
|
// The SAME matcher the sandbox hook enforces with. "Is this cited file inside this scope's
|
|
49
49
|
// substrate" has to mean exactly what the guard means, or a bug is addressed to a scope that is
|
|
50
50
|
// then denied the write that fixes it.
|
|
@@ -748,6 +748,42 @@ export function t0ArtifactsFor(cwd, slug, round) {
|
|
|
748
748
|
return { artifacts, missing };
|
|
749
749
|
}
|
|
750
750
|
|
|
751
|
+
// --- the launch evidence the judge grades `[ui]` rows against ---------------------------------
|
|
752
|
+
//
|
|
753
|
+
// WHY THE KERNEL DERIVES IT. A `[ui]` criterion is graded on the RUNNING app, and the evaluator's
|
|
754
|
+
// contract for getting one is `payload.run_cmd`; absent, an orchestrated evaluator escalates rather
|
|
755
|
+
// than guess. The ledger's `run_cmd` is what the round build gate runs FIRST, as the build — on a
|
|
756
|
+
// toolchain where building and launching are different acts it is a build, and the ledger carries
|
|
757
|
+
// none at all when nobody pinned one. Meanwhile the gate had already run the project's launch probe
|
|
758
|
+
// (install, start, assert the first screen) moments earlier, recorded its output, and left the
|
|
759
|
+
// app up. None of that reached the order: the gate's artifact was forwarded only when RED, as the
|
|
760
|
+
// next round's bug list, so a green launch was proven and then unavailable to the one reader who
|
|
761
|
+
// needed it. Every `[ui]` row was graded "no evidence on the running app" over a build that launched.
|
|
762
|
+
//
|
|
763
|
+
// Two fields, both derived from disk, both optional, both absent when there is nothing to say —
|
|
764
|
+
// the same non-regression rule as `t0_artifacts`: a project that declares no launch probe compiles
|
|
765
|
+
// exactly the order it always did.
|
|
766
|
+
|
|
767
|
+
/**
|
|
768
|
+
* What an evaluate order should carry about launching the app.
|
|
769
|
+
*
|
|
770
|
+
* @param {string} cwd - Project root.
|
|
771
|
+
* @param {string} slug - Feature slug.
|
|
772
|
+
* @param {number} [round] - The round being evaluated. Omitted, the newest gate artifact of any
|
|
773
|
+
* round — a standalone evaluation has no round.
|
|
774
|
+
* @returns {{build_gate?: string, launch_cmd?: string}} `build_gate`: repo-relative path of this
|
|
775
|
+
* run's newest round build gate artifact (each step's exit and output tail). `launch_cmd`: the
|
|
776
|
+
* profile's launch probe. A key is omitted, never null, when its source is absent.
|
|
777
|
+
*/
|
|
778
|
+
export function launchEvidenceFor(cwd, slug, round) {
|
|
779
|
+
const out = {};
|
|
780
|
+
const gate = latestRoundBuildFile(cwd, slug, round);
|
|
781
|
+
if (gate) out.build_gate = relLocal(slug, "build", basename(gate.path));
|
|
782
|
+
const cmd = launchProbeFor(cwd, slug);
|
|
783
|
+
if (cmd) out.launch_cmd = cmd;
|
|
784
|
+
return out;
|
|
785
|
+
}
|
|
786
|
+
|
|
751
787
|
/**
|
|
752
788
|
* Assemble a WorkOrder envelope. Pure given its inputs — the CLI wrapper does the disk reads.
|
|
753
789
|
* @param {object} opts - The order inputs (destructured):
|
|
@@ -1093,6 +1129,13 @@ export async function cli(rawArgv) {
|
|
|
1093
1129
|
`${missing.join(", ")}; the evaluator has nothing to cite for ${missing.length === 1 ? "that scope" : "those scopes"}`);
|
|
1094
1130
|
}
|
|
1095
1131
|
}
|
|
1132
|
+
// The launch evidence, for every lane (see launchEvidenceFor). Same rule as above: an explicit
|
|
1133
|
+
// `--payload` value outranks the derivation.
|
|
1134
|
+
if (operation === "evaluate") {
|
|
1135
|
+
const ev = launchEvidenceFor(cwd, slug, round);
|
|
1136
|
+
if (ev.build_gate !== undefined && payloadExtra.build_gate === undefined) payloadExtra.build_gate = ev.build_gate;
|
|
1137
|
+
if (ev.launch_cmd !== undefined && payloadExtra.launch_cmd === undefined) payloadExtra.launch_cmd = ev.launch_cmd;
|
|
1138
|
+
}
|
|
1096
1139
|
|
|
1097
1140
|
const order = compileOrder({
|
|
1098
1141
|
slug, worker, operation, round, attempt, scope, tasks, decisions, digestedErrors, trialHistory, bugs,
|
package/kernel/init/run.mjs
CHANGED
|
@@ -75,12 +75,12 @@ import { join, dirname, resolve, relative, sep } from "node:path";
|
|
|
75
75
|
import { createHash } from "node:crypto";
|
|
76
76
|
import { decideLane, treeSize } from "./fit.mjs";
|
|
77
77
|
import { runArgs } from "../lib/argv.mjs";
|
|
78
|
-
import { uncoerce } from "../lib/contract.mjs";
|
|
78
|
+
import { uncoerce, splitFrontmatter } from "../lib/contract.mjs";
|
|
79
79
|
import { deriveSnapshot } from "../reduce/snapshot.mjs";
|
|
80
80
|
import { mintRunId } from "../lib/paths.mjs";
|
|
81
81
|
import {
|
|
82
82
|
localRoot, activeScope, activeOrder, globLocal, globShared, ordersDir, resultsDir,
|
|
83
|
-
workflowsStage, globWorkflowsStage, sharedRoot, shapingDir, breadboard as stagedBreadboard,
|
|
83
|
+
workflowsStage, globWorkflowsStage, sharedRoot, shapingDir, breadboard as stagedBreadboard, harnessRun,
|
|
84
84
|
} from "../lib/paths.mjs";
|
|
85
85
|
import { parseBreadboard, hasBreadboardTables, idCounts } from "../lib/breadboard.mjs";
|
|
86
86
|
import { resolveWorkers } from "../verify/skills.mjs";
|
|
@@ -603,8 +603,22 @@ export function cli(rawArgv) {
|
|
|
603
603
|
if (existsSync(receiptPath) && !args.force) {
|
|
604
604
|
let resume = null;
|
|
605
605
|
try { resume = deriveSnapshot(cwd); } catch { /* a broken run must still produce the refusal */ }
|
|
606
|
+
// A receipt says a run EXISTS, not that it is open: a terminal close leaves the receipt in place.
|
|
607
|
+
// Calling a closed run "open" hid the one fact that changes what resuming it means.
|
|
608
|
+
let closed = null;
|
|
609
|
+
try {
|
|
610
|
+
const fm = splitFrontmatter(readFileSync(harnessRun(cwd, slug), "utf8")).meta || {};
|
|
611
|
+
if (fm.closed_status && fm.closed_status !== "~") closed = { status: fm.closed_status, at: fm.closed_at, cause: fm.close_cause };
|
|
612
|
+
} catch { /* no ledger — the receipt alone decides, as before */ }
|
|
606
613
|
fail(3, [
|
|
607
|
-
|
|
614
|
+
closed
|
|
615
|
+
? `✋ init-run: this run is CLOSED as "${closed.status}" at ${closed.at ?? "?"} — receipt exists at ${receiptPath}.`
|
|
616
|
+
: `✋ init-run: a run is ALREADY OPEN — receipt exists at ${receiptPath}.`,
|
|
617
|
+
...(closed ? [
|
|
618
|
+
` Its close: ${closed.cause && closed.cause !== "~" ? closed.cause : "no cause recorded"}`,
|
|
619
|
+
" Resuming it REOPENS it: the first phase that runs moves the ledger off the closed state, keeps",
|
|
620
|
+
" this close under prior_closes, and the run's next terminal close is recorded as its own.",
|
|
621
|
+
] : []),
|
|
608
622
|
"",
|
|
609
623
|
"Do NOT re-initialise and do NOT restart the pipeline from phase 1. Re-opening would discard",
|
|
610
624
|
"the round history the circuit breaker counts against, and the board, ledger and receipt below",
|
package/kernel/probe/resume.mjs
CHANGED
|
@@ -557,14 +557,53 @@ export function setRunStatus(cwd, slug, status) {
|
|
|
557
557
|
if (!/^status:.*$/m.test(body)) {
|
|
558
558
|
return { ok: false, path: p, status, reason: `harness-run.md carries no "status:" line to replace — the ledger's frontmatter is malformed (references/protocol.md)` };
|
|
559
559
|
}
|
|
560
|
+
// A CLOSED RUN THAT MOVES AGAIN IS REOPENED, AND SAYS SO. A relaunch resumes the same run by
|
|
561
|
+
// design, so a run closed `aborted` can be carried on to a later ship. Its close was written once
|
|
562
|
+
// and the once-only guard in closeRun refuses to flip an outcome, which is right for a close
|
|
563
|
+
// nobody took back — but here somebody did, by resuming it. Leaving the old close in place made
|
|
564
|
+
// the ledger read `status: shipped` over `closed_status: aborted`, and every reader of the close
|
|
565
|
+
// reported an abort for a run that shipped. The move to a live status is the one act that takes a
|
|
566
|
+
// close back, so it is recorded here: the prior close joins `prior_closes`, the three close
|
|
567
|
+
// fields return to `~`, and the next terminal close is written as the run's own. Nothing is lost
|
|
568
|
+
// and nothing is flipped silently — the abort stays on the record beside the outcome that
|
|
569
|
+
// replaced it.
|
|
570
|
+
const fm = parseFrontmatter(body);
|
|
571
|
+
const priorStatus = fm.closed_status && fm.closed_status !== "~" ? String(fm.closed_status) : null;
|
|
572
|
+
let next = body.replace(/^status:.*$/m, `status: ${status}`);
|
|
573
|
+
let reopened = null;
|
|
574
|
+
if (priorStatus && TERMINAL_STATUSES.includes(priorStatus) && !TERMINAL_STATUSES.includes(status)) {
|
|
575
|
+
const priorAt = fm.closed_at && fm.closed_at !== "~" ? String(fm.closed_at) : null;
|
|
576
|
+
const priorCause = fm.close_cause && fm.close_cause !== "~" ? String(fm.close_cause) : null;
|
|
577
|
+
const reopenedAt = new Date().toISOString();
|
|
578
|
+
reopened = { closed_status: priorStatus, closed_at: priorAt, cause: priorCause, reopened_at: reopenedAt };
|
|
579
|
+
const entry = `${priorStatus} at ${priorAt ?? "?"} (reopened ${reopenedAt}): ${priorCause ?? "no cause recorded"}`;
|
|
580
|
+
const history = fm.prior_closes && fm.prior_closes !== "~" ? `${fm.prior_closes} | ${entry}` : entry;
|
|
581
|
+
const historyLine = `prior_closes: ${uncoerce(history)}`;
|
|
582
|
+
next = next
|
|
583
|
+
.replace(/^closed_at:.*$/m, "closed_at: ~")
|
|
584
|
+
.replace(/^closed_status:.*$/m, "closed_status: ~")
|
|
585
|
+
.replace(/^close_cause:.*$/m, "close_cause: ~");
|
|
586
|
+
next = /^prior_closes:.*$/m.test(next)
|
|
587
|
+
? next.replace(/^prior_closes:.*$/m, historyLine)
|
|
588
|
+
: next.replace(/^closed_at:.*$/m, (m) => `${m}\n${historyLine}`);
|
|
589
|
+
}
|
|
560
590
|
try {
|
|
561
|
-
writeFileSync(p,
|
|
591
|
+
writeFileSync(p, next);
|
|
562
592
|
} catch (e) {
|
|
563
593
|
return { ok: false, path: p, status, reason: `could not write the ledger: ${e.message}` };
|
|
564
594
|
}
|
|
565
|
-
const after = parseFrontmatter(readFileSync(p, "utf8"))
|
|
566
|
-
if (after !== status) {
|
|
567
|
-
return { ok: false, path: p, status, reason: `wrote "status: ${status}" but the ledger reads "${after}" — the write did not take` };
|
|
595
|
+
const after = parseFrontmatter(readFileSync(p, "utf8"));
|
|
596
|
+
if (after.status !== status) {
|
|
597
|
+
return { ok: false, path: p, status, reason: `wrote "status: ${status}" but the ledger reads "${after.status}" — the write did not take` };
|
|
598
|
+
}
|
|
599
|
+
if (reopened) {
|
|
600
|
+
// The breadcrumb names a run that is OVER; this one no longer is. Removed only when it names
|
|
601
|
+
// this run, so another run's close is left alone.
|
|
602
|
+
try {
|
|
603
|
+
const crumb = JSON.parse(readFileSync(lastRun(cwd), "utf8"));
|
|
604
|
+
if (crumb?.run_id && crumb.run_id === readRunId(cwd, slug)) rmSync(lastRun(cwd), { force: true });
|
|
605
|
+
} catch { /* no breadcrumb, or unreadable — nothing to retire */ }
|
|
606
|
+
return { ok: true, path: p, status, reopened, decision: "reopened" };
|
|
568
607
|
}
|
|
569
608
|
return { ok: true, path: p, status };
|
|
570
609
|
}
|
package/kernel/probe/t0.mjs
CHANGED
|
@@ -1,8 +1,16 @@
|
|
|
1
1
|
// probe t0 — "has this scope already reached T0-green in this round?"
|
|
2
2
|
//
|
|
3
3
|
// CONTRACT. A bounded, read-only query over the verdict artifacts on disk. Prints
|
|
4
|
-
// `{green, path, round, scope_id}` on stdout; exits 0 when green, 1 when not, 2 on a bad
|
|
5
|
-
// Writes nothing.
|
|
4
|
+
// `{green, path, sha256?, round, scope_id}` on stdout; exits 0 when green, 1 when not, 2 on a bad
|
|
5
|
+
// argv. Writes nothing.
|
|
6
|
+
//
|
|
7
|
+
// WHY IT PRINTS A DIGEST. A verdict on a scoped spec must cite a T0 artifact with the sha256 of the
|
|
8
|
+
// file as it is on disk now, and the judge is asked to obtain that itself rather than accept one it
|
|
9
|
+
// was handed. Some sessions have no shell hasher they may run, and a grant that covers this entry
|
|
10
|
+
// point and not `shasum` is the ordinary case, not an oversight. The digest here is computed from
|
|
11
|
+
// the bytes at the moment of the call, with the same reading the ingest step re-checks it against,
|
|
12
|
+
// so a judge that asks for it is asking the file, not a caller. `sha256` is present only when the
|
|
13
|
+
// verdict is green and readable; absent — never null — otherwise.
|
|
6
14
|
//
|
|
7
15
|
// WHY IT IS A SUBCOMMAND AND NOT AN INLINE SNIPPET. The control plane has no filesystem of its
|
|
8
16
|
// own, so the alternative is a `node -e` blob assembled inside the workflow script. Such a blob is
|
|
@@ -14,6 +22,7 @@
|
|
|
14
22
|
// the only durable evidence of that is the verdict artifact the evaluator is required to cite.
|
|
15
23
|
|
|
16
24
|
import { existsSync, readdirSync, readFileSync } from "node:fs";
|
|
25
|
+
import { createHash } from "node:crypto";
|
|
17
26
|
import { join, resolve } from "node:path";
|
|
18
27
|
import { runArgs } from "../lib/argv.mjs";
|
|
19
28
|
import { verdictsDir, readRunId } from "../lib/paths.mjs";
|
|
@@ -89,6 +98,18 @@ export function cli(rawArgv) {
|
|
|
89
98
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
90
99
|
const cwd = resolve(args.cwd || process.cwd());
|
|
91
100
|
const { green, path } = greenVerdict(cwd, args.slug, args.scope, args.round);
|
|
92
|
-
|
|
101
|
+
const sha256 = green ? digestOf(path) : null;
|
|
102
|
+
console.log(JSON.stringify({ green, path, ...(sha256 ? { sha256 } : {}), scope_id: args.scope, round: args.round }));
|
|
93
103
|
process.exit(green ? 0 : 1);
|
|
94
104
|
}
|
|
105
|
+
|
|
106
|
+
/**
|
|
107
|
+
* The sha256 of a verdict artifact, read as `reduce ingest` re-reads it (utf8 text, then hashed) so
|
|
108
|
+
* the two can never disagree about the same bytes.
|
|
109
|
+
* @param {string} path - Absolute path of the verdict file.
|
|
110
|
+
* @returns {(string|null)} Lowercase hex digest, or null when the file cannot be read.
|
|
111
|
+
*/
|
|
112
|
+
function digestOf(path) {
|
|
113
|
+
try { return createHash("sha256").update(readFileSync(path, "utf8")).digest("hex"); }
|
|
114
|
+
catch { return null; }
|
|
115
|
+
}
|
package/kernel/reduce/board.mjs
CHANGED
|
@@ -293,7 +293,15 @@ export function derive({ cwd, slug, appetiteHours = null }) {
|
|
|
293
293
|
}
|
|
294
294
|
|
|
295
295
|
/**
|
|
296
|
-
* Persist derived `unlocks` into task frontmatter —
|
|
296
|
+
* Persist derived `unlocks` into task frontmatter — and give a task with NO status the one status a
|
|
297
|
+
* task on a freshly written board can have.
|
|
298
|
+
*
|
|
299
|
+
* `unlocks` is derived, so writing it is safe by construction. `status: todo` is the other field a
|
|
300
|
+
* board is written with, and it is added only where the line is absent — an existing status is never
|
|
301
|
+
* touched, so a task in progress or done is never reset. A regenerated board once came back with
|
|
302
|
+
* `depends_on` and neither field on every task; spec-lint failed all of them at L1b and an unattended
|
|
303
|
+
* run had nobody to add a line the worker's own template already prescribes.
|
|
304
|
+
*
|
|
297
305
|
* @param {{_tasks:Array<object>, unlocks:Object<string,string[]>}} report - A {@link derive} report.
|
|
298
306
|
* @returns {string[]} The ids of task files actually rewritten (unchanged files are skipped).
|
|
299
307
|
* Side effect: writes those task files.
|
|
@@ -305,7 +313,8 @@ export function writeUnlocks(report) {
|
|
|
305
313
|
const fmMatch = t.body.match(/^---\r?\n([\s\S]*?)\r?\n---/);
|
|
306
314
|
if (!fmMatch) continue;
|
|
307
315
|
const fm = fmMatch[1];
|
|
308
|
-
|
|
316
|
+
let next = /^unlocks:.*$/m.test(fm) ? fm.replace(/^unlocks:.*$/m, `unlocks: ${want}`) : `${fm}\nunlocks: ${want}`;
|
|
317
|
+
if (!/^status:/m.test(next)) next = `${next}\nstatus: todo`;
|
|
309
318
|
if (next !== fm) {
|
|
310
319
|
writeFileSync(t.file, t.body.replace(fm, next));
|
|
311
320
|
written.push(t.id);
|
package/kernel/reduce/ship.mjs
CHANGED
|
@@ -32,7 +32,7 @@ import { runArgs } from "../lib/argv.mjs";
|
|
|
32
32
|
import {
|
|
33
33
|
report as reportPath, tasksDir, verdictsDir, trials, evaluationDir, qaDir,
|
|
34
34
|
roundLedger, discoveryLedger, receipt as receiptPath, harnessRun, relShared,
|
|
35
|
-
activeOrder, runArgsPath, readReceipt, runIdFromReceipt,
|
|
35
|
+
activeOrder, runArgsPath, readReceipt, runIdFromReceipt, hammerCensus,
|
|
36
36
|
} from "../lib/paths.mjs";
|
|
37
37
|
import { readTrials } from "../verify/t0.mjs";
|
|
38
38
|
import { ratchetReport } from "../probe/stats.mjs";
|
|
@@ -191,6 +191,31 @@ export function section(md, heading) {
|
|
|
191
191
|
return body || null;
|
|
192
192
|
}
|
|
193
193
|
|
|
194
|
+
/**
|
|
195
|
+
* The report's opening line when the ship is not a passing one — null for a PASS.
|
|
196
|
+
*
|
|
197
|
+
* @param {string} verdict - The run's final verdict (`PASS`, `FAIL`, `not-evaluated`, …).
|
|
198
|
+
* @param {({verdict?:string, cut_list?:Array}|null)} census - GATE H's census as data, or null.
|
|
199
|
+
* @returns {(string|null)} One markdown line naming the verdict, and the cut list that cleared it
|
|
200
|
+
* when a census is on record; null when the verdict is PASS.
|
|
201
|
+
*/
|
|
202
|
+
export function shipHeadline(verdict, census) {
|
|
203
|
+
if (verdict === "PASS") return null;
|
|
204
|
+
const cuts = Array.isArray(census?.cut_list) ? census.cut_list.length : null;
|
|
205
|
+
const why = census
|
|
206
|
+
? `GATE H's baseline comparison cleared it${census.verdict ? ` (census: ${census.verdict})` : ""}` +
|
|
207
|
+
(cuts != null ? ` with ${cuts} item${cuts === 1 ? "" : "s"} cut` : "")
|
|
208
|
+
: "no GATE H census was on record when it shipped";
|
|
209
|
+
// A FAIL is a judgement that some criterion failed; anything else here means no judgement was
|
|
210
|
+
// made at all, and the line must not claim criteria failed that nobody graded.
|
|
211
|
+
if (verdict === "FAIL") {
|
|
212
|
+
return `> **Shipped with a FAIL verdict.** Not every criterion passed; ${why}. Read the failed ` +
|
|
213
|
+
"criteria and the cut list below before treating this feature as done.";
|
|
214
|
+
}
|
|
215
|
+
return `> **Shipped without a passing verdict (${verdict}).** No criterion was graded as passing; ${why}. ` +
|
|
216
|
+
"Nothing below is evidence that this feature works.";
|
|
217
|
+
}
|
|
218
|
+
|
|
194
219
|
/**
|
|
195
220
|
* Assemble the report. Pure given its inputs, so the structural tests can assert its shape
|
|
196
221
|
* without a filesystem.
|
|
@@ -201,6 +226,7 @@ export function buildReport(facts) {
|
|
|
201
226
|
const {
|
|
202
227
|
slug, at, verdict, qa, rounds, roundsJudged, board, t0, artifacts, ratchet,
|
|
203
228
|
evalCriteria, evalBugs, qaFindings, decisions, discovered, intakeSha, leftovers, requirements,
|
|
229
|
+
census = null,
|
|
204
230
|
} = facts;
|
|
205
231
|
|
|
206
232
|
const L = [];
|
|
@@ -208,6 +234,12 @@ export function buildReport(facts) {
|
|
|
208
234
|
`verdict: ${verdict}`, `rounds_used: ${rounds ?? "~"}`, `rounds_judged: ${roundsJudged ?? "~"}`, `qa: ${qa}`,
|
|
209
235
|
`intake_sha256: ${intakeSha ?? "~"}`, "---", "");
|
|
210
236
|
L.push(`# ${slug} — ship report`, "");
|
|
237
|
+
// A SHIP IS NOT A PASS, AND THE FIRST LINE SAYS WHICH ONE THIS IS. A run whose verdict FAILed can
|
|
238
|
+
// still ship — GATE H compares against the baseline, not the ideal, and a cut list can clear it —
|
|
239
|
+
// and its status then reads `shipped` like any other. The verdict sat in a table cell below the
|
|
240
|
+
// fold, so a reader who stopped at the title took a shipped FAIL for a passing build.
|
|
241
|
+
const headline = shipHeadline(verdict, census);
|
|
242
|
+
if (headline) L.push(headline, "");
|
|
211
243
|
L.push("Frozen at GATE L4. Every figure below is derived from run artifacts on disk — the trial",
|
|
212
244
|
"ledger, the verdict artifacts, the board — never from a summary of the run.", "");
|
|
213
245
|
|
|
@@ -361,6 +393,7 @@ export function generate({ cwd, slug, verdict, qa }) {
|
|
|
361
393
|
slug,
|
|
362
394
|
at: today(),
|
|
363
395
|
verdict: verdict || run.final_verdict || "not-evaluated",
|
|
396
|
+
census: (() => { try { return JSON.parse(readIf(hammerCensus(cwd, slug)) || "null"); } catch { return null; } })(),
|
|
364
397
|
qa: qa || (huntReport ? "run" : "skipped"),
|
|
365
398
|
rounds: derivedRounds.rounds_used,
|
|
366
399
|
roundsJudged: derivedRounds.rounds_judged,
|
|
@@ -301,6 +301,13 @@
|
|
|
301
301
|
"via": "discovered_tasks[]",
|
|
302
302
|
"note": "a red gate's output, digested — surfaces as payload.bugs on the next round's orders"
|
|
303
303
|
},
|
|
304
|
+
{
|
|
305
|
+
"from": "WorkOrder",
|
|
306
|
+
"to": "RoundBuildVerdict",
|
|
307
|
+
"cardinality": "N:1",
|
|
308
|
+
"via": "payload.build_gate",
|
|
309
|
+
"note": "an evaluate order names this run's newest gate artifact for its round, green or red, so the judge knows the build ran and the app launched before grading a [ui] row; payload.launch_cmd carries the profile's launch_probe beside it"
|
|
310
|
+
},
|
|
304
311
|
{
|
|
305
312
|
"from": "RoundBuildVerdict",
|
|
306
313
|
"to": "HillShard",
|
|
@@ -353,6 +360,8 @@
|
|
|
353
360
|
"feature",
|
|
354
361
|
"dimensions",
|
|
355
362
|
"run_cmd",
|
|
363
|
+
"launch_cmd",
|
|
364
|
+
"build_gate",
|
|
356
365
|
"t0_artifacts",
|
|
357
366
|
"browser",
|
|
358
367
|
"tasks"
|
|
@@ -2416,6 +2425,14 @@
|
|
|
2416
2425
|
"type": "string",
|
|
2417
2426
|
"description": "spec-evaluator: how to start the running app. Absent orchestrated → ESCALATE, never guess."
|
|
2418
2427
|
},
|
|
2428
|
+
"launch_cmd": {
|
|
2429
|
+
"type": "string",
|
|
2430
|
+
"description": "spec-evaluator: the project profile's launch probe — installs the built artifact, starts it and asserts the first screen. Derived by `harness compile` from project-profile.md, absent when the profile declares none. Where `run_cmd` is only a build, this is how the app is brought up; a non-zero exit is a finding, not a reason to guess another way."
|
|
2431
|
+
},
|
|
2432
|
+
"build_gate": {
|
|
2433
|
+
"type": "string",
|
|
2434
|
+
"description": "spec-evaluator: this run's newest round build gate artifact (build/r<N>-t<T>.json) — each step's exit code and output tail, the record that the build ran and the app launched. Derived by `harness compile`, absent when the gate never ran. Read it before grading a [ui] row NO EVIDENCE."
|
|
2435
|
+
},
|
|
2419
2436
|
"t0_artifacts": {
|
|
2420
2437
|
"type": "array",
|
|
2421
2438
|
"items": {
|
package/kernel/verify/build.mjs
CHANGED
|
@@ -220,6 +220,62 @@ export function latestRoundBuild(cwd, slug, round) {
|
|
|
220
220
|
return all.length ? all[all.length - 1].body : null;
|
|
221
221
|
}
|
|
222
222
|
|
|
223
|
+
/**
|
|
224
|
+
* The gate artifact an evaluate order should name — the newest one THIS run wrote, with the file it
|
|
225
|
+
* lives in.
|
|
226
|
+
*
|
|
227
|
+
* {@link latestRoundBuild} answers "what did the gate say about round N" and hands back a body; an
|
|
228
|
+
* order has to point at a file, so the path travels with it here. Scoped to the current run when a
|
|
229
|
+
* receipt is readable: a prior run over the same slug leaves its own `r1-t1.json` on disk, and the
|
|
230
|
+
* judge must not be pointed at another run's launch evidence. Scoped to one round when given; with
|
|
231
|
+
* no round it is the newest of any round — a standalone evaluation has none, the same reading
|
|
232
|
+
* `t0ArtifactsFor` uses for the T0 verdicts.
|
|
233
|
+
*
|
|
234
|
+
* @param {string} cwd - Project root.
|
|
235
|
+
* @param {string} slug - Feature slug.
|
|
236
|
+
* @param {number} [round] - The round being evaluated.
|
|
237
|
+
* @returns {({path:string, body:object}|null)} null when the gate never ran, or nothing readable
|
|
238
|
+
* belongs to this run. Never throws: an order must compile without it.
|
|
239
|
+
*/
|
|
240
|
+
export function latestRoundBuildFile(cwd, slug, round) {
|
|
241
|
+
let files;
|
|
242
|
+
const dir = roundBuildDir(cwd, slug);
|
|
243
|
+
try { files = readdirSync(dir); } catch { return null; }
|
|
244
|
+
const runId = readRunId(cwd, slug);
|
|
245
|
+
let best = null;
|
|
246
|
+
for (const f of files) {
|
|
247
|
+
const m = f.match(/^r(\d+)-t(\d+)\.json$/);
|
|
248
|
+
if (!m) continue;
|
|
249
|
+
const r = Number(m[1]), t = Number(m[2]);
|
|
250
|
+
if (round != null && r !== Number(round)) continue;
|
|
251
|
+
let body;
|
|
252
|
+
try { body = JSON.parse(readFileSync(join(dir, f), "utf8")); } catch { continue; }
|
|
253
|
+
if (runId && body?.run_id && body.run_id !== runId) continue;
|
|
254
|
+
if (!best || r > best.r || (r === best.r && t > best.t)) best = { r, t, path: join(dir, f), body };
|
|
255
|
+
}
|
|
256
|
+
return best && { path: best.path, body: best.body };
|
|
257
|
+
}
|
|
258
|
+
|
|
259
|
+
/**
|
|
260
|
+
* The command the project declared for bringing the built app up — `launch_probe` in the committed
|
|
261
|
+
* profile: install the artifact, start it, assert the first screen. Read on its own rather than
|
|
262
|
+
* through {@link declaredSteps}, which also reads the ledger and computes warnings an order has no
|
|
263
|
+
* use for.
|
|
264
|
+
*
|
|
265
|
+
* @param {string} cwd - Project root.
|
|
266
|
+
* @param {string} slug - Feature slug.
|
|
267
|
+
* @returns {(string|null)} The command, or null when the profile is absent or declares none.
|
|
268
|
+
* Never throws.
|
|
269
|
+
*/
|
|
270
|
+
export function launchProbeFor(cwd, slug) {
|
|
271
|
+
try {
|
|
272
|
+
const pp = projectProfile(cwd, slug);
|
|
273
|
+
if (!existsSync(pp)) return null;
|
|
274
|
+
const v = readContract(pp, PROJECT_PROFILE)?.contract?.launch_probe;
|
|
275
|
+
return typeof v === "string" && v.trim() ? v.trim() : null;
|
|
276
|
+
} catch { return null; }
|
|
277
|
+
}
|
|
278
|
+
|
|
223
279
|
/**
|
|
224
280
|
* The rounds whose latest gate artifact is red — the set `reduce hill` subtracts from.
|
|
225
281
|
* @param {string} cwd - Project root.
|
|
@@ -275,13 +331,40 @@ export function writeRoundBuild(cwd, slug, round, body) {
|
|
|
275
331
|
}
|
|
276
332
|
|
|
277
333
|
export const ARGV_SPEC = {
|
|
278
|
-
usage: "harness.mjs verify build --slug <slug> --round <N> [--cwd <dir>]",
|
|
334
|
+
usage: "harness.mjs verify build --slug <slug> (--round <N> | --preflight) [--cwd <dir>]",
|
|
279
335
|
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
280
336
|
slug: { type: "str", required: true },
|
|
281
|
-
round: { type: "int", min: 1
|
|
337
|
+
round: { type: "int", min: 1 },
|
|
338
|
+
preflight: { type: "flag" },
|
|
282
339
|
cwd: { type: "path" },
|
|
283
340
|
};
|
|
284
341
|
|
|
342
|
+
/**
|
|
343
|
+
* Classify a preflight: the declared steps run, nothing is written.
|
|
344
|
+
*
|
|
345
|
+
* WHY A PREFLIGHT, AND WHY IT WRITES NOTHING. Every environment fault a run has died of was
|
|
346
|
+
* discoverable in seconds and discovered after most of an hour: a checkout missing its package
|
|
347
|
+
* install or its local SDK pointer fails the build in about a second, a probe outside the session's
|
|
348
|
+
* grant is refused, a device that is not attached leaves every `[ui]` row ungraded. Each surfaced
|
|
349
|
+
* only at the first round gate, after planning had been paid for. The same commands, run once before
|
|
350
|
+
* anything is dispatched, turn that into one report up front. It is not a round, so it writes no
|
|
351
|
+
* round artifact: a red preflight read back as round 0's gate would reach round 1's orders as bugs,
|
|
352
|
+
* handing a feature worker an environment fault it cannot fix.
|
|
353
|
+
*
|
|
354
|
+
* A step that exits 2 is a probe that COULD NOT RUN — the fixture convention, e.g. no device
|
|
355
|
+
* attached — which is a gap in the environment rather than a red build, and is reported apart.
|
|
356
|
+
*
|
|
357
|
+
* @param {{overall:string, steps:Array<{kind:string, exit?:number, pass?:boolean, skipped?:boolean}>}} gate
|
|
358
|
+
* What {@link runGate} returned.
|
|
359
|
+
* @returns {{status:("green"|"red"|"cannot-run"), failed_step:(string|null)}} `cannot-run` when the
|
|
360
|
+
* first failing step exited 2.
|
|
361
|
+
*/
|
|
362
|
+
export function classifyPreflight(gate) {
|
|
363
|
+
const failing = gate.steps.find((s) => !s.skipped && !s.pass);
|
|
364
|
+
if (!failing) return { status: "green", failed_step: null };
|
|
365
|
+
return { status: failing.exit === 2 ? "cannot-run" : "red", failed_step: failing.kind };
|
|
366
|
+
}
|
|
367
|
+
|
|
285
368
|
/**
|
|
286
369
|
* Run the round build gate and write its artifact.
|
|
287
370
|
*
|
|
@@ -292,9 +375,29 @@ export async function cli(rawArgv) {
|
|
|
292
375
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
293
376
|
const cwd = resolve(args.cwd || process.cwd());
|
|
294
377
|
const { slug, round } = args;
|
|
378
|
+
if (Boolean(args.preflight) === (round != null)) {
|
|
379
|
+
console.error("verify build: pass exactly one of --round <N> (the round gate, writes its artifact) or --preflight (runs the same steps, writes nothing)");
|
|
380
|
+
process.exit(2);
|
|
381
|
+
}
|
|
295
382
|
const { archetype, steps, warnings } = declaredSteps(cwd, slug);
|
|
296
383
|
for (const w of warnings) console.error(`build-gate: ${w}`);
|
|
297
384
|
|
|
385
|
+
if (args.preflight) {
|
|
386
|
+
if (steps.length === 0) {
|
|
387
|
+
console.log(JSON.stringify({ preflight: true, status: "undeclared", steps: [], warnings }, null, 2));
|
|
388
|
+
process.exit(3);
|
|
389
|
+
}
|
|
390
|
+
const gate = runGate(steps, cwd);
|
|
391
|
+
const { status, failed_step } = classifyPreflight(gate);
|
|
392
|
+
const failing = gate.steps.find((s) => s.kind === failed_step);
|
|
393
|
+
console.log(JSON.stringify({
|
|
394
|
+
preflight: true, status, archetype, warnings,
|
|
395
|
+
steps: gate.steps.map((s) => (s.skipped ? { kind: s.kind, skipped: true } : { kind: s.kind, exit: s.exit, pass: s.pass })),
|
|
396
|
+
...(failing ? { failed_step, stderr_tail: `${failing.stdout_tail || ""}\n${failing.stderr_tail || ""}`.trim().slice(-1200) } : {}),
|
|
397
|
+
}, null, 2));
|
|
398
|
+
process.exit(status === "green" ? 0 : status === "cannot-run" ? 4 : 1);
|
|
399
|
+
}
|
|
400
|
+
|
|
298
401
|
if (steps.length === 0) {
|
|
299
402
|
console.log(JSON.stringify({ round, overall: "skipped", steps: [], warnings,
|
|
300
403
|
reason: "nothing declared — no run_cmd in the run ledger, no build_probe or launch_probe in project-profile.md" }, null, 2));
|
package/package.json
CHANGED
|
@@ -36,7 +36,9 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
|
|
|
36
36
|
| `payload.spec_folder` | The committed grading truth: `usecases/` + `domain-model.md` (+ `contracts/`, `scope-summary.md`, `_index.md`). No `usecases/` → HARD STOP, nothing to grade against |
|
|
37
37
|
| `payload.feature` | Feature slug — scopes the probe and names the report |
|
|
38
38
|
| `payload.dimensions[]` | The active dimension set (the caller resolved precedence). Absent → `[spec-conformance]` + the auto-enable rules below |
|
|
39
|
-
| `payload.run_cmd` | How to start the running app. Absent standalone → ask; absent orchestrated → ESCALATE, do not guess |
|
|
39
|
+
| `payload.run_cmd` | How to start the running app. On a stack where building and launching are different acts it is only the build — prefer `payload.launch_cmd` when the order carries one. Absent standalone → ask; absent orchestrated with no `launch_cmd` either → ESCALATE, do not guess |
|
|
40
|
+
| `payload.launch_cmd` | The project profile's launch probe: installs the built artifact, starts it and asserts the first screen. This is how the app is brought up for `[ui]` probing. A non-zero exit is a finding to cite, not a reason to try another way. Absent → the profile declares none; fall back to `run_cmd` |
|
|
41
|
+
| `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
|
|
40
42
|
| `payload.t0_artifacts[]` | Per-scope T0 verdict paths for this round (scoped specs), compiled from each scope's green verdict. An artifact listed but missing/red on disk, or a scoped spec with none listed → the round is NOT gradeable: return `status: failed` with the reason, naming the scope, as your FIRST deviation — a structural precondition, not a criterion |
|
|
41
43
|
| `payload.browser` | `cli` (default, ~4x cheaper) \| `mcp` \| `none` |
|
|
42
44
|
| `payload.tasks[]` | Traceability only (which UCs a task claims): NEVER a grading source — the committed UC text is the criterion, a paraphrase mismatch is a finding |
|
|
@@ -94,7 +96,10 @@ Done-when statements; `_index.md` Non-Go list. Which UCs are in scope comes from
|
|
|
94
96
|
contract triplet + Non-Go). Overall PASS only if ALL active dimensions pass — the halo
|
|
95
97
|
effect is banned; a strong dimension never lifts a failing one.
|
|
96
98
|
- **T0 citation (scoped specs).** Recompute each cited artifact's sha256 from disk — never
|
|
97
|
-
trust a handed hash.
|
|
99
|
+
trust a handed hash. Use any hasher you may run; when none is available to you,
|
|
100
|
+
`harness probe t0 --slug <slug> --scope <scope_id> --round <r>` prints the digest of that
|
|
101
|
+
scope's green verdict as the file is at that moment. A digest from any other place — a ledger
|
|
102
|
+
line, an earlier report, an error message — is a handed hash, however correct it looks. A verdict on a scoped spec without a T0 citation is structurally
|
|
98
103
|
invalid, regardless of how convincing your own probing looked; generator prose ("tests
|
|
99
104
|
pass", "verified locally") is never admissible evidence.
|
|
100
105
|
- **When a criterion names a command, run THAT command.** Not the one that works, not the
|
|
@@ -39,9 +39,9 @@ derived RESUME STATE (slug, status, round, board counts) — read it and go stra
|
|
|
39
39
|
this session's memory. `--force` re-opens deliberately and discards the round history the breaker
|
|
40
40
|
counts.
|
|
41
41
|
|
|
42
|
-
**If this command comes back "requires approval", stop and say so
|
|
43
|
-
|
|
44
|
-
|
|
42
|
+
**If this command comes back "requires approval", stop and say so** — the scripts need a one-time grant
|
|
43
|
+
(`npx shapeup-sdlc init`); never route around it or hand-build the feature. Know a grant by running the
|
|
44
|
+
command, never by reading `.claude/settings.json` — Preflight runs the project's probes and reports.
|
|
45
45
|
|
|
46
46
|
**Language gate (delegated to `translator`, not this skill):** before Step 1 opens the run, dispatch
|
|
47
47
|
an Agent (model: exec) calling `Skill(shapeup-sdlc-plugin:translator) --check` on the pitch *and* its
|
|
@@ -477,6 +477,13 @@ Read back: stdout JSON — {overall: green|red, steps[], warnings[], failed_step
|
|
|
477
477
|
A `mobile` profile with no launch_probe is warned about on stderr every round, and so is
|
|
478
478
|
every scope none of whose fixtures invoke the tool run_cmd builds with — advisory, because a
|
|
479
479
|
green T0 from such fixtures is not evidence the scope compiles.
|
|
480
|
+
Preflight form — the same steps, run once before anything is dispatched, WRITING NOTHING:
|
|
481
|
+
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify build --slug <slug> --preflight
|
|
482
|
+
stdout JSON {preflight: true, status: green|red|cannot-run|undeclared, steps[], failed_step?,
|
|
483
|
+
stderr_tail?}. Exit 0 green · 1 a step red · 4 a probe could not run (it exited 2 — no
|
|
484
|
+
device, no artifact) · 3 nothing declared. No round artifact: a red preflight read back as
|
|
485
|
+
round 0's gate would reach round 1's orders as bugs. The workflow warns on 1 and 4, never
|
|
486
|
+
aborts.
|
|
480
487
|
Consequences, both mechanical and both read off the artifact, never off this prose:
|
|
481
488
|
red → EVAL is not dispatched this round; `harness compile` turns each failing step into a
|
|
482
489
|
`payload.bugs` entry for round N+1, addressed to the scope whose substrate holds the
|
|
@@ -490,6 +497,9 @@ compile-order --operation evaluate --slug <slug> --worker spec-evaluator --round
|
|
|
490
497
|
--payload '{"dimensions": ["spec-conformance"], "run_cmd": "<cmd>"}'
|
|
491
498
|
t0_artifacts is compiled from each scope's green T0 verdict for round <r> — pass it only to
|
|
492
499
|
override. A scope with no green verdict is named on stderr: the judge has nothing to cite for it.
|
|
500
|
+
build_gate and launch_cmd are compiled too: this run's newest round build gate artifact, and the
|
|
501
|
+
profile's launch_probe. Neither is passed here — a `[ui]` row is graded on the running app, and
|
|
502
|
+
these are how the judge finds out the app was launched and how to bring it up again.
|
|
493
503
|
Invoke via Agent (model: eval), ONCE, after GATE L2:
|
|
494
504
|
Skill(shapeup-sdlc-plugin:spec-evaluator) --order <path>
|
|
495
505
|
Effect: one feature-level pass over the running app against all AC + Done-when; writes
|
|
@@ -548,7 +558,7 @@ Read back: the proposed cut list + verdict (SHIP now | SHIP after fixing ship-bl
|
|
|
548
558
|
| `harness-run.md` | **tech lead (sole writer)** | tech lead (round ledger + Hill + run-state), PO (audit) |
|
|
549
559
|
| `scopes/<scope-id>.md` | `scope-architect` (sole writer) | tech lead (substrate/sequence), sandbox hook (write-whitelist), compile-order (inlined into orders) |
|
|
550
560
|
| `t0/verdicts/r<N>-a<M>-t<T>.json` | `harness verify t0` (skill-local, mechanical — not a worker) | spec-evaluator (required citation), tech lead (hill derivation), compile-order (digested errors) |
|
|
551
|
-
| `build/r<N>-t<T>.json` (the round build gate) | `harness verify build` (mechanical — run_cmd + build_probe + launch_probe, once per round before EVAL) | tech lead (GATE L2 block), `harness reduce hill` (a red round moves no dot), compile-order (`payload.bugs` for round N+1) |
|
|
561
|
+
| `build/r<N>-t<T>.json` (the round build gate) | `harness verify build` (mechanical — run_cmd + build_probe + launch_probe, once per round before EVAL) | tech lead (GATE L2 block), `harness reduce hill` (a red round moves no dot), compile-order (`payload.bugs` for round N+1 when red; `payload.build_gate` on the same round's evaluate order) |
|
|
552
562
|
| `t0/trials.jsonl` (the ratchet ledger, append-only, `baseline_trial` as the parent link) | `harness verify t0` (one row per attempt: score, status, delta, tree_ref) | compile-order (`trial_history` into the next order), ship-report (T0 + Ratchet sections), `harness probe stats --ratchet` |
|
|
553
563
|
| `round-ledger.md` | **tech lead (sole writer)** | compile-order (decisions into every order), PO (audit) |
|
|
554
564
|
| `hill/<scope-id>.yml` + `hill-chart.md` | **tech lead (sole writer)** | PO ("status without asking"), scope-hammer (H0 census) |
|
|
@@ -54,7 +54,7 @@ export const meta = {
|
|
|
54
54
|
name: "shapeup-run",
|
|
55
55
|
description: "BUILD-phase pipeline: ORIENT → ANALYZE → WIRE → MAP SCOPES → rounds of BUILD/EVAL → QA → GATE H → ship. Gates resolve by the kernel's exit code; every dispatch is WorkOrder in / WorkResult out; the fast-forward is derived from artifacts on disk.",
|
|
56
56
|
phases: [
|
|
57
|
-
{ title: "Preflight", detail: "one canary dispatch — can this session resolve a worker skill" },
|
|
57
|
+
{ title: "Preflight", detail: "one canary dispatch — can this session resolve a worker skill; then the project's build and launch probes, run once and recorded nowhere" },
|
|
58
58
|
{ title: "Orient" }, { title: "Analyze" }, { title: "Wire" }, { title: "MapScopes" },
|
|
59
59
|
{ title: "Build" }, { title: "Eval" }, { title: "Refute" }, { title: "QA" }, { title: "Ship" },
|
|
60
60
|
],
|
|
@@ -1117,6 +1117,11 @@ async function setRunStatus(status, phaseName) {
|
|
|
1117
1117
|
`artifacts, not from this field), but the snapshot and the the ship report's census hook will read ` +
|
|
1118
1118
|
`this run as unfinished.`);
|
|
1119
1119
|
stateWarnings.push(`status="${status}" did not take: ${why}`);
|
|
1120
|
+
} else if (r.decision === "reopened") {
|
|
1121
|
+
// The ledger now carries the earlier close under `prior_closes`; say so where a reader of the
|
|
1122
|
+
// run's own log will see it, because resuming a closed run changes what its close will say.
|
|
1123
|
+
log(`RUN STATE — this run had been closed; moving it to "${status}" reopened it. The earlier close is ` +
|
|
1124
|
+
`kept in the ledger under prior_closes, and the run's next terminal close is recorded as its own.`);
|
|
1120
1125
|
}
|
|
1121
1126
|
}
|
|
1122
1127
|
|
|
@@ -1235,6 +1240,33 @@ if (!canary.ok) {
|
|
|
1235
1240
|
`(${canary.detail || `exit ${canary.exit_code}`})`));
|
|
1236
1241
|
}
|
|
1237
1242
|
|
|
1243
|
+
// ENVIRONMENT CANARY — do the project's own build and launch probes run in THIS session?
|
|
1244
|
+
//
|
|
1245
|
+
// The skill canary proves the workers resolve; it says nothing about whether the commands the round
|
|
1246
|
+
// gate will run can. A missing package install or local SDK pointer fails the build in a second, a
|
|
1247
|
+
// probe outside the grant is refused, a device that is not attached leaves every `[ui]` row
|
|
1248
|
+
// ungraded — and each used to surface only at the first round gate, after planning was paid for.
|
|
1249
|
+
// `verify build --preflight` runs the same declared steps from a sub-agent, through the same grant,
|
|
1250
|
+
// and writes nothing. It WARNS rather than aborts: a baseline can be red for a reason the feature is
|
|
1251
|
+
// meant to fix, and an unattended run should still say so up front rather than stop.
|
|
1252
|
+
//
|
|
1253
|
+
// The evidence is this command's exit code. Whether a grant "looks present" in a settings file is
|
|
1254
|
+
// not evidence of anything: a grant can arrive by other routes, and one in the file can still be
|
|
1255
|
+
// refused above it.
|
|
1256
|
+
const envCanary = await cmd(`verify build --slug ${slug} --preflight`, "Preflight", "canary-env");
|
|
1257
|
+
if (envCanary.exit_code === 1 || envCanary.exit_code === 4) {
|
|
1258
|
+
const kind = envCanary.exit_code === 4 ? "a probe could not run (exit 2 — typically no device attached, or a missing artifact)" : "a declared step is red";
|
|
1259
|
+
const msg = `ENVIRONMENT — preflight: ${kind}${envCanary.detail ? `: ${envCanary.detail}` : ""}. The same step runs in every ` +
|
|
1260
|
+
`round's build gate, so fix the environment now if it is not the feature's to fix.`;
|
|
1261
|
+
log(msg);
|
|
1262
|
+
stateWarnings.push(msg);
|
|
1263
|
+
} else if (envCanary.exit_code !== 0 && envCanary.exit_code !== 3) {
|
|
1264
|
+
const msg = `ENVIRONMENT — preflight did not run (exit ${envCanary.exit_code}${envCanary.detail ? `: ${envCanary.detail}` : ""}); ` +
|
|
1265
|
+
`nothing is known about the build or launch probes until the first round gate.`;
|
|
1266
|
+
log(msg);
|
|
1267
|
+
stateWarnings.push(msg);
|
|
1268
|
+
}
|
|
1269
|
+
|
|
1238
1270
|
// GATE L0.9b's launch record must exist before anything past Preflight dispatches; see
|
|
1239
1271
|
// requireLaunchRecord()'s own banner for why this cannot be left to Step 2's prose alone.
|
|
1240
1272
|
const launchRecordAbort = await requireLaunchRecord();
|
|
@@ -1374,6 +1406,10 @@ if (!rs.has_spec_tree) {
|
|
|
1374
1406
|
"carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder.",
|
|
1375
1407
|
});
|
|
1376
1408
|
if (b.__failed) return await withWarnings(diedAt("ANALYZE", b));
|
|
1409
|
+
// The board a worker writes carries `depends_on`; its inverse (`unlocks`) and a new task's
|
|
1410
|
+
// `status: todo` are mechanical, so the kernel writes them rather than trusting every regeneration
|
|
1411
|
+
// to remember — a board that came back without both failed spec-lint on every task at L1b.
|
|
1412
|
+
await advisory(`reduce board --slug ${slug} --write`, "Analyze", "board:derive");
|
|
1377
1413
|
const post = await requirePhase("ANALYZE", "analyze", "Analyze", "board");
|
|
1378
1414
|
if (post) return await withWarnings(post);
|
|
1379
1415
|
await advisory(`reduce graph --slug ${slug}`, "Analyze", "graph:board");
|
|
@@ -1837,6 +1873,9 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1837
1873
|
model: evalModel, round,
|
|
1838
1874
|
// No `t0_artifacts` here, deliberately: `harness compile` derives them from the round's green
|
|
1839
1875
|
// T0 verdicts on disk, for every lane — this script could only name paths it was told about.
|
|
1876
|
+
// The same goes for `build_gate` and `launch_cmd`: compile reads them off the round's gate
|
|
1877
|
+
// artifact and the project profile, so the judge gets the launch evidence without this script
|
|
1878
|
+
// having to carry it.
|
|
1840
1879
|
payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
|
|
1841
1880
|
extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself.",
|
|
1842
1881
|
});
|