shapeup-sdlc 3.7.4 → 3.7.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +3 -3
- package/kernel/compile.mjs +15 -1
- package/kernel/gate.mjs +34 -2
- package/kernel/lib/paths.mjs +10 -0
- package/kernel/probe/leg.mjs +31 -4
- package/kernel/probe/resume.mjs +117 -7
- package/kernel/reduce/ingest.mjs +48 -7
- package/kernel/schemas/domain.schema.json +18 -1
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/SKILL.md +3 -2
- package/skills/scope-hammer/SKILL.md +13 -0
- package/skills/tech-lead/references/gates.md +6 -3
- package/skills/tech-lead/workflows/shapeup-run.js +129 -15
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.7.
|
|
4
|
+
"version": "3.7.6",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -13,7 +13,7 @@ Skills and commands are named short throughout this file; every one of them reso
|
|
|
13
13
|
- **The single writer is asked, never assumed.** A dispatch's result is applied by exactly one act — the ingest step — and that act leaves a leg row. Every phase post-condition and every settled build scope asks the leg ledger whether the row exists before any verdict on the phase or the scope, whatever the worker reported: a result on disk that nothing applied is ingested there if the scope is green, and otherwise named at the close as `unapplied_results=N`, never folded into "not green". A round that ends with nothing green names `breaker=none` and `stalled=no_green` unless the attested attempt census says a scope tripped — the close cannot name a breaker the census denies. The run graph carries a `Leg` node with an `INGESTED` edge from its `Result`, and the export a `leg` table, so "which results were never read" is one query.
|
|
14
14
|
- GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
|
|
15
15
|
- Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
|
|
16
|
-
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
|
|
16
|
+
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. The same close derives the ledger's verdict, its round count and its Rounds and Decisions tables from the run's own records — the ones the export reads — so the file a teammate opens first cannot contradict its own close line. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
|
|
17
17
|
- **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
|
|
18
18
|
- The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
|
|
19
19
|
|
|
@@ -32,7 +32,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
32
32
|
| Kick-off | ⏸ **L0** — Intake & Config (L0.8 model/budget matrix) + worker roster ✧ | `/translator` if non-English |
|
|
33
33
|
| Orient (Scout) | ⏸ **L1a** — Orient Review | `/orient` |
|
|
34
34
|
| Requirements | — (reviewed at L1b) | `/ba-pitch-analyzer` (`coverage`): the pitch's clauses → committed `requirements.md`, one atomic clause per `REQ-<n>` row naming the pitch clause it came from, ids assigned once and frozen; dispatched once, ahead of Analyze because the acceptance criteria are what cite its ids. A registry already on disk is not re-dispatched |
|
|
35
|
-
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
|
|
35
|
+
| Analyze | — (reviewed at L1b) | `/ba-pitch-analyzer` (`analyze`): spec tree + board (UC + Invariants + Test Surface ★); before Wire (needs its use cases). The tree is committed and the board is per-machine: a run that finds the tree on disk and no board dispatches `board`, which regenerates the board from the tree without re-deriving it — GATE L2 refuses `proceed` over a board with zero tasks. An acceptance criterion that grades a requirement carries `(covers: REQ-…)` — that clause is the edge the verdict travels back along |
|
|
36
36
|
| Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
38
|
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦; the seesaw regression arm is declared and not yet wired), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
@@ -48,7 +48,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
48
48
|
**Q0** Preflight → **Q1** Charter (6 lenses − EVAL-covered) → **Hunt** (repro required, findings `~` → ledger) → report (no verdict, no score). Skip with `--no-qa`.
|
|
49
49
|
|
|
50
50
|
### Ship & Triage
|
|
51
|
-
- **SHIP S.0 / GATE H** — `/scope-hammer`: census (QA findings + discovered ledger + attempt-budget proposals ✦; every "no scope owns X" cites `probe owner`, which derives ownership from the contracts) → baseline comparison (never the ideal) → cut list; TL/PO promotes selected items only.
|
|
51
|
+
- **SHIP S.0 / GATE H** — `/scope-hammer`: census (QA findings + discovered ledger + attempt-budget proposals ✦; every "no scope owns X" cites `probe owner`, which derives ownership from the contracts) → baseline comparison (never the ideal) → cut list; TL/PO promotes selected items only. The census is written as data beside the report, and GATE L4 reads that file — and nothing else — before any answer set may say `ship`. A breaker does not end the run: the run itself dispatches the census, crosses GATE H, writes the ship report with the verdict as it is, crosses L4, and only then closes — `shipped` with the cut list when the census clears it, `escalated` naming the census when it does not.
|
|
52
52
|
- ⏸ **L4** — Ship Sign-off (shows QA status ★).
|
|
53
53
|
- **Coach retro** — L4 feedback → `/coach`; GATE COACH-1 asks the PO which skill owns each rule (never assumes) → committed `shapeup/knowledge-base/<skill>.md` (team inherits on pull). Coachable: `/task-executor`, `/ba-pitch-analyzer`, `/qa-edge-hunter`, `/orient`, `/scope-architect`, `/solution-architect`, and `/tech-lead` (workflow guidance read at L0, plus suggested L0 values it confirms before pinning); `/spec-evaluator` is not (single judge) and neither is `/scope-hammer` (its census cites `probe owner`). **Guidance never decides a gate**: a rule may add a question or a check to a gate block, never an answer, a skip or a wider substrate. `/retro --scan` seeds the same files from the project on disk before the first run, and `/retro --research <stack>` from the platform's official documentation when there is nothing on disk yet (a source, never a verification: nothing it reads runs until the tech lead pins it at L0), optional, every rule confirmed at COACH-1. Mechanism defects file to `knowledge-base/harness-defects.md` as Betting Table raw ideas, never worker steering.
|
|
54
54
|
- Post-fix: `eval --single-pass` → remaining `~` + new feedback → new raw idea.
|
package/kernel/compile.mjs
CHANGED
|
@@ -175,6 +175,9 @@ export function ledgerDecisions(ledgerText, scopeId) {
|
|
|
175
175
|
*/
|
|
176
176
|
export const OP_OWNER = {
|
|
177
177
|
analyze: "ba-pitch-analyzer", reconcile: "ba-pitch-analyzer",
|
|
178
|
+
// `board` regenerates the per-machine board from a committed spec tree without re-deriving the
|
|
179
|
+
// tree — the half of ANALYZE that does not survive a clone. Same worker, narrower write surface.
|
|
180
|
+
board: "ba-pitch-analyzer",
|
|
178
181
|
"retrofit-surface": "ba-pitch-analyzer", coverage: "ba-pitch-analyzer",
|
|
179
182
|
"map-scopes": "scope-architect",
|
|
180
183
|
wire: "solution-architect", evaluate: "spec-evaluator", orient: "orient",
|
|
@@ -243,7 +246,10 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
|
|
|
243
246
|
// not cover is a leg forging a SIBLING's result, which is a real hole with a different fix — the
|
|
244
247
|
// glob form here cannot say "every result except this order's own", so closing it needs a
|
|
245
248
|
// mechanism rather than one more entry on this list.
|
|
246
|
-
|
|
249
|
+
// The board index is ingest's projection of the task results, not the doer's bookkeeping — the
|
|
250
|
+
// task files are. The executor's contract already forbids editing it; the freeze makes that a
|
|
251
|
+
// denial rather than a rule, on the same terms as the attestation channels.
|
|
252
|
+
const FROZEN_ATTESTATION = [`${local}/receipts/**`, `${local}/legs.jsonl`, `${local}/t0/verdicts/**`, `${local}/tasks/_index.md`];
|
|
247
253
|
switch (operation) {
|
|
248
254
|
case "execute": case "fix": case "spike":
|
|
249
255
|
// Build legs are the widest window on FROZEN_INTAKE, not an exemption from it: they are the
|
|
@@ -259,6 +265,14 @@ export function substrateFor(operation, { slug, specDir, scope } = {}) {
|
|
|
259
265
|
case "analyze":
|
|
260
266
|
return { allowed: [`${spec}/**`, `${local}/**`], frozen: [...FROZEN_INTAKE] };
|
|
261
267
|
|
|
268
|
+
case "board":
|
|
269
|
+
// The spec is READ, never written: this operation exists because the tree is already on disk
|
|
270
|
+
// and the board is not. Task ids renumber per machine by design, so nothing here is committed.
|
|
271
|
+
return {
|
|
272
|
+
allowed: [`${local}/tasks/**`, `${working}/**`],
|
|
273
|
+
frozen: [...FROZEN_SPEC_CORE, ...FROZEN_INTAKE, `${spec}/usecases/*.md`, `${spec}/scope-summary.md`],
|
|
274
|
+
};
|
|
275
|
+
|
|
262
276
|
case "reconcile":
|
|
263
277
|
return {
|
|
264
278
|
allowed: [`${local}/tasks/**`, `${spec}/scope-summary.md`, `${working}/**`],
|
package/kernel/gate.mjs
CHANGED
|
@@ -51,9 +51,10 @@
|
|
|
51
51
|
// 2 usage / validation error
|
|
52
52
|
|
|
53
53
|
import { readFileSync, writeFileSync, appendFileSync, existsSync, mkdirSync } from "node:fs";
|
|
54
|
+
import { parseBoard } from "./reduce/board.mjs";
|
|
54
55
|
import { join, dirname } from "node:path";
|
|
55
56
|
import { runArgs } from "./lib/argv.mjs";
|
|
56
|
-
import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir } from "./lib/paths.mjs";
|
|
57
|
+
import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir, tasksDir, hammerCensus } from "./lib/paths.mjs";
|
|
57
58
|
|
|
58
59
|
export const GATE_IDS = ["L0", "L1a", "L1a.5", "L1b", "L2", "L3", "QA", "H", "L4", "COACH-1"];
|
|
59
60
|
|
|
@@ -110,7 +111,7 @@ export const PRESETS = {
|
|
|
110
111
|
"L1a": { decision: "proceed", note: "Orient review — advisory read." },
|
|
111
112
|
"L1a.5": { decision: "proceed", note: "Wiring review — checked by trace-lint." },
|
|
112
113
|
"L1b": { decision: "ask", note: "Board review is where scope is actually decided. Not pre-approvable." },
|
|
113
|
-
"L2": { decision: "proceed", note: "The board facts travel in the gate block itself — green_scopes and hammer_proposals — so a preset answering here is not answering blind.
|
|
114
|
+
"L2": { decision: "proceed", note: "The board facts travel in the gate block itself — green_scopes and hammer_proposals — so a preset answering here is not answering blind. A board with zero tasks is refused here regardless of the answer — the resolver narrows `proceed` to `ask` over an empty board." },
|
|
114
115
|
"L3": { decision: "loop", max_rounds: 3, note: "Loop on FAIL; the breaker ends it." },
|
|
115
116
|
"QA": { decision: "run" },
|
|
116
117
|
"H": { decision: "ask", note: "The cut list changes what ships." },
|
|
@@ -186,6 +187,17 @@ export const HAMMER_VERDICTS = ["ship-now", "ship-after-fixes", "cannot-ship"];
|
|
|
186
187
|
*/
|
|
187
188
|
export function censusVerdict(cwd, slug) {
|
|
188
189
|
if (!slug) return null;
|
|
190
|
+
// The census artifact — written by the hammer at a path its order's substrate permits. The
|
|
191
|
+
// WorkResult fallback below is kept for a lane that wrote one the old way; a WorkResult may not
|
|
192
|
+
// carry a verdict under the schema, so in practice the artifact is the only source.
|
|
193
|
+
try {
|
|
194
|
+
const c = hammerCensus(cwd, slug);
|
|
195
|
+
if (existsSync(c)) {
|
|
196
|
+
const r = JSON.parse(readFileSync(c, "utf8"));
|
|
197
|
+
const v = r?.verdict ?? null;
|
|
198
|
+
if (HAMMER_VERDICTS.includes(v)) return v;
|
|
199
|
+
}
|
|
200
|
+
} catch { /* unreadable census proves nothing — fall through */ }
|
|
189
201
|
try {
|
|
190
202
|
const p = join(resultsDir(cwd, slug), "hammer.json");
|
|
191
203
|
if (!existsSync(p)) return null;
|
|
@@ -218,6 +230,26 @@ export function censusVerdict(cwd, slug) {
|
|
|
218
230
|
* @returns {object} The result, or an `ask` carrying why `ship` was not available.
|
|
219
231
|
*/
|
|
220
232
|
export function narrowToEvidence(r, cwd, slug) {
|
|
233
|
+
// GATE L2's question is "is the board done", and a preset used to answer it over a board with
|
|
234
|
+
// zero tasks — 100% of nothing. Measured on a consumer: a committed spec fast-forwarded ANALYZE,
|
|
235
|
+
// the per-machine board was never regenerated, L2 crossed `proceed` from the `ci` preset and the
|
|
236
|
+
// frozen report printed `Board 0/0 tasks done`. Same rule as L4 below: an answer set chooses
|
|
237
|
+
// among allowed answers and cannot supply the evidence that makes one allowed.
|
|
238
|
+
if (r?.gate === "L2" && r?.status === "ok" && r?.decision === "proceed" && slug) {
|
|
239
|
+
let tasks = 0;
|
|
240
|
+
try { tasks = parseBoard(tasksDir(cwd, slug)).length; } catch { tasks = 0; }
|
|
241
|
+
if (tasks === 0) {
|
|
242
|
+
return {
|
|
243
|
+
gate: r.gate, status: "ask", source: r.source, decision: "ask", note: r.note,
|
|
244
|
+
refused: "proceed", board_tasks: 0,
|
|
245
|
+
reason: `GATE L2 cannot be answered "proceed": the board has no tasks, so "board 100%" is a ` +
|
|
246
|
+
`hundred percent of nothing. A committed spec with no per-machine board needs the board ` +
|
|
247
|
+
`regenerated (operation \`board\`) before this gate means anything. Put the block to ` +
|
|
248
|
+
`the PO, or regenerate the board first.`,
|
|
249
|
+
};
|
|
250
|
+
}
|
|
251
|
+
return r;
|
|
252
|
+
}
|
|
221
253
|
if (r?.gate !== "L4" || r?.status !== "ok" || r?.decision !== "ship") return r;
|
|
222
254
|
const verdict = censusVerdict(cwd, slug);
|
|
223
255
|
if (verdict === "ship-now" || verdict === "ship-after-fixes") return r;
|
package/kernel/lib/paths.mjs
CHANGED
|
@@ -306,6 +306,16 @@ export const activeOrder = (cwd) => join(localDir(cwd), "active-order");
|
|
|
306
306
|
*/
|
|
307
307
|
export const lastRun = (cwd) => join(localDir(cwd), "last-run");
|
|
308
308
|
|
|
309
|
+
/**
|
|
310
|
+
* Scope-hammer's census as DATA — `{verdict, cut_list, …}` — the one artifact GATE L4's resolver
|
|
311
|
+
* reads before it lets an answer set say `ship`. The hammer's WorkResult may not carry a verdict
|
|
312
|
+
* (its census is a proposal, never envelope data ingest acts on), and its printed H0/H1/H2 blocks
|
|
313
|
+
* are prose; this file is the same proposal, written by the same hand, in a shape a gate can read.
|
|
314
|
+
* Measured before it existed: a green census returned as the worker's report, and L4 could only
|
|
315
|
+
* record `ask`, because nothing on disk said so.
|
|
316
|
+
*/
|
|
317
|
+
export const hammerCensus = (cwd, slug) => join(localRoot(cwd, slug), "reports", "hammer-census.json");
|
|
318
|
+
|
|
309
319
|
/**
|
|
310
320
|
* Where the run scripts are staged for launch, inside the project.
|
|
311
321
|
*
|
package/kernel/probe/leg.mjs
CHANGED
|
@@ -27,7 +27,7 @@
|
|
|
27
27
|
import { existsSync, readdirSync, readFileSync } from "node:fs";
|
|
28
28
|
import { join, resolve } from "node:path";
|
|
29
29
|
import { runArgs } from "../lib/argv.mjs";
|
|
30
|
-
import { ordersDir, resultsDir, legLedger } from "../lib/paths.mjs";
|
|
30
|
+
import { ordersDir, resultsDir, legLedger, dispatchReceipts, readRunId } from "../lib/paths.mjs";
|
|
31
31
|
|
|
32
32
|
/** `<scope>-r<round>-a<attempt>.json` — the only address a build order is written under. */
|
|
33
33
|
const BUILD_ORDER = /^(.+)-r(\d+)-a(\d+)\.json$/;
|
|
@@ -80,6 +80,12 @@ export function legsOf(cwd, slug) {
|
|
|
80
80
|
const ingested = new Set(readLegs(legLedger(cwd, slug))
|
|
81
81
|
.filter((r) => r.ingested_at)
|
|
82
82
|
.map((r) => String(r.order_id)));
|
|
83
|
+
// Dispatch receipts are the hook layer's record that a Skill was invoked against an order —
|
|
84
|
+
// scoped to this run, since receipts over one slug accumulate across runs and order ids repeat.
|
|
85
|
+
const runId = readRunId(cwd, slug);
|
|
86
|
+
const receipted = new Set(readLegs(dispatchReceipts(cwd, slug))
|
|
87
|
+
.filter((r) => r?.order_id && (!runId || !r.run_id || r.run_id === runId))
|
|
88
|
+
.map((r) => String(r.order_id)));
|
|
83
89
|
const out = [];
|
|
84
90
|
for (const f of (existsSync(oDir) ? readdirSync(oDir) : []).filter((x) => x.endsWith(".json")).sort()) {
|
|
85
91
|
let orderId = null;
|
|
@@ -88,6 +94,7 @@ export function legsOf(cwd, slug) {
|
|
|
88
94
|
order: join(oDir, f),
|
|
89
95
|
name: f.replace(/\.json$/, ""),
|
|
90
96
|
order_id: orderId,
|
|
97
|
+
has_receipt: orderId !== null && receipted.has(orderId),
|
|
91
98
|
has_result: existsSync(join(rDir, f)),
|
|
92
99
|
applied: orderId !== null && ingested.has(orderId),
|
|
93
100
|
});
|
|
@@ -104,10 +111,25 @@ export function legsOf(cwd, slug) {
|
|
|
104
111
|
*/
|
|
105
112
|
export function orderLegState(cwd, slug, name) {
|
|
106
113
|
const o = legsOf(cwd, slug).find((x) => x.name === name);
|
|
107
|
-
if (!o) return { closed: false, found: false, order: null, order_id: null, has_result: false, applied: false };
|
|
114
|
+
if (!o) return { closed: false, found: false, order: null, order_id: null, has_receipt: false, has_result: false, applied: false };
|
|
108
115
|
return { closed: o.applied, found: true, ...o };
|
|
109
116
|
}
|
|
110
117
|
|
|
118
|
+
/**
|
|
119
|
+
* Orders that were dispatched — a receipt says a Skill ran against them — and never answered:
|
|
120
|
+
* no WorkResult on disk. Measured on a live run: the first dispatch wrote its four artifacts and
|
|
121
|
+
* no envelope, every later phase read the artifacts, and the run closed `shipped` with the order
|
|
122
|
+
* still open by construction — the export's own dispatch row said `answered: false` and nothing
|
|
123
|
+
* upstream of it had noticed. Distinct from {@link openLegs}: that is work nobody applied; this is
|
|
124
|
+
* a leg that never came back.
|
|
125
|
+
* @param {string} cwd - Project root.
|
|
126
|
+
* @param {string} slug - Feature slug.
|
|
127
|
+
* @returns {{order:string, name:string, order_id:(string|null)}[]}
|
|
128
|
+
*/
|
|
129
|
+
export function unansweredOrders(cwd, slug) {
|
|
130
|
+
return legsOf(cwd, slug).filter((o) => o.has_receipt && !o.has_result);
|
|
131
|
+
}
|
|
132
|
+
|
|
111
133
|
/**
|
|
112
134
|
* The results on disk that no leg row applied — finished work the board never saw.
|
|
113
135
|
* @param {string} cwd - Project root.
|
|
@@ -171,8 +193,13 @@ export function cli(rawArgv) {
|
|
|
171
193
|
// `--open`: every result nothing applied, across the whole run, any phase. Exit 0 when none.
|
|
172
194
|
if (args.open) {
|
|
173
195
|
const open = openLegs(cwd, args.slug);
|
|
174
|
-
|
|
175
|
-
|
|
196
|
+
const unanswered = unansweredOrders(cwd, args.slug);
|
|
197
|
+
console.log(JSON.stringify({
|
|
198
|
+
closed: open.length === 0 && unanswered.length === 0,
|
|
199
|
+
open_total: open.length, open: open.map((o) => o.order), open_ids: open.map((o) => o.order_id),
|
|
200
|
+
unanswered_total: unanswered.length, unanswered: unanswered.map((o) => o.order), unanswered_ids: unanswered.map((o) => o.order_id),
|
|
201
|
+
}));
|
|
202
|
+
process.exit(open.length === 0 && unanswered.length === 0 ? 0 : 1);
|
|
176
203
|
}
|
|
177
204
|
// `--order <stem>`: one order by file stem (`analyze`, `alpha-r1-a1`). Exit 0 when its leg closed.
|
|
178
205
|
if (args.order) {
|
package/kernel/probe/resume.mjs
CHANGED
|
@@ -61,8 +61,10 @@ import { dirname, join, resolve } from "node:path";
|
|
|
61
61
|
import { runArgs } from "../lib/argv.mjs";
|
|
62
62
|
import { splitFrontmatter, uncoerce } from "../lib/contract.mjs";
|
|
63
63
|
import { globToRegExp } from "../verify/spec.mjs";
|
|
64
|
-
import {
|
|
64
|
+
import { parseBoard } from "../reduce/board.mjs";
|
|
65
|
+
import { intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir, orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements, exportRunDir, lastRun, readRunId, tasksDir, gates, verdictsDir, roundBuildDir } from "../lib/paths.mjs";
|
|
65
66
|
import { evalVerdict } from "./eval.mjs";
|
|
67
|
+
import { deriveRounds } from "./rounds.mjs";
|
|
66
68
|
import { collectRun, writeRun } from "../report/export.mjs";
|
|
67
69
|
|
|
68
70
|
/** The run-state values `references/protocol.md` (Part 4 — State) defines. A typo'd status is a rejection,
|
|
@@ -371,6 +373,21 @@ export function usecasesPath(cwd, slug, specFolder) {
|
|
|
371
373
|
* @param {string|null} specFolder - The ledger's `spec_folder`, if it names one.
|
|
372
374
|
* @returns {boolean} True when at least one use case (not the index) is on disk.
|
|
373
375
|
*/
|
|
376
|
+
/**
|
|
377
|
+
* Is the per-machine board on disk with at least one task? ANALYZE writes two artifacts in two
|
|
378
|
+
* tiers — the spec tree, committed, and the board, gitignored — and a fast-forward that asked only
|
|
379
|
+
* about the committed half walked every later run on a machine, and every run in a fresh checkout,
|
|
380
|
+
* into a build with a spec and no board: GATE L2 crossed over 0/0 tasks, the evaluator could read no
|
|
381
|
+
* `covers:` clause, and the requirements projection printed "no evidence" on a round whose static
|
|
382
|
+
* criteria all passed. Measured on a live consumer, fourth run of one pitch.
|
|
383
|
+
* @param {string} cwd - Project root.
|
|
384
|
+
* @param {string} slug - Feature slug.
|
|
385
|
+
* @returns {boolean} True when the board directory holds at least one task file.
|
|
386
|
+
*/
|
|
387
|
+
export function hasBoard(cwd, slug) {
|
|
388
|
+
try { return parseBoard(tasksDir(cwd, slug)).length > 0; } catch { return false; }
|
|
389
|
+
}
|
|
390
|
+
|
|
374
391
|
export function hasSpecTree(cwd, slug, specFolder) {
|
|
375
392
|
const dir = usecasesPath(cwd, slug, specFolder);
|
|
376
393
|
if (!existsSync(dir)) return false;
|
|
@@ -385,7 +402,10 @@ export function hasSpecTree(cwd, slug, specFolder) {
|
|
|
385
402
|
*/
|
|
386
403
|
export const PHASE_ARTIFACT = {
|
|
387
404
|
orient: { fact: "has_orient_artifacts", artifact: "orient/{code-surface,discovered-seed,hill-signal}.md + spike-*.md" },
|
|
388
|
-
|
|
405
|
+
// BOTH HALVES. The spec tree is committed and survives a clone; the board is per-machine and does
|
|
406
|
+
// not. A phase is complete only when both are on disk — a committed spec with no board resumes AT
|
|
407
|
+
// analyze, where the workflow dispatches the board-only operation rather than the whole phase.
|
|
408
|
+
analyze: { fact: "has_spec_tree", also: "has_board", artifact: "spec/usecases/*.md + tasks/TASK-*.md" },
|
|
389
409
|
wire: { fact: "has_wiring_map", artifact: "wiring-map.md" },
|
|
390
410
|
"map-scopes": { fact: "scope_files", artifact: "scopes/*.md" },
|
|
391
411
|
};
|
|
@@ -399,8 +419,9 @@ export const PHASES = Object.keys(PHASE_ARTIFACT);
|
|
|
399
419
|
* @returns {boolean} True when the phase's artifact exists.
|
|
400
420
|
*/
|
|
401
421
|
export function phaseSatisfied(state, phase) {
|
|
402
|
-
const v = state[
|
|
403
|
-
|
|
422
|
+
const present = (k) => { const v = state[k]; return Array.isArray(v) ? v.length > 0 : Boolean(v); };
|
|
423
|
+
const p = PHASE_ARTIFACT[phase];
|
|
424
|
+
return present(p.fact) && (!p.also || present(p.also));
|
|
404
425
|
}
|
|
405
426
|
|
|
406
427
|
/**
|
|
@@ -461,6 +482,7 @@ export function deriveResumeState(cwd, slug) {
|
|
|
461
482
|
orient_dir: `.shapeup/${slug}/orient/`,
|
|
462
483
|
has_orient_artifacts: hasOrientArtifacts(cwd, slug),
|
|
463
484
|
has_spec_tree: hasSpecTree(cwd, slug, hr.spec_folder || null),
|
|
485
|
+
has_board: hasBoard(cwd, slug),
|
|
464
486
|
// A PLAIN FACT, DELIBERATELY NOT A PHASE. The requirements registry is dispatched once, before
|
|
465
487
|
// ANALYZE, and the orchestrator guards that one dispatch on this boolean. It is NOT an entry in
|
|
466
488
|
// PHASE_ARTIFACT, and adding it there would be a migration hazard rather than a tidier shape:
|
|
@@ -539,7 +561,83 @@ export function setRunStatus(cwd, slug, status) {
|
|
|
539
561
|
* already normalized prose (newlines collapsed, truncated) — this function only `uncoerce`s it.
|
|
540
562
|
* @returns {string} The rewritten text.
|
|
541
563
|
*/
|
|
542
|
-
|
|
564
|
+
/**
|
|
565
|
+
* What the ledger's front matter and tables should say at a close, derived from the run's own
|
|
566
|
+
* records — the same ones the export reads. The counters (`final_verdict`, `rounds_used`) and the
|
|
567
|
+
* Rounds/Decisions tables used to be hand-shaped placeholders nothing filled: a run closed
|
|
568
|
+
* `shipped` beside `final_verdict: ~`, `rounds_used: 0`, an empty Decisions table and seven rows in
|
|
569
|
+
* `gates.jsonl`. A reader who trusted the counters concluded no round ran. Measured live.
|
|
570
|
+
*
|
|
571
|
+
* @param {string} cwd - Project root.
|
|
572
|
+
* @param {string} slug - Feature slug.
|
|
573
|
+
* @returns {{final_verdict:string, rounds_used:(number|null), decisions:object[], roundRows:string[]}}
|
|
574
|
+
*/
|
|
575
|
+
export function deriveLedgerFacts(cwd, slug) {
|
|
576
|
+
const readJson = (p) => { try { return JSON.parse(readFileSync(p, "utf8")); } catch { return null; } };
|
|
577
|
+
const evalRows = [];
|
|
578
|
+
const rDir = resultsDir(cwd, slug);
|
|
579
|
+
if (existsSync(rDir)) {
|
|
580
|
+
for (const f of readdirSync(rDir)) {
|
|
581
|
+
const m = f.match(/^evaluate-r(\d+)\.json$/);
|
|
582
|
+
if (!m) continue;
|
|
583
|
+
const r = readJson(join(rDir, f));
|
|
584
|
+
const overall = r?.verdict?.overall;
|
|
585
|
+
if (overall) evalRows.push({ round: Number(m[1]), overall: String(overall), criteria: Array.isArray(r.verdict.criteria) ? r.verdict.criteria : [] });
|
|
586
|
+
}
|
|
587
|
+
}
|
|
588
|
+
evalRows.sort((a, b) => a.round - b.round);
|
|
589
|
+
const finalVerdict = evalRows.length ? evalRows[evalRows.length - 1].overall : "not-evaluated";
|
|
590
|
+
const decisions = [];
|
|
591
|
+
const gp = gates(cwd, slug);
|
|
592
|
+
if (existsSync(gp)) {
|
|
593
|
+
for (const line of readFileSync(gp, "utf8").split("\n")) {
|
|
594
|
+
if (!line.trim()) continue;
|
|
595
|
+
try { decisions.push(JSON.parse(line)); } catch { /* a torn line proves nothing */ }
|
|
596
|
+
}
|
|
597
|
+
}
|
|
598
|
+
const t0ByRound = {};
|
|
599
|
+
const vDir = verdictsDir(cwd, slug);
|
|
600
|
+
if (existsSync(vDir)) {
|
|
601
|
+
for (const f of readdirSync(vDir).filter((x) => x.endsWith(".json")).sort()) {
|
|
602
|
+
const v = readJson(join(vDir, f));
|
|
603
|
+
if (typeof v?.round !== "number") continue;
|
|
604
|
+
const t = (t0ByRound[v.round] ||= { green: 0, red: 0 });
|
|
605
|
+
if (v.overall === "green") t.green++; else t.red++;
|
|
606
|
+
}
|
|
607
|
+
}
|
|
608
|
+
const gateByRound = {};
|
|
609
|
+
const bDir = roundBuildDir(cwd, slug);
|
|
610
|
+
if (existsSync(bDir)) {
|
|
611
|
+
for (const f of readdirSync(bDir).sort()) {
|
|
612
|
+
const m = f.match(/^r(\d+)-t\d+\.json$/);
|
|
613
|
+
if (m) gateByRound[Number(m[1])] = readJson(join(bDir, f))?.overall ?? "?";
|
|
614
|
+
}
|
|
615
|
+
}
|
|
616
|
+
const roundNums = [...new Set([...Object.keys(t0ByRound), ...Object.keys(gateByRound), ...evalRows.map((e) => e.round)].map(Number))].sort((a, b) => a - b);
|
|
617
|
+
const roundRows = [];
|
|
618
|
+
for (const r of roundNums) {
|
|
619
|
+
const t0 = t0ByRound[r], bg = gateByRound[r], ev = evalRows.find((e) => e.round === r);
|
|
620
|
+
roundRows.push(`| Build | ${r} | ${t0 ? `T0 ${t0.green} green / ${t0.red} red` : "no T0 verdict"} | — | round build gate: ${bg ?? "not run"} |`);
|
|
621
|
+
if (ev) {
|
|
622
|
+
const p = ev.criteria.filter((c) => c?.verdict === "PASS").length, f = ev.criteria.filter((c) => c?.verdict === "FAIL").length;
|
|
623
|
+
roundRows.push(`| Eval | ${r} | ${ev.overall} | — | ${p} PASS / ${f} FAIL criteria |`);
|
|
624
|
+
}
|
|
625
|
+
}
|
|
626
|
+
let roundsUsed = null;
|
|
627
|
+
try { roundsUsed = deriveRounds(cwd, slug, null)?.rounds_used ?? null; } catch { roundsUsed = null; }
|
|
628
|
+
if (roundsUsed == null && roundNums.length) roundsUsed = roundNums[roundNums.length - 1];
|
|
629
|
+
return { final_verdict: finalVerdict, rounds_used: roundsUsed, decisions, roundRows };
|
|
630
|
+
}
|
|
631
|
+
|
|
632
|
+
/** Replace the rows of one markdown table (identified by its header line) with `rows`, keeping the header, the separator and any row the caller marks as kept. */
|
|
633
|
+
function rewriteTable(body, headerRe, rows, keepRow = null) {
|
|
634
|
+
return body.replace(headerRe, (m, header, sep, oldRows) => {
|
|
635
|
+
const kept = keepRow && oldRows ? oldRows.split("\n").filter((l) => l.startsWith(keepRow)) : [];
|
|
636
|
+
return header + sep + [...kept, ...rows].map((l) => `${l}\n`).join("");
|
|
637
|
+
});
|
|
638
|
+
}
|
|
639
|
+
|
|
640
|
+
function writeCloseLines(body, { status, closedAt, cause, derived = null }) {
|
|
543
641
|
const causeLine = `close_cause: ${uncoerce(cause || null)}`;
|
|
544
642
|
const closedStatusLine = `closed_status: ${status}`;
|
|
545
643
|
let out = body
|
|
@@ -553,6 +651,15 @@ function writeCloseLines(body, { status, closedAt, cause }) {
|
|
|
553
651
|
out = /^close_cause:.*$/m.test(out)
|
|
554
652
|
? out.replace(/^close_cause:.*$/m, causeLine)
|
|
555
653
|
: out.replace(/^closed_at:.*$/m, (m) => `${m}\n${causeLine}`);
|
|
654
|
+
if (derived) {
|
|
655
|
+
// THE COUNTERS AND TABLES ARE DERIVED, IN THE SAME PASS AS THE CLOSE LINE. Written from the
|
|
656
|
+
// run's own records so they cannot disagree with the export that reads the same records.
|
|
657
|
+
if (/^final_verdict:.*$/m.test(out)) out = out.replace(/^final_verdict:.*$/m, `final_verdict: ${derived.final_verdict}`);
|
|
658
|
+
if (typeof derived.rounds_used === "number" && /^rounds_used:.*$/m.test(out)) out = out.replace(/^rounds_used:.*$/m, `rounds_used: ${derived.rounds_used}`);
|
|
659
|
+
out = rewriteTable(out, /(\| Phase \| Round \| Result \| Duration \| Notes \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, derived.roundRows, "| Init");
|
|
660
|
+
const decisionRows = derived.decisions.map((g) => `| ${g.gate ?? "?"} | ${g.decision ?? g.status ?? "?"} | ${g.source ?? "?"} | ${String(g.note ?? "").replace(/\|/g, "\\|").replace(/\s+/g, " ").slice(0, 160)} |`);
|
|
661
|
+
out = rewriteTable(out, /(\| Gate \| Decision \| Source \| Note \|\n)(\|[-| ]+\|\n)((?:\|[^\n]*\n)*)/, decisionRows);
|
|
662
|
+
}
|
|
556
663
|
return out;
|
|
557
664
|
}
|
|
558
665
|
|
|
@@ -675,6 +782,9 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
675
782
|
// The Ship phase (`shapeup-run.js`) already exports a shipped run itself, before this call ever
|
|
676
783
|
// runs — Stage 2 adds the endings that wrote nothing, and leaves that path untouched.
|
|
677
784
|
const shouldExport = withExport && status !== "shipped";
|
|
785
|
+
// Derived once, from the run's own records, and written with the close line (see deriveLedgerFacts).
|
|
786
|
+
let derived = null;
|
|
787
|
+
try { derived = deriveLedgerFacts(cwd, slug); } catch { derived = null; }
|
|
678
788
|
|
|
679
789
|
/**
|
|
680
790
|
* Everything a close owes the checkout once the ledger line is written: export the run's
|
|
@@ -757,7 +867,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
757
867
|
// new line rather than lost, and the return says so explicitly.
|
|
758
868
|
const closedAt = new Date().toISOString();
|
|
759
869
|
const foldedCause = `${normCause || "no reason recorded"} — supersedes an earlier close recorded ${priorClosedAt} (cause: ${JSON.stringify(priorCause)})`.slice(0, 4000);
|
|
760
|
-
body = writeCloseLines(body, { status, closedAt, cause: foldedCause });
|
|
870
|
+
body = writeCloseLines(body, { status, closedAt, cause: foldedCause, derived });
|
|
761
871
|
try { writeFileSync(p, body); } catch (e) {
|
|
762
872
|
return { ok: false, path: p, status, reason: `could not write the ledger: ${e.message}` };
|
|
763
873
|
}
|
|
@@ -772,7 +882,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
772
882
|
}
|
|
773
883
|
|
|
774
884
|
const closedAt = new Date().toISOString();
|
|
775
|
-
body = writeCloseLines(body, { status, closedAt, cause: normCause });
|
|
885
|
+
body = writeCloseLines(body, { status, closedAt, cause: normCause, derived });
|
|
776
886
|
try { writeFileSync(p, body); } catch (e) {
|
|
777
887
|
return { ok: false, path: p, status, reason: `could not write the ledger: ${e.message}` };
|
|
778
888
|
}
|
package/kernel/reduce/ingest.mjs
CHANGED
|
@@ -187,11 +187,48 @@ export function setCheckbox(body, ac, checked) {
|
|
|
187
187
|
* @param {boolean} done - When true, rewrite the row's status emoji/word to done; false is a no-op.
|
|
188
188
|
* @returns {string} The board text with the matching row updated (unchanged when no row matches).
|
|
189
189
|
*/
|
|
190
|
-
|
|
190
|
+
/**
|
|
191
|
+
* Is this index line the row FOR `taskId` — its first cell — rather than a row that merely
|
|
192
|
+
* mentions it? A task id appears in other rows' `Depends On` column, and matching by substring
|
|
193
|
+
* anywhere in the line flipped every dependent of a finished task to done along with it. Measured
|
|
194
|
+
* on a live run: the executor reported one task `skipped`, its task file still said `ready`, and
|
|
195
|
+
* the index showed it ✅ because the task it depended on had just been ticked. The census read the
|
|
196
|
+
* index.
|
|
197
|
+
* @param {string} line - One line of `tasks/_index.md`.
|
|
198
|
+
* @param {string} taskId - `TASK-NNN`.
|
|
199
|
+
* @returns {boolean}
|
|
200
|
+
*/
|
|
201
|
+
export function rowIs(line, taskId) {
|
|
202
|
+
const id = taskId.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
|
203
|
+
return new RegExp(`^\\|\\s*(?:\\[\\[)?${id}(?:\\\\\\|${id}\\]\\])?\\s*\\|`).test(line);
|
|
204
|
+
}
|
|
205
|
+
|
|
206
|
+
/**
|
|
207
|
+
* The status cell for one TaskResult status, or null when the index should not change.
|
|
208
|
+
* @param {(string|boolean)} status - A TaskResult status (`true` is accepted as `done` for callers written before skipped existed).
|
|
209
|
+
* @returns {({icon:string, word:string}|null)} The icon and word the index row should carry.
|
|
210
|
+
*/
|
|
211
|
+
export function boardCell(status) {
|
|
212
|
+
if (status === true || status === "done") return { icon: "✅", word: "done" };
|
|
213
|
+
if (status === "skipped") return { icon: "⏭", word: "skipped" };
|
|
214
|
+
if (status === "partial" || status === "failed") return { icon: "🔄", word: "in-progress" };
|
|
215
|
+
return null;
|
|
216
|
+
}
|
|
217
|
+
|
|
218
|
+
/**
|
|
219
|
+
* Rewrite one task's status cell in `tasks/_index.md` — the row whose id cell is `taskId`, never a
|
|
220
|
+
* row that merely mentions it in `Depends On`.
|
|
221
|
+
* @param {string} indexBody - The index file's text.
|
|
222
|
+
* @param {string} taskId - `TASK-NNN`.
|
|
223
|
+
* @param {(string|boolean)} status - A TaskResult status; statuses the index does not render leave it unchanged.
|
|
224
|
+
* @returns {string} The rewritten index text.
|
|
225
|
+
*/
|
|
226
|
+
export function updateBoardRow(indexBody, taskId, status) {
|
|
227
|
+
const cell = boardCell(status);
|
|
228
|
+
if (!cell) return indexBody;
|
|
191
229
|
return indexBody.split(/\r?\n/).map((line) => {
|
|
192
|
-
if (!line
|
|
193
|
-
|
|
194
|
-
return line;
|
|
230
|
+
if (!rowIs(line, taskId)) return line;
|
|
231
|
+
return line.replace(/⬜|🔄|⏳|🚫|✅|⏭/g, cell.icon).replace(/\b(ready|in-progress|blocked|done|skipped)\b/gi, cell.word);
|
|
195
232
|
}).join("\n");
|
|
196
233
|
}
|
|
197
234
|
|
|
@@ -240,14 +277,18 @@ function applyResultLocked(result, { cwd, slug }) {
|
|
|
240
277
|
body = setFrontmatter(body, "completed_at", today());
|
|
241
278
|
} else if (tr.status === "partial" || tr.status === "failed") {
|
|
242
279
|
body = setFrontmatter(body, "status", "in-progress");
|
|
280
|
+
} else if (tr.status === "skipped") {
|
|
281
|
+
// Rendered as what it is. A skipped task used to leave its file at `ready` and, through the
|
|
282
|
+
// substring match above, could show ✅ on the index — the census read the index.
|
|
283
|
+
body = setFrontmatter(body, "status", "skipped");
|
|
243
284
|
}
|
|
244
285
|
// Execution Log (append; the checkbox list must never disagree with it).
|
|
245
286
|
const logLines = (tr.ac_results || []).map((a) => `- ${a.ac}: ${a.result}${a.evidence ? ` (${a.evidence})` : ""}`).join("\n");
|
|
246
287
|
body += `\n\n## Execution Log — ${today()} (${result.order_id})\n- executor: ${result.worker || "task-executor"} via ingest-result\n- status: ${tr.status}\n${logLines}${tr.notes ? `\n- notes: ${tr.notes}` : ""}\n`;
|
|
247
288
|
writeFileSync(path, body);
|
|
248
289
|
summary.tasks_updated.push(tr.task_id);
|
|
249
|
-
if (
|
|
250
|
-
writeFileSync(boardIndex, updateBoardRow(readFileSync(boardIndex, "utf8"), tr.task_id,
|
|
290
|
+
if (existsSync(boardIndex) && boardCell(tr.status)) {
|
|
291
|
+
writeFileSync(boardIndex, updateBoardRow(readFileSync(boardIndex, "utf8"), tr.task_id, tr.status));
|
|
251
292
|
}
|
|
252
293
|
}
|
|
253
294
|
|
|
@@ -271,7 +312,7 @@ function applyResultLocked(result, { cwd, slug }) {
|
|
|
271
312
|
summary.unblocked.push(t.id);
|
|
272
313
|
if (existsSync(boardIndex)) {
|
|
273
314
|
const idx = readFileSync(boardIndex, "utf8").split(/\r?\n/).map((line) =>
|
|
274
|
-
line
|
|
315
|
+
rowIs(line, t.id)
|
|
275
316
|
? line.replace(/🚫|⏳/g, "⬜").replace(/\bblocked\b/gi, "ready")
|
|
276
317
|
: line).join("\n");
|
|
277
318
|
writeFileSync(boardIndex, idx);
|
|
@@ -494,7 +494,7 @@
|
|
|
494
494
|
"execute",
|
|
495
495
|
"fix",
|
|
496
496
|
"spike",
|
|
497
|
-
"analyze",
|
|
497
|
+
"analyze", "board",
|
|
498
498
|
"reconcile",
|
|
499
499
|
"retrofit-surface",
|
|
500
500
|
"coverage",
|
|
@@ -2537,6 +2537,22 @@
|
|
|
2537
2537
|
}
|
|
2538
2538
|
}
|
|
2539
2539
|
},
|
|
2540
|
+
"HammerCensus": {
|
|
2541
|
+
"description": "Scope-hammer's census as data — the proposal GATE L4's resolver reads before it lets an answer set say `ship`. Written by the hammer at `.shapeup/<slug>/reports/hammer-census.json`, the only path its order's substrate permits besides the committed report. A proposal, never a decision: promotion and shipping stay the PO's call, and ingest never acts on it.",
|
|
2542
|
+
"x-tier": "LOCAL",
|
|
2543
|
+
"x-location": ".shapeup/<slug>/reports/hammer-census.json",
|
|
2544
|
+
"type": "object",
|
|
2545
|
+
"required": ["schema_version", "verdict", "cut_list"],
|
|
2546
|
+
"properties": {
|
|
2547
|
+
"schema_version": { "const": 1 },
|
|
2548
|
+
"order_id": { "type": "string" },
|
|
2549
|
+
"verdict": { "type": "string", "enum": ["ship-now", "ship-after-fixes", "cannot-ship"] },
|
|
2550
|
+
"cut_list": { "type": "array", "items": { "type": "string" } },
|
|
2551
|
+
"ship_blocking": { "type": "array", "items": { "type": "string" }, "description": "Must-have items that failed the baseline comparison — non-empty exactly when the verdict is cannot-ship." },
|
|
2552
|
+
"breaker": { "type": ["string", "null"] },
|
|
2553
|
+
"baseline": { "type": ["string", "null"], "description": "The baseline the comparison ran against, or null when the pitch's problem statement stood in for it." }
|
|
2554
|
+
}
|
|
2555
|
+
},
|
|
2540
2556
|
"ResumeState": {
|
|
2541
2557
|
"description": "The fast-forward derivation: which phase a launch resumes at, derived from artifacts on disk and NEVER from stored state or conversation memory. Produced by kernel/probe/resume.mjs on stdout and consumed by shapeup-run.js's preamble on every launch, fresh or relaunch alike. Every phase predicate is an artifact test — `has_orient_artifacts` exists because the ORIENT branch was once gated on the ledger's stored `status` instead, which a silent courier failure left stale, and a completed ORIENT phase was re-dispatched on resume. `next_phase` is a convenience derived from the same booleans, which all travel too: a caller is never forced to trust a summary it cannot re-derive.",
|
|
2542
2558
|
"x-tier": "EMBEDDED",
|
|
@@ -2635,6 +2651,7 @@
|
|
|
2635
2651
|
"type": "boolean",
|
|
2636
2652
|
"description": "ANALYZE finished: the spec folder's usecases/ carries at least one use case that is not _index.md. WIRE reads these — one wiring-map entry per use case — which is why ANALYZE precedes WIRE in the phase chain: dispatched against an empty spec folder, WIRE escalates on every launch."
|
|
2637
2653
|
},
|
|
2654
|
+
"has_board": { "type": "boolean", "description": "The per-machine board (tasks/TASK-*.md) holds at least one task. ANALYZE is complete only with both the committed spec tree and this; a committed tree with no board resumes at analyze, where the board-only operation regenerates it." },
|
|
2638
2655
|
"has_requirements": {
|
|
2639
2656
|
"type": "boolean",
|
|
2640
2657
|
"description": "The requirements registry is on disk: shapeup/<slug>/requirements.md exists. A PLAIN FACT, not a phase — the orchestrator guards its single `coverage` dispatch on this boolean, and it is deliberately absent from kernel/probe/resume.mjs's PHASE_ARTIFACT map, which doubles as nextPhase()'s ordered list: an entry there would fast-forward every run recorded before the registry existed to the registry instead of to build."
|
package/package.json
CHANGED
|
@@ -24,7 +24,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown; surface i
|
|
|
24
24
|
|
|
25
25
|
| Field | What it is |
|
|
26
26
|
|---|---|
|
|
27
|
-
| `operation` | `analyze` (pitch → full spec tree + board) · `reconcile` (fold discovered-ledger items into the board + UC invariants) · `retrofit-surface` (append `## Test Surface` to a pre-surface spec) · `coverage` (extract atomic requirement clauses → the SHARED `requirements.md` registry) |
|
|
27
|
+
| `operation` | `analyze` (pitch → full spec tree + board) · `board` (committed spec tree → the per-machine board only; the tree is frozen) · `reconcile` (fold discovered-ledger items into the board + UC invariants) · `retrofit-surface` (append `## Test Surface` to a pre-surface spec) · `coverage` (extract atomic requirement clauses → the SHARED `requirements.md` registry) |
|
|
28
28
|
| `payload.pitch` | The pitch/PRD path (analyze) |
|
|
29
29
|
| `payload.breadboard` | The breadboard (analyze): its Places are your screens; its U# and N# are the affordances you place and cite. Absent = none separate; never inferred |
|
|
30
30
|
| `payload.requirements` | (coverage) the REQ source to extract atomic clauses from — pitch / a customer-requirements doc / the use-case bodies. Absent → default to the pitch and record the choice in `assumptions[]` |
|
|
@@ -115,10 +115,11 @@ covered AC is a requirement the run can be measured against.
|
|
|
115
115
|
|
|
116
116
|
---
|
|
117
117
|
|
|
118
|
-
## The other
|
|
118
|
+
## The other four operations — same craft, different payload + whitelist
|
|
119
119
|
|
|
120
120
|
| Operation | Essence | Never |
|
|
121
121
|
|---|---|---|
|
|
122
|
+
| `board` | The spec tree is already on disk and committed; the board is per-machine and is not. Regenerate `tasks/` from the use cases as they are — same task derivation as `analyze`, every AC carrying its `(covers: REQ-…)` clause, ids numbered fresh for this machine | touch the spec folder (frozen), re-derive or "improve" a use case, invent a task no UC step or Test Surface row sources |
|
|
122
123
|
| `reconcile` | Verify `ledger.feature == payload.feature` (mismatch → STOP). Map each `[+]` Keep item → its owning UC; new task continues numbering (never renumber); `~`/Cut → synthesis "Hammered Out" row, no file. A Keep item asserting a new invariant → APPEND `[INV-NN]` + TS-INV row to that UC (append-only sections in your substrate). A new actor/action with no UC → `status: "escalated"` + a `deviations[]` spec-ambiguity entry: spawning a UC mid-cycle is silent re-shaping, the PO decides. Finish with board-derive (appetite overflow → report) + spec-lint | re-run phases 1–5; edit UC Steps; resolve the appetite HAMMER yourself |
|
|
123
124
|
| `retrofit-surface` | Append `## Test Surface` (derived rows only, after Error Cases) to each UC of a pre-surface spec; an all-sources-empty UC gets the explicit empty-sources line | touch anything else — append-only substrate |
|
|
124
125
|
| `coverage` | Extract **atomic** customer requirement clauses from `payload.requirements` (default: the pitch) and write the SHARED `shapeup/<slug>/requirements.md` registry: one `\| REQ-id \| clause (verbatim) \| source \| status \| note \|` row per clause. **The `source` cell, and any other prose in this file, cites the committed pitch or shaping doc — never the gitignored run-tier path** (`.shapeup/<slug>/intake.md`, or any other `.shapeup/` path): spec-lint's `TIER-DIRECTION` rule reds a committed file naming that path in ANY form, a bare path in a sentence exactly as much as a `[[tasks/...]]` wikilink, because it dangles on every other clone. When the intake has no committed original to name, describe the run tier without a path (`"the pitch staged for this run"`) rather than citing where it actually lives. Split compound sentences into one testable clause each — a clause lost *inside* a bigger sentence is a requirement nothing can be traced to. **Assign REQ-ids ONCE and freeze them** (they behave like scope_id, never TASK-NNN — every `covers:` link rots otherwise): re-running, append new clauses with fresh ids, mark a removed clause `CUT (PO-approved)`, never renumber or delete. Status starts `covered` (a live requirement); only the PO sets `CUT`. The REQ source itself is frozen — the registry is a separate derived file. **Numbering.** A source clause already carrying an `R<n>` keeps its number — `R12` → `REQ-12` — and its `source` cell records where it came from verbatim (`shaping.md R12`), because that cell is the only thing that survives a re-run. A clause with no R-id takes the next free number ABOVE the highest `R<n>` in the source, so it can never collide with one added later. Splitting a compound clause keeps `REQ-12` for the first atomic part and records `shaping.md R12 (split 2/3)` for the rest — a requirement graded in parts is why splitting matters at all. On a re-run, match an existing id by its frozen `source` cell and clause text, **never** by re-deriving the number from the source's current order | edit the REQ source; renumber existing REQ-ids; delete a dropped clause instead of marking it CUT; invent a requirement not in the source; re-point an existing REQ-id because the source's R-numbers shifted |
|
|
@@ -171,6 +171,19 @@ The WorkResult may carry only `files_touched`, `artifacts`, `assumptions`, `devi
|
|
|
171
171
|
(`x-result-by-worker`): the census, cut list, and ship verdict live in the report artifact as a
|
|
172
172
|
**proposal** — promotion and shipping stay a human call, never envelope data ingest acts on.
|
|
173
173
|
|
|
174
|
+
**The census is also written as data**, at `.shapeup/<slug>/reports/hammer-census.json` (inside
|
|
175
|
+
this operation's substrate), shape `HammerCensus` in the domain registry:
|
|
176
|
+
|
|
177
|
+
```json
|
|
178
|
+
{ "schema_version": 1, "order_id": "<the order>", "verdict": "ship-now | ship-after-fixes | cannot-ship",
|
|
179
|
+
"cut_list": ["…"], "ship_blocking": ["…"], "breaker": "outer | inner | deadline | null", "baseline": "<path or null>" }
|
|
180
|
+
```
|
|
181
|
+
|
|
182
|
+
GATE L4's resolver reads this file — and nothing else — before it lets any answer set say `ship`.
|
|
183
|
+
It is the same proposal as the H2 block, written by the same hand, in a shape a gate can read:
|
|
184
|
+
a census that lives only in printed blocks cannot be read by a gate, and a run that reached L4
|
|
185
|
+
with a green census could only be recorded as `ask`. Still a proposal: the file grants nothing.
|
|
186
|
+
|
|
174
187
|
---
|
|
175
188
|
|
|
176
189
|
## Invocation
|
|
@@ -560,9 +560,12 @@ and without it the trace holds no record of that decision at all: `node
|
|
|
560
560
|
"${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" gate --resolve L4 --slug <slug>
|
|
561
561
|
[--file <path>|--preset <name>]`. Exit 0 (`decision=ship|hold`) — render the block above and close
|
|
562
562
|
the run: `node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe resume --slug <slug> --close shipped
|
|
563
|
-
--cause "verdict=<verdict> rounds=<r> decision=<ship|hold>"`.
|
|
564
|
-
|
|
565
|
-
|
|
563
|
+
--cause "verdict=<verdict> rounds=<r> decision=<ship|hold>"`. **This applies to the prose lane
|
|
564
|
+
only.** A launched run (`shapeup-run.js`) resolves L4 itself, on both the PASS path and a breaker
|
|
565
|
+
path: it dispatches the census, crosses GATE H, writes the ship report with the verdict as it is,
|
|
566
|
+
crosses L4, and only then closes — `shipped` when the census and L4 clear it, `escalated` naming
|
|
567
|
+
the census when they do not. A run that returned `{status: "shipped"}` or `{status: "gate_h"}` is
|
|
568
|
+
already closed; do not close it again, and do not run the census by hand after it.
|
|
566
569
|
The close itself is a once-only fact IN THE KERNEL (`closeRun`'s own guard reads a `closed_status:`
|
|
567
570
|
line that only `closeRun` ever writes — never the mutable `status:` line every phase rewrites, this
|
|
568
571
|
call included), not a conditional this instruction has to get right: if this run_id was NOT already
|
|
@@ -42,7 +42,9 @@
|
|
|
42
42
|
// { status: "paused", paused_at, block, valid_decisions, context }
|
|
43
43
|
// { status: "aborted", aborted_at, reason }
|
|
44
44
|
// { status: "gate_h", breaker: "outer"|"attempt_budget"|"none"|"deadline", hammer_proposals, green_scopes,
|
|
45
|
-
// tripped_scopes?, unapplied_results? }
|
|
45
|
+
// tripped_scopes?, unapplied_results?, census?, cut_list?, l4? }
|
|
46
|
+
// — returned only when the census said cannot-ship, L4 said hold, or the ship report was
|
|
47
|
+
// refused; a breaker whose census clears the run ends as `shipped` with `after: "gate_h"`.
|
|
46
48
|
|
|
47
49
|
// meta must be a PURE LITERAL — the runtime parses it statically, before the body ever runs, and
|
|
48
50
|
// rejects the whole script on anything it has to evaluate. A `+`-joined description is a
|
|
@@ -411,6 +413,7 @@ const RESUME = {
|
|
|
411
413
|
eval_dimensions: { type: "array", items: { type: "string" } },
|
|
412
414
|
has_orient_artifacts: { type: "boolean" },
|
|
413
415
|
has_spec_tree: { type: "boolean" },
|
|
416
|
+
has_board: { type: "boolean" },
|
|
414
417
|
// The requirements registry — a fact, not a phase. See the COVERAGE block below for why it is
|
|
415
418
|
// guarded on this bare boolean and never asked about through `probe resume --require`.
|
|
416
419
|
has_requirements: { type: "boolean" },
|
|
@@ -479,6 +482,7 @@ const ORDERLEG = {
|
|
|
479
482
|
closed: { type: "boolean" },
|
|
480
483
|
found: { type: "boolean" },
|
|
481
484
|
order: nullable("string"),
|
|
485
|
+
has_receipt: { type: "boolean" },
|
|
482
486
|
has_result: { type: "boolean" },
|
|
483
487
|
applied: { type: "boolean" },
|
|
484
488
|
},
|
|
@@ -624,6 +628,69 @@ const HAMMER = {
|
|
|
624
628
|
required: ["ok", "verdict", "cut_list"],
|
|
625
629
|
};
|
|
626
630
|
|
|
631
|
+
/** What every hammer dispatch is told, on the PASS path and the breaker path alike. */
|
|
632
|
+
const HAMMER_EXTRA =
|
|
633
|
+
"Run the census, compare against the BASELINE and never the ideal, and produce the cut list. " +
|
|
634
|
+
"Write the census as data to `reports/hammer-census.json` under the run's local root — the " +
|
|
635
|
+
"`reports/**` entry in your order's substrate names the directory — with schema_version 1, " +
|
|
636
|
+
"order_id, verdict (ship-now | ship-after-fixes | cannot-ship), cut_list, ship_blocking, breaker, " +
|
|
637
|
+
"baseline — beside the human-readable report. GATE L4's resolver reads that file and nothing " +
|
|
638
|
+
"else before it lets any answer set say `ship`; a census that lives only in your printed blocks " +
|
|
639
|
+
"cannot be read by a gate.";
|
|
640
|
+
|
|
641
|
+
/**
|
|
642
|
+
* A breaker routed the run to GATE H. This used to be a terminal return: the loop handed back
|
|
643
|
+
* `{status: "gate_h"}`, the close-out stamped the ledger `escalated` and retired the pointers, and
|
|
644
|
+
* the census, GATE H, the ship report and GATE L4 all happened AFTER the close, in the tech
|
|
645
|
+
* lead's prose — so the census reached no artifact, `gates.jsonl` held no H and no L4, and a later
|
|
646
|
+
* `--close shipped` was refused over the `escalated` fact already on the ledger. Measured on three
|
|
647
|
+
* consumer runs. AGENTS.md has always said what a breaker means: "ship what's green, never kill the
|
|
648
|
+
* run from outside". So the run does that itself, in this launch: the hammer census (an artifact),
|
|
649
|
+
* GATE H (a row), the ship report (with the verdict as it is, FAIL or not-evaluated included), and
|
|
650
|
+
* GATE L4 (a row) — and only then a close, `shipped` or `escalated`, that names the census.
|
|
651
|
+
*
|
|
652
|
+
* QA never ran on this path (it sits after a PASS), so the census is told so plainly. `qaFindings`
|
|
653
|
+
* is declared after the round loop and must not be read from here.
|
|
654
|
+
*
|
|
655
|
+
* @param {object} ret - The gate_h return the round loop built.
|
|
656
|
+
* @returns {Promise<object>} The RunReturn that actually ends the run.
|
|
657
|
+
*/
|
|
658
|
+
async function settleAtGateH(ret) {
|
|
659
|
+
phase("Ship");
|
|
660
|
+
const h = await worker({
|
|
661
|
+
skill: "scope-hammer", operation: "hammer", schema: HAMMER, phase: "Ship", label: "hammer",
|
|
662
|
+
payload: { feature: slug, qa_findings: 0, hammer_proposals: ret.hammer_proposals || [], breaker: ret.breaker ?? null },
|
|
663
|
+
extra: HAMMER_EXTRA,
|
|
664
|
+
});
|
|
665
|
+
if (h.__failed) return await withWarnings(diedAt("H", h));
|
|
666
|
+
{
|
|
667
|
+
const g = await crossGate("H", "Ship", ["accept-cut-list", "ship-all", "ask"],
|
|
668
|
+
{ verdict: h.verdict, cut_list: h.cut_list, breaker: ret.breaker ?? null, green_scopes: ret.green_scopes, hammer_proposals: ret.hammer_proposals });
|
|
669
|
+
if (g.stop) return await withWarnings(g.stop);
|
|
670
|
+
}
|
|
671
|
+
if (h.verdict === "cannot-ship") {
|
|
672
|
+
return await withWarnings({ ...ret, census: "cannot-ship", cut_list: h.cut_list });
|
|
673
|
+
}
|
|
674
|
+
const hVerdict = verdict === "pass" ? "PASS" : verdict === "fail" ? "FAIL" : "not-evaluated";
|
|
675
|
+
const ship = await cmd(`reduce ship --slug ${slug} --verdict ${hVerdict} --qa skipped`, "Ship", "ship-report");
|
|
676
|
+
if (!ship.ok) {
|
|
677
|
+
return await withWarnings({ ...ret, census: h.verdict, cut_list: h.cut_list, ship_report: `refused: ${ship.detail || `exit ${ship.exit_code}`}` });
|
|
678
|
+
}
|
|
679
|
+
{
|
|
680
|
+
const g = await crossGate("L4", "Ship", ["ship", "hold", "ask"], { verdict: hVerdict, census: h.verdict, cut_list: h.cut_list, breaker: ret.breaker ?? null });
|
|
681
|
+
if (g.stop) return await withWarnings(g.stop);
|
|
682
|
+
if (g.decision === "hold") return await withWarnings({ ...ret, census: h.verdict, cut_list: h.cut_list, l4: "hold" });
|
|
683
|
+
}
|
|
684
|
+
await advisory(`report export --slug ${slug}`, "Ship", "export-run");
|
|
685
|
+
await setRunStatus("shipped", "Ship");
|
|
686
|
+
return await withWarnings({
|
|
687
|
+
status: "shipped", verdict, rounds_used: round, after: "gate_h", breaker: ret.breaker ?? null,
|
|
688
|
+
census: h.verdict, cut_list: h.cut_list, green_scopes: ret.green_scopes,
|
|
689
|
+
unapplied_results: ret.unapplied_results || [], qa_findings: 0,
|
|
690
|
+
report: ship.detail || `shapeup/${slug}/REPORT.md`,
|
|
691
|
+
});
|
|
692
|
+
}
|
|
693
|
+
|
|
627
694
|
// ---------------------------------------------------------------------------------------------
|
|
628
695
|
// DISPATCH — three shapes, and none of them parses text.
|
|
629
696
|
//
|
|
@@ -859,16 +926,29 @@ const attest = (phaseKey, phaseName, label) =>
|
|
|
859
926
|
* rather than folding the gap into a phase that "completed".
|
|
860
927
|
*
|
|
861
928
|
* @param {string} gate - The gate name to report the abort under.
|
|
862
|
-
* @param {string} phaseKey - The phase
|
|
929
|
+
* @param {string} phaseKey - The phase.
|
|
863
930
|
* @param {string} phaseName - Progress group.
|
|
931
|
+
* @param {string} [orderStem=phaseKey] - The order's file stem, when the phase was dispatched
|
|
932
|
+
* under a different operation (a board-only ANALYZE compiles `board.json`).
|
|
864
933
|
* @returns {Promise<(object|null)>} An aborted RunReturn, or null when the leg closed (or the
|
|
865
934
|
* question could not be asked — a probe that did not run proves nothing, and is logged as such).
|
|
866
935
|
*/
|
|
867
|
-
async function requireLeg(gate, phaseKey, phaseName) {
|
|
868
|
-
const ask = () => query(`probe leg --slug ${slug} --order "${
|
|
936
|
+
async function requireLeg(gate, phaseKey, phaseName, orderStem = phaseKey) {
|
|
937
|
+
const ask = () => query(`probe leg --slug ${slug} --order "${orderStem}"`, ORDERLEG, phaseName, `legcheck:${orderStem}`);
|
|
869
938
|
let leg = await ask();
|
|
870
939
|
if (!leg || !leg.found) { log(`${gate} — could not ask the leg ledger about "${phaseKey}" (probe returned ${leg ? "no order" : "nothing"}); proceeding on the artifact alone.`); return null; }
|
|
871
|
-
if (!leg.has_result
|
|
940
|
+
if (!leg.has_result) {
|
|
941
|
+
// The artifact is on disk and the envelope never came back. Measured live: the first dispatch of
|
|
942
|
+
// a run wrote its four artifacts and no WorkResult, every later phase read the artifacts, and
|
|
943
|
+
// the run closed `shipped` over an order still open by construction. The phase stands on its
|
|
944
|
+
// artifact — that is what the post-condition checks — but the order is named at the close, and
|
|
945
|
+
// the next run over this slug will find it unanswered rather than be surprised by it.
|
|
946
|
+
log(`${gate} — "${phaseKey}" ${leg.has_receipt ? "was dispatched (receipt on disk) and" : "has an order but no receipt, and"} never answered: no WorkResult at all. ` +
|
|
947
|
+
`Its artifacts are on disk and the phase proceeds on them; the order stays unanswered and is named at the close.`);
|
|
948
|
+
if (leg.order && !unansweredOrders.includes(leg.order)) unansweredOrders.push(leg.order);
|
|
949
|
+
return null;
|
|
950
|
+
}
|
|
951
|
+
if (leg.applied) return null;
|
|
872
952
|
log(`${gate} — "${phaseKey}" came back with a result nothing applied (no leg row). Ingesting it here: ${leg.order}.`);
|
|
873
953
|
await advisory(`reduce ingest --order "${leg.order}"`, phaseName, `late-ingest:${phaseKey}`);
|
|
874
954
|
leg = await ask();
|
|
@@ -891,9 +971,9 @@ async function requireLeg(gate, phaseKey, phaseName) {
|
|
|
891
971
|
* @param {string} phaseName - Progress group.
|
|
892
972
|
* @returns {Promise<(object|null)>} An aborted RunReturn, or null when the artifact is there.
|
|
893
973
|
*/
|
|
894
|
-
async function requirePhase(gate, phaseKey, phaseName) {
|
|
974
|
+
async function requirePhase(gate, phaseKey, phaseName, orderStem = phaseKey) {
|
|
895
975
|
const r = await attest(phaseKey, phaseName, `require:${phaseKey}`);
|
|
896
|
-
if (r.exit_code === 0) return await requireLeg(gate, phaseKey, phaseName);
|
|
976
|
+
if (r.exit_code === 0) return await requireLeg(gate, phaseKey, phaseName, orderStem);
|
|
897
977
|
// Exit 6 is `probe resume --require`'s OWN documented code for "the artifact really is not on
|
|
898
978
|
// disk" (kernel/probe/resume.mjs banner). Any other value — including -1, the courier's sentinel
|
|
899
979
|
// for a tool call that never ran — is not that predicate answering "no"; it is the predicate never
|
|
@@ -1016,6 +1096,9 @@ async function requireLaunchRecord() {
|
|
|
1016
1096
|
// it warns and continues — and the warning travels in the RunReturn, because a headless stdout
|
|
1017
1097
|
// carries only the final message and a diagnostic on a channel nobody reads is not a diagnostic.
|
|
1018
1098
|
const stateWarnings = [];
|
|
1099
|
+
// Orders a receipt says were dispatched and no WorkResult ever answered — a leg that did its craft
|
|
1100
|
+
// and never came back. Named at every close, whatever the close; never folded into "complete".
|
|
1101
|
+
const unansweredOrders = [];
|
|
1019
1102
|
async function setRunStatus(status, phaseName) {
|
|
1020
1103
|
const r = await cmd(`probe resume --slug ${slug} --set-status ${status}`, phaseName, `status:${status}`);
|
|
1021
1104
|
if (!r.ok) {
|
|
@@ -1069,11 +1152,14 @@ async function closeIfTerminal(ret) {
|
|
|
1069
1152
|
+ (Array.isArray(ret.tripped_scopes) ? ` tripped_scopes=${ret.tripped_scopes.length}` : "")
|
|
1070
1153
|
// A result the single writer never applied is named at the close, not folded into "not green".
|
|
1071
1154
|
+ (Array.isArray(ret.unapplied_results) && ret.unapplied_results.length ? ` unapplied_results=${ret.unapplied_results.length}` : "")
|
|
1072
|
-
|
|
1155
|
+
+ (ret.census ? ` census=${ret.census}` : "") + (ret.l4 ? ` l4=${ret.l4}` : "") + (ret.ship_report ? ` ship_report=${ret.ship_report}` : "")
|
|
1156
|
+
: `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}`
|
|
1157
|
+
+ (ret.after ? ` after=${ret.after} breaker=${ret.breaker ?? "?"} census=${ret.census ?? "?"} cut_list=${Array.isArray(ret.cut_list) ? ret.cut_list.length : "?"}` : "");
|
|
1073
1158
|
// `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
|
|
1074
1159
|
// a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
|
|
1075
1160
|
// stay silent for it exactly as they did when this file's own guard returned early.
|
|
1076
|
-
const
|
|
1161
|
+
const unanswered = unansweredOrders.length ? ` unanswered_orders=${unansweredOrders.length}` : "";
|
|
1162
|
+
const r = await cmd(`probe resume --slug ${slug} --close-arm ${ret.status} --cause "${causeArg(cause + unanswered)}"`, "Ship", `close:${ret.status}`);
|
|
1077
1163
|
if (!r.ok) {
|
|
1078
1164
|
const why = (r.detail || `exit ${r.exit_code}`).trim();
|
|
1079
1165
|
log(`RUN STATE — close(${ret.status}) did not take: ${why}. This return's own status and reason still ` +
|
|
@@ -1254,8 +1340,28 @@ if (!rs.has_spec_tree) {
|
|
|
1254
1340
|
const post = await requirePhase("ANALYZE", "analyze", "Analyze");
|
|
1255
1341
|
if (post) return await withWarnings(post);
|
|
1256
1342
|
await advisory(`reduce graph --slug ${slug}`, "Analyze", "graph:analyze");
|
|
1343
|
+
} else if (!rs.has_board) {
|
|
1344
|
+
// THE HALF THAT DOES NOT SURVIVE. The spec tree is committed; the board is per-machine and
|
|
1345
|
+
// gitignored, and ANALYZE writes both. A run after the first on a machine — and every run in a
|
|
1346
|
+
// fresh checkout — used to fast-forward on the committed half alone and build over no board:
|
|
1347
|
+
// GATE L2 crossed 0/0 tasks, the evaluator could read no `covers:` clause, and the requirements
|
|
1348
|
+
// projection said "no evidence" for a round whose static criteria all passed. The board-only
|
|
1349
|
+
// operation regenerates it from the tree without re-deriving the tree.
|
|
1350
|
+
log(`ANALYZE — spec tree on disk, no board: dispatching the board-only operation (slug ${slug})`);
|
|
1351
|
+
await setRunStatus("mapping", "Analyze");
|
|
1352
|
+
const b = await worker({
|
|
1353
|
+
skill: "ba-pitch-analyzer", operation: "board", schema: PHASE_OK, phase: "Analyze", label: "board",
|
|
1354
|
+
payload: { spec_folder: specFolder, feature: slug, lens: rs.lens },
|
|
1355
|
+
extra: "The spec tree is committed and FROZEN for this dispatch. Regenerate the per-machine board " +
|
|
1356
|
+
"under the run's tasks/ directory from the use cases on disk — every acceptance criterion " +
|
|
1357
|
+
"carrying its `(covers: REQ-…)` clause — and write nothing under the spec folder.",
|
|
1358
|
+
});
|
|
1359
|
+
if (b.__failed) return await withWarnings(diedAt("ANALYZE", b));
|
|
1360
|
+
const post = await requirePhase("ANALYZE", "analyze", "Analyze", "board");
|
|
1361
|
+
if (post) return await withWarnings(post);
|
|
1362
|
+
await advisory(`reduce graph --slug ${slug}`, "Analyze", "graph:board");
|
|
1257
1363
|
} else {
|
|
1258
|
-
const post = await fastForward("ANALYZE", "analyze", "Analyze", "spec tree already on disk");
|
|
1364
|
+
const post = await fastForward("ANALYZE", "analyze", "Analyze", "spec tree and board already on disk");
|
|
1259
1365
|
if (post) return await withWarnings(post);
|
|
1260
1366
|
}
|
|
1261
1367
|
|
|
@@ -1476,7 +1582,7 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1476
1582
|
const budget = await cmd(`verify budget --slug ${slug} --strict`, "Build", `budget:r${round}`);
|
|
1477
1583
|
if (budget.exit_code === 6) {
|
|
1478
1584
|
await advisory(`reduce hill --slug ${slug}`, "Build", "hill-derive");
|
|
1479
|
-
return await
|
|
1585
|
+
return await settleAtGateH({ status: "gate_h", breaker: "deadline", unapplied_results: allUnapplied, hammer_proposals: allHammer, green_scopes: allGreen });
|
|
1480
1586
|
}
|
|
1481
1587
|
|
|
1482
1588
|
log(`BUILD round ${round} — ${scopes.length} scope(s), up to ${maxParallelScopes} at once, attempt budget ${attemptBudget}`);
|
|
@@ -1640,7 +1746,7 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1640
1746
|
if (census?.tripped) tripped.push(sid);
|
|
1641
1747
|
}
|
|
1642
1748
|
await advisory(`reduce hill --slug ${slug}`, "Build", "hill-derive");
|
|
1643
|
-
return await
|
|
1749
|
+
return await settleAtGateH({
|
|
1644
1750
|
status: "gate_h", breaker: tripped.length ? "attempt_budget" : "none", stalled: "no_green",
|
|
1645
1751
|
tripped_scopes: tripped, unapplied_results: allUnapplied,
|
|
1646
1752
|
hammer_proposals: allHammer, green_scopes: allGreen,
|
|
@@ -1755,13 +1861,13 @@ while (verdict !== "pass" && round <= maxRounds) {
|
|
|
1755
1861
|
// SHIP" once EVAL is skipped, never spends another round waiting on a verdict nobody is producing.
|
|
1756
1862
|
if (verdict === "pass" || verdict === "not-evaluated") break; // → QA → GATE H → ship
|
|
1757
1863
|
if (g3.decision === "stop" || round >= maxRounds) {
|
|
1758
|
-
return await
|
|
1864
|
+
return await settleAtGateH({ status: "gate_h", breaker: "outer", unapplied_results: allUnapplied, hammer_proposals: allHammer, green_scopes: allGreen });
|
|
1759
1865
|
}
|
|
1760
1866
|
round += 1;
|
|
1761
1867
|
}
|
|
1762
1868
|
|
|
1763
1869
|
if (verdict !== "pass" && verdict !== "not-evaluated") {
|
|
1764
|
-
return await
|
|
1870
|
+
return await settleAtGateH({ status: "gate_h", breaker: "outer", unapplied_results: allUnapplied, hammer_proposals: allHammer, green_scopes: allGreen });
|
|
1765
1871
|
}
|
|
1766
1872
|
|
|
1767
1873
|
// ---- QA (post-PASS, pre-ship) — a level-up, never a gate. `--no-qa` answers it "skip". --------
|
|
@@ -1786,7 +1892,7 @@ phase("Ship");
|
|
|
1786
1892
|
const h = await worker({
|
|
1787
1893
|
skill: "scope-hammer", operation: "hammer", schema: HAMMER, phase: "Ship", label: "hammer",
|
|
1788
1894
|
payload: { feature: slug, qa_findings: qaFindings, hammer_proposals: allHammer },
|
|
1789
|
-
extra:
|
|
1895
|
+
extra: HAMMER_EXTRA,
|
|
1790
1896
|
});
|
|
1791
1897
|
if (h.__failed) return await withWarnings(diedAt("H", h));
|
|
1792
1898
|
if (h.verdict === "cannot-ship") {
|
|
@@ -1804,6 +1910,14 @@ if (h.verdict === "cannot-ship") {
|
|
|
1804
1910
|
// default). Hardcoding PASS here is exactly the silent upgrade protocol.md's Rules forbid.
|
|
1805
1911
|
const shipVerdict = verdict === "not-evaluated" ? "not-evaluated" : "PASS";
|
|
1806
1912
|
const ship = await cmd(`reduce ship --slug ${slug} --verdict ${shipVerdict} --qa ${qaRan ? "run" : "skipped"}`, "Ship", "ship-report");
|
|
1913
|
+
// GATE L4 HAS A CALL SITE. It used to be a line of prose in the tech lead's references — "resolve
|
|
1914
|
+
// the gate itself before any of the above" — and a run reached its ship decision with no ledger row
|
|
1915
|
+
// for it. The resolver narrows `ship` on the census artifact the hammer just wrote.
|
|
1916
|
+
{
|
|
1917
|
+
const g = await crossGate("L4", "Ship", ["ship", "hold", "ask"], { verdict: shipVerdict, census: h.verdict, cut_list: h.cut_list });
|
|
1918
|
+
if (g.stop) return await withWarnings(g.stop);
|
|
1919
|
+
if (g.decision === "hold") return await withWarnings({ status: "aborted", aborted_at: "L4", reason: `GATE L4 answered "hold" (census ${h.verdict}, verdict ${shipVerdict})` });
|
|
1920
|
+
}
|
|
1807
1921
|
await advisory(`report export --slug ${slug}`, "Ship", "export-run");
|
|
1808
1922
|
// The run's own concurrency, printed once where the records are complete and before the next run
|
|
1809
1923
|
// supersedes the trace. It is a projection over `receipts/dispatch.jsonl` and `legs.jsonl`, so it
|