shapeup-sdlc 3.6.0 → 3.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +3 -2
- package/kernel/compile.mjs +37 -1
- package/kernel/harness.mjs +8 -2
- package/kernel/probe/attempts.mjs +135 -0
- package/kernel/probe/resume.mjs +145 -19
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/SKILL.md +1 -1
- package/skills/ba-pitch-analyzer/references/doc-schemas.md +6 -0
- package/skills/scope-hammer/SKILL.md +10 -2
- package/skills/tech-lead/workflows/shapeup-run.js +38 -10
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.7.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -11,8 +11,9 @@ Skills and commands are named short throughout this file; every one of them reso
|
|
|
11
11
|
- Attested, not assumed: a dispatch leaves a receipt naming the skill that ran, and ingest refuses an orchestrated result that has none. A dispatch that fails is answered by the sub-agent improvising the craft, which every other check accepts — so "the artifact exists" is not evidence that the shipped skill produced it. `--no-receipt-check` is the way through when the receipt channel itself fails.
|
|
12
12
|
- GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
|
|
13
13
|
- Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
|
|
14
|
+
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
|
|
14
15
|
- **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
|
|
15
|
-
- The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round.
|
|
16
|
+
- The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
|
|
16
17
|
|
|
17
18
|
### Phase 1 — Shaping (`/shapeup`)
|
|
18
19
|
1. Set Boundaries → `/shapeup shaping`
|
|
@@ -80,7 +81,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
|
|
|
80
81
|
- Two storage tiers (ADR-0001): COMMITTED `shapeup/<slug>/` (shaping, spec, scopes, wiring-map, project-profile, requirements, hill, `REPORT.md` frozen at L4) vs GITIGNORED `.shapeup/` (board, orders/results, T0/eval/QA artifacts, ledgers, metrics, gate answers, and the run scripts staged for launch).
|
|
81
82
|
- **The run launches from a copy inside your project, and it has to.** The Workflow tool loads a script only from a directory the session may already read; the plugin installs outside your project, so naming the shipped path is refused before the run begins and no permission rule repairs it — the grant authorises the tool, not what it may read. Opening a run therefore re-copies the run scripts to `.shapeup/workflows/` and reports the path the launch names. A run in flight keeps the copy it started with: an upgrade reaches the next run, not the current round.
|
|
82
83
|
- **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — can leave one still open, and while the run's pointer is still on disk the fence holds on that order exactly as if the run were live. The pointer is what actually switches it off, though: the fence is enforced only while that pointer exists on disk, so a close that removes it releases the fence whatever the order set still says — restoring the pointer flips it straight back to denying. A **ship** close does not itself answer what it leaves outstanding — a ship close can retire the pointer over an order that never got a result — so it is not special because nothing is left unanswered; it is special only because retiring the pointer is the one lever every close needs pulled, and `reduce ship` pulls it for you. Whichever way a run closes, an order it leaves unanswered stays genuinely unresolved — not merely un-fenced — until `init run --force` runs: it writes a synthetic result for every order the closed run left unanswered, so the fence lifts without waiting on a worker that is never coming back. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
|
|
83
|
-
- Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows, agent-call journal rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded
|
|
84
|
+
- Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows, agent-call journal rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded — and so does **every other terminal ending**: a run that aborts, escalates or stops at GATE H exports on its close, because the runs whose records are worth most are the ones that never reach the Ship phase. The export is advisory at every one of them: a failed export degrades the trace and is reported in the run's state warnings, but never turns a close into a non-close. A WorkResult carries no `run_id` and reaches it through `order_id`.
|
|
84
85
|
- Every run projects a **run graph** — `.shapeup/<slug>/graph.jsonl`, append-only, written only by
|
|
85
86
|
`reduce graph`. Two families kept separate: work lineage (Run, Order, Result, Verdict, Trial,
|
|
86
87
|
GateDecision) and domain (Scope, UseCase, Requirement, Seam). It is derived from the artifacts,
|
package/kernel/compile.mjs
CHANGED
|
@@ -31,7 +31,7 @@ import { fileURLToPath } from "node:url";
|
|
|
31
31
|
import { validate } from "./verify/envelope.mjs";
|
|
32
32
|
import { readTrials } from "./verify/t0.mjs";
|
|
33
33
|
import { runArgs } from "./lib/argv.mjs";
|
|
34
|
-
import { readRunId } from "./lib/paths.mjs";
|
|
34
|
+
import { readRunId, dispatchReceipts, legLedger } from "./lib/paths.mjs";
|
|
35
35
|
// `specDir` is aliased: this module has a local `let specDir` holding the resolved, possibly
|
|
36
36
|
// --spec-overridden directory, and the import is the convention-derived default.
|
|
37
37
|
import {
|
|
@@ -41,6 +41,8 @@ import {
|
|
|
41
41
|
import { readContract, readAllContracts, tasksForScope, SCOPE_CONTRACT } from "./lib/contract.mjs";
|
|
42
42
|
import { writeActiveOrder } from "./probe/resume.mjs";
|
|
43
43
|
import { greenVerdict } from "./probe/t0.mjs";
|
|
44
|
+
import { attemptEvidence, readReceipts } from "./probe/attempts.mjs";
|
|
45
|
+
import { readLegs } from "./probe/leg.mjs";
|
|
44
46
|
import { latestRoundBuild } from "./verify/build.mjs";
|
|
45
47
|
// The SAME matcher the sandbox hook enforces with. "Is this cited file inside this scope's
|
|
46
48
|
// substrate" has to mean exactly what the guard means, or a bug is addressed to a scope that is
|
|
@@ -816,6 +818,40 @@ export async function cli(rawArgv) {
|
|
|
816
818
|
const round = flag("round");
|
|
817
819
|
const attempt = flag("attempt");
|
|
818
820
|
|
|
821
|
+
// AN ATTEMPT MAY NOT OPEN OVER AN UNANSWERED ONE — the write half of the attested-channel rule.
|
|
822
|
+
//
|
|
823
|
+
// `--attempt` is a flag: the CALLER picks the number, and nothing here used to check that the
|
|
824
|
+
// previous one had come back. Measured on a consumer run: attempt 2 was compiled and T0-verified
|
|
825
|
+
// while attempt 1 was still in flight, so the round graded the tree attempt 1 was still writing,
|
|
826
|
+
// counted the attempt, and stopped at GATE H with most of its budget and clock unspent. An order
|
|
827
|
+
// and a T0 verdict are both writable by the very scope being judged; only a dispatch receipt, a
|
|
828
|
+
// leg row or a WorkResult attests that work actually happened.
|
|
829
|
+
//
|
|
830
|
+
// FAILS OPEN, NEVER CLOSED, unless the bad state is positively proven. The proof required is the
|
|
831
|
+
// receipts ledger EXISTING while carrying no row for the previous attempt: a lane that does not
|
|
832
|
+
// attest dispatches at all (no ledger on disk) cannot be judged by this rule and is waved
|
|
833
|
+
// through, so `--tiny`, a prose round loop and a standalone build are untouched.
|
|
834
|
+
if (scope?.scope_id && round && attempt > 1) {
|
|
835
|
+
const receiptsPath = dispatchReceipts(cwd, slug);
|
|
836
|
+
if (existsSync(receiptsPath)) {
|
|
837
|
+
const prev = attemptEvidence(
|
|
838
|
+
cwd, slug, scope.scope_id, round, attempt - 1,
|
|
839
|
+
readReceipts(receiptsPath), readLegs(legLedger(cwd, slug)),
|
|
840
|
+
);
|
|
841
|
+
if (prev.state !== "spent") {
|
|
842
|
+
const why = prev.state === "unattested"
|
|
843
|
+
? "no dispatch receipt was ever written for it"
|
|
844
|
+
: "it was dispatched but has neither a leg-completion row nor a WorkResult";
|
|
845
|
+
console.error(
|
|
846
|
+
`compile-order: refusing to open attempt ${attempt} for "${scope.scope_id}" in round ${round} — ` +
|
|
847
|
+
`attempt ${attempt - 1} (${prev.orderId}) is unanswered: ${why}. Grading a tree the previous ` +
|
|
848
|
+
`attempt may still be writing counts an attempt that never ran, and spends a budget on work ` +
|
|
849
|
+
`nobody did. Wait for it to return, or record its outcome, before opening the next one.`);
|
|
850
|
+
process.exit(3);
|
|
851
|
+
}
|
|
852
|
+
}
|
|
853
|
+
}
|
|
854
|
+
|
|
819
855
|
// The fix round's inbound evidence. Derived here, from the ledgered verdict, for every lane —
|
|
820
856
|
// the workflow, `--tiny`, the prose round loop and a standalone `/build` all compile through
|
|
821
857
|
// this line, and none of them can pass a payload to a build order (see the banner above).
|
package/kernel/harness.mjs
CHANGED
|
@@ -32,7 +32,7 @@
|
|
|
32
32
|
// gate An answer file with a source, not a vibe.
|
|
33
33
|
// probe resume · t0 · stats · digest · Read-only queries over run state. `concurrency`
|
|
34
34
|
// concurrency · leg · eval · answers how many legs ran at once and what the
|
|
35
|
-
// owner · requirements
|
|
35
|
+
// owner · requirements · attempts
|
|
36
36
|
// fan-out bought, and refuses a figure the record set
|
|
37
37
|
// cannot support rather than printing a plausible one.
|
|
38
38
|
// `leg` answers whether a scope's work reached the
|
|
@@ -50,6 +50,12 @@
|
|
|
50
50
|
// answers which pitch clause a verdict reached, joined
|
|
51
51
|
// through the plan's own covers: edge — the L4 line and
|
|
52
52
|
// GATE H's census cite it for the same reason.
|
|
53
|
+
// `attempts` answers how many of a scope's attempts are
|
|
54
|
+
// ATTESTED (a dispatch receipt AND a leg row or a
|
|
55
|
+
// WorkResult), never the order set or the T0 verdict
|
|
56
|
+
// set alone — the round loop's inner breaker and
|
|
57
|
+
// scope-hammer's census both cite it, so they cannot
|
|
58
|
+
// disagree about the same exhaustion again.
|
|
53
59
|
// init run · fit · run-args Opens a run, or refuses it (exit 3). `run-args`
|
|
54
60
|
// writes GATE L0.9b's launch record and echoes it.
|
|
55
61
|
// report export Projects the run's records as fact tables.
|
|
@@ -86,7 +92,7 @@ export const ROUTES = {
|
|
|
86
92
|
resume: "./probe/resume.mjs", t0: "./probe/t0.mjs", stats: "./probe/stats.mjs",
|
|
87
93
|
digest: "./probe/digest.mjs", concurrency: "./probe/concurrency.mjs",
|
|
88
94
|
leg: "./probe/leg.mjs", eval: "./probe/eval.mjs", owner: "./probe/owner.mjs",
|
|
89
|
-
requirements: "./probe/requirements.mjs",
|
|
95
|
+
requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs",
|
|
90
96
|
},
|
|
91
97
|
init: { run: "./init/run.mjs", fit: "./init/fit.mjs", "run-args": "./init/run-args.mjs" },
|
|
92
98
|
report: { export: "./report/export.mjs", _default: "export" },
|
|
@@ -0,0 +1,135 @@
|
|
|
1
|
+
// probe attempts — "how many of this scope's attempts are ATTESTED work, not writable artifacts?"
|
|
2
|
+
//
|
|
3
|
+
// CONTRACT. A bounded, read-only query over one scope's attested channels for one round: dispatch
|
|
4
|
+
// receipts (`receipts/dispatch.jsonl`), leg-completion rows (`legs.jsonl`) and WorkResults
|
|
5
|
+
// (`results/`). Prints `{scope_id, round, attempt_budget, spent, in_flight, unattested, green,
|
|
6
|
+
// tripped, attempts}` on stdout; exits 0 when the breaker holds, 1 when it has tripped, 2 on a bad
|
|
7
|
+
// argv. Writes nothing.
|
|
8
|
+
//
|
|
9
|
+
// WHY IT EXISTS. Measured on a real run: attempt 2 of a scope was compiled and T0-verified
|
|
10
|
+
// several minutes BEFORE attempt 1 ingested — while attempt 1 was still in flight. No worker was
|
|
11
|
+
// ever dispatched for attempt 2:
|
|
12
|
+
// `receipts/dispatch.jsonl`, `legs.jsonl` and `results/` carried no row for it. A derivation keyed
|
|
13
|
+
// off the order set (`orders/`) or the T0 verdict set (`t0/verdicts/`) alone counts a compiled
|
|
14
|
+
// order or a green trial as a spent attempt regardless of whether a worker ever ran — both are
|
|
15
|
+
// WRITABLE by the very leg whose exhaustion is being judged. This module counts an attempt as
|
|
16
|
+
// SPENT only when a dispatch receipt attests it started AND either a leg-completion row or a
|
|
17
|
+
// WorkResult on disk attests it closed. Neither channel alone is enough: a receipt with no result
|
|
18
|
+
// is a leg still in flight (open, not spent — the breaker must not trip on unanswered work), and a
|
|
19
|
+
// result or verdict with no receipt is unattested — no worker ran, which is the shape that
|
|
20
|
+
// produced this module.
|
|
21
|
+
//
|
|
22
|
+
// WHY THE SAME FUNCTION SERVES THE BREAKER AND THE CENSUS. Before this module, the round loop's
|
|
23
|
+
// inner breaker and `scope-hammer`'s GATE H0 census read different evidence for the same question
|
|
24
|
+
// — the loop trusted the worker's own self-reported `attempts_used`/`breaker` (schema-shaped, not
|
|
25
|
+
// re-verified), the census read `t0/verdicts/*.json` directly — and disagreed on the measured run.
|
|
26
|
+
// One shared, attested derivation is what makes "the census and the breaker agree" true by
|
|
27
|
+
// construction rather than by coincidence: `skills/scope-hammer/SKILL.md` cites this same probe
|
|
28
|
+
// the way it already cites `probe owner` for ownership claims.
|
|
29
|
+
|
|
30
|
+
import { existsSync, readdirSync, readFileSync } from "node:fs";
|
|
31
|
+
import { join, resolve } from "node:path";
|
|
32
|
+
import { runArgs } from "../lib/argv.mjs";
|
|
33
|
+
import { dispatchReceipts, legLedger, resultsDir } from "../lib/paths.mjs";
|
|
34
|
+
import { readLegs } from "./leg.mjs";
|
|
35
|
+
import { greenVerdict } from "./t0.mjs";
|
|
36
|
+
|
|
37
|
+
/**
|
|
38
|
+
* Every dispatch-receipt row on disk, tolerant of a torn last line (mirrors {@link readLegs}).
|
|
39
|
+
* @param {string} path - `receipts/dispatch.jsonl`.
|
|
40
|
+
* @returns {object[]} Parsed rows; an unparsable line is skipped rather than fatal.
|
|
41
|
+
*/
|
|
42
|
+
export function readReceipts(path) {
|
|
43
|
+
if (!existsSync(path)) return [];
|
|
44
|
+
return readFileSync(path, "utf8").split("\n").filter(Boolean)
|
|
45
|
+
.map((l) => { try { return JSON.parse(l); } catch { return null; } })
|
|
46
|
+
.filter(Boolean);
|
|
47
|
+
}
|
|
48
|
+
|
|
49
|
+
/**
|
|
50
|
+
* One attempt's evidence across the three attested channels, and the state it derives to.
|
|
51
|
+
*
|
|
52
|
+
* `unattested` (no receipt at all) counts as though the attempt never happened — the order and any
|
|
53
|
+
* T0 verdict may still be sitting on disk, written by something other than a dispatched worker, and
|
|
54
|
+
* neither is asked here. `in-flight` (a receipt, but no leg row and no result) is a real dispatch
|
|
55
|
+
* whose leg has not yet closed: open, not spent. `spent` is closed either way a leg closes — via
|
|
56
|
+
* `reduce ingest`'s own leg-completion row, or a WorkResult already on disk pending ingest.
|
|
57
|
+
*
|
|
58
|
+
* @param {string} cwd - Project root.
|
|
59
|
+
* @param {string} slug - Feature slug.
|
|
60
|
+
* @param {string} scopeId - Scope contract id (the order id's own address, e.g. `shell-r1-a2`).
|
|
61
|
+
* @param {number} round - Build round.
|
|
62
|
+
* @param {number} attempt - Attempt number within the round.
|
|
63
|
+
* @param {object[]} receipts - Pre-read `receipts/dispatch.jsonl` rows (avoids re-reading per attempt).
|
|
64
|
+
* @param {object[]} legs - Pre-read `legs.jsonl` rows.
|
|
65
|
+
* @returns {{orderId:string, hasReceipt:boolean, hasResult:boolean, hasLeg:boolean, state:("unattested"|"in-flight"|"spent")}}
|
|
66
|
+
*/
|
|
67
|
+
export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, legs) {
|
|
68
|
+
const orderId = `${slug}/${scopeId}-r${round}-a${attempt}`;
|
|
69
|
+
const hasReceipt = receipts.some((r) => r?.order_id === orderId);
|
|
70
|
+
const hasLeg = legs.some((r) => r?.order_id === orderId);
|
|
71
|
+
const hasResult = existsSync(join(resultsDir(cwd, slug), `${scopeId}-r${round}-a${attempt}.json`));
|
|
72
|
+
const state = !hasReceipt ? "unattested" : (hasResult || hasLeg) ? "spent" : "in-flight";
|
|
73
|
+
return { orderId, hasReceipt, hasResult, hasLeg, state };
|
|
74
|
+
}
|
|
75
|
+
|
|
76
|
+
/**
|
|
77
|
+
* A scope's attempt census for one round, derived ONLY from attested channels — the function both
|
|
78
|
+
* the round loop's inner breaker and scope-hammer's GATE H0 census call, so they cannot drift apart
|
|
79
|
+
* again the way a measured run found them.
|
|
80
|
+
*
|
|
81
|
+
* @param {string} cwd - Project root.
|
|
82
|
+
* @param {string} slug - Feature slug.
|
|
83
|
+
* @param {string} scopeId - Scope contract id.
|
|
84
|
+
* @param {number} round - Build round.
|
|
85
|
+
* @param {number} attemptBudget - The scope's configured `attempt_budget`.
|
|
86
|
+
* @returns {{scope_id:string, round:number, attempt_budget:number, spent:number, in_flight:number,
|
|
87
|
+
* unattested:number, green:boolean, tripped:boolean, attempts:object[]}} `tripped` is true only
|
|
88
|
+
* when every attempt within budget is genuinely SPENT and none produced a green T0 — an attempt
|
|
89
|
+
* still in flight holds the breaker open regardless of how many slots are nominally used.
|
|
90
|
+
*/
|
|
91
|
+
export function scopeAttempts(cwd, slug, scopeId, round, attemptBudget) {
|
|
92
|
+
const receipts = readReceipts(dispatchReceipts(cwd, slug));
|
|
93
|
+
const legs = readLegs(legLedger(cwd, slug));
|
|
94
|
+
const attempts = [];
|
|
95
|
+
let spent = 0, inFlight = 0, unattested = 0;
|
|
96
|
+
for (let a = 1; a <= attemptBudget; a++) {
|
|
97
|
+
const ev = attemptEvidence(cwd, slug, scopeId, round, a, receipts, legs);
|
|
98
|
+
if (ev.state === "spent") spent++;
|
|
99
|
+
else if (ev.state === "in-flight") inFlight++;
|
|
100
|
+
else unattested++;
|
|
101
|
+
attempts.push({ attempt: a, ...ev });
|
|
102
|
+
}
|
|
103
|
+
const { green } = greenVerdict(cwd, slug, scopeId, round);
|
|
104
|
+
return {
|
|
105
|
+
scope_id: scopeId, round, attempt_budget: attemptBudget,
|
|
106
|
+
spent, in_flight: inFlight, unattested, green,
|
|
107
|
+
tripped: !green && spent >= attemptBudget,
|
|
108
|
+
attempts,
|
|
109
|
+
};
|
|
110
|
+
}
|
|
111
|
+
|
|
112
|
+
export const ARGV_SPEC = {
|
|
113
|
+
usage: "harness.mjs probe attempts --slug <slug> --scope <scope-id> --round N --attempt-budget N [--cwd <dir>]",
|
|
114
|
+
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
115
|
+
slug: { type: "str", required: true },
|
|
116
|
+
scope: { type: "str", required: true },
|
|
117
|
+
round: { type: "int", min: 1, required: true },
|
|
118
|
+
"attempt-budget": { type: "int", min: 1, required: true },
|
|
119
|
+
cwd: { type: "path" },
|
|
120
|
+
};
|
|
121
|
+
|
|
122
|
+
/**
|
|
123
|
+
* Report one scope's attested attempt census for one round.
|
|
124
|
+
*
|
|
125
|
+
* @param {string[]} rawArgv - The subcommand's own arguments (harness.mjs strips the verb words).
|
|
126
|
+
* @returns {void} Exits 0 when the breaker holds, 1 when it has tripped — the shape a caller (or a
|
|
127
|
+
* census citing this row) can branch on without parsing prose.
|
|
128
|
+
*/
|
|
129
|
+
export function cli(rawArgv) {
|
|
130
|
+
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
131
|
+
const cwd = resolve(args.cwd || process.cwd());
|
|
132
|
+
const r = scopeAttempts(cwd, args.slug, args.scope, args.round, args.attemptBudget);
|
|
133
|
+
console.log(JSON.stringify(r));
|
|
134
|
+
process.exit(r.tripped ? 1 : 0);
|
|
135
|
+
}
|
package/kernel/probe/resume.mjs
CHANGED
|
@@ -64,8 +64,10 @@ import { globToRegExp } from "../verify/spec.mjs";
|
|
|
64
64
|
import {
|
|
65
65
|
intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir,
|
|
66
66
|
orientDir, activeOrder, usecasesDir, breadboard, receipt, readReceipt, requirements,
|
|
67
|
+
exportRunDir,
|
|
67
68
|
} from "../lib/paths.mjs";
|
|
68
69
|
import { evalVerdict } from "./eval.mjs";
|
|
70
|
+
import { collectRun, writeRun } from "../report/export.mjs";
|
|
69
71
|
|
|
70
72
|
/** The run-state values `references/protocol.md` (Part 4 — State) defines. A typo'd status is a rejection,
|
|
71
73
|
* not a write — the whole point of this file is that a write nobody validates is a write nobody
|
|
@@ -81,6 +83,66 @@ export const RUN_STATUSES = ["orienting", "mapping", "building", "evaluating", "
|
|
|
81
83
|
*/
|
|
82
84
|
export const TERMINAL_STATUSES = ["shipped", "aborted", "escalated"];
|
|
83
85
|
|
|
86
|
+
/**
|
|
87
|
+
* The RunReturn union (`kernel/schemas/domain.schema.json` `$defs/RunReturn.properties.status.enum`)
|
|
88
|
+
* mapped to what closing the run means for each arm — the derivation `closeIfTerminal`
|
|
89
|
+
* (`skills/tech-lead/workflows/shapeup-run.js`) now reads instead of a hand-typed
|
|
90
|
+
* `status !== "aborted" && status !== "shipped"` pair that referenced {@link TERMINAL_STATUSES}
|
|
91
|
+
* zero times and so could not see when a new arm went unhandled.
|
|
92
|
+
*
|
|
93
|
+
* A lookup answers one of three ways, and the distinction is load-bearing for what this map must
|
|
94
|
+
* catch: an arm ABSENT from this object (never listed as a key) returns `undefined` — an arm the
|
|
95
|
+
* schema carries that nobody has mapped, which a caller must treat as a defect, never as "fine to
|
|
96
|
+
* skip". An arm mapped to a terminal status closes the run as that status. An arm mapped to `null`
|
|
97
|
+
* is EXPLICITLY non-terminal — `paused` resumes on relaunch and `ok` is one inner round finishing,
|
|
98
|
+
* not the run — so it is a key with a falsy value, not an omission a reader could mistake for "not
|
|
99
|
+
* decided yet".
|
|
100
|
+
*
|
|
101
|
+
* `gate_h` → `escalated`: the breaker that tripped (`outer`/`inner`/`deadline`) travels in
|
|
102
|
+
* `close_cause`, never as a new member of {@link TERMINAL_STATUSES} or {@link RUN_STATUSES} — a
|
|
103
|
+
* circuit breaker tripping is not a new way a run ends, it is the reason an existing one
|
|
104
|
+
* (`escalated`) fires this time.
|
|
105
|
+
*
|
|
106
|
+
* @type {Object<string, (string|null)>}
|
|
107
|
+
*/
|
|
108
|
+
export const RUN_RETURN_CLOSE = {
|
|
109
|
+
shipped: "shipped",
|
|
110
|
+
aborted: "aborted",
|
|
111
|
+
gate_h: "escalated",
|
|
112
|
+
paused: null,
|
|
113
|
+
ok: null,
|
|
114
|
+
};
|
|
115
|
+
|
|
116
|
+
/**
|
|
117
|
+
* Resolve one RunReturn arm to a close outcome and, when the arm is terminal, perform the close —
|
|
118
|
+
* the single call `closeIfTerminal` makes instead of deciding locally which arms are terminal.
|
|
119
|
+
*
|
|
120
|
+
* @param {string} cwd - Project root.
|
|
121
|
+
* @param {string} slug - Feature slug.
|
|
122
|
+
* @param {string} arm - A `RunReturn.status` value.
|
|
123
|
+
* @param {(string|null)} [cause] - Why the run ended there; only used when `arm` is terminal.
|
|
124
|
+
* @param {boolean} [withExport] - Forwarded to {@link closeRun} — defaults true. The one caller
|
|
125
|
+
* that ever passes `false` is a test fixture proving the export assertion is real — a check that
|
|
126
|
+
* cannot fail did not happen; production call sites never set this.
|
|
127
|
+
* @returns {({ok:true, arm:string, terminal:false, reason:string} |
|
|
128
|
+
* {ok:false, arm:string, reason:string} |
|
|
129
|
+
* ({ok:boolean, arm:string, terminal:true} & ReturnType<typeof closeRun>))} `terminal:false` when
|
|
130
|
+
* `arm` maps to `null` (nothing closed, not an error). `ok:false` with no `terminal` field when
|
|
131
|
+
* `arm` is not a key of {@link RUN_RETURN_CLOSE} at all — a schema arm this map has not been
|
|
132
|
+
* taught, which must never be silently treated as non-terminal. Otherwise the {@link closeRun}
|
|
133
|
+
* outcome, tagged with the arm that produced it.
|
|
134
|
+
*/
|
|
135
|
+
export function closeArm(cwd, slug, arm, cause = null, withExport = true) {
|
|
136
|
+
if (!Object.hasOwn(RUN_RETURN_CLOSE, arm)) {
|
|
137
|
+
return { ok: false, arm, reason: `closeArm: "${arm}" is not a RunReturn arm this kernel maps — known arms: ${Object.keys(RUN_RETURN_CLOSE).join(", ")}` };
|
|
138
|
+
}
|
|
139
|
+
const status = RUN_RETURN_CLOSE[arm];
|
|
140
|
+
if (!status) {
|
|
141
|
+
return { ok: true, arm, terminal: false, reason: `"${arm}" is explicitly non-terminal — no close` };
|
|
142
|
+
}
|
|
143
|
+
return { ...closeRun(cwd, slug, { status, cause, withExport }), arm, terminal: true };
|
|
144
|
+
}
|
|
145
|
+
|
|
84
146
|
/** ORIENT's four artifacts (skills/orient/SKILL.md §Outputs): three by exact name, plus a spike
|
|
85
147
|
* whose filename carries the area it spiked (`spike-<area>.md`, or `spike-not-needed.md` when
|
|
86
148
|
* the risk scan came back rank 0 — both count, because both are ORIENT having finished). */
|
|
@@ -498,6 +560,40 @@ function writeCloseLines(body, { status, closedAt, cause }) {
|
|
|
498
560
|
return out;
|
|
499
561
|
}
|
|
500
562
|
|
|
563
|
+
/**
|
|
564
|
+
* Export a just-closed run's own records into fact tables, best-effort
|
|
565
|
+
* (`report export`, `kernel/report/export.mjs`). {@link closeRun} calls this for every terminal
|
|
566
|
+
* status except `"shipped"`: the Ship phase (`skills/tech-lead/workflows/shapeup-run.js`) already
|
|
567
|
+
* calls `report export` itself, several lines before this file's own close-out runs, and that call
|
|
568
|
+
* site is left untouched on purpose — the shipped path's export stays byte-comparable to what it
|
|
569
|
+
* wrote before Stage 2. `aborted`, `escalated` and `gate_h` (which closes as `escalated`, see
|
|
570
|
+
* {@link RUN_RETURN_CLOSE}) never reached that call site at all, because it sits inside a phase
|
|
571
|
+
* those endings never enter — so a run ending any of them left its whole trace (orders, results,
|
|
572
|
+
* T0 verdicts, hook decisions, `graph.jsonl`) in the gitignored LOCAL tier with nothing durable
|
|
573
|
+
* surviving the next `init run`'s wipe. This is the read that was missing for those endings, run at
|
|
574
|
+
* the one point every one of them passes through: the close itself.
|
|
575
|
+
*
|
|
576
|
+
* FAIL-OPEN, the hook discipline (CLAUDE.md) extended to a write that is not a hook: an export that
|
|
577
|
+
* cannot write — a blocked or missing exports directory, a full disk — must never turn a close that
|
|
578
|
+
* DID take into one that looks like it did not. The three facts {@link closeRun} just wrote
|
|
579
|
+
* (`closed_status`/`close_cause`/`closed_at`) are never touched by this function; a failure here is
|
|
580
|
+
* handed back to the caller as a warning string, never thrown.
|
|
581
|
+
*
|
|
582
|
+
* @param {string} cwd - Project root.
|
|
583
|
+
* @param {string} slug - Feature slug.
|
|
584
|
+
* @returns {(string|null)} A one-line warning when the export did not complete; `null` on success.
|
|
585
|
+
*/
|
|
586
|
+
function exportOnClose(cwd, slug) {
|
|
587
|
+
try {
|
|
588
|
+
const collected = collectRun(cwd, slug);
|
|
589
|
+
if (!collected) return `export on close: no readable receipt for "${slug}" — nothing to export`;
|
|
590
|
+
writeRun(collected, exportRunDir(cwd, collected.run_id ?? slug));
|
|
591
|
+
return null;
|
|
592
|
+
} catch (e) {
|
|
593
|
+
return `export on close: ${e.message}`;
|
|
594
|
+
}
|
|
595
|
+
}
|
|
596
|
+
|
|
501
597
|
/**
|
|
502
598
|
* Close the run: a terminal status, its cause, and a close timestamp, written together in ONE
|
|
503
599
|
* pass.
|
|
@@ -548,20 +644,26 @@ function writeCloseLines(body, { status, closedAt, cause }) {
|
|
|
548
644
|
*
|
|
549
645
|
* @param {string} cwd - Project root.
|
|
550
646
|
* @param {string} slug - Feature slug.
|
|
551
|
-
* @param {{status:string, cause:(string|null)}} o - The terminal status (one
|
|
552
|
-
* {@link TERMINAL_STATUSES}) and why the run ended there. `cause` travels through `uncoerce`
|
|
553
|
-
* one dialect `harness-run.md`'s frontmatter is read and written in), so free prose — quotes
|
|
554
|
-
* colons included — round-trips as one frontmatter line; an embedded newline is collapsed to
|
|
555
|
-
* space first, because this dialect is line-based and could not carry one either way.
|
|
647
|
+
* @param {{status:string, cause:(string|null), withExport?:boolean}} o - The terminal status (one
|
|
648
|
+
* of {@link TERMINAL_STATUSES}) and why the run ended there. `cause` travels through `uncoerce`
|
|
649
|
+
* (the one dialect `harness-run.md`'s frontmatter is read and written in), so free prose — quotes
|
|
650
|
+
* and colons included — round-trips as one frontmatter line; an embedded newline is collapsed to
|
|
651
|
+
* a space first, because this dialect is line-based and could not carry one either way.
|
|
652
|
+
* `withExport` (default true) gates {@link exportOnClose} — every status except `"shipped"` runs
|
|
653
|
+
* it on a successful close; `false` exists only for a test fixture proving the export assertion
|
|
654
|
+
* is real, never for a production call site.
|
|
556
655
|
* @returns {{ok:boolean, path:string, status:string, closed_at?:string, cause?:(string|null),
|
|
557
656
|
* reason?:string, closed_status?:string, superseded?:boolean, decision?:string,
|
|
558
|
-
* prior_cause?:(string|null), prior_closed_at?:string}} Outcome. A refused
|
|
559
|
-
* closed with a DIFFERENT terminal status) carries
|
|
560
|
-
* is actually on disk. A successful supersede
|
|
561
|
-
* `superseded:true`, `decision:"superseded"` (the
|
|
562
|
-
* `shapeup-run.js` relays verbatim — see its own
|
|
657
|
+
* prior_cause?:(string|null), prior_closed_at?:string, export_warning?:string}} Outcome. A refused
|
|
658
|
+
* overwrite (already closed with a DIFFERENT terminal status) carries
|
|
659
|
+
* `closed_status`/`closed_at`/`cause` naming what is actually on disk. A successful supersede
|
|
660
|
+
* (same status, different cause) carries `superseded:true`, `decision:"superseded"` (the
|
|
661
|
+
* one-token signal the courier boundary in `shapeup-run.js` relays verbatim — see its own
|
|
662
|
+
* `cmd()` banner) and the prior close it folded in. Any successful, non-`"shipped"` close carries
|
|
663
|
+
* `export_warning` when {@link exportOnClose} could not write the run's fact tables — the close
|
|
664
|
+
* itself still stands; this is advisory only.
|
|
563
665
|
*/
|
|
564
|
-
export function closeRun(cwd, slug, { status, cause = null } = {}) {
|
|
666
|
+
export function closeRun(cwd, slug, { status, cause = null, withExport = true } = {}) {
|
|
565
667
|
const p = harnessRun(cwd, slug);
|
|
566
668
|
if (!TERMINAL_STATUSES.includes(status)) {
|
|
567
669
|
return { ok: false, path: p, status, reason: `closeRun: "${status}" is not terminal — expected one of ${TERMINAL_STATUSES.join(" | ")}` };
|
|
@@ -574,6 +676,15 @@ export function closeRun(cwd, slug, { status, cause = null } = {}) {
|
|
|
574
676
|
return { ok: false, path: p, status, reason: `harness-run.md carries no "status:"/"closed_at:" line to replace — the ledger's frontmatter is malformed (references/protocol.md)` };
|
|
575
677
|
}
|
|
576
678
|
|
|
679
|
+
// The Ship phase (`shapeup-run.js`) already exports a shipped run itself, before this call ever
|
|
680
|
+
// runs — Stage 2 adds the endings that wrote nothing, and leaves that path untouched.
|
|
681
|
+
const shouldExport = withExport && status !== "shipped";
|
|
682
|
+
const withExportWarning = (result) => {
|
|
683
|
+
if (!shouldExport) return result;
|
|
684
|
+
const warning = exportOnClose(cwd, slug);
|
|
685
|
+
return warning ? { ...result, export_warning: warning } : result;
|
|
686
|
+
};
|
|
687
|
+
|
|
577
688
|
// Truncated, not elided: a cause this long has already done its job in the run's own log — the
|
|
578
689
|
// ledger line is a pointer back to it, not the full transcript. Newlines are collapsed to spaces
|
|
579
690
|
// FIRST — this dialect is line-based, so a raw embedded newline would split one field into a value
|
|
@@ -588,7 +699,7 @@ export function closeRun(cwd, slug, { status, cause = null } = {}) {
|
|
|
588
699
|
if (priorClosedStatus && priorClosedAt) {
|
|
589
700
|
if (priorClosedStatus === status && normCause === priorCause) {
|
|
590
701
|
// The identical fact, restated — a retried or duplicated call costs nothing.
|
|
591
|
-
return { ok: true, path: p, status, closed_at: priorClosedAt, cause: priorCause, decision: "idempotent", reason: `already closed as "${status}" at ${priorClosedAt} — idempotent no-op` };
|
|
702
|
+
return withExportWarning({ ok: true, path: p, status, closed_at: priorClosedAt, cause: priorCause, decision: "idempotent", reason: `already closed as "${status}" at ${priorClosedAt} — idempotent no-op` });
|
|
592
703
|
}
|
|
593
704
|
if (priorClosedStatus !== status) {
|
|
594
705
|
// A DIFFERENT terminal status over an already-closed run — refused outright, the original
|
|
@@ -613,10 +724,10 @@ export function closeRun(cwd, slug, { status, cause = null } = {}) {
|
|
|
613
724
|
if (afterSup.status !== status || !afterSup.closed_at || afterSup.closed_at === "~") {
|
|
614
725
|
return { ok: false, path: p, status, reason: `wrote the superseding close but the ledger reads back status="${afterSup.status}" closed_at="${afterSup.closed_at}" — the write did not take` };
|
|
615
726
|
}
|
|
616
|
-
return {
|
|
727
|
+
return withExportWarning({
|
|
617
728
|
ok: true, path: p, status, closed_at: afterSup.closed_at, cause: afterSup.close_cause ?? null,
|
|
618
729
|
superseded: true, decision: "superseded", prior_cause: priorCause, prior_closed_at: priorClosedAt,
|
|
619
|
-
};
|
|
730
|
+
});
|
|
620
731
|
}
|
|
621
732
|
|
|
622
733
|
const closedAt = new Date().toISOString();
|
|
@@ -628,7 +739,7 @@ export function closeRun(cwd, slug, { status, cause = null } = {}) {
|
|
|
628
739
|
if (after.status !== status || !after.closed_at || after.closed_at === "~") {
|
|
629
740
|
return { ok: false, path: p, status, reason: `wrote the close but the ledger reads back status="${after.status}" closed_at="${after.closed_at}" — the write did not take` };
|
|
630
741
|
}
|
|
631
|
-
return { ok: true, path: p, status, closed_at: after.closed_at, cause: after.close_cause ?? null, decision: "closed" };
|
|
742
|
+
return withExportWarning({ ok: true, path: p, status, closed_at: after.closed_at, cause: after.close_cause ?? null, decision: "closed" });
|
|
632
743
|
}
|
|
633
744
|
|
|
634
745
|
/**
|
|
@@ -658,7 +769,8 @@ export function writeActiveOrder(cwd, slug, orderPath) {
|
|
|
658
769
|
/** The typed argv contract (see `./lib/argv.mjs`). */
|
|
659
770
|
export const ARGV_SPEC = {
|
|
660
771
|
usage: "harness.mjs probe resume --slug <slug> [--cwd <dir>] " +
|
|
661
|
-
"[--require <phase> | --set-status <status> | --set-active-order <path> |
|
|
772
|
+
"[--require <phase> | --set-status <status> | --set-active-order <path> | " +
|
|
773
|
+
"--close <status> | --close-arm <RunReturn.status> [--cause <text>]]",
|
|
662
774
|
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
663
775
|
slug: { type: "str", required: true },
|
|
664
776
|
cwd: { type: "path" },
|
|
@@ -667,6 +779,9 @@ export const ARGV_SPEC = {
|
|
|
667
779
|
"set-active-order": { type: "str" },
|
|
668
780
|
// The one call site that stamps a terminal status, its cause and closed_at together.
|
|
669
781
|
close: { type: "enum", values: TERMINAL_STATUSES },
|
|
782
|
+
// Not `--close`: the caller (shapeup-run.js's closeIfTerminal) hands over a RunReturn arm, never
|
|
783
|
+
// a status it decided was terminal itself — RUN_RETURN_CLOSE/closeArm above make that call.
|
|
784
|
+
"close-arm": { type: "str" },
|
|
670
785
|
cause: { type: "str" },
|
|
671
786
|
};
|
|
672
787
|
|
|
@@ -681,13 +796,13 @@ export function cli(rawArgv) {
|
|
|
681
796
|
const args = runArgs(ARGV_SPEC, rawArgv);
|
|
682
797
|
const cwd = args.cwd || process.cwd();
|
|
683
798
|
|
|
684
|
-
const ops = [args.require && "--require", args.setStatus && "--set-status", args.setActiveOrder && "--set-active-order", args.close && "--close"].filter(Boolean);
|
|
799
|
+
const ops = [args.require && "--require", args.setStatus && "--set-status", args.setActiveOrder && "--set-active-order", args.close && "--close", args.closeArm && "--close-arm"].filter(Boolean);
|
|
685
800
|
if (ops.length > 1) {
|
|
686
801
|
process.stderr.write(JSON.stringify({ error: "conflicting_flags", flags: ops, expected: "one operation per invocation" }) + "\n");
|
|
687
802
|
process.exit(2);
|
|
688
803
|
}
|
|
689
|
-
if (args.cause !== undefined && !args.close) {
|
|
690
|
-
process.stderr.write(JSON.stringify({ error: "conflicting_flags", flags: ["--cause"], expected: "--cause is only meaningful with --close" }) + "\n");
|
|
804
|
+
if (args.cause !== undefined && !args.close && !args.closeArm) {
|
|
805
|
+
process.stderr.write(JSON.stringify({ error: "conflicting_flags", flags: ["--cause"], expected: "--cause is only meaningful with --close or --close-arm" }) + "\n");
|
|
691
806
|
process.exit(2);
|
|
692
807
|
}
|
|
693
808
|
|
|
@@ -724,6 +839,17 @@ export function cli(rawArgv) {
|
|
|
724
839
|
process.exit(r.ok ? 0 : 3);
|
|
725
840
|
}
|
|
726
841
|
|
|
842
|
+
// The arm-derived close: the caller hands a RunReturn arm and this kernel decides — via
|
|
843
|
+
// RUN_RETURN_CLOSE/closeArm above — whether it is terminal and, if so, what status it closes as.
|
|
844
|
+
// Exit 0 for both a real close AND a correctly-declined non-terminal arm (`terminal:false`) —
|
|
845
|
+
// neither is an error the caller (shapeup-run.js's closeIfTerminal) should treat as failed; only
|
|
846
|
+
// an arm this map does not recognize at all, or a close `closeRun` itself refuses, exits non-zero.
|
|
847
|
+
if (args.closeArm) {
|
|
848
|
+
const r = closeArm(cwd, args.slug, args.closeArm, args.cause ?? null);
|
|
849
|
+
console.log(JSON.stringify(r));
|
|
850
|
+
process.exit(r.ok ? 0 : 3);
|
|
851
|
+
}
|
|
852
|
+
|
|
727
853
|
console.log(JSON.stringify(deriveResumeState(cwd, args.slug)));
|
|
728
854
|
process.exit(0);
|
|
729
855
|
}
|
package/package.json
CHANGED
|
@@ -121,7 +121,7 @@ covered AC is a requirement the run can be measured against.
|
|
|
121
121
|
|---|---|---|
|
|
122
122
|
| `reconcile` | Verify `ledger.feature == payload.feature` (mismatch → STOP). Map each `[+]` Keep item → its owning UC; new task continues numbering (never renumber); `~`/Cut → synthesis "Hammered Out" row, no file. A Keep item asserting a new invariant → APPEND `[INV-NN]` + TS-INV row to that UC (append-only sections in your substrate). A new actor/action with no UC → `status: "escalated"` + a `deviations[]` spec-ambiguity entry: spawning a UC mid-cycle is silent re-shaping, the PO decides. Finish with board-derive (appetite overflow → report) + spec-lint | re-run phases 1–5; edit UC Steps; resolve the appetite HAMMER yourself |
|
|
123
123
|
| `retrofit-surface` | Append `## Test Surface` (derived rows only, after Error Cases) to each UC of a pre-surface spec; an all-sources-empty UC gets the explicit empty-sources line | touch anything else — append-only substrate |
|
|
124
|
-
| `coverage` | Extract **atomic** customer requirement clauses from `payload.requirements` (default: the pitch) and write the SHARED `shapeup/<slug>/requirements.md` registry: one `\| REQ-id \| clause (verbatim) \| source \| status \| note \|` row per clause. Split compound sentences into one testable clause each — a clause lost *inside* a bigger sentence is a requirement nothing can be traced to. **Assign REQ-ids ONCE and freeze them** (they behave like scope_id, never TASK-NNN — every `covers:` link rots otherwise): re-running, append new clauses with fresh ids, mark a removed clause `CUT (PO-approved)`, never renumber or delete. Status starts `covered` (a live requirement); only the PO sets `CUT`. The REQ source itself is frozen — the registry is a separate derived file. **Numbering.** A source clause already carrying an `R<n>` keeps its number — `R12` → `REQ-12` — and its `source` cell records where it came from verbatim (`shaping.md R12`), because that cell is the only thing that survives a re-run. A clause with no R-id takes the next free number ABOVE the highest `R<n>` in the source, so it can never collide with one added later. Splitting a compound clause keeps `REQ-12` for the first atomic part and records `shaping.md R12 (split 2/3)` for the rest — a requirement graded in parts is why splitting matters at all. On a re-run, match an existing id by its frozen `source` cell and clause text, **never** by re-deriving the number from the source's current order | edit the REQ source; renumber existing REQ-ids; delete a dropped clause instead of marking it CUT; invent a requirement not in the source; re-point an existing REQ-id because the source's R-numbers shifted |
|
|
124
|
+
| `coverage` | Extract **atomic** customer requirement clauses from `payload.requirements` (default: the pitch) and write the SHARED `shapeup/<slug>/requirements.md` registry: one `\| REQ-id \| clause (verbatim) \| source \| status \| note \|` row per clause. **The `source` cell, and any other prose in this file, cites the committed pitch or shaping doc — never the gitignored run-tier path** (`.shapeup/<slug>/intake.md`, or any other `.shapeup/` path): spec-lint's `TIER-DIRECTION` rule reds a committed file naming that path in ANY form, a bare path in a sentence exactly as much as a `[[tasks/...]]` wikilink, because it dangles on every other clone. When the intake has no committed original to name, describe the run tier without a path (`"the pitch staged for this run"`) rather than citing where it actually lives. Split compound sentences into one testable clause each — a clause lost *inside* a bigger sentence is a requirement nothing can be traced to. **Assign REQ-ids ONCE and freeze them** (they behave like scope_id, never TASK-NNN — every `covers:` link rots otherwise): re-running, append new clauses with fresh ids, mark a removed clause `CUT (PO-approved)`, never renumber or delete. Status starts `covered` (a live requirement); only the PO sets `CUT`. The REQ source itself is frozen — the registry is a separate derived file. **Numbering.** A source clause already carrying an `R<n>` keeps its number — `R12` → `REQ-12` — and its `source` cell records where it came from verbatim (`shaping.md R12`), because that cell is the only thing that survives a re-run. A clause with no R-id takes the next free number ABOVE the highest `R<n>` in the source, so it can never collide with one added later. Splitting a compound clause keeps `REQ-12` for the first atomic part and records `shaping.md R12 (split 2/3)` for the rest — a requirement graded in parts is why splitting matters at all. On a re-run, match an existing id by its frozen `source` cell and clause text, **never** by re-deriving the number from the source's current order | edit the REQ source; renumber existing REQ-ids; delete a dropped clause instead of marking it CUT; invent a requirement not in the source; re-point an existing REQ-id because the source's R-numbers shifted |
|
|
125
125
|
---
|
|
126
126
|
|
|
127
127
|
## Anti-rationalization table
|
|
@@ -275,6 +275,12 @@ Always use wikilinks (double brackets), never relative paths like `../domain-mod
|
|
|
275
275
|
`.shapeup/` is gitignored, so a committed task link dangles on every fresh clone.
|
|
276
276
|
spec-lint flags it as a red `TIER-DIRECTION` finding. Coverage views (synthesis
|
|
277
277
|
traceability) record derived counts/status, not task ids.
|
|
278
|
+
- **The rule is not only about wikilinks.** `TIER-DIRECTION` reds *any* line in a SHARED
|
|
279
|
+
doc that names a `.shapeup/` path — a bare path cited in a sentence or a table cell,
|
|
280
|
+
not only a `[[tasks/...]]` link. A provenance sentence that names its real source
|
|
281
|
+
(`"extracted from .shapeup/<slug>/intake.md"`) reds for the same reason a task
|
|
282
|
+
wikilink does: the path dangles on every other clone. Cite the committed pitch or
|
|
283
|
+
shaping doc instead, or describe the run tier without a path.
|
|
278
284
|
- `[[tasks/...]]` wikilinks are valid only inside LOCAL documents (task files, the board,
|
|
279
285
|
EVAL reports), where they resolve against the LOCAL root (`.shapeup/<slug>/`);
|
|
280
286
|
every wikilink in a SHARED doc stays `spec_folder`-relative.
|
|
@@ -67,8 +67,16 @@ H0.0 Ownership is DERIVED, never stated. Before the census says "no scope owns
|
|
|
67
67
|
H0.1 Unresolved scopes (breaker cases only):
|
|
68
68
|
- uphill/downhill scopes when round_budget hit 0 → CARRY candidates (their own hill
|
|
69
69
|
phase + open unknowns, from hill/<scope-id>.yml)
|
|
70
|
-
- scopes with hammer_proposals (attempt_budget exhausted) → CARRY candidates
|
|
71
|
-
|
|
70
|
+
- scopes with hammer_proposals (attempt_budget exhausted) → CARRY candidates. Exhaustion
|
|
71
|
+
is DERIVED, never read off `t0/verdicts/*.json` directly — a compiled order or a T0
|
|
72
|
+
verdict is writable by the very scope being judged and proves nothing on its own. Run
|
|
73
|
+
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe attempts --slug <slug> \
|
|
74
|
+
--scope <scope-id> --round <n> --attempt-budget <n>
|
|
75
|
+
and cite its `spent`/`tripped` fields (exit 1 = tripped) — an attempt counts only when a
|
|
76
|
+
dispatch receipt AND either a leg-completion row or a WorkResult attest it, so a leg still
|
|
77
|
+
in flight holds it open rather than reading as exhausted. This is the SAME derivation the
|
|
78
|
+
round loop's own inner breaker reads, so the census and the breaker cannot disagree about
|
|
79
|
+
the same exhaustion the way a live run once measured.
|
|
72
80
|
H0.2 QA findings (qa-edge-hunter's hunt-report.md, when present) — all `~` by default.
|
|
73
81
|
H0.3 Discovered-task ledger entries still open (discovery/ledger.md, `[+]`/`~` unresolved).
|
|
74
82
|
H0.4 Attempt-budget hammer proposals (scopes that exhausted their T0 attempts during BUILD).
|
|
@@ -391,6 +391,9 @@ const CMD = {
|
|
|
391
391
|
// from `detail` on purpose: `detail` is prose for a human to read, this is a token the control
|
|
392
392
|
// plane branches on, and collapsing the two is what made every gate comparison silently false.
|
|
393
393
|
decision: { type: "string" },
|
|
394
|
+
// A close's own advisory failure (kernel/probe/resume.mjs's `exportOnClose`) — copied verbatim, the same discipline as `decision`, so "the export failed" is a
|
|
395
|
+
// fact `closeIfTerminal` can act on rather than a line buried inside free-text `detail`.
|
|
396
|
+
export_warning: { type: "string" },
|
|
394
397
|
},
|
|
395
398
|
required: ["exit_code", "ok"],
|
|
396
399
|
};
|
|
@@ -617,7 +620,9 @@ async function cmd(verbs, phaseName, label) {
|
|
|
617
620
|
`Report its exit code as exit_code, ok=true if and only if exit_code is 0, and one line of ` +
|
|
618
621
|
`detail. If the command printed JSON carrying a top-level "decision" key, copy that value into ` +
|
|
619
622
|
`decision EXACTLY as it appears — one bare token, no sentence, no quotes, no rephrasing. ` +
|
|
620
|
-
`Otherwise omit decision.
|
|
623
|
+
`Otherwise omit decision. If the command printed JSON carrying a top-level "export_warning" ` +
|
|
624
|
+
`key with a non-empty string value, copy that string into export_warning verbatim. Otherwise ` +
|
|
625
|
+
`omit export_warning. Do not interpret, summarise or act on the command's output beyond that.\n\n` +
|
|
621
626
|
`If the tool call itself is refused or blocked before the command ever runs — a permission or ` +
|
|
622
627
|
`policy denial, not the command's own exit — that is NOT an exit code, and you must never invent ` +
|
|
623
628
|
`one to fill the field: report exit_code as -1 and put the denial's own wording verbatim in ` +
|
|
@@ -968,11 +973,18 @@ async function setRunStatus(status, phaseName) {
|
|
|
968
973
|
const causeArg = (s) => (String(s ?? "").replace(/[`"'$\\\n\r]/g, " ").replace(/\s+/g, " ").trim().slice(0, 300) || "no reason recorded");
|
|
969
974
|
|
|
970
975
|
// A terminal RunReturn closes the run's own ledger — a terminal status, its cause, and a close
|
|
971
|
-
// timestamp, in one write
|
|
972
|
-
//
|
|
973
|
-
//
|
|
974
|
-
//
|
|
975
|
-
//
|
|
976
|
+
// timestamp, in one write. WHICH arms are terminal, and what status each closes as, is no longer
|
|
977
|
+
// decided here: `probe resume --close-arm <ret.status>` hands the kernel the RunReturn arm itself
|
|
978
|
+
// and it answers from `RUN_RETURN_CLOSE` (kernel/probe/resume.mjs), derived from the schema's own
|
|
979
|
+
// enum rather than a pair of statuses this file used to compare `ret.status` against by hand — a
|
|
980
|
+
// literal comparison that could not see a new arm go unhandled, because it never consulted the
|
|
981
|
+
// kernel's own list of terminal statuses at all. `gate_h` now closes the run as `escalated` — the
|
|
982
|
+
// breaker that tripped travels in the close's cause text, never as a status of its own; the
|
|
983
|
+
// tech-lead skill's own GATE H → L4 orchestration still runs the census and the ship decision, it
|
|
984
|
+
// just no longer finds the ledger's close fields unset when it gets there. `paused` and `ok` remain
|
|
985
|
+
// explicitly non-terminal — the kernel says so, this call site no longer needs to know why.
|
|
986
|
+
// Best-effort, the same discipline as `setRunStatus` above: a lost write degrades the trace's own
|
|
987
|
+
// record of why the run ended, it does not change what this return reports.
|
|
976
988
|
//
|
|
977
989
|
// A close can also come back `ok:true` and still be a degraded outcome: `closeRun`
|
|
978
990
|
// (kernel/probe/resume.mjs) refuses a DIFFERENT terminal status outright (that failure already hits
|
|
@@ -986,11 +998,15 @@ const causeArg = (s) => (String(s ?? "").replace(/[`"'$\\\n\r]/g, " ").replace(/
|
|
|
986
998
|
// reports — the ledger already folded the prior cause in — but the run's own RunReturn must say the
|
|
987
999
|
// trace is degraded.
|
|
988
1000
|
async function closeIfTerminal(ret) {
|
|
989
|
-
if (ret.status !== "aborted" && ret.status !== "shipped") return;
|
|
990
1001
|
const cause = ret.status === "aborted"
|
|
991
1002
|
? `${ret.aborted_at || "?"}: ${ret.reason || "no reason recorded"}`
|
|
1003
|
+
: ret.status === "gate_h"
|
|
1004
|
+
? `breaker=${ret.breaker ?? "?"} green_scopes=${Array.isArray(ret.green_scopes) ? ret.green_scopes.length : "?"} hammer_proposals=${Array.isArray(ret.hammer_proposals) ? ret.hammer_proposals.length : "?"}`
|
|
992
1005
|
: `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}`;
|
|
993
|
-
|
|
1006
|
+
// `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
|
|
1007
|
+
// a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
|
|
1008
|
+
// stay silent for it exactly as they did when this file's own guard returned early.
|
|
1009
|
+
const r = await cmd(`probe resume --slug ${slug} --close-arm ${ret.status} --cause "${causeArg(cause)}"`, "Ship", `close:${ret.status}`);
|
|
994
1010
|
if (!r.ok) {
|
|
995
1011
|
const why = (r.detail || `exit ${r.exit_code}`).trim();
|
|
996
1012
|
log(`RUN STATE — close(${ret.status}) did not take: ${why}. This return's own status and reason still ` +
|
|
@@ -1004,6 +1020,14 @@ async function closeIfTerminal(ret) {
|
|
|
1004
1020
|
`own close_cause line; this return's trace is degraded, not corrupted.`);
|
|
1005
1021
|
stateWarnings.push(`close(${ret.status}) superseded an earlier close of this run_id — see harness-run.md's close_cause for both reasons`);
|
|
1006
1022
|
}
|
|
1023
|
+
// The close itself took (r.ok above) — an export_warning here is `exportOnClose`
|
|
1024
|
+
// (kernel/probe/resume.mjs) reporting it could not project this run's
|
|
1025
|
+
// fact tables. Advisory, same as every other line in this function: the close stands, the run's
|
|
1026
|
+
// own return says the trace is short one export rather than swallowing the fact.
|
|
1027
|
+
if (r.export_warning) {
|
|
1028
|
+
log(`RUN STATE — close(${ret.status}) took, but its export did not: ${r.export_warning}`);
|
|
1029
|
+
stateWarnings.push(`close(${ret.status}): ${r.export_warning}`);
|
|
1030
|
+
}
|
|
1007
1031
|
}
|
|
1008
1032
|
const withWarnings = async (ret) => {
|
|
1009
1033
|
await closeIfTerminal(ret);
|
|
@@ -1739,11 +1763,15 @@ async function buildScope(scope, roundNo) {
|
|
|
1739
1763
|
`and keep T0 green. An entry marked \`unowned\` cites no file any scope owns — fix it only ` +
|
|
1740
1764
|
`if it falls inside your substrate. `
|
|
1741
1765
|
: "") +
|
|
1742
|
-
`Re-compile the order for every attempt after the first, with --attempt <n>. ` +
|
|
1766
|
+
`Re-compile the order for every attempt after the first, with --attempt <n>. An attempt will ` +
|
|
1767
|
+
`be REFUSED (exit 3) while the previous one is unanswered — dispatched with no leg row and no ` +
|
|
1768
|
+
`WorkResult — because grading a tree the previous attempt may still be writing counts an ` +
|
|
1769
|
+
`attempt nobody ran. Let it come back rather than opening the next one. ` +
|
|
1743
1770
|
`Run the attempt ratchet for THIS scope only: up to ${attemptBudget} attempts of implement → ` +
|
|
1744
1771
|
`\`node "${KERNEL}" verify t0 "${scope.path}" --round ${roundNo} --attempt <n>\`, each scored against ` +
|
|
1745
1772
|
`the last kept trial. Stop on the first green T0, or when the attempt budget or the stagnation ` +
|
|
1746
1773
|
`breaker trips. Write only inside this scope's substrate whitelist — the sandbox hook enforces it. ` +
|
|
1747
|
-
`Report green, attempts_used, which breaker (if any) tripped, and the T0 artifact path
|
|
1774
|
+
`Report green, attempts_used, which breaker (if any) tripped, and the T0 artifact path — ` +
|
|
1775
|
+
`attempts_used is what the attested channels carry, not a count of the orders you compiled.`,
|
|
1748
1776
|
});
|
|
1749
1777
|
}
|