shapeup-sdlc 3.5.0 → 3.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/AGENTS.md +15 -4
  3. package/README.md +1 -1
  4. package/kernel/compile.mjs +55 -7
  5. package/kernel/harness.mjs +11 -4
  6. package/kernel/init/run-args.mjs +206 -0
  7. package/kernel/init/run.mjs +10 -0
  8. package/kernel/lib/paths.mjs +10 -0
  9. package/kernel/probe/attempts.mjs +135 -0
  10. package/kernel/probe/concurrency.mjs +31 -6
  11. package/kernel/probe/digest.mjs +15 -1
  12. package/kernel/probe/owner.mjs +4 -1
  13. package/kernel/probe/resume.mjs +314 -5
  14. package/kernel/probe/rounds.mjs +104 -0
  15. package/kernel/reduce/ingest.mjs +53 -12
  16. package/kernel/reduce/ship.mjs +15 -30
  17. package/kernel/reduce/snapshot.mjs +23 -2
  18. package/kernel/report/export.mjs +54 -2
  19. package/kernel/report/facts.mjs +24 -2
  20. package/{skills/tech-lead → kernel}/schemas/domain.schema.json +15 -12
  21. package/kernel/verify/envelope.mjs +2 -2
  22. package/kernel/verify/skills.mjs +1 -1
  23. package/package.json +1 -1
  24. package/skills/ba-pitch-analyzer/SKILL.md +1 -1
  25. package/skills/ba-pitch-analyzer/references/doc-schemas.md +6 -0
  26. package/skills/coach/SKILL.md +8 -2
  27. package/skills/hill-chart/SKILL.md +3 -4
  28. package/skills/scope-hammer/SKILL.md +11 -3
  29. package/skills/tech-lead/SKILL.md +10 -10
  30. package/skills/tech-lead/references/gates.md +48 -11
  31. package/skills/tech-lead/references/protocol.md +4 -2
  32. package/skills/tech-lead/workflows/shapeup-run.js +164 -40
  33. package/skills/translator/SKILL.md +1 -1
  34. /package/{skills/tech-lead → kernel}/schemas/gate-answers.schema.json +0 -0
  35. /package/{skills/tech-lead → kernel}/schemas/work-order.schema.json +0 -0
  36. /package/{skills/tech-lead → kernel}/schemas/work-result.schema.json +0 -0
@@ -60,15 +60,15 @@ check the lane:
60
60
  legacy loop instead — `references/protocol.md` (BUILD(r)/EVAL) + `references/protocol.md`
61
61
  carry the full step-by-step for both the tiny lane and a scope-less BUILD loop, verbatim, non-
62
62
  regression. Stop reading this file here for that run.
63
- - **Otherwise** (the common case — a scoped spec, any auto level): build `RunArgs`
64
- (`domain.schema.json` `$defs/RunArgs` — `{slug, runId, autoLevel, answers, lane,
65
- models:{exec,eval,qa}, budgets:{maxRounds,attemptBudget,wallClockS}, pluginRoot, startedAt}`,
66
- plus every switch the operator typed — `references/gates.md` GATE L0.9 has the flag→field table,
67
- and a flag that stops here is a flag that was accepted and ignored). **Write that exact object to
68
- `.shapeup/<slug>/run-args.json` before launching**, fresh on every launch and relaunch: the flags
69
- reach the workflow as a value in memory, so it is the run's only evidence of what it was launched
70
- with, and a run that cannot state its own configuration cannot have a claim about it checked.
71
- Then launch with the **`Workflow` tool** — naming `init run`'s staged copy, never the install path:
63
+ - **Otherwise** (the common case — a scoped spec, any auto level): resolve every switch the
64
+ operator typed (`references/gates.md` GATE L0.9b has the flag→field table) and run the kernel's
65
+ sole `RunArgs` writer: `node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" init run-args --slug
66
+ <slug> --auto-level <level> --exec-model <n> [--eval-model <n>] [--qa-model <n>] --max-rounds <N>
67
+ --attempts <N> --plugin-root "${CLAUDE_PLUGIN_ROOT}" [--answers <a>] [--lane <l>] [--no-eval]
68
+ [--no-qa] [--adversarial-verify] [--parallel-scopes <N>]`. It writes `.shapeup/<slug>/run-args.json`
69
+ fresh on every launch/relaunch and prints that identical object — **pass it to `Workflow`
70
+ verbatim, never re-type it**. Then launch with the **`Workflow` tool** — naming `init run`'s
71
+ staged copy, never the install path:
72
72
 
73
73
  ```
74
74
  Workflow({
@@ -114,7 +114,7 @@ did not actually receive from the PO — an unattended lane with no answer for a
114
114
 
115
115
  FIRST freeze the evidence — run state is gitignored, so `shapeup/<slug>/REPORT.md` (already
116
116
  written by `shapeup-run.js` via `harness reduce ship`, or write it now on a `gate_h` close) is all a
117
- teammate sees. Then emit:
117
+ teammate sees. Then RESOLVE the gate — `references/gates.md` GATE L4 has the call — and emit:
118
118
 
119
119
  ```
120
120
  ⏸ GATE L4 — Ship Sign-Off
@@ -113,11 +113,13 @@ Collect (explicit — never inferred):
113
113
  gate to be its first execution.
114
114
  ```
115
115
 
116
- **L0.9b — the launch record.** Every switch the operator typed becomes a `RunArgs` field, or it
117
- does nothing at all: the workflow cannot read a config file and cannot ask a follow-up, so a flag
118
- that stops at the skill boundary was accepted and ignored. That is not hypothetical — `--no-qa` was
119
- documented in seven places across the shipped set and inert in all of them, because no line of this
120
- protocol ever put `noQa` into the record.
116
+ **L0.9b — the launch record.** Every switch the operator typed to *this launch* becomes a `RunArgs`
117
+ field, or it does nothing at all: the workflow cannot read a config file and cannot ask a follow-up,
118
+ so a flag that stops at the skill boundary was accepted and ignored. That is not hypothetical —
119
+ `--no-qa` was documented in seven places across the shipped set and inert in all of them, because no
120
+ line of this protocol ever put `noQa` into the record. `--wall-clock-budget` is the one flag below
121
+ that is not a `RunArgs` field at all — it is consumed earlier, at `init run` itself, and never
122
+ needed to reach this launch; see its row for where it actually lands.
121
123
 
122
124
  | Flag | `RunArgs` field |
123
125
  |---|---|
@@ -125,14 +127,19 @@ protocol ever put `noQa` into the record.
125
127
  | `--no-qa` | `noQa: true` |
126
128
  | `--parallel-scopes N` | `maxParallelScopes: N` — how many scopes build at once (default 4; `1` = sequential) |
127
129
  | `--adversarial-verify` | `adversarialVerify: true` |
128
- | `--rounds N` / `--attempts N` / `--wall-clock-budget S` | `budgets.{maxRounds,attemptBudget,wallClockS}` |
130
+ | `--rounds N` / `--attempts N` | `budgets.{maxRounds,attemptBudget}` |
129
131
  | `--gate-answers <set>` | `answers` |
132
+ | `--wall-clock-budget S` | *(not a `RunArgs` field)* — typed once, on the `harness init run` command line itself, not on this launch; it lands straight in the run receipt as `wall_clock_budget_s`, and the deadline breaker reads that receipt field directly — consumed by `kernel/verify/budget.mjs` as `wall_clock_budget_s`. `budgets` declares only `maxRounds`/`attemptBudget` — the schema, `SKILL.md`'s own RunArgs contract line and this script's own header comment all agree there is no third member |
130
133
  | `--orch-model/--exec-model/--eval-model/--qa-model` | `models.{…}` (L0.8) |
131
134
 
132
- The assembled object is written to `.shapeup/<slug>/run-args.json` before the launch, fresh on every
133
- launch and relaunch. It is the only artifact that records what a run was configured with; the ship
134
- report, a resumed session and any later measurement all read it, and none of them can recover a
135
- value that only ever existed as an argument.
135
+ `harness init run-args` (invoked at Step 2 of `SKILL.md`) is the sole writer of the assembled
136
+ object: it takes the resolved values above, writes `.shapeup/<slug>/run-args.json` fresh on every
137
+ launch and relaunch, and prints the same object back so the launch never re-assembles it by hand. It
138
+ is the only artifact that records what a run was configured with; the ship report, a resumed session
139
+ and any later measurement all read it, and none of them can recover a value that only ever existed
140
+ as an argument. Step 2 is not merely advisory: `shapeup-run.js`'s own Preflight refuses to dispatch
141
+ ORIENT (or anything past it) when this file is missing at the run's local root — a launch that
142
+ skipped this step aborts there rather than proceeding on a silent default.
136
143
 
137
144
  **L0.0 — intake precondition (before any other L0 collection):**
138
145
  ```
@@ -161,7 +168,14 @@ Model matrix : orch=[model] exec=[model] eval=[model] qa=[model] digester=[scrip
161
168
  Budgets : round_budget=[N] (outer) attempt_budget=[N] (inner, per scope)
162
169
  Knowledge : [tech-lead.md — N workflow rules, M suggested values (confirmed above) | none — `/retro --scan` or `/retro --research <stack>` seeds it (optional)]
163
170
  ```
164
- Do NOT start ORIENT until confirmed (interactive/auto). Under --unattended, proceed.
171
+ **Resolve it** — this gate is this skill's own (the workflow never sees it), so it is this skill
172
+ that runs the same tool every other gate resolves through, not a paragraph read as a stand-in for
173
+ one: `node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" gate --resolve L0 --slug <slug>
174
+ [--file <path>|--preset <name>]`. Exit 4 (`ask`) is the confirmation this block already asks for —
175
+ put it to the PO and wait, same as the paragraph above always meant. Exit 0 (`decision=proceed`) —
176
+ continue straight to ORIENT, which is what `--unattended`'s pre-answered set resolves to. Exit 5
177
+ (`abort`) — stop; do not launch. Either way, the gate's own ledger row is what lets a later reader
178
+ see the decision that opened the run, not only the decisions that closed it.
165
179
 
166
180
  ---
167
181
 
@@ -541,5 +555,28 @@ On confirm:
541
555
  - If the PO provides substantive feedback (not just 'y' or empty) → automatically delegate via Agent (model: exec — see references/protocol.md "Invocation mechanism"): Skill(shapeup-sdlc-plugin:coach) with the provided feedback for RLHF. The coach runs its own GATE COACH-1 to have the PO categorize each rule, then files it under the responsible skill in `shapeup/knowledge-base/<skill>.md` (committed → team-shared). Coachable: `task-executor`, `ba-pitch-analyzer`, `qa-edge-hunter`, `orient`, `scope-architect`, `solution-architect` (each reads its own file at the top of its next run) and `tech-lead` (workflow guidance, read at the next GATE L0). Guidance never decides a gate: a filed rule may add a question or a check to a gate block, never an answer. The tech lead does not categorize the feedback itself — that is the coach's gate, by design (no assumptions).
542
556
  - Then output → `✅ [slug] [shipped & deployed | built & verified, deploy pending] — [r] rounds, verdict PASS.`
543
557
 
558
+ **Resolve the gate itself before any of the above** — this is the decision that shipped the run,
559
+ and without it the trace holds no record of that decision at all: `node
560
+ "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" gate --resolve L4 --slug <slug>
561
+ [--file <path>|--preset <name>]`. Exit 0 (`decision=ship|hold`) — render the block above and close
562
+ the run: `node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe resume --slug <slug> --close shipped
563
+ --cause "verdict=<verdict> rounds=<r> decision=<ship|hold>"`. Always issue this call — a `gate_h`
564
+ close is the ordinary case where `shapeup-run.js` handed off without closing the run, and the
565
+ GATE H → L4 path (scope-hammer's census, then this gate) is the one this instruction exists for.
566
+ The close itself is a once-only fact IN THE KERNEL (`closeRun`'s own guard reads a `closed_status:`
567
+ line that only `closeRun` ever writes — never the mutable `status:` line every phase rewrites, this
568
+ call included), not a conditional this instruction has to get right: if this run_id was NOT already
569
+ closed, this call performs the close, fresh. If it was already closed `shipped` (or `aborted`) and
570
+ the cause text is byte-identical to what is already on the ledger, this call is a true idempotent
571
+ no-op. If it was already closed with the SAME status but a genuinely different cause — a run closed
572
+ more than once across relaunches, the ordinary shape a `gate_h` hand-off after an earlier abort takes
573
+ — this call SUPERSEDES it: the new cause is written, the prior one is folded into the same
574
+ `close_cause` line rather than lost, and the kernel call itself still exits 0 (only the RunReturn a
575
+ launch's own `withWarnings` wraps carries the resulting `state_warning` — a prose-driven close like
576
+ this one has no RunReturn to attach it to, so read `close_cause` by hand if this branch matters to
577
+ you). If it was already closed with a DIFFERENT status altogether, this call is refused outright —
578
+ cause intact, never silently flipped. Exit 4 (`ask`) — the block above IS that stop; put it to the
579
+ PO and wait, same as always. L4's answer set carries no `abort`.
580
+
544
581
  ---
545
582
 
@@ -630,7 +630,7 @@ trace. See `references/gates.md` — GATE L0.1.
630
630
  ## Central domain registry
631
631
 
632
632
  Every record type and payload field that crosses a skill boundary is defined exactly once in
633
- `skills/tech-lead/schemas/domain.schema.json` — the envelope schemas (`work-order.schema.json`,
633
+ `kernel/schemas/domain.schema.json` — the envelope schemas (`work-order.schema.json`,
634
634
  `work-result.schema.json`) only `$ref` it. The registry annotates each entity's tier
635
635
  (SHARED/LOCAL), location, sole writer, and readers, carries the machine-readable ERD (`x-erd`),
636
636
  and maps which payload fields each worker may rely on (`x-payload-by-worker`).
@@ -686,13 +686,15 @@ lens: lite | standard | cross-context
686
686
  eval_dimensions: [spec-conformance] # the set from GATE L0.5 (init-run --dimensions); every EVAL order is compiled from THIS line
687
687
  max_rounds: 3
688
688
  auto_level: interactive | auto | unattended
689
- status: orienting | mapping | building | evaluating | shipped | escalated
689
+ status: orienting | mapping | building | evaluating | shipped | escalated | aborted
690
690
  final_verdict: ~ | pass | fail | not-evaluated
691
691
  rounds_used: [N]
692
692
  discovered_rounds: [N]
693
693
  deploy: ~ | deployed | pending-po
694
694
  started_at: [ISO]
695
695
  closed_at: ~ | [ISO]
696
+ close_cause: ~ | [why the run ended at that terminal status — `probe resume --close` writes this and closed_at together]
697
+ closed_status: ~ | [the terminal status actually closed — written ONLY by `probe resume --close`, never by `--set-status`, so it is immune to `status:` above being rewritten by ordinary phase traffic after the close]
696
698
  ---
697
699
  ```
698
700
 
@@ -33,7 +33,7 @@
33
33
  //
34
34
  // args — RunArgs (domain.schema.json $defs/RunArgs):
35
35
  // slug, autoLevel (interactive|auto|unattended), answers (preset name or path),
36
- // models {exec, eval, qa?}, budgets {maxRounds, attemptBudget, wallClockS?}, pluginRoot,
36
+ // models {exec, eval, qa?}, budgets {maxRounds, attemptBudget}, pluginRoot,
37
37
  // startedAt, and the optional switches noEval / noQa / adversarialVerify /
38
38
  // maxParallelScopes (default 4).
39
39
  //
@@ -391,6 +391,9 @@ const CMD = {
391
391
  // from `detail` on purpose: `detail` is prose for a human to read, this is a token the control
392
392
  // plane branches on, and collapsing the two is what made every gate comparison silently false.
393
393
  decision: { type: "string" },
394
+ // A close's own advisory failure (kernel/probe/resume.mjs's `exportOnClose`) — copied verbatim, the same discipline as `decision`, so "the export failed" is a
395
+ // fact `closeIfTerminal` can act on rather than a line buried inside free-text `detail`.
396
+ export_warning: { type: "string" },
394
397
  },
395
398
  required: ["exit_code", "ok"],
396
399
  };
@@ -617,7 +620,9 @@ async function cmd(verbs, phaseName, label) {
617
620
  `Report its exit code as exit_code, ok=true if and only if exit_code is 0, and one line of ` +
618
621
  `detail. If the command printed JSON carrying a top-level "decision" key, copy that value into ` +
619
622
  `decision EXACTLY as it appears — one bare token, no sentence, no quotes, no rephrasing. ` +
620
- `Otherwise omit decision. Do not interpret, summarise or act on the command's output beyond that.\n\n` +
623
+ `Otherwise omit decision. If the command printed JSON carrying a top-level "export_warning" ` +
624
+ `key with a non-empty string value, copy that string into export_warning verbatim. Otherwise ` +
625
+ `omit export_warning. Do not interpret, summarise or act on the command's output beyond that.\n\n` +
621
626
  `If the tool call itself is refused or blocked before the command ever runs — a permission or ` +
622
627
  `policy denial, not the command's own exit — that is NOT an exit code, and you must never invent ` +
623
628
  `one to fill the field: report exit_code as -1 and put the denial's own wording verbatim in ` +
@@ -901,6 +906,47 @@ async function fastForward(gate, phaseKey, phaseName, what) {
901
906
  `${phaseKey} artifact by hand before assuming it regressed, then relaunch.`);
902
907
  }
903
908
 
909
+ /**
910
+ * The positive enforcer for GATE L0.9b's launch record: the RunArgs object this run was
911
+ * configured with, kernel-written to the run's own local root, is on disk before this file
912
+ * dispatches anything past Preflight.
913
+ *
914
+ * Nothing upstream of this call can be trusted to have written it. `SKILL.md` Step 2 tells the
915
+ * orchestrating session to run `harness init run-args` right before `Workflow(...)` is invoked, but
916
+ * that is prose the session reads, not a check anything runs — and a session that launches this
917
+ * script without having done it left no trace anywhere else: the receipt, `intake.md` and
918
+ * `harness-run.md` all exist regardless, and `dialFrom()` (`kernel/probe/concurrency.mjs`) answers
919
+ * every reader with a silent default rather than an error. So the record's existence would depend
920
+ * entirely on whether a prior step of prose was followed — exactly the shape this file's own header
921
+ * exists to retire (no rule with no enforcer). `probe concurrency --require-run-args` is a presence
922
+ * check, not a measurement: it shares `dialFrom()`'s run-root resolution and file path so there is
923
+ * one definition of where the record lives, not two.
924
+ *
925
+ * @returns {Promise<(object|null)>} An aborted RunReturn, or null when the record is there.
926
+ */
927
+ async function requireLaunchRecord() {
928
+ const r = await cmd(`probe concurrency --slug ${slug} --require-run-args`, "Preflight", "launch-record");
929
+ if (r.exit_code === 0) return null;
930
+ // Same distinction requirePhase/fastForward draw: only exit 6 is the predicate genuinely
931
+ // answering "not there". Anything else means the check itself did not run.
932
+ if (r.exit_code === 6) {
933
+ return aborted("preflight",
934
+ `the run's launch record does not exist — GATE L0.9b's RunArgs (the model matrix, budgets and ` +
935
+ `every operator switch this run was typed with) was never written to disk before this launch, ` +
936
+ `so nothing downstream can be attested against what the operator actually configured. Run ` +
937
+ `\`harness init run-args\` (references/gates.md L0.9b names every flag it resolves — \`SKILL.md\` ` +
938
+ `Step 2 runs it right before this workflow launches) and relaunch.`);
939
+ }
940
+ return aborted("preflight",
941
+ `the launch-record check did not run to a verdict — exit_code ${r.exit_code}, not the 0 (present) ` +
942
+ `or 6 (absent) \`probe concurrency --require-run-args\` documents.` +
943
+ `${r.detail ? ` Courier reported: ${r.detail}.` : ""} This is NOT evidence the record is missing — ` +
944
+ `most often a Bash call denied above this plugin's own hooks and permission grant (an untrusted ` +
945
+ `workspace, or Claude Code's auto-mode classifier). Verify the run's launch record by hand — ` +
946
+ `\`harness probe concurrency --slug ${slug} --require-run-args\` — before assuming it is absent, ` +
947
+ `then relaunch.`);
948
+ }
949
+
904
950
  // The ledger's `status` field is bookkeeping, not this file's resume oracle — the fast-forward reads
905
951
  // artifacts. It survives because `reduce snapshot` and the the ship report's census hook read it to tell
906
952
  // a run in flight from a finished one. A lost write is a degraded digest, not a corrupted build, so
@@ -917,7 +963,76 @@ async function setRunStatus(status, phaseName) {
917
963
  stateWarnings.push(`status="${status}" did not take: ${why}`);
918
964
  }
919
965
  }
920
- const withWarnings = (ret) => (stateWarnings.length ? { ...ret, state_warnings: stateWarnings } : ret);
966
+
967
+ // A free-text reason, made safe to spell into a sub-agent's shell instruction: no quotes, no
968
+ // newlines, no backticks or `$` — every one of those risks breaking the command the sub-agent
969
+ // itself constructs from this file's own instruction text (see `cmd()`'s banner: every kernel call
970
+ // in this script is spelled into a prompt, not spawned directly). Truncated, not elided: the exit
971
+ // code and the ledger's own close_cause line are still the source of record — this is only what
972
+ // crosses the prompt boundary.
973
+ const causeArg = (s) => (String(s ?? "").replace(/[`"'$\\\n\r]/g, " ").replace(/\s+/g, " ").trim().slice(0, 300) || "no reason recorded");
974
+
975
+ // A terminal RunReturn closes the run's own ledger — a terminal status, its cause, and a close
976
+ // timestamp, in one write. WHICH arms are terminal, and what status each closes as, is no longer
977
+ // decided here: `probe resume --close-arm <ret.status>` hands the kernel the RunReturn arm itself
978
+ // and it answers from `RUN_RETURN_CLOSE` (kernel/probe/resume.mjs), derived from the schema's own
979
+ // enum rather than a pair of statuses this file used to compare `ret.status` against by hand — a
980
+ // literal comparison that could not see a new arm go unhandled, because it never consulted the
981
+ // kernel's own list of terminal statuses at all. `gate_h` now closes the run as `escalated` — the
982
+ // breaker that tripped travels in the close's cause text, never as a status of its own; the
983
+ // tech-lead skill's own GATE H → L4 orchestration still runs the census and the ship decision, it
984
+ // just no longer finds the ledger's close fields unset when it gets there. `paused` and `ok` remain
985
+ // explicitly non-terminal — the kernel says so, this call site no longer needs to know why.
986
+ // Best-effort, the same discipline as `setRunStatus` above: a lost write degrades the trace's own
987
+ // record of why the run ended, it does not change what this return reports.
988
+ //
989
+ // A close can also come back `ok:true` and still be a degraded outcome: `closeRun`
990
+ // (kernel/probe/resume.mjs) refuses a DIFFERENT terminal status outright (that failure already hits
991
+ // the `!r.ok` branch below, exit 3), but the SAME status with a DIFFERENT cause — a run_id closed
992
+ // more than once across relaunches, e.g. `aborted` at Preflight then `aborted` again for a different
993
+ // reason after a later relaunch — is *superseded* rather than refused: `ok:true`, the new cause
994
+ // folded together with the old one on disk, and `decision:"superseded"` (the one bare token `cmd()`
995
+ // relays verbatim across the courier boundary — see its own banner) naming that outcome. Reporting
996
+ // that as a clean, silent success would be a state fact the trace needs with nothing surfacing it —
997
+ // the same shape as leaving a run's ledger open after it ended. It costs nothing this return itself
998
+ // reports — the ledger already folded the prior cause in — but the run's own RunReturn must say the
999
+ // trace is degraded.
1000
+ async function closeIfTerminal(ret) {
1001
+ const cause = ret.status === "aborted"
1002
+ ? `${ret.aborted_at || "?"}: ${ret.reason || "no reason recorded"}`
1003
+ : ret.status === "gate_h"
1004
+ ? `breaker=${ret.breaker ?? "?"} green_scopes=${Array.isArray(ret.green_scopes) ? ret.green_scopes.length : "?"} hammer_proposals=${Array.isArray(ret.hammer_proposals) ? ret.hammer_proposals.length : "?"}`
1005
+ : `verdict=${ret.verdict ?? "?"} rounds=${ret.rounds_used ?? "?"} qa_findings=${ret.qa_findings ?? "?"}`;
1006
+ // `--close-arm` hands the kernel the arm itself (not a status this file decided was terminal) —
1007
+ // a non-terminal arm (`paused`, `ok`) still exits 0 with no "decision" key, so the branches below
1008
+ // stay silent for it exactly as they did when this file's own guard returned early.
1009
+ const r = await cmd(`probe resume --slug ${slug} --close-arm ${ret.status} --cause "${causeArg(cause)}"`, "Ship", `close:${ret.status}`);
1010
+ if (!r.ok) {
1011
+ const why = (r.detail || `exit ${r.exit_code}`).trim();
1012
+ log(`RUN STATE — close(${ret.status}) did not take: ${why}. This return's own status and reason still ` +
1013
+ `stand; only the ledger's own closed_at/close_cause record is degraded.`);
1014
+ stateWarnings.push(`close(${ret.status}) did not take: ${why}`);
1015
+ return;
1016
+ }
1017
+ if (String(r.decision ?? "").trim() === "superseded") {
1018
+ log(`RUN STATE — close(${ret.status}) superseded an earlier close recorded under the same status ` +
1019
+ `but a different cause — this run_id was closed more than once. Both causes are on the ledger's ` +
1020
+ `own close_cause line; this return's trace is degraded, not corrupted.`);
1021
+ stateWarnings.push(`close(${ret.status}) superseded an earlier close of this run_id — see harness-run.md's close_cause for both reasons`);
1022
+ }
1023
+ // The close itself took (r.ok above) — an export_warning here is `exportOnClose`
1024
+ // (kernel/probe/resume.mjs) reporting it could not project this run's
1025
+ // fact tables. Advisory, same as every other line in this function: the close stands, the run's
1026
+ // own return says the trace is short one export rather than swallowing the fact.
1027
+ if (r.export_warning) {
1028
+ log(`RUN STATE — close(${ret.status}) took, but its export did not: ${r.export_warning}`);
1029
+ stateWarnings.push(`close(${ret.status}): ${r.export_warning}`);
1030
+ }
1031
+ }
1032
+ const withWarnings = async (ret) => {
1033
+ await closeIfTerminal(ret);
1034
+ return stateWarnings.length ? { ...ret, state_warnings: stateWarnings } : ret;
1035
+ };
921
1036
 
922
1037
  // =============================================================================================
923
1038
  // THE RUN
@@ -950,20 +1065,25 @@ await agent(
950
1065
  );
951
1066
  const canary = await cmd(`verify dispatch --skill ${canarySkill} --within 900`, "Preflight", "canary-evidence");
952
1067
  if (!canary.ok) {
953
- return aborted("preflight",
1068
+ return await withWarnings(aborted("preflight",
954
1069
  `the ${canarySkill} skill did not resolve in this session — no dispatch reached the hook layer. ` +
955
1070
  `A run would report phases completing while the sub-agents improvised every worker's craft. ` +
956
1071
  `Load the plugin (\`claude --plugin-dir <repo>\`, or install and enable it) and relaunch. ` +
957
- `(${canary.detail || `exit ${canary.exit_code}`})`);
1072
+ `(${canary.detail || `exit ${canary.exit_code}`})`));
958
1073
  }
959
1074
 
1075
+ // GATE L0.9b's launch record must exist before anything past Preflight dispatches; see
1076
+ // requireLaunchRecord()'s own banner for why this cannot be left to Step 2's prose alone.
1077
+ const launchRecordAbort = await requireLaunchRecord();
1078
+ if (launchRecordAbort) return await withWarnings(launchRecordAbort);
1079
+
960
1080
  phase("Orient");
961
1081
 
962
1082
  const rs = await query(`probe resume --slug ${slug}`, RESUME, "Orient", "resume-state");
963
1083
  // A probe that produced nothing is not an EMPTY run — it is an unknown one. Treating it as empty
964
1084
  // would re-dispatch every phase from the top, over a run that may be in progress.
965
1085
  if (!rs) {
966
- return aborted("probe", "the fast-forward derivation returned no state — refusing to re-dispatch a run that may already be in progress");
1086
+ return await withWarnings(aborted("probe", "the fast-forward derivation returned no state — refusing to re-dispatch a run that may already be in progress"));
967
1087
  }
968
1088
 
969
1089
  const specFolder = rs.spec_folder || `shapeup/${slug}/spec/`;
@@ -991,14 +1111,14 @@ if (!rs.has_orient_artifacts) {
991
1111
  "when the risk scan came back rank 0). Any other filename leaves the phase incomplete and " +
992
1112
  "the run aborts, however good the contents are.",
993
1113
  });
994
- if (o.__failed) return diedAt("ORIENT", o);
1114
+ if (o.__failed) return await withWarnings(diedAt("ORIENT", o));
995
1115
  const post = await requirePhase("ORIENT", "orient", "Orient");
996
- if (post) return withWarnings(post);
1116
+ if (post) return await withWarnings(post);
997
1117
  await advisory(`reduce graph --slug ${slug}`, "Orient", "graph:orient");
998
1118
  spikedArea = o.spiked_area; spikeResult = o.spike_result; riskiest = o.riskiest_unknowns || [];
999
1119
  } else {
1000
1120
  const post = await fastForward("ORIENT", "orient", "Orient", "artifacts already on disk");
1001
- if (post) return withWarnings(post);
1121
+ if (post) return await withWarnings(post);
1002
1122
  }
1003
1123
 
1004
1124
  {
@@ -1006,7 +1126,7 @@ if (!rs.has_orient_artifacts) {
1006
1126
  // downstream artifact reads the same whether or not the pitch's second half reached the run.
1007
1127
  const g = await crossGate("L1a", "Orient", ["proceed", "ask", "abort"],
1008
1128
  { breadboard: rs.breadboard_source ?? "none", spiked_area: spikedArea, spike_result: spikeResult, riskiest_unknowns: riskiest });
1009
- if (g.stop) return withWarnings(g.stop);
1129
+ if (g.stop) return await withWarnings(g.stop);
1010
1130
  }
1011
1131
 
1012
1132
  // ---- COVERAGE (the requirements registry) — ahead of ANALYZE, whose ACs cite its ids ----------
@@ -1043,7 +1163,7 @@ if (!rs.has_requirements) {
1043
1163
  "clause per row. A clause carrying an R-id keeps its number as REQ-<n> and records the R-id " +
1044
1164
  "verbatim in its source cell; ids are assigned once and never renumbered.",
1045
1165
  });
1046
- if (c.__failed) return diedAt("COVERAGE", c);
1166
+ if (c.__failed) return await withWarnings(diedAt("COVERAGE", c));
1047
1167
  await advisory(`reduce graph --slug ${slug}`, "Analyze", "graph:coverage");
1048
1168
  } else {
1049
1169
  log(`COVERAGE — a requirements registry is already on disk; not re-dispatching it`);
@@ -1063,13 +1183,13 @@ if (!rs.has_spec_tree) {
1063
1183
  payload: { pitch: rs.intake_path, breadboard: rs.breadboard_path, spec_folder: specFolder, feature: slug, lens: rs.lens, orient_dir: rs.orient_dir },
1064
1184
  extra: "Write the spec tree and the board from the orient artifacts — do not re-scan the code.",
1065
1185
  });
1066
- if (a.__failed) return diedAt("ANALYZE", a);
1186
+ if (a.__failed) return await withWarnings(diedAt("ANALYZE", a));
1067
1187
  const post = await requirePhase("ANALYZE", "analyze", "Analyze");
1068
- if (post) return withWarnings(post);
1188
+ if (post) return await withWarnings(post);
1069
1189
  await advisory(`reduce graph --slug ${slug}`, "Analyze", "graph:analyze");
1070
1190
  } else {
1071
1191
  const post = await fastForward("ANALYZE", "analyze", "Analyze", "spec tree already on disk");
1072
- if (post) return withWarnings(post);
1192
+ if (post) return await withWarnings(post);
1073
1193
  }
1074
1194
 
1075
1195
  // ---- WIRE + GATE L1a.5 ------------------------------------------------------------------------
@@ -1083,7 +1203,7 @@ if (!rs.has_wiring_map) {
1083
1203
  // stays false and every relaunch re-dispatches and re-escalates identically. The orchestrator
1084
1204
  // holds the state a gate needs; it should not hand the check to the LLM it is about to pay for.
1085
1205
  if (!rs.has_project_profile) {
1086
- return withWarnings(aborted("WIRE",
1206
+ return await withWarnings(aborted("WIRE",
1087
1207
  `missing SHARED project-profile.md at ${rs.project_profile_path} — GATE L0 writes it ` +
1088
1208
  `({schema_version:1, archetype, entry_point}; references/gates.md GATE L0 §PROFILE) before ` +
1089
1209
  `this workflow launches. WIRE cannot resolve an entry_call_site without an entry_point to ` +
@@ -1095,18 +1215,18 @@ if (!rs.has_wiring_map) {
1095
1215
  payload: { feature: slug, spec_folder: specFolder, project_profile: rs.project_profile_path, breadboard: rs.breadboard_path },
1096
1216
  extra: "Write the wiring map: per use case, engine → seam → entry-point call site → affordance.",
1097
1217
  });
1098
- if (w.__failed) return diedAt("WIRE", w);
1218
+ if (w.__failed) return await withWarnings(diedAt("WIRE", w));
1099
1219
  const post = await requirePhase("WIRE", "wire", "Wire");
1100
- if (post) return withWarnings(post);
1220
+ if (post) return await withWarnings(post);
1101
1221
  await advisory(`reduce graph --slug ${slug}`, "Wire", "graph:wire");
1102
1222
  } else {
1103
1223
  const post = await fastForward("WIRE", "wire", "Wire", "wiring map already on disk");
1104
- if (post) return withWarnings(post);
1224
+ if (post) return await withWarnings(post);
1105
1225
  }
1106
1226
 
1107
1227
  {
1108
1228
  const g = await crossGate("L1a.5", "Wire", ["proceed", "ask", "abort"], { wiring_map: "written" });
1109
- if (g.stop) return withWarnings(g.stop);
1229
+ if (g.stop) return await withWarnings(g.stop);
1110
1230
  }
1111
1231
 
1112
1232
  // ---- MAP SCOPES + GATE L1b --------------------------------------------------------------------
@@ -1141,15 +1261,15 @@ if (scopes.length === 0) {
1141
1261
  "with prose appended to it is not runnable. A scope whose fixtures do not parse has nothing " +
1142
1262
  "to verify it and is refused at the board review.",
1143
1263
  });
1144
- if (m.__failed) return diedAt("MAP SCOPES", m);
1264
+ if (m.__failed) return await withWarnings(diedAt("MAP SCOPES", m));
1145
1265
  const post = await requirePhase("MAP SCOPES", "map-scopes", "MapScopes");
1146
- if (post) return withWarnings(post);
1266
+ if (post) return await withWarnings(post);
1147
1267
  await advisory(`reduce graph --slug ${slug}`, "MapScopes", "graph:map-scopes");
1148
1268
  scopes = m.scopes;
1149
1269
  } else {
1150
1270
  const post = await fastForward("MAP SCOPES", "map-scopes", "MapScopes",
1151
1271
  `${scopes.length} scope contract(s) already on disk`);
1152
- if (post) return withWarnings(post);
1272
+ if (post) return await withWarnings(post);
1153
1273
  }
1154
1274
 
1155
1275
  // DEPENDENCY ORDER — a scope is never built beside a scope it consumes.
@@ -1214,14 +1334,14 @@ if (waves.length > 1 || excluded.added || ceiling < maxParallelScopes) {
1214
1334
  // has more than one kind of red, and the detail says which.
1215
1335
  const specLint = await cmd(`verify spec --slug ${slug}`, "MapScopes", "spec-lint");
1216
1336
  if (!specLint.ok) {
1217
- return aborted("L1b", `spec-lint reported red findings before BUILD: ${specLint.detail || `exit ${specLint.exit_code}`}`);
1337
+ return await withWarnings(aborted("L1b", `spec-lint reported red findings before BUILD: ${specLint.detail || `exit ${specLint.exit_code}`}`));
1218
1338
  }
1219
1339
  await advisory(`verify trace --slug ${slug} --quiet`, "MapScopes", "trace-lint");
1220
1340
  await advisory(`reduce hill --slug ${slug}`, "MapScopes", "hill-derive");
1221
1341
 
1222
1342
  {
1223
1343
  const g = await crossGate("L1b", "MapScopes", ["proceed", "ask", "abort"], { scopes: scopes.map((s) => s.scope_id) });
1224
- if (g.stop) return withWarnings(g.stop);
1344
+ if (g.stop) return await withWarnings(g.stop);
1225
1345
  }
1226
1346
 
1227
1347
  // =============================================================================================
@@ -1288,7 +1408,7 @@ while (verdict !== "pass" && round <= maxRounds) {
1288
1408
  const budget = await cmd(`verify budget --slug ${slug} --strict`, "Build", `budget:r${round}`);
1289
1409
  if (budget.exit_code === 6) {
1290
1410
  await advisory(`reduce hill --slug ${slug}`, "Build", "hill-derive");
1291
- return withWarnings({ status: "gate_h", breaker: "deadline", hammer_proposals: allHammer, green_scopes: allGreen });
1411
+ return await withWarnings({ status: "gate_h", breaker: "deadline", hammer_proposals: allHammer, green_scopes: allGreen });
1292
1412
  }
1293
1413
 
1294
1414
  log(`BUILD round ${round} — ${scopes.length} scope(s), up to ${maxParallelScopes} at once, attempt budget ${attemptBudget}`);
@@ -1424,7 +1544,7 @@ while (verdict !== "pass" && round <= maxRounds) {
1424
1544
  // INNER breaker: nothing green and something queued → GATE H. The census is scope-hammer's job.
1425
1545
  if (roundGreen.length === 0 && roundHammer.length > 0) {
1426
1546
  await advisory(`reduce hill --slug ${slug}`, "Build", "hill-derive");
1427
- return withWarnings({ status: "gate_h", breaker: "inner", hammer_proposals: allHammer, green_scopes: allGreen });
1547
+ return await withWarnings({ status: "gate_h", breaker: "inner", hammer_proposals: allHammer, green_scopes: allGreen });
1428
1548
  }
1429
1549
 
1430
1550
  // ---- ROUND BUILD GATE — the feature builds and launches, measured before anyone is asked --------
@@ -1461,7 +1581,7 @@ while (verdict !== "pass" && round <= maxRounds) {
1461
1581
  {
1462
1582
  const g = await crossGate("L2", "Build", ["proceed", "ask", "abort"],
1463
1583
  { round, green_scopes: roundGreen, hammer_proposals: roundHammer, build_gate: buildGate });
1464
- if (g.stop) return withWarnings(g.stop);
1584
+ if (g.stop) return await withWarnings(g.stop);
1465
1585
  }
1466
1586
 
1467
1587
  // ---- EVAL — exactly one feature-level pass per round (the single-judge invariant) ------------
@@ -1486,16 +1606,16 @@ while (verdict !== "pass" && round <= maxRounds) {
1486
1606
  payload: { dimensions: evalDims, run_cmd: rs.run_cmd, round },
1487
1607
  extra: "Evaluate the running feature against every acceptance criterion and Done-when. One feature-level pass; cite every artifact the order lists under t0_artifacts, re-hashing each yourself.",
1488
1608
  });
1489
- if (e.__failed) return diedAt("L3", e);
1609
+ if (e.__failed) return await withWarnings(diedAt("L3", e));
1490
1610
  // The pass/fail branch is decided from the WorkResult on disk, not from the dispatching
1491
1611
  // agent's own summary of it (`e.overall`) — see EVAL_VERDICT's comment for why.
1492
1612
  const ev = await query(`probe eval --slug ${slug} --round ${round}`, EVAL_VERDICT, "Eval", `verdict:r${round}`);
1493
- if (!ev) return diedAt("L3", nullFail(`verdict:r${round}`));
1613
+ if (!ev) return await withWarnings(diedAt("L3", nullFail(`verdict:r${round}`)));
1494
1614
  // A round with no verdict to act on is NOT a dead worker. An evaluator that refused the round
1495
1615
  // wrote a result saying why, and `probe eval` carries it as `reason`; reported as "died after
1496
1616
  // retries", the one sentence naming the cause stayed in a file nobody was pointed at.
1497
1617
  if (!ev.ok || !ev.overall) {
1498
- return diedAt("L3", { __failed: `verdict:r${round}: no verdict this round can act on — ${ev.reason || `status ${ev.status || "unknown"}`}` });
1618
+ return await withWarnings(diedAt("L3", { __failed: `verdict:r${round}: no verdict this round can act on — ${ev.reason || `status ${ev.status || "unknown"}`}` }));
1499
1619
  }
1500
1620
  verdict = ev.overall === "PASS" ? "pass" : "fail";
1501
1621
  findings = e.findings || [];
@@ -1525,24 +1645,24 @@ while (verdict !== "pass" && round <= maxRounds) {
1525
1645
  await advisory(`reduce graph --slug ${slug}`, "Eval", `graph:eval-r${round}`);
1526
1646
  await advisory(`reduce hill --slug ${slug}`, "Eval", "hill-derive");
1527
1647
  const g3 = await crossGate("L3", "Eval", ["loop", "stop", "ask"], { round, verdict, build_gate: buildGate });
1528
- if (g3.stop) return withWarnings(g3.stop);
1648
+ if (g3.stop) return await withWarnings(g3.stop);
1529
1649
 
1530
1650
  if (verdict === "pass") break; // → QA → GATE H → ship
1531
1651
  if (g3.decision === "stop" || round >= maxRounds) {
1532
- return withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
1652
+ return await withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
1533
1653
  }
1534
1654
  round += 1;
1535
1655
  }
1536
1656
 
1537
1657
  if (verdict !== "pass") {
1538
- return withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
1658
+ return await withWarnings({ status: "gate_h", breaker: "outer", hammer_proposals: allHammer, green_scopes: allGreen });
1539
1659
  }
1540
1660
 
1541
1661
  // ---- QA (post-PASS, pre-ship) — a level-up, never a gate. `--no-qa` answers it "skip". --------
1542
1662
  phase("QA");
1543
1663
  let qaFindings = 0;
1544
1664
  const qaG = await crossGate("QA", "QA", ["run", "skip", "ask"], { round, verdict });
1545
- if (qaG.stop) return withWarnings(qaG.stop);
1665
+ if (qaG.stop) return await withWarnings(qaG.stop);
1546
1666
  const qaRan = !args.noQa && qaG.decision === "run";
1547
1667
  if (qaRan) {
1548
1668
  const q = await worker({
@@ -1562,14 +1682,14 @@ const h = await worker({
1562
1682
  payload: { feature: slug, qa_findings: qaFindings, hammer_proposals: allHammer },
1563
1683
  extra: "Run the census, compare against the BASELINE and never the ideal, and produce the cut list.",
1564
1684
  });
1565
- if (h.__failed) return diedAt("H", h);
1685
+ if (h.__failed) return await withWarnings(diedAt("H", h));
1566
1686
  if (h.verdict === "cannot-ship") {
1567
- return withWarnings(aborted("H", `scope-hammer: CANNOT SHIP — ${h.cut_list.join(", ") || "a must-have failed"}`));
1687
+ return await withWarnings(aborted("H", `scope-hammer: CANNOT SHIP — ${h.cut_list.join(", ") || "a must-have failed"}`));
1568
1688
  }
1569
1689
 
1570
1690
  {
1571
1691
  const g = await crossGate("H", "Ship", ["accept-cut-list", "ship-all", "ask"], { verdict: h.verdict, cut_list: h.cut_list });
1572
- if (g.stop) return withWarnings(g.stop);
1692
+ if (g.stop) return await withWarnings(g.stop);
1573
1693
  }
1574
1694
 
1575
1695
  const ship = await cmd(`reduce ship --slug ${slug} --verdict PASS --qa ${qaRan ? "run" : "skipped"}`, "Ship", "ship-report");
@@ -1585,7 +1705,7 @@ await setRunStatus("shipped", "Ship");
1585
1705
  const ALL_DIMS = ["spec-conformance", "tdd-surface", "integration", "completeness",
1586
1706
  "test-surface-conformance", "security", "performance"];
1587
1707
 
1588
- return withWarnings({
1708
+ return await withWarnings({
1589
1709
  status: "shipped",
1590
1710
  verdict: "pass",
1591
1711
  rounds_used: round,
@@ -1643,11 +1763,15 @@ async function buildScope(scope, roundNo) {
1643
1763
  `and keep T0 green. An entry marked \`unowned\` cites no file any scope owns — fix it only ` +
1644
1764
  `if it falls inside your substrate. `
1645
1765
  : "") +
1646
- `Re-compile the order for every attempt after the first, with --attempt <n>. ` +
1766
+ `Re-compile the order for every attempt after the first, with --attempt <n>. An attempt will ` +
1767
+ `be REFUSED (exit 3) while the previous one is unanswered — dispatched with no leg row and no ` +
1768
+ `WorkResult — because grading a tree the previous attempt may still be writing counts an ` +
1769
+ `attempt nobody ran. Let it come back rather than opening the next one. ` +
1647
1770
  `Run the attempt ratchet for THIS scope only: up to ${attemptBudget} attempts of implement → ` +
1648
1771
  `\`node "${KERNEL}" verify t0 "${scope.path}" --round ${roundNo} --attempt <n>\`, each scored against ` +
1649
1772
  `the last kept trial. Stop on the first green T0, or when the attempt budget or the stagnation ` +
1650
1773
  `breaker trips. Write only inside this scope's substrate whitelist — the sandbox hook enforces it. ` +
1651
- `Report green, attempts_used, which breaker (if any) tripped, and the T0 artifact path.`,
1774
+ `Report green, attempts_used, which breaker (if any) tripped, and the T0 artifact path — ` +
1775
+ `attempts_used is what the attested channels carry, not a count of the orders you compiled.`,
1652
1776
  });
1653
1777
  }
@@ -188,7 +188,7 @@ shared vocabulary → write it to `shapeup/<slug>/shaping/glossary.md`.
188
188
  Orchestrated, this skill is dispatched like every worker: a **WorkOrder** in (`--order <path>`,
189
189
  operation `translate`), a **WorkResult** out. The standalone arguments below map 1:1 onto the
190
190
  payload fields registered for this worker in the central domain registry
191
- (`skills/tech-lead/schemas/domain.schema.json`, `x-payload-by-worker`):
191
+ (`kernel/schemas/domain.schema.json`, `x-payload-by-worker`):
192
192
 
193
193
  | Payload field | Standalone form | Meaning |
194
194
  |---|---|---|