mandrel 2.58.0 → 2.60.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (124) hide show
  1. package/.agents/README.md +17 -12
  2. package/.agents/agents/acceptance-critic.md +24 -43
  3. package/.agents/agents/story-worker.md +18 -19
  4. package/.agents/docs/SDLC.md +12 -13
  5. package/.agents/docs/agentrc-reference.json +1 -2
  6. package/.agents/docs/configuration.md +29 -46
  7. package/.agents/docs/quality-gates.md +9 -5
  8. package/.agents/docs/workflows.md +1 -1
  9. package/.agents/instructions.md +5 -7
  10. package/.agents/rules/ci-remediation.md +41 -8
  11. package/.agents/rules/known-tooling-behavior.md +65 -15
  12. package/.agents/runtime-deps.json +7 -2
  13. package/.agents/schemas/acceptance-eval-verdict.schema.json +1 -1
  14. package/.agents/schemas/agentrc.schema.json +6 -11
  15. package/.agents/schemas/crap-baseline.schema.json +1 -1
  16. package/.agents/schemas/crap-report.schema.json +1 -1
  17. package/.agents/schemas/story-deliver-terminal.schema.json +3 -3
  18. package/.agents/scripts/README.md +11 -1
  19. package/.agents/scripts/acceptance-eval.js +25 -27
  20. package/.agents/scripts/ceremony-derive.js +15 -10
  21. package/.agents/scripts/check-context-budget.js +148 -228
  22. package/.agents/scripts/check-schema-references.js +5 -3
  23. package/.agents/scripts/check-workflow-citations.js +33 -147
  24. package/.agents/scripts/coverage-capture.js +7 -4
  25. package/.agents/scripts/deliver-light.js +41 -100
  26. package/.agents/scripts/deliver-run.js +631 -0
  27. package/.agents/scripts/file-ci-gap.js +59 -11
  28. package/.agents/scripts/install-matrix-assert.js +48 -3
  29. package/.agents/scripts/lib/audit-to-stories/seed-from-findings.js +51 -33
  30. package/.agents/scripts/lib/baselines/crap-preview-incremental.js +6 -2
  31. package/.agents/scripts/lib/baselines/kinds/_crap-read.js +0 -8
  32. package/.agents/scripts/lib/baselines/kinds/crap.js +35 -18
  33. package/.agents/scripts/lib/changed-files.js +30 -0
  34. package/.agents/scripts/lib/config/delivery-routing.js +5 -4
  35. package/.agents/scripts/lib/config/explain.js +1 -3
  36. package/.agents/scripts/lib/config/gates/crap-incremental-coverage.schema.js +1 -1
  37. package/.agents/scripts/lib/config-resolver.js +1 -0
  38. package/.agents/scripts/lib/config-settings-schema-delivery.js +28 -21
  39. package/.agents/scripts/lib/coverage-capture-fullscope.js +10 -2
  40. package/.agents/scripts/lib/coverage-capture-incremental.js +3 -2
  41. package/.agents/scripts/lib/coverage-capture-usage.js +4 -1
  42. package/.agents/scripts/lib/crap-engine.js +2 -2
  43. package/.agents/scripts/lib/crap-utils.js +21 -5
  44. package/.agents/scripts/lib/doc-tiers.js +4 -2
  45. package/.agents/scripts/lib/escomplex-ast-compat.js +39 -17
  46. package/.agents/scripts/lib/escomplex-kernel.js +298 -0
  47. package/.agents/scripts/lib/feedback-loop/graduator-core.js +7 -6
  48. package/.agents/scripts/lib/feedback-loop/retro-proposals-graduator.js +7 -5
  49. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  50. package/.agents/scripts/lib/gh-exec.js +160 -0
  51. package/.agents/scripts/lib/maintainability-engine.js +3 -3
  52. package/.agents/scripts/lib/observability/source-classifier.js +1 -0
  53. package/.agents/scripts/lib/orchestration/ceremony-routing.js +74 -132
  54. package/.agents/scripts/lib/orchestration/ci-rerun-guard.js +123 -12
  55. package/.agents/scripts/lib/orchestration/complexity-gate.js +180 -352
  56. package/.agents/scripts/lib/orchestration/light-suitability.js +71 -136
  57. package/.agents/scripts/lib/orchestration/plan-context.js +44 -50
  58. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +8 -6
  59. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +104 -119
  60. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +41 -25
  61. package/.agents/scripts/lib/orchestration/plan-persist/summary.js +11 -11
  62. package/.agents/scripts/lib/orchestration/plan-persist/supersede-ops.js +63 -29
  63. package/.agents/scripts/lib/orchestration/plan-persist/wave-collision-gate.js +107 -0
  64. package/.agents/scripts/lib/orchestration/review-depth.js +14 -11
  65. package/.agents/scripts/lib/orchestration/run-epilogue.js +260 -182
  66. package/.agents/scripts/lib/orchestration/run-scoped-config.js +63 -99
  67. package/.agents/scripts/lib/orchestration/single-story-close/phases/base-sync.js +3 -3
  68. package/.agents/scripts/lib/orchestration/single-story-close/phases/graphql-preflight.js +137 -0
  69. package/.agents/scripts/lib/orchestration/single-story-close/runner.js +105 -18
  70. package/.agents/scripts/lib/orchestration/story-deliver-terminal.js +4 -3
  71. package/.agents/scripts/lib/orchestration/story-follow-ups.js +156 -39
  72. package/.agents/scripts/lib/orchestration/story-init-envelope.js +71 -0
  73. package/.agents/scripts/lib/orchestration/task-body-validator.js +8 -17
  74. package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +25 -209
  75. package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +8 -5
  76. package/.agents/scripts/lib/orchestration/ticket-validator.js +44 -183
  77. package/.agents/scripts/lib/orchestration/ticketing/reads.js +14 -25
  78. package/.agents/scripts/lib/runtime-deps/dep-resolution.js +155 -0
  79. package/.agents/scripts/lib/runtime-deps/ensure-installed.js +44 -9
  80. package/.agents/scripts/lib/runtime-deps/parser-major.js +110 -0
  81. package/.agents/scripts/lib/runtime-deps/preflight.js +6 -25
  82. package/.agents/scripts/lib/runtime-deps/scan-imports.js +46 -1
  83. package/.agents/scripts/lib/skills/walk-skill-files.js +1 -1
  84. package/.agents/scripts/lib/story-body/body-format-lints.js +58 -12
  85. package/.agents/scripts/lib/story-body/story-body.js +83 -29
  86. package/.agents/scripts/lib/templates/decomposer-prompts.js +28 -33
  87. package/.agents/scripts/lib/wave-runner/live-probe.js +31 -5
  88. package/.agents/scripts/merge-baseline.js +4 -5
  89. package/.agents/scripts/plan-context.js +117 -28
  90. package/.agents/scripts/plan-persist.js +79 -39
  91. package/.agents/scripts/plan-run-epilogue.js +11 -8
  92. package/.agents/scripts/pr-watch-with-update.js +9 -2
  93. package/.agents/scripts/run-verify.js +13 -6
  94. package/.agents/scripts/single-story-init.js +7 -57
  95. package/.agents/scripts/stories-wave-tick.js +160 -26
  96. package/.agents/skills/core/gates-and-baselines/reference.md +0 -1
  97. package/.agents/skills/skills.index.json +2 -12
  98. package/.agents/skills/stack/qa/playwright/SKILL.md +26 -0
  99. package/.agents/workflows/audit-to-stories.md +14 -11
  100. package/.agents/workflows/helpers/acceptance-self-eval.md +84 -157
  101. package/.agents/workflows/helpers/code-review.md +4 -2
  102. package/.agents/workflows/helpers/deliver-digest.md +31 -24
  103. package/.agents/workflows/helpers/deliver-light.md +92 -101
  104. package/.agents/workflows/helpers/deliver-reference.md +116 -100
  105. package/.agents/workflows/helpers/deliver-story-reference.md +58 -124
  106. package/.agents/workflows/helpers/deliver-story.md +17 -18
  107. package/.agents/workflows/helpers/plan-reference.md +82 -60
  108. package/.agents/workflows/mandrel-deliver.md +47 -31
  109. package/.agents/workflows/mandrel-plan.md +32 -30
  110. package/.agents/workflows/mandrel-update.md +36 -21
  111. package/README.md +3 -3
  112. package/docs/CHANGELOG.md +43 -0
  113. package/lib/cli/registry.js +45 -25
  114. package/lib/cli/update.js +376 -17
  115. package/lib/migrations/index.js +2 -0
  116. package/lib/migrations/steps/2.60.0-retire-audit-results-autofile.js +40 -0
  117. package/package.json +8 -2
  118. package/.agents/schemas/model-attribution.schema.json +0 -53
  119. package/.agents/scripts/lib/orchestration/model-attribution.js +0 -418
  120. package/.agents/scripts/lib/orchestration/split-policy-validator.js +0 -188
  121. package/.agents/scripts/lib/orchestration/story-plan-state.js +0 -33
  122. package/.agents/scripts/lib/orchestration/structured-comment-parser.js +0 -67
  123. package/.agents/scripts/lib/templates/spec-author-prompts.js +0 -76
  124. package/.agents/skills/core/scope-triage/SKILL.md +0 -48
@@ -1,9 +1,9 @@
1
1
  ---
2
2
  description:
3
3
  The unplanned prompt path /mandrel-deliver takes for a free-text prompt. Judges a
4
- prompt's predicted footprint, authors a receipt Story, then lands it through
5
- the same single-story-init / single-story-close engine — every close gate
6
- unchanged.
4
+ prompt's predicted footprint for risk, authors a receipt Story, then lands it
5
+ through the same single-story-init / single-story-close engine — every close
6
+ gate unchanged.
7
7
  ---
8
8
 
9
9
  # Unplanned delivery (the prompt path)
@@ -24,53 +24,45 @@ straight to execution from a prompt, landing through the unchanged close path.
24
24
  It never relaxes a close gate, never bypasses the PR to `main`, and never lands
25
25
  over-scope work silently.
26
26
 
27
- One gate, whatever the door: the suitability gate below is the decision, read
28
- against the predicted work's *effort and risk* (`STORY_SHAPE_CEILINGS`: change
29
- kinds, magnitude, uncertainty, deployable span). `/mandrel-plan` no longer
30
- suggests this path at its Gate #1 (Story #5312) — a prompt reaches it through
31
- `/mandrel-deliver`, and the gate runs there every time.
32
-
33
- ## Scope by effort, not by artifact count {#scope-by-effort}
34
-
35
- **Counting the footprint is the wrong axis.** Three identical one-line edits
36
- across three files is trivial work with a high count; a 200-line rewrite of one
37
- module is a single change. The axes are therefore effort and risk: distinct
38
- change **kinds** (N instances of one mechanical edit is one kind at N sites), a
39
- coarse **magnitude** bucket, and **uncertainty** — is the shape determined by
40
- the request, or does it still need the design decisions `/mandrel-plan` exists to
41
- resolve?
42
-
43
- Because the predicted footprint is a *declaration* — a guess, and a gameable one
44
- — this gate is deliberately **coarse**: it rejects clearly-epic work only
45
- (multiple deployables, a migration plus its consumers, an explicit
46
- multi-capability enumeration). Size is enforced where ground truth is available:
47
- the diff backstop in step 4. Do not talk yourself past that one.
48
-
49
- **The backstop counts by the same principle.** It reads magnitude — changed
50
- lines over implementation files — not artifacts, and exempts the test and doc
51
- companions the framework itself mandates. A ceiling that punishes a repo for
52
- obeying its own test-first rule is a ceiling that over-fires.
27
+ ## One gate, and it reads evidence {#one-gate}
28
+
29
+ There is exactly **one gate** at prediction time, and it asks one question:
30
+ does the predicted footprint trip an absolute **risk** rule? Two do — a path in
31
+ a registered sensitive class, and a migration paired with its consumers — and
32
+ both are derived from the paths you name, not from anything you assert about
33
+ the size of your own request.
34
+
35
+ **A size you declare about your own request is not a measurement.** The path
36
+ used to ask for four more axes (distinct change kinds, a magnitude bucket, an
37
+ uncertainty bucket, a deployable span) and judge them against framework
38
+ ceilings. Story #5313 demoted them to warnings, which meant they decided
39
+ nothing; Story #5344 deleted them. Every one was supplied by the same agent
40
+ asking to proceed, so the axis and the answer had a single author.
41
+
42
+ **Size is enforced where ground truth is available:** the diff backstop in
43
+ step 4, against the actual committed change set. `LIGHT_DIFF_CEILINGS` —
44
+ implementation lines plus a file-sprawl tripwire — is the only size block on
45
+ this path, and it reads a diff rather than a declaration.
46
+
47
+ **The backstop counts by the right principle too.** It reads magnitude —
48
+ changed lines over implementation files — not artifacts, and exempts the test
49
+ and doc companions the framework itself mandates. A ceiling that punishes a
50
+ repo for obeying its own test-first rule is a ceiling that over-fires.
53
51
 
54
52
  Sensitivity is the exception and stays absolute: a footprint touching an auth,
55
- crypto, billing, or migration class routes `full` however small or mechanical —
56
- and unlike a ceiling, it is never a warning. The shape ceilings themselves
57
- stopped gating in Story #5313: a prediction past one is carried as a
58
- `warnings[]` entry on the envelope, naming the exceeded axis, and the run
59
- proceeds — `LIGHT_DIFF_CEILINGS` in the backstop is the only size block.
53
+ crypto, billing, or migration class routes `full` however small or mechanical,
54
+ at prediction time and again at the backstop. That one was never a ceiling and
55
+ is not being loosened.
60
56
 
61
57
  ## Four invariants (do not skip one)
62
58
 
63
- 1. **Suitability gate.** The prompt's predicted footprint is judged by the
64
- shared effort/risk machinery (`deriveStoryShape` / `deriveChangeLevel`)
65
- **and** a ledgered model verdict with a recorded reason. Both must agree on
66
- `lite`.
67
- 2. **The predicted shape warns; risk escalates.** A prompt past a shape
68
- ceiling proceeds light with a `warnings[]` entry (Story #5313) — the
69
- prediction is a guess the backstop bounds for real. Only an un-ledgered
70
- verdict or an un-waivable risk rule (a sensitive-path class, a migration
71
- span) refuses, and it refuses the same way attended or not: an
72
- **`escalated` terminal envelope** that ends the session (§ Escalation is
73
- terminal). There is no question to wait for and no answer flag.
59
+ 1. **Risk gate.** The predicted footprint is judged by the shared risk
60
+ machinery (`deriveStoryShape` / `deriveChangeLevel`) **and** a ledgered
61
+ verdict carrying a recorded reason.
62
+ 2. **Only risk and the ledger refuse.** An un-waivable risk rule (a
63
+ sensitive-path class, a migration span) or an unrecorded reason emits an
64
+ **`escalated` terminal envelope** (§ Escalation ends this path), attended or
65
+ not. Nothing else refuses at prediction time.
74
66
  3. **Diff-derived backstop.** After implementation the ACTUAL change set is
75
67
  re-checked — the diff is the real scope signal — and an over-ceiling diff is
76
68
  blocked rather than landed.
@@ -80,31 +72,26 @@ proceeds — `LIGHT_DIFF_CEILINGS` in the backstop is the only size block.
80
72
 
81
73
  ## Procedure
82
74
 
83
- 1. **Predict + gate.** Form the predicted footprint (new files, edited files,
84
- acceptance count), judge its effort honestly (`--kinds` / `--magnitude` /
85
- `--uncertainty`, per § Scope by effort), and record your ledgered verdict (a
86
- recorded reason for `lite`), then run the gate — it documents every flag
87
- itself, so run it with `--help` rather than guessing:
75
+ 1. **Predict + gate.** Form the predicted footprint (new files, edited files)
76
+ and record the reason you are taking this path, then run the gate — it
77
+ documents every flag itself, so run it with `--help` rather than guessing:
88
78
 
89
79
  ```bash
90
80
  node .agents/scripts/deliver-light.js --prompt "<prompt>" \
91
- --creates <csv> --refactors <csv> --acceptance <n> \
92
- --kinds <csv> --magnitude trivial|moderate|substantial \
93
- --uncertainty determined|needs-design \
94
- --route lite --reason "<why this is trivial>" [--amends '#<id>'] [--yes]
81
+ --creates <csv> --refactors <csv> \
82
+ --reason "<why this is small>" [--amends '#<id>']
95
83
  ```
96
84
 
97
85
  Branch on `action` in the JSON envelope:
98
86
  - **`proceed-light`** — the receipt Story is authored; read `storyId` and
99
- `nextCommands`. Relay every `warnings[]` entry to the operator verbatim
100
- (each names the shape axis the prediction exceeded), then continue to
101
- step 2 — the diff backstop in step 4 is what bounds the actual change.
87
+ `nextCommands`, then continue to step 2. The diff backstop in step 4 is
88
+ what bounds the actual change.
102
89
  - **escalation** — no `action` to branch on: the gate emits an
103
90
  **`escalated` terminal envelope** instead (exit 2), attended or not.
104
- § Escalation is terminal governs; you are finished.
91
+ § Escalation ends this path governs.
105
92
 
106
- `--amends '#<id>'` is the canonical light case — shape-checked identically; a
107
- heavy amendment escalates to `/mandrel-plan` like any other over-scope prompt.
93
+ `--amends '#<id>'` is the canonical light case — judged identically; an
94
+ amendment touching a sensitive class escalates like any other prompt.
108
95
 
109
96
  2. **Init (same engine).** From the main checkout, synchronously, with the
110
97
  maximum Bash timeout:
@@ -120,7 +107,7 @@ proceeds — `LIGHT_DIFF_CEILINGS` in the backstop is the only size block.
120
107
  3. **Implement + self-eval.** `cd` into `workCwd`, implement the change, run
121
108
  the full suite once in the worktree **so close can credit it** — the
122
109
  crediting invocation and the freshness contract are
123
- [`deliver-story.md`](deliver-story.md) Step 1.3, unchanged here — then run
110
+ [`deliver-digest.md`](deliver-digest.md) § 5, unchanged here — then run
124
111
  the bounded acceptance self-eval loop
125
112
  ([`deliver-story.md`](deliver-story.md) Step 1a). Commit
126
113
  on `story-<id>` with `(refs #<storyId>)`.
@@ -132,7 +119,7 @@ proceeds — `LIGHT_DIFF_CEILINGS` in the backstop is the only size block.
132
119
  ```
133
120
 
134
121
  This is the pass that actually bounds size, which is why the prediction gate
135
- above can afford to be coarse. It measures **magnitude on the change's
122
+ above can afford to judge risk only. It measures **magnitude on the change's
136
123
  implementation half** — changed lines (additions + deletions) plus a file
137
124
  sprawl tripwire — never raw artifact count. Tests, `docs/**`, `**/*.md`,
138
125
  `baselines/**`, and lockfiles are exempt from the counts, because the
@@ -161,62 +148,66 @@ proceeds — `LIGHT_DIFF_CEILINGS` in the backstop is the only size block.
161
148
  ```
162
149
 
163
150
  Branch on the terminal envelope's `status` per
164
- [`deliver-digest.md`](deliver-digest.md) § 5 — every close
151
+ [`deliver-digest.md`](deliver-digest.md) § 6 — every close
165
152
  gate runs byte-identical to the full path.
166
153
 
167
- ## Escalation is terminal {#escalation-is-terminal}
154
+ ## Escalation ends this path {#escalation-is-terminal}
168
155
 
169
- A refused gate — an un-ledgered verdict or an un-waivable risk rule, under
170
- `--yes` or attended alike — emits a schema-validated `story-deliver-terminal`
156
+ A refused gate — an un-ledgered verdict or an un-waivable risk rule, attended
157
+ or unattended alike — emits a schema-validated `story-deliver-terminal`
171
158
  envelope with **`status: "escalated"`**, `storyId: null`, and a `nextCommand`
172
159
  naming the `/mandrel-plan` invocation that owns the work.
173
160
 
174
- **That envelope IS this session's terminal output.** Relay it and stop. There is
175
- no remaining step, no degraded fallback, and no smaller version of the work to
176
- attempt.
177
-
178
- **Invoking `/mandrel-plan` in this same session is forbidden.** Hand the operator the
179
- `nextCommand`; `/mandrel-plan` runs in a **fresh** session.
180
-
181
- This is not style — it is the empirical finding that motivated the envelope.
182
- A mandrel-bench 2.13.0 light-arm run read the escalation and continued anyway:
183
- it invoked `/mandrel-plan` in-session and delivered. The in-session plan authored **one**
184
- Story against the scenario's 3–5 contract, where a fresh `/mandrel-plan` session on the
185
- identical seed authored **four**. Planning inside a session already framed as
186
- small work under-decomposes, so walking past the escalation silently produced
187
- the very outcome the guard exists to prevent. The gate's decision was right both
188
- times; only the outcome's finality was missing.
161
+ **That envelope IS this session's terminal output for the light path.** There
162
+ is no remaining light step, no degraded fallback, and no smaller version of the
163
+ work to attempt. Relay the envelope.
189
164
 
190
165
  Nothing is left half-started: an escalated run creates **no receipt Story, no
191
166
  `story-<id>` branch, and no worktree** — the escalation path returns before
192
167
  every creation call site, and `escalation.created` records all three as `false`
193
168
  in a shape the schema pins, so a later run finds nothing to trip over.
194
169
 
195
- ## Why escalation breaks the session {#why-the-two-directions-differ}
196
-
197
- Traffic runs one way between this path and `/mandrel-plan` — light →
198
- `/mandrel-plan` on an over-scope prompt — and that direction has a deliberate
199
- session rule:
200
-
201
- > **The direction whose guard is model judgment must break the session.**
202
-
203
- **Light → `/mandrel-plan` must be a fresh session.** What is being protected is
204
- *authoring judgment*, and the empirical finding above is that a session already
205
- framed as small work under-decomposes — one Story against a 3–5 contract where
206
- a fresh session on the identical seed authored four. The frame is the hazard,
207
- so only a new session removes it. Everything on this path's own side is
208
- mechanical — `STORY_SHAPE_CEILINGS`, the ledgered verdict, the diff backstop —
209
- and none of it degrades because the context is large.
210
-
211
- Do not "fix" this by letting light → `/mandrel-plan` run in-session: that
212
- reintroduces the exact failure the `escalated` envelope exists to prevent.
170
+ ## Continuing into `/mandrel-plan` {#continuing-into-plan}
171
+
172
+ You may run the `nextCommand` **in this same session**, and when you do, seed
173
+ it deliberately:
174
+
175
+ - hand `/mandrel-plan` the **original prompt** and the envelope's
176
+ `escalation.reasons` verbatim — the reasons name the risk class or the
177
+ missing ledger, which is planning input, not noise;
178
+ - state plainly that the light path refused it, so the planning pass starts
179
+ from "this is not small" rather than from the framing that produced the
180
+ prompt.
181
+
182
+ **A fresh session is still the safer default** when the work is clearly larger
183
+ than the prompt admitted, or when the reasons name a sensitive class you had
184
+ not considered: the cost is one boot, and it removes the frame entirely.
185
+
186
+ ### Why in-session planning was barred, and why it is not any more
187
+
188
+ Story #4746 required the *session* to end here, on one mandrel-bench 2.13.0
189
+ light-arm observation: a run read the escalation, invoked `/mandrel-plan`
190
+ in-session, and the in-session plan authored **one** Story against the
191
+ scenario's 3–5 contract, where a fresh `/mandrel-plan` session on the identical
192
+ seed authored **four**. The reading was that a session already framed as small
193
+ work under-decomposes.
194
+
195
+ That is one observation on one scenario, and the rule it bought cost every
196
+ escalated prompt a full cold boot — the exact session multiplication this path
197
+ exists to remove. Story #5344 keeps the finding and drops the ban: the
198
+ under-decomposition risk is real, so the seeding instructions above exist to
199
+ counter the frame directly, and the **light-arm cell of mandrel-bench** is the
200
+ measurement that decides whether the loosening stays. If that cell's
201
+ decomposition counts regress against the fresh-session arm, the ban comes back
202
+ and this section is the record of why. See
203
+ [`docs/decisions.md`](../../../docs/decisions.md) ADR `20260917-5344`.
213
204
 
214
205
  ## See also
215
206
 
216
207
  - [`/mandrel-deliver`](../mandrel-deliver.md) — the delivery entry point; routes here on a
217
208
  free-text prompt.
218
- - [`/mandrel-plan`](../mandrel-plan.md) — owns the work an over-scope prompt
219
- escalates to.
209
+ - [`/mandrel-plan`](../mandrel-plan.md) — owns the work an escalated prompt
210
+ goes to next.
220
211
  - [`deliver-story.md`](deliver-story.md) — the one Story delivery engine every
221
212
  path shares.
222
213
  - [`deliver-digest.md`](deliver-digest.md) — engine invariants, gates, and the
@@ -20,7 +20,7 @@ means exactly the five ids in it.
20
20
 
21
21
  **Pass the span through; never expand it by hand.** Every id-list flag on the
22
22
  delivery path takes range tokens — `resolve-stories.js --ids`,
23
- `stories-wave-tick.js --stories` and `--dispatched`, and
23
+ `deliver-run.js --stories` and `--handoff`, and
24
24
  `plan-run-epilogue.js --stories`. Normalize the operator's spacing away and hand
25
25
  the scripts one unspaced token (`--ids 4922-4926`), mixed freely with singles
26
26
  and commas (`--ids 4901,4922-4926`); overlaps dedupe. A hand-typed enumeration
@@ -40,7 +40,7 @@ The cap is per range token, not per run: a genuine 60-Story delivery is still
40
40
  expressible as two ranges, but a slipped digit cannot fan out into a live
41
41
  resolution sweep of thousands of issues.
42
42
 
43
- ## Sequencing edge cases (`stories-wave-tick.js`)
43
+ ## Sequencing edge cases (`deliver-run.js` over `stories-wave-tick.js`)
44
44
 
45
45
  **What "discovered, not declared" means concretely.** `resolve-stories.js` reads
46
46
  the graph from live state as the union of the Story bodies' `depends_on` edges
@@ -56,6 +56,21 @@ worker at an unenriched body, after taking its lease. Route it through
56
56
  `/mandrel-plan` first, which applies `agent::ready` at the end of planning.
57
57
  `--allow-unlabelled` is the deliberate escape hatch.
58
58
 
59
+ **Inspecting a beat without taking one.** `deliver-run.js` writes — the ledger,
60
+ the prompts — so it is not the tool for "what *would* the next beat do?". The
61
+ tick underneath it is read-only and answers exactly that, with no ledger and no
62
+ prompt files:
63
+
64
+ ```bash
65
+ node .agents/scripts/stories-wave-tick.js --stories <id,id,...> --probe-live
66
+ ```
67
+
68
+ Its envelope is the beat's scheduling half verbatim — `ready`, `inFlight`,
69
+ `inFlightReservation`, `footprintGuard`, `foreignHeld`, `wedged`. Read it when
70
+ a slot is unfilled and you want to know why before acting. Do **not** drive a
71
+ run from it: without the ledger the init window reopens, and the same Story is
72
+ handed out twice.
73
+
59
74
  **The non-zero exit codes.** **2** — `cycleError`: the graph is
60
75
  self-referential; fix `depends_on`, do not retry. **3** — `wedged`: nothing
61
76
  dispatchable and nothing in flight, with the undone Stories and their unmet
@@ -77,22 +92,45 @@ another run), and derives **in-flight** from live `agent::executing` /
77
92
  `agent::closing` labels. You never compute `done` or `in-flight` — that
78
93
  accounting is read from reality every beat.
79
94
 
80
- **`--dispatched` is the one thing you must tell it.** List every
81
- Story id you have spawned this run. Live state cannot instantly report a Story
82
- you dispatched moments ago: `single-story-init.js` publishes `agent::executing`
83
- before the worktree install (ahead of the multi-minute install, so the
84
- window is short rather than minutes-long), but it is not
85
- zero — until the label lands the Story still reads `agent::ready` and, without
86
- `--dispatched`, the next beat would hand it back and a second sub-agent would
87
- join the first on the same branch and worktree, interleaving commits.
88
- `--dispatched` closes that residual same-run window. The rule is
89
- **append-only: add each id as you dispatch it and never remove one.** The flag
90
- is additive, not authoritative — the probe unions it into the label-derived set
91
- and then filters it against live state, so an id that has since gone
92
- `agent::done` is dropped for you. Re-listing an id costs nothing and cannot
93
- double-count a slot; *omitting* one is the only way to get this wrong. This is
94
- why `--dispatched` is not the retired `--done` bookkeeping, and why
95
- `--in-flight` remains rejected under `--probe-live`.
95
+ **The run ledger closes the init window, and you maintain nothing.** Live state
96
+ cannot instantly report a Story dispatched moments ago:
97
+ `single-story-init.js` publishes `agent::executing` before the worktree install
98
+ (ahead of the multi-minute install, so the window is short rather than
99
+ minutes-long), but it is not zero — until the label lands the Story still reads
100
+ `agent::ready`, and an unaugmented beat would hand it back so a second
101
+ sub-agent joined the first on the same branch and worktree, interleaving
102
+ commits. `deliver-run.js` closes that window from its own ledger
103
+ (`<tempRoot>/run-<id>/ledger.json`): every id it hands out as ready is
104
+ recorded, and the next beat reads the file back and seeds the tick with it.
105
+ Append-only by construction, so the "forgot to re-list one" failure the
106
+ hand-maintained list had cannot occur. The ledger is additive, not
107
+ authoritative — the probe unions it into the label-derived set and then filters
108
+ it against live state, so an id that has since gone `agent::done` is dropped
109
+ for you. A missing or corrupt ledger costs one extra beat of the init window,
110
+ never the run. The run id is a stable digest of the Story id set, so every beat
111
+ of one run finds the same ledger and two concurrent runs never share one;
112
+ `--run-id` pins it explicitly.
113
+
114
+ **A spawn that never reached init is the ledger's one sharp edge.** Append-only
115
+ is what makes the list safe to keep, and it is also why a bad entry never
116
+ leaves: if a spawn dies before `single-story-init.js` runs, the id stays
117
+ ledgered, is withheld as in flight on every later beat, and the run returns an
118
+ empty `ready[]` with a non-zero `inFlight` forever — which reads exactly like a
119
+ healthy wait. The beat names it instead: every ledgered id live state still
120
+ reports as `agent::ready` appears in `stalledDispatch[]`, with the recovery in
121
+ `stalledDispatchReason`. That is its **own** reason — not a footprint withhold
122
+ (which names a blocking peer and the colliding paths) and not a foreign lease
123
+ (which names a holder and clears itself when their run ends).
124
+
125
+ It is a report, not a release. A slow init and a dead spawn are the same
126
+ observation at the beat's altitude, and auto-releasing would re-dispatch a live
127
+ Story onto its own branch — the failure the ledger exists to prevent. So the
128
+ operator owns the call: confirm no worker is running, remove the id from
129
+ `dispatched` in `<tempRoot>/run-<id>/ledger.json` (deleting the file works too,
130
+ at the cost of reopening the init window for the rest), and beat again with
131
+ `--run-id <id>` so the same run directory is reused. An id that has since
132
+ picked up `agent::executing`, `agent::closing` or `agent::done` is never
133
+ reported here — live state has moved on and the ledger entry is already inert.
96
134
 
97
135
  **Cross-run de-confliction is automatic.** A Story another
98
136
  operator is delivering is withheld without any bookkeeping from you: the probe
@@ -122,21 +160,12 @@ report is `available: false` and selection de-conflicts within the beat only.
122
160
  advisory, note }`. They used to be an unreported skip, so a Story simply
123
161
  vanished from `ready[]` and an unfilled slot read exactly like a cap that was
124
162
  never reached. Every entry in **either** report carries the colliding `paths`
125
- and a `source` tag:
126
-
127
- - `declared-overlap` — both Stories' `changes[]` named the path (or a declared
128
- glob). Intended serialization; two Stories rewriting the same generated
129
- baseline must not co-dispatch.
130
- - `scraped-overlap` — only the text evidence produced it. Real signal — a
131
- declaration is only a lower bound — but the class where a false positive is
132
- possible.
133
-
134
- **The evidence scrape excludes exactly three token sources**, each structurally
135
- incapable of naming an edit target: `audit-fingerprints` /
136
- `audit-semantic-keys` provenance footers, paths under `project.paths.tempRoot`,
137
- and markdown-link URL interiors. Nothing else is stripped — a
138
- `<!-- DECOMPOSITION -->` block's paths are genuine intent
139
- ([`instructions.md` § 7](../../instructions.md)) and still count.
163
+ and one `source` tag, `declared-overlap`: both Stories' `changes[]` named the
164
+ path, or one declared a glob. Intended serialization — two Stories rewriting
165
+ the same generated baseline must not co-dispatch. A declared footprint is the
166
+ whole footprint: Story #5313 retired the body scrape that used to widen it,
167
+ and the second source class it produced, so `changes[]` is the only evidence
168
+ a collision is scored against.
140
169
 
141
170
  **`delivery.deliverRunner.footprintGuard`** selects what a collision does:
142
171
 
@@ -163,23 +192,23 @@ anything, read the Story's `dispatchMode` from the resolver envelope
163
192
  `lib/orchestration/complexity-gate.js`, which decides on the resolved set size
164
193
  alone — it does not read the Story body). A Story with `dispatchMode: "inline"`
165
194
  executes [`deliver-story.md`](deliver-story.md) **inline in this session** — no
166
- `story-worker` sub-agent boot and no fresh acceptance-critic sub-agents
167
- (sub-agent boots are the dominant deliver-phase token cost at trivial scope) —
168
- threading the same `docsDigestPath` / `checklistPath` / change-set discipline
169
- as a spawned worker. Inline removes model-side fan-out only: every
170
- `single-story-close.js` gate, the PR to `main`, and the terminal envelope are
171
- identical.
195
+ `story-worker` sub-agent boot (sub-agent boots are the dominant deliver-phase
196
+ token cost at trivial scope) — threading the same `docsDigestPath` /
197
+ `checklistPath` / change-set discipline as a spawned worker. It does not touch
198
+ the acceptance verdict owner, which the ceremony profile alone names
199
+ ([`deliver-digest.md`](deliver-digest.md) § 3). Every `single-story-close.js`
200
+ gate, the PR to `main`, and the terminal envelope are identical.
172
201
 
173
202
  **A trivial shape does not buy that session.** Only the
174
203
  one-Story rule above yields `inline`; every Story of a multi-Story run comes
175
204
  back `subagent` however lite its body, because the ready set below may offer
176
205
  several Stories on one beat and a session cannot be split between them. The
177
- Story's derived shape is still reported (it sets ceremony, and a sensitive
178
- footprint keeps the fresh acceptance critic), and the `route::lite` label
179
- remains a human-visible hint only, never the control signal — a lost or
180
- never-written label cannot misroute delivery.
206
+ Story's shape does not enter the decision at all — `resolveStoryDispatchMode`
207
+ reads the resolved set size and nothing else, and the `route::lite` hint label
208
+ was retired with the plan-side route claim (Story #5312), so there is no label
209
+ left to lose or misread.
181
210
 
182
- **Issue a beat's spawns in one turn.** A wave tick hands you a ready set, not a
211
+ **Issue a beat's spawns in one turn.** A beat hands you a ready set, not a
183
212
  queue: those Stories have no dependency edge between them (the resolver already
184
213
  withheld any that do) and no shared write paths (each owns its own worktree and
185
214
  branch). Dispatch them the way
@@ -200,43 +229,25 @@ exposes agent dispatch, spawn each ready Story as its own
200
229
  sub-agent executes [`deliver-story.md`](deliver-story.md) Steps 0–2.5
201
230
  (init → implement → acceptance self-eval → **push**) and stops there; **you**
202
231
  own Step 3, serialized — see `/mandrel-deliver` § Closing what the workers hand back.
203
- Thread into its prompt: `storyId`; `docsDigestPath` (the per-run docs digest, null when
204
- `project.docsContextFiles` is unset); `checklistPath` (the footprint-matched
205
- write-time audit checklist, produced at dispatch, below); and the
206
- **change-set discipline** — the worker computes the change set once with
207
- `computeChangeSet` and hands that one list to every acceptance critic; it
208
- never lets a critic re-derive the diff.
209
-
210
- **Produce `checklistPath` before the spawn.** Compute the payload
211
- from the Story's predicted footprint (its `changes[]` / `references[]` path
212
- entries) with `buildDispatchChecklist` and write it to the run temp dir, then
213
- thread the resulting path (empty when nothing matched):
214
232
 
215
- ```bash
216
- node --input-type=module -e '
217
- import { buildDispatchChecklist } from "<main-repo>/.agents/scripts/lib/audit-suite/index.js";
218
- import { parse } from "<main-repo>/.agents/scripts/lib/story-body/story-body.js";
219
- // storyBody is the fetched Story issue body.
220
- const { changes, references } = parse(process.env.STORY_BODY);
221
- const { checklistPath } = buildDispatchChecklist({
222
- storyId: <storyId>, changes, references, runTempDir: "temp/run-<id>",
223
- });
224
- console.log(checklistPath ?? "");
225
- '
226
- ```
227
-
228
- `buildDispatchChecklist` (`lib/audit-suite/dispatch-checklist.js`) is a pure
229
- function of the footprint and the on-disk checklists; an empty match prints
230
- nothing and the worker runs with no write-time checklist — the maker-blind
231
- close-scope pass still covers it.
233
+ **The beat writes the prompt; you pass the file.** Each `ready[]` entry carries
234
+ a `promptPath` under `<tempRoot>/run-<id>/`, and that file is the whole spawn
235
+ payload: the Story id, the `workCwd` conventions, the `docsDigestPath` (null
236
+ when `project.docsContextFiles` is unset), the `checklistPath` — the
237
+ footprint-matched write-time audit checklist built from the Story's declared
238
+ `changes[]` / `references[]`, empty when nothing matched — and the
239
+ **change-set discipline** (derive the change set once with `ceremony-derive.js`
240
+ and hand that one list to the verdict owner; never let a critic re-derive the
241
+ diff). An unmatched checklist costs nothing: the maker-blind close-scope pass
242
+ still covers the Story.
232
243
 
233
244
  **Inline fallback (`roleScopedAgents: false` / no-nesting harness).** When the
234
245
  kill-switch is off, or the host cannot spawn a sub-agent at this nesting depth,
235
246
  do **not** stall: read [`deliver-story.md`](deliver-story.md) **in full** and
236
- execute it directly, in this turn, threading the same `docsDigestPath` /
237
- `checklistPath` / change-set discipline. Under `--yes` / injected helper
238
- content, execute directly without a re-read turn. The engine, gates, and
239
- terminal envelope are identical either way — only the isolation differs.
247
+ execute it directly, in this turn, following that same dispatch prompt. Under
248
+ `--yes` / injected helper content, execute directly without a re-read turn. The
249
+ engine, gates, and terminal envelope are identical either way — only the
250
+ isolation differs.
240
251
 
241
252
  ## Intent phrases (what replaced the flag table)
242
253
 
@@ -292,18 +303,22 @@ node .agents/scripts/plan-run-epilogue.js --stories 101,102
292
303
 
293
304
  This executes, in order:
294
305
 
295
- - `audit-roster` — selects cross-Story audit lenses over the combined landed
296
- tip and posts `plan-run-audit-roster` on the primary Story; the host MUST
297
- walk each listed lens against the combined diff.
298
306
  - `follow-up-rollup` — friction follow-ups across every Story in the run
299
307
  (files issues when auto-file is on; posts `follow-ups`).
300
- - `sibling-coherence` — Spec/Acceptance coherence check across sibling bodies
301
- (`plan-run-sibling-coherence`).
302
308
  - `epic-close` — **reports** which of the run's container Epics its land tails
303
309
  left closed and which are still open. **Read-only** — it derives nothing:
304
310
  every child state change is already a rollup edge, so the container was
305
311
  derived from a complete child set by the last Story's own land tail.
306
312
 
313
+ **The audit roster is opt-in** (Story #5343). Add `--audit-roster` and the run
314
+ also selects cross-Story audit lenses over the combined landed tip and posts
315
+ `plan-run-audit-roster` on the primary Story — and the host MUST then walk
316
+ every listed lens against the combined diff, one `auditor` sub-agent per lens.
317
+ That walk is the expensive half, and it only pays for itself when someone is
318
+ going to read it, so **the operator asks for it** — exactly as for the
319
+ pre-mortem plan critic. Without the flag no roster comment is posted and no
320
+ auditor is spawned.
321
+
307
322
  A single-Story run skips the epilogue — follow-ups are captured on merge
308
323
  confirm instead (`captureStoryFollowUps`).
309
324
 
@@ -353,22 +368,25 @@ strands its container open above finished work.
353
368
  ## Ceremony (profiles + two scopes)
354
369
 
355
370
  Ceremony depth is selected by `delivery.routing.ceremonyProfile`
356
- (`minimal` | `standard` | `strict`, default `standard`) and the **change level
357
- derived from the Story's own diff** — the changed files' intersection with the
358
- sensitive-path classes in `audit-rules.json`
359
- (`review-depth.js#deriveChangeLevel`), not a planner-authored verdict:
360
-
361
- | Profile | Acceptance critic | When to use |
371
+ (`minimal` | `standard` | `strict`, default `standard`) — and by that alone
372
+ since Story #5343. The **change level derived from the Story's own diff** (the
373
+ changed files' intersection with the sensitive-path classes in
374
+ `audit-rules.json`, `review-depth.js#deriveChangeLevel`, never a
375
+ planner-authored verdict) still selects **review depth**, which is the second
376
+ row of the scope table below:
377
+
378
+ | Profile | Acceptance verdict owner | When to use |
362
379
  | --- | --- | --- |
363
- | `minimal` | Always inline | Tiny trusted N=1 Stories |
364
- | `standard` | Derived-level routed | Default |
365
- | `strict` | Always fresh-context | High-assurance / regulated surfaces |
380
+ | `minimal` | Inline self-eval | Tiny trusted N=1 Stories |
381
+ | `standard` | Inline self-eval | Default |
382
+ | `strict` | Fresh-context critic | High-assurance / regulated surfaces |
366
383
 
367
384
  | Scope | What runs | Mechanism |
368
385
  | --- | --- | --- |
369
386
  | **Per-Story (always)** | Gates, branch discipline, close-and-land | `deliver-story` / `single-story-close` |
370
- | **Per-Story (profile + derived level)** | Acceptance critic mode; review depth | `ceremony-routing.js` + `review-depth.js` + `code-review.js` |
371
- | **Per-run (N>1)** | Audit roster · follow-up roll-up · sibling coherence | `plan-run-epilogue.js` once at run end |
387
+ | **Per-Story (profile)** | Acceptance verdict owner (the rule: digest § 3) | `ceremony-routing.js` |
388
+ | **Per-Story (derived level)** | Review depth | `review-depth.js` + `code-review.js` |
389
+ | **Per-run (N>1)** | Follow-up roll-up · container-Epic report (· audit roster on `--audit-roster`) | `plan-run-epilogue.js` once at run end |
372
390
  | **Per-Story land tail** | Follow-up capture · status resync · Epic rollup · ref cleanup · base fast-forward | `single-story-close/phases/post-land.js` (in-process, per-step reported) |
373
391
 
374
392
  ## Async merge-confirm mode (`delivery.mergeWatch.mode: "async"`)
@@ -383,17 +401,15 @@ agent) and move on to the next Story; `single-story-confirm-merge.js` is
383
401
  idempotent and owns the whole tail. Do not foreground-poll the merge. The
384
402
  default `"sync"` behaviour is unchanged.
385
403
 
386
- **On a multi-Story run, pass `--merge-watch-mode async` on every close.** Close
404
+ **On a multi-Story run the beat adds `--merge-watch-mode async` for you.** Close
387
405
  sees one Story and cannot see run topology, so it cannot make this call for
388
- itself — you can. Implementation runs in parallel but the close tail is
406
+ itself — `deliver-run.js` can, and does: every `close[]` command it renders
407
+ carries the flag when the run holds more than one Story and omits it for a run
408
+ of one. Run the command it printed verbatim rather than composing your own.
409
+ The reason it matters: implementation runs in parallel but the close tail is
389
410
  serialized one at a time, and under `sync` each of those closes holds the
390
411
  foreground for its full merge wait before the next Story's close may start.
391
- That is the run's dominant serialized cost, and it is paid per sibling:
392
-
393
- ```bash
394
- node <main-repo>/.agents/scripts/single-story-close.js \
395
- --story <storyId> --cwd <main-repo> --merge-watch-mode async
396
- ```
412
+ That is the run's dominant serialized cost, and it is paid per sibling.
397
413
 
398
414
  The flag overrides `delivery.mergeWatch.mode` for that invocation only — the
399
415
  config default stays `"sync"`, which is right for the solo delivery that has no