mandrel 2.59.0 → 2.60.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (97) hide show
  1. package/.agents/README.md +11 -9
  2. package/.agents/agents/acceptance-critic.md +24 -43
  3. package/.agents/agents/story-worker.md +18 -19
  4. package/.agents/docs/SDLC.md +6 -6
  5. package/.agents/docs/agentrc-reference.json +1 -2
  6. package/.agents/docs/configuration.md +29 -46
  7. package/.agents/docs/quality-gates.md +8 -4
  8. package/.agents/docs/workflows.md +1 -1
  9. package/.agents/instructions.md +4 -5
  10. package/.agents/rules/ci-remediation.md +41 -8
  11. package/.agents/rules/known-tooling-behavior.md +65 -15
  12. package/.agents/schemas/acceptance-eval-verdict.schema.json +1 -1
  13. package/.agents/schemas/agentrc.schema.json +6 -11
  14. package/.agents/schemas/story-deliver-terminal.schema.json +3 -3
  15. package/.agents/scripts/README.md +11 -1
  16. package/.agents/scripts/acceptance-eval.js +25 -27
  17. package/.agents/scripts/ceremony-derive.js +15 -10
  18. package/.agents/scripts/check-context-budget.js +148 -228
  19. package/.agents/scripts/check-schema-references.js +5 -3
  20. package/.agents/scripts/check-workflow-citations.js +33 -147
  21. package/.agents/scripts/coverage-capture.js +7 -4
  22. package/.agents/scripts/deliver-light.js +41 -100
  23. package/.agents/scripts/deliver-run.js +631 -0
  24. package/.agents/scripts/file-ci-gap.js +59 -11
  25. package/.agents/scripts/lib/baselines/crap-preview-incremental.js +6 -2
  26. package/.agents/scripts/lib/changed-files.js +30 -0
  27. package/.agents/scripts/lib/config/delivery-routing.js +5 -4
  28. package/.agents/scripts/lib/config/explain.js +1 -3
  29. package/.agents/scripts/lib/config/gates/crap-incremental-coverage.schema.js +1 -1
  30. package/.agents/scripts/lib/config-resolver.js +1 -0
  31. package/.agents/scripts/lib/config-settings-schema-delivery.js +28 -21
  32. package/.agents/scripts/lib/coverage-capture-fullscope.js +10 -2
  33. package/.agents/scripts/lib/coverage-capture-incremental.js +3 -2
  34. package/.agents/scripts/lib/coverage-capture-usage.js +4 -1
  35. package/.agents/scripts/lib/doc-tiers.js +4 -2
  36. package/.agents/scripts/lib/feedback-loop/graduator-core.js +7 -6
  37. package/.agents/scripts/lib/feedback-loop/retro-proposals-graduator.js +7 -5
  38. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  39. package/.agents/scripts/lib/gh-exec.js +160 -0
  40. package/.agents/scripts/lib/observability/source-classifier.js +1 -0
  41. package/.agents/scripts/lib/orchestration/ceremony-routing.js +74 -132
  42. package/.agents/scripts/lib/orchestration/ci-rerun-guard.js +123 -12
  43. package/.agents/scripts/lib/orchestration/complexity-gate.js +180 -352
  44. package/.agents/scripts/lib/orchestration/light-suitability.js +71 -136
  45. package/.agents/scripts/lib/orchestration/plan-context.js +13 -25
  46. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +8 -6
  47. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +76 -95
  48. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +35 -18
  49. package/.agents/scripts/lib/orchestration/plan-persist/summary.js +11 -11
  50. package/.agents/scripts/lib/orchestration/plan-persist/supersede-ops.js +63 -29
  51. package/.agents/scripts/lib/orchestration/review-depth.js +14 -11
  52. package/.agents/scripts/lib/orchestration/run-epilogue.js +260 -182
  53. package/.agents/scripts/lib/orchestration/run-scoped-config.js +63 -99
  54. package/.agents/scripts/lib/orchestration/single-story-close/phases/base-sync.js +3 -3
  55. package/.agents/scripts/lib/orchestration/single-story-close/phases/graphql-preflight.js +137 -0
  56. package/.agents/scripts/lib/orchestration/single-story-close/runner.js +105 -18
  57. package/.agents/scripts/lib/orchestration/story-deliver-terminal.js +4 -3
  58. package/.agents/scripts/lib/orchestration/story-follow-ups.js +156 -39
  59. package/.agents/scripts/lib/orchestration/story-init-envelope.js +71 -0
  60. package/.agents/scripts/lib/orchestration/task-body-validator.js +8 -17
  61. package/.agents/scripts/lib/orchestration/ticket-validator.js +44 -183
  62. package/.agents/scripts/lib/orchestration/ticketing/reads.js +14 -25
  63. package/.agents/scripts/lib/story-body/body-format-lints.js +58 -12
  64. package/.agents/scripts/lib/story-body/story-body.js +83 -29
  65. package/.agents/scripts/lib/templates/decomposer-prompts.js +7 -15
  66. package/.agents/scripts/lib/wave-runner/live-probe.js +31 -5
  67. package/.agents/scripts/merge-baseline.js +4 -5
  68. package/.agents/scripts/plan-context.js +117 -28
  69. package/.agents/scripts/plan-persist.js +79 -28
  70. package/.agents/scripts/plan-run-epilogue.js +11 -8
  71. package/.agents/scripts/pr-watch-with-update.js +9 -2
  72. package/.agents/scripts/run-verify.js +13 -6
  73. package/.agents/scripts/single-story-init.js +7 -57
  74. package/.agents/scripts/stories-wave-tick.js +160 -26
  75. package/.agents/skills/core/gates-and-baselines/reference.md +0 -1
  76. package/.agents/skills/skills.index.json +2 -2
  77. package/.agents/skills/stack/qa/playwright/SKILL.md +26 -0
  78. package/.agents/workflows/helpers/acceptance-self-eval.md +84 -157
  79. package/.agents/workflows/helpers/code-review.md +4 -2
  80. package/.agents/workflows/helpers/deliver-digest.md +31 -24
  81. package/.agents/workflows/helpers/deliver-light.md +92 -101
  82. package/.agents/workflows/helpers/deliver-reference.md +116 -100
  83. package/.agents/workflows/helpers/deliver-story-reference.md +58 -124
  84. package/.agents/workflows/helpers/deliver-story.md +17 -18
  85. package/.agents/workflows/helpers/plan-reference.md +65 -54
  86. package/.agents/workflows/mandrel-deliver.md +47 -31
  87. package/.agents/workflows/mandrel-plan.md +22 -21
  88. package/.agents/workflows/mandrel-update.md +36 -21
  89. package/docs/CHANGELOG.md +35 -0
  90. package/lib/cli/update.js +376 -17
  91. package/lib/migrations/index.js +2 -0
  92. package/lib/migrations/steps/2.60.0-retire-audit-results-autofile.js +40 -0
  93. package/package.json +2 -1
  94. package/.agents/schemas/model-attribution.schema.json +0 -53
  95. package/.agents/scripts/lib/orchestration/model-attribution.js +0 -418
  96. package/.agents/scripts/lib/orchestration/story-plan-state.js +0 -33
  97. package/.agents/scripts/lib/orchestration/structured-comment-parser.js +0 -67
@@ -112,7 +112,7 @@ not a substitute for prefixing paths correctly.
112
112
 
113
113
  ---
114
114
 
115
- ## Engine invariants and the lite route
115
+ ## Engine invariants and ceremony
116
116
 
117
117
  **Prerequisites before Step 0.** A `type::story` issue, a clean
118
118
  `gh auth status`, and `project.baseBranch` present both locally and on
@@ -128,27 +128,26 @@ The v2 engine's trait table:
128
128
  | Branch | `story-<id>` seeded from `project.baseBranch` (`main`) |
129
129
  | Merge target | `main` via PR (squash + required checks) |
130
130
  | Spec / slices | Folded `## Spec` + optional `## Slicing` checkpoints in-session |
131
- | Ceremony | Per-Story, routed off the derived change level via `ceremony-routing.js` |
132
-
133
- **Ceremony-lite Stories still land through this engine unchanged.** A
134
- lite-routed Story collapses only the _advisory_ plan/mandrel-deliver
135
- ceremony — the fresh-critic / Tech-Spec authoring a one-artifact scope does
136
- not earn. It does **not** get a cheaper landing: the close-validation gates
137
- (lint / test / format / coverage / CRAP / maintainability), the PR to `main`,
138
- and the `rules/security-baseline.md` MUSTs all run exactly as for a
139
- full-ceremony Story. The lite route's `preserves` field is the machine-readable
140
- record of those non-negotiables; there is no lite-specific gate bypass.
141
-
142
- **Ceremony comes from the landed diff; the dispatch mode comes from the
143
- run.** Persist stamps no route label (Story #5312 retired the plan-side lite
144
- claim with its `route::lite` hint). Ceremony is resolved from the **derived
145
- change level** (`deriveChangeLevel` over the computed change set — digest § 3),
146
- not from a body-shape read: a footprint intersecting a sensitive-path class
147
- derives `high`, so the Story keeps its fresh acceptance critic. The light path
148
- is the one caller that reads the authored body's shape, through
149
- `deriveStoryShape` (`lib/orchestration/complexity-gate.js`).
150
-
151
- That derived level sets ceremony. It does **not** set the dispatch mode,
131
+ | Ceremony | Per-Story, resolved from the ceremony profile alone via `ceremony-routing.js` (digest § 3) |
132
+
133
+ **A cheap shape never buys a cheaper landing.** A small Story collapses only
134
+ the _advisory_ ceremony — the fresh-critic / Tech-Spec authoring a
135
+ one-artifact scope does not earn. The close-validation gates (lint / test /
136
+ format / coverage / CRAP / maintainability), the PR to `main`, and the
137
+ `rules/security-baseline.md` MUSTs run exactly as for any other Story. There
138
+ is no gate bypass to opt into, and nothing in the ticket can declare one:
139
+ Story #5312 retired the plan-side lite claim, its `route::lite` hint and the
140
+ machine-readable field that used to enumerate the non-negotiables.
141
+
142
+ **The ceremony rule has one home: digest § 3.** The profile alone names the
143
+ verdict owner; the derived change level (`deriveChangeLevel` over the computed
144
+ change set) sets **review depth**, so a footprint intersecting a sensitive-path
145
+ class buys a deep review rather than a fresh acceptance critic. Persist stamps
146
+ no route label. The light path is the one caller that reads the authored body's
147
+ shape, through `deriveStoryShape`
148
+ (`lib/orchestration/complexity-gate.js`).
149
+
150
+ The derived level does **not** set the dispatch mode either,
152
151
  because `inline` names one indivisible resource — the router's own session —
153
152
  and only run topology can say whether it is free: a **single-Story run**
154
153
  executes inline, and every Story of a multi-Story run dispatches as a
@@ -235,46 +234,9 @@ before it spawns the worker — see [`/mandrel-deliver`](../mandrel-deliver.md).
235
234
  drift-guard and schema tests living outside the Story's scoped greps — are
236
235
  the failure class that actually bounces deliveries: close-validation
237
236
  discovers them only after the whole close pipeline has run, at several times
238
- the cost of one full-suite run in the worktree.
239
-
240
- **Run it once, last, through the depositor.** The run belongs **after** the
241
- self-eval loop's last fix commit; redraft rounds run scoped tests. Run it in
242
- the worktree on `story-<id>` as
243
-
244
- ```bash
245
- node <main-repo>/.agents/scripts/evidence-gate.js \
246
- --standalone --scope-id <storyId> --gate test \
247
- --worktree <workCwd> -- npm test
248
- ```
249
-
250
- The wrapper spawns the project's own `npm test` — whatever that resolves to —
251
- and records the pass into the Story evidence keyspace, so the credit is
252
- runner-agnostic by construction: it stamps only what it just ran. The record
253
- is keyed on HEAD and the tree fingerprint and hashed on the exact command
254
- close spawns, so close reports the gate as credited at unchanged HEAD instead
255
- of re-running the suite. The credit expires the moment it stops describing the
256
- tree: any later commit invalidates it and close re-runs the suite for real, so
257
- this never trades away the gate. The CRAP gate still runs
258
- `coverage-capture.js` itself when it needs a fresh artifact — the capture
259
- stamp is a claim about `coverage/coverage-final.json`, which `npm test` alone
260
- does not produce.
261
-
262
- **A bare `npm test` is a bonus, not the contract.** It deposits the same
263
- record only where the project's `test` script routes through mandrel's own
264
- runner (`run-tests.js` → `lib/test-run-credit.js`, Story #5313), which prints
265
- the outcome. A project whose `npm test` is `vitest run`, `jest` or any other
266
- runner never reaches that code, so it prints nothing and deposits nothing —
267
- silence is not a signal, and nothing here asks you to confirm the credit by
268
- reading for a line that cannot appear. `mandrel doctor`'s `test-credit-path`
269
- check reports which shape a project is and names the command above as its
270
- remedy.
271
-
272
- **`verify[]` reuses the same credit.** A `verify[]` entry that is itself a
273
- full-suite command is reported **credited** against that record rather than
274
- respawned (`resolveVerifyCredit` in
275
- [`verify-credit.js`](../../scripts/lib/orchestration/verify-credit.js)), and the
276
- self-eval gate warns when it sees one: the intended shape is scoped `verify[]`
277
- entries **plus** the single credited run.
237
+ the cost of one full-suite run in the worktree. The run itself — when, how
238
+ and what it credits — is stated once, in
239
+ [`deliver-digest.md`](deliver-digest.md) § 5.
278
240
 
279
241
  **Conflict with `main` mid-implementation** → resolve as you would any branch
280
242
  rebase. There is no `epic/<id>` intermediate, so the rebase base is `main`
@@ -282,53 +244,33 @@ directly.
282
244
 
283
245
  ### Step 1a — self-eval mechanics
284
246
 
285
- **One verdict-owner per cluster.** The ceremony routing's
286
- resolved decision names each cluster's single verdict owner
247
+ **One verdict owner per Story.** The ceremony decision names it
287
248
  (`verdictOwner: 'fresh-critic' | 'inline-self-eval'` from
288
- `resolveCeremonyForRisk`): the fresh maker-blind critic when sensitivity
289
- routes the cluster `fresh`, the contract-identical inline self-eval when it
290
- routes `inline`. Exactly one pass authors the verdict — never both, and
291
- never a preliminary self-assessment pass before dispatching the fresh
292
- critic (the redundant pre-pass buys no measurable quality and roughly
293
- triples the acceptance-block cost). `acceptance-eval.js` is the
294
- deterministic **scorer** of that one authored verdict — schema validation,
295
- round cap, proceed / redraft / block — not an independent additional pass
296
- over the criteria. The M4-B floor holds: one verdict per cluster, with the
297
- cluster count owned by the dispatching caller and never by routing.
298
-
299
- **One round = N cluster critics → ONE merged verdict → ONE gate call** (fresh
300
- critics; an inline-owned verdict is one file scored in one call). The
301
- clusters are how a round is _authored_; they are not how it is _scored_.
302
- Concatenate every cluster's records into a single `criteria[]` ordered by
303
- `index` — exactly one per `acceptance[]` item, under one `storyId`,
304
- `schemaVersion`, `round` and `commitSha` — and hand that merged file to the
305
- gate once:
249
+ `resolveCeremonyForRisk`), and it follows the **ceremony profile alone**
250
+ (digest § 3): the contract-identical inline self-eval under `minimal` /
251
+ `standard` (the default), the fresh maker-blind critic under `strict`.
252
+ Exactly one pass authors the verdict — never both, and never a preliminary
253
+ self-assessment before dispatching a fresh critic (the redundant pre-pass
254
+ buys no measurable quality and roughly triples the acceptance-block cost).
255
+ `acceptance-eval.js` is the deterministic **scorer** of that one authored
256
+ verdict — schema validation, round cap, proceed / redraft / block — not an
257
+ independent additional pass over the criteria.
258
+
259
+ **One round = ONE verdict file → ONE gate call.** The verdict covers every
260
+ `acceptance[]` item exactly once, ordered by `index`, under one `storyId`,
261
+ `schemaVersion`, `round` and `commitSha`:
306
262
 
307
263
  ```bash
308
264
  node <main-repo>/.agents/scripts/acceptance-eval.js \
309
- --story <storyId> --verdict <merged-verdict-path>
265
+ --story <storyId> --verdict <verdict-path>
310
266
  ```
311
267
 
312
268
  The gate reads the Story's `acceptance[]` count itself (Story #5313), so a
313
- single cluster's verdict handed over unmerged is rejected before scoring and
314
- consumes no round; `--expected-criteria` is accepted but redundant. Calling
315
- the gate once per cluster instead spends a round
316
- _per cluster_ — a Story past the cluster ceiling would burn its whole redraft
317
- budget on cluster arithmetic — and N concurrent calls race the Story-scoped
318
- round ledger. Full per-round mechanics, including the parallel dispatch and the
319
- merge shape: [`acceptance-self-eval.md`](acceptance-self-eval.md).
320
-
321
- **Critic evidence-share.** When the critic runs a `verify[]`
322
- command that is byte-identical to a close gate (`lint` / `typecheck`), it
323
- records the pass into the Story evidence keyspace via `--standalone` so
324
- close short-circuits the gate at unchanged HEAD. Run it in the **Story
325
- worktree** (`workCwd` from Step 0):
326
-
327
- ```bash
328
- node <main-repo>/.agents/scripts/evidence-gate.js \
329
- --standalone --scope-id <storyId> --gate lint \
330
- --worktree <workCwd> -- npm run lint
331
- ```
269
+ partial verdict is rejected before scoring and consumes no round;
270
+ `--expected-criteria` is accepted but redundant. A second gate call inside one
271
+ round spends a round for nothing, and concurrent calls race the Story-scoped
272
+ round ledger. Full per-round mechanics:
273
+ [`acceptance-self-eval.md`](acceptance-self-eval.md).
332
274
 
333
275
  **On `decision: "block"`** — post a `friction` comment naming the unmet
334
276
  criteria, then transition the Story to `agent::blocked`:
@@ -357,26 +299,17 @@ Its `level` comes from
357
299
  the one computed change-set list: a diff touching a sensitive path registered
358
300
  in `.agents/schemas/audit-rules.json` derives `high`, one touching none
359
301
  derives `low`, and an unenumerable diff (`files: null`) derives `null`.
360
- Hand the **same** `files` list to every acceptance critic you spawn (Step 1a)
361
- — a critic that re-ran its own `git diff` could score against a different
362
- set than the one that routed it.
363
-
364
- Its `mode` / `verdictOwner` come from
365
- [`resolveCeremonyForRisk`](../../scripts/lib/orchestration/ceremony-routing.js)
366
- (`minimal` → always inline; `strict` → always fresh; `standard` →
367
- `high`/`null` → `fresh`, `low` → `inline`). Review depth reads the same
368
- derived level via `review-depth.js` inside close, so the two decisions cannot
369
- disagree.
370
-
371
- **Inline-dispatch override.** When the Story dispatches
372
- `inline` (`resolveStoryDispatchMode` → `inline`, which is exactly a
373
- single-Story run — the function reads the resolved set size and nothing
374
- else), run
375
- every acceptance critic **inline** — do not spawn fresh-context critic
376
- sub-agents regardless of what the profile would otherwise resolve. The self-eval rigor
377
- (scoring each `acceptance[]` item against the one computed change set, with
378
- `verify[]` output as evidence) is unchanged; only the sub-agent boot is
379
- removed. Hard gates are untouched.
302
+ Hand the **same** `files` list to the verdict owner (Step 1a) — an evaluator
303
+ that re-ran its own `git diff` could score against a different set than the
304
+ one that routed it.
305
+
306
+ The derived level drives **review depth** and nothing else: `review-depth.js`
307
+ reads it inside close and still resolves `deep` for any sensitive class.
308
+ `mode` / `verdictOwner` come from
309
+ [`resolveCeremonyForRisk`](../../scripts/lib/orchestration/ceremony-routing.js),
310
+ which since Story #5366 accepts the profile and nothing else. The rule — and
311
+ what an `inline` dispatch does and does not change about it — is stated once
312
+ in **digest § 3**; do not restate it here.
380
313
 
381
314
  ---
382
315
 
@@ -716,9 +649,10 @@ resolves — that is the mechanism by which you wait. You MUST keep your turn al
716
649
  across the wait: watch → (fix + push + re-watch on red) → confirm the merge
717
650
  (Step 5) → flip `agent::done` → run the post-merge steps → and only then
718
651
  return the terminal JSON status contract. The CI wait NEVER terminates your
719
- turn; **only** a confirmed-`MERGED` PR (→ `status: "done"`), an
652
+ turn; **only** a confirmed-`MERGED` PR (→ `status: "landed"`), an
720
653
  `agent::blocked` transition (→ `status: "blocked"`), or an unrecoverable
721
- failure (→ `status: "failed"`) does. Ending your turn with prose and an
654
+ failure (→ `status: "failed"`) does — the statuses the shipped
655
+ [terminal schema](../../schemas/story-deliver-terminal.schema.json) accepts. Ending your turn with prose and an
722
656
  unconfirmed merge is a contract violation — it is the very bug this workflow
723
657
  exists to prevent.
724
658
 
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  description:
3
3
  Execute one Story end-to-end: story-<id> from main, implemented in a worktree
4
- (optional ## Slicing checkpoints), derived-level ceremony, PR against main.
4
+ (optional ## Slicing checkpoints), profile-resolved ceremony, PR against main.
5
5
  mandatoryReads: [deliver-digest.md]
6
6
  ---
7
7
 
@@ -18,13 +18,12 @@ mandatoryReads: [deliver-digest.md]
18
18
  The **one** delivery engine in v2:
19
19
 
20
20
  ```text
21
- single-story-init.js → implement + commits → derived-level ceremony → push
21
+ single-story-init.js → implement + commits → ceremony → push
22
22
  ──hand-off──▶ single-story-close.js (gates, PR → main, agent::closing)
23
23
  → CI watch + merge → single-story-confirm-merge.js (agent::done)
24
24
  ```
25
25
 
26
- An `Epic: #N` reference marks a v1 ticket — **stop** and re-plan. Engine
27
- traits and prerequisites: reference § Engine invariants.
26
+ Engine traits and prerequisites: reference § Engine invariants.
28
27
 
29
28
  ## Who owns which step
30
29
 
@@ -33,8 +32,9 @@ that dispatched the work**, never to a spawned worker.
33
32
 
34
33
  - **Inline dispatch** (a one-Story run — digest § 1): one session is both
35
34
  roles and walks Steps 0→7, no hand-off.
36
- - **Sub-agent dispatch**: the `story-worker` stops at Step 2.5 with the branch
37
- pushed and returns a hand-off; the dispatching `/mandrel-deliver` session runs Step 3
35
+ - **Sub-agent dispatch**: the `story-worker` boots on the `promptPath`
36
+ `deliver-run.js` wrote for it, stops at Step 2.5 with the branch pushed and
37
+ returns a hand-off; the dispatching `/mandrel-deliver` session runs Step 3
38
38
  **in its own turn** and **serializes the tail — one close at a time across
39
39
  the run**, even though implementation ran in parallel (reference § Step 3).
40
40
 
@@ -84,22 +84,21 @@ live in [`acceptance-self-eval.md`](acceptance-self-eval.md). **`proceed`** →
84
84
  Step 2. **`block`** → **do not close**: post a `friction` comment and flip
85
85
  `agent::blocked` (reference § Step 1a).
86
86
 
87
- ## Step 2 — Ceremony (profile + derived level)
87
+ ## Step 2 — Ceremony (the profile alone)
88
88
 
89
- Ceremony is `delivery.routing.ceremonyProfile` × the **derived change level**,
90
- never a planner-authored verdict. **Digest § 3** is the incantation (change set
91
- once, derive the level, resolve critics with `ceremony-routing.js`); edge cases
92
- are reference § Step 2. Hard gates always run in Step 3 — the derived level
93
- never disables them; do **not** pre-run the chain here — Step 2.5's credited
94
- suite run is the sole exception.
89
+ The acceptance verdict owner comes from `delivery.routing.ceremonyProfile` and
90
+ nothing else — not the diff, not the dispatch mode, never a planner-authored
91
+ verdict. **Digest § 3** is the incantation (change set once, derive the level
92
+ for review depth, resolve the owner with `ceremony-routing.js`) and the one
93
+ home of the rule; edge cases are reference § Step 2. Hard gates always run in
94
+ Step 3 — nothing here disables them; do **not** pre-run the chain here —
95
+ Step 2.5's credited suite run is the sole exception.
95
96
 
96
97
  ### Step 2.5 — The one credited suite run, the push, then hand off
97
98
 
98
- After the self-eval loop's last fix commit, run the suite **once** in the
99
- worktree through the depositor — `evidence-gate.js … --gate test -- npm test`,
100
- spelled out in **digest § 5**. It runs whatever `npm test` resolves to and
101
- stamps that, so the `test` credit is earned on any runner and only a *later*
102
- commit invalidates it. Red → fix, commit, re-run.
99
+ After the self-eval loop's last fix commit, run the one credited suite run
100
+ in the worktree — **digest § 5** is its only home and carries the
101
+ invocation. Red → fix, commit, re-run.
103
102
 
104
103
  Push `story-<storyId>` to `origin`, confirming the remote ref moved. Then
105
104
  (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
@@ -153,7 +153,8 @@ template's pre-resolved `changes[]` and the `/prototype` offer, nothing else.
153
153
  Story #5312 deleted the plan-side lite claim that used to read them
154
154
  (`--route-downgrade-reason`, the persist shape backstop, the `route::lite`
155
155
  hint): every Story lands through the same engine and the same close gates,
156
- and ceremony is derived from the landed diff at close.
156
+ and the acceptance verdict owner follows the operator's ceremony profile —
157
+ never anything a plan authored.
157
158
 
158
159
  ## Correct-by-construction authoring template
159
160
 
@@ -169,15 +170,15 @@ and ceremony is derived from the landed diff at close.
169
170
  are omitted: the work regenerates them, the refresh is a close-gate concern,
170
171
  and a declared shared artifact path reserves a footprint that needlessly
171
172
  serializes sibling Stories at dispatch.
172
- - **`changes[]` arrive pre-resolved to creates-vs-refactors.** Every path
173
- the seed predicted is probed against the repo: an existing path is
174
- emitted with `assumption: "refactors-existing"`, a missing one with
175
- `assumption: "creates"`. The persist gates stay authoritative — they probe
176
- the base branch ref, not the working tree — but a `creates` on a path that
177
- exists at base, or a `refactors-existing` on one that does not, is a
178
- dry-run **warning**, not a rejection; only a `deletes` naming an absent
179
- path is refused. A plain-string bullet or a trailing parenthetical is
180
- repaired into the object form by probing base, and the repair is reported.
173
+ - **`changes[]` entries are bare paths.** The skeleton emits each seed-predicted
174
+ path as a bare string and persist derives its assumption by probing the base
175
+ branch ref — present is a `refactors-existing`, absent a `creates` — and
176
+ reports each derivation on the dry-run's repair list. Pin
177
+ `{path, assumption}` only when the probe would get it wrong; a `creates` on
178
+ a path that exists at base, or a `refactors-existing` on one that does not,
179
+ is then a dry-run **warning**, not a rejection. `deletes` stays explicit —
180
+ a bare path can never express a removal — and a `deletes` naming a path
181
+ absent at base is the one `changes[]` shape still refused.
181
182
  - **Keep `## Spec` at contract-level prose** — interfaces, invariants,
182
183
  load-bearing constraints; no per-file behavior narration — and as long as
183
184
  the work needs. There is no word or token budget.
@@ -188,11 +189,12 @@ kept — passes the persist ticket validators with no round-trip.
188
189
  ### Authored entry shape
189
190
 
190
191
  Each `stories.json` entry: `slug` (`^[a-z0-9][a-z0-9-]*$`), `type: "story"`,
191
- `title`, `body` (`goal`, optional `spec`, `changes[{path, assumption}]` —
192
- `creates|refactors-existing|deletes`, `non_goals`, `reason_to_exist`),
193
- top-level `acceptance[]`, `verify[]` (each a **bare command** — there is no
194
- tier suffix), and `depends_on[]` (a sibling slug, or `#<id>` for an existing
195
- open Story).
192
+ `title`, `body` (`goal`, optional `spec`, `changes[]` — a **bare path string**
193
+ by default, or `{path, assumption}` with
194
+ `creates|refactors-existing|deletes` to pin one, `non_goals`,
195
+ `reason_to_exist`), top-level `acceptance[]`, `verify[]` (each a **bare
196
+ command** — there is no tier suffix), and `depends_on[]` (a sibling slug, or
197
+ `#<id>` for an existing open Story).
196
198
 
197
199
  Author `acceptance[]` **without** the `AC-<n>:` handle: the body renderer
198
200
  numbers each checkbox from its array position, so a carried handle renders
@@ -200,11 +202,11 @@ doubled. Persist normalises one off rather than refusing, and names the strip
200
202
  on the dry-run's repair list.
201
203
 
202
204
  Nothing in that shape inventories the repo for the author. `changes[]` arrives
203
- pre-resolved against the working tree, and Phase 8's
204
- `validateStoryFileAssumptions` re-probes every `{path, assumption}` at persist
205
- as a hard error — so the grounding contract is the author's own targeted reads
206
- plus that gate. There is no pre-computed codebase snapshot to fall back on,
207
- and no manifest-derived replacement to build.
205
+ filled in from the working-tree probe, and Phase 8's
206
+ `validateStoryFileAssumptions` re-probes every resolved `{path, assumption}`
207
+ at persist — so the grounding contract is the author's own targeted reads plus
208
+ that gate. There is no pre-computed codebase snapshot to fall back on, and no
209
+ manifest-derived replacement to build.
208
210
 
209
211
  ### Per-Story audit provenance (`provenance`)
210
212
 
@@ -264,8 +266,8 @@ substring-match advisories that depended on it are retired, leaving
264
266
  `shared-editor` as the one conflict kind. Both passes complete before the first
265
267
  `createIssue`, so a refusal still costs no writes.
266
268
 
267
- `shared-editor` findings are rendered into the posted `plan-summary` comment,
268
- directly beneath the wave table: the table promises which Stories can run
269
+ `shared-editor` findings are rendered into the plan-summary section of every
270
+ Story's posted `story-plan-state` comment, directly beneath the wave table: the table promises which Stories can run
269
271
  together, and a path two same-wave Stories both write is exactly where that
270
272
  promise breaks. Promise and caveat belong on one durable surface — previously
271
273
  the caveat was a stderr warning nobody kept.
@@ -318,13 +320,15 @@ template-only prose.
318
320
 
319
321
  ### Supersede-map partition
320
322
 
321
- `plan-persist` refuses a partial supersede map **before** it creates any
322
- Story, the same fail-closed shape as the collision refusal: every id passed to
323
- `--tickets` must be claimed by **exactly one** Story, and no Story may
324
- claim an id that was not a source ticket. With N>1 the mapping is not
325
- total by default — an authored map is the only thing that can say
326
- `#11-#14 → #20` while `#15 → #21`, which a blanket "superseded by
327
- this plan-run" reference could not.
323
+ `plan-persist` completes the supersede map **before** it creates any Story,
324
+ splitting it by who can be right. **Refused, fail-closed:** a Story claiming
325
+ an id that was never a source ticket, and two Stories claiming the same id —
326
+ the first would close an issue nobody asked about, the second cannot say which
327
+ Story replaced it. **Assigned with a warning:** a source id no Story claimed
328
+ goes to the primary Story, because this plan is replacing it either way and
329
+ only the bookkeeping was open. With N>1 an authored map is still the only
330
+ thing that can say `#11-#14 → #20` while `#15 → #21`, so author one whenever
331
+ the default is not what you mean.
328
332
 
329
333
  ## The pre-mortem critic — operator-invoked
330
334
 
@@ -369,33 +373,41 @@ a re-author round.
369
373
 
370
374
  ## What `--dry-run` actually gates
371
375
 
372
- `plan-persist.js --dry-run` is the same command with GitHub writes suppressed,
373
- and every gate runs before the first `createIssue` would fire. Since
374
- Story #5312 the gates split two ways, and the dry-run is where the second
375
- half is read:
376
+ A bare `plan-persist.js` runs the whole gate list write-free and then, on a
377
+ clean list, persists in the same invocation (Story #5342); `--dry-run` is that
378
+ same first half with the second suppressed. Either way every gate runs before
379
+ the first `createIssue` would fire. Since Story #5312 the gates split two
380
+ ways, and the dry-run output is where the second half is read:
376
381
 
377
382
  **Hard — the run refuses:** a body that does not parse, a ticket that is not
378
- a Story, an empty `acceptance[]` or `verify[]`, an unknown or cyclic
379
- `depends_on`, the acceptance partition at N>1, the supersede partition, a
380
- forbidden commit-subject prefix, and a `deletes` entry naming a path absent
381
- at base.
383
+ a Story, an empty `acceptance[]`, an unknown or cyclic `depends_on`, the
384
+ same-wave collision refusal at N>1, a supersede claim on a non-source id or
385
+ one claimed twice, and a `deletes` entry naming a path absent at base.
386
+ Story #5342 retired two: the commit-subject-prefix scan is gone entirely —
387
+ the `commit-msg` hook and `normalize-pr-title.js` enforce subjects — and an
388
+ empty `verify[]` is now a warning.
382
389
 
383
390
  **Warnings — listed, then the persist proceeds:** a `creates` on a path that
384
391
  exists at base or a `refactors-existing` on one that does not (including a
385
392
  path the base branch deleted or renamed, named with the removing commit), a
386
- goal or acceptance path absent at base, a `verify[]` command naming an absent
387
- test file, an `open-question` in a body (`Flag if…`, `TBD`, a trailing `?`),
388
- and a `pinned-identifier` in an acceptance item — a backticked bare symbol
393
+ goal or acceptance path absent at base, an empty `verify[]`, a source id
394
+ assigned to the primary Story by default, a `verify[]` command naming an
395
+ absent test file, an `open-question` in a body (`Flag if…`, `TBD`, a trailing
396
+ `?`), and a `pinned-identifier` in an acceptance item — a backticked bare symbol
389
397
  that is not a path, a label, a kebab token, a flag or a command, which the
390
398
  advisory `changes[]` is free to reshape out from under the criterion. The
391
399
  list also names every **repair** the run applied — a plain-string bullet or a
392
400
  trailing parenthetical rewritten into `{ path, assumption }` by probing base,
393
- and an `AC-<n>:` handle normalised off an acceptance item. The same list rides the result
394
- envelope as `warnings[]` and `repairs[]`, so a `--chain-on-clean` run loses
395
- nothing.
401
+ and an `AC-<n>:` handle normalised off an acceptance item. The same list rides
402
+ the result envelope as `warnings[]` and `repairs[]` — and because the repairs
403
+ are applied in the write-free pass, the chained run carries that pass's
404
+ `repairs[]` and `warnings[]` onto the envelope it returns, so the evidence
405
+ survives the hand-off rather than being re-derived from an already-repaired
406
+ draft.
396
407
 
397
408
  A dry run that comes back clean has paid for every deterministic refusal, so
398
- the real persist has nothing left to discover except network failure.
409
+ the real persist has nothing left to discover except network failure — which
410
+ is why the clean case no longer waits for a second operator invocation.
399
411
 
400
412
  ## The container Epic (Gate #3)
401
413
 
@@ -458,10 +470,8 @@ body lint), a `## Spec`, an `acceptance[]` / `verify[]`, or any path, finding
458
470
  or rationale a child does not already hold. It is a container; unique content
459
471
  here is content no delivering agent reads.
460
472
 
461
- What the **Stories** never gain is an `Epic: #N` footer. Linkage is
462
- parent→child only, which is exactly why every existing refusal of that footer
463
- still stands and each Story stays independently deliverable (ADR
464
- `20260905-5139`).
473
+ Linkage runs parent→child only, which is why each Story stays independently
474
+ deliverable (ADR `20260905-5139`).
465
475
 
466
476
  Degradation is deliberate: an unensurable `type::epic` label skips the Epic
467
477
  entirely (an unlabelled container is not a container), while a failed
@@ -508,13 +518,14 @@ state.
508
518
  ## Ready means fully persisted
509
519
 
510
520
  `agent::ready` is the **terminal** step, not part of the creating POST.
511
- The order is: create unlabelled → upsert `story-plan-state` on
512
- every Story → upsert `plan-summary` on the primary → flip every Story to
513
- `agent::ready`.
521
+ The order is: create unlabelled → upsert `story-plan-state` on every Story —
522
+ since Story #5343 one comment, carrying the plan summary (story set, delivery
523
+ order, deliver command) and nothing else, since Story #5367 deleted the machine
524
+ checkpoint no reader consumed — → flip every Story to `agent::ready`.
514
525
 
515
526
  This is what lets `/mandrel-deliver` trust the label: a Story carrying
516
- `agent::ready` always has its persist receipt on the ticket, so nothing can
517
- pick it up mid-write and read a half-persisted plan.
527
+ `agent::ready` always has the operator's delivery instructions on the ticket,
528
+ so nothing can pick it up mid-write and act on a half-persisted plan.
518
529
 
519
530
  ## Resuming a failed persist
520
531
 
@@ -546,7 +557,7 @@ them **envelope-first**:
546
557
 
547
558
  | Channel | When it wins |
548
559
  | --- | --- |
549
- | Envelope `sourceTickets[]` | **The normal path.** Written by step 1's `--out`, then read from `--plan-context <file>` or auto-discovered at `<plan-dir>/plan-context.json`. No ids to re-type. |
560
+ | Envelope `sourceTickets[]` | **The normal path.** Written by step 1 to `<tempRoot>/plan-<slug>/` (or wherever `--out` points), then read from `--plan-context <file>` or auto-discovered at `<plan-dir>/plan-context.json`. No ids to re-type. |
550
561
  | `--source-tickets <ids>` | Explicit **override** for hand-driven runs (no captured envelope, or deliberately narrowing the set). Wins over the envelope; a disagreement is warned about, not silently reconciled. |
551
562
 
552
563
  The result envelope's `supersede.sourceTicketOrigin` reports which channel was