mandrel 2.56.0 → 2.57.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (106) hide show
  1. package/.agents/agents/plan-critic.md +13 -18
  2. package/.agents/agents/story-worker.md +25 -34
  3. package/.agents/docs/agentrc-reference.json +0 -30
  4. package/.agents/docs/configuration.md +8 -28
  5. package/.agents/docs/execution-reference.md +5 -5
  6. package/.agents/docs/quality-gates.md +8 -7
  7. package/.agents/instructions.md +9 -10
  8. package/.agents/schemas/agentrc.schema.json +9 -185
  9. package/.agents/schemas/story-deliver-terminal.schema.json +1 -1
  10. package/.agents/scripts/acceptance-eval.js +107 -17
  11. package/.agents/scripts/ceremony-derive.js +191 -0
  12. package/.agents/scripts/check-context-budget.js +28 -33
  13. package/.agents/scripts/check-cyclomatic.js +4 -3
  14. package/.agents/scripts/deliver-light.js +31 -94
  15. package/.agents/scripts/lib/audit-suite/checklist-threading.js +15 -2
  16. package/.agents/scripts/lib/baselines/coverage-updater-cli.js +110 -0
  17. package/.agents/scripts/lib/baselines/crap-preview-scan.js +25 -0
  18. package/.agents/scripts/lib/baselines/crap-updater-cli.js +223 -0
  19. package/.agents/scripts/lib/bdd-scenario-budget.js +21 -3
  20. package/.agents/scripts/lib/bootstrap/quality-bootstrap.js +0 -1
  21. package/.agents/scripts/lib/close-validation/gates.js +52 -1
  22. package/.agents/scripts/lib/config/acceptance-eval.js +25 -57
  23. package/.agents/scripts/lib/config/delivery-routing.js +7 -33
  24. package/.agents/scripts/lib/config/explain.js +0 -19
  25. package/.agents/scripts/lib/config/limits.js +18 -78
  26. package/.agents/scripts/lib/config/quality.js +6 -3
  27. package/.agents/scripts/lib/config/runners.js +3 -2
  28. package/.agents/scripts/lib/config-settings-schema-delivery.js +15 -68
  29. package/.agents/scripts/lib/config-settings-schema-quality.js +0 -14
  30. package/.agents/scripts/lib/config-settings-schema.js +16 -143
  31. package/.agents/scripts/lib/crap-engine.js +35 -4
  32. package/.agents/scripts/lib/crap-utils.js +17 -1
  33. package/.agents/scripts/lib/cyclomatic-ceiling.js +19 -7
  34. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  35. package/.agents/scripts/lib/observability/runtime-friction.js +1 -1
  36. package/.agents/scripts/lib/observability/source-classifier.js +1 -0
  37. package/.agents/scripts/lib/orchestration/acceptance-eval-decision.js +5 -4
  38. package/.agents/scripts/lib/orchestration/ceremony-routing.js +19 -73
  39. package/.agents/scripts/lib/orchestration/complexity-gate.js +46 -212
  40. package/.agents/scripts/lib/orchestration/file-assumptions.js +32 -17
  41. package/.agents/scripts/lib/orchestration/light-escalation.js +3 -3
  42. package/.agents/scripts/lib/orchestration/light-suitability.js +66 -233
  43. package/.agents/scripts/lib/orchestration/plan-context.js +181 -387
  44. package/.agents/scripts/lib/orchestration/plan-critic-conditions.js +42 -153
  45. package/.agents/scripts/lib/orchestration/plan-critics-evaluate.js +14 -70
  46. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +300 -0
  47. package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +131 -168
  48. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +118 -297
  49. package/.agents/scripts/lib/orchestration/plan-persist/soft-findings.js +55 -0
  50. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +16 -65
  51. package/.agents/scripts/lib/orchestration/plan-persist/wave-serialisation.js +22 -35
  52. package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +30 -139
  53. package/.agents/scripts/lib/orchestration/planning/memory-pool-advisory.js +61 -223
  54. package/.agents/scripts/lib/orchestration/single-story-close/phases/close-validation.js +5 -0
  55. package/.agents/scripts/lib/orchestration/single-story-close/phases/pre-gate-steps.js +46 -16
  56. package/.agents/scripts/lib/orchestration/story-close/context-budget-writeback.js +213 -0
  57. package/.agents/scripts/lib/orchestration/task-body-validator.js +10 -63
  58. package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +33 -539
  59. package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +21 -414
  60. package/.agents/scripts/lib/orchestration/ticket-validator.js +54 -118
  61. package/.agents/scripts/lib/orchestration/verify-credit.js +69 -24
  62. package/.agents/scripts/lib/story-body/body-format-lints.js +15 -85
  63. package/.agents/scripts/lib/story-body/story-body.js +17 -237
  64. package/.agents/scripts/lib/templates/decomposer-prompts.js +84 -121
  65. package/.agents/scripts/lib/test-isolate/cli-options.js +93 -0
  66. package/.agents/scripts/lib/test-isolate/progress-log.js +45 -0
  67. package/.agents/scripts/lib/test-isolate/render-report.js +97 -0
  68. package/.agents/scripts/lib/test-isolate/run-isolate.js +87 -0
  69. package/.agents/scripts/lib/test-run-credit.js +266 -0
  70. package/.agents/scripts/lib/wave-runner/footprint.js +48 -358
  71. package/.agents/scripts/lib/wave-runner/ready-set.js +6 -5
  72. package/.agents/scripts/lib/workers/crap-worker.js +32 -41
  73. package/.agents/scripts/plan-context.js +7 -9
  74. package/.agents/scripts/plan-critics.js +28 -54
  75. package/.agents/scripts/plan-persist.js +25 -68
  76. package/.agents/scripts/quality-preview.js +51 -0
  77. package/.agents/scripts/run-tests.js +12 -0
  78. package/.agents/scripts/stories-wave-tick.js +23 -45
  79. package/.agents/scripts/test-isolate.js +13 -180
  80. package/.agents/scripts/update-coverage-baseline.js +25 -70
  81. package/.agents/scripts/update-crap-baseline.js +19 -123
  82. package/.agents/skills/core/scope-triage/SKILL.md +3 -3
  83. package/.agents/workflows/audit-clean-code.md +4 -3
  84. package/.agents/workflows/helpers/acceptance-self-eval.md +41 -41
  85. package/.agents/workflows/helpers/code-quality-guardrails.md +4 -4
  86. package/.agents/workflows/helpers/code-review.md +2 -3
  87. package/.agents/workflows/helpers/deliver-digest.md +41 -57
  88. package/.agents/workflows/helpers/deliver-light.md +40 -105
  89. package/.agents/workflows/helpers/deliver-reference.md +1 -1
  90. package/.agents/workflows/helpers/deliver-story-reference.md +37 -58
  91. package/.agents/workflows/helpers/deliver-story.md +9 -13
  92. package/.agents/workflows/helpers/plan-reference.md +132 -219
  93. package/.agents/workflows/mandrel-plan.md +27 -40
  94. package/.agents/workflows/memory-consolidate.md +9 -13
  95. package/docs/CHANGELOG.md +23 -0
  96. package/lib/migrations/index.js +4 -0
  97. package/lib/migrations/steps/2.57.0-retire-delivery-limit-knobs.js +45 -0
  98. package/lib/migrations/steps/2.57.0-retire-planning-limit-knobs.js +59 -0
  99. package/package.json +1 -1
  100. package/.agents/scripts/lib/framework-version.js +0 -39
  101. package/.agents/scripts/lib/orchestration/consolidation-precondition.js +0 -223
  102. package/.agents/scripts/lib/orchestration/plan-persist/fan-out-gate.js +0 -97
  103. package/.agents/scripts/lib/orchestration/planning/decomposer-context.js +0 -26
  104. package/.agents/scripts/lib/orchestration/spec-budget.js +0 -89
  105. package/.agents/scripts/lib/orchestration/spec-spill.js +0 -74
  106. package/.agents/scripts/lib/orchestration/verify-tier-repair.js +0 -107
@@ -140,10 +140,8 @@ full-ceremony Story. The lite route's `preserves` field is the machine-readable
140
140
  record of those non-negotiables; there is no lite-specific gate bypass.
141
141
 
142
142
  **Ceremony comes from the landed diff; the dispatch mode comes from the
143
- run.** Persist stamps a lite cohort's Stories with the `route::lite` label as a
144
- _human-visible hint only_ (and ledgers the authored verdict — recorded reason
145
- plus per-Story shape evidence — on the `story-plan-state` checkpoint); the
146
- label is never the control signal. Ceremony is resolved from the **derived
143
+ run.** Persist stamps no route label (Story #5312 retired the plan-side lite
144
+ claim with its `route::lite` hint). Ceremony is resolved from the **derived
147
145
  change level** (`deriveChangeLevel` over the computed change set — digest § 3),
148
146
  not from a body-shape read: a footprint intersecting a sensitive-path class
149
147
  derives `high`, so the Story keeps its fresh acceptance critic. The light path
@@ -240,35 +238,19 @@ discovers them only after the whole close pipeline has run, at several times
240
238
  the cost of one full-suite run in the worktree.
241
239
 
242
240
  **Run it once, last, so close can credit it.** The run belongs **after** the
243
- self-eval loop's last fix commit and **after** the hand-off push — the credit
244
- is keyed on the tree, not on push state, so pushing first keeps the stamp and
245
- buys the ordering Step 2.5 needs (the capture is backgrounded, and its
246
- completion ends the turn); redraft rounds run scoped tests. Close skips a gate that already passed at the current HEAD, but a bare
247
- `npm test` deposits no such record — the suite then runs twice per delivery,
248
- once here and once in the close gate chain. Pick the invocation by the same
249
- predicate `close-validation/gates.js` uses to choose its test gate:
250
-
251
- ```bash
252
- # CRAP gate enabled (default) + a `test:coverage` script — writes the stamp
253
- # the close `coverage-capture` gate reads:
254
- node <main-repo>/.agents/scripts/coverage-capture.js --cwd <workCwd>
255
- # otherwise — the evidence record the close `test` gate reads. <workCwd> must
256
- # be ABSOLUTE and the runner exactly `npm test`: both sides hash
257
- # {cmd, args, cwd}, so a relative path or a wrapper misses the credit.
258
- node <main-repo>/.agents/scripts/evidence-gate.js --standalone \
259
- --scope-id <storyId> --gate test --worktree <workCwd> -- npm test
260
- ```
261
-
262
- The credit expires the moment it stops describing the tree: evidence is keyed
263
- on HEAD, the capture stamp on a content digest of `crap.targetDirs`. A
264
- self-eval fix — or any commit — invalidates it and close re-runs the suite for
265
- real, so this never trades away the gate. That keying is exactly why the run
266
- comes last, and why the push before it is free. Close's own base-sync can
267
- spend the stamp too when it lands base commits; it now says so out loud rather
268
- than silently re-running the suite.
269
-
270
- **`verify[]` reuses the same stamp.** A `verify[]` entry that is itself a
271
- full-suite command is reported **credited** against that stamp rather than
241
+ self-eval loop's last fix commit; redraft rounds run scoped tests. A green
242
+ `npm test` in the worktree on `story-<id>` deposits the `test` evidence
243
+ record close reads (Story #5313 — `lib/test-run-credit.js`), keyed on HEAD
244
+ and the tree fingerprint and hashed on the exact command close spawns, so
245
+ close reports the gate as credited at unchanged HEAD instead of re-running
246
+ the suite. The credit expires the moment it stops describing the tree: any
247
+ later commit invalidates it and close re-runs the suite for real, so this
248
+ never trades away the gate. The CRAP gate still runs `coverage-capture.js`
249
+ itself when it needs a fresh artifact — the capture stamp is a claim about
250
+ `coverage/coverage-final.json`, which a bare `npm test` does not produce.
251
+
252
+ **`verify[]` reuses the same credit.** A `verify[]` entry that is itself a
253
+ full-suite command is reported **credited** against that record rather than
272
254
  respawned (`resolveVerifyCredit` in
273
255
  [`verify-credit.js`](../../scripts/lib/orchestration/verify-credit.js)), and the
274
256
  self-eval gate warns when it sees one: the intended shape is scoped `verify[]`
@@ -294,24 +276,23 @@ round cap, proceed / redraft / block — not an independent additional pass
294
276
  over the criteria. The M4-B floor holds: one verdict per cluster, with the
295
277
  cluster count owned by the dispatching caller and never by routing.
296
278
 
297
- **One round = N cluster critics → ONE merged verdict → ONE gate call.** The
279
+ **One round = N cluster critics → ONE merged verdict → ONE gate call** (fresh
280
+ critics; an inline-owned verdict is one file scored in one call). The
298
281
  clusters are how a round is _authored_; they are not how it is _scored_.
299
282
  Concatenate every cluster's records into a single `criteria[]` ordered by
300
283
  `index` — exactly one per `acceptance[]` item, under one `storyId`,
301
284
  `schemaVersion`, `round` and `commitSha` — and hand that merged file to the
302
- gate once, with `--expected-criteria` set to the Story's `acceptance[]` count:
285
+ gate once:
303
286
 
304
287
  ```bash
305
288
  node <main-repo>/.agents/scripts/acceptance-eval.js \
306
- --story <storyId> --verdict <merged-verdict-path> \
307
- --expected-criteria <acceptance[] count>
289
+ --story <storyId> --verdict <merged-verdict-path>
308
290
  ```
309
291
 
310
- The flag is what makes the merge enforceable: `assertCriteriaCoverage` returns
311
- early on the `null` default, so **omitting it leaves the guard inert** and a
312
- single cluster's verdict handed over unmerged scores a fraction of the criteria
313
- and still reports `proceed`. A length mismatch is rejected before scoring and
314
- consumes no round. Calling the gate once per cluster instead spends a round
292
+ The gate reads the Story's `acceptance[]` count itself (Story #5313), so a
293
+ single cluster's verdict handed over unmerged is rejected before scoring and
294
+ consumes no round; `--expected-criteria` is accepted but redundant. Calling
295
+ the gate once per cluster instead spends a round
315
296
  _per cluster_ — a Story past the cluster ceiling would burn its whole redraft
316
297
  budget on cluster arithmetic — and N concurrent calls race the Story-scoped
317
298
  round ledger. Full per-round mechanics, including the parallel dispatch and the
@@ -342,32 +323,30 @@ node .agents/scripts/update-ticket-state.js --ticket <storyId> --state agent::bl
342
323
 
343
324
  ## Step 2 — Ceremony detail
344
325
 
345
- **Compute the change set once** with the shared enumerator —
346
- the same module close uses — and reuse that one list downstream:
326
+ **Compute the change set once** with the shared enumerator — the same module
327
+ close uses — and reuse that one list downstream. `ceremony-derive.js`
328
+ (Story #5313) is that enumeration, the level derivation and the ceremony
329
+ resolution in one call:
347
330
 
348
331
  ```bash
349
- node --input-type=module -e '
350
- import { computeChangeSet } from "<main-repo>/.agents/scripts/lib/orchestration/change-set.js";
351
- const { files } = computeChangeSet({ baseRef: "main", headRef: "story-<storyId>" });
352
- console.log(JSON.stringify(files));
353
- '
332
+ node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
354
333
  ```
355
334
 
356
- Derive the level with
335
+ Its `level` comes from
357
336
  [`deriveChangeLevel`](../../scripts/lib/orchestration/review-depth.js) over
358
337
  the one computed change-set list: a diff touching a sensitive path registered
359
338
  in `.agents/schemas/audit-rules.json` derives `high`, one touching none
360
- derives `low`, and an unenumerable diff (`files === null`) derives `null`.
361
- Hand the **same** list to every acceptance critic you spawn (Step 1a) — a
362
- critic that re-ran its own `git diff` could score against a different set
363
- than the one that routed it.
339
+ derives `low`, and an unenumerable diff (`files: null`) derives `null`.
340
+ Hand the **same** `files` list to every acceptance critic you spawn (Step 1a)
341
+ — a critic that re-ran its own `git diff` could score against a different
342
+ set than the one that routed it.
364
343
 
365
- Resolve fresh-vs-inline acceptance critics per AC-cluster with
344
+ Its `mode` / `verdictOwner` come from
366
345
  [`resolveCeremonyForRisk`](../../scripts/lib/orchestration/ceremony-routing.js)
367
346
  (`minimal` → always inline; `strict` → always fresh; `standard` →
368
- `high`/`null` → `fresh`, `low` → `inline` unless the `freshCriticSampleRate`
369
- floor forces `fresh`). Review depth reads the same derived level via
370
- `review-depth.js` inside close, so the two decisions cannot disagree.
347
+ `high`/`null` → `fresh`, `low` → `inline`). Review depth reads the same
348
+ derived level via `review-depth.js` inside close, so the two decisions cannot
349
+ disagree.
371
350
 
372
351
  **Inline-dispatch override.** When the Story dispatches
373
352
  `inline` (`resolveStoryDispatchMode` → `inline`, which is exactly a
@@ -74,7 +74,7 @@ One branch, one PR to `main`, commits against the inline `acceptance[]` /
74
74
  `## Slicing` rows as **intra-session checkpoints** (reference § Step 1).
75
75
  2. Implement and commit on the Story branch, iterating with quick advisory
76
76
  gates (`typecheck`, `lint`, scoped tests) — the full chain runs in Step 3,
77
- and the **one** creditable full-suite run at Step 2.5.
77
+ and the **one** full-suite run at Step 2.5.
78
78
 
79
79
  ### Step 1a — Bounded acceptance self-eval loop (**required**)
80
80
 
@@ -93,21 +93,17 @@ are reference § Step 2. Hard gates always run in Step 3 — the derived level
93
93
  never disables them; do **not** pre-run the chain here — Step 2.5's credited
94
94
  suite run is the sole exception.
95
95
 
96
- ### Step 2.5 — Push, then the creditable full-suite run, then hand off
96
+ ### Step 2.5 — The one full-suite run, the push, then hand off
97
97
 
98
- **Push first.** After the self-eval loop's last fix commit, push
99
- `story-<storyId>` to `origin`, confirming the remote ref moved: the capture
100
- below is backgrounded, so *its* completion ends the turn.
98
+ After the self-eval loop's last fix commit, run `npm test` **once** in the
99
+ worktree (**digest § 5**): a green full run deposits the `test` credit close
100
+ reads, keyed on the tree, so only a *later* commit invalidates it. Red →
101
+ fix, commit, re-run.
101
102
 
102
- Then run the full suite **once**, after the push, in the shape close credits
103
- (**digest § 5**): the credit is keyed on the tree, not on push state, so only
104
- a *later* commit invalidates it; a bare `npm test` deposits none. Red →
105
- fix, commit, push, re-capture.
106
-
107
- Then (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
103
+ Push `story-<storyId>` to `origin`, confirming the remote ref moved. Then
104
+ (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
108
105
  branch, pushed head SHA, self-eval verdict, `verify[]` evidence — and stop.
109
- Do not open the PR or compose a terminal envelope. An inline run captures
110
- before Step 3.
106
+ Do not open the PR or compose a terminal envelope.
111
107
 
112
108
  ## Step 3 — Close and land (`single-story-close.js`)
113
109
 
@@ -1,10 +1,10 @@
1
1
  # /mandrel-plan — on-demand reference appendix
2
2
 
3
3
  > **Applies when:** you are executing [`/mandrel-plan`](../mandrel-plan.md) and hit one of the
4
- > situations below — input-mode derivation, the Gate #1 light handoff,
5
- > shape-derived complexity routing, tickets-mode supersede authoring, critic
6
- > dispatch detail, a failed persist, or source-id resolution. The spine stays
7
- > resident; this file is read on demand.
4
+ > situations below — input-mode derivation, the Gate #1 advisory line,
5
+ > tickets-mode supersede authoring, the operator-invoked pre-mortem, a
6
+ > failed persist, or source-id resolution. The spine stays resident; this
7
+ > file is read on demand.
8
8
 
9
9
  ## Deriving the input mode
10
10
 
@@ -92,140 +92,58 @@ The marker keeps the operator's undelegated decisions findable after the
92
92
  fact: reviewing a `--yes` plan means scanning its decisions-made-by-default,
93
93
  not re-deriving which assumptions were really the agent's to make.
94
94
 
95
- ## Gate #1 → the memory-pool advisory (`memoryPoolAdvisory`)
96
-
97
- On a truthy `memoryPoolAdvisory.recommend`, name
98
- [`/memory-consolidate`](../memory-consolidate.md) at Gate #1, quoting its
99
- `reasons[]`. Purely advisory: a stale pool degrades recall, it does not make
100
- the plan wrong, so it never blocks and never reroutes.
101
-
102
- ## Gate #1 → graduating a CI-gap intake filing (`intake`)
103
-
104
- The envelope's `priorFeedback` arrays carry the open `meta::*` feedback issues.
105
- A row flagged `intake: true` is a **CI-gap intake filing** — written by
106
- [`file-ci-gap.js`](../../scripts/file-ci-gap.js) when a delivery reached an
107
- Option-2 verdict in [`ci-remediation.md`](../../rules/ci-remediation.md) and the
108
- root cause was outside its scope. It carries evidence (failure signature, run
109
- link, occurrence history, ownership routing) but **no `## Spec`, no
110
- `acceptance[]` / `verify[]` and no `agent::*` label**, so `/mandrel-deliver`
111
- cannot take it: it is intake awaiting graduation, by design. Delivery files it
112
- and moves on rather than blocking on a planning pass nobody is present for.
113
-
114
- Graduating one is exactly **tickets mode**: `/mandrel-plan <issue number>`
115
- rewrites it into a Story, `supersedes[]` claims it, and persist closes it.
116
-
117
- At Gate #1, in **ask** and **seed** mode, name any open intake rows and offer
118
- that instead of the seed in front of you — a filing that keeps recurring
119
- (its `## Occurrences` table is the count) is usually the better next Story than
120
- whatever prompted this run. It is **advisory**: never reroute automatically, and
121
- skip the offer entirely under `--yes`, where nobody is at the keyboard to take
122
- it. A `platformGaps[]` row is the same shape with a different owner — the
123
- Story it graduates into may well be a config or runbook change rather than code.
124
-
125
- ## Gate #1 → the light path (in-session handoff)
126
-
127
- On a confirmed `deliverLightSuggestion`, `/mandrel-plan` routes into
128
- [`deliver-light.md`](deliver-light.md) **without ending the session**. Two
129
- things make that safe, and both are worth understanding before changing it:
130
-
131
- 1. **The handoff carries the envelope, not the seed.** Gate #1 already holds
132
- the interrogated `complexitySignals`; fill the light gate's `--creates`
133
- / `--refactors` / `--acceptance` / `--reason` from those. Re-deriving from
134
- raw seed text throws away the better signal and can disagree with the
135
- suggestion that routed you.
136
- 2. **The gate still runs.** The suggestion is read against seed-time ceilings
137
- (`DELIVER_LIGHT_SUGGESTION_CEILINGS` — artifacts, risk hits, sensitive-path
138
- classes); the light gate is read against the predicted work's effort and risk
139
- (`STORY_SHAPE_CEILINGS` — change kinds, magnitude, uncertainty, deployable
140
- span). Two different checks on purpose, so a confirm is not a bypass.
141
-
142
- **When the light gate answers `ask-operator`**, the two ceiling sets disagreed.
143
- Resume `/mandrel-plan` at step 2 (Author) **in this same session** — the interrogation
144
- is still valid and re-paying for it buys nothing. This bounce-back is not an
145
- escalation.
146
-
147
- **Under `--yes` the offer is recorded and planning proceeds** — it is *never*
148
- auto-downgraded to light. An unattended run has nobody to confirm the reroute,
149
- and a suggestion is not a confirmation. The same rule governs unknown triage
150
- unattended: AFK unknowns are still researched, but no free-form operator
151
- question is asked — each HITL unknown lands in Key Assumptions marked a
152
- decision-made-by-default, so the record shows what was decided for the operator
153
- rather than pretending it was decided with them.
154
-
155
- Escalation in the *other* direction — an over-scope prompt on the light path —
156
- is terminal and requires a fresh session. The rule that separates the two, and
157
- why it must not be flattened into symmetry:
158
- [`deliver-light.md` § Why the two directions differ](deliver-light.md).
159
-
160
- ## Gate #1 → the `/prototype` offer (`uiSurface`)
161
-
162
- `complexitySignals.uiSurface` is the second advisory Gate #1 offer, and the
163
- weaker of the two on purpose: it carries **no routing authority and adds no
164
- gate**. Both halves are derived from observables already in the checkout — the
165
- `hasWebSurface` applicability predicate the `target: "web"` audit lenses gate
166
- on, and whether any predicted path matches a web lens `filePattern` registered
167
- in `audit-rules.json`. There is no configuration key to set: a project with no
168
- rendered frontend resolves falsey and the offer never fires.
169
-
170
- When it does fire, **name [`/prototype`](../prototype.md) and stop there.**
171
- `/mandrel-plan` must never invoke it — operator invocation is the entire design, because
172
- the value is a human looking at a layout before its UI acceptance criteria are
173
- frozen.
174
-
175
- **Under `--yes` the offer is recorded and planning proceeds** — no reroute, no
176
- prototype written, no gate raised. This is exactly how `deliverLightSuggestion`
177
- behaves unattended, and for the same reason: an unattended run has nobody to
178
- review an artifact, so recording the offer is the whole of the right behaviour.
179
-
180
- ## Shape-derived complexity routing (`complexitySignals`)
181
-
182
- Complexity routes on the **objective shape of the authored work**, never on
183
- seed word count — a detailed prompt can describe trivial work, a terse one
184
- complex work. The pipeline stages the
185
- decision:
186
-
187
- - **Signals, not routing.** The envelope's `complexitySignals` field is
188
- advisory only (`routingAuthority: false`): enumerated-artifact count (with
189
- the configured `maxArtifacts` threshold beside it as one input),
190
- `planning.riskHeuristics` phrases present in the seed, the repo state of
191
- predicted paths (existing paths predict refactors; missing predict
192
- creates), and the `audit-rules.json` sensitive-path classes the predicted
193
- footprint intersects.
194
- - **You author the verdict.** Judge the signals: a genuinely trivial scope
195
- (small additive footprint, no risk hits, no sensitive class) earns a `lite`
196
- claim via `plan-persist.js --route-downgrade-reason "<why>"`. The reason is
197
- recorded on every created Story's `story-plan-state` checkpoint, making the
198
- judgment auditable; without a recorded reason the conservative default
199
- (`full`) stands.
200
- - **Persist backstops the claim deterministically.** After authoring, the
201
- work has measurable shape, so persist validates the `lite` claim against
202
- each Story's own shape — distinct change kinds, declared magnitude,
203
- uncertainty, deployable/migration span, glob-free footprint, and
204
- sensitive-path classes, against the framework `STORY_SHAPE_CEILINGS` (effort
205
- and risk, never artifact counts) — and **fails closed to
206
- `full`** when any Story exceeds them (the refusal is ledgered on the
207
- checkpoint too). The lite route is **not** licence to drop a
208
- non-negotiable — every decision's `preserves` field enumerates what still
209
- holds: the Story ticket, the PR-to-`main` landing, every repo quality gate,
210
- and the security baseline. Those gates run in `single-story-close.js`
211
- regardless of route.
212
-
213
- **The label is a hint; deliver re-derives.** Persist labels a
214
- lite cohort's Stories with **`route::lite`** as a *human-visible hint only* —
215
- `/mandrel-deliver` computes the route from each fetched Story body via the same shape
216
- function at dispatch, so neither a lost label nor an unread marker can
217
- misroute delivery: a lite-shaped Story derives `lite` even with the label
218
- absent, and a sensitive-footprint Story routes `full` and keeps its fresh
219
- critic even with the label present. The derived route sets ceremony, not
220
- where the engine runs — sub-agent boots are collapsed by a **single-Story
221
- run**, never by a trivial shape. The `route::*` axis stays runtime-derived: hand-authored
222
- `route::*` entries in `labels[]` are dropped by persist.
223
-
224
- The knobs (`planning.complexityGate.{enabled, maxArtifacts}`) are documented
225
- in [`.agents/docs/configuration.md`](../../docs/configuration.md) under
226
- `### planning`; the defaults live on `DEFAULT_COMPLEXITY_GATE` and the shape
227
- ceilings on `STORY_SHAPE_CEILINGS` in
228
- [`lib/orchestration/complexity-gate.js`](../../scripts/lib/orchestration/complexity-gate.js).
95
+ ## Gate #1 → the one advisory line
96
+
97
+ Gate #1 stops for exactly two things — the sharpened plan intent and any HITL
98
+ unknown — and everything else the envelope surfaced collapses to **one
99
+ advisory line** beneath it (Story #5312). Nothing on that line stops the run,
100
+ reroutes it, or is invoked by `/mandrel-plan`; each item names something the
101
+ operator may prefer to do instead, and the run proceeds either way. Under
102
+ `--yes` the line is recorded and planning continues — an unattended run has
103
+ nobody to take an offer.
104
+
105
+ The line names, in order, whichever of these the envelope carries:
106
+
107
+ - **`duplicates[]`** — open Stories the seed resembles (never Epics). Name
108
+ the top one or two by id and title; a plan that duplicates open work is
109
+ still the operator's call.
110
+ - **Open `intake` rows** (`priorFeedback`) — CI-gap intake filings written by
111
+ [`file-ci-gap.js`](../../scripts/file-ci-gap.js) when a delivery reached an
112
+ Option-2 verdict in [`ci-remediation.md`](../../rules/ci-remediation.md).
113
+ They carry evidence but no `## Spec`, no `acceptance[]` / `verify[]` and no
114
+ `agent::*` label, so `/mandrel-deliver` cannot take one: graduating it is
115
+ exactly **tickets mode** (`/mandrel-plan <issue number>`), and a filing
116
+ that keeps recurring (its `## Occurrences` table is the count) is often the
117
+ better next Story than the seed in front of you. A `platformGaps[]` row is
118
+ the same shape with a different owner.
119
+ - **`memoryPoolAdvisory.recommend`** — name
120
+ [`/memory-consolidate`](../memory-consolidate.md), quoting its
121
+ `reasons[]`. The one arm left measures the `MEMORY.md` index against the
122
+ harness's byte cap; a stale pool degrades recall, it does not make the plan
123
+ wrong.
124
+ - **`complexitySignals.uiSurface`** — name [`/prototype`](../prototype.md)
125
+ and stop there. The signal carries **no routing authority and adds no
126
+ gate**: both halves are derived from observables already in the checkout —
127
+ the `hasWebSurface` applicability predicate the `target: "web"` audit
128
+ lenses gate on, and whether any predicted path matches a web lens
129
+ `filePattern` in `audit-rules.json`; a project with no rendered frontend
130
+ resolves falsey and the offer never fires. `/mandrel-plan` must never invoke
131
+ it — operator invocation is the entire design, because the value is a human
132
+ looking at a layout before its UI acceptance criteria are frozen. Under
133
+ `--yes` the offer is recorded and planning proceeds — no reroute, no
134
+ prototype written, no gate raised.
135
+
136
+ ## `complexitySignals` are advisory
137
+
138
+ The envelope's `complexitySignals` field carries the paths the seed predicts,
139
+ their repo state (existing paths predict refactors; missing predict creates)
140
+ and the `audit-rules.json` sensitive-path classes the footprint intersects —
141
+ `routingAuthority: false`, no `route` field. They ground the authoring
142
+ template's pre-resolved `changes[]` and the `/prototype` offer, nothing else.
143
+ Story #5312 deleted the plan-side lite claim that used to read them
144
+ (`--route-downgrade-reason`, the persist shape backstop, the `route::lite`
145
+ hint): every Story lands through the same engine and the same close gates,
146
+ and ceremony is derived from the landed diff at close.
229
147
 
230
148
  ## Correct-by-construction authoring template
231
149
 
@@ -233,30 +151,24 @@ ceilings on `STORY_SHAPE_CEILINGS` in
233
151
  **correct-by-construction** skeleton, built from the same repo probe the
234
152
  `complexitySignals` ran:
235
153
 
236
- - **`verify[]` placeholders already end with a valid `(tier)` tag.** Keep
237
- every filled entry's trailing tag one of `(unit)` / `(contract)` /
238
- `(e2e)` / `(validate)` (or use the `manual:<reason>` escape) — a tierless
239
- entry is exactly the mechanical persist round-trip the template exists to
240
- prevent.
154
+ - **`verify[]` entries are commands.** There is no tier suffix and no
155
+ `manual:<reason>` escape (Story #5312): write the exact command or test
156
+ path the deliverer runs and the acceptance critic reads as evidence.
241
157
  - **`changes[]` arrive pre-resolved to creates-vs-refactors.** Every path
242
158
  the seed predicted is probed against the repo: an existing path is
243
159
  emitted with `assumption: "refactors-existing"`, a missing one with
244
- `assumption: "creates"`. Trust the pre-resolved assumption — verify
245
- against the repo before overriding one (authoring `creates` for a file
246
- that exists at base is a validator rejection). The persist gates stay
247
- authoritative: they probe the base branch ref, not the working tree.
248
- - **Keep `## Spec` near contract-level prose.** **Aim for ~250 words; an
249
- advisory warning fires past 350** (`SPEC_SOFT_WORD_BUDGET`). Two numbers,
250
- two jobs: ~250 is the authoring target — the nudge toward a contract-level
251
- Spec (interfaces, invariants, load-bearing constraints; no per-file
252
- behavior narration) — while 350 is the slacker threshold at which persist
253
- actually warns, so the warning marks a real outlier instead of ordinary
254
- variance. Neither fails the persist. The hard fail-closed ceiling
255
- (~1500 tokens, `spec-spill.js`) is unchanged.
160
+ `assumption: "creates"`. The persist gates stay authoritative — they probe
161
+ the base branch ref, not the working tree — but a `creates` on a path that
162
+ exists at base, or a `refactors-existing` on one that does not, is a
163
+ dry-run **warning**, not a rejection; only a `deletes` naming an absent
164
+ path is refused. A plain-string bullet or a trailing parenthetical is
165
+ repaired into the object form by probing base, and the repair is reported.
166
+ - **Keep `## Spec` at contract-level prose** — interfaces, invariants,
167
+ load-bearing constraints; no per-file behavior narration — and as long as
168
+ the work needs. There is no word or token budget.
256
169
 
257
170
  A faithfully-filled skeleton — placeholders replaced, pre-resolved entries
258
- kept, tags valid — passes the persist ticket validators with no
259
- round-trip.
171
+ kept — passes the persist ticket validators with no round-trip.
260
172
 
261
173
  ### Authored entry shape
262
174
 
@@ -336,18 +248,13 @@ together, and a path two same-wave Stories both write is exactly where that
336
248
  promise breaks. Promise and caveat belong on one durable surface — previously
337
249
  the caveat was a stderr warning nobody kept.
338
250
 
339
- Two `planning.*` knobs upgrade a conflict class from advisory to a hard
340
- refusal. **Both default to `false` and are documented, not recommended:**
341
-
342
- | Knob | Upgrades | Why it is off |
343
- | --- | --- | --- |
344
- | `planning.failOnSharedEditors` | `shared-editor` → `hard` | Co-editing one file is routine and often correct; the delivery scheduler already serializes file-overlapping Stories. |
345
- | `planning.requireExplicitCrossStoryDeps` | `implicit-cross-story-dep` → `hard` | Path references are matched by substring, so a legitimate mention in prose can read as a dependency. |
346
-
347
- Turn one on for a repo where the class is genuinely fatal; expect a refusal to
348
- name the Stories and the fix (a `depends_on` edge, or folding the shared edit
349
- into one Story). The sibling knobs `failOnRegistryConflicts`,
350
- `failOnMissingBddScaffold` and `failOnLargeFanOut` behave the same way.
251
+ Every conflict class is advisory (Story #5312 retired the
252
+ `planning.failOn*` / `requireExplicitCrossStoryDeps` upgrade knobs with the
253
+ registry and fan-out findings): co-editing one file is routine and often
254
+ correct — the delivery scheduler already serializes file-overlapping Stories —
255
+ and a path reference matched by substring can read as a dependency a prose
256
+ mention never meant. A finding names the Stories and the fix (a `depends_on`
257
+ edge, or folding the shared edit into one Story) for the operator to weigh.
351
258
 
352
259
  ## Tickets mode — authoring `supersedes[]`
353
260
 
@@ -382,67 +289,73 @@ total by default — an authored map is the only thing that can say
382
289
  `#11-#14 → #20` while `#15 → #21`, which a blanket "superseded by
383
290
  this plan-run" reference could not.
384
291
 
385
- ## Critic dispatch detail
292
+ ## The pre-mortem critic — operator-invoked
293
+
294
+ The maker-blind **pre-mortem** critic is not a step of the spine (Story #5312
295
+ retired step 2.5 with the consolidation critic, whose one deterministic input
296
+ was a `## Delivery Slicing` table no Story carries). Run it when the operator
297
+ asks for it, after Author and before Persist — the last point a finding folds
298
+ into a re-author:
299
+
300
+ ```bash
301
+ node .agents/scripts/plan-critics.js \
302
+ --stories temp/plan-<slug>/stories.json \
303
+ [--tech-spec temp/plan-<slug>/techspec.md]
304
+ ```
386
305
 
387
- The **pre-mortem** critic fires on any of three deterministic triggers: the
388
- draft ticket count reaching half the reviewability budget, a
389
- `planning.riskHeuristics` phrase matching the plan text, or the
390
- **external-dependency** probe finding an out-of-repo marker — a
391
- scoped package the plan names that no repo manifest declares, a cross-repo
392
- `github.com/<owner>/<repo>` reference, or an endpoint named as a service
393
- prerequisite. That third trigger is what gives the default N=1 plan a cheap
394
- viability check, since the size trigger is unreachable at one ticket and this
395
- repo's resolved `riskHeuristics` is empty. The probe is conservative — explicit
396
- markers only, so a plan naming no such artifact dispatches exactly as before.
306
+ It fires on one deterministic trigger: the **external-dependency** probe
307
+ finding an out-of-repo marker — a scoped package the plan names that no repo
308
+ manifest declares, a cross-repo `github.com/<owner>/<repo>` reference, or an
309
+ endpoint named as a service prerequisite. The probe is conservative —
310
+ explicit markers only. It exits 0 on **any** verdict (verdicts route work,
311
+ they do not gate) and exits **1** only on a usage/IO error — no critic ran.
397
312
 
398
313
  ```jsonc
399
314
  {
400
- "consolidation": { "critic": "consolidation", "dispatch": false, "reasons": ["…"] },
401
- "premortem": { "critic": "pre-mortem", "dispatch": true, "reasons": ["…"] },
402
- "textHygiene": { "critic": "text-hygiene", "findings": [] }
315
+ "premortem": { "critic": "pre-mortem", "dispatch": true, "reasons": ["…"] }
403
316
  }
404
317
  ```
405
318
 
406
- The verdict's third entry, `textHygiene`, is advisory-only: it
407
- carries deterministic body lints (`dangling-citation` / `open-question` /
408
- `slicing-mass`) with no dispatch semantics — it spawns nothing and never
409
- gates the run. Fold `textHygiene.findings[]` into the re-author round the
410
- same way critic findings fold in: fix each named defect in `stories.json`
411
- (anchor or inline the citation, resolve the question into a declarative
412
- assumption, thin the Slicing checkpoint) and re-run the critic step. Empty
413
- `findings` add nothing to the round.
414
-
415
- **Dispatch shape.** When `delivery.routing.roleScopedAgents` is enabled (the
416
- **default**), dispatch each firing critic with `subagent_type: plan-critic` —
417
- it boots on the role-scoped [`plan-critic`](../../agents/plan-critic.md)
418
- context (its own system prompt, no `CLAUDE.md` @-closure) that carries the
419
- maker-blind invariant, the `consolidation` and `pre-mortem` charters, and the
420
- output shape standalone. When the kill-switch is off
421
- (`roleScopedAgents: false`) or the host cannot spawn at this depth, fall back
422
- to a generic sub-agent and hand it the same charter (the `consolidation` /
423
- `pre-mortem` definitions in [`plan-critic.md`](../../agents/plan-critic.md)).
424
- **When both critics fire, dispatch them in a single turn.** Consolidation and
425
- pre-mortem read the same immutable draft, share no write path, and neither
426
- consumes the other's verdict — the textbook independent fan-out of
427
- [`parallel-tooling.md`](parallel-tooling.md) Rule 3. Issue both `Agent` calls
428
- together in one assistant turn rather than awaiting the first verdict before
429
- spawning the second; serialized critics double the round's wall clock and buy
430
- nothing, because you fold both verdicts into the same re-author round anyway.
431
-
432
- Either way the critic is **maker-blind**: hand it the draft artifacts
433
- (`stories.json`, and `techspec.md` when present) — never the authoring
434
- transcript or the reasons the planner believed its own draft is sound. A
435
- critic that reads the maker's case grades the case, not the draft.
319
+ On `dispatch: true`, dispatch **one fresh-context, maker-blind sub-agent**.
320
+ When `delivery.routing.roleScopedAgents` is enabled (the **default**), use
321
+ `subagent_type: plan-critic` — it boots on the role-scoped
322
+ [`plan-critic`](../../agents/plan-critic.md) context (its own system prompt,
323
+ no `CLAUDE.md` @-closure) that carries the maker-blind invariant, the
324
+ `pre-mortem` charter, and the output shape standalone. When the kill-switch
325
+ is off (`roleScopedAgents: false`) or the host cannot spawn at this depth,
326
+ fall back to a generic sub-agent and hand it the same charter. Either way the
327
+ critic is **maker-blind**: hand it the draft artifacts (`stories.json`, and
328
+ `techspec.md` when present) — never the authoring transcript or the reasons
329
+ the planner believed its own draft is sound. A critic that reads the maker's
330
+ case grades the case, not the draft. Fold surviving findings into Gate #2 or
331
+ a re-author round.
436
332
 
437
333
  ## What `--dry-run` actually gates
438
334
 
439
335
  `plan-persist.js --dry-run` is the same command with GitHub writes suppressed,
440
- and every gate runs before the first `createIssue` would fire — the validator,
441
- the body parse, the DAG, the capacity and Spec-budget ceilings, the
442
- reachability check, the split and supersede partitions, and the Tech Spec fold.
443
- That is the whole point of running it first: a dry run that comes back clean
444
- has already paid for every deterministic refusal, so the real persist has
445
- nothing left to discover except network failure.
336
+ and every gate runs before the first `createIssue` would fire. Since
337
+ Story #5312 the gates split two ways, and the dry-run is where the second
338
+ half is read:
339
+
340
+ **Hard — the run refuses:** a body that does not parse, a ticket that is not
341
+ a Story, an empty `acceptance[]` or `verify[]`, an unknown or cyclic
342
+ `depends_on`, the acceptance partition at N>1, the supersede partition, a
343
+ forbidden commit-subject prefix, and a `deletes` entry naming a path absent
344
+ at base.
345
+
346
+ **Warnings — listed, then the persist proceeds:** a `creates` on a path that
347
+ exists at base or a `refactors-existing` on one that does not (including a
348
+ path the base branch deleted or renamed, named with the removing commit), a
349
+ goal or acceptance path absent at base, a `verify[]` command naming an absent
350
+ test file, and an `open-question` in a body (`Flag if…`, `TBD`, a trailing
351
+ `?`). The list also names every `changes[]` **repair** the run applied — a
352
+ plain-string bullet or a trailing parenthetical rewritten into
353
+ `{ path, assumption }` by probing base. The same list rides the result
354
+ envelope as `warnings[]` and `repairs[]`, so a `--chain-on-clean` run loses
355
+ nothing.
356
+
357
+ A dry run that comes back clean has paid for every deterministic refusal, so
358
+ the real persist has nothing left to discover except network failure.
446
359
 
447
360
  ## The container Epic (Gate #3)
448
361