mandrel 2.55.0 → 2.57.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. package/.agents/agents/plan-critic.md +13 -18
  2. package/.agents/agents/story-worker.md +25 -34
  3. package/.agents/docs/agentrc-reference.json +4 -30
  4. package/.agents/docs/configuration.md +11 -28
  5. package/.agents/docs/execution-reference.md +5 -5
  6. package/.agents/docs/quality-gates.md +8 -7
  7. package/.agents/instructions.md +9 -10
  8. package/.agents/rules/ci-remediation.md +39 -21
  9. package/.agents/schemas/agentrc.schema.json +28 -185
  10. package/.agents/schemas/story-deliver-terminal.schema.json +1 -1
  11. package/.agents/scripts/acceptance-eval.js +107 -17
  12. package/.agents/scripts/audit-to-stories.js +222 -75
  13. package/.agents/scripts/ceremony-derive.js +191 -0
  14. package/.agents/scripts/check-context-budget.js +28 -33
  15. package/.agents/scripts/check-cyclomatic.js +4 -3
  16. package/.agents/scripts/deliver-light.js +31 -94
  17. package/.agents/scripts/file-ci-gap.js +306 -0
  18. package/.agents/scripts/lib/audit-suite/checklist-threading.js +15 -2
  19. package/.agents/scripts/lib/audit-to-stories/audit-label-taxonomy.js +25 -1
  20. package/.agents/scripts/lib/audit-to-stories/dedupe-against-github.js +40 -52
  21. package/.agents/scripts/lib/audit-to-stories/finding-adapter.js +5 -1
  22. package/.agents/scripts/lib/audit-to-stories/issue-corpus.js +162 -0
  23. package/.agents/scripts/lib/audit-to-stories/issues-file.js +121 -0
  24. package/.agents/scripts/lib/audit-to-stories/ledger-commit.js +1 -1
  25. package/.agents/scripts/lib/audit-to-stories/ledger-record.js +126 -0
  26. package/.agents/scripts/lib/audit-to-stories/seed-from-findings.js +11 -0
  27. package/.agents/scripts/lib/baselines/coverage-updater-cli.js +110 -0
  28. package/.agents/scripts/lib/baselines/crap-preview-scan.js +25 -0
  29. package/.agents/scripts/lib/baselines/crap-updater-cli.js +223 -0
  30. package/.agents/scripts/lib/bdd-scenario-budget.js +21 -3
  31. package/.agents/scripts/lib/bootstrap/quality-bootstrap.js +0 -1
  32. package/.agents/scripts/lib/close-validation/gates.js +52 -1
  33. package/.agents/scripts/lib/config/acceptance-eval.js +25 -57
  34. package/.agents/scripts/lib/config/delivery-routing.js +7 -33
  35. package/.agents/scripts/lib/config/explain.js +0 -19
  36. package/.agents/scripts/lib/config/limits.js +18 -78
  37. package/.agents/scripts/lib/config/quality.js +6 -3
  38. package/.agents/scripts/lib/config/runners.js +3 -2
  39. package/.agents/scripts/lib/config-settings-schema-delivery.js +15 -68
  40. package/.agents/scripts/lib/config-settings-schema-quality.js +0 -14
  41. package/.agents/scripts/lib/config-settings-schema.js +49 -143
  42. package/.agents/scripts/lib/crap-engine.js +35 -4
  43. package/.agents/scripts/lib/crap-utils.js +17 -1
  44. package/.agents/scripts/lib/cyclomatic-ceiling.js +19 -7
  45. package/.agents/scripts/lib/feedback-loop/graduator-core.js +53 -13
  46. package/.agents/scripts/lib/feedback-loop/prior-feedback-fetcher.js +71 -25
  47. package/.agents/scripts/lib/feedback-loop/retro-proposals-graduator.js +18 -25
  48. package/.agents/scripts/lib/{audit-to-stories/ledger.js → findings/audit-ledger.js} +131 -24
  49. package/.agents/scripts/lib/findings/route-finding.js +38 -0
  50. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  51. package/.agents/scripts/lib/github/framework-repo.js +148 -2
  52. package/.agents/scripts/lib/label-constants.js +6 -1
  53. package/.agents/scripts/lib/observability/runtime-friction.js +1 -1
  54. package/.agents/scripts/lib/observability/source-classifier.js +2 -0
  55. package/.agents/scripts/lib/orchestration/acceptance-eval-decision.js +5 -4
  56. package/.agents/scripts/lib/orchestration/ceremony-routing.js +19 -73
  57. package/.agents/scripts/lib/orchestration/ci-gap-intake.js +605 -0
  58. package/.agents/scripts/lib/orchestration/ci-rerun-guard.js +13 -8
  59. package/.agents/scripts/lib/orchestration/complexity-gate.js +46 -212
  60. package/.agents/scripts/lib/orchestration/file-assumptions.js +32 -17
  61. package/.agents/scripts/lib/orchestration/light-escalation.js +3 -3
  62. package/.agents/scripts/lib/orchestration/light-suitability.js +66 -233
  63. package/.agents/scripts/lib/orchestration/plan-context.js +181 -387
  64. package/.agents/scripts/lib/orchestration/plan-critic-conditions.js +42 -153
  65. package/.agents/scripts/lib/orchestration/plan-critics-evaluate.js +14 -70
  66. package/.agents/scripts/lib/orchestration/plan-persist/audit-provenance.js +197 -0
  67. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +300 -0
  68. package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +131 -168
  69. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +133 -299
  70. package/.agents/scripts/lib/orchestration/plan-persist/soft-findings.js +55 -0
  71. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +16 -65
  72. package/.agents/scripts/lib/orchestration/plan-persist/wave-serialisation.js +22 -35
  73. package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +30 -139
  74. package/.agents/scripts/lib/orchestration/planning/memory-pool-advisory.js +61 -223
  75. package/.agents/scripts/lib/orchestration/run-epilogue.js +4 -4
  76. package/.agents/scripts/lib/orchestration/single-story-close/phases/close-validation.js +5 -0
  77. package/.agents/scripts/lib/orchestration/single-story-close/phases/pre-gate-steps.js +46 -16
  78. package/.agents/scripts/lib/orchestration/story-close/context-budget-writeback.js +213 -0
  79. package/.agents/scripts/lib/orchestration/story-follow-ups.js +32 -20
  80. package/.agents/scripts/lib/orchestration/task-body-validator.js +10 -63
  81. package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +33 -539
  82. package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +21 -414
  83. package/.agents/scripts/lib/orchestration/ticket-validator.js +54 -118
  84. package/.agents/scripts/lib/orchestration/verify-credit.js +69 -24
  85. package/.agents/scripts/lib/story-body/body-format-lints.js +15 -85
  86. package/.agents/scripts/lib/story-body/story-body.js +17 -237
  87. package/.agents/scripts/lib/templates/decomposer-prompts.js +84 -121
  88. package/.agents/scripts/lib/test-isolate/cli-options.js +93 -0
  89. package/.agents/scripts/lib/test-isolate/progress-log.js +45 -0
  90. package/.agents/scripts/lib/test-isolate/render-report.js +97 -0
  91. package/.agents/scripts/lib/test-isolate/run-isolate.js +87 -0
  92. package/.agents/scripts/lib/test-run-credit.js +266 -0
  93. package/.agents/scripts/lib/wave-runner/footprint.js +48 -358
  94. package/.agents/scripts/lib/wave-runner/ready-set.js +6 -5
  95. package/.agents/scripts/lib/workers/crap-worker.js +32 -41
  96. package/.agents/scripts/plan-context.js +7 -9
  97. package/.agents/scripts/plan-critics.js +28 -54
  98. package/.agents/scripts/plan-persist.js +25 -68
  99. package/.agents/scripts/pr-watch-with-update.js +3 -2
  100. package/.agents/scripts/quality-preview.js +51 -0
  101. package/.agents/scripts/run-tests.js +12 -0
  102. package/.agents/scripts/stories-wave-tick.js +23 -45
  103. package/.agents/scripts/test-isolate.js +13 -180
  104. package/.agents/scripts/update-coverage-baseline.js +25 -70
  105. package/.agents/scripts/update-crap-baseline.js +19 -123
  106. package/.agents/skills/core/scope-triage/SKILL.md +3 -3
  107. package/.agents/workflows/audit-clean-code.md +4 -3
  108. package/.agents/workflows/audit-to-stories.md +63 -27
  109. package/.agents/workflows/helpers/acceptance-self-eval.md +41 -41
  110. package/.agents/workflows/helpers/code-quality-guardrails.md +4 -4
  111. package/.agents/workflows/helpers/code-review.md +2 -3
  112. package/.agents/workflows/helpers/deliver-digest.md +41 -57
  113. package/.agents/workflows/helpers/deliver-light.md +40 -105
  114. package/.agents/workflows/helpers/deliver-reference.md +1 -1
  115. package/.agents/workflows/helpers/deliver-story-reference.md +56 -62
  116. package/.agents/workflows/helpers/deliver-story.md +9 -13
  117. package/.agents/workflows/helpers/plan-reference.md +132 -196
  118. package/.agents/workflows/mandrel-plan.md +28 -41
  119. package/.agents/workflows/memory-consolidate.md +9 -13
  120. package/docs/CHANGELOG.md +33 -0
  121. package/lib/migrations/index.js +4 -0
  122. package/lib/migrations/steps/2.57.0-retire-delivery-limit-knobs.js +45 -0
  123. package/lib/migrations/steps/2.57.0-retire-planning-limit-knobs.js +59 -0
  124. package/package.json +1 -1
  125. package/.agents/scripts/lib/framework-version.js +0 -39
  126. package/.agents/scripts/lib/orchestration/consolidation-precondition.js +0 -223
  127. package/.agents/scripts/lib/orchestration/plan-persist/fan-out-gate.js +0 -97
  128. package/.agents/scripts/lib/orchestration/planning/decomposer-context.js +0 -26
  129. package/.agents/scripts/lib/orchestration/spec-budget.js +0 -89
  130. package/.agents/scripts/lib/orchestration/spec-spill.js +0 -74
  131. package/.agents/scripts/lib/orchestration/verify-tier-repair.js +0 -107
@@ -140,10 +140,8 @@ full-ceremony Story. The lite route's `preserves` field is the machine-readable
140
140
  record of those non-negotiables; there is no lite-specific gate bypass.
141
141
 
142
142
  **Ceremony comes from the landed diff; the dispatch mode comes from the
143
- run.** Persist stamps a lite cohort's Stories with the `route::lite` label as a
144
- _human-visible hint only_ (and ledgers the authored verdict — recorded reason
145
- plus per-Story shape evidence — on the `story-plan-state` checkpoint); the
146
- label is never the control signal. Ceremony is resolved from the **derived
143
+ run.** Persist stamps no route label (Story #5312 retired the plan-side lite
144
+ claim with its `route::lite` hint). Ceremony is resolved from the **derived
147
145
  change level** (`deriveChangeLevel` over the computed change set — digest § 3),
148
146
  not from a body-shape read: a footprint intersecting a sensitive-path class
149
147
  derives `high`, so the Story keeps its fresh acceptance critic. The light path
@@ -240,35 +238,19 @@ discovers them only after the whole close pipeline has run, at several times
240
238
  the cost of one full-suite run in the worktree.
241
239
 
242
240
  **Run it once, last, so close can credit it.** The run belongs **after** the
243
- self-eval loop's last fix commit and **after** the hand-off push — the credit
244
- is keyed on the tree, not on push state, so pushing first keeps the stamp and
245
- buys the ordering Step 2.5 needs (the capture is backgrounded, and its
246
- completion ends the turn); redraft rounds run scoped tests. Close skips a gate that already passed at the current HEAD, but a bare
247
- `npm test` deposits no such record — the suite then runs twice per delivery,
248
- once here and once in the close gate chain. Pick the invocation by the same
249
- predicate `close-validation/gates.js` uses to choose its test gate:
250
-
251
- ```bash
252
- # CRAP gate enabled (default) + a `test:coverage` script — writes the stamp
253
- # the close `coverage-capture` gate reads:
254
- node <main-repo>/.agents/scripts/coverage-capture.js --cwd <workCwd>
255
- # otherwise — the evidence record the close `test` gate reads. <workCwd> must
256
- # be ABSOLUTE and the runner exactly `npm test`: both sides hash
257
- # {cmd, args, cwd}, so a relative path or a wrapper misses the credit.
258
- node <main-repo>/.agents/scripts/evidence-gate.js --standalone \
259
- --scope-id <storyId> --gate test --worktree <workCwd> -- npm test
260
- ```
261
-
262
- The credit expires the moment it stops describing the tree: evidence is keyed
263
- on HEAD, the capture stamp on a content digest of `crap.targetDirs`. A
264
- self-eval fix — or any commit — invalidates it and close re-runs the suite for
265
- real, so this never trades away the gate. That keying is exactly why the run
266
- comes last, and why the push before it is free. Close's own base-sync can
267
- spend the stamp too when it lands base commits; it now says so out loud rather
268
- than silently re-running the suite.
269
-
270
- **`verify[]` reuses the same stamp.** A `verify[]` entry that is itself a
271
- full-suite command is reported **credited** against that stamp rather than
241
+ self-eval loop's last fix commit; redraft rounds run scoped tests. A green
242
+ `npm test` in the worktree on `story-<id>` deposits the `test` evidence
243
+ record close reads (Story #5313 — `lib/test-run-credit.js`), keyed on HEAD
244
+ and the tree fingerprint and hashed on the exact command close spawns, so
245
+ close reports the gate as credited at unchanged HEAD instead of re-running
246
+ the suite. The credit expires the moment it stops describing the tree: any
247
+ later commit invalidates it and close re-runs the suite for real, so this
248
+ never trades away the gate. The CRAP gate still runs `coverage-capture.js`
249
+ itself when it needs a fresh artifact — the capture stamp is a claim about
250
+ `coverage/coverage-final.json`, which a bare `npm test` does not produce.
251
+
252
+ **`verify[]` reuses the same credit.** A `verify[]` entry that is itself a
253
+ full-suite command is reported **credited** against that record rather than
272
254
  respawned (`resolveVerifyCredit` in
273
255
  [`verify-credit.js`](../../scripts/lib/orchestration/verify-credit.js)), and the
274
256
  self-eval gate warns when it sees one: the intended shape is scoped `verify[]`
@@ -294,24 +276,23 @@ round cap, proceed / redraft / block — not an independent additional pass
294
276
  over the criteria. The M4-B floor holds: one verdict per cluster, with the
295
277
  cluster count owned by the dispatching caller and never by routing.
296
278
 
297
- **One round = N cluster critics → ONE merged verdict → ONE gate call.** The
279
+ **One round = N cluster critics → ONE merged verdict → ONE gate call** (fresh
280
+ critics; an inline-owned verdict is one file scored in one call). The
298
281
  clusters are how a round is _authored_; they are not how it is _scored_.
299
282
  Concatenate every cluster's records into a single `criteria[]` ordered by
300
283
  `index` — exactly one per `acceptance[]` item, under one `storyId`,
301
284
  `schemaVersion`, `round` and `commitSha` — and hand that merged file to the
302
- gate once, with `--expected-criteria` set to the Story's `acceptance[]` count:
285
+ gate once:
303
286
 
304
287
  ```bash
305
288
  node <main-repo>/.agents/scripts/acceptance-eval.js \
306
- --story <storyId> --verdict <merged-verdict-path> \
307
- --expected-criteria <acceptance[] count>
289
+ --story <storyId> --verdict <merged-verdict-path>
308
290
  ```
309
291
 
310
- The flag is what makes the merge enforceable: `assertCriteriaCoverage` returns
311
- early on the `null` default, so **omitting it leaves the guard inert** and a
312
- single cluster's verdict handed over unmerged scores a fraction of the criteria
313
- and still reports `proceed`. A length mismatch is rejected before scoring and
314
- consumes no round. Calling the gate once per cluster instead spends a round
292
+ The gate reads the Story's `acceptance[]` count itself (Story #5313), so a
293
+ single cluster's verdict handed over unmerged is rejected before scoring and
294
+ consumes no round; `--expected-criteria` is accepted but redundant. Calling
295
+ the gate once per cluster instead spends a round
315
296
  _per cluster_ — a Story past the cluster ceiling would burn its whole redraft
316
297
  budget on cluster arithmetic — and N concurrent calls race the Story-scoped
317
298
  round ledger. Full per-round mechanics, including the parallel dispatch and the
@@ -342,32 +323,30 @@ node .agents/scripts/update-ticket-state.js --ticket <storyId> --state agent::bl
342
323
 
343
324
  ## Step 2 — Ceremony detail
344
325
 
345
- **Compute the change set once** with the shared enumerator —
346
- the same module close uses — and reuse that one list downstream:
326
+ **Compute the change set once** with the shared enumerator — the same module
327
+ close uses — and reuse that one list downstream. `ceremony-derive.js`
328
+ (Story #5313) is that enumeration, the level derivation and the ceremony
329
+ resolution in one call:
347
330
 
348
331
  ```bash
349
- node --input-type=module -e '
350
- import { computeChangeSet } from "<main-repo>/.agents/scripts/lib/orchestration/change-set.js";
351
- const { files } = computeChangeSet({ baseRef: "main", headRef: "story-<storyId>" });
352
- console.log(JSON.stringify(files));
353
- '
332
+ node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
354
333
  ```
355
334
 
356
- Derive the level with
335
+ Its `level` comes from
357
336
  [`deriveChangeLevel`](../../scripts/lib/orchestration/review-depth.js) over
358
337
  the one computed change-set list: a diff touching a sensitive path registered
359
338
  in `.agents/schemas/audit-rules.json` derives `high`, one touching none
360
- derives `low`, and an unenumerable diff (`files === null`) derives `null`.
361
- Hand the **same** list to every acceptance critic you spawn (Step 1a) — a
362
- critic that re-ran its own `git diff` could score against a different set
363
- than the one that routed it.
339
+ derives `low`, and an unenumerable diff (`files: null`) derives `null`.
340
+ Hand the **same** `files` list to every acceptance critic you spawn (Step 1a)
341
+ — a critic that re-ran its own `git diff` could score against a different
342
+ set than the one that routed it.
364
343
 
365
- Resolve fresh-vs-inline acceptance critics per AC-cluster with
344
+ Its `mode` / `verdictOwner` come from
366
345
  [`resolveCeremonyForRisk`](../../scripts/lib/orchestration/ceremony-routing.js)
367
346
  (`minimal` → always inline; `strict` → always fresh; `standard` →
368
- `high`/`null` → `fresh`, `low` → `inline` unless the `freshCriticSampleRate`
369
- floor forces `fresh`). Review depth reads the same derived level via
370
- `review-depth.js` inside close, so the two decisions cannot disagree.
347
+ `high`/`null` → `fresh`, `low` → `inline`). Review depth reads the same
348
+ derived level via `review-depth.js` inside close, so the two decisions cannot
349
+ disagree.
371
350
 
372
351
  **Inline-dispatch override.** When the Story dispatches
373
352
  `inline` (`resolveStoryDispatchMode` → `inline`, which is exactly a
@@ -682,13 +661,28 @@ When the watch exits, branch on the exit code:
682
661
 
683
662
  **Triage authority.** How to classify and remediate a red (or repeatedly slow)
684
663
  check — the root-cause-only decision tree for infra/transient and flaky failures
685
- (reproduce → check `main` → bisect env vs code → fix in-scope or file a
686
- `meta::framework-gap` issue), the never-rerun / never-quarantine prohibitions,
687
- and the escalation criteria (three-strikes, the 30-minute wall-clock timebox,
688
- and the clearly-environmental fast path) — is defined once in
664
+ (reproduce → check `main` → bisect env vs code → fix in-scope, or reach an
665
+ Option-2 verdict and file the intake issue), the never-rerun / never-quarantine
666
+ prohibitions, and the escalation criteria (three-strikes, the 30-minute
667
+ wall-clock timebox, and the clearly-environmental fast path) — is defined once in
689
668
  [`.agents/rules/ci-remediation.md`](../../rules/ci-remediation.md). Read it
690
669
  before remediating a red check.
691
670
 
671
+ **Filing an out-of-scope root cause is one command, never a hand-run `gh issue
672
+ create`:**
673
+
674
+ ```bash
675
+ node <agentRoot>/scripts/file-ci-gap.js --story <storyId> \
676
+ --verdict <pre-existing|capacity|unreproducible-tier> \
677
+ --owner <consumer|framework|platform> --evidence "<proof reading>" [--block]
678
+ ```
679
+
680
+ It reads the digest, routes the filing to the repo that owns the fault, updates
681
+ the existing ticket when the signature is a repeat, posts the `friction`
682
+ comment, and with `--block` flips the Story. What it files is an **intake**
683
+ issue, not a Story — `/mandrel-plan <issue number>` graduates it on the next
684
+ planning pass, so the delivery never waits on planning.
685
+
692
686
  ### The auto-merge wait is an internally-blocking step
693
687
 
694
688
  This is the single most important contract of this workflow, and the seam
@@ -74,7 +74,7 @@ One branch, one PR to `main`, commits against the inline `acceptance[]` /
74
74
  `## Slicing` rows as **intra-session checkpoints** (reference § Step 1).
75
75
  2. Implement and commit on the Story branch, iterating with quick advisory
76
76
  gates (`typecheck`, `lint`, scoped tests) — the full chain runs in Step 3,
77
- and the **one** creditable full-suite run at Step 2.5.
77
+ and the **one** full-suite run at Step 2.5.
78
78
 
79
79
  ### Step 1a — Bounded acceptance self-eval loop (**required**)
80
80
 
@@ -93,21 +93,17 @@ are reference § Step 2. Hard gates always run in Step 3 — the derived level
93
93
  never disables them; do **not** pre-run the chain here — Step 2.5's credited
94
94
  suite run is the sole exception.
95
95
 
96
- ### Step 2.5 — Push, then the creditable full-suite run, then hand off
96
+ ### Step 2.5 — The one full-suite run, the push, then hand off
97
97
 
98
- **Push first.** After the self-eval loop's last fix commit, push
99
- `story-<storyId>` to `origin`, confirming the remote ref moved: the capture
100
- below is backgrounded, so *its* completion ends the turn.
98
+ After the self-eval loop's last fix commit, run `npm test` **once** in the
99
+ worktree (**digest § 5**): a green full run deposits the `test` credit close
100
+ reads, keyed on the tree, so only a *later* commit invalidates it. Red →
101
+ fix, commit, re-run.
101
102
 
102
- Then run the full suite **once**, after the push, in the shape close credits
103
- (**digest § 5**): the credit is keyed on the tree, not on push state, so only
104
- a *later* commit invalidates it; a bare `npm test` deposits none. Red →
105
- fix, commit, push, re-capture.
106
-
107
- Then (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
103
+ Push `story-<storyId>` to `origin`, confirming the remote ref moved. Then
104
+ (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
108
105
  branch, pushed head SHA, self-eval verdict, `verify[]` evidence — and stop.
109
- Do not open the PR or compose a terminal envelope. An inline run captures
110
- before Step 3.
106
+ Do not open the PR or compose a terminal envelope.
111
107
 
112
108
  ## Step 3 — Close and land (`single-story-close.js`)
113
109
 
@@ -1,10 +1,10 @@
1
1
  # /mandrel-plan — on-demand reference appendix
2
2
 
3
3
  > **Applies when:** you are executing [`/mandrel-plan`](../mandrel-plan.md) and hit one of the
4
- > situations below — input-mode derivation, the Gate #1 light handoff,
5
- > shape-derived complexity routing, tickets-mode supersede authoring, critic
6
- > dispatch detail, a failed persist, or source-id resolution. The spine stays
7
- > resident; this file is read on demand.
4
+ > situations below — input-mode derivation, the Gate #1 advisory line,
5
+ > tickets-mode supersede authoring, the operator-invoked pre-mortem, a
6
+ > failed persist, or source-id resolution. The spine stays resident; this
7
+ > file is read on demand.
8
8
 
9
9
  ## Deriving the input mode
10
10
 
@@ -92,117 +92,58 @@ The marker keeps the operator's undelegated decisions findable after the
92
92
  fact: reviewing a `--yes` plan means scanning its decisions-made-by-default,
93
93
  not re-deriving which assumptions were really the agent's to make.
94
94
 
95
- ## Gate #1 → the memory-pool advisory (`memoryPoolAdvisory`)
96
-
97
- On a truthy `memoryPoolAdvisory.recommend`, name
98
- [`/memory-consolidate`](../memory-consolidate.md) at Gate #1, quoting its
99
- `reasons[]`. Purely advisory: a stale pool degrades recall, it does not make
100
- the plan wrong, so it never blocks and never reroutes.
101
-
102
- ## Gate #1 → the light path (in-session handoff)
103
-
104
- On a confirmed `deliverLightSuggestion`, `/mandrel-plan` routes into
105
- [`deliver-light.md`](deliver-light.md) **without ending the session**. Two
106
- things make that safe, and both are worth understanding before changing it:
107
-
108
- 1. **The handoff carries the envelope, not the seed.** Gate #1 already holds
109
- the interrogated `complexitySignals`; fill the light gate's `--creates`
110
- / `--refactors` / `--acceptance` / `--reason` from those. Re-deriving from
111
- raw seed text throws away the better signal and can disagree with the
112
- suggestion that routed you.
113
- 2. **The gate still runs.** The suggestion is read against seed-time ceilings
114
- (`DELIVER_LIGHT_SUGGESTION_CEILINGS` — artifacts, risk hits, sensitive-path
115
- classes); the light gate is read against the predicted work's effort and risk
116
- (`STORY_SHAPE_CEILINGS` — change kinds, magnitude, uncertainty, deployable
117
- span). Two different checks on purpose, so a confirm is not a bypass.
118
-
119
- **When the light gate answers `ask-operator`**, the two ceiling sets disagreed.
120
- Resume `/mandrel-plan` at step 2 (Author) **in this same session** — the interrogation
121
- is still valid and re-paying for it buys nothing. This bounce-back is not an
122
- escalation.
123
-
124
- **Under `--yes` the offer is recorded and planning proceeds** — it is *never*
125
- auto-downgraded to light. An unattended run has nobody to confirm the reroute,
126
- and a suggestion is not a confirmation. The same rule governs unknown triage
127
- unattended: AFK unknowns are still researched, but no free-form operator
128
- question is asked — each HITL unknown lands in Key Assumptions marked a
129
- decision-made-by-default, so the record shows what was decided for the operator
130
- rather than pretending it was decided with them.
131
-
132
- Escalation in the *other* direction — an over-scope prompt on the light path —
133
- is terminal and requires a fresh session. The rule that separates the two, and
134
- why it must not be flattened into symmetry:
135
- [`deliver-light.md` § Why the two directions differ](deliver-light.md).
136
-
137
- ## Gate #1 → the `/prototype` offer (`uiSurface`)
138
-
139
- `complexitySignals.uiSurface` is the second advisory Gate #1 offer, and the
140
- weaker of the two on purpose: it carries **no routing authority and adds no
141
- gate**. Both halves are derived from observables already in the checkout — the
142
- `hasWebSurface` applicability predicate the `target: "web"` audit lenses gate
143
- on, and whether any predicted path matches a web lens `filePattern` registered
144
- in `audit-rules.json`. There is no configuration key to set: a project with no
145
- rendered frontend resolves falsey and the offer never fires.
146
-
147
- When it does fire, **name [`/prototype`](../prototype.md) and stop there.**
148
- `/mandrel-plan` must never invoke it — operator invocation is the entire design, because
149
- the value is a human looking at a layout before its UI acceptance criteria are
150
- frozen.
151
-
152
- **Under `--yes` the offer is recorded and planning proceeds** — no reroute, no
153
- prototype written, no gate raised. This is exactly how `deliverLightSuggestion`
154
- behaves unattended, and for the same reason: an unattended run has nobody to
155
- review an artifact, so recording the offer is the whole of the right behaviour.
156
-
157
- ## Shape-derived complexity routing (`complexitySignals`)
158
-
159
- Complexity routes on the **objective shape of the authored work**, never on
160
- seed word count — a detailed prompt can describe trivial work, a terse one
161
- complex work. The pipeline stages the
162
- decision:
163
-
164
- - **Signals, not routing.** The envelope's `complexitySignals` field is
165
- advisory only (`routingAuthority: false`): enumerated-artifact count (with
166
- the configured `maxArtifacts` threshold beside it as one input),
167
- `planning.riskHeuristics` phrases present in the seed, the repo state of
168
- predicted paths (existing paths predict refactors; missing predict
169
- creates), and the `audit-rules.json` sensitive-path classes the predicted
170
- footprint intersects.
171
- - **You author the verdict.** Judge the signals: a genuinely trivial scope
172
- (small additive footprint, no risk hits, no sensitive class) earns a `lite`
173
- claim via `plan-persist.js --route-downgrade-reason "<why>"`. The reason is
174
- recorded on every created Story's `story-plan-state` checkpoint, making the
175
- judgment auditable; without a recorded reason the conservative default
176
- (`full`) stands.
177
- - **Persist backstops the claim deterministically.** After authoring, the
178
- work has measurable shape, so persist validates the `lite` claim against
179
- each Story's own shape — distinct change kinds, declared magnitude,
180
- uncertainty, deployable/migration span, glob-free footprint, and
181
- sensitive-path classes, against the framework `STORY_SHAPE_CEILINGS` (effort
182
- and risk, never artifact counts) — and **fails closed to
183
- `full`** when any Story exceeds them (the refusal is ledgered on the
184
- checkpoint too). The lite route is **not** licence to drop a
185
- non-negotiable — every decision's `preserves` field enumerates what still
186
- holds: the Story ticket, the PR-to-`main` landing, every repo quality gate,
187
- and the security baseline. Those gates run in `single-story-close.js`
188
- regardless of route.
189
-
190
- **The label is a hint; deliver re-derives.** Persist labels a
191
- lite cohort's Stories with **`route::lite`** as a *human-visible hint only* —
192
- `/mandrel-deliver` computes the route from each fetched Story body via the same shape
193
- function at dispatch, so neither a lost label nor an unread marker can
194
- misroute delivery: a lite-shaped Story derives `lite` even with the label
195
- absent, and a sensitive-footprint Story routes `full` and keeps its fresh
196
- critic even with the label present. The derived route sets ceremony, not
197
- where the engine runs — sub-agent boots are collapsed by a **single-Story
198
- run**, never by a trivial shape. The `route::*` axis stays runtime-derived: hand-authored
199
- `route::*` entries in `labels[]` are dropped by persist.
200
-
201
- The knobs (`planning.complexityGate.{enabled, maxArtifacts}`) are documented
202
- in [`.agents/docs/configuration.md`](../../docs/configuration.md) under
203
- `### planning`; the defaults live on `DEFAULT_COMPLEXITY_GATE` and the shape
204
- ceilings on `STORY_SHAPE_CEILINGS` in
205
- [`lib/orchestration/complexity-gate.js`](../../scripts/lib/orchestration/complexity-gate.js).
95
+ ## Gate #1 → the one advisory line
96
+
97
+ Gate #1 stops for exactly two things — the sharpened plan intent and any HITL
98
+ unknown — and everything else the envelope surfaced collapses to **one
99
+ advisory line** beneath it (Story #5312). Nothing on that line stops the run,
100
+ reroutes it, or is invoked by `/mandrel-plan`; each item names something the
101
+ operator may prefer to do instead, and the run proceeds either way. Under
102
+ `--yes` the line is recorded and planning continues — an unattended run has
103
+ nobody to take an offer.
104
+
105
+ The line names, in order, whichever of these the envelope carries:
106
+
107
+ - **`duplicates[]`** — open Stories the seed resembles (never Epics). Name
108
+ the top one or two by id and title; a plan that duplicates open work is
109
+ still the operator's call.
110
+ - **Open `intake` rows** (`priorFeedback`) — CI-gap intake filings written by
111
+ [`file-ci-gap.js`](../../scripts/file-ci-gap.js) when a delivery reached an
112
+ Option-2 verdict in [`ci-remediation.md`](../../rules/ci-remediation.md).
113
+ They carry evidence but no `## Spec`, no `acceptance[]` / `verify[]` and no
114
+ `agent::*` label, so `/mandrel-deliver` cannot take one: graduating it is
115
+ exactly **tickets mode** (`/mandrel-plan <issue number>`), and a filing
116
+ that keeps recurring (its `## Occurrences` table is the count) is often the
117
+ better next Story than the seed in front of you. A `platformGaps[]` row is
118
+ the same shape with a different owner.
119
+ - **`memoryPoolAdvisory.recommend`** — name
120
+ [`/memory-consolidate`](../memory-consolidate.md), quoting its
121
+ `reasons[]`. The one arm left measures the `MEMORY.md` index against the
122
+ harness's byte cap; a stale pool degrades recall, it does not make the plan
123
+ wrong.
124
+ - **`complexitySignals.uiSurface`** — name [`/prototype`](../prototype.md)
125
+ and stop there. The signal carries **no routing authority and adds no
126
+ gate**: both halves are derived from observables already in the checkout —
127
+ the `hasWebSurface` applicability predicate the `target: "web"` audit
128
+ lenses gate on, and whether any predicted path matches a web lens
129
+ `filePattern` in `audit-rules.json`; a project with no rendered frontend
130
+ resolves falsey and the offer never fires. `/mandrel-plan` must never invoke
131
+ it — operator invocation is the entire design, because the value is a human
132
+ looking at a layout before its UI acceptance criteria are frozen. Under
133
+ `--yes` the offer is recorded and planning proceeds — no reroute, no
134
+ prototype written, no gate raised.
135
+
136
+ ## `complexitySignals` are advisory
137
+
138
+ The envelope's `complexitySignals` field carries the paths the seed predicts,
139
+ their repo state (existing paths predict refactors; missing predict creates)
140
+ and the `audit-rules.json` sensitive-path classes the footprint intersects —
141
+ `routingAuthority: false`, no `route` field. They ground the authoring
142
+ template's pre-resolved `changes[]` and the `/prototype` offer, nothing else.
143
+ Story #5312 deleted the plan-side lite claim that used to read them
144
+ (`--route-downgrade-reason`, the persist shape backstop, the `route::lite`
145
+ hint): every Story lands through the same engine and the same close gates,
146
+ and ceremony is derived from the landed diff at close.
206
147
 
207
148
  ## Correct-by-construction authoring template
208
149
 
@@ -210,30 +151,24 @@ ceilings on `STORY_SHAPE_CEILINGS` in
210
151
  **correct-by-construction** skeleton, built from the same repo probe the
211
152
  `complexitySignals` ran:
212
153
 
213
- - **`verify[]` placeholders already end with a valid `(tier)` tag.** Keep
214
- every filled entry's trailing tag one of `(unit)` / `(contract)` /
215
- `(e2e)` / `(validate)` (or use the `manual:<reason>` escape) — a tierless
216
- entry is exactly the mechanical persist round-trip the template exists to
217
- prevent.
154
+ - **`verify[]` entries are commands.** There is no tier suffix and no
155
+ `manual:<reason>` escape (Story #5312): write the exact command or test
156
+ path the deliverer runs and the acceptance critic reads as evidence.
218
157
  - **`changes[]` arrive pre-resolved to creates-vs-refactors.** Every path
219
158
  the seed predicted is probed against the repo: an existing path is
220
159
  emitted with `assumption: "refactors-existing"`, a missing one with
221
- `assumption: "creates"`. Trust the pre-resolved assumption — verify
222
- against the repo before overriding one (authoring `creates` for a file
223
- that exists at base is a validator rejection). The persist gates stay
224
- authoritative: they probe the base branch ref, not the working tree.
225
- - **Keep `## Spec` near contract-level prose.** **Aim for ~250 words; an
226
- advisory warning fires past 350** (`SPEC_SOFT_WORD_BUDGET`). Two numbers,
227
- two jobs: ~250 is the authoring target — the nudge toward a contract-level
228
- Spec (interfaces, invariants, load-bearing constraints; no per-file
229
- behavior narration) — while 350 is the slacker threshold at which persist
230
- actually warns, so the warning marks a real outlier instead of ordinary
231
- variance. Neither fails the persist. The hard fail-closed ceiling
232
- (~1500 tokens, `spec-spill.js`) is unchanged.
160
+ `assumption: "creates"`. The persist gates stay authoritative — they probe
161
+ the base branch ref, not the working tree — but a `creates` on a path that
162
+ exists at base, or a `refactors-existing` on one that does not, is a
163
+ dry-run **warning**, not a rejection; only a `deletes` naming an absent
164
+ path is refused. A plain-string bullet or a trailing parenthetical is
165
+ repaired into the object form by probing base, and the repair is reported.
166
+ - **Keep `## Spec` at contract-level prose** — interfaces, invariants,
167
+ load-bearing constraints; no per-file behavior narration — and as long as
168
+ the work needs. There is no word or token budget.
233
169
 
234
170
  A faithfully-filled skeleton — placeholders replaced, pre-resolved entries
235
- kept, tags valid — passes the persist ticket validators with no
236
- round-trip.
171
+ kept — passes the persist ticket validators with no round-trip.
237
172
 
238
173
  ### Authored entry shape
239
174
 
@@ -313,18 +248,13 @@ together, and a path two same-wave Stories both write is exactly where that
313
248
  promise breaks. Promise and caveat belong on one durable surface — previously
314
249
  the caveat was a stderr warning nobody kept.
315
250
 
316
- Two `planning.*` knobs upgrade a conflict class from advisory to a hard
317
- refusal. **Both default to `false` and are documented, not recommended:**
318
-
319
- | Knob | Upgrades | Why it is off |
320
- | --- | --- | --- |
321
- | `planning.failOnSharedEditors` | `shared-editor` → `hard` | Co-editing one file is routine and often correct; the delivery scheduler already serializes file-overlapping Stories. |
322
- | `planning.requireExplicitCrossStoryDeps` | `implicit-cross-story-dep` → `hard` | Path references are matched by substring, so a legitimate mention in prose can read as a dependency. |
323
-
324
- Turn one on for a repo where the class is genuinely fatal; expect a refusal to
325
- name the Stories and the fix (a `depends_on` edge, or folding the shared edit
326
- into one Story). The sibling knobs `failOnRegistryConflicts`,
327
- `failOnMissingBddScaffold` and `failOnLargeFanOut` behave the same way.
251
+ Every conflict class is advisory (Story #5312 retired the
252
+ `planning.failOn*` / `requireExplicitCrossStoryDeps` upgrade knobs with the
253
+ registry and fan-out findings): co-editing one file is routine and often
254
+ correct — the delivery scheduler already serializes file-overlapping Stories —
255
+ and a path reference matched by substring can read as a dependency a prose
256
+ mention never meant. A finding names the Stories and the fix (a `depends_on`
257
+ edge, or folding the shared edit into one Story) for the operator to weigh.
328
258
 
329
259
  ## Tickets mode — authoring `supersedes[]`
330
260
 
@@ -359,67 +289,73 @@ total by default — an authored map is the only thing that can say
359
289
  `#11-#14 → #20` while `#15 → #21`, which a blanket "superseded by
360
290
  this plan-run" reference could not.
361
291
 
362
- ## Critic dispatch detail
292
+ ## The pre-mortem critic — operator-invoked
293
+
294
+ The maker-blind **pre-mortem** critic is not a step of the spine (Story #5312
295
+ retired step 2.5 with the consolidation critic, whose one deterministic input
296
+ was a `## Delivery Slicing` table no Story carries). Run it when the operator
297
+ asks for it, after Author and before Persist — the last point a finding folds
298
+ into a re-author:
299
+
300
+ ```bash
301
+ node .agents/scripts/plan-critics.js \
302
+ --stories temp/plan-<slug>/stories.json \
303
+ [--tech-spec temp/plan-<slug>/techspec.md]
304
+ ```
363
305
 
364
- The **pre-mortem** critic fires on any of three deterministic triggers: the
365
- draft ticket count reaching half the reviewability budget, a
366
- `planning.riskHeuristics` phrase matching the plan text, or the
367
- **external-dependency** probe finding an out-of-repo marker — a
368
- scoped package the plan names that no repo manifest declares, a cross-repo
369
- `github.com/<owner>/<repo>` reference, or an endpoint named as a service
370
- prerequisite. That third trigger is what gives the default N=1 plan a cheap
371
- viability check, since the size trigger is unreachable at one ticket and this
372
- repo's resolved `riskHeuristics` is empty. The probe is conservative — explicit
373
- markers only, so a plan naming no such artifact dispatches exactly as before.
306
+ It fires on one deterministic trigger: the **external-dependency** probe
307
+ finding an out-of-repo marker — a scoped package the plan names that no repo
308
+ manifest declares, a cross-repo `github.com/<owner>/<repo>` reference, or an
309
+ endpoint named as a service prerequisite. The probe is conservative —
310
+ explicit markers only. It exits 0 on **any** verdict (verdicts route work,
311
+ they do not gate) and exits **1** only on a usage/IO error — no critic ran.
374
312
 
375
313
  ```jsonc
376
314
  {
377
- "consolidation": { "critic": "consolidation", "dispatch": false, "reasons": ["…"] },
378
- "premortem": { "critic": "pre-mortem", "dispatch": true, "reasons": ["…"] },
379
- "textHygiene": { "critic": "text-hygiene", "findings": [] }
315
+ "premortem": { "critic": "pre-mortem", "dispatch": true, "reasons": ["…"] }
380
316
  }
381
317
  ```
382
318
 
383
- The verdict's third entry, `textHygiene`, is advisory-only: it
384
- carries deterministic body lints (`dangling-citation` / `open-question` /
385
- `slicing-mass`) with no dispatch semantics — it spawns nothing and never
386
- gates the run. Fold `textHygiene.findings[]` into the re-author round the
387
- same way critic findings fold in: fix each named defect in `stories.json`
388
- (anchor or inline the citation, resolve the question into a declarative
389
- assumption, thin the Slicing checkpoint) and re-run the critic step. Empty
390
- `findings` add nothing to the round.
391
-
392
- **Dispatch shape.** When `delivery.routing.roleScopedAgents` is enabled (the
393
- **default**), dispatch each firing critic with `subagent_type: plan-critic` —
394
- it boots on the role-scoped [`plan-critic`](../../agents/plan-critic.md)
395
- context (its own system prompt, no `CLAUDE.md` @-closure) that carries the
396
- maker-blind invariant, the `consolidation` and `pre-mortem` charters, and the
397
- output shape standalone. When the kill-switch is off
398
- (`roleScopedAgents: false`) or the host cannot spawn at this depth, fall back
399
- to a generic sub-agent and hand it the same charter (the `consolidation` /
400
- `pre-mortem` definitions in [`plan-critic.md`](../../agents/plan-critic.md)).
401
- **When both critics fire, dispatch them in a single turn.** Consolidation and
402
- pre-mortem read the same immutable draft, share no write path, and neither
403
- consumes the other's verdict — the textbook independent fan-out of
404
- [`parallel-tooling.md`](parallel-tooling.md) Rule 3. Issue both `Agent` calls
405
- together in one assistant turn rather than awaiting the first verdict before
406
- spawning the second; serialized critics double the round's wall clock and buy
407
- nothing, because you fold both verdicts into the same re-author round anyway.
408
-
409
- Either way the critic is **maker-blind**: hand it the draft artifacts
410
- (`stories.json`, and `techspec.md` when present) — never the authoring
411
- transcript or the reasons the planner believed its own draft is sound. A
412
- critic that reads the maker's case grades the case, not the draft.
319
+ On `dispatch: true`, dispatch **one fresh-context, maker-blind sub-agent**.
320
+ When `delivery.routing.roleScopedAgents` is enabled (the **default**), use
321
+ `subagent_type: plan-critic` — it boots on the role-scoped
322
+ [`plan-critic`](../../agents/plan-critic.md) context (its own system prompt,
323
+ no `CLAUDE.md` @-closure) that carries the maker-blind invariant, the
324
+ `pre-mortem` charter, and the output shape standalone. When the kill-switch
325
+ is off (`roleScopedAgents: false`) or the host cannot spawn at this depth,
326
+ fall back to a generic sub-agent and hand it the same charter. Either way the
327
+ critic is **maker-blind**: hand it the draft artifacts (`stories.json`, and
328
+ `techspec.md` when present) — never the authoring transcript or the reasons
329
+ the planner believed its own draft is sound. A critic that reads the maker's
330
+ case grades the case, not the draft. Fold surviving findings into Gate #2 or
331
+ a re-author round.
413
332
 
414
333
  ## What `--dry-run` actually gates
415
334
 
416
335
  `plan-persist.js --dry-run` is the same command with GitHub writes suppressed,
417
- and every gate runs before the first `createIssue` would fire — the validator,
418
- the body parse, the DAG, the capacity and Spec-budget ceilings, the
419
- reachability check, the split and supersede partitions, and the Tech Spec fold.
420
- That is the whole point of running it first: a dry run that comes back clean
421
- has already paid for every deterministic refusal, so the real persist has
422
- nothing left to discover except network failure.
336
+ and every gate runs before the first `createIssue` would fire. Since
337
+ Story #5312 the gates split two ways, and the dry-run is where the second
338
+ half is read:
339
+
340
+ **Hard — the run refuses:** a body that does not parse, a ticket that is not
341
+ a Story, an empty `acceptance[]` or `verify[]`, an unknown or cyclic
342
+ `depends_on`, the acceptance partition at N>1, the supersede partition, a
343
+ forbidden commit-subject prefix, and a `deletes` entry naming a path absent
344
+ at base.
345
+
346
+ **Warnings — listed, then the persist proceeds:** a `creates` on a path that
347
+ exists at base or a `refactors-existing` on one that does not (including a
348
+ path the base branch deleted or renamed, named with the removing commit), a
349
+ goal or acceptance path absent at base, a `verify[]` command naming an absent
350
+ test file, and an `open-question` in a body (`Flag if…`, `TBD`, a trailing
351
+ `?`). The list also names every `changes[]` **repair** the run applied — a
352
+ plain-string bullet or a trailing parenthetical rewritten into
353
+ `{ path, assumption }` by probing base. The same list rides the result
354
+ envelope as `warnings[]` and `repairs[]`, so a `--chain-on-clean` run loses
355
+ nothing.
356
+
357
+ A dry run that comes back clean has paid for every deterministic refusal, so
358
+ the real persist has nothing left to discover except network failure.
423
359
 
424
360
  ## The container Epic (Gate #3)
425
361