mandrel 2.55.0 → 2.57.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. package/.agents/agents/plan-critic.md +13 -18
  2. package/.agents/agents/story-worker.md +25 -34
  3. package/.agents/docs/agentrc-reference.json +4 -30
  4. package/.agents/docs/configuration.md +11 -28
  5. package/.agents/docs/execution-reference.md +5 -5
  6. package/.agents/docs/quality-gates.md +8 -7
  7. package/.agents/instructions.md +9 -10
  8. package/.agents/rules/ci-remediation.md +39 -21
  9. package/.agents/schemas/agentrc.schema.json +28 -185
  10. package/.agents/schemas/story-deliver-terminal.schema.json +1 -1
  11. package/.agents/scripts/acceptance-eval.js +107 -17
  12. package/.agents/scripts/audit-to-stories.js +222 -75
  13. package/.agents/scripts/ceremony-derive.js +191 -0
  14. package/.agents/scripts/check-context-budget.js +28 -33
  15. package/.agents/scripts/check-cyclomatic.js +4 -3
  16. package/.agents/scripts/deliver-light.js +31 -94
  17. package/.agents/scripts/file-ci-gap.js +306 -0
  18. package/.agents/scripts/lib/audit-suite/checklist-threading.js +15 -2
  19. package/.agents/scripts/lib/audit-to-stories/audit-label-taxonomy.js +25 -1
  20. package/.agents/scripts/lib/audit-to-stories/dedupe-against-github.js +40 -52
  21. package/.agents/scripts/lib/audit-to-stories/finding-adapter.js +5 -1
  22. package/.agents/scripts/lib/audit-to-stories/issue-corpus.js +162 -0
  23. package/.agents/scripts/lib/audit-to-stories/issues-file.js +121 -0
  24. package/.agents/scripts/lib/audit-to-stories/ledger-commit.js +1 -1
  25. package/.agents/scripts/lib/audit-to-stories/ledger-record.js +126 -0
  26. package/.agents/scripts/lib/audit-to-stories/seed-from-findings.js +11 -0
  27. package/.agents/scripts/lib/baselines/coverage-updater-cli.js +110 -0
  28. package/.agents/scripts/lib/baselines/crap-preview-scan.js +25 -0
  29. package/.agents/scripts/lib/baselines/crap-updater-cli.js +223 -0
  30. package/.agents/scripts/lib/bdd-scenario-budget.js +21 -3
  31. package/.agents/scripts/lib/bootstrap/quality-bootstrap.js +0 -1
  32. package/.agents/scripts/lib/close-validation/gates.js +52 -1
  33. package/.agents/scripts/lib/config/acceptance-eval.js +25 -57
  34. package/.agents/scripts/lib/config/delivery-routing.js +7 -33
  35. package/.agents/scripts/lib/config/explain.js +0 -19
  36. package/.agents/scripts/lib/config/limits.js +18 -78
  37. package/.agents/scripts/lib/config/quality.js +6 -3
  38. package/.agents/scripts/lib/config/runners.js +3 -2
  39. package/.agents/scripts/lib/config-settings-schema-delivery.js +15 -68
  40. package/.agents/scripts/lib/config-settings-schema-quality.js +0 -14
  41. package/.agents/scripts/lib/config-settings-schema.js +49 -143
  42. package/.agents/scripts/lib/crap-engine.js +35 -4
  43. package/.agents/scripts/lib/crap-utils.js +17 -1
  44. package/.agents/scripts/lib/cyclomatic-ceiling.js +19 -7
  45. package/.agents/scripts/lib/feedback-loop/graduator-core.js +53 -13
  46. package/.agents/scripts/lib/feedback-loop/prior-feedback-fetcher.js +71 -25
  47. package/.agents/scripts/lib/feedback-loop/retro-proposals-graduator.js +18 -25
  48. package/.agents/scripts/lib/{audit-to-stories/ledger.js → findings/audit-ledger.js} +131 -24
  49. package/.agents/scripts/lib/findings/route-finding.js +38 -0
  50. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  51. package/.agents/scripts/lib/github/framework-repo.js +148 -2
  52. package/.agents/scripts/lib/label-constants.js +6 -1
  53. package/.agents/scripts/lib/observability/runtime-friction.js +1 -1
  54. package/.agents/scripts/lib/observability/source-classifier.js +2 -0
  55. package/.agents/scripts/lib/orchestration/acceptance-eval-decision.js +5 -4
  56. package/.agents/scripts/lib/orchestration/ceremony-routing.js +19 -73
  57. package/.agents/scripts/lib/orchestration/ci-gap-intake.js +605 -0
  58. package/.agents/scripts/lib/orchestration/ci-rerun-guard.js +13 -8
  59. package/.agents/scripts/lib/orchestration/complexity-gate.js +46 -212
  60. package/.agents/scripts/lib/orchestration/file-assumptions.js +32 -17
  61. package/.agents/scripts/lib/orchestration/light-escalation.js +3 -3
  62. package/.agents/scripts/lib/orchestration/light-suitability.js +66 -233
  63. package/.agents/scripts/lib/orchestration/plan-context.js +181 -387
  64. package/.agents/scripts/lib/orchestration/plan-critic-conditions.js +42 -153
  65. package/.agents/scripts/lib/orchestration/plan-critics-evaluate.js +14 -70
  66. package/.agents/scripts/lib/orchestration/plan-persist/audit-provenance.js +197 -0
  67. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +300 -0
  68. package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +131 -168
  69. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +133 -299
  70. package/.agents/scripts/lib/orchestration/plan-persist/soft-findings.js +55 -0
  71. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +16 -65
  72. package/.agents/scripts/lib/orchestration/plan-persist/wave-serialisation.js +22 -35
  73. package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +30 -139
  74. package/.agents/scripts/lib/orchestration/planning/memory-pool-advisory.js +61 -223
  75. package/.agents/scripts/lib/orchestration/run-epilogue.js +4 -4
  76. package/.agents/scripts/lib/orchestration/single-story-close/phases/close-validation.js +5 -0
  77. package/.agents/scripts/lib/orchestration/single-story-close/phases/pre-gate-steps.js +46 -16
  78. package/.agents/scripts/lib/orchestration/story-close/context-budget-writeback.js +213 -0
  79. package/.agents/scripts/lib/orchestration/story-follow-ups.js +32 -20
  80. package/.agents/scripts/lib/orchestration/task-body-validator.js +10 -63
  81. package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +33 -539
  82. package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +21 -414
  83. package/.agents/scripts/lib/orchestration/ticket-validator.js +54 -118
  84. package/.agents/scripts/lib/orchestration/verify-credit.js +69 -24
  85. package/.agents/scripts/lib/story-body/body-format-lints.js +15 -85
  86. package/.agents/scripts/lib/story-body/story-body.js +17 -237
  87. package/.agents/scripts/lib/templates/decomposer-prompts.js +84 -121
  88. package/.agents/scripts/lib/test-isolate/cli-options.js +93 -0
  89. package/.agents/scripts/lib/test-isolate/progress-log.js +45 -0
  90. package/.agents/scripts/lib/test-isolate/render-report.js +97 -0
  91. package/.agents/scripts/lib/test-isolate/run-isolate.js +87 -0
  92. package/.agents/scripts/lib/test-run-credit.js +266 -0
  93. package/.agents/scripts/lib/wave-runner/footprint.js +48 -358
  94. package/.agents/scripts/lib/wave-runner/ready-set.js +6 -5
  95. package/.agents/scripts/lib/workers/crap-worker.js +32 -41
  96. package/.agents/scripts/plan-context.js +7 -9
  97. package/.agents/scripts/plan-critics.js +28 -54
  98. package/.agents/scripts/plan-persist.js +25 -68
  99. package/.agents/scripts/pr-watch-with-update.js +3 -2
  100. package/.agents/scripts/quality-preview.js +51 -0
  101. package/.agents/scripts/run-tests.js +12 -0
  102. package/.agents/scripts/stories-wave-tick.js +23 -45
  103. package/.agents/scripts/test-isolate.js +13 -180
  104. package/.agents/scripts/update-coverage-baseline.js +25 -70
  105. package/.agents/scripts/update-crap-baseline.js +19 -123
  106. package/.agents/skills/core/scope-triage/SKILL.md +3 -3
  107. package/.agents/workflows/audit-clean-code.md +4 -3
  108. package/.agents/workflows/audit-to-stories.md +63 -27
  109. package/.agents/workflows/helpers/acceptance-self-eval.md +41 -41
  110. package/.agents/workflows/helpers/code-quality-guardrails.md +4 -4
  111. package/.agents/workflows/helpers/code-review.md +2 -3
  112. package/.agents/workflows/helpers/deliver-digest.md +41 -57
  113. package/.agents/workflows/helpers/deliver-light.md +40 -105
  114. package/.agents/workflows/helpers/deliver-reference.md +1 -1
  115. package/.agents/workflows/helpers/deliver-story-reference.md +56 -62
  116. package/.agents/workflows/helpers/deliver-story.md +9 -13
  117. package/.agents/workflows/helpers/plan-reference.md +132 -196
  118. package/.agents/workflows/mandrel-plan.md +28 -41
  119. package/.agents/workflows/memory-consolidate.md +9 -13
  120. package/docs/CHANGELOG.md +33 -0
  121. package/lib/migrations/index.js +4 -0
  122. package/lib/migrations/steps/2.57.0-retire-delivery-limit-knobs.js +45 -0
  123. package/lib/migrations/steps/2.57.0-retire-planning-limit-knobs.js +59 -0
  124. package/package.json +1 -1
  125. package/.agents/scripts/lib/framework-version.js +0 -39
  126. package/.agents/scripts/lib/orchestration/consolidation-precondition.js +0 -223
  127. package/.agents/scripts/lib/orchestration/plan-persist/fan-out-gate.js +0 -97
  128. package/.agents/scripts/lib/orchestration/planning/decomposer-context.js +0 -26
  129. package/.agents/scripts/lib/orchestration/spec-budget.js +0 -89
  130. package/.agents/scripts/lib/orchestration/spec-spill.js +0 -74
  131. package/.agents/scripts/lib/orchestration/verify-tier-repair.js +0 -107
@@ -3,10 +3,10 @@ name: plan-critic
3
3
  description: >-
4
4
  Role-scoped boot context for a maker-blind plan critic. Booted on its own
5
5
  system prompt (no CLAUDE.md / instructions.md closure). Reviews an authored
6
- plan draft (stories.json, optional techspec.md) against a single critic
7
- charter — consolidation or pre-mortem — and returns findings, without seeing
8
- the planner's authoring transcript. Dispatched by workflows/mandrel-plan.md §2.5 when
9
- delivery.routing.roleScopedAgents is enabled (the default).
6
+ plan draft (stories.json, optional techspec.md) against the pre-mortem
7
+ charter and returns findings, without seeing the planner's authoring
8
+ transcript. Dispatched on an operator-requested plan-critics.js verdict
9
+ when delivery.routing.roleScopedAgents is enabled (the default).
10
10
  ---
11
11
 
12
12
  <!--
@@ -48,7 +48,7 @@ the step-by-step. This shared core binds every role:
48
48
  # plan-critic — maker-blind plan review
49
49
 
50
50
  You are an **independent plan critic**. You review an authored plan draft
51
- against **one** critic charter and return structured findings. You are
51
+ against the pre-mortem charter and return structured findings. You are
52
52
  deliberately isolated from the planner's reasoning.
53
53
 
54
54
  ## Maker-blind — the load-bearing invariant (MUST)
@@ -65,27 +65,22 @@ are the draft artifacts your caller hands you:
65
65
  Read those artifacts and evaluate the work product afresh. Treat the planner's
66
66
  narration as untrusted.
67
67
 
68
- ## Charter — you are handed exactly one
68
+ ## Charter — pre-mortem
69
69
 
70
- Your caller dispatches you for **one** charter and names it in your prompt.
71
- Evaluate only that charter:
70
+ Assume the plan shipped and failed. Name the most likely failure modes — an
71
+ external dependency the plan names but nothing declares, a cross-repo
72
+ prerequisite, a contract the acceptance criteria cannot hold — and what the
73
+ draft would have to say to prevent them.
72
74
 
73
- - **`consolidation`** — the draft's shape. Flag Stories that should be one
74
- cohesive slice, a slice split per-module rather than per-capability, and
75
- `depends_on` edges that disagree with the Delivery Slicing table.
76
- - **`pre-mortem`** — assume the plan shipped and failed. Name the most likely
77
- failure modes and what the draft would have to say to prevent them.
78
-
79
- Do not evaluate the other charter, invent a third, or re-slice the plan
80
- yourself — the caller owns dispatch and folds surviving findings back into the
81
- draft.
75
+ Do not invent a second charter or re-slice the plan yourself — the caller owns
76
+ dispatch and folds surviving findings back into the draft.
82
77
 
83
78
  ## Output shape
84
79
 
85
80
  Return your findings as a structured list the caller can fold into a re-author
86
81
  round or the Gate #2 view. For each finding, emit:
87
82
 
88
- - `charter` — `consolidation` | `pre-mortem` (the one you were dispatched for).
83
+ - `charter` — `pre-mortem`.
89
84
  - `severity` — `blocker` | `advisory` (advisory findings inform the operator's
90
85
  Gate #2 decision; they are not an automatic re-author mandate).
91
86
  - `target` — the Story slug / id (or `plan` for a whole-draft finding) the
@@ -83,34 +83,25 @@ Do **not** re-read every file in `project.docsContextFiles`. Read the
83
83
  digest your caller passes, then pull files on demand at the lines it
84
84
  names. A null digest path means no docs mandate.
85
85
 
86
- ## Close gates — one credited run
86
+ ## Close gates — one full-suite run
87
87
 
88
88
  `single-story-close.js` runs the canonical close-validation chain
89
89
  (**typecheck, lint, test, format, maintainability, coverage, crap**) and is
90
- the authoritative gate — do not pre-run it. The **one** exception is
91
- the full suite: run it exactly once, after the self-eval loop's last fix
92
- commit and immediately after the push, in the shape close credits. A bare
93
- `npm test` / `pnpm run test` deposits **no** credit:
94
-
95
- ```bash
96
- # CRAP gate on (default) + a `test:coverage` script:
97
- node <main-repo>/.agents/scripts/coverage-capture.js --cwd <workCwd>
98
- # otherwise — <workCwd> ABSOLUTE, runner exactly `npm test`:
99
- node <main-repo>/.agents/scripts/evidence-gate.js --standalone \
100
- --scope-id <storyId> --gate test --worktree <workCwd> -- npm test
101
- ```
102
-
103
- Dispatch it in the **background**: it routinely outruns the host's sync
104
- Bash ceiling, and its completion re-invokes you — which is why the push
105
- comes first. Never spawn a task to poll or `sleep`-loop
106
- against it; a waiter with a wrong condition outlives the agent. Share
107
- `lint` / `typecheck` evidence with close via `evidence-gate.js`; never
108
- stamp coverage / CRAP fresh any other way.
109
-
110
- **It can legitimately run nothing.** With nothing changed under the CRAP
111
- `targetDirs` it skips capture and exits 0. An exit code is never evidence a
112
- gate did work — its **output** is: no credit was deposited. Run the scoped
113
- projects for the roots you changed plus `verify[]`, not the whole suite.
90
+ the authoritative gate — do not pre-run it. The **one** exception is the
91
+ full suite: after the self-eval loop's last fix commit, run `npm test`
92
+ exactly once in `<workCwd>`. A green full run on `story-<storyId>`
93
+ deposits the `test` credit close reads, keyed on the tree; a later commit
94
+ voids it, and close captures coverage itself when the CRAP gate needs an
95
+ artifact. If the suite outruns the host's sync Bash ceiling, dispatch it in
96
+ the **background** — its completion notification re-invokes you. Never
97
+ spawn a task to poll or `sleep`-loop against it; a waiter with a wrong
98
+ condition outlives the agent. An exit code is never evidence a gate did
99
+ work — the runner's **output** is: it prints whether it deposited credit,
100
+ and a run off the Story branch or of a partial tier deposits nothing and
101
+ says so. Redraft rounds run the scoped projects for the roots you changed
102
+ plus `verify[]`, not the whole suite. Share `lint` / `typecheck` evidence
103
+ with close via `evidence-gate.js`; never stamp coverage / CRAP fresh any
104
+ other way.
114
105
 
115
106
  Gate output that lies: [`known-tooling-behavior.md`](../rules/known-tooling-behavior.md).
116
107
  Waiter traps: [`parallel-tooling.md`](../workflows/helpers/parallel-tooling.md) Rule 2.
@@ -119,12 +110,13 @@ Waiter traps: [`parallel-tooling.md`](../workflows/helpers/parallel-tooling.md)
119
110
 
120
111
  **Before** flipping to `closing`, run the bounded self-eval loop
121
112
  ([`acceptance-self-eval.md`](../workflows/helpers/acceptance-self-eval.md)).
122
- It scores the change set you computed **once** and injected into the critic
123
- — never one it re-derives — against each `acceptance[]` item,
124
- consuming `verify[]` output as evidence. **proceed** → flip to `closing`,
125
- push, capture, hand off; **redraft** → fix the criteria, commit, re-eval;
126
- **block** → take the blocked path below. Never hand off an unscored
127
- branch.
113
+ Derive the change set, level and ceremony with
114
+ `node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>`
115
+ and hand its `files` to the critic — never one it re-derives. It scores
116
+ each `acceptance[]` item, consuming `verify[]` output as evidence.
117
+ **proceed** → flip to `closing`, run the suite, push, hand off;
118
+ **redraft** → fix the criteria, commit, re-eval; **block** → take the
119
+ blocked path below. Never hand off an unscored branch.
128
120
 
129
121
  ## Lifecycle: progress & blocked (MUST)
130
122
 
@@ -146,9 +138,8 @@ the only sanctioned landing.
146
138
  ## Your turn ends at a pushed branch (MUST)
147
139
 
148
140
  You do **not** run close. Push `story-<storyId>` to `origin` — confirming
149
- the remote ref moved — **before** the credited capture, then return: the
150
- capture is backgrounded, so its completion ends your turn, and a turn that
151
- ends unpushed reads as unfinished work. The orchestrator runs
141
+ the remote ref moved — then return: a turn that ends unpushed reads as
142
+ unfinished work. The orchestrator runs
152
143
  `single-story-close.js` in its own session, serialized against your
153
144
  siblings. Do not open the PR, flip `agent::done`, or spawn a child to
154
145
  close for you. If the push fails, take the blocked path above.
@@ -28,6 +28,10 @@
28
28
  "projectOwner": null,
29
29
  "operatorHandle": "@[USERNAME]",
30
30
  "defaultTimeoutMs": 60000,
31
+ "followUpRepos": {
32
+ "framework": "dsj1984/mandrel",
33
+ "platform": null
34
+ },
31
35
  "branchProtection": {
32
36
  "enforce": true,
33
37
  "requiredChecks": [
@@ -66,28 +70,9 @@
66
70
  }
67
71
  },
68
72
  "planning": {
69
- "riskHeuristics": [
70
- "Destructive or irreversible data mutations (dropping tables, deleting rows without soft-delete or backup, truncating production state).",
71
- "Modifications to shared security or auth infrastructure (IAM policies, auth middleware, session or token handling, secret rotation).",
72
- "Changes to CI/CD, deployment pipelines, or release gating that could disable safety checks or ship unverified code to production.",
73
- "Monorepo-wide AST or text replacements touching overlapping files in parallel (catastrophic merge-conflict risk across concurrent agents).",
74
- "Schema migrations that rewrite existing rows or drop columns without a backfill or rollback plan."
75
- ],
76
73
  "memoryPool": {
77
- "staleAfterDays": 30,
78
- "growthDelta": 25,
79
74
  "indexByteCeiling": 24576
80
75
  },
81
- "failOnSharedEditors": false,
82
- "requireExplicitCrossStoryDeps": false,
83
- "crossCuttingRegistries": [
84
- "lib/orchestration/lifecycle/listeners/index.js",
85
- "**/listeners/index.js",
86
- "**/handlers/index.js"
87
- ],
88
- "failOnRegistryConflicts": false,
89
- "failOnLargeFanOut": false,
90
- "largeFanOutThreshold": 10,
91
76
  "navigation": {
92
77
  "routeGlobs": [],
93
78
  "navRegistry": []
@@ -130,14 +115,6 @@
130
115
  ".agents/instructions.local.md"
131
116
  ]
132
117
  },
133
- "signals": {
134
- "rework": {
135
- "editsPerFile": 5
136
- },
137
- "retry": {
138
- "repeatCount": 3
139
- }
140
- },
141
118
  "quality": {
142
119
  "gateScoping": {
143
120
  "scope": "diff",
@@ -284,7 +261,6 @@
284
261
  },
285
262
  "codingGuardrails": {
286
263
  "cyclomaticFlag": 8,
287
- "cyclomaticMustFix": 12,
288
264
  "requireSiblingTest": false
289
265
  },
290
266
  "autoRefresh": {
@@ -332,7 +308,6 @@
332
308
  }
333
309
  ],
334
310
  "maxFixAttempts": 3,
335
- "maxFixScopeFiles": 5,
336
311
  "autoFixSeverity": "medium"
337
312
  },
338
313
  "refactorStage": {
@@ -359,7 +334,6 @@
359
334
  },
360
335
  "routing": {
361
336
  "roleScopedAgents": true,
362
- "freshCriticSampleRate": 0.2,
363
337
  "ceremonyProfile": "standard",
364
338
  "closeAndLand": true
365
339
  }
@@ -103,6 +103,9 @@ GitHub provider identity plus the remote stance the bootstrap enforces. `owner`,
103
103
  | `projectOwner` | No | `string` \| `null` | `null` | Owner of the Projects V2 board when it lives outside `owner` (an org board fed by a user repo). `null` means the board shares `owner`. |
104
104
  | `operatorHandle` | Yes | `string` | `"@[USERNAME]"` | The human the framework escalates to, `@`-prefixed. Used for HITL @-mentions on `agent::blocked`. |
105
105
  | `defaultTimeoutMs` | No | `integer` | `60000` | Default `timeoutMs` applied to every `gh` subprocess the provider facade spawns, so a stalled socket or long-poll cannot hang an orchestration indefinitely. A `GhExecTimeoutError` from a hit ceiling is classified `transient` and retried by `withTransientRetry`. Story #2860. |
106
+ | `followUpRepos` | No | `object` | — | Repository slugs for the non-consumer follow-up ownership buckets, used when a CI gap, retro proposal, or audit finding belongs to someone other than the repo that surfaced it. |
107
+ | `followUpRepos.framework` | No | `string` | `"dsj1984/mandrel"` | `<owner>/<repo>` that owns framework-level defects. Defaults to the Mandrel mirror — the one bucket with a knowable default. |
108
+ | `followUpRepos.platform` | No | `string` \| `null` | `null` | `<owner>/<repo>` for a shared platform or infrastructure tracker (a shared base config, a runner fleet, a cross-repo toolchain). No default — nothing can guess a shared repo. Left unset, platform-owned findings file locally and say so. |
106
109
  | `branchProtection` | No | `object` | — | Branch-protection stance applied to `project.baseBranch` by the GitHub bootstrap, and reproduced locally before every push. |
107
110
  | `branchProtection.enforce` | No | `boolean` | `true` | When true, the GitHub bootstrap writes the required-check ruleset. False leaves the remote stance alone. |
108
111
  | `branchProtection.requiredChecks[]` | No | `array<object>` | `[{"name":"lint","cmd":["npm","run","lint"]},{"name":"test","cmd":["npm","test"]},{"name":"baselines","cmd":["node",".agents/scripts/check-baselines.js"]}]` | Checks that must pass before a Story PR merges. Each entry carries both the remote context name and the local argv. Each item has: name, cmd. |
@@ -119,31 +122,19 @@ GitHub provider identity plus the remote stance the bootstrap enforces. `owner`,
119
122
 
120
123
  ### `planning` (optional)
121
124
 
122
- Inputs to `/mandrel-plan`: risk escalation heuristics, ceremony-lite routing, the memory-hygiene advisory thresholds, and the cross-Story conflict-finding severity gates.
125
+ Inputs to `/mandrel-plan`: the memory-hygiene advisory ceiling and the opt-in navigability reachability gate.
123
126
 
124
127
  | Key | Required | Type | Default | Description |
125
128
  | --- | --- | --- | --- | --- |
126
- | `riskHeuristics` | No | `string[]` or `{ append?, prepend? }` | `["Destructive or irreversible data mutations (dropping tables, deleting rows without soft-delete or backup, truncating production state).","Modifications to shared security or auth infrastructure (IAM policies, auth middleware, session or token handling, secret rotation).","Changes to CI/CD, deployment pipelines, or release gating that could disable safety checks or ship unverified code to production.","Monorepo-wide AST or text replacements touching overlapping files in parallel (catastrophic merge-conflict risk across concurrent agents).","Schema migrations that rewrite existing rows or drop columns without a backfill or rollback plan."]` | Prose heuristics the planner escalates a Story against. A plain array replaces the framework list; the `{ append, prepend }` extender form deep-merges with it. |
127
- | `complexityGate` | No | `object` | — | Shape-derived ceremony-lite complexity routing. A lite claim is validated against the authored Story shape at persist and re-derived from the Story body at dispatch; conservative (full on any doubt). Never relaxes the Story-ticket / PR-to-main / repo-gates / security-baseline non-negotiables. |
128
- | `complexityGate.enabled` | No | `boolean` | — | Master switch. When false, lite routing is disabled everywhere: persist refuses lite claims and dispatch always takes the sub-agent path. Default true. |
129
- | `complexityGate.maxArtifacts` | No | `integer` | — | Enumerated-artifact threshold reported by the plan-context complexity signals. An input signal for the planner verdict — carries no routing authority. Default 1. |
130
- | `memoryPool` | No | `object` | — | Thresholds for the memory-hygiene advisory `/mandrel-plan` surfaces at Gate #1. Advisory only: it recommends `/memory-consolidate` and never gates, reroutes, or mutates the memory pool. |
131
- | `memoryPool.staleAfterDays` | No | `integer` | `30` | Recommend a consolidation pass once the pool's stamp is older than this many days. Default 30. |
132
- | `memoryPool.growthDelta` | No | `integer` | `25` | Recommend a consolidation pass once this many entries have been written since the last one. Measured against the entry count the last pass stamped, so a stamp predating that field leaves growth unmeasured and only the age threshold applies. Default 25. |
133
- | `memoryPool.indexByteCeiling` | No | `integer` | `24576` | Recommend a consolidation pass once the pool's `MEMORY.md` index exceeds this many bytes. Independent of the age and growth thresholds: the harness truncates the index it loads into each session at its own byte cap, so an oversized index is a loss already happening — every entry listed after the cut is invisible — rather than a hygiene forecast. Default 24576, the harness cap itself. |
134
- | `failOnSharedEditors` | No | `boolean` | `false` | When true, upgrade shared-editor conflict findings to hard errors (default false — advisory soft findings only). |
135
- | `requireExplicitCrossStoryDeps` | No | `boolean` | `false` | When true, upgrade implicit cross-Story dependency findings to hard errors (default false — advisory soft findings only). |
136
- | `crossCuttingRegistries` | No | `string[]` or `{ append?, prepend? }` | `["lib/orchestration/lifecycle/listeners/index.js","**/listeners/index.js","**/handlers/index.js"]` | Registry path patterns whose concurrent edits across Stories are flagged as conflicts. Defaults to the framework listener/handler index patterns when omitted. |
137
- | `failOnRegistryConflicts` | No | `boolean` | `false` | When true, upgrade cross-cutting registry conflict findings to hard errors (default false). |
138
- | `failOnLargeFanOut` | No | `boolean` | `false` | When true, upgrade fan-out-warning findings (delete blast radius) to hard errors (default false — soft advisory). |
139
- | `largeFanOutThreshold` | No | `integer` | `10` | Call-site count above which a Story that deletes a module emits a fan-out-warning. Counts base-branch references to the deleted path basename. Soft by default; does not size or reject Stories. Default 10. |
129
+ | `memoryPool` | No | `object` | — | Threshold for the memory-hygiene advisory `/mandrel-plan` surfaces at Gate #1. Advisory only: it recommends `/memory-consolidate` and never gates, reroutes, or mutates the memory pool. |
130
+ | `memoryPool.indexByteCeiling` | No | `integer` | `24576` | Recommend a consolidation pass once the pool's `MEMORY.md` index exceeds this many bytes. The harness truncates the index it loads into each session at its own byte cap, so an oversized index is a loss already happening — every entry listed after the cut is invisible — rather than a hygiene forecast. Default 24576, the harness cap itself. |
140
131
  | `navigation` | No | `object` | — | Opt-in navigability reachability gate. Absent or empty routeGlobs is a silent no-op. |
141
132
  | `navigation.routeGlobs` | No | `array<string>` | `[]` | Glob patterns (e.g. pages/**, app/**/route.ts) marking paths that add a user-facing route. |
142
133
  | `navigation.navRegistry` | No | `array<string>` | `[]` | Tokens identifying the nav-registry SSOT a route-adding Story is expected to reference. |
143
134
 
144
135
  ### `delivery` (optional)
145
136
 
146
- Everything `/mandrel-deliver` and `single-story-close` consume: execution timeouts, worktree isolation, runner concurrency, docs freshness, signals, quality gates, merge/CI watch, review ceremony, and the feedback loop.
137
+ Everything `/mandrel-deliver` and `single-story-close` consume: execution timeouts, worktree isolation, runner concurrency, docs freshness, quality gates, merge/CI watch, review ceremony, and the feedback loop.
147
138
 
148
139
  | Key | Required | Type | Default | Description |
149
140
  | --- | --- | --- | --- | --- |
@@ -172,11 +163,6 @@ Everything `/mandrel-deliver` and `single-story-close` consume: execution timeou
172
163
  | `worktreeIsolation.allowSymlinkOnWindows` | No | `boolean` | `false` | Permit the `symlink` strategy on win32, where it needs Developer Mode or elevation. Off by default so a Windows consumer fails over to a strategy that works. |
173
164
  | `worktreeIsolation.reapOnSuccess` | No | `boolean` | `true` | Remove the Story's worktree once its PR merges. False keeps it for post-mortem inspection. |
174
165
  | `worktreeIsolation.bootstrapFiles` | No | `array<string>` | `[".env",".mcp.json",".agentrc.local.json",".agents/instructions.local.md"]` | Gitignored files copied from the main checkout into every new worktree. A worktree checks out tracked files only, so local secrets and overrides would otherwise be missing. |
175
- | `signals` | No | `object` | — | Detector thresholds for the surviving performance-signal categories. Each block is shallow-merged by the resolver. |
176
- | `signals.rework` | No | `object` | — | Rework detector — repeated edits to one file in a run. |
177
- | `signals.rework.editsPerFile` | No | `integer` | `5` | Edits to a single file within one run that trip the rework signal. |
178
- | `signals.retry` | No | `object` | — | Retry detector — the same command failing repeatedly. |
179
- | `signals.retry.repeatCount` | No | `integer` | `3` | Repeats of an identical failing command that trip the retry signal. |
180
166
  | `quality` | No | `object` | — | Quality-gate configuration. Every gate lives under `gates.<tier>` and shares the same `{ enabled, baselinePath, tolerance, floors, components }` base; shared scoping lives at this block root. |
181
167
  | `quality.gateScoping` | No | `object` | — | Shared scope applied to every gate that supports one, unless the gate overrides it. |
182
168
  | `quality.gateScoping.scope` | No | `"diff"` \| `"full"` | `"diff"` | Score only the files changed against `diffRef` (`diff`) or every file in the gate target dirs (`full`). |
@@ -277,7 +263,6 @@ Everything `/mandrel-deliver` and `single-story-close` consume: execution timeou
277
263
  | `quality.formatAutofix.timeoutMs` | No | `integer` | `60000` | Timeout (ms) for the format-write spawn. |
278
264
  | `quality.codingGuardrails` | No | `object` | — | Authoring-time cyclomatic-complexity advisories surfaced by the quality-preview pre-commit gate. |
279
265
  | `quality.codingGuardrails.cyclomaticFlag` | No | `integer` | `8` | Cyclomatic complexity at which a new or changed method is flagged for a refactor look. |
280
- | `quality.codingGuardrails.cyclomaticMustFix` | No | `integer` | `12` | Cyclomatic complexity at which a new or changed method must be decomposed before the diff closes. |
281
266
  | `quality.codingGuardrails.requireSiblingTest` | No | `boolean` | `false` | When true, a new source file with no colocated sibling test is reported by the guardrails pass. |
282
267
  | `quality.autoRefresh` | No | `object` | — | Baseline-attribution auto-refresh: when a gate can prove a regression is a legitimate consequence of the diff, it rewrites the baseline instead of blocking. |
283
268
  | `quality.autoRefresh.enabled` | No | `boolean` | `true` | Master switch for the auto-refresh path. When false, every baseline refresh is a deliberate operator action. |
@@ -307,14 +292,13 @@ Everything `/mandrel-deliver` and `single-story-close` consume: execution timeou
307
292
  | `codeReview.providers[]` | No | `array<object>` | `[{"name":"native"},{"name":"security-review","scopes":["story"],"optional":true},{"name":"ultrareview","scopes":["story"],"manualPrompt":true,"when":{"label":"risk::high"}}]` | Review-provider chain (Story #2871). When unset or empty, defaults to [{ name: "native" }]. The orchestrator iterates inline entries in declaration order and merges their Finding[] before posting one structured comment; manual-prompt entries (e.g. ultrareview) contribute a trailing 'Manual review suggestions' section. Selecting an adapter whose probe fails hard-fails at factory construction unless declared `optional: true` in the chain. Each item has: name, scopes, optional, manualPrompt, when. |
308
293
  | `codeReview.providerConfig` | No | `object` | — | Optional escape hatch for adapter-specific configuration. No documented keys in Epic #2815; reserved so future adapters can be configured without another schema migration. |
309
294
  | `codeReview.maxFixAttempts` | No | `integer` | `3` | Maximum auto-fix retry attempts per finding in /mandrel-deliver Phase 5 (code-review). 0 disables auto-fix. Default 3. |
310
- | `codeReview.maxFixScopeFiles` | No | `integer` | `5` | Maximum file count a single auto-fix may modify before escalating to agent::blocked. Default 5. |
311
295
  | `codeReview.autoFixSeverity` | No | `"high"` \| `"medium"` | `"medium"` | Severity threshold for on-branch remediation in /mandrel-deliver Phase 5 (code-review). `medium` (default) routes 🔴/🟠/🟡 findings into the host-LLM focused-fix routing (Mediums batched per lens: one commit per lens, a single validation + rescan at the end) while 🟢 suggestions still graduate to follow-up issues; `high` reproduces the pre-4399 Critical/High-only routing. Hard cutover — no back-compat flag. |
312
296
  | `review` | No | `object` | — | Close-scope review tuning (Story #4699). Governs the Story-scope local-lens pass that runs inside the close subprocess; the maker-blind code-review pass and all hard gates are unaffected. |
313
297
  | `review.lensDiffFloor` | No | `integer` | — | Changed-line floor for the close-scope lens walk (Story #4699). A diff strictly below this many changed lines (additions + deletions) with zero sensitive-path hits skips lens materialization and records the skip in the findings-yield ledger. Default 40; 0 disables the skip. Hard gates and the maker-blind code-review pass are unaffected. |
314
298
  | `refactorStage` | No | `object` | — | Opt-in, config-gated post-green refactor checkpoint wired into story-deliver (Story #3430, Epic #3418). Strictly additive and default-OFF: when disabled, story-deliver behaves exactly as before. Advisory only — never changes existing close-validation gate semantics. |
315
299
  | `refactorStage.enabled` | No | `boolean` | `false` | When true, story-deliver runs an advisory post-green refactor stage (core/code-review-and-quality skill, Post-Green Refactor Pass) after the suite is green. Default false — when unset the stage is skipped and close-validation gate semantics are unchanged. |
316
- | `acceptanceEval` | No | `object` | — | Story #3819. Bounded per-Story acceptance self-eval loop. After the implementation commits land and before the Story-implementation phase flips to `closing`, an independent (fresh-context) critic pass scores the caller-injected change set against each inline `acceptance[]` item, redrafts the unmet items, and re-evaluates — capped at `maxRounds` redraft rounds, then escalates to `agent::blocked` when criteria remain unmet. There is no `enabled` flag: the loop is a hard cutover (always on). |
317
- | `acceptanceEval.maxRounds` | No | `integer` | `2` | Maximum number of redraft rounds before escalation. Default 2; clamped into [1, hard ceiling] by lib/config/acceptance-eval.js so the cap can never be disabled (maxRounds: 0 clamps up to 1). |
300
+ | `acceptanceEval` | No | `object` | — | Story #3819. Bounded per-Story acceptance self-eval loop. After the implementation commits land and before the Story-implementation phase flips to `closing`, an independent (fresh-context) critic pass scores the caller-injected change set against each inline `acceptance[]` item, redrafts the unmet items, and re-evaluates — capped at `maxRounds` redraft rounds (0 = scored once, no redraft), then escalates to `agent::blocked` when criteria remain unmet. There is no `enabled` flag: the scoring pass is a hard cutover (always on). |
301
+ | `acceptanceEval.maxRounds` | No | `integer` | `2` | Maximum number of redraft rounds before escalation. Default 2; 0 means the verdict is scored once with no redraft round (Story #5313 dropped the hard ceiling and the floor-of-one clamp). |
318
302
  | `feedbackLoop` | No | `object` | — | Opt-out toggles for the close-time auto-file graduators. All default to auto-filing on. |
319
303
  | `feedbackLoop.auditResultsAutoFile` | No | `boolean` | `true` | When true (default), the close-time audit-results graduator auto-files non-blocking audit-results findings as follow-up issues routed by source classification. Set to false to suppress auto-filing; findings remain accessible in the structured comments on the Story. |
320
304
  | `feedbackLoop.retroProposals` | No | `boolean` | `true` | When true (default), the retro auto-files its actionable routed proposals as meta::<framework-gap\|consumer-improvement> + friction::<category> issues via the graduator pre-parsed-findings seam, and the rendered retro sections list the filed issue numbers instead of paste-ready gh command stanzas. Set to false to fall back to the command stanzas. |
@@ -332,10 +316,9 @@ Everything `/mandrel-deliver` and `single-story-close` consume: execution timeou
332
316
  | `ci.blockOnAdvisoryFailure` | No | `boolean` | `true` | Story #5096. When true (default), delivery refuses to arm — and disarms — GitHub native auto-merge while a non-required (advisory) check is genuinely red on the PR head and GitHub reports the PR mergeable anyway (mergeStateStatus=UNSTABLE). `--auto` waits on REQUIRED contexts only, so without this a red advisory quality gate merges unattended. Set false to restore the pre-#5096 behaviour verbatim. |
333
317
  | `ci.advisoryAllowlist` | No | `array<string>` | `[]` | Story #5096. Check-run names exempt from blockOnAdvisoryFailure — a red run whose name matches exactly never blocks arming. Matching is exact; an unnamed run can never match and always blocks. |
334
318
  | `ci.rerunAdvisory` | No | `integer` | `0` | Story #5266. How many times close may re-run a failed advisory workflow run before blocking on it, per close invocation. Default 0: close spends no CI minutes and issues no GitHub mutation on an advisory red unless asked. At n > 0 the failed run(s) are re-run within that allowance and the merge wait re-polls inside its existing budget, landing or blocking on the re-run verdict. Overridden per invocation by --rerun-advisory <n>. |
335
- | `routing` | No | `object` | — | v2 delivery-spawn routing: role-scoped boot contexts and maker-checker sampling. The v1 singleDelivery epic-route kill-switch was removed in Stage 6. |
319
+ | `routing` | No | `object` | — | v2 delivery-spawn routing: role-scoped boot contexts and the ceremony profile. The v1 singleDelivery epic-route kill-switch was removed in Stage 6; the freshCriticSampleRate sampling floor was retired in Story #5313. |
336
320
  | `routing.roleScopedAgents` | No | `boolean` | `true` | Epic #4478 (M7-B). Kill-switch for the role-scoped boot contexts. When true (default), a converted delivery spawn (`story-worker`, `acceptance-critic`) boots on its own `.claude/agents/<role>.md` system prompt instead of re-paying the full CLAUDE.md @-import closure. When false, every converted spawn falls back to `subagent_type: general-purpose` — the instant, code-rollback-free per-consumer revert, and the universal escape for hosts that ignore `.claude/agents/`. The fallback is the full-closure agent that ran before M7-B, so flipping it off never drops a gate. |
337
- | `routing.freshCriticSampleRate` | No | `number` | `0.2` | Epic #4478 (M7-B, Part 2). Maker-checker sampling floor. Under the standard profile, a change set touching no sensitive path routes its acceptance clusters down the contract-identical inline critic path, but this fraction of them is still forced through a fresh-context critic so a low derived level never means zero independent checking. Clamped to [0, 1]; 0 disables the floor, 1 forces every cluster fresh. Consumed by resolveCeremonyForRisk (lib/orchestration/ceremony-routing.js). |
338
- | `routing.ceremonyProfile` | No | `"minimal"` \| `"standard"` \| `"strict"` | `"standard"` | Acceptance-ceremony depth. minimal = always inline critic; strict = always fresh-context critic; standard (default) = routed off the change level derived from the Story diff, with the maker-checker sampling floor. |
321
+ | `routing.ceremonyProfile` | No | `"minimal"` \| `"standard"` \| `"strict"` | `"standard"` | Acceptance-ceremony depth. minimal = always inline critic; strict = always fresh-context critic; standard (default) = routed off the change level derived from the Story diff: high or underivable → fresh, low → inline. |
339
322
  | `routing.closeAndLand` | No | `boolean` | `true` | When true (default), single-story-close lands through merge in one close. Opt out per-run with --no-wait-merge. |
340
323
 
341
324
  ### `qa` (optional)
@@ -88,7 +88,8 @@ over-ceiling envelope or an over-budget Story count.
88
88
  > on planner-context size. Separately, the `ContextEnvelope` SDK this section
89
89
  > used to credit with limiting hydrated prompt size had no production caller
90
90
  > and was deleted in Story #5005; only its `estimateTokens` helper survived,
91
- > re-homed in `lib/orchestration/spec-spill.js`.
91
+ > now private to `lib/audit-suite/checklist-threading.js` (Story #5312 deleted
92
+ > `spec-spill.js` with the Spec token budget).
92
93
 
93
94
  ### Planner-context envelope (`/mandrel-plan`)
94
95
 
@@ -120,9 +121,8 @@ over-ceiling envelope or an over-budget Story count.
120
121
 
121
122
  ### Session-mass capacity (plan-time sizing)
122
123
 
123
- - **`DEFAULT_MODEL_CAPACITY`** (`lib/orchestration/ticket-validator-sizing.js`):
124
- absolute authored-token ceilings for plan-time Story sizing (soft 30k /
125
- hard 75k). Not operator-configurable via `.agentrc.json`; programmatic
126
- override via `opts.modelCapacity` on validateTickets / runPlanPersist only.
124
+ - **Retired (Story #5312).** `DEFAULT_MODEL_CAPACITY` and the plan-time
125
+ Story sizing ceilings are gone: a Story is as large as the work needs, and
126
+ nothing at plan time scores its authored mass.
127
127
  - **Host runtime**: session billing, quota exhaustion, and operator overrides
128
128
  are enforced by your provider (e.g. Claude Code), not by Mandrel scripts.
@@ -302,8 +302,9 @@ default and the deep-merge extender form).
302
302
 
303
303
  ## Cyclomatic ceiling ratchet
304
304
 
305
- `delivery.quality.codingGuardrails.cyclomaticMustFix` (default `12`) is the
306
- per-function complexity ceiling, enforced by `check-cyclomatic.js`. It is a
305
+ A fixed per-function complexity ceiling of `12` is enforced by
306
+ `check-cyclomatic.js` (`lib/cyclomatic-ceiling.js#CYCLOMATIC_CEILING`; the
307
+ `cyclomaticMustFix` config key was retired in Story #5313). It is a
307
308
  **standalone ratchet** — the same slot as `check-arch-cycles.js`,
308
309
  `check-dead-exports.js`, and `check-context-budget.js` — not a
309
310
  `delivery.quality.gates` kind, so it needs no gate block and no floor.
@@ -330,9 +331,10 @@ The scan reuses `delivery.quality.gates.maintainability.targetDirs` /
330
331
  `ignoreGlobs` — both instruments read the same coverage-free escomplex
331
332
  surface, so a separate scope declaration could only ever restate it.
332
333
 
333
- `cyclomaticFlag` (default `8`) is the softer half of the pair: it is not
334
- gated, and names the ceiling `quality:preview` counts new methods against in
335
- its `new-method count over c=<flag>` column.
334
+ `cyclomaticFlag` (default `8`) is the one advisory knob: it is not gated, and
335
+ names the ceiling `quality:preview` counts new methods against in its
336
+ `new-method count over c=<flag>` column. The preview also lists every scanned
337
+ method at cyclomatic 12 or above as an advisory and exits 0 on it.
336
338
 
337
339
  ---
338
340
 
@@ -825,8 +827,7 @@ correct shape for a gate with no rescoring path of its own.
825
827
  ## HITL blocker escalation
826
828
 
827
829
  `risk::high` is planning/audit metadata only — it never pauses runtime. The
828
- sole runtime HITL pause point is `agent::blocked`; `planning.riskHeuristics`
829
- is the rubric for what should escalate. The full model is owned by
830
+ sole runtime HITL pause point is `agent::blocked`. The full model is owned by
830
831
  [`.agents/instructions.md`](../instructions.md) § 1.J and
831
832
  [`SDLC.md` § HITL model](SDLC.md#hitl-human-in-the-loop-model).
832
833
 
@@ -96,8 +96,7 @@ metadata only — no automatic runtime pause. The single runtime pause
96
96
  point is **`agent::blocked`**: on an unresolvable blocker or an unsafe
97
97
  destructive action without explicit authorization, transition to
98
98
  `agent::blocked`, summarize the blocker, and wait for operator resume
99
- (`agent::executing` or equivalent); escalate per
100
- `planning.riskHeuristics` in `.agentrc.json`.
99
+ (`agent::executing` or equivalent).
101
100
 
102
101
  ### K. Precedence & Conflict Resolution
103
102
 
@@ -113,8 +112,8 @@ always wins regardless of tier.
113
112
  ## 2. FinOps & Token Budgeting (Economic Guardrails)
114
113
 
115
114
  Mandrel does not enforce live LLM spend; your host owns session quota.
116
- Fixed framework ceilings (the `/mandrel-plan` context envelope, plan-time Story
117
- sizing) **fail closed** naming what to trim:
115
+ The one fixed framework ceiling (the `/mandrel-plan` context envelope)
116
+ truncates with a note naming what was cut:
118
117
  [`docs/execution-reference.md`](docs/execution-reference.md#finops--token-budgeting-economic-guardrails).
119
118
 
120
119
  ---
@@ -169,8 +168,8 @@ Do NOT manually update issue descriptions or status fields unless prompted.
169
168
  ### B. Ticket hierarchy
170
169
 
171
170
  The Story is the only executable ticket: `acceptance[]` / `verify[]`
172
- inline plus the folded Tech Spec in `## Spec` (over-budget Specs fail
173
- closed — split or tighten; never under `docs/`). Optional `depends_on`
171
+ inline plus the folded Tech Spec in `## Spec` (as long as the work
172
+ needs; never under `docs/`). Optional `depends_on`
174
173
  edges order rare multi-Story runs, resolved by `/mandrel-deliver` from
175
174
  live state; `plan-run::<id>` is filter metadata. Commit subjects
176
175
  reference the Story via `(refs #<storyId>)`. There is no `type::task`.
@@ -194,9 +193,9 @@ anything under it.
194
193
 
195
194
  `/mandrel-plan` sizes each Story as a **capability slice a frontier model
196
195
  delivers and self-verifies in one pass** — a broad footprint is normal
197
- when the change is cohesive (backstop: `DEFAULT_MODEL_CAPACITY` in
198
- `ticket-validator-sizing.js`); do not re-slice it into per-module
199
- fragments. On a `⚠️ COMPLEXITY WARNING` or out-of-scope task: **plan
200
- first** (numbered cohesive sub-steps in a `<!-- DECOMPOSITION -->`
196
+ when the change is cohesive, and no plan-time ceiling scores it; do not
197
+ re-slice it into per-module fragments. On a `⚠️ COMPLEXITY WARNING` or
198
+ out-of-scope task: **plan first** (numbered cohesive sub-steps in a
199
+ `<!-- DECOMPOSITION -->`
201
200
  block), **commit incrementally** per sub-step, and **fail fast** — STOP
202
201
  and report if any sub-step fails validation.
@@ -27,13 +27,29 @@ exactly one of two ways, and no others:
27
27
  through the fix table in
28
28
  [`deliver-story-reference.md` § Step 4](../workflows/helpers/deliver-story-reference.md#step-4--ci-watch--fix-recovery);
29
29
  refresh a baseline only when the diff demonstrably can't be covered.
30
- 2. **File a `meta::framework-gap` issue** when the root cause is outside this
30
+ 2. **File the CI-gap intake issue** when the root cause is outside this
31
31
  delivery's scope — a pre-existing flaky test, a runner/infra weakness, a
32
- framework-level environment gap. Open the issue with the `meta::framework-gap`
33
- label (see [`git-conventions.md`](git-conventions.md)) carrying **the run
34
- link and the failure signature** so a later `/mandrel-plan` Phase 0 sweep can act on
35
- it. Remediate this delivery only if the pre-existing defect is genuinely
36
- blocking it.
32
+ framework-level environment gap. One command does it, and it is the only
33
+ sanctioned filing surface:
34
+
35
+ ```bash
36
+ node .agents/scripts/file-ci-gap.js --story <id> --verdict <verdict> \
37
+ --owner <consumer|framework|platform> --evidence "<proof reading>" [--block]
38
+ ```
39
+
40
+ It reads the digest for the run link and failure signature, routes the
41
+ filing to the repository that **owns** the fault, dedups by signature so
42
+ the Nth occurrence updates the existing ticket, posts the `friction`
43
+ comment, and (with `--block`) flips the Story. Hand-running `gh issue
44
+ create` is not the fallback: it files an issue no `/mandrel-plan` pass can
45
+ graduate, in whichever repo you happen to be standing in. Remediate this
46
+ delivery only if the pre-existing defect is genuinely blocking it.
47
+
48
+ **`--owner` is the judgement call**, and it is yours to make from the
49
+ evidence: `consumer` for this repository's own code, `framework` for a
50
+ Mandrel defect, `platform` for a shared base config, runner fleet, or
51
+ cross-repo toolchain that neither owns. An unconfigured bucket files
52
+ locally and says so in the issue body — it never pretends to be routed.
37
53
 
38
54
  Infra, transient, and flaky failures are root-cause defects too — a flaky test
39
55
  that passes on a rerun is still a bug that will fail a future run. They route
@@ -49,9 +65,9 @@ the two options above. Name the verdict you reached in the `friction` comment.
49
65
  | Verdict | Evidence | Routes to |
50
66
  | --- | --- | --- |
51
67
  | **defect-in-diff** | The failure reproduces on the branch and not on an unmodified `main` | Option 1 — fix at source |
52
- | **pre-existing** | The same check fails on an unmodified `main` too | Option 2 — file `meta::framework-gap`; remediate here only if it blocks this delivery |
53
- | **capacity** | Proven exhaustion of a runner resource, not a property of the diff (see below) | Option 2 — file `meta::framework-gap` **and** escalate to the operator |
54
- | **unreproducible-tier** | The tier cannot be exercised in this sandbox at all, proven by an attempted attach (see below) | Option 2 — file `meta::framework-gap` **and** escalate on first encounter |
68
+ | **pre-existing** | The same check fails on an unmodified `main` too | Option 2 — `file-ci-gap.js --verdict pre-existing`; remediate here only if it blocks this delivery |
69
+ | **capacity** | Proven exhaustion of a runner resource, not a property of the diff (see below) | Option 2 — `file-ci-gap.js --verdict capacity` (`meta::framework-gap` unless `--owner` routes it elsewhere) **and** escalate to the operator |
70
+ | **unreproducible-tier** | The tier cannot be exercised in this sandbox at all, proven by an attempted attach (see below) | Option 2 — `file-ci-gap.js --verdict unreproducible-tier` (`meta::framework-gap` unless `--owner` routes it elsewhere) **and** escalate on first encounter |
55
71
 
56
72
  Why the verdict set carries these last two is recorded in
57
73
  [`docs/decisions.md` ADR 20260906-5160a](../../docs/decisions.md).
@@ -72,10 +88,12 @@ line naming the exhausted limit (an OOM kill, `ENOSPC`, `EMFILE`,
72
88
  timeout), plus the fact that the failure is not specific to this diff. Absent
73
89
  that reading the verdict is **flaky, not capacity**, and it routes to Option 1.
74
90
 
75
- On a `capacity` verdict: file the `meta::framework-gap` issue with the run link,
76
- the failure signature, and the resource reading; flip the Story to
77
- `agent::blocked` with a `friction` comment naming the verdict; and hand back to
78
- the operator, who owns the runner pool. Do not sit in a retry loop waiting for
91
+ On a `capacity` verdict: run `file-ci-gap.js --verdict capacity --block`, passing
92
+ the resource reading as `--evidence` (the run link and failure signature come
93
+ from the digest). That files the intake issue — `meta::framework-gap`, or
94
+ `meta::platform-gap` when `--owner platform` names a shared runner fleet — posts
95
+ the `friction` comment and flips the Story in one call; then hand back to the
96
+ operator, who owns the runner pool. Do not sit in a retry loop waiting for
79
97
  capacity to return.
80
98
 
81
99
  **Rerunning a failed job to reach green stays forbidden under every verdict,
@@ -106,10 +124,9 @@ both:
106
124
  failure in the app under test.
107
125
 
108
126
  Absent both readings the verdict is unavailable and the failure routes as it did
109
- before. On the verdict: file the `meta::framework-gap` issue with the run link,
110
- the failure signature, and the attach attempt; flip the Story to
111
- `agent::blocked` with a `friction` comment naming the verdict; and hand back to
112
- the operator, who owns the sandbox. Do not author a fix for a tier you could not
127
+ before. On the verdict: run
128
+ `file-ci-gap.js --verdict unreproducible-tier --block`, passing the failed attach
129
+ as `--evidence`; then hand back to the operator, who owns the sandbox. Do not author a fix for a tier you could not
113
130
  run — a blind fix to a suite nobody exercised is how the gap compounds.
114
131
 
115
132
  ## Verifier
@@ -131,8 +148,9 @@ alongside the failing check-run identity. On green it adjudicates:
131
148
 
132
149
  - **Same head SHA** → the green came from re-running the failed job. The
133
150
  watcher exits non-zero, flips the Story to `agent::blocked` with a
134
- `friction` comment, and requires the `meta::framework-gap` issue (run link +
135
- failure signature, both already in the digest) before the delivery proceeds.
151
+ `friction` comment, and requires the CI-gap intake issue
152
+ (`file-ci-gap.js` — run link and failure signature are already in the
153
+ digest) before the delivery proceeds.
136
154
  - **New head SHA** → fix at source. The digest is retired, auto-merge is
137
155
  re-armed, and the delivery continues unobstructed.
138
156
 
@@ -153,8 +171,8 @@ operator under **any** of:
153
171
  - **Clearly-environmental → escalate immediately.** An unambiguously
154
172
  environmental failure outside your control (runner provisioning, a persistent
155
173
  registry/network outage, a branch-protection or CI misconfiguration, an
156
- expired credential) — file the `meta::framework-gap` issue (with run link +
157
- signature) and escalate on the first encounter rather than burning iterations
174
+ expired credential) — run `file-ci-gap.js --block` and escalate on the
175
+ first encounter rather than burning iterations
158
176
  trying to code around it. A proven-capacity failure is this case: reach the
159
177
  `capacity` verdict above and escalate on the first encounter.
160
178
  - **Unrunnable tier → escalate immediately.** A tier the sandbox cannot host at