mandrel 2.56.0 → 2.57.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (106) hide show
  1. package/.agents/agents/plan-critic.md +13 -18
  2. package/.agents/agents/story-worker.md +25 -34
  3. package/.agents/docs/agentrc-reference.json +0 -30
  4. package/.agents/docs/configuration.md +8 -28
  5. package/.agents/docs/execution-reference.md +5 -5
  6. package/.agents/docs/quality-gates.md +8 -7
  7. package/.agents/instructions.md +9 -10
  8. package/.agents/schemas/agentrc.schema.json +9 -185
  9. package/.agents/schemas/story-deliver-terminal.schema.json +1 -1
  10. package/.agents/scripts/acceptance-eval.js +107 -17
  11. package/.agents/scripts/ceremony-derive.js +191 -0
  12. package/.agents/scripts/check-context-budget.js +28 -33
  13. package/.agents/scripts/check-cyclomatic.js +4 -3
  14. package/.agents/scripts/deliver-light.js +31 -94
  15. package/.agents/scripts/lib/audit-suite/checklist-threading.js +15 -2
  16. package/.agents/scripts/lib/baselines/coverage-updater-cli.js +110 -0
  17. package/.agents/scripts/lib/baselines/crap-preview-scan.js +25 -0
  18. package/.agents/scripts/lib/baselines/crap-updater-cli.js +223 -0
  19. package/.agents/scripts/lib/bdd-scenario-budget.js +21 -3
  20. package/.agents/scripts/lib/bootstrap/quality-bootstrap.js +0 -1
  21. package/.agents/scripts/lib/close-validation/gates.js +52 -1
  22. package/.agents/scripts/lib/config/acceptance-eval.js +25 -57
  23. package/.agents/scripts/lib/config/delivery-routing.js +7 -33
  24. package/.agents/scripts/lib/config/explain.js +0 -19
  25. package/.agents/scripts/lib/config/limits.js +18 -78
  26. package/.agents/scripts/lib/config/quality.js +6 -3
  27. package/.agents/scripts/lib/config/runners.js +3 -2
  28. package/.agents/scripts/lib/config-settings-schema-delivery.js +15 -68
  29. package/.agents/scripts/lib/config-settings-schema-quality.js +0 -14
  30. package/.agents/scripts/lib/config-settings-schema.js +16 -143
  31. package/.agents/scripts/lib/crap-engine.js +35 -4
  32. package/.agents/scripts/lib/crap-utils.js +17 -1
  33. package/.agents/scripts/lib/cyclomatic-ceiling.js +19 -7
  34. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  35. package/.agents/scripts/lib/observability/runtime-friction.js +1 -1
  36. package/.agents/scripts/lib/observability/source-classifier.js +1 -0
  37. package/.agents/scripts/lib/orchestration/acceptance-eval-decision.js +5 -4
  38. package/.agents/scripts/lib/orchestration/ceremony-routing.js +19 -73
  39. package/.agents/scripts/lib/orchestration/complexity-gate.js +46 -212
  40. package/.agents/scripts/lib/orchestration/file-assumptions.js +32 -17
  41. package/.agents/scripts/lib/orchestration/light-escalation.js +3 -3
  42. package/.agents/scripts/lib/orchestration/light-suitability.js +66 -233
  43. package/.agents/scripts/lib/orchestration/plan-context.js +181 -387
  44. package/.agents/scripts/lib/orchestration/plan-critic-conditions.js +42 -153
  45. package/.agents/scripts/lib/orchestration/plan-critics-evaluate.js +14 -70
  46. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +300 -0
  47. package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +131 -168
  48. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +118 -297
  49. package/.agents/scripts/lib/orchestration/plan-persist/soft-findings.js +55 -0
  50. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +16 -65
  51. package/.agents/scripts/lib/orchestration/plan-persist/wave-serialisation.js +22 -35
  52. package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +30 -139
  53. package/.agents/scripts/lib/orchestration/planning/memory-pool-advisory.js +61 -223
  54. package/.agents/scripts/lib/orchestration/single-story-close/phases/close-validation.js +5 -0
  55. package/.agents/scripts/lib/orchestration/single-story-close/phases/pre-gate-steps.js +46 -16
  56. package/.agents/scripts/lib/orchestration/story-close/context-budget-writeback.js +213 -0
  57. package/.agents/scripts/lib/orchestration/task-body-validator.js +10 -63
  58. package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +33 -539
  59. package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +21 -414
  60. package/.agents/scripts/lib/orchestration/ticket-validator.js +54 -118
  61. package/.agents/scripts/lib/orchestration/verify-credit.js +69 -24
  62. package/.agents/scripts/lib/story-body/body-format-lints.js +15 -85
  63. package/.agents/scripts/lib/story-body/story-body.js +17 -237
  64. package/.agents/scripts/lib/templates/decomposer-prompts.js +84 -121
  65. package/.agents/scripts/lib/test-isolate/cli-options.js +93 -0
  66. package/.agents/scripts/lib/test-isolate/progress-log.js +45 -0
  67. package/.agents/scripts/lib/test-isolate/render-report.js +97 -0
  68. package/.agents/scripts/lib/test-isolate/run-isolate.js +87 -0
  69. package/.agents/scripts/lib/test-run-credit.js +266 -0
  70. package/.agents/scripts/lib/wave-runner/footprint.js +48 -358
  71. package/.agents/scripts/lib/wave-runner/ready-set.js +6 -5
  72. package/.agents/scripts/lib/workers/crap-worker.js +32 -41
  73. package/.agents/scripts/plan-context.js +7 -9
  74. package/.agents/scripts/plan-critics.js +28 -54
  75. package/.agents/scripts/plan-persist.js +25 -68
  76. package/.agents/scripts/quality-preview.js +51 -0
  77. package/.agents/scripts/run-tests.js +12 -0
  78. package/.agents/scripts/stories-wave-tick.js +23 -45
  79. package/.agents/scripts/test-isolate.js +13 -180
  80. package/.agents/scripts/update-coverage-baseline.js +25 -70
  81. package/.agents/scripts/update-crap-baseline.js +19 -123
  82. package/.agents/skills/core/scope-triage/SKILL.md +3 -3
  83. package/.agents/workflows/audit-clean-code.md +4 -3
  84. package/.agents/workflows/helpers/acceptance-self-eval.md +41 -41
  85. package/.agents/workflows/helpers/code-quality-guardrails.md +4 -4
  86. package/.agents/workflows/helpers/code-review.md +2 -3
  87. package/.agents/workflows/helpers/deliver-digest.md +41 -57
  88. package/.agents/workflows/helpers/deliver-light.md +40 -105
  89. package/.agents/workflows/helpers/deliver-reference.md +1 -1
  90. package/.agents/workflows/helpers/deliver-story-reference.md +37 -58
  91. package/.agents/workflows/helpers/deliver-story.md +9 -13
  92. package/.agents/workflows/helpers/plan-reference.md +132 -219
  93. package/.agents/workflows/mandrel-plan.md +27 -40
  94. package/.agents/workflows/memory-consolidate.md +9 -13
  95. package/docs/CHANGELOG.md +23 -0
  96. package/lib/migrations/index.js +4 -0
  97. package/lib/migrations/steps/2.57.0-retire-delivery-limit-knobs.js +45 -0
  98. package/lib/migrations/steps/2.57.0-retire-planning-limit-knobs.js +59 -0
  99. package/package.json +1 -1
  100. package/.agents/scripts/lib/framework-version.js +0 -39
  101. package/.agents/scripts/lib/orchestration/consolidation-precondition.js +0 -223
  102. package/.agents/scripts/lib/orchestration/plan-persist/fan-out-gate.js +0 -97
  103. package/.agents/scripts/lib/orchestration/planning/decomposer-context.js +0 -26
  104. package/.agents/scripts/lib/orchestration/spec-budget.js +0 -89
  105. package/.agents/scripts/lib/orchestration/spec-spill.js +0 -74
  106. package/.agents/scripts/lib/orchestration/verify-tier-repair.js +0 -107
@@ -22,9 +22,9 @@ maintainability / coverage / crap) proves the code is *healthy*, not that it
22
22
  satisfies *this Story's* acceptance criteria.
23
23
 
24
24
  The loop is **always on** (a hard cutover — there is no flag to disable it) and
25
- **bounded** by `delivery.acceptanceEval.maxRounds` (default 2), which the
26
- resolver clamps into `[1, hard ceiling]` so the cap can never be switched off or
27
- run unbounded. It is **distinct from** the per-run `sibling-coherence` epilogue
25
+ **bounded** by `delivery.acceptanceEval.maxRounds` (default 2; `0` means the
26
+ verdict is scored once with no redraft round). It is **distinct from** the
27
+ per-run `sibling-coherence` epilogue
28
28
  step (`planRunEpilogue` for N>1): this loop is per-Story, per-criterion,
29
29
  mid-delivery, and evaluates the actual work product.
30
30
 
@@ -55,23 +55,21 @@ mid-delivery, and evaluates the actual work product.
55
55
  >
56
56
  > **Whether to spawn fresh at all is routed off the derived change level**
57
57
  > — the same signal `review-depth.js` resolves depth from, so the two
58
- > decisions cannot disagree. Derive it with `deriveChangeLevel` from
59
- > [`review-depth.js`](../../scripts/lib/orchestration/review-depth.js) over
60
- > the **change set your caller computed once** for this Story
61
- > (`computeChangeSet` from
62
- > [`change-set.js`](../../scripts/lib/orchestration/change-set.js); see
63
- > [`deliver-story.md`](deliver-story.md) Step 2), then
64
- > resolve the ceremony per cluster with `resolveCeremonyForRisk` from
65
- > [`ceremony-routing.js`](../../scripts/lib/orchestration/ceremony-routing.js)
66
- > using that `derivedLevel` and `delivery.routing.freshCriticSampleRate`:
67
- > **`high` (the diff touches a sensitive path registered in
68
- > `audit-rules.json`) → `fresh`** (spawn the critic); **`low` (it touches
69
- > none) → `inline`** (the contract-identical inline fallback below),
70
- > **except** the `freshCriticSampleRate` fraction of low-level clusters the
71
- > sampling floor forces `fresh` so a low level never means zero independent
72
- > checking; **`null` / unknown (the diff could not be enumerated) → `fresh` +
73
- > full ceremony** (fail-safe). This chooses fresh-vs-inline **per cluster
74
- > only — it never changes the cluster count**.
58
+ > decisions cannot disagree. Derive it with one script over the Story
59
+ > branch (Story #5313) — never a hand-carried import block:
60
+ >
61
+ > ```bash
62
+ > node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
63
+ > ```
64
+ >
65
+ > It computes the change set once (`files`), derives the level and
66
+ > classes, and resolves the ceremony per cluster (`mode`, `reason`,
67
+ > `verdictOwner`): **`high` (the diff touches a sensitive path registered
68
+ > in `audit-rules.json`) → `fresh`** (spawn the critic); **`low` (it
69
+ > touches none) → `inline`** (the contract-identical inline fallback
70
+ > below); **`null` / unknown (the diff could not be enumerated) → `fresh`
71
+ > + full ceremony** (fail-safe). This chooses fresh-vs-inline **per
72
+ > cluster only — it never changes the cluster count**.
75
73
  >
76
74
  > The routing signal is deliberately **not** a planner-authored risk
77
75
  > verdict: a level the plan asserted about itself was exactly the signal
@@ -80,7 +78,7 @@ mid-delivery, and evaluates the actual work product.
80
78
  >
81
79
  > **Inline-critic path (low-level-routed OR nesting-absent harness).** The
82
80
  > verdict is authored **inline** whenever the risk router above resolves to
83
- > `inline` (a low-risk cluster not caught by the sampling floor), and also as
81
+ > `inline` (a low-risk cluster), and also as
84
82
  > a **fallback** on any harness that cannot spawn the fresh critic.
85
83
  > Dispatching the critic as a nested `Agent` is the fresh-context shape and
86
84
  > works on any harness that carries `Agent` into sub-agents (Claude Code ≥
@@ -102,18 +100,18 @@ mid-delivery, and evaluates the actual work product.
102
100
  > comment (if you block) that the inline fallback was used.
103
101
 
104
102
  The critic:
105
- - Inspects the **change set handed to it in its spawn context** — the one
103
+ + Inspects the **change set handed to it in its spawn context** — the one
106
104
  list computed above — and the Story's inline `acceptance[]` / `verify[]`
107
105
  arrays. Pass the file list explicitly when you dispatch the critic; it
108
106
  does not re-enumerate the diff for itself, so a commit
109
107
  landing mid-ceremony cannot leave the critic scoring a different change
110
108
  than the one that routed it.
111
- - **Runs the `verify[]` commands** and consumes their output as **required
109
+ + **Runs the `verify[]` commands** and consumes their output as **required
112
110
  evidence** when scoring the relevant acceptance items. `verify[]` is not
113
111
  optional advisory pre-flight — a criterion cannot be scored `met` without
114
112
  the supporting `verify[]` evidence where a `verify[]` command is relevant
115
113
  to it.
116
- - **Reuses the credited full-suite run instead of re-paying for it.**
114
+ + **Reuses the credited full-suite run instead of re-paying for it.**
117
115
  Before spawning a `verify[]` entry, classify it with `resolveVerifyCredit`
118
116
  from
119
117
  [`verify-credit.js`](../../scripts/lib/orchestration/verify-credit.js): an
@@ -125,7 +123,7 @@ mid-delivery, and evaluates the actual work product.
125
123
  manufacture a pass. The gate warns on any such entry: the intended shape
126
124
  is scoped `verify[]` entries **plus** the one credited run
127
125
  ([`deliver-digest.md`](deliver-digest.md) § 5).
128
- - **Shares `lint` / `typecheck` evidence with close.** When a
126
+ + **Shares `lint` / `typecheck` evidence with close.** When a
129
127
  `verify[]` command is **byte-identical** to a close-validation gate — in
130
128
  practice only the cheap, command-identical `lint` and `typecheck` gates
131
129
  (`npm run lint` and the resolved `project.commands.typecheck`) — the
@@ -146,19 +144,21 @@ mid-delivery, and evaluates the actual work product.
146
144
  fresh — a false-fresh coverage record without `coverage-final.json`
147
145
  silently weakens the floor. Limit the evidence-share to `lint` and
148
146
  `typecheck`.
149
- - Emits a **cluster** verdict file under `temp/` conforming to
147
+ + Emits a **cluster** verdict file under `temp/` conforming to
150
148
  [`acceptance-eval-verdict.schema.json`](../../schemas/acceptance-eval-verdict.schema.json):
151
149
  one `{ index, criterion, verdict: met|partial|unmet, evidence,
152
150
  verifyEvidence[] }` record per acceptance item **in that cluster**, each
153
151
  `index` being the item's position in the Story's full `acceptance[]`
154
152
  array. A fresh critic **returns that path to you** rather than calling the
155
153
  gate itself.
156
- 2. **Dispatch the round's clusters in parallel, then merge into one verdict.**
157
- The clusters of a round are independent, so dispatch **all** of the round's
158
- fresh critics as N `Agent` calls **in a single assistant turn** —
159
- [`parallel-tooling.md`](parallel-tooling.md) **Rule 3** — never serially,
160
- and never one round per cluster. Clusters routed `inline` are authored in
161
- the same round alongside them.
154
+ 2. **Dispatch the round's clusters in parallel, then merge into one verdict
155
+ (fresh critics only).** When the verdict owner is the **inline self-eval**,
156
+ author **one** verdict file covering every `acceptance[]` item and score it
157
+ in one gate call — the cluster merge below does not apply (Story #5313).
158
+ For fresh critics, the clusters of a round are independent, so dispatch
159
+ **all** of the round's critics as N `Agent` calls **in a single assistant
160
+ turn** — [`parallel-tooling.md`](parallel-tooling.md) **Rule 3** — never
161
+ serially, and never one round per cluster.
162
162
 
163
163
  Then **merge** the cluster verdicts into **one** verdict file under `temp/`:
164
164
  concatenate every cluster's `criteria[]` records and order the merged array
@@ -187,24 +187,24 @@ mid-delivery, and evaluates the actual work product.
187
187
 
188
188
  ```bash
189
189
  node <main-repo>/.agents/scripts/acceptance-eval.js \
190
- --story <storyId> --verdict <merged-verdict-path> \
191
- --expected-criteria <number of acceptance[] items>
190
+ --story <storyId> --verdict <verdict-path>
192
191
  ```
193
192
 
194
- Pass `--expected-criteria` from the `acceptance[]` count you already read
195
- off the Story body: a verdict whose `criteria[]` length differs — a single
196
- cluster's verdict handed over unmerged — is rejected **before scoring**,
197
- with an error naming the merge contract and consuming **no round**.
193
+ The gate reads the Story's `acceptance[]` count itself (Story #5313): a
194
+ verdict whose `criteria[]` length differs — a single cluster's verdict
195
+ handed over unmerged — is rejected **before scoring**, with an error naming
196
+ the merge contract and consuming **no round**. `--expected-criteria` is
197
+ still accepted but redundant; when passed it must agree with that count.
198
198
 
199
199
  The gate validates the verdict against the schema, applies the round cap,
200
200
  emits the per-criterion `acceptance-eval` signal into the retro / feedback
201
201
  substrate, prints a JSON envelope, and exits with one of three decisions:
202
- - **`decision: "proceed"`** (every criterion `met`) → exit 0. Proceed to
202
+ + **`decision: "proceed"`** (every criterion `met`) → exit 0. Proceed to
203
203
  close.
204
- - **`decision: "redraft"`** (some `partial`/`unmet`, rounds remaining) →
204
+ + **`decision: "redraft"`** (some `partial`/`unmet`, rounds remaining) →
205
205
  exit 0. Redraft the flagged criteria (named in `unmetCriteria[]`), commit
206
206
  the fix, and start another round.
207
- - **`decision: "block"`** (round cap reached, criteria still unmet) → exit
207
+ + **`decision: "block"`** (round cap reached, criteria still unmet) → exit
208
208
  non-zero. **Do not proceed to close.** Take the caller's blocked path
209
209
  (transition to `agent::blocked`) and post a `friction` comment naming the
210
210
  unmet criteria and their evidence. Never silently proceed to close.
@@ -28,15 +28,15 @@ hook calls the same script.
28
28
 
29
29
  Cyclomatic complexity (CC) is measured per function by `escomplex` (the same
30
30
  engine the maintainability axis of `check-baselines.js` runs). Two
31
- thresholds, sourced from
32
- `delivery.quality.codingGuardrails.cyclomaticFlag` /
33
- `cyclomaticMustFix`:
31
+ thresholds — the advisory `delivery.quality.codingGuardrails.cyclomaticFlag`
32
+ and the fixed ceiling of 12 (`cyclomaticMustFix` was retired as a config
33
+ key in Story #5313):
34
34
 
35
35
  | CC range | Action |
36
36
  | --- | --- |
37
37
  | ≤ 8 | Pass — no annotation required. |
38
38
  | > 8 (default `cyclomaticFlag`) | **Flag** — `quality:preview` counts the function in its `new-method count over c=<flag>` column. The function is allowed to land but the report names it. |
39
- | > 12 (default `cyclomaticMustFix`) | **Must-fix**: `check-cyclomatic.js` fails when a file gains a function above the ceiling, or when its worst function gets worse than the recorded baseline. |
39
+ | ≥ 12 | **Advisory in `quality:preview`** — listed by file, method and reading; the preview exits 0 on it. **Ratchet in `check-cyclomatic.js`**: fails when a file gains a function above 12, or when its worst function gets worse than the recorded baseline. |
40
40
 
41
41
  `check-cyclomatic.js` is a **ratchet**, not a cliff: `baselines/cyclomatic.json`
42
42
  records the over-ceiling functions a repository already carries, so adopting
@@ -384,9 +384,8 @@ ceremony above:
384
384
  `[HEAD_REF]`, stage explicit paths only, and make **one focused
385
385
  conventional commit per lens** (`fix(<scope>): <description> (review findings batch)`).
386
386
  3. Bounded-attempt semantics extend to the batch: each finding gets **at
387
- most one** attempt, and a lens's batch commit that would exceed
388
- `delivery.codeReview.maxFixScopeFiles` routes that lens's findings to
389
- escalation (`scope-exceeded`) instead of committing.
387
+ most one** attempt (the `maxFixScopeFiles` file-count ceiling was retired
388
+ in Story #5313 — a fix is bounded by attempts, not by file count).
390
389
  4. After **all** lens batches are committed, run a **single** validation
391
390
  pass (`npm run lint` plus the relevant `npm test` slice) and a
392
391
  **single** targeted rescan over the touched files. Surviving batched
@@ -49,37 +49,24 @@ the only sanctioned landing. A silent local build is not a delivery.
49
49
  ## 3. Change set — computed once, handed to everyone
50
50
 
51
51
  One enumeration per Story. A critic that re-runs its own `git diff`
52
- can score a different set than the one that routed it. Both routing calls take
53
- a single options object and are **total — they never throw**, so a wrong-shaped
54
- argument is silently absorbed into the `null` fail-safe:
52
+ can score a different set than the one that routed it. Derive the change
53
+ set, the level and the ceremony with **one script** — never a hand-carried
54
+ import block:
55
55
 
56
56
  ```bash
57
- node --input-type=module -e '
58
- const lib = "<main-repo>/.agents/scripts/lib/orchestration";
59
- const { computeChangeSet } = await import(`${lib}/change-set.js`);
60
- const { deriveChangeLevel } = await import(`${lib}/review-depth.js`);
61
- const { resolveCeremonyForRisk } = await import(`${lib}/ceremony-routing.js`);
62
- const { files } = computeChangeSet({ baseRef: "main", headRef: "story-<storyId>" });
63
- // deriveChangeLevel({ changedFiles, injectedRules?, selectSensitivePathClassesFn? })
64
- // -> { level, classes } — an OBJECT, never a bare level.
65
- const { level, classes } = deriveChangeLevel({ changedFiles: files });
66
- // resolveCeremonyForRisk({ derivedLevel, clusterIndex?, freshCriticSampleRate?,
67
- // ceremonyProfile? }) -> { mode, reason, sampled, profile, verdictOwner }.
68
- // derivedLevel is that level STRING — the object above matches no tier and
69
- // routes to the null fail-safe: a fresh critic, silently.
70
- const ceremony = resolveCeremonyForRisk({ derivedLevel: level, clusterIndex: 0 });
71
- console.log(JSON.stringify({ files, level, classes, ...ceremony }));
72
- '
57
+ node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
73
58
  ```
74
59
 
75
- Level rules ([`review-depth.js`](../../scripts/lib/orchestration/review-depth.js)):
76
- a sensitive path registered in `audit-rules.json` → `high`, none → `low`, an
77
- unenumerable diff (`files === null`) → `null`. Ceremony rules
78
- ([`ceremony-routing.js`](../../scripts/lib/orchestration/ceremony-routing.js)):
79
- `minimal` → always inline, `strict` → always fresh, `standard` → `high`/`null`
80
- → fresh and `low` → inline unless the `freshCriticSampleRate` floor forces
81
- fresh. An `inline` dispatch mode overrides all of it to inline critics. Close's
82
- `review-depth.js` reads the same derived level, so the two cannot disagree.
60
+ It prints one JSON object: `files` — the one change set every critic is
61
+ handed (`null` when the diff could not be enumerated) — plus `level` and
62
+ `classes` from `review-depth.js`, and `mode`, `reason` and `verdictOwner`
63
+ from `ceremony-routing.js`. Level rules: a sensitive path registered in
64
+ `audit-rules.json` → `high`, none → `low`, an unenumerable diff → `null`.
65
+ Ceremony rules: `minimal` → always inline, `strict` → always fresh,
66
+ `standard` → `high`/`null` → fresh and `low` → inline. An `inline` dispatch
67
+ mode overrides all of it to inline critics. Close's `review-depth.js` reads
68
+ the same derived level, so the two cannot disagree. `--base <ref>` overrides
69
+ `project.baseBranch`.
83
70
 
84
71
  ## 4. Acceptance self-eval (Step 1a, required)
85
72
 
@@ -87,52 +74,49 @@ fresh. An `inline` dispatch mode overrides all of it to inline critics. Close's
87
74
  self-eval, named by `verdictOwner`, never both and never a warm-up pass. Each
88
75
  scores its cluster's `acceptance[]` items against the change set above, with
89
76
  `verify[]` output as evidence. Bounded by `delivery.acceptanceEval.maxRounds`
90
- (default 2).
77
+ (default 2; `0` scores once with no redraft).
91
78
 
92
- **One round = N cluster critics → ONE merged verdict → ONE gate call.** Merge
93
- every cluster's records into a single `criteria[]` in `acceptance[]` order, one
94
- per acceptance item, and score that once. A gate call per cluster spends a
95
- round *per cluster* and races the round ledger:
79
+ **Inline owner:** author **one** verdict file covering every `acceptance[]`
80
+ item and score it in **one** gate call — there is no cluster merge.
81
+ **Fresh critics:** one round = N cluster critics → ONE merged verdict → ONE
82
+ gate call. Merge every cluster's records into a single `criteria[]` in
83
+ `acceptance[]` order and score that once; a gate call per cluster spends a
84
+ round *per cluster* and races the round ledger.
96
85
 
97
86
  `node <main-repo>/.agents/scripts/acceptance-eval.js --story <storyId>
98
- --verdict <merged-verdict-path> --expected-criteria <acceptance[] count>`
87
+ --verdict <verdict-path>`
99
88
 
100
- Pass `--expected-criteria` — **without it the coverage assertion is inert**, so
101
- an unmerged cluster verdict scores a fraction of the criteria and still reports
102
- `proceed`. A mismatch is rejected before scoring and costs no round.
89
+ The gate reads the Story's `acceptance[]` count itself and rejects a verdict
90
+ whose `criteria[]` length differs **before** scoring, consuming no round;
91
+ `--expected-criteria` is accepted but redundant.
103
92
 
104
93
  `proceed` → close. `redraft` → one more round inside the cap. `block` → **do
105
94
  not close**: post a `friction` comment and flip `agent::blocked`.
106
95
  Per-round mechanics: [`acceptance-self-eval.md`](acceptance-self-eval.md).
107
96
 
108
- ## 5. The one creditable full-suite run
97
+ ## 5. The one full-suite run
109
98
 
110
- **After the self-eval loop's last fix commit, immediately after the push** —
111
- the credit is keyed on the tree, not push state: a later commit voids it, a
112
- push does not, and the backgrounded capture (below) ends the turn. Redraft
113
- rounds run scoped tests; only this run needs credit, and a bare `npm test` /
114
- `pnpm run test` deposits **none**, so close re-runs it. Shape it by what
115
- `close-validation/gates.js` runs:
99
+ After the self-eval loop's last fix commit, run the project test runner
100
+ **once** in the worktree:
116
101
 
117
102
  ```bash
118
- # CRAP gate on (default) + a `test:coverage` script — writes close's stamp:
119
- node <main-repo>/.agents/scripts/coverage-capture.js --cwd <workCwd>
120
- # otherwise — the record close's `test` gate reads; <workCwd> ABSOLUTE,
121
- # runner exactly `npm test` (both sides hash {cmd, args, cwd}):
122
- node <main-repo>/.agents/scripts/evidence-gate.js --standalone \
123
- --scope-id <storyId> --gate test --worktree <workCwd> -- npm test
103
+ npm test # in <workCwd>
124
104
  ```
125
105
 
126
- Dispatch it in the **background**: it outruns the host's sync Bash ceiling, and its completion re-invokes you. Never spawn a task to poll or
127
- `sleep`-loop against it ([`parallel-tooling.md`](parallel-tooling.md)
128
- Rule 2).
129
-
130
- Read the **output**, not the exit code: capture skips — no test run, no
131
- credit — when nothing changed under CRAP `targetDirs`. Run the scoped
132
- projects for the roots you changed plus `verify[]`, not the whole suite.
106
+ A green full run on `story-<id>` deposits the `test` evidence close reads,
107
+ keyed on the tree, so close reports the gate as **credited** at unchanged
108
+ HEAD — a later commit voids it. The CRAP gate still captures coverage itself
109
+ when it needs an artifact. If the suite outruns the host's sync Bash
110
+ ceiling, dispatch it in the **background** — its completion re-invokes you;
111
+ never spawn a task to poll or `sleep`-loop against it
112
+ ([`parallel-tooling.md`](parallel-tooling.md) Rule 2). Read the **output**,
113
+ not the exit code: the runner prints whether it deposited credit, and a run
114
+ off the Story branch or of a partial tier deposits nothing and says so.
115
+ Redraft rounds run the scoped projects for the roots you changed plus
116
+ `verify[]`, not the whole suite; only this run needs credit.
133
117
 
134
118
  `verify[]` is scoped entries **plus** this one run: an entry that is itself a
135
- full-suite command is reported credited against the same stamp, never
119
+ full-suite command is reported credited against the same record, never
136
120
  respawned.
137
121
 
138
122
  ## 6. Terminal envelope — the return contract
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  description:
3
- The unplanned prompt path shared by /mandrel-deliver and /mandrel-plan Gate #1. Judges a
3
+ The unplanned prompt path /mandrel-deliver takes for a free-text prompt. Judges a
4
4
  prompt's predicted footprint, authors a receipt Story, then lands it through
5
5
  the same single-story-init / single-story-close engine — every close gate
6
6
  unchanged.
@@ -9,9 +9,9 @@ description:
9
9
  # Unplanned delivery (the prompt path)
10
10
 
11
11
  > **A path, not a command.** There is no `/deliver-light` to type. This file is
12
- > reached two ways — `/mandrel-deliver "<prompt>"` (an operator describing small work)
13
- > and `/mandrel-plan` Gate #1 (a seed the suggestion says fits the light ceilings, once
14
- > the operator confirms). Read
12
+ > reached one way — `/mandrel-deliver "<prompt>"` (an operator describing small
13
+ > work); Story #5312 retired the `/mandrel-plan` Gate #1 suggestion that used
14
+ > to be the second door. Read
15
15
  > [`deliver-digest.md`](deliver-digest.md) once first — the engine invariants,
16
16
  > gates, and terminal-envelope contract below are its.
17
17
 
@@ -24,13 +24,11 @@ straight to execution from a prompt, landing through the unchanged close path.
24
24
  It never relaxes a close gate, never bypasses the PR to `main`, and never lands
25
25
  over-scope work silently.
26
26
 
27
- Two callers, one gate: whichever door you arrived through, the suitability gate
28
- below is the decision. A `/mandrel-plan` Gate #1 suggestion is a *suggestion* — it is
29
- read against seed-time signals (`DELIVER_LIGHT_SUGGESTION_CEILINGS`: artifacts,
30
- risk hits, sensitive-path classes), while the gate here is read against the
31
- predicted work's *effort and risk* (`STORY_SHAPE_CEILINGS`: change kinds,
32
- magnitude, uncertainty, deployable span). They are deliberately two different
33
- checks, so the gate still runs after a confirm.
27
+ One gate, whatever the door: the suitability gate below is the decision, read
28
+ against the predicted work's *effort and risk* (`STORY_SHAPE_CEILINGS`: change
29
+ kinds, magnitude, uncertainty, deployable span). `/mandrel-plan` no longer
30
+ suggests this path at its Gate #1 (Story #5312) — a prompt reaches it through
31
+ `/mandrel-deliver`, and the gate runs there every time.
34
32
 
35
33
  ## Scope by effort, not by artifact count {#scope-by-effort}
36
34
 
@@ -55,8 +53,10 @@ obeying its own test-first rule is a ceiling that over-fires.
55
53
 
56
54
  Sensitivity is the exception and stays absolute: a footprint touching an auth,
57
55
  crypto, billing, or migration class routes `full` however small or mechanical —
58
- and unlike a ceiling, it is **not overridable** (§ Recording a proceed-light
59
- answer).
56
+ and unlike a ceiling, it is never a warning. The shape ceilings themselves
57
+ stopped gating in Story #5313: a prediction past one is carried as a
58
+ `warnings[]` entry on the envelope, naming the exceeded axis, and the run
59
+ proceeds — `LIGHT_DIFF_CEILINGS` in the backstop is the only size block.
60
60
 
61
61
  ## Four invariants (do not skip one)
62
62
 
@@ -64,12 +64,13 @@ answer).
64
64
  shared effort/risk machinery (`deriveStoryShape` / `deriveChangeLevel`)
65
65
  **and** a ledgered model verdict with a recorded reason. Both must agree on
66
66
  `lite`.
67
- 2. **Over-scope stops — it never hard-fails.** An over-ceiling prompt STOPS and
68
- asks the operator to escalate to `/mandrel-plan` or proceed light. **Both answers
69
- are executable** — `--operator-proceed-light` records the second one
70
- (§ Recording a proceed-light answer). Under `--yes` it fails closed to an
67
+ 2. **The predicted shape warns; risk escalates.** A prompt past a shape
68
+ ceiling proceeds light with a `warnings[]` entry (Story #5313) — the
69
+ prediction is a guess the backstop bounds for real. Only an un-ledgered
70
+ verdict or an un-waivable risk rule (a sensitive-path class, a migration
71
+ span) refuses, and it refuses the same way attended or not: an
71
72
  **`escalated` terminal envelope** that ends the session (§ Escalation is
72
- terminal).
73
+ terminal). There is no question to wait for and no answer flag.
73
74
  3. **Diff-derived backstop.** After implementation the ACTUAL change set is
74
75
  re-checked — the diff is the real scope signal — and an over-ceiling diff is
75
76
  blocked rather than landed.
@@ -95,30 +96,16 @@ answer).
95
96
 
96
97
  Branch on `action` in the JSON envelope:
97
98
  - **`proceed-light`** — the receipt Story is authored; read `storyId` and
98
- `nextCommands`. Continue to step 2.
99
- - **`ask-operator`** — predicted scope exceeds the light ceilings. STOP and
100
- ask the operator to escalate to `/mandrel-plan` or proceed light. Do not proceed
101
- on your own. This is a **question, not a terminal** — wait for the answer,
102
- then act on it: *escalate* leaves for `/mandrel-plan`, *proceed light* re-runs the
103
- same command with `--operator-proceed-light "<their reason>"`
104
- (§ Recording a proceed-light answer).
105
- - **over-scope under `--yes`** — no `action` to branch on: the gate emits an
106
- **`escalated` terminal envelope** instead (exit 2). § Escalation is
107
- terminal governs; you are finished.
99
+ `nextCommands`. Relay every `warnings[]` entry to the operator verbatim
100
+ (each names the shape axis the prediction exceeded), then continue to
101
+ step 2 — the diff backstop in step 4 is what bounds the actual change.
102
+ - **escalation** — no `action` to branch on: the gate emits an
103
+ **`escalated` terminal envelope** instead (exit 2), attended or not.
104
+ § Escalation is terminal governs; you are finished.
108
105
 
109
106
  `--amends '#<id>'` is the canonical light case — shape-checked identically; a
110
107
  heavy amendment escalates to `/mandrel-plan` like any other over-scope prompt.
111
108
 
112
- **Entered from `/mandrel-plan` Gate #1?** Fill `--creates` / `--refactors` /
113
- `--acceptance` / `--reason` from the plan-context envelope's codebase
114
- snapshot and `complexitySignals` rather than re-deriving them from the seed
115
- text — Gate #1 has already done that work, and re-deriving throws away the
116
- better signal. An `ask-operator` verdict here means the two ceiling sets
117
- disagreed: **return to [`../mandrel-plan.md`](../mandrel-plan.md) step 2 (Author) in the same
118
- session**, carrying the interrogation you already paid for. That bounce-back
119
- is not an escalation and does not need a fresh session (§ Why the two
120
- directions differ).
121
-
122
109
  2. **Init (same engine).** From the main checkout, synchronously, with the
123
110
  maximum Bash timeout:
124
111
 
@@ -177,42 +164,10 @@ answer).
177
164
  [`deliver-digest.md`](deliver-digest.md) § 5 — every close
178
165
  gate runs byte-identical to the full path.
179
166
 
180
- ## Recording a proceed-light answer {#recording-a-proceed-light-answer}
181
-
182
- The gate offers the operator two options, so **both** have to be executable.
183
- Re-run the identical gate command with their answer appended:
184
-
185
- ```bash
186
- node .agents/scripts/deliver-light.js --prompt "<prompt>" … \
187
- --operator-proceed-light "<the operator's reason, in their words>"
188
- ```
189
-
190
- The gate then proceeds light, records the decision in the receipt Story, and
191
- returns it on the envelope's `override`. Do **not** instead re-shape the
192
- prediction — shrinking `--refactors` until the gate stops objecting is
193
- under-declaring the footprint, which is the one thing the coarse design must
194
- not reward.
195
-
196
- It is deliberately narrow, and a refusal is printed rather than silent:
197
-
198
- - **Only a size prediction is waivable** — change kinds, magnitude,
199
- uncertainty, deployable span. A sensitive-path class, a
200
- migration-with-consumers span, and an unknown footprint (undeclared, glob,
201
- no acceptance, unclassifiable) are refused: those are risk, not size, and
202
- § Scope by effort keeps them absolute.
203
- - **The ledgered verdict still stands on its own.** The override substitutes
204
- for the predicted *shape* only; `--route lite --reason "<why>"` is still
205
- required.
206
- - **Attended-only.** With `--yes` it is a usage error, not a quiet no-op —
207
- an unattended run has no operator whose answer this could be, and over-scope
208
- there still fails closed (§ Escalation is terminal).
209
-
210
- What licenses this at all is step 4: the operator waives a *guess*, never the
211
- diff backstop, which re-checks the actual change set against ground truth.
212
-
213
167
  ## Escalation is terminal {#escalation-is-terminal}
214
168
 
215
- Over-scope under `--yes` emits a schema-validated `story-deliver-terminal`
169
+ A refused gate — an un-ledgered verdict or an un-waivable risk rule, under
170
+ `--yes` or attended alike — emits a schema-validated `story-deliver-terminal`
216
171
  envelope with **`status: "escalated"`**, `storyId: null`, and a `nextCommand`
217
172
  naming the `/mandrel-plan` invocation that owns the work.
218
173
 
@@ -237,51 +192,31 @@ Nothing is left half-started: an escalated run creates **no receipt Story, no
237
192
  every creation call site, and `escalation.created` records all three as `false`
238
193
  in a shape the schema pins, so a later run finds nothing to trip over.
239
194
 
240
- ## Why the two directions differ {#why-the-two-directions-differ}
195
+ ## Why escalation breaks the session {#why-the-two-directions-differ}
241
196
 
242
- Traffic runs both ways between this path and `/mandrel-plan`, and the two directions
243
- have **deliberately different session rules**. It reads like an inconsistency;
244
- it is not. The rule:
197
+ Traffic runs one way between this path and `/mandrel-plan` — light →
198
+ `/mandrel-plan` on an over-scope prompt — and that direction has a deliberate
199
+ session rule:
245
200
 
246
- > **The direction whose guard is model judgment must break the session. The
247
- > direction whose guard is mechanical need not.**
201
+ > **The direction whose guard is model judgment must break the session.**
248
202
 
249
203
  **Light → `/mandrel-plan` must be a fresh session.** What is being protected is
250
204
  *authoring judgment*, and the empirical finding above is that a session already
251
205
  framed as small work under-decomposes — one Story against a 3–5 contract where
252
206
  a fresh session on the identical seed authored four. The frame is the hazard,
253
- so only a new session removes it.
254
-
255
- **`/mandrel-plan` → light may stay in-session.** Gate #1 fires **before** authoring, so
256
- there is no authoring to corrupt, and the frame at that point is "plan this
257
- seed" — the neutral one, not the small one. Everything on the receiving side is
258
- mechanical: `STORY_SHAPE_CEILINGS`, the ledgered verdict, the diff backstop.
259
- None of them degrade because the context is large, so nothing is gained by
260
- paying for a fresh session.
261
-
262
- Do not "fix" this into symmetry in either direction. Making `/mandrel-plan` → light
263
- require a fresh session throws away a paid-for interrogation for no guard.
264
- Letting light → `/mandrel-plan` run in-session reintroduces the exact failure the
265
- `escalated` envelope exists to prevent.
266
-
267
- ## Constraints
268
-
269
- - **Land, block, or escalate — never a silent local build.** The close push is
270
- the only sanctioned landing; an `escalated` terminal is the only sanctioned
271
- ending that delivers nothing, and it ends the session
272
- (§ Escalation is terminal).
273
- - **No parallel engine.** This path invokes `single-story-init.js` and
274
- `single-story-close.js`; it never reimplements worktree, branch, PR, or merge
275
- mechanics.
276
- - **State only via `update-ticket-state.js`.** Drive every `agent::*`
277
- transition through the script; report state, not process.
207
+ so only a new session removes it. Everything on this path's own side is
208
+ mechanical — `STORY_SHAPE_CEILINGS`, the ledgered verdict, the diff backstop —
209
+ and none of it degrades because the context is large.
210
+
211
+ Do not "fix" this by letting light → `/mandrel-plan` run in-session: that
212
+ reintroduces the exact failure the `escalated` envelope exists to prevent.
278
213
 
279
214
  ## See also
280
215
 
281
216
  - [`/mandrel-deliver`](../mandrel-deliver.md) — the delivery entry point; routes here on a
282
217
  free-text prompt.
283
- - [`/mandrel-plan`](../mandrel-plan.md) — routes here from Gate #1 on a confirmed suggestion,
284
- and owns the work an over-scope prompt escalates to.
218
+ - [`/mandrel-plan`](../mandrel-plan.md) — owns the work an over-scope prompt
219
+ escalates to.
285
220
  - [`deliver-story.md`](deliver-story.md) — the one Story delivery engine every
286
221
  path shares.
287
222
  - [`deliver-digest.md`](deliver-digest.md) — engine invariants, gates, and the
@@ -361,7 +361,7 @@ sensitive-path classes in `audit-rules.json`
361
361
  | Profile | Acceptance critic | When to use |
362
362
  | --- | --- | --- |
363
363
  | `minimal` | Always inline | Tiny trusted N=1 Stories |
364
- | `standard` | Derived-level routed (+ sampling floor) | Default |
364
+ | `standard` | Derived-level routed | Default |
365
365
  | `strict` | Always fresh-context | High-assurance / regulated surfaces |
366
366
 
367
367
  | Scope | What runs | Mechanism |