mandrel 2.58.0 → 2.60.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (124) hide show
  1. package/.agents/README.md +17 -12
  2. package/.agents/agents/acceptance-critic.md +24 -43
  3. package/.agents/agents/story-worker.md +18 -19
  4. package/.agents/docs/SDLC.md +12 -13
  5. package/.agents/docs/agentrc-reference.json +1 -2
  6. package/.agents/docs/configuration.md +29 -46
  7. package/.agents/docs/quality-gates.md +9 -5
  8. package/.agents/docs/workflows.md +1 -1
  9. package/.agents/instructions.md +5 -7
  10. package/.agents/rules/ci-remediation.md +41 -8
  11. package/.agents/rules/known-tooling-behavior.md +65 -15
  12. package/.agents/runtime-deps.json +7 -2
  13. package/.agents/schemas/acceptance-eval-verdict.schema.json +1 -1
  14. package/.agents/schemas/agentrc.schema.json +6 -11
  15. package/.agents/schemas/crap-baseline.schema.json +1 -1
  16. package/.agents/schemas/crap-report.schema.json +1 -1
  17. package/.agents/schemas/story-deliver-terminal.schema.json +3 -3
  18. package/.agents/scripts/README.md +11 -1
  19. package/.agents/scripts/acceptance-eval.js +25 -27
  20. package/.agents/scripts/ceremony-derive.js +15 -10
  21. package/.agents/scripts/check-context-budget.js +148 -228
  22. package/.agents/scripts/check-schema-references.js +5 -3
  23. package/.agents/scripts/check-workflow-citations.js +33 -147
  24. package/.agents/scripts/coverage-capture.js +7 -4
  25. package/.agents/scripts/deliver-light.js +41 -100
  26. package/.agents/scripts/deliver-run.js +631 -0
  27. package/.agents/scripts/file-ci-gap.js +59 -11
  28. package/.agents/scripts/install-matrix-assert.js +48 -3
  29. package/.agents/scripts/lib/audit-to-stories/seed-from-findings.js +51 -33
  30. package/.agents/scripts/lib/baselines/crap-preview-incremental.js +6 -2
  31. package/.agents/scripts/lib/baselines/kinds/_crap-read.js +0 -8
  32. package/.agents/scripts/lib/baselines/kinds/crap.js +35 -18
  33. package/.agents/scripts/lib/changed-files.js +30 -0
  34. package/.agents/scripts/lib/config/delivery-routing.js +5 -4
  35. package/.agents/scripts/lib/config/explain.js +1 -3
  36. package/.agents/scripts/lib/config/gates/crap-incremental-coverage.schema.js +1 -1
  37. package/.agents/scripts/lib/config-resolver.js +1 -0
  38. package/.agents/scripts/lib/config-settings-schema-delivery.js +28 -21
  39. package/.agents/scripts/lib/coverage-capture-fullscope.js +10 -2
  40. package/.agents/scripts/lib/coverage-capture-incremental.js +3 -2
  41. package/.agents/scripts/lib/coverage-capture-usage.js +4 -1
  42. package/.agents/scripts/lib/crap-engine.js +2 -2
  43. package/.agents/scripts/lib/crap-utils.js +21 -5
  44. package/.agents/scripts/lib/doc-tiers.js +4 -2
  45. package/.agents/scripts/lib/escomplex-ast-compat.js +39 -17
  46. package/.agents/scripts/lib/escomplex-kernel.js +298 -0
  47. package/.agents/scripts/lib/feedback-loop/graduator-core.js +7 -6
  48. package/.agents/scripts/lib/feedback-loop/retro-proposals-graduator.js +7 -5
  49. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  50. package/.agents/scripts/lib/gh-exec.js +160 -0
  51. package/.agents/scripts/lib/maintainability-engine.js +3 -3
  52. package/.agents/scripts/lib/observability/source-classifier.js +1 -0
  53. package/.agents/scripts/lib/orchestration/ceremony-routing.js +74 -132
  54. package/.agents/scripts/lib/orchestration/ci-rerun-guard.js +123 -12
  55. package/.agents/scripts/lib/orchestration/complexity-gate.js +180 -352
  56. package/.agents/scripts/lib/orchestration/light-suitability.js +71 -136
  57. package/.agents/scripts/lib/orchestration/plan-context.js +44 -50
  58. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +8 -6
  59. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +104 -119
  60. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +41 -25
  61. package/.agents/scripts/lib/orchestration/plan-persist/summary.js +11 -11
  62. package/.agents/scripts/lib/orchestration/plan-persist/supersede-ops.js +63 -29
  63. package/.agents/scripts/lib/orchestration/plan-persist/wave-collision-gate.js +107 -0
  64. package/.agents/scripts/lib/orchestration/review-depth.js +14 -11
  65. package/.agents/scripts/lib/orchestration/run-epilogue.js +260 -182
  66. package/.agents/scripts/lib/orchestration/run-scoped-config.js +63 -99
  67. package/.agents/scripts/lib/orchestration/single-story-close/phases/base-sync.js +3 -3
  68. package/.agents/scripts/lib/orchestration/single-story-close/phases/graphql-preflight.js +137 -0
  69. package/.agents/scripts/lib/orchestration/single-story-close/runner.js +105 -18
  70. package/.agents/scripts/lib/orchestration/story-deliver-terminal.js +4 -3
  71. package/.agents/scripts/lib/orchestration/story-follow-ups.js +156 -39
  72. package/.agents/scripts/lib/orchestration/story-init-envelope.js +71 -0
  73. package/.agents/scripts/lib/orchestration/task-body-validator.js +8 -17
  74. package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +25 -209
  75. package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +8 -5
  76. package/.agents/scripts/lib/orchestration/ticket-validator.js +44 -183
  77. package/.agents/scripts/lib/orchestration/ticketing/reads.js +14 -25
  78. package/.agents/scripts/lib/runtime-deps/dep-resolution.js +155 -0
  79. package/.agents/scripts/lib/runtime-deps/ensure-installed.js +44 -9
  80. package/.agents/scripts/lib/runtime-deps/parser-major.js +110 -0
  81. package/.agents/scripts/lib/runtime-deps/preflight.js +6 -25
  82. package/.agents/scripts/lib/runtime-deps/scan-imports.js +46 -1
  83. package/.agents/scripts/lib/skills/walk-skill-files.js +1 -1
  84. package/.agents/scripts/lib/story-body/body-format-lints.js +58 -12
  85. package/.agents/scripts/lib/story-body/story-body.js +83 -29
  86. package/.agents/scripts/lib/templates/decomposer-prompts.js +28 -33
  87. package/.agents/scripts/lib/wave-runner/live-probe.js +31 -5
  88. package/.agents/scripts/merge-baseline.js +4 -5
  89. package/.agents/scripts/plan-context.js +117 -28
  90. package/.agents/scripts/plan-persist.js +79 -39
  91. package/.agents/scripts/plan-run-epilogue.js +11 -8
  92. package/.agents/scripts/pr-watch-with-update.js +9 -2
  93. package/.agents/scripts/run-verify.js +13 -6
  94. package/.agents/scripts/single-story-init.js +7 -57
  95. package/.agents/scripts/stories-wave-tick.js +160 -26
  96. package/.agents/skills/core/gates-and-baselines/reference.md +0 -1
  97. package/.agents/skills/skills.index.json +2 -12
  98. package/.agents/skills/stack/qa/playwright/SKILL.md +26 -0
  99. package/.agents/workflows/audit-to-stories.md +14 -11
  100. package/.agents/workflows/helpers/acceptance-self-eval.md +84 -157
  101. package/.agents/workflows/helpers/code-review.md +4 -2
  102. package/.agents/workflows/helpers/deliver-digest.md +31 -24
  103. package/.agents/workflows/helpers/deliver-light.md +92 -101
  104. package/.agents/workflows/helpers/deliver-reference.md +116 -100
  105. package/.agents/workflows/helpers/deliver-story-reference.md +58 -124
  106. package/.agents/workflows/helpers/deliver-story.md +17 -18
  107. package/.agents/workflows/helpers/plan-reference.md +82 -60
  108. package/.agents/workflows/mandrel-deliver.md +47 -31
  109. package/.agents/workflows/mandrel-plan.md +32 -30
  110. package/.agents/workflows/mandrel-update.md +36 -21
  111. package/README.md +3 -3
  112. package/docs/CHANGELOG.md +43 -0
  113. package/lib/cli/registry.js +45 -25
  114. package/lib/cli/update.js +376 -17
  115. package/lib/migrations/index.js +2 -0
  116. package/lib/migrations/steps/2.60.0-retire-audit-results-autofile.js +40 -0
  117. package/package.json +8 -2
  118. package/.agents/schemas/model-attribution.schema.json +0 -53
  119. package/.agents/scripts/lib/orchestration/model-attribution.js +0 -418
  120. package/.agents/scripts/lib/orchestration/split-policy-validator.js +0 -188
  121. package/.agents/scripts/lib/orchestration/story-plan-state.js +0 -33
  122. package/.agents/scripts/lib/orchestration/structured-comment-parser.js +0 -67
  123. package/.agents/scripts/lib/templates/spec-author-prompts.js +0 -76
  124. package/.agents/skills/core/scope-triage/SKILL.md +0 -48
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  description: >-
3
3
  Shared include for the bounded acceptance self-eval loop run during Story
4
- delivery (`helpers/deliver-story`). Defines the per-round critic mechanic;
4
+ delivery (`helpers/deliver-story`). Defines the per-round verdict mechanic;
5
5
  the caller supplies its gate-decision wrapper (label transitions).
6
6
  ---
7
7
 
@@ -9,181 +9,103 @@ description: >-
9
9
 
10
10
  > **Include module.** Not a slash command. Referenced from
11
11
  > [`deliver-story.md`](deliver-story.md) at Step 1a. This file is the
12
- > **single prose home** for the per-round critic mechanic; the caller layers
12
+ > **single prose home** for the per-round verdict mechanic; the caller layers
13
13
  > only its wrapper (Story label transitions).
14
14
 
15
15
  After the implementation commits land and **before** the Story proceeds to
16
- close, run an explicit, **independent** eval pass that scores the change set
17
- computed once for this Story and injected into the critic — never one the
18
- critic re-derives — against **each** `acceptance[]` item
19
- individually. This is the acceptance gate
20
- the close-validation chain does not provide: that chain (lint / test / format /
16
+ close, run an explicit eval pass that scores the change set computed once for
17
+ this Story — never one the evaluator re-derives — against **each**
18
+ `acceptance[]` item individually. This is the acceptance gate the
19
+ close-validation chain does not provide: that chain (lint / test / format /
21
20
  maintainability / coverage / crap) proves the code is *healthy*, not that it
22
21
  satisfies *this Story's* acceptance criteria.
23
22
 
24
23
  The loop is **always on** (a hard cutover — there is no flag to disable it) and
25
24
  **bounded** by `delivery.acceptanceEval.maxRounds` (default 2; `0` means the
26
25
  verdict is scored once with no redraft round). It is **distinct from** the
27
- per-run `sibling-coherence` epilogue
28
- step (`planRunEpilogue` for N>1): this loop is per-Story, per-criterion,
29
- mid-delivery, and evaluates the actual work product.
26
+ per-run epilogue (`planRunEpilogue` for N>1): this loop is per-Story,
27
+ per-criterion, mid-delivery, and evaluates the actual work product.
30
28
 
31
29
  ## Per round
32
30
 
33
- 1. **Eval pass — one verdict-owner per cluster.** Exactly
34
- **one** pass authors each cluster's verdict: the **fresh-context critic**
35
- when the ceremony routing below resolves `fresh` (a sub-agent via the
36
- `Agent` tool, *not* a continuation of your implementing turn — the
37
- evaluator does not grade its own homework), or the **inline self-eval**
38
- when it resolves `inline`. The resolved decision names the owner
39
- explicitly (`verdictOwner: 'fresh-critic' | 'inline-self-eval'` from
40
- `resolveCeremonyForRisk`). **Never run both**, and never run a
41
- preliminary self-assessment pass before dispatching the fresh critic —
42
- the redundant pre-pass buys no measurable quality and roughly triples
43
- the acceptance-block cost. Step 3's gate is the deterministic **scorer**
44
- of the one merged verdict, not a second (or third) pass over the
45
- criteria.
31
+ 1. **Eval pass — one verdict owner, one verdict file.** Exactly **one** pass
32
+ authors the Story's verdict, and it covers **every** `acceptance[]` item in
33
+ one file. Which pass is named by the ceremony decision
34
+ (`verdictOwner: 'fresh-critic' | 'inline-self-eval'` from
35
+ `resolveCeremonyForRisk`), and since Story #5343 that follows the
36
+ **ceremony profile alone**:
46
37
 
47
- > **Sub-agent type + derived-level ceremony.** When
48
- > `delivery.routing.roleScopedAgents` is enabled (the **default**), dispatch
49
- > the critic with `subagent_type: acceptance-critic` — it boots on the
50
- > role-scoped [`acceptance-critic`](../../agents/acceptance-critic.md) context
51
- > (its own system prompt, no `CLAUDE.md` @-closure) that carries the
52
- > maker-blind invariant and the verdict schema standalone. When the
53
- > kill-switch is **off** (`roleScopedAgents: false`), fall back to
54
- > `subagent_type: general-purpose`.
55
- >
56
- > **Whether to spawn fresh at all is routed off the derived change level**
57
- > — the same signal `review-depth.js` resolves depth from, so the two
58
- > decisions cannot disagree. Derive it with one script over the Story
59
- > branch (Story #5313) — never a hand-carried import block:
60
- >
61
38
  > ```bash
62
39
  > node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
63
40
  > ```
64
41
  >
65
- > It computes the change set once (`files`), derives the level and
66
- > classes, and resolves the ceremony per cluster (`mode`, `reason`,
67
- > `verdictOwner`): **`high` (the diff touches a sensitive path registered
68
- > in `audit-rules.json`) → `fresh`** (spawn the critic); **`low` (it
69
- > touches none) → `inline`** (the contract-identical inline fallback
70
- > below); **`null` / unknown (the diff could not be enumerated) → `fresh`
71
- > + full ceremony** (fail-safe). This chooses fresh-vs-inline **per
72
- > cluster only — it never changes the cluster count**.
73
- >
74
- > The routing signal is deliberately **not** a planner-authored risk
75
- > verdict: a level the plan asserted about itself was exactly the signal
76
- > that could *reduce* independent checking, and nothing verified it
77
- > against the diff.
78
- >
79
- > **Inline-critic path (low-level-routed OR nesting-absent harness).** The
80
- > verdict is authored **inline** whenever the risk router above resolves to
81
- > `inline` (a low-risk cluster), and also as
82
- > a **fallback** on any harness that cannot spawn the fresh critic.
83
- > Dispatching the critic as a nested `Agent` is the fresh-context shape and
84
- > works on any harness that carries `Agent` into sub-agents (Claude Code ≥
85
- > 2.1.202). This eval loop itself runs inside a Story delivery
86
- > sub-agent, so the nested
87
- > critic sits at nesting depth 2. If the host does **not** support nested
88
- > `Agent` dispatch at that depth — the tool is absent, or a spawn attempt
89
- > returns an unsupported-capability error — do **not** stall the Story
90
- > regardless of the risk verdict. Author the verdict **inline**: in a
91
- > deliberately scoped,
92
- > self-critical pass (re-read only the diff, the `acceptance[]` /
93
- > `verify[]` arrays, and the `verify[]` command output — treat the
94
- > implementation reasoning as untrusted and score against the criteria
95
- > afresh), write the same verdict file described below and hand it to the
96
- > same `acceptance-eval.js` gate. The fresh-context isolation is weaker in
97
- > the inline path, but the gate, the schema, the round cap, and the
98
- > proceed / redraft / block decision are identical — a Story is **never**
99
- > stranded on a nesting-absent harness. Note in the blocked/friction
100
- > comment (if you block) that the inline fallback was used.
42
+ > One call computes the change set once (`files`), derives the level and
43
+ > classes **for review depth**, and resolves the owner (`mode`, `reason`,
44
+ > `verdictOwner`): **`minimal` / `standard` → `inline`** (the default — you
45
+ > author the verdict yourself), **`strict` → `fresh`** (dispatch the
46
+ > maker-blind critic). The derived level no longer routes this decision;
47
+ > it escalates `review-depth.js` instead, which still resolves `deep` for
48
+ > any sensitive path.
49
+
50
+ **Never run both**, and never run a preliminary self-assessment before
51
+ dispatching a fresh critic — the redundant pre-pass buys no measurable
52
+ quality and roughly triples the acceptance-block cost. Step 3's gate is the
53
+ deterministic **scorer** of that one verdict, not a second pass over the
54
+ criteria.
101
55
 
102
- The critic:
103
- + Inspects the **change set handed to it in its spawn context** — the one
104
- list computed above — and the Story's inline `acceptance[]` / `verify[]`
105
- arrays. Pass the file list explicitly when you dispatch the critic; it
106
- does not re-enumerate the diff for itself, so a commit
107
- landing mid-ceremony cannot leave the critic scoring a different change
108
- than the one that routed it.
56
+ > **Inline owner (the default).** Author the verdict in a deliberately
57
+ > scoped, self-critical pass: re-read only the diff, the `acceptance[]` /
58
+ > `verify[]` arrays, and the `verify[]` command output — treat your own
59
+ > implementation reasoning as untrusted and score each criterion afresh
60
+ > from the evidence. Write one verdict file covering every item and hand it
61
+ > to the gate. This is also the **fallback** on any harness that cannot
62
+ > spawn the fresh critic, so a Story is never stranded: the gate, the
63
+ > schema, the round cap and the proceed / redraft / block decision are
64
+ > identical either way. Note in the friction comment (if you block) when
65
+ > the fallback was used in place of a `strict` critic.
66
+ >
67
+ > **Fresh critic (`strict` only).** Dispatch a sub-agent via the `Agent`
68
+ > tool — *not* a continuation of your implementing turn. When
69
+ > `delivery.routing.roleScopedAgents` is enabled (the **default**), use
70
+ > `subagent_type: acceptance-critic`: it boots on the role-scoped
71
+ > [`acceptance-critic`](../../agents/acceptance-critic.md) context (its own
72
+ > system prompt, no `CLAUDE.md` @-closure) carrying the maker-blind
73
+ > invariant and the verdict schema standalone. With the kill-switch off
74
+ > (`roleScopedAgents: false`), fall back to
75
+ > `subagent_type: general-purpose`. This loop already runs inside a Story
76
+ > delivery sub-agent, so the critic sits at nesting depth 2 — supported by
77
+ > any harness that carries `Agent` into sub-agents (Claude Code ≥ 2.1.202).
78
+
79
+ Whichever pass owns it, the verdict:
80
+ + Inspects the **change set it was handed** — the one `files` list above —
81
+ and the Story's inline `acceptance[]` / `verify[]` arrays. Pass the file
82
+ list explicitly when dispatching a fresh critic; it does not re-enumerate
83
+ the diff for itself, so a commit landing mid-ceremony cannot leave it
84
+ scoring a different change than the one that routed it.
109
85
  + **Runs the `verify[]` commands** and consumes their output as **required
110
- evidence** when scoring the relevant acceptance items. `verify[]` is not
111
- optional advisory pre-flight — a criterion cannot be scored `met` without
112
- the supporting `verify[]` evidence where a `verify[]` command is relevant
113
- to it.
86
+ evidence**. `verify[]` is not optional advisory pre-flight — a criterion
87
+ cannot be scored `met` without the supporting `verify[]` evidence where a
88
+ `verify[]` command is relevant to it.
114
89
  + **Reuses the credited full-suite run instead of re-paying for it.**
115
90
  Before spawning a `verify[]` entry, classify it with `resolveVerifyCredit`
116
91
  from
117
92
  [`verify-credit.js`](../../scripts/lib/orchestration/verify-credit.js): an
118
- entry that is itself a full-suite command (`npm test`, `pnpm run test`,
119
- a bare `node --test`) is consulted against the **same stamp close reads**
120
- and, when that stamp is fresh, recorded as `pass` with a `detail` naming
121
- the credit — **never respawned**. A stale or absent stamp reports
122
- `spawn: true` and the command runs for real, so the credit can never
123
- manufacture a pass. The gate warns on any such entry: the intended shape
124
- is scoped `verify[]` entries **plus** the one credited run
125
- ([`deliver-digest.md`](deliver-digest.md) § 5).
126
- + **Shares `lint` / `typecheck` evidence with close.** When a
127
- `verify[]` command is **byte-identical** to a close-validation gate — in
128
- practice only the cheap, command-identical `lint` and `typecheck` gates
129
- (`npm run lint` and the resolved `project.commands.typecheck`) — the
130
- critic MUST run it through `evidence-gate.js` so a passing run records an
131
- evidence entry in the **same keyspace** `close-validation/runner.js`
132
- consults. Run it in the **same Story worktree** the close validates (the
133
- HEAD-sha key enforces "unchanged HEAD") and pass the exact gate name:
134
-
135
- ```bash
136
- node <main-repo>/.agents/scripts/evidence-gate.js \
137
- --standalone --scope-id <storyId> --gate lint \
138
- --worktree <worktree> -- npm run lint
139
- ```
140
-
141
- Close's `shouldSkip` then short-circuits that gate when HEAD is
142
- unchanged; a redraft round (HEAD moves) correctly busts it. **Never**
143
- run the coverage / CRAP suite through `evidence-gate.js` to stamp it
144
- fresh — a false-fresh coverage record without `coverage-final.json`
145
- silently weakens the floor. Limit the evidence-share to `lint` and
146
- `typecheck`.
147
- + Emits a **cluster** verdict file under `temp/` conforming to
93
+ entry that is itself a full-suite command is consulted against the same
94
+ stamp close reads and, when that stamp is fresh, recorded as `pass`
95
+ without being respawned; a stale or absent stamp reports `spawn: true` and
96
+ the command runs for real. The credited run itself is stated once, in
97
+ [`deliver-digest.md`](deliver-digest.md) § 5.
98
+ + Emits **one** verdict file under `temp/` conforming to
148
99
  [`acceptance-eval-verdict.schema.json`](../../schemas/acceptance-eval-verdict.schema.json):
149
100
  one `{ index, criterion, verdict: met|partial|unmet, evidence,
150
- verifyEvidence[] }` record per acceptance item **in that cluster**, each
151
- `index` being the item's position in the Story's full `acceptance[]`
152
- array. A fresh critic **returns that path to you** rather than calling the
153
- gate itself.
154
- 2. **Dispatch the round's clusters in parallel, then merge into one verdict
155
- (fresh critics only).** When the verdict owner is the **inline self-eval**,
156
- author **one** verdict file covering every `acceptance[]` item and score it
157
- in one gate call — the cluster merge below does not apply (Story #5313).
158
- For fresh critics, the clusters of a round are independent, so dispatch
159
- **all** of the round's critics as N `Agent` calls **in a single assistant
160
- turn** — [`parallel-tooling.md`](parallel-tooling.md) **Rule 3** — never
161
- serially, and never one round per cluster.
162
-
163
- Then **merge** the cluster verdicts into **one** verdict file under `temp/`:
164
- concatenate every cluster's `criteria[]` records and order the merged array
165
- by `index`, so it holds exactly one record per `acceptance[]` item in
166
- **acceptance-array order**, under a single top-level `storyId`,
167
- `schemaVersion`, `round` and `commitSha`. The verdict schema deliberately
168
- carries **no `clusterId`** — the round's artifact is the merged verdict, and
169
- which critic scored which record is not part of the contract.
170
-
171
- > **Why one gate call and not N.** The round counter is **Story-scoped** —
172
- > derived by counting `acceptance-eval` signals in the Story's
173
- > `signals.ndjson` — and each cluster verdict has a distinct fingerprint, so
174
- > the replay guard never collapses them. A gate call per cluster would spend
175
- > one of the (default 2) rounds *per cluster*, so a Story with more than 8
176
- > acceptance criteria would exhaust its redraft budget on cluster arithmetic
177
- > alone; N concurrent calls would also race that same ledger. Cluster-scoped
178
- > round counting exists in
179
- > [`acceptance-eval-decision.js`](../../scripts/lib/orchestration/acceptance-eval-decision.js)
180
- > but requires an integer `epicId`, which v2 pins `null` — it is not a way
181
- > around the merge.
182
- 3. **Decide — exactly one gate call per round.** Run the gate against the
183
- **merged** verdict (the caller's Step 1a names the exact invocation — omit
184
- `--epic`). The gate **scores the single verdict the round produced** —
185
- schema validation, round cap, decision — and never re-scores the criteria
186
- itself:
101
+ verifyEvidence[] }` record per `acceptance[]` item, in acceptance-array
102
+ order, under a single top-level `storyId`, `schemaVersion`, `round` and
103
+ `commitSha`. A fresh critic **returns that path to you** rather than
104
+ calling the gate itself.
105
+ 2. **Decide — exactly one gate call per round.** Run the gate against that one
106
+ verdict file (the caller's Step 1a names the exact invocation — omit
107
+ `--epic`). The gate **scores the verdict the round produced** — schema
108
+ validation, round cap, decision — and never re-scores the criteria itself:
187
109
 
188
110
  ```bash
189
111
  node <main-repo>/.agents/scripts/acceptance-eval.js \
@@ -191,10 +113,16 @@ mid-delivery, and evaluates the actual work product.
191
113
  ```
192
114
 
193
115
  The gate reads the Story's `acceptance[]` count itself (Story #5313): a
194
- verdict whose `criteria[]` length differs — a single cluster's verdict
195
- handed over unmerged — is rejected **before scoring**, with an error naming
196
- the merge contract and consuming **no round**. `--expected-criteria` is
197
- still accepted but redundant; when passed it must agree with that count.
116
+ verdict whose `criteria[]` length differs — one covering only part of the
117
+ Story — is rejected **before scoring**, with an error naming the count and
118
+ consuming **no round**. `--expected-criteria` is still accepted but
119
+ redundant; when passed it must agree with that count.
120
+
121
+ > **Why one call per round and not several.** The round counter is
122
+ > **Story-scoped** — derived by counting `acceptance-eval` signals in the
123
+ > Story's `signals.ndjson` — so a second call in the same round spends one
124
+ > of the (default 2) rounds for nothing, and concurrent calls race that
125
+ > ledger.
198
126
 
199
127
  The gate validates the verdict against the schema, applies the round cap,
200
128
  emits the per-criterion `acceptance-eval` signal into the retro / feedback
@@ -209,5 +137,4 @@ mid-delivery, and evaluates the actual work product.
209
137
  (transition to `agent::blocked`) and post a `friction` comment naming the
210
138
  unmet criteria and their evidence. Never silently proceed to close.
211
139
 
212
- Write both the per-cluster verdicts and the merged verdict under `temp/` only —
213
- they are scratch artifacts.
140
+ Write the verdict under `temp/` only — it is a scratch artifact.
@@ -87,8 +87,10 @@ review yourself, honor the `depth` semantics above directly.
87
87
  1. Resolve `[TICKET_ID]` from `ticketId` (the Story for live `scope: story`).
88
88
  2. Resolve `[BASE_REF]` from `baseRef` and `[HEAD_REF]` from `headRef`.
89
89
  3. Fetch the Story ticket and resolve the planning context from its own
90
- body: folded `## Spec` / `## Slicing`, acceptance criteria, and the
91
- `story-plan-state` structured comment when present.
90
+ body: folded `## Spec` / `## Slicing`, and acceptance criteria. The
91
+ `story-plan-state` comment beside it carries the operator's plan summary
92
+ (story set, delivery order, deliver command), not planning context — it
93
+ holds no machine payload to read.
92
94
  4. Read that Spec fully to understand the intended scope, architectural
93
95
  decisions, and acceptance criteria. Do **not** look for a parent Epic.
94
96
 
@@ -25,18 +25,19 @@ rule produces it:
25
25
  shape — sub-agent isolation only matters against a *concurrent* sibling
26
26
  racing the same checkout, and a one-Story run has none.
27
27
  2. **Every other run is `subagent`.** A multi-Story run dispatches every Story
28
- as a sub-agent however trivial its shape. Shape still sets ceremony; the
29
- `route::lite` label is a human-visible hint, never the control signal.
28
+ as a sub-agent however trivial its shape — the run's size is the whole
29
+ premise, and nothing about a Story's own shape enters it.
30
30
 
31
- `inline` removes model-side fan-out only — no `story-worker` boot, no fresh
32
- acceptance-critic spawn. **`subagent` and `inline` run the same engine**: same
33
- gates, same PR to `main`, same terminal envelope, byte for byte.
31
+ `inline` removes the `story-worker` boot and nothing else. It does **not**
32
+ change the acceptance verdict owner — that is § 3's decision, and the profile
33
+ alone makes it. **`subagent` and `inline` run the same engine**: same gates,
34
+ same PR to `main`, same terminal envelope, byte for byte.
34
35
 
35
36
  ## 2. Engine invariants
36
37
 
37
38
  | Trait | Contract |
38
39
  | --- | --- |
39
- | Ticket type | `type::story` only; an `Epic: #N` footer means **stop and re-plan** |
40
+ | Ticket type | `type::story` only — `resolve-stories.js` validates it |
40
41
  | Branch | `story-<id>`, seeded from `project.baseBranch` (`main`) |
41
42
  | Merge target | `main` via PR (squash + required checks) — never a direct push |
42
43
  | Gates | Every close gate runs regardless of route; no route bypasses one |
@@ -57,38 +58,44 @@ import block:
57
58
  node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
58
59
  ```
59
60
 
60
- It prints one JSON object: `files` — the one change set every critic is
61
+ It prints one JSON object: `files` — the one change set the verdict owner is
61
62
  handed (`null` when the diff could not be enumerated) — plus `level` and
62
63
  `classes` from `review-depth.js`, and `mode`, `reason` and `verdictOwner`
63
64
  from `ceremony-routing.js`. Level rules: a sensitive path registered in
64
65
  `audit-rules.json` → `high`, none → `low`, an unenumerable diff → `null`.
65
- Ceremony rules: `minimal` → always inline, `strict` → always fresh,
66
- `standard` → `high`/`null` → fresh and `low` → inline. An `inline` dispatch
67
- mode overrides all of it to inline critics. Close's `review-depth.js` reads
68
- the same derived level, so the two cannot disagree. `--base <ref>` overrides
69
- `project.baseBranch`.
66
+ The level drives **review depth** only; close's `review-depth.js` reads the
67
+ same derived level, so the two cannot disagree. A sensitive footprint
68
+ therefore buys a **deep review**, not a fresh acceptance critic.
69
+
70
+ > **The ceremony rule, stated once.** The **profile alone** names the verdict
71
+ > owner (Story #5343, narrowed to that one input by #5366): `minimal` /
72
+ > `standard` → `inline`, `strict` → `fresh`. Nothing else moves it — not the
73
+ > derived change level, not the footprint's sensitivity, and **not the
74
+ > dispatch mode**: an `inline` Story under `strict` still spawns the fresh
75
+ > maker-blind critic, one nesting level shallower than a dispatched one. The
76
+ > only sanctioned inline authoring under `strict` is the harness fallback in
77
+ > [`acceptance-self-eval.md`](acceptance-self-eval.md) — a host that cannot
78
+ > spawn the critic at all, noted in the friction comment if you block.
79
+
80
+ `--base <ref>` overrides `project.baseBranch`.
70
81
 
71
82
  ## 4. Acceptance self-eval (Step 1a, required)
72
83
 
73
- **One verdict-owner per cluster** — the fresh critic *or* the inline
74
- self-eval, named by `verdictOwner`, never both and never a warm-up pass. Each
75
- scores its cluster's `acceptance[]` items against the change set above, with
76
- `verify[]` output as evidence. Bounded by `delivery.acceptanceEval.maxRounds`
84
+ **One verdict owner per Story** — named by `verdictOwner`: the inline
85
+ self-eval under `minimal` / `standard` (the default), a fresh maker-blind
86
+ critic under `strict`. Never both, and never a warm-up pass. The owner
87
+ authors **one** verdict file covering every `acceptance[]` item, scored
88
+ against the change set above with `verify[]` output as evidence, and it is
89
+ scored in **one** gate call. Bounded by `delivery.acceptanceEval.maxRounds`
77
90
  (default 2; `0` scores once with no redraft).
78
91
 
79
- **Inline owner:** author **one** verdict file covering every `acceptance[]`
80
- item and score it in **one** gate call — there is no cluster merge.
81
- **Fresh critics:** one round = N cluster critics → ONE merged verdict → ONE
82
- gate call. Merge every cluster's records into a single `criteria[]` in
83
- `acceptance[]` order and score that once; a gate call per cluster spends a
84
- round *per cluster* and races the round ledger.
85
-
86
92
  `node <main-repo>/.agents/scripts/acceptance-eval.js --story <storyId>
87
93
  --verdict <verdict-path>`
88
94
 
89
95
  The gate reads the Story's `acceptance[]` count itself and rejects a verdict
90
96
  whose `criteria[]` length differs **before** scoring, consuming no round;
91
- `--expected-criteria` is accepted but redundant.
97
+ `--expected-criteria` is accepted but redundant. A second gate call in the
98
+ same round spends a round for nothing and races the Story-scoped ledger.
92
99
 
93
100
  `proceed` → close. `redraft` → one more round inside the cap. `block` → **do
94
101
  not close**: post a `friction` comment and flip `agent::blocked`.