mandrel 2.57.0 → 2.59.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (58) hide show
  1. package/.agents/README.md +6 -3
  2. package/.agents/agents/story-worker.md +12 -11
  3. package/.agents/docs/SDLC.md +6 -7
  4. package/.agents/docs/quality-gates.md +1 -1
  5. package/.agents/instructions.md +2 -3
  6. package/.agents/runtime-deps.json +7 -2
  7. package/.agents/schemas/crap-baseline.schema.json +1 -1
  8. package/.agents/schemas/crap-report.schema.json +1 -1
  9. package/.agents/scripts/evidence-gate.js +17 -1
  10. package/.agents/scripts/install-matrix-assert.js +48 -3
  11. package/.agents/scripts/lib/audit-to-stories/seed-from-findings.js +51 -33
  12. package/.agents/scripts/lib/baselines/kinds/_crap-read.js +0 -8
  13. package/.agents/scripts/lib/baselines/kinds/crap.js +35 -18
  14. package/.agents/scripts/lib/crap-engine.js +2 -2
  15. package/.agents/scripts/lib/crap-utils.js +21 -5
  16. package/.agents/scripts/lib/escomplex-ast-compat.js +39 -17
  17. package/.agents/scripts/lib/escomplex-kernel.js +298 -0
  18. package/.agents/scripts/lib/maintainability-engine.js +3 -3
  19. package/.agents/scripts/lib/orchestration/code-review.js +7 -3
  20. package/.agents/scripts/lib/orchestration/pinned-identifier-lint.js +137 -0
  21. package/.agents/scripts/lib/orchestration/plan-context.js +41 -27
  22. package/.agents/scripts/lib/orchestration/plan-persist/acceptance-handle-repair.js +107 -0
  23. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +6 -1
  24. package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +14 -9
  25. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +45 -31
  26. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +8 -9
  27. package/.agents/scripts/lib/orchestration/plan-persist/supersede-ops.js +1 -1
  28. package/.agents/scripts/lib/orchestration/plan-persist/wave-collision-gate.js +107 -0
  29. package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +15 -5
  30. package/.agents/scripts/lib/orchestration/review-base-ref.js +138 -0
  31. package/.agents/scripts/lib/orchestration/single-story-close/phases/code-review.js +37 -5
  32. package/.agents/scripts/lib/orchestration/single-story-close/runner.js +6 -1
  33. package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +25 -209
  34. package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +8 -5
  35. package/.agents/scripts/lib/runtime-deps/dep-resolution.js +155 -0
  36. package/.agents/scripts/lib/runtime-deps/ensure-installed.js +44 -9
  37. package/.agents/scripts/lib/runtime-deps/parser-major.js +110 -0
  38. package/.agents/scripts/lib/runtime-deps/preflight.js +6 -25
  39. package/.agents/scripts/lib/runtime-deps/scan-imports.js +46 -1
  40. package/.agents/scripts/lib/skills/walk-skill-files.js +1 -1
  41. package/.agents/scripts/lib/story-body/story-body.js +36 -2
  42. package/.agents/scripts/lib/templates/decomposer-prompts.js +73 -21
  43. package/.agents/scripts/lib/test-run-credit.js +23 -12
  44. package/.agents/scripts/plan-persist.js +0 -11
  45. package/.agents/skills/skills.index.json +1 -11
  46. package/.agents/workflows/audit-to-stories.md +14 -11
  47. package/.agents/workflows/helpers/deliver-digest.md +22 -15
  48. package/.agents/workflows/helpers/deliver-story-reference.md +31 -11
  49. package/.agents/workflows/helpers/deliver-story.md +6 -5
  50. package/.agents/workflows/helpers/plan-reference.md +53 -13
  51. package/.agents/workflows/mandrel-plan.md +19 -14
  52. package/README.md +3 -3
  53. package/docs/CHANGELOG.md +21 -0
  54. package/lib/cli/registry.js +143 -27
  55. package/package.json +7 -2
  56. package/.agents/scripts/lib/orchestration/split-policy-validator.js +0 -188
  57. package/.agents/scripts/lib/templates/spec-author-prompts.js +0 -76
  58. package/.agents/skills/core/scope-triage/SKILL.md +0 -48
@@ -24,14 +24,19 @@ import { BODY_FORMAT_LINTS } from '../story-body/body-format-lints.js';
24
24
  * no verify-tier suffix — every one of those either scored a shape the
25
25
  * authoring model already judges or prescribed a proxy that became the
26
26
  * goal.
27
- * - **The N>1 rules** ({@link renderStorySplitRules}) — the schedule and
28
- * partition rules that only mean anything once a draft has siblings:
29
- * every Story must earn its slot in the wave schedule, and every
30
- * acceptance criterion belongs to exactly one Story.
27
+ * - **The N>1 rules** ({@link renderStorySplitRules}) — the schedule rules
28
+ * that only mean anything once a draft has siblings: every Story must
29
+ * earn its slot in the wave schedule, and no same-wave pair may collide
30
+ * on a declared path (Story #5332 replaced the acceptance partition with
31
+ * the dispatcher's own collision predicate, armed as a refusal).
32
+ * - **The tickets-mode rules** ({@link ticketsModePromptField}, Story
33
+ * #5323) — what to re-derive rather than carry when the seed is an
34
+ * existing ticket whose body is already in Story shape.
31
35
  *
32
- * The envelope carries the core as `systemPrompts.story` and the split rules
33
- * as `systemPrompts.storySplitRules`; a planner reads the second only when
34
- * the default-single split policy clears.
36
+ * The envelope carries the core as `systemPrompts.story`, the split rules as
37
+ * `systemPrompts.storySplitRules` and the tickets rules as
38
+ * `systemPrompts.storyTicketsRules`; a planner reads the second only when the
39
+ * default-single split policy clears, and the third only in tickets mode.
35
40
  */
36
41
 
37
42
  /**
@@ -131,9 +136,9 @@ The **persisted** \`body\` renders these markdown sections (in order) — you au
131
136
 
132
137
  - **goal** (in body string): One sentence stating WHY this Story exists.
133
138
  - **spec** (optional, in body string as \`## Spec\`): The technical approach at the altitude the SPEC PROSE CONTRACT below fixes — contract and invariants, never implementation narration. Write as much as the work needs and no more; persist keeps Specs inline at any length and never writes them under \`docs/\`.
134
- - **slicing** (optional): Ordered intra-session checkpoints for one Story, one line each. Not a fan-out table and not a duplicate of Acceptance.
135
- - **changes** (in body string): Each entry is an object \`{ path, assumption }\` where \`assumption\` is one of \`creates | refactors-existing | deletes\`. Acceptable path shapes include explicit files (\`src/components/Foo.tsx\`), glob patterns (\`tests/e2e/*.spec.ts\`, \`**/*.astro\`), and module identifiers that resolve to files. Use \`refactors-existing\` for in-place edits to a file already on \`main\`; \`creates\` for net-new files; \`deletes\` for removals. Persist probes every path against the base branch and repairs a plain-string bullet or a trailing parenthetical into the object form for you; a \`creates\` on an existing path or a \`refactors-existing\` on an absent one is a dry-run warning, and only a \`deletes\` naming an absent path is refused.
136
- - **acceptance** (top-level array on the ticket object): Each item is an **outcome a PR reviewer can confirm from the diff and the verify output** — what is true of the codebase once the Story lands, stated at the altitude of the capability (a command that now exits 0 against a named input, a behavior a named test now asserts, a config that now fails validation on a retired key, a document that now records a decision). Aim for **three to six** items: fewer than three usually means the outcome is under-specified; more than six usually means acceptance is re-listing the footprint or the mechanical checks that belong in \`verify[]\`. Push grep-shaped probes, file-exists checks and exit-code tests down into \`verify[]\`; never pin an internal helper name or a private file path into an acceptance item the advisory \`changes[]\` is free to reshape. UNACCEPTABLE: "verify by reading the diff", "looks good", "matches the spec".
139
+ - **slicing** (optional): Ordered intra-session checkpoints for one Story, one line each. A checkpoint is a **stage of the work** — a commit boundary the deliverer passes through inside one session, stated as the step it performs. An acceptance item is a **state of the codebase** a PR reviewer confirms once the Story has landed. The same Story therefore carries both: the checkpoints say in what order it is built, \`acceptance[]\` says what must then be true. Never a fan-out table, never a second acceptance list, and never sibling tickets — a broad sweep with many stages is still one Story, sliced here.
140
+ - **changes** (in body string): Each entry is an object \`{ path, assumption }\` where \`assumption\` is one of \`creates | refactors-existing | deletes\`. **Name the files the deliverer authors, and omit generated artifacts** — quality baselines, generated test indexes, migration journals, lockfiles and the like are regenerated by the work itself, the refresh is a close-gate concern, and declaring one needlessly reserves a footprint that serializes sibling Stories at dispatch. Acceptable path shapes include explicit files (\`src/components/Foo.tsx\`), glob patterns (\`tests/e2e/*.spec.ts\`, \`**/*.astro\`), and module identifiers that resolve to files. Use \`refactors-existing\` for in-place edits to a file already on \`main\`; \`creates\` for net-new files; \`deletes\` for removals. Persist probes every path against the base branch and repairs a plain-string bullet or a trailing parenthetical into the object form for you; a \`creates\` on an existing path or a \`refactors-existing\` on an absent one is a dry-run warning, and only a \`deletes\` naming an absent path is refused.
141
+ - **acceptance** (top-level array on the ticket object): Each item is an **outcome a PR reviewer can confirm from the diff and the verify output** — what is true of the codebase once the Story lands, stated at the altitude of the capability (a command that now exits 0 against a named input, a behavior a named test now asserts, a config that now fails validation on a retired key, a document that now records a decision). State as many outcomes as the capability has and no more — the list has no target, floor or ceiling, and a long one is never a reason to split the Story. Push grep-shaped probes, file-exists checks and exit-code tests down into \`verify[]\`; never pin an internal helper name or a private file path into an acceptance item the advisory \`changes[]\` is free to reshape. UNACCEPTABLE: "verify by reading the diff", "looks good", "matches the spec".
137
142
  - **verify** (top-level array on the ticket object): The **mechanical checks** — exact commands or test paths the deliverer runs and the acceptance critic consumes as evidence: \`node --test tests/x.test.js\`, \`npm run lint\`, \`npm run validate\`, a scoped grep. Every acceptance item should be confirmable from at least one verify entry's output plus the diff. Stories with zero verify entries fail validation.
138
143
  - **Bodies record decisions, never questions to the operator.** Never persist an open question ("Flag if…", "TBD", "confirm with the operator") into a Story body — the executing sub-agent is non-interactive and cannot answer it, and the dry-run warns on every one it finds. Triage each unknown by who can resolve it: an AFK-shaped unknown (a fact in docs, a third-party API surface, observable repo behavior) MUST be resolved by your own research before authoring — never restated as an assumption; only a HITL-shaped unknown (a genuine product or architecture call the operator owns) may be restated as a declarative Key Assumption the agent can act on, stating the default chosen (a decision-made-by-default).
139
144
  - **non_goals** (OPTIONAL, in body string as the \`## Non-Goals\` section): A short list of capabilities or changes this Story explicitly does NOT deliver — an advisory negative-scope bound that fences the executing agent away from adjacent work. It is **advisory and NON-GATING**: the validator does not require, count, or reject on it, and an absent or empty section renders nothing. Use the EXACT single-word hyphenated heading spelling \`## Non-Goals\` (a space-separated heading like \`## Out of Scope\` is NOT recognized by the parser and will be dropped). Reach for it when a Story's negative boundary is non-obvious from its \`acceptance[]\` alone; omit it otherwise.
@@ -167,13 +172,13 @@ ${advisoryCaveat}
167
172
 
168
173
  **Decompose at deliverable granularity, not module/task level.** ${granularityDefinition}
169
174
 
170
- The only sizing question is **cohesion**: *is this one coherent change with one reason to exist?* There is no ceiling on a Story's footprint, Spec length or acceptance count — a broad contract cutover is one Story when every changed site changes for the same reason. Frontier models one-shot capability-sized work in a single pass; do not fragment a coherent capability into dependent slices to stay "small", and do not pad a Story with adjacent work to look "complete".
175
+ The only sizing question is **cohesion**: *is this one coherent change with one reason to exist?* There is no target, floor or ceiling on a Story's footprint, Spec length or acceptance count, and a long acceptance list is a description of a broad capability, never a reason to split. A broad contract cutover is one Story when every changed site changes for the same reason. Frontier models one-shot capability-sized work in a single pass; do not fragment a coherent capability into dependent slices to stay "small", and do not pad a Story with adjacent work to look "complete".
171
176
 
172
177
  ${envelopeFloor}
173
178
 
174
- - **One Story = one coherent change with one reason to exist.** If you cannot state that reason in a sentence, the Story is probably two Stories — or two Stories that should be one.
179
+ - **One Story = one coherent change with one reason to exist.**
180
+ - **A remediation sweep over one subsystem is one Story.** A batch of findings in the same subsystem shares one reason to exist — the subsystem is wrong — so it arrives as one Story whose \`## Slicing\` checkpoints carry the stages, not as one Story per finding.
175
181
  - ${singleConsumerRule}
176
- - **Split independent, parallelizable work** into sibling Stories — but only when the pieces genuinely have separate reasons to exist.
177
182
 
178
183
  #### UI / TESTID INVARIANCE (per CLAUDE.md safety rule):
179
184
 
@@ -197,7 +202,7 @@ IMPORTANT DEPENDENCY RULE: Story-to-Story dependencies are expressed via \`depen
197
202
  /**
198
203
  * The rules that only apply once a draft has more than one Story: the
199
204
  * delivery-schedule simulation that makes each Story earn its slot, and the
200
- * acceptance partition persist enforces at N>1.
205
+ * same-wave collision refusal persist enforces at N>1.
201
206
  *
202
207
  * @returns {string}
203
208
  */
@@ -209,16 +214,44 @@ You are splitting past the default-single policy, so simulate the delivery sched
209
214
  1. **Build the wave schedule.** A Story runs only after every \`depends_on\` completes, and two Stories that name the same file in \`changes[]\` cannot run in the same wave (the scheduler serializes file-overlapping Stories even when no \`depends_on\` edge links them).
210
215
  2. **Every Story must earn its slot** by at least one of:
211
216
  - **(a) parallelism** — it actually runs concurrently with a sibling in the schedule you just built ("logically independent" does not count; *schedule*-independent does);
212
- - **(b) risk isolation** — it isolates a consumer-facing behavior change or high-risk cutover into its own reviewable, revertable unit;
213
- - **(c) cohesion break** — merged into its neighbor it would no longer be one coherent change with one reason to exist.
214
- 3. **A dependent link with none of those justifications merges into its consumer.** This generalizes the single-consumer merge rule from pairs to chains: N Stories that deliver no faster than one Story pay N delivery sessions (branch, PR, review, CI) for nothing.
217
+ - **(b) cohesion break** — merged into its neighbor it would no longer be one coherent change with one reason to exist.
218
+ 3. **A dependent link with neither of those justifications merges into its consumer.** This generalizes the single-consumer merge rule from pairs to chains: N Stories that deliver no faster than one Story pay N delivery sessions (branch, PR, review, CI) for nothing.
215
219
  4. **When one file appears in the \`changes[]\` of most of your Stories, the slicing axis cuts across a shared seam** — merge the Stories that co-edit it, or re-slice along the seam so each Story owns its files.
216
220
 
217
- #### ACCEPTANCE PARTITION (persist-enforced at N>1):
221
+ #### THE COLLISION REFUSAL (persist-enforced at N>1):
218
222
 
219
- - Every acceptance criterion of the plan belongs to **exactly one** Story — no criterion is shared, and none is dropped. Persist refuses a draft whose criteria overlap or leave a plan-level criterion unclaimed.
220
- - Each Story carries its **own** \`## Spec\`; a shared \`techspec.md\` cannot be folded into N>1 Stories.
221
- - Express ordering with \`depends_on\` (a sibling slug, or \`#<id>\` for an open Story from an earlier plan). A Story whose \`verify[]\` runs against a file a sibling creates MUST \`depends_on\` that sibling, so the file exists when verification runs.`;
223
+ Persist runs the **dispatcher's own** collision predicate pairwise over your draft, before it creates a single issue, and **refuses** the plan when any two same-wave Stories collide — both declaring a path in \`changes[]\`, or one declaring a glob that covers the other's path. Such a pair cannot be co-dispatched, so the split buys no parallelism and costs a delivery session. Two remedies, both yours to choose at authoring time:
224
+
225
+ - **Merge the pair** into the one Story they already are, with \`## Slicing\` checkpoints for the stages; or
226
+ - **Order them** with \`depends_on\` so they sit in different waves, when they genuinely have separate reasons to exist.
227
+
228
+ Each Story carries its **own** \`## Spec\`; a shared \`techspec.md\` cannot be folded into N>1 Stories. Express ordering with \`depends_on\` (a sibling slug, or \`#<id>\` for an open Story from an earlier plan). A Story whose \`verify[]\` runs against a file a sibling creates MUST \`depends_on\` that sibling, so the file exists when verification runs.`;
229
+ }
230
+
231
+ /**
232
+ * The rules that only apply when the seed is an existing ticket (Story
233
+ * #5323).
234
+ *
235
+ * A `--tickets` seed arrives already in Story shape — rendered `AC-<n>:`
236
+ * checkboxes, a `## Verify` list, a `## Changes` footprint — and an author
237
+ * reading it as a template carries that shape forward instead of re-deriving
238
+ * it. The observed failure (swarm-os #2707 / #2708, planned from #2542 under
239
+ * mandrel 2.57.0) was a Story whose acceptance list was the source's, handles
240
+ * and all, and whose verify entries carried a tier suffix retired two
241
+ * releases earlier. The source ticket is **evidence**, not a draft.
242
+ *
243
+ * @returns {string}
244
+ */
245
+ function renderStoryTicketsRules() {
246
+ return `#### TICKETS-MODE DRAFT — the source ticket is evidence, not a template:
247
+
248
+ You are planning from one or more existing tickets. Read them for **what the work is** — the problem, the constraints, the commands that verify it — and re-derive everything else. Specifically:
249
+
250
+ 1. **Re-derive \`acceptance[]\` from the goal.** Do not copy the source's \`## Acceptance\` list, and never carry its \`AC-<n>:\` handles — the body renderer numbers the checkboxes itself, so a copied handle renders doubled. A source ticket carrying fifteen criteria is telling you its acceptance was over-specified, not that yours must be: state the outcomes a PR reviewer can confirm, and let the mechanical checks fall to \`verify[]\`.
251
+ 2. **A mechanical check is a \`verify[]\` command, not an acceptance item.** "Baselines refreshed", "lint exits 0", "the generated index is regenerated", "the quality gate passes" are commands the deliverer runs and the critic reads as evidence. Carrying them as acceptance items inflates the binding contract with work every close already gates.
252
+ 3. **Read the source's \`verify[]\` for the commands it names, not for its shape.** Take the test paths and scripts; drop any trailing tier suffix (\`(unit)\`, \`(contract)\`, \`(e2e)\`, \`(validate)\`) and any \`manual:<reason>\` escape — a verify entry is a bare command.
253
+ 4. **Re-derive the footprint against the tree as it is now.** The source ticket's \`## Changes\` predicted a repository that has since moved; probe the paths you cite and omit the generated artifacts it listed.
254
+ 5. **Do not carry the source's prose wholesale.** Its current-state narration and per-file walkthroughs are exactly what the SPEC PROSE CONTRACT above forbids. Restate the contract and the invariants; the deliverer reads the code for the rest.`;
222
255
  }
223
256
 
224
257
  /**
@@ -233,3 +266,22 @@ export function renderStoryAuthorPrompt({ storyCount = 1 } = {}) {
233
266
  const core = renderStoryAuthorCore();
234
267
  return storyCount > 1 ? `${core}\n\n${renderStorySplitRules()}` : core;
235
268
  }
269
+
270
+ /**
271
+ * The mode-conditional slice of `systemPrompts`.
272
+ *
273
+ * `storyTicketsRules` only means anything when the seed is an existing
274
+ * ticket, and an envelope carrying it in every mode teaches the author to
275
+ * look for a source ticket a `--seed` run does not have. Returning a
276
+ * spreadable object rather than a nullable string keeps the decision here,
277
+ * beside the prompt it selects, instead of as a branch in the envelope
278
+ * builder.
279
+ *
280
+ * @param {string|undefined} mode The plan-context mode.
281
+ * @returns {{ storyTicketsRules?: string }}
282
+ */
283
+ export function ticketsModePromptField(mode) {
284
+ return mode === 'tickets'
285
+ ? { storyTicketsRules: renderStoryTicketsRules() }
286
+ : {};
287
+ }
@@ -1,15 +1,23 @@
1
1
  /**
2
- * lib/test-run-credit.js — let a green bare `npm test` earn the credit close
3
- * reads (Story #5313).
2
+ * lib/test-run-credit.js — let a green `npm test` **that routes through
3
+ * mandrel's own runner** earn the credit close reads (Story #5313, scoped by
4
+ * Story #5324).
4
5
  *
5
- * Until this module the only suite run that deposited credit was the one
6
- * shaped exactly like the close gate — `coverage-capture.js --cwd <worktree>`
7
- * or `evidence-gate.js --standalone … -- npm test` — and the digest, the
8
- * worker boot context and the reference all carried prose explaining which
9
- * invocation to type. A worker that ran the project's own test runner paid
10
- * for the suite and then close paid for it again. The runner is the natural
11
- * depositor: it knows the tree it ran against, whether the run was green,
12
- * and whether it ran the whole suite.
6
+ * This is a **bonus, not the contract.** The deposit every project can rely
7
+ * on is `evidence-gate.js --standalone --scope-id <id> --gate test --worktree
8
+ * <workCwd> -- npm test`: it spawns whatever `npm test` resolves to and
9
+ * stamps what it just ran, so it is honest on any runner. What this module
10
+ * adds is that a repo whose `test` script *is* `run-tests.js` need not type
11
+ * that wrapper — the runner already knows the tree it ran against, whether
12
+ * the run was green, and whether it ran the whole suite, so it deposits on
13
+ * the way out.
14
+ *
15
+ * The reach is therefore exactly one call site: `run-tests.js`. A consumer
16
+ * whose `npm test` is `vitest run` or `jest` never loads this module, so it
17
+ * deposits nothing **and prints nothing** — silence is not a signal, and no
18
+ * delivery surface may tell an agent to confirm credit by reading for the
19
+ * line below. `mandrel doctor`'s `test-credit-path` check reports which of
20
+ * the two shapes a project is and names the wrapper as the remedy.
13
21
  *
14
22
  * On a green **full-tier** run inside a `story-<id>` checkout the runner
15
23
  * records the `test` gate's evidence in the same keyspace
@@ -143,8 +151,11 @@ export function depositTestRunCredit({
143
151
  }
144
152
 
145
153
  /**
146
- * Deposit and say so on stderr — the runner's one-line hook. The line is
147
- * the only surface a worker sees, so it names the outcome by reason.
154
+ * Deposit and say so on stderr — the runner's one-line hook, printed only
155
+ * when this runner is the one running. It names the outcome by reason, so a
156
+ * green run that deposited nothing (wrong branch, partial tier) says so
157
+ * rather than passing silently; a project on another runner prints no line
158
+ * at all, which is why absence of this line is never evidence either way.
148
159
  *
149
160
  * @param {Parameters<typeof depositTestRunCredit>[0] & { log?: (line: string) => void }} args
150
161
  * @returns {ReturnType<typeof depositTestRunCredit>}
@@ -29,7 +29,6 @@
29
29
  * --plan-context <file> Optional explicit path to the `plan-context.js`
30
30
  * envelope. Its `sourceTickets[]` is what makes
31
31
  * `--tickets` superseding work without a flag
32
- * --plan-acceptance <file> Optional JSON string[] for partition coverage
33
32
  * --source-tickets <ids> Explicit OVERRIDE of the envelope-derived source
34
33
  * ids, for hand-driven runs. Each id must be
35
34
  * claimed by exactly one Story's `supersedes[]`;
@@ -106,7 +105,6 @@ const CLI_OPTIONS = {
106
105
  'tech-spec': { type: 'string' },
107
106
  'plan-dir': { type: 'string' },
108
107
  'plan-context': { type: 'string' },
109
- 'plan-acceptance': { type: 'string' },
110
108
  'source-tickets': { type: 'string' },
111
109
  'close-superseded': { type: 'boolean', default: true },
112
110
  'no-close-superseded': { type: 'boolean', default: false },
@@ -121,7 +119,6 @@ const CLI_OPTIONS = {
121
119
  const USAGE =
122
120
  'Usage: plan-persist.js --stories <file> ' +
123
121
  '[--tech-spec <file>] [--plan-dir <dir>] [--plan-context <file>] ' +
124
- '[--plan-acceptance <file>] ' +
125
122
  '[--source-tickets <ids>] [--no-close-superseded] ' +
126
123
  '[--dry-run] [--chain-on-clean] [--force-review] ' +
127
124
  '[--epic-title <text> --epic-goal <text> | --epic <id>]';
@@ -159,9 +156,6 @@ export function resolveInputPaths(values) {
159
156
  techSpecPath: values['tech-spec']
160
157
  ? path.resolve(values['tech-spec'])
161
158
  : null,
162
- planAcceptancePath: values['plan-acceptance']
163
- ? path.resolve(values['plan-acceptance'])
164
- : null,
165
159
  planDir,
166
160
  planContextPath: resolvePlanContextPath(values['plan-context'], planDir),
167
161
  };
@@ -172,9 +166,6 @@ async function loadArtifacts(paths) {
172
166
  const techSpecContent = paths.techSpecPath
173
167
  ? await readOptional(paths.techSpecPath, { required: true })
174
168
  : null;
175
- const planAcceptance = paths.planAcceptancePath
176
- ? await readJsonFile(paths.planAcceptancePath, 'plan-acceptance')
177
- : null;
178
169
  const planContextEnvelope = await loadPlanContextEnvelope(
179
170
  paths.planContextPath,
180
171
  );
@@ -182,7 +173,6 @@ async function loadArtifacts(paths) {
182
173
  return {
183
174
  stories,
184
175
  techSpecContent,
185
- planAcceptance,
186
176
  planContextEnvelope,
187
177
  };
188
178
  }
@@ -501,7 +491,6 @@ runAsCli(import.meta.url, main, {
501
491
  '--plan-context <file>',
502
492
  'The plan-context envelope this draft was authored against.',
503
493
  ],
504
- ['--plan-acceptance <file>', 'Acceptance artifact to attach.'],
505
494
  ['--source-tickets <ids>', 'Ticket ids this plan supersedes.'],
506
495
  ['--dry-run', 'Validate and report; create nothing.'],
507
496
  ['--chain-on-clean', 'Persist immediately when the dry run is clean.'],
@@ -1,5 +1,5 @@
1
1
  {
2
- "generatedAt": "2026-09-06T12:59:23.069Z",
2
+ "generatedAt": "2026-09-14T11:57:26.498Z",
3
3
  "generator": "generate-skills-index.js@1",
4
4
  "skills": [
5
5
  {
@@ -52,16 +52,6 @@
52
52
  "allowedTools": null,
53
53
  "vendor": null
54
54
  },
55
- {
56
- "name": "scope-triage",
57
- "tier": "core",
58
- "category": "core",
59
- "path": ".agents/skills/core/scope-triage/SKILL.md",
60
- "description": "Optional split-advisory for `/mandrel-plan`. Under v2 there is no epic|story routing verdict — `/mandrel-plan` always authors Stories. Use this skill only when judging whether a draft should stay one Story or legitimately split (near-zero overlap or an architectural seam).",
61
- "policyCapsuleBullets": 5,
62
- "allowedTools": null,
63
- "vendor": null
64
- },
65
55
  {
66
56
  "name": "security-and-hardening",
67
57
  "tier": "core",
@@ -152,15 +152,17 @@ node .agents/scripts/audit-to-stories.js --emit-plan-seed \
152
152
 
153
153
  The seed renders the canonical one-pager sections — Problem Statement,
154
154
  Recommended Direction, Key Assumptions (with links to every source
155
- report), MVP Scope (the M proposed Stories), Key Files (so `/mandrel-plan`'s
156
- authoring step has concrete anchors), Grouping, Not Doing.
157
-
158
- **Grouping is the container-Epic directive.** Above 2 proposed Stories the
159
- seed instructs `/mandrel-plan` to group them under one Epic — a sweep is the
160
- clearest case for a container, since every Story shares a provenance and an
161
- operator usually delivers them together. It is a directive in the text, not an
162
- automatic write: Phase 4 above is where an operator declines it. Below the
163
- threshold the section says so and asks for nothing.
155
+ report), MVP Scope (**the findings, flat**), Key Files (so `/mandrel-plan`'s
156
+ authoring step has concrete anchors), Not Doing.
157
+
158
+ **The seed states findings, not a partition.** MVP Scope used
159
+ to render one numbered bullet per group beneath a `## Grouping` container
160
+ directive — a plan the seed had already decided, at the grouping grain, before
161
+ `/mandrel-plan` read a word of it. N now reaches the planner **undecided**: it
162
+ applies its own cohesion judgment, and container grouping is `/mandrel-plan`'s
163
+ Gate #3 call at persist, where N is known. The grouping still drives the
164
+ **standalone-Stories** path (Phase 5b), which needs one issue payload per
165
+ group.
164
166
 
165
167
  Chain into the existing planning entrypoint:
166
168
 
@@ -172,9 +174,10 @@ Chain into the existing planning entrypoint:
172
174
  then runs its author → persist path, as documented in its workflow.
173
175
 
174
176
  **Dedup provenance is carried mechanically — do not hand-copy it.** The seed's
175
- MVP Scope bullets carry each group's `audit-fingerprints` and
177
+ MVP Scope section carries each group's `audit-fingerprints` and
176
178
  `audit-semantic-keys` footers as HTML comments (invisible in the rendered
177
- one-pager). `plan-persist` harvests them out of the seed on the
179
+ one-pager, and per-group even though the visible list is flat — they are the
180
+ identity the next sweep matches on). `plan-persist` harvests them out of the seed on the
178
181
  `plan-context.json` envelope and appends them to **every** Story body it
179
182
  persists, via `carryProvenanceFooters`
180
183
  ([`lib/findings/route-finding.js`](../scripts/lib/findings/route-finding.js)).
@@ -94,26 +94,33 @@ whose `criteria[]` length differs **before** scoring, consuming no round;
94
94
  not close**: post a `friction` comment and flip `agent::blocked`.
95
95
  Per-round mechanics: [`acceptance-self-eval.md`](acceptance-self-eval.md).
96
96
 
97
- ## 5. The one full-suite run
97
+ ## 5. The one credited suite run
98
98
 
99
- After the self-eval loop's last fix commit, run the project test runner
100
- **once** in the worktree:
99
+ After the self-eval loop's last fix commit, run the suite **once** in the
100
+ worktree through the depositor — it spawns the project's own `npm test`,
101
+ whatever that resolves to, and stamps the result, so any runner earns the
102
+ credit:
101
103
 
102
104
  ```bash
103
- npm test # in <workCwd>
105
+ node <main-repo>/.agents/scripts/evidence-gate.js --standalone \
106
+ --scope-id <storyId> --gate test --worktree <workCwd> -- npm test
104
107
  ```
105
108
 
106
- A green full run on `story-<id>` deposits the `test` evidence close reads,
107
- keyed on the tree, so close reports the gate as **credited** at unchanged
108
- HEAD — a later commit voids it. The CRAP gate still captures coverage itself
109
- when it needs an artifact. If the suite outruns the host's sync Bash
110
- ceiling, dispatch it in the **background** — its completion re-invokes you;
111
- never spawn a task to poll or `sleep`-loop against it
112
- ([`parallel-tooling.md`](parallel-tooling.md) Rule 2). Read the **output**,
113
- not the exit code: the runner prints whether it deposited credit, and a run
114
- off the Story branch or of a partial tier deposits nothing and says so.
115
- Redraft rounds run the scoped projects for the roots you changed plus
116
- `verify[]`, not the whole suite; only this run needs credit.
109
+ Green deposits the `test` evidence close reads, keyed on the tree, so close
110
+ reports the gate as **credited** at unchanged HEAD — a later commit voids it.
111
+ Read its **output**, not the exit code: `✓ test passed` is the signal. The
112
+ CRAP gate still captures coverage itself when it needs an artifact.
113
+
114
+ A bare `npm test` earns the same credit **only** where the project's test
115
+ script routes through mandrel's own runner, which prints the outcome. On any
116
+ other runner it deposits nothing and prints nothing, so silence is never
117
+ evidence of credit; `mandrel doctor`'s `test-credit-path` check names which
118
+ shape this project is. If the suite outruns the host's sync Bash ceiling,
119
+ dispatch it in the **background** — its completion re-invokes you; never spawn
120
+ a task to poll or `sleep`-loop against it
121
+ ([`parallel-tooling.md`](parallel-tooling.md) Rule 2). Redraft rounds run the
122
+ scoped projects for the roots you changed plus `verify[]`, not the whole
123
+ suite; only this run needs credit.
117
124
 
118
125
  `verify[]` is scoped entries **plus** this one run: an entry that is itself a
119
126
  full-suite command is reported credited against the same record, never
@@ -237,17 +237,37 @@ the failure class that actually bounces deliveries: close-validation
237
237
  discovers them only after the whole close pipeline has run, at several times
238
238
  the cost of one full-suite run in the worktree.
239
239
 
240
- **Run it once, last, so close can credit it.** The run belongs **after** the
241
- self-eval loop's last fix commit; redraft rounds run scoped tests. A green
242
- `npm test` in the worktree on `story-<id>` deposits the `test` evidence
243
- record close reads (Story #5313 — `lib/test-run-credit.js`), keyed on HEAD
244
- and the tree fingerprint and hashed on the exact command close spawns, so
245
- close reports the gate as credited at unchanged HEAD instead of re-running
246
- the suite. The credit expires the moment it stops describing the tree: any
247
- later commit invalidates it and close re-runs the suite for real, so this
248
- never trades away the gate. The CRAP gate still runs `coverage-capture.js`
249
- itself when it needs a fresh artifact — the capture stamp is a claim about
250
- `coverage/coverage-final.json`, which a bare `npm test` does not produce.
240
+ **Run it once, last, through the depositor.** The run belongs **after** the
241
+ self-eval loop's last fix commit; redraft rounds run scoped tests. Run it in
242
+ the worktree on `story-<id>` as
243
+
244
+ ```bash
245
+ node <main-repo>/.agents/scripts/evidence-gate.js \
246
+ --standalone --scope-id <storyId> --gate test \
247
+ --worktree <workCwd> -- npm test
248
+ ```
249
+
250
+ The wrapper spawns the project's own `npm test` — whatever that resolves to —
251
+ and records the pass into the Story evidence keyspace, so the credit is
252
+ runner-agnostic by construction: it stamps only what it just ran. The record
253
+ is keyed on HEAD and the tree fingerprint and hashed on the exact command
254
+ close spawns, so close reports the gate as credited at unchanged HEAD instead
255
+ of re-running the suite. The credit expires the moment it stops describing the
256
+ tree: any later commit invalidates it and close re-runs the suite for real, so
257
+ this never trades away the gate. The CRAP gate still runs
258
+ `coverage-capture.js` itself when it needs a fresh artifact — the capture
259
+ stamp is a claim about `coverage/coverage-final.json`, which `npm test` alone
260
+ does not produce.
261
+
262
+ **A bare `npm test` is a bonus, not the contract.** It deposits the same
263
+ record only where the project's `test` script routes through mandrel's own
264
+ runner (`run-tests.js` → `lib/test-run-credit.js`, Story #5313), which prints
265
+ the outcome. A project whose `npm test` is `vitest run`, `jest` or any other
266
+ runner never reaches that code, so it prints nothing and deposits nothing —
267
+ silence is not a signal, and nothing here asks you to confirm the credit by
268
+ reading for a line that cannot appear. `mandrel doctor`'s `test-credit-path`
269
+ check reports which shape a project is and names the command above as its
270
+ remedy.
251
271
 
252
272
  **`verify[]` reuses the same credit.** A `verify[]` entry that is itself a
253
273
  full-suite command is reported **credited** against that record rather than
@@ -93,12 +93,13 @@ are reference § Step 2. Hard gates always run in Step 3 — the derived level
93
93
  never disables them; do **not** pre-run the chain here — Step 2.5's credited
94
94
  suite run is the sole exception.
95
95
 
96
- ### Step 2.5 — The one full-suite run, the push, then hand off
96
+ ### Step 2.5 — The one credited suite run, the push, then hand off
97
97
 
98
- After the self-eval loop's last fix commit, run `npm test` **once** in the
99
- worktree (**digest § 5**): a green full run deposits the `test` credit close
100
- reads, keyed on the tree, so only a *later* commit invalidates it. Red →
101
- fix, commit, re-run.
98
+ After the self-eval loop's last fix commit, run the suite **once** in the
99
+ worktree through the depositor — `evidence-gate.js … --gate test -- npm test`,
100
+ spelled out in **digest § 5**. It runs whatever `npm test` resolves to and
101
+ stamps that, so the `test` credit is earned on any runner and only a *later*
102
+ commit invalidates it. Red → fix, commit, re-run.
102
103
 
103
104
  Push `story-<storyId>` to `origin`, confirming the remote ref moved. Then
104
105
  (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
@@ -47,10 +47,20 @@ The spine's two escape hatches from N=1 are narrow on purpose:
47
47
  sitting unverifiable behind the other.
48
48
 
49
49
  Everything else is one Story with `## Slicing` checkpoints. When N>1 does
50
- apply, **every acceptance criterion belongs to exactly one Story** —
51
- `assertAcceptancePartition` refuses a split whose criteria repeat across
52
- siblings, because a verbatim-shared criterion is the signature of coupled work
53
- cut in half rather than genuinely separable work.
50
+ apply, the split has to survive the **same-wave collision refusal**: persist
51
+ runs the dispatcher's own `detectCollision` pairwise over the draft, ahead of
52
+ the first `createIssue`, and refuses any pair of same-wave siblings that both
53
+ declare a path (or one of which declares a covering glob) — naming the pair,
54
+ the paths and the two remedies (merge them, or order them with `depends_on`).
55
+ Such a pair cannot be co-dispatched, so the split buys no parallelism and
56
+ costs a delivery session per Story. It replaced the acceptance partition,
57
+ which refused only byte-identical acceptance text across siblings — a shape
58
+ model output does not produce — so it never fired on the fragmentation it was
59
+ meant to catch. **N=1 can never trip the refusal.**
60
+
61
+ A draft of more than one Story also stops at **Gate #2** for operator
62
+ approval, `--force-review` or not: a split always earns eyes. `--yes`
63
+ auto-proceeds, as at every other gate.
54
64
 
55
65
  ## Unknown triage — AFK vs HITL
56
66
 
@@ -154,6 +164,11 @@ and ceremony is derived from the landed diff at close.
154
164
  - **`verify[]` entries are commands.** There is no tier suffix and no
155
165
  `manual:<reason>` escape (Story #5312): write the exact command or test
156
166
  path the deliverer runs and the acceptance critic reads as evidence.
167
+ - **`changes[]` names what the deliverer authors.** Generated artifacts —
168
+ quality baselines, generated test indexes, migration journals, lockfiles —
169
+ are omitted: the work regenerates them, the refresh is a close-gate concern,
170
+ and a declared shared artifact path reserves a footprint that needlessly
171
+ serializes sibling Stories at dispatch.
157
172
  - **`changes[]` arrive pre-resolved to creates-vs-refactors.** Every path
158
173
  the seed predicted is probed against the repo: an existing path is
159
174
  emitted with `assumption: "refactors-existing"`, a missing one with
@@ -175,8 +190,14 @@ kept — passes the persist ticket validators with no round-trip.
175
190
  Each `stories.json` entry: `slug` (`^[a-z0-9][a-z0-9-]*$`), `type: "story"`,
176
191
  `title`, `body` (`goal`, optional `spec`, `changes[{path, assumption}]` —
177
192
  `creates|refactors-existing|deletes`, `non_goals`, `reason_to_exist`),
178
- top-level `acceptance[]`, `verify[]` (`… (unit|contract|e2e|validate)`), and
179
- `depends_on[]` (a sibling slug, or `#<id>` for an existing open Story).
193
+ top-level `acceptance[]`, `verify[]` (each a **bare command** — there is no
194
+ tier suffix), and `depends_on[]` (a sibling slug, or `#<id>` for an existing
195
+ open Story).
196
+
197
+ Author `acceptance[]` **without** the `AC-<n>:` handle: the body renderer
198
+ numbers each checkbox from its array position, so a carried handle renders
199
+ doubled. Persist normalises one off rather than refusing, and names the strip
200
+ on the dry-run's repair list.
180
201
 
181
202
  Nothing in that shape inventories the repo for the author. `changes[]` arrives
182
203
  pre-resolved against the working tree, and Phase 8's
@@ -238,8 +259,9 @@ The conflict passes run **twice**: once over the raw `stories.json` payload
238
259
  The second pass is not belt-and-braces. The canonical authoring shape carries
239
260
  `acceptance[]` / `verify[]` at the ticket's top level and assembly folds them
240
261
  into the body, so the passes that scan `body.acceptance` / `body.verify`
241
- (`implicit-cross-story-dep`, `missing-bdd-scaffold`) saw two empty arrays on
242
- the real payload and emitted nothing. Both passes complete before the first
262
+ saw two empty arrays on the real payload and emitted nothing; the two
263
+ substring-match advisories that depended on it are retired, leaving
264
+ `shared-editor` as the one conflict kind. Both passes complete before the first
243
265
  `createIssue`, so a refusal still costs no writes.
244
266
 
245
267
  `shared-editor` findings are rendered into the posted `plan-summary` comment,
@@ -256,6 +278,21 @@ and a path reference matched by substring can read as a dependency a prose
256
278
  mention never meant. A finding names the Stories and the fix (a `depends_on`
257
279
  edge, or folding the shared edit into one Story) for the operator to weigh.
258
280
 
281
+ ## Tickets mode — the source ticket is evidence
282
+
283
+ A `--tickets` envelope carries a third author prompt beside
284
+ `systemPrompts.story` and `systemPrompts.storySplitRules`:
285
+ **`systemPrompts.storyTicketsRules`**. It exists because a
286
+ source ticket arrives already in Story shape — rendered `AC-<n>:` checkboxes,
287
+ a `## Verify` list, a `## Changes` footprint — and an author reading it as a
288
+ template carries that shape forward instead of re-deriving it. The addendum
289
+ binds the author to re-derive `acceptance[]` from the goal, to express
290
+ mechanical checks (a refreshed baseline, a lint exiting 0, a regenerated
291
+ index) as `verify[]` commands rather than acceptance items, and to take the
292
+ source's verify entries for the commands they name rather than their shape.
293
+ Read it whenever the mode is `tickets`; the other three modes do not carry
294
+ the field.
295
+
259
296
  ## Tickets mode — authoring `supersedes[]`
260
297
 
261
298
  In `--tickets` mode each Story carries a top-level `supersedes` array claiming
@@ -282,7 +319,7 @@ template-only prose.
282
319
  ### Supersede-map partition
283
320
 
284
321
  `plan-persist` refuses a partial supersede map **before** it creates any
285
- Story (mirroring `assertAcceptancePartition`): every id passed to
322
+ Story, the same fail-closed shape as the collision refusal: every id passed to
286
323
  `--tickets` must be claimed by **exactly one** Story, and no Story may
287
324
  claim an id that was not a source ticket. With N>1 the mapping is not
288
325
  total by default — an authored map is the only thing that can say
@@ -347,10 +384,13 @@ at base.
347
384
  exists at base or a `refactors-existing` on one that does not (including a
348
385
  path the base branch deleted or renamed, named with the removing commit), a
349
386
  goal or acceptance path absent at base, a `verify[]` command naming an absent
350
- test file, and an `open-question` in a body (`Flag if…`, `TBD`, a trailing
351
- `?`). The list also names every `changes[]` **repair** the run applied — a
352
- plain-string bullet or a trailing parenthetical rewritten into
353
- `{ path, assumption }` by probing base. The same list rides the result
387
+ test file, an `open-question` in a body (`Flag if…`, `TBD`, a trailing `?`),
388
+ and a `pinned-identifier` in an acceptance item — a backticked bare symbol
389
+ that is not a path, a label, a kebab token, a flag or a command, which the
390
+ advisory `changes[]` is free to reshape out from under the criterion. The
391
+ list also names every **repair** the run applied — a plain-string bullet or a
392
+ trailing parenthetical rewritten into `{ path, assumption }` by probing base,
393
+ and an `AC-<n>:` handle normalised off an acceptance item. The same list rides the result
354
394
  envelope as `warnings[]` and `repairs[]`, so a `--chain-on-clean` run loses
355
395
  nothing.
356
396