mandrel 2.57.0 → 2.59.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/README.md +6 -3
- package/.agents/agents/story-worker.md +12 -11
- package/.agents/docs/SDLC.md +6 -7
- package/.agents/docs/quality-gates.md +1 -1
- package/.agents/instructions.md +2 -3
- package/.agents/runtime-deps.json +7 -2
- package/.agents/schemas/crap-baseline.schema.json +1 -1
- package/.agents/schemas/crap-report.schema.json +1 -1
- package/.agents/scripts/evidence-gate.js +17 -1
- package/.agents/scripts/install-matrix-assert.js +48 -3
- package/.agents/scripts/lib/audit-to-stories/seed-from-findings.js +51 -33
- package/.agents/scripts/lib/baselines/kinds/_crap-read.js +0 -8
- package/.agents/scripts/lib/baselines/kinds/crap.js +35 -18
- package/.agents/scripts/lib/crap-engine.js +2 -2
- package/.agents/scripts/lib/crap-utils.js +21 -5
- package/.agents/scripts/lib/escomplex-ast-compat.js +39 -17
- package/.agents/scripts/lib/escomplex-kernel.js +298 -0
- package/.agents/scripts/lib/maintainability-engine.js +3 -3
- package/.agents/scripts/lib/orchestration/code-review.js +7 -3
- package/.agents/scripts/lib/orchestration/pinned-identifier-lint.js +137 -0
- package/.agents/scripts/lib/orchestration/plan-context.js +41 -27
- package/.agents/scripts/lib/orchestration/plan-persist/acceptance-handle-repair.js +107 -0
- package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +6 -1
- package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +14 -9
- package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +45 -31
- package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +8 -9
- package/.agents/scripts/lib/orchestration/plan-persist/supersede-ops.js +1 -1
- package/.agents/scripts/lib/orchestration/plan-persist/wave-collision-gate.js +107 -0
- package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +15 -5
- package/.agents/scripts/lib/orchestration/review-base-ref.js +138 -0
- package/.agents/scripts/lib/orchestration/single-story-close/phases/code-review.js +37 -5
- package/.agents/scripts/lib/orchestration/single-story-close/runner.js +6 -1
- package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +25 -209
- package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +8 -5
- package/.agents/scripts/lib/runtime-deps/dep-resolution.js +155 -0
- package/.agents/scripts/lib/runtime-deps/ensure-installed.js +44 -9
- package/.agents/scripts/lib/runtime-deps/parser-major.js +110 -0
- package/.agents/scripts/lib/runtime-deps/preflight.js +6 -25
- package/.agents/scripts/lib/runtime-deps/scan-imports.js +46 -1
- package/.agents/scripts/lib/skills/walk-skill-files.js +1 -1
- package/.agents/scripts/lib/story-body/story-body.js +36 -2
- package/.agents/scripts/lib/templates/decomposer-prompts.js +73 -21
- package/.agents/scripts/lib/test-run-credit.js +23 -12
- package/.agents/scripts/plan-persist.js +0 -11
- package/.agents/skills/skills.index.json +1 -11
- package/.agents/workflows/audit-to-stories.md +14 -11
- package/.agents/workflows/helpers/deliver-digest.md +22 -15
- package/.agents/workflows/helpers/deliver-story-reference.md +31 -11
- package/.agents/workflows/helpers/deliver-story.md +6 -5
- package/.agents/workflows/helpers/plan-reference.md +53 -13
- package/.agents/workflows/mandrel-plan.md +19 -14
- package/README.md +3 -3
- package/docs/CHANGELOG.md +21 -0
- package/lib/cli/registry.js +143 -27
- package/package.json +7 -2
- package/.agents/scripts/lib/orchestration/split-policy-validator.js +0 -188
- package/.agents/scripts/lib/templates/spec-author-prompts.js +0 -76
- package/.agents/skills/core/scope-triage/SKILL.md +0 -48
|
@@ -24,14 +24,19 @@ import { BODY_FORMAT_LINTS } from '../story-body/body-format-lints.js';
|
|
|
24
24
|
* no verify-tier suffix — every one of those either scored a shape the
|
|
25
25
|
* authoring model already judges or prescribed a proxy that became the
|
|
26
26
|
* goal.
|
|
27
|
-
* - **The N>1 rules** ({@link renderStorySplitRules}) — the schedule
|
|
28
|
-
*
|
|
29
|
-
*
|
|
30
|
-
*
|
|
27
|
+
* - **The N>1 rules** ({@link renderStorySplitRules}) — the schedule rules
|
|
28
|
+
* that only mean anything once a draft has siblings: every Story must
|
|
29
|
+
* earn its slot in the wave schedule, and no same-wave pair may collide
|
|
30
|
+
* on a declared path (Story #5332 replaced the acceptance partition with
|
|
31
|
+
* the dispatcher's own collision predicate, armed as a refusal).
|
|
32
|
+
* - **The tickets-mode rules** ({@link ticketsModePromptField}, Story
|
|
33
|
+
* #5323) — what to re-derive rather than carry when the seed is an
|
|
34
|
+
* existing ticket whose body is already in Story shape.
|
|
31
35
|
*
|
|
32
|
-
* The envelope carries the core as `systemPrompts.story
|
|
33
|
-
*
|
|
34
|
-
* the
|
|
36
|
+
* The envelope carries the core as `systemPrompts.story`, the split rules as
|
|
37
|
+
* `systemPrompts.storySplitRules` and the tickets rules as
|
|
38
|
+
* `systemPrompts.storyTicketsRules`; a planner reads the second only when the
|
|
39
|
+
* default-single split policy clears, and the third only in tickets mode.
|
|
35
40
|
*/
|
|
36
41
|
|
|
37
42
|
/**
|
|
@@ -131,9 +136,9 @@ The **persisted** \`body\` renders these markdown sections (in order) — you au
|
|
|
131
136
|
|
|
132
137
|
- **goal** (in body string): One sentence stating WHY this Story exists.
|
|
133
138
|
- **spec** (optional, in body string as \`## Spec\`): The technical approach at the altitude the SPEC PROSE CONTRACT below fixes — contract and invariants, never implementation narration. Write as much as the work needs and no more; persist keeps Specs inline at any length and never writes them under \`docs/\`.
|
|
134
|
-
- **slicing** (optional): Ordered intra-session checkpoints for one Story, one line each.
|
|
135
|
-
- **changes** (in body string): Each entry is an object \`{ path, assumption }\` where \`assumption\` is one of \`creates | refactors-existing | deletes\`. Acceptable path shapes include explicit files (\`src/components/Foo.tsx\`), glob patterns (\`tests/e2e/*.spec.ts\`, \`**/*.astro\`), and module identifiers that resolve to files. Use \`refactors-existing\` for in-place edits to a file already on \`main\`; \`creates\` for net-new files; \`deletes\` for removals. Persist probes every path against the base branch and repairs a plain-string bullet or a trailing parenthetical into the object form for you; a \`creates\` on an existing path or a \`refactors-existing\` on an absent one is a dry-run warning, and only a \`deletes\` naming an absent path is refused.
|
|
136
|
-
- **acceptance** (top-level array on the ticket object): Each item is an **outcome a PR reviewer can confirm from the diff and the verify output** — what is true of the codebase once the Story lands, stated at the altitude of the capability (a command that now exits 0 against a named input, a behavior a named test now asserts, a config that now fails validation on a retired key, a document that now records a decision).
|
|
139
|
+
- **slicing** (optional): Ordered intra-session checkpoints for one Story, one line each. A checkpoint is a **stage of the work** — a commit boundary the deliverer passes through inside one session, stated as the step it performs. An acceptance item is a **state of the codebase** a PR reviewer confirms once the Story has landed. The same Story therefore carries both: the checkpoints say in what order it is built, \`acceptance[]\` says what must then be true. Never a fan-out table, never a second acceptance list, and never sibling tickets — a broad sweep with many stages is still one Story, sliced here.
|
|
140
|
+
- **changes** (in body string): Each entry is an object \`{ path, assumption }\` where \`assumption\` is one of \`creates | refactors-existing | deletes\`. **Name the files the deliverer authors, and omit generated artifacts** — quality baselines, generated test indexes, migration journals, lockfiles and the like are regenerated by the work itself, the refresh is a close-gate concern, and declaring one needlessly reserves a footprint that serializes sibling Stories at dispatch. Acceptable path shapes include explicit files (\`src/components/Foo.tsx\`), glob patterns (\`tests/e2e/*.spec.ts\`, \`**/*.astro\`), and module identifiers that resolve to files. Use \`refactors-existing\` for in-place edits to a file already on \`main\`; \`creates\` for net-new files; \`deletes\` for removals. Persist probes every path against the base branch and repairs a plain-string bullet or a trailing parenthetical into the object form for you; a \`creates\` on an existing path or a \`refactors-existing\` on an absent one is a dry-run warning, and only a \`deletes\` naming an absent path is refused.
|
|
141
|
+
- **acceptance** (top-level array on the ticket object): Each item is an **outcome a PR reviewer can confirm from the diff and the verify output** — what is true of the codebase once the Story lands, stated at the altitude of the capability (a command that now exits 0 against a named input, a behavior a named test now asserts, a config that now fails validation on a retired key, a document that now records a decision). State as many outcomes as the capability has and no more — the list has no target, floor or ceiling, and a long one is never a reason to split the Story. Push grep-shaped probes, file-exists checks and exit-code tests down into \`verify[]\`; never pin an internal helper name or a private file path into an acceptance item the advisory \`changes[]\` is free to reshape. UNACCEPTABLE: "verify by reading the diff", "looks good", "matches the spec".
|
|
137
142
|
- **verify** (top-level array on the ticket object): The **mechanical checks** — exact commands or test paths the deliverer runs and the acceptance critic consumes as evidence: \`node --test tests/x.test.js\`, \`npm run lint\`, \`npm run validate\`, a scoped grep. Every acceptance item should be confirmable from at least one verify entry's output plus the diff. Stories with zero verify entries fail validation.
|
|
138
143
|
- **Bodies record decisions, never questions to the operator.** Never persist an open question ("Flag if…", "TBD", "confirm with the operator") into a Story body — the executing sub-agent is non-interactive and cannot answer it, and the dry-run warns on every one it finds. Triage each unknown by who can resolve it: an AFK-shaped unknown (a fact in docs, a third-party API surface, observable repo behavior) MUST be resolved by your own research before authoring — never restated as an assumption; only a HITL-shaped unknown (a genuine product or architecture call the operator owns) may be restated as a declarative Key Assumption the agent can act on, stating the default chosen (a decision-made-by-default).
|
|
139
144
|
- **non_goals** (OPTIONAL, in body string as the \`## Non-Goals\` section): A short list of capabilities or changes this Story explicitly does NOT deliver — an advisory negative-scope bound that fences the executing agent away from adjacent work. It is **advisory and NON-GATING**: the validator does not require, count, or reject on it, and an absent or empty section renders nothing. Use the EXACT single-word hyphenated heading spelling \`## Non-Goals\` (a space-separated heading like \`## Out of Scope\` is NOT recognized by the parser and will be dropped). Reach for it when a Story's negative boundary is non-obvious from its \`acceptance[]\` alone; omit it otherwise.
|
|
@@ -167,13 +172,13 @@ ${advisoryCaveat}
|
|
|
167
172
|
|
|
168
173
|
**Decompose at deliverable granularity, not module/task level.** ${granularityDefinition}
|
|
169
174
|
|
|
170
|
-
The only sizing question is **cohesion**: *is this one coherent change with one reason to exist?* There is no ceiling on a Story's footprint, Spec length or acceptance count
|
|
175
|
+
The only sizing question is **cohesion**: *is this one coherent change with one reason to exist?* There is no target, floor or ceiling on a Story's footprint, Spec length or acceptance count, and a long acceptance list is a description of a broad capability, never a reason to split. A broad contract cutover is one Story when every changed site changes for the same reason. Frontier models one-shot capability-sized work in a single pass; do not fragment a coherent capability into dependent slices to stay "small", and do not pad a Story with adjacent work to look "complete".
|
|
171
176
|
|
|
172
177
|
${envelopeFloor}
|
|
173
178
|
|
|
174
|
-
- **One Story = one coherent change with one reason to exist.**
|
|
179
|
+
- **One Story = one coherent change with one reason to exist.**
|
|
180
|
+
- **A remediation sweep over one subsystem is one Story.** A batch of findings in the same subsystem shares one reason to exist — the subsystem is wrong — so it arrives as one Story whose \`## Slicing\` checkpoints carry the stages, not as one Story per finding.
|
|
175
181
|
- ${singleConsumerRule}
|
|
176
|
-
- **Split independent, parallelizable work** into sibling Stories — but only when the pieces genuinely have separate reasons to exist.
|
|
177
182
|
|
|
178
183
|
#### UI / TESTID INVARIANCE (per CLAUDE.md safety rule):
|
|
179
184
|
|
|
@@ -197,7 +202,7 @@ IMPORTANT DEPENDENCY RULE: Story-to-Story dependencies are expressed via \`depen
|
|
|
197
202
|
/**
|
|
198
203
|
* The rules that only apply once a draft has more than one Story: the
|
|
199
204
|
* delivery-schedule simulation that makes each Story earn its slot, and the
|
|
200
|
-
*
|
|
205
|
+
* same-wave collision refusal persist enforces at N>1.
|
|
201
206
|
*
|
|
202
207
|
* @returns {string}
|
|
203
208
|
*/
|
|
@@ -209,16 +214,44 @@ You are splitting past the default-single policy, so simulate the delivery sched
|
|
|
209
214
|
1. **Build the wave schedule.** A Story runs only after every \`depends_on\` completes, and two Stories that name the same file in \`changes[]\` cannot run in the same wave (the scheduler serializes file-overlapping Stories even when no \`depends_on\` edge links them).
|
|
210
215
|
2. **Every Story must earn its slot** by at least one of:
|
|
211
216
|
- **(a) parallelism** — it actually runs concurrently with a sibling in the schedule you just built ("logically independent" does not count; *schedule*-independent does);
|
|
212
|
-
- **(b)
|
|
213
|
-
|
|
214
|
-
3. **A dependent link with none of those justifications merges into its consumer.** This generalizes the single-consumer merge rule from pairs to chains: N Stories that deliver no faster than one Story pay N delivery sessions (branch, PR, review, CI) for nothing.
|
|
217
|
+
- **(b) cohesion break** — merged into its neighbor it would no longer be one coherent change with one reason to exist.
|
|
218
|
+
3. **A dependent link with neither of those justifications merges into its consumer.** This generalizes the single-consumer merge rule from pairs to chains: N Stories that deliver no faster than one Story pay N delivery sessions (branch, PR, review, CI) for nothing.
|
|
215
219
|
4. **When one file appears in the \`changes[]\` of most of your Stories, the slicing axis cuts across a shared seam** — merge the Stories that co-edit it, or re-slice along the seam so each Story owns its files.
|
|
216
220
|
|
|
217
|
-
####
|
|
221
|
+
#### THE COLLISION REFUSAL (persist-enforced at N>1):
|
|
218
222
|
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
-
|
|
223
|
+
Persist runs the **dispatcher's own** collision predicate pairwise over your draft, before it creates a single issue, and **refuses** the plan when any two same-wave Stories collide — both declaring a path in \`changes[]\`, or one declaring a glob that covers the other's path. Such a pair cannot be co-dispatched, so the split buys no parallelism and costs a delivery session. Two remedies, both yours to choose at authoring time:
|
|
224
|
+
|
|
225
|
+
- **Merge the pair** into the one Story they already are, with \`## Slicing\` checkpoints for the stages; or
|
|
226
|
+
- **Order them** with \`depends_on\` so they sit in different waves, when they genuinely have separate reasons to exist.
|
|
227
|
+
|
|
228
|
+
Each Story carries its **own** \`## Spec\`; a shared \`techspec.md\` cannot be folded into N>1 Stories. Express ordering with \`depends_on\` (a sibling slug, or \`#<id>\` for an open Story from an earlier plan). A Story whose \`verify[]\` runs against a file a sibling creates MUST \`depends_on\` that sibling, so the file exists when verification runs.`;
|
|
229
|
+
}
|
|
230
|
+
|
|
231
|
+
/**
|
|
232
|
+
* The rules that only apply when the seed is an existing ticket (Story
|
|
233
|
+
* #5323).
|
|
234
|
+
*
|
|
235
|
+
* A `--tickets` seed arrives already in Story shape — rendered `AC-<n>:`
|
|
236
|
+
* checkboxes, a `## Verify` list, a `## Changes` footprint — and an author
|
|
237
|
+
* reading it as a template carries that shape forward instead of re-deriving
|
|
238
|
+
* it. The observed failure (swarm-os #2707 / #2708, planned from #2542 under
|
|
239
|
+
* mandrel 2.57.0) was a Story whose acceptance list was the source's, handles
|
|
240
|
+
* and all, and whose verify entries carried a tier suffix retired two
|
|
241
|
+
* releases earlier. The source ticket is **evidence**, not a draft.
|
|
242
|
+
*
|
|
243
|
+
* @returns {string}
|
|
244
|
+
*/
|
|
245
|
+
function renderStoryTicketsRules() {
|
|
246
|
+
return `#### TICKETS-MODE DRAFT — the source ticket is evidence, not a template:
|
|
247
|
+
|
|
248
|
+
You are planning from one or more existing tickets. Read them for **what the work is** — the problem, the constraints, the commands that verify it — and re-derive everything else. Specifically:
|
|
249
|
+
|
|
250
|
+
1. **Re-derive \`acceptance[]\` from the goal.** Do not copy the source's \`## Acceptance\` list, and never carry its \`AC-<n>:\` handles — the body renderer numbers the checkboxes itself, so a copied handle renders doubled. A source ticket carrying fifteen criteria is telling you its acceptance was over-specified, not that yours must be: state the outcomes a PR reviewer can confirm, and let the mechanical checks fall to \`verify[]\`.
|
|
251
|
+
2. **A mechanical check is a \`verify[]\` command, not an acceptance item.** "Baselines refreshed", "lint exits 0", "the generated index is regenerated", "the quality gate passes" are commands the deliverer runs and the critic reads as evidence. Carrying them as acceptance items inflates the binding contract with work every close already gates.
|
|
252
|
+
3. **Read the source's \`verify[]\` for the commands it names, not for its shape.** Take the test paths and scripts; drop any trailing tier suffix (\`(unit)\`, \`(contract)\`, \`(e2e)\`, \`(validate)\`) and any \`manual:<reason>\` escape — a verify entry is a bare command.
|
|
253
|
+
4. **Re-derive the footprint against the tree as it is now.** The source ticket's \`## Changes\` predicted a repository that has since moved; probe the paths you cite and omit the generated artifacts it listed.
|
|
254
|
+
5. **Do not carry the source's prose wholesale.** Its current-state narration and per-file walkthroughs are exactly what the SPEC PROSE CONTRACT above forbids. Restate the contract and the invariants; the deliverer reads the code for the rest.`;
|
|
222
255
|
}
|
|
223
256
|
|
|
224
257
|
/**
|
|
@@ -233,3 +266,22 @@ export function renderStoryAuthorPrompt({ storyCount = 1 } = {}) {
|
|
|
233
266
|
const core = renderStoryAuthorCore();
|
|
234
267
|
return storyCount > 1 ? `${core}\n\n${renderStorySplitRules()}` : core;
|
|
235
268
|
}
|
|
269
|
+
|
|
270
|
+
/**
|
|
271
|
+
* The mode-conditional slice of `systemPrompts`.
|
|
272
|
+
*
|
|
273
|
+
* `storyTicketsRules` only means anything when the seed is an existing
|
|
274
|
+
* ticket, and an envelope carrying it in every mode teaches the author to
|
|
275
|
+
* look for a source ticket a `--seed` run does not have. Returning a
|
|
276
|
+
* spreadable object rather than a nullable string keeps the decision here,
|
|
277
|
+
* beside the prompt it selects, instead of as a branch in the envelope
|
|
278
|
+
* builder.
|
|
279
|
+
*
|
|
280
|
+
* @param {string|undefined} mode The plan-context mode.
|
|
281
|
+
* @returns {{ storyTicketsRules?: string }}
|
|
282
|
+
*/
|
|
283
|
+
export function ticketsModePromptField(mode) {
|
|
284
|
+
return mode === 'tickets'
|
|
285
|
+
? { storyTicketsRules: renderStoryTicketsRules() }
|
|
286
|
+
: {};
|
|
287
|
+
}
|
|
@@ -1,15 +1,23 @@
|
|
|
1
1
|
/**
|
|
2
|
-
* lib/test-run-credit.js — let a green
|
|
3
|
-
* reads (Story #5313
|
|
2
|
+
* lib/test-run-credit.js — let a green `npm test` **that routes through
|
|
3
|
+
* mandrel's own runner** earn the credit close reads (Story #5313, scoped by
|
|
4
|
+
* Story #5324).
|
|
4
5
|
*
|
|
5
|
-
*
|
|
6
|
-
*
|
|
7
|
-
*
|
|
8
|
-
*
|
|
9
|
-
*
|
|
10
|
-
*
|
|
11
|
-
*
|
|
12
|
-
*
|
|
6
|
+
* This is a **bonus, not the contract.** The deposit every project can rely
|
|
7
|
+
* on is `evidence-gate.js --standalone --scope-id <id> --gate test --worktree
|
|
8
|
+
* <workCwd> -- npm test`: it spawns whatever `npm test` resolves to and
|
|
9
|
+
* stamps what it just ran, so it is honest on any runner. What this module
|
|
10
|
+
* adds is that a repo whose `test` script *is* `run-tests.js` need not type
|
|
11
|
+
* that wrapper — the runner already knows the tree it ran against, whether
|
|
12
|
+
* the run was green, and whether it ran the whole suite, so it deposits on
|
|
13
|
+
* the way out.
|
|
14
|
+
*
|
|
15
|
+
* The reach is therefore exactly one call site: `run-tests.js`. A consumer
|
|
16
|
+
* whose `npm test` is `vitest run` or `jest` never loads this module, so it
|
|
17
|
+
* deposits nothing **and prints nothing** — silence is not a signal, and no
|
|
18
|
+
* delivery surface may tell an agent to confirm credit by reading for the
|
|
19
|
+
* line below. `mandrel doctor`'s `test-credit-path` check reports which of
|
|
20
|
+
* the two shapes a project is and names the wrapper as the remedy.
|
|
13
21
|
*
|
|
14
22
|
* On a green **full-tier** run inside a `story-<id>` checkout the runner
|
|
15
23
|
* records the `test` gate's evidence in the same keyspace
|
|
@@ -143,8 +151,11 @@ export function depositTestRunCredit({
|
|
|
143
151
|
}
|
|
144
152
|
|
|
145
153
|
/**
|
|
146
|
-
* Deposit and say so on stderr — the runner's one-line hook
|
|
147
|
-
*
|
|
154
|
+
* Deposit and say so on stderr — the runner's one-line hook, printed only
|
|
155
|
+
* when this runner is the one running. It names the outcome by reason, so a
|
|
156
|
+
* green run that deposited nothing (wrong branch, partial tier) says so
|
|
157
|
+
* rather than passing silently; a project on another runner prints no line
|
|
158
|
+
* at all, which is why absence of this line is never evidence either way.
|
|
148
159
|
*
|
|
149
160
|
* @param {Parameters<typeof depositTestRunCredit>[0] & { log?: (line: string) => void }} args
|
|
150
161
|
* @returns {ReturnType<typeof depositTestRunCredit>}
|
|
@@ -29,7 +29,6 @@
|
|
|
29
29
|
* --plan-context <file> Optional explicit path to the `plan-context.js`
|
|
30
30
|
* envelope. Its `sourceTickets[]` is what makes
|
|
31
31
|
* `--tickets` superseding work without a flag
|
|
32
|
-
* --plan-acceptance <file> Optional JSON string[] for partition coverage
|
|
33
32
|
* --source-tickets <ids> Explicit OVERRIDE of the envelope-derived source
|
|
34
33
|
* ids, for hand-driven runs. Each id must be
|
|
35
34
|
* claimed by exactly one Story's `supersedes[]`;
|
|
@@ -106,7 +105,6 @@ const CLI_OPTIONS = {
|
|
|
106
105
|
'tech-spec': { type: 'string' },
|
|
107
106
|
'plan-dir': { type: 'string' },
|
|
108
107
|
'plan-context': { type: 'string' },
|
|
109
|
-
'plan-acceptance': { type: 'string' },
|
|
110
108
|
'source-tickets': { type: 'string' },
|
|
111
109
|
'close-superseded': { type: 'boolean', default: true },
|
|
112
110
|
'no-close-superseded': { type: 'boolean', default: false },
|
|
@@ -121,7 +119,6 @@ const CLI_OPTIONS = {
|
|
|
121
119
|
const USAGE =
|
|
122
120
|
'Usage: plan-persist.js --stories <file> ' +
|
|
123
121
|
'[--tech-spec <file>] [--plan-dir <dir>] [--plan-context <file>] ' +
|
|
124
|
-
'[--plan-acceptance <file>] ' +
|
|
125
122
|
'[--source-tickets <ids>] [--no-close-superseded] ' +
|
|
126
123
|
'[--dry-run] [--chain-on-clean] [--force-review] ' +
|
|
127
124
|
'[--epic-title <text> --epic-goal <text> | --epic <id>]';
|
|
@@ -159,9 +156,6 @@ export function resolveInputPaths(values) {
|
|
|
159
156
|
techSpecPath: values['tech-spec']
|
|
160
157
|
? path.resolve(values['tech-spec'])
|
|
161
158
|
: null,
|
|
162
|
-
planAcceptancePath: values['plan-acceptance']
|
|
163
|
-
? path.resolve(values['plan-acceptance'])
|
|
164
|
-
: null,
|
|
165
159
|
planDir,
|
|
166
160
|
planContextPath: resolvePlanContextPath(values['plan-context'], planDir),
|
|
167
161
|
};
|
|
@@ -172,9 +166,6 @@ async function loadArtifacts(paths) {
|
|
|
172
166
|
const techSpecContent = paths.techSpecPath
|
|
173
167
|
? await readOptional(paths.techSpecPath, { required: true })
|
|
174
168
|
: null;
|
|
175
|
-
const planAcceptance = paths.planAcceptancePath
|
|
176
|
-
? await readJsonFile(paths.planAcceptancePath, 'plan-acceptance')
|
|
177
|
-
: null;
|
|
178
169
|
const planContextEnvelope = await loadPlanContextEnvelope(
|
|
179
170
|
paths.planContextPath,
|
|
180
171
|
);
|
|
@@ -182,7 +173,6 @@ async function loadArtifacts(paths) {
|
|
|
182
173
|
return {
|
|
183
174
|
stories,
|
|
184
175
|
techSpecContent,
|
|
185
|
-
planAcceptance,
|
|
186
176
|
planContextEnvelope,
|
|
187
177
|
};
|
|
188
178
|
}
|
|
@@ -501,7 +491,6 @@ runAsCli(import.meta.url, main, {
|
|
|
501
491
|
'--plan-context <file>',
|
|
502
492
|
'The plan-context envelope this draft was authored against.',
|
|
503
493
|
],
|
|
504
|
-
['--plan-acceptance <file>', 'Acceptance artifact to attach.'],
|
|
505
494
|
['--source-tickets <ids>', 'Ticket ids this plan supersedes.'],
|
|
506
495
|
['--dry-run', 'Validate and report; create nothing.'],
|
|
507
496
|
['--chain-on-clean', 'Persist immediately when the dry run is clean.'],
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"generatedAt": "2026-09-
|
|
2
|
+
"generatedAt": "2026-09-14T11:57:26.498Z",
|
|
3
3
|
"generator": "generate-skills-index.js@1",
|
|
4
4
|
"skills": [
|
|
5
5
|
{
|
|
@@ -52,16 +52,6 @@
|
|
|
52
52
|
"allowedTools": null,
|
|
53
53
|
"vendor": null
|
|
54
54
|
},
|
|
55
|
-
{
|
|
56
|
-
"name": "scope-triage",
|
|
57
|
-
"tier": "core",
|
|
58
|
-
"category": "core",
|
|
59
|
-
"path": ".agents/skills/core/scope-triage/SKILL.md",
|
|
60
|
-
"description": "Optional split-advisory for `/mandrel-plan`. Under v2 there is no epic|story routing verdict — `/mandrel-plan` always authors Stories. Use this skill only when judging whether a draft should stay one Story or legitimately split (near-zero overlap or an architectural seam).",
|
|
61
|
-
"policyCapsuleBullets": 5,
|
|
62
|
-
"allowedTools": null,
|
|
63
|
-
"vendor": null
|
|
64
|
-
},
|
|
65
55
|
{
|
|
66
56
|
"name": "security-and-hardening",
|
|
67
57
|
"tier": "core",
|
|
@@ -152,15 +152,17 @@ node .agents/scripts/audit-to-stories.js --emit-plan-seed \
|
|
|
152
152
|
|
|
153
153
|
The seed renders the canonical one-pager sections — Problem Statement,
|
|
154
154
|
Recommended Direction, Key Assumptions (with links to every source
|
|
155
|
-
report), MVP Scope (the
|
|
156
|
-
authoring step has concrete anchors),
|
|
157
|
-
|
|
158
|
-
**
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
155
|
+
report), MVP Scope (**the findings, flat**), Key Files (so `/mandrel-plan`'s
|
|
156
|
+
authoring step has concrete anchors), Not Doing.
|
|
157
|
+
|
|
158
|
+
**The seed states findings, not a partition.** MVP Scope used
|
|
159
|
+
to render one numbered bullet per group beneath a `## Grouping` container
|
|
160
|
+
directive — a plan the seed had already decided, at the grouping grain, before
|
|
161
|
+
`/mandrel-plan` read a word of it. N now reaches the planner **undecided**: it
|
|
162
|
+
applies its own cohesion judgment, and container grouping is `/mandrel-plan`'s
|
|
163
|
+
Gate #3 call at persist, where N is known. The grouping still drives the
|
|
164
|
+
**standalone-Stories** path (Phase 5b), which needs one issue payload per
|
|
165
|
+
group.
|
|
164
166
|
|
|
165
167
|
Chain into the existing planning entrypoint:
|
|
166
168
|
|
|
@@ -172,9 +174,10 @@ Chain into the existing planning entrypoint:
|
|
|
172
174
|
then runs its author → persist path, as documented in its workflow.
|
|
173
175
|
|
|
174
176
|
**Dedup provenance is carried mechanically — do not hand-copy it.** The seed's
|
|
175
|
-
MVP Scope
|
|
177
|
+
MVP Scope section carries each group's `audit-fingerprints` and
|
|
176
178
|
`audit-semantic-keys` footers as HTML comments (invisible in the rendered
|
|
177
|
-
one-pager
|
|
179
|
+
one-pager, and per-group even though the visible list is flat — they are the
|
|
180
|
+
identity the next sweep matches on). `plan-persist` harvests them out of the seed on the
|
|
178
181
|
`plan-context.json` envelope and appends them to **every** Story body it
|
|
179
182
|
persists, via `carryProvenanceFooters`
|
|
180
183
|
([`lib/findings/route-finding.js`](../scripts/lib/findings/route-finding.js)).
|
|
@@ -94,26 +94,33 @@ whose `criteria[]` length differs **before** scoring, consuming no round;
|
|
|
94
94
|
not close**: post a `friction` comment and flip `agent::blocked`.
|
|
95
95
|
Per-round mechanics: [`acceptance-self-eval.md`](acceptance-self-eval.md).
|
|
96
96
|
|
|
97
|
-
## 5. The one
|
|
97
|
+
## 5. The one credited suite run
|
|
98
98
|
|
|
99
|
-
After the self-eval loop's last fix commit, run the
|
|
100
|
-
|
|
99
|
+
After the self-eval loop's last fix commit, run the suite **once** in the
|
|
100
|
+
worktree through the depositor — it spawns the project's own `npm test`,
|
|
101
|
+
whatever that resolves to, and stamps the result, so any runner earns the
|
|
102
|
+
credit:
|
|
101
103
|
|
|
102
104
|
```bash
|
|
103
|
-
|
|
105
|
+
node <main-repo>/.agents/scripts/evidence-gate.js --standalone \
|
|
106
|
+
--scope-id <storyId> --gate test --worktree <workCwd> -- npm test
|
|
104
107
|
```
|
|
105
108
|
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
when it needs an artifact.
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
109
|
+
Green deposits the `test` evidence close reads, keyed on the tree, so close
|
|
110
|
+
reports the gate as **credited** at unchanged HEAD — a later commit voids it.
|
|
111
|
+
Read its **output**, not the exit code: `✓ test passed` is the signal. The
|
|
112
|
+
CRAP gate still captures coverage itself when it needs an artifact.
|
|
113
|
+
|
|
114
|
+
A bare `npm test` earns the same credit **only** where the project's test
|
|
115
|
+
script routes through mandrel's own runner, which prints the outcome. On any
|
|
116
|
+
other runner it deposits nothing and prints nothing, so silence is never
|
|
117
|
+
evidence of credit; `mandrel doctor`'s `test-credit-path` check names which
|
|
118
|
+
shape this project is. If the suite outruns the host's sync Bash ceiling,
|
|
119
|
+
dispatch it in the **background** — its completion re-invokes you; never spawn
|
|
120
|
+
a task to poll or `sleep`-loop against it
|
|
121
|
+
([`parallel-tooling.md`](parallel-tooling.md) Rule 2). Redraft rounds run the
|
|
122
|
+
scoped projects for the roots you changed plus `verify[]`, not the whole
|
|
123
|
+
suite; only this run needs credit.
|
|
117
124
|
|
|
118
125
|
`verify[]` is scoped entries **plus** this one run: an entry that is itself a
|
|
119
126
|
full-suite command is reported credited against the same record, never
|
|
@@ -237,17 +237,37 @@ the failure class that actually bounces deliveries: close-validation
|
|
|
237
237
|
discovers them only after the whole close pipeline has run, at several times
|
|
238
238
|
the cost of one full-suite run in the worktree.
|
|
239
239
|
|
|
240
|
-
**Run it once, last,
|
|
241
|
-
self-eval loop's last fix commit; redraft rounds run scoped tests.
|
|
242
|
-
|
|
243
|
-
|
|
244
|
-
|
|
245
|
-
|
|
246
|
-
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
240
|
+
**Run it once, last, through the depositor.** The run belongs **after** the
|
|
241
|
+
self-eval loop's last fix commit; redraft rounds run scoped tests. Run it in
|
|
242
|
+
the worktree on `story-<id>` as
|
|
243
|
+
|
|
244
|
+
```bash
|
|
245
|
+
node <main-repo>/.agents/scripts/evidence-gate.js \
|
|
246
|
+
--standalone --scope-id <storyId> --gate test \
|
|
247
|
+
--worktree <workCwd> -- npm test
|
|
248
|
+
```
|
|
249
|
+
|
|
250
|
+
The wrapper spawns the project's own `npm test` — whatever that resolves to —
|
|
251
|
+
and records the pass into the Story evidence keyspace, so the credit is
|
|
252
|
+
runner-agnostic by construction: it stamps only what it just ran. The record
|
|
253
|
+
is keyed on HEAD and the tree fingerprint and hashed on the exact command
|
|
254
|
+
close spawns, so close reports the gate as credited at unchanged HEAD instead
|
|
255
|
+
of re-running the suite. The credit expires the moment it stops describing the
|
|
256
|
+
tree: any later commit invalidates it and close re-runs the suite for real, so
|
|
257
|
+
this never trades away the gate. The CRAP gate still runs
|
|
258
|
+
`coverage-capture.js` itself when it needs a fresh artifact — the capture
|
|
259
|
+
stamp is a claim about `coverage/coverage-final.json`, which `npm test` alone
|
|
260
|
+
does not produce.
|
|
261
|
+
|
|
262
|
+
**A bare `npm test` is a bonus, not the contract.** It deposits the same
|
|
263
|
+
record only where the project's `test` script routes through mandrel's own
|
|
264
|
+
runner (`run-tests.js` → `lib/test-run-credit.js`, Story #5313), which prints
|
|
265
|
+
the outcome. A project whose `npm test` is `vitest run`, `jest` or any other
|
|
266
|
+
runner never reaches that code, so it prints nothing and deposits nothing —
|
|
267
|
+
silence is not a signal, and nothing here asks you to confirm the credit by
|
|
268
|
+
reading for a line that cannot appear. `mandrel doctor`'s `test-credit-path`
|
|
269
|
+
check reports which shape a project is and names the command above as its
|
|
270
|
+
remedy.
|
|
251
271
|
|
|
252
272
|
**`verify[]` reuses the same credit.** A `verify[]` entry that is itself a
|
|
253
273
|
full-suite command is reported **credited** against that record rather than
|
|
@@ -93,12 +93,13 @@ are reference § Step 2. Hard gates always run in Step 3 — the derived level
|
|
|
93
93
|
never disables them; do **not** pre-run the chain here — Step 2.5's credited
|
|
94
94
|
suite run is the sole exception.
|
|
95
95
|
|
|
96
|
-
### Step 2.5 — The one
|
|
96
|
+
### Step 2.5 — The one credited suite run, the push, then hand off
|
|
97
97
|
|
|
98
|
-
After the self-eval loop's last fix commit, run
|
|
99
|
-
worktree
|
|
100
|
-
|
|
101
|
-
|
|
98
|
+
After the self-eval loop's last fix commit, run the suite **once** in the
|
|
99
|
+
worktree through the depositor — `evidence-gate.js … --gate test -- npm test`,
|
|
100
|
+
spelled out in **digest § 5**. It runs whatever `npm test` resolves to and
|
|
101
|
+
stamps that, so the `test` credit is earned on any runner and only a *later*
|
|
102
|
+
commit invalidates it. Red → fix, commit, re-run.
|
|
102
103
|
|
|
103
104
|
Push `story-<storyId>` to `origin`, confirming the remote ref moved. Then
|
|
104
105
|
(sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
|
|
@@ -47,10 +47,20 @@ The spine's two escape hatches from N=1 are narrow on purpose:
|
|
|
47
47
|
sitting unverifiable behind the other.
|
|
48
48
|
|
|
49
49
|
Everything else is one Story with `## Slicing` checkpoints. When N>1 does
|
|
50
|
-
apply,
|
|
51
|
-
`
|
|
52
|
-
|
|
53
|
-
|
|
50
|
+
apply, the split has to survive the **same-wave collision refusal**: persist
|
|
51
|
+
runs the dispatcher's own `detectCollision` pairwise over the draft, ahead of
|
|
52
|
+
the first `createIssue`, and refuses any pair of same-wave siblings that both
|
|
53
|
+
declare a path (or one of which declares a covering glob) — naming the pair,
|
|
54
|
+
the paths and the two remedies (merge them, or order them with `depends_on`).
|
|
55
|
+
Such a pair cannot be co-dispatched, so the split buys no parallelism and
|
|
56
|
+
costs a delivery session per Story. It replaced the acceptance partition,
|
|
57
|
+
which refused only byte-identical acceptance text across siblings — a shape
|
|
58
|
+
model output does not produce — so it never fired on the fragmentation it was
|
|
59
|
+
meant to catch. **N=1 can never trip the refusal.**
|
|
60
|
+
|
|
61
|
+
A draft of more than one Story also stops at **Gate #2** for operator
|
|
62
|
+
approval, `--force-review` or not: a split always earns eyes. `--yes`
|
|
63
|
+
auto-proceeds, as at every other gate.
|
|
54
64
|
|
|
55
65
|
## Unknown triage — AFK vs HITL
|
|
56
66
|
|
|
@@ -154,6 +164,11 @@ and ceremony is derived from the landed diff at close.
|
|
|
154
164
|
- **`verify[]` entries are commands.** There is no tier suffix and no
|
|
155
165
|
`manual:<reason>` escape (Story #5312): write the exact command or test
|
|
156
166
|
path the deliverer runs and the acceptance critic reads as evidence.
|
|
167
|
+
- **`changes[]` names what the deliverer authors.** Generated artifacts —
|
|
168
|
+
quality baselines, generated test indexes, migration journals, lockfiles —
|
|
169
|
+
are omitted: the work regenerates them, the refresh is a close-gate concern,
|
|
170
|
+
and a declared shared artifact path reserves a footprint that needlessly
|
|
171
|
+
serializes sibling Stories at dispatch.
|
|
157
172
|
- **`changes[]` arrive pre-resolved to creates-vs-refactors.** Every path
|
|
158
173
|
the seed predicted is probed against the repo: an existing path is
|
|
159
174
|
emitted with `assumption: "refactors-existing"`, a missing one with
|
|
@@ -175,8 +190,14 @@ kept — passes the persist ticket validators with no round-trip.
|
|
|
175
190
|
Each `stories.json` entry: `slug` (`^[a-z0-9][a-z0-9-]*$`), `type: "story"`,
|
|
176
191
|
`title`, `body` (`goal`, optional `spec`, `changes[{path, assumption}]` —
|
|
177
192
|
`creates|refactors-existing|deletes`, `non_goals`, `reason_to_exist`),
|
|
178
|
-
top-level `acceptance[]`, `verify[]` (
|
|
179
|
-
`depends_on[]` (a sibling slug, or `#<id>` for an existing
|
|
193
|
+
top-level `acceptance[]`, `verify[]` (each a **bare command** — there is no
|
|
194
|
+
tier suffix), and `depends_on[]` (a sibling slug, or `#<id>` for an existing
|
|
195
|
+
open Story).
|
|
196
|
+
|
|
197
|
+
Author `acceptance[]` **without** the `AC-<n>:` handle: the body renderer
|
|
198
|
+
numbers each checkbox from its array position, so a carried handle renders
|
|
199
|
+
doubled. Persist normalises one off rather than refusing, and names the strip
|
|
200
|
+
on the dry-run's repair list.
|
|
180
201
|
|
|
181
202
|
Nothing in that shape inventories the repo for the author. `changes[]` arrives
|
|
182
203
|
pre-resolved against the working tree, and Phase 8's
|
|
@@ -238,8 +259,9 @@ The conflict passes run **twice**: once over the raw `stories.json` payload
|
|
|
238
259
|
The second pass is not belt-and-braces. The canonical authoring shape carries
|
|
239
260
|
`acceptance[]` / `verify[]` at the ticket's top level and assembly folds them
|
|
240
261
|
into the body, so the passes that scan `body.acceptance` / `body.verify`
|
|
241
|
-
|
|
242
|
-
|
|
262
|
+
saw two empty arrays on the real payload and emitted nothing; the two
|
|
263
|
+
substring-match advisories that depended on it are retired, leaving
|
|
264
|
+
`shared-editor` as the one conflict kind. Both passes complete before the first
|
|
243
265
|
`createIssue`, so a refusal still costs no writes.
|
|
244
266
|
|
|
245
267
|
`shared-editor` findings are rendered into the posted `plan-summary` comment,
|
|
@@ -256,6 +278,21 @@ and a path reference matched by substring can read as a dependency a prose
|
|
|
256
278
|
mention never meant. A finding names the Stories and the fix (a `depends_on`
|
|
257
279
|
edge, or folding the shared edit into one Story) for the operator to weigh.
|
|
258
280
|
|
|
281
|
+
## Tickets mode — the source ticket is evidence
|
|
282
|
+
|
|
283
|
+
A `--tickets` envelope carries a third author prompt beside
|
|
284
|
+
`systemPrompts.story` and `systemPrompts.storySplitRules`:
|
|
285
|
+
**`systemPrompts.storyTicketsRules`**. It exists because a
|
|
286
|
+
source ticket arrives already in Story shape — rendered `AC-<n>:` checkboxes,
|
|
287
|
+
a `## Verify` list, a `## Changes` footprint — and an author reading it as a
|
|
288
|
+
template carries that shape forward instead of re-deriving it. The addendum
|
|
289
|
+
binds the author to re-derive `acceptance[]` from the goal, to express
|
|
290
|
+
mechanical checks (a refreshed baseline, a lint exiting 0, a regenerated
|
|
291
|
+
index) as `verify[]` commands rather than acceptance items, and to take the
|
|
292
|
+
source's verify entries for the commands they name rather than their shape.
|
|
293
|
+
Read it whenever the mode is `tickets`; the other three modes do not carry
|
|
294
|
+
the field.
|
|
295
|
+
|
|
259
296
|
## Tickets mode — authoring `supersedes[]`
|
|
260
297
|
|
|
261
298
|
In `--tickets` mode each Story carries a top-level `supersedes` array claiming
|
|
@@ -282,7 +319,7 @@ template-only prose.
|
|
|
282
319
|
### Supersede-map partition
|
|
283
320
|
|
|
284
321
|
`plan-persist` refuses a partial supersede map **before** it creates any
|
|
285
|
-
Story
|
|
322
|
+
Story, the same fail-closed shape as the collision refusal: every id passed to
|
|
286
323
|
`--tickets` must be claimed by **exactly one** Story, and no Story may
|
|
287
324
|
claim an id that was not a source ticket. With N>1 the mapping is not
|
|
288
325
|
total by default — an authored map is the only thing that can say
|
|
@@ -347,10 +384,13 @@ at base.
|
|
|
347
384
|
exists at base or a `refactors-existing` on one that does not (including a
|
|
348
385
|
path the base branch deleted or renamed, named with the removing commit), a
|
|
349
386
|
goal or acceptance path absent at base, a `verify[]` command naming an absent
|
|
350
|
-
test file,
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
|
|
387
|
+
test file, an `open-question` in a body (`Flag if…`, `TBD`, a trailing `?`),
|
|
388
|
+
and a `pinned-identifier` in an acceptance item — a backticked bare symbol
|
|
389
|
+
that is not a path, a label, a kebab token, a flag or a command, which the
|
|
390
|
+
advisory `changes[]` is free to reshape out from under the criterion. The
|
|
391
|
+
list also names every **repair** the run applied — a plain-string bullet or a
|
|
392
|
+
trailing parenthetical rewritten into `{ path, assumption }` by probing base,
|
|
393
|
+
and an `AC-<n>:` handle normalised off an acceptance item. The same list rides the result
|
|
354
394
|
envelope as `warnings[]` and `repairs[]`, so a `--chain-on-clean` run loses
|
|
355
395
|
nothing.
|
|
356
396
|
|