mandrel 2.57.0 → 2.58.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (24) hide show
  1. package/.agents/agents/story-worker.md +12 -11
  2. package/.agents/scripts/evidence-gate.js +17 -1
  3. package/.agents/scripts/lib/orchestration/code-review.js +7 -3
  4. package/.agents/scripts/lib/orchestration/pinned-identifier-lint.js +137 -0
  5. package/.agents/scripts/lib/orchestration/plan-context.js +11 -3
  6. package/.agents/scripts/lib/orchestration/plan-persist/acceptance-handle-repair.js +107 -0
  7. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +6 -1
  8. package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +14 -9
  9. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +17 -7
  10. package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +15 -5
  11. package/.agents/scripts/lib/orchestration/review-base-ref.js +138 -0
  12. package/.agents/scripts/lib/orchestration/single-story-close/phases/code-review.js +37 -5
  13. package/.agents/scripts/lib/orchestration/single-story-close/runner.js +6 -1
  14. package/.agents/scripts/lib/story-body/story-body.js +36 -2
  15. package/.agents/scripts/lib/templates/decomposer-prompts.js +53 -4
  16. package/.agents/scripts/lib/test-run-credit.js +23 -12
  17. package/.agents/workflows/helpers/deliver-digest.md +22 -15
  18. package/.agents/workflows/helpers/deliver-story-reference.md +31 -11
  19. package/.agents/workflows/helpers/deliver-story.md +6 -5
  20. package/.agents/workflows/helpers/plan-reference.md +35 -6
  21. package/.agents/workflows/mandrel-plan.md +6 -2
  22. package/docs/CHANGELOG.md +13 -0
  23. package/lib/cli/registry.js +98 -2
  24. package/package.json +1 -1
@@ -154,6 +154,11 @@ and ceremony is derived from the landed diff at close.
154
154
  - **`verify[]` entries are commands.** There is no tier suffix and no
155
155
  `manual:<reason>` escape (Story #5312): write the exact command or test
156
156
  path the deliverer runs and the acceptance critic reads as evidence.
157
+ - **`changes[]` names what the deliverer authors.** Generated artifacts —
158
+ quality baselines, generated test indexes, migration journals, lockfiles —
159
+ are omitted: the work regenerates them, the refresh is a close-gate concern,
160
+ and a declared shared artifact path reserves a footprint that needlessly
161
+ serializes sibling Stories at dispatch.
157
162
  - **`changes[]` arrive pre-resolved to creates-vs-refactors.** Every path
158
163
  the seed predicted is probed against the repo: an existing path is
159
164
  emitted with `assumption: "refactors-existing"`, a missing one with
@@ -175,8 +180,14 @@ kept — passes the persist ticket validators with no round-trip.
175
180
  Each `stories.json` entry: `slug` (`^[a-z0-9][a-z0-9-]*$`), `type: "story"`,
176
181
  `title`, `body` (`goal`, optional `spec`, `changes[{path, assumption}]` —
177
182
  `creates|refactors-existing|deletes`, `non_goals`, `reason_to_exist`),
178
- top-level `acceptance[]`, `verify[]` (`… (unit|contract|e2e|validate)`), and
179
- `depends_on[]` (a sibling slug, or `#<id>` for an existing open Story).
183
+ top-level `acceptance[]`, `verify[]` (each a **bare command** — there is no
184
+ tier suffix), and `depends_on[]` (a sibling slug, or `#<id>` for an existing
185
+ open Story).
186
+
187
+ Author `acceptance[]` **without** the `AC-<n>:` handle: the body renderer
188
+ numbers each checkbox from its array position, so a carried handle renders
189
+ doubled. Persist normalises one off rather than refusing, and names the strip
190
+ on the dry-run's repair list.
180
191
 
181
192
  Nothing in that shape inventories the repo for the author. `changes[]` arrives
182
193
  pre-resolved against the working tree, and Phase 8's
@@ -256,6 +267,21 @@ and a path reference matched by substring can read as a dependency a prose
256
267
  mention never meant. A finding names the Stories and the fix (a `depends_on`
257
268
  edge, or folding the shared edit into one Story) for the operator to weigh.
258
269
 
270
+ ## Tickets mode — the source ticket is evidence
271
+
272
+ A `--tickets` envelope carries a third author prompt beside
273
+ `systemPrompts.story` and `systemPrompts.storySplitRules`:
274
+ **`systemPrompts.storyTicketsRules`**. It exists because a
275
+ source ticket arrives already in Story shape — rendered `AC-<n>:` checkboxes,
276
+ a `## Verify` list, a `## Changes` footprint — and an author reading it as a
277
+ template carries that shape forward instead of re-deriving it. The addendum
278
+ binds the author to re-derive `acceptance[]` from the goal, to express
279
+ mechanical checks (a refreshed baseline, a lint exiting 0, a regenerated
280
+ index) as `verify[]` commands rather than acceptance items, and to take the
281
+ source's verify entries for the commands they name rather than their shape.
282
+ Read it whenever the mode is `tickets`; the other three modes do not carry
283
+ the field.
284
+
259
285
  ## Tickets mode — authoring `supersedes[]`
260
286
 
261
287
  In `--tickets` mode each Story carries a top-level `supersedes` array claiming
@@ -347,10 +373,13 @@ at base.
347
373
  exists at base or a `refactors-existing` on one that does not (including a
348
374
  path the base branch deleted or renamed, named with the removing commit), a
349
375
  goal or acceptance path absent at base, a `verify[]` command naming an absent
350
- test file, and an `open-question` in a body (`Flag if…`, `TBD`, a trailing
351
- `?`). The list also names every `changes[]` **repair** the run applied — a
352
- plain-string bullet or a trailing parenthetical rewritten into
353
- `{ path, assumption }` by probing base. The same list rides the result
376
+ test file, an `open-question` in a body (`Flag if…`, `TBD`, a trailing `?`),
377
+ and a `pinned-identifier` in an acceptance item — a backticked bare symbol
378
+ that is not a path, a label, a kebab token, a flag or a command, which the
379
+ advisory `changes[]` is free to reshape out from under the criterion. The
380
+ list also names every **repair** the run applied — a plain-string bullet or a
381
+ trailing parenthetical rewritten into `{ path, assumption }` by probing base,
382
+ and an `AC-<n>:` handle normalised off an acceptance item. The same list rides the result
354
383
  envelope as `warnings[]` and `repairs[]`, so a `--chain-on-clean` run loses
355
384
  nothing.
356
385
 
@@ -57,7 +57,8 @@ and derives source ids from its `sourceTickets[]`; it also writes
57
57
  **`stories.template.json`**, step 2's skeleton.
58
58
 
59
59
  The envelope carries docs context, the story-author prompt (`systemPrompts.story`,
60
- plus `systemPrompts.storySplitRules` for an N>1 draft), `sourceTickets[]`,
60
+ plus `systemPrompts.storySplitRules` for an N>1 draft and
61
+ `systemPrompts.storyTicketsRules` in tickets mode), `sourceTickets[]`,
61
62
  `duplicates[]` (open **Stories**, never Epics), `epicCandidates[]` +
62
63
  `dependencyCandidates[]` (Gate #3; path collisions), `priorFeedback` and
63
64
  advisory `complexitySignals` (**no routing authority**). An envelope over the
@@ -97,7 +98,10 @@ a Spec is as long as the work needs, inline, never under `docs/`); optional
97
98
  `acceptance-manifest.json` (N>1 — `--plan-acceptance`). Use the envelope
98
99
  `systemPrompts.story`; split only under the policy above, and when you do,
99
100
  read `systemPrompts.storySplitRules` too — it carries the schedule and
100
- partition rules the core omits.
101
+ partition rules the core omits. In **tickets mode** also read
102
+ `systemPrompts.storyTicketsRules`: the source ticket is evidence, not a
103
+ template — re-derive `acceptance[]` rather than carrying its list, handles
104
+ and tier suffixes forward.
101
105
 
102
106
  **Tickets mode:** every Story authors a top-level `supersedes[]`; persist
103
107
  refuses a partial map ([shape](helpers/plan-reference.md)).
package/docs/CHANGELOG.md CHANGED
@@ -15,6 +15,19 @@ All notable changes to this project will be documented in this file.
15
15
  -->
16
16
  <!-- markdownlint-disable-file MD004 MD012 MD037 -->
17
17
 
18
+ ## [2.58.0](https://github.com/dsj1984/mandrel/compare/mandrel-v2.57.0...mandrel-v2.58.0) (2026-09-12)
19
+
20
+
21
+ ### Added
22
+
23
+ * make the close test credit earnable on any test runner, and surface the gap at setup instead of mid-close ([#5324](https://github.com/dsj1984/mandrel/issues/5324)) ([#5329](https://github.com/dsj1984/mandrel/issues/5329)) ([04db145](https://github.com/dsj1984/mandrel/commit/04db145b8f1017121e129897f9207ab51db1c3f8))
24
+ * tickets-mode planning re-derives the Story instead of carrying the source ticket's shape: acceptance handles, tier suffixes, generated-artifact footprints and pinned identifiers ([#5323](https://github.com/dsj1984/mandrel/issues/5323)) ([#5326](https://github.com/dsj1984/mandrel/issues/5326)) ([5563e7a](https://github.com/dsj1984/mandrel/commit/5563e7a264864f78001beb19ce58160a1b483cda))
25
+
26
+
27
+ ### Fixed
28
+
29
+ * give base-sync and the Story-scope code review one base ref, so a stale local base branch cannot raise false blockers ([#5325](https://github.com/dsj1984/mandrel/issues/5325)) ([#5328](https://github.com/dsj1984/mandrel/issues/5328)) ([c4032b6](https://github.com/dsj1984/mandrel/commit/c4032b63a63c4cb8cc3f9c9956b6c646e5acd460))
30
+
18
31
  ## [2.57.0](https://github.com/dsj1984/mandrel/compare/mandrel-v2.56.0...mandrel-v2.57.0) (2026-09-12)
19
32
 
20
33
 
@@ -4,8 +4,11 @@
4
4
  *
5
5
  * Exports an ordered array of check objects each shaped `{ name, run() }`.
6
6
  * `run()` returns `{ ok, detail, remedy? }` — `remedy` is present and
7
- * non-empty only when `ok` is false. The registry is the single source of
8
- * truth for which checks the doctor command runs and in what order.
7
+ * non-empty whenever `ok` is false, and on an `advisory` check that passes
8
+ * while still having something actionable to say (`test-credit-path`), which
9
+ * repeats it in `detail` because that is the field the doctor prints for a
10
+ * passing check. The registry is the single source of truth for which checks
11
+ * the doctor command runs and in what order.
9
12
  *
10
13
  * Checks run sequentially in the doctor runner (not in parallel) because
11
14
  * some checks are meaningless without a prerequisite having passed first
@@ -1176,6 +1179,90 @@ function probeConfiguredDriver(command, runner, projectRoot) {
1176
1179
  };
1177
1180
  }
1178
1181
 
1182
+ // ---------------------------------------------------------------------------
1183
+ // check: test-credit-path
1184
+ // ---------------------------------------------------------------------------
1185
+
1186
+ /**
1187
+ * The runner whose own green full run deposits the close `test` credit as a
1188
+ * side effect (`.agents/scripts/run-tests.js` → `lib/test-run-credit.js`).
1189
+ */
1190
+ const MANDREL_TEST_RUNNER = 'run-tests.js';
1191
+
1192
+ /**
1193
+ * The deposit that works whatever `npm test` resolves to: it spawns the
1194
+ * project's own suite and stamps the result, so it is honest by construction.
1195
+ */
1196
+ const TEST_CREDIT_DEPOSIT_COMMAND =
1197
+ 'node .agents/scripts/evidence-gate.js --standalone --scope-id <storyId> --gate test --worktree <workCwd> -- npm test';
1198
+
1199
+ /** Remedy shared by every shape that does not earn the credit on its own. */
1200
+ const TEST_CREDIT_REMEDY = `run the suite through the depositor instead of bare \`npm test\`: ${TEST_CREDIT_DEPOSIT_COMMAND}`;
1201
+
1202
+ /**
1203
+ * The project's `test` script — what `npm test` runs, and therefore what the
1204
+ * close `test` gate spawns. Deliberately read from `package.json` rather than
1205
+ * `project.commands.test`: the close gate's argv is the literal `npm test`, so
1206
+ * the npm script is the command whose shape decides whether a bare run
1207
+ * deposits anything.
1208
+ *
1209
+ * @param {string} projectRoot
1210
+ * @param {typeof fs.readFileSync} readFileImpl
1211
+ * @returns {string|null} The trimmed script, or null when there is none.
1212
+ */
1213
+ function readProjectTestScript(projectRoot, readFileImpl) {
1214
+ try {
1215
+ const raw = readFileImpl(path.join(projectRoot, 'package.json'), 'utf8');
1216
+ const script = JSON.parse(raw)?.scripts?.test;
1217
+ return typeof script === 'string' && script.trim().length > 0
1218
+ ? script.trim()
1219
+ : null;
1220
+ } catch {
1221
+ return null;
1222
+ }
1223
+ }
1224
+
1225
+ /**
1226
+ * Report whether this project's test command reaches mandrel's own runner,
1227
+ * and therefore whether a bare `npm test` deposits the close `test` credit
1228
+ * by itself (Story #5324).
1229
+ *
1230
+ * **Always `ok: true`.** A project on `vitest`, `jest` or any other runner is
1231
+ * a supported setup, not a broken install — the only thing it lacks is the
1232
+ * runner-side bonus deposit, and the remedy command covers it. So the check
1233
+ * informs rather than gates, and (unlike every other entry here) carries its
1234
+ * `remedy` alongside a passing verdict; the same guidance also rides in
1235
+ * `detail`, which is the field the doctor prints for a passing check.
1236
+ *
1237
+ * The runner is detected by name, never by executing anything: a script that
1238
+ * reaches `run-tests.js` indirectly reads as "own runner", which errs toward
1239
+ * naming the deposit command — correct on either shape.
1240
+ *
1241
+ * @param {{ projectRoot?: string, readFile?: typeof fs.readFileSync }} [opts]
1242
+ * @returns {{ ok: boolean, detail: string, remedy?: string }}
1243
+ */
1244
+ function runTestCreditPath({ projectRoot, readFile = fs.readFileSync } = {}) {
1245
+ const script = readProjectTestScript(projectRoot ?? process.cwd(), readFile);
1246
+ if (script === null) {
1247
+ return {
1248
+ ok: true,
1249
+ detail: `no \`test\` script in package.json — close still spawns \`npm test\`, so ${TEST_CREDIT_REMEDY}`,
1250
+ remedy: TEST_CREDIT_REMEDY,
1251
+ };
1252
+ }
1253
+ if (script.includes(MANDREL_TEST_RUNNER)) {
1254
+ return {
1255
+ ok: true,
1256
+ detail: `\`npm test\` → \`${script}\` reaches mandrel's runner — a green full run on a story branch deposits the close \`test\` credit itself`,
1257
+ };
1258
+ }
1259
+ return {
1260
+ ok: true,
1261
+ detail: `\`npm test\` → \`${script}\` is this project's own runner and never reaches \`${MANDREL_TEST_RUNNER}\`, so it deposits no close \`test\` credit and prints nothing — ${TEST_CREDIT_REMEDY}`,
1262
+ remedy: TEST_CREDIT_REMEDY,
1263
+ };
1264
+ }
1265
+
1179
1266
  /**
1180
1267
  * Ordered array of doctor checks. Each entry follows the
1181
1268
  * `{ name: string, run(opts?): { ok: boolean, detail: string, remedy?: string } }` contract.
@@ -1243,6 +1330,15 @@ export const registry = [
1243
1330
  advisory: true,
1244
1331
  run: (opts) => runVersionCurrent(opts),
1245
1332
  },
1333
+ {
1334
+ name: 'test-credit-path',
1335
+ // Non-fatal: `run()` always returns ok:true. A project on its own test
1336
+ // runner is a supported setup that simply has to deposit the close
1337
+ // `test` credit explicitly, so this reports a condition rather than
1338
+ // failing the install (Story #5324).
1339
+ advisory: true,
1340
+ run: (opts) => runTestCreditPath(opts),
1341
+ },
1246
1342
  ];
1247
1343
 
1248
1344
  export default registry;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "mandrel",
3
- "version": "2.57.0",
3
+ "version": "2.58.0",
4
4
  "description": "Claude Code-first opinionated workflow framework: instructions, skills, rules, and SDLC workflows that govern AI coding assistants.",
5
5
  "files": [
6
6
  ".agents/",