mandrel 2.56.0 → 2.58.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (114) hide show
  1. package/.agents/agents/plan-critic.md +13 -18
  2. package/.agents/agents/story-worker.md +25 -33
  3. package/.agents/docs/agentrc-reference.json +0 -30
  4. package/.agents/docs/configuration.md +8 -28
  5. package/.agents/docs/execution-reference.md +5 -5
  6. package/.agents/docs/quality-gates.md +8 -7
  7. package/.agents/instructions.md +9 -10
  8. package/.agents/schemas/agentrc.schema.json +9 -185
  9. package/.agents/schemas/story-deliver-terminal.schema.json +1 -1
  10. package/.agents/scripts/acceptance-eval.js +107 -17
  11. package/.agents/scripts/ceremony-derive.js +191 -0
  12. package/.agents/scripts/check-context-budget.js +28 -33
  13. package/.agents/scripts/check-cyclomatic.js +4 -3
  14. package/.agents/scripts/deliver-light.js +31 -94
  15. package/.agents/scripts/evidence-gate.js +17 -1
  16. package/.agents/scripts/lib/audit-suite/checklist-threading.js +15 -2
  17. package/.agents/scripts/lib/baselines/coverage-updater-cli.js +110 -0
  18. package/.agents/scripts/lib/baselines/crap-preview-scan.js +25 -0
  19. package/.agents/scripts/lib/baselines/crap-updater-cli.js +223 -0
  20. package/.agents/scripts/lib/bdd-scenario-budget.js +21 -3
  21. package/.agents/scripts/lib/bootstrap/quality-bootstrap.js +0 -1
  22. package/.agents/scripts/lib/close-validation/gates.js +52 -1
  23. package/.agents/scripts/lib/config/acceptance-eval.js +25 -57
  24. package/.agents/scripts/lib/config/delivery-routing.js +7 -33
  25. package/.agents/scripts/lib/config/explain.js +0 -19
  26. package/.agents/scripts/lib/config/limits.js +18 -78
  27. package/.agents/scripts/lib/config/quality.js +6 -3
  28. package/.agents/scripts/lib/config/runners.js +3 -2
  29. package/.agents/scripts/lib/config-settings-schema-delivery.js +15 -68
  30. package/.agents/scripts/lib/config-settings-schema-quality.js +0 -14
  31. package/.agents/scripts/lib/config-settings-schema.js +16 -143
  32. package/.agents/scripts/lib/crap-engine.js +35 -4
  33. package/.agents/scripts/lib/crap-utils.js +17 -1
  34. package/.agents/scripts/lib/cyclomatic-ceiling.js +19 -7
  35. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  36. package/.agents/scripts/lib/observability/runtime-friction.js +1 -1
  37. package/.agents/scripts/lib/observability/source-classifier.js +1 -0
  38. package/.agents/scripts/lib/orchestration/acceptance-eval-decision.js +5 -4
  39. package/.agents/scripts/lib/orchestration/ceremony-routing.js +19 -73
  40. package/.agents/scripts/lib/orchestration/code-review.js +7 -3
  41. package/.agents/scripts/lib/orchestration/complexity-gate.js +46 -212
  42. package/.agents/scripts/lib/orchestration/file-assumptions.js +32 -17
  43. package/.agents/scripts/lib/orchestration/light-escalation.js +3 -3
  44. package/.agents/scripts/lib/orchestration/light-suitability.js +66 -233
  45. package/.agents/scripts/lib/orchestration/pinned-identifier-lint.js +137 -0
  46. package/.agents/scripts/lib/orchestration/plan-context.js +189 -387
  47. package/.agents/scripts/lib/orchestration/plan-critic-conditions.js +42 -153
  48. package/.agents/scripts/lib/orchestration/plan-critics-evaluate.js +14 -70
  49. package/.agents/scripts/lib/orchestration/plan-persist/acceptance-handle-repair.js +107 -0
  50. package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +305 -0
  51. package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +138 -170
  52. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +128 -297
  53. package/.agents/scripts/lib/orchestration/plan-persist/soft-findings.js +55 -0
  54. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +16 -65
  55. package/.agents/scripts/lib/orchestration/plan-persist/wave-serialisation.js +22 -35
  56. package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +36 -135
  57. package/.agents/scripts/lib/orchestration/planning/memory-pool-advisory.js +61 -223
  58. package/.agents/scripts/lib/orchestration/review-base-ref.js +138 -0
  59. package/.agents/scripts/lib/orchestration/single-story-close/phases/close-validation.js +5 -0
  60. package/.agents/scripts/lib/orchestration/single-story-close/phases/code-review.js +37 -5
  61. package/.agents/scripts/lib/orchestration/single-story-close/phases/pre-gate-steps.js +46 -16
  62. package/.agents/scripts/lib/orchestration/single-story-close/runner.js +6 -1
  63. package/.agents/scripts/lib/orchestration/story-close/context-budget-writeback.js +213 -0
  64. package/.agents/scripts/lib/orchestration/task-body-validator.js +10 -63
  65. package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +33 -539
  66. package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +21 -414
  67. package/.agents/scripts/lib/orchestration/ticket-validator.js +54 -118
  68. package/.agents/scripts/lib/orchestration/verify-credit.js +69 -24
  69. package/.agents/scripts/lib/story-body/body-format-lints.js +15 -85
  70. package/.agents/scripts/lib/story-body/story-body.js +54 -240
  71. package/.agents/scripts/lib/templates/decomposer-prompts.js +133 -121
  72. package/.agents/scripts/lib/test-isolate/cli-options.js +93 -0
  73. package/.agents/scripts/lib/test-isolate/progress-log.js +45 -0
  74. package/.agents/scripts/lib/test-isolate/render-report.js +97 -0
  75. package/.agents/scripts/lib/test-isolate/run-isolate.js +87 -0
  76. package/.agents/scripts/lib/test-run-credit.js +277 -0
  77. package/.agents/scripts/lib/wave-runner/footprint.js +48 -358
  78. package/.agents/scripts/lib/wave-runner/ready-set.js +6 -5
  79. package/.agents/scripts/lib/workers/crap-worker.js +32 -41
  80. package/.agents/scripts/plan-context.js +7 -9
  81. package/.agents/scripts/plan-critics.js +28 -54
  82. package/.agents/scripts/plan-persist.js +25 -68
  83. package/.agents/scripts/quality-preview.js +51 -0
  84. package/.agents/scripts/run-tests.js +12 -0
  85. package/.agents/scripts/stories-wave-tick.js +23 -45
  86. package/.agents/scripts/test-isolate.js +13 -180
  87. package/.agents/scripts/update-coverage-baseline.js +25 -70
  88. package/.agents/scripts/update-crap-baseline.js +19 -123
  89. package/.agents/skills/core/scope-triage/SKILL.md +3 -3
  90. package/.agents/workflows/audit-clean-code.md +4 -3
  91. package/.agents/workflows/helpers/acceptance-self-eval.md +41 -41
  92. package/.agents/workflows/helpers/code-quality-guardrails.md +4 -4
  93. package/.agents/workflows/helpers/code-review.md +2 -3
  94. package/.agents/workflows/helpers/deliver-digest.md +46 -55
  95. package/.agents/workflows/helpers/deliver-light.md +40 -105
  96. package/.agents/workflows/helpers/deliver-reference.md +1 -1
  97. package/.agents/workflows/helpers/deliver-story-reference.md +54 -55
  98. package/.agents/workflows/helpers/deliver-story.md +10 -13
  99. package/.agents/workflows/helpers/plan-reference.md +163 -221
  100. package/.agents/workflows/mandrel-plan.md +31 -40
  101. package/.agents/workflows/memory-consolidate.md +9 -13
  102. package/docs/CHANGELOG.md +36 -0
  103. package/lib/cli/registry.js +98 -2
  104. package/lib/migrations/index.js +4 -0
  105. package/lib/migrations/steps/2.57.0-retire-delivery-limit-knobs.js +45 -0
  106. package/lib/migrations/steps/2.57.0-retire-planning-limit-knobs.js +59 -0
  107. package/package.json +1 -1
  108. package/.agents/scripts/lib/framework-version.js +0 -39
  109. package/.agents/scripts/lib/orchestration/consolidation-precondition.js +0 -223
  110. package/.agents/scripts/lib/orchestration/plan-persist/fan-out-gate.js +0 -97
  111. package/.agents/scripts/lib/orchestration/planning/decomposer-context.js +0 -26
  112. package/.agents/scripts/lib/orchestration/spec-budget.js +0 -89
  113. package/.agents/scripts/lib/orchestration/spec-spill.js +0 -74
  114. package/.agents/scripts/lib/orchestration/verify-tier-repair.js +0 -107
@@ -140,10 +140,8 @@ full-ceremony Story. The lite route's `preserves` field is the machine-readable
140
140
  record of those non-negotiables; there is no lite-specific gate bypass.
141
141
 
142
142
  **Ceremony comes from the landed diff; the dispatch mode comes from the
143
- run.** Persist stamps a lite cohort's Stories with the `route::lite` label as a
144
- _human-visible hint only_ (and ledgers the authored verdict — recorded reason
145
- plus per-Story shape evidence — on the `story-plan-state` checkpoint); the
146
- label is never the control signal. Ceremony is resolved from the **derived
143
+ run.** Persist stamps no route label (Story #5312 retired the plan-side lite
144
+ claim with its `route::lite` hint). Ceremony is resolved from the **derived
147
145
  change level** (`deriveChangeLevel` over the computed change set — digest § 3),
148
146
  not from a body-shape read: a footprint intersecting a sensitive-path class
149
147
  derives `high`, so the Story keeps its fresh acceptance critic. The light path
@@ -239,36 +237,40 @@ the failure class that actually bounces deliveries: close-validation
239
237
  discovers them only after the whole close pipeline has run, at several times
240
238
  the cost of one full-suite run in the worktree.
241
239
 
242
- **Run it once, last, so close can credit it.** The run belongs **after** the
243
- self-eval loop's last fix commit and **after** the hand-off push — the credit
244
- is keyed on the tree, not on push state, so pushing first keeps the stamp and
245
- buys the ordering Step 2.5 needs (the capture is backgrounded, and its
246
- completion ends the turn); redraft rounds run scoped tests. Close skips a gate that already passed at the current HEAD, but a bare
247
- `npm test` deposits no such record — the suite then runs twice per delivery,
248
- once here and once in the close gate chain. Pick the invocation by the same
249
- predicate `close-validation/gates.js` uses to choose its test gate:
240
+ **Run it once, last, through the depositor.** The run belongs **after** the
241
+ self-eval loop's last fix commit; redraft rounds run scoped tests. Run it in
242
+ the worktree on `story-<id>` as
250
243
 
251
244
  ```bash
252
- # CRAP gate enabled (default) + a `test:coverage` script — writes the stamp
253
- # the close `coverage-capture` gate reads:
254
- node <main-repo>/.agents/scripts/coverage-capture.js --cwd <workCwd>
255
- # otherwise — the evidence record the close `test` gate reads. <workCwd> must
256
- # be ABSOLUTE and the runner exactly `npm test`: both sides hash
257
- # {cmd, args, cwd}, so a relative path or a wrapper misses the credit.
258
- node <main-repo>/.agents/scripts/evidence-gate.js --standalone \
259
- --scope-id <storyId> --gate test --worktree <workCwd> -- npm test
245
+ node <main-repo>/.agents/scripts/evidence-gate.js \
246
+ --standalone --scope-id <storyId> --gate test \
247
+ --worktree <workCwd> -- npm test
260
248
  ```
261
249
 
262
- The credit expires the moment it stops describing the tree: evidence is keyed
263
- on HEAD, the capture stamp on a content digest of `crap.targetDirs`. A
264
- self-eval fix — or any commit — invalidates it and close re-runs the suite for
265
- real, so this never trades away the gate. That keying is exactly why the run
266
- comes last, and why the push before it is free. Close's own base-sync can
267
- spend the stamp too when it lands base commits; it now says so out loud rather
268
- than silently re-running the suite.
269
-
270
- **`verify[]` reuses the same stamp.** A `verify[]` entry that is itself a
271
- full-suite command is reported **credited** against that stamp rather than
250
+ The wrapper spawns the project's own `npm test` — whatever that resolves to —
251
+ and records the pass into the Story evidence keyspace, so the credit is
252
+ runner-agnostic by construction: it stamps only what it just ran. The record
253
+ is keyed on HEAD and the tree fingerprint and hashed on the exact command
254
+ close spawns, so close reports the gate as credited at unchanged HEAD instead
255
+ of re-running the suite. The credit expires the moment it stops describing the
256
+ tree: any later commit invalidates it and close re-runs the suite for real, so
257
+ this never trades away the gate. The CRAP gate still runs
258
+ `coverage-capture.js` itself when it needs a fresh artifact — the capture
259
+ stamp is a claim about `coverage/coverage-final.json`, which `npm test` alone
260
+ does not produce.
261
+
262
+ **A bare `npm test` is a bonus, not the contract.** It deposits the same
263
+ record only where the project's `test` script routes through mandrel's own
264
+ runner (`run-tests.js` → `lib/test-run-credit.js`, Story #5313), which prints
265
+ the outcome. A project whose `npm test` is `vitest run`, `jest` or any other
266
+ runner never reaches that code, so it prints nothing and deposits nothing —
267
+ silence is not a signal, and nothing here asks you to confirm the credit by
268
+ reading for a line that cannot appear. `mandrel doctor`'s `test-credit-path`
269
+ check reports which shape a project is and names the command above as its
270
+ remedy.
271
+
272
+ **`verify[]` reuses the same credit.** A `verify[]` entry that is itself a
273
+ full-suite command is reported **credited** against that record rather than
272
274
  respawned (`resolveVerifyCredit` in
273
275
  [`verify-credit.js`](../../scripts/lib/orchestration/verify-credit.js)), and the
274
276
  self-eval gate warns when it sees one: the intended shape is scoped `verify[]`
@@ -294,24 +296,23 @@ round cap, proceed / redraft / block — not an independent additional pass
294
296
  over the criteria. The M4-B floor holds: one verdict per cluster, with the
295
297
  cluster count owned by the dispatching caller and never by routing.
296
298
 
297
- **One round = N cluster critics → ONE merged verdict → ONE gate call.** The
299
+ **One round = N cluster critics → ONE merged verdict → ONE gate call** (fresh
300
+ critics; an inline-owned verdict is one file scored in one call). The
298
301
  clusters are how a round is _authored_; they are not how it is _scored_.
299
302
  Concatenate every cluster's records into a single `criteria[]` ordered by
300
303
  `index` — exactly one per `acceptance[]` item, under one `storyId`,
301
304
  `schemaVersion`, `round` and `commitSha` — and hand that merged file to the
302
- gate once, with `--expected-criteria` set to the Story's `acceptance[]` count:
305
+ gate once:
303
306
 
304
307
  ```bash
305
308
  node <main-repo>/.agents/scripts/acceptance-eval.js \
306
- --story <storyId> --verdict <merged-verdict-path> \
307
- --expected-criteria <acceptance[] count>
309
+ --story <storyId> --verdict <merged-verdict-path>
308
310
  ```
309
311
 
310
- The flag is what makes the merge enforceable: `assertCriteriaCoverage` returns
311
- early on the `null` default, so **omitting it leaves the guard inert** and a
312
- single cluster's verdict handed over unmerged scores a fraction of the criteria
313
- and still reports `proceed`. A length mismatch is rejected before scoring and
314
- consumes no round. Calling the gate once per cluster instead spends a round
312
+ The gate reads the Story's `acceptance[]` count itself (Story #5313), so a
313
+ single cluster's verdict handed over unmerged is rejected before scoring and
314
+ consumes no round; `--expected-criteria` is accepted but redundant. Calling
315
+ the gate once per cluster instead spends a round
315
316
  _per cluster_ — a Story past the cluster ceiling would burn its whole redraft
316
317
  budget on cluster arithmetic — and N concurrent calls race the Story-scoped
317
318
  round ledger. Full per-round mechanics, including the parallel dispatch and the
@@ -342,32 +343,30 @@ node .agents/scripts/update-ticket-state.js --ticket <storyId> --state agent::bl
342
343
 
343
344
  ## Step 2 — Ceremony detail
344
345
 
345
- **Compute the change set once** with the shared enumerator —
346
- the same module close uses — and reuse that one list downstream:
346
+ **Compute the change set once** with the shared enumerator — the same module
347
+ close uses — and reuse that one list downstream. `ceremony-derive.js`
348
+ (Story #5313) is that enumeration, the level derivation and the ceremony
349
+ resolution in one call:
347
350
 
348
351
  ```bash
349
- node --input-type=module -e '
350
- import { computeChangeSet } from "<main-repo>/.agents/scripts/lib/orchestration/change-set.js";
351
- const { files } = computeChangeSet({ baseRef: "main", headRef: "story-<storyId>" });
352
- console.log(JSON.stringify(files));
353
- '
352
+ node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
354
353
  ```
355
354
 
356
- Derive the level with
355
+ Its `level` comes from
357
356
  [`deriveChangeLevel`](../../scripts/lib/orchestration/review-depth.js) over
358
357
  the one computed change-set list: a diff touching a sensitive path registered
359
358
  in `.agents/schemas/audit-rules.json` derives `high`, one touching none
360
- derives `low`, and an unenumerable diff (`files === null`) derives `null`.
361
- Hand the **same** list to every acceptance critic you spawn (Step 1a) — a
362
- critic that re-ran its own `git diff` could score against a different set
363
- than the one that routed it.
359
+ derives `low`, and an unenumerable diff (`files: null`) derives `null`.
360
+ Hand the **same** `files` list to every acceptance critic you spawn (Step 1a)
361
+ — a critic that re-ran its own `git diff` could score against a different
362
+ set than the one that routed it.
364
363
 
365
- Resolve fresh-vs-inline acceptance critics per AC-cluster with
364
+ Its `mode` / `verdictOwner` come from
366
365
  [`resolveCeremonyForRisk`](../../scripts/lib/orchestration/ceremony-routing.js)
367
366
  (`minimal` → always inline; `strict` → always fresh; `standard` →
368
- `high`/`null` → `fresh`, `low` → `inline` unless the `freshCriticSampleRate`
369
- floor forces `fresh`). Review depth reads the same derived level via
370
- `review-depth.js` inside close, so the two decisions cannot disagree.
367
+ `high`/`null` → `fresh`, `low` → `inline`). Review depth reads the same
368
+ derived level via `review-depth.js` inside close, so the two decisions cannot
369
+ disagree.
371
370
 
372
371
  **Inline-dispatch override.** When the Story dispatches
373
372
  `inline` (`resolveStoryDispatchMode` → `inline`, which is exactly a
@@ -74,7 +74,7 @@ One branch, one PR to `main`, commits against the inline `acceptance[]` /
74
74
  `## Slicing` rows as **intra-session checkpoints** (reference § Step 1).
75
75
  2. Implement and commit on the Story branch, iterating with quick advisory
76
76
  gates (`typecheck`, `lint`, scoped tests) — the full chain runs in Step 3,
77
- and the **one** creditable full-suite run at Step 2.5.
77
+ and the **one** full-suite run at Step 2.5.
78
78
 
79
79
  ### Step 1a — Bounded acceptance self-eval loop (**required**)
80
80
 
@@ -93,21 +93,18 @@ are reference § Step 2. Hard gates always run in Step 3 — the derived level
93
93
  never disables them; do **not** pre-run the chain here — Step 2.5's credited
94
94
  suite run is the sole exception.
95
95
 
96
- ### Step 2.5 — Push, then the creditable full-suite run, then hand off
96
+ ### Step 2.5 — The one credited suite run, the push, then hand off
97
97
 
98
- **Push first.** After the self-eval loop's last fix commit, push
99
- `story-<storyId>` to `origin`, confirming the remote ref moved: the capture
100
- below is backgrounded, so *its* completion ends the turn.
98
+ After the self-eval loop's last fix commit, run the suite **once** in the
99
+ worktree through the depositor — `evidence-gate.js … --gate test -- npm test`,
100
+ spelled out in **digest § 5**. It runs whatever `npm test` resolves to and
101
+ stamps that, so the `test` credit is earned on any runner and only a *later*
102
+ commit invalidates it. Red → fix, commit, re-run.
101
103
 
102
- Then run the full suite **once**, after the push, in the shape close credits
103
- (**digest § 5**): the credit is keyed on the tree, not on push state, so only
104
- a *later* commit invalidates it; a bare `npm test` deposits none. Red →
105
- fix, commit, push, re-capture.
106
-
107
- Then (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
104
+ Push `story-<storyId>` to `origin`, confirming the remote ref moved. Then
105
+ (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
108
106
  branch, pushed head SHA, self-eval verdict, `verify[]` evidence — and stop.
109
- Do not open the PR or compose a terminal envelope. An inline run captures
110
- before Step 3.
107
+ Do not open the PR or compose a terminal envelope.
111
108
 
112
109
  ## Step 3 — Close and land (`single-story-close.js`)
113
110