mandrel 2.54.0 → 2.55.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (114) hide show
  1. package/.agents/agents/story-worker.md +24 -23
  2. package/.agents/audit-checklists/accessibility.md +0 -3
  3. package/.agents/audit-checklists/mobile.md +0 -4
  4. package/.agents/docs/agentrc-reference.json +4 -2
  5. package/.agents/docs/configuration.md +2 -0
  6. package/.agents/schemas/agentrc.schema.json +15 -1
  7. package/.agents/schemas/lifecycle/merge.unlanded.schema.json +2 -1
  8. package/.agents/schemas/story-deliver-terminal.schema.json +1 -0
  9. package/.agents/scripts/audit-to-stories.js +158 -7
  10. package/.agents/scripts/check-audit-attribution.js +119 -62
  11. package/.agents/scripts/check-test-portability.js +512 -0
  12. package/.agents/scripts/coverage-capture.js +17 -10
  13. package/.agents/scripts/evidence-gate.js +31 -4
  14. package/.agents/scripts/generate-workflows-doc.js +65 -14
  15. package/.agents/scripts/git-cleanup.js +4 -0
  16. package/.agents/scripts/lib/ITicketingProvider.js +78 -0
  17. package/.agents/scripts/lib/audit-advisories.js +195 -0
  18. package/.agents/scripts/lib/audit-attribution.js +22 -0
  19. package/.agents/scripts/lib/audit-to-stories/dedupe-against-github.js +68 -5
  20. package/.agents/scripts/lib/audit-to-stories/issue-index.js +83 -0
  21. package/.agents/scripts/lib/audit-to-stories/ledger-commit.js +60 -114
  22. package/.agents/scripts/lib/audit-to-stories/ledger-pr.js +347 -0
  23. package/.agents/scripts/lib/audit-to-stories/parse-audit-md.js +169 -44
  24. package/.agents/scripts/lib/baselines/merge-envelopes.js +298 -32
  25. package/.agents/scripts/lib/bootstrap/baseline-merge-driver.js +180 -14
  26. package/.agents/scripts/lib/cli-args.js +26 -0
  27. package/.agents/scripts/lib/close-validation/gates.js +113 -7
  28. package/.agents/scripts/lib/close-validation/process.js +7 -3
  29. package/.agents/scripts/lib/close-validation/runner.js +62 -11
  30. package/.agents/scripts/lib/config/ci.js +28 -9
  31. package/.agents/scripts/lib/config-settings-schema-delivery.js +7 -0
  32. package/.agents/scripts/lib/config-settings-schema.js +19 -1
  33. package/.agents/scripts/lib/coverage-capture-fullscope.js +23 -11
  34. package/.agents/scripts/lib/coverage-capture-incremental.js +22 -16
  35. package/.agents/scripts/lib/coverage-capture-usage.js +5 -1
  36. package/.agents/scripts/lib/coverage-capture.js +77 -3
  37. package/.agents/scripts/lib/findings/route-finding.js +4 -2
  38. package/.agents/scripts/lib/full-suite-lock.js +232 -6
  39. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  40. package/.agents/scripts/lib/git/sync-from-base.js +130 -13
  41. package/.agents/scripts/lib/observability/source-classifier.js +1 -0
  42. package/.agents/scripts/lib/orchestration/check-baselines/phases/compare.js +10 -2
  43. package/.agents/scripts/lib/orchestration/check-baselines/phases/refresh-ack.js +75 -15
  44. package/.agents/scripts/lib/orchestration/deliver-recover.js +82 -43
  45. package/.agents/scripts/lib/orchestration/dependency-candidates.js +8 -4
  46. package/.agents/scripts/lib/orchestration/epic-candidates.js +9 -4
  47. package/.agents/scripts/lib/orchestration/epic-container.js +66 -4
  48. package/.agents/scripts/lib/orchestration/epic-rollup.js +233 -84
  49. package/.agents/scripts/lib/orchestration/file-assumptions.js +218 -16
  50. package/.agents/scripts/lib/orchestration/git-cleanup/phases/branches.js +93 -7
  51. package/.agents/scripts/lib/orchestration/git-cleanup/phases/git-probes.js +22 -6
  52. package/.agents/scripts/lib/orchestration/git-cleanup/phases/parse-args.js +26 -5
  53. package/.agents/scripts/lib/orchestration/git-cleanup/phases/phase-drivers.js +13 -2
  54. package/.agents/scripts/lib/orchestration/git-cleanup/phases/render.js +35 -5
  55. package/.agents/scripts/lib/orchestration/merge-block-class.js +18 -3
  56. package/.agents/scripts/lib/orchestration/merge-poll.js +284 -40
  57. package/.agents/scripts/lib/orchestration/plan-persist/epic-adoption.js +49 -2
  58. package/.agents/scripts/lib/orchestration/plan-persist/epic-ops.js +43 -7
  59. package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +24 -1
  60. package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +5 -0
  61. package/.agents/scripts/lib/orchestration/plan-persist/summary.js +3 -0
  62. package/.agents/scripts/lib/orchestration/plan-persist/supersede-ops.js +63 -0
  63. package/.agents/scripts/lib/orchestration/plan-persist/wave-serialisation.js +110 -0
  64. package/.agents/scripts/lib/orchestration/planning/memory-pool-advisory.js +130 -40
  65. package/.agents/scripts/lib/orchestration/resolve-stories.js +44 -1
  66. package/.agents/scripts/lib/orchestration/review-providers/native.js +31 -11
  67. package/.agents/scripts/lib/orchestration/review-providers/scoped-lint.js +27 -24
  68. package/.agents/scripts/lib/orchestration/run-epilogue.js +59 -38
  69. package/.agents/scripts/lib/orchestration/single-story-close/close-note.js +81 -0
  70. package/.agents/scripts/lib/orchestration/single-story-close/failed-terminal.js +40 -51
  71. package/.agents/scripts/lib/orchestration/single-story-close/phases/auto-merge.js +10 -2
  72. package/.agents/scripts/lib/orchestration/single-story-close/phases/base-sync.js +101 -0
  73. package/.agents/scripts/lib/orchestration/single-story-close/phases/confirm-merge.js +351 -28
  74. package/.agents/scripts/lib/orchestration/single-story-close/phases/options.js +27 -6
  75. package/.agents/scripts/lib/orchestration/single-story-close/runner.js +117 -22
  76. package/.agents/scripts/lib/orchestration/story-close/baseline-upward-writeback.js +94 -12
  77. package/.agents/scripts/lib/orchestration/story-close/format-autofix.js +6 -1
  78. package/.agents/scripts/lib/orchestration/ticket-validator.js +25 -14
  79. package/.agents/scripts/lib/orchestration/ticketing/bulk.js +30 -0
  80. package/.agents/scripts/lib/orchestration/verify-credit.js +37 -0
  81. package/.agents/scripts/lib/pinned-override-notes.js +41 -53
  82. package/.agents/scripts/lib/pinned-override-resolve.js +212 -0
  83. package/.agents/scripts/lib/qa/resolve-qa-contract.js +18 -0
  84. package/.agents/scripts/lib/single-story-sweep/sweep-lock.js +173 -9
  85. package/.agents/scripts/lib/skills/walk-skill-files.js +24 -7
  86. package/.agents/scripts/lib/test-temp.js +167 -30
  87. package/.agents/scripts/lib/validation-evidence.js +37 -0
  88. package/.agents/scripts/lib/wave-runner/footprint.js +167 -14
  89. package/.agents/scripts/lib/wave-runner/live-probe.js +7 -1
  90. package/.agents/scripts/lib/wave-runner/ready-set.js +1 -1
  91. package/.agents/scripts/merge-baseline.js +175 -21
  92. package/.agents/scripts/providers/github/errors.js +22 -1
  93. package/.agents/scripts/providers/github/issues.js +106 -1
  94. package/.agents/scripts/providers/github/sub-issue-add.js +18 -1
  95. package/.agents/scripts/providers/github.js +6 -0
  96. package/.agents/scripts/resolve-stories.js +44 -34
  97. package/.agents/scripts/single-story-close.js +5 -0
  98. package/.agents/scripts/stories-wave-tick.js +37 -13
  99. package/.agents/templates/docs/audit-sweep-runbook.md +41 -7
  100. package/.agents/workflows/audit-accessibility.md +16 -31
  101. package/.agents/workflows/audit-mobile.md +20 -37
  102. package/.agents/workflows/git-cleanup.md +17 -3
  103. package/.agents/workflows/helpers/audit-lens-core.md +45 -0
  104. package/.agents/workflows/helpers/deliver-digest.md +7 -6
  105. package/.agents/workflows/helpers/deliver-reference.md +35 -14
  106. package/.agents/workflows/helpers/deliver-story-reference.md +7 -4
  107. package/.agents/workflows/helpers/deliver-story.md +15 -12
  108. package/.agents/workflows/helpers/plan-reference.md +7 -0
  109. package/.agents/workflows/mandrel-plan.md +4 -7
  110. package/.agents/workflows/memory-consolidate.md +14 -9
  111. package/docs/CHANGELOG.md +27 -0
  112. package/lib/cli/registry.js +64 -21
  113. package/lib/cli/sync.js +27 -2
  114. package/package.json +7 -4
@@ -55,8 +55,8 @@
55
55
  * inFlight: number,
56
56
  * cycleError: string | null,
57
57
  * wedged: { reason, stories: [{ id, unmetBlockers }] } | null,
58
- * inFlightReservation: { available, withheld: [{ id, blockedBy, reason, source, paths }], note },
59
- * footprintGuard: { mode, withheld: [{ id, blockedBy, scope, source, paths }], advisory, note }
58
+ * inFlightReservation: { available, withheld: [{ id, blockedBy, reason, source, paths, attribution }], note },
59
+ * footprintGuard: { mode, withheld: [{ id, blockedBy, scope, source, paths, attribution }], advisory, note }
60
60
  * }
61
61
  *
62
62
  * `inFlightReservation` reports the cross-beat half of the co-dispatch guard
@@ -130,7 +130,10 @@ import { AGENT_LABELS } from './lib/label-constants.js';
130
130
  import { parseIds } from './lib/orchestration/resolve-stories.js';
131
131
  import { buildStoryAdjacency } from './lib/story-adjacency.js';
132
132
  import { expandIdList } from './lib/util/parse-id-list.js';
133
- import { OVERLAP_SOURCES } from './lib/wave-runner/footprint.js';
133
+ import {
134
+ OVERLAP_SOURCES,
135
+ renderScrapeAttribution,
136
+ } from './lib/wave-runner/footprint.js';
134
137
  import {
135
138
  createProbeContext,
136
139
  probeLiveState,
@@ -237,7 +240,10 @@ Output envelope:
237
240
  "blockedBy": 4949,
238
241
  "reason": "in-flight-earlier-beat",
239
242
  "source": "declared-overlap",
240
- "paths": ["lib/shared.js"]
243
+ "paths": ["lib/shared.js"],
244
+ "attribution": [
245
+ { "path": "lib/shared.js", "declared": true, "fields": [] }
246
+ ]
241
247
  }
242
248
  ],
243
249
  "note": "..."
@@ -250,7 +256,14 @@ Output envelope:
250
256
  "blockedBy": 4951,
251
257
  "scope": "beat",
252
258
  "source": "scraped-overlap",
253
- "paths": ["lib/other.js"]
259
+ "paths": ["lib/other.js"],
260
+ "attribution": [
261
+ {
262
+ "path": "lib/other.js",
263
+ "declared": false,
264
+ "fields": ["body:Verify"]
265
+ }
266
+ ]
254
267
  }
255
268
  ],
256
269
  "advisory": [],
@@ -268,7 +281,10 @@ footprintGuard names each Story withheld from THIS beat by a peer already
268
281
  admitted on it — the half that used to be an unreported skip — and every
269
282
  entry in either report carries the colliding paths plus a source tag
270
283
  (declared-overlap when both changes[] declarations named the path, else
271
- scraped-overlap from the text evidence). Its "mode" echoes
284
+ scraped-overlap from the text evidence) and an "attribution" list naming, per
285
+ path, the field the scrape read it from ("title", "spec", or "body:<section>"
286
+ — so a path that reached the comparison only because every Story RUNS it in
287
+ "## Verify" says so). Its "mode" echoes
272
288
  delivery.deliverRunner.footprintGuard: under "advisory" the collisions are
273
289
  detected and listed in "advisory" but never withhold, and dispatch follows the
274
290
  declared depends_on edges alone.
@@ -359,7 +375,7 @@ const RESERVATION_REASONS = Object.freeze({
359
375
  * @param {object[]|null|undefined} inFlightRecords
360
376
  * @param {Array<{id: number, blockedBy: number, source?: string, paths?: string[]}>} withheld
361
377
  * @param {Iterable<number>} [foreignHeldIds] Ids held by a foreign lease.
362
- * @returns {{ available: boolean, withheld: Array<{id: number, blockedBy: number, reason: string, source: string, paths: string[]}>, note: string|null }}
378
+ * @returns {{ available: boolean, withheld: Array<{id: number, blockedBy: number, reason: string, source: string, paths: string[], attribution: object[]}>, note: string|null }}
363
379
  */
364
380
  export function buildReservationReport(
365
381
  inFlightRecords,
@@ -390,6 +406,7 @@ export function buildReservationReport(
390
406
  : RESERVATION_REASONS.EARLIER_BEAT,
391
407
  source: w.source ?? OVERLAP_SOURCES.DECLARED,
392
408
  paths: w.paths ?? [],
409
+ attribution: w.attribution ?? [],
393
410
  }));
394
411
  return {
395
412
  available: true,
@@ -421,12 +438,16 @@ export function buildReservationReport(
421
438
  */
422
439
  export function buildFootprintGuardReport(footprintWithholds, mode) {
423
440
  const ledger = Array.isArray(footprintWithholds) ? footprintWithholds : [];
424
- const project = ({ id, blockedBy, scope, source, paths }) => ({
441
+ const project = ({ id, blockedBy, scope, source, paths, attribution }) => ({
425
442
  id,
426
443
  blockedBy,
427
444
  scope,
428
445
  source,
429
446
  paths,
447
+ // Story #5265: the per-path field attribution rides the entry itself, so
448
+ // a consumer reading the envelope never has to re-derive where a scraped
449
+ // path came from (and cannot get a different answer than the note did).
450
+ attribution: attribution ?? [],
430
451
  });
431
452
  const beat = ledger
432
453
  .filter((w) => w.scope === WITHHOLD_SCOPES.BEAT && w.enforced)
@@ -453,10 +474,11 @@ export function buildFootprintGuardReport(footprintWithholds, mode) {
453
474
  function footprintGuardNote(beat, advisory, mode) {
454
475
  const detail = (entries) =>
455
476
  entries
456
- .map(
457
- (w) =>
458
- `#${w.id} ← #${w.blockedBy} on ${w.paths.join(', ')} (${w.source})`,
459
- )
477
+ .map((w) => {
478
+ const scraped = renderScrapeAttribution(w.attribution);
479
+ const provenance = scraped ? `, scraped from ${scraped}` : '';
480
+ return `#${w.id} ← #${w.blockedBy} on ${w.paths.join(', ')} (${w.source}${provenance})`;
481
+ })
460
482
  .join('; ');
461
483
  if (beat.length > 0) {
462
484
  return (
@@ -465,7 +487,9 @@ function footprintGuardNote(beat, advisory, mode) {
465
487
  `Each is still eligible and re-admits on a later beat once its peer ` +
466
488
  `lands. A ${OVERLAP_SOURCES.SCRAPED} source means the collision came ` +
467
489
  `from path evidence in the Story text rather than from either ` +
468
- `changes[] declaration.`
490
+ `changes[] declaration — the 'scraped from' clause names the field ` +
491
+ `each such path was read out of, so a path only cited in '## Verify' ` +
492
+ `is distinguishable from an unpredicted edit target.`
469
493
  );
470
494
  }
471
495
  if (advisory.length > 0) {
@@ -51,6 +51,17 @@ to that declared tally. A mismatch — or a missing tally line — means the rep
51
51
  is not trustworthy: a finding was malformed, a severity did not resolve onto the
52
52
  canonical scale, or the lens truncated its own output.
53
53
 
54
+ The line is read **from the Executive Summary**, and a report declaring it more
55
+ than once fails as `duplicate-tally` naming both lines rather than adopting
56
+ whichever the scan reached first. Prose elsewhere in the report that quotes the
57
+ tally format is a second declaration as far as the check is concerned; move it
58
+ or reword it.
59
+
60
+ A `###` heading with no severity axis and no field bullets is read as a
61
+ **grouping header**, not as a finding — so a lens that emits `### Robust` with
62
+ `_No findings._` under it contributes zero findings and still cross-checks
63
+ against a zero tally.
64
+
54
65
  `--auto` **fails closed** on any such failure. It exits non-zero having opened
55
66
  no Issue and written no ledger, and names the offending report in
56
67
  `summary.reportFailures[]`. `--allow-missing-tally` is a `--scan` affordance
@@ -70,6 +81,10 @@ stop surprising you:
70
81
  node .agents/scripts/audit-to-stories.js --auto --dry-run
71
82
  ```
72
83
 
84
+ `--severity` is validated against the canonical scale: a typo (`--severity Hgh`)
85
+ exits non-zero naming the accepted levels rather than silently widening the run
86
+ to every finding.
87
+
73
88
  `--dry-run` performs zero GitHub writes and skips the ledger write, printing
74
89
  only the run summary. Read `totals.create` before you let the sweep file
75
90
  anything: a first full-scope run over an un-audited repository can propose more
@@ -94,11 +109,24 @@ rejected.
94
109
  `--ledger-commit` closes that loop. After the run summary has printed, and only
95
110
  when the ledger actually changed, it:
96
111
 
97
- 1. creates `chore/audit-ledger-<YYYY-MM-DD>` from the current HEAD,
98
- 2. commits **only** the ledger file, subject
112
+ 1. refuses, naming its step, if HEAD is not the base branch or there is no
113
+ `origin` — before writing anything,
114
+ 2. creates `chore/audit-ledger-<YYYY-MM-DD>-<shortsha>` from `origin/<base>`,
115
+ 3. commits **only** the ledger file, subject
99
116
  `chore(audit): reconcile audit ledger <date>`,
100
- 3. pushes the branch, and
101
- 4. opens a PR against your base branch.
117
+ 4. pushes the branch,
118
+ 5. opens a PR against your base branch, and
119
+ 6. puts the checkout back on the branch it started from.
120
+
121
+ It prints one line on stderr saying what happened: the branch and the PR URL on
122
+ success, or the skip reason otherwise.
123
+
124
+ **It is re-runnable.** The `<shortsha>` is the base commit, so a retry on the
125
+ same day does not collide with the branch a failed run left behind — it
126
+ recognises it. A ledger already committed on an **unpushed** ledger branch
127
+ resumes at the push rather than reporting `ledger-unchanged` (the ledger file is
128
+ clean: it is committed, just not pushed). Across a failed push and its retry you
129
+ get exactly one PR.
102
130
 
103
131
  **Auto-merge is never requested.** The ledger records machine-derived lifecycle
104
132
  state, so a human glance before it lands is the point — nominate that reviewer
@@ -123,8 +151,10 @@ Those bodies are **audit prose**, not delivery-ready Specs: they describe a
123
151
  symptom and a recommendation, not a scoped change with acceptance criteria a
124
152
  worker can verify against.
125
153
 
126
- Do not point `/mandrel-deliver` at a freshly-filed audit Story. Route it through
127
- `/mandrel-plan` first — the planning pass is where the finding becomes a
154
+ Do not point `/mandrel-deliver` at a freshly-filed audit Story — and as of the
155
+ label guard below, you cannot: `resolve-stories.js` refuses a Story carrying no
156
+ `agent::*` label, naming this step, unless `--allow-unlabelled` is passed. Route
157
+ it through `/mandrel-plan` first — the planning pass is where the finding becomes a
128
158
  capability slice with a `## Spec`, real `acceptance[]` items and runnable
129
159
  `verify[]` lines. Planning is deliberately not automated here: deciding what a
130
160
  finding is worth, and how far the fix should reach, is the judgement the sweep
@@ -172,6 +202,10 @@ to mint labels no taxonomy defines.
172
202
  | --- | --- | --- |
173
203
  | Non-zero exit, `summary.reportFailures[]` populated | A lens report's tally is missing or disagrees with the parse | Re-run that lens; never downgrade with `--allow-missing-tally` |
174
204
  | `ledger.unpersisted: true` in the summary | No `origin`, or HEAD off the base branch | Re-run with `--ledger-commit`, or commit the ledger by hand |
175
- | `--ledger-commit failed at step "..."` | git or `gh` failed at the named step | Fix the remote/auth and re-run; the summary above it is still valid |
205
+ | `--ledger-commit failed at step "..."` | git or `gh` failed at the named step | Fix the cause and re-run; a retry resumes rather than duplicating, and the summary above it is still valid |
206
+ | `--ledger-commit failed at step "verify-base-branch"` | HEAD is on a feature branch | Check out the base branch and re-run; nothing was committed |
207
+ | `[resolve-stories] ... carries no "agent::*" label` | An unenriched audit Story was named for delivery | Plan it (Step 5), or pass `--allow-unlabelled` deliberately |
208
+ | `[audit-to-stories] --severity "..." is not a severity` | A typo in the floor | Use one of the canonical levels |
209
+ | `duplicate-tally` in `summary.reportFailures[]` | The report declares the tally twice | Reword the prose copy; the Executive Summary line is the declaration |
176
210
  | Same findings re-proposed every cycle | The ledger is not being committed | Adopt Step 4 |
177
211
  | `totals.create` far larger than the team can absorb | Severity floor too low for a first full-scope run | Raise `--severity` and re-dry-run |
@@ -82,9 +82,9 @@ the project is configured.** Before any detection:
82
82
  - **Design tokens:** the colour tokens (`tailwind.config.*`, CSS custom
83
83
  properties, a theme object) whose literal values you need to compute contrast
84
84
  ratios statically.
85
- - **Runtime target (optional):** the `qa.environments` map (see
86
- [_Runtime verification mode_](#step-2-runtime-verification-mode-optional-corroboration))
87
- and the navigability route SSOT.
85
+ - **Runtime target (optional):** whatever the shared runtime scaffold in
86
+ [`helpers/audit-lens-core.md`](helpers/audit-lens-core.md#runtime-pass)
87
+ resolves. Note only whether one exists; the scaffold owns how it is found.
88
88
 
89
89
  Record what exists. Every finding downstream is measured against _this
90
90
  discovered surface and config_, not a generic ideal. If **no** frontend surface
@@ -135,31 +135,16 @@ Cover every static WCAG dimension:
135
135
 
136
136
  ## Step 2: Runtime verification mode (optional corroboration)
137
137
 
138
- Static detection is the default and always runs. The runtime pass is
139
- **conditional** — it runs only when a live target is configured; its absence
140
- never blocks the static report.
141
-
142
- 1. **Resolve the target from config — never a hardcoded URL.** Resolve the
143
- target through the consumer's `qa.environments.<env>.baseUrl` (via
144
- [`resolveQaEnvironment`](../scripts/lib/qa/resolve-qa-contract.js), the same
145
- resolver `/qa-run` uses): an `<env>` argument resolves by exact name or
146
- origin match; with no argument, enumerate `name → baseUrl` and let the
147
- operator pick. If **no** `qa.environments` target is configured, **skip this
148
- step** and note in the report that runtime corroboration was unavailable —
149
- do not invent a URL and do not start an arbitrary dev server.
150
- 2. **Sample routes from the navigability SSOT.** Draw the routes to exercise
151
- from the consumer's route/nav registry (`planning.navigation.navRegistry` /
152
- `routeGlobs` — the same SSOT [`/audit-navigability`](audit-navigability.md)
153
- reads), sampling a representative set (key personas' landing routes plus any
154
- route in the change-set scope) rather than a single hardcoded page.
155
- 3. **Run an accessibility engine per sampled route.** Use the
156
- `mcp__chrome-devtools__lighthouse_audit` tool's **Accessibility category**,
157
- or run **axe** via the browser tooling, against each sampled `baseUrl`-rooted
158
- route. Prefer a production-mode build.
159
- 4. **Median-of-3 or provisional.** Any runtime score or metric is subject to
160
- run-to-run variance: capture a **median-of-3** (three runs per route, report
161
- the median) before treating a number as authoritative. A single-run value is
162
- reported **provisional** and never drives a Critical/High verdict on its own.
138
+ Run the shared runtime scaffold in
139
+ [`helpers/audit-lens-core.md`](helpers/audit-lens-core.md#runtime-pass) — it
140
+ owns target resolution, route sampling, median-of-3, and the two skip reasons
141
+ (no configured target; browser tooling unavailable). This lens's probe is its
142
+ step 3:
143
+
144
+ - **Run an accessibility engine per sampled route.** Use the
145
+ `mcp__chrome-devtools__lighthouse_audit` tool's **Accessibility category**, or
146
+ run **axe** via the browser tooling, against each sampled `baseUrl`-rooted
147
+ route. Prefer a production-mode build.
163
148
 
164
149
  Corroborate static findings against the runtime results (a statically-flagged
165
150
  contrast defect confirmed by the engine graduates from provisional to
@@ -168,8 +153,8 @@ confirmed), and surface runtime-only violations the static pass could not see.
168
153
  ## Report additions
169
154
 
170
155
  Beyond the shared skeleton, the Executive Summary states the runtime mode's
171
- status (ran against `<env>` / skipped — no target configured), and the report
156
+ status (ran against `<env>`, or the scaffold's skip reason), and the report
172
157
  ends with a **Runtime Verification** section: per-route median-of-3
173
- accessibility scores when the runtime mode ran, or "_Runtime corroboration
174
- unavailable — no `qa.environments` target configured._" Drop every claimed
158
+ accessibility scores when the runtime mode ran, or that skip reason verbatim.
159
+ Drop every claimed
175
160
  violation that names no concrete element and no specific WCAG success criterion.
@@ -88,9 +88,9 @@ what they declare:
88
88
  explicit `viewport`, a Cypress `viewportWidth`/`viewportHeight`, a
89
89
  visual-regression viewport list. This is the project's own statement of which
90
90
  form factors it holds itself to.
91
- - **Runtime target (optional):** the `qa.environments` map (see
92
- [*Runtime viewport pass*](#step-3-runtime-viewport-pass-optional-corroboration))
93
- and the navigability route SSOT.
91
+ - **Runtime target (optional):** whatever the shared runtime scaffold in
92
+ [`helpers/audit-lens-core.md`](helpers/audit-lens-core.md#runtime-pass)
93
+ resolves. Note only whether one exists; the scaffold owns how it is found.
94
94
 
95
95
  Record the breakpoints, the viewport contract, and the device matrix. Every
96
96
  finding downstream is measured against *this discovered baseline*. If the
@@ -192,37 +192,20 @@ fix is a new test, name the viewport and the assertion it should make, not just
192
192
 
193
193
  ## Step 3: Runtime viewport pass (optional corroboration)
194
194
 
195
- Static detection is the default and always runs. The runtime pass is
196
- **conditional** — it runs only when a live target is configured; its absence
197
- never blocks the static report.
198
-
199
- 1. **Resolve the target from config — never a hardcoded URL.** Resolve the
200
- target through the consumer's `qa.environments.<env>.baseUrl` (via
201
- [`resolveQaEnvironment`](../scripts/lib/qa/resolve-qa-contract.js), the same
202
- resolver `/qa-run` uses): an `<env>` argument resolves by exact name or
203
- origin match; with no argument, enumerate `name → baseUrl` and let the
204
- operator pick. If **no** `qa.environments` target is configured, **skip this
205
- step** and note in the report that runtime corroboration was unavailable —
206
- do not invent a URL and do not start an arbitrary dev server.
207
- 2. **Sample routes from the navigability SSOT.** Draw the routes to exercise
208
- from the consumer's route/nav registry (`planning.navigation.navRegistry` /
209
- `routeGlobs` — the same SSOT [`/audit-navigability`](audit-navigability.md)
210
- reads), sampling a representative set (key personas' landing routes plus any
211
- route in the change-set scope) rather than a single hardcoded page.
212
- 3. **Drive two form factors per route.** Emulate a **phone** and a **tablet**
213
- viewport — `mcp__chrome-devtools__emulate` for a device profile, or
214
- `resize_page` for an explicit width/height — then, per route and viewport:
215
- take a screenshot, and evaluate the two observables static analysis cannot
216
- resolve — whether `document.scrollingElement.scrollWidth` exceeds the
217
- viewport width (horizontal overflow), and the rendered box of the
218
- interactive controls Step 1 flagged as candidates. Reload after switching
219
- form factor so load-time device gates re-run.
220
- 4. **Median-of-3 or provisional.** Any runtime measurement is subject to
221
- run-to-run variance: capture a **median-of-3** (three runs per route, report
222
- the median) before treating a number as authoritative. A single-run value is
223
- reported **provisional** and never drives a Critical/High verdict on its own.
224
- 5. **Leave the viewport as you found it.** Reset the emulation before finishing
225
- so a following lens or QA run does not inherit a phone viewport.
195
+ Run the shared runtime scaffold in
196
+ [`helpers/audit-lens-core.md`](helpers/audit-lens-core.md#runtime-pass) — it
197
+ owns target resolution, route sampling, median-of-3, and the two skip reasons
198
+ (no configured target; browser tooling unavailable). This lens's probe is its
199
+ step 3:
200
+
201
+ - **Drive two form factors per route.** Emulate a **phone** and a **tablet**
202
+ viewport — `mcp__chrome-devtools__emulate` for a device profile, or
203
+ `resize_page` for an explicit width/height — then, per route and viewport:
204
+ take a screenshot, and evaluate the two observables static analysis cannot
205
+ resolve — whether `document.scrollingElement.scrollWidth` exceeds the viewport
206
+ width (horizontal overflow), and the rendered box of the interactive controls
207
+ Step 1 flagged as candidates. Reload after switching form factor so load-time
208
+ device gates re-run.
226
209
 
227
210
  Corroborate static findings against the runtime observations (a statically
228
211
  flagged fixed width confirmed by a real horizontal overflow graduates from
@@ -233,10 +216,10 @@ that opens off-screen.
233
216
  ## Report additions
234
217
 
235
218
  Beyond the shared skeleton, the Executive Summary states the runtime mode's
236
- status (ran against `<env>` / skipped — no target configured) and names the
219
+ status (ran against `<env>`, or the scaffold's skip reason) and names the
237
220
  narrowest breakpoint the project declares, so a reader can tell what "mobile"
238
221
  meant for this run. The report ends with a **Runtime Viewport Pass** section:
239
- per-route, per-form-factor observations when the runtime mode ran, or
240
- "*Runtime corroboration unavailable — no `qa.environments` target configured.*"
222
+ per-route, per-form-factor observations when the runtime mode ran, or the
223
+ scaffold's skip reason verbatim.
241
224
  Drop every claimed finding that names no concrete element, style rule, or test
242
225
  file.
@@ -5,7 +5,7 @@ description: >-
5
5
  `git stash` entries — each step gated by operator confirmation.
6
6
  ---
7
7
 
8
- # /git-cleanup [--fast-forward-main] [--prune-remotes] [--branches] [--stashes] [--execute] [--remote] [--yes] [--drop-stashes <ref>] [--exclude <pattern>] [--json]
8
+ # /git-cleanup [--fast-forward-main] [--prune-remotes] [--branches] [--stashes] [--execute] [--remote] [--yes] [--include-content-merged] [--drop-stashes <ref>] [--exclude <pattern>] [--json]
9
9
 
10
10
  `/git-cleanup` folds the four cleanup steps operators routinely run by hand
11
11
  after a busy session into a single pipeline with per-step confirmation. It is a
@@ -36,7 +36,7 @@ flags: `node .agents/scripts/git-cleanup.js --help`.
36
36
  | --- | --- | --- |
37
37
  | **fast-forward-main** | `git fetch origin <base>` then `git merge --ff-only origin/<base>`. | Skipped silently on a dirty tree or a non-fast-forward; otherwise prompts `Fast-forward main by N commit(s)?`. Checks out `<base>` first when HEAD is elsewhere and does **not** restore the prior branch. |
38
38
  | **prune-remotes** | `git fetch --prune origin` to drop `refs/remotes/origin/*` GitHub already deleted. | Prompts before pruning. Runs as its own phase regardless of `--remote`. |
39
- | **branches** | Reaps merged local branches (squash-aware: merged-PR, git-ancestry, and content-equivalence signals), removing an attached worktree first. Also enumerates **remote-only** merged branches. | Prints the candidate list, then prompts `Reap N merged branch(es)?`. `--remote` is required **on top of** `--execute` to delete any `origin/<branch>`. `content-merged` candidates carry a weaker-signal warning. |
39
+ | **branches** | Reaps merged local branches (squash-aware: merged-PR, git-ancestry, and content-equivalence signals), removing an attached worktree first. Also enumerates **remote-only** merged branches. | Prints the candidate list, then prompts `Reap N merged branch(es)?`. `--remote` is required **on top of** `--execute` to delete any `origin/<branch>`. `content-merged` candidates carry a weaker-signal warning, and under `--yes` their **remote** ref is withheld unless `--include-content-merged` is passed (their local ref still goes — it is recoverable from the remote). |
40
40
  | **stashes** | Lists every stash and triages it. | Interactive: `drop / keep / quit` per entry (default `keep`). Under `--yes` / `--json`, drops require an explicit `--drop-stashes <ref>` allowlist (repeatable) — there is no "drop all". |
41
41
 
42
42
  ## Constraint
@@ -49,6 +49,14 @@ Two consequences the flag list alone does not carry: `--remote` deletions cannot
49
49
  be undone without re-pushing, and `--exclude '<pattern>'` is the **only** way to
50
50
  protect an in-scope merged-PR branch you want to keep.
51
51
 
52
+ That irreversibility is why `--yes` and `content-merged` do not combine on their
53
+ own. Content-equivalence says "applying this branch to the base changes nothing"
54
+ — which is true of a squash-merged branch and equally true of one whose every
55
+ change was reverted. An operator answering the prompt sees the weaker-signal
56
+ note and decides; an unattended run has nobody to decide, so it withholds the
57
+ remote delete and says so, and `--include-content-merged` is the decision made
58
+ in advance.
59
+
52
60
  Do **not** run with `--execute` if there is unmerged work that needs saving. The
53
61
  fast-forward phase skips on a dirty tree (safe), but the branches phase reaps any
54
62
  merged-PR branch in scope unless `--exclude`d.
@@ -59,9 +67,15 @@ merged-PR branch in scope unless `--exclude`d.
59
67
  # Preview all four phases (no mutation).
60
68
  node .agents/scripts/git-cleanup.js
61
69
 
62
- # Run everything non-interactively, including origin refs.
70
+ # Run everything non-interactively, including origin refs. Branches detected
71
+ # only by content-equivalence keep their origin ref — see the note below.
63
72
  node .agents/scripts/git-cleanup.js --execute --remote --yes
64
73
 
74
+ # Same, but also delete the origin refs of content-merged branches. Nobody is
75
+ # watching, so opting in is the whole confirmation this delete ever gets.
76
+ node .agents/scripts/git-cleanup.js --execute --remote --yes \
77
+ --include-content-merged
78
+
65
79
  # Only fast-forward main.
66
80
  node .agents/scripts/git-cleanup.js --fast-forward-main --execute
67
81
 
@@ -208,6 +208,51 @@ extracted and names any disagreement as a **report failure**
208
208
  report whose line is missing or wrong. A parse that silently drops findings is
209
209
  otherwise indistinguishable from a clean audit.
210
210
 
211
+ ## Runtime pass scaffold {#runtime-pass}
212
+
213
+ Some lenses corroborate their static findings against a live target. Static
214
+ detection is the default and always runs; the runtime pass is **conditional** —
215
+ it runs only when a live target and the browser tooling are both available, and
216
+ its absence never blocks the static report. The steps below are shared by every
217
+ lens that has one, so a lens documents only its own probe.
218
+
219
+ 1. **Resolve the target from config — never a hardcoded URL.** Resolve through
220
+ the consumer's `qa.environments.<env>.baseUrl` (via
221
+ [`resolveQaEnvironment`](../../scripts/lib/qa/resolve-qa-contract.js), the
222
+ same resolver `/qa-run` uses): an `<env>` argument resolves by exact name or
223
+ origin match; with no argument, enumerate `name → baseUrl` and let the
224
+ operator pick.
225
+ 2. **Sample routes from the navigability SSOT.** Draw the routes to exercise
226
+ from the consumer's route/nav registry (`planning.navigation.navRegistry` /
227
+ `routeGlobs` — the same SSOT
228
+ [`/audit-navigability`](../audit-navigability.md) reads), sampling a
229
+ representative set (key personas' landing routes plus any route in the
230
+ change-set scope) rather than a single hardcoded page.
231
+ 3. **Run the lens's own probe** against each sampled route. That step, and only
232
+ that step, is the lens's to document.
233
+ 4. **Median-of-3 or provisional.** Any runtime measurement is subject to
234
+ run-to-run variance: capture a **median-of-3** (three runs per route, report
235
+ the median) before treating a number as authoritative. A single-run value is
236
+ reported **provisional** and never drives a Critical/High verdict on its own.
237
+ 5. **Leave the environment as you found it.** Reset any emulation — viewport,
238
+ device profile, colour scheme — before finishing, so a following lens or QA
239
+ run does not inherit it.
240
+
241
+ **Skip reasons, and how to report them.** The pass is skipped, never faked, for
242
+ either of two reasons, and the Executive Summary says which:
243
+
244
+ | Condition | Reported as |
245
+ | --- | --- |
246
+ | No `qa.environments` target is configured | `skipped — no target configured` |
247
+ | The browser tooling is unavailable in this runtime | `skipped — browser tooling unavailable` |
248
+
249
+ The second is the one an unattended sweep meets: a scheduled run has no browser
250
+ MCP server attached, so a lens that treated "cannot drive" as "nothing found"
251
+ would report a clean runtime section it never ran. Do not invent a URL, do not
252
+ start an arbitrary dev server, and do not fall back to a static-only claim
253
+ dressed as a runtime one — name the skip and let the static findings stand on
254
+ their own.
255
+
211
256
  ## Execution strategy {#execution-strategy}
212
257
 
213
258
  A lens is a self-contained, read-only unit of work — exactly the shape a
@@ -107,11 +107,12 @@ Per-round mechanics: [`acceptance-self-eval.md`](acceptance-self-eval.md).
107
107
 
108
108
  ## 5. The one creditable full-suite run
109
109
 
110
- **After the self-eval loop's last fix commit, immediately before the push** —
111
- the credit is keyed on the tree, so any later commit invalidates it. Redraft
112
- rounds run scoped tests; only this final run needs credit, and a bare
113
- `npm test` / `pnpm run test` deposits **none**, so close re-runs it. Shape it
114
- by what `close-validation/gates.js` runs:
110
+ **After the self-eval loop's last fix commit, immediately after the push** —
111
+ the credit is keyed on the tree, not push state: a later commit voids it, a
112
+ push does not, and the backgrounded capture (below) ends the turn. Redraft
113
+ rounds run scoped tests; only this run needs credit, and a bare `npm test` /
114
+ `pnpm run test` deposits **none**, so close re-runs it. Shape it by what
115
+ `close-validation/gates.js` runs:
115
116
 
116
117
  ```bash
117
118
  # CRAP gate on (default) + a `test:coverage` script — writes close's stamp:
@@ -127,7 +128,7 @@ Dispatch it in the **background**: it outruns the host's sync Bash ceiling, and
127
128
  Rule 2).
128
129
 
129
130
  Read the **output**, not the exit code: capture skips — no test run, no
130
- credit — when nothing changed under the CRAP `targetDirs`. Run the scoped
131
+ credit — when nothing changed under CRAP `targetDirs`. Run the scoped
131
132
  projects for the roots you changed plus `verify[]`, not the whole suite.
132
133
 
133
134
  `verify[]` is scoped entries **plus** this one run: an entry that is itself a
@@ -49,6 +49,13 @@ issue state rather than against anything you hand it. That is why there is no
49
49
  batch label to pass and why a blocker that landed in an unrelated run is simply
50
50
  seen as done.
51
51
 
52
+ **A Story with no `agent::*` label is refused.** The audit sweep files Stories
53
+ deliberately without one — their bodies are audit prose, not a scoped change
54
+ with verifiable acceptance criteria — so resolving one means dispatching a
55
+ worker at an unenriched body, after taking its lease. Route it through
56
+ `/mandrel-plan` first, which applies `agent::ready` at the end of planning.
57
+ `--allow-unlabelled` is the deliberate escape hatch.
58
+
52
59
  **The non-zero exit codes.** **2** — `cycleError`: the graph is
53
60
  self-referential; fix `depends_on`, do not retry. **3** — `wedged`: nothing
54
61
  dispatchable and nothing in flight, with the undone Stories and their unmet
@@ -292,10 +299,10 @@ This executes, in order:
292
299
  (files issues when auto-file is on; posts `follow-ups`).
293
300
  - `sibling-coherence` — Spec/Acceptance coherence check across sibling bodies
294
301
  (`plan-run-sibling-coherence`).
295
- - `epic-close` — **reports** which container Epics this run closed and which
296
- are still pending. It derives nothing itself: every step here and the
297
- per-Story land tail alike delegate to `epic-rollup.js`, so one rule decides
298
- a container's state.
302
+ - `epic-close` — **reports** which of the run's container Epics its land tails
303
+ left closed and which are still open. **Read-only** — it derives nothing:
304
+ every child state change is already a rollup edge, so the container was
305
+ derived from a complete child set by the last Story's own land tail.
299
306
 
300
307
  A single-Story run skips the epilogue — follow-ups are captured on merge
301
308
  confirm instead (`captureStoryFollowUps`).
@@ -304,9 +311,12 @@ confirm instead (`captureStoryFollowUps`).
304
311
 
305
312
  A container Epic is never delivered, so nothing used to write to it during
306
313
  the run it was the subject of. `epic-rollup.js` derives its state from its
307
- children at both per-Story lifecycle edges — the `agent::executing` flip in
308
- `single-story-init.js` and the post-land tail (reported as the tail's
309
- `epicRollup` step) — which is why it holds at **N=1**, where no epilogue runs.
314
+ children at **every edge that changes a child's state** — the
315
+ `agent::executing` flip in `single-story-init.js`, the post-land tail
316
+ (reported as the tail's `epicRollup` step), and `plan-persist`'s supersede
317
+ close (reported as `supersede.epicRollup`) — which is why it holds at **N=1**,
318
+ where no epilogue runs, and why a cohort superseded by a re-plan no longer
319
+ strands its container open above finished work.
310
320
 
311
321
  - **Status** follows the children's composition (`deriveParentState` mapped
312
322
  onto the board's three options): any **open** child executing or blocked →
@@ -319,13 +329,24 @@ children at both per-Story lifecycle edges — the `agent::executing` flip in
319
329
  ready list.
320
330
  - **Owner** — `github.operatorHandle` is added to the Epic while any child is
321
331
  in flight, through the additive assignees endpoint, and is never removed.
322
- - **Closure** is one-way: every child landed closes the container as
323
- `completed`; a reopened child moves Status back to `In Progress` and does
324
- **not** reopen it.
325
- - The parent lookup scans open `type::epic` issues, because linkage is
326
- parent→child only, and reads children as the body checklist **union** the
327
- native sub-issue edges — the same reader `/mandrel-deliver`'s expansion
328
- uses, so an Epic can never be expandable but unclosable.
332
+ - **Closure** is one-way: a container whose children are all finished closes,
333
+ as `completed` when at least one child landed and as `not_planned` when none
334
+ did (a cohort superseded by a re-plan is finished, but nothing merged). A
335
+ reopened child moves Status back to `In Progress` and does **not** reopen the
336
+ issue — which is why the lookup reads `state: 'all'`, since an open-only
337
+ listing cannot see the container it would have to correct.
338
+ - The parent lookup resolves the native parent edge in **one** call
339
+ (`getParentIssue`), because linkage is parent→child only; it falls back to a
340
+ `type::epic` scan for a child linked by checklist alone. Children are read as
341
+ the body checklist **union** the native sub-issue edges — literally the same
342
+ reader `/mandrel-deliver`'s expansion uses, so an Epic can never be
343
+ expandable but unclosable — bounded at 5 concurrent reads.
344
+ - A checklist row citing an id that resolves to nothing is **dropped with a
345
+ warning** when the native read succeeded: hand-edited prose can cite a
346
+ deleted or mistyped issue, and no re-run will make it resolve. An
347
+ unresolvable *native* edge still fails the read. An Epic-typed child is
348
+ refused by name (`epic-typed-child`) and neither blocks nor advances the
349
+ parent.
329
350
  - Every step is best-effort and never throws: a stale container costs
330
351
  tidiness, not a landed Story's envelope.
331
352
 
@@ -240,9 +240,10 @@ discovers them only after the whole close pipeline has run, at several times
240
240
  the cost of one full-suite run in the worktree.
241
241
 
242
242
  **Run it once, last, so close can credit it.** The run belongs **after** the
243
- self-eval loop's last fix commit and immediately **before** the hand-off push,
244
- so its stamp describes the tree that is pushed; redraft rounds run scoped
245
- tests. Close skips a gate that already passed at the current HEAD, but a bare
243
+ self-eval loop's last fix commit and **after** the hand-off push — the credit
244
+ is keyed on the tree, not on push state, so pushing first keeps the stamp and
245
+ buys the ordering Step 2.5 needs (the capture is backgrounded, and its
246
+ completion ends the turn); redraft rounds run scoped tests. Close skips a gate that already passed at the current HEAD, but a bare
246
247
  `npm test` deposits no such record — the suite then runs twice per delivery,
247
248
  once here and once in the close gate chain. Pick the invocation by the same
248
249
  predicate `close-validation/gates.js` uses to choose its test gate:
@@ -262,7 +263,9 @@ The credit expires the moment it stops describing the tree: evidence is keyed
262
263
  on HEAD, the capture stamp on a content digest of `crap.targetDirs`. A
263
264
  self-eval fix — or any commit — invalidates it and close re-runs the suite for
264
265
  real, so this never trades away the gate. That keying is exactly why the run
265
- comes last.
266
+ comes last, and why the push before it is free. Close's own base-sync can
267
+ spend the stamp too when it lands base commits; it now says so out loud rather
268
+ than silently re-running the suite.
266
269
 
267
270
  **`verify[]` reuses the same stamp.** A `verify[]` entry that is itself a
268
271
  full-suite command is reported **credited** against that stamp rather than
@@ -93,18 +93,21 @@ are reference § Step 2. Hard gates always run in Step 3 — the derived level
93
93
  never disables them; do **not** pre-run the chain here — Step 2.5's credited
94
94
  suite run is the sole exception.
95
95
 
96
- ### Step 2.5 — The creditable full-suite run, then push and hand off
97
-
98
- Run the full suite **once**, after the self-eval loop's last fix commit and
99
- immediately **before** the push, in the shape close credits (**digest § 5**):
100
- the credit is keyed on the tree, so any later commit invalidates it, and a bare
101
- `npm test` deposits none. Red → fix, commit, re-run. An inline run makes the
102
- same run before Step 3.
103
-
104
- Then (sub-agent dispatch only) push `story-<storyId>` to `origin`, confirm the
105
- remote ref moved, and return the hand-off — Story id, `workCwd`, branch, pushed
106
- head SHA, self-eval verdict, `verify[]` evidence — then stop. Do not open the
107
- PR; do not compose a terminal envelope.
96
+ ### Step 2.5 — Push, then the creditable full-suite run, then hand off
97
+
98
+ **Push first.** After the self-eval loop's last fix commit, push
99
+ `story-<storyId>` to `origin`, confirming the remote ref moved: the capture
100
+ below is backgrounded, so *its* completion ends the turn.
101
+
102
+ Then run the full suite **once**, after the push, in the shape close credits
103
+ (**digest § 5**): the credit is keyed on the tree, not on push state, so only
104
+ a *later* commit invalidates it; a bare `npm test` deposits none. Red →
105
+ fix, commit, push, re-capture.
106
+
107
+ Then (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
108
+ branch, pushed head SHA, self-eval verdict, `verify[]` evidence — and stop.
109
+ Do not open the PR or compose a terminal envelope. An inline run captures
110
+ before Step 3.
108
111
 
109
112
  ## Step 3 — Close and land (`single-story-close.js`)
110
113