mandrel 2.64.0 → 2.66.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (105) hide show
  1. package/.agents/agents/acceptance-critic.md +8 -7
  2. package/.agents/agents/auditor.md +20 -20
  3. package/.agents/agents/plan-critic.md +8 -7
  4. package/.agents/agents/story-worker.md +7 -7
  5. package/.agents/audit-checklists/quality.md +3 -0
  6. package/.agents/docs/agentrc-reference.json +1 -9
  7. package/.agents/docs/configuration.md +8 -7
  8. package/.agents/docs/execution-reference.md +27 -5
  9. package/.agents/instructions.md +10 -12
  10. package/.agents/rules/ci-remediation.md +3 -3
  11. package/.agents/rules/gherkin-standards.md +3 -2
  12. package/.agents/rules/git-conventions-reference.md +12 -3
  13. package/.agents/rules/git-conventions.md +9 -7
  14. package/.agents/rules/testing-standards.md +8 -7
  15. package/.agents/runtime-deps.json +1 -1
  16. package/.agents/schemas/agentrc.schema.json +6 -13
  17. package/.agents/schemas/audit-rules.schema.json +1 -1
  18. package/.agents/schemas/story-deliver-terminal.schema.json +5 -0
  19. package/.agents/scripts/bootstrap.js +102 -91
  20. package/.agents/scripts/check-context-budget.js +1 -1
  21. package/.agents/scripts/lib/ITicketingProvider.js +1 -3
  22. package/.agents/scripts/lib/audit-suite/findings.js +1 -17
  23. package/.agents/scripts/lib/audit-suite/frontmatter.js +0 -28
  24. package/.agents/scripts/lib/audit-suite/index.js +0 -6
  25. package/.agents/scripts/lib/audit-suite/selector.js +0 -31
  26. package/.agents/scripts/lib/baselines/duplication-scanner.js +17 -7
  27. package/.agents/scripts/lib/bootstrap/agents-md-fold.js +156 -0
  28. package/.agents/scripts/lib/bootstrap/commit-push.js +1 -1
  29. package/.agents/scripts/lib/bootstrap/manifest.js +2 -2
  30. package/.agents/scripts/lib/bootstrap/project-bootstrap.js +91 -107
  31. package/.agents/scripts/lib/cli/standard-args.js +60 -76
  32. package/.agents/scripts/lib/cli-args.js +26 -0
  33. package/.agents/scripts/lib/config/gates/shared.js +3 -3
  34. package/.agents/scripts/lib/config/review-chain-default.js +13 -0
  35. package/.agents/scripts/lib/config-settings-schema-delivery.js +2 -2
  36. package/.agents/scripts/lib/config-settings-schema-quality.js +11 -13
  37. package/.agents/scripts/lib/doc-tiers.js +25 -6
  38. package/.agents/scripts/lib/feedback-loop/graduate-steps.js +205 -0
  39. package/.agents/scripts/lib/feedback-loop/graduator-core.js +47 -782
  40. package/.agents/scripts/lib/feedback-loop/graduator-gh.js +449 -0
  41. package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
  42. package/.agents/scripts/lib/observability/close-telemetry.js +330 -0
  43. package/.agents/scripts/lib/observability/metrics-ledger.js +0 -72
  44. package/.agents/scripts/lib/observability/runtime-friction.js +2 -0
  45. package/.agents/scripts/lib/observability/signal-validator.js +17 -5
  46. package/.agents/scripts/lib/orchestration/code-review.js +33 -6
  47. package/.agents/scripts/lib/orchestration/epic-rollup.js +29 -12
  48. package/.agents/scripts/lib/orchestration/merge-block-class.js +20 -4
  49. package/.agents/scripts/lib/orchestration/merge-poll.js +41 -22
  50. package/.agents/scripts/lib/orchestration/plan-metrics.js +76 -63
  51. package/.agents/scripts/lib/orchestration/required-checks.js +147 -0
  52. package/.agents/scripts/lib/orchestration/review-providers/code-review.js +203 -0
  53. package/.agents/scripts/lib/orchestration/review-providers/review-provider-factory.js +29 -4
  54. package/.agents/scripts/lib/orchestration/review-providers/security-review.js +3 -2
  55. package/.agents/scripts/lib/orchestration/run-epilogue.js +6 -0
  56. package/.agents/scripts/lib/orchestration/single-story-close/failed-terminal.js +1 -0
  57. package/.agents/scripts/lib/orchestration/single-story-close/phases/code-review.js +2 -12
  58. package/.agents/scripts/lib/orchestration/single-story-close/phases/confirm-merge.js +370 -268
  59. package/.agents/scripts/lib/orchestration/single-story-close/phases/options.js +21 -7
  60. package/.agents/scripts/lib/orchestration/single-story-close/phases/post-land.js +112 -82
  61. package/.agents/scripts/lib/orchestration/single-story-close/phases/review-override.js +4 -0
  62. package/.agents/scripts/lib/orchestration/single-story-close/runner.js +393 -313
  63. package/.agents/scripts/lib/orchestration/story-close/phases/review-core.js +12 -87
  64. package/.agents/scripts/lib/orchestration/story-deliver-terminal.js +3 -0
  65. package/.agents/scripts/lib/orchestration/ticket-validator.js +19 -36
  66. package/.agents/scripts/lib/signals/detectors/common.js +63 -51
  67. package/.agents/scripts/lib/templates/decomposer-prompts.js +5 -24
  68. package/.agents/scripts/lib/transpile.js +28 -3
  69. package/.agents/scripts/providers/github/issues.js +14 -23
  70. package/.agents/scripts/single-story-close.js +10 -2
  71. package/.agents/scripts/single-story-confirm-merge.js +267 -238
  72. package/.agents/scripts/sync-claude-agents.js +1 -1
  73. package/.agents/skills/core/idea-refinement/SKILL.md +6 -6
  74. package/.agents/skills/stack/qa/qa-harness/SKILL.md +1 -2
  75. package/.agents/workflows/audit-architecture.md +5 -4
  76. package/.agents/workflows/audit-documentation.md +5 -5
  77. package/.agents/workflows/audit-performance.md +10 -10
  78. package/.agents/workflows/audit-quality.md +42 -7
  79. package/.agents/workflows/helpers/acceptance-self-eval.md +9 -9
  80. package/.agents/workflows/helpers/audit-lens-core.md +30 -57
  81. package/.agents/workflows/helpers/code-review.md +15 -38
  82. package/.agents/workflows/helpers/deliver-digest.md +2 -2
  83. package/.agents/workflows/helpers/deliver-reference.md +7 -3
  84. package/.agents/workflows/helpers/deliver-story.md +9 -1
  85. package/.agents/workflows/helpers/parallel-tooling.md +16 -18
  86. package/.agents/workflows/helpers/plan-reference.md +9 -8
  87. package/.agents/workflows/mandrel-deliver.md +3 -2
  88. package/.agents/workflows/mandrel-plan.md +11 -7
  89. package/.agents/workflows/mandrel-update.md +5 -3
  90. package/docs/CHANGELOG.md +57 -0
  91. package/lib/cli/claude-code-version.js +73 -0
  92. package/lib/cli/doctor.js +2 -2
  93. package/lib/cli/guarded-sync.js +87 -0
  94. package/lib/cli/registry.js +9 -0
  95. package/lib/cli/sync-agents.js +9 -92
  96. package/lib/cli/sync-commands.js +9 -101
  97. package/lib/cli/uninstall.js +37 -9
  98. package/lib/migrations/index.js +2 -0
  99. package/lib/migrations/steps/2.65.0-fold-claude-md-into-agents-md.js +38 -0
  100. package/package.json +3 -2
  101. package/.agents/scripts/lib/audit-suite/lens-diff-floor.js +0 -99
  102. package/.agents/scripts/lib/audit-suite/runner.js +0 -205
  103. package/.agents/scripts/lib/audit-suite/substitutions.js +0 -96
  104. package/.agents/scripts/lib/audit-suite/workflow-loader.js +0 -37
  105. package/.agents/scripts/lib/orchestration/story-close/phases/local-lens-review.js +0 -234
@@ -27,18 +27,18 @@ Per the core's Scope interpretation:
27
27
 
28
28
  ## Execution strategy
29
29
 
30
- This is a **heavyweight lens**: dispatch it as a single `subagent_type: auditor`
31
- call, or fan its resource dimensions out per-dimension across parallel `auditor`
32
- subagents (parallel-tooling Rule 3) and merge under the self-cross-check.
33
- Sequential inline execution is the fallback (see the core's Execution strategy).
30
+ Dispatch this lens as one `subagent_type: auditor` call. Fan its resource
31
+ dimensions out across parallel `auditor` subagents (parallel-tooling Rule 3),
32
+ merging under the self-cross-check, only when the operator explicitly asks for
33
+ per-dimension fan-out. Sequential inline execution is the fallback (see the
34
+ core's Execution strategy).
34
35
 
35
36
  > **Measurement is non-mutating, not forbidden.** This lens is read-only with
36
- > respect to source, but it MUST be allowed to *run* measurements. The
37
- > orchestrated path grants its measurement agents a `Bash` tool restricted to a
38
- > **non-mutating command allowlist** (profilers, timers, bundle-stat and
39
- > file-size probes — never a command that writes source, installs, or mutates
40
- > git/labels). See the allowlist in the harness-generated
41
- > `.claude/workflows/audit-performance.workflow.js`.
37
+ > respect to source, but the auditor MUST be allowed to *run* measurements. It
38
+ > runs only **non-mutating** commands — profilers, timers, bundle-stat and
39
+ > file-size probes — and never a command that writes source, installs
40
+ > packages, or mutates git state or labels. The one write is the report
41
+ > artifact.
42
42
 
43
43
  ## Step 0: Measure before you judge (mandatory)
44
44
 
@@ -12,7 +12,8 @@ the Story under audit. The shared lens machinery lives in
12
12
  `{{auditOutputDir}}/audit-quality-results.md`. Each finding carries a
13
13
  **Category:** (`Flakiness | Coverage | Performance | Mocking | Test Plans`); the
14
14
  report adds a **Test Strategy Assessment** table (Unit / Integration / E2E /
15
- Test Plans: Healthy / Needs Work / Missing).
15
+ Test Plans / Property-Based Testing: Healthy / Needs Work / Missing, or `N/A`
16
+ for Property-Based Testing when no module is a candidate).
16
17
 
17
18
  ## Scope
18
19
 
@@ -130,6 +131,36 @@ Evaluate the gathered context against the following test quality dimensions:
130
131
  finding here. Route the *architectural* framing of the same defect to
131
132
  [`audit-architecture`](audit-architecture.md)'s Shipped-But-Never-Wired
132
133
  dimension; this lens owns the **missing-test** framing.
134
+ 8. **Property-Based Coverage — Invariants Tested Only by Examples.** Flag a
135
+ module whose correctness rests on an invariant but whose tests are all
136
+ hand-picked examples, which structurally cannot reach the inputs nobody
137
+ thought to pick. A module is a candidate **only** on code evidence: a
138
+ documented invariant or "never"/"always" claim in a header comment; an
139
+ explicit state machine or status/label transition table; an
140
+ encode/decode, parse/serialize or normalise pair (round-trip); a
141
+ merge/dedup/sort/scheduler function; bounded-concurrency or retry
142
+ coordination over async I/O; or an idempotency claim. A module with no
143
+ stated or implied invariant is **never** a finding. Rank candidates by
144
+ Step 0's churn × coverage gap and cite their `baselines/` coverage/CRAP row
145
+ where one exists.
146
+
147
+ - **Toolchain by ecosystem, never one library.** Detect an existing
148
+ property-testing library from the consumer's manifests (e.g.
149
+ `fast-check` for JS/TS, `hypothesis` for Python, `proptest`/`quickcheck`
150
+ for Rust, `jqwik` for the JVM, `rapid`/`gopter` for Go); recommend the
151
+ ecosystem-idiomatic one only when none is present.
152
+ - Severity: an async/concurrency coordinator, or a guard whose
153
+ invariant protects an irreversible write, tested only by examples →
154
+ **High**; any other invariant-bearing module with example-only tests →
155
+ **Medium**; toolchain absent with no High/Medium candidate → **one Low**
156
+ roll-up finding, not one per module.
157
+ - Category: file under `Coverage`. A property test whose seed is
158
+ neither pinned nor printed on failure goes under `Flakiness`: a red that
159
+ cannot be reproduced breaks the reproducibility the rubric demands.
160
+ - **Name the property.** Each finding states the invariant as a testable
161
+ property (e.g. `decode(encode(x)) === x`; "no transition leaves a
162
+ terminal state") plus its generator shape (the input domain to draw
163
+ from) — never a bare "add property tests".
133
164
 
134
165
  ## Constraint (lens-specific carve-out)
135
166
 
@@ -150,10 +181,14 @@ table:
150
181
 
151
182
  ## Test Strategy Assessment
152
183
 
153
- | Layer | Status | Notes |
154
- | ------------------- | -------------------------------- | -------------- |
155
- | Unit Testing | [Healthy / Needs Work / Missing] | [Brief reason] |
156
- | Integration Testing | [Healthy / Needs Work / Missing] | [Brief reason] |
157
- | E2E Testing | [Healthy / Needs Work / Missing] | [Brief reason] |
158
- | Test Plans | [Healthy / Needs Work / Missing] | [Brief reason] |
184
+ | Layer | Status | Notes |
185
+ | ---------------------- | -------------------------------------- | -------------- |
186
+ | Unit Testing | [Healthy / Needs Work / Missing] | [Brief reason] |
187
+ | Integration Testing | [Healthy / Needs Work / Missing] | [Brief reason] |
188
+ | E2E Testing | [Healthy / Needs Work / Missing] | [Brief reason] |
189
+ | Test Plans | [Healthy / Needs Work / Missing] | [Brief reason] |
190
+ | Property-Based Testing | [Healthy / Needs Work / Missing / N/A] | [Brief reason] |
159
191
  ```
192
+
193
+ `Property-Based Testing` reads `N/A` when the repo has no candidate module
194
+ (dimension 8), so a repo with no invariant-bearing code is not nagged.
@@ -32,8 +32,8 @@ per-criterion, mid-delivery, and evaluates the actual work product.
32
32
  authors the Story's verdict, and it covers **every** `acceptance[]` item in
33
33
  one file. Which pass is named by the ceremony decision
34
34
  (`verdictOwner: 'fresh-critic' | 'inline-self-eval'` from
35
- `resolveCeremonyForRisk`), and since Story #5343 that follows the
36
- **ceremony profile alone**:
35
+ `resolveCeremonyForRisk`), which follows the **ceremony profile
36
+ alone**:
37
37
 
38
38
  > ```bash
39
39
  > node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
@@ -43,9 +43,8 @@ per-criterion, mid-delivery, and evaluates the actual work product.
43
43
  > classes **for review depth**, and resolves the owner (`mode`, `reason`,
44
44
  > `verdictOwner`): **`minimal` / `standard` → `inline`** (the default — you
45
45
  > author the verdict yourself), **`strict` → `fresh`** (dispatch the
46
- > maker-blind critic). The derived level no longer routes this decision;
47
- > it escalates `review-depth.js` instead, which still resolves `deep` for
48
- > any sensitive path.
46
+ > maker-blind critic). The derived level feeds `review-depth.js`, not
47
+ > this decision; review depth resolves `deep` for any sensitive path.
49
48
 
50
49
  **Never run both**, and never run a preliminary self-assessment before
51
50
  dispatching a fresh critic — the redundant pre-pass buys no measurable
@@ -69,12 +68,13 @@ per-criterion, mid-delivery, and evaluates the actual work product.
69
68
  > `delivery.routing.roleScopedAgents` is enabled (the **default**), use
70
69
  > `subagent_type: acceptance-critic`: it boots on the role-scoped
71
70
  > [`acceptance-critic`](../../agents/acceptance-critic.md) context (its own
72
- > system prompt, no `CLAUDE.md` @-closure) carrying the maker-blind
71
+ > system prompt, no entry-doc @-closure) carrying the maker-blind
73
72
  > invariant and the verdict schema standalone. With the kill-switch off
74
73
  > (`roleScopedAgents: false`), fall back to
75
- > `subagent_type: general-purpose`. This loop already runs inside a Story
76
- > delivery sub-agent, so the critic sits at nesting depth 2 — supported by
77
- > any harness that carries `Agent` into sub-agents (Claude Code ≥ 2.1.202).
74
+ > `subagent_type: general-purpose`. Under sub-agent dispatch this loop
75
+ > runs inside a `story-worker`, so the critic sits at nesting depth 2
76
+ > (depth 1 inline) — supported by any harness that carries `Agent` into
77
+ > sub-agents (Claude Code ≥ 2.1.202).
78
78
 
79
79
  Whichever pass owns it, the verdict:
80
80
  + Inspects the **change set it was handed** — the one `files` list above —
@@ -118,14 +118,12 @@ dropped finding is indistinguishable from a finding you never wrote.
118
118
  Use it instead of inventing a below-`Low` word of your own; a finding that
119
119
  cannot clear the evidence bar below is **dropped**, not filed as `Info`.
120
120
 
121
- ## Self-cross-check (mandatory — filter false positives before you finalize) {#self-cross-check}
121
+ ## Self-cross-check (the false-positive bar) {#self-cross-check}
122
122
 
123
- You are your own adversarial reviewer. After you have drafted the Detailed
124
- Findings but **before** you write the report artifact, re-open every finding
125
- and hold it to the bar below. This pass is **read-only** — it filters and
126
- tightens the findings you already have; it never invents new ones. It gives the
127
- sequential single-pass path the same false-positive filter the orchestrated
128
- path's independent adversarial reviewer applies.
123
+ A finding goes in the report only when it clears the bar and the exclusion
124
+ list below. The bar filters and tightens findings; it never invents new ones.
125
+ It is the one false-positive filter every execution path applies — no separate
126
+ adversarial reviewer runs after it.
129
127
 
130
128
  ### Per-finding evidence bar (keep or drop)
131
129
 
@@ -173,23 +171,18 @@ that rests on one of them:
173
171
  > delivery shipped and nothing in production ever calls. When a candidate is
174
172
  > genuinely one of the exclusions, cite the exclusion and drop it.
175
173
 
176
- ### Final re-open-and-drop pass (mandatory)
174
+ ### Recording the outcome
177
175
 
178
- 1. Walk your Detailed Findings once more, applying the bar and the exclusion
179
- list above. Remove every finding that fails.
180
- 2. Count what you kept (`k`) and what you dropped (`d`).
181
- 3. Record the outcome in the report's **Executive Summary** as a single line:
176
+ Record what you kept (`k`) and dropped (`d`) in the report's **Executive
177
+ Summary** as a single line:
182
178
 
183
- ```text
184
- Self-cross-check: kept <k> / dropped <d>.
185
- ```
186
-
187
- When `d > 0`, name the dropped findings (title + the bar/exclusion reason)
188
- in one short list under that line, so the filtering is auditable and never
189
- silent.
179
+ ```text
180
+ Self-cross-check: kept <k> / dropped <d>.
181
+ ```
190
182
 
191
- A lens that keeps every finding still records `dropped 0` — the line's absence
192
- is itself a defect (it means the pass did not run).
183
+ When `d > 0`, name the dropped findings (title + the bar/exclusion reason) in
184
+ one short list under that line, so the filtering is auditable and never
185
+ silent. A lens that keeps every finding still records `dropped 0`.
193
186
 
194
187
  ## Severity tally (mandatory, machine-readable) {#severity-tally}
195
188
 
@@ -261,55 +254,35 @@ available; every path emits the **identical** report contract (the finding-block
261
254
  skeleton above), so downstream consumers (`audit-to-stories`) are agnostic to
262
255
  which path produced it.
263
256
 
264
- 1. **Subagent dispatch (first-class).** Dispatch the lens as a single
257
+ 1. **One auditor per lens (the default).** Dispatch the lens as exactly one
265
258
  `subagent_type: auditor` call — the standalone boot context in
266
259
  [`../../agents/auditor.md`](../../agents/auditor.md) carries the read-only
267
260
  MUSTs, the finding-block skeleton, the severity scale, and the
268
261
  self-cross-check bar, so the child needs only the lens's own dimensions to
269
262
  run. The subagent returns the **report path plus the Executive Summary**
270
263
  (including the self-cross-check line); the parent never needs the full
271
- findings inline. This is the default: the auditor boots without the full
272
- project closure, so the spawn is cheap relative to running the lens inline
273
- in the parent's context.
274
-
275
- - **Per-dimension fan-out (heavyweight lenses).** `audit-architecture`,
276
- `audit-performance`, and `audit-documentation` carry enough independent
277
- dimensions to be worth fanning out: dispatch one `subagent_type: auditor`
278
- call **per dimension** in a single turn via
279
- [`parallel-tooling.md`](parallel-tooling.md) Rule 3, then **merge** the
264
+ findings inline. One auditor reads the repo once; per-dimension agents each
265
+ re-read it, so the single dispatch is the cheap path, not a compromise.
266
+
267
+ - **Per-dimension fan-out (operator request only).** Fan a lens out only
268
+ when the operator's invocation explicitly asks for it. Then dispatch one
269
+ `subagent_type: auditor` call **per dimension** in a single turn via
270
+ [`parallel-tooling.md`](parallel-tooling.md) Rule 3, and **merge** the
280
271
  per-dimension findings under this file's self-cross-check (the merge is
281
- where cross-dimension duplicates and false positives are dropped). Respect
282
- the nesting-depth budget and the concurrency cap that Rule 3 documents.
272
+ where cross-dimension duplicates and false positives are dropped). Never
273
+ fan out on your own judgment of a lens's size.
274
+ - **No nested fan-out.** An auditor never dispatches sub-agents of its own.
275
+ The fan-out, when requested, happens once, at the caller.
283
276
 
284
277
  2. **Sequential inline execution (documented fallback).** When subagent
285
278
  dispatch is unavailable, run the lens's steps turn-by-turn in the current
286
279
  context exactly as written, ending with the self-cross-check. This changes
287
280
  nothing about the report contract.
288
281
 
289
- > **Orchestrated dynamic-workflow path (optimization note).** Six lenses ship a
290
- > saved project workflow at `.claude/workflows/audit-<lens>.workflow.js` that,
291
- > **when Claude Code dynamic workflows are available** (runtime is Claude Code,
292
- > `disableWorkflows` unset, version `>= 2.1.154`), fans the dimensions out as
293
- > parallel read-only subagents and runs an independent adversarial cross-check
294
- > stage before synthesising the report. It derives its per-dimension prompts
295
- > from the *lens* markdown at run time — the lens stays the single source of
296
- > truth. This is a performance optimization over path 1, **not** a separate
297
- > contract, and it is not covered by the No-Shim / hard-cutover rule in
298
- > [`../../rules/git-conventions.md`](../../rules/git-conventions.md) because
299
- > there is one report contract and only the execution strategy varies — the
300
- > same capability-degradation pattern the protocol endorses for live-docs
301
- > fallback. **The host owns the choice.** Mandrel ships no in-repo strategy
302
- > selector and no force-override env var: Claude Code launches the saved
303
- > workflow when it can, and you get path 1 or 2 above when it cannot.
304
- > Suppress the orchestrated path with `CLAUDE_CODE_DISABLE_WORKFLOWS=1`
305
- > or `disableWorkflows: true` in `.claude/settings.json`. On the orchestrated
306
- > path the analysis subagents are granted only read/search tools (`Read`,
307
- > `Grep`, `Glob`) — the single write is the final report artifact.
308
-
309
282
  ## Parallel tooling {#parallel-tooling}
310
283
 
311
284
  When a lens batches independent reads/greps, runs a long shell (a scanner, a
312
- profiler, a suite time), or fans out per-dimension, apply
313
- [`parallel-tooling.md`](parallel-tooling.md): batch independent reads in one
314
- turn (Rule 1), run long shells via `run_in_background` + `Monitor` (Rule 2),
315
- and dispatch N independent units as N `Agent` calls in one turn (Rule 3).
285
+ profiler, a suite time), apply [`parallel-tooling.md`](parallel-tooling.md):
286
+ batch independent reads in one turn (Rule 1) and run long shells via
287
+ `run_in_background` (Rule 2). Rule 3 applies only to the caller of
288
+ an operator-requested per-dimension fan-out — never inside an auditor.
@@ -74,8 +74,8 @@ How each tier changes the review protocol:
74
74
  adversarial pass over the diff hunting for integration regressions and
75
75
  security-relevant edges before findings are finalized.
76
76
 
77
- The LLM-backed review providers (codex, security-review, ultrareview) render
78
- the resolved `depth` into the prompt/instructions they emit so the underlying
77
+ The LLM-backed review providers (code-review, codex, security-review,
78
+ ultrareview) render the resolved `depth` into the prompt/instructions they emit so the underlying
79
79
  model actually changes thoroughness. The native provider deliberately ignores
80
80
  `depth` — its mechanical lint + maintainability sweep already scales with diff
81
81
  size, and there is no "review harder" knob a deterministic scorer can turn (its
@@ -110,44 +110,21 @@ The pipeline will:
110
110
  - Run a focused lint check on the change set.
111
111
  - Post a structured summary report to the `[TICKET_ID]` issue.
112
112
 
113
- ### Step 1a — Story-scope local-lens pass (`scope: story` only)
114
-
115
- When `scope === 'story'`, the shared review spine
116
- [`runStoryReviewCore`](../../scripts/lib/orchestration/story-close/phases/review-core.js)
117
- runs a **shift-left local-lens pass** in the same close subprocess, *before*
118
- returning the review envelope. It:
119
-
120
- 1. Enumerates the actual Story diff (`baseRef...headRef` via
121
- `git diff --name-only`).
122
- 2. Selects the **local-tier** lenses that own a concern decidable from a single
123
- Story's diff — `resolveLensTier(lens) === 'local'` **plus** the pure
124
- `matchesAnyFilePattern` matcher against the diff (the audit-suite SDK's
125
- [`selectLocalLenses`](../../scripts/lib/audit-suite/selector.js)). This is
126
- deliberately **not** `selectAudits`: `selectAudits` unions in keyword and
127
- gate matches and has no per-tier gate, so it would widen the roster past the
128
- footprint-matched local set this tier owns.
129
- 3. Materializes the matched roster at **`light`** depth
130
- (`STORY_SCOPE_LENS_DEPTH`) via `runAuditSuite`, surfacing the outcome on the
131
- review envelope's `localLensReview` field.
132
-
133
- A diff that matches no local lens adds **no** lens work (the roster is empty and
134
- `runAuditSuite` is never invoked). The pass is advisory and best-effort: a git
135
- or materialization failure degrades to a skipped envelope and never blocks the
136
- close.
137
-
138
- The live close entry point —
139
- [`runStoryScopeReview`](../../scripts/lib/orchestration/single-story-close/phases/code-review.js)
140
- — reaches this pass through the shared `runStoryReviewCore` spine. Because
141
- the pass lives inside the close subprocess (invoked after the delivering
142
- child exits), it honors the maker-blind invariant above: a maker never runs
143
- its own local-lens review.
113
+ ### Story scope runs no lens pass
114
+
115
+ Close runs **no** audit-lens pass of its own (Story #5416 retired it: it
116
+ materialized prompt files no workflow read, then armed auto-merge without
117
+ waiting). The Story-scope review is this pipeline plus CI. Local-tier lens
118
+ concerns are covered shift-left by the write-time authoring checklists
119
+ threaded into the Story prompt; the on-demand `/audit-*` workflows remain the
120
+ way to run a full lens over a change.
144
121
 
145
122
  ## Step 2 — Review Pillars
146
123
 
147
124
  For each changed file, execute a strict review against four pillars. The
148
125
  second pillar (**Integration Review**) deliberately defers the security /
149
126
  performance / quality / coverage sweeps to the change-set-scoped lenses —
150
- those ran shift-left in the Story-scope local-lens pass (Step 1a).
127
+ those are covered shift-left by the write-time lens checklists.
151
128
  Re-walking those sweeps a second time in this pillar is duplication, not
152
129
  defense-in-depth.
153
130
 
@@ -173,9 +150,9 @@ Does the implementation match the Story's acceptance criteria and folded Spec?
173
150
 
174
151
  The diff under review is `baseRef..headRef`
175
152
  (`main..story-<storyId>`, or the configured base branch to the Story
176
- branch). The Story-scope local-lens pass (Step 1a) has already covered the
177
- local-tier concerns. Lens findings and pillar findings share the single
178
- `verification-results` comment this pass posts. The
153
+ branch). The write-time lens checklists have already covered the local-tier
154
+ concerns. Pillar findings land in the single `verification-results` comment
155
+ this pass posts. The
179
156
  integration view here focuses on cross-cutting ripple within the Story and
180
157
  contract drift against the base branch. Look for:
181
158
 
@@ -260,7 +237,7 @@ prior baseline before merging.
260
237
 
261
238
  Findings are **persisted as a `verification-results` structured comment on
262
239
  the `[TICKET_ID]` issue** by `runCodeReview` (the unified findings contract —
263
- this single comment carries the Story-scope lens findings). The target
240
+ this single comment carries the Story-scope review findings). The target
264
241
  ticket is the Story. The comment
265
242
  is idempotent — re-runs replace the prior one — and its body includes
266
243
  severity-tier counts plus the full findings list so downstream workflows
@@ -68,8 +68,8 @@ same derived level, so the two cannot disagree. A sensitive footprint
68
68
  therefore buys a **deep review**, not a fresh acceptance critic.
69
69
 
70
70
  > **The ceremony rule, stated once.** The **profile alone** names the verdict
71
- > owner (Story #5343, narrowed to that one input by #5366): `minimal` /
72
- > `standard` → `inline`, `strict` → `fresh`. Nothing else moves it — not the
71
+ > owner: `minimal` / `standard` → `inline`, `strict` → `fresh`. Nothing else
72
+ > moves it — not the
73
73
  > derived change level, not the footprint's sensitivity, and **not the
74
74
  > dispatch mode**: an `inline` Story under `strict` still spawns the fresh
75
75
  > maker-blind critic, one nesting level shallower than a dispatched one. The
@@ -78,7 +78,9 @@ but not a cycle.
78
78
  unblocked it:
79
79
  `node .agents/scripts/update-ticket-state.js --ticket <id> --state agent::ready`.
80
80
  Do not poll the label yourself while waiting — the HITL pause is the operator's
81
- turn, not a slow beat.
81
+ turn, not a slow beat. Before resuming, the operator raises session effort one
82
+ step, and raises effort before switching models
83
+ ([effort escalation](../../docs/execution-reference.md#session-effort-and-model)).
82
84
 
83
85
  Each beat re-probes live state: it re-resolves the graph, classifies **done**
84
86
  (`agent::done` or a closed issue — including foreign blockers that landed in
@@ -175,7 +177,7 @@ into batches of `cap` and dispatch each batch in its own turn.
175
177
  exposes agent dispatch, spawn each ready Story as its own
176
178
  `subagent_type: story-worker` sub-agent — it boots on the role-scoped
177
179
  [`story-worker`](../../agents/story-worker.md) context (its own system prompt, no
178
- `CLAUDE.md` @-closure) carrying the load-bearing delivery MUSTs standalone. The
180
+ entry-doc @-closure) carrying the load-bearing delivery MUSTs standalone. The
179
181
  sub-agent executes [`deliver-story.md`](deliver-story.md) Steps 0–2.5
180
182
  (init → implement → acceptance self-eval → **push**) and stops there; **you**
181
183
  own Step 3, serialized — see `/mandrel-deliver` § Closing what the workers hand back.
@@ -291,7 +293,9 @@ confirm instead (`captureStoryFollowUps`).
291
293
  reopen the issue.
292
294
  - The parent lookup resolves the native parent edge in **one** call
293
295
  (`getParentIssue`), falling back to a `type::epic` scan for a child linked by
294
- checklist alone. Children are the body checklist **union** the native
296
+ checklist alone. An authoritative "no parent" narrows that scan to Epics
297
+ whose body checklist names the Story (no per-Epic native read); only a
298
+ degraded lookup reads every scanned Epic's native children. Children are the body checklist **union** the native
295
299
  sub-issue edges — the same reader `/mandrel-deliver`'s expansion uses.
296
300
  - A checklist row citing an id that resolves to nothing is **dropped with a
297
301
  warning** when the native read succeeded; an unresolvable *native* edge
@@ -72,6 +72,9 @@ One branch, one PR to `main`, commits against the inline `acceptance[]` /
72
72
  1. Read the Story body; its acceptance criteria are the contract. Docs are
73
73
  digest-first; read a caller-provided `checklistPath` first, and walk any
74
74
  `## Slicing` rows as **intra-session checkpoints** (reference § Step 1).
75
+ After a context summary, re-derive progress from `git log` on
76
+ `story-<id>` against the `## Slicing` rows (each checkpoint is a commit
77
+ boundary) before continuing.
75
78
  2. Implement and commit on the Story branch, iterating with quick advisory
76
79
  gates (`typecheck`, `lint`, scoped tests) — the full chain runs in Step 3,
77
80
  and the **one** full-suite run at Step 2.5.
@@ -111,9 +114,14 @@ Do not open the PR or compose a terminal envelope.
111
114
  serialized against sibling Stories:
112
115
 
113
116
  ```bash
114
- node <main-repo>/.agents/scripts/single-story-close.js --story <storyId> --cwd <main-repo>
117
+ node <main-repo>/.agents/scripts/single-story-close.js --story <storyId> --cwd <main-repo> \
118
+ [--worker-tokens <n>]
115
119
  ```
116
120
 
121
+ When the host reported a total-token figure for the story-worker's Agent
122
+ dispatch, pass it as `--worker-tokens <n>` — close records it in its
123
+ result's local `telemetry`; omit it when the host reports none.
124
+
117
125
  **The whole delivery tail** — gates, PR, merge wait, `agent::done` flip,
118
126
  post-land tail in one process. Never background it, never delegate it to a
119
127
  child, and never end your turn while it is still running: "close is running"
@@ -29,15 +29,17 @@ the batch in parallel; serial calls cost N round-trips for no gain.
29
29
  - **Bounded fan-out:** keep the batch ≤ 10 calls per turn. Larger batches
30
30
  blow the context budget and obscure the failure surface if one call errors.
31
31
 
32
- ## Rule 2 — `run_in_background` + `Monitor` for long shells
32
+ ## Rule 2 — `run_in_background` for long shells
33
33
 
34
- Shell commands that exceed roughly 30 seconds (test suites, installs,
35
- multi-file lints, `git fetch --all`, container builds) **must** use the
36
- `Bash` tool's `run_in_background: true` flag and stream events via the
37
- `Monitor` tool. A synchronous `Bash` call holds the assistant turn open for
38
- the full duration and blocks every other parallel opportunity.
34
+ A shell command that can outrun the host's synchronous Bash ceiling, or that
35
+ would idle the turn while independent work waits (test suites, installs,
36
+ multi-file lints, `git fetch --all`, container builds), runs with the `Bash`
37
+ tool's `run_in_background: true` flag; its completion notification is the
38
+ signal to proceed. Attach `Monitor` only when you must act on output before
39
+ the command exits.
39
40
 
40
- - **Tool primitives:** `Bash(run_in_background: true)` + `Monitor`.
41
+ - **Tool primitives:** `Bash(run_in_background: true)`; `Monitor` only when
42
+ mid-run output matters.
41
43
  - **When:** `npm test`, `npm ci`, full-repo `eslint`/`biome` runs, long
42
44
  fetches, anything you would have prefixed with `nohup` in a terminal.
43
45
  - **Anti-pattern:** synchronous `Bash` with a 600 000 ms timeout used as a
@@ -90,17 +92,13 @@ the same shape as Rule 1 but at the sub-agent layer.
90
92
  ## When the rules conflict
91
93
 
92
94
  If a unit of work is both long (Rule 2) and independent (Rule 1 or 3),
93
- prefer the higher-numbered rule — the parallelism gain compounds the
94
- background-shell gain. Concretely: dispatch the `Agent` calls in one turn
95
- (Rule 3), and **inside** each sub-agent let it apply Rule 2 to its own
96
- long-running shells — and, within the supported nesting depth budget
97
- (verified depth 2, announced max depth 5), let it apply
98
- **Rule 3** to its own independent sub-units as well, not only Rule 2
99
- background shells. A sub-agent is a full orchestrator at its own level:
100
- recursive `Agent` fan-out is available to it, so the host does not need to
101
- micromanage the child's shell **or** dispatch strategy. Mind the depth
102
- budget and the compounding cost — every nesting level re-pays the
103
- always-loaded context (see [`instructions.md` § 4](../../instructions.md)).
95
+ dispatch the `Agent` calls in one turn (Rule 3), and **inside** each sub-agent let it apply Rule 2 to its own
96
+ long-running shells. A sub-agent does **not** fan out again on its own
97
+ initiative: every nesting level re-pays the always-loaded context (see
98
+ [`instructions.md` § 4](../../instructions.md)), and the cost compounds with
99
+ depth. The one exception is a dispatch the sub-agent's own workflow names
100
+ explicitly — for example the maker-blind acceptance critic a `strict`-profile
101
+ Story worker spawns — which stays legal at that depth.
104
102
 
105
103
  ## Constraints
106
104
 
@@ -100,19 +100,20 @@ not re-deriving which assumptions were really the agent's to make.
100
100
 
101
101
  ## Gate #1 → the one advisory line
102
102
 
103
- Gate #1 stops for exactly two things — the sharpened plan intent and any HITL
104
- unknown — and everything else the envelope surfaced collapses to **one
105
- advisory line** beneath it. Nothing on that line stops the run,
106
- reroutes it, or is invoked by `/mandrel-plan`; each item names something the
107
- operator may prefer to do instead, and the run proceeds either way. Under
103
+ Gate #1 stops only when there is at least one HITL unknown or
104
+ `duplicates[]` is non-empty; otherwise the run announces the sharpened plan
105
+ intent plus the advisory line and continues to authoring. Everything else the
106
+ envelope surfaced collapses to **one advisory line** beneath the gate.
107
+ Nothing on that line reroutes the run or is invoked by `/mandrel-plan`; each
108
+ item names something the operator may prefer to do instead. Under
108
109
  `--yes` the line is recorded and planning continues — an unattended run has
109
110
  nobody to take an offer.
110
111
 
111
112
  The line names, in order, whichever of these the envelope carries:
112
113
 
113
114
  - **`duplicates[]`** — open Stories the seed resembles (never Epics). Name
114
- the top one or two by id and title; a plan that duplicates open work is
115
- still the operator's call.
115
+ the top one or two by id and title. A plan that duplicates open work is
116
+ still the operator's call, so a non-empty list also stops Gate #1.
116
117
  - **Open `intake` rows** (`priorFeedback`) — CI-gap intake filings written by
117
118
  [`file-ci-gap.js`](../../scripts/file-ci-gap.js) when a delivery reached an
118
119
  Option-2 verdict in [`ci-remediation.md`](../../rules/ci-remediation.md).
@@ -339,7 +340,7 @@ On `dispatch: true`, dispatch **one fresh-context, maker-blind sub-agent**.
339
340
  When `delivery.routing.roleScopedAgents` is enabled (the **default**), use
340
341
  `subagent_type: plan-critic` — it boots on the role-scoped
341
342
  [`plan-critic`](../../agents/plan-critic.md) context (its own system prompt,
342
- no `CLAUDE.md` @-closure) that carries the maker-blind invariant, the
343
+ no entry-doc @-closure) that carries the maker-blind invariant, the
343
344
  `pre-mortem` charter, and the output shape standalone. When the kill-switch
344
345
  is off (`roleScopedAgents: false`) or the host cannot spawn at this depth,
345
346
  fall back to a generic sub-agent and hand it the same charter. Either way the
@@ -76,7 +76,8 @@ to an attended run.
76
76
  it cannot read — a missing gate would co-dispatch against an unlanded
77
77
  blocker.
78
78
 
79
- 2. **Confirm (N>1).** Present the order; wait unless `--yes`.
79
+ 2. **Announce (N>1).** Present the resolved order and proceed — do not wait
80
+ for confirmation; the operator can interject mid-run to change it.
80
81
 
81
82
  3. **Run the beat.** One command per beat, repeated until the envelope reports
82
83
  the run `done`:
@@ -143,7 +144,7 @@ resume what it names.
143
144
 
144
145
  **Reading the outcome.** Each close ends the Story in one schema-validated
145
146
  envelope — `landed` | `pending` | `blocked` | `failed`; statuses, exits and
146
- fields are digest § 5. `pending` is **not** a failure — run its `nextCommand`.
147
+ fields are digest § 6. `pending` is **not** a failure — run its `nextCommand`.
147
148
 
148
149
  **Branch model (authoritative).** `story-<id>` → PR → `main` (squash +
149
150
  required checks), per digest § 2; dependent Stories land sequentially. The
@@ -71,13 +71,17 @@ assumed; a **HITL** unknown goes to Gate #1. Under `--yes` do not ask free-form
71
71
  operator questions — AFK unknowns are still researched; only HITL unknowns land
72
72
  in Key Assumptions, each a decision-made-by-default.
73
73
 
74
- **Gate #1** — STOP for exactly two things: confirm the sharpened plan intent,
75
- and settle any HITL unknown the operator owns. Everything else the envelope
76
- surfaced — `duplicates[]`, open `intake` rows, a truthy
77
- `memoryPoolAdvisory.recommend`, a truthy `complexitySignals.uiSurface` naming
78
- [`/prototype`](prototype.md) (never invoke it here) — collapses to
79
- **one advisory line** under the gate; none of it stops the run or reroutes
80
- it ([ref](helpers/plan-reference.md)). Under `--yes`, auto-proceed.
74
+ **Gate #1** — STOP only when a HITL unknown the operator owns exists or
75
+ `duplicates[]` is non-empty (planning a duplicate of open work stays the
76
+ operator's call): confirm the sharpened plan intent and settle it. Otherwise
77
+ announce the sharpened intent and the advisory line, and continue to
78
+ authoring. The advisory line names what the envelope surfaced — any
79
+ `duplicates[]` (the stop above), open `intake` rows, a truthy
80
+ `memoryPoolAdvisory.recommend`, a truthy `complexitySignals.uiSurface`
81
+ naming [`/prototype`](prototype.md) (never invoke it here) — as
82
+ **one advisory line** under the gate; the line itself never reroutes the run
83
+ ([ref](helpers/plan-reference.md)).
84
+ Under `--yes`, auto-proceed.
81
85
 
82
86
  ### 2. Author
83
87
 
@@ -200,13 +200,15 @@ non-optional.
200
200
 
201
201
  ## Step 4 — Review the surfaced changelog and update consumer-side guidance
202
202
 
203
- Framework upgrades change behaviour the consumer's own `AGENTS.md` /
204
- `CLAUDE.md` and runbooks often encode. Step 1 already printed the changelog
203
+ Framework upgrades change behaviour the consumer's own `AGENTS.md` and
204
+ runbooks often encode. Step 1 already printed the changelog
205
205
  for the applied range — that output is your source of truth (re-read the
206
206
  transcript or the GitHub Releases page if it scrolled past). For each entry
207
207
  between the installed and target versions:
208
208
 
209
- 1. **Consumer `AGENTS.md` / `CLAUDE.md`.** Update instructions so a fresh
209
+ 1. **Consumer `AGENTS.md`.** It is the entry doc — `mandrel update` folds a
210
+ root `CLAUDE.md` into it and deletes `CLAUDE.md`, so reconcile the folded
211
+ content here. Update instructions so a fresh
210
212
  agent reading them in isolation produces output that passes the
211
213
  framework's new validators; remove or rewrite instructions that
212
214
  contradict a tightened rule.