mandrel 2.64.0 → 2.66.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/agents/acceptance-critic.md +8 -7
- package/.agents/agents/auditor.md +20 -20
- package/.agents/agents/plan-critic.md +8 -7
- package/.agents/agents/story-worker.md +7 -7
- package/.agents/audit-checklists/quality.md +3 -0
- package/.agents/docs/agentrc-reference.json +1 -9
- package/.agents/docs/configuration.md +8 -7
- package/.agents/docs/execution-reference.md +27 -5
- package/.agents/instructions.md +10 -12
- package/.agents/rules/ci-remediation.md +3 -3
- package/.agents/rules/gherkin-standards.md +3 -2
- package/.agents/rules/git-conventions-reference.md +12 -3
- package/.agents/rules/git-conventions.md +9 -7
- package/.agents/rules/testing-standards.md +8 -7
- package/.agents/runtime-deps.json +1 -1
- package/.agents/schemas/agentrc.schema.json +6 -13
- package/.agents/schemas/audit-rules.schema.json +1 -1
- package/.agents/schemas/story-deliver-terminal.schema.json +5 -0
- package/.agents/scripts/bootstrap.js +102 -91
- package/.agents/scripts/check-context-budget.js +1 -1
- package/.agents/scripts/lib/ITicketingProvider.js +1 -3
- package/.agents/scripts/lib/audit-suite/findings.js +1 -17
- package/.agents/scripts/lib/audit-suite/frontmatter.js +0 -28
- package/.agents/scripts/lib/audit-suite/index.js +0 -6
- package/.agents/scripts/lib/audit-suite/selector.js +0 -31
- package/.agents/scripts/lib/baselines/duplication-scanner.js +17 -7
- package/.agents/scripts/lib/bootstrap/agents-md-fold.js +156 -0
- package/.agents/scripts/lib/bootstrap/commit-push.js +1 -1
- package/.agents/scripts/lib/bootstrap/manifest.js +2 -2
- package/.agents/scripts/lib/bootstrap/project-bootstrap.js +91 -107
- package/.agents/scripts/lib/cli/standard-args.js +60 -76
- package/.agents/scripts/lib/cli-args.js +26 -0
- package/.agents/scripts/lib/config/gates/shared.js +3 -3
- package/.agents/scripts/lib/config/review-chain-default.js +13 -0
- package/.agents/scripts/lib/config-settings-schema-delivery.js +2 -2
- package/.agents/scripts/lib/config-settings-schema-quality.js +11 -13
- package/.agents/scripts/lib/doc-tiers.js +25 -6
- package/.agents/scripts/lib/feedback-loop/graduate-steps.js +205 -0
- package/.agents/scripts/lib/feedback-loop/graduator-core.js +47 -782
- package/.agents/scripts/lib/feedback-loop/graduator-gh.js +449 -0
- package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
- package/.agents/scripts/lib/observability/close-telemetry.js +330 -0
- package/.agents/scripts/lib/observability/metrics-ledger.js +0 -72
- package/.agents/scripts/lib/observability/runtime-friction.js +2 -0
- package/.agents/scripts/lib/observability/signal-validator.js +17 -5
- package/.agents/scripts/lib/orchestration/code-review.js +33 -6
- package/.agents/scripts/lib/orchestration/epic-rollup.js +29 -12
- package/.agents/scripts/lib/orchestration/merge-block-class.js +20 -4
- package/.agents/scripts/lib/orchestration/merge-poll.js +41 -22
- package/.agents/scripts/lib/orchestration/plan-metrics.js +76 -63
- package/.agents/scripts/lib/orchestration/required-checks.js +147 -0
- package/.agents/scripts/lib/orchestration/review-providers/code-review.js +203 -0
- package/.agents/scripts/lib/orchestration/review-providers/review-provider-factory.js +29 -4
- package/.agents/scripts/lib/orchestration/review-providers/security-review.js +3 -2
- package/.agents/scripts/lib/orchestration/run-epilogue.js +6 -0
- package/.agents/scripts/lib/orchestration/single-story-close/failed-terminal.js +1 -0
- package/.agents/scripts/lib/orchestration/single-story-close/phases/code-review.js +2 -12
- package/.agents/scripts/lib/orchestration/single-story-close/phases/confirm-merge.js +370 -268
- package/.agents/scripts/lib/orchestration/single-story-close/phases/options.js +21 -7
- package/.agents/scripts/lib/orchestration/single-story-close/phases/post-land.js +112 -82
- package/.agents/scripts/lib/orchestration/single-story-close/phases/review-override.js +4 -0
- package/.agents/scripts/lib/orchestration/single-story-close/runner.js +393 -313
- package/.agents/scripts/lib/orchestration/story-close/phases/review-core.js +12 -87
- package/.agents/scripts/lib/orchestration/story-deliver-terminal.js +3 -0
- package/.agents/scripts/lib/orchestration/ticket-validator.js +19 -36
- package/.agents/scripts/lib/signals/detectors/common.js +63 -51
- package/.agents/scripts/lib/templates/decomposer-prompts.js +5 -24
- package/.agents/scripts/lib/transpile.js +28 -3
- package/.agents/scripts/providers/github/issues.js +14 -23
- package/.agents/scripts/single-story-close.js +10 -2
- package/.agents/scripts/single-story-confirm-merge.js +267 -238
- package/.agents/scripts/sync-claude-agents.js +1 -1
- package/.agents/skills/core/idea-refinement/SKILL.md +6 -6
- package/.agents/skills/stack/qa/qa-harness/SKILL.md +1 -2
- package/.agents/workflows/audit-architecture.md +5 -4
- package/.agents/workflows/audit-documentation.md +5 -5
- package/.agents/workflows/audit-performance.md +10 -10
- package/.agents/workflows/audit-quality.md +42 -7
- package/.agents/workflows/helpers/acceptance-self-eval.md +9 -9
- package/.agents/workflows/helpers/audit-lens-core.md +30 -57
- package/.agents/workflows/helpers/code-review.md +15 -38
- package/.agents/workflows/helpers/deliver-digest.md +2 -2
- package/.agents/workflows/helpers/deliver-reference.md +7 -3
- package/.agents/workflows/helpers/deliver-story.md +9 -1
- package/.agents/workflows/helpers/parallel-tooling.md +16 -18
- package/.agents/workflows/helpers/plan-reference.md +9 -8
- package/.agents/workflows/mandrel-deliver.md +3 -2
- package/.agents/workflows/mandrel-plan.md +11 -7
- package/.agents/workflows/mandrel-update.md +5 -3
- package/docs/CHANGELOG.md +57 -0
- package/lib/cli/claude-code-version.js +73 -0
- package/lib/cli/doctor.js +2 -2
- package/lib/cli/guarded-sync.js +87 -0
- package/lib/cli/registry.js +9 -0
- package/lib/cli/sync-agents.js +9 -92
- package/lib/cli/sync-commands.js +9 -101
- package/lib/cli/uninstall.js +37 -9
- package/lib/migrations/index.js +2 -0
- package/lib/migrations/steps/2.65.0-fold-claude-md-into-agents-md.js +38 -0
- package/package.json +3 -2
- package/.agents/scripts/lib/audit-suite/lens-diff-floor.js +0 -99
- package/.agents/scripts/lib/audit-suite/runner.js +0 -205
- package/.agents/scripts/lib/audit-suite/substitutions.js +0 -96
- package/.agents/scripts/lib/audit-suite/workflow-loader.js +0 -37
- package/.agents/scripts/lib/orchestration/story-close/phases/local-lens-review.js +0 -234
|
@@ -27,18 +27,18 @@ Per the core's Scope interpretation:
|
|
|
27
27
|
|
|
28
28
|
## Execution strategy
|
|
29
29
|
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
Sequential inline execution is the fallback (see the
|
|
30
|
+
Dispatch this lens as one `subagent_type: auditor` call. Fan its resource
|
|
31
|
+
dimensions out across parallel `auditor` subagents (parallel-tooling Rule 3),
|
|
32
|
+
merging under the self-cross-check, only when the operator explicitly asks for
|
|
33
|
+
per-dimension fan-out. Sequential inline execution is the fallback (see the
|
|
34
|
+
core's Execution strategy).
|
|
34
35
|
|
|
35
36
|
> **Measurement is non-mutating, not forbidden.** This lens is read-only with
|
|
36
|
-
> respect to source, but
|
|
37
|
-
>
|
|
38
|
-
>
|
|
39
|
-
>
|
|
40
|
-
>
|
|
41
|
-
> `.claude/workflows/audit-performance.workflow.js`.
|
|
37
|
+
> respect to source, but the auditor MUST be allowed to *run* measurements. It
|
|
38
|
+
> runs only **non-mutating** commands — profilers, timers, bundle-stat and
|
|
39
|
+
> file-size probes — and never a command that writes source, installs
|
|
40
|
+
> packages, or mutates git state or labels. The one write is the report
|
|
41
|
+
> artifact.
|
|
42
42
|
|
|
43
43
|
## Step 0: Measure before you judge (mandatory)
|
|
44
44
|
|
|
@@ -12,7 +12,8 @@ the Story under audit. The shared lens machinery lives in
|
|
|
12
12
|
`{{auditOutputDir}}/audit-quality-results.md`. Each finding carries a
|
|
13
13
|
**Category:** (`Flakiness | Coverage | Performance | Mocking | Test Plans`); the
|
|
14
14
|
report adds a **Test Strategy Assessment** table (Unit / Integration / E2E /
|
|
15
|
-
Test Plans: Healthy / Needs Work / Missing
|
|
15
|
+
Test Plans / Property-Based Testing: Healthy / Needs Work / Missing, or `N/A`
|
|
16
|
+
for Property-Based Testing when no module is a candidate).
|
|
16
17
|
|
|
17
18
|
## Scope
|
|
18
19
|
|
|
@@ -130,6 +131,36 @@ Evaluate the gathered context against the following test quality dimensions:
|
|
|
130
131
|
finding here. Route the *architectural* framing of the same defect to
|
|
131
132
|
[`audit-architecture`](audit-architecture.md)'s Shipped-But-Never-Wired
|
|
132
133
|
dimension; this lens owns the **missing-test** framing.
|
|
134
|
+
8. **Property-Based Coverage — Invariants Tested Only by Examples.** Flag a
|
|
135
|
+
module whose correctness rests on an invariant but whose tests are all
|
|
136
|
+
hand-picked examples, which structurally cannot reach the inputs nobody
|
|
137
|
+
thought to pick. A module is a candidate **only** on code evidence: a
|
|
138
|
+
documented invariant or "never"/"always" claim in a header comment; an
|
|
139
|
+
explicit state machine or status/label transition table; an
|
|
140
|
+
encode/decode, parse/serialize or normalise pair (round-trip); a
|
|
141
|
+
merge/dedup/sort/scheduler function; bounded-concurrency or retry
|
|
142
|
+
coordination over async I/O; or an idempotency claim. A module with no
|
|
143
|
+
stated or implied invariant is **never** a finding. Rank candidates by
|
|
144
|
+
Step 0's churn × coverage gap and cite their `baselines/` coverage/CRAP row
|
|
145
|
+
where one exists.
|
|
146
|
+
|
|
147
|
+
- **Toolchain by ecosystem, never one library.** Detect an existing
|
|
148
|
+
property-testing library from the consumer's manifests (e.g.
|
|
149
|
+
`fast-check` for JS/TS, `hypothesis` for Python, `proptest`/`quickcheck`
|
|
150
|
+
for Rust, `jqwik` for the JVM, `rapid`/`gopter` for Go); recommend the
|
|
151
|
+
ecosystem-idiomatic one only when none is present.
|
|
152
|
+
- Severity: an async/concurrency coordinator, or a guard whose
|
|
153
|
+
invariant protects an irreversible write, tested only by examples →
|
|
154
|
+
**High**; any other invariant-bearing module with example-only tests →
|
|
155
|
+
**Medium**; toolchain absent with no High/Medium candidate → **one Low**
|
|
156
|
+
roll-up finding, not one per module.
|
|
157
|
+
- Category: file under `Coverage`. A property test whose seed is
|
|
158
|
+
neither pinned nor printed on failure goes under `Flakiness`: a red that
|
|
159
|
+
cannot be reproduced breaks the reproducibility the rubric demands.
|
|
160
|
+
- **Name the property.** Each finding states the invariant as a testable
|
|
161
|
+
property (e.g. `decode(encode(x)) === x`; "no transition leaves a
|
|
162
|
+
terminal state") plus its generator shape (the input domain to draw
|
|
163
|
+
from) — never a bare "add property tests".
|
|
133
164
|
|
|
134
165
|
## Constraint (lens-specific carve-out)
|
|
135
166
|
|
|
@@ -150,10 +181,14 @@ table:
|
|
|
150
181
|
|
|
151
182
|
## Test Strategy Assessment
|
|
152
183
|
|
|
153
|
-
| Layer
|
|
154
|
-
|
|
|
155
|
-
| Unit Testing
|
|
156
|
-
| Integration Testing
|
|
157
|
-
| E2E Testing
|
|
158
|
-
| Test Plans
|
|
184
|
+
| Layer | Status | Notes |
|
|
185
|
+
| ---------------------- | -------------------------------------- | -------------- |
|
|
186
|
+
| Unit Testing | [Healthy / Needs Work / Missing] | [Brief reason] |
|
|
187
|
+
| Integration Testing | [Healthy / Needs Work / Missing] | [Brief reason] |
|
|
188
|
+
| E2E Testing | [Healthy / Needs Work / Missing] | [Brief reason] |
|
|
189
|
+
| Test Plans | [Healthy / Needs Work / Missing] | [Brief reason] |
|
|
190
|
+
| Property-Based Testing | [Healthy / Needs Work / Missing / N/A] | [Brief reason] |
|
|
159
191
|
```
|
|
192
|
+
|
|
193
|
+
`Property-Based Testing` reads `N/A` when the repo has no candidate module
|
|
194
|
+
(dimension 8), so a repo with no invariant-bearing code is not nagged.
|
|
@@ -32,8 +32,8 @@ per-criterion, mid-delivery, and evaluates the actual work product.
|
|
|
32
32
|
authors the Story's verdict, and it covers **every** `acceptance[]` item in
|
|
33
33
|
one file. Which pass is named by the ceremony decision
|
|
34
34
|
(`verdictOwner: 'fresh-critic' | 'inline-self-eval'` from
|
|
35
|
-
`resolveCeremonyForRisk`),
|
|
36
|
-
|
|
35
|
+
`resolveCeremonyForRisk`), which follows the **ceremony profile
|
|
36
|
+
alone**:
|
|
37
37
|
|
|
38
38
|
> ```bash
|
|
39
39
|
> node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
|
|
@@ -43,9 +43,8 @@ per-criterion, mid-delivery, and evaluates the actual work product.
|
|
|
43
43
|
> classes **for review depth**, and resolves the owner (`mode`, `reason`,
|
|
44
44
|
> `verdictOwner`): **`minimal` / `standard` → `inline`** (the default — you
|
|
45
45
|
> author the verdict yourself), **`strict` → `fresh`** (dispatch the
|
|
46
|
-
> maker-blind critic). The derived level
|
|
47
|
-
>
|
|
48
|
-
> any sensitive path.
|
|
46
|
+
> maker-blind critic). The derived level feeds `review-depth.js`, not
|
|
47
|
+
> this decision; review depth resolves `deep` for any sensitive path.
|
|
49
48
|
|
|
50
49
|
**Never run both**, and never run a preliminary self-assessment before
|
|
51
50
|
dispatching a fresh critic — the redundant pre-pass buys no measurable
|
|
@@ -69,12 +68,13 @@ per-criterion, mid-delivery, and evaluates the actual work product.
|
|
|
69
68
|
> `delivery.routing.roleScopedAgents` is enabled (the **default**), use
|
|
70
69
|
> `subagent_type: acceptance-critic`: it boots on the role-scoped
|
|
71
70
|
> [`acceptance-critic`](../../agents/acceptance-critic.md) context (its own
|
|
72
|
-
> system prompt, no
|
|
71
|
+
> system prompt, no entry-doc @-closure) carrying the maker-blind
|
|
73
72
|
> invariant and the verdict schema standalone. With the kill-switch off
|
|
74
73
|
> (`roleScopedAgents: false`), fall back to
|
|
75
|
-
> `subagent_type: general-purpose`.
|
|
76
|
-
>
|
|
77
|
-
> any harness that carries `Agent` into
|
|
74
|
+
> `subagent_type: general-purpose`. Under sub-agent dispatch this loop
|
|
75
|
+
> runs inside a `story-worker`, so the critic sits at nesting depth 2
|
|
76
|
+
> (depth 1 inline) — supported by any harness that carries `Agent` into
|
|
77
|
+
> sub-agents (Claude Code ≥ 2.1.202).
|
|
78
78
|
|
|
79
79
|
Whichever pass owns it, the verdict:
|
|
80
80
|
+ Inspects the **change set it was handed** — the one `files` list above —
|
|
@@ -118,14 +118,12 @@ dropped finding is indistinguishable from a finding you never wrote.
|
|
|
118
118
|
Use it instead of inventing a below-`Low` word of your own; a finding that
|
|
119
119
|
cannot clear the evidence bar below is **dropped**, not filed as `Info`.
|
|
120
120
|
|
|
121
|
-
## Self-cross-check (
|
|
121
|
+
## Self-cross-check (the false-positive bar) {#self-cross-check}
|
|
122
122
|
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
sequential single-pass path the same false-positive filter the orchestrated
|
|
128
|
-
path's independent adversarial reviewer applies.
|
|
123
|
+
A finding goes in the report only when it clears the bar and the exclusion
|
|
124
|
+
list below. The bar filters and tightens findings; it never invents new ones.
|
|
125
|
+
It is the one false-positive filter every execution path applies — no separate
|
|
126
|
+
adversarial reviewer runs after it.
|
|
129
127
|
|
|
130
128
|
### Per-finding evidence bar (keep or drop)
|
|
131
129
|
|
|
@@ -173,23 +171,18 @@ that rests on one of them:
|
|
|
173
171
|
> delivery shipped and nothing in production ever calls. When a candidate is
|
|
174
172
|
> genuinely one of the exclusions, cite the exclusion and drop it.
|
|
175
173
|
|
|
176
|
-
###
|
|
174
|
+
### Recording the outcome
|
|
177
175
|
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
2. Count what you kept (`k`) and what you dropped (`d`).
|
|
181
|
-
3. Record the outcome in the report's **Executive Summary** as a single line:
|
|
176
|
+
Record what you kept (`k`) and dropped (`d`) in the report's **Executive
|
|
177
|
+
Summary** as a single line:
|
|
182
178
|
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
When `d > 0`, name the dropped findings (title + the bar/exclusion reason)
|
|
188
|
-
in one short list under that line, so the filtering is auditable and never
|
|
189
|
-
silent.
|
|
179
|
+
```text
|
|
180
|
+
Self-cross-check: kept <k> / dropped <d>.
|
|
181
|
+
```
|
|
190
182
|
|
|
191
|
-
|
|
192
|
-
|
|
183
|
+
When `d > 0`, name the dropped findings (title + the bar/exclusion reason) in
|
|
184
|
+
one short list under that line, so the filtering is auditable and never
|
|
185
|
+
silent. A lens that keeps every finding still records `dropped 0`.
|
|
193
186
|
|
|
194
187
|
## Severity tally (mandatory, machine-readable) {#severity-tally}
|
|
195
188
|
|
|
@@ -261,55 +254,35 @@ available; every path emits the **identical** report contract (the finding-block
|
|
|
261
254
|
skeleton above), so downstream consumers (`audit-to-stories`) are agnostic to
|
|
262
255
|
which path produced it.
|
|
263
256
|
|
|
264
|
-
1. **
|
|
257
|
+
1. **One auditor per lens (the default).** Dispatch the lens as exactly one
|
|
265
258
|
`subagent_type: auditor` call — the standalone boot context in
|
|
266
259
|
[`../../agents/auditor.md`](../../agents/auditor.md) carries the read-only
|
|
267
260
|
MUSTs, the finding-block skeleton, the severity scale, and the
|
|
268
261
|
self-cross-check bar, so the child needs only the lens's own dimensions to
|
|
269
262
|
run. The subagent returns the **report path plus the Executive Summary**
|
|
270
263
|
(including the self-cross-check line); the parent never needs the full
|
|
271
|
-
findings inline.
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
275
|
-
|
|
276
|
-
`
|
|
277
|
-
|
|
278
|
-
call **per dimension** in a single turn via
|
|
279
|
-
[`parallel-tooling.md`](parallel-tooling.md) Rule 3, then **merge** the
|
|
264
|
+
findings inline. One auditor reads the repo once; per-dimension agents each
|
|
265
|
+
re-read it, so the single dispatch is the cheap path, not a compromise.
|
|
266
|
+
|
|
267
|
+
- **Per-dimension fan-out (operator request only).** Fan a lens out only
|
|
268
|
+
when the operator's invocation explicitly asks for it. Then dispatch one
|
|
269
|
+
`subagent_type: auditor` call **per dimension** in a single turn via
|
|
270
|
+
[`parallel-tooling.md`](parallel-tooling.md) Rule 3, and **merge** the
|
|
280
271
|
per-dimension findings under this file's self-cross-check (the merge is
|
|
281
|
-
where cross-dimension duplicates and false positives are dropped).
|
|
282
|
-
|
|
272
|
+
where cross-dimension duplicates and false positives are dropped). Never
|
|
273
|
+
fan out on your own judgment of a lens's size.
|
|
274
|
+
- **No nested fan-out.** An auditor never dispatches sub-agents of its own.
|
|
275
|
+
The fan-out, when requested, happens once, at the caller.
|
|
283
276
|
|
|
284
277
|
2. **Sequential inline execution (documented fallback).** When subagent
|
|
285
278
|
dispatch is unavailable, run the lens's steps turn-by-turn in the current
|
|
286
279
|
context exactly as written, ending with the self-cross-check. This changes
|
|
287
280
|
nothing about the report contract.
|
|
288
281
|
|
|
289
|
-
> **Orchestrated dynamic-workflow path (optimization note).** Six lenses ship a
|
|
290
|
-
> saved project workflow at `.claude/workflows/audit-<lens>.workflow.js` that,
|
|
291
|
-
> **when Claude Code dynamic workflows are available** (runtime is Claude Code,
|
|
292
|
-
> `disableWorkflows` unset, version `>= 2.1.154`), fans the dimensions out as
|
|
293
|
-
> parallel read-only subagents and runs an independent adversarial cross-check
|
|
294
|
-
> stage before synthesising the report. It derives its per-dimension prompts
|
|
295
|
-
> from the *lens* markdown at run time — the lens stays the single source of
|
|
296
|
-
> truth. This is a performance optimization over path 1, **not** a separate
|
|
297
|
-
> contract, and it is not covered by the No-Shim / hard-cutover rule in
|
|
298
|
-
> [`../../rules/git-conventions.md`](../../rules/git-conventions.md) because
|
|
299
|
-
> there is one report contract and only the execution strategy varies — the
|
|
300
|
-
> same capability-degradation pattern the protocol endorses for live-docs
|
|
301
|
-
> fallback. **The host owns the choice.** Mandrel ships no in-repo strategy
|
|
302
|
-
> selector and no force-override env var: Claude Code launches the saved
|
|
303
|
-
> workflow when it can, and you get path 1 or 2 above when it cannot.
|
|
304
|
-
> Suppress the orchestrated path with `CLAUDE_CODE_DISABLE_WORKFLOWS=1`
|
|
305
|
-
> or `disableWorkflows: true` in `.claude/settings.json`. On the orchestrated
|
|
306
|
-
> path the analysis subagents are granted only read/search tools (`Read`,
|
|
307
|
-
> `Grep`, `Glob`) — the single write is the final report artifact.
|
|
308
|
-
|
|
309
282
|
## Parallel tooling {#parallel-tooling}
|
|
310
283
|
|
|
311
284
|
When a lens batches independent reads/greps, runs a long shell (a scanner, a
|
|
312
|
-
profiler, a suite time),
|
|
313
|
-
|
|
314
|
-
|
|
315
|
-
|
|
285
|
+
profiler, a suite time), apply [`parallel-tooling.md`](parallel-tooling.md):
|
|
286
|
+
batch independent reads in one turn (Rule 1) and run long shells via
|
|
287
|
+
`run_in_background` (Rule 2). Rule 3 applies only to the caller of
|
|
288
|
+
an operator-requested per-dimension fan-out — never inside an auditor.
|
|
@@ -74,8 +74,8 @@ How each tier changes the review protocol:
|
|
|
74
74
|
adversarial pass over the diff hunting for integration regressions and
|
|
75
75
|
security-relevant edges before findings are finalized.
|
|
76
76
|
|
|
77
|
-
The LLM-backed review providers (codex, security-review,
|
|
78
|
-
the resolved `depth` into the prompt/instructions they emit so the underlying
|
|
77
|
+
The LLM-backed review providers (code-review, codex, security-review,
|
|
78
|
+
ultrareview) render the resolved `depth` into the prompt/instructions they emit so the underlying
|
|
79
79
|
model actually changes thoroughness. The native provider deliberately ignores
|
|
80
80
|
`depth` — its mechanical lint + maintainability sweep already scales with diff
|
|
81
81
|
size, and there is no "review harder" knob a deterministic scorer can turn (its
|
|
@@ -110,44 +110,21 @@ The pipeline will:
|
|
|
110
110
|
- Run a focused lint check on the change set.
|
|
111
111
|
- Post a structured summary report to the `[TICKET_ID]` issue.
|
|
112
112
|
|
|
113
|
-
###
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
`git diff --name-only`).
|
|
122
|
-
2. Selects the **local-tier** lenses that own a concern decidable from a single
|
|
123
|
-
Story's diff — `resolveLensTier(lens) === 'local'` **plus** the pure
|
|
124
|
-
`matchesAnyFilePattern` matcher against the diff (the audit-suite SDK's
|
|
125
|
-
[`selectLocalLenses`](../../scripts/lib/audit-suite/selector.js)). This is
|
|
126
|
-
deliberately **not** `selectAudits`: `selectAudits` unions in keyword and
|
|
127
|
-
gate matches and has no per-tier gate, so it would widen the roster past the
|
|
128
|
-
footprint-matched local set this tier owns.
|
|
129
|
-
3. Materializes the matched roster at **`light`** depth
|
|
130
|
-
(`STORY_SCOPE_LENS_DEPTH`) via `runAuditSuite`, surfacing the outcome on the
|
|
131
|
-
review envelope's `localLensReview` field.
|
|
132
|
-
|
|
133
|
-
A diff that matches no local lens adds **no** lens work (the roster is empty and
|
|
134
|
-
`runAuditSuite` is never invoked). The pass is advisory and best-effort: a git
|
|
135
|
-
or materialization failure degrades to a skipped envelope and never blocks the
|
|
136
|
-
close.
|
|
137
|
-
|
|
138
|
-
The live close entry point —
|
|
139
|
-
[`runStoryScopeReview`](../../scripts/lib/orchestration/single-story-close/phases/code-review.js)
|
|
140
|
-
— reaches this pass through the shared `runStoryReviewCore` spine. Because
|
|
141
|
-
the pass lives inside the close subprocess (invoked after the delivering
|
|
142
|
-
child exits), it honors the maker-blind invariant above: a maker never runs
|
|
143
|
-
its own local-lens review.
|
|
113
|
+
### Story scope runs no lens pass
|
|
114
|
+
|
|
115
|
+
Close runs **no** audit-lens pass of its own (Story #5416 retired it: it
|
|
116
|
+
materialized prompt files no workflow read, then armed auto-merge without
|
|
117
|
+
waiting). The Story-scope review is this pipeline plus CI. Local-tier lens
|
|
118
|
+
concerns are covered shift-left by the write-time authoring checklists
|
|
119
|
+
threaded into the Story prompt; the on-demand `/audit-*` workflows remain the
|
|
120
|
+
way to run a full lens over a change.
|
|
144
121
|
|
|
145
122
|
## Step 2 — Review Pillars
|
|
146
123
|
|
|
147
124
|
For each changed file, execute a strict review against four pillars. The
|
|
148
125
|
second pillar (**Integration Review**) deliberately defers the security /
|
|
149
126
|
performance / quality / coverage sweeps to the change-set-scoped lenses —
|
|
150
|
-
those
|
|
127
|
+
those are covered shift-left by the write-time lens checklists.
|
|
151
128
|
Re-walking those sweeps a second time in this pillar is duplication, not
|
|
152
129
|
defense-in-depth.
|
|
153
130
|
|
|
@@ -173,9 +150,9 @@ Does the implementation match the Story's acceptance criteria and folded Spec?
|
|
|
173
150
|
|
|
174
151
|
The diff under review is `baseRef..headRef`
|
|
175
152
|
(`main..story-<storyId>`, or the configured base branch to the Story
|
|
176
|
-
branch). The
|
|
177
|
-
|
|
178
|
-
|
|
153
|
+
branch). The write-time lens checklists have already covered the local-tier
|
|
154
|
+
concerns. Pillar findings land in the single `verification-results` comment
|
|
155
|
+
this pass posts. The
|
|
179
156
|
integration view here focuses on cross-cutting ripple within the Story and
|
|
180
157
|
contract drift against the base branch. Look for:
|
|
181
158
|
|
|
@@ -260,7 +237,7 @@ prior baseline before merging.
|
|
|
260
237
|
|
|
261
238
|
Findings are **persisted as a `verification-results` structured comment on
|
|
262
239
|
the `[TICKET_ID]` issue** by `runCodeReview` (the unified findings contract —
|
|
263
|
-
this single comment carries the Story-scope
|
|
240
|
+
this single comment carries the Story-scope review findings). The target
|
|
264
241
|
ticket is the Story. The comment
|
|
265
242
|
is idempotent — re-runs replace the prior one — and its body includes
|
|
266
243
|
severity-tier counts plus the full findings list so downstream workflows
|
|
@@ -68,8 +68,8 @@ same derived level, so the two cannot disagree. A sensitive footprint
|
|
|
68
68
|
therefore buys a **deep review**, not a fresh acceptance critic.
|
|
69
69
|
|
|
70
70
|
> **The ceremony rule, stated once.** The **profile alone** names the verdict
|
|
71
|
-
> owner
|
|
72
|
-
>
|
|
71
|
+
> owner: `minimal` / `standard` → `inline`, `strict` → `fresh`. Nothing else
|
|
72
|
+
> moves it — not the
|
|
73
73
|
> derived change level, not the footprint's sensitivity, and **not the
|
|
74
74
|
> dispatch mode**: an `inline` Story under `strict` still spawns the fresh
|
|
75
75
|
> maker-blind critic, one nesting level shallower than a dispatched one. The
|
|
@@ -78,7 +78,9 @@ but not a cycle.
|
|
|
78
78
|
unblocked it:
|
|
79
79
|
`node .agents/scripts/update-ticket-state.js --ticket <id> --state agent::ready`.
|
|
80
80
|
Do not poll the label yourself while waiting — the HITL pause is the operator's
|
|
81
|
-
turn, not a slow beat.
|
|
81
|
+
turn, not a slow beat. Before resuming, the operator raises session effort one
|
|
82
|
+
step, and raises effort before switching models
|
|
83
|
+
([effort escalation](../../docs/execution-reference.md#session-effort-and-model)).
|
|
82
84
|
|
|
83
85
|
Each beat re-probes live state: it re-resolves the graph, classifies **done**
|
|
84
86
|
(`agent::done` or a closed issue — including foreign blockers that landed in
|
|
@@ -175,7 +177,7 @@ into batches of `cap` and dispatch each batch in its own turn.
|
|
|
175
177
|
exposes agent dispatch, spawn each ready Story as its own
|
|
176
178
|
`subagent_type: story-worker` sub-agent — it boots on the role-scoped
|
|
177
179
|
[`story-worker`](../../agents/story-worker.md) context (its own system prompt, no
|
|
178
|
-
|
|
180
|
+
entry-doc @-closure) carrying the load-bearing delivery MUSTs standalone. The
|
|
179
181
|
sub-agent executes [`deliver-story.md`](deliver-story.md) Steps 0–2.5
|
|
180
182
|
(init → implement → acceptance self-eval → **push**) and stops there; **you**
|
|
181
183
|
own Step 3, serialized — see `/mandrel-deliver` § Closing what the workers hand back.
|
|
@@ -291,7 +293,9 @@ confirm instead (`captureStoryFollowUps`).
|
|
|
291
293
|
reopen the issue.
|
|
292
294
|
- The parent lookup resolves the native parent edge in **one** call
|
|
293
295
|
(`getParentIssue`), falling back to a `type::epic` scan for a child linked by
|
|
294
|
-
checklist alone.
|
|
296
|
+
checklist alone. An authoritative "no parent" narrows that scan to Epics
|
|
297
|
+
whose body checklist names the Story (no per-Epic native read); only a
|
|
298
|
+
degraded lookup reads every scanned Epic's native children. Children are the body checklist **union** the native
|
|
295
299
|
sub-issue edges — the same reader `/mandrel-deliver`'s expansion uses.
|
|
296
300
|
- A checklist row citing an id that resolves to nothing is **dropped with a
|
|
297
301
|
warning** when the native read succeeded; an unresolvable *native* edge
|
|
@@ -72,6 +72,9 @@ One branch, one PR to `main`, commits against the inline `acceptance[]` /
|
|
|
72
72
|
1. Read the Story body; its acceptance criteria are the contract. Docs are
|
|
73
73
|
digest-first; read a caller-provided `checklistPath` first, and walk any
|
|
74
74
|
`## Slicing` rows as **intra-session checkpoints** (reference § Step 1).
|
|
75
|
+
After a context summary, re-derive progress from `git log` on
|
|
76
|
+
`story-<id>` against the `## Slicing` rows (each checkpoint is a commit
|
|
77
|
+
boundary) before continuing.
|
|
75
78
|
2. Implement and commit on the Story branch, iterating with quick advisory
|
|
76
79
|
gates (`typecheck`, `lint`, scoped tests) — the full chain runs in Step 3,
|
|
77
80
|
and the **one** full-suite run at Step 2.5.
|
|
@@ -111,9 +114,14 @@ Do not open the PR or compose a terminal envelope.
|
|
|
111
114
|
serialized against sibling Stories:
|
|
112
115
|
|
|
113
116
|
```bash
|
|
114
|
-
node <main-repo>/.agents/scripts/single-story-close.js --story <storyId> --cwd <main-repo>
|
|
117
|
+
node <main-repo>/.agents/scripts/single-story-close.js --story <storyId> --cwd <main-repo> \
|
|
118
|
+
[--worker-tokens <n>]
|
|
115
119
|
```
|
|
116
120
|
|
|
121
|
+
When the host reported a total-token figure for the story-worker's Agent
|
|
122
|
+
dispatch, pass it as `--worker-tokens <n>` — close records it in its
|
|
123
|
+
result's local `telemetry`; omit it when the host reports none.
|
|
124
|
+
|
|
117
125
|
**The whole delivery tail** — gates, PR, merge wait, `agent::done` flip,
|
|
118
126
|
post-land tail in one process. Never background it, never delegate it to a
|
|
119
127
|
child, and never end your turn while it is still running: "close is running"
|
|
@@ -29,15 +29,17 @@ the batch in parallel; serial calls cost N round-trips for no gain.
|
|
|
29
29
|
- **Bounded fan-out:** keep the batch ≤ 10 calls per turn. Larger batches
|
|
30
30
|
blow the context budget and obscure the failure surface if one call errors.
|
|
31
31
|
|
|
32
|
-
## Rule 2 — `run_in_background`
|
|
32
|
+
## Rule 2 — `run_in_background` for long shells
|
|
33
33
|
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
34
|
+
A shell command that can outrun the host's synchronous Bash ceiling, or that
|
|
35
|
+
would idle the turn while independent work waits (test suites, installs,
|
|
36
|
+
multi-file lints, `git fetch --all`, container builds), runs with the `Bash`
|
|
37
|
+
tool's `run_in_background: true` flag; its completion notification is the
|
|
38
|
+
signal to proceed. Attach `Monitor` only when you must act on output before
|
|
39
|
+
the command exits.
|
|
39
40
|
|
|
40
|
-
- **Tool primitives:** `Bash(run_in_background: true)`
|
|
41
|
+
- **Tool primitives:** `Bash(run_in_background: true)`; `Monitor` only when
|
|
42
|
+
mid-run output matters.
|
|
41
43
|
- **When:** `npm test`, `npm ci`, full-repo `eslint`/`biome` runs, long
|
|
42
44
|
fetches, anything you would have prefixed with `nohup` in a terminal.
|
|
43
45
|
- **Anti-pattern:** synchronous `Bash` with a 600 000 ms timeout used as a
|
|
@@ -90,17 +92,13 @@ the same shape as Rule 1 but at the sub-agent layer.
|
|
|
90
92
|
## When the rules conflict
|
|
91
93
|
|
|
92
94
|
If a unit of work is both long (Rule 2) and independent (Rule 1 or 3),
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
recursive `Agent` fan-out is available to it, so the host does not need to
|
|
101
|
-
micromanage the child's shell **or** dispatch strategy. Mind the depth
|
|
102
|
-
budget and the compounding cost — every nesting level re-pays the
|
|
103
|
-
always-loaded context (see [`instructions.md` § 4](../../instructions.md)).
|
|
95
|
+
dispatch the `Agent` calls in one turn (Rule 3), and **inside** each sub-agent let it apply Rule 2 to its own
|
|
96
|
+
long-running shells. A sub-agent does **not** fan out again on its own
|
|
97
|
+
initiative: every nesting level re-pays the always-loaded context (see
|
|
98
|
+
[`instructions.md` § 4](../../instructions.md)), and the cost compounds with
|
|
99
|
+
depth. The one exception is a dispatch the sub-agent's own workflow names
|
|
100
|
+
explicitly — for example the maker-blind acceptance critic a `strict`-profile
|
|
101
|
+
Story worker spawns — which stays legal at that depth.
|
|
104
102
|
|
|
105
103
|
## Constraints
|
|
106
104
|
|
|
@@ -100,19 +100,20 @@ not re-deriving which assumptions were really the agent's to make.
|
|
|
100
100
|
|
|
101
101
|
## Gate #1 → the one advisory line
|
|
102
102
|
|
|
103
|
-
Gate #1 stops
|
|
104
|
-
|
|
105
|
-
advisory line
|
|
106
|
-
|
|
107
|
-
|
|
103
|
+
Gate #1 stops only when there is at least one HITL unknown or
|
|
104
|
+
`duplicates[]` is non-empty; otherwise the run announces the sharpened plan
|
|
105
|
+
intent plus the advisory line and continues to authoring. Everything else the
|
|
106
|
+
envelope surfaced collapses to **one advisory line** beneath the gate.
|
|
107
|
+
Nothing on that line reroutes the run or is invoked by `/mandrel-plan`; each
|
|
108
|
+
item names something the operator may prefer to do instead. Under
|
|
108
109
|
`--yes` the line is recorded and planning continues — an unattended run has
|
|
109
110
|
nobody to take an offer.
|
|
110
111
|
|
|
111
112
|
The line names, in order, whichever of these the envelope carries:
|
|
112
113
|
|
|
113
114
|
- **`duplicates[]`** — open Stories the seed resembles (never Epics). Name
|
|
114
|
-
the top one or two by id and title
|
|
115
|
-
still the operator's call.
|
|
115
|
+
the top one or two by id and title. A plan that duplicates open work is
|
|
116
|
+
still the operator's call, so a non-empty list also stops Gate #1.
|
|
116
117
|
- **Open `intake` rows** (`priorFeedback`) — CI-gap intake filings written by
|
|
117
118
|
[`file-ci-gap.js`](../../scripts/file-ci-gap.js) when a delivery reached an
|
|
118
119
|
Option-2 verdict in [`ci-remediation.md`](../../rules/ci-remediation.md).
|
|
@@ -339,7 +340,7 @@ On `dispatch: true`, dispatch **one fresh-context, maker-blind sub-agent**.
|
|
|
339
340
|
When `delivery.routing.roleScopedAgents` is enabled (the **default**), use
|
|
340
341
|
`subagent_type: plan-critic` — it boots on the role-scoped
|
|
341
342
|
[`plan-critic`](../../agents/plan-critic.md) context (its own system prompt,
|
|
342
|
-
no
|
|
343
|
+
no entry-doc @-closure) that carries the maker-blind invariant, the
|
|
343
344
|
`pre-mortem` charter, and the output shape standalone. When the kill-switch
|
|
344
345
|
is off (`roleScopedAgents: false`) or the host cannot spawn at this depth,
|
|
345
346
|
fall back to a generic sub-agent and hand it the same charter. Either way the
|
|
@@ -76,7 +76,8 @@ to an attended run.
|
|
|
76
76
|
it cannot read — a missing gate would co-dispatch against an unlanded
|
|
77
77
|
blocker.
|
|
78
78
|
|
|
79
|
-
2. **
|
|
79
|
+
2. **Announce (N>1).** Present the resolved order and proceed — do not wait
|
|
80
|
+
for confirmation; the operator can interject mid-run to change it.
|
|
80
81
|
|
|
81
82
|
3. **Run the beat.** One command per beat, repeated until the envelope reports
|
|
82
83
|
the run `done`:
|
|
@@ -143,7 +144,7 @@ resume what it names.
|
|
|
143
144
|
|
|
144
145
|
**Reading the outcome.** Each close ends the Story in one schema-validated
|
|
145
146
|
envelope — `landed` | `pending` | `blocked` | `failed`; statuses, exits and
|
|
146
|
-
fields are digest §
|
|
147
|
+
fields are digest § 6. `pending` is **not** a failure — run its `nextCommand`.
|
|
147
148
|
|
|
148
149
|
**Branch model (authoritative).** `story-<id>` → PR → `main` (squash +
|
|
149
150
|
required checks), per digest § 2; dependent Stories land sequentially. The
|
|
@@ -71,13 +71,17 @@ assumed; a **HITL** unknown goes to Gate #1. Under `--yes` do not ask free-form
|
|
|
71
71
|
operator questions — AFK unknowns are still researched; only HITL unknowns land
|
|
72
72
|
in Key Assumptions, each a decision-made-by-default.
|
|
73
73
|
|
|
74
|
-
**Gate #1** — STOP
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
74
|
+
**Gate #1** — STOP only when a HITL unknown the operator owns exists or
|
|
75
|
+
`duplicates[]` is non-empty (planning a duplicate of open work stays the
|
|
76
|
+
operator's call): confirm the sharpened plan intent and settle it. Otherwise
|
|
77
|
+
announce the sharpened intent and the advisory line, and continue to
|
|
78
|
+
authoring. The advisory line names what the envelope surfaced — any
|
|
79
|
+
`duplicates[]` (the stop above), open `intake` rows, a truthy
|
|
80
|
+
`memoryPoolAdvisory.recommend`, a truthy `complexitySignals.uiSurface`
|
|
81
|
+
naming [`/prototype`](prototype.md) (never invoke it here) — as
|
|
82
|
+
**one advisory line** under the gate; the line itself never reroutes the run
|
|
83
|
+
([ref](helpers/plan-reference.md)).
|
|
84
|
+
Under `--yes`, auto-proceed.
|
|
81
85
|
|
|
82
86
|
### 2. Author
|
|
83
87
|
|
|
@@ -200,13 +200,15 @@ non-optional.
|
|
|
200
200
|
|
|
201
201
|
## Step 4 — Review the surfaced changelog and update consumer-side guidance
|
|
202
202
|
|
|
203
|
-
Framework upgrades change behaviour the consumer's own `AGENTS.md`
|
|
204
|
-
|
|
203
|
+
Framework upgrades change behaviour the consumer's own `AGENTS.md` and
|
|
204
|
+
runbooks often encode. Step 1 already printed the changelog
|
|
205
205
|
for the applied range — that output is your source of truth (re-read the
|
|
206
206
|
transcript or the GitHub Releases page if it scrolled past). For each entry
|
|
207
207
|
between the installed and target versions:
|
|
208
208
|
|
|
209
|
-
1. **Consumer `AGENTS.md
|
|
209
|
+
1. **Consumer `AGENTS.md`.** It is the entry doc — `mandrel update` folds a
|
|
210
|
+
root `CLAUDE.md` into it and deletes `CLAUDE.md`, so reconcile the folded
|
|
211
|
+
content here. Update instructions so a fresh
|
|
210
212
|
agent reading them in isolation produces output that passes the
|
|
211
213
|
framework's new validators; remove or rewrite instructions that
|
|
212
214
|
contradict a tightened rule.
|