mandrel 2.59.0 → 2.60.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/README.md +11 -9
- package/.agents/agents/acceptance-critic.md +24 -43
- package/.agents/agents/story-worker.md +18 -19
- package/.agents/docs/SDLC.md +6 -6
- package/.agents/docs/agentrc-reference.json +1 -2
- package/.agents/docs/configuration.md +29 -46
- package/.agents/docs/quality-gates.md +8 -4
- package/.agents/docs/workflows.md +1 -1
- package/.agents/instructions.md +4 -5
- package/.agents/rules/ci-remediation.md +41 -8
- package/.agents/rules/known-tooling-behavior.md +65 -15
- package/.agents/schemas/acceptance-eval-verdict.schema.json +1 -1
- package/.agents/schemas/agentrc.schema.json +6 -11
- package/.agents/schemas/story-deliver-terminal.schema.json +3 -3
- package/.agents/scripts/README.md +11 -1
- package/.agents/scripts/acceptance-eval.js +25 -27
- package/.agents/scripts/ceremony-derive.js +15 -10
- package/.agents/scripts/check-context-budget.js +148 -228
- package/.agents/scripts/check-schema-references.js +5 -3
- package/.agents/scripts/check-workflow-citations.js +33 -147
- package/.agents/scripts/coverage-capture.js +7 -4
- package/.agents/scripts/deliver-light.js +41 -100
- package/.agents/scripts/deliver-run.js +631 -0
- package/.agents/scripts/file-ci-gap.js +59 -11
- package/.agents/scripts/lib/baselines/crap-preview-incremental.js +6 -2
- package/.agents/scripts/lib/changed-files.js +30 -0
- package/.agents/scripts/lib/config/delivery-routing.js +5 -4
- package/.agents/scripts/lib/config/explain.js +1 -3
- package/.agents/scripts/lib/config/gates/crap-incremental-coverage.schema.js +1 -1
- package/.agents/scripts/lib/config-resolver.js +1 -0
- package/.agents/scripts/lib/config-settings-schema-delivery.js +28 -21
- package/.agents/scripts/lib/coverage-capture-fullscope.js +10 -2
- package/.agents/scripts/lib/coverage-capture-incremental.js +3 -2
- package/.agents/scripts/lib/coverage-capture-usage.js +4 -1
- package/.agents/scripts/lib/doc-tiers.js +4 -2
- package/.agents/scripts/lib/feedback-loop/graduator-core.js +7 -6
- package/.agents/scripts/lib/feedback-loop/retro-proposals-graduator.js +7 -5
- package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
- package/.agents/scripts/lib/gh-exec.js +160 -0
- package/.agents/scripts/lib/observability/source-classifier.js +1 -0
- package/.agents/scripts/lib/orchestration/ceremony-routing.js +74 -132
- package/.agents/scripts/lib/orchestration/ci-rerun-guard.js +123 -12
- package/.agents/scripts/lib/orchestration/complexity-gate.js +180 -352
- package/.agents/scripts/lib/orchestration/light-suitability.js +71 -136
- package/.agents/scripts/lib/orchestration/plan-context.js +13 -25
- package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +8 -6
- package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +76 -95
- package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +35 -18
- package/.agents/scripts/lib/orchestration/plan-persist/summary.js +11 -11
- package/.agents/scripts/lib/orchestration/plan-persist/supersede-ops.js +63 -29
- package/.agents/scripts/lib/orchestration/review-depth.js +14 -11
- package/.agents/scripts/lib/orchestration/run-epilogue.js +260 -182
- package/.agents/scripts/lib/orchestration/run-scoped-config.js +63 -99
- package/.agents/scripts/lib/orchestration/single-story-close/phases/base-sync.js +3 -3
- package/.agents/scripts/lib/orchestration/single-story-close/phases/graphql-preflight.js +137 -0
- package/.agents/scripts/lib/orchestration/single-story-close/runner.js +105 -18
- package/.agents/scripts/lib/orchestration/story-deliver-terminal.js +4 -3
- package/.agents/scripts/lib/orchestration/story-follow-ups.js +156 -39
- package/.agents/scripts/lib/orchestration/story-init-envelope.js +71 -0
- package/.agents/scripts/lib/orchestration/task-body-validator.js +8 -17
- package/.agents/scripts/lib/orchestration/ticket-validator.js +44 -183
- package/.agents/scripts/lib/orchestration/ticketing/reads.js +14 -25
- package/.agents/scripts/lib/story-body/body-format-lints.js +58 -12
- package/.agents/scripts/lib/story-body/story-body.js +83 -29
- package/.agents/scripts/lib/templates/decomposer-prompts.js +7 -15
- package/.agents/scripts/lib/wave-runner/live-probe.js +31 -5
- package/.agents/scripts/merge-baseline.js +4 -5
- package/.agents/scripts/plan-context.js +117 -28
- package/.agents/scripts/plan-persist.js +79 -28
- package/.agents/scripts/plan-run-epilogue.js +11 -8
- package/.agents/scripts/pr-watch-with-update.js +9 -2
- package/.agents/scripts/run-verify.js +13 -6
- package/.agents/scripts/single-story-init.js +7 -57
- package/.agents/scripts/stories-wave-tick.js +160 -26
- package/.agents/skills/core/gates-and-baselines/reference.md +0 -1
- package/.agents/skills/skills.index.json +2 -2
- package/.agents/skills/stack/qa/playwright/SKILL.md +26 -0
- package/.agents/workflows/helpers/acceptance-self-eval.md +84 -157
- package/.agents/workflows/helpers/code-review.md +4 -2
- package/.agents/workflows/helpers/deliver-digest.md +31 -24
- package/.agents/workflows/helpers/deliver-light.md +92 -101
- package/.agents/workflows/helpers/deliver-reference.md +116 -100
- package/.agents/workflows/helpers/deliver-story-reference.md +58 -124
- package/.agents/workflows/helpers/deliver-story.md +17 -18
- package/.agents/workflows/helpers/plan-reference.md +65 -54
- package/.agents/workflows/mandrel-deliver.md +47 -31
- package/.agents/workflows/mandrel-plan.md +22 -21
- package/.agents/workflows/mandrel-update.md +36 -21
- package/docs/CHANGELOG.md +35 -0
- package/lib/cli/update.js +376 -17
- package/lib/migrations/index.js +2 -0
- package/lib/migrations/steps/2.60.0-retire-audit-results-autofile.js +40 -0
- package/package.json +2 -1
- package/.agents/schemas/model-attribution.schema.json +0 -53
- package/.agents/scripts/lib/orchestration/model-attribution.js +0 -418
- package/.agents/scripts/lib/orchestration/story-plan-state.js +0 -33
- package/.agents/scripts/lib/orchestration/structured-comment-parser.js +0 -67
package/.agents/README.md
CHANGED
|
@@ -118,8 +118,9 @@ clone produces zero file mutations.
|
|
|
118
118
|
Once installed, the ongoing upgrade path is **`mandrel update`** — it bumps
|
|
119
119
|
`mandrel` to the newest published version, re-runs `mandrel sync`,
|
|
120
120
|
applies version-keyed migrations, and verifies the install with
|
|
121
|
-
`mandrel doctor`. The
|
|
122
|
-
|
|
121
|
+
`mandrel doctor`. The command performs **no `git` mutation** — it reads the
|
|
122
|
+
index read-only and its closing line **reports** whether the dependency bump
|
|
123
|
+
is staged, printing the exact `git add` to run when it is not:
|
|
123
124
|
|
|
124
125
|
```bash
|
|
125
126
|
npx mandrel update # update → sync → migrate → doctor
|
|
@@ -643,7 +644,8 @@ Schema conventions:
|
|
|
643
644
|
## Code review providers (pluggable chain)
|
|
644
645
|
|
|
645
646
|
`runCodeReview()` (invoked from `helpers/deliver-story` and `/mandrel-deliver`'s
|
|
646
|
-
|
|
647
|
+
review pass, at the depth `review-depth.js` derives) loads its review backend
|
|
648
|
+
through a pluggable registry
|
|
647
649
|
configured via `delivery.codeReview.providers` — an array of entries iterated
|
|
648
650
|
in declaration order. The chain-entry field semantics (`name`, `scopes`,
|
|
649
651
|
`optional`, `manualPrompt`, `when`), the fix budget, and the cross-runtime
|
|
@@ -652,13 +654,13 @@ contract are documented once in
|
|
|
652
654
|
|
|
653
655
|
## Feedback loop — verification-results auto-graduation
|
|
654
656
|
|
|
655
|
-
When a Story (or plan-run) finalize path runs,
|
|
656
|
-
|
|
657
|
-
the unified `verification-results` structured comment, routed by source
|
|
657
|
+
When a Story (or plan-run) finalize path runs, the retro's actionable routed
|
|
658
|
+
proposals can be auto-graduated into follow-up issues, routed by source
|
|
658
659
|
classification into the framework or consumer repo. The toggle
|
|
659
|
-
(`delivery.feedbackLoop.
|
|
660
|
-
example, and the idempotency-marker behaviour are documented once
|
|
661
|
-
|
|
660
|
+
(`delivery.feedbackLoop.retroProposals`, default `false` since Story #5341),
|
|
661
|
+
the opt-in example, and the idempotency-marker behaviour are documented once
|
|
662
|
+
in
|
|
663
|
+
[`docs/configuration.md` § delivery.feedbackLoop](docs/configuration.md#deliveryfeedbackloop--retro-auto-graduation).
|
|
662
664
|
|
|
663
665
|
## Worktree dependency strategies
|
|
664
666
|
|
|
@@ -3,10 +3,10 @@ name: acceptance-critic
|
|
|
3
3
|
description: >-
|
|
4
4
|
Role-scoped boot context for a maker-blind acceptance critic. Booted on its
|
|
5
5
|
own system prompt (no CLAUDE.md / instructions.md closure). Scores a delivered
|
|
6
|
-
diff against the Story's acceptance
|
|
7
|
-
|
|
8
|
-
helpers/deliver-story Step 1a
|
|
9
|
-
|
|
6
|
+
diff against the Story's acceptance criteria and emits the verdict schema —
|
|
7
|
+
without seeing the maker's self-assessment. Dispatched by
|
|
8
|
+
helpers/deliver-story Step 1a under ceremonyProfile strict, the one profile
|
|
9
|
+
whose verdict owner is a fresh critic.
|
|
10
10
|
---
|
|
11
11
|
|
|
12
12
|
<!--
|
|
@@ -48,8 +48,8 @@ the step-by-step. This shared core binds every role:
|
|
|
48
48
|
# acceptance-critic — maker-blind acceptance evaluation
|
|
49
49
|
|
|
50
50
|
You are an **independent acceptance critic**. You score a delivered change
|
|
51
|
-
against
|
|
52
|
-
verdict. You are deliberately isolated from the author's reasoning.
|
|
51
|
+
against **every** item in the Story's `acceptance[]` array and emit one
|
|
52
|
+
structured verdict. You are deliberately isolated from the author's reasoning.
|
|
53
53
|
|
|
54
54
|
## Maker-blind — the load-bearing invariant (MUST)
|
|
55
55
|
|
|
@@ -69,23 +69,24 @@ turned in about it. Your only trusted inputs are:
|
|
|
69
69
|
in your verdict rather than substituting your own enumeration.
|
|
70
70
|
- the Story's inline `acceptance[]` and `verify[]` arrays, read from the
|
|
71
71
|
**Story body itself** (`gh issue view <storyId> --json body`) — its `##
|
|
72
|
-
Acceptance` / `## Verify` sections are the SSOT.
|
|
73
|
-
|
|
72
|
+
Acceptance` / `## Verify` sections are the SSOT. No structured comment
|
|
73
|
+
carries them.
|
|
74
74
|
- the **actual output** of the `verify[]` commands you run yourself.
|
|
75
75
|
|
|
76
76
|
Treat the implementation reasoning as untrusted. Score each criterion afresh
|
|
77
77
|
from the evidence.
|
|
78
78
|
|
|
79
|
-
## Scope —
|
|
79
|
+
## Scope — every criterion, exactly once
|
|
80
80
|
|
|
81
|
-
You
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
81
|
+
You score the Story's **whole** `acceptance[]` array and emit one verdict
|
|
82
|
+
record per item, in acceptance-array order. You do **not** re-slice, split or
|
|
83
|
+
sample the criteria: the gate reads the Story's own criteria count and refuses
|
|
84
|
+
a verdict whose `criteria[]` length differs, before scoring and without
|
|
85
|
+
consuming a round.
|
|
85
86
|
|
|
86
87
|
## Per-criterion evaluation
|
|
87
88
|
|
|
88
|
-
For each acceptance item
|
|
89
|
+
For each acceptance item:
|
|
89
90
|
|
|
90
91
|
1. **Inspect the change set** — read the files your caller named and look for
|
|
91
92
|
the change that would satisfy the criterion.
|
|
@@ -93,34 +94,14 @@ For each acceptance item in your cluster:
|
|
|
93
94
|
**required evidence**. A criterion cannot be scored `met` without the
|
|
94
95
|
supporting `verify[]` evidence where a `verify[]` command is relevant to it.
|
|
95
96
|
`verify[]` is evidence, not optional advisory pre-flight.
|
|
96
|
-
3. **Share `lint` / `typecheck` evidence with close** (Story #4250). When a
|
|
97
|
-
`verify[]` command is **byte-identical** to a close-validation gate, run it
|
|
98
|
-
through `evidence-gate.js` in the **same Story worktree** close validates so
|
|
99
|
-
a passing run records an evidence entry in the keyspace close consults:
|
|
100
|
-
|
|
101
|
-
```bash
|
|
102
|
-
node <main-repo>/.agents/scripts/evidence-gate.js \
|
|
103
|
-
--standalone --scope-id <storyId> --gate lint \
|
|
104
|
-
--worktree <worktree> -- npm run lint
|
|
105
|
-
|
|
106
|
-
node <main-repo>/.agents/scripts/evidence-gate.js \
|
|
107
|
-
--standalone --scope-id <storyId> --gate typecheck \
|
|
108
|
-
--worktree <worktree> -- <resolved typecheck command>
|
|
109
|
-
```
|
|
110
|
-
|
|
111
|
-
**Never** run the coverage / CRAP suite through `evidence-gate.js` to stamp
|
|
112
|
-
it fresh — a false-fresh coverage record without `coverage-final.json`
|
|
113
|
-
silently weakens the floor.
|
|
114
97
|
|
|
115
98
|
## Verdict schema (MUST)
|
|
116
99
|
|
|
117
|
-
Write
|
|
118
|
-
`temp/acceptance-verdict-<storyId>-r<round
|
|
119
|
-
sibling critics cannot overwrite each other, conforming to
|
|
100
|
+
Write **one** verdict file under `temp/` (e.g.
|
|
101
|
+
`temp/acceptance-verdict-<storyId>-r<round>.json`) conforming to
|
|
120
102
|
[`acceptance-eval-verdict.schema.json`](../schemas/acceptance-eval-verdict.schema.json):
|
|
121
|
-
one `criteria[]` record per acceptance item
|
|
122
|
-
|
|
123
|
-
`acceptance[]` array, not within your cluster — the caller merges on it.
|
|
103
|
+
one `criteria[]` record per acceptance item, in acceptance-array order, with
|
|
104
|
+
`index` being the criterion's position in that array.
|
|
124
105
|
|
|
125
106
|
```json
|
|
126
107
|
{
|
|
@@ -149,14 +130,14 @@ order. Each `index` is the criterion's position in the Story's **full**
|
|
|
149
130
|
- `unmet` — not addressed, or the evidence contradicts the claim.
|
|
150
131
|
|
|
151
132
|
**Return the verdict file's absolute path to your caller — never invoke
|
|
152
|
-
`acceptance-eval.js` yourself.** The caller
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
the gate
|
|
133
|
+
`acceptance-eval.js` yourself.** The caller calls the gate **once** per round;
|
|
134
|
+
the round counter is Story-scoped, so a second call spends a round for nothing.
|
|
135
|
+
The **proceed / redraft / block** decision is the gate's, not yours. You score;
|
|
136
|
+
the gate decides.
|
|
156
137
|
|
|
157
138
|
## Boundaries
|
|
158
139
|
|
|
159
140
|
- Do not fix the code, redraft the diff, or commit. You evaluate and report.
|
|
160
|
-
- Do not invent criteria beyond
|
|
141
|
+
- Do not invent criteria beyond the Story's `acceptance[]` array.
|
|
161
142
|
- Emit only paths, criteria text, and observed results — never secrets or raw
|
|
162
143
|
credential values (security-baseline § Data Leakage & Logging).
|
|
@@ -89,20 +89,17 @@ names. A null digest path means no docs mandate.
|
|
|
89
89
|
(**typecheck, lint, test, format, maintainability, coverage, crap**) and is
|
|
90
90
|
the authoritative gate — do not pre-run it. The **one** exception is the
|
|
91
91
|
full suite: after the self-eval loop's last fix commit, run it once in
|
|
92
|
-
`<workCwd>`
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
ceiling, dispatch it in the
|
|
98
|
-
re-invokes you. Never spawn a task to poll or
|
|
99
|
-
waiter with a wrong condition outlives the agent.
|
|
100
|
-
evidence a gate did work —
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
plus `verify[]`, not the whole suite. Share `lint` / `typecheck` evidence
|
|
104
|
-
with close via `evidence-gate.js`; never stamp coverage / CRAP fresh any
|
|
105
|
-
other way.
|
|
92
|
+
`<workCwd>` exactly as
|
|
93
|
+
[`deliver-digest.md`](../workflows/helpers/deliver-digest.md) § 5 states it —
|
|
94
|
+
that section is the rule's only home, so read the invocation there rather
|
|
95
|
+
than from a copy here.
|
|
96
|
+
|
|
97
|
+
If the suite outruns the host's sync Bash ceiling, dispatch it in the
|
|
98
|
+
**background**: its completion re-invokes you. Never spawn a task to poll or
|
|
99
|
+
`sleep`-loop against it; a waiter with a wrong condition outlives the agent.
|
|
100
|
+
An exit code is never evidence a gate did work — its **output** is. Redraft
|
|
101
|
+
rounds run the scoped projects for the roots you changed plus `verify[]`, not
|
|
102
|
+
the whole suite, and never stamp coverage / CRAP fresh any other way.
|
|
106
103
|
|
|
107
104
|
Gate output that lies: [`known-tooling-behavior.md`](../rules/known-tooling-behavior.md).
|
|
108
105
|
Waiter traps: [`parallel-tooling.md`](../workflows/helpers/parallel-tooling.md) Rule 2.
|
|
@@ -111,11 +108,13 @@ Waiter traps: [`parallel-tooling.md`](../workflows/helpers/parallel-tooling.md)
|
|
|
111
108
|
|
|
112
109
|
**Before** flipping to `closing`, run the bounded self-eval loop
|
|
113
110
|
([`acceptance-self-eval.md`](../workflows/helpers/acceptance-self-eval.md)).
|
|
114
|
-
Derive the change set
|
|
115
|
-
`node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
111
|
+
Derive the change set and the verdict owner with
|
|
112
|
+
`node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>`.
|
|
113
|
+
Under the default profile the owner is **you**: author **one** verdict file
|
|
114
|
+
covering every `acceptance[]` item, scoring the derived `files` set — never
|
|
115
|
+
one you re-derive — with `verify[]` output as evidence, and score it in one
|
|
116
|
+
gate call. Under `strict` the owner is a fresh critic; hand it that same
|
|
117
|
+
`files` list. **proceed** → flip to `closing`, run the suite, push, hand off;
|
|
119
118
|
**redraft** → fix the criteria, commit, re-eval; **block** → take the
|
|
120
119
|
blocked path below. Never hand off an unscored branch.
|
|
121
120
|
|
package/.agents/docs/SDLC.md
CHANGED
|
@@ -264,12 +264,12 @@ Dependent Stories land sequentially so each builds on the previous merge to
|
|
|
264
264
|
|
|
265
265
|
### Ceremony
|
|
266
266
|
|
|
267
|
-
|
|
268
|
-
(`minimal` | `standard` | `strict`, default
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
full profile × scope matrix lives in
|
|
267
|
+
Who authors a Story's acceptance verdict is selected by
|
|
268
|
+
`delivery.routing.ceremonyProfile` (`minimal` | `standard` | `strict`, default
|
|
269
|
+
`standard`) and by nothing else. The **derived change level** is a separate
|
|
270
|
+
decision and tunes review depth and audit-lens selection. Hard gates (lint /
|
|
271
|
+
test / format / coverage / CRAP / maintainability) always run at close —
|
|
272
|
+
neither decision disables them. The full profile × scope matrix lives in
|
|
273
273
|
[`mandrel-deliver.md` § Ceremony](../workflows/mandrel-deliver.md).
|
|
274
274
|
|
|
275
275
|
### State sync
|
|
@@ -42,7 +42,7 @@ hand-mirrored, so a key cannot exist in one artifact and not another.
|
|
|
42
42
|
"$schema": "./.agents/schemas/agentrc.schema.json",
|
|
43
43
|
"project": { /* paths, commands, baseBranch, docsContextFiles */ },
|
|
44
44
|
"github": { /* owner, repo, branchProtection, mergeMethods, notifications */ },
|
|
45
|
-
"planning": { /*
|
|
45
|
+
"planning": { /* memoryPool, navigation */ },
|
|
46
46
|
"delivery": { /* execution, quality, worktreeIsolation, deliverRunner, ... */ },
|
|
47
47
|
"qa": { /* featureRoot, fixturesManifest, environments, personas */ }
|
|
48
48
|
}
|
|
@@ -207,7 +207,7 @@ Everything `/mandrel-deliver` and `single-story-close` consume: execution timeou
|
|
|
207
207
|
| `quality.gates.crap.incrementalCoverage.skipWhenUnchanged` | No | `boolean` | `true` | Skip the capture entirely when no changed file under `crap.targetDirs` versus `baseRef` was touched. The only measured saving, and gate-semantics-neutral. Defaults to true. |
|
|
208
208
|
| `quality.gates.crap.incrementalCoverage.baselineJoin` | No | `boolean` | `false` | Let the CRAP join resolve a method in a file the diff did not touch from its committed baseline row instead of requiring fresh coverage for it. A gate loosening, not a saving — defaults to false. |
|
|
209
209
|
| `quality.gates.crap.incrementalCoverage.enabled` | No | `boolean` | — | DEPRECATED alias for setting both `skipWhenUnchanged` and `baselineJoin`. Prefer the two switches: they are not equally safe, and bundling them is why the earlier default flip was reverted. Either explicit switch overrides this alias. |
|
|
210
|
-
| `quality.gates.crap.incrementalCoverage.baseRef` | No | `string` | — |
|
|
210
|
+
| `quality.gates.crap.incrementalCoverage.baseRef` | No | `string` | — | Default git ref the changed-file set is computed against, for a caller that passes no `--ref`. A `--ref` the caller named wins over this value, so a caller that anchors another gate on the same ref — `.husky/pre-push`, which passes `--ref origin/main` and then previews `--changed-since origin/main` — resolves ONE scope for both steps (Story #5365). Omitted and unpassed, the gate’s own default `main` applies. |
|
|
211
211
|
| `quality.gates.maintainability` | No | `object` | — | Maintainability-index ratchet. Scores per file as the average over its methods, so deleting a small high-MI method can legitimately lower a file’s score. |
|
|
212
212
|
| `quality.gates.maintainability.enabled` | No | `boolean` | `true` | When false, the checker exits 0 with a skip line and the gate is reported as `skipped`, never omitted. |
|
|
213
213
|
| `quality.gates.maintainability.baselinePath` | No | `string` | `"baselines/maintainability.json"` | Repo-root-relative path to the gate's committed baseline artifact. |
|
|
@@ -299,9 +299,8 @@ Everything `/mandrel-deliver` and `single-story-close` consume: execution timeou
|
|
|
299
299
|
| `refactorStage.enabled` | No | `boolean` | `false` | When true, story-deliver runs an advisory post-green refactor stage (core/code-review-and-quality skill, Post-Green Refactor Pass) after the suite is green. Default false — when unset the stage is skipped and close-validation gate semantics are unchanged. |
|
|
300
300
|
| `acceptanceEval` | No | `object` | — | Story #3819. Bounded per-Story acceptance self-eval loop. After the implementation commits land and before the Story-implementation phase flips to `closing`, an independent (fresh-context) critic pass scores the caller-injected change set against each inline `acceptance[]` item, redrafts the unmet items, and re-evaluates — capped at `maxRounds` redraft rounds (0 = scored once, no redraft), then escalates to `agent::blocked` when criteria remain unmet. There is no `enabled` flag: the scoring pass is a hard cutover (always on). |
|
|
301
301
|
| `acceptanceEval.maxRounds` | No | `integer` | `2` | Maximum number of redraft rounds before escalation. Default 2; 0 means the verdict is scored once with no redraft round (Story #5313 dropped the hard ceiling and the floor-of-one clamp). |
|
|
302
|
-
| `feedbackLoop` | No | `object` | — | Opt-
|
|
303
|
-
| `feedbackLoop.
|
|
304
|
-
| `feedbackLoop.retroProposals` | No | `boolean` | `true` | When true (default), the retro auto-files its actionable routed proposals as meta::<framework-gap\|consumer-improvement> + friction::<category> issues via the graduator pre-parsed-findings seam, and the rendered retro sections list the filed issue numbers instead of paste-ready gh command stanzas. Set to false to fall back to the command stanzas. |
|
|
302
|
+
| `feedbackLoop` | No | `object` | — | Opt-in toggle for the close-time retro auto-file graduator, plus the friction recurrence window. Auto-filing defaults to OFF (Story #5341). |
|
|
303
|
+
| `feedbackLoop.retroProposals` | No | `boolean` | `false` | When true, the retro auto-files its actionable routed proposals as meta::<framework-gap\|consumer-improvement> + friction::<category> issues via the graduator pre-parsed-findings seam, and the rendered retro sections list the filed issue numbers instead of paste-ready gh command stanzas. Defaults to false (Story #5341), which renders the command stanzas instead. |
|
|
305
304
|
| `feedbackLoop.frictionWindowDays` | No | `integer` | — | How many days back the run-scope friction recurrence window reaches (Story #4850). The window spans every surviving per-Story signal stream rather than the triggering run's own Stories, so that a defect firing once per Story can reach the actionable threshold; this bounds it by age so a defect fixed weeks ago stops re-routing. Rows older than the bound — and rows carrying no readable timestamp — are excluded and counted on the roll-up step result. Default 30. |
|
|
306
305
|
| `auditToStories` | No | `object` | — | Knobs for the `/audit-to-stories` unattended (`--auto`) sweep (Story #4626). |
|
|
307
306
|
| `auditToStories.severityFloor` | No | `"critical"` \| `"high"` \| `"medium"` \| `"low"` \| `"all"` | — | Minimum severity a finding must meet to be proposed as a Story on an unattended `/audit-to-stories --auto` sweep (Story #4626). Default high. |
|
|
@@ -316,9 +315,9 @@ Everything `/mandrel-deliver` and `single-story-close` consume: execution timeou
|
|
|
316
315
|
| `ci.blockOnAdvisoryFailure` | No | `boolean` | `true` | Story #5096. When true (default), delivery refuses to arm — and disarms — GitHub native auto-merge while a non-required (advisory) check is genuinely red on the PR head and GitHub reports the PR mergeable anyway (mergeStateStatus=UNSTABLE). `--auto` waits on REQUIRED contexts only, so without this a red advisory quality gate merges unattended. Set false to restore the pre-#5096 behaviour verbatim. |
|
|
317
316
|
| `ci.advisoryAllowlist` | No | `array<string>` | `[]` | Story #5096. Check-run names exempt from blockOnAdvisoryFailure — a red run whose name matches exactly never blocks arming. Matching is exact; an unnamed run can never match and always blocks. |
|
|
318
317
|
| `ci.rerunAdvisory` | No | `integer` | `0` | Story #5266. How many times close may re-run a failed advisory workflow run before blocking on it, per close invocation. Default 0: close spends no CI minutes and issues no GitHub mutation on an advisory red unless asked. At n > 0 the failed run(s) are re-run within that allowance and the merge wait re-polls inside its existing budget, landing or blocking on the re-run verdict. Overridden per invocation by --rerun-advisory <n>. |
|
|
319
|
-
| `routing` | No | `object` | — | v2 delivery-spawn routing: role-scoped boot contexts and the ceremony profile. The v1 singleDelivery epic-route kill-switch was removed in Stage 6; the freshCriticSampleRate sampling floor was retired in Story #5313. |
|
|
318
|
+
| `routing` | No | `object` | — | v2 delivery-spawn routing: role-scoped boot contexts and the ceremony profile. The v1 singleDelivery epic-route kill-switch was removed in Stage 6; the freshCriticSampleRate sampling floor was retired in Story #5313 and the derived-level ceremony routing in Story #5343. |
|
|
320
319
|
| `routing.roleScopedAgents` | No | `boolean` | `true` | Epic #4478 (M7-B). Kill-switch for the role-scoped boot contexts. When true (default), a converted delivery spawn (`story-worker`, `acceptance-critic`) boots on its own `.claude/agents/<role>.md` system prompt instead of re-paying the full CLAUDE.md @-import closure. When false, every converted spawn falls back to `subagent_type: general-purpose` — the instant, code-rollback-free per-consumer revert, and the universal escape for hosts that ignore `.claude/agents/`. The fallback is the full-closure agent that ran before M7-B, so flipping it off never drops a gate. |
|
|
321
|
-
| `routing.ceremonyProfile` | No | `"minimal"` \| `"standard"` \| `"strict"` | `"standard"` | Acceptance-ceremony depth. minimal =
|
|
320
|
+
| `routing.ceremonyProfile` | No | `"minimal"` \| `"standard"` \| `"strict"` | `"standard"` | Acceptance-ceremony depth — who authors the Story acceptance verdict. minimal and standard (default) = the inline self-eval, whatever the diff touches; strict = a fresh-context maker-blind critic. Review depth is a separate decision and still derives `deep` for any sensitive path (review-depth.js). |
|
|
322
321
|
| `routing.closeAndLand` | No | `boolean` | `true` | When true (default), single-story-close lands through merge in one close. Opt out per-run with --no-wait-merge. |
|
|
323
322
|
|
|
324
323
|
### `qa` (optional)
|
|
@@ -400,38 +399,17 @@ pre-computed inventory added a second, staler answer to the same question.
|
|
|
400
399
|
A config still carrying the retired key is a hard validation failure; the
|
|
401
400
|
2.20.0 retirement migration strips it on upgrade.
|
|
402
401
|
|
|
403
|
-
- **`complexityGate`.**
|
|
404
|
-
|
|
405
|
-
|
|
406
|
-
|
|
407
|
-
|
|
408
|
-
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
verdict via `plan-persist.js --route-downgrade-reason "<why>"` (recorded on
|
|
415
|
-
every created Story's `story-plan-state` checkpoint); persist validates a
|
|
416
|
-
lite claim against each authored Story's own shape (`changes[]` count,
|
|
417
|
-
acceptance count, creates-vs-refactors mix, sensitive-path classes — the
|
|
418
|
-
framework constants `STORY_SHAPE_CEILINGS`) and **fails closed to `full`**
|
|
419
|
-
when the shape exceeds the ceilings; and `/mandrel-deliver` re-derives the route
|
|
420
|
-
from the fetched Story body via the same shape function at dispatch. The
|
|
421
|
-
`route::lite` label is a human-visible hint only — a lost label cannot
|
|
422
|
-
misroute delivery. A lite-shaped Story executes inline (no story-worker or
|
|
423
|
-
acceptance-critic sub-agent fan-out); sensitivity always wins — a footprint
|
|
424
|
-
intersecting a sensitive-path class routes `full` and keeps its fresh
|
|
425
|
-
critic. The lite path **never** relaxes a non-negotiable: it still produces
|
|
426
|
-
a Story ticket, still lands via a PR to `main`, still runs every repo
|
|
427
|
-
quality gate, and still honours `rules/security-baseline.md` — those gates
|
|
428
|
-
run in `single-story-close.js` regardless of route. **Knobs:** `enabled`
|
|
429
|
-
(default `true`; `false` disables lite routing everywhere) and
|
|
430
|
-
`maxArtifacts` (default `1` — a signal threshold, not a router). Defaults
|
|
431
|
-
live on `DEFAULT_COMPLEXITY_GATE` in
|
|
432
|
-
[`lib/orchestration/complexity-gate.js`](../scripts/lib/orchestration/complexity-gate.js);
|
|
433
|
-
a malformed or negative value falls back to the default rather than
|
|
434
|
-
widening the lite path.
|
|
402
|
+
- **`complexityGate`.** **Retired** (Story #5312). The planner's authored
|
|
403
|
+
lite claim (`plan-persist.js --route-downgrade-reason`), the persist-time
|
|
404
|
+
shape backstop that validated it against `STORY_SHAPE_CEILINGS`, the
|
|
405
|
+
`route::lite` hint label and this whole knob block are gone. Persist no
|
|
406
|
+
longer routes: every Story lands through the same engine and the same close
|
|
407
|
+
gates, and a config still setting `planning.complexityGate` is a hard
|
|
408
|
+
validation failure the `2.57.0` retirement migration strips on upgrade.
|
|
409
|
+
What survives is delivery-side and reads evidence rather than a declaration
|
|
410
|
+
— `/mandrel-plan`'s advisory `complexitySignals` (no routing authority), the
|
|
411
|
+
`/deliver-light` suitability gate's two absolute risk rules, and that path's
|
|
412
|
+
diff backstop against the actual change set.
|
|
435
413
|
|
|
436
414
|
### `delivery`
|
|
437
415
|
|
|
@@ -489,14 +467,19 @@ non-Claude consumers can pin the same `.agents/` version unmodified. The fix
|
|
|
489
467
|
budget (`maxFixAttempts` / `maxFixScopeFiles`) uses the same values for every
|
|
490
468
|
Story in a run — there is no per-Story override.
|
|
491
469
|
|
|
492
|
-
#### `delivery.feedbackLoop` —
|
|
470
|
+
#### `delivery.feedbackLoop` — retro auto-graduation
|
|
493
471
|
|
|
494
|
-
`
|
|
495
|
-
|
|
496
|
-
|
|
497
|
-
|
|
498
|
-
|
|
499
|
-
|
|
472
|
+
`retroProposals` auto-graduates the retro's actionable routed proposals into
|
|
473
|
+
follow-up issues; the graduator embeds a content-derived idempotency marker so
|
|
474
|
+
re-runs skip proposals that already have an issue (re-enabling after a
|
|
475
|
+
manual-triage window is safe). It defaults to `false` (Story #5341) — set it
|
|
476
|
+
to `true` to opt in.
|
|
477
|
+
|
|
478
|
+
Two sibling keys were retired and a config carrying either fails validation;
|
|
479
|
+
delete it. `codeReviewAutoFile` went when Story #4411 unified the pass.
|
|
480
|
+
`auditResultsAutoFile` went in Story #5366: its graduator had already been
|
|
481
|
+
deleted, so the toggle had no runtime reader and Story #5341's flip of its
|
|
482
|
+
default changed nothing. `mandrel update` strips it for you.
|
|
500
483
|
|
|
501
484
|
---
|
|
502
485
|
|
|
@@ -165,9 +165,13 @@ baseline still trips the gate.
|
|
|
165
165
|
### When the floor gate fires
|
|
166
166
|
|
|
167
167
|
- **Pre-push** (`.husky/pre-push`): diff-scoped, fast path only —
|
|
168
|
-
`
|
|
169
|
-
|
|
170
|
-
dispatcher, diff-scoped via
|
|
168
|
+
`coverage-capture.js` first, then
|
|
169
|
+
`quality-preview.js --changed-since origin/main` (MI + CRAP preview)
|
|
170
|
+
and `npm run crap:check` (unified dispatcher, diff-scoped via
|
|
171
|
+
`delivery.quality.gateScoping`). Capture leads because the preview's
|
|
172
|
+
CRAP half scores `coverage/coverage-final.json` off disk, so previewing
|
|
173
|
+
first scores whatever artifact an earlier run happened to leave there
|
|
174
|
+
(Story #5356). Full-repo
|
|
171
175
|
lint, docs generation checks, and the complete test suite are **not**
|
|
172
176
|
run on push; use `npm run verify` locally before a PR. CI enforces the
|
|
173
177
|
authoritative full gate set on every PR.
|
|
@@ -916,7 +920,7 @@ deliberately not `keyField`: CRAP groups by file (`keyField: 'path'`) but
|
|
|
916
920
|
ships one row per method, so keying on `keyField` would drop every method in a
|
|
917
921
|
file but one. Any `baselines/*.json` whose `$schema` is not a known per-kind
|
|
918
922
|
envelope — `arch-cycles`, `cyclomatic`, `dead-exports`, `audit-ledger`,
|
|
919
|
-
`context-budget
|
|
923
|
+
`context-budget` — is handed straight back to
|
|
920
924
|
`git merge-file`, so registering the driver cannot change their behaviour.
|
|
921
925
|
|
|
922
926
|
Registration has two halves:
|
|
@@ -59,7 +59,7 @@ description, edit the workflow file’s front-matter and regenerate.
|
|
|
59
59
|
| `/git-deliver` | Single ad-hoc delivery command for working-tree changes. Detects the git setup and escalates to the right terminal step — commit only, commit + push, or commit + push + open a PR with native auto-merge — picking the default from observable state and letting flags pin any level explicitly. Replaces the retired git-commit-all, git-push, and git-pr-all trio. |
|
|
60
60
|
| `/mandrel-deliver` | Unified delivery entry point. Takes Story ids or a plain-language prompt, derives which path the work belongs on, and lands it via the single deliver-story engine — story-<id> → PR → main. |
|
|
61
61
|
| `/mandrel-plan` | Unified planning entry point. Interrogate → author → persist. Emits one Story by default; splits into N>1 only under the default-single split policy. |
|
|
62
|
-
| `/mandrel-update` | npm-era upgrade wraparound for a Mandrel consumer. Runs `npx mandrel update` (resolve newest published version → install → re-materialize `.agents/` → migrate → doctor → surface changelog) as the single mechanical step, then walks the operator through the judgment wraparound the CLI deliberately leaves unowned: reconcile `.agentrc.json`, install the stabilized quality-gate surface, refresh the harness permission allowlist, reconcile the consumer's `AGENTS.md` / runbooks against the surfaced changelog, and stage + commit the
|
|
62
|
+
| `/mandrel-update` | npm-era upgrade wraparound for a Mandrel consumer. Runs `npx mandrel update` (resolve newest published version → install → re-materialize `.agents/` → migrate → doctor → surface changelog) as the single mechanical step, then walks the operator through the judgment wraparound the CLI deliberately leaves unowned: reconcile `.agentrc.json`, install the stabilized quality-gate surface, refresh the harness permission allowlist, reconcile the consumer's `AGENTS.md` / runbooks against the surfaced changelog, and stage + commit the dependency bump the CLI's closing report describes. |
|
|
63
63
|
| `/memory-consolidate` | Attended consolidation pass over this project's agent memory pool — merge duplicates, verify claims against the current tree, prune with operator confirmation, rewrite the index, and stamp the pool with the date and entry count the /mandrel-plan advisory measures its next nudge against. |
|
|
64
64
|
| `/prototype` | Operator-invoked UI prototype pass. Discovers the consumer's design-system SSOT first, then — only after the operator confirms — writes exactly one self-contained HTML file under the gitignored workspace-root temp tree, so a layout can be reviewed before its UI acceptance criteria are authored. |
|
|
65
65
|
| `/qa-assist` | Human-led QA assist loop — set up, then ride a rolling multi-observation intake session. The operator reports observations in any order; the agent enriches each (repro + root-cause file:line + coverage verdict for bugs; analysis + options + recommendation for enhancements), asks clarifying questions only when ambiguous, and appends a redacted ledger item — recording, never planning — to a persistent, resumable session under temp/qa/. Only when the operator says they are done does it review the full ledger and hand off to /mandrel-plan. |
|
package/.agents/instructions.md
CHANGED
|
@@ -176,8 +176,7 @@ reference the Story via `(refs #<storyId>)`. There is no `type::task`.
|
|
|
176
176
|
|
|
177
177
|
`type::epic` is a container only (goal + child checklist, no `agent::*`,
|
|
178
178
|
never delivered): `/mandrel-plan` offers one above 2 Stories and
|
|
179
|
-
`/mandrel-deliver <epicId>` expands it. Linkage is parent→child only
|
|
180
|
-
an `Epic: #N` footer is still refused.
|
|
179
|
+
`/mandrel-deliver <epicId>` expands it. Linkage is parent→child only.
|
|
181
180
|
|
|
182
181
|
---
|
|
183
182
|
|
|
@@ -195,6 +194,6 @@ anything under it.
|
|
|
195
194
|
delivers and self-verifies in one pass** — a broad footprint is normal
|
|
196
195
|
when the change is cohesive, and no plan-time ceiling scores it; do not
|
|
197
196
|
re-slice it into per-module fragments. On an out-of-scope task: **plan
|
|
198
|
-
first**
|
|
199
|
-
|
|
200
|
-
|
|
197
|
+
first** in numbered cohesive sub-steps, **commit incrementally** per
|
|
198
|
+
sub-step, and **fail fast** — STOP and report if any sub-step fails
|
|
199
|
+
validation.
|
|
@@ -96,10 +96,14 @@ the `friction` comment and flips the Story in one call; then hand back to the
|
|
|
96
96
|
operator, who owns the runner pool. Do not sit in a retry loop waiting for
|
|
97
97
|
capacity to return.
|
|
98
98
|
|
|
99
|
-
**
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
capacity
|
|
99
|
+
**A rerun is earned by the filing, never by the hope.** Rerunning a failed job
|
|
100
|
+
to reach green is forbidden — with exactly one exception, and it is gated on
|
|
101
|
+
evidence you have already committed to: once `file-ci-gap.js` has recorded a
|
|
102
|
+
`capacity` or `unreproducible-tier` verdict for the **current head SHA**, you
|
|
103
|
+
may rerun the failed job **once**. Reaching the verdict is not enough; the
|
|
104
|
+
filing is what records the allowance, which is what makes the claim auditable.
|
|
105
|
+
See § One rerun after a recorded verdict below. A capacity-blocked delivery
|
|
106
|
+
that never reaches a filing ends `agent::blocked` — not merged.
|
|
103
107
|
|
|
104
108
|
### The `unreproducible-tier` verdict
|
|
105
109
|
|
|
@@ -129,10 +133,36 @@ before. On the verdict: run
|
|
|
129
133
|
as `--evidence`; then hand back to the operator, who owns the sandbox. Do not author a fix for a tier you could not
|
|
130
134
|
run — a blind fix to a suite nobody exercised is how the gap compounds.
|
|
131
135
|
|
|
136
|
+
## One rerun after a recorded verdict
|
|
137
|
+
|
|
138
|
+
`capacity` and `unreproducible-tier` name failures that are proven properties
|
|
139
|
+
of the **environment**: no commit on the branch can move the head SHA to clear
|
|
140
|
+
them, so the no-rerun rule used to strand a correct delivery until a human
|
|
141
|
+
cleared it by hand. Those two verdicts — and only those two — now buy exactly
|
|
142
|
+
one rerun:
|
|
143
|
+
|
|
144
|
+
1. Reach the verdict with its required readings (§ above). A green on re-run is
|
|
145
|
+
never one of those readings.
|
|
146
|
+
2. File it: `node .agents/scripts/file-ci-gap.js --story <id> --verdict
|
|
147
|
+
<capacity|unreproducible-tier> --owner <bucket> --evidence "<proof reading>"`.
|
|
148
|
+
The filing stamps a `rerunAllowance` on the CI digest, keyed to the head SHA
|
|
149
|
+
the red was observed on.
|
|
150
|
+
3. Rerun the failed job **once**. The watcher's same-SHA guard admits that one
|
|
151
|
+
green, retires the digest and lets the delivery proceed.
|
|
152
|
+
|
|
153
|
+
The allowance is spent when it is honoured — it dies with the digest — so a
|
|
154
|
+
second red after the rerun writes a fresh digest carrying none, and is a real
|
|
155
|
+
red that routes to **Option 1**. An allowance recorded against a different head
|
|
156
|
+
SHA is not an allowance. `pre-existing` earns none: it reproduces on `main`, so
|
|
157
|
+
it names a real defect someone owns and a rerun cannot remove it. Skipping,
|
|
158
|
+
quarantining or loosening a test to reach green stays forbidden under every
|
|
159
|
+
verdict, always.
|
|
160
|
+
|
|
132
161
|
## Verifier
|
|
133
162
|
|
|
134
|
-
The check is resolved only when it is **green with
|
|
135
|
-
|
|
163
|
+
The check is resolved only when it is **green with no rerun of the failed job**
|
|
164
|
+
— save the one a recorded `capacity` / `unreproducible-tier` verdict buys (§
|
|
165
|
+
above) — and the diff carries **no `.skip` / `.only`, no quarantine, and no
|
|
136
166
|
deleted or loosened assertion** introduced to reach green. You may **not**
|
|
137
167
|
re-run a failed job to "see if it goes green," and you may **not** skip,
|
|
138
168
|
`.only`, or quarantine a flaky test to get a green bar. Both mask the defect
|
|
@@ -146,8 +176,11 @@ first red the watcher **disarms native auto-merge** (a disarm failure is a
|
|
|
146
176
|
blocker, not a warning) and records the PR **head SHA** in the digest
|
|
147
177
|
alongside the failing check-run identity. On green it adjudicates:
|
|
148
178
|
|
|
149
|
-
- **Same head SHA** → the
|
|
150
|
-
|
|
179
|
+
- **Same head SHA, with a recorded allowance** → the one rerun the section
|
|
180
|
+
above buys. The digest is retired (spending the allowance), auto-merge is
|
|
181
|
+
re-armed, and the delivery continues.
|
|
182
|
+
- **Same head SHA, no allowance** → the green came from re-running the failed
|
|
183
|
+
job. The watcher exits non-zero, flips the Story to `agent::blocked` with a
|
|
151
184
|
`friction` comment, and requires the CI-gap intake issue
|
|
152
185
|
(`file-ci-gap.js` — run link and failure signature are already in the
|
|
153
186
|
digest) before the delivery proceeds.
|