mandrel 2.56.0 → 2.57.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/agents/plan-critic.md +13 -18
- package/.agents/agents/story-worker.md +25 -34
- package/.agents/docs/agentrc-reference.json +0 -30
- package/.agents/docs/configuration.md +8 -28
- package/.agents/docs/execution-reference.md +5 -5
- package/.agents/docs/quality-gates.md +8 -7
- package/.agents/instructions.md +9 -10
- package/.agents/schemas/agentrc.schema.json +9 -185
- package/.agents/schemas/story-deliver-terminal.schema.json +1 -1
- package/.agents/scripts/acceptance-eval.js +107 -17
- package/.agents/scripts/ceremony-derive.js +191 -0
- package/.agents/scripts/check-context-budget.js +28 -33
- package/.agents/scripts/check-cyclomatic.js +4 -3
- package/.agents/scripts/deliver-light.js +31 -94
- package/.agents/scripts/lib/audit-suite/checklist-threading.js +15 -2
- package/.agents/scripts/lib/baselines/coverage-updater-cli.js +110 -0
- package/.agents/scripts/lib/baselines/crap-preview-scan.js +25 -0
- package/.agents/scripts/lib/baselines/crap-updater-cli.js +223 -0
- package/.agents/scripts/lib/bdd-scenario-budget.js +21 -3
- package/.agents/scripts/lib/bootstrap/quality-bootstrap.js +0 -1
- package/.agents/scripts/lib/close-validation/gates.js +52 -1
- package/.agents/scripts/lib/config/acceptance-eval.js +25 -57
- package/.agents/scripts/lib/config/delivery-routing.js +7 -33
- package/.agents/scripts/lib/config/explain.js +0 -19
- package/.agents/scripts/lib/config/limits.js +18 -78
- package/.agents/scripts/lib/config/quality.js +6 -3
- package/.agents/scripts/lib/config/runners.js +3 -2
- package/.agents/scripts/lib/config-settings-schema-delivery.js +15 -68
- package/.agents/scripts/lib/config-settings-schema-quality.js +0 -14
- package/.agents/scripts/lib/config-settings-schema.js +16 -143
- package/.agents/scripts/lib/crap-engine.js +35 -4
- package/.agents/scripts/lib/crap-utils.js +17 -1
- package/.agents/scripts/lib/cyclomatic-ceiling.js +19 -7
- package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
- package/.agents/scripts/lib/observability/runtime-friction.js +1 -1
- package/.agents/scripts/lib/observability/source-classifier.js +1 -0
- package/.agents/scripts/lib/orchestration/acceptance-eval-decision.js +5 -4
- package/.agents/scripts/lib/orchestration/ceremony-routing.js +19 -73
- package/.agents/scripts/lib/orchestration/complexity-gate.js +46 -212
- package/.agents/scripts/lib/orchestration/file-assumptions.js +32 -17
- package/.agents/scripts/lib/orchestration/light-escalation.js +3 -3
- package/.agents/scripts/lib/orchestration/light-suitability.js +66 -233
- package/.agents/scripts/lib/orchestration/plan-context.js +181 -387
- package/.agents/scripts/lib/orchestration/plan-critic-conditions.js +42 -153
- package/.agents/scripts/lib/orchestration/plan-critics-evaluate.js +14 -70
- package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +300 -0
- package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +131 -168
- package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +118 -297
- package/.agents/scripts/lib/orchestration/plan-persist/soft-findings.js +55 -0
- package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +16 -65
- package/.agents/scripts/lib/orchestration/plan-persist/wave-serialisation.js +22 -35
- package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +30 -139
- package/.agents/scripts/lib/orchestration/planning/memory-pool-advisory.js +61 -223
- package/.agents/scripts/lib/orchestration/single-story-close/phases/close-validation.js +5 -0
- package/.agents/scripts/lib/orchestration/single-story-close/phases/pre-gate-steps.js +46 -16
- package/.agents/scripts/lib/orchestration/story-close/context-budget-writeback.js +213 -0
- package/.agents/scripts/lib/orchestration/task-body-validator.js +10 -63
- package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +33 -539
- package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +21 -414
- package/.agents/scripts/lib/orchestration/ticket-validator.js +54 -118
- package/.agents/scripts/lib/orchestration/verify-credit.js +69 -24
- package/.agents/scripts/lib/story-body/body-format-lints.js +15 -85
- package/.agents/scripts/lib/story-body/story-body.js +17 -237
- package/.agents/scripts/lib/templates/decomposer-prompts.js +84 -121
- package/.agents/scripts/lib/test-isolate/cli-options.js +93 -0
- package/.agents/scripts/lib/test-isolate/progress-log.js +45 -0
- package/.agents/scripts/lib/test-isolate/render-report.js +97 -0
- package/.agents/scripts/lib/test-isolate/run-isolate.js +87 -0
- package/.agents/scripts/lib/test-run-credit.js +266 -0
- package/.agents/scripts/lib/wave-runner/footprint.js +48 -358
- package/.agents/scripts/lib/wave-runner/ready-set.js +6 -5
- package/.agents/scripts/lib/workers/crap-worker.js +32 -41
- package/.agents/scripts/plan-context.js +7 -9
- package/.agents/scripts/plan-critics.js +28 -54
- package/.agents/scripts/plan-persist.js +25 -68
- package/.agents/scripts/quality-preview.js +51 -0
- package/.agents/scripts/run-tests.js +12 -0
- package/.agents/scripts/stories-wave-tick.js +23 -45
- package/.agents/scripts/test-isolate.js +13 -180
- package/.agents/scripts/update-coverage-baseline.js +25 -70
- package/.agents/scripts/update-crap-baseline.js +19 -123
- package/.agents/skills/core/scope-triage/SKILL.md +3 -3
- package/.agents/workflows/audit-clean-code.md +4 -3
- package/.agents/workflows/helpers/acceptance-self-eval.md +41 -41
- package/.agents/workflows/helpers/code-quality-guardrails.md +4 -4
- package/.agents/workflows/helpers/code-review.md +2 -3
- package/.agents/workflows/helpers/deliver-digest.md +41 -57
- package/.agents/workflows/helpers/deliver-light.md +40 -105
- package/.agents/workflows/helpers/deliver-reference.md +1 -1
- package/.agents/workflows/helpers/deliver-story-reference.md +37 -58
- package/.agents/workflows/helpers/deliver-story.md +9 -13
- package/.agents/workflows/helpers/plan-reference.md +132 -219
- package/.agents/workflows/mandrel-plan.md +27 -40
- package/.agents/workflows/memory-consolidate.md +9 -13
- package/docs/CHANGELOG.md +23 -0
- package/lib/migrations/index.js +4 -0
- package/lib/migrations/steps/2.57.0-retire-delivery-limit-knobs.js +45 -0
- package/lib/migrations/steps/2.57.0-retire-planning-limit-knobs.js +59 -0
- package/package.json +1 -1
- package/.agents/scripts/lib/framework-version.js +0 -39
- package/.agents/scripts/lib/orchestration/consolidation-precondition.js +0 -223
- package/.agents/scripts/lib/orchestration/plan-persist/fan-out-gate.js +0 -97
- package/.agents/scripts/lib/orchestration/planning/decomposer-context.js +0 -26
- package/.agents/scripts/lib/orchestration/spec-budget.js +0 -89
- package/.agents/scripts/lib/orchestration/spec-spill.js +0 -74
- package/.agents/scripts/lib/orchestration/verify-tier-repair.js +0 -107
|
@@ -140,10 +140,8 @@ full-ceremony Story. The lite route's `preserves` field is the machine-readable
|
|
|
140
140
|
record of those non-negotiables; there is no lite-specific gate bypass.
|
|
141
141
|
|
|
142
142
|
**Ceremony comes from the landed diff; the dispatch mode comes from the
|
|
143
|
-
run.** Persist stamps
|
|
144
|
-
|
|
145
|
-
plus per-Story shape evidence — on the `story-plan-state` checkpoint); the
|
|
146
|
-
label is never the control signal. Ceremony is resolved from the **derived
|
|
143
|
+
run.** Persist stamps no route label (Story #5312 retired the plan-side lite
|
|
144
|
+
claim with its `route::lite` hint). Ceremony is resolved from the **derived
|
|
147
145
|
change level** (`deriveChangeLevel` over the computed change set — digest § 3),
|
|
148
146
|
not from a body-shape read: a footprint intersecting a sensitive-path class
|
|
149
147
|
derives `high`, so the Story keeps its fresh acceptance critic. The light path
|
|
@@ -240,35 +238,19 @@ discovers them only after the whole close pipeline has run, at several times
|
|
|
240
238
|
the cost of one full-suite run in the worktree.
|
|
241
239
|
|
|
242
240
|
**Run it once, last, so close can credit it.** The run belongs **after** the
|
|
243
|
-
self-eval loop's last fix commit
|
|
244
|
-
|
|
245
|
-
|
|
246
|
-
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
|
|
252
|
-
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
# be ABSOLUTE and the runner exactly `npm test`: both sides hash
|
|
257
|
-
# {cmd, args, cwd}, so a relative path or a wrapper misses the credit.
|
|
258
|
-
node <main-repo>/.agents/scripts/evidence-gate.js --standalone \
|
|
259
|
-
--scope-id <storyId> --gate test --worktree <workCwd> -- npm test
|
|
260
|
-
```
|
|
261
|
-
|
|
262
|
-
The credit expires the moment it stops describing the tree: evidence is keyed
|
|
263
|
-
on HEAD, the capture stamp on a content digest of `crap.targetDirs`. A
|
|
264
|
-
self-eval fix — or any commit — invalidates it and close re-runs the suite for
|
|
265
|
-
real, so this never trades away the gate. That keying is exactly why the run
|
|
266
|
-
comes last, and why the push before it is free. Close's own base-sync can
|
|
267
|
-
spend the stamp too when it lands base commits; it now says so out loud rather
|
|
268
|
-
than silently re-running the suite.
|
|
269
|
-
|
|
270
|
-
**`verify[]` reuses the same stamp.** A `verify[]` entry that is itself a
|
|
271
|
-
full-suite command is reported **credited** against that stamp rather than
|
|
241
|
+
self-eval loop's last fix commit; redraft rounds run scoped tests. A green
|
|
242
|
+
`npm test` in the worktree on `story-<id>` deposits the `test` evidence
|
|
243
|
+
record close reads (Story #5313 — `lib/test-run-credit.js`), keyed on HEAD
|
|
244
|
+
and the tree fingerprint and hashed on the exact command close spawns, so
|
|
245
|
+
close reports the gate as credited at unchanged HEAD instead of re-running
|
|
246
|
+
the suite. The credit expires the moment it stops describing the tree: any
|
|
247
|
+
later commit invalidates it and close re-runs the suite for real, so this
|
|
248
|
+
never trades away the gate. The CRAP gate still runs `coverage-capture.js`
|
|
249
|
+
itself when it needs a fresh artifact — the capture stamp is a claim about
|
|
250
|
+
`coverage/coverage-final.json`, which a bare `npm test` does not produce.
|
|
251
|
+
|
|
252
|
+
**`verify[]` reuses the same credit.** A `verify[]` entry that is itself a
|
|
253
|
+
full-suite command is reported **credited** against that record rather than
|
|
272
254
|
respawned (`resolveVerifyCredit` in
|
|
273
255
|
[`verify-credit.js`](../../scripts/lib/orchestration/verify-credit.js)), and the
|
|
274
256
|
self-eval gate warns when it sees one: the intended shape is scoped `verify[]`
|
|
@@ -294,24 +276,23 @@ round cap, proceed / redraft / block — not an independent additional pass
|
|
|
294
276
|
over the criteria. The M4-B floor holds: one verdict per cluster, with the
|
|
295
277
|
cluster count owned by the dispatching caller and never by routing.
|
|
296
278
|
|
|
297
|
-
**One round = N cluster critics → ONE merged verdict → ONE gate call
|
|
279
|
+
**One round = N cluster critics → ONE merged verdict → ONE gate call** (fresh
|
|
280
|
+
critics; an inline-owned verdict is one file scored in one call). The
|
|
298
281
|
clusters are how a round is _authored_; they are not how it is _scored_.
|
|
299
282
|
Concatenate every cluster's records into a single `criteria[]` ordered by
|
|
300
283
|
`index` — exactly one per `acceptance[]` item, under one `storyId`,
|
|
301
284
|
`schemaVersion`, `round` and `commitSha` — and hand that merged file to the
|
|
302
|
-
gate once
|
|
285
|
+
gate once:
|
|
303
286
|
|
|
304
287
|
```bash
|
|
305
288
|
node <main-repo>/.agents/scripts/acceptance-eval.js \
|
|
306
|
-
--story <storyId> --verdict <merged-verdict-path>
|
|
307
|
-
--expected-criteria <acceptance[] count>
|
|
289
|
+
--story <storyId> --verdict <merged-verdict-path>
|
|
308
290
|
```
|
|
309
291
|
|
|
310
|
-
The
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
consumes no round. Calling the gate once per cluster instead spends a round
|
|
292
|
+
The gate reads the Story's `acceptance[]` count itself (Story #5313), so a
|
|
293
|
+
single cluster's verdict handed over unmerged is rejected before scoring and
|
|
294
|
+
consumes no round; `--expected-criteria` is accepted but redundant. Calling
|
|
295
|
+
the gate once per cluster instead spends a round
|
|
315
296
|
_per cluster_ — a Story past the cluster ceiling would burn its whole redraft
|
|
316
297
|
budget on cluster arithmetic — and N concurrent calls race the Story-scoped
|
|
317
298
|
round ledger. Full per-round mechanics, including the parallel dispatch and the
|
|
@@ -342,32 +323,30 @@ node .agents/scripts/update-ticket-state.js --ticket <storyId> --state agent::bl
|
|
|
342
323
|
|
|
343
324
|
## Step 2 — Ceremony detail
|
|
344
325
|
|
|
345
|
-
**Compute the change set once** with the shared enumerator —
|
|
346
|
-
|
|
326
|
+
**Compute the change set once** with the shared enumerator — the same module
|
|
327
|
+
close uses — and reuse that one list downstream. `ceremony-derive.js`
|
|
328
|
+
(Story #5313) is that enumeration, the level derivation and the ceremony
|
|
329
|
+
resolution in one call:
|
|
347
330
|
|
|
348
331
|
```bash
|
|
349
|
-
node
|
|
350
|
-
import { computeChangeSet } from "<main-repo>/.agents/scripts/lib/orchestration/change-set.js";
|
|
351
|
-
const { files } = computeChangeSet({ baseRef: "main", headRef: "story-<storyId>" });
|
|
352
|
-
console.log(JSON.stringify(files));
|
|
353
|
-
'
|
|
332
|
+
node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
|
|
354
333
|
```
|
|
355
334
|
|
|
356
|
-
|
|
335
|
+
Its `level` comes from
|
|
357
336
|
[`deriveChangeLevel`](../../scripts/lib/orchestration/review-depth.js) over
|
|
358
337
|
the one computed change-set list: a diff touching a sensitive path registered
|
|
359
338
|
in `.agents/schemas/audit-rules.json` derives `high`, one touching none
|
|
360
|
-
derives `low`, and an unenumerable diff (`files
|
|
361
|
-
Hand the **same** list to every acceptance critic you spawn (Step 1a)
|
|
362
|
-
critic that re-ran its own `git diff` could score against a different
|
|
363
|
-
than the one that routed it.
|
|
339
|
+
derives `low`, and an unenumerable diff (`files: null`) derives `null`.
|
|
340
|
+
Hand the **same** `files` list to every acceptance critic you spawn (Step 1a)
|
|
341
|
+
— a critic that re-ran its own `git diff` could score against a different
|
|
342
|
+
set than the one that routed it.
|
|
364
343
|
|
|
365
|
-
|
|
344
|
+
Its `mode` / `verdictOwner` come from
|
|
366
345
|
[`resolveCeremonyForRisk`](../../scripts/lib/orchestration/ceremony-routing.js)
|
|
367
346
|
(`minimal` → always inline; `strict` → always fresh; `standard` →
|
|
368
|
-
`high`/`null` → `fresh`, `low` → `inline`
|
|
369
|
-
|
|
370
|
-
|
|
347
|
+
`high`/`null` → `fresh`, `low` → `inline`). Review depth reads the same
|
|
348
|
+
derived level via `review-depth.js` inside close, so the two decisions cannot
|
|
349
|
+
disagree.
|
|
371
350
|
|
|
372
351
|
**Inline-dispatch override.** When the Story dispatches
|
|
373
352
|
`inline` (`resolveStoryDispatchMode` → `inline`, which is exactly a
|
|
@@ -74,7 +74,7 @@ One branch, one PR to `main`, commits against the inline `acceptance[]` /
|
|
|
74
74
|
`## Slicing` rows as **intra-session checkpoints** (reference § Step 1).
|
|
75
75
|
2. Implement and commit on the Story branch, iterating with quick advisory
|
|
76
76
|
gates (`typecheck`, `lint`, scoped tests) — the full chain runs in Step 3,
|
|
77
|
-
and the **one**
|
|
77
|
+
and the **one** full-suite run at Step 2.5.
|
|
78
78
|
|
|
79
79
|
### Step 1a — Bounded acceptance self-eval loop (**required**)
|
|
80
80
|
|
|
@@ -93,21 +93,17 @@ are reference § Step 2. Hard gates always run in Step 3 — the derived level
|
|
|
93
93
|
never disables them; do **not** pre-run the chain here — Step 2.5's credited
|
|
94
94
|
suite run is the sole exception.
|
|
95
95
|
|
|
96
|
-
### Step 2.5 —
|
|
96
|
+
### Step 2.5 — The one full-suite run, the push, then hand off
|
|
97
97
|
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
98
|
+
After the self-eval loop's last fix commit, run `npm test` **once** in the
|
|
99
|
+
worktree (**digest § 5**): a green full run deposits the `test` credit close
|
|
100
|
+
reads, keyed on the tree, so only a *later* commit invalidates it. Red →
|
|
101
|
+
fix, commit, re-run.
|
|
101
102
|
|
|
102
|
-
|
|
103
|
-
(
|
|
104
|
-
a *later* commit invalidates it; a bare `npm test` deposits none. Red →
|
|
105
|
-
fix, commit, push, re-capture.
|
|
106
|
-
|
|
107
|
-
Then (sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
|
|
103
|
+
Push `story-<storyId>` to `origin`, confirming the remote ref moved. Then
|
|
104
|
+
(sub-agent dispatch only) return the hand-off — Story id, `workCwd`,
|
|
108
105
|
branch, pushed head SHA, self-eval verdict, `verify[]` evidence — and stop.
|
|
109
|
-
Do not open the PR or compose a terminal envelope.
|
|
110
|
-
before Step 3.
|
|
106
|
+
Do not open the PR or compose a terminal envelope.
|
|
111
107
|
|
|
112
108
|
## Step 3 — Close and land (`single-story-close.js`)
|
|
113
109
|
|
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
# /mandrel-plan — on-demand reference appendix
|
|
2
2
|
|
|
3
3
|
> **Applies when:** you are executing [`/mandrel-plan`](../mandrel-plan.md) and hit one of the
|
|
4
|
-
> situations below — input-mode derivation, the Gate #1
|
|
5
|
-
>
|
|
6
|
-
>
|
|
7
|
-
>
|
|
4
|
+
> situations below — input-mode derivation, the Gate #1 advisory line,
|
|
5
|
+
> tickets-mode supersede authoring, the operator-invoked pre-mortem, a
|
|
6
|
+
> failed persist, or source-id resolution. The spine stays resident; this
|
|
7
|
+
> file is read on demand.
|
|
8
8
|
|
|
9
9
|
## Deriving the input mode
|
|
10
10
|
|
|
@@ -92,140 +92,58 @@ The marker keeps the operator's undelegated decisions findable after the
|
|
|
92
92
|
fact: reviewing a `--yes` plan means scanning its decisions-made-by-default,
|
|
93
93
|
not re-deriving which assumptions were really the agent's to make.
|
|
94
94
|
|
|
95
|
-
## Gate #1 → the
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
`
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
**Under `--yes` the offer is recorded and planning proceeds** — it is *never*
|
|
148
|
-
auto-downgraded to light. An unattended run has nobody to confirm the reroute,
|
|
149
|
-
and a suggestion is not a confirmation. The same rule governs unknown triage
|
|
150
|
-
unattended: AFK unknowns are still researched, but no free-form operator
|
|
151
|
-
question is asked — each HITL unknown lands in Key Assumptions marked a
|
|
152
|
-
decision-made-by-default, so the record shows what was decided for the operator
|
|
153
|
-
rather than pretending it was decided with them.
|
|
154
|
-
|
|
155
|
-
Escalation in the *other* direction — an over-scope prompt on the light path —
|
|
156
|
-
is terminal and requires a fresh session. The rule that separates the two, and
|
|
157
|
-
why it must not be flattened into symmetry:
|
|
158
|
-
[`deliver-light.md` § Why the two directions differ](deliver-light.md).
|
|
159
|
-
|
|
160
|
-
## Gate #1 → the `/prototype` offer (`uiSurface`)
|
|
161
|
-
|
|
162
|
-
`complexitySignals.uiSurface` is the second advisory Gate #1 offer, and the
|
|
163
|
-
weaker of the two on purpose: it carries **no routing authority and adds no
|
|
164
|
-
gate**. Both halves are derived from observables already in the checkout — the
|
|
165
|
-
`hasWebSurface` applicability predicate the `target: "web"` audit lenses gate
|
|
166
|
-
on, and whether any predicted path matches a web lens `filePattern` registered
|
|
167
|
-
in `audit-rules.json`. There is no configuration key to set: a project with no
|
|
168
|
-
rendered frontend resolves falsey and the offer never fires.
|
|
169
|
-
|
|
170
|
-
When it does fire, **name [`/prototype`](../prototype.md) and stop there.**
|
|
171
|
-
`/mandrel-plan` must never invoke it — operator invocation is the entire design, because
|
|
172
|
-
the value is a human looking at a layout before its UI acceptance criteria are
|
|
173
|
-
frozen.
|
|
174
|
-
|
|
175
|
-
**Under `--yes` the offer is recorded and planning proceeds** — no reroute, no
|
|
176
|
-
prototype written, no gate raised. This is exactly how `deliverLightSuggestion`
|
|
177
|
-
behaves unattended, and for the same reason: an unattended run has nobody to
|
|
178
|
-
review an artifact, so recording the offer is the whole of the right behaviour.
|
|
179
|
-
|
|
180
|
-
## Shape-derived complexity routing (`complexitySignals`)
|
|
181
|
-
|
|
182
|
-
Complexity routes on the **objective shape of the authored work**, never on
|
|
183
|
-
seed word count — a detailed prompt can describe trivial work, a terse one
|
|
184
|
-
complex work. The pipeline stages the
|
|
185
|
-
decision:
|
|
186
|
-
|
|
187
|
-
- **Signals, not routing.** The envelope's `complexitySignals` field is
|
|
188
|
-
advisory only (`routingAuthority: false`): enumerated-artifact count (with
|
|
189
|
-
the configured `maxArtifacts` threshold beside it as one input),
|
|
190
|
-
`planning.riskHeuristics` phrases present in the seed, the repo state of
|
|
191
|
-
predicted paths (existing paths predict refactors; missing predict
|
|
192
|
-
creates), and the `audit-rules.json` sensitive-path classes the predicted
|
|
193
|
-
footprint intersects.
|
|
194
|
-
- **You author the verdict.** Judge the signals: a genuinely trivial scope
|
|
195
|
-
(small additive footprint, no risk hits, no sensitive class) earns a `lite`
|
|
196
|
-
claim via `plan-persist.js --route-downgrade-reason "<why>"`. The reason is
|
|
197
|
-
recorded on every created Story's `story-plan-state` checkpoint, making the
|
|
198
|
-
judgment auditable; without a recorded reason the conservative default
|
|
199
|
-
(`full`) stands.
|
|
200
|
-
- **Persist backstops the claim deterministically.** After authoring, the
|
|
201
|
-
work has measurable shape, so persist validates the `lite` claim against
|
|
202
|
-
each Story's own shape — distinct change kinds, declared magnitude,
|
|
203
|
-
uncertainty, deployable/migration span, glob-free footprint, and
|
|
204
|
-
sensitive-path classes, against the framework `STORY_SHAPE_CEILINGS` (effort
|
|
205
|
-
and risk, never artifact counts) — and **fails closed to
|
|
206
|
-
`full`** when any Story exceeds them (the refusal is ledgered on the
|
|
207
|
-
checkpoint too). The lite route is **not** licence to drop a
|
|
208
|
-
non-negotiable — every decision's `preserves` field enumerates what still
|
|
209
|
-
holds: the Story ticket, the PR-to-`main` landing, every repo quality gate,
|
|
210
|
-
and the security baseline. Those gates run in `single-story-close.js`
|
|
211
|
-
regardless of route.
|
|
212
|
-
|
|
213
|
-
**The label is a hint; deliver re-derives.** Persist labels a
|
|
214
|
-
lite cohort's Stories with **`route::lite`** as a *human-visible hint only* —
|
|
215
|
-
`/mandrel-deliver` computes the route from each fetched Story body via the same shape
|
|
216
|
-
function at dispatch, so neither a lost label nor an unread marker can
|
|
217
|
-
misroute delivery: a lite-shaped Story derives `lite` even with the label
|
|
218
|
-
absent, and a sensitive-footprint Story routes `full` and keeps its fresh
|
|
219
|
-
critic even with the label present. The derived route sets ceremony, not
|
|
220
|
-
where the engine runs — sub-agent boots are collapsed by a **single-Story
|
|
221
|
-
run**, never by a trivial shape. The `route::*` axis stays runtime-derived: hand-authored
|
|
222
|
-
`route::*` entries in `labels[]` are dropped by persist.
|
|
223
|
-
|
|
224
|
-
The knobs (`planning.complexityGate.{enabled, maxArtifacts}`) are documented
|
|
225
|
-
in [`.agents/docs/configuration.md`](../../docs/configuration.md) under
|
|
226
|
-
`### planning`; the defaults live on `DEFAULT_COMPLEXITY_GATE` and the shape
|
|
227
|
-
ceilings on `STORY_SHAPE_CEILINGS` in
|
|
228
|
-
[`lib/orchestration/complexity-gate.js`](../../scripts/lib/orchestration/complexity-gate.js).
|
|
95
|
+
## Gate #1 → the one advisory line
|
|
96
|
+
|
|
97
|
+
Gate #1 stops for exactly two things — the sharpened plan intent and any HITL
|
|
98
|
+
unknown — and everything else the envelope surfaced collapses to **one
|
|
99
|
+
advisory line** beneath it (Story #5312). Nothing on that line stops the run,
|
|
100
|
+
reroutes it, or is invoked by `/mandrel-plan`; each item names something the
|
|
101
|
+
operator may prefer to do instead, and the run proceeds either way. Under
|
|
102
|
+
`--yes` the line is recorded and planning continues — an unattended run has
|
|
103
|
+
nobody to take an offer.
|
|
104
|
+
|
|
105
|
+
The line names, in order, whichever of these the envelope carries:
|
|
106
|
+
|
|
107
|
+
- **`duplicates[]`** — open Stories the seed resembles (never Epics). Name
|
|
108
|
+
the top one or two by id and title; a plan that duplicates open work is
|
|
109
|
+
still the operator's call.
|
|
110
|
+
- **Open `intake` rows** (`priorFeedback`) — CI-gap intake filings written by
|
|
111
|
+
[`file-ci-gap.js`](../../scripts/file-ci-gap.js) when a delivery reached an
|
|
112
|
+
Option-2 verdict in [`ci-remediation.md`](../../rules/ci-remediation.md).
|
|
113
|
+
They carry evidence but no `## Spec`, no `acceptance[]` / `verify[]` and no
|
|
114
|
+
`agent::*` label, so `/mandrel-deliver` cannot take one: graduating it is
|
|
115
|
+
exactly **tickets mode** (`/mandrel-plan <issue number>`), and a filing
|
|
116
|
+
that keeps recurring (its `## Occurrences` table is the count) is often the
|
|
117
|
+
better next Story than the seed in front of you. A `platformGaps[]` row is
|
|
118
|
+
the same shape with a different owner.
|
|
119
|
+
- **`memoryPoolAdvisory.recommend`** — name
|
|
120
|
+
[`/memory-consolidate`](../memory-consolidate.md), quoting its
|
|
121
|
+
`reasons[]`. The one arm left measures the `MEMORY.md` index against the
|
|
122
|
+
harness's byte cap; a stale pool degrades recall, it does not make the plan
|
|
123
|
+
wrong.
|
|
124
|
+
- **`complexitySignals.uiSurface`** — name [`/prototype`](../prototype.md)
|
|
125
|
+
and stop there. The signal carries **no routing authority and adds no
|
|
126
|
+
gate**: both halves are derived from observables already in the checkout —
|
|
127
|
+
the `hasWebSurface` applicability predicate the `target: "web"` audit
|
|
128
|
+
lenses gate on, and whether any predicted path matches a web lens
|
|
129
|
+
`filePattern` in `audit-rules.json`; a project with no rendered frontend
|
|
130
|
+
resolves falsey and the offer never fires. `/mandrel-plan` must never invoke
|
|
131
|
+
it — operator invocation is the entire design, because the value is a human
|
|
132
|
+
looking at a layout before its UI acceptance criteria are frozen. Under
|
|
133
|
+
`--yes` the offer is recorded and planning proceeds — no reroute, no
|
|
134
|
+
prototype written, no gate raised.
|
|
135
|
+
|
|
136
|
+
## `complexitySignals` are advisory
|
|
137
|
+
|
|
138
|
+
The envelope's `complexitySignals` field carries the paths the seed predicts,
|
|
139
|
+
their repo state (existing paths predict refactors; missing predict creates)
|
|
140
|
+
and the `audit-rules.json` sensitive-path classes the footprint intersects —
|
|
141
|
+
`routingAuthority: false`, no `route` field. They ground the authoring
|
|
142
|
+
template's pre-resolved `changes[]` and the `/prototype` offer, nothing else.
|
|
143
|
+
Story #5312 deleted the plan-side lite claim that used to read them
|
|
144
|
+
(`--route-downgrade-reason`, the persist shape backstop, the `route::lite`
|
|
145
|
+
hint): every Story lands through the same engine and the same close gates,
|
|
146
|
+
and ceremony is derived from the landed diff at close.
|
|
229
147
|
|
|
230
148
|
## Correct-by-construction authoring template
|
|
231
149
|
|
|
@@ -233,30 +151,24 @@ ceilings on `STORY_SHAPE_CEILINGS` in
|
|
|
233
151
|
**correct-by-construction** skeleton, built from the same repo probe the
|
|
234
152
|
`complexitySignals` ran:
|
|
235
153
|
|
|
236
|
-
- **`verify[]`
|
|
237
|
-
|
|
238
|
-
|
|
239
|
-
entry is exactly the mechanical persist round-trip the template exists to
|
|
240
|
-
prevent.
|
|
154
|
+
- **`verify[]` entries are commands.** There is no tier suffix and no
|
|
155
|
+
`manual:<reason>` escape (Story #5312): write the exact command or test
|
|
156
|
+
path the deliverer runs and the acceptance critic reads as evidence.
|
|
241
157
|
- **`changes[]` arrive pre-resolved to creates-vs-refactors.** Every path
|
|
242
158
|
the seed predicted is probed against the repo: an existing path is
|
|
243
159
|
emitted with `assumption: "refactors-existing"`, a missing one with
|
|
244
|
-
`assumption: "creates"`.
|
|
245
|
-
|
|
246
|
-
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
|
|
252
|
-
|
|
253
|
-
actually warns, so the warning marks a real outlier instead of ordinary
|
|
254
|
-
variance. Neither fails the persist. The hard fail-closed ceiling
|
|
255
|
-
(~1500 tokens, `spec-spill.js`) is unchanged.
|
|
160
|
+
`assumption: "creates"`. The persist gates stay authoritative — they probe
|
|
161
|
+
the base branch ref, not the working tree — but a `creates` on a path that
|
|
162
|
+
exists at base, or a `refactors-existing` on one that does not, is a
|
|
163
|
+
dry-run **warning**, not a rejection; only a `deletes` naming an absent
|
|
164
|
+
path is refused. A plain-string bullet or a trailing parenthetical is
|
|
165
|
+
repaired into the object form by probing base, and the repair is reported.
|
|
166
|
+
- **Keep `## Spec` at contract-level prose** — interfaces, invariants,
|
|
167
|
+
load-bearing constraints; no per-file behavior narration — and as long as
|
|
168
|
+
the work needs. There is no word or token budget.
|
|
256
169
|
|
|
257
170
|
A faithfully-filled skeleton — placeholders replaced, pre-resolved entries
|
|
258
|
-
kept
|
|
259
|
-
round-trip.
|
|
171
|
+
kept — passes the persist ticket validators with no round-trip.
|
|
260
172
|
|
|
261
173
|
### Authored entry shape
|
|
262
174
|
|
|
@@ -336,18 +248,13 @@ together, and a path two same-wave Stories both write is exactly where that
|
|
|
336
248
|
promise breaks. Promise and caveat belong on one durable surface — previously
|
|
337
249
|
the caveat was a stderr warning nobody kept.
|
|
338
250
|
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
|
|
342
|
-
|
|
343
|
-
|
|
344
|
-
|
|
345
|
-
|
|
346
|
-
|
|
347
|
-
Turn one on for a repo where the class is genuinely fatal; expect a refusal to
|
|
348
|
-
name the Stories and the fix (a `depends_on` edge, or folding the shared edit
|
|
349
|
-
into one Story). The sibling knobs `failOnRegistryConflicts`,
|
|
350
|
-
`failOnMissingBddScaffold` and `failOnLargeFanOut` behave the same way.
|
|
251
|
+
Every conflict class is advisory (Story #5312 retired the
|
|
252
|
+
`planning.failOn*` / `requireExplicitCrossStoryDeps` upgrade knobs with the
|
|
253
|
+
registry and fan-out findings): co-editing one file is routine and often
|
|
254
|
+
correct — the delivery scheduler already serializes file-overlapping Stories —
|
|
255
|
+
and a path reference matched by substring can read as a dependency a prose
|
|
256
|
+
mention never meant. A finding names the Stories and the fix (a `depends_on`
|
|
257
|
+
edge, or folding the shared edit into one Story) for the operator to weigh.
|
|
351
258
|
|
|
352
259
|
## Tickets mode — authoring `supersedes[]`
|
|
353
260
|
|
|
@@ -382,67 +289,73 @@ total by default — an authored map is the only thing that can say
|
|
|
382
289
|
`#11-#14 → #20` while `#15 → #21`, which a blanket "superseded by
|
|
383
290
|
this plan-run" reference could not.
|
|
384
291
|
|
|
385
|
-
##
|
|
292
|
+
## The pre-mortem critic — operator-invoked
|
|
293
|
+
|
|
294
|
+
The maker-blind **pre-mortem** critic is not a step of the spine (Story #5312
|
|
295
|
+
retired step 2.5 with the consolidation critic, whose one deterministic input
|
|
296
|
+
was a `## Delivery Slicing` table no Story carries). Run it when the operator
|
|
297
|
+
asks for it, after Author and before Persist — the last point a finding folds
|
|
298
|
+
into a re-author:
|
|
299
|
+
|
|
300
|
+
```bash
|
|
301
|
+
node .agents/scripts/plan-critics.js \
|
|
302
|
+
--stories temp/plan-<slug>/stories.json \
|
|
303
|
+
[--tech-spec temp/plan-<slug>/techspec.md]
|
|
304
|
+
```
|
|
386
305
|
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
prerequisite. That third trigger is what gives the default N=1 plan a cheap
|
|
394
|
-
viability check, since the size trigger is unreachable at one ticket and this
|
|
395
|
-
repo's resolved `riskHeuristics` is empty. The probe is conservative — explicit
|
|
396
|
-
markers only, so a plan naming no such artifact dispatches exactly as before.
|
|
306
|
+
It fires on one deterministic trigger: the **external-dependency** probe
|
|
307
|
+
finding an out-of-repo marker — a scoped package the plan names that no repo
|
|
308
|
+
manifest declares, a cross-repo `github.com/<owner>/<repo>` reference, or an
|
|
309
|
+
endpoint named as a service prerequisite. The probe is conservative —
|
|
310
|
+
explicit markers only. It exits 0 on **any** verdict (verdicts route work,
|
|
311
|
+
they do not gate) and exits **1** only on a usage/IO error — no critic ran.
|
|
397
312
|
|
|
398
313
|
```jsonc
|
|
399
314
|
{
|
|
400
|
-
"
|
|
401
|
-
"premortem": { "critic": "pre-mortem", "dispatch": true, "reasons": ["…"] },
|
|
402
|
-
"textHygiene": { "critic": "text-hygiene", "findings": [] }
|
|
315
|
+
"premortem": { "critic": "pre-mortem", "dispatch": true, "reasons": ["…"] }
|
|
403
316
|
}
|
|
404
317
|
```
|
|
405
318
|
|
|
406
|
-
|
|
407
|
-
|
|
408
|
-
`
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
|
|
416
|
-
|
|
417
|
-
|
|
418
|
-
|
|
419
|
-
maker-blind invariant, the `consolidation` and `pre-mortem` charters, and the
|
|
420
|
-
output shape standalone. When the kill-switch is off
|
|
421
|
-
(`roleScopedAgents: false`) or the host cannot spawn at this depth, fall back
|
|
422
|
-
to a generic sub-agent and hand it the same charter (the `consolidation` /
|
|
423
|
-
`pre-mortem` definitions in [`plan-critic.md`](../../agents/plan-critic.md)).
|
|
424
|
-
**When both critics fire, dispatch them in a single turn.** Consolidation and
|
|
425
|
-
pre-mortem read the same immutable draft, share no write path, and neither
|
|
426
|
-
consumes the other's verdict — the textbook independent fan-out of
|
|
427
|
-
[`parallel-tooling.md`](parallel-tooling.md) Rule 3. Issue both `Agent` calls
|
|
428
|
-
together in one assistant turn rather than awaiting the first verdict before
|
|
429
|
-
spawning the second; serialized critics double the round's wall clock and buy
|
|
430
|
-
nothing, because you fold both verdicts into the same re-author round anyway.
|
|
431
|
-
|
|
432
|
-
Either way the critic is **maker-blind**: hand it the draft artifacts
|
|
433
|
-
(`stories.json`, and `techspec.md` when present) — never the authoring
|
|
434
|
-
transcript or the reasons the planner believed its own draft is sound. A
|
|
435
|
-
critic that reads the maker's case grades the case, not the draft.
|
|
319
|
+
On `dispatch: true`, dispatch **one fresh-context, maker-blind sub-agent**.
|
|
320
|
+
When `delivery.routing.roleScopedAgents` is enabled (the **default**), use
|
|
321
|
+
`subagent_type: plan-critic` — it boots on the role-scoped
|
|
322
|
+
[`plan-critic`](../../agents/plan-critic.md) context (its own system prompt,
|
|
323
|
+
no `CLAUDE.md` @-closure) that carries the maker-blind invariant, the
|
|
324
|
+
`pre-mortem` charter, and the output shape standalone. When the kill-switch
|
|
325
|
+
is off (`roleScopedAgents: false`) or the host cannot spawn at this depth,
|
|
326
|
+
fall back to a generic sub-agent and hand it the same charter. Either way the
|
|
327
|
+
critic is **maker-blind**: hand it the draft artifacts (`stories.json`, and
|
|
328
|
+
`techspec.md` when present) — never the authoring transcript or the reasons
|
|
329
|
+
the planner believed its own draft is sound. A critic that reads the maker's
|
|
330
|
+
case grades the case, not the draft. Fold surviving findings into Gate #2 or
|
|
331
|
+
a re-author round.
|
|
436
332
|
|
|
437
333
|
## What `--dry-run` actually gates
|
|
438
334
|
|
|
439
335
|
`plan-persist.js --dry-run` is the same command with GitHub writes suppressed,
|
|
440
|
-
and every gate runs before the first `createIssue` would fire
|
|
441
|
-
the
|
|
442
|
-
|
|
443
|
-
|
|
444
|
-
|
|
445
|
-
|
|
336
|
+
and every gate runs before the first `createIssue` would fire. Since
|
|
337
|
+
Story #5312 the gates split two ways, and the dry-run is where the second
|
|
338
|
+
half is read:
|
|
339
|
+
|
|
340
|
+
**Hard — the run refuses:** a body that does not parse, a ticket that is not
|
|
341
|
+
a Story, an empty `acceptance[]` or `verify[]`, an unknown or cyclic
|
|
342
|
+
`depends_on`, the acceptance partition at N>1, the supersede partition, a
|
|
343
|
+
forbidden commit-subject prefix, and a `deletes` entry naming a path absent
|
|
344
|
+
at base.
|
|
345
|
+
|
|
346
|
+
**Warnings — listed, then the persist proceeds:** a `creates` on a path that
|
|
347
|
+
exists at base or a `refactors-existing` on one that does not (including a
|
|
348
|
+
path the base branch deleted or renamed, named with the removing commit), a
|
|
349
|
+
goal or acceptance path absent at base, a `verify[]` command naming an absent
|
|
350
|
+
test file, and an `open-question` in a body (`Flag if…`, `TBD`, a trailing
|
|
351
|
+
`?`). The list also names every `changes[]` **repair** the run applied — a
|
|
352
|
+
plain-string bullet or a trailing parenthetical rewritten into
|
|
353
|
+
`{ path, assumption }` by probing base. The same list rides the result
|
|
354
|
+
envelope as `warnings[]` and `repairs[]`, so a `--chain-on-clean` run loses
|
|
355
|
+
nothing.
|
|
356
|
+
|
|
357
|
+
A dry run that comes back clean has paid for every deterministic refusal, so
|
|
358
|
+
the real persist has nothing left to discover except network failure.
|
|
446
359
|
|
|
447
360
|
## The container Epic (Gate #3)
|
|
448
361
|
|