mandrel 2.56.0 → 2.57.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agents/agents/plan-critic.md +13 -18
- package/.agents/agents/story-worker.md +25 -34
- package/.agents/docs/agentrc-reference.json +0 -30
- package/.agents/docs/configuration.md +8 -28
- package/.agents/docs/execution-reference.md +5 -5
- package/.agents/docs/quality-gates.md +8 -7
- package/.agents/instructions.md +9 -10
- package/.agents/schemas/agentrc.schema.json +9 -185
- package/.agents/schemas/story-deliver-terminal.schema.json +1 -1
- package/.agents/scripts/acceptance-eval.js +107 -17
- package/.agents/scripts/ceremony-derive.js +191 -0
- package/.agents/scripts/check-context-budget.js +28 -33
- package/.agents/scripts/check-cyclomatic.js +4 -3
- package/.agents/scripts/deliver-light.js +31 -94
- package/.agents/scripts/lib/audit-suite/checklist-threading.js +15 -2
- package/.agents/scripts/lib/baselines/coverage-updater-cli.js +110 -0
- package/.agents/scripts/lib/baselines/crap-preview-scan.js +25 -0
- package/.agents/scripts/lib/baselines/crap-updater-cli.js +223 -0
- package/.agents/scripts/lib/bdd-scenario-budget.js +21 -3
- package/.agents/scripts/lib/bootstrap/quality-bootstrap.js +0 -1
- package/.agents/scripts/lib/close-validation/gates.js +52 -1
- package/.agents/scripts/lib/config/acceptance-eval.js +25 -57
- package/.agents/scripts/lib/config/delivery-routing.js +7 -33
- package/.agents/scripts/lib/config/explain.js +0 -19
- package/.agents/scripts/lib/config/limits.js +18 -78
- package/.agents/scripts/lib/config/quality.js +6 -3
- package/.agents/scripts/lib/config/runners.js +3 -2
- package/.agents/scripts/lib/config-settings-schema-delivery.js +15 -68
- package/.agents/scripts/lib/config-settings-schema-quality.js +0 -14
- package/.agents/scripts/lib/config-settings-schema.js +16 -143
- package/.agents/scripts/lib/crap-engine.js +35 -4
- package/.agents/scripts/lib/crap-utils.js +17 -1
- package/.agents/scripts/lib/cyclomatic-ceiling.js +19 -7
- package/.agents/scripts/lib/generated/agentrc-validator.js +1 -1
- package/.agents/scripts/lib/observability/runtime-friction.js +1 -1
- package/.agents/scripts/lib/observability/source-classifier.js +1 -0
- package/.agents/scripts/lib/orchestration/acceptance-eval-decision.js +5 -4
- package/.agents/scripts/lib/orchestration/ceremony-routing.js +19 -73
- package/.agents/scripts/lib/orchestration/complexity-gate.js +46 -212
- package/.agents/scripts/lib/orchestration/file-assumptions.js +32 -17
- package/.agents/scripts/lib/orchestration/light-escalation.js +3 -3
- package/.agents/scripts/lib/orchestration/light-suitability.js +66 -233
- package/.agents/scripts/lib/orchestration/plan-context.js +181 -387
- package/.agents/scripts/lib/orchestration/plan-critic-conditions.js +42 -153
- package/.agents/scripts/lib/orchestration/plan-critics-evaluate.js +14 -70
- package/.agents/scripts/lib/orchestration/plan-persist/changes-repair.js +300 -0
- package/.agents/scripts/lib/orchestration/plan-persist/persist-helpers.js +131 -168
- package/.agents/scripts/lib/orchestration/plan-persist/run-plan-persist.js +118 -297
- package/.agents/scripts/lib/orchestration/plan-persist/soft-findings.js +55 -0
- package/.agents/scripts/lib/orchestration/plan-persist/story-ops.js +16 -65
- package/.agents/scripts/lib/orchestration/plan-persist/wave-serialisation.js +22 -35
- package/.agents/scripts/lib/orchestration/plan-text-hygiene.js +30 -139
- package/.agents/scripts/lib/orchestration/planning/memory-pool-advisory.js +61 -223
- package/.agents/scripts/lib/orchestration/single-story-close/phases/close-validation.js +5 -0
- package/.agents/scripts/lib/orchestration/single-story-close/phases/pre-gate-steps.js +46 -16
- package/.agents/scripts/lib/orchestration/story-close/context-budget-writeback.js +213 -0
- package/.agents/scripts/lib/orchestration/task-body-validator.js +10 -63
- package/.agents/scripts/lib/orchestration/ticket-validator-conflicts.js +33 -539
- package/.agents/scripts/lib/orchestration/ticket-validator-sizing.js +21 -414
- package/.agents/scripts/lib/orchestration/ticket-validator.js +54 -118
- package/.agents/scripts/lib/orchestration/verify-credit.js +69 -24
- package/.agents/scripts/lib/story-body/body-format-lints.js +15 -85
- package/.agents/scripts/lib/story-body/story-body.js +17 -237
- package/.agents/scripts/lib/templates/decomposer-prompts.js +84 -121
- package/.agents/scripts/lib/test-isolate/cli-options.js +93 -0
- package/.agents/scripts/lib/test-isolate/progress-log.js +45 -0
- package/.agents/scripts/lib/test-isolate/render-report.js +97 -0
- package/.agents/scripts/lib/test-isolate/run-isolate.js +87 -0
- package/.agents/scripts/lib/test-run-credit.js +266 -0
- package/.agents/scripts/lib/wave-runner/footprint.js +48 -358
- package/.agents/scripts/lib/wave-runner/ready-set.js +6 -5
- package/.agents/scripts/lib/workers/crap-worker.js +32 -41
- package/.agents/scripts/plan-context.js +7 -9
- package/.agents/scripts/plan-critics.js +28 -54
- package/.agents/scripts/plan-persist.js +25 -68
- package/.agents/scripts/quality-preview.js +51 -0
- package/.agents/scripts/run-tests.js +12 -0
- package/.agents/scripts/stories-wave-tick.js +23 -45
- package/.agents/scripts/test-isolate.js +13 -180
- package/.agents/scripts/update-coverage-baseline.js +25 -70
- package/.agents/scripts/update-crap-baseline.js +19 -123
- package/.agents/skills/core/scope-triage/SKILL.md +3 -3
- package/.agents/workflows/audit-clean-code.md +4 -3
- package/.agents/workflows/helpers/acceptance-self-eval.md +41 -41
- package/.agents/workflows/helpers/code-quality-guardrails.md +4 -4
- package/.agents/workflows/helpers/code-review.md +2 -3
- package/.agents/workflows/helpers/deliver-digest.md +41 -57
- package/.agents/workflows/helpers/deliver-light.md +40 -105
- package/.agents/workflows/helpers/deliver-reference.md +1 -1
- package/.agents/workflows/helpers/deliver-story-reference.md +37 -58
- package/.agents/workflows/helpers/deliver-story.md +9 -13
- package/.agents/workflows/helpers/plan-reference.md +132 -219
- package/.agents/workflows/mandrel-plan.md +27 -40
- package/.agents/workflows/memory-consolidate.md +9 -13
- package/docs/CHANGELOG.md +23 -0
- package/lib/migrations/index.js +4 -0
- package/lib/migrations/steps/2.57.0-retire-delivery-limit-knobs.js +45 -0
- package/lib/migrations/steps/2.57.0-retire-planning-limit-knobs.js +59 -0
- package/package.json +1 -1
- package/.agents/scripts/lib/framework-version.js +0 -39
- package/.agents/scripts/lib/orchestration/consolidation-precondition.js +0 -223
- package/.agents/scripts/lib/orchestration/plan-persist/fan-out-gate.js +0 -97
- package/.agents/scripts/lib/orchestration/planning/decomposer-context.js +0 -26
- package/.agents/scripts/lib/orchestration/spec-budget.js +0 -89
- package/.agents/scripts/lib/orchestration/spec-spill.js +0 -74
- package/.agents/scripts/lib/orchestration/verify-tier-repair.js +0 -107
|
@@ -22,9 +22,9 @@ maintainability / coverage / crap) proves the code is *healthy*, not that it
|
|
|
22
22
|
satisfies *this Story's* acceptance criteria.
|
|
23
23
|
|
|
24
24
|
The loop is **always on** (a hard cutover — there is no flag to disable it) and
|
|
25
|
-
**bounded** by `delivery.acceptanceEval.maxRounds` (default 2
|
|
26
|
-
|
|
27
|
-
|
|
25
|
+
**bounded** by `delivery.acceptanceEval.maxRounds` (default 2; `0` means the
|
|
26
|
+
verdict is scored once with no redraft round). It is **distinct from** the
|
|
27
|
+
per-run `sibling-coherence` epilogue
|
|
28
28
|
step (`planRunEpilogue` for N>1): this loop is per-Story, per-criterion,
|
|
29
29
|
mid-delivery, and evaluates the actual work product.
|
|
30
30
|
|
|
@@ -55,23 +55,21 @@ mid-delivery, and evaluates the actual work product.
|
|
|
55
55
|
>
|
|
56
56
|
> **Whether to spawn fresh at all is routed off the derived change level**
|
|
57
57
|
> — the same signal `review-depth.js` resolves depth from, so the two
|
|
58
|
-
> decisions cannot disagree. Derive it with
|
|
59
|
-
>
|
|
60
|
-
>
|
|
61
|
-
>
|
|
62
|
-
>
|
|
63
|
-
>
|
|
64
|
-
>
|
|
65
|
-
>
|
|
66
|
-
>
|
|
67
|
-
> **`high` (the diff touches a sensitive path registered
|
|
68
|
-
> `audit-rules.json`) → `fresh`** (spawn the critic); **`low` (it
|
|
69
|
-
> none) → `inline`** (the contract-identical inline fallback
|
|
70
|
-
>
|
|
71
|
-
>
|
|
72
|
-
>
|
|
73
|
-
> full ceremony** (fail-safe). This chooses fresh-vs-inline **per cluster
|
|
74
|
-
> only — it never changes the cluster count**.
|
|
58
|
+
> decisions cannot disagree. Derive it with one script over the Story
|
|
59
|
+
> branch (Story #5313) — never a hand-carried import block:
|
|
60
|
+
>
|
|
61
|
+
> ```bash
|
|
62
|
+
> node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
|
|
63
|
+
> ```
|
|
64
|
+
>
|
|
65
|
+
> It computes the change set once (`files`), derives the level and
|
|
66
|
+
> classes, and resolves the ceremony per cluster (`mode`, `reason`,
|
|
67
|
+
> `verdictOwner`): **`high` (the diff touches a sensitive path registered
|
|
68
|
+
> in `audit-rules.json`) → `fresh`** (spawn the critic); **`low` (it
|
|
69
|
+
> touches none) → `inline`** (the contract-identical inline fallback
|
|
70
|
+
> below); **`null` / unknown (the diff could not be enumerated) → `fresh`
|
|
71
|
+
> + full ceremony** (fail-safe). This chooses fresh-vs-inline **per
|
|
72
|
+
> cluster only — it never changes the cluster count**.
|
|
75
73
|
>
|
|
76
74
|
> The routing signal is deliberately **not** a planner-authored risk
|
|
77
75
|
> verdict: a level the plan asserted about itself was exactly the signal
|
|
@@ -80,7 +78,7 @@ mid-delivery, and evaluates the actual work product.
|
|
|
80
78
|
>
|
|
81
79
|
> **Inline-critic path (low-level-routed OR nesting-absent harness).** The
|
|
82
80
|
> verdict is authored **inline** whenever the risk router above resolves to
|
|
83
|
-
> `inline` (a low-risk cluster
|
|
81
|
+
> `inline` (a low-risk cluster), and also as
|
|
84
82
|
> a **fallback** on any harness that cannot spawn the fresh critic.
|
|
85
83
|
> Dispatching the critic as a nested `Agent` is the fresh-context shape and
|
|
86
84
|
> works on any harness that carries `Agent` into sub-agents (Claude Code ≥
|
|
@@ -102,18 +100,18 @@ mid-delivery, and evaluates the actual work product.
|
|
|
102
100
|
> comment (if you block) that the inline fallback was used.
|
|
103
101
|
|
|
104
102
|
The critic:
|
|
105
|
-
|
|
103
|
+
+ Inspects the **change set handed to it in its spawn context** — the one
|
|
106
104
|
list computed above — and the Story's inline `acceptance[]` / `verify[]`
|
|
107
105
|
arrays. Pass the file list explicitly when you dispatch the critic; it
|
|
108
106
|
does not re-enumerate the diff for itself, so a commit
|
|
109
107
|
landing mid-ceremony cannot leave the critic scoring a different change
|
|
110
108
|
than the one that routed it.
|
|
111
|
-
|
|
109
|
+
+ **Runs the `verify[]` commands** and consumes their output as **required
|
|
112
110
|
evidence** when scoring the relevant acceptance items. `verify[]` is not
|
|
113
111
|
optional advisory pre-flight — a criterion cannot be scored `met` without
|
|
114
112
|
the supporting `verify[]` evidence where a `verify[]` command is relevant
|
|
115
113
|
to it.
|
|
116
|
-
|
|
114
|
+
+ **Reuses the credited full-suite run instead of re-paying for it.**
|
|
117
115
|
Before spawning a `verify[]` entry, classify it with `resolveVerifyCredit`
|
|
118
116
|
from
|
|
119
117
|
[`verify-credit.js`](../../scripts/lib/orchestration/verify-credit.js): an
|
|
@@ -125,7 +123,7 @@ mid-delivery, and evaluates the actual work product.
|
|
|
125
123
|
manufacture a pass. The gate warns on any such entry: the intended shape
|
|
126
124
|
is scoped `verify[]` entries **plus** the one credited run
|
|
127
125
|
([`deliver-digest.md`](deliver-digest.md) § 5).
|
|
128
|
-
|
|
126
|
+
+ **Shares `lint` / `typecheck` evidence with close.** When a
|
|
129
127
|
`verify[]` command is **byte-identical** to a close-validation gate — in
|
|
130
128
|
practice only the cheap, command-identical `lint` and `typecheck` gates
|
|
131
129
|
(`npm run lint` and the resolved `project.commands.typecheck`) — the
|
|
@@ -146,19 +144,21 @@ mid-delivery, and evaluates the actual work product.
|
|
|
146
144
|
fresh — a false-fresh coverage record without `coverage-final.json`
|
|
147
145
|
silently weakens the floor. Limit the evidence-share to `lint` and
|
|
148
146
|
`typecheck`.
|
|
149
|
-
|
|
147
|
+
+ Emits a **cluster** verdict file under `temp/` conforming to
|
|
150
148
|
[`acceptance-eval-verdict.schema.json`](../../schemas/acceptance-eval-verdict.schema.json):
|
|
151
149
|
one `{ index, criterion, verdict: met|partial|unmet, evidence,
|
|
152
150
|
verifyEvidence[] }` record per acceptance item **in that cluster**, each
|
|
153
151
|
`index` being the item's position in the Story's full `acceptance[]`
|
|
154
152
|
array. A fresh critic **returns that path to you** rather than calling the
|
|
155
153
|
gate itself.
|
|
156
|
-
2. **Dispatch the round's clusters in parallel, then merge into one verdict
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
the
|
|
154
|
+
2. **Dispatch the round's clusters in parallel, then merge into one verdict
|
|
155
|
+
(fresh critics only).** When the verdict owner is the **inline self-eval**,
|
|
156
|
+
author **one** verdict file covering every `acceptance[]` item and score it
|
|
157
|
+
in one gate call — the cluster merge below does not apply (Story #5313).
|
|
158
|
+
For fresh critics, the clusters of a round are independent, so dispatch
|
|
159
|
+
**all** of the round's critics as N `Agent` calls **in a single assistant
|
|
160
|
+
turn** — [`parallel-tooling.md`](parallel-tooling.md) **Rule 3** — never
|
|
161
|
+
serially, and never one round per cluster.
|
|
162
162
|
|
|
163
163
|
Then **merge** the cluster verdicts into **one** verdict file under `temp/`:
|
|
164
164
|
concatenate every cluster's `criteria[]` records and order the merged array
|
|
@@ -187,24 +187,24 @@ mid-delivery, and evaluates the actual work product.
|
|
|
187
187
|
|
|
188
188
|
```bash
|
|
189
189
|
node <main-repo>/.agents/scripts/acceptance-eval.js \
|
|
190
|
-
--story <storyId> --verdict <
|
|
191
|
-
--expected-criteria <number of acceptance[] items>
|
|
190
|
+
--story <storyId> --verdict <verdict-path>
|
|
192
191
|
```
|
|
193
192
|
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
193
|
+
The gate reads the Story's `acceptance[]` count itself (Story #5313): a
|
|
194
|
+
verdict whose `criteria[]` length differs — a single cluster's verdict
|
|
195
|
+
handed over unmerged — is rejected **before scoring**, with an error naming
|
|
196
|
+
the merge contract and consuming **no round**. `--expected-criteria` is
|
|
197
|
+
still accepted but redundant; when passed it must agree with that count.
|
|
198
198
|
|
|
199
199
|
The gate validates the verdict against the schema, applies the round cap,
|
|
200
200
|
emits the per-criterion `acceptance-eval` signal into the retro / feedback
|
|
201
201
|
substrate, prints a JSON envelope, and exits with one of three decisions:
|
|
202
|
-
|
|
202
|
+
+ **`decision: "proceed"`** (every criterion `met`) → exit 0. Proceed to
|
|
203
203
|
close.
|
|
204
|
-
|
|
204
|
+
+ **`decision: "redraft"`** (some `partial`/`unmet`, rounds remaining) →
|
|
205
205
|
exit 0. Redraft the flagged criteria (named in `unmetCriteria[]`), commit
|
|
206
206
|
the fix, and start another round.
|
|
207
|
-
|
|
207
|
+
+ **`decision: "block"`** (round cap reached, criteria still unmet) → exit
|
|
208
208
|
non-zero. **Do not proceed to close.** Take the caller's blocked path
|
|
209
209
|
(transition to `agent::blocked`) and post a `friction` comment naming the
|
|
210
210
|
unmet criteria and their evidence. Never silently proceed to close.
|
|
@@ -28,15 +28,15 @@ hook calls the same script.
|
|
|
28
28
|
|
|
29
29
|
Cyclomatic complexity (CC) is measured per function by `escomplex` (the same
|
|
30
30
|
engine the maintainability axis of `check-baselines.js` runs). Two
|
|
31
|
-
thresholds
|
|
32
|
-
`
|
|
33
|
-
|
|
31
|
+
thresholds — the advisory `delivery.quality.codingGuardrails.cyclomaticFlag`
|
|
32
|
+
and the fixed ceiling of 12 (`cyclomaticMustFix` was retired as a config
|
|
33
|
+
key in Story #5313):
|
|
34
34
|
|
|
35
35
|
| CC range | Action |
|
|
36
36
|
| --- | --- |
|
|
37
37
|
| ≤ 8 | Pass — no annotation required. |
|
|
38
38
|
| > 8 (default `cyclomaticFlag`) | **Flag** — `quality:preview` counts the function in its `new-method count over c=<flag>` column. The function is allowed to land but the report names it. |
|
|
39
|
-
|
|
|
39
|
+
| ≥ 12 | **Advisory in `quality:preview`** — listed by file, method and reading; the preview exits 0 on it. **Ratchet in `check-cyclomatic.js`**: fails when a file gains a function above 12, or when its worst function gets worse than the recorded baseline. |
|
|
40
40
|
|
|
41
41
|
`check-cyclomatic.js` is a **ratchet**, not a cliff: `baselines/cyclomatic.json`
|
|
42
42
|
records the over-ceiling functions a repository already carries, so adopting
|
|
@@ -384,9 +384,8 @@ ceremony above:
|
|
|
384
384
|
`[HEAD_REF]`, stage explicit paths only, and make **one focused
|
|
385
385
|
conventional commit per lens** (`fix(<scope>): <description> (review findings batch)`).
|
|
386
386
|
3. Bounded-attempt semantics extend to the batch: each finding gets **at
|
|
387
|
-
most one** attempt
|
|
388
|
-
|
|
389
|
-
escalation (`scope-exceeded`) instead of committing.
|
|
387
|
+
most one** attempt (the `maxFixScopeFiles` file-count ceiling was retired
|
|
388
|
+
in Story #5313 — a fix is bounded by attempts, not by file count).
|
|
390
389
|
4. After **all** lens batches are committed, run a **single** validation
|
|
391
390
|
pass (`npm run lint` plus the relevant `npm test` slice) and a
|
|
392
391
|
**single** targeted rescan over the touched files. Surviving batched
|
|
@@ -49,37 +49,24 @@ the only sanctioned landing. A silent local build is not a delivery.
|
|
|
49
49
|
## 3. Change set — computed once, handed to everyone
|
|
50
50
|
|
|
51
51
|
One enumeration per Story. A critic that re-runs its own `git diff`
|
|
52
|
-
can score a different set than the one that routed it.
|
|
53
|
-
|
|
54
|
-
|
|
52
|
+
can score a different set than the one that routed it. Derive the change
|
|
53
|
+
set, the level and the ceremony with **one script** — never a hand-carried
|
|
54
|
+
import block:
|
|
55
55
|
|
|
56
56
|
```bash
|
|
57
|
-
node
|
|
58
|
-
const lib = "<main-repo>/.agents/scripts/lib/orchestration";
|
|
59
|
-
const { computeChangeSet } = await import(`${lib}/change-set.js`);
|
|
60
|
-
const { deriveChangeLevel } = await import(`${lib}/review-depth.js`);
|
|
61
|
-
const { resolveCeremonyForRisk } = await import(`${lib}/ceremony-routing.js`);
|
|
62
|
-
const { files } = computeChangeSet({ baseRef: "main", headRef: "story-<storyId>" });
|
|
63
|
-
// deriveChangeLevel({ changedFiles, injectedRules?, selectSensitivePathClassesFn? })
|
|
64
|
-
// -> { level, classes } — an OBJECT, never a bare level.
|
|
65
|
-
const { level, classes } = deriveChangeLevel({ changedFiles: files });
|
|
66
|
-
// resolveCeremonyForRisk({ derivedLevel, clusterIndex?, freshCriticSampleRate?,
|
|
67
|
-
// ceremonyProfile? }) -> { mode, reason, sampled, profile, verdictOwner }.
|
|
68
|
-
// derivedLevel is that level STRING — the object above matches no tier and
|
|
69
|
-
// routes to the null fail-safe: a fresh critic, silently.
|
|
70
|
-
const ceremony = resolveCeremonyForRisk({ derivedLevel: level, clusterIndex: 0 });
|
|
71
|
-
console.log(JSON.stringify({ files, level, classes, ...ceremony }));
|
|
72
|
-
'
|
|
57
|
+
node <main-repo>/.agents/scripts/ceremony-derive.js --story <storyId> --cwd <workCwd>
|
|
73
58
|
```
|
|
74
59
|
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
`
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
60
|
+
It prints one JSON object: `files` — the one change set every critic is
|
|
61
|
+
handed (`null` when the diff could not be enumerated) — plus `level` and
|
|
62
|
+
`classes` from `review-depth.js`, and `mode`, `reason` and `verdictOwner`
|
|
63
|
+
from `ceremony-routing.js`. Level rules: a sensitive path registered in
|
|
64
|
+
`audit-rules.json` → `high`, none → `low`, an unenumerable diff → `null`.
|
|
65
|
+
Ceremony rules: `minimal` → always inline, `strict` → always fresh,
|
|
66
|
+
`standard` → `high`/`null` → fresh and `low` → inline. An `inline` dispatch
|
|
67
|
+
mode overrides all of it to inline critics. Close's `review-depth.js` reads
|
|
68
|
+
the same derived level, so the two cannot disagree. `--base <ref>` overrides
|
|
69
|
+
`project.baseBranch`.
|
|
83
70
|
|
|
84
71
|
## 4. Acceptance self-eval (Step 1a, required)
|
|
85
72
|
|
|
@@ -87,52 +74,49 @@ fresh. An `inline` dispatch mode overrides all of it to inline critics. Close's
|
|
|
87
74
|
self-eval, named by `verdictOwner`, never both and never a warm-up pass. Each
|
|
88
75
|
scores its cluster's `acceptance[]` items against the change set above, with
|
|
89
76
|
`verify[]` output as evidence. Bounded by `delivery.acceptanceEval.maxRounds`
|
|
90
|
-
(default 2).
|
|
77
|
+
(default 2; `0` scores once with no redraft).
|
|
91
78
|
|
|
92
|
-
**
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
79
|
+
**Inline owner:** author **one** verdict file covering every `acceptance[]`
|
|
80
|
+
item and score it in **one** gate call — there is no cluster merge.
|
|
81
|
+
**Fresh critics:** one round = N cluster critics → ONE merged verdict → ONE
|
|
82
|
+
gate call. Merge every cluster's records into a single `criteria[]` in
|
|
83
|
+
`acceptance[]` order and score that once; a gate call per cluster spends a
|
|
84
|
+
round *per cluster* and races the round ledger.
|
|
96
85
|
|
|
97
86
|
`node <main-repo>/.agents/scripts/acceptance-eval.js --story <storyId>
|
|
98
|
-
--verdict <
|
|
87
|
+
--verdict <verdict-path>`
|
|
99
88
|
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
`
|
|
89
|
+
The gate reads the Story's `acceptance[]` count itself and rejects a verdict
|
|
90
|
+
whose `criteria[]` length differs **before** scoring, consuming no round;
|
|
91
|
+
`--expected-criteria` is accepted but redundant.
|
|
103
92
|
|
|
104
93
|
`proceed` → close. `redraft` → one more round inside the cap. `block` → **do
|
|
105
94
|
not close**: post a `friction` comment and flip `agent::blocked`.
|
|
106
95
|
Per-round mechanics: [`acceptance-self-eval.md`](acceptance-self-eval.md).
|
|
107
96
|
|
|
108
|
-
## 5. The one
|
|
97
|
+
## 5. The one full-suite run
|
|
109
98
|
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
push does not, and the backgrounded capture (below) ends the turn. Redraft
|
|
113
|
-
rounds run scoped tests; only this run needs credit, and a bare `npm test` /
|
|
114
|
-
`pnpm run test` deposits **none**, so close re-runs it. Shape it by what
|
|
115
|
-
`close-validation/gates.js` runs:
|
|
99
|
+
After the self-eval loop's last fix commit, run the project test runner
|
|
100
|
+
**once** in the worktree:
|
|
116
101
|
|
|
117
102
|
```bash
|
|
118
|
-
|
|
119
|
-
node <main-repo>/.agents/scripts/coverage-capture.js --cwd <workCwd>
|
|
120
|
-
# otherwise — the record close's `test` gate reads; <workCwd> ABSOLUTE,
|
|
121
|
-
# runner exactly `npm test` (both sides hash {cmd, args, cwd}):
|
|
122
|
-
node <main-repo>/.agents/scripts/evidence-gate.js --standalone \
|
|
123
|
-
--scope-id <storyId> --gate test --worktree <workCwd> -- npm test
|
|
103
|
+
npm test # in <workCwd>
|
|
124
104
|
```
|
|
125
105
|
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
106
|
+
A green full run on `story-<id>` deposits the `test` evidence close reads,
|
|
107
|
+
keyed on the tree, so close reports the gate as **credited** at unchanged
|
|
108
|
+
HEAD — a later commit voids it. The CRAP gate still captures coverage itself
|
|
109
|
+
when it needs an artifact. If the suite outruns the host's sync Bash
|
|
110
|
+
ceiling, dispatch it in the **background** — its completion re-invokes you;
|
|
111
|
+
never spawn a task to poll or `sleep`-loop against it
|
|
112
|
+
([`parallel-tooling.md`](parallel-tooling.md) Rule 2). Read the **output**,
|
|
113
|
+
not the exit code: the runner prints whether it deposited credit, and a run
|
|
114
|
+
off the Story branch or of a partial tier deposits nothing and says so.
|
|
115
|
+
Redraft rounds run the scoped projects for the roots you changed plus
|
|
116
|
+
`verify[]`, not the whole suite; only this run needs credit.
|
|
133
117
|
|
|
134
118
|
`verify[]` is scoped entries **plus** this one run: an entry that is itself a
|
|
135
|
-
full-suite command is reported credited against the same
|
|
119
|
+
full-suite command is reported credited against the same record, never
|
|
136
120
|
respawned.
|
|
137
121
|
|
|
138
122
|
## 6. Terminal envelope — the return contract
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
description:
|
|
3
|
-
The unplanned prompt path
|
|
3
|
+
The unplanned prompt path /mandrel-deliver takes for a free-text prompt. Judges a
|
|
4
4
|
prompt's predicted footprint, authors a receipt Story, then lands it through
|
|
5
5
|
the same single-story-init / single-story-close engine — every close gate
|
|
6
6
|
unchanged.
|
|
@@ -9,9 +9,9 @@ description:
|
|
|
9
9
|
# Unplanned delivery (the prompt path)
|
|
10
10
|
|
|
11
11
|
> **A path, not a command.** There is no `/deliver-light` to type. This file is
|
|
12
|
-
> reached
|
|
13
|
-
>
|
|
14
|
-
> the
|
|
12
|
+
> reached one way — `/mandrel-deliver "<prompt>"` (an operator describing small
|
|
13
|
+
> work); Story #5312 retired the `/mandrel-plan` Gate #1 suggestion that used
|
|
14
|
+
> to be the second door. Read
|
|
15
15
|
> [`deliver-digest.md`](deliver-digest.md) once first — the engine invariants,
|
|
16
16
|
> gates, and terminal-envelope contract below are its.
|
|
17
17
|
|
|
@@ -24,13 +24,11 @@ straight to execution from a prompt, landing through the unchanged close path.
|
|
|
24
24
|
It never relaxes a close gate, never bypasses the PR to `main`, and never lands
|
|
25
25
|
over-scope work silently.
|
|
26
26
|
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
magnitude, uncertainty, deployable span). They are deliberately two different
|
|
33
|
-
checks, so the gate still runs after a confirm.
|
|
27
|
+
One gate, whatever the door: the suitability gate below is the decision, read
|
|
28
|
+
against the predicted work's *effort and risk* (`STORY_SHAPE_CEILINGS`: change
|
|
29
|
+
kinds, magnitude, uncertainty, deployable span). `/mandrel-plan` no longer
|
|
30
|
+
suggests this path at its Gate #1 (Story #5312) — a prompt reaches it through
|
|
31
|
+
`/mandrel-deliver`, and the gate runs there every time.
|
|
34
32
|
|
|
35
33
|
## Scope by effort, not by artifact count {#scope-by-effort}
|
|
36
34
|
|
|
@@ -55,8 +53,10 @@ obeying its own test-first rule is a ceiling that over-fires.
|
|
|
55
53
|
|
|
56
54
|
Sensitivity is the exception and stays absolute: a footprint touching an auth,
|
|
57
55
|
crypto, billing, or migration class routes `full` however small or mechanical —
|
|
58
|
-
and unlike a ceiling, it is
|
|
59
|
-
|
|
56
|
+
and unlike a ceiling, it is never a warning. The shape ceilings themselves
|
|
57
|
+
stopped gating in Story #5313: a prediction past one is carried as a
|
|
58
|
+
`warnings[]` entry on the envelope, naming the exceeded axis, and the run
|
|
59
|
+
proceeds — `LIGHT_DIFF_CEILINGS` in the backstop is the only size block.
|
|
60
60
|
|
|
61
61
|
## Four invariants (do not skip one)
|
|
62
62
|
|
|
@@ -64,12 +64,13 @@ answer).
|
|
|
64
64
|
shared effort/risk machinery (`deriveStoryShape` / `deriveChangeLevel`)
|
|
65
65
|
**and** a ledgered model verdict with a recorded reason. Both must agree on
|
|
66
66
|
`lite`.
|
|
67
|
-
2. **
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
67
|
+
2. **The predicted shape warns; risk escalates.** A prompt past a shape
|
|
68
|
+
ceiling proceeds light with a `warnings[]` entry (Story #5313) — the
|
|
69
|
+
prediction is a guess the backstop bounds for real. Only an un-ledgered
|
|
70
|
+
verdict or an un-waivable risk rule (a sensitive-path class, a migration
|
|
71
|
+
span) refuses, and it refuses the same way attended or not: an
|
|
71
72
|
**`escalated` terminal envelope** that ends the session (§ Escalation is
|
|
72
|
-
terminal).
|
|
73
|
+
terminal). There is no question to wait for and no answer flag.
|
|
73
74
|
3. **Diff-derived backstop.** After implementation the ACTUAL change set is
|
|
74
75
|
re-checked — the diff is the real scope signal — and an over-ceiling diff is
|
|
75
76
|
blocked rather than landed.
|
|
@@ -95,30 +96,16 @@ answer).
|
|
|
95
96
|
|
|
96
97
|
Branch on `action` in the JSON envelope:
|
|
97
98
|
- **`proceed-light`** — the receipt Story is authored; read `storyId` and
|
|
98
|
-
`nextCommands`.
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
(§ Recording a proceed-light answer).
|
|
105
|
-
- **over-scope under `--yes`** — no `action` to branch on: the gate emits an
|
|
106
|
-
**`escalated` terminal envelope** instead (exit 2). § Escalation is
|
|
107
|
-
terminal governs; you are finished.
|
|
99
|
+
`nextCommands`. Relay every `warnings[]` entry to the operator verbatim
|
|
100
|
+
(each names the shape axis the prediction exceeded), then continue to
|
|
101
|
+
step 2 — the diff backstop in step 4 is what bounds the actual change.
|
|
102
|
+
- **escalation** — no `action` to branch on: the gate emits an
|
|
103
|
+
**`escalated` terminal envelope** instead (exit 2), attended or not.
|
|
104
|
+
§ Escalation is terminal governs; you are finished.
|
|
108
105
|
|
|
109
106
|
`--amends '#<id>'` is the canonical light case — shape-checked identically; a
|
|
110
107
|
heavy amendment escalates to `/mandrel-plan` like any other over-scope prompt.
|
|
111
108
|
|
|
112
|
-
**Entered from `/mandrel-plan` Gate #1?** Fill `--creates` / `--refactors` /
|
|
113
|
-
`--acceptance` / `--reason` from the plan-context envelope's codebase
|
|
114
|
-
snapshot and `complexitySignals` rather than re-deriving them from the seed
|
|
115
|
-
text — Gate #1 has already done that work, and re-deriving throws away the
|
|
116
|
-
better signal. An `ask-operator` verdict here means the two ceiling sets
|
|
117
|
-
disagreed: **return to [`../mandrel-plan.md`](../mandrel-plan.md) step 2 (Author) in the same
|
|
118
|
-
session**, carrying the interrogation you already paid for. That bounce-back
|
|
119
|
-
is not an escalation and does not need a fresh session (§ Why the two
|
|
120
|
-
directions differ).
|
|
121
|
-
|
|
122
109
|
2. **Init (same engine).** From the main checkout, synchronously, with the
|
|
123
110
|
maximum Bash timeout:
|
|
124
111
|
|
|
@@ -177,42 +164,10 @@ answer).
|
|
|
177
164
|
[`deliver-digest.md`](deliver-digest.md) § 5 — every close
|
|
178
165
|
gate runs byte-identical to the full path.
|
|
179
166
|
|
|
180
|
-
## Recording a proceed-light answer {#recording-a-proceed-light-answer}
|
|
181
|
-
|
|
182
|
-
The gate offers the operator two options, so **both** have to be executable.
|
|
183
|
-
Re-run the identical gate command with their answer appended:
|
|
184
|
-
|
|
185
|
-
```bash
|
|
186
|
-
node .agents/scripts/deliver-light.js --prompt "<prompt>" … \
|
|
187
|
-
--operator-proceed-light "<the operator's reason, in their words>"
|
|
188
|
-
```
|
|
189
|
-
|
|
190
|
-
The gate then proceeds light, records the decision in the receipt Story, and
|
|
191
|
-
returns it on the envelope's `override`. Do **not** instead re-shape the
|
|
192
|
-
prediction — shrinking `--refactors` until the gate stops objecting is
|
|
193
|
-
under-declaring the footprint, which is the one thing the coarse design must
|
|
194
|
-
not reward.
|
|
195
|
-
|
|
196
|
-
It is deliberately narrow, and a refusal is printed rather than silent:
|
|
197
|
-
|
|
198
|
-
- **Only a size prediction is waivable** — change kinds, magnitude,
|
|
199
|
-
uncertainty, deployable span. A sensitive-path class, a
|
|
200
|
-
migration-with-consumers span, and an unknown footprint (undeclared, glob,
|
|
201
|
-
no acceptance, unclassifiable) are refused: those are risk, not size, and
|
|
202
|
-
§ Scope by effort keeps them absolute.
|
|
203
|
-
- **The ledgered verdict still stands on its own.** The override substitutes
|
|
204
|
-
for the predicted *shape* only; `--route lite --reason "<why>"` is still
|
|
205
|
-
required.
|
|
206
|
-
- **Attended-only.** With `--yes` it is a usage error, not a quiet no-op —
|
|
207
|
-
an unattended run has no operator whose answer this could be, and over-scope
|
|
208
|
-
there still fails closed (§ Escalation is terminal).
|
|
209
|
-
|
|
210
|
-
What licenses this at all is step 4: the operator waives a *guess*, never the
|
|
211
|
-
diff backstop, which re-checks the actual change set against ground truth.
|
|
212
|
-
|
|
213
167
|
## Escalation is terminal {#escalation-is-terminal}
|
|
214
168
|
|
|
215
|
-
|
|
169
|
+
A refused gate — an un-ledgered verdict or an un-waivable risk rule, under
|
|
170
|
+
`--yes` or attended alike — emits a schema-validated `story-deliver-terminal`
|
|
216
171
|
envelope with **`status: "escalated"`**, `storyId: null`, and a `nextCommand`
|
|
217
172
|
naming the `/mandrel-plan` invocation that owns the work.
|
|
218
173
|
|
|
@@ -237,51 +192,31 @@ Nothing is left half-started: an escalated run creates **no receipt Story, no
|
|
|
237
192
|
every creation call site, and `escalation.created` records all three as `false`
|
|
238
193
|
in a shape the schema pins, so a later run finds nothing to trip over.
|
|
239
194
|
|
|
240
|
-
## Why
|
|
195
|
+
## Why escalation breaks the session {#why-the-two-directions-differ}
|
|
241
196
|
|
|
242
|
-
Traffic runs
|
|
243
|
-
|
|
244
|
-
|
|
197
|
+
Traffic runs one way between this path and `/mandrel-plan` — light →
|
|
198
|
+
`/mandrel-plan` on an over-scope prompt — and that direction has a deliberate
|
|
199
|
+
session rule:
|
|
245
200
|
|
|
246
|
-
> **The direction whose guard is model judgment must break the session
|
|
247
|
-
> direction whose guard is mechanical need not.**
|
|
201
|
+
> **The direction whose guard is model judgment must break the session.**
|
|
248
202
|
|
|
249
203
|
**Light → `/mandrel-plan` must be a fresh session.** What is being protected is
|
|
250
204
|
*authoring judgment*, and the empirical finding above is that a session already
|
|
251
205
|
framed as small work under-decomposes — one Story against a 3–5 contract where
|
|
252
206
|
a fresh session on the identical seed authored four. The frame is the hazard,
|
|
253
|
-
so only a new session removes it.
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
|
|
258
|
-
|
|
259
|
-
None of them degrade because the context is large, so nothing is gained by
|
|
260
|
-
paying for a fresh session.
|
|
261
|
-
|
|
262
|
-
Do not "fix" this into symmetry in either direction. Making `/mandrel-plan` → light
|
|
263
|
-
require a fresh session throws away a paid-for interrogation for no guard.
|
|
264
|
-
Letting light → `/mandrel-plan` run in-session reintroduces the exact failure the
|
|
265
|
-
`escalated` envelope exists to prevent.
|
|
266
|
-
|
|
267
|
-
## Constraints
|
|
268
|
-
|
|
269
|
-
- **Land, block, or escalate — never a silent local build.** The close push is
|
|
270
|
-
the only sanctioned landing; an `escalated` terminal is the only sanctioned
|
|
271
|
-
ending that delivers nothing, and it ends the session
|
|
272
|
-
(§ Escalation is terminal).
|
|
273
|
-
- **No parallel engine.** This path invokes `single-story-init.js` and
|
|
274
|
-
`single-story-close.js`; it never reimplements worktree, branch, PR, or merge
|
|
275
|
-
mechanics.
|
|
276
|
-
- **State only via `update-ticket-state.js`.** Drive every `agent::*`
|
|
277
|
-
transition through the script; report state, not process.
|
|
207
|
+
so only a new session removes it. Everything on this path's own side is
|
|
208
|
+
mechanical — `STORY_SHAPE_CEILINGS`, the ledgered verdict, the diff backstop —
|
|
209
|
+
and none of it degrades because the context is large.
|
|
210
|
+
|
|
211
|
+
Do not "fix" this by letting light → `/mandrel-plan` run in-session: that
|
|
212
|
+
reintroduces the exact failure the `escalated` envelope exists to prevent.
|
|
278
213
|
|
|
279
214
|
## See also
|
|
280
215
|
|
|
281
216
|
- [`/mandrel-deliver`](../mandrel-deliver.md) — the delivery entry point; routes here on a
|
|
282
217
|
free-text prompt.
|
|
283
|
-
- [`/mandrel-plan`](../mandrel-plan.md) —
|
|
284
|
-
|
|
218
|
+
- [`/mandrel-plan`](../mandrel-plan.md) — owns the work an over-scope prompt
|
|
219
|
+
escalates to.
|
|
285
220
|
- [`deliver-story.md`](deliver-story.md) — the one Story delivery engine every
|
|
286
221
|
path shares.
|
|
287
222
|
- [`deliver-digest.md`](deliver-digest.md) — engine invariants, gates, and the
|
|
@@ -361,7 +361,7 @@ sensitive-path classes in `audit-rules.json`
|
|
|
361
361
|
| Profile | Acceptance critic | When to use |
|
|
362
362
|
| --- | --- | --- |
|
|
363
363
|
| `minimal` | Always inline | Tiny trusted N=1 Stories |
|
|
364
|
-
| `standard` | Derived-level routed
|
|
364
|
+
| `standard` | Derived-level routed | Default |
|
|
365
365
|
| `strict` | Always fresh-context | High-assurance / regulated surfaces |
|
|
366
366
|
|
|
367
367
|
| Scope | What runs | Mechanism |
|