pi-gauntlet 5.5.1 → 5.5.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +15 -0
- package/agents/code-reviewer.md +7 -3
- package/extensions/lib/gauntlet-settings.test.ts +4 -3
- package/extensions/lib/gauntlet-settings.ts +1 -1
- package/extensions/phase-tracker.test.ts +259 -1
- package/extensions/phase-tracker.ts +75 -19
- package/package.json +1 -1
- package/skills/dispatching-parallel-agents/SKILL.md +1 -1
- package/skills/verification-before-completion/reference/conformance-check.md +36 -47
- package/skills/verification-before-completion/reference/settings-precedence.md +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,20 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.5.3 - 2026-09-13
|
|
4
|
+
|
|
5
|
+
- `code-reviewer`: verdict is a stated function of severity - a Critical or Moderate finding means `FIX_FIRST`, Minor-only and clean reports mean `SHIP`, matching the orchestrating skills. (#30)
|
|
6
|
+
- `phase-tracker`: two new closure-review blocks inside a brainstorming-entered flow once `verify` is in progress and a conformance audit has been observed - a lone top-level `agent: "implementer"` `subagent` dispatch is blocked (one-task `tasks` wave is the only isolated shape), and an implementer dispatch past `closureReview.maxFixRounds` is blocked with the escalation text. Rounds are counted from non-error results carrying an implementer child; the counter resets with the audit latch. `closureReview.enforce: false` disables both.
|
|
7
|
+
- `closureReview.maxFixRounds` default 2 -> 3.
|
|
8
|
+
- `conformance-check.md` / `dispatching-parallel-agents`: retries inside the conformance loop are one-task `tasks` waves and count against the cap.
|
|
9
|
+
|
|
10
|
+
## v5.5.2 - 2026-09-13
|
|
11
|
+
|
|
12
|
+
- `release.sh <level>` promotes the CHANGELOG `## Unreleased` section to `## vX.Y.Z - <date>` and commits it with `package.json` in the single `Release X.Y.Z` commit, so `patch`/`minor`/`major` now work here (the `current`-only path is gone). New CONFIG field `CHANGELOG_HEADING`.
|
|
13
|
+
- Release skill: a user instruction naming the level is the approval - no proposal step or re-confirmation; bundled follow-ups run after `verify`.
|
|
14
|
+
- AGENTS.md rewritten to always-on essentials plus routing; shared core bumped to v3. Persona frontmatter knobs table and pin rationale moved to `doc/personas.md`. Gold rule scoped to agent-initiated writes; a user instruction naming the write is its confirmation.
|
|
15
|
+
- Added `.pi/gauntlet-overrides.md` (`tracker: github`, release path, write-gate carve-out for user-named writes).
|
|
16
|
+
- `verification-before-completion/reference/conformance-check.md`: the conformance fix round is one parallel `implementer` `tasks` wave (greedy `Gn` selection, `conflicts` partners held) -> `git apply` -> scoped tests -> delta re-audit. Per-gap `spec-reviewer`, per-round `code-reviewer`, and the per-round full test set are gone. New **Convergence** step after every `CONFORMS`: full `Verification` set once, one direct `code-reviewer` over `git diff <r0-head>..HEAD`; repairs re-enter the round as a `conformance fix CR` wave that counts against `maxFixRounds`. Finish grammar, personas, and settings unchanged.
|
|
17
|
+
|
|
3
18
|
## v5.5.1 - 2026-09-10
|
|
4
19
|
|
|
5
20
|
- `linear`: copyable, version-scoped recovery for attachment-download 401s resolves the decrypted credential through linearis instead of reading encrypted token storage. Restricts credential delivery to HTTPS Linear uploads, rejects redirects, and checks downloaded bytes; an offline regression executes the documented example.
|
package/agents/code-reviewer.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: code-reviewer
|
|
3
|
-
description: Production-readiness code review with prioritized findings (Critical
|
|
3
|
+
description: Production-readiness code review with prioritized findings (Critical and Moderate block merge, Minor is a nit). Read-only — does not edit.
|
|
4
4
|
tools: read, grep, find, ls, bash
|
|
5
5
|
defaultContext: fresh
|
|
6
6
|
inheritProjectContext: true
|
|
@@ -44,11 +44,15 @@ Parallel-safe: F1,F3 disjoint; F2 conflicts F1 (both touch auth.ts)
|
|
|
44
44
|
Behaviour-change: yes | no
|
|
45
45
|
```
|
|
46
46
|
|
|
47
|
+
Severity decides SHIP versus FIX_FIRST. A Critical or Moderate finding means
|
|
48
|
+
FIX_FIRST. Minor-only findings and clean reports mean SHIP. REJECT overrides
|
|
49
|
+
both: refuse a change that must not land at all.
|
|
50
|
+
|
|
47
51
|
Severity:
|
|
48
52
|
|
|
49
53
|
- **Critical** — must fix before merge (data loss, security, broken correctness on a common path, broken contract).
|
|
50
|
-
- **Moderate** —
|
|
51
|
-
- **Minor** — nit, style, preference, suggestion.
|
|
54
|
+
- **Moderate** — must fix before merge (significant defect or drift that does not rise to Critical).
|
|
55
|
+
- **Minor** — nit, style, preference, suggestion; the only severity declinable without a fix round or re-review.
|
|
52
56
|
|
|
53
57
|
Label every finding with a globally unique `F1..Fn` ID (no restart per severity),
|
|
54
58
|
and a `touched-files:`/`touched-resources:` pair (files/resources a fix would
|
|
@@ -122,11 +122,12 @@ test("closureReview: enforce default true; false only when explicitly false", ()
|
|
|
122
122
|
assert.equal(resolveClosureReview({ closureReview: { enforce: true } }).enforce, true);
|
|
123
123
|
});
|
|
124
124
|
|
|
125
|
-
test("closureReview: maxFixRounds default
|
|
126
|
-
assert.equal(resolveClosureReview({}).maxFixRounds,
|
|
125
|
+
test("closureReview: maxFixRounds default 3, <0 -> 0, non-int -> 3", () => {
|
|
126
|
+
assert.equal(resolveClosureReview({}).maxFixRounds, 3);
|
|
127
127
|
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 5 } }).maxFixRounds, 5);
|
|
128
|
+
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 0 } }).maxFixRounds, 0);
|
|
128
129
|
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: -3 } }).maxFixRounds, 0);
|
|
129
|
-
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 1.5 } }).maxFixRounds,
|
|
130
|
+
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 1.5 } }).maxFixRounds, 3);
|
|
130
131
|
});
|
|
131
132
|
|
|
132
133
|
test("flowGuards: defaults + overrides", () => {
|
|
@@ -84,7 +84,7 @@ export function resolveClosureReview(g: PiGauntlet): ClosureReviewResolved {
|
|
|
84
84
|
const model = nonEmptyString(cr?.model) ? cr!.model.trim() : undefined;
|
|
85
85
|
const enforce = cr?.enforce !== false;
|
|
86
86
|
const raw = cr?.maxFixRounds;
|
|
87
|
-
const maxFixRounds = typeof raw === "number" && Number.isInteger(raw) ? (raw < 0 ? 0 : raw) :
|
|
87
|
+
const maxFixRounds = typeof raw === "number" && Number.isInteger(raw) ? (raw < 0 ? 0 : raw) : 3;
|
|
88
88
|
return { model, enforce, maxFixRounds };
|
|
89
89
|
}
|
|
90
90
|
|
|
@@ -6,9 +6,14 @@ import { tmpdir } from "node:os";
|
|
|
6
6
|
import { isAbsolute, join } from "node:path";
|
|
7
7
|
import registerPhaseTracker from "./phase-tracker.ts";
|
|
8
8
|
|
|
9
|
+
const originalSubagentDepth = process.env.PI_SUBAGENT_DEPTH;
|
|
10
|
+
process.env.PI_SUBAGENT_DEPTH = "0";
|
|
11
|
+
|
|
9
12
|
const tempDirs: string[] = [];
|
|
10
13
|
after(() => {
|
|
11
14
|
for (const dir of tempDirs) rmSync(dir, { recursive: true, force: true });
|
|
15
|
+
if (originalSubagentDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
|
|
16
|
+
else process.env.PI_SUBAGENT_DEPTH = originalSubagentDepth;
|
|
12
17
|
});
|
|
13
18
|
|
|
14
19
|
const tempCwd = (settings?: unknown) => {
|
|
@@ -529,7 +534,8 @@ test("implement-phase commit with implementer newer than both reviewers warns",
|
|
|
529
534
|
const warned = (await h.emitEvent("tool_result", commitResult("c1")))[0] as { content: { text: string }[] };
|
|
530
535
|
assert.match(warned.content[0].text, /no spec-reviewer or code-reviewer observed/);
|
|
531
536
|
} finally {
|
|
532
|
-
if (priorDepth
|
|
537
|
+
if (priorDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
|
|
538
|
+
else process.env.PI_SUBAGENT_DEPTH = priorDepth;
|
|
533
539
|
}
|
|
534
540
|
});
|
|
535
541
|
|
|
@@ -1164,3 +1170,255 @@ test("gauntlet_setting escalationLoop: setting absent -> ctx-derived main-loop m
|
|
|
1164
1170
|
const res = (await set.tools.find((t) => t.name === "gauntlet_setting")!.execute("g2", { key: "escalationLoop" }, undefined, undefined, set.ctx)) as { details: { implModel?: string } };
|
|
1165
1171
|
assert.equal(res.details.implModel, "p/strong:high");
|
|
1166
1172
|
});
|
|
1173
|
+
|
|
1174
|
+
// --- Conformance fix-loop dispatch guard (spec 2026-09-13-conformance-dispatch-guard) ---
|
|
1175
|
+
|
|
1176
|
+
const verifyBranch = (extra: unknown[] = []) => [
|
|
1177
|
+
...implementBranch(),
|
|
1178
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete" })),
|
|
1179
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress" })),
|
|
1180
|
+
...extra,
|
|
1181
|
+
];
|
|
1182
|
+
|
|
1183
|
+
const subagentCall = (id: string, input: unknown) => ({ toolName: "subagent", toolCallId: id, input });
|
|
1184
|
+
const loneImplementer = (extra: Record<string, unknown> = {}) => ({ agent: "implementer", task: "fix G1", ...extra });
|
|
1185
|
+
const implementerWave = (n = 1) => ({
|
|
1186
|
+
tasks: Array.from({ length: n }, (_, i) => ({ agent: "implementer", task: `fix G${i + 1}`, worktree: true })),
|
|
1187
|
+
});
|
|
1188
|
+
|
|
1189
|
+
const firstCallResult = async (h: ReturnType<typeof harness>, id: string, input: unknown) =>
|
|
1190
|
+
(await h.emitEvent("tool_call", subagentCall(id, input)))[0] as { block?: boolean; reason?: string } | undefined;
|
|
1191
|
+
|
|
1192
|
+
test("shape guard: lone implementer blocked in verify after the audit, with and without closureReview.model, async or not", async () => {
|
|
1193
|
+
for (const settings of [
|
|
1194
|
+
{ piGauntlet: { closureReview: { enforce: true } } },
|
|
1195
|
+
{ piGauntlet: { closureReview: { enforce: true, model: "x/y" } } },
|
|
1196
|
+
]) {
|
|
1197
|
+
const h = harness({ cwd: tempCwd(settings), branch: verifyBranch([subagentResult(["conformance-reviewer"])]) });
|
|
1198
|
+
await h.emit("session_start");
|
|
1199
|
+
const sync = await firstCallResult(h, "s1", loneImplementer());
|
|
1200
|
+
assert.equal(sync?.block, true);
|
|
1201
|
+
assert.match(sync?.reason ?? "", /dispatch implementers as a one-task tasks wave/);
|
|
1202
|
+
assert.match(sync?.reason ?? "", /piGauntlet\.closureReview\.enforce: false/);
|
|
1203
|
+
const async = await firstCallResult(h, "s2", loneImplementer({ async: true }));
|
|
1204
|
+
assert.equal(async?.block, true);
|
|
1205
|
+
}
|
|
1206
|
+
});
|
|
1207
|
+
|
|
1208
|
+
test("shape guard: lone implementer passes before the audit, in implement, in ship, and on management calls", async () => {
|
|
1209
|
+
const preLatch = harness({ branch: verifyBranch() });
|
|
1210
|
+
await preLatch.emit("session_start");
|
|
1211
|
+
assert.equal(await firstCallResult(preLatch, "s1", loneImplementer()), undefined);
|
|
1212
|
+
|
|
1213
|
+
const implement = harness({ branch: implementBranch([subagentResult(["conformance-reviewer"])]) });
|
|
1214
|
+
await implement.emit("session_start");
|
|
1215
|
+
assert.equal(await firstCallResult(implement, "s1", loneImplementer()), undefined);
|
|
1216
|
+
|
|
1217
|
+
const ship = harness({
|
|
1218
|
+
branch: verifyBranch([
|
|
1219
|
+
subagentResult(["conformance-reviewer"]),
|
|
1220
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete" })),
|
|
1221
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" })),
|
|
1222
|
+
]),
|
|
1223
|
+
});
|
|
1224
|
+
await ship.emit("session_start");
|
|
1225
|
+
assert.equal(await firstCallResult(ship, "s1", loneImplementer()), undefined);
|
|
1226
|
+
|
|
1227
|
+
const mgmt = harness({ branch: verifyBranch([subagentResult(["conformance-reviewer"])]) });
|
|
1228
|
+
await mgmt.emit("session_start");
|
|
1229
|
+
assert.equal(await firstCallResult(mgmt, "s1", { action: "status", agent: "implementer" }), undefined);
|
|
1230
|
+
});
|
|
1231
|
+
|
|
1232
|
+
test("shape guard: dormant when the flow was never entered (cold start verify) and when closureReview.enforce is false", async () => {
|
|
1233
|
+
const cold = harness({
|
|
1234
|
+
cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 0 } } }),
|
|
1235
|
+
branch: [phaseResult("start", phases({ verify: "in_progress" })), subagentResult(["conformance-reviewer"])],
|
|
1236
|
+
});
|
|
1237
|
+
await cold.emit("session_start");
|
|
1238
|
+
assert.equal(await firstCallResult(cold, "s1", loneImplementer()), undefined);
|
|
1239
|
+
assert.equal(await firstCallResult(cold, "s2", implementerWave(1)), undefined);
|
|
1240
|
+
assert.equal(await firstCallResult(cold, "s3", { chain: [{ agent: "implementer", task: "fix" }] }), undefined);
|
|
1241
|
+
|
|
1242
|
+
const off = harness({
|
|
1243
|
+
cwd: tempCwd({ piGauntlet: { closureReview: { enforce: false } } }),
|
|
1244
|
+
branch: verifyBranch([subagentResult(["conformance-reviewer"])]),
|
|
1245
|
+
});
|
|
1246
|
+
await off.emit("session_start");
|
|
1247
|
+
assert.equal(await firstCallResult(off, "s1", loneImplementer()), undefined);
|
|
1248
|
+
});
|
|
1249
|
+
|
|
1250
|
+
test("shape guard: a one-task tasks wave and a chain step are not the lone shape", async () => {
|
|
1251
|
+
const h = harness({ branch: verifyBranch([subagentResult(["conformance-reviewer"])]) });
|
|
1252
|
+
await h.emit("session_start");
|
|
1253
|
+
assert.equal(await firstCallResult(h, "s1", implementerWave(1)), undefined);
|
|
1254
|
+
assert.equal(await firstCallResult(h, "s2", { chain: [{ agent: "implementer", task: "fix" }] }), undefined);
|
|
1255
|
+
});
|
|
1256
|
+
|
|
1257
|
+
const waveResult = (id: string, results: { agent: string; exitCode: number }[], isError = false) => ({
|
|
1258
|
+
toolName: "subagent",
|
|
1259
|
+
toolCallId: id,
|
|
1260
|
+
isError,
|
|
1261
|
+
content: [],
|
|
1262
|
+
details: { results },
|
|
1263
|
+
});
|
|
1264
|
+
const okWave = (id: string) => waveResult(id, [{ agent: "implementer", exitCode: 0 }]);
|
|
1265
|
+
|
|
1266
|
+
const guardedHarness = (settings: unknown, extra: unknown[] = []) =>
|
|
1267
|
+
harness({ cwd: tempCwd(settings), branch: verifyBranch([subagentResult(["conformance-reviewer"]), ...extra]) });
|
|
1268
|
+
|
|
1269
|
+
test("cap guard: default 3 waves pass, the fourth is blocked with the escalation text", async () => {
|
|
1270
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { enforce: true } } });
|
|
1271
|
+
await h.emit("session_start");
|
|
1272
|
+
for (let i = 1; i <= 3; i++) {
|
|
1273
|
+
assert.equal(await firstCallResult(h, `c${i}`, implementerWave(2)), undefined, `round ${i} passes`);
|
|
1274
|
+
await h.emitEvent("tool_result", okWave(`c${i}`));
|
|
1275
|
+
}
|
|
1276
|
+
const blocked = await firstCallResult(h, "c4", implementerWave(1));
|
|
1277
|
+
assert.equal(blocked?.block, true);
|
|
1278
|
+
assert.match(blocked?.reason ?? "", /fix round 3 of 3 already used; escalate to the human/);
|
|
1279
|
+
});
|
|
1280
|
+
|
|
1281
|
+
test("cap guard: maxFixRounds 0 blocks the first wave; 1 blocks the second; a chain implementer is cap-checked and counted", async () => {
|
|
1282
|
+
const zero = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1283
|
+
await zero.emit("session_start");
|
|
1284
|
+
assert.equal((await firstCallResult(zero, "c1", implementerWave(1)))?.block, true);
|
|
1285
|
+
|
|
1286
|
+
const one = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 1 } } });
|
|
1287
|
+
await one.emit("session_start");
|
|
1288
|
+
assert.equal(await firstCallResult(one, "c1", { chain: [{ agent: "implementer", task: "fix" }] }), undefined);
|
|
1289
|
+
await one.emitEvent("tool_result", okWave("c1"));
|
|
1290
|
+
assert.equal((await firstCallResult(one, "c2", implementerWave(1)))?.block, true);
|
|
1291
|
+
});
|
|
1292
|
+
|
|
1293
|
+
test("counter: results while closure review enforcement is off do not consume budget", async () => {
|
|
1294
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { enforce: false, maxFixRounds: 1 } } });
|
|
1295
|
+
await h.emit("session_start");
|
|
1296
|
+
await h.emitEvent("tool_result", okWave("c1"));
|
|
1297
|
+
await h.emitEvent("tool_result", okWave("c2"));
|
|
1298
|
+
writeFileSync(
|
|
1299
|
+
join(h.ctx.cwd, ".pi", "settings.json"),
|
|
1300
|
+
JSON.stringify({ piGauntlet: { closureReview: { enforce: true, maxFixRounds: 1 } } }),
|
|
1301
|
+
);
|
|
1302
|
+
assert.equal(await firstCallResult(h, "c3", implementerWave(1)), undefined);
|
|
1303
|
+
});
|
|
1304
|
+
|
|
1305
|
+
test("fix-loop guard and counter are dormant in subagent children", async () => {
|
|
1306
|
+
const priorDepth = process.env.PI_SUBAGENT_DEPTH;
|
|
1307
|
+
process.env.PI_SUBAGENT_DEPTH = "1";
|
|
1308
|
+
let live: ReturnType<typeof harness>;
|
|
1309
|
+
let replay: ReturnType<typeof harness>;
|
|
1310
|
+
try {
|
|
1311
|
+
live = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1312
|
+
replay = guardedHarness(
|
|
1313
|
+
{ piGauntlet: { closureReview: { maxFixRounds: 2 } } },
|
|
1314
|
+
[subagentResult(["implementer"]), subagentResult(["implementer"])],
|
|
1315
|
+
);
|
|
1316
|
+
} finally {
|
|
1317
|
+
if (priorDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
|
|
1318
|
+
else process.env.PI_SUBAGENT_DEPTH = priorDepth;
|
|
1319
|
+
}
|
|
1320
|
+
|
|
1321
|
+
await live.emit("session_start");
|
|
1322
|
+
assert.equal(await firstCallResult(live, "c1", loneImplementer()), undefined);
|
|
1323
|
+
assert.equal(await firstCallResult(live, "c2", implementerWave(1)), undefined);
|
|
1324
|
+
await live.emitEvent("tool_result", okWave("c2"));
|
|
1325
|
+
await replay.emit("session_start");
|
|
1326
|
+
assert.equal(await firstCallResult(replay, "c3", implementerWave(1)), undefined);
|
|
1327
|
+
});
|
|
1328
|
+
|
|
1329
|
+
test("cap guard: non-implementer dispatches never blocked; enforce false passes everything", async () => {
|
|
1330
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1331
|
+
await h.emit("session_start");
|
|
1332
|
+
assert.equal(await firstCallResult(h, "r1", { agent: "conformance-reviewer", task: "re-audit" }), undefined);
|
|
1333
|
+
assert.equal(await firstCallResult(h, "r2", { agent: "code-reviewer", task: "review" }), undefined);
|
|
1334
|
+
|
|
1335
|
+
const off = guardedHarness({ piGauntlet: { closureReview: { enforce: false, maxFixRounds: 0 } } });
|
|
1336
|
+
await off.emit("session_start");
|
|
1337
|
+
assert.equal(await firstCallResult(off, "c1", implementerWave(1)), undefined);
|
|
1338
|
+
assert.equal(await firstCallResult(off, "c2", loneImplementer()), undefined);
|
|
1339
|
+
});
|
|
1340
|
+
|
|
1341
|
+
test("counter: blocked, errored, empty, non-implementer, and out-of-window results do not count; non-zero exit does", async () => {
|
|
1342
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 1 } } });
|
|
1343
|
+
await h.emit("session_start");
|
|
1344
|
+
await h.emitEvent("tool_result", waveResult("e1", [{ agent: "implementer", exitCode: 0 }], true));
|
|
1345
|
+
await h.emitEvent("tool_result", waveResult("e2", []));
|
|
1346
|
+
await h.emitEvent("tool_result", waveResult("e3", [{ agent: "code-reviewer", exitCode: 0 }]));
|
|
1347
|
+
const blockedLone = await firstCallResult(h, "b1", loneImplementer());
|
|
1348
|
+
assert.equal(blockedLone?.block, true);
|
|
1349
|
+
assert.equal(await firstCallResult(h, "c1", implementerWave(1)), undefined, "budget untouched");
|
|
1350
|
+
await h.emitEvent("tool_result", waveResult("c1", [{ agent: "implementer", exitCode: 1 }]));
|
|
1351
|
+
assert.equal((await firstCallResult(h, "c2", implementerWave(1)))?.block, true);
|
|
1352
|
+
});
|
|
1353
|
+
|
|
1354
|
+
test("counter: a ship-phase implementer wave with the latch set does not count", async () => {
|
|
1355
|
+
const h = harness({
|
|
1356
|
+
cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 1 } } }),
|
|
1357
|
+
branch: verifyBranch([
|
|
1358
|
+
subagentResult(["conformance-reviewer"]),
|
|
1359
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete" })),
|
|
1360
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" })),
|
|
1361
|
+
subagentResult(["implementer"]),
|
|
1362
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress", ship: "in_progress" })),
|
|
1363
|
+
]),
|
|
1364
|
+
});
|
|
1365
|
+
await h.emit("session_start");
|
|
1366
|
+
assert.equal(await firstCallResult(h, "c1", implementerWave(1)), undefined);
|
|
1367
|
+
});
|
|
1368
|
+
|
|
1369
|
+
test("counter reset: start implement and reset zero it; start verify --force keeps it", async () => {
|
|
1370
|
+
const mk = () => guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 1 }, flowGuards: { enforce: false } } }, [subagentResult(["implementer"])]);
|
|
1371
|
+
|
|
1372
|
+
const force = mk();
|
|
1373
|
+
await force.emit("session_start");
|
|
1374
|
+
const forceTool = force.tools.find((t) => t.name === "phase_tracker")!;
|
|
1375
|
+
const forced = (await forceTool.execute("p1", { action: "start", phase: "verify", force: true }, undefined, undefined, force.ctx)) as { details: { error?: string } };
|
|
1376
|
+
assert.equal(forced.details.error, undefined);
|
|
1377
|
+
assert.equal((await firstCallResult(force, "c1", implementerWave(1)))?.block, true, "budget survives verify --force");
|
|
1378
|
+
|
|
1379
|
+
const impl = mk();
|
|
1380
|
+
await impl.emit("session_start");
|
|
1381
|
+
const implTool = impl.tools.find((t) => t.name === "phase_tracker")!;
|
|
1382
|
+
for (const [id, input] of [
|
|
1383
|
+
["p1", { action: "skip", phase: "verify", reason: "amendment" }],
|
|
1384
|
+
["p2", { action: "start", phase: "implement", force: true }],
|
|
1385
|
+
["p3", { action: "complete", phase: "implement" }],
|
|
1386
|
+
["p4", { action: "start", phase: "verify", force: true }],
|
|
1387
|
+
] as const) {
|
|
1388
|
+
const result = (await implTool.execute(id, input, undefined, undefined, impl.ctx)) as { details: { error?: string } };
|
|
1389
|
+
assert.equal(result.details.error, undefined);
|
|
1390
|
+
}
|
|
1391
|
+
assert.equal(await firstCallResult(impl, "c1", loneImplementer()), undefined, "latch cleared by start implement");
|
|
1392
|
+
await impl.emitEvent("tool_result", waveResult("audit", [{ agent: "conformance-reviewer", exitCode: 0 }]));
|
|
1393
|
+
assert.equal(await firstCallResult(impl, "c2", implementerWave(1)), undefined, "counter cleared by start implement");
|
|
1394
|
+
|
|
1395
|
+
const reset = mk();
|
|
1396
|
+
await reset.emit("session_start");
|
|
1397
|
+
const resetTool = reset.tools.find((t) => t.name === "phase_tracker")!;
|
|
1398
|
+
const resetResult = (await resetTool.execute("p1", { action: "reset" }, undefined, undefined, reset.ctx)) as { details: { error?: string } };
|
|
1399
|
+
assert.equal(resetResult.details.error, undefined);
|
|
1400
|
+
assert.equal(await firstCallResult(reset, "c1", loneImplementer()), undefined, "dormant after reset");
|
|
1401
|
+
});
|
|
1402
|
+
|
|
1403
|
+
test("replay: two implementer waves after the audit in verify restore fixRounds 2; the same waves in ship restore 0", async () => {
|
|
1404
|
+
const inVerify = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 2 } } }, [
|
|
1405
|
+
subagentResult(["implementer"]),
|
|
1406
|
+
subagentResult(["implementer", "code-reviewer"]),
|
|
1407
|
+
]);
|
|
1408
|
+
await inVerify.emit("session_start");
|
|
1409
|
+
assert.equal((await firstCallResult(inVerify, "c1", implementerWave(1)))?.block, true);
|
|
1410
|
+
|
|
1411
|
+
const inShip = harness({
|
|
1412
|
+
cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 2 } } }),
|
|
1413
|
+
branch: verifyBranch([
|
|
1414
|
+
subagentResult(["conformance-reviewer"]),
|
|
1415
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete" })),
|
|
1416
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" })),
|
|
1417
|
+
subagentResult(["implementer"]),
|
|
1418
|
+
subagentResult(["implementer"]),
|
|
1419
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress", ship: "in_progress" })),
|
|
1420
|
+
]),
|
|
1421
|
+
});
|
|
1422
|
+
await inShip.emit("session_start");
|
|
1423
|
+
assert.equal(await firstCallResult(inShip, "c1", implementerWave(1)), undefined);
|
|
1424
|
+
});
|
|
@@ -86,6 +86,17 @@ const qualifiesAsClosureDispatch = (details: unknown): boolean => {
|
|
|
86
86
|
return d.results.some((r) => r?.agent === "conformance-reviewer" && r?.exitCode === 0);
|
|
87
87
|
};
|
|
88
88
|
|
|
89
|
+
// One conformance fix round = one non-error subagent result carrying at least one
|
|
90
|
+
// implementer child, regardless of dispatch mode or per-child exit code (a failed
|
|
91
|
+
// implementer still spent the round; its retry is the next one). Async and
|
|
92
|
+
// management dispatches return results: [] and never count.
|
|
93
|
+
const isImplementerWave = (details: unknown, isError: boolean | undefined): boolean => {
|
|
94
|
+
if (isError === true) return false;
|
|
95
|
+
const d = details as { results?: { agent?: unknown }[] } | undefined;
|
|
96
|
+
if (!d || !Array.isArray(d.results) || d.results.length === 0) return false;
|
|
97
|
+
return d.results.some((r) => r?.agent === "implementer");
|
|
98
|
+
};
|
|
99
|
+
|
|
89
100
|
// Review-cadence guard (spec 2026-08-12-execution-fidelity-hardening): presence-only
|
|
90
101
|
// advisory ledger of the most recent completed implementer / spec-reviewer /
|
|
91
102
|
// code-reviewer dispatch. Agents completing in the SAME dispatch share a sequence
|
|
@@ -164,22 +175,14 @@ const branchBlockReason = (phase: Phase): string =>
|
|
|
164
175
|
"create/enter one with /skill:using-git-worktrees and run this there. " +
|
|
165
176
|
"To override, set piGauntlet.flowGuards.enforce: false.";
|
|
166
177
|
|
|
167
|
-
//
|
|
168
|
-
//
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
// single / tasks / chain / parallel dispatch shapes and collect the model of every
|
|
172
|
-
// conformance-reviewer entry (undefined = absent/empty/non-string). A bare omission
|
|
173
|
-
// is BLOCKED; an explicit model that differs from the configured one is WARNED
|
|
174
|
-
// (non-blocking), preserving the documented retry-with-fallback hatch while
|
|
175
|
-
// surfacing drift.
|
|
176
|
-
const conformanceModels = (input: unknown): (string | undefined)[] => {
|
|
177
|
-
const models: (string | undefined)[] = [];
|
|
178
|
-
const norm = (m: unknown) => (typeof m === "string" && m.trim() ? m.trim() : undefined);
|
|
178
|
+
// Walk the single / tasks / chain / parallel dispatch shapes and return every
|
|
179
|
+
// entry whose `agent` matches `name` (the raw node, so callers can read its model).
|
|
180
|
+
const collectAgents = (input: unknown, name: string): Record<string, unknown>[] => {
|
|
181
|
+
const hits: Record<string, unknown>[] = [];
|
|
179
182
|
const collect = (node: unknown): void => {
|
|
180
183
|
if (!node || typeof node !== "object") return;
|
|
181
184
|
const o = node as Record<string, unknown>;
|
|
182
|
-
if (o.agent ===
|
|
185
|
+
if (o.agent === name) hits.push(o);
|
|
183
186
|
for (const key of ["tasks", "chain", "parallel"] as const) {
|
|
184
187
|
const v = o[key];
|
|
185
188
|
if (Array.isArray(v)) for (const item of v) collect(item);
|
|
@@ -187,9 +190,27 @@ const conformanceModels = (input: unknown): (string | undefined)[] => {
|
|
|
187
190
|
}
|
|
188
191
|
};
|
|
189
192
|
collect(input);
|
|
190
|
-
return
|
|
193
|
+
return hits;
|
|
191
194
|
};
|
|
192
195
|
|
|
196
|
+
const conformanceModels = (input: unknown): (string | undefined)[] =>
|
|
197
|
+
collectAgents(input, "conformance-reviewer").map((o) =>
|
|
198
|
+
typeof o.model === "string" && o.model.trim() ? o.model.trim() : undefined,
|
|
199
|
+
);
|
|
200
|
+
|
|
201
|
+
// Conformance fix-loop dispatch guard (spec 2026-09-13-conformance-dispatch-guard):
|
|
202
|
+
// after the first audit, pi-cohort honours worktree: true only in tasks mode, so a
|
|
203
|
+
// lone implementer runs unisolated and yields no patch for the integrate step.
|
|
204
|
+
const loneImplementerBlockReason =
|
|
205
|
+
'Conformance fix loop: dispatch implementers as a one-task tasks wave (tasks: [{ agent: "implementer", worktree: true, ... }]); ' +
|
|
206
|
+
"a lone agent call runs unisolated and produces no worktree diff. " +
|
|
207
|
+
"To disable this gate, set piGauntlet.closureReview.enforce: false.";
|
|
208
|
+
|
|
209
|
+
const fixRoundCapBlockReason = (used: number, cap: number): string =>
|
|
210
|
+
`Conformance fix loop: fix round ${used} of ${cap} already used; escalate to the human ` +
|
|
211
|
+
"with the verdict trail instead of re-looping. " +
|
|
212
|
+
"To disable this gate, set piGauntlet.closureReview.enforce: false.";
|
|
213
|
+
|
|
193
214
|
const closureModelBlockReason = (model: string, missing: number, total: number): string =>
|
|
194
215
|
`Blocked: ${missing} of ${total} conformance-reviewer ${total === 1 ? "dispatch" : "entries"} ` +
|
|
195
216
|
`omitted a model while piGauntlet.closureReview.model is set to "${model}".\n` +
|
|
@@ -306,7 +327,23 @@ function formatStatus(phases: PhaseMap): string {
|
|
|
306
327
|
export default function (pi: ExtensionAPI) {
|
|
307
328
|
let phases: PhaseMap = emptyPhases();
|
|
308
329
|
let conformanceDispatched = false;
|
|
330
|
+
let fixRounds = 0;
|
|
309
331
|
let gauntletEntered = false;
|
|
332
|
+
// Shared by replay and the live tool_result hook so a resumed session enforces
|
|
333
|
+
// the same budget. Evaluated BEFORE the same result may set the latch, so the
|
|
334
|
+
// R0 audit result itself never counts as a wave.
|
|
335
|
+
const observeFixWave = (details: unknown, isError: boolean | undefined, ctx: ExtensionContext) => {
|
|
336
|
+
if (
|
|
337
|
+
!isSubagentChild &&
|
|
338
|
+
gauntletEntered &&
|
|
339
|
+
phases.verify.status === "in_progress" &&
|
|
340
|
+
conformanceDispatched &&
|
|
341
|
+
isImplementerWave(details, isError) &&
|
|
342
|
+
resolveClosureReview(loadGauntletSettings(ctx.cwd).gauntlet).enforce
|
|
343
|
+
) {
|
|
344
|
+
fixRounds += 1;
|
|
345
|
+
}
|
|
346
|
+
};
|
|
310
347
|
let planCheckStamp: PlanCheckStamp | undefined;
|
|
311
348
|
const attemptedRecoveryEdges = new Set<RecoveryEdge>();
|
|
312
349
|
|
|
@@ -382,6 +419,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
382
419
|
const reconstructState = (ctx: ExtensionContext) => {
|
|
383
420
|
phases = emptyPhases();
|
|
384
421
|
conformanceDispatched = false;
|
|
422
|
+
fixRounds = 0;
|
|
385
423
|
gauntletEntered = false;
|
|
386
424
|
planCheckStamp = undefined;
|
|
387
425
|
attemptedRecoveryEdges.clear();
|
|
@@ -409,9 +447,11 @@ export default function (pi: ExtensionAPI) {
|
|
|
409
447
|
gauntletEntered = nextGauntletEntered(gauntletEntered, details.action, details.phases.brainstorm.status);
|
|
410
448
|
if (details.action === "start" && details.phases.implement.status === "in_progress") {
|
|
411
449
|
conformanceDispatched = false;
|
|
450
|
+
fixRounds = 0;
|
|
412
451
|
}
|
|
413
452
|
if (details.action === "reset") {
|
|
414
453
|
conformanceDispatched = false;
|
|
454
|
+
fixRounds = 0;
|
|
415
455
|
planCheckStamp = undefined;
|
|
416
456
|
}
|
|
417
457
|
}
|
|
@@ -423,6 +463,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
423
463
|
planCheckStamp = undefined;
|
|
424
464
|
}
|
|
425
465
|
} else if (msg.toolName === "subagent") {
|
|
466
|
+
observeFixWave(msg.details, msg.isError, ctx);
|
|
426
467
|
if (qualifiesAsClosureDispatch(msg.details)) conformanceDispatched = true;
|
|
427
468
|
observeCadence(msg.details);
|
|
428
469
|
} else if (msg.toolName === "plan_tracker") {
|
|
@@ -492,10 +533,20 @@ export default function (pi: ExtensionAPI) {
|
|
|
492
533
|
// subagent dispatch never loads settings or leaks a settingsErrorWarning onto its result.
|
|
493
534
|
// Inline-matched (not called) so closureEnforced() stays a lazy second conjunct.
|
|
494
535
|
if (event.toolName === "subagent" && gauntletEntered && closureEnforced()) {
|
|
495
|
-
|
|
496
|
-
//
|
|
497
|
-
// (action: list/get/create/update/delete/status/...) execute nothing, so skip them.
|
|
536
|
+
// Only execution-mode dispatches carry a model or run agents; management/control
|
|
537
|
+
// modes (action: list/get/create/update/delete/status/...) execute nothing, so skip them.
|
|
498
538
|
const hasAction = !!(event.input as { action?: unknown })?.action;
|
|
539
|
+
// Fix-loop window: verify in progress and the R0 audit observed. Read directly,
|
|
540
|
+
// not via activeGuardPhase() - verify is not a GUARD_PHASES member.
|
|
541
|
+
const fixLoopWindow = !isSubagentChild && !hasAction && phases.verify.status === "in_progress" && conformanceDispatched;
|
|
542
|
+
if (fixLoopWindow && (event.input as { agent?: unknown })?.agent === "implementer") {
|
|
543
|
+
return { block: true, reason: loneImplementerBlockReason };
|
|
544
|
+
}
|
|
545
|
+
if (fixLoopWindow && collectAgents(event.input, "implementer").length > 0) {
|
|
546
|
+
const cap = resolveClosureReview(g()).maxFixRounds;
|
|
547
|
+
if (fixRounds >= cap) return { block: true, reason: fixRoundCapBlockReason(fixRounds, cap) };
|
|
548
|
+
}
|
|
549
|
+
const model = closureReviewModel();
|
|
499
550
|
if (model && !hasAction) {
|
|
500
551
|
const configured = model.trim();
|
|
501
552
|
const models = conformanceModels(event.input);
|
|
@@ -635,8 +686,9 @@ export default function (pi: ExtensionAPI) {
|
|
|
635
686
|
return undefined;
|
|
636
687
|
});
|
|
637
688
|
|
|
638
|
-
pi.on("tool_result", async (event) => {
|
|
689
|
+
pi.on("tool_result", async (event, ctx) => {
|
|
639
690
|
if (event.toolName === "subagent") {
|
|
691
|
+
observeFixWave(event.details, event.isError, ctx);
|
|
640
692
|
if (qualifiesAsClosureDispatch(event.details)) conformanceDispatched = true;
|
|
641
693
|
observeCadence(event.details);
|
|
642
694
|
const warning = pendingGuardWarnings.get(event.toolCallId);
|
|
@@ -859,7 +911,10 @@ export default function (pi: ExtensionAPI) {
|
|
|
859
911
|
}
|
|
860
912
|
phases = { ...phases, [params.phase]: transitionPhaseState("in_progress") as PhaseState };
|
|
861
913
|
// A rewind must not inherit the prior verify's conformance latch.
|
|
862
|
-
if (params.phase === "implement")
|
|
914
|
+
if (params.phase === "implement") {
|
|
915
|
+
conformanceDispatched = false;
|
|
916
|
+
fixRounds = 0;
|
|
917
|
+
}
|
|
863
918
|
gauntletEntered = nextGauntletEntered(gauntletEntered, "start", phases.brainstorm.status);
|
|
864
919
|
firedGuards.clear();
|
|
865
920
|
updateWidget(ctx);
|
|
@@ -1024,6 +1079,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
1024
1079
|
PHASES.map((p) => [p, transitionPhaseState("pending")]),
|
|
1025
1080
|
) as PhaseMap;
|
|
1026
1081
|
conformanceDispatched = false;
|
|
1082
|
+
fixRounds = 0;
|
|
1027
1083
|
gauntletEntered = nextGauntletEntered(gauntletEntered, "reset", phases.brainstorm.status);
|
|
1028
1084
|
firedGuards.clear();
|
|
1029
1085
|
updateWidget(ctx);
|
package/package.json
CHANGED
|
@@ -105,7 +105,7 @@ When agents return:
|
|
|
105
105
|
|
|
106
106
|
**If integrated changes apply cleanly but the suite fails (semantic conflict):** agents made incompatible assumptions across disjoint files (renamed symbol, changed shape). Diagnose the incompatible pair and re-run the offending task sequentially on the integrated HEAD.
|
|
107
107
|
|
|
108
|
-
**If some agents failed:** Integrate successful agents first (commit their work). Then retry the failed agent with fresh context that includes the integrated changes.
|
|
108
|
+
**If some agents failed:** Integrate successful agents first (commit their work). Then retry the failed agent with fresh context that includes the integrated changes. Inside the conformance loop the retry is a one-task `tasks` wave, never a lone `agent: "implementer"`.
|
|
109
109
|
|
|
110
110
|
## Fix fan-out
|
|
111
111
|
|
|
@@ -142,29 +142,26 @@ prerequisites hold.
|
|
|
142
142
|
Per round:
|
|
143
143
|
|
|
144
144
|
1. **Synchronize gap tasks** — append only a genuinely new gap that is entering remediation, named `Gn: <gap origin clause verbatim, truncated>`; never `init`. Find existing gaps by their exact `Gn:` prefix and reuse that index even if origin wording changes. Carried-OPEN inventory-only gaps add nothing. Before dispatch, mark every remediated gap's existing index `in_progress`; a re-audit needing more work reopens that same `Gn` index. The lifecycle traces `[T1,T2]`, then `[T1,T2,G1]`, then `[T1,T2,G1,G2]`; no test-retry or review-round wrapper task.
|
|
145
|
-
2. **Fix
|
|
146
|
-
group of ≥ 2 gaps (per the report's `Parallel-safe:` line) fixes in one parallel
|
|
147
|
-
foreground dispatch — one `implementer` per gap (fresh context, `async: false`,
|
|
148
|
-
`worktree: true`, `cwd` = the conformance worktree, task = the gap block verbatim
|
|
149
|
-
with `touched-files` as the ownership boundary). For an `UNAUTHORIZED` `fix` gap
|
|
145
|
+
2. **Fix wave** — select gaps greedily in `Gn` order: take each `fix` gap unless a gap it `conflicts` with (per the report's `Parallel-safe:` line) is already taken; the certificate's `disjoint` grouping is ignored, and a certificate still malformed after the one re-ask means every gap `conflicts` with every other. Held gaps carry to the next round; the wave is never empty while an eligible `fix` gap exists. Dispatch **one** call - `subagent({ context: "fresh", async: false, tasks: [...] })` - with one `implementer` task per selected gap (`worktree: true`, `cwd` = the conformance worktree, task = the gap block verbatim with `touched-files` as the ownership boundary). A single gap is a one-task `tasks` call; a lone `agent: "implementer"` call never appears in this loop. For an `UNAUTHORIZED` `fix` gap
|
|
150
146
|
whose `evidence` opens with the over-spec provenance (`spec "<section>" - "<clause>" (over-spec)`), the orchestrator adds the spec path to that gap's `touched-files` before dispatch,
|
|
151
147
|
so the implementer deletes the surface **and** the clause/AC line in the same
|
|
152
148
|
fix commit; the re-audit then has no `Rn` for it and no `MISSING` echo. The dispatch adds `SCOPED_TEST_COMMANDS`
|
|
153
|
-
to the gap block: the gap-relevant plan-declared commands, or `none
|
|
154
|
-
test gate owns execution). `conflicts` pairs serialize. Gaps outside any ≥ 2-ID
|
|
155
|
-
`disjoint` group run sequentially as before. Then dispatch foreground `spec-reviewer`
|
|
156
|
-
per gap on the gap-block reference contract below.
|
|
149
|
+
to the gap block: the gap-relevant plan-declared commands, or `none`; on an ad-hoc no-plan path, the project's canonical test command.
|
|
157
150
|
3. **Integrate** serially via `git apply` onto the worktree HEAD, one gap's
|
|
158
|
-
patch at a time.
|
|
151
|
+
patch at a time. Commit each per-gap fix with the message **`conformance fix Gn`** (durable,
|
|
152
|
+
`git log`-readable pre-squash) so the finish gate and any revert can identify
|
|
153
|
+
auto-applied fixes; a Convergence repair wave commits as one **`conformance fix CR`**.
|
|
154
|
+
Failure handling is inherited verbatim from
|
|
159
155
|
`dispatching-parallel-agents` "Review and Integrate": textual conflict →
|
|
160
156
|
re-run one agent sequentially with the other's integrated changes as
|
|
161
157
|
context; semantic conflict (applies clean, suite fails) → re-run the
|
|
162
158
|
offending task sequentially on integrated HEAD; a failed agent → integrate
|
|
163
159
|
the successes, then retry the failure with fresh context including the
|
|
164
|
-
integrated changes.
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
160
|
+
integrated changes. Every re-run or retry inside this loop is itself a
|
|
161
|
+
one-task `tasks` wave (never a lone `agent: "implementer"`) and counts as a
|
|
162
|
+
wave against `maxFixRounds`. A `BLOCKED`/`NEEDS_CONTEXT` return surfaces to the user.
|
|
163
|
+
4. **Scoped tests** on the integrated tree: the round's `SCOPED_TEST_COMMANDS` union. A failure re-enters the failure-handling rules above.
|
|
164
|
+
5. **Re-audit**: foreground re-dispatch `conformance-reviewer` with `async: false` over the fixes **plus** the
|
|
168
165
|
regression guard (any prior-`DELIVERED` requirement whose `evidence` file
|
|
169
166
|
the fix diff touched). Pass the full prior conformance report (every row,
|
|
170
167
|
including DELIVERED rows and their `evidence` `file:line`) and the round's
|
|
@@ -173,18 +170,26 @@ Per round:
|
|
|
173
170
|
`model:` when it is `undefined` to inherit the parent's model. Inside a
|
|
174
171
|
brainstorming-entered flow, the phase-tracker closure guard blocks a dispatch
|
|
175
172
|
that omits `model:` when `closureReview.model` is set, and warns (non-blocking)
|
|
176
|
-
on one whose model differs.
|
|
177
|
-
|
|
173
|
+
on one whose model differs. Mark every `Gn` the re-audit reports `DELIVERED`
|
|
174
|
+
`complete`; open ones stay `in_progress`.
|
|
175
|
+
6. **Converge or continue**: verdict `CONFORMS` → Convergence below. Open gaps
|
|
178
176
|
within the cap → re-partition (per the rule above) and start the next
|
|
179
177
|
round. Cap (`gauntlet_setting({ key: "closureReview" }).maxFixRounds`,
|
|
180
|
-
default `
|
|
181
|
-
with an open `fix` gap → **escalate to the human** with the per-gap
|
|
178
|
+
default `3`, floors negatives at `0`, coerces non-integers to `3`) reached
|
|
179
|
+
with an open `fix` gap or repair item → **escalate to the human** with the per-gap
|
|
182
180
|
round-by-round verdict trail. Escalation is the sole non-completing
|
|
183
|
-
terminal state — no silent re-loop, no auto-ship.
|
|
181
|
+
terminal state — no silent re-loop, no auto-ship. Inside a
|
|
182
|
+
brainstorming-entered flow the phase tracker enforces both rules at
|
|
183
|
+
tool-call time: a lone `agent: "implementer"` dispatch is blocked, and the
|
|
184
|
+
wave after the cap is blocked (`closureReview.enforce: false` disables).
|
|
184
185
|
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
186
|
+
**Convergence** — runs after R0 `CONFORMS` and after every `CONFORMS` re-audit. `r0-head` is HEAD when the loop was entered (the R0 dispatch, or the finish-time `fix-now` entry) - the parent of the oldest `conformance fix` commit; `audited-base` stays the last audit's HEAD SHA.
|
|
187
|
+
|
|
188
|
+
a. Run the full plan-header `Verification` set once (ad-hoc: the project's canonical test command). After R0 `CONFORMS` with no round run, the pre-R0 full run counts.
|
|
189
|
+
b. Dispatch `code-reviewer` directly (foreground, `async: false`, `SCOPED_TEST_COMMANDS: none`) over `git diff <r0-head>..HEAD`; never via `/skill:requesting-code-review`. An empty diff is nothing to review - no dispatch.
|
|
190
|
+
c. Repair items = every failing command from a + every Critical/Moderate finding from b (`Behaviour-change: yes` included; the re-audit is its origin check). None → write the closure block; done. Any at the cap → escalate per step 6 with the test/CR trail. Any under the cap → re-enter step 2 as a one-task `tasks` wave: one `implementer` whose task is every repair item verbatim (ownership boundary = the files in `git diff <r0-head>..HEAD`, the CR findings' `touched-files`, and the files each failing command's output names, `SCOPED_TEST_COMMANDS: none`, no `Gn` tracker task, no gap selection), integrate as one `conformance fix CR`, re-audit, then Convergence again. That wave counts against `maxFixRounds`. Later Convergence CRs keep the same `<r0-head>..HEAD` range.
|
|
191
|
+
|
|
192
|
+
`conformance fix CR` is not a gap fix: it is absent from the `auto-applied fix commits` index and has no `revert conformance fix Gn` action at the finish gate.
|
|
188
193
|
|
|
189
194
|
**`maxFixRounds: 0`**: skip this loop entirely; every `recommended: fix` gap is
|
|
190
195
|
carried OPEN to the finish gate per the precondition-unavailable
|
|
@@ -194,22 +199,6 @@ auto-fix, so treat `fix` gaps like any other deferred gap — unlike a cap > 0
|
|
|
194
199
|
that is *exhausted*, which escalates mid-verify because the loop tried and
|
|
195
200
|
could not converge.
|
|
196
201
|
|
|
197
|
-
### `spec-reviewer` gap-block reference contract
|
|
198
|
-
|
|
199
|
-
Per-gap `spec-reviewer` in step 2 above is a **pre-integration mechanical
|
|
200
|
-
check**, distinct from the round-level re-audit in step 6 (which still
|
|
201
|
-
references the *origin* — spec + original prompt — unchanged). Frame the
|
|
202
|
-
per-gap dispatch against the **gap block**, not a plan task:
|
|
203
|
-
|
|
204
|
-
- **Requirement** = the gap's `origin` + `remediation` (what must be true
|
|
205
|
-
after the fix).
|
|
206
|
-
- **Closure proof** = the patch satisfies that requirement within the gap's
|
|
207
|
-
`touched-files` — nothing missing, nothing extra.
|
|
208
|
-
- **Output** = `spec-reviewer`'s normal MATCH/DRIFT verdict, referenced to the
|
|
209
|
-
gap block instead of a plan task.
|
|
210
|
-
|
|
211
|
-
This is a task-framing contract in the dispatch, not a new persona.
|
|
212
|
-
|
|
213
202
|
## Concern decomposition
|
|
214
203
|
|
|
215
204
|
The main verification orchestrator — **not** `conformance-reviewer` — decomposes
|
|
@@ -459,15 +448,15 @@ concerns into one gap-scoped fix contract:
|
|
|
459
448
|
boundary).
|
|
460
449
|
|
|
461
450
|
Rescoped, accepted, and followed-up sibling concerns are **excluded** from the
|
|
462
|
-
projection. The `implementer`
|
|
463
|
-
|
|
464
|
-
|
|
465
|
-
The projected task runs the
|
|
466
|
-
`conformance fix Gn` commit name,
|
|
467
|
-
|
|
468
|
-
|
|
469
|
-
|
|
470
|
-
|
|
451
|
+
projection. The `implementer` receives this projected contract in place of the
|
|
452
|
+
original whole-gap block.
|
|
453
|
+
|
|
454
|
+
The projected task runs the round and Convergence above with `r0-head` = HEAD at
|
|
455
|
+
this entry: it retains the gap-level `conformance fix Gn` commit name, re-audits
|
|
456
|
+
against the amended spec, and reenters the gate only if concerns remain.
|
|
457
|
+
Gap-level revert stays available through the flat commit index. The gate records
|
|
458
|
+
the final result by concern ID and title before showing branch integration
|
|
459
|
+
options.
|
|
471
460
|
|
|
472
461
|
## Checklist
|
|
473
462
|
|
|
@@ -22,7 +22,7 @@ does not recurse into their leaves.
|
|
|
22
22
|
whole-object, a repo file that sets only *one* leaf of a key silently drops the
|
|
23
23
|
preset's other leaves for that key. A repo `closureReview: { "model": "..." }` with
|
|
24
24
|
no `enforce`/`maxFixRounds` makes those fall back to their code defaults
|
|
25
|
-
(`enforce` true, `maxFixRounds`
|
|
25
|
+
(`enforce` true, `maxFixRounds` 3), **not** to the preset's values. Define every
|
|
26
26
|
leaf you care about together in the file that owns the key. (Sibling keys are
|
|
27
27
|
unaffected - only the key the repo redefines is replaced.)
|
|
28
28
|
|