pi-gauntlet 5.5.1 → 5.5.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,20 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.5.3 - 2026-09-13
4
+
5
+ - `code-reviewer`: verdict is a stated function of severity - a Critical or Moderate finding means `FIX_FIRST`, Minor-only and clean reports mean `SHIP`, matching the orchestrating skills. (#30)
6
+ - `phase-tracker`: two new closure-review blocks inside a brainstorming-entered flow once `verify` is in progress and a conformance audit has been observed - a lone top-level `agent: "implementer"` `subagent` dispatch is blocked (one-task `tasks` wave is the only isolated shape), and an implementer dispatch past `closureReview.maxFixRounds` is blocked with the escalation text. Rounds are counted from non-error results carrying an implementer child; the counter resets with the audit latch. `closureReview.enforce: false` disables both.
7
+ - `closureReview.maxFixRounds` default 2 -> 3.
8
+ - `conformance-check.md` / `dispatching-parallel-agents`: retries inside the conformance loop are one-task `tasks` waves and count against the cap.
9
+
10
+ ## v5.5.2 - 2026-09-13
11
+
12
+ - `release.sh <level>` promotes the CHANGELOG `## Unreleased` section to `## vX.Y.Z - <date>` and commits it with `package.json` in the single `Release X.Y.Z` commit, so `patch`/`minor`/`major` now work here (the `current`-only path is gone). New CONFIG field `CHANGELOG_HEADING`.
13
+ - Release skill: a user instruction naming the level is the approval - no proposal step or re-confirmation; bundled follow-ups run after `verify`.
14
+ - AGENTS.md rewritten to always-on essentials plus routing; shared core bumped to v3. Persona frontmatter knobs table and pin rationale moved to `doc/personas.md`. Gold rule scoped to agent-initiated writes; a user instruction naming the write is its confirmation.
15
+ - Added `.pi/gauntlet-overrides.md` (`tracker: github`, release path, write-gate carve-out for user-named writes).
16
+ - `verification-before-completion/reference/conformance-check.md`: the conformance fix round is one parallel `implementer` `tasks` wave (greedy `Gn` selection, `conflicts` partners held) -> `git apply` -> scoped tests -> delta re-audit. Per-gap `spec-reviewer`, per-round `code-reviewer`, and the per-round full test set are gone. New **Convergence** step after every `CONFORMS`: full `Verification` set once, one direct `code-reviewer` over `git diff <r0-head>..HEAD`; repairs re-enter the round as a `conformance fix CR` wave that counts against `maxFixRounds`. Finish grammar, personas, and settings unchanged.
17
+
3
18
  ## v5.5.1 - 2026-09-10
4
19
 
5
20
  - `linear`: copyable, version-scoped recovery for attachment-download 401s resolves the decrypted credential through linearis instead of reading encrypted token storage. Restricts credential delivery to HTTPS Linear uploads, rejects redirects, and checks downloaded bytes; an offline regression executes the documented example.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: code-reviewer
3
- description: Production-readiness code review with prioritized findings (Critical blocks merge, Minor is a nit). Read-only — does not edit.
3
+ description: Production-readiness code review with prioritized findings (Critical and Moderate block merge, Minor is a nit). Read-only — does not edit.
4
4
  tools: read, grep, find, ls, bash
5
5
  defaultContext: fresh
6
6
  inheritProjectContext: true
@@ -44,11 +44,15 @@ Parallel-safe: F1,F3 disjoint; F2 conflicts F1 (both touch auth.ts)
44
44
  Behaviour-change: yes | no
45
45
  ```
46
46
 
47
+ Severity decides SHIP versus FIX_FIRST. A Critical or Moderate finding means
48
+ FIX_FIRST. Minor-only findings and clean reports mean SHIP. REJECT overrides
49
+ both: refuse a change that must not land at all.
50
+
47
51
  Severity:
48
52
 
49
53
  - **Critical** — must fix before merge (data loss, security, broken correctness on a common path, broken contract).
50
- - **Moderate** — should fix; open for discussion (significant but not strictly blocking).
51
- - **Minor** — nit, style, preference, suggestion.
54
+ - **Moderate** — must fix before merge (significant defect or drift that does not rise to Critical).
55
+ - **Minor** — nit, style, preference, suggestion; the only severity declinable without a fix round or re-review.
52
56
 
53
57
  Label every finding with a globally unique `F1..Fn` ID (no restart per severity),
54
58
  and a `touched-files:`/`touched-resources:` pair (files/resources a fix would
@@ -122,11 +122,12 @@ test("closureReview: enforce default true; false only when explicitly false", ()
122
122
  assert.equal(resolveClosureReview({ closureReview: { enforce: true } }).enforce, true);
123
123
  });
124
124
 
125
- test("closureReview: maxFixRounds default 2, <0 -> 0, non-int -> 2", () => {
126
- assert.equal(resolveClosureReview({}).maxFixRounds, 2);
125
+ test("closureReview: maxFixRounds default 3, <0 -> 0, non-int -> 3", () => {
126
+ assert.equal(resolveClosureReview({}).maxFixRounds, 3);
127
127
  assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 5 } }).maxFixRounds, 5);
128
+ assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 0 } }).maxFixRounds, 0);
128
129
  assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: -3 } }).maxFixRounds, 0);
129
- assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 1.5 } }).maxFixRounds, 2);
130
+ assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 1.5 } }).maxFixRounds, 3);
130
131
  });
131
132
 
132
133
  test("flowGuards: defaults + overrides", () => {
@@ -84,7 +84,7 @@ export function resolveClosureReview(g: PiGauntlet): ClosureReviewResolved {
84
84
  const model = nonEmptyString(cr?.model) ? cr!.model.trim() : undefined;
85
85
  const enforce = cr?.enforce !== false;
86
86
  const raw = cr?.maxFixRounds;
87
- const maxFixRounds = typeof raw === "number" && Number.isInteger(raw) ? (raw < 0 ? 0 : raw) : 2;
87
+ const maxFixRounds = typeof raw === "number" && Number.isInteger(raw) ? (raw < 0 ? 0 : raw) : 3;
88
88
  return { model, enforce, maxFixRounds };
89
89
  }
90
90
 
@@ -6,9 +6,14 @@ import { tmpdir } from "node:os";
6
6
  import { isAbsolute, join } from "node:path";
7
7
  import registerPhaseTracker from "./phase-tracker.ts";
8
8
 
9
+ const originalSubagentDepth = process.env.PI_SUBAGENT_DEPTH;
10
+ process.env.PI_SUBAGENT_DEPTH = "0";
11
+
9
12
  const tempDirs: string[] = [];
10
13
  after(() => {
11
14
  for (const dir of tempDirs) rmSync(dir, { recursive: true, force: true });
15
+ if (originalSubagentDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
16
+ else process.env.PI_SUBAGENT_DEPTH = originalSubagentDepth;
12
17
  });
13
18
 
14
19
  const tempCwd = (settings?: unknown) => {
@@ -529,7 +534,8 @@ test("implement-phase commit with implementer newer than both reviewers warns",
529
534
  const warned = (await h.emitEvent("tool_result", commitResult("c1")))[0] as { content: { text: string }[] };
530
535
  assert.match(warned.content[0].text, /no spec-reviewer or code-reviewer observed/);
531
536
  } finally {
532
- if (priorDepth !== undefined) process.env.PI_SUBAGENT_DEPTH = priorDepth;
537
+ if (priorDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
538
+ else process.env.PI_SUBAGENT_DEPTH = priorDepth;
533
539
  }
534
540
  });
535
541
 
@@ -1164,3 +1170,255 @@ test("gauntlet_setting escalationLoop: setting absent -> ctx-derived main-loop m
1164
1170
  const res = (await set.tools.find((t) => t.name === "gauntlet_setting")!.execute("g2", { key: "escalationLoop" }, undefined, undefined, set.ctx)) as { details: { implModel?: string } };
1165
1171
  assert.equal(res.details.implModel, "p/strong:high");
1166
1172
  });
1173
+
1174
+ // --- Conformance fix-loop dispatch guard (spec 2026-09-13-conformance-dispatch-guard) ---
1175
+
1176
+ const verifyBranch = (extra: unknown[] = []) => [
1177
+ ...implementBranch(),
1178
+ phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete" })),
1179
+ phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress" })),
1180
+ ...extra,
1181
+ ];
1182
+
1183
+ const subagentCall = (id: string, input: unknown) => ({ toolName: "subagent", toolCallId: id, input });
1184
+ const loneImplementer = (extra: Record<string, unknown> = {}) => ({ agent: "implementer", task: "fix G1", ...extra });
1185
+ const implementerWave = (n = 1) => ({
1186
+ tasks: Array.from({ length: n }, (_, i) => ({ agent: "implementer", task: `fix G${i + 1}`, worktree: true })),
1187
+ });
1188
+
1189
+ const firstCallResult = async (h: ReturnType<typeof harness>, id: string, input: unknown) =>
1190
+ (await h.emitEvent("tool_call", subagentCall(id, input)))[0] as { block?: boolean; reason?: string } | undefined;
1191
+
1192
+ test("shape guard: lone implementer blocked in verify after the audit, with and without closureReview.model, async or not", async () => {
1193
+ for (const settings of [
1194
+ { piGauntlet: { closureReview: { enforce: true } } },
1195
+ { piGauntlet: { closureReview: { enforce: true, model: "x/y" } } },
1196
+ ]) {
1197
+ const h = harness({ cwd: tempCwd(settings), branch: verifyBranch([subagentResult(["conformance-reviewer"])]) });
1198
+ await h.emit("session_start");
1199
+ const sync = await firstCallResult(h, "s1", loneImplementer());
1200
+ assert.equal(sync?.block, true);
1201
+ assert.match(sync?.reason ?? "", /dispatch implementers as a one-task tasks wave/);
1202
+ assert.match(sync?.reason ?? "", /piGauntlet\.closureReview\.enforce: false/);
1203
+ const async = await firstCallResult(h, "s2", loneImplementer({ async: true }));
1204
+ assert.equal(async?.block, true);
1205
+ }
1206
+ });
1207
+
1208
+ test("shape guard: lone implementer passes before the audit, in implement, in ship, and on management calls", async () => {
1209
+ const preLatch = harness({ branch: verifyBranch() });
1210
+ await preLatch.emit("session_start");
1211
+ assert.equal(await firstCallResult(preLatch, "s1", loneImplementer()), undefined);
1212
+
1213
+ const implement = harness({ branch: implementBranch([subagentResult(["conformance-reviewer"])]) });
1214
+ await implement.emit("session_start");
1215
+ assert.equal(await firstCallResult(implement, "s1", loneImplementer()), undefined);
1216
+
1217
+ const ship = harness({
1218
+ branch: verifyBranch([
1219
+ subagentResult(["conformance-reviewer"]),
1220
+ phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete" })),
1221
+ phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" })),
1222
+ ]),
1223
+ });
1224
+ await ship.emit("session_start");
1225
+ assert.equal(await firstCallResult(ship, "s1", loneImplementer()), undefined);
1226
+
1227
+ const mgmt = harness({ branch: verifyBranch([subagentResult(["conformance-reviewer"])]) });
1228
+ await mgmt.emit("session_start");
1229
+ assert.equal(await firstCallResult(mgmt, "s1", { action: "status", agent: "implementer" }), undefined);
1230
+ });
1231
+
1232
+ test("shape guard: dormant when the flow was never entered (cold start verify) and when closureReview.enforce is false", async () => {
1233
+ const cold = harness({
1234
+ cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 0 } } }),
1235
+ branch: [phaseResult("start", phases({ verify: "in_progress" })), subagentResult(["conformance-reviewer"])],
1236
+ });
1237
+ await cold.emit("session_start");
1238
+ assert.equal(await firstCallResult(cold, "s1", loneImplementer()), undefined);
1239
+ assert.equal(await firstCallResult(cold, "s2", implementerWave(1)), undefined);
1240
+ assert.equal(await firstCallResult(cold, "s3", { chain: [{ agent: "implementer", task: "fix" }] }), undefined);
1241
+
1242
+ const off = harness({
1243
+ cwd: tempCwd({ piGauntlet: { closureReview: { enforce: false } } }),
1244
+ branch: verifyBranch([subagentResult(["conformance-reviewer"])]),
1245
+ });
1246
+ await off.emit("session_start");
1247
+ assert.equal(await firstCallResult(off, "s1", loneImplementer()), undefined);
1248
+ });
1249
+
1250
+ test("shape guard: a one-task tasks wave and a chain step are not the lone shape", async () => {
1251
+ const h = harness({ branch: verifyBranch([subagentResult(["conformance-reviewer"])]) });
1252
+ await h.emit("session_start");
1253
+ assert.equal(await firstCallResult(h, "s1", implementerWave(1)), undefined);
1254
+ assert.equal(await firstCallResult(h, "s2", { chain: [{ agent: "implementer", task: "fix" }] }), undefined);
1255
+ });
1256
+
1257
+ const waveResult = (id: string, results: { agent: string; exitCode: number }[], isError = false) => ({
1258
+ toolName: "subagent",
1259
+ toolCallId: id,
1260
+ isError,
1261
+ content: [],
1262
+ details: { results },
1263
+ });
1264
+ const okWave = (id: string) => waveResult(id, [{ agent: "implementer", exitCode: 0 }]);
1265
+
1266
+ const guardedHarness = (settings: unknown, extra: unknown[] = []) =>
1267
+ harness({ cwd: tempCwd(settings), branch: verifyBranch([subagentResult(["conformance-reviewer"]), ...extra]) });
1268
+
1269
+ test("cap guard: default 3 waves pass, the fourth is blocked with the escalation text", async () => {
1270
+ const h = guardedHarness({ piGauntlet: { closureReview: { enforce: true } } });
1271
+ await h.emit("session_start");
1272
+ for (let i = 1; i <= 3; i++) {
1273
+ assert.equal(await firstCallResult(h, `c${i}`, implementerWave(2)), undefined, `round ${i} passes`);
1274
+ await h.emitEvent("tool_result", okWave(`c${i}`));
1275
+ }
1276
+ const blocked = await firstCallResult(h, "c4", implementerWave(1));
1277
+ assert.equal(blocked?.block, true);
1278
+ assert.match(blocked?.reason ?? "", /fix round 3 of 3 already used; escalate to the human/);
1279
+ });
1280
+
1281
+ test("cap guard: maxFixRounds 0 blocks the first wave; 1 blocks the second; a chain implementer is cap-checked and counted", async () => {
1282
+ const zero = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
1283
+ await zero.emit("session_start");
1284
+ assert.equal((await firstCallResult(zero, "c1", implementerWave(1)))?.block, true);
1285
+
1286
+ const one = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 1 } } });
1287
+ await one.emit("session_start");
1288
+ assert.equal(await firstCallResult(one, "c1", { chain: [{ agent: "implementer", task: "fix" }] }), undefined);
1289
+ await one.emitEvent("tool_result", okWave("c1"));
1290
+ assert.equal((await firstCallResult(one, "c2", implementerWave(1)))?.block, true);
1291
+ });
1292
+
1293
+ test("counter: results while closure review enforcement is off do not consume budget", async () => {
1294
+ const h = guardedHarness({ piGauntlet: { closureReview: { enforce: false, maxFixRounds: 1 } } });
1295
+ await h.emit("session_start");
1296
+ await h.emitEvent("tool_result", okWave("c1"));
1297
+ await h.emitEvent("tool_result", okWave("c2"));
1298
+ writeFileSync(
1299
+ join(h.ctx.cwd, ".pi", "settings.json"),
1300
+ JSON.stringify({ piGauntlet: { closureReview: { enforce: true, maxFixRounds: 1 } } }),
1301
+ );
1302
+ assert.equal(await firstCallResult(h, "c3", implementerWave(1)), undefined);
1303
+ });
1304
+
1305
+ test("fix-loop guard and counter are dormant in subagent children", async () => {
1306
+ const priorDepth = process.env.PI_SUBAGENT_DEPTH;
1307
+ process.env.PI_SUBAGENT_DEPTH = "1";
1308
+ let live: ReturnType<typeof harness>;
1309
+ let replay: ReturnType<typeof harness>;
1310
+ try {
1311
+ live = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
1312
+ replay = guardedHarness(
1313
+ { piGauntlet: { closureReview: { maxFixRounds: 2 } } },
1314
+ [subagentResult(["implementer"]), subagentResult(["implementer"])],
1315
+ );
1316
+ } finally {
1317
+ if (priorDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
1318
+ else process.env.PI_SUBAGENT_DEPTH = priorDepth;
1319
+ }
1320
+
1321
+ await live.emit("session_start");
1322
+ assert.equal(await firstCallResult(live, "c1", loneImplementer()), undefined);
1323
+ assert.equal(await firstCallResult(live, "c2", implementerWave(1)), undefined);
1324
+ await live.emitEvent("tool_result", okWave("c2"));
1325
+ await replay.emit("session_start");
1326
+ assert.equal(await firstCallResult(replay, "c3", implementerWave(1)), undefined);
1327
+ });
1328
+
1329
+ test("cap guard: non-implementer dispatches never blocked; enforce false passes everything", async () => {
1330
+ const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
1331
+ await h.emit("session_start");
1332
+ assert.equal(await firstCallResult(h, "r1", { agent: "conformance-reviewer", task: "re-audit" }), undefined);
1333
+ assert.equal(await firstCallResult(h, "r2", { agent: "code-reviewer", task: "review" }), undefined);
1334
+
1335
+ const off = guardedHarness({ piGauntlet: { closureReview: { enforce: false, maxFixRounds: 0 } } });
1336
+ await off.emit("session_start");
1337
+ assert.equal(await firstCallResult(off, "c1", implementerWave(1)), undefined);
1338
+ assert.equal(await firstCallResult(off, "c2", loneImplementer()), undefined);
1339
+ });
1340
+
1341
+ test("counter: blocked, errored, empty, non-implementer, and out-of-window results do not count; non-zero exit does", async () => {
1342
+ const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 1 } } });
1343
+ await h.emit("session_start");
1344
+ await h.emitEvent("tool_result", waveResult("e1", [{ agent: "implementer", exitCode: 0 }], true));
1345
+ await h.emitEvent("tool_result", waveResult("e2", []));
1346
+ await h.emitEvent("tool_result", waveResult("e3", [{ agent: "code-reviewer", exitCode: 0 }]));
1347
+ const blockedLone = await firstCallResult(h, "b1", loneImplementer());
1348
+ assert.equal(blockedLone?.block, true);
1349
+ assert.equal(await firstCallResult(h, "c1", implementerWave(1)), undefined, "budget untouched");
1350
+ await h.emitEvent("tool_result", waveResult("c1", [{ agent: "implementer", exitCode: 1 }]));
1351
+ assert.equal((await firstCallResult(h, "c2", implementerWave(1)))?.block, true);
1352
+ });
1353
+
1354
+ test("counter: a ship-phase implementer wave with the latch set does not count", async () => {
1355
+ const h = harness({
1356
+ cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 1 } } }),
1357
+ branch: verifyBranch([
1358
+ subagentResult(["conformance-reviewer"]),
1359
+ phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete" })),
1360
+ phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" })),
1361
+ subagentResult(["implementer"]),
1362
+ phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress", ship: "in_progress" })),
1363
+ ]),
1364
+ });
1365
+ await h.emit("session_start");
1366
+ assert.equal(await firstCallResult(h, "c1", implementerWave(1)), undefined);
1367
+ });
1368
+
1369
+ test("counter reset: start implement and reset zero it; start verify --force keeps it", async () => {
1370
+ const mk = () => guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 1 }, flowGuards: { enforce: false } } }, [subagentResult(["implementer"])]);
1371
+
1372
+ const force = mk();
1373
+ await force.emit("session_start");
1374
+ const forceTool = force.tools.find((t) => t.name === "phase_tracker")!;
1375
+ const forced = (await forceTool.execute("p1", { action: "start", phase: "verify", force: true }, undefined, undefined, force.ctx)) as { details: { error?: string } };
1376
+ assert.equal(forced.details.error, undefined);
1377
+ assert.equal((await firstCallResult(force, "c1", implementerWave(1)))?.block, true, "budget survives verify --force");
1378
+
1379
+ const impl = mk();
1380
+ await impl.emit("session_start");
1381
+ const implTool = impl.tools.find((t) => t.name === "phase_tracker")!;
1382
+ for (const [id, input] of [
1383
+ ["p1", { action: "skip", phase: "verify", reason: "amendment" }],
1384
+ ["p2", { action: "start", phase: "implement", force: true }],
1385
+ ["p3", { action: "complete", phase: "implement" }],
1386
+ ["p4", { action: "start", phase: "verify", force: true }],
1387
+ ] as const) {
1388
+ const result = (await implTool.execute(id, input, undefined, undefined, impl.ctx)) as { details: { error?: string } };
1389
+ assert.equal(result.details.error, undefined);
1390
+ }
1391
+ assert.equal(await firstCallResult(impl, "c1", loneImplementer()), undefined, "latch cleared by start implement");
1392
+ await impl.emitEvent("tool_result", waveResult("audit", [{ agent: "conformance-reviewer", exitCode: 0 }]));
1393
+ assert.equal(await firstCallResult(impl, "c2", implementerWave(1)), undefined, "counter cleared by start implement");
1394
+
1395
+ const reset = mk();
1396
+ await reset.emit("session_start");
1397
+ const resetTool = reset.tools.find((t) => t.name === "phase_tracker")!;
1398
+ const resetResult = (await resetTool.execute("p1", { action: "reset" }, undefined, undefined, reset.ctx)) as { details: { error?: string } };
1399
+ assert.equal(resetResult.details.error, undefined);
1400
+ assert.equal(await firstCallResult(reset, "c1", loneImplementer()), undefined, "dormant after reset");
1401
+ });
1402
+
1403
+ test("replay: two implementer waves after the audit in verify restore fixRounds 2; the same waves in ship restore 0", async () => {
1404
+ const inVerify = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 2 } } }, [
1405
+ subagentResult(["implementer"]),
1406
+ subagentResult(["implementer", "code-reviewer"]),
1407
+ ]);
1408
+ await inVerify.emit("session_start");
1409
+ assert.equal((await firstCallResult(inVerify, "c1", implementerWave(1)))?.block, true);
1410
+
1411
+ const inShip = harness({
1412
+ cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 2 } } }),
1413
+ branch: verifyBranch([
1414
+ subagentResult(["conformance-reviewer"]),
1415
+ phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete" })),
1416
+ phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" })),
1417
+ subagentResult(["implementer"]),
1418
+ subagentResult(["implementer"]),
1419
+ phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress", ship: "in_progress" })),
1420
+ ]),
1421
+ });
1422
+ await inShip.emit("session_start");
1423
+ assert.equal(await firstCallResult(inShip, "c1", implementerWave(1)), undefined);
1424
+ });
@@ -86,6 +86,17 @@ const qualifiesAsClosureDispatch = (details: unknown): boolean => {
86
86
  return d.results.some((r) => r?.agent === "conformance-reviewer" && r?.exitCode === 0);
87
87
  };
88
88
 
89
+ // One conformance fix round = one non-error subagent result carrying at least one
90
+ // implementer child, regardless of dispatch mode or per-child exit code (a failed
91
+ // implementer still spent the round; its retry is the next one). Async and
92
+ // management dispatches return results: [] and never count.
93
+ const isImplementerWave = (details: unknown, isError: boolean | undefined): boolean => {
94
+ if (isError === true) return false;
95
+ const d = details as { results?: { agent?: unknown }[] } | undefined;
96
+ if (!d || !Array.isArray(d.results) || d.results.length === 0) return false;
97
+ return d.results.some((r) => r?.agent === "implementer");
98
+ };
99
+
89
100
  // Review-cadence guard (spec 2026-08-12-execution-fidelity-hardening): presence-only
90
101
  // advisory ledger of the most recent completed implementer / spec-reviewer /
91
102
  // code-reviewer dispatch. Agents completing in the SAME dispatch share a sequence
@@ -164,22 +175,14 @@ const branchBlockReason = (phase: Phase): string =>
164
175
  "create/enter one with /skill:using-git-worktrees and run this there. " +
165
176
  "To override, set piGauntlet.flowGuards.enforce: false.";
166
177
 
167
- // Closure-review model guard: when piGauntlet.closureReview.model is configured,
168
- // a conformance-reviewer dispatch should inject that model call-site. The persona
169
- // ships model-free, so a bare omission silently inherits the parent session's
170
- // builder model - defeating the point of an independent closing gate. Walk the
171
- // single / tasks / chain / parallel dispatch shapes and collect the model of every
172
- // conformance-reviewer entry (undefined = absent/empty/non-string). A bare omission
173
- // is BLOCKED; an explicit model that differs from the configured one is WARNED
174
- // (non-blocking), preserving the documented retry-with-fallback hatch while
175
- // surfacing drift.
176
- const conformanceModels = (input: unknown): (string | undefined)[] => {
177
- const models: (string | undefined)[] = [];
178
- const norm = (m: unknown) => (typeof m === "string" && m.trim() ? m.trim() : undefined);
178
+ // Walk the single / tasks / chain / parallel dispatch shapes and return every
179
+ // entry whose `agent` matches `name` (the raw node, so callers can read its model).
180
+ const collectAgents = (input: unknown, name: string): Record<string, unknown>[] => {
181
+ const hits: Record<string, unknown>[] = [];
179
182
  const collect = (node: unknown): void => {
180
183
  if (!node || typeof node !== "object") return;
181
184
  const o = node as Record<string, unknown>;
182
- if (o.agent === "conformance-reviewer") models.push(norm(o.model));
185
+ if (o.agent === name) hits.push(o);
183
186
  for (const key of ["tasks", "chain", "parallel"] as const) {
184
187
  const v = o[key];
185
188
  if (Array.isArray(v)) for (const item of v) collect(item);
@@ -187,9 +190,27 @@ const conformanceModels = (input: unknown): (string | undefined)[] => {
187
190
  }
188
191
  };
189
192
  collect(input);
190
- return models;
193
+ return hits;
191
194
  };
192
195
 
196
+ const conformanceModels = (input: unknown): (string | undefined)[] =>
197
+ collectAgents(input, "conformance-reviewer").map((o) =>
198
+ typeof o.model === "string" && o.model.trim() ? o.model.trim() : undefined,
199
+ );
200
+
201
+ // Conformance fix-loop dispatch guard (spec 2026-09-13-conformance-dispatch-guard):
202
+ // after the first audit, pi-cohort honours worktree: true only in tasks mode, so a
203
+ // lone implementer runs unisolated and yields no patch for the integrate step.
204
+ const loneImplementerBlockReason =
205
+ 'Conformance fix loop: dispatch implementers as a one-task tasks wave (tasks: [{ agent: "implementer", worktree: true, ... }]); ' +
206
+ "a lone agent call runs unisolated and produces no worktree diff. " +
207
+ "To disable this gate, set piGauntlet.closureReview.enforce: false.";
208
+
209
+ const fixRoundCapBlockReason = (used: number, cap: number): string =>
210
+ `Conformance fix loop: fix round ${used} of ${cap} already used; escalate to the human ` +
211
+ "with the verdict trail instead of re-looping. " +
212
+ "To disable this gate, set piGauntlet.closureReview.enforce: false.";
213
+
193
214
  const closureModelBlockReason = (model: string, missing: number, total: number): string =>
194
215
  `Blocked: ${missing} of ${total} conformance-reviewer ${total === 1 ? "dispatch" : "entries"} ` +
195
216
  `omitted a model while piGauntlet.closureReview.model is set to "${model}".\n` +
@@ -306,7 +327,23 @@ function formatStatus(phases: PhaseMap): string {
306
327
  export default function (pi: ExtensionAPI) {
307
328
  let phases: PhaseMap = emptyPhases();
308
329
  let conformanceDispatched = false;
330
+ let fixRounds = 0;
309
331
  let gauntletEntered = false;
332
+ // Shared by replay and the live tool_result hook so a resumed session enforces
333
+ // the same budget. Evaluated BEFORE the same result may set the latch, so the
334
+ // R0 audit result itself never counts as a wave.
335
+ const observeFixWave = (details: unknown, isError: boolean | undefined, ctx: ExtensionContext) => {
336
+ if (
337
+ !isSubagentChild &&
338
+ gauntletEntered &&
339
+ phases.verify.status === "in_progress" &&
340
+ conformanceDispatched &&
341
+ isImplementerWave(details, isError) &&
342
+ resolveClosureReview(loadGauntletSettings(ctx.cwd).gauntlet).enforce
343
+ ) {
344
+ fixRounds += 1;
345
+ }
346
+ };
310
347
  let planCheckStamp: PlanCheckStamp | undefined;
311
348
  const attemptedRecoveryEdges = new Set<RecoveryEdge>();
312
349
 
@@ -382,6 +419,7 @@ export default function (pi: ExtensionAPI) {
382
419
  const reconstructState = (ctx: ExtensionContext) => {
383
420
  phases = emptyPhases();
384
421
  conformanceDispatched = false;
422
+ fixRounds = 0;
385
423
  gauntletEntered = false;
386
424
  planCheckStamp = undefined;
387
425
  attemptedRecoveryEdges.clear();
@@ -409,9 +447,11 @@ export default function (pi: ExtensionAPI) {
409
447
  gauntletEntered = nextGauntletEntered(gauntletEntered, details.action, details.phases.brainstorm.status);
410
448
  if (details.action === "start" && details.phases.implement.status === "in_progress") {
411
449
  conformanceDispatched = false;
450
+ fixRounds = 0;
412
451
  }
413
452
  if (details.action === "reset") {
414
453
  conformanceDispatched = false;
454
+ fixRounds = 0;
415
455
  planCheckStamp = undefined;
416
456
  }
417
457
  }
@@ -423,6 +463,7 @@ export default function (pi: ExtensionAPI) {
423
463
  planCheckStamp = undefined;
424
464
  }
425
465
  } else if (msg.toolName === "subagent") {
466
+ observeFixWave(msg.details, msg.isError, ctx);
426
467
  if (qualifiesAsClosureDispatch(msg.details)) conformanceDispatched = true;
427
468
  observeCadence(msg.details);
428
469
  } else if (msg.toolName === "plan_tracker") {
@@ -492,10 +533,20 @@ export default function (pi: ExtensionAPI) {
492
533
  // subagent dispatch never loads settings or leaks a settingsErrorWarning onto its result.
493
534
  // Inline-matched (not called) so closureEnforced() stays a lazy second conjunct.
494
535
  if (event.toolName === "subagent" && gauntletEntered && closureEnforced()) {
495
- const model = closureReviewModel();
496
- // Only execution-mode dispatches carry a model; management/control modes
497
- // (action: list/get/create/update/delete/status/...) execute nothing, so skip them.
536
+ // Only execution-mode dispatches carry a model or run agents; management/control
537
+ // modes (action: list/get/create/update/delete/status/...) execute nothing, so skip them.
498
538
  const hasAction = !!(event.input as { action?: unknown })?.action;
539
+ // Fix-loop window: verify in progress and the R0 audit observed. Read directly,
540
+ // not via activeGuardPhase() - verify is not a GUARD_PHASES member.
541
+ const fixLoopWindow = !isSubagentChild && !hasAction && phases.verify.status === "in_progress" && conformanceDispatched;
542
+ if (fixLoopWindow && (event.input as { agent?: unknown })?.agent === "implementer") {
543
+ return { block: true, reason: loneImplementerBlockReason };
544
+ }
545
+ if (fixLoopWindow && collectAgents(event.input, "implementer").length > 0) {
546
+ const cap = resolveClosureReview(g()).maxFixRounds;
547
+ if (fixRounds >= cap) return { block: true, reason: fixRoundCapBlockReason(fixRounds, cap) };
548
+ }
549
+ const model = closureReviewModel();
499
550
  if (model && !hasAction) {
500
551
  const configured = model.trim();
501
552
  const models = conformanceModels(event.input);
@@ -635,8 +686,9 @@ export default function (pi: ExtensionAPI) {
635
686
  return undefined;
636
687
  });
637
688
 
638
- pi.on("tool_result", async (event) => {
689
+ pi.on("tool_result", async (event, ctx) => {
639
690
  if (event.toolName === "subagent") {
691
+ observeFixWave(event.details, event.isError, ctx);
640
692
  if (qualifiesAsClosureDispatch(event.details)) conformanceDispatched = true;
641
693
  observeCadence(event.details);
642
694
  const warning = pendingGuardWarnings.get(event.toolCallId);
@@ -859,7 +911,10 @@ export default function (pi: ExtensionAPI) {
859
911
  }
860
912
  phases = { ...phases, [params.phase]: transitionPhaseState("in_progress") as PhaseState };
861
913
  // A rewind must not inherit the prior verify's conformance latch.
862
- if (params.phase === "implement") conformanceDispatched = false;
914
+ if (params.phase === "implement") {
915
+ conformanceDispatched = false;
916
+ fixRounds = 0;
917
+ }
863
918
  gauntletEntered = nextGauntletEntered(gauntletEntered, "start", phases.brainstorm.status);
864
919
  firedGuards.clear();
865
920
  updateWidget(ctx);
@@ -1024,6 +1079,7 @@ export default function (pi: ExtensionAPI) {
1024
1079
  PHASES.map((p) => [p, transitionPhaseState("pending")]),
1025
1080
  ) as PhaseMap;
1026
1081
  conformanceDispatched = false;
1082
+ fixRounds = 0;
1027
1083
  gauntletEntered = nextGauntletEntered(gauntletEntered, "reset", phases.brainstorm.status);
1028
1084
  firedGuards.clear();
1029
1085
  updateWidget(ctx);
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-gauntlet",
3
- "version": "5.5.1",
3
+ "version": "5.5.3",
4
4
  "description": "Opinionated, gated workflow skills, subagent personas, and runtime extensions for the pi coding agent.",
5
5
  "author": "Jacek Juraszek",
6
6
  "type": "module",
@@ -105,7 +105,7 @@ When agents return:
105
105
 
106
106
  **If integrated changes apply cleanly but the suite fails (semantic conflict):** agents made incompatible assumptions across disjoint files (renamed symbol, changed shape). Diagnose the incompatible pair and re-run the offending task sequentially on the integrated HEAD.
107
107
 
108
- **If some agents failed:** Integrate successful agents first (commit their work). Then retry the failed agent with fresh context that includes the integrated changes.
108
+ **If some agents failed:** Integrate successful agents first (commit their work). Then retry the failed agent with fresh context that includes the integrated changes. Inside the conformance loop the retry is a one-task `tasks` wave, never a lone `agent: "implementer"`.
109
109
 
110
110
  ## Fix fan-out
111
111
 
@@ -142,29 +142,26 @@ prerequisites hold.
142
142
  Per round:
143
143
 
144
144
  1. **Synchronize gap tasks** — append only a genuinely new gap that is entering remediation, named `Gn: <gap origin clause verbatim, truncated>`; never `init`. Find existing gaps by their exact `Gn:` prefix and reuse that index even if origin wording changes. Carried-OPEN inventory-only gaps add nothing. Before dispatch, mark every remediated gap's existing index `in_progress`; a re-audit needing more work reopens that same `Gn` index. The lifecycle traces `[T1,T2]`, then `[T1,T2,G1]`, then `[T1,T2,G1,G2]`; no test-retry or review-round wrapper task.
145
- 2. **Fix dispatch** — per `dispatching-parallel-agents` "Fix fan-out": a `disjoint`
146
- group of ≥ 2 gaps (per the report's `Parallel-safe:` line) fixes in one parallel
147
- foreground dispatch — one `implementer` per gap (fresh context, `async: false`,
148
- `worktree: true`, `cwd` = the conformance worktree, task = the gap block verbatim
149
- with `touched-files` as the ownership boundary). For an `UNAUTHORIZED` `fix` gap
145
+ 2. **Fix wave** — select gaps greedily in `Gn` order: take each `fix` gap unless a gap it `conflicts` with (per the report's `Parallel-safe:` line) is already taken; the certificate's `disjoint` grouping is ignored, and a certificate still malformed after the one re-ask means every gap `conflicts` with every other. Held gaps carry to the next round; the wave is never empty while an eligible `fix` gap exists. Dispatch **one** call - `subagent({ context: "fresh", async: false, tasks: [...] })` - with one `implementer` task per selected gap (`worktree: true`, `cwd` = the conformance worktree, task = the gap block verbatim with `touched-files` as the ownership boundary). A single gap is a one-task `tasks` call; a lone `agent: "implementer"` call never appears in this loop. For an `UNAUTHORIZED` `fix` gap
150
146
  whose `evidence` opens with the over-spec provenance (`spec "<section>" - "<clause>" (over-spec)`), the orchestrator adds the spec path to that gap's `touched-files` before dispatch,
151
147
  so the implementer deletes the surface **and** the clause/AC line in the same
152
148
  fix commit; the re-audit then has no `Rn` for it and no `MISSING` echo. The dispatch adds `SCOPED_TEST_COMMANDS`
153
- to the gap block: the gap-relevant plan-declared commands, or `none` (the round's
154
- test gate owns execution). `conflicts` pairs serialize. Gaps outside any ≥ 2-ID
155
- `disjoint` group run sequentially as before. Then dispatch foreground `spec-reviewer`
156
- per gap on the gap-block reference contract below.
149
+ to the gap block: the gap-relevant plan-declared commands, or `none`; on an ad-hoc no-plan path, the project's canonical test command.
157
150
  3. **Integrate** serially via `git apply` onto the worktree HEAD, one gap's
158
- patch at a time. Failure handling is inherited verbatim from
151
+ patch at a time. Commit each per-gap fix with the message **`conformance fix Gn`** (durable,
152
+ `git log`-readable pre-squash) so the finish gate and any revert can identify
153
+ auto-applied fixes; a Convergence repair wave commits as one **`conformance fix CR`**.
154
+ Failure handling is inherited verbatim from
159
155
  `dispatching-parallel-agents` "Review and Integrate": textual conflict →
160
156
  re-run one agent sequentially with the other's integrated changes as
161
157
  context; semantic conflict (applies clean, suite fails) → re-run the
162
158
  offending task sequentially on integrated HEAD; a failed agent → integrate
163
159
  the successes, then retry the failure with fresh context including the
164
- integrated changes. A `BLOCKED`/`NEEDS_CONTEXT` return surfaces to the user.
165
- 4. **Test gate** on the integrated tree. In a plan flow, run the full plan-header `Verification` set once here; on an ad-hoc no-plan path, use the project's canonical test command. A failure re-enters the failure-handling rules above.
166
- 5. **Round CR and completion** — run `code-reviewer` once on the round's cumulative fix delta (not per gap), foreground with `async: false` and `SCOPED_TEST_COMMANDS` = the round's gap-relevant commands, or `none` (the round's test gate owns execution). After integration, tests, and this CR accept the work, explicitly mark every remediated gap's same `Gn` index `complete`, before re-audit.
167
- 6. **Re-audit**: foreground re-dispatch `conformance-reviewer` with `async: false` over the fixes **plus** the
160
+ integrated changes. Every re-run or retry inside this loop is itself a
161
+ one-task `tasks` wave (never a lone `agent: "implementer"`) and counts as a
162
+ wave against `maxFixRounds`. A `BLOCKED`/`NEEDS_CONTEXT` return surfaces to the user.
163
+ 4. **Scoped tests** on the integrated tree: the round's `SCOPED_TEST_COMMANDS` union. A failure re-enters the failure-handling rules above.
164
+ 5. **Re-audit**: foreground re-dispatch `conformance-reviewer` with `async: false` over the fixes **plus** the
168
165
  regression guard (any prior-`DELIVERED` requirement whose `evidence` file
169
166
  the fix diff touched). Pass the full prior conformance report (every row,
170
167
  including DELIVERED rows and their `evidence` `file:line`) and the round's
@@ -173,18 +170,26 @@ Per round:
173
170
  `model:` when it is `undefined` to inherit the parent's model. Inside a
174
171
  brainstorming-entered flow, the phase-tracker closure guard blocks a dispatch
175
172
  that omits `model:` when `closureReview.model` is set, and warns (non-blocking)
176
- on one whose model differs.
177
- 7. **Converge or continue**: verdict `CONFORMS` → record it, done. Open gaps
173
+ on one whose model differs. Mark every `Gn` the re-audit reports `DELIVERED`
174
+ `complete`; open ones stay `in_progress`.
175
+ 6. **Converge or continue**: verdict `CONFORMS` → Convergence below. Open gaps
178
176
  within the cap → re-partition (per the rule above) and start the next
179
177
  round. Cap (`gauntlet_setting({ key: "closureReview" }).maxFixRounds`,
180
- default `2`, floors negatives at `0`, coerces non-integers to `2`) reached
181
- with an open `fix` gap → **escalate to the human** with the per-gap
178
+ default `3`, floors negatives at `0`, coerces non-integers to `3`) reached
179
+ with an open `fix` gap or repair item → **escalate to the human** with the per-gap
182
180
  round-by-round verdict trail. Escalation is the sole non-completing
183
- terminal state — no silent re-loop, no auto-ship.
181
+ terminal state — no silent re-loop, no auto-ship. Inside a
182
+ brainstorming-entered flow the phase tracker enforces both rules at
183
+ tool-call time: a lone `agent: "implementer"` dispatch is blocked, and the
184
+ wave after the cap is blocked (`closureReview.enforce: false` disables).
184
185
 
185
- Commit each per-gap fix with the message **`conformance fix Gn`** (durable,
186
- `git log`-readable pre-squash) so the finish gate and any revert can identify
187
- auto-applied fixes.
186
+ **Convergence** — runs after R0 `CONFORMS` and after every `CONFORMS` re-audit. `r0-head` is HEAD when the loop was entered (the R0 dispatch, or the finish-time `fix-now` entry) - the parent of the oldest `conformance fix` commit; `audited-base` stays the last audit's HEAD SHA.
187
+
188
+ a. Run the full plan-header `Verification` set once (ad-hoc: the project's canonical test command). After R0 `CONFORMS` with no round run, the pre-R0 full run counts.
189
+ b. Dispatch `code-reviewer` directly (foreground, `async: false`, `SCOPED_TEST_COMMANDS: none`) over `git diff <r0-head>..HEAD`; never via `/skill:requesting-code-review`. An empty diff is nothing to review - no dispatch.
190
+ c. Repair items = every failing command from a + every Critical/Moderate finding from b (`Behaviour-change: yes` included; the re-audit is its origin check). None → write the closure block; done. Any at the cap → escalate per step 6 with the test/CR trail. Any under the cap → re-enter step 2 as a one-task `tasks` wave: one `implementer` whose task is every repair item verbatim (ownership boundary = the files in `git diff <r0-head>..HEAD`, the CR findings' `touched-files`, and the files each failing command's output names, `SCOPED_TEST_COMMANDS: none`, no `Gn` tracker task, no gap selection), integrate as one `conformance fix CR`, re-audit, then Convergence again. That wave counts against `maxFixRounds`. Later Convergence CRs keep the same `<r0-head>..HEAD` range.
191
+
192
+ `conformance fix CR` is not a gap fix: it is absent from the `auto-applied fix commits` index and has no `revert conformance fix Gn` action at the finish gate.
188
193
 
189
194
  **`maxFixRounds: 0`**: skip this loop entirely; every `recommended: fix` gap is
190
195
  carried OPEN to the finish gate per the precondition-unavailable
@@ -194,22 +199,6 @@ auto-fix, so treat `fix` gaps like any other deferred gap — unlike a cap > 0
194
199
  that is *exhausted*, which escalates mid-verify because the loop tried and
195
200
  could not converge.
196
201
 
197
- ### `spec-reviewer` gap-block reference contract
198
-
199
- Per-gap `spec-reviewer` in step 2 above is a **pre-integration mechanical
200
- check**, distinct from the round-level re-audit in step 6 (which still
201
- references the *origin* — spec + original prompt — unchanged). Frame the
202
- per-gap dispatch against the **gap block**, not a plan task:
203
-
204
- - **Requirement** = the gap's `origin` + `remediation` (what must be true
205
- after the fix).
206
- - **Closure proof** = the patch satisfies that requirement within the gap's
207
- `touched-files` — nothing missing, nothing extra.
208
- - **Output** = `spec-reviewer`'s normal MATCH/DRIFT verdict, referenced to the
209
- gap block instead of a plan task.
210
-
211
- This is a task-framing contract in the dispatch, not a new persona.
212
-
213
202
  ## Concern decomposition
214
203
 
215
204
  The main verification orchestrator — **not** `conformance-reviewer` — decomposes
@@ -459,15 +448,15 @@ concerns into one gap-scoped fix contract:
459
448
  boundary).
460
449
 
461
450
  Rescoped, accepted, and followed-up sibling concerns are **excluded** from the
462
- projection. The `implementer` and the pre-integration `spec-reviewer` receive
463
- this projected contract in place of the original whole-gap block.
464
-
465
- The projected task runs the existing full loop above: it retains the gap-level
466
- `conformance fix Gn` commit name, reruns the project's tests, runs
467
- `code-reviewer`, re-audits against the amended spec, and reenters the gate only
468
- if concerns remain. Gap-level revert stays available through the flat commit
469
- index. The gate records the final result by concern ID and title before showing
470
- branch integration options.
451
+ projection. The `implementer` receives this projected contract in place of the
452
+ original whole-gap block.
453
+
454
+ The projected task runs the round and Convergence above with `r0-head` = HEAD at
455
+ this entry: it retains the gap-level `conformance fix Gn` commit name, re-audits
456
+ against the amended spec, and reenters the gate only if concerns remain.
457
+ Gap-level revert stays available through the flat commit index. The gate records
458
+ the final result by concern ID and title before showing branch integration
459
+ options.
471
460
 
472
461
  ## Checklist
473
462
 
@@ -22,7 +22,7 @@ does not recurse into their leaves.
22
22
  whole-object, a repo file that sets only *one* leaf of a key silently drops the
23
23
  preset's other leaves for that key. A repo `closureReview: { "model": "..." }` with
24
24
  no `enforce`/`maxFixRounds` makes those fall back to their code defaults
25
- (`enforce` true, `maxFixRounds` 2), **not** to the preset's values. Define every
25
+ (`enforce` true, `maxFixRounds` 3), **not** to the preset's values. Define every
26
26
  leaf you care about together in the file that owns the key. (Sibling keys are
27
27
  unaffected - only the key the repo redefines is replaced.)
28
28