pi-gauntlet 5.5.2 → 5.5.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/agents/code-reviewer.md +7 -3
- package/extensions/lib/gauntlet-settings.test.ts +4 -3
- package/extensions/lib/gauntlet-settings.ts +1 -1
- package/extensions/phase-tracker.test.ts +556 -3
- package/extensions/phase-tracker.ts +146 -22
- package/extensions/test-support/pi-stubs.mjs +9 -9
- package/package.json +1 -1
- package/skills/dispatching-parallel-agents/SKILL.md +1 -1
- package/skills/verification-before-completion/reference/conformance-check.md +16 -6
- package/skills/verification-before-completion/reference/settings-precedence.md +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,18 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.5.4 - 2026-09-14
|
|
4
|
+
|
|
5
|
+
- `phase-tracker`: new `phase_tracker` action `grant_fix_rounds` records an explicit human approval of N more conformance fix rounds (reason quotes the human), accepted only while the cap block is live; each qualifying implementer wave spends one granted round, credits replay with the session and reset with the audit latch. The cap-block message now leads with that action and names the exact `.pi/settings.json` for the `enforce: false` last resort (applies without restart, whole-block precedence).
|
|
6
|
+
- `gauntlet_setting` is registered with sequential execution, so a settings write and a verifying read batched in one message run in order.
|
|
7
|
+
- `conformance-check.md`: an explicit human approval re-enters the fix loop via `grant_fix_rounds`; without it escalation stays terminal.
|
|
8
|
+
|
|
9
|
+
## v5.5.3 - 2026-09-13
|
|
10
|
+
|
|
11
|
+
- `code-reviewer`: verdict is a stated function of severity - a Critical or Moderate finding means `FIX_FIRST`, Minor-only and clean reports mean `SHIP`, matching the orchestrating skills. (#30)
|
|
12
|
+
- `phase-tracker`: two new closure-review blocks inside a brainstorming-entered flow once `verify` is in progress and a conformance audit has been observed - a lone top-level `agent: "implementer"` `subagent` dispatch is blocked (one-task `tasks` wave is the only isolated shape), and an implementer dispatch past `closureReview.maxFixRounds` is blocked with the escalation text. Rounds are counted from non-error results carrying an implementer child; the counter resets with the audit latch. `closureReview.enforce: false` disables both.
|
|
13
|
+
- `closureReview.maxFixRounds` default 2 -> 3.
|
|
14
|
+
- `conformance-check.md` / `dispatching-parallel-agents`: retries inside the conformance loop are one-task `tasks` waves and count against the cap.
|
|
15
|
+
|
|
3
16
|
## v5.5.2 - 2026-09-13
|
|
4
17
|
|
|
5
18
|
- `release.sh <level>` promotes the CHANGELOG `## Unreleased` section to `## vX.Y.Z - <date>` and commits it with `package.json` in the single `Release X.Y.Z` commit, so `patch`/`minor`/`major` now work here (the `current`-only path is gone). New CONFIG field `CHANGELOG_HEADING`.
|
package/agents/code-reviewer.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: code-reviewer
|
|
3
|
-
description: Production-readiness code review with prioritized findings (Critical
|
|
3
|
+
description: Production-readiness code review with prioritized findings (Critical and Moderate block merge, Minor is a nit). Read-only — does not edit.
|
|
4
4
|
tools: read, grep, find, ls, bash
|
|
5
5
|
defaultContext: fresh
|
|
6
6
|
inheritProjectContext: true
|
|
@@ -44,11 +44,15 @@ Parallel-safe: F1,F3 disjoint; F2 conflicts F1 (both touch auth.ts)
|
|
|
44
44
|
Behaviour-change: yes | no
|
|
45
45
|
```
|
|
46
46
|
|
|
47
|
+
Severity decides SHIP versus FIX_FIRST. A Critical or Moderate finding means
|
|
48
|
+
FIX_FIRST. Minor-only findings and clean reports mean SHIP. REJECT overrides
|
|
49
|
+
both: refuse a change that must not land at all.
|
|
50
|
+
|
|
47
51
|
Severity:
|
|
48
52
|
|
|
49
53
|
- **Critical** — must fix before merge (data loss, security, broken correctness on a common path, broken contract).
|
|
50
|
-
- **Moderate** —
|
|
51
|
-
- **Minor** — nit, style, preference, suggestion.
|
|
54
|
+
- **Moderate** — must fix before merge (significant defect or drift that does not rise to Critical).
|
|
55
|
+
- **Minor** — nit, style, preference, suggestion; the only severity declinable without a fix round or re-review.
|
|
52
56
|
|
|
53
57
|
Label every finding with a globally unique `F1..Fn` ID (no restart per severity),
|
|
54
58
|
and a `touched-files:`/`touched-resources:` pair (files/resources a fix would
|
|
@@ -122,11 +122,12 @@ test("closureReview: enforce default true; false only when explicitly false", ()
|
|
|
122
122
|
assert.equal(resolveClosureReview({ closureReview: { enforce: true } }).enforce, true);
|
|
123
123
|
});
|
|
124
124
|
|
|
125
|
-
test("closureReview: maxFixRounds default
|
|
126
|
-
assert.equal(resolveClosureReview({}).maxFixRounds,
|
|
125
|
+
test("closureReview: maxFixRounds default 3, <0 -> 0, non-int -> 3", () => {
|
|
126
|
+
assert.equal(resolveClosureReview({}).maxFixRounds, 3);
|
|
127
127
|
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 5 } }).maxFixRounds, 5);
|
|
128
|
+
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 0 } }).maxFixRounds, 0);
|
|
128
129
|
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: -3 } }).maxFixRounds, 0);
|
|
129
|
-
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 1.5 } }).maxFixRounds,
|
|
130
|
+
assert.equal(resolveClosureReview({ closureReview: { maxFixRounds: 1.5 } }).maxFixRounds, 3);
|
|
130
131
|
});
|
|
131
132
|
|
|
132
133
|
test("flowGuards: defaults + overrides", () => {
|
|
@@ -84,7 +84,7 @@ export function resolveClosureReview(g: PiGauntlet): ClosureReviewResolved {
|
|
|
84
84
|
const model = nonEmptyString(cr?.model) ? cr!.model.trim() : undefined;
|
|
85
85
|
const enforce = cr?.enforce !== false;
|
|
86
86
|
const raw = cr?.maxFixRounds;
|
|
87
|
-
const maxFixRounds = typeof raw === "number" && Number.isInteger(raw) ? (raw < 0 ? 0 : raw) :
|
|
87
|
+
const maxFixRounds = typeof raw === "number" && Number.isInteger(raw) ? (raw < 0 ? 0 : raw) : 3;
|
|
88
88
|
return { model, enforce, maxFixRounds };
|
|
89
89
|
}
|
|
90
90
|
|
|
@@ -6,9 +6,14 @@ import { tmpdir } from "node:os";
|
|
|
6
6
|
import { isAbsolute, join } from "node:path";
|
|
7
7
|
import registerPhaseTracker from "./phase-tracker.ts";
|
|
8
8
|
|
|
9
|
+
const originalSubagentDepth = process.env.PI_SUBAGENT_DEPTH;
|
|
10
|
+
process.env.PI_SUBAGENT_DEPTH = "0";
|
|
11
|
+
|
|
9
12
|
const tempDirs: string[] = [];
|
|
10
13
|
after(() => {
|
|
11
14
|
for (const dir of tempDirs) rmSync(dir, { recursive: true, force: true });
|
|
15
|
+
if (originalSubagentDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
|
|
16
|
+
else process.env.PI_SUBAGENT_DEPTH = originalSubagentDepth;
|
|
12
17
|
});
|
|
13
18
|
|
|
14
19
|
const tempCwd = (settings?: unknown) => {
|
|
@@ -53,7 +58,7 @@ const resumedBranch = (rest: Partial<Record<Phase, Status>>) => [
|
|
|
53
58
|
|
|
54
59
|
function harness(options: { cwd?: string; branch?: unknown[]; idle?: boolean; beforeSettled?: (setIdle: (idle: boolean) => void) => void; sendThrows?: boolean; model?: { provider: string; id: string }; thinkingLevel?: string } = {}) {
|
|
55
60
|
const handlers = new Map<string, ((event: unknown, ctx: unknown) => unknown)[]>();
|
|
56
|
-
const tools: { name: string; execute: (...args: any[]) => unknown }[] = [];
|
|
61
|
+
const tools: { name: string; executionMode?: string; parameters?: any; execute: (...args: any[]) => unknown }[] = [];
|
|
57
62
|
const sent: { message: any; options: any }[] = [];
|
|
58
63
|
let idle = options.idle ?? true;
|
|
59
64
|
let branch = options.branch ?? [];
|
|
@@ -71,7 +76,7 @@ function harness(options: { cwd?: string; branch?: unknown[]; idle?: boolean; be
|
|
|
71
76
|
registered.push(handler);
|
|
72
77
|
handlers.set(event, registered);
|
|
73
78
|
},
|
|
74
|
-
registerTool(tool: { name: string; executionMode?: string; execute: (...args: any[]) => unknown }) {
|
|
79
|
+
registerTool(tool: { name: string; executionMode?: string; parameters?: any; execute: (...args: any[]) => unknown }) {
|
|
75
80
|
tools.push(tool);
|
|
76
81
|
},
|
|
77
82
|
sendMessage(message: unknown, sendOptions: unknown) {
|
|
@@ -529,7 +534,8 @@ test("implement-phase commit with implementer newer than both reviewers warns",
|
|
|
529
534
|
const warned = (await h.emitEvent("tool_result", commitResult("c1")))[0] as { content: { text: string }[] };
|
|
530
535
|
assert.match(warned.content[0].text, /no spec-reviewer or code-reviewer observed/);
|
|
531
536
|
} finally {
|
|
532
|
-
if (priorDepth
|
|
537
|
+
if (priorDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
|
|
538
|
+
else process.env.PI_SUBAGENT_DEPTH = priorDepth;
|
|
533
539
|
}
|
|
534
540
|
});
|
|
535
541
|
|
|
@@ -1164,3 +1170,550 @@ test("gauntlet_setting escalationLoop: setting absent -> ctx-derived main-loop m
|
|
|
1164
1170
|
const res = (await set.tools.find((t) => t.name === "gauntlet_setting")!.execute("g2", { key: "escalationLoop" }, undefined, undefined, set.ctx)) as { details: { implModel?: string } };
|
|
1165
1171
|
assert.equal(res.details.implModel, "p/strong:high");
|
|
1166
1172
|
});
|
|
1173
|
+
|
|
1174
|
+
// --- Conformance fix-loop dispatch guard (spec 2026-09-13-conformance-dispatch-guard) ---
|
|
1175
|
+
|
|
1176
|
+
const verifyBranch = (extra: unknown[] = []) => [
|
|
1177
|
+
...implementBranch(),
|
|
1178
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete" })),
|
|
1179
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress" })),
|
|
1180
|
+
...extra,
|
|
1181
|
+
];
|
|
1182
|
+
|
|
1183
|
+
const subagentCall = (id: string, input: unknown) => ({ toolName: "subagent", toolCallId: id, input });
|
|
1184
|
+
const loneImplementer = (extra: Record<string, unknown> = {}) => ({ agent: "implementer", task: "fix G1", ...extra });
|
|
1185
|
+
const implementerWave = (n = 1) => ({
|
|
1186
|
+
tasks: Array.from({ length: n }, (_, i) => ({ agent: "implementer", task: `fix G${i + 1}`, worktree: true })),
|
|
1187
|
+
});
|
|
1188
|
+
|
|
1189
|
+
const firstCallResult = async (h: ReturnType<typeof harness>, id: string, input: unknown) =>
|
|
1190
|
+
(await h.emitEvent("tool_call", subagentCall(id, input)))[0] as { block?: boolean; reason?: string } | undefined;
|
|
1191
|
+
|
|
1192
|
+
test("shape guard: lone implementer blocked in verify after the audit, with and without closureReview.model, async or not", async () => {
|
|
1193
|
+
for (const settings of [
|
|
1194
|
+
{ piGauntlet: { closureReview: { enforce: true } } },
|
|
1195
|
+
{ piGauntlet: { closureReview: { enforce: true, model: "x/y" } } },
|
|
1196
|
+
]) {
|
|
1197
|
+
const h = harness({ cwd: tempCwd(settings), branch: verifyBranch([subagentResult(["conformance-reviewer"])]) });
|
|
1198
|
+
await h.emit("session_start");
|
|
1199
|
+
const sync = await firstCallResult(h, "s1", loneImplementer());
|
|
1200
|
+
assert.equal(sync?.block, true);
|
|
1201
|
+
assert.match(sync?.reason ?? "", /dispatch implementers as a one-task tasks wave/);
|
|
1202
|
+
assert.match(sync?.reason ?? "", /piGauntlet\.closureReview\.enforce: false/);
|
|
1203
|
+
const async = await firstCallResult(h, "s2", loneImplementer({ async: true }));
|
|
1204
|
+
assert.equal(async?.block, true);
|
|
1205
|
+
}
|
|
1206
|
+
});
|
|
1207
|
+
|
|
1208
|
+
test("shape guard: lone implementer passes before the audit, in implement, in ship, and on management calls", async () => {
|
|
1209
|
+
const preLatch = harness({ branch: verifyBranch() });
|
|
1210
|
+
await preLatch.emit("session_start");
|
|
1211
|
+
assert.equal(await firstCallResult(preLatch, "s1", loneImplementer()), undefined);
|
|
1212
|
+
|
|
1213
|
+
const implement = harness({ branch: implementBranch([subagentResult(["conformance-reviewer"])]) });
|
|
1214
|
+
await implement.emit("session_start");
|
|
1215
|
+
assert.equal(await firstCallResult(implement, "s1", loneImplementer()), undefined);
|
|
1216
|
+
|
|
1217
|
+
const ship = harness({
|
|
1218
|
+
branch: verifyBranch([
|
|
1219
|
+
subagentResult(["conformance-reviewer"]),
|
|
1220
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete" })),
|
|
1221
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" })),
|
|
1222
|
+
]),
|
|
1223
|
+
});
|
|
1224
|
+
await ship.emit("session_start");
|
|
1225
|
+
assert.equal(await firstCallResult(ship, "s1", loneImplementer()), undefined);
|
|
1226
|
+
|
|
1227
|
+
const mgmt = harness({ branch: verifyBranch([subagentResult(["conformance-reviewer"])]) });
|
|
1228
|
+
await mgmt.emit("session_start");
|
|
1229
|
+
assert.equal(await firstCallResult(mgmt, "s1", { action: "status", agent: "implementer" }), undefined);
|
|
1230
|
+
});
|
|
1231
|
+
|
|
1232
|
+
test("shape guard: dormant when the flow was never entered (cold start verify) and when closureReview.enforce is false", async () => {
|
|
1233
|
+
const cold = harness({
|
|
1234
|
+
cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 0 } } }),
|
|
1235
|
+
branch: [phaseResult("start", phases({ verify: "in_progress" })), subagentResult(["conformance-reviewer"])],
|
|
1236
|
+
});
|
|
1237
|
+
await cold.emit("session_start");
|
|
1238
|
+
assert.equal(await firstCallResult(cold, "s1", loneImplementer()), undefined);
|
|
1239
|
+
assert.equal(await firstCallResult(cold, "s2", implementerWave(1)), undefined);
|
|
1240
|
+
assert.equal(await firstCallResult(cold, "s3", { chain: [{ agent: "implementer", task: "fix" }] }), undefined);
|
|
1241
|
+
|
|
1242
|
+
const off = harness({
|
|
1243
|
+
cwd: tempCwd({ piGauntlet: { closureReview: { enforce: false } } }),
|
|
1244
|
+
branch: verifyBranch([subagentResult(["conformance-reviewer"])]),
|
|
1245
|
+
});
|
|
1246
|
+
await off.emit("session_start");
|
|
1247
|
+
assert.equal(await firstCallResult(off, "s1", loneImplementer()), undefined);
|
|
1248
|
+
});
|
|
1249
|
+
|
|
1250
|
+
test("shape guard: a one-task tasks wave and a chain step are not the lone shape", async () => {
|
|
1251
|
+
const h = harness({ branch: verifyBranch([subagentResult(["conformance-reviewer"])]) });
|
|
1252
|
+
await h.emit("session_start");
|
|
1253
|
+
assert.equal(await firstCallResult(h, "s1", implementerWave(1)), undefined);
|
|
1254
|
+
assert.equal(await firstCallResult(h, "s2", { chain: [{ agent: "implementer", task: "fix" }] }), undefined);
|
|
1255
|
+
});
|
|
1256
|
+
|
|
1257
|
+
const waveResult = (id: string, results: { agent: string; exitCode: number }[], isError = false) => ({
|
|
1258
|
+
toolName: "subagent",
|
|
1259
|
+
toolCallId: id,
|
|
1260
|
+
isError,
|
|
1261
|
+
content: [],
|
|
1262
|
+
details: { results },
|
|
1263
|
+
});
|
|
1264
|
+
const okWave = (id: string) => waveResult(id, [{ agent: "implementer", exitCode: 0 }]);
|
|
1265
|
+
|
|
1266
|
+
const guardedHarness = (settings: unknown, extra: unknown[] = []) =>
|
|
1267
|
+
harness({ cwd: tempCwd(settings), branch: verifyBranch([subagentResult(["conformance-reviewer"]), ...extra]) });
|
|
1268
|
+
|
|
1269
|
+
test("cap guard: default 3 waves pass, the fourth is blocked with the escalation text", async () => {
|
|
1270
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { enforce: true } } });
|
|
1271
|
+
await h.emit("session_start");
|
|
1272
|
+
for (let i = 1; i <= 3; i++) {
|
|
1273
|
+
assert.equal(await firstCallResult(h, `c${i}`, implementerWave(2)), undefined, `round ${i} passes`);
|
|
1274
|
+
await h.emitEvent("tool_result", okWave(`c${i}`));
|
|
1275
|
+
}
|
|
1276
|
+
const blocked = await firstCallResult(h, "c4", implementerWave(1));
|
|
1277
|
+
assert.equal(blocked?.block, true);
|
|
1278
|
+
assert.match(blocked?.reason ?? "", /3 fix round\(s\) used against a cap of 3 \(granted rounds included\); escalate to the human/);
|
|
1279
|
+
});
|
|
1280
|
+
|
|
1281
|
+
test("cap guard: maxFixRounds 0 blocks the first wave; 1 blocks the second; a chain implementer is cap-checked and counted", async () => {
|
|
1282
|
+
const zero = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1283
|
+
await zero.emit("session_start");
|
|
1284
|
+
assert.equal((await firstCallResult(zero, "c1", implementerWave(1)))?.block, true);
|
|
1285
|
+
|
|
1286
|
+
const one = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 1 } } });
|
|
1287
|
+
await one.emit("session_start");
|
|
1288
|
+
assert.equal(await firstCallResult(one, "c1", { chain: [{ agent: "implementer", task: "fix" }] }), undefined);
|
|
1289
|
+
await one.emitEvent("tool_result", okWave("c1"));
|
|
1290
|
+
assert.equal((await firstCallResult(one, "c2", implementerWave(1)))?.block, true);
|
|
1291
|
+
});
|
|
1292
|
+
|
|
1293
|
+
test("counter: results while closure review enforcement is off do not consume budget", async () => {
|
|
1294
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { enforce: false, maxFixRounds: 1 } } });
|
|
1295
|
+
await h.emit("session_start");
|
|
1296
|
+
await h.emitEvent("tool_result", okWave("c1"));
|
|
1297
|
+
await h.emitEvent("tool_result", okWave("c2"));
|
|
1298
|
+
writeFileSync(
|
|
1299
|
+
join(h.ctx.cwd, ".pi", "settings.json"),
|
|
1300
|
+
JSON.stringify({ piGauntlet: { closureReview: { enforce: true, maxFixRounds: 1 } } }),
|
|
1301
|
+
);
|
|
1302
|
+
assert.equal(await firstCallResult(h, "c3", implementerWave(1)), undefined);
|
|
1303
|
+
});
|
|
1304
|
+
|
|
1305
|
+
test("fix-loop guard and counter are dormant in subagent children", async () => {
|
|
1306
|
+
const priorDepth = process.env.PI_SUBAGENT_DEPTH;
|
|
1307
|
+
process.env.PI_SUBAGENT_DEPTH = "1";
|
|
1308
|
+
let live: ReturnType<typeof harness>;
|
|
1309
|
+
let replay: ReturnType<typeof harness>;
|
|
1310
|
+
try {
|
|
1311
|
+
live = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1312
|
+
replay = guardedHarness(
|
|
1313
|
+
{ piGauntlet: { closureReview: { maxFixRounds: 2 } } },
|
|
1314
|
+
[subagentResult(["implementer"]), subagentResult(["implementer"])],
|
|
1315
|
+
);
|
|
1316
|
+
} finally {
|
|
1317
|
+
if (priorDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
|
|
1318
|
+
else process.env.PI_SUBAGENT_DEPTH = priorDepth;
|
|
1319
|
+
}
|
|
1320
|
+
|
|
1321
|
+
await live.emit("session_start");
|
|
1322
|
+
assert.equal(await firstCallResult(live, "c1", loneImplementer()), undefined);
|
|
1323
|
+
assert.equal(await firstCallResult(live, "c2", implementerWave(1)), undefined);
|
|
1324
|
+
await live.emitEvent("tool_result", okWave("c2"));
|
|
1325
|
+
await replay.emit("session_start");
|
|
1326
|
+
assert.equal(await firstCallResult(replay, "c3", implementerWave(1)), undefined);
|
|
1327
|
+
});
|
|
1328
|
+
|
|
1329
|
+
test("cap guard: non-implementer dispatches never blocked; enforce false passes everything", async () => {
|
|
1330
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1331
|
+
await h.emit("session_start");
|
|
1332
|
+
assert.equal(await firstCallResult(h, "r1", { agent: "conformance-reviewer", task: "re-audit" }), undefined);
|
|
1333
|
+
assert.equal(await firstCallResult(h, "r2", { agent: "code-reviewer", task: "review" }), undefined);
|
|
1334
|
+
|
|
1335
|
+
const off = guardedHarness({ piGauntlet: { closureReview: { enforce: false, maxFixRounds: 0 } } });
|
|
1336
|
+
await off.emit("session_start");
|
|
1337
|
+
assert.equal(await firstCallResult(off, "c1", implementerWave(1)), undefined);
|
|
1338
|
+
assert.equal(await firstCallResult(off, "c2", loneImplementer()), undefined);
|
|
1339
|
+
});
|
|
1340
|
+
|
|
1341
|
+
test("counter: blocked, errored, empty, non-implementer, and out-of-window results do not count; non-zero exit does", async () => {
|
|
1342
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 1 } } });
|
|
1343
|
+
await h.emit("session_start");
|
|
1344
|
+
await h.emitEvent("tool_result", waveResult("e1", [{ agent: "implementer", exitCode: 0 }], true));
|
|
1345
|
+
await h.emitEvent("tool_result", waveResult("e2", []));
|
|
1346
|
+
await h.emitEvent("tool_result", waveResult("e3", [{ agent: "code-reviewer", exitCode: 0 }]));
|
|
1347
|
+
const blockedLone = await firstCallResult(h, "b1", loneImplementer());
|
|
1348
|
+
assert.equal(blockedLone?.block, true);
|
|
1349
|
+
assert.equal(await firstCallResult(h, "c1", implementerWave(1)), undefined, "budget untouched");
|
|
1350
|
+
await h.emitEvent("tool_result", waveResult("c1", [{ agent: "implementer", exitCode: 1 }]));
|
|
1351
|
+
assert.equal((await firstCallResult(h, "c2", implementerWave(1)))?.block, true);
|
|
1352
|
+
});
|
|
1353
|
+
|
|
1354
|
+
test("counter: a ship-phase implementer wave with the latch set does not count", async () => {
|
|
1355
|
+
const h = harness({
|
|
1356
|
+
cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 1 } } }),
|
|
1357
|
+
branch: verifyBranch([
|
|
1358
|
+
subagentResult(["conformance-reviewer"]),
|
|
1359
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete" })),
|
|
1360
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" })),
|
|
1361
|
+
subagentResult(["implementer"]),
|
|
1362
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress", ship: "in_progress" })),
|
|
1363
|
+
]),
|
|
1364
|
+
});
|
|
1365
|
+
await h.emit("session_start");
|
|
1366
|
+
assert.equal(await firstCallResult(h, "c1", implementerWave(1)), undefined);
|
|
1367
|
+
});
|
|
1368
|
+
|
|
1369
|
+
test("counter reset: start implement and reset zero it; start verify --force keeps it", async () => {
|
|
1370
|
+
const mk = () => guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 1 }, flowGuards: { enforce: false } } }, [subagentResult(["implementer"])]);
|
|
1371
|
+
|
|
1372
|
+
const force = mk();
|
|
1373
|
+
await force.emit("session_start");
|
|
1374
|
+
const forceTool = force.tools.find((t) => t.name === "phase_tracker")!;
|
|
1375
|
+
const forced = (await forceTool.execute("p1", { action: "start", phase: "verify", force: true }, undefined, undefined, force.ctx)) as { details: { error?: string } };
|
|
1376
|
+
assert.equal(forced.details.error, undefined);
|
|
1377
|
+
assert.equal((await firstCallResult(force, "c1", implementerWave(1)))?.block, true, "budget survives verify --force");
|
|
1378
|
+
|
|
1379
|
+
const impl = mk();
|
|
1380
|
+
await impl.emit("session_start");
|
|
1381
|
+
const implTool = impl.tools.find((t) => t.name === "phase_tracker")!;
|
|
1382
|
+
for (const [id, input] of [
|
|
1383
|
+
["p1", { action: "skip", phase: "verify", reason: "amendment" }],
|
|
1384
|
+
["p2", { action: "start", phase: "implement", force: true }],
|
|
1385
|
+
["p3", { action: "complete", phase: "implement" }],
|
|
1386
|
+
["p4", { action: "start", phase: "verify", force: true }],
|
|
1387
|
+
] as const) {
|
|
1388
|
+
const result = (await implTool.execute(id, input, undefined, undefined, impl.ctx)) as { details: { error?: string } };
|
|
1389
|
+
assert.equal(result.details.error, undefined);
|
|
1390
|
+
}
|
|
1391
|
+
assert.equal(await firstCallResult(impl, "c1", loneImplementer()), undefined, "latch cleared by start implement");
|
|
1392
|
+
await impl.emitEvent("tool_result", waveResult("audit", [{ agent: "conformance-reviewer", exitCode: 0 }]));
|
|
1393
|
+
assert.equal(await firstCallResult(impl, "c2", implementerWave(1)), undefined, "counter cleared by start implement");
|
|
1394
|
+
|
|
1395
|
+
const reset = mk();
|
|
1396
|
+
await reset.emit("session_start");
|
|
1397
|
+
const resetTool = reset.tools.find((t) => t.name === "phase_tracker")!;
|
|
1398
|
+
const resetResult = (await resetTool.execute("p1", { action: "reset" }, undefined, undefined, reset.ctx)) as { details: { error?: string } };
|
|
1399
|
+
assert.equal(resetResult.details.error, undefined);
|
|
1400
|
+
assert.equal(await firstCallResult(reset, "c1", loneImplementer()), undefined, "dormant after reset");
|
|
1401
|
+
});
|
|
1402
|
+
|
|
1403
|
+
test("replay: two implementer waves after the audit in verify restore fixRounds 2; the same waves in ship restore 0", async () => {
|
|
1404
|
+
const inVerify = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 2 } } }, [
|
|
1405
|
+
subagentResult(["implementer"]),
|
|
1406
|
+
subagentResult(["implementer", "code-reviewer"]),
|
|
1407
|
+
]);
|
|
1408
|
+
await inVerify.emit("session_start");
|
|
1409
|
+
assert.equal((await firstCallResult(inVerify, "c1", implementerWave(1)))?.block, true);
|
|
1410
|
+
|
|
1411
|
+
const inShip = harness({
|
|
1412
|
+
cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 2 } } }),
|
|
1413
|
+
branch: verifyBranch([
|
|
1414
|
+
subagentResult(["conformance-reviewer"]),
|
|
1415
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete" })),
|
|
1416
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" })),
|
|
1417
|
+
subagentResult(["implementer"]),
|
|
1418
|
+
subagentResult(["implementer"]),
|
|
1419
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress", ship: "in_progress" })),
|
|
1420
|
+
]),
|
|
1421
|
+
});
|
|
1422
|
+
await inShip.emit("session_start");
|
|
1423
|
+
assert.equal(await firstCallResult(inShip, "c1", implementerWave(1)), undefined);
|
|
1424
|
+
});
|
|
1425
|
+
|
|
1426
|
+
// --- Human overrule of the fix-round cap (spec 2026-09-14-fix-round-human-overrule) ---
|
|
1427
|
+
|
|
1428
|
+
const grantResult = (rounds: number, reason = "human approved") => ({
|
|
1429
|
+
type: "message",
|
|
1430
|
+
message: { role: "toolResult", toolName: "phase_tracker", details: {
|
|
1431
|
+
action: "grant_fix_rounds", rounds, reason,
|
|
1432
|
+
phases: phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress" }),
|
|
1433
|
+
} },
|
|
1434
|
+
});
|
|
1435
|
+
|
|
1436
|
+
const grant = async (h: ReturnType<typeof harness>, id: string, input: Record<string, unknown>) =>
|
|
1437
|
+
(await h.tools.find((t) => t.name === "phase_tracker")!.execute(id, { action: "grant_fix_rounds", ...input }, undefined, undefined, h.ctx)) as {
|
|
1438
|
+
content: { type: string; text: string }[];
|
|
1439
|
+
details: { action: string; rounds?: number; reason?: string; error?: string };
|
|
1440
|
+
};
|
|
1441
|
+
|
|
1442
|
+
const exhaust = async (h: ReturnType<typeof harness>, rounds: number, prefix = "x") => {
|
|
1443
|
+
for (let i = 1; i <= rounds; i++) {
|
|
1444
|
+
assert.equal(await firstCallResult(h, `${prefix}${i}`, implementerWave(1)), undefined);
|
|
1445
|
+
await h.emitEvent("tool_result", okWave(`${prefix}${i}`));
|
|
1446
|
+
}
|
|
1447
|
+
};
|
|
1448
|
+
|
|
1449
|
+
test("grant: funds extra waves and rejects another grant while credits remain", async () => {
|
|
1450
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { enforce: true } } });
|
|
1451
|
+
await h.emit("session_start");
|
|
1452
|
+
await exhaust(h, 3);
|
|
1453
|
+
const res = await grant(h, "g1", { rounds: 2, reason: "I'm approving 2 more rounds" });
|
|
1454
|
+
assert.equal(res.details.error, undefined);
|
|
1455
|
+
assert.deepEqual([res.details.rounds, res.details.reason], [2, "I'm approving 2 more rounds"]);
|
|
1456
|
+
assert.equal((await grant(h, "g2", { rounds: 1, reason: "more" })).details.error, "grant_fix_rounds: 2 granted round(s) still unused; spend them before granting more");
|
|
1457
|
+
await exhaust(h, 2, "c");
|
|
1458
|
+
const blocked = await firstCallResult(h, "c3", implementerWave(1));
|
|
1459
|
+
assert.equal(blocked?.block, true);
|
|
1460
|
+
assert.match(blocked?.reason ?? "", /5 fix round\(s\) used against a cap of 3 \(granted rounds included\)/);
|
|
1461
|
+
});
|
|
1462
|
+
|
|
1463
|
+
test("grant: schema declares rounds as integer 1..MAX_SAFE_INTEGER; missing rounds or empty reason error without state change", async () => {
|
|
1464
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1465
|
+
await h.emit("session_start");
|
|
1466
|
+
const params = h.tools.find((t) => t.name === "phase_tracker")!.parameters;
|
|
1467
|
+
const roundOptions = params.args[0].rounds.args[0].args[0];
|
|
1468
|
+
assert.equal(params.args[0].rounds.args[0].kind, "Integer");
|
|
1469
|
+
assert.deepEqual([roundOptions.minimum, roundOptions.maximum], [1, Number.MAX_SAFE_INTEGER]);
|
|
1470
|
+
assert.ok(params.args[0].action.values.includes("grant_fix_rounds"));
|
|
1471
|
+
|
|
1472
|
+
const noRounds = await grant(h, "g1", { reason: "ok" });
|
|
1473
|
+
assert.equal(noRounds.details.error, "grant_fix_rounds requires rounds: a positive integer");
|
|
1474
|
+
assert.equal(noRounds.details.rounds, undefined);
|
|
1475
|
+
const noReason = await grant(h, "g2", { rounds: 1, reason: " " });
|
|
1476
|
+
assert.equal(noReason.details.error, "grant_fix_rounds requires reason: the human's approval, quoted");
|
|
1477
|
+
assert.equal(noReason.details.rounds, undefined);
|
|
1478
|
+
assert.equal((await firstCallResult(h, "c1", implementerWave(1)))?.block, true, "no credit was granted");
|
|
1479
|
+
});
|
|
1480
|
+
|
|
1481
|
+
test("grant: rejected when no cap block is live - before the audit, in implement, below the cap, enforce off, and in a child", async () => {
|
|
1482
|
+
const expectNotLive = async (h: ReturnType<typeof harness>, used: number, cap: number) => {
|
|
1483
|
+
const res = await grant(h, "g", { rounds: 1, reason: "ok" });
|
|
1484
|
+
assert.equal(res.details.error, `grant_fix_rounds: no fix-round cap block is active (${used} used, cap ${cap}); nothing to overrule`);
|
|
1485
|
+
};
|
|
1486
|
+
|
|
1487
|
+
const preAudit = harness({ cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 0 } } }), branch: verifyBranch() });
|
|
1488
|
+
await preAudit.emit("session_start");
|
|
1489
|
+
await expectNotLive(preAudit, 0, 0);
|
|
1490
|
+
|
|
1491
|
+
const implement = harness({ cwd: tempCwd({ piGauntlet: { closureReview: { maxFixRounds: 0 } } }), branch: implementBranch([subagentResult(["conformance-reviewer"])]) });
|
|
1492
|
+
await implement.emit("session_start");
|
|
1493
|
+
await expectNotLive(implement, 0, 0);
|
|
1494
|
+
|
|
1495
|
+
const below = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 3 } } });
|
|
1496
|
+
await below.emit("session_start");
|
|
1497
|
+
await exhaust(below, 2);
|
|
1498
|
+
await expectNotLive(below, 2, 3);
|
|
1499
|
+
|
|
1500
|
+
const off = guardedHarness({ piGauntlet: { closureReview: { enforce: false, maxFixRounds: 0 } } });
|
|
1501
|
+
await off.emit("session_start");
|
|
1502
|
+
await expectNotLive(off, 0, 0);
|
|
1503
|
+
|
|
1504
|
+
const priorDepth = process.env.PI_SUBAGENT_DEPTH;
|
|
1505
|
+
process.env.PI_SUBAGENT_DEPTH = "1";
|
|
1506
|
+
let child: ReturnType<typeof harness>;
|
|
1507
|
+
try {
|
|
1508
|
+
child = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1509
|
+
} finally {
|
|
1510
|
+
if (priorDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
|
|
1511
|
+
else process.env.PI_SUBAGENT_DEPTH = priorDepth;
|
|
1512
|
+
}
|
|
1513
|
+
await child.emit("session_start");
|
|
1514
|
+
await expectNotLive(child, 0, 0);
|
|
1515
|
+
});
|
|
1516
|
+
|
|
1517
|
+
test("grant: only qualifying synchronous parent task waves spend credits", async () => {
|
|
1518
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1519
|
+
await h.emit("session_start");
|
|
1520
|
+
await grant(h, "g", { rounds: 1, reason: "ok" });
|
|
1521
|
+
await h.emitEvent("tool_result", waveResult("e1", [{ agent: "code-reviewer", exitCode: 0 }]));
|
|
1522
|
+
await h.emitEvent("tool_result", waveResult("e2", [{ agent: "implementer", exitCode: 0 }], true));
|
|
1523
|
+
await h.emitEvent("tool_result", waveResult("e3", []));
|
|
1524
|
+
assert.equal((await firstCallResult(h, "a", { ...implementerWave(1), async: true }))?.block, true);
|
|
1525
|
+
assert.equal((await firstCallResult(h, "l", loneImplementer()))?.block, true);
|
|
1526
|
+
assert.equal(await firstCallResult(h, "c", implementerWave(1)), undefined);
|
|
1527
|
+
await h.emitEvent("tool_result", okWave("c"));
|
|
1528
|
+
assert.equal((await firstCallResult(h, "d", implementerWave(1)))?.block, true);
|
|
1529
|
+
|
|
1530
|
+
// In a child (PI_SUBAGENT_DEPTH >= 1, fixed at registration) both the cap gate and observeFixWave
|
|
1531
|
+
// are dormant, so credit consumption has no public effect; the observable contract is dormancy.
|
|
1532
|
+
const priorDepth = process.env.PI_SUBAGENT_DEPTH;
|
|
1533
|
+
process.env.PI_SUBAGENT_DEPTH = "1";
|
|
1534
|
+
let child: ReturnType<typeof harness>;
|
|
1535
|
+
try {
|
|
1536
|
+
child = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } }, [grantResult(1), subagentResult(["implementer"])]);
|
|
1537
|
+
} finally {
|
|
1538
|
+
if (priorDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
|
|
1539
|
+
else process.env.PI_SUBAGENT_DEPTH = priorDepth;
|
|
1540
|
+
}
|
|
1541
|
+
await child.emit("session_start");
|
|
1542
|
+
assert.equal(await firstCallResult(child, "child-wave", implementerWave(1)), undefined, "child session: cap gate and observer are dormant, so the replayed grant and implementer result have no observable effect");
|
|
1543
|
+
});
|
|
1544
|
+
|
|
1545
|
+
test("grant replay restores and spends the recorded pool independent of current cap; rejected grants restore no credit", async () => {
|
|
1546
|
+
const trail = [subagentResult(["implementer"]), subagentResult(["implementer"]), subagentResult(["implementer"]), grantResult(2), subagentResult(["implementer"])];
|
|
1547
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 3 } } }, trail);
|
|
1548
|
+
await h.emit("session_start");
|
|
1549
|
+
assert.equal(
|
|
1550
|
+
(await grant(h, "g", { rounds: 1, reason: "more" })).details.error,
|
|
1551
|
+
"grant_fix_rounds: 1 granted round(s) still unused; spend them before granting more",
|
|
1552
|
+
);
|
|
1553
|
+
assert.equal(await firstCallResult(h, "c", implementerWave(1)), undefined);
|
|
1554
|
+
await h.emitEvent("tool_result", okWave("c"));
|
|
1555
|
+
const blocked = await firstCallResult(h, "d", implementerWave(1));
|
|
1556
|
+
assert.equal(blocked?.block, true);
|
|
1557
|
+
assert.match(blocked?.reason ?? "", /5 fix round\(s\) used against a cap of 3 \(granted rounds included\)/);
|
|
1558
|
+
|
|
1559
|
+
// Same trail at cap 5: replay yields fixRounds = 4 (cap not live yet); lowering the live cap
|
|
1560
|
+
// exposes the one replayed credit, then wave 5 passes on the counter and wave 6 blocks.
|
|
1561
|
+
const capChanged = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 5 } } }, trail);
|
|
1562
|
+
await capChanged.emit("session_start");
|
|
1563
|
+
assert.equal(
|
|
1564
|
+
(await grant(capChanged, "g", { rounds: 1, reason: "more" })).details.error,
|
|
1565
|
+
"grant_fix_rounds: no fix-round cap block is active (4 used, cap 5); nothing to overrule",
|
|
1566
|
+
"replayed fixRounds = 4 regardless of cap on disk",
|
|
1567
|
+
);
|
|
1568
|
+
// Lower the cap on disk without spending a wave: the block goes live at 4 >= 3 and the grant
|
|
1569
|
+
// probe now reads the replayed pool - exactly one credit, independent of the cap at replay time.
|
|
1570
|
+
writeFileSync(join(capChanged.ctx.cwd, ".pi", "settings.json"), JSON.stringify({ piGauntlet: { closureReview: { maxFixRounds: 3 } } }));
|
|
1571
|
+
assert.equal(
|
|
1572
|
+
(await grant(capChanged, "g2", { rounds: 1, reason: "more" })).details.error,
|
|
1573
|
+
"grant_fix_rounds: 1 granted round(s) still unused; spend them before granting more",
|
|
1574
|
+
"same credits regardless of cap on disk",
|
|
1575
|
+
);
|
|
1576
|
+
writeFileSync(join(capChanged.ctx.cwd, ".pi", "settings.json"), JSON.stringify({ piGauntlet: { closureReview: { maxFixRounds: 5 } } }));
|
|
1577
|
+
assert.equal(await firstCallResult(capChanged, "c", implementerWave(1)), undefined);
|
|
1578
|
+
await capChanged.emitEvent("tool_result", okWave("c"));
|
|
1579
|
+
const blockedAt5 = await firstCallResult(capChanged, "d", implementerWave(1));
|
|
1580
|
+
assert.equal(blockedAt5?.block, true);
|
|
1581
|
+
assert.match(blockedAt5?.reason ?? "", /5 fix round\(s\) used against a cap of 5 \(granted rounds included\)/);
|
|
1582
|
+
|
|
1583
|
+
const rejected = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } }, [{
|
|
1584
|
+
type: "message",
|
|
1585
|
+
message: {
|
|
1586
|
+
role: "toolResult",
|
|
1587
|
+
toolName: "phase_tracker",
|
|
1588
|
+
details: { action: "grant_fix_rounds", error: "grant_fix_rounds requires rounds: a positive integer", phases: phases({ verify: "in_progress" }) },
|
|
1589
|
+
},
|
|
1590
|
+
}]);
|
|
1591
|
+
await rejected.emit("session_start");
|
|
1592
|
+
assert.equal((await firstCallResult(rejected, "c1", implementerWave(1)))?.block, true, "no credits from a rejected grant");
|
|
1593
|
+
});
|
|
1594
|
+
|
|
1595
|
+
test("grant reset: start implement and reset zero credits (live and replay); start verify --force keeps them", async () => {
|
|
1596
|
+
const settings = { piGauntlet: { closureReview: { maxFixRounds: 0 }, flowGuards: { enforce: false } } };
|
|
1597
|
+
|
|
1598
|
+
const force = guardedHarness(settings, [grantResult(1)]);
|
|
1599
|
+
await force.emit("session_start");
|
|
1600
|
+
await force.tools.find((t) => t.name === "phase_tracker")!.execute("p", { action: "start", phase: "verify", force: true }, undefined, undefined, force.ctx);
|
|
1601
|
+
assert.equal(await firstCallResult(force, "c", implementerWave(1)), undefined);
|
|
1602
|
+
|
|
1603
|
+
const liveImpl = guardedHarness(settings);
|
|
1604
|
+
await liveImpl.emit("session_start");
|
|
1605
|
+
await grant(liveImpl, "g", { rounds: 1, reason: "ok" });
|
|
1606
|
+
const tool = liveImpl.tools.find((t) => t.name === "phase_tracker")!;
|
|
1607
|
+
for (const [id, input] of [
|
|
1608
|
+
["p1", { action: "skip", phase: "verify", reason: "amendment" }],
|
|
1609
|
+
["p2", { action: "start", phase: "implement", force: true }],
|
|
1610
|
+
["p3", { action: "complete", phase: "implement" }],
|
|
1611
|
+
["p4", { action: "start", phase: "verify", force: true }],
|
|
1612
|
+
] as const) {
|
|
1613
|
+
assert.equal(((await tool.execute(id, input, undefined, undefined, liveImpl.ctx)) as { details: { error?: string } }).details.error, undefined);
|
|
1614
|
+
}
|
|
1615
|
+
await liveImpl.emitEvent("tool_result", waveResult("audit", [{ agent: "conformance-reviewer", exitCode: 0 }]));
|
|
1616
|
+
assert.equal((await firstCallResult(liveImpl, "c1", implementerWave(1)))?.block, true, "credits zeroed by start implement (cap 0 blocks)");
|
|
1617
|
+
|
|
1618
|
+
const reset = guardedHarness(settings);
|
|
1619
|
+
await reset.emit("session_start");
|
|
1620
|
+
await grant(reset, "g", { rounds: 1, reason: "ok" });
|
|
1621
|
+
const resetTool = reset.tools.find((t) => t.name === "phase_tracker")!;
|
|
1622
|
+
await resetTool.execute("p", { action: "reset" }, undefined, undefined, reset.ctx);
|
|
1623
|
+
assert.match((await grant(reset, "g2", { rounds: 1, reason: "ok" })).details.error ?? "", /no fix-round cap block is active/);
|
|
1624
|
+
for (const [id, input] of [
|
|
1625
|
+
["p1", { action: "start", phase: "brainstorm" }],
|
|
1626
|
+
["p2", { action: "complete", phase: "brainstorm" }],
|
|
1627
|
+
["p3", { action: "start", phase: "plan" }],
|
|
1628
|
+
["p4", { action: "complete", phase: "plan" }],
|
|
1629
|
+
["p5", { action: "start", phase: "implement" }],
|
|
1630
|
+
["p6", { action: "complete", phase: "implement" }],
|
|
1631
|
+
["p7", { action: "start", phase: "verify" }],
|
|
1632
|
+
] as const) {
|
|
1633
|
+
assert.equal(((await resetTool.execute(id, input, undefined, undefined, reset.ctx)) as { details: { error?: string } }).details.error, undefined);
|
|
1634
|
+
}
|
|
1635
|
+
await reset.emitEvent("tool_result", waveResult("audit", [{ agent: "conformance-reviewer", exitCode: 0 }]));
|
|
1636
|
+
assert.equal((await firstCallResult(reset, "c1", implementerWave(1)))?.block, true, "credits zeroed by reset (cap 0 blocks)");
|
|
1637
|
+
|
|
1638
|
+
const replayImpl = harness({
|
|
1639
|
+
cwd: tempCwd(settings),
|
|
1640
|
+
branch: verifyBranch([
|
|
1641
|
+
subagentResult(["conformance-reviewer"]),
|
|
1642
|
+
grantResult(1),
|
|
1643
|
+
phaseResult("skip", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "skipped" })),
|
|
1644
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "in_progress", verify: "skipped" })),
|
|
1645
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "skipped" })),
|
|
1646
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress" })),
|
|
1647
|
+
subagentResult(["conformance-reviewer"]),
|
|
1648
|
+
]),
|
|
1649
|
+
});
|
|
1650
|
+
await replayImpl.emit("session_start");
|
|
1651
|
+
assert.equal((await firstCallResult(replayImpl, "c1", implementerWave(1)))?.block, true, "replayed start implement zeroed credits");
|
|
1652
|
+
|
|
1653
|
+
const replayReset = harness({
|
|
1654
|
+
cwd: tempCwd(settings),
|
|
1655
|
+
branch: verifyBranch([
|
|
1656
|
+
subagentResult(["conformance-reviewer"]),
|
|
1657
|
+
grantResult(1),
|
|
1658
|
+
phaseResult("reset", phases({})),
|
|
1659
|
+
phaseResult("start", phases({ brainstorm: "in_progress" })),
|
|
1660
|
+
phaseResult("complete", phases({ brainstorm: "complete" })),
|
|
1661
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "in_progress" })),
|
|
1662
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete" })),
|
|
1663
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "in_progress" })),
|
|
1664
|
+
phaseResult("complete", phases({ brainstorm: "complete", plan: "complete", implement: "complete" })),
|
|
1665
|
+
phaseResult("start", phases({ brainstorm: "complete", plan: "complete", implement: "complete", verify: "in_progress" })),
|
|
1666
|
+
subagentResult(["conformance-reviewer"]),
|
|
1667
|
+
]),
|
|
1668
|
+
});
|
|
1669
|
+
await replayReset.emit("session_start");
|
|
1670
|
+
assert.equal((await firstCallResult(replayReset, "c1", implementerWave(1)))?.block, true, "replayed reset zeroed credits");
|
|
1671
|
+
});
|
|
1672
|
+
|
|
1673
|
+
test("grant: block reason gives action, settings path, and live-read guidance", async () => {
|
|
1674
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1675
|
+
await h.emit("session_start");
|
|
1676
|
+
const reason = (await firstCallResult(h, "c", implementerWave(1)))?.reason ?? "";
|
|
1677
|
+
assert.match(reason, /phase_tracker\(\{ action: "grant_fix_rounds"/);
|
|
1678
|
+
assert.ok(reason.includes(join(h.ctx.cwd, ".pi", "settings.json")));
|
|
1679
|
+
assert.match(reason, /no restart.*restate model and maxFixRounds/s);
|
|
1680
|
+
});
|
|
1681
|
+
|
|
1682
|
+
test("gauntlet_setting registration requests sequential execution", () => {
|
|
1683
|
+
assert.equal(harness().tools.find((t) => t.name === "gauntlet_setting")!.executionMode, "sequential");
|
|
1684
|
+
});
|
|
1685
|
+
|
|
1686
|
+
test("grant renderResult shows count and reason", async () => {
|
|
1687
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1688
|
+
await h.emit("session_start");
|
|
1689
|
+
const res = await grant(h, "g", { rounds: 2, reason: "go on" });
|
|
1690
|
+
const tool = h.tools.find((t) => t.name === "phase_tracker")! as any;
|
|
1691
|
+
const rendered = tool.renderResult(res, {}, { fg: (_c: string, s: string) => s, bold: (s: string) => s });
|
|
1692
|
+
assert.match(JSON.stringify(rendered), /2.*go on/);
|
|
1693
|
+
});
|
|
1694
|
+
|
|
1695
|
+
test("grant: child sessions cannot overrule even a zero cap", async () => {
|
|
1696
|
+
const priorDepth = process.env.PI_SUBAGENT_DEPTH;
|
|
1697
|
+
process.env.PI_SUBAGENT_DEPTH = "1";
|
|
1698
|
+
let child: ReturnType<typeof harness>;
|
|
1699
|
+
try {
|
|
1700
|
+
child = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1701
|
+
} finally {
|
|
1702
|
+
if (priorDepth === undefined) delete process.env.PI_SUBAGENT_DEPTH;
|
|
1703
|
+
else process.env.PI_SUBAGENT_DEPTH = priorDepth;
|
|
1704
|
+
}
|
|
1705
|
+
await child.emit("session_start");
|
|
1706
|
+
assert.match((await grant(child, "g", { rounds: 1, reason: "ok" })).details.error ?? "", /no fix-round cap block is active/);
|
|
1707
|
+
});
|
|
1708
|
+
|
|
1709
|
+
test("grant: accepts MAX_SAFE_INTEGER and current on-disk cap controls whether the block is live", async () => {
|
|
1710
|
+
const h = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1711
|
+
await h.emit("session_start");
|
|
1712
|
+
const result = await grant(h, "g", { rounds: Number.MAX_SAFE_INTEGER, reason: "approved" });
|
|
1713
|
+
assert.equal(result.details.rounds, Number.MAX_SAFE_INTEGER);
|
|
1714
|
+
|
|
1715
|
+
const changed = guardedHarness({ piGauntlet: { closureReview: { maxFixRounds: 0 } } });
|
|
1716
|
+
await changed.emit("session_start");
|
|
1717
|
+
writeFileSync(join(changed.ctx.cwd, ".pi", "settings.json"), JSON.stringify({ piGauntlet: { closureReview: { maxFixRounds: 1 } } }));
|
|
1718
|
+
assert.match((await grant(changed, "g", { rounds: 1, reason: "ok" })).details.error ?? "", /0 used, cap 1/);
|
|
1719
|
+
});
|
|
@@ -56,9 +56,11 @@ interface PhaseState {
|
|
|
56
56
|
type PhaseMap = Record<Phase, PhaseState>;
|
|
57
57
|
|
|
58
58
|
interface PhaseTrackerDetails {
|
|
59
|
-
action: "start" | "complete" | "skip" | "status" | "reset" | "substep";
|
|
59
|
+
action: "start" | "complete" | "skip" | "status" | "reset" | "substep" | "grant_fix_rounds";
|
|
60
60
|
phases: PhaseMap;
|
|
61
61
|
error?: string;
|
|
62
|
+
rounds?: number;
|
|
63
|
+
reason?: string;
|
|
62
64
|
}
|
|
63
65
|
|
|
64
66
|
interface PlanCheckStamp {
|
|
@@ -86,6 +88,17 @@ const qualifiesAsClosureDispatch = (details: unknown): boolean => {
|
|
|
86
88
|
return d.results.some((r) => r?.agent === "conformance-reviewer" && r?.exitCode === 0);
|
|
87
89
|
};
|
|
88
90
|
|
|
91
|
+
// One conformance fix round = one non-error subagent result carrying at least one
|
|
92
|
+
// implementer child, regardless of dispatch mode or per-child exit code (a failed
|
|
93
|
+
// implementer still spent the round; its retry is the next one). Async and
|
|
94
|
+
// management dispatches return results: [] and never count.
|
|
95
|
+
const isImplementerWave = (details: unknown, isError: boolean | undefined): boolean => {
|
|
96
|
+
if (isError === true) return false;
|
|
97
|
+
const d = details as { results?: { agent?: unknown }[] } | undefined;
|
|
98
|
+
if (!d || !Array.isArray(d.results) || d.results.length === 0) return false;
|
|
99
|
+
return d.results.some((r) => r?.agent === "implementer");
|
|
100
|
+
};
|
|
101
|
+
|
|
89
102
|
// Review-cadence guard (spec 2026-08-12-execution-fidelity-hardening): presence-only
|
|
90
103
|
// advisory ledger of the most recent completed implementer / spec-reviewer /
|
|
91
104
|
// code-reviewer dispatch. Agents completing in the SAME dispatch share a sequence
|
|
@@ -164,22 +177,14 @@ const branchBlockReason = (phase: Phase): string =>
|
|
|
164
177
|
"create/enter one with /skill:using-git-worktrees and run this there. " +
|
|
165
178
|
"To override, set piGauntlet.flowGuards.enforce: false.";
|
|
166
179
|
|
|
167
|
-
//
|
|
168
|
-
//
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
// single / tasks / chain / parallel dispatch shapes and collect the model of every
|
|
172
|
-
// conformance-reviewer entry (undefined = absent/empty/non-string). A bare omission
|
|
173
|
-
// is BLOCKED; an explicit model that differs from the configured one is WARNED
|
|
174
|
-
// (non-blocking), preserving the documented retry-with-fallback hatch while
|
|
175
|
-
// surfacing drift.
|
|
176
|
-
const conformanceModels = (input: unknown): (string | undefined)[] => {
|
|
177
|
-
const models: (string | undefined)[] = [];
|
|
178
|
-
const norm = (m: unknown) => (typeof m === "string" && m.trim() ? m.trim() : undefined);
|
|
180
|
+
// Walk the single / tasks / chain / parallel dispatch shapes and return every
|
|
181
|
+
// entry whose `agent` matches `name` (the raw node, so callers can read its model).
|
|
182
|
+
const collectAgents = (input: unknown, name: string): Record<string, unknown>[] => {
|
|
183
|
+
const hits: Record<string, unknown>[] = [];
|
|
179
184
|
const collect = (node: unknown): void => {
|
|
180
185
|
if (!node || typeof node !== "object") return;
|
|
181
186
|
const o = node as Record<string, unknown>;
|
|
182
|
-
if (o.agent ===
|
|
187
|
+
if (o.agent === name) hits.push(o);
|
|
183
188
|
for (const key of ["tasks", "chain", "parallel"] as const) {
|
|
184
189
|
const v = o[key];
|
|
185
190
|
if (Array.isArray(v)) for (const item of v) collect(item);
|
|
@@ -187,9 +192,31 @@ const conformanceModels = (input: unknown): (string | undefined)[] => {
|
|
|
187
192
|
}
|
|
188
193
|
};
|
|
189
194
|
collect(input);
|
|
190
|
-
return
|
|
195
|
+
return hits;
|
|
191
196
|
};
|
|
192
197
|
|
|
198
|
+
const conformanceModels = (input: unknown): (string | undefined)[] =>
|
|
199
|
+
collectAgents(input, "conformance-reviewer").map((o) =>
|
|
200
|
+
typeof o.model === "string" && o.model.trim() ? o.model.trim() : undefined,
|
|
201
|
+
);
|
|
202
|
+
|
|
203
|
+
// Conformance fix-loop dispatch guard (spec 2026-09-13-conformance-dispatch-guard):
|
|
204
|
+
// after the first audit, pi-cohort honours worktree: true only in tasks mode, so a
|
|
205
|
+
// lone implementer runs unisolated and yields no patch for the integrate step.
|
|
206
|
+
const loneImplementerBlockReason =
|
|
207
|
+
'Conformance fix loop: dispatch implementers as a one-task tasks wave (tasks: [{ agent: "implementer", worktree: true, ... }]); ' +
|
|
208
|
+
"a lone agent call runs unisolated and produces no worktree diff. " +
|
|
209
|
+
"To disable this gate, set piGauntlet.closureReview.enforce: false.";
|
|
210
|
+
|
|
211
|
+
const fixRoundCapBlockReason = (used: number, cap: number, settingsPath: string): string =>
|
|
212
|
+
`Conformance fix loop: ${used} fix round(s) used against a cap of ${cap} (granted rounds included); ` +
|
|
213
|
+
"escalate to the human with the verdict trail instead of re-looping.\n" +
|
|
214
|
+
"If the human explicitly approves more rounds, record it and retry as a tasks wave: " +
|
|
215
|
+
'phase_tracker({ action: "grant_fix_rounds", rounds: <N>, reason: "<their words>" }).\n' +
|
|
216
|
+
`Last resort: set piGauntlet.closureReview.enforce: false in ${settingsPath} (disables all closure guards; ` +
|
|
217
|
+
"applies on the next tool call, no restart - a gauntlet_setting read in the same message sees the write; " +
|
|
218
|
+
"a repo closureReview block replaces the preset's whole block, so restate model and maxFixRounds alongside).";
|
|
219
|
+
|
|
193
220
|
const closureModelBlockReason = (model: string, missing: number, total: number): string =>
|
|
194
221
|
`Blocked: ${missing} of ${total} conformance-reviewer ${total === 1 ? "dispatch" : "entries"} ` +
|
|
195
222
|
`omitted a model while piGauntlet.closureReview.model is set to "${model}".\n` +
|
|
@@ -234,7 +261,7 @@ const pathInSpecDirs = (rawPath: string, specDirs: string[]): boolean => {
|
|
|
234
261
|
};
|
|
235
262
|
|
|
236
263
|
const PhaseTrackerParams = Type.Object({
|
|
237
|
-
action: StringEnum(["start", "complete", "skip", "status", "reset", "substep"] as const, {
|
|
264
|
+
action: StringEnum(["start", "complete", "skip", "status", "reset", "substep", "grant_fix_rounds"] as const, {
|
|
238
265
|
description: "Action to perform",
|
|
239
266
|
}),
|
|
240
267
|
phase: Type.Optional(
|
|
@@ -244,7 +271,14 @@ const PhaseTrackerParams = Type.Object({
|
|
|
244
271
|
),
|
|
245
272
|
reason: Type.Optional(
|
|
246
273
|
Type.String({
|
|
247
|
-
description: "Reason
|
|
274
|
+
description: "Reason (required for skip and grant_fix_rounds; for grant, the human's approval quoted)",
|
|
275
|
+
}),
|
|
276
|
+
),
|
|
277
|
+
rounds: Type.Optional(
|
|
278
|
+
Type.Integer({
|
|
279
|
+
minimum: 1,
|
|
280
|
+
maximum: Number.MAX_SAFE_INTEGER,
|
|
281
|
+
description: "Extra fix rounds the human explicitly approved (grant_fix_rounds only)",
|
|
248
282
|
}),
|
|
249
283
|
),
|
|
250
284
|
force: Type.Optional(
|
|
@@ -306,7 +340,25 @@ function formatStatus(phases: PhaseMap): string {
|
|
|
306
340
|
export default function (pi: ExtensionAPI) {
|
|
307
341
|
let phases: PhaseMap = emptyPhases();
|
|
308
342
|
let conformanceDispatched = false;
|
|
343
|
+
let fixRounds = 0;
|
|
344
|
+
let fixRoundCredits = 0;
|
|
309
345
|
let gauntletEntered = false;
|
|
346
|
+
// Shared by replay and the live tool_result hook so a resumed session enforces
|
|
347
|
+
// the same budget. Evaluated BEFORE the same result may set the latch, so the
|
|
348
|
+
// R0 audit result itself never counts as a wave.
|
|
349
|
+
const observeFixWave = (details: unknown, isError: boolean | undefined, ctx: ExtensionContext) => {
|
|
350
|
+
if (
|
|
351
|
+
!isSubagentChild &&
|
|
352
|
+
gauntletEntered &&
|
|
353
|
+
phases.verify.status === "in_progress" &&
|
|
354
|
+
conformanceDispatched &&
|
|
355
|
+
isImplementerWave(details, isError) &&
|
|
356
|
+
resolveClosureReview(loadGauntletSettings(ctx.cwd).gauntlet).enforce
|
|
357
|
+
) {
|
|
358
|
+
if (fixRoundCredits > 0) fixRoundCredits -= 1;
|
|
359
|
+
fixRounds += 1;
|
|
360
|
+
}
|
|
361
|
+
};
|
|
310
362
|
let planCheckStamp: PlanCheckStamp | undefined;
|
|
311
363
|
const attemptedRecoveryEdges = new Set<RecoveryEdge>();
|
|
312
364
|
|
|
@@ -382,6 +434,8 @@ export default function (pi: ExtensionAPI) {
|
|
|
382
434
|
const reconstructState = (ctx: ExtensionContext) => {
|
|
383
435
|
phases = emptyPhases();
|
|
384
436
|
conformanceDispatched = false;
|
|
437
|
+
fixRounds = 0;
|
|
438
|
+
fixRoundCredits = 0;
|
|
385
439
|
gauntletEntered = false;
|
|
386
440
|
planCheckStamp = undefined;
|
|
387
441
|
attemptedRecoveryEdges.clear();
|
|
@@ -406,12 +460,17 @@ export default function (pi: ExtensionAPI) {
|
|
|
406
460
|
const details = msg.details as PhaseTrackerDetails | undefined;
|
|
407
461
|
if (details && !details.error) {
|
|
408
462
|
phases = details.phases;
|
|
463
|
+
if (details.action === "grant_fix_rounds" && typeof details.rounds === "number") fixRoundCredits = details.rounds;
|
|
409
464
|
gauntletEntered = nextGauntletEntered(gauntletEntered, details.action, details.phases.brainstorm.status);
|
|
410
465
|
if (details.action === "start" && details.phases.implement.status === "in_progress") {
|
|
411
466
|
conformanceDispatched = false;
|
|
467
|
+
fixRounds = 0;
|
|
468
|
+
fixRoundCredits = 0;
|
|
412
469
|
}
|
|
413
470
|
if (details.action === "reset") {
|
|
414
471
|
conformanceDispatched = false;
|
|
472
|
+
fixRounds = 0;
|
|
473
|
+
fixRoundCredits = 0;
|
|
415
474
|
planCheckStamp = undefined;
|
|
416
475
|
}
|
|
417
476
|
}
|
|
@@ -423,6 +482,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
423
482
|
planCheckStamp = undefined;
|
|
424
483
|
}
|
|
425
484
|
} else if (msg.toolName === "subagent") {
|
|
485
|
+
observeFixWave(msg.details, msg.isError, ctx);
|
|
426
486
|
if (qualifiesAsClosureDispatch(msg.details)) conformanceDispatched = true;
|
|
427
487
|
observeCadence(msg.details);
|
|
428
488
|
} else if (msg.toolName === "plan_tracker") {
|
|
@@ -492,10 +552,25 @@ export default function (pi: ExtensionAPI) {
|
|
|
492
552
|
// subagent dispatch never loads settings or leaks a settingsErrorWarning onto its result.
|
|
493
553
|
// Inline-matched (not called) so closureEnforced() stays a lazy second conjunct.
|
|
494
554
|
if (event.toolName === "subagent" && gauntletEntered && closureEnforced()) {
|
|
495
|
-
|
|
496
|
-
//
|
|
497
|
-
// (action: list/get/create/update/delete/status/...) execute nothing, so skip them.
|
|
555
|
+
// Only execution-mode dispatches carry a model or run agents; management/control
|
|
556
|
+
// modes (action: list/get/create/update/delete/status/...) execute nothing, so skip them.
|
|
498
557
|
const hasAction = !!(event.input as { action?: unknown })?.action;
|
|
558
|
+
// Fix-loop window: verify in progress and the R0 audit observed. Read directly,
|
|
559
|
+
// not via activeGuardPhase() - verify is not a GUARD_PHASES member.
|
|
560
|
+
const fixLoopWindow = !isSubagentChild && !hasAction && phases.verify.status === "in_progress" && conformanceDispatched;
|
|
561
|
+
if (fixLoopWindow && (event.input as { agent?: unknown })?.agent === "implementer") {
|
|
562
|
+
return { block: true, reason: loneImplementerBlockReason };
|
|
563
|
+
}
|
|
564
|
+
if (fixLoopWindow && collectAgents(event.input, "implementer").length > 0) {
|
|
565
|
+
const cap = resolveClosureReview(g()).maxFixRounds;
|
|
566
|
+
if (fixRounds >= cap) {
|
|
567
|
+
const funded = fixRoundCredits > 0 && (event.input as { async?: unknown })?.async !== true;
|
|
568
|
+
if (!funded) {
|
|
569
|
+
return { block: true, reason: fixRoundCapBlockReason(fixRounds, cap, join(ctx.cwd, ".pi", "settings.json")) };
|
|
570
|
+
}
|
|
571
|
+
}
|
|
572
|
+
}
|
|
573
|
+
const model = closureReviewModel();
|
|
499
574
|
if (model && !hasAction) {
|
|
500
575
|
const configured = model.trim();
|
|
501
576
|
const models = conformanceModels(event.input);
|
|
@@ -635,8 +710,9 @@ export default function (pi: ExtensionAPI) {
|
|
|
635
710
|
return undefined;
|
|
636
711
|
});
|
|
637
712
|
|
|
638
|
-
pi.on("tool_result", async (event) => {
|
|
713
|
+
pi.on("tool_result", async (event, ctx) => {
|
|
639
714
|
if (event.toolName === "subagent") {
|
|
715
|
+
observeFixWave(event.details, event.isError, ctx);
|
|
640
716
|
if (qualifiesAsClosureDispatch(event.details)) conformanceDispatched = true;
|
|
641
717
|
observeCadence(event.details);
|
|
642
718
|
const warning = pendingGuardWarnings.get(event.toolCallId);
|
|
@@ -678,6 +754,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
678
754
|
label: "Gauntlet Setting",
|
|
679
755
|
description: "Resolve merged piGauntlet.* settings (repo over preset); for skill use only.",
|
|
680
756
|
parameters: GauntletSettingParams,
|
|
757
|
+
executionMode: "sequential",
|
|
681
758
|
async execute(_toolCallId, params, _signal, _onUpdate, ctx) {
|
|
682
759
|
const { gauntlet, errors } = loadGauntletSettings(ctx.cwd);
|
|
683
760
|
const payload =
|
|
@@ -859,7 +936,11 @@ export default function (pi: ExtensionAPI) {
|
|
|
859
936
|
}
|
|
860
937
|
phases = { ...phases, [params.phase]: transitionPhaseState("in_progress") as PhaseState };
|
|
861
938
|
// A rewind must not inherit the prior verify's conformance latch.
|
|
862
|
-
if (params.phase === "implement")
|
|
939
|
+
if (params.phase === "implement") {
|
|
940
|
+
conformanceDispatched = false;
|
|
941
|
+
fixRounds = 0;
|
|
942
|
+
fixRoundCredits = 0;
|
|
943
|
+
}
|
|
863
944
|
gauntletEntered = nextGauntletEntered(gauntletEntered, "start", phases.brainstorm.status);
|
|
864
945
|
firedGuards.clear();
|
|
865
946
|
updateWidget(ctx);
|
|
@@ -1024,6 +1105,8 @@ export default function (pi: ExtensionAPI) {
|
|
|
1024
1105
|
PHASES.map((p) => [p, transitionPhaseState("pending")]),
|
|
1025
1106
|
) as PhaseMap;
|
|
1026
1107
|
conformanceDispatched = false;
|
|
1108
|
+
fixRounds = 0;
|
|
1109
|
+
fixRoundCredits = 0;
|
|
1027
1110
|
gauntletEntered = nextGauntletEntered(gauntletEntered, "reset", phases.brainstorm.status);
|
|
1028
1111
|
firedGuards.clear();
|
|
1029
1112
|
updateWidget(ctx);
|
|
@@ -1033,6 +1116,39 @@ export default function (pi: ExtensionAPI) {
|
|
|
1033
1116
|
};
|
|
1034
1117
|
}
|
|
1035
1118
|
|
|
1119
|
+
case "grant_fix_rounds": {
|
|
1120
|
+
const reject = (error: string) => ({
|
|
1121
|
+
content: [{ type: "text" as const, text: `Error: ${error}` }],
|
|
1122
|
+
details: { action: "grant_fix_rounds", phases: { ...phases }, error } as PhaseTrackerDetails,
|
|
1123
|
+
});
|
|
1124
|
+
if (params.rounds === undefined) return reject("grant_fix_rounds requires rounds: a positive integer");
|
|
1125
|
+
const reason = params.reason?.trim() ?? "";
|
|
1126
|
+
if (reason.length === 0) return reject("grant_fix_rounds requires reason: the human's approval, quoted");
|
|
1127
|
+
const closure = resolveClosureReview(loadGauntletSettings(ctx.cwd).gauntlet);
|
|
1128
|
+
const capLive =
|
|
1129
|
+
!isSubagentChild &&
|
|
1130
|
+
gauntletEntered &&
|
|
1131
|
+
phases.verify.status === "in_progress" &&
|
|
1132
|
+
conformanceDispatched &&
|
|
1133
|
+
closure.enforce &&
|
|
1134
|
+
fixRounds >= closure.maxFixRounds;
|
|
1135
|
+
if (!capLive) {
|
|
1136
|
+
return reject(
|
|
1137
|
+
`grant_fix_rounds: no fix-round cap block is active (${fixRounds} used, cap ${closure.maxFixRounds}); nothing to overrule`,
|
|
1138
|
+
);
|
|
1139
|
+
}
|
|
1140
|
+
if (fixRoundCredits > 0) {
|
|
1141
|
+
return reject(`grant_fix_rounds: ${fixRoundCredits} granted round(s) still unused; spend them before granting more`);
|
|
1142
|
+
}
|
|
1143
|
+
fixRoundCredits = params.rounds;
|
|
1144
|
+
return {
|
|
1145
|
+
content: [
|
|
1146
|
+
{ type: "text", text: `Granted ${params.rounds} extra fix round(s) - reason: "${reason}"\n${formatStatus(phases)}` },
|
|
1147
|
+
],
|
|
1148
|
+
details: { action: "grant_fix_rounds", rounds: params.rounds, reason, phases: { ...phases } } as PhaseTrackerDetails,
|
|
1149
|
+
};
|
|
1150
|
+
}
|
|
1151
|
+
|
|
1036
1152
|
default:
|
|
1037
1153
|
return {
|
|
1038
1154
|
content: [{ type: "text", text: `Unknown action: ${params.action}` }],
|
|
@@ -1104,6 +1220,14 @@ export default function (pi: ExtensionAPI) {
|
|
|
1104
1220
|
}
|
|
1105
1221
|
case "reset":
|
|
1106
1222
|
return new Text(theme.fg("success", "✓ ") + theme.fg("muted", "Phase tracker reset"), 0, 0);
|
|
1223
|
+
case "grant_fix_rounds":
|
|
1224
|
+
return new Text(
|
|
1225
|
+
theme.fg("success", "+ ") +
|
|
1226
|
+
theme.fg("muted", `${details.rounds} extra fix round(s) granted`) +
|
|
1227
|
+
theme.fg("dim", ` (${details.reason})`),
|
|
1228
|
+
0,
|
|
1229
|
+
0,
|
|
1230
|
+
);
|
|
1107
1231
|
default:
|
|
1108
1232
|
return new Text(theme.fg("dim", "Done"), 0, 0);
|
|
1109
1233
|
}
|
|
@@ -22,16 +22,16 @@ const sources = {
|
|
|
22
22
|
}
|
|
23
23
|
`,
|
|
24
24
|
"@sinclair/typebox": `
|
|
25
|
-
const schema = (...args) => ({ args });
|
|
25
|
+
const schema = (kind) => (...args) => ({ kind, args });
|
|
26
26
|
export const Type = {
|
|
27
|
-
Object: schema,
|
|
28
|
-
Optional: schema,
|
|
29
|
-
String: schema,
|
|
30
|
-
Boolean: schema,
|
|
31
|
-
Union: schema,
|
|
32
|
-
Null: schema,
|
|
33
|
-
Array: schema,
|
|
34
|
-
Integer: schema,
|
|
27
|
+
Object: schema("Object"),
|
|
28
|
+
Optional: schema("Optional"),
|
|
29
|
+
String: schema("String"),
|
|
30
|
+
Boolean: schema("Boolean"),
|
|
31
|
+
Union: schema("Union"),
|
|
32
|
+
Null: schema("Null"),
|
|
33
|
+
Array: schema("Array"),
|
|
34
|
+
Integer: schema("Integer"),
|
|
35
35
|
};
|
|
36
36
|
`,
|
|
37
37
|
};
|
package/package.json
CHANGED
|
@@ -105,7 +105,7 @@ When agents return:
|
|
|
105
105
|
|
|
106
106
|
**If integrated changes apply cleanly but the suite fails (semantic conflict):** agents made incompatible assumptions across disjoint files (renamed symbol, changed shape). Diagnose the incompatible pair and re-run the offending task sequentially on the integrated HEAD.
|
|
107
107
|
|
|
108
|
-
**If some agents failed:** Integrate successful agents first (commit their work). Then retry the failed agent with fresh context that includes the integrated changes.
|
|
108
|
+
**If some agents failed:** Integrate successful agents first (commit their work). Then retry the failed agent with fresh context that includes the integrated changes. Inside the conformance loop the retry is a one-task `tasks` wave, never a lone `agent: "implementer"`.
|
|
109
109
|
|
|
110
110
|
## Fix fan-out
|
|
111
111
|
|
|
@@ -126,8 +126,10 @@ partition from an earlier round.
|
|
|
126
126
|
|
|
127
127
|
Mirrors `subagent-driven-development` Parallel-Wave Mode and reuses its
|
|
128
128
|
`plan_tracker` progress surface. Runs entirely inside the gate — it invokes
|
|
129
|
-
**no** `phase_tracker` calls (`
|
|
130
|
-
|
|
129
|
+
**no** phase-transition `phase_tracker` calls (`start`/`complete`/`skip`/`reset`
|
|
130
|
+
- `phase_tracker({ phase: "implement" })` errors while verify is `in_progress`;
|
|
131
|
+
the only `phase_tracker` call inside the loop is the human-approval
|
|
132
|
+
`grant_fix_rounds` in step 6) and does **not** enter SDD's phase machinery.
|
|
131
133
|
Only the fan-out/integrate/review shape and `plan_tracker` are reused. Every execution dispatch is foreground with top-level `async: false`, including retries and prose-described dispatches; an unexpected async handle is a configuration failure: stop and report, never poll or relaunch. `forceTopLevelAsync` is incompatible; see [pi-cohort dispatch configuration](https://github.com/jjuraszek/pi-cohort/blob/main/doc/configuration.md).
|
|
132
134
|
|
|
133
135
|
**Precondition — worktree required.** The loop needs a worktree HEAD to branch
|
|
@@ -157,7 +159,9 @@ Per round:
|
|
|
157
159
|
context; semantic conflict (applies clean, suite fails) → re-run the
|
|
158
160
|
offending task sequentially on integrated HEAD; a failed agent → integrate
|
|
159
161
|
the successes, then retry the failure with fresh context including the
|
|
160
|
-
integrated changes.
|
|
162
|
+
integrated changes. Every re-run or retry inside this loop is itself a
|
|
163
|
+
one-task `tasks` wave (never a lone `agent: "implementer"`) and counts as a
|
|
164
|
+
wave against `maxFixRounds`. A `BLOCKED`/`NEEDS_CONTEXT` return surfaces to the user.
|
|
161
165
|
4. **Scoped tests** on the integrated tree: the round's `SCOPED_TEST_COMMANDS` union. A failure re-enters the failure-handling rules above.
|
|
162
166
|
5. **Re-audit**: foreground re-dispatch `conformance-reviewer` with `async: false` over the fixes **plus** the
|
|
163
167
|
regression guard (any prior-`DELIVERED` requirement whose `evidence` file
|
|
@@ -173,16 +177,22 @@ Per round:
|
|
|
173
177
|
6. **Converge or continue**: verdict `CONFORMS` → Convergence below. Open gaps
|
|
174
178
|
within the cap → re-partition (per the rule above) and start the next
|
|
175
179
|
round. Cap (`gauntlet_setting({ key: "closureReview" }).maxFixRounds`,
|
|
176
|
-
default `
|
|
180
|
+
default `3`, floors negatives at `0`, coerces non-integers to `3`) reached
|
|
177
181
|
with an open `fix` gap or repair item → **escalate to the human** with the per-gap
|
|
178
182
|
round-by-round verdict trail. Escalation is the sole non-completing
|
|
179
|
-
terminal state — no silent re-loop, no auto-ship.
|
|
183
|
+
terminal state — no silent re-loop, no auto-ship. If the human explicitly
|
|
184
|
+
approves N more rounds, record `phase_tracker({ action: "grant_fix_rounds",
|
|
185
|
+
rounds: N, reason: "<their words>" })` and re-enter step 2; without that
|
|
186
|
+
approval, escalation stays terminal. Inside a brainstorming-entered flow
|
|
187
|
+
the phase tracker enforces both rules at
|
|
188
|
+
tool-call time: a lone `agent: "implementer"` dispatch is blocked, and the
|
|
189
|
+
wave after the cap is blocked (`closureReview.enforce: false` disables).
|
|
180
190
|
|
|
181
191
|
**Convergence** — runs after R0 `CONFORMS` and after every `CONFORMS` re-audit. `r0-head` is HEAD when the loop was entered (the R0 dispatch, or the finish-time `fix-now` entry) - the parent of the oldest `conformance fix` commit; `audited-base` stays the last audit's HEAD SHA.
|
|
182
192
|
|
|
183
193
|
a. Run the full plan-header `Verification` set once (ad-hoc: the project's canonical test command). After R0 `CONFORMS` with no round run, the pre-R0 full run counts.
|
|
184
194
|
b. Dispatch `code-reviewer` directly (foreground, `async: false`, `SCOPED_TEST_COMMANDS: none`) over `git diff <r0-head>..HEAD`; never via `/skill:requesting-code-review`. An empty diff is nothing to review - no dispatch.
|
|
185
|
-
c. Repair items = every failing command from a + every Critical/Moderate finding from b (`Behaviour-change: yes` included; the re-audit is its origin check). None → write the closure block; done. Any at the cap → escalate per step 6 with the test/CR trail. Any under the cap → re-enter step 2 as a one-task `tasks` wave: one `implementer` whose task is every repair item verbatim (ownership boundary = the files in `git diff <r0-head>..HEAD`, the CR findings' `touched-files`, and the files each failing command's output names, `SCOPED_TEST_COMMANDS: none`, no `Gn` tracker task, no gap selection), integrate as one `conformance fix CR`, re-audit, then Convergence again. That wave counts against `maxFixRounds`. Later Convergence CRs keep the same `<r0-head>..HEAD` range.
|
|
195
|
+
c. Repair items = every failing command from a + every Critical/Moderate finding from b (`Behaviour-change: yes` included; the re-audit is its origin check). None → write the closure block; done. Any at the cap → escalate per step 6 with the test/CR trail (an explicit human approval re-enters via `grant_fix_rounds`, as step 6 describes). Any under the cap → re-enter step 2 as a one-task `tasks` wave: one `implementer` whose task is every repair item verbatim (ownership boundary = the files in `git diff <r0-head>..HEAD`, the CR findings' `touched-files`, and the files each failing command's output names, `SCOPED_TEST_COMMANDS: none`, no `Gn` tracker task, no gap selection), integrate as one `conformance fix CR`, re-audit, then Convergence again. That wave counts against `maxFixRounds`. Later Convergence CRs keep the same `<r0-head>..HEAD` range.
|
|
186
196
|
|
|
187
197
|
`conformance fix CR` is not a gap fix: it is absent from the `auto-applied fix commits` index and has no `revert conformance fix Gn` action at the finish gate.
|
|
188
198
|
|
|
@@ -22,7 +22,7 @@ does not recurse into their leaves.
|
|
|
22
22
|
whole-object, a repo file that sets only *one* leaf of a key silently drops the
|
|
23
23
|
preset's other leaves for that key. A repo `closureReview: { "model": "..." }` with
|
|
24
24
|
no `enforce`/`maxFixRounds` makes those fall back to their code defaults
|
|
25
|
-
(`enforce` true, `maxFixRounds`
|
|
25
|
+
(`enforce` true, `maxFixRounds` 3), **not** to the preset's values. Define every
|
|
26
26
|
leaf you care about together in the file that owns the key. (Sibling keys are
|
|
27
27
|
unaffected - only the key the repo redefines is replaced.)
|
|
28
28
|
|