pi-gauntlet 5.3.6 → 5.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/README.md +2 -2
- package/extensions/lib/gauntlet-settings.test.ts +37 -0
- package/extensions/lib/gauntlet-settings.ts +24 -0
- package/extensions/lib/plan-check.test.ts +57 -0
- package/extensions/lib/plan-check.ts +11 -2
- package/extensions/phase-tracker.test.ts +54 -1
- package/extensions/phase-tracker.ts +15 -2
- package/package.json +1 -1
- package/skills/brainstorming/SKILL.md +35 -47
- package/skills/brainstorming/gatherer.md +6 -1
- package/skills/finishing-a-development-branch/SKILL.md +1 -1
- package/skills/subagent-driven-development/SKILL.md +17 -25
- package/skills/subagent-driven-development/stop-note.md +48 -0
- package/skills/verification-before-completion/reference/settings-precedence.md +1 -1
- package/skills/writing-plans/SKILL.md +4 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,18 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## v5.4.0 - 2026-09-10
|
|
4
|
+
|
|
5
|
+
- `subagent-driven-development`: a stalled review fix loop runs one escalated fix round (`implementer`, `context: fresh`, model from new `piGauntlet.escalationLoop.implModel`, default main-loop model + thinking) before stopping; the stop is a one-screen problem note (`stop-note.md`) with concrete fix options, replacing the trajectory-log escalation report. `gauntlet_setting` gains the `escalationLoop` key. (#29)
|
|
6
|
+
|
|
7
|
+
## v5.3.7 - 2026-09-10
|
|
8
|
+
|
|
9
|
+
- `brainstorming`: the standalone "does this replace a prior spec" question is gone - the gather scout names candidate predecessor specs and round 1 states them; the design is presented in two rounds (architecture/components/data flow, then errors/testing/docs) with one approval each. (#27)
|
|
10
|
+
- `brainstorming`: new `## Amending an approved spec` section - diff + one-line impact approval, `plan_check` re-stamp, redraw test for large changes; `writing-plans`, `subagent-driven-development`, `finishing-a-development-branch` link to it instead of "frozen spec" wording. (#27)
|
|
11
|
+
- `plan_check`: `quote-integrity` treats a required literal equal to the header entrypoint as satisfied by the header, so a spec line carrying both a scoped command and the full-suite entrypoint no longer forces a spec edit. (#27)
|
|
12
|
+
- `brainstorming`, `subagent-driven-development`: Red Flags lists replaced by exact 10-bullet one-sentence lists. (#27)
|
|
13
|
+
- `phase-tracker`: restarting `implement` clears the conformance-dispatch latch, so a verify re-entered after an amendment rewind reruns the closure gate. (#27)
|
|
14
|
+
- `writing-plans`: plan-split decomposition and `brainstorming` first-feature decisions are stated, not confirmed - no standalone approval prompts. (#27)
|
|
15
|
+
|
|
3
16
|
## v5.3.6 - 2026-09-08
|
|
4
17
|
|
|
5
18
|
- `plan_check`: new `waiver-literal` check (9 checks) - a `waived:` coverage row whose requirement names an inline code literal fails; `writing-plans` restates the waiver criterion (out of scope **and** excludes work) and requires cross-cutting requirements to list every deciding task. (#25)
|
package/README.md
CHANGED
|
@@ -36,7 +36,7 @@ pi-gauntlet's only hard dependency is pi-cohort - every gate that dispatches a r
|
|
|
36
36
|
Concretely, one change through the gauntlet:
|
|
37
37
|
|
|
38
38
|
0. *(Optional)* Before there's even a spec, `/skill:shape-ticket` can create or repair a single tracker issue - shaping a raw ask into a Context/Problem/Idea/Acceptance Criteria ticket, gated by an AC integrity check, a cheap council roast, and one human-confirmed write that may include one optional gated Reporter-note comment. A failed roast is retried once, then surfaced inline at the gate if it fails again. It's a tool, not a phase: no worktree, no plan/phase tracker, runs from any repo state. It never activates on its own (`disable-model-invocation: true`) - invoke it explicitly. Similarly, `/skill:chase-bug` triages a raw bug report into an evidenced verdict - and can hand off to shape-ticket, brainstorming, or a bounded hotfix - before any spec exists.
|
|
39
|
-
1. You describe the change. **`brainstorming`** sets up an isolated worktree, explores the codebase, and turns your description into a written spec. A multi-model critique runs on it automatically. If the spec replaces a known prior spec, brainstorming marks the predecessor with a `> **Superseded by:**` banner under its title (default format, syntax overridable via `.pi/gauntlet-overrides.md`;
|
|
39
|
+
1. You describe the change. **`brainstorming`** sets up an isolated worktree, explores the codebase, and turns your description into a written spec. A multi-model critique runs on it automatically. If the spec replaces a known prior spec, brainstorming marks the predecessor with a `> **Superseded by:**` banner under its title (default format, syntax overridable via `.pi/gauntlet-overrides.md`; candidates come from brainstorming's scout recon, never a mechanical sweep). **You read and approve the spec - human gate 1.** No implementation code exists yet.
|
|
40
40
|
2. **`writing-plans`** decomposes the approved spec into atomic, independently-verifiable tasks, grouped into parallel waves where they don't touch the same files.
|
|
41
41
|
3. **`subagent-driven-development`** executes the plan one task at a time, each in a fresh subagent, behind spec-compliance review then code-quality review. TDD-locked: red, green, refactor. Independent implementation, review, and council batches remain parallel, but their dispatches are foreground: the orchestrator waits for terminal results before accepting work or advancing a phase.
|
|
42
42
|
4. **verify**: the parent first runs the plan's full verification command set; only a passing result permits whole-diff code review, then the **conformance gate**. A review fix invalidates that result, so the parent reruns full verification before the next review or conformance gate. The conformance subagent reads the finished code and docs against your *original words* from step 1, not the plan, and reports per-requirement: delivered, partial, missing, drifted, or unauthorized. Unauthorized rows - shipped surface no human input asked for, whether it crept in or was laundered through the spec - follow their recommendation like every other row: a contained removal auto-runs, anything another requirement leans on is deferred with a plain-language "I'd cut it / I'd keep it" recommendation. Inside a brainstorming-entered flow this gate is machine-blocked from being skipped. Compatible executable recommendations auto-run through an isolated fix-and-re-audit loop with no prompt; anything still open surfaces as a dense list - one line per decision, plain-language, with its recommended choice inline. Reply `1` to take every recommendation, or `2:` with per-item overrides; a current `CONFORMS` / no-concerns result goes straight to the branch options with no extra conformance sign-off.
|
|
@@ -63,7 +63,7 @@ flowchart LR
|
|
|
63
63
|
|
|
64
64
|
<!-- TODO GIF: a real gauntlet run end to end -->
|
|
65
65
|
|
|
66
|
-
Everything between gate 1 and gate 2 - task breakdown, implementation, both review passes - runs without you in the loop. That's the mechanism. What follows is the machinery behind it.
|
|
66
|
+
Everything between gate 1 and gate 2 - task breakdown, implementation, both review passes - runs without you in the loop. Changing an approved spec later is a conditional diff-approval stop (brainstorming's `Amending an approved spec`), not a third numbered gate. That's the mechanism. What follows is the machinery behind it.
|
|
67
67
|
|
|
68
68
|
## Architecture
|
|
69
69
|
|
|
@@ -4,6 +4,8 @@ import {
|
|
|
4
4
|
mergeGauntlet,
|
|
5
5
|
resolveSpecCouncil,
|
|
6
6
|
resolveClosureReview,
|
|
7
|
+
resolveEscalationLoop,
|
|
8
|
+
mainLoopModel,
|
|
7
9
|
resolveFlowGuards,
|
|
8
10
|
resolveVerifyBeforeShip,
|
|
9
11
|
settingsErrorWarning,
|
|
@@ -24,6 +26,41 @@ test("mergeGauntlet: undefined layers -> {}", () => {
|
|
|
24
26
|
assert.deepEqual(mergeGauntlet(undefined, undefined), {});
|
|
25
27
|
});
|
|
26
28
|
|
|
29
|
+
test("mergeGauntlet: repo escalationLoop replaces preset whole-object", () => {
|
|
30
|
+
const preset = { escalationLoop: { implModel: "p/preset:high" }, closureReview: { model: "m" } };
|
|
31
|
+
const repo = { escalationLoop: {} };
|
|
32
|
+
const merged = mergeGauntlet(preset, repo);
|
|
33
|
+
assert.deepEqual(merged.escalationLoop, {});
|
|
34
|
+
assert.deepEqual(merged.closureReview, { model: "m" });
|
|
35
|
+
});
|
|
36
|
+
|
|
37
|
+
test("escalationLoop: absent/empty/null/non-string -> mainLoop", () => {
|
|
38
|
+
const main = "p/main:medium";
|
|
39
|
+
assert.equal(resolveEscalationLoop({}, main).implModel, main);
|
|
40
|
+
assert.equal(resolveEscalationLoop({ escalationLoop: {} }, main).implModel, main);
|
|
41
|
+
assert.equal(resolveEscalationLoop({ escalationLoop: { implModel: "" } }, main).implModel, main);
|
|
42
|
+
assert.equal(resolveEscalationLoop({ escalationLoop: { implModel: null } }, main).implModel, main);
|
|
43
|
+
assert.equal(resolveEscalationLoop({ escalationLoop: { implModel: 42 } }, main).implModel, main);
|
|
44
|
+
});
|
|
45
|
+
|
|
46
|
+
test("escalationLoop: non-empty string wins, trimmed; undefined mainLoop passes through", () => {
|
|
47
|
+
assert.equal(
|
|
48
|
+
resolveEscalationLoop({ escalationLoop: { implModel: " p/x:high " } }, "p/main:medium").implModel,
|
|
49
|
+
"p/x:high",
|
|
50
|
+
);
|
|
51
|
+
assert.equal(resolveEscalationLoop({}, undefined).implModel, undefined);
|
|
52
|
+
});
|
|
53
|
+
|
|
54
|
+
test("mainLoopModel: always suffixed; unset -> off, max -> xhigh, recognised pass through", () => {
|
|
55
|
+
const m = { provider: "p", id: "id" };
|
|
56
|
+
assert.equal(mainLoopModel(m, undefined), "p/id:off");
|
|
57
|
+
assert.equal(mainLoopModel(m, "off"), "p/id:off");
|
|
58
|
+
assert.equal(mainLoopModel(m, "medium"), "p/id:medium");
|
|
59
|
+
assert.equal(mainLoopModel(m, "xhigh"), "p/id:xhigh");
|
|
60
|
+
assert.equal(mainLoopModel(m, "max"), "p/id:xhigh");
|
|
61
|
+
assert.equal(mainLoopModel(undefined, "medium"), undefined);
|
|
62
|
+
});
|
|
63
|
+
|
|
27
64
|
test("specCouncil: non-empty string array -> council", () => {
|
|
28
65
|
const r = resolveSpecCouncil({ specCouncil: { members: ["p/m1", " p/m2 "], chair: "p/c" } });
|
|
29
66
|
assert.equal(r.verdict, "council");
|
|
@@ -8,6 +8,7 @@ export interface PiGauntlet {
|
|
|
8
8
|
closureReview?: { enforce?: unknown; model?: unknown; maxFixRounds?: unknown };
|
|
9
9
|
flowGuards?: { enforce?: unknown; specDirs?: unknown };
|
|
10
10
|
verifyBeforeShip?: { testCommands?: unknown; warningReference?: unknown };
|
|
11
|
+
escalationLoop?: { implModel?: unknown };
|
|
11
12
|
}
|
|
12
13
|
|
|
13
14
|
// Whole-object second-level merge: each piGauntlet key present in the repo layer
|
|
@@ -87,6 +88,29 @@ export function resolveClosureReview(g: PiGauntlet): ClosureReviewResolved {
|
|
|
87
88
|
return { model, enforce, maxFixRounds };
|
|
88
89
|
}
|
|
89
90
|
|
|
91
|
+
export interface EscalationLoopResolved {
|
|
92
|
+
implModel: string | undefined;
|
|
93
|
+
}
|
|
94
|
+
|
|
95
|
+
export function resolveEscalationLoop(g: PiGauntlet, mainLoop: string | undefined): EscalationLoopResolved {
|
|
96
|
+
const raw = g.escalationLoop?.implModel;
|
|
97
|
+
return { implModel: nonEmptyString(raw) ? raw.trim() : mainLoop };
|
|
98
|
+
}
|
|
99
|
+
|
|
100
|
+
const THINKING_SUFFIXES = new Set(["off", "minimal", "low", "medium", "high", "xhigh"]);
|
|
101
|
+
|
|
102
|
+
// Always emit a suffix so pi-cohort's applyThinkingSuffix never falls back to the
|
|
103
|
+
// implementer's configured thinking; pi's "max" has no pi-cohort equivalent -> xhigh.
|
|
104
|
+
export function mainLoopModel(
|
|
105
|
+
model: { provider: string; id: string } | undefined,
|
|
106
|
+
thinkingLevel: string | undefined,
|
|
107
|
+
): string | undefined {
|
|
108
|
+
if (!model) return undefined;
|
|
109
|
+
const level =
|
|
110
|
+
thinkingLevel === "max" ? "xhigh" : THINKING_SUFFIXES.has(thinkingLevel ?? "") ? thinkingLevel : "off";
|
|
111
|
+
return `${model.provider}/${model.id}:${level}`;
|
|
112
|
+
}
|
|
113
|
+
|
|
90
114
|
export interface FlowGuardsResolved {
|
|
91
115
|
enforce: boolean;
|
|
92
116
|
specDirs: string[];
|
|
@@ -380,6 +380,63 @@ test("Verification quote-integrity: task-owned literal check unchanged", () => {
|
|
|
380
380
|
assert.ok(qi.some((f) => f.reason.includes("Task 1 body does not contain the required verbatim literal `helperFn()`")));
|
|
381
381
|
});
|
|
382
382
|
|
|
383
|
+
const ENTRYPOINT_SPEC = [
|
|
384
|
+
"# Fixture Spec", // 1
|
|
385
|
+
"", // 2
|
|
386
|
+
"## Testing", // 3
|
|
387
|
+
"Run `node --test extensions/lib/plan-check.test.ts`; the full suite is `npm run verify-all`.", // 4
|
|
388
|
+
].join("\n");
|
|
389
|
+
|
|
390
|
+
const ENTRYPOINT_PLAN = `# Fixture Plan
|
|
391
|
+
|
|
392
|
+
**Spec:** \`doc/specs/fixture-spec.md\`
|
|
393
|
+
|
|
394
|
+
**Verification:** \`npm run verify-all\`
|
|
395
|
+
|
|
396
|
+
---
|
|
397
|
+
|
|
398
|
+
## Wave 1 — Solo
|
|
399
|
+
|
|
400
|
+
Solo: lone remaining task
|
|
401
|
+
|
|
402
|
+
### Task 1: Scoped test
|
|
403
|
+
|
|
404
|
+
**Spec:** doc/specs/fixture-spec.md § "Testing" L4
|
|
405
|
+
|
|
406
|
+
**Files:**
|
|
407
|
+
- Modify: extensions/lib/fixture-task1.ts
|
|
408
|
+
|
|
409
|
+
Run node --test extensions/lib/plan-check.test.ts and confirm green.
|
|
410
|
+
|
|
411
|
+
## Spec coverage
|
|
412
|
+
|
|
413
|
+
| anchor | requirement | owner |
|
|
414
|
+
|---|---|---|
|
|
415
|
+
| § "Testing" L4 | scoped test run | Task 1 |
|
|
416
|
+
`;
|
|
417
|
+
|
|
418
|
+
test("quote-integrity: header entrypoint literal on an anchored line is satisfied by the header, not the task body", () => {
|
|
419
|
+
const findings = checkPlan(ENTRYPOINT_PLAN, ENTRYPOINT_SPEC, alwaysTruePort());
|
|
420
|
+
assert.deepEqual(findingsFor(findings, "quote-integrity"), []);
|
|
421
|
+
assert.deepEqual(findingsFor(findings, "header-entrypoint"), []);
|
|
422
|
+
});
|
|
423
|
+
|
|
424
|
+
test("quote-integrity: scoped command on the same anchored line is still required in the task body", () => {
|
|
425
|
+
const mutated = ENTRYPOINT_PLAN.replace(
|
|
426
|
+
"Run node --test extensions/lib/plan-check.test.ts and confirm green.",
|
|
427
|
+
"Run the scoped test and confirm green.",
|
|
428
|
+
);
|
|
429
|
+
const qi = findingsFor(checkPlan(mutated, ENTRYPOINT_SPEC, alwaysTruePort()), "quote-integrity");
|
|
430
|
+
assert.equal(qi.length, 1);
|
|
431
|
+
assert.ok(qi[0].reason.includes("Task 1 body does not contain the required verbatim literal `node --test extensions/lib/plan-check.test.ts`"));
|
|
432
|
+
});
|
|
433
|
+
|
|
434
|
+
test("placeholder-scan: header entrypoint is not a required literal for the task (parity with quote-integrity)", () => {
|
|
435
|
+
const findings = checkPlan(ENTRYPOINT_PLAN, ENTRYPOINT_SPEC, alwaysTruePort());
|
|
436
|
+
assert.deepEqual(findingsFor(findings, "placeholder-scan"), []);
|
|
437
|
+
assert.deepEqual(findings, []);
|
|
438
|
+
});
|
|
439
|
+
|
|
383
440
|
test("check 3 anchor-resolution: ambiguous heading match (duplicate spec heading)", () => {
|
|
384
441
|
const dupSpec = SPEC_TEXT.replace('## Testing', '## Design\n\nduplicate section body.\n\n## Testing');
|
|
385
442
|
const findings = checkPlan(VALID_PLAN, dupSpec, alwaysTruePort());
|
|
@@ -367,12 +367,20 @@ function requiredLiteralsForRow(row: CoverageRow, specLines: string[]): string[]
|
|
|
367
367
|
return extractLiterals(text);
|
|
368
368
|
}
|
|
369
369
|
|
|
370
|
+
function dropHeaderEntrypoint(literals: string[], parsed: ParsedPlan): string[] {
|
|
371
|
+
const entrypoint = (parsed.header.verificationText ?? "").replaceAll("`", "");
|
|
372
|
+
if (!entrypoint) return literals;
|
|
373
|
+
return literals.filter((lit) => lit !== entrypoint);
|
|
374
|
+
}
|
|
375
|
+
|
|
370
376
|
function computeRequiredLiteralsPerTask(parsed: ParsedPlan, specLines: string[]): Map<number, string[]> {
|
|
371
377
|
const map = new Map<number, string[]>();
|
|
372
378
|
if (!parsed.coverageTableFound) return map;
|
|
373
379
|
for (const row of parsed.coverageRows) {
|
|
374
380
|
if (row.ownerMalformed || row.isWaived || row.isMechanical) continue;
|
|
375
|
-
const literals =
|
|
381
|
+
const literals = row.isVerification
|
|
382
|
+
? requiredLiteralsForRow(row, specLines)
|
|
383
|
+
: dropHeaderEntrypoint(requiredLiteralsForRow(row, specLines), parsed);
|
|
376
384
|
if (literals.length === 0) continue;
|
|
377
385
|
for (const n of row.ownerTasks) {
|
|
378
386
|
const arr = map.get(n) ?? [];
|
|
@@ -506,11 +514,12 @@ function checkQuoteIntegrity(parsed: ParsedPlan, specLines: string[]): PlanCheck
|
|
|
506
514
|
}
|
|
507
515
|
continue;
|
|
508
516
|
}
|
|
517
|
+
const taskLiterals = dropHeaderEntrypoint(literals, parsed);
|
|
509
518
|
for (const n of row.ownerTasks) {
|
|
510
519
|
const task = taskByNumber.get(n);
|
|
511
520
|
if (!task) continue;
|
|
512
521
|
const body = taskBodyText(task, parsed.lines);
|
|
513
|
-
for (const lit of
|
|
522
|
+
for (const lit of taskLiterals) {
|
|
514
523
|
if (!body.includes(lit)) {
|
|
515
524
|
findings.push({
|
|
516
525
|
check: "quote-integrity",
|
|
@@ -51,7 +51,7 @@ const resumedBranch = (rest: Partial<Record<Phase, Status>>) => [
|
|
|
51
51
|
phaseResult("complete", phases({ brainstorm: "skipped", ...rest })),
|
|
52
52
|
];
|
|
53
53
|
|
|
54
|
-
function harness(options: { cwd?: string; branch?: unknown[]; idle?: boolean; beforeSettled?: (setIdle: (idle: boolean) => void) => void; sendThrows?: boolean } = {}) {
|
|
54
|
+
function harness(options: { cwd?: string; branch?: unknown[]; idle?: boolean; beforeSettled?: (setIdle: (idle: boolean) => void) => void; sendThrows?: boolean; model?: { provider: string; id: string }; thinkingLevel?: string } = {}) {
|
|
55
55
|
const handlers = new Map<string, ((event: unknown, ctx: unknown) => unknown)[]>();
|
|
56
56
|
const tools: { name: string; execute: (...args: any[]) => unknown }[] = [];
|
|
57
57
|
const sent: { message: any; options: any }[] = [];
|
|
@@ -62,6 +62,8 @@ function harness(options: { cwd?: string; branch?: unknown[]; idle?: boolean; be
|
|
|
62
62
|
hasUI: false,
|
|
63
63
|
isIdle: () => idle,
|
|
64
64
|
sessionManager: { getBranch: () => branch },
|
|
65
|
+
model: options.model,
|
|
66
|
+
thinkingLevel: options.thinkingLevel,
|
|
65
67
|
};
|
|
66
68
|
const pi = {
|
|
67
69
|
on(event: string, handler: (event: unknown, context: unknown) => unknown) {
|
|
@@ -290,6 +292,46 @@ test("resumed session: closure gate blocks complete verify without a conformance
|
|
|
290
292
|
assert.equal(res.details.error, "no conformance-reviewer dispatch observed");
|
|
291
293
|
});
|
|
292
294
|
|
|
295
|
+
test("restarting implement clears a resumed conformance dispatch latch", async () => {
|
|
296
|
+
const h = harness({
|
|
297
|
+
cwd: tempCwd({ piGauntlet: { flowGuards: { enforce: false }, closureReview: { enforce: true } } }),
|
|
298
|
+
branch: [
|
|
299
|
+
...resumedBranch({ plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" }),
|
|
300
|
+
subagentResult(["conformance-reviewer"]),
|
|
301
|
+
],
|
|
302
|
+
});
|
|
303
|
+
await h.emit("session_start");
|
|
304
|
+
const tool = h.tools.find((t) => t.name === "phase_tracker")!;
|
|
305
|
+
await tool.execute("t1", { action: "skip", phase: "ship", reason: "amendment reopened Task 1" }, undefined, undefined, h.ctx);
|
|
306
|
+
await tool.execute("t2", { action: "start", phase: "implement", force: true }, undefined, undefined, h.ctx);
|
|
307
|
+
await tool.execute("t3", { action: "skip", phase: "implement", reason: "amendment implementation tested separately" }, undefined, undefined, h.ctx);
|
|
308
|
+
await tool.execute("t4", { action: "start", phase: "verify", force: true }, undefined, undefined, h.ctx);
|
|
309
|
+
const res = (await tool.execute("t5", { action: "complete", phase: "verify" }, undefined, undefined, h.ctx)) as {
|
|
310
|
+
details: { error?: string };
|
|
311
|
+
};
|
|
312
|
+
assert.equal(res.details.error, "no conformance-reviewer dispatch observed");
|
|
313
|
+
});
|
|
314
|
+
|
|
315
|
+
test("replayed implement restart clears a persisted conformance dispatch latch", async () => {
|
|
316
|
+
const h = harness({
|
|
317
|
+
cwd: tempCwd({ piGauntlet: { flowGuards: { enforce: false }, closureReview: { enforce: true } } }),
|
|
318
|
+
branch: [
|
|
319
|
+
...resumedBranch({ plan: "complete", implement: "complete", verify: "complete", ship: "in_progress" }),
|
|
320
|
+
subagentResult(["conformance-reviewer"]),
|
|
321
|
+
phaseResult("skip", phases({ brainstorm: "skipped", plan: "complete", implement: "complete", verify: "complete", ship: "skipped" })),
|
|
322
|
+
phaseResult("start", phases({ brainstorm: "skipped", plan: "complete", implement: "in_progress", verify: "complete", ship: "skipped" })),
|
|
323
|
+
phaseResult("skip", phases({ brainstorm: "skipped", plan: "complete", implement: "skipped", verify: "complete", ship: "skipped" })),
|
|
324
|
+
phaseResult("start", phases({ brainstorm: "skipped", plan: "complete", implement: "skipped", verify: "in_progress", ship: "skipped" })),
|
|
325
|
+
],
|
|
326
|
+
});
|
|
327
|
+
await h.emit("session_start");
|
|
328
|
+
const tool = h.tools.find((t) => t.name === "phase_tracker")!;
|
|
329
|
+
const res = (await tool.execute("t1", { action: "complete", phase: "verify" }, undefined, undefined, h.ctx)) as {
|
|
330
|
+
details: { error?: string };
|
|
331
|
+
};
|
|
332
|
+
assert.equal(res.details.error, "no conformance-reviewer dispatch observed");
|
|
333
|
+
});
|
|
334
|
+
|
|
293
335
|
const taskSnapshot = (tasks: { name: string; status: string }[], isError = false) => ({
|
|
294
336
|
type: "message",
|
|
295
337
|
message: { role: "toolResult", toolName: "plan_tracker", isError, details: { tasks } },
|
|
@@ -1101,3 +1143,14 @@ test("replay via session_switch: pass then fail clears the stamp, rebuilt from r
|
|
|
1101
1143
|
};
|
|
1102
1144
|
assert.match(res.details.error ?? "", /plan_check/);
|
|
1103
1145
|
});
|
|
1146
|
+
|
|
1147
|
+
test("gauntlet_setting escalationLoop: setting absent -> ctx-derived main-loop model; setting wins when set", async () => {
|
|
1148
|
+
const h = harness({ cwd: tempCwd({ piGauntlet: {} }), model: { provider: "p", id: "main" }, thinkingLevel: "medium" });
|
|
1149
|
+
const tool = h.tools.find((t) => t.name === "gauntlet_setting")!;
|
|
1150
|
+
const absent = (await tool.execute("g1", { key: "escalationLoop" }, undefined, undefined, h.ctx)) as { details: { key: string; implModel?: string; errors: string[] } };
|
|
1151
|
+
assert.deepEqual(absent.details, { key: "escalationLoop", implModel: "p/main:medium", errors: [] });
|
|
1152
|
+
|
|
1153
|
+
const set = harness({ cwd: tempCwd({ piGauntlet: { escalationLoop: { implModel: "p/strong:high" } } }), model: { provider: "p", id: "main" }, thinkingLevel: "medium" });
|
|
1154
|
+
const res = (await set.tools.find((t) => t.name === "gauntlet_setting")!.execute("g2", { key: "escalationLoop" }, undefined, undefined, set.ctx)) as { details: { implModel?: string } };
|
|
1155
|
+
assert.equal(res.details.implModel, "p/strong:high");
|
|
1156
|
+
});
|
|
@@ -17,7 +17,9 @@ import type { ExtensionAPI, ExtensionContext, Theme } from "@earendil-works/pi-c
|
|
|
17
17
|
import { Text } from "@earendil-works/pi-tui";
|
|
18
18
|
import { type Static, Type } from "@sinclair/typebox";
|
|
19
19
|
import {
|
|
20
|
+
mainLoopModel,
|
|
20
21
|
resolveClosureReview,
|
|
22
|
+
resolveEscalationLoop,
|
|
21
23
|
resolveFlowGuards,
|
|
22
24
|
resolveSpecCouncil,
|
|
23
25
|
settingsErrorWarning,
|
|
@@ -405,6 +407,9 @@ export default function (pi: ExtensionAPI) {
|
|
|
405
407
|
if (details && !details.error) {
|
|
406
408
|
phases = details.phases;
|
|
407
409
|
gauntletEntered = nextGauntletEntered(gauntletEntered, details.action, details.phases.brainstorm.status);
|
|
410
|
+
if (details.action === "start" && details.phases.implement.status === "in_progress") {
|
|
411
|
+
conformanceDispatched = false;
|
|
412
|
+
}
|
|
408
413
|
if (details.action === "reset") {
|
|
409
414
|
conformanceDispatched = false;
|
|
410
415
|
planCheckStamp = undefined;
|
|
@@ -663,7 +668,7 @@ export default function (pi: ExtensionAPI) {
|
|
|
663
668
|
});
|
|
664
669
|
|
|
665
670
|
const GauntletSettingParams = Type.Object({
|
|
666
|
-
key: StringEnum(["specCouncil", "closureReview"] as const, {
|
|
671
|
+
key: StringEnum(["specCouncil", "closureReview", "escalationLoop"] as const, {
|
|
667
672
|
description: "Which gauntlet setting to resolve (merged repo-over-preset).",
|
|
668
673
|
}),
|
|
669
674
|
});
|
|
@@ -678,7 +683,13 @@ export default function (pi: ExtensionAPI) {
|
|
|
678
683
|
const payload =
|
|
679
684
|
params.key === "specCouncil"
|
|
680
685
|
? { key: "specCouncil" as const, ...resolveSpecCouncil(gauntlet), errors }
|
|
681
|
-
:
|
|
686
|
+
: params.key === "closureReview"
|
|
687
|
+
? { key: "closureReview" as const, ...resolveClosureReview(gauntlet), errors }
|
|
688
|
+
: {
|
|
689
|
+
key: "escalationLoop" as const,
|
|
690
|
+
...resolveEscalationLoop(gauntlet, mainLoopModel(ctx.model, ctx.thinkingLevel)),
|
|
691
|
+
errors,
|
|
692
|
+
};
|
|
682
693
|
return {
|
|
683
694
|
content: [{ type: "text", text: "```json\n" + JSON.stringify(payload, null, 2) + "\n```" }],
|
|
684
695
|
details: payload,
|
|
@@ -847,6 +858,8 @@ export default function (pi: ExtensionAPI) {
|
|
|
847
858
|
}
|
|
848
859
|
}
|
|
849
860
|
phases = { ...phases, [params.phase]: transitionPhaseState("in_progress") as PhaseState };
|
|
861
|
+
// A rewind must not inherit the prior verify's conformance latch.
|
|
862
|
+
if (params.phase === "implement") conformanceDispatched = false;
|
|
850
863
|
gauntletEntered = nextGauntletEntered(gauntletEntered, "start", phases.brainstorm.status);
|
|
851
864
|
firedGuards.clear();
|
|
852
865
|
updateWidget(ctx);
|
package/package.json
CHANGED
|
@@ -11,7 +11,7 @@ description: "You MUST use this before any creative work - creating features, bu
|
|
|
11
11
|
|
|
12
12
|
Help turn ideas into fully formed designs and specs through natural collaborative dialogue.
|
|
13
13
|
|
|
14
|
-
Identify the target project → set up an isolated worktree → understand current project context → ask questions one at a time → propose 2-3 approaches with trade-offs → present the design in
|
|
14
|
+
Identify the target project → set up an isolated worktree → understand current project context → ask questions one at a time → propose 2-3 approaches with trade-offs → present the design in two rounds → write spec to disk inside the worktree → user reviews before any implementation.
|
|
15
15
|
|
|
16
16
|
## HARD CONSTRAINT
|
|
17
17
|
|
|
@@ -63,7 +63,7 @@ Work through the items below **in order**. This is your own checklist to follow,
|
|
|
63
63
|
4. **Understand the idea against the draft** — `Read` the draft, verify load-bearing
|
|
64
64
|
claims against real code, ask questions one at a time, append citable findings
|
|
65
65
|
5. **Propose 2-3 approaches** — with trade-offs and a recommendation
|
|
66
|
-
6. **Present the design** — in
|
|
66
|
+
6. **Present the design** — in two rounds, one approval each
|
|
67
67
|
7. **Write the spec** — to `doc/specs/` (see [Filename Convention](#filename-convention)); then mark any known superseded predecessor(s) per [Marking superseded specs](#marking-superseded-specs), at the exact-order position defined in [Spec Self-Review](#spec-self-review-before-user-review-gate)
|
|
68
68
|
8. **Spec self-review (lint)** — placeholder scan + internal consistency + documentation named, run inline
|
|
69
69
|
9. **Critique pass (auto-dispatched)** — scope + ambiguity; the spec council via `/skill:roasting-the-spec` when `gauntlet_setting` returns verdict `council` (it applies its apply-set, including any external-ref inlining, to the spec before returning — see [Spec Council](#spec-council-optional)), else a fresh `worker` that applies its own fixes in place
|
|
@@ -143,9 +143,7 @@ path.
|
|
|
143
143
|
section starts that answer; confirm it before designing from scratch.
|
|
144
144
|
- Ask questions **one at a time** to refine the idea. Prefer multiple-choice; one
|
|
145
145
|
question per message. Focus on: purpose, constraints, success criteria, who/what
|
|
146
|
-
it touches.
|
|
147
|
-
so the supersession event is captured before spec-writing (see
|
|
148
|
-
[Marking superseded specs](#marking-superseded-specs)).
|
|
146
|
+
it touches.
|
|
149
147
|
- **Append bar:** append to the draft's `## Appended during questionary` only
|
|
150
148
|
findings the spec will cite — schema shapes, hard constraints, ticket-vs-code
|
|
151
149
|
contradictions, user answers that changed scope. Not a log of every grep.
|
|
@@ -168,18 +166,14 @@ When sketching the design, prefer:
|
|
|
168
166
|
- **Single source of truth** — point at the schema/contract that owns the data (the migration, type definition, or API contract that defines it); don't invent parallel state.
|
|
169
167
|
- **Explicit error and edge cases** — name them. "Out of scope" is a valid answer, but it has to be stated.
|
|
170
168
|
|
|
171
|
-
### 6. Present the design in
|
|
169
|
+
### 6. Present the design in two rounds
|
|
172
170
|
|
|
173
|
-
|
|
171
|
+
Two rounds, one approval each. Target 300-500 words per round. A revisit after feedback stays inside the same approval point.
|
|
174
172
|
|
|
175
|
-
|
|
173
|
+
- Round 1: architecture overview, components / responsibilities, data flow, and `supersedes <path>, <scope>` when the draft names a predecessor. Ask once. Approval without correction confirms the predecessor.
|
|
174
|
+
- Round 2: error handling and edge cases, testing approach, `## Documentation impact`. Ask once.
|
|
176
175
|
|
|
177
|
-
-
|
|
178
|
-
- Components / responsibilities
|
|
179
|
-
- Data flow (or request flow)
|
|
180
|
-
- Error handling and edge cases
|
|
181
|
-
- Testing approach
|
|
182
|
-
- Documentation impact — a required `## Documentation impact` section. Cite the materiality bar in `reference/documentation-impact.md` by relative path rather than restating its categories, and reproduce its template block verbatim:
|
|
176
|
+
The `## Documentation impact` section is required. Cite the materiality bar in `reference/documentation-impact.md` by relative path rather than restating its categories, and reproduce its template block verbatim:
|
|
183
177
|
|
|
184
178
|
```markdown
|
|
185
179
|
## Documentation impact
|
|
@@ -198,18 +192,7 @@ When a ticket ID is given, fetch the ticket and treat it as **guidance, not the
|
|
|
198
192
|
|
|
199
193
|
## First-Feature Oversight (Early Project Stages)
|
|
200
194
|
|
|
201
|
-
For the **first two features** of a new initiative
|
|
202
|
-
|
|
203
|
-
- Directory and module structure decisions
|
|
204
|
-
- Naming conventions (public types, files, routes, identifiers)
|
|
205
|
-
- New shared abstraction (location, responsibility, boundary)
|
|
206
|
-
- Persistence/schema design (entity names, field types, indexing)
|
|
207
|
-
- Proposed additions to AGENTS.md or doc/ files
|
|
208
|
-
|
|
209
|
-
If the developer hasn't provided guidance, ask explicitly:
|
|
210
|
-
> "This is one of the first features in this initiative. Before I proceed, I need your confirmation on: [list specific decisions]."
|
|
211
|
-
|
|
212
|
-
After the first two features establish patterns, follow those patterns without gating.
|
|
195
|
+
For the **first two features** of a new initiative (a new top-level module/package, long-lived component, persistence/schema area, or any pattern that will repeat), round 1 lists these decisions explicitly so the user can correct them there: directory and module structure; naming conventions (public types, files, routes, identifiers); new shared abstractions (location, responsibility, boundary); persistence/schema design (entity names, field types, indexing); proposed additions to AGENTS.md or doc/ files. No separate confirmation. After the first two features establish patterns, follow them.
|
|
213
196
|
|
|
214
197
|
## Anti-Pattern: "Too simple to need a design"
|
|
215
198
|
|
|
@@ -237,7 +220,7 @@ spec-writing: write the spec at the new path **and delete the old draft file**
|
|
|
237
220
|
|
|
238
221
|
## Marking superseded specs
|
|
239
222
|
|
|
240
|
-
When the new spec replaces a prior spec — fully or in part —
|
|
223
|
+
When the new spec replaces a prior spec — fully or in part — (from the draft's scout recon or the request), mark the predecessor. No mechanical sweep: grep or path-overlap hits never decide supersession.
|
|
241
224
|
|
|
242
225
|
- `edit` the predecessor spec (in the project's spec directory, per [Project Routing](#project-routing)) to insert, after its title line and a blank line, one banner line per successor:
|
|
243
226
|
|
|
@@ -357,12 +340,26 @@ If you believe the summary needs correcting, do **not** silently rewrite it —
|
|
|
357
340
|
|
|
358
341
|
Wait for the user. On a change request (including a revert), revise the spec and re-present — mint a **fresh** temp path for the re-dispatched summarizer (never reuse a prior round's path, so stale content can never be mistaken for the new summary). On approval, proceed immediately to `/skill:writing-plans` with no further prompt — the plan and execution mode are mechanical derivatives, so the only human gate here is spec approval itself. Don't land the spec on `main`; it stays in the worktree and ships in the same squash commit as the implementation.
|
|
359
342
|
|
|
343
|
+
Post-approval changes follow [Amending an approved spec](#amending-an-approved-spec).
|
|
344
|
+
|
|
360
345
|
After approval, mark the brainstorm phase complete:
|
|
361
346
|
|
|
362
347
|
```
|
|
363
348
|
phase_tracker({ action: "complete", phase: "brainstorm" })
|
|
364
349
|
```
|
|
365
350
|
|
|
351
|
+
## Amending an approved spec
|
|
352
|
+
|
|
353
|
+
Execute this section in place from any later phase. Do not invoke `/skill:brainstorming` (its entry resets both trackers). Worktree, spec commits, and plan survive.
|
|
354
|
+
|
|
355
|
+
1. Edit the spec. Show `git --no-pager diff -- <spec path>` and one line of impact (affected plan tasks / waves, or "no plan yet").
|
|
356
|
+
2. Wait for approval. Change request -> revise, re-show.
|
|
357
|
+
3. No plan yet -> commit the spec; continue. Plan exists -> update affected anchors and tasks: `plan_tracker` `add` for new tasks, `update` for changed ones; anchor-changed completed tasks go back to `pending` and re-run the task loop. A removed task is deleted from the plan; then re-`init` the tracker with the remaining tasks in wave order and `update` every already-completed task back to `complete` (the only permitted `init` after handoff; never `clear`). Re-run `plan_check` until it passes, commit spec + plan together; continue. A task reopened while `verify` or `ship` is in progress: `phase_tracker({ action: "skip", phase: "<current>", reason: "amendment reopened Task N" })`, then `phase_tracker({ action: "start", phase: "implement", force: true })`; later phases re-enter with `force: true` and rerun in full.
|
|
358
|
+
|
|
359
|
+
Redraw test: the diff changes the problem statement, adds or removes a component, or moves a component boundary -> redraw. A change inside one component (a persistence mechanism, a worker's HTTP client, dropping a fallback and its task) -> amend. State the call in the same message as the diff; the user overrides either way.
|
|
360
|
+
|
|
361
|
+
Redraw: keep the worktree and the approved spec file. `plan_tracker({ action: "clear" })`, `phase_tracker({ action: "reset" })`, `phase_tracker({ action: "start", phase: "brainstorm" })`, delete the plan file, resume at checklist step 4 with the approved spec as the draft (steps 2-3 skipped). Spec-writing overwrites it; the full gate follows.
|
|
362
|
+
|
|
366
363
|
## Key Principles
|
|
367
364
|
|
|
368
365
|
- **One question at a time.**
|
|
@@ -370,30 +367,21 @@ phase_tracker({ action: "complete", phase: "brainstorm" })
|
|
|
370
367
|
- **YAGNI ruthlessly.**
|
|
371
368
|
- **Design for testability** — clear boundaries enable TDD.
|
|
372
369
|
- **Explore 2-3 approaches** before settling.
|
|
373
|
-
- **
|
|
370
|
+
- **Two design rounds** — one approval per round.
|
|
374
371
|
- **Be flexible** — go back and clarify when something doesn't make sense.
|
|
375
372
|
|
|
376
373
|
## Red Flags — STOP
|
|
377
374
|
|
|
378
|
-
-
|
|
379
|
-
-
|
|
380
|
-
-
|
|
381
|
-
-
|
|
382
|
-
-
|
|
383
|
-
-
|
|
384
|
-
-
|
|
385
|
-
-
|
|
386
|
-
-
|
|
387
|
-
-
|
|
388
|
-
- About to compose the gate message when the summary `Read` was not the last content-producing tool call before it (a following `rm` of the temp file is fine) — a turn boundary between the `Read` and the render lets pi-condense prune the ~9KB read result, reproducing the original bug
|
|
389
|
-
- About to present a paraphrased, condensed, or re-sectioned version of the summarizer's output instead of pasting its returned text verbatim — rewriting the summary counts as not rendering it
|
|
390
|
-
- About to run the scope or ambiguity checks inline yourself instead of dispatching them (those two are the critique pass, not the inline lint)
|
|
391
|
-
- About to skip the self-review pass
|
|
392
|
-
- About to proceed to `/skill:writing-plans` before the user has approved the spec (proceeding *after* approval is correct; skipping the gate is the violation)
|
|
393
|
-
- About to finish spec-writing for a replacement design without marking the known predecessor (see [Marking superseded specs](#marking-superseded-specs))
|
|
394
|
-
- Spec contains `TODO`, `TBD`, or unnamed components
|
|
395
|
-
- About to offer a multi-spec split that fails the split test in `../shape-ticket/reference/split-axes.md`, or without its three-line justification per spec
|
|
396
|
-
- User said "this is just a small change" and you accepted it without applying the [Anti-Pattern](#anti-pattern-too-simple-to-need-a-design) check
|
|
375
|
+
- Writing or editing anything outside `doc/specs/` while this skill is active.
|
|
376
|
+
- Overwriting the draft without reading it in full in the same turn, or using `edit` for the spec-writing overwrite.
|
|
377
|
+
- Dispatching lint, critique, council, or summarizer while the spec file's line 1 is the context-draft marker.
|
|
378
|
+
- Running the scope or ambiguity checks inline instead of dispatching the critique pass.
|
|
379
|
+
- Reaching the gate after a failed or skipped critique pass, or without re-running the placeholder scan on the applied spec.
|
|
380
|
+
- Composing the gate without the summary `Read` as the last content-producing call, or paraphrasing the summary instead of pasting it verbatim.
|
|
381
|
+
- Inserting a human stop between gather dispatch and questionary question one.
|
|
382
|
+
- Running, deploying, or validating the proposed change before approval.
|
|
383
|
+
- Proceeding to `/skill:writing-plans` before the user approves the spec, or invoking `/skill:brainstorming` to amend an approved spec.
|
|
384
|
+
- Writing a replacement spec without the known predecessor's banner, or offering a multi-spec split that fails `../shape-ticket/reference/split-axes.md`.
|
|
397
385
|
|
|
398
386
|
## Project overrides
|
|
399
387
|
|
|
@@ -46,7 +46,12 @@ Scout (always dispatched):
|
|
|
46
46
|
> exact paths and line ranges. If a spec you cite carries a supersession marker
|
|
47
47
|
> (default: a `> **Superseded by:**` banner; the project's overrides may define
|
|
48
48
|
> another format), follow the successor for the superseded scope and cite it
|
|
49
|
-
> instead; cite the old spec only for its unsuperseded sections.
|
|
49
|
+
> instead; cite the old spec only for its unsuperseded sections. Predecessor
|
|
50
|
+
> check: list the project's spec directory, read titles and `**Goal:**` lines,
|
|
51
|
+
> open at most five whose topic matches this request, and name any whose design
|
|
52
|
+
> this request replaces or amends with the section(s) affected -
|
|
53
|
+
> `Predecessor: <path>, <scope>` - or `Predecessor: none`. Judge by topic; shared
|
|
54
|
+
> file paths never decide. End with an
|
|
50
55
|
> "Open questions that matter for the spec"
|
|
51
56
|
> section. Compact handoff, not a dump.
|
|
52
57
|
|
|
@@ -127,7 +127,7 @@ Three tiers, increasing cost — name the tier when a revert is requested:
|
|
|
127
127
|
|---|---|---|---|
|
|
128
128
|
| Cheap | Council edit, reverted at the `brainstorming` gate | Spec isn't yet plan- or code-bearing | Revise spec, re-present |
|
|
129
129
|
| Light | Conformance fix, reverted at finish | Gap re-opens for a fresh disposition | Revert the `conformance fix Gn` commit(s), re-audit |
|
|
130
|
-
| Heavy | Council edit, reverted at finish | Rewrites the already-ratified contract that drove the plan and code | Amend spec
|
|
130
|
+
| Heavy | Council edit, reverted at finish | Rewrites the already-ratified contract that drove the plan and code | Amend spec per brainstorming's [Amending an approved spec](../brainstorming/SKILL.md#amending-an-approved-spec) → regenerate affected plan/code → re-run verify before ship |
|
|
131
131
|
|
|
132
132
|
A **heavy** revert is not a menu toggle — say so explicitly to the user before proceeding, and do not present it as equivalent-effort to the light tier. The council audit that lets the human identify revert candidates lives in the `brainstorming` spec commit message body (not a committed spec section).
|
|
133
133
|
|
|
@@ -30,7 +30,7 @@ You are the **orchestrator**. You read the plan, dispatch, review the review, de
|
|
|
30
30
|
**Do not pause to check in with the user between tasks.** The plan is already approved. Pause only when:
|
|
31
31
|
|
|
32
32
|
- A subagent returns `NEEDS_CONTEXT` or `BLOCKED` (see [Implementer Status](#implementer-status))
|
|
33
|
-
-
|
|
33
|
+
- An escalated round fails (stop note per [Fix-Loop Rounds](#fix-loop-rounds))
|
|
34
34
|
- A ⚠️ workflow warning fires
|
|
35
35
|
|
|
36
36
|
Reaching the end of the plan is not a pause: continue through verification and invoke `/skill:finishing-a-development-branch` as defined in [After All Tasks](#after-all-tasks-complete).
|
|
@@ -63,7 +63,7 @@ For each task in `plan_tracker`:
|
|
|
63
63
|
6. If quality reviewer finds issues → re-dispatch implementer → re-review. Loop until ✅, within [Fix-Loop Rounds](#fix-loop-rounds).
|
|
64
64
|
7. After its existing reviews accept the work (and its required commit point), mark that same task index `complete` in `plan_tracker`. A subagent exit or green tests alone are not acceptance.
|
|
65
65
|
|
|
66
|
-
The
|
|
66
|
+
The orchestrator is the spec's only writer during execution. Amendment trigger: implementer `BLOCKED` citing a spec defect, or a review finding showing the spec (not the code) is wrong -> pause the fix loop, execute brainstorming's [Amending an approved spec](../brainstorming/SKILL.md#amending-an-approved-spec) in place, resume. Code-vs-spec mismatch stays in the SR loop. An SR unable to read the spec at a cited anchor (missing file, unresolvable heading/range) returns a blocking finding — the contract is spec+task or stop, never a silent fallback to task-only review.
|
|
67
67
|
|
|
68
68
|
After all tasks: proceed to [After All Tasks](#after-all-tasks-complete) - it owns parent full verification, then whole-diff code review, then conformance.
|
|
69
69
|
|
|
@@ -88,12 +88,12 @@ Every fix re-dispatch (implementer) and code-review re-review carries the consum
|
|
|
88
88
|
|
|
89
89
|
Every dispatched fix is verified by a re-review before escalation or task progression - the loop only ever exits on a clean review or an escalation.
|
|
90
90
|
|
|
91
|
-
**
|
|
91
|
+
**Escalate** = one escalated round; stop only on failure. Call `gauntlet_setting({ key: "escalationLoop" })` (unavailable -> stop and report). `implModel` undefined -> stop note. Otherwise re-dispatch that fix - same payload and isolation knobs (`cwd`, `worktree`, `SCOPED_TEST_COMMANDS`, status protocol, prior patch, spec anchors) plus the prior report verbatim - overriding only `model: <implModel>`, `context: "fresh"`, `async: false`; one implementer, no fan-out. Run the normal fix-round review gate (SR then CR on `Behaviour-change: yes`, else the triggering reviewer with the re-review marker). All clean -> proceed; in wave mode the escalated patch supersedes the prior one at integrate. Any review with issues, a non-`DONE` status, or a dispatch error -> stop note per `stop-note.md`; no second dispatch. Once per loop; independent of the convergence exception. No `plan_tracker` write during escalation - the task stays `in_progress` until the human decides.
|
|
92
92
|
|
|
93
93
|
**Worked examples:**
|
|
94
94
|
|
|
95
95
|
- Review-2 verdict `TRAJECTORY: STAGNANT (repeat of: unchecked error path in parser)` -> escalate now, before fix 2 - earlier than the ordinary budget.
|
|
96
|
-
- Review-3 verdict `TRAJECTORY: CONVERGING (3 -> 1, max severity Moderate)` -> dispatch fix 3; if review 4 still finds issues, escalate
|
|
96
|
+
- Review-3 verdict `TRAJECTORY: CONVERGING (3 -> 1, max severity Moderate)` -> dispatch fix 3; if review 4 still finds issues, escalate.
|
|
97
97
|
- Review-3 verdict `TRAJECTORY: CONVERGING (3 -> 2, max severity Critical)` or `TRAJECTORY: DIVERGING` or no `TRAJECTORY:` line -> escalate.
|
|
98
98
|
|
|
99
99
|
## Implementer Status
|
|
@@ -139,6 +139,9 @@ When in doubt, default. Don't downgrade reviewers — false negatives are expens
|
|
|
139
139
|
// implementer
|
|
140
140
|
subagent({ agent: "implementer", async: false, task: "<task text + context + SCOPED_TEST_COMMANDS + status protocol>" })
|
|
141
141
|
|
|
142
|
+
// escalated fix round (Fix-Loop Rounds): same fix payload, model from gauntlet_setting({ key: "escalationLoop" }).implModel
|
|
143
|
+
subagent({ agent: "implementer", model: "<implModel>", context: "fresh", async: false, task: "<the just-dispatched fix payload + prior review report verbatim>" })
|
|
144
|
+
|
|
142
145
|
// spec compliance
|
|
143
146
|
subagent({ agent: "spec-reviewer", async: false, task: "<task text + patch diff + absolute spec path + task's Spec: anchors>" })
|
|
144
147
|
|
|
@@ -237,27 +240,16 @@ For the fan-out + worktree + patch-integration + conflict mechanics, see `dispat
|
|
|
237
240
|
|
|
238
241
|
## Red Flags — STOP
|
|
239
242
|
|
|
240
|
-
- Writing code yourself instead of dispatching
|
|
241
|
-
-
|
|
242
|
-
-
|
|
243
|
-
- Making a subagent read the plan
|
|
244
|
-
-
|
|
245
|
-
- Moving to next task with either review still showing issues
|
|
246
|
-
- Dispatching fix 3 without a reviewer-emitted
|
|
247
|
-
-
|
|
248
|
-
-
|
|
249
|
-
-
|
|
250
|
-
- Pausing to "check in" between tasks (continuous execution rule)
|
|
251
|
-
- Skipping the `Implementer Status` parse — treating every response as DONE
|
|
252
|
-
- Starting on main without explicit user consent
|
|
253
|
-
- Dispatching `code-reviewer` before every one of the wave's spec-review verdicts has landed (including fusing SR+CR into one parallel call)
|
|
254
|
-
- Dispatching fixes sequentially on a clean HEAD despite a certified (probe-passing, per dispatching-parallel-agents § Fix fan-out) ≥ 2-ID `disjoint` group in the review's `Parallel-safe:` line
|
|
255
|
-
- Dispatching `code-reviewer` per task inside a wave (CR binds to the integrated wave diff)
|
|
256
|
-
- Dispatching an implementer or code-reviewer without a `SCOPED_TEST_COMMANDS` value (commands or `none`)
|
|
257
|
-
- About to run the full verification entrypoint during the implement phase — task and wave gates run scoped, plan-declared commands only; the full set belongs to verify
|
|
258
|
-
- Dispatching a verification or review repair before reopening (`in_progress`) the plan-task indices that own its touched files, or completing them on the passing rerun instead of on the accepting whole-diff review
|
|
259
|
-
- Dispatching whole-diff CR before parent full verification passes, or conformance before the foreground CR result and any invalidating repair re-verification/re-review are accepted
|
|
260
|
-
- Polling, joining, or relaunching an unexpectedly asynchronous gauntlet dispatch instead of stopping and reporting
|
|
243
|
+
- Writing code yourself instead of dispatching.
|
|
244
|
+
- Pausing between tasks for anything other than `NEEDS_CONTEXT`, `BLOCKED`, a fix-loop escalation, a workflow warning, or a spec amendment.
|
|
245
|
+
- Dispatching parallel implementers on overlapping files, on a shared mutable runtime resource, or without `worktree: true`.
|
|
246
|
+
- Making a subagent read the plan, inlining spec excerpts to the spec reviewer, or dispatching without a `SCOPED_TEST_COMMANDS` value.
|
|
247
|
+
- Dispatching `code-reviewer` before every in-scope spec-review verdict is ✅, or per task inside a wave.
|
|
248
|
+
- Moving to the next task with either review still showing issues, or skipping the `Implementer Status` parse.
|
|
249
|
+
- Dispatching fix 3 without a reviewer-emitted `CONVERGING` verdict, or continuing past `STAGNANT` instead of escalating.
|
|
250
|
+
- Running the full verification entrypoint during the implement phase.
|
|
251
|
+
- Dispatching a repair before reopening the plan-task indices that own its files, whole-diff CR before parent verification passes, or conformance before the CR result is accepted.
|
|
252
|
+
- Polling, joining, or relaunching an unexpectedly asynchronous dispatch, or starting on main without explicit user consent.
|
|
261
253
|
|
|
262
254
|
## Integration
|
|
263
255
|
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
# Stop note (subagent-driven-development companion)
|
|
2
|
+
|
|
3
|
+
Emitted when no escalation model resolves, the escalated dispatch errors, or its one escalated round fails (see `SKILL.md` "Fix-Loop Rounds"); an unavailable `gauntlet_setting` tool is a configuration error - stop and report, no stop note. Inline in the reply; the turn ends; phase stays `implement`, the task stays `in_progress`; no further tasks start. Wave mode: one note per stalled task, after the current batch returns.
|
|
4
|
+
|
|
5
|
+
## Template
|
|
6
|
+
|
|
7
|
+
```
|
|
8
|
+
Stopped on task <n> (<title>)[; escalated round on <implModel> did not resolve it].
|
|
9
|
+
|
|
10
|
+
Problem: <one sentence: what is wrong and why the fixes could not resolve it>
|
|
11
|
+
<file:line> - <quoted finding from the final review>
|
|
12
|
+
<failing test/command + 1-3 line output snippet, when present>
|
|
13
|
+
|
|
14
|
+
Fix options:
|
|
15
|
+
a) <concrete change>
|
|
16
|
+
b) <concrete change - amending spec section X / plan task n is a normal option>
|
|
17
|
+
c) <optional third>
|
|
18
|
+
|
|
19
|
+
Pick one, or give another fix.
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
## Rules
|
|
23
|
+
|
|
24
|
+
- Bracketed header clause only when an escalated round actually ran; omit it when no model was resolvable or the dispatch errored.
|
|
25
|
+
- Residual issues from the final review report only; quote, do not summarise history.
|
|
26
|
+
- When no escalated round ran, the Problem is the reason: "no escalation model resolvable", or the implementer's non-DONE status text / the dispatch error, quoted; skip the file:line and test lines.
|
|
27
|
+
- Options are actionable edits. Spec/plan amendment is first-class - stalls are usually a slightly contradictory spec, not a capability gap.
|
|
28
|
+
- Never offer "skip the task". If the task is genuinely droppable, say so and name the plan tasks that depend on it.
|
|
29
|
+
- No trajectory verdicts, round history, review counts, or paths to spec/plan/review reports.
|
|
30
|
+
- Plain words, ASCII, no headings. One screen.
|
|
31
|
+
|
|
32
|
+
## Example
|
|
33
|
+
|
|
34
|
+
```
|
|
35
|
+
Stopped on task 4 (retry policy for the outbound client); escalated round on <provider>/<model>:high did not resolve it.
|
|
36
|
+
|
|
37
|
+
Problem: `RetryPolicy.next()` returns 0 ms for the first retry, but the client treats 0 as "no
|
|
38
|
+
retry", so the first failure is never retried.
|
|
39
|
+
src/net/retry.ts:41 - `return attempt * this.baseMs;`
|
|
40
|
+
npm test -- retry > "retries once after a transient failure":
|
|
41
|
+
expected 1 call after failure, got 0
|
|
42
|
+
|
|
43
|
+
Fix options:
|
|
44
|
+
a) start the backoff at `baseMs` (`(attempt + 1) * this.baseMs`)
|
|
45
|
+
b) amend plan task 4 so the client retries on any non-negative delay, and keep the policy as is
|
|
46
|
+
|
|
47
|
+
Pick one, or give another fix.
|
|
48
|
+
```
|
|
@@ -15,7 +15,7 @@ If the repo file defines a `piGauntlet.<key>` at all, that definition **replaces
|
|
|
15
15
|
the preset's for that key entirely - the two are never merged leaf-by-leaf. If the
|
|
16
16
|
repo file does not define the key, the preset's value is used unchanged. This is
|
|
17
17
|
exactly pi's own `deepMergeSettings` behaviour: it spreads the second-level keys
|
|
18
|
-
(`specCouncil`, `closureReview`, `flowGuards`, `verifyBeforeShip`) wholesale, and
|
|
18
|
+
(`specCouncil`, `closureReview`, `flowGuards`, `verifyBeforeShip`, `escalationLoop`) wholesale, and
|
|
19
19
|
does not recurse into their leaves.
|
|
20
20
|
|
|
21
21
|
**Caveat - partial definitions drop siblings.** Because the replace is
|
|
@@ -66,6 +66,7 @@ Then continue with the normal flow below (Scope Check onward, including Recon).
|
|
|
66
66
|
- Edit or create any other files: no
|
|
67
67
|
- Write implementation code: never inside this skill. After Self-Review + `phase_tracker` complete, auto-invoke `/skill:subagent-driven-development` to execute.
|
|
68
68
|
- Land the plan on `main`: no — the plan commit goes on the worktree branch (same branch as the spec)
|
|
69
|
+
- Edit the approved spec: only via brainstorming's [Amending an approved spec](../brainstorming/SKILL.md#amending-an-approved-spec)
|
|
69
70
|
|
|
70
71
|
## Scope Check
|
|
71
72
|
|
|
@@ -74,9 +75,7 @@ Before writing the plan, check the spec one more time:
|
|
|
74
75
|
- Does an intermediate state need to be **independently deployable**, under a deploy topology documented in the gauntlet overrides file's `## Deployment` section? Fail closed: undocumented or monolithic topology -> no deployment-driven split.
|
|
75
76
|
- Is there a **review-risk isolation** reason to land part separately (e.g. a large mechanical rename apart from the behavior change that motivated it)?
|
|
76
77
|
|
|
77
|
-
If yes, decompose into separate plans and
|
|
78
|
-
|
|
79
|
-
> "The spec covers A and B. I'd split into two plans, executed in order. OK?"
|
|
78
|
+
If yes, decompose into separate plans, executed in order, and state the split and its reason in the handoff message. No approval prompt.
|
|
80
79
|
|
|
81
80
|
Otherwise one plan. Service, contract, or schema count is not a split signal - one concern routinely spans several. The concern test itself lives in `../shape-ticket/reference/split-axes.md` (resolve the path against this skill's own directory) and was applied upstream at spec time; plans do not re-litigate it. A single plan should land in one PR worth of work.
|
|
82
81
|
|
|
@@ -121,7 +120,7 @@ List the files this implementation will create, modify, or delete. Group by comp
|
|
|
121
120
|
- `src/services/legacy_foo.ts`
|
|
122
121
|
```
|
|
123
122
|
|
|
124
|
-
If you can't list the files, the spec isn't ready
|
|
123
|
+
If you can't list the files, the spec isn't ready: amend or redraw per brainstorming's [Amending an approved spec](../brainstorming/SKILL.md#amending-an-approved-spec).
|
|
125
124
|
|
|
126
125
|
## Wave Grouping
|
|
127
126
|
|
|
@@ -251,7 +250,7 @@ Each task uses `- [ ]` checkbox steps so execution tools (and humans) can track
|
|
|
251
250
|
|
|
252
251
|
Every code task carries this step (red -> green -> fmt/lint -> commit). Doc-only tasks omit it unless the project formats Markdown.
|
|
253
252
|
|
|
254
|
-
**Anchor rules.** The task's `**Spec:**` line cites the plan header's spec path; multiple anchors sit comma-separated on one line (`§ "A" L10-L18, § "C" L40-L44`). Checks key on the `§` marker, so the header's path-only `**Spec:**` line is never matched. Anchors are captured
|
|
253
|
+
**Anchor rules.** The task's `**Spec:**` line cites the plan header's spec path; multiple anchors sit comma-separated on one line (`§ "A" L10-L18, § "C" L40-L44`). Checks key on the `§` marker, so the header's path-only `**Spec:**` line is never matched. Anchors are captured against the gated spec at plan-writing time; a change to the approved spec follows brainstorming's [Amending an approved spec](../brainstorming/SKILL.md#amending-an-approved-spec), executed in place. A task with no anchorable requirement (pure-mechanics chore) omits the `**Spec:**` line entirely (never `**Spec:** none`) and carries a mechanical-task row in `## Spec coverage` — silence is never valid.
|
|
255
254
|
|
|
256
255
|
## Spec Coverage Table
|
|
257
256
|
|