pi-plans 0.6.0 → 0.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +58 -0
- package/CONTRIBUTING.md +8 -15
- package/README.md +39 -37
- package/agents/execution-reviewer.md +40 -0
- package/agents/reviewer.md +12 -3
- package/index.ts +55 -58
- package/package.json +2 -1
- package/references/pi-planning-workflow.md +50 -60
- package/references/plan-artifact-template.md +81 -60
- package/references/state-and-config.md +60 -44
- package/scripts/bench/pi-adapter/pi_plans_bench.py +33 -22
- package/scripts/bench/pi-adapter/rpc_driver.mjs +4 -4
- package/scripts/run-tests.ts +12 -1
- package/scripts/validate.ts +39 -11
- package/skills/debug-and-plan/SKILL.md +3 -3
- package/skills/plan-big/SKILL.md +3 -3
- package/skills/plan-normal/SKILL.md +3 -3
- package/skills/plan-small/SKILL.md +4 -4
- package/skills/plan-with-refs/SKILL.md +6 -6
- package/skills/planning/SKILL.md +1 -1
- package/src/ask-form.ts +4 -4
- package/src/auditor.ts +227 -0
- package/src/auto-approve.ts +1 -1
- package/src/autocomplete.ts +19 -17
- package/src/code-graph/commands.ts +8 -3
- package/src/code-graph/community.ts +1 -1
- package/src/code-graph/paths.ts +1 -1
- package/src/code-graph/watch.ts +2 -2
- package/src/compaction.ts +3 -3
- package/src/config-command.ts +146 -73
- package/src/dashboard.ts +303 -0
- package/src/exec.ts +1185 -924
- package/src/global-state.ts +304 -0
- package/src/guard.ts +18 -19
- package/src/messaging.ts +44 -0
- package/src/plan.ts +421 -112
- package/src/query-hook.ts +4 -4
- package/src/refine-prompts.ts +12 -70
- package/src/refine-ui-helpers.ts +24 -5
- package/src/refine-ui-state.ts +1 -1
- package/src/refine-ui.ts +19 -3
- package/src/resume-command.ts +45 -129
- package/src/resume.ts +5 -1
- package/src/role-panels.ts +542 -0
- package/src/run-context.ts +3 -10
- package/src/staleness.ts +53 -0
- package/src/state.ts +273 -72
- package/src/subagent.ts +19 -29
- package/src/task-tool.ts +100 -0
- package/src/tasks.ts +223 -0
- package/src/thinking-levels.ts +67 -0
- package/src/ui-language.ts +7 -54
- package/src/workflow-state.ts +76 -58
- package/tests/analyze-refs.test.ts +35 -18
- package/tests/ask-choice-schema.test.ts +0 -12
- package/tests/ask-choice.test.ts +2 -49
- package/tests/ask-form-tool.test.ts +4 -5
- package/tests/ask-form.test.ts +2 -2
- package/tests/auditor.test.ts +210 -0
- package/tests/auto-approve.test.ts +7 -10
- package/tests/autocomplete.test.ts +8 -11
- package/tests/code-graph-apply-action.test.ts +2 -2
- package/tests/code-graph-commands.test.ts +2 -2
- package/tests/code-graph-index.test.ts +2 -2
- package/tests/code-graph-loop.e2e.test.ts +1 -1
- package/tests/code-graph-mutations.test.ts +1 -1
- package/tests/code-graph-rollback.test.ts +1 -1
- package/tests/code-graph-v05.test.ts +2 -2
- package/tests/compaction.test.ts +1 -1
- package/tests/config-command.test.ts +103 -100
- package/tests/dashboard.test.ts +402 -0
- package/tests/exec-lifecycle.test.ts +181 -115
- package/tests/exec-panel-lifecycle.test.ts +106 -251
- package/tests/exec-review-loop.test.ts +331 -0
- package/tests/exec.test.ts +771 -1706
- package/tests/execute-plan.test.ts +44 -19
- package/tests/extension-load.test.ts +48 -0
- package/tests/global-state.test.ts +371 -0
- package/tests/graph-aware-file-tools.test.ts +5 -5
- package/tests/guard.test.ts +1 -1
- package/tests/multi-run.test.ts +3 -103
- package/tests/plan.test.ts +139 -62
- package/tests/plans.test.ts +7 -79
- package/tests/refine-prompts.test.ts +20 -71
- package/tests/refine-resume.test.ts +27 -22
- package/tests/refine-ui.test.ts +6 -15
- package/tests/resume-lifecycle.test.ts +41 -22
- package/tests/resume.test.ts +39 -81
- package/tests/role-panels.test.ts +391 -0
- package/tests/run-context.test.ts +1 -1
- package/tests/run-ownership.test.ts +1 -1
- package/tests/stale-ctx.test.ts +218 -0
- package/tests/staleness.test.ts +76 -0
- package/tests/state.test.ts +155 -32
- package/tests/subagent-thinking.test.ts +65 -0
- package/tests/subagent-usage.test.ts +1 -1
- package/tests/task-tool.test.ts +61 -0
- package/tests/tasks.test.ts +142 -0
- package/tests/thinking-levels.test.ts +77 -0
- package/tests/ui-language.test.ts +2 -17
- package/tests/workflow-state.test.ts +73 -90
- package/tools/analyze-refs.ts +67 -32
- package/tools/ask-choice.ts +7 -53
- package/tools/code-graph.ts +2 -2
- package/tools/execute-plan.ts +55 -99
- package/tools/graph-aware-file-tools.ts +4 -10
- package/tools/plans.ts +41 -67
- package/tools/refine.ts +101 -164
- package/agents/criticizer.md +0 -18
- package/agents/executor.md +0 -26
- package/scripts/bench/pi-adapter/__pycache__/pi_plans_bench.cpython-312.pyc +0 -0
- package/src/panel.ts +0 -473
- package/src/termination-prompt.ts +0 -73
- package/tests/goal-wait.test.ts +0 -269
- package/tests/panel-i-zero.test.ts +0 -420
- package/tests/panel.test.ts +0 -355
package/index.ts
CHANGED
|
@@ -1,16 +1,16 @@
|
|
|
1
1
|
/**
|
|
2
2
|
* pi-plans — human-in-the-loop planning for the Pi coding agent.
|
|
3
3
|
*
|
|
4
|
-
* Researched, refined Markdown plans before any code changes. The
|
|
4
|
+
* Researched, refined Markdown plans before any code changes. The six
|
|
5
5
|
* planning skills are contributed via resources_discover; the extension
|
|
6
6
|
* provides the supporting machinery:
|
|
7
7
|
*
|
|
8
|
-
* - `plans` tool — workspace state (config, runs, ledgers) in .git/
|
|
8
|
+
* - `plans` tool — workspace state (config, runs, ledgers) in .git/pi-plans/
|
|
9
9
|
* - `ask_choice` tool — the choice-prompt contract (Other / Auto-complete rules)
|
|
10
|
-
* - `refine` tool — reviewer
|
|
10
|
+
* - `refine` tool — reviewer rounds (findings + questions) via read-only pi subagents
|
|
11
11
|
* - `execute_plan` — execution handoff into the tracked execution loop
|
|
12
12
|
* - write guard — planning runs may only write planning artifacts
|
|
13
|
-
* - execution loop —
|
|
13
|
+
* - execution loop — task-tree injection, plans_update_task tracking, dashboard, audit
|
|
14
14
|
*/
|
|
15
15
|
|
|
16
16
|
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
|
|
@@ -44,7 +44,7 @@ import {
|
|
|
44
44
|
requestPlanningCompaction,
|
|
45
45
|
restoreFromSession,
|
|
46
46
|
stopExecution,
|
|
47
|
-
|
|
47
|
+
toggleDashboardExpanded,
|
|
48
48
|
updateStatusWidget,
|
|
49
49
|
shouldTriggerPlanningCompaction,
|
|
50
50
|
} from "./src/exec.ts";
|
|
@@ -82,9 +82,13 @@ import { abandonCandidates, resolveCommandRun } from "./src/run-picker.ts";
|
|
|
82
82
|
import { applyPlanWritten, mutateCheckpoint, planIdentityOf } from "./src/workflow-state.ts";
|
|
83
83
|
import { registerAskChoiceTool } from "./tools/ask-choice.ts";
|
|
84
84
|
import { executeCommand, registerExecutePlanTool } from "./tools/execute-plan.ts";
|
|
85
|
+
import { registerTaskStatusTool } from "./src/task-tool.ts";
|
|
86
|
+
import { flattenTaskViews, taskIsTerminal, taskProgress } from "./src/tasks.ts";
|
|
85
87
|
import { registerPlansTool } from "./tools/plans.ts";
|
|
86
88
|
import { registerRefineTool } from "./tools/refine.ts";
|
|
87
89
|
import { registerAnalyzeRefsTool } from "./tools/analyze-refs.ts";
|
|
90
|
+
import { messaging, setMessagingApi } from "./src/messaging.ts";
|
|
91
|
+
import { stalenessLine } from "./src/staleness.ts";
|
|
88
92
|
|
|
89
93
|
const baseDir = dirname(fileURLToPath(import.meta.url));
|
|
90
94
|
|
|
@@ -93,29 +97,11 @@ const baseDir = dirname(fileURLToPath(import.meta.url));
|
|
|
93
97
|
// (code on disk newer than the loaded copy) is immediately visible.
|
|
94
98
|
const extensionLoadedAt = new Date();
|
|
95
99
|
|
|
100
|
+
// The probe itself lives in src/staleness.ts so the execution reviewer can ask
|
|
101
|
+
// the same question when it cannot read a verdict (see exec.ts) without
|
|
102
|
+
// importing index.ts, which already imports exec.ts.
|
|
96
103
|
function extensionStalenessLine(): string {
|
|
97
|
-
|
|
98
|
-
const dirs = [baseDir, path.join(baseDir, "src"), path.join(baseDir, "tools")];
|
|
99
|
-
const stack: string[] = [...dirs];
|
|
100
|
-
let newest = 0;
|
|
101
|
-
while (stack.length) {
|
|
102
|
-
const dir = stack.pop()!;
|
|
103
|
-
for (const entry of fs.readdirSync(dir, { withFileTypes: true })) {
|
|
104
|
-
const full = path.join(dir, entry.name);
|
|
105
|
-
if (entry.isDirectory()) stack.push(full);
|
|
106
|
-
else if (entry.isFile() && entry.name.endsWith(".ts")) {
|
|
107
|
-
const mtime = fs.statSync(full).mtimeMs;
|
|
108
|
-
if (mtime > newest) newest = mtime;
|
|
109
|
-
}
|
|
110
|
-
}
|
|
111
|
-
}
|
|
112
|
-
if (newest > extensionLoadedAt.getTime() + 2000) {
|
|
113
|
-
return `⚠ extension code on disk is newer than the loaded copy (loaded ${extensionLoadedAt.toISOString()}); run /reload to pick it up`;
|
|
114
|
-
}
|
|
115
|
-
return `Extension loaded: ${extensionLoadedAt.toISOString()} (up to date)`;
|
|
116
|
-
} catch {
|
|
117
|
-
return `Extension loaded: ${extensionLoadedAt.toISOString()}`;
|
|
118
|
-
}
|
|
104
|
+
return stalenessLine(baseDir, extensionLoadedAt);
|
|
119
105
|
}
|
|
120
106
|
|
|
121
107
|
function hasActivePlanningWorkflow(ctx: Parameters<typeof updateStatusWidget>[0]): boolean {
|
|
@@ -123,15 +109,28 @@ function hasActivePlanningWorkflow(ctx: Parameters<typeof updateStatusWidget>[0]
|
|
|
123
109
|
const active = resolveActiveRun(ctx.sessionManager, ctx.cwd);
|
|
124
110
|
if (!active) return false;
|
|
125
111
|
const status = getRun(ctx.cwd, active.run_id)?.status;
|
|
126
|
-
return status === "planning" || status === "accepted" || status === "executing";
|
|
112
|
+
return status === "planning" || status === "accepted" || status === "executing" || status === "verifying";
|
|
127
113
|
}
|
|
128
114
|
|
|
129
115
|
export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
116
|
+
setMessagingApi(pi);
|
|
130
117
|
registerPlansTool(pi);
|
|
131
118
|
registerAskChoiceTool(pi);
|
|
132
119
|
registerRefineTool(pi, baseDir);
|
|
133
120
|
registerAnalyzeRefsTool(pi, baseDir);
|
|
134
121
|
registerExecutePlanTool(pi);
|
|
122
|
+
registerTaskStatusTool(pi);
|
|
123
|
+
pi.registerShortcut("ctrl+shift+t", {
|
|
124
|
+
description: "Expand/collapse the pi-plans task dashboard",
|
|
125
|
+
handler: (ctx) => toggleDashboardExpanded(ctx),
|
|
126
|
+
});
|
|
127
|
+
// v0.8: reopen the in-flight execution-review overlay after ESC (the
|
|
128
|
+
// controller is one-shot; the engine re-seeds a fresh one from its held
|
|
129
|
+
// lane state). Inert when no round is in flight.
|
|
130
|
+
pi.registerShortcut("ctrl+shift+r", {
|
|
131
|
+
description: "Reopen the pi-plans execution-review overlay",
|
|
132
|
+
handler: (ctx) => reopenReviewOverlay(ctx),
|
|
133
|
+
});
|
|
135
134
|
registerQueryInterviewHooks(pi, hasActivePlanningWorkflow);
|
|
136
135
|
registerCodeGraphTool(pi);
|
|
137
136
|
registerGraphAwareFileTools(pi);
|
|
@@ -178,7 +177,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
178
177
|
handler: async (args, ctx) => {
|
|
179
178
|
const invocation = args.trim() ? `/skill:${name} ${args.trim()}` : `/skill:${name}`;
|
|
180
179
|
try {
|
|
181
|
-
|
|
180
|
+
messaging().sendUserMessage(invocation, { expandPromptTemplates: true });
|
|
182
181
|
} catch {
|
|
183
182
|
ctx.ui.notify("Agent is busy; try again once the current turn finishes.", "error");
|
|
184
183
|
}
|
|
@@ -208,7 +207,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
208
207
|
if (active) {
|
|
209
208
|
const latest = latestPlanVersion(active.artifact_dir);
|
|
210
209
|
if (latest && path.resolve(ctx.cwd, rawPath) === path.resolve(ctx.cwd, latest.path)) {
|
|
211
|
-
|
|
210
|
+
messaging().appendEntry(PLANNING_PLAN_WRITTEN_CUSTOM_TYPE, {
|
|
212
211
|
runId: active.run_id,
|
|
213
212
|
planPath: latest.path,
|
|
214
213
|
});
|
|
@@ -254,14 +253,14 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
254
253
|
if (event.toolName !== "plans") return;
|
|
255
254
|
if (!consumePrePlanCompactPending(ctx)) return;
|
|
256
255
|
if (typeof ctx.compact !== "function") {
|
|
257
|
-
sendPrePlanCompactResume(
|
|
256
|
+
sendPrePlanCompactResume(ctx);
|
|
258
257
|
return;
|
|
259
258
|
}
|
|
260
259
|
let resumed = false;
|
|
261
260
|
const resumeOnce = () => {
|
|
262
261
|
if (resumed) return;
|
|
263
262
|
resumed = true;
|
|
264
|
-
sendPrePlanCompactResume(
|
|
263
|
+
sendPrePlanCompactResume(ctx);
|
|
265
264
|
};
|
|
266
265
|
ctx.compact({
|
|
267
266
|
customInstructions: PLANNING_PREPLAN_COMPACT_HINT,
|
|
@@ -313,18 +312,18 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
313
312
|
|
|
314
313
|
pi.on("session_before_compact", async (event, ctx) => {
|
|
315
314
|
noteCompactionStarted(ctx, event.customInstructions);
|
|
316
|
-
const executionResult = await handleExecutionBeforeCompact(
|
|
315
|
+
const executionResult = await handleExecutionBeforeCompact(ctx, event);
|
|
317
316
|
if (executionResult) return executionResult;
|
|
318
|
-
return handlePlanningBeforeCompact(
|
|
317
|
+
return handlePlanningBeforeCompact(ctx, event);
|
|
319
318
|
});
|
|
320
319
|
pi.on("session_compact", async (event, ctx) => {
|
|
321
|
-
await handleExecutionCompact(
|
|
322
|
-
await handlePlanningCompact(
|
|
320
|
+
await handleExecutionCompact(ctx, event);
|
|
321
|
+
await handlePlanningCompact(ctx, event);
|
|
323
322
|
noteCompactionEnded(ctx, event.customInstructions);
|
|
324
323
|
});
|
|
325
324
|
pi.on("session_compact_failed", async (event, ctx) => {
|
|
326
|
-
handleExecutionCompactFailed(
|
|
327
|
-
handlePlanningCompactFailed(
|
|
325
|
+
handleExecutionCompactFailed(ctx, event);
|
|
326
|
+
handlePlanningCompactFailed(ctx, event);
|
|
328
327
|
noteCompactionEnded(ctx, event.customInstructions);
|
|
329
328
|
});
|
|
330
329
|
|
|
@@ -332,7 +331,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
332
331
|
// Execution loop: inject remaining checklist each turn, track markers.
|
|
333
332
|
// -----------------------------------------------------------------------
|
|
334
333
|
pi.on("before_agent_start", async (_event, ctx) => {
|
|
335
|
-
drainExecutionFlush(
|
|
334
|
+
drainExecutionFlush(ctx);
|
|
336
335
|
const content = executionContextMessage(ctx);
|
|
337
336
|
if (!content) {
|
|
338
337
|
if (!getExecution() && shouldTriggerPlanningCompaction(ctx)) {
|
|
@@ -372,7 +371,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
372
371
|
// -----------------------------------------------------------------------
|
|
373
372
|
|
|
374
373
|
pi.registerCommand("init-graph", {
|
|
375
|
-
description: "Index the worktree into .git/
|
|
374
|
+
description: "Index the worktree into .git/pi-plans/code_graph.db. If a graph DB already exists, prompt to rebuild or sync changed paths via /update-graph; `--reindex` and non-interactive runs stay on the rebuild path.",
|
|
376
375
|
handler: async (args, ctx) => {
|
|
377
376
|
await initGraphCommand(args, ctx);
|
|
378
377
|
},
|
|
@@ -460,15 +459,16 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
460
459
|
lines.push(`State: ${resolveStateRootOrNull(ctx.cwd) ?? "(no repo)"}`);
|
|
461
460
|
}
|
|
462
461
|
// One-time active.json deprecation note (v0.6.0 registry migration).
|
|
463
|
-
if (fs.existsSync(path.join(resolveStateRootOrNull(ctx.cwd) ?? ".git/
|
|
462
|
+
if (fs.existsSync(path.join(resolveStateRootOrNull(ctx.cwd) ?? ".git/pi-plans", "active.json"))) {
|
|
464
463
|
lines.push("Note: active.json is deprecated — the run registry now derives from runs/*/run.json; the legacy file is ignored.");
|
|
465
464
|
}
|
|
466
465
|
const execution = getExecution();
|
|
467
466
|
if (execution) {
|
|
468
|
-
const
|
|
469
|
-
|
|
470
|
-
|
|
471
|
-
|
|
467
|
+
const progress = taskProgress(execution.tasks);
|
|
468
|
+
const vcDone = execution.items.filter((item) => item.done).length;
|
|
469
|
+
lines.push(`Execution: ${execution.planPath} — tasks ${progress.done}/${progress.total} · VC ${vcDone}/${execution.items.length}`);
|
|
470
|
+
for (const task of flattenTaskViews(execution.tasks)) {
|
|
471
|
+
lines.push(` ${taskIsTerminal(task) ? (task.status === "skipped" ? "~" : "☑") : "☐"} ${task.id}`);
|
|
472
472
|
}
|
|
473
473
|
}
|
|
474
474
|
lines.push(`Auto-complete: ${autoCompleteStatus(ctx)}`);
|
|
@@ -478,7 +478,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
478
478
|
});
|
|
479
479
|
|
|
480
480
|
pi.registerCommand("config-pi-plans", {
|
|
481
|
-
description: "Re-ask and update pi-plans workspace config: language, artifact root, graph,
|
|
481
|
+
description: "Re-ask and update pi-plans workspace config: language, artifact root, graph, and reviewer defaults",
|
|
482
482
|
handler: async (args, ctx) => {
|
|
483
483
|
await configPiPlansCommand(args, ctx);
|
|
484
484
|
},
|
|
@@ -554,7 +554,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
554
554
|
`${doneIds.length}/${execution.items.length} verifier item(s) already verified; their work stays. Remaining items return to planning.`,
|
|
555
555
|
);
|
|
556
556
|
if (!ok) return;
|
|
557
|
-
await stopExecution(
|
|
557
|
+
await stopExecution(ctx, "interrupted by /update-plan");
|
|
558
558
|
}
|
|
559
559
|
|
|
560
560
|
// Return the run to planning so refinement rules and guards apply again.
|
|
@@ -602,7 +602,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
602
602
|
"Follow the original planning-skill contract for revisions: collect needed clarifications via ask_choice (one question at a time, recorded), apply evidence-based revisions only, then ask the next merged accept/execute question (autoComplete: false — ✓ Accept & execute now / Accept, don't execute yet / another round) and call execute_plan pointing at the new version on accept.",
|
|
603
603
|
);
|
|
604
604
|
|
|
605
|
-
await
|
|
605
|
+
await messaging().sendUserMessage(lines.join("\n"));
|
|
606
606
|
},
|
|
607
607
|
});
|
|
608
608
|
|
|
@@ -613,12 +613,9 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
613
613
|
ctx.ui.notify("No execution in progress.", "info");
|
|
614
614
|
return;
|
|
615
615
|
}
|
|
616
|
-
const ok = await ctx.ui.confirm("Stop execution?", "
|
|
616
|
+
const ok = await ctx.ui.confirm("Stop execution?", "Open tasks will be left unfinished.");
|
|
617
617
|
if (!ok) return;
|
|
618
|
-
|
|
619
|
-
// before the shared stop path clears the state.
|
|
620
|
-
abortDelegatedExecutor();
|
|
621
|
-
await stopExecution(pi, ctx, "stopped by user via /plans-stop");
|
|
618
|
+
await stopExecution(ctx, "stopped by user via /plans-stop");
|
|
622
619
|
ctx.ui.notify("Execution stopped.", "info");
|
|
623
620
|
},
|
|
624
621
|
});
|
|
@@ -627,7 +624,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
627
624
|
description:
|
|
628
625
|
"Resume the working plan in this repository: unfinished planning / reviewing / execution / implementation review, across sessions and linked worktrees, in the current session.",
|
|
629
626
|
handler: async (_args, ctx) => {
|
|
630
|
-
await resumePlansCommand(
|
|
627
|
+
await resumePlansCommand(ctx, baseDir);
|
|
631
628
|
},
|
|
632
629
|
});
|
|
633
630
|
|
|
@@ -654,8 +651,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
654
651
|
// Abandon must end execution first so the planning model is restored.
|
|
655
652
|
disableAutoComplete(ctx, "run abandoned");
|
|
656
653
|
if (abandoningBound && getExecution()) {
|
|
657
|
-
|
|
658
|
-
await stopExecution(pi, ctx, "run abandoned via /plans-abandon");
|
|
654
|
+
await stopExecution(ctx, "run abandoned via /plans-abandon");
|
|
659
655
|
}
|
|
660
656
|
try {
|
|
661
657
|
setRunStatus(ctx.cwd, chosen.run_id, "abandoned");
|
|
@@ -671,7 +667,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
671
667
|
// Session lifecycle
|
|
672
668
|
// -----------------------------------------------------------------------
|
|
673
669
|
pi.on("session_tree", async (_event, ctx) => {
|
|
674
|
-
await restoreFromSession(
|
|
670
|
+
await restoreFromSession(ctx, ctx.sessionManager.getBranch() as unknown as Parameters<typeof restoreFromSession>[1]);
|
|
675
671
|
restoreRunBindingFromSession(
|
|
676
672
|
ctx.sessionManager,
|
|
677
673
|
ctx.cwd,
|
|
@@ -679,7 +675,7 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
679
675
|
);
|
|
680
676
|
});
|
|
681
677
|
pi.on("session_start", async (_event, ctx) => {
|
|
682
|
-
await restoreFromSession(
|
|
678
|
+
await restoreFromSession(ctx, ctx.sessionManager.getBranch() as unknown as Parameters<typeof restoreFromSession>[1]);
|
|
683
679
|
restoreAutoCompleteFromSession(ctx, ctx.sessionManager.getEntries() as unknown as Parameters<typeof restoreAutoCompleteFromSession>[1]);
|
|
684
680
|
restoreRunBindingFromSession(
|
|
685
681
|
ctx.sessionManager,
|
|
@@ -688,3 +684,4 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
|
|
|
688
684
|
);
|
|
689
685
|
});
|
|
690
686
|
}
|
|
687
|
+
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-plans",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.8.0",
|
|
4
4
|
"description": "Human-in-the-loop planning extension for the Pi coding agent: researched, refined Markdown plans before any code changes.",
|
|
5
5
|
"license": "MIT",
|
|
6
6
|
"type": "module",
|
|
@@ -41,6 +41,7 @@
|
|
|
41
41
|
"README.md",
|
|
42
42
|
"LICENSE",
|
|
43
43
|
"CONTRIBUTING.md",
|
|
44
|
+
"AGENTS.md",
|
|
44
45
|
"index.ts",
|
|
45
46
|
"agents/",
|
|
46
47
|
"docs/assets",
|
|
@@ -9,17 +9,17 @@ This skill set is written for the Pi coding agent's documented behavior:
|
|
|
9
9
|
- the five skills are contributed by the pi-plans extension and loaded as Pi skills (also invokable as `/skill:<name>`);
|
|
10
10
|
- skill references and helper sources are resolved relative to the directory containing `SKILL.md`;
|
|
11
11
|
- Planning and reference analysis run with the extension tools `plans`, `ask_choice`, `refine`, `analyze_refs`, and `execute_plan`;
|
|
12
|
-
- `refine` spawns read-only Pi subagents (`pi --mode json -p --no-session --tools read,grep,find,ls`, plus `code_graph`
|
|
12
|
+
- `refine` spawns read-only Pi subagents (`pi --mode json -p --no-session --tools read,grep,find,ls`, plus `code_graph` when workspace `graph_enabled` is true) with isolated context; delegated reviewer runs show a standalone aggregate overlay titled `Reviewer` (78% × 78% top-center, ≥72 cols, no input row), stream assistant/thinking/tool events into per-lane transcripts with follow-bottom scroll, dismiss on `Esc` (close-only — the refiner child keeps running and its result still flows back as tool output), replace any retained finished overlay when a new round begins, and return conclusions to the main session as tool output; `analyze_refs` spawns one read-only subagent per downloaded reference (cwd = that ref's directory) reusing the reviewer role gates, shows the same overlay titled `Refs` in batches of at most 3 lanes, and returns structured per-reference sections for `REF_ANALYSIS.md`;
|
|
13
13
|
- when graph mode is enabled, graph-aware `read`/`edit` overrides are active for indexed source files: `read` returns a capped function digest (≤50 lines, synthetic anonymous entries folded) by default — drill in via `offset/limit` or `code_graph get-function`, and `full: true` is the only whole-file exit (small/zero-function files return full text; safety truncation matches native read); `write`/`edit` stage DB-first mutations until materialized via the `code_graph` tool's `apply` action (same planning/accepted gate as /apply-graph; refused for read-only refiner subagents via the PI_PLANS_REFINER marker; returns a per-file report with counts and a post-apply drift summary, and never changes run status); unexpected fallbacks (`not indexed` / `runtime unavailable` / `config read failed`) are marked at the top of the result while flag-off fallbacks stay unmarked;
|
|
14
|
-
- the execution loop is extension-managed: remaining
|
|
14
|
+
- the execution loop is extension-managed and task-tree driven: the current wave and remaining tasks are injected each turn, progress is reported exclusively through the `plans_update_task` tool (status + evidence / skipReason), the task dashboard tracks every task (compact widget; Ctrl+Shift+T expands the tree), and an independent execution reviewer verifies the verification checks before the run completes;
|
|
15
15
|
- execution and planning compaction keep Pi's SessionManager as the history owner; during active pi-plans runs, `session_before_compact` uses a deterministic no-LLM VCC-style summary with `[Session Goal]`, `[Files And Changes]`, `[Commits]`, `[Outstanding Context]`, `[User Preferences]`, and a ranked brief transcript; Pi core owns manual `/compact`, threshold, and overflow scheduling, while pi-plans handles smart tail keep, `keep:N`, stats, and phase-specific run/plan/current-I/checklist context; in addition, creating a new planning run (`plans start-run`) proactively requests one pre-plan VCC compaction before the first planning question and resumes the planning turn with a hidden message (default on, `prePlanCompact` in `pi-vcc-config.json`);
|
|
16
16
|
|
|
17
17
|
## Planning Boundary
|
|
18
18
|
|
|
19
19
|
- Treat the user's request as a planning target, not as write authorization.
|
|
20
|
-
- Before the execution handoff, do not edit target source files, docs, configs, package metadata, generated assets, or tests outside the planning artifact directory and the pi-plans state under `.git/
|
|
21
|
-
- The normal pre-handoff writes are `.git/
|
|
22
|
-
- Downloaded references go to the workspace's configured refs root (`refs_root` in `.git/
|
|
20
|
+
- Before the execution handoff, do not edit target source files, docs, configs, package metadata, generated assets, or tests outside the planning artifact directory and the pi-plans state under `.git/pi-plans/`. The extension enforces this for `edit` and `write` while a run is active: only `.git/pi-plans/`, the run's artifact directory, `~/.cache/pi-plans/`, and the configured refs root are writable. Bash is not machine-guarded — keep it read-only by discipline (inspection, `git init`, downloads into the cache).
|
|
21
|
+
- The normal pre-handoff writes are `.git/pi-plans/` state plus planning artifacts under the configured artifact root (default `./.git/pi-plans/plans/...`).
|
|
22
|
+
- Downloaded references go to the workspace's configured refs root (`refs_root` in `.git/pi-plans/config.json`; unset → ask once via `ask_choice`, recommended `.git/pi-plans/refs/`, second `./refs/`, third `~/.cache/pi-plans/refs/`; persist with the `plans` tool, `set-refs-root`) and their paths and evidence are recorded in `REF_ANALYSIS.md`. References are not limited to GitHub repositories: papers (arXiv etc.), engineering blogs, and documentation sites are first-class, with per-medium qualified downloads (repo = clone; paper = full text — abstracts never qualify; blog/docs = full-article markdown) and one directory per reference; theoretical references count exactly as much as implementation references. plan-normal and plan-big may optionally cite 1–2 search-found references (URLs in the plan's Evidence section) without downloading them.
|
|
23
23
|
- After the user explicitly approves the execution handoff, leave this planning workflow and execute in the extension-managed loop (see Execution Handoff).
|
|
24
24
|
|
|
25
25
|
## State And Settings
|
|
@@ -30,13 +30,13 @@ Before the first planning question, read `references/state-and-config.md` and in
|
|
|
30
30
|
{ "action": "init", "workdir": "<target-workdir>" }
|
|
31
31
|
```
|
|
32
32
|
|
|
33
|
-
State lives under the workspace's resolved git common dir as `.git/
|
|
33
|
+
State lives under the workspace's resolved git common dir as `.git/pi-plans/` (auto-ignored, no `.gitignore` entries; the tool auto-runs `git init` when the workdir safely has no repository). The target workspace is the current working directory unless the user explicitly names another repository.
|
|
34
34
|
|
|
35
|
-
If `language.tag` is missing from `.git/
|
|
35
|
+
If `language.tag` is missing from `.git/pi-plans/config.json`, ask the language setting question (via `ask_choice`) before any product question. Persist it with `plans` (`set-language`); this question does not count against the planning-question limit.
|
|
36
36
|
|
|
37
|
-
If `artifact_root_source` is missing from `.git/
|
|
37
|
+
If `artifact_root_source` is missing from `.git/pi-plans/config.json` or is `unset`, ask the planning docs location question (via `ask_choice`) before any product question. Persist it with `plans` (`set-artifact-root`); this question does not count against the planning-question limit.
|
|
38
38
|
|
|
39
|
-
|
|
39
|
+
The reviewer role lives in the GLOBAL config (`~/.pi/pi-plans/config.json`, `PI_PLANS_GLOBAL_DIR` override — one confirmation for every workspace). If its mode is missing/invalid when a round is about to run, ask the role-setting question via `ask_choice` and persist with `plans` (`set-role`). The model + thinking level are confirmed at first use through NATIVE panels, not ask_choice: in TUI the `refine` gate itself pops a searchable model panel followed by an effort panel (Default row = no `--thinking` flag; other rows follow the chosen model's `thinkingLevelMap`) and continues the same invocation on completion; hasUI non-TUI sessions get native menus; UI-less sessions get text guidance embedding the available selectors. Esc cancels the whole gate (nothing persisted; do NOT re-ask via ask_choice — suggest `/config-pi-plans`). The `refine` tool refuses to spawn until `reviewerReady` passes (current-session, or delegated with a confirmed concrete `provider/model`). (v0.6.1: the criticizer role is gone — the reviewer emits findings AND questions in one round.)
|
|
40
40
|
|
|
41
41
|
## First-Turn Contract
|
|
42
42
|
|
|
@@ -59,7 +59,7 @@ When the recommended option depends on a web-verifiable claim, search first (web
|
|
|
59
59
|
|
|
60
60
|
Every user-facing planning or refinement question goes through the `ask_choice` tool:
|
|
61
61
|
|
|
62
|
-
- `options`: ordered options, recommended option first with `recommended: true` (exactly one), each with a `description` of the form `✓ <advantage> / ✗ <drawback>` — **every option you author states both what it gains and what it costs**, in the configured language, tersely (≈8 words per half; write `—` when a side is genuinely absent). Both halves live in the single `description` string; there are no separate `pros`/`cons` fields. This rule covers every AI-written option, including the scope confirmation
|
|
62
|
+
- `options`: ordered options, recommended option first with `recommended: true` (exactly one), each with a `description` of the form `✓ <advantage> / ✗ <drawback>` — **every option you author states both what it gains and what it costs**, in the configured language, tersely (≈8 words per half; write `—` when a side is genuinely absent). Both halves live in the single `description` string; there are no separate `pros`/`cons` fields. This rule covers every AI-written option, including the scope confirmation and the merged accept/execute handoff; only the tool-appended `Other…` / `Auto-complete` rows are exempt;
|
|
63
63
|
- do not add `Other` or `Auto-complete` yourself — the tool appends `Other…` second-last and `Auto-complete` last;
|
|
64
64
|
- pass `autoComplete: false` for the merged accept/execute question — it contains the execution approval, so Auto-complete never appears there — and for any install waiver, publishing, deployment, merge, push, credential, or external-state question. Auto-complete may choose the recommended planning or refinement option only.
|
|
65
65
|
- When the user selects Auto-complete, it remains active for the current planning run: later eligible questions use their recommended options automatically, and the extension queues one deduplicated follow-up if the model stops after an auto-completed answer. `/plans-autocomplete-stop` disables it; session restore may reactivate it only for the same active run while its status is `planning`.
|
|
@@ -74,9 +74,9 @@ Create a run only after initial read-only inspection makes the topic clear:
|
|
|
74
74
|
{ "action": "start-run", "workdir": "<target-workdir>", "topic": "<short topic>", "skill": "<skill-name>", "requestText": "<original request>" }
|
|
75
75
|
```
|
|
76
76
|
|
|
77
|
-
The tool creates the configured artifact directory root (default
|
|
77
|
+
The tool creates the configured artifact directory root (default `./.git/pi-plans/plans/YYYY-MM-DD-<topic>/`) and the private run state `.git/pi-plans/runs/<run-id>/`, and stores the active run pointer. Use a short lowercase slug for the topic. Keep paths stable once written.
|
|
78
78
|
|
|
79
|
-
`DECISIONS.md` records: original request; repository evidence inspected; each question, options, selected answer, and whether it came from the user or Auto-complete; assumptions still open; external sources consulted; language
|
|
79
|
+
`DECISIONS.md` records: original request; repository evidence inspected; each question, options, selected answer, and whether it came from the user or Auto-complete; assumptions still open; external sources consulted; language and reviewer settings used.
|
|
80
80
|
|
|
81
81
|
## Final Scope Confirmation
|
|
82
82
|
|
|
@@ -91,15 +91,15 @@ If the user adds requirements, resolve only the necessary follow-up questions, t
|
|
|
91
91
|
|
|
92
92
|
## Plan Artifact Requirements
|
|
93
93
|
|
|
94
|
-
Every `PLAN_vN.md` must include stable IDs that are never recycled across revisions:
|
|
94
|
+
Every `PLAN_vN.md` must include stable IDs that are never recycled across revisions: `Task-N` tasks (with deps/files/wave and one subtask level), `VC-###` verification checks, and the revision ledger. Everything else the workflow needs lives in the run state (decisions ledger, review rounds, refs), not in the plan file.
|
|
95
95
|
|
|
96
|
-
Every plan version
|
|
96
|
+
Every plan version's body is exactly two sections — `## Tasks` and `## Verification Checks` — plus the metadata header (see `references/plan-artifact-template.md` for the full microsyntax: `Task-N` ids, one subtask level, inline `deps:`/`files:`/`wave:` fields, the `### Execution Waves` subsection, and lint rules). Each verification check is a Markdown checkbox of the exact shape:
|
|
97
97
|
|
|
98
98
|
```markdown
|
|
99
|
-
- [ ] `VC-001` covers `
|
|
99
|
+
- [ ] `VC-001` covers `Task-1`; pass condition: ...; evidence: ...; metric: <threshold or reason not quantified>.
|
|
100
100
|
```
|
|
101
101
|
|
|
102
|
-
The execution loop parses `- [ ] \`VC-###\``
|
|
102
|
+
The execution loop parses `- [ ] \`VC-###\`` checks and the task tree, so keep IDs on the checkbox line and the task grammar exact; task progress flows through `plans_update_task` and the execution reviewer reads these checks. Use `references/plan-artifact-template.md` when drafting.
|
|
103
103
|
|
|
104
104
|
## Refinement
|
|
105
105
|
|
|
@@ -109,78 +109,68 @@ After each plan version, ask one merged accept/execute question via `ask_choice`
|
|
|
109
109
|
2. `Accept PLAN_vN, don't execute yet` — mark accepted; resume later via `/plans-execute`.
|
|
110
110
|
3. `Run another round: <the level's default next refine mode>` — only while the level's default sequence is unfinished.
|
|
111
111
|
|
|
112
|
-
The recommended option follows the skill level's default sequence: while the default
|
|
112
|
+
The recommended option follows the skill level's default sequence: while the default round is unfinished it is option 3 (`plan-small` / `plan-normal`: one reviewer round; `plan-big` / `plan-with-refs`: three concurrent reviewers via `refine` with `reviewers: 3`); once the default round is complete it is option 1. Every round returns findings (`F-###`) and up to five questions (`Q-1..Q-5`) in the same output.
|
|
113
113
|
|
|
114
|
-
If the user selects
|
|
114
|
+
If the user selects another round, run the `refine` tool with the plan path and any focus. Reviewer output consolidates into `PLAN_vN_reviewer_comments.md` with findings IDs, severity, affected plan IDs, evidence, impact, recommended fix, and disposition. Revise the next plan only for findings accepted on evidence.
|
|
115
115
|
|
|
116
116
|
### Concurrent Reviewers (big plans)
|
|
117
117
|
|
|
118
118
|
A big-plan reviewer round runs three independent reviewer subagents (`reviewers: 3`); each gets its own emphasis lens but forms its own priorities. After they return, merge and dedupe their findings into one consolidated `PLAN_vN_reviewer_comments.md`, keeping each finding's source reviewer, severity, evidence, and disposition, and surface at most five high-priority comments to the user. Treat agreement between independent reviewers as stronger evidence, not as authority; every accepted finding still needs repo or reference evidence.
|
|
119
119
|
|
|
120
|
-
###
|
|
120
|
+
### Reviewer Questions
|
|
121
121
|
|
|
122
|
-
|
|
122
|
+
Merge the rounds' `Questions` sections into one deduped list and ask EVERY question with `ask_choice` (one batched `questions: [...]` form or one call per question, in the configured language, stable questionIds). Do not revise the plan until every reviewer question has a recorded answer.
|
|
123
123
|
|
|
124
124
|
### Round Lifecycle
|
|
125
125
|
|
|
126
|
-
A refinement round is complete when all reviewer
|
|
126
|
+
A refinement round is complete when all reviewer lanes have returned; each lane's output carries findings (`F-###`) and up to five questions (`Q-1..Q-5`). In the same turn: consolidate, accept or reject each finding on evidence (the user may override any disposition), revise to `PLAN_v(N+1).md` when accepted items require it (copy, edit only the new version, update the revision ledger and verifier checklist), then immediately ask the next merged accept/execute question. Never end a turn merely because a round completed.
|
|
127
127
|
|
|
128
128
|
## Execution Handoff
|
|
129
129
|
|
|
130
|
-
When the user picks `✓ Accept PLAN_vN and execute it now` in the merged question, mark the plan accepted and call the `execute_plan` tool (or the user runs `/plans-execute`). It re-confirms with the user
|
|
130
|
+
When the user picks `✓ Accept PLAN_vN and execute it now` in the merged question, mark the plan accepted and call the `execute_plan` tool (or the user runs `/plans-execute`). It re-confirms with the user (never auto-completed), then the extension enters task-tree execution mode:
|
|
131
131
|
|
|
132
|
-
- every agent turn is injected with the remaining
|
|
133
|
-
-
|
|
134
|
-
-
|
|
132
|
+
- every agent turn is injected with the current wave's open tasks, the remaining task list, verification-check summary, and execution rules (wave order, `plans_update_task` reporting with status + evidence / skipReason, subprocess polling backoff 5s -> 10s -> 20s -> 40s -> 80s then keep polling at 80s, no stopgaps, dependency and library discipline, minimum tests);
|
|
133
|
+
- task progress flows exclusively through the `plans_update_task` tool: one call per task closing it as `complete` (with evidence) or `skipped` (with skipReason); closed statuses are immutable outside the audit-authorized rollback channel; subtasks close before their parent;
|
|
134
|
+
- the task dashboard tracks the whole tree live: a compact aboveEditor widget (current task ▸, progress bar, ✓/· counts, VC pass count, wave indicator, pause state, audit-outcome line) and the Ctrl+Shift+T expanded tree view (✓/▸/~/·/↺ markers, current-task anchor, VC list with audit state, width-adaptive layout). `↺` marks a task that was completed and then rolled back by a failed audit: it is open again but still shows the evidence from its previous attempt;
|
|
135
|
+
- a stall watchdog pauses execution after three consecutive settled rounds without any task-status change (genuine user input or `/plans-execute` resumes without losing progress). A round in which the agent ran successful tool calls counts as progress even with no status change, so investigating the codebase is never mistaken for a dead agent;
|
|
136
|
+
- when every task reaches a terminal state, the run status moves to `verifying` and an independent read-only execution reviewer (`agents/execution-reviewer.md`) verifies each check against the worktree in a detached round whose progress renders in a dedicated overlay (Esc closes it; Ctrl+Shift+R reopens the in-flight round). Each check gets one of three verdicts:
|
|
137
|
+
- `pass` — the evidence establishes the check's condition;
|
|
138
|
+
- `fail` — the condition is demonstrably not met; the covered tasks roll back to pending (children cascade, skipped tasks reopen, the `evidence` of the previous attempt is retained while `skipReason` is cleared) and the audit report is injected;
|
|
139
|
+
- `undeterminable` — the reviewer could not reach a conclusion. This is never a failure and never a completion: nothing rolls back, the loop self-schedules the retry (no agent wake), and each retry counts toward the five-round budget.
|
|
140
|
+
|
|
141
|
+
The run completes only when every pending check is affirmatively `pass`. A partial round credits the checks that passed and leaves the rest for the next round. Checks whose covered tasks are all skipped pass as skipped-pass; checks covering no task never enter the audit; checks already satisfied in an earlier round are neither re-briefed nor re-judged;
|
|
142
|
+
- the round budget is five COMMITTED rounds (pass, fail, or undeterminable; discarded fingerprint-mismatch attempts and cancellations burn nothing; two consecutive discards commit as one undeterminable round). Exhaustion pauses the run in EVERY mode — interactive and auto-approve/headless alike — with an in-band `pi-plans-review-paused` message; ordinary user input and session restores never lift the pause or refill the budget. The only fresh-budget surface is `/plans-execute`, whose explicit confirmation grants five more rounds. The watchdog counter is rebased by real tool activity;
|
|
143
|
+
- execution-phase compaction is handled only when Pi core emits manual `/compact`, threshold, or overflow events; summaries are deterministic VCC-style summaries, include session-derived plan/current-task/checklist context, use smart tail keep and `keep:N`, and never call a model;
|
|
135
144
|
- the read-only guard lifts: full write access returns;
|
|
136
|
-
- the run status moves to `executing`, then `done` when the
|
|
145
|
+
- the run status moves to `executing`, then `verifying` while the review loop owns the run, then `done` when every check passes (a failed round rolls its tasks back and returns the run to `executing` for repair);
|
|
137
146
|
- `/plans-stop` stops execution; `/plans` shows progress.
|
|
138
147
|
|
|
139
148
|
If the user declines, stay in planning (or stop, per their choice). Never start implementation without the approved handoff.
|
|
140
149
|
|
|
141
|
-
###
|
|
150
|
+
### Continuation Between Turns
|
|
142
151
|
|
|
143
|
-
In TUI/RPC, automatic
|
|
152
|
+
In TUI/RPC, automatic continuation is evaluated only at `agent_settled`, after
|
|
144
153
|
Pi has finished natural tool continuation, retries, and compaction. The
|
|
145
|
-
extension rechecks that the same execution is active
|
|
154
|
+
extension rechecks that the same execution is active with open tasks, the
|
|
146
155
|
session is idle, and neither pending input nor compaction owns continuation.
|
|
147
156
|
Each eligible settled cycle can send at most one hidden custom message with
|
|
148
|
-
the current execution rules
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
New-plan handoffs and the `execute_plan` tool still require explicit approval.
|
|
157
|
-
Completion, stop, and session replacement invalidate the extension's wake
|
|
158
|
-
identity without clearing user or other-extension queues. Print/JSON
|
|
159
|
-
single-shot sessions keep VC tracking and completion but never auto-wake;
|
|
157
|
+
the current execution rules. Tool `turn_end` events only update usage; they
|
|
158
|
+
never prequeue continuation reminders. The stall watchdog counts settled
|
|
159
|
+
cycles without task-state change (threshold 3) and pauses instead of waking;
|
|
160
|
+
user interruption and final model errors pause too. Genuine interactive/RPC
|
|
161
|
+
user input or `/plans-execute` resumes a paused active execution without
|
|
162
|
+
losing task progress; extension input cannot unpause it. New-plan handoffs
|
|
163
|
+
and the `execute_plan` tool still require explicit approval. Print/JSON
|
|
164
|
+
single-shot sessions keep task tracking and completion but never auto-wake;
|
|
160
165
|
use RPC for persistent headless execution.
|
|
161
166
|
|
|
162
|
-
### Post-Execution Continuation
|
|
163
|
-
|
|
164
|
-
When execution completes in an interactive session, the completion message attaches a goal-running continuation block and triggers a new agent turn so the model can enter the implementation-review loop immediately. The interactive-only trigger keeps headless sessions silent (no unconsented subagent cost). The same behavior applies on both completion call sites (the normal `turn_end` completion and the `restoreFromSession` recovery path).
|
|
165
|
-
|
|
166
|
-
The agent then asks two `ask_choice` questions (each single-question, `autoComplete: false`) to configure the implementation-review loop:
|
|
167
|
-
|
|
168
|
-
1. **Termination condition** (questionId `termination-condition`, recommended first: goal wait) — options: 1. goal wait: continue until no unpassed VCs remain (auto-continue each round) 2. until no high-severity finding (hard cap 5 rounds) 3. 1 round 4. 2 rounds 5. 3 rounds.
|
|
169
|
-
2. **Reviewer count** (questionId `impl-review-reviewer-count`, `allowOther: false`, pure digit labels `1`/`2`/`3`, recommended first): "How many concurrent reviewers should each implementation-review round use?" The recommended default follows the run's skill: `plan-big` / `plan-with-refs` → 3, others → 1.
|
|
170
|
-
|
|
171
|
-
Both of these setup questions are AI-written options like any other: give each one a `description` of the form `✓ <advantage> / ✗ <drawback>` in the configured language (goal-wait buys thoroughness at the cost of a long run; more concurrent reviewers buy coverage at the cost of tokens), so the trade-off is visible before the user answers.
|
|
172
|
-
|
|
173
|
-
Both answers persist TOGETHER in one `plans record-checkpoint` (`transition: "implementation-review-configured"`, `terminationCondition` + `reviewerCount`). Each refinement round calls `refine` with `role: "reviewer", target: "implementation", reviewers: <configured reviewerCount>`; an omitted `reviewers` falls back to the run's configured `reviewerCount` from the checkpoint, so restarts and worktree migrations never silently revert 2/3 to 1. Each round accepts findings on evidence, applies fixes, re-runs relevant tests, and records progress. The hard cap is 5 rounds regardless of the chosen termination condition. Multi-reviewer rounds (2 or 3) consolidate like big-plan rounds: merge and dedupe findings into one `PLAN_vN_reviewer_comments.md`, keep each finding's source reviewer, severity, evidence, and disposition, and surface at most five high-priority findings. Crash recovery: if the process dies after one or both answers were recorded in `decisions.jsonl` but before the combined write, `/resume-plans` rebuilds the answered configuration from the ledger, asks only the missing question(s), and then performs the single combined write (the re-ask guard rejects duplicate configuration writes while the fields are already set).
|
|
174
|
-
- Round audit trail: `decisions.jsonl`, `subagents.jsonl`, and `pi-plans-ameliorate` entries (one at goal start, then one per round) carry `currentRound` for post-hoc verification.
|
|
175
|
-
- Headless sessions skip the prompt entirely; no `pi-plans-ameliorate` entry is appended.
|
|
176
|
-
|
|
177
167
|
`refine` records each round and lane outcome durably (successful outputs are persisted to run-state files before the tool result returns) and accepts `resumeRoundId` to resume an interrupted round lane-by-lane: completed lanes are reused from their persisted outputs and never re-run; a round id is never reused across plan versions.
|
|
178
168
|
|
|
179
169
|
## Resuming (`/resume-plans`)
|
|
180
170
|
|
|
181
|
-
After a restart or in a fresh session, `/resume-plans` (interactive only) restores the repository's working plan in the current session: the unfinished active run wins; otherwise a unique candidate resumes directly and multiple candidates get a chooser. It resumes unfinished planning (re-asks the pending question with the same `questionId`, never re-asks answered decisions), reviewing (resumes interrupted rounds via `refine resumeRoundId`, consolidates completed ones), execution (durable approval: unchanged plan digest keeps the authorization — a changed HEAD re-
|
|
171
|
+
After a restart or in a fresh session, `/resume-plans` (interactive only) restores the repository's working plan in the current session: the unfinished active run wins; otherwise a unique candidate resumes directly and multiple candidates get a chooser. It resumes unfinished planning (re-asks the pending question with the same `questionId`, never re-asks answered decisions), reviewing (resumes interrupted rounds via `refine resumeRoundId`, consolidates completed ones), and execution (durable approval: unchanged plan digest keeps the authorization — a changed HEAD re-opens previously closed tasks for re-verification; an unverifiable approval HEAD behaves the same; legacy runs without checkpoints must re-approve; v0.6.0 delegated-executor orphans require a fresh handoff approval; a legacy `implementation-review` phase maps to done — its historical acceptance stands). Linked worktrees share candidates; cross-worktree resumes confirm, copy artifacts without overwriting, reset approval and VC validity, and restart round counts. Record semantic boundaries with `plans record-checkpoint` (`plan-written`, `review-consolidated`, `completed` with evidence).
|
|
182
172
|
|
|
183
|
-
`refine`
|
|
173
|
+
`refine` reviews the plan text (the v0.6.0 `target: "implementation"` post-execution loop is gone; the independent execution reviewer now gates delivery).
|
|
184
174
|
|
|
185
175
|
## Red Flags
|
|
186
176
|
|
|
@@ -189,9 +179,9 @@ Stop and return to the workflow if any of these happen:
|
|
|
189
179
|
- implementing before the approved execution handoff;
|
|
190
180
|
- running `refine` without first asking the merged accept/execute question, or before the role gates pass;
|
|
191
181
|
- ending a turn after a completed refinement round without asking the next merged accept/execute question;
|
|
192
|
-
- storing planning settings outside the target workspace's `.git/
|
|
182
|
+
- storing planning settings outside the target workspace's `.git/pi-plans/` state directory;
|
|
193
183
|
- asking multiple planning questions in one message, or asking them outside `ask_choice`;
|
|
194
184
|
- writing `PLAN_v1.md` before final scope confirmation;
|
|
195
185
|
- accepting vague answers that contradict repo or reference evidence;
|
|
196
|
-
- treating a reviewer
|
|
186
|
+
- treating a reviewer as authority instead of evidence;
|
|
197
187
|
- offering Auto-complete for execution, install, deploy, merge, push, or destructive cleanup approval.
|