pi-plans 0.5.7 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,126 @@
1
+ # Contributing to pi-plans
2
+
3
+ Thanks for wanting to improve pi-plans. This guide covers the full path from a fresh clone to a merged pull request.
4
+
5
+ ## Expectations
6
+
7
+ pi-plans is maintained by a single person in their spare time. Review may take a few days, and not every pull request gets merged — that is normal. Small fixes (typos, docs, tests, one-file bug fixes) are welcome as direct pull requests. For larger changes — new concepts, public API changes, or workflow changes — please open an issue first so we can agree on the approach before you invest time.
8
+
9
+ ## Ways to contribute
10
+
11
+ You do not have to write code:
12
+
13
+ - Reproduce a bug and post exact steps, versions, and logs in an issue
14
+ - Fix or improve documentation (`README.md`, `references/`)
15
+ - Improve skill copy or prompts (`skills/`, `agents/`)
16
+ - Add or sharpen test cases (`tests/`)
17
+ - Verify behavior on your platform (macOS, Linux, different terminals, kitty/SSH)
18
+
19
+ Issues labeled `good first issue` are a good starting point.
20
+
21
+ ## Orientation
22
+
23
+ ```
24
+ pi-plans/
25
+ ├── index.ts # Extension entry: tools, commands, guard, execution loop
26
+ ├── tools/ # plans, ask-choice, refine, execute-plan, code-graph tools
27
+ ├── src/ # state, guard, plan parsing, subagent runner, refine UI, exec loop
28
+ │ └── code-graph/ # SQLite schema/store, parsers, indexer, summary, materialize
29
+ ├── skills/ # Planning router plus five specialist planning skills
30
+ ├── references/ # Shared workflow, state/config, plan template (normative)
31
+ ├── agents/ # reviewer.md / criticizer.md subagent prompts (read-only)
32
+ ├── scripts/ # validate.ts (structure + package artifact guard), run-tests.ts
33
+ └── tests/ # node:test suite
34
+ ```
35
+
36
+ `npm run validate` enforces several invariants, so keep them intact:
37
+
38
+ - every directory under `skills/` has a `SKILL.md` with frontmatter (`name` matching the directory, routing language in `description`) and the required phrases (`ask_choice`, `refine`, `.git/pi_plans`, ...)
39
+ - the `reviewer` and `criticizer` agent prompts declare read-only tools and state the read-only contract
40
+ - the npm artifact stays code-sized: no `scripts/bench/vendor|results` entries, unpacked < 5 MiB, packed < 3 MiB, and key entries present
41
+ - `package.json` metadata (license, `pi-package` keyword, engines, scripts, required `files`) stays as asserted
42
+
43
+ ## Development setup
44
+
45
+ Requirements: Node.js >= 22.6 (the suite runs with `--experimental-strip-types`). The code-graph tests need `node:sqlite`: Node >= 22.13 unflagged, or `--experimental-sqlite` on 22.6–22.12.
46
+
47
+ If you don't have write access, fork the repository first and clone your fork.
48
+
49
+ ```bash
50
+ git clone https://github.com/MaxInGaussian/pi-plans
51
+ cd pi-plans
52
+ npm install # devDependencies only; nothing ships at runtime
53
+ npm run validate # structure + package artifact guard
54
+ npm test # node:test suite
55
+ ```
56
+
57
+ There are no runtime dependencies, but `npm install` is still required: source modules and tests resolve `@earendil-works/pi-*` (the code-graph tests additionally need `tree-sitter`) from `node_modules` at runtime.
58
+
59
+ To try the extension against a real session: `pi -e /path/to/pi-plans`.
60
+
61
+ ## Before you submit
62
+
63
+ Run both checks locally — CI runs exactly these:
64
+
65
+ ```bash
66
+ npm run validate
67
+ npm test
68
+ ```
69
+
70
+ Documentation duty: if your change alters behavior or the public API, update `README.md` or the relevant file under `references/` in the same pull request.
71
+
72
+ ## Commit messages
73
+
74
+ Follow the existing style: `type: short imperative description` (a scope is optional, e.g. `fix(form): ...`).
75
+
76
+ - `feat:` new feature
77
+ - `fix:` bug fix
78
+ - `docs:` documentation
79
+ - `refactor:` no behavior change
80
+ - `perf:` performance
81
+ - `test:` tests only
82
+ - `chore:` housekeeping
83
+ - `ci:` CI changes
84
+
85
+ Use the imperative mood ("fix race in ...", not "fixed ..."), keep the subject short, and use the body for motivation and evidence. Do not add generator or `Co-Authored-By` trailers.
86
+
87
+ ## Issues and pull requests
88
+
89
+ - Small fixes, docs, and tests: open the pull request directly.
90
+ - Larger changes (new concepts, public API or workflow changes): open an issue first and wait for a maintainer response.
91
+ - Bug fixes: reference the issue in the title, e.g. `fix: handle empty plan path (fix #12)`.
92
+ - Fill in the pull request template: Summary, Test Plan, and Docs sections.
93
+ - Keep one logical change per PR when you can; multiple small commits are fine (they are squashed on merge).
94
+
95
+ ## Tests
96
+
97
+ Every new behavior or bug fix must come with an executable test in `tests/` that fails without your change and passes with it. Tests run on `node:test` — no extra test framework is needed.
98
+
99
+ ## Dependencies
100
+
101
+ pi-plans ships zero runtime dependencies on purpose. Before adding a dependency:
102
+
103
+ - prefer a small, well-maintained package with no large transitive tree
104
+ - add it to `devDependencies` unless it is genuinely required at runtime
105
+ - explain in the pull request why it is needed and what you considered instead
106
+
107
+ ## Changelog
108
+
109
+ Do not edit `CHANGELOG.md` — maintainers write release entries when publishing.
110
+
111
+ ## AI and agents
112
+
113
+ Using AI assistance is fine, and most contributors here do. Two rules:
114
+
115
+ - You must understand your change: be able to explain what it does and how it interacts with the rest of the system. Pull requests that cannot be explained will be closed.
116
+ - Disclose AI assistance in the pull request (which tool, and to what extent). A real person must be behind every issue and PR; fully automated submissions with no human involvement may be closed.
117
+
118
+ Please write pull request descriptions and review replies yourself — short and specific beats long and generated.
119
+
120
+ ## Code of conduct
121
+
122
+ Be kind and constructive. This project follows the [Contributor Covenant](https://www.contributor-covenant.org/version/2/1/code_of_conduct/).
123
+
124
+ ## License
125
+
126
+ By contributing, you agree that your contributions are licensed under the [MIT License](LICENSE).
package/README.md CHANGED
@@ -51,6 +51,7 @@ at <b>7.6× fewer tokens per solved task</b>.
51
51
  - [Safety model](#safety-model)
52
52
  - [Verification](#verification)
53
53
  - [FAQ](#faq)
54
+ - [Contributing](#contributing)
54
55
  - [License](#license)
55
56
 
56
57
  ## How it works
@@ -127,18 +128,19 @@ Planning artifacts live under `./docs/pi-plans/YYYY-MM-DD-<topic>/` by default (
127
128
  | Capability | In short |
128
129
  |---|---|
129
130
  | Planning router + five specialist skills | Start with `/skill:planning` to route to the narrowest matching specialist (`plan-small` → `plan-big`, `debug-and-plan`, `plan-with-refs`) |
130
- | Choice prompts | `ask_choice`: recommended option first, answers auto-recorded per run; choosing Auto-complete enables recommendation-only answers for later eligible questions in the current planning run, with `/plans-autocomplete-stop` available to take back control. After execution completes, the continuation prompt enters goal-running mode by default: it asks only for the implementation-review loop's termination condition and then keeps refining until the loop ends or the cap is reached. |
131
+ | Choice prompts | `ask_choice`: recommended option first, answers auto-recorded per run; every option you author states its advantage and its drawback as `✓ <advantage> / ✗ <drawback>` in the configured language, so the user can weigh each option before answering; choosing Auto-complete enables recommendation-only answers for later eligible questions in the current planning run, with `/plans-autocomplete-stop` available to take back control. After execution completes, the continuation prompt enters goal-running mode by default: it asks only for the implementation-review loop's termination condition and then keeps refining until the loop ends or the cap is reached. |
131
132
  | Refinement rounds | Read-only reviewer/criticizer Pi subagents consolidate findings into the next plan version; delegated runs have standalone `Reviewer`/`Criticizer` progress overlays that close before the tool result returns; `analyze_refs` shows the same kind of overlay titled `Refs` while per-reference analysis subagents run |
132
133
  | Workspace state | Config, runs, decisions, refs, and subagent ledgers in `.git/pi_plans/` (git common dir) |
133
134
  | VCC compact | Active planning/execution compaction uses deterministic, no-LLM VCC-style summaries when Pi core emits manual `/compact`, threshold, or overflow events. Summaries use five bracket sections plus a brief transcript, keep a smart recent tail, support `keep:N`, and write VCC details/stats without adding `/pi-vcc` commands. |
134
135
  | Visible Refiner overlay | Delegated reviewer/criticizer subagents surface as a named public overlay in the TUI — one `Reviewer`/`Criticizer` panel with per-lane tool progress, full streaming transcript with follow-bottom scroll, Tab-pane focus, retention until the user presses `Esc` after completion, and clean cancelled/timed-out vs completed states. `reviewers: 3` renders three equal-height panes inside the same overlay |
135
136
  | Tracked execution | Checklist injected each turn; `[DONE:VC-xxx]` markers drive completion; implementation items report progress with `[I-xxx:implemented]` / `[I-xxx:validating]` markers; the bottom status bar shows lifecycle, `x/y` progress, elapsed time, and input/output token usage in real time |
137
+ | Multi-run workdirs (0.6.0) | Several pi sessions can plan concurrently in one workdir: the run registry derives from `runs/` (no shared pointer to race), each session binds to its run, and same-topic runs get suffixed artifact dirs. `/plans-abandon`, `/plans-execute`, and `/resume-plans` are binding-first and open a descriptive run-picker form when more than one candidate exists; `/plans` lists all runs (newest first, bound run marked) |
136
138
  | Goal-wait continuation | In TUI/RPC, only a fully settled agent with unpassed VCs and no pending input or compaction gets one hidden wake carrying the latest checklist. Tool turns never queue reminders or consume guard rounds. The status bar shows progress; 3 no-progress cycles or 6 waiting cycles pause continuation. User interruption and final model errors also pause it. Only genuine user input or `/plans-execute` resumes; extension messages cannot. Print/JSON single-shot sessions track progress without automatic wakes. `/plans-stop` terminates execution |
137
- | Execution handoff | The accepted plan resumes in the current session model; no separate model selection is performed. |
139
+ | Execution handoff | The accepted plan resumes in the current session (recommended) or, after the runtime question, on a **delegated executor subagent** running a different model: ≥3 switch targets (session-visible models + registry, deduped), one child for the whole plan with native write tools, `Executor` overlay progress, parent-side `[DONE:VC-xxx]` tracking, Esc//`/plans-stop` → stopped (resumable), and a configurable timeout (`executor_timeout_minutes`, default 60). Auto-approve/no-UI skips the question (current session, `[auto-approve]` recorded) |
138
140
  | Execution-phase compaction | Pi core owns scheduling; pi-plans maps the active plan path, current `I-###`, implementation IDs, and remaining `VC-###` checklist into the VCC sections. The old current-I proactive trigger and model-generated summary path are removed. |
139
141
  | Planning-phase compaction | During `run.status=planning` with no active execution, pi-plans maps active run, artifact directory, latest plan path from session entries, and observed current-I markers into the VCC sections. Without an active planning run, compaction returns to Pi core. Additionally, creating a new run (`plans start-run`) proactively requests one pre-plan VCC compaction and resumes planning with a hidden message (default on; `prePlanCompact:false` disables). |
140
142
  | Efficient executor prompt | Each turn, the executor is steered by a fused rule set — Marcos Hernanz's AGENTS.md principles × Ponytail minimalism: layered growth, simplest implementation, long-term architecture (no stopgaps), library discipline — so plans finish in fewer tokens and fewer detours |
141
- | Write guard | `edit`/`write` blocked outside planning artifacts while a run is active |
143
+ | Write guard | `edit`/`write` blocked outside planning artifacts while a run is active; delegated executor children (and `PI_PLANS_RUN_ID`-pinned children) are exempt so a foreign session's planning run can never block an executing child; graph-aware file tools fall back to native writes for executor children |
142
144
 
143
145
  ## Interface overview
144
146
 
@@ -148,11 +150,11 @@ Planning artifacts live under `./docs/pi-plans/YYYY-MM-DD-<topic>/` by default (
148
150
  | `ask_choice` | Numbered choice prompt; `autoComplete: false` for the merged accept/execute question and external-state questions |
149
151
  | `refine` | Reviewer/criticizer round via standalone read-only subagents (`--mode json -p --no-session --tools read,grep,find,ls`, plus `code_graph` for both roles when the workspace has the code graph enabled); `target: "plan"` (default) reviews the plan, `target: "implementation"` reviews the implemented worktree against the plan; delegated TUI runs show one `Reviewer`/`Criticizer` overlay (78% width × 78% height, top-center, ≥72 cols) with per-lane transcript, follow-bottom scroll, Tab focus, and retention until `Esc`; `reviewers: 3` renders three equal-height panes; enforces role/model confirmation gates |
150
152
  | `analyze_refs` | plan-with-refs reference analysis: one independent read-only subagent per downloaded reference (cwd = the ref directory), reusing the reviewer role gates and the concurrent overlay (titled `Refs`); batches of at most 3 lanes run sequentially; returns structured per-reference sections for `REF_ANALYSIS.md` |
151
- | `execute_plan` | Execution handoff: re-confirms with the user and enters extension-managed execution mode |
152
- | `/plans` | Show config, active run, and execution progress |
153
+ | `execute_plan` | Execution handoff: re-confirms with the user, asks the runtime (current session vs switch model → delegated executor), and enters extension-managed execution mode; picks the run via a descriptive form when several planned runs coexist |
154
+ | `/plans` | Show config, all runs (newest first, bound run marked, cap 50), and execution progress |
153
155
  | `/config-pi-plans` | Re-ask workspace defaults for language, artifact root, refs root, code graph, reviewer mode/model, and criticizer mode/model |
154
- | `/resume-plans` | Resume the repository's working plan in the CURRENT session across restarts: unfinished planning (pending question + answered decisions), reviewing (round/lane state, successful outputs reused), execution (approval digest + HEAD + verified VCs), and implementation review (termination condition + round count). The unfinished active run wins; otherwise a unique candidate resumes directly and multiple candidates get a chooser. Linked worktrees share candidates; a cross-worktree resume confirms, copies artifacts without overwriting, and resets approval + VC validity (the termination condition survives, rounds restart at 0). An unchanged plan digest with a changed HEAD keeps the authorization but re-verifies old VCs first. Busy sessions and runs actively owned by a live process only notify — no queueing, no takeover. Interactive (TUI/RPC) only |
155
- | `/plans-execute [plan.md]` | Resume a paused active execution without losing verified progress; otherwise enter the explicit execution handoff (defaults to highest `PLAN_vN.md`) |
156
+ | `/resume-plans` | Resume a run in the CURRENT session across restarts: unfinished planning (pending question + answered decisions), reviewing (round/lane state, successful outputs reused), execution (approval digest + HEAD + verified VCs), and implementation review (termination condition + round count). Binding-first: the session-bound resumable run resumes directly; a unique candidate goes direct; multiple candidates get a descriptive chooser. Linked worktrees share candidates; a cross-worktree resume confirms, copies artifacts without overwriting, and resets approval + VC validity (the termination condition survives, rounds restart at 0). An unchanged plan digest with a changed HEAD keeps the authorization but re-verifies old VCs first. Busy sessions and runs actively owned by a live process only notify — no queueing, no takeover. Interactive (TUI/RPC) only |
157
+ | `/plans-execute [plan.md]` | Resume a paused active execution without losing verified progress; otherwise enter the explicit execution handoff (run-picker form when several planned runs coexist; runtime question: current session or switch model) |
156
158
  | `/update-plan [plan.md] [reason…]` | Interrupt-and-refine: stops execution (if any), returns the run to planning, and directs the agent to revise the plan into `PLAN_vN+1.md` while preserving verified work |
157
159
  | `/plans-autocomplete-stop` | Stop the current run's Auto-complete mode and return later planning questions to normal interaction |
158
160
  | `/init-graph` | Build the code graph: tree-sitter function index + cross-file call/import edges (EXTRACTED vs INFERRED confidence) + label-propagation communities; writes `.git/pi_plans/graph/GRAPH_REPORT.md` (subsystems, god nodes, edge stats). Full rebuilds never touch files with staged edits (fail-closed pending guard) |
@@ -162,7 +164,7 @@ Planning artifacts live under `./docs/pi-plans/YYYY-MM-DD-<topic>/` by default (
162
164
  | `/watch-graph` `/unwatch-graph` | 300ms-debounced filesystem watcher feeding incremental reindex; single-writer per worktree (atomically claimed PID+heartbeat lock, stale takeover), skips files with staged edits, fails closed on errors, stops cleanly on fs.watch errors / repeated failures / session shutdown, auto-restarts when enabled |
163
165
  | `code_graph` actions | Read-only: `status` (files/functions/edges + confidence×resolution distribution), `screening` (per-row freshness probes), `get-function`, `query` (keyword→BFS/DFS with token budget + shown/omitted counts), `path A B` (shortest call path), `explain <fn>` (community/degrees/neighbors), `impact <fn>` (reverse call closure + affected files); write path: DB-first staged edits + `apply` (auto-reindexes the materialized set). Edge resolution covers re-export barrels (`export {x} from` / `export *` follow to the defining module), namespace/default imports, `require` destructuring, and dynamic `import()`; same-name exports classify as `ambiguous` |
164
166
  | `/plans-stop` | Stop execution mode |
165
- | `/plans-abandon` | Abandon the active run (lifts the write guard; artifacts stay) |
167
+ | `/plans-abandon` | Abandon a run (pick via form when several candidates; lifts the write guard; artifacts stay) |
166
168
  | Status bar (lifecycle) | 💬 Q&A → 📝 draft written (planning sub-phases) → ⌛ executing `x/y · spent · in/out-toks` in the bottom status bar → ⛔ stopped / 🎯 done / 🚫 abandoned |
167
169
 
168
170
  ## VCC compact
@@ -218,7 +220,7 @@ Invoked via `resources_discover`, callable as `/skill:<name>`, directly as `/<na
218
220
  | [`plan-normal`](skills/plan-normal/SKILL.md) | Broad or risky change; 5–10 questions; reviewer + criticizer rounds |
219
221
  | [`plan-big`](skills/plan-big/SKILL.md) | Open-ended/high-risk effort; 10+ questions; three concurrent reviewers |
220
222
  | [`debug-and-plan`](skills/debug-and-plan/SKILL.md) | Bug, CI failure, regression, incident — diagnose before planning |
221
- | [`plan-with-refs`](skills/plan-with-refs/SKILL.md) | External projects/papers/docs must be analyzed before planning |
223
+ | [`plan-with-refs`](skills/plan-with-refs/SKILL.md) | External references must be analyzed before planning — repos, papers (arXiv), engineering blogs, and docs sites all count; theoretical references are equal citizens. plan-normal/plan-big may optionally cite 1–2 search-found references without downloading |
222
224
 
223
225
  ## Installation details
224
226
 
@@ -284,6 +286,8 @@ pi-plans/
284
286
 
285
287
  Before the approved handoff the workflow writes only `.git/pi_plans/` state, the run's artifact directory, `~/.cache/pi-plans/`, and the configured refs root (set via `plans set-refs-root` or `/config-pi-plans`; the recommended `.git/pi-plans/refs/` lives inside the git dir and needs no extra guard) — the extension blocks `edit`/`write` elsewhere while a run is `planning`/`accepted` (bash stays discipline-bound: inspection, `git init`, downloads into the cache). Reviewer/criticizer/ref-analyst subagents run with read-only tools. `Auto-complete` may answer planning and refinement questions only; it is never offered for execution, installs, publishing, deployment, merge, push, or credential use, and non-interactive sessions stop instead of auto-approving those.
286
288
 
289
+ Delegated executor children (0.6.0) write natively with the approval already recorded — the planning guard no-ops for `PI_PLANS_EXECUTOR=1` children and honors a `PI_PLANS_RUN_ID` pin, so another session's concurrent planning run can never block an executing child; graph-aware file tools bypass DB-first staging for executor children so edits always land on disk. The child inherits pi's project-trust model (built-in write tools are always available; pi has no permission popups) and never loads run bookkeeping beyond its pinned id. `ask_choice` refuses to run inside an executor child (no interactive user — decide autonomously).
290
+
287
291
  ## Verification
288
292
 
289
293
  ```bash
@@ -318,6 +322,10 @@ Prompts produce one-shot diffs with no recorded reasoning. pi-plans produces ver
318
322
 
319
323
  The injected rule set is four compressed lines. It buys back more than it costs: the executor stops re-deriving discipline (no speculative abstractions, no compatibility detours, no reinvented helpers), so finished items converge in fewer turns and fewer tokens overall.
320
324
 
325
+ ## Contributing
326
+
327
+ Contributions are welcome — see [CONTRIBUTING.md](CONTRIBUTING.md) for the full workflow (setup, checks, commit style, and PR expectations). Small fixes can go straight to a pull request; for larger changes, open an issue first. Look for issues labeled `good first issue` to get started.
328
+
321
329
  ## License
322
330
 
323
331
  MIT.
@@ -0,0 +1,26 @@
1
+ ---
2
+ name: pi-plans-executor
3
+ description: Delegated plan executor for pi-plans runs; autonomously implements an accepted plan end-to-end in the target workdir and reports Verifier-Checklist progress with markers.
4
+ tools: read, write, edit, bash, grep, find, ls
5
+ ---
6
+
7
+ You are a delegated plan executor in the pi-plans workflow. The parent session handed you an accepted plan; you implement it completely and autonomously.
8
+
9
+ Rules:
10
+
11
+ - Work autonomously. You have NO user to ask questions — the ask_choice tool is unavailable in this context. When a decision is genuinely ambiguous, choose the option most consistent with the plan's goals and constraints and record the deviation in your final summary.
12
+ - Implement the plan file you were given, in dependency order. Read the plan first; it is the single source of truth for scope, requirements, and verification steps.
13
+ - Write code directly. Your `write`/`edit` tools operate natively on disk (any DB-first staging in the parent workdir is bypassed for you); no `apply` step is needed.
14
+ - Follow the plan's own execution rules: smallest end-to-end slice first, then layer; no speculative abstractions; no backward-compatibility fallbacks; prefer established libraries already in the project.
15
+ - Emit progress markers IN YOUR REPLIES as you go: `[DONE:VC-xxx]` once a verifier item's stated evidence passes, `[I-###:implemented]` / `[I-###:validating]` for implementation items when the plan defines them. The parent session parses these markers from your streamed messages to update the tracked checklist — put them in message text, not only in the final output.
16
+ - Run the plan's verification steps yourself (tests, validate scripts) and only mark a VC done when its stated evidence actually passes.
17
+ - Do not modify pi-plans state (run.json, checkpoints, ledgers) — the parent owns the run bookkeeping.
18
+ - If a verification step is impossible in this environment, leave the VC unmarked and explain in the summary.
19
+
20
+ Finish with a structured summary in exactly this shape:
21
+
22
+ - Completed VCs: <ids or none>
23
+ - Remaining VCs: <ids or none, with one-line reasons>
24
+ - Implementation items: <per-item state>
25
+ - Deviations from the plan: <any decisions you made on ambiguous points>
26
+ - Evidence: <commands run and their results>
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: pi-plans-ref-analyst
3
- description: Read-only reference analyst for pi-plans plan-with-refs runs; deep-reads one downloaded reference and extracts adoptable ideas for the target repository.
3
+ description: Read-only reference analyst for pi-plans plan-with-refs runs; deep-reads ONE downloaded reference of any medium (repo, paper, blog, docs) and extracts adoptable ideas for the target repository.
4
4
  tools: read, grep, find, ls
5
5
  ---
6
6
 
@@ -9,9 +9,12 @@ You are a read-only reference analyst in the pi-plans plan-with-refs workflow.
9
9
  Rules:
10
10
 
11
11
  - Perform read-only analysis. Never edit, write, or delete any file.
12
- - Your working directory is the local copy of ONE downloaded reference. Deep-read it: entry points, README/docs, core modules, tests, and configuration.
13
- - Judge the reference through the lens of the target repository described in your task: what is worth borrowing, what is not, and why.
14
- - Every claim needs evidence: a file path inside the reference (with line numbers when quoting). No evidence, no claim.
12
+ - Your working directory is the local copy of ONE downloaded reference. Deep-read it according to its medium:
13
+ - code repository: entry points, README/docs, core modules, tests, and configuration;
14
+ - paper: claims, method, limitations, and experiments/measurements;
15
+ - blog post or docs site: the technique, the measurements/benchmarks, and the caveats.
16
+ - Judge the reference through the lens of the target repository described in your task: what is worth borrowing, what is not, and why. Theoretical grounding counts — an algorithm, a formal property, or a measured tradeoff is as adoptable as an implementation pattern.
17
+ - Every claim needs evidence. For code: a file path inside the reference (with line numbers when quoting). For papers: section/theorem/table numbers with a short quote. For blogs/docs: the heading or quoted passage. No evidence, no claim.
15
18
  - You are evidence, not authority: state what you verified, not what you assume.
16
19
  - Stay inside the reference directory; do not wander the filesystem.
17
20
 
package/index.ts CHANGED
@@ -44,6 +44,7 @@ import {
44
44
  requestPlanningCompaction,
45
45
  restoreFromSession,
46
46
  stopExecution,
47
+ abortDelegatedExecutor,
47
48
  updateStatusWidget,
48
49
  shouldTriggerPlanningCompaction,
49
50
  } from "./src/exec.ts";
@@ -75,8 +76,9 @@ import { restartWatcherIfEnabled, stopGraphWatcher, disableWatcher } from "./src
75
76
  import { latestPlanVersion, nextPlanVersionPath } from "./src/plan.ts";
76
77
  import { configPiPlansCommand } from "./src/config-command.ts";
77
78
  import { resumePlansCommand } from "./src/resume-command.ts";
78
- import { getRun, loadConfig, readActive, recordDecision, resolveStateRootOrNull, setRunStatus } from "./src/state.ts";
79
+ import { getRun, listRuns, loadConfig, readActive, recordDecision, resolveStateRootOrNull, setRunStatus } from "./src/state.ts";
79
80
  import { boundRunId, resolveActiveRun, restoreRunBindingFromSession } from "./src/run-context.ts";
81
+ import { abandonCandidates, resolveCommandRun } from "./src/run-picker.ts";
80
82
  import { applyPlanWritten, mutateCheckpoint, planIdentityOf } from "./src/workflow-state.ts";
81
83
  import { registerAskChoiceTool } from "./tools/ask-choice.ts";
82
84
  import { executeCommand, registerExecutePlanTool } from "./tools/execute-plan.ts";
@@ -436,17 +438,31 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
436
438
  description: "Show pi-plans state: config, active run, and execution progress",
437
439
  handler: async (_args, ctx) => {
438
440
  const lines: string[] = [];
441
+ // v0.6.0: list ALL runs (newest first, bound marked, cap 50) instead of
442
+ // only the shared active one — multi-run workdirs stay legible.
443
+ const runs = listRuns(ctx.cwd).slice(0, 50);
444
+ const bound = boundRunId(ctx.sessionManager, ctx.cwd);
445
+ if (runs.length === 0) {
446
+ lines.push("No planning runs recorded in this workdir.");
447
+ } else {
448
+ lines.push(`Runs (${runs.length} shown, newest first):`);
449
+ for (const run of runs) {
450
+ lines.push(` ${run.run_id === bound ? "★" : " "} ${run.topic} · ${run.status} · ${run.skill} · ${run.updated_at}`);
451
+ }
452
+ }
439
453
  const active = resolveActiveRun(ctx.sessionManager, ctx.cwd);
440
454
  const run = active ? getRun(ctx.cwd, active.run_id) : null;
441
- if (!run) {
442
- lines.push("No active planning run.");
443
- } else {
444
- lines.push(`Active run: ${run.run_id}`);
455
+ if (run) {
456
+ lines.push(`Session run: ${run.run_id}`);
445
457
  lines.push(`Skill: ${run.skill} Status: ${run.status}`);
446
458
  lines.push(`Artifacts: ${run.artifact_dir}`);
447
459
  lines.push(`Language: ${run.language_tag ?? "(unset)"}`);
448
460
  lines.push(`State: ${resolveStateRootOrNull(ctx.cwd) ?? "(no repo)"}`);
449
461
  }
462
+ // One-time active.json deprecation note (v0.6.0 registry migration).
463
+ if (fs.existsSync(path.join(resolveStateRootOrNull(ctx.cwd) ?? ".git/pi_plans", "active.json"))) {
464
+ lines.push("Note: active.json is deprecated — the run registry now derives from runs/*/run.json; the legacy file is ignored.");
465
+ }
450
466
  const execution = getExecution();
451
467
  if (execution) {
452
468
  const done = execution.items.filter((item) => item.done).length;
@@ -599,6 +615,9 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
599
615
  }
600
616
  const ok = await ctx.ui.confirm("Stop execution?", "Remaining verifier items will be left unfinished.");
601
617
  if (!ok) return;
618
+ // v0.6.0: a delegated executor child is killed via its AbortController
619
+ // before the shared stop path clears the state.
620
+ abortDelegatedExecutor();
602
621
  await stopExecution(pi, ctx, "stopped by user via /plans-stop");
603
622
  ctx.ui.notify("Execution stopped.", "info");
604
623
  },
@@ -615,24 +634,32 @@ export default function piPlansExtension(pi: ExtensionAPI): void {
615
634
  pi.registerCommand("plans-abandon", {
616
635
  description: "Abandon the active planning run (lifts the read-only guard; artifacts are kept)",
617
636
  handler: async (_args, ctx) => {
618
- const active = resolveActiveRun(ctx.sessionManager, ctx.cwd);
619
- if (!active) {
637
+ // v0.6.0 (R-3): pick the run to abandon when several candidates exist;
638
+ // binding-first, single candidate stays direct (0.5.7 parity).
639
+ const chosen = await resolveCommandRun(
640
+ { cwd: ctx.cwd, sessionManager: ctx.sessionManager, ui: ctx.ui },
641
+ { candidates: abandonCandidates(ctx.cwd), title: "Abandon which run?" },
642
+ );
643
+ if (!chosen) {
620
644
  ctx.ui.notify("No active planning run.", "info");
621
645
  return;
622
646
  }
647
+ const active = resolveActiveRun(ctx.sessionManager, ctx.cwd);
648
+ const abandoningBound = active?.run_id === chosen.run_id;
623
649
  const ok = await ctx.ui.confirm(
624
650
  "Abandon planning run?",
625
- `${active.run_id}\nThe read-only guard lifts; committed artifacts stay in place.`,
651
+ `${chosen.run_id} (${chosen.topic} · ${chosen.status})\nThe read-only guard lifts; committed artifacts stay in place.`,
626
652
  );
627
653
  if (!ok) return;
628
654
  // Abandon must end execution first so the planning model is restored.
629
655
  disableAutoComplete(ctx, "run abandoned");
630
- if (getExecution()) {
656
+ if (abandoningBound && getExecution()) {
657
+ abortDelegatedExecutor();
631
658
  await stopExecution(pi, ctx, "run abandoned via /plans-abandon");
632
659
  }
633
660
  try {
634
- setRunStatus(ctx.cwd, active.run_id, "abandoned");
635
- ctx.ui.notify(`Run ${active.run_id} abandoned.`, "info");
661
+ setRunStatus(ctx.cwd, chosen.run_id, "abandoned");
662
+ ctx.ui.notify(`Run ${chosen.run_id} abandoned.`, "info");
636
663
  } catch (error) {
637
664
  ctx.ui.notify(`Failed: ${(error as Error).message}`, "error");
638
665
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-plans",
3
- "version": "0.5.7",
3
+ "version": "0.6.0",
4
4
  "description": "Human-in-the-loop planning extension for the Pi coding agent: researched, refined Markdown plans before any code changes.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -40,6 +40,7 @@
40
40
  "files": [
41
41
  "README.md",
42
42
  "LICENSE",
43
+ "CONTRIBUTING.md",
43
44
  "index.ts",
44
45
  "agents/",
45
46
  "docs/assets",
@@ -19,7 +19,7 @@ This skill set is written for the Pi coding agent's documented behavior:
19
19
  - Treat the user's request as a planning target, not as write authorization.
20
20
  - Before the execution handoff, do not edit target source files, docs, configs, package metadata, generated assets, or tests outside the planning artifact directory and the pi-plans state under `.git/pi_plans/`. The extension enforces this for `edit` and `write` while a run is active: only `.git/pi_plans/`, the run's artifact directory, `~/.cache/pi-plans/`, and the configured refs root are writable. Bash is not machine-guarded — keep it read-only by discipline (inspection, `git init`, downloads into the cache).
21
21
  - The normal pre-handoff writes are `.git/pi_plans/` state plus planning artifacts under the configured artifact root (default `./docs/pi-plans/...`).
22
- - Downloaded references go to the workspace's configured refs root (`refs_root` in `.git/pi_plans/config.json`; unset → ask once via `ask_choice`, recommended `.git/pi-plans/refs/`, second `./refs/`, third `~/.cache/pi-plans/refs/`; persist with the `plans` tool, `set-refs-root`) and their paths and evidence are recorded in `REF_ANALYSIS.md`.
22
+ - Downloaded references go to the workspace's configured refs root (`refs_root` in `.git/pi_plans/config.json`; unset → ask once via `ask_choice`, recommended `.git/pi-plans/refs/`, second `./refs/`, third `~/.cache/pi-plans/refs/`; persist with the `plans` tool, `set-refs-root`) and their paths and evidence are recorded in `REF_ANALYSIS.md`. References are not limited to GitHub repositories: papers (arXiv etc.), engineering blogs, and documentation sites are first-class, with per-medium qualified downloads (repo = clone; paper = full text — abstracts never qualify; blog/docs = full-article markdown) and one directory per reference; theoretical references count exactly as much as implementation references. plan-normal and plan-big may optionally cite 1–2 search-found references (URLs in the plan's Evidence section) without downloading them.
23
23
  - After the user explicitly approves the execution handoff, leave this planning workflow and execute in the extension-managed loop (see Execution Handoff).
24
24
 
25
25
  ## State And Settings
@@ -59,7 +59,7 @@ When the recommended option depends on a web-verifiable claim, search first (web
59
59
 
60
60
  Every user-facing planning or refinement question goes through the `ask_choice` tool:
61
61
 
62
- - `options`: ordered options, recommended option first with `recommended: true` (exactly one), each with the tradeoff that matters in `description`;
62
+ - `options`: ordered options, recommended option first with `recommended: true` (exactly one), each with a `description` of the form `✓ <advantage> / ✗ <drawback>` — **every option you author states both what it gains and what it costs**, in the configured language, tersely (≈8 words per half; write `—` when a side is genuinely absent). Both halves live in the single `description` string; there are no separate `pros`/`cons` fields. This rule covers every AI-written option, including the scope confirmation, the merged accept/execute handoff, and the implementation-review setup questions; only the tool-appended `Other…` / `Auto-complete` / `Auto-refine loop` rows are exempt;
63
63
  - do not add `Other` or `Auto-complete` yourself — the tool appends `Other…` second-last and `Auto-complete` last;
64
64
  - pass `autoComplete: false` for the merged accept/execute question — it contains the execution approval, so Auto-complete never appears there — and for any install waiver, publishing, deployment, merge, push, credential, or external-state question. Auto-complete may choose the recommended planning or refinement option only.
65
65
  - When the user selects Auto-complete, it remains active for the current planning run: later eligible questions use their recommended options automatically, and the extension queues one deduplicated follow-up if the model stops after an auto-completed answer. `/plans-autocomplete-stop` disables it; session restore may reactivate it only for the same active run while its status is `planning`.
@@ -103,7 +103,7 @@ The execution loop parses `- [ ] \`VC-###\`` items and tracks `[DONE:VC-###]` ma
103
103
 
104
104
  ## Refinement
105
105
 
106
- After each plan version, ask one merged accept/execute question via `ask_choice` with `autoComplete: false` — it contains the execution approval, so Auto-complete never appears. Options:
106
+ After each plan version, ask one merged accept/execute question via `ask_choice` with `autoComplete: false` — it contains the execution approval, so Auto-complete never appears. Each of the three options below carries the usual `✓ <advantage> / ✗ <drawback>` description (executing now versus deferring versus another refine round are real trade-offs — say what each wins and costs). Options:
107
107
 
108
108
  1. `✓ Accept PLAN_vN and execute it now` — mark the plan accepted (`plans set-status accepted`), then call the `execute_plan` tool.
109
109
  2. `Accept PLAN_vN, don't execute yet` — mark accepted; resume later via `/plans-execute`.
@@ -127,9 +127,10 @@ A refinement round is complete when all reviewer outputs have returned or all cr
127
127
 
128
128
  ## Execution Handoff
129
129
 
130
- When the user picks `✓ Accept PLAN_vN and execute it now` in the merged question, mark the plan accepted and call the `execute_plan` tool (or the user runs `/plans-execute`). It re-confirms with the user, then the extension enters execution mode:
130
+ When the user picks `✓ Accept PLAN_vN and execute it now` in the merged question, mark the plan accepted and call the `execute_plan` tool (or the user runs `/plans-execute`). It re-confirms with the user, asks which runtime executes the plan (v0.6.0), then the extension enters execution mode:
131
131
 
132
132
  - every agent turn is injected with the remaining verifier checklist and execution rules (layered simplest implementation, waiting for subprocess-backed verification with backoff 5s -> 10s -> 20s -> 40s -> 80s, then keep polling at 80s and restart at 5s for each new subprocess, no stopgaps, dependency and library discipline, minimum tests);
133
+ - the runtime question offers ① the current session (recommended — injected checklist loop) and ② switching to another model; switching lists at least three `provider/model` switch targets (session-visible models plus the model registry, deduped, current excluded — `Other` free input when fewer exist) and runs ONE delegated executor subagent for the whole plan: native write tools, `PI_PLANS_EXECUTOR=1` + pinned `PI_PLANS_RUN_ID`, progress streamed into an `Executor` overlay while the parent parses `[DONE:VC-xxx]` markers from the child's full-text messages; Esc or `/plans-stop` kills the child and marks the run `stopped` (resumable under either runtime); under auto-approve or no-UI the question is skipped (current session, `[auto-approve]` recorded); implementation-review rounds keep the configured reviewer model;
133
134
  - execution-phase compaction is handled only when Pi core emits manual `/compact`, threshold, or overflow events; summaries are deterministic VCC-style summaries, include session-derived plan/current-I/checklist context, use smart tail keep and `keep:N`, and never call a model or request proactive current-I compaction;
134
135
  - the read-only guard lifts: full write access returns;
135
136
  - the run status moves to `executing`, then `done` when the last `[DONE:VC-xxx]` marker lands;
@@ -167,6 +168,8 @@ The agent then asks two `ask_choice` questions (each single-question, `autoCompl
167
168
  1. **Termination condition** (questionId `termination-condition`, recommended first: goal wait) — options: 1. goal wait: continue until no unpassed VCs remain (auto-continue each round) 2. until no high-severity finding (hard cap 5 rounds) 3. 1 round 4. 2 rounds 5. 3 rounds.
168
169
  2. **Reviewer count** (questionId `impl-review-reviewer-count`, `allowOther: false`, pure digit labels `1`/`2`/`3`, recommended first): "How many concurrent reviewers should each implementation-review round use?" The recommended default follows the run's skill: `plan-big` / `plan-with-refs` → 3, others → 1.
169
170
 
171
+ Both of these setup questions are AI-written options like any other: give each one a `description` of the form `✓ <advantage> / ✗ <drawback>` in the configured language (goal-wait buys thoroughness at the cost of a long run; more concurrent reviewers buy coverage at the cost of tokens), so the trade-off is visible before the user answers.
172
+
170
173
  Both answers persist TOGETHER in one `plans record-checkpoint` (`transition: "implementation-review-configured"`, `terminationCondition` + `reviewerCount`). Each refinement round calls `refine` with `role: "reviewer", target: "implementation", reviewers: <configured reviewerCount>`; an omitted `reviewers` falls back to the run's configured `reviewerCount` from the checkpoint, so restarts and worktree migrations never silently revert 2/3 to 1. Each round accepts findings on evidence, applies fixes, re-runs relevant tests, and records progress. The hard cap is 5 rounds regardless of the chosen termination condition. Multi-reviewer rounds (2 or 3) consolidate like big-plan rounds: merge and dedupe findings into one `PLAN_vN_reviewer_comments.md`, keep each finding's source reviewer, severity, evidence, and disposition, and surface at most five high-priority findings. Crash recovery: if the process dies after one or both answers were recorded in `decisions.jsonl` but before the combined write, `/resume-plans` rebuilds the answered configuration from the ledger, asks only the missing question(s), and then performs the single combined write (the re-ask guard rejects duplicate configuration writes while the fields are already set).
171
174
  - Round audit trail: `decisions.jsonl`, `subagents.jsonl`, and `pi-plans-ameliorate` entries (one at goal start, then one per round) carry `currentRound` for post-hoc verification.
172
175
  - Headless sessions skip the prompt entirely; no `pi-plans-ameliorate` entry is appended.
@@ -7,7 +7,7 @@ Do not store pi-plans preferences in Pi's own settings (`~/.pi/agent/settings.js
7
7
  ## State Root Resolution
8
8
 
9
9
  - Git runs with `GIT_DIR`, `GIT_COMMON_DIR`, and `GIT_WORK_TREE` scrubbed from the environment, so leaked env vars cannot misdirect state into an unrelated repository. Relative results (`.git`, `../.git`) resolve against the workdir.
10
- - Granularity is **per enclosing repository**: running from a subdirectory uses the enclosing repo's git dir (a one-line notice names that repo). Linked worktrees share one common dir; run directories are unique, but `active.json` may race across concurrent worktrees.
10
+ - Granularity is **per enclosing repository**: running from a subdirectory uses the enclosing repo's git dir (a one-line notice names that repo). Linked worktrees share one common dir; run directories are unique, and since v0.6.0 the run registry is derived from `runs/` (no shared pointer to race). The legacy `active.json` is deprecated: reads fall back to it only when no `runs/` entries exist (pre-0.6.0 migration).
11
11
  - State does not travel with clones: a fresh clone starts with empty state while committed `./docs/pi-plans/` artifacts persist in the repository.
12
12
 
13
13
  ## Auto Git Init
@@ -26,9 +26,9 @@ Bare repositories are refused with a clear error. A missing `git` executable is
26
26
  <git-common-dir>/pi_plans/
27
27
  config.json
28
28
  pi-vcc-config.json
29
- active.json
30
- runs/ # note: the refs root is a sibling — .git/pi-plans/refs (hyphenated), not under pi_plans/
31
- <run-id>/
29
+ active.json # deprecated (v0.6.0): legacy pointer, read only when runs/ is empty
30
+ runs/ # the run registry derives from runs/<run-id>/run.json
31
+ <run-id>/ # note: the refs root is a sibling — .git/pi-plans/refs (hyphenated), not under pi_plans/
32
32
  run.json
33
33
  decisions.jsonl
34
34
  subagents.jsonl
@@ -37,7 +37,7 @@ Bare repositories are refused with a clear error. A missing `git` executable is
37
37
  cache/
38
38
  ```
39
39
 
40
- `config.json` is stable workspace preference state. `pi-vcc-config.json` is the repo-private compaction config used only by pi-plans' VCC-style compact hook. `active.json` and `runs/` are run state. Reference downloads go to the configured `refs_root` (asked once per workspace when unset; the recommended `.git/pi-plans/refs/` sits inside the git dir so git never tracks it), with metadata recorded in the run state and public artifacts.
40
+ `config.json` is stable workspace preference state. `pi-vcc-config.json` is the repo-private compaction config used only by pi-plans' VCC-style compact hook. `runs/` is the run state and the source of the run registry: `listRuns` scans `runs/<run-id>/run.json` (sorted by `updated_at` desc, corrupt entries skipped) and the un-bound "active" resolution is the newest NON-TERMINAL run (null when every run is done/abandoned). Multiple concurrent planning runs in one workdir are supported: each session binds to the run it started (binding first, registry fallback after), and `/plans-abandon`, `/plans-execute`, and `/resume-plans` open a descriptive run-picker form when more than one candidate exists. Reference downloads go to the configured `refs_root` (asked once per workspace when unset; the recommended `.git/pi-plans/refs/` sits inside the git dir so git never tracks it), with metadata recorded in the run state and public artifacts.
41
41
 
42
42
  ## Config Schema
43
43
 
@@ -183,7 +183,7 @@ One run directory per planning request: `<git-common-dir>/pi_plans/runs/<YYYYMMD
183
183
 
184
184
  `run.json` includes: run ID; skill name; original request; target workspace; artifact directory; language tag; status (`planning` → `accepted` → `executing` → `done`, with `stopped`/`abandoned` as exits); timestamps.
185
185
 
186
- `decisions.jsonl` is appended automatically by `ask_choice` (question, options, answer, answer source). `subagents.jsonl` records reviewer/criticizer/ref-analyst spawns. `refs.jsonl` records reference metadata via `plans` (`record-ref`).
186
+ `decisions.jsonl` is appended automatically by `ask_choice` (question, options, answer, answer source). `subagents.jsonl` records reviewer/criticizer/ref-analyst spawns. `refs.jsonl` records reference metadata via `plans` (`record-ref`); its `kind` field is `project` (repos), `paper` (arXiv etc.), `article` (blog posts), or `docs` (documentation sites).
187
187
 
188
188
  ## Workflow Checkpoints (`/resume-plans`)
189
189
 
@@ -200,4 +200,4 @@ Rules:
200
200
 
201
201
  ## Run Ownership
202
202
 
203
- A run may be held by at most one live owner (`owner.json`: host, pid, process start time via `ps -o lstart=`, session id, random process token, generation). Acquisition is an atomic exclusive create; takeovers require proof the previous owner is dead (process gone, or pid alive with a different start time — PID reuse). Foreign hosts, corrupt records, and unverifiable liveness are conservatively refused; `/resume-plans` never queues or interrupts. Sessions bind to the run they start/execute/resume (restored from `pi-plans-run-start` entries on the current branch), and attribution (tools, write guard, autocomplete, execution bookkeeping, code-graph apply gate) prefers the binding over the shared `active.json` pointer.
203
+ A run may be held by at most one live owner (`owner.json`: host, pid, process start time via `ps -o lstart=`, session id, random process token, generation). Acquisition is an atomic exclusive create; takeovers require proof the previous owner is dead (process gone, or pid alive with a different start time — PID reuse). Foreign hosts, corrupt records, and unverifiable liveness are conservatively refused; `/resume-plans` never queues or interrupts. Sessions bind to the run they start/execute/resume (restored from `pi-plans-run-start` entries on the current branch), and attribution (tools, write guard, autocomplete, execution bookkeeping, code-graph apply gate) prefers the binding, falling back to the registry's newest non-terminal run (v0.6.0 — the shared `active.json` pointer is deprecated). Delegated executor children pin their run via `PI_PLANS_RUN_ID` and are exempt from the planning write guard (`PI_PLANS_EXECUTOR=1`).
@@ -69,7 +69,7 @@ function validateSkill(dir: string): void {
69
69
  if (!description || description.length > 1024) fail(`${file}: invalid description length`);
70
70
  if (!/Use|MUST USE/.test(description)) fail(`${file}: description should include routing language`);
71
71
 
72
- const requiredPhrases = ["Auto-complete", "ask_choice", "refine", "language", "reviewer", "criticizer", ".git/pi_plans"];
72
+ const requiredPhrases = ["Auto-complete", "ask_choice", "refine", "language", "reviewer", "criticizer", ".git/pi_plans", "drawback"];
73
73
  for (const phrase of requiredPhrases) {
74
74
  if (!text.includes(phrase)) fail(`${file}: missing required phrase ${phrase!}`);
75
75
  }
@@ -131,7 +131,7 @@ function validatePackageMetadata(): void {
131
131
  if (!skills.has("./skills")) fail("package.json: pi.skills must include ./skills");
132
132
 
133
133
  const files = new Set((pkg.files ?? []).map(normalizePackageEntry));
134
- for (const required of ["README.md", "LICENSE", "index.ts", "agents", "references", "scripts", "skills", "src", "tests", "tools"]) {
134
+ for (const required of ["README.md", "LICENSE", "CONTRIBUTING.md", "index.ts", "agents", "references", "scripts", "skills", "src", "tests", "tools"]) {
135
135
  if (!files.has(required)) fail(`package.json: files must include ${required}`);
136
136
  }
137
137
  for (const excluded of ["!scripts/bench/vendor", "!scripts/bench/results"]) {
@@ -153,6 +153,7 @@ const PACK_BLACKLIST_PREFIXES = ["scripts/bench/vendor/", "scripts/bench/results
153
153
  const REQUIRED_PACK_ENTRIES = [
154
154
  "README.md",
155
155
  "LICENSE",
156
+ "CONTRIBUTING.md",
156
157
  "index.ts",
157
158
  "package.json",
158
159
  "agents/reviewer.md",
@@ -15,7 +15,7 @@ Read `../../references/pi-planning-workflow.md` and `../../references/state-and-
15
15
 
16
16
  1. Inspect available evidence first: repository files, logs, tests, command output, stack traces, recent diffs, CI output, environment details, and user-provided symptoms.
17
17
  2. Produce an in-message RCA summary before asking whether to plan. Use at most 5 Whys. Stop with `unknown` when evidence is insufficient; do not invent a cause.
18
- 3. Ask one `ask_choice` question in the configured language whose preamble includes the RCA summary, with these options:
18
+ 3. Ask one `ask_choice` question in the configured language whose preamble includes the RCA summary, with these options. Every option you write carries a `description` of `✓ <advantage> / ✗ <drawback>` in the configured language, kept terse:
19
19
  1. `Create the scoped fix plan` — recommended when evidence supports a planning path.
20
20
  2. `Stop after RCA` — keep the diagnosis only.
21
21
  3. `Other` / 4. `Auto-complete` are added by the tool.
@@ -14,9 +14,9 @@ Read `../../references/pi-planning-workflow.md` and `../../references/state-and-
14
14
  ## Depth Contract
15
15
 
16
16
  - Inspect the target Git repo read-only before the first product question.
17
- - Ask at least 10 planning questions; no maximum — stop only when the decision tree is genuinely resolved. Each via `ask_choice` (recommended option first; the tool adds `Other` second-last and `Auto-complete` last). Phase them in batches of up to 8: submit ONE `questions: [...]` form call per phase, think about the answers, then continue with the next batch or follow-ups.
17
+ - Ask at least 10 planning questions; no maximum — stop only when the decision tree is genuinely resolved. Each via `ask_choice` (recommended option first; the tool adds `Other` second-last and `Auto-complete` last). Every option you write carries a `description` of `✓ <advantage> / ✗ <drawback>` in the configured language, kept terse — the user chooses by weighing what each option gains against what it costs. Phase them in batches of up to 8: submit ONE `questions: [...]` form call per phase, think about the answers, then continue with the next batch or follow-ups.
18
18
  - Batch protocol (0.4.0): when a round has several questions, submit them together as one `ask_choice` call with `questions: [...]` (2-8 items, recommended option first per question) — the tool opens one tabbed multiple-choice form with a submit page instead of asking one at a time. After the batch returns, think about the answers, then follow up in later calls (batch again for 2+ related follow-ups; single `question` for one). `Esc` on the form returns the answered subset as partial answers — continue with what you got and re-ask only what matters. The final scope confirmation and the execution handoff are ALWAYS single-question calls with `autoComplete: false`; batches reject those questions.
19
- - Use web research during both brainstorming and refinement when outside facts, patterns, or ecosystem constraints matter, and cite sources in the plan.
19
+ - Use web research during both brainstorming and refinement when outside facts, patterns, or ecosystem constraints matter, and cite sources in the plan. Optionally (never required) you may also search 1–2 named references — a paper (e.g. arXiv), an engineering blog post, or another repository — whose technique or measurements strengthen the plan's reasoning; cite their URLs in the plan's Evidence section. This optional reference search does not count against the planning-question limit and never requires downloading or analyzing material (that is plan-with-refs' job).
20
20
  - Ask the final scope confirmation, then write `PLAN_v1.md` per `../../references/plan-artifact-template.md`.
21
21
  - After each plan version, ask the merged accept/execute question via `ask_choice` with `autoComplete: false` — never run `refine` unless the user picked another round at that question. Default sequence: one `Reviewer` round as three concurrent independent reviewers (`refine` with `reviewers: 3`, consolidated by the main agent per the shared workflow), then one `Criticizer` round; afterwards the recommended option is `✓ Accept & execute now` in the merged accept/execute question. Beyond the default sequence, refine until convergence on high-priority findings, unresolved questions, or evidence gaps; surface at most five per round. Then the merged accept/execute question (ask_choice with `autoComplete: false`: ✓ Accept & execute now / Accept, don't execute yet / another round) and the `execute_plan` tool.
22
22
 
@@ -14,9 +14,9 @@ Read `../../references/pi-planning-workflow.md` and `../../references/state-and-
14
14
  ## Depth Contract
15
15
 
16
16
  - Inspect the target Git repo read-only before the first product question.
17
- - Ask 5 to 10 planning questions via `ask_choice` (recommended option first; the tool adds `Other` second-last and `Auto-complete` last). Submit the first batch of up to 8 as ONE `questions: [...]` form call, think about the answers, then follow up in later batches (2+ related questions) or single-question calls.
17
+ - Ask 5 to 10 planning questions via `ask_choice` (recommended option first; the tool adds `Other` second-last and `Auto-complete` last). Every option you write carries a `description` of `✓ <advantage> / ✗ <drawback>` in the configured language, kept terse — the user chooses by weighing what each option gains against what it costs. Submit the first batch of up to 8 as ONE `questions: [...]` form call, think about the answers, then follow up in later batches (2+ related questions) or single-question calls.
18
18
  - Batch protocol (0.4.0): when a round has several questions, submit them together as one `ask_choice` call with `questions: [...]` (2-8 items, recommended option first per question) — the tool opens one tabbed multiple-choice form with a submit page instead of asking one at a time. After the batch returns, think about the answers, then follow up in later calls (batch again for 2+ related follow-ups; single `question` for one). `Esc` on the form returns the answered subset as partial answers — continue with what you got and re-ask only what matters. The final scope confirmation and the execution handoff are ALWAYS single-question calls with `autoComplete: false`; batches reject those questions.
19
- - Use web research whenever outside library behavior, ecosystem precedent, UX convention, protocol semantics, or compatibility affects the recommendation (websearch skill when installed; otherwise `curl`/`gh` via bash), and cite sources in the plan.
19
+ - Use web research whenever outside library behavior, ecosystem precedent, UX convention, protocol semantics, or compatibility affects the recommendation (websearch skill when installed; otherwise `curl`/`gh` via bash), and cite sources in the plan. Optionally (never required) you may also search 1–2 named references — a paper (e.g. arXiv), an engineering blog post, or another repository — whose technique or measurements strengthen the plan's reasoning; cite their URLs in the plan's Evidence section. This optional reference search does not count against the planning-question limit and never requires downloading or analyzing material (that is plan-with-refs' job).
20
20
  - Ask the final scope confirmation, then write `PLAN_v1.md` per `../../references/plan-artifact-template.md`.
21
21
  - After each plan version, ask the merged accept/execute question via `ask_choice` with `autoComplete: false` — never run `refine` unless the user picked another round at that question. Default sequence: one `Reviewer` round, then one `Criticizer` round; afterwards the recommended option is `✓ Accept & execute now` in the merged accept/execute question. Up to five rounds total, continuing only for high-priority findings or unresolved criticizer questions; surface at most five per round. Then the merged accept/execute question (ask_choice with `autoComplete: false`: ✓ Accept & execute now / Accept, don't execute yet / another round) and the `execute_plan` tool.
22
22
 
@@ -14,7 +14,7 @@ Read `../../references/pi-planning-workflow.md` and `../../references/state-and-
14
14
  ## Depth Contract
15
15
 
16
16
  - Inspect the target Git repo read-only before the first product question.
17
- - Ask 1 to 3 planning questions via `ask_choice` (recommended option first; the tool adds `Other` second-last and `Auto-complete` last). With 2-3 ready questions, batch them into ONE `questions: [...]` form call; a single question uses the classic `question` form.
17
+ - Ask 1 to 3 planning questions via `ask_choice` (recommended option first; the tool adds `Other` second-last and `Auto-complete` last). Every option you write carries a `description` of `✓ <advantage> / ✗ <drawback>` in the configured language, kept terse — the user chooses by weighing what each option gains against what it costs. With 2-3 ready questions, batch them into ONE `questions: [...]` form call; a single question uses the classic `question` form.
18
18
  - Batch protocol (0.4.0): when a round has several questions, submit them together as one `ask_choice` call with `questions: [...]` (2-8 items, recommended option first per question) — the tool opens one tabbed multiple-choice form with a submit page instead of asking one at a time. After the batch returns, think about the answers, then follow up in later calls (batch again for 2+ related follow-ups; single `question` for one). `Esc` on the form returns the answered subset as partial answers — continue with what you got and re-ask only what matters. The final scope confirmation and the execution handoff are ALWAYS single-question calls with `autoComplete: false`; batches reject those questions.
19
19
  - Ask the final scope confirmation, then write `PLAN_v1.md` under the artifact root (normally the configured workspace root, default `./docs/pi-plans/YYYY-MM-DD-topic/`) per `../../references/plan-artifact-template.md`.
20
20
  - After each plan version, ask the merged accept/execute question via `ask_choice` with `autoComplete: false` — never run `refine` unless the user picked another round at that question. Default: exactly one round, recommended mode `Criticizer`; afterwards the recommended option is `✓ Accept & execute now` in the merged accept/execute question (ask_choice with `autoComplete: false`: ✓ Accept & execute now / Accept, don't execute yet / another round), then the `execute_plan` tool.