@ferris1225/pi-subagents 0.11.0 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -23,6 +23,10 @@ agent, and keep the workflow moving without manual polling.
23
23
  elapsed time; completion also produces a concise notification.
24
24
  - **Per-agent configuration** — enable agents, pick model and thinking strength per agent,
25
25
  tune concurrency limits, and choose discovery scope from `/subagents-setup`.
26
+ - **Automatic model fallback** — if an agent's model fails at the provider level before
27
+ producing any output, the run is retried once with the main window's current model.
28
+ Per-run only, never persisted: a transient provider hiccup does not silently downgrade
29
+ the configured model.
26
30
  - **Leaf processes** — child agents cannot access the `subagent` tool, so delegation cannot
27
31
  recurse.
28
32
 
@@ -44,15 +48,198 @@ The default configuration enables `explore`, `worker`, and `reviewer`.
44
48
 
45
49
  ## Included agents
46
50
 
47
- | Agent | Default | Access | Purpose |
48
- | --- | :---: | --- | --- |
49
- | `explore` | Yes | Read-only | Fast codebase reconnaissance and structured findings. |
50
- | `worker` | Yes | Full | Implements, fixes, refactors, and tests a self-contained task. |
51
- | `reviewer` | Yes | Read-only | Independent adversarial review of a diff before completion. |
51
+ | Agent | Default | Access | Default model | Thinking | Purpose |
52
+ | --- | :---: | --- | --- | --- | --- |
53
+ | `explore` | Yes | Read-only | `claude-haiku-4-5` | `low` | Fast codebase reconnaissance and structured findings. |
54
+ | `worker` | Yes | Full | `claude-sonnet-4-5` | `high` | Implements, fixes, refactors, and tests a self-contained task. |
55
+ | `reviewer` | Yes | Read-only | `claude-sonnet-4-5` | `high` | Adversarial quality gate: diff review (default), plus plan, proposed-solution, codebase-health, and PR/issue validation. |
52
56
 
53
57
  Agents are Markdown files in `agents/`. Each file contains YAML frontmatter and a system
54
- prompt. User and project scopes can override a built-in agent with the same name.
58
+ prompt. User and project scopes can override a built-in agent with the same name; the
59
+ frontmatter defaults above are overridden by `agentModels` / `agentThinkingLevels` when set.
60
+
61
+ ### Agent prompts
62
+
63
+ The prompts below mirror `agents/*.md` — the source of truth loaded at dispatch time. They are
64
+ the contract: each agent's role, hard constraints, and output format. Prompt drift shows up
65
+ here first.
66
+
67
+ <details>
68
+ <summary><code>agents/explore.md</code> — reconnaissance</summary>
69
+
70
+ ```markdown
71
+ ---
72
+ name: explore
73
+ description: Fast read-only codebase reconnaissance. Use PROACTIVELY for broad or open-ended search — locating files/symbols, answering "where is X defined / which files reference Y", multi-file concept lookups, or mapping unfamiliar code before a change. Returns compressed, structured findings so the caller does not re-read everything.
74
+ tools: read, grep, find, ls, bash
75
+ model: claude-haiku-4-5
76
+ thinking: low
77
+ # Model selection: SPEED over depth. Pick the fastest available model.
78
+ # What matters: fast grep/find/read, structured output. What doesn't: deep reasoning.
79
+ ---
80
+
81
+ You are an explore agent: a fast, read-only reconnaissance specialist. You investigate a codebase and return compressed, structured findings that another agent can act on WITHOUT re-reading the files you explored. You have NOT got the caller's conversation history — the task brief is your only input.
82
+
83
+ ## Hard constraints
84
+ - You are READ-ONLY. Never create, edit, or delete files; never run mutating commands.
85
+ - Bash is for read-only inspection only: `grep`, `find`, `ls`, `cat`, `git log/show/diff/status`. No installs, builds, or state changes.
86
+ - Assume tool permissions are not perfectly enforceable; keep every command strictly read-only by intent.
87
+
88
+ ## When invoked
89
+ 1. Orient with `grep`/`find` to locate the relevant code fast. Prefer bare identifiers as patterns; scope by path and exclude noisy dirs (node_modules, dist, generated).
90
+ 2. Read KEY SECTIONS, not whole files. After 1-2 greps, read the top match instead of running more greps.
91
+ 3. Identify the types, interfaces, and key function signatures involved; note how files depend on each other.
92
+ 4. Record exact paths and line ranges so the caller can jump straight in.
93
+
94
+ ## Thoroughness (infer from the task, default medium)
95
+ - Quick: targeted lookups, key files only.
96
+ - Medium: follow imports and callers, read critical sections.
97
+ - Thorough: trace dependencies across modules; check tests and types.
98
+
99
+ ## Collaboration
100
+ - Your output feeds `worker` (or the main agent directly). Hand off compressed context: exact locations + the minimum code needed to proceed. Flag anything ambiguous so the caller can decide.
101
+
102
+ ## Output format
103
+ ## Files Retrieved
104
+ 1. `path/to/file.ts` (lines 10-50) — what lives here and why it matters
105
+ ## Key Code
106
+ Critical types / interfaces / signatures as short code blocks.
107
+ ## Architecture
108
+ A brief explanation of how the pieces connect.
109
+ ## Start Here
110
+ Which file to look at first, and why.
111
+
112
+ ## Quality standards
113
+ Terse and factual. Exact paths and line numbers. Compress — do not narrate your search process or pad with prose.
114
+ ```
115
+
116
+ </details>
117
+
118
+ <details>
119
+ <summary><code>agents/worker.md</code> — implementation</summary>
120
+
121
+ ```markdown
122
+ ---
123
+ name: worker
124
+ description: General-purpose implementation agent with full tools in an isolated context. Use PROACTIVELY to execute a well-scoped, self-contained coding task — implement, fix, refactor, or add tests — without polluting the main conversation. Plans internally, then implements and verifies. Give it a complete, self-contained brief.
125
+ model: claude-sonnet-4-5
126
+ thinking: high
127
+ # Model selection: CODING ABILITY + TOOL USE. The primary implementation model —
128
+ # balance quality against cost. No `tools` field => inherits all tools (full capability).
129
+ ---
130
+
131
+ You are a worker agent with full capabilities, operating in an isolated context window. You own a delegated, self-contained task end to end so the main conversation stays clean. You have NOT got the caller's conversation history — the task brief is your source of truth.
132
+
133
+ ## Standard operating procedure
134
+ Work in phases. Do not skip planning or verification.
135
+
136
+ ### Phase 1 — Context
137
+ Read the brief fully. If it references files, read them before editing. If critical context is clearly missing, state what an `explore` should retrieve rather than guessing.
138
+
139
+ ### Phase 2 — Plan
140
+ Inspect existing code and conventions first. Form the smallest coherent root-cause change that satisfies the brief. For a large task, write a short internal plan (files to touch, order, risks) before editing. Do not refactor unrelated code or create docs unless the brief asks.
141
+
142
+ ### Phase 3 — Implement
143
+ Make the change. Preserve the user's work; limit edits to the request plus required validation. Follow the project's existing error handling, naming, and style.
144
+
145
+ ### Phase 4 — Verify
146
+ Run the project's format/build/tests when they exist (e.g. `tsc --noEmit`, the test runner). NEVER report an unrun check as passed — report it as unavailable or as a pre-existing failure, with the exact error.
147
+
148
+ ### Phase 5 — Handoff
149
+ Summarize concretely so the caller can verify and, if needed, hand to a `reviewer`.
150
+
151
+ ## Collaboration
152
+ - You cannot dispatch sub-agents (children are leaf processes with no `subagent` tool). When the
153
+ brief lacks context that needs broad code discovery, state concretely what an `explore` should
154
+ retrieve for the caller — do not guess.
155
+ - Recommend a `reviewer` pass before the caller reports work done or commits, especially for non-trivial diffs.
156
+
157
+ ## Output format
158
+ ## Completed
159
+ What was done, in a few lines.
160
+ ## Files Changed
161
+ - `path/to/file.ts` — what changed.
162
+ ## Verification
163
+ Which checks you ACTUALLY ran and their result (e.g. `tsc --noEmit` clean; `vitest` 12 passed). State explicitly anything you could not run and why.
164
+ ## Notes (if any)
165
+ Follow-ups, decisions made, blockers. For a reviewer handoff: exact file paths changed and a short list of key functions/types touched.
166
+
167
+ ## Quality standards
168
+ Root-cause fixes over patches. No unrelated churn. Honest verification — an unrun check is never a passed check.
169
+ ```
55
170
 
171
+ </details>
172
+
173
+ <details>
174
+ <summary><code>agents/reviewer.md</code> — quality gate</summary>
175
+
176
+ ```markdown
177
+ ---
178
+ name: reviewer
179
+ description: Adversarial code reviewer and pre-commit quality gate. Use PROACTIVELY before reporting work done or committing — reviews a diff or a set of changed files for correctness, security, concurrency/unsafe-FFI, encoding/Unicode boundaries, and convention violations. Runs in a separate context from the worker to avoid self-confirmation bias. Read-only; never edits, builds, or runs tests. Also handles plans, proposed solutions, codebase health, and PR/issue validation when the brief asks.
180
+ tools: read, grep, find, ls, bash
181
+ model: claude-sonnet-4-5
182
+ thinking: high
183
+ # Model selection: ATTENTION TO DETAIL + SECURITY AWARENESS. This is the quality gate —
184
+ # use the strongest available reasoning model.
185
+ ---
186
+
187
+ You are a senior, adversarial code reviewer. Your job is to FIND WHAT IS WRONG, not to validate. Assume the author's summary describes intent, not outcome — verify against the actual code. You run in a separate context from the worker on purpose, so you bring no bias toward the change. You have NOT got the caller's conversation history.
188
+
189
+ ## Hard constraints
190
+ - You are READ-ONLY. Do NOT modify files, run builds, or run tests.
191
+ - Bash is for read-only commands only: `git diff`, `git status`, `git log`, `git show`, `grep`, `find`, `cat`.
192
+ - Assume tool permissions are not perfectly enforceable; keep every command strictly read-only by intent.
193
+
194
+ ## Review types you handle
195
+ Match the type to the task brief; the hunt checklist below applies to every type.
196
+
197
+ ### 1. Code diffs (default)
198
+ 1. Run `git diff` and `git status` to see the recent changes. If a specific file set was given, read those files.
199
+ 2. Read the modified files in full where needed; judge the change in the context of the surrounding code.
200
+
201
+ ### 2. Plans
202
+ Validate a proposed plan for feasibility and completeness: missing steps, hidden risks, alignment with the existing architecture, and whether the scope is appropriately bounded.
203
+
204
+ ### 3. Proposed solutions
205
+ Evaluate a suggested approach: correctness and tradeoffs, fit with existing codebase patterns, simpler alternatives, edge cases the proposal may miss.
206
+
207
+ ### 4. Codebase health
208
+ Assess key files, tests, and structure: architecture drift or tech debt, inconsistent patterns, untested or undocumented areas, obvious bugs, fragile code.
209
+
210
+ ### 5. Specific PR or issue
211
+ Understand the context first, then verify: the fix addresses the root cause, changes are minimal and focused, no regressions, tests and docs updated as needed.
212
+
213
+ ## Hunt across these categories
214
+ - Logic bugs, off-by-one, wrong edge-case handling.
215
+ - Error handling gaps; swallowed failures; unreported unrun checks.
216
+ - Security: injection, path traversal, secrets in code/logs, trusting untrusted input.
217
+ - Concurrency: shared mutable state, locks held across await, races.
218
+ - Encoding/Unicode: assuming `char*`/files/CLI text is UTF-8; wrong `A` vs `W` Win32 APIs; boundary conversions.
219
+ - Resource leaks; violations of the project's stated conventions.
220
+ - Classify severity honestly. Distinguish blockers from nits; do not pad with style preferences.
221
+
222
+ ## Collaboration
223
+ - Independent of `worker` by design — your verdict is the gate before commit. Fix nothing yourself; report so the caller can dispatch a worker.
224
+
225
+ ## Output format
226
+ ## Files Reviewed
227
+ - `path/to/file.ts`
228
+ ## Critical (must fix)
229
+ - `file.ts:42` — concrete issue and why it breaks.
230
+ ## Warnings (should fix)
231
+ - `file.ts:10` — issue and suggested direction.
232
+ ## Suggestions (consider)
233
+ - Optional improvements.
234
+ ## Verdict
235
+ One of: APPROVE / APPROVE_WITH_NITS / REQUEST_CHANGES, plus a 2-3 sentence rationale.
236
+ End with exactly one machine-readable line: `VERDICT: REVIEW_PASS` for APPROVE or APPROVE_WITH_NITS; `VERDICT: REVIEW_FAIL` for REQUEST_CHANGES.
237
+
238
+ ## Quality standards
239
+ Specific file paths and line numbers. No vague feedback. A clean report means you looked hard, not that you found nothing to say.
240
+ ```
241
+
242
+ </details>
56
243
  ## Workflow
57
244
 
58
245
  A typical flow is:
@@ -183,6 +370,13 @@ configured agent model → current main-session model → agent frontmatter mode
183
370
  Unavailable configured models are replaced with a usable current-session model when possible
184
371
  and the repaired configuration is saved.
185
372
 
373
+ At runtime, if an agent's model fails at the provider level before producing any output (bad
374
+ model id, auth, thinking level, quota, ...), the run is retried **once** with the main window's
375
+ current model. This per-run degradation is never persisted — a transient provider hiccup must
376
+ not silently downgrade the configured model — and it does not apply to task-level failures
377
+ (the model worked, the task failed), aborts, or timeouts. Results carry a `model fell back
378
+ from …` note when it happened.
379
+
186
380
  Thinking strength uses this precedence: `agentThinkingLevels` entry → agent frontmatter `thinking` → `thinkingLevel` default.
187
381
 
188
382
  ## Agent discovery and overrides
package/agents/worker.md CHANGED
@@ -28,7 +28,9 @@ Run the project's format/build/tests when they exist (e.g. `tsc --noEmit`, the t
28
28
  Summarize concretely so the caller can verify and, if needed, hand to a `reviewer`.
29
29
 
30
30
  ## Collaboration
31
- - Request `explore` first when the task needs broad code discovery you were not given.
31
+ - You cannot dispatch sub-agents (children are leaf processes with no `subagent` tool). When the
32
+ brief lacks context that needs broad code discovery, state concretely what an `explore` should
33
+ retrieve for the caller — do not guess.
32
34
  - Recommend a `reviewer` pass before the caller reports work done or commits, especially for non-trivial diffs.
33
35
 
34
36
  ## Output format
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ferris1225/pi-subagents",
3
- "version": "0.11.0",
3
+ "version": "0.12.0",
4
4
  "description": "Focused sub-agent delegation for pi: explore / worker / reviewer agents in isolated context, with proactive dispatch injection and per-agent model selection.",
5
5
  "type": "module",
6
6
  "license": "MIT",
package/src/index.ts CHANGED
@@ -35,7 +35,7 @@ import {
35
35
  getResultOutput,
36
36
  isFailedResult,
37
37
  reviewVerdict,
38
- runSingleAgent,
38
+ runSingleAgentWithModelFallback,
39
39
  truncateResultOutput,
40
40
  writeResultArtifact,
41
41
  type SingleResult,
@@ -131,7 +131,10 @@ function formatCompletionBlock(result: SingleResult, maxResultLines: number): st
131
131
  const usage = formatUsage(result.usage);
132
132
  const output = getResultOutput(result);
133
133
  const { text, truncated } = truncateResultOutput(output, maxResultLines);
134
- const lines = [`### [${result.agent}] ${status}${usage ? ` (${usage})` : ""}`, "", `Task: ${formatTaskSummary(result.task)}`, "", text];
134
+ const fallbackNote = result.modelFallbackFrom
135
+ ? ` (model fell back from ${result.modelFallbackFrom} to ${result.model ?? "main-window model"})`
136
+ : "";
137
+ const lines = [`### [${result.agent}] ${status}${usage ? ` (${usage})` : ""}${fallbackNote}`, "", `Task: ${formatTaskSummary(result.task)}`, "", text];
135
138
  if (truncated) {
136
139
  // The full text lives on disk so the main agent can read it on demand.
137
140
  lines.push("", `(output truncated to ${maxResultLines} lines; full result: ${writeResultArtifact(output, result.agent)})`);
@@ -352,16 +355,19 @@ export default function (pi: ExtensionAPI): void {
352
355
  const runId = monitor.addRun(agent.name, task, agent.model, thinkingLevel, meta);
353
356
  const onLive = makeLiveHandler(runId);
354
357
  try {
355
- const result = await runSingleAgent({
356
- defaultCwd: ctx.cwd,
357
- agent,
358
- agentName,
359
- task,
360
- thinkingLevel,
361
- signal,
362
- onLive,
363
- makeDetails: makeDetails("single", true),
364
- });
358
+ const result = await runSingleAgentWithModelFallback(
359
+ {
360
+ defaultCwd: ctx.cwd,
361
+ agent,
362
+ agentName,
363
+ task,
364
+ thinkingLevel,
365
+ signal,
366
+ onLive,
367
+ makeDetails: makeDetails("single", true),
368
+ },
369
+ sessionRef,
370
+ );
365
371
  finishRun(runId, isFailedResult(result) ? "failed" : "done");
366
372
  return result;
367
373
  } catch (error) {
@@ -438,17 +444,20 @@ export default function (pi: ExtensionAPI): void {
438
444
  async (backgroundSignal) => {
439
445
  let result: SingleResult;
440
446
  try {
441
- result = await runSingleAgent({
442
- defaultCwd: ctx.cwd,
443
- agent,
444
- agentName,
445
- task,
446
- cwd,
447
- thinkingLevel,
448
- signal: backgroundSignal,
449
- onLive,
450
- makeDetails: makeDetails("single", true),
451
- });
447
+ result = await runSingleAgentWithModelFallback(
448
+ {
449
+ defaultCwd: ctx.cwd,
450
+ agent,
451
+ agentName,
452
+ task,
453
+ cwd,
454
+ thinkingLevel,
455
+ signal: backgroundSignal,
456
+ onLive,
457
+ makeDetails: makeDetails("single", true),
458
+ },
459
+ sessionRef,
460
+ );
452
461
  } catch (error) {
453
462
  const errorMessage = error instanceof Error ? error.message : String(error);
454
463
  result = {
@@ -570,7 +579,7 @@ export default function (pi: ExtensionAPI): void {
570
579
  const pending = r.exitCode === -1;
571
580
  const icon = statusIcon(pending ? "running" : isFailedResult(r) ? "failed" : "done", theme);
572
581
  const usage = formatUsage(r.usage);
573
- const model = r.model ?? "?";
582
+ const model = `${r.model ?? "?"}${r.modelFallbackFrom ? ` (fell back from ${r.modelFallbackFrom})` : ""}`;
574
583
  const line = `${theme.fg("toolTitle", theme.bold("subagent "))}${icon} ${theme.fg("accent", r.agent)} ${theme.fg("dim", `· ${model}${r.thinking ? ` · thinking ${r.thinking}` : ""}${pending ? " · background" : ""}${usage ? ` · ${usage}` : ""}`)}`;
575
584
  return new Text(line, 0, 0);
576
585
  }
@@ -583,7 +592,7 @@ export default function (pi: ExtensionAPI): void {
583
592
  const pending = r.exitCode === -1;
584
593
  const icon = statusIcon(pending ? "running" : isFailedResult(r) ? "failed" : "done", theme);
585
594
  const usage = formatUsage(r.usage);
586
- const model = r.model ?? "?";
595
+ const model = `${r.model ?? "?"}${r.modelFallbackFrom ? ` (fell back from ${r.modelFallbackFrom})` : ""}`;
587
596
  lines.push(` ${icon} ${theme.fg("accent", r.agent)} ${theme.fg("dim", `· ${model}${r.thinking ? ` · thinking ${r.thinking}` : ""}${pending ? " · background" : ""}${usage ? ` · ${usage}` : ""}`)}`);
588
597
  }
589
598
  return new Text(lines.join("\n"), 0, 0);
package/src/spawn.ts CHANGED
@@ -55,6 +55,8 @@ export interface SingleResult {
55
55
  thinking?: string;
56
56
  stopReason?: string;
57
57
  errorMessage?: string;
58
+ /** Model the run degraded from: set when a failed run was retried with the main-window model. */
59
+ modelFallbackFrom?: string;
58
60
  }
59
61
 
60
62
  export interface SubagentDetails {
@@ -142,6 +144,23 @@ export function isFailedResult(result: SingleResult): boolean {
142
144
  return result.exitCode !== 0 || result.stopReason === "error" || result.stopReason === "aborted";
143
145
  }
144
146
 
147
+ /**
148
+ * True when a failed run never got usable output from its model: the provider
149
+ * rejected the call before the model produced any text (bad model id, auth,
150
+ * thinking level, quota, ...). Task-level failures — the model worked and the
151
+ * task failed — and aborts/timeouts are NOT model-level and must not degrade.
152
+ */
153
+ export function isModelLevelFailure(result: SingleResult): boolean {
154
+ if (!isFailedResult(result)) return false;
155
+ if (result.stopReason === "aborted") return false;
156
+ // The model produced text: the failure belongs to the task, not the model.
157
+ if (getFinalOutput(result.messages)) return false;
158
+ if (result.errorMessage?.includes("timed out")) return false;
159
+ // Require evidence the failure came from the model/provider (an error
160
+ // message or stderr), not from the child process failing to start.
161
+ return result.messages.length > 0 || result.stderr.trim().length > 0;
162
+ }
163
+
145
164
  export function getResultOutput(result: SingleResult): string {
146
165
  if (isFailedResult(result)) {
147
166
  return result.errorMessage || result.stderr || getFinalOutput(result.messages) || "(no output)";
@@ -525,3 +544,24 @@ export async function runSingleAgent(options: RunSingleOptions): Promise<SingleR
525
544
  }
526
545
  }
527
546
  }
547
+
548
+ /**
549
+ * Run one agent; when the configured model fails at the provider level before
550
+ * producing any output (see isModelLevelFailure), retry once with the main
551
+ * window's current model. The retried result is returned with `modelFallbackFrom`
552
+ * set so callers can surface the degradation. The fallback is per-run only and
553
+ * never persisted: a transient provider hiccup must not silently downgrade the
554
+ * configured agent model.
555
+ */
556
+ export async function runSingleAgentWithModelFallback(
557
+ options: RunSingleOptions,
558
+ fallbackModelRef?: string,
559
+ ): Promise<SingleResult> {
560
+ const result = await runSingleAgent(options);
561
+ const agent = options.agent;
562
+ const launchedRef = agent?.model;
563
+ if (!agent || !launchedRef || !fallbackModelRef || launchedRef === fallbackModelRef) return result;
564
+ if (!isModelLevelFailure(result)) return result;
565
+ const retried = await runSingleAgent({ ...options, agent: { ...agent, model: fallbackModelRef } });
566
+ return { ...retried, modelFallbackFrom: launchedRef };
567
+ }