pi-subagents 0.59.0 → 0.61.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (72) hide show
  1. package/CHANGELOG.md +69 -0
  2. package/docs/agents.md +2 -2
  3. package/docs/configuration.md +9 -5
  4. package/docs/extension-api.md +14 -7
  5. package/docs/models.md +1 -1
  6. package/docs/observability.md +1 -1
  7. package/docs/tool-reference.md +15 -3
  8. package/docs/workflows.md +15 -14
  9. package/install.mjs +2 -1
  10. package/package.json +1 -1
  11. package/skills/council-mode/SKILL.md +48 -243
  12. package/skills/council-mode/references/pass-contracts.md +150 -0
  13. package/skills/pi-subagents/SKILL.md +89 -37
  14. package/skills/pi-subagents/references/constraints-and-recipes.md +30 -234
  15. package/skills/pi-subagents/references/execution-controls.md +79 -7
  16. package/skills/pi-subagents/references/management-authoring-rpc.md +2 -2
  17. package/skills/pi-subagents/references/multi-lane-orchestration.md +13 -1
  18. package/skills/pi-subagents/references/prompting-and-roles.md +35 -28
  19. package/skills/pi-subagents/references/review-and-validation.md +73 -0
  20. package/src/agents/agent-management.ts +256 -88
  21. package/src/agents/agents.ts +527 -221
  22. package/src/api/background-work.ts +7 -2
  23. package/src/api/external-runs.ts +67 -4
  24. package/src/api/preflight.ts +13 -8
  25. package/src/api/shared-types.ts +3 -0
  26. package/src/extension/index.ts +7 -4
  27. package/src/extension/public-execution.ts +48 -4
  28. package/src/extension/rpc.ts +62 -4
  29. package/src/extension/schemas.ts +12 -7
  30. package/src/extension/tool-description.ts +16 -16
  31. package/src/runs/background/async-execution.ts +53 -37
  32. package/src/runs/background/async-job-tracker.ts +65 -3
  33. package/src/runs/background/async-resume.ts +3 -1
  34. package/src/runs/background/async-status.ts +104 -11
  35. package/src/runs/background/auto-drain.ts +1 -1
  36. package/src/runs/background/control-channel.ts +3 -2
  37. package/src/runs/background/fleet-view.ts +1 -1
  38. package/src/runs/background/result-watcher.ts +1 -1
  39. package/src/runs/background/resume-guidance.ts +1 -1
  40. package/src/runs/background/run-status.ts +15 -4
  41. package/src/runs/background/subagent-runner.ts +6 -3
  42. package/src/runs/background/subagent-wait.ts +30 -23
  43. package/src/runs/background/wait-completions.ts +4 -1
  44. package/src/runs/background/wait-tool.ts +24 -18
  45. package/src/runs/foreground/execution.ts +72 -4
  46. package/src/runs/foreground/subagent-executor.ts +276 -59
  47. package/src/runs/shared/acceptance.ts +43 -18
  48. package/src/runs/shared/async-status-projection.ts +138 -4
  49. package/src/runs/shared/background-process-options.ts +9 -0
  50. package/src/runs/shared/host-step-status.ts +1 -0
  51. package/src/runs/shared/mcp-direct-tool-grant.ts +2 -5
  52. package/src/runs/shared/model-fallback.ts +61 -17
  53. package/src/runs/shared/mutation-evidence.ts +52 -3
  54. package/src/runs/shared/permissions.ts +1 -1
  55. package/src/runs/shared/pi-args.ts +47 -1
  56. package/src/runs/shared/single-output.ts +45 -18
  57. package/src/runs/shared/subagent-prompt-runtime.ts +20 -2
  58. package/src/runs/shared/tool-timeout.ts +1 -1
  59. package/src/runs/shared/workflow-graph.ts +16 -0
  60. package/src/shared/types.ts +74 -4
  61. package/src/shared/workflow-child-permit.ts +91 -0
  62. package/src/slash/prompt-template-bridge.ts +37 -1
  63. package/src/slash/slash-commands.ts +18 -26
  64. package/src/tui/fleet-status.ts +11 -3
  65. package/src/tui/render-helpers.ts +31 -0
  66. package/src/tui/render.ts +673 -145
  67. package/src/watchdog/change-signature.ts +40 -1
  68. package/src/workflows/host-command.ts +6 -1
  69. package/src/workflows/scripted-workflow.ts +206 -6
  70. package/src/workflows/workflow-child-summary.ts +1 -1
  71. package/src/workflows/workflow-receipt.ts +41 -4
  72. package/src/workflows/workflow-resources.ts +150 -0
@@ -8,8 +8,8 @@ Parent extensions may register a session-scoped, out-of-band ceiling through `pi
8
8
 
9
9
  ## When to Use
10
10
 
11
- - **Complex work orchestration**: use Fable mode as the default parent-agent loop for complex work. Complex means the task has multiple moving parts, unclear acceptance, cross-cutting code, meaningful user-visible impact, expensive or irreversible validation, broad review surface, or the user asks for orchestration. Lightweight one-off delegation can stay lightweight.
12
- - **Advisory review**: use fresh-context `reviewer` agents for adversarial code review, or fork to `oracle` when inherited decisions and drift matter
11
+ - **Complex work orchestration**: keep the parent on its ordinary strong default model. Delegate only when another child materially improves evidence, independent review, or isolated execution; omission failures are cheaper than unnecessary commissions. For hard orchestration or root-cause questions, use a top-reasoning model only as a bounded read-only critic/oracle escalation, never as an autonomous root. Complex means the task has multiple moving parts, unclear acceptance, cross-cutting code, meaningful user-visible impact, expensive or irreversible validation, broad review surface, or the user asks for orchestration. Lightweight one-off delegation can stay lightweight.
12
+ - **Advisory review**: use fresh-context `reviewer` agents for adversarial code review; fork to `oracle` only for rare escalation where inherited decisions, drift, model routing, root cause, or hard tradeoffs matter
13
13
  - **Implementation handoff**: have `oracle` advise, then `worker` implement only after an approved direction
14
14
  - **Recon and planning**: use `scout`, then write a plan when needed
15
15
  - **Parallel exploration**: run multiple non-conflicting tasks concurrently
@@ -20,7 +20,7 @@ Parent extensions may register a session-scoped, out-of-band ceiling through `pi
20
20
 
21
21
  ## Tool vs Slash Commands
22
22
 
23
- Agents use the `subagent(...)` tool with `workflowScript` for execution, and `action` for management, status, and control. Humans often use the slash-command layer instead:
23
+ Agents use the `subagent(...)` tool for execution, management, status, and control. Direct `{ agent, task }` execution is enough for one bounded child task; use `workflowScript` when the parent needs JavaScript control flow or data-dependent branching, keyed, parallel, sequential, retry, retained-resume, aggregate, or explicit staged-lane behavior (`runs.lanes`). Humans often use the slash-command layer instead:
24
24
 
25
25
  - `/run` — launch a single agent
26
26
  - `workflowScript` — the sole public surface for sequence, parallelism, branching, retries, and aggregation
@@ -51,11 +51,15 @@ Packaged prompt shortcuts are also available for repeatable workflows. Treat the
51
51
 
52
52
  The prompt templates in `prompts/` encode workflows the parent agent can run on demand. If the user provides a URL, issue, PR, plan, local file, screenshot, or freeform target, treat that target as the primary scope: read or fetch it before launching children, then include it explicitly in every child task. For targets outside the parent cwd, include the exact repository, explicit `cwd`, authority boundary, and expected output path in each child task. Do not depend on the parent conversation history when the recipe calls for fresh context.
53
53
 
54
+ ### Commission-risk and cold-start packets
55
+
56
+ Delegate only when the child materially improves evidence, independent review, or isolated execution; do not manufacture parallelism. Every child packet must be cold-start complete: state the goal, exact target/cwd/ref, authority and edit boundary, relevant context/evidence, success criteria, validation, output, and stop/escalation rules. For an orchestration audit by the critic tier, make the child read-only and request at most three omissions, each cited to a file, line, or decision; high thinking is an explicit escalation, not a default.
57
+
54
58
  ### Council Mode technique
55
59
 
56
- Use Council Mode when the user asks to convene advisors, debate a material decision, cross-examine recommendations, or critique and improve a plan with several model perspectives. This includes requests such as “run a council on this architecture,” “have Sol, Fable, and Kimi critique this plan,” or “get multiple oracles to debate the tradeoffs.” Read `../council-mode/SKILL.md` and follow its bounded parent-supervised protocol instead of launching ad hoc parallel oracle calls.
60
+ Use Council Mode when the user asks to convene advisors, debate a material decision, cross-examine recommendations, or critique and improve a plan with several model perspectives. This includes requests such as “run a council on this architecture,” “have the configured advisors critique this plan,” or “get multiple oracles to debate the tradeoffs.” Read `../council-mode/SKILL.md` and follow its bounded parent-supervised protocol instead of launching ad hoc parallel oracle calls.
57
61
 
58
- Council advisors are read-only. User or project `council-*` profiles can pin models such as GPT 5.6 Sol, Fable, or Kimi and define any persistent stance in the profile body. Package advisors such as Surf's `gpt-pro` can join the roster only when the `surf-cli` Pi extension is installed and its `surf-oracle` provider is registered; treat them as external runners, omit child `async` for attached results, and do not pass `outputSchema` to them. The council question and scope provide the decision frame; do not invent per-advisor role labels. The parent collects independent reports, optionally sends curated cross-exam packets, and writes the final memo. Do not treat the council as agent-to-agent chat, implementation authority, or a writer swarm.
62
+ Council advisors are read-only. User or project `council-*` profiles choose allowed models and define any persistent stance in the profile body. A top-reasoning advisor remains bounded and read-only; it does not become the root. Package advisors such as Surf's `gpt-pro` can join the roster only when the `surf-cli` Pi extension is installed and its `surf-oracle` provider is registered; treat them as external runners, omit child `async` for attached results, and do not pass `outputSchema` to them. The council question and scope provide the decision frame; do not invent per-advisor role labels. The parent collects independent reports, optionally sends curated cross-exam packets, and writes the final memo. Do not treat the council as agent-to-agent chat, implementation authority, or a writer swarm.
59
63
 
60
64
  ### Parallel review technique
61
65
 
@@ -111,6 +115,12 @@ Use this after implementation when the user wants cleanup review or when a final
111
115
 
112
116
  Use this when a broad diff has known reviewer findings across several items and the user wants the parent to “orchestrate subagents like a boss.” Keep the active worktree safe with a three-stage `workflowScript`:
113
117
 
118
+ When staged seams are available, a low-tier writer should not receive the
119
+ end-to-end issue. Use `runs.lanes` inside `workflowScript` to keep stages narrow:
120
+ a scout/red test, helper-only change, one render seam, validation, minimality
121
+ challenge, or fresh review. Give the writer only its assigned implementation
122
+ stage; keep sequencing and synthesis with the parent.
123
+
114
124
  1. A parallel read-only planning fanout, one reviewer per issue cluster. Each child inspects the real diff and returns exact files, line refs, proposed fixes, and focused validation. They must not edit.
115
125
  2. One writer worker. It receives the reviewer summaries as the awaited planning results (or their durable output paths) interpolated into its task, plus the parent’s accepted scope, stop rules, and verification contract. It is the only child allowed to edit the active worktree.
116
126
  3. A parallel read-only validation fanout. Validators inspect the worker diff from fresh context with distinct angles, report pass/fail, remaining blockers, and missing verification.
@@ -161,19 +171,19 @@ subagent({
161
171
  Builtin agents load at the lowest priority. Project agents override user agents,
162
172
  and user/project agents override builtins with the same name.
163
173
 
164
- | Agent | Purpose | Model | Typical output / role |
174
+ | Agent | Purpose | Recommended tier | Typical output / role |
165
175
  |-------|---------|-------|------------------------|
166
- | `scout` | Fast codebase recon | inherits default | Writes `context.md` handoff material |
167
- | `worker` | Implementation and approved oracle handoffs | inherits default | Single-writer implementation with decision escalation |
168
- | `reviewer` | Review specialist | inherits default | Default recipes are review-only; tools include edit/write when a fix pass is explicit |
169
- | `researcher` | Web research brief generator | inherits default | Writes `research.md` |
170
- | `delegate` | Lightweight generic delegate | inherits default | No fixed output; generic delegated work |
171
- | `oracle` | Decision-consistency advisory review | inherits default | Advisory review, intercom coordination |
172
- | `advisor` | Claude Code-compatible alias for `oracle` | inherits default | Same advisory role as `oracle` |
176
+ | `scout` | Fast codebase recon | fast worker/scout tier | Writes `context.md` handoff material |
177
+ | `worker` | Implementation and approved oracle handoffs | capable worker tier | Single-writer implementation with decision escalation |
178
+ | `reviewer` | Review specialist | strong reviewer tier; high thinking for serious reviews | Default recipes are review-only; tools include edit/write when a fix pass is explicit |
179
+ | `researcher` | Web research brief generator | inherits configured default | Writes `research.md` |
180
+ | `delegate` | Lightweight generic delegate | inherits configured default | No fixed output; generic delegated work |
181
+ | `oracle` | Rare hard-decision/root-cause escalation | top-reasoning critic tier, bounded read-only; high thinking escalation only | Advisory trajectory review, not routine code review |
182
+ | `advisor` | Compatibility alias for `oracle` | top-reasoning critic tier, bounded read-only; high thinking escalation only | Same advisory escalation role as `oracle` |
173
183
 
174
184
  Builtin `worker` and `delegate` use strict tool allowlists and do not inherit ambient parent extension tools. To give a child an extension tool, name it in `tools` and load its provider via `extensions`, a path-like `tools` entry, or `subagentOnlyExtensions`. Custom agents without an `extensions` field follow `subagents.defaultExtensions` when set.
175
185
 
176
- Builtin agents inherit the current Pi default model unless a run, user setting, project setting, or `subagents.defaultModel` overrides `model`. Set `subagents.defaultModel` when subagents should use a different default model than the parent session. Override builtin defaults before copying full agent files when a small tweak is enough.
186
+ Builtin agents inherit the current Pi default model unless a run, user setting, project setting, or `subagents.defaultModel` overrides `model`. The table records recommended tier routing, not shipped hard defaults; explicit run, user, or project settings still win. Keep the parent/orchestrator on the ordinary strong default model unless parent/user policy says otherwise. Override builtin defaults before copying full agent files when a small tweak is enough.
177
187
 
178
188
  Set `subagents.defaultThinking` to apply a shared thinking level to builtin, package, user, and project agents whose frontmatter leaves `thinking` unset. Project settings win over user settings; explicit frontmatter (including `thinking: false`), `agentOverrides.<name>.thinking`, and per-run overrides remain more specific. This setting affects child agents only and does not change the parent session's default thinking level.
179
189
 
@@ -188,7 +198,7 @@ Set `subagents.defaultThinking` to apply a shared thinking level to builtin, pac
188
198
  For one run, use inline config:
189
199
 
190
200
  ```text
191
- /run reviewer[model=anthropic/claude-sonnet-4] "Review this diff"
201
+ /run reviewer[model=provider/review-model] "Review this diff"
192
202
  ```
193
203
 
194
204
  For persistent tweaks, edit `subagents.agentOverrides` in user or project settings. User overrides apply everywhere. Project overrides apply only in that repo and win over user overrides. Use `/subagents-models` or `subagent({ action: "models" })` to inspect the live mapping after settings and overrides load.
@@ -202,18 +212,18 @@ Provider-scoped entries can layer on top of the default override for the active
202
212
  "worker": { "thinking": "medium" }
203
213
  },
204
214
  "agentOverridesByProvider": {
205
- "github-copilot": {
206
- "worker": { "model": "github-copilot/gpt-5-mini" }
215
+ "provider-a": {
216
+ "worker": { "model": "provider-a/fast-worker-model" }
207
217
  },
208
- "openrouter": {
209
- "worker": { "model": "openrouter/openai/gpt-5-mini" }
218
+ "provider-b": {
219
+ "worker": { "model": "provider-b/fast-worker-model" }
210
220
  }
211
221
  }
212
222
  }
213
223
  }
214
224
  ```
215
225
 
216
- Model ids do not have to be exact. Separator variations (`claude-haiku-4.5` vs `claude-haiku-4-5`), case (`Claude-Sonnet-4`), and optional trailing date stamps (`claude-haiku-4-5-20251001`) all resolve to the same registry model. Exact `provider/id` wins; a qualified `provider/model` never switches providers. To constrain subagents to a budget or compliance profile, set `subagents.modelScope: { enforce: true, allow: ["anthropic/*", "openai/gpt-5-*"] }` in user or project settings. Out-of-scope models you pass explicitly error and abort; models inherited from frontmatter, `subagents.defaultModel`, agent frontmatter, or the parent session only warn.
226
+ Model ids do not have to be exact. Separator variations (`fast.worker-v1` vs `fast-worker-v1`), case (`Strong-Review-Model`), and optional trailing date stamps all resolve to the same registry model. Exact `provider/id` wins; a qualified `provider/model` never switches providers. To constrain subagents to a budget or compliance profile, set `subagents.modelScope: { enforce: true, allow: ["approved-provider/*", "second-provider/approved-*"] }` in user or project settings. Out-of-scope models you pass explicitly error and abort; models inherited from frontmatter, `subagents.defaultModel`, agent frontmatter, or the parent session only warn.
217
227
 
218
228
  For model fleets, use the profile commands instead of hand-editing repeated overrides: `/subagents-refresh-provider-models <provider>`, `/subagents-generate-profiles <provider>`, `/subagents-load-profile <name>`, and `/subagents-check-profile <name>`. Profiles live under `~/.pi/agent/profiles/pi-subagents/` and replace only `settings.subagents` when loaded.
219
229
 
@@ -249,9 +259,9 @@ Direct settings example:
249
259
  "subagents": {
250
260
  "agentOverrides": {
251
261
  "reviewer": {
252
- "model": "anthropic/claude-sonnet-4",
262
+ "model": "provider/strong-review-model",
253
263
  "thinking": "high",
254
- "fallbackModels": ["openai-codex/gpt-5.6-luna:low"],
264
+ "fallbackModels": ["backup-provider/strong-review-model"],
255
265
  "acceptanceRole": "read-only"
256
266
  }
257
267
  }
@@ -269,14 +279,11 @@ agent with the same name only when you want a substantially different agent.
269
279
 
270
280
  ### Recommended model tiering (optional)
271
281
 
272
- When several providers are available, route agents by task shape instead of one model for everything:
282
+ Keep the parent/orchestrator on the ordinary strong default model because omission failures are cheaper than unnecessary commissions. Route workers and scouts to a fast, capable worker tier, and keep serious reviews on the strong reviewer tier at high thinking. Do not use `oracle` or a top-reasoning model as the routine fresh-review default. Use that tier only for bounded, read-only critic/oracle/root-cause audits after ordinary review, CI, bot, or source evidence is insufficient; critic-tier high thinking is escalation-only and never an autonomous root. Explicit parent/user model policy wins over these recommendations.
273
283
 
274
- 1. **Fast workhorse** — cheapest capable model at low thinking for recon, lookups, and mechanical edits (for example on `scout`).
275
- 2. **Standard well-scoped** — mid-tier model at medium thinking for most delegations: routine multi-file edits, focused reviews, straightforward implementation (for example on `worker`, `reviewer`, `delegate`).
276
- 3. **Deep but bounded** — top reasoning model at high thinking only for hard tasks that arrive with explicit goals and completion criteria; these models loop on vague goals (for example on oracle-style agents).
277
- 4. **Taste and intent** — a model that reads human intent well for ambiguous work: UX/design judgment, product tradeoffs, planning from vague requirements, writing quality.
284
+ Examples are illustrative, not requirements. Map these tiers to concrete models in user/project settings or a profile. A non-OpenAI setup should choose comparable available models by capability.
278
285
 
279
- Routing rule: use tiers 1–3 when the task is well-scoped; use tier 4 when scoping or judging is the task itself. Give tier-4 agents cross-provider `fallbackModels` so subscription usage limits degrade gracefully; fallback triggers automatically on rate-limit and overload errors. Note that forked context over an Anthropic parent transcript with signed thinking blocks forces the child's thinking off, so intent-tier agents work best with fresh context.
286
+ Use `fallbackModels` when a tier has provider quota or availability risk. Prefer fresh context for cross-provider children when inherited provider-specific reasoning blocks would force thinking off.
280
287
 
281
288
  If a provider rejects model IDs with thinking suffixes, use
282
289
  `subagents.disableThinking: true` in user or project settings to clear bundled
@@ -0,0 +1,73 @@
1
+ # Pi Subagents: Review And Validation
2
+
3
+ Generic review and delivery guidance for delegated work. This file does not encode private backlog, merge, or release policy.
4
+
5
+ ## Delivery loop
6
+
7
+ Use the smallest loop that proves the change:
8
+
9
+ 1. Inspect the source, diff, issue, or plan directly.
10
+ 2. Keep one writer for each cwd or worktree.
11
+ 3. Run focused validation that can fail for the changed behavior.
12
+ 4. Use fresh-context read-only review for substantial, risky, public, or hard-to-see changes.
13
+ 5. Apply only accepted findings inside the same writer boundary.
14
+ 6. Re-run affected validation and review only the changed blast radius.
15
+ 7. Inspect the final diff and evidence before parent acceptance.
16
+
17
+ Skip review ceremony for trivial wording, renames, or local-only probes when direct parent inspection is enough.
18
+
19
+ ## Review shape
20
+
21
+ | Situation | Shape |
22
+ | --- | --- |
23
+ | One coherent diff or one risk | one reviewer |
24
+ | Independent risks, such as correctness, tests, security, or UI | parallel reviewers with distinct contracts |
25
+ | Possible over-scope or needless complexity | same-writer challenge before fresh review |
26
+ | Material design tradeoff | council mode |
27
+
28
+ Reviewers are fresh-context by default. Use the ordinary `reviewer` role for routine code review. Forked oracle/advisor runs are escalation-only for parent-history, drift, root-cause, model-routing, or hard tradeoff evidence.
29
+
30
+ ## Finding disposition
31
+
32
+ The parent classifies each finding against current HEAD:
33
+
34
+ - **Valid blocker:** concrete failure, repro, security issue, contract mismatch, or source-proven regression. Fix now.
35
+ - **Valid non-blocker:** real but outside the delivery slice. Record or defer.
36
+ - **Stale:** fixed or absent at the reviewed head. Cite current evidence.
37
+ - **Invalid:** contradicted by source, tests, docs, or user-approved scope. Cite the contradiction.
38
+ - **Out of policy/scope:** needs unapproved product, architecture, authority, release, or public-repo action. Escalate.
39
+ - **Speculative:** no contract, repro, or reachable failure. Do not block.
40
+
41
+ A clean reviewer result is evidence, not publication authority.
42
+
43
+ ## Gate-failure triage
44
+
45
+ When validation fails:
46
+
47
+ 1. Confirm the run belongs to the exact head/ref under judgment.
48
+ 2. Read the focused failing logs first.
49
+ 3. Name the failing test, assertion, contract, or thread.
50
+ 4. Classify cause: current diff, stale test, environment/setup, or existing flake.
51
+ 5. Reproduce locally when practical with the narrowest command.
52
+ 6. Patch forward when the current diff caused it.
53
+ 7. For stale/flaky failures, collect proof before one rerun or residual-risk note.
54
+ 8. Re-run the affected command or exact-head gate after every fix.
55
+
56
+ For bot comments, classify each thread as valid, stale, invalid, or out of policy before assigning severity.
57
+
58
+ ## Final checklist
59
+
60
+ Before reporting delegated work as done, verify the relevant subset:
61
+
62
+ - final diff contains only intended files
63
+ - focused validation covers changed behavior
64
+ - substantial or risky changes have fresh-review evidence
65
+ - accepted findings are fixed and revalidated
66
+ - publication authority exists before push, comment, close, merge, deploy, or release
67
+ - external checks are exact-head when used as evidence
68
+ - handoff is durable before cleanup
69
+ - residual risks, skipped validation, and blocked decisions are explicit
70
+
71
+ ## Public/private boundary
72
+
73
+ For issue/PR backlogs, releases, merge queues, contributor credit, or repo-specific policy, load the matching user/project skill when available. Keep those rules out of this public package until intentionally released.