pi-subagents 0.42.0 → 0.43.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. package/CHANGELOG.md +35 -0
  2. package/README.md +3 -5
  3. package/package.json +1 -1
  4. package/skills/pi-subagents/SKILL.md +2 -2
  5. package/skills/pi-subagents/references/constraints-and-recipes.md +35 -32
  6. package/skills/pi-subagents/references/execution-controls.md +27 -43
  7. package/skills/pi-subagents/references/management-authoring-rpc.md +20 -3
  8. package/skills/pi-subagents/references/prompting-and-roles.md +16 -49
  9. package/src/agents/agent-refinements.ts +624 -0
  10. package/src/agents/agents.ts +0 -2
  11. package/src/agents/proactive-skills.ts +1 -1
  12. package/src/api/delegation.ts +1 -2
  13. package/src/extension/control-notices.ts +2 -2
  14. package/src/extension/fanout-child.ts +46 -22
  15. package/src/extension/index.ts +29 -12
  16. package/src/extension/public-execution.ts +71 -0
  17. package/src/extension/rpc.ts +7 -8
  18. package/src/extension/schemas.ts +17 -17
  19. package/src/extension/tool-description.ts +14 -14
  20. package/src/missions/actions.ts +47 -11
  21. package/src/missions/goal-driver.ts +162 -0
  22. package/src/missions/lifecycle.ts +44 -12
  23. package/src/missions/store.ts +68 -3
  24. package/src/missions/types.ts +25 -3
  25. package/src/missions/workflow-state.ts +77 -0
  26. package/src/profiles/profiles.ts +1 -3
  27. package/src/runs/background/async-execution.ts +3 -0
  28. package/src/runs/background/async-job-tracker.ts +2 -17
  29. package/src/runs/background/control-channel.ts +50 -6
  30. package/src/runs/background/retained-children.ts +68 -0
  31. package/src/runs/background/scheduled-runs.ts +17 -13
  32. package/src/runs/background/steering.ts +7 -5
  33. package/src/runs/background/subagent-runner.ts +10 -4
  34. package/src/runs/foreground/async-steering-action.ts +36 -10
  35. package/src/runs/foreground/chain-clarify.ts +13 -16
  36. package/src/runs/foreground/execution.ts +4 -0
  37. package/src/runs/foreground/subagent-executor.ts +200 -59
  38. package/src/runs/shared/acceptance.ts +154 -9
  39. package/src/runs/shared/subagent-prompt-runtime.ts +94 -25
  40. package/src/shared/types.ts +17 -3
  41. package/src/slash/delegation-adapters.ts +0 -13
  42. package/src/slash/prompt-template-bridge.ts +14 -12
  43. package/src/slash/prompt-workflows.ts +10 -4
  44. package/src/slash/slash-bridge.ts +8 -6
  45. package/src/slash/slash-commands.ts +27 -12
  46. package/src/slash/slash-live-state.ts +4 -2
  47. package/src/tui/fleet-status.ts +8 -4
  48. package/src/tui/fleet.ts +15 -5
  49. package/src/tui/render.ts +13 -31
  50. package/src/workflows/chat-progress.ts +2 -2
  51. package/src/workflows/scripted-workflow.ts +97 -10
  52. package/agents/context-builder.md +0 -46
  53. package/agents/planner.md +0 -56
  54. package/prompts/parallel-context-build.md +0 -55
  55. package/prompts/parallel-handoff-plan.md +0 -61
@@ -11,7 +11,7 @@ Parent extensions may register a session-scoped, out-of-band ceiling through `pi
11
11
  - **Complex work orchestration**: use Fable mode as the default parent-agent loop for complex work. Complex means the task has multiple moving parts, unclear acceptance, cross-cutting code, meaningful user-visible impact, expensive or irreversible validation, broad review surface, or the user asks for orchestration. Lightweight one-off delegation can stay lightweight.
12
12
  - **Advisory review**: use fresh-context `reviewer` agents for adversarial code review, or fork to `oracle` when inherited decisions and drift matter
13
13
  - **Implementation handoff**: have `oracle` advise, then `worker` implement only after an approved direction
14
- - **Recon and planning**: use `scout` or `context-builder`, then `planner`
14
+ - **Recon and planning**: use `scout`, then write a plan when needed
15
15
  - **Parallel exploration**: run multiple non-conflicting tasks concurrently
16
16
  - **Regular skill specialists**: when discovery shows proactive skill subagent suggestions and the current work is broad enough, launch a small fresh-context fanout that asks one subagent per relevant regularly used skill to apply that skill's perspective to the task
17
17
  - **Long-running work**: launch async/background runs and inspect them later. For mutation-capable work, bound the delivery slice and elapsed runtime, then request checkpoints after active tool work returns. Reserve hard turn and tool-call caps for explicitly read-only children.
@@ -20,8 +20,7 @@ Parent extensions may register a session-scoped, out-of-band ceiling through `pi
20
20
 
21
21
  ## Tool vs Slash Commands
22
22
 
23
- Agents can use the `subagent(...)` tool directly for execution, management, status, and control.
24
- Humans often use the slash-command layer instead:
23
+ Agents use the `subagent(...)` tool with `workflowScript` for execution, and `action` for management, status, and control. Humans often use the slash-command layer instead:
25
24
 
26
25
  - `/run` — launch a single agent
27
26
  - `workflowScript` — the sole public surface for sequence, parallelism, branching, retries, and aggregation
@@ -34,7 +33,7 @@ Humans often use the slash-command layer instead:
34
33
  - `/subagents-doctor` — diagnose setup, discovery, async paths, and intercom bridge state
35
34
  - `/subagents-models [agent]` — show the live runtime-loaded builtin model mapping
36
35
  - `/subagents-profiles`, `/subagents-load-profile`, `/subagents-refresh-provider-models`, `/subagents-generate-profiles`, `/subagents-check-profile` — manage model profiles and provider catalogs
37
- - `/prompt-workflow` — run a prompt template through native single-agent or workflowScript execution
36
+ - `/prompt-workflow` — run a prompt template through native workflowScript execution
38
37
 
39
38
  Prefer the tool when you are writing agent logic. Prefer the slash commands when
40
39
  you are guiding a human through an interactive flow.
@@ -43,8 +42,6 @@ Packaged prompt shortcuts are also available for repeatable workflows. Treat the
43
42
  - `/parallel-review` — fresh-context reviewers with distinct review angles, then synthesis
44
43
  - `/review-loop` — parent-orchestrated worker, fresh-reviewer, and fix-worker cycles until clean or capped
45
44
  - `/parallel-research` — combine `researcher` and `scout` for external evidence plus local code context
46
- - `/parallel-context-build` — parallel `context-builder` passes that produce planning handoff context and meta-prompts
47
- - `/parallel-handoff-plan` — external-reference research plus local `context-builder` passes, followed by a synthesis handoff plan and implementation-ready meta-prompt
48
45
  - `/gather-context-and-clarify` — scout/research first, then ask the user clarifying questions with `interview`
49
46
  - `/parallel-cleanup` — two fresh-context reviewers (deslop + verbosity passes) for an adversarial cleanup review of the current diff
50
47
 
@@ -71,18 +68,20 @@ Example shape:
71
68
 
72
69
  ```typescript
73
70
  subagent({
74
- tasks: [
75
- { agent: "reviewer", task: "Apply the available 'deslop' skill to review the current diff for concrete cleanup findings only. Do not modify files.", skill: "deslop" },
76
- { agent: "reviewer", task: "Apply the available 'accessibility' skill to review the UI changes for concrete issues only. Do not modify files.", skill: "accessibility" }
77
- ],
78
- context: "fresh",
79
- concurrency: 2
71
+ workflowScript: `
72
+ const results = await runs.all([
73
+ { key: "deslop", agent: "reviewer", task: "Apply the available 'deslop' skill to review the current diff for concrete cleanup findings only. Do not modify files.", skill: "deslop" },
74
+ { key: "accessibility", agent: "reviewer", task: "Apply the available 'accessibility' skill to review the UI changes for concrete issues only. Do not modify files.", skill: "accessibility" }
75
+ ]);
76
+ return results.map(result => result.output);
77
+ `,
78
+ context: "fresh"
80
79
  })
81
80
  ```
82
81
 
83
82
  ### Review-loop technique
84
83
 
85
- Use this when the user wants implementation or current diff review to continue until reviewers stop finding fixes worth doing now. Keep the loop in the parent session: one async `worker` implements or fixes, fresh-context `reviewer` agents inspect the actual repo and diff, the parent synthesizes accepted fixes, and one async forked `worker` applies them. The parent can express the sequence up front as an async/background chain when the workflow is known, or continue with explicit follow-up subagent runs after each async completion. For an initial chain, pass `async: true` so the main chat is unblocked; do not set `clarify: true` unless the user explicitly wants the foreground clarify UI. Treat an async implementation worker handoff as an intermediate state, not final completion, unless the user explicitly asked for worker-only work, review-only output, or to stop after implementation. Stop when reviewers find no blockers or fixes worth doing now, remaining feedback is optional or deferred, an unapproved product/scope/architecture decision appears, or the max review-round cap is reached. Default to 3 review rounds unless the user sets a different cap. Do not loop for optional polish, and do not let children launch subagents or decide the loop outcome.
84
+ Use this when the user wants implementation or current diff review to continue until reviewers stop finding fixes worth doing now. Keep the loop in the parent session: one async `worker` implements or fixes, fresh-context `reviewer` agents inspect the actual repo and diff, the parent synthesizes accepted fixes, and one async forked `worker` applies them. The parent can express the sequence up front as an async/background `workflowScript` when the workflow is known, or continue with explicit follow-up workflowScript runs after each async completion. For an initial workflow, pass `async: true` so the main chat is unblocked. Treat an async implementation worker handoff as an intermediate state, not final completion, unless the user explicitly asked for worker-only work, review-only output, or to stop after implementation. Stop when reviewers find no blockers or fixes worth doing now, remaining feedback is optional or deferred, an unapproved product/scope/architecture decision appears, or the max review-round cap is reached. Default to 3 review rounds unless the user sets a different cap. Do not loop for optional polish, and do not let children launch subagents or decide the loop outcome.
86
85
 
87
86
  As a conservative orchestration policy, do not pass `turnBudget` or a hard `toolBudget` to an implementation worker, fix worker, reviewer with edit authority, or other mutation-capable child. The default tool budget blocks read/search tools rather than mutation tools, but count limits still do not measure delivery safety. Use a narrow task plus an outer elapsed deadline with enough margin, then request a checkpoint after the current tool returns. The checkpoint should report changed files, build/test state, remaining work, and commit or PR state. An elapsed timeout is not a mutation-safe boundary and must not be used as the checkpoint trigger.
88
87
 
@@ -90,37 +89,7 @@ As a conservative orchestration policy, do not pass `turnBudget` or a hard `tool
90
89
 
91
90
  Use this when the question needs both external evidence and local implications. Combine `researcher` for official docs, specs, ecosystem behavior, recent changes, benchmarks, and primary sources with `scout` for repository files, patterns, constraints, tests, and likely integration points. Give each child a distinct angle: external evidence, local code context, and practical tradeoffs. Ask for source links or file ranges, confidence level, gaps, and decision implications. Do not ask these children to edit unless implementation was explicitly requested.
92
91
 
93
- ### Parallel context-build technique
94
-
95
- Use this before planning or implementation when a stronger handoff is needed. Use `workflowScript` with `runs.all` to launch distinct `context-builder` lanes, each with an explicit output path. Give every task a distinct output path such as `context-build/request-and-scope.md`, `context-build/codebase-and-patterns.md`, and `context-build/validation-and-risks.md`. Choose two or three builders: request/scope, codebase/patterns, and validation/risks. Each builder must read every relevant file needed to understand its slice, follow imports/callers/tests/docs/config, conduct tool-available web research when needed, and include a compact `meta-prompt` section. The parent synthesizes the outputs into important context, recommended next meta-prompt, open questions, assumptions, and artifact paths.
96
-
97
- Example shape:
98
-
99
- ```typescript
100
- subagent({ workflowScript: `
101
- const results = await runs.all([
102
- { key: "lane-a", agent: "reviewer", task: "Inspect lane A" },
103
- { key: "lane-b", agent: "reviewer", task: "Inspect lane B" }
104
- ]);
105
- return results.map(result => result.output);
106
- ` })
107
- ```
108
92
 
109
- ### Parallel handoff-plan technique
110
-
111
- Use this when the user needs a solution brief or implementation-ready handoff from an external reference plus local code context, such as “study this library behavior, inspect our codebase, then produce a worker prompt.” Run a chain with a first parallel group and a second synthesis `context-builder` step. The first group usually includes `researcher` for external projects/docs/prompt guidance and `context-builder` for local code context; add a second `context-builder` for implementation strategy only when the scope is large enough to benefit. Use distinct output paths under `handoff/`, then have the synthesis `context-builder` read those outputs and write `handoff/final-handoff-plan.md` with the recommended approach, likely files, constraints, non-goals, validation, risks, unresolved questions, and final compact implementation-ready meta-prompt.
112
-
113
- Example shape:
114
-
115
- ```typescript
116
- subagent({ workflowScript: `
117
- const results = await runs.all([
118
- { key: "lane-a", agent: "reviewer", task: "Inspect lane A" },
119
- { key: "lane-b", agent: "reviewer", task: "Inspect lane B" }
120
- ]);
121
- return results.map(result => result.output);
122
- ` })
123
- ```
124
93
 
125
94
  ### Gather-context-and-clarify technique
126
95
 
@@ -134,11 +103,11 @@ Use this after implementation when the user wants cleanup review or when a final
134
103
 
135
104
  Use this when a broad diff has known reviewer findings across several items and the user wants the parent to “orchestrate subagents like a boss.” Keep the active worktree safe with a three-stage chain:
136
105
 
137
- 1. A parallel read-only planning fanout, one planner/reviewer per issue cluster. Each child inspects the real diff and returns exact files, line refs, proposed fixes, and focused validation. They must not edit.
138
- 2. One writer worker. It receives the planner summaries through `{previous}`, the parent’s accepted scope, stop rules, and verification contract. It is the only child allowed to edit the active worktree.
106
+ 1. A parallel read-only planning fanout, one reviewer per issue cluster. Each child inspects the real diff and returns exact files, line refs, proposed fixes, and focused validation. They must not edit.
107
+ 2. One writer worker. It receives the reviewer summaries through `{previous}`, the parent’s accepted scope, stop rules, and verification contract. It is the only child allowed to edit the active worktree.
139
108
  3. A parallel read-only validation fanout. Validators inspect the worker diff from fresh context with distinct angles, report pass/fail, remaining blockers, and missing verification.
140
109
 
141
- Prefer `async: true`, `context: "fresh"` for planners/validators, `outputMode: "file-only"` for large summaries, and per-stage output names that will not collide. Add `phase` and `label` to make async status readable, and use `as` plus `{outputs.name}` when a later step needs a specific earlier result instead of the whole `{previous}` blob. Use this pattern instead of launching several writer workers into a dirty worktree. Include non-blocking suggestions in the writer prompt only when they are small, safe, and do not expand product scope; otherwise record them as deferred.
110
+ Prefer `async: true`, `context: "fresh"` for reviewers/validators, `outputMode: "file-only"` for large summaries, and per-stage output names that will not collide. Add `phase` and `label` to make async status readable, and use `as` plus `{outputs.name}` when a later step needs a specific earlier result instead of the whole `{previous}` blob. Use this pattern instead of launching several writer workers into a dirty worktree. Include non-blocking suggestions in the writer prompt only when they are small, safe, and do not expand product scope; otherwise record them as deferred.
142
111
 
143
112
  When one child returns a structured target list, use ordinary JavaScript to validate/filter it and map bounded entries into `runs.all`; do not use the removed chain fanout DSL.
144
113
 
@@ -171,10 +140,8 @@ and user/project agents override builtins with the same name.
171
140
  | Agent | Purpose | Model | Typical output / role |
172
141
  |-------|---------|-------|------------------------|
173
142
  | `scout` | Fast codebase recon | inherits default | Writes `context.md` handoff material |
174
- | `planner` | Creates implementation plans | inherits default | Writes `plan.md` |
175
143
  | `worker` | Implementation and approved oracle handoffs | inherits default | Single-writer implementation with decision escalation |
176
144
  | `reviewer` | Review specialist | inherits default | Default recipes are review-only; tools include edit/write when a fix pass is explicit |
177
- | `context-builder` | Requirements/codebase handoff builder | inherits default | Writes structured context files |
178
145
  | `researcher` | Web research brief generator | inherits default | Writes `research.md` |
179
146
  | `delegate` | Lightweight generic delegate | inherits default | No fixed output; generic delegated work |
180
147
  | `oracle` | Decision-consistency advisory review | inherits default | Advisory review, intercom coordination |
@@ -260,7 +227,7 @@ When several providers are available, route agents by task shape instead of one
260
227
 
261
228
  1. **Fast workhorse** — cheapest capable model at low thinking for recon, lookups, and mechanical edits (for example on `scout`).
262
229
  2. **Standard well-scoped** — mid-tier model at medium thinking for most delegations: routine multi-file edits, focused reviews, straightforward implementation (for example on `worker`, `reviewer`, `delegate`).
263
- 3. **Deep but bounded** — top reasoning model at high thinking only for hard tasks that arrive with explicit goals and completion criteria; these models loop on vague goals (for example on `planner` and oracle-style agents).
230
+ 3. **Deep but bounded** — top reasoning model at high thinking only for hard tasks that arrive with explicit goals and completion criteria; these models loop on vague goals (for example on oracle-style agents).
264
231
  4. **Taste and intent** — a model that reads human intent well for ambiguous work: UX/design judgment, product tradeoffs, planning from vague requirements, writing quality.
265
232
 
266
233
  Routing rule: use tiers 1–3 when the task is well-scoped; use tier 4 when scoping or judging is the task itself. Give tier-4 agents cross-provider `fallbackModels` so subscription usage limits degrade gracefully; fallback triggers automatically on rate-limit and overload errors. Note that forked context over an Anthropic parent transcript with signed thinking blocks forces the child's thinking off, so intent-tier agents work best with fresh context.