@selesai/code 0.13.2 → 0.13.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (120) hide show
  1. package/CHANGELOG.md +20 -0
  2. package/dist/defaults/models.json +85 -13
  3. package/dist/extensions/cost-reconcile.test.ts +200 -4
  4. package/dist/extensions/cost-reconcile.ts +88 -102
  5. package/dist/extensions/pi-intercom/CHANGELOG.md +13 -0
  6. package/dist/extensions/pi-intercom/README.md +4 -5
  7. package/dist/extensions/pi-intercom/config.test.ts +3 -31
  8. package/dist/extensions/pi-intercom/config.ts +0 -15
  9. package/dist/extensions/pi-intercom/index.ts +9 -46
  10. package/dist/extensions/pi-intercom/intercom.integration.test.ts +49 -57
  11. package/dist/extensions/pi-intercom/package.json +1 -1
  12. package/dist/extensions/pi-intercom/reply-tracker.test.ts +20 -0
  13. package/dist/extensions/pi-intercom/reply-tracker.ts +8 -0
  14. package/dist/extensions/pi-subagents/CHANGELOG.md +27 -0
  15. package/dist/extensions/pi-subagents/docs/tool-reference.md +4 -1
  16. package/dist/extensions/pi-subagents/docs/workflows.md +2 -2
  17. package/dist/extensions/pi-subagents/package-lock.json +2 -2
  18. package/dist/extensions/pi-subagents/package.json +1 -1
  19. package/dist/extensions/pi-subagents/skills/council-mode/SKILL.md +48 -243
  20. package/dist/extensions/pi-subagents/skills/council-mode/references/pass-contracts.md +150 -0
  21. package/dist/extensions/pi-subagents/skills/pi-subagents/SKILL.md +87 -37
  22. package/dist/extensions/pi-subagents/skills/pi-subagents/references/constraints-and-recipes.md +29 -233
  23. package/dist/extensions/pi-subagents/skills/pi-subagents/references/execution-controls.md +49 -8
  24. package/dist/extensions/pi-subagents/skills/pi-subagents/references/management-authoring-rpc.md +2 -2
  25. package/dist/extensions/pi-subagents/skills/pi-subagents/references/multi-lane-orchestration.md +13 -1
  26. package/dist/extensions/pi-subagents/skills/pi-subagents/references/prompting-and-roles.md +34 -27
  27. package/dist/extensions/pi-subagents/skills/pi-subagents/references/review-and-validation.md +73 -0
  28. package/dist/extensions/pi-subagents/src/agents/agent-management.ts +157 -28
  29. package/dist/extensions/pi-subagents/src/api/shared-types.ts +2 -0
  30. package/dist/extensions/pi-subagents/src/extension/public-execution.ts +1 -0
  31. package/dist/extensions/pi-subagents/src/extension/schemas.ts +1 -0
  32. package/dist/extensions/pi-subagents/src/extension/tool-description.ts +4 -1
  33. package/dist/extensions/pi-subagents/src/runs/background/async-execution.ts +2 -2
  34. package/dist/extensions/pi-subagents/src/runs/background/async-job-tracker.ts +3 -0
  35. package/dist/extensions/pi-subagents/src/runs/background/async-status.ts +45 -2
  36. package/dist/extensions/pi-subagents/src/runs/background/control-channel.ts +3 -2
  37. package/dist/extensions/pi-subagents/src/runs/background/run-status.ts +13 -2
  38. package/dist/extensions/pi-subagents/src/runs/background/subagent-runner.ts +5 -1
  39. package/dist/extensions/pi-subagents/src/runs/background/subagent-wait.ts +10 -2
  40. package/dist/extensions/pi-subagents/src/runs/background/wait-completions.ts +3 -0
  41. package/dist/extensions/pi-subagents/src/runs/foreground/execution.ts +11 -2
  42. package/dist/extensions/pi-subagents/src/runs/foreground/subagent-executor.ts +98 -1
  43. package/dist/extensions/pi-subagents/src/runs/shared/async-status-projection.ts +138 -4
  44. package/dist/extensions/pi-subagents/src/runs/shared/background-process-options.ts +9 -0
  45. package/dist/extensions/pi-subagents/src/runs/shared/mcp-direct-tool-grant.ts +2 -5
  46. package/dist/extensions/pi-subagents/src/runs/shared/mutation-evidence.ts +52 -3
  47. package/dist/extensions/pi-subagents/src/runs/shared/pi-args.ts +47 -1
  48. package/dist/extensions/pi-subagents/src/runs/shared/single-output.ts +45 -18
  49. package/dist/extensions/pi-subagents/src/runs/shared/subagent-prompt-runtime.ts +21 -2
  50. package/dist/extensions/pi-subagents/src/runs/shared/workflow-graph.ts +15 -0
  51. package/dist/extensions/pi-subagents/src/shared/types.ts +34 -1
  52. package/dist/extensions/pi-subagents/src/tui/fleet-status.ts +11 -3
  53. package/dist/extensions/pi-subagents/src/tui/render-helpers.ts +31 -0
  54. package/dist/extensions/pi-subagents/src/tui/render.ts +597 -112
  55. package/dist/extensions/pi-subagents/src/watchdog/change-signature.ts +40 -1
  56. package/dist/extensions/pi-subagents/src/workflows/host-command.ts +6 -1
  57. package/dist/extensions/pi-subagents/src/workflows/scripted-workflow.ts +53 -2
  58. package/dist/extensions/pi-subagents/test/integration/async-execution.test.ts +55 -6
  59. package/dist/extensions/pi-subagents/test/integration/async-status.test.ts +111 -1
  60. package/dist/extensions/pi-subagents/test/integration/render-fork-badge.test.ts +206 -38
  61. package/dist/extensions/pi-subagents/test/integration/render-widget.test.ts +522 -31
  62. package/dist/extensions/pi-subagents/test/integration/single-execution.test.ts +123 -0
  63. package/dist/extensions/pi-subagents/test/unit/agent-management.test.ts +48 -0
  64. package/dist/extensions/pi-subagents/test/unit/async-status-projection.test.ts +57 -1
  65. package/dist/extensions/pi-subagents/test/unit/background-process-options.test.ts +17 -0
  66. package/dist/extensions/pi-subagents/test/unit/external-cli-runner.test.ts +1 -1
  67. package/dist/extensions/pi-subagents/test/unit/fleet-status.test.ts +44 -4
  68. package/dist/extensions/pi-subagents/test/unit/fork-cache-key.test.ts +91 -0
  69. package/dist/extensions/pi-subagents/test/unit/host-command.test.ts +1 -0
  70. package/dist/extensions/pi-subagents/test/unit/index-child-registration.test.ts +0 -1
  71. package/dist/extensions/pi-subagents/test/unit/mcp-direct-tool-grant.test.ts +20 -3
  72. package/dist/extensions/pi-subagents/test/unit/mutation-evidence.test.ts +27 -0
  73. package/dist/extensions/pi-subagents/test/unit/pi-args.test.ts +93 -17
  74. package/dist/extensions/pi-subagents/test/unit/public-execution.test.ts +1 -0
  75. package/dist/extensions/pi-subagents/test/unit/render-helpers.test.ts +103 -26
  76. package/dist/extensions/pi-subagents/test/unit/run-status.test.ts +58 -0
  77. package/dist/extensions/pi-subagents/test/unit/schemas.test.ts +22 -2
  78. package/dist/extensions/pi-subagents/test/unit/scripted-workflow.test.ts +21 -0
  79. package/dist/extensions/pi-subagents/test/unit/single-output.test.ts +13 -0
  80. package/dist/extensions/pi-subagents/test/unit/subagent-wait.test.ts +54 -0
  81. package/dist/extensions/pi-subagents/test/unit/tool-description.test.ts +2 -0
  82. package/dist/extensions/pi-subagents/test/unit/wait-completions.test.ts +32 -0
  83. package/dist/extensions/pi-subagents/test/unit/watchdog-change-signature.test.ts +48 -1
  84. package/dist/extensions/pi-subagents/test/unit/widget-nested-render.test.ts +21 -6
  85. package/dist/extensions/pi-subagents/test/unit/windows-hide-spawn.test.ts +11 -0
  86. package/dist/extensions/pi-web-agent/CHANGELOG.md +410 -0
  87. package/dist/extensions/pi-web-agent/README.md +131 -0
  88. package/dist/extensions/pi-web-agent/package.json +4 -2
  89. package/dist/extensions/pi-web-agent/src/backends/config.ts +62 -5
  90. package/dist/extensions/pi-web-agent/src/backends/doctor.ts +136 -0
  91. package/dist/extensions/pi-web-agent/src/backends/factory.ts +96 -6
  92. package/dist/extensions/pi-web-agent/src/commands/web-agent-config.ts +183 -47
  93. package/dist/extensions/pi-web-agent/src/extension.ts +62 -25
  94. package/dist/extensions/pi-web-agent/src/extract/readability.ts +19 -11
  95. package/dist/extensions/pi-web-agent/src/orchestration/candidate-selector.ts +5 -4
  96. package/dist/extensions/pi-web-agent/src/orchestration/direct-url.ts +2 -25
  97. package/dist/extensions/pi-web-agent/src/orchestration/evidence-quality.ts +5 -2
  98. package/dist/extensions/pi-web-agent/src/orchestration/evidence-ranker.ts +2 -0
  99. package/dist/extensions/pi-web-agent/src/orchestration/research-orchestrator.ts +72 -7
  100. package/dist/extensions/pi-web-agent/src/orchestration/research-types.ts +8 -1
  101. package/dist/extensions/pi-web-agent/src/orchestration/research-worker.ts +28 -3
  102. package/dist/extensions/pi-web-agent/src/orchestration/source-profile.ts +4 -0
  103. package/dist/extensions/pi-web-agent/src/orchestration/url.ts +35 -0
  104. package/dist/extensions/pi-web-agent/src/presentation/explore-presentation.ts +18 -6
  105. package/dist/extensions/pi-web-agent/src/presentation/search-presentation.ts +14 -2
  106. package/dist/extensions/pi-web-agent/src/readers/github-reader.ts +150 -0
  107. package/dist/extensions/pi-web-agent/src/readers/limits.ts +3 -0
  108. package/dist/extensions/pi-web-agent/src/readers/pdf-reader.ts +87 -0
  109. package/dist/extensions/pi-web-agent/src/readers/resolver.ts +25 -0
  110. package/dist/extensions/pi-web-agent/src/readers/types.ts +11 -0
  111. package/dist/extensions/pi-web-agent/src/readers/youtube-reader.ts +79 -0
  112. package/dist/extensions/pi-web-agent/src/search/duckduckgo.ts +32 -5
  113. package/dist/extensions/pi-web-agent/src/search/exa.ts +109 -0
  114. package/dist/extensions/pi-web-agent/src/search/fanout.ts +154 -0
  115. package/dist/extensions/pi-web-agent/src/search/tavily.ts +113 -0
  116. package/dist/extensions/pi-web-agent/src/search/youcom.ts +109 -0
  117. package/dist/extensions/pi-web-agent/src/tools/web-search.ts +22 -9
  118. package/dist/extensions/pi-web-agent/src/types.ts +19 -4
  119. package/dist/extensions/tokenin-onboarding.ts +302 -0
  120. package/package.json +1 -1
@@ -6,8 +6,8 @@ This file is a detailed reference loaded from `skills/pi-subagents/SKILL.md`.
6
6
 
7
7
  Agent files can live in:
8
8
  - `~/.selesai/agent/agents/**/*.md` — user scope
9
- - `.selesai/agents/**/*.md` — canonical project scope
10
- - legacy `.agents/**/*.md` — still read for compatibility, but `.selesai/agents/` wins on conflicts
9
+ - `.pi/agents/**/*.md` — canonical project scope
10
+ - legacy `.agents/**/*.md` — still read for compatibility, but `.pi/agents/` wins on conflicts
11
11
 
12
12
  Saved chain files may still be discovered for management and existing durable run state, but they are not a public execution surface. Author new orchestration with `workflowScript`.
13
13
 
@@ -28,7 +28,7 @@ External CLI profiles are async-only and one-shot. They support lifecycle artifa
28
28
 
29
29
  ### External job profiles
30
30
 
31
- An agent may set `runner.type: external-job` with a non-empty `provider` and optional JSON `options`. When `surf-cli` is installed and loaded, Surf can optionally expose a `gpt-pro` package agent through provider `surf-oracle`. Surf maps `model: pro` to ChatGPT GPT-5.6 Sol Pro web mode. pi-subagents does not own that package agent or model mapping. Remove any old `agentOverrides.gpt-pro.disabled` workaround before using Surf's package agent. The provider must be registered in the host Pi process through `pi-subagents/external-job-provider`; the async runner talks to that parent-owned registry through a local operation bridge.
31
+ An agent may set `runner.type: external-job` with a non-empty `provider` and optional JSON `options`. When `surf-cli` is installed and loaded, Surf can optionally expose a `gpt-pro` package agent through provider `surf-oracle`. Surf maps `model: pro` to its configured pro web mode. pi-subagents does not own that package agent or model mapping. Remove any old `agentOverrides.gpt-pro.disabled` workaround before using Surf's package agent. The provider must be registered in the host Pi process through `pi-subagents/external-job-provider`; the async runner talks to that parent-owned registry through a local operation bridge.
32
32
 
33
33
  External job profiles are async-only. The provider owns the remote job and Pi owns the async run record. Status persists provider name, provider job id, prompt digest, provider options, handle/conversation URLs when supplied, result artifact path, last known state, and provider failure code/message. Recovery uses existing provider job metadata to call `reattach` and `result`; it refuses to redispatch a prompt when the persisted provider job does not match the prompt digest.
34
34
 
@@ -38,10 +38,17 @@ External job profiles do not support foreground/clarify, steer/resume, Pi models
38
38
 
39
39
  ```typescript
40
40
  subagent({
41
- workflowScript: `return runs.run("oracle-check", { agent: "oracle", task: "Review my current direction and challenge assumptions." })`
41
+ agent: "oracle",
42
+ task: "Review my current direction and challenge assumptions."
42
43
  })
43
44
  ```
44
45
 
46
+ Use direct single-agent execution for one bounded task when no stable key,
47
+ branching, retained-child lookup, or aggregate workflow result is needed. Use a
48
+ `workflowScript` when the parent needs JavaScript control flow or data-dependent
49
+ branching, or when the run is part of a larger coordinated wave or a later step
50
+ must resume it by key.
51
+
45
52
  ### Forked context
46
53
 
47
54
  ```typescript
@@ -61,7 +68,17 @@ its resolved launch context as `[fresh]` or `[fork]`. Aggregate headers show
61
68
 
62
69
  ### Scripted workflows
63
70
 
64
- `workflowScript` is the sole public execution surface. Use `runs.run(key, { agent, task, ... })` for one child, `runs.all([...])` for parallel children, and ordinary JavaScript for sequence, branching, filtering, retries, and aggregation. Scripts are ordinary JavaScript statement bodies, so use an explicit return such as `workflowScript: "return runs.run('main', { agent: 'worker', task: '...' })"` for a useful one-child result. Use top-level `await`, plain helper functions, or explicit Promise chains; nested `async function` helpers, async arrows, and async methods are rejected. Prefer a single scripted workflow whenever the parent is starting a coordinated wave, such as multiple reviews, review plus gate monitor, worker then monitor setup, cross-repo prep lanes, or a fanout that the parent will consume together.
71
+ `workflowScript` is the public composition surface when the parent needs
72
+ JavaScript control flow or data-dependent branching. Use
73
+ `runs.run(key, { agent, task, ... })` for keyed children, `runs.all([...])` for
74
+ parallel children, and ordinary JavaScript for sequence, filtering, retries,
75
+ and aggregation. Scripts are ordinary JavaScript statement bodies, so use an
76
+ explicit return such as `return runs.run("main", { agent: "worker", task: "..." })` for a useful one-child result. Use top-level `await`,
77
+ plain helper functions, or explicit Promise chains; nested `async function`
78
+ helpers, async arrows, and async methods are rejected. Prefer a single scripted
79
+ workflow whenever the parent is starting a coordinated wave, such as multiple
80
+ reviews, review plus gate monitor, worker then monitor setup, cross-repo prep
81
+ lanes, or a fanout that the parent will consume together.
65
82
 
66
83
  ```js
67
84
  subagent({
@@ -82,6 +99,8 @@ If `runs.all` is missing in a running session, reload or update `pi-subagents` b
82
99
 
83
100
  For one host-run verification command, pass `gate: "npm test"` on a `runs.run`/`runs.all` item (or at the top level as a workflow default). It is shorthand for verified acceptance with that single command: the runtime executes it on the host, records the result as evidence, and memoizes it per tracked workspace state and effective environment. `gate` cannot be combined with `acceptance`; use explicit `acceptance.verify` for multiple commands or custom criteria.
84
101
 
102
+ If omitted, acceptance is inferred from role, mode, and risk. Use `level: "checked"` for ordinary writer evidence and `level: "verified"` when the runtime should run explicit validation commands. Independent review is orthogonal: use `review: { required: true, agent: "reviewer" }`; reviewer/read-only calls omit `acceptance`. `review-required` means evidence passed but review is pending; `reviewed` means an independent review found no blockers. Never request `level: "reviewed"`; it is recognized only so preflight can return an actionable correction. Disable gates with `{ level: "none", reason: "..." }`; bare `"none"` is rejected and `false` is only a deprecated shorthand. Child-reported command success is evidence, not runtime verification.
103
+
85
104
  Completed workflow children from this parent session stay addressable as retained children. `subagent({ action: "children.list" })` lists up to the last 10 with run ids and reports each row as `resumable` or `not resumable` with a reason. Resume only rows reported `resumable`. For a retained-child challenge, use `resume` instead of `steer` when the child is complete. If no retained child is resumable, launch a same-role fallback challenge and label it as fallback. A later workflow continues a resumable child with `runs.run(key, { resume: "<run-id>", task: "follow-up" })`. Inside `workflowScript`, awaiting that call waits for the revived child to finish and returns its completed output and new `runId`; top-level `{ action: "resume" }` remains detached. Pass explicit follow-up task text. Assign each returned child result back to the loop variable because every resume can return a new retained `runId`; always resume the latest returned id. `resume` and `agent` are mutually exclusive, the revived child keeps its stored agent/model/tool contract, and `gate` is rejected on retained resume items.
86
105
 
87
106
  Each workflow key identifies one result lane: use a new stable workflow key for every distinct retained resume pass; same-key calls are reused only when launch parameters are identical, and incompatible parameters are rejected.
@@ -97,9 +116,22 @@ return runs.run("cross-oracle", {
97
116
 
98
117
  Keyed resume reads that one exact receipt and revalidates the retained run at launch. It fails when the workflow or key is missing, the receipt is stale, `latest` is not `true`, or the recorded child is no longer resumable. The receipt is terminal-only: if `status.json` or `events.jsonl` exists without it, the workflow may still be active or terminal receipt writing may have failed. Use direct child run IDs from status/events for direct resume after the normal retained-child checks; do not reconstruct keyed entries from those files. Foreground workflow results expose the same receipt in `details.workflow.receipt`, but cross-workflow keyed lookup requires the durable receipt from an async workflow.
99
118
 
119
+ ### Parallel sequential lanes
120
+
121
+ For a broad plan with a known set of narrow, visible stages per lane, use
122
+ `runs.lanes(...)` inside a `workflowScript`; it is a nested helper, not a
123
+ top-level `subagent` mode. Give each lane and stage a stable key. The first
124
+ stage from every lane is launched together, then later stages sequence per lane.
125
+ `resume: "previous"` requires the retained predecessor, and a failed or blocked
126
+ stage blocks only that lane. The returned board exposes lane/stage results for
127
+ the parent. See the [canonical staged-lane example](../../../docs/workflows.md#parallel-sequential-lanes).
128
+
129
+ Use raw `runs.run(...)`/`runs.all(...)` instead when branching or rolling fanout
130
+ depends on runtime data rather than a predeclared stage plan.
131
+
100
132
  ### Async/background
101
133
 
102
- Prefer async mode for every subagent launch. Set `async: true` no matter the task unless the parent must block until completion. This applies to scouts, researchers, workers, reviewers, validators, oracle checks, one-off delegates, final review gates, backlog gates, and scripted workflows. Keep the write path single-threaded even when the run is async.
134
+ Prefer async mode for every subagent launch. Set `async: true` no matter the task unless the parent must block until completion. This applies to scouts, researchers, workers, reviewers, validators, oracle checks, one-off delegates, final review gates, publication gates, and scripted workflows. Keep the write path single-threaded even when the run is async.
103
135
 
104
136
  Use `async:false` only when the parent must block until completion. Async mode still shows progress. Do not use `async:false` because a task is short, because it is the last gate, because no other work is ready, because the user asked to finish the overall job, or because blocking is convenient.
105
137
 
@@ -107,9 +139,18 @@ Async does not mean parallel writes. Do not edit the same active worktree while
107
139
 
108
140
  Do not end your turn immediately after launching an async child if you promised to keep working. Continue the local inspection, synthesis, or validation prep, then check the async run when its result is needed. If no safe independent work remains, return control and let Pi wake the session; do not convert the child to foreground.
109
141
 
110
- In an interactive chat, normally return control when ready to yield and let Pi wake the session on completion; do not call `subagent_wait()` merely to wait. A run-to-completion user request is not by itself a reason to use foreground children. Override the normal yield-and-wake flow only when this exact turn cannot safely end without the result, such as a headless provider flow or a skill contract that must produce a same-turn artifact. Use `subagent_wait()`, not `async:false`, for that current-turn dependency. Never substitute sleep or status-polling loops.
142
+ In an ordinary interactive chat, normally return control after launching or
143
+ triaging useful async work and let Pi wake the session on completion; do not
144
+ call `subagent_wait()` merely to wait. A run-to-completion user request is not
145
+ by itself a reason to use foreground children. Override the normal yield-and-
146
+ wake flow only when this exact turn cannot safely end without the result, such
147
+ as a headless provider flow or a skill contract that must produce a same-turn
148
+ artifact. Use `subagent_wait()`, not `async:false`, for that current-turn
149
+ dependency. Never substitute sleep or status-polling loops.
150
+
151
+ `subagent_wait()` returns when the next initially active async run or registered provider item finishes or a subagent needs attention. Use `subagent_wait({ all: true })` for all work active at call time, `subagent_wait({ id: "..." })` for one async or remembered detached foreground run, and `subagent_wait({ timeoutMs })` to cap the block; active work keeps running if it elapses. `subagent_wait({ stopOnAttention: false })` keeps a blocking wait through idle or long-thinking attention, but supervisor/contact requests still stop it. In a long-lived interactive parent session, use `subagent_wait({ id: "...", nonBlocking: true })` to resolve the prefix to one exact run, persist an armed subscription, return immediately, and wake later on completion, failure, attention, reconciliation failure, or timeout. Ordinary status lists armed subscriptions separately from active children. This differs from disabling `waitTool`, which returns immediately without arming a future wake. If a foreground child detaches for supervisor coordination, reply first, then wait on its id; do not resume or launch a replacement while it remains detached. Headless sessions also auto-drain exact current-session work at `agent_end` as a final safeguard.
111
152
 
112
- `subagent_wait()` returns when the next initially active async run or registered provider item finishes or a subagent needs attention. Use `subagent_wait({ all: true })` for all work active at call time, `subagent_wait({ id: "..." })` for one async or remembered detached foreground run, and `subagent_wait({ timeoutMs })` to cap the block. In a long-lived interactive parent session, use `subagent_wait({ id: "...", nonBlocking: true })` to resolve the prefix to one exact run, persist an armed subscription, return immediately, and wake later on completion, failure, attention, reconciliation failure, or timeout. Ordinary status lists armed subscriptions separately from active children. This differs from disabling `waitTool`, which returns immediately without arming a future wake. If a foreground child detaches for supervisor coordination, reply first, then wait on its id; do not resume or launch a replacement while it remains detached. Headless sessions also auto-drain exact current-session work at `agent_end` as a final safeguard.
153
+ Providers are discovered through the `pi-subagents/background-work` registry and must expose a stable item id and owning session id. Load a provider through the child’s `extensions` or `subagentOnlyExtensions` and allow `subagent_wait` in its tools. For non-interactive fleets, launch N workers, wait for the next completion, react, and replace as needed; use `all: true` only when intentionally draining the fleet. If `SELESAI_SUBAGENT_WAIT_TOOL_ENABLED` disables blocking, direct waits return immediately, but headless `agent_end` auto-drain still surfaces provider, reconciliation, or timeout failures.
113
154
 
114
155
  ```typescript
115
156
  subagent({
@@ -41,7 +41,7 @@ subagent({
41
41
  description: "Project-specific implementation helper",
42
42
  systemPrompt: "Your system prompt here.",
43
43
  systemPromptMode: "replace",
44
- model: "openai-codex/gpt-5.4",
44
+ model: "provider/model-id",
45
45
  tools: "read,grep,find,ls,bash"
46
46
  }
47
47
  })
@@ -97,7 +97,7 @@ name: my-agent
97
97
  package: code-analysis
98
98
  description: What this agent does
99
99
  aliases: developer, coder
100
- model: openai-codex/gpt-5.4
100
+ model: provider/model-id
101
101
  thinking: high
102
102
  tools: read, grep, find, ls, bash
103
103
  systemPromptMode: replace
@@ -2,6 +2,8 @@
2
2
 
3
3
  Use this reference when several independent tasks need coordinated workers, worktrees, or repositories. It defines lane ownership; use the other pi-subagents references for run controls, prompts, and mission details. The parent remains the final decision-maker.
4
4
 
5
+ Create lanes only when delegation materially improves evidence, independent review, or isolated execution. Do not manufacture parallelism: keep dependent work serial, and only split work when each lane has a distinct decision and useful output.
6
+
5
7
  ## Lane board and authority
6
8
 
7
9
  Before multiple mutation-capable lanes start, record this board in the parent context:
@@ -20,15 +22,25 @@ For Pi extension repositories, keep lane worktrees outside auto-discovered exten
20
22
 
21
23
  Partition fanout by repository, source seam, decision, or review angle. Each run needs a stable key, lane-specific task, and a managed output path when a file is needed. Do not launch prompts that differ only by item name or broad file glob.
22
24
 
25
+ ### Cold-start packets and bounded orchestration audits
26
+
27
+ Every child packet must stand alone: include the goal, exact repository/cwd/ref, authority and edit boundary, relevant context/evidence, success criteria, validation, expected output, and stop/escalation rules. Do not rely on parent history, an issue number, or a broad glob alone. An orchestration audit by a top-reasoning critic model is read-only and returns at most three cited omissions; use high thinking only as an explicit parent/user escalation, never as an autonomous root or a parallel placeholder.
28
+
23
29
  Use one async `workflowScript` for a coordinated wave. Use `runs.all` for independent lanes and `runs.run` for dependent lane stages. Give cross-repository runs explicit `cwd` values and lane-qualified outputs. Use `outputMode: "file-only"` when a report must survive the run or feed a later stage. Keep scratch outputs relative so they live under subagent artifacts; use absolute paths only for durable memory, approved docs paths, or final handoff files.
24
30
 
25
31
  ## Keep independent work moving
26
32
 
27
33
  While one lane waits, run safe independent preparation, validation, or fresh read-only review lanes. Do not block the parent just because a run is active. If no safe lane remains, record the blocker and the event that will reopen work.
28
34
 
35
+ In an ordinary interactive session, completion wakes the parent; after useful
36
+ async lanes are launched or triaged, yield rather than use
37
+ `subagent_wait({ all: true })` as a barrier. “Continue/orchestrate/work until
38
+ done” means keep the board moving while safe immediate work remains. If only
39
+ async lanes are running, record the revisit trigger and yield.
40
+
29
41
  An ordinary coordinated workflow has one mission. Use its durable state, artifacts, run records, and receipts for recovery. Treat a receipt as evidence, not as authority or acceptance.
30
42
 
31
- After a writer produces a candidate, run the required fresh-context, read-only reviewer. The reviewer inspects the exact worktree and returns evidence-backed findings. The parent decides which findings are in scope and whether the lane is ready. Send accepted fixes to that lane's sole writer, then rerun only the affected gate.
43
+ After a writer produces a candidate, run the required fresh-context, read-only reviewer. The reviewer inspects the exact worktree and returns evidence-backed findings. The parent decides which findings are in scope and whether the lane is ready. Use `review-and-validation.md` for finding disposition, validation, and gate-failure triage. Send accepted fixes to that lane's sole writer, then rerun only the affected gate.
32
44
 
33
45
  ## Handoff, cleanup, and recovery
34
46
 
@@ -8,7 +8,7 @@ Parent extensions may register a session-scoped, out-of-band ceiling through `pi
8
8
 
9
9
  ## When to Use
10
10
 
11
- - **Complex work orchestration**: use Fable mode as the default parent-agent loop for complex work. Complex means the task has multiple moving parts, unclear acceptance, cross-cutting code, meaningful user-visible impact, expensive or irreversible validation, broad review surface, or the user asks for orchestration. Lightweight one-off delegation can stay lightweight.
11
+ - **Complex work orchestration**: keep the parent on its ordinary strong default model. Delegate only when another child materially improves evidence, independent review, or isolated execution; omission failures are cheaper than unnecessary commissions. For hard orchestration or root-cause questions, use a top-reasoning model only as a bounded read-only critic/oracle escalation, never as an autonomous root. Complex means the task has multiple moving parts, unclear acceptance, cross-cutting code, meaningful user-visible impact, expensive or irreversible validation, broad review surface, or the user asks for orchestration. Lightweight one-off delegation can stay lightweight.
12
12
  - **Advisory review**: use fresh-context `reviewer` agents for adversarial code review, or fork to `oracle` when inherited decisions and drift matter
13
13
  - **Implementation handoff**: have `oracle` advise, then `worker` implement only after an approved direction
14
14
  - **Recon and planning**: use `scout`, then write a plan when needed
@@ -20,7 +20,7 @@ Parent extensions may register a session-scoped, out-of-band ceiling through `pi
20
20
 
21
21
  ## Tool vs Slash Commands
22
22
 
23
- Agents use the `subagent(...)` tool with `workflowScript` for execution, and `action` for management, status, and control. Humans often use the slash-command layer instead:
23
+ Agents use the `subagent(...)` tool for execution, management, status, and control. Direct `{ agent, task }` execution is enough for one bounded child task; use `workflowScript` when the parent needs JavaScript control flow or data-dependent branching, keyed, parallel, sequential, retry, retained-resume, aggregate, or explicit staged-lane behavior (`runs.lanes`). Humans often use the slash-command layer instead:
24
24
 
25
25
  - `/run` — launch a single agent
26
26
  - `workflowScript` — the sole public surface for sequence, parallelism, branching, retries, and aggregation
@@ -51,11 +51,15 @@ Packaged prompt shortcuts are also available for repeatable workflows. Treat the
51
51
 
52
52
  The prompt templates in `prompts/` encode workflows the parent agent can run on demand. If the user provides a URL, issue, PR, plan, local file, screenshot, or freeform target, treat that target as the primary scope: read or fetch it before launching children, then include it explicitly in every child task. For targets outside the parent cwd, include the exact repository, explicit `cwd`, authority boundary, and expected output path in each child task. Do not depend on the parent conversation history when the recipe calls for fresh context.
53
53
 
54
+ ### Commission-risk and cold-start packets
55
+
56
+ Delegate only when the child materially improves evidence, independent review, or isolated execution; do not manufacture parallelism. Every child packet must be cold-start complete: state the goal, exact target/cwd/ref, authority and edit boundary, relevant context/evidence, success criteria, validation, output, and stop/escalation rules. For an orchestration audit by the critic tier, make the child read-only and request at most three omissions, each cited to a file, line, or decision; high thinking is an explicit escalation, not a default.
57
+
54
58
  ### Council Mode technique
55
59
 
56
- Use Council Mode when the user asks to convene advisors, debate a material decision, cross-examine recommendations, or critique and improve a plan with several model perspectives. This includes requests such as “run a council on this architecture,” “have Sol, Fable, and Kimi critique this plan,” or “get multiple oracles to debate the tradeoffs.” Read `../council-mode/SKILL.md` and follow its bounded parent-supervised protocol instead of launching ad hoc parallel oracle calls.
60
+ Use Council Mode when the user asks to convene advisors, debate a material decision, cross-examine recommendations, or critique and improve a plan with several model perspectives. This includes requests such as “run a council on this architecture,” “have the configured advisors critique this plan,” or “get multiple oracles to debate the tradeoffs.” Read `../council-mode/SKILL.md` and follow its bounded parent-supervised protocol instead of launching ad hoc parallel oracle calls.
57
61
 
58
- Council advisors are read-only. User or project `council-*` profiles can pin models such as GPT 5.6 Sol, Fable, or Kimi and define any persistent stance in the profile body. Package advisors such as Surf's `gpt-pro` can join the roster only when the `surf-cli` Pi extension is installed and its `surf-oracle` provider is registered; treat them as external runners, omit child `async` for attached results, and do not pass `outputSchema` to them. The council question and scope provide the decision frame; do not invent per-advisor role labels. The parent collects independent reports, optionally sends curated cross-exam packets, and writes the final memo. Do not treat the council as agent-to-agent chat, implementation authority, or a writer swarm.
62
+ Council advisors are read-only. User or project `council-*` profiles choose allowed models and define any persistent stance in the profile body. A top-reasoning advisor remains bounded and read-only; it does not become the root. Package advisors such as Surf's `gpt-pro` can join the roster only when the `surf-cli` Pi extension is installed and its `surf-oracle` provider is registered; treat them as external runners, omit child `async` for attached results, and do not pass `outputSchema` to them. The council question and scope provide the decision frame; do not invent per-advisor role labels. The parent collects independent reports, optionally sends curated cross-exam packets, and writes the final memo. Do not treat the council as agent-to-agent chat, implementation authority, or a writer swarm.
59
63
 
60
64
  ### Parallel review technique
61
65
 
@@ -111,6 +115,12 @@ Use this after implementation when the user wants cleanup review or when a final
111
115
 
112
116
  Use this when a broad diff has known reviewer findings across several items and the user wants the parent to “orchestrate subagents like a boss.” Keep the active worktree safe with a three-stage `workflowScript`:
113
117
 
118
+ When staged seams are available, a low-tier writer should not receive the
119
+ end-to-end issue. Use `runs.lanes` inside `workflowScript` to keep stages narrow:
120
+ a scout/red test, helper-only change, one render seam, validation, minimality
121
+ challenge, or fresh review. Give the writer only its assigned implementation
122
+ stage; keep sequencing and synthesis with the parent.
123
+
114
124
  1. A parallel read-only planning fanout, one reviewer per issue cluster. Each child inspects the real diff and returns exact files, line refs, proposed fixes, and focused validation. They must not edit.
115
125
  2. One writer worker. It receives the reviewer summaries as the awaited planning results (or their durable output paths) interpolated into its task, plus the parent’s accepted scope, stop rules, and verification contract. It is the only child allowed to edit the active worktree.
116
126
  3. A parallel read-only validation fanout. Validators inspect the worker diff from fresh context with distinct angles, report pass/fail, remaining blockers, and missing verification.
@@ -161,19 +171,19 @@ subagent({
161
171
  Builtin agents load at the lowest priority. Project agents override user agents,
162
172
  and user/project agents override builtins with the same name.
163
173
 
164
- | Agent | Purpose | Model | Typical output / role |
174
+ | Agent | Purpose | Recommended tier | Typical output / role |
165
175
  |-------|---------|-------|------------------------|
166
- | `scout` | Fast codebase recon | inherits default | Writes `context.md` handoff material |
167
- | `worker` | Implementation and approved oracle handoffs | inherits default | Single-writer implementation with decision escalation |
168
- | `reviewer` | Review specialist | inherits default | Default recipes are review-only; tools include edit/write when a fix pass is explicit |
169
- | `researcher` | Web research brief generator | inherits default | Writes `research.md` |
170
- | `delegate` | Lightweight generic delegate | inherits default | No fixed output; generic delegated work |
171
- | `oracle` | Decision-consistency advisory review | inherits default | Advisory review, intercom coordination |
172
- | `advisor` | Claude Code-compatible alias for `oracle` | inherits default | Same advisory role as `oracle` |
176
+ | `scout` | Fast codebase recon | fast worker/scout tier | Writes `context.md` handoff material |
177
+ | `worker` | Implementation and approved oracle handoffs | capable worker tier | Single-writer implementation with decision escalation |
178
+ | `reviewer` | Review specialist | strong reviewer tier; high thinking for serious reviews | Default recipes are review-only; tools include edit/write when a fix pass is explicit |
179
+ | `researcher` | Web research brief generator | inherits configured default | Writes `research.md` |
180
+ | `delegate` | Lightweight generic delegate | inherits configured default | No fixed output; generic delegated work |
181
+ | `oracle` | Decision-consistency advisory review | top-reasoning critic tier, bounded read-only; high thinking escalation only | Advisory review, intercom coordination |
182
+ | `advisor` | Compatibility alias for `oracle` | top-reasoning critic tier, bounded read-only; high thinking escalation only | Same advisory role as `oracle` |
173
183
 
174
184
  Builtin `worker` and `delegate` use strict tool allowlists and do not inherit ambient parent extension tools. To give a child an extension tool, name it in `tools` and load its provider via `extensions`, a path-like `tools` entry, or `subagentOnlyExtensions`. Custom agents without an `extensions` field follow `subagents.defaultExtensions` when set.
175
185
 
176
- Builtin agents inherit the current Pi default model unless a run, user setting, project setting, or `subagents.defaultModel` overrides `model`. Set `subagents.defaultModel` when subagents should use a different default model than the parent session. Override builtin defaults before copying full agent files when a small tweak is enough.
186
+ Builtin agents inherit the current Pi default model unless a run, user setting, project setting, or `subagents.defaultModel` overrides `model`. The table records recommended tier routing, not shipped hard defaults; explicit run, user, or project settings still win. Keep the parent/orchestrator on the ordinary strong default model unless parent/user policy says otherwise. Override builtin defaults before copying full agent files when a small tweak is enough.
177
187
 
178
188
  Set `subagents.defaultThinking` to apply a shared thinking level to builtin, package, user, and project agents whose frontmatter leaves `thinking` unset. Project settings win over user settings; explicit frontmatter (including `thinking: false`), `agentOverrides.<name>.thinking`, and per-run overrides remain more specific. This setting affects child agents only and does not change the parent session's default thinking level.
179
189
 
@@ -188,7 +198,7 @@ Set `subagents.defaultThinking` to apply a shared thinking level to builtin, pac
188
198
  For one run, use inline config:
189
199
 
190
200
  ```text
191
- /run reviewer[model=anthropic/claude-sonnet-4] "Review this diff"
201
+ /run reviewer[model=provider/review-model] "Review this diff"
192
202
  ```
193
203
 
194
204
  For persistent tweaks, edit `subagents.agentOverrides` in user or project settings. User overrides apply everywhere. Project overrides apply only in that repo and win over user overrides. Use `/subagents-models` or `subagent({ action: "models" })` to inspect the live mapping after settings and overrides load.
@@ -202,18 +212,18 @@ Provider-scoped entries can layer on top of the default override for the active
202
212
  "worker": { "thinking": "medium" }
203
213
  },
204
214
  "agentOverridesByProvider": {
205
- "github-copilot": {
206
- "worker": { "model": "github-copilot/gpt-5-mini" }
215
+ "provider-a": {
216
+ "worker": { "model": "provider-a/fast-worker-model" }
207
217
  },
208
- "openrouter": {
209
- "worker": { "model": "openrouter/openai/gpt-5-mini" }
218
+ "provider-b": {
219
+ "worker": { "model": "provider-b/fast-worker-model" }
210
220
  }
211
221
  }
212
222
  }
213
223
  }
214
224
  ```
215
225
 
216
- Model ids do not have to be exact. Separator variations (`claude-haiku-4.5` vs `claude-haiku-4-5`), case (`Claude-Sonnet-4`), and optional trailing date stamps (`claude-haiku-4-5-20251001`) all resolve to the same registry model. Exact `provider/id` wins; a qualified `provider/model` never switches providers. To constrain subagents to a budget or compliance profile, set `subagents.modelScope: { enforce: true, allow: ["anthropic/*", "openai/gpt-5-*"] }` in user or project settings. Out-of-scope models you pass explicitly error and abort; models inherited from frontmatter, `subagents.defaultModel`, agent frontmatter, or the parent session only warn.
226
+ Model ids do not have to be exact. Separator variations (`fast.worker-v1` vs `fast-worker-v1`), case (`Strong-Review-Model`), and optional trailing date stamps all resolve to the same registry model. Exact `provider/id` wins; a qualified `provider/model` never switches providers. To constrain subagents to a budget or compliance profile, set `subagents.modelScope: { enforce: true, allow: ["approved-provider/*", "second-provider/approved-*"] }` in user or project settings. Out-of-scope models you pass explicitly error and abort; models inherited from frontmatter, `subagents.defaultModel`, agent frontmatter, or the parent session only warn.
217
227
 
218
228
  For model fleets, use the profile commands instead of hand-editing repeated overrides: `/subagents-refresh-provider-models <provider>`, `/subagents-generate-profiles <provider>`, `/subagents-load-profile <name>`, and `/subagents-check-profile <name>`. Profiles live under `~/.selesai/agent/profiles/pi-subagents/` and replace only `settings.subagents` when loaded.
219
229
 
@@ -249,9 +259,9 @@ Direct settings example:
249
259
  "subagents": {
250
260
  "agentOverrides": {
251
261
  "reviewer": {
252
- "model": "anthropic/claude-sonnet-4",
262
+ "model": "provider/strong-review-model",
253
263
  "thinking": "high",
254
- "fallbackModels": ["openai-codex/gpt-5.6-luna:low"],
264
+ "fallbackModels": ["backup-provider/strong-review-model"],
255
265
  "acceptanceRole": "read-only"
256
266
  }
257
267
  }
@@ -269,14 +279,11 @@ agent with the same name only when you want a substantially different agent.
269
279
 
270
280
  ### Recommended model tiering (optional)
271
281
 
272
- When several providers are available, route agents by task shape instead of one model for everything:
282
+ Keep the parent/orchestrator on the ordinary strong default model because omission failures are cheaper than unnecessary commissions. Route workers and scouts to a fast, capable worker tier, and keep serious reviews on the strong tier at high thinking. Use a top-reasoning model only for bounded, read-only critic/oracle/root-cause audits; critic-tier high thinking is escalation-only and never an autonomous root. Explicit parent/user model policy wins over these recommendations.
273
283
 
274
- 1. **Fast workhorse** — cheapest capable model at low thinking for recon, lookups, and mechanical edits (for example on `scout`).
275
- 2. **Standard well-scoped** — mid-tier model at medium thinking for most delegations: routine multi-file edits, focused reviews, straightforward implementation (for example on `worker`, `reviewer`, `delegate`).
276
- 3. **Deep but bounded** — top reasoning model at high thinking only for hard tasks that arrive with explicit goals and completion criteria; these models loop on vague goals (for example on oracle-style agents).
277
- 4. **Taste and intent** — a model that reads human intent well for ambiguous work: UX/design judgment, product tradeoffs, planning from vague requirements, writing quality.
284
+ Examples are illustrative, not requirements. Map these tiers to concrete models in user/project settings or a profile. A non-OpenAI setup should choose comparable available models by capability.
278
285
 
279
- Routing rule: use tiers 1–3 when the task is well-scoped; use tier 4 when scoping or judging is the task itself. Give tier-4 agents cross-provider `fallbackModels` so subscription usage limits degrade gracefully; fallback triggers automatically on rate-limit and overload errors. Note that forked context over an Anthropic parent transcript with signed thinking blocks forces the child's thinking off, so intent-tier agents work best with fresh context.
286
+ Use `fallbackModels` when a tier has provider quota or availability risk. Prefer fresh context for cross-provider children when inherited provider-specific reasoning blocks would force thinking off.
280
287
 
281
288
  If a provider rejects model IDs with thinking suffixes, use
282
289
  `subagents.disableThinking: true` in user or project settings to clear bundled
@@ -0,0 +1,73 @@
1
+ # Pi Subagents: Review And Validation
2
+
3
+ Generic review and delivery guidance for delegated work. This file does not encode private backlog, merge, or release policy.
4
+
5
+ ## Delivery loop
6
+
7
+ Use the smallest loop that proves the change:
8
+
9
+ 1. Inspect the source, diff, issue, or plan directly.
10
+ 2. Keep one writer for each cwd or worktree.
11
+ 3. Run focused validation that can fail for the changed behavior.
12
+ 4. Use fresh-context read-only review for substantial, risky, public, or hard-to-see changes.
13
+ 5. Apply only accepted findings inside the same writer boundary.
14
+ 6. Re-run affected validation and review only the changed blast radius.
15
+ 7. Inspect the final diff and evidence before parent acceptance.
16
+
17
+ Skip review ceremony for trivial wording, renames, or local-only probes when direct parent inspection is enough.
18
+
19
+ ## Review shape
20
+
21
+ | Situation | Shape |
22
+ | --- | --- |
23
+ | One coherent diff or one risk | one reviewer |
24
+ | Independent risks, such as correctness, tests, security, or UI | parallel reviewers with distinct contracts |
25
+ | Possible over-scope or needless complexity | same-writer challenge before fresh review |
26
+ | Material design tradeoff | council mode |
27
+
28
+ Reviewers are fresh-context by default. Forked reviewers are for parent-history, drift, or prior-decision evidence.
29
+
30
+ ## Finding disposition
31
+
32
+ The parent classifies each finding against current HEAD:
33
+
34
+ - **Valid blocker:** concrete failure, repro, security issue, contract mismatch, or source-proven regression. Fix now.
35
+ - **Valid non-blocker:** real but outside the delivery slice. Record or defer.
36
+ - **Stale:** fixed or absent at the reviewed head. Cite current evidence.
37
+ - **Invalid:** contradicted by source, tests, docs, or user-approved scope. Cite the contradiction.
38
+ - **Out of policy/scope:** needs unapproved product, architecture, authority, release, or public-repo action. Escalate.
39
+ - **Speculative:** no contract, repro, or reachable failure. Do not block.
40
+
41
+ A clean reviewer result is evidence, not publication authority.
42
+
43
+ ## Gate-failure triage
44
+
45
+ When validation fails:
46
+
47
+ 1. Confirm the run belongs to the exact head/ref under judgment.
48
+ 2. Read the focused failing logs first.
49
+ 3. Name the failing test, assertion, contract, or thread.
50
+ 4. Classify cause: current diff, stale test, environment/setup, or existing flake.
51
+ 5. Reproduce locally when practical with the narrowest command.
52
+ 6. Patch forward when the current diff caused it.
53
+ 7. For stale/flaky failures, collect proof before one rerun or residual-risk note.
54
+ 8. Re-run the affected command or exact-head gate after every fix.
55
+
56
+ For bot comments, classify each thread as valid, stale, invalid, or out of policy before assigning severity.
57
+
58
+ ## Final checklist
59
+
60
+ Before reporting delegated work as done, verify the relevant subset:
61
+
62
+ - final diff contains only intended files
63
+ - focused validation covers changed behavior
64
+ - substantial or risky changes have fresh-review evidence
65
+ - accepted findings are fixed and revalidated
66
+ - publication authority exists before push, comment, close, merge, deploy, or release
67
+ - external checks are exact-head when used as evidence
68
+ - handoff is durable before cleanup
69
+ - residual risks, skipped validation, and blocked decisions are explicit
70
+
71
+ ## Public/private boundary
72
+
73
+ For issue/PR backlogs, releases, merge queues, contributor credit, or repo-specific policy, load the matching user/project skill when available. Keep those rules out of this public package until intentionally released.