pi-subagents 0.66.0 → 0.68.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +138 -0
- package/README.md +5 -4
- package/agents/evidence-auditor.md +34 -0
- package/agents/reviewer.md +3 -2
- package/docs/agents.md +43 -15
- package/docs/configuration.md +67 -23
- package/docs/extension-api.md +38 -19
- package/docs/missions.md +10 -2
- package/docs/models.md +11 -79
- package/docs/observability.md +22 -12
- package/docs/standalone-background.md +59 -0
- package/docs/tool-reference.md +33 -20
- package/docs/watchdog.md +39 -10
- package/docs/workflows.md +37 -13
- package/index.ts +5 -2
- package/inspector-runner.mjs +2 -2
- package/package.json +4 -2
- package/prompts/parallel-review.md +1 -1
- package/runner-peer-loader.mjs +24 -0
- package/runner-peer-preload.mjs +32 -0
- package/skills/pi-subagents/SKILL.md +32 -21
- package/skills/pi-subagents/references/constraints-and-recipes.md +3 -2
- package/skills/pi-subagents/references/execution-controls.md +11 -9
- package/skills/pi-subagents/references/management-authoring-rpc.md +0 -1
- package/skills/pi-subagents/references/multi-lane-orchestration.md +1 -1
- package/skills/pi-subagents/references/prompting-and-roles.md +17 -13
- package/skills/pi-subagents/references/review-and-validation.md +3 -3
- package/src/agents/advertised-agent-prompt.ts +34 -3
- package/src/agents/agent-management.ts +57 -58
- package/src/agents/agent-serializer.ts +4 -3
- package/src/agents/agents.ts +190 -71
- package/src/agents/builtin-names.ts +1 -0
- package/src/agents/chain-serializer.ts +5 -0
- package/src/agents/runtime-agent-registry.ts +7 -6
- package/src/api/delegation.ts +4 -0
- package/src/api/preflight.ts +94 -59
- package/src/api/required-child-extensions.ts +6 -0
- package/src/api/shared-types.ts +2 -0
- package/src/extension/config.ts +10 -37
- package/src/extension/fanout-child.ts +66 -4
- package/src/extension/herdr-pi-bridge.ts +160 -0
- package/src/extension/index.ts +62 -39
- package/src/extension/public-execution.ts +7 -5
- package/src/extension/rpc.ts +4 -0
- package/src/extension/schemas.ts +81 -81
- package/src/extension/tool-description.ts +30 -82
- package/src/inspectors/actions.ts +148 -0
- package/src/inspectors/ghostty/actions.ts +74 -0
- package/src/inspectors/ghostty/plugin.ts +17 -0
- package/src/inspectors/herdr/actions.ts +99 -179
- package/src/inspectors/herdr/plugin.ts +20 -0
- package/src/inspectors/herdr/project-panes.ts +1 -1
- package/src/inspectors/{herdr/inspector-runner.ts → inspector-runner.ts} +12 -12
- package/src/inspectors/plugins.ts +8 -0
- package/src/inspectors/{herdr/session-roots-codec.ts → session-roots-codec.ts} +3 -14
- package/src/inspectors/types.ts +51 -0
- package/src/intercom/intercom-bridge.ts +50 -8
- package/src/intercom/native-supervisor-channel.ts +44 -31
- package/src/policy/authority.ts +4 -0
- package/src/profiles/profiles.ts +12 -6
- package/src/runs/background/active-async-capacity.ts +4 -0
- package/src/runs/background/active-run-index.ts +17 -1
- package/src/runs/background/async-execution.ts +348 -176
- package/src/runs/background/async-job-tracker.ts +8 -6
- package/src/runs/background/async-resume.ts +17 -12
- package/src/runs/background/async-status.ts +15 -4
- package/src/runs/background/auto-drain.ts +23 -10
- package/src/runs/background/binary-bootstrap.ts +38 -0
- package/src/runs/background/chain-append.ts +1 -1
- package/src/runs/background/chain-root-attachment.ts +14 -33
- package/src/runs/background/fleet-view.ts +30 -2
- package/src/runs/background/notify.ts +105 -7
- package/src/runs/background/owned-process-tree.ts +29 -2
- package/src/runs/background/result-files.ts +8 -4
- package/src/runs/background/result-watcher.ts +19 -2
- package/src/runs/background/run-child-session.ts +81 -33
- package/src/runs/background/run-status.ts +3 -0
- package/src/runs/background/runner-aliases.ts +12 -33
- package/src/runs/background/runner-child-launch.ts +6 -1
- package/src/runs/background/runner-child-sessions.ts +5 -4
- package/src/runs/background/runner-http-dispatcher.ts +119 -0
- package/src/runs/background/scheduled-runs.ts +51 -18
- package/src/runs/background/stale-run-reconciler.ts +35 -11
- package/src/runs/background/steering.ts +20 -2
- package/src/runs/background/subagent-runner.ts +441 -304
- package/src/runs/background/subagent-wait.ts +176 -28
- package/src/runs/background/wait-completions.ts +75 -27
- package/src/runs/background/wait-subscriptions.ts +9 -3
- package/src/runs/background/wait-tool.ts +5 -3
- package/src/runs/foreground/async-steering-action.ts +18 -7
- package/src/runs/foreground/async-stop-action.ts +93 -3
- package/src/runs/foreground/execution.ts +134 -248
- package/src/runs/foreground/foreground-history.ts +2 -1
- package/src/runs/foreground/prompt-audit.ts +3 -1
- package/src/runs/foreground/subagent-executor.ts +374 -178
- package/src/runs/foreground/workflow-detach-reconcile.ts +2 -0
- package/src/runs/foreground/workflow-foreground-steering.ts +2 -1
- package/src/runs/shared/acceptance.ts +38 -11
- package/src/runs/shared/async-status-projection.ts +127 -33
- package/src/runs/shared/capability-ceiling.ts +2 -0
- package/src/runs/shared/child-hooks.ts +25 -10
- package/src/runs/shared/child-launch-plan.ts +15 -3
- package/src/runs/shared/child-launch.ts +28 -5
- package/src/runs/shared/child-lifecycle.ts +6 -3
- package/src/runs/shared/child-runtime-config.ts +8 -1
- package/src/runs/shared/child-session.ts +127 -52
- package/src/runs/shared/child-tool-plan.ts +142 -11
- package/src/runs/shared/completion-guard.ts +5 -3
- package/src/runs/shared/dynamic-fanout.ts +2 -2
- package/src/runs/shared/effective-system-prompt.ts +33 -0
- package/src/runs/shared/external-cli-contract.ts +11 -1
- package/src/runs/shared/external-cli-preflight.ts +6 -2
- package/src/runs/shared/external-cli-runner.ts +9 -7
- package/src/runs/shared/herdr-connection.ts +134 -0
- package/src/runs/shared/herdr-external-adapters.ts +169 -0
- package/src/runs/shared/herdr-machine.ts +279 -0
- package/src/runs/shared/herdr-pi-protocol.ts +59 -0
- package/src/runs/shared/herdr-placed-run.ts +263 -0
- package/src/runs/shared/llm-intent-arbiter.ts +12 -3
- package/src/runs/shared/model-resolution-diagnostic.ts +76 -0
- package/src/runs/shared/{model-fallback.ts → model-resolution.ts} +22 -235
- package/src/runs/shared/model-scope.ts +1 -1
- package/src/runs/shared/nested-events.ts +11 -2
- package/src/runs/shared/orca-progress-tabs.ts +1 -1
- package/src/runs/shared/parallel-utils.ts +7 -2
- package/src/runs/shared/pi-spawn.ts +10 -0
- package/src/runs/shared/subagent-prompt-runtime.ts +12 -4
- package/src/runs/shared/task-intent.ts +46 -13
- package/src/runs/shared/workflow-async-child-guidance.ts +18 -0
- package/src/runs/shared/worktree-setup-command.ts +27 -4
- package/src/runs/shared/worktree.ts +45 -15
- package/src/shared/child-cache-retention.ts +43 -0
- package/src/shared/fork-context.ts +15 -72
- package/src/shared/launch-contract.ts +68 -8
- package/src/shared/opencode-session-headers.ts +30 -0
- package/src/shared/pruned-fork.ts +1 -1
- package/src/shared/required-child-extensions.ts +81 -0
- package/src/shared/settings.ts +5 -2
- package/src/shared/shortcuts.ts +0 -4
- package/src/shared/types.ts +74 -30
- package/src/slash/delegation-adapters.ts +3 -1
- package/src/slash/delegation-request.ts +14 -0
- package/src/slash/slash-commands.ts +2 -7
- package/src/slash/subagents-admin.ts +24 -13
- package/src/tui/fleet-status.ts +164 -19
- package/src/tui/fleet.ts +16 -14
- package/src/tui/render.ts +168 -37
- package/src/watchdog/child-status.ts +28 -28
- package/src/watchdog/lsp-diagnostics.ts +1 -1
- package/src/watchdog/model-selection.ts +21 -1
- package/src/watchdog/permission-arbiter.ts +3 -1
- package/src/watchdog/register-child.ts +10 -2
- package/src/watchdog/register-main.ts +39 -35
- package/src/watchdog/render.ts +1 -1
- package/src/watchdog/review.ts +123 -74
- package/src/watchdog/rules.ts +1 -1
- package/src/watchdog/runtime.ts +100 -27
- package/src/watchdog/scope.ts +1 -1
- package/src/watchdog/settings.ts +3 -0
- package/src/watchdog/tool-actions.ts +13 -12
- package/src/watchdog/turn-delta.ts +23 -0
- package/src/watchdog/types.ts +5 -3
- package/src/watchdog/warning-format.ts +1 -1
- package/src/workflows/scripted-workflow.ts +279 -10
- package/src/workflows/workflow-checklist.ts +2 -2
- package/src/workflows/workflow-receipt.ts +21 -3
- package/src/workflows/workflow-resources.ts +13 -2
- package/runner-server-preload.mjs +0 -13
- package/src/runs/shared/model-exclusions.ts +0 -374
- package/src/runs/shared/readonly-model-continuation.ts +0 -69
- package/src/runs/shared/readonly-session-evidence.ts +0 -307
- /package/src/inspectors/{herdr/shell-command.ts → shell-command.ts} +0 -0
package/docs/observability.md
CHANGED
|
@@ -12,7 +12,7 @@ A background child is a pi session created inside the detached runner process. T
|
|
|
12
12
|
|
|
13
13
|
Live progress shows compact detail for single, chain, and parallel modes: a bounded one-line task, current tool, recent output, token counts, aggregate cost, duration, activity freshness, current-tool duration, and chain graph metadata when available. Workflow `label` metadata wins over raw task text in compact multi-child cards.
|
|
14
14
|
|
|
15
|
-
Press Pi's configured expand key (`Ctrl+O` by default) to expand the full streaming view with complete output per step.
|
|
15
|
+
Press Pi's configured expand key (`Ctrl+O` by default) to expand the full streaming view with complete output per step.
|
|
16
16
|
|
|
17
17
|
Sequential chains show a flow line like `done scout → running worker`. Chains with parallel steps show per-step cards instead. Chain status uses `label` and `phase` metadata when present, while falling back to agent names for older chains.
|
|
18
18
|
|
|
@@ -35,11 +35,22 @@ async subagent worker · background
|
|
|
35
35
|
● Step 1/1: worker · running
|
|
36
36
|
task: Review authentication boundaries
|
|
37
37
|
⎿ read: src/auth.ts | 2.0s
|
|
38
|
-
Press configured-expand-key for live detail
|
|
38
|
+
Press configured-expand-key for live detail
|
|
39
39
|
```
|
|
40
40
|
|
|
41
41
|
To inspect one background child in text, use `subagent({ action: "status", id: "...", view: "transcript" })`; add `index` for a specific child in a parallel or chain run.
|
|
42
42
|
|
|
43
|
+
In Pi fullscreen mode with mouse dispatch (verified with Pi TUI 0.85.1), left-click
|
|
44
|
+
anywhere on the async widget's header row to fold it into a live one-line status
|
|
45
|
+
summary. Click again to restore the usual layout. No knowledge of extension commands
|
|
46
|
+
or keyboard shortcuts is needed. The summary counts the widget's tracked runs,
|
|
47
|
+
including workflow parents and children, rather than unique agents.
|
|
48
|
+
|
|
49
|
+
Folding stays in effect across progress updates and does not change Pi's global
|
|
50
|
+
expand setting, run execution, or completion notifications. Task rows, drag and
|
|
51
|
+
wheel events, and modifier clicks are left unhandled. The state resets when the
|
|
52
|
+
widget is removed or Pi reloads. Regular mode keeps the existing keyboard controls.
|
|
53
|
+
|
|
43
54
|
### Reducing status display noise
|
|
44
55
|
|
|
45
56
|
Chat records tool-call history; FleetView and the async widget show live run/child updates. Separate `subagent({ action: "status", id: "..." })` calls leave separate historical entries even when their `Status target: run …` labels match. A matching run ID identifies the queried run, not the tool call, and is not evidence of duplicate execution. Live Fleet/widget refreshes do not merge those entries.
|
|
@@ -55,7 +66,7 @@ For compact chat results with FleetView as the only live editor surface, merge t
|
|
|
55
66
|
```
|
|
56
67
|
|
|
57
68
|
- `inlineToolDisplay: "summary"` keeps one static result row per call, alongside its call heading. A completed status query is not proof that the queried child has finished.
|
|
58
|
-
- `fleetView: true` retains live progress. Open `/subagents-fleet`
|
|
69
|
+
- `fleetView: true` retains live progress. Open `/subagents-fleet` for details instead of repeatedly requesting status just to watch progress. Pi's expand key does not expand summary results; keep `"rich"` if you want expandable inline output.
|
|
59
70
|
- `asyncWidget: false` hides only the additional under-editor async widget, leaving FleetView available. This configuration reduces visible surfaces; it does not guarantee ordering relative to other extensions.
|
|
60
71
|
|
|
61
72
|
Thanks to [DraconDev](https://github.com/DraconDev) for reporting the display noise and suggesting summary mode in [#1931](https://github.com/nicobailon/pi-subagents/issues/1931).
|
|
@@ -78,7 +89,7 @@ After you expand it:
|
|
|
78
89
|
reviewer · running 38s · ↓ 1.1k window · 1.4k spent
|
|
79
90
|
```
|
|
80
91
|
|
|
81
|
-
When the focused editor is empty, press `↓` or `←` to expand the summary into `main` plus active children with agent name, state, elapsed time, and token usage. When providers report usage, `window` is the latest assistant turn's input plus cache-read tokens, while `spent` keeps the cumulative input-plus-output total. Old run artifacts without window data keep the existing token-total label. The compact line counts active current-session work and Herdr project panes. Then use `↑`/`↓` or `j`/`k` to select a child and `Enter` to open the Fleet lobby; press `Enter` or `H` there to open its child-specific
|
|
92
|
+
When the focused editor is empty, press `↓` or `←` to expand the summary into `main` plus active children with agent name, state, elapsed time, and token usage. When providers report usage, `window` is the latest assistant turn's input plus cache-read tokens, while `spent` keeps the cumulative input-plus-output total. Old run artifacts without window data keep the existing token-total label. The compact line counts active current-session work and Herdr project panes. Then use `↑`/`↓` or `j`/`k` to select a child and `Enter` to open the Fleet lobby; press `Enter` or `H` there to open its child-specific inspector through an available Inspect plugin. Printable navigation keys are never intercepted before activation.
|
|
82
93
|
|
|
83
94
|
FleetView and the under-editor async widget are both enabled by default; set `asyncWidget: false` to keep only FleetView. Successful background completions stay quiet so inactive Pi tabs are not marked unread, while failed or paused completions still notify the originating session. Parallel runs show every active child independently. Chains with parallel groups keep their grouped shape in progress and results, so failed or paused agents stay visible next to completed ones. When a child is explicitly allowed to fan out with `tools: subagent` or `allowNestedSubagents: true`, its nested runs appear under that parent child in the main status tree instead of being hidden inside the child session.
|
|
84
95
|
|
|
@@ -94,16 +105,14 @@ Default keys:
|
|
|
94
105
|
- `x`/`Ctrl+O` — toggle tool details
|
|
95
106
|
- `r` — refresh
|
|
96
107
|
- `Esc` — close
|
|
97
|
-
- `Enter` — open the selected inspectable async child
|
|
108
|
+
- `Enter` — open the selected inspectable async child through the available Inspect plugin
|
|
98
109
|
- `s` — compose an acknowledged message to a selected live async child; Tab cycles `steer`, `follow_up`, and `auto`
|
|
99
110
|
- `D` — stop a selected child's top-level async run after confirmation
|
|
100
|
-
- `H` — open the selected active async child
|
|
111
|
+
- `H` — open the selected active async child through the available Inspect plugin
|
|
101
112
|
|
|
102
113
|
Set `fleetKeybindings` in the extension config to replace inspector-level keys when a terminal intercepts keys such as `PgUp`, `PgDn`, `Home`, or `End`. Prompt modes keep fixed keys such as `Esc`, `Enter`, `Tab`, and stop-confirmation `Y`/`N`.
|
|
103
114
|
|
|
104
|
-
`
|
|
105
|
-
|
|
106
|
-
Enter and `H` use the existing Herdr pane path. In a child-specific Herdr inspector, type ordinary guidance and press Enter to send it through the acknowledged steer channel; `steer <message>`, `status`, and `stop` remain available as explicit controls.
|
|
115
|
+
Enter and `H` use the available Inspect plugin. On macOS with Ghostty 1.3+ (TERM_PROGRAM=ghostty), this includes the other bundled open-only plugin using Ghostty's preview AppleScript API; status and close are unavailable because no binding is written. In a child-specific inspector, type ordinary guidance and press Enter to send it through the acknowledged steer channel; `steer <message>`, `status`, and `stop` remain available as explicit controls. The bundled Herdr plugin uses Herdr 0.7.5+.
|
|
107
116
|
|
|
108
117
|
Without a TUI, `/subagents-fleet` retains the textual `subagent({ action: "status", view: "fleet" })` fallback, and mutations use explicit commands: run `/subagents-stop` and pick from the selector, or use `/subagents-stop <run-id>` / `subagent({ action: "stop", id: "..." })` when you already know the id.
|
|
109
118
|
|
|
@@ -178,7 +187,6 @@ Async runs write machine-readable lifecycle artifacts for observability and work
|
|
|
178
187
|
- `status.json` powers the widget and `subagent({ action: "status" })` output.
|
|
179
188
|
- `events.jsonl` contains wrapper events plus child Pi JSON events annotated with run and step metadata, including correlated `subagent.steer.requested`, `scheduled`, `routed`, `queued`, `delivered`, `failed`, and `recovered` events plus failure/partial/recovery notices.
|
|
180
189
|
- `output-<n>.log` is a live human-readable tail.
|
|
181
|
-
- Fallback information is persisted so background runs are debuggable after completion.
|
|
182
190
|
|
|
183
191
|
For a top-level async run, `details.asyncDir` points at that directory; the final summary is written to Pi's subagent results directory as `<runId>.json`. Nested async runs use the same shape under the nested async root and are discoverable through status projections that read the nested-run registry. These files are append/update artifacts only; interactive foreground behavior is unchanged.
|
|
184
192
|
|
|
@@ -205,7 +213,9 @@ stop API.
|
|
|
205
213
|
|
|
206
214
|
### Status and result fields
|
|
207
215
|
|
|
208
|
-
The status/result fields are: `lifecycleArtifactVersion`, `runId`/`id`, `sessionId`, `mode`, `state`, `startedAt`, `lastUpdate`, `endedAt`, `durationMs`, `cwd`, `asyncDir`, `sessionFile`, `outputFile`, `workflowGraph`, `steps`, `results`, `totalTokens`, `totalCost`, `model`/`
|
|
216
|
+
The status/result fields are: `lifecycleArtifactVersion`, `runId`/`id`, `sessionId`, `mode`, `state`, `startedAt`, `lastUpdate`, `endedAt`, `durationMs`, `cwd`, `asyncDir`, `sessionFile`, `outputFile`, `workflowGraph`, `steps`, `results`, `totalTokens`, `totalCost`, `model`/`requestedModel`, `toolCount`, `turnCount`, optional `launchResolvedExtensions`, optional `runtimeAcknowledgedExtensions`, and nested `children` when a child is allowed to launch subagents.
|
|
217
|
+
|
|
218
|
+
`requestedModel` records the launch's requested model (the explicit `--model` override, else the agent's configured model) before registry normalization.
|
|
209
219
|
|
|
210
220
|
`launchResolvedExtensions` is parent-resolved launch intent only: it reports opaque extension identifiers and whether ambient extensions were disabled, without exposing raw extension paths or claiming the child runtime acknowledged that those extensions loaded.
|
|
211
221
|
|
|
@@ -270,7 +280,7 @@ Debug artifacts live under `{sessionDir}/subagent-artifacts/`, `.pi/subagents/ar
|
|
|
270
280
|
- `{runId}_{agent}.jsonl`
|
|
271
281
|
- `{runId}_{agent}_meta.json`
|
|
272
282
|
|
|
273
|
-
Metadata records timing, usage, exit code,
|
|
283
|
+
Metadata records timing, usage, exit code, the resolved model, and the resolved acceptance ledger with its parsed child report. A strictly guarded retained-session recovery after a verified compaction abort may continue once on that same model; it never selects another model.
|
|
274
284
|
|
|
275
285
|
For npm package projects, project-scoped artifacts need a `.npmignore` rule (or `.gitignore` when no `.npmignore` exists) or a `files` allowlist that does not include `.pi/subagents/`. pi-subagents warns at launch when these package settings can include the artifacts. Use `artifactDir: "session"` or `"temp"` to keep them outside the package worktree.
|
|
276
286
|
|
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Standalone background execution
|
|
2
|
+
|
|
3
|
+
Supported standalone target: **official Pi 0.85.1, Linux x64**. Keep its adjacent release assets with the executable. Other versions, operating systems, architectures and packagers are outside the fully validated support target; limited experimental Windows coverage is described below.
|
|
4
|
+
|
|
5
|
+
Pi's extension loader supplies its embedded SDK to `binary-bootstrap.ts`, which awaits the existing configured runner before exiting. Startup authorization, revival leases, controls, disposal and process-close observation remain shared with npm. Each independent run has its own host; native sessions inside that run share it. No per-session CLI protocol, runtime download/install, alternate SDK or foreground fallback is introduced. Npm Pi keeps its Node runner, peer aliases and detected npm `PI_PACKAGE_DIR` override (including refusal when no npm root exists).
|
|
6
|
+
|
|
7
|
+
Implementation and lifecycle fixtures derive from [@xz-dev](https://github.com/xz-dev)'s [PR #2049](https://github.com/nicobailon/pi-subagents/pull/2049), source commit `910807bfefcf9ee41d73fa25ec86dcd75ab8f4b2` (Xiangzhe, `xiangzhedev@gmail.com`). Integration retains the lifecycle contract and reduces commentary rather than removing its evidence gates.
|
|
8
|
+
|
|
9
|
+
## Experimental Windows host recognition
|
|
10
|
+
|
|
11
|
+
The resolver recognizes Bun's Windows virtual entrypoint prefixes, `B:/~BUN/` and `B:\~BUN\`, alongside `/$bunfs/`. It launches the real `process.execPath` (or the existing executable override). The `B:` prefix is virtual, not the installation drive; `pi-native.exe` is not a required executable name.
|
|
12
|
+
|
|
13
|
+
A local Windows x64 smoke passed with **xz-dev/pi `0.85.1-xz.169.1.gb5f4d0ff`, Bun 1.4.2**: a fresh async worker executed a read-only Git command, returned its result, delivered the native completion notification, and exited with code 0 and no remaining runner process. This is not validation of the official Windows distribution or every Bun-compiled Pi host. Windows remains **experimental**: the full standalone lifecycle matrix has not been validated there.
|
|
14
|
+
|
|
15
|
+
Node-hosted npm Pi keeps its existing runner path and is not affected by this virtual-entrypoint detection defect. Installing only the pi-subagents extension through npm does not change a Bun-compiled Pi host into an npm Pi host.
|
|
16
|
+
|
|
17
|
+
## Official binary gate
|
|
18
|
+
|
|
19
|
+
On Linux x64 with Node, npm, tar and bubblewrap installed, provision dependencies and the checksum-pinned release separately from execution:
|
|
20
|
+
|
|
21
|
+
```bash
|
|
22
|
+
npm ci --ignore-scripts
|
|
23
|
+
release_dir="$(mktemp -d)"
|
|
24
|
+
url="$(node -p 'require("./test/smoke/standalone-release.json").url')"
|
|
25
|
+
sha="$(node -p 'require("./test/smoke/standalone-release.json").archiveSha256')"
|
|
26
|
+
curl --fail --location --retry 3 "$url" --output "$release_dir/release.tar.gz"
|
|
27
|
+
printf '%s %s\n' "$sha" "$release_dir/release.tar.gz" | sha256sum --check -
|
|
28
|
+
tar -xzf "$release_dir/release.tar.gz" -C "$release_dir"
|
|
29
|
+
node test/smoke/standalone-matrix.mjs "$release_dir/pi/pi" "$(mktemp -d)/matrix"
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
The `official-standalone` CI job runs this gate. Both archive and executable hashes are pinned. Each of 18 modes gets a fresh stage with no filesystem core SDK/shim, empty installation caches and isolated network/PID namespaces. Bare-Bun SDK import must fail; accepted execution uses Pi's actual loader. Missing sandbox support fails rather than skips. Use disk-backed storage: retained stages can occupy several GiB.
|
|
33
|
+
|
|
34
|
+
The matrix covers public launch/notification, workflows, same-run concurrent sessions, parallel stop, targeted steer/interrupt, child/tool/run deadlines, missing bootstrap, post-spawn persistence/authorization failures, SDK initialization failure, malformed bootstrap input/EOF with an authorized positive control, and competing revival. The provider is deterministic, but SDK sessions, runner and public extension are real. Only startup-failure writes are faulted.
|
|
35
|
+
|
|
36
|
+
`matrix.json` records complete/partial results; `inputs.json` freezes source identities and every mode must use the same package hash. Inspect per-mode logs, `identity.json`, lifecycle/notification evidence, `status.json` and `process-terminal.json`. A persisted result is not exit proof: the gate separately awaits observed close, verifies dead PIDs before sandbox teardown and checks session shutdown/lease release. CI retains receipts and at most 32 MiB compressed lifecycle evidence. Contributor-head passes do not establish acceptance for a different integration snapshot.
|
|
37
|
+
|
|
38
|
+
For a focused diagnostic, use `node test/smoke/standalone-background.mjs "$release_dir/pi/pi" "$(mktemp -d)/check" bootstrap-errors` (or another matrix mode). A focused pass is not the complete gate.
|
|
39
|
+
|
|
40
|
+
## Npm regressions and local trial
|
|
41
|
+
|
|
42
|
+
Existing npm clean-install CI covers real SDK 0.85.1. The standalone CI job also checks the public npm launch path without execution-time network:
|
|
43
|
+
|
|
44
|
+
```bash
|
|
45
|
+
npm_checks="$(mktemp -d)"
|
|
46
|
+
node test/smoke/clean-install.mjs "$npm_checks/sdk" 0.85.1
|
|
47
|
+
node test/smoke/npm-background.mjs "$npm_checks/sdk" "$npm_checks/launch"
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
To try a checkout without replacing your installation, start a separate supported Pi process with an isolated agent directory:
|
|
51
|
+
|
|
52
|
+
```bash
|
|
53
|
+
PI_CODING_AGENT_DIR="$(mktemp -d)" "$release_dir/pi/pi" \
|
|
54
|
+
--no-extensions --no-skills --no-prompt-templates --extension "$PWD/index.ts"
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
This command targets a source checkout. The published npm package uses the compiled `index.js` entry instead.
|
|
58
|
+
|
|
59
|
+
Configure a provider in that isolated session, ask for a read-only background child and inspect its notification/run artifacts. This loads only the checkout for that process; it does not install the candidate or reuse normal credentials. Keep the parent alive for notifications.
|
package/docs/tool-reference.md
CHANGED
|
@@ -2,6 +2,8 @@
|
|
|
2
2
|
|
|
3
3
|
Parameters and actions for the `subagent` tool. These are what the LLM passes when it calls the tool; most users ask naturally or use slash commands instead.
|
|
4
4
|
|
|
5
|
+
Call `{ action: "guide", topic: "tool-reference" }` for this reference or `topic: "workflows"` for [workflow recipes](workflows.md). Use `topic: "agents"` for authoring, `topic: "missions"` for missions/schedules, and `topic: "watchdog"` for watchdog controls. Guide reads do not change the schema or grant authority.
|
|
6
|
+
|
|
5
7
|
## Execution examples
|
|
6
8
|
|
|
7
9
|
Chaining is code-driven through `workflowScript`. Use `await runs.run(...)` for sequential steps and `await runs.all([{ key, agent, task }, ...])` for ordinary parallel fanout. `runs.all` resolves to an ordered array, not a key map, so use indexes, destructuring, or `.map(...)`, not `results.<key>`. Do not read `.output` from an unawaited `runs.run` launch. Stored `runs.run` promises are only for the advanced rolling fanout pattern under [Workflow steering](#workflow-steering), where every promise is later observed with direct `await`, `Promise.race`, or `Promise.all`. Legacy top-level `chain`, `tasks`, and `parallel` inputs are not supported. Helper functions must be plain functions or explicit Promise chains. Nested `async function` helpers, async arrows, and async methods are rejected so child-launch tracking stays portable across Node and Bun. For permission-sensitive host calls, use an extension-owned named resource such as `{ workflow: "run-ci", args: { command: "npm test" } }`; raw public `workflowScript`/`workflowScriptPath` inputs have unknown resource provenance and cannot call `runs.host`. A resolved resource may internally use `runs.host(key, { kind: "command", command, timeoutMs, output?, role?, provider? })` within its authority ceiling; there is no per-step `cwd`, and commands and relative output paths use the workflow `cwd`. Set `cwd` on the outer `subagent({...})` request instead, or put a trusted directory change in the command (for example, `cd /path/to/worktree && npm test`).
|
|
@@ -10,6 +12,8 @@ Use `{ action: "validate", workflowScript }` to check statically decidable synta
|
|
|
10
12
|
|
|
11
13
|
Use `workflowScriptPath` instead of `workflowScript` to load the same JavaScript statement body from a file. The two fields are mutually exclusive. Relative paths resolve against the request `cwd`, and absolute paths pass through. The host reads the file before validation, scheduling, or sandbox execution. The workflow sandbox still has no filesystem access. Missing, unreadable, and empty files fail as file input errors.
|
|
12
14
|
|
|
15
|
+
Raw inline and file-backed scripts accept bounded plain-JSON `args`, including during `validate` and `schedule.create`. Omitted raw args become `{}`; supplied args are deeply frozen in the sandbox. Normalized args persist in run and schedule evidence for diagnosis and exact replay, so never include secrets. Args are data only and do not grant `runs.host` authority.
|
|
16
|
+
|
|
13
17
|
For permission-extension interoperability, use one of the package-owned named resources with bounded `args` instead of caller-supplied workflow text:
|
|
14
18
|
|
|
15
19
|
```js
|
|
@@ -20,9 +24,9 @@ For permission-extension interoperability, use one of the package-owned named re
|
|
|
20
24
|
The host resolves the script and authority internally and records bounded provenance in workflow details and receipts. Named resources cannot be combined with `agent`, `task`, `workflowScript`, or `workflowScriptPath`; user/project resource registries are not part of this first slice.
|
|
21
25
|
|
|
22
26
|
```js
|
|
23
|
-
{ workflowScriptPath: "workflows/review.js", cwd: "/path/to/project" }
|
|
24
|
-
{ action: "validate", workflowScriptPath: "workflows/review.js" }
|
|
25
|
-
{ action: "schedule.create", every: "6h", workflowScriptPath: "workflows/review.js" }
|
|
27
|
+
{ workflowScriptPath: "workflows/review.js", args: { target: "src/workflows" }, cwd: "/path/to/project" }
|
|
28
|
+
{ action: "validate", workflowScriptPath: "workflows/review.js", args: { target: "src/workflows" } }
|
|
29
|
+
{ action: "schedule.create", every: "6h", workflowScriptPath: "workflows/review.js", args: { target: "src/workflows" } }
|
|
26
30
|
```
|
|
27
31
|
|
|
28
32
|
```js
|
|
@@ -86,11 +90,13 @@ The complete plain-JSON inventory is validated before the first launch (maximum
|
|
|
86
90
|
|
|
87
91
|
| Param | Type | Default | Description |
|
|
88
92
|
|-------|------|---------|-------------|
|
|
89
|
-
| `agent` | string | - |
|
|
90
|
-
| `
|
|
93
|
+
| `agent` | string | - | One direct child or agent-management target. Workflow child agents are set inside `runs.run` or `runs.all`. |
|
|
94
|
+
| `task` | string | agent default | Direct child's task; requires `agent`, excludes `action` and workflow inputs. `agent` may also select a management target. |
|
|
95
|
+
| `action` | string | - | Offline workflow `validate`, agent management (including `guide`, `children.list`, and `refine`/`refine.show`/`refine.rollback`), lane evidence (`lane.status`, `lane.recordMerge`, `lane.recordSupersession`), mission (`mission.create/list/show/update/resolve-decision/attach-run/close`), Inspect actions (`inspector.command/open/status/close`), Herdr project pane (`project.open/status/close`), status/control, plan-only `worktree.cleanup`, schedule, watchdog, or doctor action. |
|
|
91
96
|
| `topic` | `overview \| workflows \| agents \| missions \| observability \| tool-reference \| configuration \| models \| watchdog \| extension-api` | `overview` | Packaged guide topic for `action: "guide"`. |
|
|
92
97
|
| `config` | object/string | - | Agent config for management create/update. |
|
|
93
|
-
| `context` | `fresh \| fork` | global or per-agent default, else `fresh` | Explicit `fresh` or `fork` overrides every workflow child. When omitted, [`defaultSubagentContext`](configuration.md#defaultsubagentcontext) wins over each agent's `defaultContext`;
|
|
98
|
+
| `context` | `fresh \| fork \| profile` | global or per-agent default, else `fresh` | Explicit `fresh` or `fork` overrides every workflow child. `profile` requires the selected agent's declared `defaultContext` and ignores config `defaultSubagentContext`; missing agent defaults fail. When omitted, [`defaultSubagentContext`](configuration.md#defaultsubagentcontext) wins over each agent's `defaultContext`; implicit fork falls back to fresh without a persisted parent session and leaf. Explicit fork is strict. Packaged `worker`, `oracle`, and `advisor` default to `fork`. |
|
|
99
|
+
| `model` | string | agent default | Call `{action:"models"}` first and copy an exact `provider/id`; bare ids resolve only if unique, and agent names are not model ids. A suffix such as `provider/id:high` (`off/minimal/low/medium/high/xhigh/max`) overrides agent thinking. The `thinking` field is only for `watchdog.configure`, ignored on dispatch. |
|
|
94
100
|
| `missionId` | string | - | Attach a workflow to an existing project mission instead of creating its default enclosing mission. |
|
|
95
101
|
| `mission` | object/false | auto-create | Override the default enclosing mission with `{ title \| summary, objective?, goal?, budget?, labels? }`. Set exactly one non-empty `title` or `summary`; `objective` and `labels` are optional. `goal` may only be `true`, requires `budget.tokens`, and enables continuation notices. Pass `false` for an intentionally ephemeral workflow with no mission for it or its children and no `state` global. Explicit mission persistence failures are strict. |
|
|
96
102
|
| `handoffPath` | string | - | Aggregate handoff manifest for `action: "worktree.discard"` or lane evidence actions, or optional explicit metadata for `action: "worktree.cleanup"`. |
|
|
@@ -100,20 +106,22 @@ The complete plain-JSON inventory is validated before the first launch (maximum
|
|
|
100
106
|
| `laneId` | string | - | Exact `runId` stored in the handoff manifest for `lane.status`, `lane.recordMerge`, or `lane.recordSupersession`. |
|
|
101
107
|
| `merge` | object | - | Attested merge evidence for `lane.recordMerge`; requires a positive PR number, full reviewed/merge SHAs, tree-equivalence and post-merge-check statuses, attestor, and timestamp. |
|
|
102
108
|
| `supersession` | object | - | Attested replacement-lane evidence for `lane.recordSupersession`; requires a different replacement lane id, attestor, and timestamp. |
|
|
103
|
-
| `focus` | boolean | false | Focus the newly split pane for `action: "inspector.open"` or `action: "project.open"`; not a standalone action. Panes open in the background unless you set `focus: true`. Existing saved project panes can be focused through the public project-pane API when Herdr reports a tab or workspace id. |
|
|
109
|
+
| `focus` | boolean | false | Focus the newly split host inspector pane for `action: "inspector.open"` or the new Herdr project pane for `action: "project.open"`; not a standalone action. `inspector.command` is read-only and does not contact Herdr or write a binding. Panes open in the background unless you set `focus: true`. Existing saved project panes can be focused through the public project-pane API when Herdr reports a tab or workspace id. |
|
|
104
110
|
| `view` | `fleet \| transcript` | - | Optional `status` view for the active fleet surface or transcript tail inspection. |
|
|
105
111
|
| `lines` | number | `80` | Maximum transcript lines for `action: "status", view: "transcript"`; capped at 500. |
|
|
106
112
|
| `agentScope` | `user \| project \| both` | `both` | Agent discovery scope. Project wins on collisions. |
|
|
107
113
|
| `capabilities` | boolean | `false` | With `action: "list"`, return compact prompt-free rows and `details.agentCapabilities` machine-readable records for each agent's declared/default routing capabilities. External CLI rows also include their command and passive local availability. |
|
|
108
|
-
| `async` | boolean | default-on | Background execution. Workflows default to background. `async:false` blocks the parent until completion
|
|
114
|
+
| `async` | boolean | default-on | Background execution. Workflows default to background. `async:false` blocks the parent until completion. A local foreground child runs inside the parent Pi process and never loads the parent's ambient extensions, but it does inherit the providers those extensions registered. A pane-native remote foreground child instead uses the remote machine's provider discovery and configuration. Agents that need MCP tools (`mcpDirectTools`, or MCP tools from an ambient adapter such as pi-mcp-adapter) must run as background children, which load them inside the detached runner process. |
|
|
109
115
|
| `chatProgress` | `auto \| off \| live-card` | `auto` | WorkflowScript chat projection. `auto` renders a live in-chat card only for watched foreground workflows in the same Git repository, including managed worktrees; it is off otherwise. Explicit `live-card` requires `async:false` and the same Git repository. Async workflows have no inline live card, so omit `chatProgress` or use `auto`/`off`; use `async:false` only when the parent must block. |
|
|
110
116
|
| `isolation` | `none \| worktree` | - | Workflow child isolation. `none` runs in the shared cwd and does not need Git. `worktree` requires a managed Git worktree. Do not combine it with a contradictory `worktree` value. |
|
|
111
117
|
| `baseRef` | string | `HEAD` | `HEAD` or a supported named ref such as `refs/heads/release`, `refs/tags/v1`, or `origin/main`. Full 40/64-character commit IDs and revision expressions such as `HEAD~1` are unsupported. The ref must resolve to a commit at worktree allocation; omitted values default to `HEAD` resolved at that time. Source-checkout cleanliness is still checked. For workflowScript, set it on the outer request as a default or on an individual `runs.run`/`runs.all` child to override it. |
|
|
112
|
-
| `timeoutMs` / `maxRuntimeMs` | number | config `timeoutMs`, else 30 min foreground / single-agent async | Optional run-level max runtime in milliseconds. When omitted, the global [`timeoutMs`](configuration.md#timeoutms) config provides the default; absent that, foreground and plain single-agent async runs fall back to 30 minutes, while composite async runs (chains, parallel tasks, workflows) stay unbounded at the top level. Expiration of this run-level deadline is terminal
|
|
118
|
+
| `timeoutMs` / `maxRuntimeMs` | number | config `timeoutMs`, else 30 min foreground / single-agent async | Optional run-level max runtime in milliseconds. When omitted, the global [`timeoutMs`](configuration.md#timeoutms) config provides the default; absent that, foreground and plain single-agent async runs fall back to 30 minutes, while composite async runs (chains, parallel tasks, workflows) stay unbounded at the top level. Expiration of this run-level deadline is terminal. |
|
|
113
119
|
| `toolTimeoutMs` | number | fast-tool default | Optional positive hard per-tool-call deadline in milliseconds. Precedence: call value → agent frontmatter → config → `PI_SUBAGENT_TOOL_TIMEOUT_MS`. The timer starts on `tool_execution_start`, clears on the matching `tool_execution_end`, and terminates the run with `timedOut: true` if the tool remains open. When omitted, known-fast built-in tools get a five-minute default; long-running tools get attention notices but no hard default. It never extends the run deadline; `contact_supervisor`, `intercom`, and `bg_wait` are exempt. |
|
|
120
|
+
| `checkpointBeforeDeadlineMs` | number | none | Async single-agent runs only. The runner requests that the child "checkpoint and stop" this many milliseconds before the run deadline (finish the current tool call, report changed files, build/test state, remaining work, commit/PR state; start no new work). This best-effort steer uses the normal steering lifecycle at the next tool boundary, so the receipt is visible in status and events; the ordinary deadline kill still applies. Precedence: call value → config `checkpointBeforeDeadlineMs`. Disarmed when the deadline leaves under one second before the checkpoint. |
|
|
114
121
|
| `toolBudget` | object | none | Optional child tool-call budget `{ soft?, hard, block? }`. At `soft` the child is nudged to finalize. After `hard`, configured tools are blocked; `block` defaults to `read`, `grep`, `find`, and `ls`, while `"*"` blocks every tool call. Final assistant text is never blocked. |
|
|
115
122
|
| `usageBudget` | object | none | Optional root-only reported-usage budget `{ tokens?: { soft?, hard }, costUsd?: { soft?, hard } }`. Soft limits are status-only. Hard limits prevent later child launches after reported usage is reconciled; already-running children are not stopped and no reservations are made. |
|
|
116
|
-
| `cwd` | string | runtime cwd | Override working directory. |
|
|
123
|
+
| `cwd` | string | runtime cwd | Override working directory. With `machine`, the directory on that machine. |
|
|
124
|
+
| `machine` | string | - | Herdr saved machine (label or profile id) for external-cli agents; see [agents.md](agents.md#running-external-cli-agents-on-a-herdr-saved-machine). |
|
|
117
125
|
| `maxOutput` | object | 200KB, 5000 lines | Final output truncation limits. |
|
|
118
126
|
| `artifacts` | boolean | true | Write debug artifacts. |
|
|
119
127
|
| `includeProgress` | boolean | false | Include full progress in result. |
|
|
@@ -126,13 +134,13 @@ The complete plain-JSON inventory is validated before the first launch (maximum
|
|
|
126
134
|
|
|
127
135
|
As a conservative orchestration policy, do not set a hard `toolBudget` or tight `usageBudget` on implementation workers, fix workers, reviewers with edit authority, or other mutation-capable children. A default tool budget blocks read/search tools rather than mutation tools, and reported usage has no reservation model, so neither tool-call counts nor token/cost totals measure whether a delivery slice is buildable or safe to hand off. Hard caps remain appropriate for explicitly read-only scouts, reviewers, and validators.
|
|
128
136
|
|
|
129
|
-
Bound writer work with a narrow task and an outer `timeoutMs` or `maxRuntimeMs` that leaves enough margin for the slice. An elapsed timeout is not a mutation-safe boundary and may still signal a child during tool work.
|
|
137
|
+
Bound writer work with a narrow task and an outer `timeoutMs` or `maxRuntimeMs` that leaves enough margin for the slice. An elapsed timeout is not a mutation-safe boundary and may still signal a child during tool work. Request a checkpoint after the current tool returns that records changed files, build/test state, and commit or PR state; for async single-agent runs, set `checkpointBeforeDeadlineMs`, otherwise steer by hand.
|
|
130
138
|
|
|
131
139
|
### Fork context details
|
|
132
140
|
|
|
133
141
|
Explicit `context: "fork"` fails fast when the parent session is not persisted, the current leaf is missing, or the branched child session cannot be created. By contrast, global `defaultSubagentContext: "fork"` and agent-level `defaultContext: fork` are preferences: when the parent has no persisted session file or current leaf yet, the launch uses `fresh` immediately instead of failing and requiring a retry. Global `defaultSubagentContext: "fresh"` starts fresh. Explicit `context: "fresh"` always wins over both preferences.
|
|
134
142
|
|
|
135
|
-
When the inherited transcript contains signed Anthropic `thinking` / `redacted_thinking` blocks, `pi-subagents` strips those provider-private blocks from the forked child session
|
|
143
|
+
When the inherited transcript contains signed Anthropic `thinking` / `redacted_thinking` blocks, `pi-subagents` strips those provider-private blocks from the forked child session: a thinking signature is bound to the session that produced it and cannot be replayed into a branch. The child keeps its requested thinking level and reasons fresh from its first turn; sanitizing the inherited transcript is not a downgrade. Explicit `context: "fork"` never silently downgrades to `fresh`.
|
|
136
144
|
|
|
137
145
|
In workflow runs that omit `context`, each `runs.run` child follows the global `defaultSubagentContext` when set, then its own `defaultContext`. Without the global setting, a fresh-default scout can run fresh beside a fork-default worker. If the parent session file or current leaf is not available yet, implicit fork-default children run fresh. Pass explicit `context: "fork"` or `context: "fresh"` when you intentionally want one context for every child.
|
|
138
146
|
|
|
@@ -140,7 +148,7 @@ In workflow runs that omit `context`, each `runs.run` child follows the global `
|
|
|
140
148
|
|
|
141
149
|
`runs.steer(key, message, options?)` targets a stable key already launched by `runs.run` or `runs.all`. It does not accept a raw run id. Options are `mode?: "steer" | "follow_up" | "auto"`, `index?: number`, and `ackTimeoutMs?: number`. The promise returns `{ key, state, requestId?, deliveryStatus?, targets?, error? }`, where `state` is `queued`, `delivered`, `missed`, or `failed`.
|
|
142
150
|
|
|
143
|
-
The workflow trace records the attempt and receipt. Always await, return, or include the promise in an awaited standard Promise combinator. Unawaited steering calls reject workflow completion after the side effect settles. `Promise.race` remains the rolling primitive. Foreground children are steered through their in-process session (`steer` and `auto`
|
|
151
|
+
The workflow trace records the attempt and receipt. Always await, return, or include the promise in an awaited standard Promise combinator. Unawaited steering calls reject workflow completion after the side effect settles. `Promise.race` remains the rolling primitive. Foreground children are steered through their in-process session (`steer` and `auto` report `delivered` when that transport accepts the input; `follow_up` reports `queued` when accepted into Pi's queue). Async children use the file control inbox and report correlated consumption by the child, not merely inbox acceptance. Steering recovery is disabled in both cases.
|
|
144
152
|
|
|
145
153
|
For advanced rolling fanout, keep the launched `runs.run` promises in ordinary JavaScript data only when every promise is later observed with direct `await`, `Promise.race`, or `Promise.all`. `Promise.race` gives the next completed child, `runs.steer` can challenge a still-running keyed sibling, and `Promise.all` collects the rest. No separate `runs.start`, `runs.next`, or `runs.collect` API is exposed.
|
|
146
154
|
|
|
@@ -220,7 +228,6 @@ Agent definitions are not loaded into context by default. Management actions let
|
|
|
220
228
|
inheritGlobalContext: false,
|
|
221
229
|
inheritSkills: false,
|
|
222
230
|
model: "anthropic/claude-sonnet-4",
|
|
223
|
-
fallbackModels: ["openai-codex/gpt-5.6-luna:low", "anthropic/claude-haiku-4-5"],
|
|
224
231
|
tools: "read, bash, mcp:github/search_repositories",
|
|
225
232
|
extensions: "",
|
|
226
233
|
skills: "parallel-scout",
|
|
@@ -247,7 +254,7 @@ Agent definitions are not loaded into context by default. Management actions let
|
|
|
247
254
|
|
|
248
255
|
Rules:
|
|
249
256
|
|
|
250
|
-
- `capabilities: true` changes `action: "list"` to compact one-line rows and adds `details.agentCapabilities: { agents, restrictedCount, capabilityCeilingSources? }`. Each agent row includes source, aliases, runner type/capabilities, tools, MCP direct tools, mutation tools, model/thinking
|
|
257
|
+
- `capabilities: true` changes `action: "list"` to compact one-line rows and adds `details.agentCapabilities: { agents, restrictedCount, capabilityCeilingSources? }`. Each agent row includes source, aliases, runner type/capabilities, tools, MCP direct tools, mutation tools, model/thinking, default async/timeout, declared acceptance policy/role, output path/mode, skills/extensions, and whether the current capability ceiling allows execution. External CLI rows include `runner.command`, `runner.available`, and a bounded `runner.unavailableReason` when passive PATH/PATHEXT/X_OK lookup cannot find the command. It never includes an agent's system prompt. Rows show declared/default capabilities and command discoverability, not authentication, version compatibility, or successful launch; launch preflight remains authoritative.
|
|
251
258
|
- `create` uses `config.scope`, not `agentScope`.
|
|
252
259
|
- `config.name` is the local frontmatter name; optional `config.package` registers the runtime name as `{package}.{name}` and is saved as separate `name` and `package` frontmatter.
|
|
253
260
|
- `config.aliases` accepts a comma-separated string, string array, or `false` to clear aliases. Aliases resolve to the canonical agent name for execution and are shown by `list`/`get`.
|
|
@@ -260,6 +267,10 @@ Rules:
|
|
|
260
267
|
|
|
261
268
|
`refine`, `refine.show`, and `refine.rollback` manage project-local refinement overlays for one agent. `/subagents-refine <agent>` is the slash equivalent of `refine`. See [agents.md](agents.md#refinement-overlays) for behavior and storage.
|
|
262
269
|
|
|
270
|
+
### Schedule controls
|
|
271
|
+
|
|
272
|
+
Use `schedule.create` with `workflowScript` or `workflowScriptPath`, not a direct child. `at` accepts a delay like `+10m` or an ISO timestamp with timezone; `every` accepts fixed intervals. `sessionOnly:true` binds restoration/execution to the creating session file; omitted/false is project-wide. Recurring `quiet:true` keeps successful automatic fires visible without a parent turn; failed, stopped or paused runs still wake the parent. One-shot `at` and manual `schedule.run` stay noisy unless that launch passes `quiet:true`. See [missions and schedules](missions.md#schedules) for examples and list/show/history/pause/resume/run/run-due/delete. Calendar selectors (`on`, `timezone`) and schedule mission attachment are deferred. `baseRef` resolves only at worktree allocation and still requires a clean source checkout.
|
|
273
|
+
|
|
263
274
|
## Lane merge evidence and cleanup eligibility
|
|
264
275
|
|
|
265
276
|
Lane evidence actions update an existing parallel handoff manifest at an explicit update boundary. They do not verify GitHub state, run Git commands, or remove worktrees. Pass the manifest path and its exact `runId` as `laneId`:
|
|
@@ -302,7 +313,7 @@ The manifest stores one of these fail-closed eligibility states: `active` (an ow
|
|
|
302
313
|
|
|
303
314
|
A failure in the subagent workflow, child launch, prompt runtime, extension loading, or child tooling setup is a lane infrastructure blocker, not permission to silently change execution mode. Stop and report the exact failure, run/status, and repo/cwd/worktree/branch/ref state. Before a same-protocol retry or asking the owner, verify the worktree is clean or capture the partial diff. Retry or fix the `subagent` path only through a clear same-protocol action.
|
|
304
315
|
|
|
305
|
-
For backlog lanes and other subagent-governed workflows, external/foreground/CLI fallback requires explicit owner approval. Do not silently switch to `interactive_shell`, `pi -ne`, Codex/Claude/Cursor CLI, a foreground agent, or another external mode. `interactive_shell` remains valid when the user explicitly requests visible foreground/CLI work or the task is outside the governed subagent protocol. Pi core may print a generic `pi -ne` extension-load hint; that out-of-repo hint is not protocol-approved fallback.
|
|
316
|
+
For backlog lanes and other subagent-governed workflows, external/foreground/CLI fallback requires explicit owner approval. Do not silently switch to `interactive_shell`, `pi -ne`, Codex/Claude/Cursor CLI, a foreground agent, or another external mode. `interactive_shell` remains valid when the user explicitly requests visible foreground/CLI work or the task is outside the governed subagent protocol. Pi core may print a generic `pi -ne` extension-load hint; that out-of-repo hint is not protocol-approved fallback. A verified compaction abort may continue the retained child once on its already resolved model; it never selects another model.
|
|
306
317
|
|
|
307
318
|
```ts
|
|
308
319
|
subagent({ action: "status" })
|
|
@@ -356,9 +367,9 @@ subagent({ action: "doctor" })
|
|
|
356
367
|
|
|
357
368
|
### steer
|
|
358
369
|
|
|
359
|
-
`steer` waits up to three seconds for a correlated
|
|
370
|
+
`steer` waits up to three seconds for a correlated receipt and returns a request id with `delivered`, `scheduled`, `pending`, `partial`, `recovered`, or `failed` plus per-child states. The receipt also has `deliveryStatus: "delivered" | "queued"`. For async runs, delivery means the child consumed the correlated user input; foreground delivery means the in-process Pi transport accepted it. Neither means model compliance. A pending indexed child returns `scheduled`.
|
|
360
371
|
|
|
361
|
-
The optional `mode` is `steer` by default and keeps the current interrupt behavior. `follow_up` waits for the next turn boundary. `auto` uses the same native steer delivery path as `steer`, without automatic pause-and-revive recovery after a missed acknowledgment. The retained revival-brief queue holds 20 messages and returns a clear error when full; this is not a live follow-up queue bound.
|
|
372
|
+
The optional `mode` is `steer` by default and keeps the current interrupt behavior. `follow_up` waits for the next turn boundary. `auto` uses the same native steer delivery path as `steer`, without automatic pause-and-revive recovery after a missed acknowledgment. The retained revival-brief queue holds 20 messages and returns a clear error when full; this is not a live follow-up queue bound. A live follow-up acknowledgment reports queue acceptance, not consumption. Async runs later record correlated consumption or fail unconsumed requests at settlement; foreground follow-ups have no later correlated receipt. A `follow_up` sent to a completed retained workflow child becomes the first brief for its next `resume`; it does not revive the child by itself.
|
|
362
373
|
|
|
363
374
|
Only a top-level single run may interrupt after the acknowledgment deadline and recover after a further 15-second pause/revival bound; durable multi-child and nested runs never auto-interrupt. Recovery launches a replacement only after the source is confirmed paused, a valid persisted session exists, and deadline, turn, and tool budgets remain. It preserves the original child contract and remaining limits; otherwise the source stays paused with an explicit failure. Late acceptance is recorded but cannot cancel committed recovery.
|
|
364
375
|
|
|
@@ -370,6 +381,8 @@ The `/subagents-steer <run-id> [--child <child-id>] <message>` slash command is
|
|
|
370
381
|
|
|
371
382
|
Every run resolves an effective acceptance policy. Callers may omit `acceptance` for the inferred default, or set it on single runs, top-level parallel task items, chain steps, static parallel tasks, and dynamic fanout templates.
|
|
372
383
|
|
|
384
|
+
Prefer an inline JSON object. JSON-encoded object strings are tolerated only during input normalization; invalid strings fail closed. `true` is invalid. Supported evidence kinds are `changed-files`, `tests-added`, `commands-run`, `validation-output`, `residual-risks`, `no-staged-files`, `diff-summary`, `review-findings`, and `manual-notes`. For example: `{level:"checked",evidence:["commands-run","changed-files"],review:{required:true}}`. Evidence levels end at `verified`; independent review is a separate gate, not a stronger evidence level.
|
|
385
|
+
|
|
373
386
|
```ts
|
|
374
387
|
{
|
|
375
388
|
agent: "worker",
|
|
@@ -400,7 +413,7 @@ Acceptance evidence levels are `auto`, `none`, `attested`, `checked`, and `verif
|
|
|
400
413
|
Review is a separate gate configured with `acceptance.review`:
|
|
401
414
|
|
|
402
415
|
- Async, risky, and dynamic writer contexts infer checked evidence plus `review: { agent: "reviewer", required: true }`.
|
|
403
|
-
-
|
|
416
|
+
- Tasks classified as read-only infer no acceptance by default, including reviews of release, migration, or security work; those topics do not turn a read-only task into implementation. With role metadata omitted, unknown risk-topic tasks retain their gate even when the agent name suggests a reviewer. Explicit acceptance requests still apply.
|
|
404
417
|
- Normal writer tasks infer checked evidence without review.
|
|
405
418
|
|
|
406
419
|
Agent frontmatter or `subagents.agentOverrides` may set `acceptanceRole: "read-only" | "writer"` for ambiguous tasks. Explicit task mutation or no-edit intent wins over that role, while omitted metadata preserves the existing reviewer/scout/worker name heuristics. The role affects acceptance inference only and does not change tool access.
|
|
@@ -476,7 +489,7 @@ async: true
|
|
|
476
489
|
|
|
477
490
|
Supported: status artifacts, stdout/stderr logs, timeout, and stop. Full stdout and stderr are written to log files, while the in-memory final stdout response and stderr error are limited to their last 64 KiB.
|
|
478
491
|
|
|
479
|
-
Intentionally unsupported: native Pi child options such as model override, structured output, acceptance/agent contract, tool budgets, fast mode, fork context, skills, or native Pi tools unless the runner explicitly implements them. Foreground/clarify, steer/resume/interrupt-as-pause, nested subagents
|
|
492
|
+
Intentionally unsupported: native Pi child options such as model override, structured output, acceptance/agent contract, tool budgets, fast mode, fork context, skills, or native Pi tools unless the runner explicitly implements them. Foreground/clarify, steer/resume/interrupt-as-pause, and nested subagents are also unsupported.
|
|
480
493
|
|
|
481
494
|
## Session sharing
|
|
482
495
|
|
package/docs/watchdog.md
CHANGED
|
@@ -6,11 +6,12 @@ The watchdog is an opt-in second model that reviews what the agent just did and
|
|
|
6
6
|
|
|
7
7
|
| Timing | Trigger | Gate | Delivery |
|
|
8
8
|
|---|---|---|---|
|
|
9
|
-
| Boundary review | `agent_end` of every main or child turn | Repo changed |
|
|
9
|
+
| Boundary review | `agent_end` of every main or child turn | Repo changed | Routed by finding importance; high is steered to the model, low/medium are persisted for the user only |
|
|
10
|
+
| Main activity review | `agent_end`, with `clarification: true` | New delivered orchestration evidence; at most one additional review per user prompt | Same warning/clarification path, even without local edits |
|
|
10
11
|
| Cadence review | Every `cadence.everyNTools` tool results, minimum 5 | Opt-in | Steered after the current tool, before the next step |
|
|
11
12
|
| LSP pre-pass | Before boundary review | Changed TypeScript/JavaScript files | Diagnostics become watchdog findings without a model call |
|
|
12
13
|
|
|
13
|
-
Boundary reviews coalesce a turn's edits into one final-state review. Unchanged or reverted diffs are skipped
|
|
14
|
+
Boundary reviews coalesce a turn's edits into one final-state review. Unchanged or reverted diffs are skipped unless the main-only activity opt-in below admits new evidence; `.pi/subagents/` and `tmp/` artifacts remain excluded. In orchestrated runs, each writing child reviews its own worktree and the parent reviews the aggregate diff after child changes land. There are no idle timer reviews. Cadence monitoring is inspired by [Scopey](https://github.com/ArchAstro/scopey).
|
|
14
15
|
|
|
15
16
|
Children get the same boundary, cadence, and LSP behavior. Child cadence resolves from `children.overrides.<agent>.cadence`, then `children.cadence`, then top-level `cadence`:
|
|
16
17
|
|
|
@@ -37,37 +38,37 @@ That means: main every 10 tools, worker every 5, other children every 20, review
|
|
|
37
38
|
|
|
38
39
|
## What you see
|
|
39
40
|
|
|
40
|
-
Every finding
|
|
41
|
+
Every finding requires `importance: low | medium | high`. Low and medium are persisted for the user but excluded from model context and continuations. High findings retain model-visible delivery. Severity independently controls thresholds and acceptance, so a low-importance blocker still blocks acceptance. Clean reviews show nothing.
|
|
41
42
|
|
|
42
43
|
```
|
|
43
44
|
you ─▶ agent turn ─▶ edits repo ─▶ agent_end ─▶ watchdog review
|
|
44
45
|
├─ clean: turn ends
|
|
45
|
-
└─ warning:
|
|
46
|
+
└─ warning: low/medium user entry, or high steered message
|
|
46
47
|
```
|
|
47
48
|
|
|
48
|
-
Collapsed warnings show the title and evidence line. Expanded warnings show evidence, recommended action, category, and source:
|
|
49
|
+
Collapsed warnings show the title and evidence line. Expanded warnings show evidence, recommended action, importance, category, and source:
|
|
49
50
|
|
|
50
51
|
```
|
|
51
52
|
● Subagent watchdog Blocker (displayed): Claims tests passed without running them
|
|
52
53
|
Evidence: The transcript claims `npm test` passed but no test command appears in the tool log.
|
|
53
54
|
Recommended action: Run the focused test before finishing.
|
|
54
|
-
Category: Test Gap · Source: main
|
|
55
|
+
Importance: High · Category: Test Gap · Source: main
|
|
55
56
|
```
|
|
56
57
|
|
|
57
58
|
When consecutive boundary reviews raise the same warning, the agent is not making progress. After `stalemateRepeats` identical warnings in a row (default 3), the warning is shown as `stalemate`, no continuation is triggered, and the turn ends. Your next prompt resets the count.
|
|
58
59
|
|
|
59
60
|
Child watchdog findings are lifted into the parent in three ways:
|
|
60
61
|
|
|
61
|
-
-
|
|
62
|
+
- Internal/user inspection retains the last 20 findings, including importance and full details. Parent model results may include only high-importance findings.
|
|
62
63
|
- The acceptance runtime check `watchdog-blocker` fails on blockers that are unaddressed or stalemate.
|
|
63
|
-
- Completion notices include
|
|
64
|
+
- Completion notices may include high-importance concerns and blockers; low/medium finding text is omitted.
|
|
64
65
|
|
|
65
66
|
`/subagents-watchdog status` shows setting sources, enabled state, runtime state, review trigger, scope, cadence, LSP status, selected model/thinking, child overrides, timeout, stalemate count, launch-rule count, review backend, last warning, changed paths, and config errors when present.
|
|
66
67
|
|
|
67
68
|
## What the reviewer is given
|
|
68
69
|
|
|
69
70
|
- **Turn delta** with changed repo paths. Over-long input keeps the first 6,000 characters and the tail.
|
|
70
|
-
- **Current scope** (`scope.enabled`, default on): bounded real user prompts
|
|
71
|
+
- **Current scope** (`scope.enabled`, default on): bounded real user prompts. Side questions are additive; only explicit changes supersede older requirements. Scope survives compaction within the current session.
|
|
71
72
|
- **`watchdog_diff`** when inside git: diff since the session-start commit, including later commits, plus untracked paths to inspect with `read`; accepts `path` and `stat:true`.
|
|
72
73
|
- **`WATCHDOG.md`** standing instructions, read fresh on every review: `<project>/.pi/WATCHDOG.md` first, then `~/.pi/agent/WATCHDOG.md`, capped at 8,000 characters. Set `guidance.watchdogMd: false` to ignore them.
|
|
73
74
|
- **LSP diagnostics** from `typescript-language-server`, auto-detected in `node_modules/.bin` or `PATH`; it is never installed and never run over the whole workspace. Errors become blockers, warnings concerns, and info/hints stay in status.
|
|
@@ -87,7 +88,9 @@ One model setting serves both boundary and cadence reviews per endpoint. Use a s
|
|
|
87
88
|
/subagents-watchdog on
|
|
88
89
|
```
|
|
89
90
|
|
|
90
|
-
|
|
91
|
+
When a main watchdog model is configured (including a session override), recommendations keep that model and its effective thinking level rather than judging its strength or independence. An unavailable or unauthenticated configured model is reported, not replaced. Without a configured main model, the recommendation remains Opus 4.8 or GPT 5.5 at thinking high, whichever your main session is not using and is authenticated.
|
|
92
|
+
|
|
93
|
+
`session model recommended` changes only this session. `model recommended` explicitly saves the recommendation to **user settings**, affecting other projects without overrides; it does not change project settings. Project and session overrides still take precedence. Use an explicit model to replace a configured choice. Saving a model does not enable the watchdog; use `on` separately.
|
|
91
94
|
|
|
92
95
|
```json
|
|
93
96
|
{
|
|
@@ -105,8 +108,34 @@ The recommendation is Opus 4.8 or GPT 5.5 at thinking high, whichever your main
|
|
|
105
108
|
|
|
106
109
|
Omit `main.model` to inherit the session model and thinking level. A `main.model` without a thinking suffix or `main.thinking` runs with thinking off, so prefer `:high` for the strong pairing.
|
|
107
110
|
|
|
111
|
+
The watchdog resolves one reviewer model and makes one review call. Unavailable models fail visibly; rate limits, quota, authentication, provider timeouts, findings, clarification, cancellation, and the overall watchdog deadline never switch models automatically. An inherited model keeps the current session model and thinking level.
|
|
112
|
+
|
|
108
113
|
Agents can call `subagent({ action: "watchdog.recommend-model" })` and `subagent({ action: "watchdog.configure", model: "recommended", scope: "session" | "user" | "project" })`. They should use `scope: "session"` unless you ask for a lasting default.
|
|
109
114
|
|
|
115
|
+
## Optional main-session clarification
|
|
116
|
+
|
|
117
|
+
Use `watchdog_warn` directly for evidence-backed reminders of forgotten authorized work; a question is not a prerequisite. Distinguish forgotten work from dependencies still pending or explicit holds. Use clarification when task status or intent is genuinely unclear. The orchestrator remains owner of its task/lane board.
|
|
118
|
+
|
|
119
|
+
With this opt-in, completed `turn_end` events retain a recent actual-activity tail: paired calls/results for `subagent` dispatch (no action), `subagent` actions `status`, `resume`, `interrupt`, `steer`, `stop`, `bg_wait`, and `subagent_supervisor` actions `pending`, `list`, `reply`. Pairing requires the same tool name and exact tool-call ID; raw results, unrelated tool names and watchdog management actions do not qualify. Each activity entry is bounded to 3,000 characters, with a 6,000-character recent tail retained across ordinary new prompts and skipped edit boundaries. Session replacement, compaction, shutdown or disabling clears it. This is observed text, not an inferred task board.
|
|
120
|
+
|
|
121
|
+
New unreviewed activity permits at most one additional boundary review per user prompt even with no local edit. Side questions keep earlier authorized task evidence available; they do not themselves trigger a model call. Activity gathered after that prompt's extra review remains available for the next prompt. Warning continuations cannot supply fresh triggering activity. No polling, task scheduling, cross-worktree scans or idle calls are added.
|
|
122
|
+
|
|
123
|
+
**Visibility limit:** external task/gate completions are visible when returned through those parent tool results. Standalone native completion notifications, arbitrary custom messages, direct shell/CI output, and events not delivered to the parent through these contracts are not ingested by this activity tail. Existing scope retains at most eight prompts (2,000 characters each); new streaming user input cancels an active review but is not added to scope unless `before_agent_start` fires. Reminders depend on retained evidence and model judgment, not an exhaustive view of running work.
|
|
124
|
+
|
|
125
|
+
Set `subagents.watchdog.clarification: true` in Pi settings alongside `enabled: true`. It defaults to `false` and applies **only to the main watchdog**, not child watchdogs or child permission arbitration.
|
|
126
|
+
|
|
127
|
+
```json
|
|
128
|
+
{ "subagents": { "watchdog": { "enabled": true, "clarification": true } } }
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
At an eligible activity or repo-edit boundary, the reviewer may use `watchdog_ask` for one focused question when missing orchestrator context prevents a concrete judgment. It cannot ask during cadence reviews, after an accepted warning, or during stalemate. There is at most one question per real user prompt.
|
|
132
|
+
|
|
133
|
+
The tool **yields and ends that review**. A visible question with concrete evidence steers the main session into Pi's native automatic continuation after the boundary hook returns. The orchestrator handles the context as needed and continues; no answer, receipt, deadline or follow-up review is required or tracked. Questions are **not approval, permission, or warnings**.
|
|
134
|
+
|
|
135
|
+
Asking consumes the current review evidence, so an unchanged Git-backed boundary does not immediately review it again. Later reviews use the normal edit, cadence and bounded activity triggers, with the same read-only tools, warning thresholds, budgets and stalemate protections. Non-Git observed edits can also prompt a question: there is no cross-answer evidence guarantee to verify. Disabling watchdog or clarification, new user input, model changes and session lifecycle resets cancel applicable active reviews and suppress stale results; already delivered questions remain ordinary transcript messages.
|
|
136
|
+
|
|
137
|
+
Reviews retain the existing `agentEndTimeoutMs`. Questions and evidence are capped at 1,000 and 2,000 characters; scope, activity and delta share the existing 24,000-character input limit. Enabled cost adds at most one activity boundary review and one question-triggered continuation per prompt, not a dedicated answer review. Disabled execution does not collect activity or add polling, model calls or reviewer prompt/tool content. Child warning messaging and permission decisions remain unchanged.
|
|
138
|
+
|
|
110
139
|
## Child watchdogs
|
|
111
140
|
|
|
112
141
|
Opt in under `subagents.watchdog.children`. `model` and `thinking` set the default child watchdog; `overrides.<agent>` can set `model`, `thinking`, `enabled`, or `cadence` per role.
|