@tranhoangnguyen0310/pi-flow-external 2.3.1-external.0 → 2.4.1-external.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -2,20 +2,21 @@
2
2
 
3
3
  ## pi-flow external contract
4
4
 
5
- This fork changes the original pi-flow contract: `Agent` is not a generic Pi subagent launcher. It delegates only to external Claude Code, Codex CLI, Antigravity, and Grok Build CLI harnesses, plus named Pi harness configurations (below) — in-process, per-model configs registered by the user, not a spawned CLI.
5
+ This fork changes the original pi-flow contract: `Agent` is not a generic Pi subagent launcher. It delegates only to external Claude Code, Codex CLI, Antigravity, Grok Build CLI, and Muse Code harnesses, plus named Pi harness configurations (below) — in-process, per-model configs registered by the user, not a spawned CLI.
6
6
 
7
7
  - Ordinary driver tools: `Agent`, read-only `external_help`, `external_runs`, and optional `workflow`. `pi_flow_profile_create` and `pi_flow_harness_create` are active only inside `/external profile create`.
8
8
  - User operations use `/external`; `/pi-flow-profile create` is a temporary deprecated alias.
9
9
  - Extension-owned concurrency, timeout, and global default-harness settings live in `$PI_CODING_AGENT_DIR/pi-flow-external/settings.json`. Profile backend/model/thinking metadata remains in `subagents/*.md`. Named Pi harness configurations live in the sibling `$PI_CODING_AGENT_DIR/pi-flow-external/harnesses.json` registry (one shared file, `pi-<label>` keys pinning `model`/`thinking`). Registry writes are atomic per file, but there is no inter-process lock: concurrent writers adding distinct names from separate sessions are last-writer-wins and may drop one update. This is an accepted v1 limitation shared with `settings.json` and `subagents/*.md`, not a bug to re-report. A trusted project may override `defaultHarness` via `.pi/pi-flow-external/settings.json` (that key only, read-only to the extension); explicit call harness > project default > global default. `defaultHarness` may itself name a registered `pi-*` harness.
10
- - Every `Agent` call requires `description`, `prompt`, and either `role` with optional `harness` (`agy`, `claude`, `codex`, `grok`, or a registered `pi-*` harness name) or the legacy exact-profile `subagent_type`. The selectors cannot be combined. Without `harness`, use the effective default harness; if that default names a `pi-*` harness that is no longer registered, the call fails with an actionable error rather than silently falling back to `agy`.
11
- - Valid profiles come from `~/.pi/agent/subagents/*.md` and set `backend: claude`, `backend: codex`, `backend: agy`, `backend: grok`, or (only when also declaring `harness: <registered pi-* name>`) `backend: pi`. Use matching names such as `claude-*`, `codex-*`, `agy-*`, `grok-*`, or `<harness>-*`; the assisted creator enforces that convention.
12
- - A default roster ships with the extension and seeds once on session start: five code-oriented roles (explorer, planner, implementer, reviewer, qa) plus the generalist worker, one storage profile per CLI backend across four backends (24 files, but six advertised roles). Seeding never overwrites existing files and respects later deletions. Roles outside the default roster are user-created via `/external profile create`. Every registered named Pi harness automatically and immediately gets the same six roles too, synthesized in-memory from the identical canonical role source — no per-harness file needed, and no on-disk seeding for pi harnesses. An installation seeded before Grok support existed migrates forward by adding only the six new `grok-*` profiles on its next seed pass, never resurrecting a deleted or customized claude/codex/agy profile.
10
+ - Every `Agent` call requires `description`, `prompt`, and either `role` with optional `harness` (`agy`, `claude`, `codex`, `grok`, `muse`, or a registered `pi-*` harness name) or the legacy exact-profile `subagent_type`. The selectors cannot be combined. Without `harness`, use the effective default harness; if that default names a `pi-*` harness that is no longer registered, the call fails with an actionable error rather than silently falling back to `agy`.
11
+ - Valid profiles come from `~/.pi/agent/subagents/*.md` and set `backend: claude`, `backend: codex`, `backend: agy`, `backend: grok`, `backend: muse`, or (only when also declaring `harness: <registered pi-* name>`) `backend: pi`. Use matching names such as `claude-*`, `codex-*`, `agy-*`, `grok-*`, `muse-*`, or `<harness>-*`; the assisted creator enforces that convention.
12
+ - A default roster ships with the extension and seeds once on session start: five code-oriented roles (explorer, planner, implementer, reviewer, qa) plus the generalist worker, one storage profile per CLI backend across five backends (30 files, but six advertised roles). Seeding never overwrites existing files and respects later deletions. Roles outside the default roster are user-created via `/external profile create`. Every registered named Pi harness automatically and immediately gets the same six roles too, synthesized in-memory from the identical canonical role source — no per-harness file needed, and no on-disk seeding for pi harnesses. An installation seeded before Grok support existed migrates forward by adding only the six new `grok-*` profiles on its next seed pass; one seeded before Muse support existed adds only the six new `muse-*` profiles; either migration never resurrects a deleted or customized profile from an earlier backend.
13
13
  - A `backend: pi` profile is external-delegation-eligible only when it also declares `harness: <name>` for a name present in the live `harnesses.json` registry; a bare `backend: pi` profile (no `harness`, or one naming an unregistered config) is filtered out and rejected exactly as before — it belongs to Pi's native subagent system and this extension never modifies it.
14
14
  - **Named Pi harness configuration:** a `pi-<label>` entry in `harnesses.json` pinning a `provider/model` id (resolved through Pi's own model registry) and a thinking level (`off|minimal|low|medium|high|xhigh`, always persisted explicitly). It runs in-process via Pi's own SDK, not as a spawned CLI. A pi child loads no project/user extensions, skills, prompt templates, or themes — its entire tool surface is the SDK's own builtins (`read`/`bash`/`edit`/`write`, plus `grep`/`find`/`ls` where a tier's default allow-list adds them), curated but never claimed to be an OS sandbox: `danger` tier `bash` is exactly as exposed as on any external CLI. Retry is disabled per child (in-memory, call-scoped, never touching the user's real settings) to honor this extension's no-auto-retry contract, since the underlying SDK otherwise retries transient provider errors on its own. Pi children cannot resume (no persisted session) and have no enforced budget cap — both are deliberate v1 limitations, not oversights. Custom (non-canonical) roles are still authored per harness via `/external profile create`, not shared across harnesses automatically. Follow-up capability expansion (trusted extensions/MCP/skills, shared custom roles, resumable sessions, real budget controls) is tracked in issue #43.
15
15
  - Use Pi's native subagent system for Pi-backed scout/reviewer/planner/worker/oracle work.
16
16
  - Permission resolution takes the least restrictive of the profile tier (or default when absent) and the caller's request: `readonly < edit < danger`. Profile permissions are a floor, not an overridable default; a parent cannot handcuff a `danger` worker by requesting `readonly`. Existing backend-specific execution floors still apply. Disclosure and receipts show the resolved tier.
17
- - External CLI backends use their own tools and permission mechanisms. Codex tiers map to its `--sandbox` axis. Grok tiers also map to its `--sandbox` axis (`read-only`/`workspace`/`off`), always paired with `--permission-mode bypassPermissions`; bypass only skips the interactive prompt, and the kernel sandbox remains the enforced boundary at every tier. Grok's readonly network-blocking guarantee is Linux-only (a no-op on macOS), and sandbox startup can fail closed rather than silently downgrading on some macOS hosts (for example when `/var/run/docker.sock` resolves to a symlink). Claude falls back to `--permission-mode auto` when its effective UID is 0 because Claude refuses bypass mode under root. Execution lanes (implementer, qa, worker) require shell command authority to inspect repositories and run tests; on Claude and on pi harnesses (where `edit` tier excludes `bash`), execution lanes maintain a danger floor so agents are not artificially handcuffed. Grok needs no such floor: its `edit` tier (`--sandbox workspace`) already permits shell execution, sandboxed to writes within the workspace. Antigravity (`agy`) has no granular headless permission mode — its default sandbox denies even read-only tools — so every agy run is unsandboxed (`--dangerously-skip-permissions`) and `readonly`/`edit` on agy are advisory profile-body instructions, not a boundary. Run them only in trusted repositories and state whether the task is read-only or may edit files.
17
+ - External CLI backends use their own tools and permission mechanisms. Codex tiers map to its `--sandbox` axis. Grok tiers also map to its `--sandbox` axis (`read-only`/`workspace`/`off`), always paired with `--permission-mode bypassPermissions`; bypass only skips the interactive prompt, and the kernel sandbox remains the enforced boundary at every tier. Grok's readonly network-blocking guarantee is Linux-only (a no-op on macOS), and sandbox startup can fail closed rather than silently downgrading on some macOS hosts (for example when `/var/run/docker.sock` resolves to a symlink). Claude falls back to `--permission-mode auto` when its effective UID is 0 because Claude refuses bypass mode under root. Execution lanes (implementer, qa, worker) require shell command authority to inspect repositories and run tests; on Claude and on pi harnesses (where `edit` tier excludes `bash`), execution lanes maintain a danger floor so agents are not artificially handcuffed. Grok needs no such floor: its `edit` tier (`--sandbox workspace`) already permits shell execution, sandboxed to writes within the workspace. Antigravity (`agy`) has no granular headless permission mode — its default sandbox denies even read-only tools — so every agy run is unsandboxed (`--dangerously-skip-permissions`) and `readonly`/`edit` on agy are advisory profile-body instructions, not a boundary. Run them only in trusted repositories and state whether the task is read-only or may edit files. Muse (`muse exec`) has approval and its own sandbox ON by default; every tier passes `--disable-approval` so headless runs never hang on an interactive prompt. `readonly` additionally passes `--disable-write --disable-shell`; `edit` leaves the sandbox enabled with only approval bypassed, so shell/write stay available within it (no execution-lane danger floor is needed); `danger` uses `--yolo`, which disables approval and the sandbox and additionally trusts the workspace for this run (loads its skills/rules) — a broader grant than an unsandboxed run alone.
18
18
  - Grok supports `resume` via its own `--resume <sessionId>` flag and reports native cost (`total_cost_usd`) rather than an estimate. It has no budget-enforcement mechanism, so `max_budget_usd` is recorded but unenforceable, and — like codex — it is never automatically retried (agy's one infra-failure retry exception does not apply to Grok).
19
+ - Muse supports `resume` via its own `exec --session-id <uuid>` flag: verified against the real CLI (two independent `muse exec` processes sharing the same `--session-id` reported the same session, and the second recalled a fact only told to the first). Only the root run's own `run.terminal.completed`/`run.terminal.failed`/`run.output.delta` envelopes — identified by matching `payload.run_stream.id` against the id established from the process's own `runtime.command.accepted`/`session.run.linked` bootstrap pair — can finalize, fail, or contribute partial output to the result; a nested or foreign-run envelope with the same shape is ignored, and `sessionId` is captured once from that same bootstrap rather than overwritten by every subsequent envelope. It has never been observed to report token usage or cost on any run, so usage is reported as unknown (`costKnown: false`) rather than a fabricated zero or a locally estimated cost, and it has no budget-enforcement mechanism. It is never automatically retried by this extension; muse's own `meta` provider integration performs its own internal retries (observed up to 10 attempts, disclosed via activity narration such as "retrying meta model stream in 60000ms (attempt 3/10)") entirely inside the muse process, invisible to and unrelated to this extension's no-auto-retry contract. `muse exec` has no native system-prompt flag (unlike claude/codex/grok); the profile's `systemPrompt` is folded into the prompt file content instead, the same pattern agy uses. Nested-agent (sub-delegation) detection for Muse always reports false and never extends the nested-timeout deadline: every probe run saw only internal `reminder.agent.*` skill-reminder tasks (never a genuine delegation), and matching a speculative `task_kind` prefix would let ordinary internal task activity spuriously grant the one-time deadline extension, which is worse than never extending it. Revisit only once a real muse delegation event has actually been observed.
19
20
  - Children receive no parent history by default. `Agent` and workflow `agent()` can opt into `context: {mode: "recent", turns: N}` (last N user turns, including the current one) or `{mode: "full"}` (available post-compaction conversation). Snapshots exclude system instructions, thinking, tool-result metadata, and pending calls; unsupported content/images and more than 1 MiB fail explicitly. Use the smallest sufficient snapshot plus a clear task, absolute paths, and read-only/edit intent. Shared context goes to the external harness and private local evidence; avoid unnecessary sensitive history. `resume` continues an existing child and cannot be combined with sharing. Workflow children select from one frozen parent snapshot, and replay fingerprints include the transferred context.
20
21
  - Backend-native nested agents may use a different workspace. Include explicit absolute paths and required context when asking an external backend to delegate further.
21
22
  - `Agent` and `workflow` block by default; `background: true` returns a stable handle after validation/registration. Background work is owned by the originating session, not by the launching tool call or a daemon. Blocking-call interruption cancels work; wait interruption only stops waiting. Orderly session shutdown requests cancellation with bounded cleanup. A crash or unconfirmed shutdown leaves unfinished evidence interrupted/uncertain; restart never adopts live work.
@@ -55,7 +56,7 @@ Use workflows for requested fan-out or multi-agent orchestration across Claude/C
55
56
  - Test public contracts, trust/security boundaries, process lifecycle, cancellation/timeouts, profile rollback, workflow execution/resume, and receipt integrity.
56
57
  - Do not test cosmetic rendering variations, prompt prose fragments, trivial accessors, or a model's interpretation of instructions.
57
58
  - Do not repeat the same behavior at unit, integration, and E2E levels. A regression test must replace or extend overlapping coverage.
58
- - `npm test` must remain deterministic and offline. Real-provider E2E is opt-in and change-triggered. Its default lane never prompts a root LLM to pick the tool call — it builds an in-process Pi SDK session with a faux, never-streamed root model and calls the `Agent`/`workflow`/`external_runs` tool executors directly, so only the selected external backend's own child (a real spawned CLI process, or, for `pi`, a real in-process nested Pi child) is real: `npm run e2e -- --backend <claude|codex|agy|grok> [--workflow|--interrupt]`, or `npm run e2e -- --backend pi --harness <registered-name> [--workflow|--interrupt]` for a named Pi harness the caller has already registered in their own real `harnesses.json` with real credentials configured. A separate `--routing-smoke` lane spawns a real `pi` CLI process with a real root model and a plain-language instruction to exercise natural-language tool-call routing end to end; it requires real root Pi auth/config and is never the default.
59
+ - `npm test` must remain deterministic and offline. Real-provider E2E is opt-in and change-triggered. Its default lane never prompts a root LLM to pick the tool call — it builds an in-process Pi SDK session with a faux, never-streamed root model and calls the `Agent`/`workflow`/`external_runs` tool executors directly, so only the selected external backend's own child (a real spawned CLI process, or, for `pi`, a real in-process nested Pi child) is real: `npm run e2e -- --backend <claude|codex|agy|grok|muse> [--workflow|--interrupt]`, or `npm run e2e -- --backend pi --harness <registered-name> [--workflow|--interrupt]` for a named Pi harness the caller has already registered in their own real `harnesses.json` with real credentials configured. A separate `--routing-smoke` lane spawns a real `pi` CLI process with a real root model and a plain-language instruction to exercise natural-language tool-call routing end to end; it requires real root Pi auth/config and is never the default.
59
60
 
60
61
  ## Verification and release
61
62
 
package/CHANGELOG.md CHANGED
@@ -4,6 +4,30 @@ All notable changes to pi-flow external are documented here.
4
4
 
5
5
  ## Unreleased
6
6
 
7
+ ## [2.4.1-external.0] - 2026-09-23
8
+
9
+ ### Fixed
10
+
11
+ - Muse rejects effective `thinking: off` before process launch because its `meta` provider does not support `--reasoning-effort none`. The error explains supported profile/parent levels without silently changing reasoning (#60).
12
+ - Grok runtime-socket symlink sandbox failures include actionable guidance while retaining the original error, requested permissions, and fail-closed behavior. No socket changes, retries, or sandbox downgrade (#61).
13
+ - Added an offline test of the actual Pi `openai-completions` provider payload for optional tool selectors. Client optionality is preserved; the downstream model-facing conversion remains unresolved and tracked in #62. See `docs/tool-schema-compatibility.md`.
14
+
15
+ ## [2.4.0-external.0] - 2026-09-21
16
+
17
+ Adds `muse` as a fifth external CLI backend, delegating to Muse Code alongside Claude Code, Codex CLI, Antigravity, and Grok Build CLI. Investigation notes: `docs/plans/muse-backend-prep.md`; ex-ante design issue: none filed — additive backend addition following the same shape as the Grok backend.
18
+
19
+ ### Added
20
+
21
+ - `muse` backend (`src/core/muse.ts`), verified against Muse Code `1.3.0` running its `meta` provider, installed and authenticated independently of Pi. Runs `exec --json`, parsing schema-versioned MSP JSONL envelopes (`payload_type`/`payload`) — a different event shape from claude/codex/grok's flat events. Root-run ownership is established once from the process's own `runtime.command.accepted`/`session.run.linked` bootstrap pair (`command_id` tied to `run_stream.id`) and checked explicitly on every subsequent envelope: only a `run.terminal.completed`/`run.terminal.failed`/`run.output.delta` whose own `payload.run_stream.id` matches can finalize, fail, or contribute partial output to the result, so a nested or foreign-run envelope with an identical shape can never masquerade as the root's own answer, and `sessionId` is captured once from that bootstrap rather than from every envelope. Structured output uses `--output-schema <FILE>` (a temp file path, unlike Grok's inline `--json-schema`); the terminal event's `text` is already the pre-serialized JSON document. Muse has no native system-prompt flag, so a profile's `systemPrompt` is folded into the prompt file content, the same pattern Antigravity uses.
22
+ - Native approval/sandbox permission tiers: `readonly` (`--disable-approval --disable-write --disable-shell`), `edit` (`--disable-approval` only, sandbox left enabled), and `danger` (`--yolo`, which disables approval and the sandbox and additionally trusts the workspace for this run — a broader grant than an unsandboxed run alone, disclosed in full permission-help text). Approval and sandbox are ON by default, so every tier bypasses approval to avoid hanging headlessly. `edit` already permits shell execution within the sandbox, so Muse execution-lane roles need no `edit`→`danger` floor.
23
+ - `resume` support via Muse's own `exec --session-id <uuid>` flag, verified against the real CLI: two independent `muse exec` processes sharing the same session id shared context, confirmed by the second recalling a fact only told to the first (distinct from the separate, interactive-only `muse resume` command, which is not used here). Muse has never been observed to report token usage or cost on any run, so usage is reported as unknown (`costKnown: false`) rather than a fabricated zero or a locally estimated cost, and — like codex and grok — failed or aborted runs are never automatically retried by this extension; Muse's own `meta` provider integration performs its own internal retries (observed up to 10 attempts with growing backoff on transient errors) entirely inside the `muse` process, surfaced only as activity narration.
24
+ - Default roster migration: a pre-existing installation gains only the six new `muse-*` profiles on its next seed pass (30 default profiles total across five backends), without resurrecting any profile the user previously deleted or customized; a fresh installation seeds all 30 directly.
25
+ - Offline coverage in `test/muse-backend.test.ts` (including authoritative negative tests for a nested run's own terminal.completed/terminal.failed being ignored, and a terminal event whose run_stream never matches the established root failing closed) plus extended permission/defaults/profiles/resume/spawn-observation/run-inspection regression tests; opt-in real-provider coverage via `npm run e2e -- --backend muse` and `--backend muse --workflow`/`--interrupt`.
26
+
27
+ ### Known limitation
28
+
29
+ - Nested-agent (sub-delegation) detection for Muse always reports false and never grants the one-time nested-timeout extension: every real probe run observed only internal `reminder.agent.*` skill-reminder tasks, never a genuine agent delegation (the CLI reported "Agent delegation: auto unavailable: workspace is untrusted" in every probe), so there is no confirmed real event shape to key off — a speculative `task_kind` match was deliberately rejected as unsafe (it would let ordinary internal task activity spuriously extend the deadline) and removed after review.
30
+
7
31
  ## [2.3.1-external.0] - 2026-09-21
8
32
 
9
33
  ### Fixed
package/CONTEXT.md CHANGED
@@ -24,7 +24,7 @@ This package is a fork of pi-flow whose `Agent` and `workflow` tools are reserve
24
24
  ## Routing rule
25
25
 
26
26
  - Native Pi work -> native subagent tool.
27
- - Claude Code / Codex CLI / Antigravity / Grok Build CLI work, and named Pi harness work -> this extension's `Agent` or `workflow`.
27
+ - Claude Code / Codex CLI / Antigravity / Grok Build CLI / Muse Code work, and named Pi harness work -> this extension's `Agent` or `workflow`.
28
28
 
29
29
  The split is global and intentional to avoid tool ambiguity across projects. The ordinary driver sees `Agent`, optional `workflow`, and read-only `external_help`; the profile finalizer is activated only by `/external profile create`. External children start in the requested working directory, but backend-native nested helpers may create or use a separate workspace; prompts that request further nesting should include explicit absolute paths and all required context. Named Pi harness children are the one exception to "backend-native nested helpers may use a different workspace": a pi child runs in-process with no extensions, skills, prompts, or themes loaded, so it structurally cannot itself launch a further nested agent.
30
30
 
@@ -50,10 +50,12 @@ Applied to named Pi harness configurations, the same "get out of the way" stance
50
50
 
51
51
  Applied to the Grok Build CLI, permission tiers map onto its own kernel-enforced `--sandbox` axis (`read-only`/`workspace`/`off`), always paired with `--permission-mode bypassPermissions` — bypass only removes the interactive approval prompt, never the sandbox itself, so every tier is a genuine boundary rather than an advisory one. `edit` already permits shell execution (sandboxed to workspace writes), so, unlike Claude and pi, Grok execution-lane roles need no `edit`→`danger` floor. `readonly`'s network-blocking guarantee is Linux-only, and sandbox startup can fail closed on some macOS hosts rather than silently running unsandboxed; that fail-closed behavior is preserved rather than weakened for cross-platform convenience.
52
52
 
53
+ Applied to Muse Code, approval and its own sandbox are ON by default, so every tier passes `--disable-approval` (a headless run must never hang on an interactive prompt). `readonly` additionally strips non-shell writes and shell execution (`--disable-write --disable-shell`); `edit` leaves the sandbox enabled with only approval bypassed, so shell/write stay available within it — Muse needs no `edit`→`danger` floor either, for the same reason as Grok. `danger` uses `--yolo`, which disables approval and the sandbox *and additionally trusts the workspace for this run* (loads its skills/rules) — a materially broader grant than an unsandboxed run alone, so its disclosure names that extra trust rather than collapsing it into the same terse label every other danger tier uses. `resume` was verified against the real CLI (`exec --session-id`, not the separate interactive-only `resume` subcommand): two independent processes sharing one session id shared context, confirmed by the second recalling a fact only told to the first. Usage/cost has never been observed on any run, so it is reported unknown rather than a fabricated zero or estimate. Root-run ownership is checked explicitly: only envelopes whose `payload.run_stream.id` matches the id established from that same process's own `runtime.command.accepted`/`session.run.linked` bootstrap pair can finalize, fail, or contribute partial output to the result, so a nested or foreign-run terminal event can never masquerade as the root's own answer, and `sessionId` is captured once from that bootstrap rather than from every subsequent envelope. Nested-agent (sub-delegation) detection always reports false — every probe run only ever produced internal `reminder.agent.*` skill-reminder tasks, never an actual delegation, so there is no confirmed event shape to key off, and a speculative match would risk spuriously extending the nested-timeout deadline on ordinary internal task activity.
54
+
53
55
  ## Known inelegance
54
56
 
55
- <!-- ponytail: one-backend-per-file profile format forces role x backend file duplication; extend src/profiles.ts with multi-backend profiles (e.g. backends: [claude, codex, agy, grok] plus per-backend model map) if maintaining N copies of identical role bodies ever hurts -->
56
- A standardized role roster needs one file per role per backend (e.g. `claude-qa`, `codex-qa`, `agy-qa`, `grok-qa` with identical bodies) because the profile format binds `backend:` and `model:` to a single file. The duplication is accepted for now; a future format extension could let one role file cover all backends. Named Pi harnesses partially resolve this for the six canonical roles specifically (synthesized once from a single shared source, applied identically to every registered `pi-*` config, with zero files written) — but a *custom* role outside the six still needs one file per pi harness, exactly like the four CLI backends; a user with several registered pi harnesses wanting the same custom role on all of them still duplicates one file per harness. The same applies to `permission:` tiers: one tier per profile, but backends interpret tiers differently (Claude and pi deny Bash at anything below `danger`), so command-running roles declare `danger` and maintain `danger` as their floor even if an explicit `edit` override is requested. Custom profiles outside the role-name convention get that floor only by declaring `permission: danger`; an undeclared custom profile that meets an `edit` tier still hits headless Bash denials (or, on pi, a curated tool set without `bash`), so `permission: danger` doubles as the lane's "needs shell" capability declaration. Named Pi harnesses also ship with no resume and no enforced budget cap in v1 — both deliberate limitations, tracked alongside the shared-custom-role gap in issue #43, not oversights.
57
+ <!-- ponytail: one-backend-per-file profile format forces role x backend file duplication; extend src/profiles.ts with multi-backend profiles (e.g. backends: [claude, codex, agy, grok, muse] plus per-backend model map) if maintaining N copies of identical role bodies ever hurts -->
58
+ A standardized role roster needs one file per role per backend (e.g. `claude-qa`, `codex-qa`, `agy-qa`, `grok-qa`, `muse-qa` with identical bodies) because the profile format binds `backend:` and `model:` to a single file. The duplication is accepted for now; a future format extension could let one role file cover all backends. Named Pi harnesses partially resolve this for the six canonical roles specifically (synthesized once from a single shared source, applied identically to every registered `pi-*` config, with zero files written) — but a *custom* role outside the six still needs one file per pi harness, exactly like the five CLI backends; a user with several registered pi harnesses wanting the same custom role on all of them still duplicates one file per harness. The same applies to `permission:` tiers: one tier per profile, but backends interpret tiers differently (Claude and pi deny Bash at anything below `danger`), so command-running roles declare `danger` and maintain `danger` as their floor even if an explicit `edit` override is requested. Custom profiles outside the role-name convention get that floor only by declaring `permission: danger`; an undeclared custom profile that meets an `edit` tier still hits headless Bash denials (or, on pi, a curated tool set without `bash`), so `permission: danger` doubles as the lane's "needs shell" capability declaration. Named Pi harnesses also ship with no resume and no enforced budget cap in v1 — both deliberate limitations, tracked alongside the shared-custom-role gap in issue #43, not oversights.
57
59
 
58
60
  ## Evidence boundary
59
61
 
package/README.md CHANGED
@@ -6,6 +6,7 @@ External agent delegation for [pi](https://github.com/earendil-works/pi) through
6
6
  - [Codex CLI](https://github.com/openai/codex) profiles with `backend: codex`
7
7
  - Antigravity profiles with `backend: agy`
8
8
  - [Grok Build CLI](https://github.com/xai-org/grok-build) profiles with `backend: grok`
9
+ - Muse Code profiles with `backend: muse`
9
10
  - Named Pi harness configurations (`pi-<label>`) — in-process, per-model configs you register yourself, not a spawned CLI
10
11
 
11
12
  The ordinary driver has four tools (`workflow` can be disabled):
@@ -15,7 +16,7 @@ The ordinary driver has four tools (`workflow` can be disabled):
15
16
  - `external_help` returns role details, permission behavior, or workflow guidance on demand.
16
17
  - `external_runs` lists, inspects, waits for, and cancels session-owned runs.
17
18
 
18
- `Agent` and `workflow` accept external roles with an optional harness override — one of the four CLIs, or a registered named Pi harness. Use pi's native subagent system for Pi-backed agents. The `pi_flow_profile_create` and `pi_flow_harness_create` finalizers are active only during `/external profile create`.
19
+ `Agent` and `workflow` accept external roles with an optional harness override — one of the five CLIs, or a registered named Pi harness. Use pi's native subagent system for Pi-backed agents. The `pi_flow_profile_create` and `pi_flow_harness_create` finalizers are active only during `/external profile create`.
19
20
 
20
21
  ## Install
21
22
 
@@ -50,9 +51,10 @@ claude --version
50
51
  codex --version
51
52
  agy --version
52
53
  grok --version
54
+ muse --version
53
55
  ```
54
56
 
55
- Pi's coordinator model and the external CLIs authenticate independently. A working Claude, Codex, Antigravity, or Grok login does not authenticate the root Pi model. The Grok Build CLI installs and authenticates entirely separately from Pi: install with `curl -fsSL https://x.ai/cli/install.sh | bash`, then authenticate with `grok login` or an `XAI_API_KEY` environment variable. This integration is verified against Grok Build CLI `1.0.40`.
57
+ Pi's coordinator model and the external CLIs authenticate independently. A working Claude, Codex, Antigravity, Grok, or Muse login does not authenticate the root Pi model. The Grok Build CLI installs and authenticates entirely separately from Pi: install with `curl -fsSL https://x.ai/cli/install.sh | bash`, then authenticate with `grok login` or an `XAI_API_KEY` environment variable. This integration is verified against Grok Build CLI `1.0.40`. Muse Code likewise installs and authenticates separately from Pi; this integration is verified against Muse Code `1.3.0` against its `meta` provider.
56
58
 
57
59
  External agents use the effective permission tier and each harness's native mechanism:
58
60
 
@@ -60,9 +62,10 @@ External agents use the effective permission tier and each harness's native mech
60
62
  - Codex: `--sandbox read-only`, `workspace-write`, or `danger-full-access`
61
63
  - Antigravity: always `--dangerously-skip-permissions`
62
64
  - Grok: `--sandbox read-only`, `workspace`, or `off`, always alongside `--permission-mode bypassPermissions` — bypass only skips the interactive approval prompt; the kernel sandbox remains the enforced boundary. `readonly`'s network-blocking guarantee is Linux-only (a no-op on macOS), and sandbox startup can fail closed on some macOS hosts (for example when `/var/run/docker.sock` resolves to a symlink) rather than silently running unsandboxed.
65
+ - Muse: every tier passes `--disable-approval` (approval and Muse's own sandbox are ON by default, and headless runs must not hang on an interactive prompt); `readonly` additionally passes `--disable-write --disable-shell`; `edit` leaves the sandbox enabled with only approval bypassed; `danger` uses `--yolo`, which disables approval and the sandbox and additionally trusts the workspace for this run (loads its skills/rules) — a broader grant than an unsandboxed run alone.
63
66
  - Named Pi harnesses: a curated `tools:` allow-list (`read`/`grep`/`find`/`ls` at `readonly`; those plus `edit`/`write` at `edit`), or the SDK's own default active tools (`read`/`bash`/`edit`/`write`) at `danger`
64
67
 
65
- Claude refuses bypass mode when its effective UID is `0`; in that case the extension uses `--permission-mode auto`. External execution lanes (like `implementer`, `qa`, and `worker`) require shell execution to inspect repositories, run tests, and verify code. On Claude Code and named Pi harnesses, headless/`edit`-tier access excludes shell entirely; therefore, execution lanes maintain a `danger` floor so the model is not artificially handcuffed by permission blocks. Grok needs no such floor: its `edit` tier (`--sandbox workspace`) already permits shell execution, sandboxed to writes within the workspace. Run external agents only in repositories you trust and state whether each task is read-only or may edit files.
68
+ Claude refuses bypass mode when its effective UID is `0`; in that case the extension uses `--permission-mode auto`. External execution lanes (like `implementer`, `qa`, and `worker`) require shell execution to inspect repositories, run tests, and verify code. On Claude Code and named Pi harnesses, headless/`edit`-tier access excludes shell entirely; therefore, execution lanes maintain a `danger` floor so the model is not artificially handcuffed by permission blocks. Grok and Muse need no such floor: Grok's `edit` tier (`--sandbox workspace`) and Muse's `edit` tier (sandbox left enabled, only approval bypassed) already permit shell execution, sandboxed to the workspace. Run external agents only in repositories you trust and state whether each task is read-only or may edit files.
66
69
 
67
70
  The TUI labels a direct run with its effective access, including `unsandboxed external CLI` for danger and all Agy runs, `Pi SDK child · host access · curated tools` for a danger-tier named Pi harness, and shows `external host access` while workflow work is active. These labels disclose actual execution authority; they do not turn a read-only prompt into an enforced permission boundary.
68
71
 
@@ -116,7 +119,7 @@ Profiles live in:
116
119
  ~/.pi/agent/subagents/<name>.md
117
120
  ```
118
121
 
119
- Names may contain lowercase letters, numbers, and hyphens. A profile must declare `backend: claude`, `backend: codex`, `backend: agy`, `backend: grok`, or (alongside `harness: <a registered pi-* name>`) `backend: pi`, and its name should start with the matching harness name.
122
+ Names may contain lowercase letters, numbers, and hyphens. A profile must declare `backend: claude`, `backend: codex`, `backend: agy`, `backend: grok`, `backend: muse`, or (alongside `harness: <a registered pi-* name>`) `backend: pi`, and its name should start with the matching harness name.
120
123
 
121
124
  Example Claude profile, `~/.pi/agent/subagents/claude-explorer.md`:
122
125
 
@@ -148,11 +151,16 @@ backend: grok
148
151
  model: grok-4.6
149
152
  ```
150
153
 
154
+ ```yaml
155
+ backend: muse
156
+ model: muse-spark-1.3-contributor
157
+ ```
158
+
151
159
  Profile instructions become the external agent's system instructions. A profile's `description` is also shown as the user-visible reason for its selection, so keep it concise and concrete. External CLIs use their own tools, so a profile's `tools:` field does not control them; a pi profile's `tools:` field does apply, intersected with its permission tier's curated tool table. A `backend: pi` profile is only available to this extension when it also declares `harness: <name>` for a name registered in `harnesses.json` (see below); a bare `backend: pi` profile, or no backend at all, belongs to Pi's native subagent system and is never modified by this extension.
152
160
 
153
161
  ### Default profiles
154
162
 
155
- On first session start the extension seeds a default roster — five code-oriented roles (explorer, planner, implementer, reviewer, qa) plus the generalist worker — as one storage profile per backend across four backends (24 files). The compact agent catalog advertises each role once rather than presenting 24 choices. Seeding happens once: it never overwrites existing files, and profiles you delete or customize afterwards stay that way. Default profiles leave `model` and `thinking` unpinned so they track the CLI's own model and the current Pi thinking level. Roles not in the default roster can be added with `/external profile create`. An installation seeded before Grok support existed gains only the six new `grok-*` profiles the next time it seeds, without resurrecting any claude/codex/agy profile you previously deleted or customized.
163
+ On first session start the extension seeds a default roster — five code-oriented roles (explorer, planner, implementer, reviewer, qa) plus the generalist worker — as one storage profile per backend across five backends (30 files). The compact agent catalog advertises each role once rather than presenting 30 choices. Seeding happens once: it never overwrites existing files, and profiles you delete or customize afterwards stay that way. Default profiles leave `model` and `thinking` unpinned so they track the CLI's own model and the current Pi thinking level. Roles not in the default roster can be added with `/external profile create`. An installation seeded before Grok support existed gains only the six new `grok-*` profiles the next time it seeds; one seeded before Muse support existed gains only the six new `muse-*` profiles; neither migration resurrects a profile from an earlier backend that you previously deleted or customized.
156
164
 
157
165
  Project-local profiles are not supported; global profiles are used for both global and project-only package installations.
158
166
 
@@ -360,8 +368,8 @@ Every child selects from one parent snapshot and effective settings frozen at wo
360
368
  Every `Agent` call and workflow `agent()` child also accepts these optional run parameters (`context` is covered under agent usage above):
361
369
 
362
370
  - `permission`: `readonly` | `edit` | `danger` (default `danger`). Tiers map onto native harness mechanisms — Claude permission modes and Codex's single-axis `--sandbox`. Antigravity (`agy`) is different: its headless sandbox denies even read-only tools like `read_url_content`, and its only unsandboxed mode is `--dangerously-skip-permissions`, so **every agy run is unsandboxed** and `readonly`/`edit` on agy are advisory profile-body instructions, not a boundary. Getting out of the model's way is deliberate; every agy run discloses as `unsandboxed external CLI` rather than claiming a read-only boundary it cannot keep. Claude `readonly`/`edit` runs auto-deny shell commands headlessly; denials are surfaced in the receipt.
363
- - `max_budget_usd`: a spending cap. Claude Code enforces it mid-run with its native `--max-budget-usd` flag. Codex estimates cost from a price map and agy does not report cost at all; Grok reports its own native cost (`total_cost_usd`) but exposes no enforcement flag. For codex, agy, and grok alike, the cap is recorded and marked `budget unenforceable` instead of pretended.
364
- - `resume`: a prior run id. Continues the same backend conversation (Claude `--resume`, Codex `exec resume`, agy `--conversation`, Grok `--resume`) instead of starting from scratch. The prior run must use the same backend. Claude sessions persist in Claude Code's own local storage (this extension no longer passes `--no-session-persistence`) so recorded session ids stay resumable; remove old conversations from Claude Code itself if that matters to you. `resume` cannot be combined with `context` sharing — continue an existing child, or start a new one with a snapshot.
371
+ - `max_budget_usd`: a spending cap. Claude Code enforces it mid-run with its native `--max-budget-usd` flag. Codex estimates cost from a price map and agy does not report cost at all; Grok reports its own native cost (`total_cost_usd`) but exposes no enforcement flag; Muse has never been observed to report cost at all. For codex, agy, grok, and muse alike, the cap is recorded and marked `budget unenforceable` instead of pretended.
372
+ - `resume`: a prior run id. Continues the same backend conversation (Claude `--resume`, Codex `exec resume`, agy `--conversation`, Grok `--resume`, Muse `exec --session-id`) instead of starting from scratch. The prior run must use the same backend. Claude sessions persist in Claude Code's own local storage (this extension no longer passes `--no-session-persistence`) so recorded session ids stay resumable; remove old conversations from Claude Code itself if that matters to you. Muse's `--session-id` resume was verified directly: two independent `muse exec` processes sharing the same `--session-id` reported the same session, and the second recalled a fact only told to the first. `resume` cannot be combined with `context` sharing — continue an existing child, or start a new one with a snapshot.
365
373
 
366
374
  Resolution order for tiers and budgets: call > profile frontmatter (`permission:`, `max_budget_usd:`) > settings defaults.
367
375
 
@@ -468,10 +476,12 @@ npm run e2e -- --backend claude
468
476
  npm run e2e -- --backend codex
469
477
  npm run e2e -- --backend agy
470
478
  npm run e2e -- --backend grok
479
+ npm run e2e -- --backend muse
471
480
  npm run e2e -- --backend claude --workflow
472
481
  npm run e2e -- --backend codex --workflow
473
482
  npm run e2e -- --backend agy --workflow
474
483
  npm run e2e -- --backend grok --workflow
484
+ npm run e2e -- --backend muse --workflow
475
485
  npm run e2e -- --backend pi --harness pi-deepseek
476
486
  npm run e2e -- --backend pi --harness pi-deepseek --workflow
477
487
  ```
@@ -12,6 +12,7 @@ claude --version
12
12
  codex --version
13
13
  agy --version
14
14
  grok --version
15
+ muse --version
15
16
  ```
16
17
 
17
18
  The default lane needs no root Pi authentication at all (its root model is a faux, never-prompted placeholder). Only `--routing-smoke` needs a working root model:
@@ -33,6 +34,7 @@ npm run e2e -- --backend claude
33
34
  npm run e2e -- --backend codex
34
35
  npm run e2e -- --backend agy
35
36
  npm run e2e -- --backend grok
37
+ npm run e2e -- --backend muse
36
38
  ```
37
39
 
38
40
  Each run creates a collision-safe (`randomUUID`-named) temporary profile and fixture in an isolated agent directory, calls the `Agent` tool executor directly with an explicit `role`/`harness`, requires the exact expected result in the child's real result, checks for one complete `done` receipt, verifies that the fixture stayed clean, and removes its temporary files. Isolation is unconditional: for every backend above, an inherited `PI_CODING_AGENT_DIR` (such as the one exported in [Preparation](#preparation)) is ignored unless you pass `--agent-dir` explicitly — only `--backend pi` and `--routing-smoke` ever default to your real agent directory, since they only read it and never write into it.
@@ -45,10 +47,15 @@ Defaults:
45
47
  | Codex | `gpt-5.6-sol` | `high` |
46
48
  | Agy | `gemini-3.7-flash-high` | `high` |
47
49
  | Grok | `grok-4.6` | `high` |
50
+ | Muse | `muse-spark-1.3-contributor` | `high` |
48
51
 
49
52
  Override with `--model`/`--thinking`. Use `--keep` only when evidence inspection is necessary; it preserves sensitive output and the temporary profile path printed by the runner.
50
53
 
51
- Grok's `readonly`/`edit` tiers run under its own kernel sandbox (`--sandbox read-only`/`workspace`). On a macOS host where `/var/run/docker.sock` resolves to a symlink, sandbox startup can fail closed before the child even runs the prompt — that is expected fail-closed behavior, not a bug; verify enforcement on Linux, or adjust the host's Docker socket if you need to reproduce it locally on macOS.
54
+ Grok's `readonly`/`edit` tiers run under its own kernel sandbox (`--sandbox read-only`/`workspace`). Grok 1.0.40 on macOS can refuse startup with `sandbox could not be applied: socket deny resolution failed: could not resolve runtime-socket deny path /var/run/docker.sock: endpoint is a symlink`. This is an upstream CLI/environment incompatibility, not proof that the adapter's sandbox flags are wrong. The failed receipt retains the diagnostic and requested permission metadata; the adapter never retries with sandbox off. Do **not** remove or alter the Docker socket as a runner workaround. Verify sandbox enforcement on a compatible host or check a newer Grok version with an actual sandboxed invocation, not just `--help`. An explicitly requested `danger` run has no OS isolation and does not validate readonly delegation. Track compatibility and any confirmed upstream fix in [#61](https://github.com/tranhoangnguyen03/pi-flow-external/issues/61); no fixed upstream version has been verified.
55
+
56
+ Muse Code 1.3.0 accepts `none` in help text but its `meta` provider rejects `--reasoning-effort none`. Effective `thinking: off` (including the parent session's inherited default) therefore fails before launch with remediation; it is never silently mapped to minimal. Pin the Muse profile to `minimal`, `low`, `medium`, `high`, or `xhigh`, or select a supported parent thinking level. Removing a profile pin alone does not help when the parent still inherits `off`. Compatibility probes must invoke provider validation, not merely inspect `--help`.
57
+
58
+ Muse's `meta` provider performs its own internal retries (observed up to 10 attempts with growing backoff on transient 503/504 errors) entirely inside the `muse` process; this is unrelated to and invisible from this extension's own no-auto-retry contract, and shows up only as activity narration (e.g. "retrying meta model stream in 60000ms (attempt 3/10)"). A run that never reports usage/cost is expected — Muse has never been observed to report either.
52
59
 
53
60
  ### Named Pi harness receipt
54
61
 
@@ -69,6 +76,7 @@ npm run e2e -- --backend claude --workflow
69
76
  npm run e2e -- --backend codex --workflow
70
77
  npm run e2e -- --backend agy --workflow
71
78
  npm run e2e -- --backend grok --workflow
79
+ npm run e2e -- --backend muse --workflow
72
80
  npm run e2e -- --backend pi --harness pi-deepseek --workflow
73
81
  ```
74
82
 
@@ -83,6 +91,7 @@ npm run e2e -- --backend claude --interrupt
83
91
  npm run e2e -- --backend codex --interrupt
84
92
  npm run e2e -- --backend agy --interrupt
85
93
  npm run e2e -- --backend grok --interrupt
94
+ npm run e2e -- --backend muse --interrupt
86
95
  ```
87
96
 
88
97
  The runner calls the `Agent` tool executor directly with `background: true`, polls `external_runs` until the run is observed `running`, cancels its stable ID with an explicit reason, waits for its `cancelled` outcome, and asks `external_runs` for every available output and diagnostic page. A very early cancellation may legitimately have diagnostics but no assistant text; the durable receipt must still be `aborted` with the exact reason. Do not retry a failure automatically. Use `--keep` for one deliberate evidence inspection, then remove the printed run root and temporary profile.
@@ -91,7 +100,7 @@ The runner calls the `Agent` tool executor directly with `background: true`, pol
91
100
 
92
101
  There are two levels of routing check:
93
102
 
94
- - `npm run e2e -- --routing-smoke --backend <claude|codex|agy|grok>` (optionally `--workflow`/`--interrupt`) spawns a real `pi` CLI process with a real root model and gives it a single plain-language instruction naming the exact tool, role, and harness to call. It is a mechanical, automated check that the natural-language call-and-relay path (real root LLM → real tool call → real backend child) still works end to end. It requires real root Pi authentication (see Preparation) and is opt-in only — never the default backend gate — because the root model's tool-call decision is a live, provider-billed, non-deterministic step. `--root-model`/`--root-thinking` only apply here.
103
+ - `npm run e2e -- --routing-smoke --backend <claude|codex|agy|grok|muse>` (optionally `--workflow`/`--interrupt`) spawns a real `pi` CLI process with a real root model and gives it a single plain-language instruction naming the exact tool, role, and harness to call. It is a mechanical, automated check that the natural-language call-and-relay path (real root LLM → real tool call → real backend child) still works end to end. It requires real root Pi authentication (see Preparation) and is opt-in only — never the default backend gate — because the root model's tool-call decision is a live, provider-billed, non-deterministic step. `--root-model`/`--root-thinking` only apply here.
95
104
  - The broader, still-manual qualitative check below covers *discovery*, not just one named call: when role discovery, tool descriptions, or coordinator guidance changes, run three fresh Pi sessions against the current checkout and a read-only fixture. Ask for two named harnesses to review, two named harnesses to research, and a task split between Pi plus two named harnesses. For each session, verify that each requested external harness produced one complete `done` receipt through `role` plus `harness`, no help/discovery call was needed for the built-in roles, and the fixture stayed unchanged.
96
105
 
97
106
  Neither is a statistical regression comparison or a model-independent guarantee. Neither covers workflow routing or the omitted-`harness` default; use the workflow receipt above and a separate direct role request without `harness` for those paths.
@@ -141,8 +150,8 @@ Run only when project default-harness resolution changes. In a fresh Pi session
141
150
  Before a runtime release:
142
151
 
143
152
  1. Run `npm run check`.
144
- 2. Run all four direct backend receipts.
145
- 3. Run all four supervised workflow receipts.
153
+ 2. Run all five direct backend receipts.
154
+ 3. Run all five supervised workflow receipts.
146
155
  4. Run interruption checks for adapters whose cancellation/output path changed.
147
156
  5. Run `--routing-smoke` (mechanical) and the manual qualitative routing smoke only if role discovery, tool descriptions, or coordinator guidance changed.
148
157
  6. Run the nested timeout check only if nested detection or timeout behavior changed.
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "@tranhoangnguyen0310/pi-flow-external",
3
- "version": "2.3.1-external.0",
4
- "description": "External Claude Code, Codex CLI, Antigravity, and Grok Build CLI delegation for pi.",
3
+ "version": "2.4.1-external.0",
4
+ "description": "External Claude Code, Codex CLI, Antigravity, Grok Build CLI, and Muse Code delegation for pi.",
5
5
  "type": "module",
6
6
  "main": "./index.ts",
7
7
  "exports": "./index.ts",
@@ -29,7 +29,8 @@
29
29
  "claude",
30
30
  "agy",
31
31
  "antigravity",
32
- "grok"
32
+ "grok",
33
+ "muse"
33
34
  ],
34
35
  "license": "MIT",
35
36
  "scripts": {
@@ -57,7 +57,7 @@ async function loadRecord(baseDirectory, runId) {
57
57
  const eventCount = finite(document.eventCount);
58
58
  const attemptedEventCount = finite(document.attemptedEventCount);
59
59
  const backendEventCount = events.filter((event) => event.type === "backend_event").length;
60
- if (!["claude", "codex", "agy", "grok"].includes(summary.backend)) problems.push("missing or unknown backend");
60
+ if (!["claude", "codex", "agy", "grok", "muse"].includes(summary.backend)) problems.push("missing or unknown backend");
61
61
  if (typeof summary.profile !== "string" || !summary.profile) problems.push("missing profile");
62
62
  if (!["done", "error", "aborted"].includes(summary.status)) problems.push("missing or unknown status");
63
63
  if (writeErrorCount === undefined) problems.push("missing write error count");
@@ -16,5 +16,8 @@ export function getBackendAgentLabel(backend: SubagentBackend | undefined): stri
16
16
  if (backend === "grok") {
17
17
  return "Grok CLI";
18
18
  }
19
+ if (backend === "muse") {
20
+ return "Muse Code";
21
+ }
19
22
  return "Agent";
20
23
  }
package/src/core/grok.ts CHANGED
@@ -274,6 +274,29 @@ export function extractGrokError(event: Record<string, unknown>): string | undef
274
274
  return undefined;
275
275
  }
276
276
 
277
+ /**
278
+ * Produces an actionable diagnosis when Grok's kernel sandbox refuses to start
279
+ * because a runtime socket deny path is a symlink.
280
+ *
281
+ * Preserves the exact stderr and explains:
282
+ * - Upstream Grok sandbox initialization failed because a runtime socket deny path is a symlink.
283
+ * - The socket must NOT be removed or altered as a runner workaround.
284
+ * - pi-flow-external preserves requested sandbox permissions and will never automatically downgrade
285
+ * or retry with protections disabled.
286
+ * - Recommends running on a compatible host or upgrading to an upstream Grok release.
287
+ */
288
+ export function diagnoseGrokSandboxError(stderr: string): string | undefined {
289
+ if (!stderr.includes("runtime-socket deny path") || !stderr.includes("endpoint is a symlink")) {
290
+ return undefined;
291
+ }
292
+ return (
293
+ "Grok sandbox initialization failed because a runtime socket deny path is a symlink. " +
294
+ "Do not remove or alter the socket as a runner workaround. " +
295
+ "pi-flow-external preserves requested sandbox permissions and will not automatically downgrade or retry with protections disabled. " +
296
+ "Run on a compatible host or upgrade to an upstream Grok release that resolves socket symlinks."
297
+ );
298
+ }
299
+
277
300
  function getPreviewFromRecord(record: Record<string, unknown>): string {
278
301
  const candidates = [
279
302
  record.command,
@@ -560,8 +583,10 @@ export async function spawnGrokSubagent(params: {
560
583
  }
561
584
  if (closeResult.code !== 0) {
562
585
  const stderr = stderrBuffer.text().trim();
586
+ const diagnostic = diagnoseGrokSandboxError(stderr);
587
+ const diagnosticSuffix = diagnostic ? `\n\nDiagnostic: ${diagnostic}` : "";
563
588
  throw new Error(
564
- `grok exited with code ${closeResult.code}${closeResult.signal ? ` (signal ${closeResult.signal})` : ""}${stderr ? `: ${stderr}` : ""}`,
589
+ `grok exited with code ${closeResult.code}${closeResult.signal ? ` (signal ${closeResult.signal})` : ""}${stderr ? `: ${stderr}` : ""}${diagnosticSuffix}`,
565
590
  );
566
591
  }
567
592