@tranhoangnguyen0310/pi-flow-external 2.4.2-external.0 → 2.6.0-external.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/AGENTS.md CHANGED
@@ -4,17 +4,17 @@
4
4
 
5
5
  This fork changes the original pi-flow contract: `Agent` is not a generic Pi subagent launcher. It delegates only to external Claude Code, Codex CLI, Antigravity, Grok Build CLI, and Muse Code harnesses, plus named Pi harness configurations (below) — in-process, per-model configs registered by the user, not a spawned CLI.
6
6
 
7
- - Ordinary driver tools: `Agent`, read-only `external_help`, `external_runs`, and optional `workflow`. `pi_flow_profile_create` and `pi_flow_harness_create` are active only inside `/external profile create`.
8
- - User operations use `/external`; `/pi-flow-profile create` is a temporary deprecated alias.
9
- - Extension-owned concurrency, timeout, and global default-harness settings live in `$PI_CODING_AGENT_DIR/pi-flow-external/settings.json`. Profile backend/model/thinking metadata remains in `subagents/*.md`. Named Pi harness configurations live in the sibling `$PI_CODING_AGENT_DIR/pi-flow-external/harnesses.json` registry (one shared file, `pi-<label>` keys pinning `model`/`thinking`). Registry writes are atomic per file, but there is no inter-process lock: concurrent writers adding distinct names from separate sessions are last-writer-wins and may drop one update. This is an accepted v1 limitation shared with `settings.json` and `subagents/*.md`, not a bug to re-report. A trusted project may override `defaultHarness` via `.pi/pi-flow-external/settings.json` (that key only, read-only to the extension); explicit call harness > project default > global default. `defaultHarness` may itself name a registered `pi-*` harness.
10
- - Every `Agent` call requires `description`, `prompt`, and either `role` with optional `harness` (`agy`, `claude`, `codex`, `grok`, `muse`, or a registered `pi-*` harness name) or the legacy exact-profile `subagent_type`. The selectors cannot be combined. Without `harness`, use the effective default harness; if that default names a `pi-*` harness that is no longer registered, the call fails with an actionable error rather than silently falling back to `agy`.
11
- - Valid profiles come from `~/.pi/agent/subagents/*.md` and set `backend: claude`, `backend: codex`, `backend: agy`, `backend: grok`, `backend: muse`, or (only when also declaring `harness: <registered pi-* name>`) `backend: pi`. Use matching names such as `claude-*`, `codex-*`, `agy-*`, `grok-*`, `muse-*`, or `<harness>-*`; the assisted creator enforces that convention.
12
- - A default roster ships with the extension and seeds once on session start: five code-oriented roles (explorer, planner, implementer, reviewer, qa) plus the generalist worker, one storage profile per CLI backend across five backends (30 files, but six advertised roles). Seeding never overwrites existing files and respects later deletions. Roles outside the default roster are user-created via `/external profile create`. Every registered named Pi harness automatically and immediately gets the same six roles too, synthesized in-memory from the identical canonical role source — no per-harness file needed, and no on-disk seeding for pi harnesses. An installation seeded before Grok support existed migrates forward by adding only the six new `grok-*` profiles on its next seed pass; one seeded before Muse support existed adds only the six new `muse-*` profiles; either migration never resurrects a deleted or customized profile from an earlier backend.
13
- - A `backend: pi` profile is external-delegation-eligible only when it also declares `harness: <name>` for a name present in the live `harnesses.json` registry; a bare `backend: pi` profile (no `harness`, or one naming an unregistered config) is filtered out and rejected exactly as before — it belongs to Pi's native subagent system and this extension never modifies it.
14
- - **Named Pi harness configuration:** a `pi-<label>` entry in `harnesses.json` pinning a `provider/model` id (resolved through Pi's own model registry) and a thinking level (`off|minimal|low|medium|high|xhigh`, always persisted explicitly). It runs in-process via Pi's own SDK, not as a spawned CLI. A pi child loads no project/user extensions, skills, prompt templates, or themes — its entire tool surface is the SDK's own builtins (`read`/`bash`/`edit`/`write`, plus `grep`/`find`/`ls` where a tier's default allow-list adds them), curated but never claimed to be an OS sandbox: `danger` tier `bash` is exactly as exposed as on any external CLI. Retry is disabled per child (in-memory, call-scoped, never touching the user's real settings) to honor this extension's no-auto-retry contract, since the underlying SDK otherwise retries transient provider errors on its own. Pi children cannot resume (no persisted session) and have no enforced budget cap — both are deliberate v1 limitations, not oversights. Custom (non-canonical) roles are still authored per harness via `/external profile create`, not shared across harnesses automatically. Follow-up capability expansion (trusted extensions/MCP/skills, shared custom roles, resumable sessions, real budget controls) is tracked in issue #43.
7
+ - Ordinary driver tools: `Agent`, read-only `external_help`, `external_runs`, and optional `workflow`. `pi_flow_role_create` is active only inside `/external role create`. `pi_flow_harness_create` is active only inside `/external harness create`.
8
+ - User operations use `/external`. `/external profiles`, `/external profile create`, and `/pi-flow-profile create` are removed, including registrations, completions, and interview routing. They are not aliases or redirect handlers. Unknown-command help lists the current commands. `subagent_type` on `Agent` and `workflow` remains the exact-identity API.
9
+ - The only execution-configuration file is `$PI_CODING_AGENT_DIR/pi-flow-external/settings.json` (settings v4): concurrency, timeout, permission and budget defaults, retention, `defaultHarness`, named `pi-*` harness entries (`model` / `thinking` / `preset`), and `disabledProfiles`. Reads do not write the file. Shared roles live in `pi-flow-external/roles/<role>.md` (one file, any harness; `description` only). Exact overrides live in `pi-flow-external/overrides/<harness>-<role>.md` and keep the execution profile schema as a complete replacement. `permission` and `capabilitySet` are obsolete metadata on both and are rejected, not used as authority. Both directories are created only when a file is saved. The extension does not write Pi's native `subagents/` directory. Settings mutations use one validated read-modify-write helper (private staged file, same-directory replacement, unrelated fields preserved, malformed or future-version files refused). Replacement prevents a torn write. There is no inter-process lock: concurrent writers from separate sessions are last-writer-wins and may drop one update. That limitation stays accepted. A trusted project may override `defaultHarness` via `.pi/pi-flow-external/settings.json` (that key only, read-only to the extension); explicit call harness > project default > global default > built-in default. `defaultHarness` may name a registered `pi-*` harness. The project file cannot inject roles or harnesses.
10
+ - Every `Agent` call requires `description`, `prompt`, and either `role` with optional `harness` (`agy`, `claude`, `codex`, `grok`, `muse`, or a registered `pi-*` harness name) or the legacy exact `subagent_type`. The selectors cannot be combined. Without `harness`, use the effective default harness. A default that names a `pi-*` harness that is no longer registered fails with an actionable error and does not fall back to `agy`. `disabledProfiles` blocks both `role` selection and exact `subagent_type` selection. There is no harness substitution.
11
+ - One catalog snapshot serves `Agent`, workflows, coordinator guidance, `external_help`, `/external roles`, and doctor. Precedence is exact override, then shared role, then built-in. An invalid higher-priority entry blocks that selection. Unrelated invalid entries are diagnostics and do not block other roles. Invalid settings JSON or an unsupported settings version blocks delegation; `/external settings` and `/external settings convert` stay available. Stable identities are `<harness>-<role>` plus exact legacy names. The next invocation reads settings, roles, and overrides from disk. A workflow freezes one snapshot for the whole run. A `maxConcurrentSubagents` change waits until no subagent is active and none are queued.
12
+ - The six built-in roles (explorer, planner, implementer, reviewer, qa, worker) are in memory for all five CLI harnesses and every registered named Pi harness. Session start writes zero role files and zero seed markers. A shared file named `reviewer.md` replaces that built-in on every harness. Roles outside the six are authored once with `/external role create` (offline validation and write; a role is not an authenticated connection). `/external role override <role> <harness>` materializes one full override. `/external harness create` registers a named Pi harness with a real runtime smoke test and no role interview. A harness smoke test does not judge role quality. Deleted seeded identities survive conversion as `disabledProfiles` and are not recreated.
13
+ - Native Pi subagent files stay native. This extension does not modify them. A `backend: pi` file under native `subagents/` is outside the external catalog. A Pi exact override is delegation-eligible only when its `harness` names an entry in the live settings `harnesses` map.
14
+ - **Named Pi harness configuration:** a `pi-<label>` entry in settings v4 `harnesses`, pinning a `provider/model` id (resolved through Pi's own model registry), a thinking level (`off|minimal|low|medium|high|xhigh`, always persisted explicitly), and a resource preset (`minimal` or `skills`). New writes persist the preset. A legacy registration that omits it is `minimal`. The five CLI harnesses are built in and need no entries. It runs in-process via Pi's own SDK, not as a spawned CLI. `minimal` leaves skills unloaded (`noSkills: true`). `skills` loads installed skills through `DefaultResourceLoader` (`noSkills: false`); project-scope skills load only when the caller's project is trusted. Extensions, prompt templates, and themes stay unloaded in both presets. The catalog, the workflow's frozen descriptor, the replay fingerprint, and spawn all use that registration value. The tool surface stays the SDK builtins (`read`/`bash`/`edit`/`write`, plus `grep`/`find`/`ls` where a tier's allow-list adds them), curated but never claimed to be an OS sandbox: `danger` tier `bash` is exactly as exposed as on any external CLI. Retry is disabled per child (in-memory, call-scoped, never touching the user's real settings) to honor this extension's no-auto-retry contract, since the underlying SDK otherwise retries transient provider errors on its own. Pi children cannot resume (no persisted session) and have no enforced budget cap — both are deliberate v1 limitations, not oversights. Shared roles apply to every harness from one file. Backend-specific model, tools, budget, or instruction changes belong in one exact override; `tools` is enforceable only on Pi, and a Pi override's `model`/`thinking` must match the harness entry. Follow-up capability expansion (trusted extensions/MCP, resumable sessions, real budget controls) stays tracked in issue #43. `pi-web-access` was not wired: the pinned SDK docs do not verify that extension integration, and it is left for a separate investigation. Pre-v4 installs convert once through `/external settings convert`. `/external settings` points there when conversion is ready. Originals stay on disk. Custom profiles are copied as overrides with obsolete `permission` and `capabilitySet` removed. A legacy `subagents/pi-<role>.md` file with `harness: "pi-*"` becomes an ordinary cross-harness `roles/<role>.md`. Those copies install before v4 activation, and after activation old paths are ignored. `piCapabilitySets` is not copied. `/external [danger]purge-old-files` lists the legacy inventory, including customized copies, and deletes only the paths you select and then confirm. Ordinary delegation does not require it. Downgrade after purge needs the user's own backup. The literal `[danger]` is part of the purge spelling. There is no dual-write and no `/external migrate` command.
15
15
  - Use Pi's native subagent system for Pi-backed scout/reviewer/planner/worker/oracle work.
16
- - Permission resolution takes the least restrictive of the profile tier (or default when absent) and the caller's request: `readonly < edit < danger`. Profile permissions are a floor, not an overridable default; a parent cannot handcuff a `danger` worker by requesting `readonly`. Existing backend-specific execution floors still apply. Disclosure and receipts show the resolved tier.
17
- - External CLI backends use their own tools and permission mechanisms. Codex tiers map to its `--sandbox` axis. Grok tiers also map to its `--sandbox` axis (`read-only`/`workspace`/`off`), always paired with `--permission-mode bypassPermissions`; bypass only skips the interactive prompt, and the kernel sandbox remains the enforced boundary at every tier. Grok's readonly network-blocking guarantee is Linux-only (a no-op on macOS), and sandbox startup can fail closed rather than silently downgrading on some macOS hosts (for example when `/var/run/docker.sock` resolves to a symlink). Claude falls back to `--permission-mode auto` when its effective UID is 0 because Claude refuses bypass mode under root. Execution lanes (implementer, qa, worker) require shell command authority to inspect repositories and run tests; on Claude and on pi harnesses (where `edit` tier excludes `bash`), execution lanes maintain a danger floor so agents are not artificially handcuffed. Grok needs no such floor: its `edit` tier (`--sandbox workspace`) already permits shell execution, sandboxed to writes within the workspace. Antigravity (`agy`) has no granular headless permission mode — its default sandbox denies even read-only tools — so every agy run is unsandboxed (`--dangerously-skip-permissions`) and `readonly`/`edit` on agy are advisory profile-body instructions, not a boundary. Run them only in trusted repositories and state whether the task is read-only or may edit files. Muse (`muse exec`) has approval and its own sandbox ON by default; every tier passes `--disable-approval` so headless runs never hang on an interactive prompt. `readonly` additionally passes `--disable-write --disable-shell`; `edit` leaves the sandbox enabled with only approval bypassed, so shell/write stay available within it (no execution-lane danger floor is needed); `danger` uses `--yolo`, which disables approval and the sandbox and additionally trusts the workspace for this run (loads its skills/rules) — a broader grant than an unsandboxed run alone.
16
+ - Permission is the call's explicit `permission`, otherwise settings `defaultPermission` (`danger` unless changed). A role describes intent and does not grant or limit authority. There is no profile floor and no role-name escalation. Each backend maps a tier it supports and rejects a restriction it cannot enforce. Disclosure and receipts show that resolved tier.
17
+ - External CLI backends use their own tools and permission mechanisms. Codex tiers map to its `--sandbox` axis. Grok tiers also map to its `--sandbox` axis (`read-only`/`workspace`/`off`), always paired with `--permission-mode bypassPermissions`; bypass only skips the interactive prompt, and the kernel sandbox remains the enforced boundary at every tier. Grok's readonly network-blocking guarantee is Linux-only (a no-op on macOS), and sandbox startup can fail closed rather than silently downgrading on some macOS hosts (for example when `/var/run/docker.sock` resolves to a symlink). Claude falls back to `--permission-mode auto` when its effective UID is 0 because Claude refuses bypass mode under root. Claude `edit` uses `acceptEdits` and denies Bash headlessly; a role name does not raise that tier. Antigravity (`agy`) accepts only unsandboxed `--dangerously-skip-permissions`. A `readonly` or `edit` request is rejected rather than broadened. Run agy only in trusted repositories. Muse (`muse exec`) has approval and its own sandbox ON by default; every supported tier passes `--disable-approval` so headless runs never hang on an interactive prompt. `readonly` additionally passes `--disable-write --disable-shell`; `edit` leaves the sandbox enabled with only approval bypassed; `danger` uses `--yolo`, which disables approval and the sandbox and additionally trusts the workspace for this run (loads its skills/rules). Pi curated tool lists are not an OS sandbox.
18
18
  - Grok supports `resume` via its own `--resume <sessionId>` flag and reports native cost (`total_cost_usd`) rather than an estimate. It has no budget-enforcement mechanism, so `max_budget_usd` is recorded but unenforceable, and — like codex — it is never automatically retried (agy's one infra-failure retry exception does not apply to Grok).
19
19
  - Muse supports `resume` via its own `exec --session-id <uuid>` flag: verified against the real CLI (two independent `muse exec` processes sharing the same `--session-id` reported the same session, and the second recalled a fact only told to the first). Only the root run's own `run.terminal.completed`/`run.terminal.failed`/`run.output.delta` envelopes — identified by matching `payload.run_stream.id` against the id established from the process's own `runtime.command.accepted`/`session.run.linked` bootstrap pair — can finalize, fail, or contribute partial output to the result; a nested or foreign-run envelope with the same shape is ignored, and `sessionId` is captured once from that same bootstrap rather than overwritten by every subsequent envelope. It has never been observed to report token usage or cost on any run, so usage is reported as unknown (`costKnown: false`) rather than a fabricated zero or a locally estimated cost, and it has no budget-enforcement mechanism. It is never automatically retried by this extension; muse's own `meta` provider integration performs its own internal retries (observed up to 10 attempts, disclosed via activity narration such as "retrying meta model stream in 60000ms (attempt 3/10)") entirely inside the muse process, invisible to and unrelated to this extension's no-auto-retry contract. `muse exec` has no native system-prompt flag (unlike claude/codex/grok); the profile's `systemPrompt` is folded into the prompt file content instead, the same pattern agy uses. Nested-agent (sub-delegation) detection for Muse always reports false and never extends the nested-timeout deadline: every probe run saw only internal `reminder.agent.*` skill-reminder tasks (never a genuine delegation), and matching a speculative `task_kind` prefix would let ordinary internal task activity spuriously grant the one-time deadline extension, which is worse than never extending it. Revisit only once a real muse delegation event has actually been observed.
20
20
  - Children receive no parent history by default. `Agent` and workflow `agent()` can opt into `context: {mode: "recent", turns: N}` (last N user turns, including the current one) or `{mode: "full"}` (available post-compaction conversation). Snapshots exclude system instructions, thinking, tool-result metadata, and pending calls; unsupported content/images and more than 1 MiB fail explicitly. Use the smallest sufficient snapshot plus a clear task, absolute paths, and read-only/edit intent. Shared context goes to the external harness and private local evidence; avoid unnecessary sensitive history. `resume` continues an existing child and cannot be combined with sharing. Workflow children select from one frozen parent snapshot, and replay fingerprints include the transferred context.
@@ -23,8 +23,8 @@ This fork changes the original pi-flow contract: `Agent` is not a generic Pi sub
23
23
 
24
24
  ## Delegation transparency invariants
25
25
 
26
- - Treat each `description` as a concise user-facing task label. Profile descriptions are also user-visible as the declared reason for profile selection.
27
- - `unsandboxed external CLI` and `external host access` disclose the real execution boundary for CLI backends; a pi harness delegation discloses `Pi SDK child · host access · curated tools` instead — never call an in-process pi child an "external CLI". Never present a read-only prompt as permission enforcement: on agy every run is unsandboxed regardless of tier, so the profile body — not the tier — is what asks the agent to stay read-only; on pi, the curated tool table bounds which tool *names* exist, but does not make `bash` at `danger` tier any less exposed.
26
+ - Treat each `description` as a concise user-facing task label. Role and override descriptions are also user-visible as the declared reason for selection. The resolved execution object is still an internal profile; receipts keep the stable `<harness>-<role>` identity or the exact `subagent_type` name.
27
+ - `unsandboxed external CLI` and `external host access` disclose the real execution boundary for CLI backends; a pi harness delegation discloses `Pi SDK child · host access · curated tools` instead — never call an in-process pi child an "external CLI". Never present a read-only prompt as permission enforcement. Antigravity accepts only danger and rejects narrower tiers. On pi, the curated tool table bounds which tool *names* exist and is not an OS sandbox; `bash` at `danger` stays host access.
28
28
  - Keep direct intent visible during execution. Workflow access belongs once at the workflow level, not on every child row.
29
29
  - Parent-context sharing is a disclosure, not a silent optimization: the intent card names the mode before launch, and receipts name the mode and shared/requested turns. Shared conversation content leaves for the external harness and lands in local evidence, so never describe sharing as internal or free.
30
30
  - Keep default progress bounded and human-readable. Expanded terminal output shows bounded canonical output and a full-ID `/external runs` route; record paths, backend-event counts, workflow IDs, and journal paths remain advanced evidence.
@@ -40,27 +40,27 @@ This fork changes the original pi-flow contract: `Agent` is not a generic Pi sub
40
40
  - Structured nested-agent activity may extend the wall-clock deadline once, by one fresh base timeout, capped at twice the original deadline.
41
41
  - Do not automatically retry failed or aborted external runs. Preserve the receipt and retry only when the user asks. Exception: the agy backend retries once on infrastructure-classified failures (auth, eligibility, network); the retry is disclosed in the receipt details (`retries`, `retryOf`) and never applies to agent-level failures, aborts, or timeouts.
42
42
  - `external_runs` is session/project scoped. `list` pages runs and workflow roots, each carrying a shared timing projection (`queuedAt`/`executionStartedAt`/`processStartedAt`/activity timestamps plus derived `queueDelayMs`/`elapsedMs`/`activityAgeMs`/`processDurationMs`); `activityAgeMs` and a running `elapsedMs` are computed only for a target independently confirmed live by the registry, never inferred from a durable record's status field alone (an unowned "running"-looking record may be an orphaned/crashed run, not a live one). `inspect` pages `summary`/`output`/`diagnostics`/`final` for one `runId`; `final` returns only the same canonical terminal result every other view already reads (no separate backend parser, no narration promoted to final) and is empty with `finalAvailable: false` until a verified successful terminal boundary exists — including for workflows, where an intentional JSON `null` result still counts as available. `inspect` also accepts a bounded batch `runIds` (up to 20, deduplicated, order preserved, summary-only, mutually exclusive with `runId`): every requested ID's ownership is validated before any page returns, entries share one consistent nested shape regardless of live/durable/agent/workflow origin (a workflow entry omits its unbounded children array in favor of `outputRef`/`diagnosticsRef`), reuses the existing per-inspect `limitBytes` cap with no second budget, and fails with an actionable error naming the run rather than ever silently dropping a target or serving a page over the caller's own byte budget. `wait` returns outcomes only for the selected one/any/all target set and never cancels pending work; while waiting it reports bounded live progress (watched targets, completed/pending counts, recent activity) through the tool's own update channel on a fixed heartbeat, stopped on settlement/error/interruption, that never mutates run state or affects cancellation; each settled outcome's `result` is spent from one shared byte budget (`limitBytes`, default 32768) across the whole response in the caller's requested order, never settlement race order — complete when it fits, `resultTruncated: true` plus the existing `outputRef`/`diagnosticsRef` when it does not, never a fixed-length teaser regardless of size; a target's evidence is never read from disk once that shared budget is already exhausted. `cancel` targets one stable ID and preserves its reason. An unresolvable `wf_...` ID fails as unknown/unavailable across `inspect`/`cancel`/`wait`/batch `inspect`, never falling through to the agent-only durable reader. No routine run event sends a parent message or notification.
43
- - `list`, `inspect`, and `wait` share one run projection (`src/core/run-projection.ts`: `projectLiveAgent`/`projectDurableAgent` for agents, `projectWorkflowRun` for workflows) instead of independent per-surface field extraction. Task identity includes the resolved profile and an explicit `harness` (registered pi-* config, or equal to backend for agy/claude/codex), persisted at the source (run-record metadata, live progress nodes, workflow child snapshots) by every backend — never reparsed or guessed from a profile/subagentType name.
43
+ - `list`, `inspect`, and `wait` share one run projection (`src/core/run-projection.ts`: `projectLiveAgent`/`projectDurableAgent` for agents, `projectWorkflowRun` for workflows) instead of independent per-surface field extraction. Task identity includes the resolved profile and an explicit `harness` (registered pi-* config, or equal to backend for agy/claude/codex/grok/muse), persisted at the source (run-record metadata, live progress nodes, workflow child snapshots) by every backend — never reparsed or guessed from a profile/subagentType name.
44
44
 
45
45
  ## Workflow contract
46
46
 
47
- `workflow` remains trusted JavaScript orchestration over the same external-only role roster, including registered named Pi harnesses. Its first statement declares `meta.apiVersion: 1`; missing/unsupported versions fail before child launch. Every workflow `agent()` child uses the same `role`/optional `harness` resolution as `Agent`, with legacy exact-profile `subagent_type` available as an escape hatch. A workflow run freezes one resolved profile (and model) snapshot up front — real on-disk profiles plus every registered pi harness's synthesized canonical roles — so a synthesized pi role executes exactly like a real file, and a `harnesses.json` edit mid-run cannot desync what the replay fingerprint recorded from what actually ran. Successful calls return their value; failed, cancelled, and timed-out calls throw structured catchable child errors. Explicitly handled errors permit siblings to finish; escaping errors fail the workflow and drain active siblings.
47
+ `workflow` remains trusted JavaScript orchestration over the same external-only role roster, including all five CLI harnesses and registered named Pi harnesses. Its first statement declares `meta.apiVersion: 1`; missing/unsupported versions fail before child launch. Every workflow `agent()` child uses the same `role`/optional `harness` resolution as `Agent`, with legacy exact `subagent_type` available as an escape hatch. A workflow run freezes one catalog snapshot up front — built-in roles, shared roles, exact overrides, and the harness model, thinking, and resource preset — so a later edit to `settings.json`, `roles/`, or `overrides/` cannot desync what the replay fingerprint recorded from what that run executed. Replay reuses a prefix only when the effective descriptor is identical; changed instructions invalidate reuse. Successful calls return their value; failed, cancelled, and timed-out calls throw structured catchable child errors. Explicitly handled errors permit siblings to finish; escaping errors fail the workflow and drain active siblings.
48
48
 
49
49
  Replay is explicit through a persisted `scriptPath` plus `resumeFromRunId`. Reuse only the longest unchanged successful prefix; changed or unsuccessful calls and their suffix execute again. Never describe script recomposition as making child reruns free or side-effect-free. There is no live steering or automatic repaired-script retry.
50
50
 
51
- Use workflows for requested fan-out or multi-agent orchestration across Claude/Codex/Antigravity lanes. Do not route native Pi subagents through `workflow`.
51
+ Use workflows for requested fan-out or multi-agent orchestration across agy, Claude, Codex, Grok, Muse, and named Pi harnesses. Do not route native Pi subagents through `workflow`.
52
52
 
53
53
  ## Essential test mandate
54
54
 
55
55
  - Keep one authoritative automated test per behavior at the lowest useful layer.
56
- - Test public contracts, trust/security boundaries, process lifecycle, cancellation/timeouts, profile rollback, workflow execution/resume, and receipt integrity.
56
+ - Test public contracts, trust/security boundaries, process lifecycle, cancellation/timeouts, harness-registration rollback, the zero-generated-file catalog, one-time settings conversion, the purge inventory, workflow execution/resume, and receipt integrity.
57
57
  - Do not test cosmetic rendering variations, prompt prose fragments, trivial accessors, or a model's interpretation of instructions.
58
58
  - Do not repeat the same behavior at unit, integration, and E2E levels. A regression test must replace or extend overlapping coverage.
59
- - `npm test` must remain deterministic and offline. Real-provider E2E is opt-in and change-triggered. Its default lane never prompts a root LLM to pick the tool call — it builds an in-process Pi SDK session with a faux, never-streamed root model and calls the `Agent`/`workflow`/`external_runs` tool executors directly, so only the selected external backend's own child (a real spawned CLI process, or, for `pi`, a real in-process nested Pi child) is real: `npm run e2e -- --backend <claude|codex|agy|grok|muse> [--workflow|--interrupt]`, or `npm run e2e -- --backend pi --harness <registered-name> [--workflow|--interrupt]` for a named Pi harness the caller has already registered in their own real `harnesses.json` with real credentials configured. A separate `--routing-smoke` lane spawns a real `pi` CLI process with a real root model and a plain-language instruction to exercise natural-language tool-call routing end to end; it requires real root Pi auth/config and is never the default.
59
+ - `npm test` must remain deterministic and offline. Real-provider E2E is opt-in and change-triggered. Its default lane never prompts a root LLM to pick the tool call — it builds an in-process Pi SDK session with a faux, never-streamed root model and calls the `Agent`/`workflow`/`external_runs` tool executors directly, so only the selected external backend's own child (a real spawned CLI process, or, for `pi`, a real in-process nested Pi child) is real: `npm run e2e -- --backend <claude|codex|agy|grok|muse> [--workflow|--interrupt]`, or `npm run e2e -- --backend pi --harness <registered-name> [--workflow|--interrupt]` for a named Pi harness the caller has already registered under `harnesses` in their own real `settings.json` with real credentials configured. A separate `--routing-smoke` lane spawns a real `pi` CLI process with a real root model and a plain-language instruction to exercise natural-language tool-call routing end to end; it requires real root Pi auth/config and is never the default. An auth-blocked lane is not passing evidence.
60
60
 
61
61
  ## Verification and release
62
62
 
63
63
  - Run `npm run check`, `npm pack --dry-run --json`, and `npm audit --omit=dev --audit-level=high` before release.
64
64
  - Follow `docs/field-testing.md` for real-backend checks and `docs/releasing.md` for the automated release flow, versioning, trusted publishing, and registry verification.
65
- - Releases are automated: merging a PR that bumps the version (labelled `release:patch|minor|major|prerelease`, or `release:none` to skip) tags the commit, publishes to npm via OIDC trusted publishing, and creates the GitHub release. `release:none` PRs must not change shipped files without a bump; CI enforces the label/version/CHANGELOG contract.
65
+ - Releases are automated: merging a PR that bumps the version (labelled `release:patch|minor|major|prerelease`, or `release:none` to skip) tags the commit, publishes to npm via OIDC trusted publishing, and creates the GitHub release. `release:none` PRs must not change shipped files without a bump; CI enforces the label/version/CHANGELOG contract. For this redesign, the product owner explicitly chose to remain on v2: target `2.6.0-external.0` with `release:minor`. Removing slash commands and moving to settings v4 still requires an explicit breaking-update notice: and put the upgrade notice (removed commands, one-time `/external settings convert`, originals kept, v4 ignores old paths, purge deletes only selected legacy copies, downgrade after purge needs the user's backup) in the CHANGELOG section that becomes the GitHub release notes. Do not hide that break in a patch note.
66
66
  - Never republish an existing npm version or force-push release history.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,43 @@ All notable changes to pi-flow external are documented here.
4
4
 
5
5
  ## Unreleased
6
6
 
7
+ ### Fixed
8
+
9
+ - Antigravity's unsupported-tier error tells the caller to pass `permission: "danger"` or change `defaultPermission`. Omitting the call tier keeps the global default.
10
+ - Unsupported-permission and resume failures keep resolved permission, parent context, and budget on the run receipt. The same fields are written on a launched finish.
11
+ - Conversion reports a malformed short `pi-<role>.md` wildcard template and leaves it in place. Parsed native profiles stay excluded.
12
+ - An unresolved delegation intent shows the tier as unresolved instead of borrowing another harness's enforcement label.
13
+
14
+ ## [2.6.0-external.0] - 2026-09-23
15
+
16
+ ### Breaking upgrade
17
+
18
+ - Removed `/external profiles`, `/external profile create`, and `/pi-flow-profile create` without aliases. Use `/external roles`, `/external role create`, or `/external harness create` instead. Agent/workflow `role`, `harness`, and exact `subagent_type` APIs remain supported.
19
+ - Roles describe intent. They no longer declare or enforce a permission floor, and role names no longer raise `edit` to `danger`. The effective tier is the call's `permission`, otherwise settings `defaultPermission` (`danger` unless changed). `permission` and `capabilitySet` in role or override frontmatter are obsolete and rejected. Conversion strips them from copied overrides.
20
+ - Antigravity accepts only autonomous `danger`. `readonly` and `edit` are rejected instead of being run unsandboxed. Claude, Codex, Grok, Muse, and named Pi harnesses map a requested tier onto the mode they actually support.
21
+ - Named Pi harness registrations select a resource preset, `minimal` or `skills`. `minimal` is the default, including a legacy entry that omits the field, and leaves skills unloaded. `skills` loads installed skills; project skills load only when the project is trusted. Extensions, prompt templates, and themes stay unloaded. The catalog, frozen workflow descriptor, replay fingerprint, and spawn use the registration value. Curated tool lists are not an OS sandbox.
22
+ - The 2.5.0 `pi-*` shared-role templates and `piCapabilitySets` / project capability overlays are not a second runtime. Conversion turns `subagents/pi-<role>.md` with `harness: "pi-*"` into an ordinary cross-harness role and does not copy `piCapabilitySets`. Trust threading for project resources and pre-spawn failure evidence remain.
23
+ - Execution configuration now has one authoritative file: `pi-flow-external/settings.json` version 4, including named Pi harness model/thinking registrations. Run `/external settings convert` once for an older installation; invalid or conflicting inputs block conversion rather than being discarded.
24
+ - The six built-in roles are synthesized for agy, Claude, Codex, Grok, Muse, and registered Pi harnesses. No default profile files or seed markers are written. Shared authored roles live in `pi-flow-external/roles/`; exact overrides live in `pi-flow-external/overrides/`.
25
+ - Conversion copies custom external profiles and preserves deleted-default intent, leaving original files in place. Once v4 is active, legacy `subagents/` profiles and `harnesses.json` are ignored, not fallback configuration.
26
+ - Optional `/external [danger]purge-old-files` explicitly deletes the listed legacy profiles (including customized copies), seed markers, and old harness registry. Nonstandard names require individual selection. Native/unrelated files and current configuration/evidence are excluded. Downgrading after purge requires your own backup; simultaneous old/new-version configuration is unsupported.
27
+
28
+ ### Added
29
+
30
+ - Unified role discovery, inspection, offline shared-role authoring, exact override creation, harness registration, and validated settings editing. All execution/discovery consumers use the same effective catalog.
31
+ - Frozen workflow catalog/model snapshots and explicit invalid/disabled selection errors prevent hidden fallback to a built-in role or another harness.
32
+ - Isolated v4 E2E fixtures and explicit permission selection, including Grok `danger` checks without sandbox workarounds.
33
+
34
+ ## [2.5.0-external.0] - 2026-09-22
35
+
36
+ Shipped on main as 2.5.0-external.0 (shared Pi role templates and capability sets). 2.6.0-external.0 above replaces that runtime. This entry stays as the record of what 2.5.0 shipped.
37
+
38
+ ### Added
39
+
40
+ - Shared custom Pi roles as `subagents/pi-<role>.md` with `harness: "pi-*"`, materialized onto registered Pi harnesses only.
41
+ - Named `piCapabilitySets` in global settings and trusted project settings, selected by profile `capabilitySet`.
42
+ - Pre-spawn workflow failures finish the queued run record instead of leaving it incomplete.
43
+
7
44
  ## [2.4.2-external.0] - 2026-09-23
8
45
 
9
46
  ### Fixed
package/CONTEXT.md CHANGED
@@ -4,15 +4,19 @@ This package is a fork of pi-flow whose `Agent` and `workflow` tools are reserve
4
4
 
5
5
  ## Domain language
6
6
 
7
- - **External profile:** A markdown file under `~/.pi/agent/subagents/*.md` whose filename and frontmatter select `backend: claude`, `backend: codex`, `backend: agy`, `backend: grok`, or (only alongside a registered `harness:`) `backend: pi`.
8
- - **Native Pi subagent:** A subagent exposed by Pi's native subagent system. This fork intentionally does not route it through `Agent`.
9
- - **Agent call:** A direct external delegation selected by `role` and optional `harness` (an external CLI or a registered named Pi harness), or by legacy exact-profile `subagent_type`.
10
- - **Named Pi harness configuration:** A `pi-<label>` entry in `harnesses.json` pinning a `provider/model` id and a thinking level, run in-process via Pi's own SDK rather than as a spawned CLI. `harness` selects *which* configuration; `backend` (always the literal `"pi"` for these) selects the execution mechanism. The six default roles are synthesized automatically for every registered pi harness from the same canonical source the three CLI backends' default files already come from.
7
+ - **Settings:** The only execution-configuration file, `$PI_CODING_AGENT_DIR/pi-flow-external/settings.json` version 4. It holds the default harness, concurrency, timeouts, permission and budget defaults, retention, named `pi-*` harness entries, and `disabledProfiles`. The five CLI harnesses (`agy`, `claude`, `codex`, `grok`, `muse`) are built in.
8
+ - **Harness:** Where and how a role executes. One of the five CLI harnesses, or a registered `pi-<label>` entry in settings. Discovery and readiness are separate: a listed harness may still be uninstalled or unauthenticated.
9
+ - **Role:** The work instructions and description. A role does not grant authority. Six built-ins (explorer, planner, implementer, reviewer, qa, worker) exist in memory for every CLI harness and every registered named Pi harness. A user-authored shared role is one file, `pi-flow-external/roles/<role>.md`, with `description` only, and it applies to every harness. `roles/reviewer.md` replaces that built-in everywhere.
10
+ - **Exact override:** `pi-flow-external/overrides/<harness>-<role>.md`, a complete replacement of one execution identity, using the existing profile schema. `tools` is enforced only on a named Pi harness. Nonstandard legacy names stay selectable through exact `subagent_type`.
11
+ - **Resolved profile:** The internal execution object produced by binding a role definition to a harness. It is not a user-facing command and it is not a file under Pi's native `subagents/` directory. Precedence is exact override, then shared role, then built-in.
12
+ - **Native Pi subagent:** A subagent exposed by Pi's native subagent system. This fork intentionally does not route it through `Agent` and does not modify those files.
13
+ - **Agent call:** A direct external delegation selected by `role` and optional `harness` (`agy`, `claude`, `codex`, `grok`, `muse`, or a registered named Pi harness), or by legacy exact `subagent_type`.
14
+ - **Named Pi harness configuration:** A `pi-<label>` entry in settings v4 `harnesses`, pinning a `provider/model` id and a thinking level, run in-process via Pi's own SDK rather than as a spawned CLI. `harness` selects *which* configuration; `backend` (always the literal `"pi"` for these) selects the execution mechanism. The six built-in roles are available on it immediately from the same canonical source as the five CLI harnesses, with zero generated files.
11
15
  - **Parent context:** An opt-in frozen text snapshot: `none` (default), `recent` with last N user turns including the current turn, or `full` available post-compaction conversation. It is background for a new external conversation, not a native session clone, system-prompt inheritance, or guaranteed cache reuse. Sharing excludes thinking and pending calls and cannot be combined with child `resume`. Workflow children share one invocation-time snapshot; prior child outputs remain explicit inputs.
12
16
  - **Workflow call:** Trusted JavaScript orchestration beginning with `meta.apiVersion: 1` that may fan out several explicit external `agent()` calls. Each child returns a value or throws a catchable `ChildRunError`; escaping child errors fail the workflow and drain siblings.
13
17
  - **Background run:** A validated, registered `Agent` or `workflow` invocation that returns a stable handle while work remains owned by the originating session. It is not a daemon or cross-session job.
14
18
  - **Run supervision:** `external_runs` list/inspect/wait/cancel over stable IDs. Wait observes selected terminal outcomes without cancelling pending work; output and diagnostics use opaque cursors.
15
- - **Delegation intent:** The backend, profile, task label, profile purpose, and workspace shown to the user before and during direct execution.
19
+ - **Delegation intent:** The backend, resolved identity, task label, role purpose, and workspace shown to the user before and during direct execution.
16
20
  - **Access disclosure:** A visible statement of effective external CLI access and the external-host boundary; it does not create a sandbox or read-only guarantee.
17
21
  - **Evidence ID:** The short terminal reference to a persisted external run.
18
22
  - **Expanded receipt:** Bounded canonical output, full run navigation, and advanced record or workflow-journal evidence.
@@ -26,36 +30,36 @@ This package is a fork of pi-flow whose `Agent` and `workflow` tools are reserve
26
30
  - Native Pi work -> native subagent tool.
27
31
  - Claude Code / Codex CLI / Antigravity / Grok Build CLI / Muse Code work, and named Pi harness work -> this extension's `Agent` or `workflow`.
28
32
 
29
- The split is global and intentional to avoid tool ambiguity across projects. The ordinary driver sees `Agent`, optional `workflow`, and read-only `external_help`; the profile finalizer is activated only by `/external profile create`. External children start in the requested working directory, but backend-native nested helpers may create or use a separate workspace; prompts that request further nesting should include explicit absolute paths and all required context. Named Pi harness children are the one exception to "backend-native nested helpers may use a different workspace": a pi child runs in-process with no extensions, skills, prompts, or themes loaded, so it structurally cannot itself launch a further nested agent.
33
+ The split is global and intentional to avoid tool ambiguity across projects. The ordinary driver sees `Agent`, optional `workflow`, and read-only `external_help`. `pi_flow_role_create` is activated only by `/external role create`. `pi_flow_harness_create` is activated only by `/external harness create`. External children start in the requested working directory, but backend-native nested helpers may create or use a separate workspace; prompts that request further nesting should include explicit absolute paths and all required context. Named Pi harness children are the one exception to "backend-native nested helpers may use a different workspace": a pi child runs in-process with no extensions, prompt templates, or themes loaded. The registration preset chooses skills: `minimal` leaves them unloaded, and `skills` loads installed skills, including project skills only when the project is trusted. Extensions stay unloaded, so the child still cannot launch a further nested agent through an extension.
30
34
 
31
35
  ## User control surface
32
36
 
33
- User operations are namespaced under `/external`: status, Doctor, settings, profiles, profile creation, workflows, interactive run navigation, durable run summary/pruning, and help. `/external runs` uses the same registry/readers as `external_runs`, including every cursor-paged workflow child and output/diagnostic page. Extension-owned concurrency, timeout, and global default-harness settings live in `$PI_CODING_AGENT_DIR/pi-flow-external/settings.json`; profiles remain authoritative for downstream backend/model/thinking metadata. Named Pi harness configurations live in the sibling `pi-flow-external/harnesses.json` registry. A trusted project may override only `defaultHarness`, which may itself name a registered `pi-*` harness.
37
+ User operations are namespaced under `/external`: overview, doctor, settings, `settings edit`, `settings convert` (the one-time pre-v4 preview and apply; `settings` points there), harnesses, `harness create`, roles, `role create`, `role inspect`, `role override`, the literal `[danger]purge-old-files` maintenance command, workflows, interactive run navigation, durable run summary/pruning, and help. The next invocation reads the live configuration. A concurrency-cap change waits until active and queued work drains. A workflow already running keeps its frozen snapshot. `/external profiles`, `/external profile create`, and `/pi-flow-profile create` are removed and are not aliases. `/external runs` uses the same registry/readers as `external_runs`, including every cursor-paged workflow child and output/diagnostic page. Settings v4 is the only execution-configuration file, including named Pi harness registrations. Shared roles and exact overrides are the extension-owned markdown. A trusted project may override only `defaultHarness`, which may itself name a registered `pi-*` harness. After version 4 is active, old `subagents/` profiles, seed markers, and `harnesses.json` are ignored. Purge is optional, destructive, and separate from conversion.
34
38
 
35
- > Next agent: full implementation map is `docs/ARCHITECTURE_SNAPSHOT.md` (verified 2026-09-21). Keep this file and `AGENTS.md` as the persistent breadcrumbs.
39
+ > Next agent: full implementation map is `docs/ARCHITECTURE_SNAPSHOT.md` (configuration contract 2026-09-23; execution-adapter notes last verified 2026-09-21). Keep this file and `AGENTS.md` as the persistent breadcrumbs. Historical files under `docs/plans/` are not rewritten.
36
40
 
37
41
  ## Design stance
38
42
 
39
- Guardrails exist to keep lanes from bleeding into each other (a reviewer that edits, a worker that fixes), not to constrain how a model works. Role bodies stay short: define the job, state the boundary, then get out of the way and trust the model's judgment on approach, depth, and method. The generic `worker` role covers non-coding tasks with no method constraints at all. The shipped default roster is five code-oriented roles (explorer, planner, implementer, reviewer, qa) plus `worker`; other roles are user-created through `/external profile create` and still receive the name-based execution permission floor. External agents must not be arbitrarily handcuffed just to satisfy a theoretical safety checklist: what transpires on the ground is what matters. When an external agent needs heightened permissions to inspect, run tests, and execute tools, the system defaults to providing that authority rather than blocking execution.
43
+ Guardrails exist to keep lanes from bleeding into each other (a reviewer that edits, a worker that fixes), not to constrain how a model works. Role bodies stay short: define the job, state the boundary, then get out of the way and trust the model's judgment on approach, depth, and method. The generic `worker` role covers non-coding tasks with no method constraints at all. The shipped roster is five code-oriented roles (explorer, planner, implementer, reviewer, qa) plus `worker`, in memory for `agy`, `claude`, `codex`, `grok`, `muse`, and every registered named Pi harness. Other roles are user-created once through `/external role create`. Authority is the call's `permission`, or `defaultPermission` when the call omits it. The default remains `danger`. A role file cannot raise or lower that tier.
40
44
 
41
- Role resolution is mechanical, not model-routed: a profile whose name begins with its declared backend plus `-` exposes the remaining name as its role (`agy-reviewer` -> `reviewer`). The selected role and harness must resolve to that exact profile; unavailable roles report their supported harnesses and never fall back. Nonstandard names remain available only through legacy exact `subagent_type`. The resolved profile stays authoritative for its instructions, model, and permissions.
45
+ Role resolution is mechanical, not model-routed. The identity is `<harness>-<role>` (`agy-reviewer` is role `reviewer` on harness `agy`). Precedence is exact override, then shared role, then built-in. The selected role and harness must resolve to that exact identity; unavailable roles report their supported harnesses and do not fall back. Identities listed in `disabledProfiles` are blocked for both `role` and exact `subagent_type`. Nonstandard names remain available only through legacy exact `subagent_type`. The resolved profile stays authoritative for its instructions and model. Permission is not stored on the role.
42
46
 
43
- The always-visible parent guidance is limited to a compact role catalog and essential routing/safety facts. Catalog availability reflects configured profiles, not CLI installation or authentication. `external_help` supplies descriptions, permission details, workflow APIs/examples, and trust-aware saved-workflow discovery only when requested.
47
+ The always-visible parent guidance is limited to a compact role catalog and essential routing/safety facts. Catalog availability reflects the resolved roster, not CLI installation or authentication. `external_help` supplies descriptions, permission details, workflow APIs/examples, and trust-aware saved-workflow discovery only when requested.
44
48
 
45
49
  Workflow replay is explicit and successful-prefix-only. A persisted script resumed with `resumeFromRunId` reuses the longest unchanged prefix whose fingerprints and outcomes are successful; the first changed or unsuccessful call and its suffix run again. Script recomposition is cheap, but child execution can cost money or repeat side effects. There is no automatic repaired-script replay, completion-order winner selection, or live steering.
46
50
 
47
- Applied to Antigravity (`agy`), that stance is absolute: the harness offers no granular headless permission mode — its default sandbox (`proceed-in-sandbox`) hard-denies even read-only tools like `read_url_content`, and `--dangerously-skip-permissions` is the only unsandboxed mode. So every agy run is unsandboxed, and `readonly`/`edit` on agy profiles are advisory instructions carried by the profile body, never an enforced boundary. We disclose that plainly (`unsandboxed external CLI`) rather than pretend a read-only tier restrains what the harness will actually allow.
51
+ Applied to Antigravity (`agy`), the harness offers no granular headless permission mode. `--dangerously-skip-permissions` is the only accepted mode. A `readonly` or `edit` request is rejected. An accepted agy run discloses as `unsandboxed external CLI`.
48
52
 
49
- Applied to named Pi harness configurations, the same "get out of the way" stance takes a different, honestly-scoped shape: a curated, builtins-only tool surface (no project/user extensions, skills, prompts, or themes) genuinely bounds which tool *names* a pi child can reach — that claim is mechanically true because nothing else loads — but it is not a claim that `bash` at `danger` tier is any less dangerous than on an external CLI once it is granted. `readonly`/`edit` tiers get a real default tool allow-list (`read`/`grep`/`find`/`ls`, plus `edit`/`write` at `edit`) so those tiers are actually usable, not just technically restrictive. Execution-lane roles requested at `edit` are still elevated to `danger` (the same floor as Claude), because `edit` excludes `bash` entirely and an execution lane without shell access cannot do its job.
53
+ Applied to named Pi harness configurations, a curated tool list bounds which tool names exist. It is not an OS sandbox. `readonly` and `edit` keep their tool lists, including the absence of `bash` at `edit`. The caller's tier is not raised because the role is named implementer, qa, or worker. The registration preset is `minimal` (skills stay unloaded) or `skills` (installed skills load; project skills require a trusted project). An omitted preset is `minimal`. Extensions, prompt templates, and themes stay unloaded.
50
54
 
51
- Applied to the Grok Build CLI, permission tiers map onto its own kernel-enforced `--sandbox` axis (`read-only`/`workspace`/`off`), always paired with `--permission-mode bypassPermissions` — bypass only removes the interactive approval prompt, never the sandbox itself, so every tier is a genuine boundary rather than an advisory one. `edit` already permits shell execution (sandboxed to workspace writes), so, unlike Claude and pi, Grok execution-lane roles need no `edit`→`danger` floor. `readonly`'s network-blocking guarantee is Linux-only, and sandbox startup can fail closed on some macOS hosts rather than silently running unsandboxed; that fail-closed behavior is preserved rather than weakened for cross-platform convenience.
55
+ Applied to the Grok Build CLI, permission tiers map onto its own kernel-enforced `--sandbox` axis (`read-only`/`workspace`/`off`), always paired with `--permission-mode bypassPermissions` — bypass only removes the interactive approval prompt, never the sandbox itself, so every tier is a genuine boundary rather than an advisory one. `edit` already permits shell execution (sandboxed to workspace writes). A role name does not change that mapping. `readonly`'s network-blocking guarantee is Linux-only, and sandbox startup can fail closed on some macOS hosts rather than silently running unsandboxed; that fail-closed behavior is preserved rather than weakened for cross-platform convenience.
52
56
 
53
- Applied to Muse Code, approval and its own sandbox are ON by default, so every tier passes `--disable-approval` (a headless run must never hang on an interactive prompt). `readonly` additionally strips non-shell writes and shell execution (`--disable-write --disable-shell`); `edit` leaves the sandbox enabled with only approval bypassed, so shell/write stay available within it — Muse needs no `edit`→`danger` floor either, for the same reason as Grok. `danger` uses `--yolo`, which disables approval and the sandbox *and additionally trusts the workspace for this run* (loads its skills/rules) — a materially broader grant than an unsandboxed run alone, so its disclosure names that extra trust rather than collapsing it into the same terse label every other danger tier uses. `resume` was verified against the real CLI (`exec --session-id`, not the separate interactive-only `resume` subcommand): two independent processes sharing one session id shared context, confirmed by the second recalling a fact only told to the first. Usage/cost has never been observed on any run, so it is reported unknown rather than a fabricated zero or estimate. Root-run ownership is checked explicitly: only envelopes whose `payload.run_stream.id` matches the id established from that same process's own `runtime.command.accepted`/`session.run.linked` bootstrap pair can finalize, fail, or contribute partial output to the result, so a nested or foreign-run terminal event can never masquerade as the root's own answer, and `sessionId` is captured once from that bootstrap rather than from every subsequent envelope. Nested-agent (sub-delegation) detection always reports false — every probe run only ever produced internal `reminder.agent.*` skill-reminder tasks, never an actual delegation, so there is no confirmed event shape to key off, and a speculative match would risk spuriously extending the nested-timeout deadline on ordinary internal task activity.
57
+ Applied to Muse Code, approval and its own sandbox are ON by default, so every tier passes `--disable-approval` (a headless run must never hang on an interactive prompt). `readonly` additionally strips non-shell writes and shell execution (`--disable-write --disable-shell`); `edit` leaves the sandbox enabled with only approval bypassed, so shell/write stay available within it. A role name does not change that mapping. `danger` uses `--yolo`, which disables approval and the sandbox *and additionally trusts the workspace for this run* (loads its skills/rules) — a materially broader grant than an unsandboxed run alone, so its disclosure names that extra trust rather than collapsing it into the same terse label every other danger tier uses. `resume` was verified against the real CLI (`exec --session-id`, not the separate interactive-only `resume` subcommand): two independent processes sharing one session id shared context, confirmed by the second recalling a fact only told to the first. Usage/cost has never been observed on any run, so it is reported unknown rather than a fabricated zero or estimate. Root-run ownership is checked explicitly: only envelopes whose `payload.run_stream.id` matches the id established from that same process's own `runtime.command.accepted`/`session.run.linked` bootstrap pair can finalize, fail, or contribute partial output to the result, so a nested or foreign-run terminal event can never masquerade as the root's own answer, and `sessionId` is captured once from that bootstrap rather than from every subsequent envelope. Nested-agent (sub-delegation) detection always reports false — every probe run only ever produced internal `reminder.agent.*` skill-reminder tasks, never an actual delegation, so there is no confirmed event shape to key off, and a speculative match would risk spuriously extending the nested-timeout deadline on ordinary internal task activity.
54
58
 
55
59
  ## Known inelegance
56
60
 
57
- <!-- ponytail: one-backend-per-file profile format forces role x backend file duplication; extend src/profiles.ts with multi-backend profiles (e.g. backends: [claude, codex, agy, grok, muse] plus per-backend model map) if maintaining N copies of identical role bodies ever hurts -->
58
- A standardized role roster needs one file per role per backend (e.g. `claude-qa`, `codex-qa`, `agy-qa`, `grok-qa`, `muse-qa` with identical bodies) because the profile format binds `backend:` and `model:` to a single file. The duplication is accepted for now; a future format extension could let one role file cover all backends. Named Pi harnesses partially resolve this for the six canonical roles specifically (synthesized once from a single shared source, applied identically to every registered `pi-*` config, with zero files written) — but a *custom* role outside the six still needs one file per pi harness, exactly like the five CLI backends; a user with several registered pi harnesses wanting the same custom role on all of them still duplicates one file per harness. The same applies to `permission:` tiers: one tier per profile, but backends interpret tiers differently (Claude and pi deny Bash at anything below `danger`), so command-running roles declare `danger` and maintain `danger` as their floor even if an explicit `edit` override is requested. Custom profiles outside the role-name convention get that floor only by declaring `permission: danger`; an undeclared custom profile that meets an `edit` tier still hits headless Bash denials (or, on pi, a curated tool set without `bash`), so `permission: danger` doubles as the lane's "needs shell" capability declaration. Named Pi harnesses also ship with no resume and no enforced budget cap in v1 — both deliberate limitations, tracked alongside the shared-custom-role gap in issue #43, not oversights.
61
+ <!-- ponytail: shared roles removed the role x harness file product; keep backend-specific prompts in exact overrides. Do not grow a second inheritance language. -->
62
+ Shared roles are one file for every harness. Backend-specific model, tools, budget, or instruction changes are a separate exact override, which replaces that one identity completely. Permission is the call or the global default, not a field on the role. Claude `edit` still denies Bash headlessly, and a pi `edit` child still has no `bash` tool. Ask for `danger` on the call when the task needs a shell. Named Pi harnesses still have no resume and no enforced budget cap. Extensions and MCP stay out. `tools` remains enforceable only on Pi.
59
63
 
60
64
  ## Evidence boundary
61
65