@diousk/pi-subagents-fast 0.20.0 → 0.22.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,31 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.22.0] - 2026-10-02
11
+
12
+ > **Breaking: requires Pi 1.0.0 or newer.** Update Pi before using this version.
13
+
14
+ ### Added
15
+ - **Optional Jev routing with four modes.** `routingMode` defaults to `auto`: custom agents, then a Markdown guideline, then Jev. Choose `shadow` to record suggestions without changing models, `jev` for Jev-first selection with default-priority fallback, or `off` to disable routing guidance and classification. Configure the mode, guideline, model descriptions and optional TypeSafe credentials in `subagents.json` or `/agents → Model routing`; usage is recorded separately, including shadow calls. Jev choices honor Pi's model scope and, when enabled, the extension's `scopeModels` policy.
16
+
17
+ ### Changed
18
+ - **Service-tier compatibility is documented for GPT-6.1 Sol and GPT-6 Luna.** Native transport regression tests cover `fast` and `priority` on initial and resumed turns. Pi 1.0.0's Codex cost estimate still understates responses marked `fast`; this release documents that upstream limitation without recalculating costs.
19
+
20
+ ## [0.21.0] - 2026-09-30
21
+
22
+ > **Breaking: requires Pi 0.99.1 or newer.** Update Pi before installing this version. Legacy transcript and session API compatibility paths have been removed.
23
+
24
+ ### Added
25
+ - **Custom agents accept `service_tier: fast`.** OpenAI Responses and Codex requests forward the configured tier, and `priority` remains accepted.
26
+
27
+ ### Changed
28
+ - **The minimum supported Pi version is 0.99.1.** Development dependencies and blocking CI target that minimum. Scripted model fixtures use transcript replay and physical model lookup.
29
+
30
+ ### Fixed
31
+ - **Mention clones preserve conversation history on Pi 0.99.1.** History is seeded through the session manager, and the live prompt is supplied through the resource loader. Historical system messages do not grant the clone the parent's tools.
32
+ - **Subagent tool scope covers nested execution.** A bound `tool_call` guard blocks out-of-scope deferred/codemode calls, while tool narrowing preserves exposure. Synthetic extension paths match their logical names.
33
+ - **Output transcripts preserve system messages.** System/tool-state changes receive a `system` entry type, and a leading system message no longer causes the initial user prompt to be duplicated.
34
+
10
35
  ## [0.20.0] - 2026-09-09
11
36
 
12
37
  > **Breaking — the npm package is now `@diousk/pi-subagents-fast`.** Existing installs of `@tintinweb/pi-subagents` are not automatically migrated; install this fork with `pi install npm:@diousk/pi-subagents-fast`.
package/README.md CHANGED
@@ -28,6 +28,7 @@ https://github.com/user-attachments/assets/8685261b-9338-4fea-8dfe-1c590d5df543
28
28
  - **Graceful turn limits** — agents get a "wrap up" warning before hard abort, producing clean partial results instead of cut-off output
29
29
  - **Case-insensitive agent types** — `"explore"`, `"Explore"`, `"EXPLORE"` all work. A type that doesn't resolve to exactly one *enabled* agent — unknown, disabled, or ambiguous between two agents differing only by case — falls back to general-purpose with a note, or is refused outright under [`fallbackSubagent: none`](#persistent-settings)
30
30
  - **Fuzzy model selection** — specify models by name (`"haiku"`, `"sonnet"`) instead of full IDs, with automatic filtering to only available/configured models
31
+ - **Model routing** — custom agents first, then a user-supplied Markdown guideline, then optional Jev model selection. Configure it through `/agents → Model routing` or two ordinary JSON fields; see [Model routing](#model-routing)
31
32
  - **Context inheritance** — optionally fork the parent conversation into a sub-agent so it knows what's been discussed
32
33
  - **Persistent agent memory** — three scopes (project, local, user) with automatic read-only fallback for agents without write tools
33
34
  - **Git worktree isolation** — run agents in isolated repo copies; changes auto-committed to branches on completion
@@ -76,7 +77,9 @@ npm pack --dry-run
76
77
 
77
78
  The `prepublishOnly` script runs lint, typecheck, tests, and the build before npm uploads the package. After publishing, install it with `pi install npm:@diousk/pi-subagents-fast`.
78
79
 
79
- Requires pi **0.84.0 or newer**: the [`SubagentWorkflow`](#subagentworkflow) tool builds on `constrainedSampling` (pi 0.82.0) and pi-tui's `stripTerminalSequences` (0.84.0). The `peerDependencies` range declares it, so npm flags an older pi at install time.
80
+ Requires **Pi 1.0.0 or newer** and **Node.js 22.19.0 or newer**. Update Pi before installing this extension. The `peerDependencies` range declares the minimum, so npm flags an older Pi at install time.
81
+
82
+ Development and CI use **Pi 1.0.0**. Jev uses Pi's native classifier API. Child sessions reuse the parent's configured providers and authentication; model availability, limits and pricing come from Pi's catalog.
80
83
 
81
84
  ### Other hosts
82
85
 
@@ -337,8 +340,8 @@ All fields are optional — sensible defaults for everything.
337
340
  | `disallowed_tools` | — | Comma-separated tools to deny even if extensions provide them |
338
341
  | `isolation` | — | Set to `worktree` to run in an isolated git worktree, or `off` to refuse one even when the caller passes `isolation: "worktree"` (frontmatter is authoritative). `none`, `no`, and `false` are accepted spellings of `off` |
339
342
  | `model` | inherit parent | Model — `provider/modelId` or fuzzy name (`"haiku"`, `"sonnet"`). Resolved tolerantly (`.`/`-` and a trailing date stamp are interchangeable) and falls back to the same model under another provider if the named one doesn't have it |
340
- | `service_tier` | — | OpenAI Responses/Codex processing tier: `auto`, `default`, `flex`, `priority`, or `scale`. Applied only to `openai-responses` and `openai-codex-responses`; omitted preserves the provider default, and other APIs ignore it without displaying it as active |
341
- | `thinking` | inherit | off, minimal, low, medium, high, xhigh, max — actual availability depends on your pi version and model; pi clamps unsupported levels down |
343
+ | `service_tier` | — | OpenAI Responses/Codex processing tier: `auto`, `default`, `flex`, `fast`, `priority`, or `scale`. Applied only to `openai-responses` and `openai-codex-responses`; omitted preserves the provider default, and other APIs ignore it without displaying it as active |
344
+ | `thinking` | inherit | off, minimal, low, medium, high, xhigh, max — actual availability depends on your pi version and model; pi maps or clamps unsupported levels |
342
345
  | `max_turns` | unlimited | Max agentic turns before graceful shutdown. `0` or omit for unlimited |
343
346
  | `persist_session` | `subagents.json` `rememberAgents` (default `true`) | Persist this subagent as a normal pi session instead of keeping the session in memory only; overrides the `rememberAgents` project default in both directions. It records its spawning session as parent, so it nests under it in `/resume`. The subagent's `.output` transcript is still written either way unless `output_transcript: false` |
344
347
  | `output_transcript` | `true` (or `subagents.json` `outputTranscript`) | Write this subagent's `.output` transcript; when set, overrides the `subagents.json` `outputTranscript` default. Set `false` to write no transcript file or path. Governs only the transcript — independent of `persist_session`, `isolation: worktree`, and `memory:` |
@@ -350,7 +353,19 @@ All fields are optional — sensible defaults for everything.
350
353
  | `isolated` | `false` | Hermetic specialist mode: forces `extensions: false` + `skills: false` + drops `ext:` selectors. Only built-in tools. Distinct from `isolation: worktree` (filesystem) |
351
354
  | `enabled` | `true` | Set to `false` to disable an agent (useful for hiding a default agent per-project) |
352
355
 
353
- For an OpenAI Responses or Codex agent, set `service_tier: priority` to request priority processing. The UI shows the requested tier only when the effective model uses one of those APIs; other providers keep their normal request behavior and do not display a tier as active.
356
+ For an OpenAI Responses or Codex agent, set `service_tier: fast` to request fast processing; `priority` remains a supported alias. Availability depends on the provider and account. The UI shows the requested tier only when the effective model uses one of those APIs. See [OpenAI fast mode](https://developers.openai.com/api/docs/guides/fast-mode).
357
+
358
+ The extension forwards both tiers for `gpt-6.1-sol` and `gpt-6-luna` through Pi's OpenAI Responses and Codex transports. Pi 1.0.0's Codex adapter currently estimates a response marked `fast` at the standard rate; forwarding the tier works, but its displayed cost can be understated. This extension reports Pi's cost estimate without recalculating it. OpenAI documents Fast mode as unavailable for these models with EU data residency.
359
+
360
+ GPT-6.1 Sol supports `low`, `medium`, `high`, `xhigh`, and `max` reasoning. In Pi 1.0.0, `minimal` maps to provider effort `low`; `off` is clamped to that same alias because Sol cannot disable reasoning. The UI reports Pi's logical thinking level. Prefer explicit supported levels in agent files:
361
+
362
+ ```yaml
363
+ model: openai-codex/gpt-6.1-sol
364
+ thinking: medium
365
+ service_tier: fast
366
+ ```
367
+
368
+ Sol requires the Responses API for tool use. Selecting its Pi catalog entry chooses the corresponding Responses transport; no custom provider override is needed. See the [Sol model reference](https://developers.openai.com/api/docs/models/gpt-6.1-sol).
354
369
 
355
370
  Frontmatter is authoritative. If an agent file sets `model`, `thinking`, `service_tier`, `max_turns`, `inherit_context`, `run_in_background`, `isolated`, or `isolation`, those values are locked for that agent. For fields exposed by the `Agent` tool, its parameters only fill values the agent config leaves unspecified; `service_tier` is frontmatter-only.
356
371
 
@@ -408,6 +423,8 @@ A few rules the examples don't make obvious:
408
423
  - `extensions:` is the sole loading authority. `ext:foo` in `tools:` narrows what surfaces; it can't load `foo` on its own. Mismatches fire `extension-error:…` warnings.
409
424
  - Any `ext:` entry flips extension tools to an explicit allowlist — unnamed extensions still load (handlers fire) but expose no tools. So `tools: "*, ext:mcp/search"` exposes only `search` from `mcp`, nothing from any other extension.
410
425
  - Extension names match case-insensitively (`[Mcp]` = `[mcp]`); tool names in `ext:foo/bar` stay case-sensitive.
426
+ - Synthetic extension identities match their logical names: `builtin:mcp` as `mcp`, `builtin:codemode` as `codemode`, and `<inline:foo>` as `foo`. This matches loaded extensions; it does not load Pi CLI built-in factories into SDK child sessions.
427
+ - Scope checks also apply to nested `ctx.executeTool()` calls, including deferred and codemode tools. Deferred tools keep their exposure and are not automatically promoted into model declarations; extension-activated tools remain active only within the agent's scope.
411
428
  - Extensions that register tools **lazily** work too. MCP-backed extensions typically can't enumerate their tools until their servers connect, so they register from `session_start` or `before_agent_start` rather than at load. Subagent scoping is re-derived as tools appear, so these surface normally — including under `ext:` selectors, which keep narrowing correctly no matter when a tool shows up.
412
429
  - Extensions bound into a subagent see **both ends** of that session's lifecycle: `session_start` when the agent starts, `session_shutdown` (reason `quit`) when its session is disposed — on quit, and when its record is evicted ~10 minutes after it finishes. Release per-session resources there; anything left armed outlives the session it belongs to. Handlers are given three seconds on quit, after which teardown proceeds regardless.
413
430
  - An installed **package** extension matches by its package short name (`@scope/pi-subagents` → `[pi-subagents]`), in addition to its path-derived name (a package whose entry is `src/index.ts` also answers to `[src]`). Prefer the package name — the path-derived one is incidental.
@@ -615,6 +632,72 @@ When background agents complete, they notify the main agent. The **join mode** c
615
632
  **Configuration:**
616
633
  - Configure join mode in `/agents` → Settings → Join mode
617
634
 
635
+ ## Model routing
636
+
637
+ Open `/agents → Model routing` to choose a mode and configure a guideline path, models with descriptions, or a masked TypeSafe API key. The menu shows the active mode and source and saves project settings. For machine-wide defaults, edit `~/.pi/agent/subagents.json`; project overrides go in `.pi/subagents.json`. `PI_CODING_AGENT_DIR` changes the global directory along with Pi's other configuration.
638
+
639
+ Set one top-level field to control routing; omitting it means `auto`:
640
+
641
+ ```json
642
+ { "routingMode": "auto" }
643
+ ```
644
+
645
+ | `routingMode` | Behavior |
646
+ |---|---|
647
+ | `auto` (default) | Use the priority table below |
648
+ | `shadow` | Ask Jev and record its suggestion, but keep the model chosen from agent definitions, the guideline or the existing model. Jev requests can incur charges |
649
+ | `jev` | Jev chooses first for every fresh task, even with custom agents, a guideline, explicit model parameters or agent-file model pins. Invalid configuration, unavailable credentials, low confidence or errors keep the default-priority choice |
650
+ | `off` | No custom-agent routing selection guidance, guideline injection or Jev requests. Agents remain callable and existing model/thinking settings still apply |
651
+
652
+ Mode changes apply to subsequent fresh launches. The main agent's routing guidance and tool description refresh before its next turn, so switching to `off` removes previously injected routing instructions.
653
+
654
+ The default priority is:
655
+
656
+ | Priority | Configuration | Who chooses |
657
+ |---|---|---|
658
+ | 1 | At least one enabled custom agent | The main agent chooses an agent from its description. Guideline and Jev are inactive, even if it chooses a built-in agent |
659
+ | 2 | `customGuideline` points to a Markdown file | The main agent reads the guideline and passes its model/thinking choice explicitly |
660
+ | 3 | `jev` is configured | Jev chooses a model from your descriptions for a fresh delegated task |
661
+ | 4 | None of the above | The existing model is used |
662
+
663
+ **Custom agents:** keep using `~/.pi/agent/agents/<name>.md`, `.pi/agents/<name>.md` or `.agents/agents/<name>.md`. No routing setting is needed. Built-in and disabled agents do not activate priority 1. Agent-file model/thinking pins supply the default choice over `Agent` parameters; a confident Jev choice in `jev` mode can replace the model.
664
+
665
+ **Custom guideline:** write your routing rules in `~/.pi/agent/agents/custom-route.md`, then configure:
666
+
667
+ ```json
668
+ {
669
+ "customGuideline": "~/.pi/agent/agents/custom-route.md"
670
+ }
671
+ ```
672
+
673
+ For example, the Markdown can say “Use anthropic/claude-haiku-4-5 for simple edits; use openai-codex/gpt-6.1-sol for debugging concurrency.” The guideline reaches the main agent in full, compact and custom description modes, and is refreshed before each turn, except in `off` mode. Relative paths are resolved from the settings file's directory: `agents/custom-route.md` in `.pi/subagents.json` means `.pi/agents/custom-route.md`. `custom-route.md` and the configured guideline file are excluded from agent discovery. Merely placing a guideline file in the directory does not enable it. A configured missing, empty or oversized guideline produces a diagnostic and keeps the existing model under `auto`; it does not prevent Jev requests under `jev` or `shadow`.
674
+
675
+ **Jev:** provide descriptions and exact model IDs available through your Pi login:
676
+
677
+ ```json
678
+ {
679
+ "jev": {
680
+ "TYPESAFE_API_KEY": "your-typesafe-key",
681
+ "models": [
682
+ { "model": "anthropic/claude-haiku-4-5", "description": "Simple edits, lookup and concise summaries" },
683
+ { "model": "openai-codex/gpt-6.1-sol", "description": "Complex debugging, architecture and concurrency" }
684
+ ]
685
+ }
686
+ }
687
+ ```
688
+
689
+ `TYPESAFE_API_KEY` is optional when Pi already has TypeSafe credentials or the environment variable is set. A literal key must be a nonempty token without whitespace or control characters and applies only to that classifier request; it is saved in the settings file, masked in the menu and omitted from settings events, prompts and routing records. Without a literal key, Pi's native credential availability check runs before classification. Rejected credentials fall back without retrying. The extension does not change environment variables or register a routing provider. `models` accepts 1–254 unique entries with descriptions of 1–4000 characters. Jev entries use exact `provider/model-id` spelling; fuzzy names remain available for explicit `Agent` parameters.
690
+
691
+ Under `auto`, supplying **either** `model` or `thinking` explicitly skips Jev. An agent-file pin also skips it. In workflows, either `model` or `effort` skips it, and workflow options retain their precedence over agent-file defaults. Inherited models remain eligible for Jev. Under `jev`, these choices supply the fallback model, but do not skip classification. Under `shadow`, they remain the actual choice while Jev records a comparison. Jev changes only the model; Pi still determines the effective thinking level.
692
+
693
+ Automatic choices are limited to authenticated models and the current nonempty Pi model scope, regardless of the `scopeModels` setting. When `scopeModels` is enabled, Jev candidates also respect `enabledModels`. Unavailable candidates are excluded and the selected model is checked again after classification, including any changed scope. Jev failure, confidence below 0.6 or a two-second timeout keeps the default-priority model already selected by the main agent, explicit caller or agent definition, otherwise the existing model. This fallback does not make a second Jev request. Cancellation stops startup without launching a fallback. Queued agents read the current mode and classify after dequeue; nested calls share a separate classifier concurrency limit of four and do not take another agent slot.
694
+
695
+ Jev applies to fresh tool, nested, workflow, scheduled, RPC and direct-mention spawns. Schedules read current configuration when they fire. Direct calls, schedules and RPC do not create an extra main-agent turn to interpret a custom guideline: with priorities 1 or 2 under `auto`, or when Jev falls back under `jev`, they use an explicit choice, agent definition or the existing model. Resumes and internal agent-file generation do not call Jev in any mode.
696
+
697
+ Project `routingMode` and `customGuideline` replace the global values; omission inherits them, and `customGuideline: false` disables the guideline. A project `jev` block replaces the **whole** global block, including credentials; omitted fields inside that block do not inherit. `"jev": false` disables global Jev. An invalid mode or malformed settings file disables routing. An invalid explicit Jev block disables that block instead of restoring global paid routing. Other settings changes preserve these inheritance rules.
698
+
699
+ Classifier usage is recorded separately as `routingUsage`, including in `shadow`; it does not consume coding context or workflow output budgets. Its reported cost is added once to the agent and ancestor cost totals, and `reportUsage` can return its token usage to the parent. `routing` records the mode, source, reason, applied `model` and confidence; shadow records `suggestedModel` and leaves `model` unset. `fallbackSource` identifies the default source when a suggestion is observed or Jev falls back. Shadow results show “Jev shadow”. Zero catalog prices are marked `unpriced`; the Agent result shows “Jev price unavailable” instead of treating that as proof the classifier is free.
700
+
618
701
  ## Model Scope
619
702
 
620
703
  **Opt-in:** off by default. Enable via `/agents → Settings → Scope models`.
@@ -627,6 +710,7 @@ When on, each subagent spawn's effective model is validated against pi's own `en
627
710
  |---|---|
628
711
  | Caller-supplied via `Agent({ model: "..." })` | Hard error returned to the orchestrator, listing allowed models |
629
712
  | Caller-supplied via cross-extension RPC (`subagents:rpc:spawn`, e.g. pi-tasks `TaskExecute`) | Hard error returned to the calling extension, listing allowed models |
713
+ | Jev-selected | Excluded from candidates; a scope change during classification keeps the default-priority model |
630
714
  | Pinned in agent frontmatter | Warning toast + the pinned model runs (frontmatter is authoritative) |
631
715
  | Parent-inherited (neither set) | Warning toast + parent's model runs |
632
716
 
@@ -647,6 +731,8 @@ Runtime tuning values set via `/agents` → Settings (max concurrency, max foreg
647
731
 
648
732
  **Precedence:** project overrides global on any field present in both. Missing fields fall back to the hardcoded defaults (max concurrency `10`, max foreground concurrency `0` = unlimited, default max turns unlimited, grace turns `5`, nested depth `2`, join mode `smart`, defaults enabled).
649
733
 
734
+ Routing settings: `routingMode` (`auto | shadow | jev | off`, default `auto`), `customGuideline` (`string | false`, default unset) and `jev` (`{ TYPESAFE_API_KEY?: string, models: { model, description }[] } | false`, default unset). See [Model routing](#model-routing) for examples and the whole-block override rule.
735
+
650
736
  **Nested depth** (`maxSubagentDepth`, default `2`): the hard ceiling on [nested delegation](#nested-subagents), counted from the main session (main = 0, its subagents = 1). `0` or `1` disables nesting project-wide regardless of any agent's `allowed_subagents`. Read when a subagent session is built, so a change applies to agents started after it.
651
737
 
652
738
  **Fallback agent** (`fallbackSubagent`, default `general-purpose`): the agent used when a caller-supplied `subagent_type` doesn't resolve to exactly one enabled agent — unknown, disabled, or ambiguous because two agents differ only by case. Name any enabled agent to route those calls there instead, or set `none` for **strict**, fail-closed dispatch: the call is refused with an error listing the available types, and nothing spawns. Strict mode matters most for background and scheduled calls, which would otherwise start executing a substituted agent before the caller learns anything. Also settable from `/agents → Settings → Fallback agent`. The boolean `false` is accepted as a spelling of `none`, because it would otherwise be dropped as the wrong type and silently leave the permissive default in place. Every other value is read as an agent name, so a mistaken `off` fails loudly at dispatch rather than meaning one thing in the settings file and another in the resolver. A fallback agent that is itself unknown or disabled is a misconfiguration and is reported rather than quietly replaced. Note the default is unchanged and stays permissive by design: with `disableDefaultAgents` and no `general-purpose` of your own, an unresolvable type still resolves to a built-in config carrying *all* tools — set `none` (or name one of your own agents) to close that.
@@ -667,7 +753,7 @@ Runtime tuning values set via `/agents` → Settings (max concurrency, max foreg
667
753
 
668
754
  **Report usage to session** (`reportUsage`, default `false`): whether subagent spend is added to *this* session's own totals. Subagents run in their own pi sessions, so by default pi's footer, statusline and `/cost` count only what the main model spent — a session that delegated most of its work reads as nearly free. Turn it on and each `Agent` / `get_subagent_result` / `steer_subagent` result carries the spend accumulated since the last one, which pi folds into `getSessionStats()`; `/cost` attributes it to the **Tools/summaries** bucket. Toggle via `/agents → Settings → Report usage to session`; applied live.
669
755
 
670
- Three things worth knowing about the numbers. Every token component is reported, `cacheRead` included — the cached prefix genuinely is re-read and re-billed on every call, and pi counts it the same way for the session's own messages, so withholding it would make a subagent's rows count differently from every other row in one total. (The extension's *own* token displays still leave it out, which is a different question: there it inflates a reading of how much work was done.) Cost is pi's own per-message figure, priced from the model's listed rates; a model pi has no rates for contributes zero rather than an estimate. And the context-window percentage is untouched: pi derives it from assistant messages alone, so a delegating session's context doesn't appear to fill up faster. Agents that finish in the background have no tool result of their own to ride on, so their spend is carried by the next one you make — the footer catches up on the following call, not the moment they finish.
756
+ Three things worth knowing about the numbers. Every token component is reported, `cacheRead` included — the cached prefix genuinely is re-read and re-billed on every call, and pi counts it the same way for the session's own messages, so withholding it would make a subagent's rows count differently from every other row in one total. (The extension's *own* token displays still leave it out, which is a different question: there it inflates a reading of how much work was done.) Cost is pi's own per-message figure, priced from the model's listed rates; a model pi has no rates for contributes zero rather than an estimate. Reported subagent token counts do not inflate the parent's context-window percentage. On Pi 0.99.1 the tool-result text itself occupies context and is included in Pi's projection estimate. Agents that finish in the background have no tool result of their own to ride on, so their spend is carried by the next one you make — the footer catches up on the following call, not the moment they finish.
671
757
 
672
758
  **Show cost** (`showCost`, default `false`): whether the subagent surfaces print an estimated cost beside their token counts — the widget (running *and* finished lines), [FleetView](#fleetview), the conversation viewer, foreground results, `get_subagent_result`, and completion notifications:
673
759
 
@@ -696,6 +782,8 @@ Both places report what the run *actually* used, read back from the child sessio
696
782
  ↳ anthropic/claude-haiku-4-5 · thinking: low (asked max) · background
697
783
  ```
698
784
 
785
+ For a virtual model, the displayed identity is the session's selected virtual model, rather than the physical model chosen for each response.
786
+
699
787
  Toggle via `/agents → Settings → Show model`; applied live.
700
788
 
701
789
  **Viewer markdown** (`viewerMarkdown`, default `"assistant"`): how much of the [conversation viewer](#ui)'s transcript is rendered as Markdown rather than shown verbatim.
@@ -747,7 +835,7 @@ EOF
747
835
 
748
836
  Every project now starts with concurrency 16 and grace 10, without ever touching the menu. Individual projects can still override via `/agents` → Settings.
749
837
 
750
- **Failure behavior:** missing file is silent; malformed JSON logs a `[pi-subagents] Ignoring malformed settings at …` warning to stderr; invalid/out-of-range field values are dropped per-field; write failures downgrade the `/agents` toast to a warning with `(session only; failed to persist)`.
838
+ **Failure behavior:** missing file is silent; malformed JSON logs a warning without its contents and suppresses inherited routing; invalid/out-of-range ordinary fields are dropped per-field, while invalid explicit routing fields disable that route. Write failures downgrade the `/agents` toast to a warning with `(session only; failed to persist)`.
751
839
 
752
840
  ## Events
753
841
 
@@ -773,6 +861,8 @@ The four agent-lifecycle events — `subagents:started`, `:completed`, `:failed`
773
861
 
774
862
  `usage` answers the other question — what was billed — and so does include `cacheRead`, because the prefix really is re-read and re-charged on every call. It is a pi `Usage`, the same shape pi puts on `ToolResultEvent` and `AssistantMessage`, so `usage.cost.total` is where a listener already expects the money and anything pi adds to `Usage` arrives without a change here. Neither field derives from the other; `tokens` is a view model, `usage` is the data.
775
863
 
864
+ Completed/failed payloads and persisted `subagents:record` entries also carry `routing` (`source`, `code`, `reason`, optional `model`, supplied `description`, `confidence`, `unpriced`, `guidelinePath`, `guidelineHash`) and optional `routingUsage` (classifier-only Pi `Usage`). The guideline hash is SHA-256 of the original file contents. Coding `tokens` excludes classifier tokens; total cost includes the reported classifier cost once. Settings events omit `jev.TYPESAFE_API_KEY`.
865
+
776
866
  ## Cross-Extension RPC
777
867
 
778
868
  Other pi extensions can spawn and stop subagents programmatically via the `pi.events` event bus, without importing this package directly.
@@ -989,6 +1079,8 @@ src/
989
1079
  # Invocation surface
990
1080
  invocation-config.ts # Shared tool-parameter schemas (isolation, join, thinking, ...)
991
1081
  model-resolver.ts # Model resolution: exact provider/modelId with fuzzy fallback
1082
+ model-routing.ts # Source priority and bounded native Jev classification before startup
1083
+ routing-config.ts # Minimal Jev model/description and credential config validation
992
1084
  enabled-models.ts # Read pi's enabledModels settings (project over global)
993
1085
  model-scope.ts # scopeModels allowlist policy, shared by top-level and nested tools
994
1086
  mention.ts # `@handle message` grammar: suggestion triggers and send parsing
@@ -1024,6 +1116,7 @@ src/
1024
1116
  viewer-keys.ts # Viewer scroll keys resolved through user keybindings
1025
1117
  agent-mention.ts # `@` roster (running, resumable, and startable agents) + popup rows
1026
1118
  schedule-menu.ts # /agents → Scheduled jobs submenu
1119
+ model-routing-menu.ts # Guideline/model descriptions and masked API-key editor
1027
1120
  select-item.ts # Collision-safe ctx.ui.select wrapper (numbered rows)
1028
1121
  workflow-card.ts # Inline workflow card (tool result and session entry)
1029
1122
  workflow-dialog.ts # /agents → Workflows two-pane inspector
@@ -16,7 +16,8 @@
16
16
  import type { Model } from "@earendil-works/pi-ai";
17
17
  import type { AgentSession, ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
18
18
  import { type ToolActivity } from "./agent-runner.js";
19
- import type { AgentInvocation, AgentRecord, AgentTombstone, IsolationMode, MentionResolution, SubagentType, ThinkingLevel } from "./types.js";
19
+ import { type RoutingInput } from "./model-routing.js";
20
+ import type { AgentConfig, AgentInvocation, AgentRecord, AgentTombstone, IsolationMode, MentionResolution, SubagentType, ThinkingLevel } from "./types.js";
20
21
  import { type LifetimeUsage } from "./usage.js";
21
22
  import type { CompiledSchema } from "./workflow/json-schema.js";
22
23
  export type OnAgentComplete = (record: AgentRecord) => void;
@@ -47,6 +48,10 @@ export type CompactionInfo = {
47
48
  */
48
49
  export declare function isTopLevelAgent(record: Pick<AgentRecord, "parentAgentId" | "workflowId">): boolean;
49
50
  interface SpawnOptions {
51
+ /** Internal only; stripped at the programmatic/RPC boundary. */
52
+ routing?: RoutingInput;
53
+ /** Branch-local definition selected by a trusted invocation resolver. */
54
+ agentConfig?: AgentConfig;
50
55
  description: string;
51
56
  /**
52
57
  * Optional memorable name for this instance, becoming a second handle
@@ -219,6 +224,7 @@ interface ResumeOptions {
219
224
  onStarted?: () => void;
220
225
  }
221
226
  export declare class AgentManager {
227
+ private router;
222
228
  private agents;
223
229
  private cleanupInterval;
224
230
  private onComplete?;
@@ -16,9 +16,11 @@
16
16
  import { randomUUID } from "node:crypto";
17
17
  import { statSync } from "node:fs";
18
18
  import { isAbsolute } from "node:path";
19
- import { resumeAgent, runAgent } from "./agent-runner.js";
19
+ import { resolveDefaultModel, resumeAgent, runAgent } from "./agent-runner.js";
20
+ import { getAgentConfig } from "./agent-types.js";
20
21
  import { assignHandle, handleBase } from "./mention.js";
21
22
  import { describeModel } from "./model-resolver.js";
23
+ import { loadRoutingPolicy, ModelRouter } from "./model-routing.js";
22
24
  import { addUsage } from "./usage.js";
23
25
  import { cleanupWorktree, createWorktree, isWorktreeIsolationEnabled, pruneWorktrees, } from "./worktree.js";
24
26
  /**
@@ -157,6 +159,7 @@ async function shutdownChildSession(session) {
157
159
  catch { /* ignore */ }
158
160
  }
159
161
  export class AgentManager {
162
+ router = new ModelRouter();
160
163
  agents = new Map();
161
164
  cleanupInterval;
162
165
  onComplete;
@@ -323,7 +326,7 @@ export class AgentManager {
323
326
  if (record.handle !== undefined && record.alias === undefined && options.name !== undefined) {
324
327
  record.alias = assignHandle(handleBase(options.name), this.takenHandles());
325
328
  }
326
- const args = { pi, ctx, type, prompt, options };
329
+ const args = { pi, ctx, type, prompt, options: { ...options, agentConfig: options.agentConfig ?? getAgentConfig(type) } };
327
330
  const pool = this.poolFor(record);
328
331
  if (pool !== undefined && !options.bypassQueue && !this.poolHasRoom(pool)) {
329
332
  // Queue it — started when a running agent in the same pool completes.
@@ -461,6 +464,7 @@ export class AgentManager {
461
464
  else if (pool === "foreground")
462
465
  this.runningForeground--;
463
466
  };
467
+ const wasQueued = record.startGate !== undefined;
464
468
  record.status = "running";
465
469
  record.startedAt = Date.now();
466
470
  record.startGate = undefined;
@@ -468,6 +472,61 @@ export class AgentManager {
468
472
  this.runningBackground++;
469
473
  else if (pool === "foreground")
470
474
  this.runningForeground++;
475
+ const config = options.agentConfig;
476
+ const provenance = options.routing;
477
+ const explicit = provenance
478
+ ? provenance.modelExplicit || provenance.thinkingExplicit
479
+ : options.model != null || options.thinkingLevel != null;
480
+ const currentPolicy = wasQueued || !provenance?.policy ? loadRoutingPolicy(options.configCwd ?? ctx.cwd) : undefined;
481
+ const policy = currentPolicy ?? provenance.policy;
482
+ record.routing = { mode: policy.mode, source: policy.source, fallbackSource: policy.fallbackSource,
483
+ code: policy.mode === "off" ? "off" : policy.diagnostic ? "guideline_unavailable" : "baseline",
484
+ reason: policy.mode === "off" ? "Model routing and routing guidance are off" : policy.diagnostic ?? "Using the existing model", guidelinePath: policy.guidelinePath, guidelineHash: policy.guidelineHash };
485
+ if (policy.diagnostic)
486
+ console.warn(`[pi-subagents] ${policy.diagnostic}`);
487
+ const route = policy.mode === "jev" || policy.mode === "shadow" ||
488
+ (policy.mode === "auto" && !explicit && !config?.model && !config?.thinking && policy.source === "jev");
489
+ if (!options.resumeSessionFile && provenance?.entrypoint !== "internal" && route) {
490
+ const stop = () => this.abort(id);
491
+ options.signal?.addEventListener("abort", stop, { once: true });
492
+ if (options.signal?.aborted)
493
+ stop();
494
+ try {
495
+ const routed = await this.router.choose(ctx, policy, prompt, options.description ?? type, options.model ?? resolveDefaultModel(ctx.model, ctx.modelRegistry, config?.model), record.abortController.signal, usage => {
496
+ record.routingUsage = usage;
497
+ // Classifier tokens stay separate from coding/context/output budgets.
498
+ const delta = { input: usage.input, output: usage.output, cacheWrite: usage.cacheWrite, cacheRead: usage.cacheRead, cost: usage.cost.total };
499
+ this.onUsage?.(record, delta);
500
+ for (let current = record; current;) {
501
+ current.lifetimeUsage.cost = (current.lifetimeUsage.cost ?? 0) + usage.cost.total;
502
+ current = current.parentAgentId ? this.agents.get(current.parentAgentId) : undefined;
503
+ }
504
+ });
505
+ record.routing = { ...routed.decision, guidelinePath: policy.guidelinePath, guidelineHash: policy.guidelineHash };
506
+ if (routed.model && policy.mode === "shadow") {
507
+ record.routing = { ...record.routing, code: "shadow", model: undefined, suggestedModel: routed.decision.model,
508
+ fallbackSource: policy.source, reason: "Jev suggested a model; shadow mode kept the default-priority model" };
509
+ }
510
+ else if (routed.model) {
511
+ options.model = routed.model;
512
+ }
513
+ }
514
+ catch {
515
+ record.routing.code = "classifier_error";
516
+ record.routing.reason = "Jev could not choose a model; using the existing model";
517
+ }
518
+ finally {
519
+ options.signal?.removeEventListener("abort", stop);
520
+ }
521
+ if (record.status !== "running" || record.abortController.signal.aborted) {
522
+ this.settleRun(record, true, pool);
523
+ return;
524
+ }
525
+ }
526
+ else if (policy.mode === "auto" && (explicit || config?.model || config?.thinking)) {
527
+ record.routing.code = "explicit";
528
+ record.routing.reason = "Explicit model or thinking; automatic routing skipped";
529
+ }
471
530
  // Worktree isolation: try to create a temporary git worktree. Strict —
472
531
  // fail loud if not possible (no silent fallback to main tree). Done BEFORE
473
532
  // the run is kicked off so a failure doesn't leave a half-running agent.
@@ -522,6 +581,7 @@ export class AgentManager {
522
581
  const promise = runAgent(ctx, type, prompt, {
523
582
  pi,
524
583
  agentId: id,
584
+ agentConfig: options.agentConfig,
525
585
  model: options.model,
526
586
  maxTurns: options.maxTurns,
527
587
  isolated: options.isolated,
@@ -1312,6 +1372,7 @@ export class AgentManager {
1312
1372
  */
1313
1373
  async dispose(pi) {
1314
1374
  clearInterval(this.cleanupInterval);
1375
+ this.abortAll();
1315
1376
  // Clear queue — via dequeue, so anyone blocked in spawnAndWait is woken
1316
1377
  // rather than left awaiting a gate nothing will ever resolve.
1317
1378
  this.dequeue(() => true);
@@ -5,7 +5,7 @@ import type { Model } from "@earendil-works/pi-ai";
5
5
  import type { ExtensionContext } from "@earendil-works/pi-coding-agent";
6
6
  import { type AgentSession, DefaultResourceLoader, type ExtensionAPI } from "@earendil-works/pi-coding-agent";
7
7
  import { type NestedAgentManager } from "./nested-tools.js";
8
- import type { ServiceTier, SubagentType, ThinkingLevel } from "./types.js";
8
+ import type { AgentConfig, ServiceTier, SubagentType, ThinkingLevel } from "./types.js";
9
9
  import type { LifetimeUsage } from "./usage.js";
10
10
  import type { CompiledSchema } from "./workflow/json-schema.js";
11
11
  /**
@@ -85,12 +85,13 @@ export declare function parseExtSelectors(entries: string[]): {
85
85
  * snapshotted. `registerTool` writes into the very `extension.tools` maps this reads,
86
86
  * so `inScope()` sees late arrivals on the next call.
87
87
  *
88
- * Two enforcement points, because neither covers the whole picture:
88
+ * The active set and call-time checks cover different parts of scope:
89
89
  *
90
90
  * - `turn_end` re-narrows the ACTIVE set. pi emits `turn_end` immediately before
91
91
  * `prepareNextTurn` re-snapshots `agent.state.tools`, and session listeners run
92
92
  * synchronously, so the narrow lands in time for turns 2..N.
93
- * - `beforeToolCall` blocks out-of-scope calls. Turn 1 cannot be narrowed at all:
93
+ * - The loader's bound `tool_call` handler blocks direct and nested calls.
94
+ * Turn 1 cannot be narrowed at all:
94
95
  * `before_agent_start` fires INSIDE `prompt()` and may widen the tool set, but
95
96
  * `createContextSnapshot()` freezes that turn's tools immediately after — there
96
97
  * is no hook in between. A call-time check is the only correct guard there.
@@ -117,7 +118,7 @@ export declare function installExtensionToolScope(session: AgentSession, ctx: {
117
118
  * seeded from `toolNames`.
118
119
  */
119
120
  readmitToolNames: Set<string>;
120
- }): void;
121
+ }): (toolName: string) => boolean;
121
122
  /** Normalize max turns. undefined or 0 = unlimited, otherwise minimum 1. */
122
123
  export declare function normalizeMaxTurns(n: number | undefined): number | undefined;
123
124
  /** Get the default max turns value. undefined = unlimited. */
@@ -156,6 +157,8 @@ export interface ToolActivity {
156
157
  toolName: string;
157
158
  }
158
159
  export interface RunOptions {
160
+ /** Snapshot of the selected definition for this branch. */
161
+ agentConfig?: AgentConfig;
159
162
  /** ExtensionAPI instance — used for pi.exec() instead of execSync. */
160
163
  pi: ExtensionAPI;
161
164
  /** Manager-assigned id; suffixes session name to disambiguate parallel spawns (e.g. `Explore#a1b2c3d4`). */
@@ -68,6 +68,10 @@ export function installServiceTierPayload(session, serviceTier) {
68
68
  * single-file extensions to the basename minus `.ts`/`.js`.
69
69
  */
70
70
  export function extensionCanonicalName(extPath) {
71
+ if (extPath.startsWith("builtin:"))
72
+ return extPath.slice("builtin:".length).toLowerCase();
73
+ if (extPath.startsWith("<inline:") && extPath.endsWith(">"))
74
+ return extPath.slice(8, -1).toLowerCase();
71
75
  const base = basename(extPath);
72
76
  const name = base === "index.ts" || base === "index.js"
73
77
  ? basename(dirname(extPath))
@@ -132,6 +136,8 @@ function extensionPackageName(extPath) {
132
136
  */
133
137
  export function extensionCanonicalNames(extPath) {
134
138
  const canonical = extensionCanonicalName(extPath);
139
+ if (extPath.startsWith("builtin:") || extPath.startsWith("<inline:"))
140
+ return [canonical];
135
141
  const pkg = extensionPackageName(extPath);
136
142
  return pkg && pkg !== canonical ? [canonical, pkg] : [canonical];
137
143
  }
@@ -219,12 +225,13 @@ export function parseExtSelectors(entries) {
219
225
  * snapshotted. `registerTool` writes into the very `extension.tools` maps this reads,
220
226
  * so `inScope()` sees late arrivals on the next call.
221
227
  *
222
- * Two enforcement points, because neither covers the whole picture:
228
+ * The active set and call-time checks cover different parts of scope:
223
229
  *
224
230
  * - `turn_end` re-narrows the ACTIVE set. pi emits `turn_end` immediately before
225
231
  * `prepareNextTurn` re-snapshots `agent.state.tools`, and session listeners run
226
232
  * synchronously, so the narrow lands in time for turns 2..N.
227
- * - `beforeToolCall` blocks out-of-scope calls. Turn 1 cannot be narrowed at all:
233
+ * - The loader's bound `tool_call` handler blocks direct and nested calls.
234
+ * Turn 1 cannot be narrowed at all:
228
235
  * `before_agent_start` fires INSIDE `prompt()` and may widen the tool set, but
229
236
  * `createContextSnapshot()` freezes that turn's tools immediately after — there
230
237
  * is no hook in between. A call-time check is the only correct guard there.
@@ -263,7 +270,7 @@ export function installExtensionToolScope(session, ctx) {
263
270
  for (const name of EXCLUDED_TOOL_NAMES)
264
271
  keep.delete(name);
265
272
  // Injected tools are legitimately active for this agent — re-admit them so
266
- // the renarrow keeps them in the active set and beforeToolCall doesn't
273
+ // the renarrow keeps them in the active set and the call-time guard doesn't
267
274
  // block them. Already vetted against `disallowed_tools` by the caller,
268
275
  // which is the only place that knows which kind may be taken back.
269
276
  for (const name of readmitToolNames)
@@ -272,8 +279,10 @@ export function installExtensionToolScope(session, ctx) {
272
279
  };
273
280
  const renarrow = () => {
274
281
  const allowed = inScope();
275
- const next = session.getAllTools().map((t) => t.name).filter((n) => allowed.has(n));
276
282
  const current = session.getActiveToolNames();
283
+ // Keep deferred/codemode tools callable without promoting them into model
284
+ // declarations. Explicit activation by an extension is preserved in scope.
285
+ const next = session.getAllTools().filter((tool) => allowed.has(tool.name) && (tool.exposure === "direct" || tool.exposure === "model-only" || current.includes(tool.name))).map((tool) => tool.name);
277
286
  // setActiveToolsByName unconditionally rebuilds the system prompt, so skip
278
287
  // the no-op that steady-state turns would otherwise pay for every turn.
279
288
  if (next.length !== current.length || next.some((n, i) => n !== current[i])) {
@@ -287,16 +296,7 @@ export function installExtensionToolScope(session, ctx) {
287
296
  if (event.type === "turn_end")
288
297
  renarrow();
289
298
  });
290
- const priorBeforeToolCall = session.agent.beforeToolCall;
291
- session.agent.beforeToolCall = async (context, signal) => {
292
- if (!inScope().has(context.toolCall.name)) {
293
- return {
294
- block: true,
295
- reason: `Tool "${context.toolCall.name}" is not available to this subagent.`,
296
- };
297
- }
298
- return priorBeforeToolCall?.(context, signal);
299
- };
299
+ return (toolName) => inScope().has(toolName);
300
300
  }
301
301
  /** Default max turns. undefined = unlimited (no turn limit). */
302
302
  let defaultMaxTurns;
@@ -450,8 +450,8 @@ function resolveConfiguredSessionDir(sessionDir, cwd) {
450
450
  return resolve(cwd, sessionDir);
451
451
  }
452
452
  export async function runAgent(ctx, type, prompt, options) {
453
- const config = getConfig(type);
454
- const agentConfig = getAgentConfig(type);
453
+ const agentConfig = options.agentConfig ?? getAgentConfig(type);
454
+ const config = agentConfig ?? getConfig(type);
455
455
  // Resolve working directory: worktree override > parent cwd
456
456
  const effectiveCwd = options.cwd ?? ctx.cwd;
457
457
  // Filesystem work happens in effectiveCwd; config discovery in configCwd.
@@ -479,7 +479,7 @@ export async function runAgent(ctx, type, prompt, options) {
479
479
  extras.skillBlocks = loaded;
480
480
  }
481
481
  }
482
- let toolNames = getToolNamesForType(type);
482
+ let toolNames = options.agentConfig ? options.agentConfig.builtinToolNames ?? [...BUILTIN_TOOL_NAMES] : getToolNamesForType(type);
483
483
  // Persistent memory: detect write capability and branch accordingly.
484
484
  // Account for disallowedTools — a tool in the base set but on the denylist is not truly available.
485
485
  if (agentConfig?.memory) {
@@ -557,6 +557,7 @@ export async function runAgent(ctx, type, prompt, options) {
557
557
  // must compare against this, not the surviving set (absence from survivors is
558
558
  // an exclude *succeeding*).
559
559
  let discoveredNames;
560
+ let toolInScope;
560
561
  const extensionsOverride = noExtensions || (loadAll && !hasExcludes)
561
562
  ? undefined
562
563
  : (base) => {
@@ -564,6 +565,8 @@ export async function runAgent(ctx, type, prompt, options) {
564
565
  return {
565
566
  ...base,
566
567
  extensions: base.extensions.filter((e) => {
568
+ if (e.path === "<inline:subagent-tool-scope>")
569
+ return true;
567
570
  const canons = extensionCanonicalNames(e.path);
568
571
  if (canons.some((n) => excludeNames.has(n)))
569
572
  return false; // exclude wins
@@ -576,6 +579,19 @@ export async function runAgent(ctx, type, prompt, options) {
576
579
  agentDir,
577
580
  noExtensions,
578
581
  additionalExtensionPaths,
582
+ // Pi's nested ctx.executeTool calls dispatch tool_call directly, and prompt
583
+ // setup can replace agent.beforeToolCall. Enforce scope in that dispatch too.
584
+ extensionFactories: noExtensions ? [] : [{
585
+ name: "subagent-tool-scope",
586
+ hidden: true,
587
+ factory: (pi) => {
588
+ pi.on("tool_call", (event) => {
589
+ if (toolInScope && !toolInScope(event.toolName)) {
590
+ return { block: true, reason: `Tool "${event.toolName}" is not available to this subagent.` };
591
+ }
592
+ });
593
+ },
594
+ }],
579
595
  extensionsOverride,
580
596
  noSkills,
581
597
  noPromptTemplates: true,
@@ -786,21 +802,14 @@ export async function runAgent(ctx, type, prompt, options) {
786
802
  parentSession: ctx.sessionManager?.getSessionFile?.(),
787
803
  })
788
804
  : SessionManager.inMemory(effectiveCwd);
789
- // Pi 0.80.8 replaced createAgentSession's modelRegistry option with
790
- // modelRuntime, but ExtensionContext still exposes only the registry facade.
791
- // Pass both so the full supported Pi range retains the parent's providers.
805
+ // ExtensionContext exposes the registry facade; its runtime owns provider auth.
792
806
  const parentModelRuntime = ctx.modelRegistry.runtime;
793
807
  const sessionOpts = {
794
808
  cwd: effectiveCwd,
795
809
  agentDir,
796
810
  sessionManager,
797
811
  settingsManager,
798
- modelRegistry: ctx.modelRegistry,
799
- // `as never` is what keeps this assignable across the supported Pi range:
800
- // pre-0.80.8 the field exists only via the `modelRuntime?: unknown` shim
801
- // above, while newer Pi types it as `ModelRuntime` — a shape an opaque
802
- // `unknown` read off the private facade field can never satisfy.
803
- ...(parentModelRuntime !== undefined && { modelRuntime: parentModelRuntime }),
812
+ modelRuntime: parentModelRuntime,
804
813
  model,
805
814
  tools: sessionTools,
806
815
  customTools: [...nestedTools, ...structuredTools],
@@ -835,7 +844,7 @@ export async function runAgent(ctx, type, prompt, options) {
835
844
  // handled below by re-deriving scope from the loader's live extension maps —
836
845
  // `registerTool` writes into those same maps, so late arrivals are judged too.
837
846
  if (!noExtensions) {
838
- installExtensionToolScope(session, {
847
+ toolInScope = installExtensionToolScope(session, {
839
848
  loader,
840
849
  toolNames,
841
850
  disallowedSet,