@diousk/pi-subagents-fast 0.21.0 → 0.23.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.23.0] - 2026-10-02
11
+
12
+ ### Changed
13
+ - **BREAKING: Jev model profiles require `thinkingLevel`.** Add `minimal`, `low`, `medium`, `high`, `xhigh` or `max` to every `jev.models` entry. Successful routing applies the model and thinking level together; shadow and fallback preserve the original pair. The routing menu requires a level, and invalid or missing levels disable the Jev block with a diagnostic.
14
+
15
+ ## [0.22.0] - 2026-10-02
16
+
17
+ > **Breaking: requires Pi 1.0.0 or newer.** Update Pi before using this version.
18
+
19
+ ### Added
20
+ - **Optional Jev routing with four modes.** `routingMode` defaults to `auto`: custom agents, then a Markdown guideline, then Jev. Choose `shadow` to record suggestions without changing models, `jev` for Jev-first selection with default-priority fallback, or `off` to disable routing guidance and classification. Configure the mode, guideline, model descriptions and optional TypeSafe credentials in `subagents.json` or `/agents → Model routing`; usage is recorded separately, including shadow calls. Jev choices honor Pi's model scope and, when enabled, the extension's `scopeModels` policy.
21
+
22
+ ### Changed
23
+ - **Service-tier compatibility is documented for GPT-6.1 Sol and GPT-6 Luna.** Native transport regression tests cover `fast` and `priority` on initial and resumed turns. Pi 1.0.0's Codex cost estimate still understates responses marked `fast`; this release documents that upstream limitation without recalculating costs.
24
+
10
25
  ## [0.21.0] - 2026-09-30
11
26
 
12
27
  > **Breaking: requires Pi 0.99.1 or newer.** Update Pi before installing this version. Legacy transcript and session API compatibility paths have been removed.
package/README.md CHANGED
@@ -28,6 +28,7 @@ https://github.com/user-attachments/assets/8685261b-9338-4fea-8dfe-1c590d5df543
28
28
  - **Graceful turn limits** — agents get a "wrap up" warning before hard abort, producing clean partial results instead of cut-off output
29
29
  - **Case-insensitive agent types** — `"explore"`, `"Explore"`, `"EXPLORE"` all work. A type that doesn't resolve to exactly one *enabled* agent — unknown, disabled, or ambiguous between two agents differing only by case — falls back to general-purpose with a note, or is refused outright under [`fallbackSubagent: none`](#persistent-settings)
30
30
  - **Fuzzy model selection** — specify models by name (`"haiku"`, `"sonnet"`) instead of full IDs, with automatic filtering to only available/configured models
31
+ - **Model routing** — custom agents first, then a user-supplied Markdown guideline, then optional Jev model selection. Configure it through `/agents → Model routing` or two ordinary JSON fields; see [Model routing](#model-routing)
31
32
  - **Context inheritance** — optionally fork the parent conversation into a sub-agent so it knows what's been discussed
32
33
  - **Persistent agent memory** — three scopes (project, local, user) with automatic read-only fallback for agents without write tools
33
34
  - **Git worktree isolation** — run agents in isolated repo copies; changes auto-committed to branches on completion
@@ -76,9 +77,9 @@ npm pack --dry-run
76
77
 
77
78
  The `prepublishOnly` script runs lint, typecheck, tests, and the build before npm uploads the package. After publishing, install it with `pi install npm:@diousk/pi-subagents-fast`.
78
79
 
79
- Requires **Pi 0.99.1 or newer** and **Node.js 22.19.0 or newer**. Update Pi before installing this extension. The `peerDependencies` range declares the minimum, so npm flags an older Pi at install time.
80
+ Requires **Pi 1.0.0 or newer** and **Node.js 22.19.0 or newer**. Update Pi before installing this extension. The `peerDependencies` range declares the minimum, so npm flags an older Pi at install time.
80
81
 
81
- The development and CI baseline is **Pi 0.99.1**. Use that version or newer for `openai/gpt-6.1-sol` (API key) or `openai-codex/gpt-6.1-sol` (your Pi Codex login). Child sessions reuse the parent's configured providers and authentication; model limits and pricing come from Pi's provider catalog. See the [Pi 0.99.1 release](https://pi.dev/changelog/releases/0.99.1).
82
+ Development and CI use **Pi 1.0.0**. Jev uses Pi's native classifier API. Child sessions reuse the parent's configured providers and authentication; model availability, limits and pricing come from Pi's catalog.
82
83
 
83
84
  ### Other hosts
84
85
 
@@ -354,7 +355,9 @@ All fields are optional — sensible defaults for everything.
354
355
 
355
356
  For an OpenAI Responses or Codex agent, set `service_tier: fast` to request fast processing; `priority` remains a supported alias. Availability depends on the provider and account. The UI shows the requested tier only when the effective model uses one of those APIs. See [OpenAI fast mode](https://developers.openai.com/api/docs/guides/fast-mode).
356
357
 
357
- GPT-6.1 Sol supports `low`, `medium`, `high`, `xhigh`, and `max` reasoning. In Pi 0.99.1, `minimal` maps to provider effort `low`; `off` is clamped to that same alias because Sol cannot disable reasoning. The UI reports Pi's logical thinking level. Prefer explicit supported levels in agent files:
358
+ The extension forwards both tiers for `gpt-6.1-sol` and `gpt-6-luna` through Pi's OpenAI Responses and Codex transports. Pi 1.0.0's Codex adapter currently estimates a response marked `fast` at the standard rate; forwarding the tier works, but its displayed cost can be understated. This extension reports Pi's cost estimate without recalculating it. OpenAI documents Fast mode as unavailable for these models with EU data residency.
359
+
360
+ GPT-6.1 Sol supports `low`, `medium`, `high`, `xhigh`, and `max` reasoning. In Pi 1.0.0, `minimal` maps to provider effort `low`; `off` is clamped to that same alias because Sol cannot disable reasoning. The UI reports Pi's logical thinking level. Prefer explicit supported levels in agent files:
358
361
 
359
362
  ```yaml
360
363
  model: openai-codex/gpt-6.1-sol
@@ -629,6 +632,72 @@ When background agents complete, they notify the main agent. The **join mode** c
629
632
  **Configuration:**
630
633
  - Configure join mode in `/agents` → Settings → Join mode
631
634
 
635
+ ## Model routing
636
+
637
+ Open `/agents → Model routing` to choose a mode and configure a guideline path, models with descriptions and required thinking levels, or a masked TypeSafe API key. The menu shows the active mode and source and saves project settings. For machine-wide defaults, edit `~/.pi/agent/subagents.json`; project overrides go in `.pi/subagents.json`. `PI_CODING_AGENT_DIR` changes the global directory along with Pi's other configuration.
638
+
639
+ Set one top-level field to control routing; omitting it means `auto`:
640
+
641
+ ```json
642
+ { "routingMode": "auto" }
643
+ ```
644
+
645
+ | `routingMode` | Behavior |
646
+ |---|---|
647
+ | `auto` (default) | Use the priority table below |
648
+ | `shadow` | Ask Jev and record its suggestion, but keep the model chosen from agent definitions, the guideline or the existing model. Jev requests can incur charges |
649
+ | `jev` | Jev chooses first for every fresh task, even with custom agents, a guideline, explicit model parameters or agent-file model pins. Invalid configuration, unavailable credentials, low confidence or errors keep the default-priority choice |
650
+ | `off` | No custom-agent routing selection guidance, guideline injection or Jev requests. Agents remain callable and existing model/thinking settings still apply |
651
+
652
+ Mode changes apply to subsequent fresh launches. The main agent's routing guidance and tool description refresh before its next turn, so switching to `off` removes previously injected routing instructions.
653
+
654
+ The default priority is:
655
+
656
+ | Priority | Configuration | Who chooses |
657
+ |---|---|---|
658
+ | 1 | At least one enabled custom agent | The main agent chooses an agent from its description. Guideline and Jev are inactive, even if it chooses a built-in agent |
659
+ | 2 | `customGuideline` points to a Markdown file | The main agent reads the guideline and passes its model/thinking choice explicitly |
660
+ | 3 | `jev` is configured | Jev chooses a model from your descriptions for a fresh delegated task |
661
+ | 4 | None of the above | The existing model is used |
662
+
663
+ **Custom agents:** keep using `~/.pi/agent/agents/<name>.md`, `.pi/agents/<name>.md` or `.agents/agents/<name>.md`. No routing setting is needed. Built-in and disabled agents do not activate priority 1. Agent-file model/thinking pins supply the default choice over `Agent` parameters; a confident Jev choice in `jev` mode can replace both model and thinking level.
664
+
665
+ **Custom guideline:** write your routing rules in `~/.pi/agent/agents/custom-route.md`, then configure:
666
+
667
+ ```json
668
+ {
669
+ "customGuideline": "~/.pi/agent/agents/custom-route.md"
670
+ }
671
+ ```
672
+
673
+ For example, the Markdown can say “Use anthropic/claude-haiku-4-5 for simple edits; use openai-codex/gpt-6.1-sol for debugging concurrency.” The guideline reaches the main agent in full, compact and custom description modes, and is refreshed before each turn, except in `off` mode. Relative paths are resolved from the settings file's directory: `agents/custom-route.md` in `.pi/subagents.json` means `.pi/agents/custom-route.md`. `custom-route.md` and the configured guideline file are excluded from agent discovery. Merely placing a guideline file in the directory does not enable it. A configured missing, empty or oversized guideline produces a diagnostic and keeps the existing model under `auto`; it does not prevent Jev requests under `jev` or `shadow`.
674
+
675
+ **Jev:** provide descriptions and exact model IDs available through your Pi login:
676
+
677
+ ```json
678
+ {
679
+ "jev": {
680
+ "TYPESAFE_API_KEY": "your-typesafe-key",
681
+ "models": [
682
+ { "model": "anthropic/claude-haiku-4-5", "description": "Simple edits, lookup and concise summaries", "thinkingLevel": "low" },
683
+ { "model": "openai-codex/gpt-6.1-sol", "description": "Complex debugging, architecture and concurrency", "thinkingLevel": "high" }
684
+ ]
685
+ }
686
+ }
687
+ ```
688
+
689
+ `TYPESAFE_API_KEY` is optional when Pi already has TypeSafe credentials or the environment variable is set. A literal key must be a nonempty token without whitespace or control characters and applies only to that classifier request; it is saved in the settings file, masked in the menu and omitted from settings events, prompts and routing records. Without a literal key, Pi's native credential availability check runs before classification. Rejected credentials fall back without retrying. The extension does not change environment variables or register a routing provider. `models` accepts 1–254 unique entries with descriptions of 1–4000 characters. Every entry requires `thinkingLevel`: `minimal`, `low`, `medium`, `high`, `xhigh` or `max`. Existing profiles must add this field; a missing or invalid value disables the whole Jev block and uses the default-priority fallback. Jev entries use exact `provider/model-id` spelling; fuzzy names remain available for explicit `Agent` parameters.
690
+
691
+ Under `auto`, supplying **either** `model` or `thinking` explicitly skips Jev. An agent-file pin also skips it. In workflows, either `model` or `effort` skips it, and workflow options retain their precedence over agent-file defaults. Inherited models remain eligible for Jev. Under `jev`, these choices supply the fallback model, but do not skip classification. Under `shadow`, they remain the actual choice while Jev records a comparison. A selected Jev profile applies both its model and `thinkingLevel`. Pi clamps the requested thinking level to what that model supports. Shadow and fallback preserve the original model and thinking level.
692
+
693
+ Automatic choices are limited to authenticated models and the current nonempty Pi model scope, regardless of the `scopeModels` setting. When `scopeModels` is enabled, Jev candidates also respect `enabledModels`. Unavailable candidates are excluded and the selected model is checked again after classification, including any changed scope. Jev failure, confidence below 0.6 or a two-second timeout keeps the default-priority model already selected by the main agent, explicit caller or agent definition, otherwise the existing model. This fallback does not make a second Jev request. Cancellation stops startup without launching a fallback. Queued agents read the current mode and classify after dequeue; nested calls share a separate classifier concurrency limit of four and do not take another agent slot.
694
+
695
+ Jev applies to fresh tool, nested, workflow, scheduled, RPC and direct-mention spawns. Schedules read current configuration when they fire. Direct calls, schedules and RPC do not create an extra main-agent turn to interpret a custom guideline: with priorities 1 or 2 under `auto`, or when Jev falls back under `jev`, they use an explicit choice, agent definition or the existing model. Resumes and internal agent-file generation do not call Jev in any mode.
696
+
697
+ Project `routingMode` and `customGuideline` replace the global values; omission inherits them, and `customGuideline: false` disables the guideline. A project `jev` block replaces the **whole** global block, including credentials; omitted fields inside that block do not inherit. `"jev": false` disables global Jev. An invalid mode or malformed settings file disables routing. An invalid explicit Jev block disables that block instead of restoring global paid routing. Other settings changes preserve these inheritance rules.
698
+
699
+ Classifier usage is recorded separately as `routingUsage`, including in `shadow`; it does not consume coding context or workflow output budgets. Its reported cost is added once to the agent and ancestor cost totals, and `reportUsage` can return its token usage to the parent. `routing` records the mode, source, reason, applied `model`, requested `thinkingLevel` and confidence; shadow records `suggestedModel` and `suggestedThinkingLevel` and leaves both applied fields unset. `fallbackSource` identifies the default source when a suggestion is observed or Jev falls back. Shadow results show “Jev shadow”. Zero catalog prices are marked `unpriced`; the Agent result shows “Jev price unavailable” instead of treating that as proof the classifier is free.
700
+
632
701
  ## Model Scope
633
702
 
634
703
  **Opt-in:** off by default. Enable via `/agents → Settings → Scope models`.
@@ -641,6 +710,7 @@ When on, each subagent spawn's effective model is validated against pi's own `en
641
710
  |---|---|
642
711
  | Caller-supplied via `Agent({ model: "..." })` | Hard error returned to the orchestrator, listing allowed models |
643
712
  | Caller-supplied via cross-extension RPC (`subagents:rpc:spawn`, e.g. pi-tasks `TaskExecute`) | Hard error returned to the calling extension, listing allowed models |
713
+ | Jev-selected | Excluded from candidates; a scope change during classification keeps the default-priority model |
644
714
  | Pinned in agent frontmatter | Warning toast + the pinned model runs (frontmatter is authoritative) |
645
715
  | Parent-inherited (neither set) | Warning toast + parent's model runs |
646
716
 
@@ -661,6 +731,8 @@ Runtime tuning values set via `/agents` → Settings (max concurrency, max foreg
661
731
 
662
732
  **Precedence:** project overrides global on any field present in both. Missing fields fall back to the hardcoded defaults (max concurrency `10`, max foreground concurrency `0` = unlimited, default max turns unlimited, grace turns `5`, nested depth `2`, join mode `smart`, defaults enabled).
663
733
 
734
+ Routing settings: `routingMode` (`auto | shadow | jev | off`, default `auto`), `customGuideline` (`string | false`, default unset) and `jev` (`{ TYPESAFE_API_KEY?: string, models: { model, description, thinkingLevel }[] } | false`, default unset). See [Model routing](#model-routing) for examples and the whole-block override rule.
735
+
664
736
  **Nested depth** (`maxSubagentDepth`, default `2`): the hard ceiling on [nested delegation](#nested-subagents), counted from the main session (main = 0, its subagents = 1). `0` or `1` disables nesting project-wide regardless of any agent's `allowed_subagents`. Read when a subagent session is built, so a change applies to agents started after it.
665
737
 
666
738
  **Fallback agent** (`fallbackSubagent`, default `general-purpose`): the agent used when a caller-supplied `subagent_type` doesn't resolve to exactly one enabled agent — unknown, disabled, or ambiguous because two agents differ only by case. Name any enabled agent to route those calls there instead, or set `none` for **strict**, fail-closed dispatch: the call is refused with an error listing the available types, and nothing spawns. Strict mode matters most for background and scheduled calls, which would otherwise start executing a substituted agent before the caller learns anything. Also settable from `/agents → Settings → Fallback agent`. The boolean `false` is accepted as a spelling of `none`, because it would otherwise be dropped as the wrong type and silently leave the permissive default in place. Every other value is read as an agent name, so a mistaken `off` fails loudly at dispatch rather than meaning one thing in the settings file and another in the resolver. A fallback agent that is itself unknown or disabled is a misconfiguration and is reported rather than quietly replaced. Note the default is unchanged and stays permissive by design: with `disableDefaultAgents` and no `general-purpose` of your own, an unresolvable type still resolves to a built-in config carrying *all* tools — set `none` (or name one of your own agents) to close that.
@@ -763,7 +835,7 @@ EOF
763
835
 
764
836
  Every project now starts with concurrency 16 and grace 10, without ever touching the menu. Individual projects can still override via `/agents` → Settings.
765
837
 
766
- **Failure behavior:** missing file is silent; malformed JSON logs a `[pi-subagents] Ignoring malformed settings at …` warning to stderr; invalid/out-of-range field values are dropped per-field; write failures downgrade the `/agents` toast to a warning with `(session only; failed to persist)`.
838
+ **Failure behavior:** missing file is silent; malformed JSON logs a warning without its contents and suppresses inherited routing; invalid/out-of-range ordinary fields are dropped per-field, while invalid explicit routing fields disable that route. Write failures downgrade the `/agents` toast to a warning with `(session only; failed to persist)`.
767
839
 
768
840
  ## Events
769
841
 
@@ -789,6 +861,8 @@ The four agent-lifecycle events — `subagents:started`, `:completed`, `:failed`
789
861
 
790
862
  `usage` answers the other question — what was billed — and so does include `cacheRead`, because the prefix really is re-read and re-charged on every call. It is a pi `Usage`, the same shape pi puts on `ToolResultEvent` and `AssistantMessage`, so `usage.cost.total` is where a listener already expects the money and anything pi adds to `Usage` arrives without a change here. Neither field derives from the other; `tokens` is a view model, `usage` is the data.
791
863
 
864
+ Completed/failed payloads and persisted `subagents:record` entries also carry `routing` (`source`, `code`, `reason`, optional `model`, `thinkingLevel`, `suggestedModel`, `suggestedThinkingLevel`, supplied `description`, `confidence`, `unpriced`, `guidelinePath`, `guidelineHash`) and optional `routingUsage` (classifier-only Pi `Usage`). The guideline hash is SHA-256 of the original file contents. Coding `tokens` excludes classifier tokens; total cost includes the reported classifier cost once. Settings events omit `jev.TYPESAFE_API_KEY`.
865
+
792
866
  ## Cross-Extension RPC
793
867
 
794
868
  Other pi extensions can spawn and stop subagents programmatically via the `pi.events` event bus, without importing this package directly.
@@ -1005,6 +1079,8 @@ src/
1005
1079
  # Invocation surface
1006
1080
  invocation-config.ts # Shared tool-parameter schemas (isolation, join, thinking, ...)
1007
1081
  model-resolver.ts # Model resolution: exact provider/modelId with fuzzy fallback
1082
+ model-routing.ts # Source priority and bounded native Jev classification before startup
1083
+ routing-config.ts # Jev model/description/thinking-level and credential config validation
1008
1084
  enabled-models.ts # Read pi's enabledModels settings (project over global)
1009
1085
  model-scope.ts # scopeModels allowlist policy, shared by top-level and nested tools
1010
1086
  mention.ts # `@handle message` grammar: suggestion triggers and send parsing
@@ -1040,6 +1116,7 @@ src/
1040
1116
  viewer-keys.ts # Viewer scroll keys resolved through user keybindings
1041
1117
  agent-mention.ts # `@` roster (running, resumable, and startable agents) + popup rows
1042
1118
  schedule-menu.ts # /agents → Scheduled jobs submenu
1119
+ model-routing-menu.ts # Guideline/model descriptions and masked API-key editor
1043
1120
  select-item.ts # Collision-safe ctx.ui.select wrapper (numbered rows)
1044
1121
  workflow-card.ts # Inline workflow card (tool result and session entry)
1045
1122
  workflow-dialog.ts # /agents → Workflows two-pane inspector
@@ -16,7 +16,8 @@
16
16
  import type { Model } from "@earendil-works/pi-ai";
17
17
  import type { AgentSession, ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
18
18
  import { type ToolActivity } from "./agent-runner.js";
19
- import type { AgentInvocation, AgentRecord, AgentTombstone, IsolationMode, MentionResolution, SubagentType, ThinkingLevel } from "./types.js";
19
+ import { type RoutingInput } from "./model-routing.js";
20
+ import type { AgentConfig, AgentInvocation, AgentRecord, AgentTombstone, IsolationMode, MentionResolution, SubagentType, ThinkingLevel } from "./types.js";
20
21
  import { type LifetimeUsage } from "./usage.js";
21
22
  import type { CompiledSchema } from "./workflow/json-schema.js";
22
23
  export type OnAgentComplete = (record: AgentRecord) => void;
@@ -47,6 +48,10 @@ export type CompactionInfo = {
47
48
  */
48
49
  export declare function isTopLevelAgent(record: Pick<AgentRecord, "parentAgentId" | "workflowId">): boolean;
49
50
  interface SpawnOptions {
51
+ /** Internal only; stripped at the programmatic/RPC boundary. */
52
+ routing?: RoutingInput;
53
+ /** Branch-local definition selected by a trusted invocation resolver. */
54
+ agentConfig?: AgentConfig;
50
55
  description: string;
51
56
  /**
52
57
  * Optional memorable name for this instance, becoming a second handle
@@ -219,6 +224,7 @@ interface ResumeOptions {
219
224
  onStarted?: () => void;
220
225
  }
221
226
  export declare class AgentManager {
227
+ private router;
222
228
  private agents;
223
229
  private cleanupInterval;
224
230
  private onComplete?;
@@ -16,9 +16,11 @@
16
16
  import { randomUUID } from "node:crypto";
17
17
  import { statSync } from "node:fs";
18
18
  import { isAbsolute } from "node:path";
19
- import { resumeAgent, runAgent } from "./agent-runner.js";
19
+ import { resolveDefaultModel, resumeAgent, runAgent } from "./agent-runner.js";
20
+ import { getAgentConfig } from "./agent-types.js";
20
21
  import { assignHandle, handleBase } from "./mention.js";
21
22
  import { describeModel } from "./model-resolver.js";
23
+ import { loadRoutingPolicy, ModelRouter } from "./model-routing.js";
22
24
  import { addUsage } from "./usage.js";
23
25
  import { cleanupWorktree, createWorktree, isWorktreeIsolationEnabled, pruneWorktrees, } from "./worktree.js";
24
26
  /**
@@ -157,6 +159,7 @@ async function shutdownChildSession(session) {
157
159
  catch { /* ignore */ }
158
160
  }
159
161
  export class AgentManager {
162
+ router = new ModelRouter();
160
163
  agents = new Map();
161
164
  cleanupInterval;
162
165
  onComplete;
@@ -323,7 +326,7 @@ export class AgentManager {
323
326
  if (record.handle !== undefined && record.alias === undefined && options.name !== undefined) {
324
327
  record.alias = assignHandle(handleBase(options.name), this.takenHandles());
325
328
  }
326
- const args = { pi, ctx, type, prompt, options };
329
+ const args = { pi, ctx, type, prompt, options: { ...options, agentConfig: options.agentConfig ?? getAgentConfig(type) } };
327
330
  const pool = this.poolFor(record);
328
331
  if (pool !== undefined && !options.bypassQueue && !this.poolHasRoom(pool)) {
329
332
  // Queue it — started when a running agent in the same pool completes.
@@ -461,6 +464,7 @@ export class AgentManager {
461
464
  else if (pool === "foreground")
462
465
  this.runningForeground--;
463
466
  };
467
+ const wasQueued = record.startGate !== undefined;
464
468
  record.status = "running";
465
469
  record.startedAt = Date.now();
466
470
  record.startGate = undefined;
@@ -468,6 +472,63 @@ export class AgentManager {
468
472
  this.runningBackground++;
469
473
  else if (pool === "foreground")
470
474
  this.runningForeground++;
475
+ const config = options.agentConfig;
476
+ const provenance = options.routing;
477
+ const explicit = provenance
478
+ ? provenance.modelExplicit || provenance.thinkingExplicit
479
+ : options.model != null || options.thinkingLevel != null;
480
+ const currentPolicy = wasQueued || !provenance?.policy ? loadRoutingPolicy(options.configCwd ?? ctx.cwd) : undefined;
481
+ const policy = currentPolicy ?? provenance.policy;
482
+ record.routing = { mode: policy.mode, source: policy.source, fallbackSource: policy.fallbackSource,
483
+ code: policy.mode === "off" ? "off" : policy.diagnostic ? "guideline_unavailable" : "baseline",
484
+ reason: policy.mode === "off" ? "Model routing and routing guidance are off" : policy.diagnostic ?? "Using the existing model", guidelinePath: policy.guidelinePath, guidelineHash: policy.guidelineHash };
485
+ if (policy.diagnostic)
486
+ console.warn(`[pi-subagents] ${policy.diagnostic}`);
487
+ const route = policy.mode === "jev" || policy.mode === "shadow" ||
488
+ (policy.mode === "auto" && !explicit && !config?.model && !config?.thinking && policy.source === "jev");
489
+ if (!options.resumeSessionFile && provenance?.entrypoint !== "internal" && route) {
490
+ const stop = () => this.abort(id);
491
+ options.signal?.addEventListener("abort", stop, { once: true });
492
+ if (options.signal?.aborted)
493
+ stop();
494
+ try {
495
+ const routed = await this.router.choose(ctx, policy, prompt, options.description ?? type, options.model ?? resolveDefaultModel(ctx.model, ctx.modelRegistry, config?.model), record.abortController.signal, usage => {
496
+ record.routingUsage = usage;
497
+ // Classifier tokens stay separate from coding/context/output budgets.
498
+ const delta = { input: usage.input, output: usage.output, cacheWrite: usage.cacheWrite, cacheRead: usage.cacheRead, cost: usage.cost.total };
499
+ this.onUsage?.(record, delta);
500
+ for (let current = record; current;) {
501
+ current.lifetimeUsage.cost = (current.lifetimeUsage.cost ?? 0) + usage.cost.total;
502
+ current = current.parentAgentId ? this.agents.get(current.parentAgentId) : undefined;
503
+ }
504
+ });
505
+ record.routing = { ...routed.decision, guidelinePath: policy.guidelinePath, guidelineHash: policy.guidelineHash };
506
+ if (routed.model && policy.mode === "shadow") {
507
+ record.routing = { ...record.routing, code: "shadow", model: undefined, thinkingLevel: undefined, suggestedModel: routed.decision.model,
508
+ fallbackSource: policy.source, reason: "Jev suggested a model; shadow mode kept the default-priority model" };
509
+ }
510
+ else if (routed.model) {
511
+ options.model = routed.model;
512
+ options.thinkingLevel = routed.thinkingLevel;
513
+ record.invocation = { ...record.invocation, thinking: routed.thinkingLevel, requestedThinking: undefined };
514
+ }
515
+ }
516
+ catch {
517
+ record.routing.code = "classifier_error";
518
+ record.routing.reason = "Jev could not choose a model; using the existing model";
519
+ }
520
+ finally {
521
+ options.signal?.removeEventListener("abort", stop);
522
+ }
523
+ if (record.status !== "running" || record.abortController.signal.aborted) {
524
+ this.settleRun(record, true, pool);
525
+ return;
526
+ }
527
+ }
528
+ else if (policy.mode === "auto" && (explicit || config?.model || config?.thinking)) {
529
+ record.routing.code = "explicit";
530
+ record.routing.reason = "Explicit model or thinking; automatic routing skipped";
531
+ }
471
532
  // Worktree isolation: try to create a temporary git worktree. Strict —
472
533
  // fail loud if not possible (no silent fallback to main tree). Done BEFORE
473
534
  // the run is kicked off so a failure doesn't leave a half-running agent.
@@ -522,6 +583,7 @@ export class AgentManager {
522
583
  const promise = runAgent(ctx, type, prompt, {
523
584
  pi,
524
585
  agentId: id,
586
+ agentConfig: options.agentConfig,
525
587
  model: options.model,
526
588
  maxTurns: options.maxTurns,
527
589
  isolated: options.isolated,
@@ -1312,6 +1374,7 @@ export class AgentManager {
1312
1374
  */
1313
1375
  async dispose(pi) {
1314
1376
  clearInterval(this.cleanupInterval);
1377
+ this.abortAll();
1315
1378
  // Clear queue — via dequeue, so anyone blocked in spawnAndWait is woken
1316
1379
  // rather than left awaiting a gate nothing will ever resolve.
1317
1380
  this.dequeue(() => true);
@@ -5,7 +5,7 @@ import type { Model } from "@earendil-works/pi-ai";
5
5
  import type { ExtensionContext } from "@earendil-works/pi-coding-agent";
6
6
  import { type AgentSession, DefaultResourceLoader, type ExtensionAPI } from "@earendil-works/pi-coding-agent";
7
7
  import { type NestedAgentManager } from "./nested-tools.js";
8
- import type { ServiceTier, SubagentType, ThinkingLevel } from "./types.js";
8
+ import type { AgentConfig, ServiceTier, SubagentType, ThinkingLevel } from "./types.js";
9
9
  import type { LifetimeUsage } from "./usage.js";
10
10
  import type { CompiledSchema } from "./workflow/json-schema.js";
11
11
  /**
@@ -157,6 +157,8 @@ export interface ToolActivity {
157
157
  toolName: string;
158
158
  }
159
159
  export interface RunOptions {
160
+ /** Snapshot of the selected definition for this branch. */
161
+ agentConfig?: AgentConfig;
160
162
  /** ExtensionAPI instance — used for pi.exec() instead of execSync. */
161
163
  pi: ExtensionAPI;
162
164
  /** Manager-assigned id; suffixes session name to disambiguate parallel spawns (e.g. `Explore#a1b2c3d4`). */
@@ -450,8 +450,8 @@ function resolveConfiguredSessionDir(sessionDir, cwd) {
450
450
  return resolve(cwd, sessionDir);
451
451
  }
452
452
  export async function runAgent(ctx, type, prompt, options) {
453
- const config = getConfig(type);
454
- const agentConfig = getAgentConfig(type);
453
+ const agentConfig = options.agentConfig ?? getAgentConfig(type);
454
+ const config = agentConfig ?? getConfig(type);
455
455
  // Resolve working directory: worktree override > parent cwd
456
456
  const effectiveCwd = options.cwd ?? ctx.cwd;
457
457
  // Filesystem work happens in effectiveCwd; config discovery in configCwd.
@@ -479,7 +479,7 @@ export async function runAgent(ctx, type, prompt, options) {
479
479
  extras.skillBlocks = loaded;
480
480
  }
481
481
  }
482
- let toolNames = getToolNamesForType(type);
482
+ let toolNames = options.agentConfig ? options.agentConfig.builtinToolNames ?? [...BUILTIN_TOOL_NAMES] : getToolNamesForType(type);
483
483
  // Persistent memory: detect write capability and branch accordingly.
484
484
  // Account for disallowedTools — a tool in the base set but on the denylist is not truly available.
485
485
  if (agentConfig?.memory) {
@@ -2,9 +2,10 @@
2
2
  * custom-agents.ts — Load user-defined agents from project (.pi/agents/, plus the shared .agents/agents/ workspace) and global ($PI_CODING_AGENT_DIR/agents/, default ~/.pi/agent/agents/) locations.
3
3
  */
4
4
  import { existsSync, readdirSync, readFileSync } from "node:fs";
5
- import { basename, join } from "node:path";
5
+ import { basename, join, resolve } from "node:path";
6
6
  import { getAgentDir, parseFrontmatter } from "@earendil-works/pi-coding-agent";
7
7
  import { BUILTIN_TOOL_NAMES } from "./agent-types.js";
8
+ import { loadRoutingSettings } from "./settings.js";
8
9
  /**
9
10
  * The one thing a declared `name:` may not contain, matching Claude Code
10
11
  * exactly: it reserves `:` for plugin-scoped identifiers (`my-plugin:reviewer`)
@@ -42,15 +43,16 @@ export function loadCustomAgents(cwd, strict = false) {
42
43
  const workspaceProjectDir = join(cwd, ".agents", "agents");
43
44
  const projectDir = join(cwd, ".pi", "agents");
44
45
  const agents = new Map();
45
- loadFromDir(globalDir, agents, "global", strict); // lowest priority
46
- loadFromDir(workspaceProjectDir, agents, "project", strict); // shared workspace
47
- loadFromDir(projectDir, agents, "project", strict); // highest priority (overwrites)
46
+ const excluded = loadRoutingSettings(cwd).guidelineFile;
47
+ loadFromDir(globalDir, agents, "global", strict, excluded);
48
+ loadFromDir(workspaceProjectDir, agents, "project", strict, excluded);
49
+ loadFromDir(projectDir, agents, "project", strict, excluded);
48
50
  warnedLastLoad = warnedThisLoad;
49
51
  warnedThisLoad = new Set();
50
52
  return agents;
51
53
  }
52
54
  /** Load agent configs from a directory into the map. */
53
- function loadFromDir(dir, agents, source, strict) {
55
+ function loadFromDir(dir, agents, source, strict, excluded) {
54
56
  if (!existsSync(dir))
55
57
  return;
56
58
  let files;
@@ -61,6 +63,8 @@ function loadFromDir(dir, agents, source, strict) {
61
63
  return;
62
64
  }
63
65
  for (const file of files) {
66
+ if (file.toLowerCase() === "custom-route.md" || resolve(dir, file) === excluded)
67
+ continue;
64
68
  const filenameType = basename(file, ".md");
65
69
  const path = join(dir, file);
66
70
  const parsed = readAgentFile(path, strict);
package/dist/index.js CHANGED
@@ -28,16 +28,18 @@ import { isolationParam, resolveAgentInvocationConfig, resolveJoinMode } from ".
28
28
  import { describeMention, handleBase, isReservedHandle, parseMention, resolveHandleToType, stripAgentPrefix } from "./mention.js";
29
29
  import { runMentionClone } from "./mention-clone.js";
30
30
  import { describeModel, resolveModel } from "./model-resolver.js";
31
+ import { loadRoutingPolicy, routingGuidance } from "./model-routing.js";
31
32
  import { checkModelScope, isScopeModelsEnabled, setScopeModelsEnabled } from "./model-scope.js";
32
33
  import { getMaxSubagentDepth, setMaxSubagentDepth } from "./nested-tools.js";
33
34
  import { createOutputFilePath, ensureOutputFile, getOutputTranscriptDefault, sessionTaskDir, setOutputTranscriptDefault, streamToOutputFile, writeInitialEntry } from "./output-file.js";
34
35
  import { SubagentScheduler } from "./schedule.js";
35
36
  import { resolveStorePath, ScheduleStore } from "./schedule-store.js";
36
- import { applyAndEmitLoaded, loadSettings, saveAndEmitChanged } from "./settings.js";
37
+ import { applyAndEmitLoaded, loadSettings, projectRoutingSettings, saveAndEmitChanged } from "./settings.js";
37
38
  import { getForegroundOutcomeNote, getStatusNote, partialOutputSuffix } from "./status-note.js";
38
39
  import { createMentionProvider, mentionRoster } from "./ui/agent-mention.js";
39
40
  import { AgentWidget, buildInvocationTags, describeActivity, fgPreservingNestedStyles, formatCost, formatDuration, formatMs, formatTokens, formatTurns, getDisplayName, getPromptModeLabel, SPINNER, } from "./ui/agent-widget.js";
40
41
  import { FleetList } from "./ui/fleet-list.js";
42
+ import { showRoutingMenu } from "./ui/model-routing-menu.js";
41
43
  import { showSchedulesMenu } from "./ui/schedule-menu.js";
42
44
  import { selectItem } from "./ui/select-item.js";
43
45
  import { renderWorkflowCard, renderWorkflowEntryCard } from "./ui/workflow-card.js";
@@ -344,6 +346,7 @@ export default function (pi) {
344
346
  const reloadCustomAgents = (strict = false) => {
345
347
  const userAgents = loadCustomAgents(process.cwd(), strict);
346
348
  registerAgents(userAgents);
349
+ return userAgents;
347
350
  };
348
351
  // Initial load — the only strict one. A bad edit mid-session must not kill the
349
352
  // session on the next unrelated spawn, so every later reload keeps warning.
@@ -502,6 +505,8 @@ export default function (pi) {
502
505
  durationMs,
503
506
  tokens,
504
507
  usage,
508
+ routing: record.routing,
509
+ routingUsage: record.routingUsage,
505
510
  };
506
511
  }
507
512
  // Background completion: route through group join or send individual nudge
@@ -526,6 +531,7 @@ export default function (pi) {
526
531
  id: record.id, type: record.type, description: record.description,
527
532
  status: record.status, result: record.result, error: record.error,
528
533
  startedAt: record.startedAt, completedAt: record.completedAt,
534
+ routing: record.routing, routingUsage: record.routingUsage,
529
535
  });
530
536
  // Skip notification if result was already consumed via get_subagent_result
531
537
  if (record.resultConsumed) {
@@ -636,6 +642,8 @@ export default function (pi) {
636
642
  };
637
643
  const spawnTopLevel = (piRef, ctxRef, type, prompt, options) => {
638
644
  const safeOptions = { ...(options ?? {}) };
645
+ delete safeOptions.routing;
646
+ delete safeOptions.agentConfig;
639
647
  delete safeOptions.parentAgentId;
640
648
  // Internal too: a forged value would hide an RPC-spawned agent inside
641
649
  // someone else's workflow, and take it out of the concurrency pool with it.
@@ -1258,6 +1266,16 @@ export default function (pi) {
1258
1266
  fleet.setUICtx(ctx.ui);
1259
1267
  widget.onTurnStart();
1260
1268
  });
1269
+ pi.on("before_agent_start", (event, ctx) => {
1270
+ const guidance = routingGuidance(loadRoutingPolicy(ctx.cwd));
1271
+ const description = agentToolDescription + (guidance ? "\n\n" + guidance : "");
1272
+ if (agentTool.description !== description) {
1273
+ agentTool.description = description;
1274
+ pi.registerTool(agentTool);
1275
+ }
1276
+ if (guidance)
1277
+ return { systemPrompt: event.systemPrompt + "\n\n" + guidance };
1278
+ });
1261
1279
  /** Build the full type list text dynamically from available agents only. */
1262
1280
  const buildTypeListText = () => {
1263
1281
  const available = getAvailableTypes();
@@ -1454,10 +1472,11 @@ Terse command-style prompts produce shallow, generic work.
1454
1472
  // Held rather than registered inline: the mention clone reuses this exact
1455
1473
  // definition, so the agent it starts is an ordinary top-level spawn instead
1456
1474
  // of a second implementation that has to be kept in step with this one.
1475
+ const initialRoutingGuidance = routingGuidance(loadRoutingPolicy(process.cwd()));
1457
1476
  const agentTool = defineTool({
1458
1477
  name: SUBAGENT_TOOL_NAMES.AGENT,
1459
1478
  label: "Agent",
1460
- description: agentToolDescription,
1479
+ description: agentToolDescription + (initialRoutingGuidance ? "\n\n" + initialRoutingGuidance : ""),
1461
1480
  promptSnippet: "Launch autonomous sub-agents for complex multi-step tasks",
1462
1481
  promptGuidelines: [
1463
1482
  "Use Agent with specialized agents when the task matches an agent type's description. Subagents are valuable for parallelizing independent queries or for protecting the main context window from excessive results, but should not be used excessively when not needed. Importantly, avoid duplicating work that subagents are already doing — if you delegate research to a subagent, do not also perform the same searches yourself.",
@@ -1618,7 +1637,8 @@ Terse command-style prompts produce shallow, generic work.
1618
1637
  // Ensure we have UI context for widget rendering
1619
1638
  widget.setUICtx(ctx.ui);
1620
1639
  // Reload custom agents so new project/global .md files are picked up without restart
1621
- reloadCustomAgents();
1640
+ const userAgents = reloadCustomAgents();
1641
+ const routingPolicy = loadRoutingPolicy(ctx.cwd, userAgents);
1622
1642
  const rawType = params.subagent_type;
1623
1643
  // Single decision point for dispatch (#183): unknown, disabled and
1624
1644
  // case-ambiguous types are refused here, BEFORE anything spawns, so a
@@ -1765,6 +1785,11 @@ Terse command-style prompts produce shallow, generic work.
1765
1785
  const { modelName: recModelName, tags } = buildInvocationTags(rec.invocation);
1766
1786
  const recModeLabel = getPromptModeLabel(type);
1767
1787
  const recTags = recModeLabel ? [recModeLabel, ...tags] : tags;
1788
+ if (rec.routing?.source === "jev" && (rec.routingUsage || rec.routing.unpriced !== undefined)) {
1789
+ recTags.push(rec.routing.mode === "shadow" ? "Jev shadow" : "Jev");
1790
+ if (rec.routing.unpriced)
1791
+ recTags.push("Jev price unavailable");
1792
+ }
1768
1793
  return {
1769
1794
  displayName: getDisplayName(type),
1770
1795
  description: rec.description,
@@ -1800,7 +1825,7 @@ Terse command-style prompts produce shallow, generic work.
1800
1825
  subagent_type: requestedType,
1801
1826
  prompt: params.prompt,
1802
1827
  model: params.model,
1803
- thinking: thinking,
1828
+ thinking: params.thinking,
1804
1829
  max_turns: effectiveMaxTurns,
1805
1830
  isolated: isolated,
1806
1831
  isolation: isolation,
@@ -1888,6 +1913,8 @@ Terse command-style prompts produce shallow, generic work.
1888
1913
  description: params.description,
1889
1914
  name: params.name,
1890
1915
  model,
1916
+ agentConfig: customConfig,
1917
+ routing: { policy: routingPolicy, modelExplicit: !!resolvedConfig.modelInput, thinkingExplicit: thinking !== undefined, entrypoint: "agent" },
1891
1918
  maxTurns: effectiveMaxTurns,
1892
1919
  isolated,
1893
1920
  inheritContext,
@@ -2028,6 +2055,8 @@ Terse command-style prompts produce shallow, generic work.
2028
2055
  description: params.description,
2029
2056
  name: params.name,
2030
2057
  model,
2058
+ agentConfig: customConfig,
2059
+ routing: { policy: routingPolicy, modelExplicit: !!resolvedConfig.modelInput, thinkingExplicit: thinking !== undefined, entrypoint: "agent" },
2031
2060
  maxTurns: effectiveMaxTurns,
2032
2061
  isolated,
2033
2062
  inheritContext,
@@ -2687,6 +2716,7 @@ Terse command-style prompts produce shallow, generic work.
2687
2716
  // Actions
2688
2717
  options.push("Create new agent");
2689
2718
  options.push("Settings");
2719
+ options.push("Model routing");
2690
2720
  const noAgentsMsg = allNames.length === 0 && agents.length === 0
2691
2721
  ? "No agents found. Create specialized subagents that can be delegated to.\n\n" +
2692
2722
  "Each subagent has its own context window, custom system prompt, and specific tools.\n\n" +
@@ -2721,6 +2751,15 @@ Terse command-style prompts produce shallow, generic work.
2721
2751
  await showSettings(ctx);
2722
2752
  await showAgentsMenu(ctx);
2723
2753
  }
2754
+ else if (choice === "Model routing") {
2755
+ const patch = await showRoutingMenu(ctx);
2756
+ if (patch) {
2757
+ const toast = saveAndEmitChanged({ ...snapshotSettings(), ...patch }, "Model routing settings updated", (event, payload) => pi.events.emit(event, payload), ctx.cwd);
2758
+ ctx.ui.notify(toast.message, toast.level);
2759
+ reloadCustomAgents();
2760
+ }
2761
+ await showAgentsMenu(ctx);
2762
+ }
2724
2763
  }
2725
2764
  async function showAllAgentsList(ctx) {
2726
2765
  const allNames = getAllTypes();
@@ -3065,6 +3104,7 @@ Guidelines for choosing settings:
3065
3104
  Write the file using the write tool. Only write the file, nothing else.`;
3066
3105
  const { record } = await manager.spawnAndWait(pi, ctx, "general-purpose", generatePrompt, {
3067
3106
  description: `Generate ${name} agent`,
3107
+ routing: { modelExplicit: false, thinkingExplicit: false, entrypoint: "internal" },
3068
3108
  maxTurns: 5,
3069
3109
  // Exempt from maxConcurrentForeground. This runs from a modal wizard, not
3070
3110
  // a tool call: it passes no signal, and Esc in `ctx.ui` never reaches the
@@ -3173,6 +3213,7 @@ Write the file using the write tool. Only write the file, nothing else.`;
3173
3213
  */
3174
3214
  function snapshotSettings() {
3175
3215
  return {
3216
+ ...projectRoutingSettings(process.cwd()),
3176
3217
  maxConcurrent: manager.getMaxConcurrent(),
3177
3218
  // 0 = unlimited, and the default — see SubagentsSettings.
3178
3219
  maxConcurrentForeground: manager.getMaxConcurrentForeground(),