@diousk/pi-subagents-fast 0.20.0 → 0.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -7,6 +7,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.21.0] - 2026-09-30
11
+
12
+ > **Breaking: requires Pi 0.99.1 or newer.** Update Pi before installing this version. Legacy transcript and session API compatibility paths have been removed.
13
+
14
+ ### Added
15
+ - **Custom agents accept `service_tier: fast`.** OpenAI Responses and Codex requests forward the configured tier, and `priority` remains accepted.
16
+
17
+ ### Changed
18
+ - **The minimum supported Pi version is 0.99.1.** Development dependencies and blocking CI target that minimum. Scripted model fixtures use transcript replay and physical model lookup.
19
+
20
+ ### Fixed
21
+ - **Mention clones preserve conversation history on Pi 0.99.1.** History is seeded through the session manager, and the live prompt is supplied through the resource loader. Historical system messages do not grant the clone the parent's tools.
22
+ - **Subagent tool scope covers nested execution.** A bound `tool_call` guard blocks out-of-scope deferred/codemode calls, while tool narrowing preserves exposure. Synthetic extension paths match their logical names.
23
+ - **Output transcripts preserve system messages.** System/tool-state changes receive a `system` entry type, and a leading system message no longer causes the initial user prompt to be duplicated.
24
+
10
25
  ## [0.20.0] - 2026-09-09
11
26
 
12
27
  > **Breaking — the npm package is now `@diousk/pi-subagents-fast`.** Existing installs of `@tintinweb/pi-subagents` are not automatically migrated; install this fork with `pi install npm:@diousk/pi-subagents-fast`.
package/README.md CHANGED
@@ -76,7 +76,9 @@ npm pack --dry-run
76
76
 
77
77
  The `prepublishOnly` script runs lint, typecheck, tests, and the build before npm uploads the package. After publishing, install it with `pi install npm:@diousk/pi-subagents-fast`.
78
78
 
79
- Requires pi **0.84.0 or newer**: the [`SubagentWorkflow`](#subagentworkflow) tool builds on `constrainedSampling` (pi 0.82.0) and pi-tui's `stripTerminalSequences` (0.84.0). The `peerDependencies` range declares it, so npm flags an older pi at install time.
79
+ Requires **Pi 0.99.1 or newer** and **Node.js 22.19.0 or newer**. Update Pi before installing this extension. The `peerDependencies` range declares the minimum, so npm flags an older Pi at install time.
80
+
81
+ The development and CI baseline is **Pi 0.99.1**. Use that version or newer for `openai/gpt-6.1-sol` (API key) or `openai-codex/gpt-6.1-sol` (your Pi Codex login). Child sessions reuse the parent's configured providers and authentication; model limits and pricing come from Pi's provider catalog. See the [Pi 0.99.1 release](https://pi.dev/changelog/releases/0.99.1).
80
82
 
81
83
  ### Other hosts
82
84
 
@@ -337,8 +339,8 @@ All fields are optional — sensible defaults for everything.
337
339
  | `disallowed_tools` | — | Comma-separated tools to deny even if extensions provide them |
338
340
  | `isolation` | — | Set to `worktree` to run in an isolated git worktree, or `off` to refuse one even when the caller passes `isolation: "worktree"` (frontmatter is authoritative). `none`, `no`, and `false` are accepted spellings of `off` |
339
341
  | `model` | inherit parent | Model — `provider/modelId` or fuzzy name (`"haiku"`, `"sonnet"`). Resolved tolerantly (`.`/`-` and a trailing date stamp are interchangeable) and falls back to the same model under another provider if the named one doesn't have it |
340
- | `service_tier` | — | OpenAI Responses/Codex processing tier: `auto`, `default`, `flex`, `priority`, or `scale`. Applied only to `openai-responses` and `openai-codex-responses`; omitted preserves the provider default, and other APIs ignore it without displaying it as active |
341
- | `thinking` | inherit | off, minimal, low, medium, high, xhigh, max — actual availability depends on your pi version and model; pi clamps unsupported levels down |
342
+ | `service_tier` | — | OpenAI Responses/Codex processing tier: `auto`, `default`, `flex`, `fast`, `priority`, or `scale`. Applied only to `openai-responses` and `openai-codex-responses`; omitted preserves the provider default, and other APIs ignore it without displaying it as active |
343
+ | `thinking` | inherit | off, minimal, low, medium, high, xhigh, max — actual availability depends on your pi version and model; pi maps or clamps unsupported levels |
342
344
  | `max_turns` | unlimited | Max agentic turns before graceful shutdown. `0` or omit for unlimited |
343
345
  | `persist_session` | `subagents.json` `rememberAgents` (default `true`) | Persist this subagent as a normal pi session instead of keeping the session in memory only; overrides the `rememberAgents` project default in both directions. It records its spawning session as parent, so it nests under it in `/resume`. The subagent's `.output` transcript is still written either way unless `output_transcript: false` |
344
346
  | `output_transcript` | `true` (or `subagents.json` `outputTranscript`) | Write this subagent's `.output` transcript; when set, overrides the `subagents.json` `outputTranscript` default. Set `false` to write no transcript file or path. Governs only the transcript — independent of `persist_session`, `isolation: worktree`, and `memory:` |
@@ -350,7 +352,17 @@ All fields are optional — sensible defaults for everything.
350
352
  | `isolated` | `false` | Hermetic specialist mode: forces `extensions: false` + `skills: false` + drops `ext:` selectors. Only built-in tools. Distinct from `isolation: worktree` (filesystem) |
351
353
  | `enabled` | `true` | Set to `false` to disable an agent (useful for hiding a default agent per-project) |
352
354
 
353
- For an OpenAI Responses or Codex agent, set `service_tier: priority` to request priority processing. The UI shows the requested tier only when the effective model uses one of those APIs; other providers keep their normal request behavior and do not display a tier as active.
355
+ For an OpenAI Responses or Codex agent, set `service_tier: fast` to request fast processing; `priority` remains a supported alias. Availability depends on the provider and account. The UI shows the requested tier only when the effective model uses one of those APIs. See [OpenAI fast mode](https://developers.openai.com/api/docs/guides/fast-mode).
356
+
357
+ GPT-6.1 Sol supports `low`, `medium`, `high`, `xhigh`, and `max` reasoning. In Pi 0.99.1, `minimal` maps to provider effort `low`; `off` is clamped to that same alias because Sol cannot disable reasoning. The UI reports Pi's logical thinking level. Prefer explicit supported levels in agent files:
358
+
359
+ ```yaml
360
+ model: openai-codex/gpt-6.1-sol
361
+ thinking: medium
362
+ service_tier: fast
363
+ ```
364
+
365
+ Sol requires the Responses API for tool use. Selecting its Pi catalog entry chooses the corresponding Responses transport; no custom provider override is needed. See the [Sol model reference](https://developers.openai.com/api/docs/models/gpt-6.1-sol).
354
366
 
355
367
  Frontmatter is authoritative. If an agent file sets `model`, `thinking`, `service_tier`, `max_turns`, `inherit_context`, `run_in_background`, `isolated`, or `isolation`, those values are locked for that agent. For fields exposed by the `Agent` tool, its parameters only fill values the agent config leaves unspecified; `service_tier` is frontmatter-only.
356
368
 
@@ -408,6 +420,8 @@ A few rules the examples don't make obvious:
408
420
  - `extensions:` is the sole loading authority. `ext:foo` in `tools:` narrows what surfaces; it can't load `foo` on its own. Mismatches fire `extension-error:…` warnings.
409
421
  - Any `ext:` entry flips extension tools to an explicit allowlist — unnamed extensions still load (handlers fire) but expose no tools. So `tools: "*, ext:mcp/search"` exposes only `search` from `mcp`, nothing from any other extension.
410
422
  - Extension names match case-insensitively (`[Mcp]` = `[mcp]`); tool names in `ext:foo/bar` stay case-sensitive.
423
+ - Synthetic extension identities match their logical names: `builtin:mcp` as `mcp`, `builtin:codemode` as `codemode`, and `<inline:foo>` as `foo`. This matches loaded extensions; it does not load Pi CLI built-in factories into SDK child sessions.
424
+ - Scope checks also apply to nested `ctx.executeTool()` calls, including deferred and codemode tools. Deferred tools keep their exposure and are not automatically promoted into model declarations; extension-activated tools remain active only within the agent's scope.
411
425
  - Extensions that register tools **lazily** work too. MCP-backed extensions typically can't enumerate their tools until their servers connect, so they register from `session_start` or `before_agent_start` rather than at load. Subagent scoping is re-derived as tools appear, so these surface normally — including under `ext:` selectors, which keep narrowing correctly no matter when a tool shows up.
412
426
  - Extensions bound into a subagent see **both ends** of that session's lifecycle: `session_start` when the agent starts, `session_shutdown` (reason `quit`) when its session is disposed — on quit, and when its record is evicted ~10 minutes after it finishes. Release per-session resources there; anything left armed outlives the session it belongs to. Handlers are given three seconds on quit, after which teardown proceeds regardless.
413
427
  - An installed **package** extension matches by its package short name (`@scope/pi-subagents` → `[pi-subagents]`), in addition to its path-derived name (a package whose entry is `src/index.ts` also answers to `[src]`). Prefer the package name — the path-derived one is incidental.
@@ -667,7 +681,7 @@ Runtime tuning values set via `/agents` → Settings (max concurrency, max foreg
667
681
 
668
682
  **Report usage to session** (`reportUsage`, default `false`): whether subagent spend is added to *this* session's own totals. Subagents run in their own pi sessions, so by default pi's footer, statusline and `/cost` count only what the main model spent — a session that delegated most of its work reads as nearly free. Turn it on and each `Agent` / `get_subagent_result` / `steer_subagent` result carries the spend accumulated since the last one, which pi folds into `getSessionStats()`; `/cost` attributes it to the **Tools/summaries** bucket. Toggle via `/agents → Settings → Report usage to session`; applied live.
669
683
 
670
- Three things worth knowing about the numbers. Every token component is reported, `cacheRead` included — the cached prefix genuinely is re-read and re-billed on every call, and pi counts it the same way for the session's own messages, so withholding it would make a subagent's rows count differently from every other row in one total. (The extension's *own* token displays still leave it out, which is a different question: there it inflates a reading of how much work was done.) Cost is pi's own per-message figure, priced from the model's listed rates; a model pi has no rates for contributes zero rather than an estimate. And the context-window percentage is untouched: pi derives it from assistant messages alone, so a delegating session's context doesn't appear to fill up faster. Agents that finish in the background have no tool result of their own to ride on, so their spend is carried by the next one you make — the footer catches up on the following call, not the moment they finish.
684
+ Three things worth knowing about the numbers. Every token component is reported, `cacheRead` included — the cached prefix genuinely is re-read and re-billed on every call, and pi counts it the same way for the session's own messages, so withholding it would make a subagent's rows count differently from every other row in one total. (The extension's *own* token displays still leave it out, which is a different question: there it inflates a reading of how much work was done.) Cost is pi's own per-message figure, priced from the model's listed rates; a model pi has no rates for contributes zero rather than an estimate. Reported subagent token counts do not inflate the parent's context-window percentage. On Pi 0.99.1 the tool-result text itself occupies context and is included in Pi's projection estimate. Agents that finish in the background have no tool result of their own to ride on, so their spend is carried by the next one you make — the footer catches up on the following call, not the moment they finish.
671
685
 
672
686
  **Show cost** (`showCost`, default `false`): whether the subagent surfaces print an estimated cost beside their token counts — the widget (running *and* finished lines), [FleetView](#fleetview), the conversation viewer, foreground results, `get_subagent_result`, and completion notifications:
673
687
 
@@ -696,6 +710,8 @@ Both places report what the run *actually* used, read back from the child sessio
696
710
  ↳ anthropic/claude-haiku-4-5 · thinking: low (asked max) · background
697
711
  ```
698
712
 
713
+ For a virtual model, the displayed identity is the session's selected virtual model, rather than the physical model chosen for each response.
714
+
699
715
  Toggle via `/agents → Settings → Show model`; applied live.
700
716
 
701
717
  **Viewer markdown** (`viewerMarkdown`, default `"assistant"`): how much of the [conversation viewer](#ui)'s transcript is rendered as Markdown rather than shown verbatim.
@@ -85,12 +85,13 @@ export declare function parseExtSelectors(entries: string[]): {
85
85
  * snapshotted. `registerTool` writes into the very `extension.tools` maps this reads,
86
86
  * so `inScope()` sees late arrivals on the next call.
87
87
  *
88
- * Two enforcement points, because neither covers the whole picture:
88
+ * The active set and call-time checks cover different parts of scope:
89
89
  *
90
90
  * - `turn_end` re-narrows the ACTIVE set. pi emits `turn_end` immediately before
91
91
  * `prepareNextTurn` re-snapshots `agent.state.tools`, and session listeners run
92
92
  * synchronously, so the narrow lands in time for turns 2..N.
93
- * - `beforeToolCall` blocks out-of-scope calls. Turn 1 cannot be narrowed at all:
93
+ * - The loader's bound `tool_call` handler blocks direct and nested calls.
94
+ * Turn 1 cannot be narrowed at all:
94
95
  * `before_agent_start` fires INSIDE `prompt()` and may widen the tool set, but
95
96
  * `createContextSnapshot()` freezes that turn's tools immediately after — there
96
97
  * is no hook in between. A call-time check is the only correct guard there.
@@ -117,7 +118,7 @@ export declare function installExtensionToolScope(session: AgentSession, ctx: {
117
118
  * seeded from `toolNames`.
118
119
  */
119
120
  readmitToolNames: Set<string>;
120
- }): void;
121
+ }): (toolName: string) => boolean;
121
122
  /** Normalize max turns. undefined or 0 = unlimited, otherwise minimum 1. */
122
123
  export declare function normalizeMaxTurns(n: number | undefined): number | undefined;
123
124
  /** Get the default max turns value. undefined = unlimited. */
@@ -68,6 +68,10 @@ export function installServiceTierPayload(session, serviceTier) {
68
68
  * single-file extensions to the basename minus `.ts`/`.js`.
69
69
  */
70
70
  export function extensionCanonicalName(extPath) {
71
+ if (extPath.startsWith("builtin:"))
72
+ return extPath.slice("builtin:".length).toLowerCase();
73
+ if (extPath.startsWith("<inline:") && extPath.endsWith(">"))
74
+ return extPath.slice(8, -1).toLowerCase();
71
75
  const base = basename(extPath);
72
76
  const name = base === "index.ts" || base === "index.js"
73
77
  ? basename(dirname(extPath))
@@ -132,6 +136,8 @@ function extensionPackageName(extPath) {
132
136
  */
133
137
  export function extensionCanonicalNames(extPath) {
134
138
  const canonical = extensionCanonicalName(extPath);
139
+ if (extPath.startsWith("builtin:") || extPath.startsWith("<inline:"))
140
+ return [canonical];
135
141
  const pkg = extensionPackageName(extPath);
136
142
  return pkg && pkg !== canonical ? [canonical, pkg] : [canonical];
137
143
  }
@@ -219,12 +225,13 @@ export function parseExtSelectors(entries) {
219
225
  * snapshotted. `registerTool` writes into the very `extension.tools` maps this reads,
220
226
  * so `inScope()` sees late arrivals on the next call.
221
227
  *
222
- * Two enforcement points, because neither covers the whole picture:
228
+ * The active set and call-time checks cover different parts of scope:
223
229
  *
224
230
  * - `turn_end` re-narrows the ACTIVE set. pi emits `turn_end` immediately before
225
231
  * `prepareNextTurn` re-snapshots `agent.state.tools`, and session listeners run
226
232
  * synchronously, so the narrow lands in time for turns 2..N.
227
- * - `beforeToolCall` blocks out-of-scope calls. Turn 1 cannot be narrowed at all:
233
+ * - The loader's bound `tool_call` handler blocks direct and nested calls.
234
+ * Turn 1 cannot be narrowed at all:
228
235
  * `before_agent_start` fires INSIDE `prompt()` and may widen the tool set, but
229
236
  * `createContextSnapshot()` freezes that turn's tools immediately after — there
230
237
  * is no hook in between. A call-time check is the only correct guard there.
@@ -263,7 +270,7 @@ export function installExtensionToolScope(session, ctx) {
263
270
  for (const name of EXCLUDED_TOOL_NAMES)
264
271
  keep.delete(name);
265
272
  // Injected tools are legitimately active for this agent — re-admit them so
266
- // the renarrow keeps them in the active set and beforeToolCall doesn't
273
+ // the renarrow keeps them in the active set and the call-time guard doesn't
267
274
  // block them. Already vetted against `disallowed_tools` by the caller,
268
275
  // which is the only place that knows which kind may be taken back.
269
276
  for (const name of readmitToolNames)
@@ -272,8 +279,10 @@ export function installExtensionToolScope(session, ctx) {
272
279
  };
273
280
  const renarrow = () => {
274
281
  const allowed = inScope();
275
- const next = session.getAllTools().map((t) => t.name).filter((n) => allowed.has(n));
276
282
  const current = session.getActiveToolNames();
283
+ // Keep deferred/codemode tools callable without promoting them into model
284
+ // declarations. Explicit activation by an extension is preserved in scope.
285
+ const next = session.getAllTools().filter((tool) => allowed.has(tool.name) && (tool.exposure === "direct" || tool.exposure === "model-only" || current.includes(tool.name))).map((tool) => tool.name);
277
286
  // setActiveToolsByName unconditionally rebuilds the system prompt, so skip
278
287
  // the no-op that steady-state turns would otherwise pay for every turn.
279
288
  if (next.length !== current.length || next.some((n, i) => n !== current[i])) {
@@ -287,16 +296,7 @@ export function installExtensionToolScope(session, ctx) {
287
296
  if (event.type === "turn_end")
288
297
  renarrow();
289
298
  });
290
- const priorBeforeToolCall = session.agent.beforeToolCall;
291
- session.agent.beforeToolCall = async (context, signal) => {
292
- if (!inScope().has(context.toolCall.name)) {
293
- return {
294
- block: true,
295
- reason: `Tool "${context.toolCall.name}" is not available to this subagent.`,
296
- };
297
- }
298
- return priorBeforeToolCall?.(context, signal);
299
- };
299
+ return (toolName) => inScope().has(toolName);
300
300
  }
301
301
  /** Default max turns. undefined = unlimited (no turn limit). */
302
302
  let defaultMaxTurns;
@@ -557,6 +557,7 @@ export async function runAgent(ctx, type, prompt, options) {
557
557
  // must compare against this, not the surviving set (absence from survivors is
558
558
  // an exclude *succeeding*).
559
559
  let discoveredNames;
560
+ let toolInScope;
560
561
  const extensionsOverride = noExtensions || (loadAll && !hasExcludes)
561
562
  ? undefined
562
563
  : (base) => {
@@ -564,6 +565,8 @@ export async function runAgent(ctx, type, prompt, options) {
564
565
  return {
565
566
  ...base,
566
567
  extensions: base.extensions.filter((e) => {
568
+ if (e.path === "<inline:subagent-tool-scope>")
569
+ return true;
567
570
  const canons = extensionCanonicalNames(e.path);
568
571
  if (canons.some((n) => excludeNames.has(n)))
569
572
  return false; // exclude wins
@@ -576,6 +579,19 @@ export async function runAgent(ctx, type, prompt, options) {
576
579
  agentDir,
577
580
  noExtensions,
578
581
  additionalExtensionPaths,
582
+ // Pi's nested ctx.executeTool calls dispatch tool_call directly, and prompt
583
+ // setup can replace agent.beforeToolCall. Enforce scope in that dispatch too.
584
+ extensionFactories: noExtensions ? [] : [{
585
+ name: "subagent-tool-scope",
586
+ hidden: true,
587
+ factory: (pi) => {
588
+ pi.on("tool_call", (event) => {
589
+ if (toolInScope && !toolInScope(event.toolName)) {
590
+ return { block: true, reason: `Tool "${event.toolName}" is not available to this subagent.` };
591
+ }
592
+ });
593
+ },
594
+ }],
579
595
  extensionsOverride,
580
596
  noSkills,
581
597
  noPromptTemplates: true,
@@ -786,21 +802,14 @@ export async function runAgent(ctx, type, prompt, options) {
786
802
  parentSession: ctx.sessionManager?.getSessionFile?.(),
787
803
  })
788
804
  : SessionManager.inMemory(effectiveCwd);
789
- // Pi 0.80.8 replaced createAgentSession's modelRegistry option with
790
- // modelRuntime, but ExtensionContext still exposes only the registry facade.
791
- // Pass both so the full supported Pi range retains the parent's providers.
805
+ // ExtensionContext exposes the registry facade; its runtime owns provider auth.
792
806
  const parentModelRuntime = ctx.modelRegistry.runtime;
793
807
  const sessionOpts = {
794
808
  cwd: effectiveCwd,
795
809
  agentDir,
796
810
  sessionManager,
797
811
  settingsManager,
798
- modelRegistry: ctx.modelRegistry,
799
- // `as never` is what keeps this assignable across the supported Pi range:
800
- // pre-0.80.8 the field exists only via the `modelRuntime?: unknown` shim
801
- // above, while newer Pi types it as `ModelRuntime` — a shape an opaque
802
- // `unknown` read off the private facade field can never satisfy.
803
- ...(parentModelRuntime !== undefined && { modelRuntime: parentModelRuntime }),
812
+ modelRuntime: parentModelRuntime,
804
813
  model,
805
814
  tools: sessionTools,
806
815
  customTools: [...nestedTools, ...structuredTools],
@@ -835,7 +844,7 @@ export async function runAgent(ctx, type, prompt, options) {
835
844
  // handled below by re-deriving scope from the loader's live extension maps —
836
845
  // `registerTool` writes into those same maps, so late arrivals are judged too.
837
846
  if (!noExtensions) {
838
- installExtensionToolScope(session, {
847
+ toolInScope = installExtensionToolScope(session, {
839
848
  loader,
840
849
  toolNames,
841
850
  disallowedSet,
@@ -278,7 +278,7 @@ function parseMemory(val) {
278
278
  }
279
279
  /** Parse the OpenAI Responses/Codex `service_tier` frontmatter field. */
280
280
  function parseServiceTier(val) {
281
- if (val === "auto" || val === "default" || val === "flex" || val === "priority" || val === "scale") {
281
+ if (val === "auto" || val === "default" || val === "flex" || val === "fast" || val === "priority" || val === "scale") {
282
282
  return val;
283
283
  }
284
284
  return undefined;
@@ -24,20 +24,6 @@
24
24
  * conversation with nothing in it yet clones to nothing in it yet, which is the
25
25
  * correct answer rather than a failure.
26
26
  *
27
- * It is also the oldest of the equivalent Pi APIs — `buildContextEntries` on
28
- * ReadonlySessionManager and the `sessionEntryToContextMessages` export both
29
- * arrived in 0.80.5 — where this one has been exported unchanged from before
30
- * the declared peer floor, and is the same code path (`byId` is only an index
31
- * cache, so passing it or not cannot change the result). Keeping the floor
32
- * honest costs nothing here: see the `compat-floor-pi` job.
33
- *
34
- * Its `thinkingLevel` is NOT used, and is the one place the newer API would be
35
- * better. `getSessionContextSettings` starts at "off" and moves only on an
36
- * explicit `thinking_level_change` entry, so a session where nobody ran
37
- * `/think` reports "off" rather than the level it is really using. Omitting the
38
- * field instead lets `createAgentSession` resolve it from settings, which is
39
- * that real level.
40
- *
41
27
  * Three details make the spawn belong to the real session rather than the
42
28
  * clone:
43
29
  *
@@ -24,20 +24,6 @@
24
24
  * conversation with nothing in it yet clones to nothing in it yet, which is the
25
25
  * correct answer rather than a failure.
26
26
  *
27
- * It is also the oldest of the equivalent Pi APIs — `buildContextEntries` on
28
- * ReadonlySessionManager and the `sessionEntryToContextMessages` export both
29
- * arrived in 0.80.5 — where this one has been exported unchanged from before
30
- * the declared peer floor, and is the same code path (`byId` is only an index
31
- * cache, so passing it or not cannot change the result). Keeping the floor
32
- * honest costs nothing here: see the `compat-floor-pi` job.
33
- *
34
- * Its `thinkingLevel` is NOT used, and is the one place the newer API would be
35
- * better. `getSessionContextSettings` starts at "off" and moves only on an
36
- * explicit `thinking_level_change` entry, so a session where nobody ran
37
- * `/think` reports "off" rather than the level it is really using. Omitting the
38
- * field instead lets `createAgentSession` resolve it from settings, which is
39
- * that real level.
40
- *
41
27
  * Three details make the spawn belong to the real session rather than the
42
28
  * clone:
43
29
  *
@@ -60,7 +46,7 @@
60
46
  * The clone gets one tool and one job. It cannot read, write or run anything —
61
47
  * an invisible turn with the full toolset could do invisible work.
62
48
  */
63
- import { buildSessionContext, createAgentSession, SessionManager, } from "@earendil-works/pi-coding-agent";
49
+ import { buildSessionContext, convertToLlm, createAgentSession, DefaultResourceLoader, getAgentDir, SessionManager, } from "@earendil-works/pi-coding-agent";
64
50
  import { runInChildSessionContext } from "./child-context.js";
65
51
  import { agentMentionReminder } from "./mention.js";
66
52
  /**
@@ -90,31 +76,51 @@ export async function runMentionClone(opts) {
90
76
  // false, and a foreground agent answers through its TOOL RESULT — which
91
77
  // here is delivered into a session that is disposed moments later, so the
92
78
  // agent would run, appear in the widget and the fleet, and reach nobody.
93
- return agentTool.execute(undefined, { ...params, run_in_background: true }, signal, onUpdate, ctx);
79
+ return agentTool.execute(undefined, { ...params, run_in_background: true }, signal, onUpdate, { ..._cloneCtx, ...ctx });
94
80
  },
95
81
  };
96
82
  let session;
97
83
  try {
98
- // Pi 0.80.8 moved createAgentSession from modelRegistry to modelRuntime;
99
- // agent-runner.ts carries the same shim for the same reason — pass both so
100
- // the clone keeps the parent's providers across the supported range.
84
+ // The registry facade retains the parent's configured runtime and auth.
101
85
  const parentModelRuntime = ctx.modelRegistry.runtime;
102
86
  // The conversation as the main session resolves it: compaction applied,
103
87
  // branch summaries substituted.
104
88
  const conversation = buildSessionContext(ctx.sessionManager.getEntries(), ctx.sessionManager.getLeafId());
105
- // Pi 0.82.0 added this; below it the field is absent and the clone takes
106
- // the settings level instead, which is what a session that never ran
107
- // `/think` is on anyway. Same shim shape as `modelRuntime` below.
108
- const thinkingLevel = ctx.thinkingLevel;
89
+ const sessionManager = SessionManager.inMemory(ctx.cwd);
90
+ // Replay conversation turns through the session manager. Historical system
91
+ // messages contain the parent's tool declarations, which the clone must not inherit.
92
+ for (const entry of conversation.messages) {
93
+ if (entry.role === "system")
94
+ continue;
95
+ if (entry.role === "branchSummary" || entry.role === "compactionSummary") {
96
+ for (const message of convertToLlm([entry]))
97
+ sessionManager.appendMessage(message);
98
+ }
99
+ else {
100
+ sessionManager.appendMessage(entry);
101
+ }
102
+ }
103
+ const resourceLoader = new DefaultResourceLoader({
104
+ cwd: ctx.cwd,
105
+ agentDir: getAgentDir(),
106
+ noExtensions: true,
107
+ noSkills: true,
108
+ noPromptTemplates: true,
109
+ noThemes: true,
110
+ noContextFiles: true,
111
+ systemPromptOverride: () => ctx.getSystemPrompt(),
112
+ appendSystemPromptOverride: () => [],
113
+ });
114
+ await resourceLoader.reload();
109
115
  const created = await runInChildSessionContext(() => createAgentSession({
110
116
  cwd: ctx.cwd,
111
117
  // Nothing about the copy is worth persisting, and an in-memory manager
112
118
  // is also what keeps the real session untouched.
113
- sessionManager: SessionManager.inMemory(ctx.cwd),
119
+ sessionManager,
120
+ resourceLoader,
114
121
  model: ctx.model,
115
- ...(thinkingLevel && { thinkingLevel }),
116
- modelRegistry: ctx.modelRegistry,
117
- ...(parentModelRuntime !== undefined && { modelRuntime: parentModelRuntime }),
122
+ ...(ctx.thinkingLevel && { thinkingLevel: ctx.thinkingLevel }),
123
+ modelRuntime: parentModelRuntime,
118
124
  // An allowlist naming exactly the clone's own tool. NOT `noTools:
119
125
  // "all"`, whose doc comment ("start with no tools enabled") reads like
120
126
  // it spares custom tools and does not: it resolves to an EMPTY
@@ -127,16 +133,6 @@ export async function runMentionClone(opts) {
127
133
  customTools: [cloneAgentTool],
128
134
  }));
129
135
  session = created.session;
130
- // The clone rebuilds a system prompt from cwd and agentDir, which is close
131
- // but not the live one — extensions contribute to it per turn. Copy the
132
- // real thing, so the copy reasons under the instructions the user's model
133
- // is actually working under.
134
- const systemPrompt = ctx.getSystemPrompt?.();
135
- if (systemPrompt)
136
- session.agent.state.systemPrompt = systemPrompt;
137
- // The conversation itself. Pushed rather than assigned so the array the
138
- // session was built around stays the one it goes on using.
139
- session.agent.state.messages.push(...conversation.messages);
140
136
  // User text first, reminder after — the order Claude Code's attachment
141
137
  // renderer produces, where the reminder trails the message it is about.
142
138
  await session.prompt(`${message}\n\n${agentMentionReminder(type)}`);
@@ -91,20 +91,26 @@ export function writeInitialEntry(path, agentId, prompt, cwd) {
91
91
  * Returns a cleanup function that does a final flush and unsubscribes.
92
92
  */
93
93
  export function streamToOutputFile(session, path, agentId, cwd, startIndex) {
94
- // Index of the first message this stream is responsible for. A spawn writes
95
- // messages[0] as the initial prompt entry, so it starts at 1. A resume hands
94
+ // A spawn writes its initial user prompt separately. Pi can project system
95
+ // messages before that prompt, so skip the first user message by role. A resume hands
96
96
  // in the session's length as of just before the run: the session already
97
97
  // holds every prior turn, and re-emitting those would duplicate history that
98
98
  // is already in the file.
99
- let writtenCount = startIndex ?? 1;
99
+ let writtenCount = startIndex ?? 0;
100
+ let skipInitialUser = startIndex === undefined;
100
101
  const flush = () => {
101
102
  const messages = session.messages;
102
103
  while (writtenCount < messages.length) {
103
104
  const msg = messages[writtenCount];
105
+ if (skipInitialUser && msg.role === "user") {
106
+ skipInitialUser = false;
107
+ writtenCount++;
108
+ continue;
109
+ }
104
110
  const entry = {
105
111
  isSidechain: true,
106
112
  agentId,
107
- type: msg.role === "assistant" ? "assistant" : msg.role === "user" ? "user" : "toolResult",
113
+ type: msg.role === "assistant" ? "assistant" : msg.role === "user" ? "user" : msg.role === "system" ? "system" : "toolResult",
108
114
  message: msg,
109
115
  timestamp: new Date().toISOString(),
110
116
  cwd,
package/dist/types.d.ts CHANGED
@@ -12,7 +12,7 @@ export declare const DEFAULT_AGENT_NAMES: readonly ["general-purpose", "Explore"
12
12
  /** Memory scope for persistent agent memory. */
13
13
  export type MemoryScope = "user" | "project" | "local";
14
14
  /** OpenAI Responses/Codex request processing tier. */
15
- export type ServiceTier = "auto" | "default" | "flex" | "priority" | "scale";
15
+ export type ServiceTier = "auto" | "default" | "flex" | "fast" | "priority" | "scale";
16
16
  /**
17
17
  * Isolation mode for agent execution.
18
18
  *
package/docs/workflows.md CHANGED
@@ -250,6 +250,8 @@ Any other key is rejected **by name** at the call. Note that this checks option
250
250
 
251
251
  Combination rules: `resume` cannot be combined with `agentType`, `model`, `effort`, `isolation`, `gate` or `schema` — a resumed child keeps the agent type, model and tree it was started with, and its session predates the `StructuredOutput` tool.
252
252
 
253
+ With Pi 0.99.1, select Sol with `model: "openai-codex/gpt-6.1-sol"` (Pi's Codex login) or `model: "openai/gpt-6.1-sol"` (API key). Use `effort: "low"`, `"medium"`, `"high"`, `"xhigh"`, or `"max"`; Pi maps `minimal` to provider effort `low`. A custom `agentType` can pin `service_tier: fast` in its agent file; `service_tier` is not a workflow option. See the [README's model reference](../README.md#frontmatter-fields).
254
+
253
255
  ### `pipeline()` and `parallel()`
254
256
 
255
257
  ```js
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@diousk/pi-subagents-fast",
3
- "version": "0.20.0",
3
+ "version": "0.21.0",
4
4
  "description": "A pi extension that brings Claude Code-like sub-agents and workflow orchestration to pi — parallel execution, live widget, fleet view, custom agent types, mid-run steering, dynamic workflows, Claude Code compatibility, look and feel.",
5
5
  "author": "tintinweb",
6
6
  "license": "MIT",
@@ -24,9 +24,9 @@
24
24
  "autonomous"
25
25
  ],
26
26
  "peerDependencies": {
27
- "@earendil-works/pi-ai": ">=0.84.0",
28
- "@earendil-works/pi-coding-agent": ">=0.84.0",
29
- "@earendil-works/pi-tui": ">=0.84.0"
27
+ "@earendil-works/pi-ai": ">=0.99.1",
28
+ "@earendil-works/pi-coding-agent": ">=0.99.1",
29
+ "@earendil-works/pi-tui": ">=0.99.1"
30
30
  },
31
31
  "dependencies": {
32
32
  "@sinclair/typebox": "^0.34.49",
@@ -50,9 +50,9 @@
50
50
  },
51
51
  "devDependencies": {
52
52
  "@biomejs/biome": "^2.4.14",
53
- "@earendil-works/pi-ai": "0.84.2",
54
- "@earendil-works/pi-coding-agent": "0.84.2",
55
- "@earendil-works/pi-tui": "0.84.2",
53
+ "@earendil-works/pi-ai": "0.99.1",
54
+ "@earendil-works/pi-coding-agent": "0.99.1",
55
+ "@earendil-works/pi-tui": "0.99.1",
56
56
  "@types/node": "^25.5.0",
57
57
  "@vitest/coverage-istanbul": "^4.1.10",
58
58
  "typescript": "^6.0.0",
@@ -14,6 +14,7 @@ import {
14
14
  DefaultResourceLoader,
15
15
  type ExtensionAPI,
16
16
  getAgentDir,
17
+ type ModelRuntime,
17
18
  SessionManager,
18
19
  SettingsManager,
19
20
  } from "@earendil-works/pi-coding-agent";
@@ -94,6 +95,8 @@ export function installServiceTierPayload(
94
95
  * single-file extensions to the basename minus `.ts`/`.js`.
95
96
  */
96
97
  export function extensionCanonicalName(extPath: string): string {
98
+ if (extPath.startsWith("builtin:")) return extPath.slice("builtin:".length).toLowerCase();
99
+ if (extPath.startsWith("<inline:") && extPath.endsWith(">")) return extPath.slice(8, -1).toLowerCase();
97
100
  const base = basename(extPath);
98
101
  const name = base === "index.ts" || base === "index.js"
99
102
  ? basename(dirname(extPath))
@@ -159,6 +162,7 @@ function extensionPackageName(extPath: string): string | undefined {
159
162
  */
160
163
  export function extensionCanonicalNames(extPath: string): string[] {
161
164
  const canonical = extensionCanonicalName(extPath);
165
+ if (extPath.startsWith("builtin:") || extPath.startsWith("<inline:")) return [canonical];
162
166
  const pkg = extensionPackageName(extPath);
163
167
  return pkg && pkg !== canonical ? [canonical, pkg] : [canonical];
164
168
  }
@@ -250,12 +254,13 @@ export function parseExtSelectors(entries: string[]): {
250
254
  * snapshotted. `registerTool` writes into the very `extension.tools` maps this reads,
251
255
  * so `inScope()` sees late arrivals on the next call.
252
256
  *
253
- * Two enforcement points, because neither covers the whole picture:
257
+ * The active set and call-time checks cover different parts of scope:
254
258
  *
255
259
  * - `turn_end` re-narrows the ACTIVE set. pi emits `turn_end` immediately before
256
260
  * `prepareNextTurn` re-snapshots `agent.state.tools`, and session listeners run
257
261
  * synchronously, so the narrow lands in time for turns 2..N.
258
- * - `beforeToolCall` blocks out-of-scope calls. Turn 1 cannot be narrowed at all:
262
+ * - The loader's bound `tool_call` handler blocks direct and nested calls.
263
+ * Turn 1 cannot be narrowed at all:
259
264
  * `before_agent_start` fires INSIDE `prompt()` and may widen the tool set, but
260
265
  * `createContextSnapshot()` freezes that turn's tools immediately after — there
261
266
  * is no hook in between. A call-time check is the only correct guard there.
@@ -285,7 +290,7 @@ export function installExtensionToolScope(
285
290
  */
286
291
  readmitToolNames: Set<string>;
287
292
  },
288
- ): void {
293
+ ): (toolName: string) => boolean {
289
294
  const { loader, toolNames, disallowedSet, extNames, narrowing, readmitToolNames } = ctx;
290
295
 
291
296
  // The names allowed right now. Mirrors the `ext:` opt-in flip: when any `ext:`
@@ -309,7 +314,7 @@ export function installExtensionToolScope(
309
314
  }
310
315
  for (const name of EXCLUDED_TOOL_NAMES) keep.delete(name);
311
316
  // Injected tools are legitimately active for this agent — re-admit them so
312
- // the renarrow keeps them in the active set and beforeToolCall doesn't
317
+ // the renarrow keeps them in the active set and the call-time guard doesn't
313
318
  // block them. Already vetted against `disallowed_tools` by the caller,
314
319
  // which is the only place that knows which kind may be taken back.
315
320
  for (const name of readmitToolNames) keep.add(name);
@@ -318,8 +323,14 @@ export function installExtensionToolScope(
318
323
 
319
324
  const renarrow = () => {
320
325
  const allowed = inScope();
321
- const next = session.getAllTools().map((t) => t.name).filter((n) => allowed.has(n));
322
326
  const current = session.getActiveToolNames();
327
+ // Keep deferred/codemode tools callable without promoting them into model
328
+ // declarations. Explicit activation by an extension is preserved in scope.
329
+ const next = session.getAllTools().filter((tool) =>
330
+ allowed.has(tool.name) && (
331
+ tool.exposure === "direct" || tool.exposure === "model-only" || current.includes(tool.name)
332
+ ),
333
+ ).map((tool) => tool.name);
323
334
  // setActiveToolsByName unconditionally rebuilds the system prompt, so skip
324
335
  // the no-op that steady-state turns would otherwise pay for every turn.
325
336
  if (next.length !== current.length || next.some((n, i) => n !== current[i])) {
@@ -335,16 +346,7 @@ export function installExtensionToolScope(
335
346
  if (event.type === "turn_end") renarrow();
336
347
  });
337
348
 
338
- const priorBeforeToolCall = session.agent.beforeToolCall;
339
- session.agent.beforeToolCall = async (context, signal) => {
340
- if (!inScope().has(context.toolCall.name)) {
341
- return {
342
- block: true,
343
- reason: `Tool "${context.toolCall.name}" is not available to this subagent.`,
344
- };
345
- }
346
- return priorBeforeToolCall?.(context, signal);
347
- };
349
+ return (toolName) => inScope().has(toolName);
348
350
  }
349
351
 
350
352
  /** Default max turns. undefined = unlimited (no turn limit). */
@@ -768,6 +770,7 @@ export async function runAgent(
768
770
  // must compare against this, not the surviving set (absence from survivors is
769
771
  // an exclude *succeeding*).
770
772
  let discoveredNames: Set<string> | undefined;
773
+ let toolInScope: ((toolName: string) => boolean) | undefined;
771
774
  const extensionsOverride: ((base: LoadExtensionsResult) => LoadExtensionsResult) | undefined =
772
775
  noExtensions || (loadAll && !hasExcludes)
773
776
  ? undefined
@@ -776,6 +779,7 @@ export async function runAgent(
776
779
  return {
777
780
  ...base,
778
781
  extensions: base.extensions.filter((e) => {
782
+ if (e.path === "<inline:subagent-tool-scope>") return true;
779
783
  const canons = extensionCanonicalNames(e.path);
780
784
  if (canons.some((n) => excludeNames.has(n))) return false; // exclude wins
781
785
  return loadAll || canons.some((n) => keepNames.has(n));
@@ -788,6 +792,19 @@ export async function runAgent(
788
792
  agentDir,
789
793
  noExtensions,
790
794
  additionalExtensionPaths,
795
+ // Pi's nested ctx.executeTool calls dispatch tool_call directly, and prompt
796
+ // setup can replace agent.beforeToolCall. Enforce scope in that dispatch too.
797
+ extensionFactories: noExtensions ? [] : [{
798
+ name: "subagent-tool-scope",
799
+ hidden: true,
800
+ factory: (pi) => {
801
+ pi.on("tool_call", (event) => {
802
+ if (toolInScope && !toolInScope(event.toolName)) {
803
+ return { block: true, reason: `Tool "${event.toolName}" is not available to this subagent.` };
804
+ }
805
+ });
806
+ },
807
+ }],
791
808
  extensionsOverride,
792
809
  noSkills,
793
810
  noPromptTemplates: true,
@@ -1014,24 +1031,14 @@ export async function runAgent(
1014
1031
  })
1015
1032
  : SessionManager.inMemory(effectiveCwd);
1016
1033
 
1017
- // Pi 0.80.8 replaced createAgentSession's modelRegistry option with
1018
- // modelRuntime, but ExtensionContext still exposes only the registry facade.
1019
- // Pass both so the full supported Pi range retains the parent's providers.
1020
- const parentModelRuntime = (ctx.modelRegistry as unknown as { runtime?: unknown }).runtime;
1021
- const sessionOpts: Parameters<typeof createAgentSession>[0] & {
1022
- modelRegistry: ExtensionContext["modelRegistry"];
1023
- modelRuntime?: unknown;
1024
- } = {
1034
+ // ExtensionContext exposes the registry facade; its runtime owns provider auth.
1035
+ const parentModelRuntime = (ctx.modelRegistry as unknown as { runtime?: ModelRuntime }).runtime;
1036
+ const sessionOpts: Parameters<typeof createAgentSession>[0] = {
1025
1037
  cwd: effectiveCwd,
1026
1038
  agentDir,
1027
1039
  sessionManager,
1028
1040
  settingsManager,
1029
- modelRegistry: ctx.modelRegistry,
1030
- // `as never` is what keeps this assignable across the supported Pi range:
1031
- // pre-0.80.8 the field exists only via the `modelRuntime?: unknown` shim
1032
- // above, while newer Pi types it as `ModelRuntime` — a shape an opaque
1033
- // `unknown` read off the private facade field can never satisfy.
1034
- ...(parentModelRuntime !== undefined && { modelRuntime: parentModelRuntime as never }),
1041
+ modelRuntime: parentModelRuntime,
1035
1042
  model,
1036
1043
  tools: sessionTools,
1037
1044
  customTools: [...nestedTools, ...structuredTools],
@@ -1073,7 +1080,7 @@ export async function runAgent(
1073
1080
  // handled below by re-deriving scope from the loader's live extension maps —
1074
1081
  // `registerTool` writes into those same maps, so late arrivals are judged too.
1075
1082
  if (!noExtensions) {
1076
- installExtensionToolScope(session, {
1083
+ toolInScope = installExtensionToolScope(session, {
1077
1084
  loader,
1078
1085
  toolNames,
1079
1086
  disallowedSet,
@@ -297,7 +297,7 @@ function parseMemory(val: unknown): MemoryScope | undefined {
297
297
 
298
298
  /** Parse the OpenAI Responses/Codex `service_tier` frontmatter field. */
299
299
  function parseServiceTier(val: unknown): ServiceTier | undefined {
300
- if (val === "auto" || val === "default" || val === "flex" || val === "priority" || val === "scale") {
300
+ if (val === "auto" || val === "default" || val === "flex" || val === "fast" || val === "priority" || val === "scale") {
301
301
  return val;
302
302
  }
303
303
  return undefined;
@@ -24,20 +24,6 @@
24
24
  * conversation with nothing in it yet clones to nothing in it yet, which is the
25
25
  * correct answer rather than a failure.
26
26
  *
27
- * It is also the oldest of the equivalent Pi APIs — `buildContextEntries` on
28
- * ReadonlySessionManager and the `sessionEntryToContextMessages` export both
29
- * arrived in 0.80.5 — where this one has been exported unchanged from before
30
- * the declared peer floor, and is the same code path (`byId` is only an index
31
- * cache, so passing it or not cannot change the result). Keeping the floor
32
- * honest costs nothing here: see the `compat-floor-pi` job.
33
- *
34
- * Its `thinkingLevel` is NOT used, and is the one place the newer API would be
35
- * better. `getSessionContextSettings` starts at "off" and moves only on an
36
- * explicit `thinking_level_change` entry, so a session where nobody ran
37
- * `/think` reports "off" rather than the level it is really using. Omitting the
38
- * field instead lets `createAgentSession` resolve it from settings, which is
39
- * that real level.
40
- *
41
27
  * Three details make the spawn belong to the real session rather than the
42
28
  * clone:
43
29
  *
@@ -64,14 +50,18 @@
64
50
  import type { Model } from "@earendil-works/pi-ai";
65
51
  import {
66
52
  buildSessionContext,
53
+ convertToLlm,
67
54
  createAgentSession,
55
+ DefaultResourceLoader,
68
56
  type ExtensionContext,
57
+ getAgentDir,
58
+ type ModelRuntime,
69
59
  SessionManager,
70
60
  type ToolDefinition,
71
61
  } from "@earendil-works/pi-coding-agent";
72
62
  import { runInChildSessionContext } from "./child-context.js";
73
63
  import { agentMentionReminder } from "./mention.js";
74
- import type { SubagentType, ThinkingLevel } from "./types.js";
64
+ import type { SubagentType } from "./types.js";
75
65
 
76
66
  export interface MentionCloneOptions {
77
67
  /** The MAIN session's context — what the spawn is attributed to, and the
@@ -125,37 +115,54 @@ export async function runMentionClone(opts: MentionCloneOptions): Promise<Mentio
125
115
  { ...(params as Record<string, unknown>), run_in_background: true } as typeof params,
126
116
  signal,
127
117
  onUpdate,
128
- ctx,
118
+ { ..._cloneCtx, ...ctx },
129
119
  );
130
120
  },
131
121
  };
132
122
 
133
123
  let session: Awaited<ReturnType<typeof createAgentSession>>["session"] | undefined;
134
124
  try {
135
- // Pi 0.80.8 moved createAgentSession from modelRegistry to modelRuntime;
136
- // agent-runner.ts carries the same shim for the same reason — pass both so
137
- // the clone keeps the parent's providers across the supported range.
138
- const parentModelRuntime = (ctx.modelRegistry as unknown as { runtime?: unknown }).runtime;
125
+ // The registry facade retains the parent's configured runtime and auth.
126
+ const parentModelRuntime = (ctx.modelRegistry as unknown as { runtime?: ModelRuntime }).runtime;
139
127
  // The conversation as the main session resolves it: compaction applied,
140
128
  // branch summaries substituted.
141
129
  const conversation = buildSessionContext(
142
130
  ctx.sessionManager.getEntries(),
143
131
  ctx.sessionManager.getLeafId(),
144
132
  );
145
- // Pi 0.82.0 added this; below it the field is absent and the clone takes
146
- // the settings level instead, which is what a session that never ran
147
- // `/think` is on anyway. Same shim shape as `modelRuntime` below.
148
- const thinkingLevel = (ctx as { thinkingLevel?: ThinkingLevel }).thinkingLevel;
133
+ const sessionManager = SessionManager.inMemory(ctx.cwd);
134
+ // Replay conversation turns through the session manager. Historical system
135
+ // messages contain the parent's tool declarations, which the clone must not inherit.
136
+ for (const entry of conversation.messages) {
137
+ if (entry.role === "system") continue;
138
+ if (entry.role === "branchSummary" || entry.role === "compactionSummary") {
139
+ for (const message of convertToLlm([entry])) sessionManager.appendMessage(message);
140
+ } else {
141
+ sessionManager.appendMessage(entry);
142
+ }
143
+ }
144
+ const resourceLoader = new DefaultResourceLoader({
145
+ cwd: ctx.cwd,
146
+ agentDir: getAgentDir(),
147
+ noExtensions: true,
148
+ noSkills: true,
149
+ noPromptTemplates: true,
150
+ noThemes: true,
151
+ noContextFiles: true,
152
+ systemPromptOverride: () => ctx.getSystemPrompt(),
153
+ appendSystemPromptOverride: () => [],
154
+ });
155
+ await resourceLoader.reload();
149
156
  const created = await runInChildSessionContext(() =>
150
157
  createAgentSession({
151
158
  cwd: ctx.cwd,
152
159
  // Nothing about the copy is worth persisting, and an in-memory manager
153
160
  // is also what keeps the real session untouched.
154
- sessionManager: SessionManager.inMemory(ctx.cwd),
161
+ sessionManager,
162
+ resourceLoader,
155
163
  model: ctx.model as Model<never> | undefined,
156
- ...(thinkingLevel && { thinkingLevel }),
157
- modelRegistry: ctx.modelRegistry,
158
- ...(parentModelRuntime !== undefined && { modelRuntime: parentModelRuntime as never }),
164
+ ...(ctx.thinkingLevel && { thinkingLevel: ctx.thinkingLevel }),
165
+ modelRuntime: parentModelRuntime,
159
166
  // An allowlist naming exactly the clone's own tool. NOT `noTools:
160
167
  // "all"`, whose doc comment ("start with no tools enabled") reads like
161
168
  // it spares custom tools and does not: it resolves to an EMPTY
@@ -166,21 +173,10 @@ export async function runMentionClone(opts: MentionCloneOptions): Promise<Mentio
166
173
  // agent-runner's `tools: sessionTools` beside its nested `customTools`.
167
174
  tools: [cloneAgentTool.name],
168
175
  customTools: [cloneAgentTool],
169
- } as Parameters<typeof createAgentSession>[0]),
176
+ }),
170
177
  );
171
178
  session = created.session;
172
179
 
173
- // The clone rebuilds a system prompt from cwd and agentDir, which is close
174
- // but not the live one — extensions contribute to it per turn. Copy the
175
- // real thing, so the copy reasons under the instructions the user's model
176
- // is actually working under.
177
- const systemPrompt = ctx.getSystemPrompt?.();
178
- if (systemPrompt) session.agent.state.systemPrompt = systemPrompt;
179
-
180
- // The conversation itself. Pushed rather than assigned so the array the
181
- // session was built around stays the one it goes on using.
182
- session.agent.state.messages.push(...conversation.messages);
183
-
184
180
  // User text first, reminder after — the order Claude Code's attachment
185
181
  // renderer produces, where the reminder trails the message it is about.
186
182
  await session.prompt(`${message}\n\n${agentMentionReminder(type)}`);
@@ -104,21 +104,27 @@ export function streamToOutputFile(
104
104
  cwd: string,
105
105
  startIndex?: number,
106
106
  ): () => void {
107
- // Index of the first message this stream is responsible for. A spawn writes
108
- // messages[0] as the initial prompt entry, so it starts at 1. A resume hands
107
+ // A spawn writes its initial user prompt separately. Pi can project system
108
+ // messages before that prompt, so skip the first user message by role. A resume hands
109
109
  // in the session's length as of just before the run: the session already
110
110
  // holds every prior turn, and re-emitting those would duplicate history that
111
111
  // is already in the file.
112
- let writtenCount = startIndex ?? 1;
112
+ let writtenCount = startIndex ?? 0;
113
+ let skipInitialUser = startIndex === undefined;
113
114
 
114
115
  const flush = () => {
115
116
  const messages = session.messages;
116
117
  while (writtenCount < messages.length) {
117
118
  const msg = messages[writtenCount];
119
+ if (skipInitialUser && msg.role === "user") {
120
+ skipInitialUser = false;
121
+ writtenCount++;
122
+ continue;
123
+ }
118
124
  const entry = {
119
125
  isSidechain: true,
120
126
  agentId,
121
- type: msg.role === "assistant" ? "assistant" : msg.role === "user" ? "user" : "toolResult",
127
+ type: msg.role === "assistant" ? "assistant" : msg.role === "user" ? "user" : msg.role === "system" ? "system" : "toolResult",
122
128
  message: msg,
123
129
  timestamp: new Date().toISOString(),
124
130
  cwd,
package/src/types.ts CHANGED
@@ -18,7 +18,7 @@ export const DEFAULT_AGENT_NAMES = ["general-purpose", "Explore", "Plan"] as con
18
18
  export type MemoryScope = "user" | "project" | "local";
19
19
 
20
20
  /** OpenAI Responses/Codex request processing tier. */
21
- export type ServiceTier = "auto" | "default" | "flex" | "priority" | "scale";
21
+ export type ServiceTier = "auto" | "default" | "flex" | "fast" | "priority" | "scale";
22
22
 
23
23
  /**
24
24
  * Isolation mode for agent execution.
@@ -0,0 +1,47 @@
1
+ import { defineConfig } from "vitest/config";
2
+
3
+ export default defineConfig({
4
+ // The print-mode e2e suite (test/subagents-print-mode-e2e.test.ts) drives REAL
5
+ // faux-model turns through pi-coding-agent + pi-agent-core. That requires ONE
6
+ // shared @earendil-works/pi-ai instance so the faux provider the test registers
7
+ // lands in the same api-registry the session streams through. npm physically
8
+ // duplicates pi-ai (a top-level copy and one nested under pi-coding-agent), which
9
+ // otherwise yields two registries and "No API provider registered" errors.
10
+ // Inlining the @earendil-works packages routes them through Vite's resolver so
11
+ // dedupe can collapse pi-ai to a single instance — for the parent AND for every
12
+ // subagent session the extension spawns. dedupe alone is insufficient (it only
13
+ // affects modules Vite resolves; without inline the runtime stays externalized).
14
+ test: {
15
+ server: { deps: { inline: [/@earendil-works\/pi-/] } },
16
+ // Local reporting only — deliberately no `thresholds`, and not wired into
17
+ // CI. src/index.ts is mostly the /agents wizard, which is TUI flow with
18
+ // almost no logic and is not worth a fake-TUI harness; any global floor
19
+ // would therefore either sit below what the rest of the suite achieves
20
+ // (and ratchet down as tests are deleted) or force exactly that harness.
21
+ // Coverage here measures lines touched, not behavior pinned.
22
+ //
23
+ // `istanbul`, NOT the `v8` default. v8 under-reports badly on this suite:
24
+ // `src/env.ts` came out at 54.54% for the full run but 90.9% for
25
+ // `vitest run --coverage test/env.test.ts` — coverage cannot fall as more
26
+ // tests run. Bisecting pins it on test/subagents-print-mode-e2e.test.ts,
27
+ // which boots a REAL pi session; pi's own resource loader re-loads the
28
+ // extension and its imports, so the same source file appears to v8 twice —
29
+ // once executed, once barely — and the merge takes the wrong one rather
30
+ // than the union. The damage was not confined to env.ts: v8 also reported
31
+ // module-level `const` declarations in src/settings.ts as uncovered, which
32
+ // is impossible since they run on import.
33
+ //
34
+ // istanbul instruments at transform time and accumulates per source path,
35
+ // so a second load adds to the same counters instead of shadowing them.
36
+ // Same suite, same run: v8 68.6% vs istanbul 76.5% statements, and the
37
+ // per-file numbers now match what a single-file run reports. If you switch
38
+ // this back to v8, re-check env.ts against `--coverage test/env.test.ts`
39
+ // before trusting anything the table says.
40
+ coverage: {
41
+ provider: "istanbul",
42
+ reporter: ["text", "html"],
43
+ include: ["src/**/*.ts"],
44
+ },
45
+ },
46
+ resolve: { dedupe: ["@earendil-works/pi-ai"] },
47
+ });