pi-quiver 3.3.2 → 3.4.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -8,6 +8,20 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
8
8
  via OIDC trusted publishing. The release helper at
9
9
  `.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
10
10
 
11
+ ## v3.4.1 - 2026-08-04
12
+
13
+ - **`fast-mode` now corrects reported `usage.cost` to true 2x fast pricing for Opus 4.8/5** (statusline, persisted cost, and pi-cohort `Σ$` per load order) (#6).
14
+ - **Documented Anthropic stall-retry activation.** A session trace confirmed Pi's existing 2/4/8-second exponential backoff was active and that the watchdog retained the retry budget resolved when its extension instance started. Local `retry.maxRetries: 5` now takes effect after a fresh session or explicit reload, with no watchdog cooldown, resubmission path, or runtime code change.
15
+
16
+ ## v3.4.0 - 2026-07-30
17
+
18
+ - **`provider-stall-watchdog`: new `firstEventMs` tier (default 20s) fast-fails a request that never streams.** A pre-first-event deadline is armed at every provider request and cleared by the first assistant `message_start`; on expiry the request is aborted and converted to a retryable error, cutting recovery for an unresponsive request from ~240s to ~22s (detection + Pi's 2s backoff). Previously only the 240s mid-stream `recoveryMs` timer caught this case.
19
+ - **Widened activation.** The `firstEventMs` tier arms in every mode (`tui`, `print`, `json`, `rpc`) and from every origin, including extension-triggered turns that never emit `before_agent_start`; activation is now lazy inside `before_provider_request`, so the `input` and `before_agent_start` handlers are gone. The mid-stream `warningMs`/`recoveryMs` tier stays TUI-only - an unattended abort discards billed output tokens with nobody reading the warning. Headless runs report on stderr via `console.warn` (never stdout, which `json` mode uses for its protocol).
20
+ - **Behaviour change: `maxStallRetries: 0` is now valid** and means "detect and stop, never auto-retry" - it previously failed validation and disabled the extension. When unset, `maxStallRetries` inherits the layered `retry.maxRetries` including an explicit `0`, matching Pi's own `?? 3` instead of silently resolving 0 to 3. Both tiers share the one budget.
21
+ - **Abort-hang escalation.** After any watchdog abort a fixed 10s deadline announces that the provider connection is unresponsive if the request never terminates; any post-abort stream event re-arms it, so a straggler followed by a wedge still escalates. It reduces the hang, it cannot force the provider to stop - undici's `headersTimeout`/`bodyTimeout` (pi's `httpIdleTimeoutMs`, default 300s) stay the backstop.
22
+ - **Config is read once per session,** on the first provider request, instead of per prompt. Editing `settings.json` mid-session - including repairing an invalid block that disabled the extension - now requires a session restart.
23
+ - **`"providerStallWatchdog": false` now works.** The boolean shorthand the README already documented was rejected as "must be an object"; it is now accepted as `{ "enabled": <bool> }`, matching `swordHeader`, `fastMode`, and `sessionAutoName`. Previously harmless (the warning was TUI-only), the widened activation above would otherwise have printed a config error to stderr on every headless run.
24
+
11
25
  ## v3.3.2 - 2026-07-30
12
26
 
13
27
  - **`fast-mode` supports Claude Opus 5.** Added `claude-opus-5` to the model prefix allowlist (fast mode API mechanics are identical to Opus 4.8: `speed: "fast"` + `fast-mode-2026-02-01` beta; compat verified identical in pi-ai 0.82.1, so `buildBetaHeader`'s OAuth-only preservation still holds).
package/README.md CHANGED
@@ -67,7 +67,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
67
67
  | `session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, with Ghostty tab rename. OFF by default. |
68
68
  | `sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
69
69
  | `fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
70
- | `provider-stall-watchdog.ts` | - | Opt-in semantic-silence watchdog for human interactive TUI runs. Warns after 2 minutes and recovers after 4 minutes; policy D offers each stall to Pi's retry loop until the stall retry budget (`maxStallRetries`, default = `retry.maxRetries`) is exhausted. OFF by default. |
70
+ | `provider-stall-watchdog.ts` | - | Opt-in provider-stall recovery, in two tiers: a pre-first-event deadline (`firstEventMs`, 20s) on every provider request in every mode, and the mid-stream pair (warn at 2 min, recover at 4 min) in TUI runs only. Policy D offers each stall to Pi's retry loop until the stall retry budget (`maxStallRetries`, default = `retry.maxRetries`) is exhausted. OFF by default. |
71
71
 
72
72
  Full routing rules, size-gate mechanics, and config: [doc/fetch.md](doc/fetch.md), [doc/doc-to-md.md](doc/doc-to-md.md).
73
73
 
@@ -79,20 +79,20 @@ Full routing rules, size-gate mechanics, and config: [doc/fetch.md](doc/fetch.md
79
79
  | Content routing | HTML -> Markdown, binary -> untouched file, GitHub URLs -> `gh` CLI, everything else -> the size gate. |
80
80
  | Graceful degradation | Optional binaries (`gh`, `uv`, LibreOffice) are never hard install-time deps; each has a defined, documented fallback or failure mode. |
81
81
  | Opt-in extensions | `session-name`, `sword-header`, `fast-mode`, and `provider-stall-watchdog` do nothing until explicitly enabled in `settings.json`. |
82
- | Provider stall recovery | The watchdog detects missing parsed semantic progress, not network liveness. It is limited to confirmed human interactive TUI runs. |
82
+ | Provider stall recovery | The watchdog detects a missing first stream event and missing parsed semantic progress, not network liveness. The pre-first-event tier covers every mode and origin; the mid-stream tier is TUI-only. |
83
83
 
84
84
  ## When to use
85
85
 
86
86
  - An agent needs to reason from a real web page, GitHub issue/PR, or local PDF/DOCX/PPTX instead of memory.
87
87
  - You want that ingestion to be safe by default, with no risk of a single call blowing the context budget.
88
- - A human interactive Pi session needs an opt-in guard against providers that stop making semantic progress.
88
+ - A Pi run needs an opt-in guard against provider requests that never produce a first stream event, plus mid-stream silence recovery in interactive TUI sessions.
89
89
 
90
90
  ## When NOT to use
91
91
 
92
92
  - You need a general-purpose web scraper (JS-rendered pages, pagination, auth flows) - `fetch` does plain HTTP + Readability extraction, nothing more.
93
93
  - You need spreadsheet conversion - `doc_to_md` explicitly excludes spreadsheets (they paginate badly via PDF).
94
94
  - You want automatic session naming, a custom header, fast mode, or stall recovery without opting in - all stay off until you flip the config.
95
- - You need watchdog behavior in JSON, RPC, print, or subagent runs - activation excludes them by input origin and mode, not environment or session lineage.
95
+ - You need *mid-stream* stall recovery in JSON, RPC, or print runs - only the pre-first-event tier arms there; mid-stream silence falls through to pi's transport timeout.
96
96
 
97
97
  ## Install
98
98
 
@@ -150,6 +150,8 @@ These extensions are opt-in via `settings.json` (project `.pi/settings.json` ove
150
150
 
151
151
  `sessionAutoName.enabled` makes one extra short LLM call per session (once, after the first turn) to title it; `false` (default) makes no model calls. `fastMode` only affects `claude-opus-4-8` and `claude-opus-5` requests on Anthropic's `anthropic-messages` API; enabling it opts into premium fast-mode pricing. `--fast` forces it on for one launch; `/fast on|off` toggles live. Proxy providers (opencode, cloudflare-ai-gateway) are excluded. `fastMode`'s header injection needs the `before_provider_headers` hook (pi bundling `@earendil-works/pi-coding-agent` >= 0.80.5); on older pi the beta header is silently not sent. See [doc/fetch.md](doc/fetch.md) and [doc/doc-to-md.md](doc/doc-to-md.md) for the ingestion tools' full reference; session-name/sword-header behavior above is complete.
152
152
 
153
+ `pi-ai` prices every fast request at standard rates - it has no `usage.speed` support and no request-level pricing modifier - so `fastMode` corrects the reported cost itself: a `message_end` handler scales all four `usage.cost` components by `FAST_MODE_COST_MULTIPLIER` (2x) and returns the corrected message. Persisted session JSONL and pi's own native cost display are always exact, since they're written from this corrected message. pi-cohort's live `Σ$` reflects the correction only when pi-quiver's `message_end` handler runs before pi-cohort's - best-effort, depending on extension load order - and is reconciled on pi-cohort's next `session_start` regardless. The upstream fix (teaching `pi-ai`'s `Usage`/`calculateCost` about `usage.speed`) is the better long-term path and is tracked separately.
154
+
153
155
  Recommended explicit retry and watchdog settings:
154
156
 
155
157
  ```json
@@ -161,6 +163,7 @@ Recommended explicit retry and watchdog settings:
161
163
  },
162
164
  "providerStallWatchdog": {
163
165
  "enabled": true,
166
+ "firstEventMs": 20000,
164
167
  "warningMs": 120000,
165
168
  "recoveryMs": 240000,
166
169
  "maxStallRetries": 3
@@ -168,7 +171,30 @@ Recommended explicit retry and watchdog settings:
168
171
  }
169
172
  ```
170
173
 
171
- `providerStallWatchdog` is OFF by default and runs only for confirmed human interactive TUI runs. JSON, RPC, print, and subagent runs are excluded by activation, not environment or session lineage. Verified with Pi 0.80.10: each semantic stall is aborted and offered to Pi retry until `maxStallRetries` conversions are used; further stalls stop for manual resubmission. `maxStallRetries` defaults to the layered `retry.maxRetries` (Pi default 3); consecutive stall conversions consume Pi retry attempts without a success reset in between, so keep `maxStallRetries <= retry.maxRetries`. A successful assistant turn resets the stall counter (mirroring Pi's own retry counter). The silence warning and all watchdog notices render as main-window notifications, not the bottom status line. Automatic continuation needs enabled Pi retry with remaining capacity. Disabled, exhausted, or incompatible retry degrades to manual resubmission. Pending steering or follow-ups return to the editor and are excluded from automatic continuation. Invalid merged watchdog config fails closed.
174
+ | Key | Default | Where it arms | What it measures |
175
+ | --- | --- | --- | --- |
176
+ | `enabled` | `false` | - | Master switch. OFF means the extension does nothing. |
177
+ | `firstEventMs` | `20000` | every provider request, every mode (`tui`/`print`/`json`/`rpc`), every origin | Silence between the request and the first assistant `message_start`. |
178
+ | `warningMs` | `120000` | mid-stream, `ctx.mode === "tui"` only | Silence since the last non-empty text/thinking/toolcall delta; notifies. |
179
+ | `recoveryMs` | `240000` | mid-stream, `ctx.mode === "tui"` only | Same clock; aborts and converts. Must be `> warningMs`. |
180
+ | `maxStallRetries` | layered `retry.maxRetries`, else `3` | shared by both tiers | Watchdog aborts that may convert to a retryable error before stopping. |
181
+
182
+ `providerStallWatchdog` is OFF by default. Once enabled it arms in two tiers per provider request:
183
+
184
+ - **Pre-first-event (`firstEventMs`).** Armed at every provider request, in every mode and from every origin - including extension-triggered turns that never emit `before_agent_start` - and cleared by the first assistant `message_start`. On expiry the request is aborted and, budget permitting, converted to a retryable error, so an unresponsive request recovers in ~22s (20s detection + Pi's 2s backoff) instead of the ~240s it took when only the mid-stream tier existed.
185
+ - **Mid-stream (`warningMs` / `recoveryMs`).** Armed from the first assistant `message_start` onward, and only when `ctx.mode === "tui"`. Aborting mid-generation discards billed output tokens and an unattended run has nobody to read the warning, so headless mid-stream silence deliberately falls through to the transport timeout instead.
186
+
187
+ **Raise `firstEventMs` if your provider is legitimately slow to first event.** Queueing gateways, throttled endpoints, and busy single-slot local model servers can hold the connection for well over 20s before their first stream event; every false abort re-uploads the whole context and spends one stall retry.
188
+
189
+ **Leave pi's own `httpIdleTimeoutMs` (default `300000`) at its default.** It is the transport backstop, and a single value drives undici's `headersTimeout` *and* `bodyTimeout` - lowering it to get fast pre-stream failure also truncates legitimate mid-stream gaps. `firstEventMs` is the knob for pre-stream silence.
190
+
191
+ Verified with Pi 0.80.10: each stall is aborted and offered to Pi retry until `maxStallRetries` conversions are used; further stalls stop for manual resubmission. Both tiers draw on that one budget. `maxStallRetries` defaults to the layered `retry.maxRetries` (Pi default 3, an explicit `0` honoured); `0` is valid and means "detect and stop, never auto-retry". Consecutive stall conversions consume Pi retry attempts without a success reset in between, so keep `maxStallRetries <= retry.maxRetries`. A successful assistant turn resets the stall counter (mirroring Pi's own retry counter). Automatic continuation needs enabled Pi retry with remaining capacity. Disabled, exhausted, or incompatible retry degrades to manual resubmission. Pending steering or follow-ups return to the editor and are excluded from automatic continuation. Invalid merged watchdog config fails closed.
192
+
193
+ Operational notes:
194
+
195
+ - **Settings are read once per session,** on the first provider request. Editing `settings.json` mid-session changes nothing until you restart the session - that includes repairing an invalid block that already disabled the extension.
196
+ - **A watchdog abort that the provider ignores escalates after a fixed 10s.** Any post-abort stream event re-arms that deadline (bytes prove only that the connection was alive at that instant), so a stream that emits a straggler and then wedges still escalates 10s after its last event. This reduces the hang; it cannot force the provider to stop, and undici's timeouts remain the final backstop.
197
+ - **Headless runs report on stderr.** In `print`/`json` mode pi binds a no-op UI, so watchdog notices go out via `console.warn`. Nothing is ever written to stdout, which `json` mode uses for its protocol. In TUI and RPC the notices render as main-window notifications, not the bottom status line.
172
198
 
173
199
  ## Development
174
200
 
package/fast-mode.ts CHANGED
@@ -22,6 +22,7 @@
22
22
  */
23
23
 
24
24
  import type { ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
25
+ import type { Usage } from "@earendil-works/pi-ai";
25
26
  import { resolveConfig } from "./extension-config.ts";
26
27
 
27
28
  export const FAST_MODE_BETA = "fast-mode-2026-02-01";
@@ -29,6 +30,21 @@ export const FAST_SPEED = "fast";
29
30
  // Loose prefixes: match dated snapshots (claude-opus-4-8-*, claude-opus-5-*).
30
31
  // Opus 4.7 is out of scope (D1); a future model needs a one-line addition here.
31
32
  export const FAST_MODE_MODEL_PREFIXES = ["claude-opus-4-8", "claude-opus-5"];
33
+
34
+ // Anthropic bills Opus 4.8/5 fast mode at exactly 2x standard rates; input,
35
+ // output, and cache read/write all scale by the same factor because caching
36
+ // multipliers "apply on top of fast mode pricing". Rate card:
37
+ // https://docs.claude.com/en/docs/build-with-claude/fast-mode (retrieved 2026-08-04).
38
+ export const FAST_MODE_COST_MULTIPLIER = 2;
39
+
40
+ export function scaleCost(cost: Usage["cost"], multiplier: number): Usage["cost"] {
41
+ const input = cost.input * multiplier;
42
+ const output = cost.output * multiplier;
43
+ const cacheRead = cost.cacheRead * multiplier;
44
+ const cacheWrite = cost.cacheWrite * multiplier;
45
+ return { input, output, cacheRead, cacheWrite, total: input + output + cacheRead + cacheWrite };
46
+ }
47
+
32
48
  export const OAUTH_IDENTITY_BETAS = ["claude-code-20250219", "oauth-2025-04-20"];
33
49
  const STATUS_KEY = "fast-mode";
34
50
  const BETA_HEADER = "anthropic-beta";
@@ -92,6 +108,10 @@ export function resolveEnabled(s: State): boolean {
92
108
  export default function (pi: ExtensionAPI) {
93
109
  let liveOverride: boolean | null = null;
94
110
  let enabled = false;
111
+ // Per-request snapshot. Safe as plain booleans because pi serializes provider
112
+ // requests per session (headers -> request -> message_end, awaited in order).
113
+ let pendingFastSpeed = false;
114
+ let pendingFastHeader = false;
95
115
 
96
116
  const readFlag = (): boolean => pi.getFlag("fast") === true;
97
117
 
@@ -135,16 +155,43 @@ export default function (pi: ExtensionAPI) {
135
155
  });
136
156
 
137
157
  pi.on("before_provider_request", (event, ctx) => {
138
- if (!shouldInject(enabled, ctx.model)) return;
139
- return injectSpeed(event.payload);
158
+ if (!shouldInject(enabled, ctx.model)) {
159
+ pendingFastSpeed = false;
160
+ return;
161
+ }
162
+ const next = injectSpeed(event.payload);
163
+ pendingFastSpeed =
164
+ typeof next === "object" && next !== null && (next as Record<string, unknown>).speed === FAST_SPEED;
165
+ return next;
140
166
  });
141
167
 
142
168
  pi.on("before_provider_headers", async (event, ctx) => {
143
- if (!shouldInject(enabled, ctx.model)) return;
144
- if (!event.headers) return;
169
+ if (!shouldInject(enabled, ctx.model) || !event.headers) {
170
+ pendingFastHeader = false;
171
+ return;
172
+ }
145
173
  const isOAuth = await detectOAuth(ctx);
146
- if (isOAuth === null) return;
174
+ if (isOAuth === null) {
175
+ pendingFastHeader = false;
176
+ return;
177
+ }
147
178
  event.headers[BETA_HEADER] = buildBetaHeader(event.headers[BETA_HEADER], isOAuth);
179
+ pendingFastHeader = true;
180
+ });
181
+
182
+ pi.on("message_end", (event) => {
183
+ const msg = event.message;
184
+ if (msg.role !== "assistant") return;
185
+ const corrected = pendingFastSpeed && pendingFastHeader && !!msg.usage?.cost;
186
+ pendingFastSpeed = false;
187
+ pendingFastHeader = false;
188
+ if (!corrected) return;
189
+ return {
190
+ message: {
191
+ ...msg,
192
+ usage: { ...msg.usage, cost: scaleCost(msg.usage.cost, FAST_MODE_COST_MULTIPLIER) },
193
+ },
194
+ };
148
195
  });
149
196
 
150
197
  pi.registerCommand("fast", {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-quiver",
3
- "version": "3.3.2",
3
+ "version": "3.4.1",
4
4
  "description": "Personal pack of Pi coding-agent extensions: context-safe fetch, doc_to_md PDF/DOCX/PPTX-to-Markdown conversion, session naming, a themed ASCII startup header, Opus 4.8 fast mode, and a provider-stall watchdog.",
5
5
  "author": "Jacek Juraszek",
6
6
  "license": "MIT",
@@ -4,12 +4,14 @@ import { readSettings, resolveConfig, settingsPaths } from "./extension-config.t
4
4
  export const MAX_TIMER_MS = 2_147_483_647;
5
5
  export const DEFAULT_CONFIG = {
6
6
  enabled: false,
7
+ firstEventMs: 20_000,
7
8
  warningMs: 120_000,
8
9
  recoveryMs: 240_000,
9
10
  } as const;
10
11
 
11
12
  export type WatchdogConfig = {
12
13
  enabled: boolean;
14
+ firstEventMs: number;
13
15
  warningMs: number;
14
16
  recoveryMs: number;
15
17
  maxStallRetries: number;
@@ -24,6 +26,7 @@ export type WatchdogRuntime = {
24
26
  export type ConfigCandidate = {
25
27
  blockIsObject?: unknown;
26
28
  enabled?: unknown;
29
+ firstEventMs?: unknown;
27
30
  warningMs?: unknown;
28
31
  recoveryMs?: unknown;
29
32
  maxStallRetries?: unknown;
@@ -37,11 +40,12 @@ const DEFAULT_CANDIDATE: ConfigCandidate = { blockIsObject: true, ...DEFAULT_CON
37
40
 
38
41
  export function coerce(raw: unknown): ConfigCandidate | undefined {
39
42
  if (raw === undefined) return undefined;
43
+ if (typeof raw === "boolean") return { blockIsObject: true, enabled: raw };
40
44
  if (raw === null || typeof raw !== "object" || Array.isArray(raw)) return { blockIsObject: false };
41
45
 
42
46
  const source = raw as Record<string, unknown>;
43
47
  const candidate: ConfigCandidate = { blockIsObject: true };
44
- for (const key of ["enabled", "warningMs", "recoveryMs", "maxStallRetries"] as const) {
48
+ for (const key of ["enabled", "firstEventMs", "warningMs", "recoveryMs", "maxStallRetries"] as const) {
45
49
  if (Object.hasOwn(source, key)) candidate[key] = source[key];
46
50
  }
47
51
  return candidate;
@@ -50,14 +54,16 @@ export function coerce(raw: unknown): ConfigCandidate | undefined {
50
54
  export function validateConfig(candidate: ConfigCandidate): ConfigValidation {
51
55
  if (candidate.blockIsObject !== true) return { ok: false, error: "providerStallWatchdog must be an object" };
52
56
  if (typeof candidate.enabled !== "boolean") return { ok: false, error: "enabled must be a boolean" };
57
+ if (!isTimerDelay(candidate.firstEventMs)) return { ok: false, error: "firstEventMs must be a positive timer delay" };
53
58
  if (!isTimerDelay(candidate.warningMs)) return { ok: false, error: "warningMs must be a positive timer delay" };
54
59
  if (!isTimerDelay(candidate.recoveryMs)) return { ok: false, error: "recoveryMs must be a positive timer delay" };
55
60
  if (candidate.warningMs >= candidate.recoveryMs) return { ok: false, error: "warningMs must be less than recoveryMs" };
56
- if (!isPositiveInteger(candidate.maxStallRetries)) return { ok: false, error: "maxStallRetries must be a positive integer" };
61
+ if (!isNonNegativeInteger(candidate.maxStallRetries)) return { ok: false, error: "maxStallRetries must be a non-negative integer" };
57
62
  return {
58
63
  ok: true,
59
64
  config: {
60
65
  enabled: candidate.enabled,
66
+ firstEventMs: candidate.firstEventMs,
61
67
  warningMs: candidate.warningMs,
62
68
  recoveryMs: candidate.recoveryMs,
63
69
  maxStallRetries: candidate.maxStallRetries,
@@ -69,8 +75,8 @@ function isTimerDelay(value: unknown): value is number {
69
75
  return typeof value === "number" && Number.isFinite(value) && Number.isInteger(value) && value > 0 && value <= MAX_TIMER_MS;
70
76
  }
71
77
 
72
- function isPositiveInteger(value: unknown): value is number {
73
- return typeof value === "number" && Number.isInteger(value) && value > 0;
78
+ function isNonNegativeInteger(value: unknown): value is number {
79
+ return typeof value === "number" && Number.isInteger(value) && value >= 0;
74
80
  }
75
81
 
76
82
  /** Pi's own resolution is `settings.retry?.maxRetries ?? 3` over the same layered settings.json files. */
@@ -80,7 +86,7 @@ export function resolveRetryMaxRetries(cwd: string): number {
80
86
  const retry = readSettings(path)?.retry;
81
87
  if (retry === null || typeof retry !== "object" || Array.isArray(retry)) continue;
82
88
  const value = (retry as Record<string, unknown>).maxRetries;
83
- if (isPositiveInteger(value)) maxRetries = value;
89
+ if (isNonNegativeInteger(value)) maxRetries = value;
84
90
  }
85
91
  return maxRetries;
86
92
  }
@@ -100,8 +106,11 @@ const defaultRuntime: WatchdogRuntime = {
100
106
  };
101
107
 
102
108
  const DEGRADATION_NOTICE = "The stalled request was stopped, but Pi did not start an automatic retry. Retry may be disabled, exhausted, or incompatible; submit the message again to retry manually.";
109
+ // Reduces, but cannot eliminate, the hang when an aborted provider operation never terminates;
110
+ // undici's headersTimeout/bodyTimeout stay the backstop past this point.
111
+ const ABORT_GRACE_MS = 10_000;
103
112
 
104
- type Timer = { warning?: unknown; recovery?: unknown };
113
+ type Timer = { firstEvent?: unknown; warning?: unknown; recovery?: unknown; abortGuard?: unknown };
105
114
 
106
115
  function formatElapsed(ms: number): string {
107
116
  if (ms % 60_000 === 0) return `${ms / 60_000}m`;
@@ -109,6 +118,8 @@ function formatElapsed(ms: number): string {
109
118
  return `${ms}ms`;
110
119
  }
111
120
 
121
+ const ABORT_STUCK_NOTICE = `The stalled request did not stop within ${formatElapsed(ABORT_GRACE_MS)} of being aborted; the provider connection is unresponsive. No automatic retry will run - the turn will not end until the HTTP idle timeout expires.`;
122
+
112
123
  function warningNotice(config: WatchdogConfig): string {
113
124
  return `No model progress for ${formatElapsed(config.warningMs)}; aborting and asking Pi to retry in ${formatElapsed(config.recoveryMs - config.warningMs)} (Esc aborts now)`;
114
125
  }
@@ -117,9 +128,20 @@ function exhaustedNotice(config: WatchdogConfig): string {
117
128
  return `Stall retry budget (${config.maxStallRetries}) exhausted; aborting without another automatic retry. Submit the message again manually.`;
118
129
  }
119
130
 
131
+ function firstEventRetryNotice(config: WatchdogConfig): string {
132
+ return `Provider sent no response for ${formatElapsed(config.firstEventMs)}; stopping and retrying the request.`;
133
+ }
134
+
135
+ function firstEventExhaustedNotice(config: WatchdogConfig): string {
136
+ return `Provider sent no response for ${formatElapsed(config.firstEventMs)} and the stall-retry budget is spent; the request was stopped.`;
137
+ }
138
+
120
139
  export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRuntime): (pi: ExtensionAPI) => void {
121
140
  return (pi) => {
122
- let pendingInteractive = false;
141
+ let midStreamEnabled = false;
142
+ let firstEventSeen = false;
143
+ let hasUI = false;
144
+ let pendingTimeoutReason: string | undefined;
123
145
  let activeRun = false;
124
146
  let disabled = false;
125
147
  let config: WatchdogConfig | undefined;
@@ -127,22 +149,25 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
127
149
  let activeGeneration: number | undefined;
128
150
  let lastSemanticAt = 0;
129
151
  let warned = false;
130
- let epoch = 0;
131
152
  let deadlineEpoch = 0;
132
153
  let timers: Timer = {};
133
154
  let removeSignalListener: (() => void) | undefined;
134
155
  let ui: { notify(text: string, type?: string): void } | undefined;
135
156
  let watchdogAbortedGeneration: number | undefined;
136
- let timeoutConversionPending = false;
137
157
  let stallRetriesUsed = 0;
138
158
  let continuationStarted = false;
139
159
  let convertedTimeout = false;
140
160
 
141
161
  const clearTimers = () => {
142
- if (timers.warning !== undefined) runtime.clearTimeout(timers.warning);
143
- if (timers.recovery !== undefined) runtime.clearTimeout(timers.recovery);
162
+ for (const key of ["firstEvent", "warning", "recovery", "abortGuard"] as const) {
163
+ if (timers[key] !== undefined) runtime.clearTimeout(timers[key]);
164
+ }
144
165
  timers = {};
145
166
  };
167
+ const announce = (text: string, type?: string) => {
168
+ ui?.notify(text, type);
169
+ if (!hasUI) console.warn(text);
170
+ };
146
171
  const clear = () => {
147
172
  clearTimers();
148
173
  removeSignalListener?.();
@@ -153,20 +178,81 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
153
178
  const resetRunState = () => {
154
179
  disarm();
155
180
  activeRun = false;
156
- pendingInteractive = false;
157
181
  stallRetriesUsed = 0;
158
182
  continuationStarted = false;
159
183
  convertedTimeout = false;
160
184
  watchdogAbortedGeneration = undefined;
161
- timeoutConversionPending = false;
185
+ pendingTimeoutReason = undefined;
186
+ firstEventSeen = false;
187
+ };
188
+ const armAbortGuard = (capturedGeneration: number) => {
189
+ if (timers.abortGuard !== undefined) runtime.clearTimeout(timers.abortGuard);
190
+ timers.abortGuard = runtime.setTimeout(() => {
191
+ if (capturedGeneration !== activeGeneration || !activeRun) return;
192
+ announce(ABORT_STUCK_NOTICE, "error");
193
+ // Only the timers go: the generation stays armed so a message_end arriving after the grace
194
+ // period still converts the abort the stall retry already paid for.
195
+ clearTimers();
196
+ }, ABORT_GRACE_MS);
162
197
  };
163
- const schedule = (ctx: { ui: typeof ui; abort(): void }) => {
198
+ // A stream event on a generation the watchdog already aborted: the stall cycle is over for this
199
+ // request, so nothing re-enters it. The bytes only prove the connection was alive at this instant,
200
+ // so the guard is re-armed rather than cleared - a wedge right after a straggler still escalates.
201
+ const postAbortStreamEvent = () => {
202
+ if (activeGeneration === undefined || watchdogAbortedGeneration !== activeGeneration) return false;
203
+ armAbortGuard(activeGeneration);
204
+ return true;
205
+ };
206
+ const abortStall = (
207
+ ctx: { abort(): void },
208
+ capturedGeneration: number,
209
+ notices: { retry: () => string; exhausted: () => string },
210
+ reason: string,
211
+ ) => {
212
+ if (!config) return;
213
+ clearTimers();
214
+ warned = false;
215
+ watchdogAbortedGeneration = capturedGeneration;
216
+ // Armed before ctx.abort() so a synchronous teardown inside it cannot orphan the timer;
217
+ // the generation check makes the callback a no-op if that teardown disarmed the watchdog.
218
+ armAbortGuard(capturedGeneration);
219
+ if (stallRetriesUsed >= config.maxStallRetries) {
220
+ announce(notices.exhausted());
221
+ ctx.abort();
222
+ return;
223
+ }
224
+ pendingTimeoutReason = reason;
225
+ stallRetriesUsed += 1;
226
+ announce(notices.retry());
227
+ ctx.abort();
228
+ };
229
+ const armFirstEvent = (ctx: { abort(): void }) => {
164
230
  if (activeGeneration === undefined || !config) return;
231
+ const cfg = config;
232
+ const capturedGeneration = activeGeneration;
233
+ const capturedDeadlineEpoch = ++deadlineEpoch;
234
+ const threshold = cfg.firstEventMs;
235
+ const run = () => {
236
+ if (capturedGeneration !== activeGeneration || capturedDeadlineEpoch !== deadlineEpoch || !activeRun || firstEventSeen) return;
237
+ const elapsed = runtime.now() - lastSemanticAt;
238
+ if (elapsed < threshold) {
239
+ timers.firstEvent = runtime.setTimeout(run, threshold - elapsed);
240
+ return;
241
+ }
242
+ abortStall(ctx, capturedGeneration, {
243
+ retry: () => firstEventRetryNotice(cfg),
244
+ exhausted: () => firstEventExhaustedNotice(cfg),
245
+ }, `Provider first-event timeout after ${cfg.firstEventMs} ms without a stream event`);
246
+ };
247
+ timers.firstEvent = runtime.setTimeout(run, threshold);
248
+ };
249
+ const schedule = (ctx: { abort(): void }) => {
250
+ if (activeGeneration === undefined || !config) return;
251
+ const cfg = config;
165
252
  const capturedGeneration = activeGeneration;
166
- const capturedEpoch = epoch;
167
253
  const capturedDeadlineEpoch = ++deadlineEpoch;
168
254
  const run = (kind: "warning" | "recovery", threshold: number) => () => {
169
- if (capturedEpoch !== epoch || capturedGeneration !== activeGeneration || capturedDeadlineEpoch !== deadlineEpoch || !activeRun || !config) return;
255
+ if (capturedGeneration !== activeGeneration || capturedDeadlineEpoch !== deadlineEpoch || !activeRun) return;
170
256
  const elapsed = runtime.now() - lastSemanticAt;
171
257
  if (elapsed < threshold) {
172
258
  timers[kind] = runtime.setTimeout(run(kind, threshold), threshold - elapsed);
@@ -174,51 +260,38 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
174
260
  }
175
261
  if (kind === "warning" && !warned) {
176
262
  warned = true;
177
- ctx.ui?.notify(warningNotice(config), "warning");
263
+ announce(warningNotice(cfg), "warning");
178
264
  }
179
265
  if (kind === "recovery") {
180
- clearTimers();
181
- warned = false;
182
- if (stallRetriesUsed >= config.maxStallRetries) {
183
- ui?.notify(exhaustedNotice(config));
184
- ctx.abort();
185
- return;
186
- }
187
- watchdogAbortedGeneration = capturedGeneration;
188
- timeoutConversionPending = true;
189
- stallRetriesUsed += 1;
190
- ui?.notify(`No model progress for ${formatElapsed(elapsed)}; aborting now. Pi will retry (${stallRetriesUsed}/${config.maxStallRetries}) if retry is enabled and capacity remains. Pending follow-ups are returned to the editor.`);
191
- ctx.abort();
266
+ abortStall(ctx, capturedGeneration, {
267
+ retry: () => `No model progress for ${formatElapsed(elapsed)}; aborting now. Pi will retry (${stallRetriesUsed}/${cfg.maxStallRetries}) if retry is enabled and capacity remains. Pending follow-ups are returned to the editor.`,
268
+ exhausted: () => exhaustedNotice(cfg),
269
+ }, `Provider semantic timeout after ${cfg.recoveryMs} ms without progress`);
192
270
  }
193
271
  };
194
- timers.warning = runtime.setTimeout(run("warning", config.warningMs), config.warningMs);
195
- timers.recovery = runtime.setTimeout(run("recovery", config.recoveryMs), config.recoveryMs);
272
+ timers.warning = runtime.setTimeout(run("warning", cfg.warningMs), cfg.warningMs);
273
+ timers.recovery = runtime.setTimeout(run("recovery", cfg.recoveryMs), cfg.recoveryMs);
196
274
  };
197
275
 
198
- pi.on("input", (event) => {
199
- if (!activeRun) pendingInteractive = event.source === "interactive";
200
- });
201
- pi.on("before_agent_start", (_event, ctx) => {
202
- if (!pendingInteractive) return;
203
- pendingInteractive = false;
204
- if (ctx.mode !== "tui" || disabled) return;
205
- const resolved = resolveWatchdogConfig(ctx.cwd);
206
- if (!resolved.ok) {
207
- disabled = true;
208
- console.warn(`providerStallWatchdog disabled: ${resolved.error}`);
209
- ctx.ui.notify(`providerStallWatchdog disabled: ${resolved.error}`, "warning");
210
- return;
276
+ pi.on("before_provider_request", (_event, ctx) => {
277
+ if (disabled) return;
278
+ ui = ctx.ui;
279
+ hasUI = ctx.hasUI;
280
+ if (!config) {
281
+ const resolved = resolveWatchdogConfig(ctx.cwd);
282
+ if (!resolved.ok) {
283
+ disabled = true;
284
+ announce(`providerStallWatchdog disabled: ${resolved.error}`, "warning");
285
+ return;
286
+ }
287
+ config = resolved.config;
211
288
  }
212
- config = resolved.config;
213
289
  activeRun = config.enabled;
214
- });
215
- pi.on("before_provider_request", (_event, ctx) => {
216
- if (!activeRun || ctx.mode !== "tui" || !config) return;
290
+ if (!activeRun) return;
217
291
  disarm();
218
292
  if (convertedTimeout) continuationStarted = true;
219
293
  activeGeneration = ++generation;
220
294
  lastSemanticAt = runtime.now();
221
- ui = ctx.ui;
222
295
  const target = ctx.signal;
223
296
  if (target) {
224
297
  const listener = () => {
@@ -227,11 +300,30 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
227
300
  target.addEventListener("abort", listener, { once: true });
228
301
  removeSignalListener = () => target.removeEventListener("abort", listener);
229
302
  }
303
+ midStreamEnabled = ctx.mode === "tui";
304
+ firstEventSeen = false;
305
+ armFirstEvent(ctx);
306
+ });
307
+ pi.on("message_start", (event, ctx) => {
308
+ // Pi fires message_start for user and toolResult messages too; only an assistant one is
309
+ // provider traffic, so the role check must stay ahead of the liveness re-arm below.
310
+ if (event.message.role !== "assistant") return;
311
+ if (postAbortStreamEvent()) return;
312
+ if (!activeRun || activeGeneration === undefined || firstEventSeen) return;
313
+ firstEventSeen = true;
314
+ if (timers.firstEvent !== undefined) {
315
+ runtime.clearTimeout(timers.firstEvent);
316
+ timers.firstEvent = undefined;
317
+ }
318
+ if (!midStreamEnabled || !config) return;
319
+ lastSemanticAt = runtime.now();
320
+ warned = false;
230
321
  schedule(ctx);
231
322
  });
232
323
  pi.on("message_update", (event, ctx) => {
324
+ if (postAbortStreamEvent()) return;
233
325
  const update = event.assistantMessageEvent;
234
- if (activeGeneration === undefined || !(update.type === "text_delta" || update.type === "thinking_delta" || update.type === "toolcall_delta") || update.delta.length === 0) return;
326
+ if (!midStreamEnabled || !firstEventSeen || activeGeneration === undefined || !(update.type === "text_delta" || update.type === "thinking_delta" || update.type === "toolcall_delta") || update.delta.length === 0) return;
235
327
  lastSemanticAt = runtime.now();
236
328
  warned = false;
237
329
  clearTimers();
@@ -239,25 +331,25 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
239
331
  });
240
332
  pi.on("message_end", (event) => {
241
333
  if (event.message.role !== "assistant") return;
242
- const matchesWatchdogAbort = event.message.stopReason === "aborted"
243
- && activeGeneration === watchdogAbortedGeneration
244
- && timeoutConversionPending;
334
+ const errorMessage = pendingTimeoutReason;
335
+ pendingTimeoutReason = undefined;
336
+ const matchesWatchdogAbort = errorMessage !== undefined
337
+ && event.message.stopReason === "aborted"
338
+ && activeGeneration === watchdogAbortedGeneration;
245
339
  disarm();
246
340
  // Mirror Pi's retry loop, which resets its attempt counter on any successful assistant turn.
247
341
  if (event.message.stopReason !== "aborted" && event.message.stopReason !== "error") stallRetriesUsed = 0;
248
- if (!matchesWatchdogAbort || !config) return;
249
- timeoutConversionPending = false;
342
+ if (!matchesWatchdogAbort) return;
250
343
  convertedTimeout = true;
251
344
  continuationStarted = false;
252
- return { message: { ...event.message, stopReason: "error", errorMessage: `Provider semantic timeout after ${config.recoveryMs} ms without progress` } };
345
+ return { message: { ...event.message, stopReason: "error", errorMessage } };
253
346
  });
254
347
  pi.on("agent_end", () => disarm());
255
348
  pi.on("agent_settled", () => {
256
- if (convertedTimeout && !continuationStarted) ui?.notify(DEGRADATION_NOTICE);
349
+ if (convertedTimeout && !continuationStarted) announce(DEGRADATION_NOTICE);
257
350
  resetRunState();
258
351
  });
259
352
  pi.on("session_shutdown", () => {
260
- epoch += 1;
261
353
  resetRunState();
262
354
  config = undefined;
263
355
  disabled = false;