pi-quiver 3.3.2 → 3.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +14 -0
- package/README.md +31 -5
- package/fast-mode.ts +52 -5
- package/package.json +1 -1
- package/provider-stall-watchdog.ts +150 -58
package/CHANGELOG.md
CHANGED
|
@@ -8,6 +8,20 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
|
|
|
8
8
|
via OIDC trusted publishing. The release helper at
|
|
9
9
|
`.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
|
|
10
10
|
|
|
11
|
+
## v3.4.1 - 2026-08-04
|
|
12
|
+
|
|
13
|
+
- **`fast-mode` now corrects reported `usage.cost` to true 2x fast pricing for Opus 4.8/5** (statusline, persisted cost, and pi-cohort `Σ$` per load order) (#6).
|
|
14
|
+
- **Documented Anthropic stall-retry activation.** A session trace confirmed Pi's existing 2/4/8-second exponential backoff was active and that the watchdog retained the retry budget resolved when its extension instance started. Local `retry.maxRetries: 5` now takes effect after a fresh session or explicit reload, with no watchdog cooldown, resubmission path, or runtime code change.
|
|
15
|
+
|
|
16
|
+
## v3.4.0 - 2026-07-30
|
|
17
|
+
|
|
18
|
+
- **`provider-stall-watchdog`: new `firstEventMs` tier (default 20s) fast-fails a request that never streams.** A pre-first-event deadline is armed at every provider request and cleared by the first assistant `message_start`; on expiry the request is aborted and converted to a retryable error, cutting recovery for an unresponsive request from ~240s to ~22s (detection + Pi's 2s backoff). Previously only the 240s mid-stream `recoveryMs` timer caught this case.
|
|
19
|
+
- **Widened activation.** The `firstEventMs` tier arms in every mode (`tui`, `print`, `json`, `rpc`) and from every origin, including extension-triggered turns that never emit `before_agent_start`; activation is now lazy inside `before_provider_request`, so the `input` and `before_agent_start` handlers are gone. The mid-stream `warningMs`/`recoveryMs` tier stays TUI-only - an unattended abort discards billed output tokens with nobody reading the warning. Headless runs report on stderr via `console.warn` (never stdout, which `json` mode uses for its protocol).
|
|
20
|
+
- **Behaviour change: `maxStallRetries: 0` is now valid** and means "detect and stop, never auto-retry" - it previously failed validation and disabled the extension. When unset, `maxStallRetries` inherits the layered `retry.maxRetries` including an explicit `0`, matching Pi's own `?? 3` instead of silently resolving 0 to 3. Both tiers share the one budget.
|
|
21
|
+
- **Abort-hang escalation.** After any watchdog abort a fixed 10s deadline announces that the provider connection is unresponsive if the request never terminates; any post-abort stream event re-arms it, so a straggler followed by a wedge still escalates. It reduces the hang, it cannot force the provider to stop - undici's `headersTimeout`/`bodyTimeout` (pi's `httpIdleTimeoutMs`, default 300s) stay the backstop.
|
|
22
|
+
- **Config is read once per session,** on the first provider request, instead of per prompt. Editing `settings.json` mid-session - including repairing an invalid block that disabled the extension - now requires a session restart.
|
|
23
|
+
- **`"providerStallWatchdog": false` now works.** The boolean shorthand the README already documented was rejected as "must be an object"; it is now accepted as `{ "enabled": <bool> }`, matching `swordHeader`, `fastMode`, and `sessionAutoName`. Previously harmless (the warning was TUI-only), the widened activation above would otherwise have printed a config error to stderr on every headless run.
|
|
24
|
+
|
|
11
25
|
## v3.3.2 - 2026-07-30
|
|
12
26
|
|
|
13
27
|
- **`fast-mode` supports Claude Opus 5.** Added `claude-opus-5` to the model prefix allowlist (fast mode API mechanics are identical to Opus 4.8: `speed: "fast"` + `fast-mode-2026-02-01` beta; compat verified identical in pi-ai 0.82.1, so `buildBetaHeader`'s OAuth-only preservation still holds).
|
package/README.md
CHANGED
|
@@ -67,7 +67,7 @@ A 300 KB changelog page never touches your context window - you get a preview an
|
|
|
67
67
|
| `session-name.ts` | `/session-name` | Manual + opt-in automatic session naming, with Ghostty tab rename. OFF by default. |
|
|
68
68
|
| `sword-header.ts` | `/builtin-header` | Themed ASCII startup header replacing pi's default logo. OFF by default. |
|
|
69
69
|
| `fast-mode.ts` | `/fast` | Inject Anthropic fast-mode (`speed: "fast"` + `anthropic-beta: fast-mode-2026-02-01`) into every Claude Opus 4.8 / Opus 5 request, any thinking level. `--fast` flag + `/fast [on\|off\|status]`. OFF by default. |
|
|
70
|
-
| `provider-stall-watchdog.ts` | - | Opt-in
|
|
70
|
+
| `provider-stall-watchdog.ts` | - | Opt-in provider-stall recovery, in two tiers: a pre-first-event deadline (`firstEventMs`, 20s) on every provider request in every mode, and the mid-stream pair (warn at 2 min, recover at 4 min) in TUI runs only. Policy D offers each stall to Pi's retry loop until the stall retry budget (`maxStallRetries`, default = `retry.maxRetries`) is exhausted. OFF by default. |
|
|
71
71
|
|
|
72
72
|
Full routing rules, size-gate mechanics, and config: [doc/fetch.md](doc/fetch.md), [doc/doc-to-md.md](doc/doc-to-md.md).
|
|
73
73
|
|
|
@@ -79,20 +79,20 @@ Full routing rules, size-gate mechanics, and config: [doc/fetch.md](doc/fetch.md
|
|
|
79
79
|
| Content routing | HTML -> Markdown, binary -> untouched file, GitHub URLs -> `gh` CLI, everything else -> the size gate. |
|
|
80
80
|
| Graceful degradation | Optional binaries (`gh`, `uv`, LibreOffice) are never hard install-time deps; each has a defined, documented fallback or failure mode. |
|
|
81
81
|
| Opt-in extensions | `session-name`, `sword-header`, `fast-mode`, and `provider-stall-watchdog` do nothing until explicitly enabled in `settings.json`. |
|
|
82
|
-
| Provider stall recovery | The watchdog detects missing parsed semantic progress, not network liveness.
|
|
82
|
+
| Provider stall recovery | The watchdog detects a missing first stream event and missing parsed semantic progress, not network liveness. The pre-first-event tier covers every mode and origin; the mid-stream tier is TUI-only. |
|
|
83
83
|
|
|
84
84
|
## When to use
|
|
85
85
|
|
|
86
86
|
- An agent needs to reason from a real web page, GitHub issue/PR, or local PDF/DOCX/PPTX instead of memory.
|
|
87
87
|
- You want that ingestion to be safe by default, with no risk of a single call blowing the context budget.
|
|
88
|
-
- A
|
|
88
|
+
- A Pi run needs an opt-in guard against provider requests that never produce a first stream event, plus mid-stream silence recovery in interactive TUI sessions.
|
|
89
89
|
|
|
90
90
|
## When NOT to use
|
|
91
91
|
|
|
92
92
|
- You need a general-purpose web scraper (JS-rendered pages, pagination, auth flows) - `fetch` does plain HTTP + Readability extraction, nothing more.
|
|
93
93
|
- You need spreadsheet conversion - `doc_to_md` explicitly excludes spreadsheets (they paginate badly via PDF).
|
|
94
94
|
- You want automatic session naming, a custom header, fast mode, or stall recovery without opting in - all stay off until you flip the config.
|
|
95
|
-
- You need
|
|
95
|
+
- You need *mid-stream* stall recovery in JSON, RPC, or print runs - only the pre-first-event tier arms there; mid-stream silence falls through to pi's transport timeout.
|
|
96
96
|
|
|
97
97
|
## Install
|
|
98
98
|
|
|
@@ -150,6 +150,8 @@ These extensions are opt-in via `settings.json` (project `.pi/settings.json` ove
|
|
|
150
150
|
|
|
151
151
|
`sessionAutoName.enabled` makes one extra short LLM call per session (once, after the first turn) to title it; `false` (default) makes no model calls. `fastMode` only affects `claude-opus-4-8` and `claude-opus-5` requests on Anthropic's `anthropic-messages` API; enabling it opts into premium fast-mode pricing. `--fast` forces it on for one launch; `/fast on|off` toggles live. Proxy providers (opencode, cloudflare-ai-gateway) are excluded. `fastMode`'s header injection needs the `before_provider_headers` hook (pi bundling `@earendil-works/pi-coding-agent` >= 0.80.5); on older pi the beta header is silently not sent. See [doc/fetch.md](doc/fetch.md) and [doc/doc-to-md.md](doc/doc-to-md.md) for the ingestion tools' full reference; session-name/sword-header behavior above is complete.
|
|
152
152
|
|
|
153
|
+
`pi-ai` prices every fast request at standard rates - it has no `usage.speed` support and no request-level pricing modifier - so `fastMode` corrects the reported cost itself: a `message_end` handler scales all four `usage.cost` components by `FAST_MODE_COST_MULTIPLIER` (2x) and returns the corrected message. Persisted session JSONL and pi's own native cost display are always exact, since they're written from this corrected message. pi-cohort's live `Σ$` reflects the correction only when pi-quiver's `message_end` handler runs before pi-cohort's - best-effort, depending on extension load order - and is reconciled on pi-cohort's next `session_start` regardless. The upstream fix (teaching `pi-ai`'s `Usage`/`calculateCost` about `usage.speed`) is the better long-term path and is tracked separately.
|
|
154
|
+
|
|
153
155
|
Recommended explicit retry and watchdog settings:
|
|
154
156
|
|
|
155
157
|
```json
|
|
@@ -161,6 +163,7 @@ Recommended explicit retry and watchdog settings:
|
|
|
161
163
|
},
|
|
162
164
|
"providerStallWatchdog": {
|
|
163
165
|
"enabled": true,
|
|
166
|
+
"firstEventMs": 20000,
|
|
164
167
|
"warningMs": 120000,
|
|
165
168
|
"recoveryMs": 240000,
|
|
166
169
|
"maxStallRetries": 3
|
|
@@ -168,7 +171,30 @@ Recommended explicit retry and watchdog settings:
|
|
|
168
171
|
}
|
|
169
172
|
```
|
|
170
173
|
|
|
171
|
-
|
|
174
|
+
| Key | Default | Where it arms | What it measures |
|
|
175
|
+
| --- | --- | --- | --- |
|
|
176
|
+
| `enabled` | `false` | - | Master switch. OFF means the extension does nothing. |
|
|
177
|
+
| `firstEventMs` | `20000` | every provider request, every mode (`tui`/`print`/`json`/`rpc`), every origin | Silence between the request and the first assistant `message_start`. |
|
|
178
|
+
| `warningMs` | `120000` | mid-stream, `ctx.mode === "tui"` only | Silence since the last non-empty text/thinking/toolcall delta; notifies. |
|
|
179
|
+
| `recoveryMs` | `240000` | mid-stream, `ctx.mode === "tui"` only | Same clock; aborts and converts. Must be `> warningMs`. |
|
|
180
|
+
| `maxStallRetries` | layered `retry.maxRetries`, else `3` | shared by both tiers | Watchdog aborts that may convert to a retryable error before stopping. |
|
|
181
|
+
|
|
182
|
+
`providerStallWatchdog` is OFF by default. Once enabled it arms in two tiers per provider request:
|
|
183
|
+
|
|
184
|
+
- **Pre-first-event (`firstEventMs`).** Armed at every provider request, in every mode and from every origin - including extension-triggered turns that never emit `before_agent_start` - and cleared by the first assistant `message_start`. On expiry the request is aborted and, budget permitting, converted to a retryable error, so an unresponsive request recovers in ~22s (20s detection + Pi's 2s backoff) instead of the ~240s it took when only the mid-stream tier existed.
|
|
185
|
+
- **Mid-stream (`warningMs` / `recoveryMs`).** Armed from the first assistant `message_start` onward, and only when `ctx.mode === "tui"`. Aborting mid-generation discards billed output tokens and an unattended run has nobody to read the warning, so headless mid-stream silence deliberately falls through to the transport timeout instead.
|
|
186
|
+
|
|
187
|
+
**Raise `firstEventMs` if your provider is legitimately slow to first event.** Queueing gateways, throttled endpoints, and busy single-slot local model servers can hold the connection for well over 20s before their first stream event; every false abort re-uploads the whole context and spends one stall retry.
|
|
188
|
+
|
|
189
|
+
**Leave pi's own `httpIdleTimeoutMs` (default `300000`) at its default.** It is the transport backstop, and a single value drives undici's `headersTimeout` *and* `bodyTimeout` - lowering it to get fast pre-stream failure also truncates legitimate mid-stream gaps. `firstEventMs` is the knob for pre-stream silence.
|
|
190
|
+
|
|
191
|
+
Verified with Pi 0.80.10: each stall is aborted and offered to Pi retry until `maxStallRetries` conversions are used; further stalls stop for manual resubmission. Both tiers draw on that one budget. `maxStallRetries` defaults to the layered `retry.maxRetries` (Pi default 3, an explicit `0` honoured); `0` is valid and means "detect and stop, never auto-retry". Consecutive stall conversions consume Pi retry attempts without a success reset in between, so keep `maxStallRetries <= retry.maxRetries`. A successful assistant turn resets the stall counter (mirroring Pi's own retry counter). Automatic continuation needs enabled Pi retry with remaining capacity. Disabled, exhausted, or incompatible retry degrades to manual resubmission. Pending steering or follow-ups return to the editor and are excluded from automatic continuation. Invalid merged watchdog config fails closed.
|
|
192
|
+
|
|
193
|
+
Operational notes:
|
|
194
|
+
|
|
195
|
+
- **Settings are read once per session,** on the first provider request. Editing `settings.json` mid-session changes nothing until you restart the session - that includes repairing an invalid block that already disabled the extension.
|
|
196
|
+
- **A watchdog abort that the provider ignores escalates after a fixed 10s.** Any post-abort stream event re-arms that deadline (bytes prove only that the connection was alive at that instant), so a stream that emits a straggler and then wedges still escalates 10s after its last event. This reduces the hang; it cannot force the provider to stop, and undici's timeouts remain the final backstop.
|
|
197
|
+
- **Headless runs report on stderr.** In `print`/`json` mode pi binds a no-op UI, so watchdog notices go out via `console.warn`. Nothing is ever written to stdout, which `json` mode uses for its protocol. In TUI and RPC the notices render as main-window notifications, not the bottom status line.
|
|
172
198
|
|
|
173
199
|
## Development
|
|
174
200
|
|
package/fast-mode.ts
CHANGED
|
@@ -22,6 +22,7 @@
|
|
|
22
22
|
*/
|
|
23
23
|
|
|
24
24
|
import type { ExtensionAPI, ExtensionContext } from "@earendil-works/pi-coding-agent";
|
|
25
|
+
import type { Usage } from "@earendil-works/pi-ai";
|
|
25
26
|
import { resolveConfig } from "./extension-config.ts";
|
|
26
27
|
|
|
27
28
|
export const FAST_MODE_BETA = "fast-mode-2026-02-01";
|
|
@@ -29,6 +30,21 @@ export const FAST_SPEED = "fast";
|
|
|
29
30
|
// Loose prefixes: match dated snapshots (claude-opus-4-8-*, claude-opus-5-*).
|
|
30
31
|
// Opus 4.7 is out of scope (D1); a future model needs a one-line addition here.
|
|
31
32
|
export const FAST_MODE_MODEL_PREFIXES = ["claude-opus-4-8", "claude-opus-5"];
|
|
33
|
+
|
|
34
|
+
// Anthropic bills Opus 4.8/5 fast mode at exactly 2x standard rates; input,
|
|
35
|
+
// output, and cache read/write all scale by the same factor because caching
|
|
36
|
+
// multipliers "apply on top of fast mode pricing". Rate card:
|
|
37
|
+
// https://docs.claude.com/en/docs/build-with-claude/fast-mode (retrieved 2026-08-04).
|
|
38
|
+
export const FAST_MODE_COST_MULTIPLIER = 2;
|
|
39
|
+
|
|
40
|
+
export function scaleCost(cost: Usage["cost"], multiplier: number): Usage["cost"] {
|
|
41
|
+
const input = cost.input * multiplier;
|
|
42
|
+
const output = cost.output * multiplier;
|
|
43
|
+
const cacheRead = cost.cacheRead * multiplier;
|
|
44
|
+
const cacheWrite = cost.cacheWrite * multiplier;
|
|
45
|
+
return { input, output, cacheRead, cacheWrite, total: input + output + cacheRead + cacheWrite };
|
|
46
|
+
}
|
|
47
|
+
|
|
32
48
|
export const OAUTH_IDENTITY_BETAS = ["claude-code-20250219", "oauth-2025-04-20"];
|
|
33
49
|
const STATUS_KEY = "fast-mode";
|
|
34
50
|
const BETA_HEADER = "anthropic-beta";
|
|
@@ -92,6 +108,10 @@ export function resolveEnabled(s: State): boolean {
|
|
|
92
108
|
export default function (pi: ExtensionAPI) {
|
|
93
109
|
let liveOverride: boolean | null = null;
|
|
94
110
|
let enabled = false;
|
|
111
|
+
// Per-request snapshot. Safe as plain booleans because pi serializes provider
|
|
112
|
+
// requests per session (headers -> request -> message_end, awaited in order).
|
|
113
|
+
let pendingFastSpeed = false;
|
|
114
|
+
let pendingFastHeader = false;
|
|
95
115
|
|
|
96
116
|
const readFlag = (): boolean => pi.getFlag("fast") === true;
|
|
97
117
|
|
|
@@ -135,16 +155,43 @@ export default function (pi: ExtensionAPI) {
|
|
|
135
155
|
});
|
|
136
156
|
|
|
137
157
|
pi.on("before_provider_request", (event, ctx) => {
|
|
138
|
-
if (!shouldInject(enabled, ctx.model))
|
|
139
|
-
|
|
158
|
+
if (!shouldInject(enabled, ctx.model)) {
|
|
159
|
+
pendingFastSpeed = false;
|
|
160
|
+
return;
|
|
161
|
+
}
|
|
162
|
+
const next = injectSpeed(event.payload);
|
|
163
|
+
pendingFastSpeed =
|
|
164
|
+
typeof next === "object" && next !== null && (next as Record<string, unknown>).speed === FAST_SPEED;
|
|
165
|
+
return next;
|
|
140
166
|
});
|
|
141
167
|
|
|
142
168
|
pi.on("before_provider_headers", async (event, ctx) => {
|
|
143
|
-
if (!shouldInject(enabled, ctx.model))
|
|
144
|
-
|
|
169
|
+
if (!shouldInject(enabled, ctx.model) || !event.headers) {
|
|
170
|
+
pendingFastHeader = false;
|
|
171
|
+
return;
|
|
172
|
+
}
|
|
145
173
|
const isOAuth = await detectOAuth(ctx);
|
|
146
|
-
if (isOAuth === null)
|
|
174
|
+
if (isOAuth === null) {
|
|
175
|
+
pendingFastHeader = false;
|
|
176
|
+
return;
|
|
177
|
+
}
|
|
147
178
|
event.headers[BETA_HEADER] = buildBetaHeader(event.headers[BETA_HEADER], isOAuth);
|
|
179
|
+
pendingFastHeader = true;
|
|
180
|
+
});
|
|
181
|
+
|
|
182
|
+
pi.on("message_end", (event) => {
|
|
183
|
+
const msg = event.message;
|
|
184
|
+
if (msg.role !== "assistant") return;
|
|
185
|
+
const corrected = pendingFastSpeed && pendingFastHeader && !!msg.usage?.cost;
|
|
186
|
+
pendingFastSpeed = false;
|
|
187
|
+
pendingFastHeader = false;
|
|
188
|
+
if (!corrected) return;
|
|
189
|
+
return {
|
|
190
|
+
message: {
|
|
191
|
+
...msg,
|
|
192
|
+
usage: { ...msg.usage, cost: scaleCost(msg.usage.cost, FAST_MODE_COST_MULTIPLIER) },
|
|
193
|
+
},
|
|
194
|
+
};
|
|
148
195
|
});
|
|
149
196
|
|
|
150
197
|
pi.registerCommand("fast", {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-quiver",
|
|
3
|
-
"version": "3.
|
|
3
|
+
"version": "3.4.1",
|
|
4
4
|
"description": "Personal pack of Pi coding-agent extensions: context-safe fetch, doc_to_md PDF/DOCX/PPTX-to-Markdown conversion, session naming, a themed ASCII startup header, Opus 4.8 fast mode, and a provider-stall watchdog.",
|
|
5
5
|
"author": "Jacek Juraszek",
|
|
6
6
|
"license": "MIT",
|
|
@@ -4,12 +4,14 @@ import { readSettings, resolveConfig, settingsPaths } from "./extension-config.t
|
|
|
4
4
|
export const MAX_TIMER_MS = 2_147_483_647;
|
|
5
5
|
export const DEFAULT_CONFIG = {
|
|
6
6
|
enabled: false,
|
|
7
|
+
firstEventMs: 20_000,
|
|
7
8
|
warningMs: 120_000,
|
|
8
9
|
recoveryMs: 240_000,
|
|
9
10
|
} as const;
|
|
10
11
|
|
|
11
12
|
export type WatchdogConfig = {
|
|
12
13
|
enabled: boolean;
|
|
14
|
+
firstEventMs: number;
|
|
13
15
|
warningMs: number;
|
|
14
16
|
recoveryMs: number;
|
|
15
17
|
maxStallRetries: number;
|
|
@@ -24,6 +26,7 @@ export type WatchdogRuntime = {
|
|
|
24
26
|
export type ConfigCandidate = {
|
|
25
27
|
blockIsObject?: unknown;
|
|
26
28
|
enabled?: unknown;
|
|
29
|
+
firstEventMs?: unknown;
|
|
27
30
|
warningMs?: unknown;
|
|
28
31
|
recoveryMs?: unknown;
|
|
29
32
|
maxStallRetries?: unknown;
|
|
@@ -37,11 +40,12 @@ const DEFAULT_CANDIDATE: ConfigCandidate = { blockIsObject: true, ...DEFAULT_CON
|
|
|
37
40
|
|
|
38
41
|
export function coerce(raw: unknown): ConfigCandidate | undefined {
|
|
39
42
|
if (raw === undefined) return undefined;
|
|
43
|
+
if (typeof raw === "boolean") return { blockIsObject: true, enabled: raw };
|
|
40
44
|
if (raw === null || typeof raw !== "object" || Array.isArray(raw)) return { blockIsObject: false };
|
|
41
45
|
|
|
42
46
|
const source = raw as Record<string, unknown>;
|
|
43
47
|
const candidate: ConfigCandidate = { blockIsObject: true };
|
|
44
|
-
for (const key of ["enabled", "warningMs", "recoveryMs", "maxStallRetries"] as const) {
|
|
48
|
+
for (const key of ["enabled", "firstEventMs", "warningMs", "recoveryMs", "maxStallRetries"] as const) {
|
|
45
49
|
if (Object.hasOwn(source, key)) candidate[key] = source[key];
|
|
46
50
|
}
|
|
47
51
|
return candidate;
|
|
@@ -50,14 +54,16 @@ export function coerce(raw: unknown): ConfigCandidate | undefined {
|
|
|
50
54
|
export function validateConfig(candidate: ConfigCandidate): ConfigValidation {
|
|
51
55
|
if (candidate.blockIsObject !== true) return { ok: false, error: "providerStallWatchdog must be an object" };
|
|
52
56
|
if (typeof candidate.enabled !== "boolean") return { ok: false, error: "enabled must be a boolean" };
|
|
57
|
+
if (!isTimerDelay(candidate.firstEventMs)) return { ok: false, error: "firstEventMs must be a positive timer delay" };
|
|
53
58
|
if (!isTimerDelay(candidate.warningMs)) return { ok: false, error: "warningMs must be a positive timer delay" };
|
|
54
59
|
if (!isTimerDelay(candidate.recoveryMs)) return { ok: false, error: "recoveryMs must be a positive timer delay" };
|
|
55
60
|
if (candidate.warningMs >= candidate.recoveryMs) return { ok: false, error: "warningMs must be less than recoveryMs" };
|
|
56
|
-
if (!
|
|
61
|
+
if (!isNonNegativeInteger(candidate.maxStallRetries)) return { ok: false, error: "maxStallRetries must be a non-negative integer" };
|
|
57
62
|
return {
|
|
58
63
|
ok: true,
|
|
59
64
|
config: {
|
|
60
65
|
enabled: candidate.enabled,
|
|
66
|
+
firstEventMs: candidate.firstEventMs,
|
|
61
67
|
warningMs: candidate.warningMs,
|
|
62
68
|
recoveryMs: candidate.recoveryMs,
|
|
63
69
|
maxStallRetries: candidate.maxStallRetries,
|
|
@@ -69,8 +75,8 @@ function isTimerDelay(value: unknown): value is number {
|
|
|
69
75
|
return typeof value === "number" && Number.isFinite(value) && Number.isInteger(value) && value > 0 && value <= MAX_TIMER_MS;
|
|
70
76
|
}
|
|
71
77
|
|
|
72
|
-
function
|
|
73
|
-
return typeof value === "number" && Number.isInteger(value) && value
|
|
78
|
+
function isNonNegativeInteger(value: unknown): value is number {
|
|
79
|
+
return typeof value === "number" && Number.isInteger(value) && value >= 0;
|
|
74
80
|
}
|
|
75
81
|
|
|
76
82
|
/** Pi's own resolution is `settings.retry?.maxRetries ?? 3` over the same layered settings.json files. */
|
|
@@ -80,7 +86,7 @@ export function resolveRetryMaxRetries(cwd: string): number {
|
|
|
80
86
|
const retry = readSettings(path)?.retry;
|
|
81
87
|
if (retry === null || typeof retry !== "object" || Array.isArray(retry)) continue;
|
|
82
88
|
const value = (retry as Record<string, unknown>).maxRetries;
|
|
83
|
-
if (
|
|
89
|
+
if (isNonNegativeInteger(value)) maxRetries = value;
|
|
84
90
|
}
|
|
85
91
|
return maxRetries;
|
|
86
92
|
}
|
|
@@ -100,8 +106,11 @@ const defaultRuntime: WatchdogRuntime = {
|
|
|
100
106
|
};
|
|
101
107
|
|
|
102
108
|
const DEGRADATION_NOTICE = "The stalled request was stopped, but Pi did not start an automatic retry. Retry may be disabled, exhausted, or incompatible; submit the message again to retry manually.";
|
|
109
|
+
// Reduces, but cannot eliminate, the hang when an aborted provider operation never terminates;
|
|
110
|
+
// undici's headersTimeout/bodyTimeout stay the backstop past this point.
|
|
111
|
+
const ABORT_GRACE_MS = 10_000;
|
|
103
112
|
|
|
104
|
-
type Timer = { warning?: unknown; recovery?: unknown };
|
|
113
|
+
type Timer = { firstEvent?: unknown; warning?: unknown; recovery?: unknown; abortGuard?: unknown };
|
|
105
114
|
|
|
106
115
|
function formatElapsed(ms: number): string {
|
|
107
116
|
if (ms % 60_000 === 0) return `${ms / 60_000}m`;
|
|
@@ -109,6 +118,8 @@ function formatElapsed(ms: number): string {
|
|
|
109
118
|
return `${ms}ms`;
|
|
110
119
|
}
|
|
111
120
|
|
|
121
|
+
const ABORT_STUCK_NOTICE = `The stalled request did not stop within ${formatElapsed(ABORT_GRACE_MS)} of being aborted; the provider connection is unresponsive. No automatic retry will run - the turn will not end until the HTTP idle timeout expires.`;
|
|
122
|
+
|
|
112
123
|
function warningNotice(config: WatchdogConfig): string {
|
|
113
124
|
return `No model progress for ${formatElapsed(config.warningMs)}; aborting and asking Pi to retry in ${formatElapsed(config.recoveryMs - config.warningMs)} (Esc aborts now)`;
|
|
114
125
|
}
|
|
@@ -117,9 +128,20 @@ function exhaustedNotice(config: WatchdogConfig): string {
|
|
|
117
128
|
return `Stall retry budget (${config.maxStallRetries}) exhausted; aborting without another automatic retry. Submit the message again manually.`;
|
|
118
129
|
}
|
|
119
130
|
|
|
131
|
+
function firstEventRetryNotice(config: WatchdogConfig): string {
|
|
132
|
+
return `Provider sent no response for ${formatElapsed(config.firstEventMs)}; stopping and retrying the request.`;
|
|
133
|
+
}
|
|
134
|
+
|
|
135
|
+
function firstEventExhaustedNotice(config: WatchdogConfig): string {
|
|
136
|
+
return `Provider sent no response for ${formatElapsed(config.firstEventMs)} and the stall-retry budget is spent; the request was stopped.`;
|
|
137
|
+
}
|
|
138
|
+
|
|
120
139
|
export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRuntime): (pi: ExtensionAPI) => void {
|
|
121
140
|
return (pi) => {
|
|
122
|
-
let
|
|
141
|
+
let midStreamEnabled = false;
|
|
142
|
+
let firstEventSeen = false;
|
|
143
|
+
let hasUI = false;
|
|
144
|
+
let pendingTimeoutReason: string | undefined;
|
|
123
145
|
let activeRun = false;
|
|
124
146
|
let disabled = false;
|
|
125
147
|
let config: WatchdogConfig | undefined;
|
|
@@ -127,22 +149,25 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
127
149
|
let activeGeneration: number | undefined;
|
|
128
150
|
let lastSemanticAt = 0;
|
|
129
151
|
let warned = false;
|
|
130
|
-
let epoch = 0;
|
|
131
152
|
let deadlineEpoch = 0;
|
|
132
153
|
let timers: Timer = {};
|
|
133
154
|
let removeSignalListener: (() => void) | undefined;
|
|
134
155
|
let ui: { notify(text: string, type?: string): void } | undefined;
|
|
135
156
|
let watchdogAbortedGeneration: number | undefined;
|
|
136
|
-
let timeoutConversionPending = false;
|
|
137
157
|
let stallRetriesUsed = 0;
|
|
138
158
|
let continuationStarted = false;
|
|
139
159
|
let convertedTimeout = false;
|
|
140
160
|
|
|
141
161
|
const clearTimers = () => {
|
|
142
|
-
|
|
143
|
-
|
|
162
|
+
for (const key of ["firstEvent", "warning", "recovery", "abortGuard"] as const) {
|
|
163
|
+
if (timers[key] !== undefined) runtime.clearTimeout(timers[key]);
|
|
164
|
+
}
|
|
144
165
|
timers = {};
|
|
145
166
|
};
|
|
167
|
+
const announce = (text: string, type?: string) => {
|
|
168
|
+
ui?.notify(text, type);
|
|
169
|
+
if (!hasUI) console.warn(text);
|
|
170
|
+
};
|
|
146
171
|
const clear = () => {
|
|
147
172
|
clearTimers();
|
|
148
173
|
removeSignalListener?.();
|
|
@@ -153,20 +178,81 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
153
178
|
const resetRunState = () => {
|
|
154
179
|
disarm();
|
|
155
180
|
activeRun = false;
|
|
156
|
-
pendingInteractive = false;
|
|
157
181
|
stallRetriesUsed = 0;
|
|
158
182
|
continuationStarted = false;
|
|
159
183
|
convertedTimeout = false;
|
|
160
184
|
watchdogAbortedGeneration = undefined;
|
|
161
|
-
|
|
185
|
+
pendingTimeoutReason = undefined;
|
|
186
|
+
firstEventSeen = false;
|
|
187
|
+
};
|
|
188
|
+
const armAbortGuard = (capturedGeneration: number) => {
|
|
189
|
+
if (timers.abortGuard !== undefined) runtime.clearTimeout(timers.abortGuard);
|
|
190
|
+
timers.abortGuard = runtime.setTimeout(() => {
|
|
191
|
+
if (capturedGeneration !== activeGeneration || !activeRun) return;
|
|
192
|
+
announce(ABORT_STUCK_NOTICE, "error");
|
|
193
|
+
// Only the timers go: the generation stays armed so a message_end arriving after the grace
|
|
194
|
+
// period still converts the abort the stall retry already paid for.
|
|
195
|
+
clearTimers();
|
|
196
|
+
}, ABORT_GRACE_MS);
|
|
162
197
|
};
|
|
163
|
-
|
|
198
|
+
// A stream event on a generation the watchdog already aborted: the stall cycle is over for this
|
|
199
|
+
// request, so nothing re-enters it. The bytes only prove the connection was alive at this instant,
|
|
200
|
+
// so the guard is re-armed rather than cleared - a wedge right after a straggler still escalates.
|
|
201
|
+
const postAbortStreamEvent = () => {
|
|
202
|
+
if (activeGeneration === undefined || watchdogAbortedGeneration !== activeGeneration) return false;
|
|
203
|
+
armAbortGuard(activeGeneration);
|
|
204
|
+
return true;
|
|
205
|
+
};
|
|
206
|
+
const abortStall = (
|
|
207
|
+
ctx: { abort(): void },
|
|
208
|
+
capturedGeneration: number,
|
|
209
|
+
notices: { retry: () => string; exhausted: () => string },
|
|
210
|
+
reason: string,
|
|
211
|
+
) => {
|
|
212
|
+
if (!config) return;
|
|
213
|
+
clearTimers();
|
|
214
|
+
warned = false;
|
|
215
|
+
watchdogAbortedGeneration = capturedGeneration;
|
|
216
|
+
// Armed before ctx.abort() so a synchronous teardown inside it cannot orphan the timer;
|
|
217
|
+
// the generation check makes the callback a no-op if that teardown disarmed the watchdog.
|
|
218
|
+
armAbortGuard(capturedGeneration);
|
|
219
|
+
if (stallRetriesUsed >= config.maxStallRetries) {
|
|
220
|
+
announce(notices.exhausted());
|
|
221
|
+
ctx.abort();
|
|
222
|
+
return;
|
|
223
|
+
}
|
|
224
|
+
pendingTimeoutReason = reason;
|
|
225
|
+
stallRetriesUsed += 1;
|
|
226
|
+
announce(notices.retry());
|
|
227
|
+
ctx.abort();
|
|
228
|
+
};
|
|
229
|
+
const armFirstEvent = (ctx: { abort(): void }) => {
|
|
164
230
|
if (activeGeneration === undefined || !config) return;
|
|
231
|
+
const cfg = config;
|
|
232
|
+
const capturedGeneration = activeGeneration;
|
|
233
|
+
const capturedDeadlineEpoch = ++deadlineEpoch;
|
|
234
|
+
const threshold = cfg.firstEventMs;
|
|
235
|
+
const run = () => {
|
|
236
|
+
if (capturedGeneration !== activeGeneration || capturedDeadlineEpoch !== deadlineEpoch || !activeRun || firstEventSeen) return;
|
|
237
|
+
const elapsed = runtime.now() - lastSemanticAt;
|
|
238
|
+
if (elapsed < threshold) {
|
|
239
|
+
timers.firstEvent = runtime.setTimeout(run, threshold - elapsed);
|
|
240
|
+
return;
|
|
241
|
+
}
|
|
242
|
+
abortStall(ctx, capturedGeneration, {
|
|
243
|
+
retry: () => firstEventRetryNotice(cfg),
|
|
244
|
+
exhausted: () => firstEventExhaustedNotice(cfg),
|
|
245
|
+
}, `Provider first-event timeout after ${cfg.firstEventMs} ms without a stream event`);
|
|
246
|
+
};
|
|
247
|
+
timers.firstEvent = runtime.setTimeout(run, threshold);
|
|
248
|
+
};
|
|
249
|
+
const schedule = (ctx: { abort(): void }) => {
|
|
250
|
+
if (activeGeneration === undefined || !config) return;
|
|
251
|
+
const cfg = config;
|
|
165
252
|
const capturedGeneration = activeGeneration;
|
|
166
|
-
const capturedEpoch = epoch;
|
|
167
253
|
const capturedDeadlineEpoch = ++deadlineEpoch;
|
|
168
254
|
const run = (kind: "warning" | "recovery", threshold: number) => () => {
|
|
169
|
-
if (
|
|
255
|
+
if (capturedGeneration !== activeGeneration || capturedDeadlineEpoch !== deadlineEpoch || !activeRun) return;
|
|
170
256
|
const elapsed = runtime.now() - lastSemanticAt;
|
|
171
257
|
if (elapsed < threshold) {
|
|
172
258
|
timers[kind] = runtime.setTimeout(run(kind, threshold), threshold - elapsed);
|
|
@@ -174,51 +260,38 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
174
260
|
}
|
|
175
261
|
if (kind === "warning" && !warned) {
|
|
176
262
|
warned = true;
|
|
177
|
-
|
|
263
|
+
announce(warningNotice(cfg), "warning");
|
|
178
264
|
}
|
|
179
265
|
if (kind === "recovery") {
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
|
|
184
|
-
ctx.abort();
|
|
185
|
-
return;
|
|
186
|
-
}
|
|
187
|
-
watchdogAbortedGeneration = capturedGeneration;
|
|
188
|
-
timeoutConversionPending = true;
|
|
189
|
-
stallRetriesUsed += 1;
|
|
190
|
-
ui?.notify(`No model progress for ${formatElapsed(elapsed)}; aborting now. Pi will retry (${stallRetriesUsed}/${config.maxStallRetries}) if retry is enabled and capacity remains. Pending follow-ups are returned to the editor.`);
|
|
191
|
-
ctx.abort();
|
|
266
|
+
abortStall(ctx, capturedGeneration, {
|
|
267
|
+
retry: () => `No model progress for ${formatElapsed(elapsed)}; aborting now. Pi will retry (${stallRetriesUsed}/${cfg.maxStallRetries}) if retry is enabled and capacity remains. Pending follow-ups are returned to the editor.`,
|
|
268
|
+
exhausted: () => exhaustedNotice(cfg),
|
|
269
|
+
}, `Provider semantic timeout after ${cfg.recoveryMs} ms without progress`);
|
|
192
270
|
}
|
|
193
271
|
};
|
|
194
|
-
timers.warning = runtime.setTimeout(run("warning",
|
|
195
|
-
timers.recovery = runtime.setTimeout(run("recovery",
|
|
272
|
+
timers.warning = runtime.setTimeout(run("warning", cfg.warningMs), cfg.warningMs);
|
|
273
|
+
timers.recovery = runtime.setTimeout(run("recovery", cfg.recoveryMs), cfg.recoveryMs);
|
|
196
274
|
};
|
|
197
275
|
|
|
198
|
-
pi.on("
|
|
199
|
-
if (
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
if (!
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
return;
|
|
276
|
+
pi.on("before_provider_request", (_event, ctx) => {
|
|
277
|
+
if (disabled) return;
|
|
278
|
+
ui = ctx.ui;
|
|
279
|
+
hasUI = ctx.hasUI;
|
|
280
|
+
if (!config) {
|
|
281
|
+
const resolved = resolveWatchdogConfig(ctx.cwd);
|
|
282
|
+
if (!resolved.ok) {
|
|
283
|
+
disabled = true;
|
|
284
|
+
announce(`providerStallWatchdog disabled: ${resolved.error}`, "warning");
|
|
285
|
+
return;
|
|
286
|
+
}
|
|
287
|
+
config = resolved.config;
|
|
211
288
|
}
|
|
212
|
-
config = resolved.config;
|
|
213
289
|
activeRun = config.enabled;
|
|
214
|
-
|
|
215
|
-
pi.on("before_provider_request", (_event, ctx) => {
|
|
216
|
-
if (!activeRun || ctx.mode !== "tui" || !config) return;
|
|
290
|
+
if (!activeRun) return;
|
|
217
291
|
disarm();
|
|
218
292
|
if (convertedTimeout) continuationStarted = true;
|
|
219
293
|
activeGeneration = ++generation;
|
|
220
294
|
lastSemanticAt = runtime.now();
|
|
221
|
-
ui = ctx.ui;
|
|
222
295
|
const target = ctx.signal;
|
|
223
296
|
if (target) {
|
|
224
297
|
const listener = () => {
|
|
@@ -227,11 +300,30 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
227
300
|
target.addEventListener("abort", listener, { once: true });
|
|
228
301
|
removeSignalListener = () => target.removeEventListener("abort", listener);
|
|
229
302
|
}
|
|
303
|
+
midStreamEnabled = ctx.mode === "tui";
|
|
304
|
+
firstEventSeen = false;
|
|
305
|
+
armFirstEvent(ctx);
|
|
306
|
+
});
|
|
307
|
+
pi.on("message_start", (event, ctx) => {
|
|
308
|
+
// Pi fires message_start for user and toolResult messages too; only an assistant one is
|
|
309
|
+
// provider traffic, so the role check must stay ahead of the liveness re-arm below.
|
|
310
|
+
if (event.message.role !== "assistant") return;
|
|
311
|
+
if (postAbortStreamEvent()) return;
|
|
312
|
+
if (!activeRun || activeGeneration === undefined || firstEventSeen) return;
|
|
313
|
+
firstEventSeen = true;
|
|
314
|
+
if (timers.firstEvent !== undefined) {
|
|
315
|
+
runtime.clearTimeout(timers.firstEvent);
|
|
316
|
+
timers.firstEvent = undefined;
|
|
317
|
+
}
|
|
318
|
+
if (!midStreamEnabled || !config) return;
|
|
319
|
+
lastSemanticAt = runtime.now();
|
|
320
|
+
warned = false;
|
|
230
321
|
schedule(ctx);
|
|
231
322
|
});
|
|
232
323
|
pi.on("message_update", (event, ctx) => {
|
|
324
|
+
if (postAbortStreamEvent()) return;
|
|
233
325
|
const update = event.assistantMessageEvent;
|
|
234
|
-
if (activeGeneration === undefined || !(update.type === "text_delta" || update.type === "thinking_delta" || update.type === "toolcall_delta") || update.delta.length === 0) return;
|
|
326
|
+
if (!midStreamEnabled || !firstEventSeen || activeGeneration === undefined || !(update.type === "text_delta" || update.type === "thinking_delta" || update.type === "toolcall_delta") || update.delta.length === 0) return;
|
|
235
327
|
lastSemanticAt = runtime.now();
|
|
236
328
|
warned = false;
|
|
237
329
|
clearTimers();
|
|
@@ -239,25 +331,25 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
239
331
|
});
|
|
240
332
|
pi.on("message_end", (event) => {
|
|
241
333
|
if (event.message.role !== "assistant") return;
|
|
242
|
-
const
|
|
243
|
-
|
|
244
|
-
|
|
334
|
+
const errorMessage = pendingTimeoutReason;
|
|
335
|
+
pendingTimeoutReason = undefined;
|
|
336
|
+
const matchesWatchdogAbort = errorMessage !== undefined
|
|
337
|
+
&& event.message.stopReason === "aborted"
|
|
338
|
+
&& activeGeneration === watchdogAbortedGeneration;
|
|
245
339
|
disarm();
|
|
246
340
|
// Mirror Pi's retry loop, which resets its attempt counter on any successful assistant turn.
|
|
247
341
|
if (event.message.stopReason !== "aborted" && event.message.stopReason !== "error") stallRetriesUsed = 0;
|
|
248
|
-
if (!matchesWatchdogAbort
|
|
249
|
-
timeoutConversionPending = false;
|
|
342
|
+
if (!matchesWatchdogAbort) return;
|
|
250
343
|
convertedTimeout = true;
|
|
251
344
|
continuationStarted = false;
|
|
252
|
-
return { message: { ...event.message, stopReason: "error", errorMessage
|
|
345
|
+
return { message: { ...event.message, stopReason: "error", errorMessage } };
|
|
253
346
|
});
|
|
254
347
|
pi.on("agent_end", () => disarm());
|
|
255
348
|
pi.on("agent_settled", () => {
|
|
256
|
-
if (convertedTimeout && !continuationStarted)
|
|
349
|
+
if (convertedTimeout && !continuationStarted) announce(DEGRADATION_NOTICE);
|
|
257
350
|
resetRunState();
|
|
258
351
|
});
|
|
259
352
|
pi.on("session_shutdown", () => {
|
|
260
|
-
epoch += 1;
|
|
261
353
|
resetRunState();
|
|
262
354
|
config = undefined;
|
|
263
355
|
disabled = false;
|