pi-quiver 6.3.1 → 6.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +4 -0
- package/README.md +9 -3
- package/extensions/provider-stall-watchdog.ts +70 -17
- package/lib/extension-config.ts +1 -1
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -8,6 +8,10 @@ Published to npm as `pi-quiver` (`pi install npm:pi-quiver`). Pushing a
|
|
|
8
8
|
via OIDC trusted publishing. The release helper at
|
|
9
9
|
`.agents/skills/release/scripts/release.sh` cuts the tag; CI publishes.
|
|
10
10
|
|
|
11
|
+
## v6.4.0 - 2026-09-28
|
|
12
|
+
|
|
13
|
+
- provider-stall-watchdog: per-model threshold overrides via a `models` map on `quiver.providerStallWatchdog` - glob keys like `lmstudio/*` override `firstEventMs`/`warningMs`/`recoveryMs` for matching models, so slow local servers get patience without raising the global defaults (#18).
|
|
14
|
+
|
|
11
15
|
## v6.3.1 - 2026-09-25
|
|
12
16
|
|
|
13
17
|
- `provider-stall-watchdog` recovers again on pi >= 0.86 ([#23](https://github.com/jjuraszek/pi-quiver/issues/23)): pi's session abort now fences the run so its retry loop never runs after a watchdog abort; the watchdog omits the aborted attempt from the model's context at `turn_end` and re-drives the request itself via a hidden custom message after pi's backoff, honoring pi's `retry.enabled` / `retry.baseDelayMs` / `retry.maxAgentDelayMs` live and the existing `maxStallRetries` cap. Print/json await the backoff; TUI/RPC show `Retrying (n/m) in Ns... (Esc to cancel)` in the status bar on the timer path. In TUI, Esc, a new prompt, a tree switch, or compaction cancels the pending retry; RPC cancels on a new prompt, not Esc. Behavior guide: `doc/provider-stall-watchdog.md`. Dev dependencies on `@earendil-works/*` move to `^0.87.1`.
|
package/README.md
CHANGED
|
@@ -186,7 +186,7 @@ a new setting is registered there or it warns as unknown.
|
|
|
186
186
|
```text
|
|
187
187
|
Warning: pi-quiver settings (/Users/x/.pi/agent/settings.json): unknown or misplaced keys - unknown ones fall back to defaults
|
|
188
188
|
"providerStallWatchdog" at top level - move under "quiver"
|
|
189
|
-
"quiver.providerStallWatchdog.timeoutMs" - unknown; accepted: enabled, firstEventMs, warningMs, recoveryMs, maxStallRetries
|
|
189
|
+
"quiver.providerStallWatchdog.timeoutMs" - unknown; accepted: enabled, firstEventMs, warningMs, recoveryMs, maxStallRetries, models
|
|
190
190
|
```
|
|
191
191
|
|
|
192
192
|
Worked mixed-shape example: global `settings.json` has flat
|
|
@@ -218,7 +218,10 @@ not a pi-quiver setting and is never nested):
|
|
|
218
218
|
"firstEventMs": 20000,
|
|
219
219
|
"warningMs": 120000,
|
|
220
220
|
"recoveryMs": 240000,
|
|
221
|
-
"maxStallRetries": 3
|
|
221
|
+
"maxStallRetries": 3,
|
|
222
|
+
"models": {
|
|
223
|
+
"lmstudio/*": { "firstEventMs": 600000, "recoveryMs": 600000 }
|
|
224
|
+
}
|
|
222
225
|
}
|
|
223
226
|
},
|
|
224
227
|
"retry": {
|
|
@@ -236,13 +239,16 @@ not a pi-quiver setting and is never nested):
|
|
|
236
239
|
| `warningMs` | `120000` | mid-stream, `ctx.mode === "tui"` only | Silence since the last non-empty text/thinking/toolcall delta; notifies. |
|
|
237
240
|
| `recoveryMs` | `240000` | mid-stream, `ctx.mode === "tui"` only | Same clock; aborts and converts. Must be `> warningMs`. |
|
|
238
241
|
| `maxStallRetries` | layered `retry.maxRetries`, else `3` | shared by both tiers | Watchdog aborts that may convert to a retryable error before stopping. |
|
|
242
|
+
| `models` | `{}` | per-model overrides of the three thresholds | Glob keys match `provider/model` (case-insensitive, `*` matches any run of characters, first match wins); each entry may override `firstEventMs`, `warningMs`, and/or `recoveryMs`. |
|
|
239
243
|
|
|
240
244
|
`providerStallWatchdog` is OFF by default. Once enabled it arms in two tiers per provider request:
|
|
241
245
|
|
|
242
246
|
- **Pre-first-event (`firstEventMs`).** Armed at every provider request, in every mode and from every origin - including extension-triggered turns that never emit `before_agent_start` - and cleared by the first assistant `message_start`. On expiry the request is aborted and, budget permitting, converted to a retryable error and re-driven after pi's backoff when `retry.enabled` is true, so an unresponsive request recovers in ~22s (20s detection + 2s backoff) instead of the ~240s it took when only the mid-stream tier existed.
|
|
243
247
|
- **Mid-stream (`warningMs` / `recoveryMs`).** Armed from the first assistant `message_start` onward, and only when `ctx.mode === "tui"`. Aborting mid-generation discards billed output tokens and an unattended run has nobody to read the warning, so headless mid-stream silence deliberately falls through to the transport timeout instead.
|
|
244
248
|
|
|
245
|
-
|
|
249
|
+
`models` retunes the three thresholds per model. A request's effective thresholds are the base knobs overlaid with the first entry whose glob matches its `provider/model` label - `"lmstudio/*"` covers a whole local server, `"openai/gpt-5.4"` one remote model. An override that would leave the merged `warningMs >= recoveryMs` is invalid and fails closed at startup like any other invalid watchdog config.
|
|
250
|
+
|
|
251
|
+
**Raise `firstEventMs` if your provider is legitimately slow to first event** - per model via `models` when only one provider is slow. Queueing gateways, throttled endpoints, and busy single-slot local model servers can hold the connection for well over 20s before their first stream event; every false abort re-uploads the whole context and spends one stall retry.
|
|
246
252
|
|
|
247
253
|
**Leave pi's own `httpIdleTimeoutMs` (default `300000`) at its default.** It is the transport backstop, and a single value drives undici's `headersTimeout` *and* `bodyTimeout` - lowering it to get fast pre-stream failure also truncates legitimate mid-stream gaps. `firstEventMs` is the knob for pre-stream silence.
|
|
248
254
|
|
|
@@ -7,14 +7,20 @@ export const DEFAULT_CONFIG = {
|
|
|
7
7
|
firstEventMs: 20_000,
|
|
8
8
|
warningMs: 120_000,
|
|
9
9
|
recoveryMs: 240_000,
|
|
10
|
+
models: {},
|
|
10
11
|
} as const;
|
|
11
12
|
|
|
12
|
-
export type
|
|
13
|
-
enabled: boolean;
|
|
13
|
+
export type WatchdogThresholds = {
|
|
14
14
|
firstEventMs: number;
|
|
15
15
|
warningMs: number;
|
|
16
16
|
recoveryMs: number;
|
|
17
|
+
};
|
|
18
|
+
|
|
19
|
+
export type WatchdogConfig = WatchdogThresholds & {
|
|
20
|
+
enabled: boolean;
|
|
17
21
|
maxStallRetries: number;
|
|
22
|
+
/** Per-model threshold overrides keyed by glob; see thresholdsFor. */
|
|
23
|
+
models: Record<string, Partial<WatchdogThresholds>>;
|
|
18
24
|
};
|
|
19
25
|
|
|
20
26
|
export type WatchdogRuntime = {
|
|
@@ -30,6 +36,7 @@ export type ConfigCandidate = {
|
|
|
30
36
|
warningMs?: unknown;
|
|
31
37
|
recoveryMs?: unknown;
|
|
32
38
|
maxStallRetries?: unknown;
|
|
39
|
+
models?: unknown;
|
|
33
40
|
};
|
|
34
41
|
|
|
35
42
|
export type ConfigValidation =
|
|
@@ -45,12 +52,17 @@ export function coerce(raw: unknown): ConfigCandidate | undefined {
|
|
|
45
52
|
|
|
46
53
|
const source = raw as Record<string, unknown>;
|
|
47
54
|
const candidate: ConfigCandidate = { blockIsObject: true };
|
|
48
|
-
for (const key of ["enabled", "firstEventMs", "warningMs", "recoveryMs", "maxStallRetries"] as const) {
|
|
55
|
+
for (const key of ["enabled", "firstEventMs", "warningMs", "recoveryMs", "maxStallRetries", "models"] as const) {
|
|
49
56
|
if (Object.hasOwn(source, key)) candidate[key] = source[key];
|
|
50
57
|
}
|
|
51
58
|
return candidate;
|
|
52
59
|
}
|
|
53
60
|
|
|
61
|
+
const THRESHOLD_KEYS = ["firstEventMs", "warningMs", "recoveryMs"] as const;
|
|
62
|
+
|
|
63
|
+
const isRecord = (value: unknown): value is Record<string, unknown> =>
|
|
64
|
+
value !== null && typeof value === "object" && !Array.isArray(value);
|
|
65
|
+
|
|
54
66
|
export function validateConfig(candidate: ConfigCandidate): ConfigValidation {
|
|
55
67
|
if (candidate.blockIsObject !== true) return { ok: false, error: "providerStallWatchdog must be an object" };
|
|
56
68
|
if (typeof candidate.enabled !== "boolean") return { ok: false, error: "enabled must be a boolean" };
|
|
@@ -59,6 +71,21 @@ export function validateConfig(candidate: ConfigCandidate): ConfigValidation {
|
|
|
59
71
|
if (!isTimerDelay(candidate.recoveryMs)) return { ok: false, error: "recoveryMs must be a positive timer delay" };
|
|
60
72
|
if (candidate.warningMs >= candidate.recoveryMs) return { ok: false, error: "warningMs must be less than recoveryMs" };
|
|
61
73
|
if (!isNonNegativeInteger(candidate.maxStallRetries)) return { ok: false, error: "maxStallRetries must be a non-negative integer" };
|
|
74
|
+
if (!isRecord(candidate.models)) return { ok: false, error: "models must be an object" };
|
|
75
|
+
for (const [pattern, override] of Object.entries(candidate.models)) {
|
|
76
|
+
if (!isRecord(override)) return { ok: false, error: `models["${pattern}"] must be an object` };
|
|
77
|
+
for (const key of Object.keys(override)) {
|
|
78
|
+
if (!(THRESHOLD_KEYS as readonly string[]).includes(key)) {
|
|
79
|
+
return { ok: false, error: `models["${pattern}"] has unknown key "${key}"; accepted: ${THRESHOLD_KEYS.join(", ")}` };
|
|
80
|
+
}
|
|
81
|
+
if (!isTimerDelay(override[key])) return { ok: false, error: `models["${pattern}"].${key} must be a positive timer delay` };
|
|
82
|
+
}
|
|
83
|
+
const mergedWarning = (override.warningMs ?? candidate.warningMs) as number;
|
|
84
|
+
const mergedRecovery = (override.recoveryMs ?? candidate.recoveryMs) as number;
|
|
85
|
+
if (mergedWarning >= mergedRecovery) {
|
|
86
|
+
return { ok: false, error: `models["${pattern}"] leaves warningMs (${mergedWarning}) >= recoveryMs (${mergedRecovery})` };
|
|
87
|
+
}
|
|
88
|
+
}
|
|
62
89
|
return {
|
|
63
90
|
ok: true,
|
|
64
91
|
config: {
|
|
@@ -67,6 +94,7 @@ export function validateConfig(candidate: ConfigCandidate): ConfigValidation {
|
|
|
67
94
|
warningMs: candidate.warningMs,
|
|
68
95
|
recoveryMs: candidate.recoveryMs,
|
|
69
96
|
maxStallRetries: candidate.maxStallRetries,
|
|
97
|
+
models: candidate.models as Record<string, Partial<WatchdogThresholds>>,
|
|
70
98
|
},
|
|
71
99
|
};
|
|
72
100
|
}
|
|
@@ -144,20 +172,41 @@ function formatElapsed(ms: number): string {
|
|
|
144
172
|
|
|
145
173
|
const ABORT_STUCK_NOTICE = `The stalled request did not stop within ${formatElapsed(ABORT_GRACE_MS)} of being aborted; the provider connection is unresponsive. No automatic retry will run - the turn will not end until the HTTP idle timeout expires.`;
|
|
146
174
|
|
|
147
|
-
function warningNotice(
|
|
148
|
-
return `No model progress for ${formatElapsed(
|
|
175
|
+
function warningNotice(thresholds: WatchdogThresholds): string {
|
|
176
|
+
return `No model progress for ${formatElapsed(thresholds.warningMs)}; aborting and asking Pi to retry in ${formatElapsed(thresholds.recoveryMs - thresholds.warningMs)} (Esc aborts now)`;
|
|
149
177
|
}
|
|
150
178
|
|
|
151
179
|
function exhaustedNotice(config: WatchdogConfig): string {
|
|
152
180
|
return `Stall retry budget (${config.maxStallRetries}) exhausted; aborting without another automatic retry. Submit the message again manually.`;
|
|
153
181
|
}
|
|
154
182
|
|
|
155
|
-
function firstEventRetryNotice(
|
|
156
|
-
return `Provider sent no response for ${formatElapsed(
|
|
183
|
+
function firstEventRetryNotice(thresholds: WatchdogThresholds): string {
|
|
184
|
+
return `Provider sent no response for ${formatElapsed(thresholds.firstEventMs)}; stopping and retrying the request.`;
|
|
185
|
+
}
|
|
186
|
+
|
|
187
|
+
function firstEventExhaustedNotice(thresholds: WatchdogThresholds): string {
|
|
188
|
+
return `Provider sent no response for ${formatElapsed(thresholds.firstEventMs)} and the stall-retry budget is spent; the request was stopped.`;
|
|
157
189
|
}
|
|
158
190
|
|
|
159
|
-
|
|
160
|
-
|
|
191
|
+
/** A glob where `*` matches any run of characters; matching is case-insensitive. */
|
|
192
|
+
function modelPattern(pattern: string): RegExp {
|
|
193
|
+
const source = pattern.split("*").map((part) => part.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")).join(".*");
|
|
194
|
+
return new RegExp(`^${source}$`, "i");
|
|
195
|
+
}
|
|
196
|
+
|
|
197
|
+
/**
|
|
198
|
+
* Effective thresholds for one request: the base knobs overlaid with the
|
|
199
|
+
* first `models` entry whose glob matches the request's `provider/model`
|
|
200
|
+
* label. First match wins; later entries for the same model are ignored.
|
|
201
|
+
*/
|
|
202
|
+
export function thresholdsFor(config: WatchdogConfig, model: { provider: string; id: string } | undefined): WatchdogThresholds {
|
|
203
|
+
const base: WatchdogThresholds = { firstEventMs: config.firstEventMs, warningMs: config.warningMs, recoveryMs: config.recoveryMs };
|
|
204
|
+
if (model === undefined) return base;
|
|
205
|
+
const label = `${model.provider}/${model.id}`;
|
|
206
|
+
for (const [pattern, override] of Object.entries(config.models)) {
|
|
207
|
+
if (modelPattern(pattern).test(label)) return { ...base, ...override };
|
|
208
|
+
}
|
|
209
|
+
return base;
|
|
161
210
|
}
|
|
162
211
|
|
|
163
212
|
export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRuntime): (pi: ExtensionAPI) => void {
|
|
@@ -171,6 +220,7 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
171
220
|
let config: WatchdogConfig | undefined;
|
|
172
221
|
let generation = 0;
|
|
173
222
|
let activeGeneration: number | undefined;
|
|
223
|
+
let activeModel: { provider: string; id: string } | undefined;
|
|
174
224
|
let lastSemanticAt = 0;
|
|
175
225
|
let warned = false;
|
|
176
226
|
let deadlineEpoch = 0;
|
|
@@ -281,9 +331,10 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
281
331
|
const armFirstEvent = (ctx: { abort(): void }) => {
|
|
282
332
|
if (activeGeneration === undefined || !config) return;
|
|
283
333
|
const cfg = config;
|
|
334
|
+
const thresholds = thresholdsFor(cfg, activeModel);
|
|
284
335
|
const capturedGeneration = activeGeneration;
|
|
285
336
|
const capturedDeadlineEpoch = ++deadlineEpoch;
|
|
286
|
-
const threshold =
|
|
337
|
+
const threshold = thresholds.firstEventMs;
|
|
287
338
|
const run = () => {
|
|
288
339
|
if (capturedGeneration !== activeGeneration || capturedDeadlineEpoch !== deadlineEpoch || !activeRun || firstEventSeen) return;
|
|
289
340
|
const elapsed = runtime.now() - lastSemanticAt;
|
|
@@ -292,15 +343,16 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
292
343
|
return;
|
|
293
344
|
}
|
|
294
345
|
abortStall(ctx, capturedGeneration, {
|
|
295
|
-
retry: () => firstEventRetryNotice(
|
|
296
|
-
exhausted: () => firstEventExhaustedNotice(
|
|
297
|
-
}, `Provider first-event timeout after ${
|
|
346
|
+
retry: () => firstEventRetryNotice(thresholds),
|
|
347
|
+
exhausted: () => firstEventExhaustedNotice(thresholds),
|
|
348
|
+
}, `Provider first-event timeout after ${thresholds.firstEventMs} ms without a stream event`);
|
|
298
349
|
};
|
|
299
350
|
timers.firstEvent = runtime.setTimeout(run, threshold);
|
|
300
351
|
};
|
|
301
352
|
const schedule = (ctx: { abort(): void }) => {
|
|
302
353
|
if (activeGeneration === undefined || !config) return;
|
|
303
354
|
const cfg = config;
|
|
355
|
+
const thresholds = thresholdsFor(cfg, activeModel);
|
|
304
356
|
const capturedGeneration = activeGeneration;
|
|
305
357
|
const capturedDeadlineEpoch = ++deadlineEpoch;
|
|
306
358
|
const run = (kind: "warning" | "recovery", threshold: number) => () => {
|
|
@@ -312,17 +364,17 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
312
364
|
}
|
|
313
365
|
if (kind === "warning" && !warned) {
|
|
314
366
|
warned = true;
|
|
315
|
-
announce(warningNotice(
|
|
367
|
+
announce(warningNotice(thresholds), "warning");
|
|
316
368
|
}
|
|
317
369
|
if (kind === "recovery") {
|
|
318
370
|
abortStall(ctx, capturedGeneration, {
|
|
319
371
|
retry: () => `No model progress for ${formatElapsed(elapsed)}; aborting now. Pi will retry (${stallRetriesUsed}/${cfg.maxStallRetries}) if retry is enabled and capacity remains. Pending follow-ups are returned to the editor.`,
|
|
320
372
|
exhausted: () => exhaustedNotice(cfg),
|
|
321
|
-
}, `Provider semantic timeout after ${
|
|
373
|
+
}, `Provider semantic timeout after ${thresholds.recoveryMs} ms without progress`);
|
|
322
374
|
}
|
|
323
375
|
};
|
|
324
|
-
timers.warning = runtime.setTimeout(run("warning",
|
|
325
|
-
timers.recovery = runtime.setTimeout(run("recovery",
|
|
376
|
+
timers.warning = runtime.setTimeout(run("warning", thresholds.warningMs), thresholds.warningMs);
|
|
377
|
+
timers.recovery = runtime.setTimeout(run("recovery", thresholds.recoveryMs), thresholds.recoveryMs);
|
|
326
378
|
};
|
|
327
379
|
|
|
328
380
|
pi.on("before_provider_request", (_event, ctx) => {
|
|
@@ -347,6 +399,7 @@ export function createProviderStallWatchdog(runtime: WatchdogRuntime = defaultRu
|
|
|
347
399
|
if (!activeRun) return;
|
|
348
400
|
disarm();
|
|
349
401
|
activeGeneration = ++generation;
|
|
402
|
+
activeModel = ctx.model;
|
|
350
403
|
lastSemanticAt = runtime.now();
|
|
351
404
|
const target = ctx.signal;
|
|
352
405
|
if (target) {
|
package/lib/extension-config.ts
CHANGED
|
@@ -46,7 +46,7 @@ export const QUIVER_CONFIG_KEYS: Record<string, readonly string[]> = {
|
|
|
46
46
|
fastMode: ["enabled"],
|
|
47
47
|
sessionAutoName: ["enabled", "ghosttyTab", "herdrTab", "rules", "deny", "revisitFirstTurn", "revisitEveryTurns"],
|
|
48
48
|
swordHeader: ["enabled"],
|
|
49
|
-
providerStallWatchdog: ["enabled", "firstEventMs", "warningMs", "recoveryMs", "maxStallRetries"],
|
|
49
|
+
providerStallWatchdog: ["enabled", "firstEventMs", "warningMs", "recoveryMs", "maxStallRetries", "models"],
|
|
50
50
|
slack: ["enabled", "cachePath", "policyPath", "userTokenEnv", "userTokenCommand", "userTokenCommandTimeoutSeconds", "botTokenEnv", "uploadThresholdChars"],
|
|
51
51
|
docToMd: DOC_TO_MD_OPTIONS.filter((o) => o.settable).map((o) => o.key),
|
|
52
52
|
};
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-quiver",
|
|
3
|
-
"version": "6.
|
|
3
|
+
"version": "6.4.0",
|
|
4
4
|
"description": "Personal pack of Pi coding-agent extensions: context-safe fetch, doc_to_md PDF/DOCX/PPTX-to-Markdown conversion, session naming, a themed ASCII startup header, Opus 4.8 fast mode, and a provider-stall watchdog.",
|
|
5
5
|
"author": "Jacek Juraszek",
|
|
6
6
|
"license": "MIT",
|