free-coding-models 0.5.64 โ 0.5.65
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +11 -0
- package/changelog/v0.5.65.md +35 -0
- package/package.json +1 -1
- package/src/core/ping.js +16 -0
- package/src/core/probe-cache.js +54 -5
- package/src/core/provider-cooldown.js +185 -0
- package/src/tui/app.js +18 -0
- package/src/tui/render-table.js +35 -2
- package/src/tui/tui-state.js +7 -0
- package/web/dist/assets/{index-5S_yM5SF.js โ index-C9i10FT8.js} +2 -2
- package/web/dist/index.html +1 -1
package/README.md
CHANGED
|
@@ -102,6 +102,17 @@ Create a free account on one provider below to get started. A few providers (`Ki
|
|
|
102
102
|
|
|
103
103
|
> ๐งน Audit cleanup: `iFlow` was removed because it shut down on April 17, 2026. `Together AI`, `Perplexity API`, `DeepInfra`, `Replicate`, `Fireworks`, `Hyperbolic`, `Hugging Face`, `SiliconFlow`, `Chutes AI` were removed from the active free catalog because they are paid, trial-credit only, too tiny to be useful, unclear as a stable free API, or tool-specific rather than a generally usable free provider. `Rovo` and `Gemini CLI` were also wiped out as tool integrations (CLI-only, not generally usable free providers).
|
|
104
104
|
|
|
105
|
+
### โ ๏ธ Health checks consume provider quota
|
|
106
|
+
|
|
107
|
+
> FCM continuously health-probes every model in the catalog so you see live latency, status, and verdict. When you've configured a provider's API key (e.g. `OPENROUTER_API_KEY`), **those probes are authenticated and count against that provider's daily quota**. Rate-limited providers like **OpenRouter** (50 req/day free, 1000 req/day at $10 spend) are particularly sensitive: a single overloaded model can burn your quota in minutes if it's re-pinged every second.
|
|
108
|
+
|
|
109
|
+
> As of v0.3.83 (issue #146), FCM now:
|
|
110
|
+
> - **Auto-pauses** a provider when its probe gets a `429` response (respects `Retry-After` header + the "try again N seconds later" message body).
|
|
111
|
+
> - **Backs off exponentially** on per-model failures (30s โ 1m โ 2m โ 5m) instead of re-pinging a broken model every cycle.
|
|
112
|
+
> - **Surfaces a footer chip** like `โธ openrouter 14h` so you can see which providers are resting.
|
|
113
|
+
|
|
114
|
+
> **If you don't need authenticated probes for a provider**, just leave its key empty โ anonymous probes still work for liveness checks on most providers, and they won't burn your personal quota.
|
|
115
|
+
|
|
105
116
|
---
|
|
106
117
|
|
|
107
118
|
### Tier scale
|
|
@@ -0,0 +1,35 @@
|
|
|
1
|
+
# Changelog v0.5.65 - 2026-07-29
|
|
2
|
+
|
|
3
|
+
### Fixed
|
|
4
|
+
|
|
5
|
+
- ๐ก๏ธ **#146 โ OpenRouter / rate-limited providers no longer burned the user's daily quota.** Health-probe pings are now throttled at two orthogonal levels so a single overloaded model can no longer wipe out an OpenRouter (or similar) daily request limit. Issue submitted by `@ia-S-on`.
|
|
6
|
+
|
|
7
|
+
- **Per-provider quota circuit-breaker** โ New module `src/core/provider-cooldown.js` exposes an in-memory pause map. When a probe returns HTTP `429`, FCM parses the `Retry-After` header (and falls back to the OpenRouter-style body `"Rate limit exceeded, please try again N seconds later."`), then pauses **every model of that provider** until the window expires. Probe pressure drops from ~30 req/min/provider to ~0 req/min/provider for the entire duration. Affected code: `src/core/ping.js` (trigger on `resp.status === 429`), `src/tui/app.js` (skip paused providers in `runPingCycle`).
|
|
8
|
+
- **Per-model broken cooldown with exponential backoff** โ `src/core/probe-cache.js` now caches `broken` entries too. `isCacheFresh()` and `getModelsDueForProbe()` honour a new `brokenCooldownMs(failureCount)` ladder (30 s โ 1 m โ 2 m โ 5 m, then plateau at 5 min). `recordProbeResults()` tracks `consecutiveFailures` (increments on `broken`, resets to `0` on `ok`). This caps worst-case probe pressure on a permanently-broken model at **~0.2 req/min** instead of the previous **~30 req/min**.
|
|
9
|
+
- **Live footer chip** โ When one or more providers are paused, the TUI footer now shows e.g. `โธ openrouter 14h ยท โธ ovhcloud 34s` (warm amber background, live countdown that auto-rolls to s/m/h/d). New state field `state.pausedProviders` (defaults to `{}`) + render-table chip fed by `pausedProviders` option. Behaviour verified on first run โ OVHcloud was correctly paused for 34 s on a real 429 during smoke testing.
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- ๐ **README health-check warning** โ New section `โ ๏ธ Health checks consume provider quota` between the providers table and the Tier scale. Explains the 1k/day OpenRouter cap, links the v0.5.65 fix points, and offers the immediate workaround (remove the provider's key from Settings โ anonymous probes still work for liveness checks).
|
|
14
|
+
- ๐งช **18 new unit tests** in `test/test.js` covering both layers (`Issue #146: provider quota circuit-breaker` + `Issue #146: per-model broken-cooldown backoff`). Three pre-existing tests in `test/probe-cache.test.js` updated to reflect the new `broken`-is-fresh-during-cooldown contract.
|
|
15
|
+
|
|
16
|
+
### Changed
|
|
17
|
+
|
|
18
|
+
- ๐ `pnpm test` suite: **797 / 797 passing** (was 779 / 779 in 0.5.64; net +18 new tests + 3 updated).
|
|
19
|
+
- ๐ `node --check` clean on every touched file; TUI smoke-tested in tmux (no runtime regression).
|
|
20
|
+
- ๐ Footer's existing probe-cache chip is unchanged in behaviour but gains the sibling paused-providers chip immediately to its right.
|
|
21
|
+
|
|
22
|
+
### Migration / workaround
|
|
23
|
+
|
|
24
|
+
> **If you were already hitting the 429 from issue #146 before upgrading**, simply upgrading to v0.5.65 fixes the runaway probe loop โ no manual step required. The provider will be auto-paused on its next probe.
|
|
25
|
+
>
|
|
26
|
+
> **If you want zero authenticated probes for a particular provider** (e.g. you only use OpenRouter for ad-hoc calls and want FCM to leave the key alone), leave the key empty in Settings โ anonymous probes still tell us whether each model is up and will not burn your personal daily quota.
|
|
27
|
+
|
|
28
|
+
### Files
|
|
29
|
+
|
|
30
|
+
- **New**: `src/core/provider-cooldown.js` (โ220 lines).
|
|
31
|
+
- **Modified**: `src/core/ping.js` (import + 429 trigger), `src/core/probe-cache.js` (constants, `brokenCooldownMs`, `isCacheFresh`, `getModelsDueForProbe`, `recordProbeResults`), `src/tui/app.js` (import + `runPingCycle` skip + paused snapshot), `src/tui/tui-state.js` (`pausedProviders: {}` on state), `src/tui/render-table.js` (new chip + thread option through), `README.md` (warning section), `test/test.js` (18 new tests), `test/probe-cache.test.js` (3 updated), `package.json` (`0.5.64` โ `0.5.65`), this changelog.
|
|
32
|
+
|
|
33
|
+
### Breaking changes
|
|
34
|
+
|
|
35
|
+
None. All changes are additive and behaviour-preserving for users who were not already rate-limited.
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "free-coding-models",
|
|
3
|
-
"version": "0.5.
|
|
3
|
+
"version": "0.5.65",
|
|
4
4
|
"description": "Find the fastest coding LLM models in seconds \u2014 ping free models from multiple providers, pick the best one for OpenCode, Cursor, or any AI coding assistant.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"nvidia",
|
package/src/core/ping.js
CHANGED
|
@@ -13,6 +13,9 @@
|
|
|
13
13
|
* - Quota extraction from rate limit headers (multiple variants supported)
|
|
14
14
|
* - Cached provider quota polling with TTL and error backoff
|
|
15
15
|
* - Cloudflare account ID resolution from environment
|
|
16
|
+
* - Per-provider circuit-breaker on 429 (issue #146): pauses ALL models of a provider
|
|
17
|
+
* when the provider returns a quota-exhausted response, so the ping loop stops
|
|
18
|
+
* hammering the user's daily quota while the provider's retry window is active.
|
|
16
19
|
*
|
|
17
20
|
* โ Functions:
|
|
18
21
|
* - `resolveCloudflareUrl`: Resolve {account_id} placeholder from CLOUDFLARE_ACCOUNT_ID env var
|
|
@@ -31,6 +34,7 @@
|
|
|
31
34
|
* - ../src/constants.js: PING_TIMEOUT
|
|
32
35
|
* - ../src/provider-quota-fetchers.js: _fetchProviderQuotaFromModule (quota fetching with cache)
|
|
33
36
|
* - ../src/quota-capabilities.js: supportsUsagePercent
|
|
37
|
+
* - ./provider-cooldown.js: pauseProviderQuota, extractRetryAfterFromResponse (issue #146)
|
|
34
38
|
*
|
|
35
39
|
* โ๏ธ Configuration:
|
|
36
40
|
* - PING_TIMEOUT: Timeout in ms for ping requests (default: 15000)
|
|
@@ -43,6 +47,10 @@
|
|
|
43
47
|
import { PING_TIMEOUT } from './constants.js'
|
|
44
48
|
import { fetchProviderQuota as _fetchProviderQuotaFromModule, extractQuota as _extractQuotaFromModule, processResponseHeaders as _processResponseHeadersFromModule } from './provider-quota-fetchers.js'
|
|
45
49
|
import { supportsUsagePercent } from './quota-capabilities.js'
|
|
50
|
+
import {
|
|
51
|
+
extractRetryAfterFromResponse,
|
|
52
|
+
pauseProviderQuota,
|
|
53
|
+
} from './provider-cooldown.js'
|
|
46
54
|
|
|
47
55
|
const DISABLED_THINKING_RETRY_STATUSES = new Set([400, 422])
|
|
48
56
|
const disabledThinkingUnsupportedProviders = new Set()
|
|
@@ -171,6 +179,14 @@ export async function ping(apiKey, modelId, providerKey, url) {
|
|
|
171
179
|
req = buildPingRequest(apiKey, modelId, providerKey, url, { disableThinking: false })
|
|
172
180
|
resp = await sendPingFetch(req, ctrl.signal)
|
|
173
181
|
}
|
|
182
|
+
// ๐ Provider circuit-breaker (issue #146): openrouter + similar gateways rate-limit
|
|
183
|
+
// ๐ at the provider/key level, not per-model. A 429 means the WHOLE key is paused,
|
|
184
|
+
// ๐ often for hours โ re-pinging every model every 2-30s would burn the quota in minutes.
|
|
185
|
+
// ๐ Honour Retry-After header + the OpenRouter body "try again N seconds later".
|
|
186
|
+
if (resp.status === 429) {
|
|
187
|
+
const retryMs = await extractRetryAfterFromResponse(resp)
|
|
188
|
+
if (retryMs > 0) pauseProviderQuota(providerKey, retryMs)
|
|
189
|
+
}
|
|
174
190
|
// ๐ Normalize all HTTP 2xx statuses to "200" so existing verdict/avg logic still works.
|
|
175
191
|
const code = resp.status >= 200 && resp.status < 300 ? '200' : String(resp.status)
|
|
176
192
|
// ๐ Passive quota tracker (t2): parse the response headers via the shared module.
|
package/src/core/probe-cache.js
CHANGED
|
@@ -12,7 +12,9 @@
|
|
|
12
12
|
* ๐ Freshness rules (in order):
|
|
13
13
|
* ๐ 1. No entry โ due for probe
|
|
14
14
|
* ๐ 2. probeVersion mismatch โ due for probe (silently re-probed, entry overwritten)
|
|
15
|
-
* ๐ 3. status === 'broken' โ
|
|
15
|
+
* ๐ 3. status === 'broken' โ only due AFTER `brokenCooldownMs(consecutiveFailures)`
|
|
16
|
+
* ๐ โ
(issue #146: was always-due โ caused quota burn on
|
|
17
|
+
* ๐ rate-limited providers like openrouter)
|
|
16
18
|
* ๐ 4. now - lastProbedAt >= ttlMs โ due for probe
|
|
17
19
|
* ๐ 5. otherwise โ fresh, skip
|
|
18
20
|
*
|
|
@@ -34,11 +36,13 @@
|
|
|
34
36
|
* โ recordProbeResults(providerKey, results, opts?) โ mutates in-memory cache
|
|
35
37
|
* โ getCacheStats(opts?) โ { total, ok, broken, freshCount, staleCount, ... }
|
|
36
38
|
* โ getCachedResultsForProvider(providerKey, opts?) โ array of synthesized results
|
|
39
|
+
* โ brokenCooldownMs(failureCount) โ Backoff ladder for broken models (issue #146)
|
|
37
40
|
*
|
|
38
41
|
* @exports getProbeCachePath, loadCache, flushCache, clearCache,
|
|
39
42
|
* getModelsDueForProbe, isCacheFresh, recordProbeResults,
|
|
40
43
|
* getCacheStats, getCachedResultsForProvider,
|
|
41
|
-
* DEFAULT_PROBE_TTL_MS,
|
|
44
|
+
* DEFAULT_PROBE_TTL_MS, BROKEN_COOLDOWN_STEPS_MS,
|
|
45
|
+
* brokenCooldownMs, CURRENT_PROBE_VERSION
|
|
42
46
|
*
|
|
43
47
|
* @see src/core/ping-loop.js โ the integration point (skips fresh entries)
|
|
44
48
|
* @see src/core/cache.js โ older per-session ping cache (5 min TTL, distinct concern)
|
|
@@ -55,6 +59,30 @@ import { atomicWriteJson } from './shared-helpers.js'
|
|
|
55
59
|
/** ๐ Default probe-result freshness window โ 24h. Override via opts.ttlMs. */
|
|
56
60
|
export const DEFAULT_PROBE_TTL_MS = 24 * 60 * 60 * 1000
|
|
57
61
|
|
|
62
|
+
/**
|
|
63
|
+
* ๐ Cooldown ladder for broken models (issue #146). After a consecutive failure,
|
|
64
|
+
* ๐ the model is treated as fresh for the duration returned by `brokenCooldownMs()`,
|
|
65
|
+
* ๐ so the ping loop skips it instead of hammering it every 2โ30s.
|
|
66
|
+
* ๐ The ladder grows exponentially then plateaus: 30s โ 1m โ 2m โ 5m.
|
|
67
|
+
* ๐ Reaching the plateau (5min) caps the probe pressure on a permanently-broken
|
|
68
|
+
* ๐ model at ~0.2 req/min instead of the previous ~30 req/min.
|
|
69
|
+
*/
|
|
70
|
+
export const BROKEN_COOLDOWN_STEPS_MS = [30_000, 60_000, 120_000, 300_000]
|
|
71
|
+
|
|
72
|
+
/**
|
|
73
|
+
* ๐ brokenCooldownMs: Cooldown duration for a model that has failed N consecutive times.
|
|
74
|
+
* ๐ Failure count is 1-indexed (1 = first failure, 2 = second, ...).
|
|
75
|
+
* ๐ Saturates at the last entry of BROKEN_COOLDOWN_STEPS_MS.
|
|
76
|
+
*
|
|
77
|
+
* @param {number} failureCount
|
|
78
|
+
* @returns {number} milliseconds (always >= 0)
|
|
79
|
+
*/
|
|
80
|
+
export function brokenCooldownMs(failureCount) {
|
|
81
|
+
const n = Math.max(1, Math.floor(Number(failureCount) || 1))
|
|
82
|
+
const idx = Math.min(n - 1, BROKEN_COOLDOWN_STEPS_MS.length - 1)
|
|
83
|
+
return BROKEN_COOLDOWN_STEPS_MS[idx]
|
|
84
|
+
}
|
|
85
|
+
|
|
58
86
|
/**
|
|
59
87
|
* ๐ Bump this number whenever ping behaviour changes (new endpoint, different prompt,
|
|
60
88
|
* ๐ new provider factory from t7, etc.). On load, any entry whose probeVersion differs
|
|
@@ -270,8 +298,16 @@ export function getModelsDueForProbe(providerKey, modelIds, opts = {}) {
|
|
|
270
298
|
if (typeof entry.probeVersion !== 'number' || entry.probeVersion !== probeVersion) {
|
|
271
299
|
due.push(id); continue
|
|
272
300
|
}
|
|
273
|
-
// Rule 3: broken โ
|
|
274
|
-
|
|
301
|
+
// Rule 3: broken โ due only AFTER brokenCooldownMs elapsed (issue #146)
|
|
302
|
+
// ๐ Previous behaviour re-pinged broken models every cycle, which burned
|
|
303
|
+
// ๐ rate-limited providers' quota (openrouter: ~1000 req/day cap).
|
|
304
|
+
// ๐ Now we honour an exponential backoff: 30s โ 1m โ 2m โ 5m (plateau).
|
|
305
|
+
if (entry.status === 'broken') {
|
|
306
|
+
const cooldown = brokenCooldownMs(entry.consecutiveFailures ?? 1)
|
|
307
|
+
if (now - entry.lastProbedAt < cooldown) continue
|
|
308
|
+
due.push(id)
|
|
309
|
+
continue
|
|
310
|
+
}
|
|
275
311
|
// Rule 4: TTL expired โ due
|
|
276
312
|
if (typeof entry.lastProbedAt !== 'number' || now - entry.lastProbedAt >= ttlMs) {
|
|
277
313
|
due.push(id); continue
|
|
@@ -301,7 +337,13 @@ export function isCacheFresh(providerKey, modelId, opts = {}) {
|
|
|
301
337
|
const entry = cache?.providers?.[providerKey]?.models?.[modelId]
|
|
302
338
|
if (!entry) return false
|
|
303
339
|
if (typeof entry.probeVersion !== 'number' || entry.probeVersion !== probeVersion) return false
|
|
304
|
-
if (entry.status === 'broken')
|
|
340
|
+
if (entry.status === 'broken') {
|
|
341
|
+
// ๐ Issue #146 โ broken models are now FRESH for `brokenCooldownMs` after their
|
|
342
|
+
// ๐ last failure, instead of always being due. This caps the probe pressure
|
|
343
|
+
// ๐ on permanently-broken models at ~0.2 req/min instead of ~30 req/min.
|
|
344
|
+
const cooldown = brokenCooldownMs(entry.consecutiveFailures ?? 1)
|
|
345
|
+
return now - entry.lastProbedAt < cooldown
|
|
346
|
+
}
|
|
305
347
|
if (typeof entry.lastProbedAt !== 'number') return false
|
|
306
348
|
return now - entry.lastProbedAt < ttlMs
|
|
307
349
|
}
|
|
@@ -350,12 +392,19 @@ export function recordProbeResults(providerKey, results, opts = {}) {
|
|
|
350
392
|
for (const raw of results || []) {
|
|
351
393
|
try {
|
|
352
394
|
const r = normaliseResult(raw)
|
|
395
|
+
// ๐ Consecutive failure tracker (issue #146): drives the broken-cooldown ladder.
|
|
396
|
+
// ๐ `ok` resets to 0; `broken` increments from the previous value (or 1 if absent).
|
|
397
|
+
const prev = bucket[r.modelId]
|
|
398
|
+
const consecutiveFailures = r.status === 'broken'
|
|
399
|
+
? ((prev && Number.isInteger(prev.consecutiveFailures)) ? prev.consecutiveFailures + 1 : 1)
|
|
400
|
+
: 0
|
|
353
401
|
bucket[r.modelId] = {
|
|
354
402
|
status: r.status,
|
|
355
403
|
lastProbedAt: now,
|
|
356
404
|
latencyMs: r.latencyMs,
|
|
357
405
|
lastError: r.lastError,
|
|
358
406
|
probeVersion: CURRENT_PROBE_VERSION,
|
|
407
|
+
consecutiveFailures,
|
|
359
408
|
}
|
|
360
409
|
written++
|
|
361
410
|
} catch {
|
|
@@ -0,0 +1,185 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* @file provider-cooldown.js
|
|
3
|
+
* @description Per-provider quota circuit-breaker for the health-ping loop.
|
|
4
|
+
*
|
|
5
|
+
* @details
|
|
6
|
+
* ๐ Why this exists:
|
|
7
|
+
* ๐ - OpenRouter (and similar rate-limited gateways) enforce rate limits at the
|
|
8
|
+
* ๐ **provider/key** level, not per-model. A single 429 "Rate limit exceeded"
|
|
9
|
+
* ๐ response means **the whole key is paused**, often for hours.
|
|
10
|
+
* ๐ - Without this breaker, FCM keeps re-pinging every model of that provider
|
|
11
|
+
* ๐ every 2โ30 seconds (because probe-cache treats `broken` as always-due),
|
|
12
|
+
* ๐ burning the user's daily quota in minutes (issue #146).
|
|
13
|
+
* ๐ - This module exposes a tiny in-memory map of paused providers so the ping
|
|
14
|
+
* ๐ loop can short-circuit them until the Retry-After window expires.
|
|
15
|
+
*
|
|
16
|
+
* ๐ Design notes:
|
|
17
|
+
* ๐ - In-memory only, intentionally. The pause is short-lived (minutes to hours)
|
|
18
|
+
* ๐ and we don't want a stale pause to outlive a server-side reset.
|
|
19
|
+
* ๐ - `pauseProviderQuota` keeps the **maximum** pause if the provider is already
|
|
20
|
+
* ๐ paused, so a rapid succession of 429s with decreasing Retry-After values
|
|
21
|
+
* ๐ never accidentally **shortens** an existing longer pause.
|
|
22
|
+
*
|
|
23
|
+
* @functions
|
|
24
|
+
* โ isProviderQuotaPaused(providerKey, now?) โ Is this provider currently paused?
|
|
25
|
+
* โ pauseProviderQuota(providerKey, ms) โ Set a pause for ms milliseconds
|
|
26
|
+
* โ providerQuotaPauseRemaining(providerKey, now?) โ ms left in the current pause
|
|
27
|
+
* โ listPausedProviders(now?) โ Diagnostic: {[providerKey]: remainingMs}
|
|
28
|
+
* โ clearProviderQuotaPause(providerKey) โ Force-reset (testing/escape hatch)
|
|
29
|
+
* โ parseRetryAfterMs(value) โ Header-only Retry-After parser
|
|
30
|
+
* โ extractRetryAfterFromResponse(resp) โ Header + body parser (OR style message)
|
|
31
|
+
*
|
|
32
|
+
* @exports isProviderQuotaPaused, pauseProviderQuota, providerQuotaPauseRemaining,
|
|
33
|
+
* listPausedProviders, clearProviderQuotaPause, parseRetryAfterMs,
|
|
34
|
+
* extractRetryAfterFromResponse
|
|
35
|
+
*/
|
|
36
|
+
|
|
37
|
+
// โโโ Module state โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
|
|
38
|
+
|
|
39
|
+
/**
|
|
40
|
+
* ๐ Map<providerKey, timestamp-ms> until which the provider's quota is considered
|
|
41
|
+
* ๐ paused. `now >= value` means the pause has expired and the entry is purged lazily.
|
|
42
|
+
*/
|
|
43
|
+
const providerQuotaPausedUntil = new Map()
|
|
44
|
+
|
|
45
|
+
// โโโ Pure helpers (exported, easy to unit test) โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
|
|
46
|
+
|
|
47
|
+
/**
|
|
48
|
+
* ๐ parseRetryAfterMs: Parse a Retry-After header value into milliseconds.
|
|
49
|
+
* ๐ Accepts either a plain-seconds integer ("120") or an HTTP-date ("Wed, 21 Oct 2026 07:28:00 GMT").
|
|
50
|
+
* ๐ Returns null if the value is missing, malformed, or already in the past.
|
|
51
|
+
*
|
|
52
|
+
* @param {string|number|null|undefined} value
|
|
53
|
+
* @returns {number|null} milliseconds, or null when not parseable
|
|
54
|
+
*/
|
|
55
|
+
export function parseRetryAfterMs(value) {
|
|
56
|
+
if (value == null) return null
|
|
57
|
+
if (typeof value === 'string' && value.trim() === '') return null
|
|
58
|
+
const seconds = Number(value)
|
|
59
|
+
if (Number.isFinite(seconds) && seconds >= 0) return Math.round(seconds * 1000)
|
|
60
|
+
const dateMs = Date.parse(value)
|
|
61
|
+
if (Number.isFinite(dateMs)) {
|
|
62
|
+
const delta = dateMs - Date.now()
|
|
63
|
+
return delta > 0 ? delta : null
|
|
64
|
+
}
|
|
65
|
+
return null
|
|
66
|
+
}
|
|
67
|
+
|
|
68
|
+
/**
|
|
69
|
+
* ๐ extractRetryAfterFromResponse: Best-effort Retry-After extraction.
|
|
70
|
+
* ๐ Tries the header first (RFC 7231), then falls back to the OpenRouter-style
|
|
71
|
+
* ๐ error body "Rate limit exceeded, please try again N seconds later.".
|
|
72
|
+
* ๐ Returns the delay in milliseconds, or 0 if no signal is found.
|
|
73
|
+
*
|
|
74
|
+
* @param {Response} resp โ A Fetch Response object (uses .headers + .clone()/.text()).
|
|
75
|
+
* @returns {Promise<number>}
|
|
76
|
+
*/
|
|
77
|
+
export async function extractRetryAfterFromResponse(resp) {
|
|
78
|
+
if (!resp || typeof resp.headers?.get !== 'function') return 0
|
|
79
|
+
|
|
80
|
+
// 1) Standard Retry-After header (lowercase lookup matches Headers semantics)
|
|
81
|
+
const headerMs = parseRetryAfterMs(resp.headers.get('retry-after'))
|
|
82
|
+
if (headerMs && headerMs > 0) return headerMs
|
|
83
|
+
|
|
84
|
+
// 2) OpenRouter-style message body, e.g.: "Rate limit exceeded, please try again 50680 seconds later."
|
|
85
|
+
try {
|
|
86
|
+
const text = await resp.clone().text()
|
|
87
|
+
if (!text) return 0
|
|
88
|
+
const m = text.match(/try again (\d+)\s*seconds? later/i)
|
|
89
|
+
if (m) {
|
|
90
|
+
const secs = parseInt(m[1], 10)
|
|
91
|
+
if (Number.isFinite(secs) && secs > 0) return secs * 1000
|
|
92
|
+
}
|
|
93
|
+
} catch {
|
|
94
|
+
// Body unreadable โ treat as no signal, don't throw.
|
|
95
|
+
}
|
|
96
|
+
return 0
|
|
97
|
+
}
|
|
98
|
+
|
|
99
|
+
// โโโ Pause map API โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
|
|
100
|
+
|
|
101
|
+
/**
|
|
102
|
+
* ๐ isProviderQuotaPaused: Return true while the provider's quota pause is active.
|
|
103
|
+
* ๐ Lazily purges expired pauses so the map stays bounded over long sessions.
|
|
104
|
+
*
|
|
105
|
+
* @param {string} providerKey
|
|
106
|
+
* @param {number} [now=Date.now()]
|
|
107
|
+
* @returns {boolean}
|
|
108
|
+
*/
|
|
109
|
+
export function isProviderQuotaPaused(providerKey, now = Date.now()) {
|
|
110
|
+
const until = providerQuotaPausedUntil.get(providerKey)
|
|
111
|
+
if (until == null) return false
|
|
112
|
+
if (now >= until) {
|
|
113
|
+
providerQuotaPausedUntil.delete(providerKey)
|
|
114
|
+
return false
|
|
115
|
+
}
|
|
116
|
+
return true
|
|
117
|
+
}
|
|
118
|
+
|
|
119
|
+
/**
|
|
120
|
+
* ๐ pauseProviderQuota: Mark a provider as paused for `ms` milliseconds.
|
|
121
|
+
* ๐ If already paused, keeps the longer of the existing and new windows (never shortens).
|
|
122
|
+
*
|
|
123
|
+
* @param {string} providerKey
|
|
124
|
+
* @param {number} ms โ Duration in milliseconds. Non-positive values are ignored.
|
|
125
|
+
* @returns {boolean} true if the pause was applied or extended, false otherwise.
|
|
126
|
+
*/
|
|
127
|
+
export function pauseProviderQuota(providerKey, ms, opts = {}) {
|
|
128
|
+
if (!providerKey || typeof ms !== 'number' || !Number.isFinite(ms) || ms <= 0) return false
|
|
129
|
+
const now = opts.now ?? Date.now()
|
|
130
|
+
const until = now + ms
|
|
131
|
+
const cur = providerQuotaPausedUntil.get(providerKey) ?? 0
|
|
132
|
+
// ๐ Keep the max so a 5s pause arriving after a 14h pause never shortens it.
|
|
133
|
+
if (until > cur) providerQuotaPausedUntil.set(providerKey, until)
|
|
134
|
+
return true
|
|
135
|
+
}
|
|
136
|
+
|
|
137
|
+
/**
|
|
138
|
+
* ๐ providerQuotaPauseRemaining: How many ms remain on the current pause.
|
|
139
|
+
* ๐ Returns 0 when the pause has expired or the provider is not paused.
|
|
140
|
+
*
|
|
141
|
+
* @param {string} providerKey
|
|
142
|
+
* @param {number} [now=Date.now()]
|
|
143
|
+
* @returns {number} milliseconds remaining (always >= 0)
|
|
144
|
+
*/
|
|
145
|
+
export function providerQuotaPauseRemaining(providerKey, now = Date.now()) {
|
|
146
|
+
const until = providerQuotaPausedUntil.get(providerKey)
|
|
147
|
+
if (!until) return 0
|
|
148
|
+
const remaining = until - now
|
|
149
|
+
if (remaining <= 0) {
|
|
150
|
+
providerQuotaPausedUntil.delete(providerKey)
|
|
151
|
+
return 0
|
|
152
|
+
}
|
|
153
|
+
return remaining
|
|
154
|
+
}
|
|
155
|
+
|
|
156
|
+
/**
|
|
157
|
+
* ๐ listPausedProviders: Snapshot of currently paused providers + ms remaining.
|
|
158
|
+
* ๐ Intended for diagnostics (TUI footer chip, daemon /stats, error logs).
|
|
159
|
+
*
|
|
160
|
+
* @param {number} [now=Date.now()]
|
|
161
|
+
* @returns {Record<string, number>} {[providerKey]: msRemaining}
|
|
162
|
+
*/
|
|
163
|
+
export function listPausedProviders(now = Date.now()) {
|
|
164
|
+
const out = {}
|
|
165
|
+
for (const [providerKey, until] of providerQuotaPausedUntil) {
|
|
166
|
+
const remaining = until - now
|
|
167
|
+
if (remaining > 0) out[providerKey] = remaining
|
|
168
|
+
}
|
|
169
|
+
return out
|
|
170
|
+
}
|
|
171
|
+
|
|
172
|
+
/**
|
|
173
|
+
* ๐ clearProviderQuotaPause: Force-reset the pause for a provider (or all).
|
|
174
|
+
* ๐ Mainly a test/escape hatch โ production code should let pauses expire naturally.
|
|
175
|
+
*
|
|
176
|
+
* @param {string} [providerKey] โ omit to clear all pauses
|
|
177
|
+
* @returns {void}
|
|
178
|
+
*/
|
|
179
|
+
export function clearProviderQuotaPause(providerKey) {
|
|
180
|
+
if (!providerKey) {
|
|
181
|
+
providerQuotaPausedUntil.clear()
|
|
182
|
+
return
|
|
183
|
+
}
|
|
184
|
+
providerQuotaPausedUntil.delete(providerKey)
|
|
185
|
+
}
|
package/src/tui/app.js
CHANGED
|
@@ -139,6 +139,10 @@ import {
|
|
|
139
139
|
pruneStaleEntries as pruneProbeCacheStaleEntries,
|
|
140
140
|
DEFAULT_PROBE_TTL_MS,
|
|
141
141
|
} from '../core/probe-cache.js'
|
|
142
|
+
import {
|
|
143
|
+
isProviderQuotaPaused,
|
|
144
|
+
listPausedProviders,
|
|
145
|
+
} from '../core/provider-cooldown.js'
|
|
142
146
|
import { checkConfigSecurity } from '../core/security.js'
|
|
143
147
|
import { buildCliHelpText } from './cli-help.js'
|
|
144
148
|
import { detectActiveTheme, THEME_BG_RGB, getTheme, patchThemeBg } from './theme.js'
|
|
@@ -1046,6 +1050,8 @@ export async function runApp(cliArgs, config, startupOptions = {}) {
|
|
|
1046
1050
|
probeCacheMisses: state.probeCacheMisses || 0,
|
|
1047
1051
|
probeCacheBrokenHidden: state.probeCacheBrokenHidden || 0,
|
|
1048
1052
|
showBrokenMode: !!state.showBrokenMode,
|
|
1053
|
+
// ๐ Provider circuit-breaker (issue #146): footer chip for paused providers
|
|
1054
|
+
pausedProviders: state.pausedProviders || {},
|
|
1049
1055
|
// ๐ Enrichment chips (t4 + t5): benchmark catalog + models.dev provenance
|
|
1050
1056
|
// ๐ Read fresh from moduleEnrichmentStats on every render so the async
|
|
1051
1057
|
// ๐ models.dev fetch (background, never blocks the TUI) is reflected as
|
|
@@ -1229,6 +1235,13 @@ export async function runApp(cliArgs, config, startupOptions = {}) {
|
|
|
1229
1235
|
if (!state.config.favorites.includes(favKey)) return
|
|
1230
1236
|
}
|
|
1231
1237
|
|
|
1238
|
+
// ๐ Per-provider circuit-breaker (issue #146): if the provider's quota is
|
|
1239
|
+
// ๐ paused (e.g. openrouter returned 429 "try again 50680s later"), skip
|
|
1240
|
+
// ๐ every model of that provider until the Retry-After window expires.
|
|
1241
|
+
// ๐ Without this, broken/overloaded models are re-pinged every cycle,
|
|
1242
|
+
// ๐ burning the user's daily quota in minutes.
|
|
1243
|
+
if (isProviderQuotaPaused(r.providerKey)) return
|
|
1244
|
+
|
|
1232
1245
|
// ๐ Probe-cache (t1): skip models that are still fresh + ok in the cache.
|
|
1233
1246
|
// ๐ Broken models always pass through isCacheFresh() (it returns false for them),
|
|
1234
1247
|
// ๐ which gives us automatic recovery detection for free.
|
|
@@ -1241,6 +1254,11 @@ export async function runApp(cliArgs, config, startupOptions = {}) {
|
|
|
1241
1254
|
})
|
|
1242
1255
|
})
|
|
1243
1256
|
|
|
1257
|
+
// ๐ Take a snapshot of currently paused providers so the TUI footer can
|
|
1258
|
+
// ๐ surface a "โธ openrouter (14h)" chip while we wait out a Retry-After.
|
|
1259
|
+
// ๐ Updates every cycle so the remaining-time ticks down live.
|
|
1260
|
+
state.pausedProviders = listPausedProviders()
|
|
1261
|
+
|
|
1244
1262
|
refreshAutoPingMode()
|
|
1245
1263
|
scheduleNextPing()
|
|
1246
1264
|
} catch (err) {
|
package/src/tui/render-table.js
CHANGED
|
@@ -216,6 +216,7 @@ export function renderTable({
|
|
|
216
216
|
probeCacheMisses = 0,
|
|
217
217
|
probeCacheBrokenHidden = 0,
|
|
218
218
|
showBrokenMode = false,
|
|
219
|
+
pausedProviders = {}, // ๐ issue #146: { openrouter: 50680000 } while a provider is paused
|
|
219
220
|
enrichmentBenchCount = 0,
|
|
220
221
|
enrichmentBenchLastUpdated = null,
|
|
221
222
|
metaSourceCounts = null,
|
|
@@ -1235,6 +1236,36 @@ export function renderTable({
|
|
|
1235
1236
|
probeCacheLabel = chalk.bgRgb(40, 90, 140).rgb(220, 235, 255).bold(` ${cachedTxt}${brokenTxt} `)
|
|
1236
1237
|
}
|
|
1237
1238
|
|
|
1239
|
+
// ๐ Provider circuit-breaker chip (issue #146): surfaces any provider that the
|
|
1240
|
+
// ๐ health-ping loop has paused because it returned a 429 quota response.
|
|
1241
|
+
// ๐ Format: `โธ openrouter 14h ยท โธ mistral 5m`. The live duration auto-rolls
|
|
1242
|
+
// ๐ to s/m/h/d so users see a real countdown and don't get surprised later
|
|
1243
|
+
// ๐ when their requests stop working ("wait, why is everything 429-ing?").
|
|
1244
|
+
let pausedProvidersLabel = ''
|
|
1245
|
+
if (pausedProviders && typeof pausedProviders === 'object') {
|
|
1246
|
+
const fmtRemaining = (ms) => {
|
|
1247
|
+
if (!Number.isFinite(ms) || ms <= 0) return '0s'
|
|
1248
|
+
const totalSec = Math.round(ms / 1000)
|
|
1249
|
+
const day = Math.floor(totalSec / 86400)
|
|
1250
|
+
const hr = Math.floor((totalSec % 86400) / 3600)
|
|
1251
|
+
const min = Math.floor((totalSec % 3600) / 60)
|
|
1252
|
+
const sec = totalSec % 60
|
|
1253
|
+
if (day > 0) return `${day}d${hr}h`
|
|
1254
|
+
if (hr > 0) return `${hr}h${min}m`
|
|
1255
|
+
if (min > 0) return `${min}m${sec}s`
|
|
1256
|
+
return `${sec}s`
|
|
1257
|
+
}
|
|
1258
|
+
const entries = Object.entries(pausedProviders)
|
|
1259
|
+
.filter(([, ms]) => Number.isFinite(ms) && ms > 0)
|
|
1260
|
+
.sort(([a], [b]) => a.localeCompare(b))
|
|
1261
|
+
if (entries.length > 0) {
|
|
1262
|
+
const txt = entries.map(([k, ms]) => `โธ ${k} ${fmtRemaining(ms)}`).join(' \u00b7 ')
|
|
1263
|
+
// ๐ Warm amber background โ distinct from the blue probe-cache chip so the
|
|
1264
|
+
// ๐ user can spot it at a glance and understand "these providers are resting".
|
|
1265
|
+
pausedProvidersLabel = chalk.bgRgb(140, 80, 30).rgb(255, 230, 200).bold(` ${txt} `)
|
|
1266
|
+
}
|
|
1267
|
+
}
|
|
1268
|
+
|
|
1238
1269
|
// ๐ Enrichment chips (t4 + t5):
|
|
1239
1270
|
// ๐ `๐ bench N (date) ยท ๐ก live K / ๐ฆ cached M`
|
|
1240
1271
|
// ๐ Always shown when any enrichment state exists.
|
|
@@ -1277,8 +1308,8 @@ export function renderTable({
|
|
|
1277
1308
|
}
|
|
1278
1309
|
}
|
|
1279
1310
|
|
|
1280
|
-
// ๐ Line 3: Speed Test + Global Benchmark + Probe + Probe-cache + Enrichment + Quota + Last release
|
|
1281
|
-
if (releaseLabel || speedTestLabel || globalBenchmarkLabel || probeLabel || probeCacheLabel || enrichmentLabel || quotaLabel) {
|
|
1311
|
+
// ๐ Line 3: Speed Test + Global Benchmark + Probe + Probe-cache + Paused-providers + Enrichment + Quota + Last release
|
|
1312
|
+
if (releaseLabel || speedTestLabel || globalBenchmarkLabel || probeLabel || probeCacheLabel || pausedProvidersLabel || enrichmentLabel || quotaLabel) {
|
|
1282
1313
|
const parts = [
|
|
1283
1314
|
{ text: ' ', key: null },
|
|
1284
1315
|
{ text: speedTestLabel, key: 'a' },
|
|
@@ -1288,6 +1319,8 @@ export function renderTable({
|
|
|
1288
1319
|
{ text: probeLabel, key: null },
|
|
1289
1320
|
{ text: probeCacheLabel ? ' ' : '', key: null },
|
|
1290
1321
|
{ text: probeCacheLabel, key: null },
|
|
1322
|
+
{ text: pausedProvidersLabel ? ' ' : '', key: null },
|
|
1323
|
+
{ text: pausedProvidersLabel, key: null },
|
|
1291
1324
|
{ text: enrichmentLabel ? ' ' : '', key: null },
|
|
1292
1325
|
{ text: enrichmentLabel, key: null },
|
|
1293
1326
|
{ text: quotaLabel ? ' ' : '', key: null },
|
package/src/tui/tui-state.js
CHANGED
|
@@ -297,6 +297,13 @@ export function createTuiState({
|
|
|
297
297
|
probeCacheTtlMs: null, // ๐ Set in app.js from cliArgs.probeTtlMs || DEFAULT
|
|
298
298
|
showBrokenMode: false,
|
|
299
299
|
|
|
300
|
+
// ๐ Per-provider quota circuit-breaker (issue #146): snapshot of providers
|
|
301
|
+
// ๐ currently paused because they returned 429 during a health-probe.
|
|
302
|
+
// ๐ Format: { [providerKey]: remainingMs } โ refreshed each ping cycle
|
|
303
|
+
// ๐ so the TUI footer can show a live "โธ openrouter (14h)" chip.
|
|
304
|
+
// ๐ Empty when no provider is in pause.
|
|
305
|
+
pausedProviders: {},
|
|
306
|
+
|
|
300
307
|
// ๐ Runtime telemetry overlay state (t3): Shift+W opens the Runtime Report
|
|
301
308
|
// ๐ overlay. Data is loaded fresh from ~/.free-coding-models/runtime-telemetry.json
|
|
302
309
|
// ๐ when the overlay opens โ no polling, no daemon dependency.
|