free-coding-models 0.5.69 → 0.5.71

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -484,7 +484,7 @@ Then restart Pi. The extension loads automatically. Requires Pi + `free-coding-m
484
484
  | `/fcm-router` | Connect Pi to the local FCM Smart Router daemon |
485
485
  | `/fcm-status` | Diagnostics: active model, last scan source, daemon state |
486
486
 
487
- **Composite ranking** — SWE-bench (60%) + Latency (20%) + TPS (10%) + Stability (10%). Tiny-context Cerebras models (~8k total tokens) are hidden from Pi/OpenCode pickers since they pass a `hi` probe but fail real agent prompts.
487
+ **Composite ranking** — SWE-bench (60%) + Latency (20%) + TPS (10%) + Stability (10%). Cerebras free-tier models are capped at ~64-65k total tokens (paid tier gets 131k) and are hidden from Pi/OpenCode pickers when they fail real agent prompts.
488
488
 
489
489
  → Full architecture: [`packages/fcm-pi/README.md`](./packages/fcm-pi/README.md)
490
490
 
@@ -0,0 +1,49 @@
1
+ # Changelog v0.5.70 - 2026-08-13
2
+
3
+ ### 🔍 Full provider audit (20 providers, 222 → 196 models)
4
+
5
+ Major refresh across every free-coding-model provider. Most impactful change: **GitHub Models is fully retired** as of 2026-07-30 (HTTP 410 Gone) — the entire provider entry is removed from the active catalog.
6
+
7
+ ### Added
8
+ - **nvidia (5):** `poolside/laguna-xs-2.1` (S+, 70.9%), `meta/llama-3.3-70b-instruct`, `mistralai/mistral-large-2-instruct`, `meta/codellama-70b`, `ibm/granite-34b-code-instruct`
9
+ - **qwen (1):** `qwen3-vl-flash` — qwen3.5/3.6/3.7/3.7-plus/3.6-max-preview/3.5-plus/3.5-flash/3.6-flash/3.7-max/3.6-plus/3.7-plus and `qwen3-coder-next` were already added in the prior sweep
10
+ - **ollama-cloud (0 net):** `kimi-k3`, `glm-5.1`, `glm-5.2`, `minimax-m2.7`, `deepseek-v4-pro`, `mistral-large-3:675b-cloud`, `qwen3.5` already present (no `:cloud` suffix convention used)
11
+ - **cloudflare (2):** `@cf/google/gemma-3-12b-it`, `@cf/moonshotai/kimi-k2.5`
12
+ - **scaleway (1):** `deepseek-v4-flash-0731` (256k serverless)
13
+ - **googleai (1):** `gemini-3.7-flash` (Aug 2026, 1M context)
14
+ - **openrouter (2):** `liquid/lfm-2.5-2.6b:free`, `nvidia/nemotron-3.5-lightning:free`
15
+ - **opencode-zen (2):** `hy3-free` (re-added), `nemotron-3.5-lightning-free`
16
+ - **ovhcloud (1):** `Qwen2.5-VL-72B-Instruct`
17
+ - **zai (7):** `zai/glm-5.2`, `zai/glm-5-turbo`, `zai/glm-5v-turbo`, `zai/glm-4.7`, `zai/glm-4.6`, `zai/glm-4.7-flashx`, `zai/glm-4.6v` (full GLM-5 family)
18
+ - **llm7 (1):** `mistral-Nemo-Instruct-2407`
19
+ - **codestral (2):** `codestral-2501`, `codestral-2405`
20
+ - **kilo (1):** `kilo-auto/small` (262k, routes to gemma-4-26b for free accounts)
21
+
22
+ ### Removed
23
+ - **github-models (-35):** Provider fully retired 2026-07-30 by GitHub (HTTP 410 Gone with `github_models_retirement_brownout` error). All 35 catalog entries unreachable. Provider commented out of `sources` map; `githubModels` array kept empty for backwards-compat.
24
+ - **groq (-2):** `llama-3.3-70b-versatile`, `llama-3.1-8b-instant` — Groq deprecation, shutdown 2026-08-16
25
+ - **ollama-cloud (-11 net):** 11 stale cloud IDs (deepseek-v4 umbrella, glm-5, mistral-medium-3.5, qwen3-coder-480b, step-3.7-flash, etc.) — already cleaned up in prior sweep
26
+ - **ovhcloud (-2):** `Mistral-Small-3.2-24B-Instruct-2506`, `Mistral-Nemo-Instruct-2407` — endpoints reachable but no longer in public catalog
27
+ - **openrouter (-2):** `poolside/laguna-m.1:free` (gone from catalog entirely), `inclusionai/ling-3.0-flash:free` (now paid-only)
28
+ - **opencode-zen (-2):** `north-mini-code-free`, `ling-3.0-flash-free`
29
+ - **scaleway (-2):** `mistral-large-3-675b-instruct-2512`, `gemma-4-31b-it` — Dedicated tier only, not on Serverless
30
+ - **routeway (-4):** `deepseek-v4-flash:free`, `step-3.5-flash:free`, `ling-3.0-flash:free`, `ling-2.6-flash:free` — silently retired from zero-price catalog
31
+ - **mistral (-1):** `devstral-2512` — Mistral deprecation, retirement 2026-07-31
32
+ - **codestral (-1):** `codestral-2` — fabricated ID, never existed in Mistral catalog (Mistral uses date-stamped versioning)
33
+ - **novita (-1):** `tencent/hy3` — `isFree:false` per Novita pricing page. Provider now empty (0 models).
34
+
35
+ ### Fixed
36
+ - **nvidia (2):** `deepseek-ai/deepseek-v4-flash` → `deepseek-ai/deepseek-v4-flash-0731` (NIM `/v1/models` only exposes the -0731 suffix); `deepseek-ai/deepseek-v4-pro` flagged as page-only / partner-routed (NOT in `/v1/models`, served via Fireworks/DeepInfra/Together/OpenRouter)
37
+ - **openrouter (1):** `nvidia/nemotron-3-super-120b-a12b:free` ctx `1M` → `262k`
38
+ - **opencode-zen (1):** `poolside/laguna-s-2.1-free` → `laguna-s-2.1-free` (drop `poolside/` prefix)
39
+ - **scaleway (3):** `glm-5.2` ctx `1M` → `256k`, `devstral-2-123b-instruct-2512` ctx `260k` → `200k`, `llama-3.3-70b-instruct` ctx `128k` → `100k` (Serverless tier per official catalog)
40
+ - **mistral (3):** `ministral-3-14b-25-12` → `ministral-14b-2512`, `ministral-3-8b-25-12` → `ministral-8b-2512`, `ministral-3-3b-25-12` → `ministral-3b-2512` (use real API model IDs from Mistral docs JSON, not page slugs)
41
+ - **llm7 (1):** `gemini-3.1-flash-lite` ctx `1M` → `256k` (real LLM7 ctx limit)
42
+
43
+ ### Internal
44
+ - `PREFERRED_DEFAULT_MODELS` (router pinned set) refreshed to point at current model IDs after `llama-3.3-70b-versatile` / `llama-3.1-8b-instant` removal and `deepseek-v4-flash` → `deepseek-v4-flash-0731` rename
45
+ - `PROVIDER_TEST_MODEL_OVERRIDES.nvidia` updated to `deepseek-ai/deepseek-v4-flash-0731`
46
+ - `buildDefaultRouterSet` now filters to keyed providers in the no-probeFn path (matches sync fallback behavior) so users without a probe fn see only models they can actually call
47
+
48
+ ### Tests
49
+ - 564/564 tests passing (97.5% → 100% in the impacted router/key-tester suites)
@@ -0,0 +1,11 @@
1
+ # Changelog v0.5.71 - 2026-08-15
2
+
3
+ ### Fixed
4
+ - **Goose launch wrote a hardcoded 128k context limit** — when launching a model into Goose with `--goose` (or via the TUI Enter flow), the generated `~/.config/goose/custom_providers/fcm-<provider>.json` always set `context_limit: 128000`, ignoring the model's real context window. Models with larger contexts (e.g. MiniMax M3, 1M) were silently capped at 128k, and free-tier models with smaller limits (e.g. Cerebras 64-65k) were advertised above what the API actually allows. The launcher now uses the same `parseContextWindow(model.ctx)` logic as the Install Endpoints flow, so the context limit always matches the catalog value. (Fixes GitHub issue #153)
5
+ - **Stale Cerebras context references in docs** — README.md, packages/fcm-pi/README.md, and the Pi integration docs claimed Cerebras free tier has an "~8k total token limit". Per the official Cerebras Inference docs the free tier now caps context at ~64-65k (paid tier gets 131k). Docs updated to match.
6
+
7
+ ### Changed
8
+ - `parseContextWindow` is now exported from `src/core/endpoint-installer.js` so the Goose launcher (`src/core/tool-launchers.js`) and the Install Endpoints flow share the same single source of truth for context parsing.
9
+
10
+ ### Tests
11
+ - 797/797 tests passing, including a new regression assertion that a 1M-context model launched into Goose produces `context_limit: 1000000` in the generated provider file.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "free-coding-models",
3
- "version": "0.5.69",
3
+ "version": "0.5.71",
4
4
  "description": "Find the fastest coding LLM models in seconds — ping free models from multiple providers, pick the best one for OpenCode, Cursor, or any AI coding assistant.",
5
5
  "keywords": [
6
6
  "nvidia",
package/sources.js CHANGED
@@ -44,10 +44,11 @@ export const nvidiaNim = [
44
44
  ['z-ai/glm-5.2', 'GLM 5.1', 'S+', '82.8%', '128k'],
45
45
  // Removed (2026-07-27): minimaxai/minimax-m2.7 (MiniMax M2.7) — EOL 2026-07-27 (HTTP 410 Gone)
46
46
  ['moonshotai/kimi-k2.6', 'Kimi K2.6', 'S+', '80.2%', '262k'],
47
- ['deepseek-ai/deepseek-v4-pro', 'DeepSeek V4 Pro', 'S+', '80.6%', '1M'],
48
- ['deepseek-ai/deepseek-v4-flash', 'DeepSeek V4 Flash', 'S+', '79.0%', '1M'],
47
+ ['deepseek-ai/deepseek-v4-pro', 'DeepSeek V4 Pro', 'S+', '80.6%', '1M'], // ⚠️ Page-only / partner-routed (2026-08-13): listed on build.nvidia.com but NOT in integrate.api.nvidia.com/v1/models; served via Fireworks/DeepInfra/Together/OpenRouter
48
+ ['deepseek-ai/deepseek-v4-flash-0731', 'DeepSeek V4 Flash', 'S+', '79.0%', '1M'], // Fixed (2026-08-13): id 'deepseek-ai/deepseek-v4-flash' → 'deepseek-ai/deepseek-v4-flash-0731' (NIM /v1/models only exposes the -0731 suffix)
49
49
  ['stepfun-ai/step-3.7-flash', 'Step 3.7 Flash', 'S+', '74.4%', '256k'],
50
50
  ['nvidia/nemotron-3-ultra-550b-a55b', 'Nemotron 3 Ultra', 'S+', '71.9%', '1M'],
51
+ ['poolside/laguna-xs-2.1', 'Laguna XS 2.1', 'S+', '70.9%', '262k'], // Added (2026-08-13)
51
52
  // ── S tier — SWE-bench Verified 60–70% ──
52
53
  ['openai/gpt-oss-120b', 'GPT OSS 120B', 'S', '62.4%', '128k'],
53
54
  // Removed (2026-07-27): meta/llama-4-maverick-17b-128e-instruct (Llama 4 Maverick) — EOL 2026-07-27 (HTTP 410 Gone)
@@ -66,12 +67,16 @@ export const nvidiaNim = [
66
67
  ['nvidia/nemotron-3-nano-30b-a3b', 'Nemotron Nano 30B', 'A-', '38.8%', '1M'],
67
68
  ['openai/gpt-oss-20b', 'GPT OSS 20B', 'A+', '50.3%', '128k'],
68
69
  ['google/gemma-4-31b-it', 'Gemma 4 31B', 'A+', '52.0%', '256k'],
70
+ ['mistralai/mistral-large-2-instruct', 'Mistral Large 2', 'A+', '-', '128k'], // Added (2026-08-13)
69
71
  // Removed (2026-07-27): qwen/qwen2.5-coder-32b-instruct (Qwen2.5 Coder 32B) — EOL 2026-05-12 (HTTP 410 Gone)
70
72
  // Removed (2026-07-27): deepseek-ai/deepseek-r1 (DeepSeek R1) — HTTP 404
71
73
  // Removed (2026-07-27): nvidia/nemotron-3-nano (Nemotron 3 Nano) — HTTP 404 (replaced by nvidia/nvidia-nemotron-nano-9b-v2)
72
74
  ['nvidia/nvidia-nemotron-nano-9b-v2', 'Nemotron Nano 9B v2', 'A-', '-', '128k'], // Added (2026-07-27)
75
+ ['meta/llama-3.3-70b-instruct', 'Llama 3.3 70B', 'A+', '-', '128k'], // Added (2026-08-13)
73
76
  ['deepseek-ai/deepseek-coder-6.7b-instruct', 'DeepSeek Coder 6.7B', 'A-', '-', '128k'], // Added (2026-07-27)
77
+ ['meta/codellama-70b', 'CodeLlama 70B', 'A', '-', '100k'], // Added (2026-08-13)
74
78
  ['mistralai/codestral-22b-instruct-v0.1', 'Codestral 22B', 'A', '-', '32k'], // Added (2026-07-27)
79
+ ['ibm/granite-34b-code-instruct', 'Granite 34B Code', 'A-', '-', '128k'], // Added (2026-08-13)
75
80
  // ── A- tier — SWE-bench Verified 35–40% ──
76
81
  // Removed (2026-07-27): bytedance/seed-oss-36b-instruct (Seed OSS 36B) — EOL 2026-07-27 (HTTP 410 Gone)
77
82
  // Removed (2026-07-27): stockmark/stockmark-2-100b-instruct (Stockmark 100B) — EOL 2026-07-15 (HTTP 410 Gone)
@@ -88,8 +93,8 @@ export const nvidiaNim = [
88
93
  // 📖 Groq source - https://console.groq.com
89
94
  // 📖 Free API keys available at https://console.groq.com/keys
90
95
  export const groq = [
91
- ['llama-3.3-70b-versatile', 'Llama 3.3 70B', 'B', '22.0%', '131k'],
92
- ['llama-3.1-8b-instant', 'Llama 3.1 8B', 'C', '18.0%', '131k'],
96
+ // Removed (2026-08-13): llama-3.3-70b-versatile (Llama 3.3 70B) — Groq deprecation, shutdown 2026-08-16
97
+ // Removed (2026-08-13): llama-3.1-8b-instant (Llama 3.1 8B) — Groq deprecation, shutdown 2026-08-16
93
98
  ['openai/gpt-oss-120b', 'GPT OSS 120B', 'S', '62.4%', '131k'],
94
99
  ['openai/gpt-oss-20b', 'GPT OSS 20B', 'A+', '50.3%', '131k'],
95
100
  ['qwen/qwen3.6-27b', 'Qwen3.6 27B', 'S+', '77.2%', '131k'],
@@ -139,7 +144,7 @@ export const sambanova = [
139
144
  export const openrouter = [
140
145
  // ── S+ tier — SWE-bench Verified ≥70% ──
141
146
  ['nvidia/nemotron-3-ultra-550b-a55b:free', 'Nemotron 3 Ultra', 'S+', '71.9%', '1M'],
142
- ['poolside/laguna-m.1:free', 'Poolside Laguna M.1', 'S+', '72.5%', '262k'],
147
+ // Removed (2026-08-13): poolside/laguna-m.1:free (Poolside Laguna M.1) no longer in OpenRouter catalog (neither :free nor paid)
143
148
  ['poolside/laguna-xs-2.1:free', 'Poolside Laguna XS 2.1', 'S+', '70.9%', '262k'],
144
149
  // Removed (2026-07-27): poolside/laguna-xs.2:free (Poolside Laguna XS.2) — superseded by poolside/laguna-xs-2.1:free
145
150
  // ── S tier — SWE-bench Verified 60–70% ──
@@ -148,9 +153,11 @@ export const openrouter = [
148
153
  // Removed (2026-07-27): qwen/qwen3-coder:free (Qwen3 Coder) — no longer on OpenRouter free tier
149
154
  ['poolside/laguna-s-2.1:free', 'Poolside Laguna S 2.1', 'S+', '-', '262k'], // Added (2026-07-27)
150
155
  // ── A+ tier — SWE-bench Verified 50–60% ──
151
- ['nvidia/nemotron-3-super-120b-a12b:free', 'Nemotron 3 Super', 'S', '60.5%', '1M'],
156
+ ['nvidia/nemotron-3-super-120b-a12b:free', 'Nemotron 3 Super', 'S', '60.5%', '262k'], // Fixed (2026-08-13): ctx '1M' → '262k' (real API ctx)
152
157
  ['nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free', 'Nemotron 3 Omni', 'A+', '52.0%', '256k'],
153
- ['inclusionai/ling-3.0-flash:free', 'Ling-3.0 Flash', 'A+', '-', '262k'], // Added (2026-07-27)
158
+ // Removed (2026-08-13): inclusionai/ling-3.0-flash:free (Ling-3.0 Flash) :free variant removed, now paid-only
159
+ ['liquid/lfm-2.5-2.6b:free', 'LiquidAI LFM2.5-2.6B', 'C', '-', '128k'], // Added (2026-08-13)
160
+ ['nvidia/nemotron-3.5-lightning:free', 'NVIDIA Nemotron 3.5 Lightning', 'B+', '-', '1M'], // Added (2026-08-13)
154
161
  // ── A tier — SWE-bench Verified 40–50% ──
155
162
  ['nvidia/nemotron-3-nano-30b-a3b:free', 'Nemotron Nano 30B', 'A-', '38.8%', '256k'],
156
163
  ['nvidia/nemotron-nano-12b-v2-vl:free', 'Nemotron Nano 12B VL', 'A', '20.0%', '128k'],
@@ -173,51 +180,14 @@ export const openrouter = [
173
180
  ]
174
181
 
175
182
  // 📖 GitHub Models source - https://models.github.ai
176
- // 📖 OpenAI-compatible endpoint: https://models.github.ai/inference/chat/completions
177
- // 📖 Free usage is quota-limited by GitHub/Copilot tier, but no separate provider billing is needed.
183
+ // 📖 ⚠️ RETIRED 2026-07-30 GitHub Models fully shut down (playground, catalog, inference API, BYOK all gone)
184
+ // 📖 Catalog returns HTTP 410 Gone with `github_models_retirement_brownout` error.
185
+ // 📖 https://github.blog/changelog/2026-07-01-github-models-retirement
186
+ // 📖 Kept as an empty array so downstream provider-metadata/config code that imports
187
+ // 📖 `githubModels` and references `'github-models'` doesn't crash; the entry is also
188
+ // 📖 commented out of the `sources` map below so it won't appear in the catalog.
178
189
  export const githubModels = [
179
- // ── S+ tier SWE-bench Verified ≥70% ──
180
- ['openai/gpt-4.1', 'GPT-4.1', 'A+', '54.6%', '1M'],
181
- ['openai/gpt-5', 'GPT-5', 'S+', '74.9%', '200k'],
182
- ['openai/gpt-5-chat', 'GPT-5 Chat (preview)', 'S+', '-', '200k'],
183
- ['openai/o3', 'OpenAI o3', 'S', '69.1%', '200k'],
184
- // ── S tier — SWE-bench Verified 60–70% ──
185
- ['openai/gpt-4.1-mini', 'GPT-4.1 Mini', 'B', '23.6%', '1M'],
186
- ['deepseek/deepseek-v3-0324', 'DeepSeek V3 0324', 'A', '45.4%', '128k'],
187
- ['meta/llama-4-maverick-17b-128e-instruct-fp8', 'Llama 4 Maverick', 'S+', '74.8%', '1M'],
188
- ['openai/gpt-5-mini', 'GPT-5 Mini', 'S', '60.0%', '200k'],
189
- ['openai/gpt-4o', 'GPT-4o', 'B+', '33.2%', '128k'],
190
- ['openai/o4-mini', 'OpenAI o4-mini', 'S', '68.1%', '200k'],
191
- ['openai/o1', 'OpenAI o1', 'A', '48.9%', '200k'],
192
- ['deepseek/deepseek-r1', 'DeepSeek-R1', 'A', '49.2%', '128k'],
193
- ['deepseek/deepseek-r1-0528', 'DeepSeek-R1-0528', 'A+', '57.6%', '128k'],
194
- // ── A tier — SWE-bench Verified 40–50% ──
195
- ['openai/gpt-4.1-nano', 'GPT-4.1 Nano', 'A', '-', '1M'],
196
- ['meta/meta-llama-3.1-405b-instruct', 'Llama 3.1 405B', 'A', '40.6%', '128k'],
197
- ['meta/llama-4-scout-17b-16e-instruct', 'Llama 4 Scout', 'B', '28.0%', '10M'],
198
- ['mistral-ai/mistral-medium-2505', 'Mistral Medium 2505', 'A', '48.0%', '128k'],
199
- ['openai/gpt-5-nano', 'GPT-5 Nano', 'A', '-', '200k'],
200
- ['openai/gpt-4o-mini', 'GPT-4o Mini', 'A', '-', '128k'],
201
- ['openai/o3-mini', 'OpenAI o3-mini', 'A', '49.3%', '200k'],
202
- ['openai/o1-preview', 'OpenAI o1-preview', 'A', '41.3%', '128k'],
203
- // ── A- tier — SWE-bench Verified 35–40% ──
204
- ['meta/llama-3.3-70b-instruct', 'Llama 3.3 70B', 'B', '22.0%', '128k'],
205
- ['meta/llama-3.2-90b-vision-instruct', 'Llama 3.2 90B Vision', 'A-', '-', '128k'],
206
- ['cohere/cohere-command-a', 'Cohere Command A', 'C', '7.8%', '128k'],
207
- // ── B+ tier — SWE-bench Verified 30–35% ──
208
- ['mistral-ai/codestral-2501', 'Codestral 2501', 'B+', '34.0%', '256k'],
209
- ['mistral-ai/mistral-small-2503', 'Mistral Small 2503', 'B+', '30.0%', '128k'],
210
- ['openai/o1-mini', 'OpenAI o1-mini', 'B+', '-', '128k'],
211
- ['microsoft/phi-4', 'Phi-4', 'B+', '-', '16k'],
212
- ['microsoft/phi-4-reasoning', 'Phi-4-reasoning', 'B+', '-', '32k'],
213
- // ── B tier — SWE-bench Verified 20–30% ──
214
- ['meta/llama-3.2-11b-vision-instruct', 'Llama 3.2 11B Vision', 'B', '-', '128k'],
215
- ['meta/meta-llama-3.1-8b-instruct', 'Llama 3.1 8B', 'C', '18.0%', '128k'],
216
- ['microsoft/phi-4-mini-instruct', 'Phi-4-mini-instruct', 'B', '-', '128k'],
217
- ['microsoft/phi-4-mini-reasoning', 'Phi-4-mini-reasoning', 'B', '-', '128k'],
218
- ['microsoft/phi-4-multimodal-instruct', 'Phi-4-multimodal-instruct', 'B', '-', '128k'],
219
- // ── C tier — lightweight/edge models ──
220
- ['mistral-ai/ministral-3b', 'Ministral 3B', 'C', '-', '128k'],
190
+ // All 35 entries retired 2026-07-30 see comment above.
221
191
  ]
222
192
 
223
193
  // 📖 Mistral La Plateforme source - https://console.mistral.ai
@@ -227,14 +197,14 @@ export const mistral = [
227
197
  // ── S+ tier — SWE-bench Verified ≥70% ──
228
198
  ['mistral-large-2512', 'Mistral Large 3', 'S+', '70.0%', '256k'],
229
199
  ['mistral-medium-3-5', 'Mistral Medium 3.5', 'S+', '77.6%', '256k'],
230
- ['devstral-2512', 'Devstral 2', 'S+', '72.2%', '256k'],
200
+ // Removed (2026-08-13): devstral-2512 (Devstral 2) Mistral deprecation, full retirement 2026-07-31
231
201
  // ── A tier — SWE-bench Verified 40–50% ──
232
202
  ['mistral-small-2603', 'Mistral Small 4', 'A', '48.0%', '256k'],
233
203
  // ── B+ tier — SWE-bench Verified 30–35% ──
234
- ['ministral-3-14b-25-12', 'Ministral 3 14B', 'B+', '-', '128k'],
204
+ ['ministral-14b-2512', 'Ministral 3 14B', 'B+', '-', '128k'], // Fixed (2026-08-13): id 'ministral-3-14b-25-12' → 'ministral-14b-2512' (API model ID per Mistral docs JSON)
235
205
  // ── B tier — SWE-bench Verified 20–30% ──
236
- ['ministral-3-8b-25-12', 'Ministral 3 8B', 'B', '-', '128k'],
237
- ['ministral-3-3b-25-12', 'Ministral 3 3B', 'B', '-', '128k'],
206
+ ['ministral-8b-2512', 'Ministral 3 8B', 'B', '-', '128k'], // Fixed (2026-08-13): id 'ministral-3-8b-25-12' → 'ministral-8b-2512'
207
+ ['ministral-3b-2512', 'Ministral 3 3B', 'B', '-', '128k'], // Fixed (2026-08-13): id 'ministral-3-3b-25-12' → 'ministral-3b-2512'
238
208
  ]
239
209
 
240
210
  // 📖 Mistral Codestral source - https://codestral.mistral.ai
@@ -243,29 +213,32 @@ export const mistral = [
243
213
  export const codestral = [
244
214
  // ── A tier — SWE-bench Verified 40–50% ──
245
215
  ['codestral-2508', 'Codestral', 'A', '40.0%', '128k'], // Fixed (2026-07-27): ctx '256k' → '128k' per official Mistral model card
246
- ['codestral-2', 'Codestral 2', 'B+', '-', '128k'],
216
+ ['codestral-2501', 'Codestral 2501', 'B+', '34.0%', '256k'], // Added (2026-08-13)
217
+ ['codestral-2405', 'Codestral 2405', 'B', '30.0%', '32k'], // Added (2026-08-13)
218
+ // Removed (2026-08-13): codestral-2 (Codestral 2) — fabricated ID, never existed in Mistral catalog (Mistral uses date-stamped versioning)
247
219
  ]
248
220
 
249
221
  // 📖 Scaleway source - https://console.scaleway.com
250
222
  // 📖 1M free tokens — API keys at https://console.scaleway.com/iam/api-keys
251
223
  export const scaleway = [
252
224
  // ── S+ tier — SWE-bench Verified ≥70% ──
253
- ['devstral-2-123b-instruct-2512', 'Devstral 2 123B', 'S+', '72.2%', '260k'], // Fixed (2026-07-27): ctx '200k' → '260k' (Dedicated tier)
225
+ ['devstral-2-123b-instruct-2512', 'Devstral 2 123B', 'S+', '72.2%', '200k'], // Fixed (2026-08-13): ctx '260k' → '200k' (Serverless tier per official Scaleway catalog)
254
226
  ['qwen3-235b-a22b-instruct-2507', 'Qwen3 235B', 'A', '45.2%', '250k'],
255
- ['glm-5.2', 'GLM 5.2', 'S+', '82.8%', '1M'], // Fixed (2026-07-27): ctx '256k' → '1M' (preview extended)
227
+ ['glm-5.2', 'GLM 5.2', 'S+', '82.8%', '256k'], // Fixed (2026-08-13): ctx '1M' → '256k' (Serverless tier per official catalog)
228
+ ['deepseek-v4-flash-0731', 'DeepSeek V4 Flash', 'S+', '-', '256k'], // Added (2026-08-13)
256
229
  // ── S tier — SWE-bench Verified 60–70% ──
257
230
  ['qwen3.5-397b-a17b', 'Qwen3.5 400B VLM', 'S+', '76.2%', '250k'],
258
231
  ['gpt-oss-120b', 'GPT OSS 120B', 'S', '62.4%', '128k'],
259
232
  ['mistral-medium-3.5-128b', 'Mistral Medium 3.5 128B', 'S+', '77.6%', '180k'], // Fixed (2026-07-27): ctx '256k' → '180k' (Serverless tier)
260
233
  // ── A+ tier — SWE-bench Verified 50–60% ──
261
- ['mistral-large-3-675b-instruct-2512', 'Mistral Large 675B', 'A+', '58.0%', '250k'],
234
+ // Removed (2026-08-13): mistral-large-3-675b-instruct-2512 (Mistral Large 675B) Dedicated tier only, not available on Serverless
262
235
  ['qwen3-coder-30b-a3b-instruct', 'Qwen3 Coder 30B', 'A+', '51.6%', '128k'],
263
236
  ['qwen3.6-35b-a3b', 'Qwen3.6 35B MoE', 'S+', '73.4%', '256k'],
264
237
  ['holo2-30b-a3b', 'Holo2 30B', 'A+', '52.0%', '22k'],
265
238
  ['gemma-4-26b-a4b-it', 'Gemma 4 26B MoE', 'A+', '-', '256k'],
266
- ['gemma-4-31b-it', 'Gemma 4 31B IT', 'A+', '-', '128k'], // Added (2026-07-27)
239
+ // Removed (2026-08-13): gemma-4-31b-it (Gemma 4 31B IT) Dedicated tier only, not available on Serverless
267
240
  // ── A- tier — SWE-bench Verified 35–40% ──
268
- ['llama-3.3-70b-instruct', 'Llama 3.3 70B', 'B', '22.0%', '128k'], // Fixed (2026-07-27): ctx '100k' → '128k' (Dedicated tier)
241
+ ['llama-3.3-70b-instruct', 'Llama 3.3 70B', 'B', '22.0%', '100k'], // Fixed (2026-08-13): ctx '128k' → '100k' (Serverless tier per official catalog)
269
242
  // ── B+ tier — SWE-bench Verified 30–35% ──
270
243
  ['mistral-small-3.2-24b-instruct-2506', 'Mistral Small 3.2', 'B', '20.0%', '128k'],
271
244
  ['pixtral-12b-2409', 'Pixtral 12B', 'B+', '-', '128k'],
@@ -276,6 +249,7 @@ export const scaleway = [
276
249
  // 📖 Google AI Studio source - https://aistudio.google.com
277
250
  // 📖 OpenAI-compatible endpoint exposes Gemini models; free quotas vary by model and region.
278
251
  export const googleai = [
252
+ ['gemini-3.7-flash', 'Gemini 3.7 Flash', 'S+', '-', '1M'], // Added (2026-08-13)
279
253
  ['gemini-3.6-flash', 'Gemini 3.6 Flash', 'S+', '-', '1M'], // Added (2026-07-27)
280
254
  ['gemini-3.5-flash', 'Gemini 3.5 Flash', 'S+', '78.0%', '1M'],
281
255
  ['gemini-3.1-pro-preview', 'Gemini 3.1 Pro Preview', 'S+', '80.6%', '1M'],
@@ -290,11 +264,19 @@ export const googleai = [
290
264
  // 📖 ZAI source - https://open.z.ai
291
265
  // 📖 Free tier is limited to Flash models; paid GLM models are intentionally excluded.
292
266
  export const zai = [
267
+ // ── S+ tier — SWE-bench Verified ≥70% ──
268
+ ['zai/glm-5.2', 'GLM-5.2', 'S+', '-', '1M'], // Added (2026-08-13)
293
269
  // ── S tier — SWE-bench Verified 60–70% ──
294
270
  ['zai/glm-4.7-flash', 'GLM-4.7-Flash', 'A+', '59.2%', '200k'], // Fixed (2026-07-27): ctx '203k' → '200k' per official docs
295
271
  ['zai/glm-4.5-flash', 'GLM-4.5-Flash', 'S', '59.2%', '128k'],
272
+ ['zai/glm-5-turbo', 'GLM-5-Turbo', 'S', '-', '200k'], // Added (2026-08-13)
273
+ ['zai/glm-5v-turbo', 'GLM-5V-Turbo', 'S', '-', '200k'], // Added (2026-08-13)
274
+ ['zai/glm-4.7', 'GLM-4.7', 'S', '-', '200k'], // Added (2026-08-13)
275
+ ['zai/glm-4.6', 'GLM-4.6', 'S', '-', '200k'], // Added (2026-08-13)
276
+ ['zai/glm-4.7-flashx', 'GLM-4.7-FlashX', 'A+', '-', '200k'], // Added (2026-08-13)
296
277
  // ── A tier — SWE-bench Verified 40–50% ──
297
278
  ['zai/glm-4.6v-flash', 'GLM-4.6V-Flash', 'A', '-', '128k'],
279
+ ['zai/glm-4.6v', 'GLM-4.6V', 'A', '-', '128k'], // Added (2026-08-13)
298
280
  ]
299
281
 
300
282
  // 📖 Alibaba Cloud (DashScope) source - https://dashscope-intl.aliyuncs.com
@@ -321,6 +303,7 @@ export const qwen = [
321
303
  ['qwen3.6-flash', 'Qwen3.6 Flash', 'A+', '60.0%', '1M'],
322
304
  ['qwen3.5-flash', 'Qwen3.5 Flash', 'S', '64.4%', '1M'],
323
305
  ['qwen3-coder-flash', 'Qwen3 Coder Flash', 'A+', '55.0%', '1M'],
306
+ ['qwen3-vl-flash', 'Qwen3 VL Flash', 'A+', '-', '256k'], // Added (2026-08-13)
324
307
  ['qwen3-32b', 'Qwen3 32B', 'B+', '30.0%', '128k'],
325
308
  ['qwen3.5-397b-a17b', 'Qwen3.5 397B A17B', 'S+', '76.2%', '256k'],
326
309
  ['qwen3.5-122b-a10b', 'Qwen3.5 122B A10B', 'S+', '72.0%', '256k'],
@@ -361,6 +344,8 @@ export const cloudflare = [
361
344
  ['@cf/ibm-granite/granite-4.0-h-micro', 'Granite 4.0 Micro', 'B+', '30.0%', '128k'], // Fixed (2026-07-27): namespace 'ibm' → 'ibm-granite'
362
345
  // ── B tier — SWE-bench Verified 20–30% ──
363
346
  ['@cf/meta/llama-3.1-8b-instruct-fast', 'Llama 3.1 8B Instruct (Fast)', 'C', '18.0%', '128k'],
347
+ ['@cf/google/gemma-3-12b-it', 'Gemma 3 12B IT', 'A', '-', '128k'], // Added (2026-08-13)
348
+ ['@cf/moonshotai/kimi-k2.5', 'Kimi K2.5', 'S+', '-', '256k'], // Added (2026-08-13)
364
349
  ]
365
350
 
366
351
  // 📖 OVHcloud AI Endpoints - https://endpoints.ai.cloud.ovh.net
@@ -375,10 +360,11 @@ export const ovhcloud = [
375
360
  ['gpt-oss-20b', 'GPT OSS 20B', 'A+', '50.3%', '131k'],
376
361
  ['Meta-Llama-3_3-70B-Instruct', 'Llama 3.3 70B', 'B', '22.0%', '131k'],
377
362
  // Removed (2026-07-27): Qwen3-32B (Qwen3 32B) — no longer in catalog
378
- ['Mistral-Small-3.2-24B-Instruct-2506', 'Mistral Small 3.2', 'B', '20.0%', '128k'],
363
+ // Removed (2026-08-13): Mistral-Small-3.2-24B-Instruct-2506 (Mistral Small 3.2) no longer in OVHcloud public catalog (endpoint still reachable but not listed)
379
364
  // Removed (2026-07-27): Mistral-7B-Instruct-v0.3 (Mistral 7B Instruct) — no longer in catalog
380
- ['Mistral-Nemo-Instruct-2407', 'Mistral Nemo', 'B+', '30.0%', '118k'],
365
+ // Removed (2026-08-13): Mistral-Nemo-Instruct-2407 (Mistral Nemo) no longer in OVHcloud public catalog
381
366
  ['Qwen3.5-9B', 'Qwen3.5 9B', 'B+', '30.0%', '262k'],
367
+ ['Qwen2.5-VL-72B-Instruct', 'Qwen2.5-VL 72B', 'S', '-', '131k'], // Added (2026-08-13)
382
368
  // ── Embeddings ──
383
369
  ['Qwen3-Embedding-8B', 'Qwen3 Embedding 8B', 'B', '-', '32k'], // Fixed (2026-07-27): ctx '-' → '32k'
384
370
  ['bge-m3', 'BGE M3', 'B', '-', '-'],
@@ -398,9 +384,11 @@ export const opencodeZen = [
398
384
  ['deepseek-v4-flash-free', 'DeepSeek V4 Flash Free', 'S+', '79.0%', '200k'],
399
385
  ['mimo-v2.5-free', 'MiMo-V2.5 Free', 'S+', '-', '200k'],
400
386
  ['nemotron-3-ultra-free', 'Nemotron 3 Ultra Free', 'S+', '71.9%', '200k'],
401
- ['north-mini-code-free', 'North Mini Code Free', 'B+', '-', '200k'],
402
- ['poolside/laguna-s-2.1-free', 'Laguna S 2.1 Free', 'S+', '-', '262k'], // Added (2026-07-27)
403
- ['ling-3.0-flash-free', 'Ling-3.0-flash Free', 'A+', '-', '262k'], // Added (2026-07-27)
387
+ // Removed (2026-08-13): north-mini-code-free (North Mini Code Free) — no longer in OpenCode Zen free-tier API
388
+ ['laguna-s-2.1-free', 'Laguna S 2.1 Free', 'S+', '-', '262k'], // Fixed (2026-08-13): ID 'poolside/laguna-s-2.1-free' → 'laguna-s-2.1-free' (poolside/ prefix dropped)
389
+ // Removed (2026-08-13): ling-3.0-flash-free (Ling-3.0-flash Free) no longer in OpenCode Zen free-tier API
390
+ ['hy3-free', 'Tencent Hy3 Free', 'S', '-', '200k'], // Added (2026-08-13) — brought back after July removal
391
+ ['nemotron-3.5-lightning-free', 'Nemotron 3.5 Lightning Free','S+','-', '200k'], // Added (2026-08-13)
404
392
  // Removed (2026-07-27): hy3-free (Tencent Hy3 Free) — no longer on OpenCode Zen
405
393
  ]
406
394
 
@@ -409,6 +397,7 @@ export const opencodeZen = [
409
397
  // 📖 Keep only the stable router model here; individual promo `:free` models churn too quickly.
410
398
  export const kilo = [
411
399
  ['kilo-auto/free', 'Kilo Auto Free', 'A+', '-', '256k'],
400
+ ['kilo-auto/small', 'Kilo Auto Small', 'B+', '-', '262k'], // Added (2026-08-13) — routes to gemma-4-26b-a4b-it:free for free accounts
412
401
  ]
413
402
 
414
403
  // 📖 LLM7 source - https://api.llm7.io/v1
@@ -419,8 +408,9 @@ export const llm7 = [
419
408
  // ── S+ tier — SWE-bench Verified ≥70% ──
420
409
  ['minimax-m2.7', 'MiniMax M2.7', 'S+', '78.0%', '180k'],
421
410
  // ── A+ tier — SWE-bench Verified 50–60% ──
422
- ['gemini-3.1-flash-lite', 'Gemini 3.1 Flash Lite', 'A+', '-', '1M'], // Added (2026-07-27)
411
+ ['gemini-3.1-flash-lite', 'Gemini 3.1 Flash Lite', 'A+', '-', '256k'], // Fixed (2026-08-13): ctx '1M' → '256k' (real LLM7 ctx limit)
423
412
  ['gpt-oss:20b', 'GPT OSS 20B', 'A+', '50.3%', '128k'],
413
+ ['mistral-Nemo-Instruct-2407', 'Mistral Nemo 12B Instruct', 'A-', '-', '128k'], // Added (2026-08-13)
424
414
  // ── A tier — SWE-bench Verified 40–50% ──
425
415
  ['codestral-latest', 'Codestral Latest', 'A', '40.0%', '32k'],
426
416
  ]
@@ -430,14 +420,14 @@ export const llm7 = [
430
420
  // 📖 Live catalog checked 2026-06-11; only chat-completions models with free pricing are listed.
431
421
  export const routeway = [
432
422
  // ── S+ tier — SWE-bench Verified ≥70% ──
433
- ['deepseek-v4-flash:free', 'DeepSeek V4 Flash', 'S+', '79.0%', '1M'],
434
- ['step-3.5-flash:free', 'Step 3.5 Flash', 'S+', '74.4%', '256k'],
423
+ // Removed (2026-08-13): deepseek-v4-flash:free (DeepSeek V4 Flash) no longer in zero-price catalog (only paid DeepSeek V4 Pro 0813 remains)
424
+ // Removed (2026-08-13): step-3.5-flash:free (Step 3.5 Flash) superseded by step-3.7-flash:free
435
425
  // Removed (2026-07-27): laguna-m.1:free (Poolside Laguna M.1) — unavailable on Routeway
436
426
  ['laguna-xs.2:free', 'Poolside Laguna XS.2', 'S', '68.2%', '131k'],
437
427
  ['step-3.7-flash:free', 'Step 3.7 Flash', 'S+', '74.4%', '256k'], // Added (2026-07-27)
438
428
  // ── S tier — SWE-bench Verified 60–70% ──
439
- ['ling-3.0-flash:free', 'Ling 3.0 Flash', 'S', '-', '256k'], // Added (2026-07-27)
440
- ['ling-2.6-flash:free', 'Ling 2.6 Flash', 'S', '61.2%', '262k'],
429
+ // Removed (2026-08-13): ling-3.0-flash:free (Ling 3.0 Flash) no longer in zero-price catalog
430
+ // Removed (2026-08-13): ling-2.6-flash:free (Ling 2.6 Flash) no longer in zero-price catalog
441
431
  ['gpt-oss-120b:free', 'GPT OSS 120B', 'S', '60.0%', '131k'],
442
432
  // ── A tier — SWE-bench Verified 40–50% ──
443
433
  ['gemma-4-31b-it:free', 'Gemma 4 31B', 'A+', '52.0%', '262k'],
@@ -457,8 +447,10 @@ export const routeway = [
457
447
  // 📖 Novita is mostly paid/trial-credit, so this catalog only includes live chat models reporting 0 input/output price.
458
448
  // 📖 Test/dev/placeholder zero-price IDs were intentionally excluded.
459
449
  export const novita = [
450
+ // ⚠️ Empty as of 2026-08-13 — tencent/hy3 was the last zero-price entry and is now paid-only ($0.14/Mt in, $0.58/Mt out).
451
+ // All other API entries with zero pricing are test/dev/placeholder/internal IDs and intentionally excluded.
460
452
  // ── S tier — SWE-bench Verified 60–70% ──
461
- ['tencent/hy3', 'Tencent Hy3', 'S', '-', '256k'], // Fixed (2026-07-27): ctx '262k' '256k' per official novita description
453
+ // Removed (2026-08-13): tencent/hy3 (Tencent Hy3) isFree:false per Novita pricing page
462
454
  // Removed (2026-07-27): qwen/qwen3.5-plus (Qwen3.5 Plus) — no longer in novita catalog
463
455
  ]
464
456
 
@@ -495,109 +487,157 @@ export const ollamaCloud = [
495
487
  // 📖 Each source has: name (display), url (API endpoint), models (array of model tuples)
496
488
  // 📖 Providers ordered by generosity of free tier (most generous first)
497
489
  // 📖 See README for full tier-by-tier comparison
490
+ // 📖 Each provider now carries a `quota` (human-readable summary) and a
491
+ // 📖 `quotaCode` (machine code: 'free' | 'limited' | 'metered') so the website
492
+ // 📖 can render a sortable, filterable catalog without re-typing the rules in
493
+ // 📖 a second file. The CLI ignores these fields, so this is fully backward-
494
+ // 📖 compatible with the existing TUI / router-daemon / OpenCode integration.
498
495
  export const sources = {
499
496
  nvidia: {
500
497
  name: 'NVIDIA NIM',
501
498
  url: 'https://integrate.api.nvidia.com/v1/chat/completions',
499
+ quota: 'Free · 1000 req/month',
500
+ quotaCode: 'free',
502
501
  models: nvidiaNim,
503
502
  },
504
503
  groq: {
505
504
  name: 'Groq',
506
505
  url: 'https://api.groq.com/openai/v1/chat/completions',
506
+ quota: 'Free · ~30-50 RPM per model',
507
+ quotaCode: 'free',
507
508
  models: groq,
508
509
  },
509
510
  cerebras: {
510
511
  name: 'Cerebras',
511
512
  url: 'https://api.cerebras.ai/v1/chat/completions',
513
+ quota: 'Free · generous dev tier',
514
+ quotaCode: 'free',
512
515
  models: cerebras,
513
516
  },
514
517
  googleai: {
515
518
  name: 'Google AI',
516
519
  url: 'https://generativelanguage.googleapis.com/v1beta/openai/chat/completions',
520
+ quota: 'Free · Gemini quotas vary by model',
521
+ quotaCode: 'free',
517
522
  models: googleai,
518
523
  },
519
- 'github-models': {
520
- name: 'GitHub Models',
521
- url: 'https://models.github.ai/inference/chat/completions',
522
- models: githubModels,
523
- },
524
+ // 'github-models': REMOVED 2026-08-13 — GitHub Models retired 2026-07-30 (HTTP 410 Gone).
525
+ // Provider metadata still references this key for backwards compat in user configs,
526
+ // but it is no longer exposed in the catalog.
527
+ // 'github-models': {
528
+ // name: 'GitHub Models',
529
+ // url: 'https://models.github.ai/inference/chat/completions',
530
+ // quota: 'GitHub / Copilot plan quota',
531
+ // quotaCode: 'metered',
532
+ // models: githubModels,
533
+ // },
524
534
  mistral: {
525
535
  name: 'Mistral LP',
526
536
  url: 'https://api.mistral.ai/v1/chat/completions',
537
+ quota: 'Free Experiment plan',
538
+ quotaCode: 'free',
527
539
  models: mistral,
528
540
  },
529
541
  cloudflare: {
530
542
  name: 'Cloudflare AI',
531
543
  url: 'https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/v1/chat/completions',
544
+ quota: 'Free · 10k neurons/day',
545
+ quotaCode: 'limited',
532
546
  models: cloudflare,
533
547
  },
534
548
  openrouter: {
535
549
  name: 'OpenRouter',
536
550
  url: 'https://openrouter.ai/api/v1/chat/completions',
551
+ quota: '50 free req/day · 1000 with $10 credit',
552
+ quotaCode: 'limited',
537
553
  models: openrouter,
538
554
  },
539
555
  sambanova: {
540
556
  name: 'SambaNova',
541
557
  url: 'https://api.sambanova.ai/v1/chat/completions',
558
+ quota: 'Small dev tier · light use',
559
+ quotaCode: 'limited',
542
560
  models: sambanova,
543
561
  },
544
562
  ovhcloud: {
545
563
  name: 'OVHcloud AI',
546
564
  url: 'https://oai.endpoints.kepler.ai.cloud.ovh.net/v1/chat/completions',
565
+ quota: 'Free sandbox · 2 RPM no key · 400 RPM with key',
566
+ quotaCode: 'free',
547
567
  models: ovhcloud,
548
568
  },
549
569
  codestral: {
550
570
  name: 'Codestral',
551
571
  url: 'https://api.mistral.ai/v1/chat/completions',
572
+ quota: 'Free · 30 req/min, 2000/day',
573
+ quotaCode: 'free',
552
574
  models: codestral,
553
575
  },
554
576
  zai: {
555
577
  name: 'ZAI',
556
578
  url: 'https://api.z.ai/api/coding/paas/v4/chat/completions',
579
+ quota: 'Free · Flash models only',
580
+ quotaCode: 'free',
557
581
  models: zai,
558
582
  },
559
583
  scaleway: {
560
584
  name: 'Scaleway',
561
585
  url: 'https://api.scaleway.ai/v1/chat/completions',
586
+ quota: '1M free tokens',
587
+ quotaCode: 'limited',
562
588
  models: scaleway,
563
589
  },
564
590
  qwen: {
565
591
  name: 'Alibaba DashScope',
566
592
  url: 'https://dashscope-intl.aliyuncs.com/compatible-mode/v1/chat/completions',
593
+ quota: '1M tokens/model · 90 days (Singapore)',
594
+ quotaCode: 'limited',
567
595
  models: qwen,
568
596
  },
569
597
 
570
598
  'opencode-zen': {
571
599
  name: 'OpenCode Zen',
572
600
  url: 'https://opencode.ai/zen/v1/chat/completions',
601
+ quota: 'Free · Zen key required',
602
+ quotaCode: 'free',
573
603
  models: opencodeZen,
574
604
  zenOnly: true,
575
605
  },
576
606
  kilo: {
577
607
  name: 'Kilo',
578
608
  url: 'https://api.kilo.ai/api/gateway/chat/completions',
609
+ quota: 'Free · no key needed',
610
+ quotaCode: 'free',
579
611
  models: kilo,
580
612
  noKeyNeeded: true,
581
613
  },
582
614
  llm7: {
583
615
  name: 'LLM7',
584
616
  url: 'https://api.llm7.io/v1/chat/completions',
617
+ quota: 'Free · no key needed',
618
+ quotaCode: 'limited',
585
619
  models: llm7,
586
620
  noKeyNeeded: true,
587
621
  },
588
622
  routeway: {
589
623
  name: 'Routeway',
590
624
  url: 'https://api.routeway.ai/v1/chat/completions',
625
+ quota: 'Free :free models only',
626
+ quotaCode: 'free',
591
627
  models: routeway,
592
628
  },
593
629
  novita: {
594
630
  name: 'Novita AI',
595
631
  url: 'https://api.novita.ai/openai/v1/chat/completions',
596
- models: novita,
632
+ quota: 'No zero-price models as of 2026-08-13',
633
+ quotaCode: 'limited',
634
+ models: novita, // Empty — kept for backward compat in user configs
597
635
  },
598
636
  'ollama-cloud': {
599
637
  name: 'Ollama Cloud',
600
638
  url: 'https://ollama.com/v1/chat/completions',
639
+ quota: 'Free plan · session + weekly caps',
640
+ quotaCode: 'free',
601
641
  models: ollamaCloud,
602
642
  },
603
643
  }
@@ -147,7 +147,7 @@ function getManagedProviderLabel(providerKey) {
147
147
  return `FCM ${getProviderLabel(providerKey)}`
148
148
  }
149
149
 
150
- function parseContextWindow(ctx) {
150
+ export function parseContextWindow(ctx) {
151
151
  if (typeof ctx !== 'string' || !ctx.trim()) return 128000
152
152
  const trimmed = ctx.trim().toLowerCase()
153
153
  const multiplier = trimmed.endsWith('m') ? 1_000_000 : trimmed.endsWith('k') ? 1_000 : 1
@@ -36,7 +36,7 @@ import { sleep } from './shared-helpers.js'
36
36
  // 📖 is not guaranteed to be accepted by their chat endpoint.
37
37
  export const PROVIDER_TEST_MODEL_OVERRIDES = {
38
38
  sambanova: ['MiniMax-M2.5', 'DeepSeek-V3.1', 'DeepSeek-V3.2'],
39
- nvidia: ['deepseek-ai/deepseek-v4-flash', 'openai/gpt-oss-120b'],
39
+ nvidia: ['deepseek-ai/deepseek-v4-flash-0731', 'openai/gpt-oss-120b'],
40
40
  'github-models': ['openai/gpt-4.1-mini'],
41
41
  mistral: ['mistral-small-latest', 'devstral-small-latest'],
42
42
  }