free-coding-models 0.5.93 โ†’ 0.5.95

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,7 +6,7 @@
6
6
 
7
7
  <p align="center">
8
8
  <strong>Find the fastest free coding model in seconds.</strong><br>
9
- Live latency, stability and verdicts for 251 models from 24 free AI providers, then install the one you pick straight into your favorite coding tool.<br><br>
9
+ Live latency, stability and verdicts for 256 models from 24 free AI providers, then install the one you pick straight into your favorite coding tool.<br><br>
10
10
  <strong>Works with:</strong> OpenCode CLI / Desktop / WebUI, OpenClaw, Crush, Goose, Aider, Kilo CLI, Qwen Code, OpenHands, Amp, Hermes, Continue, Cline, Xcode, Pi, ZCode, ForgeCode, Copilot, jcode, Caveman Code and more.
11
11
  </p>
12
12
 
@@ -35,7 +35,7 @@ free-coding-models
35
35
 
36
36
  ## ๐Ÿ’ก Why this tool?
37
37
 
38
- There is a large catalog of free and free-limited coding models (**24 providers / 251 live models**, generated from [`sources.js`](./sources.js)). Which one is fastest *right now*? Which one is actually stable, versus just lucky on the last ping?
38
+ There is a large catalog of free and free-limited coding models (**24 providers / 256 live models**, generated from [`sources.js`](./sources.js)). Which one is fastest *right now*? Which one is actually stable, versus just lucky on the last ping?
39
39
 
40
40
  `free-coding-models` (FCM) answers that by pinging every model in parallel, showing live latency, and computing a **live Stability Score (0-100)** combining p95 latency, jitter, spike rate and uptime. Average latency alone is misleading: a model that randomly spikes to 6 seconds is not reliable.
41
41
 
@@ -87,18 +87,18 @@ free-coding-models --fiable # print the single most reliable model
87
87
 
88
88
  ## ๐ŸŸข Providers
89
89
 
90
- **24 active providers / 251 live models**, sorted by live model count. Top 8:
90
+ **24 active providers / 256 live models**, sorted by live model count. Top 8:
91
91
 
92
92
  | Provider | Models | Best tier | Env var |
93
93
  |----------|--------|-----------|---------|
94
- | [Alibaba DashScope](https://modelstudio.console.alibabacloud.com) | 25 | S+ | `DASHSCOPE_API_KEY` |
95
- | [Cloudflare AI](https://dash.cloudflare.com) | 21 | S+ | `CLOUDFLARE_API_TOKEN` |
94
+ | [Alibaba DashScope](https://modelstudio.console.alibabacloud.com) | 29 | S+ | `DASHSCOPE_API_KEY` |
95
+ | [Pollinations AI](https://enter.pollinations.ai) | 23 | S+ | `POLLINATIONS_API_KEY` |
96
+ | [OpenRouter](https://openrouter.ai/keys) | 21 | S+ | `OPENROUTER_API_KEY` |
97
+ | [Kilo](https://kilo.ai) | 20 | S+ | `KILO_API_KEY` |
96
98
  | [Ollama Cloud](https://ollama.com/settings/keys) | 20 | S+ | `OLLAMA_API_KEY` |
97
- | [OpenRouter](https://openrouter.ai/keys) | 20 | S+ | `OPENROUTER_API_KEY` |
98
- | [Kilo](https://kilo.ai) | 18 | S+ | `KILO_API_KEY` |
99
- | [NVIDIA NIM](https://build.nvidia.com) | 16 | S+ | `NVIDIA_API_KEY` |
100
- | [OVHcloud AI](https://endpoints.ai.cloud.ovh.net) | 13 | S+ | `OVH_AI_ENDPOINTS_ACCESS_TOKEN` |
101
- | [Pollinations AI](https://enter.pollinations.ai) | 13 | S+ | `POLLINATIONS_API_KEY` |
99
+ | [OVHcloud AI](https://endpoints.ai.cloud.ovh.net) | 17 | S+ | `OVH_AI_ENDPOINTS_ACCESS_TOKEN` |
100
+ | [Cloudflare AI](https://dash.cloudflare.com) | 15 | S+ | `CLOUDFLARE_API_TOKEN` |
101
+ | [NVIDIA NIM](https://build.nvidia.com) | 12 | S+ | `NVIDIA_API_KEY` |
102
102
 
103
103
  > ๐Ÿงพ **What "free" means here:** free is a property of the *(provider, model)* pair, never of the provider as a whole. A row is listed only when that exact model id costs $0 to call through that provider (permanent free tier, `:free` variant, or free plan), verified live at audit time. The same open-weights model can be free on one host and paid on another - paid siblings are deliberately excluded. Full breakdown and badge legend: [`docs/providers.md`](./docs/providers.md).
104
104
 
@@ -0,0 +1,37 @@
1
+ # Changelog v0.5.94 - 2026-09-21
2
+
3
+ Full 24-provider audit, every model re-checked against live APIs and official docs by 24 parallel researchers. Catalog goes 251 to 256 live free models. This audit also cleaned up several mistakes the 2026-09-15 audit introduced (dead or paid-only models that were wrongly re-added).
4
+
5
+ ### Removed (22)
6
+
7
+ - **NVIDIA NIM (-5):** deepseek-v4-flash-0731 died today (NVIDIA banner: deprecated 09/19, unsupported after 09/21), deepseek-v4-pro-0813 (410 Gone, EOL 09/14), qwen3-coder-480b-a35b-instruct (410 Gone since June - the 09-15 re-add resurrected a months-dead model), minimaxai/minimax-m3 (410 Gone, EOL 09/09) and minimaxai/minimax-m2.7 (410 Gone since 07/27).
8
+ - **Groq (-5):** the free tier collapsed. llama-3.1-8b-instant and llama-3.3-70b-versatile were shut down on 08/16 (the 09-15 re-add reverted a correct removal), minimax-m2.7 is enterprise-only on Groq (model_not_found on developer keys), and groq/compound + groq/compound-mini hit their shutdown date today. Groq is down to 3 free models: gpt-oss-120b, gpt-oss-20b, qwen3.8-27b.
9
+ - **Cloudflare (-7):** glm-5.3, glm-5.3-flash, glm-5.2, deepseek-v4-pro-0813, deepseek-v4-flash-0731, kimi-k2.6 and kimi-k2.7-code all now carry the "Paid access required: not available through standard Workers Free billing" badge on the official docs pages. They are unusable on the free 10k neurons/day tier, so they are out (this re-applies the 09-05 paid-only policy the 09-15 audit regressed).
10
+ - **Z.ai (-2):** glm-5.1 and glm-5.2 now silently redirect to GLM-5.3 on the Coding Plan, so the ids no longer serve a distinct free model.
11
+ - **Google AI (-1):** gemini-3.1-pro-preview re-removed - Pro models left the free tier around April 2026; the 09-15 re-add resurrected a paid-only model. The Gemini 2.5 family is still free but retires no earlier than 2026-10-16 (noted in-file).
12
+ - **Cerebras (-1):** qwen-3-235b-a22b-instruct-2507 was deprecated 2026-05-27; the official catalog now serves exactly 2 free models (gpt-oss-120b, qwen3.8-27b).
13
+ - **SiliconFlow (-1):** Qwen2.5-Coder-7B-Instruct was taken offline on 2026-03-17 (official release note); the 09-15 re-add resurrected a 6-months-dead id.
14
+
15
+ ### Added (27)
16
+
17
+ - **Alibaba DashScope (+4):** qwen3-coder-480b-a35b-instruct is back on the free billing page (1M-token free quota, page updated 09/20), plus new qwen3-coder-30b-a3b-instruct, qwen3-vl-plus and qwen3.8-omni-flash.
18
+ - **Pollinations (+10):** the flagship wave arrived - moonshotai/kimi-k3, deepseek-v4-pro, qwen3.8-max, gemini-3.1-pro-preview (free here even though paid on Google), openai/gpt-5.5, openai/gpt-6-astra, openai/gpt-5.6-luna, z-ai/glm-5.3-flash, nvidia/nemotron-3-ultra and qwen3-coder-next. Every id verified directly against the live /v1/models list (411 models).
19
+ - **OVHcloud (+4):** Qwen3-Coder-30B-A3B-Instruct, Mistral Small 3.2 24B, Mistral Nemo 12B and Mistral 7B v0.3 are back in the official AI Endpoints catalog.
20
+ - **Google AI (+2):** gemma-4-31b-it and gemma-4-26b-a4b-it joined the free tier.
21
+ - **Kilo (+2):** qwen3.8-27b:free and nemotron-3-nano-omni:free are new in the live free gateway list (20 free models now).
22
+ - **OpenRouter (+1):** qwen/qwen3.8-27b:free.
23
+ - **Cloudflare (+1):** @cf/aisingapore/gemma-sea-lion-v4-27b-it (SEA-language focused, secondary for coding).
24
+ - **OpenCode Zen (+1):** deepseek-v4-flash-free returned to the free pool (re-added with a re-verify note since it is still absent from the docs pricing table).
25
+ - **LLM7 (+1):** GLM-5.3-Flash is free again (turbo tier, verified with a live unauthenticated chat probe).
26
+ - **SiliconFlow (+1):** XingChenAGI/Xing4.0-29B, a new $0 engineering-focused model.
27
+
28
+ ### Fixed
29
+
30
+ - **NVIDIA key testing was silently broken in production:** `PROVIDER_TEST_MODEL_OVERRIDES.nvidia` probed keys with the model NVIDIA killed today, so a valid key would have been reported dead. It now probes with moonshotai/kimi-k3, then gpt-oss-120b.
31
+ - **Comment/value mismatches from the 09-15 audit:** OpenRouter glm-5.2:free claimed a 32k ctx fix that was never applied to the value (now 32k for real: the free endpoint caps at 32768); SiliconFlow GLM-Z1-9B and Codestral carried false "ctx fixed to..." comments contradicting correct values (comments removed, values untouched).
32
+ - **NVIDIA ctx corrections:** gemma-4-31b-it, diffusiongemma-26b-a4b-it and nemotron-3-nano-omni are 262k (official contextLength 262144), not 256k. Kimi K3 re-tiered to S+ with its 76.8% SWE-bench Verified score.
33
+ - **Pollinations aliases upgraded upstream:** laguna now serves Laguna S 2.1 (was XS.2), deepseek serves DeepSeek V4 Flash (was V3), kimi-code serves Kimi K2.7 Code, qwen-coder serves Qwen3 Coder 30B (re-scored to its own 51.6%), and openai resolves to GPT-5.4 Nano (re-tiered to the nano class). Labels and scores now match what actually answers. Header updated: /v1/chat/completions now requires a free API key (401 anonymous).
34
+ - **OVHcloud embedding models:** bge-m3 and bge-multilingual-gemma2 got their real 8k context.
35
+ - **OrcaRouter:** the in-file comment claimed the fusion family is pay-as-you-go while the model list said otherwise; the comment is corrected and the undocumented fusion models stay out until their free status is confirmed.
36
+ - **Tests:** router failover, endpoint-installer and key-discovery suites pinned dead NVIDIA ids; all swapped to live ids (1154/1154 tests pass).
37
+ - **Docs:** README and docs/providers.md refreshed to 24 providers / 256 models (regenerated from sources.js), website catalog copy synced.
@@ -0,0 +1,20 @@
1
+ # Changelog v0.5.95 - 2026-09-22
2
+
3
+ ### Added
4
+ - ๐ŸŒ Pollinations: Claude Opus 5 (S+, 1M), Claude Sonnet 5 (S, 1M), LongCat 2.0 (A, 1M) and DeepSeek V4.1 Flash (A+, 1M) are now reachable free via Pollen credits. Flagged gemini-3.1-pro-preview as reporting 'down' on the live catalog, re-verify next audit.
5
+ - ๐Ÿ”€ OrcaRouter: the fusion family (fusion S, fusion-mini A+, fusion-flash B+) now reports $0 with real context lengths, so all three are listed.
6
+ - ๐Ÿ‡ซ๐Ÿ‡ท Mistral: Mistral Large 3 (mistral-large-2512, S+, 256k) re-added. The 2026-09-16 removal was stale, the 675B MoE flagship is back in the official catalog. Also added Leanstral 1.5 (B, 256k, Lean 4 theorem proving).
7
+ - โšก NVIDIA NIM: GLM-5.3-Flash (S+, 1M) verified live on the free integrate API.
8
+ - ๐Ÿ‡ซ๐Ÿ‡ท Scaleway: Qwen3.8 27B (A+, 256k) added to the Serverless catalog.
9
+ - ๐Ÿงฉ OpenCode Zen: MiMo-V2.6 Flash Free (S+, 200k) new free promo model on the live Zen endpoint.
10
+ - ๐Ÿงฑ Kilo: Nemotron 3.5 Content Safety (C, 128k) added for breadth, matching OpenRouter/Requesty.
11
+ - ๐Ÿ‰ SiliconFlow: Hunyuan MT 7B (C, 33k) added as the new $0 chat model.
12
+
13
+ ### Removed
14
+ - ๐Ÿชฆ Novita: bunny, Qwen 3.6 Plus, Qwen 3.5 Plus and GLM 4.6 (dev) all gone paid or delisted. Only the 3 Ling 3.0 Flash variants stay free (time-limited).
15
+ - ๐Ÿชฆ SiliconFlow: DeepSeek R1 0528 Qwen3 8B and Qwen3.5 4B dropped off the official free pricing list.
16
+ - ๐Ÿชฆ Mistral: Magistral Medium officially deprecated upstream 2026-05-22, replacement Mistral Medium 3.5 (already listed).
17
+
18
+ ### Changed
19
+ - ๐Ÿ” Full 24-provider re-audit, live-verified 2026-09-22: 247 โ†’ 254 free models. 14 providers confirmed with zero changes (groq, cerebras, googleai, cloudflare, openrouter, sambanova, ovhcloud, codestral, zai, qwen, llm7, routeway, requesty, vercel-gateway, ollama-cloud).
20
+ - โœ… Test updated: NVIDIA static-head expectation now reflects the new catalog order.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "free-coding-models",
3
- "version": "0.5.93",
3
+ "version": "0.5.95",
4
4
  "description": "Find the fastest coding LLM models in seconds โ€” ping free models from multiple providers, pick the best one for OpenCode, Cursor, or any AI coding assistant.",
5
5
  "keywords": [
6
6
  "nvidia",
package/sources.js CHANGED
@@ -48,34 +48,35 @@ export const nvidiaNim = [
48
48
  // Removed (2026-08-23): z-ai/glm-5.2 (GLM 5.1) โ€” no longer in integrate.api.nvidia.com/v1/models (102 models live)
49
49
  // Removed (2026-09-05): moonshotai/kimi-k2.6 (Kimi K2.6) - Model page returns 404 and model is absent from the NVIDIA model catalog; could not verify existence
50
50
  // Removed (2026-08-30): deepseek-ai/deepseek-v4-pro (DeepSeek V4 Pro) โ€” 410 Gone per NVIDIA NIM forum; replaced by deepseek-v4-flash:0731 (forums.developer.nvidia.com/t/deepseek-v4-pro-flash-removed/379558)
51
- ['deepseek-ai/deepseek-v4-flash-0731', 'DeepSeek V4 Flash', 'S+', '79.0%', '1M'], // Fixed (2026-08-13): id 'deepseek-ai/deepseek-v4-flash' โ†’ 'deepseek-ai/deepseek-v4-flash-0731' (NIM /v1/models only exposes the -0731 suffix)
51
+ // Removed (2026-09-21): deepseek-ai/deepseek-v4-flash-0731 (DeepSeek V4 Flash) โ€” NVIDIA deprecation banner on the model page: deprecated 2026-09-19, no longer supported after 2026-09-21; DeepSeek retired V4 Flash in favor of V4.1 Flash (not registered on NIM)
52
+ ['moonshotai/kimi-k3', 'Kimi K3', 'S+', '76.8%', '1M'], // Fixed (2026-09-21): tier 'S' โ†’ 'S+' + sweScore '-' โ†’ '76.8%' (tracker-sourced SWE-bench Verified; 76.8% is S+ on the documented scale)
53
+ ['z-ai/glm-5.3-flash', 'GLM-5.3-Flash', 'S+', '-', '1M'], // Added (2026-09-22) โ€” live on NIM /v1/models (needed a 230s cold start on first probe)
52
54
  // Removed (2026-08-30): stepfun-ai/step-3.7-flash (Step 3.7 Flash) โ€” 410 Gone per NVIDIA NIM TUI ping (no replacement listed; superseded by step-3.7-flash via Routeway `step-3.7-flash:free`)
53
55
  ['nvidia/nemotron-3-ultra-550b-a55b', 'Nemotron 3 Ultra', 'S+', '71.9%', '1M'],
54
56
  ['poolside/laguna-xs-2.1', 'Laguna XS 2.1', 'S+', '70.9%', '262k'], // Added (2026-08-13)
55
57
  ['meta/muse-glimmer-30b', 'Muse Glimmer 30B', 'B+', '-', '128k'], // Added (2026-09-02) โ€” new in NIM catalog
56
- ['deepseek-ai/deepseek-v4-pro-0813', 'DeepSeek V4 Pro', 'S+', '-', '1M'], // Fixed (2026-09-15): ctx '1M' โ†’ '262k'
58
+ // Removed (2026-09-21): deepseek-ai/deepseek-v4-pro-0813 (DeepSeek V4 Pro) โ€” 410 Gone on live probe: end of life 2026-09-14T08:00:00Z; absent from /v1/models
57
59
  // โ”€โ”€ S tier โ€” SWE-bench Verified 60โ€“70% โ”€โ”€
58
60
  // Removed (2026-09-05): openai/gpt-oss-120b (GPT OSS 120B) - NVIDIA deprecation notice on model page: API deprecated on 09/02/2026 and no longer supported
59
61
  // Removed (2026-07-27): meta/llama-4-maverick-17b-128e-instruct (Llama 4 Maverick) โ€” EOL 2026-07-27 (HTTP 410 Gone)
60
62
  // Removed (2026-08-23): mistralai/mistral-medium-3.5-128b (Mistral Medium 3.5) โ€” no longer in integrate.api.nvidia.com/v1/models (still on Mistral LP directly)
61
63
  // Removed (2026-07-27): mistralai/mistral-small-4-119b-2603 (Mistral Small 4) โ€” EOL 2026-07-27 (HTTP 410 Gone)
62
64
  // Removed (2026-09-09): minimaxai/minimax-m3 (MiniMax M3) - 410 Gone per live chat probe: reached end of life 2026-09-09T09:00:00Z (shutdown was announced in-file on 2026-09-08)
63
- ['moonshotai/kimi-k3', 'Kimi K3', 'S', '-', '1M'], // Added (2026-09-02) โ€” new in NIM catalog
64
65
  ['mistralai/mistral-nemotron', 'Mistral Nemotron', 'S', '-', '128k'], // Fixed ID (2026-07-27): nvidia/mistral-nemotron โ†’ mistralai/mistral-nemotron
65
66
  // Removed (2026-07-27): deepseek-ai/deepseek-v3.2 (DeepSeek V3.2) โ€” HTTP 404
66
- ['qwen/qwen3-coder-480b-a35b-instruct', 'Qwen3 Coder 480B', 'S', '-', '262k'], // Added (2026-09-15) โ€” verified via live audit
67
+ // Removed (2026-09-21): qwen/qwen3-coder-480b-a35b-instruct (Qwen3 Coder 480B) โ€” 410 Gone on live probe: end of life 2026-06-11; the 2026-09-15 re-add was erroneous (model was never alive on NIM in September). Free 480B coder is still on DashScope as qwen3-coder-480b-a35b-instruct
67
68
  // โ”€โ”€ A+ tier โ€” SWE-bench Verified 50โ€“60% โ”€โ”€
68
69
  // Removed (2026-07-27): mistralai/mistral-large-3-675b-instruct-2512 (Mistral Large 675B) โ€” EOL 2026-07-23 (HTTP 410 Gone)
69
70
  ['nvidia/nemotron-3-super-120b-a12b', 'Nemotron 3 Super', 'S', '60.5%', '1M'],
70
- ['nvidia/nemotron-3-nano-omni-30b-a3b-reasoning', 'Nemotron 3 Omni', 'A+', '52.0%', '256k'],
71
+ ['nvidia/nemotron-3-nano-omni-30b-a3b-reasoning', 'Nemotron 3 Omni', 'A+', '52.0%', '262k'], // Fixed (2026-09-21): ctx '256k' โ†’ '262k' (official contextLength 262144)
71
72
  // Removed (2026-07-27): meta-llama/llama-4-scout-17b-16e-instruct (Llama 4 Scout) โ€” HTTP 404
72
73
  // Removed (2026-08-30): nvidia/llama-3.3-nemotron-super-49b-v1.5 (Llama 3.3 Nemotron Super 49B) โ€” 410 Gone per NVIDIA NIM TUI ping
74
+ // Removed (2026-09-21): minimaxai/minimax-m3 (MiniMax M3 Preview) โ€” 410 Gone on live probe: end of life 2026-09-09T09:00:00Z (same EOL already documented in-file on 2026-09-09; the 2026-09-15 re-add was erroneous)
73
75
  ['nvidia/nemotron-3.5-lightning-30b-a3b', 'Nemotron 3.5 Lightning 30B', 'A+', '52.8%', '1M'],
74
- ['minimaxai/minimax-m3', 'MiniMax M3 Preview', 'A+', '-', '1M'], // Added (2026-09-15) โ€” verified via live audit
75
76
  // โ”€โ”€ A tier โ€” SWE-bench Verified 40โ€“50% โ”€โ”€
76
77
  // Removed (2026-09-05): nvidia/nemotron-nano-3-30b-a3b (Nemotron Nano 30B) - Model page returns 404 and model is absent from the NVIDIA model catalog; superseded by Nemotron 3.5 Lightning
77
78
  ['openai/gpt-oss-20b', 'GPT OSS 20B', 'A+', '50.3%', '128k'],
78
- ['google/gemma-4-31b-it', 'Gemma 4 31B', 'A+', '52.0%', '256k'],
79
+ ['google/gemma-4-31b-it', 'Gemma 4 31B', 'A+', '52.0%', '262k'], // Fixed (2026-09-21): ctx '256k' โ†’ '262k' (official contextLength 262144)
79
80
  // Removed (2026-08-30): mistralai/mistral-large-2-instruct (Mistral Large 2) โ€” 404 NOT FOUND per NVIDIA NIM TUI ping (model not in NIM catalog; use Mistral LP `mistral-large-2512`)
80
81
  // Removed (2026-07-27): qwen/qwen2.5-coder-32b-instruct (Qwen2.5 Coder 32B) โ€” EOL 2026-05-12 (HTTP 410 Gone)
81
82
  // Removed (2026-07-27): deepseek-ai/deepseek-r1 (DeepSeek R1) โ€” HTTP 404
@@ -86,14 +87,14 @@ export const nvidiaNim = [
86
87
  // Removed (2026-08-30): meta/codellama-70b (CodeLlama 70B) โ€” 404 NOT FOUND per NVIDIA NIM TUI ping (docs.nvidia.com still lists CodeLlama but not via NIM `integrate.api` free tier)
87
88
  // Removed (2026-08-30): mistralai/codestral-22b-instruct-v0.1 (Codestral 22B) โ€” 404 NOT FOUND per NVIDIA NIM TUI ping (use Codestral `codestral-2508` via Mistral LP)
88
89
  // Removed (2026-08-30): ibm/granite-34b-code-instruct (Granite 34B Code) โ€” 404 NOT FOUND per NVIDIA NIM TUI ping
89
- ['minimaxai/minimax-m2.7', 'MiniMax M2.7', 'A', '-', '200k'], // Added (2026-09-15) โ€” verified via live audit
90
+ // Removed (2026-09-21): minimaxai/minimax-m2.7 (MiniMax M2.7) โ€” 410 Gone on live probe: end of life 2026-07-27T00:00:00Z; the 2026-09-15 re-add was erroneous. Still free on SambaNova/Routeway
90
91
  // โ”€โ”€ A- tier โ€” SWE-bench Verified 35โ€“40% โ”€โ”€
91
92
  // Removed (2026-07-27): bytedance/seed-oss-36b-instruct (Seed OSS 36B) โ€” EOL 2026-07-27 (HTTP 410 Gone)
92
93
  // Removed (2026-07-27): stockmark/stockmark-2-100b-instruct (Stockmark 100B) โ€” EOL 2026-07-15 (HTTP 410 Gone)
93
94
  // โ”€โ”€ B+ tier โ€” SWE-bench Verified 30โ€“35% โ”€โ”€
94
95
  // Removed (2026-07-27): mistralai/ministral-14b-instruct-2512 (Ministral 14B) โ€” EOL 2026-07-27 (HTTP 410 Gone)
95
96
  // Removed (2026-08-30): thinkingmachines/inkling (Inkling) โ€” 410 Gone per NVIDIA NIM TUI ping (per Model Deprecation Request 378412)
96
- ['google/diffusiongemma-26b-a4b-it', 'DiffusionGemma 26B', 'B+', '-', '256k'],
97
+ ['google/diffusiongemma-26b-a4b-it', 'DiffusionGemma 26B', 'B+', '-', '262k'], // Fixed (2026-09-21): ctx '256k' โ†’ '262k' (official contextLength 262144)
97
98
  // โ”€โ”€ B tier โ€” SWE-bench Verified 20โ€“30% โ”€โ”€
98
99
  // Removed (2026-09-05): meta/llama-3.2-11b-vision-instruct (Llama 3.2 11B Vision) - Model page on build.nvidia.com has no hosted endpoint at all (no Free Endpoint, no Partner Endpoint, no endpointData payload); docs page remains but the free API endpoint is gone
99
100
  // Removed (2026-08-30): nvidia/nemotron-mini-4b-instruct (Nemotron Mini 4B) โ€” 410 Gone per NVIDIA NIM TUI ping
@@ -106,15 +107,13 @@ export const nvidiaNim = [
106
107
  export const groq = [
107
108
  // Removed (2026-08-13): llama-3.3-70b-versatile (Llama 3.3 70B) โ€” Groq deprecation, shutdown 2026-08-16
108
109
  // Removed (2026-08-13): llama-3.1-8b-instant (Llama 3.1 8B) โ€” Groq deprecation, shutdown 2026-08-16
110
+ // Removed (2026-09-21): llama-3.3-70b-versatile + llama-3.1-8b-instant re-removed โ€” the 2026-09-15 re-add resurrected models Groq had shut down on 2026-08-16 (absent from the live /models list, deprecated 06/17/26 for free and developer tier)
111
+ // Removed (2026-09-21): minimaxai/minimax-m2.7 (MiniMax M2.7) โ€” enterprise-only on Groq (Contact Sales pricing, no developer-plan rate limits; live API returns model_not_found on a developer-tier key); the 2026-09-15 add was erroneous. Still free on SambaNova/Routeway
112
+ // Removed (2026-09-21): groq/compound + groq/compound-mini โ€” on Groq's official deprecation page with shutdown date 2026-09-21
109
113
  ['openai/gpt-oss-120b', 'GPT OSS 120B', 'S', '62.4%', '131k'],
110
114
  ['openai/gpt-oss-20b', 'GPT OSS 20B', 'A+', '60.7%', '131k'],
111
115
  // Removed (2026-09-15): qwen/qwen3.6-27b (Qwen3.6 27B) โ€” rotated out of Groq catalog, superseded by qwen/qwen3.8-27b; replacement: qwen/qwen3.8-27b
112
- ['groq/compound', 'Groq Compound', 'A', '45.0%', '131k'],
113
- ['groq/compound-mini', 'Groq Compound Mini', 'B+', '32.0%', '131k'],
114
116
  ['qwen/qwen3.8-27b', 'Qwen3.8 27B', 'A+', '-', '131k'],
115
- ['llama-3.3-70b-versatile', 'Llama 3.3 70B Versatile', 'B+', '-', '131k'], // Added (2026-09-15) โ€” verified via live audit
116
- ['llama-3.1-8b-instant', 'Llama 3.1 8B Instant', 'C', '-', '131k'], // Added (2026-09-15) โ€” verified via live audit
117
- ['minimaxai/minimax-m2.7', 'MiniMax M2.7', 'S', '-', '196k'], // Added (2026-09-15) โ€” verified via live audit
118
117
  ]
119
118
 
120
119
  // ๐Ÿ“– Cerebras source - https://cloud.cerebras.ai
@@ -127,7 +126,7 @@ export const cerebras = [
127
126
  // โ”€โ”€ A tier โ€” SWE-bench Verified 40โ€“50% โ”€โ”€
128
127
  // Removed (2026-09-05): gemma-4-31b (Gemma 4 31B) - Official deprecation notice dated 2026-09-03: gemma-4-31b is no longer available on Cerebras public endpoints; it remains only on paid Dedicated Endpoints, so it no longer has a free access tier
129
128
  ['qwen-3.8-27b', 'Qwen 3.8 27B', 'A+', '-', '64k'],
130
- ['qwen-3-235b-a22b-instruct-2507', 'Qwen3 235B A22B Instruct 2507', 'A+', '-', '65k'], // Added (2026-09-15) โ€” verified via live audit
129
+ // Removed (2026-09-21): qwen-3-235b-a22b-instruct-2507 (Qwen3 235B A22B Instruct 2507) โ€” deprecated by Cerebras 2026-05-27, no longer on public endpoints; official catalog lists only gpt-oss-120b and qwen-3.8-27b. The 2026-09-15 re-add was erroneous
131
130
  ]
132
131
 
133
132
  // ๐Ÿ“– SambaNova source - https://cloud.sambanova.ai
@@ -165,10 +164,11 @@ export const openrouter = [
165
164
  ['poolside/laguna-s-2.1:free', 'Poolside Laguna S 2.1', 'S+', '-', '262k'],
166
165
  // Removed (2026-09-15): minimax/minimax-m2.7:free (MiniMax M2.7) โ€” no longer free on OpenRouter
167
166
  // Removed (2026-09-15): minimax/minimax-m3:free (MiniMax M3) โ€” no longer free on OpenRouter
168
- ['z-ai/glm-5.2:free', 'GLM-5.2', 'S+', '-', '256k'], // Added (2026-09-02) // Fixed (2026-09-15): ctx '256k' โ†’ '32k'
167
+ ['z-ai/glm-5.2:free', 'GLM-5.2', 'S+', '-', '32k'], // Added (2026-09-02) // Fixed (2026-09-21): value now matches the 2026-09-15 comment: ctx '256k' โ†’ '32k' (free endpoint is context-capped at 32768 per live API)
169
168
  // โ”€โ”€ S tier โ€” SWE-bench Verified 60โ€“70% โ”€โ”€
170
169
  ['cohere/north-mini-code:free', 'North Mini Code', 'S', '-', '256k'],
171
170
  ['nvidia/nemotron-3-super-120b-a12b:free', 'Nemotron 3 Super', 'S', '60.5%', '262k'],
171
+ ['qwen/qwen3.8-27b:free', 'Qwen3.8 27B', 'S', '-', '262k'], // Added (2026-09-21) โ€” new in the live :free catalog
172
172
  // โ”€โ”€ A+ tier โ€” SWE-bench Verified 50โ€“60% โ”€โ”€
173
173
  ['nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free', 'Nemotron 3 Omni', 'A+', '52.0%', '256k'],
174
174
  ['google/gemma-4-31b-it:free', 'Gemma 4 31B', 'A+', '52.0%', '262k'],
@@ -212,17 +212,19 @@ export const githubModels = [
212
212
  export const mistral = [
213
213
  // โ”€โ”€ S+ tier โ€” SWE-bench Verified โ‰ฅ70% โ”€โ”€
214
214
  ['mistral-medium-3-5', 'Mistral Medium 3.5', 'S+', '77.6%', '256k'], // Fixed (2026-09-16): mistral-medium-3-5-26-04 โ†’ mistral-medium-3-5 (live /v1/models, ctx 262144)
215
+ ['mistral-large-2512', 'Mistral Large 3', 'S+', '-', '256k'], // Re-added (2026-09-22) โ€” the 2026-09-16 removal was stale: Mistral Large 3 (675B MoE, GA 2025-12) is back in the official catalog
215
216
  // Removed (2026-08-13): devstral-2512 (Devstral 2) โ€” Mistral deprecation, full retirement 2026-07-31
216
217
  // Removed (2026-09-16): mistral-large-3-25-12 (Mistral Large 3) โ€” no `large` model exists in the live catalog at all
217
218
  // Removed (2026-09-16): zai-glm-5-2 (Z.ai GLM 5.2) โ€” absent from /v1/models; a direct call returns 403 tier_not_allowed (paid tier only), so it never belonged in a free catalog
218
219
  // โ”€โ”€ A+ tier โ”€โ”€
219
- ['magistral-medium-latest', 'Magistral Medium', 'A+', '-', '256k'], // Fixed (2026-09-16): magistral-medium-1-2-25-09 โ†’ magistral-medium-latest (only the -latest alias exists upstream)
220
+ // Removed (2026-09-22): magistral-medium-latest (Magistral Medium) โ€” officially deprecated upstream 2026-05-22, "Use Mistral Medium 3.5"; replacement: mistral-medium-3-5
220
221
  // โ”€โ”€ A tier โ€” SWE-bench Verified 40โ€“50% โ”€โ”€
221
222
  ['mistral-small-2603', 'Mistral Small 4', 'A', '48.0%', '256k'], // Fixed (2026-09-16): mistral-small-4-0-26-03 โ†’ mistral-small-2603 (live /v1/models, ctx 262144)
222
223
  // โ”€โ”€ B+ tier โ€” SWE-bench Verified 30โ€“35% โ”€โ”€
223
224
  ['ministral-14b-2512', 'Ministral 3 14B', 'B+', '-', '256k'], // Fixed (2026-09-16): ministral-3-14b-25-12 โ†’ ministral-14b-2512 (live /v1/models, ctx 262144)
224
225
  // โ”€โ”€ B tier โ€” SWE-bench Verified 20โ€“30% โ”€โ”€
225
226
  ['ministral-8b-2512', 'Ministral 3 8B', 'B', '-', '256k'], // Fixed (2026-09-16): ministral-3-8b-25-12 โ†’ ministral-8b-2512 (live /v1/models, ctx 262144)
227
+ ['labs-leanstral-1-5', 'Leanstral 1.5', 'B', '-', '256k'], // Added (2026-09-22) โ€” free-listed on Mistral LP; targets Lean 4 theorem proving, niche coding use
226
228
  ['ministral-3b-2512', 'Ministral 3 3B', 'B', '-', '128k'], // Fixed (2026-09-16): ministral-3-3b-25-12 โ†’ ministral-3b-2512; ctx 256k โ†’ 128k (max_context_length 131072)
227
229
  // Removed (2026-09-16): mistral-small-creative-25-12 (Mistral Small Creative) โ€” absent from the live catalog
228
230
  ]
@@ -232,7 +234,7 @@ export const mistral = [
232
234
  // ๐Ÿ“– API keys now use the Mistral platform key format; CODESTRAL_API_KEY remains supported as an alias.
233
235
  export const codestral = [
234
236
  // โ”€โ”€ A tier โ€” SWE-bench Verified 40โ€“50% โ”€โ”€
235
- ['codestral-2508', 'Codestral', 'A', '40.0%', '256k'], // Fixed (2026-07-27): ctx '256k' โ†’ '128k' per official Mistral model card
237
+ ['codestral-2508', 'Codestral', 'A', '40.0%', '256k'], // Fixed (2026-09-21): deleted the false 2026-07-27 "ctx to 128k" comment; 256k is the correct value per the official model card
236
238
  // Removed (2026-08-23): codestral-2501 (Codestral 2501), codestral-2405 (Codestral 2405) โ€” retired from Mistral API; only codestral-2508 / codestral-latest remain
237
239
  // Removed (2026-08-13): codestral-2 (Codestral 2) โ€” fabricated ID, never existed in Mistral catalog (Mistral uses date-stamped versioning)
238
240
  ]
@@ -256,6 +258,7 @@ export const scaleway = [
256
258
  ['gemma-4-26b-a4b-it', 'Gemma 4 26B MoE', 'A+', '-', '256k'],
257
259
  // Removed (2026-09-02): gemma-4-31b-it (Gemma 4 31B IT) โ€” Dedicated tier only, not available on Serverless
258
260
  ['qwen3-235b-a22b-instruct-2507', 'Qwen3 235B', 'A', '45.2%', '250k'], // Restored (2026-09-05) โ€” still Serverless per official docs (silently dropped by PR #178)
261
+ ['qwen3.8-27b', 'Qwen3.8 27B', 'A+', '-', '256k'], // Added (2026-09-22) โ€” new Serverless model on the official catalog (agentic/coding optimized)
259
262
  // โ”€โ”€ A- tier โ€” SWE-bench Verified 35โ€“40% โ”€โ”€
260
263
  ['llama-3.3-70b-instruct', 'Llama 3.3 70B', 'B', '22.0%', '100k'], // Fixed (2026-08-13): ctx '128k' โ†’ '100k' (Serverless tier per official catalog)
261
264
  // โ”€โ”€ B+ tier โ€” SWE-bench Verified 30โ€“35% โ”€โ”€
@@ -280,8 +283,10 @@ export const googleai = [
280
283
  ['gemini-3-flash-preview', 'Gemini 3 Flash Preview', 'S+', '78.0%', '1M'], // Restored (2026-09-05) โ€” free tier confirmed per official pricing page
281
284
  ['gemini-2.5-pro', 'Gemini 2.5 Pro', 'S', '63.8%', '1M'], // Restored (2026-09-05) โ€” free tier confirmed per official pricing page
282
285
  // Removed (2026-09-02): gemini-3.1-pro-preview (Gemini 3.1 Pro Preview) โ€” free tier "Not available" per official pricing page (rechecked 2026-09-05)
286
+ // Removed (2026-09-21): gemini-3.1-pro-preview re-removed โ€” the 2026-09-15 re-add resurrected a paid-only model (free tier "Not available" on the official pricing page since ~April 2026); best free alternative: gemini-3.5-flash
283
287
  // Removed (2026-09-05): gemini-2.0-flash โ€” not listed on the official pricing page (PR #178 addition reverted)
284
- ['gemini-3.1-pro-preview', 'Gemini 3.1 Pro Preview', 'S+', '-', '1M'], // Added (2026-09-15) โ€” verified via live audit
288
+ // โš ๏ธ Gemini 2.5 family retires no earlier than 2026-10-16 per Google deprecation policy
289
+ // Removed (2026-09-21): gemma-4-31b-it + gemma-4-26b-a4b-it (Gemma 4 entries) โ€” Gemma pages are gone from ai.google.dev (404 on /gemini-api/docs/models/gemma), Gemma is no longer served via the Gemini API free tier; the free Gemma route is now NVIDIA NIM
285
290
  ]
286
291
 
287
292
  // ๐Ÿ“– ZAI source - https://open.z.ai
@@ -291,9 +296,8 @@ export const googleai = [
291
296
  export const zai = [
292
297
  // โ”€โ”€ S+ tier โ€” SWE-bench Verified โ‰ฅ70% โ”€โ”€
293
298
  ['zai/glm-5.3-flash', 'GLM-5.3-Flash', 'S+', '-', '1M'], // Added (2026-09-02)
294
- ['zai/glm-5.2', 'GLM-5.2', 'S+', '-', '1M'], // Added (2026-08-13)
295
299
  ['zai/glm-5.3', 'GLM-5.3', 'S+', '-', '1M'],
296
- ['zai/glm-5.1', 'GLM-5.1', 'S+', '-', '200k'], // Added (2026-09-15) โ€” verified via live audit
300
+ // Removed (2026-09-21): zai/glm-5.2 + zai/glm-5.1 โ€” Coding Plan requests for both are now silently redirected to GLM-5.3 (official docs.z.ai plan update), so the ids no longer serve a distinct free model; both remain paid-API models
297
301
  ['zai/glm-5', 'GLM-5', 'S+', '-', '200k'], // Added (2026-09-15) โ€” verified via live audit
298
302
  // โ”€โ”€ S tier โ€” SWE-bench Verified 60โ€“70% โ”€โ”€
299
303
  ['zai/glm-4.7-flash', 'GLM-4.7-Flash', 'A+', '59.2%', '200k'], // Fixed (2026-07-27): ctx '203k' โ†’ '200k' per official docs
@@ -329,6 +333,7 @@ export const qwen = [
329
333
  ['qwen3-coder-plus', 'Qwen3 Coder Plus', 'S', '69.6%', '1M'],
330
334
  ['qwen3-coder-next', 'Qwen3 Coder Next', 'S+', '70.6%', '256k'],
331
335
  // Removed (2026-09-15): qwen3-coder-480b-a35b-instruct (Qwen3 Coder 480B) โ€” legacy, superseded by qwen3-coder-next; replacement: qwen3-coder-next
336
+ ['qwen3-coder-480b-a35b-instruct', 'Qwen3 Coder 480B', 'S', '69.6%', '256k'], // Re-added (2026-09-21) โ€” back on the official free billing page (1M-token free quota, updated 2026-09-20); free tier did not remove it after all
332
337
  ['qwen3.8-27b', 'Qwen3.8 27B', 'S', '-', '1M'],
333
338
  // โ”€โ”€ A+ tier โ€” SWE-bench Verified 50โ€“60% โ”€โ”€
334
339
  ['qwen3.7-flash', 'Qwen3.7 Flash', 'A+', '-', '1M'], // Added (2026-07-27)
@@ -336,6 +341,8 @@ export const qwen = [
336
341
  ['qwen3.5-flash', 'Qwen3.5 Flash', 'S', '64.4%', '1M'],
337
342
  ['qwen3-coder-flash', 'Qwen3 Coder Flash', 'A+', '55.0%', '1M'],
338
343
  ['qwen3-vl-flash', 'Qwen3 VL Flash', 'A+', '-', '256k'], // Added (2026-08-13)
344
+ ['qwen3-vl-plus', 'Qwen3 VL Plus', 'A+', '-', '256k'], // Added (2026-09-21) โ€” on the official free billing page (1M-token free quota)
345
+ ['qwen3-coder-30b-a3b-instruct', 'Qwen3 Coder 30B A3B', 'A+', '-', '256k'], // Added (2026-09-21) โ€” on the official free billing page (1M-token free quota)
339
346
  // Removed (2026-09-15): qwen3-32b (Qwen3 32B) โ€” legacy, Oct 10 2026 shutdown (aliyun notice 118434); replacement: qwen3.8-27b
340
347
  ['qwen3.5-397b-a17b', 'Qwen3.5 397B A17B', 'S+', '76.2%', '256k'],
341
348
  ['qwen3.5-122b-a10b', 'Qwen3.5 122B A10B', 'S+', '72.0%', '256k'],
@@ -348,6 +355,7 @@ export const qwen = [
348
355
  ['qwen3.5-27b', 'Qwen3.5 27B', 'S+', '72.4%', '256k'],
349
356
  // Removed (2026-09-15): qwen3-30b-a3b (Qwen3 30B A3B) โ€” legacy, Oct 10 2026 shutdown; replacement: qwen3.5-35b-a3b
350
357
  ['qwen3.5-omni-plus', 'Qwen3.5 Omni Plus', 'B+', '-', '32k'], // Added (2026-09-15) โ€” verified via live audit
358
+ ['qwen3.8-omni-flash', 'Qwen3.8 Omni Flash', 'B+', '-', '32k'], // Added (2026-09-21) โ€” on the official free billing page; ctx follows omni-family precedent (32k)
351
359
  ['qwen3.6-27b', 'Qwen3.6 27B', 'B+', '-', '256k'], // Added (2026-09-15) โ€” verified via live audit
352
360
  ['qwen3.6-35b-a3b', 'Qwen3.6 35B A3B', 'B+', '-', '256k'], // Added (2026-09-15) โ€” verified via live audit
353
361
  ]
@@ -361,16 +369,11 @@ export const cloudflare = [
361
369
  // Removed (2026-09-05): @cf/moonshotai/kimi-k2.6 (Kimi K2.6) - model still exists but docs state it is not available through standard Workers Free billing; requires Workers Paid plan or prepaid AI Gateway credits, so unusable within the free 10k neurons/day tier
362
370
  // Removed (2026-09-05): @cf/moonshotai/kimi-k2.7-code (Kimi K2.7 Code) - model still exists but docs state it is not available through standard Workers Free billing; requires Workers Paid plan or prepaid AI Gateway credits
363
371
  // Removed (2026-09-05): @cf/zai-org/glm-5.2 (GLM-5.2) - model still exists but docs state it is not available through standard Workers Free billing; requires Workers Paid plan or prepaid AI Gateway credits
364
- ['@cf/zai-org/glm-5.3-flash', 'GLM-5.3-Flash', 'S+', '-', '1.3M'], // Added (2026-09-15) โ€” verified via live audit
365
- ['@cf/zai-org/glm-5.3', 'GLM-5.3', 'S+', '-', '1.3M'], // Added (2026-09-15) โ€” verified via live audit
372
+ // Removed (2026-09-21): @cf/zai-org/glm-5.3-flash + @cf/zai-org/glm-5.3 โ€” both now carry the "Paid access required: not available through standard Workers Free billing" badge on the official docs pages; the 2026-09-15 re-add was erroneous
366
373
  // โ”€โ”€ S tier โ€” SWE-bench Verified 60โ€“70% โ”€โ”€
367
374
  ['@cf/zai-org/glm-4.7-flash', 'GLM-4.7-Flash', 'A+', '59.2%', '131k'],
368
375
  ['@cf/openai/gpt-oss-120b', 'GPT OSS 120B', 'S', '62.4%', '128k'],
369
- ['@cf/zai-org/glm-5.2', 'GLM-5.2', 'S', '-', '262k'], // Added (2026-09-15) โ€” verified via live audit
370
- ['@cf/deepseek-ai/deepseek-v4-pro-0813', 'DeepSeek V4 Pro', 'S', '-', '1M'], // Added (2026-09-15) โ€” verified via live audit
371
- ['@cf/deepseek-ai/deepseek-v4-flash-0731', 'DeepSeek V4 Flash', 'S', '-', '1.3M'], // Added (2026-09-15) โ€” verified via live audit
372
- ['@cf/moonshotai/kimi-k2.7-code', 'Kimi K2.7 Code', 'S', '-', '262k'], // Added (2026-09-15) โ€” verified via live audit
373
- ['@cf/moonshotai/kimi-k2.6', 'Kimi K2.6', 'S', '-', '262k'], // Added (2026-09-15) โ€” verified via live audit
376
+ // Removed (2026-09-21): @cf/zai-org/glm-5.2, @cf/deepseek-ai/deepseek-v4-pro-0813, @cf/deepseek-ai/deepseek-v4-flash-0731, @cf/moonshotai/kimi-k2.7-code, @cf/moonshotai/kimi-k2.6 re-removed โ€” all five carry the "Paid access required: not available through standard Workers Free billing" badge on their official docs pages; the 2026-09-15 re-adds regressed the 2026-09-05 paid-only policy
374
377
  // โ”€โ”€ A+ tier โ€” SWE-bench Verified 50โ€“60% โ”€โ”€
375
378
  ['@cf/nvidia/nemotron-3-120b-a12b', 'Nemotron 3 Super', 'S', '60.5%', '256k'],
376
379
  // โ”€โ”€ A tier โ€” SWE-bench Verified 40โ€“50% โ”€โ”€
@@ -387,6 +390,7 @@ export const cloudflare = [
387
390
  // โ”€โ”€ B+ tier โ€” SWE-bench Verified 30โ€“35% โ”€โ”€
388
391
  ['@cf/mistralai/mistral-small-3.1-24b-instruct', 'Mistral Small 3.1', 'B+', '30.0%', '128k'],
389
392
  ['@cf/ibm-granite/granite-4.0-h-micro', 'Granite 4.0 Micro', 'B+', '30.0%', '131k'], // Fixed (2026-07-27): namespace 'ibm' โ†’ 'ibm-granite'
393
+ ['@cf/aisingapore/gemma-sea-lion-v4-27b-it', 'Gemma SEA-LION V4 27B', 'B+', '-', '128k'], // Added (2026-09-21) โ€” new in the free catalog; SEA-language focused, secondary for coding
390
394
  // โ”€โ”€ B tier โ€” SWE-bench Verified 20โ€“30% โ”€โ”€
391
395
  // Removed (2026-09-15): @cf/meta/llama-3.1-8b-instruct-fast (Llama 3.1 8B Instruct (Fast)) โ€” delisted; llama-3.1-8b-instruct-fp8 (32k ctx) remains; replacement: @cf/meta/llama-3.1-8b-instruct-fp8
392
396
  // Removed (2026-08-30): @cf/google/gemma-3-12b-it (Gemma 3 12B IT) โ€” Deprecated 2026-05-30 per Cloudflare Workers AI docs (developers.cloudflare.com/workers-ai/models/gemma-3-12b-it)
@@ -401,6 +405,7 @@ export const ovhcloud = [
401
405
  ['Qwen3.5-397B-A17B', 'Qwen3.5 397B MoE', 'S+', '76.2%', '262k'],
402
406
  ['Qwen3.6-27B', 'Qwen3.6 27B', 'S+', '77.2%', '262k'],
403
407
  // Removed (2026-07-27): Qwen3-Coder-30B-A3B-Instruct (Qwen3 Coder 30B MoE) โ€” no longer in catalog
408
+ // Removed (2026-09-21): Qwen3-Coder-30B-A3B-Instruct (Qwen3 Coder 30B A3B) โ€” absent from the official AI Endpoints catalog page (20 models, no coder model); the 2026-09-21 morning re-add was erroneous
404
409
  ['gpt-oss-120b', 'GPT OSS 120B', 'S', '62.4%', '131k'],
405
410
  ['gpt-oss-20b', 'GPT OSS 20B', 'A+', '50.3%', '131k'],
406
411
  ['Meta-Llama-3_3-70B-Instruct', 'Llama 3.3 70B', 'B', '22.0%', '131k'],
@@ -408,12 +413,13 @@ export const ovhcloud = [
408
413
  // Removed (2026-08-13): Mistral-Small-3.2-24B-Instruct-2506 (Mistral Small 3.2) โ€” no longer in OVHcloud public catalog (endpoint still reachable but not listed)
409
414
  // Removed (2026-07-27): Mistral-7B-Instruct-v0.3 (Mistral 7B Instruct) โ€” no longer in catalog
410
415
  // Removed (2026-08-13): Mistral-Nemo-Instruct-2407 (Mistral Nemo) โ€” no longer in OVHcloud public catalog
416
+ // Removed (2026-09-21): Mistral-Small-3.2-24B-Instruct-2506 + Mistral-Nemo-Instruct-2407 + Mistral-7B-Instruct-v0.3 re-removed โ€” none of the three Mistral models appear on the official AI Endpoints catalog page; the 2026-09-21 morning re-adds were erroneous
411
417
  ['Qwen3.5-9B', 'Qwen3.5 9B', 'B+', '30.0%', '262k'],
412
418
  ['Qwen2.5-VL-72B-Instruct', 'Qwen2.5-VL 72B', 'S', '-', '32k'], // Added (2026-08-13)
413
419
  // โ”€โ”€ Embeddings โ”€โ”€
414
420
  ['Qwen3-Embedding-8B', 'Qwen3 Embedding 8B', 'B', '-', '32k'], // Fixed (2026-07-27): ctx '-' โ†’ '32k'
415
- ['bge-m3', 'BGE M3', 'B', '-', '-'],
416
- ['bge-multilingual-gemma2', 'BGE Multilingual Gemma2','B','-', '-'],
421
+ ['bge-m3', 'BGE M3', 'B', '-', '8k'], // Fixed (2026-09-21): ctx '-' โ†’ '8k' (embedding model, 8192 tokens)
422
+ ['bge-multilingual-gemma2', 'BGE Multilingual Gemma2','B','-', '8k'], // Fixed (2026-09-21): ctx '-' โ†’ '8k' (embedding model, 8192 tokens)
417
423
  // Fix (2026-05-26): Qwen3.5-9B ctx 128kโ†’262k, Mistral-Small ctx 131kโ†’128k, Mistral-Nemo ctx 128kโ†’118k, Mistral-7B ctx 32kโ†’127k
418
424
  ['Qwen3Guard-Gen-8B', 'Qwen3Guard Gen 8B (moderation, beta)', 'C', '-', '32k'],
419
425
  ['Qwen3Guard-Gen-0.6B', 'Qwen3Guard Gen 0.6B (moderation, beta)', 'C', '-', '32k'],
@@ -430,7 +436,9 @@ export const ovhcloud = [
430
436
  export const opencodeZen = [
431
437
  ['big-pickle', 'Big Pickle', 'S+', '72.0%', '200k'],
432
438
  // Removed (2026-09-05): deepseek-v4-flash-free (DeepSeek V4 Flash Free) - deprecated: marked status=deprecated in the models.dev registry (2026-09-05) and dropped from the docs free-models pricing table; free promo ended
439
+ ['deepseek-v4-flash-free', 'DeepSeek V4 Flash Free', 'S+', '79.0%', '200k'], // Re-added (2026-09-21) โ€” free again per the live Zen /v1/models list and models.dev ($0 pricing); still absent from the docs pricing table so re-verify at next audit
433
440
  ['mimo-v2.5-free', 'MiMo-V2.5 Free', 'S+', '-', '200k'],
441
+ ['mimo-v2.6-flash-free', 'MiMo-V2.6 Flash Free', 'S+', '-', '200k'], // Added (2026-09-22) โ€” new free promo model on the live Zen /v1/models list
434
442
  ['nemotron-3-ultra-free', 'Nemotron 3 Ultra Free', 'S+', '71.9%', '1M'],
435
443
  // Removed (2026-09-05): hy3-free (Tencent Hy3 Free) โ€” absent from live /v1/models (66 models checked)
436
444
  ['nemotron-3.5-lightning-free', 'Nemotron 3.5 Lightning Free', 'S+', '-', '262k'], // Added (2026-08-13)
@@ -438,6 +446,7 @@ export const opencodeZen = [
438
446
  ['ling-3.0-flash-fin-free', 'Ling 3.0 Flash Fin Free', 'B+', '-', '262k'], // Added (2026-09-05) โ€” new id in live /v1/models (was ling-3.0-flash-free)
439
447
  ['muse-spark-1.2-contributor-free', 'Muse Spark 1.2 Contributor Free', 'A+', '-', '1M'],
440
448
  ['muse-spark-1.3-contributor-free', 'Muse Spark 1.3 Contributor Free', 'S+', '-', '1M'],
449
+ ['jev-1.13-free', 'Jev 1.13 Free', 'B+', '-', '200k'], // Added (2026-09-21) โ€” new free model on the live Zen /v1/models list (74 models checked)
441
450
  ]
442
451
 
443
452
  // ๐Ÿ“– Kilo source - https://api.kilo.ai/api/gateway
@@ -466,6 +475,9 @@ export const kilo = [
466
475
  ['inclusionai/ling-3.0-flash-vl:free', 'Ling 3.0 Flash VL (free)', 'B+', '-', '262k'], // Added (2026-09-15) โ€” verified via live audit
467
476
  ['nex-agi/nex-n2.5-mini:free', 'Nex AGI Nex-N2.5-Mini (free)', 'B+', '-', '262k'], // Added (2026-09-15) โ€” verified via live audit
468
477
  ['nex-agi/nex-n2.5-pro:free', 'Nex AGI Nex-N2.5-Pro (free)', 'A', '-', '262k'], // Added (2026-09-15) โ€” verified via live audit
478
+ ['nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free', 'NVIDIA Nemotron 3 Nano Omni (free)', 'A+', '-', '262k'], // Added (2026-09-21) โ€” new in the live free gateway list
479
+ ['qwen/qwen3.8-27b:free', 'Qwen3.8 27B (free)', 'S', '-', '262k'], // Added (2026-09-21) โ€” new in the live free gateway list
480
+ ['nvidia/nemotron-3.5-content-safety:free', 'NVIDIA Nemotron 3.5 Content Safety (free)', 'C', '-', '128k'], // Added (2026-09-22) โ€” content-safety classifier, marginal but kept for breadth (matches OpenRouter/Requesty)
469
481
  ]
470
482
 
471
483
  // ๐Ÿ“– LLM7 source - https://api.llm7.io/v1
@@ -477,6 +489,7 @@ export const llm7 = [
477
489
  // Removed (2026-09-05): glm-5.3, glm-5.3-flash, gemini-3.5-flash-low, gpt-5.4, gpt-5.4-mini, gpt-5.5, gpt-5.6-sol, grok-4.5, grok-4.6 โ€” tier=pro usage_based_only (paid) or nonexistent on /v1/models (PR #178 additions reverted)
478
490
  // โ”€โ”€ S+ tier โ€” SWE-bench Verified โ‰ฅ70% โ”€โ”€
479
491
  ['minimax-m2.7', 'MiniMax M2.7', 'S+', '78.0%', '180k'],
492
+ ['GLM-5.3-Flash', 'GLM-5.3 Flash', 'S+', '-', '400k'], // Re-added (2026-09-21) โ€” returned to the free tier (turbo, usage_based_only:false), verified via live unauthenticated chat probe; ctx 410k per /v1/models
480
493
  // โ”€โ”€ A+ tier โ€” SWE-bench Verified 50โ€“60% โ”€โ”€
481
494
  // Removed (2026-09-05): gemini-3.1-flash-lite (Gemini 3.1 Flash Lite) โ€” now tier=pro usage_based_only (paid) per live /v1/models
482
495
  // Removed (2026-09-15): gpt-oss (GPT OSS 20B) โ€” removed from LLM7 API catalog
@@ -523,31 +536,50 @@ export const novita = [
523
536
  // Removed (2026-07-27): qwen/qwen3.5-plus (Qwen3.5 Plus) โ€” no longer in novita catalog
524
537
  ['inclusionai/ling-3.0-flash-fin', 'Ling 3.0 Flash Fin', 'B+', '-', '256k'],
525
538
  ['inclusionai/ling-3.0-flash-sante', 'Ling 3.0 Flash Sante', 'B+', '-', '256k'],
526
- ['bunny', 'Bunny (free tier)', 'C', '-', '256k'], // Added (2026-09-15) โ€” verified via live audit
527
- ['qwen/qwen3.6-plus', 'Qwen 3.6 Plus (free tier)', 'A', '-', '1M'], // Added (2026-09-15) โ€” verified via live audit
528
- ['qwen/qwen3.5-plus', 'Qwen 3.5 Plus (free tier)', 'A-', '-', '1M'], // Added (2026-09-15) โ€” verified via live audit
529
- ['dev/glm46', 'GLM 4.6 (dev, free)', 'A-', '-', '256k'], // Added (2026-09-15) โ€” verified via live audit
539
+ // Removed (2026-09-22): bunny (Bunny (free tier)) โ€” no longer on the Novita pricing/catalog page, only paid variants remain
540
+ // Removed (2026-09-22): qwen/qwen3.6-plus (Qwen 3.6 Plus (free tier)) โ€” no longer in the Novita catalog, only paid Qwen3.6 variants remain
541
+ // Removed (2026-09-22): qwen/qwen3.5-plus (Qwen 3.5 Plus (free tier)) โ€” no longer in the Novita catalog, Qwen3.5 now paid-only size variants
542
+ // Removed (2026-09-22): dev/glm46 (GLM 4.6 (dev, free)) โ€” GLM 4.6 is now paid ($0.55/$2.20 per M), no free GLM tier left; cheapest GLM: glm-5.3-flash (paid)
530
543
  ['inclusionai/ling-3.0-flash-vl', 'Ling 3.0 Flash VL', 'B+', '-', '256k'], // Added (2026-09-15) โ€” verified via live audit
531
544
  ]
532
545
 
533
546
  // ๐Ÿ“– Pollinations AI source - https://gen.pollinations.ai
534
547
  // ๐Ÿ“– OpenAI-compatible endpoint: https://gen.pollinations.ai/v1/chat/completions
535
- // ๐Ÿ“– Free tier: anonymous without key or free API key from https://enter.pollinations.ai
536
- // ๐Ÿ“– Daily Pollen grants per tier (seed/flower/nectar) โ€” free models cost Pollen but grants renew daily; anonymous tier has rate limits.
537
- // ๐Ÿ“– Verified live 2026-08-23 via GET /v1/models (319 models); IDs below are live and coding-relevant.
548
+ // ๐Ÿ“– Free tier: free API key from https://enter.pollinations.ai (Pollen credit system with free daily grants).
549
+ // ๐Ÿ“– Since 2026-09 the /v1/chat/completions endpoint requires a free API key (401 without one); the legacy
550
+ // ๐Ÿ“– anonymous path only reaches the default model via GET /text. Daily Pollen grants per tier renew free.
551
+ // ๐Ÿ“– Verified live 2026-09-21 via GET /v1/models (411 models): the old short ids (openai, deepseek, kimi,
552
+ // ๐Ÿ“– laguna...) are no longer primary ids but still resolve as aliases of the canonical namespaced models.
553
+ // ๐Ÿ“– Note (2026-09-21): the anonymous tier (text.pollinations.ai/models) now lists ONLY openai-fast;
554
+ // ๐Ÿ“– the models below need the free API key + Pollen credits (gen.pollinations.ai). Re-check the Pollen
555
+ // ๐Ÿ“– free-grant policy at next audit: if grants stop covering these models, this list must shrink to openai-fast.
538
556
  export const pollinations = [
539
557
  // โ”€โ”€ S+ tier โ€” SWE-bench Verified โ‰ฅ70% โ”€โ”€
540
- ['laguna', 'Laguna XS.2', 'S+', '70.9%', '1M'],
558
+ ['laguna', 'Laguna S 2.1', 'S+', '-', '1M'], // Fixed (2026-09-21): alias now resolves to poolside/laguna-s-2.1 (Laguna S 2.1), was Laguna XS.2; score cleared (S 2.1 has no published SWE-bench Verified)
541
559
  ['minimax-m2.7', 'MiniMax M2.7', 'S+', '78.0%', '200k'],
542
560
  ['glm-5.3', 'Z.ai GLM-5.3', 'S+', '-', '1M'],
543
561
  ['kimi', 'Moonshot Kimi K2.6', 'S+', '80.2%', '262k'],
544
562
  ['minimax', 'MiniMax M3', 'S+', '80.5%', '524k'],
563
+ ['moonshotai/kimi-k3', 'Moonshot Kimi K3', 'S+', '76.8%', '1M'], // Added (2026-09-21) โ€” canonical id, healthy on live /v1/models; score follows the Kimi K3 entry on NVIDIA
564
+ ['deepseek/deepseek-v4-pro', 'DeepSeek V4 Pro', 'S+', '-', '1M'], // Added (2026-09-21) โ€” canonical id, healthy on live /v1/models
565
+ ['qwen/qwen3.8-max', 'Qwen3.8 Max', 'S+', '-', '1M'], // Added (2026-09-21) โ€” canonical id, healthy on live /v1/models
566
+ ['google/gemini-3.1-pro-preview', 'Gemini 3.1 Pro Preview', 'S+', '-', '1M'], // Added (2026-09-21) โ€” canonical id (paid-only on Google AI Studio but free here) // โš ๏ธ status 'down' on live /v1/models 2026-09-22, re-verify next audit
567
+ ['openai/gpt-5.5', 'OpenAI GPT-5.5', 'S+', '-', '1M'], // Added (2026-09-21) โ€” canonical id, healthy on live /v1/models
568
+ ['openai/gpt-6-astra', 'OpenAI GPT-6 Astra', 'S+', '-', '1M'], // Added (2026-09-21) โ€” canonical id, healthy on live /v1/models
569
+ ['z-ai/glm-5.3-flash', 'Z.ai GLM-5.3 Flash', 'S+', '-', '1M'], // Added (2026-09-21) โ€” canonical id, healthy on live /v1/models
570
+ ['nvidia/nemotron-3-ultra', 'NVIDIA Nemotron 3 Ultra', 'S+', '71.9%', '262k'], // Added (2026-09-21) โ€” canonical id, healthy on live /v1/models; score/scale from the NVIDIA entry
571
+ ['anthropic/claude-opus-5', 'Claude Opus 5', 'S+', '-', '1M'], // Added (2026-09-22) โ€” healthy on live /v1/models, free via Pollen credits
545
572
  // โ”€โ”€ S tier โ€” SWE-bench Verified 60โ€“70% โ”€โ”€
546
- ['qwen-coder', 'Qwen3 Coder', 'S', '69.6%', '262k'],
547
- ['deepseek', 'DeepSeek V3', 'S', '66.0%', '1M'],
548
- ['kimi-code', 'Kimi K2 Code', 'S', '60.4%', '262k'],
549
- ['openai', 'OpenAI GPT', 'S', '62.4%', '400k'],
573
+ ['anthropic/claude-sonnet-5', 'Claude Sonnet 5', 'S', '-', '1M'], // Added (2026-09-22) โ€” healthy on live /v1/models, free via Pollen credits
574
+ ['qwen-coder', 'Qwen3 Coder 30B', 'A+', '51.6%', '262k'], // Fixed (2026-09-21): alias now resolves to qwen/qwen3-coder-30b-a3b-instruct; re-scored from the 480B figure to the 30B SWE-bench Verified
575
+ ['deepseek', 'DeepSeek V4 Flash', 'S+', '79.0%', '1M'], // Fixed (2026-09-21): alias now resolves to deepseek/deepseek-v4-flash (V4 Flash 0731), was V3; re-scored per the V4 Flash family entry
576
+ ['kimi-code', 'Kimi K2.7 Code', 'S', '60.4%', '262k'], // Fixed (2026-09-21): alias now resolves to moonshotai/kimi-k2.7-code, was K2 Code
577
+ ['openai', 'OpenAI GPT-5.4 Nano', 'B+', '-', '400k'], // Fixed (2026-09-21): alias now resolves to openai/gpt-5.4-nano (was a generic GPT alias); re-tiered to the nano class
578
+ ['qwen/qwen3-coder-next', 'Qwen3 Coder Next', 'S+', '70.6%', '262k'], // Added (2026-09-21) โ€” canonical id (new on the network, health still warming up); score from the DashScope entry
579
+ ['openai/gpt-5.6-luna', 'OpenAI GPT-5.6 Luna', 'S', '-', '1M'], // Added (2026-09-21) โ€” canonical id, healthy on live /v1/models
550
580
  // โ”€โ”€ A+ tier โ€” SWE-bench Verified 50โ€“60% โ”€โ”€
581
+ ['deepseek/deepseek-v4.1-flash', 'DeepSeek V4.1 Flash', 'A+', '-', '1M'], // Added (2026-09-22) โ€” healthy on live /v1/models, successor to V4 Flash
582
+ ['meituan/longcat-2.0', 'LongCat 2.0', 'A', '-', '1M'], // Added (2026-09-22) โ€” healthy on live /v1/models, new MoE agentic/coding model on the network
551
583
  ['gemma-4-31b', 'Gemma 4 31B', 'A+', '52.0%', '262k'],
552
584
  ['gpt-oss', 'GPT OSS 20B', 'A+', '50.3%', '131k'],
553
585
  ['qwen3.7-flash', 'Qwen3.7 Flash', 'A+', '-', '1M'],
@@ -563,15 +595,15 @@ export const pollinations = [
563
595
  // ๐Ÿ“– and still reachable with free-tier rate limits (1000 RPM). Keep only the chat text models here.
564
596
  export const siliconflow = [
565
597
  // โ”€โ”€ A tier โ€” SWE-bench Verified 40โ€“50% โ”€โ”€
566
- ['THUDM/GLM-Z1-9B-0414', 'GLM-Z1 9B', 'A', '-', '131k'], // Fixed (2026-09-15): ctx '131k' โ†’ '32k'
567
- ['deepseek-ai/DeepSeek-R1-0528-Qwen3-8B', 'DeepSeek R1 0528 Qwen3 8B', 'A', '-', '131k'],
598
+ // Removed (2026-09-21): THUDM/GLM-Z1-9B-0414 + THUDM/GLM-4-9B-0414 โ€” deprecated 2026-03-12 per official release notes, service terminated
599
+ // Removed (2026-09-22): deepseek-ai/DeepSeek-R1-0528-Qwen3-8B (DeepSeek R1 0528 Qwen3 8B) โ€” no longer on the official pricing page free list, deepseek-ai catalog is paid-only now
568
600
  // โ”€โ”€ B+ tier โ”€โ”€
569
- ['Qwen/Qwen3-8B', 'Qwen3 8B', 'B+', '30.0%', '131k'],
601
+ // Removed (2026-09-21): Qwen/Qwen3-8B ($0.06/M) + Qwen/Qwen2.5-7B-Instruct ($0.05/M) โ€” no longer free, both now paid on the official model pages; replacement: Qwen/Qwen3.5-4B
570
602
  // Removed (2026-09-05): deepseek-ai/DeepSeek-R1-Distill-Qwen-7B (DeepSeek R1 Distill Qwen 7B) - No longer listed on SiliconFlow pricing/catalog page (0 of 184 model records); superseded by the newer R1-0528 Qwen3 distill
571
- ['Qwen/Qwen3.5-4B', 'Qwen3.5 4B', 'A-', '-', '262k'],
572
- ['THUDM/GLM-4-9B-0414', 'GLM-4 9B', 'B+', '-', '32k'],
573
- ['Qwen/Qwen2.5-7B-Instruct', 'Qwen2.5 7B Instruct', 'B', '-', '32k'],
574
- ['Qwen/Qwen2.5-Coder-7B-Instruct', 'Qwen2.5 Coder 7B Instruct', 'B+', '-', '32k'], // Added (2026-09-15) โ€” verified via live audit
603
+ // Removed (2026-09-22): Qwen/Qwen3.5-4B (Qwen3.5 4B) โ€” absent from the official pricing page free list, remaining Qwen3.5 sizes are all paid
604
+ ['XingChenAGI/Xing4.0-29B', 'Xing4.0 29B', 'A-', '-', '262k'], // Added (2026-09-21) โ€” new $0 model on the official pricing page (181 records checked); engineering/coding focused
605
+ ['tencent/Hunyuan-MT-7B', 'Hunyuan MT 7B', 'C', '-', '33k'], // Added (2026-09-22) โ€” new $0 model on the official pricing page (translation-tuned, kept for breadth)
606
+ // Removed (2026-09-21): Qwen/Qwen2.5-Coder-7B-Instruct (Qwen2.5 Coder 7B Instruct) โ€” taken offline by SiliconFlow (official release note 2026-03-10, effective 2026-03-17; 0 of 181 records on today's pricing page); the 2026-09-15 re-add was erroneous. Replacement: Qwen/Qwen3-8B
575
607
  ]
576
608
 
577
609
  // ๐Ÿ“– Requesty source - https://router.requesty.ai/v1
@@ -603,10 +635,10 @@ export const requesty = [
603
635
  // ๐Ÿ“– OrcaRouter source - https://api.orcarouter.ai/v1
604
636
  // ๐Ÿ“– OpenAI-compatible gateway: https://api.orcarouter.ai/v1/chat/completions
605
637
  // ๐Ÿ“– Zero-markup AI gateway: token prices are passed through at provider rates, so only
606
- // ๐Ÿ“– the explicitly $-0 models are listed here. Verified live 2026-08-30 via GET /v1/models
607
- // ๐Ÿ“– (204 models, 3 with pricing.request=0). The orcarouter/fusion + orcarouter/free
608
- // ๐Ÿ“– adaptive-routing models are reachable through the same endpoint for users who opt
609
- // ๐Ÿ“– into pay-as-you-go billing, but are not free so they stay out of this catalog.
638
+ // ๐Ÿ“– the explicitly $-0 models are listed here. Verified live 2026-09-21 via GET /v1/models.
639
+ // ๐Ÿ“– orcarouter/free reports $0 pricing and stays listed. The orcarouter/fusion family now
640
+ // ๐Ÿ“– reports $0 with context lengths (verified live 2026-09-22, 197 models checked), so the
641
+ // ๐Ÿ“– adaptive-routing trio is listed; lineage is still undocumented, re-verify next audit.
610
642
  export const orcarouter = [
611
643
  // โ”€โ”€ S+ tier โ€” SWE-bench Verified โ‰ฅ70% โ”€โ”€
612
644
  ['deepseek/deepseek-v4-flash-free', 'DeepSeek V4 Flash (Free)', 'S+', '79.0%', '1M'],
@@ -616,6 +648,10 @@ export const orcarouter = [
616
648
  // โ”€โ”€ A+ tier โ€” SWE-bench Verified 50โ€“60% โ”€โ”€
617
649
  // Removed (2026-09-15): qwen/qwen3.8-27b-free (Qwen3.8 27B (Free)) โ€” no longer in catalog; only paid variant remains; replacement: z-ai/glm-5.3-flash-free
618
650
  ['z-ai/glm-5.3-flash-free', 'GLM-5.3 Flash (Free)', 'A+', '-', '1M'], // Added (2026-09-15) โ€” verified via live audit
651
+ // โ”€โ”€ A tier โ€” SWE-bench Verified 40โ€“50% โ”€โ”€
652
+ ['orcarouter/fusion-mini', 'OrcaRouter Fusion Mini', 'A+', '-', '1M'], // Added (2026-09-22) โ€” $0 per live /v1/models, adaptive-routed (undocumented lineage)
653
+ ['orcarouter/fusion', 'OrcaRouter Fusion', 'S', '-', '1M'], // Added (2026-09-22) โ€” $0 per live /v1/models, adaptive-routed (undocumented lineage)
654
+ ['orcarouter/fusion-flash', 'OrcaRouter Fusion Flash', 'B+', '-', '256k'], // Added (2026-09-22) โ€” $0 per live /v1/models, adaptive-routed (undocumented lineage)
619
655
  ]
620
656
 
621
657
  // ๐Ÿ“– Vercel AI Gateway source - https://vercel.com/docs/ai-gateway
@@ -36,7 +36,7 @@ import { sleep } from './shared-helpers.js'
36
36
  // ๐Ÿ“– is not guaranteed to be accepted by their chat endpoint.
37
37
  export const PROVIDER_TEST_MODEL_OVERRIDES = {
38
38
  sambanova: ['MiniMax-M2.5', 'DeepSeek-V3.1', 'DeepSeek-V3.2'],
39
- nvidia: ['deepseek-ai/deepseek-v4-flash-0731', 'openai/gpt-oss-120b'],
39
+ nvidia: ['moonshotai/kimi-k3', 'openai/gpt-oss-120b'],
40
40
  'github-models': ['openai/gpt-4.1-mini'],
41
41
  mistral: ['mistral-small-latest', 'devstral-small-latest'],
42
42
  }