@blockrun/llm 3.13.1 → 3.13.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,9 +2,9 @@
2
2
 
3
3
  # @blockrun/llm
4
4
 
5
- ### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
5
+ ### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
6
6
 
7
- The smart-routing SDK for <!-- br:models.chatVisible -->70<!-- /br:models.chatVisible --> models — every request goes to the cheapest model that can handle it,
7
+ The smart-routing SDK for <!-- br:models.chatVisible -->74<!-- /br:models.chatVisible --> models — every request goes to the cheapest model that can handle it,
8
8
  paid per-request in USDC. No API keys. No subscriptions. No vendor lock-in.
9
9
 
10
10
  [![npm](https://img.shields.io/npm/v/@blockrun/llm.svg?style=flat-square)](https://www.npmjs.com/package/@blockrun/llm)
@@ -36,12 +36,12 @@ console.log(r.routing.savings); // 0.96 — this exact request cost 96% less th
36
36
  console.log(r.response); // the proof
37
37
  ```
38
38
 
39
- **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** across a realistic workload on the default `auto` profile, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco` — and eco's first stop is the free tier, so simple requests cost $0.00 outright. Not an "up to" figure: the baseline, workload mix, and token ratio are published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json) so anyone can recompute the claim. Details in [Smart Routing](#smart-routing-router-core-v3).
39
+ **<!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** across a realistic workload on the default `auto` profile, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco` — and eco's first stop is the free tier, so simple requests cost $0.00 outright. Not an "up to" figure: the baseline, workload mix, and token ratio are published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json) so anyone can recompute the claim. Details in [Smart Routing](#smart-routing-router-core-v3).
40
40
 
41
41
  ## Why This SDK
42
42
 
43
43
  - 🧠 **Smart routing that pays for itself** — the bundled [Router Core V3](https://github.com/BlockRunAI/router-core) engine (shared with [ClawRouter](https://github.com/BlockRunAI/ClawRouter)) classifies every request locally in <1ms across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and routes to the cheapest capable model. The main event.
44
- - 🆓 **<!-- br:models.free -->5<!-- /br:models.free --> genuinely free models** — NVIDIA-hosted, $0 in and out, incl. 1M-context DeepSeek V4 Flash and a multimodal Nemotron. No rate-limit gimmicks.
44
+ - 🆓 **<!-- br:models.free -->7<!-- /br:models.free --> genuinely free models** — $0 in and out, incl. two 1M-context Nemotrons, a multimodal one, and free coding models from Cohere and Poolside. No rate-limit gimmicks.
45
45
  - 🔐 **No API keys** — your wallet signature is your authentication. No accounts, no dashboards, no key rotation.
46
46
  - 💸 **Pay per request in USDC** — x402 micropayments on Base or Solana. $5 covers thousands of requests; agents can pay their own way.
47
47
  - 🛡️ **Automatic failover** — transient errors (timeouts, 429, 5xx) walk the router's ranked fallback chain instead of failing your request.
@@ -52,9 +52,9 @@ console.log(r.response); // the proof
52
52
 
53
53
  | | OpenAI SDK | OpenRouter | LiteLLM | **@blockrun/llm** |
54
54
  | ------------------ | -------------- | ----------------- | ---------------- | ----------------------------------------------------------------------- |
55
- | **Cost routing** | ✗ one vendor | Manual selection | Manual selection | **Automatic — <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
56
- | **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->70<!-- /br:models.chatVisible -->, one wallet** |
57
- | **Free tier** | ✗ | Rate-limited | ✗ | **<!-- br:models.free -->5<!-- /br:models.free --> models, no signup** |
55
+ | **Cost routing** | ✗ one vendor | Manual selection | Manual selection | **Automatic — <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
56
+ | **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->74<!-- /br:models.chatVisible -->, one wallet** |
57
+ | **Free tier** | ✗ | Rate-limited | ✗ | **<!-- br:models.free -->7<!-- /br:models.free --> models, no signup** |
58
58
  | **Auth** | API key | Account + API key | Your API keys | **Wallet signature** |
59
59
  | **Payment** | Card + invoice | Credit card | BYO keys | **USDC per-request** |
60
60
  | **Agent-ready** | ✗ | ✗ | ✗ | **✓ — agents fund their own wallet** |
@@ -115,7 +115,7 @@ package to install.
115
115
 
116
116
  ### Try It Free (No USDC Required)
117
117
 
118
- Want to kick the tires before funding a wallet? Route to BlockRun's free NVIDIA tier:
118
+ Want to kick the tires before funding a wallet? Route to BlockRun's free tier:
119
119
 
120
120
  ```typescript
121
121
  import { LLMClient } from '@blockrun/llm';
@@ -123,33 +123,34 @@ import { LLMClient } from '@blockrun/llm';
123
123
  const client = new LLMClient(); // Wallet still required for signing, but $0 charged
124
124
 
125
125
  // Option 1: call a free model directly
126
- const reply = await client.chat('nvidia/step-3.7-flash', 'Explain x402 in 1 sentence');
126
+ const reply = await client.chat('nvidia/nemotron-3.5-lightning', 'Explain x402 in 1 sentence');
127
127
 
128
- // Option 2: let the smart router pick — 'eco' ranks the free NVIDIA tier first
128
+ // Option 2: let the smart router pick — 'eco' ranks the free tier first
129
129
  const result = await client.smartChat('What is 2+2?', { routingProfile: 'eco' });
130
- console.log(result.model); // 'nvidia/step-3.7-flash' ($0 — verified live)
130
+ console.log(result.model); // a free-tier model — $0 in and out
131
131
  console.log(result.response); // '4'
132
132
  console.log(result.routing.savings); // 1 (100%)
133
133
  ```
134
134
 
135
135
  There is no `free` routing profile in `smartChat()` — `routingProfile` accepts
136
136
  `'eco' | 'auto' | 'premium'`. (ClawRouter's `/model free` is a feature of its
137
- own proxy, not of this SDK's router options.) For guaranteed $0, pin a
138
- `nvidia/*` model; for smart-routed $0-first, use `eco`.
137
+ own proxy, not of this SDK's router options.) For guaranteed $0, pin one of
138
+ the free model ids below; for smart-routed $0-first, use `eco`.
139
139
 
140
- **Available free models** (input + output both $0, all NVIDIA-hosted, from the live `/v1/models` catalog, last refreshed 2026-08-12):
140
+ **Available free models** — input and output both $0. The free tier is **no
141
+ longer NVIDIA-only**, so pin these by full model id rather than by an
142
+ `nvidia/*` prefix. Full contexts and notes in [Free Tier](#free-tier);
143
+ `client.listModels()` returns the live catalog at runtime.
141
144
 
142
145
  | Model ID | Context | Best For |
143
146
  |----------|---------|----------|
144
- | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning — text + images + video + audio (ChartQA 90.3, DocVQA 95.6) |
145
- | `nvidia/mistral-nemotron` | 131K | Mistral × NVIDIA instruction model — fast (~0.2s), strong instruction following |
146
- | `nvidia/step-3.7-flash` | 131K | StepFun Step 3.7 Flash — fast lightweight reasoning |
147
- | `nvidia/nemotron-nano-9b-v2` | 131K | Compact + fast (~0.7s), good for high-volume light tasks |
148
- | `nvidia/nemotron-nano-12b-v2-vl` | 131K | Vision-language — accepts images, compact + fast |
149
- | `nvidia/gpt-oss-120b` | 128K | OpenAI open-weight 120B. Hidden from `/v1/models` for privacy but direct calls still work |
150
- | `nvidia/gpt-oss-20b` | 128K | OpenAI open-weight 20B. Hidden from `/v1/models` but direct calls still work |
151
-
152
- > Privacy note: `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` are hidden from `/v1/models` because NVIDIA's free build.nvidia.com tier reserves the right to use prompts/outputs for service improvement. Direct calls by full model ID still work — opt in only when your data isn't sensitive.
147
+ | `nvidia/nemotron-3.5-lightning` | 1M | Thinking-mode reasoning at 1M context |
148
+ | `nvidia/nemotron-3-ultra-550b` | 1M | Largest free model — 550B |
149
+ | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning — text + images |
150
+ | `nvidia/nemotron-3-nano-30b` | 128K | Compact + fast, good for high-volume light tasks |
151
+ | `nvidia/llama-3.2-11b-vision` | 128K | Vision-language — accepts images |
152
+ | `cohere/north-mini-code` | 256K | Compact coding model, sub-second responses |
153
+ | `poolside/laguna-xs-2.1` | 128K | Coding model |
153
154
 
154
155
  ## Quick Start (Solana)
155
156
 
@@ -166,7 +167,7 @@ Set `SOLANA_WALLET_KEY` to your bs58-encoded Solana secret key. Payments are aut
166
167
 
167
168
  ## Smart Routing (Router Core V3)
168
169
 
169
- Let the SDK automatically pick the cheapest capable model for each request — **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** for the same traffic on `auto`, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco`.
170
+ Let the SDK automatically pick the cheapest capable model for each request — **<!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** for the same traffic on `auto`, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco`.
170
171
 
171
172
  Not an "up to" figure. The baseline, the workload mix and the token ratio are
172
173
  published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json),
@@ -210,12 +211,12 @@ const client = new LLMClient();
210
211
  // Auto-routes to cheapest capable model
211
212
  const result = await client.smartChat('What is 2+2?');
212
213
  console.log(result.response); // '4'
213
- console.log(result.model); // 'moonshot/kimi-k2.5' (cheap, fast)
214
- console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 88%'
214
+ console.log(result.model); // 'qwen/qwen3.7-flash' (cheap, fast)
215
+ console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // this request, vs the Opus 5 baseline
215
216
 
216
217
  // Complex reasoning task -> routes to reasoning model
217
218
  const complex = await client.smartChat('Prove the Riemann hypothesis step by step');
218
- console.log(complex.model); // 'xai/grok-4-1-fast-reasoning'
219
+ console.log(complex.model); // 'xai/grok-4.3'
219
220
 
220
221
  // Inspect how the request was classified and ranked (Router v3.4 portfolio).
221
222
  console.log(complex.routing.method); // 'portfolio'
@@ -232,14 +233,16 @@ console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
232
233
  `chat()` / `chatCompletion()` walk it automatically when the primary model
233
234
  returns a transient error — timeouts, network failures, 429 rate limits, or
234
235
  5xx responses (502/503/504/522/524). Other 4xx errors and `PaymentError`
235
- propagate immediately so wallet / auth issues surface fast.
236
+ propagate immediately so wallet / auth issues surface fast. (Solana's internal
237
+ stale-blockhash re-sign is a separate, lower-level retry inside the payment
238
+ step — see [How Payment Works](#phase-2--every-request-pays-itself-automatic-x402).)
236
239
 
237
240
  ```typescript
238
241
  // Manually pass a fallback chain to chat() / chatCompletion()
239
- const reply = await client.chat('nvidia/step-3.7-flash', 'hello', {
240
- fallbackModels: ['nvidia/mistral-nemotron', 'nvidia/gpt-oss-120b'],
242
+ const reply = await client.chat('nvidia/nemotron-3.5-lightning', 'hello', {
243
+ fallbackModels: ['nvidia/nemotron-3-nano-30b', 'cohere/north-mini-code'],
241
244
  });
242
- // If step-3.7-flash times out, the SDK retries against the next model
245
+ // If nemotron-3.5-lightning times out, the SDK retries against the next model
243
246
  // and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
244
247
  ```
245
248
 
@@ -247,11 +250,11 @@ const reply = await client.chat('nvidia/step-3.7-flash', 'hello', {
247
250
 
248
251
  | Profile | Strategy | Savings vs Opus 5 | Best For |
249
252
  |---------|----------|-------------------|----------|
250
- | `eco` | Cheapest capable model — ranks the <!-- br:models.free -->5<!-- /br:models.free -->-model free NVIDIA tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
251
- | `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%** | General use |
253
+ | `eco` | Cheapest capable model — ranks the <!-- br:models.free -->7<!-- /br:models.free -->-model free tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
254
+ | `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->%** | General use |
252
255
  | `premium` | Top-tier models (OpenAI, Anthropic) | 0% | Quality-critical tasks |
253
256
 
254
- For guaranteed $0, call a `nvidia/*` model directly with `chat()` — see
257
+ For guaranteed $0, call a free model directly with `chat()` — see
255
258
  [Try It Free](#try-it-free-no-usdc-required). ClawRouter's `/model free`
256
259
  profile belongs to its own proxy; `smartChat()`'s options are the three above.
257
260
 
@@ -293,14 +296,14 @@ anchors the portfolio's candidate pool):
293
296
 
294
297
  | Tier | Example Tasks | ECO | AUTO | PREMIUM |
295
298
  |------|---------------|-----|------|---------|
296
- | SIMPLE | "What is 2+2?", definitions | free/gpt-oss-120b † (**FREE**) | gemini-2.5-flash ($0.30/$2.50) | kimi-k2.7 † ($0.95/$4.00) |
297
- | MEDIUM | Code snippets, explanations | gemini-3.1-flash-lite ($0.25/$1.50) | kimi-k2.7 ($0.95/$4.00) | gpt-5.3-codex ($1.75/$14.00) |
298
- | COMPLEX | Architecture, long documents | gemini-3.1-flash-lite ($0.25/$1.50) | gemini-3.1-pro ($2/$12) | claude-fable-5 ($10/$50) |
299
- | REASONING | Proofs, multi-step reasoning | grok-4-1-fast-reasoning † ($0.20/$0.50) | grok-4-1-fast-reasoning † ($0.20/$0.50) | claude-sonnet-4.6 ($3/$15) |
299
+ | SIMPLE | "What is 2+2?", definitions | step-3.7-flash (**FREE**) | gemini-2.5-flash ($0.30/$2.50) | gemini-3.5-flash ($1.50/$9) |
300
+ | MEDIUM | Code snippets, explanations | glm-5.3-flash ($0.15/$0.50) | gemini-3.5-flash ($1.50/$9) | gpt-5.3-codex ($1.75/$14.00) |
301
+ | COMPLEX | Architecture, long documents | glm-5.3-flash ($0.15/$0.50) | gemini-3.1-pro ($2/$12) | claude-fable-5 ($10/$50) |
302
+ | REASONING | Proofs, multi-step reasoning | deepseek-reasoner ($0.14/$0.28) | deepseek-reasoner ($0.14/$0.28) | claude-sonnet-5 ($3/$15) |
300
303
 
301
- † Withheld from `/v1/models` — the router still calls it by direct ID, but you
302
- will not find it on the public pricing page. The published savings claim is
303
- priced on visible models only.
304
+ Since Router Core V3.5 every primary and every fallback rung is a model listed
305
+ on `/v1/models` — nothing the router picks is withheld from the public pricing
306
+ page, and the published savings claim is priced on those same visible models.
304
307
 
305
308
  This table mirrors ClawRouter's tier configs at the version this SDK pins;
306
309
  the [ClawRouter README](https://github.com/BlockRunAI/ClawRouter#how-it-works)
@@ -367,7 +370,7 @@ const response = await client.chat('openai/gpt-4o', 'gm Solana');
367
370
  console.log(response);
368
371
 
369
372
  // Live Search with Grok (Solana payment)
370
- const tweet = await client.chat('xai/grok-3-mini', 'What is trending on X?', { search: true });
373
+ const tweet = await client.chat('xai/grok-4.5', 'What is trending on X?', { search: true });
371
374
  ```
372
375
 
373
376
  **Setup:**
@@ -395,7 +398,7 @@ You only do this when your balance runs low. Three ways to get USDC into your wa
395
398
 
396
399
  - **(b) Transfer existing USDC.** Send USDC you already hold to your wallet address (`client.getWalletAddress()`). On Base, send Base USDC; on Solana (`SolanaLLMClient`), send Solana SPL USDC.
397
400
 
398
- - **(c) Skip funding entirely.** Use the free NVIDIA models (e.g. `nvidia/step-3.7-flash`) — every call is **$0**, no balance required.
401
+ - **(c) Skip funding entirely.** Use the free models (e.g. `nvidia/nemotron-3.5-lightning`) — every call is **$0**, no balance required.
399
402
 
400
403
  $5 of USDC covers thousands of paid requests. Check your balance any time:
401
404
 
@@ -414,7 +417,19 @@ You just call e.g. `client.chat(...)` — the payment is invisible:
414
417
  4. The request is retried automatically with the payment proof.
415
418
  5. The gateway settles on-chain and returns the AI response.
416
419
 
417
- One call, no separate pay step. Free NVIDIA models settle at **$0** (no payment signed).
420
+ One call, no separate pay step. Free-tier models settle at **$0** (no payment signed).
421
+
422
+ On **Solana**, step 3 pins the payment to a recent blockhash that is valid for
423
+ roughly 60 seconds. If one expires between signing and verification,
424
+ `SolanaLLMClient` re-signs against a fresh blockhash and retries — up to twice,
425
+ with a short backoff — rather than surfacing a payment error you would only
426
+ have to retry by hand.
427
+
428
+ The retry is deliberately narrow. It fires only when the gateway explicitly
429
+ reports a *verification-phase* stale blockhash. Settlement failures, ambiguous
430
+ rejections, insufficient funds and malformed responses all fail immediately.
431
+ Verification runs strictly before settlement, so a retryable rejection means no
432
+ transaction was broadcast and you cannot be charged twice.
418
433
 
419
434
  ### Track spend and verify settlements
420
435
 
@@ -460,7 +475,7 @@ const video = await br.poll('/v1/videos/generations', {
460
475
 
461
476
  // Streaming SSE — chat completions
462
477
  for await (const chunk of br.stream('/v1/chat/completions', {
463
- model: 'anthropic/claude-sonnet-4-6',
478
+ model: 'anthropic/claude-sonnet-5',
464
479
  messages: [{ role: 'user', content: 'Hi' }],
465
480
  stream: true,
466
481
  })) {
@@ -481,157 +496,169 @@ shims over `BlockrunClient`) and removed in 3.0.
481
496
 
482
497
  ## Available Models
483
498
 
484
- ### OpenAI GPT-5.5 Family
485
- Released 2026-04-23 — first fully retrained base since GPT-4.5. 1M context, 128K output, native agent + computer use.
486
-
487
- | Model | Input Price | Output Price |
488
- |-------|-------------|--------------|
489
- | `openai/gpt-5.5` | $5.00/M | $30.00/M |
490
-
491
- ### OpenAI GPT-5.4 Family
492
- | Model | Input Price | Output Price |
493
- |-------|-------------|--------------|
494
- | `openai/gpt-5.4` | $2.50/M | $15.00/M |
495
- | `openai/gpt-5.4-pro` | $30.00/M | $180.00/M |
496
- | `openai/gpt-5.4-nano` | $0.20/M | $1.25/M |
497
-
498
- ### OpenAI GPT-5 Family
499
- | Model | Input Price | Output Price |
500
- |-------|-------------|--------------|
501
- | `openai/gpt-5.3` | $1.75/M | $14.00/M |
502
- | `openai/gpt-5.2` | $1.75/M | $14.00/M |
503
- | `openai/gpt-5-mini` | $0.25/M | $2.00/M |
504
- | `openai/gpt-5.2-pro` | $21.00/M | $168.00/M |
505
- | `openai/gpt-5.2-codex` | $1.75/M | $14.00/M |
499
+ Prices below are the live gateway rates, regenerated from `GET /v1/models`
500
+ (the same catalog `client.listModels()` returns). Chat models are billed per
501
+ token; image, video, music and speech are billed per unit as noted in their
502
+ own sections.
503
+
504
+ ### OpenAI GPT-5.6 Family
505
+
506
+ Three tiers on one 1.05M-context base — Sol (deepest reasoning), Terra
507
+ (balanced), Luna (cheap and fast). Each has a `-pro` sibling that thinks
508
+ longer at the same token price.
509
+
510
+ | Model | Input Price | Output Price | Context |
511
+ |-------|-------------|--------------|---------|
512
+ | `openai/gpt-5.6-sol` | $5.00/M | $30.00/M | 1.05M |
513
+ | `openai/gpt-5.6-sol-pro` | $5.00/M | $30.00/M | 1.05M |
514
+ | `openai/gpt-5.6-terra` | $2.00/M | $12.00/M | 1.05M |
515
+ | `openai/gpt-5.6-terra-pro` | $2.00/M | $12.00/M | 1.05M |
516
+ | `openai/gpt-5.6-luna` | $0.20/M | $1.20/M | 1.05M |
517
+ | `openai/gpt-5.6-luna-pro` | $0.20/M | $1.20/M | 1.05M |
518
+
519
+ ### OpenAI GPT-5.5 / 5.4 / 5.2 Families
520
+
521
+ | Model | Input Price | Output Price | Context | Notes |
522
+ |-------|-------------|--------------|---------|-------|
523
+ | `openai/gpt-5.5` | $5.00/M | $30.00/M | 1.05M | |
524
+ | `openai/gpt-5.5-pro` | $30.00/M | $180.00/M | 1.05M | |
525
+ | `openai/chat-latest` | $5.00/M | $30.00/M | 128K | ChatGPT Instant — the model behind chatgpt.com |
526
+ | `openai/gpt-5.4` | $2.50/M | $15.00/M | 1.05M | |
527
+ | `openai/gpt-5.4-pro` | $30.00/M | $180.00/M | 1.05M | |
528
+ | `openai/gpt-5.4-mini` | $0.75/M | $4.50/M | 400K | |
529
+ | `openai/gpt-5.4-nano` | $0.20/M | $1.25/M | 1.05M | |
530
+ | `openai/gpt-5.2` | $1.75/M | $14.00/M | 400K | |
531
+ | `openai/gpt-5.2-pro` | $21.00/M | $168.00/M | 400K | |
532
+ | `openai/gpt-5.3-codex` | $1.75/M | $14.00/M | 400K | Coding/agentic SKU |
533
+ | `openai/gpt-5-mini` | $0.25/M | $2.00/M | 200K | |
506
534
 
507
535
  ### OpenAI GPT-4 Family
508
- | Model | Input Price | Output Price |
509
- |-------|-------------|--------------|
510
- | `openai/gpt-4.1` | $2.00/M | $8.00/M |
511
- | `openai/gpt-4.1-mini` | $0.40/M | $1.60/M |
512
- | `openai/gpt-4.1-nano` | $0.10/M | $0.40/M |
513
- | `openai/gpt-4o` | $2.50/M | $10.00/M |
514
- | `openai/gpt-4o-mini` | $0.15/M | $0.60/M |
536
+
537
+ | Model | Input Price | Output Price | Context |
538
+ |-------|-------------|--------------|---------|
539
+ | `openai/gpt-4.1` | $2.00/M | $8.00/M | 128K |
540
+ | `openai/gpt-4.1-mini` | $0.40/M | $1.60/M | 128K |
541
+ | `openai/gpt-4.1-nano` | $0.10/M | $0.40/M | 128K |
542
+ | `openai/gpt-4o` | $2.50/M | $10.00/M | 128K |
543
+ | `openai/gpt-4o-mini` | $0.15/M | $0.60/M | 128K |
515
544
 
516
545
  ### OpenAI O-Series (Reasoning)
517
- | Model | Input Price | Output Price |
518
- |-------|-------------|--------------|
519
- | `openai/o1` | $15.00/M | $60.00/M |
520
- | `openai/o3` | $2.00/M | $8.00/M |
521
- | `openai/o3-mini` | $1.10/M | $4.40/M |
522
- | `openai/o4-mini` | $1.10/M | $4.40/M |
546
+
547
+ | Model | Input Price | Output Price | Context |
548
+ |-------|-------------|--------------|---------|
549
+ | `openai/o1` | $15.00/M | $60.00/M | 200K |
550
+ | `openai/o3` | $2.00/M | $8.00/M | 200K |
551
+ | `openai/o3-mini` | $1.10/M | $4.40/M | 128K |
552
+ | `openai/o4-mini` | $1.10/M | $4.40/M | 128K |
523
553
 
524
554
  ### Anthropic Claude
555
+
525
556
  | Model | Input Price | Output Price | Context | Notes |
526
557
  |-------|-------------|--------------|---------|-------|
527
- | `anthropic/claude-fable-5` | $10.00/M | $50.00/M | **1M** | Mythos-class flagship above Opus — always-on thinking, 128K output, fallback `claude-opus-4.8`. Alias: `claude-fable-5` |
528
- | `anthropic/claude-opus-4.8` | $5.00/M | $25.00/M | **1M** | Flagship — agentic coding + adaptive thinking, 128K output |
529
- | `anthropic/claude-opus-4.7` | $5.00/M | $25.00/M | **1M** | Agentic coding + adaptive thinking, 128K output |
530
- | `anthropic/claude-opus-4.6` | $5.00/M | $25.00/M | 200K | Hidden but still callable — kept as in-family hot-swap fallback |
531
- | `anthropic/claude-opus-4.5` | $5.00/M | $25.00/M | 200K | |
532
- | `anthropic/claude-opus-4` | $15.00/M | $75.00/M | 200K | |
533
- | `anthropic/claude-sonnet-4.6` | $3.00/M | $15.00/M | 200K | Best for reasoning/instructions |
534
- | `anthropic/claude-sonnet-4` | $3.00/M | $15.00/M | 200K | |
535
- | `anthropic/claude-haiku-4.5` | $1.00/M | $5.00/M | 200K | |
558
+ | `anthropic/claude-fable-5` | $10.00/M | $50.00/M | 1M | Mythos-class flagship above Opus — always-on thinking, 128K output |
559
+ | `anthropic/claude-opus-5` | $5.00/M | $25.00/M | 1M | Flagship — the baseline the routing savings claim is measured against |
560
+ | `anthropic/claude-opus-4.8` | $5.00/M | $25.00/M | 1M | Agentic coding + adaptive thinking, 128K output |
561
+ | `anthropic/claude-opus-4.7` | $5.00/M | $25.00/M | 1M | |
562
+ | `anthropic/claude-opus-4.5` | $5.00/M | $25.00/M | 200K | |
563
+ | `anthropic/claude-sonnet-5` | $3.00/M | $15.00/M | 1M | Best cost/quality balance for long-context agent turns |
564
+ | `anthropic/claude-sonnet-4.6` | $3.00/M | $15.00/M | 1M | |
565
+ | `anthropic/claude-sonnet-4.5` | $3.00/M | $15.00/M | 200K | |
566
+ | `anthropic/claude-haiku-4.5` | $1.00/M | $5.00/M | 200K | |
536
567
 
537
568
  ### Google Gemini
538
- | Model | Input Price | Output Price |
539
- |-------|-------------|--------------|
540
- | `google/gemini-3.1-pro` | $2.00/M | $12.00/M |
541
- | `google/gemini-3.5-flash` | $0.50/M | $3.00/M |
542
- | `google/gemini-3.1-flash-lite` | $0.25/M | $1.50/M |
543
- | `google/gemini-3-flash-preview` | $0.50/M | $3.00/M |
544
- | `google/gemini-2.5-pro` | $1.25/M | $10.00/M |
545
- | `google/gemini-2.5-flash` | $0.30/M | $2.50/M |
546
- | `google/gemini-2.5-flash-lite` | $0.10/M | $0.40/M |
569
+
570
+ | Model | Input Price | Output Price | Context |
571
+ |-------|-------------|--------------|---------|
572
+ | `google/gemini-3.1-pro` | $2.00/M | $12.00/M | 1M |
573
+ | `google/gemini-3.6-flash` | $1.50/M | $7.50/M | 1M |
574
+ | `google/gemini-3.5-flash` | $1.50/M | $9.00/M | 1M |
575
+ | `google/gemini-3-flash-preview` | $0.50/M | $3.00/M | 1M |
576
+ | `google/gemini-3.5-flash-lite` | $0.30/M | $2.50/M | 1M |
577
+ | `google/gemini-3.1-flash-lite` | $0.25/M | $1.50/M | 1M |
578
+ | `google/gemini-2.5-pro` | $1.25/M | $10.00/M | 1M |
579
+ | `google/gemini-2.5-flash` | $0.30/M | $2.50/M | 1M |
580
+ | `google/gemini-2.5-flash-lite` | $0.10/M | $0.40/M | 1M |
547
581
 
548
582
  ### DeepSeek
549
583
 
550
- V4 family launched 2026-04-24. DeepSeek upstream now serves the legacy
551
- `deepseek-chat` / `deepseek-reasoner` aliases as V4 Flash non-thinking /
552
- thinking modes. V4 Pro is the new flagship paid SKU — 1.6T MoE / 49B active,
553
- 1M context, MMLU-Pro 87.5, GPQA 90.1, SWE-bench 80.6, LiveCodeBench 93.5.
584
+ DeepSeek upstream serves the legacy `deepseek-chat` / `deepseek-reasoner`
585
+ aliases as V4 Flash non-thinking / thinking modes. V4 Pro is the flagship
586
+ paid SKU; the vision SKU is an experimental preview.
554
587
 
555
588
  | Model | Input Price | Output Price | Context | Notes |
556
589
  |-------|-------------|--------------|---------|-------|
557
- | `deepseek/deepseek-v4-pro` | $0.435/M | $0.87/M | 1M | V4 flagship — strongest open-weight reasoner. The 75% launch promo became the permanent list price after 2026-05-31 |
558
- | `deepseek/deepseek-chat` | $0.14/M | $0.28/M | 1M | V4 Flash non-thinking (paid endpoint with 5MB request bodies) |
559
- | `deepseek/deepseek-reasoner` | $0.20/M | $0.40/M | 1M | V4 Flash thinking (same upstream as `deepseek-chat`, thinking enabled by default) |
590
+ | `deepseek/deepseek-v4-pro` | $1.32/M | $3.96/M | 1M | V4 flagship — strongest open-weight reasoner |
591
+ | `deepseek/deepseek-v4-flash-vision-exp` | $0.44/M | $1.32/M | 1M | Experimental vision preview |
592
+ | `deepseek/deepseek-chat` | $0.14/M | $0.28/M | 1M | V4 Flash non-thinking |
593
+ | `deepseek/deepseek-reasoner` | $0.14/M | $0.28/M | 1M | V4 Flash thinking (same upstream, thinking on by default) |
560
594
 
561
595
  ### xAI Grok
562
596
 
563
- Grok 4.3 and Grok Build are resold through BlockRun's OpenRouter credit pool
564
- (same pattern as `deepseek/deepseek-v4-pro` and `minimax/minimax-m3`). The
565
- older Grok chat SKUs (grok-3/3-mini, grok-4-fast / 4-1-fast families,
566
- grok-code-fast-1, grok-4-0709, grok-2-vision) are now **hidden from
567
- `/v1/models`** — direct calls by full ID still work, but SmartChat won't
568
- auto-pick them.
597
+ The older Grok chat SKUs (grok-3/3-mini, the grok-4 fast families,
598
+ grok-code-fast-1, grok-2-vision) have left the catalog. Retired ids stay
599
+ callable — the gateway redirects them to a healthy model — but SmartChat
600
+ only ranks what `/v1/models` lists.
569
601
 
570
602
  | Model | Input Price | Output Price | Context | Notes |
571
603
  |-------|-------------|--------------|---------|-------|
572
- | `xai/grok-4.3` | $1.50/M | $4.00/M | 1M | Reasoning model, vision-capable, tuned for agentic workflows |
573
- | `xai/grok-build-0.1` | $1.50/M | $3.00/M | 256K | Fast agentic coding model — interactive software-engineering workflows |
574
-
575
- ### Moonshot Kimi
576
- | Model | Input Price | Output Price |
577
- |-------|-------------|--------------|
578
- | `moonshot/kimi-k2.6` | $0.95/M | $4.00/M |
579
- | `moonshot/kimi-k2.5` | $0.60/M | $3.00/M |
580
-
581
- ### MiniMax
582
- | Model | Input Price | Output Price |
583
- |-------|-------------|--------------|
584
- | `minimax/minimax-m3` | $0.30/M | $1.20/M |
585
- | `minimax/minimax-m2.7` | $0.30/M | $1.20/M |
586
-
587
- ### NVIDIA (Free) + Moonshot
588
-
589
- Free tier refreshed 2026-08-12. NVIDIA has retired (HTTP 410 end-of-life)
590
- the entire free DeepSeek family — `nvidia/deepseek-v4-flash` was the last
591
- to go — along with `llama-4-maverick`, the qwen3 SKUs, and the free
592
- Mistral small/large SKUs. Retired IDs stay callable: the gateway
593
- auto-redirects them to a healthy free model, so pinned callers still get
594
- a 200. `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` remain callable by
595
- direct ID but are hidden from `/v1/models` over the NVIDIA free tier's
596
- prompt-retention terms (so SmartChat won't auto-pick them).
597
-
598
- | Model | Input Price | Output Price | Notes |
599
- |-------|-------------|--------------|-------|
600
- | `nvidia/step-3.7-flash` | **FREE** | **FREE** | Fast general-purpose chat + reasoning, 131K |
601
- | `nvidia/mistral-nemotron` | **FREE** | **FREE** | Fast free Mistral, 131K |
602
- | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 31B / 3.2B active MoE, 256K — only vision-capable free model |
603
- | `nvidia/nemotron-nano-9b-v2` | **FREE** | **FREE** | Compact fast chat, 131K |
604
- | `nvidia/nemotron-nano-12b-v2-vl` | **FREE** | **FREE** | Compact vision, 131K |
605
- | `nvidia/gpt-oss-120b` | **FREE** | **FREE** | Hidden from `/v1/models` for privacy but direct calls still work — 123 tok/s |
606
- | `nvidia/gpt-oss-20b` | **FREE** | **FREE** | Hidden from `/v1/models` but direct calls still work — 155 tok/s |
607
- | `moonshot/kimi-k2.5` | $0.60/M | $3.00/M | Direct from Moonshot — replaces `nvidia/kimi-k2.5` |
608
-
609
- ### E2E Verified Models
610
-
611
- All models below have been tested end-to-end via the TypeScript SDK (Feb 2026):
612
-
613
- | Provider | Model | Status |
614
- |----------|-------|--------|
615
- | OpenAI | `openai/gpt-4o-mini` | Passed |
616
- | OpenAI | `openai/gpt-5.2-codex` | Passed |
617
- | Anthropic | `anthropic/claude-opus-4.6` | Passed |
618
- | Anthropic | `anthropic/claude-sonnet-4` | Passed |
619
- | Google | `google/gemini-2.5-flash` | Passed |
620
- | DeepSeek | `deepseek/deepseek-chat` | Passed |
621
- | xAI | `xai/grok-3` | Passed |
622
- | Moonshot | `moonshot/kimi-k2.6` | Passed |
604
+ | `xai/grok-4.5` | $2.50/M | $9.00/M | 500K | Flagship — reasoning + vision, native Live Search (`search: true`) |
605
+ | `xai/grok-4.3` | $1.50/M | $4.00/M | 1M | Reasoning + vision, tuned for agentic workflows |
606
+ | `xai/grok-build-0.1` | $1.50/M | $3.00/M | 256K | Fast agentic coding model |
607
+
608
+ ### Moonshot, MiniMax, Z.ai, Qwen
609
+
610
+ | Model | Input Price | Output Price | Context | Notes |
611
+ |-------|-------------|--------------|---------|-------|
612
+ | `moonshot/kimi-k3` | $3.00/M | $15.00/M | 1M | Replaces the retired `kimi-k2.5` / `k2.6` SKUs |
613
+ | `minimax/minimax-m3` | $0.30/M | $1.20/M | 1M | |
614
+ | `minimax/minimax-m2.7` | $0.30/M | $1.20/M | 200K | |
615
+ | `zai/glm-5.3` | $1.40/M | $4.40/M | 1M | |
616
+ | `zai/glm-5.3-flash` | $0.15/M | $0.50/M | 1M | Cheapest vision-capable paid SKU |
617
+ | `zai/glm-5.2` | $1.40/M | $4.40/M | 1M | |
618
+ | `zai/glm-5.1` | $1.40/M | $4.40/M | 200K | |
619
+ | `zai/glm-5` | $1.00/M | $3.20/M | 200K | |
620
+ | `zai/glm-5-turbo` | $1.20/M | $4.00/M | 200K | |
621
+ | `qwen/qwen3.7-max` | $1.475/M | $4.425/M | 1M | |
622
+ | `qwen/qwen3.7-plus` | $0.32/M | $1.28/M | 1M | |
623
+ | `qwen/qwen3.8-flash` | $0.15/M | $0.47/M | 1M | |
624
+ | `qwen/qwen3.7-flash` | $0.03/M | $0.13/M | 1M | Cheapest paid chat model in the catalog |
625
+
626
+ ### Tencent, Xiaomi
627
+
628
+ | Model | Input Price | Output Price | Context |
629
+ |-------|-------------|--------------|---------|
630
+ | `tencent/hy3` | $0.132/M | $0.528/M | 256K |
631
+ | `xiaomi/mimo-v2.5` | $0.14/M | $0.28/M | 1M |
632
+ | `xiaomi/mimo-v2.5-pro` | $0.435/M | $0.87/M | 1M |
633
+
634
+ ### Free Tier
635
+
636
+ Input and output both $0 — no promo, no rate-limit gimmick. The free tier is
637
+ **no longer NVIDIA-only**, so pin these by full model id rather than by an
638
+ `nvidia/*` prefix, or let `routingProfile: 'eco'` rank them first.
639
+
640
+ | Model | Input Price | Output Price | Context | Notes |
641
+ |-------|-------------|--------------|---------|-------|
642
+ | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 256K | Multimodal reasoning — text + images |
643
+ | `nvidia/nemotron-3.5-lightning` | **FREE** | **FREE** | 1M | Thinking-mode reasoning at 1M context |
644
+ | `nvidia/nemotron-3-nano-30b` | **FREE** | **FREE** | 128K | Compact and fast, good for high-volume light tasks |
645
+ | `nvidia/llama-3.2-11b-vision` | **FREE** | **FREE** | 128K | Vision-language — accepts images |
646
+ | `nvidia/nemotron-3-ultra-550b` | **FREE** | **FREE** | 1M | Largest free model — 550B, 1M context |
647
+ | `cohere/north-mini-code` | **FREE** | **FREE** | 256K | Compact coding model, sub-second responses |
648
+ | `poolside/laguna-xs-2.1` | **FREE** | **FREE** | 128K | Coding model |
623
649
 
624
650
  ### Image Generation
625
- | Model | Price |
626
- |-------|-------|
627
- | `openai/dall-e-3` | $0.04-0.08/image |
628
- | `openai/gpt-image-1` | $0.02-0.04/image |
629
- | `openai/gpt-image-2` | $0.06-0.12/image (reasoning-driven, multilingual text rendering, character consistency) |
630
- | `google/nano-banana` | $0.05/image |
631
- | `google/nano-banana-pro` | $0.10-0.15/image |
632
- | `xai/grok-imagine-image` | $0.02/image |
633
- | `xai/grok-imagine-image-pro` | $0.07/image |
634
- | `zai/cogview-4` | $0.015/image |
651
+ | Model | Price | Notes |
652
+ |-------|-------|-------|
653
+ | `openai/gpt-image-1` | $0.02/image | Native GPT-4o image generation |
654
+ | `openai/gpt-image-2` | $0.06/image | Reasoning-driven — multilingual text rendering, character consistency |
655
+ | `google/nano-banana` | $0.05/image | Gemini 2.5 Flash image generation — fast and efficient |
656
+ | `google/nano-banana-2` | $0.09/image | Gemini 3.1 Flash — pro-level quality at Flash speed |
657
+ | `google/nano-banana-pro` | $0.10/image | Gemini 3 Pro — highest quality, up to 4K |
658
+ | `xai/grok-imagine-image` | $0.02/image | Fast, 300 RPM |
659
+ | `xai/grok-imagine-image-pro` | $0.07/image | Quality tier, 30 RPM |
660
+ | `bytedance/seedream-5-pro` | $0.045/image | Flagship generation + editing, up to 4K-class, reference images |
661
+ | `zai/cogview-4` | $0.015/image | Up to 1440x1440 |
635
662
 
636
663
  Image editing (`client.edit`) via `/v1/images/image2image`: `openai/gpt-image-1`, `openai/gpt-image-2`, `google/nano-banana`, and `google/nano-banana-pro`. Pass a single base64 `data:image/...` URI to edit one image, or an array of 2–4 URIs to **fuse** them (e.g. a subject + a brand logo). Fusion caps: `openai/*` up to 4 source images, `google/*` up to 3. A `mask` cannot be combined with multiple source images.
637
664
 
@@ -646,12 +673,16 @@ console.log(fused.data[0].url);
646
673
  ```
647
674
 
648
675
  ### Video Generation
649
- | Model | Price |
650
- |-------|-------|
651
- | `xai/grok-imagine-video` | $0.05/sec (8s default → $0.42/clip) |
652
- | `bytedance/seedance-1.5-pro` | $0.03/sec (5s default, up to 10s, 720p) |
653
- | `bytedance/seedance-2.0-fast` | $0.15/sec (~60-80s gen, sweet-spot price/quality) |
654
- | `bytedance/seedance-2.0` | $0.30/sec (720p Pro) |
676
+ | Model | Price | Default | Max | Notes |
677
+ |-------|-------|---------|-----|-------|
678
+ | `xai/grok-imagine-video` | $0.05/sec | 8s | 15s | 480p default, 720p at $0.07/sec; text or image to video |
679
+ | `xai/grok-imagine-video-1.5` | $0.08/sec | 8s | 15s | Flagship — native synced audio; 480p default, 720p at $0.11/sec |
680
+ | `bytedance/seedance-1.5-pro` | $0.07/sec | 5s | 12s | Budget 720p with synced audio. No RealFace assets |
681
+ | `bytedance/seedance-2.0-fast` | $0.165/sec | 5s | 15s | 720p, ~60-80s to generate. RealFace assets supported |
682
+ | `bytedance/seedance-2.0-mini` | $0.0797/sec | 5s | 15s | 480p/720p at half the flagship rate. RealFace supported |
683
+ | `bytedance/seedance-2.0` | $0.227/sec | 5s | 15s | Premium 720p with synced audio. RealFace supported |
684
+ | `bytedance/seedance-2.5` | $0.315/sec | 5s | 30s | Long-form — up to 30s, multilingual, multi-asset |
685
+ | `azure/sora-2` | $0.10/sec | 4s | 12s | Sora 2 via Azure AI Foundry — 720p with synced audio; 4, 8 or 12s |
655
686
 
656
687
  ```ts
657
688
  import { VideoClient } from '@blockrun/llm';
@@ -714,8 +745,9 @@ synchronous (<1s for Flash).
714
745
  |-------|-------|-----------|-------|
715
746
  | `elevenlabs/flash-v2.5` | $0.05/1k chars | 40k chars | ~75ms latency, 32 languages (default) |
716
747
  | `elevenlabs/turbo-v2.5` | $0.05/1k chars | 40k chars | ~250ms latency, balanced quality |
717
- | `elevenlabs/multilingual-v2` | $0.10/1k chars | 10k chars | Long-form narration, audiobooks |
748
+ | `elevenlabs/multilingual-v2` | $0.10/1k chars | 10k chars | Long-form narration, audiobooks — 29 languages |
718
749
  | `elevenlabs/v3` | $0.10/1k chars | 5k chars | Max expressiveness, 70+ languages |
750
+ | `bytedance/seed-audio-1.0` | $0.30/1k chars | 3k chars | Prompt-directed — describe voice, emotion and staging in words |
719
751
  | `elevenlabs/sound-effects` | $0.05/generation | 1k chars | Sound effects up to 22s |
720
752
 
721
753
  ```ts
@@ -1030,14 +1062,14 @@ const response = await client.chat('openai/gpt-4o', 'Explain quantum computing')
1030
1062
  console.log(response);
1031
1063
 
1032
1064
  // With system prompt
1033
- const response2 = await client.chat('anthropic/claude-sonnet-4', 'Write a haiku', {
1065
+ const response2 = await client.chat('anthropic/claude-sonnet-5', 'Write a haiku', {
1034
1066
  system: 'You are a creative poet.',
1035
1067
  });
1036
1068
  ```
1037
1069
 
1038
1070
  ### Smart Routing (Router Core V3)
1039
1071
 
1040
- Save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. The bundled Router Core V3 engine classifies each request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions, applies hard capability filters, and ranks the cheapest capable models (<1ms, 100% local). Bundled — nothing extra to install.
1072
+ Save up to <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. The bundled Router Core V3 engine classifies each request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions, applies hard capability filters, and ranks the cheapest capable models (<1ms, 100% local). Bundled — nothing extra to install.
1041
1073
 
1042
1074
  ```typescript
1043
1075
  import { LLMClient } from '@blockrun/llm';
@@ -1056,15 +1088,15 @@ const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); /
1056
1088
  const auto = await client.smartChat('Code review', { routingProfile: 'auto' }); // Balanced (default)
1057
1089
  const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
1058
1090
 
1059
- // Guaranteed $0: call a free NVIDIA model directly
1060
- const free = await client.chat('nvidia/step-3.7-flash', 'Hello!');
1091
+ // Guaranteed $0: call a free model directly
1092
+ const free = await client.chat('nvidia/nemotron-3.5-lightning', 'Hello!');
1061
1093
  ```
1062
1094
 
1063
1095
  **Routing Profiles:**
1064
1096
 
1065
1097
  | Profile | Description | Best For |
1066
1098
  |---------|-------------|----------|
1067
- | `eco` | Budget-optimized — ranks the <!-- br:models.free -->5<!-- /br:models.free -->-model free NVIDIA tier first | Cost-sensitive workloads, zero-cost testing |
1099
+ | `eco` | Budget-optimized — ranks the <!-- br:models.free -->7<!-- /br:models.free -->-model free tier first | Cost-sensitive workloads, zero-cost testing |
1068
1100
  | `auto` | Intelligent routing (default) | General use |
1069
1101
  | `premium` | Best quality models | Critical tasks |
1070
1102
 
@@ -1183,7 +1215,7 @@ const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to se
1183
1215
 
1184
1216
  const [gpt, claude, gemini] = await Promise.all([
1185
1217
  client.chat('openai/gpt-4o', 'What is 2+2?'),
1186
- client.chat('anthropic/claude-sonnet-4', 'What is 3+3?'),
1218
+ client.chat('anthropic/claude-sonnet-5', 'What is 3+3?'),
1187
1219
  client.chat('google/gemini-2.5-flash', 'What is 4+4?'),
1188
1220
  ]);
1189
1221
  ```
@@ -1620,13 +1652,13 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
1620
1652
  ## Frequently Asked Questions
1621
1653
 
1622
1654
  ### What is @blockrun/llm?
1623
- @blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->70<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol — no API keys, no subscriptions, no vendor lock-in.
1655
+ @blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->74<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol — no API keys, no subscriptions, no vendor lock-in.
1624
1656
 
1625
1657
  ### How does payment work?
1626
1658
  When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
1627
1659
 
1628
1660
  ### What is smart routing?
1629
- Router Core V3 is bundled into the SDK — the same deterministic routing engine that powers ClawRouter, with nothing extra to install. It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. Use `smartChat()`, `smartChatCompletion()`, or the `blockrun/auto` model alias. It can save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
1661
+ Router Core V3 is bundled into the SDK — the same deterministic routing engine that powers ClawRouter, with nothing extra to install. It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. Use `smartChat()`, `smartChatCompletion()`, or the `blockrun/auto` model alias. It can save up to <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
1630
1662
 
1631
1663
  ### Does it support streaming?
1632
1664
  Yes — as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).