@blockrun/llm 3.13.2 → 3.13.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +219 -187
- package/dist/index.cjs +1121 -521
- package/dist/index.d.cts +16 -0
- package/dist/index.d.ts +16 -0
- package/dist/index.js +1121 -521
- package/package.json +12 -10
package/README.md
CHANGED
|
@@ -2,9 +2,9 @@
|
|
|
2
2
|
|
|
3
3
|
# @blockrun/llm
|
|
4
4
|
|
|
5
|
-
### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->
|
|
5
|
+
### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
|
|
6
6
|
|
|
7
|
-
The smart-routing SDK for <!-- br:models.chatVisible -->
|
|
7
|
+
The smart-routing SDK for <!-- br:models.chatVisible -->74<!-- /br:models.chatVisible --> models — every request goes to the cheapest model that can handle it,
|
|
8
8
|
paid per-request in USDC. No API keys. No subscriptions. No vendor lock-in.
|
|
9
9
|
|
|
10
10
|
[](https://www.npmjs.com/package/@blockrun/llm)
|
|
@@ -36,12 +36,12 @@ console.log(r.routing.savings); // 0.96 — this exact request cost 96% less th
|
|
|
36
36
|
console.log(r.response); // the proof
|
|
37
37
|
```
|
|
38
38
|
|
|
39
|
-
**<!-- br:savings.autoVsBaselinePct -->
|
|
39
|
+
**<!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** across a realistic workload on the default `auto` profile, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco` — and eco's first stop is the free tier, so simple requests cost $0.00 outright. Not an "up to" figure: the baseline, workload mix, and token ratio are published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json) so anyone can recompute the claim. Details in [Smart Routing](#smart-routing-router-core-v3).
|
|
40
40
|
|
|
41
41
|
## Why This SDK
|
|
42
42
|
|
|
43
43
|
- 🧠 **Smart routing that pays for itself** — the bundled [Router Core V3](https://github.com/BlockRunAI/router-core) engine (shared with [ClawRouter](https://github.com/BlockRunAI/ClawRouter)) classifies every request locally in <1ms across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and routes to the cheapest capable model. The main event.
|
|
44
|
-
- 🆓 **<!-- br:models.free -->
|
|
44
|
+
- 🆓 **<!-- br:models.free -->7<!-- /br:models.free --> genuinely free models** — $0 in and out, incl. two 1M-context Nemotrons, a multimodal one, and free coding models from Cohere and Poolside. No rate-limit gimmicks.
|
|
45
45
|
- 🔐 **No API keys** — your wallet signature is your authentication. No accounts, no dashboards, no key rotation.
|
|
46
46
|
- 💸 **Pay per request in USDC** — x402 micropayments on Base or Solana. $5 covers thousands of requests; agents can pay their own way.
|
|
47
47
|
- 🛡️ **Automatic failover** — transient errors (timeouts, 429, 5xx) walk the router's ranked fallback chain instead of failing your request.
|
|
@@ -52,9 +52,9 @@ console.log(r.response); // the proof
|
|
|
52
52
|
|
|
53
53
|
| | OpenAI SDK | OpenRouter | LiteLLM | **@blockrun/llm** |
|
|
54
54
|
| ------------------ | -------------- | ----------------- | ---------------- | ----------------------------------------------------------------------- |
|
|
55
|
-
| **Cost routing** | ✗ one vendor | Manual selection | Manual selection | **Automatic — <!-- br:savings.autoVsBaselinePct -->
|
|
56
|
-
| **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->
|
|
57
|
-
| **Free tier** | ✗ | Rate-limited | ✗ | **<!-- br:models.free -->
|
|
55
|
+
| **Cost routing** | ✗ one vendor | Manual selection | Manual selection | **Automatic — <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
|
|
56
|
+
| **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->74<!-- /br:models.chatVisible -->, one wallet** |
|
|
57
|
+
| **Free tier** | ✗ | Rate-limited | ✗ | **<!-- br:models.free -->7<!-- /br:models.free --> models, no signup** |
|
|
58
58
|
| **Auth** | API key | Account + API key | Your API keys | **Wallet signature** |
|
|
59
59
|
| **Payment** | Card + invoice | Credit card | BYO keys | **USDC per-request** |
|
|
60
60
|
| **Agent-ready** | ✗ | ✗ | ✗ | **✓ — agents fund their own wallet** |
|
|
@@ -115,7 +115,7 @@ package to install.
|
|
|
115
115
|
|
|
116
116
|
### Try It Free (No USDC Required)
|
|
117
117
|
|
|
118
|
-
Want to kick the tires before funding a wallet? Route to BlockRun's free
|
|
118
|
+
Want to kick the tires before funding a wallet? Route to BlockRun's free tier:
|
|
119
119
|
|
|
120
120
|
```typescript
|
|
121
121
|
import { LLMClient } from '@blockrun/llm';
|
|
@@ -123,33 +123,34 @@ import { LLMClient } from '@blockrun/llm';
|
|
|
123
123
|
const client = new LLMClient(); // Wallet still required for signing, but $0 charged
|
|
124
124
|
|
|
125
125
|
// Option 1: call a free model directly
|
|
126
|
-
const reply = await client.chat('nvidia/
|
|
126
|
+
const reply = await client.chat('nvidia/nemotron-3.5-lightning', 'Explain x402 in 1 sentence');
|
|
127
127
|
|
|
128
|
-
// Option 2: let the smart router pick — 'eco' ranks the free
|
|
128
|
+
// Option 2: let the smart router pick — 'eco' ranks the free tier first
|
|
129
129
|
const result = await client.smartChat('What is 2+2?', { routingProfile: 'eco' });
|
|
130
|
-
console.log(result.model); //
|
|
130
|
+
console.log(result.model); // a free-tier model — $0 in and out
|
|
131
131
|
console.log(result.response); // '4'
|
|
132
132
|
console.log(result.routing.savings); // 1 (100%)
|
|
133
133
|
```
|
|
134
134
|
|
|
135
135
|
There is no `free` routing profile in `smartChat()` — `routingProfile` accepts
|
|
136
136
|
`'eco' | 'auto' | 'premium'`. (ClawRouter's `/model free` is a feature of its
|
|
137
|
-
own proxy, not of this SDK's router options.) For guaranteed $0, pin
|
|
138
|
-
|
|
137
|
+
own proxy, not of this SDK's router options.) For guaranteed $0, pin one of
|
|
138
|
+
the free model ids below; for smart-routed $0-first, use `eco`.
|
|
139
139
|
|
|
140
|
-
**Available free models**
|
|
140
|
+
**Available free models** — input and output both $0. The free tier is **no
|
|
141
|
+
longer NVIDIA-only**, so pin these by full model id rather than by an
|
|
142
|
+
`nvidia/*` prefix. Full contexts and notes in [Free Tier](#free-tier);
|
|
143
|
+
`client.listModels()` returns the live catalog at runtime.
|
|
141
144
|
|
|
142
145
|
| Model ID | Context | Best For |
|
|
143
146
|
|----------|---------|----------|
|
|
144
|
-
| `nvidia/nemotron-3-
|
|
145
|
-
| `nvidia/
|
|
146
|
-
| `nvidia/
|
|
147
|
-
| `nvidia/nemotron-nano-
|
|
148
|
-
| `nvidia/
|
|
149
|
-
| `
|
|
150
|
-
| `
|
|
151
|
-
|
|
152
|
-
> Privacy note: `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` are hidden from `/v1/models` because NVIDIA's free build.nvidia.com tier reserves the right to use prompts/outputs for service improvement. Direct calls by full model ID still work — opt in only when your data isn't sensitive.
|
|
147
|
+
| `nvidia/nemotron-3.5-lightning` | 1M | Thinking-mode reasoning at 1M context |
|
|
148
|
+
| `nvidia/nemotron-3-ultra-550b` | 1M | Largest free model — 550B |
|
|
149
|
+
| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning — text + images |
|
|
150
|
+
| `nvidia/nemotron-3-nano-30b` | 128K | Compact + fast, good for high-volume light tasks |
|
|
151
|
+
| `nvidia/llama-3.2-11b-vision` | 128K | Vision-language — accepts images |
|
|
152
|
+
| `cohere/north-mini-code` | 256K | Compact coding model, sub-second responses |
|
|
153
|
+
| `poolside/laguna-xs-2.1` | 128K | Coding model |
|
|
153
154
|
|
|
154
155
|
## Quick Start (Solana)
|
|
155
156
|
|
|
@@ -166,7 +167,7 @@ Set `SOLANA_WALLET_KEY` to your bs58-encoded Solana secret key. Payments are aut
|
|
|
166
167
|
|
|
167
168
|
## Smart Routing (Router Core V3)
|
|
168
169
|
|
|
169
|
-
Let the SDK automatically pick the cheapest capable model for each request — **<!-- br:savings.autoVsBaselinePct -->
|
|
170
|
+
Let the SDK automatically pick the cheapest capable model for each request — **<!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** for the same traffic on `auto`, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco`.
|
|
170
171
|
|
|
171
172
|
Not an "up to" figure. The baseline, the workload mix and the token ratio are
|
|
172
173
|
published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json),
|
|
@@ -210,12 +211,12 @@ const client = new LLMClient();
|
|
|
210
211
|
// Auto-routes to cheapest capable model
|
|
211
212
|
const result = await client.smartChat('What is 2+2?');
|
|
212
213
|
console.log(result.response); // '4'
|
|
213
|
-
console.log(result.model); // '
|
|
214
|
-
console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); //
|
|
214
|
+
console.log(result.model); // 'qwen/qwen3.7-flash' (cheap, fast)
|
|
215
|
+
console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // this request, vs the Opus 5 baseline
|
|
215
216
|
|
|
216
217
|
// Complex reasoning task -> routes to reasoning model
|
|
217
218
|
const complex = await client.smartChat('Prove the Riemann hypothesis step by step');
|
|
218
|
-
console.log(complex.model); // 'xai/grok-4
|
|
219
|
+
console.log(complex.model); // 'xai/grok-4.3'
|
|
219
220
|
|
|
220
221
|
// Inspect how the request was classified and ranked (Router v3.4 portfolio).
|
|
221
222
|
console.log(complex.routing.method); // 'portfolio'
|
|
@@ -232,14 +233,16 @@ console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
|
|
|
232
233
|
`chat()` / `chatCompletion()` walk it automatically when the primary model
|
|
233
234
|
returns a transient error — timeouts, network failures, 429 rate limits, or
|
|
234
235
|
5xx responses (502/503/504/522/524). Other 4xx errors and `PaymentError`
|
|
235
|
-
propagate immediately so wallet / auth issues surface fast.
|
|
236
|
+
propagate immediately so wallet / auth issues surface fast. (Solana's internal
|
|
237
|
+
stale-blockhash re-sign is a separate, lower-level retry inside the payment
|
|
238
|
+
step — see [How Payment Works](#phase-2--every-request-pays-itself-automatic-x402).)
|
|
236
239
|
|
|
237
240
|
```typescript
|
|
238
241
|
// Manually pass a fallback chain to chat() / chatCompletion()
|
|
239
|
-
const reply = await client.chat('nvidia/
|
|
240
|
-
fallbackModels: ['nvidia/
|
|
242
|
+
const reply = await client.chat('nvidia/nemotron-3.5-lightning', 'hello', {
|
|
243
|
+
fallbackModels: ['nvidia/nemotron-3-nano-30b', 'cohere/north-mini-code'],
|
|
241
244
|
});
|
|
242
|
-
// If
|
|
245
|
+
// If nemotron-3.5-lightning times out, the SDK retries against the next model
|
|
243
246
|
// and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
|
|
244
247
|
```
|
|
245
248
|
|
|
@@ -247,11 +250,11 @@ const reply = await client.chat('nvidia/step-3.7-flash', 'hello', {
|
|
|
247
250
|
|
|
248
251
|
| Profile | Strategy | Savings vs Opus 5 | Best For |
|
|
249
252
|
|---------|----------|-------------------|----------|
|
|
250
|
-
| `eco` | Cheapest capable model — ranks the <!-- br:models.free -->
|
|
251
|
-
| `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->
|
|
253
|
+
| `eco` | Cheapest capable model — ranks the <!-- br:models.free -->7<!-- /br:models.free -->-model free tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
|
|
254
|
+
| `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->%** | General use |
|
|
252
255
|
| `premium` | Top-tier models (OpenAI, Anthropic) | 0% | Quality-critical tasks |
|
|
253
256
|
|
|
254
|
-
For guaranteed $0, call a
|
|
257
|
+
For guaranteed $0, call a free model directly with `chat()` — see
|
|
255
258
|
[Try It Free](#try-it-free-no-usdc-required). ClawRouter's `/model free`
|
|
256
259
|
profile belongs to its own proxy; `smartChat()`'s options are the three above.
|
|
257
260
|
|
|
@@ -293,14 +296,14 @@ anchors the portfolio's candidate pool):
|
|
|
293
296
|
|
|
294
297
|
| Tier | Example Tasks | ECO | AUTO | PREMIUM |
|
|
295
298
|
|------|---------------|-----|------|---------|
|
|
296
|
-
| SIMPLE | "What is 2+2?", definitions |
|
|
297
|
-
| MEDIUM | Code snippets, explanations |
|
|
298
|
-
| COMPLEX | Architecture, long documents |
|
|
299
|
-
| REASONING | Proofs, multi-step reasoning |
|
|
299
|
+
| SIMPLE | "What is 2+2?", definitions | step-3.7-flash (**FREE**) | gemini-2.5-flash ($0.30/$2.50) | gemini-3.5-flash ($1.50/$9) |
|
|
300
|
+
| MEDIUM | Code snippets, explanations | glm-5.3-flash ($0.15/$0.50) | gemini-3.5-flash ($1.50/$9) | gpt-5.3-codex ($1.75/$14.00) |
|
|
301
|
+
| COMPLEX | Architecture, long documents | glm-5.3-flash ($0.15/$0.50) | gemini-3.1-pro ($2/$12) | claude-fable-5 ($10/$50) |
|
|
302
|
+
| REASONING | Proofs, multi-step reasoning | deepseek-reasoner ($0.14/$0.28) | deepseek-reasoner ($0.14/$0.28) | claude-sonnet-5 ($3/$15) |
|
|
300
303
|
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
priced on visible models
|
|
304
|
+
Since Router Core V3.5 every primary and every fallback rung is a model listed
|
|
305
|
+
on `/v1/models` — nothing the router picks is withheld from the public pricing
|
|
306
|
+
page, and the published savings claim is priced on those same visible models.
|
|
304
307
|
|
|
305
308
|
This table mirrors ClawRouter's tier configs at the version this SDK pins;
|
|
306
309
|
the [ClawRouter README](https://github.com/BlockRunAI/ClawRouter#how-it-works)
|
|
@@ -367,7 +370,7 @@ const response = await client.chat('openai/gpt-4o', 'gm Solana');
|
|
|
367
370
|
console.log(response);
|
|
368
371
|
|
|
369
372
|
// Live Search with Grok (Solana payment)
|
|
370
|
-
const tweet = await client.chat('xai/grok-
|
|
373
|
+
const tweet = await client.chat('xai/grok-4.5', 'What is trending on X?', { search: true });
|
|
371
374
|
```
|
|
372
375
|
|
|
373
376
|
**Setup:**
|
|
@@ -395,7 +398,7 @@ You only do this when your balance runs low. Three ways to get USDC into your wa
|
|
|
395
398
|
|
|
396
399
|
- **(b) Transfer existing USDC.** Send USDC you already hold to your wallet address (`client.getWalletAddress()`). On Base, send Base USDC; on Solana (`SolanaLLMClient`), send Solana SPL USDC.
|
|
397
400
|
|
|
398
|
-
- **(c) Skip funding entirely.** Use the free
|
|
401
|
+
- **(c) Skip funding entirely.** Use the free models (e.g. `nvidia/nemotron-3.5-lightning`) — every call is **$0**, no balance required.
|
|
399
402
|
|
|
400
403
|
$5 of USDC covers thousands of paid requests. Check your balance any time:
|
|
401
404
|
|
|
@@ -414,7 +417,19 @@ You just call e.g. `client.chat(...)` — the payment is invisible:
|
|
|
414
417
|
4. The request is retried automatically with the payment proof.
|
|
415
418
|
5. The gateway settles on-chain and returns the AI response.
|
|
416
419
|
|
|
417
|
-
One call, no separate pay step. Free
|
|
420
|
+
One call, no separate pay step. Free-tier models settle at **$0** (no payment signed).
|
|
421
|
+
|
|
422
|
+
On **Solana**, step 3 pins the payment to a recent blockhash that is valid for
|
|
423
|
+
roughly 60 seconds. If one expires between signing and verification,
|
|
424
|
+
`SolanaLLMClient` re-signs against a fresh blockhash and retries — up to twice,
|
|
425
|
+
with a short backoff — rather than surfacing a payment error you would only
|
|
426
|
+
have to retry by hand.
|
|
427
|
+
|
|
428
|
+
The retry is deliberately narrow. It fires only when the gateway explicitly
|
|
429
|
+
reports a *verification-phase* stale blockhash. Settlement failures, ambiguous
|
|
430
|
+
rejections, insufficient funds and malformed responses all fail immediately.
|
|
431
|
+
Verification runs strictly before settlement, so a retryable rejection means no
|
|
432
|
+
transaction was broadcast and you cannot be charged twice.
|
|
418
433
|
|
|
419
434
|
### Track spend and verify settlements
|
|
420
435
|
|
|
@@ -460,7 +475,7 @@ const video = await br.poll('/v1/videos/generations', {
|
|
|
460
475
|
|
|
461
476
|
// Streaming SSE — chat completions
|
|
462
477
|
for await (const chunk of br.stream('/v1/chat/completions', {
|
|
463
|
-
model: 'anthropic/claude-sonnet-
|
|
478
|
+
model: 'anthropic/claude-sonnet-5',
|
|
464
479
|
messages: [{ role: 'user', content: 'Hi' }],
|
|
465
480
|
stream: true,
|
|
466
481
|
})) {
|
|
@@ -481,157 +496,169 @@ shims over `BlockrunClient`) and removed in 3.0.
|
|
|
481
496
|
|
|
482
497
|
## Available Models
|
|
483
498
|
|
|
484
|
-
|
|
485
|
-
|
|
486
|
-
|
|
487
|
-
|
|
488
|
-
|
|
489
|
-
|
|
490
|
-
|
|
491
|
-
|
|
492
|
-
|
|
493
|
-
|
|
494
|
-
|
|
495
|
-
|
|
|
496
|
-
|
|
497
|
-
|
|
498
|
-
|
|
499
|
-
|
|
|
500
|
-
|
|
501
|
-
| `openai/gpt-5.
|
|
502
|
-
| `openai/gpt-5.
|
|
503
|
-
|
|
504
|
-
|
|
505
|
-
|
|
499
|
+
Prices below are the live gateway rates, regenerated from `GET /v1/models`
|
|
500
|
+
(the same catalog `client.listModels()` returns). Chat models are billed per
|
|
501
|
+
token; image, video, music and speech are billed per unit as noted in their
|
|
502
|
+
own sections.
|
|
503
|
+
|
|
504
|
+
### OpenAI GPT-5.6 Family
|
|
505
|
+
|
|
506
|
+
Three tiers on one 1.05M-context base — Sol (deepest reasoning), Terra
|
|
507
|
+
(balanced), Luna (cheap and fast). Each has a `-pro` sibling that thinks
|
|
508
|
+
longer at the same token price.
|
|
509
|
+
|
|
510
|
+
| Model | Input Price | Output Price | Context |
|
|
511
|
+
|-------|-------------|--------------|---------|
|
|
512
|
+
| `openai/gpt-5.6-sol` | $5.00/M | $30.00/M | 1.05M |
|
|
513
|
+
| `openai/gpt-5.6-sol-pro` | $5.00/M | $30.00/M | 1.05M |
|
|
514
|
+
| `openai/gpt-5.6-terra` | $2.00/M | $12.00/M | 1.05M |
|
|
515
|
+
| `openai/gpt-5.6-terra-pro` | $2.00/M | $12.00/M | 1.05M |
|
|
516
|
+
| `openai/gpt-5.6-luna` | $0.20/M | $1.20/M | 1.05M |
|
|
517
|
+
| `openai/gpt-5.6-luna-pro` | $0.20/M | $1.20/M | 1.05M |
|
|
518
|
+
|
|
519
|
+
### OpenAI GPT-5.5 / 5.4 / 5.2 Families
|
|
520
|
+
|
|
521
|
+
| Model | Input Price | Output Price | Context | Notes |
|
|
522
|
+
|-------|-------------|--------------|---------|-------|
|
|
523
|
+
| `openai/gpt-5.5` | $5.00/M | $30.00/M | 1.05M | |
|
|
524
|
+
| `openai/gpt-5.5-pro` | $30.00/M | $180.00/M | 1.05M | |
|
|
525
|
+
| `openai/chat-latest` | $5.00/M | $30.00/M | 128K | ChatGPT Instant — the model behind chatgpt.com |
|
|
526
|
+
| `openai/gpt-5.4` | $2.50/M | $15.00/M | 1.05M | |
|
|
527
|
+
| `openai/gpt-5.4-pro` | $30.00/M | $180.00/M | 1.05M | |
|
|
528
|
+
| `openai/gpt-5.4-mini` | $0.75/M | $4.50/M | 400K | |
|
|
529
|
+
| `openai/gpt-5.4-nano` | $0.20/M | $1.25/M | 1.05M | |
|
|
530
|
+
| `openai/gpt-5.2` | $1.75/M | $14.00/M | 400K | |
|
|
531
|
+
| `openai/gpt-5.2-pro` | $21.00/M | $168.00/M | 400K | |
|
|
532
|
+
| `openai/gpt-5.3-codex` | $1.75/M | $14.00/M | 400K | Coding/agentic SKU |
|
|
533
|
+
| `openai/gpt-5-mini` | $0.25/M | $2.00/M | 200K | |
|
|
506
534
|
|
|
507
535
|
### OpenAI GPT-4 Family
|
|
508
|
-
|
|
509
|
-
|
|
510
|
-
|
|
511
|
-
| `openai/gpt-4.1
|
|
512
|
-
| `openai/gpt-4.1-
|
|
513
|
-
| `openai/gpt-
|
|
514
|
-
| `openai/gpt-4o
|
|
536
|
+
|
|
537
|
+
| Model | Input Price | Output Price | Context |
|
|
538
|
+
|-------|-------------|--------------|---------|
|
|
539
|
+
| `openai/gpt-4.1` | $2.00/M | $8.00/M | 128K |
|
|
540
|
+
| `openai/gpt-4.1-mini` | $0.40/M | $1.60/M | 128K |
|
|
541
|
+
| `openai/gpt-4.1-nano` | $0.10/M | $0.40/M | 128K |
|
|
542
|
+
| `openai/gpt-4o` | $2.50/M | $10.00/M | 128K |
|
|
543
|
+
| `openai/gpt-4o-mini` | $0.15/M | $0.60/M | 128K |
|
|
515
544
|
|
|
516
545
|
### OpenAI O-Series (Reasoning)
|
|
517
|
-
|
|
518
|
-
|
|
519
|
-
|
|
520
|
-
| `openai/
|
|
521
|
-
| `openai/o3
|
|
522
|
-
| `openai/
|
|
546
|
+
|
|
547
|
+
| Model | Input Price | Output Price | Context |
|
|
548
|
+
|-------|-------------|--------------|---------|
|
|
549
|
+
| `openai/o1` | $15.00/M | $60.00/M | 200K |
|
|
550
|
+
| `openai/o3` | $2.00/M | $8.00/M | 200K |
|
|
551
|
+
| `openai/o3-mini` | $1.10/M | $4.40/M | 128K |
|
|
552
|
+
| `openai/o4-mini` | $1.10/M | $4.40/M | 128K |
|
|
523
553
|
|
|
524
554
|
### Anthropic Claude
|
|
555
|
+
|
|
525
556
|
| Model | Input Price | Output Price | Context | Notes |
|
|
526
557
|
|-------|-------------|--------------|---------|-------|
|
|
527
|
-
| `anthropic/claude-fable-5` | $10.00/M | $50.00/M |
|
|
528
|
-
| `anthropic/claude-opus-
|
|
529
|
-
| `anthropic/claude-opus-4.
|
|
530
|
-
| `anthropic/claude-opus-4.
|
|
531
|
-
| `anthropic/claude-opus-4.5` | $5.00/M | $25.00/M | 200K |
|
|
532
|
-
| `anthropic/claude-
|
|
533
|
-
| `anthropic/claude-sonnet-4.6` | $3.00/M | $15.00/M |
|
|
534
|
-
| `anthropic/claude-sonnet-4` | $3.00/M | $15.00/M | 200K |
|
|
535
|
-
| `anthropic/claude-haiku-4.5` | $1.00/M | $5.00/M | 200K |
|
|
558
|
+
| `anthropic/claude-fable-5` | $10.00/M | $50.00/M | 1M | Mythos-class flagship above Opus — always-on thinking, 128K output |
|
|
559
|
+
| `anthropic/claude-opus-5` | $5.00/M | $25.00/M | 1M | Flagship — the baseline the routing savings claim is measured against |
|
|
560
|
+
| `anthropic/claude-opus-4.8` | $5.00/M | $25.00/M | 1M | Agentic coding + adaptive thinking, 128K output |
|
|
561
|
+
| `anthropic/claude-opus-4.7` | $5.00/M | $25.00/M | 1M | |
|
|
562
|
+
| `anthropic/claude-opus-4.5` | $5.00/M | $25.00/M | 200K | |
|
|
563
|
+
| `anthropic/claude-sonnet-5` | $3.00/M | $15.00/M | 1M | Best cost/quality balance for long-context agent turns |
|
|
564
|
+
| `anthropic/claude-sonnet-4.6` | $3.00/M | $15.00/M | 1M | |
|
|
565
|
+
| `anthropic/claude-sonnet-4.5` | $3.00/M | $15.00/M | 200K | |
|
|
566
|
+
| `anthropic/claude-haiku-4.5` | $1.00/M | $5.00/M | 200K | |
|
|
536
567
|
|
|
537
568
|
### Google Gemini
|
|
538
|
-
|
|
539
|
-
|
|
540
|
-
|
|
541
|
-
| `google/gemini-3.
|
|
542
|
-
| `google/gemini-3.
|
|
543
|
-
| `google/gemini-3-flash
|
|
544
|
-
| `google/gemini-
|
|
545
|
-
| `google/gemini-
|
|
546
|
-
| `google/gemini-
|
|
569
|
+
|
|
570
|
+
| Model | Input Price | Output Price | Context |
|
|
571
|
+
|-------|-------------|--------------|---------|
|
|
572
|
+
| `google/gemini-3.1-pro` | $2.00/M | $12.00/M | 1M |
|
|
573
|
+
| `google/gemini-3.6-flash` | $1.50/M | $7.50/M | 1M |
|
|
574
|
+
| `google/gemini-3.5-flash` | $1.50/M | $9.00/M | 1M |
|
|
575
|
+
| `google/gemini-3-flash-preview` | $0.50/M | $3.00/M | 1M |
|
|
576
|
+
| `google/gemini-3.5-flash-lite` | $0.30/M | $2.50/M | 1M |
|
|
577
|
+
| `google/gemini-3.1-flash-lite` | $0.25/M | $1.50/M | 1M |
|
|
578
|
+
| `google/gemini-2.5-pro` | $1.25/M | $10.00/M | 1M |
|
|
579
|
+
| `google/gemini-2.5-flash` | $0.30/M | $2.50/M | 1M |
|
|
580
|
+
| `google/gemini-2.5-flash-lite` | $0.10/M | $0.40/M | 1M |
|
|
547
581
|
|
|
548
582
|
### DeepSeek
|
|
549
583
|
|
|
550
|
-
|
|
551
|
-
|
|
552
|
-
|
|
553
|
-
1M context, MMLU-Pro 87.5, GPQA 90.1, SWE-bench 80.6, LiveCodeBench 93.5.
|
|
584
|
+
DeepSeek upstream serves the legacy `deepseek-chat` / `deepseek-reasoner`
|
|
585
|
+
aliases as V4 Flash non-thinking / thinking modes. V4 Pro is the flagship
|
|
586
|
+
paid SKU; the vision SKU is an experimental preview.
|
|
554
587
|
|
|
555
588
|
| Model | Input Price | Output Price | Context | Notes |
|
|
556
589
|
|-------|-------------|--------------|---------|-------|
|
|
557
|
-
| `deepseek/deepseek-v4-pro` | $
|
|
558
|
-
| `deepseek/deepseek-
|
|
559
|
-
| `deepseek/deepseek-
|
|
590
|
+
| `deepseek/deepseek-v4-pro` | $1.32/M | $3.96/M | 1M | V4 flagship — strongest open-weight reasoner |
|
|
591
|
+
| `deepseek/deepseek-v4-flash-vision-exp` | $0.44/M | $1.32/M | 1M | Experimental vision preview |
|
|
592
|
+
| `deepseek/deepseek-chat` | $0.14/M | $0.28/M | 1M | V4 Flash non-thinking |
|
|
593
|
+
| `deepseek/deepseek-reasoner` | $0.14/M | $0.28/M | 1M | V4 Flash thinking (same upstream, thinking on by default) |
|
|
560
594
|
|
|
561
595
|
### xAI Grok
|
|
562
596
|
|
|
563
|
-
|
|
564
|
-
|
|
565
|
-
|
|
566
|
-
|
|
567
|
-
`/v1/models`** — direct calls by full ID still work, but SmartChat won't
|
|
568
|
-
auto-pick them.
|
|
597
|
+
The older Grok chat SKUs (grok-3/3-mini, the grok-4 fast families,
|
|
598
|
+
grok-code-fast-1, grok-2-vision) have left the catalog. Retired ids stay
|
|
599
|
+
callable — the gateway redirects them to a healthy model — but SmartChat
|
|
600
|
+
only ranks what `/v1/models` lists.
|
|
569
601
|
|
|
570
602
|
| Model | Input Price | Output Price | Context | Notes |
|
|
571
603
|
|-------|-------------|--------------|---------|-------|
|
|
572
|
-
| `xai/grok-4.
|
|
573
|
-
| `xai/grok-
|
|
574
|
-
|
|
575
|
-
|
|
576
|
-
|
|
577
|
-
|
|
578
|
-
|
|
|
579
|
-
|
|
580
|
-
|
|
581
|
-
|
|
582
|
-
|
|
|
583
|
-
|
|
584
|
-
| `
|
|
585
|
-
| `
|
|
586
|
-
|
|
587
|
-
|
|
588
|
-
|
|
589
|
-
|
|
590
|
-
|
|
591
|
-
|
|
592
|
-
|
|
593
|
-
|
|
594
|
-
|
|
595
|
-
|
|
596
|
-
|
|
597
|
-
|
|
598
|
-
|
|
|
599
|
-
|
|
600
|
-
| `
|
|
601
|
-
|
|
602
|
-
|
|
603
|
-
|
|
604
|
-
|
|
605
|
-
|
|
606
|
-
|
|
607
|
-
|
|
608
|
-
|
|
609
|
-
|
|
610
|
-
|
|
611
|
-
|
|
612
|
-
|
|
613
|
-
|
|
|
614
|
-
|
|
615
|
-
|
|
|
616
|
-
|
|
|
617
|
-
| Anthropic | `anthropic/claude-opus-4.6` | Passed |
|
|
618
|
-
| Anthropic | `anthropic/claude-sonnet-4` | Passed |
|
|
619
|
-
| Google | `google/gemini-2.5-flash` | Passed |
|
|
620
|
-
| DeepSeek | `deepseek/deepseek-chat` | Passed |
|
|
621
|
-
| xAI | `xai/grok-3` | Passed |
|
|
622
|
-
| Moonshot | `moonshot/kimi-k2.6` | Passed |
|
|
604
|
+
| `xai/grok-4.5` | $2.50/M | $9.00/M | 500K | Flagship — reasoning + vision, native Live Search (`search: true`) |
|
|
605
|
+
| `xai/grok-4.3` | $1.50/M | $4.00/M | 1M | Reasoning + vision, tuned for agentic workflows |
|
|
606
|
+
| `xai/grok-build-0.1` | $1.50/M | $3.00/M | 256K | Fast agentic coding model |
|
|
607
|
+
|
|
608
|
+
### Moonshot, MiniMax, Z.ai, Qwen
|
|
609
|
+
|
|
610
|
+
| Model | Input Price | Output Price | Context | Notes |
|
|
611
|
+
|-------|-------------|--------------|---------|-------|
|
|
612
|
+
| `moonshot/kimi-k3` | $3.00/M | $15.00/M | 1M | Replaces the retired `kimi-k2.5` / `k2.6` SKUs |
|
|
613
|
+
| `minimax/minimax-m3` | $0.30/M | $1.20/M | 1M | |
|
|
614
|
+
| `minimax/minimax-m2.7` | $0.30/M | $1.20/M | 200K | |
|
|
615
|
+
| `zai/glm-5.3` | $1.40/M | $4.40/M | 1M | |
|
|
616
|
+
| `zai/glm-5.3-flash` | $0.15/M | $0.50/M | 1M | Cheapest vision-capable paid SKU |
|
|
617
|
+
| `zai/glm-5.2` | $1.40/M | $4.40/M | 1M | |
|
|
618
|
+
| `zai/glm-5.1` | $1.40/M | $4.40/M | 200K | |
|
|
619
|
+
| `zai/glm-5` | $1.00/M | $3.20/M | 200K | |
|
|
620
|
+
| `zai/glm-5-turbo` | $1.20/M | $4.00/M | 200K | |
|
|
621
|
+
| `qwen/qwen3.7-max` | $1.475/M | $4.425/M | 1M | |
|
|
622
|
+
| `qwen/qwen3.7-plus` | $0.32/M | $1.28/M | 1M | |
|
|
623
|
+
| `qwen/qwen3.8-flash` | $0.15/M | $0.47/M | 1M | |
|
|
624
|
+
| `qwen/qwen3.7-flash` | $0.03/M | $0.13/M | 1M | Cheapest paid chat model in the catalog |
|
|
625
|
+
|
|
626
|
+
### Tencent, Xiaomi
|
|
627
|
+
|
|
628
|
+
| Model | Input Price | Output Price | Context |
|
|
629
|
+
|-------|-------------|--------------|---------|
|
|
630
|
+
| `tencent/hy3` | $0.132/M | $0.528/M | 256K |
|
|
631
|
+
| `xiaomi/mimo-v2.5` | $0.14/M | $0.28/M | 1M |
|
|
632
|
+
| `xiaomi/mimo-v2.5-pro` | $0.435/M | $0.87/M | 1M |
|
|
633
|
+
|
|
634
|
+
### Free Tier
|
|
635
|
+
|
|
636
|
+
Input and output both $0 — no promo, no rate-limit gimmick. The free tier is
|
|
637
|
+
**no longer NVIDIA-only**, so pin these by full model id rather than by an
|
|
638
|
+
`nvidia/*` prefix, or let `routingProfile: 'eco'` rank them first.
|
|
639
|
+
|
|
640
|
+
| Model | Input Price | Output Price | Context | Notes |
|
|
641
|
+
|-------|-------------|--------------|---------|-------|
|
|
642
|
+
| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 256K | Multimodal reasoning — text + images |
|
|
643
|
+
| `nvidia/nemotron-3.5-lightning` | **FREE** | **FREE** | 1M | Thinking-mode reasoning at 1M context |
|
|
644
|
+
| `nvidia/nemotron-3-nano-30b` | **FREE** | **FREE** | 128K | Compact and fast, good for high-volume light tasks |
|
|
645
|
+
| `nvidia/llama-3.2-11b-vision` | **FREE** | **FREE** | 128K | Vision-language — accepts images |
|
|
646
|
+
| `nvidia/nemotron-3-ultra-550b` | **FREE** | **FREE** | 1M | Largest free model — 550B, 1M context |
|
|
647
|
+
| `cohere/north-mini-code` | **FREE** | **FREE** | 256K | Compact coding model, sub-second responses |
|
|
648
|
+
| `poolside/laguna-xs-2.1` | **FREE** | **FREE** | 128K | Coding model |
|
|
623
649
|
|
|
624
650
|
### Image Generation
|
|
625
|
-
| Model | Price |
|
|
626
|
-
|
|
627
|
-
| `openai/
|
|
628
|
-
| `openai/gpt-image-
|
|
629
|
-
| `
|
|
630
|
-
| `google/nano-banana` | $0.
|
|
631
|
-
| `google/nano-banana-pro` | $0.10
|
|
632
|
-
| `xai/grok-imagine-image` | $0.02/image |
|
|
633
|
-
| `xai/grok-imagine-image-pro` | $0.07/image |
|
|
634
|
-
| `
|
|
651
|
+
| Model | Price | Notes |
|
|
652
|
+
|-------|-------|-------|
|
|
653
|
+
| `openai/gpt-image-1` | $0.02/image | Native GPT-4o image generation |
|
|
654
|
+
| `openai/gpt-image-2` | $0.06/image | Reasoning-driven — multilingual text rendering, character consistency |
|
|
655
|
+
| `google/nano-banana` | $0.05/image | Gemini 2.5 Flash image generation — fast and efficient |
|
|
656
|
+
| `google/nano-banana-2` | $0.09/image | Gemini 3.1 Flash — pro-level quality at Flash speed |
|
|
657
|
+
| `google/nano-banana-pro` | $0.10/image | Gemini 3 Pro — highest quality, up to 4K |
|
|
658
|
+
| `xai/grok-imagine-image` | $0.02/image | Fast, 300 RPM |
|
|
659
|
+
| `xai/grok-imagine-image-pro` | $0.07/image | Quality tier, 30 RPM |
|
|
660
|
+
| `bytedance/seedream-5-pro` | $0.045/image | Flagship generation + editing, up to 4K-class, reference images |
|
|
661
|
+
| `zai/cogview-4` | $0.015/image | Up to 1440x1440 |
|
|
635
662
|
|
|
636
663
|
Image editing (`client.edit`) via `/v1/images/image2image`: `openai/gpt-image-1`, `openai/gpt-image-2`, `google/nano-banana`, and `google/nano-banana-pro`. Pass a single base64 `data:image/...` URI to edit one image, or an array of 2–4 URIs to **fuse** them (e.g. a subject + a brand logo). Fusion caps: `openai/*` up to 4 source images, `google/*` up to 3. A `mask` cannot be combined with multiple source images.
|
|
637
664
|
|
|
@@ -646,12 +673,16 @@ console.log(fused.data[0].url);
|
|
|
646
673
|
```
|
|
647
674
|
|
|
648
675
|
### Video Generation
|
|
649
|
-
| Model | Price |
|
|
650
|
-
|
|
651
|
-
| `xai/grok-imagine-video` | $0.05/sec
|
|
652
|
-
| `
|
|
653
|
-
| `bytedance/seedance-
|
|
654
|
-
| `bytedance/seedance-2.0` | $0.
|
|
676
|
+
| Model | Price | Default | Max | Notes |
|
|
677
|
+
|-------|-------|---------|-----|-------|
|
|
678
|
+
| `xai/grok-imagine-video` | $0.05/sec | 8s | 15s | 480p default, 720p at $0.07/sec; text or image to video |
|
|
679
|
+
| `xai/grok-imagine-video-1.5` | $0.08/sec | 8s | 15s | Flagship — native synced audio; 480p default, 720p at $0.11/sec |
|
|
680
|
+
| `bytedance/seedance-1.5-pro` | $0.07/sec | 5s | 12s | Budget 720p with synced audio. No RealFace assets |
|
|
681
|
+
| `bytedance/seedance-2.0-fast` | $0.165/sec | 5s | 15s | 720p, ~60-80s to generate. RealFace assets supported |
|
|
682
|
+
| `bytedance/seedance-2.0-mini` | $0.0797/sec | 5s | 15s | 480p/720p at half the flagship rate. RealFace supported |
|
|
683
|
+
| `bytedance/seedance-2.0` | $0.227/sec | 5s | 15s | Premium 720p with synced audio. RealFace supported |
|
|
684
|
+
| `bytedance/seedance-2.5` | $0.315/sec | 5s | 30s | Long-form — up to 30s, multilingual, multi-asset |
|
|
685
|
+
| `azure/sora-2` | $0.10/sec | 4s | 12s | Sora 2 via Azure AI Foundry — 720p with synced audio; 4, 8 or 12s |
|
|
655
686
|
|
|
656
687
|
```ts
|
|
657
688
|
import { VideoClient } from '@blockrun/llm';
|
|
@@ -714,8 +745,9 @@ synchronous (<1s for Flash).
|
|
|
714
745
|
|-------|-------|-----------|-------|
|
|
715
746
|
| `elevenlabs/flash-v2.5` | $0.05/1k chars | 40k chars | ~75ms latency, 32 languages (default) |
|
|
716
747
|
| `elevenlabs/turbo-v2.5` | $0.05/1k chars | 40k chars | ~250ms latency, balanced quality |
|
|
717
|
-
| `elevenlabs/multilingual-v2` | $0.10/1k chars | 10k chars | Long-form narration, audiobooks |
|
|
748
|
+
| `elevenlabs/multilingual-v2` | $0.10/1k chars | 10k chars | Long-form narration, audiobooks — 29 languages |
|
|
718
749
|
| `elevenlabs/v3` | $0.10/1k chars | 5k chars | Max expressiveness, 70+ languages |
|
|
750
|
+
| `bytedance/seed-audio-1.0` | $0.30/1k chars | 3k chars | Prompt-directed — describe voice, emotion and staging in words |
|
|
719
751
|
| `elevenlabs/sound-effects` | $0.05/generation | 1k chars | Sound effects up to 22s |
|
|
720
752
|
|
|
721
753
|
```ts
|
|
@@ -1030,14 +1062,14 @@ const response = await client.chat('openai/gpt-4o', 'Explain quantum computing')
|
|
|
1030
1062
|
console.log(response);
|
|
1031
1063
|
|
|
1032
1064
|
// With system prompt
|
|
1033
|
-
const response2 = await client.chat('anthropic/claude-sonnet-
|
|
1065
|
+
const response2 = await client.chat('anthropic/claude-sonnet-5', 'Write a haiku', {
|
|
1034
1066
|
system: 'You are a creative poet.',
|
|
1035
1067
|
});
|
|
1036
1068
|
```
|
|
1037
1069
|
|
|
1038
1070
|
### Smart Routing (Router Core V3)
|
|
1039
1071
|
|
|
1040
|
-
Save up to <!-- br:savings.autoVsBaselinePct -->
|
|
1072
|
+
Save up to <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. The bundled Router Core V3 engine classifies each request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions, applies hard capability filters, and ranks the cheapest capable models (<1ms, 100% local). Bundled — nothing extra to install.
|
|
1041
1073
|
|
|
1042
1074
|
```typescript
|
|
1043
1075
|
import { LLMClient } from '@blockrun/llm';
|
|
@@ -1056,15 +1088,15 @@ const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); /
|
|
|
1056
1088
|
const auto = await client.smartChat('Code review', { routingProfile: 'auto' }); // Balanced (default)
|
|
1057
1089
|
const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
|
|
1058
1090
|
|
|
1059
|
-
// Guaranteed $0: call a free
|
|
1060
|
-
const free = await client.chat('nvidia/
|
|
1091
|
+
// Guaranteed $0: call a free model directly
|
|
1092
|
+
const free = await client.chat('nvidia/nemotron-3.5-lightning', 'Hello!');
|
|
1061
1093
|
```
|
|
1062
1094
|
|
|
1063
1095
|
**Routing Profiles:**
|
|
1064
1096
|
|
|
1065
1097
|
| Profile | Description | Best For |
|
|
1066
1098
|
|---------|-------------|----------|
|
|
1067
|
-
| `eco` | Budget-optimized — ranks the <!-- br:models.free -->
|
|
1099
|
+
| `eco` | Budget-optimized — ranks the <!-- br:models.free -->7<!-- /br:models.free -->-model free tier first | Cost-sensitive workloads, zero-cost testing |
|
|
1068
1100
|
| `auto` | Intelligent routing (default) | General use |
|
|
1069
1101
|
| `premium` | Best quality models | Critical tasks |
|
|
1070
1102
|
|
|
@@ -1183,7 +1215,7 @@ const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to se
|
|
|
1183
1215
|
|
|
1184
1216
|
const [gpt, claude, gemini] = await Promise.all([
|
|
1185
1217
|
client.chat('openai/gpt-4o', 'What is 2+2?'),
|
|
1186
|
-
client.chat('anthropic/claude-sonnet-
|
|
1218
|
+
client.chat('anthropic/claude-sonnet-5', 'What is 3+3?'),
|
|
1187
1219
|
client.chat('google/gemini-2.5-flash', 'What is 4+4?'),
|
|
1188
1220
|
]);
|
|
1189
1221
|
```
|
|
@@ -1620,13 +1652,13 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
|
|
|
1620
1652
|
## Frequently Asked Questions
|
|
1621
1653
|
|
|
1622
1654
|
### What is @blockrun/llm?
|
|
1623
|
-
@blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->
|
|
1655
|
+
@blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->74<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol — no API keys, no subscriptions, no vendor lock-in.
|
|
1624
1656
|
|
|
1625
1657
|
### How does payment work?
|
|
1626
1658
|
When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
|
|
1627
1659
|
|
|
1628
1660
|
### What is smart routing?
|
|
1629
|
-
Router Core V3 is bundled into the SDK — the same deterministic routing engine that powers ClawRouter, with nothing extra to install. It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. Use `smartChat()`, `smartChatCompletion()`, or the `blockrun/auto` model alias. It can save up to <!-- br:savings.autoVsBaselinePct -->
|
|
1661
|
+
Router Core V3 is bundled into the SDK — the same deterministic routing engine that powers ClawRouter, with nothing extra to install. It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. Use `smartChat()`, `smartChatCompletion()`, or the `blockrun/auto` model alias. It can save up to <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
|
|
1630
1662
|
|
|
1631
1663
|
### Does it support streaming?
|
|
1632
1664
|
Yes — as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).
|