@blockrun/llm 3.10.0 โ 3.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +336 -157
- package/dist/index.cjs +3601 -84
- package/dist/index.d.cts +233 -31
- package/dist/index.d.ts +233 -31
- package/dist/index.js +3601 -84
- package/package.json +4 -7
package/README.md
CHANGED
|
@@ -1,36 +1,72 @@
|
|
|
1
|
-
|
|
1
|
+
<div align="center">
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
>
|
|
5
|
-
> ๐ **Includes 7 fully-free NVIDIA-hosted models** (5 visible in `/v1/models`, 2 hidden but directly callable) โ DeepSeek V4 Flash (1M context), Nemotron Nano Omni (vision), Qwen3 Coder, Llama 4, Mistral, plus the gpt-oss pair. Zero USDC, no rate-limit gimmicks. Use `routingProfile: 'free'` or call any `nvidia/*` model directly.
|
|
3
|
+
# @blockrun/llm
|
|
6
4
|
|
|
7
|
-
|
|
8
|
-
[](LICENSE)
|
|
5
|
+
### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
|
|
9
6
|
|
|
10
|
-
|
|
7
|
+
The smart-routing SDK for <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models โ every request goes to the cheapest model that can handle it,
|
|
8
|
+
paid per-request in USDC. No API keys. No subscriptions. No vendor lock-in.
|
|
11
9
|
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
10
|
+
[](https://www.npmjs.com/package/@blockrun/llm)
|
|
11
|
+
[](https://www.npmjs.com/package/@blockrun/llm)
|
|
12
|
+
[](https://github.com/BlockRunAI/blockrun-llm-ts/actions)
|
|
13
|
+
[](LICENSE)
|
|
14
|
+
[](package.json)
|
|
15
|
+
[](tsconfig.json)
|
|
17
16
|
|
|
17
|
+
[](https://base.org)
|
|
18
|
+
[](https://solana.com)
|
|
19
|
+
[](https://x402.org)
|
|
20
|
+
[](https://t.me/blockrunAI)
|
|
18
21
|
|
|
19
|
-
|
|
22
|
+
[Website](https://blockrun.ai) ยท [Models & Pricing](https://blockrun.ai/models) ยท [ClawRouter](https://github.com/BlockRunAI/ClawRouter) ยท [Python SDK](https://github.com/BlockRunAI/blockrun-llm) ยท [Telegram](https://t.me/blockrunAI)
|
|
23
|
+
|
|
24
|
+
</div>
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
```typescript
|
|
29
|
+
import { LLMClient } from '@blockrun/llm';
|
|
30
|
+
|
|
31
|
+
const client = new LLMClient();
|
|
32
|
+
|
|
33
|
+
const r = await client.smartChat('Prove step by step that the sum of two odd integers is even.');
|
|
34
|
+
console.log(r.model); // 'deepseek/deepseek-v4-pro' โ the right model, not the $75/M flagship
|
|
35
|
+
console.log(r.routing.savings); // 0.96 โ this exact request cost 96% less than pinning the baseline
|
|
36
|
+
console.log(r.response); // the proof
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
**<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** across a realistic workload on the default `auto` profile, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco` โ and eco's first stop is the free tier, so simple requests cost $0.00 outright. Not an "up to" figure: the baseline, workload mix, and token ratio are published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json) so anyone can recompute the claim. Details in [Smart Routing](#smart-routing-router-core-v3).
|
|
40
|
+
|
|
41
|
+
## Why This SDK
|
|
42
|
+
|
|
43
|
+
- ๐ง **Smart routing that pays for itself** โ the bundled [Router Core V3](https://github.com/BlockRunAI/router-core) engine (shared with [ClawRouter](https://github.com/BlockRunAI/ClawRouter)) classifies every request locally in <1ms across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and routes to the cheapest capable model. The main event.
|
|
44
|
+
- ๐ **<!-- br:models.free -->6<!-- /br:models.free --> genuinely free models** โ NVIDIA-hosted, $0 in and out, incl. 1M-context DeepSeek V4 Flash and a multimodal Nemotron. No rate-limit gimmicks.
|
|
45
|
+
- ๐ **No API keys** โ your wallet signature is your authentication. No accounts, no dashboards, no key rotation.
|
|
46
|
+
- ๐ธ **Pay per request in USDC** โ x402 micropayments on Base or Solana. $5 covers thousands of requests; agents can pay their own way.
|
|
47
|
+
- ๐ก๏ธ **Automatic failover** โ transient errors (timeouts, 429, 5xx) walk the router's ranked fallback chain instead of failing your request.
|
|
48
|
+
- โก **Streaming, OpenAI & Anthropic compat** โ drop-in `chat.completions` / `messages` layers, SSE streaming, strict TypeScript.
|
|
49
|
+
- ๐จ **Beyond chat** โ image, video, music, speech, live search, prediction markets, crypto data, and 40-chain RPC through the same wallet.
|
|
50
|
+
|
|
51
|
+
## How It Compares
|
|
52
|
+
|
|
53
|
+
| | OpenAI SDK | OpenRouter | LiteLLM | **@blockrun/llm** |
|
|
54
|
+
| ------------------ | -------------- | ----------------- | ---------------- | ----------------------------------------------------------------------- |
|
|
55
|
+
| **Cost routing** | โ one vendor | Manual selection | Manual selection | **Automatic โ <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
|
|
56
|
+
| **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->71<!-- /br:models.chatVisible -->, one wallet** |
|
|
57
|
+
| **Free tier** | โ | Rate-limited | โ | **<!-- br:models.free -->6<!-- /br:models.free --> models, no signup** |
|
|
58
|
+
| **Auth** | API key | Account + API key | Your API keys | **Wallet signature** |
|
|
59
|
+
| **Payment** | Card + invoice | Credit card | BYO keys | **USDC per-request** |
|
|
60
|
+
| **Agent-ready** | โ | โ | โ | **โ โ agents fund their own wallet** |
|
|
20
61
|
|
|
21
62
|
## Installation
|
|
22
63
|
|
|
23
64
|
```bash
|
|
24
|
-
# Base / EVM payments โ nothing else needed
|
|
25
|
-
npm install @blockrun/llm
|
|
26
|
-
# or
|
|
27
|
-
pnpm add @blockrun/llm
|
|
28
|
-
# or
|
|
29
|
-
yarn add @blockrun/llm
|
|
65
|
+
npm install @blockrun/llm # Base / EVM payments โ smart routing included, nothing else needed
|
|
30
66
|
```
|
|
31
67
|
|
|
32
|
-
|
|
33
|
-
|
|
68
|
+
<details>
|
|
69
|
+
<summary><strong>Solana payments</strong> โ two more optional peers</summary>
|
|
34
70
|
|
|
35
71
|
```bash
|
|
36
72
|
npm install @blockrun/llm @solana/web3.js @solana/spl-token
|
|
@@ -44,63 +80,38 @@ every consumer, including projects that only ever pay on Base. As an optional
|
|
|
44
80
|
*peer* it reaches only the projects that ask for Solana. Calling a Solana path
|
|
45
81
|
without them throws an error naming the exact install command.
|
|
46
82
|
|
|
47
|
-
|
|
83
|
+
</details>
|
|
48
84
|
|
|
49
|
-
|
|
50
|
-
|
|
85
|
+
<details>
|
|
86
|
+
<summary><strong>Supported chains</strong> โ Base (primary), Base Sepolia, Solana</summary>
|
|
51
87
|
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
88
|
+
| Chain | Network | Payment | Status |
|
|
89
|
+
|-------|---------|---------|--------|
|
|
90
|
+
| **Base** | Base Mainnet (Chain ID: 8453) | USDC | Primary |
|
|
91
|
+
| **Base Testnet** | Base Sepolia (Chain ID: 84532) | Testnet USDC | Development |
|
|
92
|
+
| **Solana** | Solana Mainnet | USDC (SPL) | New |
|
|
55
93
|
|
|
56
|
-
|
|
94
|
+
**Protocol:** x402 v2 (CDP Facilitator)
|
|
57
95
|
|
|
58
|
-
|
|
96
|
+
</details>
|
|
59
97
|
|
|
60
|
-
|
|
61
|
-
**every** BlockRun endpoint over x402. New API surfaces are intended to be
|
|
62
|
-
distributed as [Claude Code skills](https://github.com/anthropics/skills)
|
|
63
|
-
that drive this primitive โ no SDK release required to add an endpoint.
|
|
98
|
+
## Quick Start (Base - Default)
|
|
64
99
|
|
|
65
100
|
```typescript
|
|
66
|
-
import {
|
|
67
|
-
|
|
68
|
-
const br = new BlockrunClient();
|
|
69
|
-
|
|
70
|
-
// Sync GET โ Surf market price (Tier 1, $0.001)
|
|
71
|
-
const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
|
|
101
|
+
import { LLMClient } from '@blockrun/llm';
|
|
72
102
|
|
|
73
|
-
|
|
74
|
-
const rows = await br.post('/v1/surf/onchain/sql', {
|
|
75
|
-
query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
|
|
76
|
-
});
|
|
103
|
+
const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
|
|
77
104
|
|
|
78
|
-
//
|
|
79
|
-
const
|
|
80
|
-
model: 'xai/grok-imagine-video',
|
|
81
|
-
prompt: 'a red apple spinning',
|
|
82
|
-
});
|
|
105
|
+
// Recommended: let the router pick the cheapest capable model
|
|
106
|
+
const result = await client.smartChat('Hello!');
|
|
83
107
|
|
|
84
|
-
//
|
|
85
|
-
|
|
86
|
-
model: 'anthropic/claude-sonnet-4-6',
|
|
87
|
-
messages: [{ role: 'user', content: 'Hi' }],
|
|
88
|
-
stream: true,
|
|
89
|
-
})) {
|
|
90
|
-
process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
|
|
91
|
-
}
|
|
108
|
+
// Or pin a model yourself
|
|
109
|
+
const response = await client.chat('openai/gpt-4o', 'Hello!');
|
|
92
110
|
```
|
|
93
111
|
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
- `poll<T>(path, body?, { budgetMs, intervalMs })` โ submit + poll (image, video, music, voice)
|
|
98
|
-
- `stream<T>(path, body?)` โ async iterator over SSE chunks (chat)
|
|
99
|
-
|
|
100
|
-
The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
|
|
101
|
-
`PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `RpcClient`,
|
|
102
|
-
`PriceClient`, `SurfClient`) all remain โ they will be soft-deprecated in 2.6 (rewritten as
|
|
103
|
-
shims over `BlockrunClient`) and removed in 3.0.
|
|
112
|
+
That's it. The SDK handles x402 payment automatically โ and `smartChat()`
|
|
113
|
+
keeps the bill down on every request. The router is bundled: no extra
|
|
114
|
+
package to install.
|
|
104
115
|
|
|
105
116
|
### Try It Free (No USDC Required)
|
|
106
117
|
|
|
@@ -114,30 +125,33 @@ const client = new LLMClient(); // Wallet still required for signing, but $0 ch
|
|
|
114
125
|
// Option 1: call a free model directly
|
|
115
126
|
const reply = await client.chat('nvidia/deepseek-v4-flash', 'Explain x402 in 1 sentence');
|
|
116
127
|
|
|
117
|
-
// Option 2: let the smart router pick the
|
|
118
|
-
const result = await client.smartChat('What is 2+2?', { routingProfile: '
|
|
119
|
-
console.log(result.model); //
|
|
128
|
+
// Option 2: let the smart router pick โ 'eco' ranks the free NVIDIA tier first
|
|
129
|
+
const result = await client.smartChat('What is 2+2?', { routingProfile: 'eco' });
|
|
130
|
+
console.log(result.model); // 'nvidia/deepseek-v4-flash' ($0 โ verified live)
|
|
120
131
|
console.log(result.response); // '4'
|
|
132
|
+
console.log(result.routing.savings); // 1 (100%)
|
|
121
133
|
```
|
|
122
134
|
|
|
123
|
-
|
|
135
|
+
There is no `free` routing profile in `smartChat()` โ `routingProfile` accepts
|
|
136
|
+
`'eco' | 'auto' | 'premium'`. (ClawRouter's `/model free` is a feature of its
|
|
137
|
+
own proxy, not of this SDK's router options.) For guaranteed $0, pin a
|
|
138
|
+
`nvidia/*` model; for smart-routed $0-first, use `eco`.
|
|
139
|
+
|
|
140
|
+
**Available free models** (input + output both $0, all NVIDIA-hosted, from the live `/v1/models` catalog, last refreshed 2026-08-10):
|
|
124
141
|
|
|
125
142
|
| Model ID | Context | Best For |
|
|
126
143
|
|----------|---------|----------|
|
|
127
|
-
| `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash โ 284B / 13B active MoE
|
|
128
|
-
| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K |
|
|
129
|
-
| `nvidia/
|
|
130
|
-
| `nvidia/
|
|
131
|
-
| `nvidia/
|
|
132
|
-
| `nvidia/
|
|
133
|
-
| `nvidia/gpt-oss-
|
|
134
|
-
|
|
135
|
-
> Need V4-Pro-class reasoning? Use the paid `deepseek/deepseek-v4-pro` ($0.435/$0.87 โ the 75% launch promo became the permanent list price after 2026-05-31) โ `nvidia/deepseek-v4-pro` is currently hidden because NVIDIA's NIM deployment is hung; backend MODEL_REDIRECTS forwards calls to V4 Flash.
|
|
144
|
+
| `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash โ 284B / 13B active MoE. Best free chat / summarization / light reasoning. Capacity-constrained: requests may be answered by an equivalent free model |
|
|
145
|
+
| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning โ text + images + video + audio (ChartQA 90.3, DocVQA 95.6) |
|
|
146
|
+
| `nvidia/mistral-nemotron` | 131K | Mistral ร NVIDIA instruction model โ fast (~0.2s), strong instruction following |
|
|
147
|
+
| `nvidia/step-3.7-flash` | 131K | StepFun Step 3.7 Flash โ fast lightweight reasoning |
|
|
148
|
+
| `nvidia/nemotron-nano-9b-v2` | 131K | Compact + fast (~0.7s), good for high-volume light tasks |
|
|
149
|
+
| `nvidia/nemotron-nano-12b-v2-vl` | 131K | Vision-language โ accepts images, compact + fast |
|
|
150
|
+
| `nvidia/gpt-oss-120b` | 128K | OpenAI open-weight 120B. Hidden from `/v1/models` for privacy but direct calls still work |
|
|
151
|
+
| `nvidia/gpt-oss-20b` | 128K | OpenAI open-weight 20B. Hidden from `/v1/models` but direct calls still work |
|
|
136
152
|
|
|
137
153
|
> Privacy note: `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` are hidden from `/v1/models` because NVIDIA's free build.nvidia.com tier reserves the right to use prompts/outputs for service improvement. Direct calls by full model ID still work โ opt in only when your data isn't sensitive.
|
|
138
154
|
|
|
139
|
-
> Retired: `nvidia/qwen3-next-80b-a3b-thinking` hit NVIDIA end-of-life 2026-05-21 (HTTP 410). The gateway auto-redirects pinned callers to `nvidia/llama-4-maverick`.
|
|
140
|
-
|
|
141
155
|
## Quick Start (Solana)
|
|
142
156
|
|
|
143
157
|
```typescript
|
|
@@ -151,6 +165,191 @@ console.log(response);
|
|
|
151
165
|
|
|
152
166
|
Set `SOLANA_WALLET_KEY` to your bs58-encoded Solana secret key. Payments are automatic via x402 โ your key never leaves your machine.
|
|
153
167
|
|
|
168
|
+
## Smart Routing (Router Core V3)
|
|
169
|
+
|
|
170
|
+
Let the SDK automatically pick the cheapest capable model for each request โ **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** for the same traffic on `auto`, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco`.
|
|
171
|
+
|
|
172
|
+
Not an "up to" figure. The baseline, the workload mix and the token ratio are
|
|
173
|
+
published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json),
|
|
174
|
+
priced against the live catalog, so anyone can recompute the claim and get the
|
|
175
|
+
same answer.
|
|
176
|
+
|
|
177
|
+
Smart routing is powered by the product-neutral
|
|
178
|
+
[`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) V3 engine โ
|
|
179
|
+
the same deterministic portfolio router that drives
|
|
180
|
+
[ClawRouter](https://github.com/BlockRunAI/ClawRouter). It is **bundled into
|
|
181
|
+
this SDK**: no separate router package to install, and routing runs 100%
|
|
182
|
+
locally with zero external calls.
|
|
183
|
+
|
|
184
|
+
Three ways to use it:
|
|
185
|
+
|
|
186
|
+
```typescript
|
|
187
|
+
// 1. smartChat() โ one-line routed chat
|
|
188
|
+
const result = await client.smartChat('What is 2+2?');
|
|
189
|
+
|
|
190
|
+
// 2. smartChatCompletion() โ full agent/tool conversations, routed
|
|
191
|
+
const agent = await client.smartChatCompletion(messages, { tools, toolChoice: 'auto' });
|
|
192
|
+
|
|
193
|
+
// 3. blockrun/auto | blockrun/eco | blockrun/premium โ model aliases accepted
|
|
194
|
+
// by chat(), chatCompletion(), and chatCompletionStream() on both chains
|
|
195
|
+
const reply = await client.chatCompletion('blockrun/auto', messages);
|
|
196
|
+
|
|
197
|
+
// Inspect a decision without paying for anything
|
|
198
|
+
const decision = await client.route('Prove the Riemann hypothesis');
|
|
199
|
+
```
|
|
200
|
+
|
|
201
|
+
The aliases are resolved locally by `LLMClient`, `SolanaLLMClient`, and the
|
|
202
|
+
OpenAI-compat layer. The Anthropic-compat layer proxies straight to the
|
|
203
|
+
gateway's `/v1/messages` and does **not** resolve them โ pass a concrete
|
|
204
|
+
model id there.
|
|
205
|
+
|
|
206
|
+
```typescript
|
|
207
|
+
import { LLMClient } from '@blockrun/llm';
|
|
208
|
+
|
|
209
|
+
const client = new LLMClient();
|
|
210
|
+
|
|
211
|
+
// Auto-routes to cheapest capable model
|
|
212
|
+
const result = await client.smartChat('What is 2+2?');
|
|
213
|
+
console.log(result.response); // '4'
|
|
214
|
+
console.log(result.model); // 'moonshot/kimi-k2.5' (cheap, fast)
|
|
215
|
+
console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 88%'
|
|
216
|
+
|
|
217
|
+
// Complex reasoning task -> routes to reasoning model
|
|
218
|
+
const complex = await client.smartChat('Prove the Riemann hypothesis step by step');
|
|
219
|
+
console.log(complex.model); // 'xai/grok-4-1-fast-reasoning'
|
|
220
|
+
|
|
221
|
+
// Inspect how the request was classified and ranked (Router v3.4 portfolio).
|
|
222
|
+
console.log(complex.routing.method); // 'portfolio'
|
|
223
|
+
console.log(complex.routing.taskType); // 'reasoning'
|
|
224
|
+
console.log(complex.routing.candidates); // ranked, capability-eligible models
|
|
225
|
+
|
|
226
|
+
// Inspect the fallback chain SmartChat will walk on transient errors.
|
|
227
|
+
console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
|
|
228
|
+
```
|
|
229
|
+
|
|
230
|
+
### Automatic Fallback on Transient Errors
|
|
231
|
+
|
|
232
|
+
`smartChat()` populates a fallback chain from the portfolio ranking and
|
|
233
|
+
`chat()` / `chatCompletion()` walk it automatically when the primary model
|
|
234
|
+
returns a transient error โ timeouts, network failures, 429 rate limits, or
|
|
235
|
+
5xx responses (502/503/504/522/524). Other 4xx errors and `PaymentError`
|
|
236
|
+
propagate immediately so wallet / auth issues surface fast.
|
|
237
|
+
|
|
238
|
+
```typescript
|
|
239
|
+
// Manually pass a fallback chain to chat() / chatCompletion()
|
|
240
|
+
const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
|
|
241
|
+
fallbackModels: ['nvidia/llama-4-maverick', 'nvidia/mistral-small-4-119b'],
|
|
242
|
+
});
|
|
243
|
+
// If deepseek-v4-flash times out, the SDK retries against the next model
|
|
244
|
+
// and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
|
|
245
|
+
```
|
|
246
|
+
|
|
247
|
+
### Routing Profiles
|
|
248
|
+
|
|
249
|
+
| Profile | Strategy | Savings vs Opus 5 | Best For |
|
|
250
|
+
|---------|----------|-------------------|----------|
|
|
251
|
+
| `eco` | Cheapest capable model โ ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
|
|
252
|
+
| `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%** | General use |
|
|
253
|
+
| `premium` | Top-tier models (OpenAI, Anthropic) | 0% | Quality-critical tasks |
|
|
254
|
+
|
|
255
|
+
For guaranteed $0, call a `nvidia/*` model directly with `chat()` โ see
|
|
256
|
+
[Try It Free](#try-it-free-no-usdc-required). ClawRouter's `/model free`
|
|
257
|
+
profile belongs to its own proxy; `smartChat()`'s options are the three above.
|
|
258
|
+
|
|
259
|
+
```typescript
|
|
260
|
+
// Use premium models for complex tasks
|
|
261
|
+
const result = await client.smartChat(
|
|
262
|
+
'Write production-grade async TypeScript code',
|
|
263
|
+
{ routingProfile: 'premium' }
|
|
264
|
+
);
|
|
265
|
+
console.log(result.model); // 'anthropic/claude-opus-4.7'
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
### How the Router Works
|
|
269
|
+
|
|
270
|
+
```mermaid
|
|
271
|
+
flowchart LR
|
|
272
|
+
A["prompt"] --> B["classify locally<br/>15 dimensions, <1ms"]
|
|
273
|
+
B --> C["hard filters<br/>tools ยท vision ยท context ยท<br/>structured output"]
|
|
274
|
+
C --> D["rank portfolio<br/>quality ยท cost ยท speed ยท<br/>reliability"]
|
|
275
|
+
D --> E["cheapest capable model<br/>+ ranked fallback chain"]
|
|
276
|
+
E --> F["x402 USDC payment<br/>for this request only"]
|
|
277
|
+
F --> G["response<br/>+ full routing metadata"]
|
|
278
|
+
```
|
|
279
|
+
|
|
280
|
+
Since ClawRouter v0.12.242, Auto uses the deterministic **Router v3.4 portfolio
|
|
281
|
+
strategy**: it classifies the task shape locally across
|
|
282
|
+
<!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions
|
|
283
|
+
(token count, code presence, reasoning markers, technical/creative terms,
|
|
284
|
+
agentic patterns, โฆ), enforces tool / vision / structured-output / context
|
|
285
|
+
constraints as **hard filters**, then ranks an ordered candidate portfolio.
|
|
286
|
+
The winner becomes `routing.model`; the rest surface as `routing.candidates`
|
|
287
|
+
and feed SmartChat's transient-error fallback chain. Routing stays 100% local
|
|
288
|
+
and deterministic โ <1ms, no extra model call, no network hop.
|
|
289
|
+
|
|
290
|
+
Classification still maps to one of four tiers (`routing.tier`). Each
|
|
291
|
+
tier ร profile has a designated primary (what the rules strategy โ
|
|
292
|
+
`routing.method: 'rules'`, the rollback lever โ routes to directly, and what
|
|
293
|
+
anchors the portfolio's candidate pool):
|
|
294
|
+
|
|
295
|
+
| Tier | Example Tasks | ECO | AUTO | PREMIUM |
|
|
296
|
+
|------|---------------|-----|------|---------|
|
|
297
|
+
| SIMPLE | "What is 2+2?", definitions | free/gpt-oss-120b โ (**FREE**) | gemini-2.5-flash ($0.30/$2.50) | kimi-k2.7 โ ($0.95/$4.00) |
|
|
298
|
+
| MEDIUM | Code snippets, explanations | gemini-3.1-flash-lite ($0.25/$1.50) | kimi-k2.7 ($0.95/$4.00) | gpt-5.3-codex ($1.75/$14.00) |
|
|
299
|
+
| COMPLEX | Architecture, long documents | gemini-3.1-flash-lite ($0.25/$1.50) | gemini-3.1-pro ($2/$12) | claude-fable-5 ($10/$50) |
|
|
300
|
+
| REASONING | Proofs, multi-step reasoning | grok-4-1-fast-reasoning โ ($0.20/$0.50) | grok-4-1-fast-reasoning โ ($0.20/$0.50) | claude-sonnet-4.6 ($3/$15) |
|
|
301
|
+
|
|
302
|
+
โ Withheld from `/v1/models` โ the router still calls it by direct ID, but you
|
|
303
|
+
will not find it on the public pricing page. The published savings claim is
|
|
304
|
+
priced on visible models only.
|
|
305
|
+
|
|
306
|
+
This table mirrors ClawRouter's tier configs at the version this SDK pins;
|
|
307
|
+
the [ClawRouter README](https://github.com/BlockRunAI/ClawRouter#how-it-works)
|
|
308
|
+
is the live source of truth as models and prices move.
|
|
309
|
+
|
|
310
|
+
### Routing Metadata Reference
|
|
311
|
+
|
|
312
|
+
Every `smartChat()` result carries the full decision on `result.routing`
|
|
313
|
+
(type `RoutingDecision`) โ enough to log, audit, or replay why a model was
|
|
314
|
+
picked:
|
|
315
|
+
|
|
316
|
+
| Field | Description |
|
|
317
|
+
|-------|-------------|
|
|
318
|
+
| `model` | Selected model id (same as `result.model`) |
|
|
319
|
+
| `method` | `'portfolio'` (the Auto default), `'rules'` (rollback strategy), or `'llm'` |
|
|
320
|
+
| `tier` | Task tier: `'SIMPLE'`, `'MEDIUM'`, `'COMPLEX'`, or `'REASONING'` |
|
|
321
|
+
| `taskType` | Portfolio task classification: `'chat'`, `'extraction'`, `'code_edit'`, `'code_agent'`, `'tool_agent'`, `'debug'`, `'reasoning'`, `'reasoning_math'`, `'long_context'`, `'vision'`, โฆ |
|
|
322
|
+
| `candidates` | Ordered, capability-eligible models ranked by the portfolio router; the first entry is `model` |
|
|
323
|
+
| `candidateScores` | Per-candidate score breakdown (`quality` / `cost` / `speed` / `reliability`), ordered with `candidates` |
|
|
324
|
+
| `fallbacks` | The chain `chat()` walks on transient errors (timeout / network / 429 / 5xx) โ `candidates` minus the primary, with ClawRouter's proxy-namespace `free/*` ids mapped to their `nvidia/*` gateway ids (SDK-computed) |
|
|
325
|
+
| `savings` | 0โ1 fraction saved vs the premium baseline |
|
|
326
|
+
| `costEstimate` / `baselineCost` | Estimated cost of the pick vs that baseline, in USD |
|
|
327
|
+
| `confidence` | Sigmoid-calibrated classifier confidence, 0โ1 |
|
|
328
|
+
| `routerVersion` | `'v3-portfolio'` or `'v2-rules'` |
|
|
329
|
+
| `profile` | Routing profile applied: `'auto'`, `'eco'`, `'premium'`, or `'agentic'` |
|
|
330
|
+
| `reasoning` | Human-readable explanation of the decision |
|
|
331
|
+
| `tierConfigs` | The tier โ primary/fallback map the decision was made against |
|
|
332
|
+
|
|
333
|
+
### TypeScript Types
|
|
334
|
+
|
|
335
|
+
`RoutingDecision`, `RoutingProfile`, `RoutingTier`, `RoutingTaskType`,
|
|
336
|
+
`RoutingTierConfig`, `SmartChatCompletionOptions`, and
|
|
337
|
+
`SmartChatCompletionResponse` are exported from `@blockrun/llm`. They are
|
|
338
|
+
derived from
|
|
339
|
+
[`@blockrun/router-core`](https://github.com/BlockRunAI/router-core), pinned
|
|
340
|
+
to a reviewed immutable commit, and shipped **inlined in this SDK's
|
|
341
|
+
declaration files and runtime bundle** โ you install nothing extra to route
|
|
342
|
+
or to typecheck.
|
|
343
|
+
|
|
344
|
+
### Going Deeper
|
|
345
|
+
|
|
346
|
+
- [ClawRouter](https://github.com/BlockRunAI/ClawRouter) โ the router itself: OpenClaw plugin, standalone proxy for Cursor / continue.dev / any OpenAI-compatible client, Telegram integration
|
|
347
|
+
- [Routing profiles in depth](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/routing-profiles.md) โ ECO / AUTO / PREMIUM details
|
|
348
|
+
- [How the routing engine works](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/smart-llm-router-14-dimension-classifier.md) โ the classifier, dimension by dimension
|
|
349
|
+
- [Router benchmark](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/llm-router-benchmark-46-models-sub-1ms-routing.md) โ sub-1ms routing across the catalog
|
|
350
|
+
- [ClawRouter vs OpenRouter](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/clawrouter-vs-openrouter-llm-routing-comparison.md) โ head-to-head comparison
|
|
351
|
+
- [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) โ the deterministic routing engine both share
|
|
352
|
+
|
|
154
353
|
## Solana Support
|
|
155
354
|
|
|
156
355
|
Pay for AI calls with Solana USDC via [sol.blockrun.ai](https://sol.blockrun.ai):
|
|
@@ -234,83 +433,52 @@ Every paid request is a real on-chain USDC transfer โ look up your wallet addr
|
|
|
234
433
|
|
|
235
434
|
**Non-custodial by design: your private key never leaves your machine** โ it is only used for local signing, and no funds are ever held by BlockRun.
|
|
236
435
|
|
|
237
|
-
##
|
|
436
|
+
## `BlockrunClient` โ the universal primitive (recommended for new code)
|
|
238
437
|
|
|
239
|
-
|
|
438
|
+
Starting in `2.5.0`, the SDK ships a single `BlockrunClient` that speaks to
|
|
439
|
+
**every** BlockRun endpoint over x402. New API surfaces are intended to be
|
|
440
|
+
distributed as [Claude Code skills](https://github.com/anthropics/skills)
|
|
441
|
+
that drive this primitive โ no SDK release required to add an endpoint.
|
|
240
442
|
|
|
241
443
|
```typescript
|
|
242
|
-
import {
|
|
243
|
-
|
|
244
|
-
const client = new LLMClient();
|
|
245
|
-
|
|
246
|
-
// Auto-routes to cheapest capable model
|
|
247
|
-
const result = await client.smartChat('What is 2+2?');
|
|
248
|
-
console.log(result.response); // '4'
|
|
249
|
-
console.log(result.model); // 'moonshot/kimi-k2.5' (cheap, fast)
|
|
250
|
-
console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 87%'
|
|
251
|
-
|
|
252
|
-
// Complex reasoning task -> routes to reasoning model
|
|
253
|
-
const complex = await client.smartChat('Prove the Riemann hypothesis step by step');
|
|
254
|
-
console.log(complex.model); // 'xai/grok-4-1-fast-reasoning'
|
|
255
|
-
|
|
256
|
-
// Inspect the fallback chain SmartChat will walk on transient errors.
|
|
257
|
-
console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
|
|
258
|
-
```
|
|
444
|
+
import { BlockrunClient } from '@blockrun/llm';
|
|
259
445
|
|
|
260
|
-
|
|
446
|
+
const br = new BlockrunClient();
|
|
261
447
|
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
transient error โ timeouts, network failures, or 5xx responses (502/503/504/
|
|
265
|
-
522/524). 4xx errors and `PaymentError` propagate immediately so wallet /
|
|
266
|
-
auth issues surface fast.
|
|
448
|
+
// Sync GET โ Surf market price (Tier 1, $0.001)
|
|
449
|
+
const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
|
|
267
450
|
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
fallbackModels: ['nvidia/llama-4-maverick', 'nvidia/mistral-small-4-119b'],
|
|
451
|
+
// Sync POST โ raw on-chain SQL (Tier 3, $0.020)
|
|
452
|
+
const rows = await br.post('/v1/surf/onchain/sql', {
|
|
453
|
+
query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
|
|
272
454
|
});
|
|
273
|
-
// If deepseek-v4-flash times out, the SDK retries against the next model
|
|
274
|
-
// and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
|
|
275
|
-
```
|
|
276
|
-
|
|
277
|
-
### Routing Profiles
|
|
278
455
|
|
|
279
|
-
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
|
|
283
|
-
|
|
284
|
-
| `premium` | Top-tier models (OpenAI, Anthropic) | Quality-critical tasks |
|
|
456
|
+
// Submit + poll โ long-running video gen (settled only on completion)
|
|
457
|
+
const video = await br.poll('/v1/videos/generations', {
|
|
458
|
+
model: 'xai/grok-imagine-video',
|
|
459
|
+
prompt: 'a red apple spinning',
|
|
460
|
+
});
|
|
285
461
|
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
)
|
|
292
|
-
|
|
462
|
+
// Streaming SSE โ chat completions
|
|
463
|
+
for await (const chunk of br.stream('/v1/chat/completions', {
|
|
464
|
+
model: 'anthropic/claude-sonnet-4-6',
|
|
465
|
+
messages: [{ role: 'user', content: 'Hi' }],
|
|
466
|
+
stream: true,
|
|
467
|
+
})) {
|
|
468
|
+
process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
|
|
469
|
+
}
|
|
293
470
|
```
|
|
294
471
|
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
-
|
|
300
|
-
- **Code presence** - Programming keywords
|
|
301
|
-
- **Reasoning markers** - "prove", "step by step", etc.
|
|
302
|
-
- **Technical terms** - Architecture, optimization, etc.
|
|
303
|
-
- **Creative markers** - Story, poem, brainstorm, etc.
|
|
304
|
-
- **Agentic patterns** - Multi-step, tool use indicators
|
|
305
|
-
|
|
306
|
-
The classifier runs in <1ms, 100% locally, and routes to one of four tiers:
|
|
472
|
+
Four call shapes cover every endpoint type:
|
|
473
|
+
- `get<T>(path, params?)` โ synchronous GET (price, ranking, list, news)
|
|
474
|
+
- `post<T>(path, body?)` โ synchronous POST (on-chain SQL, search)
|
|
475
|
+
- `poll<T>(path, body?, { budgetMs, intervalMs })` โ submit + poll (image, video, music, voice)
|
|
476
|
+
- `stream<T>(path, body?)` โ async iterator over SSE chunks (chat)
|
|
307
477
|
|
|
308
|
-
|
|
309
|
-
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
| COMPLEX | Architecture, long documents | google/gemini-3.1-pro |
|
|
313
|
-
| REASONING | Proofs, multi-step reasoning | xai/grok-4-1-fast-reasoning |
|
|
478
|
+
The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
|
|
479
|
+
`PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `RpcClient`,
|
|
480
|
+
`PriceClient`, `SurfClient`) all remain โ they will be soft-deprecated in 2.6 (rewritten as
|
|
481
|
+
shims over `BlockrunClient`) and removed in 3.0.
|
|
314
482
|
|
|
315
483
|
## Available Models
|
|
316
484
|
|
|
@@ -871,9 +1039,9 @@ const response2 = await client.chat('anthropic/claude-sonnet-4', 'Write a haiku'
|
|
|
871
1039
|
});
|
|
872
1040
|
```
|
|
873
1041
|
|
|
874
|
-
### Smart Routing (
|
|
1042
|
+
### Smart Routing (Router Core V3)
|
|
875
1043
|
|
|
876
|
-
Save up to <!-- br:savings.autoVsBaselinePct -->
|
|
1044
|
+
Save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. The bundled Router Core V3 engine classifies each request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions, applies hard capability filters, and ranks the cheapest capable models (<1ms, 100% local). Bundled โ nothing extra to install.
|
|
877
1045
|
|
|
878
1046
|
```typescript
|
|
879
1047
|
import { LLMClient } from '@blockrun/llm';
|
|
@@ -885,21 +1053,22 @@ const result = await client.smartChat('What is 2+2?');
|
|
|
885
1053
|
console.log(result.response); // '4'
|
|
886
1054
|
console.log(result.model); // 'google/gemini-2.5-flash'
|
|
887
1055
|
console.log(result.routing.tier); // 'SIMPLE'
|
|
888
|
-
console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved
|
|
1056
|
+
console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 88%'
|
|
889
1057
|
|
|
890
|
-
// Routing profiles
|
|
891
|
-
const
|
|
892
|
-
const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Budget optimized
|
|
1058
|
+
// Routing profiles ('eco' | 'auto' | 'premium')
|
|
1059
|
+
const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Free tier first, then cheapest paid
|
|
893
1060
|
const auto = await client.smartChat('Code review', { routingProfile: 'auto' }); // Balanced (default)
|
|
894
1061
|
const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
|
|
1062
|
+
|
|
1063
|
+
// Guaranteed $0: call a free NVIDIA model directly
|
|
1064
|
+
const free = await client.chat('nvidia/deepseek-v4-flash', 'Hello!');
|
|
895
1065
|
```
|
|
896
1066
|
|
|
897
1067
|
**Routing Profiles:**
|
|
898
1068
|
|
|
899
1069
|
| Profile | Description | Best For |
|
|
900
1070
|
|---------|-------------|----------|
|
|
901
|
-
| `
|
|
902
|
-
| `eco` | Budget-optimized | Cost-sensitive workloads |
|
|
1071
|
+
| `eco` | Budget-optimized โ ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | Cost-sensitive workloads, zero-cost testing |
|
|
903
1072
|
| `auto` | Intelligent routing (default) | General use |
|
|
904
1073
|
| `premium` | Best quality models | Critical tasks |
|
|
905
1074
|
|
|
@@ -1436,13 +1605,13 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
|
|
|
1436
1605
|
## Frequently Asked Questions
|
|
1437
1606
|
|
|
1438
1607
|
### What is @blockrun/llm?
|
|
1439
|
-
@blockrun/llm is a TypeScript SDK that
|
|
1608
|
+
@blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol โ no API keys, no subscriptions, no vendor lock-in.
|
|
1440
1609
|
|
|
1441
1610
|
### How does payment work?
|
|
1442
1611
|
When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
|
|
1443
1612
|
|
|
1444
|
-
### What is smart routing
|
|
1445
|
-
|
|
1613
|
+
### What is smart routing?
|
|
1614
|
+
Router Core V3 is bundled into the SDK โ the same deterministic routing engine that powers ClawRouter, with nothing extra to install. It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. Use `smartChat()`, `smartChatCompletion()`, or the `blockrun/auto` model alias. It can save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
|
|
1446
1615
|
|
|
1447
1616
|
### Does it support streaming?
|
|
1448
1617
|
Yes โ as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).
|
|
@@ -1453,6 +1622,16 @@ Pay only for what you use. Prices start at $0.0002 per request (GPT-5 Nano). The
|
|
|
1453
1622
|
### Does it support both Base and Solana?
|
|
1454
1623
|
Yes. Use `LLMClient` for Base (EVM) payments and `SolanaLLMClient` for Solana payments. Same API, different payment chain.
|
|
1455
1624
|
|
|
1625
|
+
---
|
|
1626
|
+
|
|
1627
|
+
<div align="center">
|
|
1628
|
+
|
|
1629
|
+
**If the router just cut your bill, [give it a star โญ](https://github.com/BlockRunAI/blockrun-llm-ts)** โ it helps more agents pay less.
|
|
1630
|
+
|
|
1631
|
+
[Website](https://blockrun.ai) ยท [Models & Pricing](https://blockrun.ai/models) ยท [ClawRouter](https://github.com/BlockRunAI/ClawRouter) ยท [Python SDK](https://github.com/BlockRunAI/blockrun-llm) ยท [Telegram](https://t.me/blockrunAI)
|
|
1632
|
+
|
|
1633
|
+
</div>
|
|
1634
|
+
|
|
1456
1635
|
## License
|
|
1457
1636
|
|
|
1458
1637
|
MIT
|