@blockrun/llm 3.11.0 โ 3.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +296 -194
- package/dist/index.cjs +3598 -89
- package/dist/index.d.cts +42 -16
- package/dist/index.d.ts +42 -16
- package/dist/index.js +3598 -89
- package/package.json +2 -7
package/README.md
CHANGED
|
@@ -1,36 +1,72 @@
|
|
|
1
|
-
|
|
1
|
+
<div align="center">
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
>
|
|
5
|
-
> ๐ **Includes 7 fully-free NVIDIA-hosted models** (5 visible in `/v1/models`, 2 hidden but directly callable) โ DeepSeek V4 Flash (1M context), Nemotron Nano Omni (vision), Qwen3 Coder, Llama 4, Mistral, plus the gpt-oss pair. Zero USDC, no rate-limit gimmicks. Use `routingProfile: 'free'` or call any `nvidia/*` model directly.
|
|
3
|
+
# @blockrun/llm
|
|
6
4
|
|
|
7
|
-
|
|
8
|
-
[](LICENSE)
|
|
5
|
+
### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
|
|
9
6
|
|
|
10
|
-
|
|
7
|
+
The smart-routing SDK for <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models โ every request goes to the cheapest model that can handle it,
|
|
8
|
+
paid per-request in USDC. No API keys. No subscriptions. No vendor lock-in.
|
|
11
9
|
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
10
|
+
[](https://www.npmjs.com/package/@blockrun/llm)
|
|
11
|
+
[](https://www.npmjs.com/package/@blockrun/llm)
|
|
12
|
+
[](https://github.com/BlockRunAI/blockrun-llm-ts/actions)
|
|
13
|
+
[](LICENSE)
|
|
14
|
+
[](package.json)
|
|
15
|
+
[](tsconfig.json)
|
|
17
16
|
|
|
17
|
+
[](https://base.org)
|
|
18
|
+
[](https://solana.com)
|
|
19
|
+
[](https://x402.org)
|
|
20
|
+
[](https://t.me/blockrunAI)
|
|
18
21
|
|
|
19
|
-
|
|
22
|
+
[Website](https://blockrun.ai) ยท [Models & Pricing](https://blockrun.ai/models) ยท [ClawRouter](https://github.com/BlockRunAI/ClawRouter) ยท [Python SDK](https://github.com/BlockRunAI/blockrun-llm) ยท [Telegram](https://t.me/blockrunAI)
|
|
23
|
+
|
|
24
|
+
</div>
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
```typescript
|
|
29
|
+
import { LLMClient } from '@blockrun/llm';
|
|
30
|
+
|
|
31
|
+
const client = new LLMClient();
|
|
32
|
+
|
|
33
|
+
const r = await client.smartChat('Prove step by step that the sum of two odd integers is even.');
|
|
34
|
+
console.log(r.model); // 'deepseek/deepseek-v4-pro' โ the right model, not the $75/M flagship
|
|
35
|
+
console.log(r.routing.savings); // 0.96 โ this exact request cost 96% less than pinning the baseline
|
|
36
|
+
console.log(r.response); // the proof
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
**<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** across a realistic workload on the default `auto` profile, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco` โ and eco's first stop is the free tier, so simple requests cost $0.00 outright. Not an "up to" figure: the baseline, workload mix, and token ratio are published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json) so anyone can recompute the claim. Details in [Smart Routing](#smart-routing-router-core-v3).
|
|
40
|
+
|
|
41
|
+
## Why This SDK
|
|
42
|
+
|
|
43
|
+
- ๐ง **Smart routing that pays for itself** โ the bundled [Router Core V3](https://github.com/BlockRunAI/router-core) engine (shared with [ClawRouter](https://github.com/BlockRunAI/ClawRouter)) classifies every request locally in <1ms across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and routes to the cheapest capable model. The main event.
|
|
44
|
+
- ๐ **<!-- br:models.free -->6<!-- /br:models.free --> genuinely free models** โ NVIDIA-hosted, $0 in and out, incl. 1M-context DeepSeek V4 Flash and a multimodal Nemotron. No rate-limit gimmicks.
|
|
45
|
+
- ๐ **No API keys** โ your wallet signature is your authentication. No accounts, no dashboards, no key rotation.
|
|
46
|
+
- ๐ธ **Pay per request in USDC** โ x402 micropayments on Base or Solana. $5 covers thousands of requests; agents can pay their own way.
|
|
47
|
+
- ๐ก๏ธ **Automatic failover** โ transient errors (timeouts, 429, 5xx) walk the router's ranked fallback chain instead of failing your request.
|
|
48
|
+
- โก **Streaming, OpenAI & Anthropic compat** โ drop-in `chat.completions` / `messages` layers, SSE streaming, strict TypeScript.
|
|
49
|
+
- ๐จ **Beyond chat** โ image, video, music, speech, live search, prediction markets, crypto data, and 40-chain RPC through the same wallet.
|
|
50
|
+
|
|
51
|
+
## How It Compares
|
|
52
|
+
|
|
53
|
+
| | OpenAI SDK | OpenRouter | LiteLLM | **@blockrun/llm** |
|
|
54
|
+
| ------------------ | -------------- | ----------------- | ---------------- | ----------------------------------------------------------------------- |
|
|
55
|
+
| **Cost routing** | โ one vendor | Manual selection | Manual selection | **Automatic โ <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
|
|
56
|
+
| **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->71<!-- /br:models.chatVisible -->, one wallet** |
|
|
57
|
+
| **Free tier** | โ | Rate-limited | โ | **<!-- br:models.free -->6<!-- /br:models.free --> models, no signup** |
|
|
58
|
+
| **Auth** | API key | Account + API key | Your API keys | **Wallet signature** |
|
|
59
|
+
| **Payment** | Card + invoice | Credit card | BYO keys | **USDC per-request** |
|
|
60
|
+
| **Agent-ready** | โ | โ | โ | **โ โ agents fund their own wallet** |
|
|
20
61
|
|
|
21
62
|
## Installation
|
|
22
63
|
|
|
23
64
|
```bash
|
|
24
|
-
# Base / EVM payments โ nothing else needed
|
|
25
|
-
npm install @blockrun/llm
|
|
26
|
-
# or
|
|
27
|
-
pnpm add @blockrun/llm
|
|
28
|
-
# or
|
|
29
|
-
yarn add @blockrun/llm
|
|
65
|
+
npm install @blockrun/llm # Base / EVM payments โ smart routing included, nothing else needed
|
|
30
66
|
```
|
|
31
67
|
|
|
32
|
-
|
|
33
|
-
|
|
68
|
+
<details>
|
|
69
|
+
<summary><strong>Solana payments</strong> โ two more optional peers</summary>
|
|
34
70
|
|
|
35
71
|
```bash
|
|
36
72
|
npm install @blockrun/llm @solana/web3.js @solana/spl-token
|
|
@@ -44,63 +80,38 @@ every consumer, including projects that only ever pay on Base. As an optional
|
|
|
44
80
|
*peer* it reaches only the projects that ask for Solana. Calling a Solana path
|
|
45
81
|
without them throws an error naming the exact install command.
|
|
46
82
|
|
|
47
|
-
|
|
83
|
+
</details>
|
|
48
84
|
|
|
49
|
-
|
|
50
|
-
|
|
85
|
+
<details>
|
|
86
|
+
<summary><strong>Supported chains</strong> โ Base (primary), Base Sepolia, Solana</summary>
|
|
51
87
|
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
88
|
+
| Chain | Network | Payment | Status |
|
|
89
|
+
|-------|---------|---------|--------|
|
|
90
|
+
| **Base** | Base Mainnet (Chain ID: 8453) | USDC | Primary |
|
|
91
|
+
| **Base Testnet** | Base Sepolia (Chain ID: 84532) | Testnet USDC | Development |
|
|
92
|
+
| **Solana** | Solana Mainnet | USDC (SPL) | New |
|
|
55
93
|
|
|
56
|
-
|
|
94
|
+
**Protocol:** x402 v2 (CDP Facilitator)
|
|
57
95
|
|
|
58
|
-
|
|
96
|
+
</details>
|
|
59
97
|
|
|
60
|
-
|
|
61
|
-
**every** BlockRun endpoint over x402. New API surfaces are intended to be
|
|
62
|
-
distributed as [Claude Code skills](https://github.com/anthropics/skills)
|
|
63
|
-
that drive this primitive โ no SDK release required to add an endpoint.
|
|
98
|
+
## Quick Start (Base - Default)
|
|
64
99
|
|
|
65
100
|
```typescript
|
|
66
|
-
import {
|
|
67
|
-
|
|
68
|
-
const br = new BlockrunClient();
|
|
69
|
-
|
|
70
|
-
// Sync GET โ Surf market price (Tier 1, $0.001)
|
|
71
|
-
const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
|
|
101
|
+
import { LLMClient } from '@blockrun/llm';
|
|
72
102
|
|
|
73
|
-
|
|
74
|
-
const rows = await br.post('/v1/surf/onchain/sql', {
|
|
75
|
-
query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
|
|
76
|
-
});
|
|
103
|
+
const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
|
|
77
104
|
|
|
78
|
-
//
|
|
79
|
-
const
|
|
80
|
-
model: 'xai/grok-imagine-video',
|
|
81
|
-
prompt: 'a red apple spinning',
|
|
82
|
-
});
|
|
105
|
+
// Recommended: let the router pick the cheapest capable model
|
|
106
|
+
const result = await client.smartChat('Hello!');
|
|
83
107
|
|
|
84
|
-
//
|
|
85
|
-
|
|
86
|
-
model: 'anthropic/claude-sonnet-4-6',
|
|
87
|
-
messages: [{ role: 'user', content: 'Hi' }],
|
|
88
|
-
stream: true,
|
|
89
|
-
})) {
|
|
90
|
-
process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
|
|
91
|
-
}
|
|
108
|
+
// Or pin a model yourself
|
|
109
|
+
const response = await client.chat('openai/gpt-4o', 'Hello!');
|
|
92
110
|
```
|
|
93
111
|
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
- `poll<T>(path, body?, { budgetMs, intervalMs })` โ submit + poll (image, video, music, voice)
|
|
98
|
-
- `stream<T>(path, body?)` โ async iterator over SSE chunks (chat)
|
|
99
|
-
|
|
100
|
-
The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
|
|
101
|
-
`PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `RpcClient`,
|
|
102
|
-
`PriceClient`, `SurfClient`) all remain โ they will be soft-deprecated in 2.6 (rewritten as
|
|
103
|
-
shims over `BlockrunClient`) and removed in 3.0.
|
|
112
|
+
That's it. The SDK handles x402 payment automatically โ and `smartChat()`
|
|
113
|
+
keeps the bill down on every request. The router is bundled: no extra
|
|
114
|
+
package to install.
|
|
104
115
|
|
|
105
116
|
### Try It Free (No USDC Required)
|
|
106
117
|
|
|
@@ -114,30 +125,33 @@ const client = new LLMClient(); // Wallet still required for signing, but $0 ch
|
|
|
114
125
|
// Option 1: call a free model directly
|
|
115
126
|
const reply = await client.chat('nvidia/deepseek-v4-flash', 'Explain x402 in 1 sentence');
|
|
116
127
|
|
|
117
|
-
// Option 2: let the smart router pick the
|
|
118
|
-
const result = await client.smartChat('What is 2+2?', { routingProfile: '
|
|
119
|
-
console.log(result.model); //
|
|
128
|
+
// Option 2: let the smart router pick โ 'eco' ranks the free NVIDIA tier first
|
|
129
|
+
const result = await client.smartChat('What is 2+2?', { routingProfile: 'eco' });
|
|
130
|
+
console.log(result.model); // 'nvidia/deepseek-v4-flash' ($0 โ verified live)
|
|
120
131
|
console.log(result.response); // '4'
|
|
132
|
+
console.log(result.routing.savings); // 1 (100%)
|
|
121
133
|
```
|
|
122
134
|
|
|
123
|
-
|
|
135
|
+
There is no `free` routing profile in `smartChat()` โ `routingProfile` accepts
|
|
136
|
+
`'eco' | 'auto' | 'premium'`. (ClawRouter's `/model free` is a feature of its
|
|
137
|
+
own proxy, not of this SDK's router options.) For guaranteed $0, pin a
|
|
138
|
+
`nvidia/*` model; for smart-routed $0-first, use `eco`.
|
|
139
|
+
|
|
140
|
+
**Available free models** (input + output both $0, all NVIDIA-hosted, from the live `/v1/models` catalog, last refreshed 2026-08-10):
|
|
124
141
|
|
|
125
142
|
| Model ID | Context | Best For |
|
|
126
143
|
|----------|---------|----------|
|
|
127
|
-
| `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash โ 284B / 13B active MoE
|
|
128
|
-
| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K |
|
|
129
|
-
| `nvidia/
|
|
130
|
-
| `nvidia/
|
|
131
|
-
| `nvidia/
|
|
132
|
-
| `nvidia/
|
|
133
|
-
| `nvidia/gpt-oss-
|
|
134
|
-
|
|
135
|
-
> Need V4-Pro-class reasoning? Use the paid `deepseek/deepseek-v4-pro` ($0.435/$0.87 โ the 75% launch promo became the permanent list price after 2026-05-31) โ `nvidia/deepseek-v4-pro` is currently hidden because NVIDIA's NIM deployment is hung; backend MODEL_REDIRECTS forwards calls to V4 Flash.
|
|
144
|
+
| `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash โ 284B / 13B active MoE. Best free chat / summarization / light reasoning. Capacity-constrained: requests may be answered by an equivalent free model |
|
|
145
|
+
| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning โ text + images + video + audio (ChartQA 90.3, DocVQA 95.6) |
|
|
146
|
+
| `nvidia/mistral-nemotron` | 131K | Mistral ร NVIDIA instruction model โ fast (~0.2s), strong instruction following |
|
|
147
|
+
| `nvidia/step-3.7-flash` | 131K | StepFun Step 3.7 Flash โ fast lightweight reasoning |
|
|
148
|
+
| `nvidia/nemotron-nano-9b-v2` | 131K | Compact + fast (~0.7s), good for high-volume light tasks |
|
|
149
|
+
| `nvidia/nemotron-nano-12b-v2-vl` | 131K | Vision-language โ accepts images, compact + fast |
|
|
150
|
+
| `nvidia/gpt-oss-120b` | 128K | OpenAI open-weight 120B. Hidden from `/v1/models` for privacy but direct calls still work |
|
|
151
|
+
| `nvidia/gpt-oss-20b` | 128K | OpenAI open-weight 20B. Hidden from `/v1/models` but direct calls still work |
|
|
136
152
|
|
|
137
153
|
> Privacy note: `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` are hidden from `/v1/models` because NVIDIA's free build.nvidia.com tier reserves the right to use prompts/outputs for service improvement. Direct calls by full model ID still work โ opt in only when your data isn't sensitive.
|
|
138
154
|
|
|
139
|
-
> Retired: `nvidia/qwen3-next-80b-a3b-thinking` hit NVIDIA end-of-life 2026-05-21 (HTTP 410). The gateway auto-redirects pinned callers to `nvidia/llama-4-maverick`.
|
|
140
|
-
|
|
141
155
|
## Quick Start (Solana)
|
|
142
156
|
|
|
143
157
|
```typescript
|
|
@@ -151,90 +165,7 @@ console.log(response);
|
|
|
151
165
|
|
|
152
166
|
Set `SOLANA_WALLET_KEY` to your bs58-encoded Solana secret key. Payments are automatic via x402 โ your key never leaves your machine.
|
|
153
167
|
|
|
154
|
-
##
|
|
155
|
-
|
|
156
|
-
Pay for AI calls with Solana USDC via [sol.blockrun.ai](https://sol.blockrun.ai):
|
|
157
|
-
|
|
158
|
-
```typescript
|
|
159
|
-
import { SolanaLLMClient } from '@blockrun/llm';
|
|
160
|
-
|
|
161
|
-
// SOLANA_WALLET_KEY env var (bs58-encoded Solana secret key)
|
|
162
|
-
const client = new SolanaLLMClient();
|
|
163
|
-
|
|
164
|
-
// Or pass key directly
|
|
165
|
-
const client2 = new SolanaLLMClient({ privateKey: 'your-bs58-solana-key' });
|
|
166
|
-
|
|
167
|
-
// Same API as LLMClient
|
|
168
|
-
const response = await client.chat('openai/gpt-4o', 'gm Solana');
|
|
169
|
-
console.log(response);
|
|
170
|
-
|
|
171
|
-
// Live Search with Grok (Solana payment)
|
|
172
|
-
const tweet = await client.chat('xai/grok-3-mini', 'What is trending on X?', { search: true });
|
|
173
|
-
```
|
|
174
|
-
|
|
175
|
-
**Setup:**
|
|
176
|
-
1. Export your Solana wallet key: `export SOLANA_WALLET_KEY="your-bs58-key"`
|
|
177
|
-
2. Fund with USDC on Solana mainnet
|
|
178
|
-
3. That's it โ payments are automatic via x402
|
|
179
|
-
|
|
180
|
-
**Supported endpoint:** `https://sol.blockrun.ai/api`
|
|
181
|
-
**Payment:** Solana USDC (SPL, mainnet)
|
|
182
|
-
|
|
183
|
-
## How Payment Works
|
|
184
|
-
|
|
185
|
-
No API keys, no subscription. You hold USDC in your own wallet, and **every request pays for itself** with an on-chain micropayment. Two phases:
|
|
186
|
-
|
|
187
|
-
### Phase 1 โ Fund your wallet once
|
|
188
|
-
|
|
189
|
-
You only do this when your balance runs low. Three ways to get USDC into your wallet:
|
|
190
|
-
|
|
191
|
-
- **(a) Buy with a card (Base USDC).** Call the new `onramp()` method to mint a one-time Coinbase Onramp link, then open the returned `pay.coinbase.com` URL โ pay by card/bank in 60+ fiat currencies and the USDC lands in your wallet. The call itself is **free**. Onramp is **Base-only** (buying USDC with a card always lands Base USDC), and the funding address must equal your signing wallet:
|
|
192
|
-
|
|
193
|
-
```typescript
|
|
194
|
-
const { url } = await client.onramp(client.getWalletAddress());
|
|
195
|
-
console.log(`Fund your wallet: ${url}`); // single-use, expires ~5 min โ mint at click time
|
|
196
|
-
```
|
|
197
|
-
|
|
198
|
-
- **(b) Transfer existing USDC.** Send USDC you already hold to your wallet address (`client.getWalletAddress()`). On Base, send Base USDC; on Solana (`SolanaLLMClient`), send Solana SPL USDC.
|
|
199
|
-
|
|
200
|
-
- **(c) Skip funding entirely.** Use the free NVIDIA models (e.g. `nvidia/deepseek-v4-flash`) โ every call is **$0**, no balance required.
|
|
201
|
-
|
|
202
|
-
$5 of USDC covers thousands of paid requests. Check your balance any time:
|
|
203
|
-
|
|
204
|
-
```typescript
|
|
205
|
-
const balance = await client.getBalance(); // USDC on Base
|
|
206
|
-
console.log(`Balance: $${balance.toFixed(2)} USDC`);
|
|
207
|
-
```
|
|
208
|
-
|
|
209
|
-
### Phase 2 โ Every request pays itself (automatic x402)
|
|
210
|
-
|
|
211
|
-
You just call e.g. `client.chat(...)` โ the payment is invisible:
|
|
212
|
-
|
|
213
|
-
1. You send a request to BlockRun's API.
|
|
214
|
-
2. The gateway returns **402 Payment Required** with the price.
|
|
215
|
-
3. The SDK signs a USDC payment **locally** (EIP-712) โ on **Base** for `LLMClient`, on **Solana** for `SolanaLLMClient` โ using your wallet key.
|
|
216
|
-
4. The request is retried automatically with the payment proof.
|
|
217
|
-
5. The gateway settles on-chain and returns the AI response.
|
|
218
|
-
|
|
219
|
-
One call, no separate pay step. Free NVIDIA models settle at **$0** (no payment signed).
|
|
220
|
-
|
|
221
|
-
### Track spend and verify settlements
|
|
222
|
-
|
|
223
|
-
```typescript
|
|
224
|
-
import { getCostSummary } from '@blockrun/llm';
|
|
225
|
-
|
|
226
|
-
const spent = client.getSpending(); // this session
|
|
227
|
-
console.log(`Spent $${spent.totalUsd.toFixed(4)} across ${spent.calls} calls`);
|
|
228
|
-
|
|
229
|
-
const summary = getCostSummary(); // across sessions (~/.blockrun/data/costs.jsonl)
|
|
230
|
-
console.log(`Lifetime: $${summary.totalUsd.toFixed(2)} over ${summary.calls} calls`);
|
|
231
|
-
```
|
|
232
|
-
|
|
233
|
-
Every paid request is a real on-chain USDC transfer โ look up your wallet address on [Basescan](https://basescan.org) (or a Solana explorer) to verify each settlement independently.
|
|
234
|
-
|
|
235
|
-
**Non-custodial by design: your private key never leaves your machine** โ it is only used for local signing, and no funds are ever held by BlockRun.
|
|
236
|
-
|
|
237
|
-
## Smart Routing (ClawRouter)
|
|
168
|
+
## Smart Routing (Router Core V3)
|
|
238
169
|
|
|
239
170
|
Let the SDK automatically pick the cheapest capable model for each request โ **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** for the same traffic on `auto`, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco`.
|
|
240
171
|
|
|
@@ -243,17 +174,34 @@ published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/ma
|
|
|
243
174
|
priced against the live catalog, so anyone can recompute the claim and get the
|
|
244
175
|
same answer.
|
|
245
176
|
|
|
246
|
-
Smart routing is powered by
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
install
|
|
177
|
+
Smart routing is powered by the product-neutral
|
|
178
|
+
[`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) V3 engine โ
|
|
179
|
+
the same deterministic portfolio router that drives
|
|
180
|
+
[ClawRouter](https://github.com/BlockRunAI/ClawRouter). It is **bundled into
|
|
181
|
+
this SDK**: no separate router package to install, and routing runs 100%
|
|
182
|
+
locally with zero external calls.
|
|
251
183
|
|
|
252
|
-
|
|
253
|
-
|
|
184
|
+
Three ways to use it:
|
|
185
|
+
|
|
186
|
+
```typescript
|
|
187
|
+
// 1. smartChat() โ one-line routed chat
|
|
188
|
+
const result = await client.smartChat('What is 2+2?');
|
|
189
|
+
|
|
190
|
+
// 2. smartChatCompletion() โ full agent/tool conversations, routed
|
|
191
|
+
const agent = await client.smartChatCompletion(messages, { tools, toolChoice: 'auto' });
|
|
192
|
+
|
|
193
|
+
// 3. blockrun/auto | blockrun/eco | blockrun/premium โ model aliases accepted
|
|
194
|
+
// by chat(), chatCompletion(), and chatCompletionStream() on both chains
|
|
195
|
+
const reply = await client.chatCompletion('blockrun/auto', messages);
|
|
196
|
+
|
|
197
|
+
// Inspect a decision without paying for anything
|
|
198
|
+
const decision = await client.route('Prove the Riemann hypothesis');
|
|
254
199
|
```
|
|
255
200
|
|
|
256
|
-
|
|
201
|
+
The aliases are resolved locally by `LLMClient`, `SolanaLLMClient`, and the
|
|
202
|
+
OpenAI-compat layer. The Anthropic-compat layer proxies straight to the
|
|
203
|
+
gateway's `/v1/messages` and does **not** resolve them โ pass a concrete
|
|
204
|
+
model id there.
|
|
257
205
|
|
|
258
206
|
```typescript
|
|
259
207
|
import { LLMClient } from '@blockrun/llm';
|
|
@@ -300,11 +248,14 @@ const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
|
|
|
300
248
|
|
|
301
249
|
| Profile | Strategy | Savings vs Opus 5 | Best For |
|
|
302
250
|
|---------|----------|-------------------|----------|
|
|
303
|
-
| `
|
|
304
|
-
| `eco` | Cheapest capable model per tier | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production |
|
|
251
|
+
| `eco` | Cheapest capable model โ ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
|
|
305
252
|
| `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%** | General use |
|
|
306
253
|
| `premium` | Top-tier models (OpenAI, Anthropic) | 0% | Quality-critical tasks |
|
|
307
254
|
|
|
255
|
+
For guaranteed $0, call a `nvidia/*` model directly with `chat()` โ see
|
|
256
|
+
[Try It Free](#try-it-free-no-usdc-required). ClawRouter's `/model free`
|
|
257
|
+
profile belongs to its own proxy; `smartChat()`'s options are the three above.
|
|
258
|
+
|
|
308
259
|
```typescript
|
|
309
260
|
// Use premium models for complex tasks
|
|
310
261
|
const result = await client.smartChat(
|
|
@@ -314,7 +265,17 @@ const result = await client.smartChat(
|
|
|
314
265
|
console.log(result.model); // 'anthropic/claude-opus-4.7'
|
|
315
266
|
```
|
|
316
267
|
|
|
317
|
-
### How
|
|
268
|
+
### How the Router Works
|
|
269
|
+
|
|
270
|
+
```mermaid
|
|
271
|
+
flowchart LR
|
|
272
|
+
A["prompt"] --> B["classify locally<br/>15 dimensions, <1ms"]
|
|
273
|
+
B --> C["hard filters<br/>tools ยท vision ยท context ยท<br/>structured output"]
|
|
274
|
+
C --> D["rank portfolio<br/>quality ยท cost ยท speed ยท<br/>reliability"]
|
|
275
|
+
D --> E["cheapest capable model<br/>+ ranked fallback chain"]
|
|
276
|
+
E --> F["x402 USDC payment<br/>for this request only"]
|
|
277
|
+
F --> G["response<br/>+ full routing metadata"]
|
|
278
|
+
```
|
|
318
279
|
|
|
319
280
|
Since ClawRouter v0.12.242, Auto uses the deterministic **Router v3.4 portfolio
|
|
320
281
|
strategy**: it classifies the task shape locally across
|
|
@@ -371,14 +332,14 @@ picked:
|
|
|
371
332
|
|
|
372
333
|
### TypeScript Types
|
|
373
334
|
|
|
374
|
-
`RoutingDecision`, `RoutingProfile`, `RoutingTier`, `RoutingTaskType`,
|
|
375
|
-
`RoutingTierConfig`
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
335
|
+
`RoutingDecision`, `RoutingProfile`, `RoutingTier`, `RoutingTaskType`,
|
|
336
|
+
`RoutingTierConfig`, `SmartChatCompletionOptions`, and
|
|
337
|
+
`SmartChatCompletionResponse` are exported from `@blockrun/llm`. They are
|
|
338
|
+
derived from
|
|
339
|
+
[`@blockrun/router-core`](https://github.com/BlockRunAI/router-core), pinned
|
|
340
|
+
to a reviewed immutable commit, and shipped **inlined in this SDK's
|
|
341
|
+
declaration files and runtime bundle** โ you install nothing extra to route
|
|
342
|
+
or to typecheck.
|
|
382
343
|
|
|
383
344
|
### Going Deeper
|
|
384
345
|
|
|
@@ -389,6 +350,136 @@ runtime package is only needed to actually call `smartChat()`.
|
|
|
389
350
|
- [ClawRouter vs OpenRouter](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/clawrouter-vs-openrouter-llm-routing-comparison.md) โ head-to-head comparison
|
|
390
351
|
- [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) โ the deterministic routing engine both share
|
|
391
352
|
|
|
353
|
+
## Solana Support
|
|
354
|
+
|
|
355
|
+
Pay for AI calls with Solana USDC via [sol.blockrun.ai](https://sol.blockrun.ai):
|
|
356
|
+
|
|
357
|
+
```typescript
|
|
358
|
+
import { SolanaLLMClient } from '@blockrun/llm';
|
|
359
|
+
|
|
360
|
+
// SOLANA_WALLET_KEY env var (bs58-encoded Solana secret key)
|
|
361
|
+
const client = new SolanaLLMClient();
|
|
362
|
+
|
|
363
|
+
// Or pass key directly
|
|
364
|
+
const client2 = new SolanaLLMClient({ privateKey: 'your-bs58-solana-key' });
|
|
365
|
+
|
|
366
|
+
// Same API as LLMClient
|
|
367
|
+
const response = await client.chat('openai/gpt-4o', 'gm Solana');
|
|
368
|
+
console.log(response);
|
|
369
|
+
|
|
370
|
+
// Live Search with Grok (Solana payment)
|
|
371
|
+
const tweet = await client.chat('xai/grok-3-mini', 'What is trending on X?', { search: true });
|
|
372
|
+
```
|
|
373
|
+
|
|
374
|
+
**Setup:**
|
|
375
|
+
1. Export your Solana wallet key: `export SOLANA_WALLET_KEY="your-bs58-key"`
|
|
376
|
+
2. Fund with USDC on Solana mainnet
|
|
377
|
+
3. That's it โ payments are automatic via x402
|
|
378
|
+
|
|
379
|
+
**Supported endpoint:** `https://sol.blockrun.ai/api`
|
|
380
|
+
**Payment:** Solana USDC (SPL, mainnet)
|
|
381
|
+
|
|
382
|
+
## How Payment Works
|
|
383
|
+
|
|
384
|
+
No API keys, no subscription. You hold USDC in your own wallet, and **every request pays for itself** with an on-chain micropayment. Two phases:
|
|
385
|
+
|
|
386
|
+
### Phase 1 โ Fund your wallet once
|
|
387
|
+
|
|
388
|
+
You only do this when your balance runs low. Three ways to get USDC into your wallet:
|
|
389
|
+
|
|
390
|
+
- **(a) Buy with a card (Base USDC).** Call the new `onramp()` method to mint a one-time Coinbase Onramp link, then open the returned `pay.coinbase.com` URL โ pay by card/bank in 60+ fiat currencies and the USDC lands in your wallet. The call itself is **free**. Onramp is **Base-only** (buying USDC with a card always lands Base USDC), and the funding address must equal your signing wallet:
|
|
391
|
+
|
|
392
|
+
```typescript
|
|
393
|
+
const { url } = await client.onramp(client.getWalletAddress());
|
|
394
|
+
console.log(`Fund your wallet: ${url}`); // single-use, expires ~5 min โ mint at click time
|
|
395
|
+
```
|
|
396
|
+
|
|
397
|
+
- **(b) Transfer existing USDC.** Send USDC you already hold to your wallet address (`client.getWalletAddress()`). On Base, send Base USDC; on Solana (`SolanaLLMClient`), send Solana SPL USDC.
|
|
398
|
+
|
|
399
|
+
- **(c) Skip funding entirely.** Use the free NVIDIA models (e.g. `nvidia/deepseek-v4-flash`) โ every call is **$0**, no balance required.
|
|
400
|
+
|
|
401
|
+
$5 of USDC covers thousands of paid requests. Check your balance any time:
|
|
402
|
+
|
|
403
|
+
```typescript
|
|
404
|
+
const balance = await client.getBalance(); // USDC on Base
|
|
405
|
+
console.log(`Balance: $${balance.toFixed(2)} USDC`);
|
|
406
|
+
```
|
|
407
|
+
|
|
408
|
+
### Phase 2 โ Every request pays itself (automatic x402)
|
|
409
|
+
|
|
410
|
+
You just call e.g. `client.chat(...)` โ the payment is invisible:
|
|
411
|
+
|
|
412
|
+
1. You send a request to BlockRun's API.
|
|
413
|
+
2. The gateway returns **402 Payment Required** with the price.
|
|
414
|
+
3. The SDK signs a USDC payment **locally** (EIP-712) โ on **Base** for `LLMClient`, on **Solana** for `SolanaLLMClient` โ using your wallet key.
|
|
415
|
+
4. The request is retried automatically with the payment proof.
|
|
416
|
+
5. The gateway settles on-chain and returns the AI response.
|
|
417
|
+
|
|
418
|
+
One call, no separate pay step. Free NVIDIA models settle at **$0** (no payment signed).
|
|
419
|
+
|
|
420
|
+
### Track spend and verify settlements
|
|
421
|
+
|
|
422
|
+
```typescript
|
|
423
|
+
import { getCostSummary } from '@blockrun/llm';
|
|
424
|
+
|
|
425
|
+
const spent = client.getSpending(); // this session
|
|
426
|
+
console.log(`Spent $${spent.totalUsd.toFixed(4)} across ${spent.calls} calls`);
|
|
427
|
+
|
|
428
|
+
const summary = getCostSummary(); // across sessions (~/.blockrun/data/costs.jsonl)
|
|
429
|
+
console.log(`Lifetime: $${summary.totalUsd.toFixed(2)} over ${summary.calls} calls`);
|
|
430
|
+
```
|
|
431
|
+
|
|
432
|
+
Every paid request is a real on-chain USDC transfer โ look up your wallet address on [Basescan](https://basescan.org) (or a Solana explorer) to verify each settlement independently.
|
|
433
|
+
|
|
434
|
+
**Non-custodial by design: your private key never leaves your machine** โ it is only used for local signing, and no funds are ever held by BlockRun.
|
|
435
|
+
|
|
436
|
+
## `BlockrunClient` โ the universal primitive (recommended for new code)
|
|
437
|
+
|
|
438
|
+
Starting in `2.5.0`, the SDK ships a single `BlockrunClient` that speaks to
|
|
439
|
+
**every** BlockRun endpoint over x402. New API surfaces are intended to be
|
|
440
|
+
distributed as [Claude Code skills](https://github.com/anthropics/skills)
|
|
441
|
+
that drive this primitive โ no SDK release required to add an endpoint.
|
|
442
|
+
|
|
443
|
+
```typescript
|
|
444
|
+
import { BlockrunClient } from '@blockrun/llm';
|
|
445
|
+
|
|
446
|
+
const br = new BlockrunClient();
|
|
447
|
+
|
|
448
|
+
// Sync GET โ Surf market price (Tier 1, $0.001)
|
|
449
|
+
const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
|
|
450
|
+
|
|
451
|
+
// Sync POST โ raw on-chain SQL (Tier 3, $0.020)
|
|
452
|
+
const rows = await br.post('/v1/surf/onchain/sql', {
|
|
453
|
+
query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
|
|
454
|
+
});
|
|
455
|
+
|
|
456
|
+
// Submit + poll โ long-running video gen (settled only on completion)
|
|
457
|
+
const video = await br.poll('/v1/videos/generations', {
|
|
458
|
+
model: 'xai/grok-imagine-video',
|
|
459
|
+
prompt: 'a red apple spinning',
|
|
460
|
+
});
|
|
461
|
+
|
|
462
|
+
// Streaming SSE โ chat completions
|
|
463
|
+
for await (const chunk of br.stream('/v1/chat/completions', {
|
|
464
|
+
model: 'anthropic/claude-sonnet-4-6',
|
|
465
|
+
messages: [{ role: 'user', content: 'Hi' }],
|
|
466
|
+
stream: true,
|
|
467
|
+
})) {
|
|
468
|
+
process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
|
|
469
|
+
}
|
|
470
|
+
```
|
|
471
|
+
|
|
472
|
+
Four call shapes cover every endpoint type:
|
|
473
|
+
- `get<T>(path, params?)` โ synchronous GET (price, ranking, list, news)
|
|
474
|
+
- `post<T>(path, body?)` โ synchronous POST (on-chain SQL, search)
|
|
475
|
+
- `poll<T>(path, body?, { budgetMs, intervalMs })` โ submit + poll (image, video, music, voice)
|
|
476
|
+
- `stream<T>(path, body?)` โ async iterator over SSE chunks (chat)
|
|
477
|
+
|
|
478
|
+
The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
|
|
479
|
+
`PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `RpcClient`,
|
|
480
|
+
`PriceClient`, `SurfClient`) all remain โ they will be soft-deprecated in 2.6 (rewritten as
|
|
481
|
+
shims over `BlockrunClient`) and removed in 3.0.
|
|
482
|
+
|
|
392
483
|
## Available Models
|
|
393
484
|
|
|
394
485
|
### OpenAI GPT-5.5 Family
|
|
@@ -948,9 +1039,9 @@ const response2 = await client.chat('anthropic/claude-sonnet-4', 'Write a haiku'
|
|
|
948
1039
|
});
|
|
949
1040
|
```
|
|
950
1041
|
|
|
951
|
-
### Smart Routing (
|
|
1042
|
+
### Smart Routing (Router Core V3)
|
|
952
1043
|
|
|
953
|
-
Save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing.
|
|
1044
|
+
Save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. The bundled Router Core V3 engine classifies each request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions, applies hard capability filters, and ranks the cheapest capable models (<1ms, 100% local). Bundled โ nothing extra to install.
|
|
954
1045
|
|
|
955
1046
|
```typescript
|
|
956
1047
|
import { LLMClient } from '@blockrun/llm';
|
|
@@ -964,19 +1055,20 @@ console.log(result.model); // 'google/gemini-2.5-flash'
|
|
|
964
1055
|
console.log(result.routing.tier); // 'SIMPLE'
|
|
965
1056
|
console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 88%'
|
|
966
1057
|
|
|
967
|
-
// Routing profiles
|
|
968
|
-
const
|
|
969
|
-
const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Budget optimized
|
|
1058
|
+
// Routing profiles ('eco' | 'auto' | 'premium')
|
|
1059
|
+
const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Free tier first, then cheapest paid
|
|
970
1060
|
const auto = await client.smartChat('Code review', { routingProfile: 'auto' }); // Balanced (default)
|
|
971
1061
|
const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
|
|
1062
|
+
|
|
1063
|
+
// Guaranteed $0: call a free NVIDIA model directly
|
|
1064
|
+
const free = await client.chat('nvidia/deepseek-v4-flash', 'Hello!');
|
|
972
1065
|
```
|
|
973
1066
|
|
|
974
1067
|
**Routing Profiles:**
|
|
975
1068
|
|
|
976
1069
|
| Profile | Description | Best For |
|
|
977
1070
|
|---------|-------------|----------|
|
|
978
|
-
| `
|
|
979
|
-
| `eco` | Budget-optimized | Cost-sensitive workloads |
|
|
1071
|
+
| `eco` | Budget-optimized โ ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | Cost-sensitive workloads, zero-cost testing |
|
|
980
1072
|
| `auto` | Intelligent routing (default) | General use |
|
|
981
1073
|
| `premium` | Best quality models | Critical tasks |
|
|
982
1074
|
|
|
@@ -1513,13 +1605,13 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
|
|
|
1513
1605
|
## Frequently Asked Questions
|
|
1514
1606
|
|
|
1515
1607
|
### What is @blockrun/llm?
|
|
1516
|
-
@blockrun/llm is a TypeScript SDK that
|
|
1608
|
+
@blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol โ no API keys, no subscriptions, no vendor lock-in.
|
|
1517
1609
|
|
|
1518
1610
|
### How does payment work?
|
|
1519
1611
|
When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
|
|
1520
1612
|
|
|
1521
|
-
### What is smart routing
|
|
1522
|
-
|
|
1613
|
+
### What is smart routing?
|
|
1614
|
+
Router Core V3 is bundled into the SDK โ the same deterministic routing engine that powers ClawRouter, with nothing extra to install. It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. Use `smartChat()`, `smartChatCompletion()`, or the `blockrun/auto` model alias. It can save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
|
|
1523
1615
|
|
|
1524
1616
|
### Does it support streaming?
|
|
1525
1617
|
Yes โ as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).
|
|
@@ -1530,6 +1622,16 @@ Pay only for what you use. Prices start at $0.0002 per request (GPT-5 Nano). The
|
|
|
1530
1622
|
### Does it support both Base and Solana?
|
|
1531
1623
|
Yes. Use `LLMClient` for Base (EVM) payments and `SolanaLLMClient` for Solana payments. Same API, different payment chain.
|
|
1532
1624
|
|
|
1625
|
+
---
|
|
1626
|
+
|
|
1627
|
+
<div align="center">
|
|
1628
|
+
|
|
1629
|
+
**If the router just cut your bill, [give it a star โญ](https://github.com/BlockRunAI/blockrun-llm-ts)** โ it helps more agents pay less.
|
|
1630
|
+
|
|
1631
|
+
[Website](https://blockrun.ai) ยท [Models & Pricing](https://blockrun.ai/models) ยท [ClawRouter](https://github.com/BlockRunAI/ClawRouter) ยท [Python SDK](https://github.com/BlockRunAI/blockrun-llm) ยท [Telegram](https://t.me/blockrunAI)
|
|
1632
|
+
|
|
1633
|
+
</div>
|
|
1634
|
+
|
|
1533
1635
|
## License
|
|
1534
1636
|
|
|
1535
1637
|
MIT
|