@blockrun/llm 3.10.0 โ†’ 3.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,36 +1,72 @@
1
- # @blockrun/llm (TypeScript SDK)
1
+ <div align="center">
2
2
 
3
- > **@blockrun/llm** is a TypeScript/Node.js SDK for accessing <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> large language models (GPT-5, Claude, Gemini, Grok, DeepSeek, Kimi, and more) with automatic pay-per-request USDC micropayments via the x402 protocol. No API keys required โ€” your wallet signature is your authentication. Supports **streaming**, smart routing, Base and Solana chains.
4
- >
5
- > ๐Ÿ†“ **Includes 7 fully-free NVIDIA-hosted models** (5 visible in `/v1/models`, 2 hidden but directly callable) โ€” DeepSeek V4 Flash (1M context), Nemotron Nano Omni (vision), Qwen3 Coder, Llama 4, Mistral, plus the gpt-oss pair. Zero USDC, no rate-limit gimmicks. Use `routingProfile: 'free'` or call any `nvidia/*` model directly.
3
+ # @blockrun/llm
6
4
 
7
- [![npm](https://img.shields.io/npm/v/@blockrun/llm.svg)](https://www.npmjs.com/package/@blockrun/llm)
8
- [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE)
5
+ ### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
9
6
 
10
- ## Supported Chains
7
+ The smart-routing SDK for <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models โ€” every request goes to the cheapest model that can handle it,
8
+ paid per-request in USDC. No API keys. No subscriptions. No vendor lock-in.
11
9
 
12
- | Chain | Network | Payment | Status |
13
- |-------|---------|---------|--------|
14
- | **Base** | Base Mainnet (Chain ID: 8453) | USDC | Primary |
15
- | **Base Testnet** | Base Sepolia (Chain ID: 84532) | Testnet USDC | Development |
16
- | **Solana** | Solana Mainnet | USDC (SPL) | New |
10
+ [![npm](https://img.shields.io/npm/v/@blockrun/llm.svg?style=flat-square)](https://www.npmjs.com/package/@blockrun/llm)
11
+ [![npm downloads](https://img.shields.io/npm/dm/@blockrun/llm.svg?style=flat-square)](https://www.npmjs.com/package/@blockrun/llm)
12
+ [![CI](https://img.shields.io/github/actions/workflow/status/BlockRunAI/blockrun-llm-ts/ci.yml?branch=main&style=flat-square&label=CI)](https://github.com/BlockRunAI/blockrun-llm-ts/actions)
13
+ [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg?style=flat-square)](LICENSE)
14
+ [![Node](https://img.shields.io/badge/Node-%E2%89%A520-brightgreen?style=flat-square&logo=node.js&logoColor=white)](package.json)
15
+ [![TypeScript](https://img.shields.io/badge/TypeScript-strict-3178C6?style=flat-square&logo=typescript&logoColor=white)](tsconfig.json)
17
16
 
17
+ [![Base Network](https://img.shields.io/badge/Base-USDC-0052FF?style=flat-square&logo=coinbase&logoColor=white)](https://base.org)
18
+ [![Solana](https://img.shields.io/badge/Solana-USDC-9945FF?style=flat-square&logo=solana&logoColor=white)](https://solana.com)
19
+ [![x402](https://img.shields.io/badge/x402-micropayments-orange?style=flat-square)](https://x402.org)
20
+ [![Telegram](https://img.shields.io/badge/Telegram-Community-26A5E4?style=flat-square&logo=telegram)](https://t.me/blockrunAI)
18
21
 
19
- **Protocol:** x402 v2 (CDP Facilitator)
22
+ [Website](https://blockrun.ai) ยท [Models & Pricing](https://blockrun.ai/models) ยท [ClawRouter](https://github.com/BlockRunAI/ClawRouter) ยท [Python SDK](https://github.com/BlockRunAI/blockrun-llm) ยท [Telegram](https://t.me/blockrunAI)
23
+
24
+ </div>
25
+
26
+ ---
27
+
28
+ ```typescript
29
+ import { LLMClient } from '@blockrun/llm';
30
+
31
+ const client = new LLMClient();
32
+
33
+ const r = await client.smartChat('Prove step by step that the sum of two odd integers is even.');
34
+ console.log(r.model); // 'deepseek/deepseek-v4-pro' โ€” the right model, not the $75/M flagship
35
+ console.log(r.routing.savings); // 0.96 โ€” this exact request cost 96% less than pinning the baseline
36
+ console.log(r.response); // the proof
37
+ ```
38
+
39
+ **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** across a realistic workload on the default `auto` profile, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco` โ€” and eco's first stop is the free tier, so simple requests cost $0.00 outright. Not an "up to" figure: the baseline, workload mix, and token ratio are published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json) so anyone can recompute the claim. Details in [Smart Routing](#smart-routing-router-core-v3).
40
+
41
+ ## Why This SDK
42
+
43
+ - ๐Ÿง  **Smart routing that pays for itself** โ€” the bundled [Router Core V3](https://github.com/BlockRunAI/router-core) engine (shared with [ClawRouter](https://github.com/BlockRunAI/ClawRouter)) classifies every request locally in <1ms across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and routes to the cheapest capable model. The main event.
44
+ - ๐Ÿ†“ **<!-- br:models.free -->6<!-- /br:models.free --> genuinely free models** โ€” NVIDIA-hosted, $0 in and out, incl. 1M-context DeepSeek V4 Flash and a multimodal Nemotron. No rate-limit gimmicks.
45
+ - ๐Ÿ” **No API keys** โ€” your wallet signature is your authentication. No accounts, no dashboards, no key rotation.
46
+ - ๐Ÿ’ธ **Pay per request in USDC** โ€” x402 micropayments on Base or Solana. $5 covers thousands of requests; agents can pay their own way.
47
+ - ๐Ÿ›ก๏ธ **Automatic failover** โ€” transient errors (timeouts, 429, 5xx) walk the router's ranked fallback chain instead of failing your request.
48
+ - โšก **Streaming, OpenAI & Anthropic compat** โ€” drop-in `chat.completions` / `messages` layers, SSE streaming, strict TypeScript.
49
+ - ๐ŸŽจ **Beyond chat** โ€” image, video, music, speech, live search, prediction markets, crypto data, and 40-chain RPC through the same wallet.
50
+
51
+ ## How It Compares
52
+
53
+ | | OpenAI SDK | OpenRouter | LiteLLM | **@blockrun/llm** |
54
+ | ------------------ | -------------- | ----------------- | ---------------- | ----------------------------------------------------------------------- |
55
+ | **Cost routing** | โœ— one vendor | Manual selection | Manual selection | **Automatic โ€” <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
56
+ | **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->71<!-- /br:models.chatVisible -->, one wallet** |
57
+ | **Free tier** | โœ— | Rate-limited | โœ— | **<!-- br:models.free -->6<!-- /br:models.free --> models, no signup** |
58
+ | **Auth** | API key | Account + API key | Your API keys | **Wallet signature** |
59
+ | **Payment** | Card + invoice | Credit card | BYO keys | **USDC per-request** |
60
+ | **Agent-ready** | โœ— | โœ— | โœ— | **โœ“ โ€” agents fund their own wallet** |
20
61
 
21
62
  ## Installation
22
63
 
23
64
  ```bash
24
- # Base / EVM payments โ€” nothing else needed
25
- npm install @blockrun/llm
26
- # or
27
- pnpm add @blockrun/llm
28
- # or
29
- yarn add @blockrun/llm
65
+ npm install @blockrun/llm # Base / EVM payments โ€” smart routing included, nothing else needed
30
66
  ```
31
67
 
32
- **Solana payments** need two more packages. They are optional peer dependencies,
33
- so npm will not install them for you:
68
+ <details>
69
+ <summary><strong>Solana payments</strong> โ€” two more optional peers</summary>
34
70
 
35
71
  ```bash
36
72
  npm install @blockrun/llm @solana/web3.js @solana/spl-token
@@ -44,63 +80,38 @@ every consumer, including projects that only ever pay on Base. As an optional
44
80
  *peer* it reaches only the projects that ask for Solana. Calling a Solana path
45
81
  without them throws an error naming the exact install command.
46
82
 
47
- ## Quick Start (Base - Default)
83
+ </details>
48
84
 
49
- ```typescript
50
- import { LLMClient } from '@blockrun/llm';
85
+ <details>
86
+ <summary><strong>Supported chains</strong> โ€” Base (primary), Base Sepolia, Solana</summary>
51
87
 
52
- const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
53
- const response = await client.chat('openai/gpt-4o', 'Hello!');
54
- ```
88
+ | Chain | Network | Payment | Status |
89
+ |-------|---------|---------|--------|
90
+ | **Base** | Base Mainnet (Chain ID: 8453) | USDC | Primary |
91
+ | **Base Testnet** | Base Sepolia (Chain ID: 84532) | Testnet USDC | Development |
92
+ | **Solana** | Solana Mainnet | USDC (SPL) | New |
55
93
 
56
- That's it. The SDK handles x402 payment automatically.
94
+ **Protocol:** x402 v2 (CDP Facilitator)
57
95
 
58
- ## `BlockrunClient` โ€” the universal primitive (recommended for new code)
96
+ </details>
59
97
 
60
- Starting in `2.5.0`, the SDK ships a single `BlockrunClient` that speaks to
61
- **every** BlockRun endpoint over x402. New API surfaces are intended to be
62
- distributed as [Claude Code skills](https://github.com/anthropics/skills)
63
- that drive this primitive โ€” no SDK release required to add an endpoint.
98
+ ## Quick Start (Base - Default)
64
99
 
65
100
  ```typescript
66
- import { BlockrunClient } from '@blockrun/llm';
67
-
68
- const br = new BlockrunClient();
69
-
70
- // Sync GET โ€” Surf market price (Tier 1, $0.001)
71
- const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
101
+ import { LLMClient } from '@blockrun/llm';
72
102
 
73
- // Sync POST โ€” raw on-chain SQL (Tier 3, $0.020)
74
- const rows = await br.post('/v1/surf/onchain/sql', {
75
- query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
76
- });
103
+ const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
77
104
 
78
- // Submit + poll โ€” long-running video gen (settled only on completion)
79
- const video = await br.poll('/v1/videos/generations', {
80
- model: 'xai/grok-imagine-video',
81
- prompt: 'a red apple spinning',
82
- });
105
+ // Recommended: let the router pick the cheapest capable model
106
+ const result = await client.smartChat('Hello!');
83
107
 
84
- // Streaming SSE โ€” chat completions
85
- for await (const chunk of br.stream('/v1/chat/completions', {
86
- model: 'anthropic/claude-sonnet-4-6',
87
- messages: [{ role: 'user', content: 'Hi' }],
88
- stream: true,
89
- })) {
90
- process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
91
- }
108
+ // Or pin a model yourself
109
+ const response = await client.chat('openai/gpt-4o', 'Hello!');
92
110
  ```
93
111
 
94
- Four call shapes cover every endpoint type:
95
- - `get<T>(path, params?)` โ€” synchronous GET (price, ranking, list, news)
96
- - `post<T>(path, body?)` โ€” synchronous POST (on-chain SQL, search)
97
- - `poll<T>(path, body?, { budgetMs, intervalMs })` โ€” submit + poll (image, video, music, voice)
98
- - `stream<T>(path, body?)` โ€” async iterator over SSE chunks (chat)
99
-
100
- The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
101
- `PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `RpcClient`,
102
- `PriceClient`, `SurfClient`) all remain โ€” they will be soft-deprecated in 2.6 (rewritten as
103
- shims over `BlockrunClient`) and removed in 3.0.
112
+ That's it. The SDK handles x402 payment automatically โ€” and `smartChat()`
113
+ keeps the bill down on every request. The router is bundled: no extra
114
+ package to install.
104
115
 
105
116
  ### Try It Free (No USDC Required)
106
117
 
@@ -114,30 +125,33 @@ const client = new LLMClient(); // Wallet still required for signing, but $0 ch
114
125
  // Option 1: call a free model directly
115
126
  const reply = await client.chat('nvidia/deepseek-v4-flash', 'Explain x402 in 1 sentence');
116
127
 
117
- // Option 2: let the smart router pick the best free model per request
118
- const result = await client.smartChat('What is 2+2?', { routingProfile: 'free' });
119
- console.log(result.model); // e.g. 'nvidia/deepseek-v4-flash' (cheapest capable for SIMPLE tier)
128
+ // Option 2: let the smart router pick โ€” 'eco' ranks the free NVIDIA tier first
129
+ const result = await client.smartChat('What is 2+2?', { routingProfile: 'eco' });
130
+ console.log(result.model); // 'nvidia/deepseek-v4-flash' ($0 โ€” verified live)
120
131
  console.log(result.response); // '4'
132
+ console.log(result.routing.savings); // 1 (100%)
121
133
  ```
122
134
 
123
- **Available free models** (input + output both $0, all NVIDIA-hosted, last refreshed 2026-06-07):
135
+ There is no `free` routing profile in `smartChat()` โ€” `routingProfile` accepts
136
+ `'eco' | 'auto' | 'premium'`. (ClawRouter's `/model free` is a feature of its
137
+ own proxy, not of this SDK's router options.) For guaranteed $0, pin a
138
+ `nvidia/*` model; for smart-routed $0-first, use `eco`.
139
+
140
+ **Available free models** (input + output both $0, all NVIDIA-hosted, from the live `/v1/models` catalog, last refreshed 2026-08-10):
124
141
 
125
142
  | Model ID | Context | Best For |
126
143
  |----------|---------|----------|
127
- | `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash โ€” 284B / 13B active MoE, ~5ร— faster than V4 Pro. Best free chat / summarization / light reasoning |
128
- | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Only vision-capable free model โ€” text + images + video (โ‰ค2 min) + audio (โ‰ค1 hr) |
129
- | `nvidia/llama-4-maverick` | 131K | Meta Llama 4 Maverick MoE |
130
- | `nvidia/qwen3-coder-480b` | 131K | Coding-optimised 480B MoE |
131
- | `nvidia/mistral-small-4-119b` | 131K | โš ๏ธ Upstream timing out as of 2026-06-07 โ€” avoid until NVIDIA recovers it |
132
- | `nvidia/gpt-oss-120b` | 128K | OpenAI open-weight 120B โ€” 123 tok/s. Hidden from `/v1/models` for privacy but direct calls still work |
133
- | `nvidia/gpt-oss-20b` | 128K | OpenAI open-weight 20B โ€” 155 tok/s. Hidden from `/v1/models` but direct calls still work |
134
-
135
- > Need V4-Pro-class reasoning? Use the paid `deepseek/deepseek-v4-pro` ($0.435/$0.87 โ€” the 75% launch promo became the permanent list price after 2026-05-31) โ€” `nvidia/deepseek-v4-pro` is currently hidden because NVIDIA's NIM deployment is hung; backend MODEL_REDIRECTS forwards calls to V4 Flash.
144
+ | `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash โ€” 284B / 13B active MoE. Best free chat / summarization / light reasoning. Capacity-constrained: requests may be answered by an equivalent free model |
145
+ | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning โ€” text + images + video + audio (ChartQA 90.3, DocVQA 95.6) |
146
+ | `nvidia/mistral-nemotron` | 131K | Mistral ร— NVIDIA instruction model โ€” fast (~0.2s), strong instruction following |
147
+ | `nvidia/step-3.7-flash` | 131K | StepFun Step 3.7 Flash โ€” fast lightweight reasoning |
148
+ | `nvidia/nemotron-nano-9b-v2` | 131K | Compact + fast (~0.7s), good for high-volume light tasks |
149
+ | `nvidia/nemotron-nano-12b-v2-vl` | 131K | Vision-language โ€” accepts images, compact + fast |
150
+ | `nvidia/gpt-oss-120b` | 128K | OpenAI open-weight 120B. Hidden from `/v1/models` for privacy but direct calls still work |
151
+ | `nvidia/gpt-oss-20b` | 128K | OpenAI open-weight 20B. Hidden from `/v1/models` but direct calls still work |
136
152
 
137
153
  > Privacy note: `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` are hidden from `/v1/models` because NVIDIA's free build.nvidia.com tier reserves the right to use prompts/outputs for service improvement. Direct calls by full model ID still work โ€” opt in only when your data isn't sensitive.
138
154
 
139
- > Retired: `nvidia/qwen3-next-80b-a3b-thinking` hit NVIDIA end-of-life 2026-05-21 (HTTP 410). The gateway auto-redirects pinned callers to `nvidia/llama-4-maverick`.
140
-
141
155
  ## Quick Start (Solana)
142
156
 
143
157
  ```typescript
@@ -151,6 +165,191 @@ console.log(response);
151
165
 
152
166
  Set `SOLANA_WALLET_KEY` to your bs58-encoded Solana secret key. Payments are automatic via x402 โ€” your key never leaves your machine.
153
167
 
168
+ ## Smart Routing (Router Core V3)
169
+
170
+ Let the SDK automatically pick the cheapest capable model for each request โ€” **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** for the same traffic on `auto`, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco`.
171
+
172
+ Not an "up to" figure. The baseline, the workload mix and the token ratio are
173
+ published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json),
174
+ priced against the live catalog, so anyone can recompute the claim and get the
175
+ same answer.
176
+
177
+ Smart routing is powered by the product-neutral
178
+ [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) V3 engine โ€”
179
+ the same deterministic portfolio router that drives
180
+ [ClawRouter](https://github.com/BlockRunAI/ClawRouter). It is **bundled into
181
+ this SDK**: no separate router package to install, and routing runs 100%
182
+ locally with zero external calls.
183
+
184
+ Three ways to use it:
185
+
186
+ ```typescript
187
+ // 1. smartChat() โ€” one-line routed chat
188
+ const result = await client.smartChat('What is 2+2?');
189
+
190
+ // 2. smartChatCompletion() โ€” full agent/tool conversations, routed
191
+ const agent = await client.smartChatCompletion(messages, { tools, toolChoice: 'auto' });
192
+
193
+ // 3. blockrun/auto | blockrun/eco | blockrun/premium โ€” model aliases accepted
194
+ // by chat(), chatCompletion(), and chatCompletionStream() on both chains
195
+ const reply = await client.chatCompletion('blockrun/auto', messages);
196
+
197
+ // Inspect a decision without paying for anything
198
+ const decision = await client.route('Prove the Riemann hypothesis');
199
+ ```
200
+
201
+ The aliases are resolved locally by `LLMClient`, `SolanaLLMClient`, and the
202
+ OpenAI-compat layer. The Anthropic-compat layer proxies straight to the
203
+ gateway's `/v1/messages` and does **not** resolve them โ€” pass a concrete
204
+ model id there.
205
+
206
+ ```typescript
207
+ import { LLMClient } from '@blockrun/llm';
208
+
209
+ const client = new LLMClient();
210
+
211
+ // Auto-routes to cheapest capable model
212
+ const result = await client.smartChat('What is 2+2?');
213
+ console.log(result.response); // '4'
214
+ console.log(result.model); // 'moonshot/kimi-k2.5' (cheap, fast)
215
+ console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 88%'
216
+
217
+ // Complex reasoning task -> routes to reasoning model
218
+ const complex = await client.smartChat('Prove the Riemann hypothesis step by step');
219
+ console.log(complex.model); // 'xai/grok-4-1-fast-reasoning'
220
+
221
+ // Inspect how the request was classified and ranked (Router v3.4 portfolio).
222
+ console.log(complex.routing.method); // 'portfolio'
223
+ console.log(complex.routing.taskType); // 'reasoning'
224
+ console.log(complex.routing.candidates); // ranked, capability-eligible models
225
+
226
+ // Inspect the fallback chain SmartChat will walk on transient errors.
227
+ console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
228
+ ```
229
+
230
+ ### Automatic Fallback on Transient Errors
231
+
232
+ `smartChat()` populates a fallback chain from the portfolio ranking and
233
+ `chat()` / `chatCompletion()` walk it automatically when the primary model
234
+ returns a transient error โ€” timeouts, network failures, 429 rate limits, or
235
+ 5xx responses (502/503/504/522/524). Other 4xx errors and `PaymentError`
236
+ propagate immediately so wallet / auth issues surface fast.
237
+
238
+ ```typescript
239
+ // Manually pass a fallback chain to chat() / chatCompletion()
240
+ const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
241
+ fallbackModels: ['nvidia/llama-4-maverick', 'nvidia/mistral-small-4-119b'],
242
+ });
243
+ // If deepseek-v4-flash times out, the SDK retries against the next model
244
+ // and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
245
+ ```
246
+
247
+ ### Routing Profiles
248
+
249
+ | Profile | Strategy | Savings vs Opus 5 | Best For |
250
+ |---------|----------|-------------------|----------|
251
+ | `eco` | Cheapest capable model โ€” ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
252
+ | `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%** | General use |
253
+ | `premium` | Top-tier models (OpenAI, Anthropic) | 0% | Quality-critical tasks |
254
+
255
+ For guaranteed $0, call a `nvidia/*` model directly with `chat()` โ€” see
256
+ [Try It Free](#try-it-free-no-usdc-required). ClawRouter's `/model free`
257
+ profile belongs to its own proxy; `smartChat()`'s options are the three above.
258
+
259
+ ```typescript
260
+ // Use premium models for complex tasks
261
+ const result = await client.smartChat(
262
+ 'Write production-grade async TypeScript code',
263
+ { routingProfile: 'premium' }
264
+ );
265
+ console.log(result.model); // 'anthropic/claude-opus-4.7'
266
+ ```
267
+
268
+ ### How the Router Works
269
+
270
+ ```mermaid
271
+ flowchart LR
272
+ A["prompt"] --> B["classify locally<br/>15 dimensions, &lt;1ms"]
273
+ B --> C["hard filters<br/>tools ยท vision ยท context ยท<br/>structured output"]
274
+ C --> D["rank portfolio<br/>quality ยท cost ยท speed ยท<br/>reliability"]
275
+ D --> E["cheapest capable model<br/>+ ranked fallback chain"]
276
+ E --> F["x402 USDC payment<br/>for this request only"]
277
+ F --> G["response<br/>+ full routing metadata"]
278
+ ```
279
+
280
+ Since ClawRouter v0.12.242, Auto uses the deterministic **Router v3.4 portfolio
281
+ strategy**: it classifies the task shape locally across
282
+ <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions
283
+ (token count, code presence, reasoning markers, technical/creative terms,
284
+ agentic patterns, โ€ฆ), enforces tool / vision / structured-output / context
285
+ constraints as **hard filters**, then ranks an ordered candidate portfolio.
286
+ The winner becomes `routing.model`; the rest surface as `routing.candidates`
287
+ and feed SmartChat's transient-error fallback chain. Routing stays 100% local
288
+ and deterministic โ€” <1ms, no extra model call, no network hop.
289
+
290
+ Classification still maps to one of four tiers (`routing.tier`). Each
291
+ tier ร— profile has a designated primary (what the rules strategy โ€”
292
+ `routing.method: 'rules'`, the rollback lever โ€” routes to directly, and what
293
+ anchors the portfolio's candidate pool):
294
+
295
+ | Tier | Example Tasks | ECO | AUTO | PREMIUM |
296
+ |------|---------------|-----|------|---------|
297
+ | SIMPLE | "What is 2+2?", definitions | free/gpt-oss-120b โ€  (**FREE**) | gemini-2.5-flash ($0.30/$2.50) | kimi-k2.7 โ€  ($0.95/$4.00) |
298
+ | MEDIUM | Code snippets, explanations | gemini-3.1-flash-lite ($0.25/$1.50) | kimi-k2.7 ($0.95/$4.00) | gpt-5.3-codex ($1.75/$14.00) |
299
+ | COMPLEX | Architecture, long documents | gemini-3.1-flash-lite ($0.25/$1.50) | gemini-3.1-pro ($2/$12) | claude-fable-5 ($10/$50) |
300
+ | REASONING | Proofs, multi-step reasoning | grok-4-1-fast-reasoning โ€  ($0.20/$0.50) | grok-4-1-fast-reasoning โ€  ($0.20/$0.50) | claude-sonnet-4.6 ($3/$15) |
301
+
302
+ โ€  Withheld from `/v1/models` โ€” the router still calls it by direct ID, but you
303
+ will not find it on the public pricing page. The published savings claim is
304
+ priced on visible models only.
305
+
306
+ This table mirrors ClawRouter's tier configs at the version this SDK pins;
307
+ the [ClawRouter README](https://github.com/BlockRunAI/ClawRouter#how-it-works)
308
+ is the live source of truth as models and prices move.
309
+
310
+ ### Routing Metadata Reference
311
+
312
+ Every `smartChat()` result carries the full decision on `result.routing`
313
+ (type `RoutingDecision`) โ€” enough to log, audit, or replay why a model was
314
+ picked:
315
+
316
+ | Field | Description |
317
+ |-------|-------------|
318
+ | `model` | Selected model id (same as `result.model`) |
319
+ | `method` | `'portfolio'` (the Auto default), `'rules'` (rollback strategy), or `'llm'` |
320
+ | `tier` | Task tier: `'SIMPLE'`, `'MEDIUM'`, `'COMPLEX'`, or `'REASONING'` |
321
+ | `taskType` | Portfolio task classification: `'chat'`, `'extraction'`, `'code_edit'`, `'code_agent'`, `'tool_agent'`, `'debug'`, `'reasoning'`, `'reasoning_math'`, `'long_context'`, `'vision'`, โ€ฆ |
322
+ | `candidates` | Ordered, capability-eligible models ranked by the portfolio router; the first entry is `model` |
323
+ | `candidateScores` | Per-candidate score breakdown (`quality` / `cost` / `speed` / `reliability`), ordered with `candidates` |
324
+ | `fallbacks` | The chain `chat()` walks on transient errors (timeout / network / 429 / 5xx) โ€” `candidates` minus the primary, with ClawRouter's proxy-namespace `free/*` ids mapped to their `nvidia/*` gateway ids (SDK-computed) |
325
+ | `savings` | 0โ€“1 fraction saved vs the premium baseline |
326
+ | `costEstimate` / `baselineCost` | Estimated cost of the pick vs that baseline, in USD |
327
+ | `confidence` | Sigmoid-calibrated classifier confidence, 0โ€“1 |
328
+ | `routerVersion` | `'v3-portfolio'` or `'v2-rules'` |
329
+ | `profile` | Routing profile applied: `'auto'`, `'eco'`, `'premium'`, or `'agentic'` |
330
+ | `reasoning` | Human-readable explanation of the decision |
331
+ | `tierConfigs` | The tier โ†’ primary/fallback map the decision was made against |
332
+
333
+ ### TypeScript Types
334
+
335
+ `RoutingDecision`, `RoutingProfile`, `RoutingTier`, `RoutingTaskType`,
336
+ `RoutingTierConfig`, `SmartChatCompletionOptions`, and
337
+ `SmartChatCompletionResponse` are exported from `@blockrun/llm`. They are
338
+ derived from
339
+ [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core), pinned
340
+ to a reviewed immutable commit, and shipped **inlined in this SDK's
341
+ declaration files and runtime bundle** โ€” you install nothing extra to route
342
+ or to typecheck.
343
+
344
+ ### Going Deeper
345
+
346
+ - [ClawRouter](https://github.com/BlockRunAI/ClawRouter) โ€” the router itself: OpenClaw plugin, standalone proxy for Cursor / continue.dev / any OpenAI-compatible client, Telegram integration
347
+ - [Routing profiles in depth](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/routing-profiles.md) โ€” ECO / AUTO / PREMIUM details
348
+ - [How the routing engine works](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/smart-llm-router-14-dimension-classifier.md) โ€” the classifier, dimension by dimension
349
+ - [Router benchmark](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/llm-router-benchmark-46-models-sub-1ms-routing.md) โ€” sub-1ms routing across the catalog
350
+ - [ClawRouter vs OpenRouter](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/clawrouter-vs-openrouter-llm-routing-comparison.md) โ€” head-to-head comparison
351
+ - [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) โ€” the deterministic routing engine both share
352
+
154
353
  ## Solana Support
155
354
 
156
355
  Pay for AI calls with Solana USDC via [sol.blockrun.ai](https://sol.blockrun.ai):
@@ -234,83 +433,52 @@ Every paid request is a real on-chain USDC transfer โ€” look up your wallet addr
234
433
 
235
434
  **Non-custodial by design: your private key never leaves your machine** โ€” it is only used for local signing, and no funds are ever held by BlockRun.
236
435
 
237
- ## Smart Routing (ClawRouter)
436
+ ## `BlockrunClient` โ€” the universal primitive (recommended for new code)
238
437
 
239
- Let the SDK automatically pick the cheapest capable model for each request:
438
+ Starting in `2.5.0`, the SDK ships a single `BlockrunClient` that speaks to
439
+ **every** BlockRun endpoint over x402. New API surfaces are intended to be
440
+ distributed as [Claude Code skills](https://github.com/anthropics/skills)
441
+ that drive this primitive โ€” no SDK release required to add an endpoint.
240
442
 
241
443
  ```typescript
242
- import { LLMClient } from '@blockrun/llm';
243
-
244
- const client = new LLMClient();
245
-
246
- // Auto-routes to cheapest capable model
247
- const result = await client.smartChat('What is 2+2?');
248
- console.log(result.response); // '4'
249
- console.log(result.model); // 'moonshot/kimi-k2.5' (cheap, fast)
250
- console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 87%'
251
-
252
- // Complex reasoning task -> routes to reasoning model
253
- const complex = await client.smartChat('Prove the Riemann hypothesis step by step');
254
- console.log(complex.model); // 'xai/grok-4-1-fast-reasoning'
255
-
256
- // Inspect the fallback chain SmartChat will walk on transient errors.
257
- console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
258
- ```
444
+ import { BlockrunClient } from '@blockrun/llm';
259
445
 
260
- ### Automatic Fallback on Transient Errors
446
+ const br = new BlockrunClient();
261
447
 
262
- `smartChat()` populates a tier-specific fallback chain and `chat()` /
263
- `chatCompletion()` walk it automatically when the primary model returns a
264
- transient error โ€” timeouts, network failures, or 5xx responses (502/503/504/
265
- 522/524). 4xx errors and `PaymentError` propagate immediately so wallet /
266
- auth issues surface fast.
448
+ // Sync GET โ€” Surf market price (Tier 1, $0.001)
449
+ const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
267
450
 
268
- ```typescript
269
- // Manually pass a fallback chain to chat() / chatCompletion()
270
- const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
271
- fallbackModels: ['nvidia/llama-4-maverick', 'nvidia/mistral-small-4-119b'],
451
+ // Sync POST โ€” raw on-chain SQL (Tier 3, $0.020)
452
+ const rows = await br.post('/v1/surf/onchain/sql', {
453
+ query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
272
454
  });
273
- // If deepseek-v4-flash times out, the SDK retries against the next model
274
- // and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
275
- ```
276
-
277
- ### Routing Profiles
278
455
 
279
- | Profile | Description | Best For |
280
- |---------|-------------|----------|
281
- | `free` | NVIDIA free tier โ€” smart-routes across <!-- br:models.free -->6<!-- /br:models.free --> models (DeepSeek V4 Flash, Nemotron Nano Omni, Qwen3, Llama 4, Mistral, plus 2 hidden gpt-oss) | Zero-cost testing, dev, prod |
282
- | `eco` | Cheapest models per tier (DeepSeek, xAI) | Cost-sensitive production |
283
- | `auto` | Best balance of cost/quality (default) | General use |
284
- | `premium` | Top-tier models (OpenAI, Anthropic) | Quality-critical tasks |
456
+ // Submit + poll โ€” long-running video gen (settled only on completion)
457
+ const video = await br.poll('/v1/videos/generations', {
458
+ model: 'xai/grok-imagine-video',
459
+ prompt: 'a red apple spinning',
460
+ });
285
461
 
286
- ```typescript
287
- // Use premium models for complex tasks
288
- const result = await client.smartChat(
289
- 'Write production-grade async TypeScript code',
290
- { routingProfile: 'premium' }
291
- );
292
- console.log(result.model); // 'anthropic/claude-opus-4.7'
462
+ // Streaming SSE โ€” chat completions
463
+ for await (const chunk of br.stream('/v1/chat/completions', {
464
+ model: 'anthropic/claude-sonnet-4-6',
465
+ messages: [{ role: 'user', content: 'Hi' }],
466
+ stream: true,
467
+ })) {
468
+ process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
469
+ }
293
470
  ```
294
471
 
295
- ### How ClawRouter Works
296
-
297
- ClawRouter uses a 14-dimension rule-based classifier to analyze each request:
298
-
299
- - **Token count** - Short vs long prompts
300
- - **Code presence** - Programming keywords
301
- - **Reasoning markers** - "prove", "step by step", etc.
302
- - **Technical terms** - Architecture, optimization, etc.
303
- - **Creative markers** - Story, poem, brainstorm, etc.
304
- - **Agentic patterns** - Multi-step, tool use indicators
305
-
306
- The classifier runs in <1ms, 100% locally, and routes to one of four tiers:
472
+ Four call shapes cover every endpoint type:
473
+ - `get<T>(path, params?)` โ€” synchronous GET (price, ranking, list, news)
474
+ - `post<T>(path, body?)` โ€” synchronous POST (on-chain SQL, search)
475
+ - `poll<T>(path, body?, { budgetMs, intervalMs })` โ€” submit + poll (image, video, music, voice)
476
+ - `stream<T>(path, body?)` โ€” async iterator over SSE chunks (chat)
307
477
 
308
- | Tier | Example Tasks | Auto Profile Model |
309
- |------|---------------|-------------------|
310
- | SIMPLE | "What is 2+2?", definitions | moonshot/kimi-k2.5 |
311
- | MEDIUM | Code snippets, explanations | xai/grok-code-fast-1 |
312
- | COMPLEX | Architecture, long documents | google/gemini-3.1-pro |
313
- | REASONING | Proofs, multi-step reasoning | xai/grok-4-1-fast-reasoning |
478
+ The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
479
+ `PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `RpcClient`,
480
+ `PriceClient`, `SurfClient`) all remain โ€” they will be soft-deprecated in 2.6 (rewritten as
481
+ shims over `BlockrunClient`) and removed in 3.0.
314
482
 
315
483
  ## Available Models
316
484
 
@@ -871,9 +1039,9 @@ const response2 = await client.chat('anthropic/claude-sonnet-4', 'Write a haiku'
871
1039
  });
872
1040
  ```
873
1041
 
874
- ### Smart Routing (ClawRouter)
1042
+ ### Smart Routing (Router Core V3)
875
1043
 
876
- Save up to <!-- br:savings.autoVsBaselinePct -->87<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. ClawRouter uses a <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions -->-dimension rule-based scoring algorithm to select the cheapest model that can handle your request (<1ms, 100% local).
1044
+ Save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. The bundled Router Core V3 engine classifies each request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions, applies hard capability filters, and ranks the cheapest capable models (<1ms, 100% local). Bundled โ€” nothing extra to install.
877
1045
 
878
1046
  ```typescript
879
1047
  import { LLMClient } from '@blockrun/llm';
@@ -885,21 +1053,22 @@ const result = await client.smartChat('What is 2+2?');
885
1053
  console.log(result.response); // '4'
886
1054
  console.log(result.model); // 'google/gemini-2.5-flash'
887
1055
  console.log(result.routing.tier); // 'SIMPLE'
888
- console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 87%'
1056
+ console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 88%'
889
1057
 
890
- // Routing profiles
891
- const free = await client.smartChat('Hello!', { routingProfile: 'free' }); // Zero cost
892
- const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Budget optimized
1058
+ // Routing profiles ('eco' | 'auto' | 'premium')
1059
+ const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Free tier first, then cheapest paid
893
1060
  const auto = await client.smartChat('Code review', { routingProfile: 'auto' }); // Balanced (default)
894
1061
  const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
1062
+
1063
+ // Guaranteed $0: call a free NVIDIA model directly
1064
+ const free = await client.chat('nvidia/deepseek-v4-flash', 'Hello!');
895
1065
  ```
896
1066
 
897
1067
  **Routing Profiles:**
898
1068
 
899
1069
  | Profile | Description | Best For |
900
1070
  |---------|-------------|----------|
901
- | `free` | NVIDIA free tier (<!-- br:models.free -->6<!-- /br:models.free --> models, smart-routed) | Zero-cost testing, dev, prod |
902
- | `eco` | Budget-optimized | Cost-sensitive workloads |
1071
+ | `eco` | Budget-optimized โ€” ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | Cost-sensitive workloads, zero-cost testing |
903
1072
  | `auto` | Intelligent routing (default) | General use |
904
1073
  | `premium` | Best quality models | Critical tasks |
905
1074
 
@@ -1436,13 +1605,13 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
1436
1605
  ## Frequently Asked Questions
1437
1606
 
1438
1607
  ### What is @blockrun/llm?
1439
- @blockrun/llm is a TypeScript SDK that provides pay-per-request access to 40+ large language models from OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more. It uses the x402 protocol for automatic USDC micropayments โ€” no API keys, no subscriptions, no vendor lock-in.
1608
+ @blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol โ€” no API keys, no subscriptions, no vendor lock-in.
1440
1609
 
1441
1610
  ### How does payment work?
1442
1611
  When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
1443
1612
 
1444
- ### What is smart routing / ClawRouter?
1445
- ClawRouter is a built-in smart routing engine that analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. It can save up to <!-- br:savings.autoVsBaselinePct -->87<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
1613
+ ### What is smart routing?
1614
+ Router Core V3 is bundled into the SDK โ€” the same deterministic routing engine that powers ClawRouter, with nothing extra to install. It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. Use `smartChat()`, `smartChatCompletion()`, or the `blockrun/auto` model alias. It can save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
1446
1615
 
1447
1616
  ### Does it support streaming?
1448
1617
  Yes โ€” as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).
@@ -1453,6 +1622,16 @@ Pay only for what you use. Prices start at $0.0002 per request (GPT-5 Nano). The
1453
1622
  ### Does it support both Base and Solana?
1454
1623
  Yes. Use `LLMClient` for Base (EVM) payments and `SolanaLLMClient` for Solana payments. Same API, different payment chain.
1455
1624
 
1625
+ ---
1626
+
1627
+ <div align="center">
1628
+
1629
+ **If the router just cut your bill, [give it a star โญ](https://github.com/BlockRunAI/blockrun-llm-ts)** โ€” it helps more agents pay less.
1630
+
1631
+ [Website](https://blockrun.ai) ยท [Models & Pricing](https://blockrun.ai/models) ยท [ClawRouter](https://github.com/BlockRunAI/ClawRouter) ยท [Python SDK](https://github.com/BlockRunAI/blockrun-llm) ยท [Telegram](https://t.me/blockrunAI)
1632
+
1633
+ </div>
1634
+
1456
1635
  ## License
1457
1636
 
1458
1637
  MIT