@blockrun/llm 3.11.0 โ†’ 3.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,36 +1,72 @@
1
- # @blockrun/llm (TypeScript SDK)
1
+ <div align="center">
2
2
 
3
- > **@blockrun/llm** is a TypeScript/Node.js SDK for accessing <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> large language models (GPT-5, Claude, Gemini, Grok, DeepSeek, Kimi, and more) with automatic pay-per-request USDC micropayments via the x402 protocol. No API keys required โ€” your wallet signature is your authentication. Supports **streaming**, smart routing, Base and Solana chains.
4
- >
5
- > ๐Ÿ†“ **Includes 7 fully-free NVIDIA-hosted models** (5 visible in `/v1/models`, 2 hidden but directly callable) โ€” DeepSeek V4 Flash (1M context), Nemotron Nano Omni (vision), Qwen3 Coder, Llama 4, Mistral, plus the gpt-oss pair. Zero USDC, no rate-limit gimmicks. Use `routingProfile: 'free'` or call any `nvidia/*` model directly.
3
+ # @blockrun/llm
6
4
 
7
- [![npm](https://img.shields.io/npm/v/@blockrun/llm.svg)](https://www.npmjs.com/package/@blockrun/llm)
8
- [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE)
5
+ ### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
9
6
 
10
- ## Supported Chains
7
+ The smart-routing SDK for <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models โ€” every request goes to the cheapest model that can handle it,
8
+ paid per-request in USDC. No API keys. No subscriptions. No vendor lock-in.
11
9
 
12
- | Chain | Network | Payment | Status |
13
- |-------|---------|---------|--------|
14
- | **Base** | Base Mainnet (Chain ID: 8453) | USDC | Primary |
15
- | **Base Testnet** | Base Sepolia (Chain ID: 84532) | Testnet USDC | Development |
16
- | **Solana** | Solana Mainnet | USDC (SPL) | New |
10
+ [![npm](https://img.shields.io/npm/v/@blockrun/llm.svg?style=flat-square)](https://www.npmjs.com/package/@blockrun/llm)
11
+ [![npm downloads](https://img.shields.io/npm/dm/@blockrun/llm.svg?style=flat-square)](https://www.npmjs.com/package/@blockrun/llm)
12
+ [![CI](https://img.shields.io/github/actions/workflow/status/BlockRunAI/blockrun-llm-ts/ci.yml?branch=main&style=flat-square&label=CI)](https://github.com/BlockRunAI/blockrun-llm-ts/actions)
13
+ [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg?style=flat-square)](LICENSE)
14
+ [![Node](https://img.shields.io/badge/Node-%E2%89%A520-brightgreen?style=flat-square&logo=node.js&logoColor=white)](package.json)
15
+ [![TypeScript](https://img.shields.io/badge/TypeScript-strict-3178C6?style=flat-square&logo=typescript&logoColor=white)](tsconfig.json)
17
16
 
17
+ [![Base Network](https://img.shields.io/badge/Base-USDC-0052FF?style=flat-square&logo=coinbase&logoColor=white)](https://base.org)
18
+ [![Solana](https://img.shields.io/badge/Solana-USDC-9945FF?style=flat-square&logo=solana&logoColor=white)](https://solana.com)
19
+ [![x402](https://img.shields.io/badge/x402-micropayments-orange?style=flat-square)](https://x402.org)
20
+ [![Telegram](https://img.shields.io/badge/Telegram-Community-26A5E4?style=flat-square&logo=telegram)](https://t.me/blockrunAI)
18
21
 
19
- **Protocol:** x402 v2 (CDP Facilitator)
22
+ [Website](https://blockrun.ai) ยท [Models & Pricing](https://blockrun.ai/models) ยท [ClawRouter](https://github.com/BlockRunAI/ClawRouter) ยท [Python SDK](https://github.com/BlockRunAI/blockrun-llm) ยท [Telegram](https://t.me/blockrunAI)
23
+
24
+ </div>
25
+
26
+ ---
27
+
28
+ ```typescript
29
+ import { LLMClient } from '@blockrun/llm';
30
+
31
+ const client = new LLMClient();
32
+
33
+ const r = await client.smartChat('Prove step by step that the sum of two odd integers is even.');
34
+ console.log(r.model); // 'deepseek/deepseek-v4-pro' โ€” the right model, not the $75/M flagship
35
+ console.log(r.routing.savings); // 0.96 โ€” this exact request cost 96% less than pinning the baseline
36
+ console.log(r.response); // the proof
37
+ ```
38
+
39
+ **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** across a realistic workload on the default `auto` profile, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco` โ€” and eco's first stop is the free tier, so simple requests cost $0.00 outright. Not an "up to" figure: the baseline, workload mix, and token ratio are published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json) so anyone can recompute the claim. Details in [Smart Routing](#smart-routing-router-core-v3).
40
+
41
+ ## Why This SDK
42
+
43
+ - ๐Ÿง  **Smart routing that pays for itself** โ€” the bundled [Router Core V3](https://github.com/BlockRunAI/router-core) engine (shared with [ClawRouter](https://github.com/BlockRunAI/ClawRouter)) classifies every request locally in <1ms across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and routes to the cheapest capable model. The main event.
44
+ - ๐Ÿ†“ **<!-- br:models.free -->6<!-- /br:models.free --> genuinely free models** โ€” NVIDIA-hosted, $0 in and out, incl. 1M-context DeepSeek V4 Flash and a multimodal Nemotron. No rate-limit gimmicks.
45
+ - ๐Ÿ” **No API keys** โ€” your wallet signature is your authentication. No accounts, no dashboards, no key rotation.
46
+ - ๐Ÿ’ธ **Pay per request in USDC** โ€” x402 micropayments on Base or Solana. $5 covers thousands of requests; agents can pay their own way.
47
+ - ๐Ÿ›ก๏ธ **Automatic failover** โ€” transient errors (timeouts, 429, 5xx) walk the router's ranked fallback chain instead of failing your request.
48
+ - โšก **Streaming, OpenAI & Anthropic compat** โ€” drop-in `chat.completions` / `messages` layers, SSE streaming, strict TypeScript.
49
+ - ๐ŸŽจ **Beyond chat** โ€” image, video, music, speech, live search, prediction markets, crypto data, and 40-chain RPC through the same wallet.
50
+
51
+ ## How It Compares
52
+
53
+ | | OpenAI SDK | OpenRouter | LiteLLM | **@blockrun/llm** |
54
+ | ------------------ | -------------- | ----------------- | ---------------- | ----------------------------------------------------------------------- |
55
+ | **Cost routing** | โœ— one vendor | Manual selection | Manual selection | **Automatic โ€” <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
56
+ | **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->71<!-- /br:models.chatVisible -->, one wallet** |
57
+ | **Free tier** | โœ— | Rate-limited | โœ— | **<!-- br:models.free -->6<!-- /br:models.free --> models, no signup** |
58
+ | **Auth** | API key | Account + API key | Your API keys | **Wallet signature** |
59
+ | **Payment** | Card + invoice | Credit card | BYO keys | **USDC per-request** |
60
+ | **Agent-ready** | โœ— | โœ— | โœ— | **โœ“ โ€” agents fund their own wallet** |
20
61
 
21
62
  ## Installation
22
63
 
23
64
  ```bash
24
- # Base / EVM payments โ€” nothing else needed
25
- npm install @blockrun/llm
26
- # or
27
- pnpm add @blockrun/llm
28
- # or
29
- yarn add @blockrun/llm
65
+ npm install @blockrun/llm # Base / EVM payments โ€” smart routing included, nothing else needed
30
66
  ```
31
67
 
32
- **Solana payments** need two more packages. They are optional peer dependencies,
33
- so npm will not install them for you:
68
+ <details>
69
+ <summary><strong>Solana payments</strong> โ€” two more optional peers</summary>
34
70
 
35
71
  ```bash
36
72
  npm install @blockrun/llm @solana/web3.js @solana/spl-token
@@ -44,63 +80,38 @@ every consumer, including projects that only ever pay on Base. As an optional
44
80
  *peer* it reaches only the projects that ask for Solana. Calling a Solana path
45
81
  without them throws an error naming the exact install command.
46
82
 
47
- ## Quick Start (Base - Default)
83
+ </details>
48
84
 
49
- ```typescript
50
- import { LLMClient } from '@blockrun/llm';
85
+ <details>
86
+ <summary><strong>Supported chains</strong> โ€” Base (primary), Base Sepolia, Solana</summary>
51
87
 
52
- const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
53
- const response = await client.chat('openai/gpt-4o', 'Hello!');
54
- ```
88
+ | Chain | Network | Payment | Status |
89
+ |-------|---------|---------|--------|
90
+ | **Base** | Base Mainnet (Chain ID: 8453) | USDC | Primary |
91
+ | **Base Testnet** | Base Sepolia (Chain ID: 84532) | Testnet USDC | Development |
92
+ | **Solana** | Solana Mainnet | USDC (SPL) | New |
55
93
 
56
- That's it. The SDK handles x402 payment automatically.
94
+ **Protocol:** x402 v2 (CDP Facilitator)
57
95
 
58
- ## `BlockrunClient` โ€” the universal primitive (recommended for new code)
96
+ </details>
59
97
 
60
- Starting in `2.5.0`, the SDK ships a single `BlockrunClient` that speaks to
61
- **every** BlockRun endpoint over x402. New API surfaces are intended to be
62
- distributed as [Claude Code skills](https://github.com/anthropics/skills)
63
- that drive this primitive โ€” no SDK release required to add an endpoint.
98
+ ## Quick Start (Base - Default)
64
99
 
65
100
  ```typescript
66
- import { BlockrunClient } from '@blockrun/llm';
67
-
68
- const br = new BlockrunClient();
69
-
70
- // Sync GET โ€” Surf market price (Tier 1, $0.001)
71
- const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
101
+ import { LLMClient } from '@blockrun/llm';
72
102
 
73
- // Sync POST โ€” raw on-chain SQL (Tier 3, $0.020)
74
- const rows = await br.post('/v1/surf/onchain/sql', {
75
- query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
76
- });
103
+ const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
77
104
 
78
- // Submit + poll โ€” long-running video gen (settled only on completion)
79
- const video = await br.poll('/v1/videos/generations', {
80
- model: 'xai/grok-imagine-video',
81
- prompt: 'a red apple spinning',
82
- });
105
+ // Recommended: let the router pick the cheapest capable model
106
+ const result = await client.smartChat('Hello!');
83
107
 
84
- // Streaming SSE โ€” chat completions
85
- for await (const chunk of br.stream('/v1/chat/completions', {
86
- model: 'anthropic/claude-sonnet-4-6',
87
- messages: [{ role: 'user', content: 'Hi' }],
88
- stream: true,
89
- })) {
90
- process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
91
- }
108
+ // Or pin a model yourself
109
+ const response = await client.chat('openai/gpt-4o', 'Hello!');
92
110
  ```
93
111
 
94
- Four call shapes cover every endpoint type:
95
- - `get<T>(path, params?)` โ€” synchronous GET (price, ranking, list, news)
96
- - `post<T>(path, body?)` โ€” synchronous POST (on-chain SQL, search)
97
- - `poll<T>(path, body?, { budgetMs, intervalMs })` โ€” submit + poll (image, video, music, voice)
98
- - `stream<T>(path, body?)` โ€” async iterator over SSE chunks (chat)
99
-
100
- The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
101
- `PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `RpcClient`,
102
- `PriceClient`, `SurfClient`) all remain โ€” they will be soft-deprecated in 2.6 (rewritten as
103
- shims over `BlockrunClient`) and removed in 3.0.
112
+ That's it. The SDK handles x402 payment automatically โ€” and `smartChat()`
113
+ keeps the bill down on every request. The router is bundled: no extra
114
+ package to install.
104
115
 
105
116
  ### Try It Free (No USDC Required)
106
117
 
@@ -114,30 +125,33 @@ const client = new LLMClient(); // Wallet still required for signing, but $0 ch
114
125
  // Option 1: call a free model directly
115
126
  const reply = await client.chat('nvidia/deepseek-v4-flash', 'Explain x402 in 1 sentence');
116
127
 
117
- // Option 2: let the smart router pick the best free model per request
118
- const result = await client.smartChat('What is 2+2?', { routingProfile: 'free' });
119
- console.log(result.model); // e.g. 'nvidia/deepseek-v4-flash' (cheapest capable for SIMPLE tier)
128
+ // Option 2: let the smart router pick โ€” 'eco' ranks the free NVIDIA tier first
129
+ const result = await client.smartChat('What is 2+2?', { routingProfile: 'eco' });
130
+ console.log(result.model); // 'nvidia/deepseek-v4-flash' ($0 โ€” verified live)
120
131
  console.log(result.response); // '4'
132
+ console.log(result.routing.savings); // 1 (100%)
121
133
  ```
122
134
 
123
- **Available free models** (input + output both $0, all NVIDIA-hosted, last refreshed 2026-06-07):
135
+ There is no `free` routing profile in `smartChat()` โ€” `routingProfile` accepts
136
+ `'eco' | 'auto' | 'premium'`. (ClawRouter's `/model free` is a feature of its
137
+ own proxy, not of this SDK's router options.) For guaranteed $0, pin a
138
+ `nvidia/*` model; for smart-routed $0-first, use `eco`.
139
+
140
+ **Available free models** (input + output both $0, all NVIDIA-hosted, from the live `/v1/models` catalog, last refreshed 2026-08-10):
124
141
 
125
142
  | Model ID | Context | Best For |
126
143
  |----------|---------|----------|
127
- | `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash โ€” 284B / 13B active MoE, ~5ร— faster than V4 Pro. Best free chat / summarization / light reasoning |
128
- | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Only vision-capable free model โ€” text + images + video (โ‰ค2 min) + audio (โ‰ค1 hr) |
129
- | `nvidia/llama-4-maverick` | 131K | Meta Llama 4 Maverick MoE |
130
- | `nvidia/qwen3-coder-480b` | 131K | Coding-optimised 480B MoE |
131
- | `nvidia/mistral-small-4-119b` | 131K | โš ๏ธ Upstream timing out as of 2026-06-07 โ€” avoid until NVIDIA recovers it |
132
- | `nvidia/gpt-oss-120b` | 128K | OpenAI open-weight 120B โ€” 123 tok/s. Hidden from `/v1/models` for privacy but direct calls still work |
133
- | `nvidia/gpt-oss-20b` | 128K | OpenAI open-weight 20B โ€” 155 tok/s. Hidden from `/v1/models` but direct calls still work |
134
-
135
- > Need V4-Pro-class reasoning? Use the paid `deepseek/deepseek-v4-pro` ($0.435/$0.87 โ€” the 75% launch promo became the permanent list price after 2026-05-31) โ€” `nvidia/deepseek-v4-pro` is currently hidden because NVIDIA's NIM deployment is hung; backend MODEL_REDIRECTS forwards calls to V4 Flash.
144
+ | `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash โ€” 284B / 13B active MoE. Best free chat / summarization / light reasoning. Capacity-constrained: requests may be answered by an equivalent free model |
145
+ | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning โ€” text + images + video + audio (ChartQA 90.3, DocVQA 95.6) |
146
+ | `nvidia/mistral-nemotron` | 131K | Mistral ร— NVIDIA instruction model โ€” fast (~0.2s), strong instruction following |
147
+ | `nvidia/step-3.7-flash` | 131K | StepFun Step 3.7 Flash โ€” fast lightweight reasoning |
148
+ | `nvidia/nemotron-nano-9b-v2` | 131K | Compact + fast (~0.7s), good for high-volume light tasks |
149
+ | `nvidia/nemotron-nano-12b-v2-vl` | 131K | Vision-language โ€” accepts images, compact + fast |
150
+ | `nvidia/gpt-oss-120b` | 128K | OpenAI open-weight 120B. Hidden from `/v1/models` for privacy but direct calls still work |
151
+ | `nvidia/gpt-oss-20b` | 128K | OpenAI open-weight 20B. Hidden from `/v1/models` but direct calls still work |
136
152
 
137
153
  > Privacy note: `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` are hidden from `/v1/models` because NVIDIA's free build.nvidia.com tier reserves the right to use prompts/outputs for service improvement. Direct calls by full model ID still work โ€” opt in only when your data isn't sensitive.
138
154
 
139
- > Retired: `nvidia/qwen3-next-80b-a3b-thinking` hit NVIDIA end-of-life 2026-05-21 (HTTP 410). The gateway auto-redirects pinned callers to `nvidia/llama-4-maverick`.
140
-
141
155
  ## Quick Start (Solana)
142
156
 
143
157
  ```typescript
@@ -151,90 +165,7 @@ console.log(response);
151
165
 
152
166
  Set `SOLANA_WALLET_KEY` to your bs58-encoded Solana secret key. Payments are automatic via x402 โ€” your key never leaves your machine.
153
167
 
154
- ## Solana Support
155
-
156
- Pay for AI calls with Solana USDC via [sol.blockrun.ai](https://sol.blockrun.ai):
157
-
158
- ```typescript
159
- import { SolanaLLMClient } from '@blockrun/llm';
160
-
161
- // SOLANA_WALLET_KEY env var (bs58-encoded Solana secret key)
162
- const client = new SolanaLLMClient();
163
-
164
- // Or pass key directly
165
- const client2 = new SolanaLLMClient({ privateKey: 'your-bs58-solana-key' });
166
-
167
- // Same API as LLMClient
168
- const response = await client.chat('openai/gpt-4o', 'gm Solana');
169
- console.log(response);
170
-
171
- // Live Search with Grok (Solana payment)
172
- const tweet = await client.chat('xai/grok-3-mini', 'What is trending on X?', { search: true });
173
- ```
174
-
175
- **Setup:**
176
- 1. Export your Solana wallet key: `export SOLANA_WALLET_KEY="your-bs58-key"`
177
- 2. Fund with USDC on Solana mainnet
178
- 3. That's it โ€” payments are automatic via x402
179
-
180
- **Supported endpoint:** `https://sol.blockrun.ai/api`
181
- **Payment:** Solana USDC (SPL, mainnet)
182
-
183
- ## How Payment Works
184
-
185
- No API keys, no subscription. You hold USDC in your own wallet, and **every request pays for itself** with an on-chain micropayment. Two phases:
186
-
187
- ### Phase 1 โ€” Fund your wallet once
188
-
189
- You only do this when your balance runs low. Three ways to get USDC into your wallet:
190
-
191
- - **(a) Buy with a card (Base USDC).** Call the new `onramp()` method to mint a one-time Coinbase Onramp link, then open the returned `pay.coinbase.com` URL โ€” pay by card/bank in 60+ fiat currencies and the USDC lands in your wallet. The call itself is **free**. Onramp is **Base-only** (buying USDC with a card always lands Base USDC), and the funding address must equal your signing wallet:
192
-
193
- ```typescript
194
- const { url } = await client.onramp(client.getWalletAddress());
195
- console.log(`Fund your wallet: ${url}`); // single-use, expires ~5 min โ€” mint at click time
196
- ```
197
-
198
- - **(b) Transfer existing USDC.** Send USDC you already hold to your wallet address (`client.getWalletAddress()`). On Base, send Base USDC; on Solana (`SolanaLLMClient`), send Solana SPL USDC.
199
-
200
- - **(c) Skip funding entirely.** Use the free NVIDIA models (e.g. `nvidia/deepseek-v4-flash`) โ€” every call is **$0**, no balance required.
201
-
202
- $5 of USDC covers thousands of paid requests. Check your balance any time:
203
-
204
- ```typescript
205
- const balance = await client.getBalance(); // USDC on Base
206
- console.log(`Balance: $${balance.toFixed(2)} USDC`);
207
- ```
208
-
209
- ### Phase 2 โ€” Every request pays itself (automatic x402)
210
-
211
- You just call e.g. `client.chat(...)` โ€” the payment is invisible:
212
-
213
- 1. You send a request to BlockRun's API.
214
- 2. The gateway returns **402 Payment Required** with the price.
215
- 3. The SDK signs a USDC payment **locally** (EIP-712) โ€” on **Base** for `LLMClient`, on **Solana** for `SolanaLLMClient` โ€” using your wallet key.
216
- 4. The request is retried automatically with the payment proof.
217
- 5. The gateway settles on-chain and returns the AI response.
218
-
219
- One call, no separate pay step. Free NVIDIA models settle at **$0** (no payment signed).
220
-
221
- ### Track spend and verify settlements
222
-
223
- ```typescript
224
- import { getCostSummary } from '@blockrun/llm';
225
-
226
- const spent = client.getSpending(); // this session
227
- console.log(`Spent $${spent.totalUsd.toFixed(4)} across ${spent.calls} calls`);
228
-
229
- const summary = getCostSummary(); // across sessions (~/.blockrun/data/costs.jsonl)
230
- console.log(`Lifetime: $${summary.totalUsd.toFixed(2)} over ${summary.calls} calls`);
231
- ```
232
-
233
- Every paid request is a real on-chain USDC transfer โ€” look up your wallet address on [Basescan](https://basescan.org) (or a Solana explorer) to verify each settlement independently.
234
-
235
- **Non-custodial by design: your private key never leaves your machine** โ€” it is only used for local signing, and no funds are ever held by BlockRun.
236
-
237
- ## Smart Routing (ClawRouter)
168
+ ## Smart Routing (Router Core V3)
238
169
 
239
170
  Let the SDK automatically pick the cheapest capable model for each request โ€” **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** for the same traffic on `auto`, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco`.
240
171
 
@@ -243,17 +174,34 @@ published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/ma
243
174
  priced against the live catalog, so anyone can recompute the claim and get the
244
175
  same answer.
245
176
 
246
- Smart routing is powered by [ClawRouter](https://github.com/BlockRunAI/ClawRouter) โ€”
247
- the open-source, agent-first LLM router: wallet signatures instead of API keys,
248
- USDC micropayments instead of credit cards, and <1ms fully-local routing with
249
- zero external calls. In this SDK it is an **optional peer dependency** โ€”
250
- install it alongside the SDK:
177
+ Smart routing is powered by the product-neutral
178
+ [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) V3 engine โ€”
179
+ the same deterministic portfolio router that drives
180
+ [ClawRouter](https://github.com/BlockRunAI/ClawRouter). It is **bundled into
181
+ this SDK**: no separate router package to install, and routing runs 100%
182
+ locally with zero external calls.
251
183
 
252
- ```bash
253
- npm install @blockrun/clawrouter
184
+ Three ways to use it:
185
+
186
+ ```typescript
187
+ // 1. smartChat() โ€” one-line routed chat
188
+ const result = await client.smartChat('What is 2+2?');
189
+
190
+ // 2. smartChatCompletion() โ€” full agent/tool conversations, routed
191
+ const agent = await client.smartChatCompletion(messages, { tools, toolChoice: 'auto' });
192
+
193
+ // 3. blockrun/auto | blockrun/eco | blockrun/premium โ€” model aliases accepted
194
+ // by chat(), chatCompletion(), and chatCompletionStream() on both chains
195
+ const reply = await client.chatCompletion('blockrun/auto', messages);
196
+
197
+ // Inspect a decision without paying for anything
198
+ const decision = await client.route('Prove the Riemann hypothesis');
254
199
  ```
255
200
 
256
- Only `smartChat()` needs it: every other API works without it, and the package is loaded lazily, so a missing or broken router can never break `import '@blockrun/llm'`. Calling `smartChat()` without it throws an error naming the package instead of a cryptic module-load failure.
201
+ The aliases are resolved locally by `LLMClient`, `SolanaLLMClient`, and the
202
+ OpenAI-compat layer. The Anthropic-compat layer proxies straight to the
203
+ gateway's `/v1/messages` and does **not** resolve them โ€” pass a concrete
204
+ model id there.
257
205
 
258
206
  ```typescript
259
207
  import { LLMClient } from '@blockrun/llm';
@@ -300,11 +248,14 @@ const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
300
248
 
301
249
  | Profile | Strategy | Savings vs Opus 5 | Best For |
302
250
  |---------|----------|-------------------|----------|
303
- | `free` | NVIDIA free tier โ€” smart-routes across <!-- br:models.free -->6<!-- /br:models.free --> models (DeepSeek V4 Flash, Nemotron Nano Omni, Qwen3, Llama 4, Mistral, plus 2 hidden gpt-oss) | **100%** | Zero-cost testing, dev, prod |
304
- | `eco` | Cheapest capable model per tier | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production |
251
+ | `eco` | Cheapest capable model โ€” ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
305
252
  | `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%** | General use |
306
253
  | `premium` | Top-tier models (OpenAI, Anthropic) | 0% | Quality-critical tasks |
307
254
 
255
+ For guaranteed $0, call a `nvidia/*` model directly with `chat()` โ€” see
256
+ [Try It Free](#try-it-free-no-usdc-required). ClawRouter's `/model free`
257
+ profile belongs to its own proxy; `smartChat()`'s options are the three above.
258
+
308
259
  ```typescript
309
260
  // Use premium models for complex tasks
310
261
  const result = await client.smartChat(
@@ -314,7 +265,17 @@ const result = await client.smartChat(
314
265
  console.log(result.model); // 'anthropic/claude-opus-4.7'
315
266
  ```
316
267
 
317
- ### How ClawRouter Works
268
+ ### How the Router Works
269
+
270
+ ```mermaid
271
+ flowchart LR
272
+ A["prompt"] --> B["classify locally<br/>15 dimensions, &lt;1ms"]
273
+ B --> C["hard filters<br/>tools ยท vision ยท context ยท<br/>structured output"]
274
+ C --> D["rank portfolio<br/>quality ยท cost ยท speed ยท<br/>reliability"]
275
+ D --> E["cheapest capable model<br/>+ ranked fallback chain"]
276
+ E --> F["x402 USDC payment<br/>for this request only"]
277
+ F --> G["response<br/>+ full routing metadata"]
278
+ ```
318
279
 
319
280
  Since ClawRouter v0.12.242, Auto uses the deterministic **Router v3.4 portfolio
320
281
  strategy**: it classifies the task shape locally across
@@ -371,14 +332,14 @@ picked:
371
332
 
372
333
  ### TypeScript Types
373
334
 
374
- `RoutingDecision`, `RoutingProfile`, `RoutingTier`, `RoutingTaskType`, and
375
- `RoutingTierConfig` are exported from `@blockrun/llm`. They are derived from
376
- [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) โ€” the
377
- routing engine ClawRouter bundles โ€” pinned to the exact commit ClawRouter's
378
- published build inlines, and shipped **inlined in this SDK's declaration
379
- files**. You do not need to install `@blockrun/clawrouter` (or router-core,
380
- which is not on npm) for your project to typecheck against these types; the
381
- runtime package is only needed to actually call `smartChat()`.
335
+ `RoutingDecision`, `RoutingProfile`, `RoutingTier`, `RoutingTaskType`,
336
+ `RoutingTierConfig`, `SmartChatCompletionOptions`, and
337
+ `SmartChatCompletionResponse` are exported from `@blockrun/llm`. They are
338
+ derived from
339
+ [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core), pinned
340
+ to a reviewed immutable commit, and shipped **inlined in this SDK's
341
+ declaration files and runtime bundle** โ€” you install nothing extra to route
342
+ or to typecheck.
382
343
 
383
344
  ### Going Deeper
384
345
 
@@ -389,6 +350,136 @@ runtime package is only needed to actually call `smartChat()`.
389
350
  - [ClawRouter vs OpenRouter](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/clawrouter-vs-openrouter-llm-routing-comparison.md) โ€” head-to-head comparison
390
351
  - [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) โ€” the deterministic routing engine both share
391
352
 
353
+ ## Solana Support
354
+
355
+ Pay for AI calls with Solana USDC via [sol.blockrun.ai](https://sol.blockrun.ai):
356
+
357
+ ```typescript
358
+ import { SolanaLLMClient } from '@blockrun/llm';
359
+
360
+ // SOLANA_WALLET_KEY env var (bs58-encoded Solana secret key)
361
+ const client = new SolanaLLMClient();
362
+
363
+ // Or pass key directly
364
+ const client2 = new SolanaLLMClient({ privateKey: 'your-bs58-solana-key' });
365
+
366
+ // Same API as LLMClient
367
+ const response = await client.chat('openai/gpt-4o', 'gm Solana');
368
+ console.log(response);
369
+
370
+ // Live Search with Grok (Solana payment)
371
+ const tweet = await client.chat('xai/grok-3-mini', 'What is trending on X?', { search: true });
372
+ ```
373
+
374
+ **Setup:**
375
+ 1. Export your Solana wallet key: `export SOLANA_WALLET_KEY="your-bs58-key"`
376
+ 2. Fund with USDC on Solana mainnet
377
+ 3. That's it โ€” payments are automatic via x402
378
+
379
+ **Supported endpoint:** `https://sol.blockrun.ai/api`
380
+ **Payment:** Solana USDC (SPL, mainnet)
381
+
382
+ ## How Payment Works
383
+
384
+ No API keys, no subscription. You hold USDC in your own wallet, and **every request pays for itself** with an on-chain micropayment. Two phases:
385
+
386
+ ### Phase 1 โ€” Fund your wallet once
387
+
388
+ You only do this when your balance runs low. Three ways to get USDC into your wallet:
389
+
390
+ - **(a) Buy with a card (Base USDC).** Call the new `onramp()` method to mint a one-time Coinbase Onramp link, then open the returned `pay.coinbase.com` URL โ€” pay by card/bank in 60+ fiat currencies and the USDC lands in your wallet. The call itself is **free**. Onramp is **Base-only** (buying USDC with a card always lands Base USDC), and the funding address must equal your signing wallet:
391
+
392
+ ```typescript
393
+ const { url } = await client.onramp(client.getWalletAddress());
394
+ console.log(`Fund your wallet: ${url}`); // single-use, expires ~5 min โ€” mint at click time
395
+ ```
396
+
397
+ - **(b) Transfer existing USDC.** Send USDC you already hold to your wallet address (`client.getWalletAddress()`). On Base, send Base USDC; on Solana (`SolanaLLMClient`), send Solana SPL USDC.
398
+
399
+ - **(c) Skip funding entirely.** Use the free NVIDIA models (e.g. `nvidia/deepseek-v4-flash`) โ€” every call is **$0**, no balance required.
400
+
401
+ $5 of USDC covers thousands of paid requests. Check your balance any time:
402
+
403
+ ```typescript
404
+ const balance = await client.getBalance(); // USDC on Base
405
+ console.log(`Balance: $${balance.toFixed(2)} USDC`);
406
+ ```
407
+
408
+ ### Phase 2 โ€” Every request pays itself (automatic x402)
409
+
410
+ You just call e.g. `client.chat(...)` โ€” the payment is invisible:
411
+
412
+ 1. You send a request to BlockRun's API.
413
+ 2. The gateway returns **402 Payment Required** with the price.
414
+ 3. The SDK signs a USDC payment **locally** (EIP-712) โ€” on **Base** for `LLMClient`, on **Solana** for `SolanaLLMClient` โ€” using your wallet key.
415
+ 4. The request is retried automatically with the payment proof.
416
+ 5. The gateway settles on-chain and returns the AI response.
417
+
418
+ One call, no separate pay step. Free NVIDIA models settle at **$0** (no payment signed).
419
+
420
+ ### Track spend and verify settlements
421
+
422
+ ```typescript
423
+ import { getCostSummary } from '@blockrun/llm';
424
+
425
+ const spent = client.getSpending(); // this session
426
+ console.log(`Spent $${spent.totalUsd.toFixed(4)} across ${spent.calls} calls`);
427
+
428
+ const summary = getCostSummary(); // across sessions (~/.blockrun/data/costs.jsonl)
429
+ console.log(`Lifetime: $${summary.totalUsd.toFixed(2)} over ${summary.calls} calls`);
430
+ ```
431
+
432
+ Every paid request is a real on-chain USDC transfer โ€” look up your wallet address on [Basescan](https://basescan.org) (or a Solana explorer) to verify each settlement independently.
433
+
434
+ **Non-custodial by design: your private key never leaves your machine** โ€” it is only used for local signing, and no funds are ever held by BlockRun.
435
+
436
+ ## `BlockrunClient` โ€” the universal primitive (recommended for new code)
437
+
438
+ Starting in `2.5.0`, the SDK ships a single `BlockrunClient` that speaks to
439
+ **every** BlockRun endpoint over x402. New API surfaces are intended to be
440
+ distributed as [Claude Code skills](https://github.com/anthropics/skills)
441
+ that drive this primitive โ€” no SDK release required to add an endpoint.
442
+
443
+ ```typescript
444
+ import { BlockrunClient } from '@blockrun/llm';
445
+
446
+ const br = new BlockrunClient();
447
+
448
+ // Sync GET โ€” Surf market price (Tier 1, $0.001)
449
+ const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
450
+
451
+ // Sync POST โ€” raw on-chain SQL (Tier 3, $0.020)
452
+ const rows = await br.post('/v1/surf/onchain/sql', {
453
+ query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
454
+ });
455
+
456
+ // Submit + poll โ€” long-running video gen (settled only on completion)
457
+ const video = await br.poll('/v1/videos/generations', {
458
+ model: 'xai/grok-imagine-video',
459
+ prompt: 'a red apple spinning',
460
+ });
461
+
462
+ // Streaming SSE โ€” chat completions
463
+ for await (const chunk of br.stream('/v1/chat/completions', {
464
+ model: 'anthropic/claude-sonnet-4-6',
465
+ messages: [{ role: 'user', content: 'Hi' }],
466
+ stream: true,
467
+ })) {
468
+ process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
469
+ }
470
+ ```
471
+
472
+ Four call shapes cover every endpoint type:
473
+ - `get<T>(path, params?)` โ€” synchronous GET (price, ranking, list, news)
474
+ - `post<T>(path, body?)` โ€” synchronous POST (on-chain SQL, search)
475
+ - `poll<T>(path, body?, { budgetMs, intervalMs })` โ€” submit + poll (image, video, music, voice)
476
+ - `stream<T>(path, body?)` โ€” async iterator over SSE chunks (chat)
477
+
478
+ The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
479
+ `PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `RpcClient`,
480
+ `PriceClient`, `SurfClient`) all remain โ€” they will be soft-deprecated in 2.6 (rewritten as
481
+ shims over `BlockrunClient`) and removed in 3.0.
482
+
392
483
  ## Available Models
393
484
 
394
485
  ### OpenAI GPT-5.5 Family
@@ -948,9 +1039,9 @@ const response2 = await client.chat('anthropic/claude-sonnet-4', 'Write a haiku'
948
1039
  });
949
1040
  ```
950
1041
 
951
- ### Smart Routing (ClawRouter)
1042
+ ### Smart Routing (Router Core V3)
952
1043
 
953
- Save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. ClawRouter's deterministic portfolio router (v3.4, default since ClawRouter v0.12.242) classifies each request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions, applies hard capability filters, and ranks the cheapest capable models (<1ms, 100% local). Requires the optional peer dependency: `npm install @blockrun/clawrouter`.
1044
+ Save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. The bundled Router Core V3 engine classifies each request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions, applies hard capability filters, and ranks the cheapest capable models (<1ms, 100% local). Bundled โ€” nothing extra to install.
954
1045
 
955
1046
  ```typescript
956
1047
  import { LLMClient } from '@blockrun/llm';
@@ -964,19 +1055,20 @@ console.log(result.model); // 'google/gemini-2.5-flash'
964
1055
  console.log(result.routing.tier); // 'SIMPLE'
965
1056
  console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 88%'
966
1057
 
967
- // Routing profiles
968
- const free = await client.smartChat('Hello!', { routingProfile: 'free' }); // Zero cost
969
- const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Budget optimized
1058
+ // Routing profiles ('eco' | 'auto' | 'premium')
1059
+ const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Free tier first, then cheapest paid
970
1060
  const auto = await client.smartChat('Code review', { routingProfile: 'auto' }); // Balanced (default)
971
1061
  const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
1062
+
1063
+ // Guaranteed $0: call a free NVIDIA model directly
1064
+ const free = await client.chat('nvidia/deepseek-v4-flash', 'Hello!');
972
1065
  ```
973
1066
 
974
1067
  **Routing Profiles:**
975
1068
 
976
1069
  | Profile | Description | Best For |
977
1070
  |---------|-------------|----------|
978
- | `free` | NVIDIA free tier (<!-- br:models.free -->6<!-- /br:models.free --> models, smart-routed) | Zero-cost testing, dev, prod |
979
- | `eco` | Budget-optimized | Cost-sensitive workloads |
1071
+ | `eco` | Budget-optimized โ€” ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | Cost-sensitive workloads, zero-cost testing |
980
1072
  | `auto` | Intelligent routing (default) | General use |
981
1073
  | `premium` | Best quality models | Critical tasks |
982
1074
 
@@ -1513,13 +1605,13 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
1513
1605
  ## Frequently Asked Questions
1514
1606
 
1515
1607
  ### What is @blockrun/llm?
1516
- @blockrun/llm is a TypeScript SDK that provides pay-per-request access to 40+ large language models from OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more. It uses the x402 protocol for automatic USDC micropayments โ€” no API keys, no subscriptions, no vendor lock-in.
1608
+ @blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol โ€” no API keys, no subscriptions, no vendor lock-in.
1517
1609
 
1518
1610
  ### How does payment work?
1519
1611
  When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
1520
1612
 
1521
- ### What is smart routing / ClawRouter?
1522
- ClawRouter is the SDK's smart routing engine, shipped as the optional `@blockrun/clawrouter` peer dependency (`npm install @blockrun/clawrouter` โ€” only `smartChat()` needs it). It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. It can save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
1613
+ ### What is smart routing?
1614
+ Router Core V3 is bundled into the SDK โ€” the same deterministic routing engine that powers ClawRouter, with nothing extra to install. It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. Use `smartChat()`, `smartChatCompletion()`, or the `blockrun/auto` model alias. It can save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
1523
1615
 
1524
1616
  ### Does it support streaming?
1525
1617
  Yes โ€” as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).
@@ -1530,6 +1622,16 @@ Pay only for what you use. Prices start at $0.0002 per request (GPT-5 Nano). The
1530
1622
  ### Does it support both Base and Solana?
1531
1623
  Yes. Use `LLMClient` for Base (EVM) payments and `SolanaLLMClient` for Solana payments. Same API, different payment chain.
1532
1624
 
1625
+ ---
1626
+
1627
+ <div align="center">
1628
+
1629
+ **If the router just cut your bill, [give it a star โญ](https://github.com/BlockRunAI/blockrun-llm-ts)** โ€” it helps more agents pay less.
1630
+
1631
+ [Website](https://blockrun.ai) ยท [Models & Pricing](https://blockrun.ai/models) ยท [ClawRouter](https://github.com/BlockRunAI/ClawRouter) ยท [Python SDK](https://github.com/BlockRunAI/blockrun-llm) ยท [Telegram](https://t.me/blockrunAI)
1632
+
1633
+ </div>
1634
+
1533
1635
  ## License
1534
1636
 
1535
1637
  MIT