@blockrun/llm 2.12.0 → 3.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,1302 +1,1323 @@
1
- # @blockrun/llm (TypeScript SDK)
2
-
3
- > **@blockrun/llm** is a TypeScript/Node.js SDK for accessing 41+ large language models (GPT-5, Claude, Gemini, Grok, DeepSeek, Kimi, and more) with automatic pay-per-request USDC micropayments via the x402 protocol. No API keys required — your wallet signature is your authentication. Supports **streaming**, smart routing, Base and Solana chains.
4
- >
5
- > 🆓 **Includes 8 fully-free NVIDIA-hosted models** (6 visible in `/v1/models`, 2 hidden but directly callable) — DeepSeek V4 Flash (1M context), Nemotron Nano Omni (vision), Qwen3, Llama 4, Mistral, plus the gpt-oss pair. Zero USDC, no rate-limit gimmicks. Use `routingProfile: 'free'` or call any `nvidia/*` model directly.
6
-
7
- [![npm](https://img.shields.io/npm/v/@blockrun/llm.svg)](https://www.npmjs.com/package/@blockrun/llm)
8
- [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE)
9
-
10
- ## Supported Chains
11
-
12
- | Chain | Network | Payment | Status |
13
- |-------|---------|---------|--------|
14
- | **Base** | Base Mainnet (Chain ID: 8453) | USDC | Primary |
15
- | **Base Testnet** | Base Sepolia (Chain ID: 84532) | Testnet USDC | Development |
16
- | **Solana** | Solana Mainnet | USDC (SPL) | New |
17
-
18
- > **XRPL (RLUSD):** Use [@blockrun/llm-xrpl](https://www.npmjs.com/package/@blockrun/llm-xrpl) for XRPL payments
19
-
20
- **Protocol:** x402 v2 (CDP Facilitator)
21
-
22
- ## Installation
23
-
24
- ```bash
25
- # Base and Solana support (optional Solana deps auto-installed)
26
- npm install @blockrun/llm
27
- # or
28
- pnpm add @blockrun/llm
29
- # or
30
- yarn add @blockrun/llm
31
- ```
32
-
33
- ## Quick Start (Base - Default)
34
-
35
- ```typescript
36
- import { LLMClient } from '@blockrun/llm';
37
-
38
- const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
39
- const response = await client.chat('openai/gpt-4o', 'Hello!');
40
- ```
41
-
42
- That's it. The SDK handles x402 payment automatically.
43
-
44
- ## `BlockrunClient` — the universal primitive (recommended for new code)
45
-
46
- Starting in `2.5.0`, the SDK ships a single `BlockrunClient` that speaks to
47
- **every** BlockRun endpoint over x402. New API surfaces are intended to be
48
- distributed as [Claude Code skills](https://github.com/anthropics/skills)
49
- that drive this primitive — no SDK release required to add an endpoint.
50
-
51
- ```typescript
52
- import { BlockrunClient } from '@blockrun/llm';
53
-
54
- const br = new BlockrunClient();
55
-
56
- // Sync GET — Surf market price (Tier 1, $0.001)
57
- const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
58
-
59
- // Sync POST — raw on-chain SQL (Tier 3, $0.020)
60
- const rows = await br.post('/v1/surf/onchain/sql', {
61
- query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
62
- });
63
-
64
- // Submit + poll — long-running video gen (settled only on completion)
65
- const video = await br.poll('/v1/videos/generations', {
66
- model: 'xai/grok-imagine-video',
67
- prompt: 'a red apple spinning',
68
- });
69
-
70
- // Streaming SSE — chat completions
71
- for await (const chunk of br.stream('/v1/chat/completions', {
72
- model: 'anthropic/claude-sonnet-4-6',
73
- messages: [{ role: 'user', content: 'Hi' }],
74
- stream: true,
75
- })) {
76
- process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
77
- }
78
- ```
79
-
80
- Four call shapes cover every endpoint type:
81
- - `get<T>(path, params?)` — synchronous GET (price, ranking, list, news)
82
- - `post<T>(path, body?)` — synchronous POST (on-chain SQL, search)
83
- - `poll<T>(path, body?, { budgetMs, intervalMs })` — submit + poll (image, video, music, voice)
84
- - `stream<T>(path, body?)` — async iterator over SSE chunks (chat)
85
-
86
- The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
87
- `PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `XClient`,
88
- `PriceClient`, `SurfClient`) all remain — they will be soft-deprecated in 2.6 (rewritten as
89
- shims over `BlockrunClient`) and removed in 3.0.
90
-
91
- ### Try It Free (No USDC Required)
92
-
93
- Want to kick the tires before funding a wallet? Route to BlockRun's free NVIDIA tier:
94
-
95
- ```typescript
96
- import { LLMClient } from '@blockrun/llm';
97
-
98
- const client = new LLMClient(); // Wallet still required for signing, but $0 charged
99
-
100
- // Option 1: call a free model directly
101
- const reply = await client.chat('nvidia/qwen3-next-80b-a3b-thinking', 'Explain x402 in 1 sentence');
102
-
103
- // Option 2: let the smart router pick the best free model per request
104
- const result = await client.smartChat('What is 2+2?', { routingProfile: 'free' });
105
- console.log(result.model); // e.g. 'nvidia/deepseek-v4-flash' (cheapest capable for SIMPLE tier)
106
- console.log(result.response); // '4'
107
- ```
108
-
109
- **Available free models** (input + output both $0, all NVIDIA-hosted, last refreshed 2026-04-28):
110
-
111
- | Model ID | Context | Best For |
112
- |----------|---------|----------|
113
- | `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash — 284B / 13B active MoE, ~5× faster than V4 Pro. Best free chat / summarization / light reasoning |
114
- | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Only vision-capable free model — text + images + video (≤2 min) + audio (≤1 hr) |
115
- | `nvidia/qwen3-next-80b-a3b-thinking` | 131K | 116 tok/s reasoning with thinking mode |
116
- | `nvidia/mistral-small-4-119b` | 131K | 114 tok/s — fastest free chat |
117
- | `nvidia/llama-4-maverick` | 131K | Meta Llama 4 Maverick MoE |
118
- | `nvidia/qwen3-coder-480b` | 131K | Coding-optimised 480B MoE |
119
- | `nvidia/gpt-oss-120b` | 128K | OpenAI open-weight 120B — 123 tok/s. Hidden from `/v1/models` for privacy but direct calls still work |
120
- | `nvidia/gpt-oss-20b` | 128K | OpenAI open-weight 20B — 155 tok/s. Hidden from `/v1/models` but direct calls still work |
121
-
122
- > Need V4-Pro-class reasoning? Use the paid `deepseek/deepseek-v4-pro` ($0.435/$0.87 — the 75% launch promo became the permanent list price after 2026-05-31) — `nvidia/deepseek-v4-pro` is currently hidden because NVIDIA's NIM deployment is hung; backend MODEL_REDIRECTS forwards calls to V4 Flash.
123
-
124
- > Privacy note: `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` are hidden from `/v1/models` because NVIDIA's free build.nvidia.com tier reserves the right to use prompts/outputs for service improvement. Direct calls by full model ID still work — opt in only when your data isn't sensitive.
125
-
126
- ## Quick Start (Solana)
127
-
128
- ```typescript
129
- import { SolanaLLMClient } from '@blockrun/llm';
130
-
131
- // SOLANA_WALLET_KEY env var (bs58-encoded Solana secret key)
132
- const client = new SolanaLLMClient();
133
- const response = await client.chat('openai/gpt-4o', 'gm Solana');
134
- console.log(response);
135
- ```
136
-
137
- Set `SOLANA_WALLET_KEY` to your bs58-encoded Solana secret key. Payments are automatic via x402 — your key never leaves your machine.
138
-
139
- ## Solana Support
140
-
141
- Pay for AI calls with Solana USDC via [sol.blockrun.ai](https://sol.blockrun.ai):
142
-
143
- ```typescript
144
- import { SolanaLLMClient } from '@blockrun/llm';
145
-
146
- // SOLANA_WALLET_KEY env var (bs58-encoded Solana secret key)
147
- const client = new SolanaLLMClient();
148
-
149
- // Or pass key directly
150
- const client2 = new SolanaLLMClient({ privateKey: 'your-bs58-solana-key' });
151
-
152
- // Same API as LLMClient
153
- const response = await client.chat('openai/gpt-4o', 'gm Solana');
154
- console.log(response);
155
-
156
- // Live Search with Grok (Solana payment)
157
- const tweet = await client.chat('xai/grok-3-mini', 'What is trending on X?', { search: true });
158
- ```
159
-
160
- **Setup:**
161
- 1. Export your Solana wallet key: `export SOLANA_WALLET_KEY="your-bs58-key"`
162
- 2. Fund with USDC on Solana mainnet
163
- 3. That's it — payments are automatic via x402
164
-
165
- **Supported endpoint:** `https://sol.blockrun.ai/api`
166
- **Payment:** Solana USDC (SPL, mainnet)
167
-
168
- ## How It Works
169
-
170
- 1. You send a request to BlockRun's API
171
- 2. The API returns a 402 Payment Required with the price
172
- 3. The SDK automatically signs a USDC payment on Base
173
- 4. The request is retried with the payment proof
174
- 5. You receive the AI response
175
-
176
- **Your private key never leaves your machine** - it's only used for local signing.
177
-
178
- ## Smart Routing (ClawRouter)
179
-
180
- Let the SDK automatically pick the cheapest capable model for each request:
181
-
182
- ```typescript
183
- import { LLMClient } from '@blockrun/llm';
184
-
185
- const client = new LLMClient();
186
-
187
- // Auto-routes to cheapest capable model
188
- const result = await client.smartChat('What is 2+2?');
189
- console.log(result.response); // '4'
190
- console.log(result.model); // 'moonshot/kimi-k2.5' (cheap, fast)
191
- console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 78%'
192
-
193
- // Complex reasoning task -> routes to reasoning model
194
- const complex = await client.smartChat('Prove the Riemann hypothesis step by step');
195
- console.log(complex.model); // 'xai/grok-4-1-fast-reasoning'
196
-
197
- // Inspect the fallback chain SmartChat will walk on transient errors.
198
- console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
199
- ```
200
-
201
- ### Automatic Fallback on Transient Errors
202
-
203
- `smartChat()` populates a tier-specific fallback chain and `chat()` /
204
- `chatCompletion()` walk it automatically when the primary model returns a
205
- transient error — timeouts, network failures, or 5xx responses (502/503/504/
206
- 522/524). 4xx errors and `PaymentError` propagate immediately so wallet /
207
- auth issues surface fast.
208
-
209
- ```typescript
210
- // Manually pass a fallback chain to chat() / chatCompletion()
211
- const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
212
- fallbackModels: ['nvidia/llama-4-maverick', 'nvidia/mistral-small-4-119b'],
213
- });
214
- // If deepseek-v4-flash times out, the SDK retries against the next model
215
- // and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
216
- ```
217
-
218
- ### Routing Profiles
219
-
220
- | Profile | Description | Best For |
221
- |---------|-------------|----------|
222
- | `free` | NVIDIA free tier — smart-routes across 8 models (DeepSeek V4 Flash, Nemotron Nano Omni, Qwen3, Llama 4, Mistral, plus 2 hidden gpt-oss) | Zero-cost testing, dev, prod |
223
- | `eco` | Cheapest models per tier (DeepSeek, xAI) | Cost-sensitive production |
224
- | `auto` | Best balance of cost/quality (default) | General use |
225
- | `premium` | Top-tier models (OpenAI, Anthropic) | Quality-critical tasks |
226
-
227
- ```typescript
228
- // Use premium models for complex tasks
229
- const result = await client.smartChat(
230
- 'Write production-grade async TypeScript code',
231
- { routingProfile: 'premium' }
232
- );
233
- console.log(result.model); // 'anthropic/claude-opus-4.7'
234
- ```
235
-
236
- ### How ClawRouter Works
237
-
238
- ClawRouter uses a 14-dimension rule-based classifier to analyze each request:
239
-
240
- - **Token count** - Short vs long prompts
241
- - **Code presence** - Programming keywords
242
- - **Reasoning markers** - "prove", "step by step", etc.
243
- - **Technical terms** - Architecture, optimization, etc.
244
- - **Creative markers** - Story, poem, brainstorm, etc.
245
- - **Agentic patterns** - Multi-step, tool use indicators
246
-
247
- The classifier runs in <1ms, 100% locally, and routes to one of four tiers:
248
-
249
- | Tier | Example Tasks | Auto Profile Model |
250
- |------|---------------|-------------------|
251
- | SIMPLE | "What is 2+2?", definitions | moonshot/kimi-k2.5 |
252
- | MEDIUM | Code snippets, explanations | xai/grok-code-fast-1 |
253
- | COMPLEX | Architecture, long documents | google/gemini-3.1-pro |
254
- | REASONING | Proofs, multi-step reasoning | xai/grok-4-1-fast-reasoning |
255
-
256
- ## Available Models
257
-
258
- ### OpenAI GPT-5.5 Family
259
- Released 2026-04-23 — first fully retrained base since GPT-4.5. 1M context, 128K output, native agent + computer use.
260
-
261
- | Model | Input Price | Output Price |
262
- |-------|-------------|--------------|
263
- | `openai/gpt-5.5` | $5.00/M | $30.00/M |
264
-
265
- ### OpenAI GPT-5.4 Family
266
- | Model | Input Price | Output Price |
267
- |-------|-------------|--------------|
268
- | `openai/gpt-5.4` | $2.50/M | $15.00/M |
269
- | `openai/gpt-5.4-pro` | $30.00/M | $180.00/M |
270
- | `openai/gpt-5.4-nano` | $0.20/M | $1.25/M |
271
-
272
- ### OpenAI GPT-5 Family
273
- | Model | Input Price | Output Price |
274
- |-------|-------------|--------------|
275
- | `openai/gpt-5.3` | $1.75/M | $14.00/M |
276
- | `openai/gpt-5.2` | $1.75/M | $14.00/M |
277
- | `openai/gpt-5-mini` | $0.25/M | $2.00/M |
278
- | `openai/gpt-5.2-pro` | $21.00/M | $168.00/M |
279
- | `openai/gpt-5.2-codex` | $1.75/M | $14.00/M |
280
-
281
- ### OpenAI GPT-4 Family
282
- | Model | Input Price | Output Price |
283
- |-------|-------------|--------------|
284
- | `openai/gpt-4.1` | $2.00/M | $8.00/M |
285
- | `openai/gpt-4.1-mini` | $0.40/M | $1.60/M |
286
- | `openai/gpt-4.1-nano` | $0.10/M | $0.40/M |
287
- | `openai/gpt-4o` | $2.50/M | $10.00/M |
288
- | `openai/gpt-4o-mini` | $0.15/M | $0.60/M |
289
-
290
- ### OpenAI O-Series (Reasoning)
291
- | Model | Input Price | Output Price |
292
- |-------|-------------|--------------|
293
- | `openai/o1` | $15.00/M | $60.00/M |
294
- | `openai/o1-mini` | $1.10/M | $4.40/M |
295
- | `openai/o3` | $2.00/M | $8.00/M |
296
- | `openai/o3-mini` | $1.10/M | $4.40/M |
297
- | `openai/o4-mini` | $1.10/M | $4.40/M |
298
-
299
- ### Anthropic Claude
300
- | Model | Input Price | Output Price | Context | Notes |
301
- |-------|-------------|--------------|---------|-------|
302
- | `anthropic/claude-opus-4.8` | $5.00/M | $25.00/M | **1M** | Flagship — agentic coding + adaptive thinking, 128K output |
303
- | `anthropic/claude-opus-4.7` | $5.00/M | $25.00/M | **1M** | Agentic coding + adaptive thinking, 128K output |
304
- | `anthropic/claude-opus-4.6` | $5.00/M | $25.00/M | 200K | Hidden but still callable — kept as in-family hot-swap fallback |
305
- | `anthropic/claude-opus-4.5` | $5.00/M | $25.00/M | 200K | |
306
- | `anthropic/claude-opus-4` | $15.00/M | $75.00/M | 200K | |
307
- | `anthropic/claude-sonnet-4.6` | $3.00/M | $15.00/M | 200K | Best for reasoning/instructions |
308
- | `anthropic/claude-sonnet-4` | $3.00/M | $15.00/M | 200K | |
309
- | `anthropic/claude-haiku-4.5` | $1.00/M | $5.00/M | 200K | |
310
-
311
- ### Google Gemini
312
- | Model | Input Price | Output Price |
313
- |-------|-------------|--------------|
314
- | `google/gemini-3.1-pro` | $2.00/M | $12.00/M |
315
- | `google/gemini-3.5-flash` | $0.50/M | $3.00/M |
316
- | `google/gemini-3.1-flash-lite` | $0.25/M | $1.50/M |
317
- | `google/gemini-3-flash-preview` | $0.50/M | $3.00/M |
318
- | `google/gemini-2.5-pro` | $1.25/M | $10.00/M |
319
- | `google/gemini-2.5-flash` | $0.30/M | $2.50/M |
320
- | `google/gemini-2.5-flash-lite` | $0.10/M | $0.40/M |
321
-
322
- ### DeepSeek
323
-
324
- V4 family launched 2026-04-24. DeepSeek upstream now serves the legacy
325
- `deepseek-chat` / `deepseek-reasoner` aliases as V4 Flash non-thinking /
326
- thinking modes. V4 Pro is the new flagship paid SKU — 1.6T MoE / 49B active,
327
- 1M context, MMLU-Pro 87.5, GPQA 90.1, SWE-bench 80.6, LiveCodeBench 93.5.
328
-
329
- | Model | Input Price | Output Price | Context | Notes |
330
- |-------|-------------|--------------|---------|-------|
331
- | `deepseek/deepseek-v4-pro` | $0.435/M | $0.87/M | 1M | V4 flagship — strongest open-weight reasoner. The 75% launch promo became the permanent list price after 2026-05-31 |
332
- | `deepseek/deepseek-chat` | $0.20/M | $0.40/M | 1M | V4 Flash non-thinking (paid endpoint with 5MB request bodies; same upstream as `nvidia/deepseek-v4-flash`) |
333
- | `deepseek/deepseek-reasoner` | $0.20/M | $0.40/M | 1M | V4 Flash thinking (same upstream as `deepseek-chat`, thinking enabled by default) |
334
-
335
- ### xAI Grok
336
-
337
- Grok 4.3 and Grok Build are resold through BlockRun's OpenRouter credit pool
338
- (same pattern as `deepseek/deepseek-v4-pro` and `minimax/minimax-m3`). The
339
- older Grok chat SKUs (grok-3/3-mini, grok-4-fast / 4-1-fast families,
340
- grok-code-fast-1, grok-4-0709, grok-2-vision) are now **hidden from
341
- `/v1/models`** — direct calls by full ID still work, but SmartChat won't
342
- auto-pick them.
343
-
344
- | Model | Input Price | Output Price | Context | Notes |
345
- |-------|-------------|--------------|---------|-------|
346
- | `xai/grok-4.3` | $1.50/M | $4.00/M | 1M | Reasoning model, vision-capable, tuned for agentic workflows |
347
- | `xai/grok-build-0.1` | $1.50/M | $3.00/M | 256K | Fast agentic coding model — interactive software-engineering workflows |
348
-
349
- ### Moonshot Kimi
350
- | Model | Input Price | Output Price |
351
- |-------|-------------|--------------|
352
- | `moonshot/kimi-k2.6` | $0.95/M | $4.00/M |
353
- | `moonshot/kimi-k2.5` | $0.60/M | $3.00/M |
354
-
355
- ### MiniMax
356
- | Model | Input Price | Output Price |
357
- |-------|-------------|--------------|
358
- | `minimax/minimax-m3` | $0.30/M | $1.20/M |
359
- | `minimax/minimax-m2.7` | $0.30/M | $1.20/M |
360
-
361
- ### NVIDIA (Free) + Moonshot
362
-
363
- Free tier refreshed 2026-04-28: added `nvidia/deepseek-v4-flash` (1M context)
364
- and Nemotron Nano Omni (vision). `nvidia/gpt-oss-120b` and
365
- `nvidia/gpt-oss-20b` were briefly delisted over privacy concerns then
366
- **re-enabled 2026-04-30** with `available: true` + `hidden: true` — they
367
- no longer appear in `/v1/models` (so SmartChat won't auto-pick them) but
368
- direct calls by full ID still return HTTP 200. `nvidia/deepseek-v4-pro`,
369
- `nvidia/deepseek-v3.2`, and `nvidia/glm-4.7` are hidden because NVIDIA's
370
- NIM deployment is hung — backend MODEL_REDIRECTS forwards calls to V4
371
- Flash / qwen3-coder.
372
-
373
- | Model | Input Price | Output Price | Notes |
374
- |-------|-------------|--------------|-------|
375
- | `nvidia/deepseek-v4-flash` | **FREE** | **FREE** | 284B / 13B active MoE, 1M context — best free chat / summarization / light reasoning |
376
- | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 31B / 3.2B active MoE, 256K — only vision-capable free model |
377
- | `nvidia/qwen3-next-80b-a3b-thinking` | **FREE** | **FREE** | 116 tok/s — reasoning flagship with thinking mode |
378
- | `nvidia/mistral-small-4-119b` | **FREE** | **FREE** | 114 tok/s — fastest free chat |
379
- | `nvidia/llama-4-maverick` | **FREE** | **FREE** | Meta Llama 4 Maverick MoE |
380
- | `nvidia/qwen3-coder-480b` | **FREE** | **FREE** | Coding-optimised 480B MoE |
381
- | `nvidia/gpt-oss-120b` | **FREE** | **FREE** | Hidden from `/v1/models` for privacy but direct calls still work — 123 tok/s |
382
- | `nvidia/gpt-oss-20b` | **FREE** | **FREE** | Hidden from `/v1/models` but direct calls still work — 155 tok/s |
383
- | `moonshot/kimi-k2.5` | $0.60/M | $3.00/M | Direct from Moonshot — replaces `nvidia/kimi-k2.5` |
384
-
385
- ### E2E Verified Models
386
-
387
- All models below have been tested end-to-end via the TypeScript SDK (Feb 2026):
388
-
389
- | Provider | Model | Status |
390
- |----------|-------|--------|
391
- | OpenAI | `openai/gpt-4o-mini` | Passed |
392
- | OpenAI | `openai/gpt-5.2-codex` | Passed |
393
- | Anthropic | `anthropic/claude-opus-4.6` | Passed |
394
- | Anthropic | `anthropic/claude-sonnet-4` | Passed |
395
- | Google | `google/gemini-2.5-flash` | Passed |
396
- | DeepSeek | `deepseek/deepseek-chat` | Passed |
397
- | xAI | `xai/grok-3` | Passed |
398
- | Moonshot | `moonshot/kimi-k2.6` | Passed |
399
-
400
- ### Image Generation
401
- | Model | Price |
402
- |-------|-------|
403
- | `openai/dall-e-3` | $0.04-0.08/image |
404
- | `openai/gpt-image-1` | $0.02-0.04/image |
405
- | `openai/gpt-image-2` | $0.06-0.12/image (reasoning-driven, multilingual text rendering, character consistency) |
406
- | `google/nano-banana` | $0.05/image |
407
- | `google/nano-banana-pro` | $0.10-0.15/image |
408
- | `xai/grok-imagine-image` | $0.02/image |
409
- | `xai/grok-imagine-image-pro` | $0.07/image |
410
- | `zai/cogview-4` | $0.015/image |
411
-
412
- Image editing (`client.edit`) via `/v1/images/image2image`: `openai/gpt-image-1`, `openai/gpt-image-2`, `google/nano-banana`, and `google/nano-banana-pro`. Pass a single base64 `data:image/...` URI to edit one image, or an array of 2–4 URIs to **fuse** them (e.g. a subject + a brand logo). Fusion caps: `openai/*` up to 4 source images, `google/*` up to 3. A `mask` cannot be combined with multiple source images.
413
-
414
- ```ts
415
- // Multi-image fusion with Nano Banana
416
- const fused = await client.edit(
417
- "Place the logo on the t-shirt",
418
- [subjectDataUri, logoDataUri],
419
- { model: "google/nano-banana" }
420
- );
421
- console.log(fused.data[0].url);
422
- ```
423
-
424
- ### Video Generation
425
- | Model | Price |
426
- |-------|-------|
427
- | `xai/grok-imagine-video` | $0.05/sec (8s default → $0.42/clip) |
428
- | `bytedance/seedance-1.5-pro` | $0.03/sec (5s default, up to 10s, 720p) |
429
- | `bytedance/seedance-2.0-fast` | $0.15/sec (~60-80s gen, sweet-spot price/quality) |
430
- | `bytedance/seedance-2.0` | $0.30/sec (720p Pro) |
431
-
432
- ```ts
433
- import { VideoClient } from '@blockrun/llm';
434
-
435
- const client = new VideoClient();
436
- const result = await client.generate('a red apple slowly spinning on a wooden table');
437
- console.log(result.data[0].url); // permanent MP4 URL
438
- console.log(result.data[0].duration_seconds); // 8
439
-
440
- // Image-to-video
441
- const r2 = await client.generate('the subject turns and smiles', {
442
- imageUrl: 'https://example.com/portrait.jpg',
443
- });
444
-
445
- // Token360 / Seedance options (silently ignored by xAI Grok video)
446
- const r3 = await client.generate('aerial drone shot over a snowy mountain', {
447
- model: 'bytedance/seedance-2.0-fast',
448
- aspectRatio: '21:9',
449
- resolution: '1080p',
450
- generateAudio: true, // omit to use the model's default
451
- seed: 42,
452
- watermark: false,
453
- returnLastFrame: true, // useful for clip chaining
454
- });
455
- ```
456
-
457
- ### Text-to-Speech & Sound Effects
458
-
459
- `SpeechClient` wraps BlockRun Voice (ElevenLabs): `POST /v1/audio/speech`
460
- (OpenAI-compatible TTS), `POST /v1/audio/sound-effects`, and the free
461
- `GET /v1/audio/voices`. TTS price scales with character count:
462
- `(chars / 1000) × model rate`, minimum $0.001/request. Synthesis is
463
- synchronous (<1s for Flash).
464
-
465
- | Model | Price | Max Input | Notes |
466
- |-------|-------|-----------|-------|
467
- | `elevenlabs/flash-v2.5` | $0.05/1k chars | 40k chars | ~75ms latency, 32 languages (default) |
468
- | `elevenlabs/turbo-v2.5` | $0.05/1k chars | 40k chars | ~250ms latency, balanced quality |
469
- | `elevenlabs/multilingual-v2` | $0.10/1k chars | 10k chars | Long-form narration, audiobooks |
470
- | `elevenlabs/v3` | $0.10/1k chars | 5k chars | Max expressiveness, 70+ languages |
471
- | `elevenlabs/sound-effects` | $0.05/generation | 1k chars | Sound effects up to 22s |
472
-
473
- ```ts
474
- import { SpeechClient } from '@blockrun/llm';
475
-
476
- const client = new SpeechClient();
477
-
478
- // Text-to-speech (voice aliases: sarah, george, laura, charlie,
479
- // river, roger, callum, harry — or any raw ElevenLabs voice_id)
480
- const result = await client.generate('Welcome to BlockRun.', { voice: 'george' });
481
- console.log(result.data[0].url); // audio URL (mp3 by default)
482
-
483
- // Other formats / speed
484
- const wav = await client.generate('Breaking news from the world of micropayments.', {
485
- model: 'elevenlabs/v3',
486
- responseFormat: 'wav',
487
- speed: 1.1,
488
- });
489
-
490
- // Sound effects (flat $0.05/generation)
491
- const fx = await client.soundEffect('rain on a tin roof, distant thunder');
492
-
493
- // List voices (free, rate-limited)
494
- const voices = await client.listVoices();
495
- ```
496
-
497
- ### Virtual Portraits
498
-
499
- `PortraitClient` wraps `POST /v1/portrait/enroll` (paid, flat **$0.01** promo,
500
- no KYC). Enroll a face image by URL and get back a Token360 asset id (`ta_xxxxxx`).
501
- Pass that id as `realFaceAssetId` on a Seedance 2.0 video generation to keep the
502
- same AI character across clips. Payment settles only after Token360 confirms the
503
- enrollment, so a failed enrollment never charges your wallet. The returned
504
- `image_url` is a gateway-mirrored copy of your source image (see `mirrored` /
505
- `source_image_url`). (Real-person likeness is not supported on BlockRun —
506
- enrolled portraits are AI characters.)
507
-
508
- ```ts
509
- import { PortraitClient, VideoClient } from '@blockrun/llm';
510
-
511
- const portraits = new PortraitClient();
512
- const { asset_id } = await portraits.enroll({
513
- name: 'Spokesperson',
514
- imageUrl: 'https://example.com/face.jpg', // public https JPG/PNG/WEBP, ≤10 MB
515
- });
516
-
517
- // Reuse the same character across Seedance 2.0 clips
518
- const video = new VideoClient();
519
- const clip = await video.generate('she waves and smiles', {
520
- model: 'bytedance/seedance-2.0-fast',
521
- realFaceAssetId: asset_id,
522
- });
523
- console.log(clip.data[0].url);
524
- ```
525
-
526
- ### Voice Calls
527
-
528
- `VoiceClient` wraps `POST /v1/voice/call` (paid, $0.54/call) and
529
- `GET /v1/voice/call/{callId}` (free polling) — AI-powered outbound phone
530
- calls powered by Bland.ai. The agent dials the recipient and runs a real-time
531
- conversation based on your `task` instructions. US + Canada destinations.
532
-
533
- ```ts
534
- import { VoiceClient } from '@blockrun/llm';
535
-
536
- const client = new VoiceClient();
537
-
538
- // Initiate (paid $0.54)
539
- const result = await client.call({
540
- to: '+14155552671',
541
- task: 'You are a friendly assistant calling to confirm a 3pm dentist appointment.',
542
- voice: 'maya', // 'nat' | 'josh' | 'maya' | 'june' | 'paige' | 'derek' | 'florian'
543
- max_duration: 5, // minutes (1–30)
544
- });
545
- console.log(result.call_id);
546
-
547
- // Poll for transcript + recording (free)
548
- const status = await client.getStatus(result.call_id);
549
- console.log(status.status, status.recording_url);
550
- ```
551
-
552
- Bring your own caller-ID: pass `from: '+14155552671'` (must be a BlockRun
553
- phone number you own; buy via `/v1/phone/numbers/buy`).
554
-
555
- ### Standalone Search
556
-
557
- `SearchClient` wraps `POST /v1/search` — standalone Grok Live Search.
558
- Pricing: `$0.025/source + margin` (10 sources ≈ `$0.26`).
559
-
560
- ```ts
561
- import { SearchClient } from '@blockrun/llm';
562
-
563
- const client = new SearchClient();
564
- const result = await client.search('Latest news on x402 adoption', {
565
- sources: ['x', 'web'],
566
- maxResults: 10,
567
- });
568
- console.log(result.summary);
569
- for (const url of result.citations ?? []) console.log(url);
570
- ```
571
-
572
- ### X/Twitter (AttentionVC)
573
-
574
- `XClient` covers the full `/v1/x/*` endpoint family — previously the `X*`
575
- types were exported but there was no client to call them with.
576
-
577
- ```ts
578
- import { XClient } from '@blockrun/llm';
579
-
580
- const x = new XClient();
581
- const info = await x.userInfo('elonmusk');
582
- const followers = await x.followers('paulg');
583
- const results = await x.search('x402 micropayments', { queryType: 'Latest' });
584
- const tweets = await x.userTweets({ username: 'vitalikbuterin', includeReplies: false });
585
- ```
586
-
587
- ### Surf Crypto Data
588
-
589
- `SurfClient` exposes the full `/v1/surf/*` catalog — 84+ pay-per-call
590
- endpoints across CEX/DEX market data, on-chain SQL, wallet intelligence,
591
- prediction markets (Polymarket + Kalshi), social analytics, news, VC fund
592
- data, and an OpenAI-compatible chat surface. Flat pricing per call:
593
-
594
- | Tier | Price/call | Examples |
595
- |------|-----------|----------|
596
- | 1 | $0.001 | `/market/price`, `/market/ranking`, `/news/feed`, prediction-market reads, social tweets |
597
- | 2 | $0.005 | `/exchange/depth`, `/exchange/klines`, `/wallet/detail`, `/search/*`, `/social/ranking` |
598
- | 3 | $0.020 | `/onchain/sql`, `/onchain/query`, `/onchain/schema`, `/chat/completions` |
599
-
600
- Because the catalog is broad and evolving, the client deliberately ships a
601
- generic `get` / `post` pair instead of 84 typed wrappers. Pass the path
602
- (with or without the `/v1/surf` prefix), query params, or a JSON body —
603
- type the response via a generic if you want.
604
-
605
- ```ts
606
- import { SurfClient } from '@blockrun/llm';
607
-
608
- const surf = new SurfClient();
609
-
610
- // Tier 1 — token price ($0.001)
611
- const btc = await surf.get('/market/price', { symbol: 'BTC' });
612
-
613
- // Tier 2 — order book depth ($0.005)
614
- const book = await surf.get('/exchange/depth', {
615
- exchange: 'binance',
616
- symbol: 'BTC-USDT',
617
- });
618
-
619
- // Tier 3 — raw on-chain SQL against 80+ ClickHouse tables ($0.020)
620
- const rows = await surf.post('/onchain/sql', {
621
- query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 5',
622
- });
623
-
624
- // Typed response via generic
625
- type Price = { symbol: string; price: number; timestamp: string };
626
- const eth = await surf.get<Price>('/market/price', { symbol: 'ETH' });
627
- ```
628
-
629
- Full endpoint inventory: <https://blockrun.ai/marketplace/surf>.
630
-
631
- Methods: `userLookup`, `userInfo`, `followers`, `following`, `followings`,
632
- `verifiedFollowers`, `userTweets`, `mentions`, `tweetLookup`, `tweetReplies`,
633
- `tweetThread`, `search`, `trending`, `articlesRising`.
634
-
635
- ### Market Data (Pyth)
636
-
637
- `PriceClient` wraps the Pyth-backed market-data endpoints. Crypto, FX and
638
- commodity are fully free (price + history + list); 12 global stock markets
639
- and the `usstock` legacy alias charge `$0.001` for price + history (list is
640
- always free). Pass `requireWallet: false` to construct a free-only client.
641
-
642
- ```ts
643
- import { PriceClient } from '@blockrun/llm';
644
-
645
- const p = new PriceClient({ requireWallet: false });
646
- const btc = await p.price('crypto', 'BTC-USD');
647
- const eur = await p.price('fx', 'EUR-USD');
648
-
649
- // Paid — requires a wallet
650
- const p2 = new PriceClient();
651
- const aapl = await p2.price('stocks', 'AAPL', { market: 'us' });
652
- const bars = await p2.history('stocks', 'AAPL', {
653
- market: 'us',
654
- resolution: 'D',
655
- from: 1700000000,
656
- to: 1710000000,
657
- });
658
- const symbols = await p.listSymbols('crypto', { query: 'sol', limit: 20 });
659
- ```
660
-
661
- Supported `StockMarket` values: `us, hk, jp, kr, gb, de, fr, nl, ie, lu, cn, ca`.
662
-
663
- ### Testnet Models (Base Sepolia)
664
- | Model | Price |
665
- |-------|-------|
666
- | `openai/gpt-oss-20b` | $0.001/request |
667
- | `openai/gpt-oss-120b` | $0.002/request |
668
-
669
- *Testnet models use flat pricing (no token counting) for simplicity.*
670
-
671
- ## X/Twitter Data (Powered by AttentionVC)
672
-
673
- Access X/Twitter user profiles, followers, and followings via [AttentionVC](https://attentionvc.ai) partner API. No API keys needed — pay-per-request via x402.
674
-
675
- ```typescript
676
- import { LLMClient } from '@blockrun/llm';
677
-
678
- const client = new LLMClient();
679
-
680
- // Look up user profiles ($0.002/user, min $0.02)
681
- const users = await client.xUserLookup(['elonmusk', 'blockaborr']);
682
- for (const user of users.users) {
683
- console.log(`@${user.userName}: ${user.followers} followers`);
684
- }
685
-
686
- // Get followers ($0.05/page, ~200 accounts)
687
- let result = await client.xFollowers('blockaborr');
688
- for (const f of result.followers) {
689
- console.log(` @${f.screen_name}`);
690
- }
691
-
692
- // Paginate through all followers
693
- while (result.has_next_page) {
694
- result = await client.xFollowers('blockaborr', result.next_cursor);
695
- }
696
-
697
- // Get followings ($0.05/page)
698
- const followings = await client.xFollowings('blockaborr');
699
- ```
700
-
701
- Works on both `LLMClient` (Base) and `SolanaLLMClient`.
702
-
703
- ## Standalone Search
704
-
705
- Search web, X/Twitter, and news without using a chat model:
706
-
707
- ```typescript
708
- import { LLMClient } from '@blockrun/llm';
709
-
710
- const client = new LLMClient();
711
-
712
- const result = await client.search('latest AI agent frameworks 2026');
713
- console.log(result.summary);
714
- for (const cite of result.citations ?? []) {
715
- console.log(` - ${cite}`);
716
- }
717
-
718
- // Filter by source type and date range
719
- const filtered = await client.search('BlockRun x402', {
720
- sources: ['web', 'x'],
721
- fromDate: '2026-01-01',
722
- maxResults: 5,
723
- });
724
- ```
725
-
726
- ## Image Editing (img2img)
727
-
728
- Edit existing images with text prompts:
729
-
730
- ```typescript
731
- import { LLMClient } from '@blockrun/llm';
732
-
733
- const client = new LLMClient();
734
-
735
- const result = await client.imageEdit(
736
- 'Make the sky purple and add northern lights',
737
- 'data:image/png;base64,...', // base64 or URL
738
- { model: 'openai/gpt-image-1' }
739
- );
740
- console.log(result.data[0].url);
741
- ```
742
-
743
- ## Usage Examples
744
-
745
- ### Simple Chat
746
-
747
- ```typescript
748
- import { LLMClient } from '@blockrun/llm';
749
-
750
- const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
751
-
752
- const response = await client.chat('openai/gpt-4o', 'Explain quantum computing');
753
- console.log(response);
754
-
755
- // With system prompt
756
- const response2 = await client.chat('anthropic/claude-sonnet-4', 'Write a haiku', {
757
- system: 'You are a creative poet.',
758
- });
759
- ```
760
-
761
- ### Smart Routing (ClawRouter)
762
-
763
- Save up to 78% on inference costs with intelligent model routing. ClawRouter uses a 14-dimension rule-based scoring algorithm to select the cheapest model that can handle your request (<1ms, 100% local).
764
-
765
- ```typescript
766
- import { LLMClient } from '@blockrun/llm';
767
-
768
- const client = new LLMClient();
769
-
770
- // Auto-route to cheapest capable model
771
- const result = await client.smartChat('What is 2+2?');
772
- console.log(result.response); // '4'
773
- console.log(result.model); // 'google/gemini-2.5-flash'
774
- console.log(result.routing.tier); // 'SIMPLE'
775
- console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 78%'
776
-
777
- // Routing profiles
778
- const free = await client.smartChat('Hello!', { routingProfile: 'free' }); // Zero cost
779
- const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Budget optimized
780
- const auto = await client.smartChat('Code review', { routingProfile: 'auto' }); // Balanced (default)
781
- const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
782
- ```
783
-
784
- **Routing Profiles:**
785
-
786
- | Profile | Description | Best For |
787
- |---------|-------------|----------|
788
- | `free` | NVIDIA free tier (9 models, smart-routed) | Zero-cost testing, dev, prod |
789
- | `eco` | Budget-optimized | Cost-sensitive workloads |
790
- | `auto` | Intelligent routing (default) | General use |
791
- | `premium` | Best quality models | Critical tasks |
792
-
793
- **Tiers:**
794
-
795
- | Tier | Example Tasks | Typical Models |
796
- |------|---------------|----------------|
797
- | SIMPLE | Greetings, math, lookups | Gemini Flash, GPT-4o-mini |
798
- | MEDIUM | Explanations, summaries | GPT-4o, Claude Sonnet |
799
- | COMPLEX | Analysis, code generation | GPT-5.2, Claude Opus |
800
- | REASONING | Multi-step logic, planning | o3, DeepSeek Reasoner |
801
-
802
- ### Full Chat Completion
803
-
804
- ```typescript
805
- import { LLMClient, type ChatMessage } from '@blockrun/llm';
806
-
807
- const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
808
-
809
- const messages: ChatMessage[] = [
810
- { role: 'system', content: 'You are a helpful assistant.' },
811
- { role: 'user', content: 'How do I read a file in Node.js?' },
812
- ];
813
-
814
- const result = await client.chatCompletion('openai/gpt-4o', messages);
815
- console.log(result.choices[0].message.content);
816
- ```
817
-
818
- ### Streaming
819
-
820
- Stream responses token-by-token with automatic x402 payment. Uses a **pre-auth cache** to skip the 402 round-trip on repeat calls to the same model (~200ms saved per request after the first).
821
-
822
- #### OpenAI-compatible (recommended)
823
-
824
- ```typescript
825
- import { OpenAI } from '@blockrun/llm';
826
-
827
- const client = new OpenAI({ walletKey: process.env.BASE_CHAIN_WALLET_KEY });
828
-
829
- const stream = await client.chat.completions.create({
830
- model: 'openai/gpt-5.4',
831
- messages: [{ role: 'user', content: 'Write a short story about AI agents' }],
832
- stream: true,
833
- });
834
-
835
- for await (const chunk of stream) {
836
- process.stdout.write(chunk.choices[0]?.delta?.content || '');
837
- }
838
- ```
839
-
840
- #### Native client
841
-
842
- ```typescript
843
- import { LLMClient, type ChatMessage } from '@blockrun/llm';
844
-
845
- const client = new LLMClient();
846
-
847
- const messages: ChatMessage[] = [
848
- { role: 'user', content: 'Explain quantum computing in simple terms' },
849
- ];
850
-
851
- // Returns a raw fetch Response with SSE body
852
- const response = await client.chatCompletionStream('google/gemini-2.5-flash', messages);
853
-
854
- const reader = response.body!.getReader();
855
- const decoder = new TextDecoder();
856
-
857
- while (true) {
858
- const { done, value } = await reader.read();
859
- if (done) break;
860
-
861
- const chunk = decoder.decode(value, { stream: true });
862
- for (const line of chunk.split('\n')) {
863
- if (!line.startsWith('data: ') || line === 'data: [DONE]') continue;
864
- const data = JSON.parse(line.slice(6));
865
- process.stdout.write(data.choices?.[0]?.delta?.content || '');
866
- }
867
- }
868
- ```
869
-
870
- #### Payment + streaming flow
871
-
872
- ```
873
- First call (cache miss):
874
- 1. Send request → 402 response (BlockRun returns price)
875
- 2. Sign USDC payment locally (key never leaves machine)
876
- 3. Retry with PAYMENT-SIGNATURE header + stream: true
877
- 4. Cache payment requirements for this model (1h TTL)
878
- 5. Stream tokens as they arrive
879
-
880
- Subsequent calls (cache hit):
881
- 1. Pre-sign payment from cache — skip 402 round-trip
882
- 2. Send request with PAYMENT-SIGNATURE upfront
883
- 3. Stream tokens immediately (~200ms faster)
884
- ```
885
-
886
- ### List Available Models
887
-
888
- ```typescript
889
- import { LLMClient } from '@blockrun/llm';
890
-
891
- const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
892
- const models = await client.listModels();
893
-
894
- for (const model of models) {
895
- console.log(`${model.id}: $${model.inputPrice}/M input`);
896
- }
897
- ```
898
-
899
- ### Multiple Requests
900
-
901
- ```typescript
902
- import { LLMClient } from '@blockrun/llm';
903
-
904
- const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
905
-
906
- const [gpt, claude, gemini] = await Promise.all([
907
- client.chat('openai/gpt-4o', 'What is 2+2?'),
908
- client.chat('anthropic/claude-sonnet-4', 'What is 3+3?'),
909
- client.chat('google/gemini-2.5-flash', 'What is 4+4?'),
910
- ]);
911
- ```
912
-
913
- ## Prediction Markets (Powered by Predexon)
914
-
915
- Access real-time prediction market data from Polymarket, Kalshi, and Binance Futures via [Predexon](https://predexon.com). No API keys needed — pay-per-request via x402.
916
-
917
- ### Polymarket
918
-
919
- ```typescript
920
- import { LLMClient } from '@blockrun/llm';
921
-
922
- const client = new LLMClient();
923
-
924
- // List markets with optional filters ($0.001/request)
925
- const markets = await client.pm("polymarket/markets");
926
- const filtered = await client.pm("polymarket/markets", { status: "active", limit: 10 });
927
- const searched = await client.pm("polymarket/markets", { search: "bitcoin" });
928
-
929
- // List events ($0.001/request)
930
- const events = await client.pm("polymarket/events");
931
-
932
- // Historical trades ($0.001/request)
933
- const trades = await client.pm("polymarket/trades");
934
-
935
- // OHLCV candlestick data for a specific condition ($0.001/request)
936
- const candles = await client.pm("polymarket/candlesticks/0x1234abcd...");
937
-
938
- // Wallet profile ($0.005/request — tier 2)
939
- const profile = await client.pm("polymarket/wallet/0xABC123...");
940
-
941
- // Wallet P&L ($0.005/request — tier 2)
942
- const pnl = await client.pm("polymarket/wallet/pnl/0xABC123...");
943
-
944
- // Global leaderboard ($0.001/request)
945
- const leaderboard = await client.pm("polymarket/leaderboard");
946
- ```
947
-
948
- ### Kalshi & Binance
949
-
950
- ```typescript
951
- // Kalshi markets ($0.001/request)
952
- const kalshiMarkets = await client.pm("kalshi/markets");
953
-
954
- // Kalshi trades ($0.001/request)
955
- const kalshiTrades = await client.pm("kalshi/trades");
956
-
957
- // Binance candles for supported pairs ($0.001/request)
958
- const btcCandles = await client.pm("binance/candles/BTCUSDT");
959
- const ethCandles = await client.pm("binance/candles/ETHUSDT");
960
- // Also: SOLUSDT, XRPUSDT
961
- ```
962
-
963
- ### Cross-Platform
964
-
965
- ```typescript
966
- // Cross-platform matching pairs ($0.001/request)
967
- const pairs = await client.pm("matching-markets/pairs");
968
- ```
969
-
970
- All current endpoints are GET. The `pmQuery()` method is available for future POST endpoints.
971
-
972
- Works on both `LLMClient` (Base) and `SolanaLLMClient`.
973
-
974
- ## Exa Web Search (Powered by Exa)
975
-
976
- Access [Exa](https://exa.ai)'s neural web search via x402. No API keys needed — pay-per-request. Available on **`LLMClient` (Base USDC)** and `SolanaLLMClient` (Solana USDC). Use Base as the primary path; the Solana gateway is awaiting `EXA_API_KEY` provisioning.
977
-
978
- | Method | Description | Price |
979
- |---|---|---|
980
- | `exaSearch(query, options?)` | Neural/keyword web search | $0.01/request |
981
- | `exaFindSimilar(url, options?)` | Find semantically similar pages | $0.01/request |
982
- | `exaContents(urls, options?)` | Extract full text from URLs | $0.002/URL |
983
- | `exaAnswer(query, options?)` | AI answer grounded in web search | $0.01/request |
984
- | `exa(path, body)` | Generic proxy for any Exa endpoint | varies |
985
-
986
- ```typescript
987
- import { LLMClient } from '@blockrun/llm';
988
-
989
- const client = new LLMClient();
990
-
991
- // Neural web search ($0.01/request)
992
- const results = await client.exaSearch("latest AI safety research", { numResults: 5 });
993
- const news = await client.exaSearch("bitcoin ETF news", { category: "news", numResults: 10 });
994
-
995
- // Find similar pages ($0.01/request)
996
- const similar = await client.exaFindSimilar("https://openai.com/research/gpt-4", { numResults: 5 });
997
-
998
- // Extract content from URLs ($0.002/URL)
999
- const content = await client.exaContents(["https://arxiv.org/abs/2303.08774"]);
1000
-
1001
- // AI-generated answer from live web ($0.01/request)
1002
- const answer = await client.exaAnswer("What is the current state of AI safety research?");
1003
-
1004
- // Generic proxy for any Exa endpoint
1005
- const custom = await client.exa("search", { query: "transformer architecture", numResults: 5 });
1006
- ```
1007
-
1008
- Same surface on `SolanaLLMClient` once Solana-side `EXA_API_KEY` is provisioned.
1009
-
1010
- ## Configuration
1011
-
1012
- ```typescript
1013
- // Default: reads BASE_CHAIN_WALLET_KEY from environment
1014
- const client = new LLMClient();
1015
-
1016
- // Or pass options explicitly
1017
- const client = new LLMClient({
1018
- privateKey: '0x...', // Your wallet key (never sent to server)
1019
- apiUrl: 'https://blockrun.ai/api', // Optional
1020
- timeout: 60000, // Optional (ms)
1021
- });
1022
- ```
1023
-
1024
- ## Environment Variables
1025
-
1026
- | Variable | Description |
1027
- |----------|-------------|
1028
- | `BASE_CHAIN_WALLET_KEY` | Your Base chain wallet private key (for Base / `LLMClient`) |
1029
- | `SOLANA_WALLET_KEY` | Your Solana wallet secret key - bs58 encoded (for `SolanaLLMClient`) |
1030
- | `BLOCKRUN_API_URL` | API endpoint (optional, default: https://blockrun.ai/api) |
1031
-
1032
- ## Error Handling
1033
-
1034
- ```typescript
1035
- import { LLMClient, APIError, PaymentError } from '@blockrun/llm';
1036
-
1037
- const client = new LLMClient();
1038
-
1039
- try {
1040
- const response = await client.chat('openai/gpt-4o', 'Hello!');
1041
- } catch (error) {
1042
- if (error instanceof PaymentError) {
1043
- console.error('Payment failed - check USDC balance');
1044
- } else if (error instanceof APIError) {
1045
- console.error(`API error: ${error.message}`);
1046
- }
1047
- }
1048
- ```
1049
-
1050
- ## Testing
1051
-
1052
- ### Running Unit Tests
1053
-
1054
- Unit tests do not require API access or funded wallets:
1055
-
1056
- ```bash
1057
- npm test # Run tests in watch mode
1058
- npm test run # Run tests once
1059
- npm test -- --coverage # Run with coverage report
1060
- ```
1061
-
1062
- ### Running Integration Tests
1063
-
1064
- Integration tests call the production API and require:
1065
- - A funded Base wallet with USDC ($1+ recommended)
1066
- - `BASE_CHAIN_WALLET_KEY` environment variable set
1067
- - Estimated cost: ~$0.05 per test run
1068
-
1069
- ```bash
1070
- export BASE_CHAIN_WALLET_KEY=0x...
1071
- npm test -- test/integration # Run integration tests only
1072
- ```
1073
-
1074
- Integration tests are automatically skipped if `BASE_CHAIN_WALLET_KEY` is not set.
1075
-
1076
- ## Setting Up Your Wallet
1077
-
1078
- ### Base (EVM)
1079
- 1. Create a wallet on Base (Coinbase Wallet, MetaMask, etc.)
1080
- 2. Get USDC on Base for API payments
1081
- 3. Export your private key and set as `BASE_CHAIN_WALLET_KEY`
1082
-
1083
- ```bash
1084
- # .env
1085
- BASE_CHAIN_WALLET_KEY=0x...
1086
- ```
1087
-
1088
- ### Solana
1089
- 1. Create a Solana wallet (Phantom, Backpack, Solflare, etc.)
1090
- 2. Get USDC on Solana for API payments
1091
- 3. Export your secret key and set as `SOLANA_WALLET_KEY`
1092
-
1093
- ```bash
1094
- # .env
1095
- SOLANA_WALLET_KEY=...your_bs58_secret_key
1096
- ```
1097
-
1098
- Note: Solana transactions are gasless for the user - the CDP facilitator pays for transaction fees.
1099
-
1100
- ## Security
1101
-
1102
- ### Private Key Safety
1103
-
1104
- - **Private key stays local**: Your key is only used for signing on your machine
1105
- - **No custody**: BlockRun never holds your funds
1106
- - **Verify transactions**: All payments are on-chain and verifiable
1107
-
1108
- ### Best Practices
1109
-
1110
- **Private Key Management:**
1111
- - Use environment variables, never hard-code keys
1112
- - Use dedicated wallets for API payments (separate from main holdings)
1113
- - Set spending limits by only funding payment wallets with small amounts
1114
- - Never commit `.env` files to version control
1115
- - Rotate keys periodically
1116
-
1117
- **Input Validation:**
1118
- The SDK validates all inputs before API requests:
1119
- - Private keys (format, length, valid hex)
1120
- - API URLs (HTTPS required for production, HTTP allowed for localhost)
1121
- - Model names and parameters (ranges for max\_tokens, temperature, top\_p)
1122
-
1123
- **Error Sanitization:**
1124
- API errors are automatically sanitized to prevent sensitive information leaks.
1125
-
1126
- **Monitoring:**
1127
- ```typescript
1128
- const address = client.getWalletAddress();
1129
- console.log(`View transactions: https://basescan.org/address/${address}`);
1130
- ```
1131
-
1132
- **Keep Updated:**
1133
- ```bash
1134
- npm update @blockrun/llm # Get security patches
1135
- ```
1136
-
1137
- ## TypeScript Support
1138
-
1139
- Full TypeScript support with exported types:
1140
-
1141
- ```typescript
1142
- import {
1143
- LLMClient,
1144
- OpenAI,
1145
- type ChatMessage,
1146
- type ChatResponse,
1147
- type ChatOptions,
1148
- type ChatCompletionOptions,
1149
- type Model,
1150
- // Smart routing types
1151
- type SmartChatOptions,
1152
- type SmartChatResponse,
1153
- type RoutingDecision,
1154
- type RoutingProfile,
1155
- type RoutingTier,
1156
- APIError,
1157
- PaymentError,
1158
- } from '@blockrun/llm';
1159
-
1160
- // chatCompletionStream returns a standard fetch Response with SSE body
1161
- const streamResponse: Response = await client.chatCompletionStream(model, messages, options);
1162
-
1163
- // OpenAI-compat stream returns AsyncIterable
1164
- const stream: AsyncIterable<OpenAIChatCompletionChunk> = await openaiClient.chat.completions.create({
1165
- model, messages, stream: true
1166
- });
1167
- ```
1168
-
1169
- ## Agent Wallet Setup
1170
-
1171
- One-line setup for agent runtimes (Claude Code skills, MCP servers, etc.):
1172
-
1173
- ```typescript
1174
- import { setupAgentWallet } from '@blockrun/llm';
1175
-
1176
- // Auto-creates wallet if none exists, returns ready client
1177
- const client = setupAgentWallet();
1178
- const response = await client.chat('openai/gpt-5.4', 'Hello!');
1179
- ```
1180
-
1181
- For Solana:
1182
-
1183
- ```typescript
1184
- import { setupAgentSolanaWallet } from '@blockrun/llm';
1185
-
1186
- const client = await setupAgentSolanaWallet();
1187
- const response = await client.chat('anthropic/claude-sonnet-4.6', 'Hello!');
1188
- ```
1189
-
1190
- Check wallet status:
1191
-
1192
- ```typescript
1193
- import { status } from '@blockrun/llm';
1194
-
1195
- await status();
1196
- // Wallet: 0xCC8c...5EF8
1197
- // Balance: $5.30 USDC
1198
- ```
1199
-
1200
- ## Wallet Scanning
1201
-
1202
- The SDK auto-detects wallets from any provider on your system:
1203
-
1204
- ```typescript
1205
- import { scanWallets, scanSolanaWallets } from '@blockrun/llm';
1206
-
1207
- // Scans ~/.<dir>/wallet.json for Base wallets
1208
- const baseWallets = scanWallets();
1209
-
1210
- // Scans ~/.<dir>/solana-wallet.json and ~/.brcc/wallet.json
1211
- const solWallets = scanSolanaWallets();
1212
- ```
1213
-
1214
- `getOrCreateWallet()` checks scanned wallets first, so if you already have a wallet from another BlockRun tool, it will be reused automatically.
1215
-
1216
- ## Response Caching
1217
-
1218
- The SDK caches responses to avoid duplicate payments:
1219
-
1220
- ```typescript
1221
- import { getCachedByRequest, saveToCache, clearCache } from '@blockrun/llm';
1222
-
1223
- // Automatic TTLs by endpoint:
1224
- // - X/Twitter: 1 hour
1225
- // - Search: 15 minutes
1226
- // - Models: 24 hours
1227
- // - Chat/Image: no cache (every call is unique)
1228
-
1229
- // Manual cache management
1230
- clearCache(); // Remove all cached responses
1231
- ```
1232
-
1233
- ## Cost Logging
1234
-
1235
- Track spending across sessions:
1236
-
1237
- ```typescript
1238
- import { logCost, getCostSummary } from '@blockrun/llm';
1239
-
1240
- // Costs are logged to ~/.blockrun/data/costs.jsonl
1241
- const summary = getCostSummary();
1242
- console.log(`Total: $${summary.totalUsd.toFixed(2)}`);
1243
- console.log(`Calls: ${summary.calls}`);
1244
- console.log(`By model:`, summary.byModel);
1245
- ```
1246
-
1247
- ## Anthropic SDK Compatibility
1248
-
1249
- Use the official Anthropic SDK interface with BlockRun's pay-per-request backend:
1250
-
1251
- ```typescript
1252
- import { AnthropicClient } from '@blockrun/llm';
1253
-
1254
- const client = new AnthropicClient(); // Auto-detects wallet, auto-pays
1255
-
1256
- const response = await client.messages.create({
1257
- model: 'claude-sonnet-4-6',
1258
- max_tokens: 1024,
1259
- messages: [{ role: 'user', content: 'Hello!' }],
1260
- });
1261
- console.log(response.content[0].text);
1262
-
1263
- // Any model works in Anthropic format
1264
- const gptResponse = await client.messages.create({
1265
- model: 'openai/gpt-5.4',
1266
- max_tokens: 1024,
1267
- messages: [{ role: 'user', content: 'Hello from GPT!' }],
1268
- });
1269
- ```
1270
-
1271
- The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch that handles x402 payment automatically. Your private key never leaves your machine.
1272
-
1273
- ## Links
1274
-
1275
- - [Website](https://blockrun.ai)
1276
- - [Documentation](https://github.com/BlockRunAI/awesome-blockrun/tree/main/docs)
1277
- - [GitHub](https://github.com/blockrunai/blockrun-llm-ts)
1278
- - [Telegram](https://t.me/+mroQv4-4hGgzOGUx)
1279
-
1280
- ## Frequently Asked Questions
1281
-
1282
- ### What is @blockrun/llm?
1283
- @blockrun/llm is a TypeScript SDK that provides pay-per-request access to 40+ large language models from OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more. It uses the x402 protocol for automatic USDC micropayments — no API keys, no subscriptions, no vendor lock-in.
1284
-
1285
- ### How does payment work?
1286
- When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
1287
-
1288
- ### What is smart routing / ClawRouter?
1289
- ClawRouter is a built-in smart routing engine that analyzes your request across 14 dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. It can save up to 78% on LLM costs compared to using premium models for every request.
1290
-
1291
- ### Does it support streaming?
1292
- Yes — as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).
1293
-
1294
- ### How much does it cost?
1295
- Pay only for what you use. Prices start at $0.0002 per request (GPT-5 Nano). There are no minimums, subscriptions, or monthly fees. $5 in USDC gets you thousands of requests.
1296
-
1297
- ### Does it support both Base and Solana?
1298
- Yes. Use `LLMClient` for Base (EVM) payments and `SolanaLLMClient` for Solana payments. Same API, different payment chain.
1299
-
1300
- ## License
1301
-
1302
- MIT
1
+ # @blockrun/llm (TypeScript SDK)
2
+
3
+ > **@blockrun/llm** is a TypeScript/Node.js SDK for accessing 41+ large language models (GPT-5, Claude, Gemini, Grok, DeepSeek, Kimi, and more) with automatic pay-per-request USDC micropayments via the x402 protocol. No API keys required — your wallet signature is your authentication. Supports **streaming**, smart routing, Base and Solana chains.
4
+ >
5
+ > 🆓 **Includes 7 fully-free NVIDIA-hosted models** (5 visible in `/v1/models`, 2 hidden but directly callable) — DeepSeek V4 Flash (1M context), Nemotron Nano Omni (vision), Qwen3 Coder, Llama 4, Mistral, plus the gpt-oss pair. Zero USDC, no rate-limit gimmicks. Use `routingProfile: 'free'` or call any `nvidia/*` model directly.
6
+
7
+ [![npm](https://img.shields.io/npm/v/@blockrun/llm.svg)](https://www.npmjs.com/package/@blockrun/llm)
8
+ [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE)
9
+
10
+ ## Supported Chains
11
+
12
+ | Chain | Network | Payment | Status |
13
+ |-------|---------|---------|--------|
14
+ | **Base** | Base Mainnet (Chain ID: 8453) | USDC | Primary |
15
+ | **Base Testnet** | Base Sepolia (Chain ID: 84532) | Testnet USDC | Development |
16
+ | **Solana** | Solana Mainnet | USDC (SPL) | New |
17
+
18
+
19
+ **Protocol:** x402 v2 (CDP Facilitator)
20
+
21
+ ## Installation
22
+
23
+ ```bash
24
+ # Base and Solana support (optional Solana deps auto-installed)
25
+ npm install @blockrun/llm
26
+ # or
27
+ pnpm add @blockrun/llm
28
+ # or
29
+ yarn add @blockrun/llm
30
+ ```
31
+
32
+ ## Quick Start (Base - Default)
33
+
34
+ ```typescript
35
+ import { LLMClient } from '@blockrun/llm';
36
+
37
+ const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
38
+ const response = await client.chat('openai/gpt-4o', 'Hello!');
39
+ ```
40
+
41
+ That's it. The SDK handles x402 payment automatically.
42
+
43
+ ## `BlockrunClient` — the universal primitive (recommended for new code)
44
+
45
+ Starting in `2.5.0`, the SDK ships a single `BlockrunClient` that speaks to
46
+ **every** BlockRun endpoint over x402. New API surfaces are intended to be
47
+ distributed as [Claude Code skills](https://github.com/anthropics/skills)
48
+ that drive this primitive — no SDK release required to add an endpoint.
49
+
50
+ ```typescript
51
+ import { BlockrunClient } from '@blockrun/llm';
52
+
53
+ const br = new BlockrunClient();
54
+
55
+ // Sync GET — Surf market price (Tier 1, $0.001)
56
+ const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
57
+
58
+ // Sync POST — raw on-chain SQL (Tier 3, $0.020)
59
+ const rows = await br.post('/v1/surf/onchain/sql', {
60
+ query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
61
+ });
62
+
63
+ // Submit + poll — long-running video gen (settled only on completion)
64
+ const video = await br.poll('/v1/videos/generations', {
65
+ model: 'xai/grok-imagine-video',
66
+ prompt: 'a red apple spinning',
67
+ });
68
+
69
+ // Streaming SSE — chat completions
70
+ for await (const chunk of br.stream('/v1/chat/completions', {
71
+ model: 'anthropic/claude-sonnet-4-6',
72
+ messages: [{ role: 'user', content: 'Hi' }],
73
+ stream: true,
74
+ })) {
75
+ process.stdout.write(chunk?.choices?.[0]?.delta?.content ?? '');
76
+ }
77
+ ```
78
+
79
+ Four call shapes cover every endpoint type:
80
+ - `get<T>(path, params?)` — synchronous GET (price, ranking, list, news)
81
+ - `post<T>(path, body?)` — synchronous POST (on-chain SQL, search)
82
+ - `poll<T>(path, body?, { budgetMs, intervalMs })` — submit + poll (image, video, music, voice)
83
+ - `stream<T>(path, body?)` — async iterator over SSE chunks (chat)
84
+
85
+ The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
86
+ `PortraitClient`, `VoiceClient`, `MusicClient`, `SearchClient`, `RpcClient`,
87
+ `PriceClient`, `SurfClient`) all remain — they will be soft-deprecated in 2.6 (rewritten as
88
+ shims over `BlockrunClient`) and removed in 3.0.
89
+
90
+ ### Try It Free (No USDC Required)
91
+
92
+ Want to kick the tires before funding a wallet? Route to BlockRun's free NVIDIA tier:
93
+
94
+ ```typescript
95
+ import { LLMClient } from '@blockrun/llm';
96
+
97
+ const client = new LLMClient(); // Wallet still required for signing, but $0 charged
98
+
99
+ // Option 1: call a free model directly
100
+ const reply = await client.chat('nvidia/deepseek-v4-flash', 'Explain x402 in 1 sentence');
101
+
102
+ // Option 2: let the smart router pick the best free model per request
103
+ const result = await client.smartChat('What is 2+2?', { routingProfile: 'free' });
104
+ console.log(result.model); // e.g. 'nvidia/deepseek-v4-flash' (cheapest capable for SIMPLE tier)
105
+ console.log(result.response); // '4'
106
+ ```
107
+
108
+ **Available free models** (input + output both $0, all NVIDIA-hosted, last refreshed 2026-06-07):
109
+
110
+ | Model ID | Context | Best For |
111
+ |----------|---------|----------|
112
+ | `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash — 284B / 13B active MoE, ~5× faster than V4 Pro. Best free chat / summarization / light reasoning |
113
+ | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Only vision-capable free model — text + images + video (≤2 min) + audio (≤1 hr) |
114
+ | `nvidia/llama-4-maverick` | 131K | Meta Llama 4 Maverick MoE |
115
+ | `nvidia/qwen3-coder-480b` | 131K | Coding-optimised 480B MoE |
116
+ | `nvidia/mistral-small-4-119b` | 131K | ⚠️ Upstream timing out as of 2026-06-07 — avoid until NVIDIA recovers it |
117
+ | `nvidia/gpt-oss-120b` | 128K | OpenAI open-weight 120B — 123 tok/s. Hidden from `/v1/models` for privacy but direct calls still work |
118
+ | `nvidia/gpt-oss-20b` | 128K | OpenAI open-weight 20B — 155 tok/s. Hidden from `/v1/models` but direct calls still work |
119
+
120
+ > Need V4-Pro-class reasoning? Use the paid `deepseek/deepseek-v4-pro` ($0.435/$0.87 — the 75% launch promo became the permanent list price after 2026-05-31) — `nvidia/deepseek-v4-pro` is currently hidden because NVIDIA's NIM deployment is hung; backend MODEL_REDIRECTS forwards calls to V4 Flash.
121
+
122
+ > Privacy note: `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` are hidden from `/v1/models` because NVIDIA's free build.nvidia.com tier reserves the right to use prompts/outputs for service improvement. Direct calls by full model ID still work — opt in only when your data isn't sensitive.
123
+
124
+ > Retired: `nvidia/qwen3-next-80b-a3b-thinking` hit NVIDIA end-of-life 2026-05-21 (HTTP 410). The gateway auto-redirects pinned callers to `nvidia/llama-4-maverick`.
125
+
126
+ ## Quick Start (Solana)
127
+
128
+ ```typescript
129
+ import { SolanaLLMClient } from '@blockrun/llm';
130
+
131
+ // SOLANA_WALLET_KEY env var (bs58-encoded Solana secret key)
132
+ const client = new SolanaLLMClient();
133
+ const response = await client.chat('openai/gpt-4o', 'gm Solana');
134
+ console.log(response);
135
+ ```
136
+
137
+ Set `SOLANA_WALLET_KEY` to your bs58-encoded Solana secret key. Payments are automatic via x402 — your key never leaves your machine.
138
+
139
+ ## Solana Support
140
+
141
+ Pay for AI calls with Solana USDC via [sol.blockrun.ai](https://sol.blockrun.ai):
142
+
143
+ ```typescript
144
+ import { SolanaLLMClient } from '@blockrun/llm';
145
+
146
+ // SOLANA_WALLET_KEY env var (bs58-encoded Solana secret key)
147
+ const client = new SolanaLLMClient();
148
+
149
+ // Or pass key directly
150
+ const client2 = new SolanaLLMClient({ privateKey: 'your-bs58-solana-key' });
151
+
152
+ // Same API as LLMClient
153
+ const response = await client.chat('openai/gpt-4o', 'gm Solana');
154
+ console.log(response);
155
+
156
+ // Live Search with Grok (Solana payment)
157
+ const tweet = await client.chat('xai/grok-3-mini', 'What is trending on X?', { search: true });
158
+ ```
159
+
160
+ **Setup:**
161
+ 1. Export your Solana wallet key: `export SOLANA_WALLET_KEY="your-bs58-key"`
162
+ 2. Fund with USDC on Solana mainnet
163
+ 3. That's it — payments are automatic via x402
164
+
165
+ **Supported endpoint:** `https://sol.blockrun.ai/api`
166
+ **Payment:** Solana USDC (SPL, mainnet)
167
+
168
+ ## How It Works
169
+
170
+ 1. You send a request to BlockRun's API
171
+ 2. The API returns a 402 Payment Required with the price
172
+ 3. The SDK automatically signs a USDC payment on Base
173
+ 4. The request is retried with the payment proof
174
+ 5. You receive the AI response
175
+
176
+ **Your private key never leaves your machine** - it's only used for local signing.
177
+
178
+ ## Smart Routing (ClawRouter)
179
+
180
+ Let the SDK automatically pick the cheapest capable model for each request:
181
+
182
+ ```typescript
183
+ import { LLMClient } from '@blockrun/llm';
184
+
185
+ const client = new LLMClient();
186
+
187
+ // Auto-routes to cheapest capable model
188
+ const result = await client.smartChat('What is 2+2?');
189
+ console.log(result.response); // '4'
190
+ console.log(result.model); // 'moonshot/kimi-k2.5' (cheap, fast)
191
+ console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 78%'
192
+
193
+ // Complex reasoning task -> routes to reasoning model
194
+ const complex = await client.smartChat('Prove the Riemann hypothesis step by step');
195
+ console.log(complex.model); // 'xai/grok-4-1-fast-reasoning'
196
+
197
+ // Inspect the fallback chain SmartChat will walk on transient errors.
198
+ console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
199
+ ```
200
+
201
+ ### Automatic Fallback on Transient Errors
202
+
203
+ `smartChat()` populates a tier-specific fallback chain and `chat()` /
204
+ `chatCompletion()` walk it automatically when the primary model returns a
205
+ transient error — timeouts, network failures, or 5xx responses (502/503/504/
206
+ 522/524). 4xx errors and `PaymentError` propagate immediately so wallet /
207
+ auth issues surface fast.
208
+
209
+ ```typescript
210
+ // Manually pass a fallback chain to chat() / chatCompletion()
211
+ const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
212
+ fallbackModels: ['nvidia/llama-4-maverick', 'nvidia/mistral-small-4-119b'],
213
+ });
214
+ // If deepseek-v4-flash times out, the SDK retries against the next model
215
+ // and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
216
+ ```
217
+
218
+ ### Routing Profiles
219
+
220
+ | Profile | Description | Best For |
221
+ |---------|-------------|----------|
222
+ | `free` | NVIDIA free tier — smart-routes across 8 models (DeepSeek V4 Flash, Nemotron Nano Omni, Qwen3, Llama 4, Mistral, plus 2 hidden gpt-oss) | Zero-cost testing, dev, prod |
223
+ | `eco` | Cheapest models per tier (DeepSeek, xAI) | Cost-sensitive production |
224
+ | `auto` | Best balance of cost/quality (default) | General use |
225
+ | `premium` | Top-tier models (OpenAI, Anthropic) | Quality-critical tasks |
226
+
227
+ ```typescript
228
+ // Use premium models for complex tasks
229
+ const result = await client.smartChat(
230
+ 'Write production-grade async TypeScript code',
231
+ { routingProfile: 'premium' }
232
+ );
233
+ console.log(result.model); // 'anthropic/claude-opus-4.7'
234
+ ```
235
+
236
+ ### How ClawRouter Works
237
+
238
+ ClawRouter uses a 14-dimension rule-based classifier to analyze each request:
239
+
240
+ - **Token count** - Short vs long prompts
241
+ - **Code presence** - Programming keywords
242
+ - **Reasoning markers** - "prove", "step by step", etc.
243
+ - **Technical terms** - Architecture, optimization, etc.
244
+ - **Creative markers** - Story, poem, brainstorm, etc.
245
+ - **Agentic patterns** - Multi-step, tool use indicators
246
+
247
+ The classifier runs in <1ms, 100% locally, and routes to one of four tiers:
248
+
249
+ | Tier | Example Tasks | Auto Profile Model |
250
+ |------|---------------|-------------------|
251
+ | SIMPLE | "What is 2+2?", definitions | moonshot/kimi-k2.5 |
252
+ | MEDIUM | Code snippets, explanations | xai/grok-code-fast-1 |
253
+ | COMPLEX | Architecture, long documents | google/gemini-3.1-pro |
254
+ | REASONING | Proofs, multi-step reasoning | xai/grok-4-1-fast-reasoning |
255
+
256
+ ## Available Models
257
+
258
+ ### OpenAI GPT-5.5 Family
259
+ Released 2026-04-23 — first fully retrained base since GPT-4.5. 1M context, 128K output, native agent + computer use.
260
+
261
+ | Model | Input Price | Output Price |
262
+ |-------|-------------|--------------|
263
+ | `openai/gpt-5.5` | $5.00/M | $30.00/M |
264
+
265
+ ### OpenAI GPT-5.4 Family
266
+ | Model | Input Price | Output Price |
267
+ |-------|-------------|--------------|
268
+ | `openai/gpt-5.4` | $2.50/M | $15.00/M |
269
+ | `openai/gpt-5.4-pro` | $30.00/M | $180.00/M |
270
+ | `openai/gpt-5.4-nano` | $0.20/M | $1.25/M |
271
+
272
+ ### OpenAI GPT-5 Family
273
+ | Model | Input Price | Output Price |
274
+ |-------|-------------|--------------|
275
+ | `openai/gpt-5.3` | $1.75/M | $14.00/M |
276
+ | `openai/gpt-5.2` | $1.75/M | $14.00/M |
277
+ | `openai/gpt-5-mini` | $0.25/M | $2.00/M |
278
+ | `openai/gpt-5.2-pro` | $21.00/M | $168.00/M |
279
+ | `openai/gpt-5.2-codex` | $1.75/M | $14.00/M |
280
+
281
+ ### OpenAI GPT-4 Family
282
+ | Model | Input Price | Output Price |
283
+ |-------|-------------|--------------|
284
+ | `openai/gpt-4.1` | $2.00/M | $8.00/M |
285
+ | `openai/gpt-4.1-mini` | $0.40/M | $1.60/M |
286
+ | `openai/gpt-4.1-nano` | $0.10/M | $0.40/M |
287
+ | `openai/gpt-4o` | $2.50/M | $10.00/M |
288
+ | `openai/gpt-4o-mini` | $0.15/M | $0.60/M |
289
+
290
+ ### OpenAI O-Series (Reasoning)
291
+ | Model | Input Price | Output Price |
292
+ |-------|-------------|--------------|
293
+ | `openai/o1` | $15.00/M | $60.00/M |
294
+ | `openai/o3` | $2.00/M | $8.00/M |
295
+ | `openai/o3-mini` | $1.10/M | $4.40/M |
296
+ | `openai/o4-mini` | $1.10/M | $4.40/M |
297
+
298
+ ### Anthropic Claude
299
+ | Model | Input Price | Output Price | Context | Notes |
300
+ |-------|-------------|--------------|---------|-------|
301
+ | `anthropic/claude-opus-4.8` | $5.00/M | $25.00/M | **1M** | Flagship — agentic coding + adaptive thinking, 128K output |
302
+ | `anthropic/claude-opus-4.7` | $5.00/M | $25.00/M | **1M** | Agentic coding + adaptive thinking, 128K output |
303
+ | `anthropic/claude-opus-4.6` | $5.00/M | $25.00/M | 200K | Hidden but still callable — kept as in-family hot-swap fallback |
304
+ | `anthropic/claude-opus-4.5` | $5.00/M | $25.00/M | 200K | |
305
+ | `anthropic/claude-opus-4` | $15.00/M | $75.00/M | 200K | |
306
+ | `anthropic/claude-sonnet-4.6` | $3.00/M | $15.00/M | 200K | Best for reasoning/instructions |
307
+ | `anthropic/claude-sonnet-4` | $3.00/M | $15.00/M | 200K | |
308
+ | `anthropic/claude-haiku-4.5` | $1.00/M | $5.00/M | 200K | |
309
+
310
+ ### Google Gemini
311
+ | Model | Input Price | Output Price |
312
+ |-------|-------------|--------------|
313
+ | `google/gemini-3.1-pro` | $2.00/M | $12.00/M |
314
+ | `google/gemini-3.5-flash` | $0.50/M | $3.00/M |
315
+ | `google/gemini-3.1-flash-lite` | $0.25/M | $1.50/M |
316
+ | `google/gemini-3-flash-preview` | $0.50/M | $3.00/M |
317
+ | `google/gemini-2.5-pro` | $1.25/M | $10.00/M |
318
+ | `google/gemini-2.5-flash` | $0.30/M | $2.50/M |
319
+ | `google/gemini-2.5-flash-lite` | $0.10/M | $0.40/M |
320
+
321
+ ### DeepSeek
322
+
323
+ V4 family launched 2026-04-24. DeepSeek upstream now serves the legacy
324
+ `deepseek-chat` / `deepseek-reasoner` aliases as V4 Flash non-thinking /
325
+ thinking modes. V4 Pro is the new flagship paid SKU — 1.6T MoE / 49B active,
326
+ 1M context, MMLU-Pro 87.5, GPQA 90.1, SWE-bench 80.6, LiveCodeBench 93.5.
327
+
328
+ | Model | Input Price | Output Price | Context | Notes |
329
+ |-------|-------------|--------------|---------|-------|
330
+ | `deepseek/deepseek-v4-pro` | $0.435/M | $0.87/M | 1M | V4 flagship — strongest open-weight reasoner. The 75% launch promo became the permanent list price after 2026-05-31 |
331
+ | `deepseek/deepseek-chat` | $0.20/M | $0.40/M | 1M | V4 Flash non-thinking (paid endpoint with 5MB request bodies; same upstream as `nvidia/deepseek-v4-flash`) |
332
+ | `deepseek/deepseek-reasoner` | $0.20/M | $0.40/M | 1M | V4 Flash thinking (same upstream as `deepseek-chat`, thinking enabled by default) |
333
+
334
+ ### xAI Grok
335
+
336
+ Grok 4.3 and Grok Build are resold through BlockRun's OpenRouter credit pool
337
+ (same pattern as `deepseek/deepseek-v4-pro` and `minimax/minimax-m3`). The
338
+ older Grok chat SKUs (grok-3/3-mini, grok-4-fast / 4-1-fast families,
339
+ grok-code-fast-1, grok-4-0709, grok-2-vision) are now **hidden from
340
+ `/v1/models`** — direct calls by full ID still work, but SmartChat won't
341
+ auto-pick them.
342
+
343
+ | Model | Input Price | Output Price | Context | Notes |
344
+ |-------|-------------|--------------|---------|-------|
345
+ | `xai/grok-4.3` | $1.50/M | $4.00/M | 1M | Reasoning model, vision-capable, tuned for agentic workflows |
346
+ | `xai/grok-build-0.1` | $1.50/M | $3.00/M | 256K | Fast agentic coding model — interactive software-engineering workflows |
347
+
348
+ ### Moonshot Kimi
349
+ | Model | Input Price | Output Price |
350
+ |-------|-------------|--------------|
351
+ | `moonshot/kimi-k2.6` | $0.95/M | $4.00/M |
352
+ | `moonshot/kimi-k2.5` | $0.60/M | $3.00/M |
353
+
354
+ ### MiniMax
355
+ | Model | Input Price | Output Price |
356
+ |-------|-------------|--------------|
357
+ | `minimax/minimax-m3` | $0.30/M | $1.20/M |
358
+ | `minimax/minimax-m2.7` | $0.30/M | $1.20/M |
359
+
360
+ ### NVIDIA (Free) + Moonshot
361
+
362
+ Free tier refreshed 2026-04-28: added `nvidia/deepseek-v4-flash` (1M context)
363
+ and Nemotron Nano Omni (vision). `nvidia/gpt-oss-120b` and
364
+ `nvidia/gpt-oss-20b` were briefly delisted over privacy concerns then
365
+ **re-enabled 2026-04-30** with `available: true` + `hidden: true` — they
366
+ no longer appear in `/v1/models` (so SmartChat won't auto-pick them) but
367
+ direct calls by full ID still return HTTP 200. `nvidia/deepseek-v4-pro`,
368
+ `nvidia/deepseek-v3.2`, and `nvidia/glm-4.7` are hidden because NVIDIA's
369
+ NIM deployment is hung — backend MODEL_REDIRECTS forwards calls to V4
370
+ Flash / qwen3-coder. `nvidia/qwen3-next-80b-a3b-thinking` hit NVIDIA
371
+ end-of-life 2026-05-21 (HTTP 410) and is auto-redirected to
372
+ `nvidia/llama-4-maverick`.
373
+
374
+ | Model | Input Price | Output Price | Notes |
375
+ |-------|-------------|--------------|-------|
376
+ | `nvidia/deepseek-v4-flash` | **FREE** | **FREE** | 284B / 13B active MoE, 1M context — best free chat / summarization / light reasoning |
377
+ | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 31B / 3.2B active MoE, 256K — only vision-capable free model |
378
+ | `nvidia/mistral-small-4-119b` | **FREE** | **FREE** | ⚠️ Upstream timing out as of 2026-06-07 |
379
+ | `nvidia/llama-4-maverick` | **FREE** | **FREE** | Meta Llama 4 Maverick MoE |
380
+ | `nvidia/qwen3-coder-480b` | **FREE** | **FREE** | Coding-optimised 480B MoE |
381
+ | `nvidia/gpt-oss-120b` | **FREE** | **FREE** | Hidden from `/v1/models` for privacy but direct calls still work — 123 tok/s |
382
+ | `nvidia/gpt-oss-20b` | **FREE** | **FREE** | Hidden from `/v1/models` but direct calls still work — 155 tok/s |
383
+ | `moonshot/kimi-k2.5` | $0.60/M | $3.00/M | Direct from Moonshot — replaces `nvidia/kimi-k2.5` |
384
+
385
+ ### E2E Verified Models
386
+
387
+ All models below have been tested end-to-end via the TypeScript SDK (Feb 2026):
388
+
389
+ | Provider | Model | Status |
390
+ |----------|-------|--------|
391
+ | OpenAI | `openai/gpt-4o-mini` | Passed |
392
+ | OpenAI | `openai/gpt-5.2-codex` | Passed |
393
+ | Anthropic | `anthropic/claude-opus-4.6` | Passed |
394
+ | Anthropic | `anthropic/claude-sonnet-4` | Passed |
395
+ | Google | `google/gemini-2.5-flash` | Passed |
396
+ | DeepSeek | `deepseek/deepseek-chat` | Passed |
397
+ | xAI | `xai/grok-3` | Passed |
398
+ | Moonshot | `moonshot/kimi-k2.6` | Passed |
399
+
400
+ ### Image Generation
401
+ | Model | Price |
402
+ |-------|-------|
403
+ | `openai/dall-e-3` | $0.04-0.08/image |
404
+ | `openai/gpt-image-1` | $0.02-0.04/image |
405
+ | `openai/gpt-image-2` | $0.06-0.12/image (reasoning-driven, multilingual text rendering, character consistency) |
406
+ | `google/nano-banana` | $0.05/image |
407
+ | `google/nano-banana-pro` | $0.10-0.15/image |
408
+ | `xai/grok-imagine-image` | $0.02/image |
409
+ | `xai/grok-imagine-image-pro` | $0.07/image |
410
+ | `zai/cogview-4` | $0.015/image |
411
+
412
+ Image editing (`client.edit`) via `/v1/images/image2image`: `openai/gpt-image-1`, `openai/gpt-image-2`, `google/nano-banana`, and `google/nano-banana-pro`. Pass a single base64 `data:image/...` URI to edit one image, or an array of 2–4 URIs to **fuse** them (e.g. a subject + a brand logo). Fusion caps: `openai/*` up to 4 source images, `google/*` up to 3. A `mask` cannot be combined with multiple source images.
413
+
414
+ ```ts
415
+ // Multi-image fusion with Nano Banana
416
+ const fused = await client.edit(
417
+ "Place the logo on the t-shirt",
418
+ [subjectDataUri, logoDataUri],
419
+ { model: "google/nano-banana" }
420
+ );
421
+ console.log(fused.data[0].url);
422
+ ```
423
+
424
+ ### Video Generation
425
+ | Model | Price |
426
+ |-------|-------|
427
+ | `xai/grok-imagine-video` | $0.05/sec (8s default → $0.42/clip) |
428
+ | `bytedance/seedance-1.5-pro` | $0.03/sec (5s default, up to 10s, 720p) |
429
+ | `bytedance/seedance-2.0-fast` | $0.15/sec (~60-80s gen, sweet-spot price/quality) |
430
+ | `bytedance/seedance-2.0` | $0.30/sec (720p Pro) |
431
+
432
+ ```ts
433
+ import { VideoClient } from '@blockrun/llm';
434
+
435
+ const client = new VideoClient();
436
+ const result = await client.generate('a red apple slowly spinning on a wooden table');
437
+ console.log(result.data[0].url); // permanent MP4 URL
438
+ console.log(result.data[0].duration_seconds); // 8
439
+
440
+ // Image-to-video
441
+ const r2 = await client.generate('the subject turns and smiles', {
442
+ imageUrl: 'https://example.com/portrait.jpg',
443
+ });
444
+
445
+ // Token360 / Seedance options (silently ignored by xAI Grok video)
446
+ const r3 = await client.generate('aerial drone shot over a snowy mountain', {
447
+ model: 'bytedance/seedance-2.0-fast',
448
+ aspectRatio: '21:9',
449
+ resolution: '1080p',
450
+ generateAudio: true, // omit to use the model's default
451
+ seed: 42,
452
+ watermark: false,
453
+ returnLastFrame: true, // useful for clip chaining
454
+ });
455
+
456
+ // First-and-last-frame interpolation (Seedance only): the model tweens
457
+ // from imageUrl (first frame) to lastFrameUrl (final frame).
458
+ // Priced identically to image-to-video.
459
+ const r4 = await client.generate('the flower blooms in golden morning light', {
460
+ model: 'bytedance/seedance-1.5-pro',
461
+ imageUrl: 'https://example.com/bud.jpg',
462
+ lastFrameUrl: 'https://example.com/bloom.jpg',
463
+ });
464
+
465
+ // Omni / multi-reference (Seedance 2.0 only): up to 9 reference images
466
+ // for character/style consistency. Cite them as "image 1", "image 2" in
467
+ // the prompt. Mutually exclusive with imageUrl / lastFrameUrl /
468
+ // realFaceAssetId.
469
+ const r5 = await client.generate(
470
+ 'the character from image 1 walks through the city from image 2',
471
+ {
472
+ model: 'bytedance/seedance-2.0',
473
+ referenceImageUrls: [
474
+ 'https://example.com/character.jpg',
475
+ 'https://example.com/city.jpg',
476
+ ],
477
+ }
478
+ );
479
+ ```
480
+
481
+ ### Text-to-Speech & Sound Effects
482
+
483
+ `SpeechClient` wraps BlockRun Voice (ElevenLabs): `POST /v1/audio/speech`
484
+ (OpenAI-compatible TTS), `POST /v1/audio/sound-effects`, and the free
485
+ `GET /v1/audio/voices`. TTS price scales with character count:
486
+ `(chars / 1000) × model rate`, minimum $0.001/request. Synthesis is
487
+ synchronous (<1s for Flash).
488
+
489
+ | Model | Price | Max Input | Notes |
490
+ |-------|-------|-----------|-------|
491
+ | `elevenlabs/flash-v2.5` | $0.05/1k chars | 40k chars | ~75ms latency, 32 languages (default) |
492
+ | `elevenlabs/turbo-v2.5` | $0.05/1k chars | 40k chars | ~250ms latency, balanced quality |
493
+ | `elevenlabs/multilingual-v2` | $0.10/1k chars | 10k chars | Long-form narration, audiobooks |
494
+ | `elevenlabs/v3` | $0.10/1k chars | 5k chars | Max expressiveness, 70+ languages |
495
+ | `elevenlabs/sound-effects` | $0.05/generation | 1k chars | Sound effects up to 22s |
496
+
497
+ ```ts
498
+ import { SpeechClient } from '@blockrun/llm';
499
+
500
+ const client = new SpeechClient();
501
+
502
+ // Text-to-speech (voice aliases: sarah, george, laura, charlie,
503
+ // river, roger, callum, harry — or any raw ElevenLabs voice_id)
504
+ const result = await client.generate('Welcome to BlockRun.', { voice: 'george' });
505
+ console.log(result.data[0].url); // audio URL (mp3 by default)
506
+
507
+ // Other formats / speed
508
+ const wav = await client.generate('Breaking news from the world of micropayments.', {
509
+ model: 'elevenlabs/v3',
510
+ responseFormat: 'wav',
511
+ speed: 1.1,
512
+ });
513
+
514
+ // Sound effects (flat $0.05/generation)
515
+ const fx = await client.soundEffect('rain on a tin roof, distant thunder');
516
+
517
+ // List voices (free, rate-limited)
518
+ const voices = await client.listVoices();
519
+ ```
520
+
521
+ ### Virtual Portraits
522
+
523
+ `PortraitClient` wraps `POST /v1/portrait/enroll` (paid, flat **$0.01** promo,
524
+ no KYC). Enroll a face image by URL and get back a Token360 asset id (`ta_xxxxxx`).
525
+ Pass that id as `realFaceAssetId` on a Seedance 2.0 video generation to keep the
526
+ same AI character across clips. Payment settles only after Token360 confirms the
527
+ enrollment, so a failed enrollment never charges your wallet. The returned
528
+ `image_url` is a gateway-mirrored copy of your source image (see `mirrored` /
529
+ `source_image_url`). (Real-person likeness is not supported on BlockRun —
530
+ enrolled portraits are AI characters.)
531
+
532
+ ```ts
533
+ import { PortraitClient, VideoClient } from '@blockrun/llm';
534
+
535
+ const portraits = new PortraitClient();
536
+ const { asset_id } = await portraits.enroll({
537
+ name: 'Spokesperson',
538
+ imageUrl: 'https://example.com/face.jpg', // public https JPG/PNG/WEBP, ≤10 MB
539
+ });
540
+
541
+ // Reuse the same character across Seedance 2.0 clips
542
+ const video = new VideoClient();
543
+ const clip = await video.generate('she waves and smiles', {
544
+ model: 'bytedance/seedance-2.0-fast',
545
+ realFaceAssetId: asset_id,
546
+ });
547
+ console.log(clip.data[0].url);
548
+ ```
549
+
550
+ ### Voice Calls
551
+
552
+ `VoiceClient` wraps `POST /v1/voice/call` (paid, $0.54/call) and
553
+ `GET /v1/voice/call/{callId}` (free polling) — AI-powered outbound phone
554
+ calls powered by Bland.ai. The agent dials the recipient and runs a real-time
555
+ conversation based on your `task` instructions. US + Canada destinations.
556
+
557
+ ```ts
558
+ import { VoiceClient } from '@blockrun/llm';
559
+
560
+ const client = new VoiceClient();
561
+
562
+ // Initiate (paid $0.54)
563
+ const result = await client.call({
564
+ to: '+14155552671',
565
+ task: 'You are a friendly assistant calling to confirm a 3pm dentist appointment.',
566
+ voice: 'maya', // 'nat' | 'josh' | 'maya' | 'june' | 'paige' | 'derek' | 'florian'
567
+ max_duration: 5, // minutes (1–30)
568
+ });
569
+ console.log(result.call_id);
570
+
571
+ // Poll for transcript + recording (free)
572
+ const status = await client.getStatus(result.call_id);
573
+ console.log(status.status, status.recording_url);
574
+ ```
575
+
576
+ Bring your own caller-ID: pass `from: '+14155552671'` (must be a BlockRun
577
+ phone number you own; buy via `/v1/phone/numbers/buy`).
578
+
579
+ ### Standalone Search
580
+
581
+ `SearchClient` wraps `POST /v1/search` — standalone Grok Live Search.
582
+ Pricing: `$0.025/source + margin` (10 sources ≈ `$0.26`).
583
+
584
+ ```ts
585
+ import { SearchClient } from '@blockrun/llm';
586
+
587
+ const client = new SearchClient();
588
+ const result = await client.search('Latest news on x402 adoption', {
589
+ sources: ['x', 'web'],
590
+ maxResults: 10,
591
+ });
592
+ console.log(result.summary);
593
+ for (const url of result.citations ?? []) console.log(url);
594
+ ```
595
+
596
+ ### Surf Crypto Data
597
+
598
+ `SurfClient` exposes the full `/v1/surf/*` catalog — 84+ pay-per-call
599
+ endpoints across CEX/DEX market data, on-chain SQL, wallet intelligence,
600
+ prediction markets (Polymarket + Kalshi), social analytics, news, VC fund
601
+ data, and an OpenAI-compatible chat surface. Flat pricing per call:
602
+
603
+ | Tier | Price/call | Examples |
604
+ |------|-----------|----------|
605
+ | 1 | $0.001 | `/market/price`, `/market/ranking`, `/news/feed`, prediction-market reads, social tweets |
606
+ | 2 | $0.005 | `/exchange/depth`, `/exchange/klines`, `/wallet/detail`, `/search/*`, `/social/ranking` |
607
+ | 3 | $0.020 | `/onchain/sql`, `/onchain/query`, `/onchain/schema`, `/chat/completions` |
608
+
609
+ Because the catalog is broad and evolving, the client deliberately ships a
610
+ generic `get` / `post` pair instead of 84 typed wrappers. Pass the path
611
+ (with or without the `/v1/surf` prefix), query params, or a JSON body —
612
+ type the response via a generic if you want.
613
+
614
+ ```ts
615
+ import { SurfClient } from '@blockrun/llm';
616
+
617
+ const surf = new SurfClient();
618
+
619
+ // Tier 1 — token price ($0.001)
620
+ const btc = await surf.get('/market/price', { symbol: 'BTC' });
621
+
622
+ // Tier 2 — order book depth ($0.005)
623
+ const book = await surf.get('/exchange/depth', {
624
+ exchange: 'binance',
625
+ symbol: 'BTC-USDT',
626
+ });
627
+
628
+ // Tier 3 — raw on-chain SQL against 80+ ClickHouse tables ($0.020)
629
+ const rows = await surf.post('/onchain/sql', {
630
+ query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 5',
631
+ });
632
+
633
+ // Typed response via generic
634
+ type Price = { symbol: string; price: number; timestamp: string };
635
+ const eth = await surf.get<Price>('/market/price', { symbol: 'ETH' });
636
+ ```
637
+
638
+ Full endpoint inventory: <https://blockrun.ai/marketplace/surf>.
639
+
640
+ Methods: `userLookup`, `userInfo`, `followers`, `following`, `followings`,
641
+ `verifiedFollowers`, `userTweets`, `mentions`, `tweetLookup`, `tweetReplies`,
642
+ `tweetThread`, `search`, `trending`, `articlesRising`.
643
+
644
+ ### Market Data (Pyth)
645
+
646
+ `PriceClient` wraps the Pyth-backed market-data endpoints. Crypto, FX and
647
+ commodity are fully free (price + history + list); 12 global stock markets
648
+ and the `usstock` legacy alias charge `$0.001` for price + history (list is
649
+ always free). Pass `requireWallet: false` to construct a free-only client.
650
+
651
+ ```ts
652
+ import { PriceClient } from '@blockrun/llm';
653
+
654
+ const p = new PriceClient({ requireWallet: false });
655
+ const btc = await p.price('crypto', 'BTC-USD');
656
+ const eur = await p.price('fx', 'EUR-USD');
657
+
658
+ // Paid — requires a wallet
659
+ const p2 = new PriceClient();
660
+ const aapl = await p2.price('stocks', 'AAPL', { market: 'us' });
661
+ const bars = await p2.history('stocks', 'AAPL', {
662
+ market: 'us',
663
+ resolution: 'D',
664
+ from: 1700000000,
665
+ to: 1710000000,
666
+ });
667
+ const symbols = await p.listSymbols('crypto', { query: 'sol', limit: 20 });
668
+ ```
669
+
670
+ Supported `StockMarket` values: `us, hk, jp, kr, gb, de, fr, nl, ie, lu, cn, ca`.
671
+
672
+ ### Multi-chain RPC
673
+
674
+ `RpcClient` wraps `POST /v1/rpc/{network}` — standard JSON-RPC 2.0 access to
675
+ 40+ chains through one endpoint (Ethereum, Base, Solana, Polygon, BSC,
676
+ Arbitrum, Optimism, Avalanche, Bitcoin, Sui, and more; powered by Tatum's RPC
677
+ gateway). No API key, no per-chain endpoints: flat **$0.002 per call** in
678
+ USDC; a JSON-RPC batch charges per element.
679
+
680
+ ```ts
681
+ import { RpcClient } from '@blockrun/llm';
682
+
683
+ const client = new RpcClient();
684
+
685
+ // EVM chains speak eth_* JSON-RPC
686
+ const block = await client.call('ethereum', 'eth_blockNumber');
687
+ console.log(parseInt(block.result as string, 16));
688
+
689
+ const balance = await client.call('base', 'eth_getBalance', [
690
+ '0x4200000000000000000000000000000000000006',
691
+ 'latest',
692
+ ]);
693
+
694
+ // Non-EVM chains speak their native JSON-RPC
695
+ const slot = await client.call('solana', 'getSlot');
696
+ const tip = await client.call('bitcoin', 'getblockcount');
697
+
698
+ // Batch: one payment, per-element pricing ($0.002 x N)
699
+ const out = await client.batch('polygon', [
700
+ { method: 'eth_blockNumber' },
701
+ { method: 'eth_gasPrice' },
702
+ ]);
703
+
704
+ console.log(block.network); // 'ethereum' (canonical key from X-Network)
705
+ console.log(block.cacheHit); // true if served from the gateway's hot cache
706
+ console.log(block.txHash); // x402 settlement tx
707
+ ```
708
+
709
+ 40 curated chains are exported as `SUPPORTED_NETWORKS`; common aliases
710
+ (`eth`, `arb`, `op`, `matic`, `bnb`, `avax`, `sol`, `btc`, `xrp`, `dot`, ...)
711
+ resolve server-side (`NETWORK_ALIASES`). Unknown but well-formed slugs fall
712
+ through to a generic `{slug}-mainnet` gateway attempt, so new chains work
713
+ without an SDK update. Hot, low-volatility reads (`eth_chainId`, mined
714
+ blocks/receipts, `getTransaction`, ...) are served from a method-aware
715
+ gateway cache — same price, lower latency.
716
+
717
+ ### Testnet Models (Base Sepolia)
718
+ | Model | Price |
719
+ |-------|-------|
720
+ | `openai/gpt-oss-20b` | $0.001/request |
721
+ | `openai/gpt-oss-120b` | $0.002/request |
722
+
723
+ *Testnet models use flat pricing (no token counting) for simplicity.*
724
+
725
+ ## Standalone Search
726
+
727
+ Search web, X/Twitter, and news without using a chat model:
728
+
729
+ ```typescript
730
+ import { LLMClient } from '@blockrun/llm';
731
+
732
+ const client = new LLMClient();
733
+
734
+ const result = await client.search('latest AI agent frameworks 2026');
735
+ console.log(result.summary);
736
+ for (const cite of result.citations ?? []) {
737
+ console.log(` - ${cite}`);
738
+ }
739
+
740
+ // Filter by source type and date range
741
+ const filtered = await client.search('BlockRun x402', {
742
+ sources: ['web', 'x'],
743
+ fromDate: '2026-01-01',
744
+ maxResults: 5,
745
+ });
746
+ ```
747
+
748
+ ## Image Editing (img2img)
749
+
750
+ Edit existing images with text prompts:
751
+
752
+ ```typescript
753
+ import { LLMClient } from '@blockrun/llm';
754
+
755
+ const client = new LLMClient();
756
+
757
+ const result = await client.imageEdit(
758
+ 'Make the sky purple and add northern lights',
759
+ 'data:image/png;base64,...', // base64 or URL
760
+ { model: 'openai/gpt-image-1' }
761
+ );
762
+ console.log(result.data[0].url);
763
+ ```
764
+
765
+ ## Usage Examples
766
+
767
+ ### Simple Chat
768
+
769
+ ```typescript
770
+ import { LLMClient } from '@blockrun/llm';
771
+
772
+ const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
773
+
774
+ const response = await client.chat('openai/gpt-4o', 'Explain quantum computing');
775
+ console.log(response);
776
+
777
+ // With system prompt
778
+ const response2 = await client.chat('anthropic/claude-sonnet-4', 'Write a haiku', {
779
+ system: 'You are a creative poet.',
780
+ });
781
+ ```
782
+
783
+ ### Smart Routing (ClawRouter)
784
+
785
+ Save up to 78% on inference costs with intelligent model routing. ClawRouter uses a 14-dimension rule-based scoring algorithm to select the cheapest model that can handle your request (<1ms, 100% local).
786
+
787
+ ```typescript
788
+ import { LLMClient } from '@blockrun/llm';
789
+
790
+ const client = new LLMClient();
791
+
792
+ // Auto-route to cheapest capable model
793
+ const result = await client.smartChat('What is 2+2?');
794
+ console.log(result.response); // '4'
795
+ console.log(result.model); // 'google/gemini-2.5-flash'
796
+ console.log(result.routing.tier); // 'SIMPLE'
797
+ console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 78%'
798
+
799
+ // Routing profiles
800
+ const free = await client.smartChat('Hello!', { routingProfile: 'free' }); // Zero cost
801
+ const eco = await client.smartChat('Explain AI', { routingProfile: 'eco' }); // Budget optimized
802
+ const auto = await client.smartChat('Code review', { routingProfile: 'auto' }); // Balanced (default)
803
+ const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
804
+ ```
805
+
806
+ **Routing Profiles:**
807
+
808
+ | Profile | Description | Best For |
809
+ |---------|-------------|----------|
810
+ | `free` | NVIDIA free tier (9 models, smart-routed) | Zero-cost testing, dev, prod |
811
+ | `eco` | Budget-optimized | Cost-sensitive workloads |
812
+ | `auto` | Intelligent routing (default) | General use |
813
+ | `premium` | Best quality models | Critical tasks |
814
+
815
+ **Tiers:**
816
+
817
+ | Tier | Example Tasks | Typical Models |
818
+ |------|---------------|----------------|
819
+ | SIMPLE | Greetings, math, lookups | Gemini Flash, GPT-4o-mini |
820
+ | MEDIUM | Explanations, summaries | GPT-4o, Claude Sonnet |
821
+ | COMPLEX | Analysis, code generation | GPT-5.2, Claude Opus |
822
+ | REASONING | Multi-step logic, planning | o3, DeepSeek Reasoner |
823
+
824
+ ### Full Chat Completion
825
+
826
+ ```typescript
827
+ import { LLMClient, type ChatMessage } from '@blockrun/llm';
828
+
829
+ const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
830
+
831
+ const messages: ChatMessage[] = [
832
+ { role: 'system', content: 'You are a helpful assistant.' },
833
+ { role: 'user', content: 'How do I read a file in Node.js?' },
834
+ ];
835
+
836
+ const result = await client.chatCompletion('openai/gpt-4o', messages);
837
+ console.log(result.choices[0].message.content);
838
+ ```
839
+
840
+ ### Streaming
841
+
842
+ Stream responses token-by-token with automatic x402 payment. Uses a **pre-auth cache** to skip the 402 round-trip on repeat calls to the same model (~200ms saved per request after the first).
843
+
844
+ #### OpenAI-compatible (recommended)
845
+
846
+ ```typescript
847
+ import { OpenAI } from '@blockrun/llm';
848
+
849
+ const client = new OpenAI({ walletKey: process.env.BASE_CHAIN_WALLET_KEY });
850
+
851
+ const stream = await client.chat.completions.create({
852
+ model: 'openai/gpt-5.4',
853
+ messages: [{ role: 'user', content: 'Write a short story about AI agents' }],
854
+ stream: true,
855
+ });
856
+
857
+ for await (const chunk of stream) {
858
+ process.stdout.write(chunk.choices[0]?.delta?.content || '');
859
+ }
860
+ ```
861
+
862
+ #### Native client
863
+
864
+ ```typescript
865
+ import { LLMClient, type ChatMessage } from '@blockrun/llm';
866
+
867
+ const client = new LLMClient();
868
+
869
+ const messages: ChatMessage[] = [
870
+ { role: 'user', content: 'Explain quantum computing in simple terms' },
871
+ ];
872
+
873
+ // Returns a raw fetch Response with SSE body
874
+ const response = await client.chatCompletionStream('google/gemini-2.5-flash', messages);
875
+
876
+ const reader = response.body!.getReader();
877
+ const decoder = new TextDecoder();
878
+
879
+ while (true) {
880
+ const { done, value } = await reader.read();
881
+ if (done) break;
882
+
883
+ const chunk = decoder.decode(value, { stream: true });
884
+ for (const line of chunk.split('\n')) {
885
+ if (!line.startsWith('data: ') || line === 'data: [DONE]') continue;
886
+ const data = JSON.parse(line.slice(6));
887
+ process.stdout.write(data.choices?.[0]?.delta?.content || '');
888
+ }
889
+ }
890
+ ```
891
+
892
+ #### Payment + streaming flow
893
+
894
+ ```
895
+ First call (cache miss):
896
+ 1. Send request → 402 response (BlockRun returns price)
897
+ 2. Sign USDC payment locally (key never leaves machine)
898
+ 3. Retry with PAYMENT-SIGNATURE header + stream: true
899
+ 4. Cache payment requirements for this model (1h TTL)
900
+ 5. Stream tokens as they arrive
901
+
902
+ Subsequent calls (cache hit):
903
+ 1. Pre-sign payment from cache — skip 402 round-trip
904
+ 2. Send request with PAYMENT-SIGNATURE upfront
905
+ 3. Stream tokens immediately (~200ms faster)
906
+ ```
907
+
908
+ ### List Available Models
909
+
910
+ ```typescript
911
+ import { LLMClient } from '@blockrun/llm';
912
+
913
+ const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
914
+ const models = await client.listModels();
915
+
916
+ for (const model of models) {
917
+ console.log(`${model.id}: $${model.inputPrice}/M input`);
918
+ }
919
+ ```
920
+
921
+ ### Multiple Requests
922
+
923
+ ```typescript
924
+ import { LLMClient } from '@blockrun/llm';
925
+
926
+ const client = new LLMClient(); // Uses BASE_CHAIN_WALLET_KEY (never sent to server)
927
+
928
+ const [gpt, claude, gemini] = await Promise.all([
929
+ client.chat('openai/gpt-4o', 'What is 2+2?'),
930
+ client.chat('anthropic/claude-sonnet-4', 'What is 3+3?'),
931
+ client.chat('google/gemini-2.5-flash', 'What is 4+4?'),
932
+ ]);
933
+ ```
934
+
935
+ ## Prediction Markets (Powered by Predexon)
936
+
937
+ Access real-time prediction market data from Polymarket, Kalshi, and Binance Futures via [Predexon](https://predexon.com). No API keys needed — pay-per-request via x402.
938
+
939
+ ### Polymarket
940
+
941
+ ```typescript
942
+ import { LLMClient } from '@blockrun/llm';
943
+
944
+ const client = new LLMClient();
945
+
946
+ // List markets with optional filters ($0.001/request)
947
+ const markets = await client.pm("polymarket/markets");
948
+ const filtered = await client.pm("polymarket/markets", { status: "active", limit: 10 });
949
+ const searched = await client.pm("polymarket/markets", { search: "bitcoin" });
950
+
951
+ // List events ($0.001/request)
952
+ const events = await client.pm("polymarket/events");
953
+
954
+ // Historical trades ($0.001/request)
955
+ const trades = await client.pm("polymarket/trades");
956
+
957
+ // OHLCV candlestick data for a specific condition ($0.001/request)
958
+ const candles = await client.pm("polymarket/candlesticks/0x1234abcd...");
959
+
960
+ // Wallet profile ($0.005/request — tier 2)
961
+ const profile = await client.pm("polymarket/wallet/0xABC123...");
962
+
963
+ // Wallet P&L ($0.005/request — tier 2)
964
+ const pnl = await client.pm("polymarket/wallet/pnl/0xABC123...");
965
+
966
+ // Global leaderboard ($0.001/request)
967
+ const leaderboard = await client.pm("polymarket/leaderboard");
968
+ ```
969
+
970
+ ### Kalshi & Binance
971
+
972
+ ```typescript
973
+ // Kalshi markets ($0.001/request)
974
+ const kalshiMarkets = await client.pm("kalshi/markets");
975
+
976
+ // Kalshi trades ($0.001/request)
977
+ const kalshiTrades = await client.pm("kalshi/trades");
978
+
979
+ // Binance candles for supported pairs ($0.001/request)
980
+ const btcCandles = await client.pm("binance/candles/BTCUSDT");
981
+ const ethCandles = await client.pm("binance/candles/ETHUSDT");
982
+ // Also: SOLUSDT, XRPUSDT
983
+ ```
984
+
985
+ ### Cross-Platform
986
+
987
+ ```typescript
988
+ // Cross-platform matching pairs ($0.001/request)
989
+ const pairs = await client.pm("matching-markets/pairs");
990
+ ```
991
+
992
+ All current endpoints are GET. The `pmQuery()` method is available for future POST endpoints.
993
+
994
+ Works on both `LLMClient` (Base) and `SolanaLLMClient`.
995
+
996
+ ## Exa Web Search (Powered by Exa)
997
+
998
+ Access [Exa](https://exa.ai)'s neural web search via x402. No API keys needed — pay-per-request. Available on **`LLMClient` (Base USDC)** and `SolanaLLMClient` (Solana USDC). Use Base as the primary path; the Solana gateway is awaiting `EXA_API_KEY` provisioning.
999
+
1000
+ | Method | Description | Price |
1001
+ |---|---|---|
1002
+ | `exaSearch(query, options?)` | Neural/keyword web search | $0.01/request |
1003
+ | `exaFindSimilar(url, options?)` | Find semantically similar pages | $0.01/request |
1004
+ | `exaContents(urls, options?)` | Extract full text from URLs | $0.002/URL |
1005
+ | `exaAnswer(query, options?)` | AI answer grounded in web search | $0.01/request |
1006
+ | `exa(path, body)` | Generic proxy for any Exa endpoint | varies |
1007
+
1008
+ ```typescript
1009
+ import { LLMClient } from '@blockrun/llm';
1010
+
1011
+ const client = new LLMClient();
1012
+
1013
+ // Neural web search ($0.01/request)
1014
+ const results = await client.exaSearch("latest AI safety research", { numResults: 5 });
1015
+ const news = await client.exaSearch("bitcoin ETF news", { category: "news", numResults: 10 });
1016
+
1017
+ // Find similar pages ($0.01/request)
1018
+ const similar = await client.exaFindSimilar("https://openai.com/research/gpt-4", { numResults: 5 });
1019
+
1020
+ // Extract content from URLs ($0.002/URL)
1021
+ const content = await client.exaContents(["https://arxiv.org/abs/2303.08774"]);
1022
+
1023
+ // AI-generated answer from live web ($0.01/request)
1024
+ const answer = await client.exaAnswer("What is the current state of AI safety research?");
1025
+
1026
+ // Generic proxy for any Exa endpoint
1027
+ const custom = await client.exa("search", { query: "transformer architecture", numResults: 5 });
1028
+ ```
1029
+
1030
+ Same surface on `SolanaLLMClient` once Solana-side `EXA_API_KEY` is provisioned.
1031
+
1032
+ ## Configuration
1033
+
1034
+ ```typescript
1035
+ // Default: reads BASE_CHAIN_WALLET_KEY from environment
1036
+ const client = new LLMClient();
1037
+
1038
+ // Or pass options explicitly
1039
+ const client = new LLMClient({
1040
+ privateKey: '0x...', // Your wallet key (never sent to server)
1041
+ apiUrl: 'https://blockrun.ai/api', // Optional
1042
+ timeout: 60000, // Optional (ms)
1043
+ });
1044
+ ```
1045
+
1046
+ ## Environment Variables
1047
+
1048
+ | Variable | Description |
1049
+ |----------|-------------|
1050
+ | `BASE_CHAIN_WALLET_KEY` | Your Base chain wallet private key (for Base / `LLMClient`) |
1051
+ | `SOLANA_WALLET_KEY` | Your Solana wallet secret key - bs58 encoded (for `SolanaLLMClient`) |
1052
+ | `BLOCKRUN_API_URL` | API endpoint (optional, default: https://blockrun.ai/api) |
1053
+
1054
+ ## Error Handling
1055
+
1056
+ ```typescript
1057
+ import { LLMClient, APIError, PaymentError } from '@blockrun/llm';
1058
+
1059
+ const client = new LLMClient();
1060
+
1061
+ try {
1062
+ const response = await client.chat('openai/gpt-4o', 'Hello!');
1063
+ } catch (error) {
1064
+ if (error instanceof PaymentError) {
1065
+ console.error('Payment failed - check USDC balance');
1066
+ } else if (error instanceof APIError) {
1067
+ console.error(`API error: ${error.message}`);
1068
+ }
1069
+ }
1070
+ ```
1071
+
1072
+ ## Testing
1073
+
1074
+ ### Running Unit Tests
1075
+
1076
+ Unit tests do not require API access or funded wallets:
1077
+
1078
+ ```bash
1079
+ npm test # Run tests in watch mode
1080
+ npm test run # Run tests once
1081
+ npm test -- --coverage # Run with coverage report
1082
+ ```
1083
+
1084
+ ### Running Integration Tests
1085
+
1086
+ Integration tests call the production API and require:
1087
+ - A funded Base wallet with USDC ($1+ recommended)
1088
+ - `BASE_CHAIN_WALLET_KEY` environment variable set
1089
+ - Estimated cost: ~$0.05 per test run
1090
+
1091
+ ```bash
1092
+ export BASE_CHAIN_WALLET_KEY=0x...
1093
+ npm test -- test/integration # Run integration tests only
1094
+ ```
1095
+
1096
+ Integration tests are automatically skipped if `BASE_CHAIN_WALLET_KEY` is not set.
1097
+
1098
+ ## Setting Up Your Wallet
1099
+
1100
+ ### Base (EVM)
1101
+ 1. Create a wallet on Base (Coinbase Wallet, MetaMask, etc.)
1102
+ 2. Get USDC on Base for API payments
1103
+ 3. Export your private key and set as `BASE_CHAIN_WALLET_KEY`
1104
+
1105
+ ```bash
1106
+ # .env
1107
+ BASE_CHAIN_WALLET_KEY=0x...
1108
+ ```
1109
+
1110
+ ### Solana
1111
+ 1. Create a Solana wallet (Phantom, Backpack, Solflare, etc.)
1112
+ 2. Get USDC on Solana for API payments
1113
+ 3. Export your secret key and set as `SOLANA_WALLET_KEY`
1114
+
1115
+ ```bash
1116
+ # .env
1117
+ SOLANA_WALLET_KEY=...your_bs58_secret_key
1118
+ ```
1119
+
1120
+ Note: Solana transactions are gasless for the user - the CDP facilitator pays for transaction fees.
1121
+
1122
+ ## Security
1123
+
1124
+ ### Private Key Safety
1125
+
1126
+ - **Private key stays local**: Your key is only used for signing on your machine
1127
+ - **No custody**: BlockRun never holds your funds
1128
+ - **Verify transactions**: All payments are on-chain and verifiable
1129
+
1130
+ ### Best Practices
1131
+
1132
+ **Private Key Management:**
1133
+ - Use environment variables, never hard-code keys
1134
+ - Use dedicated wallets for API payments (separate from main holdings)
1135
+ - Set spending limits by only funding payment wallets with small amounts
1136
+ - Never commit `.env` files to version control
1137
+ - Rotate keys periodically
1138
+
1139
+ **Input Validation:**
1140
+ The SDK validates all inputs before API requests:
1141
+ - Private keys (format, length, valid hex)
1142
+ - API URLs (HTTPS required for production, HTTP allowed for localhost)
1143
+ - Model names and parameters (ranges for max\_tokens, temperature, top\_p)
1144
+
1145
+ **Error Sanitization:**
1146
+ API errors are automatically sanitized to prevent sensitive information leaks.
1147
+
1148
+ **Monitoring:**
1149
+ ```typescript
1150
+ const address = client.getWalletAddress();
1151
+ console.log(`View transactions: https://basescan.org/address/${address}`);
1152
+ ```
1153
+
1154
+ **Keep Updated:**
1155
+ ```bash
1156
+ npm update @blockrun/llm # Get security patches
1157
+ ```
1158
+
1159
+ ## TypeScript Support
1160
+
1161
+ Full TypeScript support with exported types:
1162
+
1163
+ ```typescript
1164
+ import {
1165
+ LLMClient,
1166
+ OpenAI,
1167
+ type ChatMessage,
1168
+ type ChatResponse,
1169
+ type ChatOptions,
1170
+ type ChatCompletionOptions,
1171
+ type Model,
1172
+ // Smart routing types
1173
+ type SmartChatOptions,
1174
+ type SmartChatResponse,
1175
+ type RoutingDecision,
1176
+ type RoutingProfile,
1177
+ type RoutingTier,
1178
+ APIError,
1179
+ PaymentError,
1180
+ } from '@blockrun/llm';
1181
+
1182
+ // chatCompletionStream returns a standard fetch Response with SSE body
1183
+ const streamResponse: Response = await client.chatCompletionStream(model, messages, options);
1184
+
1185
+ // OpenAI-compat stream returns AsyncIterable
1186
+ const stream: AsyncIterable<OpenAIChatCompletionChunk> = await openaiClient.chat.completions.create({
1187
+ model, messages, stream: true
1188
+ });
1189
+ ```
1190
+
1191
+ ## Agent Wallet Setup
1192
+
1193
+ One-line setup for agent runtimes (Claude Code skills, MCP servers, etc.):
1194
+
1195
+ ```typescript
1196
+ import { setupAgentWallet } from '@blockrun/llm';
1197
+
1198
+ // Auto-creates wallet if none exists, returns ready client
1199
+ const client = setupAgentWallet();
1200
+ const response = await client.chat('openai/gpt-5.4', 'Hello!');
1201
+ ```
1202
+
1203
+ For Solana:
1204
+
1205
+ ```typescript
1206
+ import { setupAgentSolanaWallet } from '@blockrun/llm';
1207
+
1208
+ const client = await setupAgentSolanaWallet();
1209
+ const response = await client.chat('anthropic/claude-sonnet-4.6', 'Hello!');
1210
+ ```
1211
+
1212
+ Check wallet status:
1213
+
1214
+ ```typescript
1215
+ import { status } from '@blockrun/llm';
1216
+
1217
+ await status();
1218
+ // Wallet: 0xCC8c...5EF8
1219
+ // Balance: $5.30 USDC
1220
+ ```
1221
+
1222
+ ## Wallet Scanning
1223
+
1224
+ The SDK auto-detects wallets from any provider on your system:
1225
+
1226
+ ```typescript
1227
+ import { scanWallets, scanSolanaWallets } from '@blockrun/llm';
1228
+
1229
+ // Scans ~/.<dir>/wallet.json for Base wallets
1230
+ const baseWallets = scanWallets();
1231
+
1232
+ // Scans ~/.<dir>/solana-wallet.json and ~/.brcc/wallet.json
1233
+ const solWallets = scanSolanaWallets();
1234
+ ```
1235
+
1236
+ `getOrCreateWallet()` checks scanned wallets first, so if you already have a wallet from another BlockRun tool, it will be reused automatically.
1237
+
1238
+ ## Response Caching
1239
+
1240
+ The SDK caches responses to avoid duplicate payments:
1241
+
1242
+ ```typescript
1243
+ import { getCachedByRequest, saveToCache, clearCache } from '@blockrun/llm';
1244
+
1245
+ // Automatic TTLs by endpoint:
1246
+ // - Search: 15 minutes
1247
+ // - Models: 24 hours
1248
+ // - Chat/Image: no cache (every call is unique)
1249
+
1250
+ // Manual cache management
1251
+ clearCache(); // Remove all cached responses
1252
+ ```
1253
+
1254
+ ## Cost Logging
1255
+
1256
+ Track spending across sessions:
1257
+
1258
+ ```typescript
1259
+ import { logCost, getCostSummary } from '@blockrun/llm';
1260
+
1261
+ // Costs are logged to ~/.blockrun/data/costs.jsonl
1262
+ const summary = getCostSummary();
1263
+ console.log(`Total: $${summary.totalUsd.toFixed(2)}`);
1264
+ console.log(`Calls: ${summary.calls}`);
1265
+ console.log(`By model:`, summary.byModel);
1266
+ ```
1267
+
1268
+ ## Anthropic SDK Compatibility
1269
+
1270
+ Use the official Anthropic SDK interface with BlockRun's pay-per-request backend:
1271
+
1272
+ ```typescript
1273
+ import { AnthropicClient } from '@blockrun/llm';
1274
+
1275
+ const client = new AnthropicClient(); // Auto-detects wallet, auto-pays
1276
+
1277
+ const response = await client.messages.create({
1278
+ model: 'claude-sonnet-4-6',
1279
+ max_tokens: 1024,
1280
+ messages: [{ role: 'user', content: 'Hello!' }],
1281
+ });
1282
+ console.log(response.content[0].text);
1283
+
1284
+ // Any model works in Anthropic format
1285
+ const gptResponse = await client.messages.create({
1286
+ model: 'openai/gpt-5.4',
1287
+ max_tokens: 1024,
1288
+ messages: [{ role: 'user', content: 'Hello from GPT!' }],
1289
+ });
1290
+ ```
1291
+
1292
+ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch that handles x402 payment automatically. Your private key never leaves your machine.
1293
+
1294
+ ## Links
1295
+
1296
+ - [Website](https://blockrun.ai)
1297
+ - [Documentation](https://github.com/BlockRunAI/awesome-blockrun/tree/main/docs)
1298
+ - [GitHub](https://github.com/blockrunai/blockrun-llm-ts)
1299
+ - [Telegram](https://t.me/+mroQv4-4hGgzOGUx)
1300
+
1301
+ ## Frequently Asked Questions
1302
+
1303
+ ### What is @blockrun/llm?
1304
+ @blockrun/llm is a TypeScript SDK that provides pay-per-request access to 40+ large language models from OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more. It uses the x402 protocol for automatic USDC micropayments — no API keys, no subscriptions, no vendor lock-in.
1305
+
1306
+ ### How does payment work?
1307
+ When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
1308
+
1309
+ ### What is smart routing / ClawRouter?
1310
+ ClawRouter is a built-in smart routing engine that analyzes your request across 14 dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. It can save up to 78% on LLM costs compared to using premium models for every request.
1311
+
1312
+ ### Does it support streaming?
1313
+ Yes — as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).
1314
+
1315
+ ### How much does it cost?
1316
+ Pay only for what you use. Prices start at $0.0002 per request (GPT-5 Nano). There are no minimums, subscriptions, or monthly fees. $5 in USDC gets you thousands of requests.
1317
+
1318
+ ### Does it support both Base and Solana?
1319
+ Yes. Use `LLMClient` for Base (EVM) payments and `SolanaLLMClient` for Solana payments. Same API, different payment chain.
1320
+
1321
+ ## License
1322
+
1323
+ MIT