@blockrun/llm 3.12.0 → 3.13.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  ### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
6
6
 
7
- The smart-routing SDK for <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models — every request goes to the cheapest model that can handle it,
7
+ The smart-routing SDK for <!-- br:models.chatVisible -->70<!-- /br:models.chatVisible --> models — every request goes to the cheapest model that can handle it,
8
8
  paid per-request in USDC. No API keys. No subscriptions. No vendor lock-in.
9
9
 
10
10
  [![npm](https://img.shields.io/npm/v/@blockrun/llm.svg?style=flat-square)](https://www.npmjs.com/package/@blockrun/llm)
@@ -41,7 +41,7 @@ console.log(r.response); // the proof
41
41
  ## Why This SDK
42
42
 
43
43
  - 🧠 **Smart routing that pays for itself** — the bundled [Router Core V3](https://github.com/BlockRunAI/router-core) engine (shared with [ClawRouter](https://github.com/BlockRunAI/ClawRouter)) classifies every request locally in <1ms across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and routes to the cheapest capable model. The main event.
44
- - 🆓 **<!-- br:models.free -->6<!-- /br:models.free --> genuinely free models** — NVIDIA-hosted, $0 in and out, incl. 1M-context DeepSeek V4 Flash and a multimodal Nemotron. No rate-limit gimmicks.
44
+ - 🆓 **<!-- br:models.free -->5<!-- /br:models.free --> genuinely free models** — NVIDIA-hosted, $0 in and out, incl. 1M-context DeepSeek V4 Flash and a multimodal Nemotron. No rate-limit gimmicks.
45
45
  - 🔐 **No API keys** — your wallet signature is your authentication. No accounts, no dashboards, no key rotation.
46
46
  - 💸 **Pay per request in USDC** — x402 micropayments on Base or Solana. $5 covers thousands of requests; agents can pay their own way.
47
47
  - 🛡️ **Automatic failover** — transient errors (timeouts, 429, 5xx) walk the router's ranked fallback chain instead of failing your request.
@@ -53,8 +53,8 @@ console.log(r.response); // the proof
53
53
  | | OpenAI SDK | OpenRouter | LiteLLM | **@blockrun/llm** |
54
54
  | ------------------ | -------------- | ----------------- | ---------------- | ----------------------------------------------------------------------- |
55
55
  | **Cost routing** | ✗ one vendor | Manual selection | Manual selection | **Automatic — <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
56
- | **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->71<!-- /br:models.chatVisible -->, one wallet** |
57
- | **Free tier** | ✗ | Rate-limited | ✗ | **<!-- br:models.free -->6<!-- /br:models.free --> models, no signup** |
56
+ | **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->70<!-- /br:models.chatVisible -->, one wallet** |
57
+ | **Free tier** | ✗ | Rate-limited | ✗ | **<!-- br:models.free -->5<!-- /br:models.free --> models, no signup** |
58
58
  | **Auth** | API key | Account + API key | Your API keys | **Wallet signature** |
59
59
  | **Payment** | Card + invoice | Credit card | BYO keys | **USDC per-request** |
60
60
  | **Agent-ready** | ✗ | ✗ | ✗ | **✓ — agents fund their own wallet** |
@@ -123,11 +123,11 @@ import { LLMClient } from '@blockrun/llm';
123
123
  const client = new LLMClient(); // Wallet still required for signing, but $0 charged
124
124
 
125
125
  // Option 1: call a free model directly
126
- const reply = await client.chat('nvidia/deepseek-v4-flash', 'Explain x402 in 1 sentence');
126
+ const reply = await client.chat('nvidia/step-3.7-flash', 'Explain x402 in 1 sentence');
127
127
 
128
128
  // Option 2: let the smart router pick — 'eco' ranks the free NVIDIA tier first
129
129
  const result = await client.smartChat('What is 2+2?', { routingProfile: 'eco' });
130
- console.log(result.model); // 'nvidia/deepseek-v4-flash' ($0 — verified live)
130
+ console.log(result.model); // 'nvidia/step-3.7-flash' ($0 — verified live)
131
131
  console.log(result.response); // '4'
132
132
  console.log(result.routing.savings); // 1 (100%)
133
133
  ```
@@ -137,11 +137,10 @@ There is no `free` routing profile in `smartChat()` — `routingProfile` accepts
137
137
  own proxy, not of this SDK's router options.) For guaranteed $0, pin a
138
138
  `nvidia/*` model; for smart-routed $0-first, use `eco`.
139
139
 
140
- **Available free models** (input + output both $0, all NVIDIA-hosted, from the live `/v1/models` catalog, last refreshed 2026-08-10):
140
+ **Available free models** (input + output both $0, all NVIDIA-hosted, from the live `/v1/models` catalog, last refreshed 2026-08-12):
141
141
 
142
142
  | Model ID | Context | Best For |
143
143
  |----------|---------|----------|
144
- | `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash — 284B / 13B active MoE. Best free chat / summarization / light reasoning. Capacity-constrained: requests may be answered by an equivalent free model |
145
144
  | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning — text + images + video + audio (ChartQA 90.3, DocVQA 95.6) |
146
145
  | `nvidia/mistral-nemotron` | 131K | Mistral × NVIDIA instruction model — fast (~0.2s), strong instruction following |
147
146
  | `nvidia/step-3.7-flash` | 131K | StepFun Step 3.7 Flash — fast lightweight reasoning |
@@ -237,10 +236,10 @@ propagate immediately so wallet / auth issues surface fast.
237
236
 
238
237
  ```typescript
239
238
  // Manually pass a fallback chain to chat() / chatCompletion()
240
- const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
241
- fallbackModels: ['nvidia/llama-4-maverick', 'nvidia/mistral-small-4-119b'],
239
+ const reply = await client.chat('nvidia/step-3.7-flash', 'hello', {
240
+ fallbackModels: ['nvidia/mistral-nemotron', 'nvidia/gpt-oss-120b'],
242
241
  });
243
- // If deepseek-v4-flash times out, the SDK retries against the next model
242
+ // If step-3.7-flash times out, the SDK retries against the next model
244
243
  // and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
245
244
  ```
246
245
 
@@ -248,7 +247,7 @@ const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
248
247
 
249
248
  | Profile | Strategy | Savings vs Opus 5 | Best For |
250
249
  |---------|----------|-------------------|----------|
251
- | `eco` | Cheapest capable model — ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
250
+ | `eco` | Cheapest capable model — ranks the <!-- br:models.free -->5<!-- /br:models.free -->-model free NVIDIA tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
252
251
  | `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%** | General use |
253
252
  | `premium` | Top-tier models (OpenAI, Anthropic) | 0% | Quality-critical tasks |
254
253
 
@@ -396,7 +395,7 @@ You only do this when your balance runs low. Three ways to get USDC into your wa
396
395
 
397
396
  - **(b) Transfer existing USDC.** Send USDC you already hold to your wallet address (`client.getWalletAddress()`). On Base, send Base USDC; on Solana (`SolanaLLMClient`), send Solana SPL USDC.
398
397
 
399
- - **(c) Skip funding entirely.** Use the free NVIDIA models (e.g. `nvidia/deepseek-v4-flash`) — every call is **$0**, no balance required.
398
+ - **(c) Skip funding entirely.** Use the free NVIDIA models (e.g. `nvidia/step-3.7-flash`) — every call is **$0**, no balance required.
400
399
 
401
400
  $5 of USDC covers thousands of paid requests. Check your balance any time:
402
401
 
@@ -556,7 +555,7 @@ thinking modes. V4 Pro is the new flagship paid SKU — 1.6T MoE / 49B active,
556
555
  | Model | Input Price | Output Price | Context | Notes |
557
556
  |-------|-------------|--------------|---------|-------|
558
557
  | `deepseek/deepseek-v4-pro` | $0.435/M | $0.87/M | 1M | V4 flagship — strongest open-weight reasoner. The 75% launch promo became the permanent list price after 2026-05-31 |
559
- | `deepseek/deepseek-chat` | $0.20/M | $0.40/M | 1M | V4 Flash non-thinking (paid endpoint with 5MB request bodies; same upstream as `nvidia/deepseek-v4-flash`) |
558
+ | `deepseek/deepseek-chat` | $0.14/M | $0.28/M | 1M | V4 Flash non-thinking (paid endpoint with 5MB request bodies) |
560
559
  | `deepseek/deepseek-reasoner` | $0.20/M | $0.40/M | 1M | V4 Flash thinking (same upstream as `deepseek-chat`, thinking enabled by default) |
561
560
 
562
561
  ### xAI Grok
@@ -587,25 +586,22 @@ auto-pick them.
587
586
 
588
587
  ### NVIDIA (Free) + Moonshot
589
588
 
590
- Free tier refreshed 2026-04-28: added `nvidia/deepseek-v4-flash` (1M context)
591
- and Nemotron Nano Omni (vision). `nvidia/gpt-oss-120b` and
592
- `nvidia/gpt-oss-20b` were briefly delisted over privacy concerns then
593
- **re-enabled 2026-04-30** with `available: true` + `hidden: true` — they
594
- no longer appear in `/v1/models` (so SmartChat won't auto-pick them) but
595
- direct calls by full ID still return HTTP 200. `nvidia/deepseek-v4-pro`,
596
- `nvidia/deepseek-v3.2`, and `nvidia/glm-4.7` are hidden because NVIDIA's
597
- NIM deployment is hung — backend MODEL_REDIRECTS forwards calls to V4
598
- Flash / qwen3-coder. `nvidia/qwen3-next-80b-a3b-thinking` hit NVIDIA
599
- end-of-life 2026-05-21 (HTTP 410) and is auto-redirected to
600
- `nvidia/llama-4-maverick`.
589
+ Free tier refreshed 2026-08-12. NVIDIA has retired (HTTP 410 end-of-life)
590
+ the entire free DeepSeek family — `nvidia/deepseek-v4-flash` was the last
591
+ to go — along with `llama-4-maverick`, the qwen3 SKUs, and the free
592
+ Mistral small/large SKUs. Retired IDs stay callable: the gateway
593
+ auto-redirects them to a healthy free model, so pinned callers still get
594
+ a 200. `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` remain callable by
595
+ direct ID but are hidden from `/v1/models` over the NVIDIA free tier's
596
+ prompt-retention terms (so SmartChat won't auto-pick them).
601
597
 
602
598
  | Model | Input Price | Output Price | Notes |
603
599
  |-------|-------------|--------------|-------|
604
- | `nvidia/deepseek-v4-flash` | **FREE** | **FREE** | 284B / 13B active MoE, 1M context — best free chat / summarization / light reasoning |
600
+ | `nvidia/step-3.7-flash` | **FREE** | **FREE** | Fast general-purpose chat + reasoning, 131K |
601
+ | `nvidia/mistral-nemotron` | **FREE** | **FREE** | Fast free Mistral, 131K |
605
602
  | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 31B / 3.2B active MoE, 256K — only vision-capable free model |
606
- | `nvidia/mistral-small-4-119b` | **FREE** | **FREE** | ⚠️ Upstream timing out as of 2026-06-07 |
607
- | `nvidia/llama-4-maverick` | **FREE** | **FREE** | Meta Llama 4 Maverick MoE |
608
- | `nvidia/qwen3-coder-480b` | **FREE** | **FREE** | Coding-optimised 480B MoE |
603
+ | `nvidia/nemotron-nano-9b-v2` | **FREE** | **FREE** | Compact fast chat, 131K |
604
+ | `nvidia/nemotron-nano-12b-v2-vl` | **FREE** | **FREE** | Compact vision, 131K |
609
605
  | `nvidia/gpt-oss-120b` | **FREE** | **FREE** | Hidden from `/v1/models` for privacy but direct calls still work — 123 tok/s |
610
606
  | `nvidia/gpt-oss-20b` | **FREE** | **FREE** | Hidden from `/v1/models` but direct calls still work — 155 tok/s |
611
607
  | `moonshot/kimi-k2.5` | $0.60/M | $3.00/M | Direct from Moonshot — replaces `nvidia/kimi-k2.5` |
@@ -1061,14 +1057,14 @@ const auto = await client.smartChat('Code review', { routingProfile: 'auto' });
1061
1057
  const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
1062
1058
 
1063
1059
  // Guaranteed $0: call a free NVIDIA model directly
1064
- const free = await client.chat('nvidia/deepseek-v4-flash', 'Hello!');
1060
+ const free = await client.chat('nvidia/step-3.7-flash', 'Hello!');
1065
1061
  ```
1066
1062
 
1067
1063
  **Routing Profiles:**
1068
1064
 
1069
1065
  | Profile | Description | Best For |
1070
1066
  |---------|-------------|----------|
1071
- | `eco` | Budget-optimized — ranks the <!-- br:models.free -->6<!-- /br:models.free -->-model free NVIDIA tier first | Cost-sensitive workloads, zero-cost testing |
1067
+ | `eco` | Budget-optimized — ranks the <!-- br:models.free -->5<!-- /br:models.free -->-model free NVIDIA tier first | Cost-sensitive workloads, zero-cost testing |
1072
1068
  | `auto` | Intelligent routing (default) | General use |
1073
1069
  | `premium` | Best quality models | Critical tasks |
1074
1070
 
@@ -1536,6 +1532,25 @@ matches on that derived address — so a wallet file cannot claim an address it
1536
1532
  cannot sign for, nor be adopted by one. `listDiscoveredWallets()` never returns
1537
1533
  private keys.
1538
1534
 
1535
+ ### One wallet across every BlockRun product
1536
+
1537
+ Base wallet resolution, discovery, and adoption are implemented in
1538
+ [`@blockrun/core`](https://www.npmjs.com/package/@blockrun/core), the shared kernel
1539
+ this SDK, the `blockrun` CLI, and clawrouter-codex all read. Defining the canonical
1540
+ order in one place is what keeps them in agreement — when each product carried its
1541
+ own copy, they drifted, and a fix made here did not reach the CLI.
1542
+
1543
+ The kernel is bundled into the SDK at build time (frozen, reviewed bytes — no
1544
+ floating dependency), so there is nothing extra to install. Set `BLOCKRUN_HOME`
1545
+ to override the base directory (`~` by default) for test isolation; unset,
1546
+ behaviour is unchanged. **Treat `BLOCKRUN_HOME` as security-sensitive**: it
1547
+ redirects where the signing key is read from and written to, so an environment
1548
+ that can set it controls the wallet as surely as one that can set
1549
+ `BLOCKRUN_WALLET_KEY`. Set it before importing the SDK — the exported
1550
+ `WALLET_FILE_PATH`/`WALLET_DIR_PATH` constants snapshot at import (all internal
1551
+ reads and writes resolve per call). Solana resolution is still SDK-local and
1552
+ does not honor `BLOCKRUN_HOME`.
1553
+
1539
1554
  For a single run without changing anything, use
1540
1555
  `export BLOCKRUN_WALLET_KEY=<private-key>`.
1541
1556
 
@@ -1605,7 +1620,7 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
1605
1620
  ## Frequently Asked Questions
1606
1621
 
1607
1622
  ### What is @blockrun/llm?
1608
- @blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->71<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol — no API keys, no subscriptions, no vendor lock-in.
1623
+ @blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->70<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol — no API keys, no subscriptions, no vendor lock-in.
1609
1624
 
1610
1625
  ### How does payment work?
1611
1626
  When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.