@blockrun/llm 3.13.0 → 3.13.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +28 -32
- package/dist/index.cjs +1 -1
- package/dist/index.js +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
|
|
6
6
|
|
|
7
|
-
The smart-routing SDK for <!-- br:models.chatVisible -->
|
|
7
|
+
The smart-routing SDK for <!-- br:models.chatVisible -->70<!-- /br:models.chatVisible --> models — every request goes to the cheapest model that can handle it,
|
|
8
8
|
paid per-request in USDC. No API keys. No subscriptions. No vendor lock-in.
|
|
9
9
|
|
|
10
10
|
[](https://www.npmjs.com/package/@blockrun/llm)
|
|
@@ -41,7 +41,7 @@ console.log(r.response); // the proof
|
|
|
41
41
|
## Why This SDK
|
|
42
42
|
|
|
43
43
|
- 🧠 **Smart routing that pays for itself** — the bundled [Router Core V3](https://github.com/BlockRunAI/router-core) engine (shared with [ClawRouter](https://github.com/BlockRunAI/ClawRouter)) classifies every request locally in <1ms across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and routes to the cheapest capable model. The main event.
|
|
44
|
-
- 🆓 **<!-- br:models.free -->
|
|
44
|
+
- 🆓 **<!-- br:models.free -->5<!-- /br:models.free --> genuinely free models** — NVIDIA-hosted, $0 in and out, incl. 1M-context DeepSeek V4 Flash and a multimodal Nemotron. No rate-limit gimmicks.
|
|
45
45
|
- 🔐 **No API keys** — your wallet signature is your authentication. No accounts, no dashboards, no key rotation.
|
|
46
46
|
- 💸 **Pay per request in USDC** — x402 micropayments on Base or Solana. $5 covers thousands of requests; agents can pay their own way.
|
|
47
47
|
- 🛡️ **Automatic failover** — transient errors (timeouts, 429, 5xx) walk the router's ranked fallback chain instead of failing your request.
|
|
@@ -53,8 +53,8 @@ console.log(r.response); // the proof
|
|
|
53
53
|
| | OpenAI SDK | OpenRouter | LiteLLM | **@blockrun/llm** |
|
|
54
54
|
| ------------------ | -------------- | ----------------- | ---------------- | ----------------------------------------------------------------------- |
|
|
55
55
|
| **Cost routing** | ✗ one vendor | Manual selection | Manual selection | **Automatic — <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
|
|
56
|
-
| **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->
|
|
57
|
-
| **Free tier** | ✗ | Rate-limited | ✗ | **<!-- br:models.free -->
|
|
56
|
+
| **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->70<!-- /br:models.chatVisible -->, one wallet** |
|
|
57
|
+
| **Free tier** | ✗ | Rate-limited | ✗ | **<!-- br:models.free -->5<!-- /br:models.free --> models, no signup** |
|
|
58
58
|
| **Auth** | API key | Account + API key | Your API keys | **Wallet signature** |
|
|
59
59
|
| **Payment** | Card + invoice | Credit card | BYO keys | **USDC per-request** |
|
|
60
60
|
| **Agent-ready** | ✗ | ✗ | ✗ | **✓ — agents fund their own wallet** |
|
|
@@ -123,11 +123,11 @@ import { LLMClient } from '@blockrun/llm';
|
|
|
123
123
|
const client = new LLMClient(); // Wallet still required for signing, but $0 charged
|
|
124
124
|
|
|
125
125
|
// Option 1: call a free model directly
|
|
126
|
-
const reply = await client.chat('nvidia/
|
|
126
|
+
const reply = await client.chat('nvidia/step-3.7-flash', 'Explain x402 in 1 sentence');
|
|
127
127
|
|
|
128
128
|
// Option 2: let the smart router pick — 'eco' ranks the free NVIDIA tier first
|
|
129
129
|
const result = await client.smartChat('What is 2+2?', { routingProfile: 'eco' });
|
|
130
|
-
console.log(result.model); // 'nvidia/
|
|
130
|
+
console.log(result.model); // 'nvidia/step-3.7-flash' ($0 — verified live)
|
|
131
131
|
console.log(result.response); // '4'
|
|
132
132
|
console.log(result.routing.savings); // 1 (100%)
|
|
133
133
|
```
|
|
@@ -137,11 +137,10 @@ There is no `free` routing profile in `smartChat()` — `routingProfile` accepts
|
|
|
137
137
|
own proxy, not of this SDK's router options.) For guaranteed $0, pin a
|
|
138
138
|
`nvidia/*` model; for smart-routed $0-first, use `eco`.
|
|
139
139
|
|
|
140
|
-
**Available free models** (input + output both $0, all NVIDIA-hosted, from the live `/v1/models` catalog, last refreshed 2026-08-
|
|
140
|
+
**Available free models** (input + output both $0, all NVIDIA-hosted, from the live `/v1/models` catalog, last refreshed 2026-08-12):
|
|
141
141
|
|
|
142
142
|
| Model ID | Context | Best For |
|
|
143
143
|
|----------|---------|----------|
|
|
144
|
-
| `nvidia/deepseek-v4-flash` | 1M | DeepSeek V4 Flash — 284B / 13B active MoE. Best free chat / summarization / light reasoning. Capacity-constrained: requests may be answered by an equivalent free model |
|
|
145
144
|
| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning — text + images + video + audio (ChartQA 90.3, DocVQA 95.6) |
|
|
146
145
|
| `nvidia/mistral-nemotron` | 131K | Mistral × NVIDIA instruction model — fast (~0.2s), strong instruction following |
|
|
147
146
|
| `nvidia/step-3.7-flash` | 131K | StepFun Step 3.7 Flash — fast lightweight reasoning |
|
|
@@ -237,10 +236,10 @@ propagate immediately so wallet / auth issues surface fast.
|
|
|
237
236
|
|
|
238
237
|
```typescript
|
|
239
238
|
// Manually pass a fallback chain to chat() / chatCompletion()
|
|
240
|
-
const reply = await client.chat('nvidia/
|
|
241
|
-
fallbackModels: ['nvidia/
|
|
239
|
+
const reply = await client.chat('nvidia/step-3.7-flash', 'hello', {
|
|
240
|
+
fallbackModels: ['nvidia/mistral-nemotron', 'nvidia/gpt-oss-120b'],
|
|
242
241
|
});
|
|
243
|
-
// If
|
|
242
|
+
// If step-3.7-flash times out, the SDK retries against the next model
|
|
244
243
|
// and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
|
|
245
244
|
```
|
|
246
245
|
|
|
@@ -248,7 +247,7 @@ const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
|
|
|
248
247
|
|
|
249
248
|
| Profile | Strategy | Savings vs Opus 5 | Best For |
|
|
250
249
|
|---------|----------|-------------------|----------|
|
|
251
|
-
| `eco` | Cheapest capable model — ranks the <!-- br:models.free -->
|
|
250
|
+
| `eco` | Cheapest capable model — ranks the <!-- br:models.free -->5<!-- /br:models.free -->-model free NVIDIA tier first | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production, zero-cost testing |
|
|
252
251
|
| `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%** | General use |
|
|
253
252
|
| `premium` | Top-tier models (OpenAI, Anthropic) | 0% | Quality-critical tasks |
|
|
254
253
|
|
|
@@ -396,7 +395,7 @@ You only do this when your balance runs low. Three ways to get USDC into your wa
|
|
|
396
395
|
|
|
397
396
|
- **(b) Transfer existing USDC.** Send USDC you already hold to your wallet address (`client.getWalletAddress()`). On Base, send Base USDC; on Solana (`SolanaLLMClient`), send Solana SPL USDC.
|
|
398
397
|
|
|
399
|
-
- **(c) Skip funding entirely.** Use the free NVIDIA models (e.g. `nvidia/
|
|
398
|
+
- **(c) Skip funding entirely.** Use the free NVIDIA models (e.g. `nvidia/step-3.7-flash`) — every call is **$0**, no balance required.
|
|
400
399
|
|
|
401
400
|
$5 of USDC covers thousands of paid requests. Check your balance any time:
|
|
402
401
|
|
|
@@ -556,7 +555,7 @@ thinking modes. V4 Pro is the new flagship paid SKU — 1.6T MoE / 49B active,
|
|
|
556
555
|
| Model | Input Price | Output Price | Context | Notes |
|
|
557
556
|
|-------|-------------|--------------|---------|-------|
|
|
558
557
|
| `deepseek/deepseek-v4-pro` | $0.435/M | $0.87/M | 1M | V4 flagship — strongest open-weight reasoner. The 75% launch promo became the permanent list price after 2026-05-31 |
|
|
559
|
-
| `deepseek/deepseek-chat` | $0.
|
|
558
|
+
| `deepseek/deepseek-chat` | $0.14/M | $0.28/M | 1M | V4 Flash non-thinking (paid endpoint with 5MB request bodies) |
|
|
560
559
|
| `deepseek/deepseek-reasoner` | $0.20/M | $0.40/M | 1M | V4 Flash thinking (same upstream as `deepseek-chat`, thinking enabled by default) |
|
|
561
560
|
|
|
562
561
|
### xAI Grok
|
|
@@ -587,25 +586,22 @@ auto-pick them.
|
|
|
587
586
|
|
|
588
587
|
### NVIDIA (Free) + Moonshot
|
|
589
588
|
|
|
590
|
-
Free tier refreshed 2026-
|
|
591
|
-
|
|
592
|
-
`
|
|
593
|
-
|
|
594
|
-
|
|
595
|
-
|
|
596
|
-
|
|
597
|
-
|
|
598
|
-
Flash / qwen3-coder. `nvidia/qwen3-next-80b-a3b-thinking` hit NVIDIA
|
|
599
|
-
end-of-life 2026-05-21 (HTTP 410) and is auto-redirected to
|
|
600
|
-
`nvidia/llama-4-maverick`.
|
|
589
|
+
Free tier refreshed 2026-08-12. NVIDIA has retired (HTTP 410 end-of-life)
|
|
590
|
+
the entire free DeepSeek family — `nvidia/deepseek-v4-flash` was the last
|
|
591
|
+
to go — along with `llama-4-maverick`, the qwen3 SKUs, and the free
|
|
592
|
+
Mistral small/large SKUs. Retired IDs stay callable: the gateway
|
|
593
|
+
auto-redirects them to a healthy free model, so pinned callers still get
|
|
594
|
+
a 200. `nvidia/gpt-oss-120b` and `nvidia/gpt-oss-20b` remain callable by
|
|
595
|
+
direct ID but are hidden from `/v1/models` over the NVIDIA free tier's
|
|
596
|
+
prompt-retention terms (so SmartChat won't auto-pick them).
|
|
601
597
|
|
|
602
598
|
| Model | Input Price | Output Price | Notes |
|
|
603
599
|
|-------|-------------|--------------|-------|
|
|
604
|
-
| `nvidia/
|
|
600
|
+
| `nvidia/step-3.7-flash` | **FREE** | **FREE** | Fast general-purpose chat + reasoning, 131K |
|
|
601
|
+
| `nvidia/mistral-nemotron` | **FREE** | **FREE** | Fast free Mistral, 131K |
|
|
605
602
|
| `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 31B / 3.2B active MoE, 256K — only vision-capable free model |
|
|
606
|
-
| `nvidia/
|
|
607
|
-
| `nvidia/
|
|
608
|
-
| `nvidia/qwen3-coder-480b` | **FREE** | **FREE** | Coding-optimised 480B MoE |
|
|
603
|
+
| `nvidia/nemotron-nano-9b-v2` | **FREE** | **FREE** | Compact fast chat, 131K |
|
|
604
|
+
| `nvidia/nemotron-nano-12b-v2-vl` | **FREE** | **FREE** | Compact vision, 131K |
|
|
609
605
|
| `nvidia/gpt-oss-120b` | **FREE** | **FREE** | Hidden from `/v1/models` for privacy but direct calls still work — 123 tok/s |
|
|
610
606
|
| `nvidia/gpt-oss-20b` | **FREE** | **FREE** | Hidden from `/v1/models` but direct calls still work — 155 tok/s |
|
|
611
607
|
| `moonshot/kimi-k2.5` | $0.60/M | $3.00/M | Direct from Moonshot — replaces `nvidia/kimi-k2.5` |
|
|
@@ -1061,14 +1057,14 @@ const auto = await client.smartChat('Code review', { routingProfile: 'auto' });
|
|
|
1061
1057
|
const premium = await client.smartChat('Write a legal brief', { routingProfile: 'premium' }); // Best quality
|
|
1062
1058
|
|
|
1063
1059
|
// Guaranteed $0: call a free NVIDIA model directly
|
|
1064
|
-
const free = await client.chat('nvidia/
|
|
1060
|
+
const free = await client.chat('nvidia/step-3.7-flash', 'Hello!');
|
|
1065
1061
|
```
|
|
1066
1062
|
|
|
1067
1063
|
**Routing Profiles:**
|
|
1068
1064
|
|
|
1069
1065
|
| Profile | Description | Best For |
|
|
1070
1066
|
|---------|-------------|----------|
|
|
1071
|
-
| `eco` | Budget-optimized — ranks the <!-- br:models.free -->
|
|
1067
|
+
| `eco` | Budget-optimized — ranks the <!-- br:models.free -->5<!-- /br:models.free -->-model free NVIDIA tier first | Cost-sensitive workloads, zero-cost testing |
|
|
1072
1068
|
| `auto` | Intelligent routing (default) | General use |
|
|
1073
1069
|
| `premium` | Best quality models | Critical tasks |
|
|
1074
1070
|
|
|
@@ -1624,7 +1620,7 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
|
|
|
1624
1620
|
## Frequently Asked Questions
|
|
1625
1621
|
|
|
1626
1622
|
### What is @blockrun/llm?
|
|
1627
|
-
@blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->
|
|
1623
|
+
@blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->70<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol — no API keys, no subscriptions, no vendor lock-in.
|
|
1628
1624
|
|
|
1629
1625
|
### How does payment work?
|
|
1630
1626
|
When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
|
package/dist/index.cjs
CHANGED
package/dist/index.js
CHANGED