@blockrun/llm 3.17.1 → 3.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  ### Cut your LLM bill by <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->%. One line of TypeScript.
6
6
 
7
- The smart-routing SDK for <!-- br:models.chatVisible -->79<!-- /br:models.chatVisible --> models — every request goes to the cheapest model that can handle it,
7
+ The smart-routing SDK for <!-- br:models.chatVisible -->82<!-- /br:models.chatVisible --> models — every request goes to the cheapest model that can handle it,
8
8
  paid with an API key or per-request USDC on Solana or Base. No vendor lock-in.
9
9
 
10
10
  [![npm](https://img.shields.io/npm/v/@blockrun/llm.svg?style=flat-square)](https://www.npmjs.com/package/@blockrun/llm)
@@ -56,7 +56,7 @@ console.log(r.response); // the proof
56
56
  | | OpenAI SDK | OpenRouter | LiteLLM | **@blockrun/llm** |
57
57
  | ------------------ | -------------- | ----------------- | ---------------- | ----------------------------------------------------------------------- |
58
58
  | **Cost routing** | ✗ one vendor | Manual selection | Manual selection | **Automatic — <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% cheaper** |
59
- | **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->79<!-- /br:models.chatVisible -->, one credential** |
59
+ | **Models** | GPT only | 200+ | 100+ (BYO keys) | **<!-- br:models.chatVisible -->82<!-- /br:models.chatVisible -->, one credential** |
60
60
  | **Free tier** | ✗ | Rate-limited | ✗ | **<!-- br:models.free -->6<!-- /br:models.free --> models, no signup** |
61
61
  | **Auth** | API key | Account + API key | Your API keys | **API key *or* wallet signature** |
62
62
  | **Payment** | Card + invoice | Credit card | BYO keys | **Account credit or USDC per-request** |
@@ -210,7 +210,6 @@ longer NVIDIA-only**, so pin these by full model id rather than by an
210
210
  | `nvidia/nemotron-3.5-lightning` | 1M | Thinking-mode reasoning at 1M context |
211
211
  | `nvidia/nemotron-3-ultra-550b` | 1M | Largest free model — 550B |
212
212
  | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning — text + images |
213
- | `nvidia/nemotron-3-nano-30b` | 128K | Compact + fast, good for high-volume light tasks |
214
213
  | `nvidia/llama-3.2-11b-vision` | 128K | Vision-language — accepts images |
215
214
  | `cohere/north-mini-code` | 256K | Compact coding model, sub-second responses |
216
215
  | `poolside/laguna-xs-2.1` | 128K | Coding model |
@@ -303,7 +302,7 @@ step — see [How Payment Works](#phase-2--every-request-pays-itself-automatic-x
303
302
  ```typescript
304
303
  // Manually pass a fallback chain to chat() / chatCompletion()
305
304
  const reply = await client.chat('nvidia/nemotron-3.5-lightning', 'hello', {
306
- fallbackModels: ['nvidia/nemotron-3-nano-30b', 'cohere/north-mini-code'],
305
+ fallbackModels: ['nvidia/nemotron-3-ultra-550b', 'cohere/north-mini-code'],
307
306
  });
308
307
  // If nemotron-3.5-lightning times out, the SDK retries against the next model
309
308
  // and logs each hop to stderr: "[@blockrun/llm] <from> -> <to> (...)".
@@ -444,6 +443,64 @@ const tweet = await client.chat('xai/grok-4.5', 'What is trending on X?', { sear
444
443
  **Supported endpoint:** `https://sol.blockrun.ai/api`
445
444
  **Payment:** Solana USDC (SPL, mainnet)
446
445
 
446
+ ### Metered billing with x402 batch-settlement (opt-in)
447
+
448
+ By default every Solana call is an `exact` payment: one SPL transfer per call, priced at the call's **ceiling** (the quote for your `maxTokens`), settled on-chain before the model answers. With `batch-settlement` you lock a small deposit in a payment channel once. After that, each call carries only a signed authorization for its ceiling. The gateway serves the call, meters what it **actually** cost, and charges that, never more than the ceiling. It then redeems the charges on-chain in batches.
449
+
450
+ Batch mode runs on three optional peer dependencies:
451
+
452
+ ```bash
453
+ npm install @x402/core@~2.28.0 @x402/svm@~2.28.0 @solana/kit
454
+ ```
455
+
456
+ ```typescript
457
+ import { SolanaLLMClient } from '@blockrun/llm';
458
+
459
+ const client = new SolanaLLMClient({
460
+ privateKey: process.env.SOLANA_WALLET_KEY,
461
+ batch: {
462
+ // BlockRun's published operator public key. The SDK ships no default:
463
+ // you decide which operator may sign vouchers against your deposit.
464
+ operators: ['<BlockRun operator public key>'],
465
+ // Most USDC this client will ever lock in the channel (deposit + top-ups),
466
+ // and so the most an operator could claim. Default "$1".
467
+ maxDeposit: '$5',
468
+ },
469
+ });
470
+
471
+ // The first call opens the channel. By default the SDK deposits 5x this
472
+ // call's ceiling (never more than maxDeposit), and the call that opens or tops
473
+ // up the channel is charged its quoted price, as it would be with exact.
474
+ await client.chat('openai/gpt-4o-mini', 'gm');
475
+
476
+ // Every later call is metered: you pay for the tokens the call actually used.
477
+ const reply = await client.chat('anthropic/claude-sonnet-4.6', 'Summarize x402 in one line', {
478
+ maxTokens: 2048, // a ceiling, not a price: a short answer is not billed for 2,048 tokens
479
+ });
480
+
481
+ console.log(client.getSpending()); // { totalUsd: <actual charges>, calls: 2 }
482
+
483
+ // Done for good? Close the channel to get the unused escrow back.
484
+ // BlockRun closes it cooperatively when it can; otherwise this starts a
485
+ // payer-forced close and the escrow returns after the channel's grace period.
486
+ await client.closeBatchChannel();
487
+ ```
488
+
489
+ How it behaves:
490
+
491
+ - **Opt-in and trust-pinned.** In server-signed mode BlockRun's operator key signs the vouchers, so it could claim up to the whole unspent deposit. The SDK only enters a channel for an operator you list in `operators`, and only up to `maxDeposit`. A 402 that asks for any other operator is paid with `exact`.
492
+ - **Never worse than `exact`.** These cases all pay with `exact` instead:
493
+ - the 402 has no batch accept;
494
+ - the deposit would exceed `maxDeposit`;
495
+ - the gateway refuses batch (payer not admitted, admission paused, verifier unavailable);
496
+ - another batch call for the same wallet is in flight.
497
+
498
+ Parallel calls never queue behind the channel. A gateway refusal charges nothing, so the `exact` retry is the only charge.
499
+ - **Scope.** Batch covers non-streaming chat: `chat`, `chatCompletion`, `smartChat`, and `smartChatCompletion`. `stream()` and the image and media jobs still pay with `exact`.
500
+ - **One channel per wallet.** Every `SolanaLLMClient` for a wallet in one process shares a single channel. A second client with different `batch` options pays `exact`. On disk, one live process owns a wallet's channel file through a pid lock; other processes pay `exact` until that process exits.
501
+ - **State.** The open channel is saved to `~/.blockrun/solana-batch/<wallet>.json` (mode `0600`), so a restart reuses it. A custom `channelStore` path must also be per wallet. With `channelStore: false` the channel is kept in memory only and found again on-chain, and nothing coordinates processes, so use it for a single process per wallet. If a channel open or top-up gets no clean answer, or after `closeBatchChannel()`, the SDK forgets the saved channel and reads the real one from the chain. The channel uses your `rpcUrl` / `SOLANA_RPC_URL`.
502
+ - **Rollout.** sol.blockrun.ai lists `exact` first, and offers `batch-settlement` only once BlockRun enables it. Until then, and for any payer not yet admitted, a client with `batch` set simply keeps paying `exact`.
503
+
447
504
  ## Arc Support
448
505
 
449
506
  The same `LLMClient` pays on [Circle's Arc](https://www.arc.network) via [arc.blockrun.ai](https://arc.blockrun.ai) — point `apiUrl` at it and hold USDC on Arc in the same EVM wallet:
@@ -531,7 +588,7 @@ const summary = getCostSummary(); // across sessions (~/.blockr
531
588
  console.log(`Lifetime: $${summary.totalUsd.toFixed(2)} over ${summary.calls} calls`);
532
589
  ```
533
590
 
534
- In wallet mode, every paid request is a real on-chain USDC transfer — look up your wallet address on [Basescan](https://basescan.org) (or a Solana explorer) to verify each settlement independently.
591
+ In wallet mode, every paid request is a real on-chain USDC transfer — look up your wallet address on [Basescan](https://basescan.org) (or a Solana explorer) to verify each settlement independently. The exception is Solana [batch-settlement](#metered-billing-with-x402-batch-settlement-opt-in): there the on-chain records are the channel deposit and BlockRun's batched redemptions, and each call is a signed voucher.
535
592
 
536
593
  **Non-custodial by design: your private key never leaves your machine** — it is only used for local signing, and no funds are ever held by BlockRun.
537
594
 
@@ -589,6 +646,18 @@ README is wrong the day after it lands. See **[blockrun.ai/models](https://block
589
646
  for live rates, or read them from the catalog at runtime — `client.listModels()`
590
647
  and `client.listImageModels()` return exactly what the gateway is charging.
591
648
 
649
+ ### OpenAI GPT-6 Family
650
+
651
+ The GPT-6 generation: Astra is the flagship for long-horizon agentic work and
652
+ computer use, Sol the cost-efficient tier for complex coding, Luna the fast
653
+ low-cost tier.
654
+
655
+ | Model | Context |
656
+ |---|---|
657
+ | `openai/gpt-6-astra` | 1.05M |
658
+ | `openai/gpt-6-sol` | 1.05M |
659
+ | `openai/gpt-6-luna` | 1.05M |
660
+
592
661
  ### OpenAI GPT-5.6 Family
593
662
 
594
663
  Three tiers on one 1.05M-context base — Sol (deepest reasoning), Terra
@@ -604,7 +673,7 @@ longer at the same token price.
604
673
  | `openai/gpt-5.6-luna` | 1.05M |
605
674
  | `openai/gpt-5.6-luna-pro` | 1.05M |
606
675
 
607
- ### OpenAI GPT-5.5 / 5.4 / 5.2 Families
676
+ ### OpenAI GPT-5.5 / 5.4 / 5.2 / 5.1 Families
608
677
 
609
678
  | Model | Context | Notes |
610
679
  |---|---|---|
@@ -617,6 +686,7 @@ longer at the same token price.
617
686
  | `openai/gpt-5.4-nano` | 1.05M | |
618
687
  | `openai/gpt-5.2` | 400K | |
619
688
  | `openai/gpt-5.2-pro` | 400K | |
689
+ | `openai/gpt-5.1` | 400K | Configurable reasoning effort |
620
690
  | `openai/gpt-5.3-codex` | 400K | Coding/agentic SKU |
621
691
  | `openai/gpt-5-mini` | 200K | |
622
692
 
@@ -643,11 +713,14 @@ longer at the same token price.
643
713
 
644
714
  | Model | Context | Notes |
645
715
  |---|---|---|
716
+ | `anthropic/claude-fable-5.1` | 1M | Most capable — successor to Fable 5 at the same tier and price |
646
717
  | `anthropic/claude-fable-5` | 1M | Mythos-class flagship above Opus — always-on thinking, 128K output |
718
+ | `anthropic/claude-opus-5.5` | 1M | Newest Opus — Opus-class reasoning at a lower price than Opus 5 |
647
719
  | `anthropic/claude-opus-5` | 1M | Flagship — the baseline the routing savings claim is measured against |
648
720
  | `anthropic/claude-opus-4.8` | 1M | Agentic coding + adaptive thinking, 128K output |
649
721
  | `anthropic/claude-opus-4.7` | 1M | |
650
722
  | `anthropic/claude-opus-4.5` | 200K | |
723
+ | `anthropic/claude-sonnet-5.5` | 1M | Newest Sonnet — everyday coding and agent work, 128K output |
651
724
  | `anthropic/claude-sonnet-5` | 1M | Best cost/quality balance for long-context agent turns |
652
725
  | `anthropic/claude-sonnet-4.6` | 1M | |
653
726
  | `anthropic/claude-sonnet-4.5` | 200K | |
@@ -658,6 +731,7 @@ longer at the same token price.
658
731
  | Model | Context |
659
732
  |---|---|
660
733
  | `google/gemini-3.1-pro` | 1M |
734
+ | `google/gemini-3.8-flash` | 1M |
661
735
  | `google/gemini-3.6-flash` | 1M |
662
736
  | `google/gemini-3.5-flash` | 1M |
663
737
  | `google/gemini-3-flash-preview` | 1M |
@@ -689,7 +763,9 @@ only ranks what `/v1/models` lists.
689
763
 
690
764
  | Model | Context | Notes |
691
765
  |---|---|---|
692
- | `xai/grok-4.5` | 500K | Flagship — reasoning + vision, native Live Search (`search: true`) |
766
+ | `xai/grok-4.7` | 500K | Flagship — reasoning + vision, selectable effort (low → xhigh), native Live Search (`search: true`) |
767
+ | `xai/grok-4.6` | 500K | Reasoning with selectable effort, native Live Search |
768
+ | `xai/grok-4.5` | 500K | Reasoning + vision, native Live Search |
693
769
  | `xai/grok-4.3` | 1M | Reasoning + vision, tuned for agentic workflows |
694
770
  | `xai/grok-build-0.1` | 256K | Fast agentic coding model |
695
771
 
@@ -711,11 +787,10 @@ only ranks what `/v1/models` lists.
711
787
  | `qwen/qwen3.8-flash` | 1M | |
712
788
  | `qwen/qwen3.7-flash` | 1M | Cheapest paid chat model in the catalog |
713
789
 
714
- ### Tencent, Xiaomi
790
+ ### Xiaomi
715
791
 
716
792
  | Model | Context |
717
793
  |---|---|
718
- | `tencent/hy3` | 256K |
719
794
  | `xiaomi/mimo-v2.5` | 1M |
720
795
  | `xiaomi/mimo-v2.5-pro` | 1M |
721
796
 
@@ -729,7 +804,6 @@ Input and output both $0 — no promo, no rate-limit gimmick. The free tier is
729
804
  |---|---|---|
730
805
  | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning — text + images |
731
806
  | `nvidia/nemotron-3.5-lightning` | 1M | Thinking-mode reasoning at 1M context |
732
- | `nvidia/nemotron-3-nano-30b` | 128K | Compact and fast, good for high-volume light tasks |
733
807
  | `nvidia/llama-3.2-11b-vision` | 128K | Vision-language — accepts images |
734
808
  | `nvidia/nemotron-3-ultra-550b` | 1M | Largest free model — 550B, 1M context |
735
809
  | `cohere/north-mini-code` | 256K | Compact coding model, sub-second responses |
@@ -740,11 +814,14 @@ Input and output both $0 — no promo, no rate-limit gimmick. The free tier is
740
814
  |---|---|
741
815
  | `openai/gpt-image-1` | Native GPT-4o image generation |
742
816
  | `openai/gpt-image-2` | Reasoning-driven — multilingual text rendering, character consistency |
817
+ | `openai/gpt-image-2.5-flare` | GPT Image 2.5 |
818
+ | `openai/gpt-image-2.5-sunburst` | GPT Image 2.5 |
743
819
  | `google/nano-banana` | Gemini 2.5 Flash image generation — fast and efficient |
744
820
  | `google/nano-banana-2` | Gemini 3.1 Flash — pro-level quality at Flash speed |
745
821
  | `google/nano-banana-pro` | Gemini 3 Pro — highest quality, up to 4K |
746
822
  | `xai/grok-imagine-image` | Fast, 300 RPM |
747
823
  | `xai/grok-imagine-image-pro` | Quality tier, 30 RPM |
824
+ | `xai/grok-imagine-image-2.0` | Grok Imagine 2.0 |
748
825
  | `bytedance/seedream-5-pro` | Flagship generation + editing, up to 4K-class, reference images |
749
826
  | `zai/cogview-4` | Up to 1440x1440 |
750
827
 
@@ -805,11 +882,40 @@ const r4 = await client.generate('the flower blooms in golden morning light', {
805
882
  lastFrameUrl: 'https://example.com/bloom.jpg',
806
883
  });
807
884
 
808
- // Omni / multi-reference (Seedance 2.0 only): up to 9 reference images
809
- // for character/style consistency. Cite them as "image 1", "image 2" in
810
- // the prompt. Mutually exclusive with imageUrl / lastFrameUrl /
811
- // realFaceAssetId.
812
- const r5 = await client.generate(
885
+ // Seedance output controls. Each is model-gated at the gateway: an
886
+ // unsupported one is a 400 before payment, never silently dropped.
887
+ const r5 = await client.generate('a paper boat drifting down a rain gutter', {
888
+ model: 'bytedance/seedance-2.5',
889
+ bitrateMode: 'high', // Seedance 2.x
890
+ outputFormat: 'mov', // Seedance 2.5 only
891
+ returnLastFrame: true,
892
+ });
893
+ console.log(r5.data[0].last_frame_url); // present when the upstream returns it
894
+ // cameraFixed: true is Seedance 1.5-pro only.
895
+ ```
896
+
897
+ #### Reference images, video and audio (account API key only)
898
+
899
+ Seedance reference media is served by `api.blockrun.ai`, so it needs an
900
+ account API key. The wallet gateways (blockrun.ai, sol.blockrun.ai) refuse
901
+ `referenceImageUrls` / `referenceVideos` / `referenceAudios` with a 400 before
902
+ any payment.
903
+
904
+ | Model | Reference images | Reference video / audio |
905
+ |---|---|---|
906
+ | `bytedance/seedance-2.0` / `-fast` / `-mini` | 1–9 | 1–3 clips of each; audio needs an image or video alongside |
907
+ | `bytedance/seedance-2.5` | 1–30 | — |
908
+
909
+ Reference mode is its own mode: it cannot be mixed with `imageUrl`,
910
+ `lastFrameUrl` or `realFaceAssetId`. Cite images as "image 1", "image 2" and
911
+ clips as "video 1" in the prompt. Reference clips add a per-clip surcharge;
912
+ reference audio must be ≤15.2s.
913
+
914
+ ```ts
915
+ const client = new VideoClient({ apiKey: process.env.BLOCKRUN_API_KEY });
916
+
917
+ // Omni / multi-reference: character/style consistency from images
918
+ const r6 = await client.generate(
813
919
  'the character from image 1 walks through the city from image 2',
814
920
  {
815
921
  model: 'bytedance/seedance-2.0',
@@ -819,6 +925,18 @@ const r5 = await client.generate(
819
925
  ],
820
926
  }
821
927
  );
928
+
929
+ // Reference-to-video: image 1 for the character, video 1 for the motion
930
+ const r7 = await client.generate(
931
+ 'use image 1 for the character and video 1 for the motion',
932
+ {
933
+ model: 'bytedance/seedance-2.0-fast',
934
+ durationSeconds: 5,
935
+ referenceImageUrls: ['https://example.com/character.png'],
936
+ referenceVideos: [{ url: 'https://example.com/motion.mp4' }],
937
+ inputType: 'reference', // optional: 400 if the fields say otherwise
938
+ }
939
+ );
822
940
  ```
823
941
 
824
942
  ### Text-to-Speech & Sound Effects
@@ -1774,7 +1892,7 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
1774
1892
  ## Frequently Asked Questions
1775
1893
 
1776
1894
  ### What is @blockrun/llm?
1777
- @blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->79<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol — with API key account billing or x402 wallet payments on Solana or Base.
1895
+ @blockrun/llm is a TypeScript SDK that cuts LLM costs by up to <!-- br:savings.autoVsBaselinePct -->84<!-- /br:savings.autoVsBaselinePct -->% with built-in smart routing: every request is routed to the cheapest of <!-- br:models.chatVisible -->82<!-- /br:models.chatVisible --> models (OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot, and more) that can handle it, then paid per-request in USDC via the x402 protocol — with API key account billing or x402 wallet payments on Solana or Base.
1778
1896
 
1779
1897
  ### How does payment work?
1780
1898
  When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.