@blockrun/llm 3.14.1 → 3.14.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -34,7 +34,7 @@ import { LLMClient } from '@blockrun/llm';
34
34
  const client = new LLMClient();
35
35
 
36
36
  const r = await client.smartChat('Prove step by step that the sum of two odd integers is even.');
37
- console.log(r.model); // 'deepseek/deepseek-v4-pro' — the right model, not the $75/M flagship
37
+ console.log(r.model); // 'deepseek/deepseek-v4-pro' — the right model, not the frontier flagship
38
38
  console.log(r.routing.savings); // 0.96 — this exact request cost 96% less than pinning the baseline
39
39
  console.log(r.response); // the proof
40
40
  ```
@@ -352,10 +352,10 @@ anchors the portfolio's candidate pool):
352
352
 
353
353
  | Tier | Example Tasks | ECO | AUTO | PREMIUM |
354
354
  |------|---------------|-----|------|---------|
355
- | SIMPLE | "What is 2+2?", definitions | nemotron-3.5-lightning (**FREE**) | gemini-2.5-flash ($0.30/$2.50) | gemini-3.5-flash ($1.50/$9) |
356
- | MEDIUM | Code snippets, explanations | glm-5.3-flash ($0.15/$0.50) | gemini-3.5-flash ($1.50/$9) | gpt-5.3-codex ($1.75/$14.00) |
357
- | COMPLEX | Architecture, long documents | glm-5.3-flash ($0.15/$0.50) | gemini-3.1-pro ($2/$12) | claude-fable-5 ($10/$50) |
358
- | REASONING | Proofs, multi-step reasoning | deepseek-reasoner ($0.14/$0.28) | deepseek-reasoner ($0.14/$0.28) | claude-sonnet-5 ($3/$15) |
355
+ | SIMPLE | "What is 2+2?", definitions | nemotron-3.5-lightning (**FREE**) | gemini-2.5-flash | gemini-3.5-flash |
356
+ | MEDIUM | Code snippets, explanations | glm-5.3-flash | gemini-3.5-flash | gpt-5.3-codex |
357
+ | COMPLEX | Architecture, long documents | glm-5.3-flash | gemini-3.1-pro | claude-fable-5 |
358
+ | REASONING | Proofs, multi-step reasoning | deepseek-reasoner | deepseek-reasoner | claude-sonnet-5 |
359
359
 
360
360
  Since Router Core V3.5 every primary and every fallback rung is a model listed
361
361
  on `/v1/models` — nothing the router picks is withheld from the public pricing
@@ -515,10 +515,10 @@ import { BlockrunClient } from '@blockrun/llm';
515
515
 
516
516
  const br = new BlockrunClient();
517
517
 
518
- // Sync GET — Surf market price (Tier 1, $0.001)
518
+ // Sync GET — Surf market price (Tier 1)
519
519
  const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
520
520
 
521
- // Sync POST — raw on-chain SQL (Tier 3, $0.020)
521
+ // Sync POST — raw on-chain SQL (Tier 3)
522
522
  const rows = await br.post('/v1/surf/onchain/sql', {
523
523
  query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
524
524
  });
@@ -552,10 +552,10 @@ shims over `BlockrunClient`) and removed in 3.0.
552
552
 
553
553
  ## Available Models
554
554
 
555
- Prices below are the live gateway rates, regenerated from `GET /v1/models`
556
- (the same catalog `client.listModels()` returns). Chat models are billed per
557
- token; image, video, music and speech are billed per unit as noted in their
558
- own sections.
555
+ **Prices are not listed here.** They change often, and a number copied into a
556
+ README is wrong the day after it lands. See **[blockrun.ai/models](https://blockrun.ai/models)**
557
+ for live rates, or read them from the catalog at runtime — `client.listModels()`
558
+ and `client.listImageModels()` return exactly what the gateway is charging.
559
559
 
560
560
  ### OpenAI GPT-5.6 Family
561
561
 
@@ -563,77 +563,77 @@ Three tiers on one 1.05M-context base — Sol (deepest reasoning), Terra
563
563
  (balanced), Luna (cheap and fast). Each has a `-pro` sibling that thinks
564
564
  longer at the same token price.
565
565
 
566
- | Model | Input Price | Output Price | Context |
567
- |-------|-------------|--------------|---------|
568
- | `openai/gpt-5.6-sol` | $5.00/M | $30.00/M | 1.05M |
569
- | `openai/gpt-5.6-sol-pro` | $5.00/M | $30.00/M | 1.05M |
570
- | `openai/gpt-5.6-terra` | $2.00/M | $12.00/M | 1.05M |
571
- | `openai/gpt-5.6-terra-pro` | $2.00/M | $12.00/M | 1.05M |
572
- | `openai/gpt-5.6-luna` | $0.20/M | $1.20/M | 1.05M |
573
- | `openai/gpt-5.6-luna-pro` | $0.20/M | $1.20/M | 1.05M |
566
+ | Model | Context |
567
+ |---|---|
568
+ | `openai/gpt-5.6-sol` | 1.05M |
569
+ | `openai/gpt-5.6-sol-pro` | 1.05M |
570
+ | `openai/gpt-5.6-terra` | 1.05M |
571
+ | `openai/gpt-5.6-terra-pro` | 1.05M |
572
+ | `openai/gpt-5.6-luna` | 1.05M |
573
+ | `openai/gpt-5.6-luna-pro` | 1.05M |
574
574
 
575
575
  ### OpenAI GPT-5.5 / 5.4 / 5.2 Families
576
576
 
577
- | Model | Input Price | Output Price | Context | Notes |
578
- |-------|-------------|--------------|---------|-------|
579
- | `openai/gpt-5.5` | $5.00/M | $30.00/M | 1.05M | |
580
- | `openai/gpt-5.5-pro` | $30.00/M | $180.00/M | 1.05M | |
581
- | `openai/chat-latest` | $5.00/M | $30.00/M | 128K | ChatGPT Instant — the model behind chatgpt.com |
582
- | `openai/gpt-5.4` | $2.50/M | $15.00/M | 1.05M | |
583
- | `openai/gpt-5.4-pro` | $30.00/M | $180.00/M | 1.05M | |
584
- | `openai/gpt-5.4-mini` | $0.75/M | $4.50/M | 400K | |
585
- | `openai/gpt-5.4-nano` | $0.20/M | $1.25/M | 1.05M | |
586
- | `openai/gpt-5.2` | $1.75/M | $14.00/M | 400K | |
587
- | `openai/gpt-5.2-pro` | $21.00/M | $168.00/M | 400K | |
588
- | `openai/gpt-5.3-codex` | $1.75/M | $14.00/M | 400K | Coding/agentic SKU |
589
- | `openai/gpt-5-mini` | $0.25/M | $2.00/M | 200K | |
577
+ | Model | Context | Notes |
578
+ |---|---|---|
579
+ | `openai/gpt-5.5` | 1.05M | |
580
+ | `openai/gpt-5.5-pro` | 1.05M | |
581
+ | `openai/chat-latest` | 128K | ChatGPT Instant — the model behind chatgpt.com |
582
+ | `openai/gpt-5.4` | 1.05M | |
583
+ | `openai/gpt-5.4-pro` | 1.05M | |
584
+ | `openai/gpt-5.4-mini` | 400K | |
585
+ | `openai/gpt-5.4-nano` | 1.05M | |
586
+ | `openai/gpt-5.2` | 400K | |
587
+ | `openai/gpt-5.2-pro` | 400K | |
588
+ | `openai/gpt-5.3-codex` | 400K | Coding/agentic SKU |
589
+ | `openai/gpt-5-mini` | 200K | |
590
590
 
591
591
  ### OpenAI GPT-4 Family
592
592
 
593
- | Model | Input Price | Output Price | Context |
594
- |-------|-------------|--------------|---------|
595
- | `openai/gpt-4.1` | $2.00/M | $8.00/M | 128K |
596
- | `openai/gpt-4.1-mini` | $0.40/M | $1.60/M | 128K |
597
- | `openai/gpt-4.1-nano` | $0.10/M | $0.40/M | 128K |
598
- | `openai/gpt-4o` | $2.50/M | $10.00/M | 128K |
599
- | `openai/gpt-4o-mini` | $0.15/M | $0.60/M | 128K |
593
+ | Model | Context |
594
+ |---|---|
595
+ | `openai/gpt-4.1` | 128K |
596
+ | `openai/gpt-4.1-mini` | 128K |
597
+ | `openai/gpt-4.1-nano` | 128K |
598
+ | `openai/gpt-4o` | 128K |
599
+ | `openai/gpt-4o-mini` | 128K |
600
600
 
601
601
  ### OpenAI O-Series (Reasoning)
602
602
 
603
- | Model | Input Price | Output Price | Context |
604
- |-------|-------------|--------------|---------|
605
- | `openai/o1` | $15.00/M | $60.00/M | 200K |
606
- | `openai/o3` | $2.00/M | $8.00/M | 200K |
607
- | `openai/o3-mini` | $1.10/M | $4.40/M | 128K |
608
- | `openai/o4-mini` | $1.10/M | $4.40/M | 128K |
603
+ | Model | Context |
604
+ |---|---|
605
+ | `openai/o1` | 200K |
606
+ | `openai/o3` | 200K |
607
+ | `openai/o3-mini` | 128K |
608
+ | `openai/o4-mini` | 128K |
609
609
 
610
610
  ### Anthropic Claude
611
611
 
612
- | Model | Input Price | Output Price | Context | Notes |
613
- |-------|-------------|--------------|---------|-------|
614
- | `anthropic/claude-fable-5` | $10.00/M | $50.00/M | 1M | Mythos-class flagship above Opus — always-on thinking, 128K output |
615
- | `anthropic/claude-opus-5` | $5.00/M | $25.00/M | 1M | Flagship — the baseline the routing savings claim is measured against |
616
- | `anthropic/claude-opus-4.8` | $5.00/M | $25.00/M | 1M | Agentic coding + adaptive thinking, 128K output |
617
- | `anthropic/claude-opus-4.7` | $5.00/M | $25.00/M | 1M | |
618
- | `anthropic/claude-opus-4.5` | $5.00/M | $25.00/M | 200K | |
619
- | `anthropic/claude-sonnet-5` | $3.00/M | $15.00/M | 1M | Best cost/quality balance for long-context agent turns |
620
- | `anthropic/claude-sonnet-4.6` | $3.00/M | $15.00/M | 1M | |
621
- | `anthropic/claude-sonnet-4.5` | $3.00/M | $15.00/M | 200K | |
622
- | `anthropic/claude-haiku-4.5` | $1.00/M | $5.00/M | 200K | |
612
+ | Model | Context | Notes |
613
+ |---|---|---|
614
+ | `anthropic/claude-fable-5` | 1M | Mythos-class flagship above Opus — always-on thinking, 128K output |
615
+ | `anthropic/claude-opus-5` | 1M | Flagship — the baseline the routing savings claim is measured against |
616
+ | `anthropic/claude-opus-4.8` | 1M | Agentic coding + adaptive thinking, 128K output |
617
+ | `anthropic/claude-opus-4.7` | 1M | |
618
+ | `anthropic/claude-opus-4.5` | 200K | |
619
+ | `anthropic/claude-sonnet-5` | 1M | Best cost/quality balance for long-context agent turns |
620
+ | `anthropic/claude-sonnet-4.6` | 1M | |
621
+ | `anthropic/claude-sonnet-4.5` | 200K | |
622
+ | `anthropic/claude-haiku-4.5` | 200K | |
623
623
 
624
624
  ### Google Gemini
625
625
 
626
- | Model | Input Price | Output Price | Context |
627
- |-------|-------------|--------------|---------|
628
- | `google/gemini-3.1-pro` | $2.00/M | $12.00/M | 1M |
629
- | `google/gemini-3.6-flash` | $1.50/M | $7.50/M | 1M |
630
- | `google/gemini-3.5-flash` | $1.50/M | $9.00/M | 1M |
631
- | `google/gemini-3-flash-preview` | $0.50/M | $3.00/M | 1M |
632
- | `google/gemini-3.5-flash-lite` | $0.30/M | $2.50/M | 1M |
633
- | `google/gemini-3.1-flash-lite` | $0.25/M | $1.50/M | 1M |
634
- | `google/gemini-2.5-pro` | $1.25/M | $10.00/M | 1M |
635
- | `google/gemini-2.5-flash` | $0.30/M | $2.50/M | 1M |
636
- | `google/gemini-2.5-flash-lite` | $0.10/M | $0.40/M | 1M |
626
+ | Model | Context |
627
+ |---|---|
628
+ | `google/gemini-3.1-pro` | 1M |
629
+ | `google/gemini-3.6-flash` | 1M |
630
+ | `google/gemini-3.5-flash` | 1M |
631
+ | `google/gemini-3-flash-preview` | 1M |
632
+ | `google/gemini-3.5-flash-lite` | 1M |
633
+ | `google/gemini-3.1-flash-lite` | 1M |
634
+ | `google/gemini-2.5-pro` | 1M |
635
+ | `google/gemini-2.5-flash` | 1M |
636
+ | `google/gemini-2.5-flash-lite` | 1M |
637
637
 
638
638
  ### DeepSeek
639
639
 
@@ -641,12 +641,12 @@ DeepSeek upstream serves the legacy `deepseek-chat` / `deepseek-reasoner`
641
641
  aliases as V4 Flash non-thinking / thinking modes. V4 Pro is the flagship
642
642
  paid SKU; the vision SKU is an experimental preview.
643
643
 
644
- | Model | Input Price | Output Price | Context | Notes |
645
- |-------|-------------|--------------|---------|-------|
646
- | `deepseek/deepseek-v4-pro` | $1.32/M | $3.96/M | 1M | V4 flagship — strongest open-weight reasoner |
647
- | `deepseek/deepseek-v4-flash-vision-exp` | $0.44/M | $1.32/M | 1M | Experimental vision preview |
648
- | `deepseek/deepseek-chat` | $0.14/M | $0.28/M | 1M | V4 Flash non-thinking |
649
- | `deepseek/deepseek-reasoner` | $0.14/M | $0.28/M | 1M | V4 Flash thinking (same upstream, thinking on by default) |
644
+ | Model | Context | Notes |
645
+ |---|---|---|
646
+ | `deepseek/deepseek-v4-pro` | 1M | V4 flagship — strongest open-weight reasoner |
647
+ | `deepseek/deepseek-v4-flash-vision-exp` | 1M | Experimental vision preview |
648
+ | `deepseek/deepseek-chat` | 1M | V4 Flash non-thinking |
649
+ | `deepseek/deepseek-reasoner` | 1M | V4 Flash thinking (same upstream, thinking on by default) |
650
650
 
651
651
  ### xAI Grok
652
652
 
@@ -655,37 +655,37 @@ grok-code-fast-1, grok-2-vision) have left the catalog. Retired ids stay
655
655
  callable — the gateway redirects them to a healthy model — but SmartChat
656
656
  only ranks what `/v1/models` lists.
657
657
 
658
- | Model | Input Price | Output Price | Context | Notes |
659
- |-------|-------------|--------------|---------|-------|
660
- | `xai/grok-4.5` | $2.00/M | $6.00/M | 500K | Flagship — reasoning + vision, native Live Search (`search: true`) |
661
- | `xai/grok-4.3` | $1.25/M | $2.50/M | 1M | Reasoning + vision, tuned for agentic workflows |
662
- | `xai/grok-build-0.1` | $1.00/M | $2.00/M | 256K | Fast agentic coding model |
658
+ | Model | Context | Notes |
659
+ |---|---|---|
660
+ | `xai/grok-4.5` | 500K | Flagship — reasoning + vision, native Live Search (`search: true`) |
661
+ | `xai/grok-4.3` | 1M | Reasoning + vision, tuned for agentic workflows |
662
+ | `xai/grok-build-0.1` | 256K | Fast agentic coding model |
663
663
 
664
664
  ### Moonshot, MiniMax, Z.ai, Qwen
665
665
 
666
- | Model | Input Price | Output Price | Context | Notes |
667
- |-------|-------------|--------------|---------|-------|
668
- | `moonshot/kimi-k3` | $3.00/M | $15.00/M | 1M | Replaces the retired `kimi-k2.5` / `k2.6` SKUs |
669
- | `minimax/minimax-m3` | $0.30/M | $1.20/M | 1M | |
670
- | `minimax/minimax-m2.7` | $0.30/M | $1.20/M | 200K | |
671
- | `zai/glm-5.3` | $1.40/M | $4.40/M | 1M | |
672
- | `zai/glm-5.3-flash` | $0.15/M | $0.50/M | 1M | Cheapest vision-capable paid SKU |
673
- | `zai/glm-5.2` | $1.40/M | $4.40/M | 1M | |
674
- | `zai/glm-5.1` | $1.40/M | $4.40/M | 200K | |
675
- | `zai/glm-5` | $1.00/M | $3.20/M | 200K | |
676
- | `zai/glm-5-turbo` | $1.20/M | $4.00/M | 200K | |
677
- | `qwen/qwen3.7-max` | $1.475/M | $4.425/M | 1M | |
678
- | `qwen/qwen3.7-plus` | $0.32/M | $1.28/M | 1M | |
679
- | `qwen/qwen3.8-flash` | $0.15/M | $0.47/M | 1M | |
680
- | `qwen/qwen3.7-flash` | $0.03/M | $0.13/M | 1M | Cheapest paid chat model in the catalog |
666
+ | Model | Context | Notes |
667
+ |---|---|---|
668
+ | `moonshot/kimi-k3` | 1M | Replaces the retired `kimi-k2.5` / `k2.6` SKUs |
669
+ | `minimax/minimax-m3` | 1M | |
670
+ | `minimax/minimax-m2.7` | 200K | |
671
+ | `zai/glm-5.3` | 1M | |
672
+ | `zai/glm-5.3-flash` | 1M | Cheapest vision-capable paid SKU |
673
+ | `zai/glm-5.2` | 1M | |
674
+ | `zai/glm-5.1` | 200K | |
675
+ | `zai/glm-5` | 200K | |
676
+ | `zai/glm-5-turbo` | 200K | |
677
+ | `qwen/qwen3.7-max` | 1M | |
678
+ | `qwen/qwen3.7-plus` | 1M | |
679
+ | `qwen/qwen3.8-flash` | 1M | |
680
+ | `qwen/qwen3.7-flash` | 1M | Cheapest paid chat model in the catalog |
681
681
 
682
682
  ### Tencent, Xiaomi
683
683
 
684
- | Model | Input Price | Output Price | Context |
685
- |-------|-------------|--------------|---------|
686
- | `tencent/hy3` | $0.132/M | $0.528/M | 256K |
687
- | `xiaomi/mimo-v2.5` | $0.14/M | $0.28/M | 1M |
688
- | `xiaomi/mimo-v2.5-pro` | $0.435/M | $0.87/M | 1M |
684
+ | Model | Context |
685
+ |---|---|
686
+ | `tencent/hy3` | 256K |
687
+ | `xiaomi/mimo-v2.5` | 1M |
688
+ | `xiaomi/mimo-v2.5-pro` | 1M |
689
689
 
690
690
  ### Free Tier
691
691
 
@@ -693,28 +693,28 @@ Input and output both $0 — no promo, no rate-limit gimmick. The free tier is
693
693
  **no longer NVIDIA-only**, so pin these by full model id rather than by an
694
694
  `nvidia/*` prefix, or let `routingProfile: 'eco'` rank them first.
695
695
 
696
- | Model | Input Price | Output Price | Context | Notes |
697
- |-------|-------------|--------------|---------|-------|
698
- | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 256K | Multimodal reasoning — text + images |
699
- | `nvidia/nemotron-3.5-lightning` | **FREE** | **FREE** | 1M | Thinking-mode reasoning at 1M context |
700
- | `nvidia/nemotron-3-nano-30b` | **FREE** | **FREE** | 128K | Compact and fast, good for high-volume light tasks |
701
- | `nvidia/llama-3.2-11b-vision` | **FREE** | **FREE** | 128K | Vision-language — accepts images |
702
- | `nvidia/nemotron-3-ultra-550b` | **FREE** | **FREE** | 1M | Largest free model — 550B, 1M context |
703
- | `cohere/north-mini-code` | **FREE** | **FREE** | 256K | Compact coding model, sub-second responses |
704
- | `poolside/laguna-xs-2.1` | **FREE** | **FREE** | 128K | Coding model |
696
+ | Model | Context | Notes |
697
+ |---|---|---|
698
+ | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning — text + images |
699
+ | `nvidia/nemotron-3.5-lightning` | 1M | Thinking-mode reasoning at 1M context |
700
+ | `nvidia/nemotron-3-nano-30b` | 128K | Compact and fast, good for high-volume light tasks |
701
+ | `nvidia/llama-3.2-11b-vision` | 128K | Vision-language — accepts images |
702
+ | `nvidia/nemotron-3-ultra-550b` | 1M | Largest free model — 550B, 1M context |
703
+ | `cohere/north-mini-code` | 256K | Compact coding model, sub-second responses |
704
+ | `poolside/laguna-xs-2.1` | 128K | Coding model |
705
705
 
706
706
  ### Image Generation
707
- | Model | Price | Notes |
708
- |-------|-------|-------|
709
- | `openai/gpt-image-1` | $0.02/image | Native GPT-4o image generation |
710
- | `openai/gpt-image-2` | $0.06/image | Reasoning-driven — multilingual text rendering, character consistency |
711
- | `google/nano-banana` | $0.05/image | Gemini 2.5 Flash image generation — fast and efficient |
712
- | `google/nano-banana-2` | $0.09/image | Gemini 3.1 Flash — pro-level quality at Flash speed |
713
- | `google/nano-banana-pro` | $0.10/image | Gemini 3 Pro — highest quality, up to 4K |
714
- | `xai/grok-imagine-image` | $0.02/image | Fast, 300 RPM |
715
- | `xai/grok-imagine-image-pro` | $0.07/image | Quality tier, 30 RPM |
716
- | `bytedance/seedream-5-pro` | $0.045/image | Flagship generation + editing, up to 4K-class, reference images |
717
- | `zai/cogview-4` | $0.015/image | Up to 1440x1440 |
707
+ | Model | Notes |
708
+ |---|---|
709
+ | `openai/gpt-image-1` | Native GPT-4o image generation |
710
+ | `openai/gpt-image-2` | Reasoning-driven — multilingual text rendering, character consistency |
711
+ | `google/nano-banana` | Gemini 2.5 Flash image generation — fast and efficient |
712
+ | `google/nano-banana-2` | Gemini 3.1 Flash — pro-level quality at Flash speed |
713
+ | `google/nano-banana-pro` | Gemini 3 Pro — highest quality, up to 4K |
714
+ | `xai/grok-imagine-image` | Fast, 300 RPM |
715
+ | `xai/grok-imagine-image-pro` | Quality tier, 30 RPM |
716
+ | `bytedance/seedream-5-pro` | Flagship generation + editing, up to 4K-class, reference images |
717
+ | `zai/cogview-4` | Up to 1440x1440 |
718
718
 
719
719
  Image editing (`client.edit`) via `/v1/images/image2image`: `openai/gpt-image-1`, `openai/gpt-image-2`, `google/nano-banana`, and `google/nano-banana-pro`. Pass a single base64 `data:image/...` URI to edit one image, or an array of 2–4 URIs to **fuse** them (e.g. a subject + a brand logo). Fusion caps: `openai/*` up to 4 source images, `google/*` up to 3. A `mask` cannot be combined with multiple source images.
720
720
 
@@ -729,16 +729,16 @@ console.log(fused.data[0].url);
729
729
  ```
730
730
 
731
731
  ### Video Generation
732
- | Model | Price | Default | Max | Notes |
733
- |-------|-------|---------|-----|-------|
734
- | `xai/grok-imagine-video` | $0.05/sec | 8s | 15s | 480p default, 720p at $0.07/sec; text or image to video |
735
- | `xai/grok-imagine-video-1.5` | $0.08/sec | 8s | 15s | Flagship — native synced audio; 480p default, 720p at $0.11/sec |
736
- | `bytedance/seedance-1.5-pro` | $0.07/sec | 5s | 12s | Budget 720p with synced audio. No RealFace assets |
737
- | `bytedance/seedance-2.0-fast` | $0.165/sec | 5s | 15s | 720p, ~60-80s to generate. RealFace assets supported |
738
- | `bytedance/seedance-2.0-mini` | $0.0797/sec | 5s | 15s | 480p/720p at half the flagship rate. RealFace supported |
739
- | `bytedance/seedance-2.0` | $0.227/sec | 5s | 15s | Premium 720p with synced audio. RealFace supported |
740
- | `bytedance/seedance-2.5` | $0.315/sec | 5s | 30s | Long-form — up to 30s, multilingual, multi-asset |
741
- | `azure/sora-2` | $0.10/sec | 4s | 12s | Sora 2 via Azure AI Foundry — 720p with synced audio; 4, 8 or 12s |
732
+ | Model | Default | Max | Notes |
733
+ |---|---|---|---|
734
+ | `xai/grok-imagine-video` | 8s | 15s | 480p default, 720p available; text or image to video |
735
+ | `xai/grok-imagine-video-1.5` | 8s | 15s | Flagship — native synced audio; 480p default, 720p available |
736
+ | `bytedance/seedance-1.5-pro` | 5s | 12s | Budget 720p with synced audio. No RealFace assets |
737
+ | `bytedance/seedance-2.0-fast` | 5s | 15s | 720p, ~60-80s to generate. RealFace assets supported |
738
+ | `bytedance/seedance-2.0-mini` | 5s | 15s | 480p/720p at half the flagship rate. RealFace supported |
739
+ | `bytedance/seedance-2.0` | 5s | 15s | Premium 720p with synced audio. RealFace supported |
740
+ | `bytedance/seedance-2.5` | 5s | 30s | Long-form — up to 30s, multilingual, multi-asset |
741
+ | `azure/sora-2` | 4s | 12s | Sora 2 via Azure AI Foundry — 720p with synced audio; 4, 8 or 12s |
742
742
 
743
743
  ```ts
744
744
  import { VideoClient } from '@blockrun/llm';
@@ -794,17 +794,17 @@ const r5 = await client.generate(
794
794
  `SpeechClient` wraps BlockRun Voice (ElevenLabs): `POST /v1/audio/speech`
795
795
  (OpenAI-compatible TTS), `POST /v1/audio/sound-effects`, and the free
796
796
  `GET /v1/audio/voices`. TTS price scales with character count:
797
- `(chars / 1000) × model rate`, minimum $0.001/request. Synthesis is
797
+ `(chars / 1000) × model rate`, with a per-request minimum. Synthesis is
798
798
  synchronous (<1s for Flash).
799
799
 
800
- | Model | Price | Max Input | Notes |
801
- |-------|-------|-----------|-------|
802
- | `elevenlabs/flash-v2.5` | $0.05/1k chars | 40k chars | ~75ms latency, 32 languages (default) |
803
- | `elevenlabs/turbo-v2.5` | $0.05/1k chars | 40k chars | ~250ms latency, balanced quality |
804
- | `elevenlabs/multilingual-v2` | $0.10/1k chars | 10k chars | Long-form narration, audiobooks — 29 languages |
805
- | `elevenlabs/v3` | $0.10/1k chars | 5k chars | Max expressiveness, 70+ languages |
806
- | `bytedance/seed-audio-1.0` | $0.30/1k chars | 3k chars | Prompt-directed — describe voice, emotion and staging in words |
807
- | `elevenlabs/sound-effects` | $0.05/generation | 1k chars | Sound effects up to 22s |
800
+ | Model | Max Input | Notes |
801
+ |---|---|---|
802
+ | `elevenlabs/flash-v2.5` | 40k chars | ~75ms latency, 32 languages (default) |
803
+ | `elevenlabs/turbo-v2.5` | 40k chars | ~250ms latency, balanced quality |
804
+ | `elevenlabs/multilingual-v2` | 10k chars | Long-form narration, audiobooks — 29 languages |
805
+ | `elevenlabs/v3` | 5k chars | Max expressiveness, 70+ languages |
806
+ | `bytedance/seed-audio-1.0` | 3k chars | Prompt-directed — describe voice, emotion and staging in words |
807
+ | `elevenlabs/sound-effects` | 1k chars | Sound effects up to 22s |
808
808
 
809
809
  ```ts
810
810
  import { SpeechClient } from '@blockrun/llm';
@@ -823,7 +823,7 @@ const wav = await client.generate('Breaking news from the world of micropayments
823
823
  speed: 1.1,
824
824
  });
825
825
 
826
- // Sound effects (flat $0.05/generation)
826
+ // Sound effects (flat per generation)
827
827
  const fx = await client.soundEffect('rain on a tin roof, distant thunder');
828
828
 
829
829
  // List voices (free, rate-limited)
@@ -832,7 +832,7 @@ const voices = await client.listVoices();
832
832
 
833
833
  ### Virtual Portraits
834
834
 
835
- `PortraitClient` wraps `POST /v1/portrait/enroll` (paid, flat **$0.01** promo,
835
+ `PortraitClient` wraps `POST /v1/portrait/enroll` (paid, flat promo rate,
836
836
  no KYC). Enroll a face image by URL and get back a Token360 asset id (`ta_xxxxxx`).
837
837
  Pass that id as `realFaceAssetId` on a Seedance 2.0 video generation to keep the
838
838
  same AI character across clips. Payment settles only after Token360 confirms the
@@ -861,7 +861,7 @@ console.log(clip.data[0].url);
861
861
 
862
862
  ### Voice Calls
863
863
 
864
- `VoiceClient` wraps `POST /v1/voice/call` (paid, $0.54/call) and
864
+ `VoiceClient` wraps `POST /v1/voice/call` (paid, flat per call) and
865
865
  `GET /v1/voice/call/{callId}` (free polling) — AI-powered outbound phone
866
866
  calls powered by Bland.ai. The agent dials the recipient and runs a real-time
867
867
  conversation based on your `task` instructions. US + Canada destinations.
@@ -871,7 +871,7 @@ import { VoiceClient } from '@blockrun/llm';
871
871
 
872
872
  const client = new VoiceClient();
873
873
 
874
- // Initiate (paid $0.54)
874
+ // Initiate (paid)
875
875
  const result = await client.call({
876
876
  to: '+14155552671',
877
877
  task: 'You are a friendly assistant calling to confirm a 3pm dentist appointment.',
@@ -891,7 +891,7 @@ phone number you own; buy via `/v1/phone/numbers/buy`).
891
891
  ### Standalone Search
892
892
 
893
893
  `SearchClient` wraps `POST /v1/search` — standalone Grok Live Search.
894
- Pricing: `$0.025/source + margin` (10 sources ≈ `$0.26`).
894
+ Pricing is per source plus margin — see [blockrun.ai/models](https://blockrun.ai/models).
895
895
 
896
896
  ```ts
897
897
  import { SearchClient } from '@blockrun/llm';
@@ -912,11 +912,11 @@ endpoints across CEX/DEX market data, on-chain SQL, wallet intelligence,
912
912
  prediction markets (Polymarket + Kalshi), social analytics, news, VC fund
913
913
  data, and an OpenAI-compatible chat surface. Flat pricing per call:
914
914
 
915
- | Tier | Price/call | Examples |
916
- |------|-----------|----------|
917
- | 1 | $0.001 | `/market/price`, `/market/ranking`, `/news/feed`, prediction-market reads, social tweets |
918
- | 2 | $0.005 | `/exchange/depth`, `/exchange/klines`, `/wallet/detail`, `/search/*`, `/social/ranking` |
919
- | 3 | $0.020 | `/onchain/sql`, `/onchain/query`, `/onchain/schema`, `/chat/completions` |
915
+ | Tier | Examples |
916
+ |---|---|
917
+ | 1 | `/market/price`, `/market/ranking`, `/news/feed`, prediction-market reads, social tweets |
918
+ | 2 | `/exchange/depth`, `/exchange/klines`, `/wallet/detail`, `/search/*`, `/social/ranking` |
919
+ | 3 | `/onchain/sql`, `/onchain/query`, `/onchain/schema`, `/chat/completions` |
920
920
 
921
921
  Because the catalog is broad and evolving, the client deliberately ships a
922
922
  generic `get` / `post` pair instead of 84 typed wrappers. Pass the path
@@ -928,16 +928,16 @@ import { SurfClient } from '@blockrun/llm';
928
928
 
929
929
  const surf = new SurfClient();
930
930
 
931
- // Tier 1 — token price ($0.001)
931
+ // Tier 1 — token price
932
932
  const btc = await surf.get('/market/price', { symbol: 'BTC' });
933
933
 
934
- // Tier 2 — order book depth ($0.005)
934
+ // Tier 2 — order book depth
935
935
  const book = await surf.get('/exchange/depth', {
936
936
  exchange: 'binance',
937
937
  symbol: 'BTC-USDT',
938
938
  });
939
939
 
940
- // Tier 3 — raw on-chain SQL against 80+ ClickHouse tables ($0.020)
940
+ // Tier 3 — raw on-chain SQL against 80+ ClickHouse tables
941
941
  const rows = await surf.post('/onchain/sql', {
942
942
  query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 5',
943
943
  });
@@ -957,7 +957,7 @@ Methods: `userLookup`, `userInfo`, `followers`, `following`, `followings`,
957
957
 
958
958
  `PriceClient` wraps the Pyth-backed market-data endpoints. Crypto, FX and
959
959
  commodity are fully free (price + history + list); 12 global stock markets
960
- and the `usstock` legacy alias charge `$0.001` for price + history (list is
960
+ and the `usstock` legacy alias are billed per call for price + history (list is
961
961
  always free). Pass `requireWallet: false` to construct a free-only client.
962
962
 
963
963
  ```ts
@@ -988,7 +988,7 @@ Three passthrough families live directly on `LLMClient` / `SolanaLLMClient`:
988
988
  ```ts
989
989
  const client = new LLMClient();
990
990
 
991
- // DefiLlama — protocols / TVL / yields / prices ($0.005/call, prices $0.001)
991
+ // DefiLlama — protocols / TVL / yields / prices
992
992
  const protocols = await client.defiProtocols();
993
993
  const aave = await client.defiProtocol('aave');
994
994
  const prices = await client.defiPrices(['coingecko:bitcoin', 'base:0x833589...']);
@@ -1003,7 +1003,7 @@ const gq = await client.dexGaslessQuote({ /* ... */ });
1003
1003
  const res = await client.dexGaslessSubmit({ trade: { /* signed eip712 */ } });
1004
1004
  const status = await client.dexGaslessStatus(res.tradeHash as string);
1005
1005
 
1006
- // Modal — sandboxed compute ($0.01 create CPU / $0.05 GPU, $0.001 exec)
1006
+ // Modal — sandboxed compute (create CPU / GPU, exec)
1007
1007
  const sb = await client.modalSandboxCreate({ image: 'python:3.11' });
1008
1008
  const out = await client.modalSandboxExec(sb.sandbox_id as string, ['python', '-c', 'print(42)']);
1009
1009
  await client.modalSandboxTerminate(sb.sandbox_id as string);
@@ -1017,7 +1017,7 @@ Generic escape hatches: `client.defi(path, params)`, `client.dex(path, params, b
1017
1017
  `RpcClient` wraps `POST /v1/rpc/{network}` — standard JSON-RPC 2.0 access to
1018
1018
  <!-- br:chains.rpc -->40<!-- /br:chains.rpc --> chains through one endpoint (Ethereum, Base, Solana, Polygon, BSC,
1019
1019
  Arbitrum, Optimism, Avalanche, Bitcoin, Sui, and more; powered by Tatum's RPC
1020
- gateway). No API key, no per-chain endpoints: flat **$0.002 per call** in
1020
+ gateway). No API key, no per-chain endpoints: one flat per-call rate in
1021
1021
  USDC; a JSON-RPC batch charges per element.
1022
1022
 
1023
1023
  ```ts
@@ -1038,7 +1038,7 @@ const balance = await client.call('base', 'eth_getBalance', [
1038
1038
  const slot = await client.call('solana', 'getSlot');
1039
1039
  const tip = await client.call('bitcoin', 'getblockcount');
1040
1040
 
1041
- // Batch: one payment, per-element pricing ($0.002 x N)
1041
+ // Batch: one payment, per-element pricing (rate x N)
1042
1042
  const out = await client.batch('polygon', [
1043
1043
  { method: 'eth_blockNumber' },
1044
1044
  { method: 'eth_gasPrice' },
@@ -1058,10 +1058,10 @@ blocks/receipts, `getTransaction`, ...) are served from a method-aware
1058
1058
  gateway cache — same price, lower latency.
1059
1059
 
1060
1060
  ### Testnet Models (Base Sepolia)
1061
- | Model | Price |
1062
- |-------|-------|
1063
- | `openai/gpt-oss-20b` | $0.001/request |
1064
- | `openai/gpt-oss-120b` | $0.002/request |
1061
+ | Model |
1062
+ |-------|
1063
+ | `openai/gpt-oss-20b` |
1064
+ | `openai/gpt-oss-120b` |
1065
1065
 
1066
1066
  *Testnet models use flat pricing (no token counting) for simplicity.*
1067
1067
 
@@ -1287,40 +1287,40 @@ import { LLMClient } from '@blockrun/llm';
1287
1287
 
1288
1288
  const client = new LLMClient();
1289
1289
 
1290
- // List markets with optional filters ($0.001/request)
1290
+ // List markets with optional filters
1291
1291
  const markets = await client.pm("polymarket/markets");
1292
1292
  const filtered = await client.pm("polymarket/markets", { status: "active", limit: 10 });
1293
1293
  const searched = await client.pm("polymarket/markets", { search: "bitcoin" });
1294
1294
 
1295
- // List events ($0.001/request)
1295
+ // List events
1296
1296
  const events = await client.pm("polymarket/events");
1297
1297
 
1298
- // Historical trades ($0.001/request)
1298
+ // Historical trades
1299
1299
  const trades = await client.pm("polymarket/trades");
1300
1300
 
1301
- // OHLCV candlestick data for a specific condition ($0.001/request)
1301
+ // OHLCV candlestick data for a specific condition
1302
1302
  const candles = await client.pm("polymarket/candlesticks/0x1234abcd...");
1303
1303
 
1304
- // Wallet profile ($0.005/request — tier 2)
1304
+ // Wallet profile (tier 2)
1305
1305
  const profile = await client.pm("polymarket/wallet/0xABC123...");
1306
1306
 
1307
- // Wallet P&L ($0.005/request — tier 2)
1307
+ // Wallet P&L (tier 2)
1308
1308
  const pnl = await client.pm("polymarket/wallet/pnl/0xABC123...");
1309
1309
 
1310
- // Global leaderboard ($0.001/request)
1310
+ // Global leaderboard
1311
1311
  const leaderboard = await client.pm("polymarket/leaderboard");
1312
1312
  ```
1313
1313
 
1314
1314
  ### Kalshi & Binance
1315
1315
 
1316
1316
  ```typescript
1317
- // Kalshi markets ($0.001/request)
1317
+ // Kalshi markets
1318
1318
  const kalshiMarkets = await client.pm("kalshi/markets");
1319
1319
 
1320
- // Kalshi trades ($0.001/request)
1320
+ // Kalshi trades
1321
1321
  const kalshiTrades = await client.pm("kalshi/trades");
1322
1322
 
1323
- // Binance candles for supported pairs ($0.001/request)
1323
+ // Binance candles for supported pairs
1324
1324
  const btcCandles = await client.pm("binance/candles/BTCUSDT");
1325
1325
  const ethCandles = await client.pm("binance/candles/ETHUSDT");
1326
1326
  // Also: SOLUSDT, XRPUSDT
@@ -1329,7 +1329,7 @@ const ethCandles = await client.pm("binance/candles/ETHUSDT");
1329
1329
  ### Cross-Platform
1330
1330
 
1331
1331
  ```typescript
1332
- // Cross-platform matching pairs ($0.001/request)
1332
+ // Cross-platform matching pairs
1333
1333
  const pairs = await client.pm("matching-markets/pairs");
1334
1334
  ```
1335
1335
 
@@ -1341,30 +1341,30 @@ Works on both `LLMClient` (Base) and `SolanaLLMClient`.
1341
1341
 
1342
1342
  Access [Exa](https://exa.ai)'s neural web search via x402. No API keys needed — pay-per-request. Available on **`LLMClient` (Base USDC)** and `SolanaLLMClient` (Solana USDC). Use Base as the primary path; the Solana gateway is awaiting `EXA_API_KEY` provisioning.
1343
1343
 
1344
- | Method | Description | Price |
1345
- |---|---|---|
1346
- | `exaSearch(query, options?)` | Neural/keyword web search | $0.01/request |
1347
- | `exaFindSimilar(url, options?)` | Find semantically similar pages | $0.01/request |
1348
- | `exaContents(urls, options?)` | Extract full text from URLs | $0.002/URL |
1349
- | `exaAnswer(query, options?)` | AI answer grounded in web search | $0.01/request |
1350
- | `exa(path, body)` | Generic proxy for any Exa endpoint | varies |
1344
+ | Method | Description |
1345
+ |---|---|
1346
+ | `exaSearch(query, options?)` | Neural/keyword web search |
1347
+ | `exaFindSimilar(url, options?)` | Find semantically similar pages |
1348
+ | `exaContents(urls, options?)` | Extract full text from URLs |
1349
+ | `exaAnswer(query, options?)` | AI answer grounded in web search |
1350
+ | `exa(path, body)` | Generic proxy for any Exa endpoint |
1351
1351
 
1352
1352
  ```typescript
1353
1353
  import { LLMClient } from '@blockrun/llm';
1354
1354
 
1355
1355
  const client = new LLMClient();
1356
1356
 
1357
- // Neural web search ($0.01/request)
1357
+ // Neural web search
1358
1358
  const results = await client.exaSearch("latest AI safety research", { numResults: 5 });
1359
1359
  const news = await client.exaSearch("bitcoin ETF news", { category: "news", numResults: 10 });
1360
1360
 
1361
- // Find similar pages ($0.01/request)
1361
+ // Find similar pages
1362
1362
  const similar = await client.exaFindSimilar("https://openai.com/research/gpt-4", { numResults: 5 });
1363
1363
 
1364
- // Extract content from URLs ($0.002/URL)
1364
+ // Extract content from URLs
1365
1365
  const content = await client.exaContents(["https://arxiv.org/abs/2303.08774"]);
1366
1366
 
1367
- // AI-generated answer from live web ($0.01/request)
1367
+ // AI-generated answer from live web
1368
1368
  const answer = await client.exaAnswer("What is the current state of AI safety research?");
1369
1369
 
1370
1370
  // Generic proxy for any Exa endpoint
@@ -1436,7 +1436,7 @@ npm test -- --coverage # Run with coverage report
1436
1436
  Integration tests call the production API and require:
1437
1437
  - A funded Base wallet with USDC ($1+ recommended)
1438
1438
  - `BASE_CHAIN_WALLET_KEY` environment variable set
1439
- - Estimated cost: ~$0.05 per test run
1439
+ - Integration tests make real paid calls; cost depends on the models exercised
1440
1440
 
1441
1441
  ```bash
1442
1442
  export BASE_CHAIN_WALLET_KEY=0x...
@@ -1726,7 +1726,7 @@ Router Core V3 is bundled into the SDK — the same deterministic routing engine
1726
1726
  Yes — as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).
1727
1727
 
1728
1728
  ### How much does it cost?
1729
- Pay only for what you use. Prices start at $0.0002 per request (GPT-5 Nano). There are no minimums, subscriptions, or monthly fees. $5 in USDC gets you thousands of requests.
1729
+ Pay only for what you use. There are no minimums, subscriptions, or monthly fees, and $5 in USDC gets you thousands of requests. Live per-model rates are at [blockrun.ai/models](https://blockrun.ai/models).
1730
1730
 
1731
1731
  ### Does it support both Solana and Base?
1732
1732
  Yes. Use `SolanaLLMClient` for Solana payments (recommended) and `LLMClient` for Base payments. Use `apiKey` for account billing without selecting a chain.