@blockrun/llm 3.14.1 → 3.14.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -34,7 +34,7 @@ import { LLMClient } from '@blockrun/llm';
34
34
  const client = new LLMClient();
35
35
 
36
36
  const r = await client.smartChat('Prove step by step that the sum of two odd integers is even.');
37
- console.log(r.model); // 'deepseek/deepseek-v4-pro' — the right model, not the $75/M flagship
37
+ console.log(r.model); // 'deepseek/deepseek-v4-pro' — the right model, not the frontier flagship
38
38
  console.log(r.routing.savings); // 0.96 — this exact request cost 96% less than pinning the baseline
39
39
  console.log(r.response); // the proof
40
40
  ```
@@ -352,10 +352,10 @@ anchors the portfolio's candidate pool):
352
352
 
353
353
  | Tier | Example Tasks | ECO | AUTO | PREMIUM |
354
354
  |------|---------------|-----|------|---------|
355
- | SIMPLE | "What is 2+2?", definitions | nemotron-3.5-lightning (**FREE**) | gemini-2.5-flash ($0.30/$2.50) | gemini-3.5-flash ($1.50/$9) |
356
- | MEDIUM | Code snippets, explanations | glm-5.3-flash ($0.15/$0.50) | gemini-3.5-flash ($1.50/$9) | gpt-5.3-codex ($1.75/$14.00) |
357
- | COMPLEX | Architecture, long documents | glm-5.3-flash ($0.15/$0.50) | gemini-3.1-pro ($2/$12) | claude-fable-5 ($10/$50) |
358
- | REASONING | Proofs, multi-step reasoning | deepseek-reasoner ($0.14/$0.28) | deepseek-reasoner ($0.14/$0.28) | claude-sonnet-5 ($3/$15) |
355
+ | SIMPLE | "What is 2+2?", definitions | nemotron-3.5-lightning (**FREE**) | gemini-2.5-flash | gemini-3.5-flash |
356
+ | MEDIUM | Code snippets, explanations | glm-5.3-flash | gemini-3.5-flash | gpt-5.3-codex |
357
+ | COMPLEX | Architecture, long documents | glm-5.3-flash | gemini-3.1-pro | claude-fable-5 |
358
+ | REASONING | Proofs, multi-step reasoning | deepseek-reasoner | deepseek-reasoner | claude-sonnet-5 |
359
359
 
360
360
  Since Router Core V3.5 every primary and every fallback rung is a model listed
361
361
  on `/v1/models` — nothing the router picks is withheld from the public pricing
@@ -515,10 +515,10 @@ import { BlockrunClient } from '@blockrun/llm';
515
515
 
516
516
  const br = new BlockrunClient();
517
517
 
518
- // Sync GET — Surf market price (Tier 1, $0.001)
518
+ // Sync GET — Surf market price (Tier 1)
519
519
  const btc = await br.get('/v1/surf/market/price', { symbol: 'BTC' });
520
520
 
521
- // Sync POST — raw on-chain SQL (Tier 3, $0.020)
521
+ // Sync POST — raw on-chain SQL (Tier 3)
522
522
  const rows = await br.post('/v1/surf/onchain/sql', {
523
523
  query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 1',
524
524
  });
@@ -552,10 +552,10 @@ shims over `BlockrunClient`) and removed in 3.0.
552
552
 
553
553
  ## Available Models
554
554
 
555
- Prices below are the live gateway rates, regenerated from `GET /v1/models`
556
- (the same catalog `client.listModels()` returns). Chat models are billed per
557
- token; image, video, music and speech are billed per unit as noted in their
558
- own sections.
555
+ **Prices are not listed here.** They change often, and a number copied into a
556
+ README is wrong the day after it lands. See **[blockrun.ai/models](https://blockrun.ai/models)**
557
+ for live rates, or read them from the catalog at runtime — `client.listModels()`
558
+ and `client.listImageModels()` return exactly what the gateway is charging.
559
559
 
560
560
  ### OpenAI GPT-5.6 Family
561
561
 
@@ -563,77 +563,77 @@ Three tiers on one 1.05M-context base — Sol (deepest reasoning), Terra
563
563
  (balanced), Luna (cheap and fast). Each has a `-pro` sibling that thinks
564
564
  longer at the same token price.
565
565
 
566
- | Model | Input Price | Output Price | Context |
567
- |-------|-------------|--------------|---------|
568
- | `openai/gpt-5.6-sol` | $5.00/M | $30.00/M | 1.05M |
569
- | `openai/gpt-5.6-sol-pro` | $5.00/M | $30.00/M | 1.05M |
570
- | `openai/gpt-5.6-terra` | $2.00/M | $12.00/M | 1.05M |
571
- | `openai/gpt-5.6-terra-pro` | $2.00/M | $12.00/M | 1.05M |
572
- | `openai/gpt-5.6-luna` | $0.20/M | $1.20/M | 1.05M |
573
- | `openai/gpt-5.6-luna-pro` | $0.20/M | $1.20/M | 1.05M |
566
+ | Model | Context |
567
+ |---|---|
568
+ | `openai/gpt-5.6-sol` | 1.05M |
569
+ | `openai/gpt-5.6-sol-pro` | 1.05M |
570
+ | `openai/gpt-5.6-terra` | 1.05M |
571
+ | `openai/gpt-5.6-terra-pro` | 1.05M |
572
+ | `openai/gpt-5.6-luna` | 1.05M |
573
+ | `openai/gpt-5.6-luna-pro` | 1.05M |
574
574
 
575
575
  ### OpenAI GPT-5.5 / 5.4 / 5.2 Families
576
576
 
577
- | Model | Input Price | Output Price | Context | Notes |
578
- |-------|-------------|--------------|---------|-------|
579
- | `openai/gpt-5.5` | $5.00/M | $30.00/M | 1.05M | |
580
- | `openai/gpt-5.5-pro` | $30.00/M | $180.00/M | 1.05M | |
581
- | `openai/chat-latest` | $5.00/M | $30.00/M | 128K | ChatGPT Instant — the model behind chatgpt.com |
582
- | `openai/gpt-5.4` | $2.50/M | $15.00/M | 1.05M | |
583
- | `openai/gpt-5.4-pro` | $30.00/M | $180.00/M | 1.05M | |
584
- | `openai/gpt-5.4-mini` | $0.75/M | $4.50/M | 400K | |
585
- | `openai/gpt-5.4-nano` | $0.20/M | $1.25/M | 1.05M | |
586
- | `openai/gpt-5.2` | $1.75/M | $14.00/M | 400K | |
587
- | `openai/gpt-5.2-pro` | $21.00/M | $168.00/M | 400K | |
588
- | `openai/gpt-5.3-codex` | $1.75/M | $14.00/M | 400K | Coding/agentic SKU |
589
- | `openai/gpt-5-mini` | $0.25/M | $2.00/M | 200K | |
577
+ | Model | Context | Notes |
578
+ |---|---|---|
579
+ | `openai/gpt-5.5` | 1.05M | |
580
+ | `openai/gpt-5.5-pro` | 1.05M | |
581
+ | `openai/chat-latest` | 128K | ChatGPT Instant — the model behind chatgpt.com |
582
+ | `openai/gpt-5.4` | 1.05M | |
583
+ | `openai/gpt-5.4-pro` | 1.05M | |
584
+ | `openai/gpt-5.4-mini` | 400K | |
585
+ | `openai/gpt-5.4-nano` | 1.05M | |
586
+ | `openai/gpt-5.2` | 400K | |
587
+ | `openai/gpt-5.2-pro` | 400K | |
588
+ | `openai/gpt-5.3-codex` | 400K | Coding/agentic SKU |
589
+ | `openai/gpt-5-mini` | 200K | |
590
590
 
591
591
  ### OpenAI GPT-4 Family
592
592
 
593
- | Model | Input Price | Output Price | Context |
594
- |-------|-------------|--------------|---------|
595
- | `openai/gpt-4.1` | $2.00/M | $8.00/M | 128K |
596
- | `openai/gpt-4.1-mini` | $0.40/M | $1.60/M | 128K |
597
- | `openai/gpt-4.1-nano` | $0.10/M | $0.40/M | 128K |
598
- | `openai/gpt-4o` | $2.50/M | $10.00/M | 128K |
599
- | `openai/gpt-4o-mini` | $0.15/M | $0.60/M | 128K |
593
+ | Model | Context |
594
+ |---|---|
595
+ | `openai/gpt-4.1` | 128K |
596
+ | `openai/gpt-4.1-mini` | 128K |
597
+ | `openai/gpt-4.1-nano` | 128K |
598
+ | `openai/gpt-4o` | 128K |
599
+ | `openai/gpt-4o-mini` | 128K |
600
600
 
601
601
  ### OpenAI O-Series (Reasoning)
602
602
 
603
- | Model | Input Price | Output Price | Context |
604
- |-------|-------------|--------------|---------|
605
- | `openai/o1` | $15.00/M | $60.00/M | 200K |
606
- | `openai/o3` | $2.00/M | $8.00/M | 200K |
607
- | `openai/o3-mini` | $1.10/M | $4.40/M | 128K |
608
- | `openai/o4-mini` | $1.10/M | $4.40/M | 128K |
603
+ | Model | Context |
604
+ |---|---|
605
+ | `openai/o1` | 200K |
606
+ | `openai/o3` | 200K |
607
+ | `openai/o3-mini` | 128K |
608
+ | `openai/o4-mini` | 128K |
609
609
 
610
610
  ### Anthropic Claude
611
611
 
612
- | Model | Input Price | Output Price | Context | Notes |
613
- |-------|-------------|--------------|---------|-------|
614
- | `anthropic/claude-fable-5` | $10.00/M | $50.00/M | 1M | Mythos-class flagship above Opus — always-on thinking, 128K output |
615
- | `anthropic/claude-opus-5` | $5.00/M | $25.00/M | 1M | Flagship — the baseline the routing savings claim is measured against |
616
- | `anthropic/claude-opus-4.8` | $5.00/M | $25.00/M | 1M | Agentic coding + adaptive thinking, 128K output |
617
- | `anthropic/claude-opus-4.7` | $5.00/M | $25.00/M | 1M | |
618
- | `anthropic/claude-opus-4.5` | $5.00/M | $25.00/M | 200K | |
619
- | `anthropic/claude-sonnet-5` | $3.00/M | $15.00/M | 1M | Best cost/quality balance for long-context agent turns |
620
- | `anthropic/claude-sonnet-4.6` | $3.00/M | $15.00/M | 1M | |
621
- | `anthropic/claude-sonnet-4.5` | $3.00/M | $15.00/M | 200K | |
622
- | `anthropic/claude-haiku-4.5` | $1.00/M | $5.00/M | 200K | |
612
+ | Model | Context | Notes |
613
+ |---|---|---|
614
+ | `anthropic/claude-fable-5` | 1M | Mythos-class flagship above Opus — always-on thinking, 128K output |
615
+ | `anthropic/claude-opus-5` | 1M | Flagship — the baseline the routing savings claim is measured against |
616
+ | `anthropic/claude-opus-4.8` | 1M | Agentic coding + adaptive thinking, 128K output |
617
+ | `anthropic/claude-opus-4.7` | 1M | |
618
+ | `anthropic/claude-opus-4.5` | 200K | |
619
+ | `anthropic/claude-sonnet-5` | 1M | Best cost/quality balance for long-context agent turns |
620
+ | `anthropic/claude-sonnet-4.6` | 1M | |
621
+ | `anthropic/claude-sonnet-4.5` | 200K | |
622
+ | `anthropic/claude-haiku-4.5` | 200K | |
623
623
 
624
624
  ### Google Gemini
625
625
 
626
- | Model | Input Price | Output Price | Context |
627
- |-------|-------------|--------------|---------|
628
- | `google/gemini-3.1-pro` | $2.00/M | $12.00/M | 1M |
629
- | `google/gemini-3.6-flash` | $1.50/M | $7.50/M | 1M |
630
- | `google/gemini-3.5-flash` | $1.50/M | $9.00/M | 1M |
631
- | `google/gemini-3-flash-preview` | $0.50/M | $3.00/M | 1M |
632
- | `google/gemini-3.5-flash-lite` | $0.30/M | $2.50/M | 1M |
633
- | `google/gemini-3.1-flash-lite` | $0.25/M | $1.50/M | 1M |
634
- | `google/gemini-2.5-pro` | $1.25/M | $10.00/M | 1M |
635
- | `google/gemini-2.5-flash` | $0.30/M | $2.50/M | 1M |
636
- | `google/gemini-2.5-flash-lite` | $0.10/M | $0.40/M | 1M |
626
+ | Model | Context |
627
+ |---|---|
628
+ | `google/gemini-3.1-pro` | 1M |
629
+ | `google/gemini-3.6-flash` | 1M |
630
+ | `google/gemini-3.5-flash` | 1M |
631
+ | `google/gemini-3-flash-preview` | 1M |
632
+ | `google/gemini-3.5-flash-lite` | 1M |
633
+ | `google/gemini-3.1-flash-lite` | 1M |
634
+ | `google/gemini-2.5-pro` | 1M |
635
+ | `google/gemini-2.5-flash` | 1M |
636
+ | `google/gemini-2.5-flash-lite` | 1M |
637
637
 
638
638
  ### DeepSeek
639
639
 
@@ -641,12 +641,12 @@ DeepSeek upstream serves the legacy `deepseek-chat` / `deepseek-reasoner`
641
641
  aliases as V4 Flash non-thinking / thinking modes. V4 Pro is the flagship
642
642
  paid SKU; the vision SKU is an experimental preview.
643
643
 
644
- | Model | Input Price | Output Price | Context | Notes |
645
- |-------|-------------|--------------|---------|-------|
646
- | `deepseek/deepseek-v4-pro` | $1.32/M | $3.96/M | 1M | V4 flagship — strongest open-weight reasoner |
647
- | `deepseek/deepseek-v4-flash-vision-exp` | $0.44/M | $1.32/M | 1M | Experimental vision preview |
648
- | `deepseek/deepseek-chat` | $0.14/M | $0.28/M | 1M | V4 Flash non-thinking |
649
- | `deepseek/deepseek-reasoner` | $0.14/M | $0.28/M | 1M | V4 Flash thinking (same upstream, thinking on by default) |
644
+ | Model | Context | Notes |
645
+ |---|---|---|
646
+ | `deepseek/deepseek-v4-pro` | 1M | V4 flagship — strongest open-weight reasoner |
647
+ | `deepseek/deepseek-v4-flash-vision-exp` | 1M | Experimental vision preview |
648
+ | `deepseek/deepseek-chat` | 1M | V4 Flash non-thinking |
649
+ | `deepseek/deepseek-reasoner` | 1M | V4 Flash thinking (same upstream, thinking on by default) |
650
650
 
651
651
  ### xAI Grok
652
652
 
@@ -655,37 +655,37 @@ grok-code-fast-1, grok-2-vision) have left the catalog. Retired ids stay
655
655
  callable — the gateway redirects them to a healthy model — but SmartChat
656
656
  only ranks what `/v1/models` lists.
657
657
 
658
- | Model | Input Price | Output Price | Context | Notes |
659
- |-------|-------------|--------------|---------|-------|
660
- | `xai/grok-4.5` | $2.00/M | $6.00/M | 500K | Flagship — reasoning + vision, native Live Search (`search: true`) |
661
- | `xai/grok-4.3` | $1.25/M | $2.50/M | 1M | Reasoning + vision, tuned for agentic workflows |
662
- | `xai/grok-build-0.1` | $1.00/M | $2.00/M | 256K | Fast agentic coding model |
658
+ | Model | Context | Notes |
659
+ |---|---|---|
660
+ | `xai/grok-4.5` | 500K | Flagship — reasoning + vision, native Live Search (`search: true`) |
661
+ | `xai/grok-4.3` | 1M | Reasoning + vision, tuned for agentic workflows |
662
+ | `xai/grok-build-0.1` | 256K | Fast agentic coding model |
663
663
 
664
664
  ### Moonshot, MiniMax, Z.ai, Qwen
665
665
 
666
- | Model | Input Price | Output Price | Context | Notes |
667
- |-------|-------------|--------------|---------|-------|
668
- | `moonshot/kimi-k3` | $3.00/M | $15.00/M | 1M | Replaces the retired `kimi-k2.5` / `k2.6` SKUs |
669
- | `minimax/minimax-m3` | $0.30/M | $1.20/M | 1M | |
670
- | `minimax/minimax-m2.7` | $0.30/M | $1.20/M | 200K | |
671
- | `zai/glm-5.3` | $1.40/M | $4.40/M | 1M | |
672
- | `zai/glm-5.3-flash` | $0.15/M | $0.50/M | 1M | Cheapest vision-capable paid SKU |
673
- | `zai/glm-5.2` | $1.40/M | $4.40/M | 1M | |
674
- | `zai/glm-5.1` | $1.40/M | $4.40/M | 200K | |
675
- | `zai/glm-5` | $1.00/M | $3.20/M | 200K | |
676
- | `zai/glm-5-turbo` | $1.20/M | $4.00/M | 200K | |
677
- | `qwen/qwen3.7-max` | $1.475/M | $4.425/M | 1M | |
678
- | `qwen/qwen3.7-plus` | $0.32/M | $1.28/M | 1M | |
679
- | `qwen/qwen3.8-flash` | $0.15/M | $0.47/M | 1M | |
680
- | `qwen/qwen3.7-flash` | $0.03/M | $0.13/M | 1M | Cheapest paid chat model in the catalog |
666
+ | Model | Context | Notes |
667
+ |---|---|---|
668
+ | `moonshot/kimi-k3` | 1M | Replaces the retired `kimi-k2.5` / `k2.6` SKUs |
669
+ | `minimax/minimax-m3` | 1M | |
670
+ | `minimax/minimax-m2.7` | 200K | |
671
+ | `zai/glm-5.3` | 1M | |
672
+ | `zai/glm-5.3-flash` | 1M | Cheapest vision-capable paid SKU |
673
+ | `zai/glm-5.2` | 1M | |
674
+ | `zai/glm-5.1` | 200K | |
675
+ | `zai/glm-5` | 200K | |
676
+ | `zai/glm-5-turbo` | 200K | |
677
+ | `qwen/qwen3.7-max` | 1M | |
678
+ | `qwen/qwen3.7-plus` | 1M | |
679
+ | `qwen/qwen3.8-flash` | 1M | |
680
+ | `qwen/qwen3.7-flash` | 1M | Cheapest paid chat model in the catalog |
681
681
 
682
682
  ### Tencent, Xiaomi
683
683
 
684
- | Model | Input Price | Output Price | Context |
685
- |-------|-------------|--------------|---------|
686
- | `tencent/hy3` | $0.132/M | $0.528/M | 256K |
687
- | `xiaomi/mimo-v2.5` | $0.14/M | $0.28/M | 1M |
688
- | `xiaomi/mimo-v2.5-pro` | $0.435/M | $0.87/M | 1M |
684
+ | Model | Context |
685
+ |---|---|
686
+ | `tencent/hy3` | 256K |
687
+ | `xiaomi/mimo-v2.5` | 1M |
688
+ | `xiaomi/mimo-v2.5-pro` | 1M |
689
689
 
690
690
  ### Free Tier
691
691
 
@@ -693,28 +693,28 @@ Input and output both $0 — no promo, no rate-limit gimmick. The free tier is
693
693
  **no longer NVIDIA-only**, so pin these by full model id rather than by an
694
694
  `nvidia/*` prefix, or let `routingProfile: 'eco'` rank them first.
695
695
 
696
- | Model | Input Price | Output Price | Context | Notes |
697
- |-------|-------------|--------------|---------|-------|
698
- | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | **FREE** | **FREE** | 256K | Multimodal reasoning — text + images |
699
- | `nvidia/nemotron-3.5-lightning` | **FREE** | **FREE** | 1M | Thinking-mode reasoning at 1M context |
700
- | `nvidia/nemotron-3-nano-30b` | **FREE** | **FREE** | 128K | Compact and fast, good for high-volume light tasks |
701
- | `nvidia/llama-3.2-11b-vision` | **FREE** | **FREE** | 128K | Vision-language — accepts images |
702
- | `nvidia/nemotron-3-ultra-550b` | **FREE** | **FREE** | 1M | Largest free model — 550B, 1M context |
703
- | `cohere/north-mini-code` | **FREE** | **FREE** | 256K | Compact coding model, sub-second responses |
704
- | `poolside/laguna-xs-2.1` | **FREE** | **FREE** | 128K | Coding model |
696
+ | Model | Context | Notes |
697
+ |---|---|---|
698
+ | `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` | 256K | Multimodal reasoning — text + images |
699
+ | `nvidia/nemotron-3.5-lightning` | 1M | Thinking-mode reasoning at 1M context |
700
+ | `nvidia/nemotron-3-nano-30b` | 128K | Compact and fast, good for high-volume light tasks |
701
+ | `nvidia/llama-3.2-11b-vision` | 128K | Vision-language — accepts images |
702
+ | `nvidia/nemotron-3-ultra-550b` | 1M | Largest free model — 550B, 1M context |
703
+ | `cohere/north-mini-code` | 256K | Compact coding model, sub-second responses |
704
+ | `poolside/laguna-xs-2.1` | 128K | Coding model |
705
705
 
706
706
  ### Image Generation
707
- | Model | Price | Notes |
708
- |-------|-------|-------|
709
- | `openai/gpt-image-1` | $0.02/image | Native GPT-4o image generation |
710
- | `openai/gpt-image-2` | $0.06/image | Reasoning-driven — multilingual text rendering, character consistency |
711
- | `google/nano-banana` | $0.05/image | Gemini 2.5 Flash image generation — fast and efficient |
712
- | `google/nano-banana-2` | $0.09/image | Gemini 3.1 Flash — pro-level quality at Flash speed |
713
- | `google/nano-banana-pro` | $0.10/image | Gemini 3 Pro — highest quality, up to 4K |
714
- | `xai/grok-imagine-image` | $0.02/image | Fast, 300 RPM |
715
- | `xai/grok-imagine-image-pro` | $0.07/image | Quality tier, 30 RPM |
716
- | `bytedance/seedream-5-pro` | $0.045/image | Flagship generation + editing, up to 4K-class, reference images |
717
- | `zai/cogview-4` | $0.015/image | Up to 1440x1440 |
707
+ | Model | Notes |
708
+ |---|---|
709
+ | `openai/gpt-image-1` | Native GPT-4o image generation |
710
+ | `openai/gpt-image-2` | Reasoning-driven — multilingual text rendering, character consistency |
711
+ | `google/nano-banana` | Gemini 2.5 Flash image generation — fast and efficient |
712
+ | `google/nano-banana-2` | Gemini 3.1 Flash — pro-level quality at Flash speed |
713
+ | `google/nano-banana-pro` | Gemini 3 Pro — highest quality, up to 4K |
714
+ | `xai/grok-imagine-image` | Fast, 300 RPM |
715
+ | `xai/grok-imagine-image-pro` | Quality tier, 30 RPM |
716
+ | `bytedance/seedream-5-pro` | Flagship generation + editing, up to 4K-class, reference images |
717
+ | `zai/cogview-4` | Up to 1440x1440 |
718
718
 
719
719
  Image editing (`client.edit`) via `/v1/images/image2image`: `openai/gpt-image-1`, `openai/gpt-image-2`, `google/nano-banana`, and `google/nano-banana-pro`. Pass a single base64 `data:image/...` URI to edit one image, or an array of 2–4 URIs to **fuse** them (e.g. a subject + a brand logo). Fusion caps: `openai/*` up to 4 source images, `google/*` up to 3. A `mask` cannot be combined with multiple source images.
720
720
 
@@ -729,16 +729,16 @@ console.log(fused.data[0].url);
729
729
  ```
730
730
 
731
731
  ### Video Generation
732
- | Model | Price | Default | Max | Notes |
733
- |-------|-------|---------|-----|-------|
734
- | `xai/grok-imagine-video` | $0.05/sec | 8s | 15s | 480p default, 720p at $0.07/sec; text or image to video |
735
- | `xai/grok-imagine-video-1.5` | $0.08/sec | 8s | 15s | Flagship — native synced audio; 480p default, 720p at $0.11/sec |
736
- | `bytedance/seedance-1.5-pro` | $0.07/sec | 5s | 12s | Budget 720p with synced audio. No RealFace assets |
737
- | `bytedance/seedance-2.0-fast` | $0.165/sec | 5s | 15s | 720p, ~60-80s to generate. RealFace assets supported |
738
- | `bytedance/seedance-2.0-mini` | $0.0797/sec | 5s | 15s | 480p/720p at half the flagship rate. RealFace supported |
739
- | `bytedance/seedance-2.0` | $0.227/sec | 5s | 15s | Premium 720p with synced audio. RealFace supported |
740
- | `bytedance/seedance-2.5` | $0.315/sec | 5s | 30s | Long-form — up to 30s, multilingual, multi-asset |
741
- | `azure/sora-2` | $0.10/sec | 4s | 12s | Sora 2 via Azure AI Foundry — 720p with synced audio; 4, 8 or 12s |
732
+ | Model | Default | Max | Notes |
733
+ |---|---|---|---|
734
+ | `xai/grok-imagine-video` | 8s | 15s | 480p default, 720p available; text or image to video |
735
+ | `xai/grok-imagine-video-1.5` | 8s | 15s | Flagship — native synced audio; 480p default, 720p available |
736
+ | `bytedance/seedance-1.5-pro` | 5s | 12s | Budget 720p with synced audio. No RealFace assets |
737
+ | `bytedance/seedance-2.0-fast` | 5s | 15s | 720p, ~60-80s to generate. RealFace assets supported |
738
+ | `bytedance/seedance-2.0-mini` | 5s | 15s | 480p/720p at half the flagship rate. RealFace supported |
739
+ | `bytedance/seedance-2.0` | 5s | 15s | Premium 720p with synced audio. RealFace supported |
740
+ | `bytedance/seedance-2.5` | 5s | 30s | Long-form — up to 30s, multilingual, multi-asset |
741
+ | `azure/sora-2` | 4s | 12s | Sora 2 via Azure AI Foundry — 720p with synced audio; 4, 8 or 12s |
742
742
 
743
743
  ```ts
744
744
  import { VideoClient } from '@blockrun/llm';
@@ -794,17 +794,17 @@ const r5 = await client.generate(
794
794
  `SpeechClient` wraps BlockRun Voice (ElevenLabs): `POST /v1/audio/speech`
795
795
  (OpenAI-compatible TTS), `POST /v1/audio/sound-effects`, and the free
796
796
  `GET /v1/audio/voices`. TTS price scales with character count:
797
- `(chars / 1000) × model rate`, minimum $0.001/request. Synthesis is
797
+ `(chars / 1000) × model rate`, with a per-request minimum. Synthesis is
798
798
  synchronous (<1s for Flash).
799
799
 
800
- | Model | Price | Max Input | Notes |
801
- |-------|-------|-----------|-------|
802
- | `elevenlabs/flash-v2.5` | $0.05/1k chars | 40k chars | ~75ms latency, 32 languages (default) |
803
- | `elevenlabs/turbo-v2.5` | $0.05/1k chars | 40k chars | ~250ms latency, balanced quality |
804
- | `elevenlabs/multilingual-v2` | $0.10/1k chars | 10k chars | Long-form narration, audiobooks — 29 languages |
805
- | `elevenlabs/v3` | $0.10/1k chars | 5k chars | Max expressiveness, 70+ languages |
806
- | `bytedance/seed-audio-1.0` | $0.30/1k chars | 3k chars | Prompt-directed — describe voice, emotion and staging in words |
807
- | `elevenlabs/sound-effects` | $0.05/generation | 1k chars | Sound effects up to 22s |
800
+ | Model | Max Input | Notes |
801
+ |---|---|---|
802
+ | `elevenlabs/flash-v2.5` | 40k chars | ~75ms latency, 32 languages (default) |
803
+ | `elevenlabs/turbo-v2.5` | 40k chars | ~250ms latency, balanced quality |
804
+ | `elevenlabs/multilingual-v2` | 10k chars | Long-form narration, audiobooks — 29 languages |
805
+ | `elevenlabs/v3` | 5k chars | Max expressiveness, 70+ languages |
806
+ | `bytedance/seed-audio-1.0` | 3k chars | Prompt-directed — describe voice, emotion and staging in words |
807
+ | `elevenlabs/sound-effects` | 1k chars | Sound effects up to 22s |
808
808
 
809
809
  ```ts
810
810
  import { SpeechClient } from '@blockrun/llm';
@@ -823,7 +823,7 @@ const wav = await client.generate('Breaking news from the world of micropayments
823
823
  speed: 1.1,
824
824
  });
825
825
 
826
- // Sound effects (flat $0.05/generation)
826
+ // Sound effects (flat per generation)
827
827
  const fx = await client.soundEffect('rain on a tin roof, distant thunder');
828
828
 
829
829
  // List voices (free, rate-limited)
@@ -832,7 +832,7 @@ const voices = await client.listVoices();
832
832
 
833
833
  ### Virtual Portraits
834
834
 
835
- `PortraitClient` wraps `POST /v1/portrait/enroll` (paid, flat **$0.01** promo,
835
+ `PortraitClient` wraps `POST /v1/portrait/enroll` (paid, flat promo rate,
836
836
  no KYC). Enroll a face image by URL and get back a Token360 asset id (`ta_xxxxxx`).
837
837
  Pass that id as `realFaceAssetId` on a Seedance 2.0 video generation to keep the
838
838
  same AI character across clips. Payment settles only after Token360 confirms the
@@ -861,7 +861,7 @@ console.log(clip.data[0].url);
861
861
 
862
862
  ### Voice Calls
863
863
 
864
- `VoiceClient` wraps `POST /v1/voice/call` (paid, $0.54/call) and
864
+ `VoiceClient` wraps `POST /v1/voice/call` (paid, flat per call) and
865
865
  `GET /v1/voice/call/{callId}` (free polling) — AI-powered outbound phone
866
866
  calls powered by Bland.ai. The agent dials the recipient and runs a real-time
867
867
  conversation based on your `task` instructions. US + Canada destinations.
@@ -871,7 +871,7 @@ import { VoiceClient } from '@blockrun/llm';
871
871
 
872
872
  const client = new VoiceClient();
873
873
 
874
- // Initiate (paid $0.54)
874
+ // Initiate (paid)
875
875
  const result = await client.call({
876
876
  to: '+14155552671',
877
877
  task: 'You are a friendly assistant calling to confirm a 3pm dentist appointment.',
@@ -891,7 +891,7 @@ phone number you own; buy via `/v1/phone/numbers/buy`).
891
891
  ### Standalone Search
892
892
 
893
893
  `SearchClient` wraps `POST /v1/search` — standalone Grok Live Search.
894
- Pricing: `$0.025/source + margin` (10 sources ≈ `$0.26`).
894
+ Pricing is per source plus margin — see [blockrun.ai/models](https://blockrun.ai/models).
895
895
 
896
896
  ```ts
897
897
  import { SearchClient } from '@blockrun/llm';
@@ -912,11 +912,11 @@ endpoints across CEX/DEX market data, on-chain SQL, wallet intelligence,
912
912
  prediction markets (Polymarket + Kalshi), social analytics, news, VC fund
913
913
  data, and an OpenAI-compatible chat surface. Flat pricing per call:
914
914
 
915
- | Tier | Price/call | Examples |
916
- |------|-----------|----------|
917
- | 1 | $0.001 | `/market/price`, `/market/ranking`, `/news/feed`, prediction-market reads, social tweets |
918
- | 2 | $0.005 | `/exchange/depth`, `/exchange/klines`, `/wallet/detail`, `/search/*`, `/social/ranking` |
919
- | 3 | $0.020 | `/onchain/sql`, `/onchain/query`, `/onchain/schema`, `/chat/completions` |
915
+ | Tier | Examples |
916
+ |---|---|
917
+ | 1 | `/market/price`, `/market/ranking`, `/news/feed`, prediction-market reads, social tweets |
918
+ | 2 | `/exchange/depth`, `/exchange/klines`, `/wallet/detail`, `/search/*`, `/social/ranking` |
919
+ | 3 | `/onchain/sql`, `/onchain/query`, `/onchain/schema`, `/chat/completions` |
920
920
 
921
921
  Because the catalog is broad and evolving, the client deliberately ships a
922
922
  generic `get` / `post` pair instead of 84 typed wrappers. Pass the path
@@ -928,16 +928,16 @@ import { SurfClient } from '@blockrun/llm';
928
928
 
929
929
  const surf = new SurfClient();
930
930
 
931
- // Tier 1 — token price ($0.001)
931
+ // Tier 1 — token price
932
932
  const btc = await surf.get('/market/price', { symbol: 'BTC' });
933
933
 
934
- // Tier 2 — order book depth ($0.005)
934
+ // Tier 2 — order book depth
935
935
  const book = await surf.get('/exchange/depth', {
936
936
  exchange: 'binance',
937
937
  symbol: 'BTC-USDT',
938
938
  });
939
939
 
940
- // Tier 3 — raw on-chain SQL against 80+ ClickHouse tables ($0.020)
940
+ // Tier 3 — raw on-chain SQL against 80+ ClickHouse tables
941
941
  const rows = await surf.post('/onchain/sql', {
942
942
  query: 'SELECT block_number FROM ethereum.blocks ORDER BY block_number DESC LIMIT 5',
943
943
  });
@@ -957,7 +957,7 @@ Methods: `userLookup`, `userInfo`, `followers`, `following`, `followings`,
957
957
 
958
958
  `PriceClient` wraps the Pyth-backed market-data endpoints. Crypto, FX and
959
959
  commodity are fully free (price + history + list); 12 global stock markets
960
- and the `usstock` legacy alias charge `$0.001` for price + history (list is
960
+ and the `usstock` legacy alias are billed per call for price + history (list is
961
961
  always free). Pass `requireWallet: false` to construct a free-only client.
962
962
 
963
963
  ```ts
@@ -988,7 +988,7 @@ Three passthrough families live directly on `LLMClient` / `SolanaLLMClient`:
988
988
  ```ts
989
989
  const client = new LLMClient();
990
990
 
991
- // DefiLlama — protocols / TVL / yields / prices ($0.005/call, prices $0.001)
991
+ // DefiLlama — protocols / TVL / yields / prices
992
992
  const protocols = await client.defiProtocols();
993
993
  const aave = await client.defiProtocol('aave');
994
994
  const prices = await client.defiPrices(['coingecko:bitcoin', 'base:0x833589...']);
@@ -1003,7 +1003,7 @@ const gq = await client.dexGaslessQuote({ /* ... */ });
1003
1003
  const res = await client.dexGaslessSubmit({ trade: { /* signed eip712 */ } });
1004
1004
  const status = await client.dexGaslessStatus(res.tradeHash as string);
1005
1005
 
1006
- // Modal — sandboxed compute ($0.01 create CPU / $0.05 GPU, $0.001 exec)
1006
+ // Modal — sandboxed compute (create CPU / GPU, exec)
1007
1007
  const sb = await client.modalSandboxCreate({ image: 'python:3.11' });
1008
1008
  const out = await client.modalSandboxExec(sb.sandbox_id as string, ['python', '-c', 'print(42)']);
1009
1009
  await client.modalSandboxTerminate(sb.sandbox_id as string);
@@ -1017,7 +1017,7 @@ Generic escape hatches: `client.defi(path, params)`, `client.dex(path, params, b
1017
1017
  `RpcClient` wraps `POST /v1/rpc/{network}` — standard JSON-RPC 2.0 access to
1018
1018
  <!-- br:chains.rpc -->40<!-- /br:chains.rpc --> chains through one endpoint (Ethereum, Base, Solana, Polygon, BSC,
1019
1019
  Arbitrum, Optimism, Avalanche, Bitcoin, Sui, and more; powered by Tatum's RPC
1020
- gateway). No API key, no per-chain endpoints: flat **$0.002 per call** in
1020
+ gateway). No API key, no per-chain endpoints: one flat per-call rate in
1021
1021
  USDC; a JSON-RPC batch charges per element.
1022
1022
 
1023
1023
  ```ts
@@ -1038,7 +1038,7 @@ const balance = await client.call('base', 'eth_getBalance', [
1038
1038
  const slot = await client.call('solana', 'getSlot');
1039
1039
  const tip = await client.call('bitcoin', 'getblockcount');
1040
1040
 
1041
- // Batch: one payment, per-element pricing ($0.002 x N)
1041
+ // Batch: one payment, per-element pricing (rate x N)
1042
1042
  const out = await client.batch('polygon', [
1043
1043
  { method: 'eth_blockNumber' },
1044
1044
  { method: 'eth_gasPrice' },
@@ -1058,10 +1058,10 @@ blocks/receipts, `getTransaction`, ...) are served from a method-aware
1058
1058
  gateway cache — same price, lower latency.
1059
1059
 
1060
1060
  ### Testnet Models (Base Sepolia)
1061
- | Model | Price |
1062
- |-------|-------|
1063
- | `openai/gpt-oss-20b` | $0.001/request |
1064
- | `openai/gpt-oss-120b` | $0.002/request |
1061
+ | Model |
1062
+ |-------|
1063
+ | `openai/gpt-oss-20b` |
1064
+ | `openai/gpt-oss-120b` |
1065
1065
 
1066
1066
  *Testnet models use flat pricing (no token counting) for simplicity.*
1067
1067
 
@@ -1287,7 +1287,7 @@ import { LLMClient } from '@blockrun/llm';
1287
1287
 
1288
1288
  const client = new LLMClient();
1289
1289
 
1290
- // List markets with optional filters ($0.001/request)
1290
+ // List markets with optional filters
1291
1291
  const markets = await client.pm("polymarket/markets");
1292
1292
  const filtered = await client.pm("polymarket/markets", { status: "active", limit: 10 });
1293
1293
  const searched = await client.pm("polymarket/markets", { search: "bitcoin" });
@@ -1341,13 +1341,13 @@ Works on both `LLMClient` (Base) and `SolanaLLMClient`.
1341
1341
 
1342
1342
  Access [Exa](https://exa.ai)'s neural web search via x402. No API keys needed — pay-per-request. Available on **`LLMClient` (Base USDC)** and `SolanaLLMClient` (Solana USDC). Use Base as the primary path; the Solana gateway is awaiting `EXA_API_KEY` provisioning.
1343
1343
 
1344
- | Method | Description | Price |
1345
- |---|---|---|
1346
- | `exaSearch(query, options?)` | Neural/keyword web search | $0.01/request |
1347
- | `exaFindSimilar(url, options?)` | Find semantically similar pages | $0.01/request |
1348
- | `exaContents(urls, options?)` | Extract full text from URLs | $0.002/URL |
1349
- | `exaAnswer(query, options?)` | AI answer grounded in web search | $0.01/request |
1350
- | `exa(path, body)` | Generic proxy for any Exa endpoint | varies |
1344
+ | Method | Description |
1345
+ |---|---|
1346
+ | `exaSearch(query, options?)` | Neural/keyword web search |
1347
+ | `exaFindSimilar(url, options?)` | Find semantically similar pages |
1348
+ | `exaContents(urls, options?)` | Extract full text from URLs |
1349
+ | `exaAnswer(query, options?)` | AI answer grounded in web search |
1350
+ | `exa(path, body)` | Generic proxy for any Exa endpoint |
1351
1351
 
1352
1352
  ```typescript
1353
1353
  import { LLMClient } from '@blockrun/llm';