@blockrun/llm 3.19.0 β†’ 3.20.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -47,7 +47,7 @@ console.log(r.response); // the proof
47
47
  - πŸ†“ **<!-- br:models.free -->6<!-- /br:models.free --> genuinely free models** β€” $0 in and out, incl. two 1M-context Nemotrons, a multimodal one, and free coding models from Cohere and Poolside. No rate-limit gimmicks.
48
48
  - πŸ” **Two ways to connect** β€” a **BlockRun API key** billed against account credit ([sign up at user.blockrun.ai](https://user.blockrun.ai), [create a key](https://user.blockrun.ai/dashboard/keys), [add credit](https://user.blockrun.ai/dashboard/credits)), or a wallet signature with x402 micropayments and no account at all. Same code either way.
49
49
  - πŸ’Έ **Pay per request in USDC** β€” x402 micropayments on Solana or Base. $5 covers thousands of requests; agents can pay their own way.
50
- - πŸ›‘οΈ **Automatic failover** β€” transient errors (timeouts, 429, 5xx) walk the router's ranked fallback chain instead of failing your request.
50
+ - πŸ›‘οΈ **Automatic failover** β€” transient errors (timeouts, 429, 5xx) before payment walk the router's ranked fallback chain instead of failing your request; a failure after a payment was sent never buys a second model.
51
51
  - ⚑ **Streaming, OpenAI & Anthropic compat** β€” drop-in `chat.completions` / `messages` layers, SSE streaming, strict TypeScript.
52
52
  - 🎨 **Beyond chat** β€” image, video, music, speech, live search, prediction markets, crypto data, and 40-chain RPC through the same API key or wallet.
53
53
 
@@ -100,7 +100,21 @@ without them throws an error naming the exact install command.
100
100
 
101
101
  The Anthropic SDK is a runtime dependency because the public compatibility wrapper exposes its types. Solana signing dependencies remain optional and are unnecessary for account billing.
102
102
 
103
- ## Quick Start: API Key
103
+ ## Quick Start
104
+
105
+ Choose one billing mode. The chat and routing APIs are the same in every mode.
106
+
107
+ | Mode | Credential | Settlement | Best for |
108
+ |------|------------|------------|----------|
109
+ | **Account** | `BLOCKRUN_API_KEY` | BlockRun account credit | Applications and teams that want the simplest setup |
110
+ | **Solana wallet** | `SOLANA_WALLET_KEY` | Per-request USDC over x402 | Agents and non-custodial workloads; recommended wallet chain |
111
+ | **Base wallet** | `BASE_CHAIN_WALLET_KEY` | Per-request USDC over x402 | Existing EVM wallets and Base-native applications |
112
+
113
+ Credential precedence is deterministic: an explicit `apiKey` or `privateKey`
114
+ wins; otherwise `BLOCKRUN_API_KEY` wins over wallet environment variables.
115
+ Passing both explicit credentials is an error.
116
+
117
+ ### Account API key
104
118
 
105
119
  1. [Sign up or sign in](https://user.blockrun.ai).
106
120
  2. Create a key on [API Keys](https://user.blockrun.ai/dashboard/keys) and add credit on [Credits](https://user.blockrun.ai/dashboard/credits).
@@ -144,7 +158,7 @@ To switch back to a wallet, unset `BLOCKRUN_API_KEY` and create a new wallet cli
144
158
  address/balance helpers require a wallet. Account credit does not sign trades or
145
159
  transfer wallet funds. Service availability depends on the account gateway and model.
146
160
 
147
- ## Quick Start: Solana (Recommended Wallet Chain)
161
+ ### Solana wallet
148
162
 
149
163
  ```typescript
150
164
  import { setupAgentClient, SolanaLLMClient } from '@blockrun/llm';
@@ -158,7 +172,7 @@ console.log(await client.smartChat('Explain photosynthesis.'));
158
172
  const solana = new SolanaLLMClient({ privateKey: process.env.SOLANA_WALLET_KEY });
159
173
  ```
160
174
 
161
- ## Quick Start: Base Wallet
175
+ ### Base wallet
162
176
 
163
177
  ```typescript
164
178
  import { LLMClient } from '@blockrun/llm';
@@ -214,18 +228,17 @@ longer NVIDIA-only**, so pin these by full model id rather than by an
214
228
  | `cohere/north-mini-code` | 256K | Compact coding model, sub-second responses |
215
229
  | `poolside/laguna-xs-2.1` | 128K | Coding model |
216
230
 
217
- ## Quick Start (Solana)
218
-
219
- ```typescript
220
- import { SolanaLLMClient } from '@blockrun/llm';
221
-
222
- // SOLANA_WALLET_KEY env var (bs58-encoded Solana secret key)
223
- const client = new SolanaLLMClient();
224
- const response = await client.chat('openai/gpt-4o', 'gm Solana');
225
- console.log(response);
226
- ```
231
+ ## Documentation map
227
232
 
228
- Set `SOLANA_WALLET_KEY` to your bs58-encoded Solana secret key. Payments are automatic via x402 β€” your key never leaves your machine.
233
+ | If you want to… | Start here |
234
+ |-----------------|------------|
235
+ | Let the SDK choose the cheapest capable model | [Smart Routing](#smart-routing-router-core-v3) |
236
+ | Pay from a Solana wallet | [Solana Support](#solana-support) |
237
+ | Enable metered batch settlement | [Batch settlement](#batch-settlement-optional-metered-billing) |
238
+ | Understand wallet payments and settlement | [How Payment Works](#how-payment-works) |
239
+ | Use the OpenAI or Anthropic SDK surface | [Streaming](#streaming) and [Anthropic SDK Compatibility](#anthropic-sdk-compatibility) |
240
+ | Generate media or query data services | [Models and service APIs](#models-and-service-apis) |
241
+ | Configure production credentials safely | [Configuration](#configuration) and [Security](#security) |
229
242
 
230
243
  ## Smart Routing (Router Core V3)
231
244
 
@@ -294,10 +307,38 @@ console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
294
307
  `smartChat()` populates a fallback chain from the portfolio ranking and
295
308
  `chat()` / `chatCompletion()` walk it automatically when the primary model
296
309
  returns a transient error β€” timeouts, network failures, 429 rate limits, or
297
- 5xx responses (502/503/504/522/524). Other 4xx errors and `PaymentError`
298
- propagate immediately so wallet / auth issues surface fast. (Solana's internal
299
- stale-blockhash re-sign is a separate, lower-level retry inside the payment
300
- step β€” see [How Payment Works](#phase-2--every-request-pays-itself-automatic-x402).)
310
+ 5xx responses (502/503/504/522/524) β€” **before any payment was sent**. Other
311
+ 4xx errors and `PaymentError` propagate immediately so wallet / auth issues
312
+ surface fast. (Solana's internal stale-blockhash re-sign is a separate,
313
+ lower-level retry inside the payment step β€” see [How Payment Works](#phase-2--every-request-pays-itself-automatic-x402).)
314
+
315
+ The next model is a new paid request, so the chain is only walked when the
316
+ failed request cannot have been charged. Every error from these calls carries
317
+ a retry disposition, which `retryDisposition(err)` returns:
318
+
319
+ - `'unpaid'`: nothing chargeable was sent β€” the unpaid first request and its
320
+ 402 challenge, or signing the payment. A 429, 5xx, timeout or network error
321
+ here moves on to the next model.
322
+ - `'paid-or-in-doubt'`: the signed payment (exact or batch) was sent and may
323
+ have been charged β€” a timeout, abort or network error after sending it, any
324
+ error status in answer to it, or a 2xx whose body could not be read. The
325
+ error propagates; no other model is bought for the call. With an API key
326
+ the request itself is billed, so only the account API's explicit 4xx
327
+ answer (such as a 429) is `'unpaid'`; a 5xx or a timeout is not.
328
+
329
+ An error without a disposition counts as `'paid-or-in-doubt'`. Use the same
330
+ check in your own retry wrapper:
331
+
332
+ ```typescript
333
+ import { retryDisposition } from '@blockrun/llm';
334
+
335
+ try {
336
+ await client.chat('openai/gpt-5.2', 'hello');
337
+ } catch (err) {
338
+ if (retryDisposition(err) !== 'unpaid') throw err; // may have been charged: never resend
339
+ // safe to retry: nothing was paid
340
+ }
341
+ ```
301
342
 
302
343
  ```typescript
303
344
  // Manually pass a fallback chain to chat() / chatCompletion()
@@ -385,7 +426,7 @@ picked:
385
426
  | `taskType` | Portfolio task classification: `'chat'`, `'extraction'`, `'code_edit'`, `'code_agent'`, `'tool_agent'`, `'debug'`, `'reasoning'`, `'reasoning_math'`, `'long_context'`, `'vision'`, … |
386
427
  | `candidates` | Ordered, capability-eligible models ranked by the portfolio router; the first entry is `model` |
387
428
  | `candidateScores` | Per-candidate score breakdown (`quality` / `cost` / `speed` / `reliability`), ordered with `candidates` |
388
- | `fallbacks` | The chain `chat()` walks on transient errors (timeout / network / 429 / 5xx) β€” `candidates` minus the primary, with ClawRouter's proxy-namespace `free/*` ids mapped to their `nvidia/*` gateway ids (SDK-computed) |
429
+ | `fallbacks` | The chain `chat()` walks on transient errors before payment (timeout / network / 429 / 5xx) β€” `candidates` minus the primary, with ClawRouter's proxy-namespace `free/*` ids mapped to their `nvidia/*` gateway ids (SDK-computed) |
389
430
  | `savings` | 0–1 fraction saved vs the premium baseline |
390
431
  | `costEstimate` / `baselineCost` | Estimated cost of the pick vs that baseline, in USD |
391
432
  | `confidence` | Sigmoid-calibrated classifier confidence, 0–1 |
@@ -443,34 +484,51 @@ const tweet = await client.chat('xai/grok-4.5', 'What is trending on X?', { sear
443
484
  **Supported endpoint:** `https://sol.blockrun.ai/api`
444
485
  **Payment:** Solana USDC (SPL, mainnet)
445
486
 
446
- ### Metered billing with x402 batch-settlement (opt-in)
487
+ ### Batch settlement (optional metered billing)
447
488
 
448
- By default every Solana call is an `exact` payment: one SPL transfer per call, priced at the call's **ceiling** (the quote for your `maxTokens`), settled on-chain before the model answers. With `batch-settlement` you lock a small deposit in a payment channel once. After that, each call carries only a signed authorization for its ceiling. The gateway serves the call, meters what it **actually** cost, and charges that, never more than the ceiling. It then redeems the charges on-chain in batches.
489
+ By default, every Solana call uses x402 `exact`: one SPL transfer per call,
490
+ priced at the request ceiling and settled before the model answers. Batch
491
+ settlement replaces those per-call transfers with a bounded payment channel.
492
+ The client locks a small deposit once, each request authorizes no more than its
493
+ quoted ceiling, and the gateway charges the metered cost after the response.
449
494
 
450
- Batch mode runs on three optional peer dependencies:
495
+ > **No operator key needs to be created, copied, or stored in an environment
496
+ > variable.** `BLOCKRUN_SOL_OPERATOR` is a public constant exported by this SDK.
497
+ > Adding it to `batch.operators` is the explicit opt-in: it tells the client that
498
+ > this wallet trusts BlockRun's operator to sign vouchers against the channel,
499
+ > up to `maxDeposit`.
500
+
501
+ Without a `batch` option, behavior is unchanged and every request uses `exact`.
502
+ If batch is unavailable or unsafe for a request, the SDK falls back to `exact`
503
+ before it sends a batch payment. Once one is sent, a call whose payment may have
504
+ been charged is raised as `BatchPaymentUnresolvedError`, never paid again (see
505
+ [Runtime behavior](#runtime-behavior)).
506
+
507
+ In addition to the Solana wallet dependencies from [Installation](#installation),
508
+ batch mode needs three optional peer dependencies:
451
509
 
452
510
  ```bash
453
511
  npm install @x402/core@~2.28.0 @x402/svm@~2.28.0 @solana/kit
454
512
  ```
455
513
 
456
514
  ```typescript
457
- import { SolanaLLMClient } from '@blockrun/llm';
515
+ import { SolanaLLMClient, BLOCKRUN_SOL_OPERATOR } from '@blockrun/llm';
458
516
 
459
517
  const client = new SolanaLLMClient({
460
518
  privateKey: process.env.SOLANA_WALLET_KEY,
461
519
  batch: {
462
- // BlockRun's published operator public key. The SDK ships no default:
463
- // you decide which operator may sign vouchers against your deposit.
464
- operators: ['<BlockRun operator public key>'],
465
- // Most USDC this client will ever lock in the channel (deposit + top-ups),
466
- // and so the most an operator could claim. Default "$1".
520
+ // Provided by @blockrun/llm. This explicit trust decision enables batch.
521
+ operators: [BLOCKRUN_SOL_OPERATOR],
522
+ // Most escrow at stake in the channel: deposited and not yet settled. Default: "$1".
467
523
  maxDeposit: '$5',
468
524
  },
469
525
  });
470
526
 
471
- // The first call opens the channel. By default the SDK deposits 5x this
472
- // call's ceiling (never more than maxDeposit), and the call that opens or tops
473
- // up the channel is charged its quoted price, as it would be with exact.
527
+ // The first call opens the channel. The deposit is the 402's extra.minDeposit
528
+ // (the @x402/svm server defaults it to 3x this call's ceiling), or 5x the
529
+ // ceiling when the 402 names none, never more than the room maxDeposit leaves.
530
+ // The call that opens or tops up the channel is charged its quoted price, as
531
+ // it would be with exact.
474
532
  await client.chat('openai/gpt-4o-mini', 'gm');
475
533
 
476
534
  // Every later call is metered: you pay for the tokens the call actually used.
@@ -486,20 +544,149 @@ console.log(client.getSpending()); // { totalUsd: <actual charges>, calls: 2 }
486
544
  await client.closeBatchChannel();
487
545
  ```
488
546
 
489
- How it behaves:
547
+ #### Trust and spending boundary
548
+
549
+ - **Explicit trust.** `operators` is required and has no implicit default. The
550
+ package exports BlockRun's production public key as `BLOCKRUN_SOL_OPERATOR`
551
+ (`5YKPQUFjw5WQqhSUkEGKNNfYYVqnRRNbpYyL71qQ1vm3`) so applications do not
552
+ duplicate it. The SDK never learns a trusted key from an untrusted 402.
553
+ - **Bounded exposure.** `maxDeposit` caps the escrow at stake: the channel's
554
+ deposit minus what has already been settled on-chain. A top-up is allowed
555
+ only while `(deposit - settled) + topUp <= maxDeposit`, so a long-lived
556
+ channel keeps topping up as BlockRun settles what it charged, and its
557
+ unsettled balance never exceeds the cap. In server-signed mode that makes
558
+ `maxDeposit` the most the selected operator could claim beyond what has
559
+ already been settled, without another authorization from the payer. Before a
560
+ top-up the SDK reads the channel's settled amount from the chain; if the
561
+ RPC cannot answer, it assumes nothing was settled. If the account there
562
+ cannot be read as a payment channel (another owner, too short, an
563
+ unsupported encoding), or the installed `@x402/svm` lays out channel
564
+ accounts differently from what the SDK decodes, it signs no top-up and the
565
+ call pays `exact` (`channel_unreadable`), keeping the record.
566
+ - **Key rotation.** `operators` accepts a list so an old and new BlockRun key can
567
+ overlap during a controlled rotation.
568
+ - **Server-signed channels only.** Batch pays only into server-signed
569
+ (operator) channels, the ones `closeBatchChannel()` can refund. A batch
570
+ accept that is client-signed (`extra.voucherSigner` omitted or `"client"`),
571
+ which a custom gateway could offer, is never paid: the SDK ignores it, and a
572
+ 402 whose only batch accepts are client-signed pays `exact`
573
+ (`client_signed_not_supported`). A client-signed channel an earlier SDK
574
+ version opened can still be closed: `closeBatchChannel()` refunds a stored
575
+ one through `@x402/svm`'s client-signed refund when the gateway's challenge
576
+ offers its client-signed accept (after the trusted server-signed channel,
577
+ when both are stored).
578
+
579
+ #### Runtime behavior
580
+
581
+ Batch never pays twice for one call. Before a batch payment is sent, any
582
+ problem falls back to `exact`. Once it is sent, the SDK decides what the
583
+ gateway's answer proves, in one place:
584
+
585
+ | The gateway's answer to the payment | What the SDK does |
586
+ | --- | --- |
587
+ | a 2xx | the call is served and booked (a missing or unreconciled receipt only makes the SDK re-read the channel before the next payment) |
588
+ | first send: a 402, or a recognised refusal (`batch_payer_not_allowed`, `batch_payer_not_admitted`, `batch_admission_paused`, `batch_server_signed_only`, `batch_unavailable`, `batch_channel_limit`, `PAYMENT_VERIFICATION_UNAVAILABLE`), with no receipt or a failed one with no transaction | nothing was charged: pays `exact`. A deposit refused with a 402 or `PAYMENT_VERIFICATION_UNAVAILABLE` may still have been broadcast, so its channel is re-read from the chain before anything pays into it again |
589
+ | first send: a 429 whose failed receipt proves nothing was broadcast (`batch_account_channel_capacity_exhausted`, `batch_channel_capacity_exhausted`, `batch_deposit_rate_limited`) | nothing was charged: backs off, then pays again with a new batch payment from a fresh 402; `exact` once `rateLimit` runs out |
590
+ | first send: a 429 with no receipt | **in doubt**: after the backoff, replays the identical payment **once**; a 2xx with a success receipt ends it, anything else raises |
591
+ | first send of an authorization: an error status with a cancelled receipt (`success: false`, `errorReason: "batch_cancelled"`, no transaction) | nothing was charged: the error is raised with retry disposition `'unpaid'`, so `fallbackModels` / `smartChat()` move on to the next model |
592
+ | anything else: a timeout, an abort or a network error after sending; a 5xx or other error status without a recognised refusal; a receipt naming a transaction, saying `settlement_pending` or saying it succeeded on an error status; a 429 with any other receipt | **in doubt**: raises `BatchPaymentUnresolvedError` |
593
+
594
+ - **Trust model.** The gateway is already trusted with the escrow up to
595
+ `maxDeposit`, so its explicit answer to a payment's first send is taken as
596
+ proof that nothing was charged. Nothing else is: not an error's type or
597
+ `cause.code` (a request can reach the gateway before the connection fails),
598
+ not a missing receipt, and not chain state. An answer to a replay only
599
+ describes the replay: a 402 or `duplicate_settlement` on a replay means the
600
+ original reached the gateway, so it stays in doubt.
601
+ - **In doubt means raised, never paid again.** A call in doubt throws
602
+ `BatchPaymentUnresolvedError` and is never paid again: no new
603
+ authorization, no new deposit, no `exact`, and no `fallbackModels` or
604
+ `smartChat()` fallback model (it is a `PaymentError` whose retry
605
+ disposition is `'paid-or-in-doubt'`). It carries `wallet`, `requestId` (the
606
+ payment that may have been charged), `channelId`, `payloadKind` (`'open'`,
607
+ `'top-up'` or `'authorization'`), `depositInDoubt`, `status` (the gateway's
608
+ last answer, when there was one), `cause` (that answer as an `APIError`, or
609
+ the transport error) and `reason`:
610
+ - `'replay_unresolved'`: a 429 without a receipt, whose one replay got no
611
+ success receipt (or `rateLimit` left no room for it);
612
+ - `'ambiguous_rate_limit'`: a 429 whose receipt does not prove nothing was
613
+ broadcast;
614
+ - `'no_response'`: sending it threw (timeout, abort, network error);
615
+ - `'duplicate_settlement'`: the first answer was a 402 naming
616
+ `duplicate_settlement`, meaning another copy of this payment had already
617
+ reached the gateway;
618
+ - `'outcome_unknown'`: any other answer that neither serves the call nor
619
+ proves nothing was charged.
620
+
621
+ A retry is up to you, and it is a new payment:
622
+
623
+ ```typescript
624
+ import type { BatchPaymentUnresolvedError } from '@blockrun/llm';
625
+
626
+ try {
627
+ await client.smartChat('Summarize x402 in one line');
628
+ } catch (err) {
629
+ // Match by name: `instanceof` fails when the error comes from another
630
+ // loaded copy of the SDK (the CJS and ESM builds side by side).
631
+ if (err instanceof Error && err.name === 'BatchPaymentUnresolvedError') {
632
+ const unresolved = err as BatchPaymentUnresolvedError;
633
+ console.warn(`payment ${unresolved.requestId} from ${unresolved.wallet} may have been charged`);
634
+ }
635
+ throw err;
636
+ }
637
+ ```
638
+
639
+ Raising is deliberately conservative. Resolving a payment in doubt
640
+ automatically needs the gateway's help: a receipt (`PAYMENT-RESPONSE`) on
641
+ every response to a paid request, and a request-status lookup that answers
642
+ for a request id and refuses (fences) one it never saw, in the spirit of
643
+ [idempotency keys](https://datatracker.ietf.org/doc/draft-ietf-httpapi-idempotency-key-header/).
644
+ The SDK will use them once the gateway offers them.
645
+ - **Safe fallback.** These pay `exact` instead, before anything is sent or
646
+ after an answer proving nothing was charged:
647
+ - the 402 has no batch accept, offers only client-signed ones (`client_signed_not_supported`), or names an operator outside `operators`;
648
+ - the deposit the call needs does not fit in the room `maxDeposit` leaves;
649
+ - another batch call for the same wallet is in flight;
650
+ - a deposit is in doubt, or a channel record is being re-read from the chain (below);
651
+ - the gateway refuses batch, or a 402 answers the first send (see the table);
652
+ - the gateway still answers 429 after the retries `rateLimit` allows, when nothing was charged.
653
+
654
+ Parallel calls never queue behind the channel, except during a 429 cooldown (below).
655
+ - **Deposit size.** An open or top-up deposits the 402's `extra.minDeposit` when it is at least the call's ceiling (the `@x402/svm` server defaults it to 3x the ceiling), otherwise 5x the ceiling. It is never less than the call needs, and never more than the room `maxDeposit` leaves: `maxDeposit - (deposit - settled)`, which for an open is all of `maxDeposit`. A call whose shortfall does not fit pays `exact` (`deposit_over_cap`).
656
+ - **429 means wait, not give up.** After a 429 the call waits for `Retry-After` (seconds or an HTTP date; without one, about 1s, 2s, 4s with jitter). A `Retry-After` that is not a finite delay of at most a day is ignored, and the default backoff applies. While the wallet cools down, its other calls wait as well instead of sending new channel opens. `batch.rateLimit` bounds this per call: `maxAttempts` (default `3` batch payments sent, the first and the one replay included) and `maxWaitMs` (default `60000` ms spent waiting). A call that cannot wait out `Retry-After` within what is left of `maxWaitMs` pays `exact` at once, and the wallet's shared cooldown never lasts longer than `maxWaitMs`, so no single header keeps batch off for longer than that. A 429 that proves nothing was broadcast is paid again with a new payment, rebuilt from a fresh 402 challenge (one unpaid request) because the original's blockhash and slot hints may have gone stale. That unpaid request is the call itself: if the gateway serves it (a 2xx), that response is the call's result and nothing more is paid (a `recovered` event, `served_unpaid_on_rechallenge`); if it gets any other answer or none, the call throws that error with retry disposition `'unpaid'` (nothing was paid for it, so `fallbackModels` may take over). It is never paid `exact` against the stale challenge. A 429 without a receipt is never rebuilt: it gets its one replay, byte for byte (same request id, same signed authorization, same signed deposit transaction).
657
+ - **A deposit in doubt blocks new deposits.** When an open or top-up is in doubt, or a deposit-bearing call was served without a receipt the SDK could reconcile, the wallet signs no batch payment at all (so no second deposit can land next to it) and its calls pay `exact` (`channel_resync_pending`) until the chain settles the deposit. The SDK reads channels at `finalized` commitment only. A deposit has landed once the finalized channel holds it. It has provably never landed once the finalized chain is 300 blocks past a block height the SDK read itself after the payment was built (a transaction expires 150 blocks after its blockhash), and a read at least that recent still does not show it: then an open's record is dropped and a top-up's channel is rewritten with the deposit it really has. Wall-clock time, and any `lastValidBlockHeight` a 402 names, are never used. This settles only the deposit, never whether the gateway charged the payment that carried it.
658
+ - **No silent fallbacks.** Every fallback, every 429 backoff, every recovery, every channel re-read and every payment raised in doubt writes one line to stderr, each time it happens:
659
+
660
+ ```text
661
+ [@blockrun/llm] batch-settlement event=backoff reason=rate_limited status=429 errorReason=batch_deposit_rate_limited retryAfterMs=30000 attempt=1 wallet=<address> next=retry
662
+ [@blockrun/llm] batch-settlement event=unresolved reason=no_response attempt=1 wallet=<address> next=raise detail="the payment got no answer (AbortError: This operation was aborted); it may have been charged, so it is not paid again"
663
+ ```
490
664
 
491
- - **Opt-in and trust-pinned.** In server-signed mode BlockRun's operator key signs the vouchers, so it could claim up to the whole unspent deposit. The SDK only enters a channel for an operator you list in `operators`, and only up to `maxDeposit`. A 402 that asks for any other operator is paid with `exact`.
492
- - **Never worse than `exact`.** These cases all pay with `exact` instead:
493
- - the 402 has no batch accept;
494
- - the deposit would exceed `maxDeposit`;
495
- - the gateway refuses batch (payer not admitted, admission paused, verifier unavailable);
496
- - another batch call for the same wallet is in flight.
665
+ `batch.onEvent` receives each one as a `SolanaBatchEvent`: `{ type: 'fallback' | 'backoff' | 'recovered' | 'resync' | 'unresolved', reason, status?, errorReason?, retryAfterMs?, attempt?, detail?, wallet, at }`. `reason` is a stable code: `rate_limited` or `cooldown` (backoffs, and the `recovered` event after one, or `served_unpaid_on_rechallenge`); for fallbacks `not_offered`, `client_signed_not_supported`, `channel_busy`, `untrusted_operator`, `deposit_over_cap`, `channel_pending`, `closed_during_call`, `peer_dependency_missing`, `payment_creation_failed`, `wallet_config_conflict`, `channel_store_locked`, `deposit_journal_unreadable`, `deposit_journal_failed`, `channel_resync_pending`, `channel_resync_failed`, `channel_unreadable`, `payment_required`, `rate_limited` or one of the recognised refusal codes above; for `unresolved`, the `BatchPaymentUnresolvedError` reasons; for `resync`, the re-read reasons below. The callback may be async; one that throws or rejects is logged and never affects the payment. `client.getBatchStats()` returns this client's counters:
497
666
 
498
- Parallel calls never queue behind the channel. A gateway refusal charges nothing, so the `exact` retry is the only charge.
667
+ ```typescript
668
+ const client = new SolanaLLMClient({
669
+ batch: {
670
+ operators: [BLOCKRUN_SOL_OPERATOR],
671
+ rateLimit: { maxAttempts: 3, maxWaitMs: 60_000 }, // the defaults
672
+ onEvent: (event) => metrics.increment(`batch.${event.type}`, { reason: event.reason }),
673
+ },
674
+ });
675
+
676
+ client.getBatchStats();
677
+ // { fallbacks: 1, fallbacksByReason: { rate_limited: 1 }, backoffs: 2, retries: 2, recoveries: 0,
678
+ // resyncs: 0, unresolved: 0, unresolvedByReason: {} }
679
+ ```
499
680
  - **Scope.** Batch covers non-streaming chat: `chat`, `chatCompletion`, `smartChat`, and `smartChatCompletion`. `stream()` and the image and media jobs still pay with `exact`.
500
- - **One channel per wallet.** Every `SolanaLLMClient` for a wallet in one process shares a single channel. A second client with different `batch` options pays `exact`. On disk, one live process owns a wallet's channel file through a pid lock; other processes pay `exact` until that process exits.
501
- - **State.** The open channel is saved to `~/.blockrun/solana-batch/<wallet>.json` (mode `0600`), so a restart reuses it. A custom `channelStore` path must also be per wallet. With `channelStore: false` the channel is kept in memory only and found again on-chain, and nothing coordinates processes, so use it for a single process per wallet. If a channel open or top-up gets no clean answer, or after `closeBatchChannel()`, the SDK forgets the saved channel and reads the real one from the chain. The channel uses your `rpcUrl` / `SOLANA_RPC_URL`.
502
- - **Rollout.** sol.blockrun.ai lists `exact` first, and offers `batch-settlement` only once BlockRun enables it. Until then, and for any payer not yet admitted, a client with `batch` set simply keeps paying `exact`.
681
+ - **One channel per wallet.** Every `SolanaLLMClient` for a wallet in one process shares a single channel, including clients built from different loaded copies of this package (its CJS and ESM builds, or two installs). A second client with different `batch` options pays `exact`. On disk, one live process owns a wallet's channel file through a lock file holding its pid, in the same format earlier releases write and read, so a process still on an older version that shares the channel store (say, during a rolling upgrade) sees the lock as live too; a random ownership token sits beside it in `<lock>.owner`. Other processes pay `exact` until that process exits. A lock that names the current process but was not taken by this SDK version is never treated as stale; if no process uses it (say, a crashed process whose pid was reused), delete the `.lock` and `.lock.owner` files.
682
+ - **State and crash safety.** The open channel is saved to `~/.blockrun/solana-batch/<wallet>.json` (mode `0600`), so a restart reuses it; a custom `channelStore` path must also be per wallet. Before an open or top-up is sent, its intent is written to `<wallet>.json.deposit-intents` beside it (the store's full path plus `.deposit-intents`, so every store has its own; mode `0600`; the file and its directory are fsynced where the platform allows), and removed once the deposit is reconciled. After a crash, every intent left there makes the new process treat that deposit as in doubt (above): it signs no deposit until the chain settles it, and it never re-sends the old payment or pays for its request again. A deposit that cannot be journaled is not sent (`deposit_journal_failed`, paid `exact`). Intents an earlier version journaled under its old name (`<wallet>.deposit-intents.json`, which a store named `<wallet>` without `.json` shared) are moved over on start, only the ones for this wallet; an old journal that cannot be read keeps batch off as below. Only a missing journal counts as empty: one that cannot be read, or is not a well-formed journal this version knows (a missing `intents`, an unknown `version`, a malformed entry), keeps batch off for the wallet and its calls pay `exact` (`deposit_journal_unreadable`); the file is left as it is. **With `channelStore: false` this crash recovery is not available:** the channel is kept in memory and found again on-chain, and a deposit that was in flight when the process died is not known to be in doubt. Nothing coordinates processes in that mode either, so use it for a single process per wallet.
683
+ - **Closing.** `closeBatchChannel()` closes the channel the gateway's current 402 names: the one for its trusted operator, receiver authorizer, fee payer, asset and receiver. It first re-reads that record if it is in doubt (for example after a process died mid-deposit) and loads that channel alone, so the escrow can be recovered without another paid call; after the close the SDK forgets that channel only. During a key rotation, a channel opened with the other operator key keeps its record and can be closed once the gateway names that operator. A client-signed channel left by an earlier SDK version is closed the same way once no trusted server-signed record is left for that challenge, through the challenge's client-signed accept. It closes nothing and throws `BatchCloseDeferredError` while a chat call for the wallet is in flight (`reason: 'call_in_flight'`, including one waiting out a 429, so that retry can never open a channel you believe is closed) or while any of its deposits is in doubt (`reason: 'deposit_in_doubt'`, with `channelIds`); call it again once the call returns or the deposit is settled (usually a few minutes: it waits for finality). A chat that was still waiting for its first answer when a close completed pays `exact` instead of opening a new channel (`closed_during_call`).
684
+ - **RPC.** Batch uses the same Solana RPC as `exact`: `rpcUrl` / `SOLANA_RPC_URL`, with `rpcHeaders` / `SOLANA_RPC_HEADERS` / `SOLANA_RPC_API_KEY` for header-auth endpoints. `@x402/svm` builds its RPC client from a URL alone, so when headers are configured the SDK adds them through a thin wrapper around the global `fetch` that only touches requests to exactly that RPC URL made during a batch payment. The RPC must support `finalized` commitment and `minContextSlot` (standard on current Solana RPCs); one that does not keeps calls on `exact`.
685
+ - **A funded channel is never forgotten.** A refused or rate-limited open or top-up that charged nothing leaves the saved channel exactly as it was confirmed. When the SDK cannot trust its record β€” a deposit in doubt; a response whose receipt is missing or was rebuilt without a voucher; a pending deposit or journaled intent left by a process that died; a chain that shows more deposit than the record before a top-up β€” it re-reads that one channel by its address (`getAccountInfo` at `finalized`, no program scan) before the next payment, and rewrites the record with the on-chain deposit and the highest of its confirmed cumulative charge, the one the saved record already holds, and the on-chain settled amount. The saved cumulative never moves backwards, even when a journaled intent outlived the receipt that reconciled its deposit (vouchers are redeemed on-chain later, so the settled amount can lag). Each re-read emits a `resync` event (`reason`: `deposit_unanswered`, `deposit_failed`, `deposit_rate_limited`, `receipt_missing`, `receipt_unreconciled`, `orphaned_deposit`, `deposit_unrecorded`, `channel_unusable`; `detail`: what the chain showed). A top-up never goes into a channel the chain shows closing or closed, or that is not your wallet's for its operator and mint: that record is re-read at once (`channel_unusable`), a closed channel's record is dropped and the call carries on as with no channel (a fresh open when it fits in `maxDeposit`), and anything else keeps the record and pays `exact` (`channel_unreadable`). Until it succeeds, calls pay `exact` (`channel_resync_pending`, or `channel_resync_failed` when the chain cannot be read). Only an absent account counts as "no channel": an account at that address that cannot be read as your channel (another owner, layout, encoding, payer, operator or mint) keeps the record and pays `exact` (`channel_unreadable`), and is never a reason to sign a new deposit. Every channel read (these re-reads and the settled read before a top-up) first checks that the installed `@x402/svm` still uses the channel layout the SDK decodes; when it does not, nothing is decoded and those calls pay `exact` (`channel_unreadable`) with the record kept.
686
+ - **Rollout.** `sol.blockrun.ai` lists `exact` first and advertises
687
+ `batch-settlement` only when it is enabled. A configured client remains fully
688
+ compatible while rollout is paused or restricted because it falls back to
689
+ `exact`.
503
690
 
504
691
  ## Arc Support
505
692
 
@@ -588,7 +775,7 @@ const summary = getCostSummary(); // across sessions (~/.blockr
588
775
  console.log(`Lifetime: $${summary.totalUsd.toFixed(2)} over ${summary.calls} calls`);
589
776
  ```
590
777
 
591
- In wallet mode, every paid request is a real on-chain USDC transfer β€” look up your wallet address on [Basescan](https://basescan.org) (or a Solana explorer) to verify each settlement independently. The exception is Solana [batch-settlement](#metered-billing-with-x402-batch-settlement-opt-in): there the on-chain records are the channel deposit and BlockRun's batched redemptions, and each call is a signed voucher.
778
+ In wallet mode, every paid request is a real on-chain USDC transfer β€” look up your wallet address on [Basescan](https://basescan.org) (or a Solana explorer) to verify each settlement independently. The exception is Solana [batch settlement](#batch-settlement-optional-metered-billing): there the on-chain records are the channel deposit and BlockRun's batched redemptions, and each call is a signed voucher.
592
779
 
593
780
  **Non-custodial by design: your private key never leaves your machine** β€” it is only used for local signing, and no funds are ever held by BlockRun.
594
781
 
@@ -639,13 +826,25 @@ The per-API client classes (`LLMClient`, `ImageClient`, `VideoClient`,
639
826
  `PriceClient`, `SurfClient`) all remain β€” they will be soft-deprecated in 2.6 (rewritten as
640
827
  shims over `BlockrunClient`) and removed in 3.0.
641
828
 
642
- ## Available Models
829
+ ## Models and service APIs
830
+
831
+ The live catalog is the source of truth for availability, context windows, and
832
+ pricing: call `client.listModels()` or visit
833
+ [Models & Pricing](https://blockrun.ai/models). The reference below documents
834
+ the SDK's model families and non-chat client surfaces; it is intentionally
835
+ secondary to the runtime catalog.
643
836
 
644
837
  **Prices are not listed here.** They change often, and a number copied into a
645
838
  README is wrong the day after it lands. See **[blockrun.ai/models](https://blockrun.ai/models)**
646
839
  for live rates, or read them from the catalog at runtime β€” `client.listModels()`
647
840
  and `client.listImageModels()` return exactly what the gateway is charging.
648
841
 
842
+ <details>
843
+ <summary><strong>Model and service catalog</strong> β€” chat families, media, search, market data, DeFi, and RPC</summary>
844
+
845
+ The tables and examples in this section are an SDK surface reference. Use the
846
+ live catalog for runtime availability and prices.
847
+
649
848
  ### OpenAI GPT-6 Family
650
849
 
651
850
  The GPT-6 generation: Astra is the flagship for long-horizon agentic work and
@@ -1055,6 +1254,9 @@ console.log(result.summary);
1055
1254
  for (const url of result.citations ?? []) console.log(url);
1056
1255
  ```
1057
1256
 
1257
+ The same endpoint is available as `client.search()` on `LLMClient` and
1258
+ `SolanaLLMClient` when the selected gateway supports it.
1259
+
1058
1260
  ### Surf Crypto Data
1059
1261
 
1060
1262
  `SurfClient` exposes the full `/v1/surf/*` catalog β€” 84+ pay-per-call
@@ -1214,28 +1416,7 @@ gateway cache β€” same price, lower latency.
1214
1416
 
1215
1417
  *Testnet models use flat pricing (no token counting) for simplicity.*
1216
1418
 
1217
- ## Standalone Search
1218
-
1219
- Search web, X/Twitter, and news without using a chat model:
1220
-
1221
- ```typescript
1222
- import { LLMClient } from '@blockrun/llm';
1223
-
1224
- const client = new LLMClient();
1225
-
1226
- const result = await client.search('latest AI agent frameworks 2026');
1227
- console.log(result.summary);
1228
- for (const cite of result.citations ?? []) {
1229
- console.log(` - ${cite}`);
1230
- }
1231
-
1232
- // Filter by source type and date range
1233
- const filtered = await client.search('BlockRun x402', {
1234
- sources: ['web', 'x'],
1235
- fromDate: '2026-01-01',
1236
- maxResults: 5,
1237
- });
1238
- ```
1419
+ </details>
1239
1420
 
1240
1421
  ## Image Editing (img2img)
1241
1422
 
@@ -1579,6 +1760,10 @@ const client = new LLMClient({
1579
1760
  you pass no explicit credential, the client runs in account mode. Pass an explicit
1580
1761
  `privateKey` to force wallet mode.
1581
1762
 
1763
+ `BLOCKRUN_SOL_OPERATOR` is **not** an environment variable or a secret. It is an
1764
+ SDK export used for the explicit batch-settlement trust configuration shown in
1765
+ [Solana Support](#batch-settlement-optional-metered-billing).
1766
+
1582
1767
  ## Error Handling
1583
1768
 
1584
1769
  ```typescript
@@ -1597,6 +1782,10 @@ try {
1597
1782
  }
1598
1783
  ```
1599
1784
 
1785
+ Before retrying a failed call yourself, check `retryDisposition(error)`
1786
+ (see [Automatic Fallback on Transient Errors](#automatic-fallback-on-transient-errors)):
1787
+ anything but `'unpaid'` may already have been charged.
1788
+
1600
1789
  ## Testing
1601
1790
 
1602
1791
  ### Running Unit Tests