@blockrun/llm 3.10.0 → 3.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -236,7 +236,24 @@ Every paid request is a real on-chain USDC transfer — look up your wallet addr
236
236
 
237
237
  ## Smart Routing (ClawRouter)
238
238
 
239
- Let the SDK automatically pick the cheapest capable model for each request:
239
+ Let the SDK automatically pick the cheapest capable model for each request — **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% cheaper than pinning Claude Opus 5** for the same traffic on `auto`, **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** on `eco`.
240
+
241
+ Not an "up to" figure. The baseline, the workload mix and the token ratio are
242
+ published in [`savings-mix.json`](https://github.com/BlockRunAI/blockrun/blob/main/src/brand/savings-mix.json),
243
+ priced against the live catalog, so anyone can recompute the claim and get the
244
+ same answer.
245
+
246
+ Smart routing is powered by [ClawRouter](https://github.com/BlockRunAI/ClawRouter) —
247
+ the open-source, agent-first LLM router: wallet signatures instead of API keys,
248
+ USDC micropayments instead of credit cards, and <1ms fully-local routing with
249
+ zero external calls. In this SDK it is an **optional peer dependency** —
250
+ install it alongside the SDK:
251
+
252
+ ```bash
253
+ npm install @blockrun/clawrouter
254
+ ```
255
+
256
+ Only `smartChat()` needs it: every other API works without it, and the package is loaded lazily, so a missing or broken router can never break `import '@blockrun/llm'`. Calling `smartChat()` without it throws an error naming the package instead of a cryptic module-load failure.
240
257
 
241
258
  ```typescript
242
259
  import { LLMClient } from '@blockrun/llm';
@@ -247,23 +264,28 @@ const client = new LLMClient();
247
264
  const result = await client.smartChat('What is 2+2?');
248
265
  console.log(result.response); // '4'
249
266
  console.log(result.model); // 'moonshot/kimi-k2.5' (cheap, fast)
250
- console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 87%'
267
+ console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 88%'
251
268
 
252
269
  // Complex reasoning task -> routes to reasoning model
253
270
  const complex = await client.smartChat('Prove the Riemann hypothesis step by step');
254
271
  console.log(complex.model); // 'xai/grok-4-1-fast-reasoning'
255
272
 
273
+ // Inspect how the request was classified and ranked (Router v3.4 portfolio).
274
+ console.log(complex.routing.method); // 'portfolio'
275
+ console.log(complex.routing.taskType); // 'reasoning'
276
+ console.log(complex.routing.candidates); // ranked, capability-eligible models
277
+
256
278
  // Inspect the fallback chain SmartChat will walk on transient errors.
257
279
  console.log(complex.routing.fallbacks); // ['anthropic/claude-opus-4.7', ...]
258
280
  ```
259
281
 
260
282
  ### Automatic Fallback on Transient Errors
261
283
 
262
- `smartChat()` populates a tier-specific fallback chain and `chat()` /
263
- `chatCompletion()` walk it automatically when the primary model returns a
264
- transient error — timeouts, network failures, or 5xx responses (502/503/504/
265
- 522/524). 4xx errors and `PaymentError` propagate immediately so wallet /
266
- auth issues surface fast.
284
+ `smartChat()` populates a fallback chain from the portfolio ranking and
285
+ `chat()` / `chatCompletion()` walk it automatically when the primary model
286
+ returns a transient error — timeouts, network failures, 429 rate limits, or
287
+ 5xx responses (502/503/504/522/524). Other 4xx errors and `PaymentError`
288
+ propagate immediately so wallet / auth issues surface fast.
267
289
 
268
290
  ```typescript
269
291
  // Manually pass a fallback chain to chat() / chatCompletion()
@@ -276,12 +298,12 @@ const reply = await client.chat('nvidia/deepseek-v4-flash', 'hello', {
276
298
 
277
299
  ### Routing Profiles
278
300
 
279
- | Profile | Description | Best For |
280
- |---------|-------------|----------|
281
- | `free` | NVIDIA free tier — smart-routes across <!-- br:models.free -->6<!-- /br:models.free --> models (DeepSeek V4 Flash, Nemotron Nano Omni, Qwen3, Llama 4, Mistral, plus 2 hidden gpt-oss) | Zero-cost testing, dev, prod |
282
- | `eco` | Cheapest models per tier (DeepSeek, xAI) | Cost-sensitive production |
283
- | `auto` | Best balance of cost/quality (default) | General use |
284
- | `premium` | Top-tier models (OpenAI, Anthropic) | Quality-critical tasks |
301
+ | Profile | Strategy | Savings vs Opus 5 | Best For |
302
+ |---------|----------|-------------------|----------|
303
+ | `free` | NVIDIA free tier — smart-routes across <!-- br:models.free -->6<!-- /br:models.free --> models (DeepSeek V4 Flash, Nemotron Nano Omni, Qwen3, Llama 4, Mistral, plus 2 hidden gpt-oss) | **100%** | Zero-cost testing, dev, prod |
304
+ | `eco` | Cheapest capable model per tier | **<!-- br:savings.ecoVsBaselinePct -->98<!-- /br:savings.ecoVsBaselinePct -->%** | Cost-sensitive production |
305
+ | `auto` | Best balance of cost/quality (default) | **<!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->%** | General use |
306
+ | `premium` | Top-tier models (OpenAI, Anthropic) | 0% | Quality-critical tasks |
285
307
 
286
308
  ```typescript
287
309
  // Use premium models for complex tasks
@@ -294,23 +316,78 @@ console.log(result.model); // 'anthropic/claude-opus-4.7'
294
316
 
295
317
  ### How ClawRouter Works
296
318
 
297
- ClawRouter uses a 14-dimension rule-based classifier to analyze each request:
298
-
299
- - **Token count** - Short vs long prompts
300
- - **Code presence** - Programming keywords
301
- - **Reasoning markers** - "prove", "step by step", etc.
302
- - **Technical terms** - Architecture, optimization, etc.
303
- - **Creative markers** - Story, poem, brainstorm, etc.
304
- - **Agentic patterns** - Multi-step, tool use indicators
305
-
306
- The classifier runs in <1ms, 100% locally, and routes to one of four tiers:
307
-
308
- | Tier | Example Tasks | Auto Profile Model |
309
- |------|---------------|-------------------|
310
- | SIMPLE | "What is 2+2?", definitions | moonshot/kimi-k2.5 |
311
- | MEDIUM | Code snippets, explanations | xai/grok-code-fast-1 |
312
- | COMPLEX | Architecture, long documents | google/gemini-3.1-pro |
313
- | REASONING | Proofs, multi-step reasoning | xai/grok-4-1-fast-reasoning |
319
+ Since ClawRouter v0.12.242, Auto uses the deterministic **Router v3.4 portfolio
320
+ strategy**: it classifies the task shape locally across
321
+ <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions
322
+ (token count, code presence, reasoning markers, technical/creative terms,
323
+ agentic patterns, …), enforces tool / vision / structured-output / context
324
+ constraints as **hard filters**, then ranks an ordered candidate portfolio.
325
+ The winner becomes `routing.model`; the rest surface as `routing.candidates`
326
+ and feed SmartChat's transient-error fallback chain. Routing stays 100% local
327
+ and deterministic — <1ms, no extra model call, no network hop.
328
+
329
+ Classification still maps to one of four tiers (`routing.tier`). Each
330
+ tier × profile has a designated primary (what the rules strategy —
331
+ `routing.method: 'rules'`, the rollback lever — routes to directly, and what
332
+ anchors the portfolio's candidate pool):
333
+
334
+ | Tier | Example Tasks | ECO | AUTO | PREMIUM |
335
+ |------|---------------|-----|------|---------|
336
+ | SIMPLE | "What is 2+2?", definitions | free/gpt-oss-120b † (**FREE**) | gemini-2.5-flash ($0.30/$2.50) | kimi-k2.7 † ($0.95/$4.00) |
337
+ | MEDIUM | Code snippets, explanations | gemini-3.1-flash-lite ($0.25/$1.50) | kimi-k2.7 ($0.95/$4.00) | gpt-5.3-codex ($1.75/$14.00) |
338
+ | COMPLEX | Architecture, long documents | gemini-3.1-flash-lite ($0.25/$1.50) | gemini-3.1-pro ($2/$12) | claude-fable-5 ($10/$50) |
339
+ | REASONING | Proofs, multi-step reasoning | grok-4-1-fast-reasoning † ($0.20/$0.50) | grok-4-1-fast-reasoning † ($0.20/$0.50) | claude-sonnet-4.6 ($3/$15) |
340
+
341
+ † Withheld from `/v1/models` — the router still calls it by direct ID, but you
342
+ will not find it on the public pricing page. The published savings claim is
343
+ priced on visible models only.
344
+
345
+ This table mirrors ClawRouter's tier configs at the version this SDK pins;
346
+ the [ClawRouter README](https://github.com/BlockRunAI/ClawRouter#how-it-works)
347
+ is the live source of truth as models and prices move.
348
+
349
+ ### Routing Metadata Reference
350
+
351
+ Every `smartChat()` result carries the full decision on `result.routing`
352
+ (type `RoutingDecision`) — enough to log, audit, or replay why a model was
353
+ picked:
354
+
355
+ | Field | Description |
356
+ |-------|-------------|
357
+ | `model` | Selected model id (same as `result.model`) |
358
+ | `method` | `'portfolio'` (the Auto default), `'rules'` (rollback strategy), or `'llm'` |
359
+ | `tier` | Task tier: `'SIMPLE'`, `'MEDIUM'`, `'COMPLEX'`, or `'REASONING'` |
360
+ | `taskType` | Portfolio task classification: `'chat'`, `'extraction'`, `'code_edit'`, `'code_agent'`, `'tool_agent'`, `'debug'`, `'reasoning'`, `'reasoning_math'`, `'long_context'`, `'vision'`, … |
361
+ | `candidates` | Ordered, capability-eligible models ranked by the portfolio router; the first entry is `model` |
362
+ | `candidateScores` | Per-candidate score breakdown (`quality` / `cost` / `speed` / `reliability`), ordered with `candidates` |
363
+ | `fallbacks` | The chain `chat()` walks on transient errors (timeout / network / 429 / 5xx) — `candidates` minus the primary, with ClawRouter's proxy-namespace `free/*` ids mapped to their `nvidia/*` gateway ids (SDK-computed) |
364
+ | `savings` | 0–1 fraction saved vs the premium baseline |
365
+ | `costEstimate` / `baselineCost` | Estimated cost of the pick vs that baseline, in USD |
366
+ | `confidence` | Sigmoid-calibrated classifier confidence, 0–1 |
367
+ | `routerVersion` | `'v3-portfolio'` or `'v2-rules'` |
368
+ | `profile` | Routing profile applied: `'auto'`, `'eco'`, `'premium'`, or `'agentic'` |
369
+ | `reasoning` | Human-readable explanation of the decision |
370
+ | `tierConfigs` | The tier → primary/fallback map the decision was made against |
371
+
372
+ ### TypeScript Types
373
+
374
+ `RoutingDecision`, `RoutingProfile`, `RoutingTier`, `RoutingTaskType`, and
375
+ `RoutingTierConfig` are exported from `@blockrun/llm`. They are derived from
376
+ [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) — the
377
+ routing engine ClawRouter bundles — pinned to the exact commit ClawRouter's
378
+ published build inlines, and shipped **inlined in this SDK's declaration
379
+ files**. You do not need to install `@blockrun/clawrouter` (or router-core,
380
+ which is not on npm) for your project to typecheck against these types; the
381
+ runtime package is only needed to actually call `smartChat()`.
382
+
383
+ ### Going Deeper
384
+
385
+ - [ClawRouter](https://github.com/BlockRunAI/ClawRouter) — the router itself: OpenClaw plugin, standalone proxy for Cursor / continue.dev / any OpenAI-compatible client, Telegram integration
386
+ - [Routing profiles in depth](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/routing-profiles.md) — ECO / AUTO / PREMIUM details
387
+ - [How the routing engine works](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/smart-llm-router-14-dimension-classifier.md) — the classifier, dimension by dimension
388
+ - [Router benchmark](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/llm-router-benchmark-46-models-sub-1ms-routing.md) — sub-1ms routing across the catalog
389
+ - [ClawRouter vs OpenRouter](https://github.com/BlockRunAI/ClawRouter/blob/main/docs/clawrouter-vs-openrouter-llm-routing-comparison.md) — head-to-head comparison
390
+ - [`@blockrun/router-core`](https://github.com/BlockRunAI/router-core) — the deterministic routing engine both share
314
391
 
315
392
  ## Available Models
316
393
 
@@ -873,7 +950,7 @@ const response2 = await client.chat('anthropic/claude-sonnet-4', 'Write a haiku'
873
950
 
874
951
  ### Smart Routing (ClawRouter)
875
952
 
876
- Save up to <!-- br:savings.autoVsBaselinePct -->87<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. ClawRouter uses a <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions -->-dimension rule-based scoring algorithm to select the cheapest model that can handle your request (<1ms, 100% local).
953
+ Save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on inference costs with intelligent model routing. ClawRouter's deterministic portfolio router (v3.4, default since ClawRouter v0.12.242) classifies each request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions, applies hard capability filters, and ranks the cheapest capable models (<1ms, 100% local). Requires the optional peer dependency: `npm install @blockrun/clawrouter`.
877
954
 
878
955
  ```typescript
879
956
  import { LLMClient } from '@blockrun/llm';
@@ -885,7 +962,7 @@ const result = await client.smartChat('What is 2+2?');
885
962
  console.log(result.response); // '4'
886
963
  console.log(result.model); // 'google/gemini-2.5-flash'
887
964
  console.log(result.routing.tier); // 'SIMPLE'
888
- console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 87%'
965
+ console.log(`Saved ${(result.routing.savings * 100).toFixed(0)}%`); // 'Saved 88%'
889
966
 
890
967
  // Routing profiles
891
968
  const free = await client.smartChat('Hello!', { routingProfile: 'free' }); // Zero cost
@@ -1442,7 +1519,7 @@ The `AnthropicClient` wraps the official `@anthropic-ai/sdk` with a custom fetch
1442
1519
  When you make an API call, the SDK automatically handles x402 payment. It signs a USDC transaction locally using your wallet private key (which never leaves your machine), and includes the payment proof in the request header. Settlement is non-custodial and instant on Base or Solana.
1443
1520
 
1444
1521
  ### What is smart routing / ClawRouter?
1445
- ClawRouter is a built-in smart routing engine that analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. It can save up to <!-- br:savings.autoVsBaselinePct -->87<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
1522
+ ClawRouter is the SDK's smart routing engine, shipped as the optional `@blockrun/clawrouter` peer dependency (`npm install @blockrun/clawrouter` — only `smartChat()` needs it). It analyzes your request across <!-- br:clawrouter.dimensions -->15<!-- /br:clawrouter.dimensions --> dimensions and automatically picks the cheapest model capable of handling it. Routing happens locally in under 1ms. It can save up to <!-- br:savings.autoVsBaselinePct -->88<!-- /br:savings.autoVsBaselinePct -->% on LLM costs compared to using premium models for every request.
1446
1523
 
1447
1524
  ### Does it support streaming?
1448
1525
  Yes — as of v1.6.1. Use `client.chatCompletionStream()` for native streaming or `stream: true` in the OpenAI-compatible client. Payment is handled automatically: the SDK signs USDC payment before streaming begins, and caches payment requirements per model so subsequent calls skip the 402 round-trip (~200ms faster).
package/dist/index.cjs CHANGED
@@ -608,14 +608,14 @@ function getCostSummary() {
608
608
  }
609
609
 
610
610
  // src/version.ts
611
- var SDK_VERSION = "3.10.0";
611
+ var SDK_VERSION = "3.11.0";
612
612
  var USER_AGENT = `blockrun-ts/${SDK_VERSION}`;
613
613
 
614
614
  // src/client.ts
615
615
  function isTransientError(err) {
616
616
  if (err instanceof PaymentError) return false;
617
617
  if (err instanceof APIError) {
618
- return [502, 503, 504, 522, 524].includes(err.statusCode);
618
+ return [429, 502, 503, 504, 522, 524].includes(err.statusCode);
619
619
  }
620
620
  if (err instanceof Error) {
621
621
  if (err.name === "AbortError") return true;
@@ -742,8 +742,11 @@ var LLMClient = class _LLMClient {
742
742
  /**
743
743
  * Smart chat with automatic model routing.
744
744
  *
745
- * Uses ClawRouter's 14-dimension rule-based scoring algorithm (<1ms, 100% local)
746
- * to select the cheapest model that can handle your request.
745
+ * Uses ClawRouter's deterministic portfolio router (Router v3.4, the Auto
746
+ * default since v0.12.242): it classifies the task shape locally (<1ms, no
747
+ * extra model call), enforces capability constraints as hard filters, and
748
+ * ranks an ordered candidate portfolio — the cheapest model that can handle
749
+ * the request wins, and the rest become the transient-error fallback chain.
747
750
  *
748
751
  * @param prompt - User message
749
752
  * @param options - Optional chat and routing parameters
@@ -753,7 +756,8 @@ var LLMClient = class _LLMClient {
753
756
  * ```ts
754
757
  * const result = await client.smartChat('What is 2+2?');
755
758
  * console.log(result.response); // '4'
756
- * console.log(result.model); // 'google/gemini-2.5-flash-lite'
759
+ * console.log(result.model); // 'google/gemini-3.5-flash'
760
+ * console.log(result.routing.method); // 'portfolio'
757
761
  * console.log(result.routing.savings); // 0.78 (78% savings)
758
762
  * ```
759
763
  *
@@ -785,11 +789,15 @@ var LLMClient = class _LLMClient {
785
789
  routingProfile: options?.routingProfile
786
790
  });
787
791
  const tierConfigs = decision.tierConfigs ?? DEFAULT_ROUTING_CONFIG.tiers;
788
- const fullChain = getFallbackChain(decision.tier, tierConfigs);
789
- const fallbacks = fullChain.filter(
790
- (id) => id !== decision.model && modelPricing.has(id)
791
- );
792
- const response = await this.chat(decision.model, prompt, {
792
+ const ranked = decision.candidates?.length ? decision.candidates : [decision.model, ...getFallbackChain(decision.tier, tierConfigs)];
793
+ const callable = [];
794
+ for (const id of ranked) {
795
+ const resolved = !id.startsWith("free/") ? id : modelPricing.has(`nvidia/${id.slice(5)}`) ? `nvidia/${id.slice(5)}` : null;
796
+ if (resolved && !callable.includes(resolved)) callable.push(resolved);
797
+ }
798
+ const primary = callable[0] ?? decision.model;
799
+ const fallbacks = callable.slice(1);
800
+ const response = await this.chat(primary, prompt, {
793
801
  system: options?.system,
794
802
  maxTokens: options?.maxTokens,
795
803
  temperature: options?.temperature,
@@ -802,8 +810,8 @@ var LLMClient = class _LLMClient {
802
810
  });
803
811
  return {
804
812
  response,
805
- model: decision.model,
806
- routing: { ...decision, fallbacks }
813
+ model: primary,
814
+ routing: { ...decision, model: primary, fallbacks }
807
815
  };
808
816
  }
809
817
  /**
package/dist/index.d.cts CHANGED
@@ -1,8 +1,183 @@
1
1
  import * as _anthropic_ai_sdk from '@anthropic-ai/sdk';
2
2
 
3
+ /**
4
+ * Smart Router Types
5
+ *
6
+ * Four classification tiers — REASONING is distinct from COMPLEX because
7
+ * reasoning tasks need different models (o3, gemini-pro) than general
8
+ * complex tasks (gpt-4o, sonnet-4).
9
+ *
10
+ * Scoring uses weighted float dimensions with sigmoid confidence calibration.
11
+ */
12
+ type Tier = "SIMPLE" | "MEDIUM" | "COMPLEX" | "REASONING";
13
+ type RoutingDecision$1 = {
14
+ model: string;
15
+ tier: Tier;
16
+ confidence: number;
17
+ method: "rules" | "llm" | "portfolio";
18
+ reasoning: string;
19
+ costEstimate: number;
20
+ baselineCost: number;
21
+ savings: number;
22
+ agenticScore?: number;
23
+ /** Which tier configs were used (auto/eco/premium/agentic) — avoids re-derivation in the host */
24
+ tierConfigs?: Record<Tier, TierConfig>;
25
+ /** Which routing profile was applied */
26
+ profile?: "auto" | "eco" | "premium" | "agentic";
27
+ /** Ordered, capability-eligible candidates. The first entry is `model`. */
28
+ candidates?: string[];
29
+ /** Explainable request classification used by the portfolio router. */
30
+ taskType?: TaskType;
31
+ /** Router implementation that made the selection. */
32
+ routerVersion?: "v2-rules" | "v3-portfolio";
33
+ /** Explainable local portfolio score breakdown, ordered with `candidates`. */
34
+ candidateScores?: Array<{
35
+ model: string;
36
+ score: number;
37
+ quality: number;
38
+ cost: number;
39
+ speed: number;
40
+ reliability: number;
41
+ }>;
42
+ };
43
+ type TaskType = "chat" | "extraction" | "code_edit" | "code_agent" | "tool_agent" | "tool_agent_parallel" | "debug" | "reasoning" | "reasoning_mcq" | "reasoning_math" | "long_context" | "vision";
44
+ type TierConfig = {
45
+ primary: string;
46
+ fallback: string[];
47
+ };
48
+ type ScoringConfig = {
49
+ tokenCountThresholds: {
50
+ simple: number;
51
+ complex: number;
52
+ };
53
+ codeKeywords: string[];
54
+ reasoningKeywords: string[];
55
+ simpleKeywords: string[];
56
+ technicalKeywords: string[];
57
+ creativeKeywords: string[];
58
+ imperativeVerbs: string[];
59
+ constraintIndicators: string[];
60
+ outputFormatKeywords: string[];
61
+ referenceKeywords: string[];
62
+ negationKeywords: string[];
63
+ domainSpecificKeywords: string[];
64
+ agenticTaskKeywords: string[];
65
+ dimensionWeights: Record<string, number>;
66
+ tierBoundaries: {
67
+ simpleMedium: number;
68
+ mediumComplex: number;
69
+ complexReasoning: number;
70
+ };
71
+ confidenceSteepness: number;
72
+ confidenceThreshold: number;
73
+ };
74
+ type ClassifierConfig = {
75
+ llmModel: string;
76
+ llmMaxTokens: number;
77
+ llmTemperature: number;
78
+ promptTruncationChars: number;
79
+ cacheTtlMs: number;
80
+ };
81
+ type OverridesConfig = {
82
+ maxTokensForceComplex: number;
83
+ structuredOutputMinTier: Tier;
84
+ ambiguousDefaultTier: Tier;
85
+ /**
86
+ * When enabled, prefer models optimized for agentic workflows.
87
+ * Agentic models continue autonomously with multi-step tasks
88
+ * instead of stopping and waiting for user input.
89
+ */
90
+ agenticMode?: boolean;
91
+ };
92
+ /**
93
+ * Time-windowed promotion that temporarily overrides tier routing.
94
+ * Active promotions are auto-applied; expired ones are ignored at runtime.
95
+ */
96
+ type Promotion = {
97
+ /** Human-readable label (e.g. "GLM-5 Launch Promo") */
98
+ name: string;
99
+ /** ISO date string, promotion starts (inclusive). e.g. "2026-04-01" */
100
+ startDate: string;
101
+ /** ISO date string, promotion ends (exclusive). e.g. "2026-04-15" */
102
+ endDate: string;
103
+ /** Partial tier overrides — merged into the active tier configs (primary/fallback) */
104
+ tierOverrides: Partial<Record<Tier, Partial<TierConfig>>>;
105
+ /** Which profiles this applies to. Default: all profiles. */
106
+ profiles?: Array<"auto" | "eco" | "premium" | "agentic">;
107
+ };
108
+ type RoutingConfig = {
109
+ version: string;
110
+ /** Enables a one-line rollback to the established V2 rules selector. */
111
+ strategy?: "rules" | "portfolio";
112
+ /**
113
+ * Locally recompute a comparison strategy without changing the model that
114
+ * actually serves the request. The host emits only decision metadata via
115
+ * `onShadowRouted`; it never persists prompt content or makes a second call.
116
+ */
117
+ shadow?: {
118
+ strategy: "rules" | "portfolio";
119
+ sampleRate?: number;
120
+ };
121
+ /** Calibratable local portfolio scoring weights; values are relative, not probabilities. */
122
+ portfolio?: {
123
+ auto: {
124
+ quality: number;
125
+ capability: number;
126
+ cost: number;
127
+ speed: number;
128
+ reliability: number;
129
+ legacy: number;
130
+ };
131
+ eco: {
132
+ quality: number;
133
+ capability: number;
134
+ cost: number;
135
+ speed: number;
136
+ reliability: number;
137
+ legacy: number;
138
+ };
139
+ premium: {
140
+ quality: number;
141
+ capability: number;
142
+ cost: number;
143
+ speed: number;
144
+ reliability: number;
145
+ legacy: number;
146
+ };
147
+ highStakesBoost: {
148
+ quality: number;
149
+ reliability: number;
150
+ };
151
+ latencySensitiveSpeedBoost: number;
152
+ /** A candidate materially below the best task affinity cannot win on cost alone. */
153
+ affinityFloorGap: {
154
+ auto: number;
155
+ eco: number;
156
+ premium: number;
157
+ };
158
+ };
159
+ classifier: ClassifierConfig;
160
+ scoring: ScoringConfig;
161
+ tiers: Record<Tier, TierConfig>;
162
+ /**
163
+ * Tier configs for agentic mode — models that excel at multi-step tasks.
164
+ * Set to `null` to disable agentic tier selection entirely (forces all
165
+ * requests through `tiers`, even when tools are present in the request).
166
+ */
167
+ agenticTiers?: Record<Tier, TierConfig> | null;
168
+ /** Tier configs for eco profile — ultra cost-optimized (blockrun/eco). `null` falls back to `tiers`. */
169
+ ecoTiers?: Record<Tier, TierConfig> | null;
170
+ /** Tier configs for premium profile — best quality (blockrun/premium). `null` falls back to `tiers`. */
171
+ premiumTiers?: Record<Tier, TierConfig> | null;
172
+ /** Time-windowed promotions that temporarily override tier routing */
173
+ promotions?: Promotion[];
174
+ overrides: OverridesConfig;
175
+ };
176
+
3
177
  /**
4
178
  * Type definitions for BlockRun LLM SDK
5
179
  */
180
+
6
181
  interface FunctionDefinition {
7
182
  name: string;
8
183
  description?: string;
@@ -313,24 +488,21 @@ interface ChatCompletionOptions {
313
488
  fallbackModels?: string[];
314
489
  }
315
490
  type RoutingProfile = "eco" | "auto" | "premium";
316
- type RoutingTier = "SIMPLE" | "MEDIUM" | "COMPLEX" | "REASONING";
317
- interface RoutingDecision {
318
- model: string;
319
- tier: RoutingTier;
320
- confidence: number;
321
- method: "rules" | "llm";
322
- reasoning: string;
323
- costEstimate: number;
324
- baselineCost: number;
325
- savings: number;
326
- /** Routing profile applied by clawrouter (may include "agentic" on gateway responses). */
327
- profile?: RoutingProfile | "agentic";
328
- /** Score used when agentic routing is active. */
329
- agenticScore?: number;
491
+ type RoutingTier = Tier;
492
+ /**
493
+ * Request classification produced by the portfolio router (ClawRouter
494
+ * v0.12.242+, Router v3.4). Present on decisions with `method: "portfolio"`.
495
+ */
496
+ type RoutingTaskType = TaskType;
497
+ /** Primary model + ordered fallbacks for one routing tier. */
498
+ type RoutingTierConfig = RoutingConfig["tiers"][Tier];
499
+ interface RoutingDecision extends RoutingDecision$1 {
330
500
  /**
331
- * Remaining tier models with known pricing, in fallback order. `chat()`
332
- * walks this list when the primary model hits a transient error
333
- * (timeout, network, 5xx). Excludes the primary itself.
501
+ * Remaining gateway-callable models, in the router's ranked order.
502
+ * `chat()` walks this list when the primary model hits a transient error
503
+ * (timeout, network, 429, 5xx). Excludes the primary itself. ClawRouter's
504
+ * proxy-namespace `free/*` ids appear here as their `nvidia/*` gateway
505
+ * ids (or are dropped when the mapping is proxy-only).
334
506
  */
335
507
  fallbacks?: string[];
336
508
  }
@@ -1023,8 +1195,11 @@ declare class LLMClient {
1023
1195
  /**
1024
1196
  * Smart chat with automatic model routing.
1025
1197
  *
1026
- * Uses ClawRouter's 14-dimension rule-based scoring algorithm (<1ms, 100% local)
1027
- * to select the cheapest model that can handle your request.
1198
+ * Uses ClawRouter's deterministic portfolio router (Router v3.4, the Auto
1199
+ * default since v0.12.242): it classifies the task shape locally (<1ms, no
1200
+ * extra model call), enforces capability constraints as hard filters, and
1201
+ * ranks an ordered candidate portfolio — the cheapest model that can handle
1202
+ * the request wins, and the rest become the transient-error fallback chain.
1028
1203
  *
1029
1204
  * @param prompt - User message
1030
1205
  * @param options - Optional chat and routing parameters
@@ -1034,7 +1209,8 @@ declare class LLMClient {
1034
1209
  * ```ts
1035
1210
  * const result = await client.smartChat('What is 2+2?');
1036
1211
  * console.log(result.response); // '4'
1037
- * console.log(result.model); // 'google/gemini-2.5-flash-lite'
1212
+ * console.log(result.model); // 'google/gemini-3.5-flash'
1213
+ * console.log(result.routing.method); // 'portfolio'
1038
1214
  * console.log(result.routing.savings); // 0.78 (78% savings)
1039
1215
  * ```
1040
1216
  *
@@ -3238,4 +3414,4 @@ declare function validateTemperature(temperature?: number): void;
3238
3414
  */
3239
3415
  declare function validateTopP(topP?: number): void;
3240
3416
 
3241
- export { APIError, AnthropicClient, type AudioModel, type AudioTrack, BASE_CHAIN_ID, type BarResolution, type BlockRunAnthropicOptions, BlockrunClient, type BlockrunClientOptions, BlockrunError, type CallInitiatedResponse, type CallModel, type CallOptions, type CallStatusResponse, type ChatChoice, type ChatCompletionOptions, type ChatMessage, type ChatOptions, type ChatResponse, type ChatResponseWithCost, type ChatUsage, type CostEntry, type CostEstimate, type CreatePaymentOptions, type FunctionCall, type FunctionDefinition, type HistoryOptions, ImageClient, type ImageClientOptions, type ImageData, type ImageEditOptions, type ImageGenerateOptions, type ImageModel, type ImageResponse, KNOWN_PROVIDERS, LLMClient, type LLMClientOptions, type ListOptions, type MarketSession, type Model, MusicClient, type MusicClientOptions, type MusicGenerateOptions, type MusicResponse, NETWORK_ALIASES, type NewsSearchSource, OpenAI, type OpenAIChatCompletionChoice, type OpenAIChatCompletionChunk, type OpenAIChatCompletionParams, type OpenAIChatCompletionResponse, type OpenAIClientOptions, PHONE_PRICES, PORTRAIT_ENROLLMENT_PRICE_USD, PaymentError, type PaymentLinks, type PhoneBuyOptions, type PhoneBuyResponse, PhoneClient, type PhoneClientOptions, type PhoneListResponse, type PhoneLookupResponse, type PhoneNumberRecord, type PhoneReleaseResponse, type PhoneRenewResponse, type PollOptions, PortraitClient, type PortraitClientOptions, type PortraitEnrollOptions, type PortraitEnrollResponse, type PriceBar, type PriceCategory, PriceClient, type PriceClientOptions, type PriceHistoryResponse, type PriceOptions, type PricePoint, RPC_PRICE_USD, type ResponseFormat, RetiredEndpointError, type RoutingDecision, type RoutingProfile, type RoutingTier, type RpcBatchRequest, RpcClient, type RpcClientOptions, type RpcError, type RpcNetwork, type RpcResponse, type RssSearchSource, SOLANA_NETWORK, SOLANA_WALLET_FILE as SOLANA_WALLET_FILE_PATH, SUPPORTED_NETWORKS, SearchClient, type SearchClientOptions, type SearchOptions, type SearchParameters, type SearchResult, type SearchSource, type SearchUsage, type SmartChatOptions, type SmartChatResponse, SolanaLLMClient, type SolanaLLMClientOptions, type SolanaWalletInfo, type SoundEffectOptions, type SpeechAudio, SpeechClient, type SpeechClientOptions, type SpeechGenerateOptions, type SpeechModel, type SpeechResponse, type SpeechVoice, type Spending, type SpendingReport, type StockMarket, SurfClient, type SurfClientOptions, type SymbolListResponse, type Tool, type ToolCall, type ToolChoice, USDC_BASE, USDC_BASE_CONTRACT, USDC_SOLANA, VideoClient, type VideoClientOptions, type VideoClip, type VideoGenerateOptions, type VideoModel, type VideoResponse, VoiceClient, type VoiceClientOptions, type VoiceInfo, type VoicePreset, WALLET_DIR_PATH, WALLET_FILE_PATH, type WalletInfo, type WebSearchSource, type XSearchSource, clearCache, createPaymentPayload, createSolanaPaymentPayload, createSolanaWallet, createWallet, LLMClient as default, extractPaymentDetails, formatFundingMessageCompact, formatNeedsFundingMessage, formatSolanaWalletMigrationNotice, formatWalletCreatedMessage, formatWalletMigrationNotice, getCached, getCachedByRequest, getCostLogSummary, getCostSummary, getEip681Uri, getOrCreateSolanaWallet, getOrCreateWallet, getPaymentLinks, getWalletAddress, importSolanaWallet, importWallet, listDiscoveredSolanaWallets, listDiscoveredWallets, loadSolanaWallet, loadWallet, logCost, parsePaymentRequired, saveSolanaWallet, saveToCache, saveWallet, scanSolanaWallets, scanWallets, setCache, setupAgentSolanaWallet, setupAgentWallet, solanaClient, solanaKeyToBytes, solanaPublicKey, status, validateMaxTokens, validateModel, validateTemperature, validateTopP };
3417
+ export { APIError, AnthropicClient, type AudioModel, type AudioTrack, BASE_CHAIN_ID, type BarResolution, type BlockRunAnthropicOptions, BlockrunClient, type BlockrunClientOptions, BlockrunError, type CallInitiatedResponse, type CallModel, type CallOptions, type CallStatusResponse, type ChatChoice, type ChatCompletionOptions, type ChatMessage, type ChatOptions, type ChatResponse, type ChatResponseWithCost, type ChatUsage, type CostEntry, type CostEstimate, type CreatePaymentOptions, type FunctionCall, type FunctionDefinition, type HistoryOptions, ImageClient, type ImageClientOptions, type ImageData, type ImageEditOptions, type ImageGenerateOptions, type ImageModel, type ImageResponse, KNOWN_PROVIDERS, LLMClient, type LLMClientOptions, type ListOptions, type MarketSession, type Model, MusicClient, type MusicClientOptions, type MusicGenerateOptions, type MusicResponse, NETWORK_ALIASES, type NewsSearchSource, OpenAI, type OpenAIChatCompletionChoice, type OpenAIChatCompletionChunk, type OpenAIChatCompletionParams, type OpenAIChatCompletionResponse, type OpenAIClientOptions, PHONE_PRICES, PORTRAIT_ENROLLMENT_PRICE_USD, PaymentError, type PaymentLinks, type PhoneBuyOptions, type PhoneBuyResponse, PhoneClient, type PhoneClientOptions, type PhoneListResponse, type PhoneLookupResponse, type PhoneNumberRecord, type PhoneReleaseResponse, type PhoneRenewResponse, type PollOptions, PortraitClient, type PortraitClientOptions, type PortraitEnrollOptions, type PortraitEnrollResponse, type PriceBar, type PriceCategory, PriceClient, type PriceClientOptions, type PriceHistoryResponse, type PriceOptions, type PricePoint, RPC_PRICE_USD, type ResponseFormat, RetiredEndpointError, type RoutingDecision, type RoutingProfile, type RoutingTaskType, type RoutingTier, type RoutingTierConfig, type RpcBatchRequest, RpcClient, type RpcClientOptions, type RpcError, type RpcNetwork, type RpcResponse, type RssSearchSource, SOLANA_NETWORK, SOLANA_WALLET_FILE as SOLANA_WALLET_FILE_PATH, SUPPORTED_NETWORKS, SearchClient, type SearchClientOptions, type SearchOptions, type SearchParameters, type SearchResult, type SearchSource, type SearchUsage, type SmartChatOptions, type SmartChatResponse, SolanaLLMClient, type SolanaLLMClientOptions, type SolanaWalletInfo, type SoundEffectOptions, type SpeechAudio, SpeechClient, type SpeechClientOptions, type SpeechGenerateOptions, type SpeechModel, type SpeechResponse, type SpeechVoice, type Spending, type SpendingReport, type StockMarket, SurfClient, type SurfClientOptions, type SymbolListResponse, type Tool, type ToolCall, type ToolChoice, USDC_BASE, USDC_BASE_CONTRACT, USDC_SOLANA, VideoClient, type VideoClientOptions, type VideoClip, type VideoGenerateOptions, type VideoModel, type VideoResponse, VoiceClient, type VoiceClientOptions, type VoiceInfo, type VoicePreset, WALLET_DIR_PATH, WALLET_FILE_PATH, type WalletInfo, type WebSearchSource, type XSearchSource, clearCache, createPaymentPayload, createSolanaPaymentPayload, createSolanaWallet, createWallet, LLMClient as default, extractPaymentDetails, formatFundingMessageCompact, formatNeedsFundingMessage, formatSolanaWalletMigrationNotice, formatWalletCreatedMessage, formatWalletMigrationNotice, getCached, getCachedByRequest, getCostLogSummary, getCostSummary, getEip681Uri, getOrCreateSolanaWallet, getOrCreateWallet, getPaymentLinks, getWalletAddress, importSolanaWallet, importWallet, listDiscoveredSolanaWallets, listDiscoveredWallets, loadSolanaWallet, loadWallet, logCost, parsePaymentRequired, saveSolanaWallet, saveToCache, saveWallet, scanSolanaWallets, scanWallets, setCache, setupAgentSolanaWallet, setupAgentWallet, solanaClient, solanaKeyToBytes, solanaPublicKey, status, validateMaxTokens, validateModel, validateTemperature, validateTopP };
package/dist/index.d.ts CHANGED
@@ -1,8 +1,183 @@
1
1
  import * as _anthropic_ai_sdk from '@anthropic-ai/sdk';
2
2
 
3
+ /**
4
+ * Smart Router Types
5
+ *
6
+ * Four classification tiers — REASONING is distinct from COMPLEX because
7
+ * reasoning tasks need different models (o3, gemini-pro) than general
8
+ * complex tasks (gpt-4o, sonnet-4).
9
+ *
10
+ * Scoring uses weighted float dimensions with sigmoid confidence calibration.
11
+ */
12
+ type Tier = "SIMPLE" | "MEDIUM" | "COMPLEX" | "REASONING";
13
+ type RoutingDecision$1 = {
14
+ model: string;
15
+ tier: Tier;
16
+ confidence: number;
17
+ method: "rules" | "llm" | "portfolio";
18
+ reasoning: string;
19
+ costEstimate: number;
20
+ baselineCost: number;
21
+ savings: number;
22
+ agenticScore?: number;
23
+ /** Which tier configs were used (auto/eco/premium/agentic) — avoids re-derivation in the host */
24
+ tierConfigs?: Record<Tier, TierConfig>;
25
+ /** Which routing profile was applied */
26
+ profile?: "auto" | "eco" | "premium" | "agentic";
27
+ /** Ordered, capability-eligible candidates. The first entry is `model`. */
28
+ candidates?: string[];
29
+ /** Explainable request classification used by the portfolio router. */
30
+ taskType?: TaskType;
31
+ /** Router implementation that made the selection. */
32
+ routerVersion?: "v2-rules" | "v3-portfolio";
33
+ /** Explainable local portfolio score breakdown, ordered with `candidates`. */
34
+ candidateScores?: Array<{
35
+ model: string;
36
+ score: number;
37
+ quality: number;
38
+ cost: number;
39
+ speed: number;
40
+ reliability: number;
41
+ }>;
42
+ };
43
+ type TaskType = "chat" | "extraction" | "code_edit" | "code_agent" | "tool_agent" | "tool_agent_parallel" | "debug" | "reasoning" | "reasoning_mcq" | "reasoning_math" | "long_context" | "vision";
44
+ type TierConfig = {
45
+ primary: string;
46
+ fallback: string[];
47
+ };
48
+ type ScoringConfig = {
49
+ tokenCountThresholds: {
50
+ simple: number;
51
+ complex: number;
52
+ };
53
+ codeKeywords: string[];
54
+ reasoningKeywords: string[];
55
+ simpleKeywords: string[];
56
+ technicalKeywords: string[];
57
+ creativeKeywords: string[];
58
+ imperativeVerbs: string[];
59
+ constraintIndicators: string[];
60
+ outputFormatKeywords: string[];
61
+ referenceKeywords: string[];
62
+ negationKeywords: string[];
63
+ domainSpecificKeywords: string[];
64
+ agenticTaskKeywords: string[];
65
+ dimensionWeights: Record<string, number>;
66
+ tierBoundaries: {
67
+ simpleMedium: number;
68
+ mediumComplex: number;
69
+ complexReasoning: number;
70
+ };
71
+ confidenceSteepness: number;
72
+ confidenceThreshold: number;
73
+ };
74
+ type ClassifierConfig = {
75
+ llmModel: string;
76
+ llmMaxTokens: number;
77
+ llmTemperature: number;
78
+ promptTruncationChars: number;
79
+ cacheTtlMs: number;
80
+ };
81
+ type OverridesConfig = {
82
+ maxTokensForceComplex: number;
83
+ structuredOutputMinTier: Tier;
84
+ ambiguousDefaultTier: Tier;
85
+ /**
86
+ * When enabled, prefer models optimized for agentic workflows.
87
+ * Agentic models continue autonomously with multi-step tasks
88
+ * instead of stopping and waiting for user input.
89
+ */
90
+ agenticMode?: boolean;
91
+ };
92
+ /**
93
+ * Time-windowed promotion that temporarily overrides tier routing.
94
+ * Active promotions are auto-applied; expired ones are ignored at runtime.
95
+ */
96
+ type Promotion = {
97
+ /** Human-readable label (e.g. "GLM-5 Launch Promo") */
98
+ name: string;
99
+ /** ISO date string, promotion starts (inclusive). e.g. "2026-04-01" */
100
+ startDate: string;
101
+ /** ISO date string, promotion ends (exclusive). e.g. "2026-04-15" */
102
+ endDate: string;
103
+ /** Partial tier overrides — merged into the active tier configs (primary/fallback) */
104
+ tierOverrides: Partial<Record<Tier, Partial<TierConfig>>>;
105
+ /** Which profiles this applies to. Default: all profiles. */
106
+ profiles?: Array<"auto" | "eco" | "premium" | "agentic">;
107
+ };
108
+ type RoutingConfig = {
109
+ version: string;
110
+ /** Enables a one-line rollback to the established V2 rules selector. */
111
+ strategy?: "rules" | "portfolio";
112
+ /**
113
+ * Locally recompute a comparison strategy without changing the model that
114
+ * actually serves the request. The host emits only decision metadata via
115
+ * `onShadowRouted`; it never persists prompt content or makes a second call.
116
+ */
117
+ shadow?: {
118
+ strategy: "rules" | "portfolio";
119
+ sampleRate?: number;
120
+ };
121
+ /** Calibratable local portfolio scoring weights; values are relative, not probabilities. */
122
+ portfolio?: {
123
+ auto: {
124
+ quality: number;
125
+ capability: number;
126
+ cost: number;
127
+ speed: number;
128
+ reliability: number;
129
+ legacy: number;
130
+ };
131
+ eco: {
132
+ quality: number;
133
+ capability: number;
134
+ cost: number;
135
+ speed: number;
136
+ reliability: number;
137
+ legacy: number;
138
+ };
139
+ premium: {
140
+ quality: number;
141
+ capability: number;
142
+ cost: number;
143
+ speed: number;
144
+ reliability: number;
145
+ legacy: number;
146
+ };
147
+ highStakesBoost: {
148
+ quality: number;
149
+ reliability: number;
150
+ };
151
+ latencySensitiveSpeedBoost: number;
152
+ /** A candidate materially below the best task affinity cannot win on cost alone. */
153
+ affinityFloorGap: {
154
+ auto: number;
155
+ eco: number;
156
+ premium: number;
157
+ };
158
+ };
159
+ classifier: ClassifierConfig;
160
+ scoring: ScoringConfig;
161
+ tiers: Record<Tier, TierConfig>;
162
+ /**
163
+ * Tier configs for agentic mode — models that excel at multi-step tasks.
164
+ * Set to `null` to disable agentic tier selection entirely (forces all
165
+ * requests through `tiers`, even when tools are present in the request).
166
+ */
167
+ agenticTiers?: Record<Tier, TierConfig> | null;
168
+ /** Tier configs for eco profile — ultra cost-optimized (blockrun/eco). `null` falls back to `tiers`. */
169
+ ecoTiers?: Record<Tier, TierConfig> | null;
170
+ /** Tier configs for premium profile — best quality (blockrun/premium). `null` falls back to `tiers`. */
171
+ premiumTiers?: Record<Tier, TierConfig> | null;
172
+ /** Time-windowed promotions that temporarily override tier routing */
173
+ promotions?: Promotion[];
174
+ overrides: OverridesConfig;
175
+ };
176
+
3
177
  /**
4
178
  * Type definitions for BlockRun LLM SDK
5
179
  */
180
+
6
181
  interface FunctionDefinition {
7
182
  name: string;
8
183
  description?: string;
@@ -313,24 +488,21 @@ interface ChatCompletionOptions {
313
488
  fallbackModels?: string[];
314
489
  }
315
490
  type RoutingProfile = "eco" | "auto" | "premium";
316
- type RoutingTier = "SIMPLE" | "MEDIUM" | "COMPLEX" | "REASONING";
317
- interface RoutingDecision {
318
- model: string;
319
- tier: RoutingTier;
320
- confidence: number;
321
- method: "rules" | "llm";
322
- reasoning: string;
323
- costEstimate: number;
324
- baselineCost: number;
325
- savings: number;
326
- /** Routing profile applied by clawrouter (may include "agentic" on gateway responses). */
327
- profile?: RoutingProfile | "agentic";
328
- /** Score used when agentic routing is active. */
329
- agenticScore?: number;
491
+ type RoutingTier = Tier;
492
+ /**
493
+ * Request classification produced by the portfolio router (ClawRouter
494
+ * v0.12.242+, Router v3.4). Present on decisions with `method: "portfolio"`.
495
+ */
496
+ type RoutingTaskType = TaskType;
497
+ /** Primary model + ordered fallbacks for one routing tier. */
498
+ type RoutingTierConfig = RoutingConfig["tiers"][Tier];
499
+ interface RoutingDecision extends RoutingDecision$1 {
330
500
  /**
331
- * Remaining tier models with known pricing, in fallback order. `chat()`
332
- * walks this list when the primary model hits a transient error
333
- * (timeout, network, 5xx). Excludes the primary itself.
501
+ * Remaining gateway-callable models, in the router's ranked order.
502
+ * `chat()` walks this list when the primary model hits a transient error
503
+ * (timeout, network, 429, 5xx). Excludes the primary itself. ClawRouter's
504
+ * proxy-namespace `free/*` ids appear here as their `nvidia/*` gateway
505
+ * ids (or are dropped when the mapping is proxy-only).
334
506
  */
335
507
  fallbacks?: string[];
336
508
  }
@@ -1023,8 +1195,11 @@ declare class LLMClient {
1023
1195
  /**
1024
1196
  * Smart chat with automatic model routing.
1025
1197
  *
1026
- * Uses ClawRouter's 14-dimension rule-based scoring algorithm (<1ms, 100% local)
1027
- * to select the cheapest model that can handle your request.
1198
+ * Uses ClawRouter's deterministic portfolio router (Router v3.4, the Auto
1199
+ * default since v0.12.242): it classifies the task shape locally (<1ms, no
1200
+ * extra model call), enforces capability constraints as hard filters, and
1201
+ * ranks an ordered candidate portfolio — the cheapest model that can handle
1202
+ * the request wins, and the rest become the transient-error fallback chain.
1028
1203
  *
1029
1204
  * @param prompt - User message
1030
1205
  * @param options - Optional chat and routing parameters
@@ -1034,7 +1209,8 @@ declare class LLMClient {
1034
1209
  * ```ts
1035
1210
  * const result = await client.smartChat('What is 2+2?');
1036
1211
  * console.log(result.response); // '4'
1037
- * console.log(result.model); // 'google/gemini-2.5-flash-lite'
1212
+ * console.log(result.model); // 'google/gemini-3.5-flash'
1213
+ * console.log(result.routing.method); // 'portfolio'
1038
1214
  * console.log(result.routing.savings); // 0.78 (78% savings)
1039
1215
  * ```
1040
1216
  *
@@ -3238,4 +3414,4 @@ declare function validateTemperature(temperature?: number): void;
3238
3414
  */
3239
3415
  declare function validateTopP(topP?: number): void;
3240
3416
 
3241
- export { APIError, AnthropicClient, type AudioModel, type AudioTrack, BASE_CHAIN_ID, type BarResolution, type BlockRunAnthropicOptions, BlockrunClient, type BlockrunClientOptions, BlockrunError, type CallInitiatedResponse, type CallModel, type CallOptions, type CallStatusResponse, type ChatChoice, type ChatCompletionOptions, type ChatMessage, type ChatOptions, type ChatResponse, type ChatResponseWithCost, type ChatUsage, type CostEntry, type CostEstimate, type CreatePaymentOptions, type FunctionCall, type FunctionDefinition, type HistoryOptions, ImageClient, type ImageClientOptions, type ImageData, type ImageEditOptions, type ImageGenerateOptions, type ImageModel, type ImageResponse, KNOWN_PROVIDERS, LLMClient, type LLMClientOptions, type ListOptions, type MarketSession, type Model, MusicClient, type MusicClientOptions, type MusicGenerateOptions, type MusicResponse, NETWORK_ALIASES, type NewsSearchSource, OpenAI, type OpenAIChatCompletionChoice, type OpenAIChatCompletionChunk, type OpenAIChatCompletionParams, type OpenAIChatCompletionResponse, type OpenAIClientOptions, PHONE_PRICES, PORTRAIT_ENROLLMENT_PRICE_USD, PaymentError, type PaymentLinks, type PhoneBuyOptions, type PhoneBuyResponse, PhoneClient, type PhoneClientOptions, type PhoneListResponse, type PhoneLookupResponse, type PhoneNumberRecord, type PhoneReleaseResponse, type PhoneRenewResponse, type PollOptions, PortraitClient, type PortraitClientOptions, type PortraitEnrollOptions, type PortraitEnrollResponse, type PriceBar, type PriceCategory, PriceClient, type PriceClientOptions, type PriceHistoryResponse, type PriceOptions, type PricePoint, RPC_PRICE_USD, type ResponseFormat, RetiredEndpointError, type RoutingDecision, type RoutingProfile, type RoutingTier, type RpcBatchRequest, RpcClient, type RpcClientOptions, type RpcError, type RpcNetwork, type RpcResponse, type RssSearchSource, SOLANA_NETWORK, SOLANA_WALLET_FILE as SOLANA_WALLET_FILE_PATH, SUPPORTED_NETWORKS, SearchClient, type SearchClientOptions, type SearchOptions, type SearchParameters, type SearchResult, type SearchSource, type SearchUsage, type SmartChatOptions, type SmartChatResponse, SolanaLLMClient, type SolanaLLMClientOptions, type SolanaWalletInfo, type SoundEffectOptions, type SpeechAudio, SpeechClient, type SpeechClientOptions, type SpeechGenerateOptions, type SpeechModel, type SpeechResponse, type SpeechVoice, type Spending, type SpendingReport, type StockMarket, SurfClient, type SurfClientOptions, type SymbolListResponse, type Tool, type ToolCall, type ToolChoice, USDC_BASE, USDC_BASE_CONTRACT, USDC_SOLANA, VideoClient, type VideoClientOptions, type VideoClip, type VideoGenerateOptions, type VideoModel, type VideoResponse, VoiceClient, type VoiceClientOptions, type VoiceInfo, type VoicePreset, WALLET_DIR_PATH, WALLET_FILE_PATH, type WalletInfo, type WebSearchSource, type XSearchSource, clearCache, createPaymentPayload, createSolanaPaymentPayload, createSolanaWallet, createWallet, LLMClient as default, extractPaymentDetails, formatFundingMessageCompact, formatNeedsFundingMessage, formatSolanaWalletMigrationNotice, formatWalletCreatedMessage, formatWalletMigrationNotice, getCached, getCachedByRequest, getCostLogSummary, getCostSummary, getEip681Uri, getOrCreateSolanaWallet, getOrCreateWallet, getPaymentLinks, getWalletAddress, importSolanaWallet, importWallet, listDiscoveredSolanaWallets, listDiscoveredWallets, loadSolanaWallet, loadWallet, logCost, parsePaymentRequired, saveSolanaWallet, saveToCache, saveWallet, scanSolanaWallets, scanWallets, setCache, setupAgentSolanaWallet, setupAgentWallet, solanaClient, solanaKeyToBytes, solanaPublicKey, status, validateMaxTokens, validateModel, validateTemperature, validateTopP };
3417
+ export { APIError, AnthropicClient, type AudioModel, type AudioTrack, BASE_CHAIN_ID, type BarResolution, type BlockRunAnthropicOptions, BlockrunClient, type BlockrunClientOptions, BlockrunError, type CallInitiatedResponse, type CallModel, type CallOptions, type CallStatusResponse, type ChatChoice, type ChatCompletionOptions, type ChatMessage, type ChatOptions, type ChatResponse, type ChatResponseWithCost, type ChatUsage, type CostEntry, type CostEstimate, type CreatePaymentOptions, type FunctionCall, type FunctionDefinition, type HistoryOptions, ImageClient, type ImageClientOptions, type ImageData, type ImageEditOptions, type ImageGenerateOptions, type ImageModel, type ImageResponse, KNOWN_PROVIDERS, LLMClient, type LLMClientOptions, type ListOptions, type MarketSession, type Model, MusicClient, type MusicClientOptions, type MusicGenerateOptions, type MusicResponse, NETWORK_ALIASES, type NewsSearchSource, OpenAI, type OpenAIChatCompletionChoice, type OpenAIChatCompletionChunk, type OpenAIChatCompletionParams, type OpenAIChatCompletionResponse, type OpenAIClientOptions, PHONE_PRICES, PORTRAIT_ENROLLMENT_PRICE_USD, PaymentError, type PaymentLinks, type PhoneBuyOptions, type PhoneBuyResponse, PhoneClient, type PhoneClientOptions, type PhoneListResponse, type PhoneLookupResponse, type PhoneNumberRecord, type PhoneReleaseResponse, type PhoneRenewResponse, type PollOptions, PortraitClient, type PortraitClientOptions, type PortraitEnrollOptions, type PortraitEnrollResponse, type PriceBar, type PriceCategory, PriceClient, type PriceClientOptions, type PriceHistoryResponse, type PriceOptions, type PricePoint, RPC_PRICE_USD, type ResponseFormat, RetiredEndpointError, type RoutingDecision, type RoutingProfile, type RoutingTaskType, type RoutingTier, type RoutingTierConfig, type RpcBatchRequest, RpcClient, type RpcClientOptions, type RpcError, type RpcNetwork, type RpcResponse, type RssSearchSource, SOLANA_NETWORK, SOLANA_WALLET_FILE as SOLANA_WALLET_FILE_PATH, SUPPORTED_NETWORKS, SearchClient, type SearchClientOptions, type SearchOptions, type SearchParameters, type SearchResult, type SearchSource, type SearchUsage, type SmartChatOptions, type SmartChatResponse, SolanaLLMClient, type SolanaLLMClientOptions, type SolanaWalletInfo, type SoundEffectOptions, type SpeechAudio, SpeechClient, type SpeechClientOptions, type SpeechGenerateOptions, type SpeechModel, type SpeechResponse, type SpeechVoice, type Spending, type SpendingReport, type StockMarket, SurfClient, type SurfClientOptions, type SymbolListResponse, type Tool, type ToolCall, type ToolChoice, USDC_BASE, USDC_BASE_CONTRACT, USDC_SOLANA, VideoClient, type VideoClientOptions, type VideoClip, type VideoGenerateOptions, type VideoModel, type VideoResponse, VoiceClient, type VoiceClientOptions, type VoiceInfo, type VoicePreset, WALLET_DIR_PATH, WALLET_FILE_PATH, type WalletInfo, type WebSearchSource, type XSearchSource, clearCache, createPaymentPayload, createSolanaPaymentPayload, createSolanaWallet, createWallet, LLMClient as default, extractPaymentDetails, formatFundingMessageCompact, formatNeedsFundingMessage, formatSolanaWalletMigrationNotice, formatWalletCreatedMessage, formatWalletMigrationNotice, getCached, getCachedByRequest, getCostLogSummary, getCostSummary, getEip681Uri, getOrCreateSolanaWallet, getOrCreateWallet, getPaymentLinks, getWalletAddress, importSolanaWallet, importWallet, listDiscoveredSolanaWallets, listDiscoveredWallets, loadSolanaWallet, loadWallet, logCost, parsePaymentRequired, saveSolanaWallet, saveToCache, saveWallet, scanSolanaWallets, scanWallets, setCache, setupAgentSolanaWallet, setupAgentWallet, solanaClient, solanaKeyToBytes, solanaPublicKey, status, validateMaxTokens, validateModel, validateTemperature, validateTopP };
package/dist/index.js CHANGED
@@ -494,14 +494,14 @@ function getCostSummary() {
494
494
  }
495
495
 
496
496
  // src/version.ts
497
- var SDK_VERSION = "3.10.0";
497
+ var SDK_VERSION = "3.11.0";
498
498
  var USER_AGENT = `blockrun-ts/${SDK_VERSION}`;
499
499
 
500
500
  // src/client.ts
501
501
  function isTransientError(err) {
502
502
  if (err instanceof PaymentError) return false;
503
503
  if (err instanceof APIError) {
504
- return [502, 503, 504, 522, 524].includes(err.statusCode);
504
+ return [429, 502, 503, 504, 522, 524].includes(err.statusCode);
505
505
  }
506
506
  if (err instanceof Error) {
507
507
  if (err.name === "AbortError") return true;
@@ -628,8 +628,11 @@ var LLMClient = class _LLMClient {
628
628
  /**
629
629
  * Smart chat with automatic model routing.
630
630
  *
631
- * Uses ClawRouter's 14-dimension rule-based scoring algorithm (<1ms, 100% local)
632
- * to select the cheapest model that can handle your request.
631
+ * Uses ClawRouter's deterministic portfolio router (Router v3.4, the Auto
632
+ * default since v0.12.242): it classifies the task shape locally (<1ms, no
633
+ * extra model call), enforces capability constraints as hard filters, and
634
+ * ranks an ordered candidate portfolio — the cheapest model that can handle
635
+ * the request wins, and the rest become the transient-error fallback chain.
633
636
  *
634
637
  * @param prompt - User message
635
638
  * @param options - Optional chat and routing parameters
@@ -639,7 +642,8 @@ var LLMClient = class _LLMClient {
639
642
  * ```ts
640
643
  * const result = await client.smartChat('What is 2+2?');
641
644
  * console.log(result.response); // '4'
642
- * console.log(result.model); // 'google/gemini-2.5-flash-lite'
645
+ * console.log(result.model); // 'google/gemini-3.5-flash'
646
+ * console.log(result.routing.method); // 'portfolio'
643
647
  * console.log(result.routing.savings); // 0.78 (78% savings)
644
648
  * ```
645
649
  *
@@ -671,11 +675,15 @@ var LLMClient = class _LLMClient {
671
675
  routingProfile: options?.routingProfile
672
676
  });
673
677
  const tierConfigs = decision.tierConfigs ?? DEFAULT_ROUTING_CONFIG.tiers;
674
- const fullChain = getFallbackChain(decision.tier, tierConfigs);
675
- const fallbacks = fullChain.filter(
676
- (id) => id !== decision.model && modelPricing.has(id)
677
- );
678
- const response = await this.chat(decision.model, prompt, {
678
+ const ranked = decision.candidates?.length ? decision.candidates : [decision.model, ...getFallbackChain(decision.tier, tierConfigs)];
679
+ const callable = [];
680
+ for (const id of ranked) {
681
+ const resolved = !id.startsWith("free/") ? id : modelPricing.has(`nvidia/${id.slice(5)}`) ? `nvidia/${id.slice(5)}` : null;
682
+ if (resolved && !callable.includes(resolved)) callable.push(resolved);
683
+ }
684
+ const primary = callable[0] ?? decision.model;
685
+ const fallbacks = callable.slice(1);
686
+ const response = await this.chat(primary, prompt, {
679
687
  system: options?.system,
680
688
  maxTokens: options?.maxTokens,
681
689
  temperature: options?.temperature,
@@ -688,8 +696,8 @@ var LLMClient = class _LLMClient {
688
696
  });
689
697
  return {
690
698
  response,
691
- model: decision.model,
692
- routing: { ...decision, fallbacks }
699
+ model: primary,
700
+ routing: { ...decision, model: primary, fallbacks }
693
701
  };
694
702
  }
695
703
  /**
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@blockrun/llm",
3
- "version": "3.10.0",
3
+ "version": "3.11.0",
4
4
  "type": "module",
5
5
  "description": "BlockRun SDK - Pay-per-request AI (LLM, Image, Video, Music, Voice) via x402 on Base and Solana",
6
6
  "main": "dist/index.cjs",
@@ -18,8 +18,8 @@
18
18
  "README.md"
19
19
  ],
20
20
  "scripts": {
21
- "build": "tsup src/index.ts --format cjs,esm --dts --clean --external @blockrun/clawrouter --external @anthropic-ai/sdk --external @solana/web3.js --external @solana/spl-token --external bs58",
22
- "dev": "tsup src/index.ts --format cjs,esm --dts --watch --external @blockrun/clawrouter --external @anthropic-ai/sdk --external @solana/web3.js --external @solana/spl-token --external bs58",
21
+ "build": "tsup",
22
+ "dev": "tsup --watch",
23
23
  "test": "vitest",
24
24
  "lint": "eslint src/",
25
25
  "typecheck": "tsc --noEmit"
@@ -55,6 +55,8 @@
55
55
  "@anthropic-ai/sdk": "^0.39.0"
56
56
  },
57
57
  "devDependencies": {
58
+ "@blockrun/clawrouter": "^0.12.244",
59
+ "@blockrun/router-core": "https://codeload.github.com/BlockRunAI/router-core/tar.gz/6a790ebec60161825bbc8c4093fd221006b4e5fb",
58
60
  "@eslint/js": "^9.39.4",
59
61
  "@types/node": "^20.19.41",
60
62
  "eslint": "^9.39.4",
@@ -75,7 +77,7 @@
75
77
  },
76
78
  "packageManager": "pnpm@9.15.4",
77
79
  "peerDependencies": {
78
- "@blockrun/clawrouter": "^0.12.222",
80
+ "@blockrun/clawrouter": "^0.12.242",
79
81
  "@solana/spl-token": "^0.4.14",
80
82
  "@solana/web3.js": "^1.98.4"
81
83
  },