pi-fireworks-provider 1.0.2 → 1.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  **31+ models through [Fireworks AI](https://fireworks.ai/)**
6
6
 
7
- _Kimi, MiniMax, GLM, DeepSeek, GPT-OSS — unified OpenAI-compatible API for [pi](https://github.com/earendil-works/pi-coding-agent)._
7
+ _Kimi, MiniMax, GLM, DeepSeek, GPT-OSS — via Fireworks AI's Anthropic Messages and OpenAI-compatible endpoints for [pi](https://github.com/earendil-works/pi-coding-agent)._
8
8
 
9
9
  [![pi extension](https://img.shields.io/badge/pi-extension-blueviolet)](https://github.com/earendil-works/pi-coding-agent)
10
10
  [![license](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)
@@ -16,7 +16,8 @@ _Kimi, MiniMax, GLM, DeepSeek, GPT-OSS — unified OpenAI-compatible API for [pi
16
16
  ## Features
17
17
 
18
18
  - **35+ AI Models** including Kimi K2.5, MiniMax M2.5, GLM 4.5/4.7/5, DeepSeek V3.1/V3.2, DeepSeek V4 Flash, and GPT-OSS
19
- - **Unified API** via Fireworks AI's OpenAI-compatible completions endpoint
19
+ - **Dual API support** via Fireworks AI's Anthropic Messages and OpenAI-compatible completions endpoints (per-model routing, matching pi core's Fireworks provider)
20
+ - **Service tiers** — toggle Fireworks `priority` vs `standard` per request on supported models (with priority pricing reflected in cost tracking), via a keybinding, `/fireworks-tier`, and a footer status area
20
21
  - **Cost Tracking** with per-model pricing for budget management
21
22
  - **Reasoning Models** support for advanced reasoning capabilities
22
23
  - **Vision Support** for image-capable models
@@ -105,6 +106,50 @@ pi
105
106
  | Qwen3 VL 30B A3B Thinking | Text + Image | 262K | 0 | Free | Free |
106
107
  *Costs are per million tokens. Prices subject to change - check [fireworks.ai](https://fireworks.ai) for current pricing.*
107
108
 
109
+ ## Service Tiers
110
+
111
+ Fireworks exposes a `service_tier` request field (`standard` | `priority`) on its chat-completions endpoint. The **priority** tier trades higher per-token pricing for higher throughput / lower latency. This is orthogonal to the `-fast`/`-turbo` router model IDs (which are separate models) — service tiers apply to the base models below.
112
+
113
+ | Model | Priority Uncached Input | Priority Cached Input | Priority Output |
114
+ | --- | --- | --- | --- |
115
+ | GLM 5.2 | $1.75/M | $0.175/M | $5.5/M |
116
+ | Kimi K2.7 Code | $1.43/M | $0.29/M | $6/M |
117
+ | Minimax M3 | $0.45/M | $0.09/M | $1.8/M |
118
+ | DeepSeek V4 Pro | $2.61/M | $0.218/M | $5.22/M |
119
+ | Kimi K2.6 | $1.5/M | $0.22/M | $6/M |
120
+ | MiniMax M2.7 | $0.45/M | $0.09/M | $1.8/M |
121
+ | GLM 5.1 | $2.1/M | $0.39/M | $6.6/M |
122
+ | GPT OSS 120B | $0.18/M | $0.018/M | $0.72/M |
123
+ | DeepSeek V4 Flash | $0.21/M | $0.045/M | $0.42/M |
124
+
125
+ *Priority pricing is roughly 1.2–1.5× the standard rate. `cacheWrite` is not tiered.*
126
+
127
+ **Switching tiers:**
128
+
129
+ - **Keybinding:** `ctrl+shift+l` (default) toggles `standard` ↔ `priority` for the active supported model. No-op with an info notice for unsupported models.
130
+ - **Command:** `/fireworks-tier standard|priority|toggle`.
131
+ - **Status area:** a dim `tier: standard` / `tier: ⚡priority` line is shown in the footer for supported models while a Fireworks model is active.
132
+
133
+ The selection is persisted per session (survives `/reload` and resume). When `priority` is active, `service_tier: "priority"` is injected into every request and finalized cost is recomputed against the priority rates above.
134
+
135
+ **Configuration** — `~/.pi/agent/extensions/fireworks.json` (created with defaults on first load):
136
+
137
+ ```json
138
+ {
139
+ "serviceTier": {
140
+ "default": "standard",
141
+ "keybinding": "ctrl+shift+l",
142
+ "display": "statusbar"
143
+ }
144
+ }
145
+ ```
146
+
147
+ - `default` — tier used until you toggle (`standard` | `priority`).
148
+ - `keybinding` — any [pi key format](https://github.com/earendil-works/pi-coding-agent/blob/main/docs/keybindings.md) (e.g. `ctrl+shift+l`, `ctrl+shift+k`). Requires `/reload` after changing. On macOS browser terminals (localterm), avoid `alt`/`ctrl+alt` (Option produces special chars) and `ctrl+shift+t/w/n/c/v` (browser/localterm tab + copy/paste shortcuts).
149
+ - `display` — `statusbar` (footer status area) or `off` (hide the tier indicator).
150
+
151
+ > **Note:** The OpenAI completions endpoint accepts `service_tier` directly (per Fireworks' API). The Anthropic Messages endpoint passes the top-level field through as an extra. If a supported Anthropic-routed model rejects it, file an issue so we can gate injection by API.
152
+
108
153
  ## Usage
109
154
 
110
155
  After loading the extension, use the `/model` command in pi to select your preferred model:
@@ -147,5 +147,72 @@
147
147
  ],
148
148
  "contextWindow": 262000,
149
149
  "maxTokens": 0
150
+ },
151
+ {
152
+ "id": "accounts/fireworks/models/qwen3p7-plus",
153
+ "name": "Qwen 3.7 Plus",
154
+ "reasoning": true,
155
+ "cost": {
156
+ "input": 0.4,
157
+ "output": 1.6,
158
+ "cacheRead": 0.08,
159
+ "cacheWrite": 0
160
+ },
161
+ "input": [
162
+ "text",
163
+ "image"
164
+ ],
165
+ "contextWindow": 262144,
166
+ "maxTokens": 65536
167
+ },
168
+ {
169
+ "id": "accounts/fireworks/routers/glm-5p2-fast",
170
+ "name": "GLM 5.2 Fast",
171
+ "reasoning": true,
172
+ "cost": {
173
+ "input": 2.1,
174
+ "output": 6.6,
175
+ "cacheRead": 0.21,
176
+ "cacheWrite": 0
177
+ },
178
+ "input": [
179
+ "text"
180
+ ],
181
+ "contextWindow": 1048575,
182
+ "maxTokens": 131072
183
+ },
184
+ {
185
+ "id": "accounts/fireworks/routers/kimi-k2p6-fast",
186
+ "name": "Kimi K2.6 Fast",
187
+ "reasoning": true,
188
+ "cost": {
189
+ "input": 2,
190
+ "output": 8,
191
+ "cacheRead": 0.3,
192
+ "cacheWrite": 0
193
+ },
194
+ "input": [
195
+ "text",
196
+ "image"
197
+ ],
198
+ "contextWindow": 262000,
199
+ "maxTokens": 262000
200
+ },
201
+ {
202
+ "id": "accounts/fireworks/routers/kimi-k2p7-code-fast",
203
+ "name": "Kimi K2.7 Code Fast",
204
+ "reasoning": true,
205
+ "cost": {
206
+ "input": 1.9,
207
+ "output": 8,
208
+ "cacheRead": 0.38,
209
+ "cacheWrite": 0
210
+ },
211
+ "input": [
212
+ "text",
213
+ "image"
214
+ ],
215
+ "contextWindow": 262000,
216
+ "maxTokens": 262000
150
217
  }
151
218
  ]
package/index.ts CHANGED
@@ -34,10 +34,15 @@ import path from "path";
34
34
 
35
35
  // ─── Types ────────────────────────────────────────────────────────────────────
36
36
 
37
+ type FireworksApi = "anthropic-messages" | "openai-completions";
38
+
37
39
  interface JsonModel {
38
40
  id: string;
39
41
  name: string;
42
+ api?: FireworksApi;
43
+ baseUrl?: string;
40
44
  reasoning: boolean;
45
+ thinkingLevelMap?: Record<string, string | null>;
41
46
  input: string[];
42
47
  cost: {
43
48
  input: number;
@@ -47,18 +52,17 @@ interface JsonModel {
47
52
  };
48
53
  contextWindow: number;
49
54
  maxTokens: number;
50
- compat?: {
51
- supportsDeveloperRole?: boolean;
52
- supportsStore?: boolean;
53
- maxTokensField?: "max_completion_tokens" | "max_tokens";
54
- thinkingFormat?: "openai" | "zai" | "qwen" | "qwen-chat-template";
55
- supportsReasoningEffort?: boolean;
56
- };
55
+ // Loose on purpose: shape depends on `api` (OpenAICompletionsCompat vs
56
+ // AnthropicMessagesCompat). pi-ai validates at runtime.
57
+ compat?: Record<string, unknown>;
57
58
  }
58
59
 
59
60
  interface PatchEntry {
60
61
  name?: string;
62
+ api?: FireworksApi;
63
+ baseUrl?: string;
61
64
  reasoning?: boolean;
65
+ thinkingLevelMap?: Record<string, string | null>;
62
66
  input?: string[];
63
67
  cost?: {
64
68
  input?: number;
@@ -79,7 +83,10 @@ function applyPatch(model: JsonModel, patch: PatchEntry): JsonModel {
79
83
  const result = { ...model };
80
84
 
81
85
  if (patch.name !== undefined) result.name = patch.name;
86
+ if (patch.api !== undefined) result.api = patch.api;
87
+ if (patch.baseUrl !== undefined) result.baseUrl = patch.baseUrl;
82
88
  if (patch.reasoning !== undefined) result.reasoning = patch.reasoning;
89
+ if (patch.thinkingLevelMap !== undefined) result.thinkingLevelMap = patch.thinkingLevelMap;
83
90
  if (patch.input !== undefined) result.input = patch.input;
84
91
  if (patch.contextWindow !== undefined) result.contextWindow = patch.contextWindow;
85
92
  if (patch.maxTokens !== undefined) result.maxTokens = patch.maxTokens;
@@ -335,9 +342,176 @@ function stripAnchorBleedInPlace(obj: Record<string, unknown>): void {
335
342
  }
336
343
  }
337
344
 
345
+ // ─── Service Tier (standard / priority) ──────────────────────────────────────
346
+
347
+ // Fireworks exposes a `service_tier` request field ("standard" | "priority") on
348
+ // its chat-completions endpoint. The priority tier trades higher per-token
349
+ // pricing for higher throughput / lower latency on supported models. This is
350
+ // orthogonal to the "fast" router model IDs (e.g. routers/...-fast), which are
351
+ // separate models; priority applies to the base models below.
352
+ //
353
+ // Per-request priority pricing (USD per million tokens) from Fireworks' tier
354
+ // reference. cacheWrite is not tiered (stays 0).
355
+ const PRIORITY_PRICING: Record<string, { input: number; output: number; cacheRead: number; cacheWrite: number }> = {
356
+ "accounts/fireworks/models/glm-5p2": { input: 1.75, output: 5.5, cacheRead: 0.175, cacheWrite: 0 },
357
+ "accounts/fireworks/models/kimi-k2p7-code": { input: 1.43, output: 6, cacheRead: 0.29, cacheWrite: 0 },
358
+ "accounts/fireworks/models/minimax-m3": { input: 0.45, output: 1.8, cacheRead: 0.09, cacheWrite: 0 },
359
+ "accounts/fireworks/models/deepseek-v4-pro": { input: 2.61, output: 5.22, cacheRead: 0.218, cacheWrite: 0 },
360
+ "accounts/fireworks/models/kimi-k2p6": { input: 1.5, output: 6, cacheRead: 0.22, cacheWrite: 0 },
361
+ "accounts/fireworks/models/minimax-m2p7": { input: 0.45, output: 1.8, cacheRead: 0.09, cacheWrite: 0 },
362
+ "accounts/fireworks/models/glm-5p1": { input: 2.1, output: 6.6, cacheRead: 0.39, cacheWrite: 0 },
363
+ "accounts/fireworks/models/gpt-oss-120b": { input: 0.18, output: 0.72, cacheRead: 0.018, cacheWrite: 0 },
364
+ "accounts/fireworks/models/deepseek-v4-flash":{ input: 0.21, output: 0.42, cacheRead: 0.045, cacheWrite: 0 },
365
+ };
366
+
367
+ type ServiceTier = "standard" | "priority";
368
+
369
+ interface ServiceTierConfig {
370
+ default: ServiceTier;
371
+ keybinding: string;
372
+ display: "statusbar" | "off";
373
+ }
374
+
375
+ interface FireworksConfig {
376
+ serviceTier: ServiceTierConfig;
377
+ }
378
+
379
+ const FIREWORKS_CONFIG_PATH = path.join(getAgentDir(), "extensions", "fireworks.json");
380
+ const TIER_ENTRY_TYPE = "fireworks-service-tier";
381
+ const TIER_STATUS_KEY = "fireworks-tier";
382
+ const DEFAULT_SERVICE_TIER_CONFIG: ServiceTierConfig = {
383
+ default: "standard",
384
+ keybinding: "ctrl+shift+l",
385
+ display: "statusbar",
386
+ };
387
+ const DEFAULT_FIREWORKS_CONFIG: FireworksConfig = { serviceTier: DEFAULT_SERVICE_TIER_CONFIG };
388
+
389
+ function isValidTier(v: unknown): v is ServiceTier {
390
+ return v === "standard" || v === "priority";
391
+ }
392
+
393
+ function loadFireworksConfig(): FireworksConfig {
394
+ try {
395
+ const raw = JSON.parse(fs.readFileSync(FIREWORKS_CONFIG_PATH, "utf8"));
396
+ const st = raw?.serviceTier ?? {};
397
+ return {
398
+ serviceTier: {
399
+ default: isValidTier(st.default) ? st.default : DEFAULT_SERVICE_TIER_CONFIG.default,
400
+ keybinding: typeof st.keybinding === "string" && st.keybinding.length > 0 ? st.keybinding : DEFAULT_SERVICE_TIER_CONFIG.keybinding,
401
+ display: st.display === "off" ? "off" : "statusbar",
402
+ },
403
+ };
404
+ } catch {
405
+ // Config missing or invalid — write defaults so the user can discover it.
406
+ try {
407
+ fs.mkdirSync(path.dirname(FIREWORKS_CONFIG_PATH), { recursive: true });
408
+ fs.writeFileSync(FIREWORKS_CONFIG_PATH, JSON.stringify(DEFAULT_FIREWORKS_CONFIG, null, 2) + "\n");
409
+ } catch {
410
+ // Write failure is non-fatal — defaults still work in memory.
411
+ }
412
+ return { serviceTier: { ...DEFAULT_SERVICE_TIER_CONFIG } };
413
+ }
414
+ }
415
+
416
+ let fireworksConfig = loadFireworksConfig();
417
+
418
+ // Held so module-scope helpers (setTier) can call pi.appendEntry.
419
+ let piRef: ExtensionAPI | null = null;
420
+
421
+ function isPriorityApplicable(id: string | undefined): boolean {
422
+ return !!id && Object.prototype.hasOwnProperty.call(PRIORITY_PRICING, id);
423
+ }
424
+
425
+ // Session state: the active service tier. Replayed from session entries on
426
+ // session_start so it survives reload / resume.
427
+ let currentTier: ServiceTier = fireworksConfig.serviceTier.default;
428
+
429
+ function replayTierState(ctx: any, defaultTier: ServiceTier): void {
430
+ let tier: ServiceTier = defaultTier;
431
+ for (const entry of ctx.sessionManager.getBranch()) {
432
+ if (entry?.type === "custom" && entry.customType === TIER_ENTRY_TYPE && entry.data) {
433
+ const t = entry.data.tier;
434
+ if (isValidTier(t)) tier = t;
435
+ }
436
+ }
437
+ currentTier = tier;
438
+ }
439
+
440
+ function updateTierStatus(ctx: any): void {
441
+ if (fireworksConfig.serviceTier.display === "off") {
442
+ ctx.ui.setStatus(TIER_STATUS_KEY, undefined);
443
+ return;
444
+ }
445
+ let model: any;
446
+ try {
447
+ model = ctx.model;
448
+ } catch {
449
+ model = undefined;
450
+ }
451
+ if (!model || model.provider !== "fireworks" || !isPriorityApplicable(model.id)) {
452
+ ctx.ui.setStatus(TIER_STATUS_KEY, undefined);
453
+ return;
454
+ }
455
+ const label = currentTier === "priority" ? "tier: ⚡priority" : "tier: standard";
456
+ try {
457
+ ctx.ui.setStatus(TIER_STATUS_KEY, ctx.ui.theme.fg("dim", label));
458
+ } catch {
459
+ // setStatus / theme are no-ops without a UI runner.
460
+ }
461
+ }
462
+
463
+ function setTier(ctx: any, tier: ServiceTier): void {
464
+ currentTier = tier;
465
+ try {
466
+ piRef?.appendEntry(TIER_ENTRY_TYPE, { tier });
467
+ } catch {
468
+ // appendEntry outside a session context is non-fatal.
469
+ }
470
+ updateTierStatus(ctx);
471
+ }
472
+
473
+ function toggleTier(ctx: any): void {
474
+ let model: any;
475
+ try {
476
+ model = ctx.model;
477
+ } catch {
478
+ model = undefined;
479
+ }
480
+ if (!model || model.provider !== "fireworks") {
481
+ try { ctx.ui.notify("Fireworks service tier only applies to Fireworks models.", "info"); } catch {}
482
+ return;
483
+ }
484
+ if (!isPriorityApplicable(model.id)) {
485
+ try { ctx.ui.notify(`Service tier is not available for ${model.name || model.id}.`, "info"); } catch {}
486
+ return;
487
+ }
488
+ const next: ServiceTier = currentTier === "priority" ? "standard" : "priority";
489
+ setTier(ctx, next);
490
+ try { ctx.ui.notify(`Fireworks service tier: ${next}`, "info"); } catch {}
491
+ }
492
+
493
+ // Recompute a finalized assistant message's cost against priority pricing so
494
+ // cost tracking reflects the premium tier instead of the model's base cost.
495
+ function recomputePriorityCost(message: any): any {
496
+ const pricing = message?.model ? PRIORITY_PRICING[message.model] : undefined;
497
+ if (!pricing) return undefined;
498
+ const usage = message?.usage;
499
+ if (!usage) return undefined;
500
+ const cost = {
501
+ input: (pricing.input / 1_000_000) * (usage.input || 0),
502
+ output: (pricing.output / 1_000_000) * (usage.output || 0),
503
+ cacheRead: (pricing.cacheRead / 1_000_000) * (usage.cacheRead || 0),
504
+ cacheWrite: 0,
505
+ total: 0,
506
+ };
507
+ cost.total = cost.input + cost.output + cost.cacheRead + cost.cacheWrite;
508
+ return { ...message, usage: { ...usage, cost } };
509
+ }
510
+
338
511
  // ─── Extension Entry Point ────────────────────────────────────────────────────
339
512
 
340
513
  export default function (pi: ExtensionAPI) {
514
+ piRef = pi;
341
515
  const embeddedModels = modelsData as JsonModel[];
342
516
  const customModels = customModelsData as JsonModel[];
343
517
  const patches = patchData as PatchData;
@@ -356,6 +530,9 @@ export default function (pi: ExtensionAPI) {
356
530
  revalidateAbort?.abort();
357
531
  revalidateAbort = new AbortController();
358
532
  const signal = revalidateAbort.signal;
533
+ fireworksConfig = loadFireworksConfig();
534
+ replayTierState(ctx, fireworksConfig.serviceTier.default);
535
+ updateTierStatus(ctx);
359
536
  resolveApiKey(ctx.modelRegistry).then(() => {
360
537
  revalidateModels(cachedApiKey, embeddedModels, signal).then((freshBase) => {
361
538
  if (freshBase && !signal.aborted) {
@@ -370,8 +547,9 @@ export default function (pi: ExtensionAPI) {
370
547
  });
371
548
  });
372
549
 
373
- pi.on("session_shutdown", () => {
550
+ pi.on("session_shutdown", (_event, ctx) => {
374
551
  revalidateAbort?.abort();
552
+ try { ctx.ui.setStatus(TIER_STATUS_KEY, undefined); } catch {}
375
553
  });
376
554
 
377
555
  // Sanitize JSON Schema patterns for Kimi models before sending to Fireworks.
@@ -380,34 +558,52 @@ export default function (pi: ExtensionAPI) {
380
558
  // present. We strip anchors from simple patterns and drop patterns that
381
559
  // combine alternation with anchors entirely.
382
560
  pi.on("before_provider_request", (event, ctx) => {
383
- if (!isFireworksKimiModel(ctx.model)) return;
561
+ const model = ctx.model;
562
+ if (!model || model.provider !== "fireworks") return;
384
563
 
385
564
  const payload = event.payload as Record<string, unknown>;
386
565
  if (!payload || typeof payload !== "object") return;
387
566
 
388
567
  let modified = false;
389
568
 
390
- const tools = payload.tools;
391
- if (Array.isArray(tools)) {
392
- payload.tools = tools.map((tool: any) => {
393
- if (tool?.function?.parameters) {
394
- return {
395
- ...tool,
396
- function: {
397
- ...tool.function,
398
- parameters: sanitizeSchemaForKimi(tool.function.parameters),
399
- },
400
- };
401
- }
402
- return tool;
403
- });
569
+ // Service tier: inject `service_tier` on supported Fireworks models when
570
+ // priority is selected. Injected at the top level of both the OpenAI
571
+ // completions and Anthropic Messages request bodies (Fireworks accepts the
572
+ // field on its chat-completions endpoint; the Anthropic-compatible endpoint
573
+ // passes it through as a top-level extra).
574
+ if (currentTier === "priority" && isPriorityApplicable(model.id)) {
575
+ payload.service_tier = "priority";
404
576
  modified = true;
405
577
  }
406
578
 
407
- const responseFormat = payload.response_format as any;
408
- if (responseFormat?.json_schema?.schema) {
409
- responseFormat.json_schema.schema = sanitizeSchemaForKimi(responseFormat.json_schema.schema);
410
- modified = true;
579
+ // Kimi anchor-bleed sanitization (Kimi K2.x pattern bug). Only applies to
580
+ // Kimi models, but a single request can be both Kimi and priority-tiered.
581
+ if (isFireworksKimiModel(model)) {
582
+ const tools = payload.tools;
583
+ if (Array.isArray(tools)) {
584
+ payload.tools = tools.map((tool: any) => {
585
+ // OpenAI completions shape: tools[].function.parameters
586
+ if (tool?.function?.parameters) {
587
+ return {
588
+ ...tool,
589
+ function: { ...tool.function, parameters: sanitizeSchemaForKimi(tool.function.parameters) },
590
+ };
591
+ }
592
+ // Anthropic messages shape: tools[].input_schema
593
+ if (tool?.input_schema) {
594
+ return { ...tool, input_schema: sanitizeSchemaForKimi(tool.input_schema) };
595
+ }
596
+ return tool;
597
+ });
598
+ modified = true;
599
+ }
600
+
601
+ // OpenAI-only: the Anthropic Messages API has no response_format equivalent.
602
+ const responseFormat = payload.response_format as any;
603
+ if (responseFormat?.json_schema?.schema) {
604
+ responseFormat.json_schema.schema = sanitizeSchemaForKimi(responseFormat.json_schema.schema);
605
+ modified = true;
606
+ }
411
607
  }
412
608
 
413
609
  if (modified) {
@@ -426,4 +622,54 @@ export default function (pi: ExtensionAPI) {
426
622
  stripAnchorBleedInPlace(input);
427
623
  }
428
624
  });
625
+
626
+ // Refresh the tier status area when the active model changes.
627
+ pi.on("model_select", async (_event, ctx) => {
628
+ updateTierStatus(ctx);
629
+ });
630
+
631
+ // Recompute finalized assistant-message cost against priority pricing so
632
+ // usage/cost reflects the premium tier rather than the model's base cost.
633
+ // message_end fires for user, assistant, and toolResult messages; we only
634
+ // touch assistant messages, and only when priority is active for that model.
635
+ pi.on("message_end", async (event, _ctx) => {
636
+ if (currentTier !== "priority") return;
637
+ const message = (event as any).message;
638
+ if (!message || message.role !== "assistant") return;
639
+ const replaced = recomputePriorityCost(message);
640
+ if (replaced) {
641
+ return { message: replaced };
642
+ }
643
+ });
644
+
645
+ // Keybinding: toggle the service tier for the active Fireworks model.
646
+ pi.registerShortcut(fireworksConfig.serviceTier.keybinding, {
647
+ description: "Toggle Fireworks service tier (standard / priority)",
648
+ handler: async (ctx) => {
649
+ toggleTier(ctx);
650
+ },
651
+ });
652
+
653
+ // Command form (non-TUI use, or explicit setting): /fireworks-tier [standard|priority|toggle]
654
+ pi.registerCommand("fireworks-tier", {
655
+ description: "Set Fireworks service tier: standard | priority | toggle (default)",
656
+ handler: async (args, ctx) => {
657
+ const arg = (args || "").trim().toLowerCase();
658
+ let model: any;
659
+ try { model = ctx.model; } catch { model = undefined; }
660
+ if (!model || model.provider !== "fireworks" || !isPriorityApplicable(model.id)) {
661
+ try { ctx.ui.notify("Fireworks service tier only applies to supported Fireworks models.", "info"); } catch {}
662
+ return;
663
+ }
664
+ if (arg === "priority") {
665
+ setTier(ctx, "priority");
666
+ try { ctx.ui.notify("Fireworks service tier: priority", "info"); } catch {}
667
+ } else if (arg === "standard") {
668
+ setTier(ctx, "standard");
669
+ try { ctx.ui.notify("Fireworks service tier: standard", "info"); } catch {}
670
+ } else {
671
+ toggleTier(ctx);
672
+ }
673
+ },
674
+ });
429
675
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-fireworks-provider",
3
- "version": "1.0.2",
3
+ "version": "1.1.0",
4
4
  "description": "Fireworks AI provider extension for pi - Access Kimi, MiniMax, GLM, DeepSeek, and GPT-OSS models through the Fireworks AI API",
5
5
  "type": "module",
6
6
  "main": "index.ts",
package/patch.json CHANGED
@@ -1,157 +1,257 @@
1
1
  {
2
+ "accounts/fireworks/models/deepseek-v3p1": {
3
+ "name": "DeepSeek V3.1",
4
+ "reasoning": true,
5
+ "input": ["text"],
6
+ "cost": {
7
+ "input": 0.56,
8
+ "output": 1.68,
9
+ "cacheRead": 0,
10
+ "cacheWrite": 0
11
+ },
12
+ "contextWindow": 163840,
13
+ "maxTokens": 163840,
14
+ "compat": {
15
+ "supportsReasoningEffort": true
16
+ }
17
+ },
18
+ "accounts/fireworks/models/deepseek-v3p2": {
19
+ "name": "DeepSeek V3.2",
20
+ "reasoning": true,
21
+ "input": ["text"],
22
+ "cost": {
23
+ "input": 0.56,
24
+ "output": 1.68,
25
+ "cacheRead": 0.28,
26
+ "cacheWrite": 0
27
+ },
28
+ "contextWindow": 163840,
29
+ "maxTokens": 160000,
30
+ "compat": {
31
+ "supportsReasoningEffort": true
32
+ }
33
+ },
2
34
  "accounts/fireworks/models/deepseek-v4-flash": {
3
35
  "name": "DeepSeek V4 Flash",
36
+ "api": "anthropic-messages",
37
+ "baseUrl": "https://api.fireworks.ai/inference",
4
38
  "reasoning": true,
39
+ "input": ["text"],
5
40
  "cost": {
6
41
  "input": 0.14,
7
42
  "output": 0.28,
8
- "cacheRead": 0.03,
43
+ "cacheRead": 0.028,
9
44
  "cacheWrite": 0
10
45
  },
11
- "maxTokens": 1048576,
46
+ "contextWindow": 1000000,
47
+ "maxTokens": 384000,
12
48
  "compat": {
13
- "supportsReasoningEffort": true
49
+ "sendSessionAffinityHeaders": true,
50
+ "supportsEagerToolInputStreaming": false,
51
+ "supportsCacheControlOnTools": false,
52
+ "supportsLongCacheRetention": false
14
53
  }
15
54
  },
16
55
  "accounts/fireworks/models/deepseek-v4-pro": {
17
56
  "name": "DeepSeek V4 Pro",
57
+ "api": "anthropic-messages",
58
+ "baseUrl": "https://api.fireworks.ai/inference",
18
59
  "reasoning": true,
19
- "maxTokens": 1048576,
60
+ "input": ["text"],
61
+ "cost": {
62
+ "input": 1.74,
63
+ "output": 3.48,
64
+ "cacheRead": 0.145,
65
+ "cacheWrite": 0
66
+ },
67
+ "contextWindow": 1000000,
68
+ "maxTokens": 384000,
69
+ "compat": {
70
+ "sendSessionAffinityHeaders": true,
71
+ "supportsEagerToolInputStreaming": false,
72
+ "supportsCacheControlOnTools": false,
73
+ "supportsLongCacheRetention": false
74
+ }
75
+ },
76
+ "accounts/fireworks/models/glm-4p5": {
77
+ "name": "GLM 4.5",
78
+ "reasoning": true,
79
+ "input": ["text"],
80
+ "cost": {
81
+ "input": 0.55,
82
+ "output": 2.19,
83
+ "cacheRead": 0,
84
+ "cacheWrite": 0
85
+ },
86
+ "contextWindow": 131072,
87
+ "maxTokens": 131072,
20
88
  "compat": {
21
89
  "supportsReasoningEffort": true
22
90
  }
23
91
  },
24
- "accounts/fireworks/routers/deepseek-v4-pro": {
25
- "name": "DeepSeek V4 Pro (router)",
92
+ "accounts/fireworks/models/glm-4p5-air": {
93
+ "name": "GLM 4.5 Air",
26
94
  "reasoning": true,
95
+ "input": ["text"],
27
96
  "cost": {
28
- "input": 1.74,
29
- "output": 3.48,
30
- "cacheRead": 0.0145,
97
+ "input": 0.22,
98
+ "output": 0.88,
99
+ "cacheRead": 0,
31
100
  "cacheWrite": 0
32
101
  },
33
- "maxTokens": 1048576,
102
+ "contextWindow": 131072,
103
+ "maxTokens": 131072,
34
104
  "compat": {
35
105
  "supportsReasoningEffort": true
36
106
  }
37
107
  },
38
- "accounts/fireworks/routers/kimi-k2p5-fast": {
39
- "name": "Kimi K2.5 Fast (router)",
108
+ "accounts/fireworks/models/glm-4p7": {
109
+ "name": "GLM 4.7",
40
110
  "reasoning": true,
111
+ "input": ["text"],
41
112
  "cost": {
42
113
  "input": 0.6,
43
- "output": 3,
44
- "cacheRead": 0.1,
114
+ "output": 2.2,
115
+ "cacheRead": 0.3,
45
116
  "cacheWrite": 0
46
117
  },
47
- "input": [
48
- "text",
49
- "image",
50
- "video"
51
- ],
52
- "maxTokens": 256000,
118
+ "contextWindow": 202752,
119
+ "maxTokens": 198000,
53
120
  "compat": {
54
121
  "supportsReasoningEffort": true
55
122
  }
56
123
  },
57
- "accounts/fireworks/routers/kimi-k2p6-turbo": {
58
- "name": "Kimi K2.6 Turbo (router)",
124
+ "accounts/fireworks/models/glm-5": {
125
+ "name": "GLM 5",
59
126
  "reasoning": true,
127
+ "input": ["text"],
60
128
  "cost": {
61
- "input": 0.95,
62
- "output": 4,
63
- "cacheRead": 0.16,
129
+ "input": 1,
130
+ "output": 3.2,
131
+ "cacheRead": 0.5,
64
132
  "cacheWrite": 0
65
133
  },
66
- "input": [
67
- "text",
68
- "image"
69
- ],
70
- "maxTokens": 262144,
134
+ "contextWindow": 202752,
135
+ "maxTokens": 131072,
71
136
  "compat": {
72
137
  "supportsReasoningEffort": true
73
138
  }
74
139
  },
75
140
  "accounts/fireworks/models/glm-5p1": {
76
141
  "name": "GLM 5.1",
142
+ "api": "anthropic-messages",
143
+ "baseUrl": "https://api.fireworks.ai/inference",
77
144
  "reasoning": true,
145
+ "input": ["text"],
78
146
  "cost": {
79
147
  "input": 1.4,
80
148
  "output": 4.4,
81
149
  "cacheRead": 0.26,
82
150
  "cacheWrite": 0
83
151
  },
152
+ "contextWindow": 202800,
84
153
  "maxTokens": 131072,
85
154
  "compat": {
86
- "supportsReasoningEffort": true
155
+ "sendSessionAffinityHeaders": true,
156
+ "supportsEagerToolInputStreaming": false,
157
+ "supportsCacheControlOnTools": false,
158
+ "supportsLongCacheRetention": false
87
159
  }
88
160
  },
89
- "accounts/fireworks/models/deepseek-v3p2": {
90
- "name": "DeepSeek V3.2",
161
+ "accounts/fireworks/models/glm-5p2": {
162
+ "name": "GLM 5.2",
163
+ "api": "openai-completions",
164
+ "baseUrl": "https://api.fireworks.ai/inference/v1",
91
165
  "reasoning": true,
166
+ "thinkingLevelMap": {
167
+ "off": "none",
168
+ "minimal": null,
169
+ "low": "high",
170
+ "medium": "high",
171
+ "xhigh": "max"
172
+ },
173
+ "input": ["text"],
92
174
  "cost": {
93
- "input": 0.56,
94
- "output": 1.68,
95
- "cacheRead": 0.28,
175
+ "input": 1.4,
176
+ "output": 4.4,
177
+ "cacheRead": 0.26,
96
178
  "cacheWrite": 0
97
179
  },
98
- "maxTokens": 160000,
180
+ "contextWindow": 1048575,
181
+ "maxTokens": 131072,
99
182
  "compat": {
100
- "supportsReasoningEffort": true
183
+ "supportsStore": false,
184
+ "supportsDeveloperRole": false
101
185
  }
102
186
  },
103
- "accounts/fireworks/models/minimax-m2p5": {
104
- "name": "MiniMax-M2.5",
187
+ "accounts/fireworks/models/gpt-oss-120b": {
188
+ "name": "GPT OSS 120B",
189
+ "api": "anthropic-messages",
190
+ "baseUrl": "https://api.fireworks.ai/inference",
105
191
  "reasoning": true,
192
+ "input": ["text"],
106
193
  "cost": {
107
- "input": 0.3,
108
- "output": 1.2,
109
- "cacheRead": 0.03,
194
+ "input": 0.15,
195
+ "output": 0.6,
196
+ "cacheRead": 0.015,
110
197
  "cacheWrite": 0
111
198
  },
112
- "maxTokens": 196608,
199
+ "contextWindow": 131072,
200
+ "maxTokens": 32768,
113
201
  "compat": {
114
- "supportsReasoningEffort": true
202
+ "sendSessionAffinityHeaders": true,
203
+ "supportsEagerToolInputStreaming": false,
204
+ "supportsCacheControlOnTools": false,
205
+ "supportsLongCacheRetention": false
115
206
  }
116
207
  },
117
- "accounts/fireworks/models/glm-4p5-air": {
118
- "name": "GLM 4.5 Air",
208
+ "accounts/fireworks/models/gpt-oss-20b": {
209
+ "name": "GPT OSS 20B",
210
+ "api": "anthropic-messages",
211
+ "baseUrl": "https://api.fireworks.ai/inference",
119
212
  "reasoning": true,
213
+ "input": ["text"],
120
214
  "cost": {
121
- "input": 0.22,
122
- "output": 0.88,
123
- "cacheRead": 0,
215
+ "input": 0.07,
216
+ "output": 0.3,
217
+ "cacheRead": 0.035,
124
218
  "cacheWrite": 0
125
219
  },
126
- "maxTokens": 131072,
220
+ "contextWindow": 131072,
221
+ "maxTokens": 32768,
127
222
  "compat": {
128
- "supportsReasoningEffort": true
223
+ "sendSessionAffinityHeaders": true,
224
+ "supportsEagerToolInputStreaming": false,
225
+ "supportsCacheControlOnTools": false,
226
+ "supportsLongCacheRetention": false
129
227
  }
130
228
  },
131
- "accounts/fireworks/models/glm-5": {
132
- "name": "GLM 5",
229
+ "accounts/fireworks/models/gemma-4-26b-a4b-it": {
230
+ "name": "Gemma 4 26B A4B IT",
133
231
  "reasoning": true,
232
+ "input": ["text", "image"],
134
233
  "cost": {
135
- "input": 1,
136
- "output": 3.2,
137
- "cacheRead": 0.5,
234
+ "input": 0,
235
+ "output": 0,
236
+ "cacheRead": 0,
138
237
  "cacheWrite": 0
139
238
  },
140
- "maxTokens": 131072,
239
+ "contextWindow": 262000,
141
240
  "compat": {
142
241
  "supportsReasoningEffort": true
143
242
  }
144
243
  },
145
- "accounts/fireworks/models/deepseek-v3p1": {
146
- "name": "DeepSeek V3.1",
244
+ "accounts/fireworks/models/gemma-4-31b-it": {
245
+ "name": "Gemma 4 31B IT",
147
246
  "reasoning": true,
247
+ "input": ["text", "image"],
148
248
  "cost": {
149
- "input": 0.56,
150
- "output": 1.68,
249
+ "input": 0,
250
+ "output": 0,
151
251
  "cacheRead": 0,
152
252
  "cacheWrite": 0
153
253
  },
154
- "maxTokens": 163840,
254
+ "contextWindow": 262000,
155
255
  "compat": {
156
256
  "supportsReasoningEffort": true
157
257
  }
@@ -159,159 +259,287 @@
159
259
  "accounts/fireworks/models/kimi-k2-instruct": {
160
260
  "name": "Kimi K2 Instruct",
161
261
  "reasoning": false,
262
+ "input": ["text"],
162
263
  "cost": {
163
264
  "input": 1,
164
265
  "output": 3,
165
266
  "cacheRead": 0,
166
267
  "cacheWrite": 0
167
268
  },
269
+ "contextWindow": 131072,
168
270
  "maxTokens": 16384
169
271
  },
170
- "accounts/fireworks/models/qwen3p6-plus": {
171
- "name": "Qwen 3.6 Plus",
272
+ "accounts/fireworks/models/kimi-k2-thinking": {
273
+ "name": "Kimi K2 Thinking",
172
274
  "reasoning": true,
275
+ "input": ["text"],
173
276
  "cost": {
174
- "input": 0.5,
277
+ "input": 0.6,
278
+ "output": 2.5,
279
+ "cacheRead": 0.3,
280
+ "cacheWrite": 0
281
+ },
282
+ "contextWindow": 262144,
283
+ "maxTokens": 256000,
284
+ "compat": {
285
+ "supportsReasoningEffort": true
286
+ }
287
+ },
288
+ "accounts/fireworks/models/kimi-k2p5": {
289
+ "name": "Kimi K2.5",
290
+ "reasoning": true,
291
+ "input": ["text", "image"],
292
+ "cost": {
293
+ "input": 0.6,
175
294
  "output": 3,
176
295
  "cacheRead": 0.1,
177
296
  "cacheWrite": 0
178
297
  },
179
- "input": [
180
- "text",
181
- "image"
182
- ],
183
- "maxTokens": 8192,
298
+ "contextWindow": 262144,
299
+ "maxTokens": 256000,
184
300
  "compat": {
185
301
  "supportsReasoningEffort": true
186
302
  }
187
303
  },
304
+ "accounts/fireworks/models/kimi-k2p6": {
305
+ "name": "Kimi K2.6",
306
+ "api": "anthropic-messages",
307
+ "baseUrl": "https://api.fireworks.ai/inference",
308
+ "reasoning": true,
309
+ "input": ["text", "image"],
310
+ "cost": {
311
+ "input": 0.95,
312
+ "output": 4,
313
+ "cacheRead": 0.16,
314
+ "cacheWrite": 0
315
+ },
316
+ "contextWindow": 262000,
317
+ "maxTokens": 262000,
318
+ "compat": {
319
+ "sendSessionAffinityHeaders": true,
320
+ "supportsEagerToolInputStreaming": false,
321
+ "supportsCacheControlOnTools": false,
322
+ "supportsLongCacheRetention": false
323
+ }
324
+ },
325
+ "accounts/fireworks/models/kimi-k2p7-code": {
326
+ "name": "Kimi K2.7 Code",
327
+ "api": "anthropic-messages",
328
+ "baseUrl": "https://api.fireworks.ai/inference",
329
+ "reasoning": true,
330
+ "input": ["text", "image"],
331
+ "cost": {
332
+ "input": 0.95,
333
+ "output": 4,
334
+ "cacheRead": 0.19,
335
+ "cacheWrite": 0
336
+ },
337
+ "contextWindow": 262000,
338
+ "maxTokens": 262000,
339
+ "compat": {
340
+ "sendSessionAffinityHeaders": true,
341
+ "supportsEagerToolInputStreaming": false,
342
+ "supportsCacheControlOnTools": false,
343
+ "supportsLongCacheRetention": false
344
+ }
345
+ },
188
346
  "accounts/fireworks/models/minimax-m2p1": {
189
347
  "name": "MiniMax-M2.1",
190
348
  "reasoning": true,
349
+ "input": ["text"],
191
350
  "cost": {
192
351
  "input": 0.3,
193
352
  "output": 1.2,
194
353
  "cacheRead": 0.03,
195
354
  "cacheWrite": 0
196
355
  },
356
+ "contextWindow": 196608,
197
357
  "maxTokens": 200000,
198
358
  "compat": {
199
359
  "supportsReasoningEffort": true
200
360
  }
201
361
  },
202
- "accounts/fireworks/models/minimax-m2p7": {
203
- "name": "MiniMax-M2.7",
362
+ "accounts/fireworks/models/minimax-m2p5": {
363
+ "name": "MiniMax-M2.5",
204
364
  "reasoning": true,
365
+ "input": ["text"],
205
366
  "cost": {
206
367
  "input": 0.3,
207
368
  "output": 1.2,
208
369
  "cacheRead": 0.03,
209
370
  "cacheWrite": 0
210
371
  },
372
+ "contextWindow": 196608,
211
373
  "maxTokens": 196608,
212
374
  "compat": {
213
375
  "supportsReasoningEffort": true
214
376
  }
215
377
  },
216
- "accounts/fireworks/models/glm-4p7": {
217
- "name": "GLM 4.7",
378
+ "accounts/fireworks/models/minimax-m2p7": {
379
+ "name": "MiniMax-M2.7",
380
+ "api": "anthropic-messages",
381
+ "baseUrl": "https://api.fireworks.ai/inference",
218
382
  "reasoning": true,
383
+ "input": ["text"],
219
384
  "cost": {
220
- "input": 0.6,
221
- "output": 2.2,
222
- "cacheRead": 0.3,
385
+ "input": 0.3,
386
+ "output": 1.2,
387
+ "cacheRead": 0.06,
223
388
  "cacheWrite": 0
224
389
  },
225
- "maxTokens": 198000,
390
+ "contextWindow": 196608,
391
+ "maxTokens": 196608,
226
392
  "compat": {
227
- "supportsReasoningEffort": true
393
+ "sendSessionAffinityHeaders": true,
394
+ "supportsEagerToolInputStreaming": false,
395
+ "supportsCacheControlOnTools": false,
396
+ "supportsLongCacheRetention": false
228
397
  }
229
398
  },
230
- "accounts/fireworks/models/glm-4p5": {
231
- "name": "GLM 4.5",
399
+ "accounts/fireworks/models/minimax-m3": {
400
+ "name": "MiniMax-M3",
401
+ "api": "anthropic-messages",
402
+ "baseUrl": "https://api.fireworks.ai/inference",
232
403
  "reasoning": true,
404
+ "input": ["text"],
233
405
  "cost": {
234
- "input": 0.55,
235
- "output": 2.19,
236
- "cacheRead": 0,
406
+ "input": 0.3,
407
+ "output": 1.2,
408
+ "cacheRead": 0.06,
237
409
  "cacheWrite": 0
238
410
  },
239
- "maxTokens": 131072,
411
+ "contextWindow": 512000,
412
+ "maxTokens": 512000,
240
413
  "compat": {
241
- "supportsReasoningEffort": true
414
+ "sendSessionAffinityHeaders": true,
415
+ "supportsEagerToolInputStreaming": false,
416
+ "supportsCacheControlOnTools": false,
417
+ "supportsLongCacheRetention": false
242
418
  }
243
419
  },
244
- "accounts/fireworks/models/kimi-k2p5": {
245
- "name": "Kimi K2.5",
420
+ "accounts/fireworks/models/qwen3p6-plus": {
421
+ "name": "Qwen 3.6 Plus",
246
422
  "reasoning": true,
423
+ "input": ["text", "image"],
247
424
  "cost": {
248
- "input": 0.6,
425
+ "input": 0.5,
249
426
  "output": 3,
250
427
  "cacheRead": 0.1,
251
428
  "cacheWrite": 0
252
429
  },
253
- "input": [
254
- "text",
255
- "image",
256
- "video"
257
- ],
258
- "maxTokens": 256000,
430
+ "contextWindow": 262144,
431
+ "maxTokens": 8192,
259
432
  "compat": {
260
433
  "supportsReasoningEffort": true
261
434
  }
262
435
  },
263
- "accounts/fireworks/models/gpt-oss-20b": {
264
- "name": "GPT OSS 20B",
436
+ "accounts/fireworks/models/qwen3p7-plus": {
437
+ "name": "Qwen 3.7 Plus",
438
+ "api": "anthropic-messages",
439
+ "baseUrl": "https://api.fireworks.ai/inference",
265
440
  "reasoning": true,
441
+ "input": ["text", "image"],
266
442
  "cost": {
267
- "input": 0.05,
268
- "output": 0.2,
269
- "cacheRead": 0,
443
+ "input": 0.4,
444
+ "output": 1.6,
445
+ "cacheRead": 0.08,
270
446
  "cacheWrite": 0
271
447
  },
272
- "maxTokens": 32768,
448
+ "contextWindow": 262144,
449
+ "maxTokens": 65536,
273
450
  "compat": {
274
- "supportsReasoningEffort": true
451
+ "sendSessionAffinityHeaders": true,
452
+ "supportsEagerToolInputStreaming": false,
453
+ "supportsCacheControlOnTools": false,
454
+ "supportsLongCacheRetention": false
275
455
  }
276
456
  },
277
- "accounts/fireworks/models/gpt-oss-120b": {
278
- "name": "GPT OSS 120B",
457
+ "accounts/fireworks/routers/deepseek-v4-pro": {
458
+ "name": "DeepSeek V4 Pro (router)",
279
459
  "reasoning": true,
460
+ "input": ["text"],
280
461
  "cost": {
281
- "input": 0.15,
282
- "output": 0.6,
283
- "cacheRead": 0,
462
+ "input": 1.74,
463
+ "output": 3.48,
464
+ "cacheRead": 0.0145,
284
465
  "cacheWrite": 0
285
466
  },
286
- "maxTokens": 32768,
467
+ "contextWindow": 1048576,
468
+ "maxTokens": 1048576,
287
469
  "compat": {
288
470
  "supportsReasoningEffort": true
289
471
  }
290
472
  },
291
- "accounts/fireworks/models/kimi-k2-thinking": {
292
- "name": "Kimi K2 Thinking",
473
+ "accounts/fireworks/routers/glm-5-fast": {
474
+ "name": "GLM 5 Fast",
293
475
  "reasoning": true,
476
+ "input": ["text"],
294
477
  "cost": {
295
- "input": 0.6,
296
- "output": 2.5,
297
- "cacheRead": 0.3,
478
+ "input": 1,
479
+ "output": 3.2,
480
+ "cacheRead": 0.5,
298
481
  "cacheWrite": 0
299
482
  },
300
- "maxTokens": 256000,
483
+ "contextWindow": 202752,
484
+ "maxTokens": 131072,
301
485
  "compat": {
302
486
  "supportsReasoningEffort": true
303
487
  }
304
488
  },
305
- "accounts/fireworks/models/kimi-k2p6": {
306
- "name": "Kimi K2.6",
489
+ "accounts/fireworks/routers/glm-5p1-fast": {
490
+ "name": "GLM 5.1 Fast",
491
+ "api": "anthropic-messages",
492
+ "baseUrl": "https://api.fireworks.ai/inference",
307
493
  "reasoning": true,
494
+ "input": ["text"],
308
495
  "cost": {
309
- "input": 0.95,
310
- "output": 4,
311
- "cacheRead": 0.16,
496
+ "input": 2.8,
497
+ "output": 8.8,
498
+ "cacheRead": 0.52,
312
499
  "cacheWrite": 0
313
500
  },
314
- "maxTokens": 262144,
501
+ "contextWindow": 202800,
502
+ "maxTokens": 131072,
503
+ "compat": {
504
+ "sendSessionAffinityHeaders": true,
505
+ "supportsEagerToolInputStreaming": false,
506
+ "supportsCacheControlOnTools": false,
507
+ "supportsLongCacheRetention": false
508
+ }
509
+ },
510
+ "accounts/fireworks/routers/glm-5p2-fast": {
511
+ "name": "GLM 5.2 Fast",
512
+ "api": "anthropic-messages",
513
+ "baseUrl": "https://api.fireworks.ai/inference",
514
+ "reasoning": true,
515
+ "input": ["text"],
516
+ "cost": {
517
+ "input": 2.1,
518
+ "output": 6.6,
519
+ "cacheRead": 0.21,
520
+ "cacheWrite": 0
521
+ },
522
+ "contextWindow": 1048575,
523
+ "maxTokens": 131072,
524
+ "compat": {
525
+ "sendSessionAffinityHeaders": true,
526
+ "supportsEagerToolInputStreaming": false,
527
+ "supportsCacheControlOnTools": false,
528
+ "supportsLongCacheRetention": false
529
+ }
530
+ },
531
+ "accounts/fireworks/routers/kimi-k2p5-fast": {
532
+ "name": "Kimi K2.5 Fast (router)",
533
+ "reasoning": true,
534
+ "input": ["text", "image"],
535
+ "cost": {
536
+ "input": 0.6,
537
+ "output": 3,
538
+ "cacheRead": 0.1,
539
+ "cacheWrite": 0
540
+ },
541
+ "contextWindow": 262144,
542
+ "maxTokens": 256000,
315
543
  "compat": {
316
544
  "supportsReasoningEffort": true
317
545
  }
@@ -333,84 +561,93 @@
333
561
  "accounts/fireworks/routers/kimi-k2p6": {
334
562
  "name": "Kimi K2.6 (router)",
335
563
  "reasoning": true,
564
+ "input": ["text", "image"],
336
565
  "cost": {
337
566
  "input": 0.95,
338
567
  "output": 4,
339
568
  "cacheRead": 0.16,
340
569
  "cacheWrite": 0
341
570
  },
342
- "input": [
343
- "text",
344
- "image"
345
- ],
571
+ "contextWindow": 262144,
346
572
  "maxTokens": 262144,
347
573
  "compat": {
348
574
  "supportsReasoningEffort": true
349
575
  }
350
576
  },
351
- "accounts/fireworks/routers/glm-5p1-fast": {
352
- "name": "GLM 5.1 Fast (router)",
577
+ "accounts/fireworks/routers/kimi-k2p6-fast": {
578
+ "name": "Kimi K2.6 Fast",
579
+ "api": "anthropic-messages",
580
+ "baseUrl": "https://api.fireworks.ai/inference",
353
581
  "reasoning": true,
582
+ "input": ["text", "image"],
354
583
  "cost": {
355
- "input": 1.4,
356
- "output": 4.4,
357
- "cacheRead": 0.26,
358
- "cacheWrite": 0
359
- },
360
- "maxTokens": 131072,
361
- "compat": {
362
- "supportsReasoningEffort": true
363
- }
364
- },
365
- "accounts/fireworks/routers/minimax-m2p7": {
366
- "name": "MiniMax M2.7 (router)",
367
- "reasoning": true,
368
- "cost": {
369
- "input": 0.3,
370
- "output": 1.2,
371
- "cacheRead": 0.06,
584
+ "input": 2,
585
+ "output": 8,
586
+ "cacheRead": 0.3,
372
587
  "cacheWrite": 0
373
588
  },
589
+ "contextWindow": 262000,
590
+ "maxTokens": 262000,
374
591
  "compat": {
375
- "supportsReasoningEffort": true
592
+ "sendSessionAffinityHeaders": true,
593
+ "supportsEagerToolInputStreaming": false,
594
+ "supportsCacheControlOnTools": false,
595
+ "supportsLongCacheRetention": false
376
596
  }
377
597
  },
378
- "accounts/fireworks/routers/glm-5-fast": {
379
- "name": "GLM 5 Fast (router)",
598
+ "accounts/fireworks/routers/kimi-k2p6-turbo": {
599
+ "name": "Kimi K2.6 Turbo",
600
+ "api": "anthropic-messages",
601
+ "baseUrl": "https://api.fireworks.ai/inference",
380
602
  "reasoning": true,
603
+ "input": ["text", "image"],
381
604
  "cost": {
382
- "input": 1,
383
- "output": 3.2,
384
- "cacheRead": 0.5,
605
+ "input": 2,
606
+ "output": 8,
607
+ "cacheRead": 0.3,
385
608
  "cacheWrite": 0
386
609
  },
387
- "maxTokens": 131072,
610
+ "contextWindow": 262000,
611
+ "maxTokens": 262000,
388
612
  "compat": {
389
- "supportsReasoningEffort": true
613
+ "sendSessionAffinityHeaders": true,
614
+ "supportsEagerToolInputStreaming": false,
615
+ "supportsCacheControlOnTools": false,
616
+ "supportsLongCacheRetention": false
390
617
  }
391
618
  },
392
- "accounts/fireworks/models/gemma-4-31b-it": {
393
- "name": "Gemma 4 31B IT",
619
+ "accounts/fireworks/routers/kimi-k2p7-code-fast": {
620
+ "name": "Kimi K2.7 Code Fast",
621
+ "api": "anthropic-messages",
622
+ "baseUrl": "https://api.fireworks.ai/inference",
394
623
  "reasoning": true,
624
+ "input": ["text", "image"],
395
625
  "cost": {
396
- "input": 0,
397
- "output": 0,
398
- "cacheRead": 0,
626
+ "input": 1.9,
627
+ "output": 8,
628
+ "cacheRead": 0.38,
399
629
  "cacheWrite": 0
400
630
  },
631
+ "contextWindow": 262000,
632
+ "maxTokens": 262000,
401
633
  "compat": {
402
- "supportsReasoningEffort": true
634
+ "sendSessionAffinityHeaders": true,
635
+ "supportsEagerToolInputStreaming": false,
636
+ "supportsCacheControlOnTools": false,
637
+ "supportsLongCacheRetention": false
403
638
  }
404
639
  },
405
- "accounts/fireworks/models/gemma-4-26b-a4b-it": {
406
- "name": "Gemma 4 26B A4B IT",
640
+ "accounts/fireworks/routers/minimax-m2p7": {
641
+ "name": "MiniMax M2.7 (router)",
407
642
  "reasoning": true,
643
+ "input": ["text"],
408
644
  "cost": {
409
- "input": 0,
410
- "output": 0,
411
- "cacheRead": 0,
645
+ "input": 0.3,
646
+ "output": 1.2,
647
+ "cacheRead": 0.06,
412
648
  "cacheWrite": 0
413
649
  },
650
+ "contextWindow": 204000,
414
651
  "compat": {
415
652
  "supportsReasoningEffort": true
416
653
  }