pi-fireworks-provider 1.0.2 → 1.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +80 -2
- package/custom-models.json +67 -0
- package/index.ts +533 -26
- package/package.json +26 -8
- package/patch.json +400 -163
- package/.github/FUNDING.yml +0 -4
- package/.pi/messenger/channels/memory.jsonl +0 -1
- package/.pi/messenger/session-id +0 -1
- package/AGENTS.md +0 -56
- package/scripts/update-models.js +0 -349
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
**31+ models through [Fireworks AI](https://fireworks.ai/)**
|
|
6
6
|
|
|
7
|
-
_Kimi, MiniMax, GLM, DeepSeek, GPT-OSS —
|
|
7
|
+
_Kimi, MiniMax, GLM, DeepSeek, GPT-OSS — via Fireworks AI's Anthropic Messages and OpenAI-compatible endpoints for [pi](https://github.com/earendil-works/pi-coding-agent)._
|
|
8
8
|
|
|
9
9
|
[](https://github.com/earendil-works/pi-coding-agent)
|
|
10
10
|
[](./LICENSE)
|
|
@@ -16,7 +16,10 @@ _Kimi, MiniMax, GLM, DeepSeek, GPT-OSS — unified OpenAI-compatible API for [pi
|
|
|
16
16
|
## Features
|
|
17
17
|
|
|
18
18
|
- **35+ AI Models** including Kimi K2.5, MiniMax M2.5, GLM 4.5/4.7/5, DeepSeek V3.1/V3.2, DeepSeek V4 Flash, and GPT-OSS
|
|
19
|
-
- **
|
|
19
|
+
- **Dual API support** via Fireworks AI's Anthropic Messages and OpenAI-compatible completions endpoints (per-model routing, matching pi core's Fireworks provider)
|
|
20
|
+
- **Service tiers** — toggle Fireworks `priority` vs `standard` per request on supported models (with priority pricing reflected in cost tracking), via a keybinding, `/fireworks-tier`, and a footer status area
|
|
21
|
+
- **Preserved thinking** — toggle Fireworks' `reasoning_history: "preserved"` so prior assistant reasoning is retained across turns (better multi-turn recall; uses more tokens), via the `/fireworks-settings` panel, with a model-select notification. Matches neuralwatt/makora's settings-only UX, adapted to Fireworks' single global `reasoning_history` knob
|
|
22
|
+
- **Settings panel** — `/fireworks-settings` (TUI) to configure preserved thinking, service tier, and display preferences; persisted to `~/.pi/agent/extensions/fireworks.json`
|
|
20
23
|
- **Cost Tracking** with per-model pricing for budget management
|
|
21
24
|
- **Reasoning Models** support for advanced reasoning capabilities
|
|
22
25
|
- **Vision Support** for image-capable models
|
|
@@ -105,6 +108,69 @@ pi
|
|
|
105
108
|
| Qwen3 VL 30B A3B Thinking | Text + Image | 262K | 0 | Free | Free |
|
|
106
109
|
*Costs are per million tokens. Prices subject to change - check [fireworks.ai](https://fireworks.ai) for current pricing.*
|
|
107
110
|
|
|
111
|
+
## Service Tiers
|
|
112
|
+
|
|
113
|
+
Fireworks exposes a `service_tier` request field (`standard` | `priority`) on its chat-completions endpoint. The **priority** tier trades higher per-token pricing for higher throughput / lower latency. This is orthogonal to the `-fast`/`-turbo` router model IDs (which are separate models) — service tiers apply to the base models below.
|
|
114
|
+
|
|
115
|
+
| Model | Priority Uncached Input | Priority Cached Input | Priority Output |
|
|
116
|
+
| --- | --- | --- | --- |
|
|
117
|
+
| GLM 5.2 | $1.75/M | $0.175/M | $5.5/M |
|
|
118
|
+
| Kimi K2.7 Code | $1.43/M | $0.29/M | $6/M |
|
|
119
|
+
| Minimax M3 | $0.45/M | $0.09/M | $1.8/M |
|
|
120
|
+
| DeepSeek V4 Pro | $2.61/M | $0.218/M | $5.22/M |
|
|
121
|
+
| Kimi K2.6 | $1.5/M | $0.22/M | $6/M |
|
|
122
|
+
| MiniMax M2.7 | $0.45/M | $0.09/M | $1.8/M |
|
|
123
|
+
| GLM 5.1 | $2.1/M | $0.39/M | $6.6/M |
|
|
124
|
+
| GPT OSS 120B | $0.18/M | $0.018/M | $0.72/M |
|
|
125
|
+
| DeepSeek V4 Flash | $0.21/M | $0.045/M | $0.42/M |
|
|
126
|
+
|
|
127
|
+
*Priority pricing is roughly 1.2–1.5× the standard rate. `cacheWrite` is not tiered.*
|
|
128
|
+
|
|
129
|
+
**Switching tiers:**
|
|
130
|
+
|
|
131
|
+
- **Keybinding:** `ctrl+shift+l` (default) toggles `standard` ↔ `priority` for the active supported model. No-op with an info notice for unsupported models.
|
|
132
|
+
- **Command:** `/fireworks-tier standard|priority|toggle`.
|
|
133
|
+
- **Status area:** a dim `tier: standard` / `tier: ⚡priority` line is shown in the footer for supported models while a Fireworks model is active.
|
|
134
|
+
|
|
135
|
+
The selection is persisted per session (survives `/reload` and resume). When `priority` is active, `service_tier: "priority"` is injected into every request and finalized cost is recomputed against the priority rates above.
|
|
136
|
+
|
|
137
|
+
**Configuration** — `~/.pi/agent/extensions/fireworks.json` (created with defaults on first load):
|
|
138
|
+
|
|
139
|
+
```json
|
|
140
|
+
{
|
|
141
|
+
"serviceTier": {
|
|
142
|
+
"default": "standard",
|
|
143
|
+
"keybinding": "ctrl+shift+l",
|
|
144
|
+
"display": "statusbar"
|
|
145
|
+
},
|
|
146
|
+
"preserveThinking": {
|
|
147
|
+
"default": false
|
|
148
|
+
}
|
|
149
|
+
}
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
- `serviceTier.default` — tier used until you toggle (`standard` | `priority`).
|
|
153
|
+
- `serviceTier.keybinding` — any [pi key format](https://github.com/earendil-works/pi-coding-agent/blob/main/docs/keybindings.md) (e.g. `ctrl+shift+l`, `ctrl+shift+k`). Requires `/reload` after changing. On macOS browser terminals (localterm), avoid `alt`/`ctrl+alt` (Option produces special chars) and `ctrl+shift+t/w/n/c/v` (browser/localterm tab + copy/paste shortcuts).
|
|
154
|
+
- `serviceTier.display` — `statusbar` (footer status area) or `off` (hide the tier indicator).
|
|
155
|
+
- `preserveThinking.default` — whether `reasoning_history: "preserved"` is injected (`true` | `false`, default `false`). Also settable via `/fireworks-settings`.
|
|
156
|
+
|
|
157
|
+
> **Note:** The OpenAI completions endpoint accepts `service_tier` directly (per Fireworks' API). The Anthropic Messages endpoint passes the top-level field through as an extra. If a supported Anthropic-routed model rejects it, file an issue so we can gate injection by API.
|
|
158
|
+
|
|
159
|
+
## Preserved Thinking
|
|
160
|
+
|
|
161
|
+
Fireworks exposes a top-level `reasoning_history` request parameter. The only accepted value is `"preserved"`; omitting it (the default) means prior assistant reasoning is **stripped** from the model's context each turn. Setting `reasoning_history: "preserved"` makes Fireworks render prior assistant reasoning into the model's context, improving multi-turn recall at the cost of extra tokens. See [the Fireworks reasoning guide](https://docs.fireworks.ai/guides/reasoning#preserved-thinking).
|
|
162
|
+
|
|
163
|
+
This works on **both** transports — the OpenAI completions endpoint (assistant `reasoning_content` field) and the Anthropic Messages endpoint (assistant `thinking` content blocks, for which Fireworks returns a `signature` so pi-ai replays them). pi-ai already replays the reasoning field/block on prior assistant turns; this extension's only job is injecting the top-level `reasoning_history: "preserved"` flag that makes Fireworks honor it.
|
|
164
|
+
|
|
165
|
+
Unlike neuralwatt/makora (which use per-model vLLM `chat_template_kwargs` flags like `preserve_thinking`/`clear_thinking`), Fireworks' knob is a single global parameter that applies to every reasoning model, so we expose it as one on/off toggle rather than a per-model submenu.
|
|
166
|
+
|
|
167
|
+
**Toggle it:**
|
|
168
|
+
|
|
169
|
+
- **`/fireworks-settings`** (TUI) → *Preserved thinking* → `on` / `off`. Takes effect immediately and persists as the default for future sessions.
|
|
170
|
+
- A dim **model-select notification** tells you the current state when you switch to a Fireworks reasoning model (`Preserved thinking ON for …` / `… OFF for …`).
|
|
171
|
+
|
|
172
|
+
Preserved thinking is **off by default** to match pi core and Fireworks' default (stripped). There is intentionally no `/fireworks-preserve` command or keybinding — it's settings-panel-only, mirroring neuralwatt/makora.
|
|
173
|
+
|
|
108
174
|
## Usage
|
|
109
175
|
|
|
110
176
|
After loading the extension, use the `/model` command in pi to select your preferred model:
|
|
@@ -145,6 +211,18 @@ Add to your pi configuration for automatic loading:
|
|
|
145
211
|
}
|
|
146
212
|
```
|
|
147
213
|
|
|
214
|
+
## Development
|
|
215
|
+
|
|
216
|
+
```bash
|
|
217
|
+
pnpm install # install dev tooling (vitest, knip, typescript)
|
|
218
|
+
pnpm test # run the test suite (vitest)
|
|
219
|
+
pnpm run test:watch # watch mode
|
|
220
|
+
pnpm run lint:dead # dead-code / unused-export scan (knip)
|
|
221
|
+
pnpm run check # typecheck (tsc) + tests + knip, all green or exit non-zero
|
|
222
|
+
```
|
|
223
|
+
|
|
224
|
+
Tests live in `tests/` and stub the `@earendil-works/pi-coding-agent` / `@earendil-works/pi-tui` peer dependencies (see `tests/__mocks__/`) so they run without the real pi packages installed. A per-run temp dir is used for `~/.pi/agent` (via `PI_CODING_AGENT_DIR` in `tests/vitest.setup.ts`) so config/cache reads and writes never touch your real environment.
|
|
225
|
+
|
|
148
226
|
## License
|
|
149
227
|
|
|
150
228
|
MIT
|
package/custom-models.json
CHANGED
|
@@ -147,5 +147,72 @@
|
|
|
147
147
|
],
|
|
148
148
|
"contextWindow": 262000,
|
|
149
149
|
"maxTokens": 0
|
|
150
|
+
},
|
|
151
|
+
{
|
|
152
|
+
"id": "accounts/fireworks/models/qwen3p7-plus",
|
|
153
|
+
"name": "Qwen 3.7 Plus",
|
|
154
|
+
"reasoning": true,
|
|
155
|
+
"cost": {
|
|
156
|
+
"input": 0.4,
|
|
157
|
+
"output": 1.6,
|
|
158
|
+
"cacheRead": 0.08,
|
|
159
|
+
"cacheWrite": 0
|
|
160
|
+
},
|
|
161
|
+
"input": [
|
|
162
|
+
"text",
|
|
163
|
+
"image"
|
|
164
|
+
],
|
|
165
|
+
"contextWindow": 262144,
|
|
166
|
+
"maxTokens": 65536
|
|
167
|
+
},
|
|
168
|
+
{
|
|
169
|
+
"id": "accounts/fireworks/routers/glm-5p2-fast",
|
|
170
|
+
"name": "GLM 5.2 Fast",
|
|
171
|
+
"reasoning": true,
|
|
172
|
+
"cost": {
|
|
173
|
+
"input": 2.1,
|
|
174
|
+
"output": 6.6,
|
|
175
|
+
"cacheRead": 0.21,
|
|
176
|
+
"cacheWrite": 0
|
|
177
|
+
},
|
|
178
|
+
"input": [
|
|
179
|
+
"text"
|
|
180
|
+
],
|
|
181
|
+
"contextWindow": 1048575,
|
|
182
|
+
"maxTokens": 131072
|
|
183
|
+
},
|
|
184
|
+
{
|
|
185
|
+
"id": "accounts/fireworks/routers/kimi-k2p6-fast",
|
|
186
|
+
"name": "Kimi K2.6 Fast",
|
|
187
|
+
"reasoning": true,
|
|
188
|
+
"cost": {
|
|
189
|
+
"input": 2,
|
|
190
|
+
"output": 8,
|
|
191
|
+
"cacheRead": 0.3,
|
|
192
|
+
"cacheWrite": 0
|
|
193
|
+
},
|
|
194
|
+
"input": [
|
|
195
|
+
"text",
|
|
196
|
+
"image"
|
|
197
|
+
],
|
|
198
|
+
"contextWindow": 262000,
|
|
199
|
+
"maxTokens": 262000
|
|
200
|
+
},
|
|
201
|
+
{
|
|
202
|
+
"id": "accounts/fireworks/routers/kimi-k2p7-code-fast",
|
|
203
|
+
"name": "Kimi K2.7 Code Fast",
|
|
204
|
+
"reasoning": true,
|
|
205
|
+
"cost": {
|
|
206
|
+
"input": 1.9,
|
|
207
|
+
"output": 8,
|
|
208
|
+
"cacheRead": 0.38,
|
|
209
|
+
"cacheWrite": 0
|
|
210
|
+
},
|
|
211
|
+
"input": [
|
|
212
|
+
"text",
|
|
213
|
+
"image"
|
|
214
|
+
],
|
|
215
|
+
"contextWindow": 262000,
|
|
216
|
+
"maxTokens": 262000
|
|
150
217
|
}
|
|
151
218
|
]
|