slash-tokens 1.6.4 → 1.6.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +34 -1
- package/README.md +6 -2
- package/dist/cli.js +182 -23
- package/dist/intercept.d.ts +8 -0
- package/dist/intercept.js +59 -14
- package/dist/models.d.ts +7 -0
- package/dist/models.js +111 -36
- package/dist/patterns.d.ts +1 -1
- package/dist/patterns.js +4 -4
- package/dist/preflight.js +10 -4
- package/dist/providers.d.ts +15 -17
- package/dist/providers.js +44 -22
- package/dist/report.js +1 -0
- package/dist/slash.js +38 -2
- package/package.json +7 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,10 +1,43 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## [1.6.6] — The Fixed Deal Edition
|
|
4
|
+
|
|
5
|
+
*2026-10-07*
|
|
6
|
+
|
|
7
|
+
`--version` / `--help` answer, and today's models price correctly.
|
|
8
|
+
|
|
9
|
+
Solo $20 mailbox, 10% waived. Team $39 for the data.
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
- **`slash-tokens --version` and `--help` print and exit.** In 1.6.5 every flag fell through to a full scan of the current folder, so `--version` run from `~` scanned the home folder. An empty `bunx` run now tells you to run it in an app.
|
|
13
|
+
- **`preflight()` no longer throws on current models or real API IDs.** IDs are canonicalised strictly (`claude-opus-4-7` → `claude-opus-4.7`, date stamps dropped). There is no family guessing: an unknown version still fails with "Unknown model", so a new model never silently gets an older model's price.
|
|
14
|
+
- **`npm publish` always builds first** (`prepublishOnly`), so a stale `dist/` can't ship again (1.6.4).
|
|
15
|
+
- **`report()` sends numbers only.** A test pins its payload to `tokens_estimated`, `tokens_saved`, `model`, `action`, `cost_saved_usd`; prompt content never leaves the machine.
|
|
16
|
+
- **`npm test` runs offline.** The live integration suite (`tests/z-integration.test.ts`, which registers keys on mcpaas.live) runs only with `SLASH_LIVE=1`; the weekly `integration.yml` sets it.
|
|
17
|
+
|
|
18
|
+
### Changed
|
|
19
|
+
- **Model table, prices as of 2026-10-07** (checked against the Anthropic, OpenAI, xAI and Google pricing pages). Added Claude Opus 5.5 ($4/$20), Sonnet 5.5 ($2/$10), Opus 4.8 / 4.6 / 4.5, Sonnet 4.6 / 4.5, Fable 5 / 5.1 and Mythos 5 / 5.1 ($10/$50); Grok 4.7 and 4.5 ($2/$6, doubling above 200K), grok-build-0.1 ($1/$2); GPT-6 Astra ($10/$50), GPT-6.1 Sol and GPT-6 Sol ($2/$10), GPT-6 Luna ($0.10/$0.50); Gemini 3.8 / 3.7 / 3.6 Flash ($0.75/$3.75 launch price through 2026-12-31), 3.5 Flash ($1.50/$9), 3.1 Flash-Lite ($0.25/$1.50), 3.1 Pro Preview.
|
|
20
|
+
- **`slash-tokens/auto` routing is unchanged within 1.6.x.** `/auto` rewrites live requests only to the models it routed to in 1.6.5 (`AUTO_ROUTE_TARGETS`), using the same identification, so a patch upgrade never changes where production calls go. New models are now recognised and priced correctly in `/auto` events (`identifyModel`: `claude-opus-5-5` at $4/$20, `gpt-6-astra` at $10/$50); /auto starts routing to them in 1.7.0.
|
|
21
|
+
- **Routing (`preflightRoute`) follows the new ladder.** Cheapest same-provider is now GPT-6 Luna for OpenAI and Gemini 3.1 Flash-Lite for Google (Anthropic and xAI unchanged). grok-build-0.1 is priced but never a routing target (`NOT_ROUTE_TARGETS`): it's a coding-agent model. When two alternatives' costs round to the same value (tiny prompts), the lower list price wins.
|
|
22
|
+
- **Calibration:** new Claude, Grok and Gemini models use their family's measured factor. GPT-6 is a new generation with an unbenchmarked tokenizer, so it takes the conservative default until a bench run adds it (slash never under-reports).
|
|
23
|
+
|
|
24
|
+
No Team/Solo price change.
|
|
25
|
+
|
|
26
|
+
## [1.6.5] — The Fixed Deal Edition
|
|
27
|
+
|
|
28
|
+
*2026-08-25*
|
|
29
|
+
|
|
30
|
+
Solo $20 mailbox, 10% waived. Team $39 for the data.
|
|
31
|
+
|
|
32
|
+
Rebuilt tarball. 1.6.4 packed a stale gitignored `dist/` (`bun test` runs `src/`). This tarball contains the live ladder: Grok **4.6** / **4.3**, GPT-5.6 **Sol / Terra / Luna**, Claude **Opus 5 / Sonnet 5 / Haiku 4.5**, Gemini **3.5 Flash-Lite**.
|
|
33
|
+
|
|
34
|
+
No Team/Solo price change.
|
|
35
|
+
|
|
3
36
|
## [1.6.4] — The Fixed Deal Edition
|
|
4
37
|
|
|
5
38
|
*2026-08-25*
|
|
6
39
|
|
|
7
|
-
Live model ladder
|
|
40
|
+
Live model ladder in source. **Tarball packed stale `dist/`** — `bun test` never rebuilds it. Use **1.6.5**.
|
|
8
41
|
|
|
9
42
|
No Team/Solo price change.
|
|
10
43
|
|
package/README.md
CHANGED
|
@@ -15,10 +15,14 @@ Know the cost before the call leaves your machine.
|
|
|
15
15
|
Models change. Windows grow. Slash adapts — you keep building.
|
|
16
16
|
Cheaper tokens haven't shrunk the bill — usage has.
|
|
17
17
|
|
|
18
|
-
## v1.6.
|
|
18
|
+
## v1.6.6 — The Fixed Deal Edition
|
|
19
|
+
|
|
20
|
+
`--version` / `--help` answer, and today's models price correctly.
|
|
19
21
|
|
|
20
22
|
Solo $20 mailbox, 10% waived. Team $39 for the data.
|
|
21
23
|
|
|
24
|
+
New in 1.6.6: Claude Opus 5.5 / Sonnet 5.5 / Fable 5.1, Grok 4.7, GPT-6 (Astra, Sol, Luna), Gemini 3.6–3.8 Flash and 3.1 Flash-Lite, priced as of 2026-10-07. Real API IDs (`claude-opus-4-7`) work in `preflight()`. `preflightRoute()` now finds GPT-6 Luna and Gemini 3.1 Flash-Lite as the cheapest same-provider options.
|
|
25
|
+
|
|
22
26
|
**Free forever is bunx** — no account. A one-person account is email → key, **$20 on the house**. We show the savings. We don't charge. 10% is the model, waived. Team is **$39 for the data** (`$390`/year).
|
|
23
27
|
|
|
24
28
|
```bash
|
|
@@ -26,7 +30,7 @@ bunx slash-tokens
|
|
|
26
30
|
# or: npx --yes slash-tokens
|
|
27
31
|
```
|
|
28
32
|
|
|
29
|
-
Run it in a project that already calls an LLM. An empty folder prints that nothing was found — that's normal.
|
|
33
|
+
Run it in a project that already calls an LLM. An empty folder prints that nothing was found, then tells you to run it in an app — that's normal. `--version` / `--help` print and exit (they do not scan). Pin proof: `slash-tokens --version` or `npm view slash-tokens version`.
|
|
30
34
|
|
|
31
35
|
```bash
|
|
32
36
|
npm install slash-tokens
|
package/dist/cli.js
CHANGED
|
@@ -1,5 +1,10 @@
|
|
|
1
1
|
#!/usr/bin/env node
|
|
2
2
|
|
|
3
|
+
// src/cli.ts
|
|
4
|
+
import { readFileSync as readFileSync2 } from "node:fs";
|
|
5
|
+
import { dirname, join as join2 } from "node:path";
|
|
6
|
+
import { fileURLToPath } from "node:url";
|
|
7
|
+
|
|
3
8
|
// src/scanner.ts
|
|
4
9
|
import { readdirSync, readFileSync, statSync } from "fs";
|
|
5
10
|
import { join, extname } from "path";
|
|
@@ -49,17 +54,162 @@ function writeToMemory(content) {
|
|
|
49
54
|
return maxLen;
|
|
50
55
|
}
|
|
51
56
|
|
|
57
|
+
// src/models.ts
|
|
58
|
+
var OPUS = { input: 5, output: 25, context: 1e6 };
|
|
59
|
+
var OPUS_55 = { input: 4, output: 20, context: 1e6 };
|
|
60
|
+
var FABLE = { input: 10, output: 50, context: 1e6 };
|
|
61
|
+
var SONNET_4X = { input: 3, output: 15, context: 1e6 };
|
|
62
|
+
var SONNET = { input: 2, output: 10, context: 1e6 };
|
|
63
|
+
var HAIKU = { input: 1, output: 5, context: 200000 };
|
|
64
|
+
var GROK_46 = {
|
|
65
|
+
input: 2,
|
|
66
|
+
output: 6,
|
|
67
|
+
context: 500000,
|
|
68
|
+
longContextThreshold: 200000,
|
|
69
|
+
longContextInput: 4,
|
|
70
|
+
longContextOutput: 12
|
|
71
|
+
};
|
|
72
|
+
var GROK_43 = {
|
|
73
|
+
input: 1.25,
|
|
74
|
+
output: 2.5,
|
|
75
|
+
context: 1e6,
|
|
76
|
+
longContextThreshold: 200000,
|
|
77
|
+
longContextInput: 2.5,
|
|
78
|
+
longContextOutput: 5
|
|
79
|
+
};
|
|
80
|
+
var GEMINI_PRO = {
|
|
81
|
+
input: 2,
|
|
82
|
+
output: 12,
|
|
83
|
+
context: 1e6,
|
|
84
|
+
longContextThreshold: 200000,
|
|
85
|
+
longContextInput: 4,
|
|
86
|
+
longContextOutput: 18
|
|
87
|
+
};
|
|
88
|
+
var GEMINI_FLASH = { input: 0.3, output: 2.5, context: 1e6 };
|
|
89
|
+
var GEMINI_FLASH_3X = { input: 0.75, output: 3.75, context: 1e6 };
|
|
90
|
+
var GEMINI_35_FLASH = { input: 1.5, output: 9, context: 1e6 };
|
|
91
|
+
var GEMINI_31_FLASH_LITE = { input: 0.25, output: 1.5, context: 1e6 };
|
|
92
|
+
var GROK_BUILD = {
|
|
93
|
+
input: 1,
|
|
94
|
+
output: 2,
|
|
95
|
+
context: 256000,
|
|
96
|
+
longContextThreshold: 200000,
|
|
97
|
+
longContextInput: 2,
|
|
98
|
+
longContextOutput: 4
|
|
99
|
+
};
|
|
100
|
+
var GPT_6_ASTRA = { input: 10, output: 50, context: 1050000 };
|
|
101
|
+
var GPT_6_SOL = { input: 2, output: 10, context: 1050000 };
|
|
102
|
+
var GPT_6_LUNA = { input: 0.1, output: 0.5, context: 1050000 };
|
|
103
|
+
var GPT_SOL = { input: 4, output: 20, context: 1050000 };
|
|
104
|
+
var GPT_TERRA = { input: 2, output: 12, context: 1050000 };
|
|
105
|
+
var GPT_LUNA = { input: 0.2, output: 1.2, context: 1050000 };
|
|
106
|
+
var GPT_54 = { input: 2.5, output: 15, context: 1e6 };
|
|
107
|
+
var GPT_54_MINI = { input: 0.75, output: 4.5, context: 128000 };
|
|
108
|
+
var GPT_54_NANO = { input: 0.2, output: 1.25, context: 128000 };
|
|
109
|
+
var MODELS = {
|
|
110
|
+
"claude-fable-5.1": { ...FABLE },
|
|
111
|
+
"claude-fable-5": { ...FABLE },
|
|
112
|
+
"claude-mythos-5.1": { ...FABLE },
|
|
113
|
+
"claude-mythos-5": { ...FABLE },
|
|
114
|
+
"claude-opus-5.5": { ...OPUS_55 },
|
|
115
|
+
"claude-opus-5": { ...OPUS },
|
|
116
|
+
"claude-opus-4.8": { ...OPUS },
|
|
117
|
+
"claude-opus": { ...OPUS },
|
|
118
|
+
"claude-opus-4.7": { ...OPUS },
|
|
119
|
+
"claude-opus-4.6": { ...OPUS },
|
|
120
|
+
"claude-opus-4.5": { ...OPUS },
|
|
121
|
+
"claude-sonnet-5.5": { ...SONNET },
|
|
122
|
+
"claude-sonnet-5": { ...SONNET },
|
|
123
|
+
"claude-sonnet": { ...SONNET },
|
|
124
|
+
"claude-sonnet-4.6": { ...SONNET_4X },
|
|
125
|
+
"claude-sonnet-4.5": { ...SONNET_4X },
|
|
126
|
+
"claude-haiku-4.5": { ...HAIKU },
|
|
127
|
+
"claude-haiku": { ...HAIKU },
|
|
128
|
+
"grok-4.7": { ...GROK_46 },
|
|
129
|
+
"grok-4.6": { ...GROK_46 },
|
|
130
|
+
"grok-4.5": { ...GROK_46 },
|
|
131
|
+
"grok-build-0.1": { ...GROK_BUILD },
|
|
132
|
+
"grok-4.3": { ...GROK_43 },
|
|
133
|
+
"grok-4.20": { ...GROK_43 },
|
|
134
|
+
"grok-4-1-fast": { ...GROK_43 },
|
|
135
|
+
"gemini-3.1-pro": { ...GEMINI_PRO },
|
|
136
|
+
"gemini-3.1-pro-preview": { ...GEMINI_PRO },
|
|
137
|
+
"gemini-3.8-flash": { ...GEMINI_FLASH_3X },
|
|
138
|
+
"gemini-3.7-flash": { ...GEMINI_FLASH_3X },
|
|
139
|
+
"gemini-3.6-flash": { ...GEMINI_FLASH_3X },
|
|
140
|
+
"gemini-3.5-flash": { ...GEMINI_35_FLASH },
|
|
141
|
+
"gemini-3.1-flash-lite": { ...GEMINI_31_FLASH_LITE },
|
|
142
|
+
"gemini-3.5-flash-lite": { ...GEMINI_FLASH },
|
|
143
|
+
"gemini-2.5-flash": { ...GEMINI_FLASH },
|
|
144
|
+
"gpt-6-astra": { ...GPT_6_ASTRA },
|
|
145
|
+
"gpt-6.1-sol": { ...GPT_6_SOL },
|
|
146
|
+
"gpt-6-sol": { ...GPT_6_SOL },
|
|
147
|
+
"gpt-6-luna": { ...GPT_6_LUNA },
|
|
148
|
+
"gpt-5.6-sol": { ...GPT_SOL },
|
|
149
|
+
"gpt-5.6-terra": { ...GPT_TERRA },
|
|
150
|
+
"gpt-5.6-luna": { ...GPT_LUNA },
|
|
151
|
+
"gpt-5.4": { ...GPT_54 },
|
|
152
|
+
"gpt-5.4-mini": { ...GPT_54_MINI },
|
|
153
|
+
"gpt-5.4-nano": { ...GPT_54_NANO }
|
|
154
|
+
};
|
|
155
|
+
function canonicalModel(name) {
|
|
156
|
+
if (MODELS[name])
|
|
157
|
+
return name;
|
|
158
|
+
let n = name.trim().toLowerCase();
|
|
159
|
+
if (MODELS[n])
|
|
160
|
+
return n;
|
|
161
|
+
n = n.replace(/-\d{8}$/, "");
|
|
162
|
+
if (MODELS[n])
|
|
163
|
+
return n;
|
|
164
|
+
const dotted = n.replace(/(\d)-(\d)(?=$|-)/g, "$1.$2");
|
|
165
|
+
if (MODELS[dotted])
|
|
166
|
+
return dotted;
|
|
167
|
+
return n;
|
|
168
|
+
}
|
|
169
|
+
function getModel(name) {
|
|
170
|
+
return MODELS[canonicalModel(name)];
|
|
171
|
+
}
|
|
172
|
+
|
|
52
173
|
// src/slash.ts
|
|
53
174
|
var WASM_INPUT_OFFSET2 = 4096;
|
|
54
175
|
var CALIBRATION = {
|
|
176
|
+
"claude-fable-5.1": 2.05,
|
|
177
|
+
"claude-fable-5": 2.05,
|
|
178
|
+
"claude-mythos-5.1": 2.05,
|
|
179
|
+
"claude-mythos-5": 2.05,
|
|
180
|
+
"claude-opus-5.5": 2.05,
|
|
181
|
+
"claude-opus-5": 2.05,
|
|
182
|
+
"claude-opus-4.8": 2.05,
|
|
183
|
+
"claude-opus-4.6": 2.05,
|
|
184
|
+
"claude-opus-4.5": 2.05,
|
|
185
|
+
"claude-sonnet-5.5": 2.05,
|
|
186
|
+
"claude-sonnet-4.6": 2.05,
|
|
187
|
+
"claude-sonnet-4.5": 2.05,
|
|
55
188
|
"claude-opus": 2.05,
|
|
56
189
|
"claude-opus-4.7": 2.05,
|
|
190
|
+
"claude-sonnet-5": 2.05,
|
|
57
191
|
"claude-sonnet": 2.05,
|
|
192
|
+
"claude-haiku-4.5": 1.45,
|
|
58
193
|
"claude-haiku": 1.45,
|
|
59
194
|
"gemini-3.1-pro": 1.45,
|
|
195
|
+
"gemini-3.1-pro-preview": 1.45,
|
|
196
|
+
"gemini-3.8-flash": 1.45,
|
|
197
|
+
"gemini-3.7-flash": 1.45,
|
|
198
|
+
"gemini-3.6-flash": 1.45,
|
|
199
|
+
"gemini-3.5-flash": 1.45,
|
|
200
|
+
"gemini-3.1-flash-lite": 1.45,
|
|
201
|
+
"gemini-3.5-flash-lite": 1.45,
|
|
60
202
|
"gemini-2.5-flash": 1.45,
|
|
203
|
+
"grok-4.7": 1.15,
|
|
204
|
+
"grok-4.6": 1.15,
|
|
205
|
+
"grok-4.5": 1.15,
|
|
206
|
+
"grok-build-0.1": 1.15,
|
|
207
|
+
"grok-4.3": 1.15,
|
|
61
208
|
"grok-4.20": 1.15,
|
|
62
209
|
"grok-4-1-fast": 1.15,
|
|
210
|
+
"gpt-5.6-sol": 1.15,
|
|
211
|
+
"gpt-5.6-terra": 1.15,
|
|
212
|
+
"gpt-5.6-luna": 1.15,
|
|
63
213
|
"gpt-5.4": 1.15,
|
|
64
214
|
"gpt-5.4-mini": 1.15,
|
|
65
215
|
"gpt-5.4-nano": 1.15
|
|
@@ -73,7 +223,7 @@ function slash(content, model) {
|
|
|
73
223
|
const raw = instance.exports.estimate_tokens(WASM_INPUT_OFFSET2, len);
|
|
74
224
|
if (!model)
|
|
75
225
|
return raw;
|
|
76
|
-
const factor = CALIBRATION[model] ?? DEFAULT_UNKNOWN_MODEL_FACTOR;
|
|
226
|
+
const factor = CALIBRATION[canonicalModel(model)] ?? DEFAULT_UNKNOWN_MODEL_FACTOR;
|
|
77
227
|
return factor === 1 ? raw : Math.ceil(raw * factor);
|
|
78
228
|
}
|
|
79
229
|
|
|
@@ -91,12 +241,12 @@ var AI_PATTERNS = [
|
|
|
91
241
|
{ name: "Mistral", regex: /from\s+['"]@mistralai|MistralClient/g }
|
|
92
242
|
];
|
|
93
243
|
var SDK_REPRESENTATIVE_MODEL = {
|
|
94
|
-
Anthropic: "claude-sonnet",
|
|
95
|
-
OpenAI: "gpt-5.
|
|
244
|
+
Anthropic: "claude-sonnet-5",
|
|
245
|
+
OpenAI: "gpt-5.6-sol",
|
|
96
246
|
Gemini: "gemini-3.1-pro",
|
|
97
|
-
Grok: "grok-4.
|
|
247
|
+
Grok: "grok-4.6"
|
|
98
248
|
};
|
|
99
|
-
var UNKNOWN_SDK_REPRESENTATIVE_MODEL = "claude-sonnet";
|
|
249
|
+
var UNKNOWN_SDK_REPRESENTATIVE_MODEL = "claude-sonnet-5";
|
|
100
250
|
var SKIP_DIRS = new Set([
|
|
101
251
|
"node_modules",
|
|
102
252
|
".git",
|
|
@@ -202,24 +352,6 @@ function scan(dir) {
|
|
|
202
352
|
};
|
|
203
353
|
}
|
|
204
354
|
|
|
205
|
-
// src/models.ts
|
|
206
|
-
var MODELS = {
|
|
207
|
-
"claude-opus": { input: 5, output: 25, context: 1e6 },
|
|
208
|
-
"claude-opus-4.7": { input: 5, output: 25, context: 1e6 },
|
|
209
|
-
"claude-sonnet": { input: 2, output: 10, context: 1e6 },
|
|
210
|
-
"claude-haiku": { input: 1, output: 5, context: 200000 },
|
|
211
|
-
"grok-4.20": { input: 1.25, output: 2.5, context: 1e6, longContextThreshold: 200000, longContextInput: 2.5, longContextOutput: 5 },
|
|
212
|
-
"grok-4-1-fast": { input: 1.25, output: 2.5, context: 1e6, longContextThreshold: 200000, longContextInput: 2.5, longContextOutput: 5 },
|
|
213
|
-
"gemini-3.1-pro": { input: 2, output: 12, context: 1e6 },
|
|
214
|
-
"gemini-2.5-flash": { input: 0.3, output: 2.5, context: 1e6 },
|
|
215
|
-
"gpt-5.4": { input: 2.5, output: 15, context: 1e6 },
|
|
216
|
-
"gpt-5.4-mini": { input: 0.75, output: 4.5, context: 128000 },
|
|
217
|
-
"gpt-5.4-nano": { input: 0.2, output: 1.25, context: 128000 }
|
|
218
|
-
};
|
|
219
|
-
function getModel(name) {
|
|
220
|
-
return MODELS[name] || MODELS[name.toLowerCase()];
|
|
221
|
-
}
|
|
222
|
-
|
|
223
355
|
// src/report.ts
|
|
224
356
|
var R = "\x1B[0m";
|
|
225
357
|
var B = "\x1B[1m";
|
|
@@ -244,6 +376,7 @@ function printReport(sites, filesScanned, timeMs, cwd) {
|
|
|
244
376
|
if (sites.length === 0) {
|
|
245
377
|
console.log(`${WHITE} No AI API call sites detected.${R}`);
|
|
246
378
|
console.log(`${GRAY} Supported: OpenAI, Anthropic, Vercel AI, LangChain, Gemini, Bedrock, Grok${R}`);
|
|
379
|
+
console.log(`${GRAY} Run this in a project that already calls an LLM.${R}`);
|
|
247
380
|
console.log("");
|
|
248
381
|
return;
|
|
249
382
|
}
|
|
@@ -315,6 +448,14 @@ function init(opts) {
|
|
|
315
448
|
|
|
316
449
|
// src/cli.ts
|
|
317
450
|
var args = process.argv.slice(2);
|
|
451
|
+
if (args.includes("--version") || args.includes("-V")) {
|
|
452
|
+
console.log(packageVersion());
|
|
453
|
+
process.exit(0);
|
|
454
|
+
}
|
|
455
|
+
if (args.includes("--help") || args.includes("-h")) {
|
|
456
|
+
printHelp();
|
|
457
|
+
process.exit(0);
|
|
458
|
+
}
|
|
318
459
|
var keyArg = args.find((a) => a.startsWith("--key="));
|
|
319
460
|
var key = keyArg?.split("=")[1] || process.env.SLASH_KEY;
|
|
320
461
|
if (key)
|
|
@@ -336,3 +477,21 @@ if (sites.length > 0) {
|
|
|
336
477
|
}
|
|
337
478
|
console.log("");
|
|
338
479
|
}
|
|
480
|
+
function packageVersion() {
|
|
481
|
+
const here = dirname(fileURLToPath(import.meta.url));
|
|
482
|
+
const pkg = JSON.parse(readFileSync2(join2(here, "..", "package.json"), "utf8"));
|
|
483
|
+
return pkg.version;
|
|
484
|
+
}
|
|
485
|
+
function printHelp() {
|
|
486
|
+
console.log(`slash-tokens — Token Optimization for Context Engineers
|
|
487
|
+
|
|
488
|
+
bunx slash-tokens try (scan this directory, no account)
|
|
489
|
+
npm install slash-tokens SDK
|
|
490
|
+
|
|
491
|
+
--version, -V print version and exit
|
|
492
|
+
--help, -h print this help and exit
|
|
493
|
+
--key=KEY optional; CLI scans never charge
|
|
494
|
+
|
|
495
|
+
Run in a project that already calls an LLM.
|
|
496
|
+
Empty folder → no call sites. That's normal.`);
|
|
497
|
+
}
|
package/dist/intercept.d.ts
CHANGED
|
@@ -11,6 +11,14 @@ export interface InterceptEvent {
|
|
|
11
11
|
routed: boolean;
|
|
12
12
|
timestamp: string;
|
|
13
13
|
}
|
|
14
|
+
/**
|
|
15
|
+
* The model a request actually names, for PRICING: the exact table entry when
|
|
16
|
+
* the strict canonical ID matches (claude-opus-5-5 → claude-opus-5.5,
|
|
17
|
+
* gpt-6-luna), else the legacy family mapping. Routing keeps using
|
|
18
|
+
* normalizeModel() so /auto decisions are unchanged within 1.6.x.
|
|
19
|
+
*/
|
|
20
|
+
export declare function identifyModel(raw: string): string;
|
|
14
21
|
export declare function normalizeModel(raw: string): string;
|
|
22
|
+
export declare function findCheapestRoute(provider: string, tokens: number, currentModel: string): string | null;
|
|
15
23
|
export declare function onIntercept(handler: (event: InterceptEvent) => void): void;
|
|
16
24
|
export declare function patchFetch(): void;
|
package/dist/intercept.js
CHANGED
|
@@ -1,20 +1,29 @@
|
|
|
1
1
|
import { slash } from './slash.js';
|
|
2
|
-
import { getModel } from './models.js';
|
|
2
|
+
import { getModel, MODELS, canonicalModel } from './models.js';
|
|
3
3
|
import { shouldRoute, isModelAllowed } from './config.js';
|
|
4
|
-
import { PROVIDER_MODELS } from './providers.js';
|
|
4
|
+
import { PROVIDER_MODELS, AUTO_ROUTE_TARGETS, NOT_ROUTE_TARGETS } from './providers.js';
|
|
5
5
|
// Reverse lookup: model name → provider model names in the API
|
|
6
6
|
// (what to put back in the request body)
|
|
7
7
|
const MODEL_API_NAMES = {
|
|
8
|
+
'claude-opus-5': 'claude-opus-5',
|
|
8
9
|
'claude-opus': 'claude-opus-5',
|
|
9
10
|
'claude-opus-4.7': 'claude-opus-4-7',
|
|
11
|
+
'claude-sonnet-5': 'claude-sonnet-5',
|
|
10
12
|
'claude-sonnet': 'claude-sonnet-5',
|
|
13
|
+
'claude-haiku-4.5': 'claude-haiku-4-5-20251001',
|
|
11
14
|
'claude-haiku': 'claude-haiku-4-5-20251001',
|
|
15
|
+
'gpt-5.6-sol': 'gpt-5.6-sol',
|
|
16
|
+
'gpt-5.6-terra': 'gpt-5.6-terra',
|
|
17
|
+
'gpt-5.6-luna': 'gpt-5.6-luna',
|
|
12
18
|
'gpt-5.4': 'gpt-5.4',
|
|
13
19
|
'gpt-5.4-mini': 'gpt-5.4-mini',
|
|
14
20
|
'gpt-5.4-nano': 'gpt-5.4-nano',
|
|
21
|
+
'grok-4.6': 'grok-4.6',
|
|
22
|
+
'grok-4.3': 'grok-4.3',
|
|
15
23
|
'grok-4.20': 'grok-4.20-0309-non-reasoning',
|
|
16
24
|
'grok-4-1-fast': 'grok-4.3',
|
|
17
25
|
'gemini-3.1-pro': 'gemini-pro-latest',
|
|
26
|
+
'gemini-3.5-flash-lite': 'gemini-3.5-flash-lite',
|
|
18
27
|
'gemini-2.5-flash': 'gemini-flash-latest',
|
|
19
28
|
};
|
|
20
29
|
// AI API endpoint detection
|
|
@@ -27,7 +36,7 @@ const AI_ENDPOINTS = [
|
|
|
27
36
|
{
|
|
28
37
|
pattern: /api\.openai\.com/,
|
|
29
38
|
provider: 'OpenAI',
|
|
30
|
-
modelExtractor: (body) => body?.model || 'gpt-5.
|
|
39
|
+
modelExtractor: (body) => body?.model || 'gpt-5.6-sol',
|
|
31
40
|
},
|
|
32
41
|
{
|
|
33
42
|
pattern: /generativelanguage\.googleapis\.com/,
|
|
@@ -35,13 +44,13 @@ const AI_ENDPOINTS = [
|
|
|
35
44
|
modelExtractor: (_body, url) => {
|
|
36
45
|
// Model is in the URL path: /v1beta/models/gemini-2.0-flash:generateContent
|
|
37
46
|
const match = url?.match(/\/models\/([^/:]+)/);
|
|
38
|
-
return match ? match[1] : 'gemini-
|
|
47
|
+
return match ? match[1] : 'gemini-3.5-flash-lite';
|
|
39
48
|
},
|
|
40
49
|
},
|
|
41
50
|
{
|
|
42
51
|
pattern: /api\.x\.ai/,
|
|
43
52
|
provider: 'xAI',
|
|
44
|
-
modelExtractor: (body) => body?.model || 'grok-4.
|
|
53
|
+
modelExtractor: (body) => body?.model || 'grok-4.6',
|
|
45
54
|
},
|
|
46
55
|
];
|
|
47
56
|
// Normalize model names to our pricing table keys.
|
|
@@ -53,35 +62,66 @@ const AI_ENDPOINTS = [
|
|
|
53
62
|
// (getModel() returning undefined → $0 reported cost, see slash.ts's
|
|
54
63
|
// DEFAULT_UNKNOWN_MODEL_FACTOR comment for why "unrecognized" defaulting
|
|
55
64
|
// to a falsely-safe-looking value is the dangerous case).
|
|
65
|
+
/**
|
|
66
|
+
* The model a request actually names, for PRICING: the exact table entry when
|
|
67
|
+
* the strict canonical ID matches (claude-opus-5-5 → claude-opus-5.5,
|
|
68
|
+
* gpt-6-luna), else the legacy family mapping. Routing keeps using
|
|
69
|
+
* normalizeModel() so /auto decisions are unchanged within 1.6.x.
|
|
70
|
+
*/
|
|
71
|
+
export function identifyModel(raw) {
|
|
72
|
+
const c = canonicalModel(raw);
|
|
73
|
+
return MODELS[c] ? c : normalizeModel(raw);
|
|
74
|
+
}
|
|
56
75
|
export function normalizeModel(raw) {
|
|
57
76
|
const lower = raw.toLowerCase();
|
|
58
|
-
// Anthropic —
|
|
77
|
+
// Anthropic — specific versions before generic family
|
|
59
78
|
if (lower.includes('opus') && (lower.includes('4-7') || lower.includes('4.7')))
|
|
60
79
|
return 'claude-opus-4.7';
|
|
80
|
+
if (lower.includes('opus-5'))
|
|
81
|
+
return 'claude-opus-5';
|
|
61
82
|
if (lower.includes('opus'))
|
|
62
83
|
return 'claude-opus';
|
|
84
|
+
if (lower.includes('sonnet-5'))
|
|
85
|
+
return 'claude-sonnet-5';
|
|
63
86
|
if (lower.includes('sonnet'))
|
|
64
87
|
return 'claude-sonnet';
|
|
88
|
+
if (lower.includes('haiku') && (lower.includes('4.5') || lower.includes('4-5')))
|
|
89
|
+
return 'claude-haiku-4.5';
|
|
65
90
|
if (lower.includes('haiku'))
|
|
66
91
|
return 'claude-haiku';
|
|
67
|
-
// xAI
|
|
92
|
+
// xAI — 4.6 / 4.3 / 4.20 before generic grok
|
|
93
|
+
if (lower.includes('grok') && (lower.includes('4.6') || lower.includes('4-6')))
|
|
94
|
+
return 'grok-4.6';
|
|
95
|
+
if (lower.includes('grok') && (lower.includes('4.3') || lower.includes('4-3')))
|
|
96
|
+
return 'grok-4.3';
|
|
97
|
+
if (lower.includes('grok') && (lower.includes('4.20') || lower.includes('4-20')))
|
|
98
|
+
return 'grok-4.20';
|
|
68
99
|
if (lower.includes('grok') && lower.includes('fast'))
|
|
69
100
|
return 'grok-4-1-fast';
|
|
70
101
|
if (lower.includes('grok'))
|
|
71
|
-
return 'grok-4.
|
|
102
|
+
return 'grok-4.6';
|
|
72
103
|
// Google
|
|
73
104
|
if (lower.includes('gemini') && lower.includes('pro'))
|
|
74
105
|
return 'gemini-3.1-pro';
|
|
75
|
-
if (lower.includes('gemini'))
|
|
106
|
+
if (lower.includes('gemini') && lower.includes('3.5') && lower.includes('lite'))
|
|
107
|
+
return 'gemini-3.5-flash-lite';
|
|
108
|
+
if (lower.includes('gemini') && lower.includes('2.5'))
|
|
76
109
|
return 'gemini-2.5-flash';
|
|
77
|
-
|
|
110
|
+
if (lower.includes('gemini'))
|
|
111
|
+
return 'gemini-3.5-flash-lite';
|
|
112
|
+
// OpenAI — 5.6 then 5.4 then legacy
|
|
113
|
+
if (lower.includes('5.6') && lower.includes('luna'))
|
|
114
|
+
return 'gpt-5.6-luna';
|
|
115
|
+
if (lower.includes('5.6') && lower.includes('terra'))
|
|
116
|
+
return 'gpt-5.6-terra';
|
|
117
|
+
if (lower.includes('5.6'))
|
|
118
|
+
return 'gpt-5.6-sol';
|
|
78
119
|
if (lower.includes('5.4') && lower.includes('nano'))
|
|
79
120
|
return 'gpt-5.4-nano';
|
|
80
121
|
if (lower.includes('5.4') && lower.includes('mini'))
|
|
81
122
|
return 'gpt-5.4-mini';
|
|
82
123
|
if (lower.includes('5.4'))
|
|
83
124
|
return 'gpt-5.4';
|
|
84
|
-
// OpenAI — legacy model names → map to closest current equivalent
|
|
85
125
|
if (lower.includes('o1-mini') || lower.includes('o1_mini'))
|
|
86
126
|
return 'gpt-5.4-mini';
|
|
87
127
|
if (lower.includes('o1'))
|
|
@@ -118,7 +158,7 @@ function extractContent(body) {
|
|
|
118
158
|
return JSON.stringify(body);
|
|
119
159
|
}
|
|
120
160
|
// Find cheapest model from same provider that fits
|
|
121
|
-
function findCheapestRoute(provider, tokens, currentModel) {
|
|
161
|
+
export function findCheapestRoute(provider, tokens, currentModel) {
|
|
122
162
|
const providerModels = PROVIDER_MODELS[provider];
|
|
123
163
|
if (!providerModels)
|
|
124
164
|
return null;
|
|
@@ -126,6 +166,10 @@ function findCheapestRoute(provider, tokens, currentModel) {
|
|
|
126
166
|
for (const model of providerModels) {
|
|
127
167
|
if (model === currentModel)
|
|
128
168
|
continue;
|
|
169
|
+
if (!AUTO_ROUTE_TARGETS.has(model))
|
|
170
|
+
continue; // 1.6.x: /auto rewrites only to its 1.6.5 targets
|
|
171
|
+
if (NOT_ROUTE_TARGETS.has(model))
|
|
172
|
+
continue; // specialised models are never targets
|
|
129
173
|
if (!isModelAllowed(model))
|
|
130
174
|
continue; // user excluded this model
|
|
131
175
|
const info = getModel(model);
|
|
@@ -158,13 +202,14 @@ export function patchFetch() {
|
|
|
158
202
|
const body = JSON.parse(bodyStr);
|
|
159
203
|
const content = extractContent(body);
|
|
160
204
|
const rawModel = match.modelExtractor(body, url);
|
|
161
|
-
const
|
|
205
|
+
const routingModel = normalizeModel(rawModel); // decides routes (unchanged in 1.6.x)
|
|
206
|
+
const originalModel = identifyModel(rawModel); // prices the request it really names
|
|
162
207
|
const tokens = slash(content, originalModel);
|
|
163
208
|
const originalInfo = getModel(originalModel);
|
|
164
209
|
const originalCost = originalInfo ? Math.round(((tokens / 1000000) * originalInfo.input) * 1000000) / 1000000 : 0;
|
|
165
210
|
const fits = originalInfo ? tokens <= originalInfo.context : true;
|
|
166
211
|
// Find cheapest route within same provider (if routing enabled)
|
|
167
|
-
const routeModel = shouldRoute() ? findCheapestRoute(match.provider, tokens,
|
|
212
|
+
const routeModel = shouldRoute() ? findCheapestRoute(match.provider, tokens, routingModel) : null;
|
|
168
213
|
const routedInfo = routeModel ? getModel(routeModel) : null;
|
|
169
214
|
const routedCost = routedInfo ? Math.round(((tokens / 1000000) * routedInfo.input) * 1000000) / 1000000 : originalCost;
|
|
170
215
|
const salvaged = routeModel ? Math.round((originalCost - routedCost) * 1000000) / 1000000 : 0;
|
package/dist/models.d.ts
CHANGED
|
@@ -7,6 +7,13 @@ export interface ModelInfo {
|
|
|
7
7
|
longContextOutput?: number;
|
|
8
8
|
}
|
|
9
9
|
export declare const MODELS: Record<string, ModelInfo>;
|
|
10
|
+
/**
|
|
11
|
+
* Real API IDs → table keys, strictly: lowercase, drop a trailing date stamp
|
|
12
|
+
* (`-20250514`), and write version numbers with dots (`claude-opus-4-7` →
|
|
13
|
+
* `claude-opus-4.7`). No family guessing: an unknown version stays unknown,
|
|
14
|
+
* so a new model never silently gets an older model's price.
|
|
15
|
+
*/
|
|
16
|
+
export declare function canonicalModel(name: string): string;
|
|
10
17
|
export declare function getModel(name: string): ModelInfo | undefined;
|
|
11
18
|
export declare function effectiveRate(tokens: number, info: ModelInfo): {
|
|
12
19
|
input: number;
|
package/dist/models.js
CHANGED
|
@@ -1,44 +1,119 @@
|
|
|
1
|
-
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
1
|
+
const OPUS = { input: 5.00, output: 25.00, context: 1000000 };
|
|
2
|
+
const OPUS_55 = { input: 4.00, output: 20.00, context: 1000000 };
|
|
3
|
+
const FABLE = { input: 10.00, output: 50.00, context: 1000000 };
|
|
4
|
+
const SONNET_4X = { input: 3.00, output: 15.00, context: 1000000 };
|
|
5
|
+
const SONNET = { input: 2.00, output: 10.00, context: 1000000 };
|
|
6
|
+
const HAIKU = { input: 1.00, output: 5.00, context: 200000 };
|
|
7
|
+
const GROK_46 = {
|
|
8
|
+
input: 2.00, output: 6.00, context: 500000,
|
|
9
|
+
longContextThreshold: 200000, longContextInput: 4.00, longContextOutput: 12.00,
|
|
10
|
+
};
|
|
11
|
+
const GROK_43 = {
|
|
12
|
+
input: 1.25, output: 2.50, context: 1000000,
|
|
13
|
+
longContextThreshold: 200000, longContextInput: 2.50, longContextOutput: 5.00,
|
|
14
|
+
};
|
|
15
|
+
const GEMINI_PRO = {
|
|
16
|
+
input: 2.00, output: 12.00, context: 1000000,
|
|
17
|
+
longContextThreshold: 200000, longContextInput: 4.00, longContextOutput: 18.00,
|
|
18
|
+
};
|
|
19
|
+
const GEMINI_FLASH = { input: 0.30, output: 2.50, context: 1000000 };
|
|
20
|
+
// Gemini 3.6–3.8 Flash: launch price through 2026-12-31; Google lists $1.50/$7.50
|
|
21
|
+
// from 2027-01-01. Update this entry before then.
|
|
22
|
+
const GEMINI_FLASH_3X = { input: 0.75, output: 3.75, context: 1000000 };
|
|
23
|
+
const GEMINI_35_FLASH = { input: 1.50, output: 9.00, context: 1000000 };
|
|
24
|
+
const GEMINI_31_FLASH_LITE = { input: 0.25, output: 1.50, context: 1000000 };
|
|
25
|
+
const GROK_BUILD = {
|
|
26
|
+
input: 1.00, output: 2.00, context: 256000,
|
|
27
|
+
longContextThreshold: 200000, longContextInput: 2.00, longContextOutput: 4.00,
|
|
28
|
+
};
|
|
29
|
+
const GPT_6_ASTRA = { input: 10.00, output: 50.00, context: 1050000 };
|
|
30
|
+
const GPT_6_SOL = { input: 2.00, output: 10.00, context: 1050000 };
|
|
31
|
+
const GPT_6_LUNA = { input: 0.10, output: 0.50, context: 1050000 };
|
|
32
|
+
const GPT_SOL = { input: 4.00, output: 20.00, context: 1050000 };
|
|
33
|
+
const GPT_TERRA = { input: 2.00, output: 12.00, context: 1050000 };
|
|
34
|
+
const GPT_LUNA = { input: 0.20, output: 1.20, context: 1050000 };
|
|
35
|
+
const GPT_54 = { input: 2.50, output: 15.00, context: 1000000 };
|
|
36
|
+
const GPT_54_MINI = { input: 0.75, output: 4.50, context: 128000 };
|
|
37
|
+
const GPT_54_NANO = { input: 0.20, output: 1.25, context: 128000 };
|
|
38
|
+
// Pricing as of 2026-10-07 — USD per million tokens (checked against the four pages below).
|
|
39
|
+
// First-party: platform.claude.com/docs/en/about-claude/pricing
|
|
40
|
+
// developers.openai.com/api/docs/models
|
|
41
|
+
// docs.x.ai/developers/models
|
|
42
|
+
// ai.google.dev/gemini-api/docs/pricing
|
|
43
|
+
// Old keys stay as aliases so existing call sites don't throw.
|
|
17
44
|
export const MODELS = {
|
|
18
|
-
// Anthropic
|
|
19
|
-
'claude-
|
|
20
|
-
'claude-
|
|
21
|
-
'claude-
|
|
22
|
-
'claude-
|
|
23
|
-
|
|
24
|
-
'
|
|
25
|
-
'
|
|
45
|
+
// Anthropic — live names + generic aliases (same rates)
|
|
46
|
+
'claude-fable-5.1': { ...FABLE },
|
|
47
|
+
'claude-fable-5': { ...FABLE },
|
|
48
|
+
'claude-mythos-5.1': { ...FABLE },
|
|
49
|
+
'claude-mythos-5': { ...FABLE },
|
|
50
|
+
'claude-opus-5.5': { ...OPUS_55 },
|
|
51
|
+
'claude-opus-5': { ...OPUS },
|
|
52
|
+
'claude-opus-4.8': { ...OPUS },
|
|
53
|
+
'claude-opus': { ...OPUS },
|
|
54
|
+
'claude-opus-4.7': { ...OPUS },
|
|
55
|
+
'claude-opus-4.6': { ...OPUS },
|
|
56
|
+
'claude-opus-4.5': { ...OPUS },
|
|
57
|
+
'claude-sonnet-5.5': { ...SONNET },
|
|
58
|
+
'claude-sonnet-5': { ...SONNET },
|
|
59
|
+
'claude-sonnet': { ...SONNET },
|
|
60
|
+
'claude-sonnet-4.6': { ...SONNET_4X },
|
|
61
|
+
'claude-sonnet-4.5': { ...SONNET_4X },
|
|
62
|
+
'claude-haiku-4.5': { ...HAIKU },
|
|
63
|
+
'claude-haiku': { ...HAIKU },
|
|
64
|
+
// xAI — flagship 4.6, cheap same-provider 4.3. 4.20 / fast are aliases.
|
|
65
|
+
'grok-4.7': { ...GROK_46 },
|
|
66
|
+
'grok-4.6': { ...GROK_46 },
|
|
67
|
+
'grok-4.5': { ...GROK_46 },
|
|
68
|
+
'grok-build-0.1': { ...GROK_BUILD },
|
|
69
|
+
'grok-4.3': { ...GROK_43 },
|
|
70
|
+
'grok-4.20': { ...GROK_43 },
|
|
71
|
+
'grok-4-1-fast': { ...GROK_43 },
|
|
26
72
|
// Google
|
|
27
|
-
'gemini-3.1-pro': {
|
|
28
|
-
'gemini-
|
|
29
|
-
|
|
30
|
-
'
|
|
31
|
-
'
|
|
32
|
-
'
|
|
73
|
+
'gemini-3.1-pro': { ...GEMINI_PRO },
|
|
74
|
+
'gemini-3.1-pro-preview': { ...GEMINI_PRO },
|
|
75
|
+
'gemini-3.8-flash': { ...GEMINI_FLASH_3X },
|
|
76
|
+
'gemini-3.7-flash': { ...GEMINI_FLASH_3X },
|
|
77
|
+
'gemini-3.6-flash': { ...GEMINI_FLASH_3X },
|
|
78
|
+
'gemini-3.5-flash': { ...GEMINI_35_FLASH },
|
|
79
|
+
'gemini-3.1-flash-lite': { ...GEMINI_31_FLASH_LITE },
|
|
80
|
+
'gemini-3.5-flash-lite': { ...GEMINI_FLASH },
|
|
81
|
+
'gemini-2.5-flash': { ...GEMINI_FLASH },
|
|
82
|
+
// OpenAI — live 5.6 ladder. 5.4 family kept as aliases (old prices).
|
|
83
|
+
'gpt-6-astra': { ...GPT_6_ASTRA },
|
|
84
|
+
'gpt-6.1-sol': { ...GPT_6_SOL },
|
|
85
|
+
'gpt-6-sol': { ...GPT_6_SOL },
|
|
86
|
+
'gpt-6-luna': { ...GPT_6_LUNA },
|
|
87
|
+
'gpt-5.6-sol': { ...GPT_SOL },
|
|
88
|
+
'gpt-5.6-terra': { ...GPT_TERRA },
|
|
89
|
+
'gpt-5.6-luna': { ...GPT_LUNA },
|
|
90
|
+
'gpt-5.4': { ...GPT_54 },
|
|
91
|
+
'gpt-5.4-mini': { ...GPT_54_MINI },
|
|
92
|
+
'gpt-5.4-nano': { ...GPT_54_NANO },
|
|
33
93
|
};
|
|
94
|
+
/**
|
|
95
|
+
* Real API IDs → table keys, strictly: lowercase, drop a trailing date stamp
|
|
96
|
+
* (`-20250514`), and write version numbers with dots (`claude-opus-4-7` →
|
|
97
|
+
* `claude-opus-4.7`). No family guessing: an unknown version stays unknown,
|
|
98
|
+
* so a new model never silently gets an older model's price.
|
|
99
|
+
*/
|
|
100
|
+
export function canonicalModel(name) {
|
|
101
|
+
if (MODELS[name])
|
|
102
|
+
return name;
|
|
103
|
+
let n = name.trim().toLowerCase();
|
|
104
|
+
if (MODELS[n])
|
|
105
|
+
return n;
|
|
106
|
+
n = n.replace(/-\d{8}$/, '');
|
|
107
|
+
if (MODELS[n])
|
|
108
|
+
return n;
|
|
109
|
+
const dotted = n.replace(/(\d)-(\d)(?=$|-)/g, '$1.$2');
|
|
110
|
+
if (MODELS[dotted])
|
|
111
|
+
return dotted;
|
|
112
|
+
return n;
|
|
113
|
+
}
|
|
34
114
|
export function getModel(name) {
|
|
35
|
-
return MODELS[name
|
|
115
|
+
return MODELS[canonicalModel(name)];
|
|
36
116
|
}
|
|
37
|
-
// Resolve the actual billable rate for a given token count — applies the
|
|
38
|
-
// long-context tier above if the model has one and tokens cross it.
|
|
39
|
-
// preflight()/preflightRoute() should go through this, not read
|
|
40
|
-
// .input/.output directly, or a Grok call over 200K tokens gets silently
|
|
41
|
-
// under-costed at the base rate.
|
|
42
117
|
export function effectiveRate(tokens, info) {
|
|
43
118
|
if (info.longContextThreshold !== undefined && tokens > info.longContextThreshold) {
|
|
44
119
|
return {
|
package/dist/patterns.d.ts
CHANGED
|
@@ -4,6 +4,6 @@ export interface Pattern {
|
|
|
4
4
|
}
|
|
5
5
|
export declare const AI_PATTERNS: Pattern[];
|
|
6
6
|
export declare const SDK_REPRESENTATIVE_MODEL: Record<string, string>;
|
|
7
|
-
export declare const UNKNOWN_SDK_REPRESENTATIVE_MODEL = "claude-sonnet";
|
|
7
|
+
export declare const UNKNOWN_SDK_REPRESENTATIVE_MODEL = "claude-sonnet-5";
|
|
8
8
|
export declare const SKIP_DIRS: Set<string>;
|
|
9
9
|
export declare const SCAN_EXTENSIONS: Set<string>;
|
package/dist/patterns.js
CHANGED
|
@@ -34,17 +34,17 @@ export const AI_PATTERNS = [
|
|
|
34
34
|
// know the exact model), but a real per-provider one instead of a single
|
|
35
35
|
// guess applied to everyone.
|
|
36
36
|
export const SDK_REPRESENTATIVE_MODEL = {
|
|
37
|
-
'Anthropic': 'claude-sonnet',
|
|
38
|
-
'OpenAI': 'gpt-5.
|
|
37
|
+
'Anthropic': 'claude-sonnet-5',
|
|
38
|
+
'OpenAI': 'gpt-5.6-sol',
|
|
39
39
|
'Gemini': 'gemini-3.1-pro',
|
|
40
|
-
'Grok': 'grok-4.
|
|
40
|
+
'Grok': 'grok-4.6',
|
|
41
41
|
};
|
|
42
42
|
// Fallback for SDKs that don't map to one specific provider (Vercel AI,
|
|
43
43
|
// LangChain, and Bedrock can all wrap any underlying provider; raw
|
|
44
44
|
// fetch-to-AI-endpoint and Cohere/Mistral have no pricing data in MODELS
|
|
45
45
|
// at all). claude-sonnet is used as a documented, honest middle-of-the-
|
|
46
46
|
// road placeholder — not a claim about which model is actually running.
|
|
47
|
-
export const UNKNOWN_SDK_REPRESENTATIVE_MODEL = 'claude-sonnet';
|
|
47
|
+
export const UNKNOWN_SDK_REPRESENTATIVE_MODEL = 'claude-sonnet-5';
|
|
48
48
|
export const SKIP_DIRS = new Set([
|
|
49
49
|
'node_modules', '.git', 'dist', 'build', '.next', '.nuxt', '.svelte-kit',
|
|
50
50
|
'coverage', '.turbo', '.cache', '__pycache__', '.venv', 'venv',
|
package/dist/preflight.js
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
import { slash } from './slash.js';
|
|
2
|
-
import { getModel, MODELS, effectiveRate } from './models.js';
|
|
3
|
-
import { PROVIDER_MODELS, providerOf } from './providers.js';
|
|
2
|
+
import { getModel, MODELS, effectiveRate, canonicalModel } from './models.js';
|
|
3
|
+
import { PROVIDER_MODELS, providerOf, NOT_ROUTE_TARGETS } from './providers.js';
|
|
4
4
|
import { shouldRoute, isModelAllowed } from './config.js';
|
|
5
5
|
/**
|
|
6
6
|
* Compute cost of a prompt of `tokens` tokens on the given ModelInfo.
|
|
@@ -99,9 +99,12 @@ export function preflightRoute(content, model) {
|
|
|
99
99
|
return null;
|
|
100
100
|
const originalCost = computeCost(tokens, info);
|
|
101
101
|
let cheapest = null;
|
|
102
|
+
const self = canonicalModel(model);
|
|
102
103
|
for (const m of providerModels) {
|
|
103
|
-
if (m ===
|
|
104
|
+
if (m === self)
|
|
104
105
|
continue;
|
|
106
|
+
if (NOT_ROUTE_TARGETS.has(m))
|
|
107
|
+
continue; // specialised: never a target
|
|
105
108
|
if (!isModelAllowed(m))
|
|
106
109
|
continue; // user excluded this model
|
|
107
110
|
const altInfo = getModel(m);
|
|
@@ -112,7 +115,10 @@ export function preflightRoute(content, model) {
|
|
|
112
115
|
if (altInfo.input >= info.input)
|
|
113
116
|
continue; // not cheaper
|
|
114
117
|
const alt = buildAlternative(m, originalCost, tokens, altInfo);
|
|
115
|
-
|
|
118
|
+
// Tiny prompts round costs to the same value; break ties on the list price
|
|
119
|
+
// so the genuinely cheaper model wins (gpt-5.6-sol → gpt-6-luna, not 5.4-nano).
|
|
120
|
+
if (!cheapest || alt.cost < cheapest.cost ||
|
|
121
|
+
(alt.cost === cheapest.cost && altInfo.input < (getModel(cheapest.model)?.input ?? Infinity))) {
|
|
116
122
|
cheapest = alt;
|
|
117
123
|
}
|
|
118
124
|
}
|
package/dist/providers.d.ts
CHANGED
|
@@ -1,26 +1,24 @@
|
|
|
1
1
|
/**
|
|
2
2
|
* Provider groups — single source of truth.
|
|
3
3
|
*
|
|
4
|
-
* Slash routing is always SAME-PROVIDER. Opus → Haiku
|
|
5
|
-
* Grok
|
|
4
|
+
* Slash routing is always SAME-PROVIDER. Opus 5 → Haiku 4.5, Sol → Luna,
|
|
5
|
+
* Grok 4.6 → 4.3, Gemini 3.1 Pro → 3.5 Flash-Lite. Never cross-provider.
|
|
6
6
|
*
|
|
7
|
-
*
|
|
8
|
-
*
|
|
9
|
-
* prediction) consume from here so they can never drift.
|
|
10
|
-
*
|
|
11
|
-
* TEST-NOTE: Whenever a new model is added, it MUST appear in exactly one
|
|
12
|
-
* provider group below. A model missing from here will:
|
|
13
|
-
* - Never be a routing target from `findCheapestRoute` / `preflightRoute`
|
|
14
|
-
* - Still appear in `preflight().options` (which is cross-provider analysis)
|
|
15
|
-
* That mismatch is by design — see preflight.ts semantics.
|
|
7
|
+
* Order inside a group matters when two models share a price: the first
|
|
8
|
+
* strictly-cheaper hit wins (findCheapestRoute / preflightRoute).
|
|
16
9
|
*/
|
|
17
10
|
export declare const PROVIDER_MODELS: Record<string, string[]>;
|
|
18
11
|
/**
|
|
19
|
-
*
|
|
20
|
-
*
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
*
|
|
12
|
+
* Priced and grouped, but never a routing TARGET: specialised models a general
|
|
13
|
+
* call shouldn't be moved to (grok-build-0.1 is a coding-agent model).
|
|
14
|
+
*/
|
|
15
|
+
export declare const NOT_ROUTE_TARGETS: ReadonlySet<string>;
|
|
16
|
+
/**
|
|
17
|
+
* Models `slash-tokens/auto` may rewrite a live request TO. Frozen at the
|
|
18
|
+
* 1.6.5 set so a patch release never changes where production calls go:
|
|
19
|
+
* models added in 1.6.6 (GPT-6, Gemini 3.1 Flash-Lite / 3.6–3.8 Flash, Grok
|
|
20
|
+
* 4.7 / 4.5, Claude 5.5 / Fable …) are priced and shown by preflight() /
|
|
21
|
+
* preflightRoute(), but /auto only starts routing to them in 1.7.0.
|
|
25
22
|
*/
|
|
23
|
+
export declare const AUTO_ROUTE_TARGETS: ReadonlySet<string>;
|
|
26
24
|
export declare function providerOf(model: string): string | null;
|
package/dist/providers.js
CHANGED
|
@@ -1,36 +1,58 @@
|
|
|
1
|
+
import { canonicalModel } from './models.js';
|
|
1
2
|
/**
|
|
2
3
|
* Provider groups — single source of truth.
|
|
3
4
|
*
|
|
4
|
-
* Slash routing is always SAME-PROVIDER. Opus → Haiku
|
|
5
|
-
* Grok
|
|
5
|
+
* Slash routing is always SAME-PROVIDER. Opus 5 → Haiku 4.5, Sol → Luna,
|
|
6
|
+
* Grok 4.6 → 4.3, Gemini 3.1 Pro → 3.5 Flash-Lite. Never cross-provider.
|
|
6
7
|
*
|
|
7
|
-
*
|
|
8
|
-
*
|
|
9
|
-
* prediction) consume from here so they can never drift.
|
|
10
|
-
*
|
|
11
|
-
* TEST-NOTE: Whenever a new model is added, it MUST appear in exactly one
|
|
12
|
-
* provider group below. A model missing from here will:
|
|
13
|
-
* - Never be a routing target from `findCheapestRoute` / `preflightRoute`
|
|
14
|
-
* - Still appear in `preflight().options` (which is cross-provider analysis)
|
|
15
|
-
* That mismatch is by design — see preflight.ts semantics.
|
|
8
|
+
* Order inside a group matters when two models share a price: the first
|
|
9
|
+
* strictly-cheaper hit wins (findCheapestRoute / preflightRoute).
|
|
16
10
|
*/
|
|
17
11
|
export const PROVIDER_MODELS = {
|
|
18
|
-
Anthropic: [
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
12
|
+
Anthropic: [
|
|
13
|
+
'claude-fable-5.1', 'claude-fable-5', 'claude-mythos-5.1', 'claude-mythos-5',
|
|
14
|
+
'claude-opus-5', 'claude-opus', 'claude-opus-4.8', 'claude-opus-4.7',
|
|
15
|
+
'claude-opus-4.6', 'claude-opus-4.5', 'claude-opus-5.5',
|
|
16
|
+
'claude-sonnet-4.6', 'claude-sonnet-4.5',
|
|
17
|
+
'claude-sonnet-5.5', 'claude-sonnet-5', 'claude-sonnet',
|
|
18
|
+
'claude-haiku', 'claude-haiku-4.5',
|
|
19
|
+
],
|
|
20
|
+
OpenAI: [
|
|
21
|
+
'gpt-6-astra',
|
|
22
|
+
'gpt-5.6-sol', 'gpt-5.6-terra',
|
|
23
|
+
'gpt-5.4', 'gpt-6.1-sol', 'gpt-6-sol', 'gpt-5.4-mini', 'gpt-5.4-nano',
|
|
24
|
+
'gpt-5.6-luna', 'gpt-6-luna',
|
|
25
|
+
],
|
|
26
|
+
xAI: ['grok-4.7', 'grok-4.6', 'grok-4.5', 'grok-4.3', 'grok-4.20', 'grok-4-1-fast', 'grok-build-0.1'],
|
|
27
|
+
Google: [
|
|
28
|
+
'gemini-3.1-pro', 'gemini-3.1-pro-preview',
|
|
29
|
+
'gemini-3.5-flash', 'gemini-3.8-flash', 'gemini-3.7-flash', 'gemini-3.6-flash',
|
|
30
|
+
'gemini-3.5-flash-lite', 'gemini-2.5-flash', 'gemini-3.1-flash-lite',
|
|
31
|
+
],
|
|
22
32
|
};
|
|
23
33
|
/**
|
|
24
|
-
*
|
|
25
|
-
*
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
*
|
|
34
|
+
* Priced and grouped, but never a routing TARGET: specialised models a general
|
|
35
|
+
* call shouldn't be moved to (grok-build-0.1 is a coding-agent model).
|
|
36
|
+
*/
|
|
37
|
+
export const NOT_ROUTE_TARGETS = new Set(['grok-build-0.1']);
|
|
38
|
+
/**
|
|
39
|
+
* Models `slash-tokens/auto` may rewrite a live request TO. Frozen at the
|
|
40
|
+
* 1.6.5 set so a patch release never changes where production calls go:
|
|
41
|
+
* models added in 1.6.6 (GPT-6, Gemini 3.1 Flash-Lite / 3.6–3.8 Flash, Grok
|
|
42
|
+
* 4.7 / 4.5, Claude 5.5 / Fable …) are priced and shown by preflight() /
|
|
43
|
+
* preflightRoute(), but /auto only starts routing to them in 1.7.0.
|
|
30
44
|
*/
|
|
45
|
+
export const AUTO_ROUTE_TARGETS = new Set([
|
|
46
|
+
'claude-opus-5', 'claude-opus', 'claude-opus-4.7', 'claude-sonnet-5', 'claude-sonnet',
|
|
47
|
+
'claude-haiku', 'claude-haiku-4.5',
|
|
48
|
+
'gpt-5.6-sol', 'gpt-5.6-terra', 'gpt-5.4', 'gpt-5.4-mini', 'gpt-5.4-nano', 'gpt-5.6-luna',
|
|
49
|
+
'grok-4.6', 'grok-4.3', 'grok-4.20', 'grok-4-1-fast',
|
|
50
|
+
'gemini-3.1-pro', 'gemini-3.5-flash-lite', 'gemini-2.5-flash',
|
|
51
|
+
]);
|
|
31
52
|
export function providerOf(model) {
|
|
53
|
+
const id = canonicalModel(model);
|
|
32
54
|
for (const [provider, models] of Object.entries(PROVIDER_MODELS)) {
|
|
33
|
-
if (models.includes(
|
|
55
|
+
if (models.includes(id))
|
|
34
56
|
return provider;
|
|
35
57
|
}
|
|
36
58
|
return null;
|
package/dist/report.js
CHANGED
|
@@ -24,6 +24,7 @@ export function printReport(sites, filesScanned, timeMs, cwd) {
|
|
|
24
24
|
if (sites.length === 0) {
|
|
25
25
|
console.log(`${WHITE} No AI API call sites detected.${R}`);
|
|
26
26
|
console.log(`${GRAY} Supported: OpenAI, Anthropic, Vercel AI, LangChain, Gemini, Bedrock, Grok${R}`);
|
|
27
|
+
console.log(`${GRAY} Run this in a project that already calls an LLM.${R}`);
|
|
27
28
|
console.log('');
|
|
28
29
|
return;
|
|
29
30
|
}
|
package/dist/slash.js
CHANGED
|
@@ -1,4 +1,5 @@
|
|
|
1
1
|
import { getInstance, writeToMemory, ensureCapacity } from './wasm.js';
|
|
2
|
+
import { canonicalModel } from './models.js';
|
|
2
3
|
const WASM_INPUT_OFFSET = 4096;
|
|
3
4
|
/**
|
|
4
5
|
* Per-model calibration factors.
|
|
@@ -65,7 +66,10 @@ const WASM_INPUT_OFFSET = 4096;
|
|
|
65
66
|
* DEFAULT_UNKNOWN_MODEL_FACTOR (1.85), a ~60% larger correction
|
|
66
67
|
* than it needed. Still safe either way (1.85 > required
|
|
67
68
|
* minimum), just needlessly inflated for real GPT users.
|
|
68
|
-
* grok-4.20 / grok-4-1-fast: 1.15 —
|
|
69
|
+
* grok-4.6 / grok-4.3 (and aliases grok-4.20 / grok-4-1-fast): 1.15 —
|
|
70
|
+
* 4.6/4.3 IDs added 2026-08-25; factor CARRIED from the
|
|
71
|
+
* 2026-08-23 corpus, not re-measured on the new wire IDs.
|
|
72
|
+
* grok-4.20 / grok-4-1-fast (original): 1.15 — re-verified 2026-08-23 against the
|
|
69
73
|
* 29-sample corpus (20 new samples: more languages, Spanish/
|
|
70
74
|
* Japanese prose, more JSON shapes) and UNCHANGED — same
|
|
71
75
|
* worst case (technical-docs prose, ratio 0.928) as the
|
|
@@ -91,14 +95,43 @@ const WASM_INPUT_OFFSET = 4096;
|
|
|
91
95
|
* Slash must NEVER under-report. Over-reporting is safe (go/no-go only).
|
|
92
96
|
*/
|
|
93
97
|
const CALIBRATION = {
|
|
98
|
+
'claude-fable-5.1': 2.05,
|
|
99
|
+
'claude-fable-5': 2.05,
|
|
100
|
+
'claude-mythos-5.1': 2.05,
|
|
101
|
+
'claude-mythos-5': 2.05,
|
|
102
|
+
'claude-opus-5.5': 2.05,
|
|
103
|
+
'claude-opus-5': 2.05,
|
|
104
|
+
'claude-opus-4.8': 2.05,
|
|
105
|
+
'claude-opus-4.6': 2.05,
|
|
106
|
+
'claude-opus-4.5': 2.05,
|
|
107
|
+
'claude-sonnet-5.5': 2.05,
|
|
108
|
+
'claude-sonnet-4.6': 2.05,
|
|
109
|
+
'claude-sonnet-4.5': 2.05,
|
|
94
110
|
'claude-opus': 2.05,
|
|
95
111
|
'claude-opus-4.7': 2.05,
|
|
112
|
+
'claude-sonnet-5': 2.05,
|
|
96
113
|
'claude-sonnet': 2.05,
|
|
114
|
+
'claude-haiku-4.5': 1.45,
|
|
97
115
|
'claude-haiku': 1.45,
|
|
98
116
|
'gemini-3.1-pro': 1.45,
|
|
117
|
+
'gemini-3.1-pro-preview': 1.45,
|
|
118
|
+
'gemini-3.8-flash': 1.45,
|
|
119
|
+
'gemini-3.7-flash': 1.45,
|
|
120
|
+
'gemini-3.6-flash': 1.45,
|
|
121
|
+
'gemini-3.5-flash': 1.45,
|
|
122
|
+
'gemini-3.1-flash-lite': 1.45,
|
|
123
|
+
'gemini-3.5-flash-lite': 1.45,
|
|
99
124
|
'gemini-2.5-flash': 1.45,
|
|
125
|
+
'grok-4.7': 1.15,
|
|
126
|
+
'grok-4.6': 1.15,
|
|
127
|
+
'grok-4.5': 1.15,
|
|
128
|
+
'grok-build-0.1': 1.15,
|
|
129
|
+
'grok-4.3': 1.15,
|
|
100
130
|
'grok-4.20': 1.15,
|
|
101
131
|
'grok-4-1-fast': 1.15,
|
|
132
|
+
'gpt-5.6-sol': 1.15,
|
|
133
|
+
'gpt-5.6-terra': 1.15,
|
|
134
|
+
'gpt-5.6-luna': 1.15,
|
|
102
135
|
'gpt-5.4': 1.15,
|
|
103
136
|
'gpt-5.4-mini': 1.15,
|
|
104
137
|
'gpt-5.4-nano': 1.15,
|
|
@@ -129,7 +162,10 @@ export function slash(content, model) {
|
|
|
129
162
|
const raw = instance.exports.estimate_tokens(WASM_INPUT_OFFSET, len);
|
|
130
163
|
if (!model)
|
|
131
164
|
return raw;
|
|
132
|
-
|
|
165
|
+
// GPT-6 (gpt-6-astra / 6.1-sol / 6-sol / 6-luna) is deliberately NOT in
|
|
166
|
+
// CALIBRATION: a new generation with an unbenchmarked tokenizer takes the
|
|
167
|
+
// conservative default (never under-report) until a bench run adds it.
|
|
168
|
+
const factor = CALIBRATION[canonicalModel(model)] ?? DEFAULT_UNKNOWN_MODEL_FACTOR;
|
|
133
169
|
return factor === 1.0 ? raw : Math.ceil(raw * factor);
|
|
134
170
|
}
|
|
135
171
|
/**
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "slash-tokens",
|
|
3
|
-
"version": "1.6.
|
|
3
|
+
"version": "1.6.6",
|
|
4
4
|
"description": "Token Optimization for Context Engineers. 4.8 KB WASM. Sub-millisecond. Zero dependencies.",
|
|
5
5
|
"main": "dist/index.js",
|
|
6
6
|
"module": "dist/index.js",
|
|
@@ -26,8 +26,10 @@
|
|
|
26
26
|
],
|
|
27
27
|
"sideEffects": false,
|
|
28
28
|
"scripts": {
|
|
29
|
+
"prepublishOnly": "npm run build",
|
|
29
30
|
"build": "npx tsc && bun build src/cli.ts --outfile=dist/cli.js --target=node",
|
|
30
|
-
"test": "bun test"
|
|
31
|
+
"test": "bun test",
|
|
32
|
+
"check:engines": "node scripts/check-engines.mjs"
|
|
31
33
|
},
|
|
32
34
|
"author": "wolfejam",
|
|
33
35
|
"license": "MIT",
|
|
@@ -54,5 +56,8 @@
|
|
|
54
56
|
"@types/node": "^25.5.0",
|
|
55
57
|
"js-tiktoken": "^1.0.21",
|
|
56
58
|
"typescript": "^6.0.2"
|
|
59
|
+
},
|
|
60
|
+
"engines": {
|
|
61
|
+
"node": ">=22.0.0"
|
|
57
62
|
}
|
|
58
63
|
}
|