slash-tokens 1.6.4 → 1.6.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,10 +1,43 @@
1
1
  # Changelog
2
2
 
3
+ ## [1.6.6] — The Fixed Deal Edition
4
+
5
+ *2026-10-07*
6
+
7
+ `--version` / `--help` answer, and today's models price correctly.
8
+
9
+ Solo $20 mailbox, 10% waived. Team $39 for the data.
10
+
11
+ ### Fixed
12
+ - **`slash-tokens --version` and `--help` print and exit.** In 1.6.5 every flag fell through to a full scan of the current folder, so `--version` run from `~` scanned the home folder. An empty `bunx` run now tells you to run it in an app.
13
+ - **`preflight()` no longer throws on current models or real API IDs.** IDs are canonicalised strictly (`claude-opus-4-7` → `claude-opus-4.7`, date stamps dropped). There is no family guessing: an unknown version still fails with "Unknown model", so a new model never silently gets an older model's price.
14
+ - **`npm publish` always builds first** (`prepublishOnly`), so a stale `dist/` can't ship again (1.6.4).
15
+ - **`report()` sends numbers only.** A test pins its payload to `tokens_estimated`, `tokens_saved`, `model`, `action`, `cost_saved_usd`; prompt content never leaves the machine.
16
+ - **`npm test` runs offline.** The live integration suite (`tests/z-integration.test.ts`, which registers keys on mcpaas.live) runs only with `SLASH_LIVE=1`; the weekly `integration.yml` sets it.
17
+
18
+ ### Changed
19
+ - **Model table, prices as of 2026-10-07** (checked against the Anthropic, OpenAI, xAI and Google pricing pages). Added Claude Opus 5.5 ($4/$20), Sonnet 5.5 ($2/$10), Opus 4.8 / 4.6 / 4.5, Sonnet 4.6 / 4.5, Fable 5 / 5.1 and Mythos 5 / 5.1 ($10/$50); Grok 4.7 and 4.5 ($2/$6, doubling above 200K), grok-build-0.1 ($1/$2); GPT-6 Astra ($10/$50), GPT-6.1 Sol and GPT-6 Sol ($2/$10), GPT-6 Luna ($0.10/$0.50); Gemini 3.8 / 3.7 / 3.6 Flash ($0.75/$3.75 launch price through 2026-12-31), 3.5 Flash ($1.50/$9), 3.1 Flash-Lite ($0.25/$1.50), 3.1 Pro Preview.
20
+ - **`slash-tokens/auto` routing is unchanged within 1.6.x.** `/auto` rewrites live requests only to the models it routed to in 1.6.5 (`AUTO_ROUTE_TARGETS`), using the same identification, so a patch upgrade never changes where production calls go. New models are now recognised and priced correctly in `/auto` events (`identifyModel`: `claude-opus-5-5` at $4/$20, `gpt-6-astra` at $10/$50); /auto starts routing to them in 1.7.0.
21
+ - **Routing (`preflightRoute`) follows the new ladder.** Cheapest same-provider is now GPT-6 Luna for OpenAI and Gemini 3.1 Flash-Lite for Google (Anthropic and xAI unchanged). grok-build-0.1 is priced but never a routing target (`NOT_ROUTE_TARGETS`): it's a coding-agent model. When two alternatives' costs round to the same value (tiny prompts), the lower list price wins.
22
+ - **Calibration:** new Claude, Grok and Gemini models use their family's measured factor. GPT-6 is a new generation with an unbenchmarked tokenizer, so it takes the conservative default until a bench run adds it (slash never under-reports).
23
+
24
+ No Team/Solo price change.
25
+
26
+ ## [1.6.5] — The Fixed Deal Edition
27
+
28
+ *2026-08-25*
29
+
30
+ Solo $20 mailbox, 10% waived. Team $39 for the data.
31
+
32
+ Rebuilt tarball. 1.6.4 packed a stale gitignored `dist/` (`bun test` runs `src/`). This tarball contains the live ladder: Grok **4.6** / **4.3**, GPT-5.6 **Sol / Terra / Luna**, Claude **Opus 5 / Sonnet 5 / Haiku 4.5**, Gemini **3.5 Flash-Lite**.
33
+
34
+ No Team/Solo price change.
35
+
3
36
  ## [1.6.4] — The Fixed Deal Edition
4
37
 
5
38
  *2026-08-25*
6
39
 
7
- Live model ladder. Grok **4.6** / **4.3**, GPT-5.6 **Sol / Terra / Luna**, Claude **Opus 5 / Sonnet 5 / Haiku 4.5**, Gemini **3.5 Flash-Lite**. Old keys stay as aliases. Calibration factors carried from the 2026-08-23 corpus — not re-measured on the new IDs.
40
+ Live model ladder in source. **Tarball packed stale `dist/`** — `bun test` never rebuilds it. Use **1.6.5**.
8
41
 
9
42
  No Team/Solo price change.
10
43
 
package/README.md CHANGED
@@ -15,10 +15,14 @@ Know the cost before the call leaves your machine.
15
15
  Models change. Windows grow. Slash adapts — you keep building.
16
16
  Cheaper tokens haven't shrunk the bill — usage has.
17
17
 
18
- ## v1.6.4 — The Fixed Deal Edition
18
+ ## v1.6.6 — The Fixed Deal Edition
19
+
20
+ `--version` / `--help` answer, and today's models price correctly.
19
21
 
20
22
  Solo $20 mailbox, 10% waived. Team $39 for the data.
21
23
 
24
+ New in 1.6.6: Claude Opus 5.5 / Sonnet 5.5 / Fable 5.1, Grok 4.7, GPT-6 (Astra, Sol, Luna), Gemini 3.6–3.8 Flash and 3.1 Flash-Lite, priced as of 2026-10-07. Real API IDs (`claude-opus-4-7`) work in `preflight()`. `preflightRoute()` now finds GPT-6 Luna and Gemini 3.1 Flash-Lite as the cheapest same-provider options.
25
+
22
26
  **Free forever is bunx** — no account. A one-person account is email → key, **$20 on the house**. We show the savings. We don't charge. 10% is the model, waived. Team is **$39 for the data** (`$390`/year).
23
27
 
24
28
  ```bash
@@ -26,7 +30,7 @@ bunx slash-tokens
26
30
  # or: npx --yes slash-tokens
27
31
  ```
28
32
 
29
- Run it in a project that already calls an LLM. An empty folder prints that nothing was found — that's normal.
33
+ Run it in a project that already calls an LLM. An empty folder prints that nothing was found, then tells you to run it in an app — that's normal. `--version` / `--help` print and exit (they do not scan). Pin proof: `slash-tokens --version` or `npm view slash-tokens version`.
30
34
 
31
35
  ```bash
32
36
  npm install slash-tokens
package/dist/cli.js CHANGED
@@ -1,5 +1,10 @@
1
1
  #!/usr/bin/env node
2
2
 
3
+ // src/cli.ts
4
+ import { readFileSync as readFileSync2 } from "node:fs";
5
+ import { dirname, join as join2 } from "node:path";
6
+ import { fileURLToPath } from "node:url";
7
+
3
8
  // src/scanner.ts
4
9
  import { readdirSync, readFileSync, statSync } from "fs";
5
10
  import { join, extname } from "path";
@@ -49,17 +54,162 @@ function writeToMemory(content) {
49
54
  return maxLen;
50
55
  }
51
56
 
57
+ // src/models.ts
58
+ var OPUS = { input: 5, output: 25, context: 1e6 };
59
+ var OPUS_55 = { input: 4, output: 20, context: 1e6 };
60
+ var FABLE = { input: 10, output: 50, context: 1e6 };
61
+ var SONNET_4X = { input: 3, output: 15, context: 1e6 };
62
+ var SONNET = { input: 2, output: 10, context: 1e6 };
63
+ var HAIKU = { input: 1, output: 5, context: 200000 };
64
+ var GROK_46 = {
65
+ input: 2,
66
+ output: 6,
67
+ context: 500000,
68
+ longContextThreshold: 200000,
69
+ longContextInput: 4,
70
+ longContextOutput: 12
71
+ };
72
+ var GROK_43 = {
73
+ input: 1.25,
74
+ output: 2.5,
75
+ context: 1e6,
76
+ longContextThreshold: 200000,
77
+ longContextInput: 2.5,
78
+ longContextOutput: 5
79
+ };
80
+ var GEMINI_PRO = {
81
+ input: 2,
82
+ output: 12,
83
+ context: 1e6,
84
+ longContextThreshold: 200000,
85
+ longContextInput: 4,
86
+ longContextOutput: 18
87
+ };
88
+ var GEMINI_FLASH = { input: 0.3, output: 2.5, context: 1e6 };
89
+ var GEMINI_FLASH_3X = { input: 0.75, output: 3.75, context: 1e6 };
90
+ var GEMINI_35_FLASH = { input: 1.5, output: 9, context: 1e6 };
91
+ var GEMINI_31_FLASH_LITE = { input: 0.25, output: 1.5, context: 1e6 };
92
+ var GROK_BUILD = {
93
+ input: 1,
94
+ output: 2,
95
+ context: 256000,
96
+ longContextThreshold: 200000,
97
+ longContextInput: 2,
98
+ longContextOutput: 4
99
+ };
100
+ var GPT_6_ASTRA = { input: 10, output: 50, context: 1050000 };
101
+ var GPT_6_SOL = { input: 2, output: 10, context: 1050000 };
102
+ var GPT_6_LUNA = { input: 0.1, output: 0.5, context: 1050000 };
103
+ var GPT_SOL = { input: 4, output: 20, context: 1050000 };
104
+ var GPT_TERRA = { input: 2, output: 12, context: 1050000 };
105
+ var GPT_LUNA = { input: 0.2, output: 1.2, context: 1050000 };
106
+ var GPT_54 = { input: 2.5, output: 15, context: 1e6 };
107
+ var GPT_54_MINI = { input: 0.75, output: 4.5, context: 128000 };
108
+ var GPT_54_NANO = { input: 0.2, output: 1.25, context: 128000 };
109
+ var MODELS = {
110
+ "claude-fable-5.1": { ...FABLE },
111
+ "claude-fable-5": { ...FABLE },
112
+ "claude-mythos-5.1": { ...FABLE },
113
+ "claude-mythos-5": { ...FABLE },
114
+ "claude-opus-5.5": { ...OPUS_55 },
115
+ "claude-opus-5": { ...OPUS },
116
+ "claude-opus-4.8": { ...OPUS },
117
+ "claude-opus": { ...OPUS },
118
+ "claude-opus-4.7": { ...OPUS },
119
+ "claude-opus-4.6": { ...OPUS },
120
+ "claude-opus-4.5": { ...OPUS },
121
+ "claude-sonnet-5.5": { ...SONNET },
122
+ "claude-sonnet-5": { ...SONNET },
123
+ "claude-sonnet": { ...SONNET },
124
+ "claude-sonnet-4.6": { ...SONNET_4X },
125
+ "claude-sonnet-4.5": { ...SONNET_4X },
126
+ "claude-haiku-4.5": { ...HAIKU },
127
+ "claude-haiku": { ...HAIKU },
128
+ "grok-4.7": { ...GROK_46 },
129
+ "grok-4.6": { ...GROK_46 },
130
+ "grok-4.5": { ...GROK_46 },
131
+ "grok-build-0.1": { ...GROK_BUILD },
132
+ "grok-4.3": { ...GROK_43 },
133
+ "grok-4.20": { ...GROK_43 },
134
+ "grok-4-1-fast": { ...GROK_43 },
135
+ "gemini-3.1-pro": { ...GEMINI_PRO },
136
+ "gemini-3.1-pro-preview": { ...GEMINI_PRO },
137
+ "gemini-3.8-flash": { ...GEMINI_FLASH_3X },
138
+ "gemini-3.7-flash": { ...GEMINI_FLASH_3X },
139
+ "gemini-3.6-flash": { ...GEMINI_FLASH_3X },
140
+ "gemini-3.5-flash": { ...GEMINI_35_FLASH },
141
+ "gemini-3.1-flash-lite": { ...GEMINI_31_FLASH_LITE },
142
+ "gemini-3.5-flash-lite": { ...GEMINI_FLASH },
143
+ "gemini-2.5-flash": { ...GEMINI_FLASH },
144
+ "gpt-6-astra": { ...GPT_6_ASTRA },
145
+ "gpt-6.1-sol": { ...GPT_6_SOL },
146
+ "gpt-6-sol": { ...GPT_6_SOL },
147
+ "gpt-6-luna": { ...GPT_6_LUNA },
148
+ "gpt-5.6-sol": { ...GPT_SOL },
149
+ "gpt-5.6-terra": { ...GPT_TERRA },
150
+ "gpt-5.6-luna": { ...GPT_LUNA },
151
+ "gpt-5.4": { ...GPT_54 },
152
+ "gpt-5.4-mini": { ...GPT_54_MINI },
153
+ "gpt-5.4-nano": { ...GPT_54_NANO }
154
+ };
155
+ function canonicalModel(name) {
156
+ if (MODELS[name])
157
+ return name;
158
+ let n = name.trim().toLowerCase();
159
+ if (MODELS[n])
160
+ return n;
161
+ n = n.replace(/-\d{8}$/, "");
162
+ if (MODELS[n])
163
+ return n;
164
+ const dotted = n.replace(/(\d)-(\d)(?=$|-)/g, "$1.$2");
165
+ if (MODELS[dotted])
166
+ return dotted;
167
+ return n;
168
+ }
169
+ function getModel(name) {
170
+ return MODELS[canonicalModel(name)];
171
+ }
172
+
52
173
  // src/slash.ts
53
174
  var WASM_INPUT_OFFSET2 = 4096;
54
175
  var CALIBRATION = {
176
+ "claude-fable-5.1": 2.05,
177
+ "claude-fable-5": 2.05,
178
+ "claude-mythos-5.1": 2.05,
179
+ "claude-mythos-5": 2.05,
180
+ "claude-opus-5.5": 2.05,
181
+ "claude-opus-5": 2.05,
182
+ "claude-opus-4.8": 2.05,
183
+ "claude-opus-4.6": 2.05,
184
+ "claude-opus-4.5": 2.05,
185
+ "claude-sonnet-5.5": 2.05,
186
+ "claude-sonnet-4.6": 2.05,
187
+ "claude-sonnet-4.5": 2.05,
55
188
  "claude-opus": 2.05,
56
189
  "claude-opus-4.7": 2.05,
190
+ "claude-sonnet-5": 2.05,
57
191
  "claude-sonnet": 2.05,
192
+ "claude-haiku-4.5": 1.45,
58
193
  "claude-haiku": 1.45,
59
194
  "gemini-3.1-pro": 1.45,
195
+ "gemini-3.1-pro-preview": 1.45,
196
+ "gemini-3.8-flash": 1.45,
197
+ "gemini-3.7-flash": 1.45,
198
+ "gemini-3.6-flash": 1.45,
199
+ "gemini-3.5-flash": 1.45,
200
+ "gemini-3.1-flash-lite": 1.45,
201
+ "gemini-3.5-flash-lite": 1.45,
60
202
  "gemini-2.5-flash": 1.45,
203
+ "grok-4.7": 1.15,
204
+ "grok-4.6": 1.15,
205
+ "grok-4.5": 1.15,
206
+ "grok-build-0.1": 1.15,
207
+ "grok-4.3": 1.15,
61
208
  "grok-4.20": 1.15,
62
209
  "grok-4-1-fast": 1.15,
210
+ "gpt-5.6-sol": 1.15,
211
+ "gpt-5.6-terra": 1.15,
212
+ "gpt-5.6-luna": 1.15,
63
213
  "gpt-5.4": 1.15,
64
214
  "gpt-5.4-mini": 1.15,
65
215
  "gpt-5.4-nano": 1.15
@@ -73,7 +223,7 @@ function slash(content, model) {
73
223
  const raw = instance.exports.estimate_tokens(WASM_INPUT_OFFSET2, len);
74
224
  if (!model)
75
225
  return raw;
76
- const factor = CALIBRATION[model] ?? DEFAULT_UNKNOWN_MODEL_FACTOR;
226
+ const factor = CALIBRATION[canonicalModel(model)] ?? DEFAULT_UNKNOWN_MODEL_FACTOR;
77
227
  return factor === 1 ? raw : Math.ceil(raw * factor);
78
228
  }
79
229
 
@@ -91,12 +241,12 @@ var AI_PATTERNS = [
91
241
  { name: "Mistral", regex: /from\s+['"]@mistralai|MistralClient/g }
92
242
  ];
93
243
  var SDK_REPRESENTATIVE_MODEL = {
94
- Anthropic: "claude-sonnet",
95
- OpenAI: "gpt-5.4",
244
+ Anthropic: "claude-sonnet-5",
245
+ OpenAI: "gpt-5.6-sol",
96
246
  Gemini: "gemini-3.1-pro",
97
- Grok: "grok-4.20"
247
+ Grok: "grok-4.6"
98
248
  };
99
- var UNKNOWN_SDK_REPRESENTATIVE_MODEL = "claude-sonnet";
249
+ var UNKNOWN_SDK_REPRESENTATIVE_MODEL = "claude-sonnet-5";
100
250
  var SKIP_DIRS = new Set([
101
251
  "node_modules",
102
252
  ".git",
@@ -202,24 +352,6 @@ function scan(dir) {
202
352
  };
203
353
  }
204
354
 
205
- // src/models.ts
206
- var MODELS = {
207
- "claude-opus": { input: 5, output: 25, context: 1e6 },
208
- "claude-opus-4.7": { input: 5, output: 25, context: 1e6 },
209
- "claude-sonnet": { input: 2, output: 10, context: 1e6 },
210
- "claude-haiku": { input: 1, output: 5, context: 200000 },
211
- "grok-4.20": { input: 1.25, output: 2.5, context: 1e6, longContextThreshold: 200000, longContextInput: 2.5, longContextOutput: 5 },
212
- "grok-4-1-fast": { input: 1.25, output: 2.5, context: 1e6, longContextThreshold: 200000, longContextInput: 2.5, longContextOutput: 5 },
213
- "gemini-3.1-pro": { input: 2, output: 12, context: 1e6 },
214
- "gemini-2.5-flash": { input: 0.3, output: 2.5, context: 1e6 },
215
- "gpt-5.4": { input: 2.5, output: 15, context: 1e6 },
216
- "gpt-5.4-mini": { input: 0.75, output: 4.5, context: 128000 },
217
- "gpt-5.4-nano": { input: 0.2, output: 1.25, context: 128000 }
218
- };
219
- function getModel(name) {
220
- return MODELS[name] || MODELS[name.toLowerCase()];
221
- }
222
-
223
355
  // src/report.ts
224
356
  var R = "\x1B[0m";
225
357
  var B = "\x1B[1m";
@@ -244,6 +376,7 @@ function printReport(sites, filesScanned, timeMs, cwd) {
244
376
  if (sites.length === 0) {
245
377
  console.log(`${WHITE} No AI API call sites detected.${R}`);
246
378
  console.log(`${GRAY} Supported: OpenAI, Anthropic, Vercel AI, LangChain, Gemini, Bedrock, Grok${R}`);
379
+ console.log(`${GRAY} Run this in a project that already calls an LLM.${R}`);
247
380
  console.log("");
248
381
  return;
249
382
  }
@@ -315,6 +448,14 @@ function init(opts) {
315
448
 
316
449
  // src/cli.ts
317
450
  var args = process.argv.slice(2);
451
+ if (args.includes("--version") || args.includes("-V")) {
452
+ console.log(packageVersion());
453
+ process.exit(0);
454
+ }
455
+ if (args.includes("--help") || args.includes("-h")) {
456
+ printHelp();
457
+ process.exit(0);
458
+ }
318
459
  var keyArg = args.find((a) => a.startsWith("--key="));
319
460
  var key = keyArg?.split("=")[1] || process.env.SLASH_KEY;
320
461
  if (key)
@@ -336,3 +477,21 @@ if (sites.length > 0) {
336
477
  }
337
478
  console.log("");
338
479
  }
480
+ function packageVersion() {
481
+ const here = dirname(fileURLToPath(import.meta.url));
482
+ const pkg = JSON.parse(readFileSync2(join2(here, "..", "package.json"), "utf8"));
483
+ return pkg.version;
484
+ }
485
+ function printHelp() {
486
+ console.log(`slash-tokens — Token Optimization for Context Engineers
487
+
488
+ bunx slash-tokens try (scan this directory, no account)
489
+ npm install slash-tokens SDK
490
+
491
+ --version, -V print version and exit
492
+ --help, -h print this help and exit
493
+ --key=KEY optional; CLI scans never charge
494
+
495
+ Run in a project that already calls an LLM.
496
+ Empty folder → no call sites. That's normal.`);
497
+ }
@@ -11,6 +11,14 @@ export interface InterceptEvent {
11
11
  routed: boolean;
12
12
  timestamp: string;
13
13
  }
14
+ /**
15
+ * The model a request actually names, for PRICING: the exact table entry when
16
+ * the strict canonical ID matches (claude-opus-5-5 → claude-opus-5.5,
17
+ * gpt-6-luna), else the legacy family mapping. Routing keeps using
18
+ * normalizeModel() so /auto decisions are unchanged within 1.6.x.
19
+ */
20
+ export declare function identifyModel(raw: string): string;
14
21
  export declare function normalizeModel(raw: string): string;
22
+ export declare function findCheapestRoute(provider: string, tokens: number, currentModel: string): string | null;
15
23
  export declare function onIntercept(handler: (event: InterceptEvent) => void): void;
16
24
  export declare function patchFetch(): void;
package/dist/intercept.js CHANGED
@@ -1,20 +1,29 @@
1
1
  import { slash } from './slash.js';
2
- import { getModel } from './models.js';
2
+ import { getModel, MODELS, canonicalModel } from './models.js';
3
3
  import { shouldRoute, isModelAllowed } from './config.js';
4
- import { PROVIDER_MODELS } from './providers.js';
4
+ import { PROVIDER_MODELS, AUTO_ROUTE_TARGETS, NOT_ROUTE_TARGETS } from './providers.js';
5
5
  // Reverse lookup: model name → provider model names in the API
6
6
  // (what to put back in the request body)
7
7
  const MODEL_API_NAMES = {
8
+ 'claude-opus-5': 'claude-opus-5',
8
9
  'claude-opus': 'claude-opus-5',
9
10
  'claude-opus-4.7': 'claude-opus-4-7',
11
+ 'claude-sonnet-5': 'claude-sonnet-5',
10
12
  'claude-sonnet': 'claude-sonnet-5',
13
+ 'claude-haiku-4.5': 'claude-haiku-4-5-20251001',
11
14
  'claude-haiku': 'claude-haiku-4-5-20251001',
15
+ 'gpt-5.6-sol': 'gpt-5.6-sol',
16
+ 'gpt-5.6-terra': 'gpt-5.6-terra',
17
+ 'gpt-5.6-luna': 'gpt-5.6-luna',
12
18
  'gpt-5.4': 'gpt-5.4',
13
19
  'gpt-5.4-mini': 'gpt-5.4-mini',
14
20
  'gpt-5.4-nano': 'gpt-5.4-nano',
21
+ 'grok-4.6': 'grok-4.6',
22
+ 'grok-4.3': 'grok-4.3',
15
23
  'grok-4.20': 'grok-4.20-0309-non-reasoning',
16
24
  'grok-4-1-fast': 'grok-4.3',
17
25
  'gemini-3.1-pro': 'gemini-pro-latest',
26
+ 'gemini-3.5-flash-lite': 'gemini-3.5-flash-lite',
18
27
  'gemini-2.5-flash': 'gemini-flash-latest',
19
28
  };
20
29
  // AI API endpoint detection
@@ -27,7 +36,7 @@ const AI_ENDPOINTS = [
27
36
  {
28
37
  pattern: /api\.openai\.com/,
29
38
  provider: 'OpenAI',
30
- modelExtractor: (body) => body?.model || 'gpt-5.4',
39
+ modelExtractor: (body) => body?.model || 'gpt-5.6-sol',
31
40
  },
32
41
  {
33
42
  pattern: /generativelanguage\.googleapis\.com/,
@@ -35,13 +44,13 @@ const AI_ENDPOINTS = [
35
44
  modelExtractor: (_body, url) => {
36
45
  // Model is in the URL path: /v1beta/models/gemini-2.0-flash:generateContent
37
46
  const match = url?.match(/\/models\/([^/:]+)/);
38
- return match ? match[1] : 'gemini-2.5-flash';
47
+ return match ? match[1] : 'gemini-3.5-flash-lite';
39
48
  },
40
49
  },
41
50
  {
42
51
  pattern: /api\.x\.ai/,
43
52
  provider: 'xAI',
44
- modelExtractor: (body) => body?.model || 'grok-4.20',
53
+ modelExtractor: (body) => body?.model || 'grok-4.6',
45
54
  },
46
55
  ];
47
56
  // Normalize model names to our pricing table keys.
@@ -53,35 +62,66 @@ const AI_ENDPOINTS = [
53
62
  // (getModel() returning undefined → $0 reported cost, see slash.ts's
54
63
  // DEFAULT_UNKNOWN_MODEL_FACTOR comment for why "unrecognized" defaulting
55
64
  // to a falsely-safe-looking value is the dangerous case).
65
+ /**
66
+ * The model a request actually names, for PRICING: the exact table entry when
67
+ * the strict canonical ID matches (claude-opus-5-5 → claude-opus-5.5,
68
+ * gpt-6-luna), else the legacy family mapping. Routing keeps using
69
+ * normalizeModel() so /auto decisions are unchanged within 1.6.x.
70
+ */
71
+ export function identifyModel(raw) {
72
+ const c = canonicalModel(raw);
73
+ return MODELS[c] ? c : normalizeModel(raw);
74
+ }
56
75
  export function normalizeModel(raw) {
57
76
  const lower = raw.toLowerCase();
58
- // Anthropic — 4.7 before generic opus check, same order as slash-models.ts
77
+ // Anthropic — specific versions before generic family
59
78
  if (lower.includes('opus') && (lower.includes('4-7') || lower.includes('4.7')))
60
79
  return 'claude-opus-4.7';
80
+ if (lower.includes('opus-5'))
81
+ return 'claude-opus-5';
61
82
  if (lower.includes('opus'))
62
83
  return 'claude-opus';
84
+ if (lower.includes('sonnet-5'))
85
+ return 'claude-sonnet-5';
63
86
  if (lower.includes('sonnet'))
64
87
  return 'claude-sonnet';
88
+ if (lower.includes('haiku') && (lower.includes('4.5') || lower.includes('4-5')))
89
+ return 'claude-haiku-4.5';
65
90
  if (lower.includes('haiku'))
66
91
  return 'claude-haiku';
67
- // xAI
92
+ // xAI — 4.6 / 4.3 / 4.20 before generic grok
93
+ if (lower.includes('grok') && (lower.includes('4.6') || lower.includes('4-6')))
94
+ return 'grok-4.6';
95
+ if (lower.includes('grok') && (lower.includes('4.3') || lower.includes('4-3')))
96
+ return 'grok-4.3';
97
+ if (lower.includes('grok') && (lower.includes('4.20') || lower.includes('4-20')))
98
+ return 'grok-4.20';
68
99
  if (lower.includes('grok') && lower.includes('fast'))
69
100
  return 'grok-4-1-fast';
70
101
  if (lower.includes('grok'))
71
- return 'grok-4.20';
102
+ return 'grok-4.6';
72
103
  // Google
73
104
  if (lower.includes('gemini') && lower.includes('pro'))
74
105
  return 'gemini-3.1-pro';
75
- if (lower.includes('gemini'))
106
+ if (lower.includes('gemini') && lower.includes('3.5') && lower.includes('lite'))
107
+ return 'gemini-3.5-flash-lite';
108
+ if (lower.includes('gemini') && lower.includes('2.5'))
76
109
  return 'gemini-2.5-flash';
77
- // OpenAI — current (5.4 family)
110
+ if (lower.includes('gemini'))
111
+ return 'gemini-3.5-flash-lite';
112
+ // OpenAI — 5.6 then 5.4 then legacy
113
+ if (lower.includes('5.6') && lower.includes('luna'))
114
+ return 'gpt-5.6-luna';
115
+ if (lower.includes('5.6') && lower.includes('terra'))
116
+ return 'gpt-5.6-terra';
117
+ if (lower.includes('5.6'))
118
+ return 'gpt-5.6-sol';
78
119
  if (lower.includes('5.4') && lower.includes('nano'))
79
120
  return 'gpt-5.4-nano';
80
121
  if (lower.includes('5.4') && lower.includes('mini'))
81
122
  return 'gpt-5.4-mini';
82
123
  if (lower.includes('5.4'))
83
124
  return 'gpt-5.4';
84
- // OpenAI — legacy model names → map to closest current equivalent
85
125
  if (lower.includes('o1-mini') || lower.includes('o1_mini'))
86
126
  return 'gpt-5.4-mini';
87
127
  if (lower.includes('o1'))
@@ -118,7 +158,7 @@ function extractContent(body) {
118
158
  return JSON.stringify(body);
119
159
  }
120
160
  // Find cheapest model from same provider that fits
121
- function findCheapestRoute(provider, tokens, currentModel) {
161
+ export function findCheapestRoute(provider, tokens, currentModel) {
122
162
  const providerModels = PROVIDER_MODELS[provider];
123
163
  if (!providerModels)
124
164
  return null;
@@ -126,6 +166,10 @@ function findCheapestRoute(provider, tokens, currentModel) {
126
166
  for (const model of providerModels) {
127
167
  if (model === currentModel)
128
168
  continue;
169
+ if (!AUTO_ROUTE_TARGETS.has(model))
170
+ continue; // 1.6.x: /auto rewrites only to its 1.6.5 targets
171
+ if (NOT_ROUTE_TARGETS.has(model))
172
+ continue; // specialised models are never targets
129
173
  if (!isModelAllowed(model))
130
174
  continue; // user excluded this model
131
175
  const info = getModel(model);
@@ -158,13 +202,14 @@ export function patchFetch() {
158
202
  const body = JSON.parse(bodyStr);
159
203
  const content = extractContent(body);
160
204
  const rawModel = match.modelExtractor(body, url);
161
- const originalModel = normalizeModel(rawModel);
205
+ const routingModel = normalizeModel(rawModel); // decides routes (unchanged in 1.6.x)
206
+ const originalModel = identifyModel(rawModel); // prices the request it really names
162
207
  const tokens = slash(content, originalModel);
163
208
  const originalInfo = getModel(originalModel);
164
209
  const originalCost = originalInfo ? Math.round(((tokens / 1000000) * originalInfo.input) * 1000000) / 1000000 : 0;
165
210
  const fits = originalInfo ? tokens <= originalInfo.context : true;
166
211
  // Find cheapest route within same provider (if routing enabled)
167
- const routeModel = shouldRoute() ? findCheapestRoute(match.provider, tokens, originalModel) : null;
212
+ const routeModel = shouldRoute() ? findCheapestRoute(match.provider, tokens, routingModel) : null;
168
213
  const routedInfo = routeModel ? getModel(routeModel) : null;
169
214
  const routedCost = routedInfo ? Math.round(((tokens / 1000000) * routedInfo.input) * 1000000) / 1000000 : originalCost;
170
215
  const salvaged = routeModel ? Math.round((originalCost - routedCost) * 1000000) / 1000000 : 0;
package/dist/models.d.ts CHANGED
@@ -7,6 +7,13 @@ export interface ModelInfo {
7
7
  longContextOutput?: number;
8
8
  }
9
9
  export declare const MODELS: Record<string, ModelInfo>;
10
+ /**
11
+ * Real API IDs → table keys, strictly: lowercase, drop a trailing date stamp
12
+ * (`-20250514`), and write version numbers with dots (`claude-opus-4-7` →
13
+ * `claude-opus-4.7`). No family guessing: an unknown version stays unknown,
14
+ * so a new model never silently gets an older model's price.
15
+ */
16
+ export declare function canonicalModel(name: string): string;
10
17
  export declare function getModel(name: string): ModelInfo | undefined;
11
18
  export declare function effectiveRate(tokens: number, info: ModelInfo): {
12
19
  input: number;
package/dist/models.js CHANGED
@@ -1,44 +1,119 @@
1
- // Pricing as of April 2026 — USD per million tokens
2
- // Claude Sonnet 5 pricing corrected 2026-08-23: was hardcoded at $3.00/
3
- // $15.00 (the previously-scheduled Sept 1, 2026 increase), but Anthropic's
4
- // own pricing page confirms that increase will NOT occur — $2.00/$10.00
5
- // (the introductory rate) is now the permanent standard price. Found via
6
- // a full-codebase review that independently re-verified every provider's
7
- // live pricing, not just the model IDs already fixed that day.
8
- // xAI pricing re-derived 2026-08-23: grok-4.20 and grok-4-1-fast (the literal
9
- // API IDs, not just the generic keys below) were both fully retired — not
10
- // just old snapshots, they 404 on the live API. Current lineup has no cheap
11
- // tier at all; grok-4.20 (generic) now targets grok-4.20-0309-non-reasoning
12
- // and grok-4-1-fast (generic) targets grok-4.3, both $1.25/$2.50/1M — see
13
- // intercept.ts MODEL_API_NAMES. There is currently no xAI model cheaper than
14
- // $1.25/M input, so routing between these two generic keys yields zero
15
- // savings (findCheapestRoute requires strictly cheaper — this is honest,
16
- // not a bug: the old $0.20/M "fast" tier no longer exists).
1
+ const OPUS = { input: 5.00, output: 25.00, context: 1000000 };
2
+ const OPUS_55 = { input: 4.00, output: 20.00, context: 1000000 };
3
+ const FABLE = { input: 10.00, output: 50.00, context: 1000000 };
4
+ const SONNET_4X = { input: 3.00, output: 15.00, context: 1000000 };
5
+ const SONNET = { input: 2.00, output: 10.00, context: 1000000 };
6
+ const HAIKU = { input: 1.00, output: 5.00, context: 200000 };
7
+ const GROK_46 = {
8
+ input: 2.00, output: 6.00, context: 500000,
9
+ longContextThreshold: 200000, longContextInput: 4.00, longContextOutput: 12.00,
10
+ };
11
+ const GROK_43 = {
12
+ input: 1.25, output: 2.50, context: 1000000,
13
+ longContextThreshold: 200000, longContextInput: 2.50, longContextOutput: 5.00,
14
+ };
15
+ const GEMINI_PRO = {
16
+ input: 2.00, output: 12.00, context: 1000000,
17
+ longContextThreshold: 200000, longContextInput: 4.00, longContextOutput: 18.00,
18
+ };
19
+ const GEMINI_FLASH = { input: 0.30, output: 2.50, context: 1000000 };
20
+ // Gemini 3.6–3.8 Flash: launch price through 2026-12-31; Google lists $1.50/$7.50
21
+ // from 2027-01-01. Update this entry before then.
22
+ const GEMINI_FLASH_3X = { input: 0.75, output: 3.75, context: 1000000 };
23
+ const GEMINI_35_FLASH = { input: 1.50, output: 9.00, context: 1000000 };
24
+ const GEMINI_31_FLASH_LITE = { input: 0.25, output: 1.50, context: 1000000 };
25
+ const GROK_BUILD = {
26
+ input: 1.00, output: 2.00, context: 256000,
27
+ longContextThreshold: 200000, longContextInput: 2.00, longContextOutput: 4.00,
28
+ };
29
+ const GPT_6_ASTRA = { input: 10.00, output: 50.00, context: 1050000 };
30
+ const GPT_6_SOL = { input: 2.00, output: 10.00, context: 1050000 };
31
+ const GPT_6_LUNA = { input: 0.10, output: 0.50, context: 1050000 };
32
+ const GPT_SOL = { input: 4.00, output: 20.00, context: 1050000 };
33
+ const GPT_TERRA = { input: 2.00, output: 12.00, context: 1050000 };
34
+ const GPT_LUNA = { input: 0.20, output: 1.20, context: 1050000 };
35
+ const GPT_54 = { input: 2.50, output: 15.00, context: 1000000 };
36
+ const GPT_54_MINI = { input: 0.75, output: 4.50, context: 128000 };
37
+ const GPT_54_NANO = { input: 0.20, output: 1.25, context: 128000 };
38
+ // Pricing as of 2026-10-07 — USD per million tokens (checked against the four pages below).
39
+ // First-party: platform.claude.com/docs/en/about-claude/pricing
40
+ // developers.openai.com/api/docs/models
41
+ // docs.x.ai/developers/models
42
+ // ai.google.dev/gemini-api/docs/pricing
43
+ // Old keys stay as aliases so existing call sites don't throw.
17
44
  export const MODELS = {
18
- // Anthropic
19
- 'claude-opus': { input: 5.00, output: 25.00, context: 1000000 },
20
- 'claude-opus-4.7': { input: 5.00, output: 25.00, context: 1000000 },
21
- 'claude-sonnet': { input: 2.00, output: 10.00, context: 1000000 },
22
- 'claude-haiku': { input: 1.00, output: 5.00, context: 200000 },
23
- // xAI
24
- 'grok-4.20': { input: 1.25, output: 2.50, context: 1000000, longContextThreshold: 200000, longContextInput: 2.50, longContextOutput: 5.00 },
25
- 'grok-4-1-fast': { input: 1.25, output: 2.50, context: 1000000, longContextThreshold: 200000, longContextInput: 2.50, longContextOutput: 5.00 },
45
+ // Anthropic — live names + generic aliases (same rates)
46
+ 'claude-fable-5.1': { ...FABLE },
47
+ 'claude-fable-5': { ...FABLE },
48
+ 'claude-mythos-5.1': { ...FABLE },
49
+ 'claude-mythos-5': { ...FABLE },
50
+ 'claude-opus-5.5': { ...OPUS_55 },
51
+ 'claude-opus-5': { ...OPUS },
52
+ 'claude-opus-4.8': { ...OPUS },
53
+ 'claude-opus': { ...OPUS },
54
+ 'claude-opus-4.7': { ...OPUS },
55
+ 'claude-opus-4.6': { ...OPUS },
56
+ 'claude-opus-4.5': { ...OPUS },
57
+ 'claude-sonnet-5.5': { ...SONNET },
58
+ 'claude-sonnet-5': { ...SONNET },
59
+ 'claude-sonnet': { ...SONNET },
60
+ 'claude-sonnet-4.6': { ...SONNET_4X },
61
+ 'claude-sonnet-4.5': { ...SONNET_4X },
62
+ 'claude-haiku-4.5': { ...HAIKU },
63
+ 'claude-haiku': { ...HAIKU },
64
+ // xAI — flagship 4.6, cheap same-provider 4.3. 4.20 / fast are aliases.
65
+ 'grok-4.7': { ...GROK_46 },
66
+ 'grok-4.6': { ...GROK_46 },
67
+ 'grok-4.5': { ...GROK_46 },
68
+ 'grok-build-0.1': { ...GROK_BUILD },
69
+ 'grok-4.3': { ...GROK_43 },
70
+ 'grok-4.20': { ...GROK_43 },
71
+ 'grok-4-1-fast': { ...GROK_43 },
26
72
  // Google
27
- 'gemini-3.1-pro': { input: 2.00, output: 12.00, context: 1000000 },
28
- 'gemini-2.5-flash': { input: 0.30, output: 2.50, context: 1000000 },
29
- // OpenAI
30
- 'gpt-5.4': { input: 2.50, output: 15.00, context: 1000000 },
31
- 'gpt-5.4-mini': { input: 0.75, output: 4.50, context: 128000 },
32
- 'gpt-5.4-nano': { input: 0.20, output: 1.25, context: 128000 },
73
+ 'gemini-3.1-pro': { ...GEMINI_PRO },
74
+ 'gemini-3.1-pro-preview': { ...GEMINI_PRO },
75
+ 'gemini-3.8-flash': { ...GEMINI_FLASH_3X },
76
+ 'gemini-3.7-flash': { ...GEMINI_FLASH_3X },
77
+ 'gemini-3.6-flash': { ...GEMINI_FLASH_3X },
78
+ 'gemini-3.5-flash': { ...GEMINI_35_FLASH },
79
+ 'gemini-3.1-flash-lite': { ...GEMINI_31_FLASH_LITE },
80
+ 'gemini-3.5-flash-lite': { ...GEMINI_FLASH },
81
+ 'gemini-2.5-flash': { ...GEMINI_FLASH },
82
+ // OpenAI — live 5.6 ladder. 5.4 family kept as aliases (old prices).
83
+ 'gpt-6-astra': { ...GPT_6_ASTRA },
84
+ 'gpt-6.1-sol': { ...GPT_6_SOL },
85
+ 'gpt-6-sol': { ...GPT_6_SOL },
86
+ 'gpt-6-luna': { ...GPT_6_LUNA },
87
+ 'gpt-5.6-sol': { ...GPT_SOL },
88
+ 'gpt-5.6-terra': { ...GPT_TERRA },
89
+ 'gpt-5.6-luna': { ...GPT_LUNA },
90
+ 'gpt-5.4': { ...GPT_54 },
91
+ 'gpt-5.4-mini': { ...GPT_54_MINI },
92
+ 'gpt-5.4-nano': { ...GPT_54_NANO },
33
93
  };
94
+ /**
95
+ * Real API IDs → table keys, strictly: lowercase, drop a trailing date stamp
96
+ * (`-20250514`), and write version numbers with dots (`claude-opus-4-7` →
97
+ * `claude-opus-4.7`). No family guessing: an unknown version stays unknown,
98
+ * so a new model never silently gets an older model's price.
99
+ */
100
+ export function canonicalModel(name) {
101
+ if (MODELS[name])
102
+ return name;
103
+ let n = name.trim().toLowerCase();
104
+ if (MODELS[n])
105
+ return n;
106
+ n = n.replace(/-\d{8}$/, '');
107
+ if (MODELS[n])
108
+ return n;
109
+ const dotted = n.replace(/(\d)-(\d)(?=$|-)/g, '$1.$2');
110
+ if (MODELS[dotted])
111
+ return dotted;
112
+ return n;
113
+ }
34
114
  export function getModel(name) {
35
- return MODELS[name] || MODELS[name.toLowerCase()];
115
+ return MODELS[canonicalModel(name)];
36
116
  }
37
- // Resolve the actual billable rate for a given token count — applies the
38
- // long-context tier above if the model has one and tokens cross it.
39
- // preflight()/preflightRoute() should go through this, not read
40
- // .input/.output directly, or a Grok call over 200K tokens gets silently
41
- // under-costed at the base rate.
42
117
  export function effectiveRate(tokens, info) {
43
118
  if (info.longContextThreshold !== undefined && tokens > info.longContextThreshold) {
44
119
  return {
@@ -4,6 +4,6 @@ export interface Pattern {
4
4
  }
5
5
  export declare const AI_PATTERNS: Pattern[];
6
6
  export declare const SDK_REPRESENTATIVE_MODEL: Record<string, string>;
7
- export declare const UNKNOWN_SDK_REPRESENTATIVE_MODEL = "claude-sonnet";
7
+ export declare const UNKNOWN_SDK_REPRESENTATIVE_MODEL = "claude-sonnet-5";
8
8
  export declare const SKIP_DIRS: Set<string>;
9
9
  export declare const SCAN_EXTENSIONS: Set<string>;
package/dist/patterns.js CHANGED
@@ -34,17 +34,17 @@ export const AI_PATTERNS = [
34
34
  // know the exact model), but a real per-provider one instead of a single
35
35
  // guess applied to everyone.
36
36
  export const SDK_REPRESENTATIVE_MODEL = {
37
- 'Anthropic': 'claude-sonnet',
38
- 'OpenAI': 'gpt-5.4',
37
+ 'Anthropic': 'claude-sonnet-5',
38
+ 'OpenAI': 'gpt-5.6-sol',
39
39
  'Gemini': 'gemini-3.1-pro',
40
- 'Grok': 'grok-4.20',
40
+ 'Grok': 'grok-4.6',
41
41
  };
42
42
  // Fallback for SDKs that don't map to one specific provider (Vercel AI,
43
43
  // LangChain, and Bedrock can all wrap any underlying provider; raw
44
44
  // fetch-to-AI-endpoint and Cohere/Mistral have no pricing data in MODELS
45
45
  // at all). claude-sonnet is used as a documented, honest middle-of-the-
46
46
  // road placeholder — not a claim about which model is actually running.
47
- export const UNKNOWN_SDK_REPRESENTATIVE_MODEL = 'claude-sonnet';
47
+ export const UNKNOWN_SDK_REPRESENTATIVE_MODEL = 'claude-sonnet-5';
48
48
  export const SKIP_DIRS = new Set([
49
49
  'node_modules', '.git', 'dist', 'build', '.next', '.nuxt', '.svelte-kit',
50
50
  'coverage', '.turbo', '.cache', '__pycache__', '.venv', 'venv',
package/dist/preflight.js CHANGED
@@ -1,6 +1,6 @@
1
1
  import { slash } from './slash.js';
2
- import { getModel, MODELS, effectiveRate } from './models.js';
3
- import { PROVIDER_MODELS, providerOf } from './providers.js';
2
+ import { getModel, MODELS, effectiveRate, canonicalModel } from './models.js';
3
+ import { PROVIDER_MODELS, providerOf, NOT_ROUTE_TARGETS } from './providers.js';
4
4
  import { shouldRoute, isModelAllowed } from './config.js';
5
5
  /**
6
6
  * Compute cost of a prompt of `tokens` tokens on the given ModelInfo.
@@ -99,9 +99,12 @@ export function preflightRoute(content, model) {
99
99
  return null;
100
100
  const originalCost = computeCost(tokens, info);
101
101
  let cheapest = null;
102
+ const self = canonicalModel(model);
102
103
  for (const m of providerModels) {
103
- if (m === model)
104
+ if (m === self)
104
105
  continue;
106
+ if (NOT_ROUTE_TARGETS.has(m))
107
+ continue; // specialised: never a target
105
108
  if (!isModelAllowed(m))
106
109
  continue; // user excluded this model
107
110
  const altInfo = getModel(m);
@@ -112,7 +115,10 @@ export function preflightRoute(content, model) {
112
115
  if (altInfo.input >= info.input)
113
116
  continue; // not cheaper
114
117
  const alt = buildAlternative(m, originalCost, tokens, altInfo);
115
- if (!cheapest || alt.cost < cheapest.cost) {
118
+ // Tiny prompts round costs to the same value; break ties on the list price
119
+ // so the genuinely cheaper model wins (gpt-5.6-sol → gpt-6-luna, not 5.4-nano).
120
+ if (!cheapest || alt.cost < cheapest.cost ||
121
+ (alt.cost === cheapest.cost && altInfo.input < (getModel(cheapest.model)?.input ?? Infinity))) {
116
122
  cheapest = alt;
117
123
  }
118
124
  }
@@ -1,26 +1,24 @@
1
1
  /**
2
2
  * Provider groups — single source of truth.
3
3
  *
4
- * Slash routing is always SAME-PROVIDER. Opus → Haiku, GPT-5.4 → Nano,
5
- * Grok-4.20 → Fast, Gemini 3.1 Pro → 2.5 Flash. Never cross-provider.
4
+ * Slash routing is always SAME-PROVIDER. Opus 5 → Haiku 4.5, Sol → Luna,
5
+ * Grok 4.6 → 4.3, Gemini 3.1 Pro → 3.5 Flash-Lite. Never cross-provider.
6
6
  *
7
- * This file is the canonical provider mapping. Both `intercept.ts`
8
- * (runtime fetch patching) and `preflight.ts` (analysis + routing
9
- * prediction) consume from here so they can never drift.
10
- *
11
- * TEST-NOTE: Whenever a new model is added, it MUST appear in exactly one
12
- * provider group below. A model missing from here will:
13
- * - Never be a routing target from `findCheapestRoute` / `preflightRoute`
14
- * - Still appear in `preflight().options` (which is cross-provider analysis)
15
- * That mismatch is by design — see preflight.ts semantics.
7
+ * Order inside a group matters when two models share a price: the first
8
+ * strictly-cheaper hit wins (findCheapestRoute / preflightRoute).
16
9
  */
17
10
  export declare const PROVIDER_MODELS: Record<string, string[]>;
18
11
  /**
19
- * Identify a model's provider from its canonical key.
20
- *
21
- * TEST-NOTE: This function MUST return a non-null provider for every model
22
- * present in `MODELS` (from models.ts). If `MODELS` adds a model without
23
- * adding it to `PROVIDER_MODELS`, this returns null and routing is disabled
24
- * for that model. Silent skip.
12
+ * Priced and grouped, but never a routing TARGET: specialised models a general
13
+ * call shouldn't be moved to (grok-build-0.1 is a coding-agent model).
14
+ */
15
+ export declare const NOT_ROUTE_TARGETS: ReadonlySet<string>;
16
+ /**
17
+ * Models `slash-tokens/auto` may rewrite a live request TO. Frozen at the
18
+ * 1.6.5 set so a patch release never changes where production calls go:
19
+ * models added in 1.6.6 (GPT-6, Gemini 3.1 Flash-Lite / 3.6–3.8 Flash, Grok
20
+ * 4.7 / 4.5, Claude 5.5 / Fable …) are priced and shown by preflight() /
21
+ * preflightRoute(), but /auto only starts routing to them in 1.7.0.
25
22
  */
23
+ export declare const AUTO_ROUTE_TARGETS: ReadonlySet<string>;
26
24
  export declare function providerOf(model: string): string | null;
package/dist/providers.js CHANGED
@@ -1,36 +1,58 @@
1
+ import { canonicalModel } from './models.js';
1
2
  /**
2
3
  * Provider groups — single source of truth.
3
4
  *
4
- * Slash routing is always SAME-PROVIDER. Opus → Haiku, GPT-5.4 → Nano,
5
- * Grok-4.20 → Fast, Gemini 3.1 Pro → 2.5 Flash. Never cross-provider.
5
+ * Slash routing is always SAME-PROVIDER. Opus 5 → Haiku 4.5, Sol → Luna,
6
+ * Grok 4.6 → 4.3, Gemini 3.1 Pro → 3.5 Flash-Lite. Never cross-provider.
6
7
  *
7
- * This file is the canonical provider mapping. Both `intercept.ts`
8
- * (runtime fetch patching) and `preflight.ts` (analysis + routing
9
- * prediction) consume from here so they can never drift.
10
- *
11
- * TEST-NOTE: Whenever a new model is added, it MUST appear in exactly one
12
- * provider group below. A model missing from here will:
13
- * - Never be a routing target from `findCheapestRoute` / `preflightRoute`
14
- * - Still appear in `preflight().options` (which is cross-provider analysis)
15
- * That mismatch is by design — see preflight.ts semantics.
8
+ * Order inside a group matters when two models share a price: the first
9
+ * strictly-cheaper hit wins (findCheapestRoute / preflightRoute).
16
10
  */
17
11
  export const PROVIDER_MODELS = {
18
- Anthropic: ['claude-opus', 'claude-opus-4.7', 'claude-sonnet', 'claude-haiku'],
19
- OpenAI: ['gpt-5.4', 'gpt-5.4-mini', 'gpt-5.4-nano'],
20
- xAI: ['grok-4.20', 'grok-4-1-fast'],
21
- Google: ['gemini-3.1-pro', 'gemini-2.5-flash'],
12
+ Anthropic: [
13
+ 'claude-fable-5.1', 'claude-fable-5', 'claude-mythos-5.1', 'claude-mythos-5',
14
+ 'claude-opus-5', 'claude-opus', 'claude-opus-4.8', 'claude-opus-4.7',
15
+ 'claude-opus-4.6', 'claude-opus-4.5', 'claude-opus-5.5',
16
+ 'claude-sonnet-4.6', 'claude-sonnet-4.5',
17
+ 'claude-sonnet-5.5', 'claude-sonnet-5', 'claude-sonnet',
18
+ 'claude-haiku', 'claude-haiku-4.5',
19
+ ],
20
+ OpenAI: [
21
+ 'gpt-6-astra',
22
+ 'gpt-5.6-sol', 'gpt-5.6-terra',
23
+ 'gpt-5.4', 'gpt-6.1-sol', 'gpt-6-sol', 'gpt-5.4-mini', 'gpt-5.4-nano',
24
+ 'gpt-5.6-luna', 'gpt-6-luna',
25
+ ],
26
+ xAI: ['grok-4.7', 'grok-4.6', 'grok-4.5', 'grok-4.3', 'grok-4.20', 'grok-4-1-fast', 'grok-build-0.1'],
27
+ Google: [
28
+ 'gemini-3.1-pro', 'gemini-3.1-pro-preview',
29
+ 'gemini-3.5-flash', 'gemini-3.8-flash', 'gemini-3.7-flash', 'gemini-3.6-flash',
30
+ 'gemini-3.5-flash-lite', 'gemini-2.5-flash', 'gemini-3.1-flash-lite',
31
+ ],
22
32
  };
23
33
  /**
24
- * Identify a model's provider from its canonical key.
25
- *
26
- * TEST-NOTE: This function MUST return a non-null provider for every model
27
- * present in `MODELS` (from models.ts). If `MODELS` adds a model without
28
- * adding it to `PROVIDER_MODELS`, this returns null and routing is disabled
29
- * for that model. Silent skip.
34
+ * Priced and grouped, but never a routing TARGET: specialised models a general
35
+ * call shouldn't be moved to (grok-build-0.1 is a coding-agent model).
36
+ */
37
+ export const NOT_ROUTE_TARGETS = new Set(['grok-build-0.1']);
38
+ /**
39
+ * Models `slash-tokens/auto` may rewrite a live request TO. Frozen at the
40
+ * 1.6.5 set so a patch release never changes where production calls go:
41
+ * models added in 1.6.6 (GPT-6, Gemini 3.1 Flash-Lite / 3.6–3.8 Flash, Grok
42
+ * 4.7 / 4.5, Claude 5.5 / Fable …) are priced and shown by preflight() /
43
+ * preflightRoute(), but /auto only starts routing to them in 1.7.0.
30
44
  */
45
+ export const AUTO_ROUTE_TARGETS = new Set([
46
+ 'claude-opus-5', 'claude-opus', 'claude-opus-4.7', 'claude-sonnet-5', 'claude-sonnet',
47
+ 'claude-haiku', 'claude-haiku-4.5',
48
+ 'gpt-5.6-sol', 'gpt-5.6-terra', 'gpt-5.4', 'gpt-5.4-mini', 'gpt-5.4-nano', 'gpt-5.6-luna',
49
+ 'grok-4.6', 'grok-4.3', 'grok-4.20', 'grok-4-1-fast',
50
+ 'gemini-3.1-pro', 'gemini-3.5-flash-lite', 'gemini-2.5-flash',
51
+ ]);
31
52
  export function providerOf(model) {
53
+ const id = canonicalModel(model);
32
54
  for (const [provider, models] of Object.entries(PROVIDER_MODELS)) {
33
- if (models.includes(model))
55
+ if (models.includes(id))
34
56
  return provider;
35
57
  }
36
58
  return null;
package/dist/report.js CHANGED
@@ -24,6 +24,7 @@ export function printReport(sites, filesScanned, timeMs, cwd) {
24
24
  if (sites.length === 0) {
25
25
  console.log(`${WHITE} No AI API call sites detected.${R}`);
26
26
  console.log(`${GRAY} Supported: OpenAI, Anthropic, Vercel AI, LangChain, Gemini, Bedrock, Grok${R}`);
27
+ console.log(`${GRAY} Run this in a project that already calls an LLM.${R}`);
27
28
  console.log('');
28
29
  return;
29
30
  }
package/dist/slash.js CHANGED
@@ -1,4 +1,5 @@
1
1
  import { getInstance, writeToMemory, ensureCapacity } from './wasm.js';
2
+ import { canonicalModel } from './models.js';
2
3
  const WASM_INPUT_OFFSET = 4096;
3
4
  /**
4
5
  * Per-model calibration factors.
@@ -65,7 +66,10 @@ const WASM_INPUT_OFFSET = 4096;
65
66
  * DEFAULT_UNKNOWN_MODEL_FACTOR (1.85), a ~60% larger correction
66
67
  * than it needed. Still safe either way (1.85 > required
67
68
  * minimum), just needlessly inflated for real GPT users.
68
- * grok-4.20 / grok-4-1-fast: 1.15 — re-verified 2026-08-23 against the
69
+ * grok-4.6 / grok-4.3 (and aliases grok-4.20 / grok-4-1-fast): 1.15 —
70
+ * 4.6/4.3 IDs added 2026-08-25; factor CARRIED from the
71
+ * 2026-08-23 corpus, not re-measured on the new wire IDs.
72
+ * grok-4.20 / grok-4-1-fast (original): 1.15 — re-verified 2026-08-23 against the
69
73
  * 29-sample corpus (20 new samples: more languages, Spanish/
70
74
  * Japanese prose, more JSON shapes) and UNCHANGED — same
71
75
  * worst case (technical-docs prose, ratio 0.928) as the
@@ -91,14 +95,43 @@ const WASM_INPUT_OFFSET = 4096;
91
95
  * Slash must NEVER under-report. Over-reporting is safe (go/no-go only).
92
96
  */
93
97
  const CALIBRATION = {
98
+ 'claude-fable-5.1': 2.05,
99
+ 'claude-fable-5': 2.05,
100
+ 'claude-mythos-5.1': 2.05,
101
+ 'claude-mythos-5': 2.05,
102
+ 'claude-opus-5.5': 2.05,
103
+ 'claude-opus-5': 2.05,
104
+ 'claude-opus-4.8': 2.05,
105
+ 'claude-opus-4.6': 2.05,
106
+ 'claude-opus-4.5': 2.05,
107
+ 'claude-sonnet-5.5': 2.05,
108
+ 'claude-sonnet-4.6': 2.05,
109
+ 'claude-sonnet-4.5': 2.05,
94
110
  'claude-opus': 2.05,
95
111
  'claude-opus-4.7': 2.05,
112
+ 'claude-sonnet-5': 2.05,
96
113
  'claude-sonnet': 2.05,
114
+ 'claude-haiku-4.5': 1.45,
97
115
  'claude-haiku': 1.45,
98
116
  'gemini-3.1-pro': 1.45,
117
+ 'gemini-3.1-pro-preview': 1.45,
118
+ 'gemini-3.8-flash': 1.45,
119
+ 'gemini-3.7-flash': 1.45,
120
+ 'gemini-3.6-flash': 1.45,
121
+ 'gemini-3.5-flash': 1.45,
122
+ 'gemini-3.1-flash-lite': 1.45,
123
+ 'gemini-3.5-flash-lite': 1.45,
99
124
  'gemini-2.5-flash': 1.45,
125
+ 'grok-4.7': 1.15,
126
+ 'grok-4.6': 1.15,
127
+ 'grok-4.5': 1.15,
128
+ 'grok-build-0.1': 1.15,
129
+ 'grok-4.3': 1.15,
100
130
  'grok-4.20': 1.15,
101
131
  'grok-4-1-fast': 1.15,
132
+ 'gpt-5.6-sol': 1.15,
133
+ 'gpt-5.6-terra': 1.15,
134
+ 'gpt-5.6-luna': 1.15,
102
135
  'gpt-5.4': 1.15,
103
136
  'gpt-5.4-mini': 1.15,
104
137
  'gpt-5.4-nano': 1.15,
@@ -129,7 +162,10 @@ export function slash(content, model) {
129
162
  const raw = instance.exports.estimate_tokens(WASM_INPUT_OFFSET, len);
130
163
  if (!model)
131
164
  return raw;
132
- const factor = CALIBRATION[model] ?? DEFAULT_UNKNOWN_MODEL_FACTOR;
165
+ // GPT-6 (gpt-6-astra / 6.1-sol / 6-sol / 6-luna) is deliberately NOT in
166
+ // CALIBRATION: a new generation with an unbenchmarked tokenizer takes the
167
+ // conservative default (never under-report) until a bench run adds it.
168
+ const factor = CALIBRATION[canonicalModel(model)] ?? DEFAULT_UNKNOWN_MODEL_FACTOR;
133
169
  return factor === 1.0 ? raw : Math.ceil(raw * factor);
134
170
  }
135
171
  /**
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "slash-tokens",
3
- "version": "1.6.4",
3
+ "version": "1.6.6",
4
4
  "description": "Token Optimization for Context Engineers. 4.8 KB WASM. Sub-millisecond. Zero dependencies.",
5
5
  "main": "dist/index.js",
6
6
  "module": "dist/index.js",
@@ -26,8 +26,10 @@
26
26
  ],
27
27
  "sideEffects": false,
28
28
  "scripts": {
29
+ "prepublishOnly": "npm run build",
29
30
  "build": "npx tsc && bun build src/cli.ts --outfile=dist/cli.js --target=node",
30
- "test": "bun test"
31
+ "test": "bun test",
32
+ "check:engines": "node scripts/check-engines.mjs"
31
33
  },
32
34
  "author": "wolfejam",
33
35
  "license": "MIT",
@@ -54,5 +56,8 @@
54
56
  "@types/node": "^25.5.0",
55
57
  "js-tiktoken": "^1.0.21",
56
58
  "typescript": "^6.0.2"
59
+ },
60
+ "engines": {
61
+ "node": ">=22.0.0"
57
62
  }
58
63
  }