slash-tokens 1.6.7 → 1.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,51 @@
1
1
  # Changelog
2
2
 
3
+ ## [1.7.1] — The Hired Agent Edition
4
+
5
+ *2026-10-10*
6
+
7
+ The scan counts each real LLM request once.
8
+
9
+ No change to `quote()`, `decide()`, `reconcile()`, `hire()`, `/auto` or prices.
10
+
11
+ ### Fixed
12
+ - **The scan (`npx slash-tokens`) counts each LLM request once.** It counted every import, client constructor and request line as its own call site, so one `client.messages.create(...)` behind `import Anthropic` and `new Anthropic()` was reported as two sites (and the request itself wasn't matched unless the client variable was named `anthropic`). The monthly estimate is per site, so it came out inflated: a two-request sample project read 4 sites and $35.77/mo in 1.7.0, and reads 2 sites and $17.89/mo now. A request is now one site; a file that only imports or builds a client counts once. Requests on any client variable are matched (`.messages.create(`, `.chat.completions.create(`, `.responses.create(`, `generateText(` …), in Python too.
13
+ - **The scan no longer counts look-alike calls.** Matching any client variable made `.messages.create(` also catch Twilio's SMS API and `.invoke(` catch Tauri / Electron IPC. A request now counts only in a file that uses that SDK (an import or a client: `import anthropic`, `from openai import`, `new Anthropic()`, `AsyncOpenAI()` …), unless the call names the provider itself (`fetch` to an AI API URL, `x.ai/api`, `cohere.chat`). Python Gemini (`generate_content`) is matched too.
14
+ - The scan's footnote reads "1 call site uses an assumed price" (was "use").
15
+
16
+ ### Docs
17
+ - **Accuracy claims now say what was measured.** Request framing is measured on Nebius only; xAI comes from the Grok bench, OpenAI from its tiktoken cookbook, Anthropic and Google are allowances. The npm README's accuracy line, the root README's accuracy gate (GPT-5.x, not all GPT; GPT-6 takes the conservative default) and the `quote()` / `FRAMING` doc comments now say so. The 1.7.0 entry below called all framing "measured" and its live check didn't name Nebius; both are corrected in place.
18
+ - **`TEST-NOTES.md` is current.** It was the v1.4.0 plan: retired models, and a "preflightRoute agrees with /auto" invariant that has been false by design since 1.6.6 (`/auto` is frozen at its 1.6.5 targets). It now maps each invariant to the test that holds it and lists what isn't tested yet. Two of its invariants had no test and now do: `preflightRoute()` never picks a model the prompt doesn't fit, and `intercept.ts` uses the shared `PROVIDER_MODELS`. The `preflightRoute()` tie-break comment now matches the code (lower list price, not list order).
19
+
20
+ ## [1.7.0] — The Hired Agent Edition
21
+
22
+ *2026-10-08*
23
+
24
+ Quote the job, book the right model, prove it with a receipt.
25
+
26
+ New: `quote()`, `decide()`, `reconcile()`, `hire().run()`, `slash-tokens quote`, and NVIDIA Nemotron on Nebius. `/auto` routing is unchanged. No price change.
27
+
28
+ ### Added
29
+ - **`quote(task)`** prices a job before it runs: input tokens (calibrated, never under-reports), an output band (`minOutputTokens`–`maxOutputTokens`, default 0–4,096 and the quote says when it assumed the ceiling), a low–high USD cost, whether input plus max output fits the context window, and the date the price was checked. Accepts a string or chat messages; real API IDs are accepted.
30
+ - **`decide(task, { budget, floor })`** returns `go`, `downgrade` or `block`. Candidates are the requested model and its same-provider siblings at or above the quality floor; the cheapest one that fits and stays within budget (on the high end) wins. The default floor is the requested model's own tier, so nothing is swapped for a smaller line unless you allow it. A substitute always costs less than the requested model; it respects `init({ route: false })` and `init({ models })` like `preflightRoute()`.
31
+ - **`reconcile(quote, usage, { baseline, feeRate, waived })`** writes the receipt after a call: estimated vs actual tokens and cost (from the provider's own `usage`, OpenAI-style or Anthropic-style), `underReported` (the estimate counted fewer input tokens than were billed — Slash should never do this), whether the cost stayed within the quote, what was saved against a baseline model, and the agent's fee (10% of measured savings by default, shown and waived by default, never charged on a loss). Cached input is priced at the full input rate, so a cached call's cost is an upper bound.
32
+ - **`hire({ budget, baseline, floor, feeRate, waived }).run(job)`**: the agent you hire. Each job is quoted, booked on the cheapest model the policy allows (or blocked), run through your own `call(model, input)`, and reconciled into a receipt. The budget covers all jobs: each one may spend only what earlier jobs left. Blocked jobs never call the model. `agent.spent`, `agent.remaining`, `agent.receipts`.
33
+ - **`slash-tokens quote --model M [--file F | --text T | stdin] [--max-output N] [--min-output N] [--budget USD] [--floor 1-4] [--json]`** prints the quote and the decision (`--json` for pipelines). Exit 0 for go, downgrade and block (read `action`); 1 for bad input.
34
+ - **Quotes include each provider's request framing.** A provider bills the chat template and system preamble on top of your content. Content-only counting quoted 12 input tokens for a real Nebius call billed 23. `quote()` now adds framing: Nebius 16 per request + 7 per extra message, measured (real Token Factory calls, recorded in `bench/results-nemotron-requests.json`), xAI's 193-token system preamble (the Grok bench baseline), OpenAI's 3-per-message rule (its tiktoken cookbook), and a conservative allowance for Anthropic and Google. Two live Nebius calls after the fix: 4% and 8% over the bill, never under. `tests/request-framing.test.ts` holds every recorded bill. `preflight()` stays a content-only go/no-go.
35
+ - **`reconcile()` reads Nebius cache hits** (`prompt_cache_hit_tokens`).
36
+ - **npm metadata:** description now says what 1.7.0 does; keywords add nemotron, nvidia, nebius, agent, cost.
37
+ - **The catalog** (`CATALOG`): one source for every model's prices, context, provider, capability tier (1 small · 2 mid · 3 flagship · 4 frontier — the vendor's own line position, compared within one provider only) and `asOf` date. `MODELS` is now derived from it, with the same keys, order and fields.
38
+ - **NVIDIA Nemotron on Nebius Token Factory:** Nemotron 3 Ultra ($1.00/$3.00, 1,048,576 context), 3 Super ($0.30/$0.90, 262,144), 3 Nano ($0.06/$0.24, 262,144) and 3.5 Lightning ($0.06/$0.24, 1,048,576), in a `Nebius` provider group. Prices and context windows are what the Token Factory API reports (`/v1/models?verbose=true`, 2026-10-08). Its four API IDs (`nvidia/Nemotron-3-Ultra-550b-a55b`, `nvidia/nemotron-3-super-120b-a12b`, `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B`, `nvidia/Nemotron-3_5-Lightning`) resolve to catalog keys. Token counts use the Nemotron factor below.
39
+
40
+ - **Nemotron calibration: 1.30**, measured against the real Nemotron tokenizer (Hugging Face, run locally; the four models share it). The worst case on the 29-sample corpus (a large JSON API response) needs ≥ 1.208; 1.30 leaves 7.6% headroom. Nemotron previously took the unknown-model default (2.05), 70% above what even the worst sample needs. `npm run bench:nemotron` reproduces it; `bench/REPORT-nemotron.md`.
41
+ - **Accuracy gate in CI** (`tests/accuracy-gate.test.ts`): every calibrated estimate must be at least the provider's real count on all 29 corpus samples, for the Claude, Gemini, Grok and Nemotron families (recorded bench results) and GPT-5.x (o200k_base, computed in the test). A change that breaks "never under-report" now fails CI.
42
+ - **Price-freshness alarm** (`npm run check:freshness`, weekly `freshness.yml`): fails when any catalog price is more than 30 days old, or an announced price change (`priceUntil`, e.g. Gemini 3.6–3.8 Flash on 2026-12-31) is within 14 days. Runs even when mcpaas.live is down.
43
+
44
+ ### Changed
45
+ - `preflight()` options (the cross-provider analysis) now include the Nemotron models, so the cheapest option can be a Nemotron model. `preflightRoute()` is unchanged: same-provider only.
46
+ - **`slash-tokens/auto` stays frozen at its 1.6.5 routing targets.** 1.6.6 said /auto would start routing to the newer models in 1.7.0; we've decided against changing where live production calls go inside a feature release. The newer models are priced and recognised in /auto events, and `preflightRoute()` and `decide()` use them. Moving /auto's targets will be its own release, announced first.
47
+ - Root README: the `report()` example now uses the real `action` values (`'prevented' | 'routed' | 'pass'`).
48
+
3
49
  ## [1.6.7] — The Fixed Deal Edition
4
50
 
5
51
  *2026-10-08*
package/README.md CHANGED
@@ -7,71 +7,71 @@
7
7
  [![license](https://img.shields.io/npm/l/slash-tokens?style=flat)](./LICENSE)
8
8
  [![⭐ Star on GitHub](https://img.shields.io/badge/%E2%AD%90_Star-black?logo=github&logoColor=white)](https://github.com/Wolfe-Jam/slash-tokens)
9
9
 
10
- Token Optimization for Context Engineers.
11
- For anyone building with LLMs. 4.8 KB WASM. Sub-millisecond. Zero dependencies.
10
+ **Know what an LLM call will cost before you make it, and prove what you saved after.**
11
+ 4.8 KB WASM · sub-millisecond · zero dependencies · Claude, GPT, Grok, Gemini, Nemotron.
12
12
 
13
- Know the cost before the call leaves your machine.
13
+ ## v1.7.1 — The Hired Agent Edition
14
14
 
15
- Models change. Windows grow. Slash adapts — you keep building.
16
- Cheaper tokens haven't shrunk the bill — usage has.
15
+ Quote the job, book the right model, prove it with a receipt.
17
16
 
18
- ## v1.6.7 — The Fixed Deal Edition
17
+ New in 1.7.1: the scan counts each real LLM request once.
19
18
 
20
- One very large request (about 1 MB of text) no longer skews every count after it.
19
+ ```js
20
+ import { hire } from 'slash-tokens'
21
21
 
22
- Solo $20 mailbox, 10% waived. Team $39 for the data.
22
+ const agent = hire({ budget: 0.50, floor: 2 }) // $0.50 to spend; nothing below a mid-tier model
23
23
 
24
- New in 1.6.7: a single request over ~1 MB of text used to overwrite the token counter's lookup tables, so every later count in that process came out wrong; prompts now go above them. Counts for normal prompts are unchanged.
24
+ const { decision, receipt } = await agent.run({
25
+ input: prompt,
26
+ model: 'nemotron-3-ultra',
27
+ maxOutputTokens: 2000,
28
+ call: (model, input) => yourClient(model, input), // returns { output, usage }
29
+ })
25
30
 
26
- New in 1.6.6: Claude Opus 5.5 / Sonnet 5.5 / Fable 5.1, Grok 4.7, GPT-6 (Astra, Sol, Luna), Gemini 3.6–3.8 Flash and 3.1 Flash-Lite, priced as of 2026-10-07. Real API IDs (`claude-opus-4-7`) work in `preflight()`. `preflightRoute()` now finds GPT-6 Luna and Gemini 3.1 Flash-Lite as the cheapest same-provider options.
31
+ decision.action // 'downgrade': nemotron-3-super does the job for less
32
+ receipt.saved // dollars saved against the model you asked for
33
+ ```
27
34
 
28
- **Free forever is bunx** — no account. A one-person account is email → key, **$20 on the house**. We show the savings. We don't charge. 10% is the model, waived. Team is **$39 for the data** (`$390`/year).
35
+ - **Quote:** `quote()` prices a job before it runs, low to high.
36
+ - **Decide:** `decide()` goes, downgrades or blocks, within your budget and quality floor. Same provider, never a pricier model.
37
+ - **Prove:** `reconcile()` checks the quote against what the provider billed.
38
+ - **Hire:** `hire()` does all three for every job, inside one budget. Its cut, 10% of measured savings, is on every receipt, waived.
29
39
 
30
- ```bash
31
- bunx slash-tokens
32
- # or: npx --yes slash-tokens
33
- ```
40
+ Counts are calibrated against each provider's real tokenizer; models not benchmarked yet (GPT-6) get a conservative default. Quotes add what the provider bills on top of your text, measured on Nebius and a conservative allowance elsewhere. CI fails if a count comes in under the real one, or a quote under a real Nebius bill.
34
41
 
35
- Run it in a project that already calls an LLM. An empty folder prints that nothing was found, then tells you to run it in an app — that's normal. `--version` / `--help` print and exit (they do not scan). Pin proof: `slash-tokens --version` or `npm view slash-tokens version`.
42
+ ## Install
36
43
 
37
44
  ```bash
38
45
  npm install slash-tokens
39
46
  ```
40
47
 
41
- ```js
42
- import { preflight, preflightRoute } from 'slash-tokens'
43
-
44
- // Analysis — cheaper alternatives across all providers. Not a route.
45
- const check = preflight(prompt, 'claude-opus-5')
46
- check.tokens
47
- check.cost
48
- check.fits
49
- check.options
48
+ ## From the terminal
50
49
 
51
- // Routing decision — same-provider only, matches the Slash proxy
52
- const route = preflightRoute(prompt, 'claude-opus-5')
53
- // { model: 'claude-haiku', cost, salvaged, salvagePercent } or null
50
+ ```bash
51
+ npx slash-tokens # find the LLM calls in this project and what they cost a month
52
+ echo "Fix the bug" | npx slash-tokens quote --model nemotron-3-ultra --floor 1
54
53
  ```
55
54
 
56
- Or one line — every LLM call checked pre-call:
55
+ ## Check every call, automatically
57
56
 
58
57
  ```js
59
58
  import 'slash-tokens/auto'
60
59
  ```
61
60
 
62
- Intercepts `fetch()` to Anthropic, OpenAI, xAI, and Google. Estimates before the call leaves your machine. Same-provider cheaper swap if one fits.
63
-
64
- ## See it work
61
+ Checks each request to Anthropic, OpenAI, xAI and Google before it leaves your machine, and swaps in a cheaper model from the same provider when one fits.
65
62
 
66
- A live chat with every call through the gate: [live demo](https://slash-nextjs-wofejams-projects.vercel.app)
63
+ ## Per-call checks
67
64
 
68
- Then `bunx slash-tokens`, [get a key](https://mcpaas.live/slash/setup) ($20 on the house), or [Team — $39 for the data](https://slashtokens.com).
65
+ ```js
66
+ import { preflight, preflightRoute } from 'slash-tokens'
69
67
 
70
- ## Dashboard
68
+ const check = preflight(prompt, 'claude-opus-5') // tokens, cost, fits, cheaper options (all providers)
69
+ const route = preflightRoute(prompt, 'claude-opus-5') // the cheapest same-provider model that fits, or null
70
+ ```
71
71
 
72
- Track savings across all your apps. One-person key (email, $20 on the house) at [mcpaas.live/slash/setup](https://mcpaas.live/slash/setup)
72
+ ## Pricing
73
73
 
74
- Full docs, examples, and model pricing at **[GitHub](https://github.com/Wolfe-Jam/slash-tokens)**
74
+ The library and CLI are free, no account needed. A one-person key is $20 on the house: we show the savings and don't charge. Team is $39/month for the data. [Live demo](https://slash-nextjs-wofejams-projects.vercel.app) · [Get a key](https://mcpaas.live/slash/setup) · [slashtokens.com](https://slashtokens.com) · [Full docs](https://github.com/Wolfe-Jam/slash-tokens)
75
75
 
76
76
  ## License
77
77
 
@@ -0,0 +1,50 @@
1
+ import { type Task, type Decision, type Message } from './quote.js';
2
+ import { type Receipt, type Usage } from './receipt.js';
3
+ import type { Tier } from './catalog.js';
4
+ export interface HireOptions {
5
+ /** Total USD the agent may spend across every job it runs. No budget = no limit. */
6
+ budget?: number;
7
+ /** Model the savings are measured against. Default: the model each job asks for. */
8
+ baseline?: string;
9
+ /** Lowest tier a substitute may have. Default: each job's requested tier. */
10
+ floor?: Tier;
11
+ /** The agent's cut of measured savings. Default 0.10. */
12
+ feeRate?: number;
13
+ /** Fee shown on the receipt but not charged. Default true. */
14
+ waived?: boolean;
15
+ }
16
+ /**
17
+ * What the agent calls to do the work: your provider client. `model` is the
18
+ * booked catalog key (e.g. `nemotron-3.5-lightning`); map it to your
19
+ * provider's API ID. Return the output and the provider's `usage` object.
20
+ */
21
+ export type CallModel = (model: string, input: string | Message[]) => Promise<{
22
+ output: string;
23
+ usage: Usage;
24
+ }>;
25
+ export interface Job extends Task {
26
+ call: CallModel;
27
+ }
28
+ export interface RunResult {
29
+ decision: Decision;
30
+ /** The model's output; undefined when blocked. */
31
+ output?: string;
32
+ /** Estimate vs actual, saved vs baseline, fee; undefined when blocked. */
33
+ receipt?: Receipt;
34
+ }
35
+ export interface Agent {
36
+ /** Quote → decide → book → run → reconcile → receipt. */
37
+ run(job: Job): Promise<RunResult>;
38
+ /** USD spent so far (actual cost, from receipts). */
39
+ readonly spent: number;
40
+ /** USD left in the budget (Infinity with no budget). */
41
+ readonly remaining: number;
42
+ readonly receipts: readonly Receipt[];
43
+ }
44
+ /**
45
+ * Hire Slash for a run of jobs. Each job is quoted, booked on the cheapest
46
+ * model the policy allows (or blocked), run through your `call`, and
47
+ * reconciled against the provider's own usage. The budget covers all jobs:
48
+ * each one may spend only what earlier jobs left.
49
+ */
50
+ export declare function hire(opts?: HireOptions): Agent;
package/dist/agent.js ADDED
@@ -0,0 +1,36 @@
1
+ import { decide } from './quote.js';
2
+ import { reconcile } from './receipt.js';
3
+ /**
4
+ * Hire Slash for a run of jobs. Each job is quoted, booked on the cheapest
5
+ * model the policy allows (or blocked), run through your `call`, and
6
+ * reconciled against the provider's own usage. The budget covers all jobs:
7
+ * each one may spend only what earlier jobs left.
8
+ */
9
+ export function hire(opts = {}) {
10
+ const receipts = [];
11
+ let spent = 0;
12
+ const budget = opts.budget ?? Infinity;
13
+ return {
14
+ get spent() { return spent; },
15
+ get remaining() { return Math.max(budget - spent, 0); },
16
+ get receipts() { return receipts; },
17
+ async run(job) {
18
+ const { call, ...task } = job;
19
+ const policy = { floor: opts.floor };
20
+ if (budget !== Infinity)
21
+ policy.budget = Math.max(budget - spent, 0);
22
+ const decision = decide(task, policy);
23
+ if (decision.action === 'block' || !decision.chosen)
24
+ return { decision };
25
+ const { output, usage } = await call(decision.chosen.model, task.input);
26
+ const receipt = reconcile(decision.chosen, usage, {
27
+ baseline: opts.baseline ?? decision.requested.model,
28
+ feeRate: opts.feeRate,
29
+ waived: opts.waived,
30
+ });
31
+ spent = Math.round((spent + receipt.actual.cost) * 1000000) / 1000000;
32
+ receipts.push(receipt);
33
+ return { decision, output, receipt };
34
+ },
35
+ };
36
+ }
@@ -0,0 +1,38 @@
1
+ /**
2
+ * The catalog — one source for every priced model.
3
+ *
4
+ * Each entry carries its prices (USD per million tokens), context window,
5
+ * provider, capability tier and the date the price was checked. `MODELS`
6
+ * (models.ts) and the quote / decide policy (quote.ts) both read from here.
7
+ *
8
+ * Tiers are the vendor's own line position, not a benchmark:
9
+ * 4 frontier — the vendor's top line above its flagship (Fable/Mythos, GPT-6 Astra)
10
+ * 3 flagship — Opus, Sol, Grok 4.5+, Gemini Pro, Nemotron Ultra, GPT-5.4
11
+ * 2 mid — Sonnet, Terra, Grok 4.3, Gemini Flash, Nemotron Super, GPT-5.4 mini
12
+ * 1 small — Haiku, Luna, Flash-Lite, GPT-5.4 nano, Nemotron Nano / Lightning
13
+ * A quality floor compares tiers within one provider only; tiers never rank
14
+ * one vendor's model against another's.
15
+ */
16
+ export type Tier = 1 | 2 | 3 | 4;
17
+ export declare const TIER_NAMES: Record<Tier, string>;
18
+ export interface CatalogEntry {
19
+ provider: string;
20
+ tier: Tier;
21
+ /** Date the price was checked against the provider's own pricing page. */
22
+ asOf: string;
23
+ /** Last day this price holds, when the provider has announced a change. */
24
+ priceUntil?: string;
25
+ input: number;
26
+ output: number;
27
+ context: number;
28
+ longContextThreshold?: number;
29
+ longContextInput?: number;
30
+ longContextOutput?: number;
31
+ }
32
+ export declare const CATALOG: Record<string, CatalogEntry>;
33
+ /**
34
+ * Real API IDs that don't follow the table's naming → table keys.
35
+ * Lowercased; canonicalModel() lowercases before looking here. The Nebius IDs
36
+ * are the ones Token Factory lists (GET /v1/models, 2026-10-08).
37
+ */
38
+ export declare const API_ALIASES: Record<string, string>;
@@ -0,0 +1,143 @@
1
+ /**
2
+ * The catalog — one source for every priced model.
3
+ *
4
+ * Each entry carries its prices (USD per million tokens), context window,
5
+ * provider, capability tier and the date the price was checked. `MODELS`
6
+ * (models.ts) and the quote / decide policy (quote.ts) both read from here.
7
+ *
8
+ * Tiers are the vendor's own line position, not a benchmark:
9
+ * 4 frontier — the vendor's top line above its flagship (Fable/Mythos, GPT-6 Astra)
10
+ * 3 flagship — Opus, Sol, Grok 4.5+, Gemini Pro, Nemotron Ultra, GPT-5.4
11
+ * 2 mid — Sonnet, Terra, Grok 4.3, Gemini Flash, Nemotron Super, GPT-5.4 mini
12
+ * 1 small — Haiku, Luna, Flash-Lite, GPT-5.4 nano, Nemotron Nano / Lightning
13
+ * A quality floor compares tiers within one provider only; tiers never rank
14
+ * one vendor's model against another's.
15
+ */
16
+ export const TIER_NAMES = {
17
+ 4: 'frontier',
18
+ 3: 'flagship',
19
+ 2: 'mid',
20
+ 1: 'small',
21
+ };
22
+ const OPUS = { input: 5.00, output: 25.00, context: 1000000 };
23
+ const OPUS_55 = { input: 4.00, output: 20.00, context: 1000000 };
24
+ const FABLE = { input: 10.00, output: 50.00, context: 1000000 };
25
+ const SONNET_4X = { input: 3.00, output: 15.00, context: 1000000 };
26
+ const SONNET = { input: 2.00, output: 10.00, context: 1000000 };
27
+ const HAIKU = { input: 1.00, output: 5.00, context: 200000 };
28
+ const GROK_46 = {
29
+ input: 2.00, output: 6.00, context: 500000,
30
+ longContextThreshold: 200000, longContextInput: 4.00, longContextOutput: 12.00,
31
+ };
32
+ const GROK_43 = {
33
+ input: 1.25, output: 2.50, context: 1000000,
34
+ longContextThreshold: 200000, longContextInput: 2.50, longContextOutput: 5.00,
35
+ };
36
+ const GEMINI_PRO = {
37
+ input: 2.00, output: 12.00, context: 1000000,
38
+ longContextThreshold: 200000, longContextInput: 4.00, longContextOutput: 18.00,
39
+ };
40
+ const GEMINI_FLASH = { input: 0.30, output: 2.50, context: 1000000 };
41
+ // Gemini 3.6–3.8 Flash: launch price through 2026-12-31; Google lists $1.50/$7.50
42
+ // from 2027-01-01. priceUntil makes the freshness check fail before then.
43
+ const GEMINI_FLASH_3X = { input: 0.75, output: 3.75, context: 1000000, priceUntil: '2026-12-31' };
44
+ const GEMINI_35_FLASH = { input: 1.50, output: 9.00, context: 1000000 };
45
+ const GEMINI_31_FLASH_LITE = { input: 0.25, output: 1.50, context: 1000000 };
46
+ const GROK_BUILD = {
47
+ input: 1.00, output: 2.00, context: 256000,
48
+ longContextThreshold: 200000, longContextInput: 2.00, longContextOutput: 4.00,
49
+ };
50
+ const GPT_6_ASTRA = { input: 10.00, output: 50.00, context: 1050000 };
51
+ const GPT_6_SOL = { input: 2.00, output: 10.00, context: 1050000 };
52
+ const GPT_6_LUNA = { input: 0.10, output: 0.50, context: 1050000 };
53
+ const GPT_SOL = { input: 4.00, output: 20.00, context: 1050000 };
54
+ const GPT_TERRA = { input: 2.00, output: 12.00, context: 1050000 };
55
+ const GPT_LUNA = { input: 0.20, output: 1.20, context: 1050000 };
56
+ const GPT_54 = { input: 2.50, output: 15.00, context: 1000000 };
57
+ const GPT_54_MINI = { input: 0.75, output: 4.50, context: 128000 };
58
+ const GPT_54_NANO = { input: 0.20, output: 1.25, context: 128000 };
59
+ // NVIDIA Nemotron on Nebius Token Factory. Prices and context windows are
60
+ // exactly what the API reports (GET /v1/models?verbose=true, 2026-10-08).
61
+ const NEMOTRON_ULTRA = { input: 1.00, output: 3.00, context: 1048576 };
62
+ const NEMOTRON_SUPER = { input: 0.30, output: 0.90, context: 262144 };
63
+ const NEMOTRON_NANO = { input: 0.06, output: 0.24, context: 262144 };
64
+ const NEMOTRON_LIGHTNING = { input: 0.06, output: 0.24, context: 1048576 };
65
+ // Prices as of 2026-10-07 — first-party pages:
66
+ // platform.claude.com/docs/en/about-claude/pricing
67
+ // developers.openai.com/api/docs/models
68
+ // docs.x.ai/developers/models
69
+ // ai.google.dev/gemini-api/docs/pricing
70
+ const ASOF = '2026-10-07';
71
+ // Nemotron: read from the Token Factory API (see above).
72
+ const ASOF_NEBIUS = '2026-10-08';
73
+ function e(provider, tier, price) {
74
+ return { provider, tier, asOf: provider === 'Nebius' ? ASOF_NEBIUS : ASOF, ...price };
75
+ }
76
+ // Order is kept from the 1.6.6 MODELS table (new entries appended), so
77
+ // preflight()'s option ordering for equal costs doesn't move.
78
+ export const CATALOG = {
79
+ // Anthropic — live names + generic aliases (same rates)
80
+ 'claude-fable-5.1': e('Anthropic', 4, FABLE),
81
+ 'claude-fable-5': e('Anthropic', 4, FABLE),
82
+ 'claude-mythos-5.1': e('Anthropic', 4, FABLE),
83
+ 'claude-mythos-5': e('Anthropic', 4, FABLE),
84
+ 'claude-opus-5.5': e('Anthropic', 3, OPUS_55),
85
+ 'claude-opus-5': e('Anthropic', 3, OPUS),
86
+ 'claude-opus-4.8': e('Anthropic', 3, OPUS),
87
+ 'claude-opus': e('Anthropic', 3, OPUS),
88
+ 'claude-opus-4.7': e('Anthropic', 3, OPUS),
89
+ 'claude-opus-4.6': e('Anthropic', 3, OPUS),
90
+ 'claude-opus-4.5': e('Anthropic', 3, OPUS),
91
+ 'claude-sonnet-5.5': e('Anthropic', 2, SONNET),
92
+ 'claude-sonnet-5': e('Anthropic', 2, SONNET),
93
+ 'claude-sonnet': e('Anthropic', 2, SONNET),
94
+ 'claude-sonnet-4.6': e('Anthropic', 2, SONNET_4X),
95
+ 'claude-sonnet-4.5': e('Anthropic', 2, SONNET_4X),
96
+ 'claude-haiku-4.5': e('Anthropic', 1, HAIKU),
97
+ 'claude-haiku': e('Anthropic', 1, HAIKU),
98
+ // xAI — flagship 4.6, cheap same-provider 4.3. 4.20 / fast are aliases.
99
+ 'grok-4.7': e('xAI', 3, GROK_46),
100
+ 'grok-4.6': e('xAI', 3, GROK_46),
101
+ 'grok-4.5': e('xAI', 3, GROK_46),
102
+ 'grok-build-0.1': e('xAI', 2, GROK_BUILD),
103
+ 'grok-4.3': e('xAI', 2, GROK_43),
104
+ 'grok-4.20': e('xAI', 2, GROK_43),
105
+ 'grok-4-1-fast': e('xAI', 2, GROK_43),
106
+ // Google
107
+ 'gemini-3.1-pro': e('Google', 3, GEMINI_PRO),
108
+ 'gemini-3.1-pro-preview': e('Google', 3, GEMINI_PRO),
109
+ 'gemini-3.8-flash': e('Google', 2, GEMINI_FLASH_3X),
110
+ 'gemini-3.7-flash': e('Google', 2, GEMINI_FLASH_3X),
111
+ 'gemini-3.6-flash': e('Google', 2, GEMINI_FLASH_3X),
112
+ 'gemini-3.5-flash': e('Google', 2, GEMINI_35_FLASH),
113
+ 'gemini-3.1-flash-lite': e('Google', 1, GEMINI_31_FLASH_LITE),
114
+ 'gemini-3.5-flash-lite': e('Google', 1, GEMINI_FLASH),
115
+ 'gemini-2.5-flash': e('Google', 2, GEMINI_FLASH),
116
+ // OpenAI — live 5.6 ladder + GPT-6. 5.4 family kept as aliases (old prices).
117
+ 'gpt-6-astra': e('OpenAI', 4, GPT_6_ASTRA),
118
+ 'gpt-6.1-sol': e('OpenAI', 3, GPT_6_SOL),
119
+ 'gpt-6-sol': e('OpenAI', 3, GPT_6_SOL),
120
+ 'gpt-6-luna': e('OpenAI', 1, GPT_6_LUNA),
121
+ 'gpt-5.6-sol': e('OpenAI', 3, GPT_SOL),
122
+ 'gpt-5.6-terra': e('OpenAI', 2, GPT_TERRA),
123
+ 'gpt-5.6-luna': e('OpenAI', 1, GPT_LUNA),
124
+ 'gpt-5.4': e('OpenAI', 3, GPT_54),
125
+ 'gpt-5.4-mini': e('OpenAI', 2, GPT_54_MINI),
126
+ 'gpt-5.4-nano': e('OpenAI', 1, GPT_54_NANO),
127
+ // NVIDIA Nemotron, served and billed by Nebius Token Factory
128
+ 'nemotron-3-ultra': e('Nebius', 3, NEMOTRON_ULTRA),
129
+ 'nemotron-3-super': e('Nebius', 2, NEMOTRON_SUPER),
130
+ 'nemotron-3-nano': e('Nebius', 1, NEMOTRON_NANO),
131
+ 'nemotron-3.5-lightning': e('Nebius', 1, NEMOTRON_LIGHTNING),
132
+ };
133
+ /**
134
+ * Real API IDs that don't follow the table's naming → table keys.
135
+ * Lowercased; canonicalModel() lowercases before looking here. The Nebius IDs
136
+ * are the ones Token Factory lists (GET /v1/models, 2026-10-08).
137
+ */
138
+ export const API_ALIASES = {
139
+ 'nvidia/nemotron-3-ultra-550b-a55b': 'nemotron-3-ultra',
140
+ 'nvidia/nemotron-3-super-120b-a12b': 'nemotron-3-super',
141
+ 'nvidia/nemotron-3_5-lightning': 'nemotron-3.5-lightning',
142
+ 'nvidia/nvidia-nemotron-3-nano-30b-a3b': 'nemotron-3-nano',
143
+ };