slash-tokens 1.6.6 → 1.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,48 @@
1
1
  # Changelog
2
2
 
3
+ ## [1.7.0] — The Hired Agent Edition
4
+
5
+ *2026-10-08*
6
+
7
+ Quote the job, book the right model, prove it with a receipt.
8
+
9
+ New: `quote()`, `decide()`, `reconcile()`, `hire().run()`, `slash-tokens quote`, and NVIDIA Nemotron on Nebius. `/auto` routing is unchanged. No price change.
10
+
11
+ ### Added
12
+ - **`quote(task)`** prices a job before it runs: input tokens (calibrated, never under-reports), an output band (`minOutputTokens`–`maxOutputTokens`, default 0–4,096 and the quote says when it assumed the ceiling), a low–high USD cost, whether input plus max output fits the context window, and the date the price was checked. Accepts a string or chat messages; real API IDs are accepted.
13
+ - **`decide(task, { budget, floor })`** returns `go`, `downgrade` or `block`. Candidates are the requested model and its same-provider siblings at or above the quality floor; the cheapest one that fits and stays within budget (on the high end) wins. The default floor is the requested model's own tier, so nothing is swapped for a smaller line unless you allow it. A substitute always costs less than the requested model; it respects `init({ route: false })` and `init({ models })` like `preflightRoute()`.
14
+ - **`reconcile(quote, usage, { baseline, feeRate, waived })`** writes the receipt after a call: estimated vs actual tokens and cost (from the provider's own `usage`, OpenAI-style or Anthropic-style), `underReported` (the estimate counted fewer input tokens than were billed — Slash should never do this), whether the cost stayed within the quote, what was saved against a baseline model, and the agent's fee (10% of measured savings by default, shown and waived by default, never charged on a loss). Cached input is priced at the full input rate, so a cached call's cost is an upper bound.
15
+ - **`hire({ budget, baseline, floor, feeRate, waived }).run(job)`**: the agent you hire. Each job is quoted, booked on the cheapest model the policy allows (or blocked), run through your own `call(model, input)`, and reconciled into a receipt. The budget covers all jobs: each one may spend only what earlier jobs left. Blocked jobs never call the model. `agent.spent`, `agent.remaining`, `agent.receipts`.
16
+ - **`slash-tokens quote --model M [--file F | --text T | stdin] [--max-output N] [--min-output N] [--budget USD] [--floor 1-4] [--json]`** prints the quote and the decision (`--json` for pipelines). Exit 0 for go, downgrade and block (read `action`); 1 for bad input.
17
+ - **Quotes include each provider's request framing.** A provider bills the chat template and system preamble on top of your content. Content-only counting quoted 12 input tokens for a real Nebius call billed 23. `quote()` now adds measured framing: Nebius 16 per request + 7 per extra message (real Token Factory calls, recorded in `bench/results-nemotron-requests.json`), xAI's 193-token system preamble (the Grok bench baseline), OpenAI's 3-per-message rule, and a conservative allowance for Anthropic and Google. Live check after the fix: 4% and 8% over the bill, never under. `tests/request-framing.test.ts` holds every recorded bill. `preflight()` stays a content-only go/no-go.
18
+ - **`reconcile()` reads Nebius cache hits** (`prompt_cache_hit_tokens`).
19
+ - **npm metadata:** description now says what 1.7.0 does; keywords add nemotron, nvidia, nebius, agent, cost.
20
+ - **The catalog** (`CATALOG`): one source for every model's prices, context, provider, capability tier (1 small · 2 mid · 3 flagship · 4 frontier — the vendor's own line position, compared within one provider only) and `asOf` date. `MODELS` is now derived from it, with the same keys, order and fields.
21
+ - **NVIDIA Nemotron on Nebius Token Factory:** Nemotron 3 Ultra ($1.00/$3.00, 1,048,576 context), 3 Super ($0.30/$0.90, 262,144), 3 Nano ($0.06/$0.24, 262,144) and 3.5 Lightning ($0.06/$0.24, 1,048,576), in a `Nebius` provider group. Prices and context windows are what the Token Factory API reports (`/v1/models?verbose=true`, 2026-10-08). Its four API IDs (`nvidia/Nemotron-3-Ultra-550b-a55b`, `nvidia/nemotron-3-super-120b-a12b`, `nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B`, `nvidia/Nemotron-3_5-Lightning`) resolve to catalog keys. Token counts use the Nemotron factor below.
22
+
23
+ - **Nemotron calibration: 1.30**, measured against the real Nemotron tokenizer (Hugging Face, run locally; the four models share it). The worst case on the 29-sample corpus (a large JSON API response) needs ≥ 1.208; 1.30 leaves 7.6% headroom. Nemotron previously took the unknown-model default (2.05), 70% above what even the worst sample needs. `npm run bench:nemotron` reproduces it; `bench/REPORT-nemotron.md`.
24
+ - **Accuracy gate in CI** (`tests/accuracy-gate.test.ts`): every calibrated estimate must be at least the provider's real count on all 29 corpus samples, for the Claude, Gemini, Grok and Nemotron families (recorded bench results) and GPT-5.x (o200k_base, computed in the test). A change that breaks "never under-report" now fails CI.
25
+ - **Price-freshness alarm** (`npm run check:freshness`, weekly `freshness.yml`): fails when any catalog price is more than 30 days old, or an announced price change (`priceUntil`, e.g. Gemini 3.6–3.8 Flash on 2026-12-31) is within 14 days. Runs even when mcpaas.live is down.
26
+
27
+ ### Changed
28
+ - `preflight()` options (the cross-provider analysis) now include the Nemotron models, so the cheapest option can be a Nemotron model. `preflightRoute()` is unchanged: same-provider only.
29
+ - **`slash-tokens/auto` stays frozen at its 1.6.5 routing targets.** 1.6.6 said /auto would start routing to the newer models in 1.7.0; we've decided against changing where live production calls go inside a feature release. The newer models are priced and recognised in /auto events, and `preflightRoute()` and `decide()` use them. Moving /auto's targets will be its own release, announced first.
30
+ - Root README: the `report()` example now uses the real `action` values (`'prevented' | 'routed' | 'pass'`).
31
+
32
+ ## [1.6.7] — The Fixed Deal Edition
33
+
34
+ *2026-10-08*
35
+
36
+ One very large request (about 1 MB of text) no longer skews every count after it.
37
+
38
+ Solo $20 mailbox, 10% waived. Team $39 for the data.
39
+
40
+ ### Fixed
41
+ - **A single request over ~1.04 MB of text no longer skews every later estimate in the same process.** Input was written into WASM memory from byte 4096, so a big one ran over the module's stack and lookup tables (at 1 MiB), and later estimates came out wrong — a 14-token prompt read as 28. Input now starts above them (byte 1,114,112). Estimates for inputs under the threshold are unchanged (checked against 1.6.6 on the 29-sample corpus and five models). Older versions are affected too (any version that writes input at byte 4096, including 1.6.6).
42
+ - **The test suite now passes file by file.** `preflight.test.ts` failed when run alone: on a tiny prompt the routed model's cost rounds to the same 6-decimal value as the original's. It passed in the full suite only because the bug above had already inflated the counts. The test now checks a lower list price, and strictly lower cost on a prompt big enough to show it.
43
+
44
+ No Team/Solo price change. No routing or pricing change.
45
+
3
46
  ## [1.6.6] — The Fixed Deal Edition
4
47
 
5
48
  *2026-10-07*
package/README.md CHANGED
@@ -7,69 +7,69 @@
7
7
  [![license](https://img.shields.io/npm/l/slash-tokens?style=flat)](./LICENSE)
8
8
  [![⭐ Star on GitHub](https://img.shields.io/badge/%E2%AD%90_Star-black?logo=github&logoColor=white)](https://github.com/Wolfe-Jam/slash-tokens)
9
9
 
10
- Token Optimization for Context Engineers.
11
- For anyone building with LLMs. 4.8 KB WASM. Sub-millisecond. Zero dependencies.
10
+ **Know what an LLM call will cost before you make it, and prove what you saved after.**
11
+ 4.8 KB WASM · sub-millisecond · zero dependencies · Claude, GPT, Grok, Gemini, Nemotron.
12
12
 
13
- Know the cost before the call leaves your machine.
13
+ ## v1.7.0 — The Hired Agent Edition
14
14
 
15
- Models change. Windows grow. Slash adapts — you keep building.
16
- Cheaper tokens haven't shrunk the bill — usage has.
15
+ Quote the job, book the right model, prove it with a receipt.
17
16
 
18
- ## v1.6.6 — The Fixed Deal Edition
17
+ ```js
18
+ import { hire } from 'slash-tokens'
19
19
 
20
- `--version` / `--help` answer, and today's models price correctly.
20
+ const agent = hire({ budget: 0.50, floor: 2 }) // $0.50 to spend; nothing below a mid-tier model
21
21
 
22
- Solo $20 mailbox, 10% waived. Team $39 for the data.
22
+ const { decision, receipt } = await agent.run({
23
+ input: prompt,
24
+ model: 'nemotron-3-ultra',
25
+ maxOutputTokens: 2000,
26
+ call: (model, input) => yourClient(model, input), // returns { output, usage }
27
+ })
23
28
 
24
- New in 1.6.6: Claude Opus 5.5 / Sonnet 5.5 / Fable 5.1, Grok 4.7, GPT-6 (Astra, Sol, Luna), Gemini 3.6–3.8 Flash and 3.1 Flash-Lite, priced as of 2026-10-07. Real API IDs (`claude-opus-4-7`) work in `preflight()`. `preflightRoute()` now finds GPT-6 Luna and Gemini 3.1 Flash-Lite as the cheapest same-provider options.
29
+ decision.action // 'downgrade': nemotron-3-super does the job for less
30
+ receipt.saved // dollars saved against the model you asked for
31
+ ```
25
32
 
26
- **Free forever is bunx** — no account. A one-person account is email → key, **$20 on the house**. We show the savings. We don't charge. 10% is the model, waived. Team is **$39 for the data** (`$390`/year).
33
+ - **Quote:** `quote()` prices a job before it runs, low to high.
34
+ - **Decide:** `decide()` goes, downgrades or blocks, within your budget and quality floor. Same provider, never a pricier model.
35
+ - **Prove:** `reconcile()` checks the quote against what the provider billed.
36
+ - **Hire:** `hire()` does all three for every job, inside one budget. Its cut, 10% of measured savings, is on every receipt, waived.
27
37
 
28
- ```bash
29
- bunx slash-tokens
30
- # or: npx --yes slash-tokens
31
- ```
38
+ Counts are calibrated against each provider's real tokenizer, quotes add what each provider bills on top of your text, and CI fails if either comes in under the real number.
32
39
 
33
- Run it in a project that already calls an LLM. An empty folder prints that nothing was found, then tells you to run it in an app — that's normal. `--version` / `--help` print and exit (they do not scan). Pin proof: `slash-tokens --version` or `npm view slash-tokens version`.
40
+ ## Install
34
41
 
35
42
  ```bash
36
43
  npm install slash-tokens
37
44
  ```
38
45
 
39
- ```js
40
- import { preflight, preflightRoute } from 'slash-tokens'
41
-
42
- // Analysis — cheaper alternatives across all providers. Not a route.
43
- const check = preflight(prompt, 'claude-opus-5')
44
- check.tokens
45
- check.cost
46
- check.fits
47
- check.options
46
+ ## From the terminal
48
47
 
49
- // Routing decision — same-provider only, matches the Slash proxy
50
- const route = preflightRoute(prompt, 'claude-opus-5')
51
- // { model: 'claude-haiku', cost, salvaged, salvagePercent } or null
48
+ ```bash
49
+ npx slash-tokens # find the LLM calls in this project and what they cost a month
50
+ echo "Fix the bug" | npx slash-tokens quote --model nemotron-3-ultra --floor 1
52
51
  ```
53
52
 
54
- Or one line — every LLM call checked pre-call:
53
+ ## Check every call, automatically
55
54
 
56
55
  ```js
57
56
  import 'slash-tokens/auto'
58
57
  ```
59
58
 
60
- Intercepts `fetch()` to Anthropic, OpenAI, xAI, and Google. Estimates before the call leaves your machine. Same-provider cheaper swap if one fits.
61
-
62
- ## See it work
59
+ Checks each request to Anthropic, OpenAI, xAI and Google before it leaves your machine, and swaps in a cheaper model from the same provider when one fits.
63
60
 
64
- A live chat with every call through the gate: [live demo](https://slash-nextjs-wofejams-projects.vercel.app)
61
+ ## Per-call checks
65
62
 
66
- Then `bunx slash-tokens`, [get a key](https://mcpaas.live/slash/setup) ($20 on the house), or [Team — $39 for the data](https://slashtokens.com).
63
+ ```js
64
+ import { preflight, preflightRoute } from 'slash-tokens'
67
65
 
68
- ## Dashboard
66
+ const check = preflight(prompt, 'claude-opus-5') // tokens, cost, fits, cheaper options (all providers)
67
+ const route = preflightRoute(prompt, 'claude-opus-5') // the cheapest same-provider model that fits, or null
68
+ ```
69
69
 
70
- Track savings across all your apps. One-person key (email, $20 on the house) at [mcpaas.live/slash/setup](https://mcpaas.live/slash/setup)
70
+ ## Pricing
71
71
 
72
- Full docs, examples, and model pricing at **[GitHub](https://github.com/Wolfe-Jam/slash-tokens)**
72
+ The library and CLI are free, no account needed. A one-person key is $20 on the house: we show the savings and don't charge. Team is $39/month for the data. [Live demo](https://slash-nextjs-wofejams-projects.vercel.app) · [Get a key](https://mcpaas.live/slash/setup) · [slashtokens.com](https://slashtokens.com) · [Full docs](https://github.com/Wolfe-Jam/slash-tokens)
73
73
 
74
74
  ## License
75
75
 
@@ -0,0 +1,50 @@
1
+ import { type Task, type Decision, type Message } from './quote.js';
2
+ import { type Receipt, type Usage } from './receipt.js';
3
+ import type { Tier } from './catalog.js';
4
+ export interface HireOptions {
5
+ /** Total USD the agent may spend across every job it runs. No budget = no limit. */
6
+ budget?: number;
7
+ /** Model the savings are measured against. Default: the model each job asks for. */
8
+ baseline?: string;
9
+ /** Lowest tier a substitute may have. Default: each job's requested tier. */
10
+ floor?: Tier;
11
+ /** The agent's cut of measured savings. Default 0.10. */
12
+ feeRate?: number;
13
+ /** Fee shown on the receipt but not charged. Default true. */
14
+ waived?: boolean;
15
+ }
16
+ /**
17
+ * What the agent calls to do the work: your provider client. `model` is the
18
+ * booked catalog key (e.g. `nemotron-3.5-lightning`); map it to your
19
+ * provider's API ID. Return the output and the provider's `usage` object.
20
+ */
21
+ export type CallModel = (model: string, input: string | Message[]) => Promise<{
22
+ output: string;
23
+ usage: Usage;
24
+ }>;
25
+ export interface Job extends Task {
26
+ call: CallModel;
27
+ }
28
+ export interface RunResult {
29
+ decision: Decision;
30
+ /** The model's output; undefined when blocked. */
31
+ output?: string;
32
+ /** Estimate vs actual, saved vs baseline, fee; undefined when blocked. */
33
+ receipt?: Receipt;
34
+ }
35
+ export interface Agent {
36
+ /** Quote → decide → book → run → reconcile → receipt. */
37
+ run(job: Job): Promise<RunResult>;
38
+ /** USD spent so far (actual cost, from receipts). */
39
+ readonly spent: number;
40
+ /** USD left in the budget (Infinity with no budget). */
41
+ readonly remaining: number;
42
+ readonly receipts: readonly Receipt[];
43
+ }
44
+ /**
45
+ * Hire Slash for a run of jobs. Each job is quoted, booked on the cheapest
46
+ * model the policy allows (or blocked), run through your `call`, and
47
+ * reconciled against the provider's own usage. The budget covers all jobs:
48
+ * each one may spend only what earlier jobs left.
49
+ */
50
+ export declare function hire(opts?: HireOptions): Agent;
package/dist/agent.js ADDED
@@ -0,0 +1,36 @@
1
+ import { decide } from './quote.js';
2
+ import { reconcile } from './receipt.js';
3
+ /**
4
+ * Hire Slash for a run of jobs. Each job is quoted, booked on the cheapest
5
+ * model the policy allows (or blocked), run through your `call`, and
6
+ * reconciled against the provider's own usage. The budget covers all jobs:
7
+ * each one may spend only what earlier jobs left.
8
+ */
9
+ export function hire(opts = {}) {
10
+ const receipts = [];
11
+ let spent = 0;
12
+ const budget = opts.budget ?? Infinity;
13
+ return {
14
+ get spent() { return spent; },
15
+ get remaining() { return Math.max(budget - spent, 0); },
16
+ get receipts() { return receipts; },
17
+ async run(job) {
18
+ const { call, ...task } = job;
19
+ const policy = { floor: opts.floor };
20
+ if (budget !== Infinity)
21
+ policy.budget = Math.max(budget - spent, 0);
22
+ const decision = decide(task, policy);
23
+ if (decision.action === 'block' || !decision.chosen)
24
+ return { decision };
25
+ const { output, usage } = await call(decision.chosen.model, task.input);
26
+ const receipt = reconcile(decision.chosen, usage, {
27
+ baseline: opts.baseline ?? decision.requested.model,
28
+ feeRate: opts.feeRate,
29
+ waived: opts.waived,
30
+ });
31
+ spent = Math.round((spent + receipt.actual.cost) * 1000000) / 1000000;
32
+ receipts.push(receipt);
33
+ return { decision, output, receipt };
34
+ },
35
+ };
36
+ }
@@ -0,0 +1,38 @@
1
+ /**
2
+ * The catalog — one source for every priced model.
3
+ *
4
+ * Each entry carries its prices (USD per million tokens), context window,
5
+ * provider, capability tier and the date the price was checked. `MODELS`
6
+ * (models.ts) and the quote / decide policy (quote.ts) both read from here.
7
+ *
8
+ * Tiers are the vendor's own line position, not a benchmark:
9
+ * 4 frontier — the vendor's top line above its flagship (Fable/Mythos, GPT-6 Astra)
10
+ * 3 flagship — Opus, Sol, Grok 4.5+, Gemini Pro, Nemotron Ultra, GPT-5.4
11
+ * 2 mid — Sonnet, Terra, Grok 4.3, Gemini Flash, Nemotron Super, GPT-5.4 mini
12
+ * 1 small — Haiku, Luna, Flash-Lite, GPT-5.4 nano, Nemotron Nano / Lightning
13
+ * A quality floor compares tiers within one provider only; tiers never rank
14
+ * one vendor's model against another's.
15
+ */
16
+ export type Tier = 1 | 2 | 3 | 4;
17
+ export declare const TIER_NAMES: Record<Tier, string>;
18
+ export interface CatalogEntry {
19
+ provider: string;
20
+ tier: Tier;
21
+ /** Date the price was checked against the provider's own pricing page. */
22
+ asOf: string;
23
+ /** Last day this price holds, when the provider has announced a change. */
24
+ priceUntil?: string;
25
+ input: number;
26
+ output: number;
27
+ context: number;
28
+ longContextThreshold?: number;
29
+ longContextInput?: number;
30
+ longContextOutput?: number;
31
+ }
32
+ export declare const CATALOG: Record<string, CatalogEntry>;
33
+ /**
34
+ * Real API IDs that don't follow the table's naming → table keys.
35
+ * Lowercased; canonicalModel() lowercases before looking here. The Nebius IDs
36
+ * are the ones Token Factory lists (GET /v1/models, 2026-10-08).
37
+ */
38
+ export declare const API_ALIASES: Record<string, string>;
@@ -0,0 +1,143 @@
1
+ /**
2
+ * The catalog — one source for every priced model.
3
+ *
4
+ * Each entry carries its prices (USD per million tokens), context window,
5
+ * provider, capability tier and the date the price was checked. `MODELS`
6
+ * (models.ts) and the quote / decide policy (quote.ts) both read from here.
7
+ *
8
+ * Tiers are the vendor's own line position, not a benchmark:
9
+ * 4 frontier — the vendor's top line above its flagship (Fable/Mythos, GPT-6 Astra)
10
+ * 3 flagship — Opus, Sol, Grok 4.5+, Gemini Pro, Nemotron Ultra, GPT-5.4
11
+ * 2 mid — Sonnet, Terra, Grok 4.3, Gemini Flash, Nemotron Super, GPT-5.4 mini
12
+ * 1 small — Haiku, Luna, Flash-Lite, GPT-5.4 nano, Nemotron Nano / Lightning
13
+ * A quality floor compares tiers within one provider only; tiers never rank
14
+ * one vendor's model against another's.
15
+ */
16
+ export const TIER_NAMES = {
17
+ 4: 'frontier',
18
+ 3: 'flagship',
19
+ 2: 'mid',
20
+ 1: 'small',
21
+ };
22
+ const OPUS = { input: 5.00, output: 25.00, context: 1000000 };
23
+ const OPUS_55 = { input: 4.00, output: 20.00, context: 1000000 };
24
+ const FABLE = { input: 10.00, output: 50.00, context: 1000000 };
25
+ const SONNET_4X = { input: 3.00, output: 15.00, context: 1000000 };
26
+ const SONNET = { input: 2.00, output: 10.00, context: 1000000 };
27
+ const HAIKU = { input: 1.00, output: 5.00, context: 200000 };
28
+ const GROK_46 = {
29
+ input: 2.00, output: 6.00, context: 500000,
30
+ longContextThreshold: 200000, longContextInput: 4.00, longContextOutput: 12.00,
31
+ };
32
+ const GROK_43 = {
33
+ input: 1.25, output: 2.50, context: 1000000,
34
+ longContextThreshold: 200000, longContextInput: 2.50, longContextOutput: 5.00,
35
+ };
36
+ const GEMINI_PRO = {
37
+ input: 2.00, output: 12.00, context: 1000000,
38
+ longContextThreshold: 200000, longContextInput: 4.00, longContextOutput: 18.00,
39
+ };
40
+ const GEMINI_FLASH = { input: 0.30, output: 2.50, context: 1000000 };
41
+ // Gemini 3.6–3.8 Flash: launch price through 2026-12-31; Google lists $1.50/$7.50
42
+ // from 2027-01-01. priceUntil makes the freshness check fail before then.
43
+ const GEMINI_FLASH_3X = { input: 0.75, output: 3.75, context: 1000000, priceUntil: '2026-12-31' };
44
+ const GEMINI_35_FLASH = { input: 1.50, output: 9.00, context: 1000000 };
45
+ const GEMINI_31_FLASH_LITE = { input: 0.25, output: 1.50, context: 1000000 };
46
+ const GROK_BUILD = {
47
+ input: 1.00, output: 2.00, context: 256000,
48
+ longContextThreshold: 200000, longContextInput: 2.00, longContextOutput: 4.00,
49
+ };
50
+ const GPT_6_ASTRA = { input: 10.00, output: 50.00, context: 1050000 };
51
+ const GPT_6_SOL = { input: 2.00, output: 10.00, context: 1050000 };
52
+ const GPT_6_LUNA = { input: 0.10, output: 0.50, context: 1050000 };
53
+ const GPT_SOL = { input: 4.00, output: 20.00, context: 1050000 };
54
+ const GPT_TERRA = { input: 2.00, output: 12.00, context: 1050000 };
55
+ const GPT_LUNA = { input: 0.20, output: 1.20, context: 1050000 };
56
+ const GPT_54 = { input: 2.50, output: 15.00, context: 1000000 };
57
+ const GPT_54_MINI = { input: 0.75, output: 4.50, context: 128000 };
58
+ const GPT_54_NANO = { input: 0.20, output: 1.25, context: 128000 };
59
+ // NVIDIA Nemotron on Nebius Token Factory. Prices and context windows are
60
+ // exactly what the API reports (GET /v1/models?verbose=true, 2026-10-08).
61
+ const NEMOTRON_ULTRA = { input: 1.00, output: 3.00, context: 1048576 };
62
+ const NEMOTRON_SUPER = { input: 0.30, output: 0.90, context: 262144 };
63
+ const NEMOTRON_NANO = { input: 0.06, output: 0.24, context: 262144 };
64
+ const NEMOTRON_LIGHTNING = { input: 0.06, output: 0.24, context: 1048576 };
65
+ // Prices as of 2026-10-07 — first-party pages:
66
+ // platform.claude.com/docs/en/about-claude/pricing
67
+ // developers.openai.com/api/docs/models
68
+ // docs.x.ai/developers/models
69
+ // ai.google.dev/gemini-api/docs/pricing
70
+ const ASOF = '2026-10-07';
71
+ // Nemotron: read from the Token Factory API (see above).
72
+ const ASOF_NEBIUS = '2026-10-08';
73
+ function e(provider, tier, price) {
74
+ return { provider, tier, asOf: provider === 'Nebius' ? ASOF_NEBIUS : ASOF, ...price };
75
+ }
76
+ // Order is kept from the 1.6.6 MODELS table (new entries appended), so
77
+ // preflight()'s option ordering for equal costs doesn't move.
78
+ export const CATALOG = {
79
+ // Anthropic — live names + generic aliases (same rates)
80
+ 'claude-fable-5.1': e('Anthropic', 4, FABLE),
81
+ 'claude-fable-5': e('Anthropic', 4, FABLE),
82
+ 'claude-mythos-5.1': e('Anthropic', 4, FABLE),
83
+ 'claude-mythos-5': e('Anthropic', 4, FABLE),
84
+ 'claude-opus-5.5': e('Anthropic', 3, OPUS_55),
85
+ 'claude-opus-5': e('Anthropic', 3, OPUS),
86
+ 'claude-opus-4.8': e('Anthropic', 3, OPUS),
87
+ 'claude-opus': e('Anthropic', 3, OPUS),
88
+ 'claude-opus-4.7': e('Anthropic', 3, OPUS),
89
+ 'claude-opus-4.6': e('Anthropic', 3, OPUS),
90
+ 'claude-opus-4.5': e('Anthropic', 3, OPUS),
91
+ 'claude-sonnet-5.5': e('Anthropic', 2, SONNET),
92
+ 'claude-sonnet-5': e('Anthropic', 2, SONNET),
93
+ 'claude-sonnet': e('Anthropic', 2, SONNET),
94
+ 'claude-sonnet-4.6': e('Anthropic', 2, SONNET_4X),
95
+ 'claude-sonnet-4.5': e('Anthropic', 2, SONNET_4X),
96
+ 'claude-haiku-4.5': e('Anthropic', 1, HAIKU),
97
+ 'claude-haiku': e('Anthropic', 1, HAIKU),
98
+ // xAI — flagship 4.6, cheap same-provider 4.3. 4.20 / fast are aliases.
99
+ 'grok-4.7': e('xAI', 3, GROK_46),
100
+ 'grok-4.6': e('xAI', 3, GROK_46),
101
+ 'grok-4.5': e('xAI', 3, GROK_46),
102
+ 'grok-build-0.1': e('xAI', 2, GROK_BUILD),
103
+ 'grok-4.3': e('xAI', 2, GROK_43),
104
+ 'grok-4.20': e('xAI', 2, GROK_43),
105
+ 'grok-4-1-fast': e('xAI', 2, GROK_43),
106
+ // Google
107
+ 'gemini-3.1-pro': e('Google', 3, GEMINI_PRO),
108
+ 'gemini-3.1-pro-preview': e('Google', 3, GEMINI_PRO),
109
+ 'gemini-3.8-flash': e('Google', 2, GEMINI_FLASH_3X),
110
+ 'gemini-3.7-flash': e('Google', 2, GEMINI_FLASH_3X),
111
+ 'gemini-3.6-flash': e('Google', 2, GEMINI_FLASH_3X),
112
+ 'gemini-3.5-flash': e('Google', 2, GEMINI_35_FLASH),
113
+ 'gemini-3.1-flash-lite': e('Google', 1, GEMINI_31_FLASH_LITE),
114
+ 'gemini-3.5-flash-lite': e('Google', 1, GEMINI_FLASH),
115
+ 'gemini-2.5-flash': e('Google', 2, GEMINI_FLASH),
116
+ // OpenAI — live 5.6 ladder + GPT-6. 5.4 family kept as aliases (old prices).
117
+ 'gpt-6-astra': e('OpenAI', 4, GPT_6_ASTRA),
118
+ 'gpt-6.1-sol': e('OpenAI', 3, GPT_6_SOL),
119
+ 'gpt-6-sol': e('OpenAI', 3, GPT_6_SOL),
120
+ 'gpt-6-luna': e('OpenAI', 1, GPT_6_LUNA),
121
+ 'gpt-5.6-sol': e('OpenAI', 3, GPT_SOL),
122
+ 'gpt-5.6-terra': e('OpenAI', 2, GPT_TERRA),
123
+ 'gpt-5.6-luna': e('OpenAI', 1, GPT_LUNA),
124
+ 'gpt-5.4': e('OpenAI', 3, GPT_54),
125
+ 'gpt-5.4-mini': e('OpenAI', 2, GPT_54_MINI),
126
+ 'gpt-5.4-nano': e('OpenAI', 1, GPT_54_NANO),
127
+ // NVIDIA Nemotron, served and billed by Nebius Token Factory
128
+ 'nemotron-3-ultra': e('Nebius', 3, NEMOTRON_ULTRA),
129
+ 'nemotron-3-super': e('Nebius', 2, NEMOTRON_SUPER),
130
+ 'nemotron-3-nano': e('Nebius', 1, NEMOTRON_NANO),
131
+ 'nemotron-3.5-lightning': e('Nebius', 1, NEMOTRON_LIGHTNING),
132
+ };
133
+ /**
134
+ * Real API IDs that don't follow the table's naming → table keys.
135
+ * Lowercased; canonicalModel() lowercases before looking here. The Nebius IDs
136
+ * are the ones Token Factory lists (GET /v1/models, 2026-10-08).
137
+ */
138
+ export const API_ALIASES = {
139
+ 'nvidia/nemotron-3-ultra-550b-a55b': 'nemotron-3-ultra',
140
+ 'nvidia/nemotron-3-super-120b-a12b': 'nemotron-3-super',
141
+ 'nvidia/nemotron-3_5-lightning': 'nemotron-3.5-lightning',
142
+ 'nvidia/nvidia-nemotron-3-nano-30b-a3b': 'nemotron-3-nano',
143
+ };