@gullabs/xai 0.4.1 → 0.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -14,26 +14,28 @@ xAI has no first-party TypeScript SDK. xAI's own quickstart recommends using the
14
14
 
15
15
  ## Key exports
16
16
 
17
- | Export | What it is |
18
- | ----------------------- | ----------------------------------------------------------------------------------------------- |
19
- | `xaiProvider(opts?)` | `ProviderPlugin` factory — bundles the adapter, `grok-4.5` model descriptor, and pricing source |
20
- | `xaiAdapter(opts?)` | Creates the `ProviderAdapter` for xAI |
21
- | `XaiAdapterOptions` | `{ client?: XaiClientLike }` — inject a pre-built or fake client |
22
- | `XaiClientLike` | Structural interface the adapter depends on (satisfied by real SDK and fakes) |
23
- | `buildXaiClient(auth)` | Builds the real `openai`-SDK-backed client from `AuthMaterial`, pointed at xAI's base URL |
24
- | `classifyXaiError(err)` | Classifies a raw thrown error into a typed `LlmError`, including xAI's 400-for-auth quirk |
25
- | `grok45ModelDescriptor` | The `grok-4.5` `ModelDescriptor` |
26
- | `xaiModelDescriptors` | Every model descriptor this package contributes (v1: just `grok45ModelDescriptor`) |
27
- | `xaiRegistry` | Pre-built `ModelRegistry` over `xaiModelDescriptors` |
28
- | `xaiPricingSource()` | Built-in xAI `PricingSource` port implementation, backed by `XAI_PRICING` |
29
- | `XAI_PRICING` | Frozen xAI pricing snapshot (µUSD per million tokens) |
30
- | `XaiModelRates` | Per-model rate entry type (`inputPerM`, `cachedPerM`, `outputPerM`, optional `gt200k`) |
31
- | `Grok45ConfigSchema` | Strict Zod config schema for `grok-4.5` |
32
- | `XaiProviderOptions` | `{ promptCacheKey? }` — typed `providerOptions.xai` extension shape |
33
- | `XaiFileStore` | Files API store: upload (TTL), get, list, idempotent delete, content |
34
- | `XaiFileHandle` | `{ id, filename?, bytes?, expiresAt?, … }` returned by the store |
35
- | `FileDeleteOptions` | `{ failClosed?, signal? }` — opt-in fail-closed delete for durable release gates |
36
- | `XAI_FILE_TTL_*` | TTL bounds (`3600`…`2592000` seconds) and `XAI_FILE_MAX_BYTES` (48 MiB) |
17
+ | Export | What it is |
18
+ | ----------------------- | ------------------------------------------------------------------------------------------------------- |
19
+ | `xaiProvider(opts?)` | `ProviderPlugin` factory — bundles the adapter, `grok-4.5` / `grok-4.6` descriptors, and pricing source |
20
+ | `xaiAdapter(opts?)` | Creates the `ProviderAdapter` for xAI |
21
+ | `XaiAdapterOptions` | `{ client?: XaiClientLike }` — inject a pre-built or fake client |
22
+ | `XaiClientLike` | Structural interface the adapter depends on (satisfied by real SDK and fakes) |
23
+ | `buildXaiClient(auth)` | Builds the real `openai`-SDK-backed client from `AuthMaterial`, pointed at xAI's base URL |
24
+ | `classifyXaiError(err)` | Classifies a raw thrown error into a typed `LlmError`, including xAI's 400-for-auth quirk |
25
+ | `grok45ModelDescriptor` | The `grok-4.5` `ModelDescriptor` |
26
+ | `grok46ModelDescriptor` | The `grok-4.6` `ModelDescriptor` |
27
+ | `xaiModelDescriptors` | Every model descriptor this package contributes (`grok-4.5`, `grok-4.6`) |
28
+ | `xaiRegistry` | Pre-built `ModelRegistry` over `xaiModelDescriptors` |
29
+ | `xaiPricingSource()` | Built-in xAI `PricingSource` port implementation, backed by `XAI_PRICING` |
30
+ | `XAI_PRICING` | Frozen xAI pricing snapshot (µUSD per million tokens) |
31
+ | `XaiModelRates` | Per-model rate entry type (`inputPerM`, `cachedPerM`, `outputPerM`, optional `gt200k`) |
32
+ | `Grok45ConfigSchema` | Strict Zod config schema for `grok-4.5` |
33
+ | `Grok46ConfigSchema` | Strict Zod config schema for `grok-4.6` |
34
+ | `XaiProviderOptions` | `{ promptCacheKey? }` — typed `providerOptions.xai` extension shape |
35
+ | `XaiFileStore` | Files API store: upload (TTL), get, list, idempotent delete, content |
36
+ | `XaiFileHandle` | `{ id, filename?, bytes?, expiresAt?, … }` returned by the store |
37
+ | `FileDeleteOptions` | `{ failClosed?, signal? }` — opt-in fail-closed delete for durable release gates |
38
+ | `XAI_FILE_TTL_*` | TTL bounds (`3600`…`2592000` seconds) and `XAI_FILE_MAX_BYTES` (48 MiB) |
37
39
 
38
40
  ## Quick example
39
41
 
@@ -49,23 +51,25 @@ const client = createClient({
49
51
  const result = await client.generate(
50
52
  {
51
53
  provider: 'xai',
52
- model: 'grok-4.5',
54
+ model: 'grok-4.6',
53
55
  messages: [{ role: 'user', parts: [{ kind: 'text', text: 'Hello' }] }],
54
56
  },
55
57
  { auth: { apiKey: 'YOUR_XAI_API_KEY' } },
56
58
  )
57
59
  ```
58
60
 
59
- ## grok-4.5
61
+ ## grok-4.5 and grok-4.6
60
62
 
61
- The default registry ships exactly one model: `grok-4.5` (500k token context window). It routes through this adapter and supports:
63
+ The default registry ships two canonical models (500k token context window each). They route through this adapter and support:
62
64
 
63
- - **Reasoning** — level-api (`reasoningApi: 'level'`), mapped to the Responses API `reasoning.effort` field. `admittedReasoningEfforts: ['low', 'high']` — this is deliberately narrower than Gemini's effort vocabulary: `'none'` and `'medium'` are **rejected by the live xAI API**, not merely unsupported by this library. The config schema (`Grok45ConfigSchema`) enforces `reasoning.effort` as `'low' | 'high'` only, and there is no `budgetTokens` field at all (xAI uses level-style reasoning, not token budgets) — passing it throws `bad_request`. The schema does not set a default effort; if `reasoning` is omitted, no `reasoning` field is sent and xAI's own server-side default applies.
65
+ - **Reasoning** — level-api (`reasoningApi: 'level'`), mapped to the Responses API `reasoning.effort` field. There is no `budgetTokens` field (xAI uses level-style reasoning) — passing it throws `bad_request`. The schema does not set a default effort; if `reasoning` is omitted, no `reasoning` field is sent and xAI's own server-side default (`high`) applies.
66
+ - `grok-4.5`: `admittedReasoningEfforts: ['low', 'high']`. `'none'`, `'medium'`, and `'xhigh'` are rejected by `Grok45ConfigSchema` (frozen 2026-07-09 live probe; do not silently widen).
67
+ - `grok-4.6`: `admittedReasoningEfforts: ['low', 'medium', 'high', 'xhigh']` (live-verified 2026-08-12). `'none'` is rejected by the live API.
64
68
  - **Structured output** — native. `output.jsonSchema` maps to the Responses API's `text.format` field with `{ type: 'json_schema', name, schema, strict: true }`, **not** `response_format` — this differs from OpenAI's own convention for the same underlying concept.
65
69
  - **`strict: true` performs no OpenAI-style compile-time schema validation, as of the 2026-07-09 live probes.** 2026-07-09 live verification against the real xAI Responses API — 13 single-variant probes plus 1 combined probe (14 calls total, all accepted HTTP 200; the combined probe is recorded as fixture `10-non-strict-schema-accepted.json`) — verified that `text.format` with `strict: true` accepted every one of the following schema shapes that OpenAI's own strict mode rejects at compile time: schemas (root and nested) missing `additionalProperties: false`; properties omitted from `required` (optional properties); `format`, `minLength`, `pattern`, and `default` keywords; `anyOf`; `$defs`/`$ref`; `enum`/`const`; and nullable unions (`type: [T, 'null']`). `strict: false` on the same surface showed no observed behavioral divergence from `strict: true`. This adapter forwards schemas to xAI verbatim — no rewriting, no preflight validation, and no injection of `additionalProperties: false` or `required` completion — so OpenAI-strict schema rewriting (including `@gullabs/codex-cli`'s `toOpenAiStrictOutputSchema` helper) is unnecessary for xai as of that verification date. (Reject-don't-map still applies to genuinely invalid input the xai schema/types layer itself rejects; this note is only about strict-mode compile-time schema-shape enforcement.) `packages/xai/src/__fixtures__/10-non-strict-schema-accepted.json` records one live example combining three of these — missing root `additionalProperties: false`, an optional property, and a `format` keyword — in a single accepted call.
66
70
  - **Sampling** — `temperature` and `topP` are forwarded verbatim. No `topK`.
67
71
  - **No penalties/stop** — `presence_penalty`, `frequency_penalty`, and `stop` are not in the config schema at all; xAI hard-rejects these on reasoning models, so the schema never admits them (reject-don't-map).
68
- - **No service tiers** — xAI has no service-tier concept for `grok-4.5`; setting `serviceTier` on the request throws `bad_request`.
72
+ - **Service tiers** — `grok-4.5` admits none; setting `serviceTier` throws `bad_request`. `grok-4.6` admits `serviceTier: 'priority'` only (Responses `service_tier: "priority"`, live-verified 2026-08-12). `'flex'` / `'standard'` / `'batch'` are rejected — xAI silently remaps unknown tiers to `default`, so this library never forwards them.
69
73
 
70
74
  ## Files store (`XaiFileStore`)
71
75
 
@@ -118,7 +122,7 @@ try {
118
122
 
119
123
  ## Vision constraints
120
124
 
121
- `grok-4.5` accepts image input as an `inline-media` or `file-uri` `Part`, and document attachments as a `file-ref` `Part`:
125
+ Both models accept image input as an `inline-media` or `file-uri` `Part`, and document attachments as a `file-ref` `Part`:
122
126
 
123
127
  - **`inline-media`** — only `image/jpeg` and `image/png` are accepted; anything else throws `bad_request`. The decoded payload must be at most 20 MiB (xAI's documented inline-image ceiling); larger images throw `bad_request` before the request is sent.
124
128
  - **`file-uri`** — only accepted when the URI is a public `http(s)://` URL **and** the declared `mimeType` is jpg/png. A provider-hosted URI from another provider — for example a Gemini Files API URI (`https://generativelanguage.googleapis.com/...`) — is technically `https://` but is not dereferenceable by xAI and is not portable across providers. The adapter rejects it rather than trying to map or proxy it (reject-don't-map).
@@ -131,14 +135,16 @@ xAI caching is automatic — there is no explicit cache-create/cache-store API c
131
135
 
132
136
  ## Pricing
133
137
 
134
- `XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-07-09'`) — a point-in-time capture, not a live lookup, following the same frozen-snapshot precedent as Gemini pricing (ADR-005). Rates are in µUSD per million tokens:
138
+ `XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-08-12'`) — a point-in-time capture from `/v1/models`, not a live lookup (ADR-005). Rates are in µUSD per million tokens:
135
139
 
136
140
  | Model | Tier | Input | Cached input | Output |
137
141
  | ---------- | ---------------------------- | ------- | ------------ | -------- |
138
- | `grok-4.5` | standard (≤200k gross input) | $2.00/M | $0.50/M | $6.00/M |
139
- | `grok-4.5` | `gt200k` (>200k gross input) | $4.00/M | $1.00/M | $12.00/M |
142
+ | `grok-4.5` | standard (≤200k gross input) | $2.00/M | $0.30/M | $6.00/M |
143
+ | `grok-4.5` | `gt200k` (>200k gross input) | $4.00/M | $0.60/M | $12.00/M |
144
+ | `grok-4.6` | standard (≤200k gross input) | $2.00/M | $0.50/M | $6.00/M |
145
+ | `grok-4.6` | `gt200k` (>200k gross input) | $4.00/M | $1.00/M | $12.00/M |
140
146
 
141
- The `gt200k` long-context tier is selected by **gross** `inputTokens` (including cached), not billable input — strictly greater than 200,000 tokens, mirroring core's `selectRates` convention. `grok-4.5` has no service-tier concept; a defined `tier` argument to `xaiPricingSource().price()` always resolves to unpriced (`microUsd: null`) rather than guessing a multiplier.
147
+ The `gt200k` long-context tier is selected by **gross** `inputTokens` (including cached), not billable input — strictly greater than 200,000 tokens, mirroring core's `selectRates` convention. The adapter now surfaces the echoed Responses `service_tier` (`'default'` or `'priority'`), so `price()` receives that served value instead of `undefined`. Custom xAI `PricingSource` implementations must price `'default'` at the standard list. Built-in `xaiPricingSource().price()` prices `grok-4.6` + `tier: 'priority'` at 2× every token type after the cache discount: uncached standard-list 2× is confirmed by fixture `12-grok-4-6-xhigh-priority.json` `cost_in_usd_ticks`; cached and `gt200k` legs follow the official 2×-after-cache-discount rule. Any other defined tier (including `priority` on `grok-4.5`) is unpriced (`microUsd: null`). Standard list rates are pinned to `packages/xai/src/__fixtures__/14-v1-models-pricing.json` (live `GET /v1/models` 2026-08-12).
142
148
 
143
149
  ## EU unavailability
144
150
 
@@ -146,7 +152,7 @@ xAI has no EU region at launch — its documented regions are `us-east-1` and `u
146
152
 
147
153
  ## Aliases are not registered
148
154
 
149
- xAI's own `/v1/models` listing surfaces `grok-4.5-latest` and `grok-build-latest` as aliases of `grok-4.5`. Neither is registered as a separate `ModelDescriptor` or `XAI_PRICING` key. Callers must use the canonical id `grok-4.5` verbatim — passing an alias resolves to "model not found" in the registry, not to `grok-4.5`'s config or pricing (reject-don't-map).
155
+ xAI's own `/v1/models` listing surfaces `grok-4.5-latest` and `grok-build-latest` as aliases of `grok-4.5`. `grok-4.6` has no aliases as of 2026-08-12. Aliases are not registered as `ModelDescriptor`s or `XAI_PRICING` keys. Callers must use the canonical id verbatim — passing an alias resolves to "model not found" (reject-don't-map).
150
156
 
151
157
  ## Explicitly deferred (not built in v1)
152
158
 
@@ -159,10 +165,11 @@ xAI's own `/v1/models` listing surfaces `grok-4.5-latest` and `grok-build-latest
159
165
  ## What it maps
160
166
 
161
167
  - `providerOptions.xai.promptCacheKey` → `prompt_cache_key`
162
- - `reasoning.effort` → `reasoning.effort` (`'low' | 'high'` only)
168
+ - `reasoning.effort` → `reasoning.effort` (per-model admitted set)
169
+ - `serviceTier: 'priority'` → `service_tier: 'priority'` (`grok-4.6` only)
163
170
  - `output.jsonSchema` → `text.format: { type: 'json_schema', name, schema, strict: true }`
164
171
  - Usage: `usage.input_tokens` → `inputTokens`, `usage.output_tokens` → `outputTokens` (both already GROSS on xAI, unlike Gemini's sub-field summation); numeric extras (`num_sources_used`, `cost_in_usd_ticks`, etc.) surface into `usage.details` under their raw names, and the full raw payload is always in `usage.raw`
165
- - Errors: xAI's Responses API returns **HTTP 400 (not 401) for an invalid API key**. `classifyXaiError` special-cases the exact structured error-body signature (`code: 'invalid-argument'` with message prefix `"Incorrect API key provided"`, taken verbatim from a recorded live fixture) and reclassifies it as `invalid_auth`. It only inspects the STRUCTURED parsed error body — never free-form `Error.message` text — so a 400 that merely _mentions_ an API key (e.g. a schema-validation error echoing user content) stays `bad_request`. Any other 400, `429`→`rate_limited`, `5xx`→`server`, and timeouts fall through to `@gullabs/core`'s generic `classifyError`.
172
+ - Errors: HTTP status is a hint. `classifyXaiError` inspects the STRUCTURED parsed body only — never free-form `Error.message`. Two recorded overlays: HTTP **400** whose body starts with `"Incorrect API key provided"` (prefix only; the SDK may drop `code`) → `invalid_auth`; HTTP **403** whose body starts with `"Content violates usage guidelines"` (e.g. `SAFETY_CHECK_TYPE_*`) → `content_filter`. A bare 403 without that body stays `invalid_auth`. Any other 400, `429`→`rate_limited`, `5xx`→`server`, and timeouts fall through to `@gullabs/core`'s generic `classifyError`.
166
173
 
167
174
  ## Learn more
168
175
 
package/dist/index.cjs CHANGED
@@ -173,6 +173,11 @@ function isXaiAuthFailureBody(rawErr) {
173
173
  const text = extractXaiErrorBodyText(rawErr);
174
174
  return text !== void 0 && text.startsWith(XAI_AUTH_ERROR_MESSAGE_PREFIX);
175
175
  }
176
+ var XAI_SAFETY_CHECK_MESSAGE_PREFIX = "Content violates usage guidelines";
177
+ function isXaiSafetyCheckBody(rawErr) {
178
+ const text = extractXaiErrorBodyText(rawErr);
179
+ return text !== void 0 && text.startsWith(XAI_SAFETY_CHECK_MESSAGE_PREFIX);
180
+ }
176
181
  var XAI_TRANSPORT_ERROR_PATTERN = /connection error|econnreset|econnrefused|etimedout|eai_again|epipe|socket hang up|fetch failed/i;
177
182
  function matchesXaiTransportSignature(err) {
178
183
  if (!(err instanceof Error)) return false;
@@ -211,6 +216,16 @@ function classifyXaiError(rawErr) {
211
216
  cause: base.cause ?? rawErr
212
217
  });
213
218
  }
219
+ if (base.httpStatus === 403 && isXaiSafetyCheckBody(rawErr)) {
220
+ const bodyText = extractXaiErrorBodyText(rawErr);
221
+ return new core.LlmError(bodyText ?? base.message, {
222
+ kind: "content_filter",
223
+ retryable: false,
224
+ httpStatus: base.httpStatus,
225
+ provider: "xai",
226
+ cause: base.cause ?? rawErr
227
+ });
228
+ }
214
229
  if (base.kind === "unknown" && isXaiTransportError(rawErr)) {
215
230
  return new core.LlmError(base.message, {
216
231
  kind: "server",
@@ -262,10 +277,19 @@ function xaiAdapter(opts) {
262
277
  if (genConfig.maxOutputTokens !== void 0) {
263
278
  params.max_output_tokens = genConfig.maxOutputTokens;
264
279
  }
280
+ const admittedTiers = req.modelDescriptor?.capabilities?.serviceTiers;
265
281
  if (genConfig.serviceTier !== void 0) {
266
- throw badXaiRequest(
267
- `serviceTier is not supported for xai models (got "${genConfig.serviceTier}").`
268
- );
282
+ if (admittedTiers === void 0 || !admittedTiers.includes(genConfig.serviceTier)) {
283
+ throw badXaiRequest(
284
+ `serviceTier is not supported for xai model "${model}" (got "${genConfig.serviceTier}").`
285
+ );
286
+ }
287
+ if (genConfig.serviceTier !== "priority") {
288
+ throw badXaiRequest(
289
+ `serviceTier "${genConfig.serviceTier}" is not supported for xai model "${model}" (only "priority" is admitted).`
290
+ );
291
+ }
292
+ params.service_tier = "priority";
269
293
  }
270
294
  const reasoning = genConfig.reasoning;
271
295
  if (reasoning !== void 0) {
@@ -276,13 +300,13 @@ function xaiAdapter(opts) {
276
300
  }
277
301
  if (reasoning.effort !== void 0) {
278
302
  const effort = reasoning.effort;
279
- if (effort !== "low" && effort !== "high") {
303
+ if (effort === "none") {
280
304
  throw badXaiRequest(
281
- `reasoning.effort "${effort}" is not supported for xai model "${model}" (only "low" and "high" are admitted).`
305
+ `reasoning.effort "none" is not supported for xai model "${model}".`
282
306
  );
283
307
  }
284
308
  const admitted = req.modelDescriptor?.capabilities?.admittedReasoningEfforts;
285
- if (admitted !== void 0 && !admitted.includes(effort)) {
309
+ if (admitted === void 0 || !admitted.includes(effort)) {
286
310
  throw badXaiRequest(
287
311
  `reasoning.effort "${effort}" is not supported for xai model "${model}".`
288
312
  );
@@ -360,6 +384,7 @@ function xaiAdapter(opts) {
360
384
  if (isPlainRecord(response.metadata)) {
361
385
  providerMeta["metadata"] = response.metadata;
362
386
  }
387
+ const servedServiceTier = typeof response.service_tier === "string" && response.service_tier.length > 0 ? response.service_tier : void 0;
363
388
  const result = {
364
389
  model: response.model,
365
390
  usage,
@@ -369,6 +394,7 @@ function xaiAdapter(opts) {
369
394
  ...text.length > 0 ? { text } : {},
370
395
  ...reasoningText !== void 0 ? { reasoningText } : {},
371
396
  ...rawStructured !== void 0 ? { rawStructured } : {},
397
+ ...servedServiceTier !== void 0 ? { servedServiceTier } : {},
372
398
  ...Object.keys(providerMeta).length > 0 ? { providerMetadata: providerMeta } : {}
373
399
  };
374
400
  return result;
@@ -875,6 +901,55 @@ var Grok45ConfigSchema = zod.z.strictObject({
875
901
  description: "Strict Responses API config for model grok-4.5. Level reasoning (low/high only), tunable sampling, no service tiers, structured output, vision, priced.",
876
902
  examples: [{ reasoning: { effort: "high" } }]
877
903
  });
904
+ var Grok46ConfigSchema = zod.z.strictObject({
905
+ temperature: zod.z.number().optional().meta({
906
+ title: "Temperature",
907
+ description: "Sampling temperature forwarded verbatim to grok-4.6."
908
+ }),
909
+ topP: zod.z.number().optional().meta({
910
+ title: "Top P",
911
+ description: "Nucleus sampling parameter forwarded verbatim to grok-4.6."
912
+ }),
913
+ maxOutputTokens: zod.z.number().int().positive().optional().meta({
914
+ title: "Max Output Tokens",
915
+ description: "Maximum output token cap for grok-4.6. No artificial ceiling \u2014 xAI accepts arbitrarily large values; truncation surfaces as finishReason:'length', not an error."
916
+ }),
917
+ reasoning: zod.z.strictObject({
918
+ effort: zod.z.enum(["low", "medium", "high", "xhigh"]).meta({
919
+ title: "Reasoning Effort",
920
+ description: 'Reasoning effort for grok-4.6. Live-verified 2026-08-12: "low", "medium", "high", and "xhigh" are accepted; "none" is rejected. Vendor default when omitted is "high".'
921
+ })
922
+ }).optional().meta({
923
+ title: "Reasoning",
924
+ description: "grok-4.6 effort-level reasoning configuration. No budgetTokens field \u2014 xAI uses level-style reasoning, not token budgets."
925
+ }),
926
+ serviceTier: zod.z.literal("priority").optional().meta({
927
+ title: "Service Tier",
928
+ description: 'xAI priority processing for grok-4.6 (Responses `service_tier: "priority"`). Echo live-verified 2026-08-12. Bills at 2\xD7 after the cache discount (uncached standard-list 2\xD7 confirmed by live ticks; cached/long-context legs follow the official 2\xD7 rule). Omitted requests stay on xAI default. "flex"/"standard"/"batch" are rejected.'
929
+ }),
930
+ timeoutMs: zod.z.number().int().positive().optional().meta({
931
+ title: "Timeout",
932
+ description: "Logical request timeout in milliseconds."
933
+ }),
934
+ providerOptions: zod.z.strictObject({
935
+ xai: zod.z.strictObject({
936
+ promptCacheKey: zod.z.string().min(1).optional().meta({
937
+ title: "Prompt Cache Key",
938
+ description: "xAI conversation-routing cache key \u2014 maps to Responses API `prompt_cache_key`."
939
+ })
940
+ }).optional().meta({
941
+ title: "xAI Provider Options",
942
+ description: "Allowlisted xAI provider options for grok-4.6."
943
+ })
944
+ }).optional().meta({
945
+ title: "Provider Options",
946
+ description: "Provider-specific options accepted for grok-4.6."
947
+ })
948
+ }).meta({
949
+ title: "Grok46Config",
950
+ description: "Strict Responses API config for model grok-4.6. Level reasoning (low/medium/high/xhigh), optional priority service tier, tunable sampling, structured output, vision, priced.",
951
+ examples: [{ reasoning: { effort: "high" } }]
952
+ });
878
953
 
879
954
  // src/models.ts
880
955
  var grok45ModelDescriptor = {
@@ -892,20 +967,55 @@ var grok45ModelDescriptor = {
892
967
  sampling: "tunable",
893
968
  caching: { explicit: false, minTokens: 0 },
894
969
  grounding: false
895
- // No serviceTiers key — xai has no service-tier concept.
970
+ // No serviceTiers key — grok-4.5 has no admitted service-tier vocabulary.
896
971
  },
897
972
  configSchema: Grok45ConfigSchema,
898
973
  configJsonSchema: core.toConfigJsonSchema(Grok45ConfigSchema),
899
974
  validateConfig: core.zodToStandardSchema(Grok45ConfigSchema)
900
975
  };
901
- var xaiModelDescriptors = [grok45ModelDescriptor];
976
+ var grok46ModelDescriptor = {
977
+ model: "grok-4.6",
978
+ provider: "xai",
979
+ pricingFamily: "grok-4.6",
980
+ capabilities: {
981
+ reasoning: true,
982
+ reasoningApi: "level",
983
+ admittedReasoningEfforts: ["low", "medium", "high", "xhigh"],
984
+ structuredOutput: true,
985
+ nativeStructuredOutput: true,
986
+ vision: true,
987
+ audioInput: false,
988
+ sampling: "tunable",
989
+ caching: { explicit: false, minTokens: 0 },
990
+ grounding: false,
991
+ serviceTiers: ["priority"]
992
+ },
993
+ configSchema: Grok46ConfigSchema,
994
+ configJsonSchema: core.toConfigJsonSchema(Grok46ConfigSchema),
995
+ validateConfig: core.zodToStandardSchema(Grok46ConfigSchema)
996
+ };
997
+ var xaiModelDescriptors = [
998
+ grok45ModelDescriptor,
999
+ grok46ModelDescriptor
1000
+ ];
902
1001
  var xaiRegistry = core.createModelRegistry(xaiModelDescriptors);
903
1002
 
904
1003
  // src/pricing.ts
905
- var xaiPricingVersion = "xai-2026-07-09";
1004
+ var xaiPricingVersion = "xai-2026-08-12";
906
1005
  var XAI_PRICING = Object.freeze({
907
- // ── grok-4.5 ── $2.00/$6.00 (≤200k), $4.00/$12.00 (>200k); cached $0.50/$1.00
1006
+ // ── grok-4.5 ── $2.00/$6.00 (≤200k), $4.00/$12.00 (>200k); cached $0.30/$0.60
908
1007
  "grok-4.5": {
1008
+ inputPerM: 2e6,
1009
+ cachedPerM: 3e5,
1010
+ outputPerM: 6e6,
1011
+ gt200k: {
1012
+ inputPerM: 4e6,
1013
+ cachedPerM: 6e5,
1014
+ outputPerM: 12e6
1015
+ }
1016
+ },
1017
+ // ── grok-4.6 ── $2.00/$6.00 (≤200k), $4.00/$12.00 (>200k); cached $0.50/$1.00
1018
+ "grok-4.6": {
909
1019
  inputPerM: 2e6,
910
1020
  cachedPerM: 5e5,
911
1021
  outputPerM: 6e6,
@@ -913,7 +1023,9 @@ var XAI_PRICING = Object.freeze({
913
1023
  inputPerM: 4e6,
914
1024
  cachedPerM: 1e6,
915
1025
  outputPerM: 12e6
916
- }
1026
+ },
1027
+ // Confirmed 2026-08-12 by fixture 12 cost_in_usd_ticks (2× list).
1028
+ priorityFactor: 2
917
1029
  }
918
1030
  });
919
1031
  var LONG_CONTEXT_THRESHOLD = 2e5;
@@ -942,22 +1054,29 @@ function computeXaiCost(model, usage, tier) {
942
1054
  unpricedReason: `Unknown model "${model}"; no pricing entry found.`
943
1055
  };
944
1056
  }
945
- if (tier !== void 0) {
946
- return {
947
- microUsd: null,
948
- usd: null,
949
- pricingVersion: xaiPricingVersion,
950
- confidence: "estimated",
951
- details: { input: 0, cached: 0, output: 0 },
952
- unpricedReason: `Unknown service tier "${tier}"; xai has no service tiers, refusing to guess a pricing multiplier.`
953
- };
1057
+ let factor = 1;
1058
+ if (tier !== void 0 && tier !== "default") {
1059
+ if (tier === "priority" && rates.priorityFactor !== void 0) {
1060
+ factor = rates.priorityFactor;
1061
+ } else {
1062
+ return {
1063
+ microUsd: null,
1064
+ usd: null,
1065
+ pricingVersion: xaiPricingVersion,
1066
+ confidence: "estimated",
1067
+ details: { input: 0, cached: 0, output: 0 },
1068
+ unpricedReason: `Unknown service tier "${tier}"; xai model "${model}" has no such tier, refusing to guess a pricing multiplier.`
1069
+ };
1070
+ }
954
1071
  }
955
1072
  const base = selectRates(rates, usage.inputTokens);
956
1073
  const cached = usage.cachedInputTokens ?? 0;
957
1074
  const billableInput = Math.max(0, usage.inputTokens - cached);
958
- const inputCost = Math.round(billableInput * base.inputPerM / 1e6);
959
- const cachedCost = Math.round(cached * base.cachedPerM / 1e6);
960
- const outputCost = Math.round(usage.outputTokens * base.outputPerM / 1e6);
1075
+ const inputCost = Math.round(billableInput * base.inputPerM * factor / 1e6);
1076
+ const cachedCost = Math.round(cached * base.cachedPerM * factor / 1e6);
1077
+ const outputCost = Math.round(
1078
+ usage.outputTokens * base.outputPerM * factor / 1e6
1079
+ );
961
1080
  const microUsd = inputCost + cachedCost + outputCost;
962
1081
  return {
963
1082
  microUsd,
@@ -996,6 +1115,7 @@ function xaiProvider(opts) {
996
1115
  }
997
1116
 
998
1117
  exports.Grok45ConfigSchema = Grok45ConfigSchema;
1118
+ exports.Grok46ConfigSchema = Grok46ConfigSchema;
999
1119
  exports.XAI_FILES_DEFAULT_BASE_URL = XAI_FILES_DEFAULT_BASE_URL;
1000
1120
  exports.XAI_FILE_MAX_BYTES = XAI_FILE_MAX_BYTES;
1001
1121
  exports.XAI_FILE_TTL_MAX_SECONDS = XAI_FILE_TTL_MAX_SECONDS;
@@ -1006,6 +1126,7 @@ exports.buildXaiClient = buildXaiClient;
1006
1126
  exports.classifyXaiError = classifyXaiError;
1007
1127
  exports.computeXaiCost = computeXaiCost;
1008
1128
  exports.grok45ModelDescriptor = grok45ModelDescriptor;
1129
+ exports.grok46ModelDescriptor = grok46ModelDescriptor;
1009
1130
  exports.requireApiKey = requireApiKey;
1010
1131
  exports.xaiAdapter = xaiAdapter;
1011
1132
  exports.xaiModelDescriptors = xaiModelDescriptors;