@gullabs/xai 0.7.1 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -14,28 +14,30 @@ xAI has no first-party TypeScript SDK. xAI's own quickstart recommends using the
14
14
 
15
15
  ## Key exports
16
16
 
17
- | Export | What it is |
18
- | ----------------------- | ------------------------------------------------------------------------------------------------------- |
19
- | `xaiProvider(opts?)` | `ProviderPlugin` factory — bundles the adapter, `grok-4.5` / `grok-4.6` descriptors, and pricing source |
20
- | `xaiAdapter(opts?)` | Creates the `ProviderAdapter` for xAI |
21
- | `XaiAdapterOptions` | `{ client?: XaiClientLike }` — inject a pre-built or fake client |
22
- | `XaiClientLike` | Structural interface the adapter depends on (satisfied by real SDK and fakes) |
23
- | `buildXaiClient(auth)` | Builds the real `openai`-SDK-backed client from `AuthMaterial`, pointed at xAI's base URL |
24
- | `classifyXaiError(err)` | Classifies a raw thrown error into a typed `LlmError`, including xAI's 400-for-auth quirk |
25
- | `grok45ModelDescriptor` | The `grok-4.5` `ModelDescriptor` |
26
- | `grok46ModelDescriptor` | The `grok-4.6` `ModelDescriptor` |
27
- | `xaiModelDescriptors` | Every model descriptor this package contributes (`grok-4.5`, `grok-4.6`) |
28
- | `xaiRegistry` | Pre-built `ModelRegistry` over `xaiModelDescriptors` |
29
- | `xaiPricingSource()` | Built-in xAI `PricingSource` port implementation, backed by `XAI_PRICING` |
30
- | `XAI_PRICING` | Frozen xAI pricing snapshot (µUSD per million tokens) |
31
- | `XaiModelRates` | Per-model rate entry type (`inputPerM`, `cachedPerM`, `outputPerM`, optional `gt200k`) |
32
- | `Grok45ConfigSchema` | Strict Zod config schema for `grok-4.5` |
33
- | `Grok46ConfigSchema` | Strict Zod config schema for `grok-4.6` |
34
- | `XaiProviderOptions` | `{ promptCacheKey? }` — typed `providerOptions.xai` extension shape |
35
- | `XaiFileStore` | Files API store: upload (TTL), get, list, idempotent delete, content |
36
- | `XaiFileHandle` | `{ id, filename?, bytes?, expiresAt?, … }` returned by the store |
37
- | `FileDeleteOptions` | `{ failClosed?, signal? }` — opt-in fail-closed delete for durable release gates |
38
- | `XAI_FILE_TTL_*` | TTL bounds (`3600`…`2592000` seconds) and `XAI_FILE_MAX_BYTES` (48 MiB) |
17
+ | Export | What it is |
18
+ | ----------------------- | -------------------------------------------------------------------------------------------------------------------- |
19
+ | `xaiProvider(opts?)` | `ProviderPlugin` factory — bundles the adapter, `grok-4.5` / `grok-4.6` / `grok-4.7` descriptors, and pricing source |
20
+ | `xaiAdapter(opts?)` | Creates the `ProviderAdapter` for xAI |
21
+ | `XaiAdapterOptions` | `{ client?: XaiClientLike }` — inject a pre-built or fake client |
22
+ | `XaiClientLike` | Structural interface the adapter depends on (satisfied by real SDK and fakes) |
23
+ | `buildXaiClient(auth)` | Builds the real `openai`-SDK-backed client from `AuthMaterial`, pointed at xAI's base URL |
24
+ | `classifyXaiError(err)` | Classifies a raw thrown error into a typed `LlmError`, including xAI's 400-for-auth quirk |
25
+ | `grok45ModelDescriptor` | The `grok-4.5` `ModelDescriptor` |
26
+ | `grok46ModelDescriptor` | The `grok-4.6` `ModelDescriptor` |
27
+ | `grok47ModelDescriptor` | The `grok-4.7` `ModelDescriptor` |
28
+ | `xaiModelDescriptors` | Every model descriptor this package contributes (`grok-4.5`, `grok-4.6`, `grok-4.7`) |
29
+ | `xaiRegistry` | Pre-built `ModelRegistry` over `xaiModelDescriptors` |
30
+ | `xaiPricingSource()` | Built-in xAI `PricingSource` port implementation, backed by `XAI_PRICING` |
31
+ | `XAI_PRICING` | Frozen xAI pricing snapshot (µUSD per million tokens) |
32
+ | `XaiModelRates` | Per-model rate entry type (`inputPerM`, `cachedPerM`, `outputPerM`, optional `gt200k`) |
33
+ | `Grok45ConfigSchema` | Strict Zod config schema for `grok-4.5` |
34
+ | `Grok46ConfigSchema` | Strict Zod config schema for `grok-4.6` |
35
+ | `Grok47ConfigSchema` | Strict Zod config schema for `grok-4.7` |
36
+ | `XaiProviderOptions` | Typed `providerOptions.xai` shape for cache key, search tools, and parallel calls |
37
+ | `XaiFileStore` | Files API store: upload (TTL), get, list, idempotent delete, content |
38
+ | `XaiFileHandle` | `{ id, filename?, bytes?, expiresAt?, … }` returned by the store |
39
+ | `FileDeleteOptions` | `{ failClosed?, signal? }` — opt-in fail-closed delete for durable release gates |
40
+ | `XAI_FILE_TTL_*` | TTL bounds (`3600`…`2592000` seconds) and `XAI_FILE_MAX_BYTES` (48 MiB) |
39
41
 
40
42
  ## Quick example
41
43
 
@@ -123,18 +125,19 @@ const replay = await client.generate(
123
125
  // replay.text — model answer after the host dispatched the tool
124
126
  ```
125
127
 
126
- ## grok-4.5 and grok-4.6
128
+ ## grok-4.5, grok-4.6, and grok-4.7
127
129
 
128
- The default registry ships two canonical models (500k token context window each). They route through this adapter and support:
130
+ The default registry ships three canonical models (500k token context window each). They route through this adapter and support:
129
131
 
130
132
  - **Reasoning** — level-api (`reasoningApi: 'level'`), mapped to the Responses API `reasoning.effort` field. There is no `budgetTokens` field (xAI uses level-style reasoning) — passing it throws `bad_request`. The schema does not set a default effort; if `reasoning` is omitted, no `reasoning` field is sent and xAI's own server-side default (`high`) applies.
131
133
  - `grok-4.5`: `admittedReasoningEfforts: ['low', 'medium', 'high']` (live-verified 2026-08-24; `'medium'` is now accepted). `'none'` and `'xhigh'` are rejected. `'none'` remains rejected ("reasoning cannot be disabled").
132
- - `grok-4.6`: `admittedReasoningEfforts: ['low', 'medium', 'high', 'xhigh']` (live-verified 2026-08-12). `'none'` is rejected by the live API.
134
+ - `grok-4.6` and `grok-4.7`: `admittedReasoningEfforts: ['low', 'medium', 'high', 'xhigh']`. `'none'` is rejected.
133
135
  - **Structured output** — native. `output.jsonSchema` maps to the Responses API's `text.format` field with `{ type: 'json_schema', name, schema, strict: true }`, **not** `response_format` — this differs from OpenAI's own convention for the same underlying concept.
136
+ - **Structured output with built-in search** — admitted on `grok-4.6`, as captured in fixture 18. The adapter rejects this combination on descriptors without `structuredOutputWithTools`.
134
137
  - **`strict: true` performs no OpenAI-style compile-time schema validation, as of the 2026-07-09 live probes.** 2026-07-09 live verification against the real xAI Responses API — 13 single-variant probes plus 1 combined probe (14 calls total, all accepted HTTP 200; the combined probe is recorded as fixture `10-non-strict-schema-accepted.json`) — verified that `text.format` with `strict: true` accepted every one of the following schema shapes that OpenAI's own strict mode rejects at compile time: schemas (root and nested) missing `additionalProperties: false`; properties omitted from `required` (optional properties); `format`, `minLength`, `pattern`, and `default` keywords; `anyOf`; `$defs`/`$ref`; `enum`/`const`; and nullable unions (`type: [T, 'null']`). `strict: false` on the same surface showed no observed behavioral divergence from `strict: true`. This adapter forwards schemas to xAI verbatim — no rewriting, no preflight validation, and no injection of `additionalProperties: false` or `required` completion — so OpenAI-strict schema rewriting (including `@gullabs/codex-cli`'s `toOpenAiStrictOutputSchema` helper) is unnecessary for xai as of that verification date. (Reject-don't-map still applies to genuinely invalid input the xai schema/types layer itself rejects; this note is only about strict-mode compile-time schema-shape enforcement.) `packages/xai/src/__fixtures__/10-non-strict-schema-accepted.json` records one live example combining three of these — missing root `additionalProperties: false`, an optional property, and a `format` keyword — in a single accepted call.
135
138
  - **Sampling** — `temperature` and `topP` are forwarded verbatim. No `topK`.
136
139
  - **No penalties/stop** — `presence_penalty`, `frequency_penalty`, and `stop` are not in the config schema at all; xAI hard-rejects these on reasoning models, so the schema never admits them (reject-don't-map).
137
- - **Service tiers** — `grok-4.5` admits none; setting `serviceTier` throws `bad_request`. `grok-4.6` admits `serviceTier: 'priority'` only (Responses `service_tier: "priority"`, live-verified 2026-08-12). `'flex'` / `'standard'` / `'batch'` are rejected — xAI silently remaps unknown tiers to `default`, so this library never forwards them.
140
+ - **Service tiers** — all three models admit `serviceTier: 'priority'` (Responses `service_tier: "priority"`, billed at 2×). Grok 4.5 was captured live on 2026-09-25 and Grok 4.6 on 2026-08-12. `'flex'` / `'standard'` / `'batch'` are rejected — xAI silently remaps unknown tiers to `default`, so this library never forwards them.
138
141
 
139
142
  ## Files store (`XaiFileStore`)
140
143
 
@@ -181,13 +184,42 @@ try {
181
184
  | ZDR teams | New uploads and `file_id` attachments are blocked by xAI; errors mention Zero Data Retention when detectable |
182
185
  | Max size | 48 MiB (conservative vs docs 48–50 MB) |
183
186
 
184
- **Billing note:** attaching files on Responses implicitly enables xAI's `attachment_search` agentic tool. Live 2026-08-24 pins `web_search_calls` and `x_search_calls` in `usage.server_side_tool_usage_details` (flattened into `usage.details`). The attachment_search counter is **not** live-pinned (ZDR blocks file attach on this key); a `file-ref` call sets synthetic `usage.details.attachment_search_unpinned = 1` and `Cost.confidence: 'estimated'` — it is never reported as exact `$0`. `server_tools_requested = 1` is adapter-owned. Missing expected web/X counters → `tools: 0`, `estimated`, plus an adapter warning.
187
+ **Hidden input tokens:** grok-4.5, grok-4.6, and grok-4.7 bill about 1.3k hidden input tokens per request (1,532 for a one-line prompt versus 208 in July; cached on repeats).
188
+
189
+ **grok-4.7 replay:** the adapter sends `store: false`. Each result returns
190
+ `result.transientProviderState`, containing the complete wire input and
191
+ response output in provider order, including opaque `encrypted_content`,
192
+ messages, and server-tool items. Pass that object unchanged as
193
+ `request.transientProviderState` on the next request. This state is not written
194
+ to the call ledger. It is returned even for one-shot calls and can contain the
195
+ full prompt, inline media, and encrypted reasoning. Strip it before logging or
196
+ caching a whole result; store it securely only when continuation is needed.
197
+ When passing state, provide only new user or tool-result messages; the state
198
+ already contains prior turns. Use the new state returned by each subsequent
199
+ result. The adapter rejects assistant history alongside state, an empty new
200
+ message list, an unknown tool-result id, or a mismatched model. Without state,
201
+ a request starts a fresh conversation and may include text-only assistant
202
+ examples; function-call history requires state. Live fixtures
203
+ `28-grok-4-7-replay.json`, `30-grok-4-7-search-replay.json`, and
204
+ `31-grok-4-7-third-turn.json` cover function replay and follow-ups that replay
205
+ assistant message and web-search items.
206
+
207
+ **Billing note:** attaching files on Responses implicitly enables xAI's `attachment_search` agentic tool. `web_search_calls` is billed per call. Since 2026-09-21, x_search is billed from `x_posts_fetched` and `x_users_fetched`, not `x_search_calls`. The attachment_search counter is **not** live-pinned (P-X2); a `file-ref` call sets synthetic `usage.details.attachment_search_unpinned = 1` and `Cost.confidence: 'estimated'`. When a required server-tool counter is absent, the snapshot cost is unpriced (`microUsd: null`) rather than understating an unknown fee. The provider's billed `cost_in_usd_ticks` remains in raw usage for separate reconciliation; it is not represented as a rate-snapshot-derived `Cost`.
208
+
209
+ Fixture `19-x-search.json` was captured on 2026-08-24, before the billing
210
+ change. It has only `x_search_calls`; the fixture test retains its actual
211
+ billed total in usage but leaves snapshot cost unpriced.
212
+ Live 2026-09-26 fixtures `26-x-posts.json` and `27-x-users.json` pin both
213
+ item counters, including explicit zero counts, and reconcile snapshot cost to
214
+ the provider's billed ticks. P-X2 attachment counter verification remains
215
+ blocked: the available Zero Data Retention key returned 403 for file upload
216
+ and 400 for a public URL attachment (`29-attachment-zdr-blocked.json`).
185
217
 
186
218
  **Host tests:** `@gullabs/testing` exports `FakeXaiFileStore` (in-memory upload/get/delete with optional TTL clock and `failClosed`).
187
219
 
188
220
  ## Vision constraints
189
221
 
190
- Both models accept image input as an `inline-media` or `file-uri` `Part`, and document attachments as a `file-ref` `Part`:
222
+ All three models accept image input as an `inline-media` or `file-uri` `Part`, and document attachments as a `file-ref` `Part`:
191
223
 
192
224
  - **`inline-media`** — only `image/jpeg` and `image/png` are accepted; anything else throws `bad_request`. The decoded payload must be at most 20 MiB (xAI's documented inline-image ceiling); larger images throw `bad_request` before the request is sent.
193
225
  - **`file-uri`** — only accepted when the URI is a public `http(s)://` URL **and** the declared `mimeType` is jpg/png. A provider-hosted URI from another provider — for example a Gemini Files API URI (`https://generativelanguage.googleapis.com/...`) — is technically `https://` but is not dereferenceable by xAI and is not portable across providers. The adapter rejects it rather than trying to map or proxy it (reject-don't-map).
@@ -200,23 +232,26 @@ xAI caching is automatic — there is no explicit cache-create/cache-store API c
200
232
 
201
233
  ## Pricing
202
234
 
203
- `XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-08-24'`) — a point-in-time capture from `/v1/models`, not a live lookup (ADR-005). Rates are in µUSD per million tokens. Tool invocations add `Cost.details.tools` (`microUsd = input + cached + output + tools`):
235
+ `XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-09-25'`) — a point-in-time capture from `/v1/models`, not a live lookup (ADR-005). Rates are in µUSD per million tokens. Tool invocations add `Cost.details.tools` (`microUsd = input + cached + output + tools`):
204
236
 
205
- | Counter (raw `usage.details` key) | Rate |
206
- | --------------------------------- | ---------- |
207
- | `web_search_calls` | $5 / 1,000 |
208
- | `x_search_calls` | $5 / 1,000 |
237
+ | Counter (raw `usage.details` key) | Rate |
238
+ | --------------------------------- | -------------------- |
239
+ | `web_search_calls` | $5 / 1,000 calls |
240
+ | `x_posts_fetched` | $5 / 1,000 posts |
241
+ | `x_users_fetched` | $10 / 1,000 profiles |
209
242
 
210
243
  Enable Live Search with `providerOptions.xai.tools` (`web_search` / `x_search`). Citations land on `result.citations`. `countTokens` uses `POST /v1/tokenize-text` and returns `accuracy: 'lower-bound'` (text parts only; media / file parts are `bad_request`).
211
244
 
212
245
  | Model | Tier | Input | Cached input | Output |
213
246
  | ---------- | ---------------------------- | ------- | ------------ | -------- |
214
- | `grok-4.5` | standard (≤200k gross input) | $2.00/M | $0.30/M | $6.00/M |
215
- | `grok-4.5` | `gt200k` (>200k gross input) | $4.00/M | $0.60/M | $12.00/M |
216
- | `grok-4.6` | standard (≤200k gross input) | $2.00/M | $0.50/M | $6.00/M |
217
- | `grok-4.6` | `gt200k` (>200k gross input) | $4.00/M | $1.00/M | $12.00/M |
247
+ | `grok-4.5` | standard (<200k gross input) | $2.00/M | $0.30/M | $6.00/M |
248
+ | `grok-4.5` | `gt200k` (≥200k gross input) | $4.00/M | $0.60/M | $12.00/M |
249
+ | `grok-4.6` | standard (<200k gross input) | $2.00/M | $0.50/M | $6.00/M |
250
+ | `grok-4.6` | `gt200k` (≥200k gross input) | $4.00/M | $1.00/M | $12.00/M |
251
+ | `grok-4.7` | standard (<200k gross input) | $2.00/M | $0.50/M | $6.00/M |
252
+ | `grok-4.7` | `gt200k` (≥200k gross input) | $4.00/M | $1.00/M | $12.00/M |
218
253
 
219
- The `gt200k` long-context tier is selected by **gross** `inputTokens` (including cached), not billable input — strictly greater than 200,000 tokens, mirroring core's `selectRates` convention. The adapter now surfaces the echoed Responses `service_tier` (`'default'` or `'priority'`), so `price()` receives that served value instead of `undefined`. Custom xAI `PricingSource` implementations must price `'default'` at the standard list. Built-in `xaiPricingSource().price()` prices `grok-4.6` + `tier: 'priority'` at 2× every token type after the cache discount: uncached standard-list 2× is confirmed by fixture `12-grok-4-6-xhigh-priority.json` `cost_in_usd_ticks`; cached and `gt200k` legs follow the official 2×-after-cache-discount rule. Any other defined tier (including `priority` on `grok-4.5`) is unpriced (`microUsd: null`). Standard list rates are pinned to `packages/xai/src/__fixtures__/14-v1-models-pricing.json` (live `GET /v1/models` 2026-08-12).
254
+ The `gt200k` long-context tier is selected by **gross** `inputTokens` (including cached), not billable input — at or above 200,000 tokens (`long_context_threshold`), as stated on [xAI's pricing page](https://docs.x.ai/developers/pricing). The adapter surfaces the echoed Responses `service_tier` (`'default'` or `'priority'`), so `price()` receives that served value instead of `undefined`. Custom xAI `PricingSource` implementations must price `'default'` at the standard list. Built-in `xaiPricingSource().price()` prices priority at 2× every token type after the cache discount. Fixture `23-grok-4-5-priority.json` confirms Grok 4.5's 2× total, and fixture `12-grok-4-6-xhigh-priority.json` confirms Grok 4.6; cached and `gt200k` legs follow the official 2×-after-cache-discount rule. `fast` is not admitted. Any other defined tier is unpriced (`microUsd: null`). Grok 4.5/4.6 list rates are pinned to `packages/xai/src/__fixtures__/14-v1-models-pricing.json` (live `GET /v1/models` 2026-08-12); Grok 4.7 rates come from the [September 21 release notes](https://docs.x.ai/developers/release-notes).
220
255
 
221
256
  ## EU unavailability
222
257
 
@@ -238,7 +273,7 @@ xAI's own `/v1/models` listing surfaces `grok-4.5-latest` and `grok-build-latest
238
273
 
239
274
  - `providerOptions.xai.promptCacheKey` → `prompt_cache_key`
240
275
  - `reasoning.effort` → `reasoning.effort` (per-model admitted set)
241
- - `serviceTier: 'priority'` → `service_tier: 'priority'` (`grok-4.6` only)
276
+ - `serviceTier: 'priority'` → `service_tier: 'priority'` (all three models)
242
277
  - `output.jsonSchema` → `text.format: { type: 'json_schema', name, schema, strict: true }`
243
278
  - Usage: `usage.input_tokens` → `inputTokens`, `usage.output_tokens` → `outputTokens` (both already GROSS on xAI, unlike Gemini's sub-field summation); numeric extras (`num_sources_used`, `cost_in_usd_ticks`, etc.) surface into `usage.details` under their raw names, and the full raw payload is always in `usage.raw`
244
279
  - Errors: HTTP status is a hint. `classifyXaiError` inspects the STRUCTURED parsed body only — never free-form `Error.message`. Two recorded overlays: HTTP **400** whose body starts with `"Incorrect API key provided"` (prefix only; the SDK may drop `code`) → `invalid_auth`; HTTP **403** whose body starts with `"Content violates usage guidelines"` (e.g. `SAFETY_CHECK_TYPE_*`) → `content_filter`. A bare 403 without that body stays `invalid_auth`. Any other 400, `429`→`rate_limited`, `5xx`→`server`, and timeouts fall through to `@gullabs/core`'s generic `classifyError`.