@gullabs/xai 0.7.1 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +100 -43
- package/dist/index.cjs +884 -551
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +233 -42
- package/dist/index.d.ts +233 -42
- package/dist/index.js +884 -553
- package/dist/index.js.map +1 -1
- package/package.json +4 -4
package/README.md
CHANGED
|
@@ -14,28 +14,30 @@ xAI has no first-party TypeScript SDK. xAI's own quickstart recommends using the
|
|
|
14
14
|
|
|
15
15
|
## Key exports
|
|
16
16
|
|
|
17
|
-
| Export | What it is
|
|
18
|
-
| ----------------------- |
|
|
19
|
-
| `xaiProvider(opts?)` | `ProviderPlugin` factory — bundles the adapter, `grok-4.5` / `grok-4.6` descriptors, and pricing source |
|
|
20
|
-
| `xaiAdapter(opts?)` | Creates the `ProviderAdapter` for xAI
|
|
21
|
-
| `XaiAdapterOptions` | `{ client?: XaiClientLike }` — inject a pre-built or fake client
|
|
22
|
-
| `XaiClientLike` | Structural interface the adapter depends on (satisfied by real SDK and fakes)
|
|
23
|
-
| `buildXaiClient(auth)` | Builds the real `openai`-SDK-backed client from `AuthMaterial`, pointed at xAI's base URL
|
|
24
|
-
| `classifyXaiError(err)` | Classifies a raw thrown error into a typed `LlmError`, including xAI's 400-for-auth quirk
|
|
25
|
-
| `grok45ModelDescriptor` | The `grok-4.5` `ModelDescriptor`
|
|
26
|
-
| `grok46ModelDescriptor` | The `grok-4.6` `ModelDescriptor`
|
|
27
|
-
| `
|
|
28
|
-
| `
|
|
29
|
-
| `
|
|
30
|
-
| `
|
|
31
|
-
| `
|
|
32
|
-
| `
|
|
33
|
-
| `
|
|
34
|
-
| `
|
|
35
|
-
| `
|
|
36
|
-
| `
|
|
37
|
-
| `
|
|
38
|
-
| `
|
|
17
|
+
| Export | What it is |
|
|
18
|
+
| ----------------------- | -------------------------------------------------------------------------------------------------------------------- |
|
|
19
|
+
| `xaiProvider(opts?)` | `ProviderPlugin` factory — bundles the adapter, `grok-4.5` / `grok-4.6` / `grok-4.7` descriptors, and pricing source |
|
|
20
|
+
| `xaiAdapter(opts?)` | Creates the `ProviderAdapter` for xAI |
|
|
21
|
+
| `XaiAdapterOptions` | `{ client?: XaiClientLike }` — inject a pre-built or fake client |
|
|
22
|
+
| `XaiClientLike` | Structural interface the adapter depends on (satisfied by real SDK and fakes) |
|
|
23
|
+
| `buildXaiClient(auth)` | Builds the real `openai`-SDK-backed client from `AuthMaterial`, pointed at xAI's base URL |
|
|
24
|
+
| `classifyXaiError(err)` | Classifies a raw thrown error into a typed `LlmError`, including xAI's 400-for-auth quirk |
|
|
25
|
+
| `grok45ModelDescriptor` | The `grok-4.5` `ModelDescriptor` |
|
|
26
|
+
| `grok46ModelDescriptor` | The `grok-4.6` `ModelDescriptor` |
|
|
27
|
+
| `grok47ModelDescriptor` | The `grok-4.7` `ModelDescriptor` |
|
|
28
|
+
| `xaiModelDescriptors` | Every model descriptor this package contributes (`grok-4.5`, `grok-4.6`, `grok-4.7`) |
|
|
29
|
+
| `xaiRegistry` | Pre-built `ModelRegistry` over `xaiModelDescriptors` |
|
|
30
|
+
| `xaiPricingSource()` | Built-in xAI `PricingSource` port implementation, backed by `XAI_PRICING` |
|
|
31
|
+
| `XAI_PRICING` | Frozen xAI pricing snapshot (µUSD per million tokens) |
|
|
32
|
+
| `XaiModelRates` | Per-model rate entry type (`inputPerM`, `cachedPerM`, `outputPerM`, optional `gt200k`) |
|
|
33
|
+
| `Grok45ConfigSchema` | Strict Zod config schema for `grok-4.5` |
|
|
34
|
+
| `Grok46ConfigSchema` | Strict Zod config schema for `grok-4.6` |
|
|
35
|
+
| `Grok47ConfigSchema` | Strict Zod config schema for `grok-4.7` |
|
|
36
|
+
| `XaiProviderOptions` | Typed `providerOptions.xai` shape for cache key, search tools, tool choice, turn cap, and parallel calls |
|
|
37
|
+
| `XaiFileStore` | Files API store: upload (TTL), get, list, idempotent delete, content |
|
|
38
|
+
| `XaiFileHandle` | `{ id, filename?, bytes?, expiresAt?, … }` returned by the store |
|
|
39
|
+
| `FileDeleteOptions` | `{ failClosed?, signal? }` — opt-in fail-closed delete for durable release gates |
|
|
40
|
+
| `XAI_FILE_TTL_*` | TTL bounds (`3600`…`2592000` seconds) and `XAI_FILE_MAX_BYTES` (48 MiB) |
|
|
39
41
|
|
|
40
42
|
## Quick example
|
|
41
43
|
|
|
@@ -123,18 +125,21 @@ const replay = await client.generate(
|
|
|
123
125
|
// replay.text — model answer after the host dispatched the tool
|
|
124
126
|
```
|
|
125
127
|
|
|
126
|
-
## grok-4.5 and grok-4.
|
|
128
|
+
## grok-4.5, grok-4.6, and grok-4.7
|
|
127
129
|
|
|
128
|
-
The default registry ships
|
|
130
|
+
The default registry ships three canonical models (500k token context window each). They route through this adapter and support:
|
|
129
131
|
|
|
130
132
|
- **Reasoning** — level-api (`reasoningApi: 'level'`), mapped to the Responses API `reasoning.effort` field. There is no `budgetTokens` field (xAI uses level-style reasoning) — passing it throws `bad_request`. The schema does not set a default effort; if `reasoning` is omitted, no `reasoning` field is sent and xAI's own server-side default (`high`) applies.
|
|
131
133
|
- `grok-4.5`: `admittedReasoningEfforts: ['low', 'medium', 'high']` (live-verified 2026-08-24; `'medium'` is now accepted). `'none'` and `'xhigh'` are rejected. `'none'` remains rejected ("reasoning cannot be disabled").
|
|
132
|
-
- `grok-4.6`: `admittedReasoningEfforts: ['low', 'medium', 'high', 'xhigh']
|
|
134
|
+
- `grok-4.6` and `grok-4.7`: `admittedReasoningEfforts: ['low', 'medium', 'high', 'xhigh']`. `'none'` is rejected.
|
|
133
135
|
- **Structured output** — native. `output.jsonSchema` maps to the Responses API's `text.format` field with `{ type: 'json_schema', name, schema, strict: true }`, **not** `response_format` — this differs from OpenAI's own convention for the same underlying concept.
|
|
134
|
-
-
|
|
136
|
+
- **Structured output with built-in search** — admitted on all three models (`grok-4.6` in fixture 18; `grok-4.5` and `grok-4.7` live-verified 2026-10-02 in fixture 32). The adapter rejects this combination on descriptors without `structuredOutputWithTools`.
|
|
137
|
+
- **Output schemas are standard JSON Schema.** A nullable field lists `'null'` in `type` (`type: ['string', 'null']`). The adapter rejects the OpenAPI `nullable` keyword and uppercase type names (`STRING`, `OBJECT`) with `bad_request` before dispatch, naming the path, and never rewrites a schema. Live on 2026-10-02 (fixture 34): xAI accepted `nullable: true` and ignored it on all three models, so the model could not return `null` and wrote `""`, `0` or the string `"null"`; uppercase type names failed at xAI with HTTP 400.
|
|
138
|
+
- **`strict: true` performs no OpenAI-style compile-time schema validation, as of the 2026-07-09 live probes.** 2026-07-09 live verification against the real xAI Responses API — 13 single-variant probes plus 1 combined probe (14 calls total, all accepted HTTP 200; the combined probe is recorded as fixture `10-non-strict-schema-accepted.json`) — verified that `text.format` with `strict: true` accepted every one of the following schema shapes that OpenAI's own strict mode rejects at compile time: schemas (root and nested) missing `additionalProperties: false`; properties omitted from `required` (optional properties); `format`, `minLength`, `pattern`, and `default` keywords; `anyOf`; `$defs`/`$ref`; `enum`/`const`; and nullable unions (`type: [T, 'null']`). `strict: false` on the same surface showed no observed behavioral divergence from `strict: true`. This adapter forwards schemas to xAI verbatim — no rewriting, no OpenAI-strict preflight, and no injection of `additionalProperties: false` or `required` completion — so OpenAI-strict schema rewriting (including `@gullabs/codex-cli`'s `toOpenAiStrictOutputSchema` helper) is unnecessary for xai as of that verification date. The one preflight added since is the dialect check in the bullet above. (Reject-don't-map still applies to genuinely invalid input the xai schema/types layer itself rejects; this note is only about strict-mode compile-time schema-shape enforcement.) `packages/xai/src/__fixtures__/10-non-strict-schema-accepted.json` records one live example combining three of these — missing root `additionalProperties: false`, an optional property, and a `format` keyword — in a single accepted call.
|
|
135
139
|
- **Sampling** — `temperature` and `topP` are forwarded verbatim. No `topK`.
|
|
140
|
+
- **`max_output_tokens`** — forwarded only when the caller sets `maxOutputTokens`. The value includes both output and reasoning tokens and defaults to 128,000 when unset (docs read 2026-10-02). Truncation surfaces as `finishReason: 'length'`, not an error.
|
|
136
141
|
- **No penalties/stop** — `presence_penalty`, `frequency_penalty`, and `stop` are not in the config schema at all; xAI hard-rejects these on reasoning models, so the schema never admits them (reject-don't-map).
|
|
137
|
-
- **Service tiers** —
|
|
142
|
+
- **Service tiers** — all three models admit `serviceTier: 'priority'` (Responses `service_tier: "priority"`, billed at 2×). Grok 4.5 was captured live on 2026-09-25 and Grok 4.6 on 2026-08-12. `'flex'` / `'standard'` / `'batch'` are rejected — xAI silently remaps unknown tiers to `default`, so this library never forwards them.
|
|
138
143
|
|
|
139
144
|
## Files store (`XaiFileStore`)
|
|
140
145
|
|
|
@@ -181,13 +186,42 @@ try {
|
|
|
181
186
|
| ZDR teams | New uploads and `file_id` attachments are blocked by xAI; errors mention Zero Data Retention when detectable |
|
|
182
187
|
| Max size | 48 MiB (conservative vs docs 48–50 MB) |
|
|
183
188
|
|
|
184
|
-
**
|
|
189
|
+
**Hidden input tokens:** grok-4.5, grok-4.6, and grok-4.7 bill about 1.3k hidden input tokens per request (1,532 for a one-line prompt versus 208 in July; cached on repeats).
|
|
190
|
+
|
|
191
|
+
**grok-4.7 replay:** the adapter sends `store: false`. Each result returns
|
|
192
|
+
`result.transientProviderState`, containing the complete wire input and
|
|
193
|
+
response output in provider order, including opaque `encrypted_content`,
|
|
194
|
+
messages, and server-tool items. Pass that object unchanged as
|
|
195
|
+
`request.transientProviderState` on the next request. This state is not written
|
|
196
|
+
to the call ledger. It is returned even for one-shot calls and can contain the
|
|
197
|
+
full prompt, inline media, and encrypted reasoning. Strip it before logging or
|
|
198
|
+
caching a whole result; store it securely only when continuation is needed.
|
|
199
|
+
When passing state, provide only new user or tool-result messages; the state
|
|
200
|
+
already contains prior turns. Use the new state returned by each subsequent
|
|
201
|
+
result. The adapter rejects assistant history alongside state, an empty new
|
|
202
|
+
message list, an unknown tool-result id, or a mismatched model. Without state,
|
|
203
|
+
a request starts a fresh conversation and may include text-only assistant
|
|
204
|
+
examples; function-call history requires state. Live fixtures
|
|
205
|
+
`28-grok-4-7-replay.json`, `30-grok-4-7-search-replay.json`, and
|
|
206
|
+
`31-grok-4-7-third-turn.json` cover function replay and follow-ups that replay
|
|
207
|
+
assistant message and web-search items.
|
|
208
|
+
|
|
209
|
+
**Billing note:** attaching files on Responses implicitly enables xAI's `attachment_search` agentic tool. `web_search_calls` is billed per call. Since 2026-09-21, x_search is billed from `x_posts_fetched` and `x_users_fetched`, not `x_search_calls`. The attachment_search counter is **not** live-pinned (P-X2); a `file-ref` call sets synthetic `usage.details.attachment_search_unpinned = 1` and `Cost.confidence: 'estimated'`. When a required server-tool counter is absent, the snapshot cost is unpriced (`microUsd: null`) rather than understating an unknown fee. The provider's billed `cost_in_usd_ticks` remains in raw usage for separate reconciliation; it is not represented as a rate-snapshot-derived `Cost`.
|
|
210
|
+
|
|
211
|
+
Fixture `19-x-search.json` was captured on 2026-08-24, before the billing
|
|
212
|
+
change. It has only `x_search_calls`; the fixture test retains its actual
|
|
213
|
+
billed total in usage but leaves snapshot cost unpriced.
|
|
214
|
+
Live 2026-09-26 fixtures `26-x-posts.json` and `27-x-users.json` pin both
|
|
215
|
+
item counters, including explicit zero counts, and reconcile snapshot cost to
|
|
216
|
+
the provider's billed ticks. P-X2 attachment counter verification remains
|
|
217
|
+
blocked: the available Zero Data Retention key returned 403 for file upload
|
|
218
|
+
and 400 for a public URL attachment (`29-attachment-zdr-blocked.json`).
|
|
185
219
|
|
|
186
220
|
**Host tests:** `@gullabs/testing` exports `FakeXaiFileStore` (in-memory upload/get/delete with optional TTL clock and `failClosed`).
|
|
187
221
|
|
|
188
222
|
## Vision constraints
|
|
189
223
|
|
|
190
|
-
|
|
224
|
+
All three models accept image input as an `inline-media` or `file-uri` `Part`, and document attachments as a `file-ref` `Part`:
|
|
191
225
|
|
|
192
226
|
- **`inline-media`** — only `image/jpeg` and `image/png` are accepted; anything else throws `bad_request`. The decoded payload must be at most 20 MiB (xAI's documented inline-image ceiling); larger images throw `bad_request` before the request is sent.
|
|
193
227
|
- **`file-uri`** — only accepted when the URI is a public `http(s)://` URL **and** the declared `mimeType` is jpg/png. A provider-hosted URI from another provider — for example a Gemini Files API URI (`https://generativelanguage.googleapis.com/...`) — is technically `https://` but is not dereferenceable by xAI and is not portable across providers. The adapter rejects it rather than trying to map or proxy it (reject-don't-map).
|
|
@@ -200,27 +234,50 @@ xAI caching is automatic — there is no explicit cache-create/cache-store API c
|
|
|
200
234
|
|
|
201
235
|
## Pricing
|
|
202
236
|
|
|
203
|
-
`XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-
|
|
237
|
+
`XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-09-25'`) — a point-in-time capture from `/v1/models`, not a live lookup (ADR-005). Rates are in µUSD per million tokens. Tool invocations add `Cost.details.tools` (`microUsd = input + cached + output + tools`):
|
|
204
238
|
|
|
205
|
-
| Counter (raw `usage.details` key) | Rate
|
|
206
|
-
| --------------------------------- |
|
|
207
|
-
| `web_search_calls` | $5 / 1,000 |
|
|
208
|
-
| `
|
|
239
|
+
| Counter (raw `usage.details` key) | Rate |
|
|
240
|
+
| --------------------------------- | -------------------- |
|
|
241
|
+
| `web_search_calls` | $5 / 1,000 calls |
|
|
242
|
+
| `x_posts_fetched` | $5 / 1,000 posts |
|
|
243
|
+
| `x_users_fetched` | $10 / 1,000 profiles |
|
|
209
244
|
|
|
210
|
-
Enable Live Search with `providerOptions.xai.tools` (`web_search` / `x_search`). Citations land on `result.citations`.
|
|
245
|
+
Enable Live Search with `providerOptions.xai.tools` (`web_search` / `x_search`). Citations land on `result.citations`.
|
|
246
|
+
|
|
247
|
+
### Controlling the search tools
|
|
248
|
+
|
|
249
|
+
```ts
|
|
250
|
+
config: {
|
|
251
|
+
providerOptions: {
|
|
252
|
+
xai: {
|
|
253
|
+
tools: [{ type: 'web_search' }, { type: 'x_search' }],
|
|
254
|
+
toolChoice: 'required', // 'auto' | 'required' | 'none'
|
|
255
|
+
maxTurns: 3,
|
|
256
|
+
},
|
|
257
|
+
},
|
|
258
|
+
}
|
|
259
|
+
```
|
|
260
|
+
|
|
261
|
+
- **`toolChoice`** maps to the Responses `tool_choice` for the search tools. Left on `auto`, a model can answer without searching; `required` forces at least one search and `none` disables the declared tools. Live on 2026-10-02 (fixture 32), `required` ran 3 / 2 / 2 searches on grok-4.5 / 4.6 / 4.7 and `none` ran 0. It needs a non-empty `tools`, and the adapter rejects it together with function tools, file attachments or the request-level `toolChoice`: xAI takes one `tool_choice` per request and `required` means "at least one tool", which a function call or the implicit `attachment_search` would satisfy. Send it on every request; nothing carries over between calls.
|
|
262
|
+
- **`maxTurns`** maps to the Responses `max_turns` (integer ≥ 1, needs `tools`). xAI documents it as the cap on agentic tool-calling turns. A turn can run several searches, so it is not a search count. **xAI did not enforce it as of 2026-10-02** (fixture 33): with `max_turns: 1` the three models still ran 10 to 17 searches over several rounds. The option is forwarded verbatim so hosts get the cap when xAI enforces it. Until then, state the search budget in the prompt and assert on the observed count.
|
|
263
|
+
- **Observed count.** `result.usage.details.web_search_calls` is the number of web searches billed; `x_posts_fetched` and `x_users_fetched` are the X Search billing counters (items, not calls). All three persist to the ledger's token details. When no server tool ran, xAI reports `num_server_side_tools_used: 0` and omits the counters; the adapter prices that call exactly with no tool fee.
|
|
264
|
+
- **Cost.** There is no enforceable search cap, and every search result is fed back as input. One uncapped grok-4.7 research call used 362k input tokens, which crosses the 200k long-context threshold, and cost about $1.07.
|
|
265
|
+
`countTokens` uses `POST /v1/tokenize-text` and returns `accuracy: 'lower-bound'` (text parts only; media / file parts are `bad_request`).
|
|
211
266
|
|
|
212
267
|
| Model | Tier | Input | Cached input | Output |
|
|
213
268
|
| ---------- | ---------------------------- | ------- | ------------ | -------- |
|
|
214
|
-
| `grok-4.5` | standard (
|
|
215
|
-
| `grok-4.5` | `gt200k` (
|
|
216
|
-
| `grok-4.6` | standard (
|
|
217
|
-
| `grok-4.6` | `gt200k` (
|
|
269
|
+
| `grok-4.5` | standard (<200k gross input) | $2.00/M | $0.30/M | $6.00/M |
|
|
270
|
+
| `grok-4.5` | `gt200k` (≥200k gross input) | $4.00/M | $0.60/M | $12.00/M |
|
|
271
|
+
| `grok-4.6` | standard (<200k gross input) | $2.00/M | $0.50/M | $6.00/M |
|
|
272
|
+
| `grok-4.6` | `gt200k` (≥200k gross input) | $4.00/M | $1.00/M | $12.00/M |
|
|
273
|
+
| `grok-4.7` | standard (<200k gross input) | $2.00/M | $0.50/M | $6.00/M |
|
|
274
|
+
| `grok-4.7` | `gt200k` (≥200k gross input) | $4.00/M | $1.00/M | $12.00/M |
|
|
218
275
|
|
|
219
|
-
The `gt200k` long-context tier is selected by **gross** `inputTokens` (including cached), not billable input —
|
|
276
|
+
The `gt200k` long-context tier is selected by **gross** `inputTokens` (including cached), not billable input — at or above 200,000 tokens (`long_context_threshold`), as stated on [xAI's pricing page](https://docs.x.ai/developers/pricing). The adapter surfaces the echoed Responses `service_tier` (`'default'` or `'priority'`), so `price()` receives that served value instead of `undefined`. Custom xAI `PricingSource` implementations must price `'default'` at the standard list. Built-in `xaiPricingSource().price()` prices priority at 2× every token type after the cache discount. Fixture `23-grok-4-5-priority.json` confirms Grok 4.5's 2× total, and fixture `12-grok-4-6-xhigh-priority.json` confirms Grok 4.6; cached and `gt200k` legs follow the official 2×-after-cache-discount rule. `fast` is not admitted. Any other defined tier is unpriced (`microUsd: null`). Grok 4.5/4.6 list rates are pinned to `packages/xai/src/__fixtures__/14-v1-models-pricing.json` (live `GET /v1/models` 2026-08-12); Grok 4.7 rates come from the [September 21 release notes](https://docs.x.ai/developers/release-notes).
|
|
220
277
|
|
|
221
|
-
##
|
|
278
|
+
## Regions
|
|
222
279
|
|
|
223
|
-
|
|
280
|
+
The grok-4.6 and grok-4.7 model pages list `us-east-1`, `us-west-2`, and `us-central-1` (docs read 2026-10-02). The release notes say Grok 4.5 is available in the API console for EU users (docs read 2026-10-02). This is a hosting/deployment concern for callers, not something this library can route around; it is documented here so consumers are not surprised by data-residency constraints.
|
|
224
281
|
|
|
225
282
|
## Aliases are not registered
|
|
226
283
|
|
|
@@ -238,7 +295,7 @@ xAI's own `/v1/models` listing surfaces `grok-4.5-latest` and `grok-build-latest
|
|
|
238
295
|
|
|
239
296
|
- `providerOptions.xai.promptCacheKey` → `prompt_cache_key`
|
|
240
297
|
- `reasoning.effort` → `reasoning.effort` (per-model admitted set)
|
|
241
|
-
- `serviceTier: 'priority'` → `service_tier: 'priority'` (
|
|
298
|
+
- `serviceTier: 'priority'` → `service_tier: 'priority'` (all three models)
|
|
242
299
|
- `output.jsonSchema` → `text.format: { type: 'json_schema', name, schema, strict: true }`
|
|
243
300
|
- Usage: `usage.input_tokens` → `inputTokens`, `usage.output_tokens` → `outputTokens` (both already GROSS on xAI, unlike Gemini's sub-field summation); numeric extras (`num_sources_used`, `cost_in_usd_ticks`, etc.) surface into `usage.details` under their raw names, and the full raw payload is always in `usage.raw`
|
|
244
301
|
- Errors: HTTP status is a hint. `classifyXaiError` inspects the STRUCTURED parsed body only — never free-form `Error.message`. Two recorded overlays: HTTP **400** whose body starts with `"Incorrect API key provided"` (prefix only; the SDK may drop `code`) → `invalid_auth`; HTTP **403** whose body starts with `"Content violates usage guidelines"` (e.g. `SAFETY_CHECK_TYPE_*`) → `content_filter`. A bare 403 without that body stays `invalid_auth`. Any other 400, `429`→`rate_limited`, `5xx`→`server`, and timeouts fall through to `@gullabs/core`'s generic `classifyError`.
|