@gullabs/xai 0.5.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -63,7 +63,7 @@ const result = await client.generate(
63
63
  The default registry ships two canonical models (500k token context window each). They route through this adapter and support:
64
64
 
65
65
  - **Reasoning** — level-api (`reasoningApi: 'level'`), mapped to the Responses API `reasoning.effort` field. There is no `budgetTokens` field (xAI uses level-style reasoning) — passing it throws `bad_request`. The schema does not set a default effort; if `reasoning` is omitted, no `reasoning` field is sent and xAI's own server-side default (`high`) applies.
66
- - `grok-4.5`: `admittedReasoningEfforts: ['low', 'high']`. `'none'`, `'medium'`, and `'xhigh'` are rejected by `Grok45ConfigSchema` (frozen 2026-07-09 live probe; do not silently widen).
66
+ - `grok-4.5`: `admittedReasoningEfforts: ['low', 'medium', 'high']` (live-verified 2026-08-24; `'medium'` is now accepted). `'none'` and `'xhigh'` are rejected. `'none'` remains rejected ("reasoning cannot be disabled").
67
67
  - `grok-4.6`: `admittedReasoningEfforts: ['low', 'medium', 'high', 'xhigh']` (live-verified 2026-08-12). `'none'` is rejected by the live API.
68
68
  - **Structured output** — native. `output.jsonSchema` maps to the Responses API's `text.format` field with `{ type: 'json_schema', name, schema, strict: true }`, **not** `response_format` — this differs from OpenAI's own convention for the same underlying concept.
69
69
  - **`strict: true` performs no OpenAI-style compile-time schema validation, as of the 2026-07-09 live probes.** 2026-07-09 live verification against the real xAI Responses API — 13 single-variant probes plus 1 combined probe (14 calls total, all accepted HTTP 200; the combined probe is recorded as fixture `10-non-strict-schema-accepted.json`) — verified that `text.format` with `strict: true` accepted every one of the following schema shapes that OpenAI's own strict mode rejects at compile time: schemas (root and nested) missing `additionalProperties: false`; properties omitted from `required` (optional properties); `format`, `minLength`, `pattern`, and `default` keywords; `anyOf`; `$defs`/`$ref`; `enum`/`const`; and nullable unions (`type: [T, 'null']`). `strict: false` on the same surface showed no observed behavioral divergence from `strict: true`. This adapter forwards schemas to xAI verbatim — no rewriting, no preflight validation, and no injection of `additionalProperties: false` or `required` completion — so OpenAI-strict schema rewriting (including `@gullabs/codex-cli`'s `toOpenAiStrictOutputSchema` helper) is unnecessary for xai as of that verification date. (Reject-don't-map still applies to genuinely invalid input the xai schema/types layer itself rejects; this note is only about strict-mode compile-time schema-shape enforcement.) `packages/xai/src/__fixtures__/10-non-strict-schema-accepted.json` records one live example combining three of these — missing root `additionalProperties: false`, an optional property, and a `format` keyword — in a single accepted call.
@@ -85,7 +85,7 @@ const store = new XaiFileStore({
85
85
 
86
86
  const handle = await store.upload({
87
87
  data: pdfBytes,
88
- filename: 'matter.pdf',
88
+ filename: 'document.pdf',
89
89
  mimeType: 'application/pdf',
90
90
  expiresAfterSeconds: 86_400, // 24h; range 3600…2592000
91
91
  })
@@ -116,7 +116,7 @@ try {
116
116
  | ZDR teams | New uploads and `file_id` attachments are blocked by xAI; errors mention Zero Data Retention when detectable |
117
117
  | Max size | 48 MiB (conservative vs docs 48–50 MB) |
118
118
 
119
- **Billing note:** attaching files on Responses implicitly enables xAI's `attachment_search` agentic tool. Expect tool-invocation fees and reasoning tokens beyond a plain completion. When xAI returns numeric counters such as `num_server_side_tools_used` / `num_sources_used`, they appear on `usage.details` under those raw names (and full payload in `usage.raw`) for host visibility — they are **not** folded into `computeXaiCost` token lanes yet (no tool Cost lane). Collections / public URL minting are out of scope for this store.
119
+ **Billing note:** attaching files on Responses implicitly enables xAI's `attachment_search` agentic tool. Per-invocation fees land in `Cost.details.tools`. Live 2026-08-24 `/v1/responses` pins counters at `usage.server_side_tool_usage_details` (`web_search_calls`, `x_search_calls`, `document_search_calls`, …); the adapter flattens those names into `usage.details`. File-attach confirmation of the attachment lane is blocked on ZDR keys (uploads return Zero Data Retention); `document_search_calls` is the live-payload key used for the $10/1k attachment lane. When the adapter requested server tools (`providerOptions.xai.tools` or `file-ref`) it also sets synthetic `usage.details.server_tools_requested = 1` (adapter-owned, not a provider field). Missing counters → `tools: 0`, `confidence: 'estimated'`, plus an adapter warning.
120
120
 
121
121
  **Host tests:** `@gullabs/testing` exports `FakeXaiFileStore` (in-memory upload/get/delete with optional TTL clock and `failClosed`).
122
122
 
@@ -135,7 +135,14 @@ xAI caching is automatic — there is no explicit cache-create/cache-store API c
135
135
 
136
136
  ## Pricing
137
137
 
138
- `XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-08-12'`) — a point-in-time capture from `/v1/models`, not a live lookup (ADR-005). Rates are in µUSD per million tokens:
138
+ `XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-08-24'`) — a point-in-time capture from `/v1/models`, not a live lookup (ADR-005). Rates are in µUSD per million tokens. Tool invocations add `Cost.details.tools` (`microUsd = input + cached + output + tools`):
139
+
140
+ | Counter (raw `usage.details` key) | Rate |
141
+ | --------------------------------- | ---------- |
142
+ | `web_search_calls` | $5 / 1,000 |
143
+ | `x_search_calls` | $5 / 1,000 |
144
+
145
+ Enable Live Search with `providerOptions.xai.tools` (`web_search` / `x_search`). Citations land on `result.citations`. `countTokens` uses `POST /v1/tokenize-text` and returns `accuracy: 'lower-bound'` (text parts only; media / file parts are `bad_request`).
139
146
 
140
147
  | Model | Tier | Input | Cached input | Output |
141
148
  | ---------- | ---------------------------- | ------- | ------------ | -------- |
@@ -169,7 +176,7 @@ xAI's own `/v1/models` listing surfaces `grok-4.5-latest` and `grok-build-latest
169
176
  - `serviceTier: 'priority'` → `service_tier: 'priority'` (`grok-4.6` only)
170
177
  - `output.jsonSchema` → `text.format: { type: 'json_schema', name, schema, strict: true }`
171
178
  - Usage: `usage.input_tokens` → `inputTokens`, `usage.output_tokens` → `outputTokens` (both already GROSS on xAI, unlike Gemini's sub-field summation); numeric extras (`num_sources_used`, `cost_in_usd_ticks`, etc.) surface into `usage.details` under their raw names, and the full raw payload is always in `usage.raw`
172
- - Errors: xAI's Responses API returns **HTTP 400 (not 401) for an invalid API key**. `classifyXaiError` special-cases the exact structured error-body signature (`code: 'invalid-argument'` with message prefix `"Incorrect API key provided"`, taken verbatim from a recorded live fixture) and reclassifies it as `invalid_auth`. It only inspects the STRUCTURED parsed error body — never free-form `Error.message` text — so a 400 that merely _mentions_ an API key (e.g. a schema-validation error echoing user content) stays `bad_request`. Any other 400, `429`→`rate_limited`, `5xx`→`server`, and timeouts fall through to `@gullabs/core`'s generic `classifyError`.
179
+ - Errors: HTTP status is a hint. `classifyXaiError` inspects the STRUCTURED parsed body only — never free-form `Error.message`. Two recorded overlays: HTTP **400** whose body starts with `"Incorrect API key provided"` (prefix only; the SDK may drop `code`) → `invalid_auth`; HTTP **403** whose body starts with `"Content violates usage guidelines"` (e.g. `SAFETY_CHECK_TYPE_*`) → `content_filter`. A bare 403 without that body stays `invalid_auth`. Any other 400, `429`→`rate_limited`, `5xx`→`server`, and timeouts fall through to `@gullabs/core`'s generic `classifyError`.
173
180
 
174
181
  ## Learn more
175
182