@gullabs/xai 0.5.0 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +12 -5
- package/dist/index.cjs +479 -39
- package/dist/index.cjs.map +1 -1
- package/dist/index.d.cts +310 -29
- package/dist/index.d.ts +310 -29
- package/dist/index.js +480 -41
- package/dist/index.js.map +1 -1
- package/package.json +6 -6
package/README.md
CHANGED
|
@@ -63,7 +63,7 @@ const result = await client.generate(
|
|
|
63
63
|
The default registry ships two canonical models (500k token context window each). They route through this adapter and support:
|
|
64
64
|
|
|
65
65
|
- **Reasoning** — level-api (`reasoningApi: 'level'`), mapped to the Responses API `reasoning.effort` field. There is no `budgetTokens` field (xAI uses level-style reasoning) — passing it throws `bad_request`. The schema does not set a default effort; if `reasoning` is omitted, no `reasoning` field is sent and xAI's own server-side default (`high`) applies.
|
|
66
|
-
- `grok-4.5`: `admittedReasoningEfforts: ['low', 'high']
|
|
66
|
+
- `grok-4.5`: `admittedReasoningEfforts: ['low', 'medium', 'high']` (live-verified 2026-08-24; `'medium'` is now accepted). `'none'` and `'xhigh'` are rejected. `'none'` remains rejected ("reasoning cannot be disabled").
|
|
67
67
|
- `grok-4.6`: `admittedReasoningEfforts: ['low', 'medium', 'high', 'xhigh']` (live-verified 2026-08-12). `'none'` is rejected by the live API.
|
|
68
68
|
- **Structured output** — native. `output.jsonSchema` maps to the Responses API's `text.format` field with `{ type: 'json_schema', name, schema, strict: true }`, **not** `response_format` — this differs from OpenAI's own convention for the same underlying concept.
|
|
69
69
|
- **`strict: true` performs no OpenAI-style compile-time schema validation, as of the 2026-07-09 live probes.** 2026-07-09 live verification against the real xAI Responses API — 13 single-variant probes plus 1 combined probe (14 calls total, all accepted HTTP 200; the combined probe is recorded as fixture `10-non-strict-schema-accepted.json`) — verified that `text.format` with `strict: true` accepted every one of the following schema shapes that OpenAI's own strict mode rejects at compile time: schemas (root and nested) missing `additionalProperties: false`; properties omitted from `required` (optional properties); `format`, `minLength`, `pattern`, and `default` keywords; `anyOf`; `$defs`/`$ref`; `enum`/`const`; and nullable unions (`type: [T, 'null']`). `strict: false` on the same surface showed no observed behavioral divergence from `strict: true`. This adapter forwards schemas to xAI verbatim — no rewriting, no preflight validation, and no injection of `additionalProperties: false` or `required` completion — so OpenAI-strict schema rewriting (including `@gullabs/codex-cli`'s `toOpenAiStrictOutputSchema` helper) is unnecessary for xai as of that verification date. (Reject-don't-map still applies to genuinely invalid input the xai schema/types layer itself rejects; this note is only about strict-mode compile-time schema-shape enforcement.) `packages/xai/src/__fixtures__/10-non-strict-schema-accepted.json` records one live example combining three of these — missing root `additionalProperties: false`, an optional property, and a `format` keyword — in a single accepted call.
|
|
@@ -85,7 +85,7 @@ const store = new XaiFileStore({
|
|
|
85
85
|
|
|
86
86
|
const handle = await store.upload({
|
|
87
87
|
data: pdfBytes,
|
|
88
|
-
filename: '
|
|
88
|
+
filename: 'document.pdf',
|
|
89
89
|
mimeType: 'application/pdf',
|
|
90
90
|
expiresAfterSeconds: 86_400, // 24h; range 3600…2592000
|
|
91
91
|
})
|
|
@@ -116,7 +116,7 @@ try {
|
|
|
116
116
|
| ZDR teams | New uploads and `file_id` attachments are blocked by xAI; errors mention Zero Data Retention when detectable |
|
|
117
117
|
| Max size | 48 MiB (conservative vs docs 48–50 MB) |
|
|
118
118
|
|
|
119
|
-
**Billing note:** attaching files on Responses implicitly enables xAI's `attachment_search` agentic tool.
|
|
119
|
+
**Billing note:** attaching files on Responses implicitly enables xAI's `attachment_search` agentic tool. Per-invocation fees land in `Cost.details.tools`. Live 2026-08-24 `/v1/responses` pins counters at `usage.server_side_tool_usage_details` (`web_search_calls`, `x_search_calls`, `document_search_calls`, …); the adapter flattens those names into `usage.details`. File-attach confirmation of the attachment lane is blocked on ZDR keys (uploads return Zero Data Retention); `document_search_calls` is the live-payload key used for the $10/1k attachment lane. When the adapter requested server tools (`providerOptions.xai.tools` or `file-ref`) it also sets synthetic `usage.details.server_tools_requested = 1` (adapter-owned, not a provider field). Missing counters → `tools: 0`, `confidence: 'estimated'`, plus an adapter warning.
|
|
120
120
|
|
|
121
121
|
**Host tests:** `@gullabs/testing` exports `FakeXaiFileStore` (in-memory upload/get/delete with optional TTL clock and `failClosed`).
|
|
122
122
|
|
|
@@ -135,7 +135,14 @@ xAI caching is automatic — there is no explicit cache-create/cache-store API c
|
|
|
135
135
|
|
|
136
136
|
## Pricing
|
|
137
137
|
|
|
138
|
-
`XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-08-
|
|
138
|
+
`XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-08-24'`) — a point-in-time capture from `/v1/models`, not a live lookup (ADR-005). Rates are in µUSD per million tokens. Tool invocations add `Cost.details.tools` (`microUsd = input + cached + output + tools`):
|
|
139
|
+
|
|
140
|
+
| Counter (raw `usage.details` key) | Rate |
|
|
141
|
+
| --------------------------------- | ---------- |
|
|
142
|
+
| `web_search_calls` | $5 / 1,000 |
|
|
143
|
+
| `x_search_calls` | $5 / 1,000 |
|
|
144
|
+
|
|
145
|
+
Enable Live Search with `providerOptions.xai.tools` (`web_search` / `x_search`). Citations land on `result.citations`. `countTokens` uses `POST /v1/tokenize-text` and returns `accuracy: 'lower-bound'` (text parts only; media / file parts are `bad_request`).
|
|
139
146
|
|
|
140
147
|
| Model | Tier | Input | Cached input | Output |
|
|
141
148
|
| ---------- | ---------------------------- | ------- | ------------ | -------- |
|
|
@@ -169,7 +176,7 @@ xAI's own `/v1/models` listing surfaces `grok-4.5-latest` and `grok-build-latest
|
|
|
169
176
|
- `serviceTier: 'priority'` → `service_tier: 'priority'` (`grok-4.6` only)
|
|
170
177
|
- `output.jsonSchema` → `text.format: { type: 'json_schema', name, schema, strict: true }`
|
|
171
178
|
- Usage: `usage.input_tokens` → `inputTokens`, `usage.output_tokens` → `outputTokens` (both already GROSS on xAI, unlike Gemini's sub-field summation); numeric extras (`num_sources_used`, `cost_in_usd_ticks`, etc.) surface into `usage.details` under their raw names, and the full raw payload is always in `usage.raw`
|
|
172
|
-
- Errors:
|
|
179
|
+
- Errors: HTTP status is a hint. `classifyXaiError` inspects the STRUCTURED parsed body only — never free-form `Error.message`. Two recorded overlays: HTTP **400** whose body starts with `"Incorrect API key provided"` (prefix only; the SDK may drop `code`) → `invalid_auth`; HTTP **403** whose body starts with `"Content violates usage guidelines"` (e.g. `SAFETY_CHECK_TYPE_*`) → `content_filter`. A bare 403 without that body stays `invalid_auth`. Any other 400, `429`→`rate_limited`, `5xx`→`server`, and timeouts fall through to `@gullabs/core`'s generic `classifyError`.
|
|
173
180
|
|
|
174
181
|
## Learn more
|
|
175
182
|
|