@gullabs/xai 0.5.1 → 0.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -58,12 +58,77 @@ const result = await client.generate(
58
58
  )
59
59
  ```
60
60
 
61
+ ### Function-calling seam (no agent loop)
62
+
63
+ ```ts
64
+ const tools = [
65
+ {
66
+ name: 'get_temperature',
67
+ description: 'Get current temperature for a location',
68
+ inputJsonSchema: {
69
+ type: 'object',
70
+ properties: { location: { type: 'string' } },
71
+ required: ['location'],
72
+ },
73
+ },
74
+ ]
75
+
76
+ const first = await client.generate(
77
+ {
78
+ provider: 'xai',
79
+ model: 'grok-4.6',
80
+ messages: [{ role: 'user', parts: [{ kind: 'text', text: 'Temperature in SF?' }] }],
81
+ tools,
82
+ toolChoice: 'required',
83
+ },
84
+ { auth: { apiKey: 'YOUR_XAI_API_KEY' } },
85
+ )
86
+ // first.finishReason === 'tool_calls'
87
+ // first.toolCalls === [{ toolCallId, toolName, args }]
88
+
89
+ const call = first.toolCalls![0]!
90
+ const replay = await client.generate(
91
+ {
92
+ provider: 'xai',
93
+ model: 'grok-4.6',
94
+ messages: [
95
+ { role: 'user', parts: [{ kind: 'text', text: 'Temperature in SF?' }] },
96
+ {
97
+ role: 'assistant',
98
+ parts: [
99
+ {
100
+ kind: 'tool-call',
101
+ toolCallId: call.toolCallId,
102
+ toolName: call.toolName,
103
+ args: call.args,
104
+ },
105
+ ],
106
+ },
107
+ {
108
+ role: 'user',
109
+ parts: [
110
+ {
111
+ kind: 'tool-result',
112
+ toolCallId: call.toolCallId,
113
+ toolName: call.toolName,
114
+ result: { temperature: 59 },
115
+ },
116
+ ],
117
+ },
118
+ ],
119
+ tools,
120
+ },
121
+ { auth: { apiKey: 'YOUR_XAI_API_KEY' } },
122
+ )
123
+ // replay.text — model answer after the host dispatched the tool
124
+ ```
125
+
61
126
  ## grok-4.5 and grok-4.6
62
127
 
63
128
  The default registry ships two canonical models (500k token context window each). They route through this adapter and support:
64
129
 
65
130
  - **Reasoning** — level-api (`reasoningApi: 'level'`), mapped to the Responses API `reasoning.effort` field. There is no `budgetTokens` field (xAI uses level-style reasoning) — passing it throws `bad_request`. The schema does not set a default effort; if `reasoning` is omitted, no `reasoning` field is sent and xAI's own server-side default (`high`) applies.
66
- - `grok-4.5`: `admittedReasoningEfforts: ['low', 'high']`. `'none'`, `'medium'`, and `'xhigh'` are rejected by `Grok45ConfigSchema` (frozen 2026-07-09 live probe; do not silently widen).
131
+ - `grok-4.5`: `admittedReasoningEfforts: ['low', 'medium', 'high']` (live-verified 2026-08-24; `'medium'` is now accepted). `'none'` and `'xhigh'` are rejected. `'none'` remains rejected ("reasoning cannot be disabled").
67
132
  - `grok-4.6`: `admittedReasoningEfforts: ['low', 'medium', 'high', 'xhigh']` (live-verified 2026-08-12). `'none'` is rejected by the live API.
68
133
  - **Structured output** — native. `output.jsonSchema` maps to the Responses API's `text.format` field with `{ type: 'json_schema', name, schema, strict: true }`, **not** `response_format` — this differs from OpenAI's own convention for the same underlying concept.
69
134
  - **`strict: true` performs no OpenAI-style compile-time schema validation, as of the 2026-07-09 live probes.** 2026-07-09 live verification against the real xAI Responses API — 13 single-variant probes plus 1 combined probe (14 calls total, all accepted HTTP 200; the combined probe is recorded as fixture `10-non-strict-schema-accepted.json`) — verified that `text.format` with `strict: true` accepted every one of the following schema shapes that OpenAI's own strict mode rejects at compile time: schemas (root and nested) missing `additionalProperties: false`; properties omitted from `required` (optional properties); `format`, `minLength`, `pattern`, and `default` keywords; `anyOf`; `$defs`/`$ref`; `enum`/`const`; and nullable unions (`type: [T, 'null']`). `strict: false` on the same surface showed no observed behavioral divergence from `strict: true`. This adapter forwards schemas to xAI verbatim — no rewriting, no preflight validation, and no injection of `additionalProperties: false` or `required` completion — so OpenAI-strict schema rewriting (including `@gullabs/codex-cli`'s `toOpenAiStrictOutputSchema` helper) is unnecessary for xai as of that verification date. (Reject-don't-map still applies to genuinely invalid input the xai schema/types layer itself rejects; this note is only about strict-mode compile-time schema-shape enforcement.) `packages/xai/src/__fixtures__/10-non-strict-schema-accepted.json` records one live example combining three of these — missing root `additionalProperties: false`, an optional property, and a `format` keyword — in a single accepted call.
@@ -85,7 +150,7 @@ const store = new XaiFileStore({
85
150
 
86
151
  const handle = await store.upload({
87
152
  data: pdfBytes,
88
- filename: 'matter.pdf',
153
+ filename: 'document.pdf',
89
154
  mimeType: 'application/pdf',
90
155
  expiresAfterSeconds: 86_400, // 24h; range 3600…2592000
91
156
  })
@@ -116,7 +181,7 @@ try {
116
181
  | ZDR teams | New uploads and `file_id` attachments are blocked by xAI; errors mention Zero Data Retention when detectable |
117
182
  | Max size | 48 MiB (conservative vs docs 48–50 MB) |
118
183
 
119
- **Billing note:** attaching files on Responses implicitly enables xAI's `attachment_search` agentic tool. Expect tool-invocation fees and reasoning tokens beyond a plain completion. When xAI returns numeric counters such as `num_server_side_tools_used` / `num_sources_used`, they appear on `usage.details` under those raw names (and full payload in `usage.raw`) for host visibility — they are **not** folded into `computeXaiCost` token lanes yet (no tool Cost lane). Collections / public URL minting are out of scope for this store.
184
+ **Billing note:** attaching files on Responses implicitly enables xAI's `attachment_search` agentic tool. Live 2026-08-24 pins `web_search_calls` and `x_search_calls` in `usage.server_side_tool_usage_details` (flattened into `usage.details`). The attachment_search counter is **not** live-pinned (ZDR blocks file attach on this key); a `file-ref` call sets synthetic `usage.details.attachment_search_unpinned = 1` and `Cost.confidence: 'estimated'` — it is never reported as exact `$0`. `server_tools_requested = 1` is adapter-owned. Missing expected web/X counters → `tools: 0`, `estimated`, plus an adapter warning.
120
185
 
121
186
  **Host tests:** `@gullabs/testing` exports `FakeXaiFileStore` (in-memory upload/get/delete with optional TTL clock and `failClosed`).
122
187
 
@@ -135,7 +200,14 @@ xAI caching is automatic — there is no explicit cache-create/cache-store API c
135
200
 
136
201
  ## Pricing
137
202
 
138
- `XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-08-12'`) — a point-in-time capture from `/v1/models`, not a live lookup (ADR-005). Rates are in µUSD per million tokens:
203
+ `XAI_PRICING` is a frozen, versioned snapshot (`xaiPricingVersion: 'xai-2026-08-24'`) — a point-in-time capture from `/v1/models`, not a live lookup (ADR-005). Rates are in µUSD per million tokens. Tool invocations add `Cost.details.tools` (`microUsd = input + cached + output + tools`):
204
+
205
+ | Counter (raw `usage.details` key) | Rate |
206
+ | --------------------------------- | ---------- |
207
+ | `web_search_calls` | $5 / 1,000 |
208
+ | `x_search_calls` | $5 / 1,000 |
209
+
210
+ Enable Live Search with `providerOptions.xai.tools` (`web_search` / `x_search`). Citations land on `result.citations`. `countTokens` uses `POST /v1/tokenize-text` and returns `accuracy: 'lower-bound'` (text parts only; media / file parts are `bad_request`).
139
211
 
140
212
  | Model | Tier | Input | Cached input | Output |
141
213
  | ---------- | ---------------------------- | ------- | ------------ | -------- |