@gullabs/google 0.13.0 → 0.16.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/NOTICE ADDED
@@ -0,0 +1,6 @@
1
+ any-llm
2
+ Copyright 2026 Gul Labs
3
+
4
+ This product includes software developed at Gul Labs (https://github.com/gul-labs).
5
+
6
+ Licensed under the Apache License, Version 2.0. See LICENSE.
package/README.md CHANGED
@@ -8,32 +8,47 @@ Gemini provider adapter for any-llm. A thin mapping layer over `@google/genai` t
8
8
  pnpm add @gullabs/google @gullabs/core @google/genai
9
9
  ```
10
10
 
11
- **Peer dependency:** `@google/genai ^1 || ^2`
11
+ **Peer dependency:** `@google/genai ^2`
12
12
 
13
13
  ## Key exports
14
14
 
15
- | Export | What it is |
16
- | -------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |
17
- | `googleProvider(opts?)` | `ProviderPlugin` factory — bundles the adapter, model descriptors, and pricing source |
18
- | `geminiAdapter(opts?)` | Creates the `ProviderAdapter` for Gemini |
19
- | `GeminiAdapterOptions` | `{ client?: GeminiClientLike }` — inject a pre-built or fake client |
20
- | `GeminiClientLike` | Structural interface the adapter depends on (satisfied by real SDK and fakes) |
21
- | `buildGoogleClient(auth)` | Builds the real `@google/genai` client from `AuthMaterial` |
22
- | `isGeminiCapacityError(err)` | Detects Gemini Flex shared-capacity errors for built-in fallback |
23
- | `geminiModelDescriptors`, `gemmaModelDescriptors`, `defaultGeminiRegistry` | Built-in model descriptors + pre-built registry |
24
- | `geminiPricingSource()`, `GEMINI_PRICING`, `resolveGeminiRates` | Built-in Gemini pricing snapshot (concrete standard / flex / batch rates) |
25
- | `GoogleFileStore` | Files API: upload + poll ACTIVE + delete |
26
- | `FileDeleteOptions` | `{ failClosed?, signal? }` — opt-in fail-closed delete (parity with `@gullabs/xai`) |
15
+ | Export | What it is |
16
+ | -------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
17
+ | `googleProvider(opts?)` | `ProviderPlugin` factory — bundles the adapter, model descriptors, and pricing source |
18
+ | `geminiAdapter(opts?)` | Creates the `ProviderAdapter` for Gemini |
19
+ | `GeminiAdapterOptions` | `{ client?: GeminiClientLike }` — inject a pre-built or fake client |
20
+ | `GeminiClientLike` | Structural interface the adapter depends on (satisfied by real SDK and fakes) |
21
+ | `buildGoogleClient(auth)` | Builds the real `@google/genai` client from `AuthMaterial` |
22
+ | `isGeminiCapacityError(err)` | Detects Gemini Flex capacity errors (HTTP 503 only) for fallback |
23
+ | `classifyGoogleError(err)` | The classifier every Gemini error goes through (`@gullabs/testing`'s fakes call it for you) |
24
+ | `GEMINI_INPUT_MIME_TYPES` | The media types a Gemini file upload and a Gemini request admit (`text/*`, `image/*`, `application/pdf`, ...) |
25
+ | `geminiModelDescriptors`, `gemmaModelDescriptors`, `defaultGeminiRegistry` | Built-in model descriptors + pre-built registry |
26
+ | `geminiPricingSource()`, `GEMINI_PRICING`, `resolveGeminiRates` | Built-in Gemini pricing snapshot (concrete standard / flex rates, audio input rates where published) |
27
+ | `GoogleFileStore` | Files API: upload + poll ACTIVE + delete |
28
+ | `FileDeleteOptions` | `{ failClosed?, signal? }` — opt-in fail-closed delete (parity with `@gullabs/xai`) |
27
29
 
28
30
  ## File store delete modes
29
31
 
30
- `GoogleFileStore.delete` defaults to **fail-open** (errors → `onDeleteError`, resolve). Pass `{ failClosed: true }` when the host gates durable state on known success; HTTP/SDK not-found remains success (idempotent). Empty `handle.name` always throws `bad_request`.
32
+ `GoogleFileStore.upload` takes an `AbortSignal` (an abort releases the caller at once, also while a poll request is pending, and the polling deadline ends a stalled poll request too; bytes already sent may still be stored by Google), keeps Google's own `File.error` when a file ends `FAILED` (a transient status code, `DEADLINE_EXCEEDED`, `INTERNAL` or `UNAVAILABLE`, is a retryable `server` error; any other is `bad_request`), and its polling timeout is `kind: 'server'`, **not retryable** (the upload succeeded; a retry would upload again and orphan the file). Store errors (`GoogleFileStore`, `GoogleCacheStore`) are classified exactly like `generate()` errors, with `provider: 'google'`.
33
+
34
+ `GoogleFileStore` takes a `scheduler` (the timer source of the poll wait; pass the client's `FakeClock` in tests, with `now: () => clock.now()` for the poll timeout), and the Gemini flex/standard client-side ceiling runs on the engine's `scheduler` too, so a `FakeClock` fires both. The default poll wait is cleared the moment the deadline or an abort ends the upload first, so no timer outlives `upload()`. `sleep` still replaces the poll wait wholesale, and a host's own `sleep` cannot be cancelled: a timer it sets stays pending until it fires.
35
+
36
+ `GoogleFileStore.delete` defaults to **fail-open** (errors → `onDeleteError`, resolve). Pass `{ failClosed: true }` when the host gates durable state on known success; not-found remains success (idempotent): HTTP 404, `NOT_FOUND`, and the HTTP 403 Google sends for an unknown or expired file id ("You do not have permission to access the File ... or it may not exist"). That 403 wording comes from public bug reports, not from a live capture here, and a real permission failure carries the same status, so the shape is treated as gone. `GoogleCacheStore.delete` follows the same rule (a 404, or the captured 403 `CachedContent not found`, is success and never reaches `onDeleteError`). Empty `handle.name` always throws `bad_request`.
31
37
 
32
38
  ```ts
39
+ import type { GoogleFileHandle, GoogleFileStore } from '@gullabs/google'
40
+
41
+ declare const store: GoogleFileStore
42
+ declare const handle: GoogleFileHandle
43
+
33
44
  await store.delete(handle) // fail-open
34
45
  await store.delete(handle, { failClosed: true }) // throw on non-not-found failure
35
46
  ```
36
47
 
48
+ ## Model lifecycle
49
+
50
+ `gemini-3.1-flash-lite` carries `shutdownDate: '2027-05-07'` (replacement: `gemini-3.5-flash-lite`), read on Google's [deprecations page](https://ai.google.dev/gemini-api/docs/deprecations) (last updated 2026-10-01) on 2026-10-03. From 2027-02-06 on, the first successful call to it on a client has a `{ type: 'shutdown' }` warning saying so (once per client, not on every call). Google also limits Gemini 2.5 access to users who have used those models; no shutdown date is announced for them, so none is set. Gemma 4 is not on the page.
51
+
37
52
  ## Quick example
38
53
 
39
54
  ```ts
@@ -55,17 +70,282 @@ const result = await client.generate(
55
70
  )
56
71
  ```
57
72
 
73
+ ## Function calling and Gemini 3 thought signatures
74
+
75
+ Gemini 3.x returns an opaque `thoughtSignature` on the first function call of each model turn
76
+ (and sometimes on a text part) and answers HTTP 400 if a replayed function call has lost it
77
+ (live capture, 2026-10-03: all six registered 3.x models; Google signs only the **first** call of a
78
+ parallel set, and every sequential step needs its own). The library handles this without keeping a
79
+ copy of your history. Continuation is `'history'`: append `result.message`, send the full history,
80
+ and pass `result.transientProviderState` back.
81
+
82
+ ```ts
83
+ import { composeProviders, createClient } from '@gullabs/core'
84
+ import type { JsonValue, Message } from '@gullabs/core'
85
+ import { googleProvider } from '@gullabs/google'
86
+
87
+ const client = createClient({ ...composeProviders([googleProvider()]) })
88
+
89
+ const tools = [
90
+ {
91
+ name: 'get_weather',
92
+ description: 'Current weather for a city',
93
+ inputJsonSchema: { type: 'object', properties: { city: { type: 'string' } } },
94
+ },
95
+ ]
96
+ const base = { provider: 'google', model: 'gemini-3.6-flash', tools } as const
97
+ let messages: Message[] = [
98
+ { role: 'user', parts: [{ kind: 'text', text: 'Weather in Paris?' }] },
99
+ ]
100
+ let state: JsonValue | undefined
101
+
102
+ for (;;) {
103
+ const result = await client.generate(
104
+ {
105
+ ...base,
106
+ messages,
107
+ ...(state !== undefined ? { transientProviderState: state } : {}),
108
+ },
109
+ { auth: { apiKey: 'YOUR_GEMINI_API_KEY' } },
110
+ )
111
+ if (result.toolCalls === undefined) break // result.text is the answer
112
+
113
+ const toolResults = {
114
+ role: 'user' as const,
115
+ parts: result.toolCalls.map((c) => ({
116
+ kind: 'tool-result' as const,
117
+ toolCallId: c.toolCallId,
118
+ toolName: c.toolName,
119
+ result: { tempC: 18 }, // run the tool; any JSON value (non-objects are wrapped as { output })
120
+ })),
121
+ }
122
+ messages = [...messages, result.message, toolResults] // unedited
123
+ state = result.transientProviderState // the signature overlay; carries every earlier entry
124
+ }
125
+ ```
126
+
127
+ `result.transientProviderState` is `{ google: { signatures: [{ messageIndex, partIndex, kind, model,
128
+ partSha256, signature }] } }`, an **overlay** that says which part of _your_ messages gets which
129
+ signature. `kind` is the signed part's kind (`'text'` or `'tool-call'`); `messageIndex` indexes the
130
+ messages the adapter receives (a middleware that adds or removes messages shifts them). `partSha256`
131
+ is the SHA-256 of the part's RFC 8785 canonical JSON, so an edited text or tool argument is
132
+ detected and key order does not matter (history stored in Postgres `jsonb` still verifies). Persist
133
+ it with the history it belongs to; it contains opaque provider tokens, not prompt text, and is never
134
+ written to the ledger.
135
+
136
+ **What is rejected and what is dropped.** A function-call signature is required, so the next request
137
+ is `bad_request` before dispatch when a function-call entry is stale (the call was edited, reordered
138
+ or removed, the message moved, or the entry was issued for another model string; signatures are not
139
+ replayed across models, and a declared alias is a different string from its canonical id), when an
140
+ assistant message with tool calls has no entry for its first call (this is also what rejects history
141
+ produced by another provider, or by Gemini 2.5), when two entries name the same part, or when the
142
+ state is malformed. A **text** signature is optional (Google accepts the next turn without it), so a
143
+ stale text entry (the text was trimmed, edited or moved, the message is gone, or it was issued for
144
+ another model) is dropped with a result warning and nothing else is lost; the dropped entry is not
145
+ carried into the next state. A host that `.trim()`s the final answer, or rebuilds it from
146
+ `result.text`, keeps working.
147
+
148
+ Google validates the signature only on function calls in the current turn (live capture, 2026-10-03:
149
+ an unsigned call in an older turn was accepted on `gemini-3.1-pro-preview` and `gemini-3.8-flash`).
150
+ The library still requires an entry for every replayed tool-call message, so history without
151
+ signatures never reaches Google by accident; keep a conversation that began on another provider on
152
+ that provider.
153
+
154
+ The library does not offer Google's dummy signature that bypasses validation: it degrades quality and
155
+ is a [BACKLOG](../../BACKLOG.md) item, not a default. Gemini 2.5 and Gemma need none of this.
156
+
157
+ ### Trimming, compacting and rewinding the history
158
+
159
+ The overlay is addressed by message index, so removing messages from your history moves what the
160
+ indices mean. After you remove messages, pass their positions (in the history the state was issued
161
+ for) to `dropMessagesFromSignatureState(state, indices)`. It drops the entries of the removed
162
+ messages, shifts the later `messageIndex`es down, and returns `undefined` when nothing is left:
163
+
164
+ ```ts
165
+ import type { JsonValue, Message } from '@gullabs/core'
166
+ import { dropMessagesFromSignatureState } from '@gullabs/google'
167
+
168
+ let messages: Message[] = []
169
+ let state: JsonValue | undefined // the previous result's transientProviderState
170
+
171
+ // Front-trim: drop the oldest turn (messages 0-3) to fit the context window.
172
+ messages = messages.slice(4)
173
+ state = dropMessagesFromSignatureState(state, [0, 1, 2, 3])
174
+
175
+ // Rewind: undo the last tool step (messages 5-6), then ask again.
176
+ messages = messages.slice(0, 5)
177
+ state = dropMessagesFromSignatureState(state, [5, 6])
178
+ ```
179
+
180
+ The rule is whole turns only: remove a tool-call message together with its tool-result message, and
181
+ never keep a message that holds a function call while removing its entry. Compaction (replacing
182
+ several messages by one summary) works the same way if the summary takes the place of the range's first
183
+ message and you pass the rest of the range: replacing messages 0-3 with one user summary is
184
+ `messages = [summary, ...messages.slice(4)]` and `dropMessagesFromSignatureState(state, [1, 2, 3])`
185
+ (message 0 is a user message and never has an entry). Do not insert messages in front of kept ones; that
186
+ shifts indices upward and the helper only shifts them down. Without the helper, a stale function-call
187
+ entry is `bad_request`.
188
+
189
+ ### Tool-call ids
190
+
191
+ Gemini returns a `functionCall.id` (`call_<number>` on the Developer API in the live capture); the
192
+ library uses it as `toolCallId` and sends it back on the `functionCall` and `functionResponse`. If a
193
+ response has no id, the library synthesizes `anyllm_call_<name>_<n>`, unique among the ids already in
194
+ your history, and **never sends it** to Gemini (a response is matched to its call by name and order,
195
+ which is what Gemini documents). Gemini 3 accepted a replay with provider ids, without ids, with
196
+ synthesized ids and with the same id repeated across steps (capture
197
+ `__fixtures__/function-call-ids-2026-10-03.json`).
198
+
199
+ ### Empty responses and `countTokens`
200
+
201
+ A response with nothing to replay (the model produced only thoughts, for example when
202
+ `maxOutputTokens` was spent on reasoning) has `result.message.parts === []`. Do not append it to your
203
+ history: an assistant message with no parts is `bad_request`. Retry the call instead.
204
+
205
+ `countTokens` sends the history without signatures (Gemini accepts that), and `generate()` bills each
206
+ replayed signature (up to about 110 prompt tokens each, and model-dependent: 0 on some models). So
207
+ when the counted history holds function calls on a Gemini 3 model, `accuracy` is `'estimated'`, not
208
+ `'exact'`: the real count can be higher by up to one signature per signed call or text part. Without function calls, or on Gemini 2.5, it is
209
+ `'exact'`.
210
+
211
+ Provider output that cannot be hashed never fails a call that was billed: a lone surrogate in a
212
+ function call's arguments or in text returns the result with a warning naming the part and no
213
+ signature entry for it (the next turn is then `bad_request`, because the function call lacks its
214
+ signature), and `-0` is treated as `0`.
215
+
216
+ `geminiContentToMessages({ contents, model })` imports signatures from hand-authored
217
+ `@google/genai` history into the same overlay, returned as `transientProviderState` beside `messages`; `model` is required
218
+ when any part carries one.
219
+
220
+ ## Thinking budgets and the output cap
221
+
222
+ `maxOutputTokens` **includes thinking tokens**. A model that thinks for the whole cap returns no answer
223
+ (`finishReason: 'length'`, empty text, `thinkingTokens > 0`), and the library warns about it after the fact
224
+ ("maxOutputTokens (M) was used up by reasoning (T tokens); no answer was produced"). Gemini 2.5 models
225
+ take a thinking budget, and the adapter turns `reasoning.effort` into one:
226
+
227
+ | `reasoning.effort` | `thinkingBudget` | Note |
228
+ | ------------------ | ---------------- | --------------------------------------------------------------------------- |
229
+ | `none` | 0 | Thinking off; only the models that admit `none` (2.5 Flash, 2.5 Flash-Lite) |
230
+ | `low` | 1,024 | |
231
+ | `medium` | 8,192 | |
232
+ | `high` | 24,576 | The Flash maximum, and 75% of 2.5 Pro's 32,768 |
233
+
234
+ `xhigh` and `max` are rejected: they have no Gemini budget. `reasoning.budgetTokens` sets the budget
235
+ directly (128 to 32,768 on 2.5 Pro, per its schema) and cannot be combined with `effort`. Gemini 3.x
236
+ and Gemma 4 use `thinkingLevel` and have no token budget.
237
+
238
+ When a budget model's `thinkingBudget` (from `effort` or `budgetTokens`) is **at or above**
239
+ `maxOutputTokens`, the result carries a warning ("thinkingBudget (B) is not below maxOutputTokens (M);
240
+ thinking may consume the whole cap and leave no answer"). It is a warning, not a rejection: Google says actual
241
+ thinking can under- or overflow the budget, so the combination is a risk, not an invalid request. The
242
+ common trap is `effort: 'high'` (24,576) with a small `maxOutputTokens`.
243
+
244
+ Gemini 3.x models take a level, not a budget, so there is nothing to compare with the cap. The adapter warns
245
+ when `reasoning.effort` is `high` and `maxOutputTokens` is below 4,096 (measured: thinking reached 4,000 tokens
246
+ in 7 of 72 3.x calls at `high`, up to 8,859). It does not warn at `low` or `medium` (no 3.x call reached 4,000
247
+ tokens there, and a smaller figure would rest on a prompt-dependent tail the sample cannot place), nor when
248
+ `reasoning` is omitted (the default level was not measured). Warnings never reject.
249
+
250
+ Measured thinking is prompt-driven and heavy-tailed: `docs/thinking-token-distribution.md` has p50, p95
251
+ and max per model and effort (336 calls, 2026-10-03; the raw records are kept outside this repository). At `high`, p95 was 2.5 to 6 times the p50, and the largest value was
252
+ 8,859 tokens, so size `maxOutputTokens` for the answer **plus** thinking: under 4,096 is unsafe at `high`
253
+ and 1,024 was used up by thinking in most `high` calls.
254
+
255
+ ## Model limits and input media types
256
+
257
+ Every descriptor states `limits: { contextWindow, maxOutputTokens }`. The Gemini model pages
258
+ (`ai.google.dev/gemini-api/docs/models/<id>`, read 2026-10-03) give 1,048,576 input and 65,536 output tokens
259
+ for every registered Gemini model, and their config schemas cap `maxOutputTokens` at 65,536 (a larger value is
260
+ `bad_request` before dispatch). The Gemma 4 model card gives a 256K window (262,144) and **no output limit**, so
261
+ Gemma's `limits.maxOutputTokens` is `null`: no figure is invented, its schema applies no cap, and Google decides
262
+ what it accepts.
263
+
264
+ `capabilities.inputMimeTypes` is what a model takes in `inline-media` and `file-uri` parts; anything else, and an
265
+ empty type, is `bad_request` naming `messages[i].parts[j]` before dispatch (and in `countTokens`). Matching
266
+ ignores case and `; parameters` (`IMAGE/PNG`, `text/plain; charset=utf-8` pass) and the string you sent goes to
267
+ Google unchanged.
268
+
269
+ - **Gemini:** `application/pdf` and the families `text/*`, `image/*`, `audio/*`, `video/*`. Google lists image
270
+ types (PNG, JPEG, WebP, HEIC, HEIF), audio and video types, but publishes no closed list for documents: its
271
+ document page says PDF is understood natively and "you can pass other MIME types for document understanding,
272
+ like TXT, Markdown, HTML, XML, etc." (extracted as plain text). The library therefore admits the documented
273
+ families by prefix and leaves a type inside a family that Google does not take (for example an image format
274
+ it does not decode) to Google's own error. `application/json`, `application/xml` and other `application/*`
275
+ types are in no documented family and are rejected. A YouTube URL is a `file-uri` part: give it any
276
+ `video/*` type (for example `video/mp4`).
277
+ - **Gemma 4:** `image/*` and `video/*`. The model card lists image input and video as frames (up to 60
278
+ seconds at one frame per second) for the 31B and 26B A4B models and names no media types; audio input
279
+ belongs to other Gemma sizes. That the Gemini API's Gemma endpoint takes a video part has not been probed.
280
+ - **`GoogleFileStore.upload`** applies the same Gemini rule (one shared function), so an empty or unadmitted type
281
+ is `bad_request` before any bytes are sent, and a file that uploads can be used in `generate`.
282
+
58
283
  ## What it maps
59
284
 
60
285
  - `serviceTier: 'flex'` → Gemini Flex service tier when the model descriptor supports it
61
286
  - omitted `serviceTier` → provider-default request behavior
62
287
  - `reasoning.includeThoughts` → `thinkingConfig.includeThoughts`; thought parts become `reasoningText`
63
288
  - `reasoning.effort` → `thinkingBudget` (Gemini 2.5) or `thinkingLevel` (Gemini 3 / Gemma 4)
64
- - `reasoning.budgetTokens` → admitted only on Gemini 2.5 budget-api models; strict descriptors reject it on level-api models
65
- - `output.jsonSchema` → `responseMimeType: 'application/json'` + verbatim `responseSchema` when native structured output is enabled; the engine returns parsed output and `outputParsed` without validating shape
289
+ - `reasoning.budgetTokens` → admitted only on Gemini 2.5 budget-api models; strict descriptors reject it on level-api models (see "Thinking budgets and the output cap")
290
+ - `output.jsonSchema` → `responseMimeType: 'application/json'` + verbatim `responseJsonSchema` when native structured output is enabled, and `tools[].inputJsonSchema` → `parametersJsonSchema` (both standard JSON Schema, in your key order; see "JSON Schema" below); the engine returns parsed output and `outputParsed` without validating shape
66
291
  - `providerOptions.google.*` → typed provider-extension lane for admitted keys such as `cachedContent`, `safetySettings`, and exact tool declarations
67
292
  - Usage: `promptTokenCount`→`inputTokens`, `candidatesTokenCount`+`thoughtsTokenCount`→`outputTokens` (GROSS)
68
- - Errors: `401` and a bare `403` default to `invalid_auth`; `429`→`rate_limited`; `5xx`→`server`; timeouts; Gemini safety blocks are a 200-path `content_filter` when `promptFeedback.blockReason` is set. A candidate-less 200 without a block reason is retryable `server`.
293
+ - Errors: `401` and a bare `403` default to `invalid_auth`; `429`→`rate_limited`; `5xx`→`server`; timeouts; Gemini safety blocks are a 200-path `content_filter` when `promptFeedback.blockReason` is set, and so is an output filter stop (`SAFETY`, `RECITATION`, `BLOCKLIST`, `PROHIBITED_CONTENT`, `SPII`, `IMAGE_*`) that produced no text and no complete tool call (not retryable, usage attached; a stop that kept partial text is a success with `finishReason: 'content_filter'`). A function call is complete only when the candidate finished with `STOP` (or no finish reason): beside `MAX_TOKENS` (cut by the output cap), a filter stop or any other finish it is dropped from `toolCalls` and from `result.message`, a warning names it, and `finishReason` is `length`, `content_filter` or `other`, so a tool loop never runs a call Google stopped. A candidate-less 200 without a block reason is retryable `server`, unless it billed reasoning tokens (the output cap was spent on thinking: the same request fails the same way), which is not retryable. A call that reaches the adapter's own client-side ceiling (5 minutes standard, 25 minutes flex, when no `timeoutMs` is set) or the SDK's transport timer is `timeout`, not retryable, `reason: 'transport_timeout'`, so `retryMiddleware` makes one attempt and the ledger holds one row. The structured body adds: `RetryInfo.retryDelay` → `retryAfterMs`; a per-day quota (`QuotaFailure` quota id containing `PerDay`) → `rate_limited`, not retryable, `reason: 'daily_quota'`; `API_KEY_INVALID` / `API_KEY_EXPIRED` (Google sends the first as HTTP 400) → `invalid_auth`; a stale `cachedContent` (HTTP 403, "CachedContent not found") → `bad_request`, `reason: 'cache_not_found'`. `retryMiddleware` (default `maxDelayMs` 60 s) sleeps a typical per-minute `retryDelay` and retries; a delay over 60 s stops the retry and the 429 surfaces with `retryAfterMs` for a scheduler (raise `maxDelayMs` to wait longer in process).
294
+ - `providerOptions.google.safetySettings` → `category` is one of `HARM_CATEGORY_HARASSMENT`, `_HATE_SPEECH`, `_SEXUALLY_EXPLICIT`, `_DANGEROUS_CONTENT`, `_CIVIC_INTEGRITY`, `_JAILBREAK` and `threshold` one of `HARM_BLOCK_THRESHOLD_UNSPECIFIED`, `BLOCK_LOW_AND_ABOVE`, `BLOCK_MEDIUM_AND_ABOVE`, `BLOCK_ONLY_HIGH`, `BLOCK_NONE`, `OFF` (Google's safety-settings guide, dated 2026-09-17); anything else is `bad_request` before dispatch.
295
+ - `providerOptions.google.cachedContent` cannot be sent with `system`, `tools` or `providerOptions.google.tools`: Gemini needs them stored in the cache, so pass `tools` / `toolConfig` (and `systemInstruction`) to `GoogleCacheStore.create` instead. The adapter rejects the combination before dispatch. `cachedContent` is admitted only on the models that cache explicitly (Gemini 2.5 and 3.x; Gemma 4 has no caching capability and rejects it). It takes the cache name, or `{ cacheName: handle.cacheName, toolKinds: handle.toolKinds }` from a `GoogleCacheHandle` (not the whole handle: the schema is strict and the type refuses it) (`handle.toolKinds` lists the kinds of tool the cache holds, empty when none); the request sent to Google carries only the name. A cache that holds `googleSearch` sends no search tool in the request, so a handle whose `toolKinds` includes it marks the call as a Search call (priced like a sent tool); with a bare name the library cannot know, so a `groundingMetadata` in the response is the evidence that Search ran (see "Search facts, grounding price and `requireGrounding`").
296
+ - Inline media is checked against Google's request limits before dispatch: an inline PDF over 50 MB (matched on the media type whatever its case or `; parameters`), or a request whose inline data and text exceed 100 MB, is `bad_request`; upload it with `GoogleFileStore` and send a `file-uri` part. The result's `providerMetadata.google.candidate` holds the candidate's raw `finishReason`, `finishMessage`, `safetyRatings`, `citationMetadata` and `urlContextMetadata` when Google sent them.
297
+
298
+ ## JSON Schema
299
+
300
+ `output.jsonSchema` and `tools[].inputJsonSchema` are standard JSON Schema, sent as
301
+ `responseJsonSchema` and `parametersJsonSchema` exactly as you wrote them, key order included (put
302
+ `reasoning` before `answer` and the model generates in that order). There is no OpenAPI
303
+ `responseSchema` / `parameters` path, and `nullable` and uppercase type names are rejected.
304
+
305
+ Google accepts every keyword and silently ignores the ones it does not support, so the adapter
306
+ only lets through the ones it enforces and rejects the rest with `bad_request` and the JSON path
307
+ before dispatch (ADR-034; live probe on every Gemini and Gemma model, 2026-10-03, and Google's
308
+ structured-output guide read the same day):
309
+
310
+ - **Enforced:** `type` (a type array only as one type plus `'null'`), `properties`, `required`,
311
+ `additionalProperties` (boolean or schema), `enum`, `anyOf`, `$ref` / `$defs` (recursive
312
+ schemas too), `items`, `prefixItems` (and `items: false` to close a tuple), `minItems` /
313
+ `maxItems`, `minimum` / `maximum`, and `format` for `date-time`, `date` and `email` (the
314
+ values the probe exercised; `time` is named in Google's guide but no capture exercised it, so
315
+ it is rejected).
316
+ - **Accepted, but soft:** `pattern`, `minLength` and `maxLength` are supported by Google yet obeyed
317
+ only probabilistically (the probe saw violations on several models, at worst 4 of 7 samples).
318
+ They are not guarantees: validate `output` yourself. The library never validates the result.
319
+ `pattern` is held to a regex subset (no backreferences, property escapes, word boundaries,
320
+ lookaround or inline modifiers) because no capture shows Google enforcing them.
321
+ - **Rejected, because Google ignores them:** `const`, `allOf`, `exclusiveMinimum` /
322
+ `exclusiveMaximum`, `multipleOf`, `uniqueItems`, and `oneOf`, which Google reads as `anyOf`.
323
+ `not`, `if`/`then`/`else`, `minProperties` and other `format` values are outside the enforced
324
+ set too, and so is any `propertyNames` except `{ type: 'string' }` (what
325
+ `z.record(z.string(), X)` emits; it constrains nothing and is accepted).
326
+ - **Gemma 4 is stricter.** It ignored `format` (7 of 7 samples on both models) and
327
+ `minLength` / `maxLength` (7 of 7 and 6 of 7), so those three keywords are rejected on a Gemma
328
+ model. The profile follows the resolved model descriptor, so a declared alias of a Gemma model
329
+ gets it too.
330
+ - **Annotations** (`$schema`, `$id`, `$comment`, `title`, `description`, `examples`, `default`,
331
+ `deprecated`, `readOnly`, `writeOnly`) are always accepted.
332
+ - **Malformed schemas** (a value in a schema position that is not a schema, `maxLength: '3000'`,
333
+ an invalid `pattern`, a cyclic JavaScript object, nesting deeper than 128) are `bad_request`
334
+ with the path, before dispatch.
335
+
336
+ **Tool schemas rest on the output-schema probe.** The live probe (P3) ran `responseJsonSchema`
337
+ only; the only live `parametersJsonSchema` evidence is two trivial schemas (an object with one
338
+ string property and `additionalProperties: false`, and an empty `properties`) on the six 3.x
339
+ models. `$schema`, `$ref` / `$defs`, `anyOf`, `items: false` and type arrays are verified for
340
+ output schemas only, and the same profile is applied to tools. Do not read tool-schema acceptance
341
+ beyond trivial schemas as live-verified.
342
+
343
+ Zod: `z.literal('x')` emits `const`, which Google ignores. Write `z.enum(['x'])` instead; the
344
+ library does not rewrite it for you. `z.discriminatedUnion` emits `oneOf`; use `z.union`, which
345
+ emits `anyOf`. `startsWith`, `endsWith` and `includes` emit a non-standard `format` next to a
346
+ `pattern`: chain `.meta({ format: undefined })` after the check, or write `z.string().regex(...)`.
347
+ To lint a schema against what both Gemini 3.x and xAI enforce, call `assertPortableJsonSchema`
348
+ from `@gullabs/core` (it does not cover Gemma's stricter profile).
69
349
 
70
350
  ## Strict model-config expectations
71
351
 
@@ -81,9 +361,113 @@ descriptor boundary:
81
361
  recording, and tests.
82
362
 
83
363
  The Developer API accepted structured JSON plus `googleSearch` on all registered
84
- Gemini 3.x models in the 2026-09-26 live probes. The structured responses did
85
- not include `groundingMetadata`, even when asked to search; callers must not
86
- assume that an accepted tool means Search ran or that citations are available.
364
+ Gemini 3.x models in the 2026-09-26 live probes, but the structured responses did
365
+ not include `groundingMetadata` even when asked to search, and Flash-Lite often
366
+ skipped Search. An accepted request is not proof that Search ran, so
367
+ `structuredOutputWithTools` is `false` on every descriptor: the combination fails
368
+ with `bad_request` before dispatch. Make two calls instead (grounded research, then
369
+ structured synthesis); see [`docs/grounded-structured.md`](../../docs/grounded-structured.md).
370
+ To send both in one call anyway, set `providerOptions.google.allowSchemaWithSearch: true`.
371
+ That also turns on `requireGrounding` (override with `requireGrounding: false`), because a
372
+ schema'd call can skip Search without saying so. The opt-in exists only for the Gemini 3.x models,
373
+ the only ones with a capture: on Gemini 2.5 and Gemma the pair is rejected with or without it.
374
+ The rule covers Search held by a cache too: `cachedContent: { cacheName, toolKinds: ['googleSearch'] }`
375
+ with a schema needs the same opt-in. A bare cache name is not blocked (its contents are unknown); if the
376
+ response of such a schema call reports search queries it is returned with a warning, priced as `estimated`.
377
+
378
+ ### Search facts, grounding price and `requireGrounding`
379
+
380
+ A call that sends `googleSearch` reports two facts in `usage.details`:
381
+ `web_search_requested` (`1`) and `web_search_calls`, the number of queries in
382
+ `groundingMetadata.webSearchQueries` counted as occurrences (a repeated query counts each
383
+ time; absent when the response has no metadata or no query list). `tool_use_prompt` records
384
+ `toolUsePromptTokenCount` when Google reports it (Gemini 2.5). Those tokens are not priced: Google's
385
+ pricing page says retrieved search results are not charged as input tokens (read 2026-10), but no live billing reconciliation has confirmed it, so the total mismatch they cause still
386
+ marks the cost `'estimated'` (ADR-035).
387
+
388
+ Search counts as requested when the request sends `googleSearch` or the `cachedContent` handle lists
389
+ it in `toolKinds`. When neither says so (a bare cache name, or no cache at all) and the response
390
+ carries `groundingMetadata`, the response is the evidence: `web_search_requested` is `1`,
391
+ `web_search_calls` is the observed query count (left out when the metadata names none), the fee is
392
+ priced from those queries, the cost is `'estimated'` and a warning says the request did not declare
393
+ `googleSearch`. A grounded response is never priced `'exact'` just because the request was silent.
394
+
395
+ The pricing source puts the grounding fee on `cost.details.tools`: Gemini 3 bills per query
396
+ (`web_search_calls × $0.014`), Gemini 2.5 per grounded prompt (`$0.035`, once however many
397
+ queries ran), from Google's pricing page read 2026-10-03. A call that ran Search is always
398
+ `cost.confidence: 'estimated'`: Google's free allowance (as read 2026-10, 5,000
399
+ requests per month shared across Gemini 3.x and 1,500 per day on Gemini 2.5) is shared across a
400
+ project, so no single call can know it was free, and every fee is charged in full. When Search was
401
+ requested but the count is unknown, the tools lane is `0`, the cost is estimated and a warning
402
+ says so. Google's billing of repeated queries and of tool-use tokens is not established.
403
+
404
+ ### Audio input and cache storage
405
+
406
+ Gemini 2.5 Flash, 2.5 Flash-Lite and 3.1 Flash-Lite bill audio input at a higher rate than text, image and
407
+ video (cached audio apart from cached text), on both the standard and the flex tier. The adapter records the
408
+ prompt's per-modality counts from `usageMetadata.promptTokensDetails` and `cacheTokensDetails` as
409
+ `usage.details.input_<modality>` and `cached_<modality>` (`input_audio`, `cached_audio`, ...), and the
410
+ pricing source bills the audio tokens at the audio rates and every other token at the text rate. The other
411
+ models list one rate for all modalities. The rates are from Google's pricing page, read 2026-10-03
412
+ (page last updated 2026-10-01). A request that carries audio but whose response reports no `AUDIO` tokens
413
+ (the entry is absent or `{ AUDIO, 0 }`) carries a warning, and on a model that prices audio apart its cost
414
+ is `'estimated'` (it can understate); the warning and the estimate read the same predicate. Cached tokens
415
+ are priced as text only when the audio share of the cache is known: a `cacheTokensDetails` that lists no
416
+ audio and covers every cached token records `cached_audio: 0` (a text cache beside new audio stays
417
+ `'exact'`), a listing that covers fewer tokens than were cached leaves it unknown, and cached tokens with
418
+ neither `promptTokensDetails` nor `cacheTokensDetails` cannot rule out audio in the cached content, so they
419
+ carry a warning and the cost is `'estimated'`. A reported cached `AUDIO` count is priced at the cached audio
420
+ rate even when `promptTokensDetails` is absent. Counts that contradict each other (cached audio above the
421
+ cache, prompt lanes above the prompt) are clamped, warned about and `'estimated'`.
422
+ There is no batch tier: `'batch'` is an unpriced tier.
423
+
424
+ `GoogleCacheStore.create` and `getOrCreate` return a handle with `totalTokenCount`, the create response's
425
+ `usageMetadata.totalTokenCount`. Cache storage is billed per token-hour and appears in no usage record;
426
+ price it from that count and the time you keep the cache. The handle also carries `toolKinds` (the kinds
427
+ of tool given to `create`, empty when none); pass `{ cacheName, toolKinds }` as `cachedContent` so a
428
+ cached `googleSearch` is priced as Search. `ttlSeconds` must be a positive integer (`bad_request`
429
+ otherwise, before any call); an `expireTime` Google sends that does not parse falls back to now plus
430
+ the TTL, so a live cache is reused; and `getOrCreate` drops an expired entry from its in-process map
431
+ when its key is asked for again.
432
+
433
+ `providerOptions.google.requireGrounding: true` fails the call unless the response proves Search
434
+ ran (`groundingMetadata` with at least one non-empty query): a `server` error with
435
+ `reason: 'grounding_missing'` and the attempt's usage attached, so the billed tokens reach the
436
+ ledger. It is `retryable: true` only when no response schema is attached (4 of 4 captured calls
437
+ grounded); with a schema the same request keeps missing, so it is `retryable: false` and a retry
438
+ middleware makes one billed attempt, not three. Use the two-call recipe, or override `shouldRetry`
439
+ and accept the spend. The check applies only to a candidate that finished with `STOP`: a filtered
440
+ candidate (`SAFETY`, `RECITATION`, `BLOCKLIST`, `PROHIBITED_CONTENT`, `IMAGE_SAFETY`) with no evidence
441
+ throws `content_filter` (not retryable), and a `MAX_TOKENS` or other abnormal finish with no evidence is
442
+ not judged: it returns, with `finishReason: 'length'` (or `other`), so "a response without proof throws"
443
+ holds only for a normally finished candidate. It needs
444
+ `googleSearch` in the same request. `result.cost` of a call that succeeded on a retry covers the
445
+ last attempt only; the earlier attempts' spend is in their ledger rows.
446
+
447
+ `result.citations` entries carry `cited` (a `groundingSupports` segment points at the source) and
448
+ `textRange` (the first supported span of `result.text`, UTF-16 offsets). Gemini measures its segment
449
+ offsets in UTF-8 bytes into the answer part (verified on a live Japanese and emoji answer), and its
450
+ `partIndex` does not count thought parts. Every range is checked against `segment.text`: when the
451
+ answer at the converted range is not exactly that text, the range is dropped, the source stays
452
+ `cited: true`, and the result carries a warning, so a `textRange` is never a guess.
453
+
454
+ Google requires a grounded answer to display its Search Suggestions: the widget is at
455
+ `result.providerMetadata.google.searchEntryPoint` and is stored only there (the raw
456
+ `providerMetadata.groundingMetadata` omits it, so persisted rows hold the HTML once). Its
457
+ `renderedContent` is HTML and CSS that Google generates around model-chosen query strings: treat it
458
+ as untrusted markup. Render it in a sandboxed `<iframe>` (for example `sandbox` with no
459
+ `allow-scripts` and `srcdoc`), never inject it into your page's DOM with `innerHTML`.
460
+
461
+ `countTokens` counts `messages`, `system` and `tools`. The SDK's Developer API method cannot carry
462
+ `system` or `tools`, so with either present the library calls the REST `countTokens` with a full
463
+ `generateContentRequest` (a messages-only count still goes through the SDK). The count covers
464
+ `messages`, `system` and `tools` only: it carries no response schema, thinking config, `toolConfig`,
465
+ safety settings or Search tool, which `generate()` also sends, so it is not the whole prompt `generate()`
466
+ bills (whether Google bills schema tokens as prompt tokens is unverified). Tool schemas are held to the
467
+ same JSON Schema profile as `generate()`. `cachedContent` is not part of a count. The SDK client and
468
+ the REST count both go to one pinned endpoint (`https://generativelanguage.googleapis.com`); the SDK's
469
+ `GOOGLE_GEMINI_BASE_URL` environment override is not honoured, and a count's timeout is the
470
+ `timeoutMs` of `client.countTokens` (the abort signal the adapter receives carries it).
87
471
 
88
472
  ## Registered models
89
473
 
@@ -92,12 +476,12 @@ assume that an accepted tool means Search ran or that citations are available.
92
476
  | `gemini-2.5-pro` | `low`, `medium`, `high` | no | 2048 | flex, standard |
93
477
  | `gemini-2.5-flash` | `none`, `low`, `medium`, `high` | no | 2048 | flex, standard |
94
478
  | `gemini-2.5-flash-lite` | `none`, `low`, `medium`, `high` | no | 2048 | flex, standard |
95
- | `gemini-3.1-pro-preview` | `low`, `medium`, `high` | yes | 1024 | flex, standard |
96
- | `gemini-3.1-flash-lite` | `none`, `low`, `medium`, `high` | yes | 1024 | flex, standard |
97
- | `gemini-3.5-flash-lite` | `none`, `low`, `medium`, `high` | yes | 1024 | flex, standard |
98
- | `gemini-3.6-flash` | `none`, `low`, `medium`, `high` | yes | 1024 | flex, standard |
99
- | `gemini-3.7-flash` | `low`, `medium`, `high` | yes | 1024 | flex, standard |
100
- | `gemini-3.8-flash` | `low`, `medium`, `high` | yes | 1024 | flex, standard |
479
+ | `gemini-3.1-pro-preview` | `low`, `medium`, `high` | no | 1024 | flex, standard |
480
+ | `gemini-3.1-flash-lite` | `none`, `low`, `medium`, `high` | no | 1024 | flex, standard |
481
+ | `gemini-3.5-flash-lite` | `none`, `low`, `medium`, `high` | no | 1024 | flex, standard |
482
+ | `gemini-3.6-flash` | `none`, `low`, `medium`, `high` | no | 1024 | flex, standard |
483
+ | `gemini-3.7-flash` | `low`, `medium`, `high` | no | 1024 | flex, standard |
484
+ | `gemini-3.8-flash` | `low`, `medium`, `high` | no | 1024 | flex, standard |
101
485
  | `gemma-4-31b-it` | `none`, `high` | no | n/a | none |
102
486
  | `gemma-4-26b-a4b-it` | `none`, `high` | no | n/a | none |
103
487
 
@@ -113,17 +497,44 @@ Migrate both to `gemini-3.6-flash`. `servedServiceTier` reads the provider's
113
497
  actually dispatched when the echo is absent.
114
498
  An echo that differs from the requested tier emits a warning so callers can
115
499
  see a provider-side remap. `flexFallback: false` disables the adapter's retry
116
- at standard tier; it cannot prevent a provider-side remap after dispatch.
117
- A candidate-less HTTP 200 without a safety block is a retryable provider error;
118
- its reported usage and snapshot cost are saved on that failed attempt.
500
+ at standard tier (it fires only on an HTTP 503; a Flex 429 is an ordinary rate limit and honours `RetryInfo`); it cannot prevent a provider-side remap after dispatch.
501
+ A candidate-less HTTP 200 without a safety block is a retryable provider error, except when it billed
502
+ reasoning tokens (not retryable: the cap was spent on thinking); its reported usage and snapshot cost are
503
+ saved on that failed attempt.
504
+
505
+ A 200 that carries no `usageMetadata` is unknown usage, not a free call: `usage.details.usage_missing`
506
+ is `1`, the usage is recorded as zero tokens, a warning says so, and the pricing source reports the cost
507
+ unpriced (`microUsd: null`, `confidence: 'estimated'`, with an `unpricedReason`), never an exact $0.
508
+
509
+ `providerOptions.google.httpOptions.timeout` is the SDK's own timer. It is rejected above 2147483647 ms
510
+ (a Node timer that long fires after 1 ms and the call would abort at once), and with `timeoutMs` set it
511
+ must be at least `timeoutMs + 5000`, because a shorter SDK timer would end the call before the engine's
512
+ deadline with a raw SDK abort instead of the clean `timeout`. A flex call the adapter sends again at the
513
+ standard tier (HTTP 503) carries a warning saying so; with no `timeoutMs` the standard attempt then
514
+ runs under the 300000 ms client-side ceiling, not the 25-minute flex one. A retry never gives a request
515
+ that named no `serviceTier` one: only a request that asked for a tier is pinned to the tier an attempt
516
+ was served at, and a failed attempt's ledger row keeps the tier it asked for (`serviceTier`) beside the
517
+ tier it was served at (`servedServiceTier`).
518
+
519
+ `providerMetadata.groundingMetadata` and `providerMetadata.promptFeedback` are copied with the same
520
+ bounds as `providerMetadata.google.candidate` (50 list entries, 2048 characters per string, 8 levels,
521
+ a warning when something is cut); the citations are built from the full response. `GEMINI_PRICING`
522
+ is deep-frozen.
119
523
 
120
524
  ## Gemma 4
121
525
 
122
526
  The default registry includes two API-verified Gemma 4 models: `gemma-4-31b-it`
123
527
  and `gemma-4-26b-a4b-it`. Both route through this adapter and support:
124
528
 
125
- - **Native structured output** — `responseMimeType` + verbatim `responseSchema` are sent
126
- automatically when `output.jsonSchema` is set.
529
+ - **Native structured output** — `responseMimeType` + verbatim `responseJsonSchema` are sent
530
+ automatically when `output.jsonSchema` is set. Gemma ignored `format`, `minLength` and
531
+ `maxLength` in live probes, so those three keywords are rejected for Gemma models. **Caveat:**
532
+ Gemma wrapped 67 of 162 schema answers (41%) in a markdown ` ```json ` fence in the live probe
533
+ (2026-10-03), which is not parseable JSON: `outputParsed` is `false` and `result.text` holds the fenced
534
+ text. The adapter does not unwrap it (that would be repair, not rejection); it adds a warning
535
+ `gemma_fenced_json` naming the cause. Hosts that use Gemma for structured output should expect
536
+ about four in ten calls to need handling: call again, or strip a single enclosing fence from
537
+ `result.text` themselves and parse it. Gemini models were not seen to fence.
127
538
  - **Grounding** — `tools:[{googleSearch:{}}]` via `providerOptions.google`.
128
539
  - **Vision** — `inline-media` and `file-uri` multimodal message parts.
129
540
  - **Thinking** — `reasoning.effort` maps to `thinkingLevel` (`reasoningApi: 'level'`).
@@ -134,6 +545,10 @@ and `gemma-4-26b-a4b-it`. Both route through this adapter and support:
134
545
  (rejected by the API with HTTP 400).
135
546
  - **Tunable sampling** — `temperature`, `topP`, `topK` are accepted.
136
547
 
548
+ Gemma 4 has no caching capability: `providerOptions.google.cachedContent` is rejected for it, and
549
+ `allowSchemaWithSearch` is not in its schema (schema plus Search is rejected on Gemma with or
550
+ without it). Google's pricing page marks grounding "Not available" for Gemma 4 (read 2026-10); the `grounding: true` flag follows the live capture instead.
551
+
137
552
  This follows the library-wide **reject, don't map** rule: unsupported or incorrect input throws a
138
553
  typed `bad_request` `LlmError` at validation time rather than being silently clamped or coerced
139
554
  into something the model happens to accept.