@gullabs/google 0.12.1 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -21,7 +21,7 @@ pnpm add @gullabs/google @gullabs/core @google/genai
21
21
  | `buildGoogleClient(auth)` | Builds the real `@google/genai` client from `AuthMaterial` |
22
22
  | `isGeminiCapacityError(err)` | Detects Gemini Flex shared-capacity errors for built-in fallback |
23
23
  | `geminiModelDescriptors`, `gemmaModelDescriptors`, `defaultGeminiRegistry` | Built-in model descriptors + pre-built registry |
24
- | `geminiPricingSource()`, `GEMINI_PRICING`, `TIER_FACTOR` | Built-in Gemini pricing snapshot and tier-factor map |
24
+ | `geminiPricingSource()`, `GEMINI_PRICING`, `resolveGeminiRates` | Built-in Gemini pricing snapshot (concrete standard / flex / batch rates) |
25
25
  | `GoogleFileStore` | Files API: upload + poll ACTIVE + delete |
26
26
  | `FileDeleteOptions` | `{ failClosed?, signal? }` — opt-in fail-closed delete (parity with `@gullabs/xai`) |
27
27
 
@@ -65,7 +65,7 @@ const result = await client.generate(
65
65
  - `output.jsonSchema` → `responseMimeType: 'application/json'` + verbatim `responseSchema` when native structured output is enabled; the engine returns parsed output and `outputParsed` without validating shape
66
66
  - `providerOptions.google.*` → typed provider-extension lane for admitted keys such as `cachedContent`, `safetySettings`, and exact tool declarations
67
67
  - Usage: `promptTokenCount`→`inputTokens`, `candidatesTokenCount`+`thoughtsTokenCount`→`outputTokens` (GROSS)
68
- - Errors: `401` and a bare `403` default to `invalid_auth`; `429`→`rate_limited`; `5xx`→`server`; timeouts; Gemini safety blocks are a 200-path `content_filter` (`promptFeedback.blockReason` / no candidates), not an HTTP 403
68
+ - Errors: `401` and a bare `403` default to `invalid_auth`; `429`→`rate_limited`; `5xx`→`server`; timeouts; Gemini safety blocks are a 200-path `content_filter` when `promptFeedback.blockReason` is set. A candidate-less 200 without a block reason is retryable `server`.
69
69
 
70
70
  ## Strict model-config expectations
71
71
 
@@ -80,10 +80,42 @@ descriptor boundary:
80
80
  library has not yet shipped the matching schema, pricing, served-tier
81
81
  recording, and tests.
82
82
 
83
- For structured output with built-in tools, follow the exact public
84
- `generateContent` evidence: the current docs only admit that combination for
85
- `gemini-3.1-pro-preview` and `gemini-3.5-flash`. Other models should fail early
86
- instead of relying on adapter repair or provider-side surprises.
83
+ The Developer API accepted structured JSON plus `googleSearch` on all registered
84
+ Gemini 3.x models in the 2026-09-26 live probes. The structured responses did
85
+ not include `groundingMetadata`, even when asked to search; callers must not
86
+ assume that an accepted tool means Search ran or that citations are available.
87
+
88
+ ## Registered models
89
+
90
+ | id | Efforts | SO + search | caching.minTokens | Tiers |
91
+ | ------------------------ | ------------------------------- | ----------- | ----------------- | -------------- |
92
+ | `gemini-2.5-pro` | `low`, `medium`, `high` | no | 2048 | flex, standard |
93
+ | `gemini-2.5-flash` | `none`, `low`, `medium`, `high` | no | 2048 | flex, standard |
94
+ | `gemini-2.5-flash-lite` | `none`, `low`, `medium`, `high` | no | 2048 | flex, standard |
95
+ | `gemini-3.1-pro-preview` | `low`, `medium`, `high` | yes | 1024 | flex, standard |
96
+ | `gemini-3.1-flash-lite` | `none`, `low`, `medium`, `high` | yes | 1024 | flex, standard |
97
+ | `gemini-3.5-flash-lite` | `none`, `low`, `medium`, `high` | yes | 1024 | flex, standard |
98
+ | `gemini-3.6-flash` | `none`, `low`, `medium`, `high` | yes | 1024 | flex, standard |
99
+ | `gemini-3.7-flash` | `low`, `medium`, `high` | yes | 1024 | flex, standard |
100
+ | `gemini-3.8-flash` | `low`, `medium`, `high` | yes | 1024 | flex, standard |
101
+ | `gemma-4-31b-it` | `none`, `high` | no | n/a | none |
102
+ | `gemma-4-26b-a4b-it` | `none`, `high` | no | n/a | none |
103
+
104
+ All six Gemini 3.x cache-create minimums above were live checked: 103 tokens
105
+ returned `min_total_token_count=1024`; exactly 1024 tokens succeeded.
106
+ Google's 4096-token table in the caching guide describes **implicit** caching;
107
+ this column is the **explicit cache-create** floor.
108
+
109
+ `gemini-3.7-flash` and `gemini-3.8-flash` never emit `thinkingLevel` MINIMAL.
110
+ `gemini-3-flash-preview` and `gemini-3.5-flash` are deleted and are not aliased.
111
+ Migrate both to `gemini-3.6-flash`. `servedServiceTier` reads the provider's
112
+ `usageMetadata.serviceTier` echo when present, then falls back to the tier
113
+ actually dispatched when the echo is absent.
114
+ An echo that differs from the requested tier emits a warning so callers can
115
+ see a provider-side remap. `flexFallback: false` disables the adapter's retry
116
+ at standard tier; it cannot prevent a provider-side remap after dispatch.
117
+ A candidate-less HTTP 200 without a safety block is a retryable provider error;
118
+ its reported usage and snapshot cost are saved on that failed attempt.
87
119
 
88
120
  ## Gemma 4
89
121