@gullabs/any-llm 0.4.1 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -2,6 +2,11 @@
2
2
 
3
3
  The default, batteries-included any-llm client package.
4
4
 
5
+ "Batteries-included" means concretely this: `@gullabs/any-llm` bundles `@gullabs/core` (the
6
+ engine), `@gullabs/google` (the Gemini adapter), and `@google/genai` (the Gemini SDK) as
7
+ dependencies and re-exports their full public API, so a single `pnpm add` gets you a working
8
+ client instead of three separate installs.
9
+
5
10
  ## Install
6
11
 
7
12
  ```bash
@@ -39,6 +44,7 @@ const summarize = defineCallSite({
39
44
  userTemplate: 'Summarize this:\n\n{{text}}',
40
45
  })
41
46
 
47
+ // Auth is required per call — the library never reads environment variables itself.
42
48
  const result = await client.runStructured(
43
49
  summarize,
44
50
  { text: documentText },
@@ -46,4 +52,32 @@ const result = await client.runStructured(
46
52
  )
47
53
  ```
48
54
 
55
+ ## Key exports
56
+
57
+ This package re-exports the full public API of `@gullabs/core` and `@gullabs/google` verbatim —
58
+ `createClient`, `defineCallSite`, `geminiAdapter`, `geminiPricingSource`, `LlmError`, and every
59
+ other named export from both packages. See their READMEs for details:
60
+
61
+ | Export | What it is |
62
+ | ----------------------- | ---------------------------------------------------------- |
63
+ | `createClient(config)` | Wires ports into a `{ generate, runStructured }` client |
64
+ | `defineCallSite(opts)` | Defines a typed, reusable prompt template bound to a model |
65
+ | `geminiAdapter(opts?)` | The Gemini `ProviderAdapter`, from `@gullabs/google` |
66
+ | `geminiPricingSource()` | Built-in Gemini pricing snapshot |
67
+ | `LlmError` | Typed error class — always thrown on call failure |
68
+ | `ANY_LLM_VERSION` | This package's version, sourced from `package.json` |
69
+
49
70
  Use `@gullabs/core` and `@gullabs/google` directly only when you want modular dependency control.
71
+
72
+ ## AI-agent skill
73
+
74
+ This package ships an [Agent Skill](https://code.claude.com/docs/en/skills) at
75
+ [`skills/any-llm/SKILL.md`](./skills/any-llm/SKILL.md) for AI coding tools that support the
76
+ convention (e.g. Claude Code). Point a compatible assistant at it to get accurate, up-to-date
77
+ usage guidance for this library instead of relying on its training data.
78
+
79
+ ## Learn more
80
+
81
+ - [Monorepo root README](../../README.md) — full architecture, auth model, and package overview
82
+ - [`docs/structured-output-validation.md`](../../docs/structured-output-validation.md) — validating `result.output` after `outputParsed`
83
+ - [`docs/multi-runtime.md`](../../docs/multi-runtime.md) — web route + Temporal worker integration pattern
package/dist/index.cjs CHANGED
@@ -6,7 +6,7 @@ var google = require('@gullabs/google');
6
6
  // src/index.ts
7
7
 
8
8
  // package.json
9
- var version = "0.4.1";
9
+ var version = "0.5.0";
10
10
 
11
11
  // src/index.ts
12
12
  var ANY_LLM_VERSION = version;
@@ -1 +1 @@
1
- {"version":3,"sources":["../package.json","../src/index.ts"],"names":[],"mappings":";;;;;;;;AAEE,IAAA,OAAA,GAAW,OAAA;;;ACcN,IAAM,eAAA,GAA0B","file":"index.cjs","sourcesContent":["{\n \"name\": \"@gullabs/any-llm\",\n \"version\": \"0.4.1\",\n \"description\": \"Batteries-included any-llm client for Gemini: engine, adapter, and Google SDK in one install.\",\n \"type\": \"module\",\n \"license\": \"Apache-2.0\",\n \"repository\": {\n \"type\": \"git\",\n \"url\": \"git+https://github.com/gullabs/any-llm.git\",\n \"directory\": \"packages/any-llm\"\n },\n \"main\": \"./dist/index.cjs\",\n \"module\": \"./dist/index.js\",\n \"types\": \"./dist/index.d.ts\",\n \"exports\": {\n \".\": {\n \"types\": \"./dist/index.d.ts\",\n \"import\": \"./dist/index.js\",\n \"require\": \"./dist/index.cjs\"\n }\n },\n \"files\": [\n \"dist\"\n ],\n \"scripts\": {\n \"build\": \"tsup\"\n },\n \"dependencies\": {\n \"@google/genai\": \"^1.45.0 || ^2\",\n \"@gullabs/core\": \"workspace:*\",\n \"@gullabs/google\": \"workspace:*\"\n },\n \"engines\": {\n \"node\": \">=20.9.0\"\n },\n \"sideEffects\": false,\n \"keywords\": [\n \"llm\",\n \"gemini\",\n \"google-genai\",\n \"ai\",\n \"tokens\",\n \"cost\",\n \"usage\",\n \"observability\",\n \"typescript\"\n ],\n \"publishConfig\": {\n \"access\": \"public\"\n },\n \"homepage\": \"https://github.com/gullabs/any-llm/tree/main/packages/any-llm#readme\",\n \"bugs\": \"https://github.com/gullabs/any-llm/issues\"\n}\n","/**\n * @gullabs/any-llm — batteries-included public entrypoint.\n *\n * This package is the default client install path. It re-exports the core\n * engine and Gemini adapter while depending on the Gemini SDK for a one-package\n * setup.\n *\n * @module\n */\n\nexport * from '@gullabs/core'\nexport * from '@gullabs/google'\n\nimport { version } from '../package.json'\n\n/** Library version, sourced from package.json at build time. */\nexport const ANY_LLM_VERSION: string = version\n"]}
1
+ {"version":3,"sources":["../package.json","../src/index.ts"],"names":[],"mappings":";;;;;;;;AAEE,IAAA,OAAA,GAAW,OAAA;;;ACcN,IAAM,eAAA,GAA0B","file":"index.cjs","sourcesContent":["{\n \"name\": \"@gullabs/any-llm\",\n \"version\": \"0.5.0\",\n \"description\": \"Batteries-included any-llm client for Gemini: engine, adapter, and Google SDK in one install.\",\n \"type\": \"module\",\n \"license\": \"Apache-2.0\",\n \"repository\": {\n \"type\": \"git\",\n \"url\": \"git+https://github.com/gullabs/any-llm.git\",\n \"directory\": \"packages/any-llm\"\n },\n \"main\": \"./dist/index.cjs\",\n \"module\": \"./dist/index.js\",\n \"types\": \"./dist/index.d.ts\",\n \"exports\": {\n \".\": {\n \"types\": \"./dist/index.d.ts\",\n \"import\": \"./dist/index.js\",\n \"require\": \"./dist/index.cjs\"\n }\n },\n \"files\": [\n \"dist\",\n \"skills\"\n ],\n \"scripts\": {\n \"build\": \"tsup\"\n },\n \"dependencies\": {\n \"@google/genai\": \"^1.45.0 || ^2\",\n \"@gullabs/core\": \"workspace:*\",\n \"@gullabs/google\": \"workspace:*\"\n },\n \"engines\": {\n \"node\": \">=20.9.0\"\n },\n \"sideEffects\": false,\n \"keywords\": [\n \"llm\",\n \"gemini\",\n \"google-genai\",\n \"ai\",\n \"tokens\",\n \"cost\",\n \"usage\",\n \"observability\",\n \"typescript\"\n ],\n \"publishConfig\": {\n \"access\": \"public\"\n },\n \"homepage\": \"https://github.com/gullabs/any-llm/tree/main/packages/any-llm#readme\",\n \"bugs\": \"https://github.com/gullabs/any-llm/issues\"\n}\n","/**\n * @gullabs/any-llm — batteries-included public entrypoint.\n *\n * This package is the default client install path. It re-exports the core\n * engine and Gemini adapter while depending on the Gemini SDK for a one-package\n * setup.\n *\n * @module\n */\n\nexport * from '@gullabs/core'\nexport * from '@gullabs/google'\n\nimport { version } from '../package.json'\n\n/** Library version, sourced from package.json at build time. */\nexport const ANY_LLM_VERSION: string = version\n"]}
package/dist/index.js CHANGED
@@ -4,7 +4,7 @@ export * from '@gullabs/google';
4
4
  // src/index.ts
5
5
 
6
6
  // package.json
7
- var version = "0.4.1";
7
+ var version = "0.5.0";
8
8
 
9
9
  // src/index.ts
10
10
  var ANY_LLM_VERSION = version;
package/dist/index.js.map CHANGED
@@ -1 +1 @@
1
- {"version":3,"sources":["../package.json","../src/index.ts"],"names":[],"mappings":";;;;;;AAEE,IAAA,OAAA,GAAW,OAAA;;;ACcN,IAAM,eAAA,GAA0B","file":"index.js","sourcesContent":["{\n \"name\": \"@gullabs/any-llm\",\n \"version\": \"0.4.1\",\n \"description\": \"Batteries-included any-llm client for Gemini: engine, adapter, and Google SDK in one install.\",\n \"type\": \"module\",\n \"license\": \"Apache-2.0\",\n \"repository\": {\n \"type\": \"git\",\n \"url\": \"git+https://github.com/gullabs/any-llm.git\",\n \"directory\": \"packages/any-llm\"\n },\n \"main\": \"./dist/index.cjs\",\n \"module\": \"./dist/index.js\",\n \"types\": \"./dist/index.d.ts\",\n \"exports\": {\n \".\": {\n \"types\": \"./dist/index.d.ts\",\n \"import\": \"./dist/index.js\",\n \"require\": \"./dist/index.cjs\"\n }\n },\n \"files\": [\n \"dist\"\n ],\n \"scripts\": {\n \"build\": \"tsup\"\n },\n \"dependencies\": {\n \"@google/genai\": \"^1.45.0 || ^2\",\n \"@gullabs/core\": \"workspace:*\",\n \"@gullabs/google\": \"workspace:*\"\n },\n \"engines\": {\n \"node\": \">=20.9.0\"\n },\n \"sideEffects\": false,\n \"keywords\": [\n \"llm\",\n \"gemini\",\n \"google-genai\",\n \"ai\",\n \"tokens\",\n \"cost\",\n \"usage\",\n \"observability\",\n \"typescript\"\n ],\n \"publishConfig\": {\n \"access\": \"public\"\n },\n \"homepage\": \"https://github.com/gullabs/any-llm/tree/main/packages/any-llm#readme\",\n \"bugs\": \"https://github.com/gullabs/any-llm/issues\"\n}\n","/**\n * @gullabs/any-llm — batteries-included public entrypoint.\n *\n * This package is the default client install path. It re-exports the core\n * engine and Gemini adapter while depending on the Gemini SDK for a one-package\n * setup.\n *\n * @module\n */\n\nexport * from '@gullabs/core'\nexport * from '@gullabs/google'\n\nimport { version } from '../package.json'\n\n/** Library version, sourced from package.json at build time. */\nexport const ANY_LLM_VERSION: string = version\n"]}
1
+ {"version":3,"sources":["../package.json","../src/index.ts"],"names":[],"mappings":";;;;;;AAEE,IAAA,OAAA,GAAW,OAAA;;;ACcN,IAAM,eAAA,GAA0B","file":"index.js","sourcesContent":["{\n \"name\": \"@gullabs/any-llm\",\n \"version\": \"0.5.0\",\n \"description\": \"Batteries-included any-llm client for Gemini: engine, adapter, and Google SDK in one install.\",\n \"type\": \"module\",\n \"license\": \"Apache-2.0\",\n \"repository\": {\n \"type\": \"git\",\n \"url\": \"git+https://github.com/gullabs/any-llm.git\",\n \"directory\": \"packages/any-llm\"\n },\n \"main\": \"./dist/index.cjs\",\n \"module\": \"./dist/index.js\",\n \"types\": \"./dist/index.d.ts\",\n \"exports\": {\n \".\": {\n \"types\": \"./dist/index.d.ts\",\n \"import\": \"./dist/index.js\",\n \"require\": \"./dist/index.cjs\"\n }\n },\n \"files\": [\n \"dist\",\n \"skills\"\n ],\n \"scripts\": {\n \"build\": \"tsup\"\n },\n \"dependencies\": {\n \"@google/genai\": \"^1.45.0 || ^2\",\n \"@gullabs/core\": \"workspace:*\",\n \"@gullabs/google\": \"workspace:*\"\n },\n \"engines\": {\n \"node\": \">=20.9.0\"\n },\n \"sideEffects\": false,\n \"keywords\": [\n \"llm\",\n \"gemini\",\n \"google-genai\",\n \"ai\",\n \"tokens\",\n \"cost\",\n \"usage\",\n \"observability\",\n \"typescript\"\n ],\n \"publishConfig\": {\n \"access\": \"public\"\n },\n \"homepage\": \"https://github.com/gullabs/any-llm/tree/main/packages/any-llm#readme\",\n \"bugs\": \"https://github.com/gullabs/any-llm/issues\"\n}\n","/**\n * @gullabs/any-llm — batteries-included public entrypoint.\n *\n * This package is the default client install path. It re-exports the core\n * engine and Gemini adapter while depending on the Gemini SDK for a one-package\n * setup.\n *\n * @module\n */\n\nexport * from '@gullabs/core'\nexport * from '@gullabs/google'\n\nimport { version } from '../package.json'\n\n/** Library version, sourced from package.json at build time. */\nexport const ANY_LLM_VERSION: string = version\n"]}
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@gullabs/any-llm",
3
- "version": "0.4.1",
3
+ "version": "0.5.0",
4
4
  "description": "Batteries-included any-llm client for Gemini: engine, adapter, and Google SDK in one install.",
5
5
  "type": "module",
6
6
  "license": "Apache-2.0",
@@ -20,12 +20,13 @@
20
20
  }
21
21
  },
22
22
  "files": [
23
- "dist"
23
+ "dist",
24
+ "skills"
24
25
  ],
25
26
  "dependencies": {
26
27
  "@google/genai": "^1.45.0 || ^2",
27
- "@gullabs/core": "0.4.1",
28
- "@gullabs/google": "0.5.0"
28
+ "@gullabs/core": "0.4.2",
29
+ "@gullabs/google": "0.5.1"
29
30
  },
30
31
  "engines": {
31
32
  "node": ">=20.9.0"
@@ -0,0 +1,260 @@
1
+ ---
2
+ name: any-llm
3
+ description: >-
4
+ Guidance for writing, reviewing, or debugging TypeScript code that calls
5
+ @gullabs/any-llm, @gullabs/core, or @gullabs/google to talk to Gemini models.
6
+ Applies when adding a new LLM call site, wiring createClient/generate/runStructured,
7
+ defining a defineCallSite prompt template, requesting structured JSON output,
8
+ catching or narrowing an LlmError, configuring reasoning/thinking budgets, or wiring
9
+ a UsageSink for cost tracking. Also applies whenever the user mentions any-llm, the
10
+ Gemini adapter, Gemini Flex tier, structured-output validation, or per-call auth for
11
+ this library. Covers the mandatory per-call `{ auth: { apiKey } }` pattern (there is
12
+ no env-var or ambient auth), the caller-owned output-validation contract, and the
13
+ reject-don't-map error philosophy — the things a developer used to other LLM SDKs
14
+ would otherwise get wrong by default.
15
+ ---
16
+
17
+ # any-llm
18
+
19
+ Typed, provider-agnostic-by-design (currently Gemini-only) LLM call engine with cost
20
+ tracking, retries, rate limiting, structured output, and per-call observability.
21
+
22
+ Three packages:
23
+
24
+ - `@gullabs/core` — engine (`createClient`), types, errors, `defineCallSite`.
25
+ - `@gullabs/google` — Gemini adapter (`geminiAdapter`) over `@google/genai`.
26
+ - `@gullabs/any-llm` — batteries-included: re-exports both of the above plus `@google/genai` as a dependency.
27
+
28
+ Install `@gullabs/any-llm` for a one-package setup, or the two modular packages for
29
+ explicit dependency control. Import names are identical either way.
30
+
31
+ ## #1 gotcha: auth is per-call, always
32
+
33
+ **There is no env-var auth, no ambient/singleton auth, no `AuthProvider` port.** Every
34
+ `generate()` and `runStructured()` call requires `opts.auth = { apiKey: string }`
35
+ explicitly. `createClient()` itself takes no credentials.
36
+
37
+ ```ts
38
+ // WRONG — GenerateOptions.auth is a required field; this will not type-check, and if
39
+ // bypassed with `as any` it throws LlmError({ kind: 'invalid_auth' }) before any I/O.
40
+ const client = createClient({
41
+ adapters: [geminiAdapter()],
42
+ pricing: geminiPricingSource(),
43
+ })
44
+ await client.generate(request, {} as GenerateOptions)
45
+
46
+ // RIGHT — bring the key from wherever your app resolves it, pass it on every call.
47
+ const auth = { apiKey: myResolvedGeminiKey }
48
+ await client.generate(request, { auth })
49
+ ```
50
+
51
+ A missing, empty, or non-string `apiKey` throws `LlmError` with `kind: 'invalid_auth'`
52
+ before any network call is made. Vertex AI (ADC/service-account auth) is **not
53
+ supported** — it was removed and is a roadmap item only; do not write code assuming
54
+ `{ vertex: { project, location } }` works.
55
+
56
+ ## Quickstart
57
+
58
+ ```ts
59
+ import { createClient, geminiPricingSource, geminiAdapter } from '@gullabs/any-llm'
60
+ // (or: from '@gullabs/core' / '@gullabs/google' respectively, if using modular install)
61
+
62
+ const client = createClient({
63
+ adapters: [geminiAdapter()],
64
+ pricing: geminiPricingSource(),
65
+ })
66
+
67
+ const result = await client.generate(
68
+ {
69
+ model: 'gemini-2.5-flash',
70
+ messages: [{ role: 'user', parts: [{ kind: 'text', text: 'Hello!' }] }],
71
+ },
72
+ { auth: { apiKey: myResolvedGeminiKey } },
73
+ )
74
+
75
+ console.log(result.text) // raw text
76
+ console.log(result.usage) // { inputTokens, outputTokens, cachedInputTokens?, thinkingTokens?, details, raw }
77
+ console.log(result.cost?.microUsd) // integer micro-USD, or undefined if unpriced
78
+ ```
79
+
80
+ `Message.parts` is `TextPart | InlineMediaPart | FileUriPart` — multimodal input mixes
81
+ `{ kind: 'text', text }`, `{ kind: 'inline-media', mimeType, data /* raw base64, no data: prefix */ }`,
82
+ and `{ kind: 'file-uri', uri, mimeType }` freely in one `parts` array.
83
+
84
+ ## `defineCallSite` — reusable prompt templates
85
+
86
+ ```ts
87
+ import { defineCallSite } from '@gullabs/core'
88
+
89
+ const summarize = defineCallSite({
90
+ id: 'summarize-article', // persisted as callSiteId on every record
91
+ model: 'gemini-2.5-flash',
92
+ system: 'You are a concise summarizer.',
93
+ userTemplate: 'Summarize this article in 3 sentences:\n\n{{article}}',
94
+ config: { temperature: 0.3 },
95
+ })
96
+
97
+ const result = await client.runStructured(summarize, { article: text }, { auth })
98
+ ```
99
+
100
+ `{{var}}` interpolation is non-recursive (substituted values are never re-scanned for
101
+ further `{{...}}`, preventing template injection) and applies to both `system` and
102
+ `userTemplate`. A missing var is left as the literal `{{var}}` placeholder, not an
103
+ empty string. `runStructured` also accepts a two-arg form, `(callSite, opts)`, when the
104
+ template has no vars. Config resolution order everywhere is
105
+ `clientDefaults → callSite.config → opts.config` (later wins; objects like `reasoning`
106
+ and `providerOptions` merge key-by-key, everything else is last-write-wins).
107
+
108
+ ## Structured output — auth + validation together
109
+
110
+ `request.output = { jsonSchema }` (or `callSite.jsonSchema`) is forwarded to the
111
+ provider as a **hint**, not enforced by the library. The engine JSON-parses the
112
+ response and sets `outputParsed`; `result.output` is always `unknown`. **The caller
113
+ owns shape validation** — this library does not validate output shape itself.
114
+
115
+ ```ts
116
+ import { createClient, geminiPricingSource, geminiAdapter } from '@gullabs/any-llm'
117
+ import type { StandardSchemaV1 } from '@gullabs/core'
118
+
119
+ const client = createClient({
120
+ adapters: [geminiAdapter()],
121
+ pricing: geminiPricingSource(),
122
+ })
123
+
124
+ const result = await client.generate(
125
+ {
126
+ model: 'gemini-2.5-flash',
127
+ messages: [{ role: 'user', parts: [{ kind: 'text', text: 'Rate this PR 1-10.' }] }],
128
+ output: {
129
+ jsonSchema: {
130
+ type: 'object',
131
+ properties: { rating: { type: 'number' }, summary: { type: 'string' } },
132
+ required: ['rating', 'summary'],
133
+ },
134
+ },
135
+ },
136
+ { auth: { apiKey: myResolvedGeminiKey } }, // auth is still required
137
+ )
138
+
139
+ // 1. Gate on outputParsed first — cheap boolean signal.
140
+ if (result.outputParsed !== true) {
141
+ // provider didn't return parseable JSON — retry, escalate, or fall back
142
+ } else {
143
+ // 2. Then validate shape with any Standard Schema v1 validator (zod, valibot, etc.
144
+ // all implement '~standard'). Do NOT trust result.output's shape without this.
145
+ const validated = await mySchema['~standard'].validate(result.output)
146
+ if ('issues' in validated && validated.issues !== undefined) {
147
+ // shape invalid but JSON-parseable — inspect validated.issues
148
+ } else {
149
+ // validated.value is now typed
150
+ }
151
+ }
152
+ ```
153
+
154
+ Full caller-owned validation helper: `docs/structured-output-validation.md` in the
155
+ source repo (see "For more detail" below — this file may not ship in `node_modules`).
156
+
157
+ ## Error handling
158
+
159
+ Every rejection from `generate()` / `runStructured()` is an `LlmError` with a `.kind`
160
+ discriminant (from `packages/core/src/errors.ts`):
161
+
162
+ | `kind` | Meaning | Retryable |
163
+ | ---------------- | ----------------------------------------------------- | --------- |
164
+ | `invalid_auth` | 401/403, or missing/empty `apiKey` | No |
165
+ | `rate_limited` | 429 — provider quota exceeded | Yes |
166
+ | `server` | 5xx — transient provider error | Yes |
167
+ | `timeout` | exceeded `config.timeoutMs` or a network timeout | Yes |
168
+ | `aborted` | caller's `AbortSignal` fired | No |
169
+ | `bad_request` | 400/422, or a request the library rejected before I/O | No |
170
+ | `content_filter` | provider refused output for safety reasons | No |
171
+ | `unknown` | uncategorized — inspect `.cause` | No |
172
+
173
+ ```ts
174
+ import { LlmError } from '@gullabs/core'
175
+
176
+ try {
177
+ const result = await client.generate(request, { auth })
178
+ } catch (e) {
179
+ if (e instanceof LlmError) {
180
+ if (e.retryable) scheduleRetry(e.retryAfterMs)
181
+ else if (e.kind === 'invalid_auth') /* surface a credentials error */
182
+ else throw e
183
+ } else {
184
+ throw e // never expected — the engine always throws LlmError
185
+ }
186
+ }
187
+ ```
188
+
189
+ `LlmError` also carries `httpStatus?`, `retryAfterMs?`, `provider?`, `callId?`,
190
+ `attemptId?`, `servedServiceTier?`, and `cause` (the original thrown value).
191
+
192
+ ## Reject, don't map
193
+
194
+ Bad input or config throws `bad_request` (or `invalid_auth`) **before any I/O** —
195
+ nothing is silently clamped, coerced, or defaulted around a typo. Examples already
196
+ enforced by the engine/adapter: an unrouteable `model` string, a config value that
197
+ fails a model's schema (`validateConfig`), a `reasoning.budgetTokens` set on a model
198
+ whose API only supports `reasoning.effort`, duplicate adapter/middleware `id`s, and
199
+ grounding (`providerOptions.google.tools` search) combined with `output.jsonSchema` in
200
+ the same request (mutually exclusive on Gemini — see `docs/grounded-structured.md`).
201
+
202
+ **Do not add defensive fallback/clamping code around this library.** If a call throws
203
+ `bad_request`, fix the input — do not catch-and-retry with a "safer" guessed value; the
204
+ library is telling you the config is invalid, not transiently rejected.
205
+
206
+ ## Reasoning / thinking budgets
207
+
208
+ ```ts
209
+ config: {
210
+ reasoning: { effort: 'medium', includeThoughts: true }
211
+ }
212
+ ```
213
+
214
+ `ReasoningEffort` is `'none' | 'low' | 'medium' | 'high'`. Two provider APIs exist
215
+ under the hood: Gemini 2.5 models take a token `budgetTokens`; Gemini 3.x / Gemma 4
216
+ models take a discrete `effort` level (`thinkingLevel`). `EFFORT_BUDGET` (from
217
+ `@gullabs/core`) is the shared effort→token mapping (`none: 0, low: 1024, medium: 8192,
218
+ high: 24576`) and `resolveReasoning({ model, budgetTokens, registry })` — also exported
219
+ from `@gullabs/core` — buckets a raw token budget into the nearest admitted effort
220
+ level for a given model, throwing `bad_request` if the model doesn't support reasoning
221
+ at all or has no admitted effort at or above the requested bucket. Prefer setting
222
+ `reasoning.effort` directly for level-API models rather than a token budget; the
223
+ adapter throws `bad_request` if you pass `budgetTokens` to a level-API model.
224
+
225
+ ## Rate limiting and cost tracking
226
+
227
+ - Pre-send backpressure is a `RateLimiter` port (`ClientConfig.rateLimiter`); default
228
+ is a no-op. `@gullabs/core` ships a dependency-free `inMemoryRateLimiter`, and the
229
+ companion `@gullabs/quota` package provides shared/cross-process quota primitives —
230
+ see that package's README for setup.
231
+ - Every call computes `result.cost` (micro-USD) via the configured `PricingSource`
232
+ (`geminiPricingSource()`), and, when `sink` is configured on `createClient`, persists
233
+ a full per-attempt record (usage, cost, warnings, error classification) — fail-open,
234
+ so a broken sink never fails the LLM call. See `docs/ledger.md` for the record shape
235
+ and the `@gullabs/drizzle` package for a ready-made Postgres `UsageSink`.
236
+
237
+ ## Common mistakes
238
+
239
+ - Forgetting `opts.auth` on a `generate()`/`runStructured()` call — it is required on
240
+ every call, not just once at `createClient()` time.
241
+ - Assuming `process.env.GEMINI_API_KEY` (or similar) is read automatically — it is
242
+ never read by this library; the host must resolve and pass the key itself.
243
+ - Assuming `result.output`'s shape is validated — it is `unknown`; validate it yourself
244
+ (Standard Schema v1) after checking `result.outputParsed === true`.
245
+ - Assuming Vertex AI is a supported target — it currently is not (roadmap only); only
246
+ the Gemini Developer API (API-key auth) is supported.
247
+ - Catching a `bad_request` `LlmError` and retrying with a clamped/guessed value instead
248
+ of fixing the call — the library never silently coerces invalid config.
249
+ - Setting `reasoning.budgetTokens` on a `thinkingLevel`-API model (Gemini 3.x / Gemma 4) — use `reasoning.effort` instead; the adapter throws `bad_request` otherwise.
250
+
251
+ ## For more detail
252
+
253
+ These docs live in the source repo and may not ship in the installed `node_modules`
254
+ copy — the auth + structured-output example above is intentionally self-contained.
255
+
256
+ - `README.md` — full quickstart, install options, Gemma 4 notes.
257
+ - `docs/structured-output-validation.md` — full caller-owned validation helper.
258
+ - `docs/multi-runtime.md` — client construction across request/response vs. worker code.
259
+ - `docs/ledger.md` — persisted record field reference.
260
+ - `docs/grounded-structured.md` — grounding + structured output two-call pattern.