@gullabs/any-llm 0.4.1 → 0.5.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +34 -0
- package/dist/index.cjs +1 -1
- package/dist/index.cjs.map +1 -1
- package/dist/index.js +1 -1
- package/dist/index.js.map +1 -1
- package/package.json +5 -4
- package/skills/any-llm/SKILL.md +260 -0
package/README.md
CHANGED
|
@@ -2,6 +2,11 @@
|
|
|
2
2
|
|
|
3
3
|
The default, batteries-included any-llm client package.
|
|
4
4
|
|
|
5
|
+
"Batteries-included" means concretely this: `@gullabs/any-llm` bundles `@gullabs/core` (the
|
|
6
|
+
engine), `@gullabs/google` (the Gemini adapter), and `@google/genai` (the Gemini SDK) as
|
|
7
|
+
dependencies and re-exports their full public API, so a single `pnpm add` gets you a working
|
|
8
|
+
client instead of three separate installs.
|
|
9
|
+
|
|
5
10
|
## Install
|
|
6
11
|
|
|
7
12
|
```bash
|
|
@@ -39,6 +44,7 @@ const summarize = defineCallSite({
|
|
|
39
44
|
userTemplate: 'Summarize this:\n\n{{text}}',
|
|
40
45
|
})
|
|
41
46
|
|
|
47
|
+
// Auth is required per call — the library never reads environment variables itself.
|
|
42
48
|
const result = await client.runStructured(
|
|
43
49
|
summarize,
|
|
44
50
|
{ text: documentText },
|
|
@@ -46,4 +52,32 @@ const result = await client.runStructured(
|
|
|
46
52
|
)
|
|
47
53
|
```
|
|
48
54
|
|
|
55
|
+
## Key exports
|
|
56
|
+
|
|
57
|
+
This package re-exports the full public API of `@gullabs/core` and `@gullabs/google` verbatim —
|
|
58
|
+
`createClient`, `defineCallSite`, `geminiAdapter`, `geminiPricingSource`, `LlmError`, and every
|
|
59
|
+
other named export from both packages. See their READMEs for details:
|
|
60
|
+
|
|
61
|
+
| Export | What it is |
|
|
62
|
+
| ----------------------- | ---------------------------------------------------------- |
|
|
63
|
+
| `createClient(config)` | Wires ports into a `{ generate, runStructured }` client |
|
|
64
|
+
| `defineCallSite(opts)` | Defines a typed, reusable prompt template bound to a model |
|
|
65
|
+
| `geminiAdapter(opts?)` | The Gemini `ProviderAdapter`, from `@gullabs/google` |
|
|
66
|
+
| `geminiPricingSource()` | Built-in Gemini pricing snapshot |
|
|
67
|
+
| `LlmError` | Typed error class — always thrown on call failure |
|
|
68
|
+
| `ANY_LLM_VERSION` | This package's version, sourced from `package.json` |
|
|
69
|
+
|
|
49
70
|
Use `@gullabs/core` and `@gullabs/google` directly only when you want modular dependency control.
|
|
71
|
+
|
|
72
|
+
## AI-agent skill
|
|
73
|
+
|
|
74
|
+
This package ships an [Agent Skill](https://code.claude.com/docs/en/skills) at
|
|
75
|
+
[`skills/any-llm/SKILL.md`](./skills/any-llm/SKILL.md) for AI coding tools that support the
|
|
76
|
+
convention (e.g. Claude Code). Point a compatible assistant at it to get accurate, up-to-date
|
|
77
|
+
usage guidance for this library instead of relying on its training data.
|
|
78
|
+
|
|
79
|
+
## Learn more
|
|
80
|
+
|
|
81
|
+
- [Monorepo root README](../../README.md) — full architecture, auth model, and package overview
|
|
82
|
+
- [`docs/structured-output-validation.md`](../../docs/structured-output-validation.md) — validating `result.output` after `outputParsed`
|
|
83
|
+
- [`docs/multi-runtime.md`](../../docs/multi-runtime.md) — web route + Temporal worker integration pattern
|
package/dist/index.cjs
CHANGED
package/dist/index.cjs.map
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"sources":["../package.json","../src/index.ts"],"names":[],"mappings":";;;;;;;;AAEE,IAAA,OAAA,GAAW,OAAA;;;ACcN,IAAM,eAAA,GAA0B","file":"index.cjs","sourcesContent":["{\n \"name\": \"@gullabs/any-llm\",\n \"version\": \"0.
|
|
1
|
+
{"version":3,"sources":["../package.json","../src/index.ts"],"names":[],"mappings":";;;;;;;;AAEE,IAAA,OAAA,GAAW,OAAA;;;ACcN,IAAM,eAAA,GAA0B","file":"index.cjs","sourcesContent":["{\n \"name\": \"@gullabs/any-llm\",\n \"version\": \"0.5.1\",\n \"description\": \"Batteries-included any-llm client for Gemini: engine, adapter, and Google SDK in one install.\",\n \"type\": \"module\",\n \"license\": \"Apache-2.0\",\n \"repository\": {\n \"type\": \"git\",\n \"url\": \"git+https://github.com/gullabs/any-llm.git\",\n \"directory\": \"packages/any-llm\"\n },\n \"main\": \"./dist/index.cjs\",\n \"module\": \"./dist/index.js\",\n \"types\": \"./dist/index.d.ts\",\n \"exports\": {\n \".\": {\n \"types\": \"./dist/index.d.ts\",\n \"import\": \"./dist/index.js\",\n \"require\": \"./dist/index.cjs\"\n }\n },\n \"files\": [\n \"dist\",\n \"skills\"\n ],\n \"scripts\": {\n \"build\": \"tsup\"\n },\n \"dependencies\": {\n \"@google/genai\": \"^1.45.0 || ^2\",\n \"@gullabs/core\": \"workspace:*\",\n \"@gullabs/google\": \"workspace:*\"\n },\n \"engines\": {\n \"node\": \">=20.9.0\"\n },\n \"sideEffects\": false,\n \"keywords\": [\n \"llm\",\n \"gemini\",\n \"google-genai\",\n \"ai\",\n \"tokens\",\n \"cost\",\n \"usage\",\n \"observability\",\n \"typescript\"\n ],\n \"publishConfig\": {\n \"access\": \"public\"\n },\n \"homepage\": \"https://github.com/gullabs/any-llm/tree/main/packages/any-llm#readme\",\n \"bugs\": \"https://github.com/gullabs/any-llm/issues\"\n}\n","/**\n * @gullabs/any-llm — batteries-included public entrypoint.\n *\n * This package is the default client install path. It re-exports the core\n * engine and Gemini adapter while depending on the Gemini SDK for a one-package\n * setup.\n *\n * @module\n */\n\nexport * from '@gullabs/core'\nexport * from '@gullabs/google'\n\nimport { version } from '../package.json'\n\n/** Library version, sourced from package.json at build time. */\nexport const ANY_LLM_VERSION: string = version\n"]}
|
package/dist/index.js
CHANGED
package/dist/index.js.map
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
{"version":3,"sources":["../package.json","../src/index.ts"],"names":[],"mappings":";;;;;;AAEE,IAAA,OAAA,GAAW,OAAA;;;ACcN,IAAM,eAAA,GAA0B","file":"index.js","sourcesContent":["{\n \"name\": \"@gullabs/any-llm\",\n \"version\": \"0.
|
|
1
|
+
{"version":3,"sources":["../package.json","../src/index.ts"],"names":[],"mappings":";;;;;;AAEE,IAAA,OAAA,GAAW,OAAA;;;ACcN,IAAM,eAAA,GAA0B","file":"index.js","sourcesContent":["{\n \"name\": \"@gullabs/any-llm\",\n \"version\": \"0.5.1\",\n \"description\": \"Batteries-included any-llm client for Gemini: engine, adapter, and Google SDK in one install.\",\n \"type\": \"module\",\n \"license\": \"Apache-2.0\",\n \"repository\": {\n \"type\": \"git\",\n \"url\": \"git+https://github.com/gullabs/any-llm.git\",\n \"directory\": \"packages/any-llm\"\n },\n \"main\": \"./dist/index.cjs\",\n \"module\": \"./dist/index.js\",\n \"types\": \"./dist/index.d.ts\",\n \"exports\": {\n \".\": {\n \"types\": \"./dist/index.d.ts\",\n \"import\": \"./dist/index.js\",\n \"require\": \"./dist/index.cjs\"\n }\n },\n \"files\": [\n \"dist\",\n \"skills\"\n ],\n \"scripts\": {\n \"build\": \"tsup\"\n },\n \"dependencies\": {\n \"@google/genai\": \"^1.45.0 || ^2\",\n \"@gullabs/core\": \"workspace:*\",\n \"@gullabs/google\": \"workspace:*\"\n },\n \"engines\": {\n \"node\": \">=20.9.0\"\n },\n \"sideEffects\": false,\n \"keywords\": [\n \"llm\",\n \"gemini\",\n \"google-genai\",\n \"ai\",\n \"tokens\",\n \"cost\",\n \"usage\",\n \"observability\",\n \"typescript\"\n ],\n \"publishConfig\": {\n \"access\": \"public\"\n },\n \"homepage\": \"https://github.com/gullabs/any-llm/tree/main/packages/any-llm#readme\",\n \"bugs\": \"https://github.com/gullabs/any-llm/issues\"\n}\n","/**\n * @gullabs/any-llm — batteries-included public entrypoint.\n *\n * This package is the default client install path. It re-exports the core\n * engine and Gemini adapter while depending on the Gemini SDK for a one-package\n * setup.\n *\n * @module\n */\n\nexport * from '@gullabs/core'\nexport * from '@gullabs/google'\n\nimport { version } from '../package.json'\n\n/** Library version, sourced from package.json at build time. */\nexport const ANY_LLM_VERSION: string = version\n"]}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@gullabs/any-llm",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.5.1",
|
|
4
4
|
"description": "Batteries-included any-llm client for Gemini: engine, adapter, and Google SDK in one install.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "Apache-2.0",
|
|
@@ -20,12 +20,13 @@
|
|
|
20
20
|
}
|
|
21
21
|
},
|
|
22
22
|
"files": [
|
|
23
|
-
"dist"
|
|
23
|
+
"dist",
|
|
24
|
+
"skills"
|
|
24
25
|
],
|
|
25
26
|
"dependencies": {
|
|
26
27
|
"@google/genai": "^1.45.0 || ^2",
|
|
27
|
-
"@gullabs/core": "0.4.
|
|
28
|
-
"@gullabs/google": "0.5.
|
|
28
|
+
"@gullabs/core": "0.4.3",
|
|
29
|
+
"@gullabs/google": "0.5.2"
|
|
29
30
|
},
|
|
30
31
|
"engines": {
|
|
31
32
|
"node": ">=20.9.0"
|
|
@@ -0,0 +1,260 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: any-llm
|
|
3
|
+
description: >-
|
|
4
|
+
Guidance for writing, reviewing, or debugging TypeScript code that calls
|
|
5
|
+
@gullabs/any-llm, @gullabs/core, or @gullabs/google to talk to Gemini models.
|
|
6
|
+
Applies when adding a new LLM call site, wiring createClient/generate/runStructured,
|
|
7
|
+
defining a defineCallSite prompt template, requesting structured JSON output,
|
|
8
|
+
catching or narrowing an LlmError, configuring reasoning/thinking budgets, or wiring
|
|
9
|
+
a UsageSink for cost tracking. Also applies whenever the user mentions any-llm, the
|
|
10
|
+
Gemini adapter, Gemini Flex tier, structured-output validation, or per-call auth for
|
|
11
|
+
this library. Covers the mandatory per-call `{ auth: { apiKey } }` pattern (there is
|
|
12
|
+
no env-var or ambient auth), the caller-owned output-validation contract, and the
|
|
13
|
+
reject-don't-map error philosophy — the things a developer used to other LLM SDKs
|
|
14
|
+
would otherwise get wrong by default.
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
# any-llm
|
|
18
|
+
|
|
19
|
+
Typed, provider-agnostic-by-design (currently Gemini-only) LLM call engine with cost
|
|
20
|
+
tracking, retries, rate limiting, structured output, and per-call observability.
|
|
21
|
+
|
|
22
|
+
Three packages:
|
|
23
|
+
|
|
24
|
+
- `@gullabs/core` — engine (`createClient`), types, errors, `defineCallSite`.
|
|
25
|
+
- `@gullabs/google` — Gemini adapter (`geminiAdapter`) over `@google/genai`.
|
|
26
|
+
- `@gullabs/any-llm` — batteries-included: re-exports both of the above plus `@google/genai` as a dependency.
|
|
27
|
+
|
|
28
|
+
Install `@gullabs/any-llm` for a one-package setup, or the two modular packages for
|
|
29
|
+
explicit dependency control. Import names are identical either way.
|
|
30
|
+
|
|
31
|
+
## #1 gotcha: auth is per-call, always
|
|
32
|
+
|
|
33
|
+
**There is no env-var auth, no ambient/singleton auth, no `AuthProvider` port.** Every
|
|
34
|
+
`generate()` and `runStructured()` call requires `opts.auth = { apiKey: string }`
|
|
35
|
+
explicitly. `createClient()` itself takes no credentials.
|
|
36
|
+
|
|
37
|
+
```ts
|
|
38
|
+
// WRONG — GenerateOptions.auth is a required field; this will not type-check, and if
|
|
39
|
+
// bypassed with `as any` it throws LlmError({ kind: 'invalid_auth' }) before any I/O.
|
|
40
|
+
const client = createClient({
|
|
41
|
+
adapters: [geminiAdapter()],
|
|
42
|
+
pricing: geminiPricingSource(),
|
|
43
|
+
})
|
|
44
|
+
await client.generate(request, {} as GenerateOptions)
|
|
45
|
+
|
|
46
|
+
// RIGHT — bring the key from wherever your app resolves it, pass it on every call.
|
|
47
|
+
const auth = { apiKey: myResolvedGeminiKey }
|
|
48
|
+
await client.generate(request, { auth })
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
A missing, empty, or non-string `apiKey` throws `LlmError` with `kind: 'invalid_auth'`
|
|
52
|
+
before any network call is made. Vertex AI (ADC/service-account auth) is **not
|
|
53
|
+
supported** — it was removed and is a roadmap item only; do not write code assuming
|
|
54
|
+
`{ vertex: { project, location } }` works.
|
|
55
|
+
|
|
56
|
+
## Quickstart
|
|
57
|
+
|
|
58
|
+
```ts
|
|
59
|
+
import { createClient, geminiPricingSource, geminiAdapter } from '@gullabs/any-llm'
|
|
60
|
+
// (or: from '@gullabs/core' / '@gullabs/google' respectively, if using modular install)
|
|
61
|
+
|
|
62
|
+
const client = createClient({
|
|
63
|
+
adapters: [geminiAdapter()],
|
|
64
|
+
pricing: geminiPricingSource(),
|
|
65
|
+
})
|
|
66
|
+
|
|
67
|
+
const result = await client.generate(
|
|
68
|
+
{
|
|
69
|
+
model: 'gemini-2.5-flash',
|
|
70
|
+
messages: [{ role: 'user', parts: [{ kind: 'text', text: 'Hello!' }] }],
|
|
71
|
+
},
|
|
72
|
+
{ auth: { apiKey: myResolvedGeminiKey } },
|
|
73
|
+
)
|
|
74
|
+
|
|
75
|
+
console.log(result.text) // raw text
|
|
76
|
+
console.log(result.usage) // { inputTokens, outputTokens, cachedInputTokens?, thinkingTokens?, details, raw }
|
|
77
|
+
console.log(result.cost?.microUsd) // integer micro-USD, or null if unpriced
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
`Message.parts` is `TextPart | InlineMediaPart | FileUriPart` — multimodal input mixes
|
|
81
|
+
`{ kind: 'text', text }`, `{ kind: 'inline-media', mimeType, data /* raw base64, no data: prefix */ }`,
|
|
82
|
+
and `{ kind: 'file-uri', uri, mimeType }` freely in one `parts` array.
|
|
83
|
+
|
|
84
|
+
## `defineCallSite` — reusable prompt templates
|
|
85
|
+
|
|
86
|
+
```ts
|
|
87
|
+
import { defineCallSite } from '@gullabs/core'
|
|
88
|
+
|
|
89
|
+
const summarize = defineCallSite({
|
|
90
|
+
id: 'summarize-article', // persisted as callSiteId on every record
|
|
91
|
+
model: 'gemini-2.5-flash',
|
|
92
|
+
system: 'You are a concise summarizer.',
|
|
93
|
+
userTemplate: 'Summarize this article in 3 sentences:\n\n{{article}}',
|
|
94
|
+
config: { temperature: 0.3 },
|
|
95
|
+
})
|
|
96
|
+
|
|
97
|
+
const result = await client.runStructured(summarize, { article: text }, { auth })
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
`{{var}}` interpolation is non-recursive (substituted values are never re-scanned for
|
|
101
|
+
further `{{...}}`, preventing template injection) and applies to both `system` and
|
|
102
|
+
`userTemplate`. A missing var is left as the literal `{{var}}` placeholder, not an
|
|
103
|
+
empty string. `runStructured` also accepts a two-arg form, `(callSite, opts)`, when the
|
|
104
|
+
template has no vars. Config resolution order everywhere is
|
|
105
|
+
`clientDefaults → callSite.config → opts.config` (later wins; objects like `reasoning`
|
|
106
|
+
and `providerOptions` merge key-by-key, everything else is last-write-wins).
|
|
107
|
+
|
|
108
|
+
## Structured output — auth + validation together
|
|
109
|
+
|
|
110
|
+
`request.output = { jsonSchema }` (or `callSite.jsonSchema`) is forwarded to the
|
|
111
|
+
provider as a **hint**, not enforced by the library. The engine JSON-parses the
|
|
112
|
+
response and sets `outputParsed`; `result.output` is always `unknown`. **The caller
|
|
113
|
+
owns shape validation** — this library does not validate output shape itself.
|
|
114
|
+
|
|
115
|
+
```ts
|
|
116
|
+
import { createClient, geminiPricingSource, geminiAdapter } from '@gullabs/any-llm'
|
|
117
|
+
import type { StandardSchemaV1 } from '@gullabs/core'
|
|
118
|
+
|
|
119
|
+
const client = createClient({
|
|
120
|
+
adapters: [geminiAdapter()],
|
|
121
|
+
pricing: geminiPricingSource(),
|
|
122
|
+
})
|
|
123
|
+
|
|
124
|
+
const result = await client.generate(
|
|
125
|
+
{
|
|
126
|
+
model: 'gemini-2.5-flash',
|
|
127
|
+
messages: [{ role: 'user', parts: [{ kind: 'text', text: 'Rate this PR 1-10.' }] }],
|
|
128
|
+
output: {
|
|
129
|
+
jsonSchema: {
|
|
130
|
+
type: 'object',
|
|
131
|
+
properties: { rating: { type: 'number' }, summary: { type: 'string' } },
|
|
132
|
+
required: ['rating', 'summary'],
|
|
133
|
+
},
|
|
134
|
+
},
|
|
135
|
+
},
|
|
136
|
+
{ auth: { apiKey: myResolvedGeminiKey } }, // auth is still required
|
|
137
|
+
)
|
|
138
|
+
|
|
139
|
+
// 1. Gate on outputParsed first — cheap boolean signal.
|
|
140
|
+
if (result.outputParsed !== true) {
|
|
141
|
+
// provider didn't return parseable JSON — retry, escalate, or fall back
|
|
142
|
+
} else {
|
|
143
|
+
// 2. Then validate shape with any Standard Schema v1 validator (zod, valibot, etc.
|
|
144
|
+
// all implement '~standard'). Do NOT trust result.output's shape without this.
|
|
145
|
+
const validated = await mySchema['~standard'].validate(result.output)
|
|
146
|
+
if ('issues' in validated && validated.issues !== undefined) {
|
|
147
|
+
// shape invalid but JSON-parseable — inspect validated.issues
|
|
148
|
+
} else {
|
|
149
|
+
// validated.value is now typed
|
|
150
|
+
}
|
|
151
|
+
}
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
Full caller-owned validation helper: `docs/structured-output-validation.md` in the
|
|
155
|
+
source repo (see "For more detail" below — this file may not ship in `node_modules`).
|
|
156
|
+
|
|
157
|
+
## Error handling
|
|
158
|
+
|
|
159
|
+
Every rejection from `generate()` / `runStructured()` is an `LlmError` with a `.kind`
|
|
160
|
+
discriminant (from `packages/core/src/errors.ts`):
|
|
161
|
+
|
|
162
|
+
| `kind` | Meaning | Retryable |
|
|
163
|
+
| ---------------- | ----------------------------------------------------- | --------- |
|
|
164
|
+
| `invalid_auth` | 401/403, or missing/empty `apiKey` | No |
|
|
165
|
+
| `rate_limited` | 429 — provider quota exceeded | Yes |
|
|
166
|
+
| `server` | 5xx — transient provider error | Yes |
|
|
167
|
+
| `timeout` | exceeded `config.timeoutMs` or a network timeout | Yes |
|
|
168
|
+
| `aborted` | caller's `AbortSignal` fired | No |
|
|
169
|
+
| `bad_request` | 400/422, or a request the library rejected before I/O | No |
|
|
170
|
+
| `content_filter` | provider refused output for safety reasons | No |
|
|
171
|
+
| `unknown` | uncategorized — inspect `.cause` | No |
|
|
172
|
+
|
|
173
|
+
```ts
|
|
174
|
+
import { LlmError } from '@gullabs/core'
|
|
175
|
+
|
|
176
|
+
try {
|
|
177
|
+
const result = await client.generate(request, { auth })
|
|
178
|
+
} catch (e) {
|
|
179
|
+
if (e instanceof LlmError) {
|
|
180
|
+
if (e.retryable) scheduleRetry(e.retryAfterMs)
|
|
181
|
+
else if (e.kind === 'invalid_auth') /* surface a credentials error */
|
|
182
|
+
else throw e
|
|
183
|
+
} else {
|
|
184
|
+
throw e // never expected — the engine always throws LlmError
|
|
185
|
+
}
|
|
186
|
+
}
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
`LlmError` also carries `httpStatus?`, `retryAfterMs?`, `provider?`, `callId?`,
|
|
190
|
+
`attemptId?`, `servedServiceTier?`, and `cause` (the original thrown value).
|
|
191
|
+
|
|
192
|
+
## Reject, don't map
|
|
193
|
+
|
|
194
|
+
Bad input or config throws `bad_request` (or `invalid_auth`) **before any I/O** —
|
|
195
|
+
nothing is silently clamped, coerced, or defaulted around a typo. Examples already
|
|
196
|
+
enforced by the engine/adapter: an unrouteable `model` string, a config value that
|
|
197
|
+
fails a model's schema (`validateConfig`), a `reasoning.budgetTokens` set on a model
|
|
198
|
+
whose API only supports `reasoning.effort`, duplicate adapter/middleware `id`s, and
|
|
199
|
+
grounding (`providerOptions.google.tools` search) combined with `output.jsonSchema` in
|
|
200
|
+
the same request (mutually exclusive on Gemini — see `docs/grounded-structured.md`).
|
|
201
|
+
|
|
202
|
+
**Do not add defensive fallback/clamping code around this library.** If a call throws
|
|
203
|
+
`bad_request`, fix the input — do not catch-and-retry with a "safer" guessed value; the
|
|
204
|
+
library is telling you the config is invalid, not transiently rejected.
|
|
205
|
+
|
|
206
|
+
## Reasoning / thinking budgets
|
|
207
|
+
|
|
208
|
+
```ts
|
|
209
|
+
config: {
|
|
210
|
+
reasoning: { effort: 'medium', includeThoughts: true }
|
|
211
|
+
}
|
|
212
|
+
```
|
|
213
|
+
|
|
214
|
+
`ReasoningEffort` is `'none' | 'low' | 'medium' | 'high'`. Two provider APIs exist
|
|
215
|
+
under the hood: Gemini 2.5 models take a token `budgetTokens`; Gemini 3.x / Gemma 4
|
|
216
|
+
models take a discrete `effort` level (`thinkingLevel`). `EFFORT_BUDGET` (from
|
|
217
|
+
`@gullabs/core`) is the shared effort→token mapping (`none: 0, low: 1024, medium: 8192,
|
|
218
|
+
high: 24576`) and `resolveReasoning({ model, budgetTokens, registry })` — also exported
|
|
219
|
+
from `@gullabs/core` — buckets a raw token budget into the nearest admitted effort
|
|
220
|
+
level for a given model, throwing `bad_request` if the model doesn't support reasoning
|
|
221
|
+
at all or has no admitted effort at or above the requested bucket. Prefer setting
|
|
222
|
+
`reasoning.effort` directly for level-API models rather than a token budget; the
|
|
223
|
+
adapter throws `bad_request` if you pass `budgetTokens` to a level-API model.
|
|
224
|
+
|
|
225
|
+
## Rate limiting and cost tracking
|
|
226
|
+
|
|
227
|
+
- Pre-send backpressure is a `RateLimiter` port (`ClientConfig.rateLimiter`); default
|
|
228
|
+
is a no-op. `@gullabs/core` ships a dependency-free `inMemoryRateLimiter`, and the
|
|
229
|
+
companion `@gullabs/quota` package provides shared/cross-process quota primitives —
|
|
230
|
+
see that package's README for setup.
|
|
231
|
+
- Every call computes `result.cost` (micro-USD) via the configured `PricingSource`
|
|
232
|
+
(`geminiPricingSource()`), and, when `sink` is configured on `createClient`, persists
|
|
233
|
+
a full per-attempt record (usage, cost, warnings, error classification) — fail-open,
|
|
234
|
+
so a broken sink never fails the LLM call. See `docs/ledger.md` for the record shape
|
|
235
|
+
and the `@gullabs/drizzle` package for a ready-made Postgres `UsageSink`.
|
|
236
|
+
|
|
237
|
+
## Common mistakes
|
|
238
|
+
|
|
239
|
+
- Forgetting `opts.auth` on a `generate()`/`runStructured()` call — it is required on
|
|
240
|
+
every call, not just once at `createClient()` time.
|
|
241
|
+
- Assuming `process.env.GEMINI_API_KEY` (or similar) is read automatically — it is
|
|
242
|
+
never read by this library; the host must resolve and pass the key itself.
|
|
243
|
+
- Assuming `result.output`'s shape is validated — it is `unknown`; validate it yourself
|
|
244
|
+
(Standard Schema v1) after checking `result.outputParsed === true`.
|
|
245
|
+
- Assuming Vertex AI is a supported target — it currently is not (roadmap only); only
|
|
246
|
+
the Gemini Developer API (API-key auth) is supported.
|
|
247
|
+
- Catching a `bad_request` `LlmError` and retrying with a clamped/guessed value instead
|
|
248
|
+
of fixing the call — the library never silently coerces invalid config.
|
|
249
|
+
- Setting `reasoning.budgetTokens` on a `thinkingLevel`-API model (Gemini 3.x / Gemma 4) — use `reasoning.effort` instead; the adapter throws `bad_request` otherwise.
|
|
250
|
+
|
|
251
|
+
## For more detail
|
|
252
|
+
|
|
253
|
+
These docs live in the source repo and may not ship in the installed `node_modules`
|
|
254
|
+
copy — the auth + structured-output example above is intentionally self-contained.
|
|
255
|
+
|
|
256
|
+
- `README.md` — full quickstart, install options, Gemma 4 notes.
|
|
257
|
+
- `docs/structured-output-validation.md` — full caller-owned validation helper.
|
|
258
|
+
- `docs/multi-runtime.md` — client construction across request/response vs. worker code.
|
|
259
|
+
- `docs/ledger.md` — persisted record field reference.
|
|
260
|
+
- `docs/grounded-structured.md` — grounding + structured output two-call pattern.
|