@zhivex-ai/vertex 1.0.2 → 1.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +995 -10
- package/dist/anthropic.d.ts +9 -0
- package/dist/anthropic.d.ts.map +1 -1
- package/dist/anthropic.js +86 -16
- package/dist/anthropic.js.map +1 -1
- package/dist/capabilities.d.ts +1 -0
- package/dist/capabilities.d.ts.map +1 -1
- package/dist/capabilities.js +12 -3
- package/dist/capabilities.js.map +1 -1
- package/dist/chat-profiles.d.ts +8 -0
- package/dist/chat-profiles.d.ts.map +1 -0
- package/dist/chat-profiles.js +61 -0
- package/dist/chat-profiles.js.map +1 -0
- package/dist/chat.d.ts +10 -0
- package/dist/chat.d.ts.map +1 -0
- package/dist/chat.js +120 -0
- package/dist/chat.js.map +1 -0
- package/dist/embeddings.d.ts +71 -0
- package/dist/embeddings.d.ts.map +1 -0
- package/dist/embeddings.js +166 -0
- package/dist/embeddings.js.map +1 -0
- package/dist/endpoints.d.ts +76 -0
- package/dist/endpoints.d.ts.map +1 -0
- package/dist/endpoints.js +152 -0
- package/dist/endpoints.js.map +1 -0
- package/dist/grpc.d.ts +33 -0
- package/dist/grpc.d.ts.map +1 -0
- package/dist/grpc.js +161 -0
- package/dist/grpc.js.map +1 -0
- package/dist/index.d.ts +40 -2
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +653 -292
- package/dist/index.js.map +1 -1
- package/dist/inline-thinking.d.ts +3 -0
- package/dist/inline-thinking.d.ts.map +1 -0
- package/dist/inline-thinking.js +106 -0
- package/dist/inline-thinking.js.map +1 -0
- package/dist/interactions.d.ts +21 -0
- package/dist/interactions.d.ts.map +1 -0
- package/dist/interactions.js +240 -0
- package/dist/interactions.js.map +1 -0
- package/dist/multimodal-embeddings.d.ts +53 -0
- package/dist/multimodal-embeddings.d.ts.map +1 -0
- package/dist/multimodal-embeddings.js +135 -0
- package/dist/multimodal-embeddings.js.map +1 -0
- package/dist/responses.d.ts +3 -0
- package/dist/responses.d.ts.map +1 -0
- package/dist/responses.js +120 -0
- package/dist/responses.js.map +1 -0
- package/dist/specialized.d.ts +14 -0
- package/dist/specialized.d.ts.map +1 -0
- package/dist/specialized.js +117 -0
- package/dist/specialized.js.map +1 -0
- package/dist/tensors.d.ts +22 -0
- package/dist/tensors.d.ts.map +1 -0
- package/dist/tensors.js +49 -0
- package/dist/tensors.js.map +1 -0
- package/dist/token-counting.d.ts +37 -0
- package/dist/token-counting.d.ts.map +1 -0
- package/dist/token-counting.js +33 -0
- package/dist/token-counting.js.map +1 -0
- package/dist/transcription.d.ts +60 -0
- package/dist/transcription.d.ts.map +1 -0
- package/dist/transcription.js +88 -0
- package/dist/transcription.js.map +1 -0
- package/dist/virtual-try-on.d.ts +23 -0
- package/dist/virtual-try-on.d.ts.map +1 -0
- package/dist/virtual-try-on.js +80 -0
- package/dist/virtual-try-on.js.map +1 -0
- package/package.json +5 -3
package/README.md
CHANGED
|
@@ -2,29 +2,45 @@
|
|
|
2
2
|
|
|
3
3
|
Vertex AI / Gemini Enterprise Agent Platform adapter for Zhivex AI SDK.
|
|
4
4
|
|
|
5
|
-
Supports Claude text, tools, and streaming through the Anthropic publisher, plus Vertex Gemini text, multimodal embeddings, speech, realtime sessions, grounded generation, Context Caching, Batch API, raw prediction calls, and current Google generative media endpoints for Gemini Image, Veo 3.1, and Lyria
|
|
5
|
+
Supports Claude text, tools, and streaming through the Anthropic publisher, plus Vertex Gemini text, multimodal embeddings, speech, realtime sessions, grounded generation, Context Caching, Batch API, raw prediction calls, and current Google generative media endpoints for Gemini Image, Veo 3.1, Lyria 2, and Gemini Omni / Lyria 3 through Interactions.
|
|
6
6
|
|
|
7
7
|
Google is transitioning Vertex AI into Gemini Enterprise Agent Platform. The SDK keeps the package name `@zhivex-ai/vertex`, the factory `createVertex()`, and provider id `"vertex"` for backwards compatibility and because the public API endpoints still use `aiplatform.googleapis.com`. Treat "Vertex" in this package as the Google Cloud Agent Platform / Vertex API surface, not as a separate deprecated wire contract.
|
|
8
8
|
|
|
9
|
+
This provider covers Google and partner model routes. A model's author and its
|
|
10
|
+
API host are separate: Claude, E5 and other publisher models invoked here still
|
|
11
|
+
use the `vertex` provider, Google Cloud authentication and Google billing.
|
|
12
|
+
Implemented capabilities do not imply that every model is enabled in your project.
|
|
13
|
+
See [verification and remaining limits](#verification-and-remaining-limits).
|
|
14
|
+
|
|
9
15
|
## Install
|
|
10
16
|
|
|
17
|
+
For HTTP operations, request deadlines and abort signals also bound waiting for
|
|
18
|
+
ADC or custom `getAccessToken()` credentials. A late token does not send a request
|
|
19
|
+
after cancellation. The credential resolver itself may continue in the background
|
|
20
|
+
because its interface does not accept an abort signal.
|
|
21
|
+
|
|
11
22
|
Requires Node.js 22 or newer when running on Node, matching Google Auth Library 11.
|
|
12
23
|
|
|
13
24
|
```bash
|
|
14
|
-
bun add @zhivex-ai/core @zhivex-ai/vertex
|
|
25
|
+
bun add @zhivex-ai/core @zhivex-ai/vertex @zhivex-ai/gateway
|
|
15
26
|
```
|
|
16
27
|
|
|
17
28
|
| Surface | Support |
|
|
18
29
|
| --- | --- |
|
|
19
30
|
| Text, tools, structured output, audio input | `generateText()` |
|
|
20
31
|
| Multimodal embeddings | `embeddingModel("gemini-embedding-2")` |
|
|
32
|
+
| E5 text embeddings | `embeddingModel("intfloat/multilingual-e5-small-maas")`; OAuth required |
|
|
21
33
|
| Speech and realtime sessions | `generateSpeech()`, `streamSpeech()`, and `realtimeModel()`; model and location dependent |
|
|
22
34
|
| Context Caching and Batch API | high-level |
|
|
23
35
|
| Google Search, Google Maps, URL Context, Code Execution, Computer Use | hosted tool helpers where the selected endpoint supports them |
|
|
24
36
|
| Image, video, music generation | high-level |
|
|
25
37
|
| Claude on Vertex | `vertex("claude-...")`: text, client tools, streaming, reasoning, native structured output on supported models |
|
|
26
38
|
| Publisher models / Model Garden | `predictionModel("publishers/<publisher>/models/<id>")`: explicit raw contract; bare IDs default to Google |
|
|
27
|
-
|
|
|
39
|
+
| Partner chat | `vertex("publisher/model")`: normalized chat, tools and streaming according to model capabilities |
|
|
40
|
+
| Interactions / Lyria 3 | `interactions` (experimental project-scoped API) |
|
|
41
|
+
| Mistral OCR / Codestral FIM | `ocr.process()` / `fim.generate()` / `fim.stream()` |
|
|
42
|
+
| DeepSeek OCR | `ocr.process()` with one image and optional extraction prompt |
|
|
43
|
+
| Gemini Files API, Gemini File Search stores | explicit unsupported surface in this adapter |
|
|
28
44
|
|
|
29
45
|
```ts
|
|
30
46
|
import {
|
|
@@ -116,7 +132,13 @@ await createContextCache({
|
|
|
116
132
|
await createBatch({
|
|
117
133
|
provider: productionVertex,
|
|
118
134
|
modelId: "gemini-3.7-flash",
|
|
119
|
-
fileName: "
|
|
135
|
+
fileName: "gs://my-bucket/batch-input.jsonl",
|
|
136
|
+
providerOptions: {
|
|
137
|
+
outputConfig: {
|
|
138
|
+
predictionsFormat: "jsonl",
|
|
139
|
+
gcsDestination: { outputUriPrefix: "gs://my-bucket/batch-output/" }
|
|
140
|
+
}
|
|
141
|
+
}
|
|
120
142
|
});
|
|
121
143
|
|
|
122
144
|
await predictRaw({
|
|
@@ -140,8 +162,9 @@ Current model guidance:
|
|
|
140
162
|
- Video: use the Google Cloud IDs `veo-3.1-generate-001`, `veo-3.1-fast-generate-001`, and `veo-3.1-lite-generate-001`. The Gemini Developer API uses different Veo `*-preview` IDs.
|
|
141
163
|
- Embeddings: `gemini-embedding-2` is the current multimodal model and is available on `global`, `us`, and `eu`.
|
|
142
164
|
- Speech: `gemini-3.1-flash-tts-preview` supports buffered `generateSpeech()` and incremental `streamSpeech()` output. It is currently available through the Vertex AI API on `global`; older Gemini 2.5 TTS models have broader regional coverage.
|
|
143
|
-
- Music: `lyria-002` is the GA model supported by `musicGenerationModel()`. Lyria 3
|
|
165
|
+
- Music: `lyria-002` is the GA model supported by `musicGenerationModel()`. Lyria 3 uses the experimental project-scoped Interactions API with bearer credentials and `location: "global"`.
|
|
144
166
|
- Imagen 4 and older Veo endpoints are intentionally no longer recommended here; Google Cloud required migration away from them by June 30, 2026.
|
|
167
|
+
- Legacy Imagen `outputMimeType` maps to `parameters.outputOptions.mimeType`; native `providerOptions.outputOptions.compressionQuality` is preserved. Conflicting MIME settings are rejected. The text-to-image factory rejects `images` rather than silently ignoring them; native editing requires the separate `referenceImages` prediction contract. This does not restore access to retired models.
|
|
145
168
|
|
|
146
169
|
Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite use provider-managed sampling. Do not pass `temperature`, `topP` / `top_p`, `topK` / `top_k`, `candidateCount` / `candidate_count`, or frequency/presence penalties; the adapter rejects those controls locally for these model IDs. Gemini 3.7 accepts `reasoning.effort` values `low`, `medium`, and `high`; Gemini 3.6 and Gemini 3.5 Flash-Lite also accept `minimal`. All three reject a final assistant/model-output prefill. The current Google Cloud endpoint does not expose Computer Use for these models, so their Vertex model capabilities report it as unsupported.
|
|
147
170
|
|
|
@@ -149,7 +172,7 @@ The mutable aliases `gemini-flash-latest` and `gemini-flash-lite-latest` are ava
|
|
|
149
172
|
|
|
150
173
|
The built-in catalog's Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite rates represent Standard global text-token pricing. Non-global Vertex endpoints can cost more, and media, tools, Batch/Flex, Priority, tuning, and Provisioned Throughput use separate pricing.
|
|
151
174
|
|
|
152
|
-
|
|
175
|
+
`videoGenerationModel("gemini-omni-flash-preview")` and `videoGenerationModel("gemini-omni-1.1-flash-preview")` route text-to-video and image-to-video through Vertex Interactions with bearer credentials and `location: "global"`. They generate one video per synchronous call, accept integer durations from 3 to 10 seconds, aspect ratios `16:9` / `9:16`, and optional `outputStorageUri` for GCS delivery. `providerOptions.resolution` supports `720p` on Omni and `360p`, `720p`, `1080p`, `4k` on Omni 1.1. Negative prompts and polling controls are rejected. Use `vertex.interactions` for reference-to-video, first/last-frame, editing and asynchronous workflows. Managed-agent inference is accessible through `interactions.create({ agent, input, background: true })`; provisioning and deployment administration are separate APIs.
|
|
153
176
|
|
|
154
177
|
When Google Maps grounding is enabled, retain the provider response metadata and render the returned source names and Google Maps links directly after the grounded content. Google requires those sources and its text attribution to remain visible to the end user.
|
|
155
178
|
|
|
@@ -157,7 +180,11 @@ See Google's current [Agent Platform model lifecycle](https://docs.cloud.google.
|
|
|
157
180
|
|
|
158
181
|
Google's current product page labels this surface as [Gemini Enterprise Agent Platform, formerly Vertex AI](https://cloud.google.com/products/gemini-enterprise-agent-platform), and Google's migration docs say Vertex AI is transitioning to become part of Agent Platform. This package intentionally does not rename the provider id yet; doing so would be a breaking API change without a corresponding endpoint-level migration requirement.
|
|
159
182
|
|
|
160
|
-
Model Garden raw prediction accepts explicit `publishers/<publisher>/models/<id>` resources, relative to the configured project and location. Bare prediction IDs retain the Google publisher default. Supply the model-specific `body` and `providerOptions.action` (for example `rawPredict`) to `predictRaw()`. This is transport access, not a promise of normalized tools, streaming, or support for every Model Garden deployment. Self-deployed
|
|
183
|
+
Model Garden raw prediction accepts explicit `publishers/<publisher>/models/<id>` resources, relative to the configured project and location. Bare prediction IDs retain the Google publisher default. Supply the model-specific `body` and `providerOptions.action` (for example `rawPredict`) to `predictRaw()`. This is transport access, not a promise of normalized tools, streaming, or support for every Model Garden deployment. Self-deployed endpoints use `predictionModel("endpoints/<id>")` or a fully qualified `projects/<project>/locations/<location>/endpoints/<id>` resource. Both require bearer credentials and use the deployed model's raw request/response contract; they do not automatically provide normalized chat or streaming.
|
|
184
|
+
|
|
185
|
+
Callable tool inputs are validated locally against their Zod schemas. Gemini
|
|
186
|
+
requests map those schemas to Vertex parameters, removing unsupported JSON
|
|
187
|
+
Schema metadata and `additionalProperties`, including nested schemas.
|
|
161
188
|
|
|
162
189
|
## Claude on Vertex
|
|
163
190
|
|
|
@@ -180,18 +207,976 @@ Enable the selected Claude model in Model Garden and choose a supported location
|
|
|
180
207
|
|
|
181
208
|
Supported through the shared language-model API: text, image/document input using Anthropic message mapping, client tool loops, streaming, usage, reasoning, and native structured output for Claude 4.5 and later families. Structured output additionally requires the Google organization policy to allow `structured_outputs`. Capabilities describe the implemented contract, not a guarantee of account entitlement or live certification.
|
|
182
209
|
|
|
183
|
-
|
|
210
|
+
Supported Claude native tools include web search (`web_search_20250305`), computer use (`computer_20250124`), Bash, text editor, memory and tool search. Use `hostedTool({ provider: "vertex", type, name, config })`; computer use adds its required beta to the Vertex request body. Unsupported server tools (web fetch/code execution/advisor), direct Anthropic Files API IDs and URL input sources, remote MCP, unrecognized betas, fast mode, server-side fallbacks and direct-API context-management options are explicitly rejected. SDK-managed MCP tools can still execute as ordinary client tools. Google grounding, Gemini cache/batch/media APIs, Interactions, and managed Agent Platform runtime/session/deployment APIs are not Claude language-model features exposed here.
|
|
184
211
|
|
|
185
212
|
The SDK catalog includes Claude entries under `vertex` separately from `anthropic`, without copying direct-API prices or automatic recommendations. Contract tests use mocked HTTP; the opt-in `VERTEX_CLAUDE_INTEGRATION_MODEL` suite validates the actual Google route when credentials and model access are available.
|
|
186
213
|
|
|
187
214
|
Sources: [Claude requests on Vertex](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/claude/use-claude), [Claude structured outputs on Vertex](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/claude/structured-outputs).
|
|
188
215
|
|
|
189
|
-
|
|
216
|
+
### Claude prompt caching
|
|
217
|
+
|
|
218
|
+
`providerOptions.cache_control = { type: "ephemeral", ttl: "1h" }` enables
|
|
219
|
+
supported automatic prompt caching; `5m` is also accepted. Explicit block-level
|
|
220
|
+
breakpoints can be supplied in Anthropic protocol `provider-data` blocks. Older
|
|
221
|
+
Claude 3.7 Sonnet / 3.5 Sonnet / 3 Opus reject a one-hour TTL. Cache read and write
|
|
222
|
+
tokens are normalized in usage. On `global`, set `providerOptions.sessionId` to a
|
|
223
|
+
stable application session ID for the `X-Vertex-Ai-Session-Id` routing header.
|
|
224
|
+
Neither this ID nor Anthropic credentials are serialized as model inputs.
|
|
225
|
+
`VertexClaudeOptions` provides the typed options contract.
|
|
226
|
+
|
|
227
|
+
References: [Claude feature availability](https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai),
|
|
228
|
+
[Google prompt caching](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/claude/prompt-caching),
|
|
229
|
+
[Google Claude web search](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/claude/web-search).
|
|
230
|
+
|
|
231
|
+
### Claude context management
|
|
232
|
+
|
|
233
|
+
Automatic context compaction and context editing use Vertex beta flags in the
|
|
234
|
+
request body. The adapter adds the appropriate flags for these native options:
|
|
235
|
+
|
|
236
|
+
```ts
|
|
237
|
+
await vertex("claude-sonnet-4-6").generate({
|
|
238
|
+
messages: [{ role: "user", parts: [{ type: "text", text: "Continue the task." }] }],
|
|
239
|
+
providerOptions: {
|
|
240
|
+
context_management: {
|
|
241
|
+
edits: [{ type: "compact_20260112", trigger: { type: "input_tokens", value: 100000 } }],
|
|
242
|
+
},
|
|
243
|
+
},
|
|
244
|
+
});
|
|
245
|
+
```
|
|
246
|
+
|
|
247
|
+
Supported strategies are `compact_20260112`, `clear_tool_uses_20250919` and
|
|
248
|
+
`clear_thinking_20251015`. Thinking clearing must come first when combined with
|
|
249
|
+
other edits. Automatic compaction requires a supported Claude model and a trigger
|
|
250
|
+
of at least 50,000 input tokens. Preserve returned provider-data blocks in the
|
|
251
|
+
conversation history. These features are beta and their availability depends on
|
|
252
|
+
the model. On-demand `compaction` and `compact-2026-09-04` are not available on
|
|
253
|
+
Vertex and are rejected. See Anthropic's
|
|
254
|
+
[context management availability](https://platform.claude.com/docs/en/build-with-claude/overview)
|
|
255
|
+
and [compaction contract](https://platform.claude.com/docs/en/build-with-claude/compaction).
|
|
256
|
+
|
|
257
|
+
### Count Claude input tokens
|
|
258
|
+
|
|
259
|
+
```ts
|
|
260
|
+
const count = await vertex.claude.countTokens({
|
|
261
|
+
modelId: "claude-sonnet-4-6",
|
|
262
|
+
messages: [{ role: "user", content: "How many tokens are in this request?" }],
|
|
263
|
+
});
|
|
264
|
+
console.log(count.inputTokens);
|
|
265
|
+
```
|
|
266
|
+
|
|
267
|
+
This method accepts native Claude message content (text or content-block arrays),
|
|
268
|
+
plus `system`, `tools`, `thinking` and `toolChoice`. It calls the dedicated
|
|
269
|
+
`publishers/anthropic/models/count-tokens:rawPredict` endpoint, with the target
|
|
270
|
+
model in the body. Configure `global`, `us`, `eu` or `asia-southeast1` and OAuth
|
|
271
|
+
credentials. It does not fall back to a generation request. HTTP failures and
|
|
272
|
+
invalid count responses are surfaced explicitly. See Google's
|
|
273
|
+
[Claude token-counting reference](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/claude/count-tokens).
|
|
274
|
+
|
|
275
|
+
### Claude browser toolsets
|
|
276
|
+
|
|
277
|
+
Compatible Claude models accept the native `browser_toolset_20260801` declaration:
|
|
278
|
+
|
|
279
|
+
```ts
|
|
280
|
+
import { hostedTool } from "@zhivex-ai/core";
|
|
281
|
+
|
|
282
|
+
const browser = hostedTool({
|
|
283
|
+
provider: "vertex", type: "browser_toolset_20260801", name: "browser",
|
|
284
|
+
});
|
|
285
|
+
const model = vertex("claude-opus-5");
|
|
286
|
+
const turn = await model.generate({
|
|
287
|
+
messages: [{ role: "user", parts: [{ type: "text", text: "Open example.com." }] }],
|
|
288
|
+
tools: { browser },
|
|
289
|
+
});
|
|
290
|
+
```
|
|
291
|
+
|
|
292
|
+
Use the low-level `generate()` / `stream()` loop with your browser executor. The
|
|
293
|
+
SDK returns each member call with `providerMetadata.toolset_name = "browser"`.
|
|
294
|
+
Execute members sequentially in response order. Return a `tool-result` with that
|
|
295
|
+
same metadata and either text or a native content-block array in `output` (for
|
|
296
|
+
example `text`, `image` and `browser_state` blocks). If one action fails, mark it
|
|
297
|
+
and the remaining actions in that batch as errors instead of executing later
|
|
298
|
+
steps. Append the assistant message and all results before requesting the next
|
|
299
|
+
turn. The application owns browser execution and permissions; declaring this
|
|
300
|
+
native toolset does not register callable member executors with `generateText()`.
|
|
190
301
|
|
|
191
|
-
|
|
302
|
+
The SDK's local tool name is omitted from the native declaration. Supported
|
|
303
|
+
models are Opus 4.8/5, Sonnet 5 and Fable/Mythos 5/5.1 where available on Vertex.
|
|
304
|
+
Other model IDs reject the declaration before network access. See the official
|
|
305
|
+
[browser toolset contract](https://platform.claude.com/docs/en/agents-and-tools/tool-use/browser-use-tool)
|
|
306
|
+
for member parameters and native result blocks. Offline tests cover declaration,
|
|
307
|
+
streaming identity and the complete call/result replay; live validation is pending.
|
|
192
308
|
|
|
193
309
|
## Gemini 3.8 Flash
|
|
194
310
|
|
|
195
311
|
`gemini-3.8-flash` is included in the catalog and uses the same local sampling, prefill, and reasoning validation as the current Vertex Flash family. Accepted reasoning efforts are `low`, `medium`, and `high`. Unlike the earlier Flash releases, 3.8 exposes Computer Use (Preview) and supports `global`, `us`, and `eu` locations. Availability must be verified for the selected project, endpoint, and location. The Vertex adapter does not inherit Gemini API-only surfaces or pricing.
|
|
196
312
|
|
|
197
313
|
See the [Google Cloud model card](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-8-flash).
|
|
314
|
+
|
|
315
|
+
## Native Vertex resources
|
|
316
|
+
|
|
317
|
+
### Arbitrary HTTP predictions on deployed endpoints
|
|
318
|
+
|
|
319
|
+
`vertex.endpoints.rawPredict()` preserves binary or text payloads and returns
|
|
320
|
+
response bytes without JSON serialization or parsing. For example:
|
|
321
|
+
|
|
322
|
+
```ts
|
|
323
|
+
const endpoint = createVertex({ projectId: "my-project", location: "us-central1" });
|
|
324
|
+
const response = await endpoint.endpoints.rawPredict({
|
|
325
|
+
endpoint: "endpoints/my-endpoint-id",
|
|
326
|
+
body: new Uint8Array([0, 255, 128]),
|
|
327
|
+
contentType: "application/octet-stream",
|
|
328
|
+
maxResponseBytes: 4 * 1024 * 1024,
|
|
329
|
+
timeoutMs: 30_000,
|
|
330
|
+
maxRetries: 0,
|
|
331
|
+
});
|
|
332
|
+
// response.body, contentType, status, endpointId and deployedModelId
|
|
333
|
+
```
|
|
334
|
+
|
|
335
|
+
The deployed container defines the input and output formats. Use Google bearer
|
|
336
|
+
credentials/ADC and the endpoint's region; full project-qualified endpoint names
|
|
337
|
+
are also accepted, without changing the configured host. The default response
|
|
338
|
+
limit is 16 MiB, enforced before and during consumption. Shared deadlines,
|
|
339
|
+
cancellation and explicit HTTP retries apply. This unary method does not decode
|
|
340
|
+
streams; existing `predictionModel(...).rawPredict()` remains the JSON contract.
|
|
341
|
+
For streaming containers, `vertex.endpoints.streamRawPredict()` accepts the same
|
|
342
|
+
input and returns an async iterable: a `response` event with HTTP metadata,
|
|
343
|
+
followed by `chunk` events containing raw `data: Uint8Array`. Chunk boundaries
|
|
344
|
+
are transport boundaries; applications must decode their container's protocol.
|
|
345
|
+
The size limit applies to the accumulated stream. Breaking the loop cancels the
|
|
346
|
+
body; abort and deadline also interrupt a stalled read. HTTP errors can be
|
|
347
|
+
retried before yielding events, but an interrupted response body is never replayed.
|
|
348
|
+
Contract and installed-package tests cover binary payloads; live validation still
|
|
349
|
+
requires a deployed endpoint. See [Google's rawPredict API](https://docs.cloud.google.com/gemini-enterprise-agent-platform/reference/rest/v1/projects.locations.endpoints/rawPredict).
|
|
350
|
+
The streaming route follows [streamRawPredict](https://docs.cloud.google.com/gemini-enterprise-agent-platform/reference/rest/v1/projects.locations.endpoints/streamRawPredict).
|
|
351
|
+
|
|
352
|
+
For a custom container exposing a gRPC model server, use
|
|
353
|
+
`vertex.endpoints.directRawPredict({ endpoint, methodName, input })`, where
|
|
354
|
+
`methodName` is a fully qualified method such as
|
|
355
|
+
`/tensorflow.serving.PredictionService/Predict` and `input` is its serialized
|
|
356
|
+
request as `Uint8Array`. The client sends the REST base64 envelope and returns
|
|
357
|
+
`output: Uint8Array` plus HTTP status. Applications own protobuf serialization;
|
|
358
|
+
this is not a native gRPC channel. `maxResponseBytes` bounds decoded output
|
|
359
|
+
(16 MiB default); the JSON envelope is bounded separately with base64 overhead.
|
|
360
|
+
Empty bytes, including an omitted default output field, remain empty bytes.
|
|
361
|
+
The same auth, timeout and retry options apply. See
|
|
362
|
+
[directRawPredict](https://docs.cloud.google.com/gemini-enterprise-agent-platform/reference/rest/v1/projects.locations.endpoints/directRawPredict).
|
|
363
|
+
|
|
364
|
+
For bidirectional gRPC containers, use
|
|
365
|
+
`vertex.endpoints.streamDirectRawPredict({ endpoint, methodName, inputs })`
|
|
366
|
+
with an iterable or async iterable of `Uint8Array` messages, or
|
|
367
|
+
`vertex.endpoints.streamDirectPredict({ endpoint, inputs })` with tensor frames
|
|
368
|
+
`{ inputs: VertexTensor[], parameters?: VertexTensor }`. Iterate the returned
|
|
369
|
+
async iterable to receive bytes or `{ outputs, parameters? }` respectively.
|
|
370
|
+
`vertex.endpoints.streamingRawPredict()` and
|
|
371
|
+
`vertex.endpoints.streamingPredict()` expose the separate StreamingRawPredict
|
|
372
|
+
and StreamingPredict RPCs with the same byte and tensor input contracts. Select
|
|
373
|
+
the RPC supported by your deployed container; no automatic fallback is applied.
|
|
374
|
+
|
|
375
|
+
For a single request followed by multiple responses, use
|
|
376
|
+
`vertex.endpoints.serverStreamingPredict({ endpoint, inputs, parameters })`.
|
|
377
|
+
Here `inputs` is a `VertexTensor[]`, not an iterable of request frames. The
|
|
378
|
+
result is an async iterable of `{ outputs, parameters? }`. This client uses the
|
|
379
|
+
server-streaming gRPC RPC for deployed endpoints or publisher model resources
|
|
380
|
+
(`publishers/<publisher>/models/<model>` or the project-qualified form), using
|
|
381
|
+
the configured API host. A model resource does not imply that the model supports
|
|
382
|
+
this RPC. Other direct/bidirectional methods still require deployed endpoints.
|
|
383
|
+
It shares the same response
|
|
384
|
+
limits, cancellation and no-replay behavior.
|
|
385
|
+
|
|
386
|
+
These methods use the official Google gRPC client and bearer authentication;
|
|
387
|
+
they do not use the configured HTTP `fetch` implementation. Native service
|
|
388
|
+
failures retain the Google gRPC error fields (`code`, `details`, `metadata`);
|
|
389
|
+
they are not `ProviderHTTPError` instances. SDK configuration, cancellation and
|
|
390
|
+
response-limit errors retain their existing contracts. The first request
|
|
391
|
+
carries endpoint routing metadata; subsequent requests carry input frames.
|
|
392
|
+
|
|
393
|
+
Set `timeoutMs` or `abortSignal` to bound the session. Leaving the response loop
|
|
394
|
+
cancels both directions and closes the client. `maxResponseBytes` bounds the
|
|
395
|
+
cumulative raw output bytes or serialized tensor output (16 MiB by default),
|
|
396
|
+
and also configures the gRPC per-message receive limit. Streaming inputs cannot
|
|
397
|
+
be replayed: `maxRetries` must be omitted or zero. These methods require a
|
|
398
|
+
deployed endpoint supporting the corresponding bidirectional protocol.
|
|
399
|
+
|
|
400
|
+
`vertex.endpoints.directPredict({ endpoint, inputs, parameters })` provides the
|
|
401
|
+
native REST tensor contract for compatible gRPC model servers. `inputs` and
|
|
402
|
+
returned `outputs` are `VertexTensor[]`; optional `parameters` is a tensor too.
|
|
403
|
+
`shape`, `int64Val` and `uint64Val` use decimal strings to avoid JavaScript number
|
|
404
|
+
precision loss. Byte fields retain base64 strings; floating fields accept finite
|
|
405
|
+
numbers or the ProtoJSON strings `NaN`, `Infinity` and `-Infinity`. Nested
|
|
406
|
+
`listVal`/`structVal` tensors are preserved with a maximum nesting depth of 32.
|
|
407
|
+
Invalid field types, 64-bit ranges and incompatible scalar representations are
|
|
408
|
+
rejected; the model server owns tensor shape and model-specific validation.
|
|
409
|
+
`maxResponseBytes` bounds the JSON response (16 MiB default). See the native
|
|
410
|
+
[directPredict](https://docs.cloud.google.com/gemini-enterprise-agent-platform/reference/rest/v1/projects.locations.endpoints/directPredict)
|
|
411
|
+
and [Tensor](https://docs.cloud.google.com/gemini-enterprise-agent-platform/reference/rest/v1/Tensor) contracts.
|
|
412
|
+
|
|
413
|
+
`vertex.endpoints.explain({ endpoint, instances, parameters, deployedModelId,
|
|
414
|
+
explanationSpecOverride })` returns native `explanations`, `predictions` and the
|
|
415
|
+
serving `deployedModelId`. It preserves attribution/example details and checks
|
|
416
|
+
that the response has one explanation for each input instance. Overrides may
|
|
417
|
+
set native `parameters`, `metadata` or `examplesOverride`; their contents depend
|
|
418
|
+
on the model. The selected deployed model must already have an `explanationSpec`
|
|
419
|
+
configured; without a selected ID, all deployed models must have it. Response
|
|
420
|
+
JSON is bounded to 16 MiB by default, configurable through `maxResponseBytes`.
|
|
421
|
+
This API does not configure the deployment or implement attribution locally.
|
|
422
|
+
See [online explanations](https://docs.cloud.google.com/gemini-enterprise-agent-platform/reference/rest/v1/projects.locations.endpoints/explain).
|
|
423
|
+
|
|
424
|
+
### Batch, cache and generated media operations
|
|
425
|
+
|
|
426
|
+
Native image, music and video generation honor configured HTTP retries and bound
|
|
427
|
+
backoff by the overall deadline. Video polling retries the existing operation
|
|
428
|
+
without resubmitting generation. Submission retries remain opt-in and can incur
|
|
429
|
+
additional generation work; they do not imply server-side deduplication.
|
|
430
|
+
|
|
431
|
+
Batch input selects exactly one source: `fileName` or `providerOptions.inputConfig`.
|
|
432
|
+
For BigQuery, `fileName` identifies a table as `bq://project.dataset.table`.
|
|
433
|
+
Configure `outputConfig.predictionsFormat` as `bigquery` and
|
|
434
|
+
`outputConfig.bigqueryDestination.outputUri` for BigQuery output. The request
|
|
435
|
+
mapping has contract coverage; a real BigQuery batch lifecycle remains unverified.
|
|
436
|
+
|
|
437
|
+
`predictionModel().predictRaw()` uses `providerOptions.action` to select the URL
|
|
438
|
+
action and omits it from the generated body. Supply model-native fields directly
|
|
439
|
+
in `body` to preserve the raw payload unchanged. Without an explicit body,
|
|
440
|
+
provider options cannot override dedicated `instances` or `parameters` fields.
|
|
441
|
+
Operation polling takes its identity from `name` and rejects native overrides.
|
|
442
|
+
|
|
443
|
+
Native prediction methods and operation polling honor `maxRetries` for transient
|
|
444
|
+
HTTP errors, with retry backoff bounded by the request timeout. Retries are off
|
|
445
|
+
by default. A retry of a submission can execute it again; this API does not add
|
|
446
|
+
an idempotency key or guarantee deduplication for deployed models.
|
|
447
|
+
|
|
448
|
+
`gemini-embedding-2` uses `embedContent` and accepts text or `MediaInput` values
|
|
449
|
+
(inline data or Cloud Storage URIs). Legacy text embedding models continue to use
|
|
450
|
+
`predict`. Each value produces one vector in caller order.
|
|
451
|
+
|
|
452
|
+
Batch model selectors accept `publisher/model` as well as full publisher resources.
|
|
453
|
+
Transient HTTP failures respect the configured retry policy. A bounded ADC/GCS
|
|
454
|
+
check verified Gemini 2.5 Flash batch creation, completion, expected output and
|
|
455
|
+
job deletion in `us-central1`; its temporary storage was removed and absence
|
|
456
|
+
confirmed. Partner batch and BigQuery need separate live evidence. A separate live check
|
|
457
|
+
verified cancellation through JOB_STATE_CANCELLED and cleanup with 404 checks.
|
|
458
|
+
The Claude smoke uses `start STATE_FILE --claude` for one Sonnet 4.6 request
|
|
459
|
+
on `us-east5`, with a temporary private bucket in `us-central1`. The tested
|
|
460
|
+
project returned 404 at job creation; no job ID was returned. This does not
|
|
461
|
+
establish whether the cause is model access or service availability. Follow-up
|
|
462
|
+
commands use the same state file without the flag. Its temporary storage was
|
|
463
|
+
removed; Claude batch completion remains unverified.
|
|
464
|
+
|
|
465
|
+
Batch operations use the project-scoped `batchPredictionJobs` API with Google
|
|
466
|
+
Cloud bearer credentials. Supply `fileName` as a `gs://` JSONL or `bq://` table
|
|
467
|
+
URI, or pass `providerOptions.inputConfig`; also supply
|
|
468
|
+
`providerOptions.outputConfig`. Gemini Developer API `files/*` IDs and inline
|
|
469
|
+
`requests` are not Vertex batch inputs. Model availability for batch must be
|
|
470
|
+
checked separately from online prediction locations. Claude IDs route to the
|
|
471
|
+
Anthropic publisher; explicit publisher resources are accepted for other models.
|
|
472
|
+
Partner batch creation requires a supported regional endpoint: `global` is
|
|
473
|
+
rejected locally, including when the model resource contains a regional prefix.
|
|
474
|
+
The client's location determines where the job is created. See the
|
|
475
|
+
[Claude batch contract](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/claude/batch).
|
|
476
|
+
Cancellation returns the job's current state after requesting cancellation;
|
|
477
|
+
deletion returns the raw long-running operation for inspection.
|
|
478
|
+
|
|
479
|
+
US/EU jurisdictional endpoints use `aiplatform.us.rep.googleapis.com` and
|
|
480
|
+
`aiplatform.eu.rep.googleapis.com`. Full cache/job resource names returned by
|
|
481
|
+
Google can be passed directly to their get/delete/cancel methods.
|
|
482
|
+
|
|
483
|
+
## Partner chat and deployed chat endpoints
|
|
484
|
+
|
|
485
|
+
Use publisher-qualified IDs for managed open models. These requests use Google
|
|
486
|
+
Cloud bearer credentials and billing; direct provider API keys are not used.
|
|
487
|
+
|
|
488
|
+
```ts
|
|
489
|
+
const cloud = createVertex({ projectId: "my-project", location: "global" });
|
|
490
|
+
const answer = await generateText({
|
|
491
|
+
model: cloud("xai/grok-4.3"),
|
|
492
|
+
prompt: "Explain this architecture"
|
|
493
|
+
});
|
|
494
|
+
|
|
495
|
+
const deployed = cloud.chatModel("my-deployed-model", {
|
|
496
|
+
endpoint: "endpoints/123456789",
|
|
497
|
+
capabilities: { tools: true, toolChoice: true, structuredOutput: true }
|
|
498
|
+
});
|
|
499
|
+
```
|
|
500
|
+
|
|
501
|
+
The callable factory recognizes `xai/`, `meta/`, `deepseek-ai/`, `qwen/`,
|
|
502
|
+
`zai-org/`, `moonshotai/`, `minimaxai/`, `openai/` and `google/gemma*` selectors
|
|
503
|
+
and routes them through `endpoints/openapi/chat/completions`. `mistralai/` and
|
|
504
|
+
`ai21/` use publisher `rawPredict` / `streamRawPredict` with the chat contract.
|
|
505
|
+
Bare `grok-*`, `mistral-*`, `codestral*` and `jamba-*` IDs are qualified with their
|
|
506
|
+
publisher. Jamba 1.5 Mini and Large retired on February 27, 2026; their catalog entries retain this lifecycle history and do not indicate current availability. See the [partner retirement schedule](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/deprecations/partner-models). Mistral OCR is a separate document API and is rejected by chat.
|
|
507
|
+
Mistral publisher calls map the shared `toolChoice: "required"` to native `any`.
|
|
508
|
+
Managed Mistral rejects `providerOptions.safe_prompt`, including `false`.
|
|
509
|
+
The legacy Jamba profile exposes JSON-object mode without claiming native
|
|
510
|
+
JSON-schema output; requests combining streaming and tools are rejected locally.
|
|
511
|
+
These managed-host restrictions do not override explicitly configured custom
|
|
512
|
+
deployment capabilities.
|
|
513
|
+
Model access and location support remain Google-account dependent; a recognized
|
|
514
|
+
selector is not live certification or a guarantee that every publisher model
|
|
515
|
+
speaks the chat protocol. Advanced capabilities and thinking controls are selected by exact managed model IDs; unknown IDs retain text/stream transport but do not inherit tools, schema output, vision or reasoning from a name prefix. Mistral/AI21 `@revision` selectors retain the base model profile. Codestral FIM and OCR require their specialized APIs.
|
|
516
|
+
Google-only factories (grounding, media, Live and explicit context caching)
|
|
517
|
+
reject recognized partner selectors before sending a request. E5 embeddings
|
|
518
|
+
use their own embedding route rather than the Google embedding API.
|
|
519
|
+
|
|
520
|
+
The SDK catalog includes GLM 5.2 Preview, Gemma 4 26B, both Llama 4 variants
|
|
521
|
+
and gpt-oss 120B. A Preview minimum-availability date is not a retirement date;
|
|
522
|
+
GLM 5.2 does not inherit the October retirement of GLM 5.
|
|
523
|
+
|
|
524
|
+
Client tool loops, native JSON-schema output, image inputs on supported models,
|
|
525
|
+
usage and text streaming use the shared SDK contract. Hosted direct-provider
|
|
526
|
+
tools and API-only features do not carry over. Separate `reasoning_content`
|
|
527
|
+
fields are preserved as Vertex provider-data and replayed in tool history;
|
|
528
|
+
reasoning embedded by most hosts in text is retained
|
|
529
|
+
verbatim. Grok does not accept effort controls on Vertex. GPT OSS accepts
|
|
530
|
+
`low`, `medium`, or `high`; DeepSeek V3.1/V3.2, Gemma 4 and GLM 4.7/5/5.2 map effort to their
|
|
531
|
+
hosted thinking toggle (`none` disables it; `low`, `medium` and `high` enable the
|
|
532
|
+
same toggle). No direct-provider thinking contract is assumed.
|
|
533
|
+
|
|
534
|
+
GPT OSS accepts only `auto` or `none` tool choice on Vertex; required and named
|
|
535
|
+
choices are rejected locally. When tools are provided without a choice, GPT OSS
|
|
536
|
+
and Qwen requests explicitly use `auto`. This also prevents a current GPT OSS
|
|
537
|
+
host template error when `tool_choice` is omitted. Missing GPT OSS tool
|
|
538
|
+
descriptions are serialized as empty strings for its Harmony serializer. See Google's
|
|
539
|
+
[function-calling guidance](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/maas/capabilities/function-calling).
|
|
540
|
+
|
|
541
|
+
Grok also exposes a separate Responses API on Vertex/global:
|
|
542
|
+
|
|
543
|
+
```ts
|
|
544
|
+
const grokResponses = createVertex({ location: "global" }).responsesModel("xai/grok-4.20-reasoning");
|
|
545
|
+
const answer = await generateText({ model: grokResponses, prompt: "Explain vector search briefly." });
|
|
546
|
+
```
|
|
547
|
+
|
|
548
|
+
This factory supports text/image input, streaming, callable tools and native JSON
|
|
549
|
+
Schema output. It uses Google bearer credentials, preserves Vertex provider-data
|
|
550
|
+
and replays conversation/tool history locally with `store: false`. Google does
|
|
551
|
+
not currently support `store: true` or `previous_response_id` on this route.
|
|
552
|
+
Use the shared `temperature`, `maxTokens`, `toolChoice` and `structuredOutput`
|
|
553
|
+
fields; the additional provider options are `top_p`, `parallel_tool_calls` and
|
|
554
|
+
`store: false`. Hosted tools, reasoning controls and arbitrary OpenAI options are
|
|
555
|
+
rejected before network access. `chatModel()` remains the Chat Completions route.
|
|
556
|
+
Live validation passed for `xai/grok-4.20-reasoning` on global: streaming, native
|
|
557
|
+
schema, a single-execution function loop and synthetic-invoice vision (JSON and
|
|
558
|
+
streaming). Other IDs and broader vision quality remain separately unverified. Reproduce with the live smoke
|
|
559
|
+
`--responses-only` flag and `VERTEX_INTEGRATION_MODEL=xai/grok-4.20-reasoning`.
|
|
560
|
+
For vision, use `bun scripts/vertex-vision-live-smoke.ts --responses` with ADC.
|
|
561
|
+
Local functions named `shell`, `computer` or `apply_patch` retain ordinary function
|
|
562
|
+
semantics. The internal OpenAI package supplies only the Responses wire parser;
|
|
563
|
+
requests go to Google, never to the OpenAI API.
|
|
564
|
+
Sources: [Responses](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/grok/responses),
|
|
565
|
+
[function calling](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/grok/capabilities/function-calling),
|
|
566
|
+
[structured output](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/partner-models/grok/capabilities/structured-output).
|
|
567
|
+
|
|
568
|
+
For self-deployed models, `chatModel()` accepts an endpoint ID, `endpoints/<id>`
|
|
569
|
+
or full project/location endpoint resource. Supply `baseURL` and `apiVersion`
|
|
570
|
+
when using a dedicated prediction host or a deployment requiring `v1beta1`.
|
|
571
|
+
Deployment capabilities default conservatively and do not inherit hosted reasoning controls from a publisher-like model name. Native tool-choice, parallel-call and response-format options respect these capability restrictions. Explicitly enable the features
|
|
572
|
+
supported by your deployed model. This invokes an existing endpoint and does
|
|
573
|
+
not provision infrastructure.
|
|
574
|
+
|
|
575
|
+
Many older open-model MaaS IDs retire on October 21, 2026. Check Google's
|
|
576
|
+
[retirement schedule](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/deprecations/open-models)
|
|
577
|
+
before selecting one for a new workload; deployed endpoints are supported as a
|
|
578
|
+
migration path. Contract and SDK-consumer tests cover partner routes. Bounded
|
|
579
|
+
ADC checks also verified gpt-oss 120B streaming, native schema and a client tool
|
|
580
|
+
loop; those results do not certify all models or account entitlements.
|
|
581
|
+
|
|
582
|
+
|
|
583
|
+
### Implicit caching on managed open models
|
|
584
|
+
|
|
585
|
+
Vertex manages implicit context-cache hits on eligible MaaS models. This does
|
|
586
|
+
not use the explicit Google `cachedContents` resource client. Normalized usage
|
|
587
|
+
retains cache-read token counts from OpenAI-style `prompt_tokens_details` or
|
|
588
|
+
Vertex `cachedContentTokenCount`, including terminal streaming usage. Cache
|
|
589
|
+
hits remain service-dependent; see the [host's supported models and conditions](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/maas/use-open-models#context-caching).
|
|
590
|
+
|
|
591
|
+
### Routing publishers through the gateway
|
|
592
|
+
|
|
593
|
+
Register Vertex once; publisher-qualified model IDs remain within that host:
|
|
594
|
+
|
|
595
|
+
```ts
|
|
596
|
+
import { createGateway } from "@zhivex-ai/gateway";
|
|
597
|
+
import { createVertex } from "@zhivex-ai/vertex";
|
|
598
|
+
|
|
599
|
+
const vertex = createVertex({ projectId: "my-project", location: "global" });
|
|
600
|
+
const gateway = createGateway({
|
|
601
|
+
adapters: { vertex },
|
|
602
|
+
scoreTarget: ({ isPrimary }) => isPrimary ? 1 : 0,
|
|
603
|
+
});
|
|
604
|
+
const result = await gateway.generate({
|
|
605
|
+
primary: { provider: "vertex", modelId: "meta/llama-4-maverick-17b-128e-instruct-maas" },
|
|
606
|
+
fallbacks: [{ provider: "vertex", modelId: "claude-sonnet-4-6" }],
|
|
607
|
+
messages: [{ role: "user", content: "Explain this delivery delay." }],
|
|
608
|
+
});
|
|
609
|
+
```
|
|
610
|
+
|
|
611
|
+
The explicit score keeps the primary first; the gateway's default scoring may
|
|
612
|
+
prefer a fallback model. Both destinations require access in your Google Cloud
|
|
613
|
+
project and must be available at the configured location. `providerUsed` remains
|
|
614
|
+
`vertex`; the model ID identifies the publisher. This example's cross-publisher
|
|
615
|
+
fallback is covered by offline SDK tests, not a live availability claim.
|
|
616
|
+
|
|
617
|
+
## Embedding configuration and specialized partner APIs
|
|
618
|
+
|
|
619
|
+
The shared `embed()` / `embedMany()` helpers forward `providerOptions` to Vertex.
|
|
620
|
+
|
|
621
|
+
E5 publisher embeddings use Google OAuth and the OpenMaaS embeddings endpoint:
|
|
622
|
+
|
|
623
|
+
```ts
|
|
624
|
+
const vertex = createVertex({ projectId: "my-project", location: "us-central1" });
|
|
625
|
+
const result = await vertex.embeddingModel("intfloat/multilingual-e5-small-maas")
|
|
626
|
+
.embed({ values: ["query: available shipping methods", "passage: Express shipping takes two days."] });
|
|
627
|
+
```
|
|
628
|
+
|
|
629
|
+
The supported selectors are `intfloat/multilingual-e5-small-maas` and
|
|
630
|
+
`intfloat/multilingual-e5-large-instruct-maas`; publisher resource names are also
|
|
631
|
+
accepted. Supply the model's query/document formatting yourself: small uses
|
|
632
|
+
`query: ` and `passage: ` prefixes; large-instruct uses
|
|
633
|
+
`Instruct: <task description>\nQuery: <query>` for queries and plain documents.
|
|
634
|
+
See the [large-instruct model card](https://huggingface.co/intfloat/multilingual-e5-large-instruct).
|
|
635
|
+
These models
|
|
636
|
+
accept text and reject Google-specific embedding controls. Returned indices are
|
|
637
|
+
validated and vectors are restored to input order. Their MaaS retirement date is
|
|
638
|
+
October 21, 2026, recorded in the SDK catalog. Use a documented regional endpoint:
|
|
639
|
+
`us-central1` or `europe-west4`. Bounded ADC checks in `us-central1` returned
|
|
640
|
+
384 dimensions for small and 1,024 for large; the same small-model request on
|
|
641
|
+
`global` returned HTTP 500. The adapter preserves the caller's location and does
|
|
642
|
+
not silently move requests between regions. See the
|
|
643
|
+
[E5 model card](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/maas/e5/multilingual-e5-small).
|
|
644
|
+
Legacy text models support `outputDimensionality`, `autoTruncate`, `taskType` and
|
|
645
|
+
`title` (retrieval documents only). Embedding 2 supports `outputDimensionality`,
|
|
646
|
+
`documentOcr` and `audioTrackExtraction` in `embedContentConfig`; text-only task
|
|
647
|
+
controls are rejected. Embedding 001 is sent one text at a time; other legacy text
|
|
648
|
+
models are split into batches of five, preserving order and aggregate token usage.
|
|
649
|
+
Google responses must match an explicitly requested `outputDimensionality`.
|
|
650
|
+
Google and E5 vectors must also have consistent dimensions across the entire
|
|
651
|
+
call, including split requests; malformed responses raise `ConfigurationError`.
|
|
652
|
+
|
|
653
|
+
```ts
|
|
654
|
+
const extracted = await vertex.ocr.process({
|
|
655
|
+
modelId: "mistralai/mistral-ocr-2505",
|
|
656
|
+
document: { uri: "https://example.com/report.pdf", mediaType: "application/pdf" },
|
|
657
|
+
includeImages: false
|
|
658
|
+
});
|
|
659
|
+
const completion = await vertex.fim.generate({
|
|
660
|
+
modelId: "codestral-2", prompt: "function answer() {", suffix: "}", maxTokens: 64
|
|
661
|
+
});
|
|
662
|
+
```
|
|
663
|
+
|
|
664
|
+
These Mistral clients require bearer credentials and model access in the selected
|
|
665
|
+
region. OCR accepts inline PDF/image data or HTTP(S) URLs and preserves page
|
|
666
|
+
metadata. FIM uses `rawPredict` / `streamRawPredict` with prompt and suffix.
|
|
667
|
+
The contract tests do not certify account access or every model revision.
|
|
668
|
+
|
|
669
|
+
### DeepSeek image extraction
|
|
670
|
+
|
|
671
|
+
```ts
|
|
672
|
+
const extraction = await vertex.ocr.process({
|
|
673
|
+
modelId: "deepseek-ai/deepseek-ocr-maas",
|
|
674
|
+
document: { uri: "https://example.com/invoice.png", mediaType: "image/png" },
|
|
675
|
+
prompt: "Free OCR",
|
|
676
|
+
});
|
|
677
|
+
```
|
|
678
|
+
|
|
679
|
+
DeepSeek OCR uses the OpenMaaS chat transport with image input; the result represents
|
|
680
|
+
one input image as page index zero. PDF input, page selection and image extraction
|
|
681
|
+
options are rejected for this model. Rasterize a PDF before supplying an image,
|
|
682
|
+
with one call per page, or use Mistral OCR for PDF documents. Truncated or filtered
|
|
683
|
+
responses are rejected instead of being returned as a complete extraction.
|
|
684
|
+
The prompt defaults to `Free OCR`; Mistral OCR rejects this prompt option because
|
|
685
|
+
its document endpoint has a different contract. This facade has offline contract
|
|
686
|
+
coverage and successful OAuth transport, but a live synthetic invoice check
|
|
687
|
+
omitted its heading; OCR content verification is not yet passing. Mistral OCR
|
|
688
|
+
returned 404 in the test project at us-central1. Run the dedicated OCR smoke
|
|
689
|
+
against your enabled models before relying on extraction completeness. DeepSeek OCR MaaS retires on
|
|
690
|
+
October 21, 2026 according to the SDK catalog.
|
|
691
|
+
|
|
692
|
+
OCR and FIM reject reserved wire fields in `providerOptions` (such as `document`,
|
|
693
|
+
`messages`, `pages`, `prompt` or `suffix`) that conflict with their dedicated input
|
|
694
|
+
fields. Other native options remain available through `providerOptions`.
|
|
695
|
+
|
|
696
|
+
Live embedding checks also verified task type, disabled truncation and 256
|
|
697
|
+
dimensions on text-embedding-005, plus inline PNG input with 768 dimensions on
|
|
698
|
+
gemini-embedding-2. Additional bounded calls verified a one-page PDF with `documentOcr: true` and one-second WAV audio at 768 dimensions. A separate live check embedded one second of Google’s public highway video from GCS at 1 FPS with audio extraction disabled: 128 dimensions, unit norm and 66 input tokens. This verifies that configuration, not semantic retrieval quality or inline video. Reproduce with `bun scripts/vertex-video-embedding-live-smoke.ts` using ADC credentials. Embedding 2 advertises image, document and audio input capabilities; legacy Google and E5 models remain text-only.
|
|
699
|
+
|
|
700
|
+
|
|
701
|
+
### Legacy multimodal embeddings
|
|
702
|
+
|
|
703
|
+
Basic image/text retrieval at 128 dimensions passed live for this model and
|
|
704
|
+
Gemini Embedding 2 using the synthetic invoice fixture. Reproduce with
|
|
705
|
+
`bun scripts/vertex-image-retrieval-live-smoke.ts` and ADC credentials from the
|
|
706
|
+
repository root. This checks one matching description against two distractors;
|
|
707
|
+
it does not establish general multimodal retrieval quality.
|
|
708
|
+
|
|
709
|
+
`embeddingModel("multimodalembedding@001")` supports text and PNG/JPEG image
|
|
710
|
+
values, preserving one vector per input. Set `providerOptions.outputDimensionality`
|
|
711
|
+
to 128, 256, 512 or 1408. Each input uses its own native prediction request.
|
|
712
|
+
Google Cloud bearer credentials and a supported regional location are required.
|
|
713
|
+
|
|
714
|
+
For combined modalities or video, use the native client:
|
|
715
|
+
|
|
716
|
+
```ts
|
|
717
|
+
const result = await vertex.multimodalEmbeddings.embed({
|
|
718
|
+
text: "A road with vehicles",
|
|
719
|
+
video: { uri: "gs://my-bucket/road.mp4", mediaType: "video/mp4" },
|
|
720
|
+
videoSegmentConfig: { startOffsetSec: 0, endOffsetSec: 8, intervalSec: 4 },
|
|
721
|
+
});
|
|
722
|
+
// result.textEmbedding: 1408 dimensions
|
|
723
|
+
// result.videoEmbeddings: individual 1408-dimensional vectors with
|
|
724
|
+
// startOffsetSec and endOffsetSec for every returned segment.
|
|
725
|
+
```
|
|
726
|
+
|
|
727
|
+
`outputDimensionality` is available for text/image-only requests. Any request
|
|
728
|
+
containing video uses 1408 dimensions and rejects that option locally, and the unified `embed` method directs video callers to the native client
|
|
729
|
+
so no segments are silently discarded. Media accepts inline bytes or `gs://`
|
|
730
|
+
object URIs. Video audio is not embedded by this model. No token usage is invented
|
|
731
|
+
when its response provides none. See the [native API reference](https://docs.cloud.google.com/gemini-enterprise-agent-platform/reference/models/multimodal-embeddings-api).
|
|
732
|
+
|
|
733
|
+
## Virtual Try-On
|
|
734
|
+
|
|
735
|
+
```ts
|
|
736
|
+
const cloud = createVertex({ location: "us-central1" }); // Google Cloud ADC
|
|
737
|
+
const result = await cloud.virtualTryOn.generate({
|
|
738
|
+
personImage: { uri: "gs://my-bucket/person.png", mediaType: "image/png" },
|
|
739
|
+
productImage: { uri: "gs://my-bucket/shirt.jpg", mediaType: "image/jpeg" },
|
|
740
|
+
count: 1,
|
|
741
|
+
outputMimeType: "image/jpeg",
|
|
742
|
+
providerOptions: { outputOptions: { compressionQuality: 85 } }
|
|
743
|
+
});
|
|
744
|
+
// result.images contains inline bytes or GCS URIs; result.filtered retains reasons.
|
|
745
|
+
```
|
|
746
|
+
|
|
747
|
+
This dedicated client calls `virtual-try-on-001:predict` using Google bearer
|
|
748
|
+
authentication. Person and product are named inputs, not positional chat images.
|
|
749
|
+
It supports PNG/JPEG as inline `data` or GCS `uri`, a product mask/configuration,
|
|
750
|
+
1–4 outputs, native prediction parameters and optional `outputStorageUri`.
|
|
751
|
+
Inline inputs are limited to 7 MiB each. Unknown/malformed image responses fail
|
|
752
|
+
explicitly; filtering reasons remain available even when no image is returned.
|
|
753
|
+
Shared deadlines, cancellation and retries apply through the prediction transport.
|
|
754
|
+
Use `virtualTryOn.generate()` instead of the Gemini language/image factories.
|
|
755
|
+
A bounded ADC smoke passed in `us-central1` with the two public images from
|
|
756
|
+
Google's notebook, returning one JPEG (311,148 bytes). This proves transport and
|
|
757
|
+
output-format behavior. A second ADC request with both images inline and
|
|
758
|
+
`outputOptions.compressionQuality:85` returned a valid JPEG (347,627 bytes).
|
|
759
|
+
Visual inspection of that example confirmed the blue V-neck sweater replaced
|
|
760
|
+
the hoodie while preserving the subject's pose and field background. This is
|
|
761
|
+
one inspected example, not a quality benchmark; masks remain unverified live.
|
|
762
|
+
Reproduce with `bun scripts/vertex-virtual-try-on-live-smoke.ts` and ADC configured;
|
|
763
|
+
add `--inline --save-artifacts` to exercise inline input and save the two public
|
|
764
|
+
fixtures and generated JPEG to a new local temporary directory for inspection.
|
|
765
|
+
It creates no persistent cloud resources.
|
|
766
|
+
See [Google's Virtual Try-On guide](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/capabilities/generate-virtual-try-on-images).
|
|
767
|
+
|
|
768
|
+
## Vertex Interactions (experimental)
|
|
769
|
+
|
|
770
|
+
```ts
|
|
771
|
+
const vertex = createVertex({ projectId: "my-project", location: "global", getAccessToken });
|
|
772
|
+
const music = await vertex.interactions.create({
|
|
773
|
+
modelId: "lyria-3-clip-preview", input: "A short instrumental jazz piece", store: false
|
|
774
|
+
});
|
|
775
|
+
// Consume music.outputs directly; this Lyria route does not support storage.
|
|
776
|
+
```
|
|
777
|
+
|
|
778
|
+
Lyria 3 Clip requires `store: false` on the tested Vertex route; the adapter sends
|
|
779
|
+
that default explicitly and rejects `store: true`. Do not use its returned ID as
|
|
780
|
+
proof of stored retrieval or resumption. Other model/agent storage options retain
|
|
781
|
+
their own contracts.
|
|
782
|
+
|
|
783
|
+
Interactions model availability is separate from `generateContent`: a live call
|
|
784
|
+
with `gemini-3.7-flash` was rejected as unsupported. The Vertex reference names
|
|
785
|
+
Lyria 3 and Deep Research, while the video guides document Gemini Omni. Choose
|
|
786
|
+
a model or agent explicitly supported by this API. Multiple model-output steps
|
|
787
|
+
are preserved, including Lyria lyrics, captions and audio.
|
|
788
|
+
|
|
789
|
+
The client uses the project-scoped `v1beta1` Interactions endpoint and supports
|
|
790
|
+
create/get/list/cancel/delete, streaming and resumption with `lastEventId`. Streaming
|
|
791
|
+
preserves native events, event IDs and multimedia blocks as `provider-data`, in
|
|
792
|
+
addition to normalized text, tool calls and terminal status. Interrupted streams
|
|
793
|
+
throw instead of reporting successful completion. `cancel()` targets background interactions using the project-scoped
|
|
794
|
+
`interactions/{id}/cancel` route verified in the official Google Gen AI SDK.
|
|
795
|
+
|
|
796
|
+
Live validation passed for stored Omni 1.1 background video: close the initial
|
|
797
|
+
stream, resume the same ID using its cursor, retrieve completed outputs and
|
|
798
|
+
delete the interaction (absence verified with HTTP 404). A background GET stream
|
|
799
|
+
may exhaust currently available events before completion; the client raises its
|
|
800
|
+
missing-terminal error. Retain the latest cursor and observe that same interaction
|
|
801
|
+
again with an application deadline instead of creating another generation.
|
|
802
|
+
The bounded smoke demonstrates this flow:
|
|
803
|
+
`VERTEX_INTERACTIONS_RESUME_MODEL=gemini-omni-1.1-flash-preview bun scripts/vertex-interactions-resume-live-smoke.ts`
|
|
804
|
+
with ADC configured. Its native video settings are 3 seconds and 360p, and it
|
|
805
|
+
cleans up the owned interaction. Mid-tool-argument recovery remains covered by
|
|
806
|
+
contract tests, not this video smoke.
|
|
807
|
+
|
|
808
|
+
For text resumption, persist the interaction ID and the latest delivered
|
|
809
|
+
`provider-data.data.event_id` together with the output already consumed, then
|
|
810
|
+
call `resume({ id, lastEventId })`. Keep event IDs opaque: pass the original value;
|
|
811
|
+
the client encodes the query parameter. It does not persist application output or
|
|
812
|
+
automatically replay an interrupted stream. Contract tests cover a text stream
|
|
813
|
+
ending before its terminal event and resuming without duplicating prior text.
|
|
814
|
+
To resume during tool arguments, also pass `previousEvents`: the ordered native
|
|
815
|
+
Vertex event objects already consumed, ending at the event whose `event_id`
|
|
816
|
+
exactly matches `lastEventId`. Include all prior tool start/delta/stop events so
|
|
817
|
+
pending arguments and completed-call IDs can be reconstructed. The history is
|
|
818
|
+
used locally and never sent to Google; prior text, metadata and completed tool
|
|
819
|
+
calls are not emitted again. Limits are 16,384 events and 8 MiB of serialized
|
|
820
|
+
history, with the existing per-call argument bounds. `VertexInteractionResumeInput`
|
|
821
|
+
is exported for typed consumers. Persist cursors together with application output
|
|
822
|
+
and tool execution state; this API does not provide durable exactly-once execution.
|
|
823
|
+
|
|
824
|
+
`musicGenerationModel("lyria-3-clip-preview")` and Lyria 3 Pro route synchronous
|
|
825
|
+
music generation through Interactions, including optional image input. Native
|
|
826
|
+
asynchronous workflows should use `interactions` directly. Model access, location,
|
|
827
|
+
preview availability and billing remain governed by Google Cloud. Bounded ADC
|
|
828
|
+
checks verified Lyria 3 Clip creation and streaming on global, including text,
|
|
829
|
+
inline MP3 audio and successful terminal status. Persistence, resumption, Pro
|
|
830
|
+
and other model/agent variants require separate live verification.
|
|
831
|
+
|
|
832
|
+
Hosted tools use the Interactions contract: Google Maps maps `enableWidget` to
|
|
833
|
+
`enable_widget`, and `vertexSearch` accepts native `engine`/`datastores` config
|
|
834
|
+
and maps to Vertex AI Search retrieval. Native `retrieval` config can also be
|
|
835
|
+
supplied. Gemini Developer File Search is rejected on this Vertex route.
|
|
836
|
+
Callable tool streams accept the documented `arguments_delta` discriminator.
|
|
837
|
+
|
|
838
|
+
References: [Vertex Interactions](https://docs.cloud.google.com/gemini-enterprise-agent-platform/reference/models/interactions-api),
|
|
839
|
+
[Lyria music generation](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/music/generate-music),
|
|
840
|
+
[Mistral on Vertex](https://docs.mistral.ai/inference/deployment/cloud-deployments/vertex).
|
|
841
|
+
|
|
842
|
+
|
|
843
|
+
## Vertex Live transport
|
|
844
|
+
|
|
845
|
+
Buffered transcription and speech generation, plus speech stream setup, honor
|
|
846
|
+
configured HTTP retries with deadline-bound backoff. Once a speech stream delivers
|
|
847
|
+
audio, parsing or transport failures propagate without replaying the response.
|
|
848
|
+
This HTTP speech behavior is separate from Live WebSocket reconnection.
|
|
849
|
+
|
|
850
|
+
The connection `timeoutMs` includes credential acquisition. The transport receives
|
|
851
|
+
the remaining time after ADC or a custom token resolver returns. Cancellation
|
|
852
|
+
while waiting for credentials prevents a later WebSocket connection. The original
|
|
853
|
+
caller signal is retained for transport/session cancellation; the temporary
|
|
854
|
+
credential deadline does not abort an established session.
|
|
855
|
+
|
|
856
|
+
Native tool cancellations emit `realtime-tool-call-cancellation` with
|
|
857
|
+
`toolCallIds`. Callback sessions reject results for those IDs and suppress
|
|
858
|
+
cancelled call replays. `session.toolCallSignal(id)` aborts independently of event
|
|
859
|
+
consumption. `streamLiveAgent` uses it to stop approval waits and signal running
|
|
860
|
+
executors through `context.abortSignal`, then continues the conversation. Custom
|
|
861
|
+
session implementations must expose this optional method for the same behavior.
|
|
862
|
+
Executors must cooperate with abort; already-applied side effects are not undone.
|
|
863
|
+
An interrupted durable execution retains its indeterminate `running` journal entry.
|
|
864
|
+
|
|
865
|
+
For explicit `session.interrupt()`, connect with
|
|
866
|
+
`providerOptions: { realtimeInputConfig: { automaticActivityDetection: { disabled: true } } }`.
|
|
867
|
+
The method sends a manual activity start/end pair and preserves the connection.
|
|
868
|
+
With automatic VAD enabled, speech drives interruptions; calling `interrupt()`
|
|
869
|
+
in that mode fails locally. Dedicated Live Translate rejects this operation.
|
|
870
|
+
Clients must still discard queued playback when they receive an interrupted event.
|
|
871
|
+
|
|
872
|
+
`session.sendMedia()` sends visual frames as native `realtimeInput.mediaChunks`.
|
|
873
|
+
Supply discrete image frames, rather than a video container. The synthetic invoice reading smoke passed with a one-second gap before the
|
|
874
|
+
text question. Immediate image/text sends completed but misread the amount;
|
|
875
|
+
applications must account for realtime media processing and turn ordering.
|
|
876
|
+
|
|
877
|
+
Dedicated Live Translate sends `generationConfig.translationConfig`, with
|
|
878
|
+
`translation.targetLanguage` and optional
|
|
879
|
+
`providerOptions.translationConfig.echoTargetLanguage`. Source language is
|
|
880
|
+
detected automatically; explicit `translation.sourceLanguage` is rejected.
|
|
881
|
+
Output transcription requests AUDIO and TEXT, matching Google's introductory
|
|
882
|
+
notebook. Use `session.setInputMuted(true)` after the final audio chunk to send
|
|
883
|
+
`audioStreamEnd`; unmuting allows subsequent audio input.
|
|
884
|
+
Use location `global`; other standard regions fail locally for this model.
|
|
885
|
+
The corrected setup is accepted and one bounded live probe received audio,
|
|
886
|
+
but translation text and completion are not yet live-certified. A subsequent
|
|
887
|
+
v1beta1 probe returned a quota-exceeded message inside a text part instead of
|
|
888
|
+
translation output. Diagnostics report this separately from protocol failures.
|
|
889
|
+
See the maintainer smoke guide for the reproducible `--translate` check.
|
|
890
|
+
|
|
891
|
+
`session.update({ instructions: "New instructions" })` sends a system-content
|
|
892
|
+
update without starting a response. Other configuration changes require a new
|
|
893
|
+
connection and are rejected locally. Repeating the same instructions or sending
|
|
894
|
+
an empty update is a no-op; removing instructions requires reconnecting.
|
|
895
|
+
|
|
896
|
+
Enable resumption with `providerOptions: { sessionResumption: {} }` when
|
|
897
|
+
connecting. Save handles only from `realtime-session-resumption` events with
|
|
898
|
+
`resumable: true`, then reconnect with
|
|
899
|
+
`providerOptions: { sessionResumption: { handle } }`. Reconnection is managed by
|
|
900
|
+
the caller. A `realtime-go-away` event exposes `timeLeftMs` when the server
|
|
901
|
+
announces impending closure. A `realtime-response-complete` event with
|
|
902
|
+
`reason: "interrupted"` signals that playback should stop and queued audio
|
|
903
|
+
should be discarded; a subsequent `turn-complete` can still follow.
|
|
904
|
+
|
|
905
|
+
Live sessions use a full `projects/.../locations/.../publishers/google/models/...`
|
|
906
|
+
model resource and OAuth bearer headers. Node and Bun use the shared authenticated
|
|
907
|
+
WebSocket transport by default. Other runtimes can supply a `realtimeConnectionFactory`;
|
|
908
|
+
the browser transport cannot attach bearer headers. The project scope comes from `baseURL` when it
|
|
909
|
+
contains a project/location resource, otherwise from the provider project and
|
|
910
|
+
location. The live smoke uses `ws` with a bounded queue and validates a synthetic
|
|
911
|
+
text turn, audio output and transcription in `us-central1`.
|
|
912
|
+
|
|
913
|
+
Live callable tools arrive as native `toolCall.functionCalls` messages and are
|
|
914
|
+
normalized to `realtime-tool-call` events. Send results with `sendToolResult()`
|
|
915
|
+
using the received call ID. `generation-complete` can precede a tool follow-up;
|
|
916
|
+
a tool-only turn may also complete without audio. When collecting a spoken tool
|
|
917
|
+
response, wait for the subsequent `turn-complete` with audio after the tool result.
|
|
918
|
+
|
|
919
|
+
## Grounded generation
|
|
920
|
+
|
|
921
|
+
`groundedLanguageModel()` uses Google Search and returns source URLs, normalized
|
|
922
|
+
`usage`, and the original response containing grounding supports and search
|
|
923
|
+
entry-point metadata. Retryable HTTP errors respect `maxRetries` and the overall
|
|
924
|
+
request deadline. A live `gemini-3.7-flash` check returned seven sources and
|
|
925
|
+
attribution supports; reproduce with `bun scripts/vertex-grounding-live-smoke.ts`
|
|
926
|
+
using ADC. The separate Maps smoke also passed; private Vertex AI Search still
|
|
927
|
+
requires a configured test datastore and separate validation.
|
|
928
|
+
|
|
929
|
+
Gemini generation and stream setup honor `maxRetries` for retryable HTTP errors.
|
|
930
|
+
`timeoutMs` bounds the request including retry backoff. Once a stream starts
|
|
931
|
+
delivering output, subsequent stream errors propagate without automatic replay.
|
|
932
|
+
|
|
933
|
+
Google Maps coordinates must be finite, with latitude within [-90, 90] and longitude
|
|
934
|
+
within [-180, 180]; `enableWidget` must be boolean. Invalid configurations fail
|
|
935
|
+
before sending a request. `bun scripts/vertex-maps-live-smoke.ts` checks place
|
|
936
|
+
sources and attribution metadata. The latest live check on gemini-3.7-flash/global
|
|
937
|
+
passed with two Maps places and four attribution supports. An earlier 429 was
|
|
938
|
+
transient in the tested project; availability elsewhere is not implied.
|
|
939
|
+
|
|
940
|
+
## Context caching
|
|
941
|
+
|
|
942
|
+
Cache creation and deletion honor explicit `maxRetries`, with backoff bounded by
|
|
943
|
+
`timeoutMs`; retries are disabled by default. Creation has no deduplication key,
|
|
944
|
+
so an uncertain result can require reconciliation before retrying. Deletion
|
|
945
|
+
accepts HTTP 204 and preserves 404 as an error rather than claiming it performed
|
|
946
|
+
the deletion. A separate GET 404 can establish resource absence during cleanup.
|
|
947
|
+
|
|
948
|
+
```ts
|
|
949
|
+
import { updateContextCache } from "@zhivex-ai/core";
|
|
950
|
+
|
|
951
|
+
await updateContextCache({ provider: vertex, name: cache.name, ttl: "3600s" });
|
|
952
|
+
// Alternatively set expireTime to an RFC 3339 timestamp; do not set both.
|
|
953
|
+
```
|
|
954
|
+
|
|
955
|
+
The helper is also exported by `@zhivex-ai/sdk`. The direct
|
|
956
|
+
`vertex.caches.update()` method remains available. Other providers without this
|
|
957
|
+
optional operation throw `UnsupportedFeatureError` through the helper.
|
|
958
|
+
|
|
959
|
+
Context-cache `get()` and `list()` honor `maxRetries` for transient HTTP failures;
|
|
960
|
+
`timeoutMs` also bounds their retry backoff. Pagination tokens remain unchanged
|
|
961
|
+
across attempts.
|
|
962
|
+
|
|
963
|
+
The context-cache lifecycle smoke is `bun scripts/vertex-cache-live-smoke.ts --gcs`.
|
|
964
|
+
It uses Google's public sample PDF plus a synthetic verification code, a short
|
|
965
|
+
TTL and cleanup. The live run passed creation, read, expiration update, code
|
|
966
|
+
retrieval, nonzero cached-token usage and deletion confirmed by a subsequent 404.
|
|
967
|
+
Use `--diverse-text` instead of `--gcs` for the verified synthetic text lifecycle,
|
|
968
|
+
which also passed both `expireTime` and `ttl` updates. The default repetitive
|
|
969
|
+
text fixture still has an unresolved creation error.
|
|
970
|
+
For uncertain creation outcomes, use
|
|
971
|
+
`bun scripts/vertex-cache-reconcile-smoke.ts <state-file>` to enumerate caches and
|
|
972
|
+
remove only the uniquely named smoke resource.
|
|
973
|
+
|
|
974
|
+
When creating a context cache, set either `ttl` or `expireTime`. Supply model,
|
|
975
|
+
contents, system instructions, tools, display name and expiry through their
|
|
976
|
+
dedicated fields; conflicting `providerOptions` fail locally. Other native
|
|
977
|
+
options, such as `kmsKeyName`, are preserved.
|
|
978
|
+
|
|
979
|
+
Cache creation accepts bare Google model IDs, `publishers/google/models/<id>` and
|
|
980
|
+
fully qualified project model resources. The repetitive-text live follow-up reached HTTP
|
|
981
|
+
with 28,752 text characters but received a one-token/minimum-size error from
|
|
982
|
+
Vertex. That discrepancy remains under investigation. Both varied text and GCS
|
|
983
|
+
PDF lifecycles passed on gemini-2.5-flash/us-central1; other model and region
|
|
984
|
+
combinations remain unverified.
|
|
985
|
+
|
|
986
|
+
For cache encryption, `providerOptions.kmsKeyName` is a convenience field mapped
|
|
987
|
+
to the REST `encryptionSpec.kmsKeyName` object. You can instead pass native
|
|
988
|
+
`providerOptions.encryptionSpec`; supplying both forms is rejected. The KMS key
|
|
989
|
+
must be a full `projects/.../locations/.../keyRings/.../cryptoKeys/...` resource.
|
|
990
|
+
This mapping has contract coverage; use with an actual KMS key remains unverified.
|
|
991
|
+
|
|
992
|
+
## Gemini token counting
|
|
993
|
+
|
|
994
|
+
```ts
|
|
995
|
+
const count = await vertex.gemini.countTokens({
|
|
996
|
+
modelId: "gemini-2.5-flash",
|
|
997
|
+
messages: [{ role: "user", parts: [{ type: "text", text: "Hello" }] }],
|
|
998
|
+
timeoutMs: 15_000,
|
|
999
|
+
});
|
|
1000
|
+
console.log(count.inputTokens);
|
|
1001
|
+
```
|
|
1002
|
+
|
|
1003
|
+
The client accepts system instructions, SDK tools, multimodal message parts and
|
|
1004
|
+
native `generationConfig`. It returns validated `inputTokens`, optional
|
|
1005
|
+
`totalBillableCharacters` and `rawResponse`. Claude uses the separate
|
|
1006
|
+
`vertex.claude.countTokens()` native message contract. Token counting does not
|
|
1007
|
+
create a cache or prove that a subsequent cache creation will succeed.
|
|
1008
|
+
|
|
1009
|
+
## Dedicated audio transcription
|
|
1010
|
+
|
|
1011
|
+
Use `transcriptionModel("gemini-3.5-transcribe-preview")` with bearer credentials
|
|
1012
|
+
and `location: "global"` for recorded audio. The adapter sends audio-only contents
|
|
1013
|
+
and native recognition configuration; `prompt` is rejected for this model.
|
|
1014
|
+
|
|
1015
|
+
Dedicated Transcribe and Live Translate models reject `vertex(modelId)`,
|
|
1016
|
+
`languageModel()` and `groundedLanguageModel()` locally. Use the transcription
|
|
1017
|
+
factory above or `realtimeModel()` for Live variants. Transcription and speech
|
|
1018
|
+
capabilities describe those audio adapters: they do not advertise chat tools,
|
|
1019
|
+
vision, grounding, URL context, cache, batch or endpoint prediction operations.
|
|
1020
|
+
Provider-level resource clients remain separate surfaces.
|
|
1021
|
+
|
|
1022
|
+
```ts
|
|
1023
|
+
const result = await vertex.transcriptionModel("gemini-3.5-transcribe-preview").transcribe({
|
|
1024
|
+
audio: { data: audioBytes, mediaType: "audio/wav" },
|
|
1025
|
+
language: "en-US",
|
|
1026
|
+
providerOptions: {
|
|
1027
|
+
audioTranscriptionConfig: {
|
|
1028
|
+
customVocabulary: ["Zhivex"],
|
|
1029
|
+
wordTimestamp: true,
|
|
1030
|
+
diarization: true,
|
|
1031
|
+
mode: "VERBATIM"
|
|
1032
|
+
}
|
|
1033
|
+
},
|
|
1034
|
+
timeoutMs: 45_000
|
|
1035
|
+
});
|
|
1036
|
+
console.log(result.text);
|
|
1037
|
+
for (const part of result.transcriptions) {
|
|
1038
|
+
console.log(part.speakerLabel, part.languageCode, part.words);
|
|
1039
|
+
}
|
|
1040
|
+
```
|
|
1041
|
+
|
|
1042
|
+
`text` combines all response fragments. Native `transcriptions` preserve speaker
|
|
1043
|
+
labels, language codes and word offsets as duration strings; `rawResponse`
|
|
1044
|
+
retains the original response. The shared `transcribeAudio()` helper exposes its
|
|
1045
|
+
shared result contract; use the provider model directly for typed native details.
|
|
1046
|
+
`language` maps to `languageCodes`; conflicting hints fail before sending.
|
|
1047
|
+
`SMART` mode cannot combine with `wordTimestamp` or `diarization`, and custom
|
|
1048
|
+
vocabulary accepts up to 1,000 nonempty terms. Existing Gemini audio-understanding
|
|
1049
|
+
models retain prompted transcription behavior.
|
|
1050
|
+
|
|
1051
|
+
One live v1/global check passed synthetic speech and word timestamps. Diarization,
|
|
1052
|
+
SMART formatting and custom vocabulary quality remain unverified. The
|
|
1053
|
+
synchronous factory rejects the separate Live model. See [Google's transcription guide](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-5-transcribe).
|
|
1054
|
+
|
|
1055
|
+
### Live transcription
|
|
1056
|
+
|
|
1057
|
+
Use `realtimeModel("gemini-3.5-transcribe-live-preview")` on global with bearer
|
|
1058
|
+
authentication. The session waits for setup acknowledgement before accepting
|
|
1059
|
+
audio. It requests text output and defaults to input transcription enabled.
|
|
1060
|
+
|
|
1061
|
+
```ts
|
|
1062
|
+
const session = await vertex.realtimeModel!("gemini-3.5-transcribe-live-preview").connect({
|
|
1063
|
+
mode: "transcription",
|
|
1064
|
+
inputAudioTranscription: { languageCodes: ["en-US"], customVocabulary: ["Zhivex"] }
|
|
1065
|
+
}, { timeoutMs: 15_000 });
|
|
1066
|
+
try {
|
|
1067
|
+
await session.sendAudio({ data: pcmBytes, mediaType: "audio/pcm;rate=16000" });
|
|
1068
|
+
await session.setInputMuted(true);
|
|
1069
|
+
for await (const event of session.eventStream()) {
|
|
1070
|
+
if (event.type === "realtime-provider-data") console.log(event.data);
|
|
1071
|
+
if (event.type === "realtime-transcript" && event.isFinal) {
|
|
1072
|
+
console.log(event.text);
|
|
1073
|
+
break;
|
|
1074
|
+
}
|
|
1075
|
+
}
|
|
1076
|
+
} finally {
|
|
1077
|
+
await session.close();
|
|
1078
|
+
}
|
|
1079
|
+
```
|
|
1080
|
+
|
|
1081
|
+
Interim hypotheses arrive as `realtime-provider-data` with
|
|
1082
|
+
`data.type: "vertex_transcription_interim"` and the native `transcription` object.
|
|
1083
|
+
They replace the previous hypothesis; do not append them as text deltas. Final
|
|
1084
|
+
segments arrive as `realtime-transcript`, `role: "user"`, `isFinal: true`.
|
|
1085
|
+
Native metadata remains available on the events.
|
|
1086
|
+
|
|
1087
|
+
`setInputMuted(true)` sends `audioStreamEnd` once and discards subsequent audio
|
|
1088
|
+
frames while muted. `setInputMuted(false)` permits audio for the next segment.
|
|
1089
|
+
This model produces no generated audio and rejects text/image input, tools,
|
|
1090
|
+
system instructions, reasoning, word timestamps and diarization. Shared
|
|
1091
|
+
`inputTranscription.language` maps to native language hints. Native language
|
|
1092
|
+
codes, custom vocabulary and VERBATIM/SMART modes are accepted; quality depends
|
|
1093
|
+
on the selected language and audio.
|
|
1094
|
+
|
|
1095
|
+
A bounded real v1/global session passed with three interim hypotheses and a
|
|
1096
|
+
final hello-world transcript, without generated audio. Run
|
|
1097
|
+
`bun scripts/vertex-transcription-realtime-smoke.ts` for the owned audio fixture.
|
|
1098
|
+
|
|
1099
|
+
## Verification and remaining limits
|
|
1100
|
+
|
|
1101
|
+
Live results are scoped to the tested model, location and project. The
|
|
1102
|
+
[readiness summary](../../docs/maintainers/VERTEX_READINESS.md) separates
|
|
1103
|
+
implementation, external blockers, quality evaluation and delivery status.
|
|
1104
|
+
|
|
1105
|
+
| Area | Verification evidence | Remaining limitations |
|
|
1106
|
+
| --- | --- | --- |
|
|
1107
|
+
| Gemini and partners | Gemini text/tools/stream/schema; GPT OSS, Gemma 4, DeepSeek V3.2, Kimi K2 and MiniMax M2 tool loop/stream/schema; Qwen3-Next Instruct schema | Other models need separate checks; Claude returns 429, Llama 4 Scout returns 404 in us-east5 |
|
|
1108
|
+
| Embeddings | Google text and selected media, including inline/GCS video with audio extraction; regional E5 vectors; basic text retrieval for Google and both E5 variants | No broad quality benchmark; additional media retrieval remains unverified |
|
|
1109
|
+
| Google caches | Varied text and public PDF create/read/use/delete; ttl and expireTime updates | Repetitive-text fixture error unresolved; real CMEK unverified |
|
|
1110
|
+
| Google batch | GCS output and separate cancellation; job/bucket removal confirmed by 404 | BigQuery API is disabled in the test project; partner batch unverified |
|
|
1111
|
+
| Grounding | Google Search and Maps sources and attribution | Private Search needs a test datastore |
|
|
1112
|
+
| Specialized APIs | Lyria 3 Clip and Omni video | OCR content/access, Codestral and additional Interactions workflows remain pending |
|
|
1113
|
+
| Live | Audio, tools, resume, instruction update and interruption | Tool cancellation and dedicated translation lack live certification |
|
|
1114
|
+
| Deployed endpoints | Contracts and installed Node/Bun local TLS tests; real Google gRPC auth reached a missing-resource error | Successful inference requires a real deployment and schema; none found in global/us-central1 |
|
|
1115
|
+
|
|
1116
|
+
Gemma 4 additionally passed image-based invoice extraction with native schema
|
|
1117
|
+
and streaming on global. The test uses a repository-owned synthetic PNG and
|
|
1118
|
+
asserts its amount/currency without putting the answers in the prompt or schema.
|
|
1119
|
+
Run `bun scripts/vertex-vision-live-smoke.ts` with ADC. This is one fixture, not
|
|
1120
|
+
a broad OCR quality benchmark.
|
|
1121
|
+
|
|
1122
|
+
Gemma 4 and DeepSeek V3.2 additionally passed explicit thinking on/off and
|
|
1123
|
+
streamed reasoning checks on global. Reasoning remains in Vertex provider-data
|
|
1124
|
+
parts, separate from answer text. Run `bun scripts/vertex-reasoning-live-smoke.ts`
|
|
1125
|
+
with either exact model ID as its argument; the script prints counts, not reasoning
|
|
1126
|
+
content. GLM thinking toggle mapping has contract tests; live access is unresolved.
|
|
1127
|
+
|
|
1128
|
+
MiniMax M2 now passes streaming, native schema and tool-loop checks. Its documented
|
|
1129
|
+
leading `<think>` envelope is separated from answer text into Vertex provider-data
|
|
1130
|
+
with `type: "vertex_inline_thinking"` and the exact envelope in `content`.
|
|
1131
|
+
Subsequent assistant history restores that envelope for the host. Streaming
|
|
1132
|
+
supports split delimiters; literal tags within answer text remain unchanged.
|
|
1133
|
+
Incomplete envelopes and envelopes exceeding 1,048,576 characters raise an error.
|
|
1134
|
+
This normalization applies only to managed `minimaxai/minimax-m2-maas`, not to
|
|
1135
|
+
custom deployments or arbitrary model IDs.
|
|
1136
|
+
|
|
1137
|
+
Publisher checks are operation-specific. In the current test project, GLM 5.2
|
|
1138
|
+
returned 429 for all three chat scenarios, Qwen3-Next Instruct passed native schema
|
|
1139
|
+
but returned 429 for streaming/tools, and Grok 4.3 returned 404 on global.
|
|
1140
|
+
Mistral Medium 3 and Small 2503 also returned 404 in us-central1; Google reported
|
|
1141
|
+
NOT_FOUND and an ambiguous missing-model/access message for Small. Verify model
|
|
1142
|
+
enablement and project access in Model Garden before attempting certification.
|
|
1143
|
+
These statuses do not establish a permanent model limitation. The implementation
|
|
1144
|
+
ledger records exact IDs and locations; successful calls do not override catalog
|
|
1145
|
+
retirement dates or certify untested vision/reasoning features.
|
|
1146
|
+
|
|
1147
|
+
Repository smoke commands require configured credentials and can incur usage.
|
|
1148
|
+
Use `VERTEX_INTEGRATION_USE_ADC=1` and remove API-key/token environment overrides
|
|
1149
|
+
when selecting ADC. No access token needs to be pasted into source code.
|
|
1150
|
+
|
|
1151
|
+
| Check | Command from the repository root |
|
|
1152
|
+
| --- | --- |
|
|
1153
|
+
| Selected MaaS chat | `bun scripts/vertex-live-smoke.ts --chat-only` (set `VERTEX_INTEGRATION_MODEL` and `VERTEX_LOCATION`) |
|
|
1154
|
+
| Claude generation/counting | `bun scripts/vertex-claude-live-smoke.ts` |
|
|
1155
|
+
| Text retrieval | `bun scripts/vertex-retrieval-live-smoke.ts` |
|
|
1156
|
+
| Inline video embeddings | `bun scripts/vertex-video-embedding-live-smoke.ts --inline` |
|
|
1157
|
+
| Video audio extraction | `bun scripts/vertex-video-embedding-live-smoke.ts --extract-audio` (add `--inline` for inline bytes) |
|
|
1158
|
+
| Cache lifecycle | `bun scripts/vertex-cache-live-smoke.ts --diverse-text` or `--gcs` |
|
|
1159
|
+
| Batch lifecycle | `bun scripts/vertex-batch-live-smoke.ts start STATE_FILE`, then `status`, `cancel`, `cleanup`, or `verify-cleanup` with the same state file |
|
|
1160
|
+
|
|
1161
|
+
E5 small retrieval uses explicit `query: ` and `passage: ` prefixes supplied by
|
|
1162
|
+
the caller. The retrieval smoke checks a tiny synthetic corpus, not model quality.
|
|
1163
|
+
Batch cancellation is complete only after a GET reports JOB_STATE_CANCELLED;
|
|
1164
|
+
a successful cancel request alone does not establish that terminal state.
|
|
1165
|
+
|
|
1166
|
+
The SDK-owned catalog now includes partner and specialized model IDs. Optional
|
|
1167
|
+
`entry.lifecycle` carries source-backed deprecation and retirement dates. It is
|
|
1168
|
+
snapshot metadata, not a live availability check or proof of account access;
|
|
1169
|
+
the frozen legacy core catalog remains unchanged.
|
|
1170
|
+
|
|
1171
|
+
The generated [catalog and adapter matrix](../../docs/maintainers/VERTEX_CATALOG_MATRIX.md)
|
|
1172
|
+
lists all 62 inventory entries with their selected factory, lifecycle and declared
|
|
1173
|
+
capabilities. It is an offline contract audit; use the live evidence table above
|
|
1174
|
+
to assess validation of a specific route.
|
|
1175
|
+
|
|
1176
|
+
|
|
1177
|
+
For local artifact validation, run `bun run build` followed by
|
|
1178
|
+
`bun run scripts/vertex-package-smoke.ts`. The latter installs the unreleased
|
|
1179
|
+
local package cohort into an isolated consumer and uses synthetic responses;
|
|
1180
|
+
release dependency resolution and live model access are separate checks.
|
|
1181
|
+
|
|
1182
|
+
Repository: <https://github.com/Zhivex/zhivex-ai-sdk>
|