@maci0/dsh-google-vertex 0.0.0-stage → 0.12.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +166 -2
- package/cordis.patch.yml +34 -0
- package/icon.svg +6 -0
- package/lib/adapter.js +398 -0
- package/lib/auth.js +246 -0
- package/lib/client.js +206 -0
- package/lib/discovery.js +180 -0
- package/lib/gemini.js +561 -0
- package/lib/gemini_adapter.js +61 -0
- package/lib/host.js +18 -0
- package/lib/index.js +196 -0
- package/lib/types/adapter.d.ts +219 -0
- package/lib/types/auth.d.ts +111 -0
- package/lib/types/discovery.d.ts +57 -0
- package/lib/types/gemini.d.ts +210 -0
- package/lib/types/gemini_adapter.d.ts +54 -0
- package/lib/types/host.d.ts +202 -0
- package/lib/types/index.d.ts +161 -0
- package/lib/types/wire-shared.d.ts +19 -0
- package/lib/types/wire.d.ts +271 -0
- package/lib/wire-shared.js +59 -0
- package/lib/wire.js +722 -0
- package/locale/en.json +6 -0
- package/locale/zh.json +6 -0
- package/package.json +93 -4
package/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
package/README.md
CHANGED
|
@@ -1,3 +1,167 @@
|
|
|
1
|
-
#
|
|
1
|
+
# dsh-google-vertex
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
One Google service-account JSON file unlocks eleven models inside DeepSeek Harness: Claude Opus 4.6, Sonnet 4.6, Opus 4.5, Sonnet 4.5, Haiku 4.5, plus Gemini 3.5 Flash, 3.1 Pro Preview, 3 Flash Preview, 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite.
|
|
4
|
+
Install it when your model capacity already lives in a Google Cloud project and you would rather point Harness at a credential file than paste an API key.
|
|
5
|
+
Nothing is copied into the harness credential store: the file is read once at mount, and each request trades a signed RS256 assertion for a cached OAuth token.
|
|
6
|
+
|
|
7
|
+
## What you get
|
|
8
|
+
|
|
9
|
+
- Two provider routes from one configuration row (`google-vertex-anthropic` for Claude and `google-vertex-gemini` for Gemini), both in the Web model picker.
|
|
10
|
+
- Service-account auth with no API key: a path in config, or `GOOGLE_APPLICATION_CREDENTIALS`.
|
|
11
|
+
- Streaming with tool calling on both routes. Claude emits raw JSON argument deltas; Gemini emits complete function calls and replays the thought signature Gemini 3 requires.
|
|
12
|
+
- The Gemini catalog discovered from Vertex at runtime behind a five-minute cache, with the built-in list as the fallback when the provider cannot be reached. The Claude catalog is served from the configured (or built-in) list.
|
|
13
|
+
- A **Refresh models** control on this plugin's row page, which drops both caches and re-reads the model picker without restarting `dsh web`.
|
|
14
|
+
- Claude prompt-cache breakpoints on tools, system, and the final conversation block, with cache read/write counts reported as usage.
|
|
15
|
+
- Failure codes the harness can act on: `429` → `RATE_LIMIT`, `5xx` → `SERVER`, `RESOURCE_EXHAUSTED` → `QUOTA`, oversized → `CONTEXT_WINDOW_EXCEEDED`, a stalled stream → `TIMEOUT`.
|
|
16
|
+
- Environment fallbacks for the credential, project, and region.
|
|
17
|
+
|
|
18
|
+
## Install
|
|
19
|
+
|
|
20
|
+
> **Install it as a bundle.** `dsh plugin add …` mounts the row from the
|
|
21
|
+
> package's own patch layer, which is what the settings editor can write to. A
|
|
22
|
+
> row added with `--patch` is an overlay: it disappears at the next start, and
|
|
23
|
+
> the Plugins card cannot save into it (the editor refuses a write an overlay
|
|
24
|
+
> would win).
|
|
25
|
+
|
|
26
|
+
```sh
|
|
27
|
+
dsh plugin --profile web add @maci0/dsh-google-vertex@0.12.4
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
This installs the public npm package; no GitHub token or `~/.secrets` setup is needed.
|
|
31
|
+
The version is pinned. To upgrade, run the same command with a newer version,
|
|
32
|
+
then restart `dsh web` (bundle layers compose at boot).
|
|
33
|
+
|
|
34
|
+
The package declares `dsh.bundle`, so `dsh plugin` appends it to `dsh.profile.bundles` and the package's own `cordis.patch.yml` inserts the `google-vertex` row: there is no row to paste. Adding that row by hand while the package is a bundle mounts the plugin twice, because `insert` does not dedupe ids.
|
|
35
|
+
|
|
36
|
+
## Configure
|
|
37
|
+
|
|
38
|
+
Override the row by id in `~/.dsh/profiles/web/cordis.patch.yml`:
|
|
39
|
+
|
|
40
|
+
```yaml
|
|
41
|
+
- id: google-vertex
|
|
42
|
+
config:
|
|
43
|
+
serviceAccountFile: ~/.secrets/global-claude-project.json
|
|
44
|
+
project: global-claude-project
|
|
45
|
+
location: global
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
`~` is expanded, and the file is read once at mount, so a typo fails loudly instead of mid-turn. A patch replaces the targeted row's whole `config`, so restate every key you keep. This file is live-watched: saving it remounts the plugin, no restart.
|
|
49
|
+
|
|
50
|
+
| Key | Default | Meaning |
|
|
51
|
+
|---|---|---|
|
|
52
|
+
| `serviceAccountFile` | `$GOOGLE_APPLICATION_CREDENTIALS` | Path to the service-account JSON; `~` is expanded. Nothing is copied into the harness credential store: the file is the credential, and it is read once at mount. |
|
|
53
|
+
| `project` | `$GOOGLE_CLOUD_PROJECT`, `$GCLOUD_PROJECT`, then the file's `project_id` | Google Cloud project id, used by both routes. |
|
|
54
|
+
| `location` | `$GOOGLE_CLOUD_LOCATION`, then `global` | Region for both endpoints, or `global` for `aiplatform.googleapis.com`. |
|
|
55
|
+
| `models` | the Claude catalog below | Claude ids to advertise, replacing the built-in catalog. Any id is accepted at request time regardless. |
|
|
56
|
+
| `geminiModels` | the Gemini catalog below | Gemini ids to advertise, served as written: a configured list turns live discovery off. Every entry serves the same context/output pair. |
|
|
57
|
+
| `contextWindow` | `200000` | Claude context window reported per model. |
|
|
58
|
+
| `maxTokens` | `32000` | Claude output cap applied when a caller omits one. |
|
|
59
|
+
| `streamIdleTimeoutMs` | `300000` | Bound on the interval between two stream reads, on both routes. |
|
|
60
|
+
|
|
61
|
+
Environment variables are read from `process.env` of the harness process.
|
|
62
|
+
|
|
63
|
+
## Routes and models
|
|
64
|
+
|
|
65
|
+
Vertex serves the two publishers over different paths, which is why one row registers two routes. The service-account path, project, and region are written once.
|
|
66
|
+
|
|
67
|
+
| Route | Provider path | Serves |
|
|
68
|
+
|---|---|---|
|
|
69
|
+
| `google-vertex-anthropic` | `publishers/anthropic/models/{model}:streamRawPredict` | Claude, picker group **Google Vertex AI (Anthropic)** |
|
|
70
|
+
| `google-vertex-gemini` | `publishers/google/models/{model}:streamGenerateContent?alt=sse` | Gemini, picker group **Google Vertex AI (Gemini)** |
|
|
71
|
+
|
|
72
|
+
Claude defaults. Every listed family serves 200k context and caps output at 32k unless a caller asks otherwise:
|
|
73
|
+
|
|
74
|
+
| Picker label | Model id |
|
|
75
|
+
|---|---|
|
|
76
|
+
| Claude Opus 4.6 (Vertex) | `claude-opus-4-6` |
|
|
77
|
+
| Claude Sonnet 4.6 (Vertex) | `claude-sonnet-4-6` |
|
|
78
|
+
| Claude Opus 4.5 (Vertex) | `claude-opus-4-5` |
|
|
79
|
+
| Claude Sonnet 4.5 (Vertex) | `claude-sonnet-4-5` |
|
|
80
|
+
| Claude Haiku 4.5 (Vertex) | `claude-haiku-4-5` |
|
|
81
|
+
|
|
82
|
+
Gemini defaults. Every entry serves the same context/output pair:
|
|
83
|
+
|
|
84
|
+
| Picker label | Model id | Context | Max output |
|
|
85
|
+
|---|---|---|---|
|
|
86
|
+
| Gemini 3.5 Flash (Vertex) | `gemini-3.5-flash` | 1,048,576 | 65,535 |
|
|
87
|
+
| Gemini 3.1 Pro Preview (Vertex) | `gemini-3.1-pro-preview` | 1,048,576 | 65,535 |
|
|
88
|
+
| Gemini 3 Flash Preview (Vertex) | `gemini-3-flash-preview` | 1,048,576 | 65,535 |
|
|
89
|
+
| Gemini 2.5 Pro (Vertex) | `gemini-2.5-pro` | 1,048,576 | 65,535 |
|
|
90
|
+
| Gemini 2.5 Flash (Vertex) | `gemini-2.5-flash` | 1,048,576 | 65,535 |
|
|
91
|
+
| Gemini 2.5 Flash-Lite (Vertex) | `gemini-2.5-flash-lite` | 1,048,576 | 65,535 |
|
|
92
|
+
|
|
93
|
+
The Gemini cap is one below Vertex's own ceiling on purpose: `maxOutputTokens: 65536` is refused with "supported range is from 1 (inclusive) to 65536 (exclusive)".
|
|
94
|
+
|
|
95
|
+
Ids are Vertex's aliases, so a promoted release needs no edit here. Both catalogs are overridable from configuration for a deployment that pins dated versions.
|
|
96
|
+
|
|
97
|
+
### Discovery and refresh
|
|
98
|
+
|
|
99
|
+
The lists above are the fallback, not the whole catalog. Gemini models are read from Vertex's `publishers/google/models` list (a `v1beta1` method: v1 has no list), keeping `gemini-` ids except the embedding, TTS, image, Live, and audio variants this text route cannot drive. A built-in id keeps its picker label; any other is labelled by its id. The result is cached for five minutes so opening the model picker is not a network call. The Claude list is served exactly as configured, because Vertex exposes no Anthropic model listing endpoint.
|
|
100
|
+
|
|
101
|
+
The picker re-reads that catalog when the host says a model input changed, so a model Google published a minute ago stays invisible until then. This plugin's row page carries the control that forces it: open **Plugins** in the sidebar, open the `dsh-google-vertex` bundle, and configure the `google-vertex` row. **Refresh models** drops the cached Gemini catalog and makes the picker re-read it from this process. The card records the last manual refresh in its own `google-vertex` settings namespace, and the write is also the signal: a browser half has no other channel to the host.
|
|
102
|
+
|
|
103
|
+
## Try it
|
|
104
|
+
|
|
105
|
+
1. Restart `dsh web`, then open a session.
|
|
106
|
+
2. Type `/model` in the composer, or click the model seat beside it.
|
|
107
|
+
3. Pick **Google Vertex AI (Anthropic)** → **Claude Sonnet 4.6 (Vertex)**, or the Gemini group for a Gemini model. Selecting a model also makes it the default for new sessions; a session that already sent a request keeps the model recorded in its own log.
|
|
108
|
+
4. Type a prompt that needs a tool:
|
|
109
|
+
|
|
110
|
+
```
|
|
111
|
+
List the files in this directory and tell me which one is largest.
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
The agent runs a shell tool and answers from its output. The same works on the Gemini group, including the signed replay of the model's function call.
|
|
115
|
+
|
|
116
|
+
## How it works
|
|
117
|
+
|
|
118
|
+
Failed HTTP requests forward valid `Retry-After` seconds or HTTP dates to the harness retry policy. Invalid, non-positive, and past delays are omitted.
|
|
119
|
+
|
|
120
|
+
- **Auth.** The service-account JSON is read once at mount. Per request, a JWT is signed with its private key, sent to the file's `token_uri` (or Google's public token endpoint when the file omits one), and traded for an access token (`https://www.googleapis.com/auth/cloud-platform`), which is cached and refreshed five minutes before expiry. The request carries it as a bearer token. A credential problem reports `AUTH`; a token-endpoint transport problem reports `TRANSPORT`.
|
|
121
|
+
- **Streaming.** Both routes bound every read by `streamIdleTimeoutMs`, including HTTP error bodies. The watchdog owns its own controller, so a stalled read is torn down and the turn ends with a single `TIMEOUT` failure instead of hanging. Caller cancellation ends as an `aborted` finish, not a provider error.
|
|
122
|
+
- **Failure classification.** A refused HTTP response and an in-band provider error envelope classify the same way: a named cause first (`RESOURCE_EXHAUSTED` → `QUOTA`, oversized prompt → `CONTEXT_WINDOW_EXCEEDED`), then the numeric code (`429` → `RATE_LIMIT`, `5xx` → `SERVER`, `404` → `NOT_FOUND`). Gemini has no terminal event, so a body that ends without a finish reason is a truncated response unless the watchdog or an in-band error already ended the turn.
|
|
123
|
+
- **Replay.** Gemini 3 signs its function calls, and a replay that drops the signature is refused with `400 INVALID_ARGUMENT: Function call is missing a thought_signature in functionCall parts`. The adapter stores each emitted block's `thoughtSignature` in the harness replay envelope and echoes it on the next request. Signatures are per model, so a cross-model replay sends the call unsigned rather than failing. No `thinkingConfig` is ever sent: Gemini's dynamic-thinking default is what keeps `gemini-2.5-pro` working. Thinking tokens still arrive in `usageMetadata` and count as output.
|
|
124
|
+
- **Attribution.** Requests carry `attributionHeaders()` from `@deepseek-ai/dsh-llm` rather than a pinned release line, so `User-Agent` cannot drift from the installed harness.
|
|
125
|
+
|
|
126
|
+
## Limits
|
|
127
|
+
|
|
128
|
+
- **Text only.** Both routes declare `inputModalities: ['text']`, so the harness projects images and files to placeholder text before dispatch. Claude and Gemini on Vertex both accept images; wiring the attachment service into these adapters is the work that would add it.
|
|
129
|
+
- **No extended thinking on Claude.** That route never sends a `thinking` block, and replaying one needs a thinking signature the harness reasoning block cannot hold. Enabling thinking without it makes every tool round trip fail.
|
|
130
|
+
- **Gemini thought summaries are dropped** for the same signature reason, so only answer text is shown.
|
|
131
|
+
- **Vertex AI API must be enabled** on the project, and the service account needs `roles/aiplatform.user`.
|
|
132
|
+
- **Claude is not servable in every region.** `global`, `us-east5`, and `europe-west1` answered; `us-central1` returned `400 FAILED_PRECONDITION: … is not servable in region us-central1`. A `404 NOT_FOUND` naming the publisher model means the id does not exist on Vertex; `429 RESOURCE_EXHAUSTED` is the project's quota, not a bad request.
|
|
133
|
+
- **Claude prompt caching has provider-side minimums.** Blocks under the cacheable minimum are accepted and simply not cached.
|
|
134
|
+
|
|
135
|
+
### The pi-ai `google-vertex` route
|
|
136
|
+
|
|
137
|
+
Harness already ships a pi-ai-backed `google-vertex` route that serves Gemini from a credential record carrying the same three environment values. It accepts images and surfaces thinking summaries; this plugin's Gemini route is text-only and drops them. Choose it when those matter. Its record in `~/.dsh/.credentials.yaml` (owner-only, mode 600, live-watched):
|
|
138
|
+
|
|
139
|
+
```yaml
|
|
140
|
+
records:
|
|
141
|
+
llm-pi-ai/google-vertex:
|
|
142
|
+
kind: api-key
|
|
143
|
+
env:
|
|
144
|
+
GOOGLE_APPLICATION_CREDENTIALS: /path/to/service-account.json
|
|
145
|
+
GOOGLE_CLOUD_PROJECT: global-claude-project
|
|
146
|
+
GOOGLE_CLOUD_LOCATION: global
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
Two of its catalog entries disagree with Vertex's own limits, so those models need an override in `~/.dsh/settings.yaml` under `llm-pi-ai.providers.google-vertex.modelOverrides`: `gemini-2.5-pro` needs `reasoningEfforts: false` (Vertex rejects `thinkingBudget: 0` with `400 INVALID_ARGUMENT`), and `gemini-2.5-flash-lite` needs `maxTokens: 65535`. This plugin's own Gemini route needs neither.
|
|
150
|
+
|
|
151
|
+
## Development
|
|
152
|
+
|
|
153
|
+
```sh
|
|
154
|
+
bun test # hermetic, stubbed transport, no network
|
|
155
|
+
bun run build # tsc -p tsconfig.build.json → lib/index.js + lib/types/
|
|
156
|
+
bun run typecheck # tsc -p tsconfig.json
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
The package ships the built `lib/` and declares `dsh.bundle`, so a change to `src/` needs `bun run build` before it takes effect. For local development, run `bun run build`, then `dsh plugin --profile <name> add <path-to-checkout>` and restart `dsh web`. `lib/client.js` is the exception: the browser half is authored directly as plain JavaScript in the client module loader's factory format and is not produced by `tsc`.
|
|
160
|
+
|
|
161
|
+
Coverage: the signed assertion, token caching and refresh, both auth failure classes, request projection, cache breakpoints, Gemini request projection with signature replay, SSE framing across split chunks, every terminal finish class (including the stream idle bound and an in-band provider error envelope), usage reported only when the provider reported it, and configuration validation. A real Cordis `Context` mount proves both routes are registered and withdrawn with the fiber, and that the plugin's settings change drops both cached catalogs. The browser half is imported from `lib/client.js` under a stub of the module loader's own registration format, which is how its slot, its Refresh control, and its write are covered without a browser.
|
|
162
|
+
|
|
163
|
+
dsh loads plugins on Node `^22.19.0 || >=24.0.0`; development and tests run on bun.
|
|
164
|
+
|
|
165
|
+
## Licence
|
|
166
|
+
|
|
167
|
+
MIT. See [LICENSE](LICENSE).
|
package/cordis.patch.yml
ADDED
|
@@ -0,0 +1,34 @@
|
|
|
1
|
+
# The dsh-google-vertex bundle patch: applied automatically when a profile
|
|
2
|
+
# lists this bundle. Users override this row from their profile's own
|
|
3
|
+
# cordis.patch.yml (live-watched; dsh.profile.bundles is frozen at boot) with a
|
|
4
|
+
# `- id: google-vertex` row, which replaces the row's whole `config`.
|
|
5
|
+
#
|
|
6
|
+
# One row registers two routes:
|
|
7
|
+
# google-vertex-anthropic: Google-hosted Claude (publishers/anthropic)
|
|
8
|
+
# google-vertex-gemini: Gemini (publishers/google)
|
|
9
|
+
- insert:
|
|
10
|
+
- id: google-vertex
|
|
11
|
+
name: '@maci0/dsh-google-vertex'
|
|
12
|
+
config:
|
|
13
|
+
# Service-account JSON path ('~' allowed). Left unset here so the
|
|
14
|
+
# $GOOGLE_APPLICATION_CREDENTIALS fallback applies; a deployment that
|
|
15
|
+
# keeps the file elsewhere overrides this row with the path. Mount
|
|
16
|
+
# fails loudly when neither is set.
|
|
17
|
+
# serviceAccountFile: ~/.secrets/your-service-account.json
|
|
18
|
+
# Google Cloud project. Optional: falls back to $GOOGLE_CLOUD_PROJECT,
|
|
19
|
+
# $GCLOUD_PROJECT, then the credentials file's own project_id.
|
|
20
|
+
# project: my-project
|
|
21
|
+
# Region for both Vertex endpoints, or 'global' (default) for
|
|
22
|
+
# aiplatform.googleapis.com. Optional: falls back to
|
|
23
|
+
# $GOOGLE_CLOUD_LOCATION. Claude is not servable in every region.
|
|
24
|
+
# location: us-east5
|
|
25
|
+
# Claude model ids, replacing the built-in catalog.
|
|
26
|
+
# models:
|
|
27
|
+
# - claude-sonnet-4-5
|
|
28
|
+
# Gemini model ids, replacing the built-in catalog and live discovery.
|
|
29
|
+
# geminiModels:
|
|
30
|
+
# - gemini-3.5-flash
|
|
31
|
+
# - gemini-2.5-pro
|
|
32
|
+
# Claude context window and output cap reported for every model.
|
|
33
|
+
# contextWindow: 200000
|
|
34
|
+
# maxTokens: 32000
|
package/icon.svg
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
<svg width="36" height="36" viewBox="0 0 36 36" fill="none" xmlns="http://www.w3.org/2000/svg">
|
|
2
|
+
<path d="M16.4 11.2L11.2 23.4M19.6 11.2L24.8 23.4M12.4 26.2h11.2" stroke="#3D5AFE" stroke-width="1.8" stroke-linecap="round"/>
|
|
3
|
+
<circle cx="18" cy="8.2" r="3.2" fill="#5B6CFF"/>
|
|
4
|
+
<circle cx="9" cy="26.6" r="3.2" fill="#45D9E7"/>
|
|
5
|
+
<circle cx="27" cy="26.6" r="3.2" fill="#7CB7FF"/>
|
|
6
|
+
</svg>
|
package/lib/adapter.js
ADDED
|
@@ -0,0 +1,398 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* The machinery both Vertex publisher adapters share, plus the Claude route's
|
|
3
|
+
* own adapter: one metadata face the harness calls on every dispatch, and one
|
|
4
|
+
* streaming pipeline carrying the per-read idle watchdog, the token mint, the
|
|
5
|
+
* fetch, the refusal classification, the SSE loop, and the single-terminal-chunk
|
|
6
|
+
* guarantee.
|
|
7
|
+
*
|
|
8
|
+
* Two things separate the Claude route from the pi-ai-backed `google-vertex`
|
|
9
|
+
* route the harness already ships. Its credential is a service-account file,
|
|
10
|
+
* turned into a bearer token per request by {@link ServiceAccountTokens} rather
|
|
11
|
+
* than typed in as an API key. And its models are Claude, which Vertex serves
|
|
12
|
+
* through `publishers/anthropic`, a path and body the Gemini protocol cannot
|
|
13
|
+
* express.
|
|
14
|
+
*
|
|
15
|
+
* Both routes are text-only: `inputModalities: ['text']` makes `LlmRuntime`
|
|
16
|
+
* project images and files to placeholder text before dispatch, which is honest
|
|
17
|
+
* about what these adapters send.
|
|
18
|
+
*
|
|
19
|
+
* The two routes differ in exactly three things: how a request is addressed
|
|
20
|
+
* and built, what the provider's payloads mean, and what a body that ends
|
|
21
|
+
* without a terminal event should be called. Each is an injected callback;
|
|
22
|
+
* everything else lives here once.
|
|
23
|
+
*
|
|
24
|
+
* @module dsh-google-vertex/adapter
|
|
25
|
+
*/
|
|
26
|
+
import { attributionHeaders } from '@deepseek-ai/dsh-llm';
|
|
27
|
+
import { ServiceAccountTokens, VertexAuthError } from './auth.js';
|
|
28
|
+
import { ModelCache } from './discovery.js';
|
|
29
|
+
import { buildRequestBody, endpointFor, failureForStatus, parseSseRecord, SseRecordReader, StreamTranslator, } from './wire.js';
|
|
30
|
+
/** Terminal error finish carrying one failure. */
|
|
31
|
+
function errorFinish(failure) {
|
|
32
|
+
return { type: 'finish', reason: { kind: 'error', failure } };
|
|
33
|
+
}
|
|
34
|
+
/** Terminal aborted finish carrying one failure. */
|
|
35
|
+
function abortedFinish(failure) {
|
|
36
|
+
return { type: 'finish', reason: { kind: 'aborted', failure } };
|
|
37
|
+
}
|
|
38
|
+
/**
|
|
39
|
+
* The failure for a stream that produced no bytes within the configured bound.
|
|
40
|
+
* @param idleTimeoutMs - the configured per-read bound.
|
|
41
|
+
* @returns the terminal failure, coded `TIMEOUT`.
|
|
42
|
+
*/
|
|
43
|
+
function idleTimeoutFailure(idleTimeoutMs) {
|
|
44
|
+
return {
|
|
45
|
+
message: `google-vertex: no stream data for ${idleTimeoutMs}ms (streamIdleTimeoutMs)`,
|
|
46
|
+
code: 'TIMEOUT',
|
|
47
|
+
};
|
|
48
|
+
}
|
|
49
|
+
/** Classify a credential failure raised before the request was sent. */
|
|
50
|
+
function credentialFailure(error) {
|
|
51
|
+
if (error instanceof VertexAuthError) {
|
|
52
|
+
return { message: error.message, code: error.code };
|
|
53
|
+
}
|
|
54
|
+
return {
|
|
55
|
+
message: `google-vertex: ${error instanceof Error ? error.message : String(error)}`,
|
|
56
|
+
code: 'AUTH',
|
|
57
|
+
};
|
|
58
|
+
}
|
|
59
|
+
/** Classify a fetch or body-read failure, distinguishing cancellation. */
|
|
60
|
+
function transportFinish(options, error) {
|
|
61
|
+
if (options.signal?.aborted === true) {
|
|
62
|
+
return abortedFinish({ message: 'google-vertex: request aborted', code: 'ABORTED' });
|
|
63
|
+
}
|
|
64
|
+
return errorFinish({
|
|
65
|
+
message: `google-vertex: ${error instanceof Error ? error.message : String(error)}`,
|
|
66
|
+
code: 'TRANSPORT',
|
|
67
|
+
});
|
|
68
|
+
}
|
|
69
|
+
/**
|
|
70
|
+
* One streaming completion, shared by both publisher routes.
|
|
71
|
+
*
|
|
72
|
+
* The bearer token is minted before the request, so a credentials failure is
|
|
73
|
+
* reported as a terminal finish rather than as a thrown error escaping the
|
|
74
|
+
* generator; `LlmRuntime` would turn a throw into a finish too, but would code
|
|
75
|
+
* it `UNKNOWN` and lose the `AUTH` versus `TRANSPORT` classification.
|
|
76
|
+
*
|
|
77
|
+
* Every read is bounded by `streamIdleTimeoutMs`: a provider that stops sending
|
|
78
|
+
* is a terminal `TIMEOUT` rather than a turn that never ends. The watchdog owns
|
|
79
|
+
* its own controller so the stalled read can be torn down; the caller's signal
|
|
80
|
+
* is combined with it when present. The token mint is armed and cleared with
|
|
81
|
+
* the same bound, because it is a network call too.
|
|
82
|
+
*
|
|
83
|
+
* Exactly one terminal chunk leaves this stream. A body that ended without the
|
|
84
|
+
* provider's terminal event is a truncated response, unless the watchdog or the
|
|
85
|
+
* caller ended the turn, which is the more specific account of the same missing
|
|
86
|
+
* event.
|
|
87
|
+
* @param config - the resolved configuration this adapter serves.
|
|
88
|
+
* @param options - the harness request.
|
|
89
|
+
* @param fetch - transport, already defaulted by the adapter.
|
|
90
|
+
* @param tokens - token source, already defaulted by the adapter.
|
|
91
|
+
* @param makeTranslator - builds this route's payload translator for one model.
|
|
92
|
+
* @param pump - this route's endpoint, body, and terminal wording.
|
|
93
|
+
* @yields every chunk the provider's stream completes.
|
|
94
|
+
*/
|
|
95
|
+
export async function* streamVertex(config, options, fetch, tokens, makeTranslator, pump) {
|
|
96
|
+
const model = options.model.length > 0 ? options.model : config.models[0]?.id ?? '';
|
|
97
|
+
const consumer = new AbortController();
|
|
98
|
+
const signal = options.signal === undefined
|
|
99
|
+
? consumer.signal
|
|
100
|
+
: AbortSignal.any([options.signal, consumer.signal]);
|
|
101
|
+
let idleTimedOut = false;
|
|
102
|
+
let idleTimer;
|
|
103
|
+
/** True while an operation of the provider's is outstanding. */
|
|
104
|
+
let idleReading = false;
|
|
105
|
+
let idleStartedAt = 0;
|
|
106
|
+
/**
|
|
107
|
+
* One timer wakes when the outstanding read would have outlived the bound;
|
|
108
|
+
* while the provider keeps answering it re-arms itself for the remainder, so
|
|
109
|
+
* a stream costs one timer instead of one `setTimeout`/`clearTimeout` pair per
|
|
110
|
+
* transport read.
|
|
111
|
+
*/
|
|
112
|
+
const checkIdle = () => {
|
|
113
|
+
idleTimer = undefined;
|
|
114
|
+
// Nothing outstanding: leave the timer unarmed; the next arm creates one.
|
|
115
|
+
// This also keeps a stream a consumer abandoned collectable.
|
|
116
|
+
if (!idleReading)
|
|
117
|
+
return;
|
|
118
|
+
const elapsed = performance.now() - idleStartedAt;
|
|
119
|
+
if (elapsed >= config.streamIdleTimeoutMs) {
|
|
120
|
+
idleTimedOut = true;
|
|
121
|
+
consumer.abort('google-vertex: stream idle timeout');
|
|
122
|
+
return;
|
|
123
|
+
}
|
|
124
|
+
// The wake-up preceded this read's own deadline: sleep the remainder.
|
|
125
|
+
idleTimer = setTimeout(checkIdle, config.streamIdleTimeoutMs - elapsed);
|
|
126
|
+
};
|
|
127
|
+
const armIdle = () => {
|
|
128
|
+
idleReading = true;
|
|
129
|
+
idleStartedAt = performance.now();
|
|
130
|
+
idleTimer ??= setTimeout(checkIdle, config.streamIdleTimeoutMs);
|
|
131
|
+
};
|
|
132
|
+
/** A read resolved: nothing of the provider's is outstanding any more. */
|
|
133
|
+
const clearIdle = () => {
|
|
134
|
+
idleReading = false;
|
|
135
|
+
};
|
|
136
|
+
/** No operation will be armed again: leave no timer holding the event loop. */
|
|
137
|
+
const disarmIdle = () => {
|
|
138
|
+
idleReading = false;
|
|
139
|
+
if (idleTimer !== undefined)
|
|
140
|
+
clearTimeout(idleTimer);
|
|
141
|
+
idleTimer = undefined;
|
|
142
|
+
};
|
|
143
|
+
let token;
|
|
144
|
+
armIdle();
|
|
145
|
+
try {
|
|
146
|
+
// The token mint is a network call too: an unbounded one stalls the same
|
|
147
|
+
// way a stalled body does.
|
|
148
|
+
token = await tokens.get(signal);
|
|
149
|
+
}
|
|
150
|
+
catch (error) {
|
|
151
|
+
if (idleTimedOut) {
|
|
152
|
+
yield errorFinish(idleTimeoutFailure(config.streamIdleTimeoutMs));
|
|
153
|
+
return;
|
|
154
|
+
}
|
|
155
|
+
if (options.signal?.aborted === true) {
|
|
156
|
+
yield transportFinish(options, error);
|
|
157
|
+
return;
|
|
158
|
+
}
|
|
159
|
+
yield errorFinish(credentialFailure(error));
|
|
160
|
+
return;
|
|
161
|
+
}
|
|
162
|
+
finally {
|
|
163
|
+
disarmIdle();
|
|
164
|
+
}
|
|
165
|
+
let response;
|
|
166
|
+
armIdle();
|
|
167
|
+
try {
|
|
168
|
+
response = await fetch(pump.endpoint(model, config), {
|
|
169
|
+
method: 'POST',
|
|
170
|
+
headers: {
|
|
171
|
+
'content-type': 'application/json',
|
|
172
|
+
'accept': 'text/event-stream',
|
|
173
|
+
...attributionHeaders(),
|
|
174
|
+
'authorization': `Bearer ${token}`,
|
|
175
|
+
},
|
|
176
|
+
body: JSON.stringify(pump.body(options, config)),
|
|
177
|
+
signal,
|
|
178
|
+
});
|
|
179
|
+
}
|
|
180
|
+
catch (error) {
|
|
181
|
+
if (idleTimedOut) {
|
|
182
|
+
yield errorFinish(idleTimeoutFailure(config.streamIdleTimeoutMs));
|
|
183
|
+
return;
|
|
184
|
+
}
|
|
185
|
+
yield transportFinish(options, error);
|
|
186
|
+
return;
|
|
187
|
+
}
|
|
188
|
+
finally {
|
|
189
|
+
disarmIdle();
|
|
190
|
+
}
|
|
191
|
+
if (!response.ok) {
|
|
192
|
+
let body;
|
|
193
|
+
armIdle();
|
|
194
|
+
try {
|
|
195
|
+
body = await response.text().catch(() => '');
|
|
196
|
+
}
|
|
197
|
+
finally {
|
|
198
|
+
disarmIdle();
|
|
199
|
+
}
|
|
200
|
+
yield idleTimedOut ? errorFinish(idleTimeoutFailure(config.streamIdleTimeoutMs))
|
|
201
|
+
: options.signal?.aborted === true ? transportFinish(options, options.signal.reason)
|
|
202
|
+
: errorFinish(failureForStatus(response.status, body, `model "${model}" in ${config.location}`, response.headers));
|
|
203
|
+
return;
|
|
204
|
+
}
|
|
205
|
+
if (response.body === null) {
|
|
206
|
+
yield errorFinish({ message: 'google-vertex: response carried no body', code: 'TRANSPORT' });
|
|
207
|
+
return;
|
|
208
|
+
}
|
|
209
|
+
const translator = makeTranslator(model);
|
|
210
|
+
try {
|
|
211
|
+
// The reader is driven directly rather than through an async generator:
|
|
212
|
+
// one record costs one loop turn instead of a suspended generator frame,
|
|
213
|
+
// a promise, and a microtask.
|
|
214
|
+
const reader = new SseRecordReader(response.body, { armIdle, clearIdle });
|
|
215
|
+
try {
|
|
216
|
+
for (let records = await reader.read(); records !== undefined; records = await reader.read()) {
|
|
217
|
+
for (const record of records) {
|
|
218
|
+
const event = parseSseRecord(record);
|
|
219
|
+
if (event === undefined)
|
|
220
|
+
continue;
|
|
221
|
+
yield* translator.handle(event);
|
|
222
|
+
// The provider ended the turn mid-body (`message_stop`, or an
|
|
223
|
+
// in-band error envelope). Its finish has been emitted, so neither the
|
|
224
|
+
// end of the body nor a close fault may add a second one.
|
|
225
|
+
if (translator.terminal)
|
|
226
|
+
return;
|
|
227
|
+
}
|
|
228
|
+
}
|
|
229
|
+
}
|
|
230
|
+
finally {
|
|
231
|
+
await reader.close();
|
|
232
|
+
}
|
|
233
|
+
}
|
|
234
|
+
catch (error) {
|
|
235
|
+
// The provider already ended the turn, so a fault while closing the body
|
|
236
|
+
// cannot add a second terminal chunk.
|
|
237
|
+
if (translator.terminal)
|
|
238
|
+
return;
|
|
239
|
+
if (idleTimedOut) {
|
|
240
|
+
yield errorFinish(idleTimeoutFailure(config.streamIdleTimeoutMs));
|
|
241
|
+
return;
|
|
242
|
+
}
|
|
243
|
+
yield transportFinish(options, error);
|
|
244
|
+
return;
|
|
245
|
+
}
|
|
246
|
+
finally {
|
|
247
|
+
disarmIdle();
|
|
248
|
+
}
|
|
249
|
+
if (idleTimedOut) {
|
|
250
|
+
yield errorFinish(idleTimeoutFailure(config.streamIdleTimeoutMs));
|
|
251
|
+
return;
|
|
252
|
+
}
|
|
253
|
+
if (options.signal?.aborted === true) {
|
|
254
|
+
yield transportFinish(options, options.signal?.reason);
|
|
255
|
+
return;
|
|
256
|
+
}
|
|
257
|
+
// The body ended without the provider's own finish: a truncated response,
|
|
258
|
+
// which is the more specific account of a turn the caller or the watchdog
|
|
259
|
+
// already ended.
|
|
260
|
+
if (!(translator.sawFinish ?? translator.terminal)) {
|
|
261
|
+
yield errorFinish({ message: pump.truncatedMessage(model), code: 'TRANSPORT' });
|
|
262
|
+
return;
|
|
263
|
+
}
|
|
264
|
+
yield* translator.finish?.() ?? [];
|
|
265
|
+
}
|
|
266
|
+
/**
|
|
267
|
+
* Duck-typed base for both Vertex publisher adapters.
|
|
268
|
+
*
|
|
269
|
+
* `LlmRuntime` reaches adapters through plain method calls, so these objects
|
|
270
|
+
* need no harness base class; the plugin's only runtime `@deepseek-ai/*`
|
|
271
|
+
* dependency is `@deepseek-ai/dsh-llm`'s pure `attributionHeaders()` helper.
|
|
272
|
+
* The metadata face below is identical for both routes, so it is written once
|
|
273
|
+
* and parameterized by {@link AdapterMetadata}.
|
|
274
|
+
*/
|
|
275
|
+
export class VertexPublisherAdapter {
|
|
276
|
+
#metadata;
|
|
277
|
+
#modelCache;
|
|
278
|
+
#discover;
|
|
279
|
+
/** The config this adapter serves, for the stream pipeline below. */
|
|
280
|
+
config;
|
|
281
|
+
/** The transport this adapter was built with. */
|
|
282
|
+
fetch;
|
|
283
|
+
/** The token source this adapter asks before each request. */
|
|
284
|
+
tokens;
|
|
285
|
+
/**
|
|
286
|
+
* @param metadata - display name, capacities, and description text.
|
|
287
|
+
* @param config - the resolved configuration this adapter serves.
|
|
288
|
+
* @param options - transport, token-source, and discovery overrides.
|
|
289
|
+
*/
|
|
290
|
+
constructor(metadata, config, options = {}) {
|
|
291
|
+
this.#metadata = metadata;
|
|
292
|
+
this.config = config;
|
|
293
|
+
this.fetch = options.fetch ?? ((input, init) => globalThis.fetch(input, init));
|
|
294
|
+
this.tokens = options.tokens ?? new ServiceAccountTokens(config.serviceAccount, { fetch: this.fetch });
|
|
295
|
+
this.#discover = options.discover;
|
|
296
|
+
this.#modelCache = new ModelCache(config.models);
|
|
297
|
+
}
|
|
298
|
+
/** {@inheritDoc LlmAdapterLike.providerInfo} */
|
|
299
|
+
providerInfo(provider) {
|
|
300
|
+
return { id: provider, name: this.#metadata.providerName };
|
|
301
|
+
}
|
|
302
|
+
/**
|
|
303
|
+
* No provider-owned retry policy: Vertex's own 429s carry a quota code the
|
|
304
|
+
* harness already classifies, and a per-request backoff belongs to the
|
|
305
|
+
* harness defaults.
|
|
306
|
+
*
|
|
307
|
+
* This and {@link imageRequestPricing} exist because `LlmRuntime` calls them
|
|
308
|
+
* on every dispatch; a duck-typed adapter must supply them or the first
|
|
309
|
+
* registration throws.
|
|
310
|
+
*/
|
|
311
|
+
providerRetryPolicy(_provider) {
|
|
312
|
+
return undefined;
|
|
313
|
+
}
|
|
314
|
+
/** No route charges visual tokens: these adapters are text-only. */
|
|
315
|
+
imageRequestPricing(_provider, _model) {
|
|
316
|
+
return undefined;
|
|
317
|
+
}
|
|
318
|
+
/**
|
|
319
|
+
* The current model catalog, fetched from the provider when discovery is
|
|
320
|
+
* configured, falling back to the static catalog on failure.
|
|
321
|
+
*
|
|
322
|
+
* The result is cached with a five-minute TTL so the model picker does
|
|
323
|
+
* not make a network call on every open.
|
|
324
|
+
*/
|
|
325
|
+
async listModels(provider) {
|
|
326
|
+
const models = await this.#modelCache.get(this.#discover);
|
|
327
|
+
return models.map(model => this.#info(provider, model.id, model.name));
|
|
328
|
+
}
|
|
329
|
+
/**
|
|
330
|
+
* Drop the cached catalog, so the next `listModels` re-discovers from the
|
|
331
|
+
* provider. Both adapters answer the plugin's manual refresh with this.
|
|
332
|
+
*/
|
|
333
|
+
invalidateModels() {
|
|
334
|
+
this.#modelCache.invalidate();
|
|
335
|
+
}
|
|
336
|
+
/** {@inheritDoc LlmAdapterLike.resolveModel} */
|
|
337
|
+
resolveModel(provider, model, _signal) {
|
|
338
|
+
const capacity = this.#metadata.capacity;
|
|
339
|
+
return Promise.resolve({
|
|
340
|
+
...this.#info(provider, model),
|
|
341
|
+
context: { contextWindow: capacity.contextWindow },
|
|
342
|
+
defaultMaxTokens: capacity.defaultMaxTokens,
|
|
343
|
+
});
|
|
344
|
+
}
|
|
345
|
+
/** {@inheritDoc LlmAdapterLike.prepareCall} */
|
|
346
|
+
async prepareCall(provider, model, signal) {
|
|
347
|
+
return {
|
|
348
|
+
model: await this.resolveModel(provider, model, signal),
|
|
349
|
+
stream: (options) => this.stream(options),
|
|
350
|
+
};
|
|
351
|
+
}
|
|
352
|
+
/** Display metadata for one model id, named from the catalog or the given name. */
|
|
353
|
+
#info(provider, model, overrideName) {
|
|
354
|
+
const known = this.config.models.find(entry => entry.id === model);
|
|
355
|
+
const name = overrideName ?? known?.name ?? model;
|
|
356
|
+
return {
|
|
357
|
+
provider,
|
|
358
|
+
id: model,
|
|
359
|
+
name,
|
|
360
|
+
description: this.#metadata.describe(this.config),
|
|
361
|
+
inputModalities: ['text'],
|
|
362
|
+
};
|
|
363
|
+
}
|
|
364
|
+
}
|
|
365
|
+
/**
|
|
366
|
+
* Duck-typed adapter over Vertex's Anthropic publisher endpoint.
|
|
367
|
+
*
|
|
368
|
+
* The route is text-only and declares no server-executed tools, and every
|
|
369
|
+
* Claude family in the default catalog serves the same 200k context, so a
|
|
370
|
+
* model's capacities are the configured pair rather than a per-model entry.
|
|
371
|
+
*/
|
|
372
|
+
export class GoogleVertexAnthropicAdapter extends VertexPublisherAdapter {
|
|
373
|
+
/**
|
|
374
|
+
* @param config - the resolved configuration this adapter serves.
|
|
375
|
+
* @param options - transport, token-source, and discovery overrides for tests.
|
|
376
|
+
*/
|
|
377
|
+
constructor(config, options = {}) {
|
|
378
|
+
super({
|
|
379
|
+
providerName: 'Google Vertex AI (Anthropic)',
|
|
380
|
+
capacity: { contextWindow: config.contextWindow, defaultMaxTokens: config.maxTokens },
|
|
381
|
+
describe: row => `Google-hosted Anthropic model on Vertex AI (project ${row.project}, ${row.location}).`,
|
|
382
|
+
}, config, options);
|
|
383
|
+
}
|
|
384
|
+
/**
|
|
385
|
+
* Stream one completion through `:streamRawPredict`.
|
|
386
|
+
*
|
|
387
|
+
* The shared pump owns the watchdog, the token mint, the SSE loop, and the
|
|
388
|
+
* single terminal chunk; `message_stop` closes the turn mid-body, so this
|
|
389
|
+
* route's finish rides that event.
|
|
390
|
+
*/
|
|
391
|
+
stream(options) {
|
|
392
|
+
return streamVertex(this.config, options, this.fetch, this.tokens, () => new StreamTranslator(), {
|
|
393
|
+
endpoint: (model, config) => endpointFor(config.project, config.location, model),
|
|
394
|
+
body: (request, config) => buildRequestBody(request, config),
|
|
395
|
+
truncatedMessage: model => `google-vertex: model "${model}" stream ended before message_stop`,
|
|
396
|
+
});
|
|
397
|
+
}
|
|
398
|
+
}
|