zuplo 7.8.15 → 7.8.16
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/ai-gateway/bedrock-runtime.mdx +18 -10
- package/docs/ai-gateway/custom-providers.mdx +13 -2
- package/docs/ai-gateway/integrations/ai-sdk.mdx +42 -5
- package/docs/ai-gateway/managing-providers.mdx +3 -1
- package/docs/ai-gateway/openrouter.mdx +265 -0
- package/docs/ai-gateway/overview.mdx +5 -4
- package/docs/ai-gateway/providers.mdx +31 -2
- package/docs/ai-gateway/universal-api.mdx +6 -6
- package/package.json +5 -5
|
@@ -202,8 +202,12 @@ bedrockruntime/us.anthropic.claude-haiku-4-5-20251001-v1:0
|
|
|
202
202
|
```
|
|
203
203
|
|
|
204
204
|
The gateway strips its own prefix and sends AWS the model ID exactly as you
|
|
205
|
-
wrote it
|
|
206
|
-
|
|
205
|
+
wrote it. As with every other provider, the model must be one you enabled for
|
|
206
|
+
the provider. Cross-region inference profile IDs appear in the model list as
|
|
207
|
+
models of their own. A foundation-model or inference-profile ARN counts as the
|
|
208
|
+
model ID it contains, so it works when that model is enabled. The gateway
|
|
209
|
+
refuses any other model before it calls AWS, including a provisioned-throughput
|
|
210
|
+
ARN and an application inference profile, which don't name a model in the list.
|
|
207
211
|
|
|
208
212
|
:::tip{title="Many models need an inference profile ID"}
|
|
209
213
|
|
|
@@ -231,12 +235,9 @@ priced too: the `us.`, `eu.`, `apac.`, `au.`, `jp.`, `us-gov.`, and `global.`
|
|
|
231
235
|
IDs are catalog rows in their own right, each at its own rate. A
|
|
232
236
|
foundation-model or inference-profile ARN, such as
|
|
233
237
|
`arn:aws:bedrock:eu-central-1:123456789012:inference-profile/eu.anthropic.claude-opus-4-8`,
|
|
234
|
-
is priced as the ID it contains.
|
|
235
|
-
|
|
236
|
-
|
|
237
|
-
token and request [budgets](./usage-limits.mdx), but is priced at zero. It
|
|
238
|
-
doesn't count toward spending budgets, and its response carries no `X-Cost-USD`
|
|
239
|
-
header.
|
|
238
|
+
is priced as the ID it contains. Every model you can call has a catalog price,
|
|
239
|
+
so each request counts toward token, request, and spending
|
|
240
|
+
[budgets](./usage-limits.mdx).
|
|
240
241
|
|
|
241
242
|
When a budget is exhausted, these routes answer
|
|
242
243
|
`400 ServiceQuotaExceededException`—not the `429` the Universal API returns. AWS
|
|
@@ -296,6 +297,12 @@ baked into the deployed gateway. The dialog also rejects an access key ID that
|
|
|
296
297
|
isn't 16 to 128 uppercase letters and digits, a secret containing spaces or line
|
|
297
298
|
breaks, and either field filled in without the other.
|
|
298
299
|
|
|
300
|
+
**An error saying the model `is not included in model selections`.** The model
|
|
301
|
+
isn't enabled for the provider. Open the provider in
|
|
302
|
+
[**Settings → AI Providers**](https://portal.zuplo.com/+/account/project/ai/settings/data-models),
|
|
303
|
+
enable the model, and save. If the model is missing from the list, the gateway
|
|
304
|
+
can't serve it through this provider yet.
|
|
305
|
+
|
|
299
306
|
**A `400` naming the supported operations.** You called `CountTokens` or
|
|
300
307
|
`InvokeModelWithBidirectionalStream`, which the gateway doesn't serve.
|
|
301
308
|
|
|
@@ -305,8 +312,9 @@ these routes with that exception—check the app's usage limits before you open
|
|
|
305
312
|
quota ticket with AWS.
|
|
306
313
|
|
|
307
314
|
**Your app's model list is empty of Bedrock models.** Expected—Bedrock Runtime
|
|
308
|
-
models don't appear in `GET /v1/models`, because
|
|
309
|
-
(
|
|
315
|
+
models don't appear in `GET /v1/models`, because that listing describes the
|
|
316
|
+
[Universal API](./universal-api.mdx), and this provider serves only Bedrock's
|
|
317
|
+
native paths.
|
|
310
318
|
|
|
311
319
|
## Next steps
|
|
312
320
|
|
|
@@ -57,8 +57,10 @@ To add a custom AI provider to your Zuplo AI Gateway, follow these steps:
|
|
|
57
57
|
provider named `acme-llm` serves models as `acme-llm/<model>`. Names are
|
|
58
58
|
lowercase (letters, numbers, dots, dashes, and underscores), and Zuplo
|
|
59
59
|
reserves the built-in provider names (`openai`, `anthropic`, `google`,
|
|
60
|
-
`mistral`, `xai`, `moonshot`, `zuplo`, `zuplodemo`, `zuplo-demo
|
|
61
|
-
|
|
60
|
+
`mistral`, `xai`, `moonshot`, `zuplo`, `zuplodemo`, `zuplo-demo`,
|
|
61
|
+
`bedrockmantle`, `bedrock-mantle`, `azureai`, `azure-ai`, `vertexai`,
|
|
62
|
+
`vertex-ai`, `openrouter`, `open-router`). The name is permanent after
|
|
63
|
+
creation.
|
|
62
64
|
|
|
63
65
|
1. Enter the provider's **API URL** as its origin root—for example
|
|
64
66
|
`https://api.together.xyz`. The gateway appends `/v1/chat/completions` (and
|
|
@@ -80,6 +82,15 @@ To add a custom AI provider to your Zuplo AI Gateway, follow these steps:
|
|
|
80
82
|
| Groq | `https://api.groq.com/openai/v1` | `https://api.groq.com/openai` |
|
|
81
83
|
| Fireworks | `https://api.fireworks.ai/inference/v1` | `https://api.fireworks.ai/inference` |
|
|
82
84
|
|
|
85
|
+
:::tip{title="OpenRouter is built in"}
|
|
86
|
+
|
|
87
|
+
OpenRouter is now a built-in provider—pick it from the provider list instead
|
|
88
|
+
of configuring it here, and you get model discovery and per-model cost
|
|
89
|
+
tracking. See [Using OpenRouter](./openrouter.mdx). The OpenRouter row above
|
|
90
|
+
still applies to any OpenRouter provider you added before it became built-in.
|
|
91
|
+
|
|
92
|
+
:::
|
|
93
|
+
|
|
83
94
|
1. Enter the API Key for the selected provider (if there is no API key required,
|
|
84
95
|
you can leave this blank).
|
|
85
96
|
|
|
@@ -53,9 +53,10 @@ configure.
|
|
|
53
53
|
Each provider package appends its own operation path to `baseURL`, so pick the
|
|
54
54
|
model factory that targets an endpoint the gateway serves:
|
|
55
55
|
`/v1/chat/completions` for OpenAI-compatible requests, `/v1/messages` for
|
|
56
|
-
Anthropic, `/v1/embeddings` for every
|
|
57
|
-
`/v1/responses` for OpenAI and
|
|
58
|
-
|
|
56
|
+
Anthropic and other providers' Claude models, `/v1/embeddings` for every
|
|
57
|
+
provider with embedding models, and `/v1/responses` for OpenAI and the other
|
|
58
|
+
providers whose models serve it (see [AI Providers](../providers.mdx) for the
|
|
59
|
+
matrix). The examples below use the right factory for each provider.
|
|
59
60
|
|
|
60
61
|
### OpenAI
|
|
61
62
|
|
|
@@ -143,8 +144,10 @@ const { text } = await generateText({
|
|
|
143
144
|
### xAI
|
|
144
145
|
|
|
145
146
|
Call `xai.chat(...)` rather than `xai(...)`. The bare callable targets xAI's
|
|
146
|
-
Responses API, and
|
|
147
|
-
|
|
147
|
+
Responses API, and xAI's models do not serve `/v1/responses` through the
|
|
148
|
+
gateway—an xAI model sent there returns a `400`. The endpoint is
|
|
149
|
+
capability-gated per model rather than restricted to one provider, so which
|
|
150
|
+
models reach it is the matrix in [AI Providers](../providers.mdx).
|
|
148
151
|
|
|
149
152
|
```typescript
|
|
150
153
|
import { createXai } from "@ai-sdk/xai";
|
|
@@ -287,6 +290,35 @@ const { text } = await generateText({
|
|
|
287
290
|
});
|
|
288
291
|
```
|
|
289
292
|
|
|
293
|
+
### OpenRouter
|
|
294
|
+
|
|
295
|
+
Use `@ai-sdk/openai`. OpenRouter model references have two slashes,
|
|
296
|
+
`providerName/vendor/model`—see
|
|
297
|
+
[Model references include the vendor prefix](../openrouter.mdx#model-references-include-the-vendor-prefix).
|
|
298
|
+
|
|
299
|
+
Which factory works depends on the model. The bare `openai(id)` callable targets
|
|
300
|
+
the Responses API, which OpenRouter's Claude models don't serve through the
|
|
301
|
+
gateway, so a Claude model sent that way returns a `400`. Call `openai.chat(id)`
|
|
302
|
+
for Claude—the gateway translates the chat completion to the Messages API—or use
|
|
303
|
+
`@ai-sdk/anthropic` with `authToken` for the native Messages API. Every other
|
|
304
|
+
chat model works with either form.
|
|
305
|
+
|
|
306
|
+
```typescript
|
|
307
|
+
import { createOpenAI } from "@ai-sdk/openai";
|
|
308
|
+
import { generateText } from "ai";
|
|
309
|
+
|
|
310
|
+
const openai = createOpenAI({
|
|
311
|
+
apiKey: process.env.ZUPLO_AI_GATEWAY_API_KEY,
|
|
312
|
+
baseURL:
|
|
313
|
+
"https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e/v1",
|
|
314
|
+
});
|
|
315
|
+
|
|
316
|
+
const { text } = await generateText({
|
|
317
|
+
model: openai.chat("openrouter/anthropic/claude-sonnet-4.5"),
|
|
318
|
+
prompt: "Write a one-sentence bedtime story about a unicorn.",
|
|
319
|
+
});
|
|
320
|
+
```
|
|
321
|
+
|
|
290
322
|
## Provider options the gateway doesn't forward
|
|
291
323
|
|
|
292
324
|
On the OpenAI-shaped endpoints (`/v1/chat/completions`, `/v1/embeddings`, and
|
|
@@ -298,6 +330,11 @@ forwarded. Options outside that list are dropped without a warning. For example,
|
|
|
298
330
|
`seed` and `providerOptions.mistral.safePrompt` don't reach Mistral, and
|
|
299
331
|
`temperature` is capped at Mistral's maximum of `1`.
|
|
300
332
|
|
|
333
|
+
OpenRouter's model-fallback options are the exception: the gateway rejects
|
|
334
|
+
`models` and `route` on chat completions, and `fallbacks` on `/v1/messages`,
|
|
335
|
+
with a `400` rather than dropping them. See
|
|
336
|
+
[Using OpenRouter](../openrouter.mdx#provider-routing-preferences).
|
|
337
|
+
|
|
301
338
|
Anthropic is different: `/v1/messages` is a native passthrough, so
|
|
302
339
|
`@ai-sdk/anthropic` requests reach Anthropic unchanged apart from the model and
|
|
303
340
|
credential, which the gateway sets from your app configuration.
|
|
@@ -58,7 +58,9 @@ To add a new AI provider to your Zuplo AI Gateway, follow these steps:
|
|
|
58
58
|
Project ID** and takes a service account JSON key file instead of an API key,
|
|
59
59
|
and [Azure AI](./azure-ai.mdx) asks for an **Azure Resource Name** and a host
|
|
60
60
|
family, plus a mapping for any deployment named differently from the model it
|
|
61
|
-
serves.
|
|
61
|
+
serves. [OpenRouter](./openrouter.mdx) is the opposite case: it asks for
|
|
62
|
+
nothing beyond the key, being a single global host with no endpoint, region
|
|
63
|
+
or project to enter.
|
|
62
64
|
|
|
63
65
|
1. Select the model or models you want to use with this provider. The available
|
|
64
66
|
models will depend on the selected provider. This can be changed later.
|
|
@@ -0,0 +1,265 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Using OpenRouter
|
|
3
|
+
sidebar_label: OpenRouter
|
|
4
|
+
description:
|
|
5
|
+
Reach hundreds of models from every major vendor through one OpenRouter API
|
|
6
|
+
key. Model references carry the vendor prefix, Claude models serve the native
|
|
7
|
+
Messages API, and reported spend matches your OpenRouter invoice.
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
Reach hundreds of models from every major vendor through one OpenRouter API key,
|
|
11
|
+
with gateway spend that matches your OpenRouter invoice.
|
|
12
|
+
|
|
13
|
+
**OpenRouter** is a marketplace rather than a single vendor: one key reaches
|
|
14
|
+
models from OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen, xAI and many more,
|
|
15
|
+
and OpenRouter chooses an upstream host for each request. Adding it as a
|
|
16
|
+
provider serves those models to your [apps](./apps.mdx) through the
|
|
17
|
+
[Universal API](./universal-api.mdx), on your OpenRouter account and your
|
|
18
|
+
OpenRouter billing.
|
|
19
|
+
|
|
20
|
+
Three things set this provider apart, and each one changes something you'll
|
|
21
|
+
type:
|
|
22
|
+
|
|
23
|
+
- **Model references carry the vendor prefix**, so they have two slashes.
|
|
24
|
+
- **Claude models serve the Messages API**, not the Responses API.
|
|
25
|
+
- **Reported spend matches your OpenRouter invoice**, because OpenRouter prices
|
|
26
|
+
each request and the gateway bills from that figure.
|
|
27
|
+
|
|
28
|
+
## Model references include the vendor prefix
|
|
29
|
+
|
|
30
|
+
Every OpenRouter model ID is vendor-qualified — `anthropic/claude-sonnet-4.5`,
|
|
31
|
+
`openai/gpt-5`, `meta-llama/llama-3.3-70b-instruct`. A gateway model reference
|
|
32
|
+
prefixes your provider name onto that, so it has **two** slashes:
|
|
33
|
+
|
|
34
|
+
```text
|
|
35
|
+
openrouter/anthropic/claude-sonnet-4.5
|
|
36
|
+
openrouter/openai/gpt-5
|
|
37
|
+
openrouter/meta-llama/llama-3.3-70b-instruct
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
The first segment is your provider name, and everything after the first slash is
|
|
41
|
+
the model ID exactly as OpenRouter publishes it. This is the same shape
|
|
42
|
+
[Vertex AI](./vertex-ai.mdx) uses.
|
|
43
|
+
|
|
44
|
+
## Supported endpoints
|
|
45
|
+
|
|
46
|
+
Which endpoints a model serves depends on what kind of model it is. Claude is
|
|
47
|
+
served on a different API from every other chat model, and embedding models
|
|
48
|
+
serve only embeddings:
|
|
49
|
+
|
|
50
|
+
| Endpoint | Claude models (`anthropic/…`) | Other chat models | Embedding models |
|
|
51
|
+
| --------------------------- | ----------------------------- | ----------------- | ---------------- |
|
|
52
|
+
| `/v1/chat/completions` | ✅ (translated) | ✅ | ❌ |
|
|
53
|
+
| `/v1/messages` | ✅ | ❌ | ❌ |
|
|
54
|
+
| `/v1/messages/count_tokens` | ❌ | ❌ | ❌ |
|
|
55
|
+
| `/v1/responses` | ❌ | ✅ | ❌ |
|
|
56
|
+
| `/v1/embeddings` | ❌ | ❌ | ✅ |
|
|
57
|
+
|
|
58
|
+
:::note{title="The Responses API is stateless on OpenRouter"}
|
|
59
|
+
|
|
60
|
+
OpenRouter doesn't store responses. You can create one, but you can't retrieve
|
|
61
|
+
or delete it or list its input items, and OpenRouter rejects requests that set
|
|
62
|
+
`store: true` or `previous_response_id`. Send the full conversation with each
|
|
63
|
+
request instead.
|
|
64
|
+
|
|
65
|
+
OpenRouter doesn't offer token counting either, so `/v1/messages/count_tokens`
|
|
66
|
+
returns a `400` for every OpenRouter model.
|
|
67
|
+
|
|
68
|
+
:::
|
|
69
|
+
|
|
70
|
+
:::note{title="Claude models don't serve the Responses API here"}
|
|
71
|
+
|
|
72
|
+
OpenRouter's own Responses endpoint will accept a Claude model. The gateway
|
|
73
|
+
doesn't allow it, so that every provider follows the same rule: `/v1/messages`
|
|
74
|
+
means Claude, and `/v1/responses` means everything else. To call a Claude model
|
|
75
|
+
from an OpenAI-compatible client, use `/v1/chat/completions`: the gateway
|
|
76
|
+
translates the request onto the Messages API for you.
|
|
77
|
+
|
|
78
|
+
:::
|
|
79
|
+
|
|
80
|
+
## Before you begin
|
|
81
|
+
|
|
82
|
+
You need an OpenRouter API key. Create one at
|
|
83
|
+
[openrouter.ai/keys](https://openrouter.ai/keys); it starts with `sk-or-v1-`.
|
|
84
|
+
|
|
85
|
+
You also need credit on the OpenRouter account the key belongs to, or requests
|
|
86
|
+
come back with OpenRouter's own insufficient-credit error.
|
|
87
|
+
|
|
88
|
+
## Add the provider
|
|
89
|
+
|
|
90
|
+
Adding or editing providers requires the **Edit** permission, granted to Zuplo
|
|
91
|
+
account and project **Admins**—see
|
|
92
|
+
[Managing Providers](./managing-providers.mdx).
|
|
93
|
+
|
|
94
|
+
<Stepper>
|
|
95
|
+
|
|
96
|
+
1. Open
|
|
97
|
+
[**Settings → AI Providers**](https://portal.zuplo.com/+/account/project/ai/settings/data-models)
|
|
98
|
+
in your AI Gateway project in the Zuplo Portal.
|
|
99
|
+
|
|
100
|
+
2. Select **Add Provider**, then choose **OpenRouter**.
|
|
101
|
+
|
|
102
|
+
3. Give the provider a **Provider Name**. This is the first segment of every
|
|
103
|
+
model reference, so a provider named `openrouter` serves
|
|
104
|
+
`openrouter/openai/gpt-5`. The name is permanent after creation.
|
|
105
|
+
|
|
106
|
+
4. Paste your **API Key**. There is no endpoint or region to enter — OpenRouter
|
|
107
|
+
is a single global host.
|
|
108
|
+
|
|
109
|
+
5. Select the models this provider should serve, then **Save**.
|
|
110
|
+
|
|
111
|
+
</Stepper>
|
|
112
|
+
|
|
113
|
+
## Routing shortcuts
|
|
114
|
+
|
|
115
|
+
OpenRouter lets you append a routing preference to any model ID:
|
|
116
|
+
|
|
117
|
+
| Suffix | Effect |
|
|
118
|
+
| --------- | -------------------------------------------- |
|
|
119
|
+
| `:nitro` | Prefer the fastest upstream host |
|
|
120
|
+
| `:floor` | Prefer the cheapest upstream host |
|
|
121
|
+
| `:online` | Add web search to the request |
|
|
122
|
+
| `:exacto` | Prefer hosts that call tools most accurately |
|
|
123
|
+
|
|
124
|
+
These work through the gateway and are forwarded to OpenRouter unchanged. They
|
|
125
|
+
don't appear in the model picker, because they aren't separate models — the
|
|
126
|
+
gateway resolves `openrouter/openai/gpt-5:nitro` against the `openai/gpt-5`
|
|
127
|
+
catalog entry and sends the suffix upstream.
|
|
128
|
+
|
|
129
|
+
Two consequences worth knowing:
|
|
130
|
+
|
|
131
|
+
- **A model-filtering allow list matches literally.** Allowing
|
|
132
|
+
`openrouter/openai/gpt-5` does not allow `openrouter/openai/gpt-5:nitro`. Add
|
|
133
|
+
the suffixed reference explicitly if you want it.
|
|
134
|
+
- **Other suffixes aren't supported.** OpenRouter's `:free` variants are
|
|
135
|
+
separate models, and the gateway's OpenRouter catalog doesn't include them, so
|
|
136
|
+
a request for one is rejected like any model outside your selections.
|
|
137
|
+
|
|
138
|
+
## Call the models
|
|
139
|
+
|
|
140
|
+
Point any OpenAI-compatible client at your app URL with `/v1` appended:
|
|
141
|
+
|
|
142
|
+
```ts
|
|
143
|
+
import OpenAI from "openai";
|
|
144
|
+
|
|
145
|
+
const client = new OpenAI({
|
|
146
|
+
baseURL:
|
|
147
|
+
"https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e/v1",
|
|
148
|
+
apiKey: process.env.ZUPLO_APP_API_KEY,
|
|
149
|
+
});
|
|
150
|
+
|
|
151
|
+
const completion = await client.chat.completions.create({
|
|
152
|
+
model: "openrouter/openai/gpt-5",
|
|
153
|
+
messages: [{ role: "user", content: "Hello!" }],
|
|
154
|
+
});
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
### Provider routing preferences
|
|
158
|
+
|
|
159
|
+
OpenRouter's `provider` option is forwarded, so you can express upstream
|
|
160
|
+
preferences per request:
|
|
161
|
+
|
|
162
|
+
```ts
|
|
163
|
+
const completion = await client.chat.completions.create({
|
|
164
|
+
model: "openrouter/openai/gpt-5",
|
|
165
|
+
messages: [{ role: "user", content: "Hello!" }],
|
|
166
|
+
// @ts-expect-error -- an OpenRouter extension, not part of the OpenAI types
|
|
167
|
+
provider: { sort: "throughput" },
|
|
168
|
+
});
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
`provider` is forwarded on embeddings and the Responses API too. On chat
|
|
172
|
+
completions, `transforms`, `reasoning` and `plugins` are forwarded the same way.
|
|
173
|
+
The Responses API forwards `plugins`, and `reasoning` is a standard Responses
|
|
174
|
+
parameter.
|
|
175
|
+
|
|
176
|
+
:::caution{title="The models fallback array is never forwarded"}
|
|
177
|
+
|
|
178
|
+
OpenRouter's `models` array (and `route: "fallback"`) asks OpenRouter to
|
|
179
|
+
substitute a different model when the first one is unavailable. On the Messages
|
|
180
|
+
API the same array is called `fallbacks`. On chat completions and the Messages
|
|
181
|
+
API, the gateway **rejects** these parameters with a `400` rather than dropping
|
|
182
|
+
them silently. On embeddings and the Responses API, it drops them before the
|
|
183
|
+
request reaches OpenRouter.
|
|
184
|
+
|
|
185
|
+
They would route around the gateway itself: your model filtering wouldn't see
|
|
186
|
+
the substitute, your budgets and per-model costs would be attributed to a model
|
|
187
|
+
that never ran, and it would collide with the gateway's own fallback handling.
|
|
188
|
+
Configure a backup model on the route instead.
|
|
189
|
+
|
|
190
|
+
:::
|
|
191
|
+
|
|
192
|
+
## Call Claude models on the Messages API
|
|
193
|
+
|
|
194
|
+
Claude models on OpenRouter serve Anthropic's native Messages API, so the
|
|
195
|
+
Anthropic SDK works against them directly. Set `baseURL` **without** `/v1` — the
|
|
196
|
+
Anthropic SDK appends that itself — and pass your key as `authToken`:
|
|
197
|
+
|
|
198
|
+
```ts
|
|
199
|
+
import Anthropic from "@anthropic-ai/sdk";
|
|
200
|
+
|
|
201
|
+
const client = new Anthropic({
|
|
202
|
+
baseURL:
|
|
203
|
+
"https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e",
|
|
204
|
+
authToken: process.env.ZUPLO_APP_API_KEY,
|
|
205
|
+
});
|
|
206
|
+
|
|
207
|
+
const message = await client.messages.create({
|
|
208
|
+
model: "openrouter/anthropic/claude-sonnet-4.5",
|
|
209
|
+
max_tokens: 1024,
|
|
210
|
+
messages: [{ role: "user", content: "Hello!" }],
|
|
211
|
+
});
|
|
212
|
+
```
|
|
213
|
+
|
|
214
|
+
`authToken`, not `apiKey`: the SDK's `apiKey` sends an `x-api-key` header, which
|
|
215
|
+
the gateway doesn't read.
|
|
216
|
+
|
|
217
|
+
For the Vercel AI SDK, see
|
|
218
|
+
[the OpenRouter section of the AI SDK guide](./integrations/ai-sdk.mdx#openrouter).
|
|
219
|
+
|
|
220
|
+
## Costs
|
|
221
|
+
|
|
222
|
+
OpenRouter prices every request itself and reports the figure in the response,
|
|
223
|
+
and the gateway records **that** figure rather than computing one from a stored
|
|
224
|
+
rate card. Your reported spend therefore reconciles with your OpenRouter
|
|
225
|
+
invoice.
|
|
226
|
+
|
|
227
|
+
This matters more here than for a single-vendor provider. A `:floor` request and
|
|
228
|
+
a `:nitro` request for the same model can run on different upstream hosts at
|
|
229
|
+
different prices, so there is no single per-model rate that would be correct for
|
|
230
|
+
both.
|
|
231
|
+
|
|
232
|
+
Responses carry an `X-Cost-USD` header alongside `X-Cost-Source: provider`,
|
|
233
|
+
which tells you the figure came from OpenRouter rather than from a rate card.
|
|
234
|
+
|
|
235
|
+
## Troubleshooting
|
|
236
|
+
|
|
237
|
+
**`400` saying `/v1/responses` is not supported** — you sent a Claude model to
|
|
238
|
+
`/v1/responses`. Use `/v1/messages`, or `/v1/chat/completions`, which the
|
|
239
|
+
gateway translates onto the Messages API; see
|
|
240
|
+
[the table above](#supported-endpoints).
|
|
241
|
+
|
|
242
|
+
**`400` naming `/v1/messages`** — you sent a model other than Claude to
|
|
243
|
+
`/v1/messages`, which serves only Claude models here.
|
|
244
|
+
|
|
245
|
+
**An error saying the model is not included in your model selections** — the
|
|
246
|
+
model isn't selected on this provider. Open the provider on the **AI Providers**
|
|
247
|
+
page and enable it.
|
|
248
|
+
|
|
249
|
+
**`400` rejecting `models`, `route` or `fallbacks`** — see the caution above;
|
|
250
|
+
use a route backup model instead.
|
|
251
|
+
|
|
252
|
+
**`400` from `/v1/messages/count_tokens`** — OpenRouter doesn't offer token
|
|
253
|
+
counting.
|
|
254
|
+
|
|
255
|
+
**`404` from OpenRouter** — usually a model ID typo. The ID after your provider
|
|
256
|
+
name must match OpenRouter's exactly, vendor prefix and all.
|
|
257
|
+
|
|
258
|
+
**A key that won't save** — OpenRouter keys start with `sk-or-v1-` and contain
|
|
259
|
+
no whitespace. Re-copy it from [openrouter.ai/keys](https://openrouter.ai/keys).
|
|
260
|
+
|
|
261
|
+
## Next steps
|
|
262
|
+
|
|
263
|
+
- [Universal API](./universal-api.mdx) — the endpoints every provider shares
|
|
264
|
+
- [Managing providers](./managing-providers.mdx) — editing models and keys
|
|
265
|
+
- [AI Providers](./providers.mdx) — the full provider and capability matrix
|
|
@@ -78,10 +78,11 @@ Configure multiple LLM providers within a single AI Gateway project. Supported
|
|
|
78
78
|
providers include OpenAI, Anthropic, Google, Mistral, xAI, Microsoft Azure
|
|
79
79
|
(through [Azure AI](./azure-ai.mdx)), Amazon Bedrock (through
|
|
80
80
|
[Bedrock Mantle](./bedrock-mantle.mdx)), Google Cloud
|
|
81
|
-
([Vertex AI](./vertex-ai.mdx)),
|
|
82
|
-
[AI Providers](./providers.mdx) for the
|
|
83
|
-
capabilities. Apps reference models as
|
|
84
|
-
`openai/gpt-5-mini`—so a single app can use
|
|
81
|
+
([Vertex AI](./vertex-ai.mdx)), [OpenRouter](./openrouter.mdx), and
|
|
82
|
+
OpenAI-compatible custom providers. See [AI Providers](./providers.mdx) for the
|
|
83
|
+
full list of providers and supported capabilities. Apps reference models as
|
|
84
|
+
`providerName/model`—for example `openai/gpt-5-mini`—so a single app can use
|
|
85
|
+
models from several providers.
|
|
85
86
|
|
|
86
87
|
### Source-Controlled Gateway
|
|
87
88
|
|
|
@@ -3,8 +3,8 @@ title: AI Providers
|
|
|
3
3
|
sidebar_label: Overview
|
|
4
4
|
description:
|
|
5
5
|
The AI providers and capabilities the Zuplo AI Gateway supports, including
|
|
6
|
-
OpenAI, Anthropic, Google, Vertex AI, Mistral, xAI, and
|
|
7
|
-
custom providers.
|
|
6
|
+
OpenAI, Anthropic, Google, Vertex AI, Mistral, xAI, OpenRouter, and
|
|
7
|
+
OpenAI-compatible custom providers.
|
|
8
8
|
---
|
|
9
9
|
|
|
10
10
|
Zuplo's AI Gateway supports integration with various AI providers, allowing you
|
|
@@ -27,6 +27,8 @@ Zuplo currently supports the following AI providers:
|
|
|
27
27
|
that already calls Bedrock with an AWS SDK
|
|
28
28
|
- [Vertex AI](./vertex-ai.mdx)—Google Cloud's managed model platform, serving
|
|
29
29
|
Gemini and Model Garden models on your own Google Cloud project
|
|
30
|
+
- [OpenRouter](./openrouter.mdx)—hundreds of models from every major vendor
|
|
31
|
+
behind a single API key
|
|
30
32
|
- [Zuplo Demo](#zuplo-demo)—a free, keyless provider for trying the gateway
|
|
31
33
|
- OpenAI-compatible [Custom Providers](./custom-providers.mdx) (such as Qwen,
|
|
32
34
|
Kimi, etc)
|
|
@@ -44,6 +46,7 @@ The following capabilities are supported across providers:
|
|
|
44
46
|
| Bedrock Mantle | ✅ | ❌ | ✅ | ✅ |
|
|
45
47
|
| Bedrock Runtime | ❌ | ❌ | ❌ | ❌ |
|
|
46
48
|
| Vertex AI | ✅ | ✅ | ❌ | ✅ |
|
|
49
|
+
| OpenRouter | ✅ | ✅ | ✅ | ✅ |
|
|
47
50
|
| Zuplo Demo | ✅ | ❌ | ❌ | ❌ |
|
|
48
51
|
| OpenAI-compatible (Custom) | ✅ | ✅ | ❌ | ❌ |
|
|
49
52
|
|
|
@@ -179,6 +182,32 @@ Model Garden for your Google Cloud project before it answers.
|
|
|
179
182
|
For prerequisites, Google Cloud setup, provider steps, code examples, and
|
|
180
183
|
troubleshooting, see [Using Vertex AI](./vertex-ai.mdx).
|
|
181
184
|
|
|
185
|
+
## OpenRouter
|
|
186
|
+
|
|
187
|
+
**OpenRouter** is a marketplace, not a single vendor: one API key reaches
|
|
188
|
+
hundreds of models from OpenAI, Anthropic, Google, Meta, DeepSeek, Qwen, xAI and
|
|
189
|
+
many more, and OpenRouter picks an upstream host for each request.
|
|
190
|
+
|
|
191
|
+
Three things are specific to it:
|
|
192
|
+
|
|
193
|
+
- **Model references have two slashes.** Every OpenRouter model ID is
|
|
194
|
+
vendor-qualified, so a provider named `openrouter` serves
|
|
195
|
+
`openrouter/anthropic/claude-sonnet-4.5` and `openrouter/openai/gpt-5`. This
|
|
196
|
+
is the same shape Vertex AI uses.
|
|
197
|
+
- **Claude models are on the Messages API.** Models whose ID begins `anthropic/`
|
|
198
|
+
serve the native Anthropic Messages API (`/v1/messages`) and a translated
|
|
199
|
+
`/v1/chat/completions`, but not `/v1/responses`, even though OpenRouter's own
|
|
200
|
+
Responses endpoint would accept them. Other chat models serve chat completions
|
|
201
|
+
and the Responses API, and embedding models serve embeddings.
|
|
202
|
+
- **Reported spend matches your OpenRouter invoice.** OpenRouter prices each
|
|
203
|
+
request itself and reports the figure inline, and the gateway bills from that
|
|
204
|
+
rather than from a stored rate card. That matters because OpenRouter's routing
|
|
205
|
+
shortcuts can send the same model to different hosts at different prices from
|
|
206
|
+
one request to the next.
|
|
207
|
+
|
|
208
|
+
For prerequisites, provider steps, code examples, the routing shortcuts, and
|
|
209
|
+
troubleshooting, see [Using OpenRouter](./openrouter.mdx).
|
|
210
|
+
|
|
182
211
|
## Zuplo Demo
|
|
183
212
|
|
|
184
213
|
**Zuplo Demo** is a free provider that Zuplo operates so you can try the AI
|
|
@@ -63,12 +63,12 @@ list—so clients that can't set a model still work.
|
|
|
63
63
|
|
|
64
64
|
## Supported endpoints
|
|
65
65
|
|
|
66
|
-
| Endpoint | Notes
|
|
67
|
-
| ---------------------- |
|
|
68
|
-
| `/v1/chat/completions` | Chat completions for every provider
|
|
69
|
-
| `/v1/embeddings` | Embeddings for the providers marked in the [capability matrix](./providers.mdx#supported-providers)
|
|
70
|
-
| `/v1/responses` | OpenAI Responses API: OpenAI, and the OpenAI-compatible models that serve it on [Azure AI](./azure-ai.mdx)
|
|
71
|
-
| `/v1/messages` | Anthropic Messages API: Anthropic, and the Claude models of [Azure AI](./azure-ai.mdx), [Bedrock Mantle](./bedrock-mantle.mdx)
|
|
66
|
+
| Endpoint | Notes |
|
|
67
|
+
| ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
68
|
+
| `/v1/chat/completions` | Chat completions for every provider |
|
|
69
|
+
| `/v1/embeddings` | Embeddings for the providers marked in the [capability matrix](./providers.mdx#supported-providers) |
|
|
70
|
+
| `/v1/responses` | OpenAI Responses API: OpenAI, and the OpenAI-compatible models that serve it on [Azure AI](./azure-ai.mdx), [Bedrock Mantle](./bedrock-mantle.mdx) and [OpenRouter](./openrouter.mdx) |
|
|
71
|
+
| `/v1/messages` | Anthropic Messages API: Anthropic, and the Claude models of [Azure AI](./azure-ai.mdx), [Bedrock Mantle](./bedrock-mantle.mdx), [Vertex AI](./vertex-ai.mdx) and [OpenRouter](./openrouter.mdx) |
|
|
72
72
|
|
|
73
73
|
## Request headers
|
|
74
74
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "zuplo",
|
|
3
|
-
"version": "7.8.
|
|
3
|
+
"version": "7.8.16",
|
|
4
4
|
"type": "module",
|
|
5
5
|
"description": "The official Zuplo CLI for local development and platform management",
|
|
6
6
|
"homepage": "https://zuplo.com/docs/cli/overview",
|
|
@@ -32,9 +32,9 @@
|
|
|
32
32
|
"zuplo": "zuplo.js"
|
|
33
33
|
},
|
|
34
34
|
"dependencies": {
|
|
35
|
-
"@zuplo/cli": "7.8.
|
|
36
|
-
"@zuplo/core": "7.8.
|
|
37
|
-
"@zuplo/runtime": "7.8.
|
|
38
|
-
"@zuplo/test": "7.8.
|
|
35
|
+
"@zuplo/cli": "7.8.16",
|
|
36
|
+
"@zuplo/core": "7.8.16",
|
|
37
|
+
"@zuplo/runtime": "7.8.16",
|
|
38
|
+
"@zuplo/test": "7.8.16"
|
|
39
39
|
}
|
|
40
40
|
}
|