zuplo 7.9.5 → 7.9.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/docs/ai-gateway/apps.mdx +28 -16
- package/docs/ai-gateway/custom-policies.mdx +7 -7
- package/docs/ai-gateway/custom-providers.mdx +1 -1
- package/docs/ai-gateway/fallback.mdx +5 -5
- package/docs/ai-gateway/integrations/ai-sdk.mdx +2 -2
- package/docs/ai-gateway/integrations/claude-code.mdx +41 -33
- package/docs/ai-gateway/integrations/claude-desktop.mdx +37 -34
- package/docs/ai-gateway/integrations/codex.mdx +42 -31
- package/docs/ai-gateway/integrations/github-copilot.mdx +39 -37
- package/docs/ai-gateway/integrations/goose.mdx +27 -30
- package/docs/ai-gateway/integrations/langchain.mdx +2 -2
- package/docs/ai-gateway/integrations/openai.mdx +2 -2
- package/docs/ai-gateway/jev.mdx +344 -0
- package/docs/ai-gateway/managing-apps.mdx +22 -22
- package/docs/ai-gateway/managing-pools.mdx +87 -0
- package/docs/ai-gateway/managing-providers.mdx +4 -4
- package/docs/ai-gateway/overview.mdx +38 -26
- package/docs/ai-gateway/policy-chains.mdx +7 -7
- package/docs/ai-gateway/policy-templates.mdx +22 -22
- package/docs/ai-gateway/pools.mdx +53 -0
- package/docs/ai-gateway/providers.mdx +32 -2
- package/docs/ai-gateway/source-control.mdx +2 -2
- package/docs/ai-gateway/universal-api.mdx +4 -0
- package/docs/ai-gateway/usage-limits.mdx +49 -47
- package/docs/ai-gateway/user-apps.mdx +166 -0
- package/docs/articles/accounts/roles-and-permissions.mdx +92 -63
- package/docs/concepts/ai-gateway.mdx +24 -20
- package/docs/policies/ai-gateway-akamai-firewall-inbound/doc.md +1 -1
- package/docs/policies/ai-gateway-metering-inbound/schema.json +16 -2
- package/docs/policies/ai-gateway-smart-router-inbound/doc.md +172 -29
- package/docs/policies/ai-gateway-smart-router-inbound/intro.md +4 -3
- package/docs/policies/ai-gateway-smart-router-inbound/schema.json +177 -30
- package/package.json +5 -5
- package/docs/ai-gateway/managing-teams.mdx +0 -87
- package/docs/ai-gateway/teams.mdx +0 -49
|
@@ -2,8 +2,8 @@
|
|
|
2
2
|
title: GitHub Copilot
|
|
3
3
|
sidebar_label: GitHub Copilot
|
|
4
4
|
description:
|
|
5
|
-
Route GitHub Copilot in VS Code and the Copilot CLI through
|
|
6
|
-
|
|
5
|
+
Route GitHub Copilot in VS Code and the Copilot CLI through the AI Gateway
|
|
6
|
+
with a personal API key, so every Copilot model request is metered per person.
|
|
7
7
|
---
|
|
8
8
|
|
|
9
9
|
[GitHub Copilot](https://github.com/features/copilot) can use your own model
|
|
@@ -11,26 +11,26 @@ endpoint instead of GitHub's hosted models — the
|
|
|
11
11
|
[Copilot CLI](https://docs.github.com/en/copilot/how-tos/copilot-cli/customize-copilot/use-byok-models)
|
|
12
12
|
through environment variables, and VS Code through its
|
|
13
13
|
[Custom Endpoint provider](https://code.visualstudio.com/docs/agent-customization/language-models#_add-a-custom-endpoint-model).
|
|
14
|
-
Point that endpoint at
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
spend as they work.
|
|
14
|
+
Point that endpoint at the AI Gateway, and every Copilot chat and agent request
|
|
15
|
+
is authenticated, metered against a budget, and limited to the models you allow.
|
|
16
|
+
Developers keep using Copilot exactly as before, and their requests, tokens, and
|
|
17
|
+
spend show up in the Zuplo Portal as they work.
|
|
19
18
|
|
|
20
19
|
## Before you start
|
|
21
20
|
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
1. Create a [provider](../managing-providers.mdx) for the models Copilot will
|
|
25
|
-
use, such as OpenAI or Anthropic
|
|
26
|
-
|
|
27
|
-
2. [Create a team](../managing-teams.mdx) for the developers using Copilot
|
|
21
|
+
Create a [provider](../managing-providers.mdx) for the models Copilot will use,
|
|
22
|
+
such as OpenAI or Anthropic, then get a gateway URL and key in one of two ways:
|
|
28
23
|
|
|
29
|
-
|
|
24
|
+
- **Use your personal API key (recommended).** Each developer calls the
|
|
25
|
+
project's [User App](../user-apps.mdx) with their own key. Open the
|
|
26
|
+
[**Home**](https://portal.zuplo.com/+/account/project/ai/home) tab of your AI
|
|
27
|
+
Gateway project, copy the **Gateway URL**, and create a key under **My API
|
|
28
|
+
keys**. Each developer's requests count against their own budget.
|
|
29
|
+
- **Use an app.** For a shared setup with one key,
|
|
30
|
+
[create a pool](../managing-pools.mdx) and [an app](../managing-apps.mdx) for
|
|
31
|
+
Copilot, then copy the app's **API URL** and **API Key**.
|
|
30
32
|
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
</Stepper>
|
|
33
|
+
The rest of this guide calls these your **gateway URL** and **gateway key**.
|
|
34
34
|
|
|
35
35
|
The gateway names models as `providerName/model`, where `providerName` is the
|
|
36
36
|
name you gave the provider. The examples below use `openai` and `anthropic`;
|
|
@@ -39,12 +39,12 @@ substitute your own provider names.
|
|
|
39
39
|
## Copilot CLI
|
|
40
40
|
|
|
41
41
|
The CLI reads its provider from four environment variables. Set the base URL to
|
|
42
|
-
|
|
42
|
+
your gateway URL plus `/v1`, and name a model that supports tool calls, such as
|
|
43
43
|
GPT-4.1:
|
|
44
44
|
|
|
45
45
|
```bash
|
|
46
|
-
export COPILOT_PROVIDER_BASE_URL="https://my-gateway-main-2e18f50.zuplo.app/
|
|
47
|
-
export COPILOT_PROVIDER_API_KEY="<your-
|
|
46
|
+
export COPILOT_PROVIDER_BASE_URL="https://my-gateway-main-2e18f50.zuplo.app/u/741d375d631b429293481d6d0458bb64/v1"
|
|
47
|
+
export COPILOT_PROVIDER_API_KEY="<your-gateway-key>"
|
|
48
48
|
export COPILOT_PROVIDER_TYPE="openai"
|
|
49
49
|
export COPILOT_MODEL="openai/gpt-4.1-mini"
|
|
50
50
|
```
|
|
@@ -52,13 +52,13 @@ export COPILOT_MODEL="openai/gpt-4.1-mini"
|
|
|
52
52
|
Start `copilot` from the same shell. Every request in the session now passes
|
|
53
53
|
through the gateway.
|
|
54
54
|
|
|
55
|
-
For a Claude model, use the `anthropic` provider type with
|
|
55
|
+
For a Claude model, use the `anthropic` provider type with your gateway URL
|
|
56
56
|
without `/v1` — the CLI appends `/v1/messages` itself — and pass the key as a
|
|
57
57
|
bearer token, which is the header the gateway reads:
|
|
58
58
|
|
|
59
59
|
```bash
|
|
60
|
-
export COPILOT_PROVIDER_BASE_URL="https://my-gateway-main-2e18f50.zuplo.app/
|
|
61
|
-
export COPILOT_PROVIDER_BEARER_TOKEN="<your-
|
|
60
|
+
export COPILOT_PROVIDER_BASE_URL="https://my-gateway-main-2e18f50.zuplo.app/u/741d375d631b429293481d6d0458bb64"
|
|
61
|
+
export COPILOT_PROVIDER_BEARER_TOKEN="<your-gateway-key>"
|
|
62
62
|
export COPILOT_PROVIDER_TYPE="anthropic"
|
|
63
63
|
export COPILOT_MODEL="anthropic/claude-sonnet-4-6"
|
|
64
64
|
```
|
|
@@ -79,8 +79,8 @@ configuration file VS Code opens.
|
|
|
79
79
|
|
|
80
80
|
2. Select **Add Models**, then **Custom Endpoint**
|
|
81
81
|
|
|
82
|
-
3. Enter a name for the group, such as `Zuplo AI Gateway`, and paste
|
|
83
|
-
|
|
82
|
+
3. Enter a name for the group, such as `Zuplo AI Gateway`, and paste your
|
|
83
|
+
gateway key
|
|
84
84
|
|
|
85
85
|
4. Select **Chat Completions** as the API type
|
|
86
86
|
|
|
@@ -89,7 +89,8 @@ configuration file VS Code opens.
|
|
|
89
89
|
|
|
90
90
|
</Stepper>
|
|
91
91
|
|
|
92
|
-
This configuration offers a GPT model and a Claude model
|
|
92
|
+
This configuration offers a GPT model and a Claude model through one gateway
|
|
93
|
+
URL:
|
|
93
94
|
|
|
94
95
|
```json
|
|
95
96
|
[
|
|
@@ -102,7 +103,7 @@ This configuration offers a GPT model and a Claude model from one app:
|
|
|
102
103
|
"id": "openai/gpt-5.6",
|
|
103
104
|
"name": "GPT-5.6 (Zuplo)",
|
|
104
105
|
"apiType": "responses",
|
|
105
|
-
"url": "https://my-gateway-main-2e18f50.zuplo.app/
|
|
106
|
+
"url": "https://my-gateway-main-2e18f50.zuplo.app/u/741d375d631b429293481d6d0458bb64/v1/responses",
|
|
106
107
|
"toolCalling": true,
|
|
107
108
|
"vision": true,
|
|
108
109
|
"maxInputTokens": 128000,
|
|
@@ -113,7 +114,7 @@ This configuration offers a GPT model and a Claude model from one app:
|
|
|
113
114
|
"id": "anthropic/claude-sonnet-5",
|
|
114
115
|
"name": "Claude Sonnet 5 (Zuplo)",
|
|
115
116
|
"apiType": "messages",
|
|
116
|
-
"url": "https://my-gateway-main-2e18f50.zuplo.app/
|
|
117
|
+
"url": "https://my-gateway-main-2e18f50.zuplo.app/u/741d375d631b429293481d6d0458bb64/v1/messages",
|
|
117
118
|
"toolCalling": true,
|
|
118
119
|
"vision": true,
|
|
119
120
|
"maxInputTokens": 200000,
|
|
@@ -145,9 +146,9 @@ Save the file and pick a model from the chat model picker.
|
|
|
145
146
|
## What routes through the gateway
|
|
146
147
|
|
|
147
148
|
Chat and agent requests, on both surfaces. Copilot's inline code completions
|
|
148
|
-
stay on GitHub's infrastructure. The
|
|
149
|
+
stay on GitHub's infrastructure. The
|
|
149
150
|
[Model Filtering](../../policies/ai-gateway-model-filtering-inbound.mdx) policy
|
|
150
|
-
decides which models Copilot may use.
|
|
151
|
+
on the User App or the app decides which models Copilot may use.
|
|
151
152
|
|
|
152
153
|
## Troubleshooting
|
|
153
154
|
|
|
@@ -156,10 +157,11 @@ Copilot reports gateway errors as a retry failure, such as
|
|
|
156
157
|
**GitHub Copilot Chat** channel of VS Code's Output panel, or in the CLI's error
|
|
157
158
|
text.
|
|
158
159
|
|
|
159
|
-
| The gateway responds | Cause
|
|
160
|
-
| ----------------------------------------------------------------- |
|
|
161
|
-
| `
|
|
162
|
-
| `401` `
|
|
163
|
-
| `
|
|
164
|
-
| `
|
|
165
|
-
| `
|
|
160
|
+
| The gateway responds | Cause | Fix |
|
|
161
|
+
| ----------------------------------------------------------------- | ---------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
|
|
162
|
+
| `403` `This User App route requires a personal API key` | A personal key that's missing, wrong, from another project, or sent as `x-api-key` | VS Code: add `requestHeaders`. CLI: use `COPILOT_PROVIDER_BEARER_TOKEN`. Check the key on **Home** |
|
|
163
|
+
| `401` `Header configured by options.authHeader is missing` | With an app key: the key was sent as `x-api-key` | VS Code: add `requestHeaders`. CLI: use `COPILOT_PROVIDER_BEARER_TOKEN` |
|
|
164
|
+
| `401` `Invalid Authorization Scheme` | VS Code's `apiKey` was replaced with the key itself, so the key sent was empty | Re-enter the key with **Chat: Manage Language Models** |
|
|
165
|
+
| `404` | The `url` stops at the gateway URL | Append `/v1/responses`, `/v1/messages`, or `/v1/chat/completions` |
|
|
166
|
+
| `Function tools with reasoning_effort are not supported` | A GPT-5.x model on Chat Completions | VS Code: `"apiType": "responses"`. CLI: use a GPT-4.1 or Claude model |
|
|
167
|
+
| `'temperature' does not support 0` or `temperature is deprecated` | The CLI's temperature setting on a GPT-5.x or Claude Sonnet 5 model | Use GPT-4.1, GPT-4o, Claude Sonnet 4.6, or Claude Haiku 4.5 |
|
|
@@ -2,8 +2,9 @@
|
|
|
2
2
|
title: goose
|
|
3
3
|
sidebar_label: Goose
|
|
4
4
|
description:
|
|
5
|
-
Add
|
|
6
|
-
desktop app so goose's requests route through the
|
|
5
|
+
Add the AI Gateway as an OpenAI-compatible provider in the goose CLI or
|
|
6
|
+
desktop app, with your personal API key, so goose's requests route through the
|
|
7
|
+
gateway.
|
|
7
8
|
---
|
|
8
9
|
|
|
9
10
|
[goose](https://block.github.io/goose/) is a local AI agent and CLI tool for
|
|
@@ -14,22 +15,20 @@ powerful extensibility through recipes.
|
|
|
14
15
|
|
|
15
16
|
## Prerequisites
|
|
16
17
|
|
|
17
|
-
|
|
18
|
-
|
|
18
|
+
Create a [provider](../managing-providers.mdx) in the AI Gateway for the
|
|
19
|
+
provider you want to use with goose, then get a gateway URL and key in one of
|
|
20
|
+
two ways:
|
|
19
21
|
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
22
|
+
- **Use your personal API key (recommended).** goose runs on your own machine,
|
|
23
|
+
so call the project's [User App](../user-apps.mdx). Open the
|
|
24
|
+
[**Home**](https://portal.zuplo.com/+/account/project/ai/home) tab of your AI
|
|
25
|
+
Gateway project, copy the **Gateway URL**, and create a key under **My API
|
|
26
|
+
keys**. Your requests count against your own budget.
|
|
27
|
+
- **Use an app.** For a shared or automated setup,
|
|
28
|
+
[create a pool](../managing-pools.mdx) and [an app](../managing-apps.mdx) for
|
|
29
|
+
goose, then copy the app's **API URL** and **API Key**.
|
|
26
30
|
|
|
27
|
-
|
|
28
|
-
assign it to the team you created
|
|
29
|
-
|
|
30
|
-
4. Copy the **API URL** and **API Key** shown at the top of the app page
|
|
31
|
-
|
|
32
|
-
</Stepper>
|
|
31
|
+
The rest of this guide calls these your **gateway URL** and **gateway key**.
|
|
33
32
|
|
|
34
33
|
## Configure goose
|
|
35
34
|
|
|
@@ -47,12 +46,11 @@ configuration approaches depending on which version you choose.
|
|
|
47
46
|
3. Choose **OpenAI** as the provider you want to add (this will work for any
|
|
48
47
|
OpenAI compatible provider and model)
|
|
49
48
|
|
|
50
|
-
4. Set the `OPENAI_API_KEY` to
|
|
51
|
-
Zuplo
|
|
49
|
+
4. Set the `OPENAI_API_KEY` to your gateway key
|
|
52
50
|
|
|
53
|
-
5. Set the `OPENAI_HOST` to your
|
|
54
|
-
|
|
55
|
-
`https://my-gateway-main-2e18f50.zuplo.app/
|
|
51
|
+
5. Set the `OPENAI_HOST` to your gateway URL _without_ the `/v1` suffix (for
|
|
52
|
+
example
|
|
53
|
+
`https://my-gateway-main-2e18f50.zuplo.app/u/741d375d631b429293481d6d0458bb64`)—the
|
|
56
54
|
base path in the next step supplies it
|
|
57
55
|
|
|
58
56
|
6. The `OPENAI_BASE_PATH` may already be configured. If it's not, enter
|
|
@@ -65,10 +63,9 @@ configuration approaches depending on which version you choose.
|
|
|
65
63
|
|
|
66
64
|
:::note
|
|
67
65
|
|
|
68
|
-
The
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
that policy allows.
|
|
66
|
+
The [Model Filtering](../../policies/ai-gateway-model-filtering-inbound.mdx)
|
|
67
|
+
policy on the User App or the app controls which models you may use, so the
|
|
68
|
+
model you enter here must be one that policy allows.
|
|
72
69
|
|
|
73
70
|
:::
|
|
74
71
|
|
|
@@ -96,13 +93,13 @@ Gateway by following these steps:
|
|
|
96
93
|
|
|
97
94
|
6. Enter a **Display Name** (for example Zuplo AI Gateway)
|
|
98
95
|
|
|
99
|
-
7. Set the **API URL** to your
|
|
96
|
+
7. Set the **API URL** to your gateway URL
|
|
100
97
|
|
|
101
|
-
8. Set the **API Key** to
|
|
98
|
+
8. Set the **API Key** to your gateway key
|
|
102
99
|
|
|
103
|
-
9. Set the list of **Available Models** to the models the
|
|
104
|
-
|
|
105
|
-
`openai/gpt-5
|
|
100
|
+
9. Set the list of **Available Models** to the models the Model Filtering policy
|
|
101
|
+
allows, named as `providerName/model` (for example `openai/gpt-5-mini`,
|
|
102
|
+
`openai/gpt-5`, `openai/gpt-5-nano`)
|
|
106
103
|
|
|
107
104
|
10. Click on **Create Provider**
|
|
108
105
|
|
|
@@ -21,10 +21,10 @@ complete these steps first:
|
|
|
21
21
|
1. Create a [new provider](../managing-providers.mdx) in the AI Gateway for the
|
|
22
22
|
provider you want to use with LangChain
|
|
23
23
|
|
|
24
|
-
2. [Set up a new
|
|
24
|
+
2. [Set up a new pool](../managing-pools.mdx)
|
|
25
25
|
|
|
26
26
|
3. Create a [new app](../managing-apps.mdx) to use specifically with LangChain
|
|
27
|
-
and assign it to the
|
|
27
|
+
and assign it to the pool you created
|
|
28
28
|
|
|
29
29
|
4. Copy the **API URL** and **API Key** shown at the top of the app page
|
|
30
30
|
|
|
@@ -21,10 +21,10 @@ complete these steps first:
|
|
|
21
21
|
1. Create a [new provider](../managing-providers.mdx) in the AI Gateway for
|
|
22
22
|
OpenAI
|
|
23
23
|
|
|
24
|
-
2. [Set up a new
|
|
24
|
+
2. [Set up a new pool](../managing-pools.mdx)
|
|
25
25
|
|
|
26
26
|
3. Create a [new app](../managing-apps.mdx) to use specifically with the OpenAI
|
|
27
|
-
SDK and assign it to the
|
|
27
|
+
SDK and assign it to the pool you created
|
|
28
28
|
|
|
29
29
|
4. Copy the **API URL** and **API Key** shown at the top of the app page
|
|
30
30
|
|
|
@@ -0,0 +1,344 @@
|
|
|
1
|
+
---
|
|
2
|
+
title: Using Jev
|
|
3
|
+
sidebar_label: Jev
|
|
4
|
+
description:
|
|
5
|
+
Call TypeSafe's Jev System One model through your AI Gateway. Send a state and
|
|
6
|
+
typed questions to your app's System One endpoint and get calibrated answers.
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
**Jev** is [TypeSafe's](https://typesafe.ai) System One model. It doesn't write
|
|
10
|
+
text. You send it a piece of text—the _state_—and a set of typed questions, and
|
|
11
|
+
it returns a calibrated answer for each one: a probability, a choice from
|
|
12
|
+
options you define, or a score on a scale you define. Adding it as a provider
|
|
13
|
+
puts the gateway in front of TypeSafe's API for authentication, model filtering,
|
|
14
|
+
usage metering, and cost tracking, on your own TypeSafe account and billing.
|
|
15
|
+
|
|
16
|
+
Jev differs from every other provider in three ways that change what you type:
|
|
17
|
+
|
|
18
|
+
- **It has its own endpoint.** Requests go to `/v1/systemone`, not to chat
|
|
19
|
+
completions.
|
|
20
|
+
- **Chat clients can't call it.** Every other endpoint refuses a Jev model, and
|
|
21
|
+
the default model list doesn't include it.
|
|
22
|
+
- **Only input tokens cost money.** TypeSafe charges $0.042 per million input
|
|
23
|
+
tokens, and output tokens are free.
|
|
24
|
+
|
|
25
|
+
## What Jev answers
|
|
26
|
+
|
|
27
|
+
Each request carries a `state` and a set of named `questions`. A question's type
|
|
28
|
+
decides the shape of its answer:
|
|
29
|
+
|
|
30
|
+
| Question type | Asks | Answer |
|
|
31
|
+
| ------------- | ----------------------------------- | -------------------------------------------------------------------- |
|
|
32
|
+
| `noul` | A yes-or-no question | A probability between 0 and 1 |
|
|
33
|
+
| `choice` | Which of your options applies | The chosen option, a probability for each option, and a confidence |
|
|
34
|
+
| `score` | Where the state falls on your scale | A score on the scale, a probability for each level, and a confidence |
|
|
35
|
+
|
|
36
|
+
The `state` is text: a string, a JSON object, or an array of text values. Jev
|
|
37
|
+
doesn't accept images, audio, or video. TypeSafe's documentation explains
|
|
38
|
+
[state](https://docs.typesafe.ai/concepts/state), the
|
|
39
|
+
[question types](https://docs.typesafe.ai/primitives), and how to read a
|
|
40
|
+
[confidence](https://docs.typesafe.ai/confidence).
|
|
41
|
+
|
|
42
|
+
## Supported endpoint
|
|
43
|
+
|
|
44
|
+
A Jev model serves one endpoint:
|
|
45
|
+
|
|
46
|
+
| Endpoint | Jev |
|
|
47
|
+
| ------------------------------------------------------------------------- | --- |
|
|
48
|
+
| `/v1/systemone` | ✅ |
|
|
49
|
+
| `/v1/chat/completions`, `/v1/responses`, `/v1/messages`, `/v1/embeddings` | ❌ |
|
|
50
|
+
|
|
51
|
+
The gateway forwards your request body to TypeSafe unchanged, except for the
|
|
52
|
+
`model` field, which it replaces with the model ID your provider serves. The
|
|
53
|
+
response body comes back unchanged too, including the `model` field, which
|
|
54
|
+
reports the versioned model that answered. The `x-typesafe-request-id` header
|
|
55
|
+
comes back as well. TypeSafe's own error responses (`401`, `422`, `429`, and
|
|
56
|
+
`529`) pass through with TypeSafe's body, so the field a `422` names is the one
|
|
57
|
+
to fix.
|
|
58
|
+
|
|
59
|
+
The refusal works in both directions. A Jev model sent to a chat endpoint
|
|
60
|
+
(`/v1/chat/completions`, `/v1/responses`, or `/v1/messages`) returns a `400`
|
|
61
|
+
that names the endpoint, and a model from any other provider sent to
|
|
62
|
+
`/v1/systemone` returns a `400` that names Jev. `/v1/embeddings` fails
|
|
63
|
+
differently: Jev has no embedding models, so the gateway returns a `500` saying
|
|
64
|
+
the model `is not included in model selections`.
|
|
65
|
+
|
|
66
|
+
## Before you begin
|
|
67
|
+
|
|
68
|
+
You need:
|
|
69
|
+
|
|
70
|
+
- A TypeSafe account and an API key. Create a key in the
|
|
71
|
+
[TypeSafe console](https://console.typesafe.ai/keys).
|
|
72
|
+
- An AI Gateway project in the Zuplo Portal, and an [app](./apps.mdx) to call
|
|
73
|
+
from. The app page shows its API URL, and its API key lives on the app's **API
|
|
74
|
+
Key** tab.
|
|
75
|
+
|
|
76
|
+
## Add the provider
|
|
77
|
+
|
|
78
|
+
Adding or editing providers requires the **Edit** permission, granted to Zuplo
|
|
79
|
+
account and project **Admins**—see
|
|
80
|
+
[Managing Providers](./managing-providers.mdx).
|
|
81
|
+
|
|
82
|
+
<Stepper>
|
|
83
|
+
|
|
84
|
+
1. Open
|
|
85
|
+
[**Settings → AI Providers**](https://portal.zuplo.com/+/account/project/ai/settings/data-models)
|
|
86
|
+
in your AI Gateway project in the Zuplo Portal.
|
|
87
|
+
|
|
88
|
+
1. Click the **Add Provider** button.
|
|
89
|
+
|
|
90
|
+
1. In the **AI Provider** list, select **Jev**.
|
|
91
|
+
|
|
92
|
+
1. Review the **Provider Name**, which fills in as `jev`. You can replace it
|
|
93
|
+
with your own name, but only now—the name is permanent after creation, and
|
|
94
|
+
it's the prefix in every model reference.
|
|
95
|
+
|
|
96
|
+
1. Paste your TypeSafe **API Key**. There's no endpoint or region to enter,
|
|
97
|
+
because TypeSafe is a single global host.
|
|
98
|
+
|
|
99
|
+
1. Select the models to enable, then click **Create**.
|
|
100
|
+
|
|
101
|
+
</Stepper>
|
|
102
|
+
|
|
103
|
+
:::note
|
|
104
|
+
|
|
105
|
+
Saving provider settings triggers an automatic production deployment of your
|
|
106
|
+
gateway, because provider credentials are part of the deployed gateway. The
|
|
107
|
+
change is live once the deployment completes.
|
|
108
|
+
|
|
109
|
+
:::
|
|
110
|
+
|
|
111
|
+
## Model references
|
|
112
|
+
|
|
113
|
+
Apps reference models as `providerName/model`, where `providerName` is the name
|
|
114
|
+
you gave the provider. The Jev catalog has three models, so a provider named
|
|
115
|
+
`jev` can serve:
|
|
116
|
+
|
|
117
|
+
| Model reference | What it is |
|
|
118
|
+
| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
|
|
119
|
+
| `jev/jev-latest` | TypeSafe's most recent stable release, and the default in TypeSafe's SDKs. |
|
|
120
|
+
| `jev/jev-1.13.0` | A specific release. Pin it when you tune confidence thresholds against one version. |
|
|
121
|
+
| `jev/jev-preview` | TypeSafe's most recent release, official or not. It points to the same model as `jev-latest` until TypeSafe publishes a preview build. |
|
|
122
|
+
|
|
123
|
+
An alias moves when TypeSafe ships a release, so the answers behind it can
|
|
124
|
+
change without a change on your side. The response's `model` field reports the
|
|
125
|
+
versioned ID that answered, so you can log which release produced each result.
|
|
126
|
+
|
|
127
|
+
## Call the model
|
|
128
|
+
|
|
129
|
+
Send requests to `/v1/systemone` under your app's URL, with the app's API key as
|
|
130
|
+
a bearer token. Include the `model` in every request, as `jev/<model>`.
|
|
131
|
+
|
|
132
|
+
### With TypeSafe's SDK
|
|
133
|
+
|
|
134
|
+
TypeSafe's JavaScript SDK works against your app. Set `baseURL` to the app URL
|
|
135
|
+
**without** `/v1`, because the SDK appends `/v1/systemone` itself. The `apiKey`
|
|
136
|
+
is the app's API key, not your TypeSafe key.
|
|
137
|
+
|
|
138
|
+
```ts
|
|
139
|
+
import { choice, noul, score, TypeSafeClient } from "@typesafe-ai/sdk";
|
|
140
|
+
|
|
141
|
+
const client = new TypeSafeClient({
|
|
142
|
+
baseURL:
|
|
143
|
+
"https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e",
|
|
144
|
+
apiKey: process.env.ZUPLO_APP_API_KEY,
|
|
145
|
+
defaultModel: "jev/jev-latest",
|
|
146
|
+
});
|
|
147
|
+
|
|
148
|
+
const response = await client.systemOne({
|
|
149
|
+
state: "Help! My payouts have been failing for 3 days.",
|
|
150
|
+
questions: {
|
|
151
|
+
urgent: noul("Does this convey urgency?"),
|
|
152
|
+
department: choice("Which team should handle this?", {
|
|
153
|
+
billing: "Payments and invoices",
|
|
154
|
+
technical: null,
|
|
155
|
+
}),
|
|
156
|
+
frustration: score("How frustrated is the customer?", [
|
|
157
|
+
"Calm",
|
|
158
|
+
"Frustrated",
|
|
159
|
+
"Very angry",
|
|
160
|
+
]),
|
|
161
|
+
},
|
|
162
|
+
});
|
|
163
|
+
|
|
164
|
+
console.log(response.answers.department.choice);
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
To use a different model for one call, pass `model: "jev/jev-1.13.0"` to
|
|
168
|
+
`systemOne()`.
|
|
169
|
+
|
|
170
|
+
### With HTTP
|
|
171
|
+
|
|
172
|
+
```bash
|
|
173
|
+
curl https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e/v1/systemone \
|
|
174
|
+
-H "Authorization: Bearer $ZUPLO_APP_API_KEY" \
|
|
175
|
+
-H "Content-Type: application/json" \
|
|
176
|
+
-d '{
|
|
177
|
+
"model": "jev/jev-latest",
|
|
178
|
+
"state": "Help! My payouts have been failing for 3 days.",
|
|
179
|
+
"questions": {
|
|
180
|
+
"urgent": {
|
|
181
|
+
"type": "noul",
|
|
182
|
+
"instructions": "Does this convey urgency?"
|
|
183
|
+
},
|
|
184
|
+
"department": {
|
|
185
|
+
"type": "choice",
|
|
186
|
+
"instructions": "Which team should handle this?",
|
|
187
|
+
"criteria": { "billing": "Payments and invoices", "technical": null }
|
|
188
|
+
},
|
|
189
|
+
"frustration": {
|
|
190
|
+
"type": "score",
|
|
191
|
+
"instructions": "How frustrated is the customer?",
|
|
192
|
+
"criteria": ["Calm", "Frustrated", "Very angry"]
|
|
193
|
+
}
|
|
194
|
+
}
|
|
195
|
+
}'
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
The response is TypeSafe's, unchanged. The values below are illustrative:
|
|
199
|
+
|
|
200
|
+
```json
|
|
201
|
+
{
|
|
202
|
+
"model": "jev-1.13.0",
|
|
203
|
+
"answers": {
|
|
204
|
+
"urgent": { "type": "noul", "noul": 0.95 },
|
|
205
|
+
"department": {
|
|
206
|
+
"type": "choice",
|
|
207
|
+
"choice": "billing",
|
|
208
|
+
"probabilities": { "billing": 0.88, "technical": 0.12 },
|
|
209
|
+
"confidence": 0.81
|
|
210
|
+
},
|
|
211
|
+
"frustration": {
|
|
212
|
+
"type": "score",
|
|
213
|
+
"score": 1.05,
|
|
214
|
+
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
|
|
215
|
+
"probabilities": { "0": 0, "1": 0.95, "2": 0.05 },
|
|
216
|
+
"confidence": 0.92
|
|
217
|
+
}
|
|
218
|
+
},
|
|
219
|
+
"usage": { "input_tokens": 296, "output_tokens": 20 }
|
|
220
|
+
}
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
### List the Jev models
|
|
224
|
+
|
|
225
|
+
The default `GET /v1/models` list holds chat models, so it leaves Jev out. To
|
|
226
|
+
list your Jev models, send `x-zuplo-models-endpoint: systemone`:
|
|
227
|
+
|
|
228
|
+
```bash
|
|
229
|
+
curl https://my-gateway-main-2e18f50.zuplo.app/config_fe0a04972d2848e0a94ae4b8bcd1497e/v1/models \
|
|
230
|
+
-H "Authorization: Bearer $ZUPLO_APP_API_KEY" \
|
|
231
|
+
-H "x-zuplo-models-endpoint: systemone"
|
|
232
|
+
```
|
|
233
|
+
|
|
234
|
+
The list has one entry for each enabled model, such as `jev/jev-latest`.
|
|
235
|
+
|
|
236
|
+
## Pricing and budgets
|
|
237
|
+
|
|
238
|
+
The gateway prices each request from the model catalog at the rate on TypeSafe's
|
|
239
|
+
[Models page](https://docs.typesafe.ai/models): $0.042 per million input tokens,
|
|
240
|
+
and nothing for output. The response carries `X-Cost-USD` with the cost and
|
|
241
|
+
`X-Cost-Source: catalog`, which tells you the figure came from the catalog
|
|
242
|
+
rather than from TypeSafe. Each request counts toward token, request, and
|
|
243
|
+
spending [budgets](./usage-limits.mdx).
|
|
244
|
+
|
|
245
|
+
A token budget counts output tokens too, even though they're free, so it's
|
|
246
|
+
slightly stricter than your TypeSafe invoice. A spending budget matches it.
|
|
247
|
+
|
|
248
|
+
When a budget is exhausted, the gateway refuses the request with a `429` before
|
|
249
|
+
it calls TypeSafe. A System One request never falls back to another model,
|
|
250
|
+
because no chat model can answer a set of typed questions.
|
|
251
|
+
|
|
252
|
+
TypeSafe's own rate limits apply to your key. A request over a limit returns
|
|
253
|
+
TypeSafe's `429` with its `retry-after` header. The Models page lists the
|
|
254
|
+
current limits.
|
|
255
|
+
|
|
256
|
+
## Policies on System One routes
|
|
257
|
+
|
|
258
|
+
These routes run your app's [policy chain](./policy-chains.mdx) like any other.
|
|
259
|
+
What changes is the request body: it's a System One request, not a chat message,
|
|
260
|
+
so a policy that reads chat content applies its `onUnknownShape` setting
|
|
261
|
+
instead. Each policy's default follows from its job: a guardrail fails closed,
|
|
262
|
+
while an observer or an optimization fails open.
|
|
263
|
+
|
|
264
|
+
| Policy | On System One routes |
|
|
265
|
+
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
266
|
+
| [Authentication](../policies/ai-gateway-auth-inbound.mdx) | Works. Reads the app key from `Authorization: Bearer`, which is how TypeSafe's SDK sends it. |
|
|
267
|
+
| [Model Filtering](../policies/ai-gateway-model-filtering-inbound.mdx) | Works. Add the Jev models to the allow list. A Jev model outside the list returns `403`, and a listed model from another provider returns `400`. |
|
|
268
|
+
| [Model Override](../policies/ai-gateway-model-override-inbound.mdx) | Works when the configured model is a Jev model. A chat model returns a configuration error that names the option, such as `options.models.completions.force`. |
|
|
269
|
+
| [Metering and budgets](../policies/ai-gateway-metering-inbound.mdx) | Works. Counts each request with its tokens and cost. An exhausted budget returns `429`. |
|
|
270
|
+
| [DLP](../policies/ai-gateway-dlp-inbound.mdx), [Akamai AI Firewall](../policies/ai-gateway-akamai-firewall-inbound.mdx), and [Prompt Injection](../policies/ai-gateway-prompt-injection.mdx) | **Blocks by default.** The policy can't inspect a System One body, so it denies the request with a `400` and the error code `guardrail_uninspectable`. Set `onUnknownShape` to `skip` to forward it uninspected. |
|
|
271
|
+
| [Galileo](../policies/ai-gateway-galileo-tracing-inbound.mdx) and [Comet Opik](../policies/ai-gateway-opik-tracing-inbound.mdx) tracing | **Skipped by default.** The request goes through untraced. Set `onUnknownShape` to `deny` to make a trace a hard requirement. |
|
|
272
|
+
| Semantic Cache and Smart Router | Skipped. |
|
|
273
|
+
| [Fallback Model](./fallback.mdx) | Doesn't apply. A System One request goes to the one model it names, so a configured backup isn't used. |
|
|
274
|
+
|
|
275
|
+
:::tip{title="Give System One its own app"}
|
|
276
|
+
|
|
277
|
+
An app can serve chat models and Jev side by side, as long as every request
|
|
278
|
+
names its model and no policy overrides it. Model Override's `force` replaces
|
|
279
|
+
the model in every request, and its `default` and the first entry of a Model
|
|
280
|
+
Filtering allow list supply one when a request omits it. Each picks a single
|
|
281
|
+
model, which can't be both a chat model and a Jev model. If you use one of them,
|
|
282
|
+
give System One its own app.
|
|
283
|
+
|
|
284
|
+
:::
|
|
285
|
+
|
|
286
|
+
## Troubleshooting
|
|
287
|
+
|
|
288
|
+
**A `400` that names a chat endpoint and your Jev provider.** You sent a Jev
|
|
289
|
+
model to `/v1/chat/completions`, `/v1/responses`, or `/v1/messages`. Call
|
|
290
|
+
`/v1/systemone` instead.
|
|
291
|
+
|
|
292
|
+
**A `400` saying `/v1/systemone` is only supported by providers that serve the
|
|
293
|
+
TypeSafe System One API.** The model in the request belongs to another provider.
|
|
294
|
+
Use a `jev/<model>` reference.
|
|
295
|
+
|
|
296
|
+
**A `400` saying the route requires the `jev` API.** The Model Filtering policy
|
|
297
|
+
found a model that isn't a Jev model. Send a `jev/<model>` reference, and put a
|
|
298
|
+
Jev model first in the allow list if requests omit `model`.
|
|
299
|
+
|
|
300
|
+
**A `403` from the Model Filtering policy.** The Jev model isn't in the app's
|
|
301
|
+
allow list. Add it, for example `jev/jev-latest`.
|
|
302
|
+
|
|
303
|
+
**A `400` with the error code `guardrail_uninspectable`.** A guardrail policy in
|
|
304
|
+
the app's chain can't inspect System One requests and denies them by default.
|
|
305
|
+
See [the policy table](#policies-on-system-one-routes).
|
|
306
|
+
|
|
307
|
+
**A `401`.** Two different keys can cause one. A `401` with a Problem Details
|
|
308
|
+
body (`application/problem+json`) comes from the gateway's authentication, so
|
|
309
|
+
check the app's API key. A `401` whose `detail.message` begins
|
|
310
|
+
`Cannot authenticate with the server` comes from TypeSafe, which rejected the
|
|
311
|
+
provider's API key. Update it in
|
|
312
|
+
[**Settings → AI Providers**](https://portal.zuplo.com/+/account/project/ai/settings/data-models)
|
|
313
|
+
with a key from the [TypeSafe console](https://console.typesafe.ai/keys).
|
|
314
|
+
|
|
315
|
+
**A `422` naming a field.** TypeSafe rejected the request body, and the gateway
|
|
316
|
+
passed its answer through. The `detail` array names the field that failed
|
|
317
|
+
validation.
|
|
318
|
+
|
|
319
|
+
**An error saying the model `is not included in model selections`.** On
|
|
320
|
+
`/v1/embeddings`, this is expected: Jev has no embedding models, so call
|
|
321
|
+
`/v1/systemone` instead. On `/v1/systemone`, the model isn't enabled for the
|
|
322
|
+
provider. Open the provider in
|
|
323
|
+
[**Settings → AI Providers**](https://portal.zuplo.com/+/account/project/ai/settings/data-models),
|
|
324
|
+
enable the model, and save.
|
|
325
|
+
|
|
326
|
+
**Your app's model list has no Jev models.** Expected. The default list holds
|
|
327
|
+
chat models—see [List the Jev models](#list-the-jev-models).
|
|
328
|
+
|
|
329
|
+
**TypeSafe's `client.models.list()` throws.** The SDK expects TypeSafe's own
|
|
330
|
+
list format, and the gateway answers `GET /v1/models` in the OpenAI format. Name
|
|
331
|
+
the model in `defaultModel` or in each call instead, or list the models with the
|
|
332
|
+
`x-zuplo-models-endpoint` header.
|
|
333
|
+
|
|
334
|
+
## Next steps
|
|
335
|
+
|
|
336
|
+
- [AI Providers](./providers.mdx)—the capability matrix across every supported
|
|
337
|
+
provider.
|
|
338
|
+
- [Managing Providers](./managing-providers.mdx)—edit models and keys, and
|
|
339
|
+
understand when changes deploy.
|
|
340
|
+
- [AI Gateway Apps](./apps.mdx)—create the apps that call your models.
|
|
341
|
+
- [Model Filtering policy](../policies/ai-gateway-model-filtering-inbound.mdx)—control
|
|
342
|
+
which models each app can call.
|
|
343
|
+
- [Usage limits](./usage-limits.mdx)—set budgets on tokens, requests, and
|
|
344
|
+
spending.
|