@agentproto/llm-endpoint 0.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +202 -0
- package/README.md +574 -0
- package/dist/cli.mjs +3738 -0
- package/dist/cli.mjs.map +1 -0
- package/dist/index.d.ts +427 -0
- package/dist/index.mjs +3714 -0
- package/dist/index.mjs.map +1 -0
- package/package.json +69 -0
package/README.md
ADDED
|
@@ -0,0 +1,574 @@
|
|
|
1
|
+
# @agentproto/llm-endpoint
|
|
2
|
+
|
|
3
|
+
A lightweight **multi-surface LLM proxy gateway**. It exposes three API
|
|
4
|
+
surfaces from a single port:
|
|
5
|
+
|
|
6
|
+
- `POST /v1/messages` — Anthropic Messages compatibility. Claude-shaped model
|
|
7
|
+
aliases are supported **only** through an explicit local compatibility pack.
|
|
8
|
+
- `POST /v1/chat/completions` — OpenAI Chat Completions compatibility with
|
|
9
|
+
transparent `provider/model` routing.
|
|
10
|
+
- `POST /v1/responses` — OpenAI Responses API facade for Codex custom providers.
|
|
11
|
+
|
|
12
|
+
Requests are fanned out to upstream providers — **Moonshot, OpenRouter, ZAI/Zhipu,
|
|
13
|
+
Groq, xAI, direct OpenAI, Nebius AI Studio, a self-hosted "forge" server for
|
|
14
|
+
fine-tunes, and any number of named local/LAN model servers** (Ollama,
|
|
15
|
+
llama-server, vLLM, …, configured from `~/.agentproto/llm-endpoints.json`) —
|
|
16
|
+
using provider-native model references. The proxy also handles
|
|
17
|
+
Anthropic↔OpenAI schema translation, per-provider tool caps, orphaned-tool-call
|
|
18
|
+
repair, and thinking-block stripping where needed.
|
|
19
|
+
|
|
20
|
+
- **Package:** `@agentproto/llm-endpoint`
|
|
21
|
+
- **Entry:** `src/cli.ts` → `start()` in `src/index.ts`
|
|
22
|
+
- **Default port:** `18090` (override with `LLM_ENDPOINT_PORT`, or `PORT`)
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## Credentials this proxy handles
|
|
27
|
+
|
|
28
|
+
Read this before pointing the proxy at anything you care about. The package
|
|
29
|
+
ships **no credentials** — every key below is one you supply — but it is a
|
|
30
|
+
credential-forwarding component, so its handling rules are part of its
|
|
31
|
+
contract.
|
|
32
|
+
|
|
33
|
+
- **Per-provider API keys** resolve from one env var each
|
|
34
|
+
(`resolveSecretKeys`), or from `~/.agentproto/providers.json`. They are
|
|
35
|
+
injected at proxy boot and never logged.
|
|
36
|
+
- **Anthropic subscription OAuth tokens** (`sk-ant-oat…`) are supported as an
|
|
37
|
+
alternative to an API key on the `anthropic` upstream. When one is resolved,
|
|
38
|
+
the proxy sends `Authorization: Bearer <oat>` plus
|
|
39
|
+
`anthropic-beta: oauth-2025-04-20` and `anthropic-version: 2023-06-01`.
|
|
40
|
+
This path is **fail-closed**: an OAuth credential is *only* ever sent to the
|
|
41
|
+
`anthropic` upstream, and every other route `401`s rather than leaking it to
|
|
42
|
+
a third-party provider. Batch creation also refuses an OAuth token outright,
|
|
43
|
+
because the Anthropic Batches API only accepts API keys.
|
|
44
|
+
- A **stored** OAuth token works while it is valid. The proxy does not yet
|
|
45
|
+
drive a self-refreshing source (`source: "claude-code-oauth"`) to re-mint an
|
|
46
|
+
expired one.
|
|
47
|
+
- `KeychainStore` is **macOS-only**; elsewhere the proxy falls back to the
|
|
48
|
+
per-provider env key.
|
|
49
|
+
|
|
50
|
+
Fronting a personal Claude subscription through a shared gateway may not be
|
|
51
|
+
compatible with your Anthropic plan's terms — that is a question about **your
|
|
52
|
+
deployment**, not about this package. If you expose the proxy beyond localhost,
|
|
53
|
+
read [Securing a public deployment](#securing-a-public-deployment) first: the
|
|
54
|
+
access gate is **off by default**.
|
|
55
|
+
|
|
56
|
+
---
|
|
57
|
+
|
|
58
|
+
## Quick start (client config)
|
|
59
|
+
|
|
60
|
+
Point any Anthropic- or OpenAI-compatible client at the proxy:
|
|
61
|
+
|
|
62
|
+
| Setting | Value |
|
|
63
|
+
| :--- | :--- |
|
|
64
|
+
| **Base URL** | `http://localhost:18090/v1` (local), or your own public origin + `/v1` |
|
|
65
|
+
| **API key** | The proxy injects the *real* upstream key server-side, so the client key is never your provider key. If the [inbound access gate](#securing-a-public-deployment) is enabled, send one of its tokens as the bearer; if it is not, any non-empty value passes. |
|
|
66
|
+
|
|
67
|
+
The client asks for a model by its **provider-transparent reference**
|
|
68
|
+
(`provider/model`, e.g. `moonshot/kimi-k2.7-code`, `openai/gpt-4.1`); the proxy
|
|
69
|
+
routes it to the right upstream endpoint.
|
|
70
|
+
|
|
71
|
+
---
|
|
72
|
+
|
|
73
|
+
## Model catalog
|
|
74
|
+
|
|
75
|
+
### Transparent routing (`provider/model`)
|
|
76
|
+
|
|
77
|
+
On the OpenAI surfaces (`/v1/chat/completions` and `/v1/responses`) the model
|
|
78
|
+
field is parsed as `provider/model`:
|
|
79
|
+
|
|
80
|
+
| Provider | Example reference | Upstream endpoint |
|
|
81
|
+
| :--- | :--- | :--- |
|
|
82
|
+
| Moonshot | `moonshot/kimi-k2.7-code` | `api.moonshot.ai/v1/chat/completions` |
|
|
83
|
+
| OpenRouter | `openrouter/anthropic/claude-3-5-sonnet-20241022` | `openrouter.ai/api/v1/chat/completions` |
|
|
84
|
+
| Requesty | `requesty/sference/thinkingcap-qwen3.6-27b` | `router.requesty.ai/v1/chat/completions` |
|
|
85
|
+
| ZAI | `zai/glm-5.2` | `open.bigmodel.cn/api/paas/v4/chat/completions` |
|
|
86
|
+
| Groq | `groq/llama-3.3-70b-versatile` | `api.groq.com/openai/v1/chat/completions` |
|
|
87
|
+
| xAI | `xai/grok-4.5` | `api.x.ai/v1/chat/completions` |
|
|
88
|
+
| OpenAI | `openai/gpt-4.1` | `api.openai.com/v1/chat/completions` |
|
|
89
|
+
| Forge (self-hosted) | `forge/my-lora-v3` | `$FORGE_BASE_URL/chat/completions` |
|
|
90
|
+
| Nebius AI Studio | `nebius/meta-llama/Llama-3.1-8B-Instruct` | `api.studio.nebius.com/v1/chat/completions` (or `$NEBIUS_BASE_URL`) |
|
|
91
|
+
|
|
92
|
+
You can also force the provider with `?p=<provider>` and send a bare model id.
|
|
93
|
+
|
|
94
|
+
### Adding an OpenAI-compatible upstream provider
|
|
95
|
+
|
|
96
|
+
`forge` and `nebius` are both **configurable providers**: any OpenAI-compatible
|
|
97
|
+
upstream wired up from exactly two env vars, `<PROVIDER>_BASE_URL` (scheme,
|
|
98
|
+
host, port, path prefix) and `<PROVIDER>_API_KEY` (sent as `Authorization:
|
|
99
|
+
Bearer <key>`) — no code change needed to point either one at a different
|
|
100
|
+
host. They differ in one way: whether the base URL has a working default.
|
|
101
|
+
|
|
102
|
+
| Provider | `..._BASE_URL` | Default when unset | `..._API_KEY` |
|
|
103
|
+
| :--- | :--- | :--- | :--- |
|
|
104
|
+
| `forge` (self-hosted) | `FORGE_BASE_URL` | *(none — provider only exists once set)* | `FORGE_API_KEY` — **optional**; omitted entirely, no `Authorization` header sent, for a server with no auth (private network) |
|
|
105
|
+
| `nebius` (Nebius AI Studio) | `NEBIUS_BASE_URL` | `https://api.studio.nebius.com/v1` | `NEBIUS_API_KEY` — **required**, like every other provider (401 when missing) |
|
|
106
|
+
|
|
107
|
+
Both `http://` and `https://` are supported, with any host/port/path prefix —
|
|
108
|
+
unlike every fixed-hostname provider above, these two read their wire format
|
|
109
|
+
from the env var rather than a hardcoded `https://` + well-known host. An
|
|
110
|
+
unset `FORGE_BASE_URL` (no default) or a malformed override on either
|
|
111
|
+
provider makes `forge/...`/`nebius/...` requests fail with a clear 4xx, never
|
|
112
|
+
a crash.
|
|
113
|
+
|
|
114
|
+
Both work on all three surfaces (`/v1/messages`, `/v1/chat/completions`,
|
|
115
|
+
`/v1/responses`) with the same Anthropic↔OpenAI translation, streaming, and
|
|
116
|
+
tool-cap handling as every other OpenAI-compatible provider. `GET /v1/models`
|
|
117
|
+
additionally proxies `GET ${FORGE_BASE_URL}/models` and merges the results
|
|
118
|
+
into the default pack's listing (ids prefixed `forge/`) — forge-only, since
|
|
119
|
+
its LoRA adapters are registered on the server itself rather than known ahead
|
|
120
|
+
of time; nebius's catalog is the well-known set of ids you already pass in.
|
|
121
|
+
|
|
122
|
+
```sh
|
|
123
|
+
# forge — self-hosted vLLM, no auth
|
|
124
|
+
curl http://localhost:18090/v1/chat/completions \
|
|
125
|
+
-H 'Content-Type: application/json' \
|
|
126
|
+
-d '{"model":"forge/my-lora-v3","messages":[{"role":"user","content":"hi"}]}'
|
|
127
|
+
|
|
128
|
+
# nebius — hosted, requires NEBIUS_API_KEY
|
|
129
|
+
curl http://localhost:18090/v1/chat/completions \
|
|
130
|
+
-H 'Content-Type: application/json' \
|
|
131
|
+
-d '{"model":"nebius/meta-llama/Llama-3.1-8B-Instruct","messages":[{"role":"user","content":"hi"}]}'
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
### Named endpoints — N local/LAN model servers (`~/.agentproto/llm-endpoints.json`)
|
|
135
|
+
|
|
136
|
+
`forge` is one keyless self-hosted server. Real setups often have several —
|
|
137
|
+
Ollama, llama-server, vLLM — each on its own host/port, possibly on another
|
|
138
|
+
LAN machine. Named endpoints generalize `forge`'s mechanism to N of them,
|
|
139
|
+
configured from a JSON file instead of a pair of env vars per server:
|
|
140
|
+
|
|
141
|
+
```jsonc
|
|
142
|
+
// ~/.agentproto/llm-endpoints.json (path overridable via LLM_ENDPOINT_ENDPOINTS_FILE)
|
|
143
|
+
{
|
|
144
|
+
"endpoints": [
|
|
145
|
+
{
|
|
146
|
+
"id": "bonsai",
|
|
147
|
+
"kind": "openai",
|
|
148
|
+
"baseUrl": "http://192.168.1.20:8081/v1",
|
|
149
|
+
"apiKeyEnv": "BONSAI_API_KEY",
|
|
150
|
+
"defaultRequestFields": { "chat_template_kwargs": { "enable_thinking": false } },
|
|
151
|
+
"timeoutMs": { "firstTokenMs": 180000 }
|
|
152
|
+
},
|
|
153
|
+
{ "id": "ollama", "kind": "openai", "baseUrl": "http://192.168.1.20:11434/v1" },
|
|
154
|
+
{
|
|
155
|
+
"id": "lmstudio",
|
|
156
|
+
"kind": "openai",
|
|
157
|
+
"baseUrl": "http://127.0.0.1:1234/v1",
|
|
158
|
+
"defaultRequestFields": { "reasoning_effort": "none" }
|
|
159
|
+
}
|
|
160
|
+
]
|
|
161
|
+
}
|
|
162
|
+
```
|
|
163
|
+
|
|
164
|
+
The field names deliberately mirror `openagentik/router`'s `providers[]`
|
|
165
|
+
schema (`kind`, `baseUrl`, `apiKeyEnv`, `defaultRequestFields`, `timeoutMs`) so
|
|
166
|
+
a config is portable between the two. Each entry becomes a routable provider
|
|
167
|
+
`<id>/<model>` — `bonsai/bonsai-27b`, `ollama/qwen2.5-coder` — going through
|
|
168
|
+
the exact same dispatch path as `forge`/`nebius` above: all three surfaces,
|
|
169
|
+
Anthropic↔OpenAI translation, streaming, tool-cap handling, and `GET
|
|
170
|
+
/v1/models` merging (an unreachable endpoint is skipped with a warning, never
|
|
171
|
+
a 500 for the whole listing). `forge` itself keeps working unchanged — it's
|
|
172
|
+
the implicit endpoint that `FORGE_BASE_URL`/`FORGE_API_KEY` configure; a file
|
|
173
|
+
entry may not reuse the id `"forge"`.
|
|
174
|
+
|
|
175
|
+
- **`apiKeyEnv`** names an env var (never the key itself). Absent, or the env
|
|
176
|
+
var unset, means the endpoint is always keyless — no `Authorization` header
|
|
177
|
+
sent, same as `forge` with no `FORGE_API_KEY`.
|
|
178
|
+
- **`defaultRequestFields`** is any set of top-level OpenAI-compatible request
|
|
179
|
+
fields, merged UNDER the client's own request (a client-supplied top-level
|
|
180
|
+
key always wins; for an object-valued key present on both sides — e.g.
|
|
181
|
+
`chat_template_kwargs` — the merge goes one level deep, client sub-keys
|
|
182
|
+
winning). It cannot set `model`, `messages`, `stream`, `tools`, or `input`
|
|
183
|
+
— those would override routing/auth, not just default a parameter.
|
|
184
|
+
- `chat_template_kwargs` is the vLLM OpenAI-compatible extension `forge`
|
|
185
|
+
already documents above. Needed in practice: Qwen3.6-based models (e.g. a
|
|
186
|
+
Bonsai-served 27B) default to a "thinking" chat template and, on a tight
|
|
187
|
+
`max_tokens` budget, can spend the whole budget reasoning and return
|
|
188
|
+
empty content — `enable_thinking: false` avoids that unless the caller
|
|
189
|
+
explicitly opts back in.
|
|
190
|
+
- A plain top-level field works the same way for a server that doesn't
|
|
191
|
+
honour `chat_template_kwargs` or Anthropic's `thinking`/OpenAI's
|
|
192
|
+
`reasoning` params at all: LM Studio serving `prism-ml/bonsai-27b`
|
|
193
|
+
ignores both and only respects a top-level `reasoning_effort: "none"` to
|
|
194
|
+
cut reasoning to 0 tokens. On the Anthropic `/v1/messages` surface, an
|
|
195
|
+
explicit client `thinking: {type: "enabled"}` withholds a default
|
|
196
|
+
`reasoning_effort` rather than forcing it off.
|
|
197
|
+
- When the upstream *still* returns only `reasoning_content` (no visible
|
|
198
|
+
`content`) — a model-level defect, not a config one — set
|
|
199
|
+
`LLM_ENDPOINT_PASSTHROUGH_THINKING=1` to surface it as an Anthropic
|
|
200
|
+
`thinking` block instead of an empty message.
|
|
201
|
+
- **`GET /v1/endpoints`** (and `/endpoints`) reports live health per endpoint —
|
|
202
|
+
forge (if configured) plus every named endpoint, never nebius (a hosted
|
|
203
|
+
provider with a well-known catalog, not a local/LAN server): `{id, baseUrl
|
|
204
|
+
(no credential), reachable, models, latencyMs}`. Gated by the same
|
|
205
|
+
access-token check as `/v1/upstreams` (not the `/v1/models` public
|
|
206
|
+
exemption).
|
|
207
|
+
- A `packs.local.json` route may target a named endpoint id as its
|
|
208
|
+
`provider` — an id that resolves to neither a canonical upstream, `forge`,
|
|
209
|
+
`nebius`, nor a configured endpoint is a clear load-time error instead of a
|
|
210
|
+
confusing 400 the first time a client hits that code.
|
|
211
|
+
|
|
212
|
+
```sh
|
|
213
|
+
curl http://localhost:18090/v1/chat/completions \
|
|
214
|
+
-H 'Content-Type: application/json' \
|
|
215
|
+
-d '{"model":"bonsai/bonsai-27b","messages":[{"role":"user","content":"hi"}]}'
|
|
216
|
+
|
|
217
|
+
curl http://localhost:18090/v1/endpoints
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
### Anthropic Messages surface (`/v1/messages`)
|
|
221
|
+
|
|
222
|
+
The default public pack lists provider-transparent model IDs (the same values
|
|
223
|
+
you would send to the upstream provider). Real Claude model IDs are preserved
|
|
224
|
+
only when they route to an actual Anthropic target.
|
|
225
|
+
|
|
226
|
+
If you need the `claude` CLI or another Anthropic-only client to accept
|
|
227
|
+
non-Anthropic backends, either flip on the **Anthropic-style format** (below —
|
|
228
|
+
works with any pack) or create a **local compatibility pack** (further below).
|
|
229
|
+
Both surface `equivalentClaudeName` aliases that are matched **only** when
|
|
230
|
+
enabled.
|
|
231
|
+
|
|
232
|
+
### `coding` pack (curated OpenRouter coding models)
|
|
233
|
+
|
|
234
|
+
The committed `coding` pack is a small, production-only portfolio of coding
|
|
235
|
+
models, all routed through OpenRouter's native Anthropic-compatible endpoint:
|
|
236
|
+
|
|
237
|
+
| Code (transparent id) | Tier → family |
|
|
238
|
+
| :--- | :--- |
|
|
239
|
+
| `openai/gpt-5.5` | extra-high → fable |
|
|
240
|
+
| `anthropic/claude-opus-4.8` | high → opus |
|
|
241
|
+
| `deepseek/deepseek-v4-pro` | high → opus |
|
|
242
|
+
| `anthropic/claude-sonnet-5` | medium → sonnet |
|
|
243
|
+
| `z-ai/glm-5.2` | medium → sonnet |
|
|
244
|
+
| `minimax/minimax-m3` | small → haiku |
|
|
245
|
+
|
|
246
|
+
Select it like any pack (`X-Proxy-Pack: coding`, `?pack=coding`, or
|
|
247
|
+
`/v1/coding/messages`). The list is curated by hand against OpenRouter's live
|
|
248
|
+
Models API/rankings (`GET https://openrouter.ai/api/v1/models?supported_parameters=tools`);
|
|
249
|
+
availability and pricing drift, so re-check before relying on a route. Anything
|
|
250
|
+
outside the pack is still reachable via transparent `openrouter/vendor/model`
|
|
251
|
+
routing.
|
|
252
|
+
|
|
253
|
+
### Anthropic-style format (`?format=anthropic`)
|
|
254
|
+
|
|
255
|
+
Any pack can be relabeled on the fly so an Anthropic-only client gets
|
|
256
|
+
Claude-shaped model ids — **without impersonating a real Anthropic model**.
|
|
257
|
+
Send `X-Proxy-Format: anthropic` (or `?format=anthropic`) and each route's id
|
|
258
|
+
becomes an opaque, deterministic `claude-<family>-<sha>` value, where the family
|
|
259
|
+
comes from the route's tier (`extra-high→fable`, `high→opus`, `medium→sonnet`,
|
|
260
|
+
`small→haiku`) and the suffix is a sha of the upstream id. The real route stays
|
|
261
|
+
as the model's `display_name`.
|
|
262
|
+
|
|
263
|
+
Discover the current ids, then use one as the model:
|
|
264
|
+
|
|
265
|
+
```bash
|
|
266
|
+
# Discover (Anthropic-formatted model list)
|
|
267
|
+
curl -s -H 'anthropic-version: 2023-06-01' -H 'X-Proxy-Format: anthropic' \
|
|
268
|
+
http://localhost:18090/v1/coding/models
|
|
269
|
+
# → { "data": [ { "id": "claude-opus-5246108", "display_name": "anthropic/claude-opus-4.8", … }, … ] }
|
|
270
|
+
|
|
271
|
+
# Use it (Messages path resolves the id back to the real OpenRouter route)
|
|
272
|
+
env -u ANTHROPIC_API_KEY \
|
|
273
|
+
ANTHROPIC_BASE_URL="http://localhost:18090" \
|
|
274
|
+
ANTHROPIC_CUSTOM_HEADERS="X-Proxy-Pack: coding, X-Proxy-Format: anthropic" \
|
|
275
|
+
ANTHROPIC_AUTH_TOKEN="unused-the-proxy-holds-the-real-key" \
|
|
276
|
+
ANTHROPIC_MODEL="claude-opus-5246108" \
|
|
277
|
+
claude -p "…"
|
|
278
|
+
```
|
|
279
|
+
|
|
280
|
+
The ids are stable across restarts (they are derived from the upstream id, not
|
|
281
|
+
random), so a discovered id keeps working until the underlying route changes.
|
|
282
|
+
|
|
283
|
+
---
|
|
284
|
+
|
|
285
|
+
## Local compatibility packs
|
|
286
|
+
|
|
287
|
+
Create a `packs.local.json` file (gitignored) in the workspace root, the package
|
|
288
|
+
root, or next to `src/index.ts`:
|
|
289
|
+
|
|
290
|
+
```json
|
|
291
|
+
{
|
|
292
|
+
"packs": {
|
|
293
|
+
"local-claude": {
|
|
294
|
+
"id": "local-claude",
|
|
295
|
+
"label": "Local Claude compat",
|
|
296
|
+
"description": "Claude-shaped aliases for my preferred backends",
|
|
297
|
+
"models": {
|
|
298
|
+
"my-opus": {
|
|
299
|
+
"provider": "moonshot",
|
|
300
|
+
"model": "kimi-k2.7-code",
|
|
301
|
+
"equivalentClaudeName": "claude-opus-4-8"
|
|
302
|
+
}
|
|
303
|
+
}
|
|
304
|
+
}
|
|
305
|
+
}
|
|
306
|
+
}
|
|
307
|
+
```
|
|
308
|
+
|
|
309
|
+
Then select the pack via header (`X-Proxy-Pack: local-claude`), query param
|
|
310
|
+
(`?pack=local-claude`), or URL path (`/v1/local-claude/messages`). The alias
|
|
311
|
+
`claude-opus-4-8` will route to `moonshot/kimi-k2.7-code` **only** on the
|
|
312
|
+
Messages path and **only** when `local-claude` is active.
|
|
313
|
+
|
|
314
|
+
### Driving the `claude` CLI through a local pack
|
|
315
|
+
|
|
316
|
+
Use the **header**, not the URL path. The claude binary appends `/v1/messages`
|
|
317
|
+
to `ANTHROPIC_BASE_URL` itself, so a base of `…/v1/local-claude` becomes
|
|
318
|
+
`/v1/local-claude/v1/messages`, which matches no pack route — the request
|
|
319
|
+
silently falls back to the default pack and 400s with "Unable to resolve model".
|
|
320
|
+
Point the base at the proxy root and select the pack by header:
|
|
321
|
+
|
|
322
|
+
```bash
|
|
323
|
+
env -u ANTHROPIC_API_KEY \
|
|
324
|
+
ANTHROPIC_BASE_URL="http://localhost:18090" \
|
|
325
|
+
ANTHROPIC_CUSTOM_HEADERS="X-Proxy-Pack: local-claude" \
|
|
326
|
+
ANTHROPIC_AUTH_TOKEN="unused-the-proxy-holds-the-real-key" \
|
|
327
|
+
ANTHROPIC_MODEL="claude-opus-4-8" \
|
|
328
|
+
ANTHROPIC_SMALL_FAST_MODEL="claude-haiku-4-5" \
|
|
329
|
+
claude -p "…"
|
|
330
|
+
```
|
|
331
|
+
|
|
332
|
+
Pin `ANTHROPIC_SMALL_FAST_MODEL` too, or the harness's background calls request
|
|
333
|
+
a Claude tier the pack does not alias. Give reasoning models real `max_tokens`
|
|
334
|
+
headroom: a thinking model can spend a small budget entirely inside its thinking
|
|
335
|
+
block, and since those blocks are stripped (see below) the client then sees an
|
|
336
|
+
empty `content` with `stop_reason: max_tokens`.
|
|
337
|
+
|
|
338
|
+
---
|
|
339
|
+
|
|
340
|
+
## Per-request overrides (query string)
|
|
341
|
+
|
|
342
|
+
| Param | Effect |
|
|
343
|
+
| :--- | :--- |
|
|
344
|
+
| `?p=<provider>` | Force the provider (`moonshot`, `openrouter`, `zai`, `groq`, `xai`, `openai`) |
|
|
345
|
+
| `?m=<code>` | Force a pack code on the **Messages** path |
|
|
346
|
+
| `?format=anthropic` | Relabel the active pack to opaque `claude-<family>-<sha>` ids (also via `X-Proxy-Format: anthropic`) |
|
|
347
|
+
| `?tools=<names>` | Tool allow-list (e.g. `?tools=Bash,Read,Write`) — drop everything else |
|
|
348
|
+
| `?notools=1` | Strip **all** tools (+ `tool_choice`) → "lean" mode for strict-cap backends |
|
|
349
|
+
|
|
350
|
+
---
|
|
351
|
+
|
|
352
|
+
## OpenAI Responses API facade (Codex custom providers)
|
|
353
|
+
|
|
354
|
+
The proxy exposes a focused `POST /v1/responses` endpoint that implements the
|
|
355
|
+
OpenAI Responses API on top of the existing OpenAI-compatible chat/completions
|
|
356
|
+
providers. Codex custom providers can set `wire_api = "responses"` and point their
|
|
357
|
+
base URL at this proxy.
|
|
358
|
+
|
|
359
|
+
The facade is intentionally narrow: it supports the constructs that map cleanly to
|
|
360
|
+
a chat/completions request and rejects everything else up front. It routes through
|
|
361
|
+
the **transparent** `provider/model` surface, not through alias packs.
|
|
362
|
+
|
|
363
|
+
### Supported
|
|
364
|
+
|
|
365
|
+
- `model` — transparent `provider/model` reference (e.g. `openai/gpt-4.1`).
|
|
366
|
+
- `input` — a plain string or an array of `message` items (`input_text`) and
|
|
367
|
+
`function_call_output` items.
|
|
368
|
+
- `instructions` — injected as a leading `system` message.
|
|
369
|
+
- `tools` — only `type: "function"` tools are accepted.
|
|
370
|
+
- `tool_choice` — `"auto"`, `"none"`, `"required"`, or `{ type: "function", name }`.
|
|
371
|
+
- `stream` — when `true`, upstream SSE is re-emitted as Responses API SSE events
|
|
372
|
+
(`response.created`, `response.output_text.delta`, `response.completed`, …).
|
|
373
|
+
- Standard sampling params: `max_output_tokens` / `max_tokens`, `temperature`,
|
|
374
|
+
`top_p`, `parallel_tool_calls`.
|
|
375
|
+
- `reasoning.effort` — mapped to the upstream `reasoning_effort` parameter.
|
|
376
|
+
|
|
377
|
+
### Explicitly unsupported (returns 400)
|
|
378
|
+
|
|
379
|
+
- `previous_response_id` — the facade is stateless; each request is translated
|
|
380
|
+
independently.
|
|
381
|
+
- `text.format` / structured output.
|
|
382
|
+
- Non-`function` tool types (e.g. `web_search`).
|
|
383
|
+
- Image, audio, or other non-text content items.
|
|
384
|
+
|
|
385
|
+
---
|
|
386
|
+
|
|
387
|
+
## Batches
|
|
388
|
+
|
|
389
|
+
`POST /v1/messages/batches` (and `/v1/{pack}/messages/batches`) exposes the
|
|
390
|
+
Anthropic Message Batches API on top of the proxy's own routing — a client
|
|
391
|
+
pointed at the proxy can use the Anthropic SDK's batches surface unchanged
|
|
392
|
+
(`client.messages.batches.create/retrieve/results/cancel/list/delete`) with
|
|
393
|
+
`params.model` being **any** model the active pack routes. Batch is a delivery
|
|
394
|
+
mode, not a model: each item's `params` is the same Messages body the sync
|
|
395
|
+
`/v1/messages` route understands, so per-item routing, tool-trimming, and
|
|
396
|
+
translation are all reused, not reimplemented.
|
|
397
|
+
|
|
398
|
+
| Route | Behaviour |
|
|
399
|
+
| :--- | :--- |
|
|
400
|
+
| `POST /v1/messages/batches` | `{ requests: [{ custom_id, params }] }`. Validated (unique `custom_id`, no `stream`/`speed`/`fallbacks`/forced `tool_choice`, resolvable `model`) — 400 lists every offending item. Returns a `message_batch` object. |
|
|
401
|
+
| `GET /v1/messages/batches/{id}` | Aggregate status across sub-batches: `processing_status`, `request_counts`, `results_url` (once ended). |
|
|
402
|
+
| `GET /v1/messages/batches/{id}/results` | 404 until ended, then JSONL — one `{ custom_id, result }` line per item, cached after the first fetch so a repeat GET doesn't re-hit the provider. |
|
|
403
|
+
| `POST /v1/messages/batches/{id}/cancel` | Fans out to every sub-batch; a provider that doesn't support cancel (OpenRouter) is recorded, not fatal. |
|
|
404
|
+
| `GET /v1/messages/batches` | List, newest first (`?limit=`). |
|
|
405
|
+
| `DELETE /v1/messages/batches/{id}` | Only once ended (else `409`); best-effort forwarded to Anthropic for native sub-batches. |
|
|
406
|
+
|
|
407
|
+
### Native vs emulated
|
|
408
|
+
|
|
409
|
+
A batch whose items resolve to several providers is split into per-provider
|
|
410
|
+
sub-batches and re-aggregated by `custom_id`:
|
|
411
|
+
|
|
412
|
+
- **`anthropic`** and **`openrouter`** run on that provider's own async Batch
|
|
413
|
+
API — **50% of token price**, up to a 24h window.
|
|
414
|
+
- **Everything else** (`moonshot`, `requesty`, `zai`, `groq`, `xai`, `openai`)
|
|
415
|
+
has no batch API of its own, so items run through a local-queue emulation:
|
|
416
|
+
the proxy submits each item to its **own** `/v1/messages` over loopback, so
|
|
417
|
+
tool caps, thinking-strip, and empty-turn retry all still apply. **Full
|
|
418
|
+
price** — there is no provider-side discount to draw from.
|
|
419
|
+
|
|
420
|
+
### Credentials
|
|
421
|
+
|
|
422
|
+
Batches reuse the same per-provider credential resolution as `/v1/messages`.
|
|
423
|
+
One exception: if the resolved `anthropic` credential is a subscription OAuth
|
|
424
|
+
token (`sk-ant-oat…`) rather than an API key, batch creation fails closed with
|
|
425
|
+
a `401` — the Anthropic Batches API only accepts API keys, so the proxy never
|
|
426
|
+
attempts to send a subscription token to it.
|
|
427
|
+
|
|
428
|
+
### Config
|
|
429
|
+
|
|
430
|
+
| Env var | Effect |
|
|
431
|
+
| :--- | :--- |
|
|
432
|
+
| `LLM_ENDPOINT_STATE_DIR` | Where batch records + local-queue results are persisted. Default `~/.agentproto/llm-endpoint`. |
|
|
433
|
+
| `LLM_ENDPOINT_BATCH_CONCURRENCY` | Concurrent in-flight loopback requests per local-queue sub-batch. Default `4`. |
|
|
434
|
+
|
|
435
|
+
Batch records outlive the process — on restart, a native sub-batch simply
|
|
436
|
+
re-polls the provider by its stored id, and an unfinished local-queue
|
|
437
|
+
sub-batch resumes only the items still missing a result.
|
|
438
|
+
|
|
439
|
+
---
|
|
440
|
+
|
|
441
|
+
## Tool handling
|
|
442
|
+
|
|
443
|
+
### Automatic trimming
|
|
444
|
+
|
|
445
|
+
The `claude` CLI loads its full MCP config (`~/.claude` + `.mcp.json` + skills), which
|
|
446
|
+
often exceeds a provider's tool limit (e.g. **Groq: 128 max** →
|
|
447
|
+
`400 'tools': maximum number of items is 128`). The proxy truncates `payload.tools` to
|
|
448
|
+
the provider cap (`PROVIDER_MAX_TOOLS` in [`src/index.ts`](src/index.ts), currently
|
|
449
|
+
`groq: 128`) **before** reshaping tools for the provider. Providers without a cap
|
|
450
|
+
(moonshot, openrouter, zai, openai) are untouched. Use `?tools=` / `?notools=1` for finer control.
|
|
451
|
+
|
|
452
|
+
### Orphaned tool calls
|
|
453
|
+
|
|
454
|
+
When `tools` is truncated, conversation history can still contain `tool_use` blocks for
|
|
455
|
+
**undeclared** tools (e.g. the CLI's `Agent` sub-agent). Groq validates strictly and
|
|
456
|
+
rejects: `400 tool call validation failed: attempted to call tool 'X' which was not in
|
|
457
|
+
request.tools`. During Anthropic→OpenAI conversion the proxy converts those orphaned
|
|
458
|
+
`tool_use` (and their matching `tool_result`) into **plain text**
|
|
459
|
+
(`[Used tool X with args …]` / `[Tool result: …]`), preserving context without breaking
|
|
460
|
+
validation. Declared-tool `tool_use` blocks pass through normally as OpenAI `tool_calls`.
|
|
461
|
+
|
|
462
|
+
---
|
|
463
|
+
|
|
464
|
+
## Running it
|
|
465
|
+
|
|
466
|
+
No bespoke start script — it's a normal workspace package, driven by `pnpm` scripts:
|
|
467
|
+
|
|
468
|
+
```bash
|
|
469
|
+
# Dev: run the server with hot-reload (tsx watch)
|
|
470
|
+
pnpm --filter @agentproto/llm-endpoint serve
|
|
471
|
+
|
|
472
|
+
# Watch-rebuild the bundle (tsup --watch) — for consumers importing the package
|
|
473
|
+
pnpm --filter @agentproto/llm-endpoint dev
|
|
474
|
+
|
|
475
|
+
# Build the bundle (dist/index.mjs + dist/cli.mjs + types)
|
|
476
|
+
pnpm --filter @agentproto/llm-endpoint build
|
|
477
|
+
|
|
478
|
+
# Run the built server
|
|
479
|
+
pnpm --filter @agentproto/llm-endpoint start
|
|
480
|
+
|
|
481
|
+
# Type-check
|
|
482
|
+
pnpm --filter @agentproto/llm-endpoint check-types
|
|
483
|
+
```
|
|
484
|
+
|
|
485
|
+
Override the port with `LLM_ENDPOINT_PORT=18099 pnpm --filter @agentproto/llm-endpoint serve`.
|
|
486
|
+
|
|
487
|
+
The package also exports `start()` and the underlying `server` for embedding:
|
|
488
|
+
|
|
489
|
+
```ts
|
|
490
|
+
import { start } from '@agentproto/llm-endpoint'
|
|
491
|
+
start(18090)
|
|
492
|
+
```
|
|
493
|
+
|
|
494
|
+
### Live end-to-end suite
|
|
495
|
+
|
|
496
|
+
`src/__tests__/e2e.live.ts` hits a **running** proxy (`localhost:18090`) with **real**
|
|
497
|
+
provider keys, so it is deliberately kept off the vitest glob (it is not a unit test).
|
|
498
|
+
Start the server first, then:
|
|
499
|
+
|
|
500
|
+
```bash
|
|
501
|
+
pnpm --filter @agentproto/llm-endpoint test:e2e
|
|
502
|
+
```
|
|
503
|
+
|
|
504
|
+
---
|
|
505
|
+
|
|
506
|
+
## Securing a public deployment
|
|
507
|
+
|
|
508
|
+
The proxy holds **real upstream provider keys**, so any host that can reach it can
|
|
509
|
+
spend your credits. Never expose it publicly without an inbound gate.
|
|
510
|
+
|
|
511
|
+
### Inbound access gate
|
|
512
|
+
|
|
513
|
+
Set `LLM_ENDPOINT_ACCESS_TOKENS` to a comma-separated allow-list of secret tokens.
|
|
514
|
+
When it is set, every request must present a listed token as either
|
|
515
|
+
`Authorization: Bearer <token>` **or** an `X-Proxy-Access: <token>` header — anything
|
|
516
|
+
else gets `401`. When the variable is **unset the gate is open** (no inbound auth):
|
|
517
|
+
fine for `localhost`, unsafe for a public origin.
|
|
518
|
+
|
|
519
|
+
> **`x-api-key` is not accepted.** Some Anthropic-compatible clients default to
|
|
520
|
+
> sending the credential as `x-api-key`; the gate only reads `Authorization:
|
|
521
|
+
> Bearer` and `X-Proxy-Access`. Set the client's auth scheme to **Bearer**.
|
|
522
|
+
|
|
523
|
+
```bash
|
|
524
|
+
LLM_ENDPOINT_ACCESS_TOKENS="$(openssl rand -hex 24)" \
|
|
525
|
+
pnpm --filter @agentproto/llm-endpoint start
|
|
526
|
+
```
|
|
527
|
+
|
|
528
|
+
### Public model discovery (optional)
|
|
529
|
+
|
|
530
|
+
Clients that auto-discover models (e.g. Claude Desktop's launch-time model
|
|
531
|
+
fetch) probe `GET /v1/models` **without** a credential, so the access gate
|
|
532
|
+
`401`s them and the connection test fails. Set `LLM_ENDPOINT_PUBLIC_MODELS=1` to
|
|
533
|
+
exempt **only** the default model-list path (`/v1/models`, `/models`) from every
|
|
534
|
+
gate. Pack-scoped lists (`/v1/<pack>/models`) and all other paths stay gated, so
|
|
535
|
+
no pack config leaks. Unset (default) keeps discovery gated too.
|
|
536
|
+
|
|
537
|
+
### Edge / WAF token layer
|
|
538
|
+
|
|
539
|
+
`LLM_ENDPOINT_EDGE_TOKENS` is a second, independent allow-list checked via the
|
|
540
|
+
`X-Edge-Auth: <token>` header. It's meant to be enforced **at the edge** (a
|
|
541
|
+
Cloudflare WAF rule in front of the tunnel) so unauthenticated traffic never
|
|
542
|
+
reaches the origin — and it's also re-checked in-process as a fallback. Each
|
|
543
|
+
layer is independent; unset means off. Reusing the same secret as
|
|
544
|
+
`LLM_ENDPOINT_ACCESS_TOKENS` (via `Authorization: Bearer`) is fine too — one
|
|
545
|
+
secret, both the edge rule and the app gate.
|
|
546
|
+
|
|
547
|
+
### Generating the Cloudflare rule (`print-waf-rule`)
|
|
548
|
+
|
|
549
|
+
`llm-endpoint print-waf-rule` prints a Cloudflare custom-rule (wirefilter)
|
|
550
|
+
expression that **blocks** any request lacking a valid token, so the secret
|
|
551
|
+
lives in one place and the edge rule is generated, not hand-typed. It reads
|
|
552
|
+
`LLM_ENDPOINT_EDGE_TOKENS` (→ `X-Edge-Auth`) when set, else
|
|
553
|
+
`LLM_ENDPOINT_ACCESS_TOKENS` (→ `Authorization: Bearer`); `--host <h>` (or
|
|
554
|
+
`LLM_ENDPOINT_PUBLIC_HOST`) scopes the rule to one hostname. `OPTIONS` preflight
|
|
555
|
+
is always allowed.
|
|
556
|
+
|
|
557
|
+
```bash
|
|
558
|
+
LLM_ENDPOINT_ACCESS_TOKENS="$SECRET" llm-endpoint print-waf-rule --host llm.example.com
|
|
559
|
+
# → (http.host eq "llm.example.com" and http.request.method ne "OPTIONS"
|
|
560
|
+
# and not any(http.request.headers["authorization"][*] eq "Bearer $SECRET"))
|
|
561
|
+
```
|
|
562
|
+
|
|
563
|
+
Paste the output into a Cloudflare **Block** custom rule. If you also enabled
|
|
564
|
+
`LLM_ENDPOINT_PUBLIC_MODELS`, add a carve-out so discovery bypasses the edge as
|
|
565
|
+
well: `… and http.request.uri.path ne "/v1/models" and http.request.uri.path ne
|
|
566
|
+
"/models" and …`.
|
|
567
|
+
|
|
568
|
+
### Exposing the port
|
|
569
|
+
|
|
570
|
+
Any HTTP tunnel or reverse proxy that forwards to `http://localhost:18090` works
|
|
571
|
+
(Cloudflare Tunnel, ngrok, a VPS + nginx, …). Whichever you pick: keep the access
|
|
572
|
+
gate enabled, and consider an **edge control** as defense-in-depth (e.g. a Cloudflare
|
|
573
|
+
WAF rule or Cloudflare Access policy keyed on the same token) so unauthenticated
|
|
574
|
+
traffic is rejected before it ever reaches the origin.
|