@agentproto/llm-endpoint 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md ADDED
@@ -0,0 +1,574 @@
1
+ # @agentproto/llm-endpoint
2
+
3
+ A lightweight **multi-surface LLM proxy gateway**. It exposes three API
4
+ surfaces from a single port:
5
+
6
+ - `POST /v1/messages` — Anthropic Messages compatibility. Claude-shaped model
7
+ aliases are supported **only** through an explicit local compatibility pack.
8
+ - `POST /v1/chat/completions` — OpenAI Chat Completions compatibility with
9
+ transparent `provider/model` routing.
10
+ - `POST /v1/responses` — OpenAI Responses API facade for Codex custom providers.
11
+
12
+ Requests are fanned out to upstream providers — **Moonshot, OpenRouter, ZAI/Zhipu,
13
+ Groq, xAI, direct OpenAI, Nebius AI Studio, a self-hosted "forge" server for
14
+ fine-tunes, and any number of named local/LAN model servers** (Ollama,
15
+ llama-server, vLLM, …, configured from `~/.agentproto/llm-endpoints.json`) —
16
+ using provider-native model references. The proxy also handles
17
+ Anthropic↔OpenAI schema translation, per-provider tool caps, orphaned-tool-call
18
+ repair, and thinking-block stripping where needed.
19
+
20
+ - **Package:** `@agentproto/llm-endpoint`
21
+ - **Entry:** `src/cli.ts` → `start()` in `src/index.ts`
22
+ - **Default port:** `18090` (override with `LLM_ENDPOINT_PORT`, or `PORT`)
23
+
24
+ ---
25
+
26
+ ## Credentials this proxy handles
27
+
28
+ Read this before pointing the proxy at anything you care about. The package
29
+ ships **no credentials** — every key below is one you supply — but it is a
30
+ credential-forwarding component, so its handling rules are part of its
31
+ contract.
32
+
33
+ - **Per-provider API keys** resolve from one env var each
34
+ (`resolveSecretKeys`), or from `~/.agentproto/providers.json`. They are
35
+ injected at proxy boot and never logged.
36
+ - **Anthropic subscription OAuth tokens** (`sk-ant-oat…`) are supported as an
37
+ alternative to an API key on the `anthropic` upstream. When one is resolved,
38
+ the proxy sends `Authorization: Bearer <oat>` plus
39
+ `anthropic-beta: oauth-2025-04-20` and `anthropic-version: 2023-06-01`.
40
+ This path is **fail-closed**: an OAuth credential is *only* ever sent to the
41
+ `anthropic` upstream, and every other route `401`s rather than leaking it to
42
+ a third-party provider. Batch creation also refuses an OAuth token outright,
43
+ because the Anthropic Batches API only accepts API keys.
44
+ - A **stored** OAuth token works while it is valid. The proxy does not yet
45
+ drive a self-refreshing source (`source: "claude-code-oauth"`) to re-mint an
46
+ expired one.
47
+ - `KeychainStore` is **macOS-only**; elsewhere the proxy falls back to the
48
+ per-provider env key.
49
+
50
+ Fronting a personal Claude subscription through a shared gateway may not be
51
+ compatible with your Anthropic plan's terms — that is a question about **your
52
+ deployment**, not about this package. If you expose the proxy beyond localhost,
53
+ read [Securing a public deployment](#securing-a-public-deployment) first: the
54
+ access gate is **off by default**.
55
+
56
+ ---
57
+
58
+ ## Quick start (client config)
59
+
60
+ Point any Anthropic- or OpenAI-compatible client at the proxy:
61
+
62
+ | Setting | Value |
63
+ | :--- | :--- |
64
+ | **Base URL** | `http://localhost:18090/v1` (local), or your own public origin + `/v1` |
65
+ | **API key** | The proxy injects the *real* upstream key server-side, so the client key is never your provider key. If the [inbound access gate](#securing-a-public-deployment) is enabled, send one of its tokens as the bearer; if it is not, any non-empty value passes. |
66
+
67
+ The client asks for a model by its **provider-transparent reference**
68
+ (`provider/model`, e.g. `moonshot/kimi-k2.7-code`, `openai/gpt-4.1`); the proxy
69
+ routes it to the right upstream endpoint.
70
+
71
+ ---
72
+
73
+ ## Model catalog
74
+
75
+ ### Transparent routing (`provider/model`)
76
+
77
+ On the OpenAI surfaces (`/v1/chat/completions` and `/v1/responses`) the model
78
+ field is parsed as `provider/model`:
79
+
80
+ | Provider | Example reference | Upstream endpoint |
81
+ | :--- | :--- | :--- |
82
+ | Moonshot | `moonshot/kimi-k2.7-code` | `api.moonshot.ai/v1/chat/completions` |
83
+ | OpenRouter | `openrouter/anthropic/claude-3-5-sonnet-20241022` | `openrouter.ai/api/v1/chat/completions` |
84
+ | Requesty | `requesty/sference/thinkingcap-qwen3.6-27b` | `router.requesty.ai/v1/chat/completions` |
85
+ | ZAI | `zai/glm-5.2` | `open.bigmodel.cn/api/paas/v4/chat/completions` |
86
+ | Groq | `groq/llama-3.3-70b-versatile` | `api.groq.com/openai/v1/chat/completions` |
87
+ | xAI | `xai/grok-4.5` | `api.x.ai/v1/chat/completions` |
88
+ | OpenAI | `openai/gpt-4.1` | `api.openai.com/v1/chat/completions` |
89
+ | Forge (self-hosted) | `forge/my-lora-v3` | `$FORGE_BASE_URL/chat/completions` |
90
+ | Nebius AI Studio | `nebius/meta-llama/Llama-3.1-8B-Instruct` | `api.studio.nebius.com/v1/chat/completions` (or `$NEBIUS_BASE_URL`) |
91
+
92
+ You can also force the provider with `?p=<provider>` and send a bare model id.
93
+
94
+ ### Adding an OpenAI-compatible upstream provider
95
+
96
+ `forge` and `nebius` are both **configurable providers**: any OpenAI-compatible
97
+ upstream wired up from exactly two env vars, `<PROVIDER>_BASE_URL` (scheme,
98
+ host, port, path prefix) and `<PROVIDER>_API_KEY` (sent as `Authorization:
99
+ Bearer <key>`) — no code change needed to point either one at a different
100
+ host. They differ in one way: whether the base URL has a working default.
101
+
102
+ | Provider | `..._BASE_URL` | Default when unset | `..._API_KEY` |
103
+ | :--- | :--- | :--- | :--- |
104
+ | `forge` (self-hosted) | `FORGE_BASE_URL` | *(none — provider only exists once set)* | `FORGE_API_KEY` — **optional**; omitted entirely, no `Authorization` header sent, for a server with no auth (private network) |
105
+ | `nebius` (Nebius AI Studio) | `NEBIUS_BASE_URL` | `https://api.studio.nebius.com/v1` | `NEBIUS_API_KEY` — **required**, like every other provider (401 when missing) |
106
+
107
+ Both `http://` and `https://` are supported, with any host/port/path prefix —
108
+ unlike every fixed-hostname provider above, these two read their wire format
109
+ from the env var rather than a hardcoded `https://` + well-known host. An
110
+ unset `FORGE_BASE_URL` (no default) or a malformed override on either
111
+ provider makes `forge/...`/`nebius/...` requests fail with a clear 4xx, never
112
+ a crash.
113
+
114
+ Both work on all three surfaces (`/v1/messages`, `/v1/chat/completions`,
115
+ `/v1/responses`) with the same Anthropic↔OpenAI translation, streaming, and
116
+ tool-cap handling as every other OpenAI-compatible provider. `GET /v1/models`
117
+ additionally proxies `GET ${FORGE_BASE_URL}/models` and merges the results
118
+ into the default pack's listing (ids prefixed `forge/`) — forge-only, since
119
+ its LoRA adapters are registered on the server itself rather than known ahead
120
+ of time; nebius's catalog is the well-known set of ids you already pass in.
121
+
122
+ ```sh
123
+ # forge — self-hosted vLLM, no auth
124
+ curl http://localhost:18090/v1/chat/completions \
125
+ -H 'Content-Type: application/json' \
126
+ -d '{"model":"forge/my-lora-v3","messages":[{"role":"user","content":"hi"}]}'
127
+
128
+ # nebius — hosted, requires NEBIUS_API_KEY
129
+ curl http://localhost:18090/v1/chat/completions \
130
+ -H 'Content-Type: application/json' \
131
+ -d '{"model":"nebius/meta-llama/Llama-3.1-8B-Instruct","messages":[{"role":"user","content":"hi"}]}'
132
+ ```
133
+
134
+ ### Named endpoints — N local/LAN model servers (`~/.agentproto/llm-endpoints.json`)
135
+
136
+ `forge` is one keyless self-hosted server. Real setups often have several —
137
+ Ollama, llama-server, vLLM — each on its own host/port, possibly on another
138
+ LAN machine. Named endpoints generalize `forge`'s mechanism to N of them,
139
+ configured from a JSON file instead of a pair of env vars per server:
140
+
141
+ ```jsonc
142
+ // ~/.agentproto/llm-endpoints.json (path overridable via LLM_ENDPOINT_ENDPOINTS_FILE)
143
+ {
144
+ "endpoints": [
145
+ {
146
+ "id": "bonsai",
147
+ "kind": "openai",
148
+ "baseUrl": "http://192.168.1.20:8081/v1",
149
+ "apiKeyEnv": "BONSAI_API_KEY",
150
+ "defaultRequestFields": { "chat_template_kwargs": { "enable_thinking": false } },
151
+ "timeoutMs": { "firstTokenMs": 180000 }
152
+ },
153
+ { "id": "ollama", "kind": "openai", "baseUrl": "http://192.168.1.20:11434/v1" },
154
+ {
155
+ "id": "lmstudio",
156
+ "kind": "openai",
157
+ "baseUrl": "http://127.0.0.1:1234/v1",
158
+ "defaultRequestFields": { "reasoning_effort": "none" }
159
+ }
160
+ ]
161
+ }
162
+ ```
163
+
164
+ The field names deliberately mirror `openagentik/router`'s `providers[]`
165
+ schema (`kind`, `baseUrl`, `apiKeyEnv`, `defaultRequestFields`, `timeoutMs`) so
166
+ a config is portable between the two. Each entry becomes a routable provider
167
+ `<id>/<model>` — `bonsai/bonsai-27b`, `ollama/qwen2.5-coder` — going through
168
+ the exact same dispatch path as `forge`/`nebius` above: all three surfaces,
169
+ Anthropic↔OpenAI translation, streaming, tool-cap handling, and `GET
170
+ /v1/models` merging (an unreachable endpoint is skipped with a warning, never
171
+ a 500 for the whole listing). `forge` itself keeps working unchanged — it's
172
+ the implicit endpoint that `FORGE_BASE_URL`/`FORGE_API_KEY` configure; a file
173
+ entry may not reuse the id `"forge"`.
174
+
175
+ - **`apiKeyEnv`** names an env var (never the key itself). Absent, or the env
176
+ var unset, means the endpoint is always keyless — no `Authorization` header
177
+ sent, same as `forge` with no `FORGE_API_KEY`.
178
+ - **`defaultRequestFields`** is any set of top-level OpenAI-compatible request
179
+ fields, merged UNDER the client's own request (a client-supplied top-level
180
+ key always wins; for an object-valued key present on both sides — e.g.
181
+ `chat_template_kwargs` — the merge goes one level deep, client sub-keys
182
+ winning). It cannot set `model`, `messages`, `stream`, `tools`, or `input`
183
+ — those would override routing/auth, not just default a parameter.
184
+ - `chat_template_kwargs` is the vLLM OpenAI-compatible extension `forge`
185
+ already documents above. Needed in practice: Qwen3.6-based models (e.g. a
186
+ Bonsai-served 27B) default to a "thinking" chat template and, on a tight
187
+ `max_tokens` budget, can spend the whole budget reasoning and return
188
+ empty content — `enable_thinking: false` avoids that unless the caller
189
+ explicitly opts back in.
190
+ - A plain top-level field works the same way for a server that doesn't
191
+ honour `chat_template_kwargs` or Anthropic's `thinking`/OpenAI's
192
+ `reasoning` params at all: LM Studio serving `prism-ml/bonsai-27b`
193
+ ignores both and only respects a top-level `reasoning_effort: "none"` to
194
+ cut reasoning to 0 tokens. On the Anthropic `/v1/messages` surface, an
195
+ explicit client `thinking: {type: "enabled"}` withholds a default
196
+ `reasoning_effort` rather than forcing it off.
197
+ - When the upstream *still* returns only `reasoning_content` (no visible
198
+ `content`) — a model-level defect, not a config one — set
199
+ `LLM_ENDPOINT_PASSTHROUGH_THINKING=1` to surface it as an Anthropic
200
+ `thinking` block instead of an empty message.
201
+ - **`GET /v1/endpoints`** (and `/endpoints`) reports live health per endpoint —
202
+ forge (if configured) plus every named endpoint, never nebius (a hosted
203
+ provider with a well-known catalog, not a local/LAN server): `{id, baseUrl
204
+ (no credential), reachable, models, latencyMs}`. Gated by the same
205
+ access-token check as `/v1/upstreams` (not the `/v1/models` public
206
+ exemption).
207
+ - A `packs.local.json` route may target a named endpoint id as its
208
+ `provider` — an id that resolves to neither a canonical upstream, `forge`,
209
+ `nebius`, nor a configured endpoint is a clear load-time error instead of a
210
+ confusing 400 the first time a client hits that code.
211
+
212
+ ```sh
213
+ curl http://localhost:18090/v1/chat/completions \
214
+ -H 'Content-Type: application/json' \
215
+ -d '{"model":"bonsai/bonsai-27b","messages":[{"role":"user","content":"hi"}]}'
216
+
217
+ curl http://localhost:18090/v1/endpoints
218
+ ```
219
+
220
+ ### Anthropic Messages surface (`/v1/messages`)
221
+
222
+ The default public pack lists provider-transparent model IDs (the same values
223
+ you would send to the upstream provider). Real Claude model IDs are preserved
224
+ only when they route to an actual Anthropic target.
225
+
226
+ If you need the `claude` CLI or another Anthropic-only client to accept
227
+ non-Anthropic backends, either flip on the **Anthropic-style format** (below —
228
+ works with any pack) or create a **local compatibility pack** (further below).
229
+ Both surface `equivalentClaudeName` aliases that are matched **only** when
230
+ enabled.
231
+
232
+ ### `coding` pack (curated OpenRouter coding models)
233
+
234
+ The committed `coding` pack is a small, production-only portfolio of coding
235
+ models, all routed through OpenRouter's native Anthropic-compatible endpoint:
236
+
237
+ | Code (transparent id) | Tier → family |
238
+ | :--- | :--- |
239
+ | `openai/gpt-5.5` | extra-high → fable |
240
+ | `anthropic/claude-opus-4.8` | high → opus |
241
+ | `deepseek/deepseek-v4-pro` | high → opus |
242
+ | `anthropic/claude-sonnet-5` | medium → sonnet |
243
+ | `z-ai/glm-5.2` | medium → sonnet |
244
+ | `minimax/minimax-m3` | small → haiku |
245
+
246
+ Select it like any pack (`X-Proxy-Pack: coding`, `?pack=coding`, or
247
+ `/v1/coding/messages`). The list is curated by hand against OpenRouter's live
248
+ Models API/rankings (`GET https://openrouter.ai/api/v1/models?supported_parameters=tools`);
249
+ availability and pricing drift, so re-check before relying on a route. Anything
250
+ outside the pack is still reachable via transparent `openrouter/vendor/model`
251
+ routing.
252
+
253
+ ### Anthropic-style format (`?format=anthropic`)
254
+
255
+ Any pack can be relabeled on the fly so an Anthropic-only client gets
256
+ Claude-shaped model ids — **without impersonating a real Anthropic model**.
257
+ Send `X-Proxy-Format: anthropic` (or `?format=anthropic`) and each route's id
258
+ becomes an opaque, deterministic `claude-<family>-<sha>` value, where the family
259
+ comes from the route's tier (`extra-high→fable`, `high→opus`, `medium→sonnet`,
260
+ `small→haiku`) and the suffix is a sha of the upstream id. The real route stays
261
+ as the model's `display_name`.
262
+
263
+ Discover the current ids, then use one as the model:
264
+
265
+ ```bash
266
+ # Discover (Anthropic-formatted model list)
267
+ curl -s -H 'anthropic-version: 2023-06-01' -H 'X-Proxy-Format: anthropic' \
268
+ http://localhost:18090/v1/coding/models
269
+ # → { "data": [ { "id": "claude-opus-5246108", "display_name": "anthropic/claude-opus-4.8", … }, … ] }
270
+
271
+ # Use it (Messages path resolves the id back to the real OpenRouter route)
272
+ env -u ANTHROPIC_API_KEY \
273
+ ANTHROPIC_BASE_URL="http://localhost:18090" \
274
+ ANTHROPIC_CUSTOM_HEADERS="X-Proxy-Pack: coding, X-Proxy-Format: anthropic" \
275
+ ANTHROPIC_AUTH_TOKEN="unused-the-proxy-holds-the-real-key" \
276
+ ANTHROPIC_MODEL="claude-opus-5246108" \
277
+ claude -p "…"
278
+ ```
279
+
280
+ The ids are stable across restarts (they are derived from the upstream id, not
281
+ random), so a discovered id keeps working until the underlying route changes.
282
+
283
+ ---
284
+
285
+ ## Local compatibility packs
286
+
287
+ Create a `packs.local.json` file (gitignored) in the workspace root, the package
288
+ root, or next to `src/index.ts`:
289
+
290
+ ```json
291
+ {
292
+ "packs": {
293
+ "local-claude": {
294
+ "id": "local-claude",
295
+ "label": "Local Claude compat",
296
+ "description": "Claude-shaped aliases for my preferred backends",
297
+ "models": {
298
+ "my-opus": {
299
+ "provider": "moonshot",
300
+ "model": "kimi-k2.7-code",
301
+ "equivalentClaudeName": "claude-opus-4-8"
302
+ }
303
+ }
304
+ }
305
+ }
306
+ }
307
+ ```
308
+
309
+ Then select the pack via header (`X-Proxy-Pack: local-claude`), query param
310
+ (`?pack=local-claude`), or URL path (`/v1/local-claude/messages`). The alias
311
+ `claude-opus-4-8` will route to `moonshot/kimi-k2.7-code` **only** on the
312
+ Messages path and **only** when `local-claude` is active.
313
+
314
+ ### Driving the `claude` CLI through a local pack
315
+
316
+ Use the **header**, not the URL path. The claude binary appends `/v1/messages`
317
+ to `ANTHROPIC_BASE_URL` itself, so a base of `…/v1/local-claude` becomes
318
+ `/v1/local-claude/v1/messages`, which matches no pack route — the request
319
+ silently falls back to the default pack and 400s with "Unable to resolve model".
320
+ Point the base at the proxy root and select the pack by header:
321
+
322
+ ```bash
323
+ env -u ANTHROPIC_API_KEY \
324
+ ANTHROPIC_BASE_URL="http://localhost:18090" \
325
+ ANTHROPIC_CUSTOM_HEADERS="X-Proxy-Pack: local-claude" \
326
+ ANTHROPIC_AUTH_TOKEN="unused-the-proxy-holds-the-real-key" \
327
+ ANTHROPIC_MODEL="claude-opus-4-8" \
328
+ ANTHROPIC_SMALL_FAST_MODEL="claude-haiku-4-5" \
329
+ claude -p "…"
330
+ ```
331
+
332
+ Pin `ANTHROPIC_SMALL_FAST_MODEL` too, or the harness's background calls request
333
+ a Claude tier the pack does not alias. Give reasoning models real `max_tokens`
334
+ headroom: a thinking model can spend a small budget entirely inside its thinking
335
+ block, and since those blocks are stripped (see below) the client then sees an
336
+ empty `content` with `stop_reason: max_tokens`.
337
+
338
+ ---
339
+
340
+ ## Per-request overrides (query string)
341
+
342
+ | Param | Effect |
343
+ | :--- | :--- |
344
+ | `?p=<provider>` | Force the provider (`moonshot`, `openrouter`, `zai`, `groq`, `xai`, `openai`) |
345
+ | `?m=<code>` | Force a pack code on the **Messages** path |
346
+ | `?format=anthropic` | Relabel the active pack to opaque `claude-<family>-<sha>` ids (also via `X-Proxy-Format: anthropic`) |
347
+ | `?tools=<names>` | Tool allow-list (e.g. `?tools=Bash,Read,Write`) — drop everything else |
348
+ | `?notools=1` | Strip **all** tools (+ `tool_choice`) → "lean" mode for strict-cap backends |
349
+
350
+ ---
351
+
352
+ ## OpenAI Responses API facade (Codex custom providers)
353
+
354
+ The proxy exposes a focused `POST /v1/responses` endpoint that implements the
355
+ OpenAI Responses API on top of the existing OpenAI-compatible chat/completions
356
+ providers. Codex custom providers can set `wire_api = "responses"` and point their
357
+ base URL at this proxy.
358
+
359
+ The facade is intentionally narrow: it supports the constructs that map cleanly to
360
+ a chat/completions request and rejects everything else up front. It routes through
361
+ the **transparent** `provider/model` surface, not through alias packs.
362
+
363
+ ### Supported
364
+
365
+ - `model` — transparent `provider/model` reference (e.g. `openai/gpt-4.1`).
366
+ - `input` — a plain string or an array of `message` items (`input_text`) and
367
+ `function_call_output` items.
368
+ - `instructions` — injected as a leading `system` message.
369
+ - `tools` — only `type: "function"` tools are accepted.
370
+ - `tool_choice` — `"auto"`, `"none"`, `"required"`, or `{ type: "function", name }`.
371
+ - `stream` — when `true`, upstream SSE is re-emitted as Responses API SSE events
372
+ (`response.created`, `response.output_text.delta`, `response.completed`, …).
373
+ - Standard sampling params: `max_output_tokens` / `max_tokens`, `temperature`,
374
+ `top_p`, `parallel_tool_calls`.
375
+ - `reasoning.effort` — mapped to the upstream `reasoning_effort` parameter.
376
+
377
+ ### Explicitly unsupported (returns 400)
378
+
379
+ - `previous_response_id` — the facade is stateless; each request is translated
380
+ independently.
381
+ - `text.format` / structured output.
382
+ - Non-`function` tool types (e.g. `web_search`).
383
+ - Image, audio, or other non-text content items.
384
+
385
+ ---
386
+
387
+ ## Batches
388
+
389
+ `POST /v1/messages/batches` (and `/v1/{pack}/messages/batches`) exposes the
390
+ Anthropic Message Batches API on top of the proxy's own routing — a client
391
+ pointed at the proxy can use the Anthropic SDK's batches surface unchanged
392
+ (`client.messages.batches.create/retrieve/results/cancel/list/delete`) with
393
+ `params.model` being **any** model the active pack routes. Batch is a delivery
394
+ mode, not a model: each item's `params` is the same Messages body the sync
395
+ `/v1/messages` route understands, so per-item routing, tool-trimming, and
396
+ translation are all reused, not reimplemented.
397
+
398
+ | Route | Behaviour |
399
+ | :--- | :--- |
400
+ | `POST /v1/messages/batches` | `{ requests: [{ custom_id, params }] }`. Validated (unique `custom_id`, no `stream`/`speed`/`fallbacks`/forced `tool_choice`, resolvable `model`) — 400 lists every offending item. Returns a `message_batch` object. |
401
+ | `GET /v1/messages/batches/{id}` | Aggregate status across sub-batches: `processing_status`, `request_counts`, `results_url` (once ended). |
402
+ | `GET /v1/messages/batches/{id}/results` | 404 until ended, then JSONL — one `{ custom_id, result }` line per item, cached after the first fetch so a repeat GET doesn't re-hit the provider. |
403
+ | `POST /v1/messages/batches/{id}/cancel` | Fans out to every sub-batch; a provider that doesn't support cancel (OpenRouter) is recorded, not fatal. |
404
+ | `GET /v1/messages/batches` | List, newest first (`?limit=`). |
405
+ | `DELETE /v1/messages/batches/{id}` | Only once ended (else `409`); best-effort forwarded to Anthropic for native sub-batches. |
406
+
407
+ ### Native vs emulated
408
+
409
+ A batch whose items resolve to several providers is split into per-provider
410
+ sub-batches and re-aggregated by `custom_id`:
411
+
412
+ - **`anthropic`** and **`openrouter`** run on that provider's own async Batch
413
+ API — **50% of token price**, up to a 24h window.
414
+ - **Everything else** (`moonshot`, `requesty`, `zai`, `groq`, `xai`, `openai`)
415
+ has no batch API of its own, so items run through a local-queue emulation:
416
+ the proxy submits each item to its **own** `/v1/messages` over loopback, so
417
+ tool caps, thinking-strip, and empty-turn retry all still apply. **Full
418
+ price** — there is no provider-side discount to draw from.
419
+
420
+ ### Credentials
421
+
422
+ Batches reuse the same per-provider credential resolution as `/v1/messages`.
423
+ One exception: if the resolved `anthropic` credential is a subscription OAuth
424
+ token (`sk-ant-oat…`) rather than an API key, batch creation fails closed with
425
+ a `401` — the Anthropic Batches API only accepts API keys, so the proxy never
426
+ attempts to send a subscription token to it.
427
+
428
+ ### Config
429
+
430
+ | Env var | Effect |
431
+ | :--- | :--- |
432
+ | `LLM_ENDPOINT_STATE_DIR` | Where batch records + local-queue results are persisted. Default `~/.agentproto/llm-endpoint`. |
433
+ | `LLM_ENDPOINT_BATCH_CONCURRENCY` | Concurrent in-flight loopback requests per local-queue sub-batch. Default `4`. |
434
+
435
+ Batch records outlive the process — on restart, a native sub-batch simply
436
+ re-polls the provider by its stored id, and an unfinished local-queue
437
+ sub-batch resumes only the items still missing a result.
438
+
439
+ ---
440
+
441
+ ## Tool handling
442
+
443
+ ### Automatic trimming
444
+
445
+ The `claude` CLI loads its full MCP config (`~/.claude` + `.mcp.json` + skills), which
446
+ often exceeds a provider's tool limit (e.g. **Groq: 128 max** →
447
+ `400 'tools': maximum number of items is 128`). The proxy truncates `payload.tools` to
448
+ the provider cap (`PROVIDER_MAX_TOOLS` in [`src/index.ts`](src/index.ts), currently
449
+ `groq: 128`) **before** reshaping tools for the provider. Providers without a cap
450
+ (moonshot, openrouter, zai, openai) are untouched. Use `?tools=` / `?notools=1` for finer control.
451
+
452
+ ### Orphaned tool calls
453
+
454
+ When `tools` is truncated, conversation history can still contain `tool_use` blocks for
455
+ **undeclared** tools (e.g. the CLI's `Agent` sub-agent). Groq validates strictly and
456
+ rejects: `400 tool call validation failed: attempted to call tool 'X' which was not in
457
+ request.tools`. During Anthropic→OpenAI conversion the proxy converts those orphaned
458
+ `tool_use` (and their matching `tool_result`) into **plain text**
459
+ (`[Used tool X with args …]` / `[Tool result: …]`), preserving context without breaking
460
+ validation. Declared-tool `tool_use` blocks pass through normally as OpenAI `tool_calls`.
461
+
462
+ ---
463
+
464
+ ## Running it
465
+
466
+ No bespoke start script — it's a normal workspace package, driven by `pnpm` scripts:
467
+
468
+ ```bash
469
+ # Dev: run the server with hot-reload (tsx watch)
470
+ pnpm --filter @agentproto/llm-endpoint serve
471
+
472
+ # Watch-rebuild the bundle (tsup --watch) — for consumers importing the package
473
+ pnpm --filter @agentproto/llm-endpoint dev
474
+
475
+ # Build the bundle (dist/index.mjs + dist/cli.mjs + types)
476
+ pnpm --filter @agentproto/llm-endpoint build
477
+
478
+ # Run the built server
479
+ pnpm --filter @agentproto/llm-endpoint start
480
+
481
+ # Type-check
482
+ pnpm --filter @agentproto/llm-endpoint check-types
483
+ ```
484
+
485
+ Override the port with `LLM_ENDPOINT_PORT=18099 pnpm --filter @agentproto/llm-endpoint serve`.
486
+
487
+ The package also exports `start()` and the underlying `server` for embedding:
488
+
489
+ ```ts
490
+ import { start } from '@agentproto/llm-endpoint'
491
+ start(18090)
492
+ ```
493
+
494
+ ### Live end-to-end suite
495
+
496
+ `src/__tests__/e2e.live.ts` hits a **running** proxy (`localhost:18090`) with **real**
497
+ provider keys, so it is deliberately kept off the vitest glob (it is not a unit test).
498
+ Start the server first, then:
499
+
500
+ ```bash
501
+ pnpm --filter @agentproto/llm-endpoint test:e2e
502
+ ```
503
+
504
+ ---
505
+
506
+ ## Securing a public deployment
507
+
508
+ The proxy holds **real upstream provider keys**, so any host that can reach it can
509
+ spend your credits. Never expose it publicly without an inbound gate.
510
+
511
+ ### Inbound access gate
512
+
513
+ Set `LLM_ENDPOINT_ACCESS_TOKENS` to a comma-separated allow-list of secret tokens.
514
+ When it is set, every request must present a listed token as either
515
+ `Authorization: Bearer <token>` **or** an `X-Proxy-Access: <token>` header — anything
516
+ else gets `401`. When the variable is **unset the gate is open** (no inbound auth):
517
+ fine for `localhost`, unsafe for a public origin.
518
+
519
+ > **`x-api-key` is not accepted.** Some Anthropic-compatible clients default to
520
+ > sending the credential as `x-api-key`; the gate only reads `Authorization:
521
+ > Bearer` and `X-Proxy-Access`. Set the client's auth scheme to **Bearer**.
522
+
523
+ ```bash
524
+ LLM_ENDPOINT_ACCESS_TOKENS="$(openssl rand -hex 24)" \
525
+ pnpm --filter @agentproto/llm-endpoint start
526
+ ```
527
+
528
+ ### Public model discovery (optional)
529
+
530
+ Clients that auto-discover models (e.g. Claude Desktop's launch-time model
531
+ fetch) probe `GET /v1/models` **without** a credential, so the access gate
532
+ `401`s them and the connection test fails. Set `LLM_ENDPOINT_PUBLIC_MODELS=1` to
533
+ exempt **only** the default model-list path (`/v1/models`, `/models`) from every
534
+ gate. Pack-scoped lists (`/v1/<pack>/models`) and all other paths stay gated, so
535
+ no pack config leaks. Unset (default) keeps discovery gated too.
536
+
537
+ ### Edge / WAF token layer
538
+
539
+ `LLM_ENDPOINT_EDGE_TOKENS` is a second, independent allow-list checked via the
540
+ `X-Edge-Auth: <token>` header. It's meant to be enforced **at the edge** (a
541
+ Cloudflare WAF rule in front of the tunnel) so unauthenticated traffic never
542
+ reaches the origin — and it's also re-checked in-process as a fallback. Each
543
+ layer is independent; unset means off. Reusing the same secret as
544
+ `LLM_ENDPOINT_ACCESS_TOKENS` (via `Authorization: Bearer`) is fine too — one
545
+ secret, both the edge rule and the app gate.
546
+
547
+ ### Generating the Cloudflare rule (`print-waf-rule`)
548
+
549
+ `llm-endpoint print-waf-rule` prints a Cloudflare custom-rule (wirefilter)
550
+ expression that **blocks** any request lacking a valid token, so the secret
551
+ lives in one place and the edge rule is generated, not hand-typed. It reads
552
+ `LLM_ENDPOINT_EDGE_TOKENS` (→ `X-Edge-Auth`) when set, else
553
+ `LLM_ENDPOINT_ACCESS_TOKENS` (→ `Authorization: Bearer`); `--host <h>` (or
554
+ `LLM_ENDPOINT_PUBLIC_HOST`) scopes the rule to one hostname. `OPTIONS` preflight
555
+ is always allowed.
556
+
557
+ ```bash
558
+ LLM_ENDPOINT_ACCESS_TOKENS="$SECRET" llm-endpoint print-waf-rule --host llm.example.com
559
+ # → (http.host eq "llm.example.com" and http.request.method ne "OPTIONS"
560
+ # and not any(http.request.headers["authorization"][*] eq "Bearer $SECRET"))
561
+ ```
562
+
563
+ Paste the output into a Cloudflare **Block** custom rule. If you also enabled
564
+ `LLM_ENDPOINT_PUBLIC_MODELS`, add a carve-out so discovery bypasses the edge as
565
+ well: `… and http.request.uri.path ne "/v1/models" and http.request.uri.path ne
566
+ "/models" and …`.
567
+
568
+ ### Exposing the port
569
+
570
+ Any HTTP tunnel or reverse proxy that forwards to `http://localhost:18090` works
571
+ (Cloudflare Tunnel, ngrok, a VPS + nginx, …). Whichever you pick: keep the access
572
+ gate enabled, and consider an **edge control** as defense-in-depth (e.g. a Cloudflare
573
+ WAF rule or Cloudflare Access policy keyed on the same token) so unauthenticated
574
+ traffic is rejected before it ever reaches the origin.