llm-relay 0.7.0 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,120 @@
1
+ ---
2
+ name: llm-relay
3
+ description: >-
4
+ Operate llm-relay, the loopback multi-provider LLM proxy (default 127.0.0.1:8791) that
5
+ validates/repairs tool calls and can offload Claude Code subagents to non-Anthropic
6
+ providers. Use when offloading bulk work to a subagent on another provider, choosing an
7
+ offload target, addressing a pool or model through the relay, toggling subagent offload,
8
+ or diagnosing a request that failed at or behind the relay.
9
+ ---
10
+
11
+ # llm-relay — operating guide
12
+
13
+ llm-relay is a **loopback-only** reverse proxy for the Anthropic `/v1/messages` API. It routes
14
+ each request to a configured provider (Anthropic passthrough, NIM, OpenRouter, Gemini, Groq,
15
+ Mistral, …), translating to/from OpenAI-compatible backends, and **validates + repairs malformed
16
+ tool calls** so agent harnesses can run on models that are weaker at tool use. Config, keys and
17
+ caches live in `~/.llm-relay/` (`config.json`, `.env`, `models-cache.json`, …).
18
+
19
+ One boundary governs everything it does: the proxy fixes **protocol form** (tool-call args that
20
+ violate the schema), never **judgment**. It refuses to fabricate destructive tool calls, and an
21
+ unrepairable response fails loudly (502 / mid-stream SSE error) rather than passing through broken.
22
+
23
+ ## Addressing a model
24
+
25
+ Three forms in a request's `model` field, resolved in this order:
26
+
27
+ | Form | Goes to | Ranked / failover? |
28
+ |---|---|---|
29
+ | `pool/<name>` | every candidate in `routing.pools[<name>]` | yes — benchmark-ranked, walks candidates on failure |
30
+ | `<provider>/<model>` (e.g. `nim/z-ai/glm-5.2`) | that exact deployment, verbatim | no — deliberately pinned |
31
+ | a Claude model id (`claude-opus-5`, …) | `routing.tiers` → Anthropic passthrough | n/a |
32
+
33
+ **Prefer `pool/<name>` over a pinned spec** — a pool survives one model being de-listed; a pin does
34
+ not. An unknown pool or provider is a loud 400 listing valid names, never a silent fallback.
35
+ Pool refs also work inside `routing.tiers`, `routing.default` and `routing.subagents`; all of them
36
+ are validated at config load, so a typo fails at startup, not on the first request.
37
+
38
+ An unnamespaced/unknown model id lands on `routing.default` — in the standard setup that is the
39
+ Anthropic passthrough, so it reaches real Anthropic (spending real quota), never a silently weaker
40
+ model.
41
+
42
+ ## Subagent offload (OPT-IN — off by default)
43
+
44
+ Claude Code stamps `cc_is_subagent=true` into the `system` block of subagent requests. When the
45
+ offload switch is ON, those requests (and only those) route through `routing.subagents`
46
+ (tier → spec); the human's own conversation never consults that map.
47
+
48
+ ```bash
49
+ llm-relay offload status # where things stand
50
+ llm-relay offload on # takes effect on the next request, no restart, persisted
51
+ llm-relay offload off
52
+ ```
53
+
54
+ Three ways to steer a subagent, in precedence order:
55
+
56
+ 1. **`@relay: <spec>` directive** — put it on its own line at the START of the subagent's prompt
57
+ (`@relay: pool/coding` or `@relay: nim/z-ai/glm-5.2`). Stripped before forwarding, so the model
58
+ never sees it. **Works with the switch OFF** — this is the per-call opt-in.
59
+ 2. **Tier** *(switch must be on)* — the Agent tool's `model` param maps through
60
+ `routing.subagents` (e.g. opus→`pool/reasoning`, sonnet→`pool/coding`, haiku→`pool/fast`).
61
+ 3. **Nothing** *(switch on)* — the inherited model id matches a tier, else `subagents.default`.
62
+
63
+ ⚠ Dispatching a subagent does NOT offload it by itself. With the switch off, a subagent runs on
64
+ Anthropic like any other request. Check with `llm-relay offload status`, don't assume.
65
+
66
+ Offloaded output is **advisory** — verify claims against source files before acting on them.
67
+
68
+ ## Choosing a target
69
+
70
+ ```bash
71
+ llm-relay candidates # one row per offload target, all dimensions side by side
72
+ ```
73
+
74
+ The table is deliberately **un-blended** — capability from each leaderboard separately (AA
75
+ agentic/coding, BFCL tool-use, Aider polyglot, LMArena), price, context, live health (verdict,
76
+ p95), quota, breaker state, and traffic observed through this proxy. Weigh the columns yourself:
77
+
78
+ - `str` is the one scalar (pool ordering needs an order) and always carries provenance:
79
+ `83.3/4` = four published signals; `obs` = ranked on this proxy's own traffic; `neut` = nothing
80
+ known. A blank cell means **not measured**, never "bad".
81
+ - `~` on ctx/$ means the figure belongs to a **different host** serving the same model id
82
+ (e.g. NIM publishes nothing, so OpenRouter's numbers are shown as reference). Never quote a `~`
83
+ figure as the serving provider's real ceiling or rate.
84
+ - Capability is synced (`npm run sync:tiers` in the repo), never hand-typed.
85
+
86
+ `GET 127.0.0.1:8791/candidates` returns the full JSON (every raw score, jitter, observed calls).
87
+
88
+ ## Everyday commands
89
+
90
+ ```bash
91
+ llm-relay models -p nim # live roster per provider (listed ≠ servable — some listed ids 404)
92
+ llm-relay keys # provider key health + quota
93
+ llm-relay ping # latency/stability probe across providers
94
+ llm-relay telemetry # JSON health/quota report
95
+ ```
96
+
97
+ Runtime endpoints on the running proxy: `/registry`, `/candidates`, `/offload` (GET/POST),
98
+ `/telemetry`, `/ping`, `/health`.
99
+
100
+ ## Failure modes worth knowing
101
+
102
+ - **Relay down** → clients pointed at it fail to start. It must be running before anything routes.
103
+ - **404 from an openai backend** is nearly always the model id: a model can be listed in `/models`
104
+ and still not be served (NIM does this). The error says so; pick another candidate or a pool.
105
+ - **429s pass through** — the client's retry/backoff handles them; the relay's circuit breaker
106
+ cools that target down and failover walks the next pool candidate.
107
+ - **400 "exceeds the context limit"** fires only when the serving provider itself published a
108
+ limit. Unknown limit = no guardrail; the backend answers with its own authoritative error.
109
+ - **Repair refused/failed** → 502 `tool call could not be repaired (…)`. That is fail-clean by
110
+ design: a refusal is a judgement and is never retried on another model.
111
+ - **Document blocks** to openai backends are converted to markdown via MarkItDown
112
+ (`pip install 'markitdown[all]'`); without it, requests carrying documents fail with a clear
113
+ error instead of injecting base64 into the prompt.
114
+
115
+ ## Safety invariants (do not work around these)
116
+
117
+ - Loopback bind only — it holds provider keys and does no auth.
118
+ - Logs are metadata-only; never ask it to log request/response bodies.
119
+ - Destructive tool calls are refused, never fabricated — repair output may run under
120
+ `--dangerously-skip-permissions`.