llm-relay 0.7.0 → 0.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "llm-relay",
3
- "version": "0.7.0",
3
+ "version": "0.8.1",
4
4
  "description": "Loopback Anthropic-Messages-API proxy that validates and repairs tool-call responses in flight, routing across multi-provider LLM backends.",
5
5
  "type": "module",
6
6
  "engines": {
@@ -20,6 +20,8 @@
20
20
  "files": [
21
21
  "dist",
22
22
  "docs/tier-data.json",
23
+ "skills",
24
+ "scripts/install-skill.mjs",
23
25
  "README.md",
24
26
  "config.example.json"
25
27
  ],
@@ -30,6 +32,7 @@
30
32
  "sync:tiers": "node scripts/sync-tiers.mjs",
31
33
  "test": "vitest run",
32
34
  "typecheck": "tsc -p tsconfig.json --noEmit",
35
+ "postinstall": "node scripts/install-skill.mjs || exit 0",
33
36
  "prepublishOnly": "npm run build && npm test"
34
37
  },
35
38
  "dependencies": {
@@ -0,0 +1,43 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * Install/refresh the Claude Code skill that ships with llm-relay into the user's
4
+ * global skills directory (~/.claude/skills/llm-relay/SKILL.md).
5
+ *
6
+ * Runs from npm `postinstall`, but only acts on GLOBAL installs (`npm i -g llm-relay`)
7
+ * so that a repo-local `npm install` (dev checkout, CI) never touches the developer's
8
+ * ~/.claude. Because the self-updater reinstalls the global package on a new version,
9
+ * the skill refreshes itself on every upgrade with no extra step.
10
+ *
11
+ * `--force` installs regardless of install context (for manual runs and tests).
12
+ * Best-effort by design: a failure here must never fail the package install.
13
+ */
14
+ import { copyFileSync, mkdirSync, existsSync } from "node:fs";
15
+ import { join, dirname } from "node:path";
16
+ import { homedir } from "node:os";
17
+ import { fileURLToPath } from "node:url";
18
+
19
+ const force = process.argv.includes("--force");
20
+
21
+ // Two signals, either suffices: npm's lifecycle env var, or the package physically living in
22
+ // an npm global tree (…/npm/node_modules/llm-relay on Windows, …/lib/node_modules/llm-relay
23
+ // elsewhere) — the env var alone is npm-version-dependent.
24
+ const pkgDir = join(dirname(fileURLToPath(import.meta.url)), "..");
25
+ const inGlobalTree = /(^|[\\/])(npm|lib)[\\/]node_modules[\\/]llm-relay$/i.test(pkgDir.replace(/[\\/]+$/, ""));
26
+ const isGlobal = process.env.npm_config_global === "true" || inGlobalTree;
27
+
28
+ if (!isGlobal && !force) {
29
+ process.exit(0); // local/dev install — leave the user's ~/.claude alone
30
+ }
31
+
32
+ try {
33
+ const src = join(dirname(fileURLToPath(import.meta.url)), "..", "skills", "llm-relay", "SKILL.md");
34
+ if (!existsSync(src)) process.exit(0); // packed without the skill — nothing to do
35
+
36
+ const dest = join(homedir(), ".claude", "skills", "llm-relay", "SKILL.md");
37
+ mkdirSync(dirname(dest), { recursive: true });
38
+ copyFileSync(src, dest);
39
+ process.stderr.write(`llm-relay: installed Claude Code skill at ${dest}\n`);
40
+ } catch (e) {
41
+ // Never fail the install over the skill — say why and move on.
42
+ process.stderr.write(`llm-relay: skill install skipped (${e?.message ?? e})\n`);
43
+ }
@@ -0,0 +1,175 @@
1
+ ---
2
+ name: llm-relay
3
+ description: >-
4
+ Operate llm-relay, the loopback multi-provider LLM proxy (default 127.0.0.1:8791) that
5
+ validates/repairs tool calls and can offload Claude Code subagents to non-Anthropic
6
+ providers. Use when offloading bulk work to a subagent on another provider, choosing an
7
+ offload target, addressing a pool or model through the relay, toggling subagent offload,
8
+ dispatching to peer agent CLIs (Antigravity/Codex) as fallback lanes, reordering dispatch,
9
+ or diagnosing a request that failed at or behind the relay.
10
+ ---
11
+
12
+ # llm-relay — operating guide
13
+
14
+ llm-relay is a **loopback-only** reverse proxy for the Anthropic `/v1/messages` API. It routes
15
+ each request to a configured provider (Anthropic passthrough, NIM, OpenRouter, Gemini, Groq,
16
+ Mistral, …), translating to/from OpenAI-compatible backends, and **validates + repairs malformed
17
+ tool calls** so agent harnesses can run on models that are weaker at tool use. Config, keys and
18
+ caches live in `~/.llm-relay/` (`config.json`, `.env`, `models-cache.json`, …).
19
+
20
+ One boundary governs everything it does: the proxy fixes **protocol form** (tool-call args that
21
+ violate the schema), never **judgment**. It refuses to fabricate destructive tool calls, and an
22
+ unrepairable response fails loudly (502 / mid-stream SSE error) rather than passing through broken.
23
+
24
+ ## Addressing a model
25
+
26
+ Three forms in a request's `model` field, resolved in this order:
27
+
28
+ | Form | Goes to | Ranked / failover? |
29
+ |---|---|---|
30
+ | `pool/<name>` | every candidate in `routing.pools[<name>]` | yes — benchmark-ranked, walks candidates on failure |
31
+ | `<provider>/<model>` (e.g. `nim/z-ai/glm-5.2`) | that exact deployment, verbatim | no — deliberately pinned |
32
+ | a Claude model id (`claude-opus-5`, …) | `routing.tiers` → Anthropic passthrough | n/a |
33
+
34
+ **Prefer `pool/<name>` over a pinned spec** — a pool survives one model being de-listed; a pin does
35
+ not. An unknown pool or provider is a loud 400 listing valid names, never a silent fallback.
36
+ Pool refs also work inside `routing.tiers`, `routing.default` and `routing.subagents`; all of them
37
+ are validated at config load, so a typo fails at startup, not on the first request.
38
+
39
+ An unnamespaced/unknown model id lands on `routing.default` — in the standard setup that is the
40
+ Anthropic passthrough, so it reaches real Anthropic (spending real quota), never a silently weaker
41
+ model.
42
+
43
+ ## Subagent offload (OPT-IN — off by default)
44
+
45
+ Claude Code stamps `cc_is_subagent=true` into the `system` block of subagent requests. When the
46
+ offload switch is ON, those requests (and only those) route through `routing.subagents`
47
+ (tier → spec); the human's own conversation never consults that map.
48
+
49
+ ```bash
50
+ llm-relay offload status # where things stand
51
+ llm-relay offload on # takes effect on the next request, no restart, persisted
52
+ llm-relay offload off
53
+ ```
54
+
55
+ Three ways to steer a subagent, in precedence order:
56
+
57
+ 1. **`@relay: <spec>` directive** — put it on its own line at the START of the subagent's prompt
58
+ (`@relay: pool/coding` or `@relay: nim/z-ai/glm-5.2`). Stripped before forwarding, so the model
59
+ never sees it. **Works with the switch OFF** — this is the per-call opt-in.
60
+ 2. **Tier** *(switch must be on)* — the Agent tool's `model` param maps through
61
+ `routing.subagents` (e.g. opus→`pool/reasoning`, sonnet→`pool/coding`, haiku→`pool/fast`).
62
+ 3. **Nothing** *(switch on)* — the inherited model id matches a tier, else `subagents.default`.
63
+
64
+ ⚠ Dispatching a subagent does NOT offload it by itself. With the switch off, a subagent runs on
65
+ Anthropic like any other request. Check with `llm-relay offload status`, don't assume.
66
+
67
+ Offloaded output is **advisory** — verify claims against source files before acting on them.
68
+
69
+ ## Choosing a target
70
+
71
+ ```bash
72
+ llm-relay candidates # one row per offload target, all dimensions side by side
73
+ ```
74
+
75
+ The table is deliberately **un-blended** — capability from each leaderboard separately (AA
76
+ agentic/coding, BFCL tool-use, Aider polyglot, LMArena), price, context, live health (verdict,
77
+ p95), quota, breaker state, and traffic observed through this proxy. Weigh the columns yourself:
78
+
79
+ - `str` is the one scalar (pool ordering needs an order) and always carries provenance:
80
+ `83.3/4` = four published signals; `obs` = ranked on this proxy's own traffic; `neut` = nothing
81
+ known. A blank cell means **not measured**, never "bad".
82
+ - `~` on ctx/$ means the figure belongs to a **different host** serving the same model id
83
+ (e.g. NIM publishes nothing, so OpenRouter's numbers are shown as reference). Never quote a `~`
84
+ figure as the serving provider's real ceiling or rate.
85
+ - Capability is synced (`npm run sync:tiers` in the repo), never hand-typed.
86
+
87
+ `GET 127.0.0.1:8791/candidates` returns the full JSON (every raw score, jitter, observed calls).
88
+
89
+ ## The dispatch ladder — including agent-CLI lanes (Antigravity, Codex)
90
+
91
+ Subscription and CLI-credit quotas are **client-bound**: only the vendor's own client can spend
92
+ them, so the relay cannot front them as providers (Antigravity's endpoint is compiled into its
93
+ binary; Codex's ChatGPT path uses client-bound OAuth on `/v1/responses`). They are still dispatch
94
+ targets — the host agent reaches them by shelling out to the vendor CLI, and they participate in
95
+ **one ordered ladder** together with the relay's pools. Walk it top to bottom; each rung falls
96
+ back to the next on failure or quota exhaustion, exactly like candidates inside a relay pool.
97
+
98
+ **Default ladder for offloadable work** (bulk recon, extraction, analysis):
99
+
100
+ 1. **Relay pool** — `@relay: pool/coding` subagent (or tier via `routing.subagents` when offload
101
+ is on). Free API-key capacity. *Exhausted when:* the pool 4xx/5xxs after failover walks every
102
+ candidate, or `llm-relay candidates` shows the breaker open / quota drained across the pool.
103
+ 2. **Antigravity CLI** — `agy -p "<task>" --output-format json` (spends AGY CLI credits; the
104
+ `gemini` CLI is its deprecated former name). Useful flags: `--model` (`agy models` lists),
105
+ `--effort low|medium|high`, `--add-dir <path>` to scope the workspace, `--json-schema` for
106
+ structured output, `--print-timeout` (default 5m), `--mode plan` for analysis-only runs.
107
+ *Exhausted when:* the CLI reports credits/quota exhausted or rate-limits.
108
+ 3. **Codex CLI** — `codex exec "<task>"` (spends the ChatGPT subscription). Also
109
+ `codex exec review` for repo review. *Exhausted when:* it reports usage-limit errors.
110
+ 4. **Anthropic subagent** — plain `Agent(...)`, no directive. Spends primary quota; always works.
111
+
112
+ Rules for walking it:
113
+
114
+ - **Skip a rung whose CLI is not installed** (`Get-Command agy` / `codex` or `command -v`) — this
115
+ ladder degrades gracefully to "relay, then Anthropic" on machines without the peer CLIs.
116
+ - Both CLIs are **full agents with their own tool loops** — hand them a self-contained prompt with
117
+ file paths, run long tasks in the background, and treat output as advisory (verify against
118
+ source) exactly like relay-offloaded output. Do NOT wrap them in a bare one-shot HTTP helper.
119
+ - A **refusal or a wrong answer is not a transport failure** — do not walk the ladder to shop for
120
+ a more compliant model. Only availability failures (errors, quota, rate limits) advance a rung.
121
+ - Interactive-only quotas (e.g. an IDE-bound plan with no CLI) are unreachable by any dispatcher;
122
+ don't try to MITM them into the ladder.
123
+
124
+ ## Reordering dispatch
125
+
126
+ Ordering exists at three levels; change the right one:
127
+
128
+ - **The ladder above** (which lane is tried first): it is instructions, not code — edit the
129
+ numbered list in this skill file (`skills/llm-relay/SKILL.md` in the repo; the installed copy
130
+ lives in `~/.claude/skills/llm-relay/`, refreshed on package upgrade). E.g. to burn AGY credits
131
+ before free API keys, swap rungs 1 and 2. The user can also reorder per-request in chat
132
+ ("try codex first for this").
133
+ - **Which pool a tier lands on** (`routing.subagents` in `~/.llm-relay/config.json`): maps the
134
+ Agent tool's `model` param (opus/sonnet/haiku/…) to a pool or pinned spec. Takes effect on the
135
+ next request; no restart.
136
+ - **Candidate order inside a pool** (`routing.pools`): with `"benchmarkSort": true` (default
137
+ setup) failover order is by synced benchmark strength and the config array only breaks ties.
138
+ To make the array order authoritative, set `"benchmarkSort": false`. Either way the circuit
139
+ breaker still demotes unhealthy targets — that is live health, not preference, and it is what
140
+ you want. For an absolutely fixed destination, pin `<provider>/<model>`; a pin is never
141
+ reordered and never fails over.
142
+
143
+ ## Everyday commands
144
+
145
+ ```bash
146
+ llm-relay models -p nim # live roster per provider (listed ≠ servable — some listed ids 404)
147
+ llm-relay keys # provider key health + quota
148
+ llm-relay ping # latency/stability probe across providers
149
+ llm-relay telemetry # JSON health/quota report
150
+ ```
151
+
152
+ Runtime endpoints on the running proxy: `/registry`, `/candidates`, `/offload` (GET/POST),
153
+ `/telemetry`, `/ping`, `/health`.
154
+
155
+ ## Failure modes worth knowing
156
+
157
+ - **Relay down** → clients pointed at it fail to start. It must be running before anything routes.
158
+ - **404 from an openai backend** is nearly always the model id: a model can be listed in `/models`
159
+ and still not be served (NIM does this). The error says so; pick another candidate or a pool.
160
+ - **429s pass through** — the client's retry/backoff handles them; the relay's circuit breaker
161
+ cools that target down and failover walks the next pool candidate.
162
+ - **400 "exceeds the context limit"** fires only when the serving provider itself published a
163
+ limit. Unknown limit = no guardrail; the backend answers with its own authoritative error.
164
+ - **Repair refused/failed** → 502 `tool call could not be repaired (…)`. That is fail-clean by
165
+ design: a refusal is a judgement and is never retried on another model.
166
+ - **Document blocks** to openai backends are converted to markdown via MarkItDown
167
+ (`pip install 'markitdown[all]'`); without it, requests carrying documents fail with a clear
168
+ error instead of injecting base64 into the prompt.
169
+
170
+ ## Safety invariants (do not work around these)
171
+
172
+ - Loopback bind only — it holds provider keys and does no auth.
173
+ - Logs are metadata-only; never ask it to log request/response bodies.
174
+ - Destructive tool calls are refused, never fabricated — repair output may run under
175
+ `--dangerously-skip-permissions`.