llm-relay 0.8.3 → 0.8.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "llm-relay",
3
- "version": "0.8.3",
3
+ "version": "0.8.4",
4
4
  "description": "Loopback Anthropic-Messages-API proxy that validates and repairs tool-call responses in flight, routing across multi-provider LLM backends.",
5
5
  "type": "module",
6
6
  "engines": {
@@ -95,47 +95,41 @@ targets — the host agent reaches them by shelling out to the vendor CLI, and t
95
95
  **one ordered ladder** together with the relay's pools. Walk it top to bottom; each rung falls
96
96
  back to the next on failure or quota exhaustion, exactly like candidates inside a relay pool.
97
97
 
98
- **Default ladder for offloadable work** (bulk recon, extraction, analysis). Included
99
- subscription/credit allowances are spent before metered-or-free API capacity, so the cheap lanes
100
- are the *first* resort, not the fallback:
101
-
102
- 1. **Antigravity Gemini 3.6 Flash, any tier.** The default workhorse for every tier; pick the
103
- reasoning level to suit the task, since AGY bakes it into the model id:
104
- ```bash
105
- agy -p "<task>" --model gemini-3.6-flash-medium --output-format json
106
- ```
107
- `-high` for analysis and tracing, `-medium` for ordinary recon, `-low` for mechanical sweeps.
108
- Other flags: `--add-dir <path>` to scope the workspace, `--json-schema` for structured output,
109
- `--print-timeout` (default 5m), `--mode plan` for analysis-only runs.
110
- *Exhausted when:* the CLI reports the **Gemini** credit balance exhausted, or rate-limits.
111
- 2. **AntigravityClaude, same CLI, SEPARATE quota.** AGY meters Gemini and Claude against
112
- **two independent credit balances**, so exhausting rung 1 does *not* exhaust this rung — that
113
- is precisely why it is a real fallback and not a duplicate of rung 1. Step up here when Flash
114
- is not strong enough, or when Gemini credits are gone: `--model claude-opus-4-6-thinking` for
115
- hard reasoning, `--model claude-sonnet-4-6` for everything else. (Neither id carries a level
116
- suffix use the session flag `--effort low|medium|high` to tune them.)
117
- *Exhausted when:* the **Claude** balance is exhausted — independently of rung 1, in either
118
- direction.
119
- 3. **Codex Sol, then Terra, then Luna.** Spends the ChatGPT subscription:
120
- ```bash
121
- codex exec --model gpt-5.6-sol -c model_reasoning_effort="medium" "<task>"
122
- ```
123
- Walk `gpt-5.6-sol` → `gpt-5.6-terra` → `gpt-5.6-luna` in that order. Reasoning level is a
124
- config override, not a flag — Codex has no `--effort`: `-c model_reasoning_effort=` accepts
125
- `minimal|low|medium|high|xhigh` (the config default in `~/.codex/config.toml` is `high`;
126
- `plan_mode_reasoning_effort` sets it separately for plan mode). Codex's own guidance is that
127
- high effort burns subscription rate limits fast match the level to the task rather than
128
- leaving it at `high` for mechanical work. `codex exec review` runs a repo review.
129
- *Exhausted when:* it reports usage-limit errors.
130
- 4. **Relay pools — free API-key capacity, benchmark order.** `@relay: pool/coding` (or the tier
131
- mapping in `routing.subagents` when offload is on). Ordering *inside* a pool is by synced
132
- benchmark strength, not config order see "Reordering dispatch" below. *Exhausted when:* the
133
- pool 4xx/5xxs after failover walks every candidate, or `llm-relay candidates` shows the breaker
134
- open / quota drained across the pool.
135
- 5. **Anthropic subagent** — plain `Agent(...)`, no directive. Spends primary quota; always works.
136
-
137
- `agy models` and the Codex model list are the authority on what exists — re-check them rather than
138
- trusting these ids after a CLI upgrade, since a de-listed id fails the whole rung.
98
+ ### The lanes, and how to drive each
99
+
100
+ - **Relay pools** — `@relay: pool/coding` on a subagent prompt, or the tier mapping in
101
+ `routing.subagents` when offload is on. Spends provider API keys. *Exhausted when:* the pool
102
+ 4xx/5xxs after failover walks every candidate, or `llm-relay candidates` shows the breaker open
103
+ / quota drained across the pool.
104
+ - **Antigravity (`agy`)** — `agy -p "<task>" --model <id> --output-format json`. Ask
105
+ `agy models` for the roster; ids may carry a reasoning-level suffix (`…-high|-medium|-low`),
106
+ and `--effort low|medium|high` tunes ids that don't. Other flags: `--add-dir <path>` to scope
107
+ the workspace, `--json-schema` for structured output, `--print-timeout` (default 5m),
108
+ `--mode plan` for analysis-only runs. **AGY meters its Gemini and its Claude models against
109
+ two independent credit balances** exhausting one leaves the other fully available, so they
110
+ are two distinct rungs, not one.
111
+ - **Codex**`codex exec --model <id> "<task>"`, spending the ChatGPT subscription. Reasoning
112
+ level is a config override, not a flag (there is no `--effort`):
113
+ `-c model_reasoning_effort="minimal|low|medium|high|xhigh"`, defaulting to whatever
114
+ `~/.codex/config.toml` sets; `plan_mode_reasoning_effort` sets it separately for plan mode.
115
+ Codex's own guidance is that high effort burns subscription rate limits fast — match the
116
+ level to the task rather than leaving it high for mechanical work. `codex exec review` runs a
117
+ repo review.
118
+ - **Anthropic subagent** — plain `Agent(...)`, no directive. Spends primary quota; always works,
119
+ so it is the natural bottom of any ladder.
120
+
121
+ The CLIs' own model lists are the authority on what exists — re-check them rather than trusting
122
+ ids written down anywhere, since a de-listed id fails a whole rung.
123
+
124
+ ### Order
125
+
126
+ **A user's own ordering wins. Look for it first**, in `~/.claude/CLAUDE.md` (or a project
127
+ `CLAUDE.md`) that is where a personal ladder belongs, because **this file is overwritten by
128
+ `npm i -g llm-relay` on every install and upgrade** and any ordering edited into it would be
129
+ silently lost. Never write a user's personal preference here; write it there.
130
+
131
+ Absent such an instruction, the package default is simply: **relay pools peer agent CLIs (if
132
+ installed) Anthropic subagent**cheapest metered capacity first, primary quota last.
139
133
 
140
134
  Rules for walking it:
141
135
 
@@ -156,11 +150,12 @@ Rules for walking it:
156
150
 
157
151
  Ordering exists at three levels; change the right one:
158
152
 
159
- - **The ladder above** (which lane is tried first): it is instructions, not code edit the
160
- numbered list in this skill file (`skills/llm-relay/SKILL.md` in the repo; the installed copy
161
- lives in `~/.claude/skills/llm-relay/`, refreshed on package upgrade). E.g. to preserve AGY
162
- credits and spend free API capacity first, move rung 4 to the top. The user can also reorder
163
- per-request in chat ("try codex first for this").
153
+ - **Which lane is tried first**: it is instructions, not code. Put the ordered list in
154
+ **`~/.claude/CLAUDE.md`** (all projects) or a project `CLAUDE.md` (one repo) both are loaded
155
+ every session and **neither is touched by an npm install**, so the ordering survives reinstalls
156
+ and upgrades. Do *not* put it in `~/.claude/skills/llm-relay/SKILL.md`: `postinstall`
157
+ overwrites that file from the package on every global install, silently discarding the edit.
158
+ The user can also reorder per-request in chat ("try codex first for this").
164
159
  - **Which pool a tier lands on** (`routing.subagents` in `~/.llm-relay/config.json`): maps the
165
160
  Agent tool's `model` param (opus/sonnet/haiku/…) to a pool or pinned spec. Takes effect on the
166
161
  next request; no restart.