llm-relay 0.8.1 → 0.8.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "llm-relay",
3
- "version": "0.8.1",
3
+ "version": "0.8.3",
4
4
  "description": "Loopback Anthropic-Messages-API proxy that validates and repairs tool-call responses in flight, routing across multi-provider LLM backends.",
5
5
  "type": "module",
6
6
  "engines": {
@@ -95,29 +95,60 @@ targets — the host agent reaches them by shelling out to the vendor CLI, and t
95
95
  **one ordered ladder** together with the relay's pools. Walk it top to bottom; each rung falls
96
96
  back to the next on failure or quota exhaustion, exactly like candidates inside a relay pool.
97
97
 
98
- **Default ladder for offloadable work** (bulk recon, extraction, analysis):
99
-
100
- 1. **Relay pool** `@relay: pool/coding` subagent (or tier via `routing.subagents` when offload
101
- is on). Free API-key capacity. *Exhausted when:* the pool 4xx/5xxs after failover walks every
102
- candidate, or `llm-relay candidates` shows the breaker open / quota drained across the pool.
103
- 2. **Antigravity CLI** `agy -p "<task>" --output-format json` (spends AGY CLI credits; the
104
- `gemini` CLI is its deprecated former name). Useful flags: `--model` (`agy models` lists),
105
- `--effort low|medium|high`, `--add-dir <path>` to scope the workspace, `--json-schema` for
106
- structured output, `--print-timeout` (default 5m), `--mode plan` for analysis-only runs.
107
- *Exhausted when:* the CLI reports credits/quota exhausted or rate-limits.
108
- 3. **Codex CLI** `codex exec "<task>"` (spends the ChatGPT subscription). Also
109
- `codex exec review` for repo review. *Exhausted when:* it reports usage-limit errors.
110
- 4. **Anthropic subagent** plain `Agent(...)`, no directive. Spends primary quota; always works.
98
+ **Default ladder for offloadable work** (bulk recon, extraction, analysis). Included
99
+ subscription/credit allowances are spent before metered-or-free API capacity, so the cheap lanes
100
+ are the *first* resort, not the fallback:
101
+
102
+ 1. **Antigravity Gemini 3.6 Flash, any tier.** The default workhorse for every tier; pick the
103
+ reasoning level to suit the task, since AGY bakes it into the model id:
104
+ ```bash
105
+ agy -p "<task>" --model gemini-3.6-flash-medium --output-format json
106
+ ```
107
+ `-high` for analysis and tracing, `-medium` for ordinary recon, `-low` for mechanical sweeps.
108
+ Other flags: `--add-dir <path>` to scope the workspace, `--json-schema` for structured output,
109
+ `--print-timeout` (default 5m), `--mode plan` for analysis-only runs.
110
+ *Exhausted when:* the CLI reports the **Gemini** credit balance exhausted, or rate-limits.
111
+ 2. **Antigravity — Claude, same CLI, SEPARATE quota.** ⚠ AGY meters Gemini and Claude against
112
+ **two independent credit balances**, so exhausting rung 1 does *not* exhaust this rung — that
113
+ is precisely why it is a real fallback and not a duplicate of rung 1. Step up here when Flash
114
+ is not strong enough, or when Gemini credits are gone: `--model claude-opus-4-6-thinking` for
115
+ hard reasoning, `--model claude-sonnet-4-6` for everything else. (Neither id carries a level
116
+ suffix — use the session flag `--effort low|medium|high` to tune them.)
117
+ *Exhausted when:* the **Claude** balance is exhausted — independently of rung 1, in either
118
+ direction.
119
+ 3. **Codex — Sol, then Terra, then Luna.** Spends the ChatGPT subscription:
120
+ ```bash
121
+ codex exec --model gpt-5.6-sol -c model_reasoning_effort="medium" "<task>"
122
+ ```
123
+ Walk `gpt-5.6-sol` → `gpt-5.6-terra` → `gpt-5.6-luna` in that order. Reasoning level is a
124
+ config override, not a flag — Codex has no `--effort`: `-c model_reasoning_effort=` accepts
125
+ `minimal|low|medium|high|xhigh` (the config default in `~/.codex/config.toml` is `high`;
126
+ `plan_mode_reasoning_effort` sets it separately for plan mode). ⚠ Codex's own guidance is that
127
+ high effort burns subscription rate limits fast — match the level to the task rather than
128
+ leaving it at `high` for mechanical work. `codex exec review` runs a repo review.
129
+ *Exhausted when:* it reports usage-limit errors.
130
+ 4. **Relay pools — free API-key capacity, benchmark order.** `@relay: pool/coding` (or the tier
131
+ mapping in `routing.subagents` when offload is on). Ordering *inside* a pool is by synced
132
+ benchmark strength, not config order — see "Reordering dispatch" below. *Exhausted when:* the
133
+ pool 4xx/5xxs after failover walks every candidate, or `llm-relay candidates` shows the breaker
134
+ open / quota drained across the pool.
135
+ 5. **Anthropic subagent** — plain `Agent(...)`, no directive. Spends primary quota; always works.
136
+
137
+ `agy models` and the Codex model list are the authority on what exists — re-check them rather than
138
+ trusting these ids after a CLI upgrade, since a de-listed id fails the whole rung.
111
139
 
112
140
  Rules for walking it:
113
141
 
114
142
  - **Skip a rung whose CLI is not installed** (`Get-Command agy` / `codex` or `command -v`) — this
115
- ladder degrades gracefully to "relay, then Anthropic" on machines without the peer CLIs.
143
+ ladder degrades gracefully to "relay pools, then Anthropic" on machines without the peer CLIs.
116
144
  - Both CLIs are **full agents with their own tool loops** — hand them a self-contained prompt with
117
145
  file paths, run long tasks in the background, and treat output as advisory (verify against
118
146
  source) exactly like relay-offloaded output. Do NOT wrap them in a bare one-shot HTTP helper.
119
147
  - A **refusal or a wrong answer is not a transport failure** — do not walk the ladder to shop for
120
148
  a more compliant model. Only availability failures (errors, quota, rate limits) advance a rung.
149
+ - **The rungs draw on five independent buckets**, so exhausting one never implies the next is
150
+ gone: AGY-Gemini, AGY-Claude (separate balances despite one CLI), ChatGPT, provider API keys,
151
+ Anthropic primary. Re-check the rung you skipped on the next dispatch — a reset refills it.
121
152
  - Interactive-only quotas (e.g. an IDE-bound plan with no CLI) are unreachable by any dispatcher;
122
153
  don't try to MITM them into the ladder.
123
154
 
@@ -127,9 +158,9 @@ Ordering exists at three levels; change the right one:
127
158
 
128
159
  - **The ladder above** (which lane is tried first): it is instructions, not code — edit the
129
160
  numbered list in this skill file (`skills/llm-relay/SKILL.md` in the repo; the installed copy
130
- lives in `~/.claude/skills/llm-relay/`, refreshed on package upgrade). E.g. to burn AGY credits
131
- before free API keys, swap rungs 1 and 2. The user can also reorder per-request in chat
132
- ("try codex first for this").
161
+ lives in `~/.claude/skills/llm-relay/`, refreshed on package upgrade). E.g. to preserve AGY
162
+ credits and spend free API capacity first, move rung 4 to the top. The user can also reorder
163
+ per-request in chat ("try codex first for this").
133
164
  - **Which pool a tier lands on** (`routing.subagents` in `~/.llm-relay/config.json`): maps the
134
165
  Agent tool's `model` param (opus/sonnet/haiku/…) to a pool or pinned spec. Takes effect on the
135
166
  next request; no restart.