llm-relay 0.8.1 → 0.8.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/skills/llm-relay/SKILL.md +48 -17
package/package.json
CHANGED
|
@@ -95,29 +95,60 @@ targets — the host agent reaches them by shelling out to the vendor CLI, and t
|
|
|
95
95
|
**one ordered ladder** together with the relay's pools. Walk it top to bottom; each rung falls
|
|
96
96
|
back to the next on failure or quota exhaustion, exactly like candidates inside a relay pool.
|
|
97
97
|
|
|
98
|
-
**Default ladder for offloadable work** (bulk recon, extraction, analysis)
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
`
|
|
110
|
-
|
|
98
|
+
**Default ladder for offloadable work** (bulk recon, extraction, analysis). Included
|
|
99
|
+
subscription/credit allowances are spent before metered-or-free API capacity, so the cheap lanes
|
|
100
|
+
are the *first* resort, not the fallback:
|
|
101
|
+
|
|
102
|
+
1. **Antigravity — Gemini 3.6 Flash, any tier.** The default workhorse for every tier; pick the
|
|
103
|
+
reasoning level to suit the task, since AGY bakes it into the model id:
|
|
104
|
+
```bash
|
|
105
|
+
agy -p "<task>" --model gemini-3.6-flash-medium --output-format json
|
|
106
|
+
```
|
|
107
|
+
`-high` for analysis and tracing, `-medium` for ordinary recon, `-low` for mechanical sweeps.
|
|
108
|
+
Other flags: `--add-dir <path>` to scope the workspace, `--json-schema` for structured output,
|
|
109
|
+
`--print-timeout` (default 5m), `--mode plan` for analysis-only runs.
|
|
110
|
+
*Exhausted when:* the CLI reports the **Gemini** credit balance exhausted, or rate-limits.
|
|
111
|
+
2. **Antigravity — Claude, same CLI, SEPARATE quota.** ⚠ AGY meters Gemini and Claude against
|
|
112
|
+
**two independent credit balances**, so exhausting rung 1 does *not* exhaust this rung — that
|
|
113
|
+
is precisely why it is a real fallback and not a duplicate of rung 1. Step up here when Flash
|
|
114
|
+
is not strong enough, or when Gemini credits are gone: `--model claude-opus-4-6-thinking` for
|
|
115
|
+
hard reasoning, `--model claude-sonnet-4-6` for everything else. (Neither id carries a level
|
|
116
|
+
suffix — use the session flag `--effort low|medium|high` to tune them.)
|
|
117
|
+
*Exhausted when:* the **Claude** balance is exhausted — independently of rung 1, in either
|
|
118
|
+
direction.
|
|
119
|
+
3. **Codex — Sol, then Terra, then Luna.** Spends the ChatGPT subscription:
|
|
120
|
+
```bash
|
|
121
|
+
codex exec --model gpt-5.6-sol -c model_reasoning_effort="medium" "<task>"
|
|
122
|
+
```
|
|
123
|
+
Walk `gpt-5.6-sol` → `gpt-5.6-terra` → `gpt-5.6-luna` in that order. Reasoning level is a
|
|
124
|
+
config override, not a flag — Codex has no `--effort`: `-c model_reasoning_effort=` accepts
|
|
125
|
+
`minimal|low|medium|high|xhigh` (the config default in `~/.codex/config.toml` is `high`;
|
|
126
|
+
`plan_mode_reasoning_effort` sets it separately for plan mode). ⚠ Codex's own guidance is that
|
|
127
|
+
high effort burns subscription rate limits fast — match the level to the task rather than
|
|
128
|
+
leaving it at `high` for mechanical work. `codex exec review` runs a repo review.
|
|
129
|
+
*Exhausted when:* it reports usage-limit errors.
|
|
130
|
+
4. **Relay pools — free API-key capacity, benchmark order.** `@relay: pool/coding` (or the tier
|
|
131
|
+
mapping in `routing.subagents` when offload is on). Ordering *inside* a pool is by synced
|
|
132
|
+
benchmark strength, not config order — see "Reordering dispatch" below. *Exhausted when:* the
|
|
133
|
+
pool 4xx/5xxs after failover walks every candidate, or `llm-relay candidates` shows the breaker
|
|
134
|
+
open / quota drained across the pool.
|
|
135
|
+
5. **Anthropic subagent** — plain `Agent(...)`, no directive. Spends primary quota; always works.
|
|
136
|
+
|
|
137
|
+
`agy models` and the Codex model list are the authority on what exists — re-check them rather than
|
|
138
|
+
trusting these ids after a CLI upgrade, since a de-listed id fails the whole rung.
|
|
111
139
|
|
|
112
140
|
Rules for walking it:
|
|
113
141
|
|
|
114
142
|
- **Skip a rung whose CLI is not installed** (`Get-Command agy` / `codex` or `command -v`) — this
|
|
115
|
-
ladder degrades gracefully to "relay, then Anthropic" on machines without the peer CLIs.
|
|
143
|
+
ladder degrades gracefully to "relay pools, then Anthropic" on machines without the peer CLIs.
|
|
116
144
|
- Both CLIs are **full agents with their own tool loops** — hand them a self-contained prompt with
|
|
117
145
|
file paths, run long tasks in the background, and treat output as advisory (verify against
|
|
118
146
|
source) exactly like relay-offloaded output. Do NOT wrap them in a bare one-shot HTTP helper.
|
|
119
147
|
- A **refusal or a wrong answer is not a transport failure** — do not walk the ladder to shop for
|
|
120
148
|
a more compliant model. Only availability failures (errors, quota, rate limits) advance a rung.
|
|
149
|
+
- **The rungs draw on five independent buckets**, so exhausting one never implies the next is
|
|
150
|
+
gone: AGY-Gemini, AGY-Claude (separate balances despite one CLI), ChatGPT, provider API keys,
|
|
151
|
+
Anthropic primary. Re-check the rung you skipped on the next dispatch — a reset refills it.
|
|
121
152
|
- Interactive-only quotas (e.g. an IDE-bound plan with no CLI) are unreachable by any dispatcher;
|
|
122
153
|
don't try to MITM them into the ladder.
|
|
123
154
|
|
|
@@ -127,9 +158,9 @@ Ordering exists at three levels; change the right one:
|
|
|
127
158
|
|
|
128
159
|
- **The ladder above** (which lane is tried first): it is instructions, not code — edit the
|
|
129
160
|
numbered list in this skill file (`skills/llm-relay/SKILL.md` in the repo; the installed copy
|
|
130
|
-
lives in `~/.claude/skills/llm-relay/`, refreshed on package upgrade). E.g. to
|
|
131
|
-
|
|
132
|
-
("try codex first for this").
|
|
161
|
+
lives in `~/.claude/skills/llm-relay/`, refreshed on package upgrade). E.g. to preserve AGY
|
|
162
|
+
credits and spend free API capacity first, move rung 4 to the top. The user can also reorder
|
|
163
|
+
per-request in chat ("try codex first for this").
|
|
133
164
|
- **Which pool a tier lands on** (`routing.subagents` in `~/.llm-relay/config.json`): maps the
|
|
134
165
|
Agent tool's `model` param (opus/sonnet/haiku/…) to a pool or pinned spec. Takes effect on the
|
|
135
166
|
next request; no restart.
|