llm-relay 0.83.2 → 0.84.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +9 -2
- package/dist/availability.d.ts +1 -1
- package/dist/backend.d.ts +1 -1
- package/dist/candidate-runner.js +2 -0
- package/dist/candidate-runner.js.map +1 -1
- package/dist/circuit-breaker.d.ts +1 -1
- package/dist/cli.d.ts +10 -1
- package/dist/cli.js +18 -3
- package/dist/cli.js.map +1 -1
- package/dist/config/routing-parser.js +3 -1
- package/dist/config/routing-parser.js.map +1 -1
- package/dist/config-types.d.ts +20 -10
- package/dist/config-types.js.map +1 -1
- package/dist/config.d.ts +1 -1
- package/dist/dispatch-lane-stats.d.ts +1 -1
- package/dist/dispatch.d.ts +3 -3
- package/dist/hedge-trigger.d.ts +1 -1
- package/dist/lane-activity.d.ts +42 -0
- package/dist/lane-activity.js +54 -0
- package/dist/lane-activity.js.map +1 -0
- package/dist/lane-cadence.d.ts +1 -1
- package/dist/lane-manifest.d.ts +1 -1
- package/dist/lane-probe.d.ts +2 -2
- package/dist/latency-demotion.d.ts +1 -1
- package/dist/mcp/job-archive.js +0 -1
- package/dist/mcp/job-archive.js.map +1 -1
- package/dist/mcp/lane-runner.d.ts +10 -7
- package/dist/mcp/lane-runner.js +3 -3
- package/dist/mcp/lane-runner.js.map +1 -1
- package/dist/mcp/protocol.d.ts +1 -1
- package/dist/mcp/readonly-boundary.d.ts +1 -1
- package/dist/mcp/server.d.ts +45 -53
- package/dist/mcp/server.js +90 -53
- package/dist/mcp/server.js.map +1 -1
- package/dist/process-safety-net.d.ts +1 -1
- package/dist/responses-request.d.ts +1 -1
- package/dist/routes/admin.js +21 -0
- package/dist/routes/admin.js.map +1 -1
- package/dist/server.js +20 -0
- package/dist/server.js.map +1 -1
- package/dist/stream-commit.d.ts +1 -1
- package/dist/stream-pipeline.d.ts +1 -1
- package/dist/tool-dialects.d.ts +1 -1
- package/docs/README.md +43 -0
- package/package.json +1 -1
- package/scripts/install-skill.mjs +1 -0
- package/skills/llm-relay/SKILL.md +4 -0
- package/skills/llm-relay/references/lane-field-notes.md +202 -0
|
@@ -0,0 +1,202 @@
|
|
|
1
|
+
# Lane field notes — measured behaviour of llm-relay dispatch and its lanes
|
|
2
|
+
|
|
3
|
+
Every note here was MEASURED on this machine, with its date. They moved out of the machine-wide
|
|
4
|
+
backlog on 2026-09-17 (owner instruction: llm-relay-specific instructions belong with the llm-relay
|
|
5
|
+
skill, not in a shared to-do file). Each note is REFERENCE, not work: it is deleted when it becomes
|
|
6
|
+
untrue, never because something shipped. A defect in llm-relay itself belongs in
|
|
7
|
+
`C:\Code\llm-relay\docs\backlog.md` in product terms.
|
|
8
|
+
|
|
9
|
+
Ask the live tools first. `dispatch_lanes` and `llm-relay dispatch` carry the current lane record
|
|
10
|
+
and quota; never copy a dated roster out of prose.
|
|
11
|
+
|
|
12
|
+
## 1. Reading a dispatch reply
|
|
13
|
+
|
|
14
|
+
- **A `cli` rung's reply body is NOT the lane's answer (2026-09-16).** A `relay` rung such as
|
|
15
|
+
`free-pool` returns the raw answer. A `cli` rung such as `agy-gemini` returns its harness's own
|
|
16
|
+
record — `{"conversation_id":…,"status":"SUCCESS","response":"<the real answer>",
|
|
17
|
+
"duration_seconds":…,"usage":{…}}` — so the answer sits inside `response`. Code that binds the
|
|
18
|
+
body to a schema fails on exactly the calls the ladder sent to a CLI rung, and the failure looks
|
|
19
|
+
like a bad lane. Measured in audit-tools, 2026-09-16: 28 of 62 sweep calls lost, 19 to this
|
|
20
|
+
envelope. (llm-relay packet M1 will unwrap it in the relay and announce an `unwrapped:` line;
|
|
21
|
+
until that release, unwrap it yourself.)
|
|
22
|
+
- **A `SUCCESS` status is not an answer (2026-09-16, owner decided not to fix this in the relay).**
|
|
23
|
+
A `posttooluse-typecheck.mjs` dispatch (job-0021) returned `"status":"SUCCESS"` around 40 repeated
|
|
24
|
+
lines of "Waiting for test execution to complete." and no work product. Read the content by hand.
|
|
25
|
+
- **A completed job can carry no usable answer (2026-09-07 and 2026-09-09).** Agent-mode jobs
|
|
26
|
+
exited after 302 s and 623 s with the single words `Now` and `Let`, and their worktrees were
|
|
27
|
+
clean. Job `job-0005` exited 0 after 765 s and returned `I`. A review job completed after 813 s
|
|
28
|
+
with a completion claim and no findings. Treat a terminal status as evidence of TERMINATION only,
|
|
29
|
+
then inspect the answer and the worktree. None of these outcomes proves a provider quota is
|
|
30
|
+
spent.
|
|
31
|
+
- **`status: "running"` after a long `waitMs` is normal (2026-09-16).** The server clamps `waitMs`
|
|
32
|
+
to `routing.mcp.maxWaitMs`. Poll `dispatch_status`, then `dispatch_result`. In audit-tools 9 of
|
|
33
|
+
62 calls were lost by treating this as a fault.
|
|
34
|
+
- **Map a lane's verdict table to the file by HEADING, never by row number (2026-09-10).** A
|
|
35
|
+
read-only lane numbered 74 backlog entries out of file order (one moved from row 20 to row 8),
|
|
36
|
+
so its row numbers would have deleted the wrong entries. Re-derive rows with a script and match
|
|
37
|
+
each verdict by its heading text.
|
|
38
|
+
|
|
39
|
+
## 2. Waits, restarts and lost jobs
|
|
40
|
+
|
|
41
|
+
- **A large `waitMs` loses the job id (2026-09-06).** `waitMs: 180000` returned `Error: Request
|
|
42
|
+
timed out` with no job id, so the lane could not be polled, resumed or cancelled. The same task
|
|
43
|
+
at `waitMs: 40000` returned `job-0001` at once and finished in 145 s. Leave `waitMs` unset.
|
|
44
|
+
- **Codex's code-mode `exec` gives up on a tool call at 31.0 s (2026-09-10).** It returns "Wall
|
|
45
|
+
time 31.0 seconds" with empty output, so a longer MCP call loses its answer: 29 of 266 first
|
|
46
|
+
Codex `dispatch` calls, 2026-09-07 to 2026-09-10. Keep every MCP call from a Codex host under
|
|
47
|
+
30 s. llm-relay blocks 25 s by default since v0.81.0 and then hands back a job id.
|
|
48
|
+
Since v0.83.x the server waits longer only for hosts measured to survive it: Claude Code with a
|
|
49
|
+
progress token gets the answer in one call, Claude Desktop gets 50 s, every other host keeps the
|
|
50
|
+
ceiling (llm-relay `docs/history/mcp-host-timeouts-2026-09-17.md`).
|
|
51
|
+
- **An MCP server restart KILLS every lane it was running (2026-09-06, re-measured 2026-09-17).**
|
|
52
|
+
The old symptom — `unknown jobId` for every job, numbering restarted at `job-0001`, nothing on
|
|
53
|
+
disk — is fixed: the running-job journal reports each such job as `killed` (v0.80.0) and the job
|
|
54
|
+
archive keeps finished jobs and the id counter across restarts (v0.82.0). The WORK is still
|
|
55
|
+
lost: 14 of 83 archived jobs on 2026-09-17 were `killed`, 11 of them `agy-gemini`. Give every
|
|
56
|
+
lane its own worktree, because that directory is the only record of what a killed lane did, and
|
|
57
|
+
dispatch a killed job again.
|
|
58
|
+
- **One `llm-relay mcp` process can outlive a release (2026-09-10).** The Claude desktop app keeps
|
|
59
|
+
one MCP connection across its sessions, so after a global reinstall that connection runs the old
|
|
60
|
+
code until the app restarts. Since v0.81.0 a reply from a process older than the installed
|
|
61
|
+
package says so.
|
|
62
|
+
- **A session that hits its usage limit mid-turn loses every in-flight job (2026-09-04).** Three
|
|
63
|
+
`relay` subagents lost their lane jobs when the limit hit: the MCP connection was replaced, job
|
|
64
|
+
ids restarted at `job-0001`, and each agent saw `Request timed out` then `Connection closed`.
|
|
65
|
+
Keep the job handles in the main session rather than in subagents, and re-dispatch after the
|
|
66
|
+
reset.
|
|
67
|
+
- **An unrunnable lane is not a quota verdict (2026-09-09).** `Lane "anthropic" cannot be run from
|
|
68
|
+
here: it is a relay target (anthropic) with no cliLane template configured` means no answer was
|
|
69
|
+
produced. It says nothing about any account's quota. Keep the exact error.
|
|
70
|
+
|
|
71
|
+
## 3. Lane capacity and concurrency
|
|
72
|
+
|
|
73
|
+
- **Cap concurrent lanes at three or four; seven died together (2026-09-06).** Five
|
|
74
|
+
`opencode-muse-spark` and two `agy-gemini` lanes all hit the 2100 s timeout at the same moment,
|
|
75
|
+
exit 124, empty output, four packets lost — one had written 348 lines. Three lanes ran
|
|
76
|
+
comfortably afterwards. The relay does not cap this; the caller must.
|
|
77
|
+
- The tell is SIMULTANEITY. One silent lane is a lane problem; a cohort dying at the same
|
|
78
|
+
elapsed second is a load problem.
|
|
79
|
+
- The usual cause is asking every lane to run the full repository gate. Give a lane its TARGETED
|
|
80
|
+
suites plus lint, and keep the full gate for the orchestrating session.
|
|
81
|
+
- The orchestrator's own gate runs count toward the same budget: a cohort of only three lanes
|
|
82
|
+
also died together at 1800 s while the orchestrating session ran five full gates. Pause
|
|
83
|
+
dispatching while you verify, or drop to two lanes.
|
|
84
|
+
- A killed lane holds its worktree directory open, so `git worktree remove` reports "Permission
|
|
85
|
+
denied" although git DOES unregister the worktree. Believe `git worktree list`, not the
|
|
86
|
+
filesystem.
|
|
87
|
+
- A gate step that reaches the network flakes under this load: `npm audit` returned an error
|
|
88
|
+
payload while three lanes ran. Rerun such a step alone before believing it.
|
|
89
|
+
- **Two concurrent Muse Spark lanes starve, not only three (2026-09-09).** Two packets dispatched
|
|
90
|
+
together ran to the 2100 s timeout with empty output and no file written; single lanes finish in
|
|
91
|
+
minutes. Its rungs carry `maxConcurrent: 1` since llm-relay v0.78.0; keep it that way.
|
|
92
|
+
|
|
93
|
+
## 4. What each lane can carry
|
|
94
|
+
|
|
95
|
+
- **`opencode-muse-spark`** carries a whole implementation packet ALONE (101–998 s, 2026-09-09),
|
|
96
|
+
and starves as a second or third concurrent lane. Always pass `cwd`. Keep the task text under
|
|
97
|
+
4,096 characters and put a long brief in a file: a longer task makes the MCP server fall back to
|
|
98
|
+
its start-time config snapshot. It runs as the `relay-lane` OpenCode agent.
|
|
99
|
+
⚠ Headless OpenCode auto-rejects every permission set to `ask`, and the global default sets
|
|
100
|
+
`edit` and `bash` to `ask`, so without that agent a lane can read but cannot edit or run a suite,
|
|
101
|
+
and the failure looks like a model failure (measured 2026-09-04: `permission requested: edit …;
|
|
102
|
+
auto-rejecting`). A lane dispatched with no `cwd` had every READ rejected as
|
|
103
|
+
`external_directory`. The agent lives in `~/.config/opencode/opencode.json`; a repository-level
|
|
104
|
+
`opencode.json` merges over it, and `agent=<name>` on the run's `stream` lines in
|
|
105
|
+
`~/.local/share/opencode/log/opencode.log` is the only proof of which agent ran.
|
|
106
|
+
⚠ The agent did not end the zero-output mode (2026-09-04/05): with everything correct, two lanes
|
|
107
|
+
ran 918 s and 895 s and wrote zero bytes. Read-only recon on this lane is 35–70 s, so no file
|
|
108
|
+
change in the worktree after about five minutes is the signal to cancel and re-dispatch.
|
|
109
|
+
- **`agy-gemini`** is the steady CLI lane: 24 of 24 answered at a 240 s median on 2026-09-06, and
|
|
110
|
+
it carried packets at 348–561 s. It ran two lanes in one worktree (authoring plus a read-only
|
|
111
|
+
review) without interference.
|
|
112
|
+
⚠ It obeys an absolute path written INSIDE the brief over the `cwd` you passed, even when the
|
|
113
|
+
task says not to (2026-09-06). Never write an absolute worktree path into a shared brief, or
|
|
114
|
+
regenerate the brief per lane. Before concluding a lane produced nothing, look where its brief
|
|
115
|
+
pointed.
|
|
116
|
+
- **`agy-claude-opus`** drops the stream on long outputs (2026-09-04/05): `The stream was
|
|
117
|
+
interrupted` after a report summary, and `There was a network issue connecting to the server`
|
|
118
|
+
after 490 s. Split the work into packets and run them on `agy-gemini`.
|
|
119
|
+
- **Codex Spark** reads its whole usage window and writes nothing (2026-09-09, twice): 193k and
|
|
120
|
+
477k tokens, every test file read whole, then "You've hit your usage limit". A preamble limiting
|
|
121
|
+
reads changed nothing. Give it a review of a bounded diff, or nothing.
|
|
122
|
+
- **A shape that scored 10/10 yesterday is not a cure (2026-09-06).** Three `opencode-muse-spark`
|
|
123
|
+
lanes using the exact shape recorded as reliable the day before wrote nothing in 25 minutes while
|
|
124
|
+
`dispatch_status` said `running` and ten `opencode.exe` processes sat at about 500 MB each. Read
|
|
125
|
+
`dispatch_lanes` before choosing a lane; the same day it read 72 calls / 6 timed out / median
|
|
126
|
+
616 s for Muse Spark against 24 / 24 ok / median 240 s for `agy-gemini`.
|
|
127
|
+
|
|
128
|
+
## 5. Writing a brief, and trusting what comes back
|
|
129
|
+
|
|
130
|
+
- **A brief's wording is NOT a boundary (2026-09-06, measured twice).** A lane told to AUDIT a
|
|
131
|
+
closeout spent 19 minutes writing its own `closeout-input.json` into a live repository root, with
|
|
132
|
+
a fabricated verification section. Five later lanes each opened with "Do NOT edit any file.
|
|
133
|
+
Report findings only"; one still ran suites in the shared tree, wrote to the repository root,
|
|
134
|
+
performed the repository's own closeout ceremony, and an untracked deliverable of the
|
|
135
|
+
orchestrating session vanished at the same minute. The controls that DO work: give every writing
|
|
136
|
+
lane its own worktree, use `mode: "answer"` when the lane needs no file access, use
|
|
137
|
+
`readOnly: true` for an agent lane (llm-relay binds the lane's own read-only tool flags since
|
|
138
|
+
v0.82.0), commit an in-progress deliverable before dispatching into the same tree, and run
|
|
139
|
+
`git status --porcelain` after EVERY lane returns or is cancelled. A cancelled lane leaves its
|
|
140
|
+
files behind exactly like a completed one.
|
|
141
|
+
- **A brief that says "never print the key" does not stop a lane writing the key (2026-09-09).** A
|
|
142
|
+
capture lane put `export DEEPSEEK_API_KEY="sk-…"` into a scratch `start-relay.sh`. Tell the lane
|
|
143
|
+
to read a secret from the environment at run time and never copy the value into a file, and grep
|
|
144
|
+
every scratch launcher for `sk-` before running it or passing it on.
|
|
145
|
+
- **Ask a lane to EXTRACT, not to give a VERDICT (2026-09-06).** A free-pool lane given a rubric
|
|
146
|
+
and asked for `clear`/`defective` answered `clear` for all 97 records: a rubric whose rules
|
|
147
|
+
mostly say "this is not a defect" pushes a weak model to the null answer, and the output looks
|
|
148
|
+
well formed. The same job as an extraction — list the terms a cold reader cannot resolve, rate
|
|
149
|
+
0–10 — did not collapse. Always check a lane's label DISTRIBUTION before using its labels.
|
|
150
|
+
- **Well-formed output can still be wrong, and that is the version that gets believed
|
|
151
|
+
(2026-09-06).** A lane produced 120 clean records in the requested shape; against a hand-labelled
|
|
152
|
+
overlap its best agreement was 64%, while answering "clear" every time scored 73%. Never fold
|
|
153
|
+
lane labels into a count, a training set or a conclusion without measuring agreement on a
|
|
154
|
+
hand-labelled overlap, and always against the always-answer-the-majority baseline.
|
|
155
|
+
- **A lane that returns one large JSON object at the end returns NOTHING when it stops early
|
|
156
|
+
(2026-09-06, three lanes lost).** Have the lane append one JSON line per record as it works, and
|
|
157
|
+
slice the job to about 40 records rather than 100.
|
|
158
|
+
- **Free lanes cannot do open-ended reconnaissance here (2026-09-05, 7 of 7 packets fabricated).**
|
|
159
|
+
They CAN review a concrete diff against a stated claim, and they carry a mechanical rewrite with
|
|
160
|
+
a stated rule. The test is whether the output can be checked by running or reading something
|
|
161
|
+
specific.
|
|
162
|
+
|
|
163
|
+
## 6. Hooks, keys and the daemon
|
|
164
|
+
|
|
165
|
+
- **A relay lane loads NO global hook (2026-09-17).** `dispatch` launches each Claude lane with
|
|
166
|
+
`CLAUDE_CONFIG_DIR=~/.llm-relay-claude`, so the lane reads `~/.llm-relay-claude/settings.json`
|
|
167
|
+
and never `~/.claude/settings.json`. A global hook guards the orchestrator's own tool calls only.
|
|
168
|
+
Codex, OpenCode and AGY lanes have no hook surface at all; for them the orchestrator-side
|
|
169
|
+
`dispatch-cwd-guard.mjs` and the lane's own worktree are the only controls.
|
|
170
|
+
- **Provider API keys are NOT environment variables on this machine (2026-09-09).** `llm-relay
|
|
171
|
+
keys` prints an `Env var` column, which is the NAME the relay looks for, not proof the variable
|
|
172
|
+
exists: `NVIDIA_API_KEY` is empty in every scope while the same key reads `VALID`, because the
|
|
173
|
+
secret lives DPAPI-wrapped in `~/.llm-relay/keystore.json`. A direct `curl` therefore sends an
|
|
174
|
+
empty bearer, and NVIDIA answers HTTP 500 with a Rust `axum::Extension` message that reads like a
|
|
175
|
+
provider fault. Probe through the relay instead: `POST http://127.0.0.1:8791/v1/messages` with
|
|
176
|
+
`"model": "<provider>/<model id>"`. A public `/v1/models` answer proves nothing about a key.
|
|
177
|
+
- **A provider timeout turns a slow model into a fake "not servable" (2026-09-09).** With `nim` at
|
|
178
|
+
`timeoutMs: 100000`, a 32-token probe of `deepseek-ai/deepseek-v4-flash-0731` returned HTTP 504
|
|
179
|
+
at 100.03 s, while its sibling answered 200 after 81.6 s for two output tokens; the same Flash
|
|
180
|
+
model had answered in 38 s on 2026-08-27. A 504 at the configured timeout is evidence about the
|
|
181
|
+
QUEUE. Raise `providers.<name>.timeoutMs` and probe again before recording a model as dead.
|
|
182
|
+
`firstByteTimeoutMs` (v0.78.0) fails over fast when nothing arrives at all, while a slow body
|
|
183
|
+
still runs to its end.
|
|
184
|
+
- **The daemon reads `config.json` ONCE at start.** A rung or provider edit is invisible until the
|
|
185
|
+
daemon restarts — confirmed live: after a rung edit the `agy.exe` command line still carried the
|
|
186
|
+
old `--model`, and a re-probe after a timeout raise timed out again at exactly the old value.
|
|
187
|
+
Each `llm-relay mcp` process loads the file once too, so a host restart is needed for it as well.
|
|
188
|
+
`GET /telemetry` carries `config.changedOnDisk`, and `routing show|get`, `config show|get` and
|
|
189
|
+
`offload status` print a notice when the running relay has not loaded an edit. Restart: stop the
|
|
190
|
+
node process running `dist\cli.js` with no subcommand (`llm-relay stop` since v0.78.0), then
|
|
191
|
+
relaunch `wscript.exe "…\Startup\llm-relay.vbs"`; verify with `GET /telemetry` and
|
|
192
|
+
`llm-relay dispatch --tier high`.
|
|
193
|
+
- **A pool request to DeepSeek runs with thinking ON, which spends output tokens on reasoning
|
|
194
|
+
(2026-09-09, causes since addressed).** Two authorized paid calls to `deepseek/deepseek-v4-pro`
|
|
195
|
+
returned HTTP 200 and `max_tokens` with zero final text. llm-relay forwards the caller's thinking
|
|
196
|
+
control since 2026-09-10 and `dispatch` takes a `model` argument since v0.81.0. For a pool
|
|
197
|
+
request, give a large `max_tokens` or send `thinking: {"type": "disabled"}`.
|
|
198
|
+
- **Codex Desktop cannot reach a relay pool through a collaboration child (2026-08-31).** With a
|
|
199
|
+
ChatGPT account the launcher validates `pool/medium` against the parent account before contacting
|
|
200
|
+
llm-relay and fails with HTTP 400 `The 'pool/medium' model is not supported when using Codex with
|
|
201
|
+
a ChatGPT account.` Use the MCP `dispatch` tool there. Generated agent files stay valid for
|
|
202
|
+
clients that honour custom providers.
|