llm-relay 0.83.2 → 0.84.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. package/README.md +9 -2
  2. package/dist/availability.d.ts +1 -1
  3. package/dist/backend.d.ts +1 -1
  4. package/dist/candidate-runner.js +2 -0
  5. package/dist/candidate-runner.js.map +1 -1
  6. package/dist/circuit-breaker.d.ts +1 -1
  7. package/dist/cli.d.ts +10 -1
  8. package/dist/cli.js +18 -3
  9. package/dist/cli.js.map +1 -1
  10. package/dist/config/routing-parser.js +3 -1
  11. package/dist/config/routing-parser.js.map +1 -1
  12. package/dist/config-types.d.ts +20 -10
  13. package/dist/config-types.js.map +1 -1
  14. package/dist/config.d.ts +1 -1
  15. package/dist/dispatch-lane-stats.d.ts +1 -1
  16. package/dist/dispatch.d.ts +3 -3
  17. package/dist/hedge-trigger.d.ts +1 -1
  18. package/dist/lane-activity.d.ts +42 -0
  19. package/dist/lane-activity.js +54 -0
  20. package/dist/lane-activity.js.map +1 -0
  21. package/dist/lane-cadence.d.ts +1 -1
  22. package/dist/lane-manifest.d.ts +1 -1
  23. package/dist/lane-probe.d.ts +2 -2
  24. package/dist/latency-demotion.d.ts +1 -1
  25. package/dist/mcp/job-archive.js +0 -1
  26. package/dist/mcp/job-archive.js.map +1 -1
  27. package/dist/mcp/lane-runner.d.ts +10 -7
  28. package/dist/mcp/lane-runner.js +3 -3
  29. package/dist/mcp/lane-runner.js.map +1 -1
  30. package/dist/mcp/protocol.d.ts +1 -1
  31. package/dist/mcp/readonly-boundary.d.ts +1 -1
  32. package/dist/mcp/server.d.ts +45 -53
  33. package/dist/mcp/server.js +90 -53
  34. package/dist/mcp/server.js.map +1 -1
  35. package/dist/process-safety-net.d.ts +1 -1
  36. package/dist/responses-request.d.ts +1 -1
  37. package/dist/routes/admin.js +21 -0
  38. package/dist/routes/admin.js.map +1 -1
  39. package/dist/server.js +20 -0
  40. package/dist/server.js.map +1 -1
  41. package/dist/stream-commit.d.ts +1 -1
  42. package/dist/stream-pipeline.d.ts +1 -1
  43. package/dist/tool-dialects.d.ts +1 -1
  44. package/docs/README.md +43 -0
  45. package/package.json +1 -1
  46. package/scripts/install-skill.mjs +1 -0
  47. package/skills/llm-relay/SKILL.md +4 -0
  48. package/skills/llm-relay/references/lane-field-notes.md +202 -0
@@ -0,0 +1,202 @@
1
+ # Lane field notes — measured behaviour of llm-relay dispatch and its lanes
2
+
3
+ Every note here was MEASURED on this machine, with its date. They moved out of the machine-wide
4
+ backlog on 2026-09-17 (owner instruction: llm-relay-specific instructions belong with the llm-relay
5
+ skill, not in a shared to-do file). Each note is REFERENCE, not work: it is deleted when it becomes
6
+ untrue, never because something shipped. A defect in llm-relay itself belongs in
7
+ `C:\Code\llm-relay\docs\backlog.md` in product terms.
8
+
9
+ Ask the live tools first. `dispatch_lanes` and `llm-relay dispatch` carry the current lane record
10
+ and quota; never copy a dated roster out of prose.
11
+
12
+ ## 1. Reading a dispatch reply
13
+
14
+ - **A `cli` rung's reply body is NOT the lane's answer (2026-09-16).** A `relay` rung such as
15
+ `free-pool` returns the raw answer. A `cli` rung such as `agy-gemini` returns its harness's own
16
+ record — `{"conversation_id":…,"status":"SUCCESS","response":"<the real answer>",
17
+ "duration_seconds":…,"usage":{…}}` — so the answer sits inside `response`. Code that binds the
18
+ body to a schema fails on exactly the calls the ladder sent to a CLI rung, and the failure looks
19
+ like a bad lane. Measured in audit-tools, 2026-09-16: 28 of 62 sweep calls lost, 19 to this
20
+ envelope. (llm-relay packet M1 will unwrap it in the relay and announce an `unwrapped:` line;
21
+ until that release, unwrap it yourself.)
22
+ - **A `SUCCESS` status is not an answer (2026-09-16, owner decided not to fix this in the relay).**
23
+ A `posttooluse-typecheck.mjs` dispatch (job-0021) returned `"status":"SUCCESS"` around 40 repeated
24
+ lines of "Waiting for test execution to complete." and no work product. Read the content by hand.
25
+ - **A completed job can carry no usable answer (2026-09-07 and 2026-09-09).** Agent-mode jobs
26
+ exited after 302 s and 623 s with the single words `Now` and `Let`, and their worktrees were
27
+ clean. Job `job-0005` exited 0 after 765 s and returned `I`. A review job completed after 813 s
28
+ with a completion claim and no findings. Treat a terminal status as evidence of TERMINATION only,
29
+ then inspect the answer and the worktree. None of these outcomes proves a provider quota is
30
+ spent.
31
+ - **`status: "running"` after a long `waitMs` is normal (2026-09-16).** The server clamps `waitMs`
32
+ to `routing.mcp.maxWaitMs`. Poll `dispatch_status`, then `dispatch_result`. In audit-tools 9 of
33
+ 62 calls were lost by treating this as a fault.
34
+ - **Map a lane's verdict table to the file by HEADING, never by row number (2026-09-10).** A
35
+ read-only lane numbered 74 backlog entries out of file order (one moved from row 20 to row 8),
36
+ so its row numbers would have deleted the wrong entries. Re-derive rows with a script and match
37
+ each verdict by its heading text.
38
+
39
+ ## 2. Waits, restarts and lost jobs
40
+
41
+ - **A large `waitMs` loses the job id (2026-09-06).** `waitMs: 180000` returned `Error: Request
42
+ timed out` with no job id, so the lane could not be polled, resumed or cancelled. The same task
43
+ at `waitMs: 40000` returned `job-0001` at once and finished in 145 s. Leave `waitMs` unset.
44
+ - **Codex's code-mode `exec` gives up on a tool call at 31.0 s (2026-09-10).** It returns "Wall
45
+ time 31.0 seconds" with empty output, so a longer MCP call loses its answer: 29 of 266 first
46
+ Codex `dispatch` calls, 2026-09-07 to 2026-09-10. Keep every MCP call from a Codex host under
47
+ 30 s. llm-relay blocks 25 s by default since v0.81.0 and then hands back a job id.
48
+ Since v0.83.x the server waits longer only for hosts measured to survive it: Claude Code with a
49
+ progress token gets the answer in one call, Claude Desktop gets 50 s, every other host keeps the
50
+ ceiling (llm-relay `docs/history/mcp-host-timeouts-2026-09-17.md`).
51
+ - **An MCP server restart KILLS every lane it was running (2026-09-06, re-measured 2026-09-17).**
52
+ The old symptom — `unknown jobId` for every job, numbering restarted at `job-0001`, nothing on
53
+ disk — is fixed: the running-job journal reports each such job as `killed` (v0.80.0) and the job
54
+ archive keeps finished jobs and the id counter across restarts (v0.82.0). The WORK is still
55
+ lost: 14 of 83 archived jobs on 2026-09-17 were `killed`, 11 of them `agy-gemini`. Give every
56
+ lane its own worktree, because that directory is the only record of what a killed lane did, and
57
+ dispatch a killed job again.
58
+ - **One `llm-relay mcp` process can outlive a release (2026-09-10).** The Claude desktop app keeps
59
+ one MCP connection across its sessions, so after a global reinstall that connection runs the old
60
+ code until the app restarts. Since v0.81.0 a reply from a process older than the installed
61
+ package says so.
62
+ - **A session that hits its usage limit mid-turn loses every in-flight job (2026-09-04).** Three
63
+ `relay` subagents lost their lane jobs when the limit hit: the MCP connection was replaced, job
64
+ ids restarted at `job-0001`, and each agent saw `Request timed out` then `Connection closed`.
65
+ Keep the job handles in the main session rather than in subagents, and re-dispatch after the
66
+ reset.
67
+ - **An unrunnable lane is not a quota verdict (2026-09-09).** `Lane "anthropic" cannot be run from
68
+ here: it is a relay target (anthropic) with no cliLane template configured` means no answer was
69
+ produced. It says nothing about any account's quota. Keep the exact error.
70
+
71
+ ## 3. Lane capacity and concurrency
72
+
73
+ - **Cap concurrent lanes at three or four; seven died together (2026-09-06).** Five
74
+ `opencode-muse-spark` and two `agy-gemini` lanes all hit the 2100 s timeout at the same moment,
75
+ exit 124, empty output, four packets lost — one had written 348 lines. Three lanes ran
76
+ comfortably afterwards. The relay does not cap this; the caller must.
77
+ - The tell is SIMULTANEITY. One silent lane is a lane problem; a cohort dying at the same
78
+ elapsed second is a load problem.
79
+ - The usual cause is asking every lane to run the full repository gate. Give a lane its TARGETED
80
+ suites plus lint, and keep the full gate for the orchestrating session.
81
+ - The orchestrator's own gate runs count toward the same budget: a cohort of only three lanes
82
+ also died together at 1800 s while the orchestrating session ran five full gates. Pause
83
+ dispatching while you verify, or drop to two lanes.
84
+ - A killed lane holds its worktree directory open, so `git worktree remove` reports "Permission
85
+ denied" although git DOES unregister the worktree. Believe `git worktree list`, not the
86
+ filesystem.
87
+ - A gate step that reaches the network flakes under this load: `npm audit` returned an error
88
+ payload while three lanes ran. Rerun such a step alone before believing it.
89
+ - **Two concurrent Muse Spark lanes starve, not only three (2026-09-09).** Two packets dispatched
90
+ together ran to the 2100 s timeout with empty output and no file written; single lanes finish in
91
+ minutes. Its rungs carry `maxConcurrent: 1` since llm-relay v0.78.0; keep it that way.
92
+
93
+ ## 4. What each lane can carry
94
+
95
+ - **`opencode-muse-spark`** carries a whole implementation packet ALONE (101–998 s, 2026-09-09),
96
+ and starves as a second or third concurrent lane. Always pass `cwd`. Keep the task text under
97
+ 4,096 characters and put a long brief in a file: a longer task makes the MCP server fall back to
98
+ its start-time config snapshot. It runs as the `relay-lane` OpenCode agent.
99
+ ⚠ Headless OpenCode auto-rejects every permission set to `ask`, and the global default sets
100
+ `edit` and `bash` to `ask`, so without that agent a lane can read but cannot edit or run a suite,
101
+ and the failure looks like a model failure (measured 2026-09-04: `permission requested: edit …;
102
+ auto-rejecting`). A lane dispatched with no `cwd` had every READ rejected as
103
+ `external_directory`. The agent lives in `~/.config/opencode/opencode.json`; a repository-level
104
+ `opencode.json` merges over it, and `agent=<name>` on the run's `stream` lines in
105
+ `~/.local/share/opencode/log/opencode.log` is the only proof of which agent ran.
106
+ ⚠ The agent did not end the zero-output mode (2026-09-04/05): with everything correct, two lanes
107
+ ran 918 s and 895 s and wrote zero bytes. Read-only recon on this lane is 35–70 s, so no file
108
+ change in the worktree after about five minutes is the signal to cancel and re-dispatch.
109
+ - **`agy-gemini`** is the steady CLI lane: 24 of 24 answered at a 240 s median on 2026-09-06, and
110
+ it carried packets at 348–561 s. It ran two lanes in one worktree (authoring plus a read-only
111
+ review) without interference.
112
+ ⚠ It obeys an absolute path written INSIDE the brief over the `cwd` you passed, even when the
113
+ task says not to (2026-09-06). Never write an absolute worktree path into a shared brief, or
114
+ regenerate the brief per lane. Before concluding a lane produced nothing, look where its brief
115
+ pointed.
116
+ - **`agy-claude-opus`** drops the stream on long outputs (2026-09-04/05): `The stream was
117
+ interrupted` after a report summary, and `There was a network issue connecting to the server`
118
+ after 490 s. Split the work into packets and run them on `agy-gemini`.
119
+ - **Codex Spark** reads its whole usage window and writes nothing (2026-09-09, twice): 193k and
120
+ 477k tokens, every test file read whole, then "You've hit your usage limit". A preamble limiting
121
+ reads changed nothing. Give it a review of a bounded diff, or nothing.
122
+ - **A shape that scored 10/10 yesterday is not a cure (2026-09-06).** Three `opencode-muse-spark`
123
+ lanes using the exact shape recorded as reliable the day before wrote nothing in 25 minutes while
124
+ `dispatch_status` said `running` and ten `opencode.exe` processes sat at about 500 MB each. Read
125
+ `dispatch_lanes` before choosing a lane; the same day it read 72 calls / 6 timed out / median
126
+ 616 s for Muse Spark against 24 / 24 ok / median 240 s for `agy-gemini`.
127
+
128
+ ## 5. Writing a brief, and trusting what comes back
129
+
130
+ - **A brief's wording is NOT a boundary (2026-09-06, measured twice).** A lane told to AUDIT a
131
+ closeout spent 19 minutes writing its own `closeout-input.json` into a live repository root, with
132
+ a fabricated verification section. Five later lanes each opened with "Do NOT edit any file.
133
+ Report findings only"; one still ran suites in the shared tree, wrote to the repository root,
134
+ performed the repository's own closeout ceremony, and an untracked deliverable of the
135
+ orchestrating session vanished at the same minute. The controls that DO work: give every writing
136
+ lane its own worktree, use `mode: "answer"` when the lane needs no file access, use
137
+ `readOnly: true` for an agent lane (llm-relay binds the lane's own read-only tool flags since
138
+ v0.82.0), commit an in-progress deliverable before dispatching into the same tree, and run
139
+ `git status --porcelain` after EVERY lane returns or is cancelled. A cancelled lane leaves its
140
+ files behind exactly like a completed one.
141
+ - **A brief that says "never print the key" does not stop a lane writing the key (2026-09-09).** A
142
+ capture lane put `export DEEPSEEK_API_KEY="sk-…"` into a scratch `start-relay.sh`. Tell the lane
143
+ to read a secret from the environment at run time and never copy the value into a file, and grep
144
+ every scratch launcher for `sk-` before running it or passing it on.
145
+ - **Ask a lane to EXTRACT, not to give a VERDICT (2026-09-06).** A free-pool lane given a rubric
146
+ and asked for `clear`/`defective` answered `clear` for all 97 records: a rubric whose rules
147
+ mostly say "this is not a defect" pushes a weak model to the null answer, and the output looks
148
+ well formed. The same job as an extraction — list the terms a cold reader cannot resolve, rate
149
+ 0–10 — did not collapse. Always check a lane's label DISTRIBUTION before using its labels.
150
+ - **Well-formed output can still be wrong, and that is the version that gets believed
151
+ (2026-09-06).** A lane produced 120 clean records in the requested shape; against a hand-labelled
152
+ overlap its best agreement was 64%, while answering "clear" every time scored 73%. Never fold
153
+ lane labels into a count, a training set or a conclusion without measuring agreement on a
154
+ hand-labelled overlap, and always against the always-answer-the-majority baseline.
155
+ - **A lane that returns one large JSON object at the end returns NOTHING when it stops early
156
+ (2026-09-06, three lanes lost).** Have the lane append one JSON line per record as it works, and
157
+ slice the job to about 40 records rather than 100.
158
+ - **Free lanes cannot do open-ended reconnaissance here (2026-09-05, 7 of 7 packets fabricated).**
159
+ They CAN review a concrete diff against a stated claim, and they carry a mechanical rewrite with
160
+ a stated rule. The test is whether the output can be checked by running or reading something
161
+ specific.
162
+
163
+ ## 6. Hooks, keys and the daemon
164
+
165
+ - **A relay lane loads NO global hook (2026-09-17).** `dispatch` launches each Claude lane with
166
+ `CLAUDE_CONFIG_DIR=~/.llm-relay-claude`, so the lane reads `~/.llm-relay-claude/settings.json`
167
+ and never `~/.claude/settings.json`. A global hook guards the orchestrator's own tool calls only.
168
+ Codex, OpenCode and AGY lanes have no hook surface at all; for them the orchestrator-side
169
+ `dispatch-cwd-guard.mjs` and the lane's own worktree are the only controls.
170
+ - **Provider API keys are NOT environment variables on this machine (2026-09-09).** `llm-relay
171
+ keys` prints an `Env var` column, which is the NAME the relay looks for, not proof the variable
172
+ exists: `NVIDIA_API_KEY` is empty in every scope while the same key reads `VALID`, because the
173
+ secret lives DPAPI-wrapped in `~/.llm-relay/keystore.json`. A direct `curl` therefore sends an
174
+ empty bearer, and NVIDIA answers HTTP 500 with a Rust `axum::Extension` message that reads like a
175
+ provider fault. Probe through the relay instead: `POST http://127.0.0.1:8791/v1/messages` with
176
+ `"model": "<provider>/<model id>"`. A public `/v1/models` answer proves nothing about a key.
177
+ - **A provider timeout turns a slow model into a fake "not servable" (2026-09-09).** With `nim` at
178
+ `timeoutMs: 100000`, a 32-token probe of `deepseek-ai/deepseek-v4-flash-0731` returned HTTP 504
179
+ at 100.03 s, while its sibling answered 200 after 81.6 s for two output tokens; the same Flash
180
+ model had answered in 38 s on 2026-08-27. A 504 at the configured timeout is evidence about the
181
+ QUEUE. Raise `providers.<name>.timeoutMs` and probe again before recording a model as dead.
182
+ `firstByteTimeoutMs` (v0.78.0) fails over fast when nothing arrives at all, while a slow body
183
+ still runs to its end.
184
+ - **The daemon reads `config.json` ONCE at start.** A rung or provider edit is invisible until the
185
+ daemon restarts — confirmed live: after a rung edit the `agy.exe` command line still carried the
186
+ old `--model`, and a re-probe after a timeout raise timed out again at exactly the old value.
187
+ Each `llm-relay mcp` process loads the file once too, so a host restart is needed for it as well.
188
+ `GET /telemetry` carries `config.changedOnDisk`, and `routing show|get`, `config show|get` and
189
+ `offload status` print a notice when the running relay has not loaded an edit. Restart: stop the
190
+ node process running `dist\cli.js` with no subcommand (`llm-relay stop` since v0.78.0), then
191
+ relaunch `wscript.exe "…\Startup\llm-relay.vbs"`; verify with `GET /telemetry` and
192
+ `llm-relay dispatch --tier high`.
193
+ - **A pool request to DeepSeek runs with thinking ON, which spends output tokens on reasoning
194
+ (2026-09-09, causes since addressed).** Two authorized paid calls to `deepseek/deepseek-v4-pro`
195
+ returned HTTP 200 and `max_tokens` with zero final text. llm-relay forwards the caller's thinking
196
+ control since 2026-09-10 and `dispatch` takes a `model` argument since v0.81.0. For a pool
197
+ request, give a large `max_tokens` or send `thinking: {"type": "disabled"}`.
198
+ - **Codex Desktop cannot reach a relay pool through a collaboration child (2026-08-31).** With a
199
+ ChatGPT account the launcher validates `pool/medium` against the parent account before contacting
200
+ llm-relay and fails with HTTP 400 `The 'pool/medium' model is not supported when using Codex with
201
+ a ChatGPT account.` Use the MCP `dispatch` tool there. Generated agent files stay valid for
202
+ clients that honour custom providers.