llm-runtime-dock 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +708 -0
- package/bin/lrd.mjs +12 -0
- package/dist/index.js +26692 -0
- package/package.json +86 -0
package/README.md
ADDED
|
@@ -0,0 +1,708 @@
|
|
|
1
|
+
# llm-runtime-dock
|
|
2
|
+
|
|
3
|
+
One local endpoint for your coding agents, with the right model loaded behind
|
|
4
|
+
it — whatever runtime that model happens to live in.
|
|
5
|
+
|
|
6
|
+
`lrd` was built with AI assistance — Claude Code and OpenCode, working against
|
|
7
|
+
both local and remote models. The agent instruction files it was developed with
|
|
8
|
+
ship with the repository (`AGENTS.md`, `CLAUDE.md`, `.claude/skills/`), where
|
|
9
|
+
they double as the contributor extension procedures.
|
|
10
|
+
|
|
11
|
+
## Why this exists
|
|
12
|
+
|
|
13
|
+
If you run models locally, you probably run more than one runtime: MTPLX for
|
|
14
|
+
one model, LM Studio for another, something else for the third. Each has its own
|
|
15
|
+
CLI, its own model names, its own idea of when it is ready — and none of them
|
|
16
|
+
agrees with the others about what "switch to another model" even means.
|
|
17
|
+
|
|
18
|
+
MTPLX serves one model per process, so switching means stopping the server and
|
|
19
|
+
starting it again with different flags. LM Studio keeps one shared server and
|
|
20
|
+
loads models into it, so switching is `lms unload` and then `lms load`. Your
|
|
21
|
+
machine, meanwhile, has exactly one pool of memory: leaving the old model
|
|
22
|
+
resident while the next one loads is how you run out of it.
|
|
23
|
+
|
|
24
|
+
Which leaves you doing the orchestration by hand — stop that one, start this
|
|
25
|
+
one, wait until it is really ready, hope nothing was mid-request — and then
|
|
26
|
+
editing three coding-agent config files because the model name changed.
|
|
27
|
+
|
|
28
|
+
`lrd` is the thing that does that for you. Your agents talk to one address, and
|
|
29
|
+
it starts, stops and switches the runtime underneath so that the model a request
|
|
30
|
+
asked for is the model that answers it: one switch at a time, in-flight requests
|
|
31
|
+
drained first, a live stream never cut. And since it already has to know what
|
|
32
|
+
your backends serve, it will write that into your agents' configuration too.
|
|
33
|
+
|
|
34
|
+
```text
|
|
35
|
+
OpenCode · Claude Code · Codex
|
|
36
|
+
│ OpenAI or Anthropic API
|
|
37
|
+
▼
|
|
38
|
+
llm-runtime-dock · http://127.0.0.1:8787
|
|
39
|
+
│ CLI / process control, plus documented HTTP lifecycle
|
|
40
|
+
▼
|
|
41
|
+
MTPLX · LM Studio · oMLX · Ollama · your own runtime
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
Three commands, and none of them asks you to learn a new model format:
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
lrd probe mtplx --save # ask the backend what it serves, save it as config
|
|
48
|
+
lrd serve # one endpoint, the right model behind it
|
|
49
|
+
lrd apply claude # point a coding agent at it
|
|
50
|
+
```
|
|
51
|
+
|
|
52
|
+
It is a gateway and orchestrator, not an inference engine. Everything it knows
|
|
53
|
+
about your machine comes from one YAML file you can read.
|
|
54
|
+
|
|
55
|
+
**Contents**
|
|
56
|
+
|
|
57
|
+
- [What it works with](#what-it-works-with)
|
|
58
|
+
- [Install](#install)
|
|
59
|
+
- [Quickstart](#quickstart)
|
|
60
|
+
- [Configuration](#configuration)
|
|
61
|
+
- [Coding agents](#coding-agents)
|
|
62
|
+
- [Commands](#commands)
|
|
63
|
+
- [API](#api)
|
|
64
|
+
- [Troubleshooting](#troubleshooting)
|
|
65
|
+
- [Security notes](#security-notes)
|
|
66
|
+
- [Extending it](#extending-it)
|
|
67
|
+
- [Documentation](#documentation)
|
|
68
|
+
- [License](#license)
|
|
69
|
+
|
|
70
|
+
## What it works with
|
|
71
|
+
|
|
72
|
+
| runtime | `adapter` | to free memory, `lrd` has to |
|
|
73
|
+
| -------------------------------- | ----------- | ------------------------------------------------------- |
|
|
74
|
+
| [MTPLX](https://mtplx.com) | `mtplx` | stop the server, then restart it |
|
|
75
|
+
| [LM Studio](https://lmstudio.ai) | `lm-studio` | run `lms unload`; the server stays up |
|
|
76
|
+
| [oMLX](https://omlx.ai) | `omlx` | POST an unload; the server stays up |
|
|
77
|
+
| [Ollama](https://ollama.com) | `ollama` | POST `keep_alive: 0`; the server stays up (OpenAI only) |
|
|
78
|
+
| your own | `custom` | run the stop command you declare in YAML |
|
|
79
|
+
|
|
80
|
+
Coding agents: **[OpenCode](https://opencode.ai)**,
|
|
81
|
+
**[Claude Code](https://code.claude.com/docs/en/overview)** and
|
|
82
|
+
**[Codex](https://github.com/openai/codex)** — `lrd apply` writes
|
|
83
|
+
the gateway into each one's own configuration file, merging with what is already
|
|
84
|
+
there.
|
|
85
|
+
|
|
86
|
+
Four official backends, four different mechanisms — and one endpoint in front of them.
|
|
87
|
+
Anything not in that table can still be driven through the `custom` adapter,
|
|
88
|
+
declaratively, without writing any TypeScript; a first-class adapter is a
|
|
89
|
+
package plus one line in the composition root. Contributions welcome — see
|
|
90
|
+
[Extending it](#extending-it).
|
|
91
|
+
|
|
92
|
+
## Install
|
|
93
|
+
|
|
94
|
+
```bash
|
|
95
|
+
npm i -g llm-runtime-dock
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
Node 24 or newer. This puts `lrd` on your `PATH`.
|
|
99
|
+
|
|
100
|
+
The package is a single bundled file with no runtime dependencies, so the
|
|
101
|
+
install pulls nothing else in.
|
|
102
|
+
|
|
103
|
+
## Quickstart
|
|
104
|
+
|
|
105
|
+
Five commands from nothing to a working setup:
|
|
106
|
+
|
|
107
|
+
```bash
|
|
108
|
+
lrd probe # what is running on this machine right now?
|
|
109
|
+
lrd probe mtplx --save # write those findings into a config file
|
|
110
|
+
lrd doctor # check the result, and say which file it read
|
|
111
|
+
lrd serve # start the gateway on 127.0.0.1:8787
|
|
112
|
+
lrd apply opencode # point a coding agent at it
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
### 1. See what you have
|
|
116
|
+
|
|
117
|
+
`lrd probe` needs no gateway and no configuration — it is how a configuration
|
|
118
|
+
gets written in the first place. It asks each backend what it serves:
|
|
119
|
+
|
|
120
|
+
```text
|
|
121
|
+
mtplx http://127.0.0.1:8000 running, serving 1 model (mtplx)
|
|
122
|
+
Qwen3.8-27B (context 131072)
|
|
123
|
+
lm-studio http://127.0.0.1:1234 not running
|
|
124
|
+
omlx http://127.0.0.1:8000 not this runtime
|
|
125
|
+
a server answered /health, but it does not serve oMLX /v1/models/status
|
|
126
|
+
ollama http://127.0.0.1:11434 running, serving 0 models, 2 installed
|
|
127
|
+
|
|
128
|
+
# configuration for the discovered models
|
|
129
|
+
# write it with: lrd probe <adapter> --save
|
|
130
|
+
runtimes:
|
|
131
|
+
mtplx:
|
|
132
|
+
adapter: mtplx
|
|
133
|
+
port: 8000
|
|
134
|
+
models:
|
|
135
|
+
Qwen3.8-27B:
|
|
136
|
+
runtime: mtplx
|
|
137
|
+
backend_model: Qwen3.8-27B
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
Every outcome is useful. **Not running** is an answer, not an error.
|
|
141
|
+
**Running, credential required** means the server answered health and then
|
|
142
|
+
refused discovery — give it a key with `--api-key-env <VAR>`. **Not installed**
|
|
143
|
+
means the runtime's own executable is missing, so nothing could be started here.
|
|
144
|
+
For most runtimes that is answered without asking anything over HTTP, because the
|
|
145
|
+
executable is how the gateway drives them at all; Ollama is driven entirely over
|
|
146
|
+
HTTP, so it asks first and a remote server is reported as running whether or not
|
|
147
|
+
`ollama` is on this machine. **Not this runtime** means something else has that
|
|
148
|
+
port: MTPLX and oMLX both default to 8000, while Ollama defaults to 11434. Each
|
|
149
|
+
adapter asks a server whether it is its own before claiming it, so discovery
|
|
150
|
+
never writes a config aimed at the wrong backend.
|
|
151
|
+
|
|
152
|
+
Two more flags for the awkward cases: `--interactive` asks per runtime what to
|
|
153
|
+
probe and with what, and `--start` brings up a backend that supports daemon-style
|
|
154
|
+
startup and probes it again. Ollama's foreground `serve` command is started by
|
|
155
|
+
`lrd serve`, not by `lrd probe --start`.
|
|
156
|
+
|
|
157
|
+
Nothing is written without `--save`, so that block is yours to copy and edit.
|
|
158
|
+
|
|
159
|
+
### 2. Save it
|
|
160
|
+
|
|
161
|
+
```bash
|
|
162
|
+
lrd probe mtplx --save # or --dry-run first, to see the result
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
`--save` prints the path it is about to write, backs up the previous file,
|
|
166
|
+
refreshes only that runtime's entries, and leaves every other runtime's alone.
|
|
167
|
+
Name collisions fail the command and leave the file untouched rather than
|
|
168
|
+
picking a winner for you.
|
|
169
|
+
|
|
170
|
+
Re-probing later is safe. An entry that comes back keeps everything you put on
|
|
171
|
+
it — options, `extra_args`, `auth`, your comments — and only the fields
|
|
172
|
+
discovery owns (`adapter`, `host` and `port` on the runtime; `runtime` and
|
|
173
|
+
`backend_model` on the model) are brought up to date. An entry the probe _doesn't_ find is usually just an idle backend, so it
|
|
174
|
+
is never deleted behind your back: `--save` lists them and asks, enter keeps
|
|
175
|
+
them all, and `--force` removes them without asking. With no terminal
|
|
176
|
+
(`--json`, a pipe, CI) they are kept and named.
|
|
177
|
+
|
|
178
|
+
### 3. Check it
|
|
179
|
+
|
|
180
|
+
```bash
|
|
181
|
+
lrd doctor
|
|
182
|
+
```
|
|
183
|
+
|
|
184
|
+
`doctor` names the file it loaded, then validates adapters, options, reserved
|
|
185
|
+
arguments, executables on your `PATH`, endpoint connectivity, and every
|
|
186
|
+
`agents:` role. Run it whenever something looks wrong — it is the fastest way to
|
|
187
|
+
find out what.
|
|
188
|
+
|
|
189
|
+
### 4. Run it
|
|
190
|
+
|
|
191
|
+
```bash
|
|
192
|
+
lrd serve
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
The gateway listens on `127.0.0.1:8787` and stays in the foreground. Runtime
|
|
196
|
+
process output goes to its stdout, so this is the window to watch.
|
|
197
|
+
|
|
198
|
+
## Configuration
|
|
199
|
+
|
|
200
|
+
The first existing match wins:
|
|
201
|
+
|
|
202
|
+
1. `--config <path>`
|
|
203
|
+
2. `$LRD_CONFIG`
|
|
204
|
+
3. `./llm-runtime-dock.yaml` (project-local)
|
|
205
|
+
4. `$XDG_CONFIG_HOME/llm-runtime-dock/config.yaml`
|
|
206
|
+
5. `~/.config/llm-runtime-dock/config.yaml`
|
|
207
|
+
|
|
208
|
+
Every command reports which file it actually loaded. A complete, commented
|
|
209
|
+
example is in [`examples/config.yaml`](examples/config.yaml).
|
|
210
|
+
|
|
211
|
+
Two maps. `runtimes:` declares the servers this machine can talk to; `models:`
|
|
212
|
+
declares what to ask them for. The rule is one sentence: anything describing the
|
|
213
|
+
**server** is a runtime field, anything describing the **model** is a model
|
|
214
|
+
field.
|
|
215
|
+
|
|
216
|
+
> **Make sure your backends run on different ports.** Two adapters can name the
|
|
217
|
+
> same default — MTPLX and oMLX both land on `8000` — and only one server can
|
|
218
|
+
> have it. Move one in that backend's own settings and record it under
|
|
219
|
+
> `runtimes:`.
|
|
220
|
+
|
|
221
|
+
```yaml
|
|
222
|
+
server:
|
|
223
|
+
host: 127.0.0.1
|
|
224
|
+
port: 8787
|
|
225
|
+
|
|
226
|
+
runtimes:
|
|
227
|
+
mtplx:
|
|
228
|
+
adapter: mtplx
|
|
229
|
+
port: 8001 # moved: oMLX keeps the 8000 default
|
|
230
|
+
omlx:
|
|
231
|
+
adapter: omlx
|
|
232
|
+
port: 8000
|
|
233
|
+
auth:
|
|
234
|
+
api_key_env: OMLX_API_KEY # a variable name, never a value
|
|
235
|
+
lmstudio:
|
|
236
|
+
adapter: lm-studio
|
|
237
|
+
port: 1234
|
|
238
|
+
ollama:
|
|
239
|
+
adapter: ollama
|
|
240
|
+
port: 11434
|
|
241
|
+
|
|
242
|
+
models:
|
|
243
|
+
coding-quality:
|
|
244
|
+
runtime: mtplx
|
|
245
|
+
backend_model: Qwen3.8-27B
|
|
246
|
+
options:
|
|
247
|
+
reasoning: 'on'
|
|
248
|
+
reasoning_effort: high
|
|
249
|
+
context_window: 131072
|
|
250
|
+
max_tokens: 32768
|
|
251
|
+
extra_args: [--batching-preset, agent]
|
|
252
|
+
|
|
253
|
+
# Same runtime as coding-quality: one server, one resident model at a time.
|
|
254
|
+
coding-fast:
|
|
255
|
+
runtime: mtplx
|
|
256
|
+
backend_model: Qwen3.6-35B-A3B
|
|
257
|
+
|
|
258
|
+
omlx-small:
|
|
259
|
+
runtime: omlx
|
|
260
|
+
backend_model: llama-3b
|
|
261
|
+
|
|
262
|
+
ollama-coder:
|
|
263
|
+
runtime: ollama
|
|
264
|
+
backend_model: llama3.2
|
|
265
|
+
|
|
266
|
+
agents:
|
|
267
|
+
claude:
|
|
268
|
+
opus: coding-quality
|
|
269
|
+
sonnet: coding-quality
|
|
270
|
+
haiku: coding-fast
|
|
271
|
+
codex:
|
|
272
|
+
model: coding-quality
|
|
273
|
+
reasoning_effort: high
|
|
274
|
+
opencode:
|
|
275
|
+
default: coding-quality
|
|
276
|
+
```
|
|
277
|
+
|
|
278
|
+
That is the whole shape. [`examples/config.yaml`](examples/config.yaml) is the
|
|
279
|
+
same thing with every field commented, plus LM Studio, Ollama and a `custom` runtime.
|
|
280
|
+
|
|
281
|
+
Each agent gets the roles its own format has: `claude` names Anthropic's three,
|
|
282
|
+
`codex` and `opencode` a single model. Every value is a logical model id from
|
|
283
|
+
`models:` — except `reasoning_effort`, which is a codex setting rather than a
|
|
284
|
+
role, and is written through to its config as-is.
|
|
285
|
+
|
|
286
|
+
Clients send the logical id and never see the backend model:
|
|
287
|
+
|
|
288
|
+
```json
|
|
289
|
+
{ "model": "coding-quality", "messages": [] }
|
|
290
|
+
```
|
|
291
|
+
|
|
292
|
+
### Fields you will actually set
|
|
293
|
+
|
|
294
|
+
On a runtime:
|
|
295
|
+
|
|
296
|
+
| field | meaning |
|
|
297
|
+
| ---------------------------- | ---------------------------------------------------------------------------- |
|
|
298
|
+
| the key | the runtime's name; what a model's `runtime:` points at |
|
|
299
|
+
| `adapter` | `mtplx`, `lm-studio`, `omlx`, `ollama` or `custom` |
|
|
300
|
+
| `port` (and optional `host`) | where it serves. Omit it to take the adapter's own default |
|
|
301
|
+
| `options` | server-scoped options only — oMLX flags or Ollama environment-backed options |
|
|
302
|
+
| `auth` | optional upstream credential, as `api_key_env` or `api_key_file` |
|
|
303
|
+
|
|
304
|
+
On a model:
|
|
305
|
+
|
|
306
|
+
| field | meaning |
|
|
307
|
+
| ----------------- | ------------------------------------------------------------------------- |
|
|
308
|
+
| the key | the logical model id clients send and `/v1/models` returns |
|
|
309
|
+
| `runtime` | the runtime that serves it. One that is not declared fails validation |
|
|
310
|
+
| `backend_model` | the model reference the runtime itself understands |
|
|
311
|
+
| `name` (optional) | name written into agent configs; falls back to `backend_model` when unset |
|
|
312
|
+
| `options` | model-scoped launch options, validated against that adapter's schema |
|
|
313
|
+
| `extra_args` | raw argv escape hatch for flags the adapter does not name |
|
|
314
|
+
| `keep_resident` | keep this model loaded instead of unloading it on a switch (below) |
|
|
315
|
+
| `disabled` | keep the entry in the file and never serve it (below) |
|
|
316
|
+
|
|
317
|
+
Put a server-scoped option on a model, or a model-scoped one on a runtime, and
|
|
318
|
+
validation says so and names the block it belongs in. `extra_args` needs a
|
|
319
|
+
command line to escape into, so the Ollama adapter refuses it outright: `ollama
|
|
320
|
+
serve` takes no flags, and everything it documents is an environment variable
|
|
321
|
+
reached through `options`.
|
|
322
|
+
|
|
323
|
+
`options` describe **how the runtime process starts**. Context and output limits
|
|
324
|
+
are start-time flags: the gateway never rewrites the fields of your request,
|
|
325
|
+
injects defaults, or adds anything you did not send. Changing an option means
|
|
326
|
+
changing configuration, and the runtime restarts on the next switch.
|
|
327
|
+
|
|
328
|
+
Quote `on`/`off`/`yes`/`no` option values. Some YAML parsers read them as
|
|
329
|
+
booleans and these runtimes expect the literal strings; validation rejects a
|
|
330
|
+
boolean rather than quietly coercing it.
|
|
331
|
+
|
|
332
|
+
### Arguments you cannot set
|
|
333
|
+
|
|
334
|
+
Some arguments belong to the gateway and are rejected rather than silently
|
|
335
|
+
merged: `--host`, `--port`, `--model`, served-id flags like `--identifier`,
|
|
336
|
+
`--api-key`, and idle-unload flags such as LM Studio's `--ttl` — an idle
|
|
337
|
+
auto-unload would drop the model behind the gateway's back and leave its view of
|
|
338
|
+
the world wrong. The check covers `extra_args` too, and the error names the
|
|
339
|
+
field to use instead. The full table, with the reasoning per flag, is in
|
|
340
|
+
[§12](docs/05-configuration.md#configuration).
|
|
341
|
+
|
|
342
|
+
### One model at a time
|
|
343
|
+
|
|
344
|
+
At most one model is resident in memory at any moment, across every adapter.
|
|
345
|
+
That is the default, and it is not something you tune by accident: one machine
|
|
346
|
+
has one memory pool.
|
|
347
|
+
|
|
348
|
+
For you that means two entries sharing a port is fine and normal — switching
|
|
349
|
+
between them restarts the runtime with different flags. It also means the
|
|
350
|
+
gateway will stop a single-model server it did not start, when that is the only
|
|
351
|
+
way to free memory, and will say so in the log:
|
|
352
|
+
|
|
353
|
+
```text
|
|
354
|
+
releasing foreign mtplx server on :8001 to free the resident slot
|
|
355
|
+
```
|
|
356
|
+
|
|
357
|
+
The reasoning is in [§8 of the specification](docs/03-lifecycle.md#lifecycle).
|
|
358
|
+
|
|
359
|
+
### Keeping one model always loaded
|
|
360
|
+
|
|
361
|
+
Sometimes one model should never leave memory — a small one a coding agent
|
|
362
|
+
reaches for constantly, to summarise a file or scan a repository, while the large
|
|
363
|
+
model you are actually working with stays put. Reloading a 27B every time the
|
|
364
|
+
agent wants a two-line summary costs more than the summary.
|
|
365
|
+
|
|
366
|
+
`keep_resident: true` on a model entry does that:
|
|
367
|
+
|
|
368
|
+
```yaml
|
|
369
|
+
runtimes:
|
|
370
|
+
mtplx: { adapter: mtplx, port: 8001 }
|
|
371
|
+
mtplx-resident: { adapter: mtplx, port: 8002 }
|
|
372
|
+
|
|
373
|
+
models:
|
|
374
|
+
coding-quality:
|
|
375
|
+
runtime: mtplx
|
|
376
|
+
backend_model: Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality
|
|
377
|
+
summariser:
|
|
378
|
+
runtime: mtplx-resident
|
|
379
|
+
backend_model: Youssofal/Qwen3.5-4B-MTPLX-Optimized-Quality
|
|
380
|
+
keep_resident: true
|
|
381
|
+
```
|
|
382
|
+
|
|
383
|
+
The small model is loaded the first time something asks for it and is not
|
|
384
|
+
unloaded again for as long as the gateway runs, and switching to it no longer
|
|
385
|
+
evicts `coding-quality` either. Both stay in memory; requests still take turns,
|
|
386
|
+
one model answering at a time,
|
|
387
|
+
because they share a GPU. So this buys you **no reloading**, not parallelism.
|
|
388
|
+
|
|
389
|
+
Note the second runtime, on its own port. MTPLX serves one model per server, so
|
|
390
|
+
freeing its memory means stopping the server — which would take the kept model
|
|
391
|
+
with it. A kept model on MTPLX or a custom runtime therefore needs a runtime of
|
|
392
|
+
its own; put another entry on `mtplx-resident` and validation refuses the config
|
|
393
|
+
and tells you why. LM Studio, oMLX and Ollama hold several models on one server,
|
|
394
|
+
so there a kept entry can share the runtime with anything. Ollama requires
|
|
395
|
+
`max_loaded_models: 2` or greater when a runtime uses `keep_resident: true`.
|
|
396
|
+
|
|
397
|
+
Because that second runtime exists to hold one chosen model, you usually want it
|
|
398
|
+
out of discovery as well. `lrd probe` reads MTPLX's installed catalogue, which
|
|
399
|
+
belongs to the installation rather than to a port, so it would keep offering to
|
|
400
|
+
add every other model to a server that must serve exactly one:
|
|
401
|
+
|
|
402
|
+
```yaml
|
|
403
|
+
runtimes:
|
|
404
|
+
mtplx-resident:
|
|
405
|
+
adapter: mtplx
|
|
406
|
+
port: 8002
|
|
407
|
+
discovery: false
|
|
408
|
+
```
|
|
409
|
+
|
|
410
|
+
`lrd probe` then skips it entirely — and says so — while `lrd probe
|
|
411
|
+
mtplx-resident` still probes it when you ask by name. Nothing else changes: the
|
|
412
|
+
gateway starts, stops and proxies to it exactly as before.
|
|
413
|
+
|
|
414
|
+
`lrd status` shows what is loaded and which model is currently answering:
|
|
415
|
+
|
|
416
|
+
```text
|
|
417
|
+
resident: coding-quality (mtplx, spawned)
|
|
418
|
+
kept: summariser (mtplx, ready, 0 active)
|
|
419
|
+
serving: summariser
|
|
420
|
+
```
|
|
421
|
+
|
|
422
|
+
Its lifetime is the gateway's. Stopping `lrd serve` releases everything loaded,
|
|
423
|
+
kept entries included — nothing would be left to free the memory otherwise, and
|
|
424
|
+
a model outliving the process that loaded it is exactly the leak this gateway
|
|
425
|
+
exists to prevent.
|
|
426
|
+
|
|
427
|
+
Nothing measures what this costs while it runs — that budget is yours. `lrd
|
|
428
|
+
doctor` warns if more than one entry is kept.
|
|
429
|
+
|
|
430
|
+
### Models you keep configured but never serve
|
|
431
|
+
|
|
432
|
+
A probe reports everything a backend has, and not all of it is something a
|
|
433
|
+
client should ask for: embedding models, a sidecar the backend loads on its own,
|
|
434
|
+
or simply a model that is installed and not wanted right now.
|
|
435
|
+
|
|
436
|
+
Deleting the entry does not help — the next `lrd probe --save` finds the model
|
|
437
|
+
again and writes it straight back. `disabled: true` is the way to say it:
|
|
438
|
+
|
|
439
|
+
```yaml
|
|
440
|
+
models:
|
|
441
|
+
embeddings:
|
|
442
|
+
runtime: lmstudio
|
|
443
|
+
backend_model: nomic-embed-text-v1.5
|
|
444
|
+
disabled: true
|
|
445
|
+
```
|
|
446
|
+
|
|
447
|
+
The entry stays in the file, `lrd models` and `lrd doctor` still show it, and a
|
|
448
|
+
re-probe refreshes it and leaves the flag alone. Everything else treats it as if
|
|
449
|
+
it were not there: it is missing from `/v1/models`, a request naming it gets a
|
|
450
|
+
404 saying it is disabled, no runtime is ever started for it, and `lrd apply`
|
|
451
|
+
does not write it into any coding agent's configuration.
|
|
452
|
+
|
|
453
|
+
Pointing an `agents:` role at a disabled model is not a config error — the rest
|
|
454
|
+
of the CLI keeps working — but `lrd apply` refuses to write that mapping, and
|
|
455
|
+
`lrd doctor` reports it first.
|
|
456
|
+
|
|
457
|
+
## Coding agents
|
|
458
|
+
|
|
459
|
+
```bash
|
|
460
|
+
lrd apply opencode
|
|
461
|
+
lrd apply claude
|
|
462
|
+
lrd apply codex
|
|
463
|
+
lrd apply --all # every agent named in agents:
|
|
464
|
+
lrd apply claude --opus coding-fast # override the mapping for one run
|
|
465
|
+
lrd apply codex --dry-run # print the result, write nothing
|
|
466
|
+
```
|
|
467
|
+
|
|
468
|
+
`apply` reads the `agents:` section and writes each agent's own configuration
|
|
469
|
+
file. It **merges** rather than replaces: every key it does not own survives —
|
|
470
|
+
your other providers, plugin lists, schema references and settings. It backs the
|
|
471
|
+
file up first, reports the path it wrote, and is idempotent.
|
|
472
|
+
|
|
473
|
+
It never writes a secret value. Where a format supports an environment-variable
|
|
474
|
+
reference it writes that; where none exists it refuses and tells you why.
|
|
475
|
+
|
|
476
|
+
Declaring the mapping in `agents:` rather than in flags is what makes this
|
|
477
|
+
survivable: after `lrd probe --save` changes your model list, re-running
|
|
478
|
+
`lrd apply --all` is one command with nothing to remember.
|
|
479
|
+
|
|
480
|
+
**You do not need an `agents:` block to start.** Applying an agent the config
|
|
481
|
+
does not mention still works:
|
|
482
|
+
|
|
483
|
+
- `lrd apply opencode` registers the provider with all your models and **leaves
|
|
484
|
+
your default model alone** — if you already picked one in OpenCode, it stays;
|
|
485
|
+
- `lrd apply claude` and `lrd apply codex` ask which model each role should use,
|
|
486
|
+
offering your configured model ids, and print the `agents:` block to paste in
|
|
487
|
+
so they stop asking.
|
|
488
|
+
|
|
489
|
+
A role can be left unset, and choosing nothing writes nothing. In a pipeline or
|
|
490
|
+
under `--json` there is no terminal to ask on, so the command fails with a clear
|
|
491
|
+
message rather than hanging.
|
|
492
|
+
|
|
493
|
+
`--all` is the exception: it applies to the agents named in `agents:` and no
|
|
494
|
+
others, so it never stops to ask. Name an agent explicitly to set it up without
|
|
495
|
+
a mapping.
|
|
496
|
+
|
|
497
|
+
An entry can only fill an Anthropic role if its runtime serves that surface.
|
|
498
|
+
`apply` and `doctor` both check before a client ever hits it.
|
|
499
|
+
|
|
500
|
+
## Commands
|
|
501
|
+
|
|
502
|
+
```text
|
|
503
|
+
lrd serve start the gateway
|
|
504
|
+
lrd status what a running gateway is doing
|
|
505
|
+
lrd probe [runtime] ask the backends what they serve
|
|
506
|
+
lrd apply <agent> | --all write the gateway into an agent's config
|
|
507
|
+
lrd doctor validate config, adapters, executables, roles
|
|
508
|
+
lrd models configured logical ids
|
|
509
|
+
lrd runtimes declared runtimes, their models and state
|
|
510
|
+
lrd switch <model> make a model resident, through the scheduler
|
|
511
|
+
lrd logs <runtime> captured runtime process output
|
|
512
|
+
```
|
|
513
|
+
|
|
514
|
+
`--config <path>`, `--json` and `--no-color` work on every command;
|
|
515
|
+
`--endpoint <url>` on the ones that talk to a running gateway.
|
|
516
|
+
|
|
517
|
+
Output is coloured only when `lrd` owns the terminal: `--no-color`, `NO_COLOR`,
|
|
518
|
+
`--json`, `TERM=dumb` and any pipe or redirect turn it off, and `FORCE_COLOR`
|
|
519
|
+
turns it back on where there is no terminal to detect. The colour is decoration
|
|
520
|
+
only — strip the escapes and the bytes match the uncoloured run.
|
|
521
|
+
|
|
522
|
+
`status` and `switch` are HTTP clients of a running gateway. With no gateway they
|
|
523
|
+
say so plainly and exit non-zero — never a stack trace:
|
|
524
|
+
|
|
525
|
+
```text
|
|
526
|
+
error [GATEWAY_NOT_RUNNING] gateway not running at http://127.0.0.1:8787
|
|
527
|
+
hint: start it with `lrd serve`
|
|
528
|
+
```
|
|
529
|
+
|
|
530
|
+
`lrd apply --all` applies to every agent named in `agents:`. An agent that is not
|
|
531
|
+
installed is reported and skipped so the others still get written; the command
|
|
532
|
+
still exits non-zero.
|
|
533
|
+
|
|
534
|
+
Runtime process output is held by the `serve` process that spawned it, so
|
|
535
|
+
`lrd logs` run as a separate command has nothing to show. Follow `lrd serve`'s
|
|
536
|
+
own output instead. There is deliberately no `--follow`: streaming logs out of
|
|
537
|
+
the gateway would mean a new HTTP endpoint, and the surface is kept to the one in
|
|
538
|
+
the specification.
|
|
539
|
+
|
|
540
|
+
## API
|
|
541
|
+
|
|
542
|
+
```http
|
|
543
|
+
GET /health
|
|
544
|
+
GET /v1/models
|
|
545
|
+
POST /v1/chat/completions
|
|
546
|
+
POST /v1/messages
|
|
547
|
+
POST /v1/messages/count_tokens
|
|
548
|
+
GET /status loopback only
|
|
549
|
+
POST /switch loopback only
|
|
550
|
+
```
|
|
551
|
+
|
|
552
|
+
Both supported surfaces are **proxied, never translated** — compatible backends
|
|
553
|
+
serve them themselves. MTPLX, LM Studio and oMLX expose both OpenAI and Anthropic;
|
|
554
|
+
Ollama exposes OpenAI only. SSE streaming passes through unaltered, and tool calls
|
|
555
|
+
work because nothing rewrites them.
|
|
556
|
+
|
|
557
|
+
Request bodies are touched in exactly one place: the `model` field, swapped from
|
|
558
|
+
your logical id to the one the backend answers to. Not `tools`, not
|
|
559
|
+
`tool_choice`, not `messages`. `GET /v1/models` returns your configured logical
|
|
560
|
+
ids whether or not they are loaded, and your `Authorization` header goes
|
|
561
|
+
upstream verbatim — a per-model `auth:` block fills in only when you sent none,
|
|
562
|
+
and the gateway never generates a key.
|
|
563
|
+
|
|
564
|
+
The full surface, both protocols, response-header handling and `/switch`
|
|
565
|
+
semantics are in [§14](docs/06-gateway-api.md#gateway-api).
|
|
566
|
+
|
|
567
|
+
## Troubleshooting
|
|
568
|
+
|
|
569
|
+
**`lrd status` says the gateway is not running.** It is an HTTP client; start
|
|
570
|
+
`lrd serve`, or point it somewhere else with `--endpoint`.
|
|
571
|
+
|
|
572
|
+
**A command reads a different config than you are editing.** Every command prints
|
|
573
|
+
the file it loaded — check that line first. A project-local
|
|
574
|
+
`./llm-runtime-dock.yaml` that did not exist when you last saved means the save
|
|
575
|
+
landed in your home config instead.
|
|
576
|
+
|
|
577
|
+
**`RUNTIME_MODEL_MISMATCH`.** The runtime came up but is not serving the
|
|
578
|
+
`backend_model` this entry asked for. Run `lrd probe <adapter>` to see what it
|
|
579
|
+
actually serves, and fix `backend_model` to match.
|
|
580
|
+
|
|
581
|
+
**`RUNTIME_MODEL_PINNED`.** A pinned model is holding memory and cannot be
|
|
582
|
+
evicted, so the switch would have left two models resident. Unpin it in the
|
|
583
|
+
runtime — the error names which one.
|
|
584
|
+
|
|
585
|
+
**`UPSTREAM_SURFACE_UNSUPPORTED`.** You sent an Anthropic request to a runtime
|
|
586
|
+
that only serves OpenAI. A `custom` entry serves OpenAI only unless it opts in
|
|
587
|
+
with `surfaces: [openai, anthropic]`.
|
|
588
|
+
|
|
589
|
+
**`RUNTIME_OPTION_RESERVED` at startup.** An option or `extra_args` entry uses a
|
|
590
|
+
flag the gateway owns. The message names the canonical field to use instead.
|
|
591
|
+
|
|
592
|
+
**LM Studio: a server answers on a different port.** `lms server start --port`
|
|
593
|
+
only decides the port when no server is running; otherwise LM Studio reuses the
|
|
594
|
+
port from its last start. The error names both ports — either point the entry at
|
|
595
|
+
the running one or restart LM Studio's server.
|
|
596
|
+
|
|
597
|
+
**A switch hangs, then fails with `RUNTIME_SLOT_BUSY`.** Something is holding a
|
|
598
|
+
stream open. Cancel the client, or raise the drain timeout.
|
|
599
|
+
|
|
600
|
+
**Tool calls arrive as text instead of being executed.** Your agent shows raw
|
|
601
|
+
`<function_calls>` or similar markup in the reply. The gateway does not touch
|
|
602
|
+
`tools` or `tool_choice`, so this is the runtime rendering tools in a format its
|
|
603
|
+
own parser did not read back. For MTPLX the relevant launch options are
|
|
604
|
+
`tool_prompt_mode`, `chat_template_profile` and `agent_rewrites` — none of them
|
|
605
|
+
is set for you, because the right value depends on the model. To confirm where
|
|
606
|
+
the problem is rather than guess, capture the traffic instead:
|
|
607
|
+
|
|
608
|
+
```bash
|
|
609
|
+
lrd serve --debug # one file per run; path printed on start
|
|
610
|
+
lrd serve --debug --debug-dir ./cap # or put it where you choose
|
|
611
|
+
```
|
|
612
|
+
|
|
613
|
+
The capture is a tee — it cannot alter what is sent — and it records both hops,
|
|
614
|
+
so you can compare what reached the runtime with what the client got. Header
|
|
615
|
+
credentials are redacted; **bodies are not**, so the file holds whole
|
|
616
|
+
conversations. Treat it as sensitive. The event format and the `jq` recipes for
|
|
617
|
+
reading one are in [§14](docs/06-gateway-api.md#gateway-api); the flags, their
|
|
618
|
+
defaults and the `LRD_DEBUG` environment variables are in
|
|
619
|
+
[§27](docs/10-cli.md#cli).
|
|
620
|
+
|
|
621
|
+
When in doubt, `lrd doctor` checks the whole configuration in one pass and names
|
|
622
|
+
the file it read.
|
|
623
|
+
|
|
624
|
+
## Security notes
|
|
625
|
+
|
|
626
|
+
The gateway binds to `127.0.0.1` and has no authentication. `/status` and
|
|
627
|
+
`/switch` are lifecycle controls: loopback-only, and refused outright when the
|
|
628
|
+
gateway is bound to a non-loopback address.
|
|
629
|
+
|
|
630
|
+
Lifecycle commands are **trusted configuration only**. Nothing derived from an
|
|
631
|
+
HTTP request is ever interpolated into a command — a request selects which
|
|
632
|
+
configured entry runs, never what it runs.
|
|
633
|
+
|
|
634
|
+
Commands are argv arrays, used verbatim:
|
|
635
|
+
|
|
636
|
+
```yaml
|
|
637
|
+
command: [some-runtime, serve, --model, some-model-7b] # preferred
|
|
638
|
+
```
|
|
639
|
+
|
|
640
|
+
Shell execution exists but must be opted into explicitly:
|
|
641
|
+
|
|
642
|
+
```yaml
|
|
643
|
+
process:
|
|
644
|
+
start:
|
|
645
|
+
command: ['some-runtime serve | tee log']
|
|
646
|
+
shell: true
|
|
647
|
+
```
|
|
648
|
+
|
|
649
|
+
With `shell: true` the string is handed to the system shell, so every shell
|
|
650
|
+
metacharacter in it is live — quoting, globbing, redirection, command
|
|
651
|
+
substitution. Use it only for a command you wrote yourself, and prefer the argv
|
|
652
|
+
form, which cannot be reinterpreted.
|
|
653
|
+
|
|
654
|
+
`lrd apply` never writes a credential into a coding agent's configuration, only
|
|
655
|
+
an environment-variable reference.
|
|
656
|
+
|
|
657
|
+
The normative rules are in [§28](docs/01-overview.md#security).
|
|
658
|
+
|
|
659
|
+
## Extending it
|
|
660
|
+
|
|
661
|
+
Contributions are welcome, and there are two ways in — one of which is not code
|
|
662
|
+
at all.
|
|
663
|
+
|
|
664
|
+
**A runtime, without writing TypeScript.** Any runtime the built-in adapters do
|
|
665
|
+
not cover can be described declaratively: the start and stop commands, the
|
|
666
|
+
health and model-discovery URLs, and the endpoint to proxy to. That is the
|
|
667
|
+
`custom` adapter, and it is a `runtimes:` entry like any other. See `legacy`
|
|
668
|
+
in [`examples/config.yaml`](examples/config.yaml), and
|
|
669
|
+
[§11](docs/04-adapters.md#custom-adapter) for the full field list.
|
|
670
|
+
|
|
671
|
+
**A first-class adapter or agent.** Worth it when a runtime needs real logic —
|
|
672
|
+
mapped launch options, its own readiness check, a CLI to drive. The workspace is
|
|
673
|
+
laid out so this stays small:
|
|
674
|
+
|
|
675
|
+
1. Create the package under `packages/runtimes/<name>` or
|
|
676
|
+
`packages/agents/<name>`, implementing `RuntimeAdapter` or `AgentIntegration`
|
|
677
|
+
from `@llm-runtime-dock/core`.
|
|
678
|
+
2. Add it as `workspace:*` where it is used, and run `pnpm install`. There is no
|
|
679
|
+
build graph to update — it is derived from those dependencies.
|
|
680
|
+
3. Register it in the composition root, `src/index.ts`, which is a list:
|
|
681
|
+
|
|
682
|
+
```ts
|
|
683
|
+
export const defaultAdapters = (logger?: Logger): RuntimeAdapter[] => {
|
|
684
|
+
return [
|
|
685
|
+
createMtplxAdapter(logger ? { logger } : {}),
|
|
686
|
+
createLmStudioAdapter(logger ? { logger } : {}),
|
|
687
|
+
// your adapter here
|
|
688
|
+
];
|
|
689
|
+
};
|
|
690
|
+
```
|
|
691
|
+
|
|
692
|
+
That is the whole registration. Nothing in `packages/core` knows an adapter
|
|
693
|
+
exists — the composition root is the only file that imports one, which is what
|
|
694
|
+
keeps a new backend from touching the core at all.
|
|
695
|
+
|
|
696
|
+
The `RuntimeAdapter` and `AgentIntegration` contracts are in
|
|
697
|
+
[§7](docs/04-adapters.md#runtime-adapter-interface) and
|
|
698
|
+
[§23](docs/08-agents.md#agent-integrations); the build, test and style
|
|
699
|
+
workflow is in [Architecture](docs/02-architecture.md#developer-experience).
|
|
700
|
+
|
|
701
|
+
## Documentation
|
|
702
|
+
|
|
703
|
+
The complete specification lives in [`docs/`](docs/README.md) — the design, the
|
|
704
|
+
invariants, and the reasoning behind the decisions that are easy to get wrong.
|
|
705
|
+
|
|
706
|
+
## License
|
|
707
|
+
|
|
708
|
+
MIT — see [LICENSE](LICENSE).
|