llm-runtime-dock 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md ADDED
@@ -0,0 +1,708 @@
1
+ # llm-runtime-dock
2
+
3
+ One local endpoint for your coding agents, with the right model loaded behind
4
+ it — whatever runtime that model happens to live in.
5
+
6
+ `lrd` was built with AI assistance — Claude Code and OpenCode, working against
7
+ both local and remote models. The agent instruction files it was developed with
8
+ ship with the repository (`AGENTS.md`, `CLAUDE.md`, `.claude/skills/`), where
9
+ they double as the contributor extension procedures.
10
+
11
+ ## Why this exists
12
+
13
+ If you run models locally, you probably run more than one runtime: MTPLX for
14
+ one model, LM Studio for another, something else for the third. Each has its own
15
+ CLI, its own model names, its own idea of when it is ready — and none of them
16
+ agrees with the others about what "switch to another model" even means.
17
+
18
+ MTPLX serves one model per process, so switching means stopping the server and
19
+ starting it again with different flags. LM Studio keeps one shared server and
20
+ loads models into it, so switching is `lms unload` and then `lms load`. Your
21
+ machine, meanwhile, has exactly one pool of memory: leaving the old model
22
+ resident while the next one loads is how you run out of it.
23
+
24
+ Which leaves you doing the orchestration by hand — stop that one, start this
25
+ one, wait until it is really ready, hope nothing was mid-request — and then
26
+ editing three coding-agent config files because the model name changed.
27
+
28
+ `lrd` is the thing that does that for you. Your agents talk to one address, and
29
+ it starts, stops and switches the runtime underneath so that the model a request
30
+ asked for is the model that answers it: one switch at a time, in-flight requests
31
+ drained first, a live stream never cut. And since it already has to know what
32
+ your backends serve, it will write that into your agents' configuration too.
33
+
34
+ ```text
35
+ OpenCode · Claude Code · Codex
36
+ │ OpenAI or Anthropic API
37
+ ▼
38
+ llm-runtime-dock · http://127.0.0.1:8787
39
+ │ CLI / process control, plus documented HTTP lifecycle
40
+ ▼
41
+ MTPLX · LM Studio · oMLX · Ollama · your own runtime
42
+ ```
43
+
44
+ Three commands, and none of them asks you to learn a new model format:
45
+
46
+ ```bash
47
+ lrd probe mtplx --save # ask the backend what it serves, save it as config
48
+ lrd serve # one endpoint, the right model behind it
49
+ lrd apply claude # point a coding agent at it
50
+ ```
51
+
52
+ It is a gateway and orchestrator, not an inference engine. Everything it knows
53
+ about your machine comes from one YAML file you can read.
54
+
55
+ **Contents**
56
+
57
+ - [What it works with](#what-it-works-with)
58
+ - [Install](#install)
59
+ - [Quickstart](#quickstart)
60
+ - [Configuration](#configuration)
61
+ - [Coding agents](#coding-agents)
62
+ - [Commands](#commands)
63
+ - [API](#api)
64
+ - [Troubleshooting](#troubleshooting)
65
+ - [Security notes](#security-notes)
66
+ - [Extending it](#extending-it)
67
+ - [Documentation](#documentation)
68
+ - [License](#license)
69
+
70
+ ## What it works with
71
+
72
+ | runtime | `adapter` | to free memory, `lrd` has to |
73
+ | -------------------------------- | ----------- | ------------------------------------------------------- |
74
+ | [MTPLX](https://mtplx.com) | `mtplx` | stop the server, then restart it |
75
+ | [LM Studio](https://lmstudio.ai) | `lm-studio` | run `lms unload`; the server stays up |
76
+ | [oMLX](https://omlx.ai) | `omlx` | POST an unload; the server stays up |
77
+ | [Ollama](https://ollama.com) | `ollama` | POST `keep_alive: 0`; the server stays up (OpenAI only) |
78
+ | your own | `custom` | run the stop command you declare in YAML |
79
+
80
+ Coding agents: **[OpenCode](https://opencode.ai)**,
81
+ **[Claude Code](https://code.claude.com/docs/en/overview)** and
82
+ **[Codex](https://github.com/openai/codex)** — `lrd apply` writes
83
+ the gateway into each one's own configuration file, merging with what is already
84
+ there.
85
+
86
+ Four official backends, four different mechanisms — and one endpoint in front of them.
87
+ Anything not in that table can still be driven through the `custom` adapter,
88
+ declaratively, without writing any TypeScript; a first-class adapter is a
89
+ package plus one line in the composition root. Contributions welcome — see
90
+ [Extending it](#extending-it).
91
+
92
+ ## Install
93
+
94
+ ```bash
95
+ npm i -g llm-runtime-dock
96
+ ```
97
+
98
+ Node 24 or newer. This puts `lrd` on your `PATH`.
99
+
100
+ The package is a single bundled file with no runtime dependencies, so the
101
+ install pulls nothing else in.
102
+
103
+ ## Quickstart
104
+
105
+ Five commands from nothing to a working setup:
106
+
107
+ ```bash
108
+ lrd probe # what is running on this machine right now?
109
+ lrd probe mtplx --save # write those findings into a config file
110
+ lrd doctor # check the result, and say which file it read
111
+ lrd serve # start the gateway on 127.0.0.1:8787
112
+ lrd apply opencode # point a coding agent at it
113
+ ```
114
+
115
+ ### 1. See what you have
116
+
117
+ `lrd probe` needs no gateway and no configuration — it is how a configuration
118
+ gets written in the first place. It asks each backend what it serves:
119
+
120
+ ```text
121
+ mtplx http://127.0.0.1:8000 running, serving 1 model (mtplx)
122
+ Qwen3.8-27B (context 131072)
123
+ lm-studio http://127.0.0.1:1234 not running
124
+ omlx http://127.0.0.1:8000 not this runtime
125
+ a server answered /health, but it does not serve oMLX /v1/models/status
126
+ ollama http://127.0.0.1:11434 running, serving 0 models, 2 installed
127
+
128
+ # configuration for the discovered models
129
+ # write it with: lrd probe <adapter> --save
130
+ runtimes:
131
+ mtplx:
132
+ adapter: mtplx
133
+ port: 8000
134
+ models:
135
+ Qwen3.8-27B:
136
+ runtime: mtplx
137
+ backend_model: Qwen3.8-27B
138
+ ```
139
+
140
+ Every outcome is useful. **Not running** is an answer, not an error.
141
+ **Running, credential required** means the server answered health and then
142
+ refused discovery — give it a key with `--api-key-env <VAR>`. **Not installed**
143
+ means the runtime's own executable is missing, so nothing could be started here.
144
+ For most runtimes that is answered without asking anything over HTTP, because the
145
+ executable is how the gateway drives them at all; Ollama is driven entirely over
146
+ HTTP, so it asks first and a remote server is reported as running whether or not
147
+ `ollama` is on this machine. **Not this runtime** means something else has that
148
+ port: MTPLX and oMLX both default to 8000, while Ollama defaults to 11434. Each
149
+ adapter asks a server whether it is its own before claiming it, so discovery
150
+ never writes a config aimed at the wrong backend.
151
+
152
+ Two more flags for the awkward cases: `--interactive` asks per runtime what to
153
+ probe and with what, and `--start` brings up a backend that supports daemon-style
154
+ startup and probes it again. Ollama's foreground `serve` command is started by
155
+ `lrd serve`, not by `lrd probe --start`.
156
+
157
+ Nothing is written without `--save`, so that block is yours to copy and edit.
158
+
159
+ ### 2. Save it
160
+
161
+ ```bash
162
+ lrd probe mtplx --save # or --dry-run first, to see the result
163
+ ```
164
+
165
+ `--save` prints the path it is about to write, backs up the previous file,
166
+ refreshes only that runtime's entries, and leaves every other runtime's alone.
167
+ Name collisions fail the command and leave the file untouched rather than
168
+ picking a winner for you.
169
+
170
+ Re-probing later is safe. An entry that comes back keeps everything you put on
171
+ it — options, `extra_args`, `auth`, your comments — and only the fields
172
+ discovery owns (`adapter`, `host` and `port` on the runtime; `runtime` and
173
+ `backend_model` on the model) are brought up to date. An entry the probe _doesn't_ find is usually just an idle backend, so it
174
+ is never deleted behind your back: `--save` lists them and asks, enter keeps
175
+ them all, and `--force` removes them without asking. With no terminal
176
+ (`--json`, a pipe, CI) they are kept and named.
177
+
178
+ ### 3. Check it
179
+
180
+ ```bash
181
+ lrd doctor
182
+ ```
183
+
184
+ `doctor` names the file it loaded, then validates adapters, options, reserved
185
+ arguments, executables on your `PATH`, endpoint connectivity, and every
186
+ `agents:` role. Run it whenever something looks wrong — it is the fastest way to
187
+ find out what.
188
+
189
+ ### 4. Run it
190
+
191
+ ```bash
192
+ lrd serve
193
+ ```
194
+
195
+ The gateway listens on `127.0.0.1:8787` and stays in the foreground. Runtime
196
+ process output goes to its stdout, so this is the window to watch.
197
+
198
+ ## Configuration
199
+
200
+ The first existing match wins:
201
+
202
+ 1. `--config <path>`
203
+ 2. `$LRD_CONFIG`
204
+ 3. `./llm-runtime-dock.yaml` (project-local)
205
+ 4. `$XDG_CONFIG_HOME/llm-runtime-dock/config.yaml`
206
+ 5. `~/.config/llm-runtime-dock/config.yaml`
207
+
208
+ Every command reports which file it actually loaded. A complete, commented
209
+ example is in [`examples/config.yaml`](examples/config.yaml).
210
+
211
+ Two maps. `runtimes:` declares the servers this machine can talk to; `models:`
212
+ declares what to ask them for. The rule is one sentence: anything describing the
213
+ **server** is a runtime field, anything describing the **model** is a model
214
+ field.
215
+
216
+ > **Make sure your backends run on different ports.** Two adapters can name the
217
+ > same default — MTPLX and oMLX both land on `8000` — and only one server can
218
+ > have it. Move one in that backend's own settings and record it under
219
+ > `runtimes:`.
220
+
221
+ ```yaml
222
+ server:
223
+ host: 127.0.0.1
224
+ port: 8787
225
+
226
+ runtimes:
227
+ mtplx:
228
+ adapter: mtplx
229
+ port: 8001 # moved: oMLX keeps the 8000 default
230
+ omlx:
231
+ adapter: omlx
232
+ port: 8000
233
+ auth:
234
+ api_key_env: OMLX_API_KEY # a variable name, never a value
235
+ lmstudio:
236
+ adapter: lm-studio
237
+ port: 1234
238
+ ollama:
239
+ adapter: ollama
240
+ port: 11434
241
+
242
+ models:
243
+ coding-quality:
244
+ runtime: mtplx
245
+ backend_model: Qwen3.8-27B
246
+ options:
247
+ reasoning: 'on'
248
+ reasoning_effort: high
249
+ context_window: 131072
250
+ max_tokens: 32768
251
+ extra_args: [--batching-preset, agent]
252
+
253
+ # Same runtime as coding-quality: one server, one resident model at a time.
254
+ coding-fast:
255
+ runtime: mtplx
256
+ backend_model: Qwen3.6-35B-A3B
257
+
258
+ omlx-small:
259
+ runtime: omlx
260
+ backend_model: llama-3b
261
+
262
+ ollama-coder:
263
+ runtime: ollama
264
+ backend_model: llama3.2
265
+
266
+ agents:
267
+ claude:
268
+ opus: coding-quality
269
+ sonnet: coding-quality
270
+ haiku: coding-fast
271
+ codex:
272
+ model: coding-quality
273
+ reasoning_effort: high
274
+ opencode:
275
+ default: coding-quality
276
+ ```
277
+
278
+ That is the whole shape. [`examples/config.yaml`](examples/config.yaml) is the
279
+ same thing with every field commented, plus LM Studio, Ollama and a `custom` runtime.
280
+
281
+ Each agent gets the roles its own format has: `claude` names Anthropic's three,
282
+ `codex` and `opencode` a single model. Every value is a logical model id from
283
+ `models:` — except `reasoning_effort`, which is a codex setting rather than a
284
+ role, and is written through to its config as-is.
285
+
286
+ Clients send the logical id and never see the backend model:
287
+
288
+ ```json
289
+ { "model": "coding-quality", "messages": [] }
290
+ ```
291
+
292
+ ### Fields you will actually set
293
+
294
+ On a runtime:
295
+
296
+ | field | meaning |
297
+ | ---------------------------- | ---------------------------------------------------------------------------- |
298
+ | the key | the runtime's name; what a model's `runtime:` points at |
299
+ | `adapter` | `mtplx`, `lm-studio`, `omlx`, `ollama` or `custom` |
300
+ | `port` (and optional `host`) | where it serves. Omit it to take the adapter's own default |
301
+ | `options` | server-scoped options only — oMLX flags or Ollama environment-backed options |
302
+ | `auth` | optional upstream credential, as `api_key_env` or `api_key_file` |
303
+
304
+ On a model:
305
+
306
+ | field | meaning |
307
+ | ----------------- | ------------------------------------------------------------------------- |
308
+ | the key | the logical model id clients send and `/v1/models` returns |
309
+ | `runtime` | the runtime that serves it. One that is not declared fails validation |
310
+ | `backend_model` | the model reference the runtime itself understands |
311
+ | `name` (optional) | name written into agent configs; falls back to `backend_model` when unset |
312
+ | `options` | model-scoped launch options, validated against that adapter's schema |
313
+ | `extra_args` | raw argv escape hatch for flags the adapter does not name |
314
+ | `keep_resident` | keep this model loaded instead of unloading it on a switch (below) |
315
+ | `disabled` | keep the entry in the file and never serve it (below) |
316
+
317
+ Put a server-scoped option on a model, or a model-scoped one on a runtime, and
318
+ validation says so and names the block it belongs in. `extra_args` needs a
319
+ command line to escape into, so the Ollama adapter refuses it outright: `ollama
320
+ serve` takes no flags, and everything it documents is an environment variable
321
+ reached through `options`.
322
+
323
+ `options` describe **how the runtime process starts**. Context and output limits
324
+ are start-time flags: the gateway never rewrites the fields of your request,
325
+ injects defaults, or adds anything you did not send. Changing an option means
326
+ changing configuration, and the runtime restarts on the next switch.
327
+
328
+ Quote `on`/`off`/`yes`/`no` option values. Some YAML parsers read them as
329
+ booleans and these runtimes expect the literal strings; validation rejects a
330
+ boolean rather than quietly coercing it.
331
+
332
+ ### Arguments you cannot set
333
+
334
+ Some arguments belong to the gateway and are rejected rather than silently
335
+ merged: `--host`, `--port`, `--model`, served-id flags like `--identifier`,
336
+ `--api-key`, and idle-unload flags such as LM Studio's `--ttl` — an idle
337
+ auto-unload would drop the model behind the gateway's back and leave its view of
338
+ the world wrong. The check covers `extra_args` too, and the error names the
339
+ field to use instead. The full table, with the reasoning per flag, is in
340
+ [§12](docs/05-configuration.md#configuration).
341
+
342
+ ### One model at a time
343
+
344
+ At most one model is resident in memory at any moment, across every adapter.
345
+ That is the default, and it is not something you tune by accident: one machine
346
+ has one memory pool.
347
+
348
+ For you that means two entries sharing a port is fine and normal — switching
349
+ between them restarts the runtime with different flags. It also means the
350
+ gateway will stop a single-model server it did not start, when that is the only
351
+ way to free memory, and will say so in the log:
352
+
353
+ ```text
354
+ releasing foreign mtplx server on :8001 to free the resident slot
355
+ ```
356
+
357
+ The reasoning is in [§8 of the specification](docs/03-lifecycle.md#lifecycle).
358
+
359
+ ### Keeping one model always loaded
360
+
361
+ Sometimes one model should never leave memory — a small one a coding agent
362
+ reaches for constantly, to summarise a file or scan a repository, while the large
363
+ model you are actually working with stays put. Reloading a 27B every time the
364
+ agent wants a two-line summary costs more than the summary.
365
+
366
+ `keep_resident: true` on a model entry does that:
367
+
368
+ ```yaml
369
+ runtimes:
370
+ mtplx: { adapter: mtplx, port: 8001 }
371
+ mtplx-resident: { adapter: mtplx, port: 8002 }
372
+
373
+ models:
374
+ coding-quality:
375
+ runtime: mtplx
376
+ backend_model: Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality
377
+ summariser:
378
+ runtime: mtplx-resident
379
+ backend_model: Youssofal/Qwen3.5-4B-MTPLX-Optimized-Quality
380
+ keep_resident: true
381
+ ```
382
+
383
+ The small model is loaded the first time something asks for it and is not
384
+ unloaded again for as long as the gateway runs, and switching to it no longer
385
+ evicts `coding-quality` either. Both stay in memory; requests still take turns,
386
+ one model answering at a time,
387
+ because they share a GPU. So this buys you **no reloading**, not parallelism.
388
+
389
+ Note the second runtime, on its own port. MTPLX serves one model per server, so
390
+ freeing its memory means stopping the server — which would take the kept model
391
+ with it. A kept model on MTPLX or a custom runtime therefore needs a runtime of
392
+ its own; put another entry on `mtplx-resident` and validation refuses the config
393
+ and tells you why. LM Studio, oMLX and Ollama hold several models on one server,
394
+ so there a kept entry can share the runtime with anything. Ollama requires
395
+ `max_loaded_models: 2` or greater when a runtime uses `keep_resident: true`.
396
+
397
+ Because that second runtime exists to hold one chosen model, you usually want it
398
+ out of discovery as well. `lrd probe` reads MTPLX's installed catalogue, which
399
+ belongs to the installation rather than to a port, so it would keep offering to
400
+ add every other model to a server that must serve exactly one:
401
+
402
+ ```yaml
403
+ runtimes:
404
+ mtplx-resident:
405
+ adapter: mtplx
406
+ port: 8002
407
+ discovery: false
408
+ ```
409
+
410
+ `lrd probe` then skips it entirely — and says so — while `lrd probe
411
+ mtplx-resident` still probes it when you ask by name. Nothing else changes: the
412
+ gateway starts, stops and proxies to it exactly as before.
413
+
414
+ `lrd status` shows what is loaded and which model is currently answering:
415
+
416
+ ```text
417
+ resident: coding-quality (mtplx, spawned)
418
+ kept: summariser (mtplx, ready, 0 active)
419
+ serving: summariser
420
+ ```
421
+
422
+ Its lifetime is the gateway's. Stopping `lrd serve` releases everything loaded,
423
+ kept entries included — nothing would be left to free the memory otherwise, and
424
+ a model outliving the process that loaded it is exactly the leak this gateway
425
+ exists to prevent.
426
+
427
+ Nothing measures what this costs while it runs — that budget is yours. `lrd
428
+ doctor` warns if more than one entry is kept.
429
+
430
+ ### Models you keep configured but never serve
431
+
432
+ A probe reports everything a backend has, and not all of it is something a
433
+ client should ask for: embedding models, a sidecar the backend loads on its own,
434
+ or simply a model that is installed and not wanted right now.
435
+
436
+ Deleting the entry does not help — the next `lrd probe --save` finds the model
437
+ again and writes it straight back. `disabled: true` is the way to say it:
438
+
439
+ ```yaml
440
+ models:
441
+ embeddings:
442
+ runtime: lmstudio
443
+ backend_model: nomic-embed-text-v1.5
444
+ disabled: true
445
+ ```
446
+
447
+ The entry stays in the file, `lrd models` and `lrd doctor` still show it, and a
448
+ re-probe refreshes it and leaves the flag alone. Everything else treats it as if
449
+ it were not there: it is missing from `/v1/models`, a request naming it gets a
450
+ 404 saying it is disabled, no runtime is ever started for it, and `lrd apply`
451
+ does not write it into any coding agent's configuration.
452
+
453
+ Pointing an `agents:` role at a disabled model is not a config error — the rest
454
+ of the CLI keeps working — but `lrd apply` refuses to write that mapping, and
455
+ `lrd doctor` reports it first.
456
+
457
+ ## Coding agents
458
+
459
+ ```bash
460
+ lrd apply opencode
461
+ lrd apply claude
462
+ lrd apply codex
463
+ lrd apply --all # every agent named in agents:
464
+ lrd apply claude --opus coding-fast # override the mapping for one run
465
+ lrd apply codex --dry-run # print the result, write nothing
466
+ ```
467
+
468
+ `apply` reads the `agents:` section and writes each agent's own configuration
469
+ file. It **merges** rather than replaces: every key it does not own survives —
470
+ your other providers, plugin lists, schema references and settings. It backs the
471
+ file up first, reports the path it wrote, and is idempotent.
472
+
473
+ It never writes a secret value. Where a format supports an environment-variable
474
+ reference it writes that; where none exists it refuses and tells you why.
475
+
476
+ Declaring the mapping in `agents:` rather than in flags is what makes this
477
+ survivable: after `lrd probe --save` changes your model list, re-running
478
+ `lrd apply --all` is one command with nothing to remember.
479
+
480
+ **You do not need an `agents:` block to start.** Applying an agent the config
481
+ does not mention still works:
482
+
483
+ - `lrd apply opencode` registers the provider with all your models and **leaves
484
+ your default model alone** — if you already picked one in OpenCode, it stays;
485
+ - `lrd apply claude` and `lrd apply codex` ask which model each role should use,
486
+ offering your configured model ids, and print the `agents:` block to paste in
487
+ so they stop asking.
488
+
489
+ A role can be left unset, and choosing nothing writes nothing. In a pipeline or
490
+ under `--json` there is no terminal to ask on, so the command fails with a clear
491
+ message rather than hanging.
492
+
493
+ `--all` is the exception: it applies to the agents named in `agents:` and no
494
+ others, so it never stops to ask. Name an agent explicitly to set it up without
495
+ a mapping.
496
+
497
+ An entry can only fill an Anthropic role if its runtime serves that surface.
498
+ `apply` and `doctor` both check before a client ever hits it.
499
+
500
+ ## Commands
501
+
502
+ ```text
503
+ lrd serve start the gateway
504
+ lrd status what a running gateway is doing
505
+ lrd probe [runtime] ask the backends what they serve
506
+ lrd apply <agent> | --all write the gateway into an agent's config
507
+ lrd doctor validate config, adapters, executables, roles
508
+ lrd models configured logical ids
509
+ lrd runtimes declared runtimes, their models and state
510
+ lrd switch <model> make a model resident, through the scheduler
511
+ lrd logs <runtime> captured runtime process output
512
+ ```
513
+
514
+ `--config <path>`, `--json` and `--no-color` work on every command;
515
+ `--endpoint <url>` on the ones that talk to a running gateway.
516
+
517
+ Output is coloured only when `lrd` owns the terminal: `--no-color`, `NO_COLOR`,
518
+ `--json`, `TERM=dumb` and any pipe or redirect turn it off, and `FORCE_COLOR`
519
+ turns it back on where there is no terminal to detect. The colour is decoration
520
+ only — strip the escapes and the bytes match the uncoloured run.
521
+
522
+ `status` and `switch` are HTTP clients of a running gateway. With no gateway they
523
+ say so plainly and exit non-zero — never a stack trace:
524
+
525
+ ```text
526
+ error [GATEWAY_NOT_RUNNING] gateway not running at http://127.0.0.1:8787
527
+ hint: start it with `lrd serve`
528
+ ```
529
+
530
+ `lrd apply --all` applies to every agent named in `agents:`. An agent that is not
531
+ installed is reported and skipped so the others still get written; the command
532
+ still exits non-zero.
533
+
534
+ Runtime process output is held by the `serve` process that spawned it, so
535
+ `lrd logs` run as a separate command has nothing to show. Follow `lrd serve`'s
536
+ own output instead. There is deliberately no `--follow`: streaming logs out of
537
+ the gateway would mean a new HTTP endpoint, and the surface is kept to the one in
538
+ the specification.
539
+
540
+ ## API
541
+
542
+ ```http
543
+ GET /health
544
+ GET /v1/models
545
+ POST /v1/chat/completions
546
+ POST /v1/messages
547
+ POST /v1/messages/count_tokens
548
+ GET /status loopback only
549
+ POST /switch loopback only
550
+ ```
551
+
552
+ Both supported surfaces are **proxied, never translated** — compatible backends
553
+ serve them themselves. MTPLX, LM Studio and oMLX expose both OpenAI and Anthropic;
554
+ Ollama exposes OpenAI only. SSE streaming passes through unaltered, and tool calls
555
+ work because nothing rewrites them.
556
+
557
+ Request bodies are touched in exactly one place: the `model` field, swapped from
558
+ your logical id to the one the backend answers to. Not `tools`, not
559
+ `tool_choice`, not `messages`. `GET /v1/models` returns your configured logical
560
+ ids whether or not they are loaded, and your `Authorization` header goes
561
+ upstream verbatim — a per-model `auth:` block fills in only when you sent none,
562
+ and the gateway never generates a key.
563
+
564
+ The full surface, both protocols, response-header handling and `/switch`
565
+ semantics are in [§14](docs/06-gateway-api.md#gateway-api).
566
+
567
+ ## Troubleshooting
568
+
569
+ **`lrd status` says the gateway is not running.** It is an HTTP client; start
570
+ `lrd serve`, or point it somewhere else with `--endpoint`.
571
+
572
+ **A command reads a different config than you are editing.** Every command prints
573
+ the file it loaded — check that line first. A project-local
574
+ `./llm-runtime-dock.yaml` that did not exist when you last saved means the save
575
+ landed in your home config instead.
576
+
577
+ **`RUNTIME_MODEL_MISMATCH`.** The runtime came up but is not serving the
578
+ `backend_model` this entry asked for. Run `lrd probe <adapter>` to see what it
579
+ actually serves, and fix `backend_model` to match.
580
+
581
+ **`RUNTIME_MODEL_PINNED`.** A pinned model is holding memory and cannot be
582
+ evicted, so the switch would have left two models resident. Unpin it in the
583
+ runtime — the error names which one.
584
+
585
+ **`UPSTREAM_SURFACE_UNSUPPORTED`.** You sent an Anthropic request to a runtime
586
+ that only serves OpenAI. A `custom` entry serves OpenAI only unless it opts in
587
+ with `surfaces: [openai, anthropic]`.
588
+
589
+ **`RUNTIME_OPTION_RESERVED` at startup.** An option or `extra_args` entry uses a
590
+ flag the gateway owns. The message names the canonical field to use instead.
591
+
592
+ **LM Studio: a server answers on a different port.** `lms server start --port`
593
+ only decides the port when no server is running; otherwise LM Studio reuses the
594
+ port from its last start. The error names both ports — either point the entry at
595
+ the running one or restart LM Studio's server.
596
+
597
+ **A switch hangs, then fails with `RUNTIME_SLOT_BUSY`.** Something is holding a
598
+ stream open. Cancel the client, or raise the drain timeout.
599
+
600
+ **Tool calls arrive as text instead of being executed.** Your agent shows raw
601
+ `<function_calls>` or similar markup in the reply. The gateway does not touch
602
+ `tools` or `tool_choice`, so this is the runtime rendering tools in a format its
603
+ own parser did not read back. For MTPLX the relevant launch options are
604
+ `tool_prompt_mode`, `chat_template_profile` and `agent_rewrites` — none of them
605
+ is set for you, because the right value depends on the model. To confirm where
606
+ the problem is rather than guess, capture the traffic instead:
607
+
608
+ ```bash
609
+ lrd serve --debug # one file per run; path printed on start
610
+ lrd serve --debug --debug-dir ./cap # or put it where you choose
611
+ ```
612
+
613
+ The capture is a tee — it cannot alter what is sent — and it records both hops,
614
+ so you can compare what reached the runtime with what the client got. Header
615
+ credentials are redacted; **bodies are not**, so the file holds whole
616
+ conversations. Treat it as sensitive. The event format and the `jq` recipes for
617
+ reading one are in [§14](docs/06-gateway-api.md#gateway-api); the flags, their
618
+ defaults and the `LRD_DEBUG` environment variables are in
619
+ [§27](docs/10-cli.md#cli).
620
+
621
+ When in doubt, `lrd doctor` checks the whole configuration in one pass and names
622
+ the file it read.
623
+
624
+ ## Security notes
625
+
626
+ The gateway binds to `127.0.0.1` and has no authentication. `/status` and
627
+ `/switch` are lifecycle controls: loopback-only, and refused outright when the
628
+ gateway is bound to a non-loopback address.
629
+
630
+ Lifecycle commands are **trusted configuration only**. Nothing derived from an
631
+ HTTP request is ever interpolated into a command — a request selects which
632
+ configured entry runs, never what it runs.
633
+
634
+ Commands are argv arrays, used verbatim:
635
+
636
+ ```yaml
637
+ command: [some-runtime, serve, --model, some-model-7b] # preferred
638
+ ```
639
+
640
+ Shell execution exists but must be opted into explicitly:
641
+
642
+ ```yaml
643
+ process:
644
+ start:
645
+ command: ['some-runtime serve | tee log']
646
+ shell: true
647
+ ```
648
+
649
+ With `shell: true` the string is handed to the system shell, so every shell
650
+ metacharacter in it is live — quoting, globbing, redirection, command
651
+ substitution. Use it only for a command you wrote yourself, and prefer the argv
652
+ form, which cannot be reinterpreted.
653
+
654
+ `lrd apply` never writes a credential into a coding agent's configuration, only
655
+ an environment-variable reference.
656
+
657
+ The normative rules are in [§28](docs/01-overview.md#security).
658
+
659
+ ## Extending it
660
+
661
+ Contributions are welcome, and there are two ways in — one of which is not code
662
+ at all.
663
+
664
+ **A runtime, without writing TypeScript.** Any runtime the built-in adapters do
665
+ not cover can be described declaratively: the start and stop commands, the
666
+ health and model-discovery URLs, and the endpoint to proxy to. That is the
667
+ `custom` adapter, and it is a `runtimes:` entry like any other. See `legacy`
668
+ in [`examples/config.yaml`](examples/config.yaml), and
669
+ [§11](docs/04-adapters.md#custom-adapter) for the full field list.
670
+
671
+ **A first-class adapter or agent.** Worth it when a runtime needs real logic —
672
+ mapped launch options, its own readiness check, a CLI to drive. The workspace is
673
+ laid out so this stays small:
674
+
675
+ 1. Create the package under `packages/runtimes/<name>` or
676
+ `packages/agents/<name>`, implementing `RuntimeAdapter` or `AgentIntegration`
677
+ from `@llm-runtime-dock/core`.
678
+ 2. Add it as `workspace:*` where it is used, and run `pnpm install`. There is no
679
+ build graph to update — it is derived from those dependencies.
680
+ 3. Register it in the composition root, `src/index.ts`, which is a list:
681
+
682
+ ```ts
683
+ export const defaultAdapters = (logger?: Logger): RuntimeAdapter[] => {
684
+ return [
685
+ createMtplxAdapter(logger ? { logger } : {}),
686
+ createLmStudioAdapter(logger ? { logger } : {}),
687
+ // your adapter here
688
+ ];
689
+ };
690
+ ```
691
+
692
+ That is the whole registration. Nothing in `packages/core` knows an adapter
693
+ exists — the composition root is the only file that imports one, which is what
694
+ keeps a new backend from touching the core at all.
695
+
696
+ The `RuntimeAdapter` and `AgentIntegration` contracts are in
697
+ [§7](docs/04-adapters.md#runtime-adapter-interface) and
698
+ [§23](docs/08-agents.md#agent-integrations); the build, test and style
699
+ workflow is in [Architecture](docs/02-architecture.md#developer-experience).
700
+
701
+ ## Documentation
702
+
703
+ The complete specification lives in [`docs/`](docs/README.md) — the design, the
704
+ invariants, and the reasoning behind the decisions that are easy to get wrong.
705
+
706
+ ## License
707
+
708
+ MIT — see [LICENSE](LICENSE).