llm-runtime-dock 0.1.0 → 0.1.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +103 -11
- package/dist/index.js +7086 -6616
- package/package.json +26 -26
package/README.md
CHANGED
|
@@ -222,6 +222,7 @@ field.
|
|
|
222
222
|
server:
|
|
223
223
|
host: 127.0.0.1
|
|
224
224
|
port: 8787
|
|
225
|
+
idle_unload: 60m # unload after an hour of quiet; 0 to never
|
|
225
226
|
|
|
226
227
|
runtimes:
|
|
227
228
|
mtplx:
|
|
@@ -291,6 +292,14 @@ Clients send the logical id and never see the backend model:
|
|
|
291
292
|
|
|
292
293
|
### Fields you will actually set
|
|
293
294
|
|
|
295
|
+
On the gateway itself:
|
|
296
|
+
|
|
297
|
+
| field | meaning |
|
|
298
|
+
| --------------- | ------------------------------------------------------------------------- |
|
|
299
|
+
| `host` / `port` | where `lrd` listens. Defaults to `127.0.0.1:8787` |
|
|
300
|
+
| `idle_unload` | unload the loaded model after this long with nothing to do (below) |
|
|
301
|
+
| `auth` | require an API key of clients, as `api_key_env` or `api_key_file` (below) |
|
|
302
|
+
|
|
294
303
|
On a runtime:
|
|
295
304
|
|
|
296
305
|
| field | meaning |
|
|
@@ -335,7 +344,9 @@ Some arguments belong to the gateway and are rejected rather than silently
|
|
|
335
344
|
merged: `--host`, `--port`, `--model`, served-id flags like `--identifier`,
|
|
336
345
|
`--api-key`, and idle-unload flags such as LM Studio's `--ttl` — an idle
|
|
337
346
|
auto-unload would drop the model behind the gateway's back and leave its view of
|
|
338
|
-
the world wrong.
|
|
347
|
+
the world wrong. That is about _who owns the timer_, not about the idea: the
|
|
348
|
+
gateway runs one itself, in front of its own bookkeeping, and you set it with
|
|
349
|
+
`idle_unload` (below). The check covers `extra_args` too, and the error names the
|
|
339
350
|
field to use instead. The full table, with the reasoning per flag, is in
|
|
340
351
|
[§12](docs/05-configuration.md#configuration).
|
|
341
352
|
|
|
@@ -356,6 +367,51 @@ releasing foreign mtplx server on :8001 to free the resident slot
|
|
|
356
367
|
|
|
357
368
|
The reasoning is in [§8 of the specification](docs/03-lifecycle.md#lifecycle).
|
|
358
369
|
|
|
370
|
+
### Unloading when you stop using it
|
|
371
|
+
|
|
372
|
+
`lrd serve` is a long-lived process, and until something asks for a different
|
|
373
|
+
model nothing frees the one it is holding. Leave the gateway running after a
|
|
374
|
+
morning's work and that 27B is still in memory at midnight.
|
|
375
|
+
|
|
376
|
+
So it unloads by itself after an hour of quiet:
|
|
377
|
+
|
|
378
|
+
```yaml
|
|
379
|
+
server:
|
|
380
|
+
idle_unload: 60m # 0 never unloads
|
|
381
|
+
```
|
|
382
|
+
|
|
383
|
+
Write it as `60m`, `1h`, `90s`, `3600000ms`, or a bare number of milliseconds.
|
|
384
|
+
`lrd serve --idle-unload 90s` overrides it for one run, and `lrd serve` prints
|
|
385
|
+
the window it is using on startup:
|
|
386
|
+
|
|
387
|
+
```text
|
|
388
|
+
idle: 1h then unload
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
The next request loads the model again exactly as the first one did — the cost
|
|
392
|
+
is one reload after an hour of not working, which is the trade the default is
|
|
393
|
+
picked for. Turn it off with `0` if you would rather keep the model warm
|
|
394
|
+
indefinitely.
|
|
395
|
+
|
|
396
|
+
Two things it will not do:
|
|
397
|
+
|
|
398
|
+
- **it never unloads a `keep_resident` model.** That flag means "for as long as
|
|
399
|
+
the gateway runs", and an idle spell is not the gateway stopping;
|
|
400
|
+
- **it will not stop a server it did not start.** If `lrd` attached to an MTPLX
|
|
401
|
+
or custom server you launched yourself, stopping it is the only way to free
|
|
402
|
+
that memory — and doing that because `lrd` went quiet is not its call. It says
|
|
403
|
+
so in the log and leaves it alone. On LM Studio, oMLX and Ollama the question
|
|
404
|
+
never arises: unloading one model leaves the server, and everyone else on it,
|
|
405
|
+
untouched.
|
|
406
|
+
|
|
407
|
+
When it has unloaded something, `lrd status` says so rather than just showing an
|
|
408
|
+
empty slot:
|
|
409
|
+
|
|
410
|
+
```text
|
|
411
|
+
resident: none
|
|
412
|
+
released: coding-quality (idle, stop_server)
|
|
413
|
+
```
|
|
414
|
+
|
|
359
415
|
### Keeping one model always loaded
|
|
360
416
|
|
|
361
417
|
Sometimes one model should never leave memory — a small one a coding agent
|
|
@@ -419,8 +475,8 @@ kept: summariser (mtplx, ready, 0 active)
|
|
|
419
475
|
serving: summariser
|
|
420
476
|
```
|
|
421
477
|
|
|
422
|
-
Its lifetime is the gateway's
|
|
423
|
-
kept entries included — nothing would be left to free the memory otherwise, and
|
|
478
|
+
Its lifetime is the gateway's — the idle unload above does not touch it either.
|
|
479
|
+
Stopping `lrd serve` releases everything loaded, kept entries included — nothing would be left to free the memory otherwise, and
|
|
424
480
|
a model outliving the process that loaded it is exactly the leak this gateway
|
|
425
481
|
exists to prevent.
|
|
426
482
|
|
|
@@ -557,9 +613,11 @@ work because nothing rewrites them.
|
|
|
557
613
|
Request bodies are touched in exactly one place: the `model` field, swapped from
|
|
558
614
|
your logical id to the one the backend answers to. Not `tools`, not
|
|
559
615
|
`tool_choice`, not `messages`. `GET /v1/models` returns your configured logical
|
|
560
|
-
ids whether or not they are loaded, and
|
|
561
|
-
upstream verbatim
|
|
562
|
-
|
|
616
|
+
ids whether or not they are loaded, and — when the gateway itself requires no
|
|
617
|
+
key — your `Authorization` header goes upstream verbatim; a per-model `auth:`
|
|
618
|
+
block fills in only when you sent none. If `server.auth` is set (below), that
|
|
619
|
+
header authenticates to the gateway instead and is not forwarded; the runtime's
|
|
620
|
+
own `auth:` block is what reaches it, resolved server-side.
|
|
563
621
|
|
|
564
622
|
The full surface, both protocols, response-header handling and `/switch`
|
|
565
623
|
semantics are in [§14](docs/06-gateway-api.md#gateway-api).
|
|
@@ -623,9 +681,41 @@ the file it read.
|
|
|
623
681
|
|
|
624
682
|
## Security notes
|
|
625
683
|
|
|
626
|
-
The gateway binds to `127.0.0.1` and
|
|
627
|
-
`/switch` are lifecycle controls: loopback-only,
|
|
628
|
-
gateway is bound to a non-loopback address
|
|
684
|
+
The gateway binds to `127.0.0.1` and requires no authentication by default.
|
|
685
|
+
`/status` and `/switch` are lifecycle controls: loopback-only, refused outright
|
|
686
|
+
when the gateway is bound to a non-loopback address, and that rule holds
|
|
687
|
+
regardless of whether an API key is configured — a leaked key must not hand out
|
|
688
|
+
remote process control on top of remote inference.
|
|
689
|
+
|
|
690
|
+
**Requiring an API key.** Generate one and wire it into `server.auth`:
|
|
691
|
+
|
|
692
|
+
```bash
|
|
693
|
+
lrd key generate # writes server.auth.api_key_file, plus a secret file next to the config
|
|
694
|
+
lrd key generate --env # prints the key once; writes server.auth.api_key_env: LRD_API_KEY instead
|
|
695
|
+
```
|
|
696
|
+
|
|
697
|
+
or set it by hand:
|
|
698
|
+
|
|
699
|
+
```yaml
|
|
700
|
+
server:
|
|
701
|
+
auth:
|
|
702
|
+
api_key_env: LRD_API_KEY # or: api_key_file: ~/.config/llm-runtime-dock/api_key
|
|
703
|
+
```
|
|
704
|
+
|
|
705
|
+
With no `server.auth` at all, the `LRD_API_KEY` environment variable is checked
|
|
706
|
+
automatically — useful for a Docker or systemd deployment that already injects
|
|
707
|
+
secrets that way, with no YAML edit needed. Once a key resolves, every route but
|
|
708
|
+
`/health` requires it (`Authorization: Bearer <key>` or `x-api-key: <key>`);
|
|
709
|
+
`lrd status`/`lrd switch` send it automatically, and `lrd apply` wires it into
|
|
710
|
+
whichever agent you configure — Claude Code's `settings.json` gets the literal
|
|
711
|
+
value (it has no reference syntax), OpenCode and Codex get an environment
|
|
712
|
+
variable reference instead.
|
|
713
|
+
|
|
714
|
+
An `api_key_env`/`api_key_file` that is set but resolves to nothing is a hard
|
|
715
|
+
failure at `lrd serve`: an admin who configured a key source must never end up
|
|
716
|
+
with a gateway that silently started unauthenticated. Binding to a non-loopback
|
|
717
|
+
address with no key at all is only a warning, from `doctor` and the `serve`
|
|
718
|
+
startup banner — not a hard stop.
|
|
629
719
|
|
|
630
720
|
Lifecycle commands are **trusted configuration only**. Nothing derived from an
|
|
631
721
|
HTTP request is ever interpolated into a command — a request selects which
|
|
@@ -651,8 +741,10 @@ metacharacter in it is live — quoting, globbing, redirection, command
|
|
|
651
741
|
substitution. Use it only for a command you wrote yourself, and prefer the argv
|
|
652
742
|
form, which cannot be reinterpreted.
|
|
653
743
|
|
|
654
|
-
`lrd apply` never writes
|
|
655
|
-
an environment-variable reference
|
|
744
|
+
`lrd apply` never writes an _upstream_ credential into a coding agent's
|
|
745
|
+
configuration, only an environment-variable reference — except the gateway's
|
|
746
|
+
own key, which some agent formats (Claude Code) can only receive as a literal
|
|
747
|
+
value, since that agent must present it to reach the gateway at all.
|
|
656
748
|
|
|
657
749
|
The normative rules are in [§28](docs/01-overview.md#security).
|
|
658
750
|
|