llm-runtime-dock 0.1.0 → 0.1.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +103 -11
  2. package/dist/index.js +7086 -6616
  3. package/package.json +26 -26
package/README.md CHANGED
@@ -222,6 +222,7 @@ field.
222
222
  server:
223
223
  host: 127.0.0.1
224
224
  port: 8787
225
+ idle_unload: 60m # unload after an hour of quiet; 0 to never
225
226
 
226
227
  runtimes:
227
228
  mtplx:
@@ -291,6 +292,14 @@ Clients send the logical id and never see the backend model:
291
292
 
292
293
  ### Fields you will actually set
293
294
 
295
+ On the gateway itself:
296
+
297
+ | field | meaning |
298
+ | --------------- | ------------------------------------------------------------------------- |
299
+ | `host` / `port` | where `lrd` listens. Defaults to `127.0.0.1:8787` |
300
+ | `idle_unload` | unload the loaded model after this long with nothing to do (below) |
301
+ | `auth` | require an API key of clients, as `api_key_env` or `api_key_file` (below) |
302
+
294
303
  On a runtime:
295
304
 
296
305
  | field | meaning |
@@ -335,7 +344,9 @@ Some arguments belong to the gateway and are rejected rather than silently
335
344
  merged: `--host`, `--port`, `--model`, served-id flags like `--identifier`,
336
345
  `--api-key`, and idle-unload flags such as LM Studio's `--ttl` — an idle
337
346
  auto-unload would drop the model behind the gateway's back and leave its view of
338
- the world wrong. The check covers `extra_args` too, and the error names the
347
+ the world wrong. That is about _who owns the timer_, not about the idea: the
348
+ gateway runs one itself, in front of its own bookkeeping, and you set it with
349
+ `idle_unload` (below). The check covers `extra_args` too, and the error names the
339
350
  field to use instead. The full table, with the reasoning per flag, is in
340
351
  [§12](docs/05-configuration.md#configuration).
341
352
 
@@ -356,6 +367,51 @@ releasing foreign mtplx server on :8001 to free the resident slot
356
367
 
357
368
  The reasoning is in [§8 of the specification](docs/03-lifecycle.md#lifecycle).
358
369
 
370
+ ### Unloading when you stop using it
371
+
372
+ `lrd serve` is a long-lived process, and until something asks for a different
373
+ model nothing frees the one it is holding. Leave the gateway running after a
374
+ morning's work and that 27B is still in memory at midnight.
375
+
376
+ So it unloads by itself after an hour of quiet:
377
+
378
+ ```yaml
379
+ server:
380
+ idle_unload: 60m # 0 never unloads
381
+ ```
382
+
383
+ Write it as `60m`, `1h`, `90s`, `3600000ms`, or a bare number of milliseconds.
384
+ `lrd serve --idle-unload 90s` overrides it for one run, and `lrd serve` prints
385
+ the window it is using on startup:
386
+
387
+ ```text
388
+ idle: 1h then unload
389
+ ```
390
+
391
+ The next request loads the model again exactly as the first one did — the cost
392
+ is one reload after an hour of not working, which is the trade the default is
393
+ picked for. Turn it off with `0` if you would rather keep the model warm
394
+ indefinitely.
395
+
396
+ Two things it will not do:
397
+
398
+ - **it never unloads a `keep_resident` model.** That flag means "for as long as
399
+ the gateway runs", and an idle spell is not the gateway stopping;
400
+ - **it will not stop a server it did not start.** If `lrd` attached to an MTPLX
401
+ or custom server you launched yourself, stopping it is the only way to free
402
+ that memory — and doing that because `lrd` went quiet is not its call. It says
403
+ so in the log and leaves it alone. On LM Studio, oMLX and Ollama the question
404
+ never arises: unloading one model leaves the server, and everyone else on it,
405
+ untouched.
406
+
407
+ When it has unloaded something, `lrd status` says so rather than just showing an
408
+ empty slot:
409
+
410
+ ```text
411
+ resident: none
412
+ released: coding-quality (idle, stop_server)
413
+ ```
414
+
359
415
  ### Keeping one model always loaded
360
416
 
361
417
  Sometimes one model should never leave memory — a small one a coding agent
@@ -419,8 +475,8 @@ kept: summariser (mtplx, ready, 0 active)
419
475
  serving: summariser
420
476
  ```
421
477
 
422
- Its lifetime is the gateway's. Stopping `lrd serve` releases everything loaded,
423
- kept entries included — nothing would be left to free the memory otherwise, and
478
+ Its lifetime is the gateway's — the idle unload above does not touch it either.
479
+ Stopping `lrd serve` releases everything loaded, kept entries included — nothing would be left to free the memory otherwise, and
424
480
  a model outliving the process that loaded it is exactly the leak this gateway
425
481
  exists to prevent.
426
482
 
@@ -557,9 +613,11 @@ work because nothing rewrites them.
557
613
  Request bodies are touched in exactly one place: the `model` field, swapped from
558
614
  your logical id to the one the backend answers to. Not `tools`, not
559
615
  `tool_choice`, not `messages`. `GET /v1/models` returns your configured logical
560
- ids whether or not they are loaded, and your `Authorization` header goes
561
- upstream verbatim — a per-model `auth:` block fills in only when you sent none,
562
- and the gateway never generates a key.
616
+ ids whether or not they are loaded, and — when the gateway itself requires no
617
+ key — your `Authorization` header goes upstream verbatim; a per-model `auth:`
618
+ block fills in only when you sent none. If `server.auth` is set (below), that
619
+ header authenticates to the gateway instead and is not forwarded; the runtime's
620
+ own `auth:` block is what reaches it, resolved server-side.
563
621
 
564
622
  The full surface, both protocols, response-header handling and `/switch`
565
623
  semantics are in [§14](docs/06-gateway-api.md#gateway-api).
@@ -623,9 +681,41 @@ the file it read.
623
681
 
624
682
  ## Security notes
625
683
 
626
- The gateway binds to `127.0.0.1` and has no authentication. `/status` and
627
- `/switch` are lifecycle controls: loopback-only, and refused outright when the
628
- gateway is bound to a non-loopback address.
684
+ The gateway binds to `127.0.0.1` and requires no authentication by default.
685
+ `/status` and `/switch` are lifecycle controls: loopback-only, refused outright
686
+ when the gateway is bound to a non-loopback address, and that rule holds
687
+ regardless of whether an API key is configured — a leaked key must not hand out
688
+ remote process control on top of remote inference.
689
+
690
+ **Requiring an API key.** Generate one and wire it into `server.auth`:
691
+
692
+ ```bash
693
+ lrd key generate # writes server.auth.api_key_file, plus a secret file next to the config
694
+ lrd key generate --env # prints the key once; writes server.auth.api_key_env: LRD_API_KEY instead
695
+ ```
696
+
697
+ or set it by hand:
698
+
699
+ ```yaml
700
+ server:
701
+ auth:
702
+ api_key_env: LRD_API_KEY # or: api_key_file: ~/.config/llm-runtime-dock/api_key
703
+ ```
704
+
705
+ With no `server.auth` at all, the `LRD_API_KEY` environment variable is checked
706
+ automatically — useful for a Docker or systemd deployment that already injects
707
+ secrets that way, with no YAML edit needed. Once a key resolves, every route but
708
+ `/health` requires it (`Authorization: Bearer <key>` or `x-api-key: <key>`);
709
+ `lrd status`/`lrd switch` send it automatically, and `lrd apply` wires it into
710
+ whichever agent you configure — Claude Code's `settings.json` gets the literal
711
+ value (it has no reference syntax), OpenCode and Codex get an environment
712
+ variable reference instead.
713
+
714
+ An `api_key_env`/`api_key_file` that is set but resolves to nothing is a hard
715
+ failure at `lrd serve`: an admin who configured a key source must never end up
716
+ with a gateway that silently started unauthenticated. Binding to a non-loopback
717
+ address with no key at all is only a warning, from `doctor` and the `serve`
718
+ startup banner — not a hard stop.
629
719
 
630
720
  Lifecycle commands are **trusted configuration only**. Nothing derived from an
631
721
  HTTP request is ever interpolated into a command — a request selects which
@@ -651,8 +741,10 @@ metacharacter in it is live — quoting, globbing, redirection, command
651
741
  substitution. Use it only for a command you wrote yourself, and prefer the argv
652
742
  form, which cannot be reinterpreted.
653
743
 
654
- `lrd apply` never writes a credential into a coding agent's configuration, only
655
- an environment-variable reference.
744
+ `lrd apply` never writes an _upstream_ credential into a coding agent's
745
+ configuration, only an environment-variable reference — except the gateway's
746
+ own key, which some agent formats (Claude Code) can only receive as a literal
747
+ value, since that agent must present it to reach the gateway at all.
656
748
 
657
749
  The normative rules are in [§28](docs/01-overview.md#security).
658
750