infer-stack 0.6.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (44) hide show
  1. infer_stack/__init__.py +2 -0
  2. infer_stack/backends/__init__.py +7 -0
  3. infer_stack/backends/compose_renderer.py +243 -0
  4. infer_stack/backends/kubeai_renderer.py +202 -0
  5. infer_stack/benchmark.py +38 -0
  6. infer_stack/catalog.py +438 -0
  7. infer_stack/cli/__init__.py +169 -0
  8. infer_stack/cli/__main__.py +4 -0
  9. infer_stack/cli/commands_profile.py +467 -0
  10. infer_stack/cli/commands_runtime.py +719 -0
  11. infer_stack/cli/commands_smoke.py +691 -0
  12. infer_stack/cli/compose.py +755 -0
  13. infer_stack/cli/context.py +471 -0
  14. infer_stack/cli/options.py +134 -0
  15. infer_stack/cli/probes.py +178 -0
  16. infer_stack/config.py +450 -0
  17. infer_stack/contracts.py +223 -0
  18. infer_stack/diff_prompt.py +117 -0
  19. infer_stack/docker_utils.py +230 -0
  20. infer_stack/env_utils.py +97 -0
  21. infer_stack/experimental/model_catalog_discover.py +1155 -0
  22. infer_stack/experimental/model_memory_estimator.py +1264 -0
  23. infer_stack/experimental/stress_test_long_context.py +397 -0
  24. infer_stack/hardware.py +70 -0
  25. infer_stack/kubeai_ops.py +76 -0
  26. infer_stack/paths.py +87 -0
  27. infer_stack/profile_runtime.py +46 -0
  28. infer_stack/renderer.py +19 -0
  29. infer_stack/resolver.py +1092 -0
  30. infer_stack/templates/default-models.yaml +674 -0
  31. infer_stack/templates/default-ollama-models.yaml +31 -0
  32. infer_stack/templates/default-profiles.yaml +1731 -0
  33. infer_stack/templates/default-vllm-models.yaml +714 -0
  34. infer_stack/templates/docker-compose.yml.j2 +430 -0
  35. infer_stack/templates/litellm_config.yaml.j2 +44 -0
  36. infer_stack/templates/nginx.conf.j2 +84 -0
  37. infer_stack/tuning.py +3 -0
  38. infer_stack/validator.py +314 -0
  39. infer_stack/verification.py +46 -0
  40. infer_stack-0.6.0.dist-info/METADATA +1034 -0
  41. infer_stack-0.6.0.dist-info/RECORD +44 -0
  42. infer_stack-0.6.0.dist-info/WHEEL +5 -0
  43. infer_stack-0.6.0.dist-info/entry_points.txt +2 -0
  44. infer_stack-0.6.0.dist-info/top_level.txt +1 -0
@@ -0,0 +1,1034 @@
1
+ Metadata-Version: 2.4
2
+ Name: infer-stack
3
+ Version: 0.6.0
4
+ Summary: Profile-driven compose and KubeAI deployment compiler for inference stacks
5
+ Author-email: "jon.crall" <jon.crall@kitware.com>
6
+ License-Expression: Apache-2.0
7
+ Project-URL: Homepage, https://github.com/AIQ-Kitware/infer_stack
8
+ Classifier: Development Status :: 1 - Planning
9
+ Classifier: Intended Audience :: Developers
10
+ Classifier: Programming Language :: Python :: 3.10
11
+ Classifier: Programming Language :: Python :: 3.11
12
+ Classifier: Programming Language :: Python :: 3.12
13
+ Classifier: Programming Language :: Python :: 3.13
14
+ Classifier: Programming Language :: Python :: 3.14
15
+ Classifier: Topic :: Software Development :: Libraries :: Python Modules
16
+ Classifier: Topic :: Utilities
17
+ Requires-Python: >=3.10
18
+ Description-Content-Type: text/markdown
19
+ Requires-Dist: jinja2>=3.1
20
+ Requires-Dist: pyyaml>=6.0
21
+ Requires-Dist: requests>=2.31
22
+ Requires-Dist: rich>=13.0
23
+ Requires-Dist: scriptconfig>=0.9
24
+ Requires-Dist: ubelt>=1.3
25
+ Provides-Extra: tests
26
+ Requires-Dist: pytest>=7.0; extra == "tests"
27
+ Requires-Dist: pytest-codeblocks>=0.17; extra == "tests"
28
+
29
+ # Infer Stack
30
+
31
+ `infer_stack` manages **named stack profiles** for local and Kubernetes-backed inference.
32
+
33
+ A stack profile is a small graph made from:
34
+
35
+ * **providers** — inference runtimes such as vLLM and Ollama
36
+ * **gateways** — optional API routers such as LiteLLM
37
+ * **frontends** — optional UIs such as Open WebUI
38
+ * **routes** — optional public model aliases exposed through a gateway
39
+
40
+ This repo can render those profiles through two backends:
41
+
42
+ * **Compose** for local single-host serving. Compose supports vLLM, Ollama, optional LiteLLM, and optional Open WebUI.
43
+ * **KubeAI** for Kubernetes-backed vLLM serving. KubeAI support is vLLM-only for now.
44
+
45
+ The direct Ollama path can run without LiteLLM and without predeclaring models. vLLM profiles still use explicit runtimes, placement, and runtime settings.
46
+
47
+ ## Main commands
48
+
49
+ ```bash
50
+ infer-stack setup --backend compose --profile ollama-direct
51
+ # or: infer-stack setup --backend compose --profile qwen2-5-7b-instruct-turbo-default
52
+ infer-stack list-profiles
53
+ infer-stack describe-profile <profile>
54
+ infer-stack validate
55
+ infer-stack render
56
+ infer-stack up -d
57
+ infer-stack deploy
58
+ infer-stack switch <profile> --apply # re-render and converge; no separate up needed
59
+ infer-stack status
60
+ infer-stack smoke-test
61
+ ```
62
+
63
+ The CLI is built on [`scriptconfig`](https://gitlab.kitware.com/utils/scriptconfig),
64
+ so every subcommand is also importable as a Python class — useful for
65
+ notebooks, tests, and other scripts:
66
+
67
+ ```python
68
+ from infer_stack.cli import RenderCLI, SmokeTestCLI
69
+
70
+ RenderCLI.main(argv=False, profile="qwen2-5-7b-instruct-turbo-default", yes=True)
71
+ SmokeTestCLI.main(argv=False, model="qwen/qwen2.5-7b-instruct-turbo")
72
+ ```
73
+
74
+ `manage.py` and `infer-stack` are aliases for the same entry point;
75
+ shell examples below use `infer-stack`.
76
+
77
+ ## Operating the rendered Compose stack
78
+
79
+ Once the stack is up, common docker compose operations are available as
80
+ `infer-stack` subcommands so you don't have to `cd` into the rendered
81
+ output directory or repeat the `-f docker-compose.yml --env-file .env`
82
+ flags. They all resolve the rendered location via the same
83
+ `output.generated_dir` chain as the rest of the CLI.
84
+
85
+ ```bash
86
+ infer-stack ps # docker compose ps
87
+ infer-stack ps -a # include stopped
88
+ infer-stack logs -f open-webui # follow one service
89
+ infer-stack logs --tail=200 litellm vllm-* # tailored backlog
90
+ infer-stack restart open-webui # restart specific services
91
+ infer-stack stop # stop everything (no remove)
92
+ infer-stack start # start back up
93
+ infer-stack pull # refresh images
94
+ ```
95
+
96
+ For Ollama model management inside the rendered Ollama service, prefer the
97
+ CLI wrappers:
98
+
99
+ ```bash
100
+ infer-stack ollama-pull smollm2:135m
101
+ infer-stack ollama-list
102
+ infer-stack ollama-ps
103
+ ```
104
+
105
+ For other interactive one-shot commands inside a container, use
106
+ `infer-stack logs`, `infer-stack ps`, `infer-stack restart`, or fall back to raw
107
+ Compose only when no wrapper exists.
108
+
109
+ On the KubeAI backend these wrappers raise ``NotImplementedError`` —
110
+ use the equivalent ``kubectl`` commands in the meantime.
111
+
112
+ ## Inspect a profile before running it
113
+
114
+ ```bash
115
+ infer-stack describe-profile qwen2-5-7b-instruct-turbo-default --format yaml
116
+ ```
117
+
118
+ ## Stack profile model
119
+
120
+ Profiles are written as stack graphs. The main sections are `providers`, `gateways`, `frontends`, and `routes`. For details and examples, see [docs/stack-graph-profiles.md](docs/stack-graph-profiles.md).
121
+
122
+ Common shapes:
123
+
124
+ ```text
125
+ Open WebUI -> Ollama # ollama-direct, no LiteLLM
126
+ Open WebUI -> LiteLLM -> vLLM # classic vLLM compose profiles
127
+ Open WebUI -> LiteLLM -> Ollama # Ollama with stable aliases
128
+ Open WebUI -> LiteLLM -> Ollama + vLLM # mixed migration / test stacks
129
+ Ollama API + vLLM API directly # raw backend profiles
130
+ ```
131
+
132
+ Custom provider models and custom profiles live in the configured `catalog.user_models_file`, which defaults to `~/.config/infer_stack/models.yaml`. New files should prefer provider-specific top-level keys:
133
+
134
+ ```yaml
135
+ vllm_models:
136
+ my-vllm-model:
137
+ hf_model_id: org/model
138
+
139
+ ollama_models:
140
+ my-ollama-model:
141
+ tag: qwen3.5:4b
142
+
143
+ profiles:
144
+ my-stack:
145
+ providers: {}
146
+ gateways: {}
147
+ frontends: {}
148
+ routes: {}
149
+ ```
150
+
151
+ `models:` is still interpreted as a vLLM model catalog for convenience, but new docs and recipes use `vllm_models:` / `ollama_models:`.
152
+
153
+ ## Where config and rendered artifacts live
154
+
155
+ `infer-stack` follows XDG basedir conventions, so where you invoke it
156
+ from never changes which config it reads or where it writes rendered
157
+ artifacts:
158
+
159
+ There are exactly two path roots:
160
+
161
+ | What | Default location | How to relocate |
162
+ | --- | --- | --- |
163
+ | `config.yaml`, `models.yaml`, `kubeai-values.local.yaml` | `~/.config/infer_stack/` (resp. `$XDG_CONFIG_HOME`) | `--config-dir` (or `INFER_STACK_CONFIG_DIR`) |
164
+ | **Everything generated** — `generated/` (docker-compose.yml, .env, plan.yaml, kubeai/*) **and** `state/` (hf-cache, postgres volumes, Ollama store, runtime bind mounts) | `~/.local/share/infer_stack/` (resp. `$XDG_DATA_HOME`) | `--data-dir` (or `INFER_STACK_DATA_DIR`) |
165
+
166
+ `--data-dir` is the single knob for "put everything I generate in one
167
+ directory." Set it once at `setup`; it is baked into the absolute
168
+ `state.*` and `output.generated_dir` paths written to `config.yaml`, so
169
+ later commands don't need it again:
170
+
171
+ ```bash
172
+ # All rendered artifacts and bind-mount state land under one directory.
173
+ infer-stack setup \
174
+ --backend compose \
175
+ --profile ollama-direct \
176
+ --data-dir /data/service/docker/vllm-stack
177
+
178
+ infer-stack render --yes
179
+ ```
180
+
181
+ ```bash
182
+ # Keep config.yaml in a checkout for ad-hoc experiments.
183
+ infer-stack setup --config-dir $PWD --backend compose --profile <p> --data-dir $PWD/stack
184
+ ```
185
+
186
+ `--config-dir` / `--data-dir` live on every subcommand, so they appear
187
+ **after** the subcommand name. For "set once for the whole shell" use the
188
+ env vars instead. For a bespoke split layout (e.g. big `state/` on a data
189
+ disk, artifacts elsewhere), edit `state.*` / `output.generated_dir` in
190
+ `config.yaml` directly.
191
+
192
+ ## Constraining placement to specific GPUs
193
+
194
+ If some of your GPUs are tied up by other work, restrict the planner
195
+ (and the rendered ``device_ids``) to the subset you want it to use:
196
+
197
+ ```bash
198
+ # Only place onto GPU 1 (e.g. GPU 0 is running a display).
199
+ infer-stack render --yes --profile test-single-11gb --allowed-gpus 1
200
+
201
+ # Or pin a TP=2 profile to physical GPUs 1 and 3.
202
+ infer-stack render --yes --profile test-multi-gpu --allowed-gpus 1,3
203
+ ```
204
+
205
+ ``--allowed-gpus`` (or ``INFER_STACK_ALLOWED_GPUS=1,3``) filters the
206
+ detected inventory before placement — real indices are preserved, so
207
+ the rendered compose stack pins ``device_ids: ["1", "3"]`` to those
208
+ exact physical GPUs. Useful for integration tests that need to share a
209
+ host with other jobs.
210
+
211
+ ## Demos / integration recipes
212
+
213
+ End-to-end examples under [docs/demos/](docs/demos/) are written as
214
+ markdown tutorials. The CI smoke test is runnable with pytest-codeblocks:
215
+
216
+ ```bash
217
+ pytest --codeblocks docs/demos/ci_smoke_test.md
218
+ ```
219
+
220
+ Each ``bash`` block is a self-contained shell snippet you can also
221
+ copy-paste into a terminal. See
222
+ [docs/demos/ci_smoke_test.md](docs/demos/ci_smoke_test.md) for the
223
+ ``setup → describe → validate → render`` flow on the smallest test
224
+ profiles.
225
+
226
+ For a real running vLLM stack on a workstation, see
227
+ [docs/demos/quickstart.md](docs/demos/quickstart.md). For direct Ollama on a dual GTX 1080 Ti style host, see
228
+ [docs/demos/ollama_direct_quickstart.md](docs/demos/ollama_direct_quickstart.md). For a focused GPU-1 backend switch test, see [docs/demos/smollm2_gpu1_backend_switch.md](docs/demos/smollm2_gpu1_backend_switch.md).
229
+
230
+ User-supplied paths on the CLI (`--file`, `--from-file`,
231
+ `--resource-profiles-file`, `--output-dir`) still resolve against the
232
+ current working directory — they're meant to behave as typed.
233
+
234
+ ---
235
+
236
+ ## Backend 1: Compose
237
+
238
+ Use Compose for local single-host deployments. It can render direct Ollama stacks, vLLM stacks, mixed Ollama+vLLM stacks, and raw backend-only stacks.
239
+
240
+ ### Getting started
241
+
242
+ Prerequisite: Docker and the `docker compose` plugin must be installed.
243
+
244
+ ```bash
245
+ # Direct Ollama, no LiteLLM and no predeclared models.
246
+ infer-stack setup --backend compose --profile ollama-direct
247
+ infer-stack validate --simulate-hardware 2x11
248
+ infer-stack render --yes --simulate-hardware 2x11
249
+ infer-stack up -d
250
+
251
+ # Classic vLLM through LiteLLM/Open WebUI.
252
+ infer-stack setup --backend compose --profile qwen2-5-7b-instruct-turbo-default
253
+ infer-stack validate
254
+ infer-stack render
255
+ infer-stack up -d
256
+ ```
257
+
258
+ ### Test that it is responding
259
+
260
+ When LiteLLM is enabled, the default Compose front door is:
261
+
262
+ ```text
263
+ http://127.0.0.1:14042/v1
264
+ ```
265
+
266
+ When using `ollama-direct`, Open WebUI talks to Ollama directly and the Ollama API is available at:
267
+
268
+ ```text
269
+ http://127.0.0.1:11434
270
+ http://127.0.0.1:11434/v1
271
+ ```
272
+
273
+ unless you changed the relevant ports in config.
274
+
275
+ Wait until the active profile can serve a real request through its resolved default endpoint:
276
+
277
+ ```bash
278
+ infer-stack wait-ready
279
+ ```
280
+
281
+ `wait-ready` is stronger than Docker Compose health: it probes the user-facing
282
+ LiteLLM, Ollama, or direct vLLM access surface and, by default, requires a tiny
283
+ generation/completion to succeed. The smoke test runs this readiness probe by
284
+ default before issuing its normal test request:
285
+
286
+ ```bash
287
+ infer-stack smoke-test
288
+ ```
289
+
290
+ For direct Ollama profiles, pull a model first and then smoke-test that model:
291
+
292
+ ```bash
293
+ infer-stack ollama-pull qwen3.5:4b
294
+ infer-stack ollama-list
295
+ infer-stack smoke-test --model qwen3.5:4b
296
+ ```
297
+
298
+ For LiteLLM profiles, `smoke-test` reads the rendered `.env` automatically and
299
+ uses the active profile's resolved OpenAI-compatible front door. You can inspect
300
+ individual secrets when needed:
301
+
302
+ ```bash
303
+ infer-stack env LITELLM_MASTER_KEY
304
+ infer-stack env VLLM_BACKEND_API_KEY
305
+ ```
306
+
307
+ When you intentionally want the old quick behavior, skip the readiness wait:
308
+
309
+ ```bash
310
+ infer-stack smoke-test --no-wait --model gpt2
311
+ ```
312
+
313
+ ### Stop it
314
+
315
+ ```bash
316
+ infer-stack down
317
+ ```
318
+
319
+ `down` never removes named volumes. The Postgres data directory and the
320
+ Open WebUI volume are preserved across `down`, `up`, `switch`, and `render`.
321
+
322
+ ### Open WebUI authentication
323
+
324
+ By default Open WebUI runs with `WEBUI_AUTH=False` — no login screen,
325
+ anyone who can reach the port gets straight into the UI. This is the
326
+ expected behavior for a local dev box. To re-enable login/signup, set
327
+ in `config.yaml`:
328
+
329
+ ```yaml
330
+ open_webui:
331
+ auth: true
332
+ ```
333
+
334
+ and re-render. Existing accounts stored in the `postgres-open-webui`
335
+ volume are preserved across the toggle.
336
+
337
+ ### Reverse proxy (TLS) and LDAP
338
+
339
+ Open WebUI can be fronted by an opt-in nginx TLS reverse proxy, and its
340
+ login can be backed by an LDAP directory. Both are off by default and
341
+ configured as ordinary config fields. The built-in `openwebui-tls-ldap`
342
+ profile wires them together as a worked example (Ollama + Open WebUI
343
+ behind nginx, no public Open WebUI/Ollama ports); see
344
+ [`examples/openwebui-tls-ldap/`](examples/openwebui-tls-ldap/).
345
+
346
+ ```bash
347
+ infer-stack setup --backend compose --profile openwebui-tls-ldap
348
+ ```
349
+
350
+ **Reverse proxy.** Enable it under `frontends.reverse_proxy`. It renders
351
+ an nginx service plus a generated `state.runtime/nginx.conf`:
352
+
353
+ ```yaml
354
+ frontends:
355
+ reverse_proxy:
356
+ enabled: true
357
+ target: open_webui # or litellm / ollama / a custom upstream
358
+ server_name: host.example.com
359
+ ssl:
360
+ enabled: true
361
+ certificate: ./certs/site.crt
362
+ certificate_key: ./certs/site.key
363
+ dhparam: ./dhparam.pem # optional
364
+ ```
365
+
366
+ When `ssl.enabled` is true, port 80 redirects to HTTPS (`force_https`)
367
+ and the cert/key/dhparam host paths are bind-mounted read-only. When
368
+ `ssl.enabled` is false, only HTTP is published (HTTPS publishing is
369
+ gated on TLS so you never get a `:443` mapping with nothing listening).
370
+
371
+ > **Path caveat.** Relative `certificate`/`certificate_key`/`dhparam`/
372
+ > `config_path` values are written verbatim into the generated
373
+ > `docker-compose.yml`, so Docker Compose resolves them **relative to the
374
+ > generated directory** (where the compose file lives), not your CWD.
375
+ > `infer-stack render` warns when a referenced cert or config file is not
376
+ > found. Use absolute paths if you want to avoid the ambiguity.
377
+
378
+ **LDAP.** Enable it under `frontends.open_webui.ldap`. The directory
379
+ settings render as Open WebUI `LDAP_*` environment variables, and
380
+ secrets/site-specific values are emitted as `.env` placeholders
381
+ (`LDAP_HOST`, `LDAP_PASSWD`, `LDAP_SEARCH_BASE`, …) so you can fill them
382
+ in after the first render without re-touching the compose YAML:
383
+
384
+ ```yaml
385
+ frontends:
386
+ open_webui:
387
+ ldap:
388
+ enabled: true
389
+ env_defaults:
390
+ LDAP_PORT: '636'
391
+ LDAP_USE_TLS: 'true'
392
+ LDAP_ATTRIBUTE_FOR_USERNAME: uid
393
+ ```
394
+
395
+ **Manual escape hatches.** When the typed renderer is not enough, drop
396
+ to manual control without leaving infer-stack:
397
+
398
+ * `frontends.reverse_proxy.config_path` — mount an existing nginx config
399
+ file instead of rendering one.
400
+ * `frontends.reverse_proxy.extra_config` — inject extra directives into
401
+ the rendered HTTPS `server` block.
402
+ * Every rendered service (`ollama`, vLLM runtimes, `litellm`,
403
+ `open_webui`, `reverse_proxy`) accepts generic overrides:
404
+ `extra_env`, `env_file`, `extra_volumes`, `extra_hosts`, `labels`,
405
+ `additional_ports`, and `gpus` (scalar `all`/count or a structured
406
+ device-request list).
407
+
408
+ Field precedence (lowest to highest) is: top-level config section
409
+ (`reverse_proxy:` / `open_webui:` / `ollama:`) → the matching
410
+ `frontends.*` / `providers.*` / `gateways.*` section → the active
411
+ profile. Newer configs should prefer the `frontends.*` / `providers.*`
412
+ form shown above.
413
+
414
+ ### Persistent state and database layout
415
+
416
+ Compose renders stateful services only when their components are enabled:
417
+
418
+ * `postgres-open-webui` — rendered only when Open WebUI is enabled. It stores chats, accounts, and settings in `state.postgres_open_webui`.
419
+ * `postgres-litellm` — rendered only when LiteLLM is enabled. It stores router state in `state.postgres_litellm`.
420
+ * `ollama` — rendered only when the Ollama provider is enabled. Its model store is `state.ollama`, mounted at `/root/.ollama`.
421
+ * vLLM runtimes mount `state.hf_cache` for Hugging Face weights and `state.vllm_cache` for compiled artifacts.
422
+
423
+ Each Postgres container has its own `POSTGRES_DB`, `POSTGRES_USER`, and
424
+ `POSTGRES_PASSWORD`, sourced from component-specific `.env` keys. There is no shared Postgres instance and no `postgres-init` bootstrap service.
425
+
426
+ Open WebUI chat history is **not** tied to the model currently being served, so after a profile switch old chats may reference model IDs the current gateway no longer advertises — that is expected.
427
+
428
+ ### Operational tips
429
+
430
+ Prefer scoping commands to specific services rather than relying on
431
+ container names. Use only the services rendered by the active profile:
432
+
433
+ ```bash
434
+ # LiteLLM gateway profile
435
+ infer-stack logs -f litellm
436
+
437
+ # Direct Ollama profile
438
+ infer-stack logs -f ollama
439
+
440
+ # Ollama model store helpers
441
+ infer-stack ollama-list
442
+ infer-stack ollama-ps
443
+ ```
444
+
445
+ You do not need to delete any volume during normal operation. If you
446
+ ever want a destructive reset, do it explicitly with
447
+ `docker compose down -v` against `generated/docker-compose.yml` — the
448
+ toolchain itself never does this.
449
+
450
+ ### Custom .env values are preserved
451
+
452
+ `generated/.env` is rewritten non-destructively. Any `KEY=value` pair
453
+ you add manually (for example `VERBOSE=1`, `HF_HOME=/data/hf`, or any
454
+ key this program does not yet know about) is preserved across
455
+ `render`, `setup`, `switch`, `up`, and `deploy`. Comments and the order
456
+ of existing lines are preserved where practical.
457
+
458
+ ### Switching profiles
459
+
460
+ ```bash
461
+ infer-stack switch <profile> --apply
462
+ ```
463
+
464
+ `switch --apply` re-renders from the updated `config.yaml`, then brings the
465
+ stack up convergently with `--remove-orphans` so a separate `infer-stack up` is
466
+ not needed. Components/runtimes that are no longer in the rendered compose file
467
+ are dropped. Compose preserves existing containers whose service definitions did
468
+ not change. For vLLM-to-vLLM profile switches, unchanged Open WebUI stays up;
469
+ LiteLLM is refreshed through its admin API when possible. The live refresh
470
+ path treats LiteLLM's "model not found in db" response for config-backed
471
+ models as non-fatal, so switching aliases can add the new route without
472
+ tearing LiteLLM down. That can temporarily leave stale config-backed aliases
473
+ in `/v1/models`; restart LiteLLM manually only when you want to clean those
474
+ up. If Compose already created or recreated LiteLLM while converging the new
475
+ stack, no extra router refresh is attempted because the new container has
476
+ already loaded the freshly rendered YAML. Profiles that do not render
477
+ LiteLLM, such as direct Ollama profiles, skip the router refresh path even if
478
+ an old `runtime/litellm_config.yaml` file remains from a previous profile.
479
+ Switches that change Open WebUI's provider wiring, such as
480
+ `Open WebUI -> LiteLLM` to `Open WebUI -> Ollama`, necessarily recreate
481
+ Open WebUI because its environment changes. Postgres volumes and provider
482
+ caches are left untouched. vLLM runtime containers are named after their
483
+ Compose service, for example `vllm-chat`, so `docker ps` and
484
+ `infer-stack logs vllm-chat` clearly identify them as vLLM containers.
485
+
486
+ ### Protocol modes for base vs. instruct models
487
+
488
+ Profiles declare a `protocol_mode` (`chat` or `completions`) that the
489
+ served model must support. Models also declare which protocols they
490
+ support via `supported_protocols`. Validation runs before render and
491
+ fails with an actionable message if a profile asks for `chat` on a
492
+ completions-only model.
493
+
494
+ Practical guidance:
495
+
496
+ * Instruct/chat models (with a chat template) can use either, but
497
+ default to `chat`.
498
+ * Base models like Pythia, Llama-2 base, Mistral-v0.1 base, and Falcon
499
+ base do not define a chat template. Their HELM profiles use
500
+ `protocol_mode: completions` and the `smoke-test` command will
501
+ exercise `/v1/completions` for them.
502
+ * The rendered LiteLLM config uses `text-completion-openai/<served>`
503
+ as the upstream provider for completions-only services. That means
504
+ even chat-shaped requests sent through Open WebUI to a Pythia model
505
+ get translated by LiteLLM into upstream `/v1/completions` calls — no
506
+ second vLLM container is needed to support Open WebUI for a
507
+ completions-only model.
508
+ * Open WebUI is still a chat UI, so prompt formatting matters.
509
+ HELM/eval clients should call `/v1/completions` directly for exact
510
+ prompt control rather than going through the chat frontend.
511
+
512
+ ### Chat-shaped clients on top of completions models
513
+
514
+ Some clients (e.g. InspectAI / Inspect Evals stock MMLU tasks) only
515
+ speak `/v1/chat/completions` and cannot be reconfigured. For those
516
+ cases, profiles can opt into a LiteLLM-only adapter:
517
+
518
+ ```yaml
519
+ chat_compat:
520
+ enabled: true
521
+ strategy: flat_messages
522
+ ```
523
+
524
+ When set on a `protocol_mode: completions` service, the rendered
525
+ LiteLLM config keeps the `text-completion-openai/<served>` upstream
526
+ and adds LiteLLM's documented prompt-template fields
527
+ (`initial_prompt_value` / `roles` / `final_prompt_value`) so chat
528
+ messages get flattened into a plain prompt — no role labels, messages
529
+ joined by `\n` — before being forwarded to vLLM `/v1/completions`.
530
+
531
+ This is **not** a chat tune; the model is still a base model and
532
+ prompt formatting still matters for evaluation. Use it only when a
533
+ chat-shaped client cannot be changed. The vLLM container is not
534
+ restarted, no `--chat-template` is rendered, and the adapter takes
535
+ effect after a `litellm`-only restart:
536
+
537
+ ```bash
538
+ infer-stack render
539
+ infer-stack restart litellm
540
+ ```
541
+
542
+ The built-in `pythia-inspect-mmlu-compat` profile is a ready-made
543
+ example; see
544
+ [`recipies/compose_pythia_inspect_mmlu_compat.md`](recipies/compose_pythia_inspect_mmlu_compat.md).
545
+
546
+ ### Reasoning / thinking models
547
+
548
+ Models can declare reasoning support in the catalog:
549
+
550
+ ```yaml
551
+ reasoning:
552
+ enabled: true
553
+ parser: qwen3
554
+ expose_to_openwebui: true
555
+ ```
556
+
557
+ Profiles can override or set the same field per service. When a
558
+ service has `reasoning.enabled: true` and a `parser`, the renderer
559
+ adds `--reasoning-parser <parser>` to that vLLM container's command
560
+ line — that flag alone enables reasoning extraction in the current
561
+ vLLM CLI. You do not need to repeat it by hand in `extra_args`.
562
+
563
+ Open WebUI sees reasoning content via two paths:
564
+
565
+ 1. Inline `<think>...</think>` tags emitted by the model.
566
+ 2. Structured `reasoning_content` fields when LiteLLM normalizes them.
567
+
568
+ The LiteLLM template keeps `merge_reasoning_content_in_choices: true`
569
+ on chat-mode entries so Open WebUI can display reasoning in the
570
+ streamed response. To test reasoning end-to-end:
571
+
572
+ ```bash
573
+ # Non-streaming CLI smoke test:
574
+ infer-stack smoke-test \
575
+ --model qwen3.6-35b-a3b \
576
+ --prompt "Think step by step: 17*23"
577
+
578
+ # For streaming inspection, read the key with the CLI wrapper:
579
+ LITELLM_MASTER_KEY=$(infer-stack env LITELLM_MASTER_KEY)
580
+ curl -N http://127.0.0.1:14042/v1/chat/completions \
581
+ -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
582
+ -H 'Content-Type: application/json' \
583
+ -d '{"model":"qwen3.6-35b-a3b","stream":true,
584
+ "messages":[{"role":"user","content":"Think step by step: 17*23"}]}'
585
+ ```
586
+
587
+ In Open WebUI, reasoning shows up best with streaming enabled in the
588
+ chat settings.
589
+
590
+ ---
591
+
592
+ ## Backend 2: KubeAI
593
+
594
+ Use KubeAI when you want Kubernetes-managed serving.
595
+
596
+ ### Important rules
597
+
598
+ 1. **Use the same namespace everywhere.** The namespace in `infer-stack setup --namespace ...` must match the namespace where the KubeAI Helm release already exists.
599
+ 2. **Prefer the repo-driven path.** The normal path is `setup` -> `validate` -> `render` -> `deploy` -> `status`.
600
+ 3. **`kubectl port-forward` stays in the foreground.** Leave it running in one terminal and send requests from another.
601
+ 4. **The first request can take a while.** `/openai/v1/models` may work before chat completions work. The first completion may trigger pod creation, image pull, model load, and compile warmup.
602
+ 5. **On the current repo version, KubeAI still needs a live workaround after deploy.** The renderer currently produces a `Model` spec that needs a small manual patch to work with the KubeAI version used in these notes.
603
+
604
+ ### KubeAI prerequisites
605
+
606
+ You need:
607
+
608
+ * a working Kubernetes cluster
609
+ * `kubectl`
610
+ * Helm
611
+
612
+ If you want a quick local single-node cluster, K3s is a good option.
613
+
614
+ Install K3s:
615
+
616
+ ```bash
617
+ curl -sfL https://get.k3s.io | sh -
618
+ # or pin
619
+ curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION='v1.34.3+k3s1' sh -
620
+ ```
621
+
622
+ Make `kubectl` usable without `sudo`:
623
+
624
+ ```bash
625
+ sudo mkdir -p /etc/rancher/k3s/config.yaml.d
626
+ printf 'write-kubeconfig-mode: "0644"\n' | \
627
+ sudo tee /etc/rancher/k3s/config.yaml.d/10-kubeconfig-mode.yaml >/dev/null
628
+ sudo systemctl restart k3s
629
+ kubectl get nodes
630
+ ```
631
+
632
+ Install Helm:
633
+
634
+ ```bash
635
+ curl -fsSL -o get_helm.sh https://raw.githubusercontent.com/helm/helm/83a46119086589a593a62ca544982977a60318ca/scripts/get-helm-4
636
+ chmod 700 get_helm.sh
637
+ ./get_helm.sh
638
+ helm version
639
+ ```
640
+
641
+ ### NVIDIA GPU support
642
+
643
+ Install the NVIDIA device plugin and GPU Feature Discovery so Kubernetes can expose GPU resources and labels:
644
+
645
+ ```bash
646
+ helm repo add nvdp https://nvidia.github.io/k8s-device-plugin
647
+ helm repo update
648
+ helm upgrade -i nvdp nvdp/nvidia-device-plugin \
649
+ --version 0.17.1 \
650
+ --namespace nvidia-device-plugin \
651
+ --create-namespace \
652
+ --set gfd.enabled=true \
653
+ --set runtimeClassName=nvidia
654
+ ```
655
+
656
+ Check that GPU support is working:
657
+
658
+ ```bash
659
+ kubectl -n nvidia-device-plugin get pods
660
+ kubectl get node "$(kubectl get nodes -o jsonpath='{.items[0].metadata.name}')" \
661
+ -o jsonpath='{.status.allocatable.nvidia\.com/gpu}{"\n"}'
662
+ kubectl get nodes --show-labels | tr ',' '\n' | grep 'nvidia.com/' || true
663
+ ```
664
+
665
+ You want to see a non-empty `nvidia.com/gpu` count and `nvidia.com/*` labels such as product and memory.
666
+
667
+ ### KubeAI Helm repository
668
+
669
+ ```bash
670
+ helm repo add kubeai https://www.kubeai.org
671
+ helm repo update
672
+ ```
673
+
674
+ ## Determine which namespace to use
675
+
676
+ Before doing anything else, discover whether a `kubeai` release already exists and which namespace it uses.
677
+
678
+ ```bash
679
+ KUBEAI_NAMESPACE="$(helm list -A | awk '$1=="kubeai"{print $2; exit}')"
680
+ if [ -z "${KUBEAI_NAMESPACE}" ]; then
681
+ KUBEAI_NAMESPACE=default
682
+ fi
683
+ echo "Using KubeAI namespace: ${KUBEAI_NAMESPACE}"
684
+ ```
685
+
686
+ If a release already exists, **reuse that namespace**.
687
+
688
+ Sanity-check the cluster:
689
+
690
+ ```bash
691
+ kubectl get nodes
692
+ kubectl get crd models.kubeai.org || true
693
+ helm list -A | grep kubeai || true
694
+ kubectl -n "${KUBEAI_NAMESPACE}" get pods || true
695
+ kubectl get node "$(kubectl get nodes -o jsonpath='{.items[0].metadata.name}')" \
696
+ -o jsonpath='{.status.allocatable.nvidia\.com/gpu}{"\n"}'
697
+ ```
698
+
699
+ ---
700
+
701
+ ## Generate the local KubeAI resource-profile file
702
+
703
+ Generate a local KubeAI resource-profile file from the labels on this machine.
704
+
705
+ For the built-in serving profiles in this repo, keep these names aligned:
706
+
707
+ * `gpu-single-default`
708
+ * `gpu-tp2-balanced`
709
+ * `gpu-tp2-maxctx`
710
+
711
+ **Important:** include GPU `requests`, GPU `limits`, and `runtimeClassName: nvidia`. Without those, the model pod can land on the GPU node but still start without `libcuda.so.1` available inside the container.
712
+
713
+ ```bash
714
+ PRODUCT="$(kubectl get nodes -o jsonpath='{.items[0].metadata.labels.nvidia\.com/gpu\.product}')"
715
+ MEMORY="$(kubectl get nodes -o jsonpath='{.items[0].metadata.labels.nvidia\.com/gpu\.memory}')"
716
+
717
+ cat > values-kubeai-local-gpu.yaml <<EOF
718
+ resourceProfiles:
719
+ gpu-single-default:
720
+ nodeSelector:
721
+ nvidia.com/gpu.product: "${PRODUCT}"
722
+ nvidia.com/gpu.memory: "${MEMORY}"
723
+ requests:
724
+ nvidia.com/gpu: 1
725
+ limits:
726
+ nvidia.com/gpu: 1
727
+ runtimeClassName: nvidia
728
+
729
+ gpu-tp2-balanced:
730
+ nodeSelector:
731
+ nvidia.com/gpu.product: "${PRODUCT}"
732
+ nvidia.com/gpu.memory: "${MEMORY}"
733
+ requests:
734
+ nvidia.com/gpu: 2
735
+ limits:
736
+ nvidia.com/gpu: 2
737
+ runtimeClassName: nvidia
738
+
739
+ gpu-tp2-maxctx:
740
+ nodeSelector:
741
+ nvidia.com/gpu.product: "${PRODUCT}"
742
+ nvidia.com/gpu.memory: "${MEMORY}"
743
+ requests:
744
+ nvidia.com/gpu: 2
745
+ limits:
746
+ nvidia.com/gpu: 2
747
+ runtimeClassName: nvidia
748
+ EOF
749
+
750
+ cat values-kubeai-local-gpu.yaml
751
+ ```
752
+
753
+ Sync that file so `validate` and `render` use the same local resource-profile data:
754
+
755
+ ```bash
756
+ infer-stack kubeai-sync-resource-profiles --from-file values-kubeai-local-gpu.yaml
757
+ ```
758
+
759
+ ---
760
+
761
+ ## Example 1: single-GPU system
762
+
763
+ Use this example on a 1-GPU workstation.
764
+
765
+ ```bash
766
+ infer-stack setup \
767
+ --backend kubeai \
768
+ --profile qwen2-5-7b-instruct-turbo-default \
769
+ --namespace "${KUBEAI_NAMESPACE}"
770
+
771
+ infer-stack list-profiles
772
+ infer-stack describe-profile qwen2-5-7b-instruct-turbo-default --format yaml
773
+ infer-stack validate
774
+ infer-stack render
775
+ infer-stack deploy
776
+ infer-stack status
777
+ ```
778
+
779
+ ### Current live workaround for the single-GPU example
780
+
781
+ On the current repo version, apply this live patch after `deploy`.
782
+
783
+ This patch does four things:
784
+
785
+ * keeps the model warm with `minReplicas: 1`
786
+ * changes `resourceProfile` from `gpu-single-default` to `gpu-single-default:1`
787
+ * makes the served model name match the public profile name
788
+ * avoids the duplicate `--served-model-name` mismatch that causes 404s on completions
789
+
790
+ ```bash
791
+ kubectl -n "${KUBEAI_NAMESPACE}" patch model qwen2-5-7b-instruct-turbo-default --type merge -p '{
792
+ "spec": {
793
+ "minReplicas": 1,
794
+ "resourceProfile": "gpu-single-default:1",
795
+ "args": [
796
+ "--served-model-name=qwen2-5-7b-instruct-turbo-default",
797
+ "--tensor-parallel-size=1",
798
+ "--data-parallel-size=1",
799
+ "--max-model-len=32768",
800
+ "--gpu-memory-utilization=0.9",
801
+ "--max-num-batched-tokens=8192",
802
+ "--max-num-seqs=16",
803
+ "--disable-log-requests",
804
+ "--enable-prefix-caching"
805
+ ]
806
+ }
807
+ }'
808
+
809
+ kubectl -n "${KUBEAI_NAMESPACE}" delete pod -l model=qwen2-5-7b-instruct-turbo-default
810
+ ```
811
+
812
+ If you run `infer-stack render` or `infer-stack deploy` again on the current repo version, re-apply this live patch.
813
+
814
+ ---
815
+
816
+ ## Example 2: four-GPU system
817
+
818
+ On a 4-GPU host, do the same **single-GPU smoke test first** to verify the cluster, KubeAI, runtime class, and model plumbing. That exact sequence worked on a 4-GPU machine during bring-up.
819
+
820
+ ```bash
821
+ infer-stack setup \
822
+ --backend kubeai \
823
+ --profile qwen2-5-7b-instruct-turbo-default \
824
+ --namespace "${KUBEAI_NAMESPACE}"
825
+
826
+ infer-stack validate
827
+ infer-stack render
828
+ infer-stack deploy
829
+ infer-stack status
830
+
831
+ kubectl -n "${KUBEAI_NAMESPACE}" patch model qwen2-5-7b-instruct-turbo-default --type merge -p '{
832
+ "spec": {
833
+ "minReplicas": 1,
834
+ "resourceProfile": "gpu-single-default:1",
835
+ "args": [
836
+ "--served-model-name=qwen2-5-7b-instruct-turbo-default",
837
+ "--tensor-parallel-size=1",
838
+ "--data-parallel-size=1",
839
+ "--max-model-len=32768",
840
+ "--gpu-memory-utilization=0.9",
841
+ "--max-num-batched-tokens=8192",
842
+ "--max-num-seqs=16",
843
+ "--disable-log-requests",
844
+ "--enable-prefix-caching"
845
+ ]
846
+ }
847
+ }'
848
+
849
+ kubectl -n "${KUBEAI_NAMESPACE}" delete pod -l model=qwen2-5-7b-instruct-turbo-default
850
+ ```
851
+
852
+ After the 7B smoke test works, move up to larger profiles such as `qwen2-72b-instruct-tp2-balanced`. On the current repo version, apply the same kind of live patch after deploy: keep `minReplicas: 1`, append `:1` to the chosen `resourceProfile`, and make the single effective `--served-model-name` match the public profile name.
853
+
854
+ ---
855
+
856
+ ## Test that KubeAI is responding
857
+
858
+ If you are not exposing ingress yet, port-forward the service.
859
+
860
+ **This command stays in the foreground.** Run it in one terminal and leave it there:
861
+
862
+ ```bash
863
+ kubectl -n "${KUBEAI_NAMESPACE}" port-forward svc/kubeai 8000:80
864
+ ```
865
+
866
+ Then use another terminal for requests.
867
+
868
+ ### First check: `/models`
869
+
870
+ ```bash
871
+ curl http://127.0.0.1:8000/openai/v1/models
872
+ ```
873
+
874
+ If that works, the KubeAI front door is alive.
875
+
876
+ ### Then try the smoke test
877
+
878
+ ```bash
879
+ infer-stack smoke-test \
880
+ --base-url http://127.0.0.1:8000/openai/v1 \
881
+ --model qwen2-5-7b-instruct-turbo-default
882
+ ```
883
+
884
+ ### Or test chat completions directly
885
+
886
+ ```bash
887
+ time curl http://127.0.0.1:8000/openai/v1/chat/completions \
888
+ -H 'Content-Type: application/json' \
889
+ -d '{
890
+ "model": "qwen2-5-7b-instruct-turbo-default",
891
+ "messages": [{"role": "user", "content": "Say hello in one short sentence."}],
892
+ "max_tokens": 8
893
+ }'
894
+ ```
895
+
896
+ ### What to expect on the first request
897
+
898
+ Common first-request behavior:
899
+
900
+ * `/openai/v1/models` works before completions work
901
+ * a completion request causes KubeAI to create a model-serving pod
902
+ * that pod may spend time in `ContainerCreating` while the image is pulled
903
+ * the model then spends more time loading and warming up
904
+ * the first completion can be much slower than later ones
905
+
906
+ That is not automatically a failure. Watch the system state while the first request is happening:
907
+
908
+ ```bash
909
+ watch -n 1 'kubectl -n '"${KUBEAI_NAMESPACE}"' get pods; echo; kubectl -n '"${KUBEAI_NAMESPACE}"' get models'
910
+ ```
911
+
912
+ ---
913
+
914
+ ## Debugging checks
915
+
916
+ ### Check the live Model object
917
+
918
+ ```bash
919
+ kubectl -n "${KUBEAI_NAMESPACE}" describe model qwen2-5-7b-instruct-turbo-default
920
+ kubectl -n "${KUBEAI_NAMESPACE}" get model qwen2-5-7b-instruct-turbo-default -o yaml | grep -E 'minReplicas|maxReplicas|resourceProfile'
921
+ ```
922
+
923
+ ### Check the current model pod
924
+
925
+ ```bash
926
+ kubectl -n "${KUBEAI_NAMESPACE}" describe pod "$(kubectl -n "${KUBEAI_NAMESPACE}" get pods -o name | grep 'model-qwen2-5-7b-instruct-turbo-default' | tail -n 1 | cut -d/ -f2)"
927
+ ```
928
+
929
+ ### Tail KubeAI controller logs
930
+
931
+ ```bash
932
+ kubectl -n "${KUBEAI_NAMESPACE}" logs deploy/kubeai --tail=200 -f
933
+ ```
934
+
935
+ ### Tail model-server logs
936
+
937
+ ```bash
938
+ kubectl -n "${KUBEAI_NAMESPACE}" logs -f "$(kubectl -n "${KUBEAI_NAMESPACE}" get pods -o name | grep 'model-qwen2-5-7b-instruct-turbo-default' | tail -n 1 | cut -d/ -f2)" -c server
939
+ ```
940
+
941
+ If the model pod restarted, inspect the previous crash:
942
+
943
+ ```bash
944
+ kubectl -n "${KUBEAI_NAMESPACE}" logs "$(kubectl -n "${KUBEAI_NAMESPACE}" get pods -o name | grep 'model-qwen2-5-7b-instruct-turbo-default' | tail -n 1 | cut -d/ -f2)" -c server --previous
945
+ ```
946
+
947
+ ### Check recent events
948
+
949
+ ```bash
950
+ kubectl -n "${KUBEAI_NAMESPACE}" get events --sort-by=.lastTimestamp | tail -n 40
951
+ ```
952
+
953
+ ### Common bad states and what they mean
954
+
955
+ * `invalid resource profile: "gpu-single-default", should match <name>:<multiple>`
956
+ * append `:1` in the live `Model` spec
957
+ * `libcuda.so.1: cannot open shared object file`
958
+ * the pod landed on the GPU node without actually requesting a GPU; fix the resource-profile file to include GPU requests, limits, and `runtimeClassName: nvidia`
959
+ * `/models` works but completions 404 with `The model ... does not exist.`
960
+ * the served model name does not match the public profile name; apply the live args patch above
961
+ * startup probe fails with `connection refused`
962
+ * the model pod may still be pulling the image, loading the model, or warming up
963
+
964
+ ---
965
+
966
+ ## Which backend should I start with?
967
+
968
+ Start with **Compose** if you want:
969
+
970
+ * the fastest path to a working local server
971
+ * easy inspection of generated files
972
+ * simple single-host iteration
973
+
974
+ Move to **KubeAI** when you want:
975
+
976
+ * vLLM runtimes on Kubernetes
977
+ * KubeAI’s OpenAI-compatible front door
978
+ * profile deployment through Kubernetes artifacts
979
+
980
+ KubeAI rendering is vLLM-only for now. Profiles that enable Ollama, LiteLLM, or Open WebUI are rejected for `--backend kubeai`.
981
+
982
+ A good workflow is:
983
+
984
+ 1. inspect a profile with `describe-profile`
985
+ 2. run it with Compose when you want the simplest local deployment
986
+ 3. move to KubeAI when you want Kubernetes-backed serving
987
+
988
+ Compose is the better fit when you already know which profile you want. KubeAI has more first-request overhead because it may need to create pods, pull images, load the model, and warm up the backend.
989
+
990
+
991
+
992
+ ## vLLM startup caches
993
+
994
+ Generated Compose mounts persist Hugging Face, vLLM, PyTorch/TorchInductor,
995
+ Triton, and CUDA JIT caches. Warm starts avoid redownloading and redoing many
996
+ compile/JIT steps, but a vLLM model swap still creates a new engine process and
997
+ must reload weights into GPU memory.
998
+
999
+ ### Diagnosing profile switches and readiness
1000
+
1001
+ `docker compose` health only means that a container-level healthcheck passed. It
1002
+ is not the same thing as "the routed model can answer a request through the
1003
+ active access surface." This matters most when switching between two vLLM
1004
+ profiles that reuse the same runtime service name: the old vLLM process exits,
1005
+ Docker starts the replacement process, and LiteLLM may remain up while returning
1006
+ upstream connection errors until vLLM finishes loading the new model.
1007
+
1008
+ Use the dedicated readiness and diagnostics commands after a switch:
1009
+
1010
+ ```bash
1011
+ infer-stack switch gpt2-single --apply --yes
1012
+ infer-stack wait-ready --model gpt2
1013
+ infer-stack smoke-test --model gpt2
1014
+ ```
1015
+
1016
+ For debugging, use:
1017
+
1018
+ ```bash
1019
+ infer-stack diagnose --model gpt2 --generation
1020
+ infer-stack diagnose --logs --tail 80
1021
+ ```
1022
+
1023
+ `diagnose` prints the resolved provider/gateway/frontend graph, rendered Compose
1024
+ service state, LiteLLM route probes, direct provider probes, and optional recent
1025
+ logs. It is intended to distinguish an actual LiteLLM outage from the more
1026
+ common case where LiteLLM is running but its upstream vLLM runtime is still
1027
+ booting.
1028
+
1029
+ The Compose service-state diagnostics include Docker's exit code, OOM-killed
1030
+ flag, restart count, and actual container name. This is important because
1031
+ `litellm exited with code 137` usually means Docker sent SIGKILL, commonly from
1032
+ an OOM kill or a forced container replacement, whereas LiteLLM returning HTTP
1033
+ 500 with `Cannot connect to host vllm-*` means LiteLLM is still running but the
1034
+ upstream vLLM runtime is not ready yet.