infer-stack 0.6.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- infer_stack-0.6.0/PKG-INFO +1034 -0
- infer_stack-0.6.0/README.md +1006 -0
- infer_stack-0.6.0/infer_stack/__init__.py +2 -0
- infer_stack-0.6.0/infer_stack/backends/__init__.py +7 -0
- infer_stack-0.6.0/infer_stack/backends/compose_renderer.py +243 -0
- infer_stack-0.6.0/infer_stack/backends/kubeai_renderer.py +202 -0
- infer_stack-0.6.0/infer_stack/benchmark.py +38 -0
- infer_stack-0.6.0/infer_stack/catalog.py +438 -0
- infer_stack-0.6.0/infer_stack/cli/__init__.py +169 -0
- infer_stack-0.6.0/infer_stack/cli/__main__.py +4 -0
- infer_stack-0.6.0/infer_stack/cli/commands_profile.py +467 -0
- infer_stack-0.6.0/infer_stack/cli/commands_runtime.py +719 -0
- infer_stack-0.6.0/infer_stack/cli/commands_smoke.py +691 -0
- infer_stack-0.6.0/infer_stack/cli/compose.py +755 -0
- infer_stack-0.6.0/infer_stack/cli/context.py +471 -0
- infer_stack-0.6.0/infer_stack/cli/options.py +134 -0
- infer_stack-0.6.0/infer_stack/cli/probes.py +178 -0
- infer_stack-0.6.0/infer_stack/config.py +450 -0
- infer_stack-0.6.0/infer_stack/contracts.py +223 -0
- infer_stack-0.6.0/infer_stack/diff_prompt.py +117 -0
- infer_stack-0.6.0/infer_stack/docker_utils.py +230 -0
- infer_stack-0.6.0/infer_stack/env_utils.py +97 -0
- infer_stack-0.6.0/infer_stack/experimental/model_catalog_discover.py +1155 -0
- infer_stack-0.6.0/infer_stack/experimental/model_memory_estimator.py +1264 -0
- infer_stack-0.6.0/infer_stack/experimental/stress_test_long_context.py +397 -0
- infer_stack-0.6.0/infer_stack/hardware.py +70 -0
- infer_stack-0.6.0/infer_stack/kubeai_ops.py +76 -0
- infer_stack-0.6.0/infer_stack/paths.py +87 -0
- infer_stack-0.6.0/infer_stack/profile_runtime.py +46 -0
- infer_stack-0.6.0/infer_stack/renderer.py +19 -0
- infer_stack-0.6.0/infer_stack/resolver.py +1092 -0
- infer_stack-0.6.0/infer_stack/templates/default-models.yaml +674 -0
- infer_stack-0.6.0/infer_stack/templates/default-ollama-models.yaml +31 -0
- infer_stack-0.6.0/infer_stack/templates/default-profiles.yaml +1731 -0
- infer_stack-0.6.0/infer_stack/templates/default-vllm-models.yaml +714 -0
- infer_stack-0.6.0/infer_stack/templates/docker-compose.yml.j2 +430 -0
- infer_stack-0.6.0/infer_stack/templates/litellm_config.yaml.j2 +44 -0
- infer_stack-0.6.0/infer_stack/templates/nginx.conf.j2 +84 -0
- infer_stack-0.6.0/infer_stack/tuning.py +3 -0
- infer_stack-0.6.0/infer_stack/validator.py +314 -0
- infer_stack-0.6.0/infer_stack/verification.py +46 -0
- infer_stack-0.6.0/infer_stack.egg-info/PKG-INFO +1034 -0
- infer_stack-0.6.0/infer_stack.egg-info/SOURCES.txt +52 -0
- infer_stack-0.6.0/infer_stack.egg-info/dependency_links.txt +1 -0
- infer_stack-0.6.0/infer_stack.egg-info/entry_points.txt +2 -0
- infer_stack-0.6.0/infer_stack.egg-info/requires.txt +10 -0
- infer_stack-0.6.0/infer_stack.egg-info/top_level.txt +1 -0
- infer_stack-0.6.0/pyproject.toml +168 -0
- infer_stack-0.6.0/setup.cfg +4 -0
- infer_stack-0.6.0/tests/test_cli_setup.py +723 -0
- infer_stack-0.6.0/tests/test_import.py +2 -0
- infer_stack-0.6.0/tests/test_ollama_stack_graph.py +262 -0
- infer_stack-0.6.0/tests/test_reverse_proxy_ldap.py +392 -0
- infer_stack-0.6.0/tests/test_serving_profiles.py +1152 -0
|
@@ -0,0 +1,1034 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: infer-stack
|
|
3
|
+
Version: 0.6.0
|
|
4
|
+
Summary: Profile-driven compose and KubeAI deployment compiler for inference stacks
|
|
5
|
+
Author-email: "jon.crall" <jon.crall@kitware.com>
|
|
6
|
+
License-Expression: Apache-2.0
|
|
7
|
+
Project-URL: Homepage, https://github.com/AIQ-Kitware/infer_stack
|
|
8
|
+
Classifier: Development Status :: 1 - Planning
|
|
9
|
+
Classifier: Intended Audience :: Developers
|
|
10
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
11
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
12
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
15
|
+
Classifier: Topic :: Software Development :: Libraries :: Python Modules
|
|
16
|
+
Classifier: Topic :: Utilities
|
|
17
|
+
Requires-Python: >=3.10
|
|
18
|
+
Description-Content-Type: text/markdown
|
|
19
|
+
Requires-Dist: jinja2>=3.1
|
|
20
|
+
Requires-Dist: pyyaml>=6.0
|
|
21
|
+
Requires-Dist: requests>=2.31
|
|
22
|
+
Requires-Dist: rich>=13.0
|
|
23
|
+
Requires-Dist: scriptconfig>=0.9
|
|
24
|
+
Requires-Dist: ubelt>=1.3
|
|
25
|
+
Provides-Extra: tests
|
|
26
|
+
Requires-Dist: pytest>=7.0; extra == "tests"
|
|
27
|
+
Requires-Dist: pytest-codeblocks>=0.17; extra == "tests"
|
|
28
|
+
|
|
29
|
+
# Infer Stack
|
|
30
|
+
|
|
31
|
+
`infer_stack` manages **named stack profiles** for local and Kubernetes-backed inference.
|
|
32
|
+
|
|
33
|
+
A stack profile is a small graph made from:
|
|
34
|
+
|
|
35
|
+
* **providers** — inference runtimes such as vLLM and Ollama
|
|
36
|
+
* **gateways** — optional API routers such as LiteLLM
|
|
37
|
+
* **frontends** — optional UIs such as Open WebUI
|
|
38
|
+
* **routes** — optional public model aliases exposed through a gateway
|
|
39
|
+
|
|
40
|
+
This repo can render those profiles through two backends:
|
|
41
|
+
|
|
42
|
+
* **Compose** for local single-host serving. Compose supports vLLM, Ollama, optional LiteLLM, and optional Open WebUI.
|
|
43
|
+
* **KubeAI** for Kubernetes-backed vLLM serving. KubeAI support is vLLM-only for now.
|
|
44
|
+
|
|
45
|
+
The direct Ollama path can run without LiteLLM and without predeclaring models. vLLM profiles still use explicit runtimes, placement, and runtime settings.
|
|
46
|
+
|
|
47
|
+
## Main commands
|
|
48
|
+
|
|
49
|
+
```bash
|
|
50
|
+
infer-stack setup --backend compose --profile ollama-direct
|
|
51
|
+
# or: infer-stack setup --backend compose --profile qwen2-5-7b-instruct-turbo-default
|
|
52
|
+
infer-stack list-profiles
|
|
53
|
+
infer-stack describe-profile <profile>
|
|
54
|
+
infer-stack validate
|
|
55
|
+
infer-stack render
|
|
56
|
+
infer-stack up -d
|
|
57
|
+
infer-stack deploy
|
|
58
|
+
infer-stack switch <profile> --apply # re-render and converge; no separate up needed
|
|
59
|
+
infer-stack status
|
|
60
|
+
infer-stack smoke-test
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
The CLI is built on [`scriptconfig`](https://gitlab.kitware.com/utils/scriptconfig),
|
|
64
|
+
so every subcommand is also importable as a Python class — useful for
|
|
65
|
+
notebooks, tests, and other scripts:
|
|
66
|
+
|
|
67
|
+
```python
|
|
68
|
+
from infer_stack.cli import RenderCLI, SmokeTestCLI
|
|
69
|
+
|
|
70
|
+
RenderCLI.main(argv=False, profile="qwen2-5-7b-instruct-turbo-default", yes=True)
|
|
71
|
+
SmokeTestCLI.main(argv=False, model="qwen/qwen2.5-7b-instruct-turbo")
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
`manage.py` and `infer-stack` are aliases for the same entry point;
|
|
75
|
+
shell examples below use `infer-stack`.
|
|
76
|
+
|
|
77
|
+
## Operating the rendered Compose stack
|
|
78
|
+
|
|
79
|
+
Once the stack is up, common docker compose operations are available as
|
|
80
|
+
`infer-stack` subcommands so you don't have to `cd` into the rendered
|
|
81
|
+
output directory or repeat the `-f docker-compose.yml --env-file .env`
|
|
82
|
+
flags. They all resolve the rendered location via the same
|
|
83
|
+
`output.generated_dir` chain as the rest of the CLI.
|
|
84
|
+
|
|
85
|
+
```bash
|
|
86
|
+
infer-stack ps # docker compose ps
|
|
87
|
+
infer-stack ps -a # include stopped
|
|
88
|
+
infer-stack logs -f open-webui # follow one service
|
|
89
|
+
infer-stack logs --tail=200 litellm vllm-* # tailored backlog
|
|
90
|
+
infer-stack restart open-webui # restart specific services
|
|
91
|
+
infer-stack stop # stop everything (no remove)
|
|
92
|
+
infer-stack start # start back up
|
|
93
|
+
infer-stack pull # refresh images
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
For Ollama model management inside the rendered Ollama service, prefer the
|
|
97
|
+
CLI wrappers:
|
|
98
|
+
|
|
99
|
+
```bash
|
|
100
|
+
infer-stack ollama-pull smollm2:135m
|
|
101
|
+
infer-stack ollama-list
|
|
102
|
+
infer-stack ollama-ps
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
For other interactive one-shot commands inside a container, use
|
|
106
|
+
`infer-stack logs`, `infer-stack ps`, `infer-stack restart`, or fall back to raw
|
|
107
|
+
Compose only when no wrapper exists.
|
|
108
|
+
|
|
109
|
+
On the KubeAI backend these wrappers raise ``NotImplementedError`` —
|
|
110
|
+
use the equivalent ``kubectl`` commands in the meantime.
|
|
111
|
+
|
|
112
|
+
## Inspect a profile before running it
|
|
113
|
+
|
|
114
|
+
```bash
|
|
115
|
+
infer-stack describe-profile qwen2-5-7b-instruct-turbo-default --format yaml
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
## Stack profile model
|
|
119
|
+
|
|
120
|
+
Profiles are written as stack graphs. The main sections are `providers`, `gateways`, `frontends`, and `routes`. For details and examples, see [docs/stack-graph-profiles.md](docs/stack-graph-profiles.md).
|
|
121
|
+
|
|
122
|
+
Common shapes:
|
|
123
|
+
|
|
124
|
+
```text
|
|
125
|
+
Open WebUI -> Ollama # ollama-direct, no LiteLLM
|
|
126
|
+
Open WebUI -> LiteLLM -> vLLM # classic vLLM compose profiles
|
|
127
|
+
Open WebUI -> LiteLLM -> Ollama # Ollama with stable aliases
|
|
128
|
+
Open WebUI -> LiteLLM -> Ollama + vLLM # mixed migration / test stacks
|
|
129
|
+
Ollama API + vLLM API directly # raw backend profiles
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
Custom provider models and custom profiles live in the configured `catalog.user_models_file`, which defaults to `~/.config/infer_stack/models.yaml`. New files should prefer provider-specific top-level keys:
|
|
133
|
+
|
|
134
|
+
```yaml
|
|
135
|
+
vllm_models:
|
|
136
|
+
my-vllm-model:
|
|
137
|
+
hf_model_id: org/model
|
|
138
|
+
|
|
139
|
+
ollama_models:
|
|
140
|
+
my-ollama-model:
|
|
141
|
+
tag: qwen3.5:4b
|
|
142
|
+
|
|
143
|
+
profiles:
|
|
144
|
+
my-stack:
|
|
145
|
+
providers: {}
|
|
146
|
+
gateways: {}
|
|
147
|
+
frontends: {}
|
|
148
|
+
routes: {}
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
`models:` is still interpreted as a vLLM model catalog for convenience, but new docs and recipes use `vllm_models:` / `ollama_models:`.
|
|
152
|
+
|
|
153
|
+
## Where config and rendered artifacts live
|
|
154
|
+
|
|
155
|
+
`infer-stack` follows XDG basedir conventions, so where you invoke it
|
|
156
|
+
from never changes which config it reads or where it writes rendered
|
|
157
|
+
artifacts:
|
|
158
|
+
|
|
159
|
+
There are exactly two path roots:
|
|
160
|
+
|
|
161
|
+
| What | Default location | How to relocate |
|
|
162
|
+
| --- | --- | --- |
|
|
163
|
+
| `config.yaml`, `models.yaml`, `kubeai-values.local.yaml` | `~/.config/infer_stack/` (resp. `$XDG_CONFIG_HOME`) | `--config-dir` (or `INFER_STACK_CONFIG_DIR`) |
|
|
164
|
+
| **Everything generated** — `generated/` (docker-compose.yml, .env, plan.yaml, kubeai/*) **and** `state/` (hf-cache, postgres volumes, Ollama store, runtime bind mounts) | `~/.local/share/infer_stack/` (resp. `$XDG_DATA_HOME`) | `--data-dir` (or `INFER_STACK_DATA_DIR`) |
|
|
165
|
+
|
|
166
|
+
`--data-dir` is the single knob for "put everything I generate in one
|
|
167
|
+
directory." Set it once at `setup`; it is baked into the absolute
|
|
168
|
+
`state.*` and `output.generated_dir` paths written to `config.yaml`, so
|
|
169
|
+
later commands don't need it again:
|
|
170
|
+
|
|
171
|
+
```bash
|
|
172
|
+
# All rendered artifacts and bind-mount state land under one directory.
|
|
173
|
+
infer-stack setup \
|
|
174
|
+
--backend compose \
|
|
175
|
+
--profile ollama-direct \
|
|
176
|
+
--data-dir /data/service/docker/vllm-stack
|
|
177
|
+
|
|
178
|
+
infer-stack render --yes
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
```bash
|
|
182
|
+
# Keep config.yaml in a checkout for ad-hoc experiments.
|
|
183
|
+
infer-stack setup --config-dir $PWD --backend compose --profile <p> --data-dir $PWD/stack
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
`--config-dir` / `--data-dir` live on every subcommand, so they appear
|
|
187
|
+
**after** the subcommand name. For "set once for the whole shell" use the
|
|
188
|
+
env vars instead. For a bespoke split layout (e.g. big `state/` on a data
|
|
189
|
+
disk, artifacts elsewhere), edit `state.*` / `output.generated_dir` in
|
|
190
|
+
`config.yaml` directly.
|
|
191
|
+
|
|
192
|
+
## Constraining placement to specific GPUs
|
|
193
|
+
|
|
194
|
+
If some of your GPUs are tied up by other work, restrict the planner
|
|
195
|
+
(and the rendered ``device_ids``) to the subset you want it to use:
|
|
196
|
+
|
|
197
|
+
```bash
|
|
198
|
+
# Only place onto GPU 1 (e.g. GPU 0 is running a display).
|
|
199
|
+
infer-stack render --yes --profile test-single-11gb --allowed-gpus 1
|
|
200
|
+
|
|
201
|
+
# Or pin a TP=2 profile to physical GPUs 1 and 3.
|
|
202
|
+
infer-stack render --yes --profile test-multi-gpu --allowed-gpus 1,3
|
|
203
|
+
```
|
|
204
|
+
|
|
205
|
+
``--allowed-gpus`` (or ``INFER_STACK_ALLOWED_GPUS=1,3``) filters the
|
|
206
|
+
detected inventory before placement — real indices are preserved, so
|
|
207
|
+
the rendered compose stack pins ``device_ids: ["1", "3"]`` to those
|
|
208
|
+
exact physical GPUs. Useful for integration tests that need to share a
|
|
209
|
+
host with other jobs.
|
|
210
|
+
|
|
211
|
+
## Demos / integration recipes
|
|
212
|
+
|
|
213
|
+
End-to-end examples under [docs/demos/](docs/demos/) are written as
|
|
214
|
+
markdown tutorials. The CI smoke test is runnable with pytest-codeblocks:
|
|
215
|
+
|
|
216
|
+
```bash
|
|
217
|
+
pytest --codeblocks docs/demos/ci_smoke_test.md
|
|
218
|
+
```
|
|
219
|
+
|
|
220
|
+
Each ``bash`` block is a self-contained shell snippet you can also
|
|
221
|
+
copy-paste into a terminal. See
|
|
222
|
+
[docs/demos/ci_smoke_test.md](docs/demos/ci_smoke_test.md) for the
|
|
223
|
+
``setup → describe → validate → render`` flow on the smallest test
|
|
224
|
+
profiles.
|
|
225
|
+
|
|
226
|
+
For a real running vLLM stack on a workstation, see
|
|
227
|
+
[docs/demos/quickstart.md](docs/demos/quickstart.md). For direct Ollama on a dual GTX 1080 Ti style host, see
|
|
228
|
+
[docs/demos/ollama_direct_quickstart.md](docs/demos/ollama_direct_quickstart.md). For a focused GPU-1 backend switch test, see [docs/demos/smollm2_gpu1_backend_switch.md](docs/demos/smollm2_gpu1_backend_switch.md).
|
|
229
|
+
|
|
230
|
+
User-supplied paths on the CLI (`--file`, `--from-file`,
|
|
231
|
+
`--resource-profiles-file`, `--output-dir`) still resolve against the
|
|
232
|
+
current working directory — they're meant to behave as typed.
|
|
233
|
+
|
|
234
|
+
---
|
|
235
|
+
|
|
236
|
+
## Backend 1: Compose
|
|
237
|
+
|
|
238
|
+
Use Compose for local single-host deployments. It can render direct Ollama stacks, vLLM stacks, mixed Ollama+vLLM stacks, and raw backend-only stacks.
|
|
239
|
+
|
|
240
|
+
### Getting started
|
|
241
|
+
|
|
242
|
+
Prerequisite: Docker and the `docker compose` plugin must be installed.
|
|
243
|
+
|
|
244
|
+
```bash
|
|
245
|
+
# Direct Ollama, no LiteLLM and no predeclared models.
|
|
246
|
+
infer-stack setup --backend compose --profile ollama-direct
|
|
247
|
+
infer-stack validate --simulate-hardware 2x11
|
|
248
|
+
infer-stack render --yes --simulate-hardware 2x11
|
|
249
|
+
infer-stack up -d
|
|
250
|
+
|
|
251
|
+
# Classic vLLM through LiteLLM/Open WebUI.
|
|
252
|
+
infer-stack setup --backend compose --profile qwen2-5-7b-instruct-turbo-default
|
|
253
|
+
infer-stack validate
|
|
254
|
+
infer-stack render
|
|
255
|
+
infer-stack up -d
|
|
256
|
+
```
|
|
257
|
+
|
|
258
|
+
### Test that it is responding
|
|
259
|
+
|
|
260
|
+
When LiteLLM is enabled, the default Compose front door is:
|
|
261
|
+
|
|
262
|
+
```text
|
|
263
|
+
http://127.0.0.1:14042/v1
|
|
264
|
+
```
|
|
265
|
+
|
|
266
|
+
When using `ollama-direct`, Open WebUI talks to Ollama directly and the Ollama API is available at:
|
|
267
|
+
|
|
268
|
+
```text
|
|
269
|
+
http://127.0.0.1:11434
|
|
270
|
+
http://127.0.0.1:11434/v1
|
|
271
|
+
```
|
|
272
|
+
|
|
273
|
+
unless you changed the relevant ports in config.
|
|
274
|
+
|
|
275
|
+
Wait until the active profile can serve a real request through its resolved default endpoint:
|
|
276
|
+
|
|
277
|
+
```bash
|
|
278
|
+
infer-stack wait-ready
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
`wait-ready` is stronger than Docker Compose health: it probes the user-facing
|
|
282
|
+
LiteLLM, Ollama, or direct vLLM access surface and, by default, requires a tiny
|
|
283
|
+
generation/completion to succeed. The smoke test runs this readiness probe by
|
|
284
|
+
default before issuing its normal test request:
|
|
285
|
+
|
|
286
|
+
```bash
|
|
287
|
+
infer-stack smoke-test
|
|
288
|
+
```
|
|
289
|
+
|
|
290
|
+
For direct Ollama profiles, pull a model first and then smoke-test that model:
|
|
291
|
+
|
|
292
|
+
```bash
|
|
293
|
+
infer-stack ollama-pull qwen3.5:4b
|
|
294
|
+
infer-stack ollama-list
|
|
295
|
+
infer-stack smoke-test --model qwen3.5:4b
|
|
296
|
+
```
|
|
297
|
+
|
|
298
|
+
For LiteLLM profiles, `smoke-test` reads the rendered `.env` automatically and
|
|
299
|
+
uses the active profile's resolved OpenAI-compatible front door. You can inspect
|
|
300
|
+
individual secrets when needed:
|
|
301
|
+
|
|
302
|
+
```bash
|
|
303
|
+
infer-stack env LITELLM_MASTER_KEY
|
|
304
|
+
infer-stack env VLLM_BACKEND_API_KEY
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
When you intentionally want the old quick behavior, skip the readiness wait:
|
|
308
|
+
|
|
309
|
+
```bash
|
|
310
|
+
infer-stack smoke-test --no-wait --model gpt2
|
|
311
|
+
```
|
|
312
|
+
|
|
313
|
+
### Stop it
|
|
314
|
+
|
|
315
|
+
```bash
|
|
316
|
+
infer-stack down
|
|
317
|
+
```
|
|
318
|
+
|
|
319
|
+
`down` never removes named volumes. The Postgres data directory and the
|
|
320
|
+
Open WebUI volume are preserved across `down`, `up`, `switch`, and `render`.
|
|
321
|
+
|
|
322
|
+
### Open WebUI authentication
|
|
323
|
+
|
|
324
|
+
By default Open WebUI runs with `WEBUI_AUTH=False` — no login screen,
|
|
325
|
+
anyone who can reach the port gets straight into the UI. This is the
|
|
326
|
+
expected behavior for a local dev box. To re-enable login/signup, set
|
|
327
|
+
in `config.yaml`:
|
|
328
|
+
|
|
329
|
+
```yaml
|
|
330
|
+
open_webui:
|
|
331
|
+
auth: true
|
|
332
|
+
```
|
|
333
|
+
|
|
334
|
+
and re-render. Existing accounts stored in the `postgres-open-webui`
|
|
335
|
+
volume are preserved across the toggle.
|
|
336
|
+
|
|
337
|
+
### Reverse proxy (TLS) and LDAP
|
|
338
|
+
|
|
339
|
+
Open WebUI can be fronted by an opt-in nginx TLS reverse proxy, and its
|
|
340
|
+
login can be backed by an LDAP directory. Both are off by default and
|
|
341
|
+
configured as ordinary config fields. The built-in `openwebui-tls-ldap`
|
|
342
|
+
profile wires them together as a worked example (Ollama + Open WebUI
|
|
343
|
+
behind nginx, no public Open WebUI/Ollama ports); see
|
|
344
|
+
[`examples/openwebui-tls-ldap/`](examples/openwebui-tls-ldap/).
|
|
345
|
+
|
|
346
|
+
```bash
|
|
347
|
+
infer-stack setup --backend compose --profile openwebui-tls-ldap
|
|
348
|
+
```
|
|
349
|
+
|
|
350
|
+
**Reverse proxy.** Enable it under `frontends.reverse_proxy`. It renders
|
|
351
|
+
an nginx service plus a generated `state.runtime/nginx.conf`:
|
|
352
|
+
|
|
353
|
+
```yaml
|
|
354
|
+
frontends:
|
|
355
|
+
reverse_proxy:
|
|
356
|
+
enabled: true
|
|
357
|
+
target: open_webui # or litellm / ollama / a custom upstream
|
|
358
|
+
server_name: host.example.com
|
|
359
|
+
ssl:
|
|
360
|
+
enabled: true
|
|
361
|
+
certificate: ./certs/site.crt
|
|
362
|
+
certificate_key: ./certs/site.key
|
|
363
|
+
dhparam: ./dhparam.pem # optional
|
|
364
|
+
```
|
|
365
|
+
|
|
366
|
+
When `ssl.enabled` is true, port 80 redirects to HTTPS (`force_https`)
|
|
367
|
+
and the cert/key/dhparam host paths are bind-mounted read-only. When
|
|
368
|
+
`ssl.enabled` is false, only HTTP is published (HTTPS publishing is
|
|
369
|
+
gated on TLS so you never get a `:443` mapping with nothing listening).
|
|
370
|
+
|
|
371
|
+
> **Path caveat.** Relative `certificate`/`certificate_key`/`dhparam`/
|
|
372
|
+
> `config_path` values are written verbatim into the generated
|
|
373
|
+
> `docker-compose.yml`, so Docker Compose resolves them **relative to the
|
|
374
|
+
> generated directory** (where the compose file lives), not your CWD.
|
|
375
|
+
> `infer-stack render` warns when a referenced cert or config file is not
|
|
376
|
+
> found. Use absolute paths if you want to avoid the ambiguity.
|
|
377
|
+
|
|
378
|
+
**LDAP.** Enable it under `frontends.open_webui.ldap`. The directory
|
|
379
|
+
settings render as Open WebUI `LDAP_*` environment variables, and
|
|
380
|
+
secrets/site-specific values are emitted as `.env` placeholders
|
|
381
|
+
(`LDAP_HOST`, `LDAP_PASSWD`, `LDAP_SEARCH_BASE`, …) so you can fill them
|
|
382
|
+
in after the first render without re-touching the compose YAML:
|
|
383
|
+
|
|
384
|
+
```yaml
|
|
385
|
+
frontends:
|
|
386
|
+
open_webui:
|
|
387
|
+
ldap:
|
|
388
|
+
enabled: true
|
|
389
|
+
env_defaults:
|
|
390
|
+
LDAP_PORT: '636'
|
|
391
|
+
LDAP_USE_TLS: 'true'
|
|
392
|
+
LDAP_ATTRIBUTE_FOR_USERNAME: uid
|
|
393
|
+
```
|
|
394
|
+
|
|
395
|
+
**Manual escape hatches.** When the typed renderer is not enough, drop
|
|
396
|
+
to manual control without leaving infer-stack:
|
|
397
|
+
|
|
398
|
+
* `frontends.reverse_proxy.config_path` — mount an existing nginx config
|
|
399
|
+
file instead of rendering one.
|
|
400
|
+
* `frontends.reverse_proxy.extra_config` — inject extra directives into
|
|
401
|
+
the rendered HTTPS `server` block.
|
|
402
|
+
* Every rendered service (`ollama`, vLLM runtimes, `litellm`,
|
|
403
|
+
`open_webui`, `reverse_proxy`) accepts generic overrides:
|
|
404
|
+
`extra_env`, `env_file`, `extra_volumes`, `extra_hosts`, `labels`,
|
|
405
|
+
`additional_ports`, and `gpus` (scalar `all`/count or a structured
|
|
406
|
+
device-request list).
|
|
407
|
+
|
|
408
|
+
Field precedence (lowest to highest) is: top-level config section
|
|
409
|
+
(`reverse_proxy:` / `open_webui:` / `ollama:`) → the matching
|
|
410
|
+
`frontends.*` / `providers.*` / `gateways.*` section → the active
|
|
411
|
+
profile. Newer configs should prefer the `frontends.*` / `providers.*`
|
|
412
|
+
form shown above.
|
|
413
|
+
|
|
414
|
+
### Persistent state and database layout
|
|
415
|
+
|
|
416
|
+
Compose renders stateful services only when their components are enabled:
|
|
417
|
+
|
|
418
|
+
* `postgres-open-webui` — rendered only when Open WebUI is enabled. It stores chats, accounts, and settings in `state.postgres_open_webui`.
|
|
419
|
+
* `postgres-litellm` — rendered only when LiteLLM is enabled. It stores router state in `state.postgres_litellm`.
|
|
420
|
+
* `ollama` — rendered only when the Ollama provider is enabled. Its model store is `state.ollama`, mounted at `/root/.ollama`.
|
|
421
|
+
* vLLM runtimes mount `state.hf_cache` for Hugging Face weights and `state.vllm_cache` for compiled artifacts.
|
|
422
|
+
|
|
423
|
+
Each Postgres container has its own `POSTGRES_DB`, `POSTGRES_USER`, and
|
|
424
|
+
`POSTGRES_PASSWORD`, sourced from component-specific `.env` keys. There is no shared Postgres instance and no `postgres-init` bootstrap service.
|
|
425
|
+
|
|
426
|
+
Open WebUI chat history is **not** tied to the model currently being served, so after a profile switch old chats may reference model IDs the current gateway no longer advertises — that is expected.
|
|
427
|
+
|
|
428
|
+
### Operational tips
|
|
429
|
+
|
|
430
|
+
Prefer scoping commands to specific services rather than relying on
|
|
431
|
+
container names. Use only the services rendered by the active profile:
|
|
432
|
+
|
|
433
|
+
```bash
|
|
434
|
+
# LiteLLM gateway profile
|
|
435
|
+
infer-stack logs -f litellm
|
|
436
|
+
|
|
437
|
+
# Direct Ollama profile
|
|
438
|
+
infer-stack logs -f ollama
|
|
439
|
+
|
|
440
|
+
# Ollama model store helpers
|
|
441
|
+
infer-stack ollama-list
|
|
442
|
+
infer-stack ollama-ps
|
|
443
|
+
```
|
|
444
|
+
|
|
445
|
+
You do not need to delete any volume during normal operation. If you
|
|
446
|
+
ever want a destructive reset, do it explicitly with
|
|
447
|
+
`docker compose down -v` against `generated/docker-compose.yml` — the
|
|
448
|
+
toolchain itself never does this.
|
|
449
|
+
|
|
450
|
+
### Custom .env values are preserved
|
|
451
|
+
|
|
452
|
+
`generated/.env` is rewritten non-destructively. Any `KEY=value` pair
|
|
453
|
+
you add manually (for example `VERBOSE=1`, `HF_HOME=/data/hf`, or any
|
|
454
|
+
key this program does not yet know about) is preserved across
|
|
455
|
+
`render`, `setup`, `switch`, `up`, and `deploy`. Comments and the order
|
|
456
|
+
of existing lines are preserved where practical.
|
|
457
|
+
|
|
458
|
+
### Switching profiles
|
|
459
|
+
|
|
460
|
+
```bash
|
|
461
|
+
infer-stack switch <profile> --apply
|
|
462
|
+
```
|
|
463
|
+
|
|
464
|
+
`switch --apply` re-renders from the updated `config.yaml`, then brings the
|
|
465
|
+
stack up convergently with `--remove-orphans` so a separate `infer-stack up` is
|
|
466
|
+
not needed. Components/runtimes that are no longer in the rendered compose file
|
|
467
|
+
are dropped. Compose preserves existing containers whose service definitions did
|
|
468
|
+
not change. For vLLM-to-vLLM profile switches, unchanged Open WebUI stays up;
|
|
469
|
+
LiteLLM is refreshed through its admin API when possible. The live refresh
|
|
470
|
+
path treats LiteLLM's "model not found in db" response for config-backed
|
|
471
|
+
models as non-fatal, so switching aliases can add the new route without
|
|
472
|
+
tearing LiteLLM down. That can temporarily leave stale config-backed aliases
|
|
473
|
+
in `/v1/models`; restart LiteLLM manually only when you want to clean those
|
|
474
|
+
up. If Compose already created or recreated LiteLLM while converging the new
|
|
475
|
+
stack, no extra router refresh is attempted because the new container has
|
|
476
|
+
already loaded the freshly rendered YAML. Profiles that do not render
|
|
477
|
+
LiteLLM, such as direct Ollama profiles, skip the router refresh path even if
|
|
478
|
+
an old `runtime/litellm_config.yaml` file remains from a previous profile.
|
|
479
|
+
Switches that change Open WebUI's provider wiring, such as
|
|
480
|
+
`Open WebUI -> LiteLLM` to `Open WebUI -> Ollama`, necessarily recreate
|
|
481
|
+
Open WebUI because its environment changes. Postgres volumes and provider
|
|
482
|
+
caches are left untouched. vLLM runtime containers are named after their
|
|
483
|
+
Compose service, for example `vllm-chat`, so `docker ps` and
|
|
484
|
+
`infer-stack logs vllm-chat` clearly identify them as vLLM containers.
|
|
485
|
+
|
|
486
|
+
### Protocol modes for base vs. instruct models
|
|
487
|
+
|
|
488
|
+
Profiles declare a `protocol_mode` (`chat` or `completions`) that the
|
|
489
|
+
served model must support. Models also declare which protocols they
|
|
490
|
+
support via `supported_protocols`. Validation runs before render and
|
|
491
|
+
fails with an actionable message if a profile asks for `chat` on a
|
|
492
|
+
completions-only model.
|
|
493
|
+
|
|
494
|
+
Practical guidance:
|
|
495
|
+
|
|
496
|
+
* Instruct/chat models (with a chat template) can use either, but
|
|
497
|
+
default to `chat`.
|
|
498
|
+
* Base models like Pythia, Llama-2 base, Mistral-v0.1 base, and Falcon
|
|
499
|
+
base do not define a chat template. Their HELM profiles use
|
|
500
|
+
`protocol_mode: completions` and the `smoke-test` command will
|
|
501
|
+
exercise `/v1/completions` for them.
|
|
502
|
+
* The rendered LiteLLM config uses `text-completion-openai/<served>`
|
|
503
|
+
as the upstream provider for completions-only services. That means
|
|
504
|
+
even chat-shaped requests sent through Open WebUI to a Pythia model
|
|
505
|
+
get translated by LiteLLM into upstream `/v1/completions` calls — no
|
|
506
|
+
second vLLM container is needed to support Open WebUI for a
|
|
507
|
+
completions-only model.
|
|
508
|
+
* Open WebUI is still a chat UI, so prompt formatting matters.
|
|
509
|
+
HELM/eval clients should call `/v1/completions` directly for exact
|
|
510
|
+
prompt control rather than going through the chat frontend.
|
|
511
|
+
|
|
512
|
+
### Chat-shaped clients on top of completions models
|
|
513
|
+
|
|
514
|
+
Some clients (e.g. InspectAI / Inspect Evals stock MMLU tasks) only
|
|
515
|
+
speak `/v1/chat/completions` and cannot be reconfigured. For those
|
|
516
|
+
cases, profiles can opt into a LiteLLM-only adapter:
|
|
517
|
+
|
|
518
|
+
```yaml
|
|
519
|
+
chat_compat:
|
|
520
|
+
enabled: true
|
|
521
|
+
strategy: flat_messages
|
|
522
|
+
```
|
|
523
|
+
|
|
524
|
+
When set on a `protocol_mode: completions` service, the rendered
|
|
525
|
+
LiteLLM config keeps the `text-completion-openai/<served>` upstream
|
|
526
|
+
and adds LiteLLM's documented prompt-template fields
|
|
527
|
+
(`initial_prompt_value` / `roles` / `final_prompt_value`) so chat
|
|
528
|
+
messages get flattened into a plain prompt — no role labels, messages
|
|
529
|
+
joined by `\n` — before being forwarded to vLLM `/v1/completions`.
|
|
530
|
+
|
|
531
|
+
This is **not** a chat tune; the model is still a base model and
|
|
532
|
+
prompt formatting still matters for evaluation. Use it only when a
|
|
533
|
+
chat-shaped client cannot be changed. The vLLM container is not
|
|
534
|
+
restarted, no `--chat-template` is rendered, and the adapter takes
|
|
535
|
+
effect after a `litellm`-only restart:
|
|
536
|
+
|
|
537
|
+
```bash
|
|
538
|
+
infer-stack render
|
|
539
|
+
infer-stack restart litellm
|
|
540
|
+
```
|
|
541
|
+
|
|
542
|
+
The built-in `pythia-inspect-mmlu-compat` profile is a ready-made
|
|
543
|
+
example; see
|
|
544
|
+
[`recipies/compose_pythia_inspect_mmlu_compat.md`](recipies/compose_pythia_inspect_mmlu_compat.md).
|
|
545
|
+
|
|
546
|
+
### Reasoning / thinking models
|
|
547
|
+
|
|
548
|
+
Models can declare reasoning support in the catalog:
|
|
549
|
+
|
|
550
|
+
```yaml
|
|
551
|
+
reasoning:
|
|
552
|
+
enabled: true
|
|
553
|
+
parser: qwen3
|
|
554
|
+
expose_to_openwebui: true
|
|
555
|
+
```
|
|
556
|
+
|
|
557
|
+
Profiles can override or set the same field per service. When a
|
|
558
|
+
service has `reasoning.enabled: true` and a `parser`, the renderer
|
|
559
|
+
adds `--reasoning-parser <parser>` to that vLLM container's command
|
|
560
|
+
line — that flag alone enables reasoning extraction in the current
|
|
561
|
+
vLLM CLI. You do not need to repeat it by hand in `extra_args`.
|
|
562
|
+
|
|
563
|
+
Open WebUI sees reasoning content via two paths:
|
|
564
|
+
|
|
565
|
+
1. Inline `<think>...</think>` tags emitted by the model.
|
|
566
|
+
2. Structured `reasoning_content` fields when LiteLLM normalizes them.
|
|
567
|
+
|
|
568
|
+
The LiteLLM template keeps `merge_reasoning_content_in_choices: true`
|
|
569
|
+
on chat-mode entries so Open WebUI can display reasoning in the
|
|
570
|
+
streamed response. To test reasoning end-to-end:
|
|
571
|
+
|
|
572
|
+
```bash
|
|
573
|
+
# Non-streaming CLI smoke test:
|
|
574
|
+
infer-stack smoke-test \
|
|
575
|
+
--model qwen3.6-35b-a3b \
|
|
576
|
+
--prompt "Think step by step: 17*23"
|
|
577
|
+
|
|
578
|
+
# For streaming inspection, read the key with the CLI wrapper:
|
|
579
|
+
LITELLM_MASTER_KEY=$(infer-stack env LITELLM_MASTER_KEY)
|
|
580
|
+
curl -N http://127.0.0.1:14042/v1/chat/completions \
|
|
581
|
+
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
|
|
582
|
+
-H 'Content-Type: application/json' \
|
|
583
|
+
-d '{"model":"qwen3.6-35b-a3b","stream":true,
|
|
584
|
+
"messages":[{"role":"user","content":"Think step by step: 17*23"}]}'
|
|
585
|
+
```
|
|
586
|
+
|
|
587
|
+
In Open WebUI, reasoning shows up best with streaming enabled in the
|
|
588
|
+
chat settings.
|
|
589
|
+
|
|
590
|
+
---
|
|
591
|
+
|
|
592
|
+
## Backend 2: KubeAI
|
|
593
|
+
|
|
594
|
+
Use KubeAI when you want Kubernetes-managed serving.
|
|
595
|
+
|
|
596
|
+
### Important rules
|
|
597
|
+
|
|
598
|
+
1. **Use the same namespace everywhere.** The namespace in `infer-stack setup --namespace ...` must match the namespace where the KubeAI Helm release already exists.
|
|
599
|
+
2. **Prefer the repo-driven path.** The normal path is `setup` -> `validate` -> `render` -> `deploy` -> `status`.
|
|
600
|
+
3. **`kubectl port-forward` stays in the foreground.** Leave it running in one terminal and send requests from another.
|
|
601
|
+
4. **The first request can take a while.** `/openai/v1/models` may work before chat completions work. The first completion may trigger pod creation, image pull, model load, and compile warmup.
|
|
602
|
+
5. **On the current repo version, KubeAI still needs a live workaround after deploy.** The renderer currently produces a `Model` spec that needs a small manual patch to work with the KubeAI version used in these notes.
|
|
603
|
+
|
|
604
|
+
### KubeAI prerequisites
|
|
605
|
+
|
|
606
|
+
You need:
|
|
607
|
+
|
|
608
|
+
* a working Kubernetes cluster
|
|
609
|
+
* `kubectl`
|
|
610
|
+
* Helm
|
|
611
|
+
|
|
612
|
+
If you want a quick local single-node cluster, K3s is a good option.
|
|
613
|
+
|
|
614
|
+
Install K3s:
|
|
615
|
+
|
|
616
|
+
```bash
|
|
617
|
+
curl -sfL https://get.k3s.io | sh -
|
|
618
|
+
# or pin
|
|
619
|
+
curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION='v1.34.3+k3s1' sh -
|
|
620
|
+
```
|
|
621
|
+
|
|
622
|
+
Make `kubectl` usable without `sudo`:
|
|
623
|
+
|
|
624
|
+
```bash
|
|
625
|
+
sudo mkdir -p /etc/rancher/k3s/config.yaml.d
|
|
626
|
+
printf 'write-kubeconfig-mode: "0644"\n' | \
|
|
627
|
+
sudo tee /etc/rancher/k3s/config.yaml.d/10-kubeconfig-mode.yaml >/dev/null
|
|
628
|
+
sudo systemctl restart k3s
|
|
629
|
+
kubectl get nodes
|
|
630
|
+
```
|
|
631
|
+
|
|
632
|
+
Install Helm:
|
|
633
|
+
|
|
634
|
+
```bash
|
|
635
|
+
curl -fsSL -o get_helm.sh https://raw.githubusercontent.com/helm/helm/83a46119086589a593a62ca544982977a60318ca/scripts/get-helm-4
|
|
636
|
+
chmod 700 get_helm.sh
|
|
637
|
+
./get_helm.sh
|
|
638
|
+
helm version
|
|
639
|
+
```
|
|
640
|
+
|
|
641
|
+
### NVIDIA GPU support
|
|
642
|
+
|
|
643
|
+
Install the NVIDIA device plugin and GPU Feature Discovery so Kubernetes can expose GPU resources and labels:
|
|
644
|
+
|
|
645
|
+
```bash
|
|
646
|
+
helm repo add nvdp https://nvidia.github.io/k8s-device-plugin
|
|
647
|
+
helm repo update
|
|
648
|
+
helm upgrade -i nvdp nvdp/nvidia-device-plugin \
|
|
649
|
+
--version 0.17.1 \
|
|
650
|
+
--namespace nvidia-device-plugin \
|
|
651
|
+
--create-namespace \
|
|
652
|
+
--set gfd.enabled=true \
|
|
653
|
+
--set runtimeClassName=nvidia
|
|
654
|
+
```
|
|
655
|
+
|
|
656
|
+
Check that GPU support is working:
|
|
657
|
+
|
|
658
|
+
```bash
|
|
659
|
+
kubectl -n nvidia-device-plugin get pods
|
|
660
|
+
kubectl get node "$(kubectl get nodes -o jsonpath='{.items[0].metadata.name}')" \
|
|
661
|
+
-o jsonpath='{.status.allocatable.nvidia\.com/gpu}{"\n"}'
|
|
662
|
+
kubectl get nodes --show-labels | tr ',' '\n' | grep 'nvidia.com/' || true
|
|
663
|
+
```
|
|
664
|
+
|
|
665
|
+
You want to see a non-empty `nvidia.com/gpu` count and `nvidia.com/*` labels such as product and memory.
|
|
666
|
+
|
|
667
|
+
### KubeAI Helm repository
|
|
668
|
+
|
|
669
|
+
```bash
|
|
670
|
+
helm repo add kubeai https://www.kubeai.org
|
|
671
|
+
helm repo update
|
|
672
|
+
```
|
|
673
|
+
|
|
674
|
+
## Determine which namespace to use
|
|
675
|
+
|
|
676
|
+
Before doing anything else, discover whether a `kubeai` release already exists and which namespace it uses.
|
|
677
|
+
|
|
678
|
+
```bash
|
|
679
|
+
KUBEAI_NAMESPACE="$(helm list -A | awk '$1=="kubeai"{print $2; exit}')"
|
|
680
|
+
if [ -z "${KUBEAI_NAMESPACE}" ]; then
|
|
681
|
+
KUBEAI_NAMESPACE=default
|
|
682
|
+
fi
|
|
683
|
+
echo "Using KubeAI namespace: ${KUBEAI_NAMESPACE}"
|
|
684
|
+
```
|
|
685
|
+
|
|
686
|
+
If a release already exists, **reuse that namespace**.
|
|
687
|
+
|
|
688
|
+
Sanity-check the cluster:
|
|
689
|
+
|
|
690
|
+
```bash
|
|
691
|
+
kubectl get nodes
|
|
692
|
+
kubectl get crd models.kubeai.org || true
|
|
693
|
+
helm list -A | grep kubeai || true
|
|
694
|
+
kubectl -n "${KUBEAI_NAMESPACE}" get pods || true
|
|
695
|
+
kubectl get node "$(kubectl get nodes -o jsonpath='{.items[0].metadata.name}')" \
|
|
696
|
+
-o jsonpath='{.status.allocatable.nvidia\.com/gpu}{"\n"}'
|
|
697
|
+
```
|
|
698
|
+
|
|
699
|
+
---
|
|
700
|
+
|
|
701
|
+
## Generate the local KubeAI resource-profile file
|
|
702
|
+
|
|
703
|
+
Generate a local KubeAI resource-profile file from the labels on this machine.
|
|
704
|
+
|
|
705
|
+
For the built-in serving profiles in this repo, keep these names aligned:
|
|
706
|
+
|
|
707
|
+
* `gpu-single-default`
|
|
708
|
+
* `gpu-tp2-balanced`
|
|
709
|
+
* `gpu-tp2-maxctx`
|
|
710
|
+
|
|
711
|
+
**Important:** include GPU `requests`, GPU `limits`, and `runtimeClassName: nvidia`. Without those, the model pod can land on the GPU node but still start without `libcuda.so.1` available inside the container.
|
|
712
|
+
|
|
713
|
+
```bash
|
|
714
|
+
PRODUCT="$(kubectl get nodes -o jsonpath='{.items[0].metadata.labels.nvidia\.com/gpu\.product}')"
|
|
715
|
+
MEMORY="$(kubectl get nodes -o jsonpath='{.items[0].metadata.labels.nvidia\.com/gpu\.memory}')"
|
|
716
|
+
|
|
717
|
+
cat > values-kubeai-local-gpu.yaml <<EOF
|
|
718
|
+
resourceProfiles:
|
|
719
|
+
gpu-single-default:
|
|
720
|
+
nodeSelector:
|
|
721
|
+
nvidia.com/gpu.product: "${PRODUCT}"
|
|
722
|
+
nvidia.com/gpu.memory: "${MEMORY}"
|
|
723
|
+
requests:
|
|
724
|
+
nvidia.com/gpu: 1
|
|
725
|
+
limits:
|
|
726
|
+
nvidia.com/gpu: 1
|
|
727
|
+
runtimeClassName: nvidia
|
|
728
|
+
|
|
729
|
+
gpu-tp2-balanced:
|
|
730
|
+
nodeSelector:
|
|
731
|
+
nvidia.com/gpu.product: "${PRODUCT}"
|
|
732
|
+
nvidia.com/gpu.memory: "${MEMORY}"
|
|
733
|
+
requests:
|
|
734
|
+
nvidia.com/gpu: 2
|
|
735
|
+
limits:
|
|
736
|
+
nvidia.com/gpu: 2
|
|
737
|
+
runtimeClassName: nvidia
|
|
738
|
+
|
|
739
|
+
gpu-tp2-maxctx:
|
|
740
|
+
nodeSelector:
|
|
741
|
+
nvidia.com/gpu.product: "${PRODUCT}"
|
|
742
|
+
nvidia.com/gpu.memory: "${MEMORY}"
|
|
743
|
+
requests:
|
|
744
|
+
nvidia.com/gpu: 2
|
|
745
|
+
limits:
|
|
746
|
+
nvidia.com/gpu: 2
|
|
747
|
+
runtimeClassName: nvidia
|
|
748
|
+
EOF
|
|
749
|
+
|
|
750
|
+
cat values-kubeai-local-gpu.yaml
|
|
751
|
+
```
|
|
752
|
+
|
|
753
|
+
Sync that file so `validate` and `render` use the same local resource-profile data:
|
|
754
|
+
|
|
755
|
+
```bash
|
|
756
|
+
infer-stack kubeai-sync-resource-profiles --from-file values-kubeai-local-gpu.yaml
|
|
757
|
+
```
|
|
758
|
+
|
|
759
|
+
---
|
|
760
|
+
|
|
761
|
+
## Example 1: single-GPU system
|
|
762
|
+
|
|
763
|
+
Use this example on a 1-GPU workstation.
|
|
764
|
+
|
|
765
|
+
```bash
|
|
766
|
+
infer-stack setup \
|
|
767
|
+
--backend kubeai \
|
|
768
|
+
--profile qwen2-5-7b-instruct-turbo-default \
|
|
769
|
+
--namespace "${KUBEAI_NAMESPACE}"
|
|
770
|
+
|
|
771
|
+
infer-stack list-profiles
|
|
772
|
+
infer-stack describe-profile qwen2-5-7b-instruct-turbo-default --format yaml
|
|
773
|
+
infer-stack validate
|
|
774
|
+
infer-stack render
|
|
775
|
+
infer-stack deploy
|
|
776
|
+
infer-stack status
|
|
777
|
+
```
|
|
778
|
+
|
|
779
|
+
### Current live workaround for the single-GPU example
|
|
780
|
+
|
|
781
|
+
On the current repo version, apply this live patch after `deploy`.
|
|
782
|
+
|
|
783
|
+
This patch does four things:
|
|
784
|
+
|
|
785
|
+
* keeps the model warm with `minReplicas: 1`
|
|
786
|
+
* changes `resourceProfile` from `gpu-single-default` to `gpu-single-default:1`
|
|
787
|
+
* makes the served model name match the public profile name
|
|
788
|
+
* avoids the duplicate `--served-model-name` mismatch that causes 404s on completions
|
|
789
|
+
|
|
790
|
+
```bash
|
|
791
|
+
kubectl -n "${KUBEAI_NAMESPACE}" patch model qwen2-5-7b-instruct-turbo-default --type merge -p '{
|
|
792
|
+
"spec": {
|
|
793
|
+
"minReplicas": 1,
|
|
794
|
+
"resourceProfile": "gpu-single-default:1",
|
|
795
|
+
"args": [
|
|
796
|
+
"--served-model-name=qwen2-5-7b-instruct-turbo-default",
|
|
797
|
+
"--tensor-parallel-size=1",
|
|
798
|
+
"--data-parallel-size=1",
|
|
799
|
+
"--max-model-len=32768",
|
|
800
|
+
"--gpu-memory-utilization=0.9",
|
|
801
|
+
"--max-num-batched-tokens=8192",
|
|
802
|
+
"--max-num-seqs=16",
|
|
803
|
+
"--disable-log-requests",
|
|
804
|
+
"--enable-prefix-caching"
|
|
805
|
+
]
|
|
806
|
+
}
|
|
807
|
+
}'
|
|
808
|
+
|
|
809
|
+
kubectl -n "${KUBEAI_NAMESPACE}" delete pod -l model=qwen2-5-7b-instruct-turbo-default
|
|
810
|
+
```
|
|
811
|
+
|
|
812
|
+
If you run `infer-stack render` or `infer-stack deploy` again on the current repo version, re-apply this live patch.
|
|
813
|
+
|
|
814
|
+
---
|
|
815
|
+
|
|
816
|
+
## Example 2: four-GPU system
|
|
817
|
+
|
|
818
|
+
On a 4-GPU host, do the same **single-GPU smoke test first** to verify the cluster, KubeAI, runtime class, and model plumbing. That exact sequence worked on a 4-GPU machine during bring-up.
|
|
819
|
+
|
|
820
|
+
```bash
|
|
821
|
+
infer-stack setup \
|
|
822
|
+
--backend kubeai \
|
|
823
|
+
--profile qwen2-5-7b-instruct-turbo-default \
|
|
824
|
+
--namespace "${KUBEAI_NAMESPACE}"
|
|
825
|
+
|
|
826
|
+
infer-stack validate
|
|
827
|
+
infer-stack render
|
|
828
|
+
infer-stack deploy
|
|
829
|
+
infer-stack status
|
|
830
|
+
|
|
831
|
+
kubectl -n "${KUBEAI_NAMESPACE}" patch model qwen2-5-7b-instruct-turbo-default --type merge -p '{
|
|
832
|
+
"spec": {
|
|
833
|
+
"minReplicas": 1,
|
|
834
|
+
"resourceProfile": "gpu-single-default:1",
|
|
835
|
+
"args": [
|
|
836
|
+
"--served-model-name=qwen2-5-7b-instruct-turbo-default",
|
|
837
|
+
"--tensor-parallel-size=1",
|
|
838
|
+
"--data-parallel-size=1",
|
|
839
|
+
"--max-model-len=32768",
|
|
840
|
+
"--gpu-memory-utilization=0.9",
|
|
841
|
+
"--max-num-batched-tokens=8192",
|
|
842
|
+
"--max-num-seqs=16",
|
|
843
|
+
"--disable-log-requests",
|
|
844
|
+
"--enable-prefix-caching"
|
|
845
|
+
]
|
|
846
|
+
}
|
|
847
|
+
}'
|
|
848
|
+
|
|
849
|
+
kubectl -n "${KUBEAI_NAMESPACE}" delete pod -l model=qwen2-5-7b-instruct-turbo-default
|
|
850
|
+
```
|
|
851
|
+
|
|
852
|
+
After the 7B smoke test works, move up to larger profiles such as `qwen2-72b-instruct-tp2-balanced`. On the current repo version, apply the same kind of live patch after deploy: keep `minReplicas: 1`, append `:1` to the chosen `resourceProfile`, and make the single effective `--served-model-name` match the public profile name.
|
|
853
|
+
|
|
854
|
+
---
|
|
855
|
+
|
|
856
|
+
## Test that KubeAI is responding
|
|
857
|
+
|
|
858
|
+
If you are not exposing ingress yet, port-forward the service.
|
|
859
|
+
|
|
860
|
+
**This command stays in the foreground.** Run it in one terminal and leave it there:
|
|
861
|
+
|
|
862
|
+
```bash
|
|
863
|
+
kubectl -n "${KUBEAI_NAMESPACE}" port-forward svc/kubeai 8000:80
|
|
864
|
+
```
|
|
865
|
+
|
|
866
|
+
Then use another terminal for requests.
|
|
867
|
+
|
|
868
|
+
### First check: `/models`
|
|
869
|
+
|
|
870
|
+
```bash
|
|
871
|
+
curl http://127.0.0.1:8000/openai/v1/models
|
|
872
|
+
```
|
|
873
|
+
|
|
874
|
+
If that works, the KubeAI front door is alive.
|
|
875
|
+
|
|
876
|
+
### Then try the smoke test
|
|
877
|
+
|
|
878
|
+
```bash
|
|
879
|
+
infer-stack smoke-test \
|
|
880
|
+
--base-url http://127.0.0.1:8000/openai/v1 \
|
|
881
|
+
--model qwen2-5-7b-instruct-turbo-default
|
|
882
|
+
```
|
|
883
|
+
|
|
884
|
+
### Or test chat completions directly
|
|
885
|
+
|
|
886
|
+
```bash
|
|
887
|
+
time curl http://127.0.0.1:8000/openai/v1/chat/completions \
|
|
888
|
+
-H 'Content-Type: application/json' \
|
|
889
|
+
-d '{
|
|
890
|
+
"model": "qwen2-5-7b-instruct-turbo-default",
|
|
891
|
+
"messages": [{"role": "user", "content": "Say hello in one short sentence."}],
|
|
892
|
+
"max_tokens": 8
|
|
893
|
+
}'
|
|
894
|
+
```
|
|
895
|
+
|
|
896
|
+
### What to expect on the first request
|
|
897
|
+
|
|
898
|
+
Common first-request behavior:
|
|
899
|
+
|
|
900
|
+
* `/openai/v1/models` works before completions work
|
|
901
|
+
* a completion request causes KubeAI to create a model-serving pod
|
|
902
|
+
* that pod may spend time in `ContainerCreating` while the image is pulled
|
|
903
|
+
* the model then spends more time loading and warming up
|
|
904
|
+
* the first completion can be much slower than later ones
|
|
905
|
+
|
|
906
|
+
That is not automatically a failure. Watch the system state while the first request is happening:
|
|
907
|
+
|
|
908
|
+
```bash
|
|
909
|
+
watch -n 1 'kubectl -n '"${KUBEAI_NAMESPACE}"' get pods; echo; kubectl -n '"${KUBEAI_NAMESPACE}"' get models'
|
|
910
|
+
```
|
|
911
|
+
|
|
912
|
+
---
|
|
913
|
+
|
|
914
|
+
## Debugging checks
|
|
915
|
+
|
|
916
|
+
### Check the live Model object
|
|
917
|
+
|
|
918
|
+
```bash
|
|
919
|
+
kubectl -n "${KUBEAI_NAMESPACE}" describe model qwen2-5-7b-instruct-turbo-default
|
|
920
|
+
kubectl -n "${KUBEAI_NAMESPACE}" get model qwen2-5-7b-instruct-turbo-default -o yaml | grep -E 'minReplicas|maxReplicas|resourceProfile'
|
|
921
|
+
```
|
|
922
|
+
|
|
923
|
+
### Check the current model pod
|
|
924
|
+
|
|
925
|
+
```bash
|
|
926
|
+
kubectl -n "${KUBEAI_NAMESPACE}" describe pod "$(kubectl -n "${KUBEAI_NAMESPACE}" get pods -o name | grep 'model-qwen2-5-7b-instruct-turbo-default' | tail -n 1 | cut -d/ -f2)"
|
|
927
|
+
```
|
|
928
|
+
|
|
929
|
+
### Tail KubeAI controller logs
|
|
930
|
+
|
|
931
|
+
```bash
|
|
932
|
+
kubectl -n "${KUBEAI_NAMESPACE}" logs deploy/kubeai --tail=200 -f
|
|
933
|
+
```
|
|
934
|
+
|
|
935
|
+
### Tail model-server logs
|
|
936
|
+
|
|
937
|
+
```bash
|
|
938
|
+
kubectl -n "${KUBEAI_NAMESPACE}" logs -f "$(kubectl -n "${KUBEAI_NAMESPACE}" get pods -o name | grep 'model-qwen2-5-7b-instruct-turbo-default' | tail -n 1 | cut -d/ -f2)" -c server
|
|
939
|
+
```
|
|
940
|
+
|
|
941
|
+
If the model pod restarted, inspect the previous crash:
|
|
942
|
+
|
|
943
|
+
```bash
|
|
944
|
+
kubectl -n "${KUBEAI_NAMESPACE}" logs "$(kubectl -n "${KUBEAI_NAMESPACE}" get pods -o name | grep 'model-qwen2-5-7b-instruct-turbo-default' | tail -n 1 | cut -d/ -f2)" -c server --previous
|
|
945
|
+
```
|
|
946
|
+
|
|
947
|
+
### Check recent events
|
|
948
|
+
|
|
949
|
+
```bash
|
|
950
|
+
kubectl -n "${KUBEAI_NAMESPACE}" get events --sort-by=.lastTimestamp | tail -n 40
|
|
951
|
+
```
|
|
952
|
+
|
|
953
|
+
### Common bad states and what they mean
|
|
954
|
+
|
|
955
|
+
* `invalid resource profile: "gpu-single-default", should match <name>:<multiple>`
|
|
956
|
+
* append `:1` in the live `Model` spec
|
|
957
|
+
* `libcuda.so.1: cannot open shared object file`
|
|
958
|
+
* the pod landed on the GPU node without actually requesting a GPU; fix the resource-profile file to include GPU requests, limits, and `runtimeClassName: nvidia`
|
|
959
|
+
* `/models` works but completions 404 with `The model ... does not exist.`
|
|
960
|
+
* the served model name does not match the public profile name; apply the live args patch above
|
|
961
|
+
* startup probe fails with `connection refused`
|
|
962
|
+
* the model pod may still be pulling the image, loading the model, or warming up
|
|
963
|
+
|
|
964
|
+
---
|
|
965
|
+
|
|
966
|
+
## Which backend should I start with?
|
|
967
|
+
|
|
968
|
+
Start with **Compose** if you want:
|
|
969
|
+
|
|
970
|
+
* the fastest path to a working local server
|
|
971
|
+
* easy inspection of generated files
|
|
972
|
+
* simple single-host iteration
|
|
973
|
+
|
|
974
|
+
Move to **KubeAI** when you want:
|
|
975
|
+
|
|
976
|
+
* vLLM runtimes on Kubernetes
|
|
977
|
+
* KubeAI’s OpenAI-compatible front door
|
|
978
|
+
* profile deployment through Kubernetes artifacts
|
|
979
|
+
|
|
980
|
+
KubeAI rendering is vLLM-only for now. Profiles that enable Ollama, LiteLLM, or Open WebUI are rejected for `--backend kubeai`.
|
|
981
|
+
|
|
982
|
+
A good workflow is:
|
|
983
|
+
|
|
984
|
+
1. inspect a profile with `describe-profile`
|
|
985
|
+
2. run it with Compose when you want the simplest local deployment
|
|
986
|
+
3. move to KubeAI when you want Kubernetes-backed serving
|
|
987
|
+
|
|
988
|
+
Compose is the better fit when you already know which profile you want. KubeAI has more first-request overhead because it may need to create pods, pull images, load the model, and warm up the backend.
|
|
989
|
+
|
|
990
|
+
|
|
991
|
+
|
|
992
|
+
## vLLM startup caches
|
|
993
|
+
|
|
994
|
+
Generated Compose mounts persist Hugging Face, vLLM, PyTorch/TorchInductor,
|
|
995
|
+
Triton, and CUDA JIT caches. Warm starts avoid redownloading and redoing many
|
|
996
|
+
compile/JIT steps, but a vLLM model swap still creates a new engine process and
|
|
997
|
+
must reload weights into GPU memory.
|
|
998
|
+
|
|
999
|
+
### Diagnosing profile switches and readiness
|
|
1000
|
+
|
|
1001
|
+
`docker compose` health only means that a container-level healthcheck passed. It
|
|
1002
|
+
is not the same thing as "the routed model can answer a request through the
|
|
1003
|
+
active access surface." This matters most when switching between two vLLM
|
|
1004
|
+
profiles that reuse the same runtime service name: the old vLLM process exits,
|
|
1005
|
+
Docker starts the replacement process, and LiteLLM may remain up while returning
|
|
1006
|
+
upstream connection errors until vLLM finishes loading the new model.
|
|
1007
|
+
|
|
1008
|
+
Use the dedicated readiness and diagnostics commands after a switch:
|
|
1009
|
+
|
|
1010
|
+
```bash
|
|
1011
|
+
infer-stack switch gpt2-single --apply --yes
|
|
1012
|
+
infer-stack wait-ready --model gpt2
|
|
1013
|
+
infer-stack smoke-test --model gpt2
|
|
1014
|
+
```
|
|
1015
|
+
|
|
1016
|
+
For debugging, use:
|
|
1017
|
+
|
|
1018
|
+
```bash
|
|
1019
|
+
infer-stack diagnose --model gpt2 --generation
|
|
1020
|
+
infer-stack diagnose --logs --tail 80
|
|
1021
|
+
```
|
|
1022
|
+
|
|
1023
|
+
`diagnose` prints the resolved provider/gateway/frontend graph, rendered Compose
|
|
1024
|
+
service state, LiteLLM route probes, direct provider probes, and optional recent
|
|
1025
|
+
logs. It is intended to distinguish an actual LiteLLM outage from the more
|
|
1026
|
+
common case where LiteLLM is running but its upstream vLLM runtime is still
|
|
1027
|
+
booting.
|
|
1028
|
+
|
|
1029
|
+
The Compose service-state diagnostics include Docker's exit code, OOM-killed
|
|
1030
|
+
flag, restart count, and actual container name. This is important because
|
|
1031
|
+
`litellm exited with code 137` usually means Docker sent SIGKILL, commonly from
|
|
1032
|
+
an OOM kill or a forced container replacement, whereas LiteLLM returning HTTP
|
|
1033
|
+
500 with `Cannot connect to host vllm-*` means LiteLLM is still running but the
|
|
1034
|
+
upstream vLLM runtime is not ready yet.
|