literegistry-base-deployment 0.1.2__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- literegistry_base_deployment-0.1.2/LICENSE +21 -0
- literegistry_base_deployment-0.1.2/MANIFEST.in +5 -0
- literegistry_base_deployment-0.1.2/PKG-INFO +428 -0
- literegistry_base_deployment-0.1.2/README.md +409 -0
- literegistry_base_deployment-0.1.2/docker/Dockerfile.local-search +43 -0
- literegistry_base_deployment-0.1.2/docker/Dockerfile.redis +30 -0
- literegistry_base_deployment-0.1.2/docker/Dockerfile.services +27 -0
- literegistry_base_deployment-0.1.2/docker/Dockerfile.terminal +35 -0
- literegistry_base_deployment-0.1.2/docker/Dockerfile.vllm +17 -0
- literegistry_base_deployment-0.1.2/pyproject.toml +38 -0
- literegistry_base_deployment-0.1.2/scripts/build-images.sh +122 -0
- literegistry_base_deployment-0.1.2/setup.cfg +4 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/__init__.py +6 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/__main__.py +5 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/cli.py +139 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/launcher.py +646 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/py.typed +0 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/PKG-INFO +428 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/SOURCES.txt +24 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/dependency_links.txt +1 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/entry_points.txt +2 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/requires.txt +10 -0
- literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/top_level.txt +1 -0
- literegistry_base_deployment-0.1.2/tests/test_docker_images.py +92 -0
- literegistry_base_deployment-0.1.2/tests/test_launcher.py +218 -0
- literegistry_base_deployment-0.1.2/tests/test_runtime.py +23 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Goncalo Faria
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,428 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: literegistry-base-deployment
|
|
3
|
+
Version: 0.1.2
|
|
4
|
+
Summary: Deploy LiteRegistry gateway, Redis, tools, search, and vLLM services on Beaker.
|
|
5
|
+
Author: Goncalo Faria
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Requires-Python: >=3.10
|
|
8
|
+
Description-Content-Type: text/markdown
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Requires-Dist: literegistry>=1.0.44
|
|
11
|
+
Requires-Dist: fire<1,>=0.7
|
|
12
|
+
Provides-Extra: test
|
|
13
|
+
Requires-Dist: httpx; extra == "test"
|
|
14
|
+
Requires-Dist: pytest>=8; extra == "test"
|
|
15
|
+
Provides-Extra: publish
|
|
16
|
+
Requires-Dist: build>=1.2; extra == "publish"
|
|
17
|
+
Requires-Dist: twine>=5; extra == "publish"
|
|
18
|
+
Dynamic: license-file
|
|
19
|
+
|
|
20
|
+
# LiteRegistry Base Deployment
|
|
21
|
+
|
|
22
|
+
`literegistry-base-deployment` is a standalone Beaker deployment package for
|
|
23
|
+
the non-Podman LiteRegistry stack previously launched from Datadev:
|
|
24
|
+
|
|
25
|
+
```text
|
|
26
|
+
clients -> one gateway -> Python execution replicas
|
|
27
|
+
-> restricted terminal replicas
|
|
28
|
+
-> cached Serper query + Jina URL replicas
|
|
29
|
+
-> Lucene BM25 local-search replicas
|
|
30
|
+
-> vLLM generation replicas
|
|
31
|
+
-> vLLM classification replicas
|
|
32
|
+
|
|
|
33
|
+
+-> managed or external Redis registry
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
It depends strongly on `literegistry` and has no Datadev dependency. Shared
|
|
37
|
+
coordination lives in `literegistry.coop`; this package contains no copy of the
|
|
38
|
+
port allocator, Redis barrier, or artifact-build locks. It copies the proven
|
|
39
|
+
mechanics from `literegistry-podman-beaker`: Fire CLI, preview before
|
|
40
|
+
submission, host networking, collision-safe dynamic ports, CPU-cluster replica
|
|
41
|
+
spreading, managed-or-external Redis, TTL-backed Weka endpoint discovery, and direct Beaker specs.
|
|
42
|
+
|
|
43
|
+
## What is deployed
|
|
44
|
+
|
|
45
|
+
| Task | Gateway route | Default replicas | Compute |
|
|
46
|
+
|---|---|---:|---|
|
|
47
|
+
| Gateway | all routes | 1 process group, 8 workers | CPU |
|
|
48
|
+
| Python | `POST /python` | 1 | CPU |
|
|
49
|
+
| Terminal | `POST /terminal` | 1 | CPU |
|
|
50
|
+
| Web search / fetch | `POST /search` | 1 | CPU |
|
|
51
|
+
| Local BM25 search | `POST /search` with `model_path=localsearch:...` | 0 | CPU |
|
|
52
|
+
| vLLM generation | `/v1/chat/completions`, `/v1/completions` | 0 | GPU |
|
|
53
|
+
| vLLM classification | `POST /classify` | 0 | GPU |
|
|
54
|
+
| Redis | internal service discovery | 1 when `--registry` is omitted | CPU |
|
|
55
|
+
|
|
56
|
+
Generation and classification are separate vLLM pools. Generation launches
|
|
57
|
+
vLLM with `--task=generate`; classification launches it with
|
|
58
|
+
`--task=classify`, which exposes vLLM's sequence-classification endpoint. The
|
|
59
|
+
gateway chooses either pool using the request's `model` field.
|
|
60
|
+
|
|
61
|
+
## End-to-end setup
|
|
62
|
+
|
|
63
|
+
These commands assume Bash, this repository checkout, Docker, the Beaker CLI,
|
|
64
|
+
and access to the target workspace, clusters, budget, and Weka source.
|
|
65
|
+
|
|
66
|
+
### 1. Install the launcher
|
|
67
|
+
|
|
68
|
+
```bash
|
|
69
|
+
cd /weka/gfaria/literegistry/literegistry_base_deployment
|
|
70
|
+
python -m pip install -e '.[test,publish]'
|
|
71
|
+
literegistry-base-deployment --help
|
|
72
|
+
beaker config test
|
|
73
|
+
beaker account whoami
|
|
74
|
+
beaker workspace get ai2/oe-agents
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
A published install through the LiteRegistry extra is:
|
|
78
|
+
|
|
79
|
+
```bash
|
|
80
|
+
pip install "literegistry[base_deployment]"
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
The standalone distribution can also be installed directly:
|
|
84
|
+
|
|
85
|
+
```bash
|
|
86
|
+
pip install literegistry-base-deployment
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
### 2. Make LiteRegistry available to Docker
|
|
90
|
+
|
|
91
|
+
Every runtime image installs this package, which installs
|
|
92
|
+
`literegistry>=1.0.44`. Publish that LiteRegistry version to the Python index
|
|
93
|
+
used by Docker, or expose its wheel through an HTTP wheelhouse reachable from
|
|
94
|
+
inside Docker:
|
|
95
|
+
|
|
96
|
+
```bash
|
|
97
|
+
cd /weka/gfaria/literegistry
|
|
98
|
+
python -m build
|
|
99
|
+
python -m twine check dist/*
|
|
100
|
+
# python -m twine upload dist/literegistry-1.0.44*
|
|
101
|
+
|
|
102
|
+
export PIP_INDEX_URL=https://python.example/simple
|
|
103
|
+
# Or: export PIP_FIND_LINKS=https://python.example/wheels
|
|
104
|
+
# Optional for an internal HTTP endpoint:
|
|
105
|
+
# export PIP_TRUSTED_HOST=python.example
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
The four default images built by this package need no `PYTHONPATH`, Datadev
|
|
109
|
+
checkout, `~/basic_images`, or Weka source-code mount. The optional Lucene image
|
|
110
|
+
uses JTC only for its index-building assets; the server is a LiteRegistry service.
|
|
111
|
+
|
|
112
|
+
### 3. Build the four package images
|
|
113
|
+
|
|
114
|
+
```bash
|
|
115
|
+
cd /weka/gfaria/literegistry/literegistry_base_deployment
|
|
116
|
+
export IMAGE_TAG=0.1.0
|
|
117
|
+
./scripts/build-images.sh "" "$IMAGE_TAG"
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
The script builds:
|
|
121
|
+
|
|
122
|
+
```text
|
|
123
|
+
literegistry-redis:0.1.0
|
|
124
|
+
literegistry-base-services:0.1.0
|
|
125
|
+
literegistry-base-terminal:0.1.0
|
|
126
|
+
literegistry-base-vllm:0.1.0
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
The services image runs gateway, Python, or web search depending on its Beaker
|
|
130
|
+
command. Terminal has every binary allowed by the restricted pipeline server.
|
|
131
|
+
Local search is not rebuilt by default: the launcher targets the Beaker image
|
|
132
|
+
`goncalof/jtc-local-search-lucene-bm25`, which must contain LiteRegistry 1.0.44
|
|
133
|
+
or newer. vLLM uses
|
|
134
|
+
`vllm/vllm-openai:latest` by default; pin or replace it when reproducibility
|
|
135
|
+
requires a specific vLLM/CUDA combination:
|
|
136
|
+
|
|
137
|
+
```bash
|
|
138
|
+
VLLM_BASE_IMAGE=vllm/vllm-openai:0.11.0 \
|
|
139
|
+
./scripts/build-images.sh "" "$IMAGE_TAG"
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
The canonical Dockerfile for that JTC image is included at
|
|
143
|
+
`docker/Dockerfile.local-search`. It copies JTC's existing `search/` index-building assets. The running
|
|
144
|
+
server is
|
|
145
|
+
`literegistry.services.bm25_server`; neither this deployment package nor the
|
|
146
|
+
image copies `datadev.infra.bm25_server`. To reproduce the image from a JTC checkout:
|
|
147
|
+
|
|
148
|
+
```bash
|
|
149
|
+
cd /weka/gfaria/literegistry/literegistry_base_deployment
|
|
150
|
+
BUILD_LOCAL_SEARCH=1 \
|
|
151
|
+
JTC_BUILD_CONTEXT=/weka/gfaria/jtc \
|
|
152
|
+
./scripts/build-images.sh "" "$IMAGE_TAG"
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
The normal build leaves `BUILD_LOCAL_SEARCH=0` and expects the updated Beaker
|
|
156
|
+
image to have already been uploaded.
|
|
157
|
+
|
|
158
|
+
To build and push to an ordinary Docker registry:
|
|
159
|
+
|
|
160
|
+
```bash
|
|
161
|
+
PUSH_IMAGES=1 \
|
|
162
|
+
./scripts/build-images.sh registry.example/team "$IMAGE_TAG"
|
|
163
|
+
```
|
|
164
|
+
|
|
165
|
+
### 4. Upload the images to Beaker
|
|
166
|
+
|
|
167
|
+
The launcher image flags take Beaker image names or IDs, not local Docker tags.
|
|
168
|
+
|
|
169
|
+
```bash
|
|
170
|
+
export WORKSPACE=ai2/oe-agents
|
|
171
|
+
export BEAKER_TAG="${IMAGE_TAG//./-}"
|
|
172
|
+
|
|
173
|
+
beaker image create \
|
|
174
|
+
"$(docker image inspect --format '{{.Id}}' "literegistry-redis:$IMAGE_TAG")" \
|
|
175
|
+
--name "literegistry-redis-$BEAKER_TAG" --workspace "$WORKSPACE"
|
|
176
|
+
|
|
177
|
+
beaker image create \
|
|
178
|
+
"$(docker image inspect --format '{{.Id}}' "literegistry-base-services:$IMAGE_TAG")" \
|
|
179
|
+
--name "literegistry-base-services-$BEAKER_TAG" --workspace "$WORKSPACE"
|
|
180
|
+
|
|
181
|
+
beaker image create \
|
|
182
|
+
"$(docker image inspect --format '{{.Id}}' "literegistry-base-terminal:$IMAGE_TAG")" \
|
|
183
|
+
--name "literegistry-base-terminal-$BEAKER_TAG" --workspace "$WORKSPACE"
|
|
184
|
+
|
|
185
|
+
beaker image create \
|
|
186
|
+
"$(docker image inspect --format '{{.Id}}' "literegistry-base-vllm:$IMAGE_TAG")" \
|
|
187
|
+
--name "literegistry-base-vllm-$BEAKER_TAG" --workspace "$WORKSPACE"
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
### 5. Create API-key secrets
|
|
191
|
+
|
|
192
|
+
Web query search uses Serper and URL fetching uses Jina Reader. The launcher
|
|
193
|
+
accepts only Beaker secret names; it never places raw keys in the experiment
|
|
194
|
+
spec or command line.
|
|
195
|
+
|
|
196
|
+
```bash
|
|
197
|
+
beaker secret write SERPER_API_KEY
|
|
198
|
+
beaker secret write JINA_API_KEY
|
|
199
|
+
beaker secret write HF_TOKEN
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
Skip the first two when launching with `--web-search-replicas=0`. Skip the
|
|
203
|
+
Hugging Face secret by passing `--hf-token-secret=None` when all model artifacts
|
|
204
|
+
are public or already cached.
|
|
205
|
+
|
|
206
|
+
### 6. Preview a complete stack
|
|
207
|
+
|
|
208
|
+
This example runs CPU services on Jupiter and GPU model pools on a selected GPU
|
|
209
|
+
cluster. Replace the model names, corpus paths, and model cluster with real
|
|
210
|
+
values. The corpus and index are Weka paths visible inside every replica.
|
|
211
|
+
|
|
212
|
+
```bash
|
|
213
|
+
literegistry-base-deployment preview \
|
|
214
|
+
--python-replicas=2 \
|
|
215
|
+
--terminal-replicas=2 \
|
|
216
|
+
--web-search-replicas=2 \
|
|
217
|
+
--local-search-replicas=2 \
|
|
218
|
+
--local-search-corpus-jsonl=/weka/gfaria/search/corpus.jsonl \
|
|
219
|
+
--local-search-index-dir=/weka/gfaria/search/lucene-index \
|
|
220
|
+
--generation-model=allenai/example-generation-model \
|
|
221
|
+
--generation-replicas=2 \
|
|
222
|
+
--generation-tp=1 \
|
|
223
|
+
--classification-model=allenai/example-reward-model \
|
|
224
|
+
--classification-replicas=1 \
|
|
225
|
+
--classification-tp=1 \
|
|
226
|
+
--gateway-workers=8 \
|
|
227
|
+
--service-cluster=ai2/jupiter \
|
|
228
|
+
--gateway-cluster=ai2/jupiter \
|
|
229
|
+
--model-cluster=ai2/jupiter \
|
|
230
|
+
--workspace="$WORKSPACE" \
|
|
231
|
+
--budget=ai2/oe-omai \
|
|
232
|
+
--redis-image="literegistry-redis-$BEAKER_TAG" \
|
|
233
|
+
--services-image="literegistry-base-services-$BEAKER_TAG" \
|
|
234
|
+
--terminal-image="literegistry-base-terminal-$BEAKER_TAG" \
|
|
235
|
+
--local-search-image=goncalof/jtc-local-search-lucene-bm25 \
|
|
236
|
+
--vllm-image="literegistry-base-vllm-$BEAKER_TAG"
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
Preview validates and prints the exact Beaker spec without creating anything.
|
|
240
|
+
With no `--registry`, it includes one managed Redis task. To reuse Redis, add:
|
|
241
|
+
|
|
242
|
+
```bash
|
|
243
|
+
literegistry-base-deployment preview \
|
|
244
|
+
--registry=redis://jupiter-cs-aus-183.reviz.ai2.in:59936 \
|
|
245
|
+
--python-replicas=1 \
|
|
246
|
+
--terminal-replicas=1 \
|
|
247
|
+
--web-search-replicas=0
|
|
248
|
+
```
|
|
249
|
+
|
|
250
|
+
### 7. Launch and find the gateway
|
|
251
|
+
|
|
252
|
+
Change `preview` to `launch`. This smaller CPU-only example is useful for a
|
|
253
|
+
first deployment test:
|
|
254
|
+
|
|
255
|
+
```bash
|
|
256
|
+
literegistry-base-deployment launch \
|
|
257
|
+
--python-replicas=1 \
|
|
258
|
+
--terminal-replicas=1 \
|
|
259
|
+
--web-search-replicas=1 \
|
|
260
|
+
--service-cluster=ai2/jupiter \
|
|
261
|
+
--gateway-cluster=ai2/jupiter \
|
|
262
|
+
--workspace="$WORKSPACE" \
|
|
263
|
+
--budget=ai2/oe-omai \
|
|
264
|
+
--redis-image="literegistry-redis-$BEAKER_TAG" \
|
|
265
|
+
--services-image="literegistry-base-services-$BEAKER_TAG" \
|
|
266
|
+
--terminal-image="literegistry-base-terminal-$BEAKER_TAG" \
|
|
267
|
+
| tee /tmp/literegistry-base-launch.json
|
|
268
|
+
|
|
269
|
+
export EXPERIMENT_ID="$(jq -r '.beaker.id' /tmp/literegistry-base-launch.json)"
|
|
270
|
+
export EXPERIMENT_NAME="$(jq -r '.experiment_name' /tmp/literegistry-base-launch.json)"
|
|
271
|
+
export COOP_ROOT="/weka/gfaria/.literegistry-coop/${EXPERIMENT_NAME}"
|
|
272
|
+
export GATEWAY_URL="$(python -m literegistry.coop.endpoints wait \
|
|
273
|
+
--root="$COOP_ROOT" \
|
|
274
|
+
--name=gateway \
|
|
275
|
+
--healthcheck=http \
|
|
276
|
+
--timeout=600)"
|
|
277
|
+
echo "$GATEWAY_URL"
|
|
278
|
+
|
|
279
|
+
beaker experiment get "$EXPERIMENT_ID"
|
|
280
|
+
curl -fsS "$GATEWAY_URL/health"
|
|
281
|
+
curl -fsS "$GATEWAY_URL/v1/models"
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
The gateway also prints `LITEREGISTRY_ENDPOINT_GATEWAY=...` in its Beaker
|
|
285
|
+
logs. Its endpoint record is refreshed while the gateway is healthy and removed
|
|
286
|
+
on clean shutdown; after a crash, its short TTL expires automatically.
|
|
287
|
+
|
|
288
|
+
### 8. Exercise every gateway route
|
|
289
|
+
|
|
290
|
+
```bash
|
|
291
|
+
curl -fsS -X POST "$GATEWAY_URL/python" \
|
|
292
|
+
-H 'content-type: application/json' \
|
|
293
|
+
-d '{"code":"print(2 + 2)","max_runtime":1}'
|
|
294
|
+
|
|
295
|
+
curl -fsS -X POST "$GATEWAY_URL/terminal" \
|
|
296
|
+
-H 'content-type: application/json' \
|
|
297
|
+
-d '{"contents":"INFO ok\nERROR ai2 hello\n","command":"rg ERROR","max_runtime":5}'
|
|
298
|
+
|
|
299
|
+
curl -fsS -X POST "$GATEWAY_URL/search" \
|
|
300
|
+
-H 'content-type: application/json' \
|
|
301
|
+
-d '{"mode":"query","query":"Allen Institute for AI","num_results":3}'
|
|
302
|
+
|
|
303
|
+
curl -fsS -X POST "$GATEWAY_URL/search" \
|
|
304
|
+
-H 'content-type: application/json' \
|
|
305
|
+
-d '{"mode":"url","url":"https://allenai.org/"}'
|
|
306
|
+
```
|
|
307
|
+
|
|
308
|
+
When local search is enabled, select its registered pool explicitly:
|
|
309
|
+
|
|
310
|
+
```bash
|
|
311
|
+
curl -fsS -X POST "$GATEWAY_URL/search" \
|
|
312
|
+
-H 'content-type: application/json' \
|
|
313
|
+
-d '{"model_path":"localsearch:corpus","mode":"query","query":"ai2 hello","num_results":3}'
|
|
314
|
+
```
|
|
315
|
+
|
|
316
|
+
For a generation model:
|
|
317
|
+
|
|
318
|
+
```bash
|
|
319
|
+
curl -fsS -X POST "$GATEWAY_URL/v1/chat/completions" \
|
|
320
|
+
-H 'content-type: application/json' \
|
|
321
|
+
-d '{"model":"allenai/example-generation-model","messages":[{"role":"user","content":"Say ai2 hello"}],"max_tokens":32}'
|
|
322
|
+
```
|
|
323
|
+
|
|
324
|
+
For a vLLM sequence classifier, the final assistant response is already part
|
|
325
|
+
of the conversation, so `add_generation_prompt` stays false:
|
|
326
|
+
|
|
327
|
+
```bash
|
|
328
|
+
curl -fsS -X POST "$GATEWAY_URL/classify" \
|
|
329
|
+
-H 'content-type: application/json' \
|
|
330
|
+
-d '{"model":"allenai/example-reward-model","messages":[{"role":"user","content":"Say hello"},{"role":"assistant","content":"ai2 hello"}],"add_generation_prompt":false}'
|
|
331
|
+
```
|
|
332
|
+
|
|
333
|
+
### 9. Stop the deployment
|
|
334
|
+
|
|
335
|
+
```bash
|
|
336
|
+
literegistry-base-deployment stop "$EXPERIMENT_ID"
|
|
337
|
+
```
|
|
338
|
+
|
|
339
|
+
## Resumption and failure policy
|
|
340
|
+
|
|
341
|
+
| Task | `context.autoResume` | `propagateFailure` | `propagatePreemption` |
|
|
342
|
+
|---|---:|---:|---:|
|
|
343
|
+
| Managed Redis | `false` | `true` | `true` |
|
|
344
|
+
| Gateway and every worker pool | `true` | `false` | `false` |
|
|
345
|
+
|
|
346
|
+
Every non-Redis task is resumable and isolated from experiment-wide failure.
|
|
347
|
+
Managed Redis is intentionally the only non-resumable task and the only task
|
|
348
|
+
whose failure or preemption terminates the complete experiment. With external
|
|
349
|
+
`--registry`, no Redis task is created, so every task in this experiment is
|
|
350
|
+
resumable.
|
|
351
|
+
|
|
352
|
+
## Local-search behavior
|
|
353
|
+
|
|
354
|
+
Local search is implemented by LiteRegistry's first-class BM25 service. The
|
|
355
|
+
deployment invokes its `/app/search/build_lucene_index.sh` when no
|
|
356
|
+
`segments_*` file exists,
|
|
357
|
+
then starts its registered application with
|
|
358
|
+
`literegistry.services.bm25_server:create_app`. The Dockerfile uses the JTC
|
|
359
|
+
checkout only for the existing Lucene index
|
|
360
|
+
builder; there is no Datadev server dependency.
|
|
361
|
+
|
|
362
|
+
The default service name is `localsearch:<corpus filename stem>`. Override it
|
|
363
|
+
with `--local-search-service-name` and pass exactly that value as
|
|
364
|
+
`model_path` in gateway `/search` requests.
|
|
365
|
+
|
|
366
|
+
## Horizontal CPU placement
|
|
367
|
+
|
|
368
|
+
Comma-separated `--service-cluster` values produce separate Beaker task groups,
|
|
369
|
+
with Python, terminal, web-search, and local-search replicas divided as evenly
|
|
370
|
+
as possible. This forces placement across the named clusters instead of merely
|
|
371
|
+
asking Beaker for many replicas in one cluster:
|
|
372
|
+
|
|
373
|
+
```bash
|
|
374
|
+
literegistry-base-deployment preview \
|
|
375
|
+
--registry=redis://registry.example:6379 \
|
|
376
|
+
--python-replicas=16 \
|
|
377
|
+
--terminal-replicas=16 \
|
|
378
|
+
--web-search-replicas=0 \
|
|
379
|
+
--service-cluster=ai2/neptune,ai2/saturn,ai2/jupiter,ai2/ceres
|
|
380
|
+
```
|
|
381
|
+
|
|
382
|
+
GPU pools currently use the single explicit `--model-cluster`; tensor
|
|
383
|
+
parallelism maps directly to Beaker `gpuCount` for each model replica.
|
|
384
|
+
|
|
385
|
+
## Python API
|
|
386
|
+
|
|
387
|
+
```python
|
|
388
|
+
from literegistry_base_deployment import BaseDeploymentConfig, BaseDeploymentLauncher
|
|
389
|
+
|
|
390
|
+
|
|
391
|
+
config = BaseDeploymentConfig(
|
|
392
|
+
registry="redis://jupiter-cs-aus-183.reviz.ai2.in:59936",
|
|
393
|
+
python_replicas=4,
|
|
394
|
+
terminal_replicas=4,
|
|
395
|
+
web_search_replicas=2,
|
|
396
|
+
service_clusters=("ai2/jupiter", "ai2/ceres"),
|
|
397
|
+
gateway_cluster="ai2/jupiter",
|
|
398
|
+
)
|
|
399
|
+
launcher = BaseDeploymentLauncher(config)
|
|
400
|
+
print(launcher.preview()) # read-only
|
|
401
|
+
receipt = launcher.submit() # creates the Beaker experiment
|
|
402
|
+
print(receipt)
|
|
403
|
+
|
|
404
|
+
# Later:
|
|
405
|
+
BaseDeploymentLauncher.stop(receipt["beaker"]["id"])
|
|
406
|
+
```
|
|
407
|
+
|
|
408
|
+
## Development validation
|
|
409
|
+
|
|
410
|
+
```bash
|
|
411
|
+
cd /weka/gfaria/literegistry/literegistry_base_deployment
|
|
412
|
+
python -m pip install -e '.[test,publish]'
|
|
413
|
+
python -m pytest
|
|
414
|
+
python -m build
|
|
415
|
+
python -m twine check dist/*
|
|
416
|
+
```
|
|
417
|
+
|
|
418
|
+
The tests verify service composition, shell syntax, managed Redis discovery,
|
|
419
|
+
secret references, multi-cluster spreading, generation/classification vLLM
|
|
420
|
+
arguments, JTC Lucene image routing, package contents, and Dockerfile
|
|
421
|
+
self-containment. Dockerfiles are structurally tested by default; actually
|
|
422
|
+
building the four package images requires a Docker daemon and network access.
|
|
423
|
+
|
|
424
|
+
|
|
425
|
+
## Experimental mirror soft affinity
|
|
426
|
+
|
|
427
|
+
The bundled gateway enables experimental mirror soft affinity by default. If this base deployment is used with external `docker-mirror` services, disable it with
|
|
428
|
+
`--docker-mirror-soft-affinity=False` to restore normal load balancing.
|