literegistry-base-deployment 0.1.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (26) hide show
  1. literegistry_base_deployment-0.1.2/LICENSE +21 -0
  2. literegistry_base_deployment-0.1.2/MANIFEST.in +5 -0
  3. literegistry_base_deployment-0.1.2/PKG-INFO +428 -0
  4. literegistry_base_deployment-0.1.2/README.md +409 -0
  5. literegistry_base_deployment-0.1.2/docker/Dockerfile.local-search +43 -0
  6. literegistry_base_deployment-0.1.2/docker/Dockerfile.redis +30 -0
  7. literegistry_base_deployment-0.1.2/docker/Dockerfile.services +27 -0
  8. literegistry_base_deployment-0.1.2/docker/Dockerfile.terminal +35 -0
  9. literegistry_base_deployment-0.1.2/docker/Dockerfile.vllm +17 -0
  10. literegistry_base_deployment-0.1.2/pyproject.toml +38 -0
  11. literegistry_base_deployment-0.1.2/scripts/build-images.sh +122 -0
  12. literegistry_base_deployment-0.1.2/setup.cfg +4 -0
  13. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/__init__.py +6 -0
  14. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/__main__.py +5 -0
  15. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/cli.py +139 -0
  16. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/launcher.py +646 -0
  17. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment/py.typed +0 -0
  18. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/PKG-INFO +428 -0
  19. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/SOURCES.txt +24 -0
  20. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/dependency_links.txt +1 -0
  21. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/entry_points.txt +2 -0
  22. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/requires.txt +10 -0
  23. literegistry_base_deployment-0.1.2/src/literegistry_base_deployment.egg-info/top_level.txt +1 -0
  24. literegistry_base_deployment-0.1.2/tests/test_docker_images.py +92 -0
  25. literegistry_base_deployment-0.1.2/tests/test_launcher.py +218 -0
  26. literegistry_base_deployment-0.1.2/tests/test_runtime.py +23 -0
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Goncalo Faria
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,5 @@
1
+ include LICENSE
2
+ include README.md
3
+ recursive-include src *.typed
4
+ recursive-include docker *
5
+ recursive-include scripts *.sh
@@ -0,0 +1,428 @@
1
+ Metadata-Version: 2.4
2
+ Name: literegistry-base-deployment
3
+ Version: 0.1.2
4
+ Summary: Deploy LiteRegistry gateway, Redis, tools, search, and vLLM services on Beaker.
5
+ Author: Goncalo Faria
6
+ License-Expression: MIT
7
+ Requires-Python: >=3.10
8
+ Description-Content-Type: text/markdown
9
+ License-File: LICENSE
10
+ Requires-Dist: literegistry>=1.0.44
11
+ Requires-Dist: fire<1,>=0.7
12
+ Provides-Extra: test
13
+ Requires-Dist: httpx; extra == "test"
14
+ Requires-Dist: pytest>=8; extra == "test"
15
+ Provides-Extra: publish
16
+ Requires-Dist: build>=1.2; extra == "publish"
17
+ Requires-Dist: twine>=5; extra == "publish"
18
+ Dynamic: license-file
19
+
20
+ # LiteRegistry Base Deployment
21
+
22
+ `literegistry-base-deployment` is a standalone Beaker deployment package for
23
+ the non-Podman LiteRegistry stack previously launched from Datadev:
24
+
25
+ ```text
26
+ clients -> one gateway -> Python execution replicas
27
+ -> restricted terminal replicas
28
+ -> cached Serper query + Jina URL replicas
29
+ -> Lucene BM25 local-search replicas
30
+ -> vLLM generation replicas
31
+ -> vLLM classification replicas
32
+ |
33
+ +-> managed or external Redis registry
34
+ ```
35
+
36
+ It depends strongly on `literegistry` and has no Datadev dependency. Shared
37
+ coordination lives in `literegistry.coop`; this package contains no copy of the
38
+ port allocator, Redis barrier, or artifact-build locks. It copies the proven
39
+ mechanics from `literegistry-podman-beaker`: Fire CLI, preview before
40
+ submission, host networking, collision-safe dynamic ports, CPU-cluster replica
41
+ spreading, managed-or-external Redis, TTL-backed Weka endpoint discovery, and direct Beaker specs.
42
+
43
+ ## What is deployed
44
+
45
+ | Task | Gateway route | Default replicas | Compute |
46
+ |---|---|---:|---|
47
+ | Gateway | all routes | 1 process group, 8 workers | CPU |
48
+ | Python | `POST /python` | 1 | CPU |
49
+ | Terminal | `POST /terminal` | 1 | CPU |
50
+ | Web search / fetch | `POST /search` | 1 | CPU |
51
+ | Local BM25 search | `POST /search` with `model_path=localsearch:...` | 0 | CPU |
52
+ | vLLM generation | `/v1/chat/completions`, `/v1/completions` | 0 | GPU |
53
+ | vLLM classification | `POST /classify` | 0 | GPU |
54
+ | Redis | internal service discovery | 1 when `--registry` is omitted | CPU |
55
+
56
+ Generation and classification are separate vLLM pools. Generation launches
57
+ vLLM with `--task=generate`; classification launches it with
58
+ `--task=classify`, which exposes vLLM's sequence-classification endpoint. The
59
+ gateway chooses either pool using the request's `model` field.
60
+
61
+ ## End-to-end setup
62
+
63
+ These commands assume Bash, this repository checkout, Docker, the Beaker CLI,
64
+ and access to the target workspace, clusters, budget, and Weka source.
65
+
66
+ ### 1. Install the launcher
67
+
68
+ ```bash
69
+ cd /weka/gfaria/literegistry/literegistry_base_deployment
70
+ python -m pip install -e '.[test,publish]'
71
+ literegistry-base-deployment --help
72
+ beaker config test
73
+ beaker account whoami
74
+ beaker workspace get ai2/oe-agents
75
+ ```
76
+
77
+ A published install through the LiteRegistry extra is:
78
+
79
+ ```bash
80
+ pip install "literegistry[base_deployment]"
81
+ ```
82
+
83
+ The standalone distribution can also be installed directly:
84
+
85
+ ```bash
86
+ pip install literegistry-base-deployment
87
+ ```
88
+
89
+ ### 2. Make LiteRegistry available to Docker
90
+
91
+ Every runtime image installs this package, which installs
92
+ `literegistry>=1.0.44`. Publish that LiteRegistry version to the Python index
93
+ used by Docker, or expose its wheel through an HTTP wheelhouse reachable from
94
+ inside Docker:
95
+
96
+ ```bash
97
+ cd /weka/gfaria/literegistry
98
+ python -m build
99
+ python -m twine check dist/*
100
+ # python -m twine upload dist/literegistry-1.0.44*
101
+
102
+ export PIP_INDEX_URL=https://python.example/simple
103
+ # Or: export PIP_FIND_LINKS=https://python.example/wheels
104
+ # Optional for an internal HTTP endpoint:
105
+ # export PIP_TRUSTED_HOST=python.example
106
+ ```
107
+
108
+ The four default images built by this package need no `PYTHONPATH`, Datadev
109
+ checkout, `~/basic_images`, or Weka source-code mount. The optional Lucene image
110
+ uses JTC only for its index-building assets; the server is a LiteRegistry service.
111
+
112
+ ### 3. Build the four package images
113
+
114
+ ```bash
115
+ cd /weka/gfaria/literegistry/literegistry_base_deployment
116
+ export IMAGE_TAG=0.1.0
117
+ ./scripts/build-images.sh "" "$IMAGE_TAG"
118
+ ```
119
+
120
+ The script builds:
121
+
122
+ ```text
123
+ literegistry-redis:0.1.0
124
+ literegistry-base-services:0.1.0
125
+ literegistry-base-terminal:0.1.0
126
+ literegistry-base-vllm:0.1.0
127
+ ```
128
+
129
+ The services image runs gateway, Python, or web search depending on its Beaker
130
+ command. Terminal has every binary allowed by the restricted pipeline server.
131
+ Local search is not rebuilt by default: the launcher targets the Beaker image
132
+ `goncalof/jtc-local-search-lucene-bm25`, which must contain LiteRegistry 1.0.44
133
+ or newer. vLLM uses
134
+ `vllm/vllm-openai:latest` by default; pin or replace it when reproducibility
135
+ requires a specific vLLM/CUDA combination:
136
+
137
+ ```bash
138
+ VLLM_BASE_IMAGE=vllm/vllm-openai:0.11.0 \
139
+ ./scripts/build-images.sh "" "$IMAGE_TAG"
140
+ ```
141
+
142
+ The canonical Dockerfile for that JTC image is included at
143
+ `docker/Dockerfile.local-search`. It copies JTC's existing `search/` index-building assets. The running
144
+ server is
145
+ `literegistry.services.bm25_server`; neither this deployment package nor the
146
+ image copies `datadev.infra.bm25_server`. To reproduce the image from a JTC checkout:
147
+
148
+ ```bash
149
+ cd /weka/gfaria/literegistry/literegistry_base_deployment
150
+ BUILD_LOCAL_SEARCH=1 \
151
+ JTC_BUILD_CONTEXT=/weka/gfaria/jtc \
152
+ ./scripts/build-images.sh "" "$IMAGE_TAG"
153
+ ```
154
+
155
+ The normal build leaves `BUILD_LOCAL_SEARCH=0` and expects the updated Beaker
156
+ image to have already been uploaded.
157
+
158
+ To build and push to an ordinary Docker registry:
159
+
160
+ ```bash
161
+ PUSH_IMAGES=1 \
162
+ ./scripts/build-images.sh registry.example/team "$IMAGE_TAG"
163
+ ```
164
+
165
+ ### 4. Upload the images to Beaker
166
+
167
+ The launcher image flags take Beaker image names or IDs, not local Docker tags.
168
+
169
+ ```bash
170
+ export WORKSPACE=ai2/oe-agents
171
+ export BEAKER_TAG="${IMAGE_TAG//./-}"
172
+
173
+ beaker image create \
174
+ "$(docker image inspect --format '{{.Id}}' "literegistry-redis:$IMAGE_TAG")" \
175
+ --name "literegistry-redis-$BEAKER_TAG" --workspace "$WORKSPACE"
176
+
177
+ beaker image create \
178
+ "$(docker image inspect --format '{{.Id}}' "literegistry-base-services:$IMAGE_TAG")" \
179
+ --name "literegistry-base-services-$BEAKER_TAG" --workspace "$WORKSPACE"
180
+
181
+ beaker image create \
182
+ "$(docker image inspect --format '{{.Id}}' "literegistry-base-terminal:$IMAGE_TAG")" \
183
+ --name "literegistry-base-terminal-$BEAKER_TAG" --workspace "$WORKSPACE"
184
+
185
+ beaker image create \
186
+ "$(docker image inspect --format '{{.Id}}' "literegistry-base-vllm:$IMAGE_TAG")" \
187
+ --name "literegistry-base-vllm-$BEAKER_TAG" --workspace "$WORKSPACE"
188
+ ```
189
+
190
+ ### 5. Create API-key secrets
191
+
192
+ Web query search uses Serper and URL fetching uses Jina Reader. The launcher
193
+ accepts only Beaker secret names; it never places raw keys in the experiment
194
+ spec or command line.
195
+
196
+ ```bash
197
+ beaker secret write SERPER_API_KEY
198
+ beaker secret write JINA_API_KEY
199
+ beaker secret write HF_TOKEN
200
+ ```
201
+
202
+ Skip the first two when launching with `--web-search-replicas=0`. Skip the
203
+ Hugging Face secret by passing `--hf-token-secret=None` when all model artifacts
204
+ are public or already cached.
205
+
206
+ ### 6. Preview a complete stack
207
+
208
+ This example runs CPU services on Jupiter and GPU model pools on a selected GPU
209
+ cluster. Replace the model names, corpus paths, and model cluster with real
210
+ values. The corpus and index are Weka paths visible inside every replica.
211
+
212
+ ```bash
213
+ literegistry-base-deployment preview \
214
+ --python-replicas=2 \
215
+ --terminal-replicas=2 \
216
+ --web-search-replicas=2 \
217
+ --local-search-replicas=2 \
218
+ --local-search-corpus-jsonl=/weka/gfaria/search/corpus.jsonl \
219
+ --local-search-index-dir=/weka/gfaria/search/lucene-index \
220
+ --generation-model=allenai/example-generation-model \
221
+ --generation-replicas=2 \
222
+ --generation-tp=1 \
223
+ --classification-model=allenai/example-reward-model \
224
+ --classification-replicas=1 \
225
+ --classification-tp=1 \
226
+ --gateway-workers=8 \
227
+ --service-cluster=ai2/jupiter \
228
+ --gateway-cluster=ai2/jupiter \
229
+ --model-cluster=ai2/jupiter \
230
+ --workspace="$WORKSPACE" \
231
+ --budget=ai2/oe-omai \
232
+ --redis-image="literegistry-redis-$BEAKER_TAG" \
233
+ --services-image="literegistry-base-services-$BEAKER_TAG" \
234
+ --terminal-image="literegistry-base-terminal-$BEAKER_TAG" \
235
+ --local-search-image=goncalof/jtc-local-search-lucene-bm25 \
236
+ --vllm-image="literegistry-base-vllm-$BEAKER_TAG"
237
+ ```
238
+
239
+ Preview validates and prints the exact Beaker spec without creating anything.
240
+ With no `--registry`, it includes one managed Redis task. To reuse Redis, add:
241
+
242
+ ```bash
243
+ literegistry-base-deployment preview \
244
+ --registry=redis://jupiter-cs-aus-183.reviz.ai2.in:59936 \
245
+ --python-replicas=1 \
246
+ --terminal-replicas=1 \
247
+ --web-search-replicas=0
248
+ ```
249
+
250
+ ### 7. Launch and find the gateway
251
+
252
+ Change `preview` to `launch`. This smaller CPU-only example is useful for a
253
+ first deployment test:
254
+
255
+ ```bash
256
+ literegistry-base-deployment launch \
257
+ --python-replicas=1 \
258
+ --terminal-replicas=1 \
259
+ --web-search-replicas=1 \
260
+ --service-cluster=ai2/jupiter \
261
+ --gateway-cluster=ai2/jupiter \
262
+ --workspace="$WORKSPACE" \
263
+ --budget=ai2/oe-omai \
264
+ --redis-image="literegistry-redis-$BEAKER_TAG" \
265
+ --services-image="literegistry-base-services-$BEAKER_TAG" \
266
+ --terminal-image="literegistry-base-terminal-$BEAKER_TAG" \
267
+ | tee /tmp/literegistry-base-launch.json
268
+
269
+ export EXPERIMENT_ID="$(jq -r '.beaker.id' /tmp/literegistry-base-launch.json)"
270
+ export EXPERIMENT_NAME="$(jq -r '.experiment_name' /tmp/literegistry-base-launch.json)"
271
+ export COOP_ROOT="/weka/gfaria/.literegistry-coop/${EXPERIMENT_NAME}"
272
+ export GATEWAY_URL="$(python -m literegistry.coop.endpoints wait \
273
+ --root="$COOP_ROOT" \
274
+ --name=gateway \
275
+ --healthcheck=http \
276
+ --timeout=600)"
277
+ echo "$GATEWAY_URL"
278
+
279
+ beaker experiment get "$EXPERIMENT_ID"
280
+ curl -fsS "$GATEWAY_URL/health"
281
+ curl -fsS "$GATEWAY_URL/v1/models"
282
+ ```
283
+
284
+ The gateway also prints `LITEREGISTRY_ENDPOINT_GATEWAY=...` in its Beaker
285
+ logs. Its endpoint record is refreshed while the gateway is healthy and removed
286
+ on clean shutdown; after a crash, its short TTL expires automatically.
287
+
288
+ ### 8. Exercise every gateway route
289
+
290
+ ```bash
291
+ curl -fsS -X POST "$GATEWAY_URL/python" \
292
+ -H 'content-type: application/json' \
293
+ -d '{"code":"print(2 + 2)","max_runtime":1}'
294
+
295
+ curl -fsS -X POST "$GATEWAY_URL/terminal" \
296
+ -H 'content-type: application/json' \
297
+ -d '{"contents":"INFO ok\nERROR ai2 hello\n","command":"rg ERROR","max_runtime":5}'
298
+
299
+ curl -fsS -X POST "$GATEWAY_URL/search" \
300
+ -H 'content-type: application/json' \
301
+ -d '{"mode":"query","query":"Allen Institute for AI","num_results":3}'
302
+
303
+ curl -fsS -X POST "$GATEWAY_URL/search" \
304
+ -H 'content-type: application/json' \
305
+ -d '{"mode":"url","url":"https://allenai.org/"}'
306
+ ```
307
+
308
+ When local search is enabled, select its registered pool explicitly:
309
+
310
+ ```bash
311
+ curl -fsS -X POST "$GATEWAY_URL/search" \
312
+ -H 'content-type: application/json' \
313
+ -d '{"model_path":"localsearch:corpus","mode":"query","query":"ai2 hello","num_results":3}'
314
+ ```
315
+
316
+ For a generation model:
317
+
318
+ ```bash
319
+ curl -fsS -X POST "$GATEWAY_URL/v1/chat/completions" \
320
+ -H 'content-type: application/json' \
321
+ -d '{"model":"allenai/example-generation-model","messages":[{"role":"user","content":"Say ai2 hello"}],"max_tokens":32}'
322
+ ```
323
+
324
+ For a vLLM sequence classifier, the final assistant response is already part
325
+ of the conversation, so `add_generation_prompt` stays false:
326
+
327
+ ```bash
328
+ curl -fsS -X POST "$GATEWAY_URL/classify" \
329
+ -H 'content-type: application/json' \
330
+ -d '{"model":"allenai/example-reward-model","messages":[{"role":"user","content":"Say hello"},{"role":"assistant","content":"ai2 hello"}],"add_generation_prompt":false}'
331
+ ```
332
+
333
+ ### 9. Stop the deployment
334
+
335
+ ```bash
336
+ literegistry-base-deployment stop "$EXPERIMENT_ID"
337
+ ```
338
+
339
+ ## Resumption and failure policy
340
+
341
+ | Task | `context.autoResume` | `propagateFailure` | `propagatePreemption` |
342
+ |---|---:|---:|---:|
343
+ | Managed Redis | `false` | `true` | `true` |
344
+ | Gateway and every worker pool | `true` | `false` | `false` |
345
+
346
+ Every non-Redis task is resumable and isolated from experiment-wide failure.
347
+ Managed Redis is intentionally the only non-resumable task and the only task
348
+ whose failure or preemption terminates the complete experiment. With external
349
+ `--registry`, no Redis task is created, so every task in this experiment is
350
+ resumable.
351
+
352
+ ## Local-search behavior
353
+
354
+ Local search is implemented by LiteRegistry's first-class BM25 service. The
355
+ deployment invokes its `/app/search/build_lucene_index.sh` when no
356
+ `segments_*` file exists,
357
+ then starts its registered application with
358
+ `literegistry.services.bm25_server:create_app`. The Dockerfile uses the JTC
359
+ checkout only for the existing Lucene index
360
+ builder; there is no Datadev server dependency.
361
+
362
+ The default service name is `localsearch:<corpus filename stem>`. Override it
363
+ with `--local-search-service-name` and pass exactly that value as
364
+ `model_path` in gateway `/search` requests.
365
+
366
+ ## Horizontal CPU placement
367
+
368
+ Comma-separated `--service-cluster` values produce separate Beaker task groups,
369
+ with Python, terminal, web-search, and local-search replicas divided as evenly
370
+ as possible. This forces placement across the named clusters instead of merely
371
+ asking Beaker for many replicas in one cluster:
372
+
373
+ ```bash
374
+ literegistry-base-deployment preview \
375
+ --registry=redis://registry.example:6379 \
376
+ --python-replicas=16 \
377
+ --terminal-replicas=16 \
378
+ --web-search-replicas=0 \
379
+ --service-cluster=ai2/neptune,ai2/saturn,ai2/jupiter,ai2/ceres
380
+ ```
381
+
382
+ GPU pools currently use the single explicit `--model-cluster`; tensor
383
+ parallelism maps directly to Beaker `gpuCount` for each model replica.
384
+
385
+ ## Python API
386
+
387
+ ```python
388
+ from literegistry_base_deployment import BaseDeploymentConfig, BaseDeploymentLauncher
389
+
390
+
391
+ config = BaseDeploymentConfig(
392
+ registry="redis://jupiter-cs-aus-183.reviz.ai2.in:59936",
393
+ python_replicas=4,
394
+ terminal_replicas=4,
395
+ web_search_replicas=2,
396
+ service_clusters=("ai2/jupiter", "ai2/ceres"),
397
+ gateway_cluster="ai2/jupiter",
398
+ )
399
+ launcher = BaseDeploymentLauncher(config)
400
+ print(launcher.preview()) # read-only
401
+ receipt = launcher.submit() # creates the Beaker experiment
402
+ print(receipt)
403
+
404
+ # Later:
405
+ BaseDeploymentLauncher.stop(receipt["beaker"]["id"])
406
+ ```
407
+
408
+ ## Development validation
409
+
410
+ ```bash
411
+ cd /weka/gfaria/literegistry/literegistry_base_deployment
412
+ python -m pip install -e '.[test,publish]'
413
+ python -m pytest
414
+ python -m build
415
+ python -m twine check dist/*
416
+ ```
417
+
418
+ The tests verify service composition, shell syntax, managed Redis discovery,
419
+ secret references, multi-cluster spreading, generation/classification vLLM
420
+ arguments, JTC Lucene image routing, package contents, and Dockerfile
421
+ self-containment. Dockerfiles are structurally tested by default; actually
422
+ building the four package images requires a Docker daemon and network access.
423
+
424
+
425
+ ## Experimental mirror soft affinity
426
+
427
+ The bundled gateway enables experimental mirror soft affinity by default. If this base deployment is used with external `docker-mirror` services, disable it with
428
+ `--docker-mirror-soft-affinity=False` to restore normal load balancing.