cortexgrid 0.2.95__tar.gz → 0.2.96__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: cortexgrid
3
- Version: 0.2.95
3
+ Version: 0.2.96
4
4
  Summary: Connect your ML code to the RoboLab compute cluster — Ray, MLflow, and S3
5
5
  Project-URL: Homepage, https://github.com/robodatalab/cortexgrid
6
6
  Project-URL: Repository, https://github.com/robodatalab/cortexgrid
@@ -178,7 +178,7 @@ deployed = cortexgrid.deploy_model("qwen", "instruct", saved.run_name, wait=True
178
178
  print(deployed.url)
179
179
  ```
180
180
 
181
- `save_model` saves a new copy under every run - meant for weights the run produced (e.g. a fine-tune). For a model produced elsewhere (e.g. a pretrained base model), `cortexgrid.import_model(source, MyServeApp, family, suffix)` uploads it once under `run_name=cortexgrid.IMPORTED` and is a no-op on later runs; deploy it with `deploy_model(family, suffix, cortexgrid.IMPORTED)`.
181
+ `save_model` saves a new copy under every run - meant for weights the run produced (e.g. a fine-tune). For a model produced elsewhere (e.g. a pretrained base model), `cortexgrid.import_model(source, MyServeApp, family, suffix)` uploads it once under `run_name=cortexgrid.IMPORTED` and on later runs only re-bundles `MyServeApp` if its code changed; deploy it with `deploy_model(family, suffix, cortexgrid.IMPORTED)`.
182
182
 
183
183
  `save_model` is synchronous (registry lifecycle: `uploading` -> `ready`); `deploy_model` schedules the serving lifecycle (`deploying` -> `running`). With `wait=True` a failed deploy raises `cortexgrid.ModelDeployFailed`; `cortexgrid.wait_for_model_serving(family, suffix, run_name, timeout=...)` waits on a deploy started elsewhere, and re-deploying a failed model retries it from scratch. See [model-serving.md](https://github.com/robodatalab/cortexgrid/blob/main/docs/cortexgrid/model-serving.md) for both lifecycles end to end - upload/deploy/undeploy/delete, status queries (`model_registry_status`, `model_serving_status`), and error handling.
184
184
 
@@ -12,6 +12,8 @@ Bundles of several seeds combine with `BundleDesc.merge`.
12
12
 
13
13
  `stage(files, dest)` lays a bundle out under `dest` at each file's import path,
14
14
  so `dest` on sys.path (e.g. a Ray working_dir) makes every module importable.
15
+ `digest(files)` hashes that layout, so bundles that stage identically compare
16
+ equal wherever their files live.
15
17
 
16
18
  `BundleDesc.pip_requirements(worker_provides())` pins the third-party
17
19
  distributions the Ray worker image does not already have, for a Ray `pip`
@@ -24,6 +26,7 @@ import ast
24
26
  from collections.abc import Iterable, Iterator
25
27
  from dataclasses import dataclass
26
28
  import functools
29
+ import hashlib
27
30
  import importlib.machinery
28
31
  import importlib.metadata
29
32
  import importlib.util
@@ -105,11 +108,23 @@ def stage(files: set[Path], dest: Path) -> None:
105
108
  in installed distributions)."""
106
109
  dest.mkdir(parents=True, exist_ok=True)
107
110
  for file in files:
108
- target = dest / file.relative_to(_sys_path_root(file))
111
+ target = dest / _import_path(file)
109
112
  target.parent.mkdir(parents=True, exist_ok=True)
110
113
  shutil.copy2(file, target)
111
114
 
112
115
 
116
+ def digest(files: set[Path]) -> str:
117
+ """SHA-256 of `files` as `stage` lays them out: each file's import path and
118
+ contents, in import-path order. Files that stage identically digest
119
+ identically, wherever they live on disk."""
120
+ sha = hashlib.sha256()
121
+ for path, file in sorted((_import_path(file).as_posix(), file) for file in files):
122
+ content = file.read_bytes()
123
+ sha.update(f"{path}\0{len(content)}\0".encode())
124
+ sha.update(content)
125
+ return sha.hexdigest()
126
+
127
+
113
128
  # What the Ray worker image pip-installs, as the Dockerfile spells it. They and
114
129
  # their dependency trees are on the worker already, so they are never installed
115
130
  # there again -- a second copy in the job's virtualenv would shadow the image's.
@@ -335,6 +350,11 @@ def _package(file: Path) -> str:
335
350
  return ".".join(parts)
336
351
 
337
352
 
353
+ def _import_path(file: Path) -> Path:
354
+ """`file` relative to the sys.path entry it is imported from."""
355
+ return file.relative_to(_sys_path_root(file))
356
+
357
+
338
358
  def _sys_path_root(file: Path) -> Path:
339
359
  """The sys.path entry `file` is imported from: its first ancestor directory
340
360
  without an __init__.py."""
@@ -3,11 +3,12 @@
3
3
  Caller stays HTTP-only: deploy/undeploy/list talk to the Ray dashboard's
4
4
  declarative `/api/serve/applications/` endpoint via [cortexgrid.ray_util],
5
5
  never `ray.init`. The deployment class is bundled at `save_model` time, zipped,
6
- uploaded to MinIO under `serve-bundles/<run_name>/<family>__<suffix>.zip`, and
6
+ uploaded to MinIO under
7
+ `serve-bundles/<run_name>/<family>__<suffix>/<fingerprint>.zip`, and
7
8
  referenced via `runtime_env.working_dir` so Ray workers fetch it from there.
8
- The bundle URL, class import path, and pip list are persisted as MLflow tags
9
- on the ModelVersion so `deploy_model` can find them later without the caller
10
- holding the class object.
9
+ The bundle URL, class import path, pip list, and fingerprint are persisted as
10
+ MLflow tags on the ModelVersion so `deploy_model` can find them later without
11
+ the caller holding the class object.
11
12
 
12
13
  Naming: the Ray Serve application is named "<family>__<suffix>__<run_name>".
13
14
  This relies on family/suffix/run_name not containing the literal "__".
@@ -26,13 +27,14 @@ import shutil
26
27
  import tempfile
27
28
  import time
28
29
  from dataclasses import dataclass, field
30
+ import hashlib
29
31
  from pathlib import Path
30
32
  from typing import Any
31
33
 
32
34
  from mlflow.tracking import MlflowClient
33
35
  from ray.serve.schema import ApplicationStatus
34
36
 
35
- from cortexgrid._bundle import bundle, stage, worker_provides
37
+ from cortexgrid._bundle import bundle, digest, stage, worker_provides
36
38
  from cortexgrid.infra import get_mlflow_tracking_uri, get_ray_serve_uri
37
39
  from cortexgrid.ray_util import (
38
40
  get_serve_details,
@@ -95,17 +97,28 @@ class BundleMetadata:
95
97
  # pinned third-party requirements the replica pip-installs (the bundle's
96
98
  # distributions the Ray image does not already provide)
97
99
  pip_requirements: list[str] = field(default_factory=list)
100
+ # ServeBundle.fingerprint of the uploaded bundle; empty for models saved
101
+ # before bundles were fingerprinted
102
+ fingerprint: str = ""
98
103
 
99
104
 
100
- def bundle_class(
101
- cls: type, family: str, suffix: str, run_name: str
102
- ) -> BundleMetadata:
103
- """Bundle the serve-app class's code (and the serve entrypoint), zip it, and
104
- upload to MinIO.
105
+ @dataclass
106
+ class ServeBundle:
107
+ """A serve-app's bundle, resolved locally but not uploaded yet: what
108
+ `build_bundle` finds and `upload_bundle` ships."""
105
109
 
106
- Returns the metadata `deploy_model` needs later; callers (typically
107
- `save_model`) persist it on the ModelVersion so the deploy step can run
108
- without holding the class object.
110
+ files: set[Path]
111
+ class_import_path: str
112
+ pip_requirements: list[str]
113
+ # Hash of everything the replica runs: the staged files, the class it
114
+ # imports, and the requirements it installs. Equal fingerprints mean the
115
+ # same code, so an uploaded bundle can be reused.
116
+ fingerprint: str
117
+
118
+
119
+ def build_bundle(cls: type) -> ServeBundle:
120
+ """Resolve the serve-app class's code (and the serve entrypoint) into a
121
+ bundle, without uploading it.
109
122
 
110
123
  Raises ValueError for a class wrapped by `ray.serve.ingress`: that wrapper
111
124
  is a subclass Ray defines in its own module, and on older Ray (e.g. 2.9) it
@@ -120,17 +133,51 @@ def bundle_class(
120
133
  entry_file = Path(inspect.getfile(cls)).resolve()
121
134
  serve_entry = Path(__file__).with_name("_serve_entry.py")
122
135
  desc = bundle(entry_file).merge(bundle(serve_entry))
136
+ class_import_path = f"{cls.__module__}:{cls.__name__}"
123
137
  pip_requirements = desc.pip_requirements(worker_provides())
138
+ fingerprint = hashlib.sha256(
139
+ json.dumps(
140
+ [digest(desc.local_files), class_import_path, pip_requirements]
141
+ ).encode()
142
+ ).hexdigest()
143
+ return ServeBundle(
144
+ files=desc.local_files,
145
+ class_import_path=class_import_path,
146
+ pip_requirements=pip_requirements,
147
+ fingerprint=fingerprint,
148
+ )
149
+
150
+
151
+ def bundle_class(
152
+ cls: type, family: str, suffix: str, run_name: str
153
+ ) -> BundleMetadata:
154
+ """Bundle the serve-app class's code (and the serve entrypoint), zip it, and
155
+ upload to MinIO: `build_bundle` followed by `upload_bundle`.
156
+
157
+ Returns the metadata `deploy_model` needs later; callers (typically
158
+ `save_model`) persist it on the ModelVersion so the deploy step can run
159
+ without holding the class object."""
160
+ return upload_bundle(build_bundle(cls), family, suffix, run_name)
161
+
162
+
163
+ def upload_bundle(
164
+ serve_bundle: ServeBundle, family: str, suffix: str, run_name: str
165
+ ) -> BundleMetadata:
166
+ """Zip a built bundle and upload it under its fingerprint.
167
+
168
+ The fingerprint is part of the URL because Ray keeps a remote working_dir
169
+ it has downloaded and reuses it for the same URL: new code at an old URL
170
+ would never reach a replica."""
124
171
  with tempfile.TemporaryDirectory() as tmp:
125
172
  code_root = Path(tmp) / "code"
126
- stage(desc.local_files, code_root)
173
+ stage(serve_bundle.files, code_root)
127
174
  log.info(
128
175
  "Serve bundle for %s/%s/%s: %d files, pip: %s",
129
176
  family,
130
177
  suffix,
131
178
  run_name,
132
- len(desc.local_files),
133
- pip_requirements,
179
+ len(serve_bundle.files),
180
+ serve_bundle.pip_requirements,
134
181
  )
135
182
  # Ray unpacks a remote (s3://) working_dir zip by stripping its
136
183
  # top-level directory when there is exactly one, so a bundle of a single
@@ -146,12 +193,16 @@ def bundle_class(
146
193
  )
147
194
  bundle_url = upload(
148
195
  archive,
149
- dest_path=f"serve-bundles/{run_name}/{family}__{suffix}.zip",
196
+ dest_path=(
197
+ f"serve-bundles/{run_name}/{family}__{suffix}/"
198
+ f"{serve_bundle.fingerprint}.zip"
199
+ ),
150
200
  )
151
201
  return BundleMetadata(
152
202
  bundle_url=bundle_url,
153
- class_import_path=f"{cls.__module__}:{cls.__name__}",
154
- pip_requirements=pip_requirements,
203
+ class_import_path=serve_bundle.class_import_path,
204
+ pip_requirements=serve_bundle.pip_requirements,
205
+ fingerprint=serve_bundle.fingerprint,
155
206
  )
156
207
 
157
208
 
@@ -189,19 +240,34 @@ def _build_application_spec(
189
240
  _CLASS_IMPORT_PATH_TAG = "class_import_path"
190
241
  _BUNDLE_URL_TAG = "serve_bundle_url"
191
242
  _PIP_REQUIREMENTS_TAG = "serve_pip_requirements"
243
+ _BUNDLE_FINGERPRINT_TAG = "serve_bundle_fingerprint"
192
244
 
193
245
 
194
246
  def metadata_to_tags(meta: BundleMetadata) -> dict[str, str]:
195
247
  """Serialise BundleMetadata to MLflow tags. The inverse of
196
- `_load_bundle_metadata`; lives here next to the consumer so the tag schema
248
+ `metadata_from_tags`; lives here next to the consumer so the tag schema
197
249
  stays in one place."""
198
250
  return {
199
251
  _CLASS_IMPORT_PATH_TAG: meta.class_import_path,
200
252
  _BUNDLE_URL_TAG: meta.bundle_url,
201
253
  _PIP_REQUIREMENTS_TAG: json.dumps(meta.pip_requirements),
254
+ _BUNDLE_FINGERPRINT_TAG: meta.fingerprint,
202
255
  }
203
256
 
204
257
 
258
+ def metadata_from_tags(tags: dict[str, str]) -> BundleMetadata:
259
+ """Deserialise BundleMetadata from a ModelVersion's MLflow tags. Raises
260
+ KeyError for a missing bundle URL or class import path."""
261
+ return BundleMetadata(
262
+ bundle_url=tags[_BUNDLE_URL_TAG],
263
+ class_import_path=tags[_CLASS_IMPORT_PATH_TAG],
264
+ # Absent on models saved before dependencies were pip-installed.
265
+ pip_requirements=json.loads(tags.get(_PIP_REQUIREMENTS_TAG, "[]")),
266
+ # Absent on models saved before bundles were fingerprinted.
267
+ fingerprint=tags.get(_BUNDLE_FINGERPRINT_TAG, ""),
268
+ )
269
+
270
+
205
271
  def _load_bundle_metadata(
206
272
  family: str, suffix: str, run_name: str
207
273
  ) -> BundleMetadata:
@@ -215,14 +281,8 @@ def _load_bundle_metadata(
215
281
  raise ValueError(
216
282
  f"No saved model for {family}/{suffix}/{run_name}; cannot deploy."
217
283
  )
218
- tags = versions[0].tags or {}
219
284
  try:
220
- return BundleMetadata(
221
- bundle_url=tags[_BUNDLE_URL_TAG],
222
- class_import_path=tags[_CLASS_IMPORT_PATH_TAG],
223
- # Absent on models saved before dependencies were pip-installed.
224
- pip_requirements=json.loads(tags.get(_PIP_REQUIREMENTS_TAG, "[]")),
225
- )
285
+ return metadata_from_tags(versions[0].tags or {})
226
286
  except KeyError as exc:
227
287
  raise ValueError(
228
288
  f"Saved model {family}/{suffix}/{run_name} is missing the deployment "
@@ -11,8 +11,9 @@ Mapping cortexgrid taxonomy <-> MLflow Registry:
11
11
 
12
12
  Two ways in: `save_model` registers a fresh copy under the calling run's
13
13
  run_name every time it runs (fine-tuned output); `import_model` registers a
14
- model produced elsewhere once, under the fixed run_name IMPORTED, and is a
15
- no-op after that. Both write the same layout, so every
14
+ model produced elsewhere once, under the fixed run_name IMPORTED, and after
15
+ that only re-bundles the serve-app when its code changed. Both write the same
16
+ layout, so every
16
17
  (family, suffix, run_name) consumer - load_model, deploy_model - handles both.
17
18
 
18
19
  storage.py is pure: it takes run_id/run_name as explicit args and never reads
@@ -35,8 +36,11 @@ from mlflow.tracking import MlflowClient
35
36
  from cortexgrid import s3_util
36
37
  from cortexgrid.infra import get_mlflow_tracking_uri, get_s3_bucket
37
38
  from cortexgrid.model_serving import (
39
+ build_bundle,
38
40
  bundle_class,
41
+ metadata_from_tags,
39
42
  metadata_to_tags,
43
+ upload_bundle,
40
44
  )
41
45
 
42
46
 
@@ -184,8 +188,10 @@ def import_model(
184
188
  while it downloads.
185
189
 
186
190
  If a version is already registered under the key:
187
- - "ready": no-op, returns it. `source` and `serve_app` are ignored; to
188
- replace the weights or the serve-app, `delete_model` it first.
191
+ - "ready": returns it without calling `source`. If `serve_app`'s code no
192
+ longer matches the stored bundle, it is re-bundled first and the
193
+ weights are kept (see `_refresh_bundle`). To replace the weights,
194
+ `delete_model` it first.
189
195
  - "uploading": raises RuntimeError - another process is importing it.
190
196
  - "upload_failed" / "broken": deleted and imported again.
191
197
 
@@ -195,6 +201,7 @@ def import_model(
195
201
  existing = model_registry_status(family, suffix, IMPORTED)
196
202
  if existing is not None:
197
203
  if existing.phase == _PHASE_READY:
204
+ _refresh_bundle(serve_app, family, suffix)
198
205
  return existing
199
206
  if existing.phase == _PHASE_UPLOADING:
200
207
  raise RuntimeError(
@@ -205,6 +212,27 @@ def import_model(
205
212
  return _upload_model(source, serve_app, suffix, family, None, IMPORTED)
206
213
 
207
214
 
215
+ def _refresh_bundle(serve_app: type, family: str, suffix: str) -> None:
216
+ """Re-bundle an imported model's serve-app when its fingerprint differs
217
+ from the bundle stored on the version, leaving the weights in place.
218
+
219
+ The new bundle is uploaded next to the old one and the version's tags are
220
+ pointed at it, so the next `deploy_model` runs the new code. An app that is
221
+ already running keeps the code it started with until it is deployed again;
222
+ the old bundle stays in storage so that app can still restart."""
223
+ name = f"{family}__{suffix}"
224
+ client = MlflowClient(tracking_uri=get_mlflow_tracking_uri())
225
+ version = client.search_model_versions(
226
+ f"name='{name}' and tags.run_name='{IMPORTED}'"
227
+ )[0]
228
+ serve_bundle = build_bundle(serve_app)
229
+ if metadata_from_tags(version.tags).fingerprint == serve_bundle.fingerprint:
230
+ return
231
+ meta = upload_bundle(serve_bundle, family, suffix, IMPORTED)
232
+ for key, value in metadata_to_tags(meta).items():
233
+ client.set_model_version_tag(name, version.version, key, value)
234
+
235
+
208
236
  def _upload_model(
209
237
  weights: str | Path | Callable[[], str | Path],
210
238
  serve_app: type,
@@ -150,7 +150,7 @@ deployed = cortexgrid.deploy_model("qwen", "instruct", saved.run_name, wait=True
150
150
  print(deployed.url)
151
151
  ```
152
152
 
153
- `save_model` saves a new copy under every run - meant for weights the run produced (e.g. a fine-tune). For a model produced elsewhere (e.g. a pretrained base model), `cortexgrid.import_model(source, MyServeApp, family, suffix)` uploads it once under `run_name=cortexgrid.IMPORTED` and is a no-op on later runs; deploy it with `deploy_model(family, suffix, cortexgrid.IMPORTED)`.
153
+ `save_model` saves a new copy under every run - meant for weights the run produced (e.g. a fine-tune). For a model produced elsewhere (e.g. a pretrained base model), `cortexgrid.import_model(source, MyServeApp, family, suffix)` uploads it once under `run_name=cortexgrid.IMPORTED` and on later runs only re-bundles `MyServeApp` if its code changed; deploy it with `deploy_model(family, suffix, cortexgrid.IMPORTED)`.
154
154
 
155
155
  `save_model` is synchronous (registry lifecycle: `uploading` -> `ready`); `deploy_model` schedules the serving lifecycle (`deploying` -> `running`). With `wait=True` a failed deploy raises `cortexgrid.ModelDeployFailed`; `cortexgrid.wait_for_model_serving(family, suffix, run_name, timeout=...)` waits on a deploy started elsewhere, and re-deploying a failed model retries it from scratch. See [model-serving.md](https://github.com/robodatalab/cortexgrid/blob/main/docs/cortexgrid/model-serving.md) for both lifecycles end to end - upload/deploy/undeploy/delete, status queries (`model_registry_status`, `model_serving_status`), and error handling.
156
156
 
@@ -1,6 +1,6 @@
1
1
  [project]
2
2
  name = "cortexgrid"
3
- version = "0.2.95"
3
+ version = "0.2.96"
4
4
  description = "Connect your ML code to the RoboLab compute cluster — Ray, MLflow, and S3"
5
5
  readme = "docs/cortexgrid/README.md"
6
6
  license = "Apache-2.0"
File without changes
File without changes