polnor 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
polnor-0.1.0/PKG-INFO ADDED
@@ -0,0 +1,113 @@
1
+ Metadata-Version: 2.4
2
+ Name: polnor
3
+ Version: 0.1.0
4
+ Summary: Polnor Python SDK — query Iceberg lakehouses, manage notebooks/jobs/models, MLflow-compatible tracking.
5
+ Author-email: Polnor <contact@polnor.net>
6
+ License: Apache-2.0
7
+ Project-URL: Homepage, https://polnor.net
8
+ Project-URL: Documentation, https://docs.polnor.net/sdk/python
9
+ Project-URL: Source, https://github.com/polnor/polnor
10
+ Keywords: polnor,lakehouse,iceberg,sql,mlflow,data-platform
11
+ Classifier: Development Status :: 3 - Alpha
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3 :: Only
15
+ Classifier: Programming Language :: Python :: 3.9
16
+ Classifier: Programming Language :: Python :: 3.10
17
+ Classifier: Programming Language :: Python :: 3.11
18
+ Classifier: Programming Language :: Python :: 3.12
19
+ Classifier: License :: OSI Approved :: Apache Software License
20
+ Requires-Python: >=3.9
21
+ Description-Content-Type: text/markdown
22
+ Requires-Dist: requests>=2.28
23
+ Provides-Extra: test
24
+ Requires-Dist: pytest>=7; extra == "test"
25
+ Requires-Dist: responses>=0.24; extra == "test"
26
+
27
+ # Polnor Python SDK
28
+
29
+ MLflow-compatible tracking for jobs running on the Polnor lakehouse.
30
+
31
+ ## Install
32
+
33
+ ```bash
34
+ pip install polnor
35
+ ```
36
+
37
+ Or, from a local checkout (contributors / pre-release):
38
+
39
+ ```bash
40
+ pip install -e ./sdk/python
41
+ ```
42
+
43
+ ## Configure
44
+
45
+ Set three environment variables before importing the SDK:
46
+
47
+ | Variable | Required | Notes |
48
+ |---|---|---|
49
+ | `POLNOR_API_URL` | yes | e.g. `https://api.polnor.net` |
50
+ | `POLNOR_TOKEN` | yes | Personal access token from `/settings/tokens` |
51
+ | `POLNOR_EXPERIMENT_ID` | optional | Default experiment for `start_run()` |
52
+ | `POLNOR_WORKSPACE_SLUG` | optional | When the token isn't workspace-scoped |
53
+
54
+ Inside a Polnor job these are injected automatically.
55
+
56
+ ## Quickstart
57
+
58
+ ```python
59
+ from polnor import mlflow
60
+
61
+ with mlflow.start_run(run_name="train-resnet50") as run:
62
+ mlflow.log_params({"lr": 0.01, "batch_size": 32, "optimizer": "adam"})
63
+
64
+ for epoch in range(10):
65
+ loss, acc = train_one_epoch()
66
+ mlflow.log_metrics({"loss": loss, "accuracy": acc}, step=epoch)
67
+
68
+ mlflow.set_tag("git_commit", os.environ.get("CI_COMMIT_SHA", "dev"))
69
+ ```
70
+
71
+ The run handle is also usable without the `with` block:
72
+
73
+ ```python
74
+ run = mlflow.start_run()
75
+ try:
76
+ mlflow.log_metric("score", 0.93)
77
+ finally:
78
+ mlflow.end_run("completed")
79
+ ```
80
+
81
+ ## What's supported
82
+
83
+ | API | Status |
84
+ |---|---|
85
+ | `start_run` / `end_run` / `active_run` | ✅ |
86
+ | `log_metric` / `log_metrics` (latest value wins) | ✅ |
87
+ | `log_param` / `log_params` | ✅ |
88
+ | `set_tag` / `set_tags` | ✅ |
89
+ | `autolog()` context manager | ✅ |
90
+ | Per-step metric history (time series) | ⏳ planned |
91
+ | `log_artifact` (presigned-URL upload) | ⏳ planned |
92
+ | Sklearn / PyTorch / TF auto-instrumentation | ⏳ planned |
93
+
94
+ ## How it differs from upstream MLflow
95
+
96
+ * Server-side, metrics merge into a single JSONB blob — the latest
97
+ value per key wins. The full `(step, value, timestamp)` history is
98
+ in scope for a follow-up migration.
99
+ * No server-side parameter immutability check yet — calling
100
+ `log_param("lr", 0.02)` after `log_param("lr", 0.01)` overwrites
101
+ silently. Upstream MLflow rejects with 400.
102
+ * No `mlflow.search_runs()` — use the REST API
103
+ (`GET /api/v1/experiments/:id/runs`) directly until we wrap it.
104
+
105
+ ## Development
106
+
107
+ ```bash
108
+ pip install -e ".[test]"
109
+ pytest
110
+ ```
111
+
112
+ Tests use [responses](https://github.com/getsentry/responses) to mock
113
+ HTTP — no live API needed.
polnor-0.1.0/README.md ADDED
@@ -0,0 +1,87 @@
1
+ # Polnor Python SDK
2
+
3
+ MLflow-compatible tracking for jobs running on the Polnor lakehouse.
4
+
5
+ ## Install
6
+
7
+ ```bash
8
+ pip install polnor
9
+ ```
10
+
11
+ Or, from a local checkout (contributors / pre-release):
12
+
13
+ ```bash
14
+ pip install -e ./sdk/python
15
+ ```
16
+
17
+ ## Configure
18
+
19
+ Set three environment variables before importing the SDK:
20
+
21
+ | Variable | Required | Notes |
22
+ |---|---|---|
23
+ | `POLNOR_API_URL` | yes | e.g. `https://api.polnor.net` |
24
+ | `POLNOR_TOKEN` | yes | Personal access token from `/settings/tokens` |
25
+ | `POLNOR_EXPERIMENT_ID` | optional | Default experiment for `start_run()` |
26
+ | `POLNOR_WORKSPACE_SLUG` | optional | When the token isn't workspace-scoped |
27
+
28
+ Inside a Polnor job these are injected automatically.
29
+
30
+ ## Quickstart
31
+
32
+ ```python
33
+ from polnor import mlflow
34
+
35
+ with mlflow.start_run(run_name="train-resnet50") as run:
36
+ mlflow.log_params({"lr": 0.01, "batch_size": 32, "optimizer": "adam"})
37
+
38
+ for epoch in range(10):
39
+ loss, acc = train_one_epoch()
40
+ mlflow.log_metrics({"loss": loss, "accuracy": acc}, step=epoch)
41
+
42
+ mlflow.set_tag("git_commit", os.environ.get("CI_COMMIT_SHA", "dev"))
43
+ ```
44
+
45
+ The run handle is also usable without the `with` block:
46
+
47
+ ```python
48
+ run = mlflow.start_run()
49
+ try:
50
+ mlflow.log_metric("score", 0.93)
51
+ finally:
52
+ mlflow.end_run("completed")
53
+ ```
54
+
55
+ ## What's supported
56
+
57
+ | API | Status |
58
+ |---|---|
59
+ | `start_run` / `end_run` / `active_run` | ✅ |
60
+ | `log_metric` / `log_metrics` (latest value wins) | ✅ |
61
+ | `log_param` / `log_params` | ✅ |
62
+ | `set_tag` / `set_tags` | ✅ |
63
+ | `autolog()` context manager | ✅ |
64
+ | Per-step metric history (time series) | ⏳ planned |
65
+ | `log_artifact` (presigned-URL upload) | ⏳ planned |
66
+ | Sklearn / PyTorch / TF auto-instrumentation | ⏳ planned |
67
+
68
+ ## How it differs from upstream MLflow
69
+
70
+ * Server-side, metrics merge into a single JSONB blob — the latest
71
+ value per key wins. The full `(step, value, timestamp)` history is
72
+ in scope for a follow-up migration.
73
+ * No server-side parameter immutability check yet — calling
74
+ `log_param("lr", 0.02)` after `log_param("lr", 0.01)` overwrites
75
+ silently. Upstream MLflow rejects with 400.
76
+ * No `mlflow.search_runs()` — use the REST API
77
+ (`GET /api/v1/experiments/:id/runs`) directly until we wrap it.
78
+
79
+ ## Development
80
+
81
+ ```bash
82
+ pip install -e ".[test]"
83
+ pytest
84
+ ```
85
+
86
+ Tests use [responses](https://github.com/getsentry/responses) to mock
87
+ HTTP — no live API needed.
@@ -0,0 +1,25 @@
1
+ """Polnor Python SDK.
2
+
3
+ Today this exposes the MLflow-compatible tracking surface
4
+ (:mod:`polnor.mlflow`). More modules — auth, jobs, runs, datasets —
5
+ land here as the API stabilises.
6
+
7
+ The SDK is auto-configured from environment variables when running
8
+ inside a Polnor job. Outside a job, set them yourself::
9
+
10
+ export POLNOR_API_URL=https://api.polnor.net
11
+ export POLNOR_TOKEN=<personal access token>
12
+ export POLNOR_EXPERIMENT_ID=<uuid of an experiment>
13
+
14
+ Then::
15
+
16
+ from polnor import mlflow
17
+ with mlflow.start_run() as run:
18
+ mlflow.log_param("lr", 0.01)
19
+ mlflow.log_metric("accuracy", 0.93)
20
+ """
21
+
22
+ from . import mlflow
23
+
24
+ __all__ = ["mlflow"]
25
+ __version__ = "0.1.0"
@@ -0,0 +1,132 @@
1
+ """Thin HTTP client used by polnor.mlflow.
2
+
3
+ Kept separate so future SDK modules (jobs, datasets) can share the
4
+ same auth + retry plumbing without dragging the MLflow surface in.
5
+ """
6
+
7
+ from __future__ import annotations
8
+
9
+ import os
10
+ import time
11
+ import urllib.parse
12
+ from dataclasses import dataclass
13
+ from typing import Any, Mapping
14
+
15
+ try:
16
+ import requests
17
+ except ImportError as e: # pragma: no cover — caught early in install
18
+ raise ImportError(
19
+ "polnor SDK requires the 'requests' package. "
20
+ "Install with `pip install polnor` or `pip install requests`."
21
+ ) from e
22
+
23
+
24
+ class PolnorAPIError(RuntimeError):
25
+ """Raised on a non-2xx API response with the body verbatim."""
26
+
27
+ def __init__(self, status: int, body: str, *, method: str, url: str):
28
+ super().__init__(f"{method} {url} → {status}: {body}")
29
+ self.status = status
30
+ self.body = body
31
+
32
+
33
+ @dataclass
34
+ class Config:
35
+ """SDK configuration. Read from env by default; override for tests."""
36
+
37
+ api_url: str
38
+ token: str
39
+ workspace_slug: str | None = None
40
+ timeout: float = 30.0
41
+ max_retries: int = 3
42
+ backoff_base: float = 0.5
43
+
44
+ @classmethod
45
+ def from_env(cls) -> "Config":
46
+ api_url = os.environ.get("POLNOR_API_URL", "").rstrip("/")
47
+ token = os.environ.get("POLNOR_TOKEN", "")
48
+ workspace = os.environ.get("POLNOR_WORKSPACE_SLUG") or None
49
+ if not api_url:
50
+ raise RuntimeError(
51
+ "POLNOR_API_URL is not set — point it at https://api.polnor.net "
52
+ "(or your self-hosted control plane)."
53
+ )
54
+ if not token:
55
+ raise RuntimeError(
56
+ "POLNOR_TOKEN is not set — generate a personal access token "
57
+ "from /settings/tokens and export POLNOR_TOKEN."
58
+ )
59
+ return cls(api_url=api_url, token=token, workspace_slug=workspace)
60
+
61
+
62
+ class Client:
63
+ """HTTP client with token auth + retry on 5xx / connection errors.
64
+
65
+ Idempotency: every method here is safe to retry — they're either
66
+ pure GETs or POST/PATCH endpoints that the server side already
67
+ handles idempotently (start_run with the same body would create a
68
+ second run, but that's a deliberate design — MLflow does the same).
69
+ """
70
+
71
+ def __init__(self, cfg: Config | None = None):
72
+ self.cfg = cfg or Config.from_env()
73
+ self._session = requests.Session()
74
+ self._session.headers.update({
75
+ "Authorization": f"Bearer {self.cfg.token}",
76
+ "Content-Type": "application/json",
77
+ "User-Agent": "polnor-sdk-python/0.1.0",
78
+ })
79
+ if self.cfg.workspace_slug:
80
+ self._session.headers["X-Workspace-Slug"] = self.cfg.workspace_slug
81
+
82
+ # --- low-level ---
83
+
84
+ def request(self, method: str, path: str, *, json: Any | None = None,
85
+ params: Mapping[str, Any] | None = None) -> Any:
86
+ url = self._build_url(path, params)
87
+ last_err: Exception | None = None
88
+ for attempt in range(self.cfg.max_retries):
89
+ try:
90
+ resp = self._session.request(
91
+ method, url, json=json, timeout=self.cfg.timeout
92
+ )
93
+ except requests.RequestException as e:
94
+ last_err = e
95
+ self._sleep(attempt)
96
+ continue
97
+
98
+ if resp.status_code >= 500:
99
+ last_err = PolnorAPIError(resp.status_code, resp.text, method=method, url=url)
100
+ self._sleep(attempt)
101
+ continue
102
+ if resp.status_code >= 400:
103
+ raise PolnorAPIError(resp.status_code, resp.text, method=method, url=url)
104
+
105
+ if resp.status_code == 204 or not resp.content:
106
+ return None
107
+ return resp.json()
108
+
109
+ if last_err:
110
+ raise last_err
111
+ return None
112
+
113
+ def _build_url(self, path: str, params: Mapping[str, Any] | None) -> str:
114
+ base = f"{self.cfg.api_url}/api/v1{path}"
115
+ if not params:
116
+ return base
117
+ return f"{base}?{urllib.parse.urlencode({k: v for k, v in params.items() if v is not None})}"
118
+
119
+ def _sleep(self, attempt: int) -> None:
120
+ delay = self.cfg.backoff_base * (2 ** attempt)
121
+ time.sleep(delay)
122
+
123
+ # --- convenience wrappers ---
124
+
125
+ def post(self, path: str, json: Any | None = None) -> Any:
126
+ return self.request("POST", path, json=json)
127
+
128
+ def patch(self, path: str, json: Any | None = None) -> Any:
129
+ return self.request("PATCH", path, json=json)
130
+
131
+ def get(self, path: str, params: Mapping[str, Any] | None = None) -> Any:
132
+ return self.request("GET", path, params=params)
@@ -0,0 +1,300 @@
1
+ """MLflow-compatible tracking surface for Polnor.
2
+
3
+ Drop-in shim for the bits of MLflow most jobs use::
4
+
5
+ from polnor import mlflow
6
+
7
+ with mlflow.start_run() as run:
8
+ mlflow.log_param("learning_rate", 0.01)
9
+ mlflow.log_param("optimizer", "adam")
10
+
11
+ for epoch in range(10):
12
+ mlflow.log_metric("loss", train_loss, step=epoch)
13
+ mlflow.log_metric("accuracy", val_acc, step=epoch)
14
+
15
+ mlflow.set_tag("git_commit", os.environ["CI_COMMIT_SHA"])
16
+
17
+ The active run is stored in a thread-local so nested calls in the
18
+ same job don't fight over a global; calling ``start_run`` twice
19
+ without ``end_run`` between them stacks runs the way MLflow's nested
20
+ runs do.
21
+
22
+ The following MLflow API is *not* implemented yet and will raise
23
+ ``NotImplementedError``:
24
+
25
+ * ``log_artifact`` — needs a presigned-URL upload flow
26
+ * ``log_image`` / ``log_dict`` — derived from artifacts
27
+ * ``mlflow.search_runs`` — read-side, use the REST API directly
28
+
29
+ Server-side, metrics are stored as the *latest* value per key (JSONB
30
+ merge). A full time-series of ``(step, value, timestamp)`` per metric
31
+ is tracked separately in a future migration; until then,
32
+ ``log_metric(..., step=...)`` records the value but not the step.
33
+ """
34
+
35
+ from __future__ import annotations
36
+
37
+ import contextlib
38
+ import os
39
+ import threading
40
+ from dataclasses import dataclass
41
+ from typing import Any, Iterator, Mapping
42
+
43
+ from ._client import Client
44
+
45
+
46
+ # Thread-local stack of active runs; supports the nested-run pattern
47
+ # used by frameworks that spawn child runs per CV fold.
48
+ _local = threading.local()
49
+
50
+
51
+ def _stack() -> list["ActiveRun"]:
52
+ if not hasattr(_local, "runs"):
53
+ _local.runs = []
54
+ return _local.runs
55
+
56
+
57
+ def _client() -> Client:
58
+ """Single Client per thread, lazily constructed.
59
+
60
+ Reusing the same Session across calls means TLS handshakes and
61
+ auth resolution amortize across log_metric calls in a tight loop.
62
+ """
63
+ if not hasattr(_local, "client"):
64
+ _local.client = Client()
65
+ return _local.client
66
+
67
+
68
+ @dataclass
69
+ class ActiveRun:
70
+ """Handle to a running experiment_run on the server.
71
+
72
+ Returned by :func:`start_run`; usable as a context manager so the
73
+ run is finalized on scope exit even when an exception escapes.
74
+ """
75
+
76
+ experiment_id: str
77
+ run_id: str
78
+ name: str
79
+
80
+ def __enter__(self) -> "ActiveRun":
81
+ _stack().append(self)
82
+ return self
83
+
84
+ def __exit__(self, exc_type, exc, tb) -> None:
85
+ # End with status='failed' when an exception escaped, otherwise
86
+ # 'completed'. Errors in end_run itself are swallowed so they
87
+ # don't replace the user's exception.
88
+ status = "failed" if exc_type is not None else "completed"
89
+ try:
90
+ end_run(status=status)
91
+ except Exception:
92
+ pass
93
+
94
+
95
+ def _experiment_id_from_env() -> str:
96
+ exp_id = os.environ.get("POLNOR_EXPERIMENT_ID")
97
+ if not exp_id:
98
+ raise RuntimeError(
99
+ "No active experiment. Either:\n"
100
+ " - export POLNOR_EXPERIMENT_ID=<uuid>, or\n"
101
+ " - call mlflow.start_run(experiment_id=...)."
102
+ )
103
+ return exp_id
104
+
105
+
106
+ def start_run(
107
+ *,
108
+ experiment_id: str | None = None,
109
+ run_name: str | None = None,
110
+ params: Mapping[str, Any] | None = None,
111
+ tags: Mapping[str, str] | None = None,
112
+ ) -> ActiveRun:
113
+ """Start a new run and make it the active one.
114
+
115
+ The returned :class:`ActiveRun` is usable as a context manager. If
116
+ ``experiment_id`` is omitted, ``POLNOR_EXPERIMENT_ID`` is used.
117
+ """
118
+ exp_id = experiment_id or _experiment_id_from_env()
119
+ body: dict[str, Any] = {}
120
+ if run_name:
121
+ body["name"] = run_name
122
+ if params:
123
+ body["params"] = dict(params)
124
+ if tags:
125
+ body["tags"] = dict(tags)
126
+ resp = _client().post(f"/experiments/{exp_id}/runs", body)
127
+ run = ActiveRun(experiment_id=exp_id, run_id=resp["id"], name=resp.get("name", ""))
128
+ _stack().append(run)
129
+ return run
130
+
131
+
132
+ def active_run() -> ActiveRun | None:
133
+ """Return the run on top of the active stack, or None."""
134
+ s = _stack()
135
+ return s[-1] if s else None
136
+
137
+
138
+ def _require_active() -> ActiveRun:
139
+ run = active_run()
140
+ if run is None:
141
+ raise RuntimeError("No active run. Call start_run() first.")
142
+ return run
143
+
144
+
145
+ def end_run(status: str = "completed") -> None:
146
+ """Finish the active run with the given status.
147
+
148
+ Pops the top of the active stack — nested runs continue to use the
149
+ next-outer parent.
150
+ """
151
+ s = _stack()
152
+ if not s:
153
+ return # no-op when called outside a run
154
+ run = s.pop()
155
+ _client().patch(
156
+ f"/experiments/{run.experiment_id}/runs/{run.run_id}",
157
+ {"status": status},
158
+ )
159
+
160
+
161
+ def log_metric(key: str, value: float, step: int | None = None) -> None:
162
+ """Record a single metric value on the active run.
163
+
164
+ ``step`` is accepted for MLflow API compatibility but currently
165
+ ignored server-side (latest-value-wins). The server will track
166
+ full history once the per-step migration lands.
167
+ """
168
+ run = _require_active()
169
+ body: dict[str, Any] = {"key": key, "value": float(value)}
170
+ if step is not None:
171
+ body["step"] = int(step)
172
+ _client().post(
173
+ f"/experiments/{run.experiment_id}/runs/{run.run_id}/metrics",
174
+ body,
175
+ )
176
+
177
+
178
+ def log_metrics(metrics: Mapping[str, float], step: int | None = None) -> None:
179
+ """Batch variant — single round-trip per dict."""
180
+ run = _require_active()
181
+ _client().post(
182
+ f"/experiments/{run.experiment_id}/runs/{run.run_id}/metrics/batch",
183
+ {"metrics": {k: float(v) for k, v in metrics.items()}},
184
+ )
185
+ _ = step # see log_metric for why we accept-and-ignore
186
+
187
+
188
+ def log_param(key: str, value: Any) -> None:
189
+ """Record a hyperparameter on the active run."""
190
+ run = _require_active()
191
+ _client().post(
192
+ f"/experiments/{run.experiment_id}/runs/{run.run_id}/params",
193
+ {"key": key, "value": value},
194
+ )
195
+
196
+
197
+ def log_params(params: Mapping[str, Any]) -> None:
198
+ """Convenience: log_param in a loop."""
199
+ for k, v in params.items():
200
+ log_param(k, v)
201
+
202
+
203
+ def set_tag(key: str, value: str) -> None:
204
+ run = _require_active()
205
+ _client().post(
206
+ f"/experiments/{run.experiment_id}/runs/{run.run_id}/tags",
207
+ {"key": key, "value": str(value)},
208
+ )
209
+
210
+
211
+ def set_tags(tags: Mapping[str, str]) -> None:
212
+ for k, v in tags.items():
213
+ set_tag(k, v)
214
+
215
+
216
+ def log_artifact(local_path: str, artifact_path: str | None = None) -> None:
217
+ """Upload a local file as an artifact of the active run.
218
+
219
+ Two-step flow: ask the server for a presigned PUT URL, then
220
+ stream the file to S3 directly. The server records the artifact
221
+ path in the run's JSONB ``artifacts`` list. Files larger than
222
+ ~100 MB should ideally be chunked; we send a single PUT here and
223
+ let the underlying TLS connection take care of the streaming.
224
+ """
225
+ import os
226
+ if not os.path.isfile(local_path):
227
+ raise FileNotFoundError(local_path)
228
+ target = artifact_path or os.path.basename(local_path)
229
+ run = _require_active()
230
+ presign = _client().post(
231
+ f"/experiments/{run.experiment_id}/runs/{run.run_id}/artifacts/presigned-url",
232
+ {"path": target, "content_type": "application/octet-stream"},
233
+ )
234
+ upload_url = presign["upload_url"]
235
+
236
+ # We bypass the SDK's auth-bearing Session for the actual PUT —
237
+ # presigned URLs carry their own signature and don't accept the
238
+ # Polnor Bearer header. Use a stripped-down requests call.
239
+ import requests
240
+ with open(local_path, "rb") as f:
241
+ resp = requests.put(upload_url, data=f, timeout=300)
242
+ if resp.status_code >= 400:
243
+ raise RuntimeError(
244
+ f"artifact upload failed ({resp.status_code}): {resp.text[:200]}"
245
+ )
246
+
247
+
248
+ def get_metric_history(key: str) -> list[dict]:
249
+ """Return every (step, value, timestamp) observation for a metric
250
+ on the active run. Useful for building local plots or comparing
251
+ across runs.
252
+ """
253
+ run = _require_active()
254
+ resp = _client().get(
255
+ f"/experiments/{run.experiment_id}/runs/{run.run_id}/metrics/{key}"
256
+ )
257
+ return resp.get("data", [])
258
+
259
+
260
+ @contextlib.contextmanager
261
+ def autolog(framework: str | None = None) -> Iterator[None]:
262
+ """Lightweight autolog: open a run on enter, close on exit.
263
+
264
+ The full MLflow autolog hooks (sklearn estimator inspection, etc.)
265
+ aren't here yet. This context just gives users the convenient
266
+ ``with mlflow.autolog():`` shape that frameworks tend to wrap.
267
+ """
268
+ run = start_run()
269
+ if framework:
270
+ try:
271
+ set_tag("autolog.framework", framework)
272
+ except Exception:
273
+ pass
274
+ try:
275
+ yield
276
+ finally:
277
+ try:
278
+ end_run("completed")
279
+ except Exception:
280
+ # Make sure the stack stays clean even if the network call fails.
281
+ s = _stack()
282
+ if s and s[-1] is run:
283
+ s.pop()
284
+
285
+
286
+ __all__ = [
287
+ "ActiveRun",
288
+ "start_run",
289
+ "end_run",
290
+ "active_run",
291
+ "log_metric",
292
+ "log_metrics",
293
+ "log_param",
294
+ "log_params",
295
+ "set_tag",
296
+ "set_tags",
297
+ "log_artifact",
298
+ "get_metric_history",
299
+ "autolog",
300
+ ]
@@ -0,0 +1,113 @@
1
+ Metadata-Version: 2.4
2
+ Name: polnor
3
+ Version: 0.1.0
4
+ Summary: Polnor Python SDK — query Iceberg lakehouses, manage notebooks/jobs/models, MLflow-compatible tracking.
5
+ Author-email: Polnor <contact@polnor.net>
6
+ License: Apache-2.0
7
+ Project-URL: Homepage, https://polnor.net
8
+ Project-URL: Documentation, https://docs.polnor.net/sdk/python
9
+ Project-URL: Source, https://github.com/polnor/polnor
10
+ Keywords: polnor,lakehouse,iceberg,sql,mlflow,data-platform
11
+ Classifier: Development Status :: 3 - Alpha
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: Programming Language :: Python :: 3
14
+ Classifier: Programming Language :: Python :: 3 :: Only
15
+ Classifier: Programming Language :: Python :: 3.9
16
+ Classifier: Programming Language :: Python :: 3.10
17
+ Classifier: Programming Language :: Python :: 3.11
18
+ Classifier: Programming Language :: Python :: 3.12
19
+ Classifier: License :: OSI Approved :: Apache Software License
20
+ Requires-Python: >=3.9
21
+ Description-Content-Type: text/markdown
22
+ Requires-Dist: requests>=2.28
23
+ Provides-Extra: test
24
+ Requires-Dist: pytest>=7; extra == "test"
25
+ Requires-Dist: responses>=0.24; extra == "test"
26
+
27
+ # Polnor Python SDK
28
+
29
+ MLflow-compatible tracking for jobs running on the Polnor lakehouse.
30
+
31
+ ## Install
32
+
33
+ ```bash
34
+ pip install polnor
35
+ ```
36
+
37
+ Or, from a local checkout (contributors / pre-release):
38
+
39
+ ```bash
40
+ pip install -e ./sdk/python
41
+ ```
42
+
43
+ ## Configure
44
+
45
+ Set three environment variables before importing the SDK:
46
+
47
+ | Variable | Required | Notes |
48
+ |---|---|---|
49
+ | `POLNOR_API_URL` | yes | e.g. `https://api.polnor.net` |
50
+ | `POLNOR_TOKEN` | yes | Personal access token from `/settings/tokens` |
51
+ | `POLNOR_EXPERIMENT_ID` | optional | Default experiment for `start_run()` |
52
+ | `POLNOR_WORKSPACE_SLUG` | optional | When the token isn't workspace-scoped |
53
+
54
+ Inside a Polnor job these are injected automatically.
55
+
56
+ ## Quickstart
57
+
58
+ ```python
59
+ from polnor import mlflow
60
+
61
+ with mlflow.start_run(run_name="train-resnet50") as run:
62
+ mlflow.log_params({"lr": 0.01, "batch_size": 32, "optimizer": "adam"})
63
+
64
+ for epoch in range(10):
65
+ loss, acc = train_one_epoch()
66
+ mlflow.log_metrics({"loss": loss, "accuracy": acc}, step=epoch)
67
+
68
+ mlflow.set_tag("git_commit", os.environ.get("CI_COMMIT_SHA", "dev"))
69
+ ```
70
+
71
+ The run handle is also usable without the `with` block:
72
+
73
+ ```python
74
+ run = mlflow.start_run()
75
+ try:
76
+ mlflow.log_metric("score", 0.93)
77
+ finally:
78
+ mlflow.end_run("completed")
79
+ ```
80
+
81
+ ## What's supported
82
+
83
+ | API | Status |
84
+ |---|---|
85
+ | `start_run` / `end_run` / `active_run` | ✅ |
86
+ | `log_metric` / `log_metrics` (latest value wins) | ✅ |
87
+ | `log_param` / `log_params` | ✅ |
88
+ | `set_tag` / `set_tags` | ✅ |
89
+ | `autolog()` context manager | ✅ |
90
+ | Per-step metric history (time series) | ⏳ planned |
91
+ | `log_artifact` (presigned-URL upload) | ⏳ planned |
92
+ | Sklearn / PyTorch / TF auto-instrumentation | ⏳ planned |
93
+
94
+ ## How it differs from upstream MLflow
95
+
96
+ * Server-side, metrics merge into a single JSONB blob — the latest
97
+ value per key wins. The full `(step, value, timestamp)` history is
98
+ in scope for a follow-up migration.
99
+ * No server-side parameter immutability check yet — calling
100
+ `log_param("lr", 0.02)` after `log_param("lr", 0.01)` overwrites
101
+ silently. Upstream MLflow rejects with 400.
102
+ * No `mlflow.search_runs()` — use the REST API
103
+ (`GET /api/v1/experiments/:id/runs`) directly until we wrap it.
104
+
105
+ ## Development
106
+
107
+ ```bash
108
+ pip install -e ".[test]"
109
+ pytest
110
+ ```
111
+
112
+ Tests use [responses](https://github.com/getsentry/responses) to mock
113
+ HTTP — no live API needed.
@@ -0,0 +1,11 @@
1
+ README.md
2
+ pyproject.toml
3
+ polnor/__init__.py
4
+ polnor/_client.py
5
+ polnor/mlflow.py
6
+ polnor.egg-info/PKG-INFO
7
+ polnor.egg-info/SOURCES.txt
8
+ polnor.egg-info/dependency_links.txt
9
+ polnor.egg-info/requires.txt
10
+ polnor.egg-info/top_level.txt
11
+ tests/test_mlflow.py
@@ -0,0 +1,5 @@
1
+ requests>=2.28
2
+
3
+ [test]
4
+ pytest>=7
5
+ responses>=0.24
@@ -0,0 +1 @@
1
+ polnor
@@ -0,0 +1,38 @@
1
+ [build-system]
2
+ requires = ["setuptools>=68.0", "wheel"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "polnor"
7
+ version = "0.1.0"
8
+ description = "Polnor Python SDK — query Iceberg lakehouses, manage notebooks/jobs/models, MLflow-compatible tracking."
9
+ readme = "README.md"
10
+ requires-python = ">=3.9"
11
+ license = { text = "Apache-2.0" }
12
+ authors = [{ name = "Polnor", email = "contact@polnor.net" }]
13
+ keywords = ["polnor", "lakehouse", "iceberg", "sql", "mlflow", "data-platform"]
14
+ classifiers = [
15
+ "Development Status :: 3 - Alpha",
16
+ "Intended Audience :: Developers",
17
+ "Programming Language :: Python :: 3",
18
+ "Programming Language :: Python :: 3 :: Only",
19
+ "Programming Language :: Python :: 3.9",
20
+ "Programming Language :: Python :: 3.10",
21
+ "Programming Language :: Python :: 3.11",
22
+ "Programming Language :: Python :: 3.12",
23
+ "License :: OSI Approved :: Apache Software License",
24
+ ]
25
+ dependencies = [
26
+ "requests>=2.28",
27
+ ]
28
+
29
+ [project.urls]
30
+ Homepage = "https://polnor.net"
31
+ Documentation = "https://docs.polnor.net/sdk/python"
32
+ Source = "https://github.com/polnor/polnor"
33
+
34
+ [project.optional-dependencies]
35
+ test = ["pytest>=7", "responses>=0.24"]
36
+
37
+ [tool.setuptools]
38
+ packages = ["polnor"]
polnor-0.1.0/setup.cfg ADDED
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
@@ -0,0 +1,169 @@
1
+ """Smoke tests for polnor.mlflow.
2
+
3
+ Mocks the HTTP layer with `responses`. Exercises the contract of every
4
+ public function — happy path + a couple of error paths. Real
5
+ integration is covered by a separate e2e suite that hits the dev API.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ import os
11
+ import threading
12
+
13
+ import pytest
14
+ import responses
15
+
16
+ from polnor import mlflow
17
+ from polnor._client import Client, Config, PolnorAPIError
18
+
19
+
20
+ @pytest.fixture(autouse=True)
21
+ def env(monkeypatch):
22
+ monkeypatch.setenv("POLNOR_API_URL", "https://api.example")
23
+ monkeypatch.setenv("POLNOR_TOKEN", "test-token")
24
+ monkeypatch.setenv("POLNOR_EXPERIMENT_ID", "exp-1")
25
+ # Reset thread-local state explicitly. threading.local.__init__
26
+ # in some interpreter versions doesn't clear attributes set on
27
+ # the instance; deleting the known caches is reliable.
28
+ for attr in ("runs", "client"):
29
+ if hasattr(mlflow._local, attr):
30
+ delattr(mlflow._local, attr)
31
+ _ = threading # silence unused import in some configurations
32
+ yield
33
+
34
+
35
+ def url(path: str) -> str:
36
+ return f"https://api.example/api/v1{path}"
37
+
38
+
39
+ @responses.activate
40
+ def test_start_and_end_run_with_context_manager():
41
+ responses.post(url("/experiments/exp-1/runs"), json={"id": "r1", "name": "auto"})
42
+ end_call = responses.patch(url("/experiments/exp-1/runs/r1"), status=204)
43
+
44
+ with mlflow.start_run() as run:
45
+ assert run.run_id == "r1"
46
+ assert mlflow.active_run() is run
47
+
48
+ assert end_call.call_count == 1
49
+ body = responses.calls[1].request.body
50
+ assert b"completed" in body
51
+
52
+
53
+ @responses.activate
54
+ def test_context_manager_marks_failed_on_exception():
55
+ responses.post(url("/experiments/exp-1/runs"), json={"id": "r2", "name": "x"})
56
+ end_call = responses.patch(url("/experiments/exp-1/runs/r2"), status=204)
57
+
58
+ with pytest.raises(ZeroDivisionError):
59
+ with mlflow.start_run():
60
+ 1 / 0
61
+
62
+ assert end_call.call_count == 1
63
+ assert b"failed" in responses.calls[1].request.body
64
+
65
+
66
+ @responses.activate
67
+ def test_log_metric_and_param_and_tag():
68
+ responses.post(url("/experiments/exp-1/runs"), json={"id": "r3", "name": "x"})
69
+ metric_call = responses.post(url("/experiments/exp-1/runs/r3/metrics"), status=204)
70
+ param_call = responses.post(url("/experiments/exp-1/runs/r3/params"), status=204)
71
+ tag_call = responses.post(url("/experiments/exp-1/runs/r3/tags"), status=204)
72
+ responses.patch(url("/experiments/exp-1/runs/r3"), status=204)
73
+
74
+ with mlflow.start_run():
75
+ mlflow.log_metric("loss", 0.123)
76
+ mlflow.log_param("optimizer", "adam")
77
+ mlflow.set_tag("git_sha", "deadbeef")
78
+
79
+ assert metric_call.call_count == 1
80
+ assert param_call.call_count == 1
81
+ assert tag_call.call_count == 1
82
+
83
+
84
+ @responses.activate
85
+ def test_log_metrics_batch():
86
+ responses.post(url("/experiments/exp-1/runs"), json={"id": "r4", "name": "x"})
87
+ batch = responses.post(url("/experiments/exp-1/runs/r4/metrics/batch"), status=204)
88
+ responses.patch(url("/experiments/exp-1/runs/r4"), status=204)
89
+
90
+ with mlflow.start_run():
91
+ mlflow.log_metrics({"a": 1.0, "b": 2.0})
92
+
93
+ assert batch.call_count == 1
94
+
95
+
96
+ def test_log_without_active_run_raises():
97
+ with pytest.raises(RuntimeError, match="No active run"):
98
+ mlflow.log_metric("x", 1.0)
99
+
100
+
101
+ def test_artifact_missing_file_raises():
102
+ # log_artifact now actually uploads, but should still surface a
103
+ # FileNotFoundError before talking to the server when the local
104
+ # path doesn't exist.
105
+ with pytest.raises(FileNotFoundError):
106
+ mlflow.log_artifact("/tmp/__definitely_not_a_real_file_in_polnor__")
107
+
108
+
109
+ @responses.activate
110
+ def test_log_artifact_happy_path(tmp_path):
111
+ # Server hands out a presigned URL → SDK does PUT to the URL.
112
+ responses.post(url("/experiments/exp-1/runs"), json={"id": "r5", "name": "x"})
113
+ responses.post(
114
+ url("/experiments/exp-1/runs/r5/artifacts/presigned-url"),
115
+ json={"upload_url": "https://s3.example/upload?sig=abc",
116
+ "s3_uri": "s3://b/k", "expires_in": 3600},
117
+ )
118
+ upload = responses.put("https://s3.example/upload", status=200)
119
+ responses.patch(url("/experiments/exp-1/runs/r5"), status=204)
120
+
121
+ f = tmp_path / "model.pkl"
122
+ f.write_bytes(b"\x80\x04model")
123
+ with mlflow.start_run():
124
+ mlflow.log_artifact(str(f), artifact_path="model.pkl")
125
+ assert upload.call_count == 1
126
+
127
+
128
+ @responses.activate
129
+ def test_get_metric_history():
130
+ responses.post(url("/experiments/exp-1/runs"), json={"id": "r6", "name": "x"})
131
+ responses.get(
132
+ url("/experiments/exp-1/runs/r6/metrics/loss"),
133
+ json={"key": "loss", "data": [
134
+ {"step": 0, "value": 1.0, "timestamp": "2026-01-01T00:00:00Z"},
135
+ {"step": 1, "value": 0.5, "timestamp": "2026-01-01T00:01:00Z"},
136
+ ]},
137
+ )
138
+ responses.patch(url("/experiments/exp-1/runs/r6"), status=204)
139
+
140
+ with mlflow.start_run():
141
+ history = mlflow.get_metric_history("loss")
142
+ assert len(history) == 2
143
+ assert history[0]["step"] == 0
144
+
145
+
146
+ @responses.activate
147
+ def test_5xx_retried_then_surfaced():
148
+ # First two attempts 503, third raises; SDK gives up and re-raises.
149
+ cfg = Config(api_url="https://api.example", token="t",
150
+ max_retries=2, backoff_base=0.0, timeout=1)
151
+ client = Client(cfg)
152
+ responses.post("https://api.example/api/v1/x", status=503, body="boom")
153
+
154
+ with pytest.raises(PolnorAPIError) as ei:
155
+ client.post("/x")
156
+ assert ei.value.status == 503
157
+
158
+
159
+ def test_config_from_env_requires_url(monkeypatch):
160
+ monkeypatch.delenv("POLNOR_API_URL", raising=False)
161
+ with pytest.raises(RuntimeError, match="POLNOR_API_URL"):
162
+ Config.from_env()
163
+
164
+
165
+ def test_config_from_env_requires_token(monkeypatch):
166
+ monkeypatch.setenv("POLNOR_API_URL", "https://api.example")
167
+ monkeypatch.delenv("POLNOR_TOKEN", raising=False)
168
+ with pytest.raises(RuntimeError, match="POLNOR_TOKEN"):
169
+ Config.from_env()