periplus-python-sdk 0.6.0__py3-none-any.whl → 0.7.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: periplus-python-sdk
3
- Version: 0.6.0
3
+ Version: 0.7.0
4
4
  Summary: Read-only Python client for the public Periplus query API
5
5
  License-Expression: Apache-2.0
6
6
  Project-URL: Repository, https://github.com/elei-io/periplus
@@ -43,7 +43,7 @@ The client reuses HTTP connections; close it with a context manager or `close()`
43
43
  Install the notebook integration from PyPI:
44
44
 
45
45
  ```sh
46
- uv add "periplus-python-sdk[notebook]>=0.6.0"
46
+ uv add "periplus-python-sdk[notebook]>=0.7.0"
47
47
  ```
48
48
 
49
49
  In a Python setup cell, create a SQLAlchemy engine:
@@ -62,7 +62,7 @@ FROM public_v1.capture
62
62
  LIMIT 10
63
63
  ```
64
64
 
65
- Marimo displays the result as a table. Expand **pp → public_v1** in Data Sources
65
+ Marimo displays the result as a table. Expand **pp → periplus → public_v1** in Data Sources
66
66
  to discover views and expand a view to load its columns for SQL completion.
67
67
  Discovery uses bounded `SHOW TABLES` and `DESCRIBE` through the same public API;
68
68
  no internal catalogue or storage credentials are used. Truncated discovery fails
@@ -82,7 +82,7 @@ captures = mo.sql(
82
82
  ```
83
83
 
84
84
  Set `mode="experimental"` for the experimental service. Omit the URL to use
85
- `PERIPLUS_PUBLIC_URL`. Optional `timeout=140` and `schema_version="public_v1"`
85
+ `PERIPLUS_PUBLIC_URL`. Optional `timeout=620` and `schema_version="public_v1"`
86
86
  arguments configure the client deadline and public schema. Run `pp.dispose()` when finished. This is a read-only
87
87
  SQLAlchemy dialect for textual SQL and reflection, not a writable ORM backend.
88
88
  Each statement has its own server snapshot; SQLAlchemy transaction blocks do not
@@ -90,7 +90,7 @@ provide a shared snapshot or rollback. The adapter makes no transaction requests
90
90
 
91
91
  A complete notebook is in `examples/notebook.py`. The integration is tested with
92
92
  marimo 0.24.1 and SQLAlchemy 2.x. SQLAlchemy is included in the standard SDK install; the `notebook` extra adds
93
- marimo. Existing marimo environments only need `uv add "periplus-python-sdk>=0.6.0"`.
93
+ marimo. Existing marimo environments only need `uv add "periplus-python-sdk>=0.7.0"`.
94
94
  The returned object is a standard SQLAlchemy Engine, also usable with pandas and
95
95
  ordinary Python scripts. Engine creation is lazy; the first query opens a connection.
96
96
 
@@ -112,14 +112,17 @@ with connect("https://periplus.dev", mode="stable") as connection:
112
112
  Connections expose `cursor`, `execute`, `close`, and context managers. Cursors
113
113
  support `execute`, `fetchone`, `fetchmany`, `fetchall`, iteration, and close.
114
114
  Use positional `?` parameters. Decimal and temporal parameters are sent as
115
- strings; use explicit SQL casts. Binary and nested parameters are not supported
116
- by this adapter. Fetching only consumes the bounded result already received;
117
- it never issues pagination or retries. Connections/cursors are not thread-shared.
115
+ strings; use explicit SQL casts. One-dimensional lists/tuples of these scalar values are supported; use an explicit array cast
116
+ for empty or typed lists. Binary, mappings, and nested collection parameters are unsupported. Fetching consumes incremental batches from one HTTP response and one source snapshot;
117
+ it never issues pagination or retries. Close a cursor early to stop delivery. Connections/cursors are not thread-shared.
118
118
  `commit()` is a no-op; `rollback()` and `executemany()` are unsupported.
119
119
 
120
- `cursor.result` preserves the original query response. `connection.last_result`
120
+ `cursor.result` preserves stream metadata and completion information without retaining consumed rows. `connection.last_result`
121
121
  also retains it after marimo closes a cursor; a new execution clears it first.
122
- Truncation emits `periplus_sdk.dbapi.TruncationWarning` and sets `rowcount` to -1.
122
+ Truncation raises `OperationalError` with `code="result_limit"` when fetching reaches the terminal frame.
123
+ `rowcount` stays -1 until completion, then reports the delivered count. `result.complete` means the
124
+ terminal frame arrived; inspect `result.truncated` separately when opting into partial data.
125
+ Use `allow_partial=True` on the engine or connection only when incomplete results are intentional.
123
126
  DB-API failures use the standard exception hierarchy in `periplus_sdk.dbapi`;
124
127
  HTTP errors retain `status_code`, `code`, and `retry_after_seconds`.
125
128
 
@@ -185,7 +188,7 @@ Use `aclose()` when managing an async client's lifetime explicitly.
185
188
  20-second server deadline. Always inspect `truncated`. The SDK does not silently fetch more rows or retry.
186
189
  - `ApiError` exposes `status_code`, safe `code`, and `retry_after_seconds` when supplied.
187
190
  `TransportError` means HTTP failed; `ResponseError` means a malformed successful response.
188
- The client timeout defaults to 140 seconds and can be set with `timeout=`. A timeout or local
191
+ The client timeout defaults to 620 seconds and can be set with `timeout=`. A timeout or local
189
192
  cancellation does not guarantee server cancellation. Redirects are not followed automatically.
190
193
  - Preparation and execution are attributed to `sdk` in the existing private query history.
191
194
  Original SQL and parameters are retained for 30 days; result rows are not stored. Recording is
@@ -196,10 +199,10 @@ Use `aclose()` when managing an async client's lifetime explicitly.
196
199
  Install the public-v1 client from PyPI:
197
200
 
198
201
  ```sh
199
- python -m pip install "periplus-python-sdk>=0.6.0"
202
+ python -m pip install "periplus-python-sdk>=0.7.0"
200
203
  ```
201
204
 
202
- Version 0.6.0 supports the current public-v1 contract. For production, configure
205
+ Version 0.6.1 supports the current public-v1 contract. For production, configure
203
206
  `PERIPLUS_PUBLIC_URL=https://periplus.dev`; no API token is required.
204
207
  Run the installed package against an available public app:
205
208
 
@@ -211,11 +214,11 @@ PERIPLUS_PUBLIC_URL=http://localhost:8080 python packages/periplus-python-sdk/ex
211
214
 
212
215
  Repository CI publishes immutable releases from tags named
213
216
  `periplus-python-sdk-v<version>`. The tag must exactly match the static version
214
- in `pyproject.toml`; for example, version `0.6.0` is released with:
217
+ in `pyproject.toml`; for example, version `0.7.0` is released with:
215
218
 
216
219
  ```sh
217
- git tag periplus-python-sdk-v0.6.0
218
- git push origin periplus-python-sdk-v0.6.0
220
+ git tag periplus-python-sdk-v0.7.0
221
+ git push origin periplus-python-sdk-v0.7.0
219
222
  ```
220
223
 
221
224
  PyPI publishing uses Trusted Publishing rather than a stored API token. The
@@ -225,10 +228,74 @@ that GitHub environment with required reviewers before the first release.
225
228
 
226
229
  ## Public v1
227
230
 
228
- Install the updated SDK from PyPI with `python -m pip install "periplus-python-sdk>=0.6.0"`. The previously published 0.2.0 release predates this contract. `prepare` and `execute` accept keyword-only `schema_version="public_v1"` (the default); responses preserve `schema_version` separately from `source_snapshot`. Unavailable versions are rejected by the server.
231
+ Install the updated SDK from PyPI with `python -m pip install "periplus-python-sdk>=0.7.0"`. The previously published 0.2.0 release predates this contract. `prepare` and `execute` accept keyword-only `schema_version="public_v1"` (the default); responses preserve `schema_version` separately from `source_snapshot`. Unavailable versions are rejected by the server.
229
232
 
230
233
  ## License
231
234
 
232
235
  Copyright (c) 2026 Ekku Leivonen (elei.io). Licensed under [Apache-2.0](LICENSE);
233
236
  see [NOTICE](NOTICE). The server and other repository packages have separate
234
237
  licensing described in the root LICENSING.md.
238
+
239
+ ## Python filtering back into marimo SQL
240
+
241
+ The existing engine streams by default. For SQL → Python filtering → SQL, bind the
242
+ filtered IDs to another engine variable; it shares the original connection pool and
243
+ creates no server-side table or upload. Marimo discovers it like any other SQLAlchemy engine.
244
+
245
+ ```python
246
+ pp = sql_api.create_engine("https://periplus.dev")
247
+
248
+ # In a Python cell, after filtering your initial SQL dataframe:
249
+ selected = sql_api.bind(pp, content_ids=qualified_pages["content_id"].to_list())
250
+
251
+ # In the next SQL cell (engine: selected):
252
+ raw_elements = mo.sql(
253
+ """
254
+ SELECT content_id, node_index, parent_index, tag,
255
+ trim(text_direct) AS text, attributes['class'] AS elem_class
256
+ FROM public_v1.html_element
257
+ WHERE content_id IN (SELECT unnest(CAST(:content_ids AS VARCHAR[])))
258
+ AND text_direct IS NOT NULL AND trim(text_direct) <> ''
259
+ """,
260
+ engine=selected,
261
+ )
262
+ ```
263
+
264
+ Keep SQL strings plain: `:content_ids` is a bound parameter, not an f-string.
265
+ An empty list selects no elements; duplicate IDs do not multiply rows. Re-running the
266
+ binding cell snapshots the new Python list into its engine options. Bindings only apply
267
+ to matching SQL placeholders, so catalogue discovery remains available. Explicit
268
+ SQLAlchemy execution parameters take precedence over engine bindings.
269
+
270
+ Marimo collects the stream into a local dataframe, which must fit notebook memory.
271
+ Its dataframe conversion currently infers types and may lose empty/all-null column
272
+ information; the DB-API cursor always exposes server column names and SQL types.
273
+ Initial SQL results and subsequent element results both reject truncation by default.
274
+ Each separate SQL execution has its own snapshot; Python filtering does not pin the
275
+ first query's snapshot across the round trip.
276
+
277
+ ## Consume batches without a dataframe
278
+
279
+ ```python
280
+ with Client("https://periplus.dev") as client:
281
+ with client.stream("SELECT content_id FROM public_v1.prose") as stream:
282
+ for rows in stream:
283
+ process_batch(rows)
284
+ print(stream.result.row_count, stream.result.source_snapshot)
285
+ print(stream.result.limits)
286
+ ```
287
+
288
+ Batches retain JSON wire values with SQL types in `stream.result.types`. DB-API fetching
289
+ performs scalar Python decoding. The synchronous `Client.stream()` raises on truncation
290
+ unless `allow_partial=True`; missing completion, malformed frames, timeouts, and broken
291
+ connections always fail. Do not treat already-consumed batches as a complete dataset
292
+ until iteration finishes successfully. Use context managers when stopping early.
293
+
294
+ Admin Query limits controls requests (default 4 MiB, ceiling 16 MiB), parameter values
295
+ (default 100,000, ceiling 1,000,000), results (ceiling 10 million rows and 1 GiB), and
296
+ execution duration (ceiling 600 seconds). Parameter counting includes collection containers
297
+ and their values. Existing result/duration defaults remain 1,000 rows, 8 MiB, and 20 seconds.
298
+ Operators must raise these settings for larger workloads. No SDK parameter grants a larger
299
+ server budget. Stream metadata includes effective limits; input rejections name the budget.
300
+ Deploy the API, query service, public gateway and SDK together after applying Alembic
301
+ revision `20260911_0016`. Ingress must permit the configured request size and stream duration.
@@ -0,0 +1,16 @@
1
+ periplus_python_sdk-0.7.0.dist-info/licenses/LICENSE,sha256=z8d0m5b2O9McPEK1xHG_dWgUBT6EfBDz6wA0F7xSPTA,11358
2
+ periplus_python_sdk-0.7.0.dist-info/licenses/NOTICE,sha256=bhbYSqcUB3U_P1-XzloiT81JGniqoYaRLxNkQ1Pm9MQ,52
3
+ periplus_sdk/__init__.py,sha256=WimXYlPB6tCimBO4VSwhcp00dwSL87jMmMuQ4-kINfM,546
4
+ periplus_sdk/client.py,sha256=bqIq2bmM_bnwf0ZnMN5kTXTctnEz7ITJUAfZohGsRIQ,7544
5
+ periplus_sdk/dbapi.py,sha256=wyhtLfpgBZGxhC9rA53DNfQ7j8HegZkVp7mRq98bFSU,12256
6
+ periplus_sdk/errors.py,sha256=rB1n-v8Hc2tu2dtHivz-MlqsCoRC5pTcWogTfM7SMLw,855
7
+ periplus_sdk/py.typed,sha256=AbpHGcgLb-kRsJGnwFEktk7uzpZOCcBY74-YBdrKVGs,1
8
+ periplus_sdk/sql_api.py,sha256=o5kLzwLim3VQ5oZfRyPUGCDr7RWmDidtktGLyBdkjQk,1884
9
+ periplus_sdk/sqlalchemy.py,sha256=skyN2R5nIiWK55uBuxOHjeGRAPb8uPxm0f8nU0IIPqw,4836
10
+ periplus_sdk/stream.py,sha256=pPXWikBWnOUZuHfBRXx08swBDMWJ4Gayx6azaz_wUCU,5197
11
+ periplus_sdk/types.py,sha256=0eCwWev-Vrk0mRCIkot3Z_-dWWL3gpjljV4GyXEf8lk,1191
12
+ periplus_python_sdk-0.7.0.dist-info/METADATA,sha256=6HTXHT4WnFKYbfB6MVc1Duxni8niJ3JQ7hP7w2GYkiA,14052
13
+ periplus_python_sdk-0.7.0.dist-info/WHEEL,sha256=YVMoNqKzERt-wjUZwJ33xBGAwnFl-4cqbYkTtWa4itE,91
14
+ periplus_python_sdk-0.7.0.dist-info/entry_points.txt,sha256=Pr14L_7AhLinq-4qDxB4awFVubrvR1BEfkaFVRpdCbU,73
15
+ periplus_python_sdk-0.7.0.dist-info/top_level.txt,sha256=o41t5TzwgoxzSmbKoP6olWW1FyEAGWVYTjeoKadBK40,13
16
+ periplus_python_sdk-0.7.0.dist-info/RECORD,,
periplus_sdk/client.py CHANGED
@@ -79,7 +79,7 @@ def _payload(sql: str, parameters: Sequence[JsonValue] | None, schema_version: s
79
79
  class Client:
80
80
  """Reusable synchronous public query client. Close it or use a with block."""
81
81
 
82
- def __init__(self, base_url: str | None = None, *, timeout: float = 140, mode: Literal["stable", "experimental"] = "stable"):
82
+ def __init__(self, base_url: str | None = None, *, timeout: float = 620, mode: Literal["stable", "experimental"] = "stable"):
83
83
  if mode not in {"stable", "experimental"}:
84
84
  raise ConfigurationError("mode must be stable or experimental.")
85
85
  self._query_path = "api/query/experimental/" if mode == "experimental" else "api/query/"
@@ -107,6 +107,25 @@ class Client:
107
107
  def execute(self, sql: str, parameters: Sequence[JsonValue] | None = None, *, schema_version: str = "public_v1") -> QueryResult:
108
108
  return self._request("POST", "exec", QueryResult, json=_payload(sql, parameters, schema_version))
109
109
 
110
+ def stream(self, sql: str, parameters: Sequence[JsonValue] | None = None, *,
111
+ schema_version: str = "public_v1", allow_partial: bool = False):
112
+ """Stream batches from one snapshot; use as a context manager for early exit."""
113
+ from .stream import MEDIA_TYPE, QueryStream
114
+ try:
115
+ request = self._http.build_request("POST", self._query_path + "exec",
116
+ headers={"accept": MEDIA_TYPE}, json=_payload(sql, parameters, schema_version))
117
+ response = self._http.send(request, stream=True)
118
+ try:
119
+ if not response.is_success:
120
+ response.read()
121
+ _decode(response, QueryResult)
122
+ return QueryStream(response, allow_partial=allow_partial)
123
+ except BaseException:
124
+ response.close()
125
+ raise
126
+ except httpx.RequestError:
127
+ raise TransportError("Could not open the public query stream.") from None
128
+
110
129
  def helpers(self) -> QueryHelpers:
111
130
  return self._request("GET", "helpers", QueryHelpers)
112
131
 
@@ -114,7 +133,7 @@ class Client:
114
133
  class AsyncClient:
115
134
  """Reusable asynchronous public query client. Use an async with block."""
116
135
 
117
- def __init__(self, base_url: str | None = None, *, timeout: float = 140, mode: Literal["stable", "experimental"] = "stable"):
136
+ def __init__(self, base_url: str | None = None, *, timeout: float = 620, mode: Literal["stable", "experimental"] = "stable"):
118
137
  if mode not in {"stable", "experimental"}:
119
138
  raise ConfigurationError("mode must be stable or experimental.")
120
139
  self._query_path = "api/query/experimental/" if mode == "experimental" else "api/query/"
periplus_sdk/dbapi.py CHANGED
@@ -1,7 +1,7 @@
1
1
  """Read-only DB-API 2.0 connection over the public query API.
2
2
 
3
- Each execute is an independent server snapshot. Fetching consumes a bounded local
4
- result, never a remote cursor. Connections and cursors must not be shared by threads.
3
+ Each execute is an independent server snapshot. Fetching consumes a bounded stream.
4
+ Connections and cursors must not be shared by threads.
5
5
  """
6
6
  from __future__ import annotations
7
7
 
@@ -11,11 +11,10 @@ from collections.abc import Sequence
11
11
  from datetime import date, datetime, time
12
12
  from decimal import Decimal
13
13
  from typing import Any, Literal
14
- import warnings
15
14
 
16
15
  from .client import Client
17
16
  from .errors import ApiError, ConfigurationError, PeriplusError, ResponseError, TransportError
18
- from .types import QueryResult
17
+ from .stream import StreamResult
19
18
 
20
19
  apilevel = "2.0"
21
20
  threadsafety = 1
@@ -26,10 +25,6 @@ class Warning(builtins.Warning):
26
25
  """DB-API warning."""
27
26
 
28
27
 
29
- class TruncationWarning(Warning):
30
- """The server returned only part of the query result."""
31
-
32
-
33
28
  class Error(PeriplusError):
34
29
  """Base DB-API error; API failures preserve their safe error attributes."""
35
30
 
@@ -147,7 +142,11 @@ def _parameter(value: Any) -> Any:
147
142
  return value
148
143
  if isinstance(value, (Decimal, date, time)):
149
144
  return str(value) if isinstance(value, Decimal) else value.isoformat()
150
- raise ProgrammingError("Parameters must be scalar JSON values, Decimal, date, time or datetime; use explicit SQL casts for typed strings.")
145
+ if isinstance(value, (list, tuple)):
146
+ if any(isinstance(item, (list, tuple, dict)) for item in value):
147
+ raise ProgrammingError("Collection parameters must be one-dimensional lists of scalar values.")
148
+ return [_parameter(item) for item in value]
149
+ raise ProgrammingError("Parameters must be scalar values or scalar lists; use explicit SQL casts for typed values.")
151
150
 
152
151
 
153
152
  class Connection:
@@ -155,16 +154,18 @@ class Connection:
155
154
 
156
155
  dialect = "duckdb"
157
156
 
158
- def __init__(self, base_url: str | None = None, *, timeout: float = 140,
157
+ def __init__(self, base_url: str | None = None, *, timeout: float = 620,
159
158
  mode: Literal["stable", "experimental"] = "stable",
160
- schema_version: str = "public_v1"):
159
+ schema_version: str = "public_v1", allow_partial: bool = False):
161
160
  try:
162
161
  self._client = Client(base_url, timeout=timeout, mode=mode)
163
162
  except ConfigurationError as exc:
164
163
  raise InterfaceError(str(exc)) from exc
165
164
  self.schema_version = schema_version
165
+ self.allow_partial = allow_partial
166
+ self._cursors = set()
166
167
  self.closed = False
167
- self.last_result: QueryResult | None = None
168
+ self.last_result: StreamResult | None = None
168
169
 
169
170
  def _check(self) -> None:
170
171
  if self.closed:
@@ -172,7 +173,9 @@ class Connection:
172
173
 
173
174
  def cursor(self) -> Cursor:
174
175
  self._check()
175
- return Cursor(self)
176
+ cursor = Cursor(self)
177
+ self._cursors.add(cursor)
178
+ return cursor
176
179
 
177
180
  def execute(self, operation: str, parameters: Sequence[Any] | None = None) -> Cursor:
178
181
  cursor = self.cursor()
@@ -192,6 +195,8 @@ class Connection:
192
195
 
193
196
  def close(self) -> None:
194
197
  if not self.closed:
198
+ for cursor in list(self._cursors):
199
+ cursor.close()
195
200
  self._client.close()
196
201
  self.closed = True
197
202
  self.last_result = None
@@ -204,25 +209,26 @@ class Connection:
204
209
  self.close()
205
210
 
206
211
 
207
- def connect(base_url: str | None = None, *, timeout: float = 140,
212
+ def connect(base_url: str | None = None, *, timeout: float = 620,
208
213
  mode: Literal["stable", "experimental"] = "stable",
209
- schema_version: str = "public_v1") -> Connection:
210
- return Connection(base_url, timeout=timeout, mode=mode, schema_version=schema_version)
214
+ schema_version: str = "public_v1", allow_partial: bool = False) -> Connection:
215
+ return Connection(base_url, timeout=timeout, mode=mode, schema_version=schema_version, allow_partial=allow_partial)
211
216
 
212
217
 
213
218
  class Cursor:
214
- """A buffered result. Metadata stays on result after rows are consumed."""
219
+ """Incremental cursor. Metadata stays available without retaining consumed rows."""
215
220
 
216
221
  arraysize = 1
217
222
 
218
223
  def __init__(self, connection: Connection):
219
224
  self.connection = connection
220
225
  self.closed = False
221
- self.result: QueryResult | None = None
226
+ self.result: StreamResult | None = None
222
227
  self.description: list[tuple[Any, ...]] | None = None
223
228
  self.rowcount = -1
224
229
  self._rows: list[tuple[Any, ...]] = []
225
230
  self._position = 0
231
+ self._stream = None
226
232
 
227
233
  def _check(self, *, result: bool = False) -> None:
228
234
  self.connection._check()
@@ -233,6 +239,9 @@ class Cursor:
233
239
 
234
240
  def execute(self, operation: str, parameters: Sequence[Any] | None = None) -> Cursor:
235
241
  self._check()
242
+ if self._stream is not None:
243
+ self._stream.close()
244
+ self._stream = None
236
245
  self.result, self.description, self.rowcount = None, None, -1
237
246
  self._rows, self._position = [], 0
238
247
  self.connection.last_result = None
@@ -242,7 +251,9 @@ class Cursor:
242
251
  raise ProgrammingError("Use a positional parameter sequence with ? placeholders.")
243
252
  values = [_parameter(v) for v in parameters] if parameters is not None else []
244
253
  try:
245
- result = self.connection._client.execute(operation, values, schema_version=self.connection.schema_version)
254
+ self._stream = self.connection._client.stream(operation, values, schema_version=self.connection.schema_version,
255
+ allow_partial=self.connection.allow_partial)
256
+ result = self._stream.result
246
257
  except ApiError as exc:
247
258
  error = ProgrammingError if exc.code == "sql_invalid" else OperationalError
248
259
  raise error(str(exc), status_code=exc.status_code, code=exc.code,
@@ -251,26 +262,37 @@ class Cursor:
251
262
  raise OperationalError(str(exc)) from exc
252
263
  except ResponseError as exc:
253
264
  raise InterfaceError(str(exc)) from exc
254
- if len(result.columns) != len(result.types) or any(len(r) != len(result.columns) for r in result.rows):
255
- raise InterfaceError("Query columns, types and rows have inconsistent widths.")
256
- try:
257
- rows = [tuple(_value(v, t) for v, t in zip(row, result.types, strict=True)) for row in result.rows]
258
- except (ValueError, TypeError, ArithmeticError) as exc:
259
- raise DataError("Query value does not match its SQL type.") from exc
260
265
  self.result = self.connection.last_result = result
261
266
  self.description = [(name, kind, None, None, None, None, None)
262
267
  for name, kind in zip(result.columns, result.types, strict=True)]
263
- self._rows = rows
264
- self.rowcount = -1 if result.truncated else len(rows)
265
- if result.truncated:
266
- warnings.warn(f"Periplus returned a truncated result ({len(rows)} rows); inspect connection.last_result or cursor.result. Fetching does not retrieve additional rows.",
267
- TruncationWarning, stacklevel=2)
268
268
  return self
269
269
 
270
+ def _batch(self):
271
+ try:
272
+ batch = next(self._stream)
273
+ self._rows = [tuple(_value(v, t) for v, t in zip(row, self.result.types, strict=True)) for row in batch]
274
+ self._position = 0
275
+ return True
276
+ except StopIteration:
277
+ self.rowcount = -1 if self.result.truncated else self.result.row_count
278
+ self._rows, self._position = [], 0
279
+ return False
280
+ except ApiError as exc:
281
+ raise OperationalError(str(exc), status_code=exc.status_code, code=exc.code,
282
+ retry_after_seconds=exc.retry_after_seconds) from exc
283
+ except TransportError as exc:
284
+ raise OperationalError(str(exc)) from exc
285
+ except ResponseError as exc:
286
+ raise InterfaceError(str(exc)) from exc
287
+ except (ValueError, TypeError, ArithmeticError) as exc:
288
+ self._stream.close()
289
+ raise DataError("Query value does not match its SQL type.") from exc
290
+
270
291
  def fetchone(self) -> tuple[Any, ...] | None:
271
292
  self._check(result=True)
272
- if self._position == len(self._rows):
273
- return None
293
+ while self._position == len(self._rows):
294
+ if not self._batch():
295
+ return None
274
296
  row = self._rows[self._position]
275
297
  self._position += 1
276
298
  return row
@@ -280,14 +302,17 @@ class Cursor:
280
302
  size = self.arraysize if size is None else size
281
303
  if not isinstance(size, int) or size < 0:
282
304
  raise ProgrammingError("Fetch size must be a non-negative integer.")
283
- end = min(self._position + size, len(self._rows))
284
- rows = self._rows[self._position:end]
285
- self._position = end
305
+ rows = []
306
+ for _ in range(size):
307
+ row = self.fetchone()
308
+ if row is None:
309
+ break
310
+ rows.append(row)
286
311
  return rows
287
312
 
288
313
  def fetchall(self) -> list[tuple[Any, ...]]:
289
314
  self._check(result=True)
290
- return self.fetchmany(len(self._rows) - self._position)
315
+ return list(self)
291
316
 
292
317
  def executemany(self, operation: str, seq_of_parameters: Any) -> None:
293
318
  self._check()
@@ -300,6 +325,9 @@ class Cursor:
300
325
  self._check()
301
326
 
302
327
  def close(self) -> None:
328
+ if self._stream is not None:
329
+ self._stream.close()
330
+ self.connection._cursors.discard(self)
303
331
  self.closed = True
304
332
  self._rows = []
305
333
  self.result = None
periplus_sdk/sql_api.py CHANGED
@@ -3,7 +3,7 @@ from __future__ import annotations
3
3
 
4
4
  from typing import Literal
5
5
 
6
- from sqlalchemy import create_engine as _create_engine
6
+ from sqlalchemy import create_engine as _create_engine, event
7
7
  from sqlalchemy.engine import Engine, URL
8
8
 
9
9
 
@@ -11,15 +11,37 @@ def create_engine(
11
11
  base_url: str | None = None,
12
12
  *,
13
13
  mode: Literal["stable", "experimental"] = "stable",
14
- timeout: float = 140,
14
+ timeout: float = 620,
15
15
  schema_version: str = "public_v1",
16
+ allow_partial: bool = False,
16
17
  ) -> Engine:
17
18
  """Create a SQLAlchemy engine recognized by marimo and other SQL tools.
18
19
 
19
20
  The public URL defaults to PERIPLUS_PUBLIC_URL. Connections are opened lazily;
20
21
  dispose the engine when finished. Each query uses an independent server snapshot.
21
22
  """
22
- return _create_engine(
23
- URL.create("periplus", database=schema_version),
24
- connect_args={"base_url": base_url, "mode": mode, "timeout": timeout},
23
+ engine = _create_engine(
24
+ URL.create("periplus", database="periplus"),
25
+ connect_args={"base_url": base_url, "mode": mode, "timeout": timeout, "schema_version": schema_version, "allow_partial": allow_partial},
25
26
  )
27
+
28
+ event.listen(engine, "before_execute", _parameters, retval=True)
29
+ return engine
30
+
31
+
32
+ def _parameters(connection, clauseelement, multiparams, params, execution_options):
33
+ bindings = execution_options.get("periplus_parameters", {})
34
+ if bindings and not multiparams:
35
+ names = clauseelement.compile().params
36
+ params = {**{name: value for name, value in bindings.items() if name in names}, **params}
37
+ return clauseelement, multiparams, params
38
+
39
+
40
+ def bind(engine: Engine, **parameters) -> Engine:
41
+ """Create a marimo-discoverable engine with named SQL parameters for this cell.
42
+
43
+ Uses the original engine's pool. No query, upload, or server state is created.
44
+ Write :name placeholders in SQL cells; lists can be CAST(:ids AS VARCHAR[]).
45
+ """
46
+ from .dbapi import _parameter
47
+ return engine.execution_options(periplus_parameters={name: _parameter(value) for name, value in parameters.items()})
@@ -60,10 +60,10 @@ class PeriplusDialect(default.DefaultDialect):
60
60
 
61
61
  def create_connect_args(self, url):
62
62
  if url.username or url.password or url.host or url.port:
63
- raise exc.ArgumentError("Use periplus:///public_v1 with base_url and mode in connect_args.")
63
+ raise exc.ArgumentError("Use periplus:///periplus with base_url and mode in connect_args.")
64
64
  if url.query:
65
65
  raise exc.ArgumentError("Pass connection options in connect_args, not URL query parameters.")
66
- return [], {"schema_version": url.database or "public_v1"}
66
+ return [], {}
67
67
 
68
68
  def initialize(self, connection):
69
69
  self.default_schema_name = connection.connection.dbapi_connection.schema_version
@@ -91,11 +91,11 @@ class PeriplusDialect(default.DefaultDialect):
91
91
 
92
92
  def _complete(self, result):
93
93
  raw = result.cursor.result
94
- if raw.truncated:
95
- result.close()
96
- raise exc.InvalidRequestError("Catalogue discovery was truncated by public query limits; refusing an incomplete schema.")
97
94
  try:
98
- return result.fetchall()
95
+ rows = result.fetchall()
96
+ if raw.truncated:
97
+ raise exc.InvalidRequestError("Catalogue discovery was truncated by public query limits; refusing an incomplete schema.")
98
+ return rows
99
99
  finally:
100
100
  result.close()
101
101
 
periplus_sdk/stream.py ADDED
@@ -0,0 +1,134 @@
1
+ """Incremental query frames. EOF is never a successful completion marker."""
2
+ from __future__ import annotations
3
+
4
+ import json
5
+ from typing import Literal
6
+
7
+ import httpx
8
+ from pydantic import BaseModel, Field, JsonValue, ValidationError
9
+
10
+ from .errors import ApiError, ResponseError, TransportError
11
+ from .types import PreparedQuery
12
+
13
+ MEDIA_TYPE = "application/x-ndjson"
14
+
15
+
16
+ class StreamResult(PreparedQuery):
17
+ columns: list[str]
18
+ types: list[str]
19
+ source_snapshot: int = Field(ge=0)
20
+ limits: dict[str, int]
21
+ complete: bool = False
22
+ truncated: bool = False
23
+ truncation_reason: Literal["max_rows", "max_result_bytes"] | None = None
24
+ row_count: int = 0
25
+ result_bytes: int = 0
26
+ elapsed_ms: float = 0
27
+
28
+
29
+ class Rows(BaseModel):
30
+ type: Literal["rows"]
31
+ rows: list[list[JsonValue]]
32
+
33
+
34
+ class Completion(BaseModel):
35
+ type: Literal["complete"]
36
+ row_count: int = Field(ge=0)
37
+ result_bytes: int = Field(ge=0)
38
+ truncated: bool
39
+ truncation_reason: Literal["max_rows", "max_result_bytes"] | None
40
+ elapsed_ms: float = Field(ge=0)
41
+
42
+
43
+ class QueryStream:
44
+ """One response, consumed a batch at a time. Close early to cancel delivery."""
45
+
46
+ def __init__(self, response: httpx.Response, *, allow_partial: bool = False):
47
+ self.response = response
48
+ self.allow_partial = allow_partial
49
+ self.closed = False
50
+ self._count = 0
51
+ self._lines = response.iter_lines()
52
+ try:
53
+ if response.headers.get("content-type", "").split(";")[0] != MEDIA_TYPE:
54
+ raise ResponseError("Expected a streaming query response.")
55
+ frame = self._frame()
56
+ if frame.get("type") != "metadata":
57
+ raise ResponseError("Query stream is missing metadata.")
58
+ self.result = StreamResult.model_validate(frame)
59
+ if len(self.result.columns) != len(self.result.types):
60
+ raise ResponseError("Query columns and types have inconsistent widths.")
61
+ except ValidationError:
62
+ self.close()
63
+ raise ResponseError("Invalid query stream metadata.") from None
64
+ except BaseException:
65
+ self.close()
66
+ raise
67
+
68
+ def _frame(self):
69
+ try:
70
+ for line in self._lines:
71
+ if not line.strip():
72
+ continue
73
+ frame = json.loads(line)
74
+ if not isinstance(frame, dict):
75
+ raise ResponseError("Invalid query stream frame.")
76
+ if frame.get("type") == "error":
77
+ raise ApiError(frame.get("detail", "Query stream failed."),
78
+ status_code=frame.get("status", 500), code=frame.get("code"))
79
+ return frame
80
+ except httpx.RequestError:
81
+ raise TransportError("Query stream interrupted; the result is incomplete.") from None
82
+ except ValueError:
83
+ raise ResponseError("Invalid query stream frame.") from None
84
+ raise ResponseError("Query stream ended without completion; the result is incomplete.")
85
+
86
+ def __iter__(self):
87
+ return self
88
+
89
+ def __next__(self) -> list[list[JsonValue]]:
90
+ if self.closed:
91
+ if not self.result.complete:
92
+ raise ResponseError("Query stream closed before completion; the result is incomplete.")
93
+ if self.result.truncated and not self.allow_partial:
94
+ raise ResponseError("Query result exceeded its budget and is incomplete.")
95
+ raise StopIteration
96
+ try:
97
+ frame = self._frame()
98
+ if frame.get("type") == "rows":
99
+ rows = Rows.model_validate(frame).rows
100
+ if any(len(row) != len(self.result.columns) for row in rows):
101
+ raise ResponseError("Query row has an inconsistent width.")
102
+ self._count += len(rows)
103
+ return rows
104
+ completion = Completion.model_validate(frame)
105
+ if completion.row_count != self._count:
106
+ raise ResponseError("Query completion row count does not match delivered rows.")
107
+ for name, value in completion.model_dump(exclude={"type"}).items():
108
+ setattr(self.result, name, value)
109
+ self.result.complete = True
110
+ self.close()
111
+ if completion.truncated and not self.allow_partial:
112
+ budget = completion.truncation_reason
113
+ limit = self.result.limits.get(budget, "unknown")
114
+ raise ApiError(f"Query result is incomplete: {budget} ({limit}) reached after {self._count} rows. "
115
+ "Narrow the query or ask the administrator to increase this budget.",
116
+ status_code=422, code="result_limit")
117
+ raise StopIteration
118
+ except ValidationError:
119
+ self.close()
120
+ raise ResponseError("Invalid query stream frame.") from None
121
+ except BaseException:
122
+ self.close()
123
+ raise
124
+
125
+ def close(self):
126
+ if not self.closed:
127
+ self.closed = True
128
+ self.response.close()
129
+
130
+ def __enter__(self):
131
+ return self
132
+
133
+ def __exit__(self, *args):
134
+ self.close()
periplus_sdk/types.py CHANGED
@@ -29,6 +29,9 @@ class QueryResult(PreparedQuery):
29
29
  truncated: bool
30
30
  elapsed_ms: float
31
31
  source_snapshot: int = Field(ge=0)
32
+ row_count: int = Field(ge=0)
33
+ result_bytes: int = Field(ge=0)
34
+ truncation_reason: Literal["max_rows", "max_result_bytes"] | None = None
32
35
 
33
36
 
34
37
  class HelperField(BaseModel):
@@ -1,15 +0,0 @@
1
- periplus_python_sdk-0.6.0.dist-info/licenses/LICENSE,sha256=z8d0m5b2O9McPEK1xHG_dWgUBT6EfBDz6wA0F7xSPTA,11358
2
- periplus_python_sdk-0.6.0.dist-info/licenses/NOTICE,sha256=bhbYSqcUB3U_P1-XzloiT81JGniqoYaRLxNkQ1Pm9MQ,52
3
- periplus_sdk/__init__.py,sha256=WimXYlPB6tCimBO4VSwhcp00dwSL87jMmMuQ4-kINfM,546
4
- periplus_sdk/client.py,sha256=ZwMmVJKNF-FrGh_qwiQ5myhu8XJodb_96ohkqK47yDA,6557
5
- periplus_sdk/dbapi.py,sha256=lK9ZROaMKXm26wvC7DCKywm3qwSvqeHa7zkktnJVP80,11244
6
- periplus_sdk/errors.py,sha256=rB1n-v8Hc2tu2dtHivz-MlqsCoRC5pTcWogTfM7SMLw,855
7
- periplus_sdk/py.typed,sha256=AbpHGcgLb-kRsJGnwFEktk7uzpZOCcBY74-YBdrKVGs,1
8
- periplus_sdk/sql_api.py,sha256=iBJt6hK_m84bDMBRZUDNwz20gi07gAxqWpm-NWWsqSo,850
9
- periplus_sdk/sqlalchemy.py,sha256=axFpEMnMP9uSxbsxzVNSUHNh41q2NruSX8W_5E8F59Q,4877
10
- periplus_sdk/types.py,sha256=PTdJO6BYTY97dBd3HMEZ52k6P9s0cMvidjpnc8zUpy4,1045
11
- periplus_python_sdk-0.6.0.dist-info/METADATA,sha256=UBcdaDmCCo0O0ISVlljy947LFFQIawX36y4dOioN4mo,10199
12
- periplus_python_sdk-0.6.0.dist-info/WHEEL,sha256=YVMoNqKzERt-wjUZwJ33xBGAwnFl-4cqbYkTtWa4itE,91
13
- periplus_python_sdk-0.6.0.dist-info/entry_points.txt,sha256=Pr14L_7AhLinq-4qDxB4awFVubrvR1BEfkaFVRpdCbU,73
14
- periplus_python_sdk-0.6.0.dist-info/top_level.txt,sha256=o41t5TzwgoxzSmbKoP6olWW1FyEAGWVYTjeoKadBK40,13
15
- periplus_python_sdk-0.6.0.dist-info/RECORD,,