periplus-python-sdk 0.7.0__py3-none-any.whl → 0.8.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,70 @@
1
+ Metadata-Version: 2.4
2
+ Name: periplus-python-sdk
3
+ Version: 0.8.0
4
+ Summary: Read-only Python client for the public Periplus query API
5
+ License-Expression: Apache-2.0
6
+ Project-URL: Repository, https://github.com/elei-io/periplus
7
+ Project-URL: Issues, https://github.com/elei-io/periplus/issues
8
+ Requires-Python: >=3.11
9
+ Description-Content-Type: text/markdown
10
+ License-File: LICENSE
11
+ License-File: NOTICE
12
+ Requires-Dist: httpx>=0.28
13
+ Requires-Dist: pydantic<3,>=2.12
14
+ Requires-Dist: sqlalchemy<3,>=2.0
15
+ Provides-Extra: notebook
16
+ Requires-Dist: marimo[sql]>=0.24.1; extra == "notebook"
17
+ Dynamic: license-file
18
+
19
+ # Periplus Python SDK
20
+
21
+ Read-only access to the public ClickHouse corpus through the Periplus HTTP API.
22
+ Install the SDK from the same checkout as your deployment:
23
+
24
+ ```sh
25
+ python -m pip install ./packages/periplus-python-sdk
26
+ ```
27
+
28
+ ```python
29
+ from periplus_sdk import Client
30
+
31
+ with Client("http://localhost:8080") as client:
32
+ result = client.execute(
33
+ "SELECT capture_id, url FROM public_v1.captures LIMIT ?", [10]
34
+ )
35
+ print(result.columns, result.rows)
36
+ ```
37
+
38
+ Use `PERIPLUS_PUBLIC_URL` to omit the URL argument. No database credentials or
39
+ service token are needed for the public gateway. `AsyncClient` provides async
40
+ methods. `prepare` validates/explains SQL; `execute` returns typed columns, rows,
41
+ truncation and query metadata; `helpers` describes the installed public views.
42
+
43
+ The public schema is `public_v1`. HTML joins use `document_id` plus node index;
44
+ `document_id` identifies retained bytes and their HTML interpretation. Raw-byte hashes stay internal.
45
+ There is one query endpoint, with no experimental fallback. Client errors preserve
46
+ server categories and do not automatically retry executed queries.
47
+
48
+ For notebook/SQLAlchemy integration:
49
+
50
+ ```python
51
+ from periplus_sdk import sql_api
52
+ from sqlalchemy import text
53
+
54
+ engine = sql_api.create_engine(base_url="http://localhost:8080")
55
+ with engine.connect() as connection:
56
+ print(connection.execute(text("SELECT url FROM public_v1.captures LIMIT 5")).all())
57
+ engine.dispose()
58
+ ```
59
+
60
+ Marimo discovers the five public views through the helper catalogue. Column
61
+ reflection uses `SELECT * ... LIMIT 0` to retrieve native types without reading
62
+ corpus rows; no `SHOW`, `DESCRIBE`, or system-table access is required.
63
+
64
+ The DB-API connection advertises the ClickHouse dialect and converts native
65
+ nullable integer, decimal, date and datetime types. Nested types retain JSON wire
66
+ values. It is read-only: there are no client transactions or writable sessions.
67
+ Streaming cursors expose incomplete/truncated results explicitly; configure
68
+ `allow_partial` only when partial results suit the application.
69
+
70
+ See [the public schema](../../docs/SCHEMA.md) and [query boundary](../../docs/QUERY.md).
@@ -0,0 +1,16 @@
1
+ periplus_python_sdk-0.8.0.dist-info/licenses/LICENSE,sha256=z8d0m5b2O9McPEK1xHG_dWgUBT6EfBDz6wA0F7xSPTA,11358
2
+ periplus_python_sdk-0.8.0.dist-info/licenses/NOTICE,sha256=bhbYSqcUB3U_P1-XzloiT81JGniqoYaRLxNkQ1Pm9MQ,52
3
+ periplus_sdk/__init__.py,sha256=WimXYlPB6tCimBO4VSwhcp00dwSL87jMmMuQ4-kINfM,546
4
+ periplus_sdk/client.py,sha256=trcsOL4hnpDByMr8wZ9iz2vzxQWCW60vSYTmpqCj1Gs,7745
5
+ periplus_sdk/dbapi.py,sha256=yDrL2CV7tI3Pawp9hN2czFYqrbXF8ret0Jk4VHRHoYg,12014
6
+ periplus_sdk/errors.py,sha256=rB1n-v8Hc2tu2dtHivz-MlqsCoRC5pTcWogTfM7SMLw,855
7
+ periplus_sdk/py.typed,sha256=AbpHGcgLb-kRsJGnwFEktk7uzpZOCcBY74-YBdrKVGs,1
8
+ periplus_sdk/sql_api.py,sha256=vRQXzBIcJOUg_eZzzRs7cpE3xygxK5Jo7YfYWN4NxBg,1868
9
+ periplus_sdk/sqlalchemy.py,sha256=e9j4nI_JQ5EDwMOK6WawX1W_QrOnaY3Wa8LchIIJJio,5567
10
+ periplus_sdk/stream.py,sha256=7gOYOMpfeB7NcxQgWjJ6f5NvymYTWhYbPkVdCijSlP8,5218
11
+ periplus_sdk/types.py,sha256=yh_xHwn6TwC45r8eWhxtlrz0w86M5lvhmPOS3QJbXiU,1253
12
+ periplus_python_sdk-0.8.0.dist-info/METADATA,sha256=7A_Jpwq8NJpskMpwQORGPhx-xbMHn_nehsqslBb_6c8,2730
13
+ periplus_python_sdk-0.8.0.dist-info/WHEEL,sha256=YVMoNqKzERt-wjUZwJ33xBGAwnFl-4cqbYkTtWa4itE,91
14
+ periplus_python_sdk-0.8.0.dist-info/entry_points.txt,sha256=Pr14L_7AhLinq-4qDxB4awFVubrvR1BEfkaFVRpdCbU,73
15
+ periplus_python_sdk-0.8.0.dist-info/top_level.txt,sha256=o41t5TzwgoxzSmbKoP6olWW1FyEAGWVYTjeoKadBK40,13
16
+ periplus_python_sdk-0.8.0.dist-info/RECORD,,
periplus_sdk/client.py CHANGED
@@ -79,10 +79,11 @@ def _payload(sql: str, parameters: Sequence[JsonValue] | None, schema_version: s
79
79
  class Client:
80
80
  """Reusable synchronous public query client. Close it or use a with block."""
81
81
 
82
- def __init__(self, base_url: str | None = None, *, timeout: float = 620, mode: Literal["stable", "experimental"] = "stable"):
83
- if mode not in {"stable", "experimental"}:
84
- raise ConfigurationError("mode must be stable or experimental.")
85
- self._query_path = "api/query/experimental/" if mode == "experimental" else "api/query/"
82
+ def __init__(self, base_url: str | None = None, *, timeout: float = 620, mode: Literal["stable"] = "stable"):
83
+ if mode not in {"stable"}:
84
+ raise ConfigurationError("Only the stable public catalogue is available.")
85
+ self.schema_version = "public_v1"
86
+ self._query_path = "api/query/"
86
87
  self._http = httpx.Client(**_options(base_url, timeout))
87
88
 
88
89
  def __enter__(self) -> Client:
@@ -101,19 +102,19 @@ class Client:
101
102
  raise TransportError("Could not complete the public query request.") from None
102
103
  return _decode(response, model)
103
104
 
104
- def prepare(self, sql: str, parameters: Sequence[JsonValue] | None = None, *, schema_version: str = "public_v1") -> PreparedQuery:
105
- return self._request("POST", "prep", PreparedQuery, json=_payload(sql, parameters, schema_version))
105
+ def prepare(self, sql: str, parameters: Sequence[JsonValue] | None = None, *, schema_version: str | None = None) -> PreparedQuery:
106
+ return self._request("POST", "prep", PreparedQuery, json=_payload(sql, parameters, schema_version if schema_version is not None else self.schema_version))
106
107
 
107
- def execute(self, sql: str, parameters: Sequence[JsonValue] | None = None, *, schema_version: str = "public_v1") -> QueryResult:
108
- return self._request("POST", "exec", QueryResult, json=_payload(sql, parameters, schema_version))
108
+ def execute(self, sql: str, parameters: Sequence[JsonValue] | None = None, *, schema_version: str | None = None) -> QueryResult:
109
+ return self._request("POST", "exec", QueryResult, json=_payload(sql, parameters, schema_version if schema_version is not None else self.schema_version))
109
110
 
110
111
  def stream(self, sql: str, parameters: Sequence[JsonValue] | None = None, *,
111
- schema_version: str = "public_v1", allow_partial: bool = False):
112
+ schema_version: str | None = None, allow_partial: bool = False):
112
113
  """Stream batches from one snapshot; use as a context manager for early exit."""
113
114
  from .stream import MEDIA_TYPE, QueryStream
114
115
  try:
115
116
  request = self._http.build_request("POST", self._query_path + "exec",
116
- headers={"accept": MEDIA_TYPE}, json=_payload(sql, parameters, schema_version))
117
+ headers={"accept": MEDIA_TYPE}, json=_payload(sql, parameters, schema_version if schema_version is not None else self.schema_version))
117
118
  response = self._http.send(request, stream=True)
118
119
  try:
119
120
  if not response.is_success:
@@ -133,10 +134,11 @@ class Client:
133
134
  class AsyncClient:
134
135
  """Reusable asynchronous public query client. Use an async with block."""
135
136
 
136
- def __init__(self, base_url: str | None = None, *, timeout: float = 620, mode: Literal["stable", "experimental"] = "stable"):
137
- if mode not in {"stable", "experimental"}:
138
- raise ConfigurationError("mode must be stable or experimental.")
139
- self._query_path = "api/query/experimental/" if mode == "experimental" else "api/query/"
137
+ def __init__(self, base_url: str | None = None, *, timeout: float = 620, mode: Literal["stable"] = "stable"):
138
+ if mode not in {"stable"}:
139
+ raise ConfigurationError("Only the stable public catalogue is available.")
140
+ self.schema_version = "public_v1"
141
+ self._query_path = "api/query/"
140
142
  self._http = httpx.AsyncClient(**_options(base_url, timeout))
141
143
 
142
144
  async def __aenter__(self) -> AsyncClient:
@@ -155,11 +157,11 @@ class AsyncClient:
155
157
  raise TransportError("Could not complete the public query request.") from None
156
158
  return _decode(response, model)
157
159
 
158
- async def prepare(self, sql: str, parameters: Sequence[JsonValue] | None = None, *, schema_version: str = "public_v1") -> PreparedQuery:
159
- return await self._request("POST", "prep", PreparedQuery, json=_payload(sql, parameters, schema_version))
160
+ async def prepare(self, sql: str, parameters: Sequence[JsonValue] | None = None, *, schema_version: str | None = None) -> PreparedQuery:
161
+ return await self._request("POST", "prep", PreparedQuery, json=_payload(sql, parameters, schema_version if schema_version is not None else self.schema_version))
160
162
 
161
- async def execute(self, sql: str, parameters: Sequence[JsonValue] | None = None, *, schema_version: str = "public_v1") -> QueryResult:
162
- return await self._request("POST", "exec", QueryResult, json=_payload(sql, parameters, schema_version))
163
+ async def execute(self, sql: str, parameters: Sequence[JsonValue] | None = None, *, schema_version: str | None = None) -> QueryResult:
164
+ return await self._request("POST", "exec", QueryResult, json=_payload(sql, parameters, schema_version if schema_version is not None else self.schema_version))
163
165
 
164
166
  async def helpers(self) -> QueryHelpers:
165
167
  return await self._request("GET", "helpers", QueryHelpers)
periplus_sdk/dbapi.py CHANGED
@@ -86,54 +86,47 @@ def TimestampFromTicks(ticks: float) -> datetime:
86
86
  return datetime.fromtimestamp(ticks)
87
87
 
88
88
 
89
- _INTEGER_TYPES = {"TINYINT", "SMALLINT", "INTEGER", "BIGINT", "HUGEINT", "UTINYINT",
90
- "USMALLINT", "UINTEGER", "UBIGINT", "UHUGEINT", "BIGNUM"}
91
- _FLOAT_TYPES = {"FLOAT", "DOUBLE", "REAL"}
92
- _TIME_TYPES = {"TIME", "TIME WITH TIME ZONE", "TIMETZ"}
93
- _TIMESTAMP_TYPES = {"TIMESTAMP", "TIMESTAMP_S", "TIMESTAMP_MS", "TIMESTAMP_NS",
94
- "TIMESTAMP WITH TIME ZONE", "TIMESTAMPTZ"}
89
+ _INTEGER_TYPES = {f'{prefix}Int{bits}' for prefix in ('', 'U') for bits in (8,16,32,64,128,256)}
90
+ _FLOAT_TYPES = {'Float32', 'Float64'}
91
+
92
+
93
+ def _base_type(value: str) -> str:
94
+ while value.startswith(('Nullable(', 'LowCardinality(')):
95
+ value = value[value.index('(')+1:-1]
96
+ return value
95
97
 
96
98
 
97
99
  class _TypeCategory:
98
- def __init__(self, names: set[str], prefix: str = ""):
99
- self.names, self.prefix = names, prefix
100
+ def __init__(self, names: set[str], prefixes: tuple[str, ...] = ()):
101
+ self.names, self.prefixes = names, prefixes
100
102
 
101
103
  def __eq__(self, other: object) -> bool:
102
- return isinstance(other, str) and (other in self.names or bool(self.prefix and other.startswith(self.prefix)))
104
+ if not isinstance(other, str):return False
105
+ value = _base_type(other)
106
+ return value in self.names or value.startswith(self.prefixes)
103
107
 
104
108
 
105
- STRING = _TypeCategory({"VARCHAR", "UUID", "JSON", "ENUM"})
106
- BINARY = _TypeCategory({"BLOB"})
107
- NUMBER = _TypeCategory(_INTEGER_TYPES | _FLOAT_TYPES | {"BOOLEAN"}, "DECIMAL(")
108
- DATETIME = _TypeCategory({"DATE"} | _TIME_TYPES | _TIMESTAMP_TYPES)
109
+ STRING = _TypeCategory({'String','UUID','JSON'}, ('FixedString(', 'Enum'))
110
+ BINARY = _TypeCategory(set())
111
+ NUMBER = _TypeCategory(_INTEGER_TYPES | _FLOAT_TYPES | {'Bool'}, ('Decimal',))
112
+ DATETIME = _TypeCategory({'Date','Date32'}, ('DateTime','Time'))
109
113
  ROWID = _TypeCategory(set())
110
114
 
111
115
 
112
116
  def _value(value: Any, sql_type: str) -> Any:
113
- if value is None:
114
- return None
115
- if sql_type in _INTEGER_TYPES:
116
- return int(value)
117
- if sql_type in _FLOAT_TYPES:
118
- return float(value)
119
- if sql_type.startswith("DECIMAL("):
120
- return Decimal(str(value))
121
- # Preserve infinities and out-of-range dates rather than clipping them.
122
- if sql_type == "DATE":
123
- try:
124
- return date.fromisoformat(value)
125
- except ValueError:
126
- return value
127
- if sql_type in _TIME_TYPES:
128
- return time.fromisoformat(value)
129
- if sql_type in _TIMESTAMP_TYPES:
130
- try:
131
- return datetime.fromisoformat(value)
132
- except ValueError:
133
- return value
134
- if sql_type == "BLOB":
135
- return base64.b64decode(value, validate=True)
136
- # UUIDs remain strings, and nested/other types retain their JSON wire values.
117
+ if value is None:return None
118
+ sql_type = _base_type(sql_type)
119
+ if sql_type in _INTEGER_TYPES:return int(value)
120
+ if sql_type in _FLOAT_TYPES:return float(value)
121
+ if sql_type.startswith('Decimal'):return Decimal(str(value))
122
+ if sql_type == 'Bool':return value in (True, 1, '1', 'true')
123
+ if sql_type in ('Date','Date32'):
124
+ try:return date.fromisoformat(value)
125
+ except ValueError:return value
126
+ if sql_type.startswith('DateTime'):
127
+ try:return datetime.fromisoformat(value)
128
+ except ValueError:return value
129
+ if sql_type.startswith('Time'):return time.fromisoformat(value)
137
130
  return value
138
131
 
139
132
 
@@ -152,16 +145,16 @@ def _parameter(value: Any) -> Any:
152
145
  class Connection:
153
146
  """Marimo-discoverable, read-only connection; commit is a no-op."""
154
147
 
155
- dialect = "duckdb"
148
+ dialect = "clickhouse"
156
149
 
157
150
  def __init__(self, base_url: str | None = None, *, timeout: float = 620,
158
- mode: Literal["stable", "experimental"] = "stable",
159
- schema_version: str = "public_v1", allow_partial: bool = False):
151
+ mode: Literal["stable"] = "stable",
152
+ schema_version: str | None = None, allow_partial: bool = False):
160
153
  try:
161
154
  self._client = Client(base_url, timeout=timeout, mode=mode)
162
155
  except ConfigurationError as exc:
163
156
  raise InterfaceError(str(exc)) from exc
164
- self.schema_version = schema_version
157
+ self.schema_version = schema_version if schema_version is not None else ("public_v1")
165
158
  self.allow_partial = allow_partial
166
159
  self._cursors = set()
167
160
  self.closed = False
@@ -210,8 +203,8 @@ class Connection:
210
203
 
211
204
 
212
205
  def connect(base_url: str | None = None, *, timeout: float = 620,
213
- mode: Literal["stable", "experimental"] = "stable",
214
- schema_version: str = "public_v1", allow_partial: bool = False) -> Connection:
206
+ mode: Literal["stable"] = "stable",
207
+ schema_version: str | None = None, allow_partial: bool = False) -> Connection:
215
208
  return Connection(base_url, timeout=timeout, mode=mode, schema_version=schema_version, allow_partial=allow_partial)
216
209
 
217
210
 
periplus_sdk/sql_api.py CHANGED
@@ -10,9 +10,9 @@ from sqlalchemy.engine import Engine, URL
10
10
  def create_engine(
11
11
  base_url: str | None = None,
12
12
  *,
13
- mode: Literal["stable", "experimental"] = "stable",
13
+ mode: Literal["stable"] = "stable",
14
14
  timeout: float = 620,
15
- schema_version: str = "public_v1",
15
+ schema_version: str | None = None,
16
16
  allow_partial: bool = False,
17
17
  ) -> Engine:
18
18
  """Create a SQLAlchemy engine recognized by marimo and other SQL tools.
@@ -9,10 +9,11 @@ from sqlalchemy.engine.reflection import cache
9
9
  from sqlalchemy.sql.compiler import IdentifierPreparer
10
10
 
11
11
  from . import dbapi
12
+ from .errors import ApiError, ResponseError, TransportError
12
13
 
13
14
 
14
15
  class SQLType(types.UserDefinedType):
15
- """Preserve DuckDB type names, including nested types, during reflection."""
16
+ """Preserve ClickHouse type names, including nested types, during reflection."""
16
17
 
17
18
  cache_ok = True
18
19
 
@@ -24,28 +25,21 @@ class SQLType(types.UserDefinedType):
24
25
 
25
26
  @property
26
27
  def python_type(self) -> type:
27
- if self.name in dbapi._INTEGER_TYPES:
28
- return int
29
- if self.name in dbapi._FLOAT_TYPES:
30
- return float
31
- if self.name == "BOOLEAN":
32
- return bool
33
- if self.name.startswith("DECIMAL("):
34
- return dbapi.Decimal
35
- if self.name == "DATE":
36
- return dbapi.date
37
- if self.name in dbapi._TIMESTAMP_TYPES:
38
- return dbapi.datetime
39
- if self.name in dbapi._TIME_TYPES:
40
- return dbapi.time
41
- if self.name == "BLOB":
42
- return bytes
28
+ name = dbapi._base_type(self.name)
29
+ if name in dbapi._INTEGER_TYPES:return int
30
+ if name in dbapi._FLOAT_TYPES:return float
31
+ if name == 'Bool':return bool
32
+ if name.startswith('Decimal'):return dbapi.Decimal
33
+ if name in ('Date','Date32'):return dbapi.date
34
+ if name.startswith('DateTime'):return dbapi.datetime
35
+ if name.startswith('Time'):return dbapi.time
43
36
  return str
44
37
 
45
38
 
39
+
46
40
  class PeriplusDialect(default.DefaultDialect):
47
- # The server speaks DuckDB SQL; this enables the correct notebook SQL dialect.
48
- name = "duckdb"
41
+ # The server speaks ClickHouse SQL; this enables the correct notebook SQL dialect.
42
+ name = "clickhouse"
49
43
  driver = "periplus"
50
44
  supports_statement_cache = False
51
45
  supports_sane_rowcount = False
@@ -102,8 +96,15 @@ class PeriplusDialect(default.DefaultDialect):
102
96
  @cache
103
97
  def get_view_names(self, connection, schema=None, **kw):
104
98
  schema = self._schema(connection, schema)
105
- result = connection.exec_driver_sql(f"SHOW TABLES FROM {self.identifier_preparer.quote_identifier(schema)}")
106
- return [row[0] for row in self._complete(result)]
99
+ try:
100
+ catalogue = connection.connection.dbapi_connection._client.helpers()
101
+ except (ApiError, ResponseError, TransportError) as error:
102
+ raise exc.InvalidRequestError(f"Public catalogue discovery failed: {error}") from error
103
+ if catalogue.schema_version != schema:
104
+ raise exc.InvalidRequestError("Public catalogue does not match the configured schema.")
105
+ prefix = f"{schema}."
106
+ return [relation.name.removeprefix(prefix) for relation in catalogue.relations
107
+ if relation.kind == "view" and relation.name.startswith(prefix)]
107
108
 
108
109
  @cache
109
110
  def get_table_names(self, connection, schema=None, **kw):
@@ -119,9 +120,15 @@ class PeriplusDialect(default.DefaultDialect):
119
120
  def get_columns(self, connection, table_name, schema=None, **kw):
120
121
  schema = self._schema(connection, schema)
121
122
  quote = self.identifier_preparer.quote_identifier
122
- result = connection.exec_driver_sql(f"DESCRIBE {quote(schema)}.{quote(table_name)}")
123
- return [{"name": row[0], "type": SQLType(row[1]), "nullable": row[2] != "NO",
124
- "default": row[4]} for row in self._complete(result)]
123
+ # A zero-row SELECT obtains native types through the supported public
124
+ # query contract, without scanning the view or requiring system access.
125
+ result = connection.exec_driver_sql(f"SELECT * FROM {quote(schema)}.{quote(table_name)} LIMIT 0")
126
+ metadata = result.cursor.result
127
+ self._complete(result)
128
+ return [{"name": name, "type": SQLType(kind),
129
+ "nullable": kind.startswith("Nullable(") or kind.startswith("LowCardinality(Nullable("),
130
+ "default": None}
131
+ for name, kind in zip(metadata.columns, metadata.types, strict=True)]
125
132
 
126
133
  @cache
127
134
  def get_pk_constraint(self, connection, table_name, schema=None, **kw):
periplus_sdk/stream.py CHANGED
@@ -16,7 +16,7 @@ MEDIA_TYPE = "application/x-ndjson"
16
16
  class StreamResult(PreparedQuery):
17
17
  columns: list[str]
18
18
  types: list[str]
19
- source_snapshot: int = Field(ge=0)
19
+ source_snapshot: int | None = Field(default=None, ge=0)
20
20
  limits: dict[str, int]
21
21
  complete: bool = False
22
22
  truncated: bool = False
periplus_sdk/types.py CHANGED
@@ -11,7 +11,7 @@ class Diagnostic(BaseModel):
11
11
 
12
12
 
13
13
  class PreparedQuery(BaseModel):
14
- query_mode: Literal["stable", "experimental"]
14
+ query_mode: Literal["stable"]
15
15
  compiler_version: str
16
16
  optimizations: list[str]
17
17
  schema_version: str
@@ -28,7 +28,7 @@ class QueryResult(PreparedQuery):
28
28
  rows: list[list[JsonValue]]
29
29
  truncated: bool
30
30
  elapsed_ms: float
31
- source_snapshot: int = Field(ge=0)
31
+ source_snapshot: int | None = Field(default=None, ge=0)
32
32
  row_count: int = Field(ge=0)
33
33
  result_bytes: int = Field(ge=0)
34
34
  truncation_reason: Literal["max_rows", "max_result_bytes"] | None = None
@@ -51,4 +51,6 @@ class QueryHelper(BaseModel):
51
51
 
52
52
  class QueryHelpers(BaseModel):
53
53
  catalogue_version: str
54
+ schema_version: str
55
+ relations: list[QueryHelper]
54
56
  helpers: list[QueryHelper]
@@ -1,301 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: periplus-python-sdk
3
- Version: 0.7.0
4
- Summary: Read-only Python client for the public Periplus query API
5
- License-Expression: Apache-2.0
6
- Project-URL: Repository, https://github.com/elei-io/periplus
7
- Project-URL: Issues, https://github.com/elei-io/periplus/issues
8
- Requires-Python: >=3.11
9
- Description-Content-Type: text/markdown
10
- License-File: LICENSE
11
- License-File: NOTICE
12
- Requires-Dist: httpx>=0.28
13
- Requires-Dist: pydantic<3,>=2.12
14
- Requires-Dist: sqlalchemy<3,>=2.0
15
- Provides-Extra: notebook
16
- Requires-Dist: marimo[sql]>=0.24.1; extra == "notebook"
17
- Dynamic: license-file
18
-
19
- # Periplus Python SDK
20
-
21
- A read-only client for the public Periplus query API. Python 3.11 or later.
22
- Configure the **public web application URL**, not the internal query or control service.
23
- No API token, DuckDB installation or lake credentials are needed.
24
-
25
- ```python
26
- from periplus_sdk import Client
27
-
28
- with Client("http://localhost:8080") as client:
29
- result = client.execute(
30
- "SELECT capture_id FROM public_v1.capture LIMIT ?", [10]
31
- )
32
- print(result.columns, result.types)
33
- print(result.rows)
34
- print(result.source_snapshot, result.truncated)
35
- ```
36
-
37
- For a hosted deployment, replace the URL with its public HTTPS origin. Alternatively set
38
- `PERIPLUS_PUBLIC_URL` and use `Client()`. An optional URL path prefix is preserved.
39
- The client reuses HTTP connections; close it with a context manager or `close()`.
40
-
41
- ## Marimo SQL cells and schema browser
42
-
43
- Install the notebook integration from PyPI:
44
-
45
- ```sh
46
- uv add "periplus-python-sdk[notebook]>=0.7.0"
47
- ```
48
-
49
- In a Python setup cell, create a SQLAlchemy engine:
50
-
51
- ```python
52
- from periplus_sdk import sql_api
53
-
54
- pp = sql_api.create_engine("https://periplus.dev", mode="stable")
55
- ```
56
-
57
- Add a SQL cell, select **pp** in its connection dropdown, and enter:
58
-
59
- ```sql
60
- SELECT capture_id, requested_url
61
- FROM public_v1.capture
62
- LIMIT 10
63
- ```
64
-
65
- Marimo displays the result as a table. Expand **pp → periplus → public_v1** in Data Sources
66
- to discover views and expand a view to load its columns for SQL completion.
67
- Discovery uses bounded `SHOW TABLES` and `DESCRIBE` through the same public API;
68
- no internal catalogue or storage credentials are used. Truncated discovery fails
69
- explicitly rather than displaying a silently incomplete schema. To eagerly load
70
- schemas and views, enable their discovery in marimo's Packages & Data settings.
71
- Column discovery is on demand by default, to avoid many public API requests.
72
-
73
- The Python equivalent of a SQL cell is:
74
-
75
- ```python
76
- import marimo as mo
77
-
78
- captures = mo.sql(
79
- "SELECT capture_id FROM public_v1.capture LIMIT 10",
80
- engine=pp,
81
- )
82
- ```
83
-
84
- Set `mode="experimental"` for the experimental service. Omit the URL to use
85
- `PERIPLUS_PUBLIC_URL`. Optional `timeout=620` and `schema_version="public_v1"`
86
- arguments configure the client deadline and public schema. Run `pp.dispose()` when finished. This is a read-only
87
- SQLAlchemy dialect for textual SQL and reflection, not a writable ORM backend.
88
- Each statement has its own server snapshot; SQLAlchemy transaction blocks do not
89
- provide a shared snapshot or rollback. The adapter makes no transaction requests.
90
-
91
- A complete notebook is in `examples/notebook.py`. The integration is tested with
92
- marimo 0.24.1 and SQLAlchemy 2.x. SQLAlchemy is included in the standard SDK install; the `notebook` extra adds
93
- marimo. Existing marimo environments only need `uv add "periplus-python-sdk>=0.7.0"`.
94
- The returned object is a standard SQLAlchemy Engine, also usable with pandas and
95
- ordinary Python scripts. Engine creation is lazy; the first query opens a connection.
96
-
97
- ## DB-API connection
98
-
99
- For SQL cells without schema browsing, or standard cursor-based Python code:
100
-
101
- ```python
102
- from periplus_sdk import connect
103
-
104
- with connect("https://periplus.dev", mode="stable") as connection:
105
- with connection.cursor() as cursor:
106
- cursor.execute("SELECT capture_id FROM public_v1.capture LIMIT ?", [10])
107
- print(cursor.description)
108
- print(cursor.fetchall())
109
- print(cursor.result.source_snapshot)
110
- ```
111
-
112
- Connections expose `cursor`, `execute`, `close`, and context managers. Cursors
113
- support `execute`, `fetchone`, `fetchmany`, `fetchall`, iteration, and close.
114
- Use positional `?` parameters. Decimal and temporal parameters are sent as
115
- strings; use explicit SQL casts. One-dimensional lists/tuples of these scalar values are supported; use an explicit array cast
116
- for empty or typed lists. Binary, mappings, and nested collection parameters are unsupported. Fetching consumes incremental batches from one HTTP response and one source snapshot;
117
- it never issues pagination or retries. Close a cursor early to stop delivery. Connections/cursors are not thread-shared.
118
- `commit()` is a no-op; `rollback()` and `executemany()` are unsupported.
119
-
120
- `cursor.result` preserves stream metadata and completion information without retaining consumed rows. `connection.last_result`
121
- also retains it after marimo closes a cursor; a new execution clears it first.
122
- Truncation raises `OperationalError` with `code="result_limit"` when fetching reaches the terminal frame.
123
- `rowcount` stays -1 until completion, then reports the delivered count. `result.complete` means the
124
- terminal frame arrived; inspect `result.truncated` separately when opting into partial data.
125
- Use `allow_partial=True` on the engine or connection only when incomplete results are intentional.
126
- DB-API failures use the standard exception hierarchy in `periplus_sdk.dbapi`;
127
- HTTP errors retain `status_code`, `code`, and `retry_after_seconds`.
128
-
129
- Scalar integer, floating-point, decimal, date, time, timestamp and BLOB results
130
- are decoded to Python values. UUIDs remain strings. Nested/other SQL types keep
131
- their JSON wire representation; out-of-range dates/timestamps remain strings.
132
- Temporal precision is limited to what the server JSON transport preserves.
133
- The cursor preserves duplicate column names, but dataframe libraries/marimo may
134
- not: use unique SQL aliases. Dataframe inference can lose types for empty or
135
- all-null results; `cursor.description` retains the SQL type names.
136
-
137
- ## Stable and experimental APIs
138
-
139
- Both clients accept `mode="stable"` (the default) or `mode="experimental"` at initialization:
140
-
141
- ```python
142
- with Client("https://periplus.dev", mode="experimental") as client:
143
- result = client.execute("SELECT capture_id FROM public_v1.capture LIMIT 1")
144
- print(result.query_mode, result.compiler_version, result.optimizations)
145
- ```
146
-
147
- The selected mode applies to preparation, execution, and helper discovery. Experimental
148
- requests use the public application's `/api/query/experimental/` routes. There is no
149
- automatic fallback to stable if the experimental service is unavailable.
150
- `AsyncClient` accepts the same option. Invalid modes raise `ConfigurationError`.
151
-
152
- ## Preparation and helpers
153
-
154
- ```python
155
- with Client("http://localhost:8080") as client:
156
- prepared = client.prepare("SELECT capture_id FROM public_v1.capture LIMIT ?", [10])
157
- print(prepared.diagnostics, prepared.plan)
158
- result = client.execute(prepared.sql, prepared.parameters)
159
- helpers = client.helpers()
160
- print(helpers.catalogue_version, helpers.helpers)
161
- ```
162
-
163
- Preparation validates and explains without executing the analytical query. Execution independently
164
- validates and prepares; a prior preparation never authorizes SQL. Linting, diagnostics and future
165
- SQL optimizations belong to the server. The SDK sends SQL unchanged.
166
-
167
- ## Async use
168
-
169
- ```python
170
- from periplus_sdk import AsyncClient
171
-
172
- async def observations():
173
- async with AsyncClient("http://localhost:8080") as client:
174
- return await client.execute("SELECT capture_id FROM public_v1.capture LIMIT 10")
175
- ```
176
-
177
- Use `aclose()` when managing an async client's lifetime explicitly.
178
-
179
- ## Permissions, results and errors
180
-
181
- - The same public SQL feature switch, shared rate budget, namespace validation and read-only
182
- execution apply as in the public web workspace. The SDK provides no writes, crawling,
183
- administrative controls or direct lake attachment.
184
- - Results retain `query_id`, SQL, parameters, diagnostics, plan, columns, SQL types, JSON rows,
185
- elapsed milliseconds, `source_snapshot` and `truncated`. Decimals and large integers remain
186
- strings exactly as returned by the server. Duplicate column names are preserved.
187
- - Operator-configured execution limits default to 1,000 rows, an 8 MiB result budget and a
188
- 20-second server deadline. Always inspect `truncated`. The SDK does not silently fetch more rows or retry.
189
- - `ApiError` exposes `status_code`, safe `code`, and `retry_after_seconds` when supplied.
190
- `TransportError` means HTTP failed; `ResponseError` means a malformed successful response.
191
- The client timeout defaults to 620 seconds and can be set with `timeout=`. A timeout or local
192
- cancellation does not guarantee server cancellation. Redirects are not followed automatically.
193
- - Preparation and execution are attributed to `sdk` in the existing private query history.
194
- Original SQL and parameters are retained for 30 days; result rows are not stored. Recording is
195
- best-effort and can be lost during outages or backpressure. This label is not a user identity.
196
-
197
- ## Installation and verification
198
-
199
- Install the public-v1 client from PyPI:
200
-
201
- ```sh
202
- python -m pip install "periplus-python-sdk>=0.7.0"
203
- ```
204
-
205
- Version 0.6.1 supports the current public-v1 contract. For production, configure
206
- `PERIPLUS_PUBLIC_URL=https://periplus.dev`; no API token is required.
207
- Run the installed package against an available public app:
208
-
209
- ```sh
210
- PERIPLUS_PUBLIC_URL=http://localhost:8080 python packages/periplus-python-sdk/examples/smoke.py
211
- ```
212
-
213
- ## Releasing
214
-
215
- Repository CI publishes immutable releases from tags named
216
- `periplus-python-sdk-v<version>`. The tag must exactly match the static version
217
- in `pyproject.toml`; for example, version `0.7.0` is released with:
218
-
219
- ```sh
220
- git tag periplus-python-sdk-v0.7.0
221
- git push origin periplus-python-sdk-v0.7.0
222
- ```
223
-
224
- PyPI publishing uses Trusted Publishing rather than a stored API token. The
225
- PyPI publisher must be configured for GitHub owner `elei-io`, repository
226
- `periplus`, workflow `python-sdk-release.yml`, and environment `pypi`. Protect
227
- that GitHub environment with required reviewers before the first release.
228
-
229
- ## Public v1
230
-
231
- Install the updated SDK from PyPI with `python -m pip install "periplus-python-sdk>=0.7.0"`. The previously published 0.2.0 release predates this contract. `prepare` and `execute` accept keyword-only `schema_version="public_v1"` (the default); responses preserve `schema_version` separately from `source_snapshot`. Unavailable versions are rejected by the server.
232
-
233
- ## License
234
-
235
- Copyright (c) 2026 Ekku Leivonen (elei.io). Licensed under [Apache-2.0](LICENSE);
236
- see [NOTICE](NOTICE). The server and other repository packages have separate
237
- licensing described in the root LICENSING.md.
238
-
239
- ## Python filtering back into marimo SQL
240
-
241
- The existing engine streams by default. For SQL → Python filtering → SQL, bind the
242
- filtered IDs to another engine variable; it shares the original connection pool and
243
- creates no server-side table or upload. Marimo discovers it like any other SQLAlchemy engine.
244
-
245
- ```python
246
- pp = sql_api.create_engine("https://periplus.dev")
247
-
248
- # In a Python cell, after filtering your initial SQL dataframe:
249
- selected = sql_api.bind(pp, content_ids=qualified_pages["content_id"].to_list())
250
-
251
- # In the next SQL cell (engine: selected):
252
- raw_elements = mo.sql(
253
- """
254
- SELECT content_id, node_index, parent_index, tag,
255
- trim(text_direct) AS text, attributes['class'] AS elem_class
256
- FROM public_v1.html_element
257
- WHERE content_id IN (SELECT unnest(CAST(:content_ids AS VARCHAR[])))
258
- AND text_direct IS NOT NULL AND trim(text_direct) <> ''
259
- """,
260
- engine=selected,
261
- )
262
- ```
263
-
264
- Keep SQL strings plain: `:content_ids` is a bound parameter, not an f-string.
265
- An empty list selects no elements; duplicate IDs do not multiply rows. Re-running the
266
- binding cell snapshots the new Python list into its engine options. Bindings only apply
267
- to matching SQL placeholders, so catalogue discovery remains available. Explicit
268
- SQLAlchemy execution parameters take precedence over engine bindings.
269
-
270
- Marimo collects the stream into a local dataframe, which must fit notebook memory.
271
- Its dataframe conversion currently infers types and may lose empty/all-null column
272
- information; the DB-API cursor always exposes server column names and SQL types.
273
- Initial SQL results and subsequent element results both reject truncation by default.
274
- Each separate SQL execution has its own snapshot; Python filtering does not pin the
275
- first query's snapshot across the round trip.
276
-
277
- ## Consume batches without a dataframe
278
-
279
- ```python
280
- with Client("https://periplus.dev") as client:
281
- with client.stream("SELECT content_id FROM public_v1.prose") as stream:
282
- for rows in stream:
283
- process_batch(rows)
284
- print(stream.result.row_count, stream.result.source_snapshot)
285
- print(stream.result.limits)
286
- ```
287
-
288
- Batches retain JSON wire values with SQL types in `stream.result.types`. DB-API fetching
289
- performs scalar Python decoding. The synchronous `Client.stream()` raises on truncation
290
- unless `allow_partial=True`; missing completion, malformed frames, timeouts, and broken
291
- connections always fail. Do not treat already-consumed batches as a complete dataset
292
- until iteration finishes successfully. Use context managers when stopping early.
293
-
294
- Admin Query limits controls requests (default 4 MiB, ceiling 16 MiB), parameter values
295
- (default 100,000, ceiling 1,000,000), results (ceiling 10 million rows and 1 GiB), and
296
- execution duration (ceiling 600 seconds). Parameter counting includes collection containers
297
- and their values. Existing result/duration defaults remain 1,000 rows, 8 MiB, and 20 seconds.
298
- Operators must raise these settings for larger workloads. No SDK parameter grants a larger
299
- server budget. Stream metadata includes effective limits; input rejections name the budget.
300
- Deploy the API, query service, public gateway and SDK together after applying Alembic
301
- revision `20260911_0016`. Ingress must permit the configured request size and stream duration.
@@ -1,16 +0,0 @@
1
- periplus_python_sdk-0.7.0.dist-info/licenses/LICENSE,sha256=z8d0m5b2O9McPEK1xHG_dWgUBT6EfBDz6wA0F7xSPTA,11358
2
- periplus_python_sdk-0.7.0.dist-info/licenses/NOTICE,sha256=bhbYSqcUB3U_P1-XzloiT81JGniqoYaRLxNkQ1Pm9MQ,52
3
- periplus_sdk/__init__.py,sha256=WimXYlPB6tCimBO4VSwhcp00dwSL87jMmMuQ4-kINfM,546
4
- periplus_sdk/client.py,sha256=bqIq2bmM_bnwf0ZnMN5kTXTctnEz7ITJUAfZohGsRIQ,7544
5
- periplus_sdk/dbapi.py,sha256=wyhtLfpgBZGxhC9rA53DNfQ7j8HegZkVp7mRq98bFSU,12256
6
- periplus_sdk/errors.py,sha256=rB1n-v8Hc2tu2dtHivz-MlqsCoRC5pTcWogTfM7SMLw,855
7
- periplus_sdk/py.typed,sha256=AbpHGcgLb-kRsJGnwFEktk7uzpZOCcBY74-YBdrKVGs,1
8
- periplus_sdk/sql_api.py,sha256=o5kLzwLim3VQ5oZfRyPUGCDr7RWmDidtktGLyBdkjQk,1884
9
- periplus_sdk/sqlalchemy.py,sha256=skyN2R5nIiWK55uBuxOHjeGRAPb8uPxm0f8nU0IIPqw,4836
10
- periplus_sdk/stream.py,sha256=pPXWikBWnOUZuHfBRXx08swBDMWJ4Gayx6azaz_wUCU,5197
11
- periplus_sdk/types.py,sha256=0eCwWev-Vrk0mRCIkot3Z_-dWWL3gpjljV4GyXEf8lk,1191
12
- periplus_python_sdk-0.7.0.dist-info/METADATA,sha256=6HTXHT4WnFKYbfB6MVc1Duxni8niJ3JQ7hP7w2GYkiA,14052
13
- periplus_python_sdk-0.7.0.dist-info/WHEEL,sha256=YVMoNqKzERt-wjUZwJ33xBGAwnFl-4cqbYkTtWa4itE,91
14
- periplus_python_sdk-0.7.0.dist-info/entry_points.txt,sha256=Pr14L_7AhLinq-4qDxB4awFVubrvR1BEfkaFVRpdCbU,73
15
- periplus_python_sdk-0.7.0.dist-info/top_level.txt,sha256=o41t5TzwgoxzSmbKoP6olWW1FyEAGWVYTjeoKadBK40,13
16
- periplus_python_sdk-0.7.0.dist-info/RECORD,,