arrowbricks 3.0.3__tar.gz → 3.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/PKG-INFO +62 -5
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/README.md +61 -4
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/pyproject.toml +1 -1
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/Cargo.lock +164 -1
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/Cargo.toml +15 -1
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/README.md +1 -1
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/src/client.rs +621 -120
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/src/heartbeat.rs +67 -1
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/src/json_convert.rs +124 -1
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/src/lib.rs +373 -27
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/src/pipeline.rs +735 -88
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/src/thrift.rs +221 -4
- arrowbricks-3.1.0/rust/arrowbricks_core/tests/common/mod.rs +29 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests/wiremock_pipeline.rs +282 -21
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests/wiremock_thrift.rs +202 -1
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests_py/test_token_provider.py +13 -5
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/src/arrowbricks/__init__.py +6 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/src/arrowbricks/_core.pyi +37 -1
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/src/arrowbricks/_streaming.py +16 -2
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/src/arrowbricks/client.py +36 -1
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/src/arrowbricks/cursor.py +70 -5
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/LICENSE +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/.gitignore +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/examples/duckdb_query.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/examples/fastapi_sse.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/rustfmt.toml +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests/wiremock_volume_files.rs +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests_py/conftest.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests_py/test_ipc_stream.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests_py/test_parameters.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests_py/test_stream_ndjson_lines.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests_py/test_streaming.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests_py/test_thrift.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests_py/test_volume_files.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/rust/arrowbricks_core/tests_py/thrift_mock.py +0 -0
- {arrowbricks-3.0.3 → arrowbricks-3.1.0}/src/arrowbricks/py.typed +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: arrowbricks
|
|
3
|
-
Version: 3.0
|
|
3
|
+
Version: 3.1.0
|
|
4
4
|
Requires-Dist: arro3-core>=0.8 ; extra == 'arro3'
|
|
5
5
|
Provides-Extra: arro3
|
|
6
6
|
License-File: LICENSE
|
|
@@ -77,9 +77,14 @@ async for item in client.stream_query_json("SELECT * FROM my_catalog.my_schema.b
|
|
|
77
77
|
|
|
78
78
|
See [`examples/basic.py`](examples/basic.py) for a runnable version,
|
|
79
79
|
[`examples/cursor_paging.py`](examples/cursor_paging.py) for paging a large
|
|
80
|
-
result with `fetchmany`/`fetchmany_arrow` without buffering it all upfront,
|
|
80
|
+
result with `fetchmany`/`fetchmany_arrow` without buffering it all upfront,
|
|
81
81
|
[`examples/azure_auth.py`](examples/azure_auth.py) for a caching
|
|
82
|
-
`token_provider` built on Azure AD (`DefaultAzureCredential
|
|
82
|
+
`token_provider` built on Azure AD (`DefaultAzureCredential`, needs
|
|
83
|
+
`azure-identity`), or [`examples/oauth_m2m_auth.py`](examples/oauth_m2m_auth.py)
|
|
84
|
+
for the same idea using Databricks' own OAuth machine-to-machine
|
|
85
|
+
client-credentials flow instead -- like `databricks-sql-connector`'s
|
|
86
|
+
`auth_type="databricks-oauth"`, but stdlib-only (`urllib.request`), no
|
|
87
|
+
extra dependency.
|
|
83
88
|
|
|
84
89
|
## FastAPI SSE example
|
|
85
90
|
|
|
@@ -135,19 +140,71 @@ arrowbricks has no opinion on *how* you get a token and no cloud-SDK dependency
|
|
|
135
140
|
conn = connect(host=..., warehouse_id=..., token_provider=my_token_provider)
|
|
136
141
|
```
|
|
137
142
|
|
|
143
|
+
## Observability
|
|
144
|
+
|
|
145
|
+
Pass `on_event` to `connect`/`DatabricksClient` to get a `QueryStats` snapshot once per query, at completion (success, cancelled, timed out, or errored):
|
|
146
|
+
|
|
147
|
+
```python
|
|
148
|
+
def log_query(stats):
|
|
149
|
+
print(
|
|
150
|
+
f"{stats.statement_id}: {stats.outcome} in {stats.fetch_s:.2f}s, {stats.bytes_downloaded} bytes, {stats.retry_count} retries"
|
|
151
|
+
)
|
|
152
|
+
|
|
153
|
+
|
|
154
|
+
conn = connect(host=..., warehouse_id=..., token=..., on_event=log_query)
|
|
155
|
+
```
|
|
156
|
+
|
|
157
|
+
`on_event` is sync or async, same as `token_provider`, and applies to every query run through that client (not passed per-call). It's strictly fire-and-forget: any exception it raises is caught and swallowed on the Rust side, never surfaced to your query, and dispatching it never blocks or measurably slows the actual fetch -- a slow or broken logging/metrics callback can't make your queries slower or fail them. `QueryStats` fields: `statement_id: str`, `protocol: "thrift" | "sea"`, `warehouse_wait_s: float` (0.0 if the client already knew the warehouse was running), `submit_to_ready_s: float` (statement submit -> ready to fetch), `fetch_s: float` (download + decode), `num_chunks: int`, `bytes_downloaded: int`, `retry_count: int`, `concurrency_used: int`, and `outcome: "success" | "cancelled" | "timeout" | "error"`.
|
|
158
|
+
|
|
159
|
+
`on_event` fires once the fetch of a given query actually completes (or is abandoned) -- a `Cursor` whose result is never fully drained (e.g. only a partial `fetchmany()`, then abandoned) never fires one, same category as this package's existing documented gap around GC-based cleanup.
|
|
160
|
+
|
|
161
|
+
## Cancellation
|
|
162
|
+
|
|
163
|
+
When a chunk download times out (`total_timeout_s`) or your code cancels the surrounding coroutine (`task.cancel()`/`asyncio.wait_for`) while it's in flight, arrowbricks fires a best-effort server-side cancel in the background (Thrift's `CancelOperation`, or SEA's `POST .../cancel`) so Databricks stops running the query instead of finishing it for nobody. This is fire-and-forget: the `QueryTimeout`/cancellation still reaches you immediately, and the cancel call's own result (success or failure) is never surfaced or awaited.
|
|
164
|
+
|
|
165
|
+
This covers `Cursor.fetchall_streamed()`/`fetchall_arrow_streamed()`, `client.stream_query_json(...)`, and the lower-level `._core.Client`'s own streamed APIs -- everything that wraps the chunk-download phase in the Rust-level heartbeat. Neither `total_timeout_s` nor a bare `task.cancel()` on `Cursor.execute()`/`execute_streamed()` itself (the initial submit/poll wait) triggers a cancel -- by the time a query's chunks are being fetched at all, the statement has typically already finished running server-side, so there's usually nothing left to cancel there in practice.
|
|
166
|
+
|
|
167
|
+
## Errors
|
|
168
|
+
|
|
169
|
+
Every exception arrowbricks raises itself is an `ArrowbricksError` -- a `RuntimeError` subclass, so an `except RuntimeError` written before this hierarchy existed keeps catching everything unchanged. Three subclasses split by what actually changes what your code should do next:
|
|
170
|
+
|
|
171
|
+
- `TransientError` -- a network blip, connection reset, or 5xx that survived every internal retry (`retry_attempts`, exponential backoff up to `retry_max_wait_s` -- see below) before ever reaching you. Back off further and try again later; retrying immediately just repeats what already failed.
|
|
172
|
+
- `AuthError` -- HTTP 401/403, even after every internal retry re-fetched a token from your `token_provider`, *or* your `token_provider` itself raised (e.g. its own OAuth refresh call came back unauthorized). Treat the credential itself as bad/expired, not a transient blip.
|
|
173
|
+
- `StatementError` -- the SQL statement itself failed or was canceled server-side (bad SQL, a permissions error, a warehouse-side query failure). Retrying the identical statement will fail identically; this isn't a transport problem.
|
|
174
|
+
|
|
175
|
+
Anything else (a generic non-auth/non-statement HTTP error, an internal parse/decode failure) raises the plain `ArrowbricksError` base.
|
|
176
|
+
|
|
177
|
+
```python
|
|
178
|
+
from arrowbricks import ArrowbricksError, AuthError, StatementError, TransientError
|
|
179
|
+
|
|
180
|
+
try:
|
|
181
|
+
await cursor.execute(sql)
|
|
182
|
+
except AuthError:
|
|
183
|
+
refresh_credentials()
|
|
184
|
+
except StatementError:
|
|
185
|
+
raise # bad SQL -- don't retry
|
|
186
|
+
except TransientError:
|
|
187
|
+
await asyncio.sleep(5)
|
|
188
|
+
await cursor.execute(sql) # safe to retry
|
|
189
|
+
except ArrowbricksError:
|
|
190
|
+
raise
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
`QueryTimeout` (`total_timeout_s` exceeded -- see Cancellation above) also subclasses `ArrowbricksError`, so `except ArrowbricksError` is a genuine catch-all for every exception this package raises itself.
|
|
194
|
+
|
|
138
195
|
## API
|
|
139
196
|
|
|
140
197
|
- `connect(host, warehouse_id, *, token=None, token_provider=None, ...) -> Connection`
|
|
141
198
|
- `Connection.cursor() -> Cursor`
|
|
142
199
|
- `Connection.client -> DatabricksClient` -- the same client `cursor()` uses, for lower-level access (e.g. `stream_query_json`, `upload_volume_file`).
|
|
143
|
-
- `Cursor.execute(sql, parameters=None, *, row_limit=None, offset=None, catalog=None, schema=None, total_timeout_s=None, prefer_inline=False) -> Cursor` -- submits and waits for the statement, like a real DB-API cursor. `parameters`, if given, is Databricks' own named-parameter format -- `[{"name": ..., "value": ..., "type": ...}]` bound against `:name` markers in `sql`. `prefer_inline=True` tries fetching a small result (well under Databricks' 25 MiB inline cap) in the same round trip as the submission itself, skipping the chunk-fetch entirely
|
|
200
|
+
- `Cursor.execute(sql, parameters=None, *, row_limit=None, offset=None, catalog=None, schema=None, total_timeout_s=None, prefer_inline=False) -> Cursor` -- submits and waits for the statement, like a real DB-API cursor. `parameters`, if given, is Databricks' own named-parameter format -- `[{"name": ..., "value": ..., "type": ...}]` bound against `:name` markers in `sql`. `prefer_inline=True` tries fetching a small result (well under Databricks' 25 MiB inline cap) in the same round trip as the submission itself, skipping the chunk-fetch entirely. If the result turns out too big, the statement fails server-side before ever really executing, and arrowbricks transparently re-runs it the normal way -- safe to double-submit, since nothing committed the first time; a caller who sets `prefer_inline` without actually expecting a small result just pays for the query twice in this case. If instead the result comes back but has a column type this can't convert (nested ARRAY/MAP/STRUCT, VARIANT), the statement already *succeeded* -- arrowbricks does **not** silently re-run it (that could duplicate a write for non-idempotent SQL); it raises `ArrowbricksError` naming the statement instead, so you know it ran and can decide yourself whether re-running is safe. Leave `prefer_inline` off unless you know the result is small, and unless you know the query is a `SELECT` (or otherwise idempotent) if you want the byte-limit fallback's automatic retry to be safe too.
|
|
144
201
|
- `Cursor.execute_streamed(...)` -- same args, but an async generator yielding `HEARTBEAT` while waiting on a slow cold start, then the ready `Cursor` -- for bridging e.g. an SSE connection. Its timeout/heartbeats stop the moment the statement is ready, *before* any chunk has been downloaded -- see `fetchall_streamed` below for the download phase itself.
|
|
145
202
|
- `Cursor.fetchone() -> tuple | None`, `Cursor.fetchmany(size) -> list[tuple]`, `Cursor.fetchall() -> list[tuple]`, and iterating a `Cursor` directly -- row tuples; needs the `arro3` extra.
|
|
146
203
|
- `Cursor.fetchmany_arrow(size) -> Table`, `Cursor.fetchall_arrow() -> Table` -- an Arrow table (implements `__arrow_c_stream__`, so arro3/pyarrow/DuckDB can all consume it directly, zero-copy).
|
|
147
204
|
- `Cursor.fetchall_streamed(*, total_timeout_s=None)` / `Cursor.fetchall_arrow_streamed(*, total_timeout_s=None)` -- like `fetchall()`/`fetchall_arrow()`, but yield `HEARTBEAT` while pulling chunks instead of blocking silently, then the final rows/Table -- for a caller downloading a large result over SSE who needs heartbeats (and a timeout) through the *download*, not just the initial wait. Compose with `execute_streamed` and a shared deadline if you want one combined budget across both phases (see `examples/fastapi_sse_pivot.py`).
|
|
148
205
|
- `Cursor.description` -- DB-API-style `[(name, type_name, None, None, None, None, None), ...]` after `execute()`.
|
|
149
206
|
- `client.stream_query_json(sql, **kwargs)` (or the equivalent free function `stream_query_json(client, sql, **kwargs)`) -- yields `HEARTBEAT`, then each row as a JSON string, as soon as its chunk arrives. Timestamps come out as full ISO-8601, every column key is always present (`"col":null` for a null value, never an omitted key). JSON has no literal for NaN/Infinity/-Infinity, so those come back as `"col":null` by default -- pass `non_finite_floats="string"` to get `"col":"NaN"`/`"col":"Infinity"`/`"col":"-Infinity"` instead if you need to tell them apart from a real NULL.
|
|
150
|
-
- `DatabricksClient(host, warehouse_id, *, token=None, token_provider=None, protocol="thrift", ...)` -- the lower-level client `Connection` wraps. `client.upload_volume_file(volume_path, data)`/`client.delete_volume_file(volume_path)` for the Files API. Pass `protocol="sea"` to opt into the REST Statement Execution API backend instead of the default Thrift one (see above) -- `prefer_inline` (SEA-only) has no effect under `protocol="thrift"` (silent no-op, not an error), since Thrift's own inline-result mechanism already covers that case.
|
|
207
|
+
- `DatabricksClient(host, warehouse_id, *, token=None, token_provider=None, protocol="thrift", on_event=None, retry_attempts=6, retry_max_wait_s=20.0, ...)` -- the lower-level client `Connection` wraps. `client.upload_volume_file(volume_path, data)`/`client.delete_volume_file(volume_path)` for the Files API. Pass `protocol="sea"` to opt into the REST Statement Execution API backend instead of the default Thrift one (see above) -- `prefer_inline` (SEA-only) has no effect under `protocol="thrift"` (silent no-op, not an error), since Thrift's own inline-result mechanism already covers that case. `on_event` -- see Observability below. `retry_attempts`/`retry_max_wait_s` tune the retry policy behind `TransientError`/`AuthError` (see Errors above) -- total attempts (the first try plus `retry_attempts - 1` retries) and the exponential-backoff ceiling in seconds, for every retryable request this client makes (statement submit/poll, chunk-index resolution, chunk download, volume file ops). `retry_attempts` must be at least 1 (raises `ValueError` otherwise).
|
|
151
208
|
- `write_ipc_stream(table, buf)` -- writes any Arrow-C-Data-Interface-compatible object as an uncompressed Arrow-IPC stream (see below).
|
|
152
209
|
- `ReplayableArrowChunk(data: bytes, chunk_index, declared_row_count=None)` -- wraps raw Arrow-IPC stream bytes (e.g. previously downloaded and stored) so they can be read more than once via `__arrow_c_stream__` (a schema peek, then the actual scan -- DuckDB's registration path does this), and `.to_table()` for a one-shot parse. No extra dependency needed.
|
|
153
210
|
|
|
@@ -65,9 +65,14 @@ async for item in client.stream_query_json("SELECT * FROM my_catalog.my_schema.b
|
|
|
65
65
|
|
|
66
66
|
See [`examples/basic.py`](examples/basic.py) for a runnable version,
|
|
67
67
|
[`examples/cursor_paging.py`](examples/cursor_paging.py) for paging a large
|
|
68
|
-
result with `fetchmany`/`fetchmany_arrow` without buffering it all upfront,
|
|
68
|
+
result with `fetchmany`/`fetchmany_arrow` without buffering it all upfront,
|
|
69
69
|
[`examples/azure_auth.py`](examples/azure_auth.py) for a caching
|
|
70
|
-
`token_provider` built on Azure AD (`DefaultAzureCredential
|
|
70
|
+
`token_provider` built on Azure AD (`DefaultAzureCredential`, needs
|
|
71
|
+
`azure-identity`), or [`examples/oauth_m2m_auth.py`](examples/oauth_m2m_auth.py)
|
|
72
|
+
for the same idea using Databricks' own OAuth machine-to-machine
|
|
73
|
+
client-credentials flow instead -- like `databricks-sql-connector`'s
|
|
74
|
+
`auth_type="databricks-oauth"`, but stdlib-only (`urllib.request`), no
|
|
75
|
+
extra dependency.
|
|
71
76
|
|
|
72
77
|
## FastAPI SSE example
|
|
73
78
|
|
|
@@ -123,19 +128,71 @@ arrowbricks has no opinion on *how* you get a token and no cloud-SDK dependency
|
|
|
123
128
|
conn = connect(host=..., warehouse_id=..., token_provider=my_token_provider)
|
|
124
129
|
```
|
|
125
130
|
|
|
131
|
+
## Observability
|
|
132
|
+
|
|
133
|
+
Pass `on_event` to `connect`/`DatabricksClient` to get a `QueryStats` snapshot once per query, at completion (success, cancelled, timed out, or errored):
|
|
134
|
+
|
|
135
|
+
```python
|
|
136
|
+
def log_query(stats):
|
|
137
|
+
print(
|
|
138
|
+
f"{stats.statement_id}: {stats.outcome} in {stats.fetch_s:.2f}s, {stats.bytes_downloaded} bytes, {stats.retry_count} retries"
|
|
139
|
+
)
|
|
140
|
+
|
|
141
|
+
|
|
142
|
+
conn = connect(host=..., warehouse_id=..., token=..., on_event=log_query)
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
`on_event` is sync or async, same as `token_provider`, and applies to every query run through that client (not passed per-call). It's strictly fire-and-forget: any exception it raises is caught and swallowed on the Rust side, never surfaced to your query, and dispatching it never blocks or measurably slows the actual fetch -- a slow or broken logging/metrics callback can't make your queries slower or fail them. `QueryStats` fields: `statement_id: str`, `protocol: "thrift" | "sea"`, `warehouse_wait_s: float` (0.0 if the client already knew the warehouse was running), `submit_to_ready_s: float` (statement submit -> ready to fetch), `fetch_s: float` (download + decode), `num_chunks: int`, `bytes_downloaded: int`, `retry_count: int`, `concurrency_used: int`, and `outcome: "success" | "cancelled" | "timeout" | "error"`.
|
|
146
|
+
|
|
147
|
+
`on_event` fires once the fetch of a given query actually completes (or is abandoned) -- a `Cursor` whose result is never fully drained (e.g. only a partial `fetchmany()`, then abandoned) never fires one, same category as this package's existing documented gap around GC-based cleanup.
|
|
148
|
+
|
|
149
|
+
## Cancellation
|
|
150
|
+
|
|
151
|
+
When a chunk download times out (`total_timeout_s`) or your code cancels the surrounding coroutine (`task.cancel()`/`asyncio.wait_for`) while it's in flight, arrowbricks fires a best-effort server-side cancel in the background (Thrift's `CancelOperation`, or SEA's `POST .../cancel`) so Databricks stops running the query instead of finishing it for nobody. This is fire-and-forget: the `QueryTimeout`/cancellation still reaches you immediately, and the cancel call's own result (success or failure) is never surfaced or awaited.
|
|
152
|
+
|
|
153
|
+
This covers `Cursor.fetchall_streamed()`/`fetchall_arrow_streamed()`, `client.stream_query_json(...)`, and the lower-level `._core.Client`'s own streamed APIs -- everything that wraps the chunk-download phase in the Rust-level heartbeat. Neither `total_timeout_s` nor a bare `task.cancel()` on `Cursor.execute()`/`execute_streamed()` itself (the initial submit/poll wait) triggers a cancel -- by the time a query's chunks are being fetched at all, the statement has typically already finished running server-side, so there's usually nothing left to cancel there in practice.
|
|
154
|
+
|
|
155
|
+
## Errors
|
|
156
|
+
|
|
157
|
+
Every exception arrowbricks raises itself is an `ArrowbricksError` -- a `RuntimeError` subclass, so an `except RuntimeError` written before this hierarchy existed keeps catching everything unchanged. Three subclasses split by what actually changes what your code should do next:
|
|
158
|
+
|
|
159
|
+
- `TransientError` -- a network blip, connection reset, or 5xx that survived every internal retry (`retry_attempts`, exponential backoff up to `retry_max_wait_s` -- see below) before ever reaching you. Back off further and try again later; retrying immediately just repeats what already failed.
|
|
160
|
+
- `AuthError` -- HTTP 401/403, even after every internal retry re-fetched a token from your `token_provider`, *or* your `token_provider` itself raised (e.g. its own OAuth refresh call came back unauthorized). Treat the credential itself as bad/expired, not a transient blip.
|
|
161
|
+
- `StatementError` -- the SQL statement itself failed or was canceled server-side (bad SQL, a permissions error, a warehouse-side query failure). Retrying the identical statement will fail identically; this isn't a transport problem.
|
|
162
|
+
|
|
163
|
+
Anything else (a generic non-auth/non-statement HTTP error, an internal parse/decode failure) raises the plain `ArrowbricksError` base.
|
|
164
|
+
|
|
165
|
+
```python
|
|
166
|
+
from arrowbricks import ArrowbricksError, AuthError, StatementError, TransientError
|
|
167
|
+
|
|
168
|
+
try:
|
|
169
|
+
await cursor.execute(sql)
|
|
170
|
+
except AuthError:
|
|
171
|
+
refresh_credentials()
|
|
172
|
+
except StatementError:
|
|
173
|
+
raise # bad SQL -- don't retry
|
|
174
|
+
except TransientError:
|
|
175
|
+
await asyncio.sleep(5)
|
|
176
|
+
await cursor.execute(sql) # safe to retry
|
|
177
|
+
except ArrowbricksError:
|
|
178
|
+
raise
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
`QueryTimeout` (`total_timeout_s` exceeded -- see Cancellation above) also subclasses `ArrowbricksError`, so `except ArrowbricksError` is a genuine catch-all for every exception this package raises itself.
|
|
182
|
+
|
|
126
183
|
## API
|
|
127
184
|
|
|
128
185
|
- `connect(host, warehouse_id, *, token=None, token_provider=None, ...) -> Connection`
|
|
129
186
|
- `Connection.cursor() -> Cursor`
|
|
130
187
|
- `Connection.client -> DatabricksClient` -- the same client `cursor()` uses, for lower-level access (e.g. `stream_query_json`, `upload_volume_file`).
|
|
131
|
-
- `Cursor.execute(sql, parameters=None, *, row_limit=None, offset=None, catalog=None, schema=None, total_timeout_s=None, prefer_inline=False) -> Cursor` -- submits and waits for the statement, like a real DB-API cursor. `parameters`, if given, is Databricks' own named-parameter format -- `[{"name": ..., "value": ..., "type": ...}]` bound against `:name` markers in `sql`. `prefer_inline=True` tries fetching a small result (well under Databricks' 25 MiB inline cap) in the same round trip as the submission itself, skipping the chunk-fetch entirely
|
|
188
|
+
- `Cursor.execute(sql, parameters=None, *, row_limit=None, offset=None, catalog=None, schema=None, total_timeout_s=None, prefer_inline=False) -> Cursor` -- submits and waits for the statement, like a real DB-API cursor. `parameters`, if given, is Databricks' own named-parameter format -- `[{"name": ..., "value": ..., "type": ...}]` bound against `:name` markers in `sql`. `prefer_inline=True` tries fetching a small result (well under Databricks' 25 MiB inline cap) in the same round trip as the submission itself, skipping the chunk-fetch entirely. If the result turns out too big, the statement fails server-side before ever really executing, and arrowbricks transparently re-runs it the normal way -- safe to double-submit, since nothing committed the first time; a caller who sets `prefer_inline` without actually expecting a small result just pays for the query twice in this case. If instead the result comes back but has a column type this can't convert (nested ARRAY/MAP/STRUCT, VARIANT), the statement already *succeeded* -- arrowbricks does **not** silently re-run it (that could duplicate a write for non-idempotent SQL); it raises `ArrowbricksError` naming the statement instead, so you know it ran and can decide yourself whether re-running is safe. Leave `prefer_inline` off unless you know the result is small, and unless you know the query is a `SELECT` (or otherwise idempotent) if you want the byte-limit fallback's automatic retry to be safe too.
|
|
132
189
|
- `Cursor.execute_streamed(...)` -- same args, but an async generator yielding `HEARTBEAT` while waiting on a slow cold start, then the ready `Cursor` -- for bridging e.g. an SSE connection. Its timeout/heartbeats stop the moment the statement is ready, *before* any chunk has been downloaded -- see `fetchall_streamed` below for the download phase itself.
|
|
133
190
|
- `Cursor.fetchone() -> tuple | None`, `Cursor.fetchmany(size) -> list[tuple]`, `Cursor.fetchall() -> list[tuple]`, and iterating a `Cursor` directly -- row tuples; needs the `arro3` extra.
|
|
134
191
|
- `Cursor.fetchmany_arrow(size) -> Table`, `Cursor.fetchall_arrow() -> Table` -- an Arrow table (implements `__arrow_c_stream__`, so arro3/pyarrow/DuckDB can all consume it directly, zero-copy).
|
|
135
192
|
- `Cursor.fetchall_streamed(*, total_timeout_s=None)` / `Cursor.fetchall_arrow_streamed(*, total_timeout_s=None)` -- like `fetchall()`/`fetchall_arrow()`, but yield `HEARTBEAT` while pulling chunks instead of blocking silently, then the final rows/Table -- for a caller downloading a large result over SSE who needs heartbeats (and a timeout) through the *download*, not just the initial wait. Compose with `execute_streamed` and a shared deadline if you want one combined budget across both phases (see `examples/fastapi_sse_pivot.py`).
|
|
136
193
|
- `Cursor.description` -- DB-API-style `[(name, type_name, None, None, None, None, None), ...]` after `execute()`.
|
|
137
194
|
- `client.stream_query_json(sql, **kwargs)` (or the equivalent free function `stream_query_json(client, sql, **kwargs)`) -- yields `HEARTBEAT`, then each row as a JSON string, as soon as its chunk arrives. Timestamps come out as full ISO-8601, every column key is always present (`"col":null` for a null value, never an omitted key). JSON has no literal for NaN/Infinity/-Infinity, so those come back as `"col":null` by default -- pass `non_finite_floats="string"` to get `"col":"NaN"`/`"col":"Infinity"`/`"col":"-Infinity"` instead if you need to tell them apart from a real NULL.
|
|
138
|
-
- `DatabricksClient(host, warehouse_id, *, token=None, token_provider=None, protocol="thrift", ...)` -- the lower-level client `Connection` wraps. `client.upload_volume_file(volume_path, data)`/`client.delete_volume_file(volume_path)` for the Files API. Pass `protocol="sea"` to opt into the REST Statement Execution API backend instead of the default Thrift one (see above) -- `prefer_inline` (SEA-only) has no effect under `protocol="thrift"` (silent no-op, not an error), since Thrift's own inline-result mechanism already covers that case.
|
|
195
|
+
- `DatabricksClient(host, warehouse_id, *, token=None, token_provider=None, protocol="thrift", on_event=None, retry_attempts=6, retry_max_wait_s=20.0, ...)` -- the lower-level client `Connection` wraps. `client.upload_volume_file(volume_path, data)`/`client.delete_volume_file(volume_path)` for the Files API. Pass `protocol="sea"` to opt into the REST Statement Execution API backend instead of the default Thrift one (see above) -- `prefer_inline` (SEA-only) has no effect under `protocol="thrift"` (silent no-op, not an error), since Thrift's own inline-result mechanism already covers that case. `on_event` -- see Observability below. `retry_attempts`/`retry_max_wait_s` tune the retry policy behind `TransientError`/`AuthError` (see Errors above) -- total attempts (the first try plus `retry_attempts - 1` retries) and the exponential-backoff ceiling in seconds, for every retryable request this client makes (statement submit/poll, chunk-index resolution, chunk download, volume file ops). `retry_attempts` must be at least 1 (raises `ValueError` otherwise).
|
|
139
196
|
- `write_ipc_stream(table, buf)` -- writes any Arrow-C-Data-Interface-compatible object as an uncompressed Arrow-IPC stream (see below).
|
|
140
197
|
- `ReplayableArrowChunk(data: bytes, chunk_index, declared_row_count=None)` -- wraps raw Arrow-IPC stream bytes (e.g. previously downloaded and stored) so they can be read more than once via `__arrow_c_stream__` (a schema peek, then the actual scan -- DuckDB's registration path does this), and `.to_table()` for a one-shot parse. No extra dependency needed.
|
|
141
198
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
[project]
|
|
2
2
|
name = "arrowbricks"
|
|
3
|
-
version = "3.0
|
|
3
|
+
version = "3.1.0"
|
|
4
4
|
description = "Runs SQL against a Databricks SQL warehouse via the Statement Execution API and hands you the result as Arrow -- a DB-API-ish Cursor (fetchone/fetchmany/fetchall/fetchall_arrow) or NDJSON streaming. Rust/PyO3 core throughout -- zero required runtime dependencies."
|
|
5
5
|
readme = "README.md"
|
|
6
6
|
license = "MIT"
|
|
@@ -241,7 +241,7 @@ dependencies = [
|
|
|
241
241
|
|
|
242
242
|
[[package]]
|
|
243
243
|
name = "arrowbricks_core"
|
|
244
|
-
version = "3.0
|
|
244
|
+
version = "3.1.0"
|
|
245
245
|
dependencies = [
|
|
246
246
|
"arrow",
|
|
247
247
|
"arrow-json",
|
|
@@ -249,6 +249,7 @@ dependencies = [
|
|
|
249
249
|
"chrono",
|
|
250
250
|
"hyper-rustls",
|
|
251
251
|
"lz4_flex",
|
|
252
|
+
"proptest",
|
|
252
253
|
"pyo3",
|
|
253
254
|
"pyo3-arrow",
|
|
254
255
|
"pyo3-async-runtimes",
|
|
@@ -304,6 +305,21 @@ version = "0.23.1"
|
|
|
304
305
|
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
305
306
|
checksum = "ac07cdecf99051d9a5238b80f35af32cdeba5b336e55d957b318b50137e18da5"
|
|
306
307
|
|
|
308
|
+
[[package]]
|
|
309
|
+
name = "bit-set"
|
|
310
|
+
version = "0.8.0"
|
|
311
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
312
|
+
checksum = "08807e080ed7f9d5433fa9b275196cfc35414f66a0c79d864dc51a0d825231a3"
|
|
313
|
+
dependencies = [
|
|
314
|
+
"bit-vec",
|
|
315
|
+
]
|
|
316
|
+
|
|
317
|
+
[[package]]
|
|
318
|
+
name = "bit-vec"
|
|
319
|
+
version = "0.8.0"
|
|
320
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
321
|
+
checksum = "5e764a1d40d510daf35e07be9eb06e75770908c27d411ee6c92109c9840eaaf7"
|
|
322
|
+
|
|
307
323
|
[[package]]
|
|
308
324
|
name = "bitflags"
|
|
309
325
|
version = "2.13.1"
|
|
@@ -458,6 +474,22 @@ version = "1.0.2"
|
|
|
458
474
|
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
459
475
|
checksum = "877a4ace8713b0bcf2a4e7eec82529c029f1d0619886d18145fea96c3ffe5c0f"
|
|
460
476
|
|
|
477
|
+
[[package]]
|
|
478
|
+
name = "errno"
|
|
479
|
+
version = "0.3.14"
|
|
480
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
481
|
+
checksum = "39cab71617ae0d63f51a36d69f866391735b51691dbda63cf6f96d042b63efeb"
|
|
482
|
+
dependencies = [
|
|
483
|
+
"libc",
|
|
484
|
+
"windows-sys 0.61.2",
|
|
485
|
+
]
|
|
486
|
+
|
|
487
|
+
[[package]]
|
|
488
|
+
name = "fastrand"
|
|
489
|
+
version = "2.5.0"
|
|
490
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
491
|
+
checksum = "da7c62ceae207dd37ea5b845da6a0696c799f85e97da1ab5b7910be3c1c80223"
|
|
492
|
+
|
|
461
493
|
[[package]]
|
|
462
494
|
name = "find-msvc-tools"
|
|
463
495
|
version = "0.1.10"
|
|
@@ -1038,6 +1070,12 @@ version = "0.2.16"
|
|
|
1038
1070
|
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1039
1071
|
checksum = "b6d2cec3eae94f9f509c767b45932f1ada8350c4bdb85af2fcab4a3c14807981"
|
|
1040
1072
|
|
|
1073
|
+
[[package]]
|
|
1074
|
+
name = "linux-raw-sys"
|
|
1075
|
+
version = "0.12.1"
|
|
1076
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1077
|
+
checksum = "32a66949e030da00e8c7d4434b251670a91556f4144941d37452769c25d58a53"
|
|
1078
|
+
|
|
1041
1079
|
[[package]]
|
|
1042
1080
|
name = "litemap"
|
|
1043
1081
|
version = "0.8.2"
|
|
@@ -1232,6 +1270,15 @@ dependencies = [
|
|
|
1232
1270
|
"zerovec",
|
|
1233
1271
|
]
|
|
1234
1272
|
|
|
1273
|
+
[[package]]
|
|
1274
|
+
name = "ppv-lite86"
|
|
1275
|
+
version = "0.2.21"
|
|
1276
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1277
|
+
checksum = "85eae3c4ed2f50dcfe72643da4befc30deadb458a9b590d720cde2f2b1e97da9"
|
|
1278
|
+
dependencies = [
|
|
1279
|
+
"zerocopy",
|
|
1280
|
+
]
|
|
1281
|
+
|
|
1235
1282
|
[[package]]
|
|
1236
1283
|
name = "proc-macro2"
|
|
1237
1284
|
version = "1.0.107"
|
|
@@ -1241,6 +1288,25 @@ dependencies = [
|
|
|
1241
1288
|
"unicode-ident",
|
|
1242
1289
|
]
|
|
1243
1290
|
|
|
1291
|
+
[[package]]
|
|
1292
|
+
name = "proptest"
|
|
1293
|
+
version = "1.11.0"
|
|
1294
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1295
|
+
checksum = "4b45fcc2344c680f5025fe57779faef368840d0bd1f42f216291f0dc4ace4744"
|
|
1296
|
+
dependencies = [
|
|
1297
|
+
"bit-set",
|
|
1298
|
+
"bit-vec",
|
|
1299
|
+
"bitflags",
|
|
1300
|
+
"num-traits",
|
|
1301
|
+
"rand",
|
|
1302
|
+
"rand_chacha",
|
|
1303
|
+
"rand_xorshift",
|
|
1304
|
+
"regex-syntax",
|
|
1305
|
+
"rusty-fork",
|
|
1306
|
+
"tempfile",
|
|
1307
|
+
"unarray",
|
|
1308
|
+
]
|
|
1309
|
+
|
|
1244
1310
|
[[package]]
|
|
1245
1311
|
name = "pyo3"
|
|
1246
1312
|
version = "0.29.2"
|
|
@@ -1346,6 +1412,12 @@ dependencies = [
|
|
|
1346
1412
|
"serde",
|
|
1347
1413
|
]
|
|
1348
1414
|
|
|
1415
|
+
[[package]]
|
|
1416
|
+
name = "quick-error"
|
|
1417
|
+
version = "1.2.3"
|
|
1418
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1419
|
+
checksum = "a1d01941d82fa2ab50be1e79e6714289dd7cde78eba4c074bc5a4374f650dfe0"
|
|
1420
|
+
|
|
1349
1421
|
[[package]]
|
|
1350
1422
|
name = "quote"
|
|
1351
1423
|
version = "1.0.47"
|
|
@@ -1361,6 +1433,44 @@ version = "5.3.0"
|
|
|
1361
1433
|
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1362
1434
|
checksum = "69cdb34c158ceb288df11e18b4bd39de994f6657d83847bdffdbd7f346754b0f"
|
|
1363
1435
|
|
|
1436
|
+
[[package]]
|
|
1437
|
+
name = "rand"
|
|
1438
|
+
version = "0.9.5"
|
|
1439
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1440
|
+
checksum = "b9ef1d0d795eb7d84685bca4f72f3649f064e6641543d3a8c415898726a57b41"
|
|
1441
|
+
dependencies = [
|
|
1442
|
+
"rand_chacha",
|
|
1443
|
+
"rand_core",
|
|
1444
|
+
]
|
|
1445
|
+
|
|
1446
|
+
[[package]]
|
|
1447
|
+
name = "rand_chacha"
|
|
1448
|
+
version = "0.9.0"
|
|
1449
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1450
|
+
checksum = "d3022b5f1df60f26e1ffddd6c66e8aa15de382ae63b3a0c1bfc0e4d3e3f325cb"
|
|
1451
|
+
dependencies = [
|
|
1452
|
+
"ppv-lite86",
|
|
1453
|
+
"rand_core",
|
|
1454
|
+
]
|
|
1455
|
+
|
|
1456
|
+
[[package]]
|
|
1457
|
+
name = "rand_core"
|
|
1458
|
+
version = "0.9.5"
|
|
1459
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1460
|
+
checksum = "76afc826de14238e6e8c374ddcc1fa19e374fd8dd986b0d2af0d02377261d83c"
|
|
1461
|
+
dependencies = [
|
|
1462
|
+
"getrandom 0.3.4",
|
|
1463
|
+
]
|
|
1464
|
+
|
|
1465
|
+
[[package]]
|
|
1466
|
+
name = "rand_xorshift"
|
|
1467
|
+
version = "0.4.0"
|
|
1468
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1469
|
+
checksum = "513962919efc330f829edb2535844d1b912b0fbe2ca165d613e4e8788bb05a5a"
|
|
1470
|
+
dependencies = [
|
|
1471
|
+
"rand_core",
|
|
1472
|
+
]
|
|
1473
|
+
|
|
1364
1474
|
[[package]]
|
|
1365
1475
|
name = "rawpointer"
|
|
1366
1476
|
version = "0.2.1"
|
|
@@ -1462,6 +1572,19 @@ dependencies = [
|
|
|
1462
1572
|
"semver",
|
|
1463
1573
|
]
|
|
1464
1574
|
|
|
1575
|
+
[[package]]
|
|
1576
|
+
name = "rustix"
|
|
1577
|
+
version = "1.1.4"
|
|
1578
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1579
|
+
checksum = "b6fe4565b9518b83ef4f91bb47ce29620ca828bd32cb7e408f0062e9930ba190"
|
|
1580
|
+
dependencies = [
|
|
1581
|
+
"bitflags",
|
|
1582
|
+
"errno",
|
|
1583
|
+
"libc",
|
|
1584
|
+
"linux-raw-sys",
|
|
1585
|
+
"windows-sys 0.61.2",
|
|
1586
|
+
]
|
|
1587
|
+
|
|
1465
1588
|
[[package]]
|
|
1466
1589
|
name = "rustls"
|
|
1467
1590
|
version = "0.23.43"
|
|
@@ -1542,6 +1665,18 @@ version = "1.0.23"
|
|
|
1542
1665
|
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1543
1666
|
checksum = "cf54715a573b99ac80df0bc206da022bcd442c974952c7b9720069370852e21f"
|
|
1544
1667
|
|
|
1668
|
+
[[package]]
|
|
1669
|
+
name = "rusty-fork"
|
|
1670
|
+
version = "0.3.1"
|
|
1671
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1672
|
+
checksum = "cc6bf79ff24e648f6da1f8d1f011e9cac26491b619e6b9280f2b47f1774e6ee2"
|
|
1673
|
+
dependencies = [
|
|
1674
|
+
"fnv",
|
|
1675
|
+
"quick-error",
|
|
1676
|
+
"tempfile",
|
|
1677
|
+
"wait-timeout",
|
|
1678
|
+
]
|
|
1679
|
+
|
|
1545
1680
|
[[package]]
|
|
1546
1681
|
name = "ryu"
|
|
1547
1682
|
version = "1.0.23"
|
|
@@ -1760,6 +1895,19 @@ version = "0.13.5"
|
|
|
1760
1895
|
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1761
1896
|
checksum = "adb6935a6f5c20170eeceb1a3835a49e12e19d792f6dd344ccc76a985ca5a6ca"
|
|
1762
1897
|
|
|
1898
|
+
[[package]]
|
|
1899
|
+
name = "tempfile"
|
|
1900
|
+
version = "3.27.0"
|
|
1901
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1902
|
+
checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd"
|
|
1903
|
+
dependencies = [
|
|
1904
|
+
"fastrand",
|
|
1905
|
+
"getrandom 0.3.4",
|
|
1906
|
+
"once_cell",
|
|
1907
|
+
"rustix",
|
|
1908
|
+
"windows-sys 0.61.2",
|
|
1909
|
+
]
|
|
1910
|
+
|
|
1763
1911
|
[[package]]
|
|
1764
1912
|
name = "thiserror"
|
|
1765
1913
|
version = "1.0.69"
|
|
@@ -1945,6 +2093,12 @@ version = "2.1.3"
|
|
|
1945
2093
|
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1946
2094
|
checksum = "8464ec13c3691491391d9fce00f6416c9a48e46972f72d7865688be2080192c9"
|
|
1947
2095
|
|
|
2096
|
+
[[package]]
|
|
2097
|
+
name = "unarray"
|
|
2098
|
+
version = "0.1.4"
|
|
2099
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
2100
|
+
checksum = "eaea85b334db583fe3274d12b4cd1880032beab409c0d774be044d4480ab9a94"
|
|
2101
|
+
|
|
1948
2102
|
[[package]]
|
|
1949
2103
|
name = "unicode-ident"
|
|
1950
2104
|
version = "1.0.24"
|
|
@@ -1993,6 +2147,15 @@ version = "0.9.5"
|
|
|
1993
2147
|
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
1994
2148
|
checksum = "0b928f33d975fc6ad9f86c8f283853ad26bdd5b10b7f1542aa2fa15e2289105a"
|
|
1995
2149
|
|
|
2150
|
+
[[package]]
|
|
2151
|
+
name = "wait-timeout"
|
|
2152
|
+
version = "0.2.1"
|
|
2153
|
+
source = "registry+https://github.com/rust-lang/crates.io-index"
|
|
2154
|
+
checksum = "09ac3b126d3914f9849036f826e054cbabdc8519970b8998ddaf3b5bd3c65f11"
|
|
2155
|
+
dependencies = [
|
|
2156
|
+
"libc",
|
|
2157
|
+
]
|
|
2158
|
+
|
|
1996
2159
|
[[package]]
|
|
1997
2160
|
name = "walkdir"
|
|
1998
2161
|
version = "2.5.0"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
[package]
|
|
2
2
|
name = "arrowbricks_core"
|
|
3
|
-
version = "3.0
|
|
3
|
+
version = "3.1.0"
|
|
4
4
|
edition = "2024"
|
|
5
5
|
readme = "README.md"
|
|
6
6
|
|
|
@@ -75,6 +75,20 @@ extension-module = ["pyo3/extension-module"]
|
|
|
75
75
|
[dev-dependencies]
|
|
76
76
|
tokio = { version = "1.53.1", features = ["rt-multi-thread", "macros", "test-util", "net", "io-util"] }
|
|
77
77
|
wiremock = "0.6.5"
|
|
78
|
+
# Property-based fuzzing for the hand-rolled wire-protocol parsers
|
|
79
|
+
# (thrift.rs's Reader, json_convert.rs's STRUCT type_text tokenizer,
|
|
80
|
+
# client.rs's decompress_lz4_frame) -- dev-only, runs inside plain
|
|
81
|
+
# `cargo test` via the `proptest!` macro (`#[test]` under the hood, no
|
|
82
|
+
# separate fuzz-target/corpus infra the way cargo-fuzz/afl need, and no
|
|
83
|
+
# nightly toolchain requirement either). Never touches the shipped wheel's
|
|
84
|
+
# dependency graph -- `[dev-dependencies]` are only ever linked into `cargo
|
|
85
|
+
# test`/`cargo bench` binaries, never into the `cdylib` maturin actually
|
|
86
|
+
# builds (see this crate's own `extension-module` feature/profile split
|
|
87
|
+
# above for how that build works). Default config (256 cases/property) is
|
|
88
|
+
# left alone; per-property `ProptestConfig` below trims a couple of the
|
|
89
|
+
# slower ones (deep nesting, wide byte vectors) so `cargo test` doesn't
|
|
90
|
+
# become noticeably slower overall.
|
|
91
|
+
proptest = "1"
|
|
78
92
|
|
|
79
93
|
# Currently clean on all three by luck, not enforcement -- pinning them here
|
|
80
94
|
# so a future PR that introduces one fails CI instead of silently landing.
|
|
@@ -66,7 +66,7 @@ unlike the table object itself.
|
|
|
66
66
|
|
|
67
67
|
## API
|
|
68
68
|
|
|
69
|
-
- `Client(host, warehouse_id, *, token=None, token_provider=None, chunk_fetch_concurrency=64, http_timeout=60.0, wait_timeout="30s", warehouse_start_timeout=300.0, warehouse_confirmed_running_ttl_s=30.0, compress_results=True, protocol="thrift")` -- exactly one of `token`/`token_provider`. `token_provider` is a callable (sync or async) returning a token string, called fresh on every request, no caching. `compress_results` requests LZ4-compressed cloud-fetch chunks (see above); set `False` to opt out. `protocol="thrift"` (the default, as of this crate's own real-workspace benchmarking -- see `AGENTS.md`'s design-invariant entry) speaks the same HiveServer2-compatible Thrift-over-HTTPS protocol `databricks-sql-connector` uses by default (`thrift.rs`, a hand-rolled `TBinaryProtocol` reader/writer -- no new Cargo dependency) -- measurably faster for small queries (its `ExecuteStatement` RPC can return a small result inline via `getDirectResults`, in the same call that submits the statement) at the cost of `prefer_inline` becoming a silent no-op (Thrift has no INLINE-disposition equivalent, and doesn't need one); never slower than SEA on any query shape tested. `protocol="sea"` instead talks to the REST Statement Execution API -- still fully supported, opt in explicitly if you have a reason to prefer it. On par with SEA for a large, multi-chunk result too -- `run_thrift_fetch_loop` (`pipeline.rs`) fans its chunk downloads out across a `chunk_fetch_concurrency`-sized worker pool spanning the *whole* result (pipelined with the sequential `FetchResults` discovery calls, not serialized behind them), the same concurrency shape as SEA's own `fetch_chunks_with_backpressure` (see `AGENTS.md`'s own entry on this).
|
|
69
|
+
- `Client(host, warehouse_id, *, token=None, token_provider=None, chunk_fetch_concurrency=64, http_timeout=60.0, wait_timeout="30s", warehouse_start_timeout=300.0, warehouse_confirmed_running_ttl_s=30.0, compress_results=True, protocol="thrift", retry_attempts=6, retry_max_wait_s=20.0)` -- exactly one of `token`/`token_provider`. `retry_attempts`/`retry_max_wait_s` tune the retry policy (total attempts, exponential-backoff ceiling in seconds) behind every retryable request this client makes; `retry_attempts` must be at least 1. `token_provider` is a callable (sync or async) returning a token string, called fresh on every request, no caching. `compress_results` requests LZ4-compressed cloud-fetch chunks (see above); set `False` to opt out. `protocol="thrift"` (the default, as of this crate's own real-workspace benchmarking -- see `AGENTS.md`'s design-invariant entry) speaks the same HiveServer2-compatible Thrift-over-HTTPS protocol `databricks-sql-connector` uses by default (`thrift.rs`, a hand-rolled `TBinaryProtocol` reader/writer -- no new Cargo dependency) -- measurably faster for small queries (its `ExecuteStatement` RPC can return a small result inline via `getDirectResults`, in the same call that submits the statement) at the cost of `prefer_inline` becoming a silent no-op (Thrift has no INLINE-disposition equivalent, and doesn't need one); never slower than SEA on any query shape tested. `protocol="sea"` instead talks to the REST Statement Execution API -- still fully supported, opt in explicitly if you have a reason to prefer it. On par with SEA for a large, multi-chunk result too -- `run_thrift_fetch_loop` (`pipeline.rs`) fans its chunk downloads out across a `chunk_fetch_concurrency`-sized worker pool spanning the *whole* result (pipelined with the sequential `FetchResults` discovery calls, not serialized behind them), the same concurrency shape as SEA's own `fetch_chunks_with_backpressure` (see `AGENTS.md`'s own entry on this).
|
|
70
70
|
- `Client.execute(statement, *, catalog=None, schema=None, parameters=None, prefer_inline=False) -> ResultSet` -- submits and starts background chunk fetching without pulling anything yet. `parameters` is Databricks' own named-parameter format (`[{"name":..., "value":..., "type":...}]`), passed straight through. `prefer_inline=True` submits with `disposition=INLINE, format=JSON_ARRAY` instead, for a caller who expects a small (well under Databricks' 25 MiB inline cap) result and wants to skip the chunk-fetch round trip -- on a result too big for INLINE, or containing a column type `json_convert.rs` doesn't map (STRUCT/ARRAY-of-STRUCT/MAP/VARIANT), it transparently falls back to a second, normal `execute()` and the query runs twice. See `AGENTS.md`'s "Design invariants" section (in the root package) for the full reasoning and real-workspace verification behind this.
|
|
71
71
|
- `ResultSet.fetchmany_arrow(n) -> Table` -- pulls/decodes only as many chunks as needed for `n` rows, buffering the rest; may return fewer than `n` once exhausted.
|
|
72
72
|
- `ResultSet.fetchall_arrow() -> Table` -- drains everything remaining.
|