parafetch 0.2.0__tar.gz → 0.2.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (38) hide show
  1. {parafetch-0.2.0 → parafetch-0.2.1}/Cargo.lock +1 -1
  2. {parafetch-0.2.0 → parafetch-0.2.1}/Cargo.toml +1 -1
  3. {parafetch-0.2.0 → parafetch-0.2.1}/PKG-INFO +219 -7
  4. {parafetch-0.2.0 → parafetch-0.2.1}/README.md +218 -6
  5. parafetch-0.2.1/dist/parafetch-0.2.0.tar.gz +0 -0
  6. parafetch-0.2.1/dist/parafetch-0.2.1-cp312-abi3-win_amd64.whl +0 -0
  7. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/_core.pdb +0 -0
  8. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/_core.pyi +4 -1
  9. parafetch-0.2.1/python/parafetch/_proxy.py +74 -0
  10. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/rpc.py +11 -4
  11. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/sqlite.py +9 -0
  12. {parafetch-0.2.0 → parafetch-0.2.1}/src/http.rs +20 -7
  13. {parafetch-0.2.0 → parafetch-0.2.1}/src/lib.rs +1 -0
  14. parafetch-0.2.1/src/proxy.rs +140 -0
  15. {parafetch-0.2.0 → parafetch-0.2.1}/src/sqlite.rs +16 -5
  16. parafetch-0.2.1/tests/test_proxy.py +155 -0
  17. {parafetch-0.2.0 → parafetch-0.2.1}/tests/test_sqlite.py +38 -0
  18. {parafetch-0.2.0 → parafetch-0.2.1}/RELEASING.md +0 -0
  19. {parafetch-0.2.0 → parafetch-0.2.1}/benchmarks/bench_sqlite.py +0 -0
  20. {parafetch-0.2.0 → parafetch-0.2.1}/dist/parafetch-0.1.0-cp312-abi3-win_amd64.whl +0 -0
  21. {parafetch-0.2.0 → parafetch-0.2.1}/dist/parafetch-0.1.0.tar.gz +0 -0
  22. {parafetch-0.2.0 → parafetch-0.2.1}/dist/parafetch-0.2.0-cp312-abi3-win_amd64.whl +0 -0
  23. {parafetch-0.2.0 → parafetch-0.2.1}/pyproject.toml +0 -0
  24. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/__init__.py +0 -0
  25. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/_errors.py +0 -0
  26. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/_utils.py +0 -0
  27. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/auth.py +0 -0
  28. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/parallel.py +0 -0
  29. {parafetch-0.2.0 → parafetch-0.2.1}/python/parafetch/py.typed +0 -0
  30. {parafetch-0.2.0 → parafetch-0.2.1}/scripts/build_windows.ps1 +0 -0
  31. {parafetch-0.2.0 → parafetch-0.2.1}/src/convert.rs +0 -0
  32. {parafetch-0.2.0 → parafetch-0.2.1}/src/negotiate.rs +0 -0
  33. {parafetch-0.2.0 → parafetch-0.2.1}/src/parallel.rs +0 -0
  34. {parafetch-0.2.0 → parafetch-0.2.1}/src/signals.rs +0 -0
  35. {parafetch-0.2.0 → parafetch-0.2.1}/src/xmlrpc.rs +0 -0
  36. {parafetch-0.2.0 → parafetch-0.2.1}/tests/test_negotiate.py +0 -0
  37. {parafetch-0.2.0 → parafetch-0.2.1}/tests/test_parallel.py +0 -0
  38. {parafetch-0.2.0 → parafetch-0.2.1}/tests/test_rpc.py +0 -0
@@ -1201,7 +1201,7 @@ dependencies = [
1201
1201
 
1202
1202
  [[package]]
1203
1203
  name = "parafetch"
1204
- version = "0.2.0"
1204
+ version = "0.2.1"
1205
1205
  dependencies = [
1206
1206
  "arrow",
1207
1207
  "base64 0.22.1",
@@ -1,6 +1,6 @@
1
1
  [package]
2
2
  name = "parafetch"
3
- version = "0.2.0"
3
+ version = "0.2.1"
4
4
  edition = "2021"
5
5
  rust-version = "1.86"
6
6
  description = "Rust-powered parallel data access for Python: fast SQLite -> Arrow, parallel map, batched REST / JSON-RPC / XML-RPC."
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: parafetch
3
- Version: 0.2.0
3
+ Version: 0.2.1
4
4
  Classifier: Programming Language :: Rust
5
5
  Classifier: Programming Language :: Python :: 3
6
6
  Classifier: Programming Language :: Python :: 3.12
@@ -41,10 +41,12 @@ Rust with Python's GIL released. Your Python code only sees ordinary lists, dict
41
41
 
42
42
  - [Installation](#installation)
43
43
  - [Quick start](#quick-start)
44
+ - [How it works (diagrams)](#how-it-works)
44
45
  - [1. Reading SQLite fast — `read_sqlite`](#1-reading-sqlite-fast)
45
46
  - [2. Parallelising one-at-a-time functions — `parallel_map`](#2-parallelising-one-at-a-time-functions)
46
47
  - [3. Fetching from size-limited APIs — REST, JSON-RPC, XML-RPC](#3-fetching-from-size-limited-apis)
47
48
  - [Authentication — Basic, Bearer, Windows Negotiate (Kerberos/NTLM)](#authentication)
49
+ - [Proxies](#proxies)
48
50
  - [Error handling](#error-handling)
49
51
  - [How parafetch deals with the GIL](#how-parafetch-deals-with-the-gil)
50
52
  - [API reference](#api-reference)
@@ -114,12 +116,159 @@ trades = api.fetch_batched(trade_ids, batch_size=50, path="/trades", id_param="i
114
116
 
115
117
  ---
116
118
 
119
+ ## How it works
120
+
121
+ Every feature is split into a thin **Python** layer and a **Rust** engine. Python handles your
122
+ arguments and the final objects you get back. Rust does the heavy lifting in between, with the
123
+ GIL **released**, so it runs on many CPU cores at once.
124
+
125
+ In the diagrams below:
126
+
127
+ ```text
128
+ [PY] runs in Python (holds the GIL: one thread at a time)
129
+ [RS] runs in Rust (GIL released: truly parallel, on all cores)
130
+ ═══ the moment control crosses between Python and Rust
131
+ ```
132
+
133
+ ### The big picture
134
+
135
+ ```text
136
+ your Python code
137
+ │
138
+ ┌────────────────────────┼─────────────────────────────┐
139
+ ▼ ▼ ▼
140
+ [PY] read_sqlite() [PY] parallel_map() [PY] HttpClient / JsonRpcClient
141
+ read_sqlite_pandas() parallel_dict() XmlRpcClient
142
+ read_sqlite_arrow() batched_map() fetch_batched() …
143
+ check arguments split into items split IDs into batches,
144
+ build request bodies
145
+ │ │ │
146
+ ════════╪════════ PYTHON ──▶ RUST (GIL released) ═════════════╪════════════════
147
+ ▼ ▼ ▼
148
+ [RS] SQLite reader [RS] thread scheduler [RS] tokio HTTP engine
149
+ • N threads, one • N OS threads • connection pool, TLS
150
+ connection each • retries + back-off • concurrency + rate limit
151
+ • decode rows • rate limit • retries, Retry-After
152
+ • build Arrow • collect results • proxies, Basic/Negotiate auth
153
+ columns in input order • parse JSON / XML-RPC
154
+ │ │ ▲ │
155
+ │ ▼ │ (per item) │
156
+ │ [PY] YOUR function runs │
157
+ │ (the only Python step; │
158
+ │ GIL held just for this) │
159
+ ════════╪════════ RUST ──▶ PYTHON ════════════════════════════╪════════════════
160
+ ▼ ▼ ▼
161
+ [PY] polars / pandas / [PY] list of results [PY] flatten batches into one
162
+ pyarrow (zero copy) or ParallelError list / dict, or BatchError
163
+ ```
164
+
165
+ ### Where Rust takes over
166
+
167
+ | step | `read_sqlite` | `parallel_map` (thread mode) | `fetch_batched` / `call_batched` |
168
+ |---|---|---|---|
169
+ | prepare | [PY] validate arguments | [PY] collect items | [PY] split IDs, build request bodies |
170
+ | plan | [RS] find rowid range, split it | [RS] start N worker threads | [RS] queue requests |
171
+ | heavy work | [RS] N threads read SQLite → Arrow | [PY] **your function** (GIL held only while it runs Python code) | [RS] send, retry, authenticate, proxy |
172
+ | waiting | none (CPU-bound) | [RS] back-off sleeps, rate limiting | [RS] network waits, back-off, rate limiting |
173
+ | decode | [RS] typed Arrow buffers | none | [RS] parse JSON / XML-RPC |
174
+ | hand back | [PY] polars/pandas/pyarrow (zero copy) | [PY] results in input order | [PY] build dicts/lists, flatten |
175
+
176
+ `parallel_map` is the only feature where Python code (yours) runs in the middle. That's why it
177
+ speeds up functions that wait on I/O (the GIL is free while they wait) but not pure-Python number
178
+ crunching. For that, use `mode="process"`.
179
+
180
+ ### `read_sqlite`: parallel range scan into Arrow
181
+
182
+ ```text
183
+ [PY] read_sqlite("market.db", "prices") # partitions = CPU count
184
+ ═══════════════════════════ PYTHON ──▶ RUST ═════════════════════════════════
185
+ [RS] SELECT MIN(rowid), MAX(rowid) ─▶ 1 … 2,000,000 → split into ranges
186
+ │
187
+ ┌────────────────┬───────────────────────┼───────────────┬──────────────┐
188
+ ▼ ▼ ▼ ▼
189
+ [RS] thread 1 [RS] thread 2 [RS] thread 3 [RS] thread 4
190
+ own read-only own read-only … …
191
+ connection connection
192
+ rowid 1–500k 500k–1M 1M–1.5M 1.5M–2M
193
+ │ │ │ │
194
+ ▼ ▼ ▼ ▼
195
+ [RS] values written straight into typed Arrow buffers (int64, float64, string …)
196
+ no Python object is ever created per row or per cell
197
+ └────────────────┴───────────┬───────────┴───────────────┘
198
+ ▼ joined in rowid order
199
+ ═══════════════════════════ RUST ──▶ PYTHON ═════════════════════════════════
200
+ [PY] polars.DataFrame (zero copy: the Arrow memory is shared)
201
+ ```
202
+
203
+ ### `parallel_map`: why threads help I/O-bound functions
204
+
205
+ ```text
206
+ [PY] parallel_map(get_price, isins, workers=3)
207
+ ═══════════════════════════ PYTHON ──▶ RUST ═════════════════════════════════
208
+ [RS] 3 OS threads pull items from a shared queue
209
+
210
+ time ───────────────────────────────────────────────────▶
211
+ thread 1 ▓░░░░░░░ waiting for DB ░░░░░░░▓ ▓░░░░░░░ waiting ░░░░░▓
212
+ thread 2 ▓░░░░░░░ waiting for DB ░░░░░░░▓ ▓░░░░░░░ waiting ░░░░░▓
213
+ thread 3 ▓░░░░░░░ waiting for DB ░░░░░░░▓ ▓░░░░░░░ waiting ░░░░▓
214
+
215
+ ▓ = [PY] your function's Python code: holds the GIL, but only briefly
216
+ ░ = waiting on network / DB: GIL released, so the other threads run
217
+ → 3 workers ≈ 3× faster, 32 workers ≈ 32× (up to what the server allows)
218
+ [RS] between items: retry back-off sleeps and rate-limit waits, GIL released
219
+ ═══════════════════════════ RUST ──▶ PYTHON ═════════════════════════════════
220
+ [PY] results in input order (or ParallelError with the partial results)
221
+
222
+ CPU-bound Python is all ▓ and threads can't overlap it → use mode="process"
223
+ (one GIL per process) or mode="interpreter" on Python 3.14+.
224
+ ```
225
+
226
+ ### `fetch_batched` / `call_batched`: size-limited APIs
227
+
228
+ ```text
229
+ [PY] api.fetch_batched(10_000 ids, batch_size=50)
230
+ [PY] split into 200 batches → build 200 request bodies ("$ids" filled in)
231
+ ═══════════════════════════ PYTHON ──▶ RUST ═════════════════════════════════
232
+ [RS] ┌── at most `concurrency` in flight ── at most `rate_limit` per second ──┐
233
+ [RS] │ POST /trades {"ids": [1 … 50]} ───▶ 200 OK │
234
+ [RS] │ POST /trades {"ids": [51 … 100]} ───▶ 503 ─ back-off ─ retry ─▶ 200
235
+ [RS] │ POST /trades {"ids": [101 … 150]} ───▶ 429 ─ waits Retry-After ─▶ 200
236
+ [RS] │ … │
237
+ [RS] └───────────────────────────────────────────────────────────────────────┘
238
+ [RS] parse JSON / XML-RPC responses in parallel
239
+ ═══════════════════════════ RUST ──▶ PYTHON ═════════════════════════════════
240
+ [PY] result_key="data.trades" → flatten in input order → [trade, trade, …]
241
+ [PY] batches that still fail → BatchError with .results and .failed_ids
242
+ ```
243
+
244
+ ### What happens to each HTTP request (all inside Rust)
245
+
246
+ ```text
247
+ [RS] request ─▶ proxy selection (same rules as requests)
248
+ │ ├─ host in no_proxy? → connect directly
249
+ │ └─ first match of scheme://host → scheme → all://host → all
250
+ │ (proxy URL may carry user:password → Proxy-Authorization)
251
+ ├────────▶ authentication
252
+ │ ├─ auth=("user","pwd") / bearer_token → header on every request
253
+ │ └─ auth=HTTPNegotiateAuth() → Windows SSPI handshake:
254
+ │
255
+ │ GET /api ─▶ 401 WWW-Authenticate: Negotiate
256
+ │ GET /api Authorization: Negotiate <token> ─▶ 401 + challenge (NTLM only)
257
+ │ GET /api Authorization: Negotiate <reply> ─▶ 200 OK
258
+ │ └─ one kept-alive connection; later requests on it skip the handshake
259
+ │
260
+ ├────────▶ retry on network errors / 408 / 425 / 429 / 5xx (exponential back-off + jitter)
261
+ └────────▶ parse body (JSON / XML-RPC / text / bytes) → outcome handed back to [PY]
262
+ ```
263
+
264
+ ---
265
+
117
266
  ## 1. Reading SQLite fast
118
267
 
119
268
  ```python
120
269
  pf.read_sqlite(path, table=None, *, query=None, columns=None, where=None, params=None,
121
270
  partition_on=None, partitions=None, schema_overrides=None,
122
- batch_size=131072, immutable=False) -> polars.DataFrame
271
+ batch_size=131072, immutable=False, busy_timeout=5.0) -> polars.DataFrame
123
272
  ```
124
273
 
125
274
  `read_sqlite_pandas(...)` (pandas DataFrame) and `read_sqlite_arrow(...)` (pyarrow Table) take
@@ -215,6 +364,26 @@ pf.read_sqlite("market.db", "prices", schema_overrides={"isin": "str", "qty": "f
215
364
  Dates are stored as text in SQLite and come back as strings. Convert them after reading, e.g.
216
365
  `df.with_columns(pl.col("asof").str.to_date())` in polars, or `pd.to_datetime(pdf["asof"])` in pandas.
217
366
 
367
+ ### Locking and databases that are being written to
368
+
369
+ parafetch **never writes and never takes a write lock**: every connection is opened read-only, and
370
+ the file is left byte-for-byte unchanged. Locks are released as soon as the read ends, fails or is
371
+ interrupted. What a running read means for *other* processes depends on the journal mode:
372
+
373
+ | journal mode | while `read_sqlite` runs |
374
+ |---|---|
375
+ | **WAL** (`PRAGMA journal_mode=WAL`) | readers and writers never block each other (recommended for live databases) |
376
+ | rollback journal (SQLite's default) | the read holds a shared lock, so a writer can prepare changes but **waits to commit** until the read finishes |
377
+ | `immutable=True` | no locks at all. Fastest, but only for files nothing writes to (snapshots, copies). |
378
+
379
+ - `busy_timeout=5.0` (default): if a writer is committing just as the read starts, parafetch waits up
380
+ to this many seconds instead of failing with `database is locked`.
381
+ - Parallel partitions use one connection each. If another process commits *during* the read,
382
+ different partitions can see the data before and after that commit. For an exact point-in-time
383
+ snapshot of a live database, use `partitions=1` or read a copy.
384
+ - SQLite's locking is unreliable on network drives (SMB/NFS). Prefer local disk for databases that are
385
+ written to while being read.
386
+
218
387
  ### Limitations
219
388
 
220
389
  - Tables declared `WITHOUT ROWID` are read on a single thread unless you pass `partition_on`.
@@ -373,7 +542,7 @@ api.request_many([ # arbitrary mi
373
542
  ])
374
543
  ```
375
544
 
376
- Each `request_many` entry takes `method`, `path` (or a full URL), `params`, `json`, `data`,
545
+ Each `request_many` entry takes `method`, `path` or `url`, `params`, `json`, `data`,
377
546
  `headers` and `parse` (`"json"`, `"text"`, `"bytes"` or `"auto"`).
378
547
 
379
548
  ### JSON-RPC — `JsonRpcClient`
@@ -429,7 +598,10 @@ prices = xr.call_many("pricing.get", isins, multicall_size=100) # system.multic
429
598
  | `timeout`, `connect_timeout` | `60`, `10` | seconds (`None` disables the timeout) |
430
599
  | `verify` | `True` | set `False` to skip TLS certificate checks (not recommended) |
431
600
  | `ca_cert` | `None` | path to a PEM bundle with extra trusted CAs |
432
- | `proxy` | `None` | e.g. `"http://proxy.corp:8080"`. `HTTP(S)_PROXY` env vars are used automatically. |
601
+ | `proxy` | `None` | one proxy for everything, e.g. `"http://user:pwd@proxy.corp:8080"` |
602
+ | `proxies` | `None` | `requests`-style dict, e.g. `{"http": ..., "https": ..., "http://host": ...}`. See [Proxies](#proxies). |
603
+ | `no_proxy` | `None` | hosts to reach directly: `".corp.local,10.0.0.0/8"` or a list |
604
+ | `trust_env` | `True` | use `HTTP(S)_PROXY` / `NO_PROXY` env vars and the Windows proxy settings, like `requests` |
433
605
  | `user_agent` | `parafetch/<version>` | |
434
606
 
435
607
  Create a client once and reuse it. It keeps connections open between calls.
@@ -461,6 +633,7 @@ Every client takes the same `auth=` argument as the `requests` library:
461
633
  | `requests_negotiate_sspi.HttpNegotiateAuth()` | `auth=pf.HttpNegotiateAuth()` |
462
634
  | `requests_kerberos.HTTPKerberosAuth()` | `auth=pf.HTTPNegotiateAuth(package="Kerberos")` |
463
635
  | `requests_ntlm.HttpNtlmAuth(user, pwd)` | `auth=pf.HTTPNegotiateAuth(username=user, password=pwd, package="NTLM")` |
636
+ | `proxies={...}`, `trust_env=False` | the same `proxies=` / `trust_env=` arguments ([Proxies](#proxies)) |
464
637
 
465
638
  ### Windows integrated auth (HTTP Negotiate: Kerberos / NTLM)
466
639
 
@@ -508,6 +681,44 @@ Getting `HTTP 401` with your own (logged-in) account? See the [FAQ](#faq--troubl
508
681
 
509
682
  ---
510
683
 
684
+ ## Proxies
685
+
686
+ Proxy settings work exactly like in `requests`, including credentials in the URL:
687
+
688
+ ```python
689
+ # one proxy for everything (user/password in the URL, like requests)
690
+ api = pf.HttpClient(base, proxy="http://jdoe:s3cret@proxy.corp:8080")
691
+
692
+ # requests-style dict
693
+ api = pf.HttpClient(base, proxies={
694
+ "http": "http://jdoe:s3cret@proxy.corp:8080",
695
+ "https": "http://jdoe:s3cret@proxy.corp:8080",
696
+ "https://internal.corp.local": "", # empty / None = connect directly
697
+ "all://legacy.example.com": "http://other-proxy:3128",
698
+ })
699
+
700
+ # hosts that skip the proxy
701
+ api = pf.HttpClient(base, proxy="http://proxy.corp:8080", no_proxy=".corp.local,10.0.0.0/8,localhost")
702
+ ```
703
+
704
+ - **Lookup order per request** is the same as `requests`: `scheme://host`, then `scheme`, then
705
+ `all://host`, then `all`. The first key present wins.
706
+ - **Credentials:** `user:password@` in the proxy URL is sent as `Proxy-Authorization: Basic`, both for
707
+ plain HTTP and for HTTPS (`CONNECT`) tunnels. Percent-encode special characters in the password, as
708
+ you would for `requests`: `@` → `%40`, `:` → `%3A`, `#` → `%23`, `%` → `%25`.
709
+ Example: `urllib.parse.quote(password, safe="")`.
710
+ - **`no_proxy`** accepts host names (`intranet`), domain suffixes (`.corp.local` or `*.corp.local`),
711
+ IP addresses, CIDR ranges (`10.0.0.0/8`), `host:port`, `*` (everything) and the Windows `<local>` token.
712
+ - **Environment and system settings** (`trust_env=True`, the default): `HTTP_PROXY`, `HTTPS_PROXY`,
713
+ `ALL_PROXY` and `NO_PROXY` are used, or on Windows the system proxy from *Internet Options*,
714
+ including its bypass list. Explicit `proxies` take precedence key by key. `trust_env=False` ignores
715
+ all of them.
716
+ - All three clients (`HttpClient`, `JsonRpcClient`, `XmlRpcClient`) take these options.
717
+
718
+ Not supported yet: proxies that require Windows sign-in (Negotiate/NTLM, `407 Proxy-Authenticate: Negotiate`).
719
+
720
+ ---
721
+
511
722
  ## Error handling
512
723
 
513
724
  parafetch never throws away work that succeeded.
@@ -569,7 +780,7 @@ Rule of thumb: **waiting on I/O → `thread`** (the default), **async code → `
569
780
 
570
781
  | function | returns |
571
782
  |---|---|
572
- | `read_sqlite(path, table=None, *, query=None, columns=None, where=None, params=None, partition_on=None, partitions=None, schema_overrides=None, batch_size=131072, immutable=False)` | `polars.DataFrame` |
783
+ | `read_sqlite(path, table=None, *, query=None, columns=None, where=None, params=None, partition_on=None, partitions=None, schema_overrides=None, batch_size=131072, immutable=False, busy_timeout=5.0)` | `polars.DataFrame` |
573
784
  | `read_sqlite_pandas(path, table=None, **kwargs)` | `pandas.DataFrame` (needs pandas) |
574
785
  | `read_sqlite_arrow(path, table=None, **kwargs)` | `pyarrow.Table` |
575
786
 
@@ -661,8 +872,9 @@ Lower `concurrency` or set `rate_limit`. 429 responses are retried automatically
661
872
  That column contains both binary and numeric data. Use `schema_overrides={"col": "bytes"}` (or `"str"`).
662
873
 
663
874
  **Can I read the database while another process writes to it?**
664
- Yes, with the default settings. Readers use normal SQLite locking (WAL mode is ideal). Don't use
665
- `immutable=True` in that case.
875
+ Yes. parafetch only reads and never takes a write lock. In WAL mode nothing blocks; in the default
876
+ rollback mode a writer waits to commit until the read finishes. Don't use `immutable=True` in that
877
+ case. See [Locking](#locking-and-databases-that-are-being-written-to).
666
878
 
667
879
  ---
668
880
 
@@ -19,10 +19,12 @@ Rust with Python's GIL released. Your Python code only sees ordinary lists, dict
19
19
 
20
20
  - [Installation](#installation)
21
21
  - [Quick start](#quick-start)
22
+ - [How it works (diagrams)](#how-it-works)
22
23
  - [1. Reading SQLite fast — `read_sqlite`](#1-reading-sqlite-fast)
23
24
  - [2. Parallelising one-at-a-time functions — `parallel_map`](#2-parallelising-one-at-a-time-functions)
24
25
  - [3. Fetching from size-limited APIs — REST, JSON-RPC, XML-RPC](#3-fetching-from-size-limited-apis)
25
26
  - [Authentication — Basic, Bearer, Windows Negotiate (Kerberos/NTLM)](#authentication)
27
+ - [Proxies](#proxies)
26
28
  - [Error handling](#error-handling)
27
29
  - [How parafetch deals with the GIL](#how-parafetch-deals-with-the-gil)
28
30
  - [API reference](#api-reference)
@@ -92,12 +94,159 @@ trades = api.fetch_batched(trade_ids, batch_size=50, path="/trades", id_param="i
92
94
 
93
95
  ---
94
96
 
97
+ ## How it works
98
+
99
+ Every feature is split into a thin **Python** layer and a **Rust** engine. Python handles your
100
+ arguments and the final objects you get back. Rust does the heavy lifting in between, with the
101
+ GIL **released**, so it runs on many CPU cores at once.
102
+
103
+ In the diagrams below:
104
+
105
+ ```text
106
+ [PY] runs in Python (holds the GIL: one thread at a time)
107
+ [RS] runs in Rust (GIL released: truly parallel, on all cores)
108
+ ═══ the moment control crosses between Python and Rust
109
+ ```
110
+
111
+ ### The big picture
112
+
113
+ ```text
114
+ your Python code
115
+ │
116
+ ┌────────────────────────┼─────────────────────────────┐
117
+ ▼ ▼ ▼
118
+ [PY] read_sqlite() [PY] parallel_map() [PY] HttpClient / JsonRpcClient
119
+ read_sqlite_pandas() parallel_dict() XmlRpcClient
120
+ read_sqlite_arrow() batched_map() fetch_batched() …
121
+ check arguments split into items split IDs into batches,
122
+ build request bodies
123
+ │ │ │
124
+ ════════╪════════ PYTHON ──▶ RUST (GIL released) ═════════════╪════════════════
125
+ ▼ ▼ ▼
126
+ [RS] SQLite reader [RS] thread scheduler [RS] tokio HTTP engine
127
+ • N threads, one • N OS threads • connection pool, TLS
128
+ connection each • retries + back-off • concurrency + rate limit
129
+ • decode rows • rate limit • retries, Retry-After
130
+ • build Arrow • collect results • proxies, Basic/Negotiate auth
131
+ columns in input order • parse JSON / XML-RPC
132
+ │ │ ▲ │
133
+ │ ▼ │ (per item) │
134
+ │ [PY] YOUR function runs │
135
+ │ (the only Python step; │
136
+ │ GIL held just for this) │
137
+ ════════╪════════ RUST ──▶ PYTHON ════════════════════════════╪════════════════
138
+ ▼ ▼ ▼
139
+ [PY] polars / pandas / [PY] list of results [PY] flatten batches into one
140
+ pyarrow (zero copy) or ParallelError list / dict, or BatchError
141
+ ```
142
+
143
+ ### Where Rust takes over
144
+
145
+ | step | `read_sqlite` | `parallel_map` (thread mode) | `fetch_batched` / `call_batched` |
146
+ |---|---|---|---|
147
+ | prepare | [PY] validate arguments | [PY] collect items | [PY] split IDs, build request bodies |
148
+ | plan | [RS] find rowid range, split it | [RS] start N worker threads | [RS] queue requests |
149
+ | heavy work | [RS] N threads read SQLite → Arrow | [PY] **your function** (GIL held only while it runs Python code) | [RS] send, retry, authenticate, proxy |
150
+ | waiting | none (CPU-bound) | [RS] back-off sleeps, rate limiting | [RS] network waits, back-off, rate limiting |
151
+ | decode | [RS] typed Arrow buffers | none | [RS] parse JSON / XML-RPC |
152
+ | hand back | [PY] polars/pandas/pyarrow (zero copy) | [PY] results in input order | [PY] build dicts/lists, flatten |
153
+
154
+ `parallel_map` is the only feature where Python code (yours) runs in the middle. That's why it
155
+ speeds up functions that wait on I/O (the GIL is free while they wait) but not pure-Python number
156
+ crunching. For that, use `mode="process"`.
157
+
158
+ ### `read_sqlite`: parallel range scan into Arrow
159
+
160
+ ```text
161
+ [PY] read_sqlite("market.db", "prices") # partitions = CPU count
162
+ ═══════════════════════════ PYTHON ──▶ RUST ═════════════════════════════════
163
+ [RS] SELECT MIN(rowid), MAX(rowid) ─▶ 1 … 2,000,000 → split into ranges
164
+ │
165
+ ┌────────────────┬───────────────────────┼───────────────┬──────────────┐
166
+ ▼ ▼ ▼ ▼
167
+ [RS] thread 1 [RS] thread 2 [RS] thread 3 [RS] thread 4
168
+ own read-only own read-only … …
169
+ connection connection
170
+ rowid 1–500k 500k–1M 1M–1.5M 1.5M–2M
171
+ │ │ │ │
172
+ ▼ ▼ ▼ ▼
173
+ [RS] values written straight into typed Arrow buffers (int64, float64, string …)
174
+ no Python object is ever created per row or per cell
175
+ └────────────────┴───────────┬───────────┴───────────────┘
176
+ ▼ joined in rowid order
177
+ ═══════════════════════════ RUST ──▶ PYTHON ═════════════════════════════════
178
+ [PY] polars.DataFrame (zero copy: the Arrow memory is shared)
179
+ ```
180
+
181
+ ### `parallel_map`: why threads help I/O-bound functions
182
+
183
+ ```text
184
+ [PY] parallel_map(get_price, isins, workers=3)
185
+ ═══════════════════════════ PYTHON ──▶ RUST ═════════════════════════════════
186
+ [RS] 3 OS threads pull items from a shared queue
187
+
188
+ time ───────────────────────────────────────────────────▶
189
+ thread 1 ▓░░░░░░░ waiting for DB ░░░░░░░▓ ▓░░░░░░░ waiting ░░░░░▓
190
+ thread 2 ▓░░░░░░░ waiting for DB ░░░░░░░▓ ▓░░░░░░░ waiting ░░░░░▓
191
+ thread 3 ▓░░░░░░░ waiting for DB ░░░░░░░▓ ▓░░░░░░░ waiting ░░░░▓
192
+
193
+ ▓ = [PY] your function's Python code: holds the GIL, but only briefly
194
+ ░ = waiting on network / DB: GIL released, so the other threads run
195
+ → 3 workers ≈ 3× faster, 32 workers ≈ 32× (up to what the server allows)
196
+ [RS] between items: retry back-off sleeps and rate-limit waits, GIL released
197
+ ═══════════════════════════ RUST ──▶ PYTHON ═════════════════════════════════
198
+ [PY] results in input order (or ParallelError with the partial results)
199
+
200
+ CPU-bound Python is all ▓ and threads can't overlap it → use mode="process"
201
+ (one GIL per process) or mode="interpreter" on Python 3.14+.
202
+ ```
203
+
204
+ ### `fetch_batched` / `call_batched`: size-limited APIs
205
+
206
+ ```text
207
+ [PY] api.fetch_batched(10_000 ids, batch_size=50)
208
+ [PY] split into 200 batches → build 200 request bodies ("$ids" filled in)
209
+ ═══════════════════════════ PYTHON ──▶ RUST ═════════════════════════════════
210
+ [RS] ┌── at most `concurrency` in flight ── at most `rate_limit` per second ──┐
211
+ [RS] │ POST /trades {"ids": [1 … 50]} ───▶ 200 OK │
212
+ [RS] │ POST /trades {"ids": [51 … 100]} ───▶ 503 ─ back-off ─ retry ─▶ 200
213
+ [RS] │ POST /trades {"ids": [101 … 150]} ───▶ 429 ─ waits Retry-After ─▶ 200
214
+ [RS] │ … │
215
+ [RS] └───────────────────────────────────────────────────────────────────────┘
216
+ [RS] parse JSON / XML-RPC responses in parallel
217
+ ═══════════════════════════ RUST ──▶ PYTHON ═════════════════════════════════
218
+ [PY] result_key="data.trades" → flatten in input order → [trade, trade, …]
219
+ [PY] batches that still fail → BatchError with .results and .failed_ids
220
+ ```
221
+
222
+ ### What happens to each HTTP request (all inside Rust)
223
+
224
+ ```text
225
+ [RS] request ─▶ proxy selection (same rules as requests)
226
+ │ ├─ host in no_proxy? → connect directly
227
+ │ └─ first match of scheme://host → scheme → all://host → all
228
+ │ (proxy URL may carry user:password → Proxy-Authorization)
229
+ ├────────▶ authentication
230
+ │ ├─ auth=("user","pwd") / bearer_token → header on every request
231
+ │ └─ auth=HTTPNegotiateAuth() → Windows SSPI handshake:
232
+ │
233
+ │ GET /api ─▶ 401 WWW-Authenticate: Negotiate
234
+ │ GET /api Authorization: Negotiate <token> ─▶ 401 + challenge (NTLM only)
235
+ │ GET /api Authorization: Negotiate <reply> ─▶ 200 OK
236
+ │ └─ one kept-alive connection; later requests on it skip the handshake
237
+ │
238
+ ├────────▶ retry on network errors / 408 / 425 / 429 / 5xx (exponential back-off + jitter)
239
+ └────────▶ parse body (JSON / XML-RPC / text / bytes) → outcome handed back to [PY]
240
+ ```
241
+
242
+ ---
243
+
95
244
  ## 1. Reading SQLite fast
96
245
 
97
246
  ```python
98
247
  pf.read_sqlite(path, table=None, *, query=None, columns=None, where=None, params=None,
99
248
  partition_on=None, partitions=None, schema_overrides=None,
100
- batch_size=131072, immutable=False) -> polars.DataFrame
249
+ batch_size=131072, immutable=False, busy_timeout=5.0) -> polars.DataFrame
101
250
  ```
102
251
 
103
252
  `read_sqlite_pandas(...)` (pandas DataFrame) and `read_sqlite_arrow(...)` (pyarrow Table) take
@@ -193,6 +342,26 @@ pf.read_sqlite("market.db", "prices", schema_overrides={"isin": "str", "qty": "f
193
342
  Dates are stored as text in SQLite and come back as strings. Convert them after reading, e.g.
194
343
  `df.with_columns(pl.col("asof").str.to_date())` in polars, or `pd.to_datetime(pdf["asof"])` in pandas.
195
344
 
345
+ ### Locking and databases that are being written to
346
+
347
+ parafetch **never writes and never takes a write lock**: every connection is opened read-only, and
348
+ the file is left byte-for-byte unchanged. Locks are released as soon as the read ends, fails or is
349
+ interrupted. What a running read means for *other* processes depends on the journal mode:
350
+
351
+ | journal mode | while `read_sqlite` runs |
352
+ |---|---|
353
+ | **WAL** (`PRAGMA journal_mode=WAL`) | readers and writers never block each other (recommended for live databases) |
354
+ | rollback journal (SQLite's default) | the read holds a shared lock, so a writer can prepare changes but **waits to commit** until the read finishes |
355
+ | `immutable=True` | no locks at all. Fastest, but only for files nothing writes to (snapshots, copies). |
356
+
357
+ - `busy_timeout=5.0` (default): if a writer is committing just as the read starts, parafetch waits up
358
+ to this many seconds instead of failing with `database is locked`.
359
+ - Parallel partitions use one connection each. If another process commits *during* the read,
360
+ different partitions can see the data before and after that commit. For an exact point-in-time
361
+ snapshot of a live database, use `partitions=1` or read a copy.
362
+ - SQLite's locking is unreliable on network drives (SMB/NFS). Prefer local disk for databases that are
363
+ written to while being read.
364
+
196
365
  ### Limitations
197
366
 
198
367
  - Tables declared `WITHOUT ROWID` are read on a single thread unless you pass `partition_on`.
@@ -351,7 +520,7 @@ api.request_many([ # arbitrary mi
351
520
  ])
352
521
  ```
353
522
 
354
- Each `request_many` entry takes `method`, `path` (or a full URL), `params`, `json`, `data`,
523
+ Each `request_many` entry takes `method`, `path` or `url`, `params`, `json`, `data`,
355
524
  `headers` and `parse` (`"json"`, `"text"`, `"bytes"` or `"auto"`).
356
525
 
357
526
  ### JSON-RPC — `JsonRpcClient`
@@ -407,7 +576,10 @@ prices = xr.call_many("pricing.get", isins, multicall_size=100) # system.multic
407
576
  | `timeout`, `connect_timeout` | `60`, `10` | seconds (`None` disables the timeout) |
408
577
  | `verify` | `True` | set `False` to skip TLS certificate checks (not recommended) |
409
578
  | `ca_cert` | `None` | path to a PEM bundle with extra trusted CAs |
410
- | `proxy` | `None` | e.g. `"http://proxy.corp:8080"`. `HTTP(S)_PROXY` env vars are used automatically. |
579
+ | `proxy` | `None` | one proxy for everything, e.g. `"http://user:pwd@proxy.corp:8080"` |
580
+ | `proxies` | `None` | `requests`-style dict, e.g. `{"http": ..., "https": ..., "http://host": ...}`. See [Proxies](#proxies). |
581
+ | `no_proxy` | `None` | hosts to reach directly: `".corp.local,10.0.0.0/8"` or a list |
582
+ | `trust_env` | `True` | use `HTTP(S)_PROXY` / `NO_PROXY` env vars and the Windows proxy settings, like `requests` |
411
583
  | `user_agent` | `parafetch/<version>` | |
412
584
 
413
585
  Create a client once and reuse it. It keeps connections open between calls.
@@ -439,6 +611,7 @@ Every client takes the same `auth=` argument as the `requests` library:
439
611
  | `requests_negotiate_sspi.HttpNegotiateAuth()` | `auth=pf.HttpNegotiateAuth()` |
440
612
  | `requests_kerberos.HTTPKerberosAuth()` | `auth=pf.HTTPNegotiateAuth(package="Kerberos")` |
441
613
  | `requests_ntlm.HttpNtlmAuth(user, pwd)` | `auth=pf.HTTPNegotiateAuth(username=user, password=pwd, package="NTLM")` |
614
+ | `proxies={...}`, `trust_env=False` | the same `proxies=` / `trust_env=` arguments ([Proxies](#proxies)) |
442
615
 
443
616
  ### Windows integrated auth (HTTP Negotiate: Kerberos / NTLM)
444
617
 
@@ -486,6 +659,44 @@ Getting `HTTP 401` with your own (logged-in) account? See the [FAQ](#faq--troubl
486
659
 
487
660
  ---
488
661
 
662
+ ## Proxies
663
+
664
+ Proxy settings work exactly like in `requests`, including credentials in the URL:
665
+
666
+ ```python
667
+ # one proxy for everything (user/password in the URL, like requests)
668
+ api = pf.HttpClient(base, proxy="http://jdoe:s3cret@proxy.corp:8080")
669
+
670
+ # requests-style dict
671
+ api = pf.HttpClient(base, proxies={
672
+ "http": "http://jdoe:s3cret@proxy.corp:8080",
673
+ "https": "http://jdoe:s3cret@proxy.corp:8080",
674
+ "https://internal.corp.local": "", # empty / None = connect directly
675
+ "all://legacy.example.com": "http://other-proxy:3128",
676
+ })
677
+
678
+ # hosts that skip the proxy
679
+ api = pf.HttpClient(base, proxy="http://proxy.corp:8080", no_proxy=".corp.local,10.0.0.0/8,localhost")
680
+ ```
681
+
682
+ - **Lookup order per request** is the same as `requests`: `scheme://host`, then `scheme`, then
683
+ `all://host`, then `all`. The first key present wins.
684
+ - **Credentials:** `user:password@` in the proxy URL is sent as `Proxy-Authorization: Basic`, both for
685
+ plain HTTP and for HTTPS (`CONNECT`) tunnels. Percent-encode special characters in the password, as
686
+ you would for `requests`: `@` → `%40`, `:` → `%3A`, `#` → `%23`, `%` → `%25`.
687
+ Example: `urllib.parse.quote(password, safe="")`.
688
+ - **`no_proxy`** accepts host names (`intranet`), domain suffixes (`.corp.local` or `*.corp.local`),
689
+ IP addresses, CIDR ranges (`10.0.0.0/8`), `host:port`, `*` (everything) and the Windows `<local>` token.
690
+ - **Environment and system settings** (`trust_env=True`, the default): `HTTP_PROXY`, `HTTPS_PROXY`,
691
+ `ALL_PROXY` and `NO_PROXY` are used, or on Windows the system proxy from *Internet Options*,
692
+ including its bypass list. Explicit `proxies` take precedence key by key. `trust_env=False` ignores
693
+ all of them.
694
+ - All three clients (`HttpClient`, `JsonRpcClient`, `XmlRpcClient`) take these options.
695
+
696
+ Not supported yet: proxies that require Windows sign-in (Negotiate/NTLM, `407 Proxy-Authenticate: Negotiate`).
697
+
698
+ ---
699
+
489
700
  ## Error handling
490
701
 
491
702
  parafetch never throws away work that succeeded.
@@ -547,7 +758,7 @@ Rule of thumb: **waiting on I/O → `thread`** (the default), **async code → `
547
758
 
548
759
  | function | returns |
549
760
  |---|---|
550
- | `read_sqlite(path, table=None, *, query=None, columns=None, where=None, params=None, partition_on=None, partitions=None, schema_overrides=None, batch_size=131072, immutable=False)` | `polars.DataFrame` |
761
+ | `read_sqlite(path, table=None, *, query=None, columns=None, where=None, params=None, partition_on=None, partitions=None, schema_overrides=None, batch_size=131072, immutable=False, busy_timeout=5.0)` | `polars.DataFrame` |
551
762
  | `read_sqlite_pandas(path, table=None, **kwargs)` | `pandas.DataFrame` (needs pandas) |
552
763
  | `read_sqlite_arrow(path, table=None, **kwargs)` | `pyarrow.Table` |
553
764
 
@@ -639,8 +850,9 @@ Lower `concurrency` or set `rate_limit`. 429 responses are retried automatically
639
850
  That column contains both binary and numeric data. Use `schema_overrides={"col": "bytes"}` (or `"str"`).
640
851
 
641
852
  **Can I read the database while another process writes to it?**
642
- Yes, with the default settings. Readers use normal SQLite locking (WAL mode is ideal). Don't use
643
- `immutable=True` in that case.
853
+ Yes. parafetch only reads and never takes a write lock. In WAL mode nothing blocks; in the default
854
+ rollback mode a writer waits to commit until the read finishes. Don't use `immutable=True` in that
855
+ case. See [Locking](#locking-and-databases-that-are-being-written-to).
644
856
 
645
857
  ---
646
858
 
@@ -18,6 +18,7 @@ def read_sqlite(
18
18
  schema_overrides: dict[str, str] | None = None,
19
19
  batch_size: int = 131072,
20
20
  immutable: bool = False,
21
+ busy_timeout: float = 5.0,
21
22
  ) -> tuple[pa.Schema, list[pa.RecordBatch]]: ...
22
23
  def parallel_map(
23
24
  func: Callable[..., Any],
@@ -49,7 +50,9 @@ class HttpTransport:
49
50
  rate_limit: float | None = None,
50
51
  verify: bool = True,
51
52
  ca_cert: str | None = None,
52
- proxy: str | None = None,
53
+ proxies: list[tuple[str, str | None]] | None = None,
54
+ no_proxy: str | None = None,
55
+ trust_env: bool = True,
53
56
  basic_auth: tuple[str, str | None] | None = None,
54
57
  user_agent: str | None = None,
55
58
  negotiate: tuple[str, str, str | None, str | None, str | None, str | None, bool, bool] | None = None,