reqstorm 2.0.1__tar.gz → 2.2.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- reqstorm-2.2.0/PKG-INFO +637 -0
- reqstorm-2.2.0/README.md +596 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/pyproject.toml +7 -3
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm/__init__.py +24 -4
- reqstorm-2.2.0/reqstorm/__main__.py +5 -0
- reqstorm-2.2.0/reqstorm/_adaptive.py +118 -0
- reqstorm-2.2.0/reqstorm/_auth.py +79 -0
- reqstorm-2.2.0/reqstorm/_cache.py +141 -0
- reqstorm-2.2.0/reqstorm/_cli.py +317 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm/_client.py +207 -27
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm/_files.py +137 -3
- reqstorm-2.2.0/reqstorm/_paginate.py +144 -0
- reqstorm-2.2.0/reqstorm/_pydantic.py +103 -0
- reqstorm-2.2.0/reqstorm/_report.py +91 -0
- reqstorm-2.2.0/reqstorm/_schema.py +427 -0
- reqstorm-2.2.0/reqstorm/_schema_sinks.py +289 -0
- reqstorm-2.2.0/reqstorm/_template.py +117 -0
- reqstorm-2.2.0/reqstorm.egg-info/PKG-INFO +637 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm.egg-info/SOURCES.txt +22 -0
- reqstorm-2.2.0/reqstorm.egg-info/entry_points.txt +2 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm.egg-info/requires.txt +5 -0
- reqstorm-2.2.0/tests/test_adaptive.py +86 -0
- reqstorm-2.2.0/tests/test_auth.py +74 -0
- reqstorm-2.2.0/tests/test_cache_proxy.py +119 -0
- reqstorm-2.2.0/tests/test_cli.py +121 -0
- reqstorm-2.2.0/tests/test_paginate.py +113 -0
- reqstorm-2.2.0/tests/test_pydantic.py +115 -0
- reqstorm-2.2.0/tests/test_run_report.py +52 -0
- reqstorm-2.2.0/tests/test_schema.py +210 -0
- reqstorm-2.2.0/tests/test_schema_db.py +208 -0
- reqstorm-2.2.0/tests/test_template.py +63 -0
- reqstorm-2.0.1/PKG-INFO +0 -271
- reqstorm-2.0.1/README.md +0 -234
- reqstorm-2.0.1/reqstorm.egg-info/PKG-INFO +0 -271
- {reqstorm-2.0.1 → reqstorm-2.2.0}/LICENSE +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm/_legacy.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm/_limits.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm/_plan.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm/_progress.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm/_sync.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm/py.typed +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm.egg-info/dependency_links.txt +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/reqstorm.egg-info/top_level.txt +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/setup.cfg +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/tests/test_fetch_all.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/tests/test_files.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/tests/test_legacy.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/tests/test_limits.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/tests/test_ordered_and_db.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/tests/test_report.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/tests/test_stream.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/tests/test_sync.py +0 -0
- {reqstorm-2.0.1 → reqstorm-2.2.0}/tests/test_tls.py +0 -0
reqstorm-2.2.0/PKG-INFO
ADDED
|
@@ -0,0 +1,637 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: reqstorm
|
|
3
|
+
Version: 2.2.0
|
|
4
|
+
Summary: Send large numbers of HTTP requests concurrently with asyncio, with retries, timeouts and per-request results.
|
|
5
|
+
Author-email: Melih Colpan <colpanmelih@gmail.com>
|
|
6
|
+
License-Expression: MIT
|
|
7
|
+
Project-URL: Homepage, https://reqstorm.github.io
|
|
8
|
+
Project-URL: Source, https://github.com/melihcolpan/reqstorm
|
|
9
|
+
Project-URL: Changelog, https://github.com/melihcolpan/reqstorm/blob/master/CHANGELOG.md
|
|
10
|
+
Project-URL: Issues, https://github.com/melihcolpan/reqstorm/issues
|
|
11
|
+
Keywords: asyncio,aiohttp,http,requests,concurrent,bulk,rate-limit,retry,crawler,scraping
|
|
12
|
+
Classifier: Development Status :: 5 - Production/Stable
|
|
13
|
+
Classifier: Framework :: AsyncIO
|
|
14
|
+
Classifier: Intended Audience :: Developers
|
|
15
|
+
Classifier: Operating System :: OS Independent
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
21
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
22
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
23
|
+
Classifier: Topic :: Internet :: WWW/HTTP
|
|
24
|
+
Classifier: Typing :: Typed
|
|
25
|
+
Requires-Python: >=3.9
|
|
26
|
+
Description-Content-Type: text/markdown
|
|
27
|
+
License-File: LICENSE
|
|
28
|
+
Requires-Dist: aiohttp<4,>=3.9
|
|
29
|
+
Provides-Extra: pydantic
|
|
30
|
+
Requires-Dist: pydantic>=2; extra == "pydantic"
|
|
31
|
+
Provides-Extra: test
|
|
32
|
+
Requires-Dist: pytest>=8; extra == "test"
|
|
33
|
+
Requires-Dist: pytest-asyncio>=0.23; extra == "test"
|
|
34
|
+
Requires-Dist: trustme>=1.1; extra == "test"
|
|
35
|
+
Requires-Dist: pydantic>=2; extra == "test"
|
|
36
|
+
Provides-Extra: lint
|
|
37
|
+
Requires-Dist: ruff>=0.6; extra == "lint"
|
|
38
|
+
Requires-Dist: mypy>=1.10; extra == "lint"
|
|
39
|
+
Requires-Dist: pydantic>=2; extra == "lint"
|
|
40
|
+
Dynamic: license-file
|
|
41
|
+
|
|
42
|
+
# reqstorm
|
|
43
|
+
|
|
44
|
+
[](https://pypi.org/project/reqstorm/)
|
|
45
|
+
[](https://pypi.org/project/reqstorm/)
|
|
46
|
+
[](https://github.com/melihcolpan/reqstorm/actions/workflows/ci.yml)
|
|
47
|
+
[](https://reqstorm.github.io)
|
|
48
|
+
[](LICENSE)
|
|
49
|
+
|
|
50
|
+
**Send thousands of HTTP requests without writing the plumbing.**
|
|
51
|
+
|
|
52
|
+
reqstorm sends large numbers of HTTP requests concurrently and gives you one result per request. Rate limits, retries, progress, error reports and output to files or databases are built in.
|
|
53
|
+
|
|
54
|
+
📖 **Documentation: [reqstorm.github.io](https://reqstorm.github.io)**
|
|
55
|
+
|
|
56
|
+
> reqstorm is the new name of **reqt**. Coming from reqt? See [Coming from reqt](#coming-from-reqt).
|
|
57
|
+
|
|
58
|
+
```python
|
|
59
|
+
import reqstorm
|
|
60
|
+
|
|
61
|
+
urls = [
|
|
62
|
+
f"https://api.example.com/items/{i}"
|
|
63
|
+
for i in range(7000)
|
|
64
|
+
]
|
|
65
|
+
|
|
66
|
+
print(reqstorm.estimate(7000, rate_limit="100/min"))
|
|
67
|
+
# 7000 requests: about 1h 10m
|
|
68
|
+
# (limited by rate_limit)
|
|
69
|
+
|
|
70
|
+
results = reqstorm.fetch_all_sync(
|
|
71
|
+
urls,
|
|
72
|
+
rate_limit="100/min",
|
|
73
|
+
retries=2,
|
|
74
|
+
progress=True,
|
|
75
|
+
)
|
|
76
|
+
|
|
77
|
+
print(results.summary())
|
|
78
|
+
# {'total': 7000, 'ok': 6987, 'failed': 13,
|
|
79
|
+
# 'failures': {'HTTP 404': 9, 'TimeoutError': 4},
|
|
80
|
+
# 'latency': {'p50': 0.21, 'p95': 0.73, ...}, ...}
|
|
81
|
+
|
|
82
|
+
for error in results.errors():
|
|
83
|
+
print(error["url"], error["error"] or error["status"])
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
## Contents
|
|
87
|
+
|
|
88
|
+
- [Why reqstorm](#why-reqstorm)
|
|
89
|
+
- [Installation](#installation)
|
|
90
|
+
- [Quick start](#quick-start)
|
|
91
|
+
- [Results and reports](#results-and-reports)
|
|
92
|
+
- [Rate limits and concurrency](#rate-limits-and-concurrency)
|
|
93
|
+
- [Timeouts and retries](#timeouts-and-retries)
|
|
94
|
+
- [Writing results to a file](#writing-results-to-a-file)
|
|
95
|
+
- [Writing results to a database](#writing-results-to-a-database)
|
|
96
|
+
- [Structured data: JSON to typed columns](#structured-data-json-to-typed-columns)
|
|
97
|
+
- [Pagination](#pagination)
|
|
98
|
+
- [Requests from a CSV file or a table](#requests-from-a-csv-file-or-a-table)
|
|
99
|
+
- [Tokens, proxies and caching](#tokens-proxies-and-caching)
|
|
100
|
+
- [Requests, headers and bodies](#requests-headers-and-bodies)
|
|
101
|
+
- [Streaming results](#streaming-results)
|
|
102
|
+
- [TLS and sessions](#tls-and-sessions)
|
|
103
|
+
- [Command line](#command-line)
|
|
104
|
+
- [When to use something else](#when-to-use-something-else)
|
|
105
|
+
- [Coming from reqt](#coming-from-reqt)
|
|
106
|
+
- [Development](#development)
|
|
107
|
+
- [License](#license)
|
|
108
|
+
|
|
109
|
+
## Why reqstorm
|
|
110
|
+
|
|
111
|
+
Sending one request is easy. Sending 7000 to an API that allows 100 a minute, without losing results when a few of them fail, is where the work is:
|
|
112
|
+
|
|
113
|
+
- **Pacing:** stay under the API's rate limit, per host, including retries.
|
|
114
|
+
- **Isolating failures:** one timeout or dropped connection must not stop the other 6999 requests.
|
|
115
|
+
- **Retrying the right failures:** retry 503s and timeouts, never 404s, and back off when the server says so.
|
|
116
|
+
- **Keeping the results:** write them somewhere as they arrive, and resume where you stopped if the run is interrupted.
|
|
117
|
+
- **Knowing what happened:** which requests failed, why, and after how many attempts.
|
|
118
|
+
|
|
119
|
+
reqstorm does all of that for you, with one call.
|
|
120
|
+
|
|
121
|
+
| Feature | What you get |
|
|
122
|
+
|---|---|
|
|
123
|
+
| One result per request | Failures are recorded, never raised; every attempt is kept |
|
|
124
|
+
| Rate limits | Per host, in any unit (`"100/min"`), or `"auto"` from the server's 429s and headers |
|
|
125
|
+
| Concurrency | Overall and per host |
|
|
126
|
+
| Retries | Immediate, with backoff and `Retry-After`; plus end-of-run rounds |
|
|
127
|
+
| Reports | Failures by reason, p50/p95/p99 response times, per-host figures |
|
|
128
|
+
| Pagination | Next links, `Link` headers, cursors and page numbers |
|
|
129
|
+
| Requests from data | URL templates over CSV rows or SQL query results |
|
|
130
|
+
| Tokens and proxies | Refresh an expired token on 401; rotate through proxies |
|
|
131
|
+
| Caching | ETag / `If-None-Match`: unchanged resources cost a 304 |
|
|
132
|
+
| Output | JSONL, CSV, SQLite, PostgreSQL, MySQL; ordered or as completed |
|
|
133
|
+
| Typed columns | JSON fields to checked columns, nested paths, arrays to rows |
|
|
134
|
+
| Resume | Skip what already succeeded after an interruption |
|
|
135
|
+
| Planning | `estimate()` before you start, progress with ETA while running |
|
|
136
|
+
| API | Blocking (scripts, Jupyter), asyncio and a `reqstorm` command |
|
|
137
|
+
| Safety | TLS verified, timeouts on, bounded concurrency, fully typed |
|
|
138
|
+
|
|
139
|
+
## Installation
|
|
140
|
+
|
|
141
|
+
```console
|
|
142
|
+
$ python -m pip install reqstorm
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
reqstorm supports Python 3.9 to 3.13 on Linux, macOS and Windows. Its only dependency is [aiohttp](https://docs.aiohttp.org). For PostgreSQL or MySQL output, install the driver you already use (`psycopg`, `psycopg2`, `pymysql`, `mysqlclient` or `mysql-connector-python`). To use Pydantic models as schemas, install `reqstorm[pydantic]`.
|
|
146
|
+
|
|
147
|
+
## Quick start
|
|
148
|
+
|
|
149
|
+
Every function has a blocking version ending in `_sync` for scripts and notebooks, and an async version for code that already runs an event loop. Both take the same options.
|
|
150
|
+
|
|
151
|
+
**In a script or a Jupyter notebook:**
|
|
152
|
+
|
|
153
|
+
```python
|
|
154
|
+
import reqstorm
|
|
155
|
+
|
|
156
|
+
urls = [
|
|
157
|
+
f"https://httpbin.org/get?page={page}"
|
|
158
|
+
for page in range(1, 101)
|
|
159
|
+
]
|
|
160
|
+
|
|
161
|
+
results = reqstorm.fetch_all_sync(
|
|
162
|
+
urls,
|
|
163
|
+
concurrency=20,
|
|
164
|
+
timeout=10,
|
|
165
|
+
retries=2,
|
|
166
|
+
)
|
|
167
|
+
|
|
168
|
+
for result in results:
|
|
169
|
+
if result.ok:
|
|
170
|
+
print(result.url, result.status)
|
|
171
|
+
else:
|
|
172
|
+
print(result.url, "failed:",
|
|
173
|
+
result.error or result.status)
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
**In async code:**
|
|
177
|
+
|
|
178
|
+
```python
|
|
179
|
+
import asyncio
|
|
180
|
+
import reqstorm
|
|
181
|
+
|
|
182
|
+
async def main():
|
|
183
|
+
results = await reqstorm.fetch_all(
|
|
184
|
+
urls, concurrency=20, timeout=10
|
|
185
|
+
)
|
|
186
|
+
print(results.summary())
|
|
187
|
+
|
|
188
|
+
asyncio.run(main())
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
| Blocking | async | What it does |
|
|
192
|
+
|---|---|---|
|
|
193
|
+
| `fetch_all_sync` | `fetch_all` | Returns a `Results` list, one `Result` per request |
|
|
194
|
+
| `stream_sync` | `stream` | Yields each result as soon as it completes |
|
|
195
|
+
| `fetch_to_file_sync` | `fetch_to_file` | Writes results to JSONL, CSV or SQLite |
|
|
196
|
+
| `fetch_to_db_sync` | `fetch_to_db` | Inserts results into SQLite, PostgreSQL or MySQL |
|
|
197
|
+
| `estimate` | | Predicts how long a batch takes |
|
|
198
|
+
|
|
199
|
+
## Results and reports
|
|
200
|
+
|
|
201
|
+
`fetch_all` returns a `Results` list with one `Result` per request, in the order you gave them. A failed request never raises an exception; it is recorded on its result.
|
|
202
|
+
|
|
203
|
+
| `Result` attribute | Meaning |
|
|
204
|
+
|---|---|
|
|
205
|
+
| `ok` | `True` when a response arrived with a status below 400 |
|
|
206
|
+
| `status`, `headers`, `body` | The response; `status` is `None` when none arrived |
|
|
207
|
+
| `text()`, `json()` | The body as text (using the response charset) or as JSON |
|
|
208
|
+
| `error` | The exception that prevented a response, or `None` |
|
|
209
|
+
| `attempts`, `elapsed` | Number of attempts and total seconds |
|
|
210
|
+
| `history` | Every attempt: number, status, error and duration |
|
|
211
|
+
| `url`, `method` | What was requested |
|
|
212
|
+
| `final_url` | Where redirects ended |
|
|
213
|
+
| `index` | Position in your input |
|
|
214
|
+
|
|
215
|
+
```python
|
|
216
|
+
results.succeeded # successful results
|
|
217
|
+
results.failed # failed results
|
|
218
|
+
results.summary() # counts, failures grouped by reason
|
|
219
|
+
results.errors() # failed requests as plain dicts
|
|
220
|
+
results.to_dicts() # every result as a dict
|
|
221
|
+
results.report() # response times, statuses, hosts
|
|
222
|
+
```
|
|
223
|
+
|
|
224
|
+
```python
|
|
225
|
+
>>> results.report()["latency"]
|
|
226
|
+
{'min': 0.081, 'p50': 0.214, 'p90': 0.502,
|
|
227
|
+
'p95': 0.733, 'p99': 1.902, 'max': 10.004,
|
|
228
|
+
'mean': 0.297}
|
|
229
|
+
>>> results.report()["hosts"]["api.example.com:443"]
|
|
230
|
+
{'requests': 7000, 'ok': 6987, 'failed': 13,
|
|
231
|
+
'latency': {...}}
|
|
232
|
+
```
|
|
233
|
+
|
|
234
|
+
`errors()` returns plain data, ready for a log file or a DataFrame:
|
|
235
|
+
|
|
236
|
+
```json
|
|
237
|
+
{
|
|
238
|
+
"index": 12,
|
|
239
|
+
"method": "GET",
|
|
240
|
+
"url": "https://api.example.com/items/12",
|
|
241
|
+
"status": null,
|
|
242
|
+
"ok": false,
|
|
243
|
+
"error": "TimeoutError",
|
|
244
|
+
"attempts": 3,
|
|
245
|
+
"elapsed": 20.415,
|
|
246
|
+
"history": [
|
|
247
|
+
{"attempt": 1, "status": null,
|
|
248
|
+
"error": "TimeoutError", "elapsed": 10.001},
|
|
249
|
+
{"attempt": 2, "status": 503,
|
|
250
|
+
"error": null, "elapsed": 0.412},
|
|
251
|
+
{"attempt": 3, "status": null,
|
|
252
|
+
"error": "TimeoutError", "elapsed": 10.002}
|
|
253
|
+
]
|
|
254
|
+
}
|
|
255
|
+
```
|
|
256
|
+
|
|
257
|
+
Prefer exceptions? `result.raise_for_error()` raises the request's error, or `reqstorm.HTTPStatusError` for a status of 400 or above.
|
|
258
|
+
|
|
259
|
+
## Rate limits and concurrency
|
|
260
|
+
|
|
261
|
+
```python
|
|
262
|
+
results = reqstorm.fetch_all_sync(
|
|
263
|
+
urls,
|
|
264
|
+
rate_limit="100/min", # per host
|
|
265
|
+
concurrency=50, # in flight, in total
|
|
266
|
+
concurrency_per_host=10, # in flight, per host
|
|
267
|
+
)
|
|
268
|
+
```
|
|
269
|
+
|
|
270
|
+
| `rate_limit` value | Means |
|
|
271
|
+
|---|---|
|
|
272
|
+
| `5` or `0.5` | Requests per second |
|
|
273
|
+
| `"10/s"` | 10 per second |
|
|
274
|
+
| `"100/min"` | 100 per minute |
|
|
275
|
+
| `"30/5min"` | 30 every 5 minutes |
|
|
276
|
+
| `"1000/h"` | 1000 per hour |
|
|
277
|
+
| `"2/day"` | 2 per day |
|
|
278
|
+
| `(100, 60)` | 100 every 60 seconds |
|
|
279
|
+
|
|
280
|
+
The limit applies to each host (`host:port`) separately and counts retries too, so requests to different APIs never slow each other down. Requests to one host are spaced evenly.
|
|
281
|
+
|
|
282
|
+
**Don't know the limit?** `rate_limit="auto"` learns it from the server. A `429 Too Many Requests` pauses the host for `Retry-After` and slows it down; `X-RateLimit-Remaining` and `X-RateLimit-Reset` spread the remaining requests over the window; the rate recovers when the server stops pushing back. 429 responses are retried without using up `retries`.
|
|
283
|
+
|
|
284
|
+
**Plan before you send.** `estimate` predicts the duration and names the bottleneck, without sending anything:
|
|
285
|
+
|
|
286
|
+
```python
|
|
287
|
+
>>> print(reqstorm.estimate(7000, rate_limit="100/min"))
|
|
288
|
+
7000 requests: about 1h 10m (limited by rate_limit)
|
|
289
|
+
>>> print(reqstorm.estimate(
|
|
290
|
+
... 7000, rate_limit="100/min", hosts=7))
|
|
291
|
+
7000 requests: about 10m 00s (limited by rate_limit)
|
|
292
|
+
```
|
|
293
|
+
|
|
294
|
+
**Watch it run.** `progress=True` prints progress to stderr:
|
|
295
|
+
|
|
296
|
+
```text
|
|
297
|
+
reqstorm: 3500/7000 (50%) ok 3493 failed 7
|
|
298
|
+
1.7 req/s ETA 35m 00s
|
|
299
|
+
```
|
|
300
|
+
|
|
301
|
+
## Timeouts and retries
|
|
302
|
+
|
|
303
|
+
```python
|
|
304
|
+
results = reqstorm.fetch_all_sync(
|
|
305
|
+
urls,
|
|
306
|
+
timeout=10, # seconds per attempt
|
|
307
|
+
retries=3, # retry right away...
|
|
308
|
+
backoff=0.5, # ...0.5 s, 1 s, 2 s apart
|
|
309
|
+
retry_rounds=2, # then resend what still
|
|
310
|
+
retry_round_delay=30, # failed, 30 s later
|
|
311
|
+
)
|
|
312
|
+
```
|
|
313
|
+
|
|
314
|
+
- **`timeout`** is the time allowed for one attempt, including reading the body (default 30 s; `None` disables it). A slow server only fails its own requests.
|
|
315
|
+
- **`retries`** retries a request right away. The delay starts at `backoff` and doubles; a `Retry-After` header (up to 60 s) takes precedence.
|
|
316
|
+
- **`retry_rounds`** holds back requests that still failed for a retryable reason and sends them again after the rest of the batch, for temporary outages. Each request still produces exactly one result, with all its attempts in `history`.
|
|
317
|
+
|
|
318
|
+
| Retried | Not retried |
|
|
319
|
+
|---|---|
|
|
320
|
+
| Connection errors and dropped connections | Invalid URLs |
|
|
321
|
+
| Timeouts | 4xx such as 400, 401, 403, 404 |
|
|
322
|
+
| `retry_statuses` (default 429, 500, 502, 503, 504) | Other statuses |
|
|
323
|
+
|
|
324
|
+
## Writing results to a file
|
|
325
|
+
|
|
326
|
+
For large batches, write results as they arrive instead of keeping them in memory:
|
|
327
|
+
|
|
328
|
+
```python
|
|
329
|
+
summary = reqstorm.fetch_to_file_sync(
|
|
330
|
+
urls,
|
|
331
|
+
"results.jsonl",
|
|
332
|
+
rate_limit="10/s",
|
|
333
|
+
progress=True,
|
|
334
|
+
resume=True,
|
|
335
|
+
)
|
|
336
|
+
print(summary.ok, summary.failed, summary.skipped)
|
|
337
|
+
print(summary.errors[:5]) # failed records
|
|
338
|
+
```
|
|
339
|
+
|
|
340
|
+
- **Format** follows the file name: `.jsonl`, `.csv`, or `.db` / `.sqlite` / `.sqlite3` for SQLite. Pass `format=` to override it.
|
|
341
|
+
- **Order:** records are written as requests complete; `ordered=True` writes them in input order instead (finished records wait in memory for the ones before them).
|
|
342
|
+
- **Resume:** with `resume=True`, requests already recorded as successful are skipped and the rest are sent and appended. A line cut short by an interruption is ignored. Without `resume`, a JSONL or CSV file is overwritten.
|
|
343
|
+
- **Fields:** `index`, `method`, `url`, `status`, `ok`, `error`, `attempts`, `elapsed`, `final_url`, `history` and `body`.
|
|
344
|
+
- **Body:** `body="text"` (default), `"base64"` for binary responses, or `"none"`. `include_headers=True` adds the response headers.
|
|
345
|
+
|
|
346
|
+
## Writing results to a database
|
|
347
|
+
|
|
348
|
+
`fetch_to_db` inserts one row per request through a connection you already have:
|
|
349
|
+
|
|
350
|
+
```python
|
|
351
|
+
import psycopg
|
|
352
|
+
import reqstorm
|
|
353
|
+
|
|
354
|
+
with psycopg.connect("dbname=crawl") as connection:
|
|
355
|
+
summary = reqstorm.fetch_to_db_sync(
|
|
356
|
+
urls,
|
|
357
|
+
connection,
|
|
358
|
+
table="api_results",
|
|
359
|
+
resume=True,
|
|
360
|
+
)
|
|
361
|
+
```
|
|
362
|
+
|
|
363
|
+
Supported drivers: `sqlite3`, `psycopg` and `psycopg2` (PostgreSQL), `pymysql`, `MySQLdb` and `mysql.connector` (MySQL). For SQLite a file name is enough: `fetch_to_file_sync(urls, "results.db")`.
|
|
364
|
+
|
|
365
|
+
reqstorm creates the table if it does not exist. The fields you filter on are real columns; the parts whose shape varies are JSON:
|
|
366
|
+
|
|
367
|
+
| Column | PostgreSQL | MySQL | SQLite |
|
|
368
|
+
|---|---|---|---|
|
|
369
|
+
| `id` | `BIGSERIAL` | `BIGINT AUTO_INCREMENT` | `INTEGER` |
|
|
370
|
+
| `request_index`, `status`, `attempts` | `INTEGER` | `INT` | `INTEGER` |
|
|
371
|
+
| `method`, `url`, `error`, `final_url` | `TEXT` | `VARCHAR` / `LONGTEXT` | `TEXT` |
|
|
372
|
+
| `ok` | `BOOLEAN` | `BOOLEAN` | `INTEGER` |
|
|
373
|
+
| `elapsed` | `DOUBLE PRECISION` | `DOUBLE` | `REAL` |
|
|
374
|
+
| `history`, `headers` | `JSONB` | `JSON` | `TEXT` |
|
|
375
|
+
| `body` | `TEXT` / `BYTEA` | `LONGTEXT` / `LONGBLOB` | `TEXT` / `BLOB` |
|
|
376
|
+
| `created_at` | `TIMESTAMPTZ` | `TIMESTAMP` | `TEXT` |
|
|
377
|
+
|
|
378
|
+
So the questions you ask afterwards are plain SQL:
|
|
379
|
+
|
|
380
|
+
```sql
|
|
381
|
+
SELECT url, error FROM api_results WHERE NOT ok;
|
|
382
|
+
SELECT status, count(*) FROM api_results
|
|
383
|
+
GROUP BY status;
|
|
384
|
+
```
|
|
385
|
+
|
|
386
|
+
Rows are inserted in batches on a background thread, so a remote database does not slow the requests down. Existing rows are never deleted. `body="bytes"` stores the raw body in a binary column.
|
|
387
|
+
|
|
388
|
+
## Structured data: JSON to typed columns
|
|
389
|
+
|
|
390
|
+
Instead of storing raw responses, give a schema and each JSON field goes to its own typed column. Every value is checked first, so a number column never gets a string or a `NaN`.
|
|
391
|
+
|
|
392
|
+
```python
|
|
393
|
+
from reqstorm import Field
|
|
394
|
+
|
|
395
|
+
schema = {
|
|
396
|
+
"id": Field("id", int, required=True, key=True),
|
|
397
|
+
"name": Field("name", str, required=True),
|
|
398
|
+
"price": Field("pricing.amount", float),
|
|
399
|
+
"tags": Field("tags", "json"),
|
|
400
|
+
"updated": Field("updated_at", "datetime"),
|
|
401
|
+
"page": Field("$.meta.page", int),
|
|
402
|
+
}
|
|
403
|
+
|
|
404
|
+
summary = reqstorm.fetch_to_db_sync(
|
|
405
|
+
urls,
|
|
406
|
+
connection,
|
|
407
|
+
table="products",
|
|
408
|
+
schema=schema,
|
|
409
|
+
explode="items", # one row per array element
|
|
410
|
+
rejects_table="products_rejects",
|
|
411
|
+
)
|
|
412
|
+
print(summary.rows, summary.rejected)
|
|
413
|
+
```
|
|
414
|
+
|
|
415
|
+
- **Types:** `int`, `float`, `str`, `bool`, `"datetime"` and `"json"` become `BIGINT`, `DOUBLE PRECISION`, `TEXT`, `BOOLEAN`, `TIMESTAMPTZ` and `JSONB` in PostgreSQL, and their equivalents in MySQL and SQLite.
|
|
416
|
+
- **Strict checks:** `10.5` is rejected for an `int`, `"42"` for a number, and NaN or Infinity always. With `coerce=True`, compatible values such as `"12.5"` are converted.
|
|
417
|
+
- **Nested JSON:** dotted paths read nested values (`"pricing.amount"`, `"tags.0"`). Type `"json"` keeps a part whole. `explode` makes each array element a row, and `"$."` paths read from the response root.
|
|
418
|
+
- **Rejected records** are never written, not even partly. They are listed in `summary.errors` with their reasons and, with `rejects_table`, stored with the original record.
|
|
419
|
+
- **No duplicates:** `key=True` fields form the primary key. Running the batch again updates existing rows.
|
|
420
|
+
- **Resume works:** each row records the request it came from (`source_url`).
|
|
421
|
+
|
|
422
|
+
The same schema works for `.db`, `.jsonl` and `.csv` files, and `reqstorm.extract(results, schema)` returns the rows as Python lists. `reqstorm.infer_schema(samples)` drafts a schema from a few responses for you to review.
|
|
423
|
+
|
|
424
|
+
**Already have a Pydantic model?** Pass it as the schema; its fields become columns and Pydantic validates each record:
|
|
425
|
+
|
|
426
|
+
```python
|
|
427
|
+
class Product(BaseModel):
|
|
428
|
+
id: int = Field(
|
|
429
|
+
json_schema_extra={"key": True})
|
|
430
|
+
name: str
|
|
431
|
+
price: Optional[float] = Field(
|
|
432
|
+
None,
|
|
433
|
+
json_schema_extra={"path": "pricing.amount"})
|
|
434
|
+
|
|
435
|
+
reqstorm.fetch_to_file_sync(
|
|
436
|
+
urls, "shop.db", schema=Product, explode="items"
|
|
437
|
+
)
|
|
438
|
+
```
|
|
439
|
+
|
|
440
|
+
More in the [structured data guide](https://reqstorm.github.io/guide/structured-data/).
|
|
441
|
+
|
|
442
|
+
## Pagination
|
|
443
|
+
|
|
444
|
+
`paginate=` follows each starting URL through all of its pages, concurrently with the other URLs and under the same rate limits:
|
|
445
|
+
|
|
446
|
+
```python
|
|
447
|
+
results = reqstorm.fetch_all_sync(
|
|
448
|
+
["https://api.example.com/products"],
|
|
449
|
+
paginate=reqstorm.NextLink("links.next"),
|
|
450
|
+
)
|
|
451
|
+
```
|
|
452
|
+
|
|
453
|
+
| Strategy | Next page comes from |
|
|
454
|
+
|---|---|
|
|
455
|
+
| `NextLink("links.next")` | a URL in the JSON body |
|
|
456
|
+
| `LinkHeader()` | the `Link` header (GitHub style) |
|
|
457
|
+
| `Cursor("meta.next", param="cursor")` | a cursor in the body |
|
|
458
|
+
| `PageNumber("page", items="data")` | `?page=2, 3, ...` until empty |
|
|
459
|
+
|
|
460
|
+
Each result has `page` and `seed_index` (its starting URL). Combined with a schema and `explode`, every page of a catalogue becomes typed rows in one call. `max_pages` (1000 by default) stops an API that never ends.
|
|
461
|
+
|
|
462
|
+
## Requests from a CSV file or a table
|
|
463
|
+
|
|
464
|
+
`from_template` makes one request per row. Values are percent-encoded, and rows are read lazily:
|
|
465
|
+
|
|
466
|
+
```python
|
|
467
|
+
rows = reqstorm.read_csv("users.csv")
|
|
468
|
+
requests = reqstorm.from_template(
|
|
469
|
+
"https://api.example.com/users/{id}",
|
|
470
|
+
rows,
|
|
471
|
+
params={"country": "{country}"},
|
|
472
|
+
)
|
|
473
|
+
reqstorm.fetch_to_file_sync(requests, "users.jsonl")
|
|
474
|
+
```
|
|
475
|
+
|
|
476
|
+
`read_sql(connection, query)` reads rows from any database connection instead, and `json=` builds a request body per row.
|
|
477
|
+
|
|
478
|
+
## Tokens, proxies and caching
|
|
479
|
+
|
|
480
|
+
**Tokens that expire.** `BearerAuth` gets a new token when a response is 401 and sends the request again. Concurrent 401s share one refresh:
|
|
481
|
+
|
|
482
|
+
```python
|
|
483
|
+
auth = reqstorm.BearerAuth(refresh=get_token)
|
|
484
|
+
reqstorm.fetch_all_sync(urls, auth=auth)
|
|
485
|
+
```
|
|
486
|
+
|
|
487
|
+
**Proxies.** One proxy, a pool used in turn, or one per `Request`:
|
|
488
|
+
|
|
489
|
+
```python
|
|
490
|
+
reqstorm.fetch_all_sync(urls, proxy=[
|
|
491
|
+
"http://proxy-1.example.com:8080",
|
|
492
|
+
"http://proxy-2.example.com:8080",
|
|
493
|
+
])
|
|
494
|
+
```
|
|
495
|
+
|
|
496
|
+
**Caching.** `Cache` keeps responses in an SQLite file. The next run asks the server with `If-None-Match`; an unchanged resource comes back as a `304` with no body, and the stored response is used. With `ttl`, recent responses skip the network entirely:
|
|
497
|
+
|
|
498
|
+
```python
|
|
499
|
+
with reqstorm.Cache("responses.sqlite") as cache:
|
|
500
|
+
reqstorm.fetch_all_sync(urls, cache=cache)
|
|
501
|
+
```
|
|
502
|
+
|
|
503
|
+
## Requests, headers and bodies
|
|
504
|
+
|
|
505
|
+
Options given to `fetch_all` apply to every request; `reqstorm.Request` varies them per request:
|
|
506
|
+
|
|
507
|
+
```python
|
|
508
|
+
results = reqstorm.fetch_all_sync(
|
|
509
|
+
[
|
|
510
|
+
"https://api.example.com/items/1",
|
|
511
|
+
reqstorm.Request(
|
|
512
|
+
"https://api.example.com/items",
|
|
513
|
+
method="POST",
|
|
514
|
+
json={"name": "new"},
|
|
515
|
+
),
|
|
516
|
+
],
|
|
517
|
+
headers={"Authorization": "Bearer ..."},
|
|
518
|
+
)
|
|
519
|
+
```
|
|
520
|
+
|
|
521
|
+
- `fetch_all_sync(urls, "POST", json=...)` sends every plain URL as a POST.
|
|
522
|
+
- `params`, `json` and `data` work the same way.
|
|
523
|
+
- A `Request`'s own headers are merged on top of the shared ones; its other fields replace the shared value.
|
|
524
|
+
|
|
525
|
+
## Streaming results
|
|
526
|
+
|
|
527
|
+
`stream_sync` (or `stream` in async code) yields each result as soon as it completes. URLs can come from a generator, which is read lazily, so millions of URLs never need to fit in memory:
|
|
528
|
+
|
|
529
|
+
```python
|
|
530
|
+
def read_urls():
|
|
531
|
+
with open("urls.txt") as file:
|
|
532
|
+
for line in file:
|
|
533
|
+
yield line.strip()
|
|
534
|
+
|
|
535
|
+
for result in reqstorm.stream_sync(read_urls()):
|
|
536
|
+
save(result.url, result.status, result.body)
|
|
537
|
+
```
|
|
538
|
+
|
|
539
|
+
Leaving the loop early (`break`) cancels the requests still in flight. `fetch_all` also accepts `callback=`, called with each result as it completes.
|
|
540
|
+
|
|
541
|
+
## TLS and sessions
|
|
542
|
+
|
|
543
|
+
TLS certificates are verified by default. To trust a private certificate authority, pass an `ssl.SSLContext`:
|
|
544
|
+
|
|
545
|
+
```python
|
|
546
|
+
import ssl
|
|
547
|
+
|
|
548
|
+
context = ssl.create_default_context(
|
|
549
|
+
cafile="internal-ca.pem"
|
|
550
|
+
)
|
|
551
|
+
results = reqstorm.fetch_all_sync(urls, ssl=context)
|
|
552
|
+
```
|
|
553
|
+
|
|
554
|
+
`verify_ssl=False` turns verification off; only use it for hosts you control.
|
|
555
|
+
|
|
556
|
+
In async code, `session=` takes an existing `aiohttp.ClientSession` to share cookies, connection pools or proxy settings. reqstorm does not close it.
|
|
557
|
+
|
|
558
|
+
## Command line
|
|
559
|
+
|
|
560
|
+
The `reqstorm` command runs a batch without any Python code:
|
|
561
|
+
|
|
562
|
+
```console
|
|
563
|
+
$ reqstorm urls.txt -o results.jsonl \
|
|
564
|
+
--rate 100/min --retries 2
|
|
565
|
+
$ reqstorm urls.txt --estimate --rate 100/min
|
|
566
|
+
$ cat urls.txt | reqstorm --rate auto -q > out.jsonl
|
|
567
|
+
$ reqstorm users.csv -o users.db \
|
|
568
|
+
--template "https://api.example.com/users/{id}"
|
|
569
|
+
$ reqstorm urls.txt -o shop.db --report \
|
|
570
|
+
--paginate next:links.next \
|
|
571
|
+
--schema products.json --explode items
|
|
572
|
+
```
|
|
573
|
+
|
|
574
|
+
`reqstorm --help` lists every option, and the [command line guide](https://reqstorm.github.io/guide/cli/) explains them.
|
|
575
|
+
|
|
576
|
+
## When to use something else
|
|
577
|
+
|
|
578
|
+
reqstorm is built for batches. For other jobs, these are better fits:
|
|
579
|
+
|
|
580
|
+
| You need | Use |
|
|
581
|
+
|---|---|
|
|
582
|
+
| A few requests, simple scripts | [requests](https://requests.readthedocs.io) |
|
|
583
|
+
| A general HTTP client, sync or async, HTTP/2 | [httpx](https://www.python-httpx.org) |
|
|
584
|
+
| Full control over an asyncio HTTP client or server | [aiohttp](https://docs.aiohttp.org) |
|
|
585
|
+
| Crawling websites (following links, parsing pages) | [Scrapy](https://scrapy.org) |
|
|
586
|
+
|
|
587
|
+
reqstorm uses aiohttp underneath and adds the batch work: pacing, retries, failure isolation, reports and output.
|
|
588
|
+
|
|
589
|
+
## Coming from reqt
|
|
590
|
+
|
|
591
|
+
reqstorm 2.0.1 is the next version of reqt, renamed because another project already uses the reqt name.
|
|
592
|
+
|
|
593
|
+
```console
|
|
594
|
+
$ python -m pip install reqstorm
|
|
595
|
+
```
|
|
596
|
+
|
|
597
|
+
Then replace `import reqt` with `import reqstorm`. The `reqt` package on PyPI now only installs reqstorm and re-exports it with a deprecation warning, so existing code keeps working while you switch.
|
|
598
|
+
|
|
599
|
+
The reqt 1.x style, which called a function with each raw response, still works with a `DeprecationWarning` and will be removed in reqstorm 3.0:
|
|
600
|
+
|
|
601
|
+
```python
|
|
602
|
+
async def handle(response): # reqt 1.x style
|
|
603
|
+
print(response.status)
|
|
604
|
+
|
|
605
|
+
await reqstorm.fetch_all(urls=urls, method=handle)
|
|
606
|
+
|
|
607
|
+
# The same today:
|
|
608
|
+
for result in reqstorm.fetch_all_sync(urls):
|
|
609
|
+
print(result.status)
|
|
610
|
+
```
|
|
611
|
+
|
|
612
|
+
Two things changed even in the 1.x style:
|
|
613
|
+
|
|
614
|
+
- **TLS certificates are now verified.** reqt 1.x skipped verification. Use `verify_ssl=False` only for hosts you control.
|
|
615
|
+
- **Failures no longer stop the batch.** Every failed request is logged to the `reqstorm` logger.
|
|
616
|
+
|
|
617
|
+
## Development
|
|
618
|
+
|
|
619
|
+
```console
|
|
620
|
+
$ git clone https://github.com/melihcolpan/reqstorm
|
|
621
|
+
$ cd reqstorm
|
|
622
|
+
$ python -m pip install -e ".[test,lint]"
|
|
623
|
+
$ pytest
|
|
624
|
+
$ ruff check . && ruff format --check . && mypy reqstorm
|
|
625
|
+
```
|
|
626
|
+
|
|
627
|
+
- The tests run against a local server and a throwaway certificate authority; they need no network access.
|
|
628
|
+
- The PostgreSQL and MySQL tests run when `REQSTORM_TEST_POSTGRES` (a libpq connection string) and `REQSTORM_TEST_MYSQL` (`host:port:user:password:database`) are set.
|
|
629
|
+
- The documentation is in `docs/` and builds with `pip install -r docs/requirements.txt && mkdocs serve`.
|
|
630
|
+
|
|
631
|
+
CI runs the tests on Python 3.9 to 3.13 (Linux, plus macOS and Windows), against real PostgreSQL and MySQL, and builds the package and the documentation. A GitHub release tagged `vX.Y.Z` publishes to PyPI through trusted publishing.
|
|
632
|
+
|
|
633
|
+
Bug reports and pull requests are welcome on [GitHub](https://github.com/melihcolpan/reqstorm/issues).
|
|
634
|
+
|
|
635
|
+
## License
|
|
636
|
+
|
|
637
|
+
MIT, see [LICENSE](LICENSE).
|